Skip to content
Tech Interview Prep home

Search

482 results for “AI Engineer”

Practice MCQ

A catalog service merges duplicate products when their embeddings score above a fixed 0.85 cosine threshold, set at launch. After a new product category shipped, false merges rose from about 2 a week to 30 a day; each one costs a support ticket and a 20-minute manual undo. The service scores about 4,000 candidate pairs a day across the request, retry and nightly batch paths and auto-merges every pair above the threshold, live for all customers. Which change addresses the cause?

Role practice MCQ for AI Engineer

Practice MCQ

A churn model scored nightly at 02:00 gained offline AUC 0.72 → 0.94 after adding `open_save_cases`, a count populated by the 06:00 churn-save workflow. At scoring time the column is null for every live account, the train/test split is already time-based on scoring date, and support tickets carry their own `created_at` in a separate table. The outreach campaign cannot move past 02:00. What does the team commit to before release?

Role practice MCQ for AI Engineer

Practice MCQ

A coding assistant's eval harness gates rollout on the score from an LLM judge. After a prompt change, mean judge score rose 9% and median answer length 38%, while human thumbs-up on the same answers stayed flat; the judge is from the generator's model family. A three-week A/B of the old and new prompt starts soon, both cohorts must be scored on one comparable path, a 400-pair human-labelled set exists, and the harness can still change. Which judge policy should the team adopt before the A/B opens?

Role practice MCQ for AI Engineer

Practice MCQ

A corporate policy assistant cites handbook excerpts in its answers and serves them straight to users with no review queue. The last incident was a sentence whose wording closely matched its cited excerpt but reversed the handbook's meaning. The handbook is re-indexed weekly, and the model only ever sees the retrieved excerpts. Compliance requires that any claim the retrieved passages do not support never reaches a user. Which validation design meets that requirement?

Role practice MCQ for AI Engineer

Practice MCQ

A credit-decision model scores 200,000 applications a month; ground truth comes back from the bureau on a fixed 30-day settlement cycle, and the model retrains quarterly. Last June an upstream re-encoding shifted three features and approvals stayed wrong for a month before realised accuracy moved. The release gate asks for evidence that the June regression and regressions like it would surface before that day-30 point. Which design should the review accept?

Role practice MCQ for AI Engineer

Practice MCQ

A fraud-scoring service takes features from four producer teams that deploy independently; payloads arrive over a Kafka stream and a nightly batch backfill. One team wants `amount`, currently dollars, to carry minor units in place, and the model cannot be retrained for six weeks. Old and new payloads overlap for about a month, and a wrong value lowers recall silently instead of raising an error. Which design should the platform team adopt?

Role practice MCQ for AI Engineer

Practice MCQ

A healthcare billing copilot sends full patient support tickets to a managed LLM API. Your classification marks account numbers and national IDs as regulated, and the provider's standard tier logs prompts for 30 days and may use them for training. The task needs the provider's frontier model to hit the accuracy bar, and live traffic already carries these fields. Which control should the team ship before next week's release?

Role practice MCQ for AI Engineer

Practice MCQ

A multi-step agent sends the full transcript on every model call. A typical run makes 20 tool calls whose results average 12,000 tokens each against a 200,000-token context window, so requests fail near the end of the loop and per-conversation token spend has tripled. The team must settle truncation, summarisation, and cost control in one policy. Which should it ship?

Role practice MCQ for AI Engineer

Practice MCQ

A multi-tenant support assistant answers from a retrieval index where each user's permission set filters the evidence, so the same question returns different documents for two roles. The prompt template and model roll on a weekly release train, corpus edits land hourly, and hot questions must serve from cache to hold a 1.5 s p99 — a hit is 60 ms, a cold generation 2.1 s, and a per-hit model call pushes p99 past budget. Which caching design fits this service?

Role practice MCQ for AI Engineer

Practice MCQ

A payment-risk model scores 4M transactions a day. Ground truth lands 48 hours after scoring, so realised precision@k is measurable within days. Shifts arrive irregularly: last March a new merchant category doubled one feature's rate with no pipeline change, and each day below the 0.75 precision@k floor costs about $40k in chargebacks. Every model that reaches traffic must pass the same validation and evaluation gates as a code release. Which retraining policy fits?

Role practice MCQ for AI Engineer

Practice MCQ

A payments agent calls twelve in-house tools and settles refunds unattended. Three failures are on record: the model sent 12.50 for an amount field that wants minor units, update_account hides account closure behind a benign name, and every error returns a plain string the model retries blind. The tool boundary rejects 4% of calls. Conversations are single turn, no replay harness exists, and no tool-layer rewrite is scheduled. Which tool-calling policy should the team adopt?

Role practice MCQ for AI Engineer

Practice MCQ

A payments service has an LLM emit one JSON record per request, straight into an automated settlement ledger. Records must satisfy cross-field invariants (end after start, amount consistent with currency), and failover sends traffic to a second provider. The path runs at 200 req/s under a 3-second p99 budget and a per-request cost cap. Which output policy should the service run in production?

Role practice MCQ for AI Engineer

Practice MCQ

A payments team must validate a new chargeback-scoring model against the live incumbent before launch. Traffic shifted two weeks ago, so logged data is stale; compliance forbids candidate predictions reaching users; candidate scoring takes 250 ms on GPU and the pool has no headroom; and promotion requires a per-case win rate on outcomes that land 45 days later. Which rollout design satisfies all four constraints?

Role practice MCQ for AI Engineer

Practice MCQ

A payments team retrains a churn model weekly on 18 months of transaction history and scores the week after each training cutoff. Each customer contributes up to 300 rows. The training job reports cross-validated AUC 0.88; the served model scores 0.79, and the nightly backfill replay writes to the same dashboard the quarterly review reads. Which evaluation design should the team adopt?

Role practice MCQ for AI Engineer

Practice MCQ

A postmortem on an agent that pays rent through a tool call: the provider committed a €2,400 transfer and returned 200, the executor pod restarted, killing the response, and the loop re-dispatched 30 s later with no call in flight — the rent was paid twice. Tracing shows late acknowledgements never duplicate; only the lost-acknowledgement window does. Transfers settle final with no recall, and customers legitimately repeat identical transfers. Which change fixes the root cause?

Role practice MCQ for AI Engineer

Practice MCQ

A retail team retrains its demand model weekly, and that retraining cannot pause for the launch, from a shared ingestion service feeding four pipelines. Last quarter an upstream release nulled `promo_spend`; the model silently learned to ignore it and accuracy decayed for four weeks before anyone noticed. Batch volumes swing with the season. The new data gate must block a bad batch before any training run starts. Which design holds up?

Role practice MCQ for AI Engineer

Practice MCQ

A ride-pricing model scores each request independently: offline, `trips_last_30d` is computed from a nightly SQL table; online, from the live event stream. Offline RMSE is 0.18 against 0.26 online for the same model version. The pricing decision is fully automated at a 45 ms p99 budget, the model runs unchanged for at least a year, and a feature spec defines the maximum train/serve distribution gap. Which train/serve skew policy should the team adopt?

Role practice MCQ for AI Engineer

Practice MCQ

A supervisor/worker run chains two workers: a researcher subagent's output is injected as context for a coder subagent. Staging shows error compounding — a hallucinated package name and a guessed account ID pass silently downstream and the coder builds on them. Which handoff design most directly prevents compounding across handoffs?

Role practice MCQ for Agentic AI Engineer

Practice MCQ

A support agent drafts from a history the summariser compresses: by turn 40 the user's turn-2 rule — 'never email me' — is gone from the prompt and the reply contradicts it. Sessions run 60 turns, past the fixed per-turn context budget, so the full transcript never replays, and the chat UI, the REST API and the tool loop can each send. Each user sets one or two free-text rules and the contract needs zero misses. The fix ships this sprint. Which design is defensible?

Role practice MCQ for AI Engineer

Practice MCQ

A support assistant cited a refund clause that does not exist. Sampling 60 flagged answers put every failure in one of four classes: 44% with no supporting chunk in the corpus, 21% with the right chunk retrieved but unused, 12% contradicting a retrieved chunk, 23% with invented specifics. Retrieval traces are logged. Policy questions must be answered, the escalation queue is already 12 minutes past SLA, and p95 answer latency sits at its 1.8 s budget. What should the action item commit to?

Role practice MCQ for AI Engineer