Batch Inference System Design

System DesignAnthropicLast reported August 2026High Frequency
Reported
12× across candidate reports
First seen
May 2025
Last reported
August 2026
Also asked in
Phone Screen, OA
Reported outcome
mixed

Problem Overview

Design a batch inference system for an LLM serving backend. The core API is: batch_infer(list<string> input) -> list<string> output. Multiple incoming single API requests must be aggregated server-side into batches before being dispatched to GPU workers, because one GPU can only process one batch at a time. Key constraints/goals: minimize…

  • 6 candidate-reported follow-ups — the exact probes interviewers asked, with the trigger for each
  • The rest of the problem statement — full requirements, constraints, and edge cases
  • Approach and trade-offs — what passing candidates did, and the mistakes that sink people
Unlock the full Anthropic catalog
Full problem statements, candidate-reported follow-ups, and walkthroughs — for every Anthropic question.
Unlock with Pro
Already a member? Sign in
Verified Source
Every question is reconstructed from multiple independent candidate reports. Verbatim follow-ups, not invented ones.
Codex Fact-Checked
Technical claims, formulas, and scale numbers are reviewed against primary sources.
Interviewer Follow-ups
The exact follow-ups reported by candidates, with the trigger that prompts each one — plus the mistakes that sink people.
Is this helpful?