Format. An onsite round where you build, not just design. One report gives 90 minutes: 75 on your own, with LLM tools allowed, then 15 minutes discussing it with the interviewer. You write a candidate-search system and run it end to end within the 75 minutes.
The task. The input is a set of job descriptions, each with a title, a description and hard criteria. For each one, return the 10 best-matching candidates from a candidate database.
What is expected. The candidate who described it in detail says the workload is heavy. The expectation is to use an LLM to write code fast while still bringing your own ideas, and to show you can debug and dig into problems. That candidate leaned on Claude and GPT and built a basic retrieve-then-rank pipeline. At runtime, an LLM turned each job description's hard criteria into filtering logic for the query. In the discussion, the interviewer asked how you used the LLM, how you designed the system, and how you found and fixed problems along the way.
Outcomes. Both candidates who described the build round were rejected, and a commenter from the same onsite also got a template rejection two days later. The detailed reporter felt the interview went fine and guessed the evaluation results were not optimized enough; that is a guess, not feedback. No report says what the evaluation measures or what counts as a good result.
The rules on LLM use have been inconsistent. One candidate's pre-interview email said no LLMs in the coding session, but on the day the coordinator's links said LLMs were allowed for both coding and search. Confirm the rules with the coordinator before you start.
MLE variant. An MLE onsite asked for this as a 45-minute system design instead: given a job description, design a candidate search system. That report gives no further detail.
Ask before you build (the reports do not say): what fields the candidate records have, what the evaluation endpoint scores and how often you may call it, and whether a hard criterion is a strict filter or can be traded off against a better overall match.
Interviewer hints: The 15-minute discussion covers how you used the LLM, how you designed the system, and how you found and fixed problems.