ML Coding: Top-p Sampling and Multi-Head Attention in NumPy

MLE / ResearchScale AILast reported March 2026Low Frequency
Reported
2× across candidate reports
First seen
February 2026
Last reported
March 2026
Reported outcome
unknown

Problem Overview

You implement four functions in three graded parts, in NumPy, not PyTorch, in a notebook that already defines softmax and comes with tests you have to pass. Part 3 builds on Part 2: Part 1 — Top-p sampling. top_p_filter(logits, p) keeps the smallest set of most-probable tokens whose total probability…

  • 1 candidate-reported follow-up — the exact probes interviewers asked, with the trigger for each
  • The rest of the problem statement — full requirements, constraints, and edge cases
  • Approach and trade-offs — what passing candidates did, and the mistakes that sink people
Unlock the full Scale AI catalog
Full problem statements, candidate-reported follow-ups, and walkthroughs — for every Scale AI question.
Unlock with Pro
Already a member? Sign in
Verified Source
Every question is reconstructed from multiple independent candidate reports. Verbatim follow-ups, not invented ones.
Codex Fact-Checked
Technical claims, formulas, and scale numbers are reviewed against primary sources.
Interviewer Follow-ups
The exact follow-ups reported by candidates, with the trigger that prompts each one — plus the mistakes that sink people.
Is this helpful?