Skip to content

Anthropic · public posting read

Research Engineer, Model Evaluations

This is an analysis of the public job description. It is not a confirmed description of this company's interview process, and no company has reviewed or endorsed it.

Source: https://job-boards.greenhouse.io/anthropic/jobs/5198255008

Last verified Sep 16, 2026. Postings close without notice; check the source before you rely on it.

What the posting clearly emphasises

  • Designing evaluations and the distributed infrastructure to run them.
  • Metrics, dashboards and debugging anomalous evaluation results.
  • Prompting, sampling and scaffolding choices; on-call production support.

Paraphrased from the public posting.

What that likely means for preparation

  • The strongest single preparation is a debugging story: an eval number moved, and you found out whether the model, the harness or the data changed.
  • Expect distributed systems questions framed around eval throughput, retries, determinism and cost.
  • On-call is named in the posting, so operational judgment is fair game: alerting, triage, and what you page a human for.

Inference from the posting. Not confirmed.

What stays unknown about the loop

  • How much of the loop is infrastructure engineering versus evaluation methodology.
  • The round order, count and length. Postings describe the job, not the loop.
  • Whether AI assistants are allowed, required or banned in any coding exercise.

Anything a recruiter tells you outranks every inference on this page.

Your own map

The company and role link are carried over. Paste the full job description on the next page for the highest-fidelity map — we save the link but do not read it for you.

Build my map for this role

Free map. No card. $39 only if you unlock.