Skip to content

Anthropic · public posting read

Research Engineer, Machine Learning (Reinforcement Learning)

This is an analysis of the public job description. It is not a confirmed description of this company's interview process, and no company has reviewed or endorsed it.

Source: https://job-boards.greenhouse.io/anthropic/jobs/4613568008

Last verified Sep 16, 2026. Postings close without notice; check the source before you rely on it.

What the posting clearly emphasises

  • Reinforcement learning applied to large language models.
  • Agentic tool use, computer use and code generation.
  • Distributed systems, training environments, evaluation and performance.

Paraphrased from the public posting.

What that likely means for preparation

  • Expect systems design in the training sense: data flow, throughput, checkpointing, and where a run stalls.
  • Be able to discuss reward design and its failure modes concretely, including reward hacking you have personally seen.
  • Performance reasoning — utilisation, memory, communication overhead — is likely to come up in a design conversation.

Inference from the posting. Not confirmed.

What stays unknown about the loop

  • How deep the loop goes into RL theory versus engineering throughput, and whether a coding exercise is included.
  • The round order, count and length. Postings describe the job, not the loop.
  • Whether AI assistants are allowed, required or banned in any coding exercise.

Anything a recruiter tells you outranks every inference on this page.

Your own map

The company and role link are carried over. Paste the full job description on the next page for the highest-fidelity map — we save the link but do not read it for you.

Build my map for this role

Free map. No card. $39 only if you unlock.