Research Engineer, Frontier Evals & Environments
This is an analysis of the public job description. It is not a confirmed description of this company's interview process, and no company has reviewed or endorsed it.
What the posting clearly emphasises
- Agent post-training, with coding, tool use and computer use as target capabilities.
- Multi-agent coordination and long-horizon execution.
- Building graders, evaluations, environments and training signal.
What that likely means for preparation
- Evaluation design is the centre of gravity: grader construction, reward hacking, contamination, and what a benchmark stops measuring once models optimise against it.
- Be able to design an environment out loud: task distribution, reset semantics, partial credit, and flakiness.
- Long-horizon failure analysis matters more than single-turn accuracy. Bring one concrete trace you debugged.
What stays unknown about the loop
- Whether an evals or environment take-home exists, and how much RL theory is expected versus systems judgment.
- The round order, count and length. Postings describe the job, not the loop.
- Whether AI assistants are allowed, required or banned in any coding exercise.
The company and role link are carried over. Paste the full job description on the next page for the highest-fidelity map — we save the link but do not read it for you.
Build my map for this role