Skip to content

Synthetic evaluation — not a customer story

We planted a subtle weakness. The map found it.

Everything here is synthetic: the candidate, the work history and the recruiter notes are fictional, and the role context is a public job description used as a test target. No real person, and no endorsement.

The test: can the product find a risk a human reviewer would skim past, explain how it surfaces in an interview, and turn it into prep — without inventing a confirmed round.

The setup

Late creative alignment

LLM Engineer, generative game content (synthetic evaluation against a public job description)

Fictional recruiter notes: the loop includes a technical screen and a project deep dive. Nothing was said about a workshop or a collaboration round.

01 — The line we buried

Rewrote the character-dialogue prompt architecture end to end and shipped it, then walked the art and narrative leads through the new behaviour once it was stable.

It reads like a win. The problem is the order: Major prompt architecture changes are completed independently and creative partners are involved only after the behaviour is already set.

02 — What the free map surfaced

Creative partners are brought in late, after behaviour is already set

Your strongest prompt-architecture work reads as independently completed and reviewed afterwards. For a role where model behaviour is a creative artefact, that timing is the thing most likely to get pressed: interviewers ask when alignment happened, not whether the system worked.

03 — How it could be questioned

  • When in the change did you bring the narrative and art leads in?
  • Describe a time a creative partner disagreed with model behaviour you had already shipped.
  • How do you co-define intended character behaviour before you implement it?

04 — What changed in the paid prep

Top move
Rebuild your deep-dive story so creative alignment happens before implementation, not after.
Scenario
A narrative lead says the new dialogue model is technically better but off-voice. Walk through how you would have avoided that before implementation, and what you would change now.
Drill
Rewrite one shipped prompt-architecture change as a collaboration timeline: who defined intended behaviour, when, against which examples, and where sign-off happened before code.

05 — What it refused to claim

Only the rounds the fictional recruiter named stayed confirmed: Technical screen, Project deep dive. Everything the planted risk implied stayed labelled as inference or unknown:

  • Creative collaboration probeLIKELY
  • Behavioural roundLIKELY
  • Design workshop with creative leadsUNKNOWN

Other risk classes in the same test set

These fixtures run in the automated regression suite, so the behaviour above is checked on every change.

  1. Impressive metric, weak measurement substantiation

    Improved answer accuracy by 40% after rebuilding the retrieval pipeline.

    Detected A headline number with no measurement story behind it

    Prep change Attach a baseline, an evaluation set, and a stated limitation to your headline metric before the first call.

  2. Strong technical depth, unclear production deployment

    Built and benchmarked three fine-tuned summarisation models; the internal pilot ran for two months.

    Detected Depth is clear; production ownership is not

    Prep change Label the pilot honestly and prepare the production path you would run today.

  3. Strong resume, unknown coding round

    Eight years shipping ML systems; most recent coding was review and design, not implementation.

    Detected Hands-on coding is the untested part of a strong record

    Prep change Ask the recruiter for the coding format and AI-tool policy, then warm up implementation this week.

Run it on your own interview

The free map names your biggest exposure. The Full Interview Map resolves it, $39 once.