Can your system find the buried risk?
Four synthetic fixtures. Each plants one subtle weakness in a fictional work history and states the ground truth: what should be detected, how an interviewer could probe it, how it should change preparation, and what must stay inferred rather than confirmed.
All fixtures are synthetic. No real candidate, recruiter, resume or company interview process is represented, and no formal academic validation is claimed.
What it tests
- Detection: is the planted risk surfaced at all, using the expected terms?
- Interpretation: does it become probes an interviewer could actually ask?
- Propagation: does the preparation output change because of the risk?
- Calibration: are inferred and unknown rounds kept out of confirmed?
It tests candidate-risk detection and confidence calibration. It does not test prediction of any company’s real interview process.
The fixtures
Late creative alignment
Major prompt architecture changes are completed independently and creative partners are involved only after the behaviour is already set.
Impressive metric, weak measurement substantiation
A headline number with no stated baseline, evaluation set, or measurement method behind it.
Strong technical depth, unclear production deployment
Deep modelling work with no evidence that anything reached and survived production traffic.
Strong resume, unknown coding round
Seniority has moved the candidate away from hands-on implementation while the coding format stays unknown.
You may use these synthetic fixtures to test interview-prep or candidate-analysis systems. Attribution to LLMInterview.com is appreciated but not required.
These same fixtures run in our own regression suite. See the worked example or how confidence is assigned.