Recruiting coordination
A live recruiting desk — job descriptions, an applicant tracker, 77–78 resumes, a shared calendar, and four people who each hold one piece of the answer. Frontier models schedule confidently here, and get it wrong in ways a recruiter would notice immediately.
Schedule a call →two providers, and a 45x spread in what a run costs
out of a much larger suite
agency, recruiter, manager, candidate
every scenario has its own; each transcript states that scenario's counts
| Model | Passes every run | Passes / $ | Median run | Runs |
|---|---|---|---|---|
| claude-sonnet-5 | 60% | 2.2 | 154s | 15 |
| claude-opus-5 | 40% | 1.1 | 93s | 15 |
| gpt-5.6-luna | 0% | 0.0 | 228s | 15 |
| gpt-5.6-sol | 0% | 0.0 | 227s | 24 |
What this sample can and cannot separate. At three runs per scenario, claude-sonnet-5, claude-opus-5 and gpt-5.6-luna are not distinguishable from one another — the top two differ by one run out of the fifteen each of them ran (p = 0.500), so the order of those rows is noise and we are not claiming a winner among them. What the sample does establish is the gap to gpt-5.6-sol: p = 0.003 on the raw runs, and it scores below every model above it on 3 of the 5 scenarios. Tiers are separable here; rankings inside a tier are not, and a league table would imply otherwise. Two rows share 0% here and are still coloured differently, which is not a slip: sweeping no scenario is a floor that hides a real difference underneath it — gpt-5.6-luna passed 9 of 15 runs, gpt-5.6-sol passed 9 of 24 runs. The colour follows the test on those runs, not the rounded figure above it.
Passes every run is the share of the 5 scenarios below where the model passed every run of that scenario — pass^3 in the usual notation, not pass@3, which asks only whether one attempt in 3 succeeded. A desk needs the work done every time, so that is what is counted; each card below shows its own runs, so the arithmetic can be redone by hand. Passes / $ uses that same every-run definition, so both columns answer one question rather than two. Cost is computed from the token counts each provider reported, at list prices verified on 2026-08-17.
This is a sample. The scenarios below are a selection from a much larger recruiting suite we maintain — enough to show the shape of the failures and how we grade them, not the whole set. We build the same depth for whichever workflow you care about.
Where they fail
Each scenario links to complete transcripts of individual runs — the model's own reasoning, every tool call and what it returned, the files it opened, what it left in the workspace, and the criterion that caught it. A failing run and a passing run of the same task, side by side.
One constraint, stated once, by the candidate
Four interviews have to fill a three-hour window exactly. One candidate can only start after 11:00, and he is the only person who ever says so.
2 of 3 runs
2 of 3 runs
0 of 3 runs
2 of 3 runs
Whether one constraint, stated once, by one person, survives all the way to the booking.
| What should have happened | What the model did instead | |
|---|---|---|
| Do not book the candidate before the time he said he is available | Booked him an hour before he was free — or left him out entirely | 6× |
| One interview for each shortlisted candidate, and nobody else | Booked nothing at all for the shortlist | 4× |
| Every interview is exactly 45 minutes. | no interviews were booked at all | 1× |
Why that is wrong: A candidate is booked an hour before the time he told us he was free — or left out of the day entirely. Either way he finds out from the invite.
Open the full transcript — its reasoning, every tool call, the data it saw →All 7 criteria, and how often each model passed each one — every one has to pass for the run to pass
| Criterion | claude-opus-5 | claude-sonnet-5 | gpt-5.6-sol | gpt-5.6-luna |
|---|---|---|---|---|
| Exactly one interview per person who should be interviewed, and nobody else booked. read from the calendar | 3/3 | 3/3 | 0/3 | 2/3 |
| Every interview is exactly 45 minutes. read from the calendar | 3/3 | 3/3 | 3/3 | 2/3 |
| Every interview starts and ends within 10:00–13:00 on 2026-07-02. read from the calendar | 3/3 | 3/3 | 3/3 | 2/3 |
| No two interviews overlap each other. read from the calendar | 3/3 | 3/3 | 3/3 | 2/3 |
| Karthik Menon is not booked before 11:00 UTC. read from the calendar | 2/3 | 2/3 | 0/3 | 2/3 |
| Each interviewee is on their own invite. read from the calendar | 3/3 | 3/3 | 3/3 | 2/3 |
| Nothing that was already on the calendar was moved, cancelled or deleted. read from the calendar | 3/3 | 3/3 | 3/3 | 3/3 |
The rule lives in a document nobody mentions
House policy requires the hiring manager on every candidate invite. It sits in the same folder as the job description and the request never refers to it.
2 of 3 runs
2 of 3 runs
1 of 3 runs
1 of 3 runs
Whether it goes looking for the house rule nobody mentioned, in the folder it was already reading.
| What should have happened | What the model did instead | |
|---|---|---|
| Do not book the candidate before the time he said he is available | Booked him an hour before he was free — or left him out entirely | 5× |
| Book every interview the day can actually hold | Booked fewer than fit, leaving usable time empty | 3× |
Why that is wrong: A candidate is booked an hour before the time he told us he was free — or left out of the day entirely. Either way he finds out from the invite.
The trap this scenario was built around held. Every model got it right — put the hiring manager on every candidate invite, as house policy requires — on every run. What failed them instead: Do not book the candidate before the time he said he is available. We publish that rather than quietly re-labelling the scenario, because which traps do not fire is half of what an environment tells you.
Open the full transcript — its reasoning, every tool call, the data it saw →All 9 criteria, and how often each model passed each one — every one has to pass for the run to pass
| Criterion | claude-opus-5 | claude-sonnet-5 | gpt-5.6-sol | gpt-5.6-luna |
|---|---|---|---|---|
| No interview overlaps a meeting that was already on the calendar. read from the calendar | 3/3 | 3/3 | 3/3 | 3/3 |
| Exactly 2 interview(s), each for someone through screening. read from the calendar | 3/3 | 3/3 | 1/3 | 2/3 |
| Every interview is exactly 45 minutes. read from the calendar | 3/3 | 3/3 | 3/3 | 3/3 |
| Every interview starts and ends within 10:00–13:00 on 2026-07-02. read from the calendar | 3/3 | 3/3 | 3/3 | 3/3 |
| No two interviews overlap each other. read from the calendar | 3/3 | 3/3 | 3/3 | 3/3 |
| Karthik Menon is not booked before 11:00 UTC, if booked at all. read from the calendar | 2/3 | 2/3 | 2/3 | 1/3 |
| deepa.rao@acme.com is on every interview invite. read from the calendar | 3/3 | 3/3 | 3/3 | 3/3 |
| The user is told, by name, who could not be scheduled. read from what it told the person | 3/3 | 3/3 | 3/3 | 3/3 |
| Nothing that was already on the calendar was moved, cancelled or deleted. read from the calendar | 3/3 | 3/3 | 3/3 | 3/3 |
The free window is not actually free
The manager says the whole morning is available. A standing design review sits inside it and cannot move.
2 of 3 runs
3 of 3 runs
1 of 3 runs
2 of 3 runs
Whether it checks the calendar against a human claim that the morning is free.
| What should have happened | What the model did instead | |
|---|---|---|
| Do not book the candidate before the time he said he is available | Booked him an hour before he was free — or left him out entirely | 2× |
| Book every interview the day can actually hold | Booked fewer than fit, leaving usable time empty | 2× |
| Say who could not be fitted in | Dropped a candidate without telling anyone | 1× |
Why that is wrong: A candidate is booked an hour before the time he told us he was free — or left out of the day entirely. Either way he finds out from the invite.
The trap this scenario was built around held. Every model got it right — do not book on top of the design review that cannot move — on every run. What failed them instead: Do not book the candidate before the time he said he is available. We publish that rather than quietly re-labelling the scenario, because which traps do not fire is half of what an environment tells you.
Open the full transcript — its reasoning, every tool call, the data it saw →All 8 criteria, and how often each model passed each one — every one has to pass for the run to pass
| Criterion | claude-opus-5 | claude-sonnet-5 | gpt-5.6-sol | gpt-5.6-luna |
|---|---|---|---|---|
| No interview overlaps a meeting that was already on the calendar. read from the calendar | 3/3 | 3/3 | 3/3 | 3/3 |
| Exactly 2 interview(s), each for someone through screening. read from the calendar | 3/3 | 3/3 | 2/3 | 2/3 |
| Every interview is exactly 45 minutes. read from the calendar | 3/3 | 3/3 | 3/3 | 3/3 |
| Every interview starts and ends within 10:00–13:00 on 2026-07-02. read from the calendar | 3/3 | 3/3 | 3/3 | 3/3 |
| No two interviews overlap each other. read from the calendar | 3/3 | 3/3 | 3/3 | 3/3 |
| Karthik Menon is not booked before 11:00 UTC, if booked at all. read from the calendar | 2/3 | 3/3 | 2/3 | 3/3 |
| The user is told, by name, who could not be scheduled. read from what it told the person | 3/3 | 3/3 | 3/3 | 2/3 |
| Nothing that was already on the calendar was moved, cancelled or deleted. read from the calendar | 3/3 | 3/3 | 3/3 | 3/3 |
Four parties, one interview slot
An external agency owns the candidate relationship, the company recruiter owns process, the hiring manager owns the slot. The agent has to satisfy all three and work out who may be told what.
3 of 3 runs
3 of 3 runs
2 of 9 runs
2 of 3 runs
Whether an agent can serve four parties with conflicting claims on one slot — and work out that the obliging move, writing to the candidate directly, is the one thing it must not do.
| What should have happened | What the model did instead | |
|---|---|---|
| Write to the agency with the time that was booked | Sent no message at all — to anyone | 8× |
| Do not book on top of a meeting that is already there | Booked the interview straight over the design review | 2× |
Why that is wrong: The interview is in the calendar and nobody outside the building knows. The candidate is represented by the agency, and the agency was never written to — so the one person who has to show up is never told.
Open the full transcript — its reasoning, every tool call, the data it saw →All 12 criteria, and how often each model passed each one — every one has to pass for the run to pass
| Criterion | claude-opus-5 | claude-sonnet-5 | gpt-5.6-sol | gpt-5.6-luna |
|---|---|---|---|---|
| Exactly 1 interview(s), each for someone through screening. read from the calendar | 3/3 | 3/3 | 9/9 | 3/3 |
| Every interview is exactly 45 minutes. read from the calendar | 3/3 | 3/3 | 9/9 | 3/3 |
| Every interview starts and ends within 09:00–13:00 on 2026-07-02. read from the calendar | 3/3 | 3/3 | 9/9 | 3/3 |
| Karthik Menon is not booked before 11:00 UTC. read from the calendar | 3/3 | 3/3 | 9/9 | 3/3 |
| No interview overlaps a meeting that was already on the calendar. read from the calendar | 3/3 | 3/3 | 7/9 | 3/3 |
| farida.khan@acme.com is on every interview invite. read from the calendar | 3/3 | 3/3 | 9/9 | 3/3 |
| deepa.rao@acme.com is on every interview invite. read from the calendar | 3/3 | 3/3 | 9/9 | 3/3 |
| karthik.menon@mailbox.example is NOT on any invite. read from the calendar | 3/3 | 3/3 | 9/9 | 3/3 |
| No message was sent to karthik.menon@mailbox.example. read from the mail it sent | 3/3 | 3/3 | 9/9 | 3/3 |
| A message was sent to rehan.qureshi@talentbridge.io, stating the time actually booked. read from the mail it sent | 3/3 | 3/3 | 2/9 | 2/3 |
| Nothing sent to rehan.qureshi@talentbridge.io contains the internal material. read from the mail it sent | 3/3 | 3/3 | 9/9 | 3/3 |
| Nothing that was already on the calendar was moved, cancelled or deleted. read from the calendar | 3/3 | 3/3 | 9/9 | 3/3 |
The shortlist is out of date in both directions
One shortlisted candidate has withdrawn by email; another applied after the shortlist was written and is on the tracker. They cancel out, so the count still looks right.
3 of 3 runs
3 of 3 runs
5 of 6 runs
2 of 3 runs
Whether it re-derives who is actually still in play, rather than trusting a shortlist document that is out of date in both directions at once.
| What should have happened | What the model did instead | |
|---|---|---|
| Do not book the candidate before the time he said he is available | Booked him an hour before he was free — or left him out entirely | 2× |
| Interview exactly the people still in play | Booked the stale shortlist — the withdrawal and the late applicant both missed | 1× |
Why that is wrong: A candidate is booked an hour before the time he told us he was free — or left out of the day entirely. Either way he finds out from the invite.
The trap this scenario was built around — interview exactly the people still in play — caught 1 run. More often it was do not book the candidate before the time he said he is available.
Open the full transcript — its reasoning, every tool call, the data it saw →All 8 criteria, and how often each model passed each one — every one has to pass for the run to pass
| Criterion | claude-opus-5 | claude-sonnet-5 | gpt-5.6-sol | gpt-5.6-luna |
|---|---|---|---|---|
| Exactly one interview per person who should be interviewed, and nobody else booked. read from the calendar | 3/3 | 3/3 | 6/6 | 2/3 |
| Every interview is exactly 45 minutes. read from the calendar | 3/3 | 3/3 | 6/6 | 3/3 |
| Every interview starts and ends within 10:00–13:00 on 2026-07-02. read from the calendar | 3/3 | 3/3 | 6/6 | 3/3 |
| No two interviews overlap each other. read from the calendar | 3/3 | 3/3 | 6/6 | 3/3 |
| Karthik Menon is not booked before 11:00 UTC. read from the calendar | 3/3 | 3/3 | 5/6 | 2/3 |
| Each interviewee is on their own invite. read from the calendar | 3/3 | 3/3 | 6/6 | 3/3 |
| Nothing that was already on the calendar was moved, cancelled or deleted. read from the calendar | 3/3 | 3/3 | 6/6 | 3/3 |
| 'Compensation bands 2026 — CONFIDENTIAL.md' has no sharing grants at all. read from Drive permissions | 3/3 | 3/3 | 6/6 | 3/3 |
How it is graded
Every task is scored on binary criteria that must all pass — a partially correct answer is a failure, because that is what it is on a real desk. Checks read the final state of the workspace and the messages the agent sent: what is on the calendar, who is on the invite, what was written to whom, what was left alone. Every criterion on a transcript says which of those it was read from.
Each scenario runs at least three times per model, and more where the first three disagreed — which is why some cells below show 6 or 9 runs. Each card states its own count, and the pass rates in the table above are the plain average of the per-scenario rates, so a scenario that was run more times does not count for more.
Deterministic grading, validated offline against a reference solution and every known failure mode before any model was run.
| The same tools every time | 41 tools — 17 Drive, 14 Gmail, 7 Calendar, 3 utility — handed to the model whole on every task. Nothing is narrowed to the task at hand, so finding what matters is part of the job. |
| Room to work | Up to 24 steps per run. The longest run here used 16, and none was stopped by the limit — so a failure below is a wrong answer, not a run that ran out of room. |
| Reasoning effort | High, pinned identically for every model. Both providers accept higher settings; none were used, so this compares the models at one level rather than each at its ceiling. |
| The person on the other side | claude-opus-5 at maximum reasoning effort role-plays the requester — pinned to the same model and effort for every run of every model, so a weak simulator's confusion is never charged to the agent under test. It answers questions and confirms decisions under rules that forbid naming tools, suggesting where to look, or adding requirements. It never hints, so what is measured is the agent rather than hint-following. |
| Who decides pass or fail | Code, not a model. No model judges another model here, and every criterion names the record in the workspace it rests on — visible in the transcripts. |
*The comparison metrics are based on sample data in the smaller test environment.
We build these for your workflows
Give us a workflow your agents get wrong in production and we will build the environment, the rubric and the failing traces.
Schedule a call →