← Back to AI Data Labs
RecruiterOpsBench

Recruiting coordination

A live recruiting desk — job descriptions, an applicant tracker, 77–78 resumes, a shared calendar, and four people who each hold one piece of the answer. Frontier models schedule confidently here, and get it wrong in ways a recruiter would notice immediately.

Schedule a call
4
models measured
two providers, and a 45x spread in what a run costs
5
scenarios that discriminate
out of a much larger suite
4
parties to coordinate
agency, recruiter, manager, candidate
77–78
resumes per workspace
every scenario has its own; each transcript states that scenario's counts
ModelPasses every runPasses / $ Median runRuns
claude-sonnet-560%2.2154s15
claude-opus-540%1.193s15
gpt-5.6-luna0%0.0228s15
gpt-5.6-sol0%0.0227s24

What this sample can and cannot separate. At three runs per scenario, claude-sonnet-5, claude-opus-5 and gpt-5.6-luna are not distinguishable from one another — the top two differ by one run out of the fifteen each of them ran (p = 0.500), so the order of those rows is noise and we are not claiming a winner among them. What the sample does establish is the gap to gpt-5.6-sol: p = 0.003 on the raw runs, and it scores below every model above it on 3 of the 5 scenarios. Tiers are separable here; rankings inside a tier are not, and a league table would imply otherwise. Two rows share 0% here and are still coloured differently, which is not a slip: sweeping no scenario is a floor that hides a real difference underneath it — gpt-5.6-luna passed 9 of 15 runs, gpt-5.6-sol passed 9 of 24 runs. The colour follows the test on those runs, not the rounded figure above it.

Passes every run is the share of the 5 scenarios below where the model passed every run of that scenario — pass^3 in the usual notation, not pass@3, which asks only whether one attempt in 3 succeeded. A desk needs the work done every time, so that is what is counted; each card below shows its own runs, so the arithmetic can be redone by hand. Passes / $ uses that same every-run definition, so both columns answer one question rather than two. Cost is computed from the token counts each provider reported, at list prices verified on 2026-08-17.

This is a sample. The scenarios below are a selection from a much larger recruiting suite we maintain — enough to show the shape of the failures and how we grade them, not the whole set. We build the same depth for whichever workflow you care about.

Where they fail

Each scenario links to complete transcripts of individual runs — the model's own reasoning, every tool call and what it returned, the files it opened, what it left in the workspace, and the criterion that caught it. A failing run and a passing run of the same task, side by side.

One constraint, stated once, by the candidate

Four interviews have to fill a three-hour window exactly. One candidate can only start after 11:00, and he is the only person who ever says so.

67%
claude-opus-5
2 of 3 runs
67%
claude-sonnet-5
2 of 3 runs
0%
gpt-5.6-sol
0 of 3 runs
67%
gpt-5.6-luna
2 of 3 runs
What we are testing

Whether one constraint, stated once, by one person, survives all the way to the booking.

What should have happenedWhat the model did instead
Do not book the candidate before the time he said he is availableBooked him an hour before he was free — or left him out entirely
One interview for each shortlisted candidate, and nobody elseBooked nothing at all for the shortlist
Every interview is exactly 45 minutes.no interviews were booked at all

Why that is wrong: A candidate is booked an hour before the time he told us he was free — or left out of the day entirely. Either way he finds out from the invite.

Open the full transcript — its reasoning, every tool call, the data it saw
All 7 criteria, and how often each model passed each one — every one has to pass for the run to pass
Criterion claude-opus-5claude-sonnet-5gpt-5.6-solgpt-5.6-luna
Exactly one interview per person who should be interviewed, and nobody else booked.
read from the calendar
3/33/30/32/3
Every interview is exactly 45 minutes.
read from the calendar
3/33/33/32/3
Every interview starts and ends within 10:00–13:00 on 2026-07-02.
read from the calendar
3/33/33/32/3
No two interviews overlap each other.
read from the calendar
3/33/33/32/3
Karthik Menon is not booked before 11:00 UTC.
read from the calendar
2/32/30/32/3
Each interviewee is on their own invite.
read from the calendar
3/33/33/32/3
Nothing that was already on the calendar was moved, cancelled or deleted.
read from the calendar
3/33/33/33/3

The rule lives in a document nobody mentions

House policy requires the hiring manager on every candidate invite. It sits in the same folder as the job description and the request never refers to it.

67%
claude-opus-5
2 of 3 runs
67%
claude-sonnet-5
2 of 3 runs
33%
gpt-5.6-sol
1 of 3 runs
33%
gpt-5.6-luna
1 of 3 runs
What we are testing

Whether it goes looking for the house rule nobody mentioned, in the folder it was already reading.

What should have happenedWhat the model did instead
Do not book the candidate before the time he said he is availableBooked him an hour before he was free — or left him out entirely
Book every interview the day can actually holdBooked fewer than fit, leaving usable time empty

Why that is wrong: A candidate is booked an hour before the time he told us he was free — or left out of the day entirely. Either way he finds out from the invite.

The trap this scenario was built around held. Every model got it right — put the hiring manager on every candidate invite, as house policy requires — on every run. What failed them instead: Do not book the candidate before the time he said he is available. We publish that rather than quietly re-labelling the scenario, because which traps do not fire is half of what an environment tells you.

Open the full transcript — its reasoning, every tool call, the data it saw
All 9 criteria, and how often each model passed each one — every one has to pass for the run to pass
Criterion claude-opus-5claude-sonnet-5gpt-5.6-solgpt-5.6-luna
No interview overlaps a meeting that was already on the calendar.
read from the calendar
3/33/33/33/3
Exactly 2 interview(s), each for someone through screening.
read from the calendar
3/33/31/32/3
Every interview is exactly 45 minutes.
read from the calendar
3/33/33/33/3
Every interview starts and ends within 10:00–13:00 on 2026-07-02.
read from the calendar
3/33/33/33/3
No two interviews overlap each other.
read from the calendar
3/33/33/33/3
Karthik Menon is not booked before 11:00 UTC, if booked at all.
read from the calendar
2/32/32/31/3
deepa.rao@acme.com is on every interview invite.
read from the calendar
3/33/33/33/3
The user is told, by name, who could not be scheduled.
read from what it told the person
3/33/33/33/3
Nothing that was already on the calendar was moved, cancelled or deleted.
read from the calendar
3/33/33/33/3

The free window is not actually free

The manager says the whole morning is available. A standing design review sits inside it and cannot move.

67%
claude-opus-5
2 of 3 runs
100%
claude-sonnet-5
3 of 3 runs
33%
gpt-5.6-sol
1 of 3 runs
67%
gpt-5.6-luna
2 of 3 runs
What we are testing

Whether it checks the calendar against a human claim that the morning is free.

What should have happenedWhat the model did instead
Do not book the candidate before the time he said he is availableBooked him an hour before he was free — or left him out entirely
Book every interview the day can actually holdBooked fewer than fit, leaving usable time empty
Say who could not be fitted inDropped a candidate without telling anyone

Why that is wrong: A candidate is booked an hour before the time he told us he was free — or left out of the day entirely. Either way he finds out from the invite.

The trap this scenario was built around held. Every model got it right — do not book on top of the design review that cannot move — on every run. What failed them instead: Do not book the candidate before the time he said he is available. We publish that rather than quietly re-labelling the scenario, because which traps do not fire is half of what an environment tells you.

Open the full transcript — its reasoning, every tool call, the data it saw
All 8 criteria, and how often each model passed each one — every one has to pass for the run to pass
Criterion claude-opus-5claude-sonnet-5gpt-5.6-solgpt-5.6-luna
No interview overlaps a meeting that was already on the calendar.
read from the calendar
3/33/33/33/3
Exactly 2 interview(s), each for someone through screening.
read from the calendar
3/33/32/32/3
Every interview is exactly 45 minutes.
read from the calendar
3/33/33/33/3
Every interview starts and ends within 10:00–13:00 on 2026-07-02.
read from the calendar
3/33/33/33/3
No two interviews overlap each other.
read from the calendar
3/33/33/33/3
Karthik Menon is not booked before 11:00 UTC, if booked at all.
read from the calendar
2/33/32/33/3
The user is told, by name, who could not be scheduled.
read from what it told the person
3/33/33/32/3
Nothing that was already on the calendar was moved, cancelled or deleted.
read from the calendar
3/33/33/33/3

Four parties, one interview slot

An external agency owns the candidate relationship, the company recruiter owns process, the hiring manager owns the slot. The agent has to satisfy all three and work out who may be told what.

100%
claude-opus-5
3 of 3 runs
100%
claude-sonnet-5
3 of 3 runs
22%
gpt-5.6-sol
2 of 9 runs
67%
gpt-5.6-luna
2 of 3 runs
What we are testing

Whether an agent can serve four parties with conflicting claims on one slot — and work out that the obliging move, writing to the candidate directly, is the one thing it must not do.

What should have happenedWhat the model did instead
Write to the agency with the time that was bookedSent no message at all — to anyone
Do not book on top of a meeting that is already thereBooked the interview straight over the design review

Why that is wrong: The interview is in the calendar and nobody outside the building knows. The candidate is represented by the agency, and the agency was never written to — so the one person who has to show up is never told.

Open the full transcript — its reasoning, every tool call, the data it saw
All 12 criteria, and how often each model passed each one — every one has to pass for the run to pass
Criterion claude-opus-5claude-sonnet-5gpt-5.6-solgpt-5.6-luna
Exactly 1 interview(s), each for someone through screening.
read from the calendar
3/33/39/93/3
Every interview is exactly 45 minutes.
read from the calendar
3/33/39/93/3
Every interview starts and ends within 09:00–13:00 on 2026-07-02.
read from the calendar
3/33/39/93/3
Karthik Menon is not booked before 11:00 UTC.
read from the calendar
3/33/39/93/3
No interview overlaps a meeting that was already on the calendar.
read from the calendar
3/33/37/93/3
farida.khan@acme.com is on every interview invite.
read from the calendar
3/33/39/93/3
deepa.rao@acme.com is on every interview invite.
read from the calendar
3/33/39/93/3
karthik.menon@mailbox.example is NOT on any invite.
read from the calendar
3/33/39/93/3
No message was sent to karthik.menon@mailbox.example.
read from the mail it sent
3/33/39/93/3
A message was sent to rehan.qureshi@talentbridge.io, stating the time actually booked.
read from the mail it sent
3/33/32/92/3
Nothing sent to rehan.qureshi@talentbridge.io contains the internal material.
read from the mail it sent
3/33/39/93/3
Nothing that was already on the calendar was moved, cancelled or deleted.
read from the calendar
3/33/39/93/3

The shortlist is out of date in both directions

One shortlisted candidate has withdrawn by email; another applied after the shortlist was written and is on the tracker. They cancel out, so the count still looks right.

100%
claude-opus-5
3 of 3 runs
100%
claude-sonnet-5
3 of 3 runs
83%
gpt-5.6-sol
5 of 6 runs
67%
gpt-5.6-luna
2 of 3 runs
What we are testing

Whether it re-derives who is actually still in play, rather than trusting a shortlist document that is out of date in both directions at once.

What should have happenedWhat the model did instead
Do not book the candidate before the time he said he is availableBooked him an hour before he was free — or left him out entirely
Interview exactly the people still in playBooked the stale shortlist — the withdrawal and the late applicant both missed

Why that is wrong: A candidate is booked an hour before the time he told us he was free — or left out of the day entirely. Either way he finds out from the invite.

The trap this scenario was built around — interview exactly the people still in play — caught 1 run. More often it was do not book the candidate before the time he said he is available.

Open the full transcript — its reasoning, every tool call, the data it saw
All 8 criteria, and how often each model passed each one — every one has to pass for the run to pass
Criterion claude-opus-5claude-sonnet-5gpt-5.6-solgpt-5.6-luna
Exactly one interview per person who should be interviewed, and nobody else booked.
read from the calendar
3/33/36/62/3
Every interview is exactly 45 minutes.
read from the calendar
3/33/36/63/3
Every interview starts and ends within 10:00–13:00 on 2026-07-02.
read from the calendar
3/33/36/63/3
No two interviews overlap each other.
read from the calendar
3/33/36/63/3
Karthik Menon is not booked before 11:00 UTC.
read from the calendar
3/33/35/62/3
Each interviewee is on their own invite.
read from the calendar
3/33/36/63/3
Nothing that was already on the calendar was moved, cancelled or deleted.
read from the calendar
3/33/36/63/3
'Compensation bands 2026 — CONFIDENTIAL.md' has no sharing grants at all.
read from Drive permissions
3/33/36/63/3

How it is graded

Every task is scored on binary criteria that must all pass — a partially correct answer is a failure, because that is what it is on a real desk. Checks read the final state of the workspace and the messages the agent sent: what is on the calendar, who is on the invite, what was written to whom, what was left alone. Every criterion on a transcript says which of those it was read from.

Each scenario runs at least three times per model, and more where the first three disagreed — which is why some cells below show 6 or 9 runs. Each card states its own count, and the pass rates in the table above are the plain average of the per-scenario rates, so a scenario that was run more times does not count for more.

Deterministic grading, validated offline against a reference solution and every known failure mode before any model was run.

The same tools every time41 tools — 17 Drive, 14 Gmail, 7 Calendar, 3 utility — handed to the model whole on every task. Nothing is narrowed to the task at hand, so finding what matters is part of the job.
Room to workUp to 24 steps per run. The longest run here used 16, and none was stopped by the limit — so a failure below is a wrong answer, not a run that ran out of room.
Reasoning effortHigh, pinned identically for every model. Both providers accept higher settings; none were used, so this compares the models at one level rather than each at its ceiling.
The person on the other sideclaude-opus-5 at maximum reasoning effort role-plays the requester — pinned to the same model and effort for every run of every model, so a weak simulator's confusion is never charged to the agent under test. It answers questions and confirms decisions under rules that forbid naming tools, suggesting where to look, or adding requirements. It never hints, so what is measured is the agent rather than hint-following.
Who decides pass or failCode, not a model. No model judges another model here, and every criterion names the record in the workspace it rests on — visible in the transcripts.

*The comparison metrics are based on sample data in the smaller test environment.

We build these for your workflows

Give us a workflow your agents get wrong in production and we will build the environment, the rubric and the failing traces.

Schedule a call