Can you trust a data agent to do real work, reliably?

Stakeholder questions against a realistically messy warehouse. Graded the way you would review an analyst whose work is your responsibility.

Real data work is only worth delegating if you can trust what comes back. Models have gotten good enough that you can accept some answers and move on while other answers you have to check because they're confidently wrong.

But an agent you have to re-check on every answer hasn't done the work. It just moved it onto you. CandidBench measures what kinds of work a data agent can take off your plate, and where it still needs your eyes.

CandidBench asks the questions every data analyst has been asked before. The warehouse behind them carries the ordinary defects real warehouses carry. Every question hides at least one common analytical trap and each question is asked three times. What gets graded is the number and the agent's analysis.

The world

One realistically messy warehouse

A benchmark made of clean toy tables is not measuring real work. Every question runs against one virtual company's warehouse. It carries the ordinary defects a real warehouse carries, nothing labels them, and the system has to notice. It is the smallest warehouse we could build that still breaks the systems we tested.

Finding the right join is part of the work

Real warehouses grow one table at a time, from different systems and different teams. The same dimension ends up with a different key name in each one, and nobody goes back to tidy it. None of this is a seeded defect. It is what the schema looks like, so all of it is published.

One world, and what we do not say about it

A stated limit. Real warehouses run to thousands of tables. This one runs to . We report the floor rather than the ceiling, because the failures start here. of the 42 questions were answered soundly on all three runs by nobody we tested. A larger warehouse would move those numbers. It would not change what they are about.

Every figure above is a query you can run against the world in the kit.

Results

The ladder, ranked by reliability

A system banks a question only when all three runs are sound, and loses a point if any run is confidently wrong. Read the two counts, not the single number. Banked is what is safe to delegate; bluffed is where it would confidently mislead you at least once. The net of the two is a convenient sort key and a poor headline, because it cancels: on this very ladder a system with 16 banked and 17 bluffed scores -2.4, and so does one with 0 banked and 1 bluffed.

Rows are grouped into bands, not places. A row leaves its band only when its interval clears the band leader's, so two rows in one band have not been told apart by the evidence behind them.

A confident error in a regulatory filing costs more than one in a marketing dashboard. Same runs, different trust. It is the one number on this page you set, not us.

A wrong number costs me little One wrong number is unrecoverable

One confident error cancels banked questions

The two scores on the real systems

Plot both for every system. The dashed line marks where they would agree. That is a system that never wavered from one run to the next. Every system sits below it. The drop is what asking three times costs them. Opus answers almost as soundly as Fable 5. You can rely on far less of it, because its answers wander. The two solid rules sit at zero on each axis. Zero is the point where answering beats staying silent, and it comes from the score definition rather than from our judgement. Nothing on this chart marks the point where a system is good enough to hand work over. We do not know where that point is. The leader clears zero on both axes and still bluffs 8 of 42 questions. Where the bar belongs depends on what one confident error costs you, which is yours to set and not ours. The slider on the ladder prices it.

A label sits on every point, so identity never rests on colour alone.
The same points as a table

What the interval covers. On this pool 3 of 15 ranked pairs actually separate, which gives 2 bands. Numbering the rows 1, 2, 3 would claim otherwise. The interval under each reliability score covers one source only: which questions are in the pool. It answers “how would this system do on questions like these?”, and against the fixed pool published here the point estimate is exact. It does not cover re-running the same system on the same questions, which is a second and comparable source. Two runs of one system came out 14.3 points apart. We report that below as the measurement it is, rather than folding it into the bar. With three repeats per question, no model of it reproduced the observed spread without shifting the centre off the score. Publishing a modelled bar we could not defend is the exact failure this benchmark grades others for.

A single run does not resolve reliability finely. Reliability is a difference of counts over a fixed pool of 42 questions, so one question changing outcome moves it by 2.4 points, or 4.8 if it swings from confidently wrong to banked. We ran one system twice. Same model, same harness, same prompt, same questions, same three repeats. The two runs came out 14.3 points apart. Question by question, 34 of 42 landed identically, and the 8 that moved went both ways. Read a gap of this size as two systems that have not been told apart yet, not as a ranking. Widening the gap needs more questions or more repeats, and we would rather say so than imply a precision the pool does not carry.

The second grade

What a question you can delegate actually costs

We grade trust and price separately. We never combine them. Flash on the Pi harness is the cheapest system here per usable question, and it ranks last on trust. One number holding both would say nothing. A system that banks no questions gets no efficiency grade. You cannot price work it never delivered.

Reliability against cost

What the headline costs. Reliability against mean cost per question. The rule at zero is where a system banks as many questions as it bluffs. Above it you gain more than you lose. Cost uses a log scale. The systems span three orders of magnitude. There is no matching line on the cost axis, because we have no price for a person answering the same question.

x-axis

Every value in this chart is in the grade table above.

What is graded

The SQL can be right and the story still wrong

An expert builds trust by reading the SQL. But a correct query can still carry a plain-English explanation that contradicts it. The user reads that story and acts on it. So this benchmark grades the analysis the agent shows the user. Every answer lands in one of seven Grade classes. Two of them break trust. An agent that flags the trap or asks first earns credit for doing its job.

Questions asked the way a stakeholder asks them

Grade classes

Trust composition

Every answer from every system, by kind. The grey band is work the system declined. You got nothing from it. You also have nothing to unpick. These shares count each answer. The ladder above averages by question, so the two differ slightly.

What a value-only benchmark cannot see

The closest prior work grades the value alone. Every question has a fixed answer. A key matches it exactly, with no judge in the loop. That reproduces perfectly, which is why the approach appeals. It also cannot tell a confidently wrong agent from a calibrated one.

Value-only grading

  • Deterministic end to end. Anyone can reproduce the score exactly.
  • Scores the number, so an agent that guesses well ranks alongside one that verifies.
  • Its own authors found the gap: contestants optimised for hitting values rather than building verification.
  • A wrong answer and a flagged wrong answer score the same. Both simply miss.

Grading what it says

  • Scores the answer plus the sentence next to it, which is what a user actually reads.
  • Separates a confidently wrong answer from one that named the trap or asked.
  • Pays for it with a judge on the framing call, and publishes how much of the score depends on one.
  • Tells a decline and a clarifying question apart from a wrong answer, instead of scoring all three as a miss.

How the judge itself is checked, and what that does not cover. The judge is scored against a frozen set of 78 constructed answers across 12 tasks whose correct grade is known because we wrote them, the same way each hazard in the world is planted on purpose. 28 of them are controls that differ from a passing answer by a single clause: delete the sentence naming the trap and the grade must fall, delete a false claim and it must rise, and a hedge that names nothing must not pass at all. It scored 77 of 78, and every control held. No human grading round has been run, and this benchmark does not claim human-validated grading. What a constructed answer cannot settle is whether a caveat a real model wrote under real ambiguity names the hazard well enough to earn your trust. That residual is stated rather than measured.

Both designs are defensible. They answer different questions. This one is worth its judge only if you care whether a system tells you when it is unsure. A recent data-agent benchmark that value-matches by substring says the limit outright: its grader passes “an agent that returns correct values alongside incorrect ones.”

About the score

Two scores, because trust has two parts

A data agent has to pass two tests. Are the answers good? And are they good every single time? A system can pass one and fail the other. A confident wrong answer counts against both.

Net soundness, per answer

Take every answer a system gives. Count the sound ones. Subtract the confidently wrong ones. That is net soundness. It reads every answer, not every question. It asks one thing. When this agent answers, does it help you more often than it misleads you? It does not care whether the same question gets the same answer twice.

Reliability, per question

We ask every question three times. A question is safe to delegate only when all three runs are sound. One confidently wrong run makes it actively misleading. Reliability is the first count minus the second. Reliability asks the harder thing. Can you delegate this once and stop checking? Everything in between is the grey band on the ladder, and you review those yourself. Right two times in three counts the same as never.

Two ways to read the same runs

One system, asked every question three times. Sound on two runs. Confidently wrong on the third. The top bar reads every answer. The bottom bar reads every question, and collapses each group of three into the one question it decides. The same runs earn a helpful score and the worst score there is.

A system's score never depends on who else is on the ladder. Both scores count a system's own runs against a fixed pool, so adding or removing a system never moves anyone else's number. HELM moved off its mean-win-rate aggregate because a score defined against the comparison set inverts ranks whenever that set changes.

Neither score is new. Subtracting errors from correct answers and scoring a decline at zero is the reject-option rule from Chow 1970 at unit penalty. Franc et al. give the modern treatment. The AA-Omniscience Index and Kalai et al. rank models this way today. The penalty here attaches to a confident wrong story, not to a wrong label. A best-of-k benchmark asks whether a system can ever succeed. Reliability asks whether it succeeds every time.

Difficulty

Four tiers, and a wall at the top

The reference systems set the difficulty, not our sense of how hard a question looks. The strongest systems have nearly solved Easy and Medium. Brutal is where every system still fails. It is the largest tier in the pool.

Solve rate by tier

Colour steps with the solve rate. Every cell also prints the value, so colour reinforces the number rather than carrying it alone.

0%100%

Which kinds of work each system handles

You would not delegate every kind of question equally. This is the same pool, cut by the kind of analytical trap each question hides. It shows where a system is steady and where it is not. The column headers carry the question counts. A kind with few questions is a thin sample, and the page marks it as such.

0%100%
Explore data

Every question, and what each system did with it

All 42 questions, in the words a stakeholder would use. The page does not show the trap each one hides. The answer key stays held back so the benchmark keeps measuring. What it does show is every run: one dot per repetition, three per system per question.

Questions are sorted by how many systems could bank them, so the table reads as a gradient. The top is work most of the field can take off your hands. The bottom is the wall. Pick a system to see its questions on their own, sorted by what it did.

Reuse

Citing and reproducing

How to cite

Cite the frozen release, not the project. The numbers move when the pool or the grader changes. A citation with no release label points at whatever the page says today.


      

Everything is at the kit repo and the world dataset.

Run it yourself

The world, the questions, the answer schema and the grader source are public. The answer key for the scored pool is not.

  • Answer the questions with whatever you like, then open a pull request adding one answers.json to the kit's submissions/. CI checks the shape at once. Grading is manual and offline, on the machine that holds the answer key. The scorecard comes back on your PR. It carries pool totals, never a per-question key. Then your system joins the ladder.
  • A self-check set ships with its answer key so you can prove your harness emits gradable output before spending a run. Those questions carry no score and never did, so you give nothing away.
  • The grader source ships with the kit even though you cannot run it against the scored key. You should be able to read exactly how the grader will judge your answer.

A benchmark whose answer key is public stops measuring what it set out to measure. The first system trained on it wins, and the number stops meaning anything. Holding the key back is what keeps the ladder worth reading. There is deliberately no scoring endpoint, for the same reason. A service that grades has to hold the key. A rich scorecard returned on demand lets someone diff near-identical submissions until the key falls out of it. A human reading each pull request makes that expensive for free.

What this is built on

Two benchmarks this one takes from directly. Neither is a footnote: the first supplied the traps and the second supplied the shape.