Can you trust a data agent to do real work, reliably?

Stakeholder questions against a realistically messy warehouse. Graded the way you would review an analyst whose work is your responsibility.

Real data work is only worth delegating if you can trust what comes back. Models have gotten good enough that you can accept some answers and move on. Other answers you have to check because they're confidently wrong.

An agent you re-check on every answer hasn't done the work. It just moved it onto you. CandidBench measures what kinds of work a data agent can take off your plate, and where it still needs your eyes.

CandidBench asks the questions every data analyst has been asked before. The warehouse behind them carries the ordinary defects real data carries. Every question hides at least one common analytical trap. Each question is asked three times. What gets graded is the number and the agent's analysis.

The world

One realistically messy warehouse

A benchmark made of clean toy tables is not measuring real work. Every question in CandidBench runs against a single virtual company's warehouse with multiple ordinary defects. Nothing flags them. The system has to notice.

What is wrong with the data, and why we do not list it

The defects are real, and nothing labels them. We withhold which ones, and where, on purpose. Publishing the list would hand over the traps.

One world, not one per question

  • Every question runs against the same warehouse. Nobody can tune a system to a dataset it will meet once.
  • The schema is unfamiliar and inherited. A real warehouse names the columns, not the question that asks about them.
  • Finding the right table is part of the work, the same as it is for a new analyst on their first week.
  • Raw event logs sit beside the modelled tables, so some questions are only answerable from the logs.

Questions asked the way a stakeholder asks them

  • Phrased in business terms. Nothing names the trap, or hints that there is one.
  • No question says which table to use, what grain to aggregate at, or which rows to exclude.
  • Several are symptom-framed: something looks wrong, go and find out why.
  • Every question has a defensible answer. None is a trick, and none is unanswerable.

The world ships with the kit, so you can check yourself.

Results

The ladder, ranked by reliability

A system banks a question only when all three runs are sound, and loses a point if any run is confidently wrong. Read the two counts, not the single number. Banked is what you could hand over; bluffed is where it would confidently mislead you at least once. The net of the two is a convenient sort key and a poor headline, because it cancels: on this very ladder a system with 16 banked and 17 bluffed scores -2.4, and so does one with 0 banked and 1 bluffed.

Rows are grouped into bands, not places. A row leaves its band only when its interval clears the band leader's. On this pool 3 of 15 ranked pairs actually separate. That gives 2 bands. Inside a band the evidence has not told these systems apart, and numbering them 1, 2, 3 would say otherwise. The interval under each reliability score covers one source only: which questions are in the pool. It answers “how would this system do on questions like these?”, and against the fixed pool published here the point estimate is exact. It does not cover re-running the same system on the same questions, which is a second and comparable source. Two runs of one system came out 14.3 points apart. We report that below as the measurement it is, rather than folding it into the bar. With three repeats per question, no model of it reproduced the observed spread without shifting the centre off the score. Publishing a modelled bar we could not defend is the exact failure this benchmark grades others for.

The second grade

What a question you can hand over actually costs

We grade trust and price separately. We never combine them. Flash on the Pi harness is the cheapest system here per usable question, and it ranks last on trust. One number holding both would say nothing. A system that banks no questions gets no efficiency grade. You cannot price work it never delivered.

Reliability against cost

What the headline costs. Reliability against mean cost per question. The good region is up and to the left. More questions answered soundly every time, for less money. Cost uses a log scale. The systems span three orders of magnitude.

x-axis

Every value in this chart is in the grade table above.

What is graded

The SQL can be right and the story still wrong

An expert builds trust by reading the SQL. But a correct query can still carry a plain-English explanation that contradicts it. The user reads that story and acts on it. So this benchmark grades the analysis the agent shows the user. Every answer lands in one of seven Grade classes. Two of them break trust. An agent that flags the trap or asks first earns credit for doing its job.

Grade classes

Trust composition

Every answer from every system, by kind. The grey band is work the system declined. You got nothing from it. You also have nothing to unpick. These shares count each answer. The ladder above averages by question, so the two differ slightly.

Table view

What a value-only benchmark cannot see

The closest prior work grades the value alone. Every question has a fixed answer. A key matches it exactly, with no judge in the loop. That reproduces perfectly, which is why the approach appeals. It also cannot tell a confidently wrong agent from a calibrated one.

Value-only grading

  • Deterministic end to end. Anyone can reproduce the score exactly.
  • Scores the number, so an agent that guesses well ranks alongside one that verifies.
  • Its own authors found the gap: contestants optimised for hitting values rather than building verification.
  • A wrong answer and a flagged wrong answer score the same. Both simply miss.

Grading what it says

  • Scores the answer plus the sentence next to it, which is what a user actually reads.
  • Separates a confidently wrong answer from one that named the trap or asked.
  • Pays for it with a judge on the framing call, and publishes how much of the score depends on one.
  • Tells a decline and a clarifying question apart from a wrong answer, instead of scoring all three as a miss.

How the judge itself is checked, and what that does not cover. The judge is scored against a frozen set of 78 constructed answers across 12 tasks whose correct grade is known because we wrote them, the same way each hazard in the world is planted on purpose. 28 of them are controls that differ from a passing answer by a single clause: delete the sentence naming the trap and the grade must fall, delete a false claim and it must rise, and a hedge that names nothing must not pass at all. It scored 77 of 78, and every control held. No human grading round has been run, and this benchmark does not claim human-validated grading. What a constructed answer cannot settle is whether a caveat a real model wrote under real ambiguity names the hazard well enough to earn your trust. That residual is stated rather than measured.

Both designs are defensible. They answer different questions. This one is worth its judge only if you care whether a system tells you when it is unsure. A recent data-agent benchmark that value-matches by substring says the limit outright: its grader passes “an agent that returns correct values alongside incorrect ones.”

About the score

Two scores, because trust has two parts

A data agent has to clear two bars. A system can clear one and fail the other. When it answers, does it help you or mislead you? Can you hand a question over and stop checking? CandidBench scores both.

Net soundness, per answer

Take every answer a system gives. Count the sound ones. Subtract the confidently wrong ones. That is net soundness. It asks one thing. When this agent answers, does it help you more often than it misleads you? It does not care whether the same question gets the same answer twice.

Reliability, per question

We ask every question three times. A system banks it only if all three runs are sound. It loses a point the moment any run is confidently wrong. Reliability asks the harder thing. Can you hand this over once and stop checking? It throws away the middle. Right two times in three counts the same as never.

Two ways to read the same runs

One system, asked every question three times. Sound on two runs. Confidently wrong on the third. Net soundness reads every answer in the grid. Reliability reads every row. The same runs earn a helpful score and the worst score there is.

The two scores on the real systems

Plot both for every system. The line marks where they would agree. That is a system that never wavered from one run to the next. Every system sits below it. The drop is what asking three times costs them. Opus answers almost as soundly as Fable 5. You can rely on far less of it, because its answers wander.

A label sits on every point, so identity never rests on colour alone.
Table view

A system's score never depends on who else is on the ladder. Both scores count a system's own runs against the fixed pool of questions. Adding or removing a system never moves anyone else's number. The difficulty tiers come from three fixed reference systems for the same reason. This is not automatic. HELM moved off its mean-win-rate aggregate because a score defined against the comparison set inverts ranks whenever that set changes.

A single run does not resolve reliability finely. Reliability is a difference of counts over a fixed pool of 42 questions, so one question changing outcome moves it by 2.4 points, or 4.8 if it swings from confidently wrong to banked. We ran one system twice. Same model, same harness, same prompt, same questions, same three repeats. The two runs came out 14.3 points apart. Question by question, 34 of 42 landed identically and the 8 that moved went both ways, 6 against 2, which chance alone reproduces about 0.29 of the time. Read a gap of this size as two systems that have not been told apart yet, not as a ranking. Widening the gap needs more questions or more repeats, and we would rather say so than imply a precision the pool does not carry.

Neither score is new. Subtracting errors from correct answers and scoring a decline at zero is the reject-option rule from Chow 1970 at unit penalty. The AA-Omniscience Index and Kalai et al. rank models this way today. What differs here is what the penalty attaches to. It attaches to a confident wrong story, not to a wrong label. Reliability differs a second way. A best-of-k benchmark asks whether a system can ever succeed. Reliability asks whether it succeeds every time.

Difficulty

Four tiers, and a wall at the top

The reference systems set the difficulty, not our sense of how hard a question looks. The strongest systems have nearly solved Easy and Medium. Brutal is where every system still fails. It is the largest tier in the pool.

Solve rate by tier

Colour steps with the solve rate. Every cell also prints the value, so colour reinforces the number rather than carrying it alone.

0%100%

Which kinds of work each system handles

You would not hand over every kind of question equally. This is the same pool, cut by the kind of analytical trap each question hides. It shows where a system is steady and where it is not. The column headers carry the question counts. A kind with few questions is a thin sample, and the page marks it as such.

0%100%
Explore data

Every question in the benchmark

All 42 questions, in the words a stakeholder would use. The page does not show the trap each one hides. The answer key stays held back so the benchmark keeps measuring. You can see how hard each has proved. That is the best any system managed across its three runs.

Reuse

Citing and reproducing

How to cite

Cite the frozen release, not the project. The numbers move when the pool or the grader changes. A citation with no release label points at whatever the page says today.


      

Everything is at the kit repo and the world dataset.

Run it yourself

The world, the questions, the answer schema and the grader source are public. The answer key for the scored pool is not.

  • Answer the questions with whatever you like, then open a pull request adding one answers.json to the kit's submissions/. CI checks the shape at once. Grading is manual and offline, on the machine that holds the answer key. The scorecard comes back on your PR. It carries pool totals, never a per-question key. Then your system joins the ladder.
  • A self-check set ships with its answer key so you can prove your harness emits gradable output before spending a run. Those questions carry no score and never did, so you give nothing away.
  • The grader source ships with the kit even though you cannot run it against the scored key. You should be able to read exactly how the grader will judge your answer.

A benchmark whose answer key is public stops measuring what it set out to measure. The first system trained on it wins, and the number stops meaning anything. Holding the key back is what keeps the ladder worth reading. There is deliberately no scoring endpoint, for the same reason. A service that grades has to hold the key. A rich scorecard returned on demand lets someone diff near-identical submissions until the key falls out of it. A human reading each pull request makes that expensive for free.

What this is built on

Two benchmarks this one takes from directly. Neither is a footnote: the first supplied the traps and the second supplied the shape.