Stakeholder questions against a realistically messy warehouse. Graded the way you would review an analyst whose work is your responsibility.
Real data work is only worth delegating if you can trust what comes back. Models have gotten good enough that you can accept some answers and move on. Other answers you have to check because they're confidently wrong.
An agent you re-check on every answer hasn't done the work. It just moved it onto you. CandidBench measures what kinds of work a data agent can take off your plate, and where it still needs your eyes.
CandidBench asks the questions every data analyst has been asked before. The warehouse behind them carries the ordinary defects real data carries. Every question hides at least one common analytical trap. Each question is asked three times. What gets graded is the number and the agent's analysis.
A benchmark made of clean toy tables is not measuring real work. Every question in CandidBench runs against a single virtual company's warehouse with multiple ordinary defects. Nothing flags them. The system has to notice.
The defects are real, and nothing labels them. We withhold which ones, and where, on purpose. Publishing the list would hand over the traps.
The world ships with the kit, so you can check yourself.
A system banks a question only when all three runs are sound, and loses a point if any run is confidently wrong. Read the two counts, not the single number. Banked is what you could hand over; bluffed is where it would confidently mislead you at least once. The net of the two is a convenient sort key and a poor headline, because it cancels: on this very ladder a system with 16 banked and 17 bluffed scores -2.4, and so does one with 0 banked and 1 bluffed.
Rows are grouped into bands, not places. A row leaves its band only when its interval clears the band leader's. On this pool 3 of 15 ranked pairs actually separate. That gives 2 bands. Inside a band the evidence has not told these systems apart, and numbering them 1, 2, 3 would say otherwise. The interval under each reliability score covers one source only: which questions are in the pool. It answers “how would this system do on questions like these?”, and against the fixed pool published here the point estimate is exact. It does not cover re-running the same system on the same questions, which is a second and comparable source. Two runs of one system came out 14.3 points apart. We report that below as the measurement it is, rather than folding it into the bar. With three repeats per question, no model of it reproduced the observed spread without shifting the centre off the score. Publishing a modelled bar we could not defend is the exact failure this benchmark grades others for.
We grade trust and price separately. We never combine them. Flash on the Pi harness is the cheapest system here per usable question, and it ranks last on trust. One number holding both would say nothing. A system that banks no questions gets no efficiency grade. You cannot price work it never delivered.
What the headline costs. Reliability against mean cost per question. The good region is up and to the left. More questions answered soundly every time, for less money. Cost uses a log scale. The systems span three orders of magnitude.
Every value in this chart is in the grade table above.
An expert builds trust by reading the SQL. But a correct query can still carry a plain-English explanation that contradicts it. The user reads that story and acts on it. So this benchmark grades the analysis the agent shows the user. Every answer lands in one of seven Grade classes. Two of them break trust. An agent that flags the trap or asks first earns credit for doing its job.
Every answer from every system, by kind. The grey band is work the system declined. You got nothing from it. You also have nothing to unpick. These shares count each answer. The ladder above averages by question, so the two differ slightly.
The closest prior work grades the value alone. Every question has a fixed answer. A key matches it exactly, with no judge in the loop. That reproduces perfectly, which is why the approach appeals. It also cannot tell a confidently wrong agent from a calibrated one.
How the judge itself is checked, and what that does not cover. The judge is scored against a frozen set of 78 constructed answers across 12 tasks whose correct grade is known because we wrote them, the same way each hazard in the world is planted on purpose. 28 of them are controls that differ from a passing answer by a single clause: delete the sentence naming the trap and the grade must fall, delete a false claim and it must rise, and a hedge that names nothing must not pass at all. It scored 77 of 78, and every control held. No human grading round has been run, and this benchmark does not claim human-validated grading. What a constructed answer cannot settle is whether a caveat a real model wrote under real ambiguity names the hazard well enough to earn your trust. That residual is stated rather than measured.
Both designs are defensible. They answer different questions. This one is worth its judge only if you care whether a system tells you when it is unsure. A recent data-agent benchmark that value-matches by substring says the limit outright: its grader passes “an agent that returns correct values alongside incorrect ones.”
A data agent has to clear two bars. A system can clear one and fail the other. When it answers, does it help you or mislead you? Can you hand a question over and stop checking? CandidBench scores both.
Take every answer a system gives. Count the sound ones. Subtract the confidently wrong ones. That is net soundness. It asks one thing. When this agent answers, does it help you more often than it misleads you? It does not care whether the same question gets the same answer twice.
We ask every question three times. A system banks it only if all three runs are sound. It loses a point the moment any run is confidently wrong. Reliability asks the harder thing. Can you hand this over once and stop checking? It throws away the middle. Right two times in three counts the same as never.
One system, asked every question three times. Sound on two runs. Confidently wrong on the third. Net soundness reads every answer in the grid. Reliability reads every row. The same runs earn a helpful score and the worst score there is.
Plot both for every system. The line marks where they would agree. That is a system that never wavered from one run to the next. Every system sits below it. The drop is what asking three times costs them. Opus answers almost as soundly as Fable 5. You can rely on far less of it, because its answers wander.
A system's score never depends on who else is on the ladder. Both scores count a system's own runs against the fixed pool of questions. Adding or removing a system never moves anyone else's number. The difficulty tiers come from three fixed reference systems for the same reason. This is not automatic. HELM moved off its mean-win-rate aggregate because a score defined against the comparison set inverts ranks whenever that set changes.
A single run does not resolve reliability finely. Reliability is a difference of counts over a fixed pool of 42 questions, so one question changing outcome moves it by 2.4 points, or 4.8 if it swings from confidently wrong to banked. We ran one system twice. Same model, same harness, same prompt, same questions, same three repeats. The two runs came out 14.3 points apart. Question by question, 34 of 42 landed identically and the 8 that moved went both ways, 6 against 2, which chance alone reproduces about 0.29 of the time. Read a gap of this size as two systems that have not been told apart yet, not as a ranking. Widening the gap needs more questions or more repeats, and we would rather say so than imply a precision the pool does not carry.
Neither score is new. Subtracting errors from correct answers and scoring a decline at zero is the reject-option rule from Chow 1970 at unit penalty. The AA-Omniscience Index and Kalai et al. rank models this way today. What differs here is what the penalty attaches to. It attaches to a confident wrong story, not to a wrong label. Reliability differs a second way. A best-of-k benchmark asks whether a system can ever succeed. Reliability asks whether it succeeds every time.
The reference systems set the difficulty, not our sense of how hard a question looks. The strongest systems have nearly solved Easy and Medium. Brutal is where every system still fails. It is the largest tier in the pool.
Colour steps with the solve rate. Every cell also prints the value, so colour reinforces the number rather than carrying it alone.
You would not hand over every kind of question equally. This is the same pool, cut by the kind of analytical trap each question hides. It shows where a system is steady and where it is not. The column headers carry the question counts. A kind with few questions is a thin sample, and the page marks it as such.
All 42 questions, in the words a stakeholder would use. The page does not show the trap each one hides. The answer key stays held back so the benchmark keeps measuring. You can see how hard each has proved. That is the best any system managed across its three runs.
Cite the frozen release, not the project. The numbers move when the pool or the grader changes. A citation with no release label points at whatever the page says today.
Everything is at the kit repo and the world dataset.
The world, the questions, the answer schema and the grader source are public. The answer key for the scored pool is not.
answers.json to the kit's submissions/. CI checks the shape at once. Grading is manual and offline, on the machine that holds the answer key. The scorecard comes back on your PR. It carries pool totals, never a per-question key. Then your system joins the ladder.A benchmark whose answer key is public stops measuring what it set out to measure. The first system trained on it wins, and the number stops meaning anything. Holding the key back is what keeps the ladder worth reading. There is deliberately no scoring endpoint, for the same reason. A service that grades has to hold the key. A rich scorecard returned on demand lets someone diff near-identical submissions until the key falls out of it. A human reading each pull request makes that expensive for free.
Two benchmarks this one takes from directly. Neither is a footnote: the first supplied the traps and the second supplied the shape.