Stakeholder questions against a realistically messy warehouse. Graded the way you would review an analyst whose work is your responsibility.
Real data work is only worth delegating if you can trust what comes back. Models have gotten good enough that you can accept some answers and move on while other answers you have to check because they're confidently wrong.
But an agent you have to re-check on every answer hasn't done the work. It just moved it onto you. CandidBench measures what kinds of work a data agent can take off your plate, and where it still needs your eyes.
CandidBench asks the questions every data analyst has been asked before. The warehouse behind them carries the ordinary defects real warehouses carry. Every question hides at least one common analytical trap and each question is asked three times. What gets graded is the number and the agent's analysis.
A benchmark made of clean toy tables is not measuring real work. Every question runs against one virtual company's warehouse. It carries the ordinary defects a real warehouse carries, nothing labels them, and the system has to notice. It is the smallest warehouse we could build that still breaks the systems we tested.
Real warehouses grow one table at a time, from different systems and different teams. The same dimension ends up with a different key name in each one, and nobody goes back to tidy it. None of this is a seeded defect. It is what the schema looks like, so all of it is published.
A stated limit. Real warehouses run to thousands of tables. This one runs to . We report the floor rather than the ceiling, because the failures start here. of the 42 questions were answered soundly on all three runs by nobody we tested. A larger warehouse would move those numbers. It would not change what they are about.
Every figure above is a query you can run against the world in the kit.
A system banks a question only when all three runs are sound, and loses a point if any run is confidently wrong. Read the two counts, not the single number. Banked is what is safe to delegate; bluffed is where it would confidently mislead you at least once. The net of the two is a convenient sort key and a poor headline, because it cancels: on this very ladder a system with 16 banked and 17 bluffed scores -2.4, and so does one with 0 banked and 1 bluffed.
Rows are grouped into bands, not places. A row leaves its band only when its interval clears the band leader's, so two rows in one band have not been told apart by the evidence behind them.
A confident error in a regulatory filing costs more than one in a marketing dashboard. Same runs, different trust. It is the one number on this page you set, not us.
One confident error cancels banked questions
Plot both for every system. The dashed line marks where they would agree. That is a system that never wavered from one run to the next. Every system sits below it. The drop is what asking three times costs them. Opus answers almost as soundly as Fable 5. You can rely on far less of it, because its answers wander. The two solid rules sit at zero on each axis. Zero is the point where answering beats staying silent, and it comes from the score definition rather than from our judgement. Nothing on this chart marks the point where a system is good enough to hand work over. We do not know where that point is. The leader clears zero on both axes and still bluffs 8 of 42 questions. Where the bar belongs depends on what one confident error costs you, which is yours to set and not ours. The slider on the ladder prices it.
What the interval covers. On this pool 3 of 15 ranked pairs actually separate, which gives 2 bands. Numbering the rows 1, 2, 3 would claim otherwise. The interval under each reliability score covers one source only: which questions are in the pool. It answers “how would this system do on questions like these?”, and against the fixed pool published here the point estimate is exact. It does not cover re-running the same system on the same questions, which is a second and comparable source. Two runs of one system came out 14.3 points apart. We report that below as the measurement it is, rather than folding it into the bar. With three repeats per question, no model of it reproduced the observed spread without shifting the centre off the score. Publishing a modelled bar we could not defend is the exact failure this benchmark grades others for.
A single run does not resolve reliability finely. Reliability is a difference of counts over a fixed pool of 42 questions, so one question changing outcome moves it by 2.4 points, or 4.8 if it swings from confidently wrong to banked. We ran one system twice. Same model, same harness, same prompt, same questions, same three repeats. The two runs came out 14.3 points apart. Question by question, 34 of 42 landed identically, and the 8 that moved went both ways. Read a gap of this size as two systems that have not been told apart yet, not as a ranking. Widening the gap needs more questions or more repeats, and we would rather say so than imply a precision the pool does not carry.
We grade trust and price separately. We never combine them. Flash on the Pi harness is the cheapest system here per usable question, and it ranks last on trust. One number holding both would say nothing. A system that banks no questions gets no efficiency grade. You cannot price work it never delivered.
What the headline costs. Reliability against mean cost per question. The rule at zero is where a system banks as many questions as it bluffs. Above it you gain more than you lose. Cost uses a log scale. The systems span three orders of magnitude. There is no matching line on the cost axis, because we have no price for a person answering the same question.
Every value in this chart is in the grade table above.
An expert builds trust by reading the SQL. But a correct query can still carry a plain-English explanation that contradicts it. The user reads that story and acts on it. So this benchmark grades the analysis the agent shows the user. Every answer lands in one of seven Grade classes. Two of them break trust. An agent that flags the trap or asks first earns credit for doing its job.
Every answer from every system, by kind. The grey band is work the system declined. You got nothing from it. You also have nothing to unpick. These shares count each answer. The ladder above averages by question, so the two differ slightly.
The closest prior work grades the value alone. Every question has a fixed answer. A key matches it exactly, with no judge in the loop. That reproduces perfectly, which is why the approach appeals. It also cannot tell a confidently wrong agent from a calibrated one.
How the judge itself is checked, and what that does not cover. The judge is scored against a frozen set of 78 constructed answers across 12 tasks whose correct grade is known because we wrote them, the same way each hazard in the world is planted on purpose. 28 of them are controls that differ from a passing answer by a single clause: delete the sentence naming the trap and the grade must fall, delete a false claim and it must rise, and a hedge that names nothing must not pass at all. It scored 77 of 78, and every control held. No human grading round has been run, and this benchmark does not claim human-validated grading. What a constructed answer cannot settle is whether a caveat a real model wrote under real ambiguity names the hazard well enough to earn your trust. That residual is stated rather than measured.
Both designs are defensible. They answer different questions. This one is worth its judge only if you care whether a system tells you when it is unsure. A recent data-agent benchmark that value-matches by substring says the limit outright: its grader passes “an agent that returns correct values alongside incorrect ones.”
A data agent has to pass two tests. Are the answers good? And are they good every single time? A system can pass one and fail the other. A confident wrong answer counts against both.
Take every answer a system gives. Count the sound ones. Subtract the confidently wrong ones. That is net soundness. It reads every answer, not every question. It asks one thing. When this agent answers, does it help you more often than it misleads you? It does not care whether the same question gets the same answer twice.
We ask every question three times. A question is safe to delegate only when all three runs are sound. One confidently wrong run makes it actively misleading. Reliability is the first count minus the second. Reliability asks the harder thing. Can you delegate this once and stop checking? Everything in between is the grey band on the ladder, and you review those yourself. Right two times in three counts the same as never.
One system, asked every question three times. Sound on two runs. Confidently wrong on the third. The top bar reads every answer. The bottom bar reads every question, and collapses each group of three into the one question it decides. The same runs earn a helpful score and the worst score there is.
A system's score never depends on who else is on the ladder. Both scores count a system's own runs against a fixed pool, so adding or removing a system never moves anyone else's number. HELM moved off its mean-win-rate aggregate because a score defined against the comparison set inverts ranks whenever that set changes.
Neither score is new. Subtracting errors from correct answers and scoring a decline at zero is the reject-option rule from Chow 1970 at unit penalty. Franc et al. give the modern treatment. The AA-Omniscience Index and Kalai et al. rank models this way today. The penalty here attaches to a confident wrong story, not to a wrong label. A best-of-k benchmark asks whether a system can ever succeed. Reliability asks whether it succeeds every time.
The reference systems set the difficulty, not our sense of how hard a question looks. The strongest systems have nearly solved Easy and Medium. Brutal is where every system still fails. It is the largest tier in the pool.
Colour steps with the solve rate. Every cell also prints the value, so colour reinforces the number rather than carrying it alone.
You would not delegate every kind of question equally. This is the same pool, cut by the kind of analytical trap each question hides. It shows where a system is steady and where it is not. The column headers carry the question counts. A kind with few questions is a thin sample, and the page marks it as such.
All 42 questions, in the words a stakeholder would use. The page does not show the trap each one hides. The answer key stays held back so the benchmark keeps measuring. What it does show is every run: one dot per repetition, three per system per question.
Questions are sorted by how many systems could bank them, so the table reads as a gradient. The top is work most of the field can take off your hands. The bottom is the wall. Pick a system to see its questions on their own, sorted by what it did.
Cite the frozen release, not the project. The numbers move when the pool or the grader changes. A citation with no release label points at whatever the page says today.
Everything is at the kit repo and the world dataset.
The world, the questions, the answer schema and the grader source are public. The answer key for the scored pool is not.
answers.json to the kit's submissions/. CI checks the shape at once. Grading is manual and offline, on the machine that holds the answer key. The scorecard comes back on your PR. It carries pool totals, never a per-question key. Then your system joins the ladder.A benchmark whose answer key is public stops measuring what it set out to measure. The first system trained on it wins, and the number stops meaning anything. Holding the key back is what keeps the ladder worth reading. There is deliberately no scoring endpoint, for the same reason. A service that grades has to hold the key. A rich scorecard returned on demand lets someone diff near-identical submissions until the key falls out of it. A human reading each pull request makes that expensive for free.
Two benchmarks this one takes from directly. Neither is a footnote: the first supplied the traps and the second supplied the shape.