Five messy real workbooks (close-pack, AP ageing, sales commission, headcount plan, FP&A consolidation) tested against Claude, GPT-5, and Gemini for formula-error detection, broken references, and tab-to-tab consistency. Claude wins on long-context multi-sheet reasoning. Gemini wins on cell-level formula checks.

Somewhere in every board pack there is a number nobody can trace. It came from a consolidation tab, which took it from an entity tab, which took it from a formula somebody dragged one row too far in March. The pack was reviewed, initialled and presented, and the number is still wrong. Every head of finance we have worked with has a version of this story.

So we ran a field test. Five real finance workbooks, anonymised and messy in the way real ones are: a close pack, an AP ageing, a sales commission model, a headcount plan and an FP&A consolidation. We gave each to Claude, GPT-5 and Gemini and asked for what a controller asks of a second pair of eyes: find the formula errors, find the broken references, and tell me where the tabs disagree.

There is no single winner, and that is the useful finding. Claude wins on long-context reasoning across many sheets. Gemini wins on cell-level formula checks. GPT-5 took neither. Which one a finance team should use depends on the workbook in front of it.

Which LLM is best at finding spreadsheet errors?

It depends on which of the three error types you are hunting. A formula error lives in a single cell: a wrong range, a relative reference where an absolute one was needed, a rounding applied to the wrong operand. A broken reference is a cell pointing at something that has moved or gone. A tab-to-tab inconsistency is not in any cell at all; it is a disagreement between two places that should agree, and finding it means holding both in mind at once.

The first two reward a model that reads each formula literally. The third rewards a model that can carry the whole workbook in context and reason across it. In our test those were different strengths, held by different models. Gemini was the better formula reader. Claude was the better workbook reader. GPT-5 was competent at both and best at neither. The split held across all five workbooks, which is why we trust it.

What did the five workbooks test?

We chose the five to stress different things, because one shape of workbook only tells you about one shape of workbook. None was tidied up first: hard-coded overrides, a tab called final beside a tab called final revised, links to files nobody could locate, all left in, because that is what the model meets in production. Each model saw the same prompt, and we compared its findings against the errors we had already found by hand.

  • Close pack. Many tabs rolling into a summary, subtotals feeding subtotals, and at least one value typed over a formula. Stresses tab-to-tab consistency: does the summary still agree with the tabs beneath it?
  • AP ageing. Lookups against a vendor list and bucket formulas across the ageing columns. Stresses broken references: the lookup range shortened when a vendor row was deleted, and the formula that quietly returned nothing.
  • Sales commission. Nested conditions, tiered rates and rounding, applied down a long list of reps. Stresses cell-level formula errors: the absolute reference that should have been relative, the formula dragged past the end of the list.
  • Headcount plan. A wide time axis, start dates, pro-rata months and a link to a salary band tab. Stresses consistency along a row, where one month was hand-edited and the rest were not.
  • FP&A consolidation. Entity tabs in more than one currency feeding a group tab, intercompany eliminations and links to other workbooks. Stresses multi-sheet reasoning and broken references at once, the hardest of the five.

Why does Claude win on multi-sheet reasoning?

Because a tab-to-tab inconsistency cannot be found one cell at a time. The consolidation is the clearest case. A group total was wrong not because any formula in it was wrong, but because an entity tab had been restructured and its subtotal row had moved, so the group tab was summing the row above. Every formula was valid. Every reference resolved. The error only exists if you hold both tabs together and notice they no longer describe the same thing.

Claude did that more reliably than the other two, on the close pack and the headcount plan as well as the consolidation. It followed a number from the summary back to its source, noticed when a tab had been left out of a roll-up, and caught the value typed over a formula because it no longer matched the neighbouring tab. When it found something, it explained the chain: this cell, fed by that tab, which disagrees with this other tab. A list of cell addresses with no chain is a list of things to check.

Why does Gemini win on cell-level checks?

Because cell-level errors are literal, and Gemini was the more literal reader. The commission model is where it showed. A tier boundary was tested with the wrong comparison, so one band of reps was paid on the tier above. A rounding function was wrapped around the rate instead of the result. A formula had been dragged past the end of the list and was quietly evaluating a blank. None of that needs the rest of the workbook.

Gemini checked each formula against the ones around it and flagged the one that broke the pattern, which is how an experienced reviewer works down a column. It was as strong on the AP ageing lookups, where a wrong range is a wrong range on any tab. It fell short where the approach falls short: it flagged inconsistencies inside a tab and missed the ones between tabs, which no single cell reveals.

There is no best LLM for spreadsheet QA. There is a best one for each kind of error, and the workbook tells you which.

How should a finance team actually use this?

Not by picking one model. Route by workbook. The close pack and the consolidation go to the model that reads workbooks; the commission model and the ageing go to the model that reads formulas; the headcount plan, which carries both kinds of error, goes to both. A finance team is allowed to use two models; the teams that resist the idea were usually sold a single tool.

Run the deterministic checks first. A plain script finds every reference error, every external link and every hard-coded value in a formula column, the same way every time. The model is for what a script cannot express: does this tab agree with that one, does this formula do what its header says. Spending model attention on what a script would catch makes the QA step slow and noisy.

And put it in review, never in posting. The output is a list of findings with the chain of reasoning, delivered before the pack goes to the board, and a person decides what to do with each one. The model does not change a cell. That is the rule on every finance workflow we build: the system surfaces, the finance team decides, and the log records which was which.

What to do next

Start with the workbook that has hurt you most recently, because you already know what an error in it costs. Decide which kind it is: many tabs that must agree, or long columns of formulas that must be right. That decides the model, and whether a script should go first.

That decision is what the first call is for. A 20 or 45-minute call, no deck and no homework. Where the scope needs defining, the four-day Spec from €5,000 turns it into a written brief you own, with acceptance criteria for what the QA step must catch before any build is priced. The average build ships in eight weeks. The board pack is due again at month end, which is reason enough to start now.

Or skip ahead and talk through it directly