All articles

Playbooks

Why does AI give wrong numbers when analysing your data, and how to stop it

ChatGPT and other AI tools predict numbers the way they predict words, which is why a total from an uploaded spreadsheet can be confidently wrong. What to do when you are analysing data yourself, and the structural fix for business reports.

Amit Chopra··Updated ·7 min read

A language model does not calculate numbers. It predicts them, the same way it predicts every other word, by producing what looks likely in context. A likely-looking revenue figure and a correct revenue figure are different things, and the model cannot tell you which one it just wrote. That is the whole explanation, and it is why the fix is structural, not a better prompt.

If you have uploaded a spreadsheet to ChatGPT, Claude or Gemini, asked for a total, and later found the total was wrong, you have seen this. You are not holding it wrong. You are watching the mechanism work as designed.

Why ChatGPT and other language models invent numbers

A model produces text one piece at a time, choosing what plausibly comes next. Ask it to summarise your quarter and it will write a fluent paragraph in which the growth percentage is chosen because a number like it belongs in a sentence like that. Sometimes the number is right, because the right number was in what you gave it. Sometimes it is subtly adjusted, rounded, or filled in from nowhere. The sentence reads identically in both cases.

That last part is the dangerous part. Wrong numbers from a model arrive with the same confident fluency as right ones. There is no stumble to warn you.

Why it happens when you upload a spreadsheet and ask for a total

This is the version most people meet first. You paste in a table or upload a file, ask for the total, the average or the change since last month, and the answer comes back instantly and sounds right.

What happened underneath depends on the tool. Some assistants can write and run a short piece of code to do the arithmetic, and when they do, the number comes from the code and is usually right. When they do not, the model is reading your figures as text and writing a plausible sum. With a dozen rows it will often get it right. With four hundred rows, merged cells, a subtotal line it mistook for data, or two columns with similar headings, it will produce a number that looks exactly like the right one and is not.

Four habits stop this costing you anything:

  1. Ask it to show its working. If the answer came from code, it can show the code and the rows it used. If it cannot, treat the figure as a guess.
  2. Check one number against the source yourself. Pick a row, find it in the original file, confirm it matches. One mismatch means the whole answer is suspect.
  3. Ask how many rows it counted. A wrong row count is the fastest way to catch a skipped section or a double-counted subtotal.
  4. Keep the arithmetic in the spreadsheet. Let the spreadsheet compute the total and let the model explain what the total means. That split, code for figures and the model for words, is the same rule we build every reporting system by, and the rest of this article is about it.

Why better prompting does not fix it

Telling a model "only use the numbers provided" and "double-check your figures" reduces the error rate. It does not make it zero, and business reporting needs zero, because one invented figure in front of a customer, a board, or a regulator costs more than the whole system saved. The consequences are not hypothetical: in October 2025, Deloitte partially refunded the Australian government for a A$440,000 report found to contain AI-invented citations and a fabricated court quote.

A prompt is a request. What you need is a guarantee, and guarantees do not come from asking nicely. They come from architecture.

The fix: a boundary, enforced by a test

The rule we build every reporting system by: code computes the numbers, the model only narrates or extracts, and the boundary is enforced by a test.

Concretely, in our own reporting product:

  1. Plain code pulls the data and computes every figure: totals, comparisons, percentages. This set of computed numbers is the only source of truth.
  2. The model is handed that set and asked to write the narration: what moved, what matters, what to look at.
  3. Every sentence the model writes is then checked by code. If a sentence contains a figure that is not in the computed set, the sentence is thrown away and a plain template takes its place.

The model can still be useful, because narration is genuinely what it is good at. It just can never be the source of a number. In our learning product the same rule appears in a different form: algebra is checked by a symbolic maths engine, and the model never decides what is correct.

How to check a system you are buying

Two questions expose everything:

"Where in the system is the line between what the model writes and what code computes?" A real answer points at a specific mechanism, like the whitelist check above. A vague answer ("the model is very accurate") means there is no line.

"What happens when the model writes a number anyway?" The right answer is that a check catches it and a template replaces it, automatically, every time. If the answer is that a human reviews the output, ask what happens on the day the human is busy.

We have removed models from working systems five times when the step needed to be right rather than clever. The full list is here, and the reporting case in our step-by-step automation guide shows the boundary in a complete build.

The short version

AI will get your numbers wrong occasionally, forever, because predicting text is what it is. When you are analysing data yourself, make the tool show its working and check one figure against the source. When a system is producing reports for your business, stop trying to prompt the risk away and move the numbers out of the model's hands entirely. If a vendor cannot show you where that boundary sits in their product, the boundary does not exist. Ours is a test that fails the build, and we will happily show it to you.

If this is your situation. This is the kind of work we do under Data and reporting automation.

Run your numbers

Put your own hours and your own quote through the arithmetic before you talk to anybody, including us. It runs in your browser, uses your numbers, and stores nothing.

Or go straight to a conversation about the process