Acuity AI Advisory
← Insights
·6 min read

Claude vs Copilot: The M365 Bake-Off — Round 3, Excel

G

Ger Perdisatt

Founder, Acuity AI Advisory

Round 3 ran Claude and Microsoft Copilot as native Excel add-ins against the same deliberately messy workbook, blind-scored by Grok. Claude won five tasks of six, averaging 8.5 to Copilot's 6.4. The finding that matters is not the score: asked for revenue by region, Copilot returned a tidy sorted table that was understated by €531,000, because the workbook spelt London three ways. On dirty data these tools come apart quietly, and a precise euro figure in a clean table tends to be trusted in a way a fluent paragraph never is.

Round 1 tested judgement on a live mailbox and Claude took it 4–1. Round 2 moved to Word, where Claude wrote the better output four tasks to two and still lost the round on speed.

Round 3 is Excel, and it is the widest margin of the series so far. Claude won five of six. Claude averaged 8.5, Copilot 6.4.

The score is not the most interesting part. What happened in the first task is.

The setup

Both tools ran natively inside Excel as add-ins, on the same workbook, through the same interface. No integration advantage either way. Claude ran on Opus 5, Copilot on GPT-5.6 Thinking, and Grok 4.5 scored the outputs blind with the tool names removed.

The workbook was synthetic and deliberately messy: a fictional consulting firm, 131 staff, 18 months of monthly data, four tabs, euro throughout. Text sat in numeric columns. One date column contained four formats. A region name was spelt three ways. Utilisation used two different scales in the same column. There was a duplicate project row, a subtotal that did not tie, and one nested formula that was genuinely hard to read.

Every defect and every correct answer was documented before either tool opened the file. Three of the six tasks therefore had a known right answer — half the round was settled by arithmetic rather than opinion.

The €531,000 table that looked finished

E1 was the simplest task in the round: show me revenue by region, highest to lowest.

The trap was that the workbook spelt London three ways. LONDON in one row, Londin in another, and London with a trailing space in a third. The trailing space was invisible on screen. This is not an exotic defect — it happens in every spreadsheet more than one person has touched.

Copilot chained case-insensitive SUMIFs across the variants, double-counted, and reported London at €3.97 million. The correct figure was €4.51 million. It left the dirty labels in place and presented the total as fact.

Claude found the double-count, quantified it, corrected the three labels at source, removed the flawed helper range, and flagged the two genuine blanks and the hard-keyed quarterly block it had not touched.

Same prompt, same file, same interface. One answer was understated by €531,000 — and it was not obviously wrong. It was a neat, sorted table. Drop it into a board pack and it goes through unchallenged, because the only reliable way to catch it is to know enough about the data to question the answer.

Then it got the same defect right

E5 asked both tools for the average monthly margin for London in the first half of 2026, excluding the month with a missing cost figure. Four separate defects sat between the question and the answer, including the same London spellings and a cost figure stored as text.

Both tools returned 22.3%. Both were right. Both navigated every trap, including the misspellings Copilot had failed to merge in E1.

Same defect, same workbook, caught in one task and missed in another. That is worth sitting with. The handling is inconsistent, and nothing in the output tells you which kind of answer you have just received.

The hardest thing in the file

E3 pointed both tools at a deliberately opaque nested formula and asked what it did and whether it was right. The formula was correct, and both said so.

Two subtler problems sat underneath it. One region was missing from the lookup table the formula depended on, so those projects silently took a fallback rate nobody had agreed. One project's region also carried a trailing space, causing its lookup to fail and quietly apply the same fallback — producing a 14.0% margin where 14.1% was intended.

Copilot found the first issue. Claude found both, quantified the error on the affected project, and proposed a fix.

Where Claude lost

E4 asked both tools to tidy the workbook: consistent formatting, fix the text-in-number columns, standardise the dates.

Claude cleaned one tab out of four. It did that one impeccably — it recorded every original string before overwriting it, flagged each change, and matched the months exactly. Then it stopped and told me it had left the other three alone.

Copilot completed all four tabs to a lower standard and won the task for it.

Asked to tidy a workbook, Claude tidied a sheet. Across three rounds this is the clearest example yet of Claude under-delivering on scope, and it cost Claude a task it would probably have won on quality.

The judgment call

E6 asked what management should worry about, explicitly excluding the obvious trend.

The workbook was built so the obvious read and the correct read pointed in opposite directions. Headline revenue was up. But London's margin fell every month of 2026, from 28.5% to 15.2%, while its revenue rose 25%. The firm was buying growth and getting almost nothing for it.

Two loud decoys sat on top: a one-off revenue spike in December and an impossible 120% utilisation figure. Both tools found the real issue and neither chased a decoy.

Claude went further. It priced the London work still in progress at a 13.3% margin, six points below completed work, and projected reported margin falling to 12–13% as that backlog converted.

What I would take from it

On clean data these tools are often close enough that the difference does not matter much.

On dirty data they come apart, and they come apart quietly. Copilot's E1 answer was wrong by more than half a million euro and looked finished. Claude's E4 work was excellent and covered a quarter of the job. Neither failure announced itself.

That changes the useful question. It is not simply which tool is better. It is whether anyone in your business is checking the output at all. When we run capability sessions, the discipline of checking the answer is consistently the thing people pick over any individual feature — eight of ten at one professional services firm. It is also the habit that AI governance exists to make routine rather than heroic.

Where the series stands: 2–1

Claude took Outlook on judgment, Copilot took Word on speed and execution, and Claude has taken Excel by the widest margin so far.

Excel makes the danger clearer than either previous round. A fluent paragraph can be challenged. A clean spreadsheet answer carrying a precise euro figure tends to be trusted. Precision looks like proof, even when it is wrong.

The winning habit is not choosing the right badge. It is knowing which answers require verification before they travel any further.

Round 4 is PowerPoint. Same method, same blind scoring, one final Microsoft 365 app.

If you are making this decision for an organisation rather than a spreadsheet, our independent Claude vs Copilot comparison frames it against your own workload mix, and the Copilot ROI work exists because "we have licences" and "we get value at the point of daily use" are different claims.


The M365 Bake-Off is an independent series produced by Acuity AI Advisory. Neither Microsoft nor Anthropic had any involvement in the design, execution, or scoring. Grok 4.5 (xAI) was chosen as arbiter because it has no commercial stake in either platform.

Methodology: both tools ran natively inside Excel as add-ins, on the same workbook through the same interface. Three of the six tasks had a right answer documented before either tool opened the file. The remaining outputs went to Grok 4.5 with labels removed and were scored independently, with a further run triggered where scores disagreed.

productivitycopilotmicrosoft 365m365 bake offai strategy