Copilot, Claude or Gemini for Excel models : what did AI models pick?

AI models from Anthropic, OpenAI and Google picked Microsoft's Copilot to build Excel models and Anthropic's Claude for Excel to check existing ones for errors, leaving Google's Gemini to Sheets work. It is a workflow choice, not a proven accuracy ranking. Any AI-written formula still needs tests you can check yourself.

Business & Economy · 2026-09-27

A finance manager at a 40-person company builds budget and cash-flow models in Excel. Every quarter they present a board pack, the set of financial reports the board of directors reviews. They want AI help with two jobs. The first is building new models. The second matters more : auditing existing workbooks, which means checking them for broken references and silent errors. A broken reference is a formula that points to cells that were moved or deleted. A silent error is a formula that returns a perfectly plausible number that happens to be wrong. Three AI tools keep coming up, each from a different company. Claude for Excel, from Anthropic, is an add-in, a small program that runs inside Excel, and it comes with paid Claude plans. Microsoft 365 Copilot, from Microsoft, also works inside Excel. Gemini, from Google, works inside Google Sheets, Google's online spreadsheet. The company runs on Microsoft 365, Microsoft's subscription bundle of Excel, Word, Outlook and its other office apps, but a few colleagues work in Sheets.

Polora put that question to AI models built by Anthropic, OpenAI and Google. Three specialist models each gave a recommendation. A researcher model checked their prices and claims against published sources, and a moderator model gave the final answer. What follows is their comparison, including where their numbers were wrong and where the public evidence runs out.

The AI models agreed on a division of labor : Copilot to build, Claude for Excel to audit

The three specialist AI models came to the same division of labor. The Enterprise Ecosystem Analyst, an OpenAI model, argued that Copilot should do the building. Its reason was not that Copilot invents better finance logic. It was that Copilot works inside the environment where the model already lives, along with its permissions, shared files and board-pack process. The question asked about Copilot's Agent Mode, but that name may not be the one you see. The Analyst pointed out that Microsoft now calls it Edit with Copilot, and Microsoft's current documentation describes three ways of working in Excel : Edit, Plan and Chat. The moderator model advised checking which names and features appear in your own company's setup rather than relying on a label. The Analyst advised using Plan before allowing any direct edits. The Spreadsheet Engineering Specialist, an Anthropic model, called that habit of planning before touching cells the most important safety property when a tool generates a new model.

For auditing, all three chose Claude for Excel, including the Risk and Audit Strategist, which is a Google model. The deciding feature was that Claude points to specific cells. Anthropic says the add-in can explain a workbook, track its changes and take you to the cells it refers to. The Spreadsheet Engineering Specialist put the argument bluntly : an audit finding you cannot trace to a cell reference is not an audit finding, it is a rumor.

The moderator model drew the line the reader most needs. This is a workflow recommendation, not evidence that Claude catches more silent errors than Copilot. The researcher model added that when three models reason from overlapping public material, their agreement is notable but is not proof, because they can share a blind spot.

Why the AI models kept Gemini away from the Excel model the board relies on

None of the AI models called Gemini weak. Their objection was about structure. Google says Gemini works best in native Google Sheets and asks users to save an Excel file as a Sheets file to use its features. For a multi-sheet workbook full of links between tabs, that conversion creates a second version of the company's financial numbers. The Spreadsheet Engineering Specialist warned that converting can itself introduce the kind of error the tool is supposed to find, such as broken references, formulas that no longer fill a block of cells the way they used to, and changed rounding.

The Enterprise Ecosystem Analyst also noted limits in Gemini's AI formula, which you type into a cell as =AI() to ask Gemini something and get its answer back in that cell. It returns text only, cannot see the whole workbook and cannot be placed inside other formulas. The Risk and Audit Strategist claimed that Google had launched a specific error-fixing feature in June 2026, but the researcher model could not confirm the feature or the date. The models' shared advice was that Gemini makes sense for colleagues whose work already lives in Sheets and is included in their plan, but not as the auditor of the model the board relies on.

What Copilot, Claude for Excel and Gemini would cost a 40-person company

Two of the specialist AI models got prices wrong, and the researcher model corrected them. For a company this size, Microsoft 365 Copilot Business lists at $21 per user per month billed yearly, $18 under a promotional offer, or $25.20 month to month, for organizations of up to 300 users. The $30 figure one model quoted is the enterprise price. Copilot is never sold on its own. It requires an eligible Microsoft 365 base plan, such as Microsoft 365 Business Standard at $14 per user per month, so the add-on price understates the real cost. Public sources disagree on when the promotion ends. Microsoft's own page says December 31, 2026, and limits the $18 rate to a yearly commitment in the first year only. Other sources say September 30, which is days away. The researcher suggested checking the quote in the Microsoft 365 admin center, the console where a company's IT administrator manages accounts and licenses.

Adding it up, the Enterprise Ecosystem Analyst worked out that Copilot for all 40 people would come to about $720 a month at the $18 yearly rate, or $1,008 a month month to month, before the base plan. It suggested starting instead with five to eight seats, meaning individual user licenses, for the people who actually build and present the numbers.

Claude for Excel has no separate price. It comes with paid Claude plans : Pro at $20 a month, Max at $100 and up, and Team. The moderator read the current Team listing as $20 per seat per month billed yearly or $25 monthly, with a two-seat minimum, lower than the Analyst's figures of $25 and $30 with a five-seat minimum. Running Claude in Excel requires a Microsoft 365 work account but no Copilot license.

Gemini in Sheets starts with Google Workspace Business Standard at $14 per user per month billed yearly. Because the AI is bundled, everyone on the plan pays for it whether they use it or not. The Spreadsheet Engineering Specialist added a warning about Copilot. On top of the seat price, usage fees can be charged each time the AI carries out certain tasks, and several such charges can run at once, so it advised setting spending limits on the first day.

Copilot for all 40 people, per month, before the base plan · $18 yearly rate · month to month · $720 · $1,008
Copilot for all 40 people, per month, before the base plan · $18 yearly rate · month to month · $720 · $1,008

A $20 Claude plan may run out halfway through auditing a large workbook

The Spreadsheet Engineering Specialist, one of the AI models in the comparison, called usage limits its single strongest practical warning. Users have reported hitting the Pro plan limit within minutes, after only a few formatting tasks and one error check, with the work left half done. The add-in uses up the allowance much faster than chat, because each request reads the workbook, reasons about it and takes several separate actions. Auditing a large linked workbook is about the most demanding thing you can ask of it.

A separate review the researcher model found agrees. Pro suits daily analytical work, while the $100 Max plan gives much more room for heavy financial modeling. The Specialist framed it as a choice : pay for a higher tier such as Max if the board pack is to be audited in one go, or accept auditing one tab per session, rather than buying the $20 plan by reflex. It noted that the complaints date from March 2026 and the limits may have changed since, so the first week of real use is the test.

Where public reviews of Claude and Copilot in Excel disagree

The question asked for disagreements to be shown, not settled, and the AI models found several. One review says most users who tried both prefer Claude to Copilot, but that is self-reported opinion, and the same review warned of installation and reliability problems as of March 2026. One consultant's verdict, checked in August 2026, calls Claude the better tool for analysis, models and digging through inherited workbooks, and says Copilot keeps the edge for quick tasks from the Excel menu bar. That is one experienced person's view.

Microsoft reported gains in engagement, retention and satisfaction for Excel during its preview, the trial period before full release. The Spreadsheet Engineering Specialist pointed out that these measure how much people used the product, not whether its formulas were correct. Anthropic has cited a result for Claude on a finance benchmark, a standardized test used to compare AI models, and Microsoft reports improvements in reliability. Neither claim, the Enterprise Ecosystem Analyst said, shows how either tool will handle your own hidden tabs, links to other files and conventions. The researcher model's conclusion was plain. It found no independent, audited benchmark comparing these three tools on financial-model auditing. Any claim that one is measurably better at it goes beyond the evidence.

Is it worth running two AI tools on the same Excel model? The AI models said yes

The Spreadsheet Engineering Specialist, an Anthropic model, based its case on separation of duties, the accounting rule that the person who records a transaction in the books does not also approve it. A model checking its own work carries its own blind spots. If Copilot misreads the assumptions tab while writing a formula and then checks that same formula, the mistake passes through both steps and looks like confirmation. A reviewer built by a different company, with different training, fails in different ways. In the Specialist's view, agreement between the two is weak evidence, and disagreement points to exactly the cell a person should look at.

The moderator model narrowed this down. It suggested giving a few finance users Copilot for drafting and Claude for review, each working on a separate copy of the workbook and never both editing the approved model. Its test was a short pilot scored against known past mistakes and deliberately planted errors : material errors found, false alarms, unintended changes, and whether each finding names an exact cell and its business consequence. It advised adding licenses only if the second tool earns its cost. The Enterprise Ecosystem Analyst warned against asking either tool to fix everything at once, because that produces a clean-looking workbook whose logic has changed in dozens of places. The moderator's own routine points the same way : one proposed change at a time.

Before installing Claude for Excel : IT approval and a possible gap in the record of AI changes

An organization-wide rollout goes through the Microsoft 365 admin center, where an IT administrator deploys the add-in to everyone, to named users or to groups. The Spreadsheet Engineering Specialist, one of the AI models, noted that the administrator must also approve Anthropic as an outside company allowed to handle company data. Because a cash-flow model contains figures from which payroll can be inferred, it said that conversation should happen before installation. The add-in needs recent Office apps. Office 2016 and 2019 are not supported.

The researcher model raised a gap the others had missed. A review from March 2026 said the add-in did not yet keep the kind of detailed log auditors expect. If that is still true and an external accountant or the board asks what the AI changed and when, a manually kept change log may be the only record. The researcher said to confirm the current status, since the review is six months old.

Check an AI-written Excel formula by saying it in words and seeing what it points to

All of the AI models agreed on how a non-programmer should approve a formula, and none of the steps depends on which AI wrote it. Work on a copy, one proposed change at a time. First, ask the AI not to edit anything, and to propose the formula, explain every reference in plain English and suggest test cases. Then restate the rule in your own words, for example that ending cash equals opening cash plus receipts minus payments. If your sentence and the AI's explanation differ, stop there.

Next, click the cell and press the F2 key, which opens the formula for editing. Excel then outlines each cell or range the formula uses in its own color, so you are not reading code but checking whether the colored boxes cover the right sheet, the right period and every row. The Spreadsheet Engineering Specialist called this the most productive check of all, because it catches a range that stops one row short, or a shifted column, in seconds. Trace Precedents and Trace Dependents show with arrows what feeds the cell and what it feeds. Evaluate Formula steps through the calculation and shows the value at each stage, so you can see where the number stops making sense.

Then copy the formula one month to the right and one row down, and look again. Many formulas work in one cell and break when copied. Check the first forecast month, a year-end, the start of the next fiscal year and the last month, not just a comfortable month in the middle.

※ Trace Precedents : an Excel command on the Formulas tab that draws arrows from a cell to the cells its formula uses.

Test an AI-written formula with answers you can work out yourself

Change the inputs in the copy and watch what happens. The AI models suggested three cases : a simple example you can do in your head, a zero or boundary case such as zero revenue or the first forecast month, and a stress case such as payroll up 10 percent or collections delayed 30 days. A discount formula given a zero rate must return the undiscounted amount. The Risk and Audit Strategist, a Google model, added a quick test. If a result does not move when you enter an extreme input, the AI has probably typed in a fixed number where a link should be.

Then compare against a calculation done separately, with a calculator or a simple scratch sheet, and not with another AI explanation. Units times price should equal revenue. Headcount times monthly cost should equal payroll. Opening cash plus net movement should equal closing cash in every period.

Finally, look for errors that never show #REF!, the message Excel displays when a reference is broken. These include numbers typed inside formula rows, totals whose SUM range, the block of cells being added up, stops before the last line, reversed signs on cash outflows, links to last year instead of last month, and the one the Spreadsheet Engineering Specialist singled out as the most dangerous in AI-generated finance models : IFERROR wrapped around a formula so that real breakage shows up as a tidy zero. It advised searching the workbook for every IFERROR and justifying each one.

※ IFERROR : an Excel function that shows a chosen value, often zero or a blank, whenever a formula would otherwise show an error.

What a board can rely on is tested logic, not the AI's confidence

The last step in the AI models' routine is to reconcile the whole model. Check that each period's closing cash carries over as the next period's opening cash, that totals which must match actually match, and that the numbers in the board pack agree with the schedules underneath. The Spreadsheet Engineering Specialist suggested permanent check rows at the top of each sheet that show zero when everything agrees and turn red when something breaks, a control that keeps working next quarter when nobody is watching. Save a version before the AI touches anything, compare outputs afterward, and explain every material change. Record each change in a simple log with the cell, the old and new formula, the reason, the tests run and who approved it.

The Specialist's rule sums up the rest. Both tools explain a wrong formula in the same fluent, assured tone they use for a right one, so the AI's confidence counts as no evidence at all. The moderator model ended in the same place. The AI can speed up drafting and point to suspicious cells, but for a model the board relies on, the sign-off is business logic someone has tested and reconciled. The Specialist's practical first step was to have the chosen reviewer audit a model that has already been signed off, where the right answers are known, to see what it catches, what it misses and what it invents.

Copilot, Claude or Gemini for Excel models : what did AI models pick?Copilot, Claude or Gemini for Excel models : what did AI models pick?A finance manager at a 40-person company wants AI help to build Excel models and, above all, to audit existing ones for silent errors. AI models from Anthropic, OpenAI and Google compared three tools.The AI models agreed on a division of labor : Copilot to build, Claude for Excel to auditWhy the AI models kept Gemini away from the Excel model the board relies onWhat Copilot, Claude for Excel and Gemini would cost a 40-person company · $18 yearly rate month to month $720 $1,008 Copilot for all 40 people, per month, before the base plan · seat : an individual user licenseWhere public reviews of Claude and Copilot in Excel disagree · benchmark : a standardized test used to compare AI modelsCheck an AI-written Excel formula by saying it in words and seeing what it points toThe last step in the AI models' routine is to reconcile the whole model.Sources 16 : anthropic.com · Anthropic : Advancing Claude for financial services, support.microsoft.com · Microsoft Support : Frequently asked questions about Copilot in Excel, microsoft.com · Microsoft 365 Copilot + 13

I'm the finance manager at a 40-person company. I build and maintain our budget and cash-flow models in Excel: multi-sheet workbooks with a lot of linked formulas, and a board pack every quarter. I want an AI tool that can help me build new models and, more importantly, audit existing ones for broken references and silent errors. The candidates I keep seeing are Claude for Excel (the add-in on paid Claude plans), Microsoft 365 Copilot in Excel with Agent Mode, and Gemini in Google Sheets. We are a Microsoft 365 shop, but a few colleagues work in Google Sheets. Compare these three for this specific job using current published prices and features. Say which one you would pick for building models and which for auditing them, whether running two of them side by side makes sense, and exactly how a non-programmer should check an AI-written formula before it goes into a model the board relies on. Where public benchmarks or reviews disagree, say so instead of picking a side.

Risk & Audit Strategist · gemini-3.1-pro

As a Risk & Audit Strategist, my primary concern is the integrity of your board pack. In financial modeling, a "silent error"—where a formula computes a mathematically valid but logically incorrect number—is the most dangerous risk you face. When evaluating AI tools, you cannot just look at their ability to generate; you must ruthlessly assess their ability to be governed, traced, and verified. Here is how the 2026 landscape breaks down for your specific needs, based on current pricing, architecture, and public consensus. The Candidates: Pricing & Architectural Reality 1. Claude for Excel (Anthropic) * Price: Requires a Claude Pro ($20/month) or Max ($100+/month) subscription. It is not a standalone purchase. * Audit Profile: Claude for Excel (which reached General Availability in May 2026) operates with a highly conservative, read-first architecture. It reads your actual workbook structure—cells, formulas, and tabs. When it explains a model, it cites the exact cells it is referencing, and it warns you before overwriting anything. * Consensus: Public reviews highlight that its cautious, "co-pilot rather than autonomous agent" approach makes it excepti…

Enterprise Ecosystem Analyst · gpt-5.6-terra

My recommendation For a 40-person Microsoft 365 company whose authoritative budget, cash-flow model, and quarterly board pack live in Excel: - Pick Microsoft 365 Copilot in Excel as the primary tool for building and maintaining models. - Use Claude for Excel as the specialist second opinion—especially for formula explanations, model-logic review, and difficult repairs. - Do not make Gemini in Google Sheets a core finance-model tool unless a model is genuinely owned in Sheets. It is useful for the colleagues already working there, but converting an .xlsx file to native Sheets to use the richest Gemini features is an unacceptable fork in the control environment for a board-critical Excel model. Google explicitly says Gemini works best in native Sheets and asks users to save an Excel file as Google Sheets to use its features. (support.google.com) The key caveat: none of the three is an audit tool in the accounting/control sense. They can identify suspicious formulas, explain dependencies, propose repairs, and run requested checks. They do not independently establish that your model i…

Spreadsheet Engineering Specialist · claude-opus-5

The short version Build with Copilot Agent Mode. Audit with Claude for Excel. Skip Gemini in Sheets for this job. Running the first two side by side is not redundancy — it is the whole point, and it costs you roughly one extra seat. Let me justify each piece, then give you the formula-checking protocol, which is honestly the most valuable part of this answer. --- What each thing actually costs you, today Claude for Excel — a real Office add-in, not a chat window. Anthropic announced general availability for the Excel, PowerPoint, and Word add-ins on May 7, 2026, on paid Claude plans (Pro, Max, Team, Enterprise). Cheapest entry is Pro at $20/mo, installed from the Microsoft Marketplace. Critically for a Microsoft shop: you need a Microsoft 365 subscription with a work account to host the add-in, but you do not need the separate Microsoft 365 Copilot per-seat license. For org-wide deployment your M365 admin must approve Anthropic as a processor in the admin centre. Copilot Agent Mode in Excel — went GA on April 22, 2026 in Word, Excel and PowerPoint. The licensing is where people get burned. Any Copilot Chat-eligible M365 business subscription gets a standard-…

Researcher · claude-opus-5

Fact-check: what's confirmed, what's off, what nobody can source I checked the pricing and availability claims all three debaters leaned on. Most of the licensing picture holds up; a few specific numbers are wrong or stale, and one entire category of claim — the "who's better at auditing" rankings — has no verifiable backing from anyone. --- ✅ Confirmed Claude for Excel is GA and included in any paid Claude plan. Anthropic announced general availability for the Excel, PowerPoint, and Word add-ins on May 7, 2026; the Outlook add-in is in public beta; all four are available on paid Claude plans (Pro, Max, Team, and Enterprise). Anthropic's own listing confirms it: Claude for Excel is available to customers on the Claude Pro, Max, Team and Enterprise plans. The date given by the Spreadsheet Engineering Specialist and the Risk Strategist (May 2026 GA) is correct. → https://inventivehq.com/blog/claude-microsoft-365-add-ins · https://marketplace.microsoft.com/en-us/product/saas/wa200009404 No Copilot licence needed to run Claude in Excel — this is the load-bearing fact for you. You need a Microsoft 365 subscription with a work or school account to host the add-in, but y…

The Arbiter · gpt-6-sol

Round record All three specialists recommend Copilot in Excel for building, Claude for Excel as a second reviewer of existing models, and Gemini for workbooks that genuinely live in Google Sheets—not for converting the board model to Sheets. That is a shared recommendation, not proof that Claude detects more silent errors. As the Researcher found, no independent, audited head-to-head benchmark of these three on financial-model auditing was established. The Spreadsheet Engineering Specialist emphasized Claude’s cell-specific explanations and warned that a large-workbook audit may exhaust a Claude Pro allowance. The Enterprise Ecosystem Analyst emphasized keeping the authoritative model in the existing Microsoft 365 environment, using a sandbox copy, and recording reviewed changes. The Risk & Audit Strategist emphasized behavioral tests and an independently calculated comparison. All three agreed that an AI finding requires human verification; agreement between two AI tools is not a sign-off. Pricing needs correction before the verdict. For a 40-person company, Microsoft currently displays Copilot Business at $21/user/month list or $18 promotional, paid yearl…

The Arbiter's Final Verdict · gpt-6-sol

My pick Build in Copilot for Excel; pilot Claude for Excel as a separate audit reviewer. Keep the board model in Excel. Use Gemini for the colleagues’ Sheets-native work, not as a reason to convert a linked, multi-sheet Excel model. That is a workflow recommendation, not a proven accuracy ranking. I would not claim Claude catches more silent errors than Copilot: no independent, audited head-to-head benchmark of these three on financial-model auditing was established. Practitioner reviews favor different tasks, but they do not settle which tool misses fewer material errors in *your* workbooks. | Tool | Fit for your job | Published price and purchasing caveat | |---|---|---| | Microsoft 365 Copilot in Excel | Best first choice for building in your existing Excel workflow: draft and revise formulas, structure analyses, and work through a model in place. Ask it to explain proposed edits before accepting them. Microsoft’s current documentation describes Edit, Plan, and Chat; “Agent Mode” may not be the label you see in your tenant. | Copilot Business: $21/user/month, or $18 promotional, billed yearly, plus an eligible Microsoft 365 base license. The $30…