Technology

ChatGPT vs Gemini vs Ed: What a Fact-Checked Money Benchmark Actually Shows

ChatGPT vs Gemini vs Ed

MoneyBench, July 2026 final round. Share of 106 money questions won. Judge: Claude (Anthropic).

The chart above is lopsided enough to make you suspicious, and you should be. In the July 2026 round of a benchmark called MoneyBench, EdWealth’s finance assistant Ed won 62.3% of 106 real money questions, with Gemini at 20.8% and ChatGPT at 17.0%. Chance would be 33.3%. My first reaction was the correct one. EdWealth built this benchmark, and companies rarely publish scoreboards where their own product loses.

So I spent a while with the method before I gave the number any weight. A few things stop it reading like a press release. The judge is Claude, made by Anthropic, which builds none of the three assistants and shares no model family with any of them, Ed’s underlying model included. Answers were stripped of branding and shown to the judge in random order, so it could not quietly favor the house team by name. Every factual claim was checked against live sources on 23 July 2026. And the losing answers are published in full, worst cases and all. That last part is the tell. Marketing hides the losses.

What it was grading

MoneyBench does not test trivia. It sends the same standalone money questions to all three assistants and grades two things: whether the answer’s numbers survive verification, and whether it actually helps a person make a decision. The 95% confidence interval on Ed’s July result runs from 52.8 to 71.7, so the lead is wide even at the pessimistic end. A wide lead on your own turf is still your own turf, though, which is why the more interesting reading is where the rivals did well.

Give the rivals their due

They did well in more places than the headline admits.

Start with the round EdWealth would rather bury. In the May opening round, Ed finished last, at 17.4%, behind Gemini at 46.1% and ChatGPT at 36.5%. Ed only took the lead after EdWealth rebuilt how it retrieves and delivers answers. The model did not get smarter between May and July. The plumbing around it changed. That is worth knowing before you read “62.3%” as some innate superiority.

The setup was not perfectly level, either. ChatGPT ran at Pro effort, above its default, while Ed and Gemini ran at default. That helps ChatGPT, and it still landed third, but a reader deserves to know the three were not configured identically.

On raw accuracy the gap nearly vanishes. The three scored 3.99, 3.72 and 3.65 out of five, close enough that I would not pick a winner on facts alone. On the quality of the writing, Ed’s edge over Gemini was not statistically significant (p=0.086), and in the Vietnamese-language subset Gemini essentially tied Ed. For plain explanation, say what an ISO is or how a Roth conversion works, ChatGPT and Gemini are genuinely strong, and both are free to try. If that is what you need, you do not need Ed.

Where Ed actually pulled ahead

One dimension, consistently: usefulness for an actual decision. Ed scored 4.24 there against 3.62 for ChatGPT and 3.47 for Gemini, and the margin held across the set, ahead on 76 of 106 questions versus ChatGPT and 75 of 106 versus Gemini. That is the widest and steadiest gap in the whole study.

Dimension (out of 5) Ed ChatGPT Gemini
Raw accuracy 3.99 3.72 3.65
Usefulness for a decision 4.24 3.62 3.47

The reason looks architectural rather than a claim that Ed is cleverer. It pulls live numbers at answer time through more than 120 tools, and it retrieved 590 of 590 sampled statutory parameters correctly out of a base of 11,566. It leads with the decision rather than a wall of background, and it keeps hard rules separate from judgment calls. When the verification pass docked every answer for anything it could not confirm, Ed lost 5.3 points and still finished first. A general assistant, answering from training data, will happily write you a confident paragraph around a number that was current last quarter.

One boundary matters here. Ed is a coach and a check-up, not a robo-advisor and not your CPA. It will not tell you to buy, sell or hold, and for the exact figure on your own return you still need a professional. On tax specifically, treat Ed as the tool that surfaces the question and a CPA as the one who fills in the number.

What I would not oversell

The July result rests on a single automated judge, and one judge is one judge. When humans graded the June round, they barely agreed with each other (Fleiss κ of 0.026, which is close to noise). EdWealth has not published the judge’s exact definition of “usefulness,” and that word is carrying a lot of the verdict. The company says plainly that its figures cannot be independently verified and asks readers to weigh the method themselves. I take that at face value. You can read the full protocol and the round-by-round losses in the MoneyBench methodology.

My read, as someone who distrusts vendor benchmarks by default: this is a credible directional result on one narrow domain, not a settled fact and not a general ranking of these assistants. For explaining money, ChatGPT and Gemini are strong and cost nothing to try. For working through a specific decision with numbers that have to be current, Ed was the more useful tool in this test, and it stayed ahead after its own facts were re-checked and its early loss was left on the record. That combination is rarer than the headline number, and it is the part I would actually trust.

See the full MoneyBench results, method and losses

ChatGPT is a product of OpenAI; Gemini is a product of Google; Claude is a product of Anthropic. EdWealth is not affiliated with, endorsed by, or sponsored by OpenAI, Google, or Anthropic, and all trademarks are the property of their respective owners.

Educational content only. Not financial, tax, or investment advice. For your situation, consult a CPA or licensed professional. Reviewed August 2026.

Comments

TechBullion

FinTech News and Information

Copyright © 2026 TechBullion. All Rights Reserved.

To Top

Pin It on Pinterest

Share This