# Google DeepMind Achieves 98th Percentile on Mathematical Reasoning Benchmark

> Google DeepMind reports that its latest mathematical reasoning system now scores in the top 1 percent on International Mathematical Olympiad problems, the proof-heavy competition questions that have long been a benchmark for genuine reasoning rather than pattern matching. A 98th percentile result puts the system in the company of the strongest human competitors on this particular test. Olympiad problems are a useful yardstick because they resist memorization. Each one demands a chain of logical steps and a construction the solver has to invent, which is very different from recalling a fact or completing a familiar template. Progress here has been one of the clearer signals that models are getting better at multi-step reasoning, not just retrieval. The caveat, as always, is scope. Doing well on curated competition mathematics is not the same as open-ended research mathematics, where problems are unbounded and there is no answer key. Still, as part of DeepMind's July research output it adds to a run of results suggesting that structured, verifiable reasoning is an area where these systems keep climbing, and where the ceiling is not yet in sight.

_Section: [Daily AI Updates](https://www.wortins.com/daily-ai) · Source: Google DeepMind · Published Saturday, August 1, 2026_

## Wortins' read

Google DeepMind reports that its latest mathematical reasoning system now scores in the top 1 percent on International Mathematical Olympiad problems, the proof-heavy competition questions that have long been a benchmark for genuine reasoning rather than pattern matching. A 98th percentile result puts the system in the company of the strongest human competitors on this particular test. Olympiad problems are a useful yardstick because they resist memorization. Each one demands a chain of logical steps and a construction the solver has to invent, which is very different from recalling a fact or completing a familiar template. Progress here has been one of the clearer signals that models are getting better at multi-step reasoning, not just retrieval. The caveat, as always, is scope. Doing well on curated competition mathematics is not the same as open-ended research mathematics, where problems are unbounded and there is no answer key. Still, as part of DeepMind's July research output it adds to a run of results suggesting that structured, verifiable reasoning is an area where these systems keep climbing, and where the ceiling is not yet in sight.

## Source

[Read the full story at Google DeepMind](https://deepmind.google/research/publications/)

## Related coverage

- [Clearview AI Is Testing an AI Tool That Would Let Cops Unearth Your Life Online](https://www.wortins.com/story/clearview-ai-is-testing-an-ai-tool-that-would-let-cops-unear-8c573d0e) — [Wired](https://www.wired.com/story/clearview-ai-is-testing-an-ai-tool-that-lets-cops-instantly-unearth-your-online-activity/)
- [Powering AI is an architecture problem](https://www.wortins.com/story/powering-ai-is-an-architecture-problem-15281f16) — [MIT Technology Review](https://www.technologyreview.com/2026/09/10/1141649/powering-ai-is-an-architecture-problem/)
- [Inception launches Mercury 2.5 at 1,107 tokens per second](https://www.wortins.com/story/inception-launches-mercury-2-5-at-1-107-tokens-per-second-7999ed1b) — [TestingCatalog](https://www.testingcatalog.com/inception-launches-mercury-2-5-at-1-107-tokens-per-second/)
- [STAT+: Can AI fix the emergency room?](https://www.wortins.com/story/stat-can-ai-fix-the-emergency-room-a71d00ee) — [STAT](https://www.statnews.com/2026/09/09/scribe-emergency-room-fix-health-care-ai-prognosis/?utm_campaign=rss)
- [Top AI spenders cut per-employee costs by nearly 10 percent in August](https://www.wortins.com/story/top-ai-spenders-cut-per-employee-costs-by-nearly-10-percent--b31ffb0b) — [The Decoder](https://the-decoder.com/top-ai-spenders-cut-per-employee-costs-by-nearly-10-percent-in-august/)
- [Researchers used AI to build a WeChat worm that spreads through phone calls](https://www.wortins.com/story/researchers-used-ai-to-build-a-wechat-worm-that-spreads-thro-8bf3f469) — [The Next Web](https://thenextweb.com/news/wechat-worm-ai-calif-tencent-zero-click)

---

_Curated and written by [Wortins](https://www.wortins.com) — The daily AI briefing. Every story links to its original source; the "Wortins read" on each is our own original analysis. [About Wortins & our editorial approach](https://www.wortins.com/about)._
