# Inception launches Mercury 2.5 at 1,107 tokens per second

> Inception has released Mercury 2.5, a model it says runs at roughly 1,107 tokens per second, a striking speed that reflects its diffusion-based approach to text generation rather than the token-by-token method most large language models use. The company claims quality comparable to cost-optimized frontier models, including GPT-5, while emphasizing raw throughput as its main advantage. Mercury 2.5 is available through Inception, Baseten, and OpenRouter, with new Mercury Voice and Mercury Router products entering preview. Speed is more than a bragging point. Faster generation means lower latency for interactive apps and cheaper serving costs at scale, which matters most for agentic workloads that make many model calls in sequence. If a smaller player can match cost-optimized quality while pulling far ahead on speed, it pressures the incumbents on a dimension they do not always prioritize. For a technologist, Inception is worth watching precisely because it is not a household name. The diffusion-for-text bet is still unproven at the frontier, but results like these suggest the architecture is maturing into a real alternative.

_Section: [Daily AI Updates](https://www.wortins.com/daily-ai) · Source: TestingCatalog · Published Wednesday, September 9, 2026_

## Wortins' read

Inception has released Mercury 2.5, a model it says runs at roughly 1,107 tokens per second, a striking speed that reflects its diffusion-based approach to text generation rather than the token-by-token method most large language models use. The company claims quality comparable to cost-optimized frontier models, including GPT-5, while emphasizing raw throughput as its main advantage. Mercury 2.5 is available through Inception, Baseten, and OpenRouter, with new Mercury Voice and Mercury Router products entering preview. Speed is more than a bragging point. Faster generation means lower latency for interactive apps and cheaper serving costs at scale, which matters most for agentic workloads that make many model calls in sequence. If a smaller player can match cost-optimized quality while pulling far ahead on speed, it pressures the incumbents on a dimension they do not always prioritize. For a technologist, Inception is worth watching precisely because it is not a household name. The diffusion-for-text bet is still unproven at the frontier, but results like these suggest the architecture is maturing into a real alternative.

## Source

[Read the full story at TestingCatalog](https://www.testingcatalog.com/inception-launches-mercury-2-5-at-1-107-tokens-per-second/)

## Related coverage

- [Chinese tech giants are hiring skilled professionals as specialized AI trainers to build high-quality datasets, mirroring efforts by US platforms like Mercor (Viola Zhou/Rest of World)](https://www.wortins.com/story/chinese-tech-giants-are-hiring-skilled-professionals-as-spec-116fb4b8) — [Techmeme](https://www.techmeme.com/260910/p6#a260910p6)
- [Trump officials say AI will help save rural health care. Some leaders in the field don’t believe it](https://www.wortins.com/story/trump-officials-say-ai-will-help-save-rural-health-care-some-bf0e7d83) — [STAT](https://www.statnews.com/2026/09/10/rural-health-care-ai-adoption-challenges-part-4-unraveled-series/?utm_campaign=rss)
- [Sequoia doubles down on Cymphony as AI agents create new enterprise security risks](https://www.wortins.com/story/sequoia-doubles-down-on-cymphony-as-ai-agents-create-new-ent-e2ff7229) — [TechCrunch](https://techcrunch.com/2026/09/09/sequoia-doubles-down-on-cymphony-as-ai-agents-create-new-enterprise-security-risks/)
- [Two years ago, Meta killed CrowdTangle. Can a new AI tool fill the void?](https://www.wortins.com/story/two-years-ago-meta-killed-crowdtangle-can-a-new-ai-tool-fill-5e67f291) — [Nieman Lab](https://www.niemanlab.org/2026/09/two-years-ago-meta-killed-crowdtangle-can-a-new-ai-tool-fill-the-void/)
- [Letter: the Senate disaster management subcommittee, led by Sen. Josh Hawley, is probing OpenAI's handling of the Hugging Face breach, calling it "reckless" (Axios)](https://www.wortins.com/story/letter-the-senate-disaster-management-subcommittee-led-by-se-da54156c) — [Techmeme](https://www.techmeme.com/260910/p11#a260910p11)
- [IBM and NASA release an open-source lunar foundation model](https://www.wortins.com/story/ibm-and-nasa-release-an-open-source-lunar-foundation-model-0522b9e8) — [The Next Web](https://thenextweb.com/news/nasa-ibm-lunar-foundation-model-open-source)

---

_Curated and written by [Wortins](https://www.wortins.com) — The daily AI briefing. Every story links to its original source; the "Wortins read" on each is our own original analysis. [About Wortins & our editorial approach](https://www.wortins.com/about)._
