# Cerebras WSE-3 Chip Achieves 6.7x Faster Inference Than GPU Cloud Providers

> Cerebras says its WSE-3 chip has been independently clocked running Moonshot's trillion-parameter Kimi K2 model at 981 output tokens per second, roughly 6.7 times faster than mainstream GPU cloud providers. The benchmark was verified by Artificial Analysis, which lends it weight beyond a vendor claim. The hardware is genuinely extreme. WSE-3 is the largest AI chip ever built, a single wafer-scale part measuring 46,225 square millimeters with 4 trillion transistors and 900,000 cores, delivering 125 petaflops. Cerebras pegs it at 19 times the transistors and 28 times the compute of Nvidia's B200. The speed matters most for agents. A 10,000-token agentic request that takes about 163 seconds on the official endpoint finishes in roughly 5.6 seconds here, the difference between a tool you wait on and one that feels instant. As inference becomes the real cost center of AI, results like this position Cerebras as one of the few credible challengers to Nvidia's grip on enterprise deployment.

_Section: [Daily AI Updates](https://www.wortins.com/daily-ai) · Source: VentureBeat · Published Monday, August 3, 2026_

## Wortins' read

Cerebras says its WSE-3 chip has been independently clocked running Moonshot's trillion-parameter Kimi K2 model at 981 output tokens per second, roughly 6.7 times faster than mainstream GPU cloud providers. The benchmark was verified by Artificial Analysis, which lends it weight beyond a vendor claim. The hardware is genuinely extreme. WSE-3 is the largest AI chip ever built, a single wafer-scale part measuring 46,225 square millimeters with 4 trillion transistors and 900,000 cores, delivering 125 petaflops. Cerebras pegs it at 19 times the transistors and 28 times the compute of Nvidia's B200. The speed matters most for agents. A 10,000-token agentic request that takes about 163 seconds on the official endpoint finishes in roughly 5.6 seconds here, the difference between a tool you wait on and one that feels instant. As inference becomes the real cost center of AI, results like this position Cerebras as one of the few credible challengers to Nvidia's grip on enterprise deployment.

## Source

[Read the full story at VentureBeat](https://venturebeat.com/technology/cerebras-says-its-chips-run-a-trillion-parameter-ai-model-nearly-7-times-faster-than-gpu-clouds)

## Related coverage

- [Researchers used AI to build a WeChat worm that spreads through phone calls](https://www.wortins.com/story/researchers-used-ai-to-build-a-wechat-worm-that-spreads-thro-8bf3f469) — [The Next Web](https://thenextweb.com/news/wechat-worm-ai-calif-tencent-zero-click)
- [A look at why the oft-discussed predictions that AI will deliver double-digit GDP growth in advanced economies are extremely unlikely over the next 10-15 years (Ghosts of Electricity)](https://www.wortins.com/story/a-look-at-why-the-oft-discussed-predictions-that-ai-will-del-6241fa38) — [Techmeme](https://www.techmeme.com/260910/p10#a260910p10)
- [Suno launches v6 music models built with Warner, BMG, and Believe](https://www.wortins.com/story/suno-launches-v6-music-models-built-with-warner-bmg-and-beli-9ba43843) — [The Decoder](https://the-decoder.com/suno-launches-v6-music-models-built-with-warner-bmg-and-believe/)
- [Inception launches Mercury 2.5 at 1,107 tokens per second](https://www.wortins.com/story/inception-launches-mercury-2-5-at-1-107-tokens-per-second-7999ed1b) — [TestingCatalog](https://www.testingcatalog.com/inception-launches-mercury-2-5-at-1-107-tokens-per-second/)
- [Trump officials say AI will help save rural health care. Some leaders in the field don’t believe it](https://www.wortins.com/story/trump-officials-say-ai-will-help-save-rural-health-care-some-bf0e7d83) — [STAT](https://www.statnews.com/2026/09/10/rural-health-care-ai-adoption-challenges-part-4-unraveled-series/?utm_campaign=rss)
- [IBM and NASA release an open-source lunar foundation model](https://www.wortins.com/story/ibm-and-nasa-release-an-open-source-lunar-foundation-model-0522b9e8) — [The Next Web](https://thenextweb.com/news/nasa-ibm-lunar-foundation-model-open-source)

---

_Curated and written by [Wortins](https://www.wortins.com) — The daily AI briefing. Every story links to its original source; the "Wortins read" on each is our own original analysis. [About Wortins & our editorial approach](https://www.wortins.com/about)._
