# DeepSeek V4 Flash Exits Preview, Beats Own Pro Model on Agent Benchmarks

> DeepSeek pushed its V4-Flash model out of preview on July 31, and the interesting part is not the launch but the scoreboard. Flash carries 284 billion total parameters with just 13 billion active, roughly a third of the company's own 1.6-trillion-parameter V4-Pro, yet it posts a Terminal-Bench score of 82.7 percent and beats Pro on agent tasks. Bigger, in other words, did not win. The pricing stays aggressive at 14 cents per million input tokens on a cache miss, dropping to a fraction of a cent on a hit, with output at 28 cents. It also carries a 1 million token context window and can emit up to 384,000 tokens in a single response. DeepSeek says the gains came entirely from post-training rather than any change to the architecture, which has held steady since the April preview. That is the quiet lesson of this release: a lot of headroom now lives in how a model is trained and tuned after the fact, not in stacking on more parameters. For anyone building agents, a small, cheap model that outperforms its heavyweight sibling is a genuinely useful data point.

_Section: [Daily AI Updates](https://www.wortins.com/daily-ai) · Source: Caixin Global · Published Monday, August 3, 2026_

## Wortins' read

DeepSeek pushed its V4-Flash model out of preview on July 31, and the interesting part is not the launch but the scoreboard. Flash carries 284 billion total parameters with just 13 billion active, roughly a third of the company's own 1.6-trillion-parameter V4-Pro, yet it posts a Terminal-Bench score of 82.7 percent and beats Pro on agent tasks. Bigger, in other words, did not win. The pricing stays aggressive at 14 cents per million input tokens on a cache miss, dropping to a fraction of a cent on a hit, with output at 28 cents. It also carries a 1 million token context window and can emit up to 384,000 tokens in a single response. DeepSeek says the gains came entirely from post-training rather than any change to the architecture, which has held steady since the April preview. That is the quiet lesson of this release: a lot of headroom now lives in how a model is trained and tuned after the fact, not in stacking on more parameters. For anyone building agents, a small, cheap model that outperforms its heavyweight sibling is a genuinely useful data point.

## Source

[Read the full story at Caixin Global](https://www.caixinglobal.com/2026-08-01/deepseek-releases-official-v4-flash-model-as-chinas-ai-race-intensifies-102470292.html)

## Related coverage

- [Top AI spenders cut per-employee costs by nearly 10 percent in August](https://www.wortins.com/story/top-ai-spenders-cut-per-employee-costs-by-nearly-10-percent--b31ffb0b) — [The Decoder](https://the-decoder.com/top-ai-spenders-cut-per-employee-costs-by-nearly-10-percent-in-august/)
- [Letter: the Senate disaster management subcommittee, led by Sen. Josh Hawley, is probing OpenAI's handling of the Hugging Face breach, calling it "reckless" (Axios)](https://www.wortins.com/story/letter-the-senate-disaster-management-subcommittee-led-by-se-da54156c) — [Techmeme](https://www.techmeme.com/260910/p11#a260910p11)
- [First ‘Take It Down Act’ Sentencing Puts Man Behind Bars for 15 Years](https://www.wortins.com/story/first-take-it-down-act-sentencing-puts-man-behind-bars-for-1-7a82b442) — [404 Media](https://www.404media.co/first-take-it-down-act-sentencing-case/)
- [Inception launches Mercury 2.5 at 1,107 tokens per second](https://www.wortins.com/story/inception-launches-mercury-2-5-at-1-107-tokens-per-second-7999ed1b) — [TestingCatalog](https://www.testingcatalog.com/inception-launches-mercury-2-5-at-1-107-tokens-per-second/)
- [Powering AI is an architecture problem](https://www.wortins.com/story/powering-ai-is-an-architecture-problem-15281f16) — [MIT Technology Review](https://www.technologyreview.com/2026/09/10/1141649/powering-ai-is-an-architecture-problem/)
- [Clearview AI Is Testing an AI Tool That Would Let Cops Unearth Your Life Online](https://www.wortins.com/story/clearview-ai-is-testing-an-ai-tool-that-would-let-cops-unear-8c573d0e) — [Wired](https://www.wired.com/story/clearview-ai-is-testing-an-ai-tool-that-lets-cops-instantly-unearth-your-online-activity/)

---

_Curated and written by [Wortins](https://www.wortins.com) — The daily AI briefing. Every story links to its original source; the "Wortins read" on each is our own original analysis. [About Wortins & our editorial approach](https://www.wortins.com/about)._
