It's the latest update to DeepSeek's Flash line — the line designed to be fast and cheap, not the biggest. The result of this version is that the distinction no longer makes sense: it now approaches the full models.
According to the source, it sits one point below GLM 5.2, beats the previous V4 Pro version, and climbs about 10 points over the previous Flash. The gain comes from the DSpark architecture, the same as the previous version, focused on efficiency and throughput.
Why this matters
- Costs about US$0.03 per million tokens — the source estimates it's roughly 100× cheaper than Claude Opus.
- It's 70% smaller than GLM 5.2 and delivers comparable performance.
- It's official: you can run intelligence at this level on your own machine.
1.Where it wins
3 minThe tests mentioned in the video cover three fronts, and they're exactly the ones that matter to people building things:
- Agentic coding — the model operating tools across multiple steps, not just writing a standalone function.
- Software engineering — real repository tasks.
- Cybersecurity.
Spec sheet
- License
- Open, available on Hugging Face
- Estimated cost
- ~US$ 0.03 / million tokens
- Size
- 167 GB (82,5 GB em GGUF de 1 bit)

Continue in the full microcourse
You've read the opening of 3 classes
The microcourse covers the complete step-by-step, the selection criteria, where the tool fails, who it's really for — and, in the Expert version, the official address to start today.
- Run on your own3 min
- When to switch from the paid model to this one2 min
How we verified
We track releases straight from primary sources, transcribe what's demonstrated, check every name and number against the manufacturer's official documentation, and rewrite it in Portuguese — with what the tool no do it together, which is the part the ad leaves out.


