DeepSeek V4 Flash Is Repricing the Post-Coding-Plan Era
DeepSeek V4 Flash is the cheapest serious pay-as-you-go model, and fixed coding plans are losing ground.

¥0.074/M makes DeepSeek V4 Flash the cheapest serious pay-as-you-go model.
GLM’s repeated plan changes, Copilot’s weekly caps, and usage resets across platforms all point to the same conclusion: the age of generous coding bundles is over, and token pricing is back at the center of product strategy.
Pay-as-you-go is replacing the coding bundle
Get the latest AI news in your inbox
Weekly picks of model releases, tools, and deep dives — no spam, unsubscribe anytime.
No spam. Unsubscribe at any time.
When vendors tighten old plans, they are not being coy, they are admitting that heavy coding workloads no longer fit inside flat-rate offers. Anyone who watched Copilot add weekly limits or saw platform quotas shrink has already seen the economics break in public.

That matters because the old bundle model hid the real cost of agentic coding. Once users run long context windows, repeated tool calls, and multi-step generation loops, the bill stops looking like software subscription math and starts looking like infrastructure math.
Token mix now decides real cost
The useful way to compare models is not by headline price alone, but by the mix of input cache, input, and output tokens. The article’s Claude Code-based ratio, 95:4.5:0.5, is a practical benchmark because coding agents spend most of their budget on cached context, not fresh generation.
Using that mix, DeepSeek V4 Flash lands at roughly ¥0.074 per million tokens, based on cached input at ¥0.02, input at ¥1, and output at ¥2. That is not just cheap in absolute terms, it is cheap enough to reset expectations for what a coding assistant should cost at scale.
DeepSeek V4 Flash changes the comparison set
The important comparison is no longer premium model versus premium model. It is whether a vendor’s pay-as-you-go rate can beat the cheapest subscription tier once real usage patterns are applied, and on that test DeepSeek V4 Flash comes out ahead of GLM Lite monthly pricing.

That is a meaningful shift because it turns token economics into a product feature. If a model is cheap enough, teams can afford more agent runs, more retries, and more background automation without treating every prompt as a budget decision.
The counter-argument
The strongest objection is that raw token price is not the same as total value. A slightly more expensive model can still win if it produces better code, needs fewer retries, or integrates more cleanly into a developer workflow.
That objection is correct, and it is why price alone should never be the only selection criterion. But it does not rescue overpriced plans, because once a model is cheap enough to unlock broad usage, quality becomes the differentiator again instead of a tax on experimentation.
What to do with this
If you are an engineer, PM, or founder, stop evaluating coding models as a flat monthly seat and start modeling them as a workload cost. Measure cached input share, average output length, retry rate, and tool-call depth, then compare vendors on effective cost per successful task, not sticker price.
// Related Articles
- [IND]
Cloudflare’s AI search pilot is the stock catalyst now
- [IND]
AI Weekly: 2026-07-27 ~ 2026-08-03
- [IND]
Kimi K3 maps the new rules of test-time scaling
- [IND]
What the SALP liquidation reveals about AI trades
- [IND]
Claude’s security test became a real breach
- [IND]
X posts let execs shape the AI story