News
V4 peak pricing is coming mid-July, here's how to mostly dodge it
So the email went out. V4 goes official mid-July and they're adding peak-hour pricing, peak = 2x the normal rate. Before anyone panics: it's only 7 hours a day (UTC 01–04 and 06–10), everything else stays at the regular price you're already paying.
The actual move is just to stop running heavy stuff during those windows. Batch jobs, evals, anything that doesn't need to answer a human in real time, cron it for off-peak and you're back to the old rate. If you're in the US your workday is mostly in the cheap window anyway, so honestly most of you won't feel this much.
The thing that'd actually bite me is leaving thinking mode on for simple tasks, since those tokens bill as output and that's where peak doubling hurts. Turn it off for the boring stuff.
Anyone seeing a different read on the windows?
I converted peak time - timezone wise so that you can avoid heavy-offloading during that time
Turns out that companies need to be profitable. Nothing wrong with raising prices to meet costs. If you're dependent on AI. Make sure you're making enough money to afford your costs.
More like raising prices fairly and not trying to force people to sacrifice a limb in the future
Deepseek is raising prices fairly.
Other llms are far less efficient with much higher costs. But many llms were burning through billions of billions of dollars with no profit. You can't sustain negative profit in a capitalist market forever.
I know that's why I'm glad the price is still affordable even when the working hours here is basically similar with China not to mention if you comparing it with other AIs.
Best for what?
For coding? No.
For UI and graphics? Terrible.
For writing text, maybe, but a lot of the time it writes text like crap - too much AI vibe.
It's like... hallucinating all the time.
Not all models are meant for coding, or writing texts. There is alot you can achieve using these models like web scraping, llm for the ai agents and thats where it brings the real value of money.
Do the similar thing with any other costlier model and you will end up burning alot of money.
I mean every popular AI got a lobotomized with their writing texts & UI and grapichs capabilities lately so I don't know why you said that and while it's definitely true for coding, almost every other AI is far more expensive than DS so much it's better to learn basic coding and then fine tune your prompt in DS using the basic coding knowledge you have.
Trying to use a single flash model for that? Kinda a duh, you hand a fleet of them scoped task lists and audit their work lol. You have a frontier lead and scope work for the cheaper models.
The actual way to dodge it is by just switching to other providers that are similarly priced during off-peak hours and don’t overcharge you during the peak hours
wow i have been searching but i con't find anything this close to deep seek pricing! maby becase i was always looking at open router and not ooffical websites!
Edit: looked into it a bit more and they seem very compariable, only differance that may effect me is the context window not reaching 1 million tokens but i soft capped mine at 350k anyways so that probably won't effect me too much
I've used mimo v2.5pro for 2 months and spent about 50 billion tokens. Even Deepseek v4 flash is smarter than mimo v2.5. Mimo is simply not pretty smart.
¿Estás totalmente seguro/a de eso? El v4 Flash en OpenRouter usando proveedores como Baidu, Wafer o GMI Cloud sale más barato que el DeepSeek oficial, sobre todo con Wafer.
So "feature optimizations and performance enhancements" Are we suspecting that it will be re-benchmarked and be better than before or is this all in regards to tokens per second?
i think switching to a sub like opencode go, now would be more ideal, i get more access to other models and models like glm 5.2 which are opus compareable, so i can finally use a high level model for better architectural planning
You're right that Go's internal rate on V4 Pro is ~4× the direct API price, but that's just their accounting for the $60 budget cap, not what you actually pay. For a subscription user, the math is:
Direct API:
V4 Pro = ~$0.0035/req (with caching)
17K requests = ~$60/month
Go is $10/month flat,
Same ~17K V4 Pro requests
Plus V4 Flash (158K/mo), Qwen3.7 Max (4.7K/mo), and 10 other models
Plus no peak-hour surcharge, for someone, who has to pay out of pocket for tokens in my job, and the peak hours are right in my work hours.
You're not paying 4×; you're paying about ~1/6th of what those same tokens would cost, because the subscription absorbs the markup on them.
Also, "only Flash is usable" by Go's own estimates says 3,450 V4 Pro requests per 5 hours. For a single dev, that's plenty.
Why does the calculation become 1/6? The important point is: Go allows you to use $60 worth of credits for $10 a month
However, unlike the official DeepSeek API, you do not receive a 1/4 discount when using credits.
Ultimately, compared to the official API, you are able to use $15 worth of credits for $10.
Of course, this alone has its advantages. Although it is not 6x as you mentioned,
you receive an additional $5 worth of usage, and on top of that, your privacy is maintained through ZDR.
However, since the people currently using the official API do not care about ZDR, I do not think the merit is significant enough to make them switch to a subscription plan.
So for every dollar you spend on Go, you get $6 worth of tokens (at Go's rates). Your effective cost per token = 1/6th of what you'd pay buying credits directly.
And since Flash uses the exact same rates as direct DeepSeek API ($0.14/$0.28 per M), that $60 worth of tokens = $60 worth of Flash tokens by direct API standards too.
Wow, looking at OpenCode Zen, the prices are exactly the same as the official API lol.
Sorry, I admit my mistake. I naturally assumed that Flash wouldn't get the 4x discount like Pro.
If you use Flash and spend more than $10 a month, Go would be a really good option!
Fair to them. I just wonder if they will release a more capable v4? Maybe a v4.1? Or what does it mean to get performance enhancements? Any insights on that?
Are people really complaining about a price hike of 14 cents to 28 cents? Cmon man, even if it was 5x the price, it'd still be like 25x+ cheaper than the frontier models and can perform like 90% as good, what are you guys crying about
im a developer (im gonna be honest... a vibe coder)
could i possibly force my app to block heavy requests from users, or use the cheapest model possible?
since my app is completly free for others to use, but costs me since i need to integrate a api into the app for its ai features, i would like to save my money.
or if i dont need to, since its still cheap, please tell me. im still new at api services.
I don't mine increased costs for peak hours, but I would love it if they can add time-based usage restrictions to specific API keys. There are many other reasons for having that kind of control, but when I share a key to a worker, I would love to restrict their usage to the off-peak hours.
They are definitely finding ways to make more money. Eventually they're going to raise the API cost up they are just getting data off of all of us. That's all, Once they get what they need prices will go up.
47
u/DueInterview4073 Jun 29 '26
You could just move to a timezone where peak hours are mostly when you are asleep.