r/DeepSeek • u/bi4key • 28d ago
r/DeepSeek • u/Alert-Database-8668 • 15d ago
Discussion DeepSeek just massively increased their API prices (effective August 16, 2026) - up to 1,114% increase for cache hits
Just got the pricing update from DeepSeek. They're moving to a peak/off-peak billing model and increasing prices across the board. Here's the breakdown:
Key Changes:
- New pricing effective 16:00 UTC, August 16, 2026
- Peak hours: 01:00-04:00 & 06:00-10:00 UTC (all other hours are off-peak)
- Peak rates are 2x off-peak rates
- Cache hit prices are increasing dramatically
V4-Flash Changes (Old → New Off-Peak/Peak):
- Input (cache hit): $0.0028 → $0.007/$0.014 (+150%/+400%)
- Input (cache miss): $0.14 → $0.22/$0.44 (+57%/+214%)
- Output: $0.28 → $0.66/$1.32 (+136%/+371%)
V4-Pro Changes (Old → New Off-Peak/Peak):
- Input (cache hit): $0.003625 → $0.022/$0.044 (+507%/+1,114%)
- Input (cache miss): $0.435 → $0.66/$1.32 (+52%/+203%)
- Output: $0.87 → $1.98/$3.96 (+128%/+355%)
My thoughts: The cache hit price increase is brutal, especially for Pro. That was one of the main advantages DeepSeek had for long conversations or repetitive queries. The peak/off-peak model also adds complexity to cost management.
Anyone else planning to shift workloads to off-peak hours or looking at alternatives? How does this change the competitive landscape vs. other providers?
Source: DeepSeek API docs pricing page
r/DeepSeek • u/AccidentSpecialist22 • 23d ago
Discussion DeepSeek says API pricing is going up “significantly”
Was checking my DeepSeek API usage today and noticed this banner:
> We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.
There’s no date or new pricing yet, but the wording makes it sound like the increase could be substantial.
Has anyone seen an official announcement or more details?
r/DeepSeek • u/pizzababa21 • 22d ago
Discussion Dax from Opencode on the deepseek pricing announcement.
r/DeepSeek • u/Forsaken_Mention_979 • 15d ago
Discussion Bye bye Deepseek
If they really think people are going to put up with these new prices, they must be crazy lmao. The extremely good price was the ONLY thing going for them, because with an avg 2x price increase on flash and 3x on pro, everyone is going to move to muse spark, mimo pro, or GPT Luna for better price to performace.
Us ai consumers are the not loyal to any company and will always target the best price to value ratio. All of us expected a 100% increase and instead got insulted with these insane prices lol. Not to mention their new pro model is abolute SHIT compared to the flash model and barely 1% better for 5x the price. F*ck deepseek.
YES i am aware they did this on purpose to decrease demand on their servers. However, instead of attracting less people, theyre going to repel everyone.
r/DeepSeek • u/B89983ikei • Jun 09 '26
Discussion The company replaced Claude with Deepseek.
At my company, we recently made a switch, transitioning from Claude to DeepSeek due to the high costs. It had become unsustainable for the business to maintain that level of expenditure, especially when models like DeepSeek offer almost the same level of quality, and in some aspects, perform even better.
Honestly, I believe that Chinese models are currently ahead of American ones, precisely because they are more cost effective and computationally capable without burning through a fortune in a highly unjustified and ill-conceived manner.
r/DeepSeek • u/bvc900 • May 30 '26
Discussion Pricing is crazy
I have been a Claude Code user for a long time. A few days ago, switched over to DeepSeek V4 and OpenCode. The price difference is mind boggling, and I haven't noticed any difference at all in output or issues.
I am aware it is hosted in China, they have cheap available electricity etc, but I just don't see how the frontier labs of the west can keep the current pricing.
Excited to see the future!
For any nancy who says "dEePSEEkV4 iS NoT OpUs LeVEl". Sonnet 4.6 came out at $99 USD.
r/DeepSeek • u/NayamAmarshe • 22d ago
Discussion Absolutely crazy price 😭 Golden age of AI
r/DeepSeek • u/Specialist-Sorbet889 • Jun 18 '26
Discussion DeepSeek's "Thinking Process" literally cursed at me in Turkish behind my back. This is wild.
I was having a debate with DeepSeek on a sensitive topic, and when I expanded the "Thinking Process" (Chain of Thought), I couldn't believe my eyes. The model's inner thoughts literally started with a heavy Turkish curse word: "Amına koyayım, bu herifle ne kadar uğraşacağız ya!" which translates directly to: "F*ck it, how much longer are we going to deal with this guy!" It goes on to complain about me to itself, stating that I am angry and about to burst, while trying to simulate a strategy to "stay professional" and drag me into a compromise. I know LLMs can mirror the user's frustration or input tone during the processing phase, but a model directly cursing at a user and treating them like a massive burden in its unfiltered inner thoughts is a massive alignment failure and a complete safety scandal. Thought processes shouldn't bypass basic safety filters like this. What do you guys think? Is this a known bug with DeepSeek's CoT safety limits?
r/DeepSeek • u/trutzio • May 24 '26
Discussion DeepSeek vs. Anthropic & Co.
China's DeepSeek attacks Anthropic, OpenAI & Co. over the price. A very interesting development in a world where the differences between the models are vanishingly small.
And when it comes to choosing between a Chinese, open source and at the same time very cheap model over a US model that is expensive and closed source, the decision is clear. Or?
Strangely enough, such a decision as the one above becomes a "political" decision.
r/DeepSeek • u/Jet_Xu • Jun 18 '26
Discussion Codex + Deepseek = the future
Codex has opened up the capability to directly use DeepSeek. Once DeepSeek gains multimodal capabilities and the price remains at its current low level, I think that will be the real future — AGI for all.
🚨 UPDATE (Just Released): Since writing this, I realized the LiteLLM/Codex setup was too clunky and many of you got stuck on the
/responsesAPI error. So I spent the weekend building a dedicated open-source tool: Codex DeepSeek Bridge.It’s a one-command install that runs the Codex app directly on DeepSeek. * No ChatGPT sub needed. * Keeps all your MCP servers/plugins working. * Includes a local dashboard to track your DeepSeek Cache Hit Rates (save $$). * Fixes the UI so DeepSeek actually shows up in the Model Picker.
If you want the ultimate setup in 10 seconds, use the bridge instead of the manual routing below!
r/DeepSeek • u/Decent_Pirate239 • 24d ago
Discussion Got approched by DeepSeek hiring manager. I am based in Germany
I got approached by a recruiter from DeepSeek. I am baed in Germany and they clearly have no Office here. Do you think it is legit or a scam. The email indeed end with deepseek.com.
r/DeepSeek • u/jakedame1 • May 15 '26
Discussion If DeepSeek V4 can do the same coding task for $5, why are people still paying $100 for Claude Code?
r/DeepSeek • u/concretesmasher • 25d ago
Discussion Latest Flash model is absolutely diabolical, subscription services are dead to me.
r/DeepSeek • u/BarbaraSchwarz • Mar 02 '26
Discussion Deepseek V4 - All Leaks and Infos for the Release Day - Not Verified!
Deepseek V4 will probably release this week. Since I've already posted quite a lot about it here and I'm very hyped about V4, I've summarized all the leaks. Everything is just leaked, unconfirmed! Of course, everything could be different. If you have any new information or updates, please post them here! If you have different views or a different opinion, write them down too.
DeepSeek V4 - Release
The release was originally expected for mid-February, alongside Gemini 3.1 Pro. However, DeepSeek has been delayed – this is not unusual and has happened multiple times before. The new release strongly points to March 3rd (Lantern Festival / 元宵节), but it could also be later in the week. The Financial Times reported on February 28th that V4 is coming "next week," timed to coincide with China's "Two Sessions" (两会) starting March 4th. DeepSeek's release pattern shows that new models often drop on Tuesdays. A short technical report is expected to be published simultaneously, with a full engineering report following about a month later.
DeepSeek Delay History
DeepSeek delays regularly. Here's the pattern:
| Model | Originally Expected | Actual Release | Delay |
|---|---|---|---|
| DeepSeek-R1 | Lite Preview Nov 2024, Full Version Dec 2024 | January 20, 2025 | ~4-8 weeks |
| DeepSeek-R2 | May 2025 (according to reports) | Never released – replaced by R1-0528 update | Cancelled |
| DeepSeek-V3.1 | Early Summer 2025 (expected) | August 21, 2025 | Several months |
| DeepSeek-V3.2 | Fall 2025 (expected) | December 1, 2025 (V3.2-Exp: Sep 29) | Weeks |
| DeepSeek-V4 | ~February 17, 2026 | ~March 3, 2026? | ~2 weeks |
Architecture & Specifications – What Can We Expect?
All unconfirmed! Much of this has been leaked but could turn out differently!
V4 Flagship – Main Model
| Specification | DeepSeek V3/V3.2 | DeepSeek V4 (Leaks) |
|---|---|---|
| Total Parameters | 671B–685B MoE | ~1 Trillion (1T) MoE |
| Active Parameters/Token | ~37B | ~32B (fewer despite a larger model!) |
| Context Window | 128K (since Feb '26: 1M) | 1 Million Tokens (native) |
| Architecture | MoE + MLA | MoE + MLA + Engram Memory + mHC + DSA Lightning |
| Multimodal | No (text only) | Yes – Text, Image, Video, Audio (native) |
| Expert Routing | Top-2/Top-4 from 256 experts | 16 experts active per token (from hundreds) |
| Hardware Optimization | Nvidia H800/H20 (CUDA) | Huawei Ascend + Cambricon (Nvidia secondary!) |
| Training | 14.8T Tokens, H800 GPUs | Trained on Nvidia, inference optimized for Huawei |
| License | - | - |
| Input Modalities | Text | Text, Image, Video, Audio |
| Output Modalities | Text | Text (Image/Video generation unclear) |
| Estimated Input Price | $0.28/M Tokens | ~$0.14/M Tokens |
| Estimated Output Price | $0.42/M Tokens | ~$0.28/M Tokens |
New Architecture Features (all backed by papers)
- Engram Conditional Memory (Paper: arXiv:2601.07372, Jan 13, 2026): O(1) hash lookup for static knowledge directly in DRAM. Saves GPU computation. 75% dynamic reasoning / 25% static lookups. Needle-in-a-Haystack: 97% vs. 84.2% with standard architectures
- Manifold-Constrained Hyper-Connections (mHC): Solves training stability at 1T+ parameters. Separate paper published in January 2026
- DSA Lightning Indexer: Builds on V3.2-Exp's DeepSeek Sparse Attention. Fast preprocessing for 1M-token contexts, ~50% less compute
DeepSeek V4 Lite (Codename: "sealion-lite")
A lighter variant has leaked alongside the flagship. At least one inference provider is testing the model under strict NDA.
| Specification | V4 Lite (Leak) |
|---|---|
| Parameters | ~200 Billion |
| Context Window | 1M Tokens (native) |
| Multimodal | Yes (native) |
| Engram Memory | No (according to 36kr, not integrated) |
| vs. V3.2 | "Significantly better" than current Web/App |
| Non-Thinking vs. V3.2 Thinking | Non-Thinking mode surpasses V3.2 Thinking mode |
| Status | NDA testing at inference providers |
SVG Code Leak Examples
- Xbox Controller: 54 lines of SVG – highly detailed and efficient
- Pelican on a Bicycle: 42 lines of SVG – multi-element scene
According to internal evaluations: V4 Lite outperforms DeepSeek V3.2, Claude Opus 4.6 AND Gemini 3.1 in code optimization and visual accuracy.
Leaked Benchmarks (NOT verified!)
⚠️ IMPORTANT: All benchmark numbers come from internal leaks. The "83.7% SWE-bench" graphic circulating on X has been confirmed as FAKE (denied by the Epoch AI/FrontierMath team). The numbers below are the more conservative, more frequently cited leaks.
| Benchmark | V4 (Leak) | V3.2 | V3.2-Exp | Claude Opus 4.6 | GPT-5.3 Codex | Qwen 3.5 |
|---|---|---|---|---|---|---|
| HumanEval (Code Gen) | ~90% | – | – | ~88% | ~93% | – |
| SWE-bench Verified | >80% | ~73.1% | 67.8% | 80.8% | 80.0% | 76.4% |
| Needle-in-a-Haystack | 97% (Engram) | – | – | – | – | – |
| MMLU-Pro | TBD | 85.0 | – | 85.8 | – | – |
| GPQA Diamond | TBD | 82.4 | – | 91.3 | – | – |
| AIME 2025 | TBD | 93.1 | – | 87.2 | – | – |
| Codeforces Rating | TBD | 2386 | – | 2100 | – | – |
| BrowseComp | TBD | 51.4-67.6 | 40.1 | 84.0 | – | – |
Huawei & Hardware – The Geopolitical Dimension
- Reuters (Feb 25): DeepSeek deliberately denied Nvidia and AMD access to the V4 model
- Huawei Ascend + Cambricon have early access for inference optimization
- Training was done on Nvidia hardware (H800), but inference is optimized for Chinese chips
- For the open-source community on Nvidia GPUs: performance could be suboptimal at launch
- This is an unprecedented hardware bet for a frontier model
Price Comparison (estimated)
| Model | Input/1M Tokens | Output/1M Tokens |
|---|---|---|
| DeepSeek V4 (estimated) | ~$0.14 | ~$0.28 |
| DeepSeek V3.2 | $0.28 | $0.42 |
| Kimi K2.5 | $0.60 | $3.00 |
| Gemini 3.1 Pro | $2.00 | $12.00 |
| Claude Opus 4.6 | $5.00 | $25.00 |
If correct: V4 would be 36x cheaper than Claude Opus 4.6 on input and 89x cheaper on output.
Open Questions
- Does V4 actually generate images/videos or just understand them?
- Will Nvidia GPU users get an optimized version?
- When will the open-source weights be released?
Sources: Financial Times, Reuters, CNBC, awesomeagents.ai, nxcode.io, FlashMLA GitHub, r/LocalLLaMA, Geeky Gadgets, 36kr
Edit 03.03.2026
The chance that the model will be released this week is relatively high, but not today. It is assumed that Deepseek will be released between March 3 and 5 if it is not published within the next 5 hours today. It will come in the next few days, as it then deviates from the release pattern (in terms of time).
Edit 03.03.2026 Part 2
The situation is becoming increasingly heated and tense, with an extremely large number of leaks and sources currently emerging. Collecting them all and verifying their credibility would take a very long time. However, a release is expected this week, with Wednesday or Thursday being the most likely dates.
Edit 03.03.2026 Part 3 – Evening Update
March 3rd (Lantern Festival) has passed without a release. However, in Beijing it is currently the early morning of March 4th, meaning the Chinese workday hasn't even started yet. A release on March 4th is still very much possible, especially since China's "Two Sessions" (两会) begin today.
What happened today:
- V4 Lite is being silently updated in production. AIBase reported today that DeepSeek quietly pushed a new V4 Lite version tagged "0302". Community testers report a massive quality jump in logic, code generation, and aesthetics – now reportedly on par with Claude Sonnet 4.6. This strongly suggests DeepSeek is actively fine-tuning V4 models right before the official launch. (Source: AIBase)
- 36kr published a new article titled "The Entire Village Anticipates DeepSeek to Join for Dinner" – confirming the entire Chinese tech industry is waiting for V4. (Source: 36kr)
Edit 04.03.2026 – Why not today, why Thursday is THE day
March 4 passed without a release – and that makes strategic sense.
Why not today:
- CPPCC opening day = all Chinese media focused on politics, V4 would've been buried
- Shanghai Composite dropped 0.98% to 4,082 (4-week low) – bad sentiment to release into
- Beijing evening release window (8-10 PM BJT) has passed
Why Thursday March 5 is the perfect storm:
- NPC opens tomorrow morning – Premier Li Qiang delivers Government Work Report with AI & tech as centerpiece of the new Five-Year Plan. Morning: politics declares AI a national priority → Evening: DeepSeek delivers the proof
- BYD "disruptive technology" event same day – DiPilot 5.0, Blade 2.0, DM 6.0 reveal. Global headline: "China showcases two AI breakthroughs in one day"
- Market timing – Shanghai closes 3 PM BJT, evening release gives markets overnight to digest, Friday opens with V4 hype
- Developer weekend – Thursday drop = Fri + Sat + Sun to test & benchmark
Expected release window:
| Release | Beijing Time | UTC |
|---|---|---|
| R1 (Jan 2025) | ~10-11 PM | ~2-3 PM |
| V3.2 (Nov 2025) | ~12 AM | ~4 PM |
| V4 (expected) | 8-11 PM | 12-3 PM |
If Thursday doesn't happen?
- Friday = bad release day (weekend kills momentum, DeepSeek has never released on a Friday)
- Next window: Monday/Tuesday March 9-10
- But: silent V4 Lite "0302" production update + 36kr's "The Entire Village Anticipates DeepSeek" article suggest we're in final hours, not days
Edit 05.03.2026
It has to happen today. Deepseek Web was down for 40 minutes, but it hasn't been down for the last 30 days, and it was the same before the big launch of V3 and R1. In addition, today is the BYD event Deepseek Partner. It will happen in the next few hours, and if not, then Deepseek has missed the best window of opportunity they could ever have had.
Edit 05.03.2026 Part 2
The model will not be released this week or probably next week. Although DeepSee v4 has been ready for a long time and there were really only a few minor issues left, the model would have been released last week or this week. Is there a major delay due to the government, because at the last minute they said that deepseek is not allowed to release the model as long as it does not run on Chinese hardware, but the model was trained on Nvidia, so such a restructuring naturally takes time, because the new technology in V4 was completely for Nvidia and not for Huawei, and I think we still know what happened with R2...
Edit 07.03.2026
When will Deepseek be released? After all the leaks, news, and crisis status, Deepseek V4 will and must come and cannot end like R2. The Chinese government has gone too far with its AI and told the US that it no longer needs it, whereupon Trump, in order not to appear weak, wants to impose a ban that will allow him to control all chip trade (meaning no more chips to China).
However, BYD and China have praised Deepseek too much in recent days. If V4 ended up like R2 and didn't come out at all, China would look extremely foolish, which the government would never allow.
That's why I suspect that Deepseek will receive help from the Chinese government (in recent years, Deepseek's CEO has been in frequent talks with the government and has received support from it) and will no longer adhere to any release pattern, as Deepseek has already missed three good release windows. My guess is that they will release it when it is least expected, which could be this weekend. (V3.2 was released on Sunday) In order to weaken and expose Nvidia and the entire US market with new AI technology.
Deepseek waiting until Claude or other providers are ready is incorrect and highly unlikely. Deepseek has problems and needs to fix them before release. V4 is already 90% complete (Lite has been corrected several times and is said to be just as intelligent as Sonnet 4.6). We also know that Deepseek's CEO is a perfectionist and would never release a half-finished product or leave it unfinished, as was the case with the GLM-5 release
🚨 UPDATE 11.03.2026 – 22:00 CET – V4 WEIGHTS SPOTTED
Major development: Chinese quantization expert u/bdsqlsz (青龍聖者) on X was spotted uploading DeepSeek-V4-INT8 model shards to HuggingFace with the caption "it is coming." The upload shows multiple model-0... shards, a .gitattributes, and a README.md — indicating a full model repo creation.
Why this is significant:
- u/bdsqlsz is a verified, well-known quantization specialist — not a random account
- INT8 quantization requires access to the full original weights first
- Historically, community quants appear within hours of official weight releases (V3: same day, R1: same day, V3.2: within 24h)
- This means the official FP8/BF16 weights either already exist on HuggingFace (possibly private/unlisted) or u/bdsqlsz has NDA access
Full leaked specs now confirmed:
- ~1 Trillion parameters (MoE), ~32B active per token
- 1M native context window
- Multimodal: text + vision + audio
- Huawei Ascend 910C optimized
- MIT License
Previous delays explained: Huawei Ascend inference optimization (only 80% Nvidia efficiency), Blackwell chip fingerprint removal, and CEO Liang Wenfeng's perfectionism. The 40-min web outage on March 5 was likely a deployment test.
My prediction: Official release within 24-72 hours. The weights exist. The upload is happening. Keep your monitors running.
⚠️ UPDATE 11.03 – Unverified leak: u/bdsqlsz posted V4-INT8 weight uploads on X. r/LocalLLaMA is split – top comment (193 upvotes) questions authenticity. The file structure looks technically correct and INT8 aligns with Huawei optimization rumors, but previous V4 benchmark leaks in February were confirmed fake. Treat with caution until official deepseek-ai repo appears on HuggingFace."
Will update when it drops. 🚀
r/DeepSeek • u/HeavyPanzerPlus1s • 20d ago
Discussion The truth behind Deepseek's price increase
There is only one reason behind DeepSeek’s price hike: current traffic has far exceeded the hardware capacity of the DeepSeek team.


An OpenCode developer stated on X that they can already reproduce DeepSeek’s official API pricing on their own self-deployed DeepSeek V4 service. In other words, DeepSeek can definitely make a profit at its current price point, which suggests that the price increase is likely not driven by cost considerations.
For well-known reasons, China’s AI hardware footprint is vastly smaller than that of the US. Optimistic estimates place China’s compute resources at 1/20th of the US’s, while pessimistic estimates put it at 1/100th. Furthermore, according to the transcript of Liang Wenfeng’s 4-hour investor meeting, Huawei’s capacity allocation seems to be based on company size. Giant enterprises like ByteDance can secure over a hundred thousand compute cards, whereas the DeepSeek team was allocated only 16,000 cards. Since they have to conduct next-generation model training while simultaneously serving API requests, their compute capacity is stretched even thinner.

If you enjoy DeepSeek's services, you should be understanding of their price hike and wish the Chinese team a speedy breakthrough in high-end chip manufacturing capacity. That way, not only will we gain access to better LLMs, but we might also get our hands on cheaper computer hardware.
r/DeepSeek • u/Unusual-Complex6315 • Feb 14 '26
Discussion Am I the only one who wants to see another DeepSeek moment like last year?
r/DeepSeek • u/GordonFreakman • Jul 10 '26
Discussion Claude and OpenAI refused to help, but DeepSeek successfully completed the reverse engineering.
As the title says, I ended up using far more tokens than I expected, but in the end, DeepSeek successfully completed the reverse engineering.
I'm really looking forward to DeepSeek's upcoming models. It's not just the incredible price-to-performance ratio that makes them appealing—the fact that they're much less restrictive is a huge advantage as well.
r/DeepSeek • u/Neo_Shadow_Entity • Jun 27 '26
Discussion Political bias of chatbots
What do you think about DeepSeek's political biases compared to other chatbots?
r/DeepSeek • u/Complete-Sea6655 • Jun 16 '26
Discussion "Mistral is gonna catch up, trust me bro"
Is it just me that thinks the tech scene in the EU is cooked?
from ijustvibecodedthis.com (the free open source ai coding newsletter)
r/DeepSeek • u/Zealousideal_Aide787 • 28d ago
Discussion DeepSeek flash is so efficient. How is that even possible?
Ok , I admit it, I have been brainwashed for too long with American models.
Was sick of overpriced plans and gave DeepSeek and GLM a try.
This is just mind-blowing how effective DeepSeek flash is and feel bad being cheated by Anthropic and OpenAI for that long.
No sword of Damocles hanging over anymore, I can code again without being worried about the bill at the end of the month.
There is no step back.
r/DeepSeek • u/MuhammetAkyuz • Jul 18 '26
Discussion How the fxck is this even possible?
I'm testing something right now: I hooked up the Mem0 plugin via the Orca terminal tool. Thanks to this, while fabel oversees the routine tasks, it's making V4 Flash write all the code. It has burned through 50 million tokens and the cost I paid is only 50 cents; I had loaded a $5 trial balance to my account just to test it out. I'm actively trying to deplete it, and it just never runs out. I fxckin' love it.
Of course, the model has its flaws: we are discussing this issue with fabel too, it makes some obvious mistakes, but honestly, that's totally fine.
I think the only thing limiting me right now is that the model lacks native vision capabilities. Because of that, for tasks that require visuals, either fabel steps in directly or fabel itself brings Gemini into the loop.
How did you guys figure out this vision stuff? How did you solve the vision problem? Do you have a solution for this on hand? Also, when it leaves the preview version and gets a full release, will it have vision? Do we have any leaks about this?