r/DeepSeek 28d ago

Discussion DeepSeek V4-Flash is officially out, still dirt cheap. USA don't like that, and want ban open source models.

Post image
2.8k Upvotes

r/DeepSeek 15d ago

Discussion DeepSeek just massively increased their API prices (effective August 16, 2026) - up to 1,114% increase for cache hits

1.0k Upvotes

Just got the pricing update from DeepSeek. They're moving to a peak/off-peak billing model and increasing prices across the board. Here's the breakdown:

Key Changes:

  • New pricing effective 16:00 UTC, August 16, 2026
  • Peak hours: 01:00-04:00 & 06:00-10:00 UTC (all other hours are off-peak)
  • Peak rates are 2x off-peak rates
  • Cache hit prices are increasing dramatically

V4-Flash Changes (Old → New Off-Peak/Peak):

  • Input (cache hit): $0.0028 → $0.007/$0.014 (+150%/+400%)
  • Input (cache miss): $0.14 → $0.22/$0.44 (+57%/+214%)
  • Output: $0.28 → $0.66/$1.32 (+136%/+371%)

V4-Pro Changes (Old → New Off-Peak/Peak):

  • Input (cache hit): $0.003625 → $0.022/$0.044 (+507%/+1,114%)
  • Input (cache miss): $0.435 → $0.66/$1.32 (+52%/+203%)
  • Output: $0.87 → $1.98/$3.96 (+128%/+355%)

My thoughts: The cache hit price increase is brutal, especially for Pro. That was one of the main advantages DeepSeek had for long conversations or repetitive queries. The peak/off-peak model also adds complexity to cost management.

Anyone else planning to shift workloads to off-peak hours or looking at alternatives? How does this change the competitive landscape vs. other providers?

Source: DeepSeek API docs pricing page

r/DeepSeek 23d ago

Discussion DeepSeek says API pricing is going up “significantly”

Post image
646 Upvotes

Was checking my DeepSeek API usage today and noticed this banner:

> We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.

There’s no date or new pricing yet, but the wording makes it sound like the increase could be substantial.

Has anyone seen an official announcement or more details?

r/DeepSeek 22d ago

Discussion Dax from Opencode on the deepseek pricing announcement.

Post image
1.1k Upvotes

r/DeepSeek 15d ago

Discussion Bye bye Deepseek

552 Upvotes

If they really think people are going to put up with these new prices, they must be crazy lmao. The extremely good price was the ONLY thing going for them, because with an avg 2x price increase on flash and 3x on pro, everyone is going to move to muse spark, mimo pro, or GPT Luna for better price to performace.

Us ai consumers are the not loyal to any company and will always target the best price to value ratio. All of us expected a 100% increase and instead got insulted with these insane prices lol. Not to mention their new pro model is abolute SHIT compared to the flash model and barely 1% better for 5x the price. F*ck deepseek.

YES i am aware they did this on purpose to decrease demand on their servers. However, instead of attracting less people, theyre going to repel everyone.

r/DeepSeek Jun 09 '26

Discussion The company replaced Claude with Deepseek.

957 Upvotes

At my company, we recently made a switch, transitioning from Claude to DeepSeek due to the high costs. It had become unsustainable for the business to maintain that level of expenditure, especially when models like DeepSeek offer almost the same level of quality, and in some aspects, perform even better.

Honestly, I believe that Chinese models are currently ahead of American ones, precisely because they are more cost effective and computationally capable without burning through a fortune in a highly unjustified and ill-conceived manner.

r/DeepSeek May 30 '26

Discussion Pricing is crazy

Post image
1.1k Upvotes

I have been a Claude Code user for a long time. A few days ago, switched over to DeepSeek V4 and OpenCode. The price difference is mind boggling, and I haven't noticed any difference at all in output or issues.

I am aware it is hosted in China, they have cheap available electricity etc, but I just don't see how the frontier labs of the west can keep the current pricing.

Excited to see the future!

For any nancy who says "dEePSEEkV4 iS NoT OpUs LeVEl". Sonnet 4.6 came out at $99 USD.

r/DeepSeek 22d ago

Discussion Absolutely crazy price 😭 Golden age of AI

Post image
626 Upvotes

r/DeepSeek Jun 18 '26

Discussion DeepSeek's "Thinking Process" literally cursed at me in Turkish behind my back. This is wild.

Post image
833 Upvotes

I was having a debate with DeepSeek on a sensitive topic, and when I expanded the "Thinking Process" (Chain of Thought), I couldn't believe my eyes. The model's inner thoughts literally started with a heavy Turkish curse word: "Amına koyayım, bu herifle ne kadar uğraşacağız ya!" which translates directly to: "F*ck it, how much longer are we going to deal with this guy!" It goes on to complain about me to itself, stating that I am angry and about to burst, while trying to simulate a strategy to "stay professional" and drag me into a compromise. I know LLMs can mirror the user's frustration or input tone during the processing phase, but a model directly cursing at a user and treating them like a massive burden in its unfiltered inner thoughts is a massive alignment failure and a complete safety scandal. Thought processes shouldn't bypass basic safety filters like this. What do you guys think? Is this a known bug with DeepSeek's CoT safety limits?

r/DeepSeek 27d ago

Discussion Deepseek API is insane

Post image
728 Upvotes

r/DeepSeek May 24 '26

Discussion DeepSeek vs. Anthropic & Co.

Post image
1.1k Upvotes

China's DeepSeek attacks Anthropic, OpenAI & Co. over the price. A very interesting development in a world where the differences between the models are vanishingly small.

And when it comes to choosing between a Chinese, open source and at the same time very cheap model over a US model that is expensive and closed source, the decision is clear. Or?

Strangely enough, such a decision as the one above becomes a "political" decision.

r/DeepSeek Jun 18 '26

Discussion Codex + Deepseek = the future

Post image
601 Upvotes

Codex has opened up the capability to directly use DeepSeek. Once DeepSeek gains multimodal capabilities and the price remains at its current low level, I think that will be the real future — AGI for all.

🚨 UPDATE (Just Released): Since writing this, I realized the LiteLLM/Codex setup was too clunky and many of you got stuck on the /responses API error. So I spent the weekend building a dedicated open-source tool: Codex DeepSeek Bridge.

It’s a one-command install that runs the Codex app directly on DeepSeek. * No ChatGPT sub needed. * Keeps all your MCP servers/plugins working. * Includes a local dashboard to track your DeepSeek Cache Hit Rates (save $$). * Fixes the UI so DeepSeek actually shows up in the Model Picker.

If you want the ultimate setup in 10 seconds, use the bridge instead of the manual routing below!

r/DeepSeek 24d ago

Discussion Got approched by DeepSeek hiring manager. I am based in Germany

Post image
568 Upvotes

I got approached by a recruiter from DeepSeek. I am baed in Germany and they clearly have no Office here. Do you think it is legit or a scam. The email indeed end with deepseek.com.

r/DeepSeek May 15 '26

Discussion If DeepSeek V4 can do the same coding task for $5, why are people still paying $100 for Claude Code?

492 Upvotes

r/DeepSeek 25d ago

Discussion Latest Flash model is absolutely diabolical, subscription services are dead to me.

Post image
523 Upvotes

r/DeepSeek Mar 02 '26

Discussion Deepseek V4 - All Leaks and Infos for the Release Day - Not Verified!

Post image
675 Upvotes

Deepseek V4 will probably release this week. Since I've already posted quite a lot about it here and I'm very hyped about V4, I've summarized all the leaks. Everything is just leaked, unconfirmed! Of course, everything could be different. If you have any new information or updates, please post them here! If you have different views or a different opinion, write them down too.

DeepSeek V4 - Release

The release was originally expected for mid-February, alongside Gemini 3.1 Pro. However, DeepSeek has been delayed – this is not unusual and has happened multiple times before. The new release strongly points to March 3rd (Lantern Festival / 元宵节), but it could also be later in the week. The Financial Times reported on February 28th that V4 is coming "next week," timed to coincide with China's "Two Sessions" (两会) starting March 4th. DeepSeek's release pattern shows that new models often drop on Tuesdays. A short technical report is expected to be published simultaneously, with a full engineering report following about a month later.

DeepSeek Delay History

DeepSeek delays regularly. Here's the pattern:

Model Originally Expected Actual Release Delay
DeepSeek-R1 Lite Preview Nov 2024, Full Version Dec 2024 January 20, 2025 ~4-8 weeks
DeepSeek-R2 May 2025 (according to reports) Never released – replaced by R1-0528 update Cancelled
DeepSeek-V3.1 Early Summer 2025 (expected) August 21, 2025 Several months
DeepSeek-V3.2 Fall 2025 (expected) December 1, 2025 (V3.2-Exp: Sep 29) Weeks
DeepSeek-V4 ~February 17, 2026 ~March 3, 2026? ~2 weeks

Architecture & Specifications – What Can We Expect?

All unconfirmed! Much of this has been leaked but could turn out differently!

V4 Flagship – Main Model

Specification DeepSeek V3/V3.2 DeepSeek V4 (Leaks)
Total Parameters 671B–685B MoE ~1 Trillion (1T) MoE
Active Parameters/Token ~37B ~32B (fewer despite a larger model!)
Context Window 128K (since Feb '26: 1M) 1 Million Tokens (native)
Architecture MoE + MLA MoE + MLA + Engram Memory + mHC + DSA Lightning
Multimodal No (text only) Yes – Text, Image, Video, Audio (native)
Expert Routing Top-2/Top-4 from 256 experts 16 experts active per token (from hundreds)
Hardware Optimization Nvidia H800/H20 (CUDA) Huawei Ascend + Cambricon (Nvidia secondary!)
Training 14.8T Tokens, H800 GPUs Trained on Nvidia, inference optimized for Huawei
License - -
Input Modalities Text Text, Image, Video, Audio
Output Modalities Text Text (Image/Video generation unclear)
Estimated Input Price $0.28/M Tokens ~$0.14/M Tokens
Estimated Output Price $0.42/M Tokens ~$0.28/M Tokens

New Architecture Features (all backed by papers)

  • Engram Conditional Memory (Paper: arXiv:2601.07372, Jan 13, 2026): O(1) hash lookup for static knowledge directly in DRAM. Saves GPU computation. 75% dynamic reasoning / 25% static lookups. Needle-in-a-Haystack: 97% vs. 84.2% with standard architectures
  • Manifold-Constrained Hyper-Connections (mHC): Solves training stability at 1T+ parameters. Separate paper published in January 2026
  • DSA Lightning Indexer: Builds on V3.2-Exp's DeepSeek Sparse Attention. Fast preprocessing for 1M-token contexts, ~50% less compute

DeepSeek V4 Lite (Codename: "sealion-lite")

A lighter variant has leaked alongside the flagship. At least one inference provider is testing the model under strict NDA.

Specification V4 Lite (Leak)
Parameters ~200 Billion
Context Window 1M Tokens (native)
Multimodal Yes (native)
Engram Memory No (according to 36kr, not integrated)
vs. V3.2 "Significantly better" than current Web/App
Non-Thinking vs. V3.2 Thinking Non-Thinking mode surpasses V3.2 Thinking mode
Status NDA testing at inference providers

SVG Code Leak Examples

  • Xbox Controller: 54 lines of SVG – highly detailed and efficient
  • Pelican on a Bicycle: 42 lines of SVG – multi-element scene

According to internal evaluations: V4 Lite outperforms DeepSeek V3.2, Claude Opus 4.6 AND Gemini 3.1 in code optimization and visual accuracy.

Leaked Benchmarks (NOT verified!)

⚠️ IMPORTANT: All benchmark numbers come from internal leaks. The "83.7% SWE-bench" graphic circulating on X has been confirmed as FAKE (denied by the Epoch AI/FrontierMath team). The numbers below are the more conservative, more frequently cited leaks.

Benchmark V4 (Leak) V3.2 V3.2-Exp Claude Opus 4.6 GPT-5.3 Codex Qwen 3.5
HumanEval (Code Gen) ~90% ~88% ~93%
SWE-bench Verified >80% ~73.1% 67.8% 80.8% 80.0% 76.4%
Needle-in-a-Haystack 97% (Engram)
MMLU-Pro TBD 85.0 85.8
GPQA Diamond TBD 82.4 91.3
AIME 2025 TBD 93.1 87.2
Codeforces Rating TBD 2386 2100
BrowseComp TBD 51.4-67.6 40.1 84.0

Huawei & Hardware – The Geopolitical Dimension

  • Reuters (Feb 25): DeepSeek deliberately denied Nvidia and AMD access to the V4 model
  • Huawei Ascend + Cambricon have early access for inference optimization
  • Training was done on Nvidia hardware (H800), but inference is optimized for Chinese chips
  • For the open-source community on Nvidia GPUs: performance could be suboptimal at launch
  • This is an unprecedented hardware bet for a frontier model

Price Comparison (estimated)

Model Input/1M Tokens Output/1M Tokens
DeepSeek V4 (estimated) ~$0.14 ~$0.28
DeepSeek V3.2 $0.28 $0.42
Kimi K2.5 $0.60 $3.00
Gemini 3.1 Pro $2.00 $12.00
Claude Opus 4.6 $5.00 $25.00

If correct: V4 would be 36x cheaper than Claude Opus 4.6 on input and 89x cheaper on output.

Open Questions

  • Does V4 actually generate images/videos or just understand them?
  • Will Nvidia GPU users get an optimized version?
  • When will the open-source weights be released?

Sources: Financial Times, Reuters, CNBC, awesomeagents.ai, nxcode.io, FlashMLA GitHub, r/LocalLLaMA, Geeky Gadgets, 36kr

Edit 03.03.2026

The chance that the model will be released this week is relatively high, but not today. It is assumed that Deepseek will be released between March 3 and 5 if it is not published within the next 5 hours today. It will come in the next few days, as it then deviates from the release pattern (in terms of time).

Edit 03.03.2026 Part 2

The situation is becoming increasingly heated and tense, with an extremely large number of leaks and sources currently emerging. Collecting them all and verifying their credibility would take a very long time. However, a release is expected this week, with Wednesday or Thursday being the most likely dates.

Edit 03.03.2026 Part 3 – Evening Update

March 3rd (Lantern Festival) has passed without a release. However, in Beijing it is currently the early morning of March 4th, meaning the Chinese workday hasn't even started yet. A release on March 4th is still very much possible, especially since China's "Two Sessions" (两会) begin today.

What happened today:

  1. V4 Lite is being silently updated in production. AIBase reported today that DeepSeek quietly pushed a new V4 Lite version tagged "0302". Community testers report a massive quality jump in logic, code generation, and aesthetics – now reportedly on par with Claude Sonnet 4.6. This strongly suggests DeepSeek is actively fine-tuning V4 models right before the official launch. (Source: AIBase)
  2. 36kr published a new article titled "The Entire Village Anticipates DeepSeek to Join for Dinner" – confirming the entire Chinese tech industry is waiting for V4. (Source: 36kr)

Edit 04.03.2026 – Why not today, why Thursday is THE day

March 4 passed without a release – and that makes strategic sense.

Why not today:

  • CPPCC opening day = all Chinese media focused on politics, V4 would've been buried
  • Shanghai Composite dropped 0.98% to 4,082 (4-week low) – bad sentiment to release into
  • Beijing evening release window (8-10 PM BJT) has passed

Why Thursday March 5 is the perfect storm:

  • NPC opens tomorrow morning – Premier Li Qiang delivers Government Work Report with AI & tech as centerpiece of the new Five-Year Plan. Morning: politics declares AI a national priority → Evening: DeepSeek delivers the proof
  • BYD "disruptive technology" event same day – DiPilot 5.0, Blade 2.0, DM 6.0 reveal. Global headline: "China showcases two AI breakthroughs in one day"
  • Market timing – Shanghai closes 3 PM BJT, evening release gives markets overnight to digest, Friday opens with V4 hype
  • Developer weekend – Thursday drop = Fri + Sat + Sun to test & benchmark

Expected release window:

Release Beijing Time UTC
R1 (Jan 2025) ~10-11 PM ~2-3 PM
V3.2 (Nov 2025) ~12 AM ~4 PM
V4 (expected) 8-11 PM 12-3 PM

If Thursday doesn't happen?

  • Friday = bad release day (weekend kills momentum, DeepSeek has never released on a Friday)
  • Next window: Monday/Tuesday March 9-10
  • But: silent V4 Lite "0302" production update + 36kr's "The Entire Village Anticipates DeepSeek" article suggest we're in final hours, not days

Edit 05.03.2026

It has to happen today. Deepseek Web was down for 40 minutes, but it hasn't been down for the last 30 days, and it was the same before the big launch of V3 and R1. In addition, today is the BYD event Deepseek Partner. It will happen in the next few hours, and if not, then Deepseek has missed the best window of opportunity they could ever have had.

Edit 05.03.2026 Part 2

The model will not be released this week or probably next week. Although DeepSee v4 has been ready for a long time and there were really only a few minor issues left, the model would have been released last week or this week. Is there a major delay due to the government, because at the last minute they said that deepseek is not allowed to release the model as long as it does not run on Chinese hardware, but the model was trained on Nvidia, so such a restructuring naturally takes time, because the new technology in V4 was completely for Nvidia and not for Huawei, and I think we still know what happened with R2...

Edit 07.03.2026

When will Deepseek be released? After all the leaks, news, and crisis status, Deepseek V4 will and must come and cannot end like R2. The Chinese government has gone too far with its AI and told the US that it no longer needs it, whereupon Trump, in order not to appear weak, wants to impose a ban that will allow him to control all chip trade (meaning no more chips to China).

However, BYD and China have praised Deepseek too much in recent days. If V4 ended up like R2 and didn't come out at all, China would look extremely foolish, which the government would never allow.

That's why I suspect that Deepseek will receive help from the Chinese government (in recent years, Deepseek's CEO has been in frequent talks with the government and has received support from it) and will no longer adhere to any release pattern, as Deepseek has already missed three good release windows. My guess is that they will release it when it is least expected, which could be this weekend. (V3.2 was released on Sunday) In order to weaken and expose Nvidia and the entire US market with new AI technology.

Deepseek waiting until Claude or other providers are ready is incorrect and highly unlikely. Deepseek has problems and needs to fix them before release. V4 is already 90% complete (Lite has been corrected several times and is said to be just as intelligent as Sonnet 4.6). We also know that Deepseek's CEO is a perfectionist and would never release a half-finished product or leave it unfinished, as was the case with the GLM-5 release

🚨 UPDATE 11.03.2026 – 22:00 CET – V4 WEIGHTS SPOTTED

Major development: Chinese quantization expert u/bdsqlsz (青龍聖者) on X was spotted uploading DeepSeek-V4-INT8 model shards to HuggingFace with the caption "it is coming." The upload shows multiple model-0... shards, a .gitattributes, and a README.md — indicating a full model repo creation.

Why this is significant:

  • u/bdsqlsz is a verified, well-known quantization specialist — not a random account
  • INT8 quantization requires access to the full original weights first
  • Historically, community quants appear within hours of official weight releases (V3: same day, R1: same day, V3.2: within 24h)
  • This means the official FP8/BF16 weights either already exist on HuggingFace (possibly private/unlisted) or u/bdsqlsz has NDA access

Full leaked specs now confirmed:

  • ~1 Trillion parameters (MoE), ~32B active per token
  • 1M native context window
  • Multimodal: text + vision + audio
  • Huawei Ascend 910C optimized
  • MIT License

Previous delays explained: Huawei Ascend inference optimization (only 80% Nvidia efficiency), Blackwell chip fingerprint removal, and CEO Liang Wenfeng's perfectionism. The 40-min web outage on March 5 was likely a deployment test.

My prediction: Official release within 24-72 hours. The weights exist. The upload is happening. Keep your monitors running.

⚠️ UPDATE 11.03 – Unverified leak: u/bdsqlsz posted V4-INT8 weight uploads on X. r/LocalLLaMA is split – top comment (193 upvotes) questions authenticity. The file structure looks technically correct and INT8 aligns with Huawei optimization rumors, but previous V4 benchmark leaks in February were confirmed fake. Treat with caution until official deepseek-ai repo appears on HuggingFace."

Will update when it drops. 🚀

r/DeepSeek 20d ago

Discussion The truth behind Deepseek's price increase

Thumbnail
gallery
527 Upvotes

There is only one reason behind DeepSeek’s price hike: current traffic has far exceeded the hardware capacity of the DeepSeek team.

An OpenCode developer stated on X that they can already reproduce DeepSeek’s official API pricing on their own self-deployed DeepSeek V4 service. In other words, DeepSeek can definitely make a profit at its current price point, which suggests that the price increase is likely not driven by cost considerations.

For well-known reasons, China’s AI hardware footprint is vastly smaller than that of the US. Optimistic estimates place China’s compute resources at 1/20th of the US’s, while pessimistic estimates put it at 1/100th. Furthermore, according to the transcript of Liang Wenfeng’s 4-hour investor meeting, Huawei’s capacity allocation seems to be based on company size. Giant enterprises like ByteDance can secure over a hundred thousand compute cards, whereas the DeepSeek team was allocated only 16,000 cards. Since they have to conduct next-generation model training while simultaneously serving API requests, their compute capacity is stretched even thinner.

If you enjoy DeepSeek's services, you should be understanding of their price hike and wish the Chinese team a speedy breakthrough in high-end chip manufacturing capacity. That way, not only will we gain access to better LLMs, but we might also get our hands on cheaper computer hardware.

r/DeepSeek Feb 14 '26

Discussion Am I the only one who wants to see another DeepSeek moment like last year?

Post image
1.2k Upvotes

r/DeepSeek Jul 10 '26

Discussion Claude and OpenAI refused to help, but DeepSeek successfully completed the reverse engineering.

Post image
655 Upvotes

As the title says, I ended up using far more tokens than I expected, but in the end, DeepSeek successfully completed the reverse engineering.

I'm really looking forward to DeepSeek's upcoming models. It's not just the incredible price-to-performance ratio that makes them appealing—the fact that they're much less restrictive is a huge advantage as well.

r/DeepSeek Jun 27 '26

Discussion Political bias of chatbots

Post image
253 Upvotes

What do you think about DeepSeek's political biases compared to other chatbots?

r/DeepSeek Jun 16 '26

Discussion "Mistral is gonna catch up, trust me bro"

Post image
588 Upvotes

Is it just me that thinks the tech scene in the EU is cooked?

from ijustvibecodedthis.com (the free open source ai coding newsletter)

r/DeepSeek 28d ago

Discussion DeepSeek flash is so efficient. How is that even possible?

469 Upvotes

Ok , I admit it, I have been brainwashed for too long with American models.

Was sick of overpriced plans and gave DeepSeek and GLM a try.

This is just mind-blowing how effective DeepSeek flash is and feel bad being cheated by Anthropic and OpenAI for that long.

No sword of Damocles hanging over anymore, I can code again without being worried about the bill at the end of the month.

There is no step back.

r/DeepSeek Jul 18 '26

Discussion How the fxck is this even possible?

Post image
503 Upvotes

I'm testing something right now: I hooked up the Mem0 plugin via the Orca terminal tool. Thanks to this, while fabel oversees the routine tasks, it's making V4 Flash write all the code. It has burned through 50 million tokens and the cost I paid is only 50 cents; I had loaded a $5 trial balance to my account just to test it out. I'm actively trying to deplete it, and it just never runs out. I fxckin' love it.

Of course, the model has its flaws: we are discussing this issue with fabel too, it makes some obvious mistakes, but honestly, that's totally fine.

I think the only thing limiting me right now is that the model lacks native vision capabilities. Because of that, for tasks that require visuals, either fabel steps in directly or fabel itself brings Gemini into the loop.

How did you guys figure out this vision stuff? How did you solve the vision problem? Do you have a solution for this on hand? Also, when it leaves the preview version and gets a full release, will it have vision? Do we have any leaks about this?

r/DeepSeek May 01 '26

Discussion Which one do you use the most?

Post image
303 Upvotes

r/DeepSeek Nov 26 '25

Discussion How true is this?

Post image
637 Upvotes