$5.5 million. That's what DeepSeek spent to train a model that outperforms GPT-4 on several benchmarks. OpenAI reportedly invested around $100 million to train GPT-4, approximately 18 times more. When DeepSeek released its app in January 2025, it became the most-downloaded app in the world within hours. NVIDIA's stock lost around $600 billion in market capitalization in a single day.
Think AI is an American game? Think again.
This article covers all current numbers, data, and facts about DeepSeek: from user counts and downloads to training costs and benchmarks, API pricing, and company metrics.
- DeepSeek shook the AI market in January 2025: 173M downloads, 96.9M monthly active users, and the #1 app worldwide in the App Store
- DeepSeek-V3 was trained for only $5.5M, roughly 1/18 of estimated GPT-4 training costs, yet achieves comparable benchmark results
- External capital only since June 2026: after years without VC funding, DeepSeek raised $7.4B and is now preparing a Shanghai IPO, with headcount last reported at 150-200 employees before the June-announced doubling
1. What is DeepSeek?
DeepSeek is a Chinese AI research company founded in July 2023 in Hangzhou. Behind it is Liang Wenfeng, who previously co-founded the quantitative hedge fund High-Flyer Capital. High-Flyer invested heavily in building a large GPU cluster before the US introduced export restrictions on high-performance chips to China.
The stated goal:
Frontier AI research with a fraction of the resources used by Western labs. DeepSeek publishes its models as open source on HuggingFace and releases technical reports transparently, a practice that is rare among commercial AI providers.
What's special about the team: Instead of experienced industry AI researchers, DeepSeek deliberately hires fresh graduates and doctoral students. Only around 20-25% of employees have more than three years of professional experience.
2. Downloads and User Numbers
January 2025 was the month DeepSeek turned the AI world upside down. The numbers speak for themselves:
- 173M downloads total (by May 2025)
- 96.9M Monthly Active Users (April 2025)
- 130M Monthly Active Users by end of 2025 (Business of Apps)
- 22.15M Daily Active Users at peak (January 2025)
- 57.2M app downloads (34.6M Google Play + 22.6M App Store by May 2025)
For comparison: ChatGPT took two months after its November 2022 launch to reach 100 million users. That was a record at the time. DeepSeek came close to that in just three months, without the global marketing budget of a Silicon Valley giant.
2.1 Geographic Distribution
DeepSeek is globally distributed, with a clear home base:
The relatively low US share of 4.34% is notable. It's partly due to political concerns and the fact that American companies and agencies often avoid DeepSeek for security reasons.
3. Models and Technical Specifications
DeepSeek has released several powerful models in a short time. Key point: all major models are open source and can be run on your own hardware.
3.1 DeepSeek-V3
DeepSeek-V3 is the flagship model for general tasks. The technical specs are impressive:
- 671 billion parameters total, with 37 billion activated per token (Mixture-of-Experts architecture)
- Training cost: $5.5M (2.788M H800 GPU hours)
- Available as API and open source on HuggingFace
The fact is: DeepSeek-V3 is on par with GPT-4o on the MMLU benchmark. On coding (HumanEval-Mul) and math (MATH-500) it is ahead of both, according to DeepSeek's own comparison table, while Claude 3.5 Sonnet stays ahead on knowledge tests like GPQA-Diamond. All of this for $5.5 million in training costs.
3.2 DeepSeek-R1
DeepSeek-R1 is the reasoning model, comparable to OpenAI's o1. It's optimized for logical reasoning, mathematics, and complex coding tasks:
- MMLU: 90.8% (higher than V3 at 88.5%)
- AIME 2025: 87.5% (mathematical olympiad problems); AIME 2024 per original paper: 79.8%
- Open source, weights available on HuggingFace
3.3 DeepSeek-V3.2, V4, and V4.1 (2026)
The latest models push even further:
- V3.2-Exp: $0.28 per million input tokens, AIME 2025: 89.3%. The compute-heavy V3.2-Speciale variant reaches 96.0%, surpassing GPT-5 High (94.6%)
- V4 Pro: previewed on April 24, 2026, reached general availability as build 0813 on August 13, 2026. A 1.6 trillion parameter MoE with 49B active per token, native 1M token context window, MIT license. In its maximum reasoning mode (V4-Pro-Max), DeepSeek reported preview-time scores of 87.5% on MMLU-Pro, 90.1% on GPQA Diamond, and 93.5% on LiveCodeBench.
- V4 Flash: 284B parameter MoE with 13B active, the cost-efficient sibling of V4 Pro. Since July 31, 2026 the official DeepSeek-V4-Flash-0731 release superseded the earlier preview, with substantially enhanced agentic capabilities per DeepSeek. As of September 10, 2026, V4 Flash has been retired, with requests under that name now served by V4.1 Flash. MIT license.
- V4.1 Flash (September 10, 2026): a 552 billion parameter MoE built on a new Causal Encoder-Decoder architecture, activating 8B parameters at prefill and 16B at decode. Native 1M token context window, native multimodal image understanding, MIT license. DeepSeek says tests by multiple parties put it ahead of the larger V4 Pro on performance, cost, speed, and total runtime.
4. API Pricing
The lowest price is DeepSeek's strongest argument in the developer community. Here's a direct comparison with the leading competitors (DeepSeek's rates apply during off-peak hours outside the 01:00-04:00 and 06:00-10:00 UTC peak windows, when input and output prices double):
Just how wide the gap is becomes clear in a direct comparison of input costs. DeepSeek V4.1 Flash sits at $0.15 per million tokens, while GPT-5.6 Terra and Claude Sonnet 5.5 cost many times more:
DeepSeek V4.1 Flash costs only around 7.5% of what Claude Sonnet 5.5 charges. This keeps putting pressure on the entire industry's API pricing.
5. Benchmark Comparison with Competitors
Benchmarks aren't a perfect measure. However, they provide a good overview of a model's relative performance on standardized tasks.
It gets even more impressive on math reasoning. On AIME 2025 (mathematical olympiad problems), DeepSeek V3.2-Speciale scores 96.0%, ahead of GPT-5 High at 94.6%:
For how Meta's new flagship Muse Spark holds up in the same field of independent benchmarks, see my Meta AI statistics.
6. Revenue and Company Data
DeepSeek is an unusual company: small, efficient, yet on par with billion-dollar labs.
- Annualized revenue: per The Information (Sep 23, 2026, cited via PYMNTS), DeepSeek's run rate more than doubled within a few months to around $1B, driven by price hikes of 2.3 to 4.5 times on its flagship V4 models since August 2026. The older $400-500M (July 2026) estimate is now outdated. Revenue in the first seven months of 2026 had already roughly tenfolded compared to all of 2025
- API calls per month: 5.7B (2025)
- Funding: Internally backed by High-Flyer Capital, no external VC round until June 2026
- First external round (June 2026): roughly $7.4B raised at a $50-59B valuation (closed mid-June 2026)
- Employees: last reported at 150-200 (vs. OpenAI's 4,500+)
- Hiring push: after that first round, DeepSeek said it plans to at least double headcount across every department (Bloomberg, Jun 25, 2026)
- Pre-IPO round: per The Information (Sep 23, 2026), DeepSeek is targeting roughly $7.5B (50B yuan) at a $75B valuation (500B yuan); per the South China Morning Post, the round was not yet closed as of Sep 9, 2026, and neither source gives a firm closing date. DeepSeek is preparing a Shanghai STAR Market listing in parallel and has hired CITIC Securities among its underwriters
Dividing the roughly $1B run rate by the last-reported 150-200 employees, my own calculation puts DeepSeek at about $5-6.7M in revenue per employee. This is very high for an AI startup, clear evidence that DeepSeek's approach of working with a small, highly qualified team works not just scientifically, but economically.
7. The DeepSeek Shock: Market Impact
January 27, 2025 will go down in technology market history. That day, NVIDIA lost approximately $600 billion in market capitalization in a single trading session, a record in stock market history.
The trigger was the announcement that DeepSeek-V3 achieves results comparable to Western labs at a fraction of their resource costs. Investors asked: if AI training becomes this cheap, do we still need thousands of NVIDIA GPUs?
- NVIDIA stock decline: approximately -17% in one day (January 27, 2025)
- NVIDIA market cap loss: approximately -$600B
- Political response: Intensification of US chip export controls for China
- Regulatory response: Privacy authorities in Italy, India, and Australia investigated DeepSeek
Sam Altman, CEO of OpenAI, commented on X: "impressive model, particularly around what they're able to deliver for the price." From a direct competitor, that's almost a compliment.
8. Founders and History
Liang Wenfeng is not a typical AI founder. He studied mathematics and engineering before entering quantitative trading. With his co-founder Xu Jin, he founded High-Flyer Capital, one of China's most successful quantitative hedge funds.
High-Flyer recognized early that AI would be the next major lever in financial markets. The company began investing heavily in NVIDIA GPUs before the US introduced export restrictions. DeepSeek emerged from this GPU infrastructure as an internal research project, spun off as an independent company in 2023.
9. Availability and Open Source
DeepSeek is accessible through multiple channels:
- App: iOS and Android (free)
- Web chat: chat.deepseek.com
- API: platform.deepseek.com
- Open source (HuggingFace): V3, R1, V4 Pro, and V4.1 Flash as open-source models under the MIT license
The open-source release of model weights has sparked a wide range of community projects. Developers run DeepSeek models locally via tools like OpenClaw or Ollama. With local deployment, privacy concerns disappear entirely.
10. Interesting Facts and Records
- #1 app worldwide in the App Store in January 2025, just days after launch.
- Largest single-day market cap loss triggered by an external announcement: NVIDIA lost approximately $600B through DeepSeek's announcement.
- One of the cheapest frontier models in the world: DeepSeek V4.1 Flash at $0.15 per million input tokens ($0.003 on cache hits).
- Complete V3 training for $5.5M: Less than the budget of an average Hollywood film production.
- 150-200 employees competing with Western AI labs each employing thousands.
- Open source despite commercial success: DeepSeek releases frontier models freely, even though they would be extremely valuable commercially.
- Team of recent graduates: Most developers are fresh university graduates without years of industry experience.
- "Sputnik moment" of AI: That's how many analysts describe the DeepSeek moment, it made the Western world reconsider its assumptions about AI dominance.
For more comparisons, see our articles on Claude Statistics, ChatGPT Statistics, and Gemini Statistics.






