Liang Wenfeng: DeepSeek and the Open-Weight Earthquake

On this page10 sections

Liang Wenfeng: DeepSeek and the Open-Weight Earthquake

Liang Wenfeng
Born:
1985 (Wuchuan, Zhanjiang, Guangdong Province)
Nationality:
Chinese
Field:
AI / Quantitative finance
Key contribution:
Founded DeepSeek and led the R1 release that triggered a US stock rout in January 2025
DeepSeek

A Chinese AI company founded in 2023 by Liang Wenfeng that released the DeepSeek-R1 reasoning model in January 2025. DeepSeek R1 matched or exceeded OpenAI o1 on benchmarks while being open-weights and trained at a fraction of the cost, resetting the AI-race narrative.


The Unlikely Founder

Liang Wenfeng was born in 1985 in Wuchuan, a county in Zhanjiang, in China’s southern Guangdong Province. His father was a primary school teacher. He enrolled at Zhejiang University in Hangzhou, studying information and electronic engineering, earning a bachelor’s degree and then a master’s degree in 2010. His graduate work focused on machine vision; his thesis was on low-cost camera tracking.

After graduation, Liang stayed in Hangzhou — also the headquarters of Alibaba and one of China’s major technology hubs. In 2015 or 2016, he and classmates from Zhejiang University founded High-Flyer (幻方量化, “Magic Square Quant”) — a quantitative hedge fund that used mathematical models rather than human judgment to make trading decisions.


The Hedge Fund That Built a Supercomputer

High-Flyer’s core insight was that deep learning could be applied to financial data. To do this, the company needed GPUs — the specialised chips originally designed for video games that turned out to be ideal for deep learning’s matrix-math requirements.

In 2020, High-Flyer built its first major GPU cluster, Fire-Flyer I, with about 1,100 NVIDIA A100 chips (~200 million yuan / ~$30 million). In 2021, the company built Fire-Flyer II, with 10,000 A100 chips (~1 billion yuan / ~$150 million) — one of the largest private GPU clusters in China. The cluster was used for quantitative trading, and High-Flyer became one of China’s most successful quant funds, reporting returns of about 57% in 2025.

But the cluster had another use: the same computing power that predicted stock prices could train large language models — the kind of AI systems that powered ChatGPT. Liang believed LLMs were the path to artificial general intelligence.


The Founding of DeepSeek

In July 2023, Liang founded DeepSeek (深度求索, “deep seeking” or “deep exploration”) as a subsidiary of High-Flyer, based in Hangzhou, funded entirely by High-Flyer, focused on a single goal: building AGI. The company was different from American frontier labs in almost every respect — funded by a hedge fund, based in China, and committed from the beginning to releasing models as open weights (model parameters freely available for anyone to download, modify, and use).

In a July 2024 interview, Liang explained: “We will not change to closed source. We believe having a big moat of originality is more important than a closed-source moat.” He also said: “In disruptive tech, closed-source moats are fleeting. Even OpenAI’s closed-source model can’t prevent others from catching up.”


The Models: V2, V3, and the Path to R1

DeepSeek released its first public models in November 2023 (7B and 67B parameters). The breakthrough came in May 2024 with DeepSeek-V2 — 236 billion total parameters using Mixture of Experts (MoE) architecture (only ~21 billion parameters active per input), making it much cheaper to run than a traditional dense model of the same size.

Mixture of Experts and DeepSeek’s cost advantage

Mixture of Experts (MoE) means only a subset of parameters is active for any given input — dramatically reducing inference cost. Combined with DeepSeek’s Multi-head Latent Attention (a more efficient attention mechanism that reduced memory requirements during training) and fine-grained MoE architecture, this allowed the company to scale total parameters without proportionally increasing compute.

On December 26, 2024, DeepSeek released V3 — 671 billion total parameters, 37 billion active per token, trained on 14.8 trillion tokens (the units of text that language models process). The technical report stated it required only 2.788 million H800 GPU-hours for its full training run.


The Training Cost Question

The 2.788 million GPU-hours figure became the most controversial number in the AI industry. Multiplied by a rough rental rate of $2 per GPU-hour, it yields about $5.6 million — widely cited as DeepSeek-V3’s training cost. But this figure is misleading:

  1. It refers only to V3’s final training run — not researcher salaries, data costs, failed experiments, or the $150 million Fire-Flyer II cluster.
  2. It is for V3, not for R1 (the reasoning model built on top of V3, released January 20, 2025). Reuters reported in September 2025 that R1’s specific reinforcement-learning fine-tuning cost only $294,000 using 512 H800 chips — but this excludes V3’s training cost.
  3. CNBC reported DeepSeek’s total hardware spend was about $500 million; other estimates exceeded $1 billion.

The honest summary: DeepSeek’s final training run for V3 cost about $5.6 million in compute. The total cost of building DeepSeek’s capability was much higher. But even the higher figure was striking, because American frontier labs were believed to be spending billions.


R1 and the Reasoning Breakthrough

On January 20, 2025, DeepSeek released R1 — a reasoning model (trained to generate explicit chains of reasoning before answering, in the manner of OpenAI’s o1). R1 matched or exceeded o1 on many benchmarks: 79.8% on AIME 2024 (vs o1’s 79.2%), 97.3% on MATH-500, and an Elo rating of 2029 on Codeforces (top 1.5% of human competitors).

R1 was released under the MIT license — anyone could download, modify, and use it commercially, for free. This was unprecedented: OpenAI’s o1 was available only through a paid API.

The R1 paper also described R1-Zero, a variant trained using pure reinforcement learning without human-generated reasoning examples — suggesting reasoning could be incentivised without expensive human-labelled data.

DeepSeek released six smaller distilled versions of R1 (1.5B to 70B parameters), built on Qwen and Llama base models. Within days, Hugging Face created an “Open-R1” project, and Perplexity released “R1-1776.”


January 27, 2025: DeepSeek Monday

The release on January 20 did not immediately cause panic. That changed on Monday, January 27, when the DeepSeek app became the most-downloaded free app on Apple’s US App Store, displacing ChatGPT. The combination of technical achievement, the $5.6 million figure, and the app’s popularity triggered a reassessment of AI investment assumptions.

NVIDIA’s stock fell ~17%, erasing approximately $589 billion — the largest single-day market-cap loss in US stock market history. The NASDAQ dropped 3.1%; the S&P 500 dropped 1.5%. The total value wiped out across the AI sector was estimated at ~$1 trillion.

President Trump called DeepSeek a “wake-up call.” Sam Altman posted that R1 was “impressive.” Demis Hassabis said the work was “impressive” but that DeepSeek had not done anything fundamentally new. David Sacks, the White House AI advisor, suggested DeepSeek might have distilled R1 from OpenAI’s models — a claim that was not proven.


The Geopolitical Backlash

The DeepSeek app collected user data and sent it to servers in China, raising the same concerns as TikTok. Italy’s Garante ordered ISPs to block DeepSeek on January 28, 2025. South Korea blocked employee access on February 5; DeepSeek voluntarily removed its app from Korean app stores on February 15. Australia, Taiwan, and Belgium restricted government use. The US Pentagon blocked employee access.

In June 2025, a US official told Reuters that DeepSeek aided China’s military and had used Southeast Asian shell companies to evade US chip export controls. The accusations were contested.

The backlash created a split: in the research community, R1 was widely adopted and built upon (Hugging Face hosted the weights; Meta and Microsoft used R1 for distillation). In the consumer market, the app faced growing restrictions.


The Open-Source Question

DeepSeek’s MIT license release made it a hero of the open-source AI community. But it raised the question of what “open source” means for AI. The Open Source Initiative’s Open Source AI Definition (OSAID) requires that open-source models not restrict uses — and R1 had built-in political filtering that some argued violated this principle. The debate continues.

DeepSeek’s open-weight approach also created a strategic dilemma for the US government. Export controls on chips were designed to slow China’s AI development, but DeepSeek had achieved frontier-level results with older, legally acquired H800 chips — suggesting the controls might be less effective than intended.


The Legacy

Liang Wenfeng’s legacy is, as of 2026, still being written. DeepSeek’s R1 release was one of the most consequential events in the history of the AI industry. It demonstrated that the gap between the American frontier and Chinese AI was narrower than assumed. It raised questions about the cost of training frontier models, the effectiveness of export controls, and the strategic value of open-weight releases. And it made Liang — a quiet, media-averse quant fund manager from Guangdong — one of the most consequential figures in the global AI race.

The deeper legacy may be about the relationship between openness and power. DeepSeek’s decision to release R1 as open weights was, in some ways, a strategic masterstroke — it made R1 ubiquitous, it built goodwill in the research community, and it undermined the narrative that Chinese AI was merely copying the West. But it also raised uncomfortable questions about who benefits from openness, and whether the open-weight approach could be used as a tool of geopolitical influence as much as a tool of scientific collaboration.


Further reading
  • DeepSeek-R1 technical report — arXiv, January 20, 2025. The R1 paper.
  • DeepSeek-V3 technical report — arXiv, December 26, 2024. The V3 paper.
  • “DeepSeek: The Chinese AI lab that shook Silicon Valley”Wired, January 2025.
  • Liang Wenfeng’s July 2024 interview — one of the few on-record interviews he has given.

Series Companions

This piece is part of Minds & Machines: Beyond the Series. The companion pieces B85 — DeepSeek-R1 and the Reset of the AI Race and B08 — The AI Chip Wars cover the market impact and the hardware context. B12 — Open vs Closed covers the open-weight debate.

Was Liang Wenfeng inevitable — the product of forces too large to redirect — or was it a series of choices, each of which could have gone differently? The answer matters, because it determines whether the future is something that happens to us or something we make.