TodaySunday, September 20, 2026

Naive AI Raised $400 Million to Build a Model It Won’t Pre-Train

Tsinghua professor Dai Jifeng raised $400 million from Tencent without plans to pre-train, betting that refining an existing Chinese open-weight model is worth more than building from zero.
September 20, 2026
3 mins read
Naive AI logo with Tencent funding announcement for Beijing LLM startup
Naive AI raised $400 million from Tencent at a $1.42 billion valuation to refine an open-weight large language model without pre-training from scratch. [Image Source: CryptoBriefing]

BEIJING — The number that matters in Naive AI’s announcement is not the figure at the top. It is the one that is missing.

Since February, Dai Jifeng has raised $400 million at a $1.42 billion valuation from Tencent, IDG Capital, and HSG, the firm formerly known as Sequoia Capital China, to build a large language model. None of that money, according to reporting by The Information, is going toward pre-training from scratch. That absence is the structural choice that makes this funding round worth examining.

Pre-training a frontier model is the most expensive phase of building one. Training on hundreds of billions of tokens requires warehouse-scale GPU clusters, months of compute time, and electricity costs that make the word startup generous. DeepSeek spent years and enormous computational resources reaching the performance it demonstrated in January. StepFun, the Shanghai startup that this week launched its 600-billion-parameter Step 5 Preview, has raised $3.2 billion and operates an industrial-scale training facility. Dai Jifeng is betting he can skip that entirely.

The strategy Naive AI is pursuing is sometimes called post-training refinement or mid-training. Rather than starting from scratch, the company is taking an existing open-weight Chinese model as its base (the specific model has not been publicly disclosed), modifying its architecture, running additional training passes, and applying reinforcement learning to improve performance across specific tasks. The model it plans to release, named Naive, is due out this month as an open-weight that anyone can download and fine-tune.

The economic case for this approach is straightforward. Frontier pretraining is consolidating among companies that have raised billions and can access either state-backed computing infrastructure or Nvidia hardware that U.S. export controls are attempting—imperfectly—to restrict.

Post-training allows a smaller team to compete in areas that remain less concentrated: specialization, alignment, task-specific refinement, and model behavior that benchmark-optimized base systems often fail to capture.

The risk is that the team is building on a foundation it does not own. If the base model has structural weaknesses in reasoning or world knowledge, post-training improvements will face a hard ceiling imposed by those limitations.

That ceiling depends entirely on which Chinese open-weight model Naive AI selected as its starting point, and the company has not disclosed that information. The quality of the model released this month will provide the first independent evidence of whether the strategy works.

Tencent headquarters and AI investment strategy in China's technology ecosystem
Tencent has positioned itself as a broad investor across China’s AI ecosystem, seeding multiple LLM startups alongside its own Hunyuan model development. [Image Source: CGTN]
Dai Jifeng’s background is credible for this kind of bet. He is an associate professor at Tsinghua University whose research over the past decade spans deep learning, computer vision, foundation models, and agentic AI. Before returning to academia, he held a position at Microsoft Research Asia, the lab that produced a generation of Chinese AI researchers who now lead major companies across the sector, and spent time at SenseTime, one of China’s largest computer vision companies before it was placed on the US entity list. The team he has assembled reportedly numbers fewer than a hundred people, which is either a sign of deliberate capital efficiency or an indication of how early the research actually is.

The three funding tranches tell their own story about how conviction built. Naive AI raised $100 million in its first round, $180 million in the second, then $120 million to close the current round. Its valuation sat around $800 million in April, after roughly $300 million had been committed. The most recent tranche nearly doubled that to $1.42 billion, a 78 percent jump for a company that has not yet publicly released a model.

That acceleration in valuation relative to demonstrated output reflects both the heat of China’s AI funding environment and a specific judgement about Dai Jifeng: that a Tsinghua professor with the right technical background and backing from Tencent and HSG is worth an increasing premium before the weights are public.

Tencent’s involvement functions, as it does with most of its AI portfolio, as a broad position across the ecosystem rather than a strategic endorsement of one company’s approach. Alibaba’s Qwen team released its own new model this week, a multimodal system built for agentic workflows, and the pace of Chinese LLM releases has made broad exposure more rational than concentrated bets. CGTN reported in April that both Tencent and Alibaba were in discussions to invest in DeepSeek at a valuation above $20 billion, part of a wider pattern of Chinese technology conglomerates seeding every credible domestic AI team before consolidation begins.

The open-weight release, if it arrives on schedule, matters for reasons that extend beyond Naive AI’s own business. In New York this weekend, US Treasury Secretary Scott Bessent and China’s Vice Premier He Lifeng opened pre-summit talks that include AI guardrails and the chip export restrictions that shape infrastructure costs for every Chinese AI company. A new Chinese open-weight model, freely downloadable and locally deployable, adds a dimension to that contest: not just lower-cost frontier performance, but a model built through refinement rather than compute scale, available to any organization that wants to run it without an American cloud provider in the chain.

What Dai Jifeng has not yet shown is whether the mid-training strategy works at the level his valuation implies. That answer ships with the weights.

Technology Desk

Technology Desk

The Technology Desk leads The Eastern Herald's coverage of consumer technology, online platforms, artificial intelligence, and internet policy.

Leave a Reply

Don't Miss