ByteDance is training an AI model with as many as 10 trillion parameters, according to sources familiar with the matter first reported by the Financial Times on August 7, 2026. The model, which would be more than three times the size of Moonshot AI's Kimi K3 at 2.8 trillion parameters and potentially larger than industry estimates for Anthropic's Mythos 5 at around 8 trillion, is in pre-training and not expected to be publicly available for at least six months. ByteDance founder Zhang Yiming reportedly told his AI teams earlier this year to stop chasing quick wins through model distillation and to focus on building world-class infrastructure from scratch. The directive resulted in this training run, the largest ever initiated by a Chinese company, and one that signals a fundamental strategic shift in how ByteDance thinks about its position in the global AI race.
What Actually Happened
The report, carried by The News International citing the Financial Times, describes a model in early pre-training using a Mixture of Experts (MoE) architecture. In an MoE system, the full parameter count does not activate for every token; instead, a routing mechanism activates only a subset of the total parameters for any given inference pass, making 10 trillion parameters physically trainable and economically deployable in ways that a dense 10-trillion-parameter model would not be. GPT-4 was widely estimated at roughly 1.8 trillion parameters using MoE; Anthropic's Mythos series uses a dense architecture that industry sources estimate at around 8 trillion parameters for the top model in the family, though Anthropic does not officially disclose parameter counts for any of its models. ByteDance's 10 trillion MoE model would activate a fraction of those parameters per inference pass, but the total parameter budget allows for a representation capacity that rivals or exceeds what any publicly known model currently offers.
ByteDance has made the hardware investment to support this ambition. According to South China Morning Post, the company planned to spend $14 billion on Nvidia AI chips in 2026, up from roughly $12 billion in 2025, as part of a total AI capital expenditure plan exceeding $23 billion for the year. That GPU budget, if deployed toward a single pre-training run, would represent one of the largest individual model training investments in history, approaching the scale of expenditure that OpenAI applied to GPT-5 training. The Nvidia purchases are made possible partly by a relaxation in US AI export controls during the Trump administration, which allowed ByteDance to source H200 and B200 series chips that were previously restricted for Chinese entities under Biden-era semiconductor rules.
ByteDance's consumer AI footprint gives the company a clear commercial motivation for frontier model investment. Doubao, the company's primary AI chatbot, has reached 324 million monthly active users, making it the dominant AI assistant in China by a wide margin. However, Doubao generates less than 1 million yuan in daily revenue while consuming tens of millions in compute resources daily, a structural imbalance that has led some analysts to warn about the sustainability of ByteDance's AI subsidy model. A proprietary frontier model that replaces purchased inference from third-party providers with internally generated compute would dramatically improve that equation, provided the model's capability justifies the switching cost from the external models currently powering Doubao's most demanding tasks.
Why This Matters More Than People Think
The conventional narrative about China's AI position has been built on a specific assumption: that US export controls on advanced chips would deny Chinese companies the compute resources required to train frontier models, creating a hardware ceiling that would keep Chinese AI development two to three years behind US labs regardless of software engineering talent. ByteDance's 10-trillion-parameter training run challenges that assumption directly. The company has sourced enough Nvidia compute to attempt a training run at a scale that matches or exceeds what US frontier labs are working on. If the training succeeds and the resulting model performs at the level its parameter budget suggests it should, the hardware-ceiling thesis will need revision across every major AI lab's planning assumptions.
The strategic significance extends beyond ByteDance itself. Zhang Yiming's reported directive to abandon distillation shortcuts and invest in original frontier research is a signal that China's leading AI labs believe they can compete on capability rather than just on cost and deployment speed. Distillation, the technique of training smaller models to mimic the outputs of larger ones, has been the dominant strategy for Chinese AI development since DeepSeek demonstrated that highly efficient smaller models could match much larger ones on many benchmarks. That strategy has real merit for cost-competitive deployment. It does not produce models that can set new capability frontiers. A ByteDance decision to fund original frontier research at this scale suggests the company has concluded that the distillation ceiling is real and that the only way to advance beyond it is to invest in training runs that US labs have historically been the only entities capable of funding.
The timing matters as well. ByteDance is making this investment at a moment when the US is actively debating further restrictions on AI chip exports to China. If the current Nvidia-sourcing window closes due to policy changes, the training run underway represents one of the last opportunities for a Chinese company to access the compute scale required for frontier model training using the world's best chips. The model being trained now could define ByteDance's AI capability ceiling for several years if export restrictions tighten materially in the next congressional cycle.
The Competitive Landscape
Anthropic is the explicit target. Industry estimates, consistently unconfirmed by Anthropic, place Mythos 5 at around 8 trillion parameters using a dense architecture, making it currently the largest model by parameter count that is regularly deployed for commercial inference. Fable 5, the more widely available model in the same family, is estimated at around 5 trillion parameters. Claude Code, which is the most commercially active deployment of Anthropic's technology, runs primarily on Fable 5. A ByteDance MoE model at 10 trillion total parameters would not necessarily outperform Mythos 5 on every benchmark, because MoE models activate only a subset of parameters per inference and their performance depends heavily on how well the routing mechanism is trained. But the parameter budget represents a real capability investment that would put ByteDance in a different competitive tier from where it has operated historically.
OpenAI's GPT-5 and its successors represent the other pole of the competitive landscape. OpenAI has not disclosed GPT-5's architecture or parameter count, but the model has set new benchmarks across coding, reasoning, and multimodal tasks since its release. ByteDance's Doubao currently relies on a mix of internal models and, for some tasks, API access to external providers. A proprietary frontier model would eliminate the dependency on external providers entirely, which matters for both cost and for the kind of fine-tuning and product integration that gives a consumer AI assistant durable competitive advantages over time.
DeepSeek remains the most interesting Chinese competitor. The Hangzhou-based lab has demonstrated that smaller, more efficiently trained models can match frontier capability on many benchmarks, and its R1 reasoning model has been adopted by hardware companies, including Unitree Robotics, for embedded AI applications. DeepSeek's approach and ByteDance's approach are now diverging sharply: DeepSeek is optimizing for efficiency at a given capability level, while ByteDance under Zhang Yiming's directive is betting that raw scale still produces qualitative capability jumps that efficiency-focused training cannot replicate. As Slashdot noted in its coverage, the two strategies represent a genuine empirical disagreement about whether scaling laws continue to hold at the 10-trillion-parameter range, a question the industry has not yet resolved.
Hidden Insight: The Scale Wars Are Not Over
The conventional wisdom after DeepSeek's R1 release was that the scaling era was effectively over, that the returns to additional compute were diminishing to the point where clever engineering and efficient training could substitute for raw parameter count. ByteDance's 10-trillion-parameter training run is a direct empirical test of that thesis. Zhang Yiming's reported instruction to his teams reads as a deliberate rejection of the efficiency-first consensus. If the resulting model produces capability jumps that efficient smaller models cannot match, scaling is not dead. If it does not, the thesis that compute-efficient training has made raw scale obsolete will gain the strongest empirical support yet from the most consequential test case China's AI industry has yet run.
The compute economics of this decision are worth examining carefully. A 10-trillion-parameter MoE model trained on a cluster of Nvidia H200s at ByteDance's scale would consume training compute that approaches or exceeds $1 billion in raw inference cost equivalent, even before counting the hardware procurement. For a company whose flagship AI product generates less than 1 million yuan in daily revenue against tens of millions in daily compute cost, this is a real additional burden, approaching an estimated $1 billion in training compute cost alone. ByteDance can absorb the cost because its social media and advertising businesses generate strong and growing cash flow. But the implied assumption is that a proprietary frontier model will eventually generate revenue that justifies the training spend, which requires either dramatically improving Doubao's monetization or finding enterprise and API customers willing to pay Anthropic-scale prices for access to a Chinese frontier model.
The bear case here deserves careful examination. Critics argue that even if ByteDance successfully trains a 10-trillion-parameter model, the inference cost of running it at Doubao's 324-million-user scale makes commercial deployment economically irrational. MoE architectures are more inference-efficient than their total parameter count suggests, but they still require compute per query that scales with model size, and the gap between training a frontier model and deploying it profitably at consumer scale is where most Chinese AI companies have struggled. Skeptics point out that ByteDance's GPU purchases are already straining its compute budget before this training run reaches completion, and that adding a frontier inference workload on top of its existing Doubao serving infrastructure could push compute costs beyond what even ByteDance's advertising revenue can sustain. The path from "we trained a 10T model" to "we are profitable on it" has no clear precedent in the industry.
The longer-term implication, however, may matter more than the near-term economics. ByteDance's SeeDance video generation model already competes with top Silicon Valley tools on output quality. If a 10-trillion-parameter foundation model gives ByteDance a capability base that powers not just Doubao but also multimodal generation, coding, enterprise AI, and eventually robotics inference, the revenue diversification changes the unit economics calculation entirely. A frontier model is not a single product; it's infrastructure that can underpin dozens of products. The question is whether ByteDance can build the product surface area fast enough to monetize the capability before the compute costs exhaust the company's capacity to subsidize the development period.
What to Watch Next
The 90-day indicator is whether ByteDance releases any benchmark results or preview demonstrations of the model before year-end 2026. Pre-training a 10-trillion-parameter model takes three to six months; if the run started in July or August, early benchmark leaks could appear by October or November. Watch for any announcements from ByteDance's research division about new model releases under the Doubao or Seed brand names, which the company uses for its research model family. A new Seed model announcement in Q4 2026 would almost certainly be derivative of this training run.
The 180-day indicator is the state of US chip export controls at the end of 2026. The current Nvidia sourcing window has been unusually permissive by historical standards, enabled by Trump administration policy changes that relaxed Biden-era restrictions. If those restrictions tighten again, either through congressional action or executive order, ByteDance's ability to procure the additional compute required for the fine-tuning, post-training, and RLHF stages of model development would be constrained in ways that could materially degrade the final model's performance relative to its parameter count. The training run happening now is only the beginning of a compute investment cycle that extends well beyond pre-training.
Watch also for any changes to Doubao's API pricing or enterprise product offerings in Q1 2027. If ByteDance launches an enterprise API for a new model generation at prices competitive with Anthropic's Claude API or OpenAI's GPT-5, it would signal that the frontier training run succeeded at a capability level the company is willing to charge enterprise prices for. A price below Anthropic and OpenAI's tiers would signal either that the model is competitive and ByteDance is using aggressive pricing as a market-entry strategy, or that capability questions remain that prevent premium pricing. The launch economics will tell the story of whether scale still rules.
ByteDance is spending $14 billion on Nvidia chips this year and training a model that rivals Anthropic's best: the two-year hardware ceiling thesis just took the most serious empirical challenge it has faced.
Key Takeaways
- 10 trillion parameters, 3x larger than Kimi K3: ByteDance's training run uses a Mixture of Experts architecture that makes this parameter budget physically trainable, targeting a capability tier that potentially exceeds industry estimates for Anthropic's Mythos 5 at roughly 8 trillion parameters.
- $14 billion in Nvidia GPU spending in 2026: ByteDance's chip procurement budget, part of a $23 billion total AI capex plan, is among the largest of any non-US company and was made possible by relaxed export controls under current US policy.
- Zhang Yiming's "no shortcuts" directive: ByteDance's founder explicitly told AI teams to abandon distillation-based development and invest in original frontier research, a strategic reversal of the approach that defined Chinese AI development since DeepSeek's R1 breakthrough.
- Doubao at 324 million MAU generates less than 1M yuan daily: The structural imbalance between Doubao's user base and its revenue makes a proprietary frontier model economically necessary to reduce reliance on expensive third-party inference, but also makes the additional training compute burden materially risky.
- Export control window is the critical variable: ByteDance is training now partly because Nvidia chips are available; if US policy tightens before fine-tuning and post-training stages are complete, the final model's capability could be materially degraded relative to its architecture's potential.
Questions Worth Asking
- If ByteDance's 10-trillion-parameter MoE model produces capability jumps that efficient smaller models cannot match, does that validate scaling laws at the frontier and force every well-funded lab to restart a compute arms race that most of the industry assumed was winding down?
- ByteDance generates the training compute budget from TikTok and Douyin advertising revenue: what happens to Chinese frontier AI development if that revenue stream is disrupted by further TikTok regulatory action in the US, which currently accounts for a large portion of ByteDance's global ad income?
- If a Chinese lab demonstrates a model that genuinely rivals or surpasses Anthropic's Mythos 5, does that change the calculus of US AI export controls, potentially triggering tighter restrictions that would paradoxically give Chinese labs an incentive to accelerate domestic chip development as a substitute?