News: 0184892242

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

ByteDance Is Training a 10-Trillion-Parameter Model To Chase the Frontier

(Friday August 07, 2026 @05:00PM (BeauHD) from the bigger-is-better dept.)


ByteDance is reportedly [1]training an AI model with roughly 10 trillion parameters as it tries to close the gap with leading frontier systems such as Anthropic's Mythos. The model is still in early pre-training, and its eventual performance will depend on more than scale alone, but the project underscores how aggressively Chinese firms are pushing frontier AI despite limits on access to advanced chips. The Next Web reports:

> The size is itself the statement. At roughly 10 trillion parameters, the model would be more than three times as large as Moonshot's Kimi K3, which sits among the biggest Chinese models today at about 2.8 trillion. [...] Parameter count is not everything, of course. Bigger models are not automatically better, and the industry has learned that data quality, training technique and efficiency often matter as much as raw scale. Even so, committing the compute to train a model this size is a declaration in its own right, a signal that ByteDance wants to compete at the very top rather than ship a capable also-ran.



[1] https://thenextweb.com/news/bytedance-10-trillion-parameter-model-mythos



More Trustworthy than American Models (Score:1)

by logjon ( 1411219 )

I use Chinese models to develop investing strategies because they'll be more ruthless for the end user's benefit. They'll acknowledge that America stopped funding civilization with Reagan, and that some degree of social and civilizational collapse has to be accounted for when investing in real estate for instance. That it's worse in some places than others. That some locations are full of people who are bound and determined to let corporations suck up and poison all their water in the hopes that it kills mi

Re: More Trustworthy than American Models (Score:1)

by dfarrow ( 1683868 )

Chinese models sound sexy as hell...

Why? (Score:5, Interesting)

by pla ( 258480 )

Current top models are already undertrained by multiple orders of magnitude vs their number of parameters. Not for lack of will or money, but for lack of data - We're already training top models with effectively the entire digitally-available corpus of human knowledge and need about a thousand times more.

That said, the huge Chinese models we've seen lately aren't single models, they're MOEs - Effectively a hive-mind of many smaller models working in concert. Bytedance hasn't mentioned that detail yet but it would somewhat mitigate the training issue. This is still absurdly over-spec'd though. There's simply no way to train it any more than all the rest starving for data.

Re: (Score:2)

by Mirnotoriety ( 10462951 )

We are The Borg. Lower your shields and surrender your ships. We will add your biological and technological distinctiveness to our own. Your culture will adapt to service us. Resistance is futile.

Re: (Score:2)

by postbigbang ( 761081 )

If you try to map this to evolution of brains, consider the biggest problem.

Where garbage in == garbage out, much more garbage in == much more garbage out. Size of a model doesn't inherently make it more useful, just as a huge volume of brain mass doesn't make you hyper-intelligent.

This isn't even able to be mapped to Nyquist thinking, or von Neumann. Bigger isn't better; better is better.

Re: (Score:2)

by SpinyNorman ( 33776 )

I don't know if the Chinchilla scaling law (which is what I assume you are referring to) includes RL-based post-training, or even non-text input tokens for that matter?

The base model certainly makes a difference, but for areas like math, coding and hacking (to compete with and counter Mythos) the power of the end model mostly comes from the RLVR post-training, and these are all areas where they can generate synthetic data, as well as pay for human-generated data if they care/need to.

So much for the US attem

Re: (Score:2)

by zlives ( 2009072 )

starve? no sir, merely rerouting of transport to make sure certain new ventures for certain family members were tied in the supply chain. FTWof few

Re: (Score:2)

by allo ( 1728082 )

Probably the same reason why GPT-3.5 was 1.4T and now much better models fit into 4B (or possibly less): Scaling up by increasing the number of parameters is simple, getting the same in fewer parameters is harder.

One can hope they optimize for the next models to be smaller, because 10T is surely larger than the level of knowledge/intelligence they will achieve in a reasonable training time. And 10T will be very costly to host as well. I don't think anyone will even pay the self-cost price for inference with

Re: (Score:3)

by dinfinity ( 2300094 )

> Not for lack of will or money, but for lack of data - We're already training top models with effectively the entire digitally-available corpus of human knowledge and need about a thousand times more.

This has to be nonsense, fundamentally speaking. "Data" is not some kind of fungible resource. Humans are 'trained' (in school) using a much smaller set of data for many subjects and given that we haven't reached AGI-levels of intelligence for those yet data clearly cannot be the bottleneck.

You can argue that we are lacking data for understanding/mastering of specific aspects of the universe (touch, motor control, etc.), but for many, many things a larger network might still allow capturing more subtle rela

Re: (Score:2)

by pla ( 258480 )

We don't know the exact answer yet, but IMO, you're two-halves right.

The key difference is, humans aren't "trained" primarily via a clean library of tagged data (which in itself is a learned skill). We experience the world as a more-or-less continual stream of noisy data from our several senses. We self-train how to interpret (and primarily, filter) sensory data; and more importantly, how to derive new training data from it.

To be clear, I'm not saying we have so many neurons to support input filtering (

if some is good more is better (Score:2)

by awwshit ( 6214476 )

More is not always better.

A flattening curve (Score:4, Insightful)

by memory_register ( 6248354 )

We are reaching the flat part of the curve, where a change in the order of magnitude in model size is yielding only slight, incremental benefit. We are reaching the hard limits of what large language models can do.

Re: (Score:2)

by ledow ( 319597 )

AI is, and always has been, exclusively nothing but exponentially-increasing requirements for logarithmically-plateauing results.

Any mathematician will tell you what that means in the end.

model measuring contest (Score:2)

by Local ID10T ( 790134 )

It's not the size of the model, but what you do with it. or so I am told...

How much? (Score:2)

by Shakes Fist ( 10502847 )

How much money and resources is this new hallucination engine going to consume?

I have a very strange feeling about this...
-- Luke Skywalker