Thinking Machines scientist Lilian Weng revises AI scaling laws amid concerns over the data wall.

iconKuCoinFlash
Share
AI summary iconSummary
Thinking Machines scientist Lilian Weng has updated AI scaling laws amid growing concerns about a "data wall" limiting model growth. Her research shows that repeated training data loses value rapidly, while overfitting penalties increase with model size. Weng states that scaling laws are not fixed but depend on engineering decisions. As CFT regulations tighten, liquidity in crypto markets faces new pressures. Developers must focus on system engineering to overcome data limits, especially as crypto markets evolve under stricter compliance rules.
ME AI News: According to monitoring by Beating, Lilian Weng, co-founder and chief scientist at Thinking Machines, published a detailed article systematically recalibrating the "Scaling Laws" underpinning the expansion of the AI industry. Amid widespread industry concerns that large models are hitting a wall due to the depletion of web data, Weng argues that scaling laws have not failed—but the traditional approach of brute-force compute stacking has reached a dead end. A shift is needed toward refined modeling that accounts for data constraints, repeated training, and penalties for overfitting. In the early stages of large model development, the industry evolved from the purely parameter-focused Kaplan Law to the Chinchilla Law, which advocates proportional increases in both parameters and data. Facing the impending "data wall" of dwindling high-quality web data, recent research has turned to data-constrained scenarios: Muennighoff et al. demonstrated that the value of repeatedly trained data decays exponentially; Lovelace et al. quantified in 2026 the explicit overfitting penalty introduced by the ratio of model parameters to unique data ($N/U_D$), confirming that regularization techniques such as enhanced weight decay can effectively mitigate the damage from repeated training. Weng further emphasizes that scaling laws are not immutable physical laws, but rather observation-based guidelines highly sensitive to engineering details. The estimation biases in Kaplan and Chinchilla fundamentally stem from whether embedding layer parameters are included and the exponential sensitivity of extrapolating small models. In training runs costing hundreds of millions of dollars, minor adjustments in parameter accounting or optimizer settings can lead to significant prediction errors. Therefore, large model development must not rely blindly on single theoretical formulas; instead, it must rigorously fit its own loss curves in practice, precisely calculate the ratio of parameters to unique data capacity, and use systematic engineering to break through the "data wall." (Source: BlockBeats)
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.