
KI-Pioneer Sutton calls synthetic data a "big mistake" in the face of an infinitely complex world
Turing Award winner Richard Sutton argues synthetic data is a 'big mistake' for scaling LLMs, claiming simulations are too limited compared to the real world's complexity.
Richard Sutton, a Turing Award recipient, has voiced strong criticism against the use of synthetic data for training large language models. He contends that artificial data generation cannot adequately represent the infinite complexity found in the real world.
This critique addresses a growing trend in the AI industry where developers rely on machine-generated text to expand training datasets. Sutton warns that this strategy may create bottlenecks, as human expertise remains essential for validating and creating high-quality information.
The expert describes simulations as 'microscopic' relative to actual reality. His comments suggest that current scaling methods might hit diminishing returns if they do not engage more directly with complex, real-world environments.
Sutton's stance adds to the broader debate on data sustainability in machine learning. As models become more powerful, the quality and origin of training data remain critical factors for future development.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.