
Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut it
Former OpenAI researcher Andrew Ho predicts $100 billion will shift toward training data, arguing that scaling alone is insufficient as LLMs show signs of specialization rather than general versatility.
Former OpenAI researcher Andrew Ho and Cambridge researcher Adam Hunt argue that current large language model development faces a bottleneck. They observe that while models improve in specific domains like coding and mathematics, performance in other areas is stagnating or regressing despite increased compute resources.
Ho suggests that the industry's reliance on scaling laws is reaching its limits. Consequently, he forecasts a significant capital shift, estimating that $100 billion will flow into acquiring and curating high-quality training data rather than solely expanding hardware infrastructure.
This perspective highlights a potential pivot in AI investment strategies. If data quality becomes the primary differentiator, companies may prioritize proprietary datasets over raw compute power to achieve broader model capabilities and prevent further specialization.
The observation challenges the prevailing notion that bigger models automatically yield more versatile intelligence. It suggests that future progress depends heavily on the diversity and richness of the information fed into these systems during the training phase.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.