Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach
Published · Aug 15 · Sat Source · The Decoder

Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach

A Princeton-backed study found AI agents using Claude Opus 4.8 and GPT-5.6 Sol failed to produce publishable research. Original NeurIPS authors rated the agent-generated papers as "Reject" after six days of autonomous work.

KeywordsOpenAIAnthropicGPTClaudeStudyAIPrinceton-backedOpus

Researchers affiliated with Princeton University tested the capabilities of leading AI agents to conduct independent scientific work. The experiment provided models with Claude Opus 4.8 and GPT-5.6 Sol access to GPU resources and a budget of $3,000 in API credits over a six-day period.

The objective was to generate novel AI research papers suitable for top-tier conferences. However, when the original authors of unpublished NeurIPS submissions evaluated the agent-generated outputs, they consistently rated the work as "Reject."

These findings challenge recent assertions from major labs like Anthropic and OpenAI regarding the imminent arrival of fully autonomous AI research systems. The results suggest a significant gap between current model capabilities and the complex reasoning required for high-quality academic publication.

This study highlights the limitations of current agent frameworks in executing long-horizon, high-cognitive tasks without human oversight. While automation in research assistance is progressing, independent discovery remains beyond the reach of today's frontier models.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.