Alibaba Tongyi Launches New Reinforcement Learning Framework EAPO
Published · Apr 28 · Tue Source · 通义实验室 (CN)

Alibaba Tongyi Launches New Reinforcement Learning Framework EAPO

Alibaba Tongyi Lab has launched a new reinforcement learning framework, EAPO (Evidence-Augmented Policy Optimization), introducing an "evidence reward" mechanism. This shifts supervision from the answer down to the evidence extraction process, solving the hallucination problem of "searching correctly but answering incorrectly" in large model long-text reasoning. The model based on Qwen3-30B using this framework performed excellently in multiple authoritative long-text benchmark tests, surpassing large models with 120B parameters such as GPT-OSS and Claude-Sonnet-4.

KeywordsGPTClaudeQwenAlibabaTongyiLaunchesNewReinforcement

Alibaba Tongyi Lab has launched a new reinforcement learning framework, EAPO (Evidence-Augmented Policy Optimization), introducing an "evidence reward" mechanism. This shifts supervision from the answer down to the evidence extraction process, solving the hallucination problem of "searching correctly but answering incorrectly" in large model long-text reasoning. The model based on Qwen3-30B using this framework performed excellently in multiple authoritative long-text benchmark tests, surpassing large models with 120B parameters such as GPT-OSS and Claude-Sonnet-4.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.