Alibaba Tongyi Launches General Agent Evaluation Benchmark PawBench
Tongyi Lab launches the general agent evaluation benchmark PawBench, incorporating base models and runtime frameworks (Harness) into joint evaluation for the first time. PawBench v1.0 includes 150 real-world tasks and 4,050 test units, covering a cross-matrix of 9 models and 3 Harnesses. The evaluation found that Harness performance gaps can reach up to 6.4 points, and switching Harnesses for the same model can result in a score difference of up to 11.5 points.
Tongyi Lab launches the general agent evaluation benchmark PawBench, incorporating base models and runtime frameworks (Harness) into joint evaluation for the first time. PawBench v1.0 includes 150 real-world tasks and 4,050 test units, covering a cross-matrix of 9 models and 3 Harnesses. The evaluation found that Harness performance gaps can reach up to 6.4 points, and switching Harnesses for the same model can result in a score difference of up to 11.5 points.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.