Baichuan Intelligence Launches Medical Enhanced Large Model Baichuan-M4
Published · Jun 24 · Wed Source · 百川智能 (CN)

Baichuan Intelligence Launches Medical Enhanced Large Model Baichuan-M4

Baichuan Intelligence, in collaboration with Tsinghua University, has launched the new generation medical enhanced large model Baichuan-M4. It tops three lists on HealthBench, ranking first in the world, with a hallucination rate of only 3.3%, the lowest in the industry. Baichuan-M4 possesses four core capabilities: deep active consultation, full-course memory, evidence anchoring, and Agent scheduling. Orchestrated by Baichuan-Harness to autonomously arrange diagnosis and treatment processes, it achieves a leap from answering questions to seeing patients.

KeywordsAgentBaichuanIntelligenceLaunchesMedicalEnhancedLargeModel

IT Home, June 22 news: Baichuan Intelligence and the Tsinghua University research team jointly released the new generation medical enhanced large model Baichuan-M4 today.

The model ranks first in the world simultaneously on the HealthBench and its Hard and Professional lists, comprehensively surpassing GPT-5.5, Claude Opus 4.7, and DeepSeek-V4-Pro, with a hallucination rate as low as 3.3%.

On the medical evaluation HealthBench proposed by OpenAI, M4 scored 68.6 comprehensively, ranking first in the world, leading the second-place GPT-5.5 by over 10 points; on the Hard subset, which most tests complex clinical decision-making, M4 leads by 15.9 points.

M4 will actively ask follow-up questions about the nature and causes of symptoms, prioritizing the identification and exclusion of critical and severe conditions, rather than passively waiting for users to provide complete information, and will not skip key medical history questions that need to be asked just to give an answer quickly.

Baichuan Intelligence introduced that the company drew on the OSCE (Objective Structured Clinical Examination) method long used in medical education, and collaborated with over 150 frontline doctors to build a dynamic consultation evaluation system SCAN-bench. It does not test static memory, but uses real clinical experience as the scoring standard, fully simulating the entire process from patient reception to diagnosis by doctors through multi-round, dynamic methods.

In this evaluation, M4 scored 79.0 for initial consultation and 74.7 for follow-up consultation, both significantly leading GPT-5.5, DeepSeek-V4-Pro, and Claude Opus 4.7.

In addition, Baichuan-M4 introduces "Full-Course Memory," connecting historical medical records, multi-round consultations, laboratory test trends, and medication feedback, allowing the model to always know who the patient is, what previous diseases they have had, and how various indicators change during multiple conversations, without having to start from zero each time.

In the long-context clinical memory evaluation, M4 achieved a score of 86.9, the highest among peers, an increase of 21.1 points compared to the previous generation M3.

Baichuan also pioneered "Evidence Anchoring," requiring every medical conclusion generated by the model to correspond precisely to specific paragraphs in original papers or guidelines, rather than just marking which document it is cited from. Relying on a six-source evidence-based paradigm, the model only searches within authoritative medical sources and does not scrape data from the open network.

On top of this, M4 further breaks down authoritative guidelines, expert consensus, and real diagnosis and treatment processes into standardized, reusable clinical pathway units. There are currently over 1000 units covering more than 200 diseases, each defined and verified by senior clinical experts.

On the evidence-based medicine evaluation Baichuan-EBM built by Baichuan, M4's evidence-based citation accuracy reached 90.0, GPT-5.5 was 54.7, and OpenEvidence was 55.9.

IT Home attaches the technical report link below:.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.