Meituan Open Sources Digital Human Video Model LongCat-Video-Avatar 1.5
Published · May 22 · Fri Source · 龙猫LongCat (CN)

Meituan Open Sources Digital Human Video Model LongCat-Video-Avatar 1.5

Meituan's LongCat team has officially open-sourced the LongCat-Video-Avatar 1.5 digital human video model, moving from open-source SOTA to commercial-grade application. The model upgrades the Whisper-large audio encoder, builds a high-quality multi-scenario data system, and introduces frame-level GRPO preference alignment, achieving comprehensive improvements in lip-syncing, physical plausibility, long-video stability, and multi-person interaction. The model uses DMD distillation to achieve 8-step generation, improving efficiency by approximately 15 times.

KeywordsMeituanOpenSourcesDigitalHumanVideoModelLongCat-Video-Avatar

Meituan's LongCat team has officially open-sourced the LongCat-Video-Avatar 1.5 digital human video model, moving from open-source SOTA to commercial-grade application. The model upgrades the Whisper-large audio encoder, builds a high-quality multi-scenario data system, and introduces frame-level GRPO preference alignment, achieving comprehensive improvements in lip-syncing, physical plausibility, long-video stability, and multi-person interaction. The model uses DMD distillation to achieve 8-step generation, improving efficiency by approximately 15 times.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.