Baidu Wenxin Open Sources Text-to-Image Model ERNIE-Image
Baidu Wenxin officially open-sources the text-to-image model ERNIE-Image. With only 8B parameters, it reaches open-source SOTA levels, matching commercial closed-source models like Nano Banana in capabilities such as text rendering and complex instruction following. The model can run with 24GB of VRAM, supports precise glyph generation in multiple languages including Chinese, English, Japanese, and Korean, has been launched on ComfyUI, and offers a GGUF quantization scheme. Related weights and inference code have been open-sourced on Hugging Face.
Baidu Wenxin officially open-sources the text-to-image model ERNIE-Image. With only 8B parameters, it reaches open-source SOTA levels, matching commercial closed-source models like Nano Banana in capabilities such as text rendering and complex instruction following. The model can run with 24GB of VRAM, supports precise glyph generation in multiple languages including Chinese, English, Japanese, and Korean, has been launched on ComfyUI, and offers a GGUF quantization scheme. Related weights and inference code have been open-sourced on Hugging Face.
May 9, 2026 ERNIE-Image is an open-source text-to-image model launched by the Baidu ERNIE-Image team. Architecturally, it adopts a Latent Diffusion Model (LDM), based on a single-stream Diffusion Transformer (DiT), possessing 8B DiT parameters, and equipped with a lightweight prompt …
ERNIE-Image adopts a single-stream Diffusion Transformer architecture, specifically addressing common weak points of generative models: clear and readable text, strict layout arrangement, multi-object prompt adherence, and Chinese-English bilingual instructions. The built-in lightweight Prompt Enhancer will take short …
April 14, 2026 ERNIE-Image is an open text-to-image generation model developed by the ERNIE-Image team at Baidu. It is built on a single-stream Diffusion Transformer (DiT) and paired with a …
Based on ERNIE-Image's image generation capabilities, it supports text-to-image and multi-style visual content creation, achieving high-quality generation from text to image. Balancing performance and efficiency under a lightweight architecture, it is suitable for scenarios such as design creation, content production, and inspiration exploration, providing stable high ….
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.