
HiDream Launches Commercial Image Generation Model HiDream-O1-Image-1.5
HiDream launches commercial image generation model HiDream-O1-Image-1.5, ranking 3rd globally and 1st in China on the Artificial Analysis text-to-image leaderboard with 1265 ELO, second only to OpenAI's GPT Image series, surpassing mainstream models from Google, NVIDIA, and ByteDance. Based on the native full-modal architecture UiT, the model excels in semantic adherence, text rendering, complex layout, and multi-subject consistency.
Recently, HiDream.ai (HiDream.ai) launched the new commercial version image generation model HiDream-O1-Image-1.5, achieving SOTA again. On the Text to Image Leaderboard of the globally renowned independent AI model evaluation and analysis platform Artificial Analysis, it rose to the first place for Chinese image generation models, becoming the Chinese large model company with scores second only to OpenAI, surpassing mainstream image generation models from domestic and international giants such as Google Nano Banana 2 (Gemini 3.1 Flash Image Preview), NVIDIA Cosmos3-Super-Text2Image, and ByteDance's Seedream 4.0.
Half a month ago, HiDream's HiDream-O1 series open-source model HiDream-O1-Image-Dev-2604 just topped the global open-source model ranking on the text-to-image leaderboard. Weeks later, HiDream-O1-Image-1.5 entered the global top three of the text-to-image large model leaderboard again. Continuous topping not only verifies HiDream's hard-core strength in image generation large models, but also marks that it has firmly established itself in the global first-tier camp in the competition for visual generation large models.
Artificial Analysis's Text to Image Leaderboard adopts an anonymous comparison, user voting, and ELO dynamic ranking mechanism, minimizing the impact of brand awareness on evaluation results, closer to real user preference judgments in open generation scenarios. Under this professional evaluation system, HiDream-O1-Image-1.5 achieved 1265 ELO in over 4000 sample comparisons. The performance of HiDream-O1-Image-1.5 not only reflects the model's competitiveness in image quality, but also reflects improvements in comprehensive capabilities such as semantic adherence, complex scene generation, text rendering, and multi-subject control.
The SOTA achievement of HiDream-O1-Image-1.5 is not just another global leaderboard lead for a leading Chinese large model company. It marks that HiDream has pioneered the industry in advancing the innovative native full-modal architecture Unified Transformer (UiT) from "tech verification" to "production verification". It is a key step for HiDream to convert underlying architecture advantages into visual productivity tools: the open-source version proved that the pixel-level native full-modal architecture can run through open evaluations and the developer community, while the HiDream-O1-Image-1.5 commercial version further targets higher-requirement commercial scenarios such as advertising marketing, brand design, e-commerce visuals, game content, film storyboards, and IP creation, comprehensively demonstrating enhanced image quality, text rendering, complex layout, multi-subject consistency, and visual narrative capabilities.
Next, what truly deserves attention is its performance in real content production tasks.
01 Writing, Layout, Storyboarding: HiDream-O1-Image-1.5 Shows All-Around Image Generation Capabilities
1. Portrait Photography Generation Examples: Photographic Quality and Multi-Style Expression
In portrait generation scenarios, HiDream-O1-Image-1.5 demonstrates stable photographic quality and multi-style adaptation capabilities. From magical lighting and shadows, dual-person interaction to character close-ups, the model performs naturally in details such as skin texture, clothing texture, limb relationships, and environmental blur; even facing complex compositions such as wide-angle, low-angle, and indoor warm light, it maintains coordination of character proportions, spatial perspective, and image narrative. It reflects strong delivery capabilities for high-requirement scenarios such as commercial portraits, brand visuals, and film storyboards.
2. Animal Generation Examples: Fine Modeling of Motion Forms and Natural Environments
In animal generation scenarios, HiDream-O1-Image-1.5 demonstrates fine modeling capabilities for subject forms, motion states, and natural environments. It maintains realism and visual impact in difficult scenes such as animal structure, fur texture, dynamic performance, complex lighting, and underwater refraction, reflecting production-level delivery capabilities for scenarios such as nature imagery, brand visuals, game assets, and creative content production.
3. Natural Landscape Generation Examples: Fine Capture of Spatial and Light/Shadow Changes
In natural generation scenarios, HiDream-O1-Image-1.5 demonstrates precise control capabilities for large scene spatial layers, light/shadow changes, and environmental atmosphere. It maintains depth, cinematic feel, and detail performance in complex landforms and multi-light source scenes such as snow mountains and lakes, desert caravans, and crystal caves, reflecting stable delivery capabilities for complex commercial scenarios such as travel visuals, film concept art, game scenes, and brand communication.
4. Multiple Artistic Styles: Precise Style Understanding and Visual Expression
In multi-style artistic generation scenarios, HiDream-O1-Image-1.5 demonstrates excellent style understanding, semantic adherence, and visual expression capabilities. It can accurately switch between styles such as Japanese illustration, anime combat, cartoon posters, and Chinese martial arts, while maintaining unity in character design, composition relationships, action rhythm, and image atmosphere. It also possesses strong stability in complex poses, dynamic effects, and basic text rendering. It can provide efficient production support for IP creation, comic storyboards, game art, and brand creative visuals.
5. E-commerce Poster Generation Examples: Seamless Fusion of Complex Scenes and Text Information
In e-commerce poster generation scenarios, HiDream-O1-Image-1.5 demonstrates comprehensive control capabilities for product subjects, layout structures, and text information. It can quickly match visual styles for different categories and naturally integrate products, scenes, decorative elements, and marketing copy; in Chinese-English mixed typesetting, multi-level selling points, and complex layout tasks, it still maintains high text readability, image integrity, and commercial quality, significantly improving efficiency in advertising marketing, e-commerce launches, social media seeding, and brand material production.
6. IP Character Design: Multi-View Generation and Character Consistency
In IP character design scenarios, HiDream-O1-Image-1.5 demonstrates stable control capabilities for character settings, expression changes, and multi-view consistency. It can generate multi-angle views and various emotional expressions around the same character, maintaining unity in facial features, hairstyle, clothing, and overall style, presenting rich personality and expressiveness. It can significantly improve efficiency in IP setting, character three-view, animation pre-production, art assets, and brand mascot development.
7. Multi-Panel/Storyboard Design: Stable Narrative Understanding and Continuous Image Generation
In multi-panel and storyboard design scenarios, HiDream-O1-Image-1.5 demonstrates understanding capabilities for continuous narrative, image sequence, and information hierarchy. It can generate logically coherent storyboard images in multi-image content such as tool processes, task progression, children's picture books, and adventure stories, while maintaining unity in characters, scenes, and visual styles; it also possesses strong organization capabilities for panel layout, numbering, titles, and key text. It can provide efficient support for film storyboards, comic creation, ad scripts, educational content, and short video script visualization.
8. Multi-Level Complex Text Rendering Capabilities: Comprehensive Generation Capabilities for Multi-Language and Multi-Structure
In multi-level complex text rendering tasks, HiDream-O1-Image-1.5 demonstrates comprehensive generation capabilities for multi-language text, information structure, and visual scenes. It can naturally embed content such as posters, plans, structure diagrams, classroom whiteboards, live interfaces, and data dashboards into corresponding scenes, while balancing layout order, image-text relationships, and overall aesthetics; facing complex requirements such as Chinese-English mixed typesetting, numeric formulas, chart information, and multi-level titles, it still maintains good readability and layout stability, expanding its practical value in scenarios such as ad design, office collaboration, e-commerce detail pages, and education training.
02 Native Full-Modal Enters Production Verification Stage, HiDream-O1-Image-1.5 Continuously Amplifies UiT Architecture Advantages
The performance of HiDream-O1-Image-1.5 further proves HiDream's architecture innovation advantages and rapid iteration capabilities on the native full-modal route. The HiDream-O1 series (8B open-source version, Pro version, to 1.5 commercial version) has formed a clear and efficient capability evolution curve.
Traditional text-to-image models usually adopt a modular path of "Text Encoder + VAE + DiT / Diffusion Model", whose form is more like a tree growing with continuous branching: text has its own tokenizer, images and videos have their respective encoders/decoders, and audio, action, and spatial relationships are often processed along different paths, requiring multiple information conversions between modules. In complex tasks such as text-heavy layout, UI pages, multi-subject generation, multi-reference image control, and multi-storyboard narrative, it is also easier to bring detail loss, semantic misalignment, and structural instability.
HiDream-O1 native full-modal architecture takes another route: true "native full-modal" is not splicing after each modality grows, but blended seamlessly at the model's bottom layer from the native initial stage like "childhood sweethearts". The HiDream-O1 Image series models remove the VAE and independent text encoder in the traditional path, mapping original signals such as image pixels, text tokens, video voxels, as well as audio, action, and spatial relationships into the same shared Token space, interacting directly with the same UiT—pixel-level unified Unified Transformer, completing understanding, generation, and reasoning in a unified representation system.
Below is a set of comparison effect images released by the Artificial Analysis official account on the X platform:
This is also the key to HiDream-O1's continuous advancement in tasks such as complex image-text fusion, text rendering, multi-subject consistency, and storyboard narrative. When all modalities are truly connected at the bottom layer, the model can possibly move towards true "Any to Any": any input supports any output. This is not only a capability upgrade for image generation models, but also a basic capability required for world models—understanding, generating, and predicting different states of the real world in a unified architecture. The rapid advancement of HiDream-O1-Image-1.5 is exactly a solid verification of the scalability of the native full-modal route.
03 Continuous Architecture Innovation, Building Native Full-Modal World Model
HiDream has always believed that images are an important entry to video generation and full-modal world modeling. An image carries the subject, space, material, light/shadow, text, and relationships of a moment in the real world; only by stably understanding and generating these states can the model possibly further process motion, causality, camera, and narrative in continuous time.
The strong performance of HiDream-O1-Image-1.5 indicates that the route based on pixel-level native unified architecture is pushing the competition of image generation models from "larger parameters" and "better looking images" to a new stage determined jointly by architecture capability, production efficiency, and workflow value. It not only improves single image generation effects, but also provides more stable underlying capabilities for multi-image consistency, storyboard generation, video first frame, image editing, and even future long video generation. It further proves the strength of Chinese large model enterprises participating in global top model competition, and verifies the feasibility of the UiT native unified architecture as a solid foundation for next-generation multimodal models.
Facing the future, HiDream will continue to advance model iteration along the native full-modal technology route, accelerate the fusion of multimodal capabilities such as images, videos, and actions, and promote the deep landing of generative artificial intelligence technology into real application scenarios of full-modal intelligent agents such as content creation, commercial marketing, film creation, and game production. From the entry of single image generation to continuous world modeling, HiDream is using continuous underlying architecture innovation.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.