SenseTime Launches SenseNova U1 Infographic Enhanced Version V2
Published · Jul 6 · Mon Source · 商汤科技SenseTime (CN)

SenseTime Launches SenseNova U1 Infographic Enhanced Version V2

SenseTime's SenseNova U1 Infographic Enhanced Version V2 has been officially released. Based on the 8B-MoT architecture, the model achieves lighter, open high-information-density image generation capabilities. The new version focuses on upgrading small text clarity, complex layout structures, and visual texture, bringing AI-generated infographics closer to professional designer standards.

KeywordsSenseTimeLaunchesSenseNovaU1InfographicEnhancedVersionV2

English | Simplified Chinese

[2026.06.29] Released SenseNova-U1-8B-MoT-Infographic-V2 📊. The new infographic model improves dense small text rendering capabilities, making text edges sharper and clearer; enhances layout capabilities for complex dense images, and improves overall visual aesthetics and harmony. Additionally, the issue of black backgrounds has been fixed. Model details and visualization effects can be seen at ✨ U1 Infographic Model Series.

[2026.06.12] Released SenseNova-U1-8B-MoT-Infographic-LoRA-8step-V1.0, used for fast infographic generation. Please check the inference example script.

[2026.06.11] Released SenseNova-U1-8B-MoT-Interleaved 📖, optimized specifically for interleaved text and image generation. It shows significant improvements in narrative coherence, character and style consistency, and text-image alignment on multi-page content such as picture books, storybooks, multi-page PPTs, and text-image tutorials.

[2026.05.21] Released full parameter fine-tuning training code for SenseNova-U1.

[2026.05.15] Released SenseNova-U1-8B-MoT-Infographic 📊 model, improving infographic generation capabilities. Model details can be seen at U1 Infographic Model Series, and 100 generation cases can be seen at ✨ Infographic Showcases.

[2026.05.10] Released 🔥SenseNova-U1 Technical Report🔥, and open-sourced SenseNova-U1-A3B-MoT-SFT and SenseNova-U1-A3B-MoT model weights.

[2026.05.08] Added GGUF quantization weight support and layered loading VRAM mode, facilitating inference in single-card low VRAM environments. See Low VRAM Inference (GGUF + VRAM Mode) for details. The GGUF weights for SenseNova-U1-8B-MoT-Merger have been uploaded to 🤗 smthem/SenseNova-U1-8B-MoT-Merger-gguf. Special thanks to @smthem for contributing quantized weights to the community.

[2026.05.06] Released SenseNova-U1-8B-MoT-LoRA-8step-V1.0. Please check the inference example script.

[2026.04.30] Released preview version of the 8-step inference model SenseNova-U1-8B-MoT-8step-preview. In most cases, the image generation quality of this model is very close to the base model (see Effect Comparison and Existing Issues). To test this model, refer to the inference script, but replace the following parameters: --cfg_scale 1.0 --num_steps 8.

[2026.04.27] First release of SenseNova-U1-8B-MoT-SFT and SenseNova-U1-8B-MoT model weights.

[2026.04.27] First release of SenseNova-U1 inference code.

🚀 SenseNova U1 is a new generation of native multimodal model series, unifying multimodal understanding, reasoning, and generation within a single architecture. It represents a fundamental paradigm shift in multimodal AI: from modality integration to true unification. SenseNova U1 no longer relies on adapters to translate between different modalities, but thinks and acts natively across language and vision.

The unification of visual understanding and generation opens up huge possibilities. SenseNova U1 stands on the data-driven learning stage (such as ChatGPT) and points to the next stage—the agent learning stage (such as OpenClaw)—learning, thinking, and acting in a native multimodal way.

The core of SenseNova U1 is NEO-unify—a brand new architecture designed for multimodal AI from first principles: it completely discards Visual Encoders (VE) and Variational Autoencoders (VAE), because pixel and text information are essentially deeply correlated. Its main features are as follows:

- 🔗 End-to-end modeling of language and visual information as a unified whole.

- 🖼️ Maintaining pixel-level visual fidelity while preserving semantic richness.

- 🧠 Achieving cross-modal reasoning through native MoT, with high efficiency and few conflicts.

Based on this new core architecture, SenseNova U1 demonstrates excellent efficiency in multimodal learning:

Left image: Comparison of generation latency and average performance on OneIG (EN, ZH), LongText (EN, ZH), CVTG, BizGenEval (Easy, Hard), and IGenBench.

Right image: Comparison of generation latency and average performance on infographic benchmarks (BizGenEval (Easy, Hard), IGenBench).

🏆 Understanding and Generation reach Open Source SOTA: SenseNova U1 sets a new benchmark in unified multimodal understanding and generation, reaching the most advanced level among open-source models on various understanding, reasoning, and generation benchmarks, comparable to commercial large models.

📖 Native Interleaved Generation: SenseNova U1 can coherently produce interleaved text and image content in a single generation process using a single model, supporting scenarios such as life guides and travel diaries that require both clear expression and narrative expressiveness, condensing complex information into intuitive diagrams.

📰 High-Density Information Presentation: SenseNova U1 demonstrates strong capabilities in high-density visual information expression, able to generate content with rich structures and complex layouts, suitable for various information-intensive scenarios such as knowledge diagrams, posters, PPTs, comics, and resumes.

- 🤖 Vision-Language-Action (VLA)

- 🌐 World Modeling (WM)

In this release, we open-sourced the SenseNova U1 Lite series, with two specifications:

- SenseNova U1-8B-MoT — Dense backbone network

- SenseNova U1-A3B-MoT — MoE backbone network

The SFT model (×32 downsampling ratio) underwent four stages of training: (1) Understanding warm-up, (2) Generation pre-training, (3) Unified mid-stage training, (4) Unified supervised fine-tuning. The final model is a version obtained after one round of T2I Reinforcement Learning (RL) training on top of the base model.

Currently, these models are relatively compact in scale, but have demonstrated strong performance on various tasks, comparable to commercial models with excellent cost-performance ratio. Larger scale versions will be released in the future to further enhance capabilities.

💡

The 8B-MoT in SenseNova-U1-8B-MoT refers to ~8B understanding parameters and ~8B generation parameters. Please refer to Model Parameter Decomposition for detailed grouping details.

SenseNova-U1 Training Code

SenseNova-U1 Final Release Weights and Technical Report

🖼️ Text-to-Image (General)

🖼️ Text-to-Image (Reasoning)

🖼️ Text-to-Image (Infographic)

📸 More generation samples: See Text-to-Image Sample Collection.

✏️ Image Editing (General)

✏️ Image Editing (Reasoning)

📸 More editing samples: See Image Editing Sample Collection.

♻️ Interleaved Generation (General)

♻️ Interleaved Generation (Reasoning)

📸 More interleaved samples: See Interleaved Generation Sample Collection.

📝 Visual Understanding (General)

📝 Visual Understanding (Agent)

📸 More visual understanding samples: See Visual Understanding Sample Collection.

🦾 World Modeling

Evaluation scripts and benchmark reproduction guides are provided in evaluation.

Despite excellent performance on various tasks, the current version still has several known limitations to be improved:

Visual Understanding: The current model supports a maximum context length of 32K tokens, which may be limited in scenarios requiring longer or more complex visual contexts.

Human Generation: There are still challenges in handling fine-grained details of humans, especially when characters occupy a small portion of the image or have complex interactions with surrounding objects.

Text Generation: Text rendering sometimes suffers from spelling errors, character distortion, or inconsistent formatting, and is sensitive to prompt wording, especially in text-dense scenarios. (See Prompt Enhancement for best practices)

Interleaved Generation:

As an experimental feature, interleaved generation is still continuously evolving, and performance may not yet reach the level of dedicated Text-to-Image (T2I) processes.

Beta Status: Reinforcement Learning has not been specifically optimized for image editing, reasoning, and interleaved tasks, and current performance is comparable to the SFT model.

We list the above directions as key points for continuous iteration and look forward to continuous improvement in subsequent versions.

💡 Tip: If you encounter any issues during configuration or operation, please refer to our FAQ.

The most convenient way to experience SenseNova-U1 is through SenseNova-Studio—a 🆓 free online experience platform, requiring no installation and no GPU, allowing you to try it directly in your browser.

Note: To serve more users, U1-Fast has undergone step distillation and CFG distillation, specifically for infographic generation.

The simplest way to integrate SenseNova-U1 into your own agent or application is to use the accompanying repository SenseNova-Skills (OpenClaw) 🦞—which encapsulates SenseNova-U1 as out-of-the-box skills and provides a unified tool calling interface.

For installation and usage details, please refer to the SenseNova-Skills README.

📝 Visual Understanding

python examples/vqa/inference.py --model_path sensenova/SenseNova-U1-8B-MoT --image examples/vqa/data/images/menu.jpg --question "My friend and I are dining together tonight. Looking at this menu, can you recommend a good combination of dishes for 2 people? We want a balanced meal — a mix of mains and maybe a starter or dessert. Budget-conscious but want to try the highlights." --output outputs/answer.txt --max_new_tokens 8192 --do_sample --temperature 0.6 --top_p 0.95 --top_k 20 --repetition_pena.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.