NVIDIA Releases NemotronLabs VoiceChat 11B: An Open Full-Duplex Speech-to-Speech Model with ~450 ms Turn-Taking and Live Tool Calling
Published · Aug 10 · Mon Source · MarkTechPost

NVIDIA Releases NemotronLabs VoiceChat 11B: An Open Full-Duplex Speech-to-Speech Model with ~450 ms Turn-Taking and Live Tool Calling

NVIDIA unveiled NemotronLabs VoiceChat 11B, an open full-duplex speech-to-speech model featuring approximately 448 ms latency and live tool calling capabilities for real-time interactions.

KeywordsNVIDIAReleasesNemotronLabsVoiceChatAnOpenFull-DuplexSpeech-to-Speech

NVIDIA has introduced NemotronLabs VoiceChat 11B, an open-weight model designed for full-duplex speech-to-speech interactions. The system processes audio inputs and generates spoken outputs directly, aiming to reduce the friction typically associated with text-based intermediaries in conversational AI.

The release highlights a turn-taking latency of roughly 448 milliseconds, enabling near-real-time dialogue. Additionally, the model supports live tool calling, allowing it to execute functions or access external data during a conversation without breaking the flow of speech.

By making this model open, NVIDIA positions itself to accelerate development in voice-enabled agents. This aligns with broader industry trends toward multimodal AI that can handle audio natively rather than relying on separate transcription and synthesis pipelines.

Such advancements could lower barriers for developers building customer service bots or interactive assistants. The focus on low latency and tool integration suggests a push toward more practical, actionable voice AI applications in production environments.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.