AIToday
Large Language ModelsAudio & SpeechRobotics & Automation NewsPublished: Aug 3, 2026, 19:00 JST5 min read

AGIBOT's audio-visual model tops Daily-Omni benchmark at 85.21%

AGIBOT's audio-visual model tops Daily-Omni benchmark at 85.21%

Key takeaway

  • AGIBOT, a Shanghai-based robotics company, has released a multimodal AI foundation model called WITA-Omni Preview that scored 85.21 percent accuracy on the Daily-Omni benchmark—higher than competing models from Google, Alibaba, and ByteDance.

  • The model excels at understanding audio and visual information together, a skill the company argues is critical for robots that must interpret sounds, follow event sequences, and decide how to respond in real-world settings.

  • AGIBOT plans to embed this technology into its robotic systems to improve human-robot interaction.

3 Key Points

  1. What happened

    AGIBOT's WITA-Omni Preview foundation model achieved an average accuracy of 85.21 percent on the Daily-Omni audio-visual reasoning benchmark, outperforming competitors including Alibaba's Qwen3.5-Omni-Plus, Google's Gemini 3.1 Pro Preview, and ByteDance's Doubao Seed 2.0 Lite. The model ranked first or joint first in six of the eight reported evaluation metrics.

  2. Why it matters

    The Daily-Omni benchmark tests whether AI models can understand audio and visual information together in real-world situations—a capability AGIBOT says is essential for robots operating in dynamic environments, where they must associate sounds with actions, follow event sequences, and decide whether and when to respond. Unlike image-text benchmarks, Daily-Omni evaluates cross-modal reasoning across audio, video, and their combination.

  3. What to watch

    AGIBOT's WITA-Omni Preview integrates a "Thinker-Talker-Actor" architecture that generates speech and coordinated physical movements and facial expressions in real time. The company plans to integrate this technology into its robotic platforms, including humanoid robots and quadrupeds, to enable more natural human-robot interactions.

In Depth

Read the full story

AGIBOT announced that its WITA-Omni Preview multimodal foundation model has achieved the highest score on the Daily-Omni audio-visual reasoning benchmark, recording an average accuracy of 85.21 percent. The model outperformed Alibaba's Qwen3.5-Omni-Plus, Google's Gemini 3.1 Pro Preview, and ByteDance's Doubao Seed 2.0 Lite. It achieved the highest scores in audio-visual alignment, comparison, event sequencing, and the benchmark's 30-second and 60-second video subsets, and ranked first or joint first in six of the eight reported evaluation metrics, including tying for first place in inference.

Daily-Omni is a third-party benchmark containing 684 real-world videos and 1,197 multiple-choice questions across six task categories: audio-visual alignment, event sequencing, inference, and reasoning. Unlike benchmarks focused primarily on image-text understanding, Daily-Omni evaluates whether AI models can associate sounds with visible events, understand how situations develop over time, and perform reasoning across multiple data types. AGIBOT states these capabilities are particularly important for embodied AI systems (robots operating in dynamic environments) that must determine who is speaking, associate sounds with actions, follow sequences of events, and decide whether, when, and to whom they should respond.

WITA-Omni differentiates itself through a unified "Thinker-Talker-Actor" architecture. The Thinker functions as the multimodal reasoning engine, processing text, images, audio, and combined audio-visual inputs within a shared representation space. The Talker generates speech in real time based on the Thinker's internal state, while the Actor produces coordinated physical movements and facial expressions. This unified design allows perception, reasoning, speech, movement, and facial expression generation to operate within a shared model state and timeline, enabling robots to continue observing their surroundings while preparing and delivering responses.

The model was trained on tens of millions of hours of multimodal data during continued training, expanding its capabilities beyond text and images to include audio perception and cross-modal reasoning. AGIBOT also developed a human-centric multimodal interaction dataset spanning thousands of hours, which preserves temporal relationships between audio, video, language, body movement, and facial expressions. Training used a three-stage process: supervised fine-tuning (establishing multimodal understanding and coordinated output), on-policy distillation (transferring knowledge from a teacher model with additional contextual information to a student model without it), and reinforcement learning focused on improving interaction decisions including response accuracy, timing, target-person selection, and determining whether a response should be given. The company used Group Relative Policy Optimization (GRPO) to optimize the model's policy.

AGIBOT said WITA-Omni Preview forms the foundation of its Interaction Intelligence platform, being developed alongside its Manipulation Intelligence and Locomotion Intelligence technologies as part of its "Three Intelligences in One" architecture. The Shanghai-based company plans to continue developing the WITA family of multimodal models and integrate the technology into its robotic platforms—including humanoid robots, quadrupeds, dexterous manipulation systems, and commercial cleaning robots—to enable more natural and context-aware interactions between people and embodied AI. In June 2026, AGIBOT announced that its 15,000th robot had rolled off the production line.

Context & Analysis

AGIBOT's result on Daily-Omni reflects a shift in how foundation models for robotics are being evaluated. Traditional AI benchmarks emphasize image-text understanding, but Daily-Omni tests cross-modal reasoning—the ability to link sounds to visible events and understand how situations develop over time. This distinction matters because embodied AI systems (robots operating in physical environments) must interpret multiple data streams simultaneously. A robot that hears a voice but cannot connect it to the person speaking, or that misses the sequence of events unfolding around it, will struggle to interact naturally or make safe decisions.

AGIBOT's approach—training on tens of millions of hours of multimodal data and then refining the model with thousands of hours of human-centric interaction recordings—addresses this gap. The company's emphasis on temporal relationships (the timing and order of audio, video, movement, and facial expressions) suggests that learning WHEN and HOW to respond matters as much as learning WHETHER to respond. This is consistent with AGIBOT's stated design goal: not only to generate accurate responses but also to determine appropriateness, timing, and intended recipient.

FAQ

What is the Daily-Omni benchmark and what does it test?
Daily-Omni is a third-party benchmark designed to evaluate a model's ability to understand and reason across audio and visual information in everyday situations. It contains 684 real-world videos and 1,197 multiple-choice questions covering six task categories, including audio-visual alignment, event sequencing, inference and reasoning.
How does WITA-Omni differ from conventional robotic systems?
WITA-Omni extends the "Thinker-Talker" framework with an additional "Actor" component that treats movement and facial expressions as native outputs alongside speech. This unified architecture allows the Thinker (reasoning engine), Talker (speech generation), and Actor (physical movement and facial expressions) to operate within a shared model state, letting robots continue observing their surroundings while preparing and delivering responses.
How was WITA-Omni trained?
WITA-Omni was trained on tens of millions of hours of multimodal data during continued training, expanding its capabilities to include audio perception and cross-modal reasoning. The training process used three stages: supervised fine-tuning, on-policy distillation, and reinforcement learning (using Group Relative Policy Optimization to improve interaction decisions).
Robotics & Automation NewsRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleReimagine Robotics emerges from stealth with robots that learn on the job

The AI news that matters, in one minute each morning.

Sign up free