AGIBOT says its WITA-Omni Preview multimodal foundation model has achieved the highest score on the Daily-Omni audio-visual reasoning benchmark, outperforming models from Alibaba, Google and ByteDance.
According to the company, WITA-Omni Preview recorded an average accuracy of 85.21 percent on the third-party benchmark, ahead of Qwen3.5-Omni-Plus, Gemini 3.1 Pro Preview and Doubao Seed 2.0 Lite.
The model achieved the highest scores in audio-visual alignment, comparison, event sequencing, and the benchmark’s 30-second and 60-second video subsets. It also ranked first or joint first in six of the eight reported evaluation metrics, including tying for first place in inference.
Daily-Omni is designed to evaluate a model’s ability to understand and reason across audio and visual information in everyday situations. The benchmark contains 684 real-world videos and 1,197 multiple-choice questions covering six task categories, including audio-visual alignment, event sequencing, inference and reasoning.
Unlike benchmarks focused primarily on image-text understanding, Daily-Omni evaluates whether AI models can associate sounds with visible events, understand how situations develop over time and perform reasoning across multiple data types.
AGIBOT says these capabilities are particularly important for embodied AI systems operating in dynamic environments, where robots must determine who is speaking, associate sounds with actions, follow sequences of events and decide whether, when and to whom they should respond.
The company said WITA-Omni differs from conventional robotic interaction systems by extending the “Thinker-Talker” framework with an additional “Actor” component that treats movement and facial expressions as native outputs alongside speech.
The architecture comprises three core elements:
- the Thinker functions as the multimodal reasoning engine, processing text, images, audio and combined audio-visual inputs within a shared representation space to perform reasoning and make interaction decisions.
- the Talker generates speech in real time based on the Thinker’s internal state; and
- the Actor produces coordinated physical movements and facial expressions.
According to AGIBOT, this unified Thinker-Talker-Actor architecture enables perception, reasoning, speech, movement and facial expression generation to operate within a shared model state and timeline, allowing robots to continue observing their surroundings while preparing and delivering responses.
The company said WITA-Omni was trained on tens of millions of hours of multimodal data during continued training, expanding its capabilities beyond text and images to include audio perception and cross-modal reasoning.
To improve performance in real-world interactions, AGIBOT also developed a human-centric multimodal interaction dataset spanning thousands of hours. The dataset preserves temporal relationships between audio, video, language, body movement and facial expressions to help the model learn how human interactions unfold over time.
Training was carried out using a three-stage process consisting of supervised fine-tuning, on-policy distillation and reinforcement learning.
According to AGIBOT, supervised fine-tuning established the model’s multimodal understanding and coordinated output capabilities, while on-policy distillation transferred knowledge from a teacher model with access to additional contextual information into a student model operating without that information.
The reinforcement learning stage focused on improving interaction decisions, including response accuracy, timing, target-person selection and determining whether a response should be given. The company said it used Group Relative Policy Optimization (GRPO) to optimize the model’s policy.
AGIBOT said the resulting system is designed not only to generate accurate responses but also to determine whether a response is appropriate, when it should occur and who it should be directed toward.
The company said WITA-Omni Preview forms the foundation of its Interaction Intelligence platform, which is being developed alongside its Manipulation Intelligence and Locomotion Intelligence technologies as part of its “Three Intelligences in One” architecture.
AGIBOT said it plans to continue developing the WITA family of multimodal models and integrate the technology into its robotic platforms to enable more natural and context-aware interactions between people and embodied AI.
The Shanghai-based company develops embodied AI foundation models and robotic systems, including humanoid robots, quadrupeds, dexterous manipulation systems and commercial cleaning robots. In June 2026, AGIBOT announced that its 15,000th robot had rolled off the production line.
Main image courtesy of Pandaily.com

