AGIBOT says its WITA-Omni Preview multimodal foundation model has achieved the highest score on the Daily-Omni audio-visual reasoning benchmark, outperforming models from Alibaba, Google and ByteDance.
According to the company, WITA-Omni Preview recorded an average accuracy of 85.21 percent on the third-party benchmark, ahead of Qwen3.5-Omni-Plus, Gemini 3.1 Pro Preview and Doubao Seed 2.0 Lite.
The model achieved the highest scores in audio-visual alignment, comparison, event sequencing, and the benchmark’s 30-second and 60-second video subsets. It also ranked first or joint first in six of the eight reported evaluation metrics, including tying for first place in inference. [Read more…] about AGIBOT’s foundation model tops benchmark test for audio-visual reasoning

