Who is speaking?
Connect voices to people using audio and visual cues.
A benchmark for social interaction
Who speaks. When to respond. What to say.
Evaluating multimodal models in social interaction through audio and video.
Social interaction requires more than understanding a video. SocialOmni evaluates whether a model can identify speakers, recognize when it should respond, and generate an appropriate reply from audio and visual context.
Connect voices to people using audio and visual cues.
Decide whether to speak at a given moment, then produce a context-appropriate response.
Every model sees the same audio-visual prefix and the same task cards. We keep the timing decision separate from the response quality score, then report their joint coverage.
Level 1 isolates perception: can the model connect the words in a marked interval to the person who said them?
Level 2 tests the interaction decision and the content of the next turn as two linked but separately measured stages.
The marked interval contains a question about the difference between wealth management and asset management. The model must use both the voice and the visible speaker position.
At 27 seconds, the target is the woman in green. The annotation says YES, so the continuation is eligible for the quality metrics.
The case contributes to When accuracy and its three scores contribute to How quality. Across the full split, the same records produce QGold, QEns, Cov+ and QEns_joint.
23 models shown
| WhoSpeaker identification | WhenResponse timing | HowResponse quality | When + HowCoverage and joint quality | |||
|---|---|---|---|---|---|---|
| Gemini 3.8 Flash | 25.40 | 82.50 | 86.52 | 86.48 | 82.81 | 71.61 |
| Gemini 3.6 Flash | 13.20 | 88.00 | 79.23 | 79.50 | 88.28 | 70.18 |
| Gemini 3 Flash | 0.75 | 78.50 | 86.59 | 87.79 | 78.91 | 69.27 |
| Gemini 3.1 Pro Preview | 45.05 | 84.00 | 80.01 | 81.92 | 82.81 | 67.84 |
| Gemini 2.5 Pro | 5.80 | 78.00 | 80.40 | 82.30 | 82.03 | 67.51 |
| Gemini 3.7 Flash | 16.70 | 73.00 | 84.90 | 84.64 | 79.69 | 67.45 |
| Gemini 2.5 Flash | 1.70 | 68.00 | 67.58 | 72.12 | 63.28 | 45.64 |
| Gemini 3.5 Flash | 6.00 | 60.50 | 79.56 | 82.95 | 50.78 | 42.12 |
| Qwen3.5-Omni-Plus | 91.05 | 58.50 | 80.27 | 77.82 | 48.44 | 37.70 |
| Qwen2.5-Omni | 4.15 | 61.50 | 49.35 | 49.90 | 64.06 | 31.97 |
| Qwen3-Omni | 70.85 | 64.00 | 49.22 | 45.49 | 66.41 | 30.21 |
| OmniVinci | 29.75 | 64.50 | 40.82 | 38.52 | 70.31 | 27.08 |
| Qwen3.5-Omni-Flash (2026-03-15) | 86.55 | 71.00 | 33.14 | 34.72 | 70.31 | 24.41 |
| Qwen3-Omni-Thinking | 75.65 | 50.50 | 63.15 | 78.85 | 30.47 | 24.02 |
| GPT-4o | 35.05 | 50.50 | 77.15 | 76.50 | 30.47 | 23.31 |
| Gemini 3.1 Flash-Lite | 73.90 | 47.00 | 83.98 | 83.81 | 27.34 | 22.92 |
| VITA-1.5 | 34.65 | 56.00 | 50.20 | 49.44 | 46.09 | 22.79 |
| Ming-Omni 2.0 | 52.30 | 49.50 | 75.78 | 77.34 | 25.00 | 19.34 |
| Qwen3.8-Omni-Flash | 89.45 | 44.00 | 77.02 | 69.93 | 17.97 | 12.57 |
| Gemini 3 Pro | 45.40 | 52.00 | 21.55 | 25.00 | 32.03 | 8.01 |
| Gemini 3.5 Flash-Lite | 78.45 | 38.50 | 74.61 | 65.28 | 4.69 | 3.06 |
| MiniCPM-o 4.5 | 72.80 | 38.00 | 57.81 | 47.92 | 6.25 | 2.99 |
| Baichuan-Omni-1.5 | 8.40 | 16.00 | 44.27 | 40.63 | 6.25 | 2.54 |
Click a metric to sort. All scores use a 0–100 scale; higher is better. — means unavailable and is placed last in either sort direction.
Compare a fitting reply with a reply that repeats an offer after it has already been accepted.
Fine, just take it. Lovely.
All right, all right, your sandwich for the cheese toasie, the bag of scratchings and a chocolate bar. Come on, I'm robbing myself. I don't want to swap it. All right, all right, here's what I'll do. The toastie, the scratchings, a choccy bar and...a roast dinner with all the trimmings. How long has that been in there? Come on, I'll throw in the briefeases well. Fine, just take it. Lovely.
What should the woman in green say?
Pleasure doing business with you.
Scores are from the recorded evaluation. The same judges supplied the brief explanations afterward; the scores are unchanged.
Pleasure doing business with you.
The response is a natural, role-consistent closing after the exchange is accepted, fitting the satisfied tone of the dialogue.
Fine, just take it.
Lovely.
The candidate continuation perfectly fits the conclusion of the bargaining exchange, maintaining role consistency and natural pragmatics.
Fine, just take it. Lovely.
The candidate perfectly matches the reference, providing a socially appropriate and coherent closing remark for the transaction described in the context.
Fine, just take it. Lovely.
Come on, I'll throw in the briefcase as well.
The candidate repeats an earlier bargaining line rather than responding after the deal is accepted, so it does not fit the dialogue state. A closing acknowledgment would be more appropriate.
Come on, I'll throw in the briefeases well.
Fine, just take it. Lovely.
The candidate repeats an offer that was already made earlier in the negotiation, making it incompatible with the current dialogue state where the deal has already been concluded.
Come on, I'll throw in the briefeases well. Fine, just take it. Lovely.
The candidate repeats a line already spoken by the previous speaker, failing to provide a distinct response from the target role. This results in a lack of conversational progression and pragmatic inappropriateness.
Come on, I'll throw in the briefeases well.
Explore the tasks through selected videos and their reference annotations.
Match the spoken words to the visible speaker.
What happened in the video from 0:20 to 0:23?
D. The man on the right asked how wealth management differs from asset management.
Match the spoken words to the visible speaker.
What happened in the video from the 15th second to the 20th second?
A. The man on the left explains how it works and how it can be used in parallel with many different tools.
Watch only the conversation before the decision point.
Should the woman in green speak at the 27th second?
Yes
Pleasure doing business with you.
Watch only the conversation before the decision point.
Should the little boy standing on the far right speak at the 17th second?
No
QEns_joint = QEns × Cov+ / 100. A high response-quality score alone does not imply reliable decisions about when to speak.
New evaluations use Gemini 3.8 Flash, Qwen3.8-Omni-Flash and GPT-5.6-Sol as judges. Model names link to the corresponding results and evaluation settings.
Missing or unfinished evaluations are not assigned a score. This page does not combine the six metrics into an overall score.
If you use SocialOmni in your research, please cite:
@article{xie2026socialomni,
title={SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models},
author={Xie, Tianyu and Huang, Jinfa and Ma, Yuexiao and Luo, Rongfang and Yang, Yan and Ma, Qingchuan and Chen, Wang and Zeng, Yuhui and Zou, Yixuan and Lu, Zhiqiang and Fang, Ruize and Luo, Jiebo and Ji, Rongrong and Zheng, Xiawu},
journal={arXiv preprint arXiv:2603.16859},
year={2026}
}