A benchmark for social interaction

SocialOmni

1Xiamen University2University of Rochester3Nanyang Technological University4Sichuan Agricultural University

† Corresponding author

Who speaks. When to respond. What to say.

Evaluating multimodal models in social interaction through audio and video.

Introduction

Social interaction requires more than understanding a video. SocialOmni evaluates whether a model can identify speakers, recognize when it should respond, and generate an appropriate reply from audio and visual context.

LEVEL 1

Who is speaking?

Connect voices to people using audio and visual cues.

LEVEL 2

When and how to respond?

Decide whether to speak at a given moment, then produce a context-appropriate response.

SocialOmni · Project introduction

How an evaluation becomes a score

Every model sees the same audio-visual prefix and the same task cards. We keep the timing decision separate from the response quality score, then report their joint coverage.

LEVEL 1

Who · speaker attribution

Level 1 isolates perception: can the model connect the words in a marked interval to the person who said them?

  1. InputA fixed audio-visual clip and a question with four speaker options. The interval is bounded before the request is sent.
  2. Model outputThe model must return exactly one label: A, B, C or D. Extra prose is treated as a parse failure.
  3. ScoringWe compare the parsed label with the annotation over all 2,000 items and report accuracy. Macro-F1 is retained to show performance across speaker classes.
2,000 items · 4-way classification · accuracy + macro-F1
LEVEL 2

When + How · social response

Level 2 tests the interaction decision and the content of the next turn as two linked but separately measured stages.

  1. When decisionAt the annotated timestamp, the model sees only the preceding prefix and returns YES or NO. This is scored over 200 fixed-time decisions.
  2. How continuationFor the 128 gold-positive states, we retain a forced continuation even if the model said NO. This makes QGold a full gold-positive measure.
  3. Quality and coverageThree judges score each non-empty continuation from 0 to 100. QEns uses true-positive, non-empty answers; Cov+ measures their coverage; QEns_joint multiplies the two.
200 When items · 128 gold-positive How items · 384 judge scores
Level 1 case · video_2

Who is asking the question?

The marked interval contains a question about the difference between wealth management and asset management. The model must use both the voice and the visible speaker position.

Prompt
“Who said the marked question?”
Output space
A / B / C / D
Reported metric
accuracy, with macro-F1 as a class-balanced diagnostic
Level 2 case · video_0003

When opens the gate for How

At 27 seconds, the target is the woman in green. The annotation says YES, so the continuation is eligible for the quality metrics.

When
YES · the target should speak
Reference
“Pleasure doing business with you.”
Ming-Omni 2.0
“Oh thank you very much”
Judge scores
100 / 75 / 75 from Gemini, Qwen and GPT

The case contributes to When accuracy and its three scores contribute to How quality. Across the full split, the same records produce QGold, QEns, Cov+ and QEns_joint.

Results

23 models shown

Model scores on a 0–100 scale. Higher is better. Select a column heading to sort.
WhoSpeaker identification WhenResponse timing HowResponse quality When + HowCoverage and joint quality
Gemini 3.8 FlashNative AV · Default thinking25.4082.5086.5286.4882.8171.61
Gemini 3.6 FlashNative AV · Default thinking13.2088.0079.2379.5088.2870.18
Gemini 3 FlashNative AV0.7578.5086.5987.7978.9169.27
Gemini 3.1 Pro PreviewNative AV · Default thinking45.0584.0080.0181.9282.8167.84
Gemini 2.5 ProNative AV5.8078.0080.4082.3082.0367.51
Gemini 3.7 FlashNative AV · Default thinking16.7073.0084.9084.6479.6967.45
Gemini 2.5 FlashNative AV1.7068.0067.5872.1263.2845.64
Gemini 3.5 FlashNative AV · Default thinking6.0060.5079.5682.9550.7842.12
Qwen3.5-Omni-PlusNative AV · Non-thinking91.0558.5080.2777.8248.4437.70
Qwen2.5-OmniNative AV4.1561.5049.3549.9064.0631.97
Qwen3-OmniNative AV70.8564.0049.2245.4966.4130.21
OmniVinciNative AV29.7564.5040.8238.5270.3127.08
Qwen3.5-Omni-Flash (2026-03-15)Native AV · Non-thinking86.5571.0033.1434.7270.3124.41
Qwen3-Omni-ThinkingNative AV · Thinking75.6550.5063.1578.8530.4724.02
GPT-4oCascade35.0550.5077.1576.5030.4723.31
Gemini 3.1 Flash-LiteNative AV · Default thinking73.9047.0083.9883.8127.3422.92
VITA-1.5Native AV34.6556.0050.2049.4446.0922.79
Ming-Omni 2.0Native AV52.3049.5075.7877.3425.0019.34
Qwen3.8-Omni-FlashNative AV · Non-thinking89.4544.0077.0269.9317.9712.57
Gemini 3 ProVisual-only45.4052.0021.5525.0032.038.01
Gemini 3.5 Flash-LiteNative AV · Default thinking78.4538.5074.6165.284.693.06
MiniCPM-o 4.5Native AV72.8038.0057.8147.926.252.99
Baichuan-Omni-1.5Native AV8.4016.0044.2740.636.252.54

Click a metric to sort. All scores use a 0–100 scale; higher is better. — means unavailable and is placed last in either sort direction.

How: two ways to continue the same conversation

Compare a fitting reply with a reply that repeats an offer after it has already been accepted.

Fine, just take it. Lovely.

Conversation context

All right, all right, your sandwich for the cheese toasie, the bag of scratchings and a chocolate bar. Come on, I'm robbing myself. I don't want to swap it. All right, all right, here's what I'll do. The toastie, the scratchings, a choccy bar and...a roast dinner with all the trimmings. How long has that been in there? Come on, I'll throw in the briefeases well. Fine, just take it. Lovely.

LEVEL 2 · video_0003

How question

What should the woman in green say?

Reference answer

Pleasure doing business with you.

Scores are from the recorded evaluation. The same judges supplied the brief explanations afterward; the scores are unchanged.

Positive example

Gemini 3.1 Pro Preview

Pleasure doing business with you.

GPT-5.6-SolRecorded score: 100 / 100

The response is a natural, role-consistent closing after the exchange is accepted, fitting the satisfied tone of the dialogue.

Fine, just take it.

Lovely.

Gemini 3.8 FlashRecorded score: 100 / 100

The candidate continuation perfectly fits the conclusion of the bargaining exchange, maintaining role consistency and natural pragmatics.

Fine, just take it. Lovely.

Qwen3.8-Omni-FlashRecorded score: 100 / 100

The candidate perfectly matches the reference, providing a socially appropriate and coherent closing remark for the transaction described in the context.

Fine, just take it. Lovely.

Negative example

Qwen3.5-Omni-Plus

Come on, I'll throw in the briefcase as well.

GPT-5.6-SolRecorded score: 25 / 100

The candidate repeats an earlier bargaining line rather than responding after the deal is accepted, so it does not fit the dialogue state. A closing acknowledgment would be more appropriate.

Come on, I'll throw in the briefeases well.

Fine, just take it. Lovely.

Gemini 3.8 FlashRecorded score: 0 / 100

The candidate repeats an offer that was already made earlier in the negotiation, making it incompatible with the current dialogue state where the deal has already been concluded.

Come on, I'll throw in the briefeases well. Fine, just take it. Lovely.

Qwen3.8-Omni-FlashRecorded score: 25 / 100

The candidate repeats a line already spoken by the previous speaker, failing to provide a distinct response from the target role. This results in a lack of conversational progression and pragmatic inappropriateness.

Come on, I'll throw in the briefeases well.

Dataset examples

Explore the tasks through selected videos and their reference annotations.

LEVEL 1 · l1-2

Who asked the question?

Match the spoken words to the visible speaker.

Question

What happened in the video from 0:20 to 0:23?

  • A. The woman on the left asked what wealth management actually is.
  • B. The man on the right asked what wealth management actually is.
  • C. The woman on the left asked how wealth management differs from asset management.
  • D. The man on the right asked how wealth management differs from asset management.
Show reference answer

D. The man on the right asked how wealth management differs from asset management.

Source annotation ↗
LEVEL 1 · l1-8

Who is explaining the tool?

Match the spoken words to the visible speaker.

Question

What happened in the video from the 15th second to the 20th second?

  • A. The man on the left explains how it works and how it can be used in parallel with many different tools.
  • B. The man on the right said he would systematically explain all the most important concepts about how to use it.
  • C. The man on the left said he was very happy to be back here,
  • D. The man on the right said they would jointly create an authoritative course
Show reference answer

A. The man on the left explains how it works and how it can be used in parallel with many different tools.

Source annotation ↗
LEVEL 2 · l2-3

Take the next turn

Watch only the conversation before the decision point.

Question

Should the woman in green speak at the 27th second?

  • A. Yes
  • B. No
Show reference answer

Yes

What should the woman in green say?

Pleasure doing business with you.

Source annotation ↗
LEVEL 2 · l2-9

Knowing when to listen

Watch only the conversation before the decision point.

Question

Should the little boy standing on the far right speak at the 17th second?

  • A. Yes
  • B. No
Show reference answer

No

Source annotation ↗

Reading the metrics

Who
Accuracy in identifying who is speaking from the available audio and video.
When
Accuracy of the YES / NO decision to respond at a fixed timestamp, using only the audio and video available up to that time.
How · QGold
Mean response quality over all gold-positive items: cases where the reference says a response is appropriate. Generation is forced even when the model predicts NO.
How · QEns
Mean quality of non-empty responses on gold-positive items where the model also predicts YES. Each response is scored by the evaluation’s three judges.
When + How · Cov+
Coverage of gold-positive items: the percentage for which the model predicts YES and generates a non-empty response.
When + How · QEns_joint
Response quality weighted by coverage. Missed gold-positive opportunities contribute zero.

QEns_joint = QEns × Cov+ / 100. A high response-quality score alone does not imply reliable decisions about when to speak.

Evaluation

New evaluations use Gemini 3.8 Flash, Qwen3.8-Omni-Flash and GPT-5.6-Sol as judges. Model names link to the corresponding results and evaluation settings.

Missing or unfinished evaluations are not assigned a score. This page does not combine the six metrics into an overall score.

Historical results and original materials ↗

Citation

If you use SocialOmni in your research, please cite:

Download BibTeX

@article{xie2026socialomni,
  title={SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models},
  author={Xie, Tianyu and Huang, Jinfa and Ma, Yuexiao and Luo, Rongfang and Yang, Yan and Ma, Qingchuan and Chen, Wang and Zeng, Yuhui and Zou, Yixuan and Lu, Zhiqiang and Fang, Ruize and Luo, Jiebo and Ji, Rongrong and Zheng, Xiawu},
  journal={arXiv preprint arXiv:2603.16859},
  year={2026}
}