Preserving Acoustic Cues for Video Reasoning:
An Efficient Uniqueness-Driven Token Compression Framework
Abstract
Effective video reasoning requires the joint understanding of visual and acoustic cues. Existing multimodal large language models either discard audio entirely or rely solely on ASR transcripts, losing rich acoustic features such as prosody, environmental sounds, and speaker characteristics. Naively forwarding all audio tokens to an MLLM also introduces prohibitive computational overhead.
We propose Flash-VAReason, a framework that integrates efficient audio token compression into video reasoning through a three-stage pipeline grounded in information uniqueness theory: Audio Time Fusion, Budget Control, and Spatial Dynamic Compression. By retaining acoustically distinctive tokens, Flash-VAReason enables joint visual-audio reasoning while preserving inference efficiency.
Method Overview
Experiment Result
Table 1a. Main Results on CharadesEgo
| Method | Venue | P | R | F1 | Sem | Comp | NoHall | Overall |
|---|---|---|---|---|---|---|---|---|
| AKeyS | arXiv 2025 | 0.4151 | 0.4121 | 0.4098 | 1.9397 | 1.9359 | 1.9952 | 1.9397 |
| mPLUG-Owl3 | ICLR 2025 | 0.5218 | 0.6893 | 0.5807 | 1.9607 | 1.9365 | 2.0980 | 1.9607 |
| Grounded-Video-LLM | EMNLP 2025 | 0.5127 | 0.7110 | 0.5832 | 2.2068 | 2.1395 | 2.4190 | 2.2019 |
| FlashVID | ICLR 2026 | 0.5002 | 0.6961 | 0.5693 | 2.2822 | 2.2434 | 2.4389 | 2.2806 |
| UniComp | CVPR 2026 | 0.4887 | 0.7225 | 0.5702 | 2.2872 | 2.2231 | 2.4425 | 2.2818 |
| Flash-VAReason | Ours | 0.5066 | 0.7285 | 0.5840 | 2.3795 | 2.3062 | 2.5776 | 2.3789 |
Table 1b. Main Results on Ego4D
| Method | Venue | P | R | F1 | Sem | Comp | NoHall | Overall |
|---|---|---|---|---|---|---|---|---|
| AKeyS | arXiv 2025 | 0.4991 | 0.3618 | 0.4138 | 1.4548 | 1.4201 | 1.5457 | 1.4548 |
| mPLUG-Owl3 | ICLR 2025 | 0.6831 | 0.6291 | 0.6511 | 1.3496 | 1.2981 | 1.5079 | 1.3528 |
| Grounded-Video-LLM | EMNLP 2025 | 0.6870 | 0.6411 | 0.6580 | 1.6099 | 1.5137 | 1.8665 | 1.6083 |
| FlashVID | ICLR 2026 | 0.6707 | 0.6392 | 0.6495 | 1.5137 | 1.4558 | 1.7103 | 1.5152 |
| UniComp | CVPR 2026 | 0.6761 | 0.6479 | 0.6553 | 1.6444 | 1.5743 | 1.8609 | 1.6476 |
| Flash-VAReason | Ours | 0.6875 | 0.6485 | 0.6585 | 1.6450 | 1.5750 | 1.8670 | 1.6480 |
Table 2. Efficiency Results
| Dataset | Method | Overall | Time (s) | Speedup |
|---|---|---|---|---|
| CharadesEgo | mPLUG-Owl3 | 1.9607 | 5.32 | 21.00× |
| CharadesEgo | Grounded-Video-LLM | 2.2019 | 111.74 | 1.00× |
| CharadesEgo | AKeyS | 1.9397 | 15.68 | 7.12× |
| CharadesEgo | FlashVID | 2.2806 | 8.95 | 12.49× |
| CharadesEgo | UniComp | 2.2818 | 6.49 | 17.21× |
| CharadesEgo | Flash-VAReason | 2.3789 | 6.10 | 18.33× |
| Ego4D | mPLUG-Owl3 | 1.3528 | 6.75 | 9.82× |
| Ego4D | Grounded-Video-LLM | 1.6083 | 66.34 | 1.00× |
| Ego4D | AKeyS | 1.4548 | 21.91 | 3.03× |
| Ego4D | FlashVID | 1.5152 | 16.40 | 4.05× |
| Ego4D | UniComp | 1.6476 | 13.08 | 5.07× |
| Ego4D | Flash-VAReason | 1.6480 | 13.74 | 4.83× |
Table 3. Compression Ablation on CharadesEgo
| Method | Qwen-Overall | Avg. Inference Time (ms) |
|---|---|---|
| Full token | 2.3838 | 10664.71 |
| Uniform Sampling | 2.0121 | 6144.51 |
| w/o ATF | 2.3011 | 7221.89 |
| w/o SDC | 2.3807 | 7249.35 |
| Flash-VAReason | 2.3789 | 6096.97 |