Preserving Acoustic Cues for Video Reasoning:
An Efficient Uniqueness-Driven Token Compression Framework

Haiwei Xue · Zichao Zhao · Zhiyong Wu
Tsinghua University · The Hong Kong University of Science and Technology · The Chinese University of Hong Kong
Teaser figure rendered directly from the provided PDF attachment.

Abstract

Effective video reasoning requires the joint understanding of visual and acoustic cues. Existing multimodal large language models either discard audio entirely or rely solely on ASR transcripts, losing rich acoustic features such as prosody, environmental sounds, and speaker characteristics. Naively forwarding all audio tokens to an MLLM also introduces prohibitive computational overhead.

We propose Flash-VAReason, a framework that integrates efficient audio token compression into video reasoning through a three-stage pipeline grounded in information uniqueness theory: Audio Time Fusion, Budget Control, and Spatial Dynamic Compression. By retaining acoustically distinctive tokens, Flash-VAReason enables joint visual-audio reasoning while preserving inference efficiency.

Method Overview

Overview figure rendered directly from the provided PDF attachment. Audio tokens pass through Audio Time Fusion, Budget Control, and Spatial Dynamic Compression before entering the Omni-MLLM.

Experiment Result

0.5840CharadesEgo F1
2.3789CharadesEgo Overall
0.6585Ego4D F1
18.33×CharadesEgo Speedup

Table 1a. Main Results on CharadesEgo

MethodVenuePRF1SemCompNoHallOverall
AKeySarXiv 20250.41510.41210.40981.93971.93591.99521.9397
mPLUG-Owl3ICLR 20250.52180.68930.58071.96071.93652.09801.9607
Grounded-Video-LLMEMNLP 20250.51270.71100.58322.20682.13952.41902.2019
FlashVIDICLR 20260.50020.69610.56932.28222.24342.43892.2806
UniCompCVPR 20260.48870.72250.57022.28722.22312.44252.2818
Flash-VAReasonOurs0.50660.72850.58402.37952.30622.57762.3789

Table 1b. Main Results on Ego4D

MethodVenuePRF1SemCompNoHallOverall
AKeySarXiv 20250.49910.36180.41381.45481.42011.54571.4548
mPLUG-Owl3ICLR 20250.68310.62910.65111.34961.29811.50791.3528
Grounded-Video-LLMEMNLP 20250.68700.64110.65801.60991.51371.86651.6083
FlashVIDICLR 20260.67070.63920.64951.51371.45581.71031.5152
UniCompCVPR 20260.67610.64790.65531.64441.57431.86091.6476
Flash-VAReasonOurs0.68750.64850.65851.64501.57501.86701.6480

Table 2. Efficiency Results

DatasetMethodOverallTime (s)Speedup
CharadesEgomPLUG-Owl31.96075.3221.00×
CharadesEgoGrounded-Video-LLM2.2019111.741.00×
CharadesEgoAKeyS1.939715.687.12×
CharadesEgoFlashVID2.28068.9512.49×
CharadesEgoUniComp2.28186.4917.21×
CharadesEgoFlash-VAReason2.37896.1018.33×
Ego4DmPLUG-Owl31.35286.759.82×
Ego4DGrounded-Video-LLM1.608366.341.00×
Ego4DAKeyS1.454821.913.03×
Ego4DFlashVID1.515216.404.05×
Ego4DUniComp1.647613.085.07×
Ego4DFlash-VAReason1.648013.744.83×

Table 3. Compression Ablation on CharadesEgo

MethodQwen-OverallAvg. Inference Time (ms)
Full token2.383810664.71
Uniform Sampling2.01216144.51
w/o ATF2.30117221.89
w/o SDC2.38077249.35
Flash-VAReason2.37896096.97
Higher is better for Qwen-Overall; lower is better for inference time.

Qualitative Example

Case figure rendered directly from the provided PDF attachment. Audio resolves the visual ambiguity: after starting the timer, the person picks up a guitar and starts playing it.

Demo