CoRDS
Coreset-based Representative and Diverse Selection for streaming video understanding. One principled, training-free selector picks visual tokens that are both representative of the value-norm-weighted importance landscape and diverse across the residual subspace, compressing the KV cache block-by-block so long videos never materialize a full cache. Plugs into Qwen2-VL and Qwen2.5-VL with no per-benchmark tuning.