Sengupta et al., "Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions" (2025)

2025-11-21 → 2026-07-16

Introduces a controlled synthetic benchmark for diagnosing how vision-language models count as object numerosity, visual properties, and prompt specificity vary. Across Qwen, Kimi, and InternVL variants, counting degrades as counts and visual or linguistic complexity increase; mask-guided reweighting of decoder attention toward object-bearing visual tokens produces modest, architecture-dependent improvements that the authors also test on FSC-147. The paper treats these interventions as diagnostic evidence about cross-modal grounding rather than a general solution to visual counting.

Receive my updates

YY's Random Walks — Science, academia, and occasional rabbit holes.

YY's Bike Shed — Sustainable mobility, urbanism, and the details that matter.

×