Sengupta et al., "Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions" (2025)
2025-11-21 → 2026-07-16
- Saurav Sengupta, Nazanin Moradinasab, Jiebei Liu, Donald E. Brown
- arXiv | Code and data
- Counting, Vision language model, Visual counting, Synthetic data
Introduces a controlled synthetic benchmark for diagnosing how vision-language models count as object numerosity, visual properties, and prompt specificity vary. Across Qwen, Kimi, and InternVL variants, counting degrades as counts and visual or linguistic complexity increase; mask-guided reweighting of decoder attention toward object-bearing visual tokens produces modest, architecture-dependent improvements that the authors also test on FSC-147. The paper treats these interventions as diagnostic evidence about cross-modal grounding rather than a general solution to visual counting.