VEGAS: Human-Aligned Video Caption Evaluation via Gaze
1The University of Texas at Austin2AMD
Transactions on Machine Learning Research (TMLR) 2026
TL;DR
- VEGAS uses a viewer’s gaze to evaluate and select video captions at test time, without retraining a vision-language model.
- It favors captions supported by what the viewer attended to, improving caption alignment and video retrieval in egocentric activities.
- We also release a new dataset VEGAS pairing visual content, gaze, and human descriptions across everyday activities and instructional slides.
A generic caption may miss what matters to a particular viewer or user. People attend to different objects and actions in the same scene, so even an accurate description can omit the detail someone needs to recall an event, find a video, or act on what they see. Our core idea is to use gaze as a proxy for visual attention to evaluate and select captions that reflect the viewer’s focus. Studying this requires visual content, synchronized gaze, and descriptions from the same viewers—a combination scarce in existing datasets. We therefore curate a new dataset that brings these three signals together.
This work follows VIBE (NeurIPS 2025), which evaluates video summaries through grounding in the video and usefulness for a downstream task, without human-written references. VEGAS (TMLR 2026) carries this direction toward viewers, using gaze to evaluate whether a caption reflects their attention. Together, the papers explore how information-theoretic scores can select descriptions that are both supported by visual content and relevant to their use. See the VIBE project page for an introduction to the precursor work.
VEGAS: A Dataset of Visual Content, Gaze, and Captions
We curate a new dataset across two everyday settings. In Aria Everyday Activities (AEA), people act in the world: the video and synchronized gaze come from the wearer of Project Aria glasses. In SlideVQA, people seek information from dense visual media: online viewers read instructional slides while we collect their gaze through webcam eye tracking.
632 visual samples · 2,981 retained human captions
| Metric | AEA · Egocentric activities | SlideVQA · Instructional slides |
|---|---|---|
| Source videos / slide decks | 91 videos | 30 decks |
| Annotated clips / slides | 332 clips | 300 slides |
| Retained human captions | 1,660 | 1,321 |
| Mean caption length | 14.8 ± 5.6 words | 16.0 ± 7.1 words |
| Gaze collection | Project Aria glasses | Webcam eye tracking via RealEye |
Do people describe what they look at?
We first check the premise. For pairs of SlideVQA viewers, we plot the distance between their fixation patterns on the horizontal axis and the semantic similarity of their captions, measured using Sentence-BERT (SBERT), on the vertical axis.
When viewers look at similar regions, their captions tend to be more similar. When their gaze patterns diverge, so do their descriptions. The association is significant but partial: Pearson r = −0.23, p < 0.001.
Gaze does not explain everything someone means. But it gives us a useful signal for asking whether a caption is relevant to that particular viewer’s attention.
How much does a caption rely on what you didn’t look at?
Consider a caption about pouring liquid into a red cap. If the viewer was looking at that cap, the attended part of the scene should already provide evidence for those words. A caption about an unrelated background object would depend more heavily on regions outside the viewer’s focus.
VEGAS makes that intuition measurable. Starting with the original video V, we construct a gaze-conditioned view G that preserves attended regions and suppresses other content. We then ask the same frozen VLM how likely a candidate caption c is under each view.
A low score means the rest of the scene provides little extra predictive support: the caption is explained by the attended view. A higher score suggests greater reliance on content outside the viewer’s focus.
The likelihood contrast decomposes over caption tokens, so we can inspect which words gaze supports. This is an information-theoretic pointwise ranking score, rather than an exact estimate of mutual information.
Once we have a pool of candidate captions, selection is straightforward: score each candidate and choose the lowest VEGAS score. Gaze is used at inference time; scoring needs no human-written reference caption and no additional model training.
Results and Limitations
The same scene can call for different captions
Before looking at retrieval results, it helps to see what the score captures. In the example below, the slide and candidate captions stay fixed, while the gaze pattern changes.
Gaze pattern 1 favors Caption A, about age, income, and education, while gaze pattern 2 favors Caption B, about the proportion of women and men. In each case, the preferred caption receives the lower score. VEGAS changes its ranking with the viewer’s attention.
In an egocentric scene, the connection can be more direct. The viewer attends to the detergent and red cap, and the corresponding caption receives a lower score than a generic description or an unrelated negative control. Token-level scores show how “liquid” and “red cap” are supported by the attended view.
Can attention-aligned captions help us find a video?
Many video systems use captions as semantic indices for search. If those captions describe what a person cared about, they should make it easier to retrieve the right clip.
We test this on AEA. A human-written caption serves as the query; each video is indexed by a caption selected from a pool generated by multiple VLMs. We compare random selection with choosing the lowest VEGAS score. As a human reference, we use one annotator’s caption to query another annotator’s caption index.
VEGAS improves retrieval across all reported metrics. Recall@5 rises from 48.37% to 52.53%, and mAP@5 rises from 31.73% to 34.21%. The gains below are shown in percentage points relative to random selection.
| Caption index | Recall@1 | Recall@5 | Recall@10 | mAP@5 | mAP@10 |
|---|---|---|---|---|---|
| Random VLM | 22.41% | 48.37% | 60.30% | 31.73% | 33.31% |
| VEGAS (ours) | 23.55%(+1.14 pp) | 52.53%(+4.16 pp) | 64.40%(+4.10 pp) | 34.21%(+2.48 pp) | 35.77%(+2.46 pp) |
| Human pairwise reference | 30.02% | 58.57% | 69.98% | 40.64% | 42.19% |
The improvement is modest at rank 1 and larger at broader retrieval depths. Gaze-aligned captions help move the correct video into the top few candidates, closing part of the gap to human-written indices.
Where does gaze help most?
The strongest results appear when gaze helps disambiguate concrete objects and actions. In AEA, VEGAS-selected captions move significantly closer to human descriptions: mean SBERT similarity increases by +0.0856 (+13.53%) over naive Gemini segmentation summaries (p < 0.001).
The benefit is less direct on slides. Understanding a chart or argument may require combining evidence from several regions, and two people can interpret the same material at different levels of abstraction. On SlideVQA, the mean shift is smaller, +0.0256 (+3.88%), and is not statistically significant (p = 0.0952).
To separate better candidates from better selection, we also compare selectors within the same candidate pool. VEGAS improves AEA SBERT similarity over random selection by +0.013 (p = 0.040), while it is tied with random selection on SlideVQA.
There are practical limits. AEA deliberately emphasizes gaze-sensitive clips, so its gains should not be read as average-case improvements across all videos. VEGAS needs gaze measurements and inherits errors in the frozen VLM, including hallucinations and imperfect likelihood calibration.
BibTeX
@article{chen2026vegas,
title={VEGAS: Human-Aligned Video Caption Evaluation via Gaze},
author={Shenghui Chen and Po-han Li and Ximeng Sun and Shijia Yang
and Emad Barsoum and Zicheng Liu and Sandeep Chinchali
and Ufuk Topcu},
journal={Transactions on Machine Learning Research},
year={2026},
url={https://openreview.net/forum?id=zfM3xSZMQi}
}