CORTEXA
← Browse
arxivcs.CVcs.AIcs.HC2026-07-09

VEGAS: Human-Aligned Video Caption Evaluation via Gaze

Shenghui Chen, Po-han Li, Ximeng Sun, Shijia Yang, Emad Barsoum, Zicheng Liu, Sandeep Chinchali, Ufuk Topcu

Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation via GAze Score), a training-free metric that leverages test-time gaze to sample personalized, attention-aligned text. It is a cross-modal, information-theoretic metric that quantifies how well a candidate caption matches a viewer's focus. To evaluate VEGAS, we curate a dataset of egocentric activities and instructional slides paired with synchronized gaze and reference annotations. We then select captions based on VEGAS via rejection sampling without model retraining. Experiments show that VEGAS-selected captions align significantly better with human focus and improve downstream caption-to-video retrieval, demonstrating the practical utility of incorporating viewer attention during inference.

View free PDFSource page

Related papers

arxivcs.CLcs.AIcs.CVcs.HC2026-07-31

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

Jeffrey M. Girard, Jason Z. Zheng, Jacqueline R. Vertino, Antony D'Avirro, Benjamin Peloquin

Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same…

View free PDFSource page
arxivcs.CVcs.AIcs.HCcs.MAcs.MM2026-07-05

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

Lingao Xiao, Yalun Dai, Yangyu Huang, Qihao Zhao, Wenshan Wu, Hugo He, et al.

Despite growing automation, turning a paper into a coherent poster, talk video, and blog piece often remains a labor-intensive last mile. Recent systems increasingly generate multiple dissemination formats, but a practical workflow must also keep the outputs editable in native to…

View free PDFSource page
arxivcs.CLcs.AIcs.CVcs.HCcs.MM2026-07-13

LightMem-Ego: Your AI Memory for Everyday Life

Yijun Chen, Boyi Xiao, Yixian Zhao, Haoting Xia, Buqiang Xu, Jizhan Fang, et al.

Personal AI assistants on mobile and wearable devices continuously perceive users' daily lives through visual and audio streams. However, answering queries about past experiences requires lightweight multimodal memory that can continuously accumulate, organize, and retrieve long-…

View free PDFSource page
arxivcs.CVcs.AIcs.GRcs.HCcs.RO2026-07-17

EgoExoMoCap: Distributed Ego-Exo Human Motion Capture

Jiaxi Jiang, Bharat Lal Bhatnagar, Nan Yang, Lingni Ma, Sebastian Starke, Robin Kips, et al.

Human motion capture from head-mounted devices (HMDs) offers a scalable way to acquire real-world human motion and interaction data, which is crucial for applications in embodied AI and VR/AR. Existing approaches focus on either egocentric body tracking, estimating the motion of…

View free PDFSource page
arxivq-bio.NCcs.AIcs.CVcs.HC2026-07-03

Modeling the Impact of Visual Brand Language on Attention, Object Recognition, and Memory Retrieval

Rachel F. Heaton, John E. Hummel

Visual brand language is the set of visual properties that convey brand identity for a product. What is the impact of visual brand language on a person's ability to recognize and understand the functional identity of an object? Using an empirically supported modeling framework ba…

View free PDFSource page