CORTEXA
← Browse
arxivcs.HC2026-07-07

PERSONAJUDGE: Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration Data

Zeyu He, Xuan Qi, Subramanian Chidambaram, Zhichao Xu, Vinayak Arannil, Lydia Chilton, Alex C. Williams

Large language models increasingly serve as judges in AI evaluation, but current approaches rely on consensus preferences that ignore individual evaluator variation. We propose a novel simulation approach that combines categorical judgments with evaluator-specific auxiliary data--retrospective reasoning traces and interface telemetry--to enable LLM-based simulation of individual evaluators via in-context learning. We conduct a systematic empirical study of this approach using multi-facet data from 32 trained annotators across 4,200 preference judgments in a 4 x 4 x 4 factorial design. Our key findings: (1) The simulation approach achieves up to 9.9 percentage point improvements over the Base Judge; (2) Reasoning traces provide the largest gains with higher collection efforts, while interface telemetry often hurts rather than helps performance despite being cheaper to collect. (3) Simulation difficulty is systematic, predicted by an evaluator's neutral usage (most clearly on Helpfulness) and divergence from consensus; the neutral-usage tendency--rather than simulatability itself--is the cross-task-stable property (r = 0.728). These results establish both the potential and limits of evaluator-specific auxiliary data for personalized evaluation, offering methodological insights for scaling individual aware AI assessment.

View free PDFSource page

Related papers

arxivcs.HC2026-07-17

A Human-Centric Evaluation of a Retrieval-Augmented Generation System for Explaining Quebec Insurance Contracts

David Beauchemin, Richard Khoury

With the rise of online insurance sales, consumers face a significant \enquote{advice gap}, requiring them to navigate complex legal contracts without expert guidance. This paper presents a human-centric, extrinsic evaluation of a state-of-the-art Retrieval-Augmented Generation s…

View free PDFSource page
arxivcs.HC2026-07-13

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment

Ancuta Margondai, Julie Rader, Emma Rader, Sara Willox, Mustapha Mouloua

Between 2023 and 2026, frontier AI systems crossed documented human expert baselines on a growing set of bounded, well-specified, evaluable cognitive tasks, including graduate-level science questions, competition mathematics, software-engineering benchmarks, and structured diagno…

View free PDFSource page
arxivcs.AIcs.HC2026-07-16

Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications

Leanne Tan, Rohan Jaggi, Shaun Khoo, Roy Ka-Wei Lee

Evaluations (Evals) are a deployment bottleneck for real-world AI applications: public benchmarks rarely match a team's users, context, or policies, and human review is often tedious to scale. Motivated by our work with AI applications in the public sector, this project addresses…

View free PDFSource page
arxivcs.CYcs.AIcs.HC2026-07-06

Beyond Accuracy: How Humans Evaluate Legally Correct but Socially Controversial Legal Advice from Machines

Benjamin Minhao Chen, Zhiyu Li

AI systems are increasingly used to provide legal advice, raising questions about whether laypeople accept guidance from algorithms--especially when that advice is legally correct but socially controversial. We report a preregistered survey experiment with 3,348 adults in mainlan…

View free PDFSource page
arxivcs.CVcs.AIcs.HC2026-07-09

VEGAS: Human-Aligned Video Caption Evaluation via Gaze

Shenghui Chen, Po-han Li, Ximeng Sun, Shijia Yang, Emad Barsoum, Zicheng Liu, et al.

Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation via GAze Score), a training-free metric that leverages test-time gaze to sample personalized, atten…

View free PDFSource page