CORTEXA
← Browse
arxivcs.HC2026-07-23

Sonic Stage: Automatically Generating Interactive Spatial Soundscapes to Facilitate Dialogue Video Comprehension for Blind Viewers

Shuchang Xu, Xiaofu Jin, Gaurav Jain, Wenshuo Zhang, Huamin Qu, Brian A. Smith, Yukang Yan

Audio description (AD) makes film and television accessible to blind and low-vision (BLV) audiences by narrating characters' actions. However, in scenes with lots of dialogue, AD often omits important actions because it is constrained not to overlap with speech. It is not yet known how to convey characters' actions during dialogue. We present Sonic Stage, a system that transforms dialogue videos into interactive spatial soundscapes, enabling BLV audiences to intuitively understand characters' actions and movements through immersive auditory cues. Sonic Stage conveys essential visual information during dialogue through three auditory techniques: (1) spatialized dialogue to represent spatial layout, (2) diegetic sound to convey character actions, and (3) interactive descriptions to provide context-specific visual details. Evaluation with 12 BLV viewers showed that Sonic Stage significantly improved video comprehension, spatial presence, and narrative engagement. We highlight opportunities for enhancing video accessibility across diverse genres through immersive, interactive audio representations.

View free PDFSource page

Related papers

arxivcs.HC2026-06-29

Concept Catalyst: Exploring Scrutable Interfaces to Structure K-12 Teacher Interactions with Generative AI

Gennie Mansi, Sunni Newton, Roxanne Moore, Meltem Alemdar, Mark Riedl

Purpose: This paper explores how to align AI-based tools with teachers' classroom needs by using scrutable interfaces -- interfaces that link an easily manipulable knowledge representation to an underlying AI model, so users can change the system's outputs without understanding i…

View free PDFSource page
arxivcs.CVcs.HC2026-07-07

AlayaWorld: Long-Horizon and Playable Video World Generation

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, et al.

Game worlds have traditionally been built through labor-intensive production pipelines, making them costly to develop, difficult to customization, and expensive to modify after deployment. Recent advances in video world models offer a fundamentally different paradigm. Rather than…

View free PDFSource page
arxivcs.HC2026-06-26

HandMade: Spatial Prompting for Generative 3D Creation with Part-Labeled VR Sketches

Jialin Huang, Rana Hanocka, Ariel Shamir, Yotam Gingold

Text-to-3D generation lowers the barrier to 3D content creation, but text alone is a weak interface for specifying spatial intent: where parts should be placed, how they relate, and how an object should be organized in 3D. We present HandMade, a workflow that combines VR 3D sketc…

View free PDFSource page
arxivcs.HC2026-06-26

Drag, Infer, Reproject: Grounding LLMs through Spatial Interaction for Image Clustering

Yang Liu, Xuxin Tang, Jiahao Xu, Chris North

Dimension reduction and semantic interaction support image clustering by making similarity structure visible and manipulable. Existing semantic interaction methods encode users' clustering criterion (a user-interpretable semantic dimension, e.g., action, location, or mood) from d…

View free PDFSource page
arxivcs.HC2026-07-02

From Answer Generators to Reasoning Facilitators: Designing AI Tutors for Mathematical Reasoning in High-Stakes Environments

Yuming Feng, Yuan Tian, Erica Zhao

The rapid integration of Large Language Models (LLMs) into educational technology threatens to reduce mathematical learning to mere answer generation. This paper presents a generative study, usability study, and 12-participant field deployment of AITutor, an interactive system th…

View free PDFSource page
arxivcs.ROcs.HC2026-07-31

STAGE: STyle-controllable Action GEneration for personalized autonomous driving

Zihao Liu, Xing Liu, Yizhai Zhang, Panfeng Huang

Driving style refers to the behavioral preferences that drivers maintain during driving, shaped by their diverse experiences, habits, and needs, and is typically reflected in varying levels of aggressiveness. If humans choose to use autonomous driving systems, they would expect t…

View free PDFSource page