CORTEXA
← Browse
arxivcs.HC2026-07-22

TargetFinder: Detecting Widgets from Pixels on Desktop Interfaces

Ahmed Ben Akouche, Géry Casiez, Mathieu Nancel, Julien Gori

''Target-aware'' pointing techniques, like Bubble Cursor or Semantic Pointing, outperform traditional pointing by leveraging knowledge of target locations. Yet the lack of application-agnostic widget geometry information limits their adoption across the desktop. We present TargetFinder, a computer vision-based system for real-time detection of GUI widgets. TargetFinder leverages several fine-tuned YOLO networks trained on a new dataset of 520 annotated desktop screenshots (~38,000 annotations) spanning Windows, macOS, Ubuntu, and web interfaces. TargetFinder uses lightweight screen monitoring and low-latency detection, achieving millisecond responsiveness suitable for interactive use. Evaluations show that TargetFinder outperforms the baseline methods (OmniParser and REMAUI), while system-wide implementations of Bubble Cursor and Semantic Pointing demonstrate the feasibility of deploying universal target-aware techniques that work across applications. We release the dataset, models, annotation tool, and an open-source library for research and applications.

View free PDFSource page

Related papers

arxivcs.CVcs.HC2026-07-08

Video-Based Detection of squint and cataract for accessibility-aware adaptive web interface rendering

Amar Ranjan Dash, Manas Ranjan Patra

Squint and cataract are major ocular disorders that majorly affect visual perception and interaction capability. This paper proposes a real-time video-based automated detection system for squint and cataract detection based on computer vision and image processing methods. The pro…

View free PDFSource page
arxivcs.HCcs.CYcs.SD2026-07-24

Kutti AI: A Voice-First, Offline-Capable Learning Companion with Real-Time Struggle Detection for Visually-Impaired Children

Kadharmoideen Fadurudeen

Most educational technology for children is built around visual interfaces, which excludes the many children worldwide who live with visual impairment -- an estimated 1.4 million children are blind and many more have low vision. We present Kutti AI, a voice-first learning compani…

View free PDFSource page
arxivcs.HC2026-07-14

Towards Knitted Textile Electromechanical Systems

John Martins, Abigail Hou, Brandon Tendilla, Noah Tannas, Rishit Garg, Wenchi Liu, et al.

E-textiles and wearable sensing technologies enable flexible, customizable interfaces for human-computer interaction, with capacitive sensing offering precise touch and pressure detection. While machine knitting provides scalable, mechanically tunable structures ideal for such se…

View free PDFSource page
arxivcs.HC2026-07-10

KnitID: Machine-Knitted RFID Antennas for Battery-Free Authentication, Localization and Interaction

Weiye Xu, Yue Xu, Devin Murphy, Sen Zhang, Te-yen Wu, Yiyue Luo

Battery-free RFID systems offer a scalable and maintenance-free approach to interaction. We present KnitID, a machine-knitted textile RFID antenna design that enables on-body authentication, localization, and interaction. Unlike prior antenna designs, KnitID achieves a compact an…

View free PDFSource page
arxivcs.CLcs.AIcs.HC2026-07-21

AutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated Journalism

Himel Ghosh, Ahmed Mosharafa, Georg Groh

We present AutoJourn, a demonstration system for multi-perspective news generation and bias-aware evaluation using large language models (LLMs). The system tackles three core challenges in responsible automated journalism: extracting diverse perspectives from unstructured social…

View free PDFSource page
arxivcs.CLcs.AIcs.HCcs.MAcs.SE2026-07-15

DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments

Huatao Li, Xinwei Geng, Yuheng Wang, Yutong Li, Runde Yang, Hantao Chen, et al.

LLM-based agents have rapidly improved at operating individual digital environments such as mobile applications, desktop systems, and smart homes. However, real-world user goals often span multiple devices: information may come from a phone, be processed on a desktop, and the res…

View free PDFSource page