CORTEXA
← Browse
arxivcs.CV2026-07-01

Evaluating Agentic Harness Systems for Autonomous Computational Pathology

Jie Lin, Zongyi Chen, Qiaoling Zheng, Liuyi Wang, Hengyi Jiang, Jiabao Chen, Xiang Liu, Yinghong Yang, Liansheng Wang

Autonomous computational pathology (ACP) converts high-level pathology analysis goals into executable, traceable and clinically bounded workflows. Realizing this capability requires adapting general agentic harness systems to pathology-specific tasks, tools, evidence standards and clinical claim boundaries. We contribute ACP-Bench, a framework that adapts existing harness systems from computational pathology support toward ACP workflow capability. ACP-Bench evaluates 41 pathology workflow tasks, including 24 biomarker, 7 morphology and 10 prognosis tasks spanning 6 body-system groups and 9 endpoint families. The benchmark evaluates 9 models and 3 harness groups (Claude Code, Codex and Open Code), yielding 369 complete trajectories. ACP-Bench evaluates each trajectory across workflow execution, diagnostic performance and clinical-boundary alignment, combining expert-adjudicated process audits, diagnostic assessment and pathologist-validated safety review. Across evaluated systems, workflow initiation, task interpretation and diagnostic reporting were more mature than tool-bound execution, result binding and reflective workflow revision, and formal end-to-end completion remained rare. ACP-Bench provides a reusable standard for auditing whether agentic systems can operationalize pathology workflows before claims of reliable clinical autonomy.

View free PDFSource page

Related papers

arxivcs.CV2026-07-04

Paired Uterine Whole-Slide Images and Pathology Reports for Multimodal Computational Pathology

Han Li, Jingsong Liu, Ayako Ura, Junlin Hou, Zhengyang Xu, Azar Kazemi, et al.

Uterine diseases represent an important category of gynecologic pathology and require accurate histopathological assessment for diagnosis and treatment planning. Whole-slide images (WSI) have enabled the digital transformation of pathology workflows and provided new opportunities…

View free PDFSource page
arxivcs.CVcs.AIcs.CLcs.CYcs.LG2026-06-30

Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents

Xueqiao Sun, Xiaohan Wang, Ludwig Schmidt, Serena Yeung-Levy, Yuhui Zhang

Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted significant attention for their utility and versatility. A major challenge in developing these agents is collecting large-scale, high-quality traje…

View free PDFSource page
arxivcs.ROcs.CV2026-07-05

Agent-driven Long-tail Simulation for Autonomous Driving

Junru Gu, Lijin Yang, Jianing Huang, Shu Liu, Zhongzhan Huang, Hang Zhao

Evaluating autonomous driving systems in closed-loop settings requires realistic and interactive simulation, yet existing simulators largely rely on log replay or rule-based agents, limiting behavioral diversity and long-tail coverage. We propose an agent-driven simulation framew…

View free PDFSource page
arxivcs.CVcs.MM2026-07-07

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

Wei Dong, Tianyu Fu, Zhe Yu, Hanning Wang, Anyang Su, Zhizhou Fang, et al.

As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion performance has emerged as a critical research priority. However, existing benchmarks exhibit fundam…

View free PDFSource page
arxivcs.CV2026-07-19

Understanding From Human Perspective: A Multi-agent System for Interactive Egocentric Medical Image Segmentation

Rongjun Ge, Dongyang Wang, Heng Zhu, Zhirui Li, Yang Chen, Yuting He

Interactive egocentric medical image segmentation (IEMIS) plays an important role in smart-glasses-assisted medical image review, segmenting the medical targets a clinician refers to from their egocentric view. Once it succeeds, the object-level visual evidence it provides streng…

View free PDFSource page
arxivcs.CVcs.AI2026-07-14

Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs

Jingyu Sun, Jiachen Tu, Yuyang Xue, Yaoxin Jiang, Guoyi Xu, Zhengtao Yao, et al.

Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself. Counterfactual images provide a natural diagnostic setting for this failure mode: when visible evidence contradicts…

View free PDFSource page