CORTEXA
← Browse
openalexJournal of the Association for Information Systems2026-08-15

Classification-Induced Data Bias: How Class-Based Schemas Distort Data at Collection

Aida Nouri, Jeffrey Parsons

Data bias is a recognized problem in research and data-driven decision-making. Existing treatments focus on how data was collected: sampling bias from non-representative selection, measurement bias from recording errors, annotation bias from inconsistent labeling, and omitted variable bias from excluded variables (Mehrabi et al., 2021). These accounts share a common assumption: the data collection schema is neutral. The schema, meaning the set of categories that structure what gets recorded, is treated as a given rather than a source of bias. This paper challenges that assumption and introduces classification-induced data bias: systematic distortion arising from the imposition of a class-based schema at the point of data collection. Class-based schemas, which require contributors to assign every observation to a predefined category, are the foundational assumption of structured data modeling (Parsons & Wand, 2000) and the prevailing practice in real-world data collection platforms (Lukyanenko et al., 2019). This assumption of inherent classification means schema designers fix a set of classes before collection begins, embedding their conceptualization into the data architecture. Phenomena that do not fit this structure cannot be recorded naturally: contributors must force-fit observations, record them inaccurately, or abandon the contribution. Lukyanenko et al. (2019) demonstrate this empirically, showing that class-based schemas significantly suppress both the completeness and accuracy of contributed observations. We define classification-induced data bias as systematic distortion arising when a class-based schema requires contributors to assign observations to predefined categories, causing exclusion, misclassification, or suppression of phenomena that do not conform to the designer’s conceptualization. This construct is structurally distinct from existing bias types: unlike sampling bias, it concerns what properties of observed phenomena can be captured, not which units are selected; unlike omitted variable bias, which concerns variables absent from analysis of already-collected data, classification-induced bias operates at the moment of recording itself, before any observation enters the dataset; unlike measurement bias, distortion originates in the schema design rather than in its application. It manifests through three mechanisms: exclusion (phenomena outside predefined classes cannot be recorded), forced misclassification (observations mapped to the nearest available class regardless of fit), and participation suppression (contributors who cannot classify disengage, producing non-random attrition). The result is data that over-represents phenomena anticipated by designers and under-represents novel or atypical ones. We propose that classification-induced bias is more severe in complex, multi-dimensional domains and when contributors have lower familiarity with the class structure. This bias matters because it distorts what phenomena can enter a dataset, affecting the validity of downstream analysis; future research should develop methods to detect it and evaluate design alternatives that reduce its effects.

View free PDFSource page

Related papers

openalexJournal of the Association for Information Systems2026-08-15

Post-incident Forensic Analysis of Industrial Control Systems (ICS)

Kelly L. Roberts-Cooper

Appearing in the Harvard Business Review in 1958, the formal term "Information Technology" (IT) refers to the combination of hardware, software, services, and infrastructure that delivers data, voice, and video to domestic users. The concept of IT existed before 1958, but the inf…

View free PDFSource page
openalexJournal of the Association for Information Systems2026-08-15

AI Applications to Customer Relationship Marketing-Ethical Considerations

Edward A Wogan

This study examines the explosion of data-driven applications utilizing Artificial Intelligence and Machine Learning in recent years and the myriads of ethical issues that have arisen and continue to evolve. The review focuses on ethics and governance encompassing AI and the appl…

View free PDFSource page
openalexJournal of the Association for Information Systems2026-08-15

Functionalist Perspective on Emotions in AI: A Review of Roles, Mechanisms and Impacts

Eunice Park, Mala Kaul, Chad Anderson

Functionalist Perspective on Emotions in AI: A Review of Roles, Mechanisms and Impacts TREO Talk Paper Eun Hee Park Old Dominion University epark@odu.edu Mala Kaul University of Nevada, Reno mkaul@unr.edu Chad Anderson Miami University, Ohio ander556@miamioh.edu Abstract Recent a…

View free PDFSource page
openalexJournal of the Association for Information Systems2026-08-15

From Prediction to Policy Simulation: AI-Enabled Decision Support for Child Welfare Service Allocation

Minoo Mondaresnezhad

In social welfare systems, agencies are often required to make policy decisions under conditions of uncertainty, limited resources, and significant human consequences. The current study aims to examine how AI-enabled policy simulation can support these policy decisions in the con…

View free PDFSource page
openalexJournal of the Association for Information Systems2026-08-15

Exploring Role of Knowledge Management in Industry 5.0: ERP Systems Can be an Assistance for Business in Knowledge Era?

Yaxin Zheng, Si Liu

Advances in information technology drive digital transformation, enabling enterprises to automate and streamline operations. As foundational digital platforms, ERP systems integrate with core Industry 5.0 (I5.0) technologies—including AI and IoT (Wijesinghe et al., 2024; Sarferaz…

View free PDFSource page
openalexJournal of the Association for Information Systems2026-08-15

Contrastive Learning for Phishing Detection in Text-Based Environments

Thomas Mulaisho, Emmanuel Michael, Sanket Kanekar, Qiunan Zhang, Xi Zhang

Phishing remains one of the most prevalent and effective forms of social engineering attacks in today’s online environment. Attackers typically impersonate trusted individuals or organizations to deceive users and gain access to sensitive information. Recent studies (e.g., Dandot…

View free PDFSource page