CORTEXA
← Browse
crossrefAI2026-03-02Cited by 0

The Development of a Large Language Model-Powered Chatbot to Advance Fairness in Machine Learning

Pedro Henrique Ribeiro Santiago, Xiangqun Ju, Xavier Vasquez, Heidi Shen, Lisa Jamieson, Hawazin W. Elani

Background: Machine learning (ML) has been widely adopted in decision-making, making fairness a central ethical and scientific priority. We developed the Themis chatbot, a Large Language Model (LLM) system designed to explain concepts of ML fairness in an accessible, conversational format. Methods: The development followed four stages: (1) curating a document corpus of 286 peer-reviewed publications on ML fairness; (2) development of Themis by combining a modern LLM (OpenAI’s GPT-4o) with Retrieval Augmented Generation (RAG); (3) creation of a 340-item benchmark dataset, the FairnessQA; and (4) evaluating performance against state-of-the-art non-augmented LLMs (DeepSeek R1, GPT-4o, GPT-5, and Grok 3). Results: For the multiple-choice questions, Themis achieved an accuracy of 96.7%, outperforming DeepSeek R1 (90.0%), GPT-4o (89.3%), GPT-5 (92.0%), and Grok 3 (86.7%), and the overall difference was statistically significant (χ2(4) = 10.1, p = 0.038). In the closed-ended questions, Themis achieved the highest accuracy (96.7%), while competing models ranged from 78.0% to 84.0%, and the overall difference was significant (χ2(4) = 23.9, p < 0.001). In the open-ended questions, Themis achieved the highest mean scores for correctness (M = 4.62), completeness (M = 4.59), and usefulness (M = 4.56), and differences were statistically significant (correctness: F(4, 195) = 20.91, p < 0.001; completeness: F(4, 195) = 7.76, p < 0.001; usefulness: F(4, 195) = 2.90, p < 0.001). By consolidating scattered research into an interactive assistant, Themis makes fairness concepts more accessible to educators, researchers, and policymakers. This work demonstrates that retrieval-augmented systems can enhance the public understanding of machine learning fairness at scale.

View free PDFSource page

Related papers

crossrefAI2026-01-09Cited by 2

Wildfire Probability Mapping in Southeastern Europe Using Deep Learning and Machine Learning Models Based on Open Satellite Data

Uroš Durlević, Velibor Ilić, Bojana Aleksova

Wildfires, which encompass all fires that occur outside urban areas, represent one of the most frequent forms of natural disaster worldwide. This study presents the wildfire occurrence across the territory of Southeastern Europe, covering an area of 800,000 km2 (Greece, Romania,…

View free PDFSource page
crossrefAI2026-01-16

A Radiomics-Based Machine Learning Model for Predicting Pneumonitis During Durvalumab Treatment in Locally Advanced NSCLC

Takeshi Masuda, Daisuke Kawahara, Wakako Daido, Nobuki Imano, Naoko Matsumoto, Kosuke Hamai, et al.

Introduction: Pneumonitis represents one of the clinically significant adverse events observed in patients with non-small-cell lung cancer (NSCLC) who receive durvalumab as consolidation therapy after chemoradiotherapy (CRT). Although clinical factors such as radiation dose (e.g.…

View free PDFSource page
crossrefAI2026-07-12

eGFR-AI: A Stacked Machine-Learning Model for Early Postoperative Kidney Function Prediction—A Pilot Study

Eva Brenner, Luka Bulić, Vilena Vrbanović Mijatović

Background: Postoperative kidney dysfunction is a common and serious complication in surgical patients. Kidney function is typically assessed using the estimated glomerular filtration rate (eGFR), most often calculated with the CKD-EPI equation based on serum creatinine. While se…

View free PDFSource page
crossrefAI2026-07-01

Towards Data-Driven Weather Intelligence in Palestine: A Multi-Station Benchmark of Classical Machine Learning and Deep Learning Models

Mohammad Odeh, Ahmad Hasasneh

Precise weather forecasting plays a critical role in sectors such as agriculture, transport, energy management, and climate change adaptation, and machine learning and deep learning algorithms have been widely used for data-driven time series forecasting problems. In this work, w…

View free PDFSource page
crossrefAI2026-02-11Cited by 6

A Comprehensive Review of Deepfake Detection Techniques: From Traditional Machine Learning to Advanced Deep Learning Architectures

Ahmad Raza, Abdul Basit, Asjad Amin, Zeeshan Ahmad Arfeen, Muhammad I. Masud, Umar Fayyaz, et al.

Deepfake technology is causing unprecedented threats to the authenticity of digital media, and demand is high for reliable digital media detection systems. This systematic review focuses on an analysis of deepfake detection methods using deep learning approaches, machine learning…

View free PDFSource page
crossrefAI2025-07-09Cited by 2

Interactive Mitigation of Biases in Machine Learning Models for Undergraduate Student Admissions

Kelly Van Busum, Shiaofen Fang

Bias and fairness issues in artificial intelligence (AI) algorithms are major concerns, as people do not want to use software they cannot trust. Because these issues are intrinsically subjective and context-dependent, creating trustworthy software requires human input and feedbac…

View free PDFSource page