CORTEXA
← Browse
openalexOpen MIND2026-07-30Cited by 0

Retrieval Augmented Generation Using Multimodal Large Language Models for Real-Time Knowledge-Grounded Question Answering

Dr. K. Sujatha

The exponential growth of heterogeneous digital information across structured and unstructured repositories presents a critical challenge for large language models (LLMs): the inability to access and reason over dynamically evolving knowledge without costly model retraining. This paper introduces a comprehensive Retrieval Augmented Generation (RAG) framework that integrates multimodal large language models (MLLMs) with real-time, knowledge-grounded question answering systems. The proposed architecture — MultiRAG — combines a dense bi-encoder retrieval backbone with a cross-modal fusion module capable of jointly indexing and retrieving text, images, tables, and structured data. Retrieved multimodal evidence is processed by a vision-language model (VLM) serving as the generative backbone, conditioned on retrieved context through a novel cross-attention grounding mechanism that attenuates hallucination by enforcing faithfulness constraints at the token level. Experiments conducted on four benchmark datasets — Natural Questions, WebQA, MultiModalQA, and a custom real-time knowledge update benchmark (RKUB-2024) — demonstrate that MultiRAG achieves 87.3% Exact Match on open-domain QA, 91.4% answer faithfulness score, and 6.7× reduction in hallucination rate compared to vanilla LLM baselines. Real-time knowledge ingestion pipeline latency averages 340 ms per document, supporting continuous knowledge grounding without model fine-tuning. The system reduces hallucination by 82% over standard LLM deployment and outperforms all retrieval-augmented baselines by 4.2–9.8 percentage points across evaluation metrics

View free PDFSource page

Related papers

openalexOpen MIND2026-07-23

Friction-Guided Inference: Calibrating Correction Strategies and Abstention from Logprob Signals

Tomas Pødenphant Lund

Large language models frequently possess the knowledge needed to answer a question correctly yet commit to the wrong response. This paper presents friction-guided inference, a calibrated inference-time pipeline that uses the model's own logprob distribution — available at zero co…

View free PDFSource page
openalexOpen MIND

A UAV-Based Framework for Community-Scale Residential Building Energy Modelling in China

Mengfan Jin

With increasing emphasis on energy efficiency and carbon emission reduction in the building sector, rapid and scalable energy modelling of existing buildings is critical for retrofit projects and policy development. Conventional surveys, data collection and energy modelling proce…

Also available via: Open MIND

View free PDFSource page
openalexOpen MIND

EpiGenAI Extraction

Sergio Consoli

This data catalogue contains the generated results for "Generative AI for Epidemiological Information Extraction," comparing in-context learning and fine-tuning with large language models to enhance outbreak data extraction. It includes the generated outputs and plots for evaluat…

Also available via: Open MIND

View free PDFSource page
openalexOpen MIND

Mechs

francis lee

You can treat this as buildable with today’s tech, but you’re in “prototype MBT + experimental railgun + biped robot” cost territory for a single unit. Below is an order‑of‑magnitude cost breakdown for the first full prototype and a manufacturing/integration roadmap assuming your…

Source page