Large Language Models for Information Retrieval in Digitized Historical Collections

Main Article Content

Daiane Campos Procópio
https://orcid.org/0000-0002-9006-191X

Abstract

Technological advances that have expanded access to information in the digital environment have driven scientific, technical, artistic, and cultural production. However, the large volume of available information has brought new challenges, especially regarding the retrieval of relevant content that is accessible to people with different needs and abilities. Among these challenges are digitized textual documents, common in institutional collections, which often lack characters recognizable by reading software, hindering information retrieval and use. In this context, Large Language Models (LLMs) emerge as a promising technology due to their ability to generate, summarize, translate, interpret, and retrieve information, even from unstructured texts.
Given this scenario, this study aimed to investigate the potential use of LLMs in the process of identifying and retrieving information from digitized textual documents in institutional repositories, using a sample of theses from the Institutional Repository of the Federal University of Minas Gerais (RI-UFMG). The research is applied and exploratory in nature, employs a qualitative–quantitative approach, and includes a case study. Forty theses were analyzed—five from each of the eight major areas of knowledge—defended between 1962 and 2009, the period during which the RI-UFMG theses were digitized. The selection considered the available resources, the exploratory character of the study, and the structural diversity of the documents, which allowed the evaluation of model performance in heterogeneous contexts. The three models analyzed were Mistral 7B, Llama 3.2, and Qwen 2.5. They were evaluated throughout September 2025 using a personal computer. The analysis was based on a sample of 40 documents, from which 600 responses were generated (200 per model) to five standardized questions, which included both objective and interpretative items. The results indicated that Mistral 7B achieved the best overall performance, with a higher number of coherent responses and fewer hallucinations, followed by Llama 3.2, while Qwen 2.5 presented the lowest coherence rate. The findings also showed that response accuracy decreases as the semantic complexity of the questions increases, indicating that LLMs are more effective in tasks involving literal retrieval than in conceptual interpretation. The qualitative analysis further highlighted computational challenges in using the models and the influence of structural differences among documents on the quality of the responses. It was concluded that LLMs have potential to support information retrieval in digitized collections, provided that their use occurs in a supervised manner, with adequate infrastructure and criteria for validating responses. It is important to emphasize that this is an exploratory study and does not aim for statistical generalization. As a contribution, the research presents a replicable methodology adaptable to other contexts and points to future studies focused on the integration between artificial intelligence and information management, strengthening open science and the democratization of knowledge.

Article Details

Section

Resumo de Tese e Dissertação

Author Biography

Daiane Campos Procópio, Federal University of Minas Gerais

Doutoranda e mestra (2025) em Gestão e Organização do Conhecimento pela Universidade Federal de Minas Gerais (UFMG). Possui graduação em Biblioteconomia pela UFMG (2013) e em Análise e Desenvolvimento de Sistemas pela Pontifícia Universidade Católica de Minas Gerais (2024). Integrante do grupo de pesquisa Observatório de Dados Abertos. Atua nas áreas de Ciência da Informação, Biblioteconomia e Engenharia de Software, com interesse em gestão e recuperação da informação, dados abertos, inteligência artificial e desenvolvimento de soluções tecnológicas aplicadas à disseminação da informação.

How to Cite

CAMPOS PROCÓPIO, Daiane. Large Language Models for Information Retrieval in Digitized Historical Collections. Múltiplos Olhares em Ciência da Informação , Belo Horizonte, v. 16, p. e064051, 2026. DOI: 10.35699/2237-6658.v16i.64051. Disponível em: https://periodicos.ufmg.br/index.php/moci/article/view/64051. Acesso em: 9 oct. 2026.