Published September 30, 2026 | Version v1

LARGE LANGUAGE MODELS IN AUTOMATIC INDEXING OF EDUCATIONAL RESOURCES: A comparative analysis between author language and bibliographic representation

  • 1. Universidade Federal de Minas Gerais, UFMG
  • 2. Instituto Federal de Minas Gerais, IFMG

Description

This study investigates the ability of Large Language Models (LLMs) to automatically index educational resources. To this end, a case study was conducted involving ChatGPT Plus, DeepSeek, and Gemini, comparing their performance to that of academic authors and librarians. The methodological approach includes the adoption of matching criteria that allowed for the establishment of different levels of coverage and progression. The data were analyzed using precision, coverage, and F-measure metrics. The models produced results similar to the keywords assigned by the authors. On the other hand, the correspondence with the descriptors included by librarians showed the lowest match. In other words, the LLMs were closer to derived indexing (direct extraction from the document) than to indexing by attribution (artificial language), as employed by librarians. It is concluded that, based on the results, these models are capable of performing semantic indexing. However, there is difficulty in reproducing the standardized, institutional, or technical logic of library indexing, making it necessary to expand the studies.

Conference abstract published in the proceedings of TOI2026 — XII Congresso Internacional em Tecnologia e Organização da Informação. ISSN 2965-4130.

Original publication in Doity

Abstract (Portuguese)

Investiga-se a capacidade dos Grandes Modelos de Linguagens (do inglês Large Language Models - LLMs) na indexação automática de recursos educacionais. Para tanto, desenvolveu-se um estudo de caso envolvendo ChatGPT Plus, DeepSeek e Gemini, comparando-os aos desempenhos dos autores científicos e dos bibliotecários. A abordagem metodológica inclui a adoção de critérios de correspondência que permitiu estabelecer diferentes níveis de abrangência e de progressão. Os dados foram analisados à luz das medidas de precisão, abrangência e F-measure. Os modelos apresentaram resultados semelhantes às palavras-chave atribuídas pelos autores. Por outro lado, a correspondência com os descritores incluídos pelos bibliotecários apresentou a menor correspondência. Ou seja, os LLMs se aproximaram mais da indexação derivada (extração direta do documento) do que da indexação por atribuição (linguagem artificial), empregada pelos bibliotecários. Conclui-se que, com base nos resultados, esses modelos são capazes de realizar a indexação semântica. Contudo, há dificuldade em reproduzir a lógica padronizada, institucional ou técnica da indexação bibliotecária, sendo necessário ampliar os estudos.

Files

TOI2026_536947_Resumo.pdf

Files (106.9 kB)

Name Size Download all
md5:4e8f3ecf5ce3aaa840a08c43eaeb53db
106.9 kB Preview Download

Additional details

Additional titles

Alternative title (Portuguese)
LARGE LANGUAGE MODELS NA INDEXAÇÃO AUTOMÁTICA DE RECURSOS EDUCACIONAIS: Análise comparativa entre linguagem autoral e representação bibliotecária

Related works

Is part of
Conference proceeding: 2965-4130 (ISSN)

References

  • ACESSO ABERTO. Dados científicos: como construir metadados, descrição, README, dicionário de dados e mais, 2026.
  • ASSOCIAÇÃO BRASILEIRA DE NORMAS TÉCNICAS. NBR 12676: Métodos para análise de documentos — determinação de seus assuntos e seleção de termos de indexação. Rio de Janeiro: ABNT, 1992.
  • BANH, L.; STROBEL, G. Generative artificial intelligence. Electronic Markets, v. 33, n. 63, p. 1-17, 2023. DOI: https://doi.org/10.1007/s12525-023-00680-1. Acesso em: 22 maio 2026.
  • DENG, F.; BUSCHEK, D. Examining autocompletion as a basic concept for interaction with generative AI. I-Com, v. 19, n. 3, p. 251-264, 2020.
  • FERREIRA, M. H. W.; CORRÊA, R. F. Sistematização da obtenção de indicadores temáticos de informação científica. Encontros Bibli 28, 1-30, 2023.
  • FUJITA, M. S. L.; SOUSA, N. M. T. Vocabulário controlado e inteligência artificial na indexação: uma revisão bibliográfica. Perspectivas em Ciência da Informação 30, e56745, 2025. DOI: https://doi.org/10.1590/1981-5344/56745. Acesso em: 22 maio 2026.
  • GRECO, B. C.; EVEDOVE, P. R.; NHACUONGUE, J. A. O potencial dos modelos de linguagem de grande escala (LLMs) na indexação automática. ISKO Brasil, v. 14, n. 2, p. 59-81, 2026.
  • KE, W.; ZHENG, Y.; LI, Y.; XU, H.; NIE, D.; WANG, P.; HE, Y. Large language models in document intelligence: a comprehensive survey, recent advances, challenges, and future trends. ACM Transactions on Information Systems, v. 44, n. 1, p. 1-64, 2025.
  • MINAEE, S.; MIKOLOV, T.; NIKZAD, N.; CHENAGHLU, M.; SOCHER, R.; AMATRIAIN, X.; GAO, J. Large language models: a survey. arXiv preprint arXiv:2402.06196, 2024.
  • TOMCZAK, J. M. Deep Generative Modeling. Springer International Publishing, Cham.
  • TORRES, A. A. L.; MACULAN, B. C. M. S.; ROCHA, A. A.; ASSUNÇÃO, F. M.; MARQUES, F. B.; SILVA, G. R. Inteligência artificial e a indexação de imagens com o Método Iconográfico de Panofsky. ISKO Brasil, Canela, RS, 2025.