Published May 6, 2026
| Version v1
Journal article
Open
Distributed Biomedical Text Summarization and Knowledge Discovery Using Generative Language Models and Apache Spark
Description
Abstract
Biomedical literature is rapidly increasing exponentially, and it constitutes a severe challenge that healthcare researchers, clinicians and data scientists must manage by extracting meaningful insights within large amounts of unstructured text. In this paper, a scalable, end-to-end Artificial Intelligence (AI) pipeline that can be used to distribute the workload of biomedical text summarization and automated knowledge discovery is presented, using Apache Spark as a distributed data ingestion and processing engine, domain-adapted transformer-based Large Language Models (LLMs) as abstractive summarization engines, and Latent Dirichlet Allocation (LDA) as an unsupervised topic modeling engine. It was tested on the corpus of more than 5,000 biomedical abstracts of PubMed-Medline. The pipeline itself took 9-12 minutes to run on Apache Spark Pandas UDFs and Falconsai/medical_summarization based on the ability to run an entire workflow 6x faster than the sequential execution. The evaluation by ROUGE showed that there was great semantic retention with ROUGE-1: 0.58, ROUGE-2: 0.41 and ROUGE-L: 0.52 scores. The LDA topic modeling identified five coherent and clinically significant medical research themes, which showed that the system can scale automatically extract knowledge. It is also designed to use SQLite, FAISS vector indexing as the persistence layer, Retrieval-Augmented Generation (RAG) serving server driven by a Flask REST API and a React-based chat interface. Findings affirm that the suggested system is a sound, scalable, and generalizable model of accelerating biomedical research, clinical decision support, and healthcare analytics.
Keywords
Biomedical Text Mining, Apache Spark, Large Language Models, Abstractive Summarization, LDA Topic Modeling, Knowledge Discovery, Healthcare NLP, Distributed Computing, ROUGE Evaluation, RAG, FAISS.
Files
Distributed-Biomedical-Text-Summarization-and-Knowledge-Discovery-Using-Generative-Language-Models-and-Apache-Spark.pdf
Files
(350.8 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:622f1ba7dfe2002e8397001805a2bc00
|
350.8 kB | Preview Download |