Published February 14, 2025 | Version v1

Review of Data Pipelines and Streaming for Generative AI Integration

Authors/Creators

Contributors

Researcher:

Description

Generative AI (GenAI) is transforming industries, but its effectiveness depends on access to timely and relevant data. This paper examines the critical role of real-time data pipelines in powering GenAI applications, synthesizing existing literature across key areas, including data integration, streaming platforms, vector databases, and architectural patterns. We explore the challenges and opportunities in building scalable, high-performance pipelines, emphasizing data freshness, accuracy, and efficient processing.

The intersection of GenAI and big data infrastructure has introduced novel data management techniques, such as data streaming, integration, and vector databases, which are crucial for optimizing AI-driven decision-making. This paper reviews these techniques, their applications, and the evolving role of data pipelines in real-time AI deployment.

A key focus is the integration of data streaming platforms with GenAI, enabling real-time processing and enhancing AI applications. We analyze the state-of-the-art in Apache Kafka, vector databases, and cloud-based solutions, addressing critical challenges such as scalability, data consistency, and integration complexity. Furthermore, we explore future directions, including the use of retrieval-augmented generation (RAG) and real-time data pipelines to unlock GenAI’s full potential.

This review synthesizes insights from recent research, industry practices, and emerging trends to provide a comprehensive understanding of real-time data infrastructure for GenAI.

Files

Review of Data Pipelines and Streaming for Generative AI Integration IJRPR38799.pdf