Unlocking Enterprise AI with S3 Vectors: Bringing RAG to Object Storage with Ceph (ENG)

Enterprises are rapidly adopting AI, but scaling Retrieval-Augmented Generation (RAG) systems remains a challenge due to fragmented data pipelines, siloed vector databases, and complex infrastructure. This session introduces a transformative approach: extending S3-compatible object storage with native vector search capabilities using open source Ceph’s S3 Vectors. By embedding vector indexing and querying directly into the Ceph Object Gateway and integrating a high-performance open source LanceDB engine, organizations can turn their existing data lakes into AI-ready knowledge systems—without duplicating data or introducing new storage silos. We’ll explain how event-driven pipelines powered by bucket notifications enable automated ingestion and embedding generation, while preserving a loosely coupled, serverless architecture. The result is a scalable, cost-efficient foundation for enterprise-grade AI applications.

Redner

  • Kalpesh Pandya
    Kalpesh Pandya
    IBM

    Kalpesh begann seine Karriere bei Red Hat und arbeitet derzeit als Backend Engineer bei IBM. Seit nahezu 6 Jahren erforscht und entwickelt er verschiedene Komponenten von Ceph (einer quelloffenen, softwaredefinierten Storage-Plattform) mit. Er ist seit den Anfängen Teil des Ceph-Rados-Gateway-Teams und hat an der Entwicklung des Secure Token Service (STS), von Bucket-Benachrichtigungen, Multisite sowie eigenständigen RGW-Projekten mitgewirkt. Mit dem Fokus auf kontinuierliches Lernen neuer Technologien und einer lösungsorientierten Denkweise ist er stets bereit, innovative Projekte anzugehen, die die Open-Source-Community voranbringen.