What Building a RAG App Taught Me About Embeddings, Infrastructure, and Reality

Embedding infrastructure failures, self-hosting mistakes, and deployment realities from shipping NOVAR — a document RAG system backed by FastAPI, ChromaDB, and Gemini. The lessons you don't get from tutorials.

When I first started building my RAG (Retrieval-Augmented Generation) application, I thought the hardest part would be the AI itself — the prompting, the retrieval logic, or maybe even the frontend experience.

I was wrong.

The real battle started with embeddings. Like many developers getting into RAG systems, I underestimated how critical embedding infrastructure is. I focused heavily on the "generation" part of the app while assuming embedding would just work in the background. That assumption cost me time, deployments, debugging sessions, and a complete rethink of how I approach AI engineering.

The First Approach: Gemini Embeddings

Initially, I used the embedding model from Google's Gemini ecosystem — it was easy to integrate and the quality looked promising during testing. Small documents embedded correctly. Queries returned relevant chunks. The retrieval pipeline worked.

But once usage increased, things started breaking unpredictably. Sometimes embeddings would take too long. Sometimes requests failed entirely. Sometimes the model behaved inconsistently under load. Since RAG systems rely heavily on embeddings being consistently available, even small instability becomes a major architectural problem.

That was my first real lesson: in RAG systems, embeddings are not a side component. They are the foundation of the entire retrieval pipeline. If embedding generation becomes unstable, the entire application becomes unstable.

The Second Attempt: Self-Hosting a Hugging Face Model

After struggling with the Gemini embedding service, I decided to self-host using models from Hugging Face. The idea sounded great in theory — more control, better reliability, no API rate limits.

But I made a major architectural mistake. Instead of preparing the model beforehand, I designed the service to download the embedding model during startup. That decision became a disaster on Render. The model was large, startup became extremely slow, and the container kept timing out before initialization completed. Every restart meant downloading the model again. Every deployment became painful. Cold starts became unacceptable.

At that moment, I realized something important: AI infrastructure decisions matter just as much as AI model selection. A good model with poor deployment architecture becomes a bad production system.

Chunking Is the Most Consequential Decision You'll Make

Before you touch retrieval or generation, the way you split your documents determines the ceiling on answer quality. The default RecursiveCharacterTextSplitter with chunk_size=1000 is a reasonable starting point but a poor ending point. Here is what actually mattered for NOVAR:

  • Chunk size interacts with your embedding model's context window. For models/embedding-001 (the Gemini embedding model), staying under 512 tokens per chunk produces measurably better similarity scores. At 1000 characters you are often fine, but with dense technical content you can exceed that limit silently.
  • Overlap is load-bearing for multi-sentence answers. With zero overlap, a sentence that straddles a chunk boundary gets split and neither chunk has enough context. 150–200 character overlap recovers most of these cases.
  • Document type changes the right strategy entirely. For PDFs with clearly demarcated sections, splitting on headings outperforms character-based splitting. NOVAR detects whether the document has markdown-style headers after PyMuPDF extraction and picks the strategy accordingly.
Note Chunking decisions compound. A bad chunking strategy cannot be recovered by a better retriever. Fix chunking first before tuning anything downstream.

Retrieval: The Case Against Vanilla Top-k

The standard approach is: embed the query, find the k nearest chunks by cosine similarity, pass them to the LLM. This works until it doesn't. The two failure modes NOVAR hit regularly:

Semantic similarity is not the same as relevance

A user asking "what are the limitations of this method" gets chunks about limitations — but potentially from the wrong section. Adding a lightweight BM25 re-rank step (sparse retrieval over the same chunks) and combining scores fixes most of these cases. The hybrid approach is two extra lines with LangChain's EnsembleRetriever.

k is context-dependent

For a short factual question, 3 chunks is plenty and adding more increases hallucination risk because the LLM tries to reconcile irrelevant context. For a summarization request, 10 chunks is often too few. NOVAR classifies the query intent — factual lookup vs. synthesis — and sets k accordingly:

def get_k(query: str) -> int:
    synthesis_patterns = [
        r"\bsummar",
        r"\boverall\b",
        r"\bexplain\b.*\bin detail\b",
        r"\bwhat does.*cover\b",
    ]
    for pattern in synthesis_patterns:
        if re.search(pattern, query, re.IGNORECASE):
            return 8
    return 4

Streaming SSE Responses

Non-streaming RAG responses feel broken even when they aren't. A 3-second blank wait before text appears trains users to think something went wrong. Gemini's Python SDK supports streaming generation via stream=True, and FastAPI's StreamingResponse with text/event-stream makes the plumbing straightforward.

The critical detail is error handling inside the generator:

async def stream_answer(query: str, session_id: str):
    retriever = get_retriever(session_id)
    docs = retriever.invoke(query)
    context = format_context(docs)
    prompt = build_prompt(query, context)

    try:
        async for chunk in model.generate_content_async(
            prompt, stream=True
        ):
            if chunk.text:
                yield f"data: {json.dumps({'text': chunk.text})}\n\n"
    except Exception as e:
        yield f"data: {json.dumps({'error': str(e)})}\n\n"
    finally:
        yield "data: [DONE]\n\n"

The [DONE] sentinel is important. Without it, the frontend has no reliable signal to stop the spinner on a partial response or a timeout.

Session Isolation

NOVAR uses ChromaDB with a collection-per-session model. Each session gets its own collection keyed by a UUID generated at document upload. The tradeoffs versus a single shared collection:

  • Pro: zero risk of cross-user document leakage, trivial to delete a session's data.
  • Con: ChromaDB's in-process mode does not scale past a single process. Fine for Render's free tier, not fine if you need horizontal scaling.
  • Con: cold session startup adds latency on first query. NOVAR moves indexing to the upload endpoint so the first query is fast.
Watch out If you are using ChromaDB's in-process mode and Gunicorn with multiple workers, each worker gets its own ChromaDB instance. Sessions created in worker A are invisible to worker B. NOVAR uses a single-worker Uvicorn process to sidestep this entirely.

What I Learned About Hosting AI Systems

Traditional web applications and AI applications behave very differently in production. A normal Flask or Django app starts quickly, but AI systems introduce entirely different constraints: large model downloads, heavy RAM usage, long initialization times, GPU and CPU dependency issues, cold start problems, and serialization concerns. What works for a CRUD app may completely fail for an AI service. I learned that hosting AI workloads requires thinking more like an infrastructure engineer than just an application developer.

How I Would Approach a Similar Project Today

If I were rebuilding the same RAG application today, my approach would be completely different:

  • Separate embedding from the main app. A dedicated embedding pipeline or microservice improves scalability, fault isolation, deployment flexibility, and caching strategies.
  • Preload models during build time. Never download models during runtime startup. Preload during the Docker build stage, cache model weights, and mount persistent storage.
  • Choose infrastructure based on the workload. Platforms optimized for lightweight web apps are not always ideal for model-heavy services. Evaluate startup timeout limits, persistent storage, GPU availability, and cold start behavior.
  • Use smaller or quantized embedding models. Production reliability matters more than theoretical benchmark superiority.
  • Build for failure from day one. Plan for retries, fallback providers, local caching, queue systems, and asynchronous processing.
Lesson AI engineering is not just about models. It is about systems design, infrastructure, reliability, deployment strategy, resource management, and operational thinking. The AI model is only one piece of the puzzle.

Even though the embedding issues frustrated me, I am glad I experienced them early. Those failures forced me to understand the deeper engineering side of AI systems instead of only focusing on model integration. Now when I design AI applications, I think beyond "Does the model work?" — I think about whether it can scale, survive traffic spikes, recover from failures, deploy reliably, initialize fast enough, and operate within infrastructure limits.

The full source is on GitHub.