Skip to content
Abdullah Al Rafi

Document RAG that stores nothing

An end-to-end RAG system for PDFs and text files: bring your own Hugging Face token, get streamed answers grounded in your file, with nothing stored.

Summary

Role
Sole developer
Timeline
Mar 2026 – Present
Stack
FastAPI · LangChain · Hugging Face · pypdf · React · Vite · Cloudflare · Sentry

0

uploaded files persisted

processed in memory, per request

5 min

vector-store cache per document

keyed by file hash, model and chunking

1–10

retrieved chunks, chosen per query

On this page
  1. Summary
  2. Problem
  3. Constraints
  4. Approach and architecture
  5. Results
  6. What I'd do differently

Problem

Most "chat with your PDF" tools keep your document in a vector database you don't control. I wanted document question answering where you bring your own Hugging Face token, upload a PDF or text file, and get answers grounded only in that file, with nothing stored on the server.

Constraints

  • No storage. Uploads are processed in memory and never written to disk or a database.
  • Bring your own token. The browser keeps the user's Hugging Face token in localStorage and sends it as a bearer header with each request. The server falls back to its own token only if one is configured.
  • A public demo. Anyone can call it, so requests are rate-limited, questions are capped at 1,000 characters, files are size-limited, and only .txt and .pdf are accepted. The API docs sit behind a private token.
  • Hosted inference. Embedding and generation run on Hugging Face inference endpoints, so repeated work on the same document has to be avoided.

Approach and architecture

  1. React + Vite on Cloudflare

    client

    Hugging Face token stays in the browser's localStorage

  2. POST /rag/query · file + question + bearer token

    FastAPI gate

    api

    Rate limit · .txt/.pdf only · size and 1,000-character limits

  3. Text extraction

    extract

    pypdf for PDFs, UTF-8 for text, in memory only

  4. cache key: SHA-256(file) + model + chunking + token hash
    • Reuse the vector store

      cache hit

      Valid for five minutes

    • Split and embed

      cache miss

      Recursive splitter → HF embeddings → in-memory store

  5. Top-k similarity search

    retrieve

    k from 1 to 10, chosen per query

  6. Grounded answer

    generate

    HF chat completion · temperature 0 · 512-token cap

  7. streamed response

    Answer, then metrics

    client

    Tokens, generation time and tokens per second

Fig. 1 · Request path for one question. Nothing is written to disk at any step.
Diagram description

The React client sends the file, the question and the user's Hugging Face token to a FastAPI endpoint. The API checks rate limits and input limits, extracts text in memory, and looks up a cached vector store by a hash of the file, model, chunking settings and token. On a miss it splits and embeds the text; on a hit it reuses the cached store. It then retrieves the top-k chunks, generates a grounded answer at temperature 0, and streams it back followed by generation metrics.

Decision

Keep the vector store in memory, behind a short-lived cache

Each document's InMemoryVectorStore is cached for five minutes, with at most 64 entries. The cache key combines a SHA-256 of the file, the embedding model, the chunking settings and a hash of the user's token. Follow-up questions skip re-embedding, and users never share each other's vectors.

Considered instead

  • A hosted vector database such as pgvector or PineconePersisting embeddings would break the no-storage promise.

Decision

Run models on Hugging Face inference with the user's token

Embedding and generation model names are environment settings, so either can be swapped without a code change.

Considered instead

  • Self-hosting the embedding and generation modelsA free public demo would need GPU hosting to run them.

Decision

Answer only from the retrieved context

The prompt tells the model to say it doesn't know when the answer isn't in the context, and generation runs at temperature 0 with a 512-token cap.

Considered instead

  • Letting the model blend in its own knowledgeAnswers could no longer be checked against the document.

Decision

Stream tokens, then append the metrics

The stream ends with a small metadata block carrying generated tokens, generation time and tokens per second, so the UI can show speed without another call.

Considered instead

  • A separate metrics requestOne more round trip, with numbers computed apart from the response they describe.

Results

0

uploaded files persisted

Text is extracted and embedded in memory, per request

5 min

cache lifetime per document

Follow-up questions skip re-embedding

2

CI gates on every push

Backend pytest; frontend immutable Yarn 4 install + production build

The system is live at rag.abdullahalrafi.com, with the source on GitHub. Errors go to Sentry, and cache hits and misses are logged as breadcrumbs, so slow answers can be traced to re-embedding.

What I'd do differently

  • Count tokens with the model's tokenizer. Tokens per second is currently estimated by splitting on whitespace, which undercounts subword tokens and makes speeds hard to compare across models.
  • Evaluate retrieval against a small labelled question set, and tune chunk size and top-k against it rather than by environment variable.
  • Plan for more than one worker. The cache lives in a single process, and a shared, encrypted cache would raise hit rates, at a real cost to the no-storage promise.