Document RAG that stores nothing
An end-to-end RAG system for PDFs and text files: bring your own Hugging Face token, get streamed answers grounded in your file, with nothing stored.
Summary
- Role
- Sole developer
- Timeline
- Mar 2026 – Present
- Stack
- FastAPI · LangChain · Hugging Face · pypdf · React · Vite · Cloudflare · Sentry
0
uploaded files persisted
processed in memory, per request
5 min
vector-store cache per document
keyed by file hash, model and chunking
1–10
retrieved chunks, chosen per query
Problem
Most "chat with your PDF" tools keep your document in a vector database you don't control. I wanted document question answering where you bring your own Hugging Face token, upload a PDF or text file, and get answers grounded only in that file, with nothing stored on the server.
Constraints
- No storage. Uploads are processed in memory and never written to disk or a database.
- Bring your own token. The browser keeps the user's Hugging Face token in
localStorageand sends it as a bearer header with each request. The server falls back to its own token only if one is configured. - A public demo. Anyone can call it, so requests are rate-limited, questions are capped at 1,000 characters, files are size-limited, and only
.txtand.pdfare accepted. The API docs sit behind a private token. - Hosted inference. Embedding and generation run on Hugging Face inference endpoints, so repeated work on the same document has to be avoided.
Approach and architecture
React + Vite on Cloudflare
client
Hugging Face token stays in the browser's localStorage
- POST /rag/query · file + question + bearer token
FastAPI gate
api
Rate limit · .txt/.pdf only · size and 1,000-character limits
Text extraction
extract
pypdf for PDFs, UTF-8 for text, in memory only
- cache key: SHA-256(file) + model + chunking + token hash
Reuse the vector store
cache hit
Valid for five minutes
Split and embed
cache miss
Recursive splitter → HF embeddings → in-memory store
Top-k similarity search
retrieve
k from 1 to 10, chosen per query
Grounded answer
generate
HF chat completion · temperature 0 · 512-token cap
- streamed response
Answer, then metrics
client
Tokens, generation time and tokens per second
Diagram description
The React client sends the file, the question and the user's Hugging Face token to a FastAPI endpoint. The API checks rate limits and input limits, extracts text in memory, and looks up a cached vector store by a hash of the file, model, chunking settings and token. On a miss it splits and embeds the text; on a hit it reuses the cached store. It then retrieves the top-k chunks, generates a grounded answer at temperature 0, and streams it back followed by generation metrics.
Decision
Keep the vector store in memory, behind a short-lived cache
Each document's InMemoryVectorStore is cached for five minutes, with at most 64 entries. The cache key combines a SHA-256 of the file, the embedding model, the chunking settings and a hash of the user's token. Follow-up questions skip re-embedding, and users never share each other's vectors.
Considered instead
- A hosted vector database such as pgvector or PineconePersisting embeddings would break the no-storage promise.
Decision
Run models on Hugging Face inference with the user's token
Embedding and generation model names are environment settings, so either can be swapped without a code change.
Considered instead
- Self-hosting the embedding and generation modelsA free public demo would need GPU hosting to run them.
Decision
Answer only from the retrieved context
The prompt tells the model to say it doesn't know when the answer isn't in the context, and generation runs at temperature 0 with a 512-token cap.
Considered instead
- Letting the model blend in its own knowledgeAnswers could no longer be checked against the document.
Decision
Stream tokens, then append the metrics
The stream ends with a small metadata block carrying generated tokens, generation time and tokens per second, so the UI can show speed without another call.
Considered instead
- A separate metrics requestOne more round trip, with numbers computed apart from the response they describe.
Results
0
uploaded files persisted
Text is extracted and embedded in memory, per request
5 min
cache lifetime per document
Follow-up questions skip re-embedding
2
CI gates on every push
Backend pytest; frontend immutable Yarn 4 install + production build
The system is live at rag.abdullahalrafi.com, with the source on GitHub. Errors go to Sentry, and cache hits and misses are logged as breadcrumbs, so slow answers can be traced to re-embedding.
What I'd do differently
- Count tokens with the model's tokenizer. Tokens per second is currently estimated by splitting on whitespace, which undercounts subword tokens and makes speeds hard to compare across models.
- Evaluate retrieval against a small labelled question set, and tune chunk size and top-k against it rather than by environment variable.
- Plan for more than one worker. The cache lives in a single process, and a shared, encrypted cache would raise hit rates, at a real cost to the no-storage promise.