RAG on a Raspberry Pi: Cited Answers from a $100 Board
Chunking, MiniLM embeddings, FAISS, hybrid rerank and URL-validated citations — the trade-offs of running retrieval-augmented generation locally.
Executive summary
Stock research is a hallucination minefield: wrong citations are worse than no answer. VietProStocks answers questions about Vietnamese equities with cited sources, running entirely on a Raspberry Pi 4 — MiniLM embeddings, FAISS retrieval, hybrid reranking and a local quantized model, with per-user cost caps and honest degradation on slower hardware.
The Challenge
Build a research assistant that answers with cited sources, costs nothing per query, and runs on a Raspberry Pi 4 8GB — while Vietnamese market data is thin, and LLM answers without sources are unusable for financial research.
Key problems
- 1Cloud LLM APIs cost money per query and are unavailable offline
- 2Uncited LLM answers are worse than useless for financial research
- 3Vietnamese-market news is sparse and unevenly trustworthy
- 4A Raspberry Pi has CPU-only inference: a 7B model runs at seconds per response
Constraints
- !Raspberry Pi 4 8GB target hardware
- !Free-tier or zero-cost deployment
- !Per-user LLM call caps to keep worst-case cost bounded
- !Trilingual UI for Vietnamese-first users
A Local-First RAG Pipeline
The pipeline is deliberately boring: chunk, embed, index, rerank, answer with citations. Every stage logs enough to debug quality without a GPU.
Chunk and embed the corpus
Documents are split at 500 tokens with 100-token overlap, embedded with all-MiniLM-L6-v2 and indexed in FAISS (IndexFlatIP). On the Pi, embedding runs at roughly one to three seconds per document — acceptable for a background ingestion job.
- 500-token chunks with 100-token overlap
- MiniLM-L6 embeddings, FAISS IndexFlatIP with HNSW as the upgrade path
- Index and metadata persisted together for reproducible rebuilds
Hybrid rerank with a trust signal
Retrieval fetches 6–12 candidates, then reranks by semantic similarity plus recency plus source trust, keeping 3–5 chunks for the answer. Trust is per-source and explicit — a scraped forum and an official filing do not enter the prompt as equals.
- top-K 6–12 retrieved, 3–5 passed to the model
- Rerank = semantic score + recency decay + source trust
- Trust per source configured, not inferred
Validate every citation before display
The model is instructed to cite chunk sources; the server then validates each citation URL against the stored corpus and strips anything that does not resolve. Hallucinated sources never reach the UI.
- Citations validated against indexed documents, not model output
- Confidence field per answer
- Hallucinated-source stripping as a hard gate
Multi-backend LLM client with cost caps
One interface covers llama.cpp, any OpenAI-compatible endpoint and a mock backend for tests. Per-user call caps bound worst-case spend, and the UI states expected latency honestly: seconds for local inference, faster on stronger machines.
- llama.cpp Q4 quantized 7B on-device; OpenAI-compatible escape hatch
- Per-user call caps and usage accounting
- Mock backend makes the full pipeline testable in CI
Key technical decisions
Retrieval quality over model size
A small local model with excellent retrieval beats a large model with mediocre retrieval for factual questions — and it is the only option that fits the hardware budget.
Fail-closed citation validation
An unresolvable citation is removed from the answer. Financial research cannot afford plausible-looking fabricated sources.
Boring deployment: systemd + Cloudflare Tunnel
No Kubernetes for a single board. systemd units, cron backups (daily database, weekly index) and a tunnel for HTTPS access — operationally silent for months.
Implementation
Project timeline
Data layer
- Price and news ingestion for Vietnamese market sources
- Technical indicators: SMA, EMA, RSI, MACD, Bollinger
- SQLite persistence with async SQLAlchemy
RAG engine
- Chunking + MiniLM embedding pipeline
- FAISS index with metadata persistence
- Hybrid rerank and citation validation
LLM client
- Multi-backend client: llama.cpp, OpenAI-compatible, mock
- Finance-focused system prompt with citation rules
- Per-user caps and usage accounting
Operations
- Telegram bot and Slack webhooks for alerts
- systemd units, Cloudflare Tunnel, cron backups
- CI for backend and frontend
Implementation challenges
A 7B Q4 model on the Pi answers in seconds to minutes — too slow for interactive chat.
Solution: Set expectations in the UI, run alerts and batch analysis on schedule rather than interactively, and document the hybrid path (local for privacy, nearby machine for latency).
The FAISS index grows to hundreds of megabytes — backups and rebuilds needed a policy.
Solution: Daily database backups, weekly index backups, keep seven generations, and a documented rebuild path from the source corpus.
Vietnamese source coverage is uneven, so recency alone would surface low-trust content.
Solution: Explicit per-source trust in the rerank formula, making trust a first-class input instead of hoping similarity captures it.
Results & Impact
Key achievements
Business impact
Lessons Learned
Retrieval quality beats model size on edge hardware
Most bad answers came from bad context, not a weak model. Improving chunking and rerank moved quality more than any model change.
Latency is a product decision, not just a metric
Seconds-per-answer is fine for scheduled research and alerts, wrong for chat. The same engine serves both if the product admits which is which.
Backups are part of the ML system
An index is a build artifact with a rebuild path; losing it should cost time, not data. Treat migrations and backups as features from day one.
What we’d do differently
- Move FAISS to IndexHNSWFlat when the corpus doubles
- Add Arabic-level tokenizer tests for Vietnamese diacritics
- Offload embedding jobs to a batch window overnight
- Publish an evaluation set with held-out questions
Technologies used
The project
See the full project
More details, screenshots, and information about VietProStocks.
View project