All projects
Governed Retrieval · 2026

Enterprise RAG Assistant

A production-grade Retrieval-Augmented Generation assistant built to demonstrate Responsible AI at production standards, where the system proves its answers are grounded, or refuses.

183×
LLM-stage latency cut (195.8s → 1.07s)
$0
per month, full stack
0
ungrounded claims survive the filter
100%
answers carry page-level citations
Enterprise RAG Assistant — a grounded answer with page-level clickable citations and a faithfulness / evaluation panel.
Grounded answers with page-level citations, every claim traceable to its source.

The business problem

Retrieval-augmented assistants are easy to demo and dangerous to ship. In any setting where an answer carries consequences, "usually right" is not a standard: a confident hallucination erodes trust faster than a slow answer ever could. The hard part isn't generation; it's proving faithfulness and knowing when to say "I don't know."

This system was built to hold that line: grounded answers with citations, measured hallucination rates, and a refusal path when the evidence isn't there.

The key decision

The obvious fix for a 195.8-second response was a bigger, faster model. The real fix was measurement: observability-driven routing cut latency 183× to 1.07s at $0. And I made refusal a feature: an answerability gate that declines rather than fabricates, because in a governed system a wrong answer costs more than no answer.

Architecture

  • Hybrid retrieval: dense (pgvector/HNSW) + lexical search fused with Reciprocal Rank Fusion (RRF).
  • Local embeddings & reranking: no per-token API cost, full control over the pipeline.
  • Observability-driven model routing: the change that cut LLM-stage latency ~183× (195.8s → 1.07s) at $0.
  • CI/CD: the evaluation harness runs in the pipeline, so regressions surface before merge.

Evaluation approach

A deterministic LLM-as-Judge harness measures two things against golden answers: faithfulness (hallucination rate) and answer correctness. Because it's deterministic, results are comparable run-to-run, so evaluation becomes a gate, not a vibe.

Governance & Responsible AI

  • NLI-entailment filter: drops claims the retrieved sources don't actually support.
  • Answerability gate: refuses rather than fabricates when evidence is insufficient.
  • Page-level clickable citations: every answer is traceable to its source.

Lessons learned

The biggest latency win came not from a faster model but from observability: measuring where time actually went, then routing accordingly. And refusal is a feature: an assistant that declines on thin evidence is more trustworthy, not less. Guardrails belong in the pipeline and in CI, where they can't be quietly skipped.