Sirjan Singh
CASE STUDY / Repo-Aware RAG Assistant

Contextual

Retrieval quality is an evaluation problem before it is a generation problem.

PythonFAISSBM25Cross-EncoderGemini

Hybrid retrieval over 10k+ code chunks combining BM25, FAISS, and cross-encoder reranking with sub-second retrieval.

WHAT THIS PROVES

Hybrid retrieval / Evaluation harness / Reranking / Three shipped interfaces

LAB_04 / RAG_RETRIEVALDETERMINISTIC SAMPLE DATA

Exact symbols and lexical matches

01tool_gate.ts · 0.94KEEP
02approval-policy.ts · 0.88CANDIDATE
03run-controller.ts · 0.73CANDIDATE

Interactive system walkthrough · simulated frontend data · not live telemetry

01THE PRESSURE

Problem

Semantic similarity alone misses exact identifiers and repository-specific language.

02THE SYSTEM

Approach

Merge lexical and dense candidates, rerank with a cross-encoder, then evaluate the retrieved context.

03THE MEASURE

Experiments

Measured precision, recall, F1, keyword hit-rate, and latency against a semantic-only baseline.

04THE BREAK

Failure mode

Dense retrieval underweighted exact code symbols and project vocabulary.

05THE JUDGMENT

Engineering decision

Use complementary retrievers and evaluate retrieval independently from generation.

06THE EVIDENCE

Measured result

30% retrieval F1 improvement versus a semantic-only baseline.

Return to all selected workVERIFY THE CLAIMS. INSPECT THE WORK.