RAG Evaluation Workbench
A practical toolkit for testing retrieval quality across datasets, prompts, and rerankers.
PROJECT
These are working questions disguised as software. Each project begins with a problem I want to understand from first principles.
A practical toolkit for testing retrieval quality across datasets, prompts, and rerankers.
A voice agent that remembers callers, identifies roofing emergencies, and hands off with context.
A deliberately small email and document assistant built to make routing decisions observable.
Short English essays about the decisions, tests, and mistakes behind the software.
A practical way to separate normal chat, retrieval, and tools without hiding uncertainty behind a large classifier.
Retrieving a relevant chunk is only the beginning; a useful evaluation must follow the evidence into the final answer.
What a roofing voice agent taught me about memory, urgency, and designing a handoff that a real team can trust.