Why this exists
My day-to-day AI work is product integration: an AI voice-agent platform, a search-grounded Gemini pipeline, an AI tutor inside a learning app. This lab is for the layer underneath: how retrieval really behaves, how to measure an agent, and how MCP fits in.
What's in it
- Structured responses. Model output is parsed into a typed Pydantic schema.
- RAG. Markdown and text documents are chunked with overlap, embedded, and stored in Postgres with pgvector. Retrieval is cosine similarity with a minimum score. When nothing clears it, the assistant refuses instead of making something up.
- Tool-calling agent. The model chooses between document search and a calculator, or answers directly. Every tool call is traced, with token usage and estimated cost per run.
- Evals. A retrieval eval set, plus agent evals that compare the tools the agent actually called against the expected set, reported as precision, recall, and accuracy.
- MCP server. The same search and calculator executors, exposed over the Model Context Protocol and callable from Claude Code.
- Local models. Non-agent completions can also run against a local Ollama model.
What I'm doing now
The roadmap is public in the repo. Next up:
- Port to Gemini. Swap chat and embeddings to Gemini, as I did for the JD Analyzer, and compare retrieval quality.
- Use my own corpus. Replace the sample documents with material from my domain, such as e-learning platform and billing runbooks, and rewrite the eval set to match.
- Publish my numbers. Record retrieval and agent eval results on both providers here once they're real.
