All projects

AI Engineering Lab: RAG, Tool-Calling Agent, and MCP

In progress: a Python RAG and agent system with pgvector retrieval, grounded refusals, tool-calling, trajectory evals, and an MCP server, which I'm porting to Gemini.

Role
Owner (in progress)
System
Personal AI engineering lab
Vector embeddings feeding an agent with search, calculator, and MCP tools

Results at a glance

  • RAG over local documents: chunking, embeddings, pgvector cosine retrieval, and a score threshold that makes the assistant refuse instead of guess.
  • An agent loop where the model decides whether to search documents, use a calculator, or answer directly, capped at five tool iterations.
  • Trajectory evals that check which tools the agent chose, with precision, recall, and accuracy.
  • The same tools exposed over MCP, usable from Claude Code.

Why this exists

My day-to-day AI work is product integration: an AI voice-agent platform, a search-grounded Gemini pipeline, an AI tutor inside a learning app. This lab is for the layer underneath: how retrieval really behaves, how to measure an agent, and how MCP fits in.

What's in it

  • Structured responses. Model output is parsed into a typed Pydantic schema.
  • RAG. Markdown and text documents are chunked with overlap, embedded, and stored in Postgres with pgvector. Retrieval is cosine similarity with a minimum score. When nothing clears it, the assistant refuses instead of making something up.
  • Tool-calling agent. The model chooses between document search and a calculator, or answers directly. Every tool call is traced, with token usage and estimated cost per run.
  • Evals. A retrieval eval set, plus agent evals that compare the tools the agent actually called against the expected set, reported as precision, recall, and accuracy.
  • MCP server. The same search and calculator executors, exposed over the Model Context Protocol and callable from Claude Code.
  • Local models. Non-agent completions can also run against a local Ollama model.

What I'm doing now

The roadmap is public in the repo. Next up:

  1. Port to Gemini. Swap chat and embeddings to Gemini, as I did for the JD Analyzer, and compare retrieval quality.
  2. Use my own corpus. Replace the sample documents with material from my domain, such as e-learning platform and billing runbooks, and rewrite the eval set to match.
  3. Publish my numbers. Record retrieval and agent eval results on both providers here once they're real.

What I took from it

  • Retrieval needs a floor. Below the score threshold, 'I don't know' is the correct answer.
  • Evaluate agent behaviour (which tool, how many calls), not just final answers.
  • Keep capabilities separate from protocols: the MCP server reuses the same executors as the agent.
Henry Iddirisu

Henry Iddirisu

AI product engineer · Accra, Ghana · Remote

More projects

PythonFastAPISQLAlchemy (async)

DUBTEL AI Platform Backend

Role: Backend engineer (built solo); led the Next.js portal

The multi-tenant FastAPI backend for an AI voice-agent platform, built solo: tenants, role hierarchy, SAML SSO, pricing, recurring…