All projects
Featured case study

JD Analyzer

A live LLM pipeline on this site that scores a job description against my profile: structured extraction, a five-part rubric, evidence-checked deal-breaker flags, and a one-line verdict.

Role
Owner and maintainer
System
Live tool at /jd-analyzer
A pipeline from job post to extract, score, flags, and verdict

Results at a glance

  • Four stages streamed to the browser as they finish: extract, then score and flags in parallel, then the verdict.
  • A deal-breaker flag can only fire if the model quotes the JD, and the quote is checked against the source text in code.
  • Runs on Gemini or OpenAI through one client. Switching provider is an environment variable, not a rewrite.
  • An eval suite pins the expected flags for real job descriptions. On Gemini, every case that ran passed, including the new location rules.

Overview

Job descriptions are long, vague, and full of details that matter only to one person: this role is "remote", but only in the US; this one says "hybrid", in Munich. The JD Analyzer reads a posting the way I would and tells me in seconds whether it's worth my time, with quotes to back every concern.

It's live on this site. Recruiters can paste their own JD and see how I reason about fit.

How it works

1. Extract. The model turns the raw posting into a typed object (title, company, location, remote stance, seniority, stack, responsibilities, must-haves, nice-to-haves), validated with a Zod schema. The rule is strict: only what the JD states. An unstated salary is null, never "competitive".

2. Score and flag, in parallel.

  • Score rates five dimensions from 0 to 10: product/AI alignment, stack fit, seniority fit, role shape, and logistics. Each score comes with a rationale that must cite the JD.
  • Flags don't ask the model for a verdict. They ask it a set of narrow yes/no questions: Does the JD restrict where the person must live? Does it require office attendance outside Ghana? Is this primarily a coordination role? Each answer needs a verbatim quote.

3. Vote and ground. The flag questions run five times and each fact is decided by majority. A "yes" only counts if its quote is really in the JD. The code normalises the text and checks it for overlapping four-word windows, so an invented quote can't flip a flag.

4. Apply the rules in code. A small deterministic function turns facts into flags. For example: a location restriction that excludes Ghana is red, and a mandatory BigTech background is yellow. The rules live in TypeScript, where they're testable, not in a prompt.

5. Verdict. A final call writes one sentence in my voice. A red flag always forces a pass, however high the score.

Results stream to the browser with Server-Sent Events as each stage finishes, so the page fills in progressively instead of waiting.

  • My profile and deal-breakers. The profile describes my real stack (TypeScript/React/Next.js, Python/FastAPI, PHP/Laravel, Postgres), my 4+ years, and my honest gaps (Go, Java, .NET, Kubernetes operations, native mobile).
  • Location rules for Ghana. I replaced the flag logic with two facts that match my situation: residency or work-authorization limited to regions that exclude Ghana, and office attendance outside Ghana. Either one is a red flag. A few hours of overlap with New York is fine and doesn't raise a flag.
  • Rubric. Seniority is scored as fit for 4+ years, so "3–5 years" scores best and "8+ / staff" scores low. Product/AI alignment rewards hands-on work on AI-integrated products over ML-infrastructure roles.
  • Gemini support. One OpenAI-SDK client picks Gemini (through its OpenAI-compatible endpoint) or OpenAI from the environment, with retries and backoff on 429s. I tested model availability and quotas against the live API and set the default to a model the key can serve.
  • Evals. I rewrote the expected flags for eight job descriptions (seven real postings and one deliberately fake one) under my profile, fixed an overlap in the matching patterns that would have let one flag satisfy another's expectation, and added pacing between cases so the suite runs within API quotas.

Guardrails for a public tool

  • Rate limits. Upstash Redis enforces 2 analyses per minute per IP and 30 per day across all visitors.
  • Input limits. Empty or oversized posts are rejected before any model is called.
  • Honest output. The page says plainly that the scores reflect my priorities, not a universal rating.

What I took from it

  • Ask the model for small yes/no facts, then apply the rules in code. Holistic judgments drift; booleans with quotes don't.
  • Voting over several runs smooths out nondeterminism better than setting temperature to zero.
  • A flag without evidence is worse than no flag. Check the evidence before believing the model.
  • Quotas are part of the design. A live LLM feature needs per-IP and daily limits, and a model choice that fits them.
Henry Iddirisu

Henry Iddirisu

AI product engineer · Accra, Ghana · Remote

More projects

Next.jsReactTypeScript

Checkout Payments and Student Billing

Role: Full-stack engineer (payment step; billing API extensions; student and…

The payment step of a European e-learning checkout (Stripe, Klarna, Alma, PagoLight, SeQura), plus the billing API and student bil…

Next.jsReactTypeScript

DUBTEL AI Portal

Role: Lead frontend engineer

The white-label Next.js portal where businesses manage AI voice agents, phone numbers, billing, and users, with role-aware navigat…