All writing
AI EngineeringSeptember 24, 20264 min read

Porting an LLM Pipeline to Gemini: The Code Took Ten Minutes, the Quotas Took the Afternoon

Moving the JD Analyzer from OpenAI to Gemini was a base-URL change. Finding out why it stopped after two eval cases was the real lesson.

Henry Iddirisu

Henry Iddirisu

AI product engineer

Two model endpoints behind one client, with a quota gauge per model

The switch was one file.

The JD Analyzer on this site is a four-stage LLM pipeline: extract a job description into a typed object, score it on five dimensions, vote on deal-breaker facts five times, then write a verdict. It used the OpenAI SDK with Zod-validated structured output.

I wanted it on Gemini. Gemini exposes an OpenAI-compatible endpoint, so the OpenAI SDK can talk to it directly. The whole provider switch is one small module:

const useGemini = Boolean(process.env.GEMINI_API_KEY);

export const openai = new OpenAI(
  useGemini
    ? {
        apiKey: process.env.GEMINI_API_KEY,
        baseURL: "https://generativelanguage.googleapis.com/v1beta/openai/",
        maxRetries: 5,
      }
    : { apiKey: process.env.OPENAI_API_KEY },
);

export const MODEL = process.env.LLM_MODEL ?? (useGemini ? "gemini-3.5-flash" : "gpt-5.4-mini");

Each stage imports MODEL instead of hard-coding a model name. That was the code change.

What worked unchanged

  • Structured output. chat.completions.parse with a Zod response format worked against Gemini as-is. The extraction, scoring, and fact schemas all parsed.
  • The rules. Deal-breaker flags are applied in TypeScript from grounded yes/no facts, and every quote is checked against the job description in code. None of that cares which model produced the facts.
  • The evals. The first eval case passed on the first run, including the new location rules I had just written for my own search.

That's the payoff of keeping the important logic outside the prompt: the model is swappable.

Then it stopped

Case two passed. Then: 429, no body.

My first guess was a per-minute rate limit, since each job description fires about eight requests, five of them in parallel. So I added SDK retries and a pause between eval cases. It passed two cases, then hit 429 again, with 60-second gaps.

The SDK error had no body, so I called the API directly to read the real one:

Quota exceeded for metric: generate_content_free_tier_requests, limit: 20, model: gemini-3.5-flash quotaId: GenerateRequestsPerDayPerProjectPerModel-FreeTier

Three things in that message mattered:

  1. It's per day, not per minute. No amount of pacing fixes a daily cap.
  2. It's per model. gemini-2.5-flash was exhausted while gemini-3.5-flash still had quota, which is why switching models seemed to help for a while.
  3. It's per project, and the project was on the free tier, even though I expected a paid key. API billing is attached to the Google Cloud project behind the key, not to the account in general, so paying for Gemini somewhere else doesn't lift the API limits.

I also found a model the key couldn't use at all: gemini-2.5-flash-lite returned "no longer available to new users". Listing models through the API before choosing one saved a round of guessing.

What I changed

  • Read the raw error. When an SDK says "429, no body", call the API directly. The real message named the metric, the limit, and the tier.
  • Treat quota as a design input. Twenty requests a day is about two analyses. Shipping the live tool on that would mean a public feature that fails for most visitors, so billing has to be linked before launch.
  • Make the model configurable. LLM_MODEL overrides the default without a code change, which matters when quotas are per model.
  • Pace the evals. The eval runner takes an EVAL_DELAY_MS between cases, so the suite can run on a low tier without false failures.

The takeaway

Porting between providers through an OpenAI-compatible API really is cheap, as long as your guarantees live in code rather than in one model's quirks. The expensive part is operational: which models your key can use, what the limits are for each one, and which project is paying. Check those before you make the switch, not after the eval run dies halfway.

#gemini#openai#llm#rate-limits#evals
Henry Iddirisu

Henry Iddirisu

AI product engineer · Accra, Ghana · Remote

Keep reading