All writing
EngineeringAugust 27, 20264 min read

Count the Processes: Three Production Bugs That Only Exist Behind More Than One Worker

A rate limiter that never limited, a database that ran out of connections, and migrations that couldn't get a lock. One root cause behind all three.

Henry Iddirisu

Henry Iddirisu

AI product engineer

Four workers each holding their own memory, pointing at one shared database

Every one of these worked on my laptop.

On my laptop, the API is one process. In production it's four uvicorn workers, or two gunicorn workers, plus a Swagger container, plus a migration job, all sharing one Postgres. Most of the bugs that surprised me this year came from forgetting that.

Here are three, from two different services, and the one question that would have caught each of them.

1. The rate limiter that never limited

On the DUBTEL AI backend (FastAPI), the forgot-password endpoint needed per-email rate limiting. The first version was the obvious one: a module-level dict mapping email to recent request times.

It passed every test. It did nothing in production.

The API runs four uvicorn workers. Each worker is its own process with its own copy of that dict. Two quick requests for the same email landed on different workers, each worker saw an empty dict, and neither request was throttled. With four workers, an attacker gets roughly four times the limit, and in practice the limit barely applied at all.

The fix was to move the counter into state every worker shares. That meant a small Postgres table, since the service already talks to Postgres and didn't need Redis for this. The limiter became a query and an insert.

The question: Where does this state live, and how many copies of it are there?

2. The database that ran out of connections

On the Mango Capital API (Flask + SQLAlchemy), QA started returning 500s, and the errors pointed at the database running out of connections.

The arithmetic was simple once I wrote it down. Each gunicorn worker gets its own SQLAlchemy connection pool. The pool defaults, multiplied by the number of workers, plus the other containers connecting to the same database, came to more connections than the managed Postgres plan allowed.

The fix was to make the arithmetic explicit, and to leave it in the config for the next person:

# 2 gunicorn workers × pool_size 3 = 6 base connections; max_overflow 2 = 10 total max
SQLALCHEMY_ENGINE_OPTIONS = {
    'pool_size': 3,
    'max_overflow': 2,
}

Two more changes went in with it:

  • Fewer workers. The Dockerfile went from four workers to two. For that service's traffic, two was plenty, and it halved the connection budget.
  • Fewer surprise JOINs. A relationship declared lazy="joined" meant every User.get() pulled in a JOIN that nothing on that path needed. Switching it to select made the common query cheaper and shorter-lived.

The managed database sat behind PgBouncer, which also needed sslmode=require on every connection URL. That's a separate lesson: a pooler in front of Postgres is one more process with opinions.

The question: Workers × pool size + overflow + every other client. Is that under the server's limit?

3. The migrations that couldn't get the database

Same service, same week. Deploys ran flask db upgrade in the pipeline, and it kept failing or hanging.

The live containers were still connected. The API workers and the Swagger container all held connections while the migration tried to alter the tables they were using. My first fix stopped the API container before migrating. That wasn't enough, because the Swagger container was still there. The next fix stopped every container explicitly. The one that finally held used docker compose down to remove the whole stack before migrating, then brought it back up.

It took five commits in one afternoon to get there. Each one fixed the process I knew about and missed the one I didn't.

The question: Who else is connected right now?

The habit

All three bugs have the same shape: something that is true for one process is false for several. None of them show up in unit tests, because a test runs in one process.

What I do now before shipping anything with state or connections:

  1. Write down the process count for the target environment: workers, sidecars, jobs, cron.
  2. For every piece of state, ask where it lives. Module globals and in-memory caches are per-process. If correctness depends on them being shared, they're wrong.
  3. Do the connection arithmetic and leave it in a comment next to the pool config.
  4. Make the deploy prove it. A /health endpoint that returns the build SHA and timestamp takes five minutes to add. It ends the "is the new code actually running?" guessing that eats half of every production investigation.

None of this is advanced. It's counting. But counting is what my laptop never forced me to do.

#python#fastapi#flask#postgres#production
Henry Iddirisu

Henry Iddirisu

AI product engineer · Accra, Ghana · Remote

Keep reading