Ahmad
Back to all articles
Case Studies
04 June 2025

Lessons from Scaling to 100K Users

AP

Ahmad Parizaad

Founder & Lead AI Engineer

Lessons from Scaling to 100K Users

Six months ago, the app handled 100 concurrent users. Last week, we hit 100K. Here's what broke, what held, and what I'd do differently.

What Broke First (And Why)

Database connections. Always database connections. We were opening a new connection per request like it was 2010. Connection pooling with pgBouncer was a 20-line config that bought us 10x headroom.

The N+1 Query Nightmare

Every new feature seemed fine in development. In production, they'd tank the database. I built a query profiler that screams in Slack when any endpoint makes 50+ queries. Caught three disasters before they shipped.

Caching Strategy Evolution

  • Week 1: No caching (everything dynamic)

  • Week 8: Redis everywhere (over-engineered)

  • Week 16: Strategic caching (just what matters)

The 80/20 rule applies. 20% of your data gets requested 80% of the time. Cache that. Ignore the rest.

When to Optimize (And When Not To)

Early optimisation killed our velocity. But ignoring performance killed our growth. The sweet spot:

  • Monitor everything from day one

  • Optimize when metrics scream, not when you feel like it

  • Every optimization needs a before/after measurement

Architecture Decisions That Paid Off

  • Stateless backend (horizontal scaling became trivial)

  • Message queues for async work (decoupled critical paths)

  • CDN for static assets (50% cost reduction)

  • Database read replicas (separated read/write traffic)

What I Wish I'd Done Earlier

  • Rate limiting (we got scraped hard)

  • Automated load testing (found bottlenecks in staging, not production)

  • Feature flags (rollbacks became instant)

The Real Lesson

Scaling isn't about rewriting everything in Go. It's about identifying bottlenecks, measuring improvements, and not breaking things that work.

Most "scale problems" are actually "we didn't think about this" problems.