How we build at scale.
Deep technical writing from our engineering team. Architecture decisions, performance investigations, postmortems, and the patterns we've adopted (and rejected) over 9 years of building production systems.
Why we stopped using ORMs in latency-critical services
A field guide to writing SQL by hand without losing your mind.
Production RAG: the patterns that actually work in 2025
Citations, evals, guardrails, and the unglamorous plumbing that makes RAG useful.
Designing for 60 FPS: a field guide to buttery interfaces
The CSS, the JS, and the discipline required to keep animations smooth.
Threat modeling for product engineers: a practical primer
STRIDE, attack trees, and how to think like an attacker without becoming one.
The SLO error budget is a leadership tool, not an engineering metric
Why most teams misuse error budgets — and how to use them to ship faster.
The case for boring infrastructure (and when to break it)
Why we default to Postgres, Redis, and Kubernetes — and the rare cases where we don't.
LLM observability: measuring what actually matters in production
Latency, cost, quality, and drift — the four metrics that separate toy LLM apps from production ones.
The 47-minute outage: what we learned from a Redis failover gone wrong
A detailed postmortem of our March 2025 incident — what broke, why, and the fixes we shipped.
Shaving 200ms off cold starts: a deep dive into Lambda init
How we reduced p99 Lambda cold starts from 1.8s to 280ms without warming strategies.
Want to read more from our team?
Subscribe to our engineering newsletter — one thoughtful post every two weeks, no spam.