All ArchitecturePerformanceDistributed SystemsAI/MLSecurityPostmortems
Engineering
Architecture

Why we stopped using ORMs in latency-critical services

A field guide to writing SQL by hand without losing your mind.

8 min read·Jun 12, 2025
AI
AI/ML

Production RAG: the patterns that actually work in 2025

Citations, evals, guardrails, and the unglamorous plumbing that makes RAG useful.

12 min read·May 28, 2025
Performance
Performance

Designing for 60 FPS: a field guide to buttery interfaces

The CSS, the JS, and the discipline required to keep animations smooth.

9 min read·May 14, 2025
Security
Security

Threat modeling for product engineers: a practical primer

STRIDE, attack trees, and how to think like an attacker without becoming one.

11 min read·May 02, 2025
Distributed
Distributed Systems

The SLO error budget is a leadership tool, not an engineering metric

Why most teams misuse error budgets — and how to use them to ship faster.

7 min read·Apr 20, 2025
Architecture
Architecture

The case for boring infrastructure (and when to break it)

Why we default to Postgres, Redis, and Kubernetes — and the rare cases where we don't.

10 min read·Mar 24, 2025
AI
AI/ML

LLM observability: measuring what actually matters in production

Latency, cost, quality, and drift — the four metrics that separate toy LLM apps from production ones.

9 min read·Mar 12, 2025
Postmortem
Postmortem

The 47-minute outage: what we learned from a Redis failover gone wrong

A detailed postmortem of our March 2025 incident — what broke, why, and the fixes we shipped.

14 min read·Mar 04, 2025
Performance
Performance

Shaving 200ms off cold starts: a deep dive into Lambda init

How we reduced p99 Lambda cold starts from 1.8s to 280ms without warming strategies.

13 min read·Feb 19, 2025

Want to read more from our team?

Subscribe to our engineering newsletter — one thoughtful post every two weeks, no spam.