A control plane spends real money on machines it does not own, on behalf of people who are asleep. Most of what makes that safe is invisible until it fails. This series opens six of those mechanisms — one per page — and says plainly what each one decides, what it refuses to do, and why the obvious simpler version quietly breaks.
A signal arrives — a heartbeat, an error, a queue depth, a price. Exactly one component is entitled to decide what that signal means. It then either acts, or deliberately does nothing. The refusal is the engineering. Anything can be built to act; the hard part is a system that declines to act on a signal it cannot trust, and says so.
Read them in order or pick the one that matches the problem you currently have. Each page is self-contained.
Five watchdogs that stop paying for a dead machine — each catching a failure the other four structurally cannot see.
Read → 02How the platform knows what it knows. Every fact has exactly one component allowed to write it — two writers is how a fleet starts lying to you.
Read → 03A failover that knows when not to fire. Unavailability is worth retrying elsewhere; a considered rejection is not.
Read → 04One box, several jobs, nobody scheduling it. What has to be true before a machine can change what it serves without dropping work.
Read → 05What a rented GPU must prove before it is allowed to take traffic — and what happens to the ones that cannot.
Read → 06Half price for the work that can wait. Why deferrable work is genuinely cheaper to run, and what a queue must guarantee to earn that discount.
Read →