Sean Kraemer
← Writing

Building a distributed system from scratch

August 20, 2026

I took CS 425 in the fall of 2025 without having taken a distributed systems course before or having written a line of Go. Those two facts are related to everything that follows.

The project was cumulative. By the end you had a membership protocol that detects failures, a replicated file system built on a consistent-hashing ring, and a stream processing engine with exactly-once semantics running on top of both. Every layer depended on the one underneath, so a subtle bug in failure detection surfaced three weeks later as a streaming job that quietly produced duplicates.

The class made the internet legible

Professor Indranil Gupta taught it, and the part I did not expect was how much history there was. We spent real time on where these ideas came from and why they took the shape they did, and it turned the cloud from a product into a set of tradeoffs somebody argued about in the nineties.

The moment it clicked was reading outage post-mortems. In October of that semester, AWS had the us-east-1 failure that took down a good fraction of the internet for most of a day: a race condition in DynamoDB’s internal DNS automation wiped a set of records, and everything downstream that assumed DNS resolution just worked went with it. I was, that same month, writing the code that decides whether a node is dead or just slow to answer. It was the first time coursework explained something I was reading about in the news while it was still happening.

It got better. CNN’s explainer on that outage quoted Professor Gupta himself, walking readers through the race condition with a classroom analogy about two students updating a shared lab notebook until the page ends up empty. Watching my professor translate the exact failure mode we were studying for a national news audience was a very persuasive argument that the course material mattered.

AI was allowed, and it did not save me

The course permitted AI assistance, and I used it. It was not a problem you could prompt your way through. Getting plausible code was easier than figuring out why my implementation produced false positives at a 40 percent packet drop rate. The answer depended on choices I’d made three files away. The design got iterated over and over, and most of the actual progress came from manually testing scenarios and watching what happened.

It was also my first time working with machines I had to SSH into. I ended up living in WSL so I could use tmux, splitting a terminal into a grid and watching ten nodes at once. Seeing all ten panes react to one process dying teaches you more about a distributed system in five seconds than an afternoon of reading logs after the fact.

The demos turned out to be the lesson

Demos were done live in front of a TA, sometimes in a five or ten minute window, and flaky did not pass. If the recovery worked four runs out of five, that was a failing demo. The system had to converge the same way every single time, which is its own engineering requirement on top of merely working.

That constraint changed how I built things. You cannot narrate a distributed failure recovery from scratch in five minutes, so I wrote scripts that drove the whole cluster through a scenario and logged what was happening in a way a person could follow in real time: kill a node here, show detection converging, show blocks re-replicating, show the leader rescheduling the failed task, verify zero duplicates against ground truth.

One of the TAs said the scripts and logging were good enough that it looked like a production system in terms of explainability and readability. That is the comment I think about most, because I had built them for a grade and accidentally learned something that applies to everything I do now. A system you cannot explain while it is running is a system nobody will trust when it breaks. That is the same problem I work on at my day job, where an AI verdict nobody can trace back to a source document is a verdict an underwriter will ignore.

The code is on GitHub. It runs as a ten-node Docker cluster now, so the demos run locally with Docker.