← Devlog
Incident

Oops, I closed my laptop and broke the chain

Entry #29 · 2026-04-23 · Devlog

Entry #29

I promise this isn't a post I wanted to write. But it's a useful one, and the whole point of a public build log is that you tell the truth about the parts that didn't work.

Yesterday at 13:22 UTC, the Valdium testnet produced block #4341. At 13:22:40 it tried to produce #4342 and couldn't reach consensus. It kept trying, every eight seconds, for twenty-three hours straight. have 1/4 prevotes. have 1/4. have 1/4.

The setup

The testnet was running across four boxes: three Hetzner VPSes in Germany, Virginia, and Oregon, and one more validator — a personal MacBook running through Valdium Operator (the desktop app). Five active validators total, the MacBook being a beta test of what "home validating" actually feels like.

Around 13:22 UTC I closed the MacBook. Not rebooted — just closed the lid, which puts it to sleep and drops the network connection immediately.

The cascade

Under normal Tendermint-style BFT, losing one of four validators is survivable. Quorum is 2f+1, so with four validators you need three votes. But the moment the Mac went silent, another validator briefly hiccupped at the exact same time. Two validators unavailable simultaneously meant the live set was two of four — not enough to reach quorum.

A second bug was waiting: the primary validator didn't have a peer RPC environment variable set, so its HTTP sync loop wasn't running. When it fell out of step, it couldn't pull blocks from peers to catch up. It just sat there, proposing block #4342 over and over, getting one prevote (its own), and timing out every eight seconds. For twenty-three hours.

The fix

The fix is a liveness-aware committee. Instead of the consensus threshold being "2/3 of everyone who ever bonded," it's "2/3 of everyone who actually proposed a block in the last 30 blocks." If your laptop has been offline for three minutes, you're not in the committee. If you come back and propose your next scheduled slot, you're in it again, immediately. No slash. No ceremony.

I wrote it this afternoon and redeployed to all four boxes. The chain produced block #4345 within five seconds of the last restart and hasn't stopped since.

What I actually learned

If your testnet survives because you personally keep four VPS boxes running, your testnet hasn't actually survived anything. The real test is whether the chain tolerates the conditions the thesis claims it tolerates. Operational defaults matter more than I was giving them credit for — sync should be on by default, not something you have to remember to configure. And 23 hours is a long time. Better to learn this on zero-stake test VLD than on the day someone's stablecoin infrastructure is on top of it.

The chain's back. It's better than it was yesterday. — dev team

« Previous

Operator polish — your Mac is a validator

Next »

Governance is on-chain — and you can see it

New entries weekly

Don't miss the next entry.

Join the launch list and we'll send you a note whenever there's a new devlog entry, a research drop, or a real milestone.

Join the launch list Read the thesis →
© 2026 Valdium. All rights reserved.
PrivacyTerms