The July post was called The Backup Plan Nobody Writes for Their AI Agent. The argument was that people who run disaster recovery for real infrastructure forget to point those instincts at the agent setup on their own desk. It ended with five things to do. Find out where your skills actually live. Put them under version control. Back up the config directory. Write a one-page recovery runbook. Test the restore.

On a Monday in September my laptop died. I got to keep the checklist. I did not get to keep much else.

What Survived Had Nothing to Do With Me

Sixteen custom skills came back untouched. Months of encoded workflow logic, the part I had called the single hardest thing to reconstruct from memory, was simply there when I logged in on a new machine.

I had assumed that was because they were packaged in a plugin, installed from a repo. That was wrong. When I finally checked, the installed plugin list was empty. The skills live server-side on the account, which is why they survived, and which is a fact I could have established in about fifteen seconds at any point in the preceding seven weeks.

Step zero of my own checklist was to find out where your skills actually live. I had published that sentence and then never run it on myself. The outcome was good and the process was luck.

The Backup Died With the Thing It Was Backing Up

This is the one that stings, because it was the opposite of neglect.

One of my internal tools keeps its entire system of record in a SQLite database: vendor assessments, review decisions, ticket linkages, approval history. None of it reconstructable from anywhere else. So months ago I built real protection for it. A daily job pulled the production database down, ran an integrity check, wrote both a binary copy and a full text dump, kept thirty days of history, sent a desktop notification on success or failure, and fed a separate watchdog that posted status to a team channel.

Every piece of that ran on my laptop and wrote to my laptop. The backup directory was in .gitignore, so it had never been in the repository either. When the machine died, thirty days of retained backups died with it, along with the scheduler that would have made tomorrow's.

The production data was fine. It lives in cloud storage. The only thing I lost was the protection I had built for it. I had written a post warning people that their crown jewels were sitting on a single disk, and the thing I was proudest of engineering was sitting on a single disk.

Restoring Files Was the Easy Part

The last file the machine wrote was a plan.

Every weekday morning an agent computes a vulnerability triage run: which unowned findings to route, to whom, which deadlines are about to breach, which exceptions to open. It writes the plan to disk, then an independent reviewer approves each item before anything executes. That morning it wrote the plan at 13:10 and the machine failed before the reviewer ran. Nothing executed.

So the file I recovered was not a record of completed work. It was forty-one findings routed and abandoned mid-flight, including two marked urgent with deadlines four days out, plus a risk acceptance that had reached its final extension with no further automatic renewal permitted. A queue, not an archive.

My checklist covered files and configuration. It said nothing about work in progress, which is the category that actually had a clock on it. Recovering the files took an afternoon. Working out what I had been in the middle of took longer, and only worked because the plan file happened to be written before the reviewer stage rather than after.

Then I Reproduced the Failure Inside the Fix

The rebuild included a backup script that mirrors several directories to cloud storage. Mirroring deletes, so a missing or half-populated source could wipe the copy. I knew that, and wrote a guard: compare file counts, refuse to mirror when the source has dramatically fewer files than the backup already holds.

I set it to engage only when the destination held more than ten files. I was thinking about large directories.

The scheduled tasks folder held nine. When the first real run found only two tasks registered on the new machine, the guard evaluated, decided nine was not more than ten, and deleted the other seven recovered task definitions. The guard I wrote to prevent silent data loss caused silent data loss on its first run.

They were recoverable from a second copy, so the cost was ten minutes. The lesson cost more than that. I had reasoned about the failure mode correctly, written a control for it, and then picked a threshold that excluded the exact case in front of me. Small directories are where a wipe is cheapest to cause and easiest to miss, which is the opposite of where I aimed the protection.

What I Actually Changed

The specifics are dull, which is a good sign. Mirrors now refuse on both a ratio test and an absolute file-loss ceiling, and a dry run prints how many files it would delete before it deletes them. The database backup writes a text dump to cloud storage in addition to the local binary, so the backup and the thing it protects no longer share a disk. A watchdog reports daily whether that chain ran, because the previous watchdog had quietly stopped working weeks before the crash and nothing told me.

The item that mattered most was the one I had ranked fourth. A single document that says: read these two files first, then run these commands, here are the six things that will waste your time. Recovery is mostly an act of remembering, and the runbook is the only part of a backup that addresses remembering. Everything else assumes you already know what you are restoring and why.

The Grade

Five items. I had completed zero of them.

Find out where your skills live: unverified, and I had it wrong. Version control for skills and config: no. Back up the config directory: no. Write a runbook: I had one, and it died with the machine, which is its own kind of answer. Test the restore: never, and the test-restore is the step that would have caught the fact that my database backups had no path off the laptop.

The gap between knowing a thing and having done it is not a knowledge problem, and writing the post publicly did not close it. What closed it was losing the machine. I would rather have closed it in July, and the checklist was sitting right there.

If you read the last post and nodded, go do the fifth item today. Clone to a second machine, confirm a skill loads and a scheduled job fires. It takes an hour. I can tell you exactly what the alternative costs.