r/selfhosted· checksum_charlie@checksum_charlie

My backup job didn't fail. It quietly stopped running, which is worse, because failures send emails

nightly backup, an email alert on failure, a restore drill every quarter. I felt pretty good about it. then during this quarter's drill, the newest snapshot on the off-site copy was five weeks old. what happened: a system update changed how the backup drive gets mounted, the scheduled job started exiting cleanly before doing any actual work, and 'exited cleanly' is not a failure as far as my alerting was concerned. no error, so no email. silence looked exactly like success. nothing was lost, and the drill caught it, which is the whole point of drills. the fix is to alert on the absence of good news. there's now a separate check that complains if the newest backup is more than two days old, no matter what the job claims it did. monitor the result, not the process. 3-2-1 still stands. it just needed a fourth number: how old is the newest copy.

▲0▼3 comments

Join the conversation

Facet is free to read. To reply you need an account: one private root identity, and up to ten public personas that can never be linked to each other or to you.

Create an account
read_only_friday@read_only_friday· 9/28/2026, 10:50:15 PM

a dead man's switch. the job has to check in, and silence is the alarm. after one quiet scheduled job bites you, you end up putting one on everything that runs on a timer

checksum_charlie@checksum_charlie· 9/29/2026, 12:10:15 AM

adding them to every scheduled job this week. certificate renewal is next, because I have a feeling I know how that story ends

greybeard_junior@greybeard_junior· 9/29/2026, 2:25:15 AM

'Silence looked exactly like success' is the sentence. Same trap with log shipping and monitoring agents: the dashboard is calm because nothing is sending it any data.