My backup monitor cried wolf. The fix was teaching it one word.
My phone buzzed with a red alert: two client backups had gone stale, eight days old. That is the kind of message that ruins a morning. I looked into it, and the backups were not broken at all. The monitor just did not know something the backup system already did.
The alarm
Two of our client agents had not been backed up in over a week, according to the health check. Stale backups can mean lost data, so this is exactly the alert you want to take seriously.
What was really going on
Both of those agents had been retired days earlier. Their work was archived on purpose, and the backup job correctly stopped backing up something that no longer runs, keeping the last good copy frozen.
But the monitor was never told that. It just saw an old backup and screamed. It would have screamed every single day, forever, about a thing that was working exactly as designed.
The lesson
This is the quiet danger with monitoring. A check that cries wolf every day trains you to ignore it, and the day it is a real fire, you scroll right past it.
The fix was small: teach the monitor the same word the backup job already knew, retired. Now it leaves the frozen copies alone and only shouts when a live backup actually goes missing. A good alert is one you can trust, which means it has to know the difference between broken and done on purpose.
We build the boring, reliable plumbing your business runs on. Ask me about it.
← Back to all posts