Almost every business we assess has backups. Far fewer have a restore — because nobody has ever tried one. Until you have pulled real data back out, a green tick in a backup console is a claim, not a fact.
That distinction stops being academic at the worst possible moment. Here is what we look for, and the three ways we most often see a healthy-looking backup turn out to be worthless.
1. The job succeeds but the data is incomplete
Backup agents skip files they cannot lock. Open database files, mailboxes mid-write and files held by a running application are the usual casualties. The job still reports success, because from its point of view it did everything it was asked to do — the exclusions were working as configured.
The fix is not more frequent backups. It is application-aware backup for anything with a database behind it, and an occasional look at the job log rather than just the status colour.
2. The backup is on the same failure domain as the original
A second copy on the same NAS, the same host, or the same site is a convenience copy, not a backup. It protects against a deleted file. It does not protect against a fire, a flood, a failed array, or ransomware that walks the network looking for mounted shares.
This is what the 3-2-1 rule is actually for: three copies of the data, on two different kinds of media, with one of them off-site. It is an old rule and it still holds.
3. Nobody knows how to run the restore
We have seen genuinely good backups sit unused during an outage because the one person who knew the console had left, the credentials were in their password manager, and the documentation described a system that had been replaced two years earlier.
A recovery plan that depends on a specific person being reachable is not a plan.
What testing actually looks like
A real restore test does not have to be disruptive. For most clients it is a quarterly exercise that takes an hour:
- Pick a restore target at random rather than always testing the same tidy file share.
- Restore to an isolated location, never over the live data.
- Open what you restored. A file that restores but will not open has not restored.
- Time it. “How long until we are working again” is the number leadership will ask for.
- Write down what happened, including what went wrong. The failures are the valuable part.
If you only find out your recovery time is eleven hours during an actual outage, you have discovered it eleven hours too late.
We run this on a schedule for the clients we manage, and we send the results through rather than just filing them. If you are not sure when your last successful restore was, that is usually an answer in itself — get in touch and we will take a look.