Every organisation we audit has backups. Almost none of them has restores. The distinction sounds like consultant pedantry right up to the morning it does not, and the reason it matters is simple: the green tick on your dashboard is reporting that a process exited without an error. It is not reporting that your data survived. Those are different claims, and in most companies only the first one has ever been checked.

Two of ours

It would be dishonest to write this article using other people's failures, so here are two of our own from the last twelve months.

The twenty-byte dump. A nightly database backup ran, reported success, and wrote a file every night. The job was monitored. The monitoring was happy. What it was writing was twenty bytes: the shell had captured an error message from the dump command into the output file, and the compression step had faithfully compressed it. Every night, a perfect, tiny, worthless backup. We found it because somebody looked at the file sizes in a directory listing while doing something else entirely. Nothing in the pipeline was designed to notice that a database backup had become the size of a sentence.

The stale lock. A backup repository uses locks to prevent two operations from colliding. A process died mid-run, the lock stayed behind, and every subsequent attempt declined to start. The job did not fail loudly; it declined politely, and the alerting was configured to fire on failure. Nothing failed. Snapshots simply stopped appearing, and the most recent good one drifted further into the past while every dashboard stayed green.

We are reasonably competent, we do this for a living, and both of these ran for days. That is the point of telling you. If your reaction is that it could not happen to you, that reaction is the finding.

Five ways backups fail without telling you

Valid but empty. The file exists, the archive opens, the format is correct, and the contents are wrong or absent. Technically successful, practically useless. This is the failure our first case belongs to, and it is the most common one we find.

The job that stopped. A credential expired, a disk filled, a schedule was disabled during maintenance, a lock was left behind. Nothing errors, because nothing runs. Alerting built around failure notifications is structurally blind to this, and it is the second case above.

The key nobody has. The backups are encrypted, which is correct, and the passphrase lives in the head of the administrator, or in a password manager whose master credential is itself in the backup. We have seen both. Encryption without a tested key-recovery path is a well-organised way to lose your data permanently.

The dependency that was never in scope. The database is backed up and the uploaded files are not. The virtual machine is backed up and the certificates that let clients connect to it are not. The application is backed up and the specific version of the runtime it needs is not. Scope is defined once, at the start, and the system grows afterwards.

The restore that is too slow. Everything works, and it takes eleven hours to bring back a system the business needs in two. This one is not a failure of backup at all; it is a recovery time objective that nobody ever timed, which makes it a number in a document rather than a fact about the world.

What a restore test actually is

Most organisations that tell us they test their backups mean that they check the job status, or that they restored a single file once. Neither is a test. A restore test has five properties.

  • Someone picks the date. Not the most recent backup. A date chosen at the start of the exercise, ideally a few weeks back, because the failure you are hunting is the one that started a while ago and has been running ever since.
  • It goes somewhere isolated. A separate environment, never over the live system. A restore test that can damage production will not be run often enough to be useful, and quite right too.
  • Verification is a business check, not a checksum. A hash proves the bytes came back. It does not prove the data is there. Somebody who knows the system should open it and confirm three specific things: the record count is plausible, a known transaction from that date is present, and the most recently added feature works.
  • It is timed. From decision to service restored. This number is your real recovery time objective, and it is usually two to four times the one in your continuity plan.
  • It is written down. Date, system, restore point, who did it, elapsed time, what broke. Two pages, in the same place every quarter.

The ninety-minute quarterly drill

The reason restore testing does not happen is that it is imagined as a project. It is not. Book ninety minutes a quarter with two people.

The first fifteen minutes: pick the system and the date, and write down beforehand what you expect to find, because a test where you decide afterwards what counts as success is not a test. The next forty-five: restore into the isolated environment and find out. The last thirty: the business verification, the timing, and the note. Rotate the system each quarter so that within a year the four or five things that would actually hurt have all been exercised, and rotate the person, so that the recovery is not held exclusively in the memory of whoever built it.

One variation is worth building in deliberately: at least once a year, run the drill without the person who normally does it. The most common hidden dependency in any recovery plan is a colleague's undocumented knowledge, and the only way to discover it is to remove them from the room.

Three checks you can automate this week

Assert on size and shape, not on exit code. After every backup, check that the artefact exists, that it is within a plausible size band compared with the last ten, and where it is a database dump, that a row count or table count is in range. Our twenty-byte file would have been caught in one night by a rule as crude as "smaller than half of yesterday is an alert".

Alert on absence. Invert the monitoring. Instead of waiting for a failure notification, check every morning that a fresh backup exists and is younger than its schedule allows. This single change catches the entire second failure class, and it is usually twenty lines of script.

Restore something small, every day, automatically. Pull one random file or one table out of the most recent backup and verify it. It is cheap, it runs unattended, and it converts the question "do our backups work" from an annual anxiety into a daily fact.

The part the auditor will ask about

Backup and recovery are named measures under the NIS2 risk-management requirements and under ISO 27001, and neither framework is satisfied by a schedule. What is asked for is evidence of testing: dates, scope, results, and what was done about the gaps. The two-page note from each quarterly drill is exactly that evidence, produced as a by-product of doing the thing properly.

We put a restore-tested backup arrangement on the short list of technical spend that a company genuinely cannot avoid in the ninety-day NIS2 plan, and this is why. It is one of the few controls where the paperwork and the protection are the same activity.

The bottom line

You do not have a backup. You have a hypothesis that you have a backup, and it remains a hypothesis until somebody restores from it, in a room, with a clock running, against a date they did not choose to flatter themselves.

Ninety minutes a quarter converts the hypothesis into a fact. It also, in our experience, finds something the first two times. Both of the failures at the top of this article were found by accident, and we would much rather have found them on a Tuesday afternoon with a stopwatch than at three in the morning with a client on the phone.


Related reading: The 3am Test: Is Your Business Resilient?, Business Continuity Beyond Backup and NIS2 in Greece, Step by Step.