Guide 22 of 25 · intermediate

Write and rehearse a home lab incident runbook

Turn an outage or security alert into a calm sequence of triage, recovery, and learning.

Source-verified · Not lab-testedOfficial sources checked 2026-09-29. No PebbleRack hardware compatibility claim.

Outcome

You will have a one-page runbook for a failed service, with contacts, evidence capture, decision points, recovery steps, and a short review afterward. The structure draws on NIST's incident-response guidance, scaled down for a home lab. It does not replace professional response for a serious compromise, safety event, or regulated environment. The example is a proposed exercise, not an observed PebbleRack result.

Before you start

Choose one common, low-risk scenario, such as an application VM becoming unreachable. Have local console access, latest backup inventory, service owner, and a place to record timestamps. Keep a printed or offline copy of access and recovery instructions in case the lab, password manager, or network is unavailable. Agree which conditions require stopping experimentation and getting expert help.

Steps

  1. Write the trigger in observable terms: “The app failed two external checks in five minutes,” not “the lab feels slow.” Record who receives the alert, how they acknowledge it, and the maximum time before escalation. Assign an incident ID and start time. If the service is merely degraded, note that instead of declaring a complete outage.
  2. Establish scope without changing anything. Check whether the client network, DNS, host, VM, application, or upstream internet is at fault. Compare independent signals. In Proxmox, inspect the VM's current state and recent task history; on Linux, inspect systemctl status SERVICE and journalctl -u SERVICE --since ... for the chosen service. Copy only relevant, sanitized evidence to the incident record. Logs may include secrets or user content.
  3. Decide whether the event is operational or potentially hostile. New administrator accounts, unexpected exposed ports, altered backups, or unexplained access should stop the routine restart path. Isolate the affected system using a planned method and preserve logs and snapshots without overwriting them. For a suspected intrusion, seek competent security help; do not have an AI agent “clean” the host.
  4. For an ordinary service failure, apply the least disruptive documented action once: restart that service if it is known safe, or restore a disposable replica from a known backup. Record the operator, command, time, and result. Avoid repeated restarts, blanket reboots, or deleting data to free disk unless the runbook specifically covers those actions and a human approves them.
  5. Verify from a user's perspective, not just a process status. Can the intended request complete? Is fresh data present? Are background jobs and backups still healthy? Continue monitoring for a period suited to the service. If the action fails, roll back a recent known change or move to the documented restore path; stop and escalate if the next action could lose data.
  6. Within a day, write a short review: timeline, impact, what detected it, actual cause versus hypotheses, what fixed it, what remains unknown, and one improvement to monitoring or documentation. Do not rewrite the original notes to make the outcome look cleaner. If customer data or an external service was involved, assess notification obligations under applicable policy with qualified help.

Check it worked

Another person can follow the runbook without asking where the console, backup, or escalation contact is. In a tabletop exercise, they can classify the alert, preserve evidence, select a safe first action, and name the stop condition. The exercise record states what was simulated; it must not be described as a live recovery test.

If it fails / rollback

If the runbook lacks a needed password or network path, do not put secrets into the document; add a reference to the secure location and test recovery access separately. If a recovery action worsens service, stop, record the new state, and use the prewritten rollback or backup procedure. Update the runbook only after the incident is stabilized.

Safety and data notes

Time pressure creates risky shortcuts. Preserve an immutable copy of essential logs and avoid unnecessary writes to a possibly compromised host. Separate facts, hypotheses, and decisions. Keep incident records private, with retention appropriate to the data they contain.

Sources

Official sources checked 2026-09-29: NIST SP 800-61 Rev. 3; Proxmox VE administration guide; Ubuntu journal documentation.

Next guide

Continue with configuration drift and change logs.