Outcome
You will have a written monitoring sheet, a working alert destination, and a deliberate test for a stopped noncritical service. Monitoring is useful only when it detects a condition you can act on. A green dashboard alone does not prove that backup restoration, remote access, or a user-facing task works. These instructions are a source-verified pattern for an existing home lab; PebbleRack has not run them on a validated device.
Before you start
Have administrator access to your Proxmox VE node, a separate device that can receive alerts, and a throwaway service or VM you are comfortable stopping. Record node names, service owners, time zone, preferred contact path, and how long you can tolerate each outage. Do not put passwords, API tokens, full log payloads, or private addresses in a public dashboard. Start with Proxmox's built-in task and notification facilities before adding a metrics stack.
Steps
- List five outcomes worth observing: host reachable, essential VM or container reachable, storage has room, latest backup completed, and a restore test was completed recently. Add UPS state if one is connected. For each, write the expected state, check interval, threshold, alert destination, and first human action. Example: a failed backup job should trigger a same-day investigation; low free space may need a trend and a threshold, not an immediate page on every fluctuation.
- In Proxmox VE, inspect Datacenter → Notifications in your installed version. Create a target using a contact method you control. Prefer a dedicated operations address or local notification endpoint. Configure a matcher for errors from backup and system jobs; keep routine success messages in a digest or task history. The exact menu and supported target types vary by version, so follow the version-matched administration guide before saving.
- Send a test notification from the interface if offered. Confirm it arrived on the separate device and note its sender, subject, timestamp, and whether replies work. A saved target is not proof of delivery. If mail traverses an external provider, check spam and provider logs too.
- Record baseline CPU, memory, storage, and network behavior during an ordinary day. Set thresholds from this baseline and available capacity. A disk alert should allow enough space and time to investigate without automatic deletion. A host reachability alert should come from another device; a monitor on the failed host cannot report its own failure reliably.
- Test one reversible failure. Stop only the selected disposable service, note the exact time, verify that your external check detects the outage, then start the service and verify recovery. Preserve the alert and task record. Do not unplug your only network or power source for an alert test.
- Review alert noise after a week. Remove redundant signals, fix false positives, and add a written first response for every retained alert. Schedule a monthly review of backup status and a separate restore exercise.
Check it worked
The alert arrived at the intended destination, identified the affected host or service, and included enough context to start investigation without leaking secrets. The stopped service generated one meaningful alert and its recovery was independently observed. Backup and storage status are visible without relying on the failed host itself.
If it fails / rollback
If no alert arrives, inspect the Proxmox notification task log, target settings, DNS and outbound network path, then send a new test. Revert the matcher or target to its prior configuration if delivery is unreliable. If alerts are noisy, narrow rules before muting all errors. Restore the disposable service to its prior state and check user-facing behavior.
Safety and data notes
Never make an AI model or dashboard the sole source of health truth. Preserve the raw task result and exact timestamps for diagnosis. Avoid posting log lines publicly; they can contain hostnames, usernames, tokens, or user data. Monitoring may notify you of a failed backup, but cannot establish that a backup restores.
Sources
Official sources checked 2026-09-29: Proxmox VE administration guide; Proxmox Backup Server notifications; Proxmox Backup Server verification.
Next guide
Continue with backup of data and offsite copies.