Guide 24 of 25 · advanced

Safely auto-restart one disposable service

Use systemd's bounded restart policy, observe failure, and stop retry storms.

Source-verified · Not lab-testedOfficial sources checked 2026-09-29. No PebbleRack hardware compatibility claim.

Outcome

You will configure a bounded systemd restart policy for one service whose restart is known to be safe, then test both recovery and a persistent failure. This is the smallest useful automation loop: detect process exit, act once or a few times, and escalate when it does not recover. It cannot diagnose data corruption or decide whether a restart is safe for a database. This draft has not been executed on a live lab.

Before you start

Choose a disposable stateless service on a Linux guest, not Proxmox host daemons, storage, databases, or services that may have in-flight writes. Know its unit name and have console access. Take a snapshot or have a tested backup of the guest. Record its current unit definition using systemctl cat SERVICE.service, current status, and recent logs. Decide a maximum number of restarts and an alert path before editing. If an application needs its own recovery semantics, use those rather than a generic restart.

Steps

  1. Confirm that stopping and starting this service manually does not lose data or break dependencies. Run the application's normal health check before and after one manual restart. If you cannot prove that, stop here. Process existence is not enough; use an actual request or functional check.
  2. Add an override with sudo systemctl edit SERVICE.service. In the edit window, put the following values under [Service]: Restart=on-failure and RestartSec=10s. Under [Unit], add StartLimitIntervalSec=5min and StartLimitBurst=3. These values are an example policy, not universal defaults. The service will stop retrying after repeated starts within the interval; check your installed systemd version's unit semantics.
  3. Run sudo systemctl daemon-reload, then systemctl cat SERVICE.service and systemctl show SERVICE.service -p Restart -p RestartUSec -p StartLimitBurst -p StartLimitIntervalUSec to inspect the effective configuration. Validate the unit if your systemd version offers systemd-analyze verify. Keep the original override text so you can reverse it.
  4. Arrange an alert when the unit reaches a failed state or the application health check fails. A restart that succeeds briefly may conceal an ongoing problem, so monitor restart counts and the user-facing request. Set a daily review of the journal for this service. Do not allow an AI agent to call systemctl restart independently as a second unbounded loop.
  5. Test on the disposable service. Create a controlled failure using the service's own test mode or a temporary invalid non-secret setting that you can immediately reverse. Observe the bounded restart attempts and the final state with systemctl status and journalctl -u SERVICE.service --since .... Restore the valid setting, run sudo systemctl reset-failed SERVICE.service if the start limit was reached, and start it manually. Confirm the application works.
  6. Write the measured result: number of retries, time to recover or stop, exact error, alert delivery, and who intervened. If the controlled failure could not be performed safely, mark the restart behavior unverified and leave the policy disabled until a safe test is available.

Check it worked

A transient controlled failure recovers within the permitted retries and a persistent one stops retrying and alerts a human. The application-level health check passes after repair. A systemd “active” state alone does not pass. Verify no hidden restart loop remains in the application, container runtime, or another supervisor.

If it fails / rollback

If the service degrades or loops, remove only the override file you created (normally /etc/systemd/system/SERVICE.service.d/override.conf) after saving a copy privately, or restore its prior contents if it existed; then run sudo systemctl daemon-reload and confirm effective settings. Avoid changing unrelated drop-ins. Restore the known-good application configuration and perform the functional check. Use console access if remote connectivity depends on the service.

Safety and data notes

Automatic restart is suitable only where the operation is idempotent and bounded. Never use it as a cure for full disks, corrupt databases, malware, or repeated unexplained crashes. Preserve the first failure log; retries can overwrite useful context. Treat alerting and rollback as part of the automation, not an afterthought.

Sources

Official sources checked 2026-09-29: systemd.service manual; systemd.unit rate limiting; Ubuntu service management.

Next guide

Continue with the guarded graph loop.