Guide 25 of 25 · advanced

Design a guarded graph loop for lab operations

Connect observation, diagnosis, approval, action, verification, and rollback without granting broad autonomy.

Source-verified · Not lab-testedOfficial sources checked 2026-09-29. No PebbleRack hardware compatibility claim.

Outcome

You will have a design for a future, guarded operations loop that can propose a bounded repair and record what happened. This is architecture and safe pseudocode, not a deployable PebbleRack feature or evidence that AI can autonomously heal a home lab. A graph makes state transitions explicit: observe → diagnose → propose → approve → act → verify → rollback or close. Each edge needs a stop condition and a durable event record. The first executable action, if ever built, should be the narrowly tested single-service restart from guide 24.

Before you start

Complete monitoring, backup/restore, change logging, and incident-runbook guides first. Select one disposable workload with a deterministic functional health check and known safe restart. Have independent out-of-band access, a tested recovery point, and an operator who can approve and stop actions. Define the maximum allowed actions, attempt count, wait time, and what evidence will count as recovery. Proxmox administrative APIs can control powerful resources; a broad root token is inappropriate for an experimental agent.

Steps

  1. Observe. Ingest a limited event with source, UTC timestamp, host, service, severity, correlation ID, and raw task/log reference. Keep raw logs private. Require two independent signs where practical: e.g., service failure plus failed application request. Deduplicate repeated alerts into one incident. An AI-written summary is a convenience, not a replacement for the source event.
  2. Diagnose. Gather read-only evidence using a restricted account: recent service logs, Proxmox task state, disk availability, and latest backup status. Classify the condition as known transient, unknown, or unsafe. Return “unknown” readily. Model output must be treated as untrusted text; do not execute commands it writes or allow it to invent a host ID. Store the model/version and prompts if a model assists, without exposing secrets.
  3. Propose. Match the condition to a preapproved runbook ID. The proposal contains exact target, one allowlisted action, expected impact, verification query, timeout, rollback, and evidence supporting the match. If no runbook matches, create a human investigation task. Never let a model author arbitrary shell or Proxmox commands for execution.
  4. Approve. Show the operator the raw signals, proposed action, expected impact, last backup/restore evidence, and stop conditions. Require explicit approval tied to incident ID, action hash, target, and expiry. Approval of one action does not authorize subsequent steps. For the first version, every write requires human approval. Dry-run or simulation results should be labeled clearly.
  5. Act. An executor with minimum permissions rechecks preconditions at execution time: target still matches, incident is active, no maintenance lock exists, attempts remain, and the action has not already run. Apply the allowlisted single-service restart at most once. Record API response and audit ID. If the request outcome is ambiguous, stop rather than retrying a potentially duplicate action.
  6. Verify and rollback. After a bounded delay, run the functional health check from outside the target, inspect restart count and logs, and compare to the baseline. If verification fails, stop automation and follow the prewritten manual rollback or escalation path. Do not turn the loop into repeated reboots. Close only when a human or explicitly defined policy accepts the evidence; review false positives and failure modes after each exercise.

Safe pseudocode for a simulation:

on incident(event):
  evidence = read_only_collect(event)
  proposal = allowlisted_runbook_match(evidence)
  if proposal is unknown or unsafe: escalate_to_human(); return
  approval = request_one_action_approval(proposal, incident_id, expiry)
  if not valid(approval) or not preconditions_still_true(): return
  result = execute_once_with_audit(proposal)
  if ambiguous(result): freeze_and_escalate(); return
  if independent_health_check_passes(): close_with_evidence()
  else: stop_and_request_manual_rollback()

Check it worked

For now, review the graph on paper or with synthetic events. It should refuse an unknown host, missing backup evidence, expired approval, duplicate event, conflicting maintenance window, and ambiguous action result. It should produce a human-readable incident trail and stop safely after failure. These are design acceptance criteria, not claims that the implementation has passed them.

If it fails / rollback

If a proposed action is unsafe or based on stale state, cancel it before execution and revise the runbook. If an approved action later fails, freeze automation for that target and use the incident runbook. Revoke executor credentials if scope is uncertain. A rollback must be chosen per action; restoring an old VM snapshot can discard newer data and must never be an automatic generic fallback.

Safety and data notes

Do not automate remediation for suspected intrusion, data corruption, full or failing storage, backup failures, network lockout, power events, unknown root cause, shared services, or anything that can destroy or alter user data. Keep Proxmox credentials server-side, restrict them to an allowlisted resource/action, and retain audit logs. An LLM should advise or summarize only until the entire loop has been independently tested against failure cases. Readers can follow the build as new tests are published; no signup is required to use these guides.

Sources

Official sources checked 2026-09-29: Proxmox VE API and administration documentation; Proxmox VE API access examples; NIST SP 800-61 Rev. 3; systemd.service restart semantics.

Next guide

This completes the first 25-guide path. Revisit the incident runbook and restore exercise before expanding automation.