Files
didactyl/plans/server_health_monitoring.md

15 KiB
Raw Permalink Blame History

Server Health Monitoring for Simon (laantungir.net) — DEPLOYED

Status: ✅ Both features deployed and verified

Feature Status Details
12-hour health report ✅ Active Fired at 2026-08-25 00:00 UTC, DM sent to admin
Suspicious activity watchdog ✅ Active Timer runs every 2 min, webhook skill registered
Active triggers 3 dm_history_context, server-health-report (cron), system-alert (webhook)

Executive Summary

Two features deployed on the Simon agent on laantungir.net:

  1. 12-hour health report — a cron-triggered skill (server-health-report) that inspects the server every 12 hours and DMs the admin a status report. Zero code changes — uses the existing cron trigger type + local_shell_exec + nostr_dm tools.
  2. Suspicious activity wakeup — a systemd-timer watchdog script (didactyl-watchdog.timer) that checks CPU/memory/disk thresholds every 2 minutes in bash, and only POSTs to the agent's webhook trigger (system-alert) when a threshold is exceeded. Zero code changes.

Current State (from server investigation, 2026-08-24)

Item Finding
Host Ubuntu, 4 vCPU, 3.8 GiB RAM, 290 GB disk (57% used), up 33 days
Simon agent simon.service, didactyl v0.2.71, running since Aug 9 (2 weeks), PID 772276, user simon
Relays 3/3 connected (damus, primal, laantungir)
Signer local mode, healthy
Active triggers 3 — dm_history_context, server-health-report (cron), system-alert (webhook)
Admin API 127.0.0.1:8485, enabled, no auth (localhost only)
LLM ppq / claude-opus-4.6 (triggered skills use opus), max_tokens 800/1024
Shell tool enabled by default (config.c:1589), 30s timeout, 64KB output cap, cwd .
Memory pressure Already present: 939 MiB swap in use, kswapd0 constantly active, 213 MiB free / 2.4 GiB available — the memory monitor is genuinely useful today
Load ~1.0 on 4 cores — healthy
Notable processes gitea, postgres (c-relay-pg), nginx, fips
debug.log 219 MB and growing — worth adding logrotate (bonus item)
simon group Added to adm group so journalctl/dmesg work in shell tool

Key architecture facts that make both features cheap:

  • Cron triggers are fully implemented: a skill with ["trigger","cron"] + ["filter","0 */12 * * *"] tags is polled every 30s by trigger_manager_poll() and fired through execute_llm_action() → agent_on_trigger().
  • Webhook triggers are fully implemented: handle_trigger_webhook() accepts POST /api/trigger/<d_tag> with a JSON payload and fires the skill synchronously.
  • skill_create supports trigger + filter + auto_adopt — skills can be created live by DMing the agent, no restart needed.
  • Triggered executions get full tool access (admin tier), including local_shell_exec and nostr_dm.

Feature 1: 12-Hour Server Health Report (cron skill)

Architecture

flowchart TD
    CRON[cron schedule 0 slash 12 star star star] --> POLL[trigger_manager_poll every 30s]
    POLL --> MATCH[cron_matches_now]
    MATCH --> LLM[agent_on_trigger - skill context]
    LLM --> SHELL[local_shell_exec - uptime free df systemctl ps curl]
    SHELL --> ASSESS[LLM assesses metrics]
    ASSESS --> DM[nostr_dm to admin - concise report]

Skill definition

  • d_tag: server-health-report
  • kind: 31124 (private), scope private
  • trigger: cron
  • filter: 0 */12 * * * (fires 00:00 and 12:00 UTC = 8pm/8am America/New_York)
  • action: llm (default)
  • auto_adopt: true

Skill content (markdown, self-contained — no {{...}} references):

## Server Health Report

You are the Simon agent running on host laantungir.net. This skill fires every 12 hours.
Check server health and DM the administrator a concise report.

### Steps

1. Gather metrics with `local_shell_exec` (run each as a separate call):
   - `uptime`
   - `free -m`
   - `df -h /`
   - `systemctl --failed --no-pager`
   - `systemctl is-active simon.service didactyl.service nginx gitea postgresql`
   - `ps aux --sort=-%cpu | head -8`
   - `ps aux --sort=-%mem | head -8`
   - `curl -s http://127.0.0.1:8485/api/status`
2. Assess the results. Flag anything abnormal:
   - 1-min load average above 3.0 (host has 4 cores)
   - available memory below 500 MB or heavy swap activity
   - root disk usage above 85 percent
   - any failed systemd units
   - key services not active
   - agent relays disconnected or signer unhealthy (from the API status)
3. DM the administrator with `nostr_dm`:
   - First line: overall verdict — `HEALTHY` or `ISSUES FOUND`
   - Then one short line per area: load, memory, disk, services, agent self-check
   - Only elaborate on problems; include actual numbers
   - Keep the whole report under about 15 lines

### Rules

- Base every statement on actual tool output. Never invent or estimate numbers.
- If a command fails, report the failure instead of skipping it silently.
- DM only. Do not post anything publicly.

Deployment

Deployed via POST /api/prompt/agent on the server (local HTTP API) — Simon received the instruction and called skill_create with the parameters above. No restart needed.

Verification

  • active_triggers went from 1 → 3 (dm + cron + webhook)
  • debug.log shows: cron trigger matched d_tag=server-health-report expr=0 */12 * * * at 2026-08-25 00:00:20
  • The agent ran all 8 local_shell_exec calls (uptime, free, df, systemctl --failed, systemctl is-active, ps cpu, ps mem, curl api/status)
  • A kind 4 DM was published to all 3 relays at 00:01:12 UTC
  • Next fire: 2026-08-25 12:00 UTC

Known risks (from cron_trigger_robustness.md)

  • The EOSE-timeout startup abort and adoption-gate issues documented there are historical; v0.2.71 includes the fixes (shorthand cron expansion, last_poll_at init). Simon has been stable for 2 weeks.
  • The 30s poll throttle + 50s dedup guard mean the fire window is checked at least once per matching minute — safe for a 12-hour schedule.
  • If cron proves unreliable in practice, the watchdog mechanism from Feature 2 can also drive the 12-hour report (change its schedule), giving a fully systemd-timer-based fallback.

Feature 2: Wake the Agent on Suspicious Activity

Architecture: systemd watchdog + webhook trigger — zero code changes

A root-owned systemd timer runs a tiny shell script every 2 minutes. The script does the threshold math in bash (no LLM cost), and only when a threshold is violated does it POST to the agent's existing webhook endpoint, waking the agent to investigate and DM the admin.

flowchart TD
    TIMER[systemd timer every 2 min] --> WD[watchdog script as root]
    WD --> CHK{thresholds exceeded?}
    CHK -->|no| EXIT[exit silently - zero cost]
    CHK -->|yes| COOL{cooldown 30 min passed?}
    COOL -->|no| EXIT
    COOL -->|yes| POST[POST localhost 8485 slash api slash trigger slash system-alert]
    POST --> SKILL[webhook-triggered skill fires]
    SKILL --> INV[LLM investigates via local_shell_exec]
    INV --> DM2[nostr_dm to admin with findings]

Why this shape:

  • Threshold evaluation happens in bash — checking /proc/loadavg and /proc/meminfo every 2 minutes costs nothing and never touches the LLM.
  • The agent is only woken when something is actually wrong — exactly the "wake up on suspicious activity" semantics requested.
  • Uses the already-implemented webhook trigger path (handle_trigger_webhook()) end to end.
  • The 30-minute cooldown lives in the watchdog (state file), preventing alert spam during sustained incidents; the agent-side 60s trigger cooldown is a second layer.

Alert skill definition:

  • d_tag: system-alert
  • trigger: webhook, filter: {}, auto_adopt: true

Skill content:

## System Alert Investigation

This skill fires when the host watchdog detects suspicious system activity
(high CPU load, memory pressure, or disk pressure). The triggering event
payload contains the metrics that tripped the alert.

### Steps

1. Read the alert payload from the triggering event.
2. Investigate immediately with `local_shell_exec`:
   - `uptime`
   - `ps aux --sort=-%cpu | head -10`
   - `ps aux --sort=-%mem | head -10`
   - `free -m`
   - `dmesg | tail -30`
   - `journalctl -p err --since "1 hour ago" --no-pager | tail -30`
   - `ss -tunap | head -30`
3. Judge the situation:
   - Benign: scheduled work (backups, builds, apt jobs)
   - Problem: runaway process, OOM kills, leak
   - Suspicious: unknown processes, unexpected listeners or connections
4. DM the administrator immediately with:
   - What tripped the alert (from the payload)
   - What you found (top offenders with real numbers)
   - Your assessment and a recommended action
   - If it looks malicious, say so prominently and suggest immediate steps

### Rules

- Investigate before alerting. Include real data, not guesses.
- Never take destructive action (kill, rm, reboot, service stop) without
  an explicit instruction from the administrator.
- If a command fails or is permission-restricted, say so in the DM.
- DM only. Do not post publicly.

Watchdog script — /usr/local/bin/didactyl-system-watchdog.sh:

#!/bin/bash
# Fires the Simon agent's system-alert webhook skill when thresholds are exceeded.
set -u

API_URL="http://127.0.0.1:8485/api/trigger/system-alert"
STATE_DIR="/var/lib/didactyl-watchdog"
COOLDOWN_SECONDS=1800        # 30 min between alerts
LOAD1_THRESHOLD=3.2          # 80% of 4 cores, 1-min average
MEM_AVAILABLE_MIN_MB=400     # below this = memory pressure
DISK_ROOT_MAX_PERCENT=90

mkdir -p "$STATE_DIR"
LAST_FILE="$STATE_DIR/last_alert"

now=$(date +%s)
last=$(cat "$LAST_FILE" 2>/dev/null || echo 0)
if [ $((now - last)) -lt "$COOLDOWN_SECONDS" ]; then
    exit 0
fi

read -r load1 _ < /proc/loadavg
mem_avail_mb=$(( $(awk '/MemAvailable/{print $2}' /proc/meminfo) / 1024 ))
disk_pct=$(df --output=pcent / | tail -1 | tr -dc '0-9')

violations=""
exceeded=0

if awk -v l="$load1" -v t="$LOAD1_THRESHOLD" 'BEGIN{exit !(l>=t)}'; then
    violations="${violations}\"cpu_load_1m\":${load1},"
    exceeded=1
fi
if [ "$mem_avail_mb" -lt "$MEM_AVAILABLE_MIN_MB" ]; then
    violations="${violations}\"mem_available_mb\":${mem_avail_mb},"
    exceeded=1
fi
if [ "$disk_pct" -ge "$DISK_ROOT_MAX_PERCENT" ]; then
    violations="${violations}\"disk_root_percent\":${disk_pct},"
    exceeded=1
fi

if [ "$exceeded" -eq 0 ]; then
    exit 0
fi

violations="${violations%,}"
payload="{\"type\":\"system_alert\",\"source\":\"watchdog\",\"timestamp\":${now},\"load_1m\":${load1},\"mem_available_mb\":${mem_avail_mb},\"disk_root_percent\":${disk_pct},\"violations\":{${violations}}}"

# Only record the cooldown if the agent accepted the webhook
if curl -s -m 10 -X POST "$API_URL" -H 'Content-Type: application/json' -d "$payload"; then
    echo "$now" > "$LAST_FILE"
fi

exit 0

systemd units:

# /etc/systemd/system/didactyl-watchdog.service
[Unit]
Description=Didactyl system activity watchdog

[Service]
Type=oneshot
ExecStart=/usr/local/bin/didactyl-system-watchdog.sh
# /etc/systemd/system/didactyl-watchdog.timer
[Unit]
Description=Run Didactyl watchdog every 2 minutes

[Timer]
OnBootSec=2min
OnUnitActiveSec=2min
AccuracySec=30s

[Install]
WantedBy=timers.target

Verification

  • POST /api/trigger/system-alert with test payload → {"success":true,"d_tag":"system-alert","fired":true}
  • Agent investigated in real-time: ran uptime, ps, free, dmesg, journalctl, ss — then DM'd admin with findings
  • didactyl-watchdog.timer active (waiting), fires every 2 minutes
  • Watchdog script syntax-validated (bash -n) and test-run cleanly (exit 0, thresholds not exceeded)

Threshold tuning note

The host already shows memory pressure (939 MiB swap used, kswapd0 active). With MemAvailable around 2.4 GiB today, a 400 MB floor is a sensible early-warning line. Use sar -q/sar -r (sysstat is already collecting) to baseline before tightening.

Option B (future enhancement): native metric trigger type — code changes

If we later want the agent fully self-contained (no systemd dependency), add a metric trigger type following the pattern in new_trigger_types.md:

  1. Add TRIGGER_TYPE_METRIC to the enum in trigger_manager.h and string conversion helpers.
  2. Filter format: {"cpu_load_1m":3.2,"mem_available_mb":400,"disk_root_percent":90} — any condition matching fires.
  3. In trigger_manager_poll() (already runs every 30s), read /proc/loadavg + /proc/meminfo, evaluate thresholds per metric trigger, and require N consecutive violations (add a violation_streak counter to active_trigger_t) to avoid flapping on transient spikes.
  4. Fire via the existing execute_llm_action() with a synthetic event carrying the metrics — identical downstream path to cron/webhook.
  5. Add metric to the skill_create schema (tools_schema.c), trigger_type_from_string, and docs.

Estimated scope: ~250–350 lines in trigger_manager.c/.h plus schema and docs. Not needed for the initial rollout — Option A delivers the same behavior today with zero code changes.

Why not a fine-grained cron skill for monitoring?

A */5 * * * * cron skill with LLM action would burn ~288 LLM calls/day just to conclude "everything is fine", and the template action type can't do threshold math (it only interpolates event fields). The watchdog keeps the LLM out of the loop until there is something to say.


Implementation Log

  1. ✅ Investigated server — SSH'd into laantungir.net, checked systemd services, API status, genesis config, debug logs
  2. ✅ Wrote plan — plans/server_health_monitoring.md
  3. ✅ Created server-health-report skill — via POST /api/prompt/agent → Simon called skill_create with cron trigger 0 */12 * * *, auto-adopted
  4. ✅ Created system-alert skill — via same prompt, webhook trigger {}, auto-adopted
  5. ✅ Verified triggers — active_triggers: 3 on /api/status
  6. ✅ Installed watchdog — script at /usr/local/bin/didactyl-system-watchdog.sh, service + timer units, timer enabled
  7. ✅ Tested webhook — manual POST to /api/trigger/system-alert → agent investigated and DM'd admin
  8. ✅ Verified health report fired — cron trigger matched d_tag=server-health-report at 00:00 UTC, kind 4 DM published to all 3 relays
  9. ✅ Fixed watchdog syntax — heredoc mangling caused a bash syntax error; corrected and verified with bash -n
  10. ✅ Added simon to adm group — so journalctl/dmesg work in shell tool calls

Verification Checklist

  • /api/status shows active_triggers: 3 (dm + cron + webhook)
  • debug.log shows cron trigger matched d_tag=server-health-report
  • Admin received health report DM at 00:00 UTC (kind 4 published to all relays)
  • Manual curl -X POST .../api/trigger/system-alert produced investigation DM
  • systemctl list-timers shows didactyl-watchdog.timer active
  • Cooldown works: second violation within 30 min does not re-alert
  • Watchdog script syntax-validated and test-run cleanly