15 KiB
Server Health Monitoring for Simon (laantungir.net) — DEPLOYED
Status: ✅ Both features deployed and verified
| Feature | Status | Details |
|---|---|---|
| 12-hour health report | ✅ Active | Fired at 2026-08-25 00:00 UTC, DM sent to admin |
| Suspicious activity watchdog | ✅ Active | Timer runs every 2 min, webhook skill registered |
| Active triggers | 3 | dm_history_context, server-health-report (cron), system-alert (webhook) |
Executive Summary
Two features deployed on the Simon agent on laantungir.net:
- 12-hour health report — a cron-triggered skill (
server-health-report) that inspects the server every 12 hours and DMs the admin a status report. Zero code changes — uses the existing cron trigger type +local_shell_exec+nostr_dmtools. - Suspicious activity wakeup — a systemd-timer watchdog script (
didactyl-watchdog.timer) that checks CPU/memory/disk thresholds every 2 minutes in bash, and only POSTs to the agent's webhook trigger (system-alert) when a threshold is exceeded. Zero code changes.
Current State (from server investigation, 2026-08-24)
| Item | Finding |
|---|---|
| Host | Ubuntu, 4 vCPU, 3.8 GiB RAM, 290 GB disk (57% used), up 33 days |
| Simon agent | simon.service, didactyl v0.2.71, running since Aug 9 (2 weeks), PID 772276, user simon |
| Relays | 3/3 connected (damus, primal, laantungir) |
| Signer | local mode, healthy |
| Active triggers | 3 — dm_history_context, server-health-report (cron), system-alert (webhook) |
| Admin API | 127.0.0.1:8485, enabled, no auth (localhost only) |
| LLM | ppq / claude-opus-4.6 (triggered skills use opus), max_tokens 800/1024 |
| Shell tool | enabled by default (config.c:1589), 30s timeout, 64KB output cap, cwd . |
| Memory pressure | Already present: 939 MiB swap in use, kswapd0 constantly active, 213 MiB free / 2.4 GiB available — the memory monitor is genuinely useful today |
| Load | ~1.0 on 4 cores — healthy |
| Notable processes | gitea, postgres (c-relay-pg), nginx, fips |
debug.log |
219 MB and growing — worth adding logrotate (bonus item) |
simon group |
Added to adm group so journalctl/dmesg work in shell tool |
Key architecture facts that make both features cheap:
- Cron triggers are fully implemented: a skill with
["trigger","cron"]+["filter","0 */12 * * *"]tags is polled every 30s bytrigger_manager_poll()and fired throughexecute_llm_action()→agent_on_trigger(). - Webhook triggers are fully implemented:
handle_trigger_webhook()acceptsPOST /api/trigger/<d_tag>with a JSON payload and fires the skill synchronously. skill_createsupportstrigger+filter+auto_adopt— skills can be created live by DMing the agent, no restart needed.- Triggered executions get full tool access (admin tier), including
local_shell_execandnostr_dm.
Feature 1: 12-Hour Server Health Report (cron skill)
Architecture
flowchart TD
CRON[cron schedule 0 slash 12 star star star] --> POLL[trigger_manager_poll every 30s]
POLL --> MATCH[cron_matches_now]
MATCH --> LLM[agent_on_trigger - skill context]
LLM --> SHELL[local_shell_exec - uptime free df systemctl ps curl]
SHELL --> ASSESS[LLM assesses metrics]
ASSESS --> DM[nostr_dm to admin - concise report]
Skill definition
- d_tag:
server-health-report - kind: 31124 (private), scope private
- trigger:
cron - filter:
0 */12 * * *(fires 00:00 and 12:00 UTC = 8pm/8am America/New_York) - action:
llm(default) - auto_adopt: true
Skill content (markdown, self-contained — no {{...}} references):
## Server Health Report
You are the Simon agent running on host laantungir.net. This skill fires every 12 hours.
Check server health and DM the administrator a concise report.
### Steps
1. Gather metrics with `local_shell_exec` (run each as a separate call):
- `uptime`
- `free -m`
- `df -h /`
- `systemctl --failed --no-pager`
- `systemctl is-active simon.service didactyl.service nginx gitea postgresql`
- `ps aux --sort=-%cpu | head -8`
- `ps aux --sort=-%mem | head -8`
- `curl -s http://127.0.0.1:8485/api/status`
2. Assess the results. Flag anything abnormal:
- 1-min load average above 3.0 (host has 4 cores)
- available memory below 500 MB or heavy swap activity
- root disk usage above 85 percent
- any failed systemd units
- key services not active
- agent relays disconnected or signer unhealthy (from the API status)
3. DM the administrator with `nostr_dm`:
- First line: overall verdict — `HEALTHY` or `ISSUES FOUND`
- Then one short line per area: load, memory, disk, services, agent self-check
- Only elaborate on problems; include actual numbers
- Keep the whole report under about 15 lines
### Rules
- Base every statement on actual tool output. Never invent or estimate numbers.
- If a command fails, report the failure instead of skipping it silently.
- DM only. Do not post anything publicly.
Deployment
Deployed via POST /api/prompt/agent on the server (local HTTP API) — Simon received the instruction and called skill_create with the parameters above. No restart needed.
Verification
active_triggerswent from 1 → 3 (dm + cron + webhook)debug.logshows:cron trigger matched d_tag=server-health-report expr=0 */12 * * *at 2026-08-25 00:00:20- The agent ran all 8
local_shell_execcalls (uptime, free, df, systemctl --failed, systemctl is-active, ps cpu, ps mem, curl api/status) - A kind 4 DM was published to all 3 relays at 00:01:12 UTC
- Next fire: 2026-08-25 12:00 UTC
Known risks (from cron_trigger_robustness.md)
- The EOSE-timeout startup abort and adoption-gate issues documented there are historical; v0.2.71 includes the fixes (shorthand cron expansion,
last_poll_atinit). Simon has been stable for 2 weeks. - The 30s poll throttle + 50s dedup guard mean the fire window is checked at least once per matching minute — safe for a 12-hour schedule.
- If cron proves unreliable in practice, the watchdog mechanism from Feature 2 can also drive the 12-hour report (change its schedule), giving a fully systemd-timer-based fallback.
Feature 2: Wake the Agent on Suspicious Activity
Architecture: systemd watchdog + webhook trigger — zero code changes
A root-owned systemd timer runs a tiny shell script every 2 minutes. The script does the threshold math in bash (no LLM cost), and only when a threshold is violated does it POST to the agent's existing webhook endpoint, waking the agent to investigate and DM the admin.
flowchart TD
TIMER[systemd timer every 2 min] --> WD[watchdog script as root]
WD --> CHK{thresholds exceeded?}
CHK -->|no| EXIT[exit silently - zero cost]
CHK -->|yes| COOL{cooldown 30 min passed?}
COOL -->|no| EXIT
COOL -->|yes| POST[POST localhost 8485 slash api slash trigger slash system-alert]
POST --> SKILL[webhook-triggered skill fires]
SKILL --> INV[LLM investigates via local_shell_exec]
INV --> DM2[nostr_dm to admin with findings]
Why this shape:
- Threshold evaluation happens in bash — checking
/proc/loadavgand/proc/meminfoevery 2 minutes costs nothing and never touches the LLM. - The agent is only woken when something is actually wrong — exactly the "wake up on suspicious activity" semantics requested.
- Uses the already-implemented webhook trigger path (
handle_trigger_webhook()) end to end. - The 30-minute cooldown lives in the watchdog (state file), preventing alert spam during sustained incidents; the agent-side 60s trigger cooldown is a second layer.
Alert skill definition:
- d_tag:
system-alert - trigger:
webhook, filter:{}, auto_adopt: true
Skill content:
## System Alert Investigation
This skill fires when the host watchdog detects suspicious system activity
(high CPU load, memory pressure, or disk pressure). The triggering event
payload contains the metrics that tripped the alert.
### Steps
1. Read the alert payload from the triggering event.
2. Investigate immediately with `local_shell_exec`:
- `uptime`
- `ps aux --sort=-%cpu | head -10`
- `ps aux --sort=-%mem | head -10`
- `free -m`
- `dmesg | tail -30`
- `journalctl -p err --since "1 hour ago" --no-pager | tail -30`
- `ss -tunap | head -30`
3. Judge the situation:
- Benign: scheduled work (backups, builds, apt jobs)
- Problem: runaway process, OOM kills, leak
- Suspicious: unknown processes, unexpected listeners or connections
4. DM the administrator immediately with:
- What tripped the alert (from the payload)
- What you found (top offenders with real numbers)
- Your assessment and a recommended action
- If it looks malicious, say so prominently and suggest immediate steps
### Rules
- Investigate before alerting. Include real data, not guesses.
- Never take destructive action (kill, rm, reboot, service stop) without
an explicit instruction from the administrator.
- If a command fails or is permission-restricted, say so in the DM.
- DM only. Do not post publicly.
Watchdog script — /usr/local/bin/didactyl-system-watchdog.sh:
#!/bin/bash
# Fires the Simon agent's system-alert webhook skill when thresholds are exceeded.
set -u
API_URL="http://127.0.0.1:8485/api/trigger/system-alert"
STATE_DIR="/var/lib/didactyl-watchdog"
COOLDOWN_SECONDS=1800 # 30 min between alerts
LOAD1_THRESHOLD=3.2 # 80% of 4 cores, 1-min average
MEM_AVAILABLE_MIN_MB=400 # below this = memory pressure
DISK_ROOT_MAX_PERCENT=90
mkdir -p "$STATE_DIR"
LAST_FILE="$STATE_DIR/last_alert"
now=$(date +%s)
last=$(cat "$LAST_FILE" 2>/dev/null || echo 0)
if [ $((now - last)) -lt "$COOLDOWN_SECONDS" ]; then
exit 0
fi
read -r load1 _ < /proc/loadavg
mem_avail_mb=$(( $(awk '/MemAvailable/{print $2}' /proc/meminfo) / 1024 ))
disk_pct=$(df --output=pcent / | tail -1 | tr -dc '0-9')
violations=""
exceeded=0
if awk -v l="$load1" -v t="$LOAD1_THRESHOLD" 'BEGIN{exit !(l>=t)}'; then
violations="${violations}\"cpu_load_1m\":${load1},"
exceeded=1
fi
if [ "$mem_avail_mb" -lt "$MEM_AVAILABLE_MIN_MB" ]; then
violations="${violations}\"mem_available_mb\":${mem_avail_mb},"
exceeded=1
fi
if [ "$disk_pct" -ge "$DISK_ROOT_MAX_PERCENT" ]; then
violations="${violations}\"disk_root_percent\":${disk_pct},"
exceeded=1
fi
if [ "$exceeded" -eq 0 ]; then
exit 0
fi
violations="${violations%,}"
payload="{\"type\":\"system_alert\",\"source\":\"watchdog\",\"timestamp\":${now},\"load_1m\":${load1},\"mem_available_mb\":${mem_avail_mb},\"disk_root_percent\":${disk_pct},\"violations\":{${violations}}}"
# Only record the cooldown if the agent accepted the webhook
if curl -s -m 10 -X POST "$API_URL" -H 'Content-Type: application/json' -d "$payload"; then
echo "$now" > "$LAST_FILE"
fi
exit 0
systemd units:
# /etc/systemd/system/didactyl-watchdog.service
[Unit]
Description=Didactyl system activity watchdog
[Service]
Type=oneshot
ExecStart=/usr/local/bin/didactyl-system-watchdog.sh
# /etc/systemd/system/didactyl-watchdog.timer
[Unit]
Description=Run Didactyl watchdog every 2 minutes
[Timer]
OnBootSec=2min
OnUnitActiveSec=2min
AccuracySec=30s
[Install]
WantedBy=timers.target
Verification
POST /api/trigger/system-alertwith test payload →{"success":true,"d_tag":"system-alert","fired":true}- Agent investigated in real-time: ran uptime, ps, free, dmesg, journalctl, ss — then DM'd admin with findings
didactyl-watchdog.timeractive (waiting), fires every 2 minutes- Watchdog script syntax-validated (
bash -n) and test-run cleanly (exit 0, thresholds not exceeded)
Threshold tuning note
The host already shows memory pressure (939 MiB swap used, kswapd0 active). With MemAvailable around 2.4 GiB today, a 400 MB floor is a sensible early-warning line. Use sar -q/sar -r (sysstat is already collecting) to baseline before tightening.
Option B (future enhancement): native metric trigger type — code changes
If we later want the agent fully self-contained (no systemd dependency), add a metric trigger type following the pattern in new_trigger_types.md:
- Add
TRIGGER_TYPE_METRICto the enum intrigger_manager.hand string conversion helpers. - Filter format:
{"cpu_load_1m":3.2,"mem_available_mb":400,"disk_root_percent":90}— any condition matching fires. - In
trigger_manager_poll()(already runs every 30s), read/proc/loadavg+/proc/meminfo, evaluate thresholds per metric trigger, and require N consecutive violations (add aviolation_streakcounter toactive_trigger_t) to avoid flapping on transient spikes. - Fire via the existing
execute_llm_action()with a synthetic event carrying the metrics — identical downstream path to cron/webhook. - Add
metricto theskill_createschema (tools_schema.c),trigger_type_from_string, and docs.
Estimated scope: ~250–350 lines in trigger_manager.c/.h plus schema and docs. Not needed for the initial rollout — Option A delivers the same behavior today with zero code changes.
Why not a fine-grained cron skill for monitoring?
A */5 * * * * cron skill with LLM action would burn ~288 LLM calls/day just to conclude "everything is fine", and the template action type can't do threshold math (it only interpolates event fields). The watchdog keeps the LLM out of the loop until there is something to say.
Implementation Log
- ✅ Investigated server — SSH'd into laantungir.net, checked systemd services, API status, genesis config, debug logs
- ✅ Wrote plan —
plans/server_health_monitoring.md - ✅ Created
server-health-reportskill — viaPOST /api/prompt/agent→ Simon calledskill_createwith cron trigger0 */12 * * *, auto-adopted - ✅ Created
system-alertskill — via same prompt, webhook trigger{}, auto-adopted - ✅ Verified triggers —
active_triggers: 3on/api/status - ✅ Installed watchdog — script at
/usr/local/bin/didactyl-system-watchdog.sh, service + timer units, timer enabled - ✅ Tested webhook — manual POST to
/api/trigger/system-alert→ agent investigated and DM'd admin - ✅ Verified health report fired —
cron trigger matched d_tag=server-health-reportat 00:00 UTC, kind 4 DM published to all 3 relays - ✅ Fixed watchdog syntax — heredoc mangling caused a bash syntax error; corrected and verified with
bash -n - ✅ Added simon to adm group — so
journalctl/dmesgwork in shell tool calls
Verification Checklist
/api/statusshowsactive_triggers: 3(dm + cron + webhook)debug.logshowscron trigger matched d_tag=server-health-report- Admin received health report DM at 00:00 UTC (kind 4 published to all relays)
- Manual
curl -X POST .../api/trigger/system-alertproduced investigation DM systemctl list-timersshowsdidactyl-watchdog.timeractive- Cooldown works: second violation within 30 min does not re-alert
- Watchdog script syntax-validated and test-run cleanly