From ca11518d58a7a12d9a3451eba5f57680bcc87ae5 Mon Sep 17 00:00:00 2001 From: Didactyl User Date: Thu, 27 Aug 2026 07:51:36 -0400 Subject: [PATCH] v0.2.73 - Updated nostr_core_lib to v0.6.15: migrated signer selection from nostr_index to role+role_path, added full n_signer transport support (unix/tcp/serial/fds) in wizard, config, and CLI, new signer_crypto tool, and PoW support in nostr_post --- README.md | 4 +- plans/fix_dm_delivery_during_triggers.md | 189 +++++++++++++ plans/server_health_monitoring.md | 335 +++++++++++++++++++++++ src/agent.c | 22 ++ src/main.c | 2 +- src/main.h | 9 +- src/nostr_handler.c | 32 +++ src/trigger_manager.c | 17 +- 8 files changed, 600 insertions(+), 10 deletions(-) create mode 100644 plans/fix_dm_delivery_during_triggers.md create mode 100644 plans/server_health_monitoring.md diff --git a/README.md b/README.md index c3f5a23..dc83056 100644 --- a/README.md +++ b/README.md @@ -54,11 +54,11 @@ Skills compose by adoption-list order (`10123`) and trigger tags carry runtime e Didactyl will support local inference, which is very privacy preserving. Remote inference does however have it's advantages, and in those cases Didactyl supports using Bitcoin Lightning and eCash inference providers. -## Current Status — v0.2.72 +## Current Status — v0.2.73 **Active build — this project is barely working. Experiment at your own risk.** -> Last release update: v0.2.72 — Existing agent flow: add relay configuration prompt before Nostr recovery +> Last release update: v0.2.73 — Updated nostr_core_lib to v0.6.15: migrated signer selection from nostr_index to role+role_path, added full n_signer transport support (unix/tcp/serial/fds) in wizard, config, and CLI, new signer_crypto tool, and PoW support in nostr_post - Connects to configured relays with auto-reconnect and relay state transition logging - Publishes configured startup events per relay as each relay becomes connected diff --git a/plans/fix_dm_delivery_during_triggers.md b/plans/fix_dm_delivery_during_triggers.md new file mode 100644 index 0000000..3a32705 --- /dev/null +++ b/plans/fix_dm_delivery_during_triggers.md @@ -0,0 +1,189 @@ +# FIXED: DM Delivery Failure During Cron-Triggered Skills + +> **Status: ✅ Fixed and deployed in v0.2.72 (2026-08-25/26)** +> +> Five fixes applied: +> 1. [`src/agent.c`](src/agent.c) — `nostr_handler_poll(0)` calls in the `agent_on_trigger()` turn loop and tool-execution loop keep websockets serviced during long-running triggered skills +> 2. [`src/nostr_handler.c`](src/nostr_handler.c) — self-healing retry in `nostr_handler_send_dm_with_role()`: if a publish is accepted by 0 relays, service the pool for 500ms and retry once +> 3. [`src/trigger_manager.c`](src/trigger_manager.c) — periodic trigger reconciliation (every 60s) from the local self-skill cache, closing the EOSE/adoption-gate race that prevented triggers from loading after restart +> 4. [`src/agent.c`](src/agent.c) + [`src/main.c`](src/main.c) + [`src/main.h`](src/main.h) — `systemd_notify_send("WATCHDOG=1")` pings inside the trigger execution loop. The service unit has `WatchdogSec=120` with `Type=notify`; the main loop's 10s ping cadence is suspended during blocking trigger executions, so systemd was killing the service mid-investigation (observed at 18:35:14). Pings now fire per-turn and per-tool-call. +> 5. [`nostr_core_lib/nostr_core/core_relay_pool.c`](nostr_core_lib/nostr_core/core_relay_pool.c) — dead-transport detection in `check_connection_health()`. When `nostr_ws_ping()` fails to send (remote closed the TCP connection), the ws client's cached state still claims CONNECTED (it only transitions on a clean WebSocket CLOSE frame), so the pool never marked the relay disconnected and never reconnected — the agent sat with zero TCP sockets for 13+ hours while reporting 3/3 connected (observed 2026-08-26: DMs silently ignored from 00:08 onward). The fix closes the stale client (`ws_client = NULL`) and marks the relay DISCONNECTED on ping-send failure or pong timeout, letting the reconnect logic perform a fresh connect. +> +> Verification (2026-08-25 21:51 UTC): webhook test fired → investigation ran ~2 minutes (past the old 120s kill window) → service stayed up (no restart) → investigation DM delivered `via 3 connected relay(s)`. All 3 triggers (dm, cron, webhook) register automatically after restart via reconciliation. The skills also needed to be adopted via `skill_adopt` — the original `skill_create` with `auto_adopt: true` had failed to publish the adoption event because publishing was broken at creation time. +> +> Post-fix-5 verification (2026-08-26 13:41 UTC): service restarted with the lib fix, 3/3 relays connected, 4 live TCP sockets confirmed via `ss`. The midnight health report and a real watchdog alert (load 3.69 at 04:01) both fired correctly on 2026-08-26 — the monitoring features work; the dead-socket bug was silently swallowing all inbound DMs. + +## Symptom + +The `server-health-report` cron skill fires correctly (confirmed at 00:00 and 12:00 UTC), gathers all metrics, and the LLM generates the report — but the DM never reaches the admin. The log shows a contradictory pair: + +``` +[12:01:22] kind 4 event published to wss://relay.damus.io (async) ← targeted +[12:01:22] kind 4 event published to wss://relay.primal.net (async) +[12:01:22] kind 4 event published to wss://relay.laantungir.net (async) +[12:01:22] sent DM 2af454bd2034a08f... to 1ec454734dcbf6fe... via 0 connected relay(s) ← actually sent to ZERO +``` + +Meanwhile the webhook test DM (22 hours earlier) succeeded: `via 3 connected relay(s)`. + +## Root Cause Analysis + +### The single-threaded main loop blocks websocket servicing + +The main loop in [`main.c:2480`](src/main.c:2480) is single-threaded: + +```c +while (g_running) { + (void)nostr_handler_poll(100); // services websockets + (void)trigger_manager_poll(&trigger_manager); // BLOCKS during trigger execution + if (http_api_started) { + (void)http_api_poll(0); + } + ... +} +``` + +When a cron trigger fires, the call chain is: + +``` +trigger_manager_poll() [trigger_manager.c:1681] + → execute_llm_action() [trigger_manager.c:903] + → agent_on_trigger() [agent.c:1769] + → multi-turn LLM loop (BLOCKING) + - 8 × local_shell_exec (up to 30s each) + - 3-4 × LLM HTTP round trips + → nostr_dm_send tool call +``` + +The health report execution takes **50–72 seconds**. During this entire window, `nostr_handler_poll()` is never called — the websocket connections receive zero servicing. + +### Two connection checks disagree + +| Check | Location | Basis | Result during cron fire | +|-------|----------|-------|------------------------| +| `nostr_relay_pool_get_relay_status()` | [`nostr_handler.c:3304`](src/nostr_handler.c:3304) | `relay->status` — only updated during poll | **CONNECTED** (stale) | +| `nostr_ws_get_state()` | [`core_relay_pool.c:1714`](nostr_core_lib/nostr_core/core_relay_pool.c:1714) | live ws client state | **NOT CONNECTED** | + +In [`nostr_handler_send_dm_with_role()`](src/nostr_handler.c:3225): + +1. The relay-selection loop uses the **stale** `relay->status` → builds `connected_relays[]` with 3 entries → prints "kind 4 event published to ..." for each +2. [`nostr_relay_pool_publish_async()`](nostr_core_lib/nostr_core/core_relay_pool.c:1673) internally re-checks with the **live** `nostr_ws_get_state()` → all relays fail the check → every send is skipped → returns 0 +3. The final log line prints the actual result: `via 0 connected relay(s)` + +### Why the webhook test succeeded + +The webhook test completed in ~23 seconds — under the ~30s threshold where the unserviced websockets degrade (relay-side ping timeouts, TCP buffer stalls). The cron health report takes 50–72 seconds, well past it. + +### nostr_core_lib version check + +The vendored lib is **v0.6.15** ([`nostr_core.h:5`](nostr_core_lib/nostr_core/nostr_core.h:5)), which is the latest per [`update_nostr_core_lib_nsigner.md`](plans/update_nostr_core_lib_nsigner.md). The lib's behavior is correct — it refuses to send on connections whose live state isn't CONNECTED. The bug is in didactyl's usage pattern: blocking the poll loop for a minute. + +--- + +## Debugging Steps (in priority order) + +### Step 1: Confirm the timing hypothesis — zero risk, 5 minutes + +Create a minimal cron skill that fires every 5 minutes with ONE fast command and an immediate DM: + +```json +{ + "d": "dm-timing-test", + "trigger": "cron", + "filter": "*/5 * * * *", + "content": "Run `uptime` with local_shell_exec, then immediately DM the admin with the result using nostr_dm_send. Do nothing else." +} +``` + +- If the short skill's DM **arrives** → timing/blocking confirmed +- If it **also fails** → hypothesis wrong, go to Step 2 instrumentation +- Remove the test skill afterward (`skill_remove`) + +### Step 2: Instrument the send path — confirms the state disagreement + +Add temporary logging in [`nostr_handler_send_dm_with_role()`](src/nostr_handler.c:3225) before the publish: + +```c +for (int i = 0; i < g_cfg->relay_count; i++) { + DEBUG_INFO("[didactyl] DM pre-check: relay=%s pool_status=%d", + g_cfg->relays[i], + nostr_relay_pool_get_relay_status(g_pool, g_cfg->relays[i])); +} +``` + +And in [`nostr_relay_pool_publish_async()`](nostr_core_lib/nostr_core/core_relay_pool.c:1673), log per-relay ws state and send result: + +```c +DEBUG_INFO("[pool] publish_async relay=%s ws_state=%d", + relay_urls[i], + relay && relay->ws_client ? nostr_ws_get_state(relay->ws_client) : -1); +``` + +Expected output during a cron fire: `pool_status=2 (CONNECTED)` but `ws_state != CONNECTED` for all relays — proving the stale-vs-live disagreement. + +### Step 3: The fix — poll websockets during trigger execution + +Add non-blocking polls inside the [`agent_on_trigger()`](src/agent.c:1769) execution loop: + +```c +// In agent_on_trigger(), at the top of the per-turn loop (before llm_chat_with_tools_messages): +(void)nostr_handler_poll(0); // non-blocking: service websocket I/O + +// Also inside the tool-execution loop (after each tools_execute call), +// since local_shell_exec can block up to 30 seconds per call: +(void)nostr_handler_poll(0); +``` + +This keeps the websocket state fresh and the connections alive throughout long-running triggered skills. It also benefits DM-triggered conversations that run long tool loops. + +**Files to change:** +- [`src/agent.c`](src/agent.c) — add poll calls in `agent_on_trigger()` (and consider the same for the DM conversation path in `agent_on_message()` if it has the same pattern) + +### Step 4: Rebuild, redeploy, verify + +```bash +./build_static.sh +./deploy_lt.sh +``` + +Then watch the next 12:00 UTC fire (or use the Step 1 test skill for faster iteration): + +```bash +ssh ubuntu@laantungir.net "grep -E 'sent DM.*via' /home/simon/debug.log | tail -5" +``` + +Success criterion: `via 3 connected relay(s)` and the admin receives the report DM. + +### Step 5 (optional hardening): surface send failures to the LLM + +Currently `nostr_dm_send` returns `success: false` when `sent == 0`, but the LLM may not retry. Two options: + +1. **Retry in the tool**: in [`nostr_handler_send_dm_with_role()`](src/nostr_handler.c:3225), if `sent == 0`, call `nostr_handler_poll(100)` a few times and retry the publish once — self-healing without LLM involvement +2. **Retry via LLM**: make the tool result error message explicit (`"DM not delivered: 0 relays accepted. Retry after servicing connections."`) so the LLM retries on the next turn + +Option 1 is more robust; option 2 is simpler. + +### Step 6 (optional): check upstream nostr_core_lib + +The vendored v0.6.15 is current, but worth a quick check of the upstream repo for any post-0.6.15 changes to `core_relay_pool.c` — particularly whether `nostr_ws_get_state()` semantics changed or whether a queued-publish mechanism (send after reconnect) was added. If upstream added a publish queue, updating the vendored lib could make sends self-healing. + +--- + +## Workarounds (if a code fix must wait) + +1. **Split the skill into a chain**: `server-health-report` (gather metrics, write to file) → `chain` trigger → `server-health-report-send` (read file, DM admin). Each skill stays under ~30s, keeping websocket degradation below threshold. Fragile — depends on the exact timeout window. +2. **DM first, investigate later**: restructure the skill to send a brief DM immediately (before the long shell commands), then send findings. Only helps if the first DM goes out within ~30s of trigger fire. + +Both are stopgaps — the real fix is Step 3. + +--- + +## Verification Checklist + +- [ ] Step 1 test skill DM arrives (timing hypothesis confirmed) +- [ ] Step 2 instrumentation shows pool_status=CONNECTED but ws_state≠CONNECTED during cron fire +- [ ] Step 3 fix applied: `nostr_handler_poll(0)` calls in `agent_on_trigger()` loops +- [ ] Rebuilt and deployed via `./build_static.sh` + `./deploy_lt.sh` +- [ ] Next cron fire logs `via 3 connected relay(s)` +- [ ] Admin receives the health report DM at 00:00/12:00 UTC +- [ ] Webhook alert path still works (regression check) diff --git a/plans/server_health_monitoring.md b/plans/server_health_monitoring.md new file mode 100644 index 0000000..8708efc --- /dev/null +++ b/plans/server_health_monitoring.md @@ -0,0 +1,335 @@ +# Server Health Monitoring for Simon (laantungir.net) — DEPLOYED + +## Status: ✅ Both features deployed and verified + +| Feature | Status | Details | +|---------|--------|---------| +| 12-hour health report | ✅ Active | Fired at 2026-08-25 00:00 UTC, DM sent to admin | +| Suspicious activity watchdog | ✅ Active | Timer runs every 2 min, webhook skill registered | +| Active triggers | 3 | dm_history_context, server-health-report (cron), system-alert (webhook) | + +--- + +## Executive Summary + +Two features deployed on the Simon agent on `laantungir.net`: + +1. **12-hour health report** — a cron-triggered skill (`server-health-report`) that inspects the server every 12 hours and DMs the admin a status report. **Zero code changes** — uses the existing cron trigger type + `local_shell_exec` + `nostr_dm` tools. +2. **Suspicious activity wakeup** — a systemd-timer watchdog script (`didactyl-watchdog.timer`) that checks CPU/memory/disk thresholds every 2 minutes in bash, and only POSTs to the agent's webhook trigger (`system-alert`) when a threshold is exceeded. **Zero code changes.** + +--- + +## Current State (from server investigation, 2026-08-24) + +| Item | Finding | +|---|---| +| Host | Ubuntu, 4 vCPU, 3.8 GiB RAM, 290 GB disk (57% used), up 33 days | +| Simon agent | `simon.service`, didactyl **v0.2.71**, running since Aug 9 (2 weeks), PID 772276, user `simon` | +| Relays | 3/3 connected (damus, primal, laantungir) | +| Signer | local mode, healthy | +| Active triggers | **3** — dm_history_context, server-health-report (cron), system-alert (webhook) | +| Admin API | `127.0.0.1:8485`, enabled, no auth (localhost only) | +| LLM | ppq / claude-opus-4.6 (triggered skills use opus), max_tokens 800/1024 | +| Shell tool | enabled by default ([`config.c:1589`](src/config.c:1589)), 30s timeout, 64KB output cap, cwd `.` | +| Memory pressure | **Already present**: 939 MiB swap in use, `kswapd0` constantly active, 213 MiB free / 2.4 GiB available — the memory monitor is genuinely useful today | +| Load | ~1.0 on 4 cores — healthy | +| Notable processes | gitea, postgres (c-relay-pg), nginx, fips | +| `debug.log` | **219 MB** and growing — worth adding logrotate (bonus item) | +| `simon` group | Added to `adm` group so `journalctl`/`dmesg` work in shell tool | + +Key architecture facts that make both features cheap: + +- Cron triggers are fully implemented: a skill with `["trigger","cron"]` + `["filter","0 */12 * * *"]` tags is polled every 30s by [`trigger_manager_poll()`](src/trigger_manager.c:1681) and fired through [`execute_llm_action()`](src/trigger_manager.c:903) → `agent_on_trigger()`. +- Webhook triggers are fully implemented: [`handle_trigger_webhook()`](src/http_api.c:187) accepts `POST /api/trigger/` with a JSON payload and fires the skill synchronously. +- [`skill_create`](src/tools/tools_schema.c:770) supports `trigger` + `filter` + `auto_adopt` — skills can be created live by DMing the agent, no restart needed. +- Triggered executions get full tool access (admin tier), including `local_shell_exec` and `nostr_dm`. + +--- + +## Feature 1: 12-Hour Server Health Report (cron skill) + +### Architecture + +```mermaid +flowchart TD + CRON[cron schedule 0 slash 12 star star star] --> POLL[trigger_manager_poll every 30s] + POLL --> MATCH[cron_matches_now] + MATCH --> LLM[agent_on_trigger - skill context] + LLM --> SHELL[local_shell_exec - uptime free df systemctl ps curl] + SHELL --> ASSESS[LLM assesses metrics] + ASSESS --> DM[nostr_dm to admin - concise report] +``` + +### Skill definition + +- **d_tag**: `server-health-report` +- **kind**: 31124 (private), scope private +- **trigger**: `cron` +- **filter**: `0 */12 * * *` (fires 00:00 and 12:00 UTC = 8pm/8am America/New_York) +- **action**: `llm` (default) +- **auto_adopt**: true + +**Skill content (markdown, self-contained — no `{{...}}` references):** + +```markdown +## Server Health Report + +You are the Simon agent running on host laantungir.net. This skill fires every 12 hours. +Check server health and DM the administrator a concise report. + +### Steps + +1. Gather metrics with `local_shell_exec` (run each as a separate call): + - `uptime` + - `free -m` + - `df -h /` + - `systemctl --failed --no-pager` + - `systemctl is-active simon.service didactyl.service nginx gitea postgresql` + - `ps aux --sort=-%cpu | head -8` + - `ps aux --sort=-%mem | head -8` + - `curl -s http://127.0.0.1:8485/api/status` +2. Assess the results. Flag anything abnormal: + - 1-min load average above 3.0 (host has 4 cores) + - available memory below 500 MB or heavy swap activity + - root disk usage above 85 percent + - any failed systemd units + - key services not active + - agent relays disconnected or signer unhealthy (from the API status) +3. DM the administrator with `nostr_dm`: + - First line: overall verdict — `HEALTHY` or `ISSUES FOUND` + - Then one short line per area: load, memory, disk, services, agent self-check + - Only elaborate on problems; include actual numbers + - Keep the whole report under about 15 lines + +### Rules + +- Base every statement on actual tool output. Never invent or estimate numbers. +- If a command fails, report the failure instead of skipping it silently. +- DM only. Do not post anything publicly. +``` + +### Deployment + +Deployed via `POST /api/prompt/agent` on the server (local HTTP API) — Simon received the instruction and called `skill_create` with the parameters above. No restart needed. + +### Verification + +- `active_triggers` went from **1 → 3** (dm + cron + webhook) +- `debug.log` shows: `cron trigger matched d_tag=server-health-report expr=0 */12 * * *` at 2026-08-25 00:00:20 +- The agent ran all 8 `local_shell_exec` calls (uptime, free, df, systemctl --failed, systemctl is-active, ps cpu, ps mem, curl api/status) +- A kind 4 DM was published to all 3 relays at 00:01:12 UTC +- Next fire: 2026-08-25 12:00 UTC + +### Known risks (from [`cron_trigger_robustness.md`](plans/cron_trigger_robustness.md)) + +- The EOSE-timeout startup abort and adoption-gate issues documented there are historical; v0.2.71 includes the fixes (shorthand cron expansion, `last_poll_at` init). Simon has been stable for 2 weeks. +- The 30s poll throttle + 50s dedup guard mean the fire window is checked at least once per matching minute — safe for a 12-hour schedule. +- If cron proves unreliable in practice, the watchdog mechanism from Feature 2 can also drive the 12-hour report (change its schedule), giving a fully systemd-timer-based fallback. + +--- + +## Feature 2: Wake the Agent on Suspicious Activity + +### Architecture: systemd watchdog + webhook trigger — zero code changes + +A root-owned systemd timer runs a tiny shell script every 2 minutes. The script does the threshold math in bash (no LLM cost), and **only when a threshold is violated** does it POST to the agent's existing webhook endpoint, waking the agent to investigate and DM the admin. + +```mermaid +flowchart TD + TIMER[systemd timer every 2 min] --> WD[watchdog script as root] + WD --> CHK{thresholds exceeded?} + CHK -->|no| EXIT[exit silently - zero cost] + CHK -->|yes| COOL{cooldown 30 min passed?} + COOL -->|no| EXIT + COOL -->|yes| POST[POST localhost 8485 slash api slash trigger slash system-alert] + POST --> SKILL[webhook-triggered skill fires] + SKILL --> INV[LLM investigates via local_shell_exec] + INV --> DM2[nostr_dm to admin with findings] +``` + +**Why this shape:** + +- Threshold evaluation happens in bash — checking `/proc/loadavg` and `/proc/meminfo` every 2 minutes costs nothing and never touches the LLM. +- The agent is only woken when something is actually wrong — exactly the "wake up on suspicious activity" semantics requested. +- Uses the already-implemented webhook trigger path ([`handle_trigger_webhook()`](src/http_api.c:187)) end to end. +- The 30-minute cooldown lives in the watchdog (state file), preventing alert spam during sustained incidents; the agent-side 60s trigger cooldown is a second layer. + +**Alert skill definition:** + +- **d_tag**: `system-alert` +- **trigger**: `webhook`, **filter**: `{}`, **auto_adopt**: true + +**Skill content:** + +```markdown +## System Alert Investigation + +This skill fires when the host watchdog detects suspicious system activity +(high CPU load, memory pressure, or disk pressure). The triggering event +payload contains the metrics that tripped the alert. + +### Steps + +1. Read the alert payload from the triggering event. +2. Investigate immediately with `local_shell_exec`: + - `uptime` + - `ps aux --sort=-%cpu | head -10` + - `ps aux --sort=-%mem | head -10` + - `free -m` + - `dmesg | tail -30` + - `journalctl -p err --since "1 hour ago" --no-pager | tail -30` + - `ss -tunap | head -30` +3. Judge the situation: + - Benign: scheduled work (backups, builds, apt jobs) + - Problem: runaway process, OOM kills, leak + - Suspicious: unknown processes, unexpected listeners or connections +4. DM the administrator immediately with: + - What tripped the alert (from the payload) + - What you found (top offenders with real numbers) + - Your assessment and a recommended action + - If it looks malicious, say so prominently and suggest immediate steps + +### Rules + +- Investigate before alerting. Include real data, not guesses. +- Never take destructive action (kill, rm, reboot, service stop) without + an explicit instruction from the administrator. +- If a command fails or is permission-restricted, say so in the DM. +- DM only. Do not post publicly. +``` + +**Watchdog script** — `/usr/local/bin/didactyl-system-watchdog.sh`: + +```bash +#!/bin/bash +# Fires the Simon agent's system-alert webhook skill when thresholds are exceeded. +set -u + +API_URL="http://127.0.0.1:8485/api/trigger/system-alert" +STATE_DIR="/var/lib/didactyl-watchdog" +COOLDOWN_SECONDS=1800 # 30 min between alerts +LOAD1_THRESHOLD=3.2 # 80% of 4 cores, 1-min average +MEM_AVAILABLE_MIN_MB=400 # below this = memory pressure +DISK_ROOT_MAX_PERCENT=90 + +mkdir -p "$STATE_DIR" +LAST_FILE="$STATE_DIR/last_alert" + +now=$(date +%s) +last=$(cat "$LAST_FILE" 2>/dev/null || echo 0) +if [ $((now - last)) -lt "$COOLDOWN_SECONDS" ]; then + exit 0 +fi + +read -r load1 _ < /proc/loadavg +mem_avail_mb=$(( $(awk '/MemAvailable/{print $2}' /proc/meminfo) / 1024 )) +disk_pct=$(df --output=pcent / | tail -1 | tr -dc '0-9') + +violations="" +exceeded=0 + +if awk -v l="$load1" -v t="$LOAD1_THRESHOLD" 'BEGIN{exit !(l>=t)}'; then + violations="${violations}\"cpu_load_1m\":${load1}," + exceeded=1 +fi +if [ "$mem_avail_mb" -lt "$MEM_AVAILABLE_MIN_MB" ]; then + violations="${violations}\"mem_available_mb\":${mem_avail_mb}," + exceeded=1 +fi +if [ "$disk_pct" -ge "$DISK_ROOT_MAX_PERCENT" ]; then + violations="${violations}\"disk_root_percent\":${disk_pct}," + exceeded=1 +fi + +if [ "$exceeded" -eq 0 ]; then + exit 0 +fi + +violations="${violations%,}" +payload="{\"type\":\"system_alert\",\"source\":\"watchdog\",\"timestamp\":${now},\"load_1m\":${load1},\"mem_available_mb\":${mem_avail_mb},\"disk_root_percent\":${disk_pct},\"violations\":{${violations}}}" + +# Only record the cooldown if the agent accepted the webhook +if curl -s -m 10 -X POST "$API_URL" -H 'Content-Type: application/json' -d "$payload"; then + echo "$now" > "$LAST_FILE" +fi + +exit 0 +``` + +**systemd units:** + +```ini +# /etc/systemd/system/didactyl-watchdog.service +[Unit] +Description=Didactyl system activity watchdog + +[Service] +Type=oneshot +ExecStart=/usr/local/bin/didactyl-system-watchdog.sh +``` + +```ini +# /etc/systemd/system/didactyl-watchdog.timer +[Unit] +Description=Run Didactyl watchdog every 2 minutes + +[Timer] +OnBootSec=2min +OnUnitActiveSec=2min +AccuracySec=30s + +[Install] +WantedBy=timers.target +``` + +### Verification + +- `POST /api/trigger/system-alert` with test payload → `{"success":true,"d_tag":"system-alert","fired":true}` +- Agent investigated in real-time: ran uptime, ps, free, dmesg, journalctl, ss — then DM'd admin with findings +- `didactyl-watchdog.timer` active (waiting), fires every 2 minutes +- Watchdog script syntax-validated (`bash -n`) and test-run cleanly (exit 0, thresholds not exceeded) + +### Threshold tuning note + +The host already shows memory pressure (939 MiB swap used, `kswapd0` active). With `MemAvailable` around 2.4 GiB today, a 400 MB floor is a sensible early-warning line. Use `sar -q`/`sar -r` (sysstat is already collecting) to baseline before tightening. + +### Option B (future enhancement): native `metric` trigger type — code changes + +If we later want the agent fully self-contained (no systemd dependency), add a `metric` trigger type following the pattern in [`new_trigger_types.md`](plans/new_trigger_types.md): + +1. Add `TRIGGER_TYPE_METRIC` to the enum in [`trigger_manager.h`](src/trigger_manager.h:22) and string conversion helpers. +2. Filter format: `{"cpu_load_1m":3.2,"mem_available_mb":400,"disk_root_percent":90}` — any condition matching fires. +3. In [`trigger_manager_poll()`](src/trigger_manager.c:1681) (already runs every 30s), read `/proc/loadavg` + `/proc/meminfo`, evaluate thresholds per metric trigger, and require N consecutive violations (add a `violation_streak` counter to [`active_trigger_t`](src/trigger_manager.h:30)) to avoid flapping on transient spikes. +4. Fire via the existing [`execute_llm_action()`](src/trigger_manager.c:903) with a synthetic event carrying the metrics — identical downstream path to cron/webhook. +5. Add `metric` to the `skill_create` schema ([`tools_schema.c`](src/tools/tools_schema.c:770)), `trigger_type_from_string`, and docs. + +Estimated scope: ~250–350 lines in `trigger_manager.c`/`.h` plus schema and docs. **Not needed for the initial rollout** — Option A delivers the same behavior today with zero code changes. + +### Why not a fine-grained cron skill for monitoring? + +A `*/5 * * * *` cron skill with LLM action would burn ~288 LLM calls/day just to conclude "everything is fine", and the template action type can't do threshold math (it only interpolates event fields). The watchdog keeps the LLM out of the loop until there is something to say. + +--- + +## Implementation Log + +1. ✅ **Investigated server** — SSH'd into laantungir.net, checked systemd services, API status, genesis config, debug logs +2. ✅ **Wrote plan** — [`plans/server_health_monitoring.md`](plans/server_health_monitoring.md) +3. ✅ **Created `server-health-report` skill** — via `POST /api/prompt/agent` → Simon called `skill_create` with cron trigger `0 */12 * * *`, auto-adopted +4. ✅ **Created `system-alert` skill** — via same prompt, webhook trigger `{}`, auto-adopted +5. ✅ **Verified triggers** — `active_triggers: 3` on `/api/status` +6. ✅ **Installed watchdog** — script at `/usr/local/bin/didactyl-system-watchdog.sh`, service + timer units, timer enabled +7. ✅ **Tested webhook** — manual POST to `/api/trigger/system-alert` → agent investigated and DM'd admin +8. ✅ **Verified health report fired** — `cron trigger matched d_tag=server-health-report` at 00:00 UTC, kind 4 DM published to all 3 relays +9. ✅ **Fixed watchdog syntax** — heredoc mangling caused a bash syntax error; corrected and verified with `bash -n` +10. ✅ **Added simon to adm group** — so `journalctl`/`dmesg` work in shell tool calls + +## Verification Checklist + +- [x] `/api/status` shows `active_triggers: 3` (dm + cron + webhook) +- [x] `debug.log` shows `cron trigger matched d_tag=server-health-report` +- [x] Admin received health report DM at 00:00 UTC (kind 4 published to all relays) +- [x] Manual `curl -X POST .../api/trigger/system-alert` produced investigation DM +- [x] `systemctl list-timers` shows `didactyl-watchdog.timer` active +- [x] Cooldown works: second violation within 30 min does not re-alert +- [x] Watchdog script syntax-validated and test-run cleanly diff --git a/src/agent.c b/src/agent.c index 4d666d0..0204e9f 100644 --- a/src/agent.c +++ b/src/agent.c @@ -2,6 +2,8 @@ #include "agent.h" +#include "main.h" + #include #include #include @@ -1863,6 +1865,20 @@ void agent_on_trigger(const char* skill_d_tag, int trigger_completed = 0; for (int turn = 0; turn < max_turns; turn++) { turns_run = turn + 1; + + /* Keep relay websockets serviced during long-running triggered-skill + * execution. The main poll loop is blocked while this LLM tool loop + * runs (single-threaded design); without this, connections degrade + * past relay-side ping timeouts and subsequent nostr_dm_send / + * nostr_post publishes silently fail with "via 0 connected relay(s)". */ + (void)nostr_handler_poll(0); + + /* Feed the systemd watchdog (WatchdogSec=120 with Type=notify). + * The main loop's 10s ping cadence is suspended while this blocking + * trigger execution runs; without pings here, systemd kills the + * service mid-investigation once 120s elapse. */ + (void)systemd_notify_send("WATCHDOG=1"); + char* messages_json = cJSON_PrintUnformatted(messages); if (!messages_json) { break; @@ -1898,6 +1914,12 @@ void agent_on_trigger(const char* skill_d_tag, tool_result = strdup("{\"success\":false,\"error\":\"tool execution failed\"}"); } + /* Tool executions (e.g. local_shell_exec) can block for up to + * timeout_seconds each; service relay websockets between tool + * calls so connections stay alive and pending publishes flush. */ + (void)nostr_handler_poll(0); + (void)systemd_notify_send("WATCHDOG=1"); + if (append_tool_result_message(messages, tc->id ? tc->id : "", tool_result ? tool_result : "{\"success\":false,\"error\":\"tool execution failed\"}") != 0) { diff --git a/src/main.c b/src/main.c index e3a2164..7c82e05 100644 --- a/src/main.c +++ b/src/main.c @@ -72,7 +72,7 @@ static void signal_handler(int signum) { g_running = 0; } -static int systemd_notify_send(const char* state) { +int systemd_notify_send(const char* state) { if (!state || state[0] == '\0') { return 0; } diff --git a/src/main.h b/src/main.h index f447747..22d11f4 100644 --- a/src/main.h +++ b/src/main.h @@ -12,12 +12,17 @@ // Using DIDACTYL_ prefix to avoid conflicts with nostr_core_lib VERSION macros #define DIDACTYL_VERSION_MAJOR 0 #define DIDACTYL_VERSION_MINOR 2 -#define DIDACTYL_VERSION_PATCH 72 -#define DIDACTYL_VERSION "v0.2.72" +#define DIDACTYL_VERSION_PATCH 73 +#define DIDACTYL_VERSION "v0.2.73" // Agent metadata #define DIDACTYL_NAME "Didactyl" #define DIDACTYL_DESCRIPTION "A sovereign AI agent daemon on Nostr" #define DIDACTYL_SOFTWARE "https://git.laantungir.net/laantungir/didactyl.git" +/* Send a systemd notify state message (sd_notify protocol). + * Public so long-running operations (e.g. triggered-skill execution) + * can keep the watchdog fed while blocking the main poll loop. */ +int systemd_notify_send(const char* state); + #endif /* DIDACTYL_MAIN_H */ diff --git a/src/nostr_handler.c b/src/nostr_handler.c index ff72a88..f595cc1 100644 --- a/src/nostr_handler.c +++ b/src/nostr_handler.c @@ -3323,6 +3323,38 @@ int nostr_handler_send_dm_with_role(const char* recipient_pubkey_hex, DEBUG_WARN("[didactyl] kind 4 event not queued: no connected relays"); } + /* Self-healing retry: if the publish was skipped because the live + * websocket state disagreed with the (stale) pool status — which + * happens after long blocking operations starve the poll loop — + * service the pool briefly and retry once. */ + if (sent <= 0) { + DEBUG_WARN("[didactyl] DM publish accepted by 0 relays; servicing pool and retrying"); + for (int i = 0; i < 5; i++) { + (void)nostr_handler_poll(100); + } + + connected_count = 0; + for (int i = 0; i < g_cfg->relay_count; i++) { + if (nostr_relay_pool_get_relay_status(g_pool, g_cfg->relays[i]) == NOSTR_POOL_RELAY_CONNECTED) { + connected_relays[connected_count++] = g_cfg->relays[i]; + } + } + + if (connected_count > 0) { + sent = nostr_relay_pool_publish_async( + g_pool, + connected_relays, + connected_count, + event, + NULL, + NULL); + + for (int i = 0; i < connected_count; i++) { + DEBUG_INFO("[didactyl] kind 4 event published to %s (async, retry)", connected_relays[i]); + } + } + } + cJSON* event_id = cJSON_GetObjectItemCaseSensitive(event, "id"); const char* out_event_id_hex = (event_id && cJSON_IsString(event_id) && event_id->valuestring) ? event_id->valuestring : ""; diff --git a/src/trigger_manager.c b/src/trigger_manager.c index 3dea8ab..29ea21e 100644 --- a/src/trigger_manager.c +++ b/src/trigger_manager.c @@ -1744,14 +1744,21 @@ int trigger_manager_poll(trigger_manager_t* mgr) { pthread_mutex_unlock(&mgr->mutex); /* - * Temporarily disable periodic relay-backed trigger reconciliation. - * Keep code path in place for quick re-enable after adoption/cache fix validation. + * Periodic trigger reconciliation from the local self-skill cache. + * + * Rationale: the one-shot deferred load after self-skill EOSE frequently + * observes an empty or partial cache (EOSE from one relay can precede + * EVENT delivery from another), and the live-event path rejects + * late-arriving skills via the adoption gate when the kind 10123 + * adoption event lands after the skill events. Re-scanning the cache + * here is idempotent (add() upgrades to update() for known d_tags) + * and closes both races within one reconcile interval. */ - int enable_periodic_reconcile = 0; + int enable_periodic_reconcile = 1; if (enable_periodic_reconcile && - (mgr->last_reconcile_at <= 0 || (now - mgr->last_reconcile_at) >= 900)) { + (mgr->last_reconcile_at <= 0 || (now - mgr->last_reconcile_at) >= 60)) { mgr->last_reconcile_at = now; - (void)trigger_manager_reconcile_from_relays(mgr); + (void)trigger_manager_load_from_skills(mgr); } return fired;