From b8afa40c35dea2d5b1685b6ca7b8f2b68ed76a3a Mon Sep 17 00:00:00 2001 From: Didactyl User Date: Sat, 13 Jun 2026 06:56:04 -0400 Subject: [PATCH] v0.2.49 - Add startup-failure DM de-dup with marker, relay-state/needed-event diagnostics, and clear-on-READY rearm --- README.md | 4 +- plans/remove_dead_dm_history_helpers.md | 58 +++ plans/startup_failure_admin_dm.md | 285 ++++++++++++ src/main.c | 203 ++++++++ src/main.h | 4 +- src/nostr_handler.c | 44 ++ src/nostr_handler.h | 1 + .../results/20260428T111750Z/agent_debug.log | 1 + .../runtime_test_genesis.jsonc | 59 +++ .../results/20260428T111840Z/agent_debug.log | 1 + .../runtime_test_genesis.jsonc | 59 +++ .../results/20260428T111920Z/agent_debug.log | 434 ++++++++++++++++++ .../runtime_test_genesis.jsonc | 59 +++ .../results/20260428T112825Z/agent_debug.log | 403 ++++++++++++++++ tests/results/20260428T112825Z/results.json | 67 +++ tests/results/20260428T112825Z/results.txt | 22 + .../runtime_test_genesis.jsonc | 59 +++ .../results/20260428T112851Z/agent_debug.log | 421 +++++++++++++++++ tests/results/20260428T112851Z/results.json | 57 +++ tests/results/20260428T112851Z/results.txt | 22 + .../runtime_test_genesis.jsonc | 59 +++ .../results/20260428T113005Z/agent_debug.log | 420 +++++++++++++++++ tests/results/20260428T113005Z/results.json | 57 +++ tests/results/20260428T113005Z/results.txt | 22 + .../runtime_test_genesis.jsonc | 59 +++ .../results/20260428T113354Z/agent_debug.log | 420 +++++++++++++++++ tests/results/20260428T113354Z/results.json | 100 ++++ tests/results/20260428T113354Z/results.txt | 31 ++ .../runtime_test_genesis.jsonc | 59 +++ .../results/runtime_test_genesis_fixed.jsonc | 59 +++ 30 files changed, 3545 insertions(+), 4 deletions(-) create mode 100644 plans/remove_dead_dm_history_helpers.md create mode 100644 plans/startup_failure_admin_dm.md create mode 100644 tests/results/20260428T111750Z/agent_debug.log create mode 100644 tests/results/20260428T111750Z/runtime_test_genesis.jsonc create mode 100644 tests/results/20260428T111840Z/agent_debug.log create mode 100644 tests/results/20260428T111840Z/runtime_test_genesis.jsonc create mode 100644 tests/results/20260428T111920Z/agent_debug.log create mode 100644 tests/results/20260428T111920Z/runtime_test_genesis.jsonc create mode 100644 tests/results/20260428T112825Z/agent_debug.log create mode 100644 tests/results/20260428T112825Z/results.json create mode 100644 tests/results/20260428T112825Z/results.txt create mode 100644 tests/results/20260428T112825Z/runtime_test_genesis.jsonc create mode 100644 tests/results/20260428T112851Z/agent_debug.log create mode 100644 tests/results/20260428T112851Z/results.json create mode 100644 tests/results/20260428T112851Z/results.txt create mode 100644 tests/results/20260428T112851Z/runtime_test_genesis.jsonc create mode 100644 tests/results/20260428T113005Z/agent_debug.log create mode 100644 tests/results/20260428T113005Z/results.json create mode 100644 tests/results/20260428T113005Z/results.txt create mode 100644 tests/results/20260428T113005Z/runtime_test_genesis.jsonc create mode 100644 tests/results/20260428T113354Z/agent_debug.log create mode 100644 tests/results/20260428T113354Z/results.json create mode 100644 tests/results/20260428T113354Z/results.txt create mode 100644 tests/results/20260428T113354Z/runtime_test_genesis.jsonc create mode 100644 tests/results/runtime_test_genesis_fixed.jsonc diff --git a/README.md b/README.md index 3d2c0f0..a04baca 100644 --- a/README.md +++ b/README.md @@ -54,11 +54,11 @@ Skills compose by adoption-list order (`10123`) and trigger tags carry runtime e Didactyl will support local inference, which is very privacy preserving. Remote inference does however have it's advantages, and in those cases Didactyl supports using Bitcoin Lightning and eCash inference providers. -## Current Status — v0.2.48 +## Current Status — v0.2.49 **Active build — this project is barely working. Experiment at your own risk.** -> Last release update: v0.2.48 — Remove admin kind 10123 from subscription and upsert, add debug logging +> Last release update: v0.2.49 — Add startup-failure DM de-dup with marker, relay-state/needed-event diagnostics, and clear-on-READY rearm - Connects to configured relays with auto-reconnect and relay state transition logging - Publishes configured startup events per relay as each relay becomes connected diff --git a/plans/remove_dead_dm_history_helpers.md b/plans/remove_dead_dm_history_helpers.md new file mode 100644 index 0000000..35e5865 --- /dev/null +++ b/plans/remove_dead_dm_history_helpers.md @@ -0,0 +1,58 @@ +# Remove Dead DM-History Helper Functions in `agent.c` + +## Background + +While investigating how Didactyl maintains continuity across a DM "chat session", two functions in [`src/agent.c`](../src/agent.c) were found to be compiled-but-unreferenced, marked with `__attribute__((unused))` to suppress compiler warnings: + +- [`append_recent_admin_dm_history`](../src/agent.c:1065) +- [`build_recent_admin_dm_history_messages`](../src/agent.c:1800) (wraps the above) + +These were part of an older design where the agent automatically injected the last N admin DM turns into every LLM call. The current architecture is **skill-driven**: DM history only reaches the LLM when an adopted, triggered skill explicitly pulls it in via the `{{dm_history}}` / `{{nostr_dm_history(...)}}` template variables or the `nostr_dm_history` tool (see [`tool_agent.c`](../src/tools/tool_agent.c:284) and [`tools_schema.c`](../src/tools/tools_schema.c:1606)). The dead functions are now a source of confusion for anyone reading the code to understand chat-session behavior. + +## Goal + +Remove the dead functions and any symbols used exclusively by them, without changing runtime behavior. + +## Scope Analysis (verified before planning) + +A repo-wide search for each symbol confirmed the following: + +| Symbol | Only callers | Disposition | +|---|---|---| +| [`append_recent_admin_dm_history`](../src/agent.c:1065) | called by `build_recent_admin_dm_history_messages` (also dead) | **delete** | +| [`build_recent_admin_dm_history_messages`](../src/agent.c:1800) | no callers | **delete** | +| [`AGENT_HISTORY_TURNS`](../src/agent.c:38) | referenced only inside the dead `append_recent_admin_dm_history` | **delete** | +| [`AGENT_HISTORY_QUERY_LIMIT`](../src/agent.c:39) | defined but zero references repo-wide | **delete** | +| [`append_simple_message`](../src/agent.c:187) | many live callers throughout `agent.c` | **keep** | +| [`nostr_handler_get_dm_history_json`](../src/nostr_handler.c:4384) | live via [`tool_agent.c`](../src/tools/tool_agent.c:318) | **keep** | + +No non-source references (Makefile, tests, docs) depend on these symbols. + +## Risks + +Very low. Both functions carry `__attribute__((unused))` so they are not linked into any runtime path. The only failure modes are accidental deletion of something still live, which the scope analysis above rules out. + +## Verification Strategy + +1. Clean build with `make` — must succeed with no new warnings. +2. Existing tests under [`tests/`](../tests/) pass (`python tests/run_tests.py` or the project's usual invocation). +3. Manual sanity check: start an agent, send a DM, confirm normal reply — no regression, since the deleted code was never on the hot path. + +## Mermaid: before/after + +```mermaid +flowchart LR + subgraph Before + A1[agent.c dead funcs] -. unused attr .- B1[compiler silences warning] + A1 --> C1[reader assumes auto DM history replay] + C1 --> D1[confusion about session behavior] + end + subgraph After + A2[agent.c clean] --> E2[skill-driven history only] + E2 --> F2[intent matches code] + end +``` + +## Follow-on (out of scope for this change) + +The documentation in [`docs/CONTEXT.md`](../docs/CONTEXT.md) already describes the skill-driven assembly model correctly, but neither that file nor [`docs/SKILLS.md`](../docs/SKILLS.md) explicitly states "DM history is not automatic — it only appears when a skill references it." A brief clarifying paragraph could be added in a separate change. diff --git a/plans/startup_failure_admin_dm.md b/plans/startup_failure_admin_dm.md new file mode 100644 index 0000000..34d708e --- /dev/null +++ b/plans/startup_failure_admin_dm.md @@ -0,0 +1,285 @@ +# Startup Failure Admin DM Notification + +> **STATUS UPDATE (follow-up phase):** The base feature (send a failure DM) is +> implemented and verified on VM410. However, because the failure happens inside the +> systemd `Restart=on-failure` loop (`RestartSec=10`), the admin now receives a *new* +> failure DM roughly every ~25 seconds, indefinitely (observed in the live message box). +> This follow-up adds **throttling/de-duplication** so the failure DM is sent **once per +> failure type until the next successful startup**, and makes the message **more +> descriptive** (per-relay connect state + the event kinds that were needed). +> See the "Follow-up: Throttling + Descriptive Failure DM" section at the end. + +# Startup Failure Admin DM Notification (original plan) + +## Goal + +When the agent fails to start, send a best-effort encrypted DM to the administrator +describing which startup step failed and why, so the operator learns about the failure +without having to SSH in and read `journalctl`. + +This directly addresses the real-world incident on VM410 where both `didactyl.service` +and `simon.service` were stuck restart-looping at startup step 14 +(`Subscribe self-skill cache: failed (no skill events found after EOSE)`) and silently +stopped answering DMs. + +## Background / Code References + +- Startup is a linear checklist in [`main()`](src/main.c:1012). Each step uses helpers: + - [`startup_step_begin()`](src/main.c:303) + - [`startup_step_ok()`](src/main.c:309) + - [`startup_step_fail()`](src/main.c:320) +- On any fatal failure, the current pattern is: + `startup_step_fail(...)` -> `fprintf(stderr, ...)` -> cleanup calls -> `return 1;` + (e.g. the step-14 self-skill abort at [`src/main.c:1520`](src/main.c:1520)). +- The success path already sends an admin "started up" DM at + [`src/main.c:1663`](src/main.c:1663) via + [`nostr_handler_send_dm_auto()`](src/nostr_handler.c:3264). +- DM delivery requires (see [`nostr_handler_send_dm_with_role()`](src/nostr_handler.c:3133)): + - `g_cfg` and `g_pool` set (relay pool initialized), + - at least one **connected** relay (otherwise the kind-4 event is dropped: + "kind 4 event not queued: no connected relays" at [`src/nostr_handler.c:3203`](src/nostr_handler.c:3203)), + - a valid `cfg.admin.pubkey` (validated at step 4, [`src/main.c:1310`](src/main.c:1310)), + - the private key derived (always done before the checklist). + +### Delivery-window reality + +A failure DM is only *deliverable* from roughly step 11 onward (relay pool created + +relays connected + admin known). The step-14 failure we actually hit is inside that +window, so it will be delivered. Earlier failures (bad nsec, missing admin, LLM config, +no relay connection) have no working channel; per decision, those are **logged only** +(best-effort, no blocking retries). + +## Design Decision (confirmed) + +**Best-effort only.** Attempt the failure DM whenever relays + admin are available; if +the relay pool isn't connected or admin pubkey is unknown, skip the DM and rely on +existing stderr/journal logging. No blocking retry loop, no self-healing of the +underlying failure (that can be a separate plan). + +## Approach + +Add a single helper that builds and sends the failure DM, then call it from every fatal +startup-abort site (the points that currently do `startup_step_fail(...) ... return 1;`). + +### 1. New helper in [`src/main.c`](src/main.c) + +Add a static function near the other startup helpers (after +[`startup_step_fail()`](src/main.c:320)): + +```c +static void notify_admin_startup_failure(const didactyl_config_t* cfg, + int step, + const char* label, + const char* detail); +``` + +Behavior: +- Guard clauses (skip + DEBUG_WARN if any precondition is missing): + - `cfg == NULL` or `cfg->admin.pubkey[0] == '\0'` -> log "admin unknown; cannot notify". + - `nostr_handler_connected_relay_count() <= 0` -> log "no connected relay; cannot notify". +- Build a concise human-readable message, e.g.: + ``` + FAILED to start at . + Startup step [14] Subscribe self-skill cache failed: no skill events found after EOSE. + Version . Connected relays: /. The agent is NOT online and will not answer DMs until fixed. + ``` + - Reuse `nostr_handler_get_startup_display_name()` and `DIDACTYL_VERSION` exactly like + the success DM at [`src/main.c:1645`](src/main.c:1645) for naming/version consistency. + - Include `step`, `label`, and `detail` so each failure site is self-describing. +- Send via `nostr_handler_send_dm_auto(cfg->admin.pubkey, msg)`; if it returns non-zero, + `DEBUG_WARN` that the failure DM could not be delivered (do not change the exit code). +- Pump the poll loop briefly after sending so the async publish can flush before process + exit: call `nostr_handler_poll(...)` a few times (e.g. ~300--500 ms total) so the + queued kind-4 event is actually written to the relay sockets before cleanup/`return 1`. + (Without this, async publish may be discarded on immediate exit.) + +### 2. Wire the helper into fatal abort sites + +For each fatal `startup_step_fail(...)` that ends in `return 1;`, insert a call to +`notify_admin_startup_failure(&cfg, , "