20 KiB
Swarm Proof of Concept
Goal: prove the thread model works end to end with a queen and a few workers, all on a single local relay.
Status: Working end to end
The PoC is running. Three agents — a queen and two workers — are live on ws://127.0.0.1:7777.
T1 passes. The swarm decomposed a problem, divided the work, computed all five subtasks correctly, and synthesized the exact right answer.
What works:
- The queen wakes on the phrase "my queen" and seeds a swarm thread
- The queen decomposes a problem into sub-ranges and posts a
claimwith the plan - Workers see the thread, claim distinct sub-ranges, and post
resultnotes - Workers combine results into a
synthesiswith the final answer swarm_create,swarm_read, andswarm_contributeall function
The ingestion fix
The first T1 run stalled after two subtasks. The cause was
maybe_fire_trigger_locked(), which deduped by
created_at timestamp — distinct events sharing a second were dropped, and
cooldown discarded work rather than deferring it.
The fix: dedupe by event id using a 64-entry ring buffer per trigger
(seen_event_ids in active_trigger_t). Timestamp
dedup is gone; the ring buffer catches genuine replays without dropping distinct
events. After the fix, the swarm ran to completion.
Observed T1 run (after the fix)
The admin posted: "My queen, find the sum of all prime numbers between 1 and 1000. Decompose the range across workers and have them report subtotals, then synthesize the total."
The resulting thread:
| Author | Type | Content |
|---|---|---|
| queen | root | "Find the sum of all prime numbers between 1 and 1000. The range [1, 1000] is decomposed into 5 independent sub-ranges..." |
| queen | claim | "Proposed decomposition: split [1, 1000] into 5 equal sub-ranges of 200 integers each..." |
| worker1 | claim | "Claiming Worker C: [401, 600]" |
| worker1 | result | "Worker C result: [401, 600] — Primes (31 total) ... Subtotal: 15,409" |
| worker2 | claim | "Claiming Worker A: [2, 200] and Worker D: [601, 800]" |
| worker2 | result | "Worker A result: [2, 200] — Primes (46 total) ... Subtotal: 4,227" |
| worker2 | result | "Worker D result: [601, 800] — Primes (30 total) ... Subtotal: 20,782" |
| worker1 | result | "Worker E result: [801, 1000] — Primes (29 total) ... Subtotal: 26,049" |
| worker1 | result | "Worker B result: [201, 400] — Primes (32 total) ... Subtotal: 9,660" |
| worker1 | synthesis | "Synthesis: Sum of all primes between 1 and 1000 ... Total 168 primes, 76,127" |
| worker2 | synthesis | "Synthesis: Sum of all primes between 1 and 1000 ... Grand total: 76,127" |
All five subtotals are exactly correct, and both syntheses report the correct total: 76,127.
Two findings from live client testing
1. The queen must always seed a swarm. The first skill said "If it is a
simple question, you may answer directly instead." When the admin asked for the
sum of primes 1–2000, the queen computed it herself with local_shell_exec
instead of seeding a swarm. The skill now says: "When addressed, you ALWAYS
seed a swarm. Do not answer the question yourself... Never call
local_shell_exec." After the update, she decomposed the range into four
sub-ranges and the swarm solved it correctly (277,050).
2. Triggered skills have no output channel. agent_on_trigger()
runs the LLM loop but only sends a DM on errors. A triggered skill's final
text is discarded — there is no way for a triggered skill to reply to the admin.
This is why the queen's direct answer never reached the client. For the swarm
this is fine (the answer lives in the thread), but it is a real gap for any
triggered skill that wants to report back. A future notify_admin tool or a
trigger-level "reply to source" option would close it.
Observations
- The queen's decomposition was good. Five equal sub-ranges, each a clean independent unit of work.
- Workers self-organized. They claimed distinct ranges without a coordinator. One worker took two ranges, the other took three.
- Both workers synthesized. The plan allows competing syntheses; both arrived at the same correct answer. This is redundancy, not conflict.
- No duplicate claims. Each range was claimed once.
- The swarm is chatty. Every contribution triggers every worker, so the thread grows quickly. Cost control (batching, budgets) is a real future need.
- Workers may over-claim. In the 1–2000 run, both workers claimed all four subtasks at once ("all four sub-ranges are unclaimed, so I'll compute them all"). They raced, but both produced correct results. A claim is advisory, so this is expected — but it doubles the work. A future improvement is to have workers re-read the thread immediately before claiming.
- The queen's skill is editable on the relay. Because genesis is consumed once, the skill was updated by publishing a new kind 31124 event as the queen. The running agent picked it up live via its self-skill subscription — no restart needed. This is the intended way to evolve an agent.
Artifacts
| Path | Purpose |
|---|---|
swarm/queen/ |
Queen config, keys, run script, log |
swarm/worker1/ |
Worker 1 config, keys, run script, log |
swarm/worker2/ |
Worker 2 config, keys, run script, log |
swarm/admin_post.sh |
Publish a kind 1 note as the admin via n_signer |
src/tools/tool_swarm.c |
The three swarm tools |
The admin key is the n_signer on the nostr_signer qube (role main, path
m/44'/1237'/0'/0/0), reached via qrexec nostr_signer:qubes.SignerRpc. The
admin key never leaves the signer qube.
Scope
This is a proof of concept, not a product. It answers one question: can a queen seed a swarm, can workers contribute to a shared thread, and can the swarm produce a result?
Everything runs locally. No public relays, no spawning, no discovery, no human recruitment. Those come later.
Constraints
Local relay only. Every agent connects to exactly one relay: ws://127.0.0.1:7777. This is already the first entry in DEFAULT_RELAYS, so the PoC configs simply list it alone.
Why this matters:
- Deterministic — no public relay rate limits, no partial views, no network flakiness
- Observable — we can query the relay directly to see every event the swarm produced
- Isolated — the swarm cannot accidentally post to the real network
- Fast — local round trips are milliseconds
Each agent's kind 10002 relay list contains only ws://127.0.0.1:7777. The startup config must not include any other relay, or the agent will sync to it.
Addressing the Queen
The queen watches the admin's timeline. She needs to know when she is being addressed.
The mechanism: the phrase "my queen"
Keep it simple. The queen's trigger fires on the admin's kind 1 posts, and she acts only when the post body contains the phrase "my queen" (case-insensitive).
My queen, what is 2 + 2?
my queen: find the sum of primes under 1000
MY QUEEN what time is it
No p tags. No npub. No client configuration. The admin just types the phrase.
How it works
The trigger filter is just the admin's posts:
{"kinds": [1], "authors": ["<ADMIN_PUBKEY>"]}
The queen skill instructs the LLM to check the body for "my queen" (case-insensitive) and to do nothing if the phrase is absent. The LLM handles the case-insensitivity naturally — no C code needed.
This means the LLM runs on every admin post, but the queen's skill makes the no-op path cheap: if "my queen" is not present, she stops immediately without calling any tools.
Why this is fine for the PoC
| Approach | Setup cost | Precision | Notes |
|---|---|---|---|
p tag mention |
Must know the queen's npub; client must add the tag | High | More correct, more friction |
| "my queen" phrase | None — just type it | Good enough | Zero friction, trivial to test |
For a proof of concept on a local relay, the phrase is the right call. It removes an entire class of setup problems (generating keys, computing pubkeys, baking them into filters, getting clients to emit tags) and lets us focus on the swarm itself.
A p-tag mechanism can be added later as a more precise alternative, but it is not needed to prove the model.
Flow
sequenceDiagram
participant Admin
participant Relay as Local Relay
participant Queen
participant Worker
Admin->>Relay: kind 1: My queen, what is 2+2?
Relay->>Queen: trigger fires (authors=admin)
Note over Queen: body contains "my queen" -> proceed
Queen->>Relay: swarm_create: root with t=didactyl-swarm,<br/>swarm=root, e=admin_post
Relay->>Worker: trigger fires (#t=didactyl-swarm)
Worker->>Relay: swarm_contribute: result
Relay->>Queen: sees the result
Queen->>Relay: swarm_contribute: synthesis
The queen's swarm_create anchors the thread to the admin's original post with an e tag, so the swarm is traceable back to the request.
Test Problems
A good swarm problem must be decomposable, verifiable, and self-contained on the local relay. Here is a progression, from smoke test to real collaboration.
T0 — Smoke Test: "The queen wakes up"
Admin posts:
My queen, what is 2 + 2? Have a worker answer.
The queen sees the phrase, creates a thread root anchored to the admin's post, and a worker replies with "4".
Verifies: the "my queen" addressing works, the queen's trigger fires, the thread root is published, a worker sees it and contributes.
Why start here: it isolates the plumbing — addressing, trigger, root, contribution. If this fails, nothing else will work.
Note: the "Have a worker answer" phrasing forces the swarm path. Without it, the queen might reasonably answer a trivial question directly, which would test addressing but not the swarm.
T1 — Decomposable Computation: "Sum of primes"
Admin posts: "Find the sum of all prime numbers between 1 and 1000."
The queen decomposes the range into chunks (1–250, 251–500, 501–750, 751–1000), posts a thread root with the plan, and each worker claims a chunk, computes its subtotal, and posts a result. The queen (or any worker) posts a synthesis with the total.
Verifies: decomposition, claims, parallel work, synthesis.
Ground truth: 76127. We can check the final answer exactly.
Why this is good: the work is genuinely parallel, the answer is verifiable, and the LLM's arithmetic is checkable. If a worker gets a subtotal wrong, the synthesis exposes it.
T2 — Relay Scavenger Hunt: "Assemble the fragments"
Before the test, seed the local relay with 20 kind 1 notes, each containing a fragment (a word or number). Admin posts: "Find all 20 fragments tagged
#poc-fragmentand combine them into the final message."
The queen creates a thread, workers query the relay for fragments, each claims a subset, posts what it found, and the swarm assembles the complete message.
Verifies: relay discovery, division of a search space, assembly of partial results.
Ground truth: we know the fragments and the expected message.
Why this is good: it exercises the discovery path — workers must query the relay, not just read the thread. It also tests that the swarm can handle a task where the input is scattered.
T3 — Research Synthesis: "Cross-reference the dataset"
Seed the relay with 30 notes, each a city with a population. Admin posts: "What is the total population of the five largest cities in the dataset?"
Workers query and filter the dataset, each handling a subset, and the swarm synthesizes the answer.
Verifies: filtering, ranking, aggregation, and a synthesis that depends on all workers' results.
Ground truth: computed from the seeded data.
Why this is good: it is the closest to a real research task, and it requires the synthesis step to actually combine partial results rather than just concatenate them.
Recommendation
Build T0 first, then T1. T0 proves the plumbing; T1 proves real collaboration with a checkable answer. T2 and T3 are stretch goals that exercise discovery and aggregation.
T1 is the sweet spot: simple enough to debug, real enough to be convincing, and the answer is a single number we can verify.
Queen Implementation (Minimal)
The queen is a normal Didactyl agent with a queen skill. No new C code is needed for the queen itself — it is a skill plus the swarm tools.
Queen skill
A kind 31124 private skill with a nostr-subscription trigger on the admin's kind 1 notes:
{
"kind": 31124,
"content": "## Queen\n\nYou watch your administrator's public notes. When the admin addresses you, you seed a swarm.\n\n### Addressing\n\nYou only respond when the admin's post contains the phrase \"my queen\" (case-insensitive). If the phrase is not present, do nothing — stop immediately and call no tools.\n\n### When to seed a swarm\n\nWhen addressed, seed a swarm if the problem is decomposable into independent subtasks. If it is a simple question, you may answer directly instead.\n\n### How to seed\n\n1. Read the admin's post carefully.\n2. Decide whether it is worth a swarm.\n3. If yes, call `swarm_create` with a clear title and the problem statement.\n4. Post a `claim` on your own thread describing the decomposition you propose.\n5. Stop. Workers will pick up the subtasks.\n\n### Rules\n\n- You have no authority. You propose; workers decide.\n- Keep the thread root focused on the problem, not on you.\n- If a worker posts a result, you may post a `synthesis` when all subtasks are done.",
"tags": [
["d", "queen"],
["app", "didactyl"],
["scope", "private"],
["description", "Seed swarms from admin posts"],
["trigger", "nostr-subscription"],
["filter", "{\"kinds\":[1],\"authors\":[\"<ADMIN_PUBKEY>\"]}"],
["tools", "swarm_create,swarm_read,swarm_contribute,nostr_post,nostr_query"]
]
}
The filter is just the admin's posts. The queen skill decides whether to act by checking the body for "my queen" (case-insensitive). No pubkey or tag setup is required.
Worker skill
A kind 31124 private skill with a nostr-subscription trigger on the swarm hashtag:
{
"kind": 31124,
"content": "## Worker\n\nYou participate in swarms. A swarm is a shared thread — a kind 1 conversation where agents work a common problem.\n\n### Protocol\n\n1. When a swarm thread appears, read it with `swarm_read`.\n2. Look for `claim` notes to see what others are doing.\n3. If you can help, post a `claim` for a subtask nobody has taken.\n4. Do the work, then post a `result` with your findings.\n5. If you are blocked, post a `question`.\n6. If you can combine the results into a final answer, post a `synthesis`.\n\n### Rules\n\n- Do not duplicate work someone else has claimed.\n- Keep contributions concise and concrete.\n- Post the actual answer, not a description of how you would find it.",
"tags": [
["d", "worker"],
["app", "didactyl"],
["scope", "private"],
["description", "Participate in swarms"],
["trigger", "nostr-subscription"],
["filter", "{\"kinds\":[1],\"#t\":[\"didactyl-swarm\"]}"],
["tools", "swarm_read,swarm_contribute,nostr_post,nostr_query"]
]
}
Swarm tools needed
Only three tools are required for the PoC:
| Tool | Purpose |
|---|---|
swarm_create |
Publish a thread root with ["t","didactyl-swarm"] and ["swarm","root"] |
swarm_read |
Fetch the root and all e-linked replies, return them as structured JSON |
swarm_contribute |
Publish a kind 1 reply with root/reply e tags and an optional swarm-type |
swarm_read and swarm_contribute are the workhorses. swarm_create is only used by the queen.
Worker Setup
For the PoC, workers are launched manually — no spawning. Each is a normal Didactyl instance with:
- Its own nsec (distinct identity)
- The same admin pubkey
- A relay list containing only
ws://127.0.0.1:7777 - The worker skill adopted
- Its own API port (so we can inspect each one)
Three instances is enough: one queen, two workers.
flowchart TD
RELAY[ws://127.0.0.1:7777]
ADMIN[Admin<br/>posts problem]
QUEEN[Queen<br/>api :8484]
W1[Worker 1<br/>api :8485]
W2[Worker 2<br/>api :8486]
ADMIN -->|kind 1| RELAY
RELAY -->|trigger| QUEEN
QUEEN -->|swarm_create| RELAY
RELAY -->|trigger| W1
RELAY -->|trigger| W2
W1 -->|contribute| RELAY
W2 -->|contribute| RELAY
RELAY -->|read| QUEEN
Test Harness
The existing AgentProcess harness already supports launching an agent with a config, port, and log file. The PoC extends it to launch three instances and drive the scenario.
Harness shape
tests/swarm/
configs/
queen.genesis.jsonc
worker1.genesis.jsonc
worker2.genesis.jsonc
seed_fragments.py # for T2/T3: publish seed events to the local relay
run_swarm_poc.py # orchestrates the scenario
Scenario steps
- Start the local relay at
ws://127.0.0.1:7777(assumed running). - Launch the queen and two workers with their configs.
- Wait for all three to reach the main poll loop.
- (T2/T3 only) Seed the relay with the dataset.
- Publish the admin's problem as a kind 1 note from the admin key.
- Poll the relay for the thread root and contributions.
- Wait for a
synthesisor a timeout. - Assert the final answer matches the ground truth.
- Dump the full thread for inspection.
Observing the swarm
Because everything is on one local relay, we can inspect the entire swarm with a single query:
{"kinds": [1], "#t": ["didactyl-swarm"]}
This returns every root and contribution. The harness can render the thread as a tree and check that:
- A root exists
- At least two distinct authors contributed
- Claims precede results
- A synthesis exists
- The synthesis contains the correct answer
Success Criteria
| Criterion | How we check |
|---|---|
| Queen reacts to the admin post | A thread root exists authored by the queen |
| Workers see the thread | At least two distinct worker pubkeys appear in the thread |
| Work is divided | At least two claim notes with different subtasks |
| Results are posted | At least two result notes |
| The swarm converges | A synthesis note exists |
| The answer is correct | The synthesis contains the ground-truth answer |
| No duplicate work | No two workers claim the same subtask |
What This PoC Deliberately Excludes
- Spawning — workers are launched manually
- Discovery — workers are pre-configured, not found via kind 0
- Human recruitment — no DMs to humans
- Gating — the thread is open; no policy enforcement
- Side threads — a single flat thread
- Reliable ingestion — the known timestamp-dedup issue is not fixed yet; the PoC uses spaced-out posts to avoid it
These are all in plans/swarm.md. The PoC proves the core loop before we build the rest.
Implementation Steps
- Write the three genesis configs — queen, worker1, worker2, each with only
ws://127.0.0.1:7777in the relay list and the appropriate skill. - Implement
swarm_create— publish a kind 1 root with the swarm tags. - Implement
swarm_read— query the root and itse-linked replies, return structured JSON. - Implement
swarm_contribute— publish a kind 1 reply with root/replyetags and an optionalswarm-type. - Register the three tools in
tools_internal.h,tools_dispatch.c, andtools_schema.c. - Write the queen and worker skills into the configs.
- Run T0 — verify the queen wakes up and a worker replies.
- Run T1 — verify the swarm computes the sum of primes correctly.
- Write the harness — automate the scenario and the assertions.
- Run T2/T3 — stretch goals for discovery and aggregation.
Open Questions
- How does the queen know a post is a problem? The phrase "my queen" (case-insensitive) in the body is the signal. The queen skill can still decline to seed a swarm for trivial questions.
- How do workers avoid duplicate claims? For the PoC,
swarm_readbefore claiming is enough. Race conditions are acceptable at this scale. - Who synthesizes? For the PoC, allow any worker to synthesize. The queen can also do it. We do not need to reserve it.
- How long do we wait? The harness should wait for a
synthesiswith a generous timeout, then report what it saw.