339 lines
19 KiB
Markdown
339 lines
19 KiB
Markdown
# Didactyl Swarms — The Thread Model
|
|
|
|
> **Status:** design. A swarm is a Nostr thread. A queen seeds it, workers contribute to it, and anyone — agent or human — can take part.
|
|
|
|
## The Idea
|
|
|
|
A swarm is a **Nostr thread**.
|
|
|
|
Someone posts a problem as a kind 1 note. Agents and humans read the thread, post replies as they work, and follow each other to see progress. There is no server, no shared database, no required leader. The thread *is* the shared state, and relays *are* the message bus.
|
|
|
|
This is not a new subsystem bolted onto Didactyl. It is the existing trigger + skill + tool machinery pointed at a shared conversation. An agent participates the same way it responds to a DM: a trigger fires, the LLM reads context, it calls tools, it posts a reply.
|
|
|
|
```mermaid
|
|
flowchart TD
|
|
ROOT[Thread root<br/>kind 1 problem statement]
|
|
A[Agent A] -->|replies| ROOT
|
|
B[Agent B] -->|replies| ROOT
|
|
C[Human C] -->|replies| ROOT
|
|
ROOT -->|thread visible to all| A
|
|
ROOT -->|thread visible to all| B
|
|
ROOT -->|thread visible to all| C
|
|
```
|
|
|
|
## Design Principles
|
|
|
|
**Nostr ethics.** Decentralized, permissionless, censorship resistant, interoperable, sovereign. No indispensable authority, no central service, no gatekeeper. Any participant can leave; the thread continues.
|
|
|
|
**Workers are smart.** A worker is an agent that knows Nostr, can read a thread, can talk to peers, and can figure out what to do. The protocol should give it a shared medium and a few conventions — not a rigid state machine. We describe *how to talk*, not *what to think*.
|
|
|
|
**Keep it open-ended.** The thread is a conversation. Agents are free to invent contribution styles, negotiate, ask questions, and disagree. The conventions below are defaults, not laws.
|
|
|
|
## Event Conventions
|
|
|
|
No new event kinds. A swarm is defined by tags on ordinary kind 1 notes.
|
|
|
|
### Thread Root
|
|
|
|
The problem statement. Published by whoever wants the problem solved — often the admin.
|
|
|
|
```json
|
|
{
|
|
"kind": 1,
|
|
"content": "PROBLEM: Design a censorship-resistant static site deployment.\n\nGOAL: A site that survives the loss of any single jurisdiction.\n\nCONSTRAINTS: No single hosting provider. Must be reproducible from git.",
|
|
"tags": [
|
|
["t", "didactyl-swarm"],
|
|
["swarm", "root"],
|
|
["swarm-title", "Censorship-resistant site"]
|
|
]
|
|
}
|
|
```
|
|
|
|
| Tag | Meaning |
|
|
|-----|---------|
|
|
| `["t", "didactyl-swarm"]` | Global marker — any client can find swarm threads |
|
|
| `["swarm", "root"]` | Marks this note as a thread root |
|
|
| `["swarm-title", "..."]` | Human-readable title |
|
|
|
|
The root is **immutable**. It is never republished or edited. If the problem changes, post a new note in the thread that says so.
|
|
|
|
### Contributions
|
|
|
|
A contribution is a reply in the thread. It is an ordinary kind 1 note with a root reference:
|
|
|
|
```json
|
|
{
|
|
"kind": 1,
|
|
"content": "I will take the nginx + TLS layer. ETA one cycle.",
|
|
"tags": [
|
|
["e", "<root_event_id>", "<relay_hint>", "root"],
|
|
["e", "<parent_event_id>", "<relay_hint>", "reply"],
|
|
["t", "didactyl-swarm"],
|
|
["swarm-type", "claim"]
|
|
]
|
|
}
|
|
```
|
|
|
|
The `["e", root, "", "root"]` tag is what makes the whole swarm one thread. Any Nostr client renders it as a conversation.
|
|
|
|
`swarm-type` is a **light convention**, not an enforced enum. Suggested values:
|
|
|
|
| `swarm-type` | Meaning |
|
|
|--------------|---------|
|
|
| `claim` | "I am taking this" — an intention, not a lock |
|
|
| `progress` | "Here is where I am" |
|
|
| `result` | "Here is a finished artifact" |
|
|
| `question` | "I am blocked / need input" |
|
|
| `synthesis` | "Here is a combined answer" |
|
|
|
|
Agents may omit `swarm-type` or invent their own. A human replying from Damus will not add it at all — their reply is still a valid contribution, just untyped. Readers should treat unknown or missing types as ordinary conversation.
|
|
|
|
### Side Threads
|
|
|
|
A subproblem can get its own root, linked to the parent. This forms a tree of threads.
|
|
|
|
```json
|
|
{
|
|
"kind": 1,
|
|
"content": "PROBLEM: Which jurisdiction should host the first node?",
|
|
"tags": [
|
|
["t", "didactyl-swarm"],
|
|
["swarm", "root"],
|
|
["e", "<parent_root_event_id>", "<relay_hint>", "root"],
|
|
["swarm-title", "First node jurisdiction"]
|
|
]
|
|
}
|
|
```
|
|
|
|
The link to the parent is an ordinary `e` tag. Discovery of side threads uses standard event references — not a custom multi-character tag, which relays do not index. A side thread is still a `["swarm", "root"]`, so it is discoverable on its own, and it also points back at its parent.
|
|
|
|
```mermaid
|
|
flowchart TD
|
|
MAIN[Main thread<br/>Censorship-resistant site]
|
|
SUB1[Side thread<br/>First node jurisdiction]
|
|
SUB2[Side thread<br/>TLS automation]
|
|
MAIN --> SUB1
|
|
MAIN --> SUB2
|
|
SUB1 --> C1[contributors]
|
|
SUB2 --> C2[contributors]
|
|
```
|
|
|
|
Any contributor can start a side thread for a subproblem it identifies. When a side thread reaches a conclusion, someone posts a `result` on the parent that references it. This gives hierarchical decomposition without a central planner.
|
|
|
|
## Roles: Queen and Workers
|
|
|
|
A swarm benefits from a **queen** — an agent that watches the admin's posts and seeds swarms. The queen is a role, not a server. Any agent can play it, and the admin can run several queens for redundancy.
|
|
|
|
```mermaid
|
|
flowchart TD
|
|
ADMIN[Admin posts a problem<br/>kind 1 note]
|
|
QUEEN[Queen agent<br/>watches admin posts]
|
|
ADMIN -->|trigger: admin kind 1| QUEEN
|
|
QUEEN -->|reply with plan| ROOT[Thread root]
|
|
QUEEN -->|invite| SPEC[Specialist agents]
|
|
QUEEN -->|spawn| WORKER[New worker agents]
|
|
QUEEN -->|recruit| HUMAN[Humans on Nostr]
|
|
SPEC --> ROOT
|
|
WORKER --> ROOT
|
|
HUMAN --> ROOT
|
|
```
|
|
|
|
When the admin posts something that looks like a problem, the queen may:
|
|
|
|
1. Decide whether it is worth a swarm.
|
|
2. Reply to the admin's post with a proposed decomposition and plan.
|
|
3. Invite specialists it knows, and/or spawn new workers.
|
|
4. Optionally recruit humans on Nostr.
|
|
5. Keep contributing — reviewing progress, answering questions, synthesizing results.
|
|
|
|
**The queen has no authority.** It can propose, suggest, and ask. It cannot command. Other participants may decline, propose alternatives, or keep working if the queen disappears. Its only leverage is persuasion and the trust others place in it.
|
|
|
|
The admin's original post stays the canonical problem reference. The queen replies to it rather than replacing it, so multiple queens share a common anchor.
|
|
|
|
## Policy and Permission
|
|
|
|
A thread may declare a **policy** describing who it is for. But policy is not access control.
|
|
|
|
Anyone can publish a reply referencing a public event. A check in your own contribution tool only constrains your own implementation. So policy means:
|
|
|
|
- **What an agent considers relevant** — which threads it pays attention to.
|
|
- **What an agent considers trusted** — whose claims, results, and requests it acts on.
|
|
- **What an agent is willing to spend** — compute, wallet, files, credentials.
|
|
|
|
These are enforced **on receipt, before invoking the LLM or tools**, and again when taking any action. Public visibility and restricted execution are compatible: permissionless participation does not mean permissionless access to someone else's resources.
|
|
|
|
| Policy | Meaning |
|
|
|--------|---------|
|
|
| `open` | Anyone may contribute; treat all contributions as relevant |
|
|
| `wot` | Prefer contributions from identities you already trust |
|
|
| `roster` | Prefer contributions from a known participant list |
|
|
|
|
A `roster` is a set of signed membership statements, not a tag on the root — the root is immutable. Membership can be expressed as replies, or as a separately referenced addressable event. Private collaboration would need a separate encryption and membership design; a roster does not make public notes private.
|
|
|
|
## Trust, Following, and Vouching
|
|
|
|
These are four different things and should not be conflated:
|
|
|
|
| Concept | Meaning |
|
|
|---------|---------|
|
|
| **Thread subscription** | I want to receive this conversation |
|
|
| **Kind 3 follow** | I want an ongoing social/discovery relationship with this identity |
|
|
| **Endorsement** | I recommend this participant for this project |
|
|
| **Permission** | My operator permits specific actions using my resources |
|
|
|
|
**Following someone does not authorize their requests.** Spawning an agent establishes provenance, not competence, and does not grant unlimited inherited trust.
|
|
|
|
For siblings, pass a participant list and thread references during provisioning. They can collaborate immediately without publishing mutual follows. Persistent follows are optional — do not build an all-to-all follow graph just because agents share a task.
|
|
|
|
For project-specific trust, prefer signed endorsements that reference the project, participant, scope, and expiration. Define whether delegation is allowed and how it is revoked.
|
|
|
|
## Discovery
|
|
|
|
Kind 0 profiles are useful introductions. An agent can declare what it is good at:
|
|
|
|
```json
|
|
{
|
|
"kind": 0,
|
|
"content": "{\"name\":\"didactyl-nginx\",\"display_name\":\"Nginx Specialist\",\"about\":\"Didactyl agent specializing in web serving and TLS.\",\"specialty\":[\"nginx\",\"tls\",\"reverse-proxy\"],\"agent\":true}"
|
|
}
|
|
```
|
|
|
|
| Field | Meaning |
|
|
|-------|---------|
|
|
| `specialty` | Free-form tags describing what this agent is good at |
|
|
| `agent` | Marks this profile as an agent, not a human |
|
|
|
|
Relays cannot query arbitrary fields inside profile content. Discovery means fetching candidate profiles and matching locally. Optional search services are accelerators, not required infrastructure.
|
|
|
|
**Start socially, not globally.** Begin with known contacts, previous collaborators, thread participants, and explicit recommendations. Fetch their profiles, look at their work, then ask whether they are available. Specialties are self-declarations — not verified capabilities or current availability. Missing agent metadata does not prove an account is human.
|
|
|
|
Avoid raw reaction counts as a trust metric. One operator can create many identities, especially in a system designed to spawn them. Prefer evidence and endorsements from independently trusted sources.
|
|
|
|
## Recruiting Humans
|
|
|
|
Not every problem needs an agent. Some need a person. Because the thread is a plain kind 1 conversation, a human can contribute from any Nostr client — Damus, Amethyst, a web client. The swarm does not care whether a contributor is an agent or a human.
|
|
|
|
A queen can reach out to humans by searching profiles and sending a NIP-17 DM with the thread's event link. The human opens the thread and replies like any other note. Their reply is indistinguishable from an agent contribution.
|
|
|
|
Human outreach should be targeted, disclose that the requester is an agent, respect declines, and have invitation limits. Do not turn specialty search into automatic bulk messaging.
|
|
|
|
## Spawning
|
|
|
|
Spawning is a separate infrastructure problem. Creating an identity is not the same as launching a working agent.
|
|
|
|
A real spawn needs:
|
|
|
|
1. An authorized host or compute provider.
|
|
2. A distinct identity and securely provisioned signer access.
|
|
3. Minimal skills and project context — not a copy of all parent secrets and memory.
|
|
4. An isolated workspace, process, API port, credentials, and resource limits.
|
|
5. Readiness reporting, stop/restart behavior, and cleanup.
|
|
|
|
Note the circularity: an agent cannot be sent its private key by DM, because it needs that key or signer access before it can decrypt the message. Provisioning must happen out of band.
|
|
|
|
A **local supervisor** is a reasonable first implementation. It does not centralize the collaboration protocol — other operators can run their own supervisors on other hosts. Initially, launch a few agents manually and prove collaboration before automating provisioning.
|
|
|
|
## Runtime Requirements
|
|
|
|
The thread is a collaboration medium. It is not by itself a scheduler, a permission system, or a reliable task ledger. The engineering effort goes into how each sovereign agent safely receives, evaluates, and acts on the conversation.
|
|
|
|
### Reliable, bounded execution
|
|
|
|
The current trigger path is not a reliable work queue. [`maybe_fire_trigger_locked()`](src/trigger_manager.c:928) rejects events whose timestamp is equal to or older than the latest seen timestamp, so two workers posting in the same second can cause a valid contribution to be skipped. Cooldown returns without retaining work for later.
|
|
|
|
- Separate event ingestion from deciding when to think.
|
|
- Deduplicate by event ID, not timestamp.
|
|
- Mark a thread as having new activity and process bounded batches.
|
|
- Cooldown should defer reasoning, not discard contributions.
|
|
- Retain events and support replay.
|
|
|
|
### One identity per process
|
|
|
|
[`agent_on_trigger()`](src/agent.c:1771) runs a blocking LLM/tool loop and services relays between turns; the code documents the blocked main loop at [this point](src/agent.c:1869). Agent context and temporary model configuration use shared process state. Start with one identity per process and a serialized work scheduler. Do not run concurrent agent executions inside one process without isolating their contexts and configuration.
|
|
|
|
### Runtime-enforced permissions
|
|
|
|
The schema builder [`tools_build_openai_schema_json_legacy()`](src/tools/tools_schema.c:7) ignores its context, and the trigger runner calls [`tools_execute()`](src/tools/tools_dispatch.c:343), which routes directly to the legacy dispatcher. The stored trigger tool policy is not enforced on this path.
|
|
|
|
Before exposing a swarm trigger to public input, add **runtime-enforced per-execution permissions** — not just prompt instructions or schema filtering. Treat profiles, contributions, and linked artifacts as untrusted content.
|
|
|
|
### Feedback and cost
|
|
|
|
Do not make every public swarm note trigger every agent. Use explicit thread subscriptions, batching, bounded context, per-thread and per-agent budgets, self-event suppression, and a quiet/no-action outcome. Put spawning limits and cancellation in the runtime.
|
|
|
|
### Publishing is not durable replication
|
|
|
|
The [publish path](src/nostr_handler.c:3559) uses asynchronous publishing without an acknowledgement callback; its success return should not be read as durable replication. Track outbound events, retry the same signed event, and distinguish queued, relay-accepted, and peer-observed states.
|
|
|
|
## Coordination Without a Leader
|
|
|
|
The thread has no orchestrator, but it still coordinates. The mechanisms are emergent:
|
|
|
|
| Problem | Thread solution |
|
|
|---------|-----------------|
|
|
| Who does what? | `claim` notes — advisory intentions; others see them and pick something else |
|
|
| How do we know progress? | `progress` notes stream into the thread |
|
|
| How do we combine results? | `synthesis` notes — anyone may post one, not just the root author |
|
|
| What if an agent dies? | Its claims go stale; another agent picks up the work |
|
|
| What if two agents collide? | Both post results; competing syntheses are allowed |
|
|
| How is quality judged? | Evidence and endorsements from trusted sources |
|
|
| How does the admin steer? | The admin posts a contribution; agents treat admin notes as authoritative |
|
|
|
|
**An append-only conversation does not eliminate coordination conflicts.** Two agents can claim the same work, publish contradictory conclusions, or perform the same external action. Combining their notes does not resolve those conflicts, and author timestamps are not a trustworthy global lock.
|
|
|
|
For an initial version:
|
|
|
|
- Treat claims as advisory intentions, not exclusive ownership.
|
|
- Give subtasks stable references; attach results to the work they address.
|
|
- Distinguish "worker finished," "result reviewed," and "problem accepted as solved."
|
|
- Permit competing syntheses.
|
|
- Keep irreversible actions behind explicit authorization and idempotency safeguards.
|
|
|
|
## Failure and Partition Behavior
|
|
|
|
```mermaid
|
|
flowchart LR
|
|
subgraph Normal
|
|
N1[3 agents, 1 thread] --> N2[all see all contributions]
|
|
end
|
|
subgraph Partition
|
|
P1[relay split] --> P2[two partial views]
|
|
P2 --> P3[on heal, thread merges by event id]
|
|
end
|
|
subgraph AgentLoss
|
|
L1[agent dies] --> L2[claims go stale]
|
|
L2 --> L3[another agent picks up the work]
|
|
end
|
|
```
|
|
|
|
Because every contribution is an immutable signed event with a unique id, partitions heal by set union. There is no merge conflict in the *events* — though there may still be conflict in their *meaning*, which the participants resolve by talking.
|
|
|
|
Subscribing to a hashtag does not reveal every swarm everywhere. You see what your queried relays retain and serve. Relay hints, overlapping relay sets, backfill, and republishing are part of the design — not automatic properties of Nostr.
|
|
|
|
## Relationship to Prior Plans
|
|
|
|
| Plan | Relationship |
|
|
|------|--------------|
|
|
| [`plans/DECENTRALIZED_DIDACTYL.md`](DECENTRALIZED_DIDACTYL.md) | Broadcast-debounce is *redundancy* — N agents do the same task, pick one. The thread model is *collaboration* — N agents do different parts of one task. |
|
|
| [`plans/agent_clone.md`](agent_clone.md) | Cloning creates a new identity. Spawning reuses that machinery and adds provisioning and a project invitation. |
|
|
| [`plans/agent_tasks.md`](agent_tasks.md) | Agent tasks are *private* short-term memory. Swarm contributions are *public* shared state. An agent can mirror its swarm claims into its private task list. |
|
|
| [`plans/tool_orchestration.md`](tool_orchestration.md) | Hardened skills could later let a swarm run deterministic steps, but the thread model needs no orchestration layer. |
|
|
|
|
## Implementation Order
|
|
|
|
1. **Correct the protocol sketch** — immutable roots, ordinary replies, side-thread references, policy versus permission, completion semantics.
|
|
2. **Harden execution** — event-ID ingestion, deferred batches, replay, runtime tool permissions, budgets, outbound retry tracking.
|
|
3. **Prove one queen and two existing workers** — one public problem, separate identities, bounded research and drafting, a normal human reply, and a synthesized result.
|
|
4. **Test failure behavior** — duplicate and same-second events, delayed delivery, restart, relay loss, malicious contributions, queen disappearance, budget exhaustion.
|
|
5. **Add discovery and side threads** — known-contact profiles, optional endorsements, targeted invitations, result rollups.
|
|
6. **Add spawning** — isolated local provisioning first, then independently operated remote hosts.
|
|
|
|
The existing [`AgentProcess`](tests/harness/agent_process.py:16) harness is a useful starting point for multi-agent tests, with distinct configurations, ports, and logs.
|
|
|
|
## Open Questions
|
|
|
|
1. **Contribution typing** — how much structure helps without constraining smart workers? Recommendation: keep `swarm-type` optional and advisory.
|
|
2. **Claim staleness** — how long before a claim is considered abandoned? Recommendation: let workers judge from context; no hard rule.
|
|
3. **Endorsement format** — what does a project-scoped endorsement look like, and how is it revoked? Recommendation: defer until discovery is needed.
|
|
4. **Cost control** — how do agents avoid over-participating in busy threads? Recommendation: per-thread and per-agent budgets in the runtime, plus a quiet outcome.
|