Files
didactyl/plans/skill_tool_policy.md

10 KiB

Skill Tool Policy — Implementation Plan

Goal: Let a skill define or limit which tools are exposed to the LLM when that skill runs. A skill's tools tag becomes an enforced allowlist instead of an ignored annotation.

Problem

Skills already declare a tools tag, e.g. the swarm worker skill:

["tools", "swarm_read,swarm_claim,swarm_contribute,nostr_post,nostr_query"]

But the trigger path ignores it. In agent_on_trigger():

char* tools_json = tools_build_openai_schema_json(&g_tools_ctx);

and tools_build_openai_schema_json() returns the full 92-tool schema regardless of the skill. So a worker skill that lists five tools is actually offered all 92 — including nostr_post, which the LLM used to hand-roll malformed swarm events instead of calling swarm_claim.

This is a known gap: plans/swarm.md — "The stored trigger tool policy is not enforced on this path." The docs already promise the behavior: docs/TOOLS.md and docs/CONTEXT.md.

What already exists (reuse, do not rebuild)

The plumbing is mostly in place:

Piece Location Status
tools_policy[256] field on the trigger active_trigger_t Exists
Parse tools tag from skill parse_trigger_runtime_tags() Exists
Store policy on trigger trigger_manager_add() Exists
Expose policy in trigger list JSON trigger_manager.c Exists
Apply policy to the schema — Missing
Pass policy into agent_on_trigger() execute_llm_action() Missing

So the work is: (1) a schema-filtering function, (2) thread the policy through to the trigger execution, (3) enforce at execution time as defense in depth.

Design

Policy format

The tools tag is a comma-separated list of tool names:

swarm_read,swarm_claim,swarm_contribute,nostr_post,nostr_query

This is a whitelist (allowlist). The named tools are the only tools exposed. There is no blacklist form in this design.

Wildcard

Support an explicit allow-all token so a skill can state its intent rather than rely on omission:

  • * (or the word all) → expose every tool. Equivalent to omitting the tag.

This matters because "no tag" and "tag = *" should mean the same thing, but the explicit form is self-documenting and lets a skill author say "I deliberately want everything" instead of leaving it ambiguous.

Semantics

tools tag value Result
absent all tools (backward compatible)
empty string all tools
* or all all tools (explicit wildcard)
swarm_read,swarm_claim only those two tools
swarm_read, bogus_tool only swarm_read; bogus_tool skipped + warned
  • Tag present, non-empty, not a wildcard → allowlist. Only the named tools are exposed.
  • Tag absent or empty → all tools exposed (backward compatible; matches docs/TOOLS.md: "If a skill has no requires_tool tags, all available tools are exposed.").
  • Unknown tool name in the list → ignored (do not fail the skill). Log a warning so typos are visible.
  • Whitespace around names is trimmed.

Why whitelist, not blacklist

  • Safer default. A new tool added to the codebase is not automatically granted to every skill. With a blacklist, every new tool is silently exposed everywhere until someone remembers to deny it.
  • Matches the existing docs. docs/TOOLS.md already frames this as "only the required and optional tools declared by the skill are exposed" — a whitelist.
  • Least privilege. The swarm bug is precisely a case where a worker should not have had nostr_post. A whitelist expresses that directly.

A blacklist form (e.g. -local_shell_exec) is listed as a possible follow-up, but is not part of this plan. If both are ever supported, the rule should be: whitelist first, then subtract any blacklist entries.

Optional: requires_tool alias

docs/SKILLS.md documents a repeatable ["requires_tool", "<name>"] tag. Support it as an alias for the same policy: if a skill has one or more requires_tool tags, union them into the allowlist. This is additive and low-risk. (Note: find_tag_value_string() returns only the first match, so a new helper is needed to collect all requires_tool values.)

Where filtering happens

Add a pure function in src/tools/tools_schema.c:

/* Return a new schema array containing only tools named in `policy_csv`.
 * policy_csv == NULL, "" or "*"/"all" returns the full schema (all tools).
 * Unknown names are skipped. Caller frees the returned string. */
char* tools_build_openai_schema_json_filtered(const tools_context_t* ctx,
                                              const char* policy_csv);

Implementation: build the full schema, parse it, iterate the array, and drop any entry whose function.name is not in the allowlist. Short-circuit to the full schema when the policy is empty or a wildcard. This mirrors the existing build_sandboxed_tools_json_local() pattern, so the approach is already proven in the codebase.

A shared helper tool_allowed_by_policy(name, policy_csv) should back both the schema filter and the execution guard, so the two can never disagree. It returns 1 when the policy is empty/wildcard, or when name is in the list.

Threading the policy to execution

  1. Change the signature of agent_on_trigger() to accept the policy:

    void agent_on_trigger(const char* skill_d_tag,
                          const char* skill_content,
                          const char* tools_policy,   /* NEW */
                          cJSON* triggering_event,
                          const char* relay_url);
    
  2. In execute_llm_action(), pass t->tools_policy:

    agent_on_trigger(t->skill_d_tag, t->skill_content, t->tools_policy, event, relay_url);
    
  3. In agent_on_trigger(), replace the schema build:

    char* tools_json = tools_build_openai_schema_json_filtered(&g_tools_ctx, tools_policy);
    
  4. Update the forward declaration in trigger_manager.c and any other callers.

Defense in depth: enforce at execution

Filtering the schema is the primary control, but the LLM could still emit a tool call for a name not in the schema. Add a guard in the trigger tool loop (agent.c) before tools_execute():

if (!tool_allowed_by_policy(tc->name, tools_policy)) {
    tool_result = strdup("{\"success\":false,\"error\":\"tool not permitted by skill policy\"}");
} else {
    tool_result = tools_execute(&g_tools_ctx, tc->name, tc->arguments_json);
}

This makes the policy a real boundary, not just a prompt hint — the same principle plans/swarm.md calls for: "runtime-enforced per-execution permissions — not just prompt instructions or schema filtering."

Scope decisions

  • Trigger path only (first cut). The DM path (agent.c) and the HTTP API path (http_api.c) keep the full schema. DM already has sender-tier gating; the API is admin-only. Filtering those is a follow-up.
  • Compose with skill_run sandbox. execute_skill_run() already filters for external skills. Leave it; the new policy applies to the trigger path. A later pass can intersect the two.
  • tools_policy is 256 bytes. A long allowlist could truncate. Either raise the buffer or validate length at parse time and warn. Note this in the change.

Files to change

File Change
src/tools/tools_schema.c Add tools_build_openai_schema_json_filtered()
src/tools/tools.h Declare the new function
src/agent.c agent_on_trigger() takes tools_policy; use filtered schema; add execution guard
src/trigger_manager.c Update forward decl; pass t->tools_policy
src/trigger_manager.h (optional) raise tools_policy size
docs/TOOLS.md Correct the doc to describe the tools tag (not just requires_tool)

Verification

  1. Unit-ish: --dump-schemas still lists all 92 (no policy). Add a debug path or log line showing the filtered count for a trigger.
  2. Swarm test: restart the workers, post a "my queen" question, and confirm in the worker debug log that the LLM request's tool list contains only the five swarm tools — no nostr_post.
  3. Behavioral: confirm workers now call swarm_claim (not nostr_post) and that claims carry swarm-task tags.
  4. Regression — no tag: a skill with no tools tag still gets all tools.
  5. Wildcard: a skill with ["tools", "*"] gets all tools (same as no tag).
  6. Allowlist: a skill with ["tools", "swarm_read,swarm_claim"] gets exactly those two.
  7. Unknown name: ["tools", "swarm_read,bogus"] gets swarm_read and logs a warning for bogus.
  8. Execution guard: a forced tool call for a non-allowlisted tool returns tool not permitted by skill policy rather than executing.

Risks

  • Over-restriction. A skill that lists too few tools will fail tasks it previously handled. Mitigation: the "no tag = all tools" fallback, plus a clear error when a blocked tool is called.
  • Truncation. The 256-byte tools_policy buffer. Mitigation: raise it or validate.
  • Silent typos. An unknown tool name is skipped, so a misspelled tool silently disappears. Mitigation: log a warning listing unknown names.

Follow-ups (out of scope)

  • Apply policy to the DM and HTTP API paths.
  • Intersect skill policy with the skill_run sandbox allowlist.
  • Support a deny-list form (e.g. -local_shell_exec) if needed.
  • Surface the effective tool set in /api/status or trigger_list for observability.