Files
c-relay-pg/plans/nip09_client_deletion_and_tombstones_plan.md

7.7 KiB

NIP-09 Client-Path Deletion + Tombstone (Option A) Implementation Plan

Background

Two verified defects:

  1. NIP-09 dead code on the client path: handle_deletion_request() is a complete, correct NIP-09 implementation (e/a tag parsing, authorship enforcement, hard delete via db_delete_event_by_id()), but it is only invoked from ingest_event() in main.c (external ingress). The client WebSocket path in websockets.c declares it at line 84 with zero call sites. Kind 5 events fall into the generic regular-event branch and are merely stored + broadcast — target events are never deleted.

  2. Resurrection vulnerability: deletion is a hard DELETE FROM events with no memory. A re-posted deleted event passes the duplicate pre-check (row gone), passes validation (signature still valid), and is re-stored + re-broadcast. Re-ingestion vectors: clients re-posting, the caching inbox poller, and custom backfill jobs.

Design Decisions

  • Tombstone table (deleted_event_ids): records every deleted event ID at deletion time. Checked during the duplicate pre-check; a tombstoned ID is treated as a duplicate and rejected with duplicate: event was deleted.
  • Atomicity: tombstone insert happens in the same SQL statement as the DELETE (CTE with RETURNING), so a crash between delete and tombstone-write is impossible.
  • Schema delivery: EMBEDDED_PG_SCHEMA_SQL is applied idempotently on every startup by postgres_db_apply_schema() — no migration logic needed; just add the table to the embedded schema.
  • Threading: the async event worker opens its own thread-local connection (g_thread_pg_conn is __thread), so handle_deletion_request() is safe to call from the worker thread.
  • Broadcast responsibility: in the async worker path, broadcast stays on the lws main thread via the existing completion mechanism (completion->should_broadcast), matching the established pattern.

Changes

1. Schema — src/pg_schema.h and src/pg_schema.sql

Add to both files (idempotent):

CREATE TABLE IF NOT EXISTS deleted_event_ids (
    event_id TEXT PRIMARY KEY,
    deleted_at BIGINT NOT NULL DEFAULT EXTRACT(EPOCH FROM NOW())::BIGINT
);

No FK to events (rows must survive the event's deletion — that's the point).

2. Atomic delete + tombstone — src/db_ops_postgres.c

Rework postgres_db_delete_event_by_id() into a single CTE statement:

WITH deleted AS (
    DELETE FROM events WHERE id = $1 AND pubkey = $2 RETURNING id
)
INSERT INTO deleted_event_ids (event_id)
SELECT id FROM deleted
ON CONFLICT (event_id) DO NOTHING;
  • Use PQexecParams with PGRES_TUPLES_OK expected (CTE returns rows) — note the current code checks PGRES_COMMAND_OK; the reworked version must check for TUPLES_OK or accept both.
  • Return value: count of deleted rows (from PQcmdTuples of the outer INSERT, or run a follow-up count). Preserve the existing contract: >0 = deleted count, 0 = nothing deleted, <0 = error.

Rework postgres_db_delete_events_by_address() with the same CTE pattern (both the d_tag and no-d_tag variants).

3. New status check — db_ops layer

Add postgres_db_event_id_status() to db_ops_postgres.c:

SELECT
    EXISTS(SELECT 1 FROM events WHERE id = $1) AS exists_in_events,
    EXISTS(SELECT 1 FROM deleted_event_ids WHERE event_id = $1) AS is_tombstoned;

Returns via out-params. Wire through:

4. Kind 5 dispatch — src/websockets.c

Three sites, all getting the same branch inserted before the ephemeral check:

a) Async worker (line ~600) — the PRIMARY client path:

if (job->event_kind == 5) {
    char del_error[512] = {0};
    if (handle_deletion_request(event_obj, del_error, sizeof(del_error)) != 0) {
        completion->success = 0;
        completion->should_broadcast = 0;
        completion->run_post_actions = 0;
        /* copy del_error into completion->error_message */
    } else {
        completion->success = 1;
        completion->should_broadcast = 1;
        completion->run_post_actions = 0;
    }
} else if (job->event_kind >= 20000 && job->event_kind < 30000) {
    ...

Note: handle_deletion_request() internally calls store_event() (nip009.c:135), which stores the deletion request itself — the branch must NOT also call store_event_core().

b) Sync block 1 (line ~1873) and c) Sync block 2 (line ~2670):

} else if (event_kind == 5) {
    char del_error[512] = {0};
    if (handle_deletion_request(event, del_error, sizeof(del_error)) != 0) {
        result = -1;
        /* copy del_error into error_message */
    } else {
        broadcast_event_to_subscriptions(event);
    }
}

5. Tombstone pre-check — two sites

a) Async worker duplicate check (websockets.c:580):

if (job->event_id[0] != '\0' && event_id_exists_in_db(job->event_id)) {
    ... duplicate ...
}

becomes a status check: if tombstoned → reject with duplicate: event was deleted (success=0 or duplicate-style OK with false? — decide: return OK false with message, consistent with duplicate handling but distinct message).

b) Ingress duplicate check (main.c:1338): same treatment.

Decision — response semantics for tombstoned re-posts: The cleanest Nostr-protocol behavior is ["OK", <id>, false, "duplicate: event was deleted"]. This tells clients the relay refuses to resurrect, without leaking whether the event ever existed. The existing duplicate path returns success=1 with a "duplicate" message; for tombstones we return success=0 since the event is NOT accepted. (Alternative: silently accept-and-drop; rejected as it confuses clients.)

6. Test — tests/9_nip_delete_test.sh

Add a resurrection test after the existing by-ID deletion test:

  1. Re-post event1 (the deleted kind 1 event) verbatim
  2. Assert the OK response is false with the deletion message
  3. Assert check_event_exists "$event1_id" still returns 0

Execution Order

  1. Schema (pg_schema.h + pg_schema.sql)
  2. db_ops layer (delete CTEs + status check + wiring)
  3. websockets.c dispatch branches (3 sites)
  4. Pre-check extensions (2 sites)
  5. Test extension
  6. Build with ./make_and_restart_relay.sh (NEVER make)
  7. Run tests/9_nip_delete_test.sh — all assertions must pass

Risks & Mitigations

  • CTE result status: WITH ... INSERT ... SELECT returns PGRES_TUPLES_OK when using RETURNING, PGRES_COMMAND_OK otherwise. The outer statement is an INSERT...SELECT without RETURNING → PGRES_COMMAND_OK, and PQcmdTuples returns the inserted count. Verify with a quick psql test during implementation.
  • store_event() in worker thread: handle_deletion_request() calls store_event() → store_event_post_actions() (main-thread-only per comment). For kind 5, post-actions only trigger WoT sync for admin kind 3 events — benign. Acceptable; note in code comment.
  • Tombstone growth: unbounded table growth. Acceptable for now (64-byte IDs); a retention policy can be added later via admin config.
  • Existing DBs: schema auto-applies on startup; no manual migration.