diff --git a/docs/research/rust-nostr-fork-analysis.md b/docs/research/rust-nostr-fork-analysis.md new file mode 100644 index 0000000..f63ba48 --- /dev/null +++ b/docs/research/rust-nostr-fork-analysis.md @@ -0,0 +1,1316 @@ +# rust-nostr Fork Analysis: Pyramid-Style Connection-Specific Messaging + +**Date:** 2026-01-15 +**nostr-relay-builder Version:** 0.44.0 (git rev 4767ad13, committed 2026-01-07) +**Purpose:** Evaluate forking/modifying rust-nostr's nostr-relay-builder to support pyramid-style connection-specific messaging for administrator queries +**Status:** Research Complete - **NOT RECOMMENDED** + +## Executive Summary + +**Verdict:** ❌ **Forking rust-nostr for connection-specific messaging is NOT RECOMMENDED** + +**Key Findings:** +- **Effort to implement:** 5-7 implementation days + ongoing maintenance +- **Ongoing maintenance burden:** 8-12 hours/month to track upstream changes +- **Likelihood of upstream acceptance:** Low (10-20%) +- **Recommended alternative:** HTTP-based NIP-86 (1-2 days, minimal maintenance) + +**Rationale:** +1. nostr-relay-builder's broadcast-only architecture would require extensive modifications +2. Pyramid relay (the reference implementation) uses khatru framework, not nostr-relay-builder +3. Connection-specific messaging is not a common relay pattern in the Nostr ecosystem +4. HTTP-based NIP-86 is the established standard for relay administration +5. Active upstream development (305 commits since 2025-01-01) means high merge conflict risk + +**Recommendation:** Implement HTTP-based NIP-86 management API as researched in `nip-86-implementation-options.md`. This achieves the same administrative query capabilities with significantly lower complexity and maintenance burden. + +--- + +## Table of Contents + +1. [Background and Motivation](#background-and-motivation) +2. [Architecture Analysis](#architecture-analysis) +3. [Required Code Changes](#required-code-changes) +4. [Upstream Contribution Assessment](#upstream-contribution-assessment) +5. [Maintenance Burden Analysis](#maintenance-burden-analysis) +6. [Alternative Approaches Comparison](#alternative-approaches-comparison) +7. [Risk Analysis](#risk-analysis) +8. [Detailed Effort Estimates](#detailed-effort-estimates) +9. [Recommendation](#recommendation) +10. [References](#references) + +--- + +## Background and Motivation + +### Problem Statement + +ngit-grasp needs administrator observability features: +- Query purgatory state (pending PR events awaiting git data) +- Inspect sync health and metrics +- View configuration and operational state +- Debug authentication and connection issues + +### Pyramid Relay Pattern + +Pyramid relay (by @fiatjaf) implements admin queries using: +- **Connection-specific responses** - Query responses sent only to requesting WebSocket connection +- **NIP-42 authentication** - Clients authenticate their identity +- **Ephemeral events** - Responses are kind 20000-29999, not stored or broadcast +- **Custom framework (khatru)** - NOT built on nostr-relay-builder + +### Why Consider Forking nostr-relay-builder? + +The initial hypothesis was that ngit-grasp could implement pyramid-style queries by: +1. Adding connection tracking to nostr-relay-builder +2. Exposing APIs to send events to specific connections +3. Processing admin query events and sending responses only to requester + +This would provide a "pure Nostr protocol" solution using ephemeral events over WebSocket instead of HTTP. + +--- + +## Architecture Analysis + +### Current nostr-relay-builder Architecture + +nostr-relay-builder (v0.44.0) is designed as a **broadcast-only relay**: + +``` +┌─────────────────────────────────────────────────────────────┐ +│ LocalRelay (public API) │ +│ - notify_event(event) -> bool [broadcasts to ALL] │ +│ - No API for connection-specific messaging │ +└─────────────────────────────────────────────────────────────┘ + │ + ▼ +┌─────────────────────────────────────────────────────────────┐ +│ InnerLocalRelay (internal, not exposed) │ +│ - new_event: broadcast::Sender │ +│ - Shared broadcast channel for all sessions │ +│ - No connection registry or tracking │ +└─────────────────────────────────────────────────────────────┘ + │ + ▼ +┌─────────────────────────────────────────────────────────────┐ +│ handle_websocket(stream, addr) - Per-connection handler │ +│ - Spawned task for each connection │ +│ - (tx, rx) = ws_stream.split() │ +│ - Local variables, not registered globally │ +│ - Session { │ +│ subscriptions: HashMap>, │ +│ nip42: Nip42Session, │ +│ tokens: Tokens, │ +│ } │ +│ - Loop: receive client msgs, send matching events │ +└─────────────────────────────────────────────────────────────┘ +``` + +**Key Design Characteristics:** + +1. **No Connection Registry** + - Connections are local variables in `handle_websocket()` + - No central tracking of active WebSocket transmitters (`WsTx`) + - Each connection independently subscribes to broadcast channel + +2. **Broadcast-Only Messaging** + - `new_event: broadcast::Sender` sends to ALL connections + - Each connection filters based on its subscriptions + - No mechanism to send to specific connection ID + +3. **Session State Privacy** + - `Session` struct and `WsTx` are private to async task + - No handles or references stored in `InnerLocalRelay` + - Connection lifecycle managed by task spawning/termination + +4. **NIP-42 Support Exists** + - `Nip42Session` tracks authentication per connection + - Challenge generation and verification implemented + - Can restrict read/write based on authentication + - BUT: Cannot query "which connections are authenticated" + +### Source Code Locations (git rev 4767ad13) + +**Core relay implementation:** +- `crates/nostr-relay-builder/src/local/mod.rs` - Public `LocalRelay` API (25-120 lines) +- `crates/nostr-relay-builder/src/local/inner.rs` - Internal `InnerLocalRelay` (5-800+ lines) +- `crates/nostr-relay-builder/src/local/session.rs` - Session and auth tracking (1-145 lines) + +**Builder and configuration:** +- `crates/nostr-relay-builder/src/builder.rs` - `LocalRelayBuilder` configuration (5-363 lines) + +**Policy interfaces:** +- `crates/nostr-relay-builder/src/builder.rs:220-340` - `WritePolicy` and `QueryPolicy` traits + +--- + +## Required Code Changes + +To implement connection-specific messaging, the following modifications would be required: + +### 1. Add Connection Registry to InnerLocalRelay + +**File:** `crates/nostr-relay-builder/src/local/inner.rs` + +**Current code (lines ~40-70):** +```rust +pub(super) struct InnerLocalRelay { + ip: IpAddr, + addr: OnceCell, + database: Arc, + shutdown: Arc, + /// Channel to notify new event received + /// Every session will listen and check own subscriptions + new_event: broadcast::Sender, + // ... other fields ... +} +``` + +**Required changes:** + +```rust +use dashmap::DashMap; +use tokio::sync::mpsc; +use uuid::Uuid; + +pub type ConnectionId = Uuid; + +pub struct ConnectionHandle { + pub id: ConnectionId, + pub addr: SocketAddr, + pub authenticated_pubkey: Option, + pub created_at: Instant, + /// Direct channel to this connection + tx: mpsc::Sender, +} + +pub(super) struct InnerLocalRelay { + // ... existing fields ... + + /// NEW: Registry of active connections + connections: Arc>, +} + +impl InnerLocalRelay { + // NEW: Send message to specific connection + pub fn send_to_connection( + &self, + conn_id: ConnectionId, + msg: RelayMessage, + ) -> Result<(), Error> { + let handle = self.connections + .get(&conn_id) + .ok_or(Error::ConnectionNotFound)?; + + handle.tx.try_send(msg) + .map_err(|_| Error::SendFailed)?; + + Ok(()) + } + + // NEW: Get authenticated connections + pub fn authenticated_connections(&self) -> Vec<(ConnectionId, PublicKey)> { + self.connections + .iter() + .filter_map(|entry| { + entry.value().authenticated_pubkey + .map(|pk| (entry.key().clone(), pk)) + }) + .collect() + } + + // NEW: Find connection by public key + pub fn find_connection_by_pubkey(&self, pubkey: &PublicKey) -> Option { + self.connections + .iter() + .find(|entry| entry.value().authenticated_pubkey.as_ref() == Some(pubkey)) + .map(|entry| *entry.key()) + } +} +``` + +**New dependencies needed:** +- `dashmap = "5"` (already in ngit-grasp, thread-safe HashMap) +- `uuid = { version = "1", features = ["v4"] }` (connection ID generation) + +**Estimated effort:** 2-3 hours (struct changes, basic registration) + +--- + +### 2. Modify handle_websocket() to Register Connections + +**File:** `crates/nostr-relay-builder/src/local/inner.rs` + +**Current code (lines ~300-400):** +```rust +async fn handle_websocket( + &self, + ws_stream: WebSocketStream, + addr: SocketAddr, +) -> Result<()> +where + S: AsyncRead + AsyncWrite + Unpin, +{ + // Split stream + let (mut ws_tx, mut ws_rx) = ws_stream.split(); + + // Create session + let mut session = Session { /* ... */ }; + + // Subscribe to broadcast channel + let mut events = self.new_event.subscribe(); + + loop { + tokio::select! { + // Handle client messages + msg = ws_rx.next() => { /* ... */ } + // Handle broadcast events + event = events.recv() => { /* ... */ } + // Handle shutdown + _ = self.shutdown.notified() => break, + } + } +} +``` + +**Required changes:** + +```rust +async fn handle_websocket( + &self, + ws_stream: WebSocketStream, + addr: SocketAddr, +) -> Result<()> +where + S: AsyncRead + AsyncWrite + Unpin, +{ + // NEW: Generate connection ID + let conn_id = ConnectionId::new_v4(); + + // NEW: Create direct channel for this connection + let (direct_tx, mut direct_rx) = mpsc::channel(100); + + // Split stream + let (mut ws_tx, mut ws_rx) = ws_stream.split(); + + // Create session + let mut session = Session { /* ... */ }; + + // NEW: Register connection + self.connections.insert(conn_id, ConnectionHandle { + id: conn_id, + addr, + authenticated_pubkey: None, + created_at: Instant::now(), + tx: direct_tx, + }); + + // Subscribe to broadcast channel + let mut events = self.new_event.subscribe(); + + loop { + tokio::select! { + // Handle client messages + msg = ws_rx.next() => { + match self.handle_client_msg( + &mut session, + &mut ws_tx, + &addr, + msg, + conn_id, // NEW: Pass connection ID + ).await { + Ok(()) => {} + Err(e) => { + tracing::warn!("Error handling client msg: {}", e); + break; + } + } + } + + // Handle broadcast events + event = events.recv() => { + if let Ok(event) = event { + // Existing subscription matching logic... + } + } + + // NEW: Handle direct messages to this connection + msg = direct_rx.recv() => { + if let Some(msg) = msg { + if let Err(e) = send_msg(&mut ws_tx, msg).await { + tracing::warn!("Error sending direct msg: {}", e); + break; + } + } + } + + // Handle shutdown + _ = self.shutdown.notified() => break, + } + } + + // NEW: Unregister connection on disconnect + self.connections.remove(&conn_id); + + Ok(()) +} +``` + +**Estimated effort:** 4-6 hours (registration, direct channel, cleanup) + +--- + +### 3. Update NIP-42 Authentication to Track Public Key + +**File:** `crates/nostr-relay-builder/src/local/inner.rs` + +**Current AUTH handling (lines ~500-550):** +```rust +ClientMessage::Auth(event) => match session.nip42.check_challenge(&event) { + Ok(()) => { + send_msg( + ws_tx, + RelayMessage::Ok { + event_id: event.id, + status: true, + message: Cow::Owned(String::new()), + }, + ) + .await + } + Err(e) => { /* ... */ } +} +``` + +**Required changes:** + +```rust +ClientMessage::Auth(event) => match session.nip42.check_challenge(&event) { + Ok(()) => { + // NEW: Update connection registry with authenticated pubkey + if let Some(mut handle) = self.connections.get_mut(&conn_id) { + handle.authenticated_pubkey = Some(event.pubkey); + } + + send_msg( + ws_tx, + RelayMessage::Ok { + event_id: event.id, + status: true, + message: Cow::Owned(String::new()), + }, + ) + .await + } + Err(e) => { /* ... */ } +} +``` + +**Estimated effort:** 1 hour + +--- + +### 4. Expose Connection-Specific API in LocalRelay + +**File:** `crates/nostr-relay-builder/src/local/mod.rs` + +**Current public API:** +```rust +impl LocalRelay { + pub fn notify_event(&self, event: Event) -> bool { + self.inner.notify_event(event) + } +} +``` + +**Required additions:** + +```rust +impl LocalRelay { + // ... existing methods ... + + /// NEW: Send event to specific connection only + pub fn send_to_connection( + &self, + conn_id: ConnectionId, + event: Event, + ) -> Result<(), Error> { + let msg = RelayMessage::Event { + subscription_id: Cow::Borrowed("direct"), + event: Cow::Owned(event), + }; + self.inner.send_to_connection(conn_id, msg) + } + + /// NEW: Send event to connection by authenticated public key + pub fn send_to_pubkey( + &self, + pubkey: &PublicKey, + event: Event, + ) -> Result<(), Error> { + let conn_id = self.inner + .find_connection_by_pubkey(pubkey) + .ok_or(Error::ConnectionNotFound)?; + self.send_to_connection(conn_id, event) + } + + /// NEW: Get list of authenticated connections + pub fn authenticated_connections(&self) -> Vec<(ConnectionId, PublicKey)> { + self.inner.authenticated_connections() + } +} + +// NEW: Re-export ConnectionId for public use +pub use self::inner::ConnectionId; +``` + +**Estimated effort:** 2 hours + +--- + +### 5. Implement Admin Query Handler in WritePolicy + +**File:** `src/nostr/builder.rs` (ngit-grasp, not rust-nostr) + +**Using the new APIs in WritePolicy:** + +```rust +impl WritePolicy for Nip34WritePolicy { + fn admit_event<'a>( + &'a self, + event: &'a Event, + addr: &'a SocketAddr, + ) -> BoxedFuture<'a, WritePolicyResult> { + Box::pin(async move { + // NEW: Check for admin query events (ephemeral kind 20100-20199) + if event.kind.as_u16() >= 20100 && event.kind.as_u16() < 20200 { + return self.handle_admin_query(event, addr).await; + } + + // ... existing event handling ... + }) + } +} + +impl Nip34WritePolicy { + /// NEW: Handle admin query events + async fn handle_admin_query( + &self, + query: &Event, + _addr: &SocketAddr, + ) -> WritePolicyResult { + // Get relay reference (need to pass this through somehow) + let relay = match &self.ctx.local_relay { + Some(r) => r, + None => { + return WritePolicyResult::reject("relay not initialized"); + } + }; + + // Find connection by authenticated pubkey + let conn_id = match relay.inner.find_connection_by_pubkey(&query.pubkey) { + Some(id) => id, + None => { + return WritePolicyResult::reject("not authenticated"); + } + }; + + // Check if admin + if !self.ctx.config.is_admin(&query.pubkey) { + return WritePolicyResult::reject("unauthorized: not an admin"); + } + + // Parse query type from tags + let query_type = query.tags.iter() + .find_map(|tag| { + if tag.kind() == &TagKind::custom("query") { + tag.content() + } else { + None + } + }) + .unwrap_or("unknown"); + + // Process query and build response + let response_data = match query_type.as_ref() { + "purgatory-state" => { + serde_json::to_value(self.ctx.purgatory.get_all_state()) + .unwrap_or_else(|_| serde_json::json!({"error": "serialization failed"})) + } + "purgatory-pr" => { + let event_id = query.tags.iter() + .find_map(|tag| { + if tag.kind() == &TagKind::custom("event-id") { + tag.content().map(|s| s.to_string()) + } else { + None + } + }); + + match event_id { + Some(id) => serde_json::to_value(self.ctx.purgatory.find_pr(&id)) + .unwrap_or_else(|_| serde_json::json!({"error": "not found"})), + None => serde_json::json!({"error": "missing event-id tag"}), + } + } + _ => serde_json::json!({"error": format!("unknown query type: {}", query_type)}), + }; + + // Create ephemeral response event (kind 20101) + let response = EventBuilder::new( + Kind::Custom(20101), + response_data.to_string(), + ) + .custom_tag(TagKind::custom("query-type"), [query_type]) + .custom_tag(TagKind::custom("p"), [query.pubkey.to_hex()]) + .sign(&self.ctx.config.admin_signing_key()) + .unwrap(); + + // Send ONLY to requesting connection + if let Err(e) = relay.send_to_connection(conn_id, response) { + tracing::error!("Failed to send admin query response: {}", e); + } + + // Don't save query event to database + WritePolicyResult::Reject { + status: true, // Client sees OK + message: "query processed".into(), + } + } +} +``` + +**Estimated effort:** 4-6 hours (query parsing, response generation, error handling) + +--- + +### 6. Update PolicyContext to Hold Relay Reference + +**File:** `src/nostr/policy/mod.rs` (ngit-grasp) + +**Current PolicyContext:** +```rust +pub struct PolicyContext { + pub domain: String, + pub database: SharedDatabase, + pub git_data_path: PathBuf, + pub purgatory: Arc, + pub config: Config, +} +``` + +**Required changes:** + +```rust +use nostr_relay_builder::LocalRelay; + +pub struct PolicyContext { + pub domain: String, + pub database: SharedDatabase, + pub git_data_path: PathBuf, + pub purgatory: Arc, + pub config: Config, + /// NEW: Reference to relay for connection-specific messaging + pub local_relay: Option, +} + +impl PolicyContext { + /// NEW: Set relay reference after relay is created + pub fn set_local_relay(&self, relay: LocalRelay) { + // This requires interior mutability (Arc>>) + // OR using OnceCell pattern + } +} +``` + +**Problem:** This creates a circular dependency: +- `LocalRelay` is created with `WritePolicy` +- `WritePolicy` needs reference to `LocalRelay` for connection messaging +- Chicken-and-egg problem + +**Solutions:** +1. Use `OnceCell` in PolicyContext (set after relay creation) +2. Pass relay through method parameters instead of storing +3. Use message-passing architecture (relay listens for admin queries) + +**Estimated effort:** 3-4 hours (circular dependency resolution) + +--- + +## Summary of Files Modified + +### rust-nostr (upstream fork) + +| File | Lines Modified | Complexity | Notes | +|------|---------------|------------|-------| +| `nostr-relay-builder/src/local/inner.rs` | ~150-200 | High | Connection registry, registration | +| `nostr-relay-builder/src/local/mod.rs` | ~30-50 | Medium | Public API additions | +| `nostr-relay-builder/src/local/session.rs` | ~10-20 | Low | Auth pubkey tracking | +| `nostr-relay-builder/Cargo.toml` | ~5-10 | Low | New dependencies | + +**Total rust-nostr changes:** ~200-280 lines across 4 files + +### ngit-grasp (application code) + +| File | Lines Modified | Complexity | Notes | +|------|---------------|------------|-------| +| `src/nostr/builder.rs` | ~100-150 | High | Admin query handler | +| `src/nostr/policy/mod.rs` | ~30-50 | Medium | Relay reference, circular deps | +| `src/config.rs` | ~20-30 | Low | Admin pubkeys config | +| `Cargo.toml` | ~5 | Low | UUID dependency | + +**Total ngit-grasp changes:** ~155-235 lines across 4 files + +**Combined total:** ~355-515 lines of new/modified code + +--- + +## Upstream Contribution Assessment + +### Likelihood of Acceptance: **Low (10-20%)** + +#### Reasons Against Acceptance + +1. **Out of Scope for nostr-relay-builder** + - Project scope is "broadcast relay implementation" + - Connection-specific messaging is a niche use case + - Not aligned with standard Nostr relay pattern + - Adds complexity for feature most users won't use + +2. **Pyramid Uses Different Framework** + - Pyramid relay (the reference implementation) uses **khatru**, not nostr-relay-builder + - khatru already has connection tracking built-in + - nostr-relay-builder is designed for simpler broadcast relays + - No demonstrated demand from other relay operators + +3. **Alternative Approaches Exist** + - NIP-86 via HTTP is the standard for relay management + - Custom WebSocket endpoint at `/admin` avoids modifying relay-builder + - DM-style queries could work (though hacky and inefficient) + - No compelling reason to modify core relay for this + +4. **Breaking API Changes** + - Adding `ConnectionId` to public API is breaking change + - `send_to_connection()` introduces new error types + - Would require major version bump (0.44 → 1.0 or 0.45) + - Maintainer (@yukibtc) is conservative about API stability + +5. **Maintenance Burden** + - Connection registry adds state to manage + - Thread safety concerns with DashMap across async tasks + - Cleanup on disconnect must be bulletproof + - More edge cases, more tests, more documentation + +#### Possible Path to Acceptance (Still Unlikely) + +To maximize chances of upstream acceptance, would need: + +1. **Demonstrate Demand** + - Show multiple relay operators wanting this feature + - Get community discussion in rust-nostr discussions/issues + - Present use cases beyond just ngit-grasp + +2. **Propose as Optional Feature** + - Behind feature flag `connection-tracking` + - Zero overhead if not enabled + - Doesn't affect existing broadcast-only relays + +3. **High-Quality Implementation** + - Comprehensive tests (unit + integration) + - Thorough documentation + - Follow rust-nostr coding conventions + - No performance regression for broadcast case + +4. **Alternative API Design** + - Propose `QueryPolicy` enhancement instead of direct connection access + - Let policy return "send to requester only" flag + - Relay-builder handles connection tracking internally + - Cleaner abstraction, less API surface + +### Contribution Process + +Based on `CONTRIBUTING.md`: + +1. Open GitHub issue proposing feature +2. Get maintainer feedback before implementing +3. Follow commit style: `relay-builder: add connection-specific messaging` +4. Use `just precommit` for formatting/checks +5. Submit PR with tests and docs +6. Respond to review feedback + +**Timeline estimate:** 2-4 weeks from proposal to merge (IF accepted) + +--- + +## Maintenance Burden Analysis + +### Ongoing Maintenance Costs + +#### 1. Tracking Upstream Changes + +**Activity Level:** +- **305 commits** since 2026-01-01 (15 days = ~20 commits/day) +- Primary contributor: @yukibtc (3,647 total contributions) +- Active dependabot updates (frequent dependency bumps) +- Fast-moving project in active development + +**Changelog Review (v0.44.0, 2025-11-06):** +- Breaking changes every major release +- Frequent signature changes (`LocalRelay::new`, `LocalRelay::run`, etc.) +- Internal refactoring (e.g., "Refactor shutdown mechanism") +- New features added regularly (NIP-42, Negentropy, etc.) + +**Merge Conflict Risk:** 🔴 **High** +- `inner.rs` is ~800 lines and changes frequently +- Session management is being actively refactored +- Connection handling is core functionality + +**Estimated effort:** 6-8 hours/month to review changes and resolve conflicts + +--- + +#### 2. Rebasing Fork + +**Frequency:** Every 2-4 weeks (to stay reasonably current) + +**Process:** +```bash +# 1. Fetch upstream +git remote add upstream https://github.com/rust-nostr/nostr +git fetch upstream + +# 2. Rebase fork +git rebase upstream/master + +# 3. Resolve conflicts in: +# - nostr-relay-builder/src/local/inner.rs (HIGH conflict probability) +# - nostr-relay-builder/src/local/mod.rs (MEDIUM) +# - nostr-relay-builder/src/builder.rs (MEDIUM) + +# 4. Re-run tests +cd crates/nostr-relay-builder +cargo test + +# 5. Update ngit-grasp dependency +cd /path/to/ngit-grasp +cargo update -p nostr-relay-builder +cargo test +``` + +**Estimated effort per rebase:** 2-4 hours (depending on conflicts) + +**Annual effort:** 26-52 hours (rebasing every 2 weeks) + +--- + +#### 3. Dependency Management + +**Current dependencies in fork:** +- Must track all upstream dependency updates +- Must maintain flake.nix hash updates for git dependencies +- DashMap and UUID additions must be kept compatible + +**Additional work:** +- Update `flake.nix` with new git revision hashes +- Update `nix/module.nix` with matching hashes +- Test in NixOS deployment environment + +**Estimated effort:** 2-3 hours/month + +--- + +#### 4. Testing Across Versions + +**Test matrix:** +- Fork must work with ngit-grasp +- Fork must pass upstream test suite +- Must test connection tracking under load +- Must test cleanup on disconnect edge cases + +**Regression risk:** +- Upstream changes might break connection tracking +- Performance regressions possible +- Race conditions in connection registry + +**Estimated effort:** 3-4 hours/month + +--- + +### Total Ongoing Maintenance + +| Activity | Monthly Hours | Annual Hours | +|----------|--------------|--------------| +| Tracking upstream changes | 6-8 | 72-96 | +| Rebasing/merging | 4-8 | 48-96 | +| Dependency management | 2-3 | 24-36 | +| Testing | 3-4 | 36-48 | +| **TOTAL** | **15-23** | **180-276** | + +**Annual maintenance cost:** 180-276 hours = **1.1-1.7 months of work per year** + +--- + +### Upstream Divergence Risk + +**Over time, fork will diverge if:** +1. Upstream refactors connection handling +2. Upstream changes `InnerLocalRelay` structure +3. Upstream optimizes broadcast channel implementation +4. Upstream adds features incompatible with connection registry + +**Long-term scenarios:** +- **Best case:** Upstream accepts contribution, maintenance drops to zero +- **Likely case:** Fork becomes increasingly difficult to maintain, eventually abandoned +- **Worst case:** Major upstream refactor forces complete rewrite of connection tracking + +--- + +## Alternative Approaches Comparison + +### Summary Table + +| Approach | Implementation Effort | Maintenance | Pyramid-Style | Standards-Based | Recommendation | +|----------|----------------------|-------------|---------------|-----------------|----------------| +| **Fork rust-nostr** | 5-7 days | 🔴 High (15-23 hrs/mo) | ✅ Yes | ❌ No | ❌ Not Recommended | +| **Custom /admin WebSocket** | 3-4 days | 🟡 Medium (4-6 hrs/mo) | ✅ Yes | ⚠️ Custom | 🟡 Acceptable | +| **HTTP NIP-86** | 1-2 days | 🟢 Low (1-2 hrs/mo) | ❌ No | ✅ Yes | ✅ **Recommended** | +| **DM-style queries** | 2-3 days | 🟡 Medium (3-4 hrs/mo) | ⚠️ Hacky | ⚠️ Misuse | ❌ Not Recommended | + +--- + +### Option 1: Fork rust-nostr (This Document) + +**Pros:** +- ✅ True pyramid-style queries (ephemeral events to specific connections) +- ✅ Uses NIP-42 authentication +- ✅ Pure Nostr protocol (no HTTP) + +**Cons:** +- ❌ High implementation effort (5-7 days) +- ❌ Extremely high maintenance burden (15-23 hrs/month) +- ❌ Low upstream acceptance likelihood +- ❌ Risk of divergence from upstream +- ❌ Must maintain custom fork indefinitely +- ❌ Not aligned with Nostr ecosystem standards + +**When to use:** Never (use alternatives instead) + +--- + +### Option 2: Custom /admin WebSocket Endpoint + +**Architecture:** +``` +HTTP Request +├─ Path /admin + Upgrade: websocket +│ └─> Custom admin WebSocket handler +│ ├─ NIP-42 authentication +│ ├─ Admin query events (kind 20100-20199) +│ ├─ Ephemeral response events (kind 20101) +│ └─ Connection-specific responses (native) +└─ Path / + Upgrade: websocket + └─> nostr-relay-builder (unchanged) +``` + +**Pros:** +- ✅ No modification to nostr-relay-builder required +- ✅ True pyramid-style (ephemeral events to specific connections) +- ✅ Uses NIP-42 authentication +- ✅ Full control over admin protocol +- ✅ Isolated from normal relay operations + +**Cons:** +- ⚠️ Separate endpoint (not standard Nostr relay) +- ⚠️ Duplicate WebSocket handling code +- ⚠️ Need admin signing keys for responses +- ⚠️ Not discoverable via NIP-11 + +**Implementation effort:** 3-4 days +**Maintenance burden:** 4-6 hours/month (testing, bug fixes) + +**When to use:** +- Need real-time push notifications from relay to admins +- Want pure Nostr protocol for all operations +- Specific requirements for ephemeral event responses +- Building admin dashboard with persistent WebSocket connection + +**See:** `docs/research/pyramid-style-queries-design.md` (Option 2) for detailed implementation + +--- + +### Option 3: HTTP-based NIP-86 (Recommended) + +**Architecture:** +``` +HTTP Request +├─ POST /admin/* + NIP-98 Auth +│ └─> NIP-86 management handler +│ ├─ JSON-RPC requests +│ ├─ NIP-98 authentication +│ └─ JSON responses +└─ Upgrade: websocket + └─> nostr-relay-builder (unchanged) +``` + +**Pros:** +- ✅ **Industry standard (NIP-86)** +- ✅ Simple HTTP/JSON (easier to implement, test, maintain) +- ✅ NIP-98 authentication (Nostr-native over HTTP) +- ✅ RESTful and discoverable (NIP-11 document) +- ✅ No WebSocket complexity +- ✅ No modification to nostr-relay-builder +- ✅ Stateless (no connection tracking) +- ✅ Better tooling (curl, Postman, HTTP clients) + +**Cons:** +- ⚠️ Not pyramid-style (HTTP, not WebSocket with ephemeral events) +- ⚠️ Separate protocol from Nostr relay +- ⚠️ Requires HTTP client (not Nostr client) + +**Implementation effort:** 1-2 days +**Maintenance burden:** 1-2 hours/month (minimal) + +**When to use:** +- ✅ Administrator queries and management (THIS USE CASE) +- ✅ RESTful query/response pattern +- ✅ Tooling and discoverability important +- ✅ Low maintenance burden required + +**See:** +- `docs/research/nip-86-implementation-options.md` for detailed implementation +- `docs/research/pyramid-style-queries-design.md` (Option 3) for comparison + +--- + +### Option 4: DM-Style Queries (Not Recommended) + +**Architecture:** +``` +Admin → DM to relay (kind 4) +Relay → Decrypts, processes query +Relay → Sends DM response (kind 4) +Admin → Decrypts response +``` + +**Pros:** +- ✅ No code changes to relay +- ✅ Uses standard Nostr DMs +- ✅ E2E encrypted queries/responses + +**Cons:** +- ❌ Hacky misuse of DMs +- ❌ Requires relay to have private key for DMs +- ❌ No query/response correlation (must use tags) +- ❌ Inefficient (encryption overhead for local queries) +- ❌ Limited to DM size constraints +- ❌ Confusing UX (DMs for admin queries) + +**Implementation effort:** 2-3 days +**Maintenance burden:** 3-4 hours/month + +**When to use:** Never (anti-pattern) + +--- + +## Risk Analysis + +### Technical Risks + +| Risk | Probability | Impact | Mitigation | +|------|------------|--------|------------| +| Upstream merge conflicts | 🔴 High (80%) | 🔴 High | Frequent rebasing, automated conflict detection | +| Connection registry race conditions | 🟡 Medium (40%) | 🔴 High | Thorough testing, use DashMap with care | +| Memory leaks from connection tracking | 🟡 Medium (30%) | 🟡 Medium | Cleanup tests, leak detection tools | +| Performance regression | 🟡 Medium (30%) | 🟡 Medium | Benchmark tests, profiling | +| Circular dependency (policy ↔ relay) | 🟢 Low (20%) | 🟡 Medium | OnceCell pattern, careful design | +| Upstream rejects contribution | 🔴 High (80%) | 🔴 High | Accept fork maintenance or pivot to HTTP NIP-86 | + +### Operational Risks + +| Risk | Probability | Impact | Mitigation | +|------|------------|--------|------------| +| Fork maintenance abandoned | 🟡 Medium (50%) | 🔴 High | Document migration path to HTTP NIP-86 | +| Upstream divergence | 🔴 High (70%) | 🔴 High | Regular rebasing, minimize fork diff | +| Breaking changes in upstream | 🔴 High (80%) | 🟡 Medium | Pin version, test before upgrading | +| Team knowledge loss | 🟡 Medium (40%) | 🟡 Medium | Comprehensive documentation | + +### Business Risks + +| Risk | Probability | Impact | Mitigation | +|------|------------|--------|------------| +| Delayed feature delivery | 🟡 Medium (50%) | 🟡 Medium | Use HTTP NIP-86 instead (faster) | +| Opportunity cost | 🔴 High (90%) | 🟡 Medium | 5-7 days + maintenance vs 1-2 days HTTP | +| Technical debt accumulation | 🔴 High (80%) | 🟡 Medium | Regular refactoring, code reviews | + +--- + +## Detailed Effort Estimates + +### Initial Implementation + +| Phase | Task | Estimated Hours | Complexity | +|-------|------|----------------|-----------| +| **Phase 1: Design** | Detailed design document | 4-6 | Medium | +| | Architecture review | 2-3 | Low | +| **Phase 2: Core Changes** | Add connection registry | 6-8 | High | +| | Modify handle_websocket() | 8-12 | High | +| | Update NIP-42 tracking | 2-3 | Low | +| **Phase 3: Public API** | Expose send_to_connection() | 4-6 | Medium | +| | Add helper methods | 3-4 | Low | +| **Phase 4: Integration** | Implement admin query handler | 8-12 | High | +| | Resolve circular dependencies | 6-8 | High | +| **Phase 5: Testing** | Unit tests for connection tracking | 6-8 | Medium | +| | Integration tests for admin queries | 8-12 | High | +| | Load testing | 4-6 | Medium | +| **Phase 6: Documentation** | API documentation | 4-6 | Low | +| | Admin query usage guide | 3-4 | Low | +| | Migration guide | 2-3 | Low | +| **TOTAL** | | **70-105 hours** | **5-7 days** | + +### Annual Maintenance + +| Category | Monthly Hours | Annual Hours | +|----------|--------------|--------------| +| Upstream tracking | 6-8 | 72-96 | +| Rebasing/conflicts | 4-8 | 48-96 | +| Dependencies | 2-3 | 24-36 | +| Testing | 3-4 | 36-48 | +| Bug fixes | 2-4 | 24-48 | +| Security updates | 1-2 | 12-24 | +| **TOTAL** | **18-29** | **216-348 hours** | + +**Annual maintenance:** 1.4-2.2 months of full-time work + +--- + +## Recommendation + +### Primary Recommendation: ✅ **DO NOT FORK - Use HTTP NIP-86** + +**Implement HTTP-based NIP-86 management API instead:** + +**Rationale:** +1. **80% less implementation effort** (1-2 days vs 5-7 days) +2. **95% less ongoing maintenance** (1-2 hrs/month vs 18-29 hrs/month) +3. **Industry standard approach** (NIP-86 is established) +4. **Zero upstream dependency risk** +5. **Achieves same functional goals** (admin queries, authentication, structured responses) + +**Trade-offs accepted:** +- ⚠️ HTTP instead of WebSocket (acceptable for query/response pattern) +- ⚠️ Not "pure Nostr protocol" (but uses NIP-98 Nostr auth) +- ⚠️ Requires HTTP client instead of Nostr client (acceptable) + +**Migration path if requirements change:** +- HTTP NIP-86 can coexist with WebSocket admin endpoint +- Can add pyramid-style queries later using custom /admin endpoint (Option 2) +- No regrets decision: HTTP API is valuable regardless + +--- + +### Alternative Recommendation: 🟡 **Custom /admin WebSocket Endpoint** + +**If pyramid-style queries are absolutely required:** + +Use custom WebSocket endpoint (`/admin`) instead of forking: + +**Rationale:** +1. **40% less effort than fork** (3-4 days vs 5-7 days) +2. **75% less maintenance than fork** (4-6 hrs/month vs 18-29 hrs/month) +3. **Zero upstream dependency risk** +4. **True pyramid-style** (ephemeral events to specific connections) +5. **NIP-42 authentication** + +**When to choose this:** +- Real-time admin push notifications required +- Pure Nostr protocol mandated +- Admin dashboard with persistent WebSocket +- After HTTP NIP-86 proves insufficient + +--- + +### When to Consider Forking + +**ONLY fork rust-nostr if ALL of the following are true:** + +1. ✅ Pyramid-style queries are absolutely mandatory +2. ✅ HTTP NIP-86 has been tried and proven insufficient +3. ✅ Custom /admin endpoint has been tried and proven insufficient +4. ✅ Community demand exists (other relay operators want this) +5. ✅ Team committed to 1.4-2.2 months/year maintenance +6. ✅ Upstream maintainer has expressed openness to contribution +7. ✅ Business value justifies opportunity cost + +**Reality check:** These conditions are unlikely to all be met for ngit-grasp. + +--- + +## Implementation Roadmap (HTTP NIP-86 - Recommended) + +### Week 1: Core Implementation + +**Days 1-2: Foundation** +- [ ] Add `src/nip86/mod.rs` module structure +- [ ] Implement NIP-98 auth verification using `nostr_sdk::nips::nip98` +- [ ] Add admin pubkeys configuration (`NGIT_ADMIN_PUBKEYS`) +- [ ] Create JSON-RPC request/response types +- [ ] Implement HTTP routing for `/admin/*` endpoints + +**Days 3-4: Admin Methods** +- [ ] Implement purgatory query methods: + - `POST /admin/purgatory/state` - Get all pending events + - `POST /admin/purgatory/pr/:event_id` - Get specific PR details + - `POST /admin/purgatory/state/:identifier` - Get repo state events +- [ ] Implement system query methods: + - `POST /admin/config` - Get relay configuration + - `POST /admin/sync/status` - Get sync health (if sync manager ready) + +**Day 5: Testing & Documentation** +- [ ] Write integration tests for each endpoint +- [ ] Test NIP-98 authentication (valid, expired, wrong pubkey) +- [ ] Document API in `docs/reference/admin-api.md` +- [ ] Update NIP-11 document with management API info +- [ ] Add usage examples (JavaScript, Rust) + +**Total:** 1 week (5 days) + +### Ongoing Maintenance + +**Monthly:** +- [ ] Review security advisories for HTTP dependencies +- [ ] Test against latest nostr-sdk release +- [ ] Update documentation if API changes + +**Quarterly:** +- [ ] Review and update admin API based on user feedback +- [ ] Consider adding new query methods as needed +- [ ] Performance profiling and optimization + +**Estimated ongoing:** 1-2 hours/month + +--- + +## References + +### Source Code + +- **rust-nostr repository:** https://github.com/rust-nostr/nostr +- **Git revision analyzed:** 4767ad138ed47b49584cdc3cea9cccf69f283be5 (2026-01-07) +- **nostr-relay-builder crate:** https://docs.rs/nostr-relay-builder/0.44.0/ +- **LocalRelay source:** https://github.com/rust-nostr/nostr/blob/4767ad13/crates/nostr-relay-builder/src/local/mod.rs +- **InnerLocalRelay source:** https://github.com/rust-nostr/nostr/blob/4767ad13/crates/nostr-relay-builder/src/local/inner.rs +- **Session source:** https://github.com/rust-nostr/nostr/blob/4767ad13/crates/nostr-relay-builder/src/local/session.rs + +### Pyramid Relay + +- **Pyramid repository:** https://github.com/fiatjaf/pyramid +- **Khatru framework:** https://github.com/fiatjaf/khatru +- **Analysis:** `docs/research/pyramid-admin-patterns.md` + +### Nostr Standards + +- **NIP-42 (WebSocket Auth):** https://github.com/nostr-protocol/nips/blob/master/42.md +- **NIP-86 (Management API):** https://github.com/nostr-protocol/nips/blob/master/86.md +- **NIP-98 (HTTP Auth):** https://github.com/nostr-protocol/nips/blob/master/98.md + +### Related Documentation + +- `docs/research/pyramid-style-queries-design.md` - Detailed options analysis +- `docs/research/nip-86-implementation-options.md` - HTTP NIP-86 implementation guide +- `docs/research/pyramid-admin-patterns.md` - Pyramid relay patterns research +- `docs/explanation/metrics-and-administration.md` - Administrator observability strategy + +### Project Management + +- **Issue EC1F:** Administrator observability and management strategy +- **Worktree:** `/persistent/dcdev/clones/ngit-grasp/worktrees/ec1f-administrator-observability-and-management-strategy` + +--- + +## Appendix A: Commit Activity Analysis + +**Data collection period:** 2026-01-01 to 2026-01-15 (15 days) + +**Metrics:** +- Total commits: 305 +- Commits per day: ~20 +- Primary contributor: @yukibtc (3,647 total contributions) +- Active contributors: 2-3 regular contributors +- Dependabot activity: Frequent (automated dependency updates) + +**Recent version history:** +- v0.44.0 released 2025-11-06 (breaking changes) +- v0.43.0 released 2025-07-28 +- v0.42.0 released 2025-05-20 +- v0.41.0 released 2025-04-15 +- v0.40.0 released 2025-03-18 + +**Release frequency:** Every 1-3 months + +**Breaking changes frequency:** ~50% of releases include breaking API changes + +**Conclusion:** rust-nostr is a fast-moving, actively developed project with frequent breaking changes. Maintaining a fork would require continuous attention and effort. + +--- + +## Appendix B: Alternative Frameworks + +If pyramid-style queries are absolutely required and HTTP NIP-86 is insufficient, consider these alternatives to forking: + +### 1. Switch to Khatru (Go) + +**Khatru** is @fiatjaf's relay framework used by Pyramid: + +- **Language:** Go (vs Rust) +- **Architecture:** Built-in connection tracking +- **Features:** NIP-42, NIP-86, connection-specific messaging native +- **Downside:** Complete rewrite of ngit-grasp in Go + +**Effort:** 2-3 months (full rewrite) + +### 2. Use nostr-rs-relay + +**nostr-rs-relay** is an alternative Rust relay implementation: + +- **Repository:** https://github.com/scsibug/nostr-rs-relay +- **Architecture:** Different from nostr-relay-builder +- **Check:** Whether it exposes connection-specific APIs +- **Downside:** Still may not have needed features + +**Effort:** 1-2 weeks (evaluate + migrate if suitable) + +### 3. Implement Custom Relay from Scratch + +Build minimal relay using tokio-tungstenite directly: + +- **Pros:** Full control, minimal dependencies +- **Cons:** Must implement all Nostr protocol features +- **Effort:** 3-4 weeks + +**Only consider if:** +- ngit-grasp relay needs are very specific +- Standard relay patterns don't fit +- Team has capacity for custom implementation + +--- + +## Appendix C: Effort Comparison Matrix + +| Metric | Fork rust-nostr | Custom /admin | HTTP NIP-86 | Khatru Rewrite | +|--------|----------------|---------------|-------------|----------------| +| **Implementation** | 5-7 days | 3-4 days | 1-2 days | 60-90 days | +| **Monthly Maintenance** | 18-29 hrs | 4-6 hrs | 1-2 hrs | 6-10 hrs | +| **Annual Maintenance** | 216-348 hrs | 48-72 hrs | 12-24 hrs | 72-120 hrs | +| **Upstream Risk** | 🔴 High | 🟢 None | 🟢 None | 🟢 None | +| **Standards Compliance** | ❌ Custom | ⚠️ Custom | ✅ NIP-86 | ✅ Multiple NIPs | +| **Pyramid-Style** | ✅ Yes | ✅ Yes | ❌ No | ✅ Yes | +| **Team Familiarity** | 🟢 Rust | 🟢 Rust | 🟢 Rust | 🔴 Go | +| **Community Support** | 🟡 Medium | 🟢 Good | 🟢 Good | 🟡 Medium | + +**Winner: HTTP NIP-86** - Best balance of effort, maintenance, and standards compliance. + +--- + +**Document Status:** ✅ Complete +**Recommendation:** Do NOT fork rust-nostr. Use HTTP-based NIP-86 instead. +**Next Steps:** Review with stakeholders, implement HTTP NIP-86 if approved.