Files
ngit-grasp/docs/explanation/metrics-and-administration.md
T
DanConwayDev 78707898fd Refine metrics and administration strategy based on architecture review
- Add ConnectionTracker to target architecture as existing infrastructure
- Remove graceful degradation and quota warning events from quota enforcement scope
- Add Issue 76fe verification tasks to storage metrics implementation
- Add configuration sync reminders for all phases introducing new environment variables
- Adjust Phase 6 effort estimates to reflect reduced scope (3-4 days instead of 5-6 days)
- Update Phase 6 success criteria to match simplified implementation scope
2026-01-14 16:40:15 +00:00

50 KiB

Metrics and Administration Strategy

This document outlines the strategy for administrator observability and management of ngit-grasp relays. It defines the workflows administrators need, the information required to support those workflows, and how different audiences access that information.

Audiences

Three distinct audiences need access to relay information, each with different needs and trust levels:

1. Public Users (Relay Selection)

Who: Anyone considering using this relay Trust level: Untrusted Needs:

  • Is this relay healthy and available?
  • What are its capabilities and policies?
  • How much content does it host?

Access method:

  • NIP-11 relay information document (public, unauthenticated)
  • Relay homepage (human-friendly view of the same information)

2. Authenticated Users (Self-Service)

Who: Users with content on this relay (NIP-42 authenticated) Trust level: Authenticated, limited to own data Needs:

  • How much storage am I using?
  • What are my limits/quotas?
  • Which repositories do I own?

Access method: Replaceable Nostr events (user-specific, requires NIP-42 auth)

3. Administrators (Operations & Management)

Who: Relay operators Trust level: Fully trusted Needs:

  • Detailed operational metrics (connections, throughput, errors)
  • Per-user and per-repository breakdowns
  • Ability to take corrective action (blacklist, delete, quotas)
  • Alerting on problems

Access method:

  • Prometheus metrics (API key authenticated)
  • Management dashboard (authenticated)
  • Management API (authenticated)

Administrator Workflows

Administrators need to perform several key workflows. Each workflow requires specific information and actions.

Workflow 1: Health Monitoring

Goal: Understand if the relay is operating normally

Questions:

  • Is the relay accepting connections?
  • Are WebSocket connections stable?
  • Are git operations succeeding?
  • Is event processing working?
  • Is sync with other relays healthy?

Information needed:

  • Connection counts and rates
  • Error rates by operation type
  • Sync status with peer relays
  • Uptime and availability

Actions: None (observation only)

Metrics support:

  • ngit_websocket_connections_total
  • ngit_git_operations_total (by operation, result)
  • ngit_events_total (by kind, result)
  • ngit_sync_relay_status

Workflow 2: Capacity Planning

Goal: Understand resource utilization and plan for growth

Questions:

  • How much storage is being used?
  • How fast is storage growing?
  • How many repositories/users do we have?
  • Are we approaching any limits?

Information needed:

  • Total storage used (git + database)
  • Storage growth rate over time
  • Repository count and size distribution
  • User count and storage distribution
  • Configured limits vs current usage

Actions: None (observation only, may inform config changes)

Metrics support:

  • ngit_storage_bytes_total (by type: git, database)
  • ngit_repositories_total
  • ngit_storage_bytes_by_user (top N)
  • ngit_storage_bytes_by_repo (top N)

Workflow 3: Problem Investigation

Goal: Identify and understand specific problems

Questions:

  • Which users are consuming the most resources?
  • Which repositories are largest?
  • Who is causing errors?
  • Are there abuse patterns?

Information needed:

  • Per-user storage breakdown
  • Per-repository storage breakdown
  • Error logs with context
  • Connection patterns by source

Actions: None (investigation only)

Metrics support:

  • Per-user and per-repo gauges (top N)
  • Error counters with labels
  • Connection tracker data (aggregated, no IPs exposed)

Workflow 4: Corrective Action

Goal: Address identified problems

Scenarios:

  • User is abusing the relay (spam, oversized repos)
  • Repository needs to be removed (DMCA, policy violation)
  • User needs quota adjustment
  • IP needs to be blocked

Actions needed:

  • Blacklist user (npub) from all operations
  • Blacklist specific repository
  • Delete repository (with confirmation)
  • Set per-user storage quota
  • Set per-repository storage quota
  • Block IP address (temporary or permanent)

Access method: Management API (authenticated)

Workflow 5: Proactive Protection

Goal: Automatic enforcement without manual intervention

Scenarios:

  • Storage quota exceeded
  • Rate limit exceeded
  • Abuse pattern detected
  • Invalid signatures

Automatic responses:

  • Reject operations exceeding quotas (clear error message)
  • Rate limit connections/operations
  • Temporary ban for repeated violations
  • Permanent ban escalation for persistent abuse

Configuration:

  • Global storage limits
  • Per-user default quotas
  • Rate limit thresholds
  • Abuse detection sensitivity

Information Architecture

Public Information (NIP-11)

The NIP-11 relay information document is extended with a stats section for public metrics:

{
  "name": "My GRASP Relay",
  "description": "...",
  "supported_nips": [1, 11, ...],
  "software": "ngit-grasp",
  "version": "0.1.0",
  "stats": {
    "repositories_total": 1234,
    "events_total": 567890,
    "uptime_seconds": 864000,
    "users_total": 42
  }
}

Caching: Stats are cached and updated periodically (e.g., every 60 seconds) to avoid expensive computation on every NIP-11 request.

Security: Only aggregate, non-sensitive data. No per-user breakdowns, no storage details that could inform attacks.

User-Specific Information (Nostr Events)

Authenticated users can query their own storage usage via replaceable Nostr events:

  • User sends request (specific kind TBD)
  • Relay responds with replaceable event containing their data
  • Requires NIP-42 authentication
  • User can only see their own data

Content:

  • Total storage used
  • Storage quota/limit
  • Repository list with sizes
  • Quota warnings

Administrator Information (Prometheus)

Full operational metrics available via authenticated Prometheus endpoint:

Authentication: API key required (NGIT_METRICS_API_KEY)

Rationale for authentication:

  • Detailed metrics could inform attacks (storage limits, per-user data)
  • Fleet monitoring requires central collection (API key enables this)
  • Prometheus supports bearer token authentication natively

Metrics categories:

  • Connection metrics (counts, rates, errors)
  • Git operation metrics (by operation type, result)
  • Event metrics (by kind, result)
  • Storage metrics (total, per-user top N, per-repo top N)
  • Sync metrics (relay status, event counts)
  • System metrics (memory, CPU if available)

Configuration Model

Deployment Modes

Two deployment models with different configuration philosophies:

Declarative (NixOS, Docker):

  • All configuration via environment variables or CLI flags
  • No runtime configuration changes
  • Dashboard is read-only
  • Management API can take actions but not change config

Dynamic (Manual deployment):

  • Base configuration via environment/CLI
  • Runtime changes allowed via management API
  • Changes persist to local config file
  • Dashboard can modify settings

Configuration Options

# Metrics authentication (required)
NGIT_METRICS_API_KEY=<secret>

# Management API authentication
NGIT_MANAGEMENT_API_KEY=<secret>

# Deployment mode
NGIT_DECLARATIVE_ONLY=false  # true = no runtime config changes

# Storage limits
NGIT_MAX_STORAGE_BYTES=10737418240  # 10GB total
NGIT_DEFAULT_USER_QUOTA_BYTES=1073741824  # 1GB per user
NGIT_DEFAULT_REPO_QUOTA_BYTES=104857600  # 100MB per repo

Security Considerations

Metrics Exposure

Risk: Detailed metrics could help attackers plan resource exhaustion attacks.

Mitigation:

  • Prometheus endpoint requires API key
  • Public stats (NIP-11) contain only safe aggregates
  • Per-user/per-repo data never exposed publicly

User Privacy

Risk: Per-user metrics could reveal usage patterns.

Mitigation:

  • Users can only query their own data (NIP-42 auth)
  • Admin metrics show top-N only, not full user list
  • IP addresses never exposed in metrics (internal only)

Management API

Risk: Unauthorized access could delete data or disrupt service.

Mitigation:

  • Separate API key from metrics key
  • Destructive actions require confirmation
  • Audit log of all admin actions
  • Rate limiting on management endpoints

Target Architecture

This section defines the technical architecture for implementing the strategy above.

System Overview

┌─────────────────────────────────────────────────────────────────────────────┐
│                              ngit-grasp                                      │
├─────────────────────────────────────────────────────────────────────────────┤
│                                                                              │
│  ┌─────────────────────────────────────────────────────────────────────┐   │
│  │                         HTTP Router                                  │   │
│  │                                                                      │   │
│  │  Public Endpoints          Authenticated Endpoints                   │   │
│  │  ├─ GET /              ──► Homepage (relay stats)                   │   │
│  │  ├─ GET / (NIP-11)     ──► Relay info + stats                       │   │
│  │  ├─ WS  /              ──► Nostr relay                              │   │
│  │  │                                                                   │   │
│  │  │                         ├─ GET /metrics ──► Prometheus (API key) │   │
│  │  │                         ├─ GET /admin/* ──► Management SPA       │   │
│  │  │                         └─ NIP-86       ──► Management API       │   │
│  │  └─────────────────────────────────────────────────────────────────┘   │
│                                                                              │
│  ┌─────────────────────────────────────────────────────────────────────┐   │
│  │                      Core Components                                 │   │
│  │                                                                      │   │
│  │  ┌──────────────┐  ┌──────────────┐  ┌──────────────────────────┐  │   │
│  │  │ Stats Cache  │  │ Metrics      │  │ Management Service       │  │   │
│  │  │              │  │ Registry     │  │ (NIP-86)                 │  │   │
│  │  │ - repo count │  │ (Prometheus) │  │ - blacklist operations   │  │   │
│  │  │ - event count│  │              │  │ - quota management       │  │   │
│  │  │ - user count │  │ - counters   │  │ - repository deletion    │  │   │
│  │  │ - uptime     │  │ - gauges     │  │ - API key generation     │  │   │
│  │  │              │  │ - histograms │  │                          │  │   │
│  │  │ Throttled    │  │              │  │                          │  │   │
│  │  └──────────────┘  └──────────────┘  └──────────────────────────┘  │   │
│  │         │                  │                      │                 │   │
│  │         └──────────────────┴──────────────────────┘                 │   │
│  │                            │                                        │   │
│  │                    ┌───────▼───────┐                                │   │
│  │                    │ Event Store   │                                │   │
│  │                    │ (LMDB/NDB)    │                                │   │
│  │                    └───────────────┘                                │   │
│  └─────────────────────────────────────────────────────────────────────┘   │
│                                                                              │
└─────────────────────────────────────────────────────────────────────────────┘
         │                    │                         │
         ▼                    ▼                         ▼
    Git Clients         Nostr Clients              Prometheus
                        (+ NIP-86 Admin)            Grafana

Component Details

Stats Cache

Purpose: Provide fast, cached aggregate statistics for public consumption.

Implementation:

  • In-memory cache with throttled updates (no more frequently than 60 seconds)
  • Queries event store for counts on first request after throttle period
  • Calculates uptime from process start time
  • Thread-safe read access for NIP-11 and homepage
  • Extends existing homepage/NIP-11 infrastructure (already implemented)

Data:

struct StatsCache {
    repositories_total: u64,
    events_total: u64,
    users_total: u64,
    uptime_seconds: u64,
    last_updated: Instant,
}

Endpoints served:

  • GET / (homepage) - Existing HTML, extended with stats
  • GET / with Accept: application/nostr+json (NIP-11) - Existing JSON, extended with stats section

Prometheus Metrics Registry

Purpose: Detailed operational metrics for administrators.

Authentication: Bearer token (standard Prometheus authentication approach)

API Key Storage:

  • Multiple API keys supported (for fleet monitoring, different access levels)
  • Keys stored hashed (align with Prometheus API key best practices)
  • Keys can be generated via admin UI or provided declaratively via environment
  • Environment variable accepts comma-separated hashed keys for declarative deployments

Implementation:

  • Existing Prometheus registry (already implemented)
  • Add storage metrics (per-user, per-repo top N)
  • Add quota metrics (usage vs limits)

New metrics to add:

# Storage metrics
ngit_storage_bytes_total{type="git|database"}
ngit_storage_bytes_by_user{pubkey="<hex>"}  # top N only
ngit_storage_bytes_by_repo{repo="<identifier>"}  # top N only

# Quota metrics  
ngit_quota_user_bytes{pubkey="<hex>"}
ngit_quota_repo_bytes{repo="<identifier>"}

Endpoint:

  • GET /metrics - Prometheus format, requires API key

Management Service

Purpose: Handle administrative actions and serve management dashboard.

Components:

  1. Management API - NIP-86 Relay Management API
  2. Management SPA - Single-page application for dashboard
  3. Audit Logging - Standard application logs

Management API (NIP-86)

Protocol: NIP-86 Relay Management API (Nostr-native)

Rationale: Using NIP-86 instead of traditional REST provides:

  • Nostr-native authentication (no separate API keys for management)
  • Standardized protocol benefiting the ecosystem
  • Existing client implementations to reference/reuse

Research needed: Identify existing projects implementing NIP-86 relay management to use as reference for both server implementation and potential SPA reuse.

Capabilities (via NIP-86):

# Blacklist operations
- Add/remove user from blacklist
- Add/remove repository from blacklist
- List blacklisted users/repos

# Quota management
- Set/remove per-user quota
- Set/remove per-repo quota
- List quotas

# Repository management
- Delete repository
- Force garbage collection

# System
- View current configuration
- Generate API keys (for Prometheus access)

API Key Generation: The NIP-86 management interface will be the primary way to generate and manage Prometheus API keys. Keys are hashed before storage.

Management SPA

Purpose: Human-friendly dashboard for administrators.

Technology: Implementation detail TBD. Will research existing NIP-86 management UIs that could be reused or adapted.

Features:

  • Overview dashboard (key metrics, health status)
  • Storage breakdown (charts for per-user, per-repo)
  • User management (search, view, blacklist, quota)
  • Repository management (search, view, delete, blacklist)
  • API key management (generate, revoke Prometheus keys)
  • Configuration viewer (read-only display of current config)

Authentication: NIP-86 (Nostr signature-based)

ConnectionTracker

Purpose: Existing infrastructure that tracks connection patterns and abuse indicators.

Implementation: Already implemented in the current codebase. Tracks:

  • Connection patterns by source
  • Invalid signature attempts
  • Filter violations
  • Potential abuse indicators

Usage in this architecture: ConnectionTracker data will be exposed via the management dashboard (Phase 5) to help administrators identify problematic users and patterns. No new development required - this is existing functionality being surfaced in the UI.

Audit Logging

Purpose: Record all administrative actions for accountability.

Implementation: Standard INFO-level log messages in application logs.

Format:

2026-01-14T17:30:00Z INFO admin_audit action=blacklist_user pubkey=abc123 reason="spam"
2026-01-14T17:31:00Z INFO admin_audit action=delete_repo repo=def456
2026-01-14T17:32:00Z INFO admin_audit action=generate_api_key key_id=xyz789

Retention: Standard log rotation (same as application logs)

User Storage Info (Nostr Events)

Purpose: Allow authenticated users to query their own storage usage.

Implementation:

  1. User authenticates via NIP-42
  2. User sends request event (application-specific kind in 30000-39999 range)
  3. Relay responds with replaceable event containing user's storage data

Event kind selection: Choose random kind in application-specific range, verify not in use via nak req -k <number> -l 1 on popular relays.

Request event:

{
  "kind": 3XXXX,
  "content": "",
  "tags": [
    ["p", "<relay-pubkey>"],
    ["request", "storage-info"]
  ]
}

Response event (replaceable, from relay's pubkey):

{
  "kind": 3XXXX,
  "content": "{\"storage_used_bytes\":1234567,\"quota_bytes\":1073741824,\"repositories\":[...]}",
  "tags": [
    ["p", "<user-pubkey>"],
    ["d", "storage-info"]
  ]
}

Implementation Note: Before implementing, verify that nostr-relay-builder supports this request/response flow easily. Fallback approach if needed: encrypt response content to user's pubkey via NIP-44, ensuring only the requesting user can read their storage info. This is an implementation detail to be resolved during that phase.

Data Flow Diagrams

Public Stats Flow

                    ┌─────────────┐
                    │ Stats Cache │
                    │ (in-memory) │
                    └──────┬──────┘
                           │
              ┌────────────┼────────────┐
              │            │            │
              ▼            ▼            ▼
         ┌────────┐  ┌──────────┐  ┌─────────┐
         │Homepage│  │  NIP-11  │  │ Nostr   │
         │  HTML  │  │   JSON   │  │ Client  │
         └────────┘  └──────────┘  └─────────┘
              │            │            │
              └────────────┴────────────┘
                           │
                           ▼
                    Public Users

Admin Metrics Flow

    ┌─────────────────────────────────────────┐
    │              ngit-grasp                  │
    │                                          │
    │  ┌──────────────────────────────────┐   │
    │  │       Prometheus Registry         │   │
    │  │                                   │   │
    │  │  Connections ◄── WebSocket Handler│   │
    │  │  Git Ops     ◄── Git Handler      │   │
    │  │  Events      ◄── Nostr Relay      │   │
    │  │  Storage     ◄── Storage Scanner  │   │
    │  │  Sync        ◄── Sync Manager     │   │
    │  └──────────────────────────────────┘   │
    │                    │                     │
    │                    ▼                     │
    │            GET /metrics                  │
    │            (API key auth)                │
    └────────────────────┬────────────────────┘
                         │
                         ▼
                  ┌─────────────┐
                  │ Prometheus  │
                  │   Server    │
                  └──────┬──────┘
                         │
                         ▼
                  ┌─────────────┐
                  │   Grafana   │
                  └──────┬──────┘
                         │
                         ▼
                  Administrator

Management Action Flow

    Administrator
         │
         ▼
    ┌─────────────┐
    │ Management  │
    │ SPA / CLI   │
    └──────┬──────┘
           │
           ▼
    NIP-86 Request (signed by admin pubkey)
    {"method": "blacklistuser", "params": ["abc123", "spam"]}
           │
           ▼
    ┌─────────────────────────────────────────┐
    │              ngit-grasp                  │
    │                                          │
    │  ┌──────────────────────────────────┐   │
    │  │    Management Service (NIP-86)    │   │
    │  │                                   │   │
    │  │  1. Verify Nostr signature        │   │
    │  │  2. Check pubkey in admin list    │   │
    │  │  3. Parse request                 │   │
    │  │  4. Update blacklist              │   │
    │  │  5. Log audit entry (INFO)        │   │
    │  │  6. Return success                │   │
    │  └──────────────────────────────────┘   │
    │                    │                     │
    │         ┌──────────┴──────────┐         │
    │         ▼                     ▼         │
    │  ┌─────────────┐      ┌─────────────┐   │
    │  │  Blacklist  │      │ App Logs    │   │
    │  │   (state)   │      │ (audit)     │   │
    │  └─────────────┘      └─────────────┘   │
    └─────────────────────────────────────────┘

Configuration

New environment variables:

# Metrics authentication (required for /metrics endpoint)
# Comma-separated list of hashed API keys for declarative deployments
# Keys should be hashed using bcrypt or similar (align with Prometheus best practices)
# Most users will generate keys via the admin UI instead
NGIT_METRICS_API_KEYS=<hashed_key1>,<hashed_key2>

# Management - uses NIP-86 (Nostr signature auth), no separate API key needed
# Admin pubkeys authorized for NIP-86 management
NGIT_ADMIN_PUBKEYS=<hex_pubkey1>,<hex_pubkey2>

# Deployment mode
NGIT_DECLARATIVE_ONLY=false  # true = config changes via API disabled

# Storage limits
NGIT_MAX_STORAGE_BYTES=10737418240        # 10GB total relay storage
NGIT_DEFAULT_USER_QUOTA_BYTES=1073741824  # 1GB default per user
NGIT_DEFAULT_REPO_QUOTA_BYTES=104857600   # 100MB default per repo

Storage Architecture

Blacklist and quota storage options:

  1. In-memory + config file (for declarative deployments)

    • Blacklists/quotas defined in environment or config
    • Loaded at startup, immutable at runtime
  2. Database tables (for dynamic deployments)

    • Blacklists and quotas stored in LMDB
    • Modified via Management API
    • Persisted across restarts

Recommendation: Support both modes via NGIT_DECLARATIVE_ONLY flag.

Security Architecture

┌─────────────────────────────────────────────────────────────────┐
│                        Authentication Layers                     │
├─────────────────────────────────────────────────────────────────┤
│                                                                  │
│  Public (no auth)           │  Authenticated                    │
│  ─────────────────          │  ─────────────                    │
│  GET /                      │                                   │
│  GET / (NIP-11)             │  GET /metrics    ◄─ API key       │
│  WS / (Nostr)               │  NIP-86 mgmt     ◄─ Nostr sig     │
│  Git clone/fetch            │  GET /admin/*    ◄─ Nostr sig     │
│                             │                                   │
│  NIP-42 Auth (user-specific)│                                   │
│  ───────────────────────────│                                   │
│  User storage info events   │                                   │
│  (can only see own data)    │                                   │
│                             │                                   │
└─────────────────────────────────────────────────────────────────┘

API key management (Prometheus):

  • Multiple keys supported (fleet monitoring, rotation)
  • Keys stored hashed (bcrypt or similar, per Prometheus best practices)
  • Keys primarily generated via admin UI (NIP-86)
  • Declarative option: provide pre-hashed keys via environment variable
  • Bearer token in Authorization header
  • Constant-time hash comparison to prevent timing attacks

NIP-86 Management authentication:

  • Nostr signature-based (no separate API keys)
  • Admin pubkeys configured via NGIT_ADMIN_PUBKEYS
  • Standard Nostr authentication flow

Rate limiting:

  • Rate limit on /metrics endpoint to prevent brute-force key guessing
  • Log failed authentication attempts

Implementation Plan

This section breaks down the implementation into phases with clear dependencies and deliverables. Each phase provides independent value and can be deployed incrementally.

Important: Each phase must be completed with passing tests and committed to version control before moving to the next phase. Code comments and commit messages should describe the functionality being added, not reference phase numbers or this planning document.

Phase 0: Research & Discovery

Goal: Understand existing implementations and verify technical feasibility of key architectural decisions.

⚠️ Manual Review Required: This phase produces research findings that inform architectural decisions in later phases. Results must be reviewed and approved before proceeding to implementation phases.

Parallelizable Tasks:

Track A: NIP-86 Research

  • Survey existing NIP-86 relay management implementations
    • Identify reference implementations (server-side)
    • Document API patterns and authentication flows
    • Evaluate existing management UIs/clients that could be reused or adapted
    • Assess maturity and production readiness
  • Deliverable: Research document with recommendations for NIP-86 implementation approach

Track B: nostr-relay-builder Capabilities

  • Verify nostr-relay-builder supports user storage info request/response flow
    • Test if custom event kinds can trigger relay-generated responses
    • Verify replaceable event generation from relay's keypair
    • Document any limitations or required workarounds
  • Evaluate NIP-44 encryption fallback if direct response not supported
  • Deliverable: Technical feasibility report with implementation approach

⚠️ Manual Review Required: Track B findings determine the implementation approach for Phase 2 Track B. Must be reviewed and approved before starting Phase 2 Track B work.

Effort: 2-3 days (parallel tracks)

Dependencies: None

Related Issues: None (foundational research)

Within-Phase Parallelization: Both research tracks are completely independent and can run concurrently (e.g., two developers or agents working simultaneously).

Deliverable: Commit research findings as documentation in docs/research/ directory. No code changes in this phase.


Phase 1: Foundation - Metrics & Storage

Goal: Establish core metrics infrastructure and storage calculation capabilities.

Parallelizable Tasks:

Track A: Stats Cache Implementation

  • Implement throttled stats cache (60-second minimum refresh)
    • Create StatsCache struct with thread-safe access
    • Query event store for repository/event/user counts
    • Calculate uptime from process start time
    • Add cache invalidation on write operations (optional optimization)
  • Extend existing homepage to display stats
  • Testing: Verify cache throttling, accuracy of counts

Track B: Storage Metrics

  • Implement storage calculation utilities
    • Git repository size calculation (walk .git directory)
    • Database size calculation (LMDB stats)
    • Per-user storage aggregation (sum all user's repos)
    • Per-repository size tracking
  • Add Prometheus metrics:
    • ngit_storage_bytes_total{type="git|database"}
    • ngit_storage_bytes_by_user{pubkey="<hex>"} (top N)
    • ngit_storage_bytes_by_repo{repo="<identifier>"} (top N)
  • Implement background storage scanner (periodic updates, configurable interval)
  • Verify Issue 76fe (Repository Count Metric):
    • Test that ngit_repositories_total metric exists and increments correctly
    • Verify count matches actual repository count in git directory
    • Add regression test to prevent future breakage
  • Testing: Verify accuracy against du command, test top-N selection

Effort: 4-5 days (2-3 days per track, parallel)

Dependencies: None (builds on existing Prometheus infrastructure)

Related Issues:

  • 76fe (Repository Count Metric) - Verify this metric works correctly
  • Partially addresses 7d0b (provides metrics for dashboard to display)

Within-Phase Parallelization: Stats cache and storage metrics are independent features that can be developed in parallel by different developers or agents.

Deliverable: Commit working stats cache and storage metrics with tests. All existing tests must pass.


Phase 2: Public & User Information

Goal: Deliver information to public users (relay selection) and authenticated users (self-service).

Parallelizable Tasks:

Track A: NIP-11 Stats Extension

  • Extend NIP-11 relay info document with stats section
    • Integrate with stats cache from Phase 1
    • Add stats to JSON response
    • Document caching behavior in NIP-11 response headers
  • Update homepage HTML to display stats in human-friendly format
  • Testing: Verify NIP-11 compliance, test caching behavior

Track B: User Storage Info Events

  • Choose and verify event kind (application-specific range 30000-39999)
    • Use nak req -k <number> -l 1 to verify kind not in use on popular relays
    • Document chosen kind in code and configuration
  • Implement request/response flow (based on Phase 0 research)
    • Handle incoming storage info request events (requires NIP-42 auth)
    • Generate replaceable response event with user's storage data
    • Include: storage used, quota, repository list with sizes
  • Implement NIP-44 encryption fallback if needed (based on Phase 0 findings)
  • Testing: End-to-end test with authenticated client, verify auth enforcement

Effort: 3-4 days (1-2 days per track, parallel)

Dependencies:

  • Phase 1 (requires stats cache and storage calculation)
  • Phase 0 Track B (user storage info implementation approach)

Related Issues:

  • Partially addresses 7d0b (public stats for relay selection)
  • Addresses user self-service workflow from strategy

Within-Phase Parallelization: NIP-11 extension and user storage events are completely independent features serving different audiences. Can be developed in parallel.

Deliverable: Commit NIP-11 stats extension and user storage info event handling with tests. All existing tests must pass.


Phase 3: Prometheus Authentication

Goal: Secure administrator access to metrics endpoint.

Tasks:

  • Implement API key authentication for /metrics endpoint
    • Add NGIT_METRICS_API_KEYS environment variable (comma-separated hashed keys)
    • Implement bearer token authentication middleware
    • Use bcrypt or similar for key hashing (align with Prometheus best practices)
    • Implement constant-time hash comparison (prevent timing attacks)
    • Add rate limiting on /metrics endpoint (prevent brute-force)
    • Log failed authentication attempts
  • Testing: Test with valid/invalid keys, verify rate limiting, test timing attack resistance
  • Configuration sync: Update all four config locations per AGENTS.md:
    • src/config.rs - Add NGIT_METRICS_API_KEYS config field
    • docs/reference/configuration.md - Document the new option
    • nix/module.nix - Add NixOS module option
    • .env.example - Add example with comments

Effort: 2-3 days

Dependencies:

  • Phase 1 (Prometheus metrics must exist to protect)

Related Issues: None (infrastructure for later phases)

Within-Phase Parallelization: None (single track)

Deliverable: Commit Prometheus authentication with tests. All existing tests must pass.


Phase 4: NIP-86 Management API

Goal: Implement authenticated management API for administrative operations.

Note: This phase can run unattended IF the rust-nostr library (nostr-sdk) supports NIP-86. If manual implementation is required, this becomes a manual review phase.

Tasks:

  • Implement NIP-86 relay management protocol (based on Phase 0 Track A research)
    • Add NGIT_ADMIN_PUBKEYS environment variable (comma-separated hex pubkeys)
    • Implement Nostr signature verification for management requests
    • Implement management operations:
      • Blacklist user (add to blacklist, enforce in event handler)
      • Blacklist repository (add to blacklist, enforce in git handler)
      • Set user quota (store in database or config)
      • Set repository quota (store in database or config)
      • List blacklists and quotas
    • Implement audit logging (INFO-level structured logs)
  • Add storage for blacklists and quotas:
    • Support both declarative (env/config) and dynamic (database) modes
    • Implement NGIT_DECLARATIVE_ONLY flag behavior
  • Testing: Test each operation, verify auth, test declarative vs dynamic modes
  • Configuration sync: Update all four config locations per AGENTS.md:
    • src/config.rs - Add NGIT_ADMIN_PUBKEYS and NGIT_DECLARATIVE_ONLY config fields
    • docs/reference/configuration.md - Document the new options
    • nix/module.nix - Add NixOS module options
    • .env.example - Add examples with comments

Effort: 4-5 days (IF rust-nostr supports NIP-86), 6-8 days (if manual implementation)

Dependencies:

  • Phase 0 Track A (NIP-86 implementation approach - must be complete)
  • Phase 1 Track B (storage metrics needed for quota operations)

Related Issues:

  • Addresses 2cdc (NIP-86 Relay Management API) - PRIMARY ISSUE
  • Partially addresses 8430 (quota storage, not enforcement yet)

Within-Phase Parallelization: None (single track)

Deliverable: Commit NIP-86 management API with tests. All existing tests must pass.


Phase 5: Management Dashboard

Goal: Provide human-friendly interface for administrators.

⚠️ Manual Review Required: Detailed implementation plan for the dashboard must be reviewed and approved before starting implementation. Plan should include:

  • Selected framework/technology (based on Phase 0 research)
  • UI/UX design approach
  • Feature prioritization
  • Integration approach with NIP-86 API
  • Testing strategy

Tasks:

  • Research and select dashboard framework (based on Phase 0 Track A NIP-86 UI findings)
    • Option 1: Reuse/adapt existing NIP-86 management UI
    • Option 2: Build custom SPA (React/Vue/Svelte)
    • Option 3: Server-rendered with htmx/Alpine.js (simpler)
  • Implement dashboard features:
    • Overview page (health status, key metrics from Prometheus)
    • Storage breakdown (charts for per-user, per-repo from Prometheus)
    • User management (search, view storage, blacklist, set quota)
    • Repository management (search, view size, blacklist, set quota)
    • Configuration viewer (read-only display of current config)
    • API key management (generate Prometheus keys, list, revoke)
    • Abuse metrics display (show existing ConnectionTracker data: connection patterns, invalid signatures, filter violations)
  • Implement NIP-86 client (connect to management API from Phase 4)
  • Extend NIP-86 API with API key operations:
    • Generate new Prometheus API key (return unhashed, store hashed)
    • List API keys (show key IDs, creation dates, not actual keys)
    • Revoke API key (remove from storage)
  • Serve dashboard from /admin/* routes
  • Testing: End-to-end tests for each feature, test on mobile devices

Effort: 6-8 days

Dependencies:

  • Phase 0 Track A (NIP-86 UI research - must be complete and approved)
  • Phase 4 (requires NIP-86 API)
  • Phase 3 (requires Prometheus auth for API key management)
  • Phase 1 (requires metrics to display)

Related Issues:

  • Addresses 7d0b (Management Dashboard) - PRIMARY ISSUE
  • Completes 2cdc (NIP-86 API now has UI)
  • Partially addresses 1f4f (displays abuse metrics from ConnectionTracker)

Within-Phase Parallelization: Limited. Dashboard UI and API key operations could be developed in parallel initially, but must integrate before completion.

Deliverable: Commit working management dashboard with tests. All existing tests must pass.


Phase 6: Quota Enforcement & Repository Deletion

Goal: Enforce storage limits and enable repository cleanup.

Tasks:

  1. Repository Deletion

    • Implement delete repository operation (called from NIP-86 API)
    • Delete git repository from filesystem
    • Delete repository events from event store
    • Delete repository metadata
    • Add confirmation requirement (two-step process)
    • Add audit logging
    • Testing: Test deletion, verify cleanup, test rollback on errors
  2. Quota Enforcement

    • Implement quota checking in git push handler
      • Check per-repository quota before accepting push
      • Check per-user total quota before accepting push
      • Check relay-wide storage limit
    • Implement quota checking in event handler (for repository creation)
    • Return clear error messages when quota exceeded
      • Include current usage, limit, and suggested actions
    • Add metrics for quota violations
    • Testing: Test each quota type, test error messages, test edge cases
  3. Configuration

    • Add NGIT_MAX_STORAGE_BYTES, NGIT_DEFAULT_USER_QUOTA_BYTES, NGIT_DEFAULT_REPO_QUOTA_BYTES environment variables
    • Configuration sync: Update all four config locations per AGENTS.md:
      • src/config.rs - Add quota-related config fields
      • docs/reference/configuration.md - Document the new options
      • nix/module.nix - Add NixOS module options
      • .env.example - Add examples with comments

Effort: 3-4 days

Dependencies:

  • Phase 4 (requires quota storage from NIP-86 API)
  • Phase 5 (can run unattended once dashboard is approved and complete)
  • Phase 1 Track B (requires storage calculation)
  • Issue b905 (Deletion Request Support - NIP-09 implementation)

Related Issues:

  • Addresses 8430 (Storage Limits and Quota Management) - PRIMARY ISSUE
  • Depends on b905 (Deletion Request Support) for NIP-09 event deletion

Within-Phase Parallelization: None. Repository deletion and quota enforcement are tightly coupled to the same code paths (git handler, event handler). Must be implemented sequentially.

Deliverable: Commit quota enforcement and repository deletion with tests. All existing tests must pass.


Summary: Phase Dependencies & Parallelization

Phase 0: Research (2-3 days) ⚠️ MANUAL REVIEW REQUIRED
├─ Track A: NIP-86 Research ─────────────┐
└─ Track B: nostr-relay-builder Research ─┴─► (within-phase parallel)
   Dependencies: None
   Deliverable: Research documentation
   Manual Review: Track B before Phase 2 Track B

Phase 1: Foundation (4-5 days)
├─ Track A: Stats Cache ─────────────────┐
└─ Track B: Storage Metrics ─────────────┴─► (within-phase parallel)
   Dependencies: Phase 0 complete
   Deliverable: Working code + tests

Phase 2: Public & User Info (3-4 days)
├─ Track A: NIP-11 Extension ────────────┐
└─ Track B: User Storage Events ─────────┴─► (within-phase parallel)
   Dependencies: Phase 1 complete, Phase 0 Track B (approved)
   Deliverable: Working code + tests

Phase 3: Prometheus Auth (2-3 days)
└─ Single track
   Dependencies: Phase 1 complete
   Deliverable: Working code + tests

Phase 4: NIP-86 Management API (4-8 days)
└─ Single track (can run unattended IF rust-nostr supports NIP-86)
   Dependencies: Phase 0 Track A, Phase 1 Track B
   Deliverable: Working code + tests
   Note: 4-5 days if supported, 6-8 days if manual implementation

Phase 5: Management Dashboard (6-8 days) ⚠️ MANUAL REVIEW REQUIRED
└─ Single track (limited internal parallelization)
   Dependencies: Phase 0 Track A (approved), Phase 3, Phase 4, Phase 1
   Deliverable: Working code + tests
   Manual Review: Detailed plan before implementation

Phase 6: Quota Enforcement (3-4 days)
└─ Sequential tasks within phase (can run unattended once Phase 5 approved)
   Dependencies: Phase 4, Phase 5 (approved), Phase 1 Track B, Issue b905
   Deliverable: Working code + tests

Total Sequential Time: ~24-32 days (5-6 weeks)

Parallelization Opportunities:

  • Within-phase: Phases 0, 1, and 2 have independent tracks that can run simultaneously
  • Cross-phase: Limited. Most phases have strict dependencies on prior phases
  • Realistic speedup: With 2 developers working parallel tracks: ~20-26 days (4-5 weeks)

Recommended Approach:

  • Single developer: Follow phases sequentially (0→1→2→3→4→5→6), executing parallel tracks within each phase
  • Team of 2: One developer per track in phases with parallelizable tracks; alternate on single-track phases
  • Larger teams: Not recommended - phases have strict sequential dependencies

Manual Review Points:

  1. Phase 0 Track B results (before Phase 2 Track B)
  2. Phase 5 detailed plan (before Phase 5 implementation)

Issue Mapping

Issues Addressed by Implementation:

Issue Phase(s) Status
76fe (Repository Count Metric) Phase 1 Track B Verify existing implementation
7d0b (Management Dashboard) Phase 2 Track A, Phase 5 Primary issue for Phase 5
2cdc (NIP-86 Relay Management API) Phase 0 Track A, Phase 4, Phase 5 Primary issue for Phase 4
8430 (Storage Limits & Quota Management) Phase 4, Phase 6 Primary issue for Phase 6
1f4f (Poor Naughty List Identification) Phase 5 Partially (display existing ConnectionTracker metrics in dashboard)
b905 (Deletion Request Support) Phase 6 Dependency (NIP-09 must be implemented before Phase 6)

Issues NOT Addressed:

Issue Reason
d6ee (Defensive Relay Features) Removed from this plan (was Phase 7 Track A) - should be separate issue/plan

Issues Superseded:

None. Most existing issues remain relevant and are incorporated into this plan.

Effort Estimates

By Phase:

  • Phase 0: 2-3 days (research, requires manual review)
  • Phase 1: 4-5 days (foundation)
  • Phase 2: 3-4 days (public/user info)
  • Phase 3: 2-3 days (Prometheus auth)
  • Phase 4: 4-8 days (NIP-86 API - depends on rust-nostr support)
  • Phase 5: 6-8 days (dashboard, requires manual review of plan)
  • Phase 6: 3-4 days (quota enforcement, can run unattended once Phase 5 approved)

Total: 24-35 days (5-7 weeks) sequential, 20-28 days (4-6 weeks) with within-phase parallelization

By Issue:

  • 76fe: 0 days (verification only)
  • 7d0b: 9-12 days (Phase 2 Track A + Phase 5)
  • 2cdc: 8-13 days (Phase 0 Track A + Phase 4 + Phase 5 API key mgmt)
  • 8430: 7-12 days (Phase 4 quota storage + Phase 6)
  • 1f4f: Included in Phase 5 (dashboard displays existing ConnectionTracker data)
  • b905: Dependency (must be complete before Phase 6)
  • d6ee: NOT ADDRESSED (removed from this plan)

Risk Mitigation

Research Risks (Phase 0):

  • Risk: NIP-86 implementations may be immature or poorly documented
  • Mitigation: Budget extra time for Phase 3 if research reveals complexity
  • Fallback: Implement simpler HTTP API first, migrate to NIP-86 later

Technical Risks:

  • Risk: nostr-relay-builder may not support relay-generated response events
  • Mitigation: Phase 0 Track B validates this early
  • Fallback: Use NIP-44 encryption to user's pubkey (still secure, slightly more complex)

Integration Risks:

  • Risk: Phase 4 SPA integration with NIP-86 API may reveal API gaps
  • Mitigation: Phase 3 includes comprehensive testing of all operations
  • Adjustment: Add missing operations in Phase 4 if discovered during integration

Performance Risks:

  • Risk: Storage calculation may be expensive for large relays
  • Mitigation: Background scanner with configurable interval, top-N only for metrics
  • Monitoring: Add metrics for scanner duration, alert if too slow

Success Criteria

Phase 0 Complete:

  • NIP-86 research document published with implementation recommendations
  • nostr-relay-builder feasibility report with chosen approach for user storage events
  • Research committed to docs/research/ directory
  • Manual review completed and Phase 0 Track B approach approved

Phase 1 Complete:

  • Stats cache implemented and tested (60-second throttle)
  • Storage metrics available in Prometheus (ngit_storage_bytes_*)
  • Background storage scanner running without performance impact
  • All tests pass, changes committed

Phase 2 Complete:

  • NIP-11 relay info includes stats section
  • Homepage displays stats in human-friendly format
  • Authenticated users can query their storage usage via Nostr events
  • All tests pass, changes committed

Phase 3 Complete:

  • /metrics endpoint requires API key authentication
  • Rate limiting prevents brute-force attacks on metrics endpoint
  • Failed authentication attempts are logged
  • All tests pass, changes committed

Phase 4 Complete:

  • NIP-86 management API functional for all operations
  • Blacklists and quotas can be managed via NIP-86
  • Audit logging records all admin actions
  • Declarative and dynamic modes both work correctly
  • All tests pass, changes committed

Phase 5 Complete:

  • Detailed dashboard plan reviewed and approved before implementation
  • Management dashboard accessible and functional
  • Dashboard displays all metrics with visualizations
  • Dashboard shows existing abuse metrics from ConnectionTracker
  • Administrators can perform all management operations via UI
  • API keys can be generated and managed via UI
  • All tests pass, changes committed

Phase 6 Complete:

  • Issue b905 (NIP-09 Deletion Request Support) implemented and merged
  • Repositories can be deleted via management API/UI
  • Repository deletion integrates with NIP-09 event deletion
  • Storage quotas enforced on git push and repository creation
  • Clear error messages guide users when quotas exceeded (include current usage, limit, and suggested actions)
  • Metrics track quota violations
  • All four config locations updated for new quota-related environment variables
  • All tests pass, changes committed

Overall Success:

  • Administrator can monitor relay health without SSH access
  • Administrator can view abuse metrics (from ConnectionTracker) in dashboard
  • Administrator can take corrective action through dashboard
  • Storage quotas protect relay from resource exhaustion
  • All workflows from strategy document are supported
  • All three audiences (public, users, admins) have appropriate access to information

Decisions

  1. Event kinds for user storage info: Use application-specific kinds (30000-39999 range for replaceable). Select a random kind and verify it's not in use on popular relays via nak req -k <number> -l 1 relay.example.com.

  2. NIP-11 stats caching interval: 60 seconds.

  3. Management dashboard technology: SPA (Single Page Application).

  4. Audit log storage: Log file, following standard log retention practices.