Files
ngit-grasp/ec1f-administrator-observability-and-management-strategy.md

20 KiB

Administrator Observability and Management Strategy

ID: ec1f

Problem

Administrators need a cohesive strategy for monitoring their GRASP relay and taking action when problems arise. Currently we have:

  1. Prometheus metrics - comprehensive time-series data (connections, git ops, events, sync)
  2. Grafana dashboard - pre-built but untested and potentially out of date
  3. Fragmented issues - each addressing pieces (dashboard UI, management API, storage limits, rate limiting, abuse detection) without a unified vision

Key gaps:

  • No clear path from "I see a problem" to "I fix the problem"
  • No strategy for configuration management (declarative vs. dynamic)
  • Overlapping concerns across multiple issues without coordination
  • Untested/outdated Prometheus setup and Grafana dashboard

Vision

Administrators should be able to:

  1. Observe - See what's happening (metrics, dashboards, alerts)
  2. Understand - Identify what needs attention (storage warnings, abusive users, large repos)
  3. Act - Take corrective action (blacklist, delete, adjust quotas)
  4. Enforce - Automatic protection (rate limits, storage quotas, abuse detection)

All while respecting deployment models:

  • Declarative deployments (NixOS, Docker) - config from env/CLI only, dashboard read-only
  • Dynamic deployments (manual) - optional GUI-based config changes that persist

Configuration Model

New Configuration Options

# Allow management dashboard to modify configuration
# Default: false (read-only dashboard)
NGIT_ALLOW_DYNAMIC_CONFIG=false

# Enforce declarative-only configuration (no local config file)
# Default: false (allow local config file if ALLOW_DYNAMIC_CONFIG=true)
NGIT_DECLARATIVE_ONLY=false

Configuration Hierarchy

1. CLI args / env vars / .env     (declarative, immutable)
         ↓
2. Local config file               (dynamic, if ALLOW_DYNAMIC_CONFIG=true)
         ↓
3. Runtime state                   (ephemeral)

Behavior Matrix

DECLARATIVE_ONLY ALLOW_DYNAMIC_CONFIG Behavior
true * (ignored) Pure declarative - no local config file, dashboard read-only
false false Declarative - dashboard read-only, no local config file
false true Dynamic - dashboard can modify config, persists to local file

NixOS/Docker deployments: Set NGIT_DECLARATIVE_ONLY=true to ensure purity

Manual deployments: Can enable NGIT_ALLOW_DYNAMIC_CONFIG=true for GUI-based management

Phased Roadmap

Phase 1: Observability (See what's happening)

Goal: Administrators can see current state and trends

Tasks:

  • Test and update Prometheus setup documentation
  • Test and update Grafana dashboard (verify all metrics work)
  • Add missing storage metrics:
    • Total disk usage (git + database)
    • Available disk space percentage
    • Per-user storage usage (top N users)
    • Per-repository storage usage (largest repos)
  • Add "attention needed" indicators:
    • Storage approaching limits (80%, 90%, 95%)
    • Flagged abusive IPs (from ConnectionTracker)
    • Large repositories (potential blacklist candidates)
    • High error rates (git operations, event rejections)

Depends on: Nothing

Related issues:

  • 76fe (Repository Count Metric) - already implemented, verify it works
  • 7d0b (Management Dashboard) - Phase 2 will build on this

Phase 2: Insights (Understand what needs action)

Goal: Administrators can identify problems and understand their scope

Tasks:

  • Implement Management Dashboard UI (issue 7d0b)
    • Read-only view of all metrics
    • Storage breakdown visualization
    • Highlight problematic pubkeys/repos
    • Recent activity feed
    • System health indicators
  • Dashboard respects configuration flags:
    • Show "read-only mode" banner when ALLOW_DYNAMIC_CONFIG=false
    • Hide action buttons in read-only mode
    • Display current config source (env, file, CLI)

Depends on: Phase 1 (metrics must exist to display)

Related issues:

  • 7d0b (Management Dashboard) - PRIMARY ISSUE FOR THIS PHASE

Phase 3: Actions (Take action)

Goal: Administrators can take corrective action through authenticated API

Decision point: NIP-86 vs. simpler HTTP API?

Option A: NIP-86 Relay Management API (issue 2cdc)

  • Standardized protocol, Nostr-native authentication
  • More complex to implement
  • Benefits entire Nostr ecosystem if we contribute patterns

Option B: Simple HTTP API

  • Faster to implement
  • Basic auth or API key
  • Less standardization

Recommendation: Start with Option B (simple HTTP API), migrate to NIP-86 later if needed

Tasks:

  • Implement authenticated management API:
    • Blacklist user (npub) from all operations
    • Blacklist repository (prevent pushes)
    • Delete repository (with confirmation)
    • Set per-user storage quota
    • Set per-repo storage quota
    • Force garbage collection on repository
    • View audit log of admin actions
  • Integrate with dashboard UI (if ALLOW_DYNAMIC_CONFIG=true):
    • Action buttons for blacklist/delete/quota
    • Confirmation dialogs for destructive actions
    • Real-time feedback on action results
  • Implement local config file persistence:
    • Define config file format (TOML? JSON?)
    • Load local config on startup (after env/CLI)
    • Save config changes from dashboard
    • Validate config on load

Depends on: Phase 2 (dashboard UI to trigger actions)

Related issues:

  • 2cdc (NIP-86 Relay Management API) - DEFERRED in favor of simpler HTTP API first
  • 7d0b (Management Dashboard) - UI integration

Phase 4: Enforcement (Automatic protection)

Goal: Relay automatically protects itself from abuse and resource exhaustion

Tasks:

  • Implement storage limits and quota management (issue 8430):
    • Per-repository size limits
    • Per-user total storage limits
    • Relay-wide storage limits
    • Graceful degradation (read-only mode when approaching limits)
    • Clear error messages when quotas exceeded
  • Implement defensive relay features (issue d6ee):
    • Phase 1 only: Explicit rate limits + max connections (APPROVED)
    • Defer per-IP enforcement until abuse detected in production
  • Implement naughty list identification (issue 1f4f):
    • Detect invalid signatures, filter violations, DoS patterns
    • Automatic temporary bans for repeated violations
    • Integration with blacklist system from Phase 3

Depends on: Phase 3 (manual blacklist/quota system must exist for overrides)

Related issues:

  • 8430 (Storage Limits and Quota Management) - PRIMARY ISSUE FOR THIS PHASE
  • d6ee (Defensive Relay Features) - Phase 1 only (config-based protection)
  • 1f4f (Poor Naughty List Identification) - Abuse detection

Issue Overlap Analysis

Primary Overlaps

Storage Management:

  • 7d0b (Management Dashboard) - Display storage metrics
  • 8430 (Storage Limits) - Enforce storage quotas
  • 2cdc (NIP-86 Management API) - Set quotas via API
  • Resolution: 8430 owns enforcement, 7d0b owns display, 2cdc owns API (deferred)

Abuse Detection & Response:

  • d6ee (Defensive Relay Features) - Rate limiting, connection limits
  • 1f4f (Poor Naughty List) - Detect malicious behavior
  • 2cdc (NIP-86 Management API) - Blacklist users/repos
  • Resolution: d6ee owns prevention, 1f4f owns detection, 2cdc owns manual response (deferred)

Configuration & Monitoring:

  • 7d0b (Management Dashboard) - Display config and metrics
  • 2cdc (NIP-86 Management API) - Modify config via API
  • 76fe (Repository Count Metric) - Specific metric implementation
  • Resolution: 7d0b owns display, 2cdc owns modification (deferred), 76fe is complete

Coordination Requirements

Before starting work on any overlapping issue:

  1. Check this strategy issue (ec1f) for current phase
  2. Verify dependencies are complete
  3. Coordinate scope with related issues
  4. Update this issue with progress

Implementation Sequence

Phase 1: Observability
├── Update Prometheus/Grafana setup (1-2 days)
├── Add storage metrics (2-3 days)
└── Add attention indicators (1-2 days)
    Total: ~1 week

Phase 2: Insights
├── Management Dashboard UI (3-5 days)
└── Configuration flag integration (1 day)
    Total: ~1 week
    Depends on: Phase 1

Phase 3: Actions
├── Simple HTTP management API (3-4 days)
├── Dashboard action integration (2-3 days)
└── Local config file persistence (2-3 days)
    Total: ~1.5 weeks
    Depends on: Phase 2

Phase 4: Enforcement
├── Storage limits (issue 8430) (1-2 weeks)
├── Defensive features Phase 1 (issue d6ee) (2-3 hours)
└── Naughty list detection (issue 1f4f) (1 week)
    Total: ~2-3 weeks
    Depends on: Phase 3

Total estimated time: 5-7 weeks for complete administrator experience

Success Criteria

Phase 1 Complete

  • Prometheus metrics endpoint tested and documented
  • Grafana dashboard loads without errors
  • All metrics display current values
  • Storage metrics available (disk usage, per-user, per-repo)
  • Attention indicators highlight problems

Phase 2 Complete

  • Management dashboard accessible via HTTP endpoint
  • Dashboard displays all metrics with visualizations
  • Dashboard respects ALLOW_DYNAMIC_CONFIG flag
  • Mobile-friendly responsive design
  • No significant performance impact on relay

Phase 3 Complete

  • Authenticated management API functional
  • Can blacklist users and repositories
  • Can set storage quotas
  • Can delete repositories with confirmation
  • Audit log tracks all admin actions
  • Local config file persists changes (when enabled)

Phase 4 Complete

  • Storage quotas enforced automatically
  • Graceful degradation when approaching limits
  • Rate limiting prevents DoS attacks
  • Abuse detection flags problematic users
  • Clear error messages guide users

Overall Success

  • Administrator can monitor relay health without SSH access
  • Administrator can take corrective action through dashboard
  • Declarative deployments remain pure (NixOS, Docker)
  • Dynamic deployments support GUI-based management
  • No manual intervention needed for common abuse scenarios

Open Questions

  1. Config file format: TOML, JSON, or YAML for local config file?

    • Recommendation: TOML (matches Rust ecosystem, human-friendly)
  2. Authentication for management API: Basic auth, API key, or Nostr signature?

    • Recommendation: Start with API key (simple), migrate to Nostr signatures later
  3. Dashboard framework: Server-rendered HTML or SPA?

    • Recommendation: Server-rendered (htmx or Alpine.js) for simplicity
  4. NIP-86 timeline: When to implement full NIP-86 compliance?

    • Recommendation: After Phase 3 proves the simple API works
  5. Prometheus vs. custom metrics: Should we add non-Prometheus metrics for dashboard?

    • Recommendation: Prometheus only (single source of truth)

Directly related (must coordinate):

  • 7d0b (Management Dashboard) - Phase 2 primary issue
  • 2cdc (NIP-86 Relay Management API) - Deferred to post-Phase 3
  • 8430 (Storage Limits and Quota Management) - Phase 4 primary issue
  • d6ee (Defensive Relay Features) - Phase 4 (Phase 1 only)
  • 1f4f (Poor Naughty List Identification) - Phase 4

Indirectly related:

  • 76fe (Repository Count Metric) - Already implemented, verify in Phase 1
  • cc36 (Database Backend Evaluation) - May affect storage metrics
  • b905 (Deletion Request Support) - May integrate with delete repository action

Key Decisions

2026-01-14 Strategy Review

Prometheus Metrics:

  • Private, API key required - Always authenticated via NGIT_METRICS_API_KEY
  • Supports fleet monitoring (multiple instances reporting to central Prometheus)
  • Contains sensitive admin data (per-user storage, per-repo sizes, etc.)

Public Relay Stats:

  • Delivered via NIP-11 - Extend relay info document with stats section
  • Cache expensive metrics rather than computing on each request
  • No separate HTTP endpoint needed

User Storage Info:

  • Delivered via replaceable Nostr event - User-specific storage usage/limits
  • Requires NIP-42 authentication to query own data
  • Preferred approach for all sensitive user-specific data via Nostr

Metrics Separation:

  • Public info (NIP-11): Safe aggregate stats for relay selection
  • Private metrics (Prometheus): Detailed operational data for admins
  • User-specific (Nostr events): Personal storage usage for authenticated users

Next Steps

The current phased roadmap above is SUPERSEDED pending proper planning:

  1. Strategy Document - docs/explanation/metrics-and-administration.md

    • Thoughtful document outlining administrator workflow
    • What metrics support each workflow step
    • How different audiences (public, users, admins) access information
  2. Target Architecture - Define technical architecture based on strategy

    • Component interactions
    • Data flows
    • API designs
  3. Implementation Plan - Only after strategy and architecture are complete

    • Phased delivery
    • Dependencies
    • Estimates

Progress

2026-01-14 [Session 16:30]

  • Created strategy issue
  • Analyzed existing issues for overlap
  • Defined configuration model (DECLARATIVE_ONLY, ALLOW_DYNAMIC_CONFIG)
  • Established 4-phase roadmap
  • Identified decision points (NIP-86 vs. simple API)

2026-01-14 [Session ~17:00]

  • Security review of public metrics exposure
  • Decision: Prometheus metrics require API key authentication
  • Decision: Public relay stats via NIP-11 (cached, not Prometheus)
  • Decision: User storage info via replaceable Nostr events (NIP-42 auth)
  • Recognized need for proper planning sequence: strategy → architecture → implementation
  • Created worktree: ec1f-administrator-observability-and-management-strategy
  • Wrote strategy document: docs/explanation/metrics-and-administration.md
    • Defined three audiences (public, authenticated users, administrators)
    • Documented five administrator workflows
    • Specified information architecture for each audience
    • Outlined configuration model and security considerations
  • Added target architecture to strategy document:
    • System overview with component diagram
    • Stats Cache (60s refresh for NIP-11/homepage)
    • Prometheus metrics with API key auth
    • Management Service (REST API + SPA dashboard)
    • User storage info via Nostr events (application-specific kind)
    • Data flow diagrams
    • Configuration and security architecture
  • Resolved open questions: event kinds (app-specific), caching (60s), dashboard (SPA), audit (log file)
  • Refined architecture based on feedback:
    • Stats cache: throttled (not periodic), reuse existing homepage/NIP-11
    • Prometheus: multiple hashed API keys, generated via admin UI
    • Management API: NIP-86 (not REST), research existing implementations
    • Audit logging: standard INFO messages in app logs
    • User storage info: verify nostr-relay-builder support, NIP-44 fallback
  • Next: Write implementation plan

2026-01-14 [Session ~18:00]

  • Completed implementation plan in strategy document
  • Broke work into 7 phases (Phase 0-7) with clear dependencies
  • Identified within-phase parallelization opportunities:
    • Phase 0: 2 parallel research tracks (NIP-86, nostr-relay-builder)
    • Phase 1: 2 parallel tracks (stats cache, storage metrics)
    • Phase 2: 2 parallel tracks (NIP-11 extension, user storage events)
    • Phase 3: Single track (Prometheus auth)
    • Phase 4: Single track (NIP-86 API)
    • Phase 5: Single track (dashboard, limited internal parallelization)
    • Phase 6: Sequential tasks (quota enforcement)
    • Phase 7: 2 parallel tracks (rate limiting, abuse detection)
  • Mapped all existing issues to phases:
    • 76fe: Phase 1 Track B (verification)
    • 7d0b: Phase 2 Track A + Phase 5 (9-12 days)
    • 2cdc: Phase 0 Track A + Phase 4 + Phase 5 (8-10 days)
    • 8430: Phase 4 + Phase 6 (9-11 days)
    • d6ee: Phase 7 Track A (2-3 days)
    • 1f4f: Phase 7 Track B (2-3 days)
  • Effort estimates:
    • Sequential: 31-39 days (6-8 weeks)
    • With within-phase parallelization: 25-30 days (5-6 weeks)
  • Added risk mitigation strategies for research and technical risks
  • Defined success criteria for each phase and overall
  • Each phase must commit working code with passing tests
  • Code/commits must not reference phase numbers
  • Ready for implementation to begin with Phase 0 research

2026-01-14 [Session ~18:30]

  • Corrected implementation plan based on feedback
  • Fixed incorrect parallelization claims:
    • Phase 3 and 4 cannot run in parallel (Phase 4 depends on Phase 0 Track A research)
    • Split original Phase 3 into two sequential phases (Prometheus auth, then NIP-86 API)
    • Clarified "within-phase" vs "cross-phase" parallelization
  • Added explicit deliverable requirements:
    • Each phase must commit changes with all tests passing
    • No phase references in code comments or commit messages
    • Phase 0 commits research docs, all other phases commit working code
  • Updated phase count from 6 to 7 phases
  • Corrected dependency graph to show true sequential requirements
  • More realistic effort estimates accounting for sequential dependencies

2026-01-15 [Session 09:54]

  • Completed Phase 0 Track A: NIP-86 implementation research
  • Used subagent to research nostr-sdk 0.43 support for NIP-86
  • Key findings:
    • ✅ NIP-98 HTTP auth fully supported in nostr-sdk (can leverage for authentication)
    • ❌ NIP-86 JSON-RPC protocol NOT provided (must implement from scratch)
    • Recommended approach: Implement NIP-86 using existing NIP-98 auth
    • Estimated effort: 6-11 days for complete implementation
  • Deliverable: docs/research/nip-86-implementation-options.md committed
  • Next: Phase 0 Track B (nostr-relay-builder user storage events) - deferred for now
  • Ready to proceed with Phase 1 implementation

2026-01-15 [Session 11:05]

  • Critical architectural decision: Identified gap in design - NIP-86 is for actions, not queries
  • Researched how to return sensitive data to authenticated users (admins and regular users)
  • Launched 3 parallel research subagents:
    1. NIP-86 query response patterns in spec and rust-nostr
    2. Pyramid relay admin patterns (real-world implementation)
    3. Feasibility of pyramid-style approach with rust-nostr
  • Key findings from research:
    • ✅ NIP-86 spec uses HTTP POST/response (NOT Nostr events for query results)
    • ✅ Pyramid uses ephemeral events to specific WebSocket connections for queries
    • ❌ nostr-relay-builder does NOT support connection-specific messaging
    • ❌ Pyramid approach requires forking nostr-relay-builder or custom WebSocket handler
  • Decision: Use HTTP-based NIP-86 for all queries and actions
    • Queries: HTTP GET/POST with NIP-98 auth → JSON response
    • Actions: HTTP POST with NIP-98 auth → JSON response
    • Both admin and user queries use same pattern
    • Simpler, standard, no framework modifications needed
  • Deliverables committed:
    • docs/research/nip-86-query-responses.md - NIP-86 spec analysis
    • docs/research/pyramid-admin-patterns.md - Pyramid implementation study
    • docs/research/pyramid-style-queries-design.md - Feasibility analysis for rust-nostr
  • Next: Update strategy document to reflect HTTP-based approach for queries
  • Phase 0 research complete, ready to finalize architecture and proceed to Phase 1

2026-01-15 [Session 11:38]

  • Final Phase 0 research: Explored forking rust-nostr to support pyramid-style queries
  • Comprehensive analysis of what it would take to modify nostr-relay-builder
  • Key findings:
    • Implementation effort: 5-7 days (vs 1-2 days for HTTP NIP-86)
    • Ongoing maintenance: 18-29 hours/month (vs 1-2 hours/month for HTTP)
    • rust-nostr is extremely active: 305 commits in 15 days (~20/day)
    • Breaking changes in ~50% of releases
    • Pyramid uses khatru framework (Go), NOT nostr-relay-builder
    • Upstream acceptance likelihood: Low (10-20%)
  • Final decision confirmed: HTTP-based NIP-86 is the correct approach
    • No fork needed
    • Standards-compliant
    • Minimal maintenance burden
    • Achieves same functional goals
  • Deliverable committed:
    • docs/research/rust-nostr-fork-analysis.md - Complete fork analysis (1,316 lines)
  • Phase 0 research COMPLETE ✅
  • All 4 research documents committed and ready for review
  • Ready to update strategy document and begin Phase 1 implementation