Files
ngit-grasp/ec1f-administrator-observability-and-management-strategy.md
T

12 KiB

Administrator Observability and Management Strategy

ID: ec1f

Problem

Administrators need a cohesive strategy for monitoring their GRASP relay and taking action when problems arise. Currently we have:

  1. Prometheus metrics - comprehensive time-series data (connections, git ops, events, sync)
  2. Grafana dashboard - pre-built but untested and potentially out of date
  3. Fragmented issues - each addressing pieces (dashboard UI, management API, storage limits, rate limiting, abuse detection) without a unified vision

Key gaps:

  • No clear path from "I see a problem" to "I fix the problem"
  • No strategy for configuration management (declarative vs. dynamic)
  • Overlapping concerns across multiple issues without coordination
  • Untested/outdated Prometheus setup and Grafana dashboard

Vision

Administrators should be able to:

  1. Observe - See what's happening (metrics, dashboards, alerts)
  2. Understand - Identify what needs attention (storage warnings, abusive users, large repos)
  3. Act - Take corrective action (blacklist, delete, adjust quotas)
  4. Enforce - Automatic protection (rate limits, storage quotas, abuse detection)

All while respecting deployment models:

  • Declarative deployments (NixOS, Docker) - config from env/CLI only, dashboard read-only
  • Dynamic deployments (manual) - optional GUI-based config changes that persist

Configuration Model

New Configuration Options

# Allow management dashboard to modify configuration
# Default: false (read-only dashboard)
NGIT_ALLOW_DYNAMIC_CONFIG=false

# Enforce declarative-only configuration (no local config file)
# Default: false (allow local config file if ALLOW_DYNAMIC_CONFIG=true)
NGIT_DECLARATIVE_ONLY=false

Configuration Hierarchy

1. CLI args / env vars / .env     (declarative, immutable)
         ↓
2. Local config file               (dynamic, if ALLOW_DYNAMIC_CONFIG=true)
         ↓
3. Runtime state                   (ephemeral)

Behavior Matrix

DECLARATIVE_ONLY ALLOW_DYNAMIC_CONFIG Behavior
true * (ignored) Pure declarative - no local config file, dashboard read-only
false false Declarative - dashboard read-only, no local config file
false true Dynamic - dashboard can modify config, persists to local file

NixOS/Docker deployments: Set NGIT_DECLARATIVE_ONLY=true to ensure purity

Manual deployments: Can enable NGIT_ALLOW_DYNAMIC_CONFIG=true for GUI-based management

Phased Roadmap

Phase 1: Observability (See what's happening)

Goal: Administrators can see current state and trends

Tasks:

  • Test and update Prometheus setup documentation
  • Test and update Grafana dashboard (verify all metrics work)
  • Add missing storage metrics:
    • Total disk usage (git + database)
    • Available disk space percentage
    • Per-user storage usage (top N users)
    • Per-repository storage usage (largest repos)
  • Add "attention needed" indicators:
    • Storage approaching limits (80%, 90%, 95%)
    • Flagged abusive IPs (from ConnectionTracker)
    • Large repositories (potential blacklist candidates)
    • High error rates (git operations, event rejections)

Depends on: Nothing

Related issues:

  • 76fe (Repository Count Metric) - already implemented, verify it works
  • 7d0b (Management Dashboard) - Phase 2 will build on this

Phase 2: Insights (Understand what needs action)

Goal: Administrators can identify problems and understand their scope

Tasks:

  • Implement Management Dashboard UI (issue 7d0b)
    • Read-only view of all metrics
    • Storage breakdown visualization
    • Highlight problematic pubkeys/repos
    • Recent activity feed
    • System health indicators
  • Dashboard respects configuration flags:
    • Show "read-only mode" banner when ALLOW_DYNAMIC_CONFIG=false
    • Hide action buttons in read-only mode
    • Display current config source (env, file, CLI)

Depends on: Phase 1 (metrics must exist to display)

Related issues:

  • 7d0b (Management Dashboard) - PRIMARY ISSUE FOR THIS PHASE

Phase 3: Actions (Take action)

Goal: Administrators can take corrective action through authenticated API

Decision point: NIP-86 vs. simpler HTTP API?

Option A: NIP-86 Relay Management API (issue 2cdc)

  • Standardized protocol, Nostr-native authentication
  • More complex to implement
  • Benefits entire Nostr ecosystem if we contribute patterns

Option B: Simple HTTP API

  • Faster to implement
  • Basic auth or API key
  • Less standardization

Recommendation: Start with Option B (simple HTTP API), migrate to NIP-86 later if needed

Tasks:

  • Implement authenticated management API:
    • Blacklist user (npub) from all operations
    • Blacklist repository (prevent pushes)
    • Delete repository (with confirmation)
    • Set per-user storage quota
    • Set per-repo storage quota
    • Force garbage collection on repository
    • View audit log of admin actions
  • Integrate with dashboard UI (if ALLOW_DYNAMIC_CONFIG=true):
    • Action buttons for blacklist/delete/quota
    • Confirmation dialogs for destructive actions
    • Real-time feedback on action results
  • Implement local config file persistence:
    • Define config file format (TOML? JSON?)
    • Load local config on startup (after env/CLI)
    • Save config changes from dashboard
    • Validate config on load

Depends on: Phase 2 (dashboard UI to trigger actions)

Related issues:

  • 2cdc (NIP-86 Relay Management API) - DEFERRED in favor of simpler HTTP API first
  • 7d0b (Management Dashboard) - UI integration

Phase 4: Enforcement (Automatic protection)

Goal: Relay automatically protects itself from abuse and resource exhaustion

Tasks:

  • Implement storage limits and quota management (issue 8430):
    • Per-repository size limits
    • Per-user total storage limits
    • Relay-wide storage limits
    • Graceful degradation (read-only mode when approaching limits)
    • Clear error messages when quotas exceeded
  • Implement defensive relay features (issue d6ee):
    • Phase 1 only: Explicit rate limits + max connections (APPROVED)
    • Defer per-IP enforcement until abuse detected in production
  • Implement naughty list identification (issue 1f4f):
    • Detect invalid signatures, filter violations, DoS patterns
    • Automatic temporary bans for repeated violations
    • Integration with blacklist system from Phase 3

Depends on: Phase 3 (manual blacklist/quota system must exist for overrides)

Related issues:

  • 8430 (Storage Limits and Quota Management) - PRIMARY ISSUE FOR THIS PHASE
  • d6ee (Defensive Relay Features) - Phase 1 only (config-based protection)
  • 1f4f (Poor Naughty List Identification) - Abuse detection

Issue Overlap Analysis

Primary Overlaps

Storage Management:

  • 7d0b (Management Dashboard) - Display storage metrics
  • 8430 (Storage Limits) - Enforce storage quotas
  • 2cdc (NIP-86 Management API) - Set quotas via API
  • Resolution: 8430 owns enforcement, 7d0b owns display, 2cdc owns API (deferred)

Abuse Detection & Response:

  • d6ee (Defensive Relay Features) - Rate limiting, connection limits
  • 1f4f (Poor Naughty List) - Detect malicious behavior
  • 2cdc (NIP-86 Management API) - Blacklist users/repos
  • Resolution: d6ee owns prevention, 1f4f owns detection, 2cdc owns manual response (deferred)

Configuration & Monitoring:

  • 7d0b (Management Dashboard) - Display config and metrics
  • 2cdc (NIP-86 Management API) - Modify config via API
  • 76fe (Repository Count Metric) - Specific metric implementation
  • Resolution: 7d0b owns display, 2cdc owns modification (deferred), 76fe is complete

Coordination Requirements

Before starting work on any overlapping issue:

  1. Check this strategy issue (ec1f) for current phase
  2. Verify dependencies are complete
  3. Coordinate scope with related issues
  4. Update this issue with progress

Implementation Sequence

Phase 1: Observability
├── Update Prometheus/Grafana setup (1-2 days)
├── Add storage metrics (2-3 days)
└── Add attention indicators (1-2 days)
    Total: ~1 week

Phase 2: Insights
├── Management Dashboard UI (3-5 days)
└── Configuration flag integration (1 day)
    Total: ~1 week
    Depends on: Phase 1

Phase 3: Actions
├── Simple HTTP management API (3-4 days)
├── Dashboard action integration (2-3 days)
└── Local config file persistence (2-3 days)
    Total: ~1.5 weeks
    Depends on: Phase 2

Phase 4: Enforcement
├── Storage limits (issue 8430) (1-2 weeks)
├── Defensive features Phase 1 (issue d6ee) (2-3 hours)
└── Naughty list detection (issue 1f4f) (1 week)
    Total: ~2-3 weeks
    Depends on: Phase 3

Total estimated time: 5-7 weeks for complete administrator experience

Success Criteria

Phase 1 Complete

  • Prometheus metrics endpoint tested and documented
  • Grafana dashboard loads without errors
  • All metrics display current values
  • Storage metrics available (disk usage, per-user, per-repo)
  • Attention indicators highlight problems

Phase 2 Complete

  • Management dashboard accessible via HTTP endpoint
  • Dashboard displays all metrics with visualizations
  • Dashboard respects ALLOW_DYNAMIC_CONFIG flag
  • Mobile-friendly responsive design
  • No significant performance impact on relay

Phase 3 Complete

  • Authenticated management API functional
  • Can blacklist users and repositories
  • Can set storage quotas
  • Can delete repositories with confirmation
  • Audit log tracks all admin actions
  • Local config file persists changes (when enabled)

Phase 4 Complete

  • Storage quotas enforced automatically
  • Graceful degradation when approaching limits
  • Rate limiting prevents DoS attacks
  • Abuse detection flags problematic users
  • Clear error messages guide users

Overall Success

  • Administrator can monitor relay health without SSH access
  • Administrator can take corrective action through dashboard
  • Declarative deployments remain pure (NixOS, Docker)
  • Dynamic deployments support GUI-based management
  • No manual intervention needed for common abuse scenarios

Open Questions

  1. Config file format: TOML, JSON, or YAML for local config file?

    • Recommendation: TOML (matches Rust ecosystem, human-friendly)
  2. Authentication for management API: Basic auth, API key, or Nostr signature?

    • Recommendation: Start with API key (simple), migrate to Nostr signatures later
  3. Dashboard framework: Server-rendered HTML or SPA?

    • Recommendation: Server-rendered (htmx or Alpine.js) for simplicity
  4. NIP-86 timeline: When to implement full NIP-86 compliance?

    • Recommendation: After Phase 3 proves the simple API works
  5. Prometheus vs. custom metrics: Should we add non-Prometheus metrics for dashboard?

    • Recommendation: Prometheus only (single source of truth)

Directly related (must coordinate):

  • 7d0b (Management Dashboard) - Phase 2 primary issue
  • 2cdc (NIP-86 Relay Management API) - Deferred to post-Phase 3
  • 8430 (Storage Limits and Quota Management) - Phase 4 primary issue
  • d6ee (Defensive Relay Features) - Phase 4 (Phase 1 only)
  • 1f4f (Poor Naughty List Identification) - Phase 4

Indirectly related:

  • 76fe (Repository Count Metric) - Already implemented, verify in Phase 1
  • cc36 (Database Backend Evaluation) - May affect storage metrics
  • b905 (Deletion Request Support) - May integrate with delete repository action

Progress

2026-01-14 [Session 16:30]

  • Created strategy issue
  • Analyzed existing issues for overlap
  • Defined configuration model (DECLARATIVE_ONLY, ALLOW_DYNAMIC_CONFIG)
  • Established 4-phase roadmap
  • Identified decision points (NIP-86 vs. simple API)
  • Next: Update overlapping issues with reference to this strategy