20 KiB
Administrator Observability and Management Strategy
ID: ec1f
Problem
Administrators need a cohesive strategy for monitoring their GRASP relay and taking action when problems arise. Currently we have:
- Prometheus metrics - comprehensive time-series data (connections, git ops, events, sync)
- Grafana dashboard - pre-built but untested and potentially out of date
- Fragmented issues - each addressing pieces (dashboard UI, management API, storage limits, rate limiting, abuse detection) without a unified vision
Key gaps:
- No clear path from "I see a problem" to "I fix the problem"
- No strategy for configuration management (declarative vs. dynamic)
- Overlapping concerns across multiple issues without coordination
- Untested/outdated Prometheus setup and Grafana dashboard
Vision
Administrators should be able to:
- Observe - See what's happening (metrics, dashboards, alerts)
- Understand - Identify what needs attention (storage warnings, abusive users, large repos)
- Act - Take corrective action (blacklist, delete, adjust quotas)
- Enforce - Automatic protection (rate limits, storage quotas, abuse detection)
All while respecting deployment models:
- Declarative deployments (NixOS, Docker) - config from env/CLI only, dashboard read-only
- Dynamic deployments (manual) - optional GUI-based config changes that persist
Configuration Model
New Configuration Options
# Allow management dashboard to modify configuration
# Default: false (read-only dashboard)
NGIT_ALLOW_DYNAMIC_CONFIG=false
# Enforce declarative-only configuration (no local config file)
# Default: false (allow local config file if ALLOW_DYNAMIC_CONFIG=true)
NGIT_DECLARATIVE_ONLY=false
Configuration Hierarchy
1. CLI args / env vars / .env (declarative, immutable)
↓
2. Local config file (dynamic, if ALLOW_DYNAMIC_CONFIG=true)
↓
3. Runtime state (ephemeral)
Behavior Matrix
| DECLARATIVE_ONLY | ALLOW_DYNAMIC_CONFIG | Behavior |
|---|---|---|
| true | * (ignored) | Pure declarative - no local config file, dashboard read-only |
| false | false | Declarative - dashboard read-only, no local config file |
| false | true | Dynamic - dashboard can modify config, persists to local file |
NixOS/Docker deployments: Set NGIT_DECLARATIVE_ONLY=true to ensure purity
Manual deployments: Can enable NGIT_ALLOW_DYNAMIC_CONFIG=true for GUI-based management
Phased Roadmap
Phase 1: Observability (See what's happening)
Goal: Administrators can see current state and trends
Tasks:
- Test and update Prometheus setup documentation
- Test and update Grafana dashboard (verify all metrics work)
- Add missing storage metrics:
- Total disk usage (git + database)
- Available disk space percentage
- Per-user storage usage (top N users)
- Per-repository storage usage (largest repos)
- Add "attention needed" indicators:
- Storage approaching limits (80%, 90%, 95%)
- Flagged abusive IPs (from ConnectionTracker)
- Large repositories (potential blacklist candidates)
- High error rates (git operations, event rejections)
Depends on: Nothing
Related issues:
- 76fe (Repository Count Metric) - already implemented, verify it works
- 7d0b (Management Dashboard) - Phase 2 will build on this
Phase 2: Insights (Understand what needs action)
Goal: Administrators can identify problems and understand their scope
Tasks:
- Implement Management Dashboard UI (issue 7d0b)
- Read-only view of all metrics
- Storage breakdown visualization
- Highlight problematic pubkeys/repos
- Recent activity feed
- System health indicators
- Dashboard respects configuration flags:
- Show "read-only mode" banner when ALLOW_DYNAMIC_CONFIG=false
- Hide action buttons in read-only mode
- Display current config source (env, file, CLI)
Depends on: Phase 1 (metrics must exist to display)
Related issues:
- 7d0b (Management Dashboard) - PRIMARY ISSUE FOR THIS PHASE
Phase 3: Actions (Take action)
Goal: Administrators can take corrective action through authenticated API
Decision point: NIP-86 vs. simpler HTTP API?
Option A: NIP-86 Relay Management API (issue 2cdc)
- Standardized protocol, Nostr-native authentication
- More complex to implement
- Benefits entire Nostr ecosystem if we contribute patterns
Option B: Simple HTTP API
- Faster to implement
- Basic auth or API key
- Less standardization
Recommendation: Start with Option B (simple HTTP API), migrate to NIP-86 later if needed
Tasks:
- Implement authenticated management API:
- Blacklist user (npub) from all operations
- Blacklist repository (prevent pushes)
- Delete repository (with confirmation)
- Set per-user storage quota
- Set per-repo storage quota
- Force garbage collection on repository
- View audit log of admin actions
- Integrate with dashboard UI (if ALLOW_DYNAMIC_CONFIG=true):
- Action buttons for blacklist/delete/quota
- Confirmation dialogs for destructive actions
- Real-time feedback on action results
- Implement local config file persistence:
- Define config file format (TOML? JSON?)
- Load local config on startup (after env/CLI)
- Save config changes from dashboard
- Validate config on load
Depends on: Phase 2 (dashboard UI to trigger actions)
Related issues:
- 2cdc (NIP-86 Relay Management API) - DEFERRED in favor of simpler HTTP API first
- 7d0b (Management Dashboard) - UI integration
Phase 4: Enforcement (Automatic protection)
Goal: Relay automatically protects itself from abuse and resource exhaustion
Tasks:
- Implement storage limits and quota management (issue 8430):
- Per-repository size limits
- Per-user total storage limits
- Relay-wide storage limits
- Graceful degradation (read-only mode when approaching limits)
- Clear error messages when quotas exceeded
- Implement defensive relay features (issue d6ee):
- Phase 1 only: Explicit rate limits + max connections (APPROVED)
- Defer per-IP enforcement until abuse detected in production
- Implement naughty list identification (issue 1f4f):
- Detect invalid signatures, filter violations, DoS patterns
- Automatic temporary bans for repeated violations
- Integration with blacklist system from Phase 3
Depends on: Phase 3 (manual blacklist/quota system must exist for overrides)
Related issues:
- 8430 (Storage Limits and Quota Management) - PRIMARY ISSUE FOR THIS PHASE
- d6ee (Defensive Relay Features) - Phase 1 only (config-based protection)
- 1f4f (Poor Naughty List Identification) - Abuse detection
Issue Overlap Analysis
Primary Overlaps
Storage Management:
- 7d0b (Management Dashboard) - Display storage metrics
- 8430 (Storage Limits) - Enforce storage quotas
- 2cdc (NIP-86 Management API) - Set quotas via API
- Resolution: 8430 owns enforcement, 7d0b owns display, 2cdc owns API (deferred)
Abuse Detection & Response:
- d6ee (Defensive Relay Features) - Rate limiting, connection limits
- 1f4f (Poor Naughty List) - Detect malicious behavior
- 2cdc (NIP-86 Management API) - Blacklist users/repos
- Resolution: d6ee owns prevention, 1f4f owns detection, 2cdc owns manual response (deferred)
Configuration & Monitoring:
- 7d0b (Management Dashboard) - Display config and metrics
- 2cdc (NIP-86 Management API) - Modify config via API
- 76fe (Repository Count Metric) - Specific metric implementation
- Resolution: 7d0b owns display, 2cdc owns modification (deferred), 76fe is complete
Coordination Requirements
Before starting work on any overlapping issue:
- Check this strategy issue (ec1f) for current phase
- Verify dependencies are complete
- Coordinate scope with related issues
- Update this issue with progress
Implementation Sequence
Phase 1: Observability
├── Update Prometheus/Grafana setup (1-2 days)
├── Add storage metrics (2-3 days)
└── Add attention indicators (1-2 days)
Total: ~1 week
Phase 2: Insights
├── Management Dashboard UI (3-5 days)
└── Configuration flag integration (1 day)
Total: ~1 week
Depends on: Phase 1
Phase 3: Actions
├── Simple HTTP management API (3-4 days)
├── Dashboard action integration (2-3 days)
└── Local config file persistence (2-3 days)
Total: ~1.5 weeks
Depends on: Phase 2
Phase 4: Enforcement
├── Storage limits (issue 8430) (1-2 weeks)
├── Defensive features Phase 1 (issue d6ee) (2-3 hours)
└── Naughty list detection (issue 1f4f) (1 week)
Total: ~2-3 weeks
Depends on: Phase 3
Total estimated time: 5-7 weeks for complete administrator experience
Success Criteria
Phase 1 Complete
- Prometheus metrics endpoint tested and documented
- Grafana dashboard loads without errors
- All metrics display current values
- Storage metrics available (disk usage, per-user, per-repo)
- Attention indicators highlight problems
Phase 2 Complete
- Management dashboard accessible via HTTP endpoint
- Dashboard displays all metrics with visualizations
- Dashboard respects ALLOW_DYNAMIC_CONFIG flag
- Mobile-friendly responsive design
- No significant performance impact on relay
Phase 3 Complete
- Authenticated management API functional
- Can blacklist users and repositories
- Can set storage quotas
- Can delete repositories with confirmation
- Audit log tracks all admin actions
- Local config file persists changes (when enabled)
Phase 4 Complete
- Storage quotas enforced automatically
- Graceful degradation when approaching limits
- Rate limiting prevents DoS attacks
- Abuse detection flags problematic users
- Clear error messages guide users
Overall Success
- Administrator can monitor relay health without SSH access
- Administrator can take corrective action through dashboard
- Declarative deployments remain pure (NixOS, Docker)
- Dynamic deployments support GUI-based management
- No manual intervention needed for common abuse scenarios
Open Questions
-
Config file format: TOML, JSON, or YAML for local config file?
- Recommendation: TOML (matches Rust ecosystem, human-friendly)
-
Authentication for management API: Basic auth, API key, or Nostr signature?
- Recommendation: Start with API key (simple), migrate to Nostr signatures later
-
Dashboard framework: Server-rendered HTML or SPA?
- Recommendation: Server-rendered (htmx or Alpine.js) for simplicity
-
NIP-86 timeline: When to implement full NIP-86 compliance?
- Recommendation: After Phase 3 proves the simple API works
-
Prometheus vs. custom metrics: Should we add non-Prometheus metrics for dashboard?
- Recommendation: Prometheus only (single source of truth)
Related Issues
Directly related (must coordinate):
- 7d0b (Management Dashboard) - Phase 2 primary issue
- 2cdc (NIP-86 Relay Management API) - Deferred to post-Phase 3
- 8430 (Storage Limits and Quota Management) - Phase 4 primary issue
- d6ee (Defensive Relay Features) - Phase 4 (Phase 1 only)
- 1f4f (Poor Naughty List Identification) - Phase 4
Indirectly related:
- 76fe (Repository Count Metric) - Already implemented, verify in Phase 1
- cc36 (Database Backend Evaluation) - May affect storage metrics
- b905 (Deletion Request Support) - May integrate with delete repository action
Key Decisions
2026-01-14 Strategy Review
Prometheus Metrics:
- Private, API key required - Always authenticated via
NGIT_METRICS_API_KEY - Supports fleet monitoring (multiple instances reporting to central Prometheus)
- Contains sensitive admin data (per-user storage, per-repo sizes, etc.)
Public Relay Stats:
- Delivered via NIP-11 - Extend relay info document with stats section
- Cache expensive metrics rather than computing on each request
- No separate HTTP endpoint needed
User Storage Info:
- Delivered via replaceable Nostr event - User-specific storage usage/limits
- Requires NIP-42 authentication to query own data
- Preferred approach for all sensitive user-specific data via Nostr
Metrics Separation:
- Public info (NIP-11): Safe aggregate stats for relay selection
- Private metrics (Prometheus): Detailed operational data for admins
- User-specific (Nostr events): Personal storage usage for authenticated users
Next Steps
The current phased roadmap above is SUPERSEDED pending proper planning:
-
Strategy Document -
docs/explanation/metrics-and-administration.md- Thoughtful document outlining administrator workflow
- What metrics support each workflow step
- How different audiences (public, users, admins) access information
-
Target Architecture - Define technical architecture based on strategy
- Component interactions
- Data flows
- API designs
-
Implementation Plan - Only after strategy and architecture are complete
- Phased delivery
- Dependencies
- Estimates
Progress
2026-01-14 [Session 16:30]
- Created strategy issue
- Analyzed existing issues for overlap
- Defined configuration model (DECLARATIVE_ONLY, ALLOW_DYNAMIC_CONFIG)
- Established 4-phase roadmap
- Identified decision points (NIP-86 vs. simple API)
2026-01-14 [Session ~17:00]
- Security review of public metrics exposure
- Decision: Prometheus metrics require API key authentication
- Decision: Public relay stats via NIP-11 (cached, not Prometheus)
- Decision: User storage info via replaceable Nostr events (NIP-42 auth)
- Recognized need for proper planning sequence: strategy → architecture → implementation
- Created worktree:
ec1f-administrator-observability-and-management-strategy - Wrote strategy document:
docs/explanation/metrics-and-administration.md- Defined three audiences (public, authenticated users, administrators)
- Documented five administrator workflows
- Specified information architecture for each audience
- Outlined configuration model and security considerations
- Added target architecture to strategy document:
- System overview with component diagram
- Stats Cache (60s refresh for NIP-11/homepage)
- Prometheus metrics with API key auth
- Management Service (REST API + SPA dashboard)
- User storage info via Nostr events (application-specific kind)
- Data flow diagrams
- Configuration and security architecture
- Resolved open questions: event kinds (app-specific), caching (60s), dashboard (SPA), audit (log file)
- Refined architecture based on feedback:
- Stats cache: throttled (not periodic), reuse existing homepage/NIP-11
- Prometheus: multiple hashed API keys, generated via admin UI
- Management API: NIP-86 (not REST), research existing implementations
- Audit logging: standard INFO messages in app logs
- User storage info: verify nostr-relay-builder support, NIP-44 fallback
- Next: Write implementation plan
2026-01-14 [Session ~18:00]
- Completed implementation plan in strategy document
- Broke work into 7 phases (Phase 0-7) with clear dependencies
- Identified within-phase parallelization opportunities:
- Phase 0: 2 parallel research tracks (NIP-86, nostr-relay-builder)
- Phase 1: 2 parallel tracks (stats cache, storage metrics)
- Phase 2: 2 parallel tracks (NIP-11 extension, user storage events)
- Phase 3: Single track (Prometheus auth)
- Phase 4: Single track (NIP-86 API)
- Phase 5: Single track (dashboard, limited internal parallelization)
- Phase 6: Sequential tasks (quota enforcement)
- Phase 7: 2 parallel tracks (rate limiting, abuse detection)
- Mapped all existing issues to phases:
- 76fe: Phase 1 Track B (verification)
- 7d0b: Phase 2 Track A + Phase 5 (9-12 days)
- 2cdc: Phase 0 Track A + Phase 4 + Phase 5 (8-10 days)
- 8430: Phase 4 + Phase 6 (9-11 days)
- d6ee: Phase 7 Track A (2-3 days)
- 1f4f: Phase 7 Track B (2-3 days)
- Effort estimates:
- Sequential: 31-39 days (6-8 weeks)
- With within-phase parallelization: 25-30 days (5-6 weeks)
- Added risk mitigation strategies for research and technical risks
- Defined success criteria for each phase and overall
- Each phase must commit working code with passing tests
- Code/commits must not reference phase numbers
- Ready for implementation to begin with Phase 0 research
2026-01-14 [Session ~18:30]
- Corrected implementation plan based on feedback
- Fixed incorrect parallelization claims:
- Phase 3 and 4 cannot run in parallel (Phase 4 depends on Phase 0 Track A research)
- Split original Phase 3 into two sequential phases (Prometheus auth, then NIP-86 API)
- Clarified "within-phase" vs "cross-phase" parallelization
- Added explicit deliverable requirements:
- Each phase must commit changes with all tests passing
- No phase references in code comments or commit messages
- Phase 0 commits research docs, all other phases commit working code
- Updated phase count from 6 to 7 phases
- Corrected dependency graph to show true sequential requirements
- More realistic effort estimates accounting for sequential dependencies
2026-01-15 [Session 09:54]
- Completed Phase 0 Track A: NIP-86 implementation research
- Used subagent to research nostr-sdk 0.43 support for NIP-86
- Key findings:
- ✅ NIP-98 HTTP auth fully supported in nostr-sdk (can leverage for authentication)
- ❌ NIP-86 JSON-RPC protocol NOT provided (must implement from scratch)
- Recommended approach: Implement NIP-86 using existing NIP-98 auth
- Estimated effort: 6-11 days for complete implementation
- Deliverable:
docs/research/nip-86-implementation-options.mdcommitted - Next: Phase 0 Track B (nostr-relay-builder user storage events) - deferred for now
- Ready to proceed with Phase 1 implementation
2026-01-15 [Session 11:05]
- Critical architectural decision: Identified gap in design - NIP-86 is for actions, not queries
- Researched how to return sensitive data to authenticated users (admins and regular users)
- Launched 3 parallel research subagents:
- NIP-86 query response patterns in spec and rust-nostr
- Pyramid relay admin patterns (real-world implementation)
- Feasibility of pyramid-style approach with rust-nostr
- Key findings from research:
- ✅ NIP-86 spec uses HTTP POST/response (NOT Nostr events for query results)
- ✅ Pyramid uses ephemeral events to specific WebSocket connections for queries
- ❌ nostr-relay-builder does NOT support connection-specific messaging
- ❌ Pyramid approach requires forking nostr-relay-builder or custom WebSocket handler
- Decision: Use HTTP-based NIP-86 for all queries and actions
- Queries: HTTP GET/POST with NIP-98 auth → JSON response
- Actions: HTTP POST with NIP-98 auth → JSON response
- Both admin and user queries use same pattern
- Simpler, standard, no framework modifications needed
- Deliverables committed:
docs/research/nip-86-query-responses.md- NIP-86 spec analysisdocs/research/pyramid-admin-patterns.md- Pyramid implementation studydocs/research/pyramid-style-queries-design.md- Feasibility analysis for rust-nostr
- Next: Update strategy document to reflect HTTP-based approach for queries
- Phase 0 research complete, ready to finalize architecture and proceed to Phase 1
2026-01-15 [Session 11:38]
- Final Phase 0 research: Explored forking rust-nostr to support pyramid-style queries
- Comprehensive analysis of what it would take to modify nostr-relay-builder
- Key findings:
- Implementation effort: 5-7 days (vs 1-2 days for HTTP NIP-86)
- Ongoing maintenance: 18-29 hours/month (vs 1-2 hours/month for HTTP)
- rust-nostr is extremely active: 305 commits in 15 days (~20/day)
- Breaking changes in ~50% of releases
- Pyramid uses khatru framework (Go), NOT nostr-relay-builder
- Upstream acceptance likelihood: Low (10-20%)
- Final decision confirmed: HTTP-based NIP-86 is the correct approach
- No fork needed
- Standards-compliant
- Minimal maintenance burden
- Achieves same functional goals
- Deliverable committed:
docs/research/rust-nostr-fork-analysis.md- Complete fork analysis (1,316 lines)
- Phase 0 research COMPLETE ✅
- All 4 research documents committed and ready for review
- Ready to update strategy document and begin Phase 1 implementation