12 KiB
Administrator Observability and Management Strategy
ID: ec1f
Problem
Administrators need a cohesive strategy for monitoring their GRASP relay and taking action when problems arise. Currently we have:
- Prometheus metrics - comprehensive time-series data (connections, git ops, events, sync)
- Grafana dashboard - pre-built but untested and potentially out of date
- Fragmented issues - each addressing pieces (dashboard UI, management API, storage limits, rate limiting, abuse detection) without a unified vision
Key gaps:
- No clear path from "I see a problem" to "I fix the problem"
- No strategy for configuration management (declarative vs. dynamic)
- Overlapping concerns across multiple issues without coordination
- Untested/outdated Prometheus setup and Grafana dashboard
Vision
Administrators should be able to:
- Observe - See what's happening (metrics, dashboards, alerts)
- Understand - Identify what needs attention (storage warnings, abusive users, large repos)
- Act - Take corrective action (blacklist, delete, adjust quotas)
- Enforce - Automatic protection (rate limits, storage quotas, abuse detection)
All while respecting deployment models:
- Declarative deployments (NixOS, Docker) - config from env/CLI only, dashboard read-only
- Dynamic deployments (manual) - optional GUI-based config changes that persist
Configuration Model
New Configuration Options
# Allow management dashboard to modify configuration
# Default: false (read-only dashboard)
NGIT_ALLOW_DYNAMIC_CONFIG=false
# Enforce declarative-only configuration (no local config file)
# Default: false (allow local config file if ALLOW_DYNAMIC_CONFIG=true)
NGIT_DECLARATIVE_ONLY=false
Configuration Hierarchy
1. CLI args / env vars / .env (declarative, immutable)
↓
2. Local config file (dynamic, if ALLOW_DYNAMIC_CONFIG=true)
↓
3. Runtime state (ephemeral)
Behavior Matrix
| DECLARATIVE_ONLY | ALLOW_DYNAMIC_CONFIG | Behavior |
|---|---|---|
| true | * (ignored) | Pure declarative - no local config file, dashboard read-only |
| false | false | Declarative - dashboard read-only, no local config file |
| false | true | Dynamic - dashboard can modify config, persists to local file |
NixOS/Docker deployments: Set NGIT_DECLARATIVE_ONLY=true to ensure purity
Manual deployments: Can enable NGIT_ALLOW_DYNAMIC_CONFIG=true for GUI-based management
Phased Roadmap
Phase 1: Observability (See what's happening)
Goal: Administrators can see current state and trends
Tasks:
- Test and update Prometheus setup documentation
- Test and update Grafana dashboard (verify all metrics work)
- Add missing storage metrics:
- Total disk usage (git + database)
- Available disk space percentage
- Per-user storage usage (top N users)
- Per-repository storage usage (largest repos)
- Add "attention needed" indicators:
- Storage approaching limits (80%, 90%, 95%)
- Flagged abusive IPs (from ConnectionTracker)
- Large repositories (potential blacklist candidates)
- High error rates (git operations, event rejections)
Depends on: Nothing
Related issues:
- 76fe (Repository Count Metric) - already implemented, verify it works
- 7d0b (Management Dashboard) - Phase 2 will build on this
Phase 2: Insights (Understand what needs action)
Goal: Administrators can identify problems and understand their scope
Tasks:
- Implement Management Dashboard UI (issue 7d0b)
- Read-only view of all metrics
- Storage breakdown visualization
- Highlight problematic pubkeys/repos
- Recent activity feed
- System health indicators
- Dashboard respects configuration flags:
- Show "read-only mode" banner when ALLOW_DYNAMIC_CONFIG=false
- Hide action buttons in read-only mode
- Display current config source (env, file, CLI)
Depends on: Phase 1 (metrics must exist to display)
Related issues:
- 7d0b (Management Dashboard) - PRIMARY ISSUE FOR THIS PHASE
Phase 3: Actions (Take action)
Goal: Administrators can take corrective action through authenticated API
Decision point: NIP-86 vs. simpler HTTP API?
Option A: NIP-86 Relay Management API (issue 2cdc)
- Standardized protocol, Nostr-native authentication
- More complex to implement
- Benefits entire Nostr ecosystem if we contribute patterns
Option B: Simple HTTP API
- Faster to implement
- Basic auth or API key
- Less standardization
Recommendation: Start with Option B (simple HTTP API), migrate to NIP-86 later if needed
Tasks:
- Implement authenticated management API:
- Blacklist user (npub) from all operations
- Blacklist repository (prevent pushes)
- Delete repository (with confirmation)
- Set per-user storage quota
- Set per-repo storage quota
- Force garbage collection on repository
- View audit log of admin actions
- Integrate with dashboard UI (if ALLOW_DYNAMIC_CONFIG=true):
- Action buttons for blacklist/delete/quota
- Confirmation dialogs for destructive actions
- Real-time feedback on action results
- Implement local config file persistence:
- Define config file format (TOML? JSON?)
- Load local config on startup (after env/CLI)
- Save config changes from dashboard
- Validate config on load
Depends on: Phase 2 (dashboard UI to trigger actions)
Related issues:
- 2cdc (NIP-86 Relay Management API) - DEFERRED in favor of simpler HTTP API first
- 7d0b (Management Dashboard) - UI integration
Phase 4: Enforcement (Automatic protection)
Goal: Relay automatically protects itself from abuse and resource exhaustion
Tasks:
- Implement storage limits and quota management (issue 8430):
- Per-repository size limits
- Per-user total storage limits
- Relay-wide storage limits
- Graceful degradation (read-only mode when approaching limits)
- Clear error messages when quotas exceeded
- Implement defensive relay features (issue d6ee):
- Phase 1 only: Explicit rate limits + max connections (APPROVED)
- Defer per-IP enforcement until abuse detected in production
- Implement naughty list identification (issue 1f4f):
- Detect invalid signatures, filter violations, DoS patterns
- Automatic temporary bans for repeated violations
- Integration with blacklist system from Phase 3
Depends on: Phase 3 (manual blacklist/quota system must exist for overrides)
Related issues:
- 8430 (Storage Limits and Quota Management) - PRIMARY ISSUE FOR THIS PHASE
- d6ee (Defensive Relay Features) - Phase 1 only (config-based protection)
- 1f4f (Poor Naughty List Identification) - Abuse detection
Issue Overlap Analysis
Primary Overlaps
Storage Management:
- 7d0b (Management Dashboard) - Display storage metrics
- 8430 (Storage Limits) - Enforce storage quotas
- 2cdc (NIP-86 Management API) - Set quotas via API
- Resolution: 8430 owns enforcement, 7d0b owns display, 2cdc owns API (deferred)
Abuse Detection & Response:
- d6ee (Defensive Relay Features) - Rate limiting, connection limits
- 1f4f (Poor Naughty List) - Detect malicious behavior
- 2cdc (NIP-86 Management API) - Blacklist users/repos
- Resolution: d6ee owns prevention, 1f4f owns detection, 2cdc owns manual response (deferred)
Configuration & Monitoring:
- 7d0b (Management Dashboard) - Display config and metrics
- 2cdc (NIP-86 Management API) - Modify config via API
- 76fe (Repository Count Metric) - Specific metric implementation
- Resolution: 7d0b owns display, 2cdc owns modification (deferred), 76fe is complete
Coordination Requirements
Before starting work on any overlapping issue:
- Check this strategy issue (ec1f) for current phase
- Verify dependencies are complete
- Coordinate scope with related issues
- Update this issue with progress
Implementation Sequence
Phase 1: Observability
├── Update Prometheus/Grafana setup (1-2 days)
├── Add storage metrics (2-3 days)
└── Add attention indicators (1-2 days)
Total: ~1 week
Phase 2: Insights
├── Management Dashboard UI (3-5 days)
└── Configuration flag integration (1 day)
Total: ~1 week
Depends on: Phase 1
Phase 3: Actions
├── Simple HTTP management API (3-4 days)
├── Dashboard action integration (2-3 days)
└── Local config file persistence (2-3 days)
Total: ~1.5 weeks
Depends on: Phase 2
Phase 4: Enforcement
├── Storage limits (issue 8430) (1-2 weeks)
├── Defensive features Phase 1 (issue d6ee) (2-3 hours)
└── Naughty list detection (issue 1f4f) (1 week)
Total: ~2-3 weeks
Depends on: Phase 3
Total estimated time: 5-7 weeks for complete administrator experience
Success Criteria
Phase 1 Complete
- Prometheus metrics endpoint tested and documented
- Grafana dashboard loads without errors
- All metrics display current values
- Storage metrics available (disk usage, per-user, per-repo)
- Attention indicators highlight problems
Phase 2 Complete
- Management dashboard accessible via HTTP endpoint
- Dashboard displays all metrics with visualizations
- Dashboard respects ALLOW_DYNAMIC_CONFIG flag
- Mobile-friendly responsive design
- No significant performance impact on relay
Phase 3 Complete
- Authenticated management API functional
- Can blacklist users and repositories
- Can set storage quotas
- Can delete repositories with confirmation
- Audit log tracks all admin actions
- Local config file persists changes (when enabled)
Phase 4 Complete
- Storage quotas enforced automatically
- Graceful degradation when approaching limits
- Rate limiting prevents DoS attacks
- Abuse detection flags problematic users
- Clear error messages guide users
Overall Success
- Administrator can monitor relay health without SSH access
- Administrator can take corrective action through dashboard
- Declarative deployments remain pure (NixOS, Docker)
- Dynamic deployments support GUI-based management
- No manual intervention needed for common abuse scenarios
Open Questions
-
Config file format: TOML, JSON, or YAML for local config file?
- Recommendation: TOML (matches Rust ecosystem, human-friendly)
-
Authentication for management API: Basic auth, API key, or Nostr signature?
- Recommendation: Start with API key (simple), migrate to Nostr signatures later
-
Dashboard framework: Server-rendered HTML or SPA?
- Recommendation: Server-rendered (htmx or Alpine.js) for simplicity
-
NIP-86 timeline: When to implement full NIP-86 compliance?
- Recommendation: After Phase 3 proves the simple API works
-
Prometheus vs. custom metrics: Should we add non-Prometheus metrics for dashboard?
- Recommendation: Prometheus only (single source of truth)
Related Issues
Directly related (must coordinate):
- 7d0b (Management Dashboard) - Phase 2 primary issue
- 2cdc (NIP-86 Relay Management API) - Deferred to post-Phase 3
- 8430 (Storage Limits and Quota Management) - Phase 4 primary issue
- d6ee (Defensive Relay Features) - Phase 4 (Phase 1 only)
- 1f4f (Poor Naughty List Identification) - Phase 4
Indirectly related:
- 76fe (Repository Count Metric) - Already implemented, verify in Phase 1
- cc36 (Database Backend Evaluation) - May affect storage metrics
- b905 (Deletion Request Support) - May integrate with delete repository action
Progress
2026-01-14 [Session 16:30]
- Created strategy issue
- Analyzed existing issues for overlap
- Defined configuration model (DECLARATIVE_ONLY, ALLOW_DYNAMIC_CONFIG)
- Established 4-phase roadmap
- Identified decision points (NIP-86 vs. simple API)
- Next: Update overlapping issues with reference to this strategy