diff --git a/ec1f-administrator-observability-and-management-strategy.md b/ec1f-administrator-observability-and-management-strategy.md new file mode 100644 index 0000000..c148caa --- /dev/null +++ b/ec1f-administrator-observability-and-management-strategy.md @@ -0,0 +1,323 @@ +# Administrator Observability and Management Strategy + +**ID:** ec1f + +## Problem + +Administrators need a cohesive strategy for monitoring their GRASP relay and taking action when problems arise. Currently we have: + +1. **Prometheus metrics** - comprehensive time-series data (connections, git ops, events, sync) +2. **Grafana dashboard** - pre-built but untested and potentially out of date +3. **Fragmented issues** - each addressing pieces (dashboard UI, management API, storage limits, rate limiting, abuse detection) without a unified vision + +**Key gaps:** +- No clear path from "I see a problem" to "I fix the problem" +- No strategy for configuration management (declarative vs. dynamic) +- Overlapping concerns across multiple issues without coordination +- Untested/outdated Prometheus setup and Grafana dashboard + +## Vision + +Administrators should be able to: + +1. **Observe** - See what's happening (metrics, dashboards, alerts) +2. **Understand** - Identify what needs attention (storage warnings, abusive users, large repos) +3. **Act** - Take corrective action (blacklist, delete, adjust quotas) +4. **Enforce** - Automatic protection (rate limits, storage quotas, abuse detection) + +All while respecting deployment models: +- **Declarative deployments** (NixOS, Docker) - config from env/CLI only, dashboard read-only +- **Dynamic deployments** (manual) - optional GUI-based config changes that persist + +## Configuration Model + +### New Configuration Options + +```bash +# Allow management dashboard to modify configuration +# Default: false (read-only dashboard) +NGIT_ALLOW_DYNAMIC_CONFIG=false + +# Enforce declarative-only configuration (no local config file) +# Default: false (allow local config file if ALLOW_DYNAMIC_CONFIG=true) +NGIT_DECLARATIVE_ONLY=false +``` + +### Configuration Hierarchy + +``` +1. CLI args / env vars / .env (declarative, immutable) + ↓ +2. Local config file (dynamic, if ALLOW_DYNAMIC_CONFIG=true) + ↓ +3. Runtime state (ephemeral) +``` + +### Behavior Matrix + +| DECLARATIVE_ONLY | ALLOW_DYNAMIC_CONFIG | Behavior | +|------------------|---------------------|----------| +| true | * (ignored) | Pure declarative - no local config file, dashboard read-only | +| false | false | Declarative - dashboard read-only, no local config file | +| false | true | Dynamic - dashboard can modify config, persists to local file | + +**NixOS/Docker deployments:** Set `NGIT_DECLARATIVE_ONLY=true` to ensure purity + +**Manual deployments:** Can enable `NGIT_ALLOW_DYNAMIC_CONFIG=true` for GUI-based management + +## Phased Roadmap + +### Phase 1: Observability (See what's happening) + +**Goal:** Administrators can see current state and trends + +**Tasks:** +- [ ] Test and update Prometheus setup documentation +- [ ] Test and update Grafana dashboard (verify all metrics work) +- [ ] Add missing storage metrics: + - [ ] Total disk usage (git + database) + - [ ] Available disk space percentage + - [ ] Per-user storage usage (top N users) + - [ ] Per-repository storage usage (largest repos) +- [ ] Add "attention needed" indicators: + - [ ] Storage approaching limits (80%, 90%, 95%) + - [ ] Flagged abusive IPs (from ConnectionTracker) + - [ ] Large repositories (potential blacklist candidates) + - [ ] High error rates (git operations, event rejections) + +**Depends on:** Nothing + +**Related issues:** +- 76fe (Repository Count Metric) - already implemented, verify it works +- 7d0b (Management Dashboard) - Phase 2 will build on this + +### Phase 2: Insights (Understand what needs action) + +**Goal:** Administrators can identify problems and understand their scope + +**Tasks:** +- [ ] Implement Management Dashboard UI (issue 7d0b) + - [ ] Read-only view of all metrics + - [ ] Storage breakdown visualization + - [ ] Highlight problematic pubkeys/repos + - [ ] Recent activity feed + - [ ] System health indicators +- [ ] Dashboard respects configuration flags: + - [ ] Show "read-only mode" banner when ALLOW_DYNAMIC_CONFIG=false + - [ ] Hide action buttons in read-only mode + - [ ] Display current config source (env, file, CLI) + +**Depends on:** Phase 1 (metrics must exist to display) + +**Related issues:** +- 7d0b (Management Dashboard) - **PRIMARY ISSUE FOR THIS PHASE** + +### Phase 3: Actions (Take action) + +**Goal:** Administrators can take corrective action through authenticated API + +**Decision point:** NIP-86 vs. simpler HTTP API? + +**Option A: NIP-86 Relay Management API** (issue 2cdc) +- Standardized protocol, Nostr-native authentication +- More complex to implement +- Benefits entire Nostr ecosystem if we contribute patterns + +**Option B: Simple HTTP API** +- Faster to implement +- Basic auth or API key +- Less standardization + +**Recommendation:** Start with Option B (simple HTTP API), migrate to NIP-86 later if needed + +**Tasks:** +- [ ] Implement authenticated management API: + - [ ] Blacklist user (npub) from all operations + - [ ] Blacklist repository (prevent pushes) + - [ ] Delete repository (with confirmation) + - [ ] Set per-user storage quota + - [ ] Set per-repo storage quota + - [ ] Force garbage collection on repository + - [ ] View audit log of admin actions +- [ ] Integrate with dashboard UI (if ALLOW_DYNAMIC_CONFIG=true): + - [ ] Action buttons for blacklist/delete/quota + - [ ] Confirmation dialogs for destructive actions + - [ ] Real-time feedback on action results +- [ ] Implement local config file persistence: + - [ ] Define config file format (TOML? JSON?) + - [ ] Load local config on startup (after env/CLI) + - [ ] Save config changes from dashboard + - [ ] Validate config on load + +**Depends on:** Phase 2 (dashboard UI to trigger actions) + +**Related issues:** +- 2cdc (NIP-86 Relay Management API) - **DEFERRED** in favor of simpler HTTP API first +- 7d0b (Management Dashboard) - UI integration + +### Phase 4: Enforcement (Automatic protection) + +**Goal:** Relay automatically protects itself from abuse and resource exhaustion + +**Tasks:** +- [ ] Implement storage limits and quota management (issue 8430): + - [ ] Per-repository size limits + - [ ] Per-user total storage limits + - [ ] Relay-wide storage limits + - [ ] Graceful degradation (read-only mode when approaching limits) + - [ ] Clear error messages when quotas exceeded +- [ ] Implement defensive relay features (issue d6ee): + - [ ] Phase 1 only: Explicit rate limits + max connections (APPROVED) + - [ ] Defer per-IP enforcement until abuse detected in production +- [ ] Implement naughty list identification (issue 1f4f): + - [ ] Detect invalid signatures, filter violations, DoS patterns + - [ ] Automatic temporary bans for repeated violations + - [ ] Integration with blacklist system from Phase 3 + +**Depends on:** Phase 3 (manual blacklist/quota system must exist for overrides) + +**Related issues:** +- 8430 (Storage Limits and Quota Management) - **PRIMARY ISSUE FOR THIS PHASE** +- d6ee (Defensive Relay Features) - Phase 1 only (config-based protection) +- 1f4f (Poor Naughty List Identification) - Abuse detection + +## Issue Overlap Analysis + +### Primary Overlaps + +**Storage Management:** +- **7d0b (Management Dashboard)** - Display storage metrics +- **8430 (Storage Limits)** - Enforce storage quotas +- **2cdc (NIP-86 Management API)** - Set quotas via API +- **Resolution:** 8430 owns enforcement, 7d0b owns display, 2cdc owns API (deferred) + +**Abuse Detection & Response:** +- **d6ee (Defensive Relay Features)** - Rate limiting, connection limits +- **1f4f (Poor Naughty List)** - Detect malicious behavior +- **2cdc (NIP-86 Management API)** - Blacklist users/repos +- **Resolution:** d6ee owns prevention, 1f4f owns detection, 2cdc owns manual response (deferred) + +**Configuration & Monitoring:** +- **7d0b (Management Dashboard)** - Display config and metrics +- **2cdc (NIP-86 Management API)** - Modify config via API +- **76fe (Repository Count Metric)** - Specific metric implementation +- **Resolution:** 7d0b owns display, 2cdc owns modification (deferred), 76fe is complete + +### Coordination Requirements + +**Before starting work on any overlapping issue:** +1. Check this strategy issue (ec1f) for current phase +2. Verify dependencies are complete +3. Coordinate scope with related issues +4. Update this issue with progress + +## Implementation Sequence + +``` +Phase 1: Observability +├── Update Prometheus/Grafana setup (1-2 days) +├── Add storage metrics (2-3 days) +└── Add attention indicators (1-2 days) + Total: ~1 week + +Phase 2: Insights +├── Management Dashboard UI (3-5 days) +└── Configuration flag integration (1 day) + Total: ~1 week + Depends on: Phase 1 + +Phase 3: Actions +├── Simple HTTP management API (3-4 days) +├── Dashboard action integration (2-3 days) +└── Local config file persistence (2-3 days) + Total: ~1.5 weeks + Depends on: Phase 2 + +Phase 4: Enforcement +├── Storage limits (issue 8430) (1-2 weeks) +├── Defensive features Phase 1 (issue d6ee) (2-3 hours) +└── Naughty list detection (issue 1f4f) (1 week) + Total: ~2-3 weeks + Depends on: Phase 3 +``` + +**Total estimated time:** 5-7 weeks for complete administrator experience + +## Success Criteria + +### Phase 1 Complete +- [ ] Prometheus metrics endpoint tested and documented +- [ ] Grafana dashboard loads without errors +- [ ] All metrics display current values +- [ ] Storage metrics available (disk usage, per-user, per-repo) +- [ ] Attention indicators highlight problems + +### Phase 2 Complete +- [ ] Management dashboard accessible via HTTP endpoint +- [ ] Dashboard displays all metrics with visualizations +- [ ] Dashboard respects ALLOW_DYNAMIC_CONFIG flag +- [ ] Mobile-friendly responsive design +- [ ] No significant performance impact on relay + +### Phase 3 Complete +- [ ] Authenticated management API functional +- [ ] Can blacklist users and repositories +- [ ] Can set storage quotas +- [ ] Can delete repositories with confirmation +- [ ] Audit log tracks all admin actions +- [ ] Local config file persists changes (when enabled) + +### Phase 4 Complete +- [ ] Storage quotas enforced automatically +- [ ] Graceful degradation when approaching limits +- [ ] Rate limiting prevents DoS attacks +- [ ] Abuse detection flags problematic users +- [ ] Clear error messages guide users + +### Overall Success +- [ ] Administrator can monitor relay health without SSH access +- [ ] Administrator can take corrective action through dashboard +- [ ] Declarative deployments remain pure (NixOS, Docker) +- [ ] Dynamic deployments support GUI-based management +- [ ] No manual intervention needed for common abuse scenarios + +## Open Questions + +1. **Config file format:** TOML, JSON, or YAML for local config file? + - Recommendation: TOML (matches Rust ecosystem, human-friendly) + +2. **Authentication for management API:** Basic auth, API key, or Nostr signature? + - Recommendation: Start with API key (simple), migrate to Nostr signatures later + +3. **Dashboard framework:** Server-rendered HTML or SPA? + - Recommendation: Server-rendered (htmx or Alpine.js) for simplicity + +4. **NIP-86 timeline:** When to implement full NIP-86 compliance? + - Recommendation: After Phase 3 proves the simple API works + +5. **Prometheus vs. custom metrics:** Should we add non-Prometheus metrics for dashboard? + - Recommendation: Prometheus only (single source of truth) + +## Related Issues + +**Directly related (must coordinate):** +- 7d0b (Management Dashboard) - Phase 2 primary issue +- 2cdc (NIP-86 Relay Management API) - Deferred to post-Phase 3 +- 8430 (Storage Limits and Quota Management) - Phase 4 primary issue +- d6ee (Defensive Relay Features) - Phase 4 (Phase 1 only) +- 1f4f (Poor Naughty List Identification) - Phase 4 + +**Indirectly related:** +- 76fe (Repository Count Metric) - Already implemented, verify in Phase 1 +- cc36 (Database Backend Evaluation) - May affect storage metrics +- b905 (Deletion Request Support) - May integrate with delete repository action + +## Progress + +### 2026-01-14 [Session 16:30] +- Created strategy issue +- Analyzed existing issues for overlap +- Defined configuration model (DECLARATIVE_ONLY, ALLOW_DYNAMIC_CONFIG) +- Established 4-phase roadmap +- Identified decision points (NIP-86 vs. simple API) +- Next: Update overlapping issues with reference to this strategy