issue: create ec1f - Administrator Observability and Management Strategy

This commit is contained in:
DanConwayDev
2026-01-14 12:11:43 +00:00
parent 052a0199d8
commit cf5d5eaf2d
@@ -0,0 +1,323 @@
# Administrator Observability and Management Strategy
**ID:** ec1f
## Problem
Administrators need a cohesive strategy for monitoring their GRASP relay and taking action when problems arise. Currently we have:
1. **Prometheus metrics** - comprehensive time-series data (connections, git ops, events, sync)
2. **Grafana dashboard** - pre-built but untested and potentially out of date
3. **Fragmented issues** - each addressing pieces (dashboard UI, management API, storage limits, rate limiting, abuse detection) without a unified vision
**Key gaps:**
- No clear path from "I see a problem" to "I fix the problem"
- No strategy for configuration management (declarative vs. dynamic)
- Overlapping concerns across multiple issues without coordination
- Untested/outdated Prometheus setup and Grafana dashboard
## Vision
Administrators should be able to:
1. **Observe** - See what's happening (metrics, dashboards, alerts)
2. **Understand** - Identify what needs attention (storage warnings, abusive users, large repos)
3. **Act** - Take corrective action (blacklist, delete, adjust quotas)
4. **Enforce** - Automatic protection (rate limits, storage quotas, abuse detection)
All while respecting deployment models:
- **Declarative deployments** (NixOS, Docker) - config from env/CLI only, dashboard read-only
- **Dynamic deployments** (manual) - optional GUI-based config changes that persist
## Configuration Model
### New Configuration Options
```bash
# Allow management dashboard to modify configuration
# Default: false (read-only dashboard)
NGIT_ALLOW_DYNAMIC_CONFIG=false
# Enforce declarative-only configuration (no local config file)
# Default: false (allow local config file if ALLOW_DYNAMIC_CONFIG=true)
NGIT_DECLARATIVE_ONLY=false
```
### Configuration Hierarchy
```
1. CLI args / env vars / .env (declarative, immutable)
↓
2. Local config file (dynamic, if ALLOW_DYNAMIC_CONFIG=true)
↓
3. Runtime state (ephemeral)
```
### Behavior Matrix
| DECLARATIVE_ONLY | ALLOW_DYNAMIC_CONFIG | Behavior |
|------------------|---------------------|----------|
| true | * (ignored) | Pure declarative - no local config file, dashboard read-only |
| false | false | Declarative - dashboard read-only, no local config file |
| false | true | Dynamic - dashboard can modify config, persists to local file |
**NixOS/Docker deployments:** Set `NGIT_DECLARATIVE_ONLY=true` to ensure purity
**Manual deployments:** Can enable `NGIT_ALLOW_DYNAMIC_CONFIG=true` for GUI-based management
## Phased Roadmap
### Phase 1: Observability (See what's happening)
**Goal:** Administrators can see current state and trends
**Tasks:**
- [ ] Test and update Prometheus setup documentation
- [ ] Test and update Grafana dashboard (verify all metrics work)
- [ ] Add missing storage metrics:
- [ ] Total disk usage (git + database)
- [ ] Available disk space percentage
- [ ] Per-user storage usage (top N users)
- [ ] Per-repository storage usage (largest repos)
- [ ] Add "attention needed" indicators:
- [ ] Storage approaching limits (80%, 90%, 95%)
- [ ] Flagged abusive IPs (from ConnectionTracker)
- [ ] Large repositories (potential blacklist candidates)
- [ ] High error rates (git operations, event rejections)
**Depends on:** Nothing
**Related issues:**
- 76fe (Repository Count Metric) - already implemented, verify it works
- 7d0b (Management Dashboard) - Phase 2 will build on this
### Phase 2: Insights (Understand what needs action)
**Goal:** Administrators can identify problems and understand their scope
**Tasks:**
- [ ] Implement Management Dashboard UI (issue 7d0b)
- [ ] Read-only view of all metrics
- [ ] Storage breakdown visualization
- [ ] Highlight problematic pubkeys/repos
- [ ] Recent activity feed
- [ ] System health indicators
- [ ] Dashboard respects configuration flags:
- [ ] Show "read-only mode" banner when ALLOW_DYNAMIC_CONFIG=false
- [ ] Hide action buttons in read-only mode
- [ ] Display current config source (env, file, CLI)
**Depends on:** Phase 1 (metrics must exist to display)
**Related issues:**
- 7d0b (Management Dashboard) - **PRIMARY ISSUE FOR THIS PHASE**
### Phase 3: Actions (Take action)
**Goal:** Administrators can take corrective action through authenticated API
**Decision point:** NIP-86 vs. simpler HTTP API?
**Option A: NIP-86 Relay Management API** (issue 2cdc)
- Standardized protocol, Nostr-native authentication
- More complex to implement
- Benefits entire Nostr ecosystem if we contribute patterns
**Option B: Simple HTTP API**
- Faster to implement
- Basic auth or API key
- Less standardization
**Recommendation:** Start with Option B (simple HTTP API), migrate to NIP-86 later if needed
**Tasks:**
- [ ] Implement authenticated management API:
- [ ] Blacklist user (npub) from all operations
- [ ] Blacklist repository (prevent pushes)
- [ ] Delete repository (with confirmation)
- [ ] Set per-user storage quota
- [ ] Set per-repo storage quota
- [ ] Force garbage collection on repository
- [ ] View audit log of admin actions
- [ ] Integrate with dashboard UI (if ALLOW_DYNAMIC_CONFIG=true):
- [ ] Action buttons for blacklist/delete/quota
- [ ] Confirmation dialogs for destructive actions
- [ ] Real-time feedback on action results
- [ ] Implement local config file persistence:
- [ ] Define config file format (TOML? JSON?)
- [ ] Load local config on startup (after env/CLI)
- [ ] Save config changes from dashboard
- [ ] Validate config on load
**Depends on:** Phase 2 (dashboard UI to trigger actions)
**Related issues:**
- 2cdc (NIP-86 Relay Management API) - **DEFERRED** in favor of simpler HTTP API first
- 7d0b (Management Dashboard) - UI integration
### Phase 4: Enforcement (Automatic protection)
**Goal:** Relay automatically protects itself from abuse and resource exhaustion
**Tasks:**
- [ ] Implement storage limits and quota management (issue 8430):
- [ ] Per-repository size limits
- [ ] Per-user total storage limits
- [ ] Relay-wide storage limits
- [ ] Graceful degradation (read-only mode when approaching limits)
- [ ] Clear error messages when quotas exceeded
- [ ] Implement defensive relay features (issue d6ee):
- [ ] Phase 1 only: Explicit rate limits + max connections (APPROVED)
- [ ] Defer per-IP enforcement until abuse detected in production
- [ ] Implement naughty list identification (issue 1f4f):
- [ ] Detect invalid signatures, filter violations, DoS patterns
- [ ] Automatic temporary bans for repeated violations
- [ ] Integration with blacklist system from Phase 3
**Depends on:** Phase 3 (manual blacklist/quota system must exist for overrides)
**Related issues:**
- 8430 (Storage Limits and Quota Management) - **PRIMARY ISSUE FOR THIS PHASE**
- d6ee (Defensive Relay Features) - Phase 1 only (config-based protection)
- 1f4f (Poor Naughty List Identification) - Abuse detection
## Issue Overlap Analysis
### Primary Overlaps
**Storage Management:**
- **7d0b (Management Dashboard)** - Display storage metrics
- **8430 (Storage Limits)** - Enforce storage quotas
- **2cdc (NIP-86 Management API)** - Set quotas via API
- **Resolution:** 8430 owns enforcement, 7d0b owns display, 2cdc owns API (deferred)
**Abuse Detection & Response:**
- **d6ee (Defensive Relay Features)** - Rate limiting, connection limits
- **1f4f (Poor Naughty List)** - Detect malicious behavior
- **2cdc (NIP-86 Management API)** - Blacklist users/repos
- **Resolution:** d6ee owns prevention, 1f4f owns detection, 2cdc owns manual response (deferred)
**Configuration & Monitoring:**
- **7d0b (Management Dashboard)** - Display config and metrics
- **2cdc (NIP-86 Management API)** - Modify config via API
- **76fe (Repository Count Metric)** - Specific metric implementation
- **Resolution:** 7d0b owns display, 2cdc owns modification (deferred), 76fe is complete
### Coordination Requirements
**Before starting work on any overlapping issue:**
1. Check this strategy issue (ec1f) for current phase
2. Verify dependencies are complete
3. Coordinate scope with related issues
4. Update this issue with progress
## Implementation Sequence
```
Phase 1: Observability
├── Update Prometheus/Grafana setup (1-2 days)
├── Add storage metrics (2-3 days)
└── Add attention indicators (1-2 days)
Total: ~1 week
Phase 2: Insights
├── Management Dashboard UI (3-5 days)
└── Configuration flag integration (1 day)
Total: ~1 week
Depends on: Phase 1
Phase 3: Actions
├── Simple HTTP management API (3-4 days)
├── Dashboard action integration (2-3 days)
└── Local config file persistence (2-3 days)
Total: ~1.5 weeks
Depends on: Phase 2
Phase 4: Enforcement
├── Storage limits (issue 8430) (1-2 weeks)
├── Defensive features Phase 1 (issue d6ee) (2-3 hours)
└── Naughty list detection (issue 1f4f) (1 week)
Total: ~2-3 weeks
Depends on: Phase 3
```
**Total estimated time:** 5-7 weeks for complete administrator experience
## Success Criteria
### Phase 1 Complete
- [ ] Prometheus metrics endpoint tested and documented
- [ ] Grafana dashboard loads without errors
- [ ] All metrics display current values
- [ ] Storage metrics available (disk usage, per-user, per-repo)
- [ ] Attention indicators highlight problems
### Phase 2 Complete
- [ ] Management dashboard accessible via HTTP endpoint
- [ ] Dashboard displays all metrics with visualizations
- [ ] Dashboard respects ALLOW_DYNAMIC_CONFIG flag
- [ ] Mobile-friendly responsive design
- [ ] No significant performance impact on relay
### Phase 3 Complete
- [ ] Authenticated management API functional
- [ ] Can blacklist users and repositories
- [ ] Can set storage quotas
- [ ] Can delete repositories with confirmation
- [ ] Audit log tracks all admin actions
- [ ] Local config file persists changes (when enabled)
### Phase 4 Complete
- [ ] Storage quotas enforced automatically
- [ ] Graceful degradation when approaching limits
- [ ] Rate limiting prevents DoS attacks
- [ ] Abuse detection flags problematic users
- [ ] Clear error messages guide users
### Overall Success
- [ ] Administrator can monitor relay health without SSH access
- [ ] Administrator can take corrective action through dashboard
- [ ] Declarative deployments remain pure (NixOS, Docker)
- [ ] Dynamic deployments support GUI-based management
- [ ] No manual intervention needed for common abuse scenarios
## Open Questions
1. **Config file format:** TOML, JSON, or YAML for local config file?
- Recommendation: TOML (matches Rust ecosystem, human-friendly)
2. **Authentication for management API:** Basic auth, API key, or Nostr signature?
- Recommendation: Start with API key (simple), migrate to Nostr signatures later
3. **Dashboard framework:** Server-rendered HTML or SPA?
- Recommendation: Server-rendered (htmx or Alpine.js) for simplicity
4. **NIP-86 timeline:** When to implement full NIP-86 compliance?
- Recommendation: After Phase 3 proves the simple API works
5. **Prometheus vs. custom metrics:** Should we add non-Prometheus metrics for dashboard?
- Recommendation: Prometheus only (single source of truth)
## Related Issues
**Directly related (must coordinate):**
- 7d0b (Management Dashboard) - Phase 2 primary issue
- 2cdc (NIP-86 Relay Management API) - Deferred to post-Phase 3
- 8430 (Storage Limits and Quota Management) - Phase 4 primary issue
- d6ee (Defensive Relay Features) - Phase 4 (Phase 1 only)
- 1f4f (Poor Naughty List Identification) - Phase 4
**Indirectly related:**
- 76fe (Repository Count Metric) - Already implemented, verify in Phase 1
- cc36 (Database Backend Evaluation) - May affect storage metrics
- b905 (Deletion Request Support) - May integrate with delete repository action
## Progress
### 2026-01-14 [Session 16:30]
- Created strategy issue
- Analyzed existing issues for overlap
- Defined configuration model (DECLARATIVE_ONLY, ALLOW_DYNAMIC_CONFIG)
- Established 4-phase roadmap
- Identified decision points (NIP-86 vs. simple API)
- Next: Update overlapping issues with reference to this strategy