Files
ngit-grasp/ec1f-administrator-observability-and-management-strategy.md
T

494 lines
20 KiB
Markdown

# Administrator Observability and Management Strategy
**ID:** ec1f
## Problem
Administrators need a cohesive strategy for monitoring their GRASP relay and taking action when problems arise. Currently we have:
1. **Prometheus metrics** - comprehensive time-series data (connections, git ops, events, sync)
2. **Grafana dashboard** - pre-built but untested and potentially out of date
3. **Fragmented issues** - each addressing pieces (dashboard UI, management API, storage limits, rate limiting, abuse detection) without a unified vision
**Key gaps:**
- No clear path from "I see a problem" to "I fix the problem"
- No strategy for configuration management (declarative vs. dynamic)
- Overlapping concerns across multiple issues without coordination
- Untested/outdated Prometheus setup and Grafana dashboard
## Vision
Administrators should be able to:
1. **Observe** - See what's happening (metrics, dashboards, alerts)
2. **Understand** - Identify what needs attention (storage warnings, abusive users, large repos)
3. **Act** - Take corrective action (blacklist, delete, adjust quotas)
4. **Enforce** - Automatic protection (rate limits, storage quotas, abuse detection)
All while respecting deployment models:
- **Declarative deployments** (NixOS, Docker) - config from env/CLI only, dashboard read-only
- **Dynamic deployments** (manual) - optional GUI-based config changes that persist
## Configuration Model
### New Configuration Options
```bash
# Allow management dashboard to modify configuration
# Default: false (read-only dashboard)
NGIT_ALLOW_DYNAMIC_CONFIG=false
# Enforce declarative-only configuration (no local config file)
# Default: false (allow local config file if ALLOW_DYNAMIC_CONFIG=true)
NGIT_DECLARATIVE_ONLY=false
```
### Configuration Hierarchy
```
1. CLI args / env vars / .env (declarative, immutable)
↓
2. Local config file (dynamic, if ALLOW_DYNAMIC_CONFIG=true)
↓
3. Runtime state (ephemeral)
```
### Behavior Matrix
| DECLARATIVE_ONLY | ALLOW_DYNAMIC_CONFIG | Behavior |
|------------------|---------------------|----------|
| true | * (ignored) | Pure declarative - no local config file, dashboard read-only |
| false | false | Declarative - dashboard read-only, no local config file |
| false | true | Dynamic - dashboard can modify config, persists to local file |
**NixOS/Docker deployments:** Set `NGIT_DECLARATIVE_ONLY=true` to ensure purity
**Manual deployments:** Can enable `NGIT_ALLOW_DYNAMIC_CONFIG=true` for GUI-based management
## Phased Roadmap
### Phase 1: Observability (See what's happening)
**Goal:** Administrators can see current state and trends
**Tasks:**
- [ ] Test and update Prometheus setup documentation
- [ ] Test and update Grafana dashboard (verify all metrics work)
- [ ] Add missing storage metrics:
- [ ] Total disk usage (git + database)
- [ ] Available disk space percentage
- [ ] Per-user storage usage (top N users)
- [ ] Per-repository storage usage (largest repos)
- [ ] Add "attention needed" indicators:
- [ ] Storage approaching limits (80%, 90%, 95%)
- [ ] Flagged abusive IPs (from ConnectionTracker)
- [ ] Large repositories (potential blacklist candidates)
- [ ] High error rates (git operations, event rejections)
**Depends on:** Nothing
**Related issues:**
- 76fe (Repository Count Metric) - already implemented, verify it works
- 7d0b (Management Dashboard) - Phase 2 will build on this
### Phase 2: Insights (Understand what needs action)
**Goal:** Administrators can identify problems and understand their scope
**Tasks:**
- [ ] Implement Management Dashboard UI (issue 7d0b)
- [ ] Read-only view of all metrics
- [ ] Storage breakdown visualization
- [ ] Highlight problematic pubkeys/repos
- [ ] Recent activity feed
- [ ] System health indicators
- [ ] Dashboard respects configuration flags:
- [ ] Show "read-only mode" banner when ALLOW_DYNAMIC_CONFIG=false
- [ ] Hide action buttons in read-only mode
- [ ] Display current config source (env, file, CLI)
**Depends on:** Phase 1 (metrics must exist to display)
**Related issues:**
- 7d0b (Management Dashboard) - **PRIMARY ISSUE FOR THIS PHASE**
### Phase 3: Actions (Take action)
**Goal:** Administrators can take corrective action through authenticated API
**Decision point:** NIP-86 vs. simpler HTTP API?
**Option A: NIP-86 Relay Management API** (issue 2cdc)
- Standardized protocol, Nostr-native authentication
- More complex to implement
- Benefits entire Nostr ecosystem if we contribute patterns
**Option B: Simple HTTP API**
- Faster to implement
- Basic auth or API key
- Less standardization
**Recommendation:** Start with Option B (simple HTTP API), migrate to NIP-86 later if needed
**Tasks:**
- [ ] Implement authenticated management API:
- [ ] Blacklist user (npub) from all operations
- [ ] Blacklist repository (prevent pushes)
- [ ] Delete repository (with confirmation)
- [ ] Set per-user storage quota
- [ ] Set per-repo storage quota
- [ ] Force garbage collection on repository
- [ ] View audit log of admin actions
- [ ] Integrate with dashboard UI (if ALLOW_DYNAMIC_CONFIG=true):
- [ ] Action buttons for blacklist/delete/quota
- [ ] Confirmation dialogs for destructive actions
- [ ] Real-time feedback on action results
- [ ] Implement local config file persistence:
- [ ] Define config file format (TOML? JSON?)
- [ ] Load local config on startup (after env/CLI)
- [ ] Save config changes from dashboard
- [ ] Validate config on load
**Depends on:** Phase 2 (dashboard UI to trigger actions)
**Related issues:**
- 2cdc (NIP-86 Relay Management API) - **DEFERRED** in favor of simpler HTTP API first
- 7d0b (Management Dashboard) - UI integration
### Phase 4: Enforcement (Automatic protection)
**Goal:** Relay automatically protects itself from abuse and resource exhaustion
**Tasks:**
- [ ] Implement storage limits and quota management (issue 8430):
- [ ] Per-repository size limits
- [ ] Per-user total storage limits
- [ ] Relay-wide storage limits
- [ ] Graceful degradation (read-only mode when approaching limits)
- [ ] Clear error messages when quotas exceeded
- [ ] Implement defensive relay features (issue d6ee):
- [ ] Phase 1 only: Explicit rate limits + max connections (APPROVED)
- [ ] Defer per-IP enforcement until abuse detected in production
- [ ] Implement naughty list identification (issue 1f4f):
- [ ] Detect invalid signatures, filter violations, DoS patterns
- [ ] Automatic temporary bans for repeated violations
- [ ] Integration with blacklist system from Phase 3
**Depends on:** Phase 3 (manual blacklist/quota system must exist for overrides)
**Related issues:**
- 8430 (Storage Limits and Quota Management) - **PRIMARY ISSUE FOR THIS PHASE**
- d6ee (Defensive Relay Features) - Phase 1 only (config-based protection)
- 1f4f (Poor Naughty List Identification) - Abuse detection
## Issue Overlap Analysis
### Primary Overlaps
**Storage Management:**
- **7d0b (Management Dashboard)** - Display storage metrics
- **8430 (Storage Limits)** - Enforce storage quotas
- **2cdc (NIP-86 Management API)** - Set quotas via API
- **Resolution:** 8430 owns enforcement, 7d0b owns display, 2cdc owns API (deferred)
**Abuse Detection & Response:**
- **d6ee (Defensive Relay Features)** - Rate limiting, connection limits
- **1f4f (Poor Naughty List)** - Detect malicious behavior
- **2cdc (NIP-86 Management API)** - Blacklist users/repos
- **Resolution:** d6ee owns prevention, 1f4f owns detection, 2cdc owns manual response (deferred)
**Configuration & Monitoring:**
- **7d0b (Management Dashboard)** - Display config and metrics
- **2cdc (NIP-86 Management API)** - Modify config via API
- **76fe (Repository Count Metric)** - Specific metric implementation
- **Resolution:** 7d0b owns display, 2cdc owns modification (deferred), 76fe is complete
### Coordination Requirements
**Before starting work on any overlapping issue:**
1. Check this strategy issue (ec1f) for current phase
2. Verify dependencies are complete
3. Coordinate scope with related issues
4. Update this issue with progress
## Implementation Sequence
```
Phase 1: Observability
├── Update Prometheus/Grafana setup (1-2 days)
├── Add storage metrics (2-3 days)
└── Add attention indicators (1-2 days)
Total: ~1 week
Phase 2: Insights
├── Management Dashboard UI (3-5 days)
└── Configuration flag integration (1 day)
Total: ~1 week
Depends on: Phase 1
Phase 3: Actions
├── Simple HTTP management API (3-4 days)
├── Dashboard action integration (2-3 days)
└── Local config file persistence (2-3 days)
Total: ~1.5 weeks
Depends on: Phase 2
Phase 4: Enforcement
├── Storage limits (issue 8430) (1-2 weeks)
├── Defensive features Phase 1 (issue d6ee) (2-3 hours)
└── Naughty list detection (issue 1f4f) (1 week)
Total: ~2-3 weeks
Depends on: Phase 3
```
**Total estimated time:** 5-7 weeks for complete administrator experience
## Success Criteria
### Phase 1 Complete
- [ ] Prometheus metrics endpoint tested and documented
- [ ] Grafana dashboard loads without errors
- [ ] All metrics display current values
- [ ] Storage metrics available (disk usage, per-user, per-repo)
- [ ] Attention indicators highlight problems
### Phase 2 Complete
- [ ] Management dashboard accessible via HTTP endpoint
- [ ] Dashboard displays all metrics with visualizations
- [ ] Dashboard respects ALLOW_DYNAMIC_CONFIG flag
- [ ] Mobile-friendly responsive design
- [ ] No significant performance impact on relay
### Phase 3 Complete
- [ ] Authenticated management API functional
- [ ] Can blacklist users and repositories
- [ ] Can set storage quotas
- [ ] Can delete repositories with confirmation
- [ ] Audit log tracks all admin actions
- [ ] Local config file persists changes (when enabled)
### Phase 4 Complete
- [ ] Storage quotas enforced automatically
- [ ] Graceful degradation when approaching limits
- [ ] Rate limiting prevents DoS attacks
- [ ] Abuse detection flags problematic users
- [ ] Clear error messages guide users
### Overall Success
- [ ] Administrator can monitor relay health without SSH access
- [ ] Administrator can take corrective action through dashboard
- [ ] Declarative deployments remain pure (NixOS, Docker)
- [ ] Dynamic deployments support GUI-based management
- [ ] No manual intervention needed for common abuse scenarios
## Open Questions
1. **Config file format:** TOML, JSON, or YAML for local config file?
- Recommendation: TOML (matches Rust ecosystem, human-friendly)
2. **Authentication for management API:** Basic auth, API key, or Nostr signature?
- Recommendation: Start with API key (simple), migrate to Nostr signatures later
3. **Dashboard framework:** Server-rendered HTML or SPA?
- Recommendation: Server-rendered (htmx or Alpine.js) for simplicity
4. **NIP-86 timeline:** When to implement full NIP-86 compliance?
- Recommendation: After Phase 3 proves the simple API works
5. **Prometheus vs. custom metrics:** Should we add non-Prometheus metrics for dashboard?
- Recommendation: Prometheus only (single source of truth)
## Related Issues
**Directly related (must coordinate):**
- 7d0b (Management Dashboard) - Phase 2 primary issue
- 2cdc (NIP-86 Relay Management API) - Deferred to post-Phase 3
- 8430 (Storage Limits and Quota Management) - Phase 4 primary issue
- d6ee (Defensive Relay Features) - Phase 4 (Phase 1 only)
- 1f4f (Poor Naughty List Identification) - Phase 4
**Indirectly related:**
- 76fe (Repository Count Metric) - Already implemented, verify in Phase 1
- cc36 (Database Backend Evaluation) - May affect storage metrics
- b905 (Deletion Request Support) - May integrate with delete repository action
## Key Decisions
### 2026-01-14 Strategy Review
**Prometheus Metrics:**
- **Private, API key required** - Always authenticated via `NGIT_METRICS_API_KEY`
- Supports fleet monitoring (multiple instances reporting to central Prometheus)
- Contains sensitive admin data (per-user storage, per-repo sizes, etc.)
**Public Relay Stats:**
- **Delivered via NIP-11** - Extend relay info document with stats section
- Cache expensive metrics rather than computing on each request
- No separate HTTP endpoint needed
**User Storage Info:**
- **Delivered via replaceable Nostr event** - User-specific storage usage/limits
- Requires NIP-42 authentication to query own data
- Preferred approach for all sensitive user-specific data via Nostr
**Metrics Separation:**
- Public info (NIP-11): Safe aggregate stats for relay selection
- Private metrics (Prometheus): Detailed operational data for admins
- User-specific (Nostr events): Personal storage usage for authenticated users
## Next Steps
The current phased roadmap above is **SUPERSEDED** pending proper planning:
1. **Strategy Document** - `docs/explanation/metrics-and-administration.md`
- Thoughtful document outlining administrator workflow
- What metrics support each workflow step
- How different audiences (public, users, admins) access information
2. **Target Architecture** - Define technical architecture based on strategy
- Component interactions
- Data flows
- API designs
3. **Implementation Plan** - Only after strategy and architecture are complete
- Phased delivery
- Dependencies
- Estimates
## Progress
### 2026-01-14 [Session 16:30]
- Created strategy issue
- Analyzed existing issues for overlap
- Defined configuration model (DECLARATIVE_ONLY, ALLOW_DYNAMIC_CONFIG)
- Established 4-phase roadmap
- Identified decision points (NIP-86 vs. simple API)
### 2026-01-14 [Session ~17:00]
- Security review of public metrics exposure
- Decision: Prometheus metrics require API key authentication
- Decision: Public relay stats via NIP-11 (cached, not Prometheus)
- Decision: User storage info via replaceable Nostr events (NIP-42 auth)
- Recognized need for proper planning sequence: strategy → architecture → implementation
- Created worktree: `ec1f-administrator-observability-and-management-strategy`
- Wrote strategy document: `docs/explanation/metrics-and-administration.md`
- Defined three audiences (public, authenticated users, administrators)
- Documented five administrator workflows
- Specified information architecture for each audience
- Outlined configuration model and security considerations
- Added target architecture to strategy document:
- System overview with component diagram
- Stats Cache (60s refresh for NIP-11/homepage)
- Prometheus metrics with API key auth
- Management Service (REST API + SPA dashboard)
- User storage info via Nostr events (application-specific kind)
- Data flow diagrams
- Configuration and security architecture
- Resolved open questions: event kinds (app-specific), caching (60s), dashboard (SPA), audit (log file)
- Refined architecture based on feedback:
- Stats cache: throttled (not periodic), reuse existing homepage/NIP-11
- Prometheus: multiple hashed API keys, generated via admin UI
- Management API: NIP-86 (not REST), research existing implementations
- Audit logging: standard INFO messages in app logs
- User storage info: verify nostr-relay-builder support, NIP-44 fallback
- Next: Write implementation plan
### 2026-01-14 [Session ~18:00]
- Completed implementation plan in strategy document
- Broke work into 7 phases (Phase 0-7) with clear dependencies
- Identified within-phase parallelization opportunities:
- Phase 0: 2 parallel research tracks (NIP-86, nostr-relay-builder)
- Phase 1: 2 parallel tracks (stats cache, storage metrics)
- Phase 2: 2 parallel tracks (NIP-11 extension, user storage events)
- Phase 3: Single track (Prometheus auth)
- Phase 4: Single track (NIP-86 API)
- Phase 5: Single track (dashboard, limited internal parallelization)
- Phase 6: Sequential tasks (quota enforcement)
- Phase 7: 2 parallel tracks (rate limiting, abuse detection)
- Mapped all existing issues to phases:
- 76fe: Phase 1 Track B (verification)
- 7d0b: Phase 2 Track A + Phase 5 (9-12 days)
- 2cdc: Phase 0 Track A + Phase 4 + Phase 5 (8-10 days)
- 8430: Phase 4 + Phase 6 (9-11 days)
- d6ee: Phase 7 Track A (2-3 days)
- 1f4f: Phase 7 Track B (2-3 days)
- Effort estimates:
- Sequential: 31-39 days (6-8 weeks)
- With within-phase parallelization: 25-30 days (5-6 weeks)
- Added risk mitigation strategies for research and technical risks
- Defined success criteria for each phase and overall
- Each phase must commit working code with passing tests
- Code/commits must not reference phase numbers
- Ready for implementation to begin with Phase 0 research
### 2026-01-14 [Session ~18:30]
- Corrected implementation plan based on feedback
- Fixed incorrect parallelization claims:
- Phase 3 and 4 cannot run in parallel (Phase 4 depends on Phase 0 Track A research)
- Split original Phase 3 into two sequential phases (Prometheus auth, then NIP-86 API)
- Clarified "within-phase" vs "cross-phase" parallelization
- Added explicit deliverable requirements:
- Each phase must commit changes with all tests passing
- No phase references in code comments or commit messages
- Phase 0 commits research docs, all other phases commit working code
- Updated phase count from 6 to 7 phases
- Corrected dependency graph to show true sequential requirements
- More realistic effort estimates accounting for sequential dependencies
### 2026-01-15 [Session 09:54]
- Completed Phase 0 Track A: NIP-86 implementation research
- Used subagent to research nostr-sdk 0.43 support for NIP-86
- Key findings:
- ✅ NIP-98 HTTP auth fully supported in nostr-sdk (can leverage for authentication)
- ❌ NIP-86 JSON-RPC protocol NOT provided (must implement from scratch)
- Recommended approach: Implement NIP-86 using existing NIP-98 auth
- Estimated effort: 6-11 days for complete implementation
- Deliverable: `docs/research/nip-86-implementation-options.md` committed
- Next: Phase 0 Track B (nostr-relay-builder user storage events) - deferred for now
- Ready to proceed with Phase 1 implementation
### 2026-01-15 [Session 11:05]
- **Critical architectural decision:** Identified gap in design - NIP-86 is for actions, not queries
- Researched how to return sensitive data to authenticated users (admins and regular users)
- Launched 3 parallel research subagents:
1. NIP-86 query response patterns in spec and rust-nostr
2. Pyramid relay admin patterns (real-world implementation)
3. Feasibility of pyramid-style approach with rust-nostr
- Key findings from research:
- ✅ NIP-86 spec uses HTTP POST/response (NOT Nostr events for query results)
- ✅ Pyramid uses ephemeral events to specific WebSocket connections for queries
- ❌ nostr-relay-builder does NOT support connection-specific messaging
- ❌ Pyramid approach requires forking nostr-relay-builder or custom WebSocket handler
- **Decision: Use HTTP-based NIP-86 for all queries and actions**
- Queries: HTTP GET/POST with NIP-98 auth → JSON response
- Actions: HTTP POST with NIP-98 auth → JSON response
- Both admin and user queries use same pattern
- Simpler, standard, no framework modifications needed
- Deliverables committed:
- `docs/research/nip-86-query-responses.md` - NIP-86 spec analysis
- `docs/research/pyramid-admin-patterns.md` - Pyramid implementation study
- `docs/research/pyramid-style-queries-design.md` - Feasibility analysis for rust-nostr
- **Next:** Update strategy document to reflect HTTP-based approach for queries
- Phase 0 research complete, ready to finalize architecture and proceed to Phase 1
### 2026-01-15 [Session 11:38]
- **Final Phase 0 research:** Explored forking rust-nostr to support pyramid-style queries
- Comprehensive analysis of what it would take to modify nostr-relay-builder
- Key findings:
- Implementation effort: 5-7 days (vs 1-2 days for HTTP NIP-86)
- Ongoing maintenance: 18-29 hours/month (vs 1-2 hours/month for HTTP)
- rust-nostr is extremely active: 305 commits in 15 days (~20/day)
- Breaking changes in ~50% of releases
- Pyramid uses khatru framework (Go), NOT nostr-relay-builder
- Upstream acceptance likelihood: Low (10-20%)
- **Final decision confirmed: HTTP-based NIP-86 is the correct approach**
- No fork needed
- Standards-compliant
- Minimal maintenance burden
- Achieves same functional goals
- Deliverable committed:
- `docs/research/rust-nostr-fork-analysis.md` - Complete fork analysis (1,316 lines)
- **Phase 0 research COMPLETE** ✅
- All 4 research documents committed and ready for review
- Ready to update strategy document and begin Phase 1 implementation