mirror of
https://relay.ngit.dev/npub15qydau2hjma6ngxkl2cyar74wzyjshvl65za5k5rl69264ar2exs5cyejr/ngit-grasp.git
synced 2026-10-05 15:08:24 +00:00
issue: create ec1f - Administrator Observability and Management Strategy
This commit is contained in:
@@ -0,0 +1,323 @@
|
||||
# Administrator Observability and Management Strategy
|
||||
|
||||
**ID:** ec1f
|
||||
|
||||
## Problem
|
||||
|
||||
Administrators need a cohesive strategy for monitoring their GRASP relay and taking action when problems arise. Currently we have:
|
||||
|
||||
1. **Prometheus metrics** - comprehensive time-series data (connections, git ops, events, sync)
|
||||
2. **Grafana dashboard** - pre-built but untested and potentially out of date
|
||||
3. **Fragmented issues** - each addressing pieces (dashboard UI, management API, storage limits, rate limiting, abuse detection) without a unified vision
|
||||
|
||||
**Key gaps:**
|
||||
- No clear path from "I see a problem" to "I fix the problem"
|
||||
- No strategy for configuration management (declarative vs. dynamic)
|
||||
- Overlapping concerns across multiple issues without coordination
|
||||
- Untested/outdated Prometheus setup and Grafana dashboard
|
||||
|
||||
## Vision
|
||||
|
||||
Administrators should be able to:
|
||||
|
||||
1. **Observe** - See what's happening (metrics, dashboards, alerts)
|
||||
2. **Understand** - Identify what needs attention (storage warnings, abusive users, large repos)
|
||||
3. **Act** - Take corrective action (blacklist, delete, adjust quotas)
|
||||
4. **Enforce** - Automatic protection (rate limits, storage quotas, abuse detection)
|
||||
|
||||
All while respecting deployment models:
|
||||
- **Declarative deployments** (NixOS, Docker) - config from env/CLI only, dashboard read-only
|
||||
- **Dynamic deployments** (manual) - optional GUI-based config changes that persist
|
||||
|
||||
## Configuration Model
|
||||
|
||||
### New Configuration Options
|
||||
|
||||
```bash
|
||||
# Allow management dashboard to modify configuration
|
||||
# Default: false (read-only dashboard)
|
||||
NGIT_ALLOW_DYNAMIC_CONFIG=false
|
||||
|
||||
# Enforce declarative-only configuration (no local config file)
|
||||
# Default: false (allow local config file if ALLOW_DYNAMIC_CONFIG=true)
|
||||
NGIT_DECLARATIVE_ONLY=false
|
||||
```
|
||||
|
||||
### Configuration Hierarchy
|
||||
|
||||
```
|
||||
1. CLI args / env vars / .env (declarative, immutable)
|
||||
↓
|
||||
2. Local config file (dynamic, if ALLOW_DYNAMIC_CONFIG=true)
|
||||
↓
|
||||
3. Runtime state (ephemeral)
|
||||
```
|
||||
|
||||
### Behavior Matrix
|
||||
|
||||
| DECLARATIVE_ONLY | ALLOW_DYNAMIC_CONFIG | Behavior |
|
||||
|------------------|---------------------|----------|
|
||||
| true | * (ignored) | Pure declarative - no local config file, dashboard read-only |
|
||||
| false | false | Declarative - dashboard read-only, no local config file |
|
||||
| false | true | Dynamic - dashboard can modify config, persists to local file |
|
||||
|
||||
**NixOS/Docker deployments:** Set `NGIT_DECLARATIVE_ONLY=true` to ensure purity
|
||||
|
||||
**Manual deployments:** Can enable `NGIT_ALLOW_DYNAMIC_CONFIG=true` for GUI-based management
|
||||
|
||||
## Phased Roadmap
|
||||
|
||||
### Phase 1: Observability (See what's happening)
|
||||
|
||||
**Goal:** Administrators can see current state and trends
|
||||
|
||||
**Tasks:**
|
||||
- [ ] Test and update Prometheus setup documentation
|
||||
- [ ] Test and update Grafana dashboard (verify all metrics work)
|
||||
- [ ] Add missing storage metrics:
|
||||
- [ ] Total disk usage (git + database)
|
||||
- [ ] Available disk space percentage
|
||||
- [ ] Per-user storage usage (top N users)
|
||||
- [ ] Per-repository storage usage (largest repos)
|
||||
- [ ] Add "attention needed" indicators:
|
||||
- [ ] Storage approaching limits (80%, 90%, 95%)
|
||||
- [ ] Flagged abusive IPs (from ConnectionTracker)
|
||||
- [ ] Large repositories (potential blacklist candidates)
|
||||
- [ ] High error rates (git operations, event rejections)
|
||||
|
||||
**Depends on:** Nothing
|
||||
|
||||
**Related issues:**
|
||||
- 76fe (Repository Count Metric) - already implemented, verify it works
|
||||
- 7d0b (Management Dashboard) - Phase 2 will build on this
|
||||
|
||||
### Phase 2: Insights (Understand what needs action)
|
||||
|
||||
**Goal:** Administrators can identify problems and understand their scope
|
||||
|
||||
**Tasks:**
|
||||
- [ ] Implement Management Dashboard UI (issue 7d0b)
|
||||
- [ ] Read-only view of all metrics
|
||||
- [ ] Storage breakdown visualization
|
||||
- [ ] Highlight problematic pubkeys/repos
|
||||
- [ ] Recent activity feed
|
||||
- [ ] System health indicators
|
||||
- [ ] Dashboard respects configuration flags:
|
||||
- [ ] Show "read-only mode" banner when ALLOW_DYNAMIC_CONFIG=false
|
||||
- [ ] Hide action buttons in read-only mode
|
||||
- [ ] Display current config source (env, file, CLI)
|
||||
|
||||
**Depends on:** Phase 1 (metrics must exist to display)
|
||||
|
||||
**Related issues:**
|
||||
- 7d0b (Management Dashboard) - **PRIMARY ISSUE FOR THIS PHASE**
|
||||
|
||||
### Phase 3: Actions (Take action)
|
||||
|
||||
**Goal:** Administrators can take corrective action through authenticated API
|
||||
|
||||
**Decision point:** NIP-86 vs. simpler HTTP API?
|
||||
|
||||
**Option A: NIP-86 Relay Management API** (issue 2cdc)
|
||||
- Standardized protocol, Nostr-native authentication
|
||||
- More complex to implement
|
||||
- Benefits entire Nostr ecosystem if we contribute patterns
|
||||
|
||||
**Option B: Simple HTTP API**
|
||||
- Faster to implement
|
||||
- Basic auth or API key
|
||||
- Less standardization
|
||||
|
||||
**Recommendation:** Start with Option B (simple HTTP API), migrate to NIP-86 later if needed
|
||||
|
||||
**Tasks:**
|
||||
- [ ] Implement authenticated management API:
|
||||
- [ ] Blacklist user (npub) from all operations
|
||||
- [ ] Blacklist repository (prevent pushes)
|
||||
- [ ] Delete repository (with confirmation)
|
||||
- [ ] Set per-user storage quota
|
||||
- [ ] Set per-repo storage quota
|
||||
- [ ] Force garbage collection on repository
|
||||
- [ ] View audit log of admin actions
|
||||
- [ ] Integrate with dashboard UI (if ALLOW_DYNAMIC_CONFIG=true):
|
||||
- [ ] Action buttons for blacklist/delete/quota
|
||||
- [ ] Confirmation dialogs for destructive actions
|
||||
- [ ] Real-time feedback on action results
|
||||
- [ ] Implement local config file persistence:
|
||||
- [ ] Define config file format (TOML? JSON?)
|
||||
- [ ] Load local config on startup (after env/CLI)
|
||||
- [ ] Save config changes from dashboard
|
||||
- [ ] Validate config on load
|
||||
|
||||
**Depends on:** Phase 2 (dashboard UI to trigger actions)
|
||||
|
||||
**Related issues:**
|
||||
- 2cdc (NIP-86 Relay Management API) - **DEFERRED** in favor of simpler HTTP API first
|
||||
- 7d0b (Management Dashboard) - UI integration
|
||||
|
||||
### Phase 4: Enforcement (Automatic protection)
|
||||
|
||||
**Goal:** Relay automatically protects itself from abuse and resource exhaustion
|
||||
|
||||
**Tasks:**
|
||||
- [ ] Implement storage limits and quota management (issue 8430):
|
||||
- [ ] Per-repository size limits
|
||||
- [ ] Per-user total storage limits
|
||||
- [ ] Relay-wide storage limits
|
||||
- [ ] Graceful degradation (read-only mode when approaching limits)
|
||||
- [ ] Clear error messages when quotas exceeded
|
||||
- [ ] Implement defensive relay features (issue d6ee):
|
||||
- [ ] Phase 1 only: Explicit rate limits + max connections (APPROVED)
|
||||
- [ ] Defer per-IP enforcement until abuse detected in production
|
||||
- [ ] Implement naughty list identification (issue 1f4f):
|
||||
- [ ] Detect invalid signatures, filter violations, DoS patterns
|
||||
- [ ] Automatic temporary bans for repeated violations
|
||||
- [ ] Integration with blacklist system from Phase 3
|
||||
|
||||
**Depends on:** Phase 3 (manual blacklist/quota system must exist for overrides)
|
||||
|
||||
**Related issues:**
|
||||
- 8430 (Storage Limits and Quota Management) - **PRIMARY ISSUE FOR THIS PHASE**
|
||||
- d6ee (Defensive Relay Features) - Phase 1 only (config-based protection)
|
||||
- 1f4f (Poor Naughty List Identification) - Abuse detection
|
||||
|
||||
## Issue Overlap Analysis
|
||||
|
||||
### Primary Overlaps
|
||||
|
||||
**Storage Management:**
|
||||
- **7d0b (Management Dashboard)** - Display storage metrics
|
||||
- **8430 (Storage Limits)** - Enforce storage quotas
|
||||
- **2cdc (NIP-86 Management API)** - Set quotas via API
|
||||
- **Resolution:** 8430 owns enforcement, 7d0b owns display, 2cdc owns API (deferred)
|
||||
|
||||
**Abuse Detection & Response:**
|
||||
- **d6ee (Defensive Relay Features)** - Rate limiting, connection limits
|
||||
- **1f4f (Poor Naughty List)** - Detect malicious behavior
|
||||
- **2cdc (NIP-86 Management API)** - Blacklist users/repos
|
||||
- **Resolution:** d6ee owns prevention, 1f4f owns detection, 2cdc owns manual response (deferred)
|
||||
|
||||
**Configuration & Monitoring:**
|
||||
- **7d0b (Management Dashboard)** - Display config and metrics
|
||||
- **2cdc (NIP-86 Management API)** - Modify config via API
|
||||
- **76fe (Repository Count Metric)** - Specific metric implementation
|
||||
- **Resolution:** 7d0b owns display, 2cdc owns modification (deferred), 76fe is complete
|
||||
|
||||
### Coordination Requirements
|
||||
|
||||
**Before starting work on any overlapping issue:**
|
||||
1. Check this strategy issue (ec1f) for current phase
|
||||
2. Verify dependencies are complete
|
||||
3. Coordinate scope with related issues
|
||||
4. Update this issue with progress
|
||||
|
||||
## Implementation Sequence
|
||||
|
||||
```
|
||||
Phase 1: Observability
|
||||
├── Update Prometheus/Grafana setup (1-2 days)
|
||||
├── Add storage metrics (2-3 days)
|
||||
└── Add attention indicators (1-2 days)
|
||||
Total: ~1 week
|
||||
|
||||
Phase 2: Insights
|
||||
├── Management Dashboard UI (3-5 days)
|
||||
└── Configuration flag integration (1 day)
|
||||
Total: ~1 week
|
||||
Depends on: Phase 1
|
||||
|
||||
Phase 3: Actions
|
||||
├── Simple HTTP management API (3-4 days)
|
||||
├── Dashboard action integration (2-3 days)
|
||||
└── Local config file persistence (2-3 days)
|
||||
Total: ~1.5 weeks
|
||||
Depends on: Phase 2
|
||||
|
||||
Phase 4: Enforcement
|
||||
├── Storage limits (issue 8430) (1-2 weeks)
|
||||
├── Defensive features Phase 1 (issue d6ee) (2-3 hours)
|
||||
└── Naughty list detection (issue 1f4f) (1 week)
|
||||
Total: ~2-3 weeks
|
||||
Depends on: Phase 3
|
||||
```
|
||||
|
||||
**Total estimated time:** 5-7 weeks for complete administrator experience
|
||||
|
||||
## Success Criteria
|
||||
|
||||
### Phase 1 Complete
|
||||
- [ ] Prometheus metrics endpoint tested and documented
|
||||
- [ ] Grafana dashboard loads without errors
|
||||
- [ ] All metrics display current values
|
||||
- [ ] Storage metrics available (disk usage, per-user, per-repo)
|
||||
- [ ] Attention indicators highlight problems
|
||||
|
||||
### Phase 2 Complete
|
||||
- [ ] Management dashboard accessible via HTTP endpoint
|
||||
- [ ] Dashboard displays all metrics with visualizations
|
||||
- [ ] Dashboard respects ALLOW_DYNAMIC_CONFIG flag
|
||||
- [ ] Mobile-friendly responsive design
|
||||
- [ ] No significant performance impact on relay
|
||||
|
||||
### Phase 3 Complete
|
||||
- [ ] Authenticated management API functional
|
||||
- [ ] Can blacklist users and repositories
|
||||
- [ ] Can set storage quotas
|
||||
- [ ] Can delete repositories with confirmation
|
||||
- [ ] Audit log tracks all admin actions
|
||||
- [ ] Local config file persists changes (when enabled)
|
||||
|
||||
### Phase 4 Complete
|
||||
- [ ] Storage quotas enforced automatically
|
||||
- [ ] Graceful degradation when approaching limits
|
||||
- [ ] Rate limiting prevents DoS attacks
|
||||
- [ ] Abuse detection flags problematic users
|
||||
- [ ] Clear error messages guide users
|
||||
|
||||
### Overall Success
|
||||
- [ ] Administrator can monitor relay health without SSH access
|
||||
- [ ] Administrator can take corrective action through dashboard
|
||||
- [ ] Declarative deployments remain pure (NixOS, Docker)
|
||||
- [ ] Dynamic deployments support GUI-based management
|
||||
- [ ] No manual intervention needed for common abuse scenarios
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. **Config file format:** TOML, JSON, or YAML for local config file?
|
||||
- Recommendation: TOML (matches Rust ecosystem, human-friendly)
|
||||
|
||||
2. **Authentication for management API:** Basic auth, API key, or Nostr signature?
|
||||
- Recommendation: Start with API key (simple), migrate to Nostr signatures later
|
||||
|
||||
3. **Dashboard framework:** Server-rendered HTML or SPA?
|
||||
- Recommendation: Server-rendered (htmx or Alpine.js) for simplicity
|
||||
|
||||
4. **NIP-86 timeline:** When to implement full NIP-86 compliance?
|
||||
- Recommendation: After Phase 3 proves the simple API works
|
||||
|
||||
5. **Prometheus vs. custom metrics:** Should we add non-Prometheus metrics for dashboard?
|
||||
- Recommendation: Prometheus only (single source of truth)
|
||||
|
||||
## Related Issues
|
||||
|
||||
**Directly related (must coordinate):**
|
||||
- 7d0b (Management Dashboard) - Phase 2 primary issue
|
||||
- 2cdc (NIP-86 Relay Management API) - Deferred to post-Phase 3
|
||||
- 8430 (Storage Limits and Quota Management) - Phase 4 primary issue
|
||||
- d6ee (Defensive Relay Features) - Phase 4 (Phase 1 only)
|
||||
- 1f4f (Poor Naughty List Identification) - Phase 4
|
||||
|
||||
**Indirectly related:**
|
||||
- 76fe (Repository Count Metric) - Already implemented, verify in Phase 1
|
||||
- cc36 (Database Backend Evaluation) - May affect storage metrics
|
||||
- b905 (Deletion Request Support) - May integrate with delete repository action
|
||||
|
||||
## Progress
|
||||
|
||||
### 2026-01-14 [Session 16:30]
|
||||
- Created strategy issue
|
||||
- Analyzed existing issues for overlap
|
||||
- Defined configuration model (DECLARATIVE_ONLY, ALLOW_DYNAMIC_CONFIG)
|
||||
- Established 4-phase roadmap
|
||||
- Identified decision points (NIP-86 vs. simple API)
|
||||
- Next: Update overlapping issues with reference to this strategy
|
||||
Reference in New Issue
Block a user