issue: update 820a - baseline collection started with enhanced diagnostics

This commit is contained in:
DanConwayDev
2026-01-17 23:49:45 +00:00
parent 2b4ac4e07b
commit 781fcccfe9
+33 -6
View File
@@ -56,16 +56,18 @@ relay.ngit.dev currently runs ngit-relay (reference implementation). We want to
**Timeline:** 3-5 days total (vs 3-5 weeks original)
### Day 0-1: Minimal Baseline Collection (48-72 hours)
### Day 0-1: Minimal Baseline Collection (48-72 hours) - IN PROGRESS ⏳
**Goal:** Capture enough baseline data to detect major regressions
**Status:** Collection started 2026-01-17 23:30 UTC, running for 36 hours
**Minimum metrics needed:**
- [ ] CPU usage patterns (1 full day cycle minimum)
- [ ] Memory baseline (current RSS, growth rate)
- [ ] Peak connection counts (identify daily peak hours)
- [ ] Basic response time percentiles (p50, p95, p99)
- [ ] Error rate baseline (if any)
- [x] CPU usage patterns (1 full day cycle minimum) - COLLECTING
- [x] Memory baseline (current RSS, growth rate) - COLLECTING
- [x] Peak connection counts (identify daily peak hours) - COLLECTING
- [x] Basic response time percentiles (p50, p95, p99) - Via connection stats
- [x] Error rate baseline (if any) - Via systemd logs
**Quick setup:**
```bash
@@ -182,6 +184,31 @@ docker start ngit-relay
- Risk mitigation: Keep ngit-relay container for instant rollback
- Next: Begin minimal baseline collection (48-72h)
### 2026-01-17 [Session 23:45] - Baseline Collection Started
- Completed: Enhanced diagnostics deployed to all three relays
- Improvements:
1. ✅ Fixed connection counting (port-based, not PID-based)
2. ✅ Added container stats collection (CPU, memory, network I/O)
3. ✅ Added connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
4. ✅ Created comprehensive MONITORING.md documentation
- Baseline Collection Status:
- **Started:** 2026-01-17 23:30 UTC
- **Duration:** 36 hours (until 2026-01-19 11:30 UTC)
- **Relays monitored:** gitnostr.com, relay.ngit.dev, ngit.danconwaydev.com
- **Enhanced metrics:** Connection counts per service, container stats, state breakdown
- **Purpose:** Establish baseline before relay.ngit.dev migration
- Timeline Update:
- **Day 0 (2026-01-17 23:30):** Baseline collection started
- **Day 0+10min (2026-01-17 23:40):** Verify data quality
- **Day 1.5 (2026-01-19 11:30):** Review baseline data (36 hours)
- **Day 2:** Rapid preparation (if baseline looks good)
- **Day 3:** Migration execution
- **Day 3-5:** Intensive monitoring
- Next Steps:
1. Check baseline at 10-minute mark (verify collection working)
2. Review baseline after 36 hours
3. Make migration decision based on baseline data quality
## Notes
- **Related issue:** 573b-production-timeout-diagnosis.md (timeout investigation that led to this)