issue: update 820a - baseline collection extended, critical metrics working

This commit is contained in:
DanConwayDev
2026-01-18 13:12:53 +00:00
parent a97a9f2255
commit 775b6caa6a
+55 -14
View File
@@ -56,18 +56,24 @@ relay.ngit.dev currently runs ngit-relay (reference implementation). We want to
**Timeline:** 3-5 days total (vs 3-5 weeks original)
### Day 0-1: Minimal Baseline Collection (48-72 hours) - IN PROGRESS ⏳
### Day 0-3: Minimal Baseline Collection (2-3 days) - IN PROGRESS ⏳
**Goal:** Capture enough baseline data to detect major regressions
**Status:** Collection started 2026-01-17 23:30 UTC, running for 36 hours
**Status:** Collection started 2026-01-17 23:30 UTC, extended to 2-3 days (until 2026-01-20/21)
**Minimum metrics needed:**
- [x] CPU usage patterns (1 full day cycle minimum) - COLLECTING
- [x] Memory baseline (current RSS, growth rate) - COLLECTING
- [x] Peak connection counts (identify daily peak hours) - COLLECTING
- [x] Basic response time percentiles (p50, p95, p99) - Via connection stats
- [x] Error rate baseline (if any) - Via systemd logs
- [x] CPU usage patterns (1 full day cycle minimum) - COLLECTING ✅
- [x] Memory baseline (current RSS, growth rate) - COLLECTING ✅
- [x] Peak connection counts (identify daily peak hours) - COLLECTING ✅
- [x] Basic response time percentiles (p50, p95, p99) - Via connection stats ✅
- [x] Error rate baseline (if any) - Via systemd logs ✅
**Metrics Status (2026-01-18):**
- ✅ Critical metrics working: connections, CPU, memory, load average
- ❌ Container stats broken (cgroup access issues) - NOT CRITICAL
- ❌ FD counts broken (permission issues) - NOT CRITICAL
- Decision: Proceed with working metrics, sufficient for migration decision
**Quick setup:**
```bash
@@ -86,12 +92,12 @@ ss -s # socket statistics
```
**Acceptance criteria:**
- ✅ 48+ hours of continuous operation data
- ✅ Identified peak usage hours
- ✅ No critical errors in baseline period
- ✅ Documented current resource usage
- [ ] 2-3 days of continuous operation data (IN PROGRESS - ~10.5h collected)
- [ ] Identified peak usage hours (collecting data)
- [ ] No critical errors in baseline period (monitoring)
- [ ] Documented current resource usage (ongoing)
### Day 2: Rapid Preparation (4-6 hours)
### Day 3-4: Rapid Preparation (4-6 hours)
**Configuration review:**
- [ ] Map ngit-relay config → ngit-grasp config
@@ -117,7 +123,7 @@ ls -la /persistent/ngit-relay/
- Document exact `docker start` command
- 5-minute rollback window if issues detected
### Day 3: Migration Execution (2-4 hours)
### Day 4-5: Migration Execution (2-4 hours)
**Execution window:** Off-peak hours (based on Day 0-1 data)
@@ -143,7 +149,7 @@ docker start ngit-relay
# Total rollback time: <5 minutes
```
### Day 3-5: Intensive Monitoring (48 hours)
### Day 5-7: Intensive Monitoring (48 hours)
**Critical monitoring period:**
- [ ] Hour 1: Check every 15 minutes
@@ -209,6 +215,41 @@ docker start ngit-relay
2. Review baseline after 36 hours
3. Make migration decision based on baseline data quality
### 2026-01-18 [Session 10:00] - Baseline Collection Status Update
- Completed: 2 iterations of diagnostics fixes to improve metric reliability
- **Critical Metrics Working:**
- ✅ Connection counts per service (port-based detection working)
- ✅ System metrics (load average, memory, CPU breakdown)
- ✅ Top processes (identifying resource consumers)
- ✅ Connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
- ✅ Basic container visibility (containers appear in top processes)
- **Non-Critical Metrics Still Broken:**
- ❌ Container stats (cgroup access issues, requires privileged systemd-run)
- ❌ File descriptor counts per process (permission issues)
- Note: Container resource usage still visible in top processes section
- **Decision: Proceed with Current Metrics**
- Critical metrics (connections, CPU, memory) working reliably
- Broken metrics are nice-to-have, not essential for migration decision
- Have sufficient data to detect major regressions
- Can proceed with migration planning based on working metrics
- **Baseline Collection Extended:**
- Current duration: ~10.5 hours (started 2026-01-17 23:30 UTC)
- Extended to: 2-3 more days (until 2026-01-20 or 2026-01-21)
- Reason: Want more data to establish reliable daily patterns
- Collection continuing with working metrics
- **Updated Timeline:**
- **Day 0 (2026-01-17 23:30):** Baseline collection started
- **Day 0.5 (2026-01-18 10:00):** Diagnostics fixes complete, extending collection
- **Day 2-3 (2026-01-20/21):** Review baseline data (2-3 days total)
- **Day 3-4:** Rapid preparation (if baseline shows stable patterns)
- **Day 4-5:** Migration execution
- **Day 5-7:** Intensive monitoring
- **Next Steps:**
1. Continue baseline collection for 2-3 more days
2. Review baseline data patterns after collection period
3. Make migration decision based on baseline analysis
4. Proceed with rapid preparation if baseline looks good
## Notes
- **Related issue:** 573b-production-timeout-diagnosis.md (timeout investigation that led to this)