issue: update 820a - baseline collection extended, critical metrics working

This commit is contained in:
DanConwayDev
2026-01-18 13:12:53 +00:00
parent a97a9f2255
commit 775b6caa6a
+55 -14
View File
@@ -56,18 +56,24 @@ relay.ngit.dev currently runs ngit-relay (reference implementation). We want to
**Timeline:** 3-5 days total (vs 3-5 weeks original) **Timeline:** 3-5 days total (vs 3-5 weeks original)
### Day 0-1: Minimal Baseline Collection (48-72 hours) - IN PROGRESS ⏳ ### Day 0-3: Minimal Baseline Collection (2-3 days) - IN PROGRESS ⏳
**Goal:** Capture enough baseline data to detect major regressions **Goal:** Capture enough baseline data to detect major regressions
**Status:** Collection started 2026-01-17 23:30 UTC, running for 36 hours **Status:** Collection started 2026-01-17 23:30 UTC, extended to 2-3 days (until 2026-01-20/21)
**Minimum metrics needed:** **Minimum metrics needed:**
- [x] CPU usage patterns (1 full day cycle minimum) - COLLECTING - [x] CPU usage patterns (1 full day cycle minimum) - COLLECTING ✅
- [x] Memory baseline (current RSS, growth rate) - COLLECTING - [x] Memory baseline (current RSS, growth rate) - COLLECTING ✅
- [x] Peak connection counts (identify daily peak hours) - COLLECTING - [x] Peak connection counts (identify daily peak hours) - COLLECTING ✅
- [x] Basic response time percentiles (p50, p95, p99) - Via connection stats - [x] Basic response time percentiles (p50, p95, p99) - Via connection stats ✅
- [x] Error rate baseline (if any) - Via systemd logs - [x] Error rate baseline (if any) - Via systemd logs ✅
**Metrics Status (2026-01-18):**
- ✅ Critical metrics working: connections, CPU, memory, load average
- ❌ Container stats broken (cgroup access issues) - NOT CRITICAL
- ❌ FD counts broken (permission issues) - NOT CRITICAL
- Decision: Proceed with working metrics, sufficient for migration decision
**Quick setup:** **Quick setup:**
```bash ```bash
@@ -86,12 +92,12 @@ ss -s # socket statistics
``` ```
**Acceptance criteria:** **Acceptance criteria:**
- ✅ 48+ hours of continuous operation data - [ ] 2-3 days of continuous operation data (IN PROGRESS - ~10.5h collected)
- ✅ Identified peak usage hours - [ ] Identified peak usage hours (collecting data)
- ✅ No critical errors in baseline period - [ ] No critical errors in baseline period (monitoring)
- ✅ Documented current resource usage - [ ] Documented current resource usage (ongoing)
### Day 2: Rapid Preparation (4-6 hours) ### Day 3-4: Rapid Preparation (4-6 hours)
**Configuration review:** **Configuration review:**
- [ ] Map ngit-relay config → ngit-grasp config - [ ] Map ngit-relay config → ngit-grasp config
@@ -117,7 +123,7 @@ ls -la /persistent/ngit-relay/
- Document exact `docker start` command - Document exact `docker start` command
- 5-minute rollback window if issues detected - 5-minute rollback window if issues detected
### Day 3: Migration Execution (2-4 hours) ### Day 4-5: Migration Execution (2-4 hours)
**Execution window:** Off-peak hours (based on Day 0-1 data) **Execution window:** Off-peak hours (based on Day 0-1 data)
@@ -143,7 +149,7 @@ docker start ngit-relay
# Total rollback time: <5 minutes # Total rollback time: <5 minutes
``` ```
### Day 3-5: Intensive Monitoring (48 hours) ### Day 5-7: Intensive Monitoring (48 hours)
**Critical monitoring period:** **Critical monitoring period:**
- [ ] Hour 1: Check every 15 minutes - [ ] Hour 1: Check every 15 minutes
@@ -209,6 +215,41 @@ docker start ngit-relay
2. Review baseline after 36 hours 2. Review baseline after 36 hours
3. Make migration decision based on baseline data quality 3. Make migration decision based on baseline data quality
### 2026-01-18 [Session 10:00] - Baseline Collection Status Update
- Completed: 2 iterations of diagnostics fixes to improve metric reliability
- **Critical Metrics Working:**
- ✅ Connection counts per service (port-based detection working)
- ✅ System metrics (load average, memory, CPU breakdown)
- ✅ Top processes (identifying resource consumers)
- ✅ Connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
- ✅ Basic container visibility (containers appear in top processes)
- **Non-Critical Metrics Still Broken:**
- ❌ Container stats (cgroup access issues, requires privileged systemd-run)
- ❌ File descriptor counts per process (permission issues)
- Note: Container resource usage still visible in top processes section
- **Decision: Proceed with Current Metrics**
- Critical metrics (connections, CPU, memory) working reliably
- Broken metrics are nice-to-have, not essential for migration decision
- Have sufficient data to detect major regressions
- Can proceed with migration planning based on working metrics
- **Baseline Collection Extended:**
- Current duration: ~10.5 hours (started 2026-01-17 23:30 UTC)
- Extended to: 2-3 more days (until 2026-01-20 or 2026-01-21)
- Reason: Want more data to establish reliable daily patterns
- Collection continuing with working metrics
- **Updated Timeline:**
- **Day 0 (2026-01-17 23:30):** Baseline collection started
- **Day 0.5 (2026-01-18 10:00):** Diagnostics fixes complete, extending collection
- **Day 2-3 (2026-01-20/21):** Review baseline data (2-3 days total)
- **Day 3-4:** Rapid preparation (if baseline shows stable patterns)
- **Day 4-5:** Migration execution
- **Day 5-7:** Intensive monitoring
- **Next Steps:**
1. Continue baseline collection for 2-3 more days
2. Review baseline data patterns after collection period
3. Make migration decision based on baseline analysis
4. Proceed with rapid preparation if baseline looks good
## Notes ## Notes
- **Related issue:** 573b-production-timeout-diagnosis.md (timeout investigation that led to this) - **Related issue:** 573b-production-timeout-diagnosis.md (timeout investigation that led to this)