mirror of
https://relay.ngit.dev/npub15qydau2hjma6ngxkl2cyar74wzyjshvl65za5k5rl69264ar2exs5cyejr/ngit-grasp.git
synced 2026-10-05 15:08:24 +00:00
issue: update 820a - baseline collection extended, critical metrics working
This commit is contained in:
@@ -56,18 +56,24 @@ relay.ngit.dev currently runs ngit-relay (reference implementation). We want to
|
|||||||
|
|
||||||
**Timeline:** 3-5 days total (vs 3-5 weeks original)
|
**Timeline:** 3-5 days total (vs 3-5 weeks original)
|
||||||
|
|
||||||
### Day 0-1: Minimal Baseline Collection (48-72 hours) - IN PROGRESS ⏳
|
### Day 0-3: Minimal Baseline Collection (2-3 days) - IN PROGRESS ⏳
|
||||||
|
|
||||||
**Goal:** Capture enough baseline data to detect major regressions
|
**Goal:** Capture enough baseline data to detect major regressions
|
||||||
|
|
||||||
**Status:** Collection started 2026-01-17 23:30 UTC, running for 36 hours
|
**Status:** Collection started 2026-01-17 23:30 UTC, extended to 2-3 days (until 2026-01-20/21)
|
||||||
|
|
||||||
**Minimum metrics needed:**
|
**Minimum metrics needed:**
|
||||||
- [x] CPU usage patterns (1 full day cycle minimum) - COLLECTING
|
- [x] CPU usage patterns (1 full day cycle minimum) - COLLECTING ✅
|
||||||
- [x] Memory baseline (current RSS, growth rate) - COLLECTING
|
- [x] Memory baseline (current RSS, growth rate) - COLLECTING ✅
|
||||||
- [x] Peak connection counts (identify daily peak hours) - COLLECTING
|
- [x] Peak connection counts (identify daily peak hours) - COLLECTING ✅
|
||||||
- [x] Basic response time percentiles (p50, p95, p99) - Via connection stats
|
- [x] Basic response time percentiles (p50, p95, p99) - Via connection stats ✅
|
||||||
- [x] Error rate baseline (if any) - Via systemd logs
|
- [x] Error rate baseline (if any) - Via systemd logs ✅
|
||||||
|
|
||||||
|
**Metrics Status (2026-01-18):**
|
||||||
|
- ✅ Critical metrics working: connections, CPU, memory, load average
|
||||||
|
- ❌ Container stats broken (cgroup access issues) - NOT CRITICAL
|
||||||
|
- ❌ FD counts broken (permission issues) - NOT CRITICAL
|
||||||
|
- Decision: Proceed with working metrics, sufficient for migration decision
|
||||||
|
|
||||||
**Quick setup:**
|
**Quick setup:**
|
||||||
```bash
|
```bash
|
||||||
@@ -86,12 +92,12 @@ ss -s # socket statistics
|
|||||||
```
|
```
|
||||||
|
|
||||||
**Acceptance criteria:**
|
**Acceptance criteria:**
|
||||||
- ✅ 48+ hours of continuous operation data
|
- [ ] 2-3 days of continuous operation data (IN PROGRESS - ~10.5h collected)
|
||||||
- ✅ Identified peak usage hours
|
- [ ] Identified peak usage hours (collecting data)
|
||||||
- ✅ No critical errors in baseline period
|
- [ ] No critical errors in baseline period (monitoring)
|
||||||
- ✅ Documented current resource usage
|
- [ ] Documented current resource usage (ongoing)
|
||||||
|
|
||||||
### Day 2: Rapid Preparation (4-6 hours)
|
### Day 3-4: Rapid Preparation (4-6 hours)
|
||||||
|
|
||||||
**Configuration review:**
|
**Configuration review:**
|
||||||
- [ ] Map ngit-relay config → ngit-grasp config
|
- [ ] Map ngit-relay config → ngit-grasp config
|
||||||
@@ -117,7 +123,7 @@ ls -la /persistent/ngit-relay/
|
|||||||
- Document exact `docker start` command
|
- Document exact `docker start` command
|
||||||
- 5-minute rollback window if issues detected
|
- 5-minute rollback window if issues detected
|
||||||
|
|
||||||
### Day 3: Migration Execution (2-4 hours)
|
### Day 4-5: Migration Execution (2-4 hours)
|
||||||
|
|
||||||
**Execution window:** Off-peak hours (based on Day 0-1 data)
|
**Execution window:** Off-peak hours (based on Day 0-1 data)
|
||||||
|
|
||||||
@@ -143,7 +149,7 @@ docker start ngit-relay
|
|||||||
# Total rollback time: <5 minutes
|
# Total rollback time: <5 minutes
|
||||||
```
|
```
|
||||||
|
|
||||||
### Day 3-5: Intensive Monitoring (48 hours)
|
### Day 5-7: Intensive Monitoring (48 hours)
|
||||||
|
|
||||||
**Critical monitoring period:**
|
**Critical monitoring period:**
|
||||||
- [ ] Hour 1: Check every 15 minutes
|
- [ ] Hour 1: Check every 15 minutes
|
||||||
@@ -209,6 +215,41 @@ docker start ngit-relay
|
|||||||
2. Review baseline after 36 hours
|
2. Review baseline after 36 hours
|
||||||
3. Make migration decision based on baseline data quality
|
3. Make migration decision based on baseline data quality
|
||||||
|
|
||||||
|
### 2026-01-18 [Session 10:00] - Baseline Collection Status Update
|
||||||
|
- Completed: 2 iterations of diagnostics fixes to improve metric reliability
|
||||||
|
- **Critical Metrics Working:**
|
||||||
|
- ✅ Connection counts per service (port-based detection working)
|
||||||
|
- ✅ System metrics (load average, memory, CPU breakdown)
|
||||||
|
- ✅ Top processes (identifying resource consumers)
|
||||||
|
- ✅ Connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
|
||||||
|
- ✅ Basic container visibility (containers appear in top processes)
|
||||||
|
- **Non-Critical Metrics Still Broken:**
|
||||||
|
- ❌ Container stats (cgroup access issues, requires privileged systemd-run)
|
||||||
|
- ❌ File descriptor counts per process (permission issues)
|
||||||
|
- Note: Container resource usage still visible in top processes section
|
||||||
|
- **Decision: Proceed with Current Metrics**
|
||||||
|
- Critical metrics (connections, CPU, memory) working reliably
|
||||||
|
- Broken metrics are nice-to-have, not essential for migration decision
|
||||||
|
- Have sufficient data to detect major regressions
|
||||||
|
- Can proceed with migration planning based on working metrics
|
||||||
|
- **Baseline Collection Extended:**
|
||||||
|
- Current duration: ~10.5 hours (started 2026-01-17 23:30 UTC)
|
||||||
|
- Extended to: 2-3 more days (until 2026-01-20 or 2026-01-21)
|
||||||
|
- Reason: Want more data to establish reliable daily patterns
|
||||||
|
- Collection continuing with working metrics
|
||||||
|
- **Updated Timeline:**
|
||||||
|
- **Day 0 (2026-01-17 23:30):** Baseline collection started
|
||||||
|
- **Day 0.5 (2026-01-18 10:00):** Diagnostics fixes complete, extending collection
|
||||||
|
- **Day 2-3 (2026-01-20/21):** Review baseline data (2-3 days total)
|
||||||
|
- **Day 3-4:** Rapid preparation (if baseline shows stable patterns)
|
||||||
|
- **Day 4-5:** Migration execution
|
||||||
|
- **Day 5-7:** Intensive monitoring
|
||||||
|
- **Next Steps:**
|
||||||
|
1. Continue baseline collection for 2-3 more days
|
||||||
|
2. Review baseline data patterns after collection period
|
||||||
|
3. Make migration decision based on baseline analysis
|
||||||
|
4. Proceed with rapid preparation if baseline looks good
|
||||||
|
|
||||||
## Notes
|
## Notes
|
||||||
|
|
||||||
- **Related issue:** 573b-production-timeout-diagnosis.md (timeout investigation that led to this)
|
- **Related issue:** 573b-production-timeout-diagnosis.md (timeout investigation that led to this)
|
||||||
|
|||||||
Reference in New Issue
Block a user