diff --git a/820a-relay-ngit-dev-migration.md b/820a-relay-ngit-dev-migration.md index 59ff19a..89bc960 100644 --- a/820a-relay-ngit-dev-migration.md +++ b/820a-relay-ngit-dev-migration.md @@ -56,18 +56,24 @@ relay.ngit.dev currently runs ngit-relay (reference implementation). We want to **Timeline:** 3-5 days total (vs 3-5 weeks original) -### Day 0-1: Minimal Baseline Collection (48-72 hours) - IN PROGRESS ⏳ +### Day 0-3: Minimal Baseline Collection (2-3 days) - IN PROGRESS ⏳ **Goal:** Capture enough baseline data to detect major regressions -**Status:** Collection started 2026-01-17 23:30 UTC, running for 36 hours +**Status:** Collection started 2026-01-17 23:30 UTC, extended to 2-3 days (until 2026-01-20/21) **Minimum metrics needed:** -- [x] CPU usage patterns (1 full day cycle minimum) - COLLECTING -- [x] Memory baseline (current RSS, growth rate) - COLLECTING -- [x] Peak connection counts (identify daily peak hours) - COLLECTING -- [x] Basic response time percentiles (p50, p95, p99) - Via connection stats -- [x] Error rate baseline (if any) - Via systemd logs +- [x] CPU usage patterns (1 full day cycle minimum) - COLLECTING ✅ +- [x] Memory baseline (current RSS, growth rate) - COLLECTING ✅ +- [x] Peak connection counts (identify daily peak hours) - COLLECTING ✅ +- [x] Basic response time percentiles (p50, p95, p99) - Via connection stats ✅ +- [x] Error rate baseline (if any) - Via systemd logs ✅ + +**Metrics Status (2026-01-18):** +- ✅ Critical metrics working: connections, CPU, memory, load average +- ❌ Container stats broken (cgroup access issues) - NOT CRITICAL +- ❌ FD counts broken (permission issues) - NOT CRITICAL +- Decision: Proceed with working metrics, sufficient for migration decision **Quick setup:** ```bash @@ -86,12 +92,12 @@ ss -s # socket statistics ``` **Acceptance criteria:** -- ✅ 48+ hours of continuous operation data -- ✅ Identified peak usage hours -- ✅ No critical errors in baseline period -- ✅ Documented current resource usage +- [ ] 2-3 days of continuous operation data (IN PROGRESS - ~10.5h collected) +- [ ] Identified peak usage hours (collecting data) +- [ ] No critical errors in baseline period (monitoring) +- [ ] Documented current resource usage (ongoing) -### Day 2: Rapid Preparation (4-6 hours) +### Day 3-4: Rapid Preparation (4-6 hours) **Configuration review:** - [ ] Map ngit-relay config → ngit-grasp config @@ -117,7 +123,7 @@ ls -la /persistent/ngit-relay/ - Document exact `docker start` command - 5-minute rollback window if issues detected -### Day 3: Migration Execution (2-4 hours) +### Day 4-5: Migration Execution (2-4 hours) **Execution window:** Off-peak hours (based on Day 0-1 data) @@ -143,7 +149,7 @@ docker start ngit-relay # Total rollback time: <5 minutes ``` -### Day 3-5: Intensive Monitoring (48 hours) +### Day 5-7: Intensive Monitoring (48 hours) **Critical monitoring period:** - [ ] Hour 1: Check every 15 minutes @@ -209,6 +215,41 @@ docker start ngit-relay 2. Review baseline after 36 hours 3. Make migration decision based on baseline data quality +### 2026-01-18 [Session 10:00] - Baseline Collection Status Update +- Completed: 2 iterations of diagnostics fixes to improve metric reliability +- **Critical Metrics Working:** + - ✅ Connection counts per service (port-based detection working) + - ✅ System metrics (load average, memory, CPU breakdown) + - ✅ Top processes (identifying resource consumers) + - ✅ Connection state breakdown (ESTABLISHED, TIME_WAIT, etc.) + - ✅ Basic container visibility (containers appear in top processes) +- **Non-Critical Metrics Still Broken:** + - ❌ Container stats (cgroup access issues, requires privileged systemd-run) + - ❌ File descriptor counts per process (permission issues) + - Note: Container resource usage still visible in top processes section +- **Decision: Proceed with Current Metrics** + - Critical metrics (connections, CPU, memory) working reliably + - Broken metrics are nice-to-have, not essential for migration decision + - Have sufficient data to detect major regressions + - Can proceed with migration planning based on working metrics +- **Baseline Collection Extended:** + - Current duration: ~10.5 hours (started 2026-01-17 23:30 UTC) + - Extended to: 2-3 more days (until 2026-01-20 or 2026-01-21) + - Reason: Want more data to establish reliable daily patterns + - Collection continuing with working metrics +- **Updated Timeline:** + - **Day 0 (2026-01-17 23:30):** Baseline collection started + - **Day 0.5 (2026-01-18 10:00):** Diagnostics fixes complete, extending collection + - **Day 2-3 (2026-01-20/21):** Review baseline data (2-3 days total) + - **Day 3-4:** Rapid preparation (if baseline shows stable patterns) + - **Day 4-5:** Migration execution + - **Day 5-7:** Intensive monitoring +- **Next Steps:** + 1. Continue baseline collection for 2-3 more days + 2. Review baseline data patterns after collection period + 3. Make migration decision based on baseline analysis + 4. Proceed with rapid preparation if baseline looks good + ## Notes - **Related issue:** 573b-production-timeout-diagnosis.md (timeout investigation that led to this)