Files
ngit-grasp/820a-relay-ngit-dev-migration.md
T

456 lines
19 KiB
Markdown

# Migrate relay.ngit.dev from ngit-relay to ngit-grasp
**ID:** 820a
## Problem
relay.ngit.dev currently runs ngit-relay (reference implementation). We want to consolidate on ngit-grasp as the production implementation to:
- Reduce operational complexity (one codebase instead of two)
- Focus logging and observability improvements on a single project
- Leverage ngit-grasp's proven performance on ngit.danconwaydev.com
**Why this matters:** Running two different implementations increases maintenance burden and makes it harder to add features like comprehensive logging and performance monitoring. Consolidating on ngit-grasp enables:
- Unified logging and metrics infrastructure
- Single codebase for bug fixes and features
- Consistent behavior across all deployments
- Reduced testing and deployment complexity
## Plan
- [ ] **Phase 1: Baseline Metrics Collection** (1-2 weeks)
- [ ] Deploy nginx metrics to relay.ngit.dev (implementation ready in ngit-relay repo)
- [ ] Collect baseline metrics for 1-2 weeks:
- CPU usage (user, system, idle)
- Memory usage (RSS, available)
- Connection counts (established, TIME_WAIT, rate)
- Git operation frequency and duration
- Request rates and response times
- [ ] Document baseline performance characteristics
- [ ] Create performance comparison criteria
- [ ] **Phase 2: Migration Preparation** (1 week)
- [ ] Review ngit-grasp configuration options
- [ ] Plan NixOS service configuration for relay.ngit.dev
- [ ] Prepare rollback plan
- [ ] Test ngit-grasp with relay.ngit.dev's data (if possible)
- [ ] Document migration steps and verification procedures
- [ ] Identify monitoring checkpoints
- [ ] **Phase 3: Migration Execution** (1 day)
- [ ] Create new ngit-grasp service config for relay.ngit.dev
- [ ] Deploy to VPS
- [ ] Verify service starts and accepts connections
- [ ] Monitor for immediate issues
- [ ] Keep ngit-relay container available for quick rollback
- [ ] Test basic functionality (websocket, git operations)
- [ ] **Phase 4: Post-Migration Validation** (1-2 weeks)
- [ ] Collect same metrics as baseline
- [ ] Compare performance (CPU, memory, connections, response times)
- [ ] Verify no new timeout patterns
- [ ] Monitor error rates and connection stability
- [ ] User acceptance testing
- [ ] Remove old ngit-relay container if successful
## Accelerated Migration Plan
**Timeline:** 3-5 days total (vs 3-5 weeks original)
### Day 0-3: Minimal Baseline Collection (2-3 days) - IN PROGRESS ⏳
**Goal:** Capture enough baseline data to detect major regressions
**Status:** Collection started 2026-01-17 23:30 UTC, extended to 2-3 days (until 2026-01-20/21)
**Minimum metrics needed:**
- [x] CPU usage patterns (1 full day cycle minimum) - COLLECTING ✅
- [x] Memory baseline (current RSS, growth rate) - COLLECTING ✅
- [x] Peak connection counts (identify daily peak hours) - COLLECTING ✅
- [x] Basic response time percentiles (p50, p95, p99) - Via connection stats ✅
- [x] Error rate baseline (if any) - Via systemd logs ✅
**Metrics Status (2026-01-18):**
- ✅ Critical metrics working: connections, CPU, memory, load average
- ❌ Container stats broken (cgroup access issues) - NOT CRITICAL
- ❌ FD counts broken (permission issues) - NOT CRITICAL
- Decision: Proceed with working metrics, sufficient for migration decision
**Quick setup:**
```bash
# On relay.ngit.dev VPS
# 1. Check current metrics via systemd/journalctl
systemctl status ngit-relay
journalctl -u ngit-relay --since "24 hours ago" | grep -i error
# 2. Capture system metrics snapshot
top -b -n 1 | head -20
free -h
ss -s # socket statistics
# 3. If nginx metrics available, capture 24h sample
# Otherwise, rely on system metrics + manual testing
```
**Acceptance criteria:**
- [ ] 2-3 days of continuous operation data (IN PROGRESS - ~10.5h collected)
- [ ] Identified peak usage hours (collecting data)
- [ ] No critical errors in baseline period (monitoring)
- [ ] Documented current resource usage (ongoing)
### Day 3-4: Rapid Preparation (4-6 hours)
**Configuration review:**
- [ ] Map ngit-relay config → ngit-grasp config
- [ ] Prepare NixOS service definition
- [ ] Document rollback procedure (keep ngit-relay container)
- [ ] Pre-build ngit-grasp on VPS (cache dependencies)
**Quick prep checklist:**
```bash
# 1. Review current ngit-relay config
docker inspect ngit-relay | jq '.[0].Config.Env'
# 2. Create ngit-grasp service config (see template below)
# 3. Test build on VPS
nix build github:DanConwayDev/ngit-grasp
# 4. Verify data directory permissions
ls -la /persistent/ngit-relay/
```
**Rollback plan:**
- Keep ngit-relay container stopped but available
- Document exact `docker start` command
- 5-minute rollback window if issues detected
### Day 4-5: Migration Execution (2-4 hours)
**Execution window:** Off-peak hours (based on Day 0-1 data)
**Steps:**
1. [ ] Stop ngit-relay container (don't remove)
2. [ ] Deploy ngit-grasp via NixOS
3. [ ] Verify service starts and binds to port
4. [ ] Test WebSocket connection
5. [ ] Test git clone operation
6. [ ] Monitor for 30 minutes (critical window)
**Go/No-Go criteria (30 min checkpoint):**
- ✅ Service running and accepting connections
- ✅ No error spikes in logs
- ✅ Memory usage within 2x of baseline
- ✅ CPU usage reasonable (<80% sustained)
- ✅ Git operations functional
**If No-Go:** Rollback immediately
```bash
systemctl stop ngit-grasp-production
docker start ngit-relay
# Total rollback time: <5 minutes
```
### Day 5-7: Intensive Monitoring (48 hours)
**Critical monitoring period:**
- [ ] Hour 1: Check every 15 minutes
- [ ] Hours 2-6: Check every hour
- [ ] Hours 6-24: Check every 4 hours
- [ ] Hours 24-48: Check every 8 hours
**Monitor:**
- CPU/memory trends (compare to baseline)
- Connection counts (should be similar)
- Error logs (any new patterns?)
- Response times (p95, p99)
- Git operation success rate
**Success criteria (48h):**
- ✅ No critical errors
- ✅ Resource usage ≤ 2x baseline (acceptable for new impl)
- ✅ No user complaints
- ✅ Git operations working
- ✅ WebSocket connections stable
**If successful:** Remove ngit-relay container after 7 days
**If issues:** Rollback and investigate
## Progress
### 2026-01-17 [Session 22:00]
- Created: Initial issue based on pivot from 573b timeout diagnosis
- Context: Phase 1 of 573b is complete (connection limits increased, rate limiting fixed)
- Rationale: Better to consolidate on ngit-grasp than implement duplicate features in ngit-relay
- Blocker: Must collect baseline metrics BEFORE migration for objective comparison
- Next: Deploy nginx metrics to relay.ngit.dev and begin baseline collection
### 2026-01-17 [Session 23:30]
- Added: Accelerated migration plan (3-5 days vs 3-5 weeks)
- Decision: User wants to move quickly after minimal baseline collection
- Approach: 48-72h baseline → rapid prep → execute → intensive monitoring
- Risk mitigation: Keep ngit-relay container for instant rollback
- Next: Begin minimal baseline collection (48-72h)
### 2026-01-17 [Session 23:45] - Baseline Collection Started
- Completed: Enhanced diagnostics deployed to all three relays
- Improvements:
1. ✅ Fixed connection counting (port-based, not PID-based)
2. ✅ Added container stats collection (CPU, memory, network I/O)
3. ✅ Added connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
4. ✅ Created comprehensive MONITORING.md documentation
- Baseline Collection Status:
- **Started:** 2026-01-17 23:30 UTC
- **Duration:** 36 hours (until 2026-01-19 11:30 UTC)
- **Relays monitored:** gitnostr.com, relay.ngit.dev, ngit.danconwaydev.com
- **Enhanced metrics:** Connection counts per service, container stats, state breakdown
- **Purpose:** Establish baseline before relay.ngit.dev migration
- Timeline Update:
- **Day 0 (2026-01-17 23:30):** Baseline collection started
- **Day 0+10min (2026-01-17 23:40):** Verify data quality
- **Day 1.5 (2026-01-19 11:30):** Review baseline data (36 hours)
- **Day 2:** Rapid preparation (if baseline looks good)
- **Day 3:** Migration execution
- **Day 3-5:** Intensive monitoring
- Next Steps:
1. Check baseline at 10-minute mark (verify collection working)
2. Review baseline after 36 hours
3. Make migration decision based on baseline data quality
### 2026-01-18 [Session 10:00] - Baseline Collection Status Update
- Completed: 2 iterations of diagnostics fixes to improve metric reliability
- **Critical Metrics Working:**
- ✅ Connection counts per service (port-based detection working)
- ✅ System metrics (load average, memory, CPU breakdown)
- ✅ Top processes (identifying resource consumers)
- ✅ Connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
- ✅ Basic container visibility (containers appear in top processes)
- **Non-Critical Metrics Still Broken:**
- ❌ Container stats (cgroup access issues, requires privileged systemd-run)
- ❌ File descriptor counts per process (permission issues)
- Note: Container resource usage still visible in top processes section
- **Decision: Proceed with Current Metrics**
- Critical metrics (connections, CPU, memory) working reliably
- Broken metrics are nice-to-have, not essential for migration decision
- Have sufficient data to detect major regressions
- Can proceed with migration planning based on working metrics
- **Baseline Collection Extended:**
- Current duration: ~10.5 hours (started 2026-01-17 23:30 UTC)
- Extended to: 2-3 more days (until 2026-01-20 or 2026-01-21)
- Reason: Want more data to establish reliable daily patterns
- Collection continuing with working metrics
- **Updated Timeline:**
- **Day 0 (2026-01-17 23:30):** Baseline collection started
- **Day 0.5 (2026-01-18 10:00):** Diagnostics fixes complete, extending collection
- **Day 2-3 (2026-01-20/21):** Review baseline data (2-3 days total)
- **Day 3-4:** Rapid preparation (if baseline shows stable patterns)
- **Day 4-5:** Migration execution
- **Day 5-7:** Intensive monitoring
- **Next Steps:**
1. Continue baseline collection for 2-3 more days
2. Review baseline data patterns after collection period
3. Make migration decision based on baseline analysis
4. Proceed with rapid preparation if baseline looks good
### 2026-01-19 [Session 11:35] - Baseline Analysis Complete - VPS UPGRADE REQUIRED ⚠️
- Completed: Comprehensive analysis of 36 hours of baseline metrics
- **CRITICAL DECISION: Migration BLOCKED until VPS upgrade**
- **Baseline Findings (2026-01-17 23:30 to 2026-01-19 11:35 UTC):**
- **CPU:** Average 6.96 load (348% utilization), peak 11.83 (592%) - 🔴 SEVERELY OVERLOADED
- **Memory:** 96.8% average usage, only 130MB available - 🔴 MAXED OUT
- **Swap:** 2.0GB average (40% of 5GB), growing 0.75GB → 2.5GB - 🟡 HEAVILY USED
- **Connections:** Port 8081: 28 avg/120 peak, Port 8083: 31 avg/126 peak - ✅ ONLY 3% OF CAPACITY
- **Correlation:** Load vs connections = 0.210 (weak) - Confirms git operations are bottleneck, not connections
- **Peak Usage Patterns:**
- Peak hours: 16:00-20:00 UTC (load 6.5-9.9, connections 865-977)
- Off-peak hours: 00:00-12:00 UTC (load 6.3-7.4, connections 350-470)
- **Critical insight:** Even off-peak shows HIGH load (6.3-7.4), indicating baseline overload independent of traffic
- **Migration Risk Assessment:**
- **Current VPS (2-core, 4GB):** 🔴 **HIGH RISK** - Migration will add load spike to already overloaded system
- **Outcome:** High probability of service degradation or outages during migration
- **Recommendation:** ❌ **DO NOT MIGRATE** on current VPS
- **VPS Upgrade Requirements:**
- **Minimum (Survival):** 4 cores, 8GB RAM (~$20-40/month) - 🟡 MEDIUM RISK
- **Recommended (Healthy):** 6 cores, 8GB RAM (~$40-60/month) - 🟢 LOW RISK ✅
- **Optimal (Future-proof):** 8 cores, 16GB RAM (~$60-100/month) - 🟢 LOW RISK
- **Decision:** Upgrade to 6-core, 8GB minimum before migration
- **Updated Migration Timeline:**
- **REVISED:** 5-7 days total (was 3-5 days)
- **Day 0-1:** VPS upgrade + verification (NEW - CRITICAL)
- **Day 2-3:** Rapid preparation
- **Day 4-5:** Migration execution
- **Day 5-7:** Intensive monitoring
- **Data Quality:**
- ✅ 36 hours continuous collection, 2,160 data points
- ✅ Clear patterns identified, no anomalies
- ✅ **Sufficient for migration decision** - no additional collection needed
- ✅ Findings are conclusive: VPS upgrade required
- **Next Steps:**
1. **CRITICAL:** Upgrade VPS to 6-core, 8GB RAM (BLOCKER for migration)
2. Verify upgrade impact (24-48 hours monitoring, expect load to drop to ~60-70%)
3. Proceed with accelerated migration plan (Day 2-7)
4. Monitor intensively during and after migration
- **Files Created:**
- `/tmp/vps-baseline-report.md` - Comprehensive 36-hour analysis with migration scenarios
- **Status:** Baseline complete ✅, Migration BLOCKED ⚠️ pending VPS upgrade
- **Blocker:** VPS upgrade to 6-core, 8GB RAM required before migration can proceed
## Notes
- **Related issue:** 573b-production-timeout-diagnosis.md (timeout investigation that led to this)
- **Critical requirement:** Must collect baseline BEFORE migration for objective comparison
- **Success criteria:** Performance equal or better than ngit-relay baseline
- No increase in CPU/memory usage
- No new timeout patterns
- Connection handling remains stable
- Git operations perform equally well
- **Rollback plan:** Keep ngit-relay container for 2 weeks post-migration
- Can switch back immediately if issues arise
- Only remove after validation period passes
- **nginx metrics:** Implementation ready in ngit-relay repo
- Provides connection counts, request rates, upstream health
- Essential for before/after comparison
- **Current status:**
- relay.ngit.dev: Running ngit-relay v0.0.5 with 2000 conn/IP rate limit
- ngit.danconwaydev.com: Running ngit-grasp with 4096 max connections
- Both services stable after Phase 1 improvements
- **Migration benefits:**
- Single codebase to maintain and improve
- Can focus logging/observability work on ngit-grasp
- Proven performance (ngit.danconwaydev.com is stable)
- Simplified deployment and configuration
### Accelerated Migration: Configuration Template
**ngit-grasp service config for relay.ngit.dev:**
```nix
# services/ngit-grasp.nix
{ inputs, ... }:
{
imports = [ inputs.ngit-grasp.nixosModules.default ];
services.ngit-grasp.relay-ngit-dev = {
enable = true;
domain = "relay.ngit.dev";
# Network - bind to localhost, nginx handles TLS
bindAddress = "127.0.0.1";
port = 8082; # Or whatever port ngit-relay was using
# Storage - reuse existing data directory if possible
dataDir = "/persistent/ngit-relay"; # Or new dir: /persistent/ngit-grasp
# Identity - copy from ngit-relay config
relayName = "relay.ngit.dev";
relayDescription = "GRASP relay for git+nostr";
relayOwnerNsecFile = "/persistent/ngit-relay/relay-owner.nsec";
# Sync - bootstrap from ngit.danconwaydev.com
syncBootstrapRelayUrl = "wss://ngit.danconwaydev.com";
# Metrics - enable for monitoring
metricsEnabled = true;
# Logging - start with info, increase to debug if issues
logLevel = "info";
};
# Nginx/Caddy reverse proxy (existing config, just update upstream)
# Change upstream from ngit-relay container to 127.0.0.1:8082
}
```
**Key decisions:**
1. **Data directory:** Reuse `/persistent/ngit-relay` if compatible, or create new `/persistent/ngit-grasp`
2. **Port:** Use same port as ngit-relay container was using (check nginx config)
3. **Bootstrap relay:** Use ngit.danconwaydev.com (known good ngit-grasp instance)
4. **Metrics:** Enable from day 1 for comparison
### Accelerated Migration: Risk Assessment
**Risks with fast migration:**
| Risk | Likelihood | Impact | Mitigation |
|------|-----------|--------|------------|
| Insufficient baseline data | Medium | Medium | Collect 48-72h minimum; focus on peak hours |
| Performance regression | Low | High | ngit-grasp proven on ngit.danconwaydev.com; instant rollback available |
| Data compatibility issues | Low | Medium | Both use same event format; test with sample data first |
| Configuration errors | Medium | Low | Pre-validate config; test build before migration |
| User disruption | Low | Medium | Migrate during off-peak; monitor intensively first 48h |
| Rollback complications | Low | High | Keep ngit-relay container; document exact rollback steps |
**Why acceptable to move fast:**
1. ✅ **Proven implementation:** ngit-grasp runs ngit.danconwaydev.com successfully
2. ✅ **Easy rollback:** Container-based deployment allows 5-min rollback
3. ✅ **Low user impact:** relay.ngit.dev is not mission-critical (dev/test relay)
4. ✅ **Stable baseline:** ngit-relay v0.0.5 is stable after timeout fixes
5. ✅ **Intensive monitoring:** 48h of close monitoring catches issues early
**What we're trading off:**
- ❌ Detailed performance comparison (1-2 weeks baseline → 48-72h baseline)
- ❌ Comprehensive testing (full test suite → basic functionality tests)
- ❌ Gradual rollout (immediate cutover → no canary deployment)
**Acceptable because:**
- This is a development/test relay, not production-critical
- We have a proven implementation (ngit.danconwaydev.com)
- Rollback is trivial (restart container)
- Benefits (consolidated codebase) outweigh risks
### Accelerated Migration: Monitoring Checklist
**Pre-migration baseline (48-72h):**
```bash
# Capture these metrics before migration
echo "=== CPU Baseline ===" > baseline.txt
top -b -n 1 | grep ngit-relay >> baseline.txt
echo "=== Memory Baseline ===" >> baseline.txt
docker stats ngit-relay --no-stream >> baseline.txt
echo "=== Connection Baseline ===" >> baseline.txt
ss -s >> baseline.txt
echo "=== Error Baseline ===" >> baseline.txt
journalctl -u ngit-relay --since "24 hours ago" | grep -i error | wc -l >> baseline.txt
```
**Post-migration monitoring (48h intensive):**
```bash
# Run every 15 min (hour 1), hourly (hours 2-6), then decreasing frequency
# 1. Service health
systemctl status ngit-grasp-relay-ngit-dev
# 2. Resource usage
top -b -n 1 | grep ngit-grasp
# 3. Connection count
ss -s
# 4. Error check
journalctl -u ngit-grasp-relay-ngit-dev --since "15 minutes ago" | grep -i error
# 5. Metrics endpoint
curl http://localhost:8082/metrics | grep -E "ngit_(connections|events|git)"
# 6. Functional test
# WebSocket: Use websocat or browser
# Git: git ls-remote https://relay.ngit.dev/<npub>/<repo>.git
```
**Rollback trigger conditions:**
- 🚨 Service crashes or fails to restart
- 🚨 Error rate >10x baseline
- 🚨 Memory usage >4x baseline (sustained >30 min)
- 🚨 CPU usage >90% sustained >15 min
- 🚨 Git operations failing
- 🚨 WebSocket connections rejected
**Success indicators:**
- ✅ Service uptime >48h continuous
- ✅ Error rate ≤ baseline
- ✅ Resource usage ≤ 2x baseline
- ✅ All functional tests passing
- ✅ No user complaints