mirror of
https://relay.ngit.dev/npub15qydau2hjma6ngxkl2cyar74wzyjshvl65za5k5rl69264ar2exs5cyejr/ngit-grasp.git
synced 2026-10-05 15:08:24 +00:00
456 lines
19 KiB
Markdown
456 lines
19 KiB
Markdown
# Migrate relay.ngit.dev from ngit-relay to ngit-grasp
|
|
|
|
**ID:** 820a
|
|
|
|
## Problem
|
|
|
|
relay.ngit.dev currently runs ngit-relay (reference implementation). We want to consolidate on ngit-grasp as the production implementation to:
|
|
- Reduce operational complexity (one codebase instead of two)
|
|
- Focus logging and observability improvements on a single project
|
|
- Leverage ngit-grasp's proven performance on ngit.danconwaydev.com
|
|
|
|
**Why this matters:** Running two different implementations increases maintenance burden and makes it harder to add features like comprehensive logging and performance monitoring. Consolidating on ngit-grasp enables:
|
|
- Unified logging and metrics infrastructure
|
|
- Single codebase for bug fixes and features
|
|
- Consistent behavior across all deployments
|
|
- Reduced testing and deployment complexity
|
|
|
|
## Plan
|
|
|
|
- [ ] **Phase 1: Baseline Metrics Collection** (1-2 weeks)
|
|
- [ ] Deploy nginx metrics to relay.ngit.dev (implementation ready in ngit-relay repo)
|
|
- [ ] Collect baseline metrics for 1-2 weeks:
|
|
- CPU usage (user, system, idle)
|
|
- Memory usage (RSS, available)
|
|
- Connection counts (established, TIME_WAIT, rate)
|
|
- Git operation frequency and duration
|
|
- Request rates and response times
|
|
- [ ] Document baseline performance characteristics
|
|
- [ ] Create performance comparison criteria
|
|
|
|
- [ ] **Phase 2: Migration Preparation** (1 week)
|
|
- [ ] Review ngit-grasp configuration options
|
|
- [ ] Plan NixOS service configuration for relay.ngit.dev
|
|
- [ ] Prepare rollback plan
|
|
- [ ] Test ngit-grasp with relay.ngit.dev's data (if possible)
|
|
- [ ] Document migration steps and verification procedures
|
|
- [ ] Identify monitoring checkpoints
|
|
|
|
- [ ] **Phase 3: Migration Execution** (1 day)
|
|
- [ ] Create new ngit-grasp service config for relay.ngit.dev
|
|
- [ ] Deploy to VPS
|
|
- [ ] Verify service starts and accepts connections
|
|
- [ ] Monitor for immediate issues
|
|
- [ ] Keep ngit-relay container available for quick rollback
|
|
- [ ] Test basic functionality (websocket, git operations)
|
|
|
|
- [ ] **Phase 4: Post-Migration Validation** (1-2 weeks)
|
|
- [ ] Collect same metrics as baseline
|
|
- [ ] Compare performance (CPU, memory, connections, response times)
|
|
- [ ] Verify no new timeout patterns
|
|
- [ ] Monitor error rates and connection stability
|
|
- [ ] User acceptance testing
|
|
- [ ] Remove old ngit-relay container if successful
|
|
|
|
## Accelerated Migration Plan
|
|
|
|
**Timeline:** 3-5 days total (vs 3-5 weeks original)
|
|
|
|
### Day 0-3: Minimal Baseline Collection (2-3 days) - IN PROGRESS ⏳
|
|
|
|
**Goal:** Capture enough baseline data to detect major regressions
|
|
|
|
**Status:** Collection started 2026-01-17 23:30 UTC, extended to 2-3 days (until 2026-01-20/21)
|
|
|
|
**Minimum metrics needed:**
|
|
- [x] CPU usage patterns (1 full day cycle minimum) - COLLECTING ✅
|
|
- [x] Memory baseline (current RSS, growth rate) - COLLECTING ✅
|
|
- [x] Peak connection counts (identify daily peak hours) - COLLECTING ✅
|
|
- [x] Basic response time percentiles (p50, p95, p99) - Via connection stats ✅
|
|
- [x] Error rate baseline (if any) - Via systemd logs ✅
|
|
|
|
**Metrics Status (2026-01-18):**
|
|
- ✅ Critical metrics working: connections, CPU, memory, load average
|
|
- ❌ Container stats broken (cgroup access issues) - NOT CRITICAL
|
|
- ❌ FD counts broken (permission issues) - NOT CRITICAL
|
|
- Decision: Proceed with working metrics, sufficient for migration decision
|
|
|
|
**Quick setup:**
|
|
```bash
|
|
# On relay.ngit.dev VPS
|
|
# 1. Check current metrics via systemd/journalctl
|
|
systemctl status ngit-relay
|
|
journalctl -u ngit-relay --since "24 hours ago" | grep -i error
|
|
|
|
# 2. Capture system metrics snapshot
|
|
top -b -n 1 | head -20
|
|
free -h
|
|
ss -s # socket statistics
|
|
|
|
# 3. If nginx metrics available, capture 24h sample
|
|
# Otherwise, rely on system metrics + manual testing
|
|
```
|
|
|
|
**Acceptance criteria:**
|
|
- [ ] 2-3 days of continuous operation data (IN PROGRESS - ~10.5h collected)
|
|
- [ ] Identified peak usage hours (collecting data)
|
|
- [ ] No critical errors in baseline period (monitoring)
|
|
- [ ] Documented current resource usage (ongoing)
|
|
|
|
### Day 3-4: Rapid Preparation (4-6 hours)
|
|
|
|
**Configuration review:**
|
|
- [ ] Map ngit-relay config → ngit-grasp config
|
|
- [ ] Prepare NixOS service definition
|
|
- [ ] Document rollback procedure (keep ngit-relay container)
|
|
- [ ] Pre-build ngit-grasp on VPS (cache dependencies)
|
|
|
|
**Quick prep checklist:**
|
|
```bash
|
|
# 1. Review current ngit-relay config
|
|
docker inspect ngit-relay | jq '.[0].Config.Env'
|
|
|
|
# 2. Create ngit-grasp service config (see template below)
|
|
# 3. Test build on VPS
|
|
nix build github:DanConwayDev/ngit-grasp
|
|
|
|
# 4. Verify data directory permissions
|
|
ls -la /persistent/ngit-relay/
|
|
```
|
|
|
|
**Rollback plan:**
|
|
- Keep ngit-relay container stopped but available
|
|
- Document exact `docker start` command
|
|
- 5-minute rollback window if issues detected
|
|
|
|
### Day 4-5: Migration Execution (2-4 hours)
|
|
|
|
**Execution window:** Off-peak hours (based on Day 0-1 data)
|
|
|
|
**Steps:**
|
|
1. [ ] Stop ngit-relay container (don't remove)
|
|
2. [ ] Deploy ngit-grasp via NixOS
|
|
3. [ ] Verify service starts and binds to port
|
|
4. [ ] Test WebSocket connection
|
|
5. [ ] Test git clone operation
|
|
6. [ ] Monitor for 30 minutes (critical window)
|
|
|
|
**Go/No-Go criteria (30 min checkpoint):**
|
|
- ✅ Service running and accepting connections
|
|
- ✅ No error spikes in logs
|
|
- ✅ Memory usage within 2x of baseline
|
|
- ✅ CPU usage reasonable (<80% sustained)
|
|
- ✅ Git operations functional
|
|
|
|
**If No-Go:** Rollback immediately
|
|
```bash
|
|
systemctl stop ngit-grasp-production
|
|
docker start ngit-relay
|
|
# Total rollback time: <5 minutes
|
|
```
|
|
|
|
### Day 5-7: Intensive Monitoring (48 hours)
|
|
|
|
**Critical monitoring period:**
|
|
- [ ] Hour 1: Check every 15 minutes
|
|
- [ ] Hours 2-6: Check every hour
|
|
- [ ] Hours 6-24: Check every 4 hours
|
|
- [ ] Hours 24-48: Check every 8 hours
|
|
|
|
**Monitor:**
|
|
- CPU/memory trends (compare to baseline)
|
|
- Connection counts (should be similar)
|
|
- Error logs (any new patterns?)
|
|
- Response times (p95, p99)
|
|
- Git operation success rate
|
|
|
|
**Success criteria (48h):**
|
|
- ✅ No critical errors
|
|
- ✅ Resource usage ≤ 2x baseline (acceptable for new impl)
|
|
- ✅ No user complaints
|
|
- ✅ Git operations working
|
|
- ✅ WebSocket connections stable
|
|
|
|
**If successful:** Remove ngit-relay container after 7 days
|
|
**If issues:** Rollback and investigate
|
|
|
|
## Progress
|
|
|
|
### 2026-01-17 [Session 22:00]
|
|
- Created: Initial issue based on pivot from 573b timeout diagnosis
|
|
- Context: Phase 1 of 573b is complete (connection limits increased, rate limiting fixed)
|
|
- Rationale: Better to consolidate on ngit-grasp than implement duplicate features in ngit-relay
|
|
- Blocker: Must collect baseline metrics BEFORE migration for objective comparison
|
|
- Next: Deploy nginx metrics to relay.ngit.dev and begin baseline collection
|
|
|
|
### 2026-01-17 [Session 23:30]
|
|
- Added: Accelerated migration plan (3-5 days vs 3-5 weeks)
|
|
- Decision: User wants to move quickly after minimal baseline collection
|
|
- Approach: 48-72h baseline → rapid prep → execute → intensive monitoring
|
|
- Risk mitigation: Keep ngit-relay container for instant rollback
|
|
- Next: Begin minimal baseline collection (48-72h)
|
|
|
|
### 2026-01-17 [Session 23:45] - Baseline Collection Started
|
|
- Completed: Enhanced diagnostics deployed to all three relays
|
|
- Improvements:
|
|
1. ✅ Fixed connection counting (port-based, not PID-based)
|
|
2. ✅ Added container stats collection (CPU, memory, network I/O)
|
|
3. ✅ Added connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
|
|
4. ✅ Created comprehensive MONITORING.md documentation
|
|
- Baseline Collection Status:
|
|
- **Started:** 2026-01-17 23:30 UTC
|
|
- **Duration:** 36 hours (until 2026-01-19 11:30 UTC)
|
|
- **Relays monitored:** gitnostr.com, relay.ngit.dev, ngit.danconwaydev.com
|
|
- **Enhanced metrics:** Connection counts per service, container stats, state breakdown
|
|
- **Purpose:** Establish baseline before relay.ngit.dev migration
|
|
- Timeline Update:
|
|
- **Day 0 (2026-01-17 23:30):** Baseline collection started
|
|
- **Day 0+10min (2026-01-17 23:40):** Verify data quality
|
|
- **Day 1.5 (2026-01-19 11:30):** Review baseline data (36 hours)
|
|
- **Day 2:** Rapid preparation (if baseline looks good)
|
|
- **Day 3:** Migration execution
|
|
- **Day 3-5:** Intensive monitoring
|
|
- Next Steps:
|
|
1. Check baseline at 10-minute mark (verify collection working)
|
|
2. Review baseline after 36 hours
|
|
3. Make migration decision based on baseline data quality
|
|
|
|
### 2026-01-18 [Session 10:00] - Baseline Collection Status Update
|
|
- Completed: 2 iterations of diagnostics fixes to improve metric reliability
|
|
- **Critical Metrics Working:**
|
|
- ✅ Connection counts per service (port-based detection working)
|
|
- ✅ System metrics (load average, memory, CPU breakdown)
|
|
- ✅ Top processes (identifying resource consumers)
|
|
- ✅ Connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
|
|
- ✅ Basic container visibility (containers appear in top processes)
|
|
- **Non-Critical Metrics Still Broken:**
|
|
- ❌ Container stats (cgroup access issues, requires privileged systemd-run)
|
|
- ❌ File descriptor counts per process (permission issues)
|
|
- Note: Container resource usage still visible in top processes section
|
|
- **Decision: Proceed with Current Metrics**
|
|
- Critical metrics (connections, CPU, memory) working reliably
|
|
- Broken metrics are nice-to-have, not essential for migration decision
|
|
- Have sufficient data to detect major regressions
|
|
- Can proceed with migration planning based on working metrics
|
|
- **Baseline Collection Extended:**
|
|
- Current duration: ~10.5 hours (started 2026-01-17 23:30 UTC)
|
|
- Extended to: 2-3 more days (until 2026-01-20 or 2026-01-21)
|
|
- Reason: Want more data to establish reliable daily patterns
|
|
- Collection continuing with working metrics
|
|
- **Updated Timeline:**
|
|
- **Day 0 (2026-01-17 23:30):** Baseline collection started
|
|
- **Day 0.5 (2026-01-18 10:00):** Diagnostics fixes complete, extending collection
|
|
- **Day 2-3 (2026-01-20/21):** Review baseline data (2-3 days total)
|
|
- **Day 3-4:** Rapid preparation (if baseline shows stable patterns)
|
|
- **Day 4-5:** Migration execution
|
|
- **Day 5-7:** Intensive monitoring
|
|
- **Next Steps:**
|
|
1. Continue baseline collection for 2-3 more days
|
|
2. Review baseline data patterns after collection period
|
|
3. Make migration decision based on baseline analysis
|
|
4. Proceed with rapid preparation if baseline looks good
|
|
|
|
### 2026-01-19 [Session 11:35] - Baseline Analysis Complete - VPS UPGRADE REQUIRED ⚠️
|
|
- Completed: Comprehensive analysis of 36 hours of baseline metrics
|
|
- **CRITICAL DECISION: Migration BLOCKED until VPS upgrade**
|
|
- **Baseline Findings (2026-01-17 23:30 to 2026-01-19 11:35 UTC):**
|
|
- **CPU:** Average 6.96 load (348% utilization), peak 11.83 (592%) - 🔴 SEVERELY OVERLOADED
|
|
- **Memory:** 96.8% average usage, only 130MB available - 🔴 MAXED OUT
|
|
- **Swap:** 2.0GB average (40% of 5GB), growing 0.75GB → 2.5GB - 🟡 HEAVILY USED
|
|
- **Connections:** Port 8081: 28 avg/120 peak, Port 8083: 31 avg/126 peak - ✅ ONLY 3% OF CAPACITY
|
|
- **Correlation:** Load vs connections = 0.210 (weak) - Confirms git operations are bottleneck, not connections
|
|
- **Peak Usage Patterns:**
|
|
- Peak hours: 16:00-20:00 UTC (load 6.5-9.9, connections 865-977)
|
|
- Off-peak hours: 00:00-12:00 UTC (load 6.3-7.4, connections 350-470)
|
|
- **Critical insight:** Even off-peak shows HIGH load (6.3-7.4), indicating baseline overload independent of traffic
|
|
- **Migration Risk Assessment:**
|
|
- **Current VPS (2-core, 4GB):** 🔴 **HIGH RISK** - Migration will add load spike to already overloaded system
|
|
- **Outcome:** High probability of service degradation or outages during migration
|
|
- **Recommendation:** ❌ **DO NOT MIGRATE** on current VPS
|
|
- **VPS Upgrade Requirements:**
|
|
- **Minimum (Survival):** 4 cores, 8GB RAM (~$20-40/month) - 🟡 MEDIUM RISK
|
|
- **Recommended (Healthy):** 6 cores, 8GB RAM (~$40-60/month) - 🟢 LOW RISK ✅
|
|
- **Optimal (Future-proof):** 8 cores, 16GB RAM (~$60-100/month) - 🟢 LOW RISK
|
|
- **Decision:** Upgrade to 6-core, 8GB minimum before migration
|
|
- **Updated Migration Timeline:**
|
|
- **REVISED:** 5-7 days total (was 3-5 days)
|
|
- **Day 0-1:** VPS upgrade + verification (NEW - CRITICAL)
|
|
- **Day 2-3:** Rapid preparation
|
|
- **Day 4-5:** Migration execution
|
|
- **Day 5-7:** Intensive monitoring
|
|
- **Data Quality:**
|
|
- ✅ 36 hours continuous collection, 2,160 data points
|
|
- ✅ Clear patterns identified, no anomalies
|
|
- ✅ **Sufficient for migration decision** - no additional collection needed
|
|
- ✅ Findings are conclusive: VPS upgrade required
|
|
- **Next Steps:**
|
|
1. **CRITICAL:** Upgrade VPS to 6-core, 8GB RAM (BLOCKER for migration)
|
|
2. Verify upgrade impact (24-48 hours monitoring, expect load to drop to ~60-70%)
|
|
3. Proceed with accelerated migration plan (Day 2-7)
|
|
4. Monitor intensively during and after migration
|
|
- **Files Created:**
|
|
- `/tmp/vps-baseline-report.md` - Comprehensive 36-hour analysis with migration scenarios
|
|
- **Status:** Baseline complete ✅, Migration BLOCKED ⚠️ pending VPS upgrade
|
|
- **Blocker:** VPS upgrade to 6-core, 8GB RAM required before migration can proceed
|
|
|
|
## Notes
|
|
|
|
- **Related issue:** 573b-production-timeout-diagnosis.md (timeout investigation that led to this)
|
|
- **Critical requirement:** Must collect baseline BEFORE migration for objective comparison
|
|
- **Success criteria:** Performance equal or better than ngit-relay baseline
|
|
- No increase in CPU/memory usage
|
|
- No new timeout patterns
|
|
- Connection handling remains stable
|
|
- Git operations perform equally well
|
|
- **Rollback plan:** Keep ngit-relay container for 2 weeks post-migration
|
|
- Can switch back immediately if issues arise
|
|
- Only remove after validation period passes
|
|
- **nginx metrics:** Implementation ready in ngit-relay repo
|
|
- Provides connection counts, request rates, upstream health
|
|
- Essential for before/after comparison
|
|
- **Current status:**
|
|
- relay.ngit.dev: Running ngit-relay v0.0.5 with 2000 conn/IP rate limit
|
|
- ngit.danconwaydev.com: Running ngit-grasp with 4096 max connections
|
|
- Both services stable after Phase 1 improvements
|
|
- **Migration benefits:**
|
|
- Single codebase to maintain and improve
|
|
- Can focus logging/observability work on ngit-grasp
|
|
- Proven performance (ngit.danconwaydev.com is stable)
|
|
- Simplified deployment and configuration
|
|
|
|
### Accelerated Migration: Configuration Template
|
|
|
|
**ngit-grasp service config for relay.ngit.dev:**
|
|
|
|
```nix
|
|
# services/ngit-grasp.nix
|
|
{ inputs, ... }:
|
|
|
|
{
|
|
imports = [ inputs.ngit-grasp.nixosModules.default ];
|
|
|
|
services.ngit-grasp.relay-ngit-dev = {
|
|
enable = true;
|
|
domain = "relay.ngit.dev";
|
|
|
|
# Network - bind to localhost, nginx handles TLS
|
|
bindAddress = "127.0.0.1";
|
|
port = 8082; # Or whatever port ngit-relay was using
|
|
|
|
# Storage - reuse existing data directory if possible
|
|
dataDir = "/persistent/ngit-relay"; # Or new dir: /persistent/ngit-grasp
|
|
|
|
# Identity - copy from ngit-relay config
|
|
relayName = "relay.ngit.dev";
|
|
relayDescription = "GRASP relay for git+nostr";
|
|
relayOwnerNsecFile = "/persistent/ngit-relay/relay-owner.nsec";
|
|
|
|
# Sync - bootstrap from ngit.danconwaydev.com
|
|
syncBootstrapRelayUrl = "wss://ngit.danconwaydev.com";
|
|
|
|
# Metrics - enable for monitoring
|
|
metricsEnabled = true;
|
|
|
|
# Logging - start with info, increase to debug if issues
|
|
logLevel = "info";
|
|
};
|
|
|
|
# Nginx/Caddy reverse proxy (existing config, just update upstream)
|
|
# Change upstream from ngit-relay container to 127.0.0.1:8082
|
|
}
|
|
```
|
|
|
|
**Key decisions:**
|
|
1. **Data directory:** Reuse `/persistent/ngit-relay` if compatible, or create new `/persistent/ngit-grasp`
|
|
2. **Port:** Use same port as ngit-relay container was using (check nginx config)
|
|
3. **Bootstrap relay:** Use ngit.danconwaydev.com (known good ngit-grasp instance)
|
|
4. **Metrics:** Enable from day 1 for comparison
|
|
|
|
### Accelerated Migration: Risk Assessment
|
|
|
|
**Risks with fast migration:**
|
|
|
|
| Risk | Likelihood | Impact | Mitigation |
|
|
|------|-----------|--------|------------|
|
|
| Insufficient baseline data | Medium | Medium | Collect 48-72h minimum; focus on peak hours |
|
|
| Performance regression | Low | High | ngit-grasp proven on ngit.danconwaydev.com; instant rollback available |
|
|
| Data compatibility issues | Low | Medium | Both use same event format; test with sample data first |
|
|
| Configuration errors | Medium | Low | Pre-validate config; test build before migration |
|
|
| User disruption | Low | Medium | Migrate during off-peak; monitor intensively first 48h |
|
|
| Rollback complications | Low | High | Keep ngit-relay container; document exact rollback steps |
|
|
|
|
**Why acceptable to move fast:**
|
|
1. ✅ **Proven implementation:** ngit-grasp runs ngit.danconwaydev.com successfully
|
|
2. ✅ **Easy rollback:** Container-based deployment allows 5-min rollback
|
|
3. ✅ **Low user impact:** relay.ngit.dev is not mission-critical (dev/test relay)
|
|
4. ✅ **Stable baseline:** ngit-relay v0.0.5 is stable after timeout fixes
|
|
5. ✅ **Intensive monitoring:** 48h of close monitoring catches issues early
|
|
|
|
**What we're trading off:**
|
|
- ❌ Detailed performance comparison (1-2 weeks baseline → 48-72h baseline)
|
|
- ❌ Comprehensive testing (full test suite → basic functionality tests)
|
|
- ❌ Gradual rollout (immediate cutover → no canary deployment)
|
|
|
|
**Acceptable because:**
|
|
- This is a development/test relay, not production-critical
|
|
- We have a proven implementation (ngit.danconwaydev.com)
|
|
- Rollback is trivial (restart container)
|
|
- Benefits (consolidated codebase) outweigh risks
|
|
|
|
### Accelerated Migration: Monitoring Checklist
|
|
|
|
**Pre-migration baseline (48-72h):**
|
|
```bash
|
|
# Capture these metrics before migration
|
|
echo "=== CPU Baseline ===" > baseline.txt
|
|
top -b -n 1 | grep ngit-relay >> baseline.txt
|
|
|
|
echo "=== Memory Baseline ===" >> baseline.txt
|
|
docker stats ngit-relay --no-stream >> baseline.txt
|
|
|
|
echo "=== Connection Baseline ===" >> baseline.txt
|
|
ss -s >> baseline.txt
|
|
|
|
echo "=== Error Baseline ===" >> baseline.txt
|
|
journalctl -u ngit-relay --since "24 hours ago" | grep -i error | wc -l >> baseline.txt
|
|
```
|
|
|
|
**Post-migration monitoring (48h intensive):**
|
|
```bash
|
|
# Run every 15 min (hour 1), hourly (hours 2-6), then decreasing frequency
|
|
|
|
# 1. Service health
|
|
systemctl status ngit-grasp-relay-ngit-dev
|
|
|
|
# 2. Resource usage
|
|
top -b -n 1 | grep ngit-grasp
|
|
|
|
# 3. Connection count
|
|
ss -s
|
|
|
|
# 4. Error check
|
|
journalctl -u ngit-grasp-relay-ngit-dev --since "15 minutes ago" | grep -i error
|
|
|
|
# 5. Metrics endpoint
|
|
curl http://localhost:8082/metrics | grep -E "ngit_(connections|events|git)"
|
|
|
|
# 6. Functional test
|
|
# WebSocket: Use websocat or browser
|
|
# Git: git ls-remote https://relay.ngit.dev/<npub>/<repo>.git
|
|
```
|
|
|
|
**Rollback trigger conditions:**
|
|
- 🚨 Service crashes or fails to restart
|
|
- 🚨 Error rate >10x baseline
|
|
- 🚨 Memory usage >4x baseline (sustained >30 min)
|
|
- 🚨 CPU usage >90% sustained >15 min
|
|
- 🚨 Git operations failing
|
|
- 🚨 WebSocket connections rejected
|
|
|
|
**Success indicators:**
|
|
- ✅ Service uptime >48h continuous
|
|
- ✅ Error rate ≤ baseline
|
|
- ✅ Resource usage ≤ 2x baseline
|
|
- ✅ All functional tests passing
|
|
- ✅ No user complaints
|