19 KiB
Migrate relay.ngit.dev from ngit-relay to ngit-grasp
ID: 820a
Problem
relay.ngit.dev currently runs ngit-relay (reference implementation). We want to consolidate on ngit-grasp as the production implementation to:
- Reduce operational complexity (one codebase instead of two)
- Focus logging and observability improvements on a single project
- Leverage ngit-grasp's proven performance on ngit.danconwaydev.com
Why this matters: Running two different implementations increases maintenance burden and makes it harder to add features like comprehensive logging and performance monitoring. Consolidating on ngit-grasp enables:
- Unified logging and metrics infrastructure
- Single codebase for bug fixes and features
- Consistent behavior across all deployments
- Reduced testing and deployment complexity
Plan
-
Phase 1: Baseline Metrics Collection (1-2 weeks)
- Deploy nginx metrics to relay.ngit.dev (implementation ready in ngit-relay repo)
- Collect baseline metrics for 1-2 weeks:
- CPU usage (user, system, idle)
- Memory usage (RSS, available)
- Connection counts (established, TIME_WAIT, rate)
- Git operation frequency and duration
- Request rates and response times
- Document baseline performance characteristics
- Create performance comparison criteria
-
Phase 2: Migration Preparation (1 week)
- Review ngit-grasp configuration options
- Plan NixOS service configuration for relay.ngit.dev
- Prepare rollback plan
- Test ngit-grasp with relay.ngit.dev's data (if possible)
- Document migration steps and verification procedures
- Identify monitoring checkpoints
-
Phase 3: Migration Execution (1 day)
- Create new ngit-grasp service config for relay.ngit.dev
- Deploy to VPS
- Verify service starts and accepts connections
- Monitor for immediate issues
- Keep ngit-relay container available for quick rollback
- Test basic functionality (websocket, git operations)
-
Phase 4: Post-Migration Validation (1-2 weeks)
- Collect same metrics as baseline
- Compare performance (CPU, memory, connections, response times)
- Verify no new timeout patterns
- Monitor error rates and connection stability
- User acceptance testing
- Remove old ngit-relay container if successful
Accelerated Migration Plan
Timeline: 3-5 days total (vs 3-5 weeks original)
Day 0-3: Minimal Baseline Collection (2-3 days) - IN PROGRESS ⏳
Goal: Capture enough baseline data to detect major regressions
Status: Collection started 2026-01-17 23:30 UTC, extended to 2-3 days (until 2026-01-20/21)
Minimum metrics needed:
- CPU usage patterns (1 full day cycle minimum) - COLLECTING ✅
- Memory baseline (current RSS, growth rate) - COLLECTING ✅
- Peak connection counts (identify daily peak hours) - COLLECTING ✅
- Basic response time percentiles (p50, p95, p99) - Via connection stats ✅
- Error rate baseline (if any) - Via systemd logs ✅
Metrics Status (2026-01-18):
- ✅ Critical metrics working: connections, CPU, memory, load average
- ❌ Container stats broken (cgroup access issues) - NOT CRITICAL
- ❌ FD counts broken (permission issues) - NOT CRITICAL
- Decision: Proceed with working metrics, sufficient for migration decision
Quick setup:
# On relay.ngit.dev VPS
# 1. Check current metrics via systemd/journalctl
systemctl status ngit-relay
journalctl -u ngit-relay --since "24 hours ago" | grep -i error
# 2. Capture system metrics snapshot
top -b -n 1 | head -20
free -h
ss -s # socket statistics
# 3. If nginx metrics available, capture 24h sample
# Otherwise, rely on system metrics + manual testing
Acceptance criteria:
- 2-3 days of continuous operation data (IN PROGRESS - ~10.5h collected)
- Identified peak usage hours (collecting data)
- No critical errors in baseline period (monitoring)
- Documented current resource usage (ongoing)
Day 3-4: Rapid Preparation (4-6 hours)
Configuration review:
- Map ngit-relay config → ngit-grasp config
- Prepare NixOS service definition
- Document rollback procedure (keep ngit-relay container)
- Pre-build ngit-grasp on VPS (cache dependencies)
Quick prep checklist:
# 1. Review current ngit-relay config
docker inspect ngit-relay | jq '.[0].Config.Env'
# 2. Create ngit-grasp service config (see template below)
# 3. Test build on VPS
nix build github:DanConwayDev/ngit-grasp
# 4. Verify data directory permissions
ls -la /persistent/ngit-relay/
Rollback plan:
- Keep ngit-relay container stopped but available
- Document exact
docker startcommand - 5-minute rollback window if issues detected
Day 4-5: Migration Execution (2-4 hours)
Execution window: Off-peak hours (based on Day 0-1 data)
Steps:
- Stop ngit-relay container (don't remove)
- Deploy ngit-grasp via NixOS
- Verify service starts and binds to port
- Test WebSocket connection
- Test git clone operation
- Monitor for 30 minutes (critical window)
Go/No-Go criteria (30 min checkpoint):
- ✅ Service running and accepting connections
- ✅ No error spikes in logs
- ✅ Memory usage within 2x of baseline
- ✅ CPU usage reasonable (<80% sustained)
- ✅ Git operations functional
If No-Go: Rollback immediately
systemctl stop ngit-grasp-production
docker start ngit-relay
# Total rollback time: <5 minutes
Day 5-7: Intensive Monitoring (48 hours)
Critical monitoring period:
- Hour 1: Check every 15 minutes
- Hours 2-6: Check every hour
- Hours 6-24: Check every 4 hours
- Hours 24-48: Check every 8 hours
Monitor:
- CPU/memory trends (compare to baseline)
- Connection counts (should be similar)
- Error logs (any new patterns?)
- Response times (p95, p99)
- Git operation success rate
Success criteria (48h):
- ✅ No critical errors
- ✅ Resource usage ≤ 2x baseline (acceptable for new impl)
- ✅ No user complaints
- ✅ Git operations working
- ✅ WebSocket connections stable
If successful: Remove ngit-relay container after 7 days If issues: Rollback and investigate
Progress
2026-01-17 [Session 22:00]
- Created: Initial issue based on pivot from 573b timeout diagnosis
- Context: Phase 1 of 573b is complete (connection limits increased, rate limiting fixed)
- Rationale: Better to consolidate on ngit-grasp than implement duplicate features in ngit-relay
- Blocker: Must collect baseline metrics BEFORE migration for objective comparison
- Next: Deploy nginx metrics to relay.ngit.dev and begin baseline collection
2026-01-17 [Session 23:30]
- Added: Accelerated migration plan (3-5 days vs 3-5 weeks)
- Decision: User wants to move quickly after minimal baseline collection
- Approach: 48-72h baseline → rapid prep → execute → intensive monitoring
- Risk mitigation: Keep ngit-relay container for instant rollback
- Next: Begin minimal baseline collection (48-72h)
2026-01-17 [Session 23:45] - Baseline Collection Started
- Completed: Enhanced diagnostics deployed to all three relays
- Improvements:
- ✅ Fixed connection counting (port-based, not PID-based)
- ✅ Added container stats collection (CPU, memory, network I/O)
- ✅ Added connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
- ✅ Created comprehensive MONITORING.md documentation
- Baseline Collection Status:
- Started: 2026-01-17 23:30 UTC
- Duration: 36 hours (until 2026-01-19 11:30 UTC)
- Relays monitored: gitnostr.com, relay.ngit.dev, ngit.danconwaydev.com
- Enhanced metrics: Connection counts per service, container stats, state breakdown
- Purpose: Establish baseline before relay.ngit.dev migration
- Timeline Update:
- Day 0 (2026-01-17 23:30): Baseline collection started
- Day 0+10min (2026-01-17 23:40): Verify data quality
- Day 1.5 (2026-01-19 11:30): Review baseline data (36 hours)
- Day 2: Rapid preparation (if baseline looks good)
- Day 3: Migration execution
- Day 3-5: Intensive monitoring
- Next Steps:
- Check baseline at 10-minute mark (verify collection working)
- Review baseline after 36 hours
- Make migration decision based on baseline data quality
2026-01-18 [Session 10:00] - Baseline Collection Status Update
- Completed: 2 iterations of diagnostics fixes to improve metric reliability
- Critical Metrics Working:
- ✅ Connection counts per service (port-based detection working)
- ✅ System metrics (load average, memory, CPU breakdown)
- ✅ Top processes (identifying resource consumers)
- ✅ Connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
- ✅ Basic container visibility (containers appear in top processes)
- Non-Critical Metrics Still Broken:
- ❌ Container stats (cgroup access issues, requires privileged systemd-run)
- ❌ File descriptor counts per process (permission issues)
- Note: Container resource usage still visible in top processes section
- Decision: Proceed with Current Metrics
- Critical metrics (connections, CPU, memory) working reliably
- Broken metrics are nice-to-have, not essential for migration decision
- Have sufficient data to detect major regressions
- Can proceed with migration planning based on working metrics
- Baseline Collection Extended:
- Current duration: ~10.5 hours (started 2026-01-17 23:30 UTC)
- Extended to: 2-3 more days (until 2026-01-20 or 2026-01-21)
- Reason: Want more data to establish reliable daily patterns
- Collection continuing with working metrics
- Updated Timeline:
- Day 0 (2026-01-17 23:30): Baseline collection started
- Day 0.5 (2026-01-18 10:00): Diagnostics fixes complete, extending collection
- Day 2-3 (2026-01-20/21): Review baseline data (2-3 days total)
- Day 3-4: Rapid preparation (if baseline shows stable patterns)
- Day 4-5: Migration execution
- Day 5-7: Intensive monitoring
- Next Steps:
- Continue baseline collection for 2-3 more days
- Review baseline data patterns after collection period
- Make migration decision based on baseline analysis
- Proceed with rapid preparation if baseline looks good
2026-01-19 [Session 11:35] - Baseline Analysis Complete - VPS UPGRADE REQUIRED ⚠️
- Completed: Comprehensive analysis of 36 hours of baseline metrics
- CRITICAL DECISION: Migration BLOCKED until VPS upgrade
- Baseline Findings (2026-01-17 23:30 to 2026-01-19 11:35 UTC):
- CPU: Average 6.96 load (348% utilization), peak 11.83 (592%) - 🔴 SEVERELY OVERLOADED
- Memory: 96.8% average usage, only 130MB available - 🔴 MAXED OUT
- Swap: 2.0GB average (40% of 5GB), growing 0.75GB → 2.5GB - 🟡 HEAVILY USED
- Connections: Port 8081: 28 avg/120 peak, Port 8083: 31 avg/126 peak - ✅ ONLY 3% OF CAPACITY
- Correlation: Load vs connections = 0.210 (weak) - Confirms git operations are bottleneck, not connections
- Peak Usage Patterns:
- Peak hours: 16:00-20:00 UTC (load 6.5-9.9, connections 865-977)
- Off-peak hours: 00:00-12:00 UTC (load 6.3-7.4, connections 350-470)
- Critical insight: Even off-peak shows HIGH load (6.3-7.4), indicating baseline overload independent of traffic
- Migration Risk Assessment:
- Current VPS (2-core, 4GB): 🔴 HIGH RISK - Migration will add load spike to already overloaded system
- Outcome: High probability of service degradation or outages during migration
- Recommendation: ❌ DO NOT MIGRATE on current VPS
- VPS Upgrade Requirements:
- Minimum (Survival): 4 cores, 8GB RAM (~$20-40/month) - 🟡 MEDIUM RISK
- Recommended (Healthy): 6 cores, 8GB RAM (~$40-60/month) - 🟢 LOW RISK ✅
- Optimal (Future-proof): 8 cores, 16GB RAM (~$60-100/month) - 🟢 LOW RISK
- Decision: Upgrade to 6-core, 8GB minimum before migration
- Updated Migration Timeline:
- REVISED: 5-7 days total (was 3-5 days)
- Day 0-1: VPS upgrade + verification (NEW - CRITICAL)
- Day 2-3: Rapid preparation
- Day 4-5: Migration execution
- Day 5-7: Intensive monitoring
- Data Quality:
- ✅ 36 hours continuous collection, 2,160 data points
- ✅ Clear patterns identified, no anomalies
- ✅ Sufficient for migration decision - no additional collection needed
- ✅ Findings are conclusive: VPS upgrade required
- Next Steps:
- CRITICAL: Upgrade VPS to 6-core, 8GB RAM (BLOCKER for migration)
- Verify upgrade impact (24-48 hours monitoring, expect load to drop to ~60-70%)
- Proceed with accelerated migration plan (Day 2-7)
- Monitor intensively during and after migration
- Files Created:
/tmp/vps-baseline-report.md- Comprehensive 36-hour analysis with migration scenarios
- Status: Baseline complete ✅, Migration BLOCKED ⚠️ pending VPS upgrade
- Blocker: VPS upgrade to 6-core, 8GB RAM required before migration can proceed
Notes
- Related issue: 573b-production-timeout-diagnosis.md (timeout investigation that led to this)
- Critical requirement: Must collect baseline BEFORE migration for objective comparison
- Success criteria: Performance equal or better than ngit-relay baseline
- No increase in CPU/memory usage
- No new timeout patterns
- Connection handling remains stable
- Git operations perform equally well
- Rollback plan: Keep ngit-relay container for 2 weeks post-migration
- Can switch back immediately if issues arise
- Only remove after validation period passes
- nginx metrics: Implementation ready in ngit-relay repo
- Provides connection counts, request rates, upstream health
- Essential for before/after comparison
- Current status:
- relay.ngit.dev: Running ngit-relay v0.0.5 with 2000 conn/IP rate limit
- ngit.danconwaydev.com: Running ngit-grasp with 4096 max connections
- Both services stable after Phase 1 improvements
- Migration benefits:
- Single codebase to maintain and improve
- Can focus logging/observability work on ngit-grasp
- Proven performance (ngit.danconwaydev.com is stable)
- Simplified deployment and configuration
Accelerated Migration: Configuration Template
ngit-grasp service config for relay.ngit.dev:
# services/ngit-grasp.nix
{ inputs, ... }:
{
imports = [ inputs.ngit-grasp.nixosModules.default ];
services.ngit-grasp.relay-ngit-dev = {
enable = true;
domain = "relay.ngit.dev";
# Network - bind to localhost, nginx handles TLS
bindAddress = "127.0.0.1";
port = 8082; # Or whatever port ngit-relay was using
# Storage - reuse existing data directory if possible
dataDir = "/persistent/ngit-relay"; # Or new dir: /persistent/ngit-grasp
# Identity - copy from ngit-relay config
relayName = "relay.ngit.dev";
relayDescription = "GRASP relay for git+nostr";
relayOwnerNsecFile = "/persistent/ngit-relay/relay-owner.nsec";
# Sync - bootstrap from ngit.danconwaydev.com
syncBootstrapRelayUrl = "wss://ngit.danconwaydev.com";
# Metrics - enable for monitoring
metricsEnabled = true;
# Logging - start with info, increase to debug if issues
logLevel = "info";
};
# Nginx/Caddy reverse proxy (existing config, just update upstream)
# Change upstream from ngit-relay container to 127.0.0.1:8082
}
Key decisions:
- Data directory: Reuse
/persistent/ngit-relayif compatible, or create new/persistent/ngit-grasp - Port: Use same port as ngit-relay container was using (check nginx config)
- Bootstrap relay: Use ngit.danconwaydev.com (known good ngit-grasp instance)
- Metrics: Enable from day 1 for comparison
Accelerated Migration: Risk Assessment
Risks with fast migration:
| Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|
| Insufficient baseline data | Medium | Medium | Collect 48-72h minimum; focus on peak hours |
| Performance regression | Low | High | ngit-grasp proven on ngit.danconwaydev.com; instant rollback available |
| Data compatibility issues | Low | Medium | Both use same event format; test with sample data first |
| Configuration errors | Medium | Low | Pre-validate config; test build before migration |
| User disruption | Low | Medium | Migrate during off-peak; monitor intensively first 48h |
| Rollback complications | Low | High | Keep ngit-relay container; document exact rollback steps |
Why acceptable to move fast:
- ✅ Proven implementation: ngit-grasp runs ngit.danconwaydev.com successfully
- ✅ Easy rollback: Container-based deployment allows 5-min rollback
- ✅ Low user impact: relay.ngit.dev is not mission-critical (dev/test relay)
- ✅ Stable baseline: ngit-relay v0.0.5 is stable after timeout fixes
- ✅ Intensive monitoring: 48h of close monitoring catches issues early
What we're trading off:
- ❌ Detailed performance comparison (1-2 weeks baseline → 48-72h baseline)
- ❌ Comprehensive testing (full test suite → basic functionality tests)
- ❌ Gradual rollout (immediate cutover → no canary deployment)
Acceptable because:
- This is a development/test relay, not production-critical
- We have a proven implementation (ngit.danconwaydev.com)
- Rollback is trivial (restart container)
- Benefits (consolidated codebase) outweigh risks
Accelerated Migration: Monitoring Checklist
Pre-migration baseline (48-72h):
# Capture these metrics before migration
echo "=== CPU Baseline ===" > baseline.txt
top -b -n 1 | grep ngit-relay >> baseline.txt
echo "=== Memory Baseline ===" >> baseline.txt
docker stats ngit-relay --no-stream >> baseline.txt
echo "=== Connection Baseline ===" >> baseline.txt
ss -s >> baseline.txt
echo "=== Error Baseline ===" >> baseline.txt
journalctl -u ngit-relay --since "24 hours ago" | grep -i error | wc -l >> baseline.txt
Post-migration monitoring (48h intensive):
# Run every 15 min (hour 1), hourly (hours 2-6), then decreasing frequency
# 1. Service health
systemctl status ngit-grasp-relay-ngit-dev
# 2. Resource usage
top -b -n 1 | grep ngit-grasp
# 3. Connection count
ss -s
# 4. Error check
journalctl -u ngit-grasp-relay-ngit-dev --since "15 minutes ago" | grep -i error
# 5. Metrics endpoint
curl http://localhost:8082/metrics | grep -E "ngit_(connections|events|git)"
# 6. Functional test
# WebSocket: Use websocat or browser
# Git: git ls-remote https://relay.ngit.dev/<npub>/<repo>.git
Rollback trigger conditions:
- 🚨 Service crashes or fails to restart
- 🚨 Error rate >10x baseline
- 🚨 Memory usage >4x baseline (sustained >30 min)
- 🚨 CPU usage >90% sustained >15 min
- 🚨 Git operations failing
- 🚨 WebSocket connections rejected
Success indicators:
- ✅ Service uptime >48h continuous
- ✅ Error rate ≤ baseline
- ✅ Resource usage ≤ 2x baseline
- ✅ All functional tests passing
- ✅ No user complaints