Files
ngit-grasp/820a-relay-ngit-dev-migration.md
T

14 KiB

Migrate relay.ngit.dev from ngit-relay to ngit-grasp

ID: 820a

Problem

relay.ngit.dev currently runs ngit-relay (reference implementation). We want to consolidate on ngit-grasp as the production implementation to:

  • Reduce operational complexity (one codebase instead of two)
  • Focus logging and observability improvements on a single project
  • Leverage ngit-grasp's proven performance on ngit.danconwaydev.com

Why this matters: Running two different implementations increases maintenance burden and makes it harder to add features like comprehensive logging and performance monitoring. Consolidating on ngit-grasp enables:

  • Unified logging and metrics infrastructure
  • Single codebase for bug fixes and features
  • Consistent behavior across all deployments
  • Reduced testing and deployment complexity

Plan

  • Phase 1: Baseline Metrics Collection (1-2 weeks)

    • Deploy nginx metrics to relay.ngit.dev (implementation ready in ngit-relay repo)
    • Collect baseline metrics for 1-2 weeks:
      • CPU usage (user, system, idle)
      • Memory usage (RSS, available)
      • Connection counts (established, TIME_WAIT, rate)
      • Git operation frequency and duration
      • Request rates and response times
    • Document baseline performance characteristics
    • Create performance comparison criteria
  • Phase 2: Migration Preparation (1 week)

    • Review ngit-grasp configuration options
    • Plan NixOS service configuration for relay.ngit.dev
    • Prepare rollback plan
    • Test ngit-grasp with relay.ngit.dev's data (if possible)
    • Document migration steps and verification procedures
    • Identify monitoring checkpoints
  • Phase 3: Migration Execution (1 day)

    • Create new ngit-grasp service config for relay.ngit.dev
    • Deploy to VPS
    • Verify service starts and accepts connections
    • Monitor for immediate issues
    • Keep ngit-relay container available for quick rollback
    • Test basic functionality (websocket, git operations)
  • Phase 4: Post-Migration Validation (1-2 weeks)

    • Collect same metrics as baseline
    • Compare performance (CPU, memory, connections, response times)
    • Verify no new timeout patterns
    • Monitor error rates and connection stability
    • User acceptance testing
    • Remove old ngit-relay container if successful

Accelerated Migration Plan

Timeline: 3-5 days total (vs 3-5 weeks original)

Day 0-1: Minimal Baseline Collection (48-72 hours) - IN PROGRESS ⏳

Goal: Capture enough baseline data to detect major regressions

Status: Collection started 2026-01-17 23:30 UTC, running for 36 hours

Minimum metrics needed:

  • CPU usage patterns (1 full day cycle minimum) - COLLECTING
  • Memory baseline (current RSS, growth rate) - COLLECTING
  • Peak connection counts (identify daily peak hours) - COLLECTING
  • Basic response time percentiles (p50, p95, p99) - Via connection stats
  • Error rate baseline (if any) - Via systemd logs

Quick setup:

# On relay.ngit.dev VPS
# 1. Check current metrics via systemd/journalctl
systemctl status ngit-relay
journalctl -u ngit-relay --since "24 hours ago" | grep -i error

# 2. Capture system metrics snapshot
top -b -n 1 | head -20
free -h
ss -s  # socket statistics

# 3. If nginx metrics available, capture 24h sample
# Otherwise, rely on system metrics + manual testing

Acceptance criteria:

  • ✅ 48+ hours of continuous operation data
  • ✅ Identified peak usage hours
  • ✅ No critical errors in baseline period
  • ✅ Documented current resource usage

Day 2: Rapid Preparation (4-6 hours)

Configuration review:

  • Map ngit-relay config → ngit-grasp config
  • Prepare NixOS service definition
  • Document rollback procedure (keep ngit-relay container)
  • Pre-build ngit-grasp on VPS (cache dependencies)

Quick prep checklist:

# 1. Review current ngit-relay config
docker inspect ngit-relay | jq '.[0].Config.Env'

# 2. Create ngit-grasp service config (see template below)
# 3. Test build on VPS
nix build github:DanConwayDev/ngit-grasp

# 4. Verify data directory permissions
ls -la /persistent/ngit-relay/

Rollback plan:

  • Keep ngit-relay container stopped but available
  • Document exact docker start command
  • 5-minute rollback window if issues detected

Day 3: Migration Execution (2-4 hours)

Execution window: Off-peak hours (based on Day 0-1 data)

Steps:

  1. Stop ngit-relay container (don't remove)
  2. Deploy ngit-grasp via NixOS
  3. Verify service starts and binds to port
  4. Test WebSocket connection
  5. Test git clone operation
  6. Monitor for 30 minutes (critical window)

Go/No-Go criteria (30 min checkpoint):

  • ✅ Service running and accepting connections
  • ✅ No error spikes in logs
  • ✅ Memory usage within 2x of baseline
  • ✅ CPU usage reasonable (<80% sustained)
  • ✅ Git operations functional

If No-Go: Rollback immediately

systemctl stop ngit-grasp-production
docker start ngit-relay
# Total rollback time: <5 minutes

Day 3-5: Intensive Monitoring (48 hours)

Critical monitoring period:

  • Hour 1: Check every 15 minutes
  • Hours 2-6: Check every hour
  • Hours 6-24: Check every 4 hours
  • Hours 24-48: Check every 8 hours

Monitor:

  • CPU/memory trends (compare to baseline)
  • Connection counts (should be similar)
  • Error logs (any new patterns?)
  • Response times (p95, p99)
  • Git operation success rate

Success criteria (48h):

  • ✅ No critical errors
  • ✅ Resource usage ≤ 2x baseline (acceptable for new impl)
  • ✅ No user complaints
  • ✅ Git operations working
  • ✅ WebSocket connections stable

If successful: Remove ngit-relay container after 7 days If issues: Rollback and investigate

Progress

2026-01-17 [Session 22:00]

  • Created: Initial issue based on pivot from 573b timeout diagnosis
  • Context: Phase 1 of 573b is complete (connection limits increased, rate limiting fixed)
  • Rationale: Better to consolidate on ngit-grasp than implement duplicate features in ngit-relay
  • Blocker: Must collect baseline metrics BEFORE migration for objective comparison
  • Next: Deploy nginx metrics to relay.ngit.dev and begin baseline collection

2026-01-17 [Session 23:30]

  • Added: Accelerated migration plan (3-5 days vs 3-5 weeks)
  • Decision: User wants to move quickly after minimal baseline collection
  • Approach: 48-72h baseline → rapid prep → execute → intensive monitoring
  • Risk mitigation: Keep ngit-relay container for instant rollback
  • Next: Begin minimal baseline collection (48-72h)

2026-01-17 [Session 23:45] - Baseline Collection Started

  • Completed: Enhanced diagnostics deployed to all three relays
  • Improvements:
    1. ✅ Fixed connection counting (port-based, not PID-based)
    2. ✅ Added container stats collection (CPU, memory, network I/O)
    3. ✅ Added connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
    4. ✅ Created comprehensive MONITORING.md documentation
  • Baseline Collection Status:
    • Started: 2026-01-17 23:30 UTC
    • Duration: 36 hours (until 2026-01-19 11:30 UTC)
    • Relays monitored: gitnostr.com, relay.ngit.dev, ngit.danconwaydev.com
    • Enhanced metrics: Connection counts per service, container stats, state breakdown
    • Purpose: Establish baseline before relay.ngit.dev migration
  • Timeline Update:
    • Day 0 (2026-01-17 23:30): Baseline collection started
    • Day 0+10min (2026-01-17 23:40): Verify data quality
    • Day 1.5 (2026-01-19 11:30): Review baseline data (36 hours)
    • Day 2: Rapid preparation (if baseline looks good)
    • Day 3: Migration execution
    • Day 3-5: Intensive monitoring
  • Next Steps:
    1. Check baseline at 10-minute mark (verify collection working)
    2. Review baseline after 36 hours
    3. Make migration decision based on baseline data quality

Notes

  • Related issue: 573b-production-timeout-diagnosis.md (timeout investigation that led to this)
  • Critical requirement: Must collect baseline BEFORE migration for objective comparison
  • Success criteria: Performance equal or better than ngit-relay baseline
    • No increase in CPU/memory usage
    • No new timeout patterns
    • Connection handling remains stable
    • Git operations perform equally well
  • Rollback plan: Keep ngit-relay container for 2 weeks post-migration
    • Can switch back immediately if issues arise
    • Only remove after validation period passes
  • nginx metrics: Implementation ready in ngit-relay repo
    • Provides connection counts, request rates, upstream health
    • Essential for before/after comparison
  • Current status:
    • relay.ngit.dev: Running ngit-relay v0.0.5 with 2000 conn/IP rate limit
    • ngit.danconwaydev.com: Running ngit-grasp with 4096 max connections
    • Both services stable after Phase 1 improvements
  • Migration benefits:
    • Single codebase to maintain and improve
    • Can focus logging/observability work on ngit-grasp
    • Proven performance (ngit.danconwaydev.com is stable)
    • Simplified deployment and configuration

Accelerated Migration: Configuration Template

ngit-grasp service config for relay.ngit.dev:

# services/ngit-grasp.nix
{ inputs, ... }:

{
  imports = [ inputs.ngit-grasp.nixosModules.default ];

  services.ngit-grasp.relay-ngit-dev = {
    enable = true;
    domain = "relay.ngit.dev";
    
    # Network - bind to localhost, nginx handles TLS
    bindAddress = "127.0.0.1";
    port = 8082;  # Or whatever port ngit-relay was using
    
    # Storage - reuse existing data directory if possible
    dataDir = "/persistent/ngit-relay";  # Or new dir: /persistent/ngit-grasp
    
    # Identity - copy from ngit-relay config
    relayName = "relay.ngit.dev";
    relayDescription = "GRASP relay for git+nostr";
    relayOwnerNsecFile = "/persistent/ngit-relay/relay-owner.nsec";
    
    # Sync - bootstrap from ngit.danconwaydev.com
    syncBootstrapRelayUrl = "wss://ngit.danconwaydev.com";
    
    # Metrics - enable for monitoring
    metricsEnabled = true;
    
    # Logging - start with info, increase to debug if issues
    logLevel = "info";
  };

  # Nginx/Caddy reverse proxy (existing config, just update upstream)
  # Change upstream from ngit-relay container to 127.0.0.1:8082
}

Key decisions:

  1. Data directory: Reuse /persistent/ngit-relay if compatible, or create new /persistent/ngit-grasp
  2. Port: Use same port as ngit-relay container was using (check nginx config)
  3. Bootstrap relay: Use ngit.danconwaydev.com (known good ngit-grasp instance)
  4. Metrics: Enable from day 1 for comparison

Accelerated Migration: Risk Assessment

Risks with fast migration:

Risk Likelihood Impact Mitigation
Insufficient baseline data Medium Medium Collect 48-72h minimum; focus on peak hours
Performance regression Low High ngit-grasp proven on ngit.danconwaydev.com; instant rollback available
Data compatibility issues Low Medium Both use same event format; test with sample data first
Configuration errors Medium Low Pre-validate config; test build before migration
User disruption Low Medium Migrate during off-peak; monitor intensively first 48h
Rollback complications Low High Keep ngit-relay container; document exact rollback steps

Why acceptable to move fast:

  1. ✅ Proven implementation: ngit-grasp runs ngit.danconwaydev.com successfully
  2. ✅ Easy rollback: Container-based deployment allows 5-min rollback
  3. ✅ Low user impact: relay.ngit.dev is not mission-critical (dev/test relay)
  4. ✅ Stable baseline: ngit-relay v0.0.5 is stable after timeout fixes
  5. ✅ Intensive monitoring: 48h of close monitoring catches issues early

What we're trading off:

  • ❌ Detailed performance comparison (1-2 weeks baseline → 48-72h baseline)
  • ❌ Comprehensive testing (full test suite → basic functionality tests)
  • ❌ Gradual rollout (immediate cutover → no canary deployment)

Acceptable because:

  • This is a development/test relay, not production-critical
  • We have a proven implementation (ngit.danconwaydev.com)
  • Rollback is trivial (restart container)
  • Benefits (consolidated codebase) outweigh risks

Accelerated Migration: Monitoring Checklist

Pre-migration baseline (48-72h):

# Capture these metrics before migration
echo "=== CPU Baseline ===" > baseline.txt
top -b -n 1 | grep ngit-relay >> baseline.txt

echo "=== Memory Baseline ===" >> baseline.txt
docker stats ngit-relay --no-stream >> baseline.txt

echo "=== Connection Baseline ===" >> baseline.txt
ss -s >> baseline.txt

echo "=== Error Baseline ===" >> baseline.txt
journalctl -u ngit-relay --since "24 hours ago" | grep -i error | wc -l >> baseline.txt

Post-migration monitoring (48h intensive):

# Run every 15 min (hour 1), hourly (hours 2-6), then decreasing frequency

# 1. Service health
systemctl status ngit-grasp-relay-ngit-dev

# 2. Resource usage
top -b -n 1 | grep ngit-grasp

# 3. Connection count
ss -s

# 4. Error check
journalctl -u ngit-grasp-relay-ngit-dev --since "15 minutes ago" | grep -i error

# 5. Metrics endpoint
curl http://localhost:8082/metrics | grep -E "ngit_(connections|events|git)"

# 6. Functional test
# WebSocket: Use websocat or browser
# Git: git ls-remote https://relay.ngit.dev/<npub>/<repo>.git

Rollback trigger conditions:

  • 🚨 Service crashes or fails to restart
  • 🚨 Error rate >10x baseline
  • 🚨 Memory usage >4x baseline (sustained >30 min)
  • 🚨 CPU usage >90% sustained >15 min
  • 🚨 Git operations failing
  • 🚨 WebSocket connections rejected

Success indicators:

  • ✅ Service uptime >48h continuous
  • ✅ Error rate ≤ baseline
  • ✅ Resource usage ≤ 2x baseline
  • ✅ All functional tests passing
  • ✅ No user complaints