Files
ngit-grasp/820a-relay-ngit-dev-migration.md
T

30 KiB

Migrate relay.ngit.dev from ngit-relay to ngit-grasp

ID: 820a

Problem

relay.ngit.dev currently runs ngit-relay (reference implementation). We want to consolidate on ngit-grasp as the production implementation to:

  • Reduce operational complexity (one codebase instead of two)
  • Focus logging and observability improvements on a single project
  • Leverage ngit-grasp's proven performance on ngit.danconwaydev.com

Why this matters: Running two different implementations increases maintenance burden and makes it harder to add features like comprehensive logging and performance monitoring. Consolidating on ngit-grasp enables:

  • Unified logging and metrics infrastructure
  • Single codebase for bug fixes and features
  • Consistent behavior across all deployments
  • Reduced testing and deployment complexity

Plan

High-Level Migration Approach: Archive-based migration to ensure complete git data availability and prevent purgatory issues.

Timeline: 3-5 days total (from archive sync completion)

Migration Phases:

  1. ✅ Phase 1: Archive Sync (1-2 days) - COMPLETED

    • Deploy archive instance syncing from relay.ngit.dev
    • Archive fetches all git data during sync
    • Monitor sync progress and completion
  2. ❌ Phase 2: Validation (4-6 hours) - BLOCKED

    • Nostr events (kind 30617): 99.86% sync ✅
    • State events (kind 30618): 42.4% sync ❌ BLOCKER
    • Archive missing 338/587 repos from production (57.6%)
    • Git data verification failed (empty refs in sample repos)
  3. Phase 3: Test Instance (1 day) - PENDING

    • Deploy test instance syncing from archive
    • Validate all repos accessible
    • Verify no purgatory issues
  4. Phase 4: Production Cutover (2-4 hours) - PENDING

    • Stop ngit-relay, deploy ngit-grasp
    • Monitor 30-minute critical window
    • Rollback if issues detected
  5. Phase 5: Post-Migration Monitoring (48 hours) - PENDING

    • Intensive monitoring (15min → hourly → 4h → 8h intervals)
    • Compare metrics to baseline
    • Remove ngit-relay container after 7 days if successful

Detailed Implementation: See ngit-relay Migration Guide for:

  • Complete phase-by-phase instructions
  • Helper scripts for validation and comparison
  • Troubleshooting procedures
  • Rollback strategies
  • Monitoring checklists

Accelerated Migration Plan

Timeline: 3-5 days total (vs 3-5 weeks original)

Day 0-3: Minimal Baseline Collection (2-3 days) - IN PROGRESS ⏳

Goal: Capture enough baseline data to detect major regressions

Status: Collection started 2026-01-17 23:30 UTC, extended to 2-3 days (until 2026-01-20/21)

Minimum metrics needed:

  • CPU usage patterns (1 full day cycle minimum) - COLLECTING ✅
  • Memory baseline (current RSS, growth rate) - COLLECTING ✅
  • Peak connection counts (identify daily peak hours) - COLLECTING ✅
  • Basic response time percentiles (p50, p95, p99) - Via connection stats ✅
  • Error rate baseline (if any) - Via systemd logs ✅

Metrics Status (2026-01-18):

  • ✅ Critical metrics working: connections, CPU, memory, load average
  • ❌ Container stats broken (cgroup access issues) - NOT CRITICAL
  • ❌ FD counts broken (permission issues) - NOT CRITICAL
  • Decision: Proceed with working metrics, sufficient for migration decision

Quick setup:

# On relay.ngit.dev VPS
# 1. Check current metrics via systemd/journalctl
systemctl status ngit-relay
journalctl -u ngit-relay --since "24 hours ago" | grep -i error

# 2. Capture system metrics snapshot
top -b -n 1 | head -20
free -h
ss -s  # socket statistics

# 3. If nginx metrics available, capture 24h sample
# Otherwise, rely on system metrics + manual testing

Acceptance criteria:

  • 2-3 days of continuous operation data (IN PROGRESS - ~10.5h collected)
  • Identified peak usage hours (collecting data)
  • No critical errors in baseline period (monitoring)
  • Documented current resource usage (ongoing)

Day 3-4: Rapid Preparation (4-6 hours)

Configuration review:

  • Map ngit-relay config → ngit-grasp config
  • Prepare NixOS service definition
  • Document rollback procedure (keep ngit-relay container)
  • Pre-build ngit-grasp on VPS (cache dependencies)

Quick prep checklist:

# 1. Review current ngit-relay config
docker inspect ngit-relay | jq '.[0].Config.Env'

# 2. Create ngit-grasp service config (see template below)
# 3. Test build on VPS
nix build github:DanConwayDev/ngit-grasp

# 4. Verify data directory permissions
ls -la /persistent/ngit-relay/

Rollback plan:

  • Keep ngit-relay container stopped but available
  • Document exact docker start command
  • 5-minute rollback window if issues detected

Day 4-5: Migration Execution (2-4 hours)

Execution window: Off-peak hours (based on Day 0-1 data)

Steps:

  1. Stop ngit-relay container (don't remove)
  2. Deploy ngit-grasp via NixOS
  3. Verify service starts and binds to port
  4. Test WebSocket connection
  5. Test git clone operation
  6. Monitor for 30 minutes (critical window)

Go/No-Go criteria (30 min checkpoint):

  • ✅ Service running and accepting connections
  • ✅ No error spikes in logs
  • ✅ Memory usage within 2x of baseline
  • ✅ CPU usage reasonable (<80% sustained)
  • ✅ Git operations functional

If No-Go: Rollback immediately

systemctl stop ngit-grasp-production
docker start ngit-relay
# Total rollback time: <5 minutes

Day 5-7: Intensive Monitoring (48 hours)

Critical monitoring period:

  • Hour 1: Check every 15 minutes
  • Hours 2-6: Check every hour
  • Hours 6-24: Check every 4 hours
  • Hours 24-48: Check every 8 hours

Monitor:

  • CPU/memory trends (compare to baseline)
  • Connection counts (should be similar)
  • Error logs (any new patterns?)
  • Response times (p95, p99)
  • Git operation success rate

Success criteria (48h):

  • ✅ No critical errors
  • ✅ Resource usage ≤ 2x baseline (acceptable for new impl)
  • ✅ No user complaints
  • ✅ Git operations working
  • ✅ WebSocket connections stable

If successful: Remove ngit-relay container after 7 days If issues: Rollback and investigate

Progress

2026-01-17 [Session 22:00]

  • Created: Initial issue based on pivot from 573b timeout diagnosis
  • Context: Phase 1 of 573b is complete (connection limits increased, rate limiting fixed)
  • Rationale: Better to consolidate on ngit-grasp than implement duplicate features in ngit-relay
  • Blocker: Must collect baseline metrics BEFORE migration for objective comparison
  • Next: Deploy nginx metrics to relay.ngit.dev and begin baseline collection

2026-01-17 [Session 23:30]

  • Added: Accelerated migration plan (3-5 days vs 3-5 weeks)
  • Decision: User wants to move quickly after minimal baseline collection
  • Approach: 48-72h baseline → rapid prep → execute → intensive monitoring
  • Risk mitigation: Keep ngit-relay container for instant rollback
  • Next: Begin minimal baseline collection (48-72h)

2026-01-17 [Session 23:45] - Baseline Collection Started

  • Completed: Enhanced diagnostics deployed to all three relays
  • Improvements:
    1. ✅ Fixed connection counting (port-based, not PID-based)
    2. ✅ Added container stats collection (CPU, memory, network I/O)
    3. ✅ Added connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
    4. ✅ Created comprehensive MONITORING.md documentation
  • Baseline Collection Status:
    • Started: 2026-01-17 23:30 UTC
    • Duration: 36 hours (until 2026-01-19 11:30 UTC)
    • Relays monitored: gitnostr.com, relay.ngit.dev, ngit.danconwaydev.com
    • Enhanced metrics: Connection counts per service, container stats, state breakdown
    • Purpose: Establish baseline before relay.ngit.dev migration
  • Timeline Update:
    • Day 0 (2026-01-17 23:30): Baseline collection started
    • Day 0+10min (2026-01-17 23:40): Verify data quality
    • Day 1.5 (2026-01-19 11:30): Review baseline data (36 hours)
    • Day 2: Rapid preparation (if baseline looks good)
    • Day 3: Migration execution
    • Day 3-5: Intensive monitoring
  • Next Steps:
    1. Check baseline at 10-minute mark (verify collection working)
    2. Review baseline after 36 hours
    3. Make migration decision based on baseline data quality

2026-01-18 [Session 10:00] - Baseline Collection Status Update

  • Completed: 2 iterations of diagnostics fixes to improve metric reliability
  • Critical Metrics Working:
    • ✅ Connection counts per service (port-based detection working)
    • ✅ System metrics (load average, memory, CPU breakdown)
    • ✅ Top processes (identifying resource consumers)
    • ✅ Connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
    • ✅ Basic container visibility (containers appear in top processes)
  • Non-Critical Metrics Still Broken:
    • ❌ Container stats (cgroup access issues, requires privileged systemd-run)
    • ❌ File descriptor counts per process (permission issues)
    • Note: Container resource usage still visible in top processes section
  • Decision: Proceed with Current Metrics
    • Critical metrics (connections, CPU, memory) working reliably
    • Broken metrics are nice-to-have, not essential for migration decision
    • Have sufficient data to detect major regressions
    • Can proceed with migration planning based on working metrics
  • Baseline Collection Extended:
    • Current duration: ~10.5 hours (started 2026-01-17 23:30 UTC)
    • Extended to: 2-3 more days (until 2026-01-20 or 2026-01-21)
    • Reason: Want more data to establish reliable daily patterns
    • Collection continuing with working metrics
  • Updated Timeline:
    • Day 0 (2026-01-17 23:30): Baseline collection started
    • Day 0.5 (2026-01-18 10:00): Diagnostics fixes complete, extending collection
    • Day 2-3 (2026-01-20/21): Review baseline data (2-3 days total)
    • Day 3-4: Rapid preparation (if baseline shows stable patterns)
    • Day 4-5: Migration execution
    • Day 5-7: Intensive monitoring
  • Next Steps:
    1. Continue baseline collection for 2-3 more days
    2. Review baseline data patterns after collection period
    3. Make migration decision based on baseline analysis
    4. Proceed with rapid preparation if baseline looks good

2026-01-19 [Session 11:35] - Baseline Analysis Complete - VPS UPGRADE REQUIRED ⚠️

  • Completed: Comprehensive analysis of 36 hours of baseline metrics
  • CRITICAL DECISION: Migration BLOCKED until VPS upgrade
  • Baseline Findings (2026-01-17 23:30 to 2026-01-19 11:35 UTC):
    • CPU: Average 6.96 load (348% utilization), peak 11.83 (592%) - 🔴 SEVERELY OVERLOADED
    • Memory: 96.8% average usage, only 130MB available - 🔴 MAXED OUT
    • Swap: 2.0GB average (40% of 5GB), growing 0.75GB → 2.5GB - 🟡 HEAVILY USED
    • Connections: Port 8081: 28 avg/120 peak, Port 8083: 31 avg/126 peak - ✅ ONLY 3% OF CAPACITY
    • Correlation: Load vs connections = 0.210 (weak) - Confirms git operations are bottleneck, not connections
  • Peak Usage Patterns:
    • Peak hours: 16:00-20:00 UTC (load 6.5-9.9, connections 865-977)
    • Off-peak hours: 00:00-12:00 UTC (load 6.3-7.4, connections 350-470)
    • Critical insight: Even off-peak shows HIGH load (6.3-7.4), indicating baseline overload independent of traffic
  • Migration Risk Assessment:
    • Current VPS (2-core, 4GB): 🔴 HIGH RISK - Migration will add load spike to already overloaded system
    • Outcome: High probability of service degradation or outages during migration
    • Recommendation: ❌ DO NOT MIGRATE on current VPS
  • VPS Upgrade Requirements:
    • Minimum (Survival): 4 cores, 8GB RAM (~$20-40/month) - 🟡 MEDIUM RISK
    • Recommended (Healthy): 6 cores, 8GB RAM (~$40-60/month) - 🟢 LOW RISK ✅
    • Optimal (Future-proof): 8 cores, 16GB RAM (~$60-100/month) - 🟢 LOW RISK
    • Decision: Upgrade to 6-core, 8GB minimum before migration
  • Updated Migration Timeline:
    • REVISED: 5-7 days total (was 3-5 days)
    • Day 0-1: VPS upgrade + verification (NEW - CRITICAL)
    • Day 2-3: Rapid preparation
    • Day 4-5: Migration execution
    • Day 5-7: Intensive monitoring
  • Data Quality:
    • ✅ 36 hours continuous collection, 2,160 data points
    • ✅ Clear patterns identified, no anomalies
    • ✅ Sufficient for migration decision - no additional collection needed
    • ✅ Findings are conclusive: VPS upgrade required
  • Next Steps:
    1. CRITICAL: Upgrade VPS to 6-core, 8GB RAM (BLOCKER for migration)
    2. Verify upgrade impact (24-48 hours monitoring, expect load to drop to ~60-70%)
    3. Proceed with accelerated migration plan (Day 2-7)
    4. Monitor intensively during and after migration
  • Files Created:
    • /tmp/vps-baseline-report.md - Comprehensive 36-hour analysis with migration scenarios
  • Status: Baseline complete ✅, Migration BLOCKED ⚠️ pending VPS upgrade
  • Blocker: VPS upgrade to 6-core, 8GB RAM required before migration can proceed

2026-01-20 [Session 16:00] - Worktree Created for Migration Work

  • Started: Created dedicated worktree for migration work
  • Context: Work has progressed in /tmp directory, now moving to proper worktree
  • Next: Save relevant files from /tmp into this worktree and continue migration work

2026-01-20 [Session 18:30] - State Event Validation FAILED - Migration BLOCKED ❌

  • Completed: State event (kind 30618) validation between relay.ngit.dev and archive
  • CRITICAL BLOCKER IDENTIFIED:
    • State event sync: 42.4% (expected >99%) ❌
    • Archive missing 338/587 repos from production (57.6% gap)
    • Git data verification failed: sample repos have empty refs directories
    • Purgatory suspected: repos synced but git data incomplete
  • Comparison to kind 30617:
    • Announcements (30617): 99.86% sync ✅
    • State events (30618): 42.4% sync ❌
    • Indicates archive failed to fetch/maintain git data for majority of repos
  • Root cause analysis:
    • Archive successfully synced event metadata
    • Archive failed to fetch git data from clone URLs
    • ngit-grasp correctly withholds state events when git data missing
    • Result: 57.6% of repos would lose main/master branch visibility after migration
  • Files created:
    • work/state-validation/STATE-VALIDATION-REPORT.md - Comprehensive analysis
    • work/state-validation/missing-from-archive.txt - List of 338 missing repos
    • work/state-validation/archive-metrics.txt - Archive relay metrics
  • Migration status: ❌ BLOCKED - Cannot proceed to Phase 3 with <50% state event coverage
  • Next actions required:
    1. Investigate why archive git fetches failed for 338 repos
    2. Check archive sync logs for git fetch errors
    3. Determine if sync still in progress or permanently failed
    4. Fix git fetch issues and re-sync missing repos
    5. Re-validate state events once archive sync issues resolved
  • Blocker: Archive sync incomplete/failed - must achieve >99% coverage before migration

2026-01-20 [Session 19:00] - Archive Sync Blocker ROOT CAUSE FOUND 🔍

  • Completed: Investigation of archive sync failures for missing 338/587 repos (57.6%)
  • ROOT CAUSE IDENTIFIED:
    • Archive is attempting to clone from https://relay.ngit.dev/... URLs
    • relay.ngit.dev does NOT serve HTTP git clones - only WebSocket relay service
    • Manual test confirmed: git clone https://relay.ngit.dev/<npub>/<repo>.git returns 404
    • Archive has been running 36+ hours but silently failing on relay.ngit.dev URLs
  • Clone URL Pattern:
    ["clone","https://git.shakespeare.diy/.../repo.git","https://relay.ngit.dev/.../repo.git"]
    
    • Primary URLs (git.shakespeare.diy, github.com, etc.): ✅ Works
    • Fallback URLs (relay.ngit.dev): ❌ 404 Not Found
  • Archive Metrics (Port 7443):
    • Service status: ✅ Running normally (19+ hours uptime)
    • Events synced: 3,445 from wss://relay.ngit.dev
    • Repositories tracked: 1,856 (matches 42% coverage)
    • Disk usage: 4.8GB (40-48% of expected 10-12GB)
    • Confirms partial sync, not full sync
  • Why Only 42% Coverage:
    • Archive successfully fetches repos with working primary clone URLs
    • Archive FAILS for repos with only relay.ngit.dev clone URLs (338/587 = 57.6%)
    • Missing repos likely have relay.ngit.dev as primary or only clone URL
  • Configuration Analysis:
    • Archive config is CORRECT: NGIT_DOMAIN=archive.internal
    • Sync source is CORRECT: wss://relay.ngit.dev
    • Archive is NOT filtering relay.ngit.dev URLs (domain is different)
  • Files Created:
    • work/state-validation/ARCHIVE-SYNC-INVESTIGATION.md - Full investigation report
    • Documented service status, metrics, logs, root cause, and recommended fix
  • Recommended Fix: Option 1 (RECOMMENDED)
    • Configure relay.ngit.dev to serve HTTP git clones
    • Add nginx/HTTP service pointing to /persistent/grasp/production/git/
    • Serve at https://relay.ngit.dev/<npub>/<repo>.git/
    • Makes relay.ngit.dev a complete git hosting service
    • No event republishing required
  • Alternative Fixes:
    • Option 2: Update state events to remove non-working clone URLs (complex)
    • Option 3: Enhance archive retry logic (doesn't solve URL problem)
  • Next Steps:
    1. Decide on fix approach (Option 1 recommended)
    2. Configure HTTP git service on relay.ngit.dev
    3. Restart archive to retry failed fetches
    4. Monitor archive metrics for coverage improvement
    5. Re-run state validation (expect >95% coverage after fix)
  • Migration Status: ❌ BLOCKED - HTTP git service required on relay.ngit.dev
  • Success Criteria: Archive coverage >95% (587/587 repos), git data size ~10-12GB

2026-01-20 [Session 17:00] - Nostr Event Analysis Complete & Migration Guide Created

  • Completed: Comprehensive Nostr event analysis comparing relay.ngit.dev and localhost archive
  • Key Findings:
    • 99.86% sync accuracy between relay.ngit.dev (690 events) and localhost:7334 (2,157 events)
    • Only 1 missing repository: einundzwanzighh-gesundesgeld (created 2025-01-19)
    • localhost:7334 is authoritative source with 2-year history vs relay.ngit.dev's 8-month subset
    • boldwallet.git confirmed on localhost:7334 ✅, missing from relay.ngit.dev ❌
  • Created: Comprehensive migration guide at docs/how-to/ngit-relay-migration-guide.md
    • Documented two-relay architecture discovery
    • Included helper scripts for querying, comparing, and validating relays
    • Detailed 5-phase migration plan with validation steps
    • Troubleshooting section for common issues
  • Organized: Copied analysis files to work/analysis/ for reference
    • CORRECTED-ANALYSIS.md - Full event count comparison
    • MISSING-ANALYSIS.md - Missing repository analysis
    • missing-announcements.jsonl - Event ready for import
    • missing-announcements-summary.txt - Human-readable summary
  • Updated: Issue file reorganized to reference migration guide instead of duplicating content
    • Validation section now summarizes findings with link to detailed analysis
    • Plan section references migration guide for implementation details
    • Maintains issue focus on WHAT/WHY, guide documents HOW
  • Decision: Sync from ws://localhost:7334 for complete coverage (1,244 repos vs 602)
  • Next: Continue with Phase 2 validation using helper scripts from migration guide

2026-01-19 [Session 15:00] - REVISED MIGRATION PLAN: Archive-Based Approach ✅

  • MAJOR PIVOT: Discovered critical data integrity issue that changes migration approach
  • Problem Identified:
    • ngit-grasp only serves state events if ALL referenced git data is present
    • ngit-relay serves state events regardless of missing git data
    • Risk: Repos with missing data would lose main/master branch visibility after migration
    • ngit-grasp fetches git data from clone URLs, but filters out its own domain
  • Solution: Archive-Based Migration
    • Deploy archive instance on VPS that syncs from relay.ngit.dev
    • Archive uses different domain (archive.internal) so doesn't filter relay.ngit.dev URLs
    • Archive fetches all git data from relay.ngit.dev during sync
    • Test instance syncs from archive to validate before production cutover
  • Disk Space Assessment:
    • VPS has 49GB total, 21GB available (56% used)
    • Current ngit services: 17.2GB (relay.ngit.dev: 7.5GB, gitnostr.com: 6.7GB, ngit.danconwaydev.com: 77MB)
    • Old ngit-relay data: 2.9GB (deleted to free space)
    • Estimated archive size: 10-12GB (with doxbox blacklist)
    • After cleanup: 24GB available - sufficient for archive approach ✅
  • Archive Service Deployed:
    • Service: ngit-grasp-relay-ngit-archive (running on VPS)
    • Domain: archive.internal
    • Mode: archiveAll = true with repositoryBlacklist = ["doxbox"]
    • Sync from: wss://relay.ngit.dev
    • Port: 7443 (localhost only)
    • Data dir: /persistent/grasp/sync-archive
    • Status: ✅ ACTIVE and syncing (330MB synced as of 17:30 CET)
  • Issues Created/Fixed:
    • Issue b454: NixOS module should auto-create data directories in ExecStartPre
    • Temporary fix: Added tmpfiles rules to create directories before service starts
    • Permanent fix: Will be implemented in ngit-grasp module (tracked in b454)
  • Backup Created:
    • Full backup of relay.ngit.dev data (7.5GB) created on VPS
    • Git repos: 7.1GB compressed
    • Relay DB: 18MB compressed
    • User downloading backup to laptop (slow connection, ~2 hours)
  • Updated Migration Plan:
    1. ✅ Phase 1: Archive Sync (IN PROGRESS)
      • Archive service deployed and syncing from relay.ngit.dev
      • Estimated completion: 10-12GB sync (currently 330MB)
      • Monitor sync progress and final size
    2. Phase 2: Validation (NEXT)
      • Once archive sync complete, assess actual disk usage
      • Verify all repos synced successfully
      • Check for repos stuck in purgatory (missing git data)
      • Compare repo list: archive vs relay.ngit.dev
    3. Phase 3: Test Instance (PENDING)
      • Deploy test instance on laptop (not VPS - saves 15GB)
      • Test instance syncs from VPS archive
      • Validate all repos accessible
      • Verify no repos stuck in purgatory
    4. Phase 4: Production Cutover (PENDING)
      • Only after test instance validation passes
      • Stop ngit-relay, deploy ngit-grasp
      • Point to VPS archive or direct sync
      • Monitor intensively for 48 hours
  • Key Decisions:
    • Test locally (laptop) instead of VPS to save disk space
    • Use doxbox blacklist to reduce archive size
    • Blank start archive on VPS (no upload needed - better bandwidth)
    • Archive fetches git data from relay.ngit.dev during sync
  • Timeline Revised:
    • Archive sync: 1-2 days (depends on bandwidth and repo count)
    • Validation: 4-6 hours
    • Test instance: 1 day
    • Production cutover: 2-4 hours
    • Post-migration monitoring: 48 hours
    • Total: 3-5 days (from archive sync complete)
  • Next Steps:
    1. Monitor archive sync progress: ssh vps1 'du -sh /persistent/grasp/sync-archive'
    2. Wait for archive sync to complete (10-12GB target)
    3. Validate archive completeness (repo count, git refs)
    4. Assess if enough disk space for test instance on VPS or test locally
    5. Proceed with Phase 2 validation
  • Status: Archive syncing ✅, Migration approach validated ✅, Waiting for sync completion ⏳

Validation Results

Analysis Summary (2026-01-20)

Comprehensive Nostr event analysis completed comparing relay.ngit.dev and localhost archive:

Event Counts:

  • ws://localhost:7334 (archive relay): 2,157 events, 1,244 unique repos
  • relay.ngit.dev: 690 events, 602 unique repos

Sync Accuracy: 99.86%

  • Missing from localhost: 1 event (einundzwanzighh-gesundesgeld)
  • Outdated events: 0
  • Shared repos: 548

Key Findings:

  • localhost:7334 is the authoritative source with complete history (Jan 2024 - Jan 2026)
  • relay.ngit.dev is a filtered subset (May 2025 - Jan 2026)
  • boldwallet.git found on localhost:7334 ✅, NOT on relay.ngit.dev ❌
  • 696 repos exist ONLY on localhost (including boldwallet)
  • 54 repos exist ONLY on relay.ngit.dev (newer additions)

Recommendation: Sync from ws://localhost:7334 for complete coverage.

Detailed Analysis: See work/analysis/ for complete event analysis and missing repository details.

Migration Guide: See ngit-relay Migration Guide for detailed validation steps, helper scripts, and troubleshooting procedures.

Notes

  • Related issue: 573b-production-timeout-diagnosis.md (timeout investigation that led to this)
  • Critical requirement: Must collect baseline BEFORE migration for objective comparison
  • Success criteria: Performance equal or better than ngit-relay baseline
    • No increase in CPU/memory usage
    • No new timeout patterns
    • Connection handling remains stable
    • Git operations perform equally well
  • Rollback plan: Keep ngit-relay container for 2 weeks post-migration
    • Can switch back immediately if issues arise
    • Only remove after validation period passes
  • nginx metrics: Implementation ready in ngit-relay repo
    • Provides connection counts, request rates, upstream health
    • Essential for before/after comparison
  • Current status:
    • relay.ngit.dev: Running ngit-relay v0.0.5 with 2000 conn/IP rate limit
    • ngit.danconwaydev.com: Running ngit-grasp with 4096 max connections
    • Both services stable after Phase 1 improvements
  • Migration benefits:
    • Single codebase to maintain and improve
    • Can focus logging/observability work on ngit-grasp
    • Proven performance (ngit.danconwaydev.com is stable)
    • Simplified deployment and configuration

Accelerated Migration: Configuration Template

ngit-grasp service config for relay.ngit.dev:

# services/ngit-grasp.nix
{ inputs, ... }:

{
  imports = [ inputs.ngit-grasp.nixosModules.default ];

  services.ngit-grasp.relay-ngit-dev = {
    enable = true;
    domain = "relay.ngit.dev";
    
    # Network - bind to localhost, nginx handles TLS
    bindAddress = "127.0.0.1";
    port = 8082;  # Or whatever port ngit-relay was using
    
    # Storage - reuse existing data directory if possible
    dataDir = "/persistent/ngit-relay";  # Or new dir: /persistent/ngit-grasp
    
    # Identity - copy from ngit-relay config
    relayName = "relay.ngit.dev";
    relayDescription = "GRASP relay for git+nostr";
    relayOwnerNsecFile = "/persistent/ngit-relay/relay-owner.nsec";
    
    # Sync - bootstrap from ngit.danconwaydev.com
    syncBootstrapRelayUrl = "wss://ngit.danconwaydev.com";
    
    # Metrics - enable for monitoring
    metricsEnabled = true;
    
    # Logging - start with info, increase to debug if issues
    logLevel = "info";
  };

  # Nginx/Caddy reverse proxy (existing config, just update upstream)
  # Change upstream from ngit-relay container to 127.0.0.1:8082
}

Key decisions:

  1. Data directory: Reuse /persistent/ngit-relay if compatible, or create new /persistent/ngit-grasp
  2. Port: Use same port as ngit-relay container was using (check nginx config)
  3. Bootstrap relay: Use ngit.danconwaydev.com (known good ngit-grasp instance)
  4. Metrics: Enable from day 1 for comparison

Accelerated Migration: Risk Assessment

Risks with fast migration:

Risk Likelihood Impact Mitigation
Insufficient baseline data Medium Medium Collect 48-72h minimum; focus on peak hours
Performance regression Low High ngit-grasp proven on ngit.danconwaydev.com; instant rollback available
Data compatibility issues Low Medium Both use same event format; test with sample data first
Configuration errors Medium Low Pre-validate config; test build before migration
User disruption Low Medium Migrate during off-peak; monitor intensively first 48h
Rollback complications Low High Keep ngit-relay container; document exact rollback steps

Why acceptable to move fast:

  1. ✅ Proven implementation: ngit-grasp runs ngit.danconwaydev.com successfully
  2. ✅ Easy rollback: Container-based deployment allows 5-min rollback
  3. ✅ Low user impact: relay.ngit.dev is not mission-critical (dev/test relay)
  4. ✅ Stable baseline: ngit-relay v0.0.5 is stable after timeout fixes
  5. ✅ Intensive monitoring: 48h of close monitoring catches issues early

What we're trading off:

  • ❌ Detailed performance comparison (1-2 weeks baseline → 48-72h baseline)
  • ❌ Comprehensive testing (full test suite → basic functionality tests)
  • ❌ Gradual rollout (immediate cutover → no canary deployment)

Acceptable because:

  • This is a development/test relay, not production-critical
  • We have a proven implementation (ngit.danconwaydev.com)
  • Rollback is trivial (restart container)
  • Benefits (consolidated codebase) outweigh risks

Accelerated Migration: Monitoring Checklist

Pre-migration baseline (48-72h):

# Capture these metrics before migration
echo "=== CPU Baseline ===" > baseline.txt
top -b -n 1 | grep ngit-relay >> baseline.txt

echo "=== Memory Baseline ===" >> baseline.txt
docker stats ngit-relay --no-stream >> baseline.txt

echo "=== Connection Baseline ===" >> baseline.txt
ss -s >> baseline.txt

echo "=== Error Baseline ===" >> baseline.txt
journalctl -u ngit-relay --since "24 hours ago" | grep -i error | wc -l >> baseline.txt

Post-migration monitoring (48h intensive):

# Run every 15 min (hour 1), hourly (hours 2-6), then decreasing frequency

# 1. Service health
systemctl status ngit-grasp-relay-ngit-dev

# 2. Resource usage
top -b -n 1 | grep ngit-grasp

# 3. Connection count
ss -s

# 4. Error check
journalctl -u ngit-grasp-relay-ngit-dev --since "15 minutes ago" | grep -i error

# 5. Metrics endpoint
curl http://localhost:8082/metrics | grep -E "ngit_(connections|events|git)"

# 6. Functional test
# WebSocket: Use websocat or browser
# Git: git ls-remote https://relay.ngit.dev/<npub>/<repo>.git

Rollback trigger conditions:

  • 🚨 Service crashes or fails to restart
  • 🚨 Error rate >10x baseline
  • 🚨 Memory usage >4x baseline (sustained >30 min)
  • 🚨 CPU usage >90% sustained >15 min
  • 🚨 Git operations failing
  • 🚨 WebSocket connections rejected

Success indicators:

  • ✅ Service uptime >48h continuous
  • ✅ Error rate ≤ baseline
  • ✅ Resource usage ≤ 2x baseline
  • ✅ All functional tests passing
  • ✅ No user complaints