30 KiB
Migrate relay.ngit.dev from ngit-relay to ngit-grasp
ID: 820a
Problem
relay.ngit.dev currently runs ngit-relay (reference implementation). We want to consolidate on ngit-grasp as the production implementation to:
- Reduce operational complexity (one codebase instead of two)
- Focus logging and observability improvements on a single project
- Leverage ngit-grasp's proven performance on ngit.danconwaydev.com
Why this matters: Running two different implementations increases maintenance burden and makes it harder to add features like comprehensive logging and performance monitoring. Consolidating on ngit-grasp enables:
- Unified logging and metrics infrastructure
- Single codebase for bug fixes and features
- Consistent behavior across all deployments
- Reduced testing and deployment complexity
Plan
High-Level Migration Approach: Archive-based migration to ensure complete git data availability and prevent purgatory issues.
Timeline: 3-5 days total (from archive sync completion)
Migration Phases:
-
✅ Phase 1: Archive Sync (1-2 days) - COMPLETED
- Deploy archive instance syncing from relay.ngit.dev
- Archive fetches all git data during sync
- Monitor sync progress and completion
-
❌ Phase 2: Validation (4-6 hours) - BLOCKED
- Nostr events (kind 30617): 99.86% sync ✅
- State events (kind 30618): 42.4% sync ❌ BLOCKER
- Archive missing 338/587 repos from production (57.6%)
- Git data verification failed (empty refs in sample repos)
-
Phase 3: Test Instance (1 day) - PENDING
- Deploy test instance syncing from archive
- Validate all repos accessible
- Verify no purgatory issues
-
Phase 4: Production Cutover (2-4 hours) - PENDING
- Stop ngit-relay, deploy ngit-grasp
- Monitor 30-minute critical window
- Rollback if issues detected
-
Phase 5: Post-Migration Monitoring (48 hours) - PENDING
- Intensive monitoring (15min → hourly → 4h → 8h intervals)
- Compare metrics to baseline
- Remove ngit-relay container after 7 days if successful
Detailed Implementation: See ngit-relay Migration Guide for:
- Complete phase-by-phase instructions
- Helper scripts for validation and comparison
- Troubleshooting procedures
- Rollback strategies
- Monitoring checklists
Accelerated Migration Plan
Timeline: 3-5 days total (vs 3-5 weeks original)
Day 0-3: Minimal Baseline Collection (2-3 days) - IN PROGRESS ⏳
Goal: Capture enough baseline data to detect major regressions
Status: Collection started 2026-01-17 23:30 UTC, extended to 2-3 days (until 2026-01-20/21)
Minimum metrics needed:
- CPU usage patterns (1 full day cycle minimum) - COLLECTING ✅
- Memory baseline (current RSS, growth rate) - COLLECTING ✅
- Peak connection counts (identify daily peak hours) - COLLECTING ✅
- Basic response time percentiles (p50, p95, p99) - Via connection stats ✅
- Error rate baseline (if any) - Via systemd logs ✅
Metrics Status (2026-01-18):
- ✅ Critical metrics working: connections, CPU, memory, load average
- ❌ Container stats broken (cgroup access issues) - NOT CRITICAL
- ❌ FD counts broken (permission issues) - NOT CRITICAL
- Decision: Proceed with working metrics, sufficient for migration decision
Quick setup:
# On relay.ngit.dev VPS
# 1. Check current metrics via systemd/journalctl
systemctl status ngit-relay
journalctl -u ngit-relay --since "24 hours ago" | grep -i error
# 2. Capture system metrics snapshot
top -b -n 1 | head -20
free -h
ss -s # socket statistics
# 3. If nginx metrics available, capture 24h sample
# Otherwise, rely on system metrics + manual testing
Acceptance criteria:
- 2-3 days of continuous operation data (IN PROGRESS - ~10.5h collected)
- Identified peak usage hours (collecting data)
- No critical errors in baseline period (monitoring)
- Documented current resource usage (ongoing)
Day 3-4: Rapid Preparation (4-6 hours)
Configuration review:
- Map ngit-relay config → ngit-grasp config
- Prepare NixOS service definition
- Document rollback procedure (keep ngit-relay container)
- Pre-build ngit-grasp on VPS (cache dependencies)
Quick prep checklist:
# 1. Review current ngit-relay config
docker inspect ngit-relay | jq '.[0].Config.Env'
# 2. Create ngit-grasp service config (see template below)
# 3. Test build on VPS
nix build github:DanConwayDev/ngit-grasp
# 4. Verify data directory permissions
ls -la /persistent/ngit-relay/
Rollback plan:
- Keep ngit-relay container stopped but available
- Document exact
docker startcommand - 5-minute rollback window if issues detected
Day 4-5: Migration Execution (2-4 hours)
Execution window: Off-peak hours (based on Day 0-1 data)
Steps:
- Stop ngit-relay container (don't remove)
- Deploy ngit-grasp via NixOS
- Verify service starts and binds to port
- Test WebSocket connection
- Test git clone operation
- Monitor for 30 minutes (critical window)
Go/No-Go criteria (30 min checkpoint):
- ✅ Service running and accepting connections
- ✅ No error spikes in logs
- ✅ Memory usage within 2x of baseline
- ✅ CPU usage reasonable (<80% sustained)
- ✅ Git operations functional
If No-Go: Rollback immediately
systemctl stop ngit-grasp-production
docker start ngit-relay
# Total rollback time: <5 minutes
Day 5-7: Intensive Monitoring (48 hours)
Critical monitoring period:
- Hour 1: Check every 15 minutes
- Hours 2-6: Check every hour
- Hours 6-24: Check every 4 hours
- Hours 24-48: Check every 8 hours
Monitor:
- CPU/memory trends (compare to baseline)
- Connection counts (should be similar)
- Error logs (any new patterns?)
- Response times (p95, p99)
- Git operation success rate
Success criteria (48h):
- ✅ No critical errors
- ✅ Resource usage ≤ 2x baseline (acceptable for new impl)
- ✅ No user complaints
- ✅ Git operations working
- ✅ WebSocket connections stable
If successful: Remove ngit-relay container after 7 days If issues: Rollback and investigate
Progress
2026-01-17 [Session 22:00]
- Created: Initial issue based on pivot from 573b timeout diagnosis
- Context: Phase 1 of 573b is complete (connection limits increased, rate limiting fixed)
- Rationale: Better to consolidate on ngit-grasp than implement duplicate features in ngit-relay
- Blocker: Must collect baseline metrics BEFORE migration for objective comparison
- Next: Deploy nginx metrics to relay.ngit.dev and begin baseline collection
2026-01-17 [Session 23:30]
- Added: Accelerated migration plan (3-5 days vs 3-5 weeks)
- Decision: User wants to move quickly after minimal baseline collection
- Approach: 48-72h baseline → rapid prep → execute → intensive monitoring
- Risk mitigation: Keep ngit-relay container for instant rollback
- Next: Begin minimal baseline collection (48-72h)
2026-01-17 [Session 23:45] - Baseline Collection Started
- Completed: Enhanced diagnostics deployed to all three relays
- Improvements:
- ✅ Fixed connection counting (port-based, not PID-based)
- ✅ Added container stats collection (CPU, memory, network I/O)
- ✅ Added connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
- ✅ Created comprehensive MONITORING.md documentation
- Baseline Collection Status:
- Started: 2026-01-17 23:30 UTC
- Duration: 36 hours (until 2026-01-19 11:30 UTC)
- Relays monitored: gitnostr.com, relay.ngit.dev, ngit.danconwaydev.com
- Enhanced metrics: Connection counts per service, container stats, state breakdown
- Purpose: Establish baseline before relay.ngit.dev migration
- Timeline Update:
- Day 0 (2026-01-17 23:30): Baseline collection started
- Day 0+10min (2026-01-17 23:40): Verify data quality
- Day 1.5 (2026-01-19 11:30): Review baseline data (36 hours)
- Day 2: Rapid preparation (if baseline looks good)
- Day 3: Migration execution
- Day 3-5: Intensive monitoring
- Next Steps:
- Check baseline at 10-minute mark (verify collection working)
- Review baseline after 36 hours
- Make migration decision based on baseline data quality
2026-01-18 [Session 10:00] - Baseline Collection Status Update
- Completed: 2 iterations of diagnostics fixes to improve metric reliability
- Critical Metrics Working:
- ✅ Connection counts per service (port-based detection working)
- ✅ System metrics (load average, memory, CPU breakdown)
- ✅ Top processes (identifying resource consumers)
- ✅ Connection state breakdown (ESTABLISHED, TIME_WAIT, etc.)
- ✅ Basic container visibility (containers appear in top processes)
- Non-Critical Metrics Still Broken:
- ❌ Container stats (cgroup access issues, requires privileged systemd-run)
- ❌ File descriptor counts per process (permission issues)
- Note: Container resource usage still visible in top processes section
- Decision: Proceed with Current Metrics
- Critical metrics (connections, CPU, memory) working reliably
- Broken metrics are nice-to-have, not essential for migration decision
- Have sufficient data to detect major regressions
- Can proceed with migration planning based on working metrics
- Baseline Collection Extended:
- Current duration: ~10.5 hours (started 2026-01-17 23:30 UTC)
- Extended to: 2-3 more days (until 2026-01-20 or 2026-01-21)
- Reason: Want more data to establish reliable daily patterns
- Collection continuing with working metrics
- Updated Timeline:
- Day 0 (2026-01-17 23:30): Baseline collection started
- Day 0.5 (2026-01-18 10:00): Diagnostics fixes complete, extending collection
- Day 2-3 (2026-01-20/21): Review baseline data (2-3 days total)
- Day 3-4: Rapid preparation (if baseline shows stable patterns)
- Day 4-5: Migration execution
- Day 5-7: Intensive monitoring
- Next Steps:
- Continue baseline collection for 2-3 more days
- Review baseline data patterns after collection period
- Make migration decision based on baseline analysis
- Proceed with rapid preparation if baseline looks good
2026-01-19 [Session 11:35] - Baseline Analysis Complete - VPS UPGRADE REQUIRED ⚠️
- Completed: Comprehensive analysis of 36 hours of baseline metrics
- CRITICAL DECISION: Migration BLOCKED until VPS upgrade
- Baseline Findings (2026-01-17 23:30 to 2026-01-19 11:35 UTC):
- CPU: Average 6.96 load (348% utilization), peak 11.83 (592%) - 🔴 SEVERELY OVERLOADED
- Memory: 96.8% average usage, only 130MB available - 🔴 MAXED OUT
- Swap: 2.0GB average (40% of 5GB), growing 0.75GB → 2.5GB - 🟡 HEAVILY USED
- Connections: Port 8081: 28 avg/120 peak, Port 8083: 31 avg/126 peak - ✅ ONLY 3% OF CAPACITY
- Correlation: Load vs connections = 0.210 (weak) - Confirms git operations are bottleneck, not connections
- Peak Usage Patterns:
- Peak hours: 16:00-20:00 UTC (load 6.5-9.9, connections 865-977)
- Off-peak hours: 00:00-12:00 UTC (load 6.3-7.4, connections 350-470)
- Critical insight: Even off-peak shows HIGH load (6.3-7.4), indicating baseline overload independent of traffic
- Migration Risk Assessment:
- Current VPS (2-core, 4GB): 🔴 HIGH RISK - Migration will add load spike to already overloaded system
- Outcome: High probability of service degradation or outages during migration
- Recommendation: ❌ DO NOT MIGRATE on current VPS
- VPS Upgrade Requirements:
- Minimum (Survival): 4 cores, 8GB RAM (~$20-40/month) - 🟡 MEDIUM RISK
- Recommended (Healthy): 6 cores, 8GB RAM (~$40-60/month) - 🟢 LOW RISK ✅
- Optimal (Future-proof): 8 cores, 16GB RAM (~$60-100/month) - 🟢 LOW RISK
- Decision: Upgrade to 6-core, 8GB minimum before migration
- Updated Migration Timeline:
- REVISED: 5-7 days total (was 3-5 days)
- Day 0-1: VPS upgrade + verification (NEW - CRITICAL)
- Day 2-3: Rapid preparation
- Day 4-5: Migration execution
- Day 5-7: Intensive monitoring
- Data Quality:
- ✅ 36 hours continuous collection, 2,160 data points
- ✅ Clear patterns identified, no anomalies
- ✅ Sufficient for migration decision - no additional collection needed
- ✅ Findings are conclusive: VPS upgrade required
- Next Steps:
- CRITICAL: Upgrade VPS to 6-core, 8GB RAM (BLOCKER for migration)
- Verify upgrade impact (24-48 hours monitoring, expect load to drop to ~60-70%)
- Proceed with accelerated migration plan (Day 2-7)
- Monitor intensively during and after migration
- Files Created:
/tmp/vps-baseline-report.md- Comprehensive 36-hour analysis with migration scenarios
- Status: Baseline complete ✅, Migration BLOCKED ⚠️ pending VPS upgrade
- Blocker: VPS upgrade to 6-core, 8GB RAM required before migration can proceed
2026-01-20 [Session 16:00] - Worktree Created for Migration Work
- Started: Created dedicated worktree for migration work
- Context: Work has progressed in /tmp directory, now moving to proper worktree
- Next: Save relevant files from /tmp into this worktree and continue migration work
2026-01-20 [Session 18:30] - State Event Validation FAILED - Migration BLOCKED ❌
- Completed: State event (kind 30618) validation between relay.ngit.dev and archive
- CRITICAL BLOCKER IDENTIFIED:
- State event sync: 42.4% (expected >99%) ❌
- Archive missing 338/587 repos from production (57.6% gap)
- Git data verification failed: sample repos have empty refs directories
- Purgatory suspected: repos synced but git data incomplete
- Comparison to kind 30617:
- Announcements (30617): 99.86% sync ✅
- State events (30618): 42.4% sync ❌
- Indicates archive failed to fetch/maintain git data for majority of repos
- Root cause analysis:
- Archive successfully synced event metadata
- Archive failed to fetch git data from clone URLs
- ngit-grasp correctly withholds state events when git data missing
- Result: 57.6% of repos would lose main/master branch visibility after migration
- Files created:
work/state-validation/STATE-VALIDATION-REPORT.md- Comprehensive analysiswork/state-validation/missing-from-archive.txt- List of 338 missing reposwork/state-validation/archive-metrics.txt- Archive relay metrics
- Migration status: ❌ BLOCKED - Cannot proceed to Phase 3 with <50% state event coverage
- Next actions required:
- Investigate why archive git fetches failed for 338 repos
- Check archive sync logs for git fetch errors
- Determine if sync still in progress or permanently failed
- Fix git fetch issues and re-sync missing repos
- Re-validate state events once archive sync issues resolved
- Blocker: Archive sync incomplete/failed - must achieve >99% coverage before migration
2026-01-20 [Session 19:00] - Archive Sync Blocker ROOT CAUSE FOUND 🔍
- Completed: Investigation of archive sync failures for missing 338/587 repos (57.6%)
- ROOT CAUSE IDENTIFIED:
- Archive is attempting to clone from
https://relay.ngit.dev/...URLs - relay.ngit.dev does NOT serve HTTP git clones - only WebSocket relay service
- Manual test confirmed:
git clone https://relay.ngit.dev/<npub>/<repo>.gitreturns 404 - Archive has been running 36+ hours but silently failing on relay.ngit.dev URLs
- Archive is attempting to clone from
- Clone URL Pattern:
["clone","https://git.shakespeare.diy/.../repo.git","https://relay.ngit.dev/.../repo.git"]- Primary URLs (git.shakespeare.diy, github.com, etc.): ✅ Works
- Fallback URLs (relay.ngit.dev): ❌ 404 Not Found
- Archive Metrics (Port 7443):
- Service status: ✅ Running normally (19+ hours uptime)
- Events synced: 3,445 from wss://relay.ngit.dev
- Repositories tracked: 1,856 (matches 42% coverage)
- Disk usage: 4.8GB (40-48% of expected 10-12GB)
- Confirms partial sync, not full sync
- Why Only 42% Coverage:
- Archive successfully fetches repos with working primary clone URLs
- Archive FAILS for repos with only relay.ngit.dev clone URLs (338/587 = 57.6%)
- Missing repos likely have relay.ngit.dev as primary or only clone URL
- Configuration Analysis:
- Archive config is CORRECT:
NGIT_DOMAIN=archive.internal - Sync source is CORRECT:
wss://relay.ngit.dev - Archive is NOT filtering relay.ngit.dev URLs (domain is different)
- Archive config is CORRECT:
- Files Created:
work/state-validation/ARCHIVE-SYNC-INVESTIGATION.md- Full investigation report- Documented service status, metrics, logs, root cause, and recommended fix
- Recommended Fix: Option 1 (RECOMMENDED)
- Configure relay.ngit.dev to serve HTTP git clones
- Add nginx/HTTP service pointing to
/persistent/grasp/production/git/ - Serve at
https://relay.ngit.dev/<npub>/<repo>.git/ - Makes relay.ngit.dev a complete git hosting service
- No event republishing required
- Alternative Fixes:
- Option 2: Update state events to remove non-working clone URLs (complex)
- Option 3: Enhance archive retry logic (doesn't solve URL problem)
- Next Steps:
- Decide on fix approach (Option 1 recommended)
- Configure HTTP git service on relay.ngit.dev
- Restart archive to retry failed fetches
- Monitor archive metrics for coverage improvement
- Re-run state validation (expect >95% coverage after fix)
- Migration Status: ❌ BLOCKED - HTTP git service required on relay.ngit.dev
- Success Criteria: Archive coverage >95% (587/587 repos), git data size ~10-12GB
2026-01-20 [Session 17:00] - Nostr Event Analysis Complete & Migration Guide Created
- Completed: Comprehensive Nostr event analysis comparing relay.ngit.dev and localhost archive
- Key Findings:
- 99.86% sync accuracy between relay.ngit.dev (690 events) and localhost:7334 (2,157 events)
- Only 1 missing repository: einundzwanzighh-gesundesgeld (created 2025-01-19)
- localhost:7334 is authoritative source with 2-year history vs relay.ngit.dev's 8-month subset
- boldwallet.git confirmed on localhost:7334 ✅, missing from relay.ngit.dev ❌
- Created: Comprehensive migration guide at
docs/how-to/ngit-relay-migration-guide.md- Documented two-relay architecture discovery
- Included helper scripts for querying, comparing, and validating relays
- Detailed 5-phase migration plan with validation steps
- Troubleshooting section for common issues
- Organized: Copied analysis files to
work/analysis/for reference- CORRECTED-ANALYSIS.md - Full event count comparison
- MISSING-ANALYSIS.md - Missing repository analysis
- missing-announcements.jsonl - Event ready for import
- missing-announcements-summary.txt - Human-readable summary
- Updated: Issue file reorganized to reference migration guide instead of duplicating content
- Validation section now summarizes findings with link to detailed analysis
- Plan section references migration guide for implementation details
- Maintains issue focus on WHAT/WHY, guide documents HOW
- Decision: Sync from ws://localhost:7334 for complete coverage (1,244 repos vs 602)
- Next: Continue with Phase 2 validation using helper scripts from migration guide
2026-01-19 [Session 15:00] - REVISED MIGRATION PLAN: Archive-Based Approach ✅
- MAJOR PIVOT: Discovered critical data integrity issue that changes migration approach
- Problem Identified:
- ngit-grasp only serves state events if ALL referenced git data is present
- ngit-relay serves state events regardless of missing git data
- Risk: Repos with missing data would lose main/master branch visibility after migration
- ngit-grasp fetches git data from clone URLs, but filters out its own domain
- Solution: Archive-Based Migration
- Deploy archive instance on VPS that syncs from relay.ngit.dev
- Archive uses different domain (archive.internal) so doesn't filter relay.ngit.dev URLs
- Archive fetches all git data from relay.ngit.dev during sync
- Test instance syncs from archive to validate before production cutover
- Disk Space Assessment:
- VPS has 49GB total, 21GB available (56% used)
- Current ngit services: 17.2GB (relay.ngit.dev: 7.5GB, gitnostr.com: 6.7GB, ngit.danconwaydev.com: 77MB)
- Old ngit-relay data: 2.9GB (deleted to free space)
- Estimated archive size: 10-12GB (with doxbox blacklist)
- After cleanup: 24GB available - sufficient for archive approach ✅
- Archive Service Deployed:
- Service:
ngit-grasp-relay-ngit-archive(running on VPS) - Domain:
archive.internal - Mode:
archiveAll = truewithrepositoryBlacklist = ["doxbox"] - Sync from:
wss://relay.ngit.dev - Port:
7443(localhost only) - Data dir:
/persistent/grasp/sync-archive - Status: ✅ ACTIVE and syncing (330MB synced as of 17:30 CET)
- Service:
- Issues Created/Fixed:
- Issue b454: NixOS module should auto-create data directories in ExecStartPre
- Temporary fix: Added tmpfiles rules to create directories before service starts
- Permanent fix: Will be implemented in ngit-grasp module (tracked in b454)
- Backup Created:
- Full backup of relay.ngit.dev data (7.5GB) created on VPS
- Git repos: 7.1GB compressed
- Relay DB: 18MB compressed
- User downloading backup to laptop (slow connection, ~2 hours)
- Updated Migration Plan:
- ✅ Phase 1: Archive Sync (IN PROGRESS)
- Archive service deployed and syncing from relay.ngit.dev
- Estimated completion: 10-12GB sync (currently 330MB)
- Monitor sync progress and final size
- Phase 2: Validation (NEXT)
- Once archive sync complete, assess actual disk usage
- Verify all repos synced successfully
- Check for repos stuck in purgatory (missing git data)
- Compare repo list: archive vs relay.ngit.dev
- Phase 3: Test Instance (PENDING)
- Deploy test instance on laptop (not VPS - saves 15GB)
- Test instance syncs from VPS archive
- Validate all repos accessible
- Verify no repos stuck in purgatory
- Phase 4: Production Cutover (PENDING)
- Only after test instance validation passes
- Stop ngit-relay, deploy ngit-grasp
- Point to VPS archive or direct sync
- Monitor intensively for 48 hours
- ✅ Phase 1: Archive Sync (IN PROGRESS)
- Key Decisions:
- Test locally (laptop) instead of VPS to save disk space
- Use doxbox blacklist to reduce archive size
- Blank start archive on VPS (no upload needed - better bandwidth)
- Archive fetches git data from relay.ngit.dev during sync
- Timeline Revised:
- Archive sync: 1-2 days (depends on bandwidth and repo count)
- Validation: 4-6 hours
- Test instance: 1 day
- Production cutover: 2-4 hours
- Post-migration monitoring: 48 hours
- Total: 3-5 days (from archive sync complete)
- Next Steps:
- Monitor archive sync progress:
ssh vps1 'du -sh /persistent/grasp/sync-archive' - Wait for archive sync to complete (10-12GB target)
- Validate archive completeness (repo count, git refs)
- Assess if enough disk space for test instance on VPS or test locally
- Proceed with Phase 2 validation
- Monitor archive sync progress:
- Status: Archive syncing ✅, Migration approach validated ✅, Waiting for sync completion ⏳
Validation Results
Analysis Summary (2026-01-20)
Comprehensive Nostr event analysis completed comparing relay.ngit.dev and localhost archive:
Event Counts:
- ws://localhost:7334 (archive relay): 2,157 events, 1,244 unique repos
- relay.ngit.dev: 690 events, 602 unique repos
Sync Accuracy: 99.86%
- Missing from localhost: 1 event (einundzwanzighh-gesundesgeld)
- Outdated events: 0
- Shared repos: 548
Key Findings:
- localhost:7334 is the authoritative source with complete history (Jan 2024 - Jan 2026)
- relay.ngit.dev is a filtered subset (May 2025 - Jan 2026)
- boldwallet.git found on localhost:7334 ✅, NOT on relay.ngit.dev ❌
- 696 repos exist ONLY on localhost (including boldwallet)
- 54 repos exist ONLY on relay.ngit.dev (newer additions)
Recommendation: Sync from ws://localhost:7334 for complete coverage.
Detailed Analysis: See work/analysis/ for complete event analysis and missing repository details.
Migration Guide: See ngit-relay Migration Guide for detailed validation steps, helper scripts, and troubleshooting procedures.
Notes
- Related issue: 573b-production-timeout-diagnosis.md (timeout investigation that led to this)
- Critical requirement: Must collect baseline BEFORE migration for objective comparison
- Success criteria: Performance equal or better than ngit-relay baseline
- No increase in CPU/memory usage
- No new timeout patterns
- Connection handling remains stable
- Git operations perform equally well
- Rollback plan: Keep ngit-relay container for 2 weeks post-migration
- Can switch back immediately if issues arise
- Only remove after validation period passes
- nginx metrics: Implementation ready in ngit-relay repo
- Provides connection counts, request rates, upstream health
- Essential for before/after comparison
- Current status:
- relay.ngit.dev: Running ngit-relay v0.0.5 with 2000 conn/IP rate limit
- ngit.danconwaydev.com: Running ngit-grasp with 4096 max connections
- Both services stable after Phase 1 improvements
- Migration benefits:
- Single codebase to maintain and improve
- Can focus logging/observability work on ngit-grasp
- Proven performance (ngit.danconwaydev.com is stable)
- Simplified deployment and configuration
Accelerated Migration: Configuration Template
ngit-grasp service config for relay.ngit.dev:
# services/ngit-grasp.nix
{ inputs, ... }:
{
imports = [ inputs.ngit-grasp.nixosModules.default ];
services.ngit-grasp.relay-ngit-dev = {
enable = true;
domain = "relay.ngit.dev";
# Network - bind to localhost, nginx handles TLS
bindAddress = "127.0.0.1";
port = 8082; # Or whatever port ngit-relay was using
# Storage - reuse existing data directory if possible
dataDir = "/persistent/ngit-relay"; # Or new dir: /persistent/ngit-grasp
# Identity - copy from ngit-relay config
relayName = "relay.ngit.dev";
relayDescription = "GRASP relay for git+nostr";
relayOwnerNsecFile = "/persistent/ngit-relay/relay-owner.nsec";
# Sync - bootstrap from ngit.danconwaydev.com
syncBootstrapRelayUrl = "wss://ngit.danconwaydev.com";
# Metrics - enable for monitoring
metricsEnabled = true;
# Logging - start with info, increase to debug if issues
logLevel = "info";
};
# Nginx/Caddy reverse proxy (existing config, just update upstream)
# Change upstream from ngit-relay container to 127.0.0.1:8082
}
Key decisions:
- Data directory: Reuse
/persistent/ngit-relayif compatible, or create new/persistent/ngit-grasp - Port: Use same port as ngit-relay container was using (check nginx config)
- Bootstrap relay: Use ngit.danconwaydev.com (known good ngit-grasp instance)
- Metrics: Enable from day 1 for comparison
Accelerated Migration: Risk Assessment
Risks with fast migration:
| Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|
| Insufficient baseline data | Medium | Medium | Collect 48-72h minimum; focus on peak hours |
| Performance regression | Low | High | ngit-grasp proven on ngit.danconwaydev.com; instant rollback available |
| Data compatibility issues | Low | Medium | Both use same event format; test with sample data first |
| Configuration errors | Medium | Low | Pre-validate config; test build before migration |
| User disruption | Low | Medium | Migrate during off-peak; monitor intensively first 48h |
| Rollback complications | Low | High | Keep ngit-relay container; document exact rollback steps |
Why acceptable to move fast:
- ✅ Proven implementation: ngit-grasp runs ngit.danconwaydev.com successfully
- ✅ Easy rollback: Container-based deployment allows 5-min rollback
- ✅ Low user impact: relay.ngit.dev is not mission-critical (dev/test relay)
- ✅ Stable baseline: ngit-relay v0.0.5 is stable after timeout fixes
- ✅ Intensive monitoring: 48h of close monitoring catches issues early
What we're trading off:
- ❌ Detailed performance comparison (1-2 weeks baseline → 48-72h baseline)
- ❌ Comprehensive testing (full test suite → basic functionality tests)
- ❌ Gradual rollout (immediate cutover → no canary deployment)
Acceptable because:
- This is a development/test relay, not production-critical
- We have a proven implementation (ngit.danconwaydev.com)
- Rollback is trivial (restart container)
- Benefits (consolidated codebase) outweigh risks
Accelerated Migration: Monitoring Checklist
Pre-migration baseline (48-72h):
# Capture these metrics before migration
echo "=== CPU Baseline ===" > baseline.txt
top -b -n 1 | grep ngit-relay >> baseline.txt
echo "=== Memory Baseline ===" >> baseline.txt
docker stats ngit-relay --no-stream >> baseline.txt
echo "=== Connection Baseline ===" >> baseline.txt
ss -s >> baseline.txt
echo "=== Error Baseline ===" >> baseline.txt
journalctl -u ngit-relay --since "24 hours ago" | grep -i error | wc -l >> baseline.txt
Post-migration monitoring (48h intensive):
# Run every 15 min (hour 1), hourly (hours 2-6), then decreasing frequency
# 1. Service health
systemctl status ngit-grasp-relay-ngit-dev
# 2. Resource usage
top -b -n 1 | grep ngit-grasp
# 3. Connection count
ss -s
# 4. Error check
journalctl -u ngit-grasp-relay-ngit-dev --since "15 minutes ago" | grep -i error
# 5. Metrics endpoint
curl http://localhost:8082/metrics | grep -E "ngit_(connections|events|git)"
# 6. Functional test
# WebSocket: Use websocat or browser
# Git: git ls-remote https://relay.ngit.dev/<npub>/<repo>.git
Rollback trigger conditions:
- 🚨 Service crashes or fails to restart
- 🚨 Error rate >10x baseline
- 🚨 Memory usage >4x baseline (sustained >30 min)
- 🚨 CPU usage >90% sustained >15 min
- 🚨 Git operations failing
- 🚨 WebSocket connections rejected
Success indicators:
- ✅ Service uptime >48h continuous
- ✅ Error rate ≤ baseline
- ✅ Resource usage ≤ 2x baseline
- ✅ All functional tests passing
- ✅ No user complaints