Files
ngit-grasp/4bc5-relay-ngit-dev-migration-v2.md
T

28 KiB

Migrate relay.ngit.dev from ngit-relay to ngit-grasp (v2)

ID: 4bc5

Problem

relay.ngit.dev currently runs ngit-relay (reference implementation). We want to consolidate on ngit-grasp as the production implementation.

Goal: Replace an ngit-relay instance on a VPS running NixOS with ngit-grasp.

Context: This is a fresh start after issue 820a became too complex with extensive investigation history. We're starting from scratch with a focus on creating a small, lightweight, easy-to-implement how-to document.

Plan

Script Development (Modular Architecture)

  • Phase 1: Fetch Events (~30s, local) - 01-fetch-events.sh
    • Fetch kind 30618 (state), 30617 (announcement), 5 (deletion) from relay
    • Run for both prod and archive relays
  • Phase 2: Git Sync Check (~20 mins, VPS) - 10-check-git-sync.sh
    • Compare state event refs to actual git data on disk
    • Run for both prod and archive git directories
    • Note: Existing Jan 22 data available, script not yet created
  • Phase 3: Categorize & Compare (fast, local) - 20-categorize.sh, 21-compare-relays.sh
    • Apply 4-category logic (complete/empty/partial/no-match)
    • Find gaps between prod and archive
  • Phase 4: Log-Based Categories (VPS) - 30-extract-parse-failures.sh, 31-extract-purgatory-expiry.sh
    • Extract parse failures and purgatory expiry from logs
    • Dependency: Logging improvements in ngit-grasp - IMPLEMENTED
  • Phase 5: Final Classification (fast, local) - 40-classify-actions.sh
    • Combine all data to produce: no-action, action-required, manual-investigation
  • Orchestration Script - run-migration-analysis.sh
    • Runs all phases with proper error handling and progress reporting
    • Supports phase control (skip, only, from-phase options)
    • Dry-run mode, timing information, summary display

Migration Execution

  • Run analysis scripts on relay.ngit.dev
  • Review action-required repos, make decisions
  • Execute migration (switch domain, disable archive mode)
  • Validate migration success

Next Session Tasks

  • Investigation: Analyze why 315 repos didn't sync to archive
  • Investigation: Review 5 repos needing manual investigation
  • Decision: Choose migration approach (gradual vs full switch)
  • Decision: Determine when to merge branch to main
  • Preparation: Create migration checklist and rollback plan
  • Execution: Switch relay.ngit.dev domain to ngit-grasp
  • Execution: Disable archive mode
  • Validation: Run post-migration validation
  • Cleanup: Merge branch to main, close issue

Progress

2026-01-23 [Session 16:00]

  • Created: Fresh v2 issue to replace complex 820a migration
  • Context: Previous issue (820a) paused due to complexity
  • Approach: Start from scratch, potentially reuse scripts from 820a worktree
  • Goal: Create small, lightweight, easy-to-implement how-to document
  • Started work: Created worktree for issue 4bc5
  • Completed: Created initial how-to document at docs/how-to/migrate-ngit-relay-to-ngit-grasp.md
  • Document includes: Approach, challenges, analysis categories, gotchas
  • Next: User requested NOT to do planning for migration script yet

2026-01-23 [Session 17:30]

  • Reviewed existing scripts from 820a worktree:
    • analyze-git-state-sync.sh - monolithic, takes ~20 mins (git sync is slow part)
    • compare-categories.sh - compares prod vs archive categories
    • migration-validation-guide.md - comprehensive troubleshooting guide
  • Reviewed existing analysis output (Jan 22):
    • Prod: 654 repos (509 complete, 114 empty, 25 partial, 6 no-match)
    • Archive: 263 repos (247 complete, 9 empty, 5 partial, 2 no-match)
  • Designed modular script architecture for fast iteration:
    • Split into 5 phases with clear inputs/outputs
    • Phases 1, 3, 5 can run locally; Phases 2, 4 need VPS
    • Can use cached data from Jan 22 to develop categorization logic
  • Added log-based categories (scriptable):
    • Parse failures: [PARSE_FAIL] kind=X event_id=Y reason=Z
    • Purgatory expiry: [PURGATORY_EXPIRED] repo=X npub=Y
  • Updated how-to doc with full architecture diagram
  • Next: Implement Phase 1 (fetch events) to get fresh data

2026-01-23 [Session 18:45]

  • Reviewed Phase 2 outputs from Jan 22 (820a worktree):
    • Prod: 654 repos (509 complete, 114 empty, 25 partial, 6 no-match)
    • Archive: 263 repos (247 complete, 9 empty, 5 partial, 2 no-match)
    • Format: repo | npub | state_refs=N | git_refs=N | matches=N [| reason=X]
  • Decision: Phase 2 outputs ARE sufficient for Phase 3 processing
    • Existing data already categorized into 4 files
    • No need to create Phase 2 script immediately (can use Jan 22 data)
  • Implemented Phase 3 scripts:
    • 20-categorize.sh - Takes TSV input, outputs 4 category files
    • 21-compare-relays.sh - Compares prod vs archive categories
  • Tested both scripts successfully:
    • 20-categorize.sh correctly categorizes sample TSV data
    • 21-compare-relays.sh produces comparison with Jan 22 data:
      • Complete in both: 231 (no action needed)
      • Complete in prod, MISSING from archive: 276 (needs investigation)
      • Complete in prod, incomplete in archive: 2
      • Incomplete in both: 131
      • In archive only: 5
  • Updated how-to doc with correct script paths and output structure
  • Next: Phase 4 (log extraction) or Phase 5 (final classification)

2026-01-23 [Session 20:00]

  • Implemented structured debug logging for Phase 4 migration scripts
  • Added [PARSE_FAIL] log entries in src/nostr/builder.rs:
    • Format: [PARSE_FAIL] kind=X event_id=Y... reason="Z" repo=R npub=N
    • Logged when: announcement parsing fails, state event parsing fails, PR git data check fails
    • Includes repo identifier extracted from 'd' tag (announcements/states) or 'a' tag (PRs)
  • Added [PURGATORY_EXPIRED] log entries in src/purgatory/mod.rs:
    • Format: [PURGATORY_EXPIRED] repo=X npub=Y event_id=Z... kind=K reason="..."
    • Logged when: state events or PR events expire from purgatory without git data
    • Includes all fields needed by Phase 4 scripts
  • All 382 tests pass
  • Log format matches what Phase 4 scripts expect (30-extract-parse-failures.sh, 31-extract-purgatory-expiry.sh)
  • Next: Commit changes, then Phase 5 (final classification)

2026-01-23 [Session 11:40]

  • Implemented Phase 5 final classification script (40-classify-actions.sh)
    • Combines all data sources from Phases 1-4
    • Produces three output files: no-action-required.txt, action-required.txt, manual-investigation.txt
    • Generates summary.txt with breakdown by category and reason
  • Created orchestration script (run-migration-analysis.sh)
    • Runs all 5 phases in sequence with proper error handling
    • Parameterized inputs: relay URLs, git paths, service name, output directory
    • Phase control: --skip-phase-N, --only-phase-N, --from-phase-N
    • Dry-run mode to preview execution
    • Progress indicators and timing information
    • Auto-detects available features (git paths, journalctl)
  • Restructured migration guide (docs/how-to/migrate-ngit-relay-to-ngit-grasp.md)
    • Added Quick Start section with copy-paste commands
    • Added Prerequisites section with verification steps
    • Added Running the Analysis section with all options
    • Added Understanding Results section explaining output files
    • Added Troubleshooting section for common issues
    • Moved Architecture section (was at top) for those wanting details
    • Added Next Steps section for post-analysis workflow
  • All scripts committed:
    • 5dfd1cb - Add orchestration script for migration analysis pipeline
    • d8a88e8 - Restructure migration guide for practical usage
  • Next: Run analysis on relay.ngit.dev, review results

2026-01-23 [Session 11:50]

  • Reviewed structured logging implementation (commit 807961b)
  • Fixed multi-repo PR event handling:
    • PR events can reference multiple repositories (via multiple a tags)
    • Original code only logged the FIRST repo identifier
    • Updated extract_repos_from_pr_event to return ALL unique repos
    • Now logs once per repo for both [PARSE_FAIL] and [PURGATORY_EXPIRED]
  • Structured logging assessment:
    • Current format uses formatted strings (not tracing structured fields)
    • This is intentional - designed for grep/awk parsing by Phase 4 scripts
    • Proper structured logging would require script updates
    • Recommendation: Keep current format for migration, consider structured logging as future improvement
  • Script compatibility verified:
    • Phase 4 scripts will correctly parse multi-repo PR events as separate entries
    • No script changes needed
  • All 382 unit tests + 38 integration tests pass
  • Recommendation: Create low-priority issue for broader structured logging adoption (better observability, log aggregation)
  • Next: Commit multi-repo fix, then run analysis on relay.ngit.dev

2026-01-23 [Session 12:30]

  • Analyzed testing and deployment options for migration scripts
  • Created comprehensive strategy document: work/testing-and-deployment-strategy.md

Testing Options Analyzed:

  1. Local testing with Phase 1 data - Partial (can't test Phase 2, 4, or structured logging)
  2. Deploy VPS from branch - RECOMMENDED (complete validation possible)
  3. Local git repo in nixos-config - Not recommended (too complex, no benefit)

Structured Logging Testing:

  • Logging only triggers on parse failures and purgatory expiry (rare events)
  • Can create test scenarios: send malformed event for [PARSE_FAIL]
  • Purgatory expiry takes 30 minutes to trigger
  • Unit tests pass (382) but don't verify log format matches scripts

Deployment Strategy Recommendation:

  1. Push branch to remote: git push -u origin 4bc5-relay-ngit-dev-migration-v2
  2. Update nixos-config flake input to use branch: ?ref=4bc5-relay-ngit-dev-migration-v2
  3. Deploy to VPS: nixos-rebuild switch
  4. Run full migration analysis
  5. Validate structured logging with test event
  6. After validation: merge to main, update nixos-config back to main

Risk Assessment:

  • Low risk: changes are additive (logging only), all tests pass
  • Easy rollback: remove ?ref=... from flake input, rebuild (~5 min)
  • Timeline: ~2 hours for complete validation

2026-01-23 [Session 15:30]

  • Investigated why 315 repos didn't sync to archive

Root Cause Analysis:

  1. Event sync statistics:

    • Announcements (30617): 674/720 synced (94%) - working correctly
    • State events (30618): 230/671 synced (34%) - major gap
    • The 315 missing repos have announcements in archive but NO state events
  2. Key finding - GRASP relay correlation:

    • Repos using only git.shakespeare.diy (not gitnostr.com): 3x more likely to be missing
    • Synced repos using only git.shakespeare.diy: 60
    • Missing repos using only git.shakespeare.diy: 180
    • Repos using gitnostr.com: roughly evenly split (142 synced, 117 missing)
  3. Hypothesis - Archive domain configuration:

    • Archive relay likely configured with domain = "relay.ngit.dev"
    • SelfSubscriber skips syncing from relays containing relay_domain
    • This causes archive to skip syncing state events from wss://relay.ngit.dev
    • Archive tries to sync from other GRASP relays (git.shakespeare.diy, gitnostr.com)
    • gitnostr.com has some state events, so those repos sync
    • git.shakespeare.diy may not have state events or has connection issues
  4. Code evidence (src/sync/self_subscriber.rs:497-500):

    // Skip our own relay URL (we're subscribed to ourselves via self-subscription)
    if relay_url.contains(&self.relay_domain) {
        continue;
    }
    

Recommended Fix Options:

  1. Configure archive with different domain (e.g., archive.relay.ngit.dev)

    • Allows syncing from wss://relay.ngit.dev
    • Requires nixos-config change
  2. Modify SelfSubscriber to not skip bootstrap relay

    • Code change in ngit-grasp
    • More complex, affects all deployments
  3. Accept the gap and migrate anyway

    • Missing repos will re-sync when users push or when prod becomes archive's upstream
    • Simplest approach if gap is acceptable

Decision needed: Which approach to take for migration?

2026-01-23 [Session 14:00]

  • Deployed branch to VPS - Successfully updated nixos-config to use branch
  • Ran full migration analysis - All 5 phases completed successfully
  • Resolved issues during deployment:
    • Git not available in systemd service environment (added to PATH)
    • Script path discovery issues (fixed with proper directory detection)
    • Phase 2 script missing (created 10-check-git-sync.sh)
  • Updated migration guide with lessons learned and gotchas section

Analysis Results (relay.ngit.dev):

Category Count Notes
Complete in both 231 No action needed
Complete in prod, MISSING from archive 315 Needs investigation
Empty in both 100 Users never pushed
In archive only 4 Deleted from prod? Or new?
No match (refs differ) 1 Manual investigation
Purgatory expiry events 382 Logged successfully

Key Findings:

  1. 315 repos missing from archive - This is the main concern. Archive service is running but these repos didn't sync.
  2. 382 purgatory expiry events - Structured logging working correctly
  3. 5 repos need manual investigation - 4 in archive only, 1 with mismatched refs
  4. 100 empty repos - Expected (users created but never pushed)

Current State:

  • VPS running ngit-grasp from branch 4bc5-relay-ngit-dev-migration-v2
  • Archive service syncing from relay.ngit.dev (prod ngit-relay)
  • All migration scripts working correctly
  • Structured logging producing expected output

Outstanding Questions:

  1. Why didn't 315 repos sync to archive? Is this expected or a bug?
  2. What are the 4 repos in archive but not prod? (deleted? or new?)
  3. What's the 1 repo with mismatched refs?
  4. Should we do gradual cutover or full switch?
  5. When to merge branch to main?

2026-01-23 [Session 15:00 - Investigation & Script Fixes]

Investigated missing repos and parse failures:

  1. Fixed Phase 4 script issues:
    • Added validation to prevent using wrong service (ngit-relay vs ngit-grasp)
    • Fixed script bugs (SIGPIPE, count increment issues)
    • Added --analysis-root filter to scope parse failures to missing announcements only
    • Commits: b90c4a6, a968168, 1715e3c, 4dabb2b, 093f5ed

2026-01-26 [Session 20:00 - Parse Failure Format Fix]

Fixed critical usability bug in parse failure output:

  1. Root cause identified:

    • Phase 4 output: event_id | kind | reason | (empty repo) | (empty npub)
    • Phase 5 expected: repo | npub | kind | event_id | reason
    • Phase 5 was extracting columns 1-2 (event_id, kind) instead of columns 4-5 (repo, npub)
    • Result: action-required.txt showed unusable event IDs instead of repo names
  2. Fix implemented (commit 2e233b6):

    • Enhanced Phase 4 (30-extract-parse-failures.sh) with enrich_with_repo_npub() function
    • Builds lookup table from announcements.json mapping event_id → repo|npub
    • Uses jq to extract d-tag (repo) and pubkey from announcements
    • Optionally converts hex pubkeys to npub format using nak
    • Enriches parse failures by looking up event_id and populating repo/npub columns
    • Fixed Phase 5 column extraction from {print $1 "|" $2} to {print $4 "|" $5}
  3. Verification:

    • Reran Phases 4-5 on existing analysis data
    • 119 of 223 parse failures enriched with repo/npub (53%)
    • Remaining 104 have empty repo/npub (event_ids not in announcements.json)
    • Output now shows: bit2factor | npub13kkpy... | parse failure logged | fix event format
    • Instead of: 000014b2... | 30617 | parse failure logged | fix event format
  4. Impact:

    • Parse failure results now immediately actionable
    • Users can identify which repos have format issues
    • No need to manually look up event IDs

2026-01-26 [Session 21:00 - Classification System Redesign]

Redesigned classification system to eliminate overlap and confusion:

  1. Root cause of confusion:

    • Purgatory expiry was treated as a terminal category (no-action)
    • But purgatory is orthogonal to git sync status (it's context, not classification)
    • This caused 285 repos with complete data in prod to be marked "no action required"
  2. Design principles for new system:

    • Primary dimension: prod status (what data exists in source of truth)
    • Secondary dimension: archive status (what data exists in destination)
    • Tertiary: context flags (purgatory-expired, parse-failure, deleted)
    • Classification based on action needed, not technical state
  3. User feedback incorporated:

    • prod=cat2 (empty) is ALWAYS no action required
    • archive-only and not-in-prod moved to no-action (nothing to migrate)
    • needs-investigation moved to manual-review (requires human judgment)
    • Include purgatory context in needs-resync entries
  4. New category structure (Option B):

    • Tier 1: ready-for-migration.txt (352 repos, 50.8%)
      • Complete in both (199)
      • Deleted by user (15)
      • Empty in prod - any archive status (115)
      • Archive-only, not in prod (3)
      • Purgatory-only, not in prod (20)
    • Tier 2: needs-resync.txt (295 repos, 42.6%)
      • Complete in prod, missing from archive (294, 283 with purgatory-expired)
      • Complete in prod, incomplete in archive (1)
    • Tier 3: manual-review.txt (46 repos, 6.6%)
      • Partial in prod (24)
      • No-match in prod (5)
      • Parse failures (17)
  5. Key improvements:

    • 285 repos correctly moved from no-action to needs-resync
    • Purgatory context visible in needs-resync entries
    • No overlap between categories
    • Organized by action type, not technical state
    • Each repo appears in exactly one file
  6. Bug fixes during implementation:

    • Fixed bash arithmetic with set -e (changed ((count++)) to $((count + 1)))
    • Fixed NDJSON deletion processing (jq handling)
    • Optimized batch hex-to-npub conversion
  7. Output format:

    • repo | npub | prod_status | archive_status | context | action
    • Example: myrepo | npub1abc... | complete | missing | purgatory-expired | trigger re-sync to archive
  8. Parse failure investigation:

    • Initial count: 446 invalid announcements (WRONG - double-counting bug)
    • Root cause: Same event logged with hex ID and note1 (bech32) ID
    • Fixed deduplication: 223 unique invalid announcements
    • Further scoped to missing announcements: 18 parse failures for repos in prod but missing from archive
    • Reason: "Invalid announcement: multiple clone tags found" (NIP-34 format violation)
  9. Announcement reconciliation:

    • Production: 720 announcements
    • Archive: 674 announcements
    • Difference: 46 missing announcements
    • Parse failures (invalid format): 18 repos
    • Remaining unexplained: 28 announcements (46 - 18 = 28)
  10. Purgatory analysis:

    • 380 purgatory expiry events logged
    • 366 unique repos had announcements but git data never arrived
    • These repos are already cleaned up by the system

Latest Analysis Location:

  • VPS: /tmp/migration-analysis-20260123-152915/
  • Local: /tmp/migration-results-updated/
  • Key files:
    • logs/parse-failures.txt - 18 invalid announcements (scoped to missing)
    • logs/purgatory-expiry.txt - 380 purgatory events
    • results/summary.txt - Full breakdown
    • results/action-required.txt - Repos needing action
    • results/no-action-required.txt - Repos ready for migration

Key Findings (Analysis Results)

Latest Analysis (2026-01-23, after script fixes):

Analysis Location:

  • VPS: /tmp/migration-analysis-20260123-152915/
  • Local: /tmp/migration-results-updated/

Announcement-Level Analysis:

Metric Count Notes
Production announcements 720 Total kind 30617 events
Archive announcements 674 Total kind 30617 events
Missing from archive 46 720 - 674
Parse failures (invalid format) 18 "multiple clone tags" NIP-34 violation
Unexplained missing 28 46 - 18 = 28 (HIGH PRIORITY)

Repository-Level Analysis:

Metric Count Percentage
Total repos in prod 654 100%
Total repos in archive 268 41%
Complete in both 231 35%
Complete in prod, missing from archive 315 48%
Empty in both 100 15%
Manual investigation needed 5 1%

Purgatory Analysis:

Metric Count Notes
Purgatory expiry events 380 Total log entries
Unique repos expired 366 Git data never arrived
Reason "git data not received within 30 minutes" System cleaned up

Key Insights:

  1. 28 announcements unexplained - Not parse failures, not in purgatory, just missing
  2. 18 parse failures - Invalid format (multiple clone tags), fixable with validation change
  3. 366 purgatory-ejected repos - Expected behavior, system working correctly
  4. 231 repos ready - Complete in both systems, can migrate immediately

Next Session Plan

Investigation Priority Order

Current State:

  • Production: 720 announcements, 674 in archive
  • 46 announcements missing from archive (720 - 674)
  • 18 have parse failures (invalid format - "multiple clone tags")
  • 28 announcements unexplained (46 - 18)

Phase 1: Investigate 28 Unexplained Missing Announcements (HIGH PRIORITY)

Goal: Understand why 28 announcements are in production but missing from archive (not due to parse failures).

Tasks:

  1. Review the 28 missing announcements:
    • Get list from: /tmp/migration-analysis-20260123-152915/comparison/announcements-prod-not-archive.txt
    • Filter out the 18 with parse failures
    • Identify the remaining 28
  2. Check archive service logs for these specific announcements:
    • Were they received and rejected for other reasons?
    • Were they never synced from production?
    • Timing issues (created after archive started)?
  3. Determine root cause:
    • Sync timing/connectivity issues?
    • Different validation rules?
    • Archive mode configuration?

Phase 2: Investigate State Events Not in Purgatory (MEDIUM PRIORITY)

Goal: Understand repos that have announcements but state events didn't sync (not ejected from purgatory).

Context:

  • 315 repos complete in prod, missing from archive
  • 366 repos had purgatory expiry (git data never arrived)
  • Some repos may have state events that never entered purgatory

Tasks:

  1. Identify repos with announcements but no state events in archive
  2. Check if state events exist in production for these repos
  3. Determine why state events didn't sync:
    • Were they rejected before entering purgatory?
    • Were they never synced from production?
    • Archive mode filtering issue?

Phase 3: Review Purgatory-Ejected Repos (LOW PRIORITY)

Goal: Understand the 366 repos that were ejected from purgatory.

Context:

  • These repos had announcements but git data never arrived within 30 minutes
  • System already cleaned them up (expected behavior)

Tasks:

  1. Review purgatory expiry reasons:
    • "git data not received within 30 minutes" - most common
    • Any other patterns?
  2. Determine if these need action:
    • Are these user errors (never pushed)?
    • Are these sync failures (should retry)?
    • Are these expected (test repos, abandoned repos)?

Phase 4: Review Other Categories (LOW PRIORITY)

Goal: Address remaining edge cases.

Tasks:

  1. 5 manual investigation repos:
    • 4 repos in archive but not prod (deleted? new?)
    • 1 repo with mismatched refs (corruption?)
  2. 18 parse failures:
    • All have "multiple clone tags" format issue
    • Decision: Fix validation to accept both formats? Or notify users?
  3. Incomplete repos:
    • Repos with partial data in both systems
    • Determine if these need user action or system fixes

Phase 2: Decision Making (30 min)

Goal: Make key decisions about migration approach.

Decisions to make:

  1. Migration approach:

    • Option A: Full switch - Change DNS, disable archive mode, done
    • Option B: Gradual cutover - Run both in parallel, migrate users gradually
    • Recommendation: Full switch is simpler if we're confident in the data
  2. Handling 315 missing repos:

    • Option A: Accept the gap - These repos exist in prod, users can still access them
    • Option B: Manual sync - Copy git data from prod to archive before switch
    • Option C: Re-announce - Have users re-announce their repos after migration
  3. Branch merge timing:

    • Option A: Merge before migration - Cleaner, but can't easily rollback
    • Option B: Merge after migration - Can rollback to main if issues
    • Recommendation: Merge after successful migration validation
  4. Create migration checklist:

    • Pre-migration checks
    • Migration steps
    • Post-migration validation
    • Rollback procedure

Phase 3: Pre-Migration Preparation (1-2 hours)

Goal: Prepare everything needed for migration execution.

Tasks:

  1. Address investigation findings:

    • If 315 repos is a bug: fix it and re-run analysis
    • If expected: document and proceed
    • Handle the 5 manual investigation repos
  2. Re-run analysis if needed:

    • If any fixes were made, re-run to verify
    • Ensure numbers are stable
  3. Prepare rollback plan:

    • Document exact steps to revert to ngit-relay
    • Test rollback procedure (dry run)
    • Ensure backups are in place
  4. Create migration runbook:

    • Step-by-step commands
    • Expected outputs at each step
    • Verification checks
    • Contact info for escalation
  5. Notify stakeholders:

    • Announce maintenance window (if needed)
    • Prepare status page update

Phase 4: Migration Execution (1-2 hours)

Goal: Execute the migration and validate success.

Pre-Migration Checklist:

  • All investigation items resolved
  • Rollback plan documented and tested
  • Backups verified
  • Stakeholders notified

Migration Steps:

  1. Switch domain to ngit-grasp:

    • Update DNS or reverse proxy configuration
    • Point relay.ngit.dev to ngit-grasp service
    • Verify connectivity
  2. Disable archive mode:

    • Update ngit-grasp configuration
    • Restart service
    • Verify archive mode is disabled
  3. Validate services:

    • Test WebSocket connection to relay.ngit.dev
    • Test git clone/push operations
    • Verify existing repos are accessible
    • Check event propagation
  4. Monitor for issues:

    • Watch logs for errors
    • Monitor resource usage
    • Check for user reports

Post-Migration Validation:

  • WebSocket connections working
  • Git operations working
  • Existing repos accessible
  • New repos can be created
  • Events propagating correctly
  • No error spikes in logs

Phase 5: Post-Migration Cleanup (30 min)

Goal: Finalize migration and clean up.

Tasks:

  1. Merge branch to main:

    • Create PR from 4bc5-relay-ngit-dev-migration-v2 to main
    • Review changes
    • Merge and delete branch
  2. Update nixos-config:

    • Remove ?ref=... from flake input
    • Point back to main branch
    • Rebuild to verify
  3. Update documentation:

    • Mark migration guide as tested/validated
    • Add any lessons learned
    • Update architecture docs if needed
  4. Archive analysis results:

    • Save analysis output for future reference
    • Document final state
  5. Close issue:

    • Update progress with final status
    • Move to closed/ directory

Success Criteria:

  • relay.ngit.dev running ngit-grasp
  • All existing repos accessible
  • New repos can be created
  • No degradation in service
  • Branch merged to main
  • Issue closed

Contingency Plans

If 315 missing repos is a critical bug:

  • Pause migration
  • Fix the bug in ngit-grasp
  • Re-deploy and re-run analysis
  • Resume migration when fixed

If migration causes issues:

  • Execute rollback plan (revert to ngit-relay)
  • Investigate root cause
  • Fix and retry

If archive data is corrupted:

  • Restore from backup
  • Re-sync from prod
  • Retry migration

Notes

  • Related issue: 820a-relay-ngit-dev-migration.md (paused, in paused/ directory)
  • VPS: Running NixOS
  • Old worktree: Can reference /persistent/dcdev/clones/ngit-grasp/worktrees/820a-relay-ngit-dev-migration/ for existing scripts and learnings
  • Target: Simple, practical migration guide that works
  • Branch deployed: 4bc5-relay-ngit-dev-migration-v2 currently running on VPS