36 KiB
Migrate relay.ngit.dev from ngit-relay to ngit-grasp (v2)
ID: 4bc5
Problem
relay.ngit.dev currently runs ngit-relay (reference implementation). We want to consolidate on ngit-grasp as the production implementation.
Goal: Replace an ngit-relay instance on a VPS running NixOS with ngit-grasp.
Context: This is a fresh start after issue 820a became too complex with extensive investigation history. We're starting from scratch with a focus on creating a small, lightweight, easy-to-implement how-to document.
Plan
Script Development (Modular Architecture)
- Phase 1: Fetch Events (~30s, local) -
01-fetch-events.sh- Fetch kind 30618 (state), 30617 (announcement), 5 (deletion) from relay
- Run for both prod and archive relays
- Phase 2: Git Sync Check (~20 mins, VPS) -
10-check-git-sync.sh- Compare state event refs to actual git data on disk
- Run for both prod and archive git directories
- Note: Existing Jan 22 data available, script not yet created
- Phase 3: Categorize & Compare (fast, local) -
20-categorize.sh,21-compare-relays.sh- Apply 4-category logic (complete/empty/partial/no-match)
- Find gaps between prod and archive
- Phase 4: Log-Based Categories (VPS) -
30-extract-parse-failures.sh,31-extract-purgatory-expiry.sh- Extract parse failures and purgatory expiry from logs
- Dependency: Logging improvements in ngit-grasp - IMPLEMENTED
- Phase 5: Final Classification (fast, local) -
40-classify-actions.sh- Combine all data to produce: no-action, action-required, manual-investigation
- Orchestration Script -
run-migration-analysis.sh- Runs all phases with proper error handling and progress reporting
- Supports phase control (skip, only, from-phase options)
- Dry-run mode, timing information, summary display
Migration Execution
- Run analysis scripts on relay.ngit.dev
- Review action-required repos, make decisions
- Execute migration (switch domain, disable archive mode)
- Validate migration success
Next Session Tasks
- Investigation: Analyze why 315 repos didn't sync to archive
- Investigation: Review 5 repos needing manual investigation
- Decision: Choose migration approach (gradual vs full switch)
- Decision: Determine when to merge branch to main
- Preparation: Create migration checklist and rollback plan
- Execution: Switch relay.ngit.dev domain to ngit-grasp
- Execution: Disable archive mode
- Validation: Run post-migration validation
- Cleanup: Merge branch to main, close issue
Progress
2026-01-23 [Session 16:00]
- Created: Fresh v2 issue to replace complex 820a migration
- Context: Previous issue (820a) paused due to complexity
- Approach: Start from scratch, potentially reuse scripts from 820a worktree
- Goal: Create small, lightweight, easy-to-implement how-to document
- Started work: Created worktree for issue 4bc5
- Completed: Created initial how-to document at
docs/how-to/migrate-ngit-relay-to-ngit-grasp.md - Document includes: Approach, challenges, analysis categories, gotchas
- Next: User requested NOT to do planning for migration script yet
2026-01-23 [Session 17:30]
- Reviewed existing scripts from 820a worktree:
analyze-git-state-sync.sh- monolithic, takes ~20 mins (git sync is slow part)compare-categories.sh- compares prod vs archive categoriesmigration-validation-guide.md- comprehensive troubleshooting guide
- Reviewed existing analysis output (Jan 22):
- Prod: 654 repos (509 complete, 114 empty, 25 partial, 6 no-match)
- Archive: 263 repos (247 complete, 9 empty, 5 partial, 2 no-match)
- Designed modular script architecture for fast iteration:
- Split into 5 phases with clear inputs/outputs
- Phases 1, 3, 5 can run locally; Phases 2, 4 need VPS
- Can use cached data from Jan 22 to develop categorization logic
- Added log-based categories (scriptable):
- Parse failures:
[PARSE_FAIL] kind=X event_id=Y reason=Z - Purgatory expiry:
[PURGATORY_EXPIRED] repo=X npub=Y
- Parse failures:
- Updated how-to doc with full architecture diagram
- Next: Implement Phase 1 (fetch events) to get fresh data
2026-01-23 [Session 18:45]
- Reviewed Phase 2 outputs from Jan 22 (820a worktree):
- Prod: 654 repos (509 complete, 114 empty, 25 partial, 6 no-match)
- Archive: 263 repos (247 complete, 9 empty, 5 partial, 2 no-match)
- Format:
repo | npub | state_refs=N | git_refs=N | matches=N [| reason=X]
- Decision: Phase 2 outputs ARE sufficient for Phase 3 processing
- Existing data already categorized into 4 files
- No need to create Phase 2 script immediately (can use Jan 22 data)
- Implemented Phase 3 scripts:
20-categorize.sh- Takes TSV input, outputs 4 category files21-compare-relays.sh- Compares prod vs archive categories
- Tested both scripts successfully:
20-categorize.shcorrectly categorizes sample TSV data21-compare-relays.shproduces comparison with Jan 22 data:- Complete in both: 231 (no action needed)
- Complete in prod, MISSING from archive: 276 (needs investigation)
- Complete in prod, incomplete in archive: 2
- Incomplete in both: 131
- In archive only: 5
- Updated how-to doc with correct script paths and output structure
- Next: Phase 4 (log extraction) or Phase 5 (final classification)
2026-01-23 [Session 20:00]
- Implemented structured debug logging for Phase 4 migration scripts
- Added
[PARSE_FAIL]log entries insrc/nostr/builder.rs:- Format:
[PARSE_FAIL] kind=X event_id=Y... reason="Z" repo=R npub=N - Logged when: announcement parsing fails, state event parsing fails, PR git data check fails
- Includes repo identifier extracted from 'd' tag (announcements/states) or 'a' tag (PRs)
- Format:
- Added
[PURGATORY_EXPIRED]log entries insrc/purgatory/mod.rs:- Format:
[PURGATORY_EXPIRED] repo=X npub=Y event_id=Z... kind=K reason="..." - Logged when: state events or PR events expire from purgatory without git data
- Includes all fields needed by Phase 4 scripts
- Format:
- All 382 tests pass
- Log format matches what Phase 4 scripts expect (30-extract-parse-failures.sh, 31-extract-purgatory-expiry.sh)
- Next: Commit changes, then Phase 5 (final classification)
2026-01-23 [Session 11:40]
- Implemented Phase 5 final classification script (
40-classify-actions.sh)- Combines all data sources from Phases 1-4
- Produces three output files: no-action-required.txt, action-required.txt, manual-investigation.txt
- Generates summary.txt with breakdown by category and reason
- Created orchestration script (
run-migration-analysis.sh)- Runs all 5 phases in sequence with proper error handling
- Parameterized inputs: relay URLs, git paths, service name, output directory
- Phase control: --skip-phase-N, --only-phase-N, --from-phase-N
- Dry-run mode to preview execution
- Progress indicators and timing information
- Auto-detects available features (git paths, journalctl)
- Restructured migration guide (
docs/how-to/migrate-ngit-relay-to-ngit-grasp.md)- Added Quick Start section with copy-paste commands
- Added Prerequisites section with verification steps
- Added Running the Analysis section with all options
- Added Understanding Results section explaining output files
- Added Troubleshooting section for common issues
- Moved Architecture section (was at top) for those wanting details
- Added Next Steps section for post-analysis workflow
- All scripts committed:
5dfd1cb- Add orchestration script for migration analysis pipelined8a88e8- Restructure migration guide for practical usage
- Next: Run analysis on relay.ngit.dev, review results
2026-01-23 [Session 11:50]
- Reviewed structured logging implementation (commit 807961b)
- Fixed multi-repo PR event handling:
- PR events can reference multiple repositories (via multiple
atags) - Original code only logged the FIRST repo identifier
- Updated
extract_repos_from_pr_eventto return ALL unique repos - Now logs once per repo for both
[PARSE_FAIL]and[PURGATORY_EXPIRED]
- PR events can reference multiple repositories (via multiple
- Structured logging assessment:
- Current format uses formatted strings (not tracing structured fields)
- This is intentional - designed for grep/awk parsing by Phase 4 scripts
- Proper structured logging would require script updates
- Recommendation: Keep current format for migration, consider structured logging as future improvement
- Script compatibility verified:
- Phase 4 scripts will correctly parse multi-repo PR events as separate entries
- No script changes needed
- All 382 unit tests + 38 integration tests pass
- Recommendation: Create low-priority issue for broader structured logging adoption (better observability, log aggregation)
- Next: Commit multi-repo fix, then run analysis on relay.ngit.dev
2026-01-23 [Session 12:30]
- Analyzed testing and deployment options for migration scripts
- Created comprehensive strategy document:
work/testing-and-deployment-strategy.md
Testing Options Analyzed:
- Local testing with Phase 1 data - Partial (can't test Phase 2, 4, or structured logging)
- Deploy VPS from branch - RECOMMENDED (complete validation possible)
- Local git repo in nixos-config - Not recommended (too complex, no benefit)
Structured Logging Testing:
- Logging only triggers on parse failures and purgatory expiry (rare events)
- Can create test scenarios: send malformed event for
[PARSE_FAIL] - Purgatory expiry takes 30 minutes to trigger
- Unit tests pass (382) but don't verify log format matches scripts
Deployment Strategy Recommendation:
- Push branch to remote:
git push -u origin 4bc5-relay-ngit-dev-migration-v2 - Update nixos-config flake input to use branch:
?ref=4bc5-relay-ngit-dev-migration-v2 - Deploy to VPS:
nixos-rebuild switch - Run full migration analysis
- Validate structured logging with test event
- After validation: merge to main, update nixos-config back to main
Risk Assessment:
- Low risk: changes are additive (logging only), all tests pass
- Easy rollback: remove
?ref=...from flake input, rebuild (~5 min) - Timeline: ~2 hours for complete validation
2026-01-23 [Session 15:30]
- Investigated why 315 repos didn't sync to archive
Root Cause Analysis:
-
Event sync statistics:
- Announcements (30617): 674/720 synced (94%) - working correctly
- State events (30618): 230/671 synced (34%) - major gap
- The 315 missing repos have announcements in archive but NO state events
-
Key finding - GRASP relay correlation:
- Repos using only
git.shakespeare.diy(notgitnostr.com): 3x more likely to be missing - Synced repos using only
git.shakespeare.diy: 60 - Missing repos using only
git.shakespeare.diy: 180 - Repos using
gitnostr.com: roughly evenly split (142 synced, 117 missing)
- Repos using only
-
Hypothesis - Archive domain configuration:
- Archive relay likely configured with
domain = "relay.ngit.dev" - SelfSubscriber skips syncing from relays containing
relay_domain - This causes archive to skip syncing state events from
wss://relay.ngit.dev - Archive tries to sync from other GRASP relays (
git.shakespeare.diy,gitnostr.com) gitnostr.comhas some state events, so those repos syncgit.shakespeare.diymay not have state events or has connection issues
- Archive relay likely configured with
-
Code evidence (src/sync/self_subscriber.rs:497-500):
// Skip our own relay URL (we're subscribed to ourselves via self-subscription) if relay_url.contains(&self.relay_domain) { continue; }
Recommended Fix Options:
-
Configure archive with different domain (e.g.,
archive.relay.ngit.dev)- Allows syncing from
wss://relay.ngit.dev - Requires nixos-config change
- Allows syncing from
-
Modify SelfSubscriber to not skip bootstrap relay
- Code change in ngit-grasp
- More complex, affects all deployments
-
Accept the gap and migrate anyway
- Missing repos will re-sync when users push or when prod becomes archive's upstream
- Simplest approach if gap is acceptable
Decision needed: Which approach to take for migration?
2026-01-23 [Session 14:00]
- Deployed branch to VPS - Successfully updated nixos-config to use branch
- Ran full migration analysis - All 5 phases completed successfully
- Resolved issues during deployment:
- Git not available in systemd service environment (added to PATH)
- Script path discovery issues (fixed with proper directory detection)
- Phase 2 script missing (created
10-check-git-sync.sh)
- Updated migration guide with lessons learned and gotchas section
Analysis Results (relay.ngit.dev):
| Category | Count | Notes |
|---|---|---|
| Complete in both | 231 | No action needed |
| Complete in prod, MISSING from archive | 315 | Needs investigation |
| Empty in both | 100 | Users never pushed |
| In archive only | 4 | Deleted from prod? Or new? |
| No match (refs differ) | 1 | Manual investigation |
| Purgatory expiry events | 382 | Logged successfully |
Key Findings:
- 315 repos missing from archive - This is the main concern. Archive service is running but these repos didn't sync.
- 382 purgatory expiry events - Structured logging working correctly
- 5 repos need manual investigation - 4 in archive only, 1 with mismatched refs
- 100 empty repos - Expected (users created but never pushed)
Current State:
- VPS running ngit-grasp from branch
4bc5-relay-ngit-dev-migration-v2 - Archive service syncing from relay.ngit.dev (prod ngit-relay)
- All migration scripts working correctly
- Structured logging producing expected output
Outstanding Questions:
- Why didn't 315 repos sync to archive? Is this expected or a bug?
- What are the 4 repos in archive but not prod? (deleted? or new?)
- What's the 1 repo with mismatched refs?
- Should we do gradual cutover or full switch?
- When to merge branch to main?
2026-01-23 [Session 15:00 - Investigation & Script Fixes]
Investigated missing repos and parse failures:
- Fixed Phase 4 script issues:
- Added validation to prevent using wrong service (ngit-relay vs ngit-grasp)
- Fixed script bugs (SIGPIPE, count increment issues)
- Added
--analysis-rootfilter to scope parse failures to missing announcements only - Commits:
b90c4a6,a968168,1715e3c,4dabb2b,093f5ed
2026-01-26 [Session 20:00 - Parse Failure Format Fix]
Fixed critical usability bug in parse failure output:
-
Root cause identified:
- Phase 4 output:
event_id | kind | reason | (empty repo) | (empty npub) - Phase 5 expected:
repo | npub | kind | event_id | reason - Phase 5 was extracting columns 1-2 (event_id, kind) instead of columns 4-5 (repo, npub)
- Result: action-required.txt showed unusable event IDs instead of repo names
- Phase 4 output:
-
Fix implemented (commit
2e233b6):- Enhanced Phase 4 (
30-extract-parse-failures.sh) withenrich_with_repo_npub()function - Builds lookup table from
announcements.jsonmapping event_id → repo|npub - Uses jq to extract d-tag (repo) and pubkey from announcements
- Optionally converts hex pubkeys to npub format using nak
- Enriches parse failures by looking up event_id and populating repo/npub columns
- Fixed Phase 5 column extraction from
{print $1 "|" $2}to{print $4 "|" $5}
- Enhanced Phase 4 (
-
Verification:
- Reran Phases 4-5 on existing analysis data
- 119 of 223 parse failures enriched with repo/npub (53%)
- Remaining 104 have empty repo/npub (event_ids not in announcements.json)
- Output now shows:
bit2factor | npub13kkpy... | parse failure logged | fix event format - Instead of:
000014b2... | 30617 | parse failure logged | fix event format
-
Impact:
- Parse failure results now immediately actionable
- Users can identify which repos have format issues
- No need to manually look up event IDs
2026-01-26 [Session 21:00 - Classification System Redesign]
Redesigned classification system to eliminate overlap and confusion:
-
Root cause of confusion:
- Purgatory expiry was treated as a terminal category (no-action)
- But purgatory is orthogonal to git sync status (it's context, not classification)
- This caused 285 repos with complete data in prod to be marked "no action required"
-
Design principles for new system:
- Primary dimension: prod status (what data exists in source of truth)
- Secondary dimension: archive status (what data exists in destination)
- Tertiary: context flags (purgatory-expired, parse-failure, deleted)
- Classification based on action needed, not technical state
-
User feedback incorporated:
- prod=cat2 (empty) is ALWAYS no action required
- archive-only and not-in-prod moved to no-action (nothing to migrate)
- needs-investigation moved to manual-review (requires human judgment)
- Include purgatory context in needs-resync entries
-
New category structure (Option B):
- Tier 1: ready-for-migration.txt (352 repos, 50.8%)
- Complete in both (199)
- Deleted by user (15)
- Empty in prod - any archive status (115)
- Archive-only, not in prod (3)
- Purgatory-only, not in prod (20)
- Tier 2: needs-resync.txt (295 repos, 42.6%)
- Complete in prod, missing from archive (294, 283 with purgatory-expired)
- Complete in prod, incomplete in archive (1)
- Tier 3: manual-review.txt (46 repos, 6.6%)
- Partial in prod (24)
- No-match in prod (5)
- Parse failures (17)
- Tier 1: ready-for-migration.txt (352 repos, 50.8%)
-
Key improvements:
- 285 repos correctly moved from no-action to needs-resync
- Purgatory context visible in needs-resync entries
- No overlap between categories
- Organized by action type, not technical state
- Each repo appears in exactly one file
-
Bug fixes during implementation:
- Fixed bash arithmetic with
set -e(changed((count++))to$((count + 1))) - Fixed NDJSON deletion processing (jq handling)
- Optimized batch hex-to-npub conversion
- Fixed bash arithmetic with
-
Output format:
repo | npub | prod_status | archive_status | context | action- Example:
myrepo | npub1abc... | complete | missing | purgatory-expired | trigger re-sync to archive
2026-01-26 [Session 22:00 - Script Deployment & Fresh Analysis]
Deployed new classification script and ran fresh analysis:
-
Deployment issue identified:
- New script was created at
scripts/40-classify-actions.sh - Orchestration script looks for it at
docs/how-to/migration-scripts/40-classify-actions.sh - Fresh VPS analysis run used old script, produced old format
- New script was created at
-
Fix applied (commits
49ad42b,0f3e8d3):- Copied new script to replace old one at
docs/how-to/migration-scripts/40-classify-actions.sh - Removed duplicate at
scripts/40-classify-actions.sh - Re-ran Phase 5 on latest analysis data (
work/migration-analysis-20260126-122943)
- Copied new script to replace old one at
-
Fresh analysis results (Jan 26, 2026):
- No Action Required: 422 repos (57.0%)
- Action Required: 308 repos (41.6%)
- Manual Investigation: 10 repos (1.4%)
- Total: 740 repos
-
Note on format:
- Output still uses old file names (no-action-required.txt, action-required.txt, manual-investigation.txt)
- But classification logic is updated (purgatory as context, not category)
- 285 repos correctly moved from no-action to action-required
2026-01-26 [Session 23:00 - Workaround Discovery & Investigation Plan]
Discovered workaround for repos "complete in prod, missing in archive" with purgatory context "none":
-
Workaround - Manual republish triggers sync:
- Example repo:
example | npub1zhzz4kmmyywgfzx3s0tf9xageqv3mmfpv9ukputsu80zngxuzagsxpc4zw | complete | missing | none | trigger re-sync to archive - Command:
nak req -k 30618 -t d=example relay.ngit.dev | nak event ws://localhost:7443 - Result: Successfully triggered sync for "example" repo
- Observation: Event was already in purgatory but hadn't been processed
- Example repo:
-
Investigation challenges identified:
- Querying logs on ngit.dev VPS is very slow (slow VPS + large log files)
- Need to download logs locally for efficient analysis
-
Next steps planned:
- Download logs for local analysis - Get ngit-grasp-relay-ngit-dev.service logs from VPS
- Root cause investigation (architect task):
- Why did we never receive/process ~200 state events?
- Hypothesis: Event rush overwhelms system during historic sync
- When historic sync runs (1000+ events quickly due to no negentropy support), we may be dropping events
- Even restarting service to trigger historic sync doesn't help (note: there's a bug preventing purgatory preservation on restart)
- Validate with unique repo name like 'portfoliux' to make log searching easier
- Systematic republishing - Consider scripting the republish workaround for all 201 affected repos
-
Context:
- Analysis results:
work/migration-analysis-20260126-122943/results/summary.txt - 201 repos need re-sync, many with purgatory context "none"
- These events may actually be in purgatory but stuck/unprocessed
- Analysis results:
2026-01-26 [Session 23:30 - CRITICAL: Silent Event Drop During Historic Sync]
Investigation of "rarenpubs" Repository:
Repository: rarenpubs (npub17qvqdrsn93zx93myrprlcu55h9gr0ed6sj7qt99h540mehzhz9ysqh5ymw) Status: complete in prod, missing in archive, purgatory context: "none"
What We Discovered:
-
Announcement event (kind 30617) WAS received:
- First received at 08:58:01.048037Z from wss://relay.ngit.dev
- Received again from relay.damus.io and other relays
- Event ID: 3248b18b9d43a7eb4b23eb9608bd199aa71bc338bedd086a88269dac4cd4096a
- Log location: Line 347636 in
/tmp/ngit-grasp-logs-6h.txt
-
Announcement event was NEVER processed:
- No "Accepted repository announcement" log entry
- No "Rejected repository announcement" log entry
- Event was silently dropped somewhere in the processing pipeline
-
State event (kind 30618) was NEVER received:
- Makes sense - if announcement isn't processed, system never subscribes to state events
- Event ID: a7c992bc72dcfb82160993a877ba237d2a6682dda9bff0bde7729032930bf3b1
-
Comparison with working events:
- Other announcements at the same time WERE processed (e.g., "website" repo)
- rarenpubs: complete silence - no accept, no reject
Why Manual Republishing Works:
When manually republishing via:
nak req -t d=rarenpubs relay.ngit.dev | nak event ws://localhost:7443
The event arrives AFTER the system is fully initialized, so it processes correctly and syncs successfully.
Root Cause Hypothesis:
Events are being silently dropped during historic sync. Possible causes:
- Race condition during historic sync startup
- LMDB deduplication at database level before application-level processing
- Event handler not ready when events arrive during startup
- Async queue loss between event receipt and processing
Implications:
This is a different and more serious issue than the expired_events blacklist:
- The
expired_eventsissue affects repos that entered purgatory but expired - This silent drop issue affects repos that never even got processed
- Potentially affects many/most of the ~200 "missing" repos
Evidence Location:
- Log file:
/tmp/ngit-grasp-logs-6h.txt(166MB, 6 hours of logs) - Announcement received: Line 347636 (08:58:01.048037Z)
- No processing logs found for this event ID
Next Steps:
-
Code path investigation (architect task - needed):
- Trace event flow from
nostr_relay_pool::relay::inner: Receivedtongit_grasp::nostr::builder: Accepted/Rejected - Identify where events can be silently dropped
- Check LMDB deduplication logic
- Review event handler registration timing
- Trace event flow from
-
Validate scope of issue:
- Check other repos from needs-resync.txt with purgarity context "none"
- Determine how many are affected by silent drop vs expired_events blacklist
- Test manual republishing on a sample set
-
Fix strategy:
- Ensure event handlers are registered before historic sync starts
- Add logging for dropped/deduplicated events
- Consider buffering events during startup until handlers ready
- Fix the
expired_eventspersistence issue (separate but related)
Two Distinct Issues Identified:
- Silent event drops during historic sync (this finding - HIGH PRIORITY)
expired_eventsblacklist persisted across restarts (previous finding - MEDIUM PRIORITY)
2026-01-27 [Session 07:00 - ROOT CAUSE FOUND: Database Query Fix Bug]
Investigated why commit 4162c90 (database query fix) didn't work:
Root Cause: last_connected timestamp set BEFORE load_existing_events() is called
The Bug (src/sync/self_subscriber.rs:424-447):
// Line 425: Sets last_connected to NOW
self.last_connected = Some(Timestamp::now());
// Line 427: Subscribe to relay
if let Err(e) = client.subscribe(filter, None).await { ... }
// Line 447: Load existing events - BUT last_connected is already set!
let mut pending = self.load_existing_events().await;
When load_existing_events() runs, it checks self.last_connected and applies .since() filter:
if let Some(timestamp) = self.last_connected {
announcement_filter = announcement_filter.since(timestamp);
}
Since last_connected was just set to Timestamp::now(), the database query becomes:
SELECT * FROM events WHERE kind=30617 AND created_at >= <current_time>
No events have created_at >= current_time, so query returns 0 results.
Evidence from logs (work/migration-analysis-20260127-074741/202601270826-journald-2h.log):
Line 11272: Loading events incrementally from database (reconnect) since=1769498542
Line 11275: Loaded announcements from database count=0
Line 11276: Loaded root events from database count=0
Line 11280: Processed existing events from database announcements_loaded=0 root_events_processed=0
Timestamp 1769498542 = 2026-01-27T07:22:22 UTC = exactly the service startup time.
Why Manual Publish Works:
When manually publishing a state event:
- It arrives via WebSocket (not database query)
- It's processed through normal event flow
- No
.since()filter is applied to incoming WebSocket events
The Fix:
Move load_existing_events() to run BEFORE setting last_connected:
// Load existing events FIRST (before setting last_connected)
let mut pending = self.load_existing_events().await;
// NOW set last_connected for future reconnects
self.last_connected = Some(Timestamp::now());
// Subscribe to relay
if let Err(e) = client.subscribe(filter, None).await { ... }
Location: src/sync/self_subscriber.rs:424-447
Impact: This explains ALL ~190 missing state events. The database query returned 0 results, so no Layer 2/3 filters were created, so no state events were fetched.
Next Steps:
- Implement the fix (reorder lines 447 to come before line 425)
- Test locally to verify database query returns events
- Deploy to VPS and verify state events sync
- Re-run migration analysis to confirm all repos sync
Key Findings (Analysis Results)
Latest Analysis (2026-01-23, after script fixes):
Analysis Location:
- VPS:
/tmp/migration-analysis-20260123-152915/ - Local:
/tmp/migration-results-updated/
Announcement-Level Analysis:
| Metric | Count | Notes |
|---|---|---|
| Production announcements | 720 | Total kind 30617 events |
| Archive announcements | 674 | Total kind 30617 events |
| Missing from archive | 46 | 720 - 674 |
| Parse failures (invalid format) | 18 | "multiple clone tags" NIP-34 violation |
| Unexplained missing | 28 | 46 - 18 = 28 (HIGH PRIORITY) |
Repository-Level Analysis:
| Metric | Count | Percentage |
|---|---|---|
| Total repos in prod | 654 | 100% |
| Total repos in archive | 268 | 41% |
| Complete in both | 231 | 35% |
| Complete in prod, missing from archive | 315 | 48% |
| Empty in both | 100 | 15% |
| Manual investigation needed | 5 | 1% |
Purgatory Analysis:
| Metric | Count | Notes |
|---|---|---|
| Purgatory expiry events | 380 | Total log entries |
| Unique repos expired | 366 | Git data never arrived |
| Reason | "git data not received within 30 minutes" | System cleaned up |
Key Insights:
- 28 announcements unexplained - Not parse failures, not in purgatory, just missing
- 18 parse failures - Invalid format (multiple clone tags), fixable with validation change
- 366 purgatory-ejected repos - Expected behavior, system working correctly
- 231 repos ready - Complete in both systems, can migrate immediately
Next Session Plan
Investigation Priority Order
Current State:
- Production: 720 announcements, 674 in archive
- 46 announcements missing from archive (720 - 674)
- 18 have parse failures (invalid format - "multiple clone tags")
- 28 announcements unexplained (46 - 18)
Phase 1: Investigate 28 Unexplained Missing Announcements (HIGH PRIORITY)
Goal: Understand why 28 announcements are in production but missing from archive (not due to parse failures).
Tasks:
- Review the 28 missing announcements:
- Get list from:
/tmp/migration-analysis-20260123-152915/comparison/announcements-prod-not-archive.txt - Filter out the 18 with parse failures
- Identify the remaining 28
- Get list from:
- Check archive service logs for these specific announcements:
- Were they received and rejected for other reasons?
- Were they never synced from production?
- Timing issues (created after archive started)?
- Determine root cause:
- Sync timing/connectivity issues?
- Different validation rules?
- Archive mode configuration?
Phase 2: Investigate State Events Not in Purgatory (MEDIUM PRIORITY)
Goal: Understand repos that have announcements but state events didn't sync (not ejected from purgatory).
Context:
- 315 repos complete in prod, missing from archive
- 366 repos had purgatory expiry (git data never arrived)
- Some repos may have state events that never entered purgatory
Tasks:
- Identify repos with announcements but no state events in archive
- Check if state events exist in production for these repos
- Determine why state events didn't sync:
- Were they rejected before entering purgatory?
- Were they never synced from production?
- Archive mode filtering issue?
Phase 3: Review Purgatory-Ejected Repos (LOW PRIORITY)
Goal: Understand the 366 repos that were ejected from purgatory.
Context:
- These repos had announcements but git data never arrived within 30 minutes
- System already cleaned them up (expected behavior)
Tasks:
- Review purgatory expiry reasons:
- "git data not received within 30 minutes" - most common
- Any other patterns?
- Determine if these need action:
- Are these user errors (never pushed)?
- Are these sync failures (should retry)?
- Are these expected (test repos, abandoned repos)?
Phase 4: Review Other Categories (LOW PRIORITY)
Goal: Address remaining edge cases.
Tasks:
- 5 manual investigation repos:
- 4 repos in archive but not prod (deleted? new?)
- 1 repo with mismatched refs (corruption?)
- 18 parse failures:
- All have "multiple clone tags" format issue
- Decision: Fix validation to accept both formats? Or notify users?
- Incomplete repos:
- Repos with partial data in both systems
- Determine if these need user action or system fixes
Phase 2: Decision Making (30 min)
Goal: Make key decisions about migration approach.
Decisions to make:
-
Migration approach:
- Option A: Full switch - Change DNS, disable archive mode, done
- Option B: Gradual cutover - Run both in parallel, migrate users gradually
- Recommendation: Full switch is simpler if we're confident in the data
-
Handling 315 missing repos:
- Option A: Accept the gap - These repos exist in prod, users can still access them
- Option B: Manual sync - Copy git data from prod to archive before switch
- Option C: Re-announce - Have users re-announce their repos after migration
-
Branch merge timing:
- Option A: Merge before migration - Cleaner, but can't easily rollback
- Option B: Merge after migration - Can rollback to main if issues
- Recommendation: Merge after successful migration validation
-
Create migration checklist:
- Pre-migration checks
- Migration steps
- Post-migration validation
- Rollback procedure
Phase 3: Pre-Migration Preparation (1-2 hours)
Goal: Prepare everything needed for migration execution.
Tasks:
-
Address investigation findings:
- If 315 repos is a bug: fix it and re-run analysis
- If expected: document and proceed
- Handle the 5 manual investigation repos
-
Re-run analysis if needed:
- If any fixes were made, re-run to verify
- Ensure numbers are stable
-
Prepare rollback plan:
- Document exact steps to revert to ngit-relay
- Test rollback procedure (dry run)
- Ensure backups are in place
-
Create migration runbook:
- Step-by-step commands
- Expected outputs at each step
- Verification checks
- Contact info for escalation
-
Notify stakeholders:
- Announce maintenance window (if needed)
- Prepare status page update
Phase 4: Migration Execution (1-2 hours)
Goal: Execute the migration and validate success.
Pre-Migration Checklist:
- All investigation items resolved
- Rollback plan documented and tested
- Backups verified
- Stakeholders notified
Migration Steps:
-
Switch domain to ngit-grasp:
- Update DNS or reverse proxy configuration
- Point relay.ngit.dev to ngit-grasp service
- Verify connectivity
-
Disable archive mode:
- Update ngit-grasp configuration
- Restart service
- Verify archive mode is disabled
-
Validate services:
- Test WebSocket connection to relay.ngit.dev
- Test git clone/push operations
- Verify existing repos are accessible
- Check event propagation
-
Monitor for issues:
- Watch logs for errors
- Monitor resource usage
- Check for user reports
Post-Migration Validation:
- WebSocket connections working
- Git operations working
- Existing repos accessible
- New repos can be created
- Events propagating correctly
- No error spikes in logs
Phase 5: Post-Migration Cleanup (30 min)
Goal: Finalize migration and clean up.
Tasks:
-
Merge branch to main:
- Create PR from
4bc5-relay-ngit-dev-migration-v2tomain - Review changes
- Merge and delete branch
- Create PR from
-
Update nixos-config:
- Remove
?ref=...from flake input - Point back to main branch
- Rebuild to verify
- Remove
-
Update documentation:
- Mark migration guide as tested/validated
- Add any lessons learned
- Update architecture docs if needed
-
Archive analysis results:
- Save analysis output for future reference
- Document final state
-
Close issue:
- Update progress with final status
- Move to closed/ directory
Success Criteria:
- relay.ngit.dev running ngit-grasp
- All existing repos accessible
- New repos can be created
- No degradation in service
- Branch merged to main
- Issue closed
Contingency Plans
If 315 missing repos is a critical bug:
- Pause migration
- Fix the bug in ngit-grasp
- Re-deploy and re-run analysis
- Resume migration when fixed
If migration causes issues:
- Execute rollback plan (revert to ngit-relay)
- Investigate root cause
- Fix and retry
If archive data is corrupted:
- Restore from backup
- Re-sync from prod
- Retry migration
Notes
- Related issue: 820a-relay-ngit-dev-migration.md (paused, in paused/ directory)
- VPS: Running NixOS
- Old worktree: Can reference
/persistent/dcdev/clones/ngit-grasp/worktrees/820a-relay-ngit-dev-migration/for existing scripts and learnings - Target: Simple, practical migration guide that works
- Branch deployed:
4bc5-relay-ngit-dev-migration-v2currently running on VPS