mirror of
https://relay.ngit.dev/npub15qydau2hjma6ngxkl2cyar74wzyjshvl65za5k5rl69264ar2exs5cyejr/ngit-grasp.git
synced 2026-10-05 23:18:24 +00:00
docs: update docs with sync and purgatory and git data sync
This commit is contained in:
@@ -2,6 +2,21 @@
|
||||
|
||||
A [GRASP](https://gitworkshop.dev/danconwaydev.com/grasp) (Git Relays Authorized via Signed-Nostr Proofs) implementation in Rust.
|
||||
|
||||
## What's New 🎉
|
||||
|
||||
**Full GRASP-02 Implementation Complete!**
|
||||
|
||||
ngit-grasp now features a sophisticated proactive sync system that automatically discovers relays, syncs events using NIP-77 negentropy, and hunts for missing git data across clone URLs. Key highlights:
|
||||
|
||||
- ✨ **NIP-77 Negentropy**: Efficient set reconciliation with automatic REQ+EOSE fallback
|
||||
- ✨ **Intelligent Purgatory**: Auto-fetches missing git data from clone URLs (500ms for synced events, 3min for user pushes)
|
||||
- ✨ **Multi-Maintainer First-Class**: Pushed git data automatically syncs to all maintainer repositories
|
||||
- ✨ **Smart Throttling**: Respectful rate limiting (5 concurrent, 30/min per domain) with exponential backoff
|
||||
- ✨ **Live & Historic Sync**: Real-time event streaming plus daily full reconciliation
|
||||
- ✨ **Connection Health**: Exponential backoff, rate limit detection, dead relay handling
|
||||
|
||||
See [GRASP-02 Proactive Sync](docs/explanation/grasp-02-proactive-sync.md) and [Purgatory Git Data Sync](docs/explanation/grasp-02-proactive-sync-purgatory-git-data.md) for details.
|
||||
|
||||
## Overview
|
||||
|
||||
`ngit-grasp` is a Rust-based implementation of the GRASP protocol, which enables decentralized Git repository hosting with Nostr-based authorization. This implementation combines:
|
||||
@@ -14,16 +29,26 @@ Unlike the reference implementation ([ngit-relay](https://gitworkshop.dev/npub15
|
||||
|
||||
## Status
|
||||
|
||||
**Production Ready** - Full GRASP-01 and GRASP-02 implementation with comprehensive test coverage.
|
||||
|
||||
## Key Features
|
||||
|
||||
- **Pure Rust Implementation**: Single binary, no external dependencies beyond Git itself
|
||||
- **Integrated Authorization**: Push validation happens inline during the Git receive-pack operation
|
||||
- **GRASP-01 Compliant**: Core service requirements for Git hosting with Nostr authorization
|
||||
- **Extensible Architecture**: Designed to support GRASP-02 (Proactive Sync) and GRASP-05 (Archive) extensions
|
||||
- **GRASP-02 Proactive Sync**: Sophisticated relay-to-relay event and git data synchronization
|
||||
- **NIP-77 Negentropy**: Efficient set reconciliation with automatic fallback to REQ+EOSE
|
||||
- **Live & Historic Sync**: Real-time event streaming plus catch-up for past events
|
||||
- **Smart Throttling**: Respectful rate limiting (5 concurrent, 30/min per domain) with exponential backoff
|
||||
- **Multi-Maintainer First-Class**: Internal sync of pushed git data across all maintainer repositories
|
||||
- **Intelligent Purgatory**: Auto-fetches missing git data from clone URLs when events arrive first
|
||||
- **Discovery-Driven**: Dynamically connects to relays listed in repository announcements
|
||||
- **Developer-Friendly**: Built with modern Rust async patterns using tokio and actix-web
|
||||
|
||||
## Architecture Highlights
|
||||
|
||||
### Inline Authorization (GRASP-01)
|
||||
|
||||
The key architectural decision is **inline authorization** rather than Git hooks:
|
||||
|
||||
- Vendored and customised `git-http-backend` crate provides low-level access to the Git protocol
|
||||
@@ -38,9 +63,50 @@ This approach provides:
|
||||
- **Tighter integration**: Shared state between Git and Nostr components
|
||||
- **Easier testing**: Pure Rust unit and integration tests
|
||||
|
||||
### Sophisticated Sync System (GRASP-02)
|
||||
|
||||
The proactive sync implementation is production-grade with advanced features:
|
||||
|
||||
**NIP-77 Negentropy with Intelligent Fallback:**
|
||||
|
||||
- Attempts efficient set reconciliation via NIP-77 for full syncs
|
||||
- Automatically falls back to REQ+EOSE with pagination when negentropy unavailable
|
||||
- Combines live subscriptions (`limit:0`) with historic catch-up
|
||||
|
||||
**Multi-Layer Filter Strategy:**
|
||||
|
||||
- **Layer 1**: Repository announcements and maintainer lists (connection-level)
|
||||
- **Layer 2**: Events tagging repositories (a/A/q tags, batched per 100 repos)
|
||||
- **Layer 3**: Events tagging root events (e/E/q tags, batched per 100 IDs)
|
||||
|
||||
**Connection Health Management:**
|
||||
|
||||
- Exponential backoff for failed connections (5s → 1 hour)
|
||||
- Rate limit detection with 65-second cooldown
|
||||
- Dead relay handling (24h+ failures → minimal retry)
|
||||
- Quick reconnect (<15min) vs fresh start (>15min or daily)
|
||||
|
||||
**Intelligent Purgatory with Active Git Data Hunting:**
|
||||
|
||||
- Events without git data held in-memory for 30 minutes
|
||||
- **User events**: 3-minute delay (expect git push to follow)
|
||||
- **Synced events**: 500ms delay (batch burst arrivals, then hunt immediately)
|
||||
- Proactively fetches missing data from clone URLs every 2 minutes
|
||||
- Respectful throttling: 5 concurrent, 30 requests/min per domain
|
||||
- Round-robin fairness across repositories
|
||||
- Auto-release when data arrives, auto-expire after 30 minutes
|
||||
|
||||
**First-Class Multi-Maintainer Support:**
|
||||
|
||||
- Git data pushed to one maintainer's repo automatically syncs to all other maintainers
|
||||
- Shared object databases for storage efficiency (planned)
|
||||
- Seamless collaboration without manual coordination
|
||||
|
||||
See [GRASP-02 Proactive Sync](docs/explanation/grasp-02-proactive-sync.md) for full architectural details.
|
||||
|
||||
## GRASP Compliance
|
||||
|
||||
### GRASP-01 (Core Service Requirements)
|
||||
### GRASP-01 (Core Service Requirements) ✅
|
||||
|
||||
- ✅ NIP-01 compliant Nostr relay at `/`
|
||||
- ✅ Accepts NIP-34 repository announcements and state events
|
||||
@@ -50,12 +116,25 @@ This approach provides:
|
||||
- ✅ Support for `refs/nostr/<event-id>` for PRs
|
||||
- ✅ CORS support for web-based Git clients
|
||||
- ✅ NIP-11 relay information document
|
||||
- ✅ **Purgatory**: Events without git data held for 30 minutes, auto-released when data arrives
|
||||
|
||||
### GRASP-02 (Proactive Sync) - Planned
|
||||
### GRASP-02 (Proactive Sync) ✅
|
||||
|
||||
- 🔄 Proactive event sync from listed relays
|
||||
- 🔄 Proactive Git data sync from listed clone URLs
|
||||
- 🔄 PR data fetching and serving
|
||||
- ✅ **Relay Discovery**: Automatically connects to relays listed in repository announcements
|
||||
- ✅ **Event Sync**: Proactive sync from discovered relays using NIP-77 negentropy with REQ+EOSE fallback
|
||||
- Live subscriptions (`limit:0`) for real-time event streaming
|
||||
- Historic sync with automatic pagination for large result sets
|
||||
- Daily full reconciliation to detect drift
|
||||
- Connection health tracking with exponential backoff
|
||||
- ✅ **Git Data Sync**: Automatic fetching of missing git data from clone URLs
|
||||
- Smart timing: 3min delay for user events, 500ms for synced events
|
||||
- Respectful throttling: 5 concurrent requests, 30/min per domain
|
||||
- Round-robin fairness across repositories
|
||||
- Exponential backoff with fresh start on new events
|
||||
- ✅ **Multi-Maintainer Support**: Pushed git data automatically synced to all maintainer repositories
|
||||
- ✅ **Comprehensive Monitoring**: Prometheus metrics for sync health, bandwidth, and relay status
|
||||
|
||||
**See**: [GRASP-02 Proactive Sync](docs/explanation/grasp-02-proactive-sync.md) and [Purgatory Git Data Sync](docs/explanation/grasp-02-proactive-sync-purgatory-git-data.md)
|
||||
|
||||
### GRASP-05 (Archive) - Planned
|
||||
|
||||
@@ -64,53 +143,52 @@ This approach provides:
|
||||
|
||||
## Roadmap
|
||||
|
||||
### Purgatory
|
||||
### GRASP-02 Enhancements
|
||||
|
||||
State events / PR / PR Update events without git data should be accepted with msg: "won't be served until git data arrives" or "in puratory awaiting git data" and not served by the main relay.
|
||||
When the git data arrives, they get released from puratory. If git data doesn't arrive within 1 day, the events get deleted.
|
||||
**Proactive Sync Plus:**
|
||||
|
||||
This ensures the grasp serve only serves these events when it can provide the git data to support them.
|
||||
- 🔄 Scan read/write relays of repo/PR/Patch/Issue authors for related comments
|
||||
- 🔄 Stricter anti-spam mechanisms for author relay events
|
||||
- 🔄 Periodic scanning of relays in User Grasp Lists for announcements listing our relay
|
||||
|
||||
Why this is useful:
|
||||
### Data Efficiency
|
||||
|
||||
1. owner submits updated state event but loses connectivity before sending the new git data. The relay causing ngit-cli to fail to clone and other clients to show a warning that the git servers state doesn't align with nostr as relays only serve the latest state event (as its addressable).
|
||||
a. clients could be made more resilient if they know older versions of the state event served by a grasp server relate to the state they are currently storing.
|
||||
b. if clients just start using grasp servers (instead of other relays) then they will always be able to find the git data related to the latest versin of the event servered by a grasp server
|
||||
**Git Object Deduplication:**
|
||||
|
||||
2. serving PR events where the git data isn't accessable isn't useful.
|
||||
- 🔄 Shared object database across repositories
|
||||
- 🔄 Use `GIT_ALTERNATE_OBJECT_DIRECTORIES` or `.git/objects/info/alternates`
|
||||
- 🔄 Significant storage savings for multi-maintainer repositories
|
||||
|
||||
### GRASP-02 (Proactive Sync)
|
||||
### Monitoring & Observability
|
||||
|
||||
- rust-nostr client websocket connection to other grasp servers listening for our repo.
|
||||
- negentropy catchup
|
||||
- look for missing data (from state or PR / PR update) then try and fetch from other grasp servers. for efficency look for it from other repos (ie repos of other maintainers). Do this on new state event / PR / PR update evnet and on a timer for events we know we don't have the data for.
|
||||
ngit-grasp exposes comprehensive Prometheus metrics at `/metrics` for:
|
||||
|
||||
#### Proactive Sync +
|
||||
**Git Operations:**
|
||||
|
||||
. look for announcement events on other relays / grasp servers that list our service.
|
||||
. look on read/write relays of repo / PR / Patch / Issue author to get related comments. pass through stricter anti SPAM mechanism?
|
||||
- Clone/fetch/push rates and bandwidth
|
||||
- Authorization results (accepted/rejected)
|
||||
- Top N repositories by bandwidth
|
||||
|
||||
### Data effiency
|
||||
**Nostr Events:**
|
||||
|
||||
dedupe git data = shared object database or (GIT_ALTERNATE_OBJECT_DIRECTORIES or .git/objects/info/alternates)
|
||||
- WebSocket connections (active, unique IPs, flagged abusers)
|
||||
- Events received, stored, rejected by kind
|
||||
- Purgatory status (events waiting for git data)
|
||||
|
||||
### Monitoring
|
||||
**Sync Health (GRASP-02):**
|
||||
|
||||
ngit-grasp exposes Prometheus metrics at `/metrics` for connection tracking, Git operations, and Nostr events.
|
||||
- Per-relay connection status and health states
|
||||
- Event sync rates and bandwidth
|
||||
- Git data fetch attempts and success rates
|
||||
- Domain throttling metrics
|
||||
|
||||
**Configuration Options:**
|
||||
|
||||
| Option | CLI Flag | Environment Variable | Default |
|
||||
|--------|----------|---------------------|---------|
|
||||
| Metrics enabled | `--metrics-enabled` | `NGIT_METRICS_ENABLED` | `true` |
|
||||
| Connection abuse threshold | `--metrics-connection-per-ip-abuse-threshold` | `NGIT_METRICS_CONNECTION_PER_IP_ABUSE_THRESHOLD` | `10` |
|
||||
| Top N repos | `--metrics-top-n-repos` | `NGIT_METRICS_TOP_N_REPOS` | `10` |
|
||||
|
||||
**Key Metrics:**
|
||||
- WebSocket connections (active, unique IPs, flagged abusers)
|
||||
- Git operations (clone/fetch/push rates, bandwidth, authorization results)
|
||||
- Nostr events (received, stored, rejected by kind)
|
||||
- Top N repositories by bandwidth
|
||||
| Option | CLI Flag | Environment Variable | Default |
|
||||
| -------------------------- | --------------------------------------------- | ------------------------------------------------ | ------- |
|
||||
| Metrics enabled | `--metrics-enabled` | `NGIT_METRICS_ENABLED` | `true` |
|
||||
| Connection abuse threshold | `--metrics-connection-per-ip-abuse-threshold` | `NGIT_METRICS_CONNECTION_PER_IP_ABUSE_THRESHOLD` | `10` |
|
||||
| Top N repos | `--metrics-top-n-repos` | `NGIT_METRICS_TOP_N_REPOS` | `10` |
|
||||
|
||||
**Privacy:** IP addresses are never exposed in metrics - only aggregate counts.
|
||||
|
||||
@@ -154,6 +232,7 @@ This a useful feature of other git servers.
|
||||
```bash
|
||||
# install ngit
|
||||
curl -Ls https://ngit.dev/install.sh | bash
|
||||
|
||||
# Clone the repository
|
||||
git clone nostr://danconwaydev.com/relay.ngit.dev/ngit-grasp
|
||||
cd ngit-grasp
|
||||
@@ -164,6 +243,8 @@ nix develop -c cargo build --release
|
||||
# Configure
|
||||
cp .env.example .env
|
||||
# Edit .env with your settings
|
||||
# Required: NGIT_DOMAIN=your-domain.com
|
||||
# Optional: NGIT_SYNC_BOOTSTRAP_RELAY_URL=wss://relay.example.com
|
||||
|
||||
# Run
|
||||
nix develop -c cargo run --release
|
||||
@@ -172,6 +253,14 @@ nix develop -c cargo run --release
|
||||
nix develop -c cargo test --lib
|
||||
```
|
||||
|
||||
**What happens on startup:**
|
||||
|
||||
- Git HTTP server starts on configured bind address
|
||||
- Nostr relay begins accepting WebSocket connections
|
||||
- If bootstrap relay configured, sync system connects and discovers repositories
|
||||
- Purgatory system activates, ready to hunt for missing git data
|
||||
- Prometheus metrics exposed at `/metrics`
|
||||
|
||||
**Don't have Nix?** See [Getting Started Tutorial](docs/tutorials/getting-started.md) for alternative setup methods.
|
||||
|
||||
## Configuration
|
||||
@@ -200,6 +289,8 @@ NGIT_OWNER_NPUB=npub1... ngit-grasp --domain relay.example.com
|
||||
|
||||
### Configuration Options
|
||||
|
||||
#### Core Settings
|
||||
|
||||
| Option | CLI Flag | Environment Variable | Default |
|
||||
| ----------------- | --------------------- | ------------------------ | -------------------------------------------- |
|
||||
| Domain | `--domain` | `NGIT_DOMAIN` | (required) |
|
||||
@@ -211,6 +302,24 @@ NGIT_OWNER_NPUB=npub1... ngit-grasp --domain relay.example.com
|
||||
| Bind address | `--bind-address` | `NGIT_BIND_ADDRESS` | `127.0.0.1:8080` |
|
||||
| Database backend | `--database-backend` | `NGIT_DATABASE_BACKEND` | `lmdb` |
|
||||
|
||||
#### GRASP-02 Sync Settings
|
||||
|
||||
| Option | CLI Flag | Environment Variable | Default |
|
||||
| ------------------------- | --------------------------------------- | ------------------------------------------ | --------------- |
|
||||
| Bootstrap relay | `--sync-bootstrap-relay-url` | `NGIT_SYNC_BOOTSTRAP_RELAY_URL` | (optional) |
|
||||
| Base backoff | `--sync-base-backoff-secs` | `NGIT_SYNC_BASE_BACKOFF_SECS` | `5` seconds |
|
||||
| Max backoff | `--sync-max-backoff-secs` | `NGIT_SYNC_MAX_BACKOFF_SECS` | `3600` (1 hour) |
|
||||
| Disconnect check interval | `--sync-disconnect-check-interval-secs` | `NGIT_SYNC_DISCONNECT_CHECK_INTERVAL_SECS` | `60` seconds |
|
||||
| Disable negentropy | `--sync-disable-negentropy` | `NGIT_SYNC_DISABLE_NEGENTROPY` | `false` |
|
||||
| Batch window | N/A | `NGIT_SYNC_BATCH_WINDOW_MS` | `5000` ms |
|
||||
|
||||
**Sync Notes:**
|
||||
|
||||
- **Bootstrap relay**: Optional starting point for relay discovery. System automatically discovers additional relays from repository announcements.
|
||||
- **Backoff settings**: Controls exponential backoff for failed connections (`base * 2^(failures-1)`, capped at max).
|
||||
- **Negentropy**: Can be disabled for testing REQ+EOSE fallback behavior.
|
||||
- **Batch window**: Self-subscriber batches events for this duration before triggering sync filters.
|
||||
|
||||
### Database Backends
|
||||
|
||||
- `lmdb`: LMDB backend (default, persistent, general purpose)
|
||||
@@ -227,9 +336,24 @@ export NGIT_DOMAIN=gitnostr.com
|
||||
export NGIT_OWNER_NPUB=npub1...
|
||||
export NGIT_BIND_ADDRESS=0.0.0.0:8080
|
||||
export NGIT_DATABASE_BACKEND=lmdb
|
||||
|
||||
# Optional: Enable proactive sync from a bootstrap relay
|
||||
export NGIT_SYNC_BOOTSTRAP_RELAY_URL=wss://relay.damus.io
|
||||
|
||||
# Optional: Tune sync behavior
|
||||
export NGIT_SYNC_BASE_BACKOFF_SECS=5 # Start backoff at 5 seconds
|
||||
export NGIT_SYNC_MAX_BACKOFF_SECS=3600 # Cap backoff at 1 hour
|
||||
|
||||
ngit-grasp
|
||||
```
|
||||
|
||||
**Production Tips:**
|
||||
|
||||
- Set `NGIT_SYNC_BOOTSTRAP_RELAY_URL` to a well-connected relay for initial repository discovery
|
||||
- The system will automatically discover and connect to additional relays listed in repository announcements
|
||||
- Monitor sync health via Prometheus metrics at `/metrics`
|
||||
- Purgatory will automatically fetch missing git data from clone URLs
|
||||
|
||||
### Example: Development
|
||||
|
||||
```bash
|
||||
@@ -331,11 +455,11 @@ ngit-grasp/
|
||||
├── src/
|
||||
│ ├── main.rs # Entry point, server setup
|
||||
│ ├── lib.rs # Library exports
|
||||
│ ├── config.rs # Configuration
|
||||
│ ├── config.rs # Configuration (core + sync settings)
|
||||
│ ├── git/
|
||||
│ │ ├── mod.rs # Git module + repository operations
|
||||
│ │ ├── handlers.rs # Git HTTP handlers
|
||||
│ │ ├── authorization.rs # Push validation logic
|
||||
│ │ ├── authorization.rs # Push validation logic (checks DB + purgatory)
|
||||
│ │ ├── protocol.rs # Git protocol encoding
|
||||
│ │ └── subprocess.rs # Git subprocess management
|
||||
│ ├── nostr/
|
||||
@@ -348,30 +472,60 @@ ngit-grasp/
|
||||
│ │ ├── state.rs # State event validation + ref alignment
|
||||
│ │ ├── pr_event.rs # PR/PR Update validation
|
||||
│ │ └── related.rs # Forward/backward reference checking
|
||||
│ ├── sync/ # GRASP-02 Proactive Sync (relay-to-relay)
|
||||
│ │ ├── mod.rs # SyncManager, main loop, data structures
|
||||
│ │ ├── algorithms.rs # derive_relay_targets(), compute_actions()
|
||||
│ │ ├── filters.rs # 3-layer filter building (announcements, repos, events)
|
||||
│ │ ├── health.rs # RelayHealthTracker (backoff, rate limits)
|
||||
│ │ ├── relay_connection.rs # RelayConnection, event loop lifecycle
|
||||
│ │ ├── self_subscriber.rs # SelfSubscriber (batched event discovery)
|
||||
│ │ └── metrics.rs # SyncMetrics for Prometheus
|
||||
│ ├── purgatory/ # In-memory holding area for events awaiting git data
|
||||
│ │ ├── mod.rs # Purgatory core (state/PR storage, 30min expiry)
|
||||
│ │ ├── helpers.rs # State event ref matching, PR lookup
|
||||
│ │ ├── processing.rs # Unified git data processing (push + sync paths)
|
||||
│ │ └── sync/ # Proactive git data fetching
|
||||
│ │ ├── mod.rs # Public API (enqueue, main loop)
|
||||
│ │ ├── loop.rs # Sync loop (1s interval, debounced delays)
|
||||
│ │ ├── functions.rs # Core sync logic (try URLs, handle results)
|
||||
│ │ ├── queue.rs # SyncQueue (backoff, fresh start on new events)
|
||||
│ │ ├── throttle.rs # DomainThrottle (5 concurrent, 30/min, round-robin)
|
||||
│ │ └── context.rs # SyncContext trait + mock for testing
|
||||
│ ├── http/
|
||||
│ │ ├── mod.rs # HTTP module
|
||||
│ │ ├── landing.rs # Landing page handler
|
||||
│ │ └── nip11.rs # NIP-11 relay info document
|
||||
│ └── metrics/
|
||||
│ ├── mod.rs # Prometheus metrics
|
||||
│ ├── mod.rs # Prometheus metrics (Git, Nostr, Sync)
|
||||
│ ├── bandwidth.rs # Bandwidth tracking
|
||||
│ └── connection.rs # Connection tracking
|
||||
├── docs/ # Documentation (Diátaxis framework)
|
||||
├── tests/ # Integration tests
|
||||
│ ├── explanation/ # Architecture, decisions, GRASP-02 deep-dives
|
||||
│ ├── how-to/ # Deployment, configuration guides
|
||||
│ ├── tutorials/ # Getting started, first steps
|
||||
│ └── reference/ # API docs, test strategy
|
||||
├── tests/ # Integration tests (NIP-01, NIP-34, purgatory)
|
||||
├── grasp-audit/ # Compliance audit subproject
|
||||
└── README.md
|
||||
```
|
||||
|
||||
## Comparison with ngit-relay
|
||||
|
||||
| Feature | ngit-relay (Go) | ngit-grasp (Rust) |
|
||||
| ------------- | ----------------------------------------- | ----------------------------- |
|
||||
| Language | Go | Rust |
|
||||
| Components | nginx + git-http-backend + hooks + Khatru | Single integrated binary |
|
||||
| Authorization | Pre-receive Git hook | Inline during receive-pack |
|
||||
| Deployment | Docker + supervisord | Single binary |
|
||||
| Testing | Go tests + shell scripts | Rust unit + integration tests |
|
||||
| Performance | Good | Excellent (zero-copy, async) |
|
||||
| Feature | ngit-relay (Go) | ngit-grasp (Rust) |
|
||||
| ------------------- | ----------------------------------------- | ---------------------------------------------------- |
|
||||
| Language | Go | Rust |
|
||||
| Components | nginx + git-http-backend + hooks + Khatru | Single integrated binary |
|
||||
| Authorization | Pre-receive Git hook | Inline during receive-pack |
|
||||
| GRASP-01 | ✅ Complete | ✅ Complete |
|
||||
| GRASP-02 Event Sync | ✅ Limited | ✅ Advanced (NIP-77 negentropy + fallback) |
|
||||
| GRASP-02 Git Sync | ✅ Basic | ✅ Automatic purgatory hunting |
|
||||
| Multi-Maintainer | ✅ Supported | ✅ First-class (auto-sync across repos) |
|
||||
| Purgatory | ✅ 24-hour expiry | ✅ 30-minute expiry + proactive git data fetching |
|
||||
| Health Tracking | Basic | Advanced (exponential backoff, rate limit detection) |
|
||||
| Deployment | Docker + supervisord | Single binary |
|
||||
| Testing | Go tests + shell scripts | Rust unit + integration tests |
|
||||
| Performance | Good | Excellent (zero-copy, async) |
|
||||
| Monitoring | Basic logs | Comprehensive Prometheus metrics |
|
||||
|
||||
## Contributing
|
||||
|
||||
|
||||
@@ -1,481 +0,0 @@
|
||||
# Sync Test Refactor Options
|
||||
|
||||
## Summary of Requirements
|
||||
|
||||
From your feedback:
|
||||
|
||||
- **Tag variations**: Test with ONE sync mode (live OR historic), but make it clear which
|
||||
- **Discovery**: Needs both live AND historic examples (mechanism differs)
|
||||
- **Duplication**: Compare approaches before deciding
|
||||
|
||||
---
|
||||
|
||||
## Chosen Approach: Unified Helper with Event Slices
|
||||
|
||||
A single helper function that handles both sync modes based on which event slices have content:
|
||||
|
||||
````rust
|
||||
/// Result from sync test scenario
|
||||
pub struct SyncTestResult {
|
||||
pub source: TestRelay,
|
||||
pub syncing: TestRelay,
|
||||
pub keys: Keys,
|
||||
pub repo_coord: String,
|
||||
}
|
||||
|
||||
/// Run a sync test scenario with historic and/or live events.
|
||||
///
|
||||
/// # Arguments
|
||||
/// - `historic_events` - Events loaded on source BEFORE syncing relay connects
|
||||
/// - `live_events` - Events fed to source AFTER syncing relay connects
|
||||
///
|
||||
/// # Mode Detection
|
||||
/// - If only `historic_events` has content → Historic sync test
|
||||
/// - If only `live_events` has content → Live sync test
|
||||
/// - Both can have content for mixed scenarios
|
||||
///
|
||||
/// # Example - Historic Sync
|
||||
/// ```rust
|
||||
/// let (result, events) = run_sync_test(&[announcement, issue], &[]).await;
|
||||
/// // Source had events before syncing relay started
|
||||
/// // Verify events synced to result.syncing
|
||||
/// ```
|
||||
///
|
||||
/// # Example - Live Sync
|
||||
/// ```rust
|
||||
/// let (result, events) = run_sync_test(&[], &[issue, patch]).await;
|
||||
/// // Events were added after connection established
|
||||
/// // Verify events synced to result.syncing
|
||||
/// ```
|
||||
pub async fn run_sync_test(
|
||||
historic_events: &[Event],
|
||||
live_events: &[Event],
|
||||
) -> SyncTestResult {
|
||||
// 1. Pre-allocate syncing relay port for announcement tags
|
||||
let syncing_port = TestRelay::find_free_port();
|
||||
let syncing_domain = format!("127.0.0.1:{}", syncing_port);
|
||||
|
||||
// 2. Start source relay
|
||||
let source = TestRelay::start().await;
|
||||
|
||||
// 3. Create keys and announcement listing both relays
|
||||
let keys = Keys::generate();
|
||||
let announcement = create_repo_announcement(
|
||||
&keys,
|
||||
&[&source.domain(), &syncing_domain],
|
||||
"test-repo",
|
||||
);
|
||||
|
||||
// 4. Send announcement + historic events to source BEFORE syncing relay starts
|
||||
send_to_relay(&source, &announcement).await;
|
||||
for event in historic_events {
|
||||
send_to_relay(&source, event).await;
|
||||
}
|
||||
|
||||
// 5. Start syncing relay (connects to source)
|
||||
let syncing = TestRelay::start_on_port_with_options(
|
||||
syncing_port,
|
||||
Some(source.url().into()),
|
||||
false,
|
||||
).await;
|
||||
|
||||
// 6. Wait for sync connection to establish
|
||||
wait_for_sync_connection(syncing.url(), 1, Duration::from_secs(5)).await.ok();
|
||||
|
||||
// 7. Send live events AFTER connection established
|
||||
for event in live_events {
|
||||
send_to_relay(&source, event).await;
|
||||
}
|
||||
|
||||
// 8. Allow sync to complete
|
||||
tokio::time::sleep(Duration::from_secs(2)).await;
|
||||
|
||||
SyncTestResult {
|
||||
source,
|
||||
syncing,
|
||||
keys,
|
||||
repo_coord: repo_coord(&keys, "test-repo"),
|
||||
}
|
||||
}
|
||||
````
|
||||
|
||||
### Test Usage Examples
|
||||
|
||||
```rust
|
||||
// Historic sync - events existed before connection
|
||||
#[tokio::test]
|
||||
async fn test_historic_layer2_issue_syncs() {
|
||||
let keys = Keys::generate();
|
||||
let repo_coord = repo_coord(&keys, "test-repo");
|
||||
let issue = build_layer2_issue_event(&keys, &repo_coord, "Historic Issue")?;
|
||||
|
||||
let result = run_sync_test(&[issue.clone()], &[]).await;
|
||||
|
||||
assert!(
|
||||
wait_for_event_on_relay(result.syncing.url(), issue.id).await,
|
||||
"Historic issue should sync"
|
||||
);
|
||||
}
|
||||
|
||||
// Live sync - events arrive after connection
|
||||
#[tokio::test]
|
||||
async fn test_live_layer2_issue_syncs() {
|
||||
let keys = Keys::generate();
|
||||
let repo_coord = repo_coord(&keys, "test-repo");
|
||||
let issue = build_layer2_issue_event(&keys, &repo_coord, "Live Issue")?;
|
||||
|
||||
let result = run_sync_test(&[], &[issue.clone()]).await;
|
||||
|
||||
assert!(
|
||||
wait_for_event_on_relay(result.syncing.url(), issue.id).await,
|
||||
"Live issue should sync"
|
||||
);
|
||||
}
|
||||
|
||||
// Discovery test - historic
|
||||
#[tokio::test]
|
||||
async fn test_discovery_historic_syncs_layer2() {
|
||||
let keys = Keys::generate();
|
||||
let repo_coord = repo_coord(&keys, "test-repo");
|
||||
let issue = build_layer2_issue_event(&keys, &repo_coord, "Discovered Issue")?;
|
||||
|
||||
// Source has the issue before discovery
|
||||
let result = run_sync_test(&[issue.clone()], &[]).await;
|
||||
|
||||
assert!(wait_for_event_on_relay(result.syncing.url(), issue.id).await);
|
||||
}
|
||||
|
||||
// Discovery test - live
|
||||
#[tokio::test]
|
||||
async fn test_discovery_live_syncs_layer2() {
|
||||
// ... setup similar, events in second slice
|
||||
}
|
||||
```
|
||||
|
||||
### Why This Approach
|
||||
|
||||
1. **Single function, both modes** - No duplication of setup logic
|
||||
2. **Clear distinction** - Function signature makes it obvious which is historic vs live
|
||||
3. **No new dependencies** - Plain Rust
|
||||
4. **Readable tests** - Test body just creates events and calls helper
|
||||
5. **Flexible** - Can test mixed scenarios if needed
|
||||
|
||||
---
|
||||
|
||||
## Proposed Test File Structure
|
||||
|
||||
```
|
||||
tests/sync/
|
||||
├── mod.rs # Module declarations + overview doc
|
||||
├── historic_sync.rs # NEW: Historic sync tests (events exist before connection)
|
||||
├── live_sync.rs # REFACTORED: Live sync tests (events arrive after connection)
|
||||
├── discovery.rs # REFACTORED: Relay discovery (both modes)
|
||||
├── tag_variations.rs # REFACTORED: Tag type coverage (live sync only)
|
||||
├── metrics.rs # UNCHANGED: Prometheus metrics tests
|
||||
└── catchup.rs # UNCHANGED: Documentation only
|
||||
```
|
||||
|
||||
### `tests/sync/historic_sync.rs` (NEW - renamed from bootstrap.rs)
|
||||
|
||||
```rust
|
||||
//! Historic Sync Tests
|
||||
//!
|
||||
//! Tests for syncing events that exist on source relay BEFORE the syncing relay connects.
|
||||
//! This is the bootstrap/startup sync path.
|
||||
|
||||
/// Historic sync: Layer 1 announcements sync on startup
|
||||
/// run_sync_test with historic_events: [announcement]
|
||||
#[tokio::test]
|
||||
async fn test_historic_layer1_announcement_syncs() { ... }
|
||||
|
||||
/// Historic sync: Layer 2 issues sync on startup
|
||||
/// run_sync_test with historic_events: [issue]
|
||||
#[tokio::test]
|
||||
async fn test_historic_layer2_issue_syncs() { ... }
|
||||
|
||||
/// Historic sync: Layer 3 comments sync after Layer 2 syncs
|
||||
/// run_sync_test with historic_events: [issue, comment]
|
||||
#[tokio::test]
|
||||
async fn test_historic_layer3_comment_syncs() { ... }
|
||||
|
||||
/// Historic sync: Events not listing relay domain are rejected
|
||||
/// run_sync_test with historic_events: [announcement_missing_domain]
|
||||
/// Verify NOT synced
|
||||
#[tokio::test]
|
||||
async fn test_historic_rejects_events_not_listing_relay() { ... }
|
||||
|
||||
/// Historic sync works without NIP-77 negentropy (REQ+EOSE fallback)
|
||||
/// run_sync_test with historic_events, negentropy disabled
|
||||
#[tokio::test]
|
||||
async fn test_historic_sync_without_negentropy() { ... }
|
||||
```
|
||||
|
||||
### `tests/sync/live_sync.rs` (REFACTORED)
|
||||
|
||||
```rust
|
||||
//! Live Sync Tests
|
||||
//!
|
||||
//! Tests for syncing events that arrive on source relay AFTER the syncing relay connects.
|
||||
//! This is the real-time/subscription-based sync path.
|
||||
|
||||
/// Live sync: Layer 2 issues sync in real-time
|
||||
/// run_sync_test with live_events: [issue]
|
||||
#[tokio::test]
|
||||
async fn test_live_layer2_issue_syncs() { ... }
|
||||
|
||||
/// Live sync: Layer 3 comments sync after Layer 2 syncs
|
||||
/// run_sync_test with live_events: [issue, comment]
|
||||
#[tokio::test]
|
||||
async fn test_live_layer3_comment_syncs() { ... }
|
||||
|
||||
/// Live sync: Events arrive in chronological order
|
||||
/// run_sync_test with live_events: [issue1, issue2, issue3]
|
||||
#[tokio::test]
|
||||
async fn test_live_sync_event_ordering() { ... }
|
||||
```
|
||||
|
||||
### `tests/sync/discovery.rs` (REFACTORED)
|
||||
|
||||
```rust
|
||||
//! Relay Discovery Tests
|
||||
//!
|
||||
//! Tests for discovering other relays from announcement events.
|
||||
//! Discovery can happen via historic sync or live sync paths.
|
||||
|
||||
// === HISTORIC DISCOVERY ===
|
||||
// Relay discovers another relay from an announcement that existed before connection
|
||||
|
||||
/// Historic discovery: Discovers relay from announcement, syncs Layer 2
|
||||
/// 1. relay_a has announcement + issue
|
||||
/// 2. relay_b starts with sync from relay_a
|
||||
/// 3. relay_b syncs announcement, discovers other relays listed, syncs issue
|
||||
#[tokio::test]
|
||||
async fn test_historic_discovery_syncs_layer2() { ... }
|
||||
|
||||
/// Historic discovery: 3-relay recursive discovery chain
|
||||
/// 1. relay_b has announcement listing relay_c
|
||||
/// 2. relay_c has separate announcement
|
||||
/// 3. relay_a starts syncing from relay_b
|
||||
/// 4. relay_a gets announcement, discovers relay_c, syncs from relay_c
|
||||
#[tokio::test]
|
||||
async fn test_historic_recursive_discovery() { ... }
|
||||
|
||||
// === LIVE DISCOVERY ===
|
||||
// Relay discovers another relay from an announcement that arrives after connection
|
||||
|
||||
/// Live discovery: Discovers relay from new announcement, syncs Layer 2
|
||||
/// 1. Both relays running
|
||||
/// 2. Announcement submitted to relay_a
|
||||
/// 3. relay_b discovers relay_a from announcement, syncs Layer 2 events
|
||||
#[tokio::test]
|
||||
async fn test_live_discovery_syncs_layer2() { ... }
|
||||
|
||||
/// Live discovery: Multi-hop discovery with layer chain
|
||||
/// Similar to recursive but with live submission of announcement
|
||||
#[tokio::test]
|
||||
async fn test_live_recursive_discovery() { ... }
|
||||
```
|
||||
|
||||
### `tests/sync/tag_variations.rs` (REFACTORED - Live sync only)
|
||||
|
||||
```rust
|
||||
//! Tag Variation Tests (Live Sync Mode)
|
||||
//!
|
||||
//! Tests that all valid tag types are correctly processed during sync.
|
||||
//! Uses LIVE sync mode - tag parsing is mode-independent, so testing one mode is sufficient.
|
||||
|
||||
// === LAYER 2 TAG VARIATIONS ===
|
||||
|
||||
/// Layer 2 with lowercase 'a' tag (standard NIP-01)
|
||||
#[tokio::test]
|
||||
async fn test_layer2_lowercase_a_tag() { ... }
|
||||
|
||||
/// Layer 2 with uppercase 'A' tag (NIP-33 style)
|
||||
#[tokio::test]
|
||||
async fn test_layer2_uppercase_a_tag() { ... }
|
||||
|
||||
/// Layer 2 with 'q' quote tag (NIP-18 style)
|
||||
#[tokio::test]
|
||||
async fn test_layer2_q_tag() { ... }
|
||||
|
||||
// === LAYER 3 TAG VARIATIONS ===
|
||||
|
||||
/// Layer 3 with lowercase 'e' tag (standard NIP-01)
|
||||
#[tokio::test]
|
||||
async fn test_layer3_lowercase_e_tag() { ... }
|
||||
|
||||
/// Layer 3 with uppercase 'E' tag (NIP-22 comment)
|
||||
#[tokio::test]
|
||||
async fn test_layer3_uppercase_e_tag() { ... }
|
||||
|
||||
/// Layer 3 with 'q' quote tag (NIP-18 style)
|
||||
#[tokio::test]
|
||||
async fn test_layer3_q_tag() { ... }
|
||||
```
|
||||
|
||||
### `tests/sync/mod.rs` (UPDATED)
|
||||
|
||||
```rust
|
||||
//! Sync Integration Tests
|
||||
//!
|
||||
//! Tests for ngit-grasp's proactive sync functionality, organized by sync mode:
|
||||
//!
|
||||
//! ## Sync Modes
|
||||
//!
|
||||
//! - **Historic Sync** (`historic_sync.rs`): Events exist BEFORE syncing relay connects
|
||||
//! - Also called bootstrap/startup sync
|
||||
//! - Tests the REQ+EOSE or negentropy-based initial sync
|
||||
//!
|
||||
//! - **Live Sync** (`live_sync.rs`): Events arrive AFTER syncing relay connects
|
||||
//! - Also called real-time/subscription-based sync
|
||||
//! - Tests the event forwarding via active subscriptions
|
||||
//!
|
||||
//! ## Other Test Categories
|
||||
//!
|
||||
//! - **Discovery** (`discovery.rs`): Relay discovers other relays from announcements
|
||||
//! - Has both historic and live variants
|
||||
//!
|
||||
//! - **Tag Variations** (`tag_variations.rs`): All valid tag types work correctly
|
||||
//! - Uses live sync (tag parsing is mode-independent)
|
||||
//!
|
||||
//! - **Metrics** (`metrics.rs`): Prometheus metrics for sync operations
|
||||
//!
|
||||
//! - **Catchup** (`catchup.rs`): Documentation only (not integration-testable)
|
||||
|
||||
pub mod catchup;
|
||||
pub mod discovery;
|
||||
pub mod historic_sync; // Renamed from bootstrap
|
||||
pub mod live_sync;
|
||||
pub mod metrics;
|
||||
pub mod tag_variations;
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Test Count Summary
|
||||
|
||||
| File | Before | After | Notes |
|
||||
| --------------------------------- | ------ | ------ | ------------------------ |
|
||||
| `historic_sync.rs` (bootstrap.rs) | 4 | 5 | Renamed, minor additions |
|
||||
| `live_sync.rs` | 3 | 3 | Simplified using helper |
|
||||
| `discovery.rs` | 3 | 4 | Split into historic/live |
|
||||
| `tag_variations.rs` | 6 | 6 | Simplified using helper |
|
||||
| `metrics.rs` | 9 | 9 | Unchanged |
|
||||
| `catchup.rs` | 0 | 0 | Documentation only |
|
||||
| **Total** | **25** | **27** | +2 discovery tests |
|
||||
|
||||
---
|
||||
|
||||
## Implementation Plan
|
||||
|
||||
### Context
|
||||
|
||||
- Two metrics tests are currently failing: `test_live_sync_event_count` and `test_multi_source_aggregate_counts`
|
||||
- This may indicate implementation bugs in sync functionality
|
||||
- Plan must account for distinguishing test bugs from implementation bugs
|
||||
|
||||
### Phase 1: Establish Baseline
|
||||
|
||||
**Goal:** Understand current state before making changes
|
||||
|
||||
1. Run `cargo test --test sync` and capture full output
|
||||
2. Identify all passing vs failing tests
|
||||
3. For each failing test, investigate whether:
|
||||
- Test logic is incorrect (test bug)
|
||||
- Implementation is broken (impl bug)
|
||||
- Test was never working (aspirational test)
|
||||
4. Document findings in Known Issues section below
|
||||
|
||||
### Phase 2: Add Test Infrastructure
|
||||
|
||||
**Goal:** Add new helpers without breaking existing tests
|
||||
|
||||
1. Add `SyncTestResult` struct to sync_helpers.rs:
|
||||
|
||||
```rust
|
||||
pub struct SyncTestResult {
|
||||
pub source: TestRelay,
|
||||
pub syncing: TestRelay,
|
||||
pub keys: Keys,
|
||||
pub repo_coord: String,
|
||||
}
|
||||
```
|
||||
|
||||
2. Add `run_sync_test(historic_events, live_events)` helper function
|
||||
|
||||
3. Add unit tests for the helper itself:
|
||||
|
||||
- `test_run_sync_test_historic_mode` - verify events sent before connection
|
||||
- `test_run_sync_test_live_mode` - verify events sent after connection
|
||||
|
||||
4. Run `cargo test` to confirm no regressions
|
||||
|
||||
### Phase 3: Refactor Historic Sync Tests
|
||||
|
||||
**Goal:** Migrate bootstrap.rs → historic_sync.rs incrementally
|
||||
|
||||
1. Rename file: `bootstrap.rs` → `historic_sync.rs`
|
||||
2. Update `mod.rs` to reference new module name
|
||||
3. Refactor tests one at a time:
|
||||
- `test_bootstrap_syncs_existing_layer2_events` → `test_historic_layer2_issue_syncs`
|
||||
- `test_relay_replays_events_after_restart` → keep or remove (tests restart, not sync mode)
|
||||
- `test_announcement_not_listing_relay_is_not_synced` → `test_historic_rejects_unlisted_relay`
|
||||
- `test_history_sync_without_negentropy` → `test_historic_sync_without_negentropy`
|
||||
4. Test after each refactor: `cargo test --test sync historic_sync`
|
||||
|
||||
### Phase 4: Refactor Live Sync Tests
|
||||
|
||||
**Goal:** Simplify live_sync.rs using run_sync_test helper
|
||||
|
||||
1. Refactor tests one at a time:
|
||||
- `test_live_sync_layer2_events` → use `run_sync_test(&[], &[issue])`
|
||||
- `test_live_sync_layer3_events` → use `run_sync_test(&[], &[issue, comment])`
|
||||
- `test_live_sync_event_ordering` → may need custom setup for ordering test
|
||||
2. Test after each refactor: `cargo test --test sync live_sync`
|
||||
|
||||
### Phase 5: Refactor Discovery Tests
|
||||
|
||||
**Goal:** Split discovery.rs into historic and live sections
|
||||
|
||||
1. Add section comment: `// === HISTORIC DISCOVERY ===`
|
||||
2. Refactor existing tests to use run_sync_test
|
||||
3. Add new tests:
|
||||
- `test_historic_discovery_syncs_layer2`
|
||||
- `test_live_discovery_syncs_layer2`
|
||||
4. Test after each change: `cargo test --test sync discovery`
|
||||
|
||||
### Phase 6: Refactor Tag Variations Tests
|
||||
|
||||
**Goal:** Simplify tag_variations.rs using run_sync_test (live mode)
|
||||
|
||||
1. Add header doc comment explaining live sync mode choice
|
||||
2. Refactor each test to use `run_sync_test(&[], &[event])`
|
||||
3. Test after all changes: `cargo test --test sync tag_variations`
|
||||
|
||||
### Phase 7: Final Verification and Cleanup
|
||||
|
||||
1. Update `mod.rs` documentation
|
||||
2. Run full test suite: `cargo test --test sync`
|
||||
3. Compare results to Phase 1 baseline:
|
||||
- New failures = regressions from refactor (must fix)
|
||||
- Same failures as baseline = pre-existing issues (document)
|
||||
4. Delete `work/sync-test-refactor-options.md`
|
||||
|
||||
---
|
||||
|
||||
## Known Issues
|
||||
|
||||
_To be filled in during Phase 1_
|
||||
|
||||
### Failing Tests Before Refactor
|
||||
|
||||
| Test | Status | Root Cause | Action |
|
||||
| --------------------------------------------- | ---------- | ---------- | ------ |
|
||||
| `metrics::test_live_sync_event_count` | ❌ Failing | TBD | TBD |
|
||||
| `metrics::test_multi_source_aggregate_counts` | ❌ Failing | TBD | TBD |
|
||||
|
||||
### Implementation Notes
|
||||
|
||||
_Any discoveries about how sync actually works vs how tests expect it to work_
|
||||
|
||||
---
|
||||
@@ -80,6 +80,77 @@ Explanation documentation helps you **understand concepts** and design decisions
|
||||
|
||||
---
|
||||
|
||||
### [Purgatory Design](purgatory-design.md)
|
||||
**In-memory holding area for events awaiting git data**
|
||||
|
||||
**Topics:**
|
||||
- The "which arrives first?" problem
|
||||
- Separate storage for state vs PR events
|
||||
- Late binding for state events
|
||||
- Bidirectional waiting for PR events
|
||||
- Authorization during push
|
||||
|
||||
**Read when:** You want to understand how ngit-grasp handles out-of-order event/git data arrival
|
||||
|
||||
---
|
||||
|
||||
### [GRASP-02 Proactive Sync](grasp-02-proactive-sync.md)
|
||||
**Relay-to-relay synchronization for repository discovery**
|
||||
|
||||
**Topics:**
|
||||
- Negentropy-based event sync
|
||||
- Repository announcement discovery
|
||||
- Relay management and reconnection
|
||||
- Layer 2 filtering
|
||||
- Bootstrap and dynamic relay discovery
|
||||
|
||||
**Read when:** You want to understand how ngit-grasp discovers and syncs repositories across relays
|
||||
|
||||
---
|
||||
|
||||
### [GRASP-02 Purgatory Git Data Fetching](grasp-02-proactive-sync-purgatory-git-data.md)
|
||||
**Proactive git data fetching from remote servers**
|
||||
|
||||
**Topics:**
|
||||
- Identifier-based batching
|
||||
- Exponential backoff with fresh start
|
||||
- Domain throttling (5 concurrent, 30/min)
|
||||
- Debounced delays (3min user, 500ms sync)
|
||||
- 30-minute expiry
|
||||
- Mock-based testability
|
||||
|
||||
**Read when:** You want to understand how purgatory automatically fetches missing git data
|
||||
|
||||
---
|
||||
|
||||
### [Unified Git Data Sync](unify-git-data-sync.md)
|
||||
**Shared processing for git push and purgatory sync paths**
|
||||
|
||||
**Topics:**
|
||||
- Why unify push and sync processing
|
||||
- OID syncing to owner repos
|
||||
- Ref alignment logic
|
||||
- Event release from purgatory
|
||||
- WebSocket notification
|
||||
|
||||
**Read when:** You want to understand how git data is processed consistently regardless of arrival method
|
||||
|
||||
---
|
||||
|
||||
### [Monitoring Overview](monitoring.md)
|
||||
**Prometheus metrics and observability**
|
||||
|
||||
**Topics:**
|
||||
- Metrics philosophy
|
||||
- Connection tracking
|
||||
- Git operation metrics
|
||||
- Nostr event metrics
|
||||
- Privacy considerations
|
||||
|
||||
**Read when:** You want to understand how to monitor ngit-grasp in production
|
||||
|
||||
---
|
||||
|
||||
## Planned Explanation Documentation
|
||||
|
||||
### GRASP Protocol Design
|
||||
|
||||
@@ -0,0 +1,675 @@
|
||||
# GRASP-02 Proactive Sync: Purgatory Git Data Fetching
|
||||
|
||||
**Status**: ✅ Implemented
|
||||
**Implementation**: [`src/purgatory/sync/`](../../src/purgatory/sync/)
|
||||
**Related**:
|
||||
|
||||
- [Purgatory Design](purgatory-design.md) - Core purgatory concepts
|
||||
- [GRASP-02 Proactive Sync](grasp-02-proactive-sync.md) - Full GRASP-02 implementation
|
||||
- [Unified Git Data Sync](unify-git-data-sync.md) - Shared processing logic
|
||||
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
When Nostr events arrive before their git data, they enter **purgatory** waiting to be served. But they don't wait passively—ngit-grasp **actively hunts** for the missing git data across all git servers assoicated with the repo until it finds what it needs.
|
||||
|
||||
### How It Works
|
||||
|
||||
**If the data exists, we'll find it.**
|
||||
|
||||
The system scours git servers listed in repository announcements and PR events, checking every **2 minutes** for **30 minutes**. If we find the data, events are released immediately. If not, they expire from purgatory after 30 minutes.
|
||||
|
||||
**Smart timing based on how events arrive:**
|
||||
|
||||
- **User-submitted events**: Wait **3 minutes** before hunting—we expect a `git push` to follow shortly
|
||||
- **Sync-received events**: Start hunting after just **500ms**—batch burst arrivals, then get to work
|
||||
|
||||
**Playing nicely with other servers:**
|
||||
|
||||
We respect remote server capacity with:
|
||||
|
||||
- **Throttling**: Max 5 concurrent requests per domain, 30 requests/minute
|
||||
- **Backoff**: Start at 20 seconds, double each attempt, cap at 2 minutes
|
||||
- **Round-robin**: Fair distribution across repositories waiting for the same domain
|
||||
- **Fresh start**: New events reset retry count—recent updates often mean fresh data
|
||||
|
||||
**The result**: If git data is available anywhere in the clone URL list, we'll find it within minutes. If it's not available within 30 minutes, the events expire cleanly.
|
||||
|
||||
### Key Features
|
||||
|
||||
✅ **Proactive hunting** - Scours git servers every 2 min (backoff), finds data automatically
|
||||
✅ **Respectful throttling** - 5 concurrent + 30/min per domain, plays nice with other implementations
|
||||
✅ **Smart timing** - 3min delay for user pushes, 500ms for synced events
|
||||
✅ **30min expiry** - Auto-cleanup of events when data never arrives
|
||||
✅ **Fully testable** - Mock-based architecture for reliable unit tests
|
||||
|
||||
---
|
||||
|
||||
## The Problem: Out-of-Order Arrival
|
||||
|
||||
In a distributed system, git data and Nostr events can arrive in any order:
|
||||
|
||||
```
|
||||
Timeline A: Event arrives first (user push expected)
|
||||
t=0s: State event received → enters purgatory
|
||||
t=180s: (3min wait - expecting git push)
|
||||
t=30s: Git push arrives → event released ✅
|
||||
|
||||
Timeline B: Git arrives first
|
||||
t=0s: Git push received → data available
|
||||
t=30s: State event received → immediately served ✅
|
||||
|
||||
Timeline C: Sync scenario (hunt for data)
|
||||
t=0s: State event received from relay X → enters purgatory
|
||||
t=0.5s: (500ms delay to batch bursts)
|
||||
t=0.5s: Start hunting git servers → check server1, server2, server3...
|
||||
t=45s: Git data found on server2 → event released ✅
|
||||
|
||||
Timeline D: Data never arrives
|
||||
t=0s: State event received → enters purgatory
|
||||
t=0.5s: Start hunting → server1 (not found), server2 (timeout), server3 (not found)
|
||||
t=20s: Retry → server1 (not found), server2 (not found), server3 (not found)
|
||||
t=60s: Retry → all servers checked, no data
|
||||
...
|
||||
t=1800s: 30 minutes expired → event discarded, purgatory cleaned up 🗑️
|
||||
```
|
||||
|
||||
**Without proactive sync**: Events in Timeline C would wait indefinitely (or until manual git push).
|
||||
**With proactive sync**: System automatically hunts for data across all known servers, releasing events as soon as the data is found.
|
||||
|
||||
---
|
||||
|
||||
## Architecture: Two-Path Sync Design
|
||||
|
||||
The system uses **two independent execution paths** that work together:
|
||||
|
||||
### Path 1: Main Sync Loop (Non-Throttled URLs)
|
||||
|
||||
Runs every **1 second**, processes identifiers ready for sync:
|
||||
|
||||
1. Find ready identifiers (where `!in_progress && next_attempt <= now`)
|
||||
2. Spawn parallel tasks for each identifier
|
||||
3. Each task tries non-throttled URLs until:
|
||||
- ✅ All OIDs fetched (complete) → remove from queue
|
||||
- ⏸️ Only throttled URLs remain → enqueue with throttled domains, apply backoff
|
||||
- ❌ No URLs left (all tried/throttled) → apply backoff, retry later
|
||||
|
||||
**Key insight**: Main loop doesn't wait for throttled domains. It quickly tries available servers, then hands off to domain queues for rate-limited processing.
|
||||
|
||||
### Path 2: Domain Throttle Queues (Throttled URLs)
|
||||
|
||||
**Trigger-based** (no polling), processes when capacity frees:
|
||||
|
||||
1. Identifier enqueued with throttled domain (from main loop)
|
||||
2. When domain has capacity (slot frees or rate limit window passes):
|
||||
- Pick next identifier (round-robin for fairness)
|
||||
- Try one URL from that domain
|
||||
- Mark URL as tried, release slot
|
||||
3. Trigger repeats until queue empty or capacity exhausted
|
||||
|
||||
**Key insight**: Each domain independently manages its queue, ensuring we respect rate limits while maximizing throughput.
|
||||
|
||||
---
|
||||
|
||||
## Data Flow: From Event to Release
|
||||
|
||||
```mermaid
|
||||
graph TB
|
||||
A[Event Arrives] --> B{Git Data<br/>Available?}
|
||||
B -->|Yes| C[Serve Immediately]
|
||||
B -->|No| D[Enter Purgatory]
|
||||
|
||||
D --> E[Enqueue for Sync]
|
||||
E --> F{Event Source?}
|
||||
F -->|User Submit| G[3min Delay<br/>expect push]
|
||||
F -->|Relay Sync| H[500ms Delay<br/>batch burst]
|
||||
|
||||
G --> I[Main Sync Loop<br/>1s interval]
|
||||
H --> I
|
||||
|
||||
I --> J{Ready?}
|
||||
J -->|Not Yet| I
|
||||
J -->|Yes| K[Spawn Sync Task]
|
||||
|
||||
K --> L[Try Non-Throttled URLs]
|
||||
L --> M{Got All OIDs?}
|
||||
M -->|Yes| N[Process & Release]
|
||||
M -->|Partial| O[Enqueue Throttled Domains]
|
||||
M -->|None| P[Apply Backoff]
|
||||
|
||||
O --> Q[Domain Queue]
|
||||
Q --> R{Has Capacity?}
|
||||
R -->|No| Q
|
||||
R -->|Yes| S[Try Domain URL]
|
||||
S --> T{Got OIDs?}
|
||||
T -->|Yes| N
|
||||
T -->|No| U[Try Next in Queue]
|
||||
|
||||
P --> I
|
||||
N --> V[Event Served]
|
||||
|
||||
style D fill:#fff3cd
|
||||
style N fill:#d4edda
|
||||
style V fill:#d1ecf1
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Retry Strategy: Exponential Backoff with Fresh Start
|
||||
|
||||
### Backoff Schedule
|
||||
|
||||
When sync attempts don't complete (OIDs still needed), backoff increases:
|
||||
|
||||
| Attempt | Delay | Formula |
|
||||
| ------- | ------------- | ---------------------- |
|
||||
| 1 | 20s | `20s * 2^0` |
|
||||
| 2 | 40s | `20s * 2^1` |
|
||||
| 3 | 80s | `20s * 2^2` |
|
||||
| 4+ | 120s (capped) | `min(20s * 2^n, 120s)` |
|
||||
|
||||
**Implementation**: [`src/purgatory/sync/queue.rs:SyncQueueEntry::backoff()`](../../src/purgatory/sync/queue.rs)
|
||||
|
||||
### Fresh Start on New Events
|
||||
|
||||
**Critical feature**: When a new event arrives for an identifier already in the sync queue, the `attempt_count` resets to 0.
|
||||
|
||||
**Why?** New events often mean:
|
||||
|
||||
- A maintainer just updated the repository
|
||||
- Fresh git data might be available at new clone URLs
|
||||
- Previous failures might have been temporary
|
||||
|
||||
**Example**:
|
||||
|
||||
```
|
||||
t=0s: State A arrives → queue with 3min delay, attempt_count=0
|
||||
t=180s: First sync attempt fails → backoff 20s, attempt_count=1
|
||||
t=200s: Second attempt fails → backoff 40s, attempt_count=2
|
||||
t=210s: State B arrives (same identifier) → attempt_count=0 ✨
|
||||
t=210s: Immediate retry (new event delay) → success!
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Debounced Delays: Smart Timing
|
||||
|
||||
### User-Submitted Events: 3 Minutes
|
||||
|
||||
When a user submits an event via `EVENT` message, we expect a `git push` to follow shortly:
|
||||
|
||||
```
|
||||
t=0s: User submits state event → purgatory + 3min delay
|
||||
t=30s: User runs `git push` → data arrives → event released ✅
|
||||
```
|
||||
|
||||
**Why 3 minutes?** Gives users time to:
|
||||
|
||||
- Finish composing their commit message
|
||||
- Run `git push` command
|
||||
- Handle network delays
|
||||
|
||||
**Configuration**: Hardcoded in [`src/purgatory/mod.rs:DEFAULT_SYNC_DELAY`](../../src/purgatory/mod.rs)
|
||||
|
||||
### Sync-Triggered Events: 500ms
|
||||
|
||||
When events arrive during relay sync (e.g., negentropy catchup), they often come in bursts:
|
||||
|
||||
```
|
||||
t=0s: State A arrives → purgatory + 500ms delay
|
||||
t=0.1s: State B arrives → purgatory + 500ms delay (same repo)
|
||||
t=0.2s: State C arrives → purgatory + 500ms delay (same repo)
|
||||
t=0.5s: Single sync attempt fetches data for all three ✅
|
||||
```
|
||||
|
||||
**Why 500ms?** Batches burst arrivals without excessive delay.
|
||||
|
||||
**Configuration**: Hardcoded in [`src/purgatory/mod.rs:IMMEDIATE_SYNC_DELAY`](../../src/purgatory/mod.rs)
|
||||
|
||||
### Debouncing Mechanism
|
||||
|
||||
Multiple events for the same identifier **don't create multiple sync tasks**. The `enqueue_sync` method:
|
||||
|
||||
1. If identifier not in queue → create new entry with delay
|
||||
2. If identifier already queued → reset `attempt_count`, update `next_attempt` if sooner
|
||||
|
||||
**Result**: Rapid event arrivals → single sync attempt after debounce window.
|
||||
|
||||
**Implementation**: [`src/purgatory/mod.rs:Purgatory::enqueue_sync()`](../../src/purgatory/mod.rs)
|
||||
|
||||
---
|
||||
|
||||
## Domain Throttling: Respectful Rate Limiting
|
||||
|
||||
### Why Throttle?
|
||||
|
||||
Git servers have finite resources. Without throttling:
|
||||
|
||||
- ❌ We could overwhelm small servers with concurrent requests
|
||||
- ❌ Servers might rate-limit or ban us
|
||||
- ❌ Other clients sharing the server suffer degraded performance
|
||||
|
||||
With throttling:
|
||||
|
||||
- ✅ Respect server capacity (5 concurrent max per domain)
|
||||
- ✅ Stay under rate limits (30 requests/min per domain)
|
||||
- ✅ Fair access for all clients
|
||||
|
||||
### Two-Level Limits
|
||||
|
||||
Each domain has **two independent limits**:
|
||||
|
||||
#### 1. Concurrent Request Limit (Default: 5)
|
||||
|
||||
Maximum in-flight requests to a domain at any moment.
|
||||
|
||||
**Example**:
|
||||
|
||||
```
|
||||
Domain: github.com
|
||||
In-flight: [fetch-1, fetch-2, fetch-3, fetch-4, fetch-5]
|
||||
Status: AT CAPACITY (throttled)
|
||||
|
||||
fetch-3 completes → in-flight: 4
|
||||
Status: HAS CAPACITY (process next queued identifier)
|
||||
```
|
||||
|
||||
#### 2. Rate Limit (Default: 30/min)
|
||||
|
||||
Maximum requests in any 60-second sliding window.
|
||||
|
||||
**Example**:
|
||||
|
||||
```
|
||||
t=0s: Request 1 → request_times: [0s]
|
||||
t=1s: Request 2 → request_times: [0s, 1s]
|
||||
...
|
||||
t=30s: Request 30 → request_times: [0s, 1s, ..., 30s]
|
||||
t=31s: Request 31? → THROTTLED (30 requests in last 60s)
|
||||
t=61s: Request at t=0s aged out → request_times: [1s, ..., 30s]
|
||||
t=61s: Request 31 → ALLOWED (only 29 in last 60s)
|
||||
```
|
||||
|
||||
**Implementation**: [`src/purgatory/sync/throttle.rs:DomainThrottle::has_capacity()`](../../src/purgatory/sync/throttle.rs)
|
||||
|
||||
### Round-Robin Fairness
|
||||
|
||||
When multiple identifiers are queued for a throttled domain, we use **round-robin** to ensure fairness:
|
||||
|
||||
```
|
||||
Queue: [repo-A, repo-B, repo-C]
|
||||
Round-robin index: 0
|
||||
|
||||
Attempt 1: Try repo-A (index=0) → fetch → index=1
|
||||
Attempt 2: Try repo-B (index=1) → fetch → index=2
|
||||
Attempt 3: Try repo-C (index=2) → fetch → index=0
|
||||
Attempt 4: Try repo-A (index=0) → ...
|
||||
```
|
||||
|
||||
**Why round-robin?** Prevents head-of-line blocking. Without it, repo-A might consume all slots while repo-B and repo-C wait indefinitely.
|
||||
|
||||
**Implementation**: [`src/purgatory/sync/throttle.rs:DomainThrottle::next_ready_identifier()`](../../src/purgatory/sync/throttle.rs)
|
||||
|
||||
### Trigger-Based Processing (Not Polling)
|
||||
|
||||
Domain queues **don't poll** for capacity. Instead, processing is triggered by two events:
|
||||
|
||||
1. **`complete_request()`** - A request finishes, slot frees
|
||||
2. **`enqueue_identifier()`** - New identifier added to queue
|
||||
|
||||
Both methods check `has_capacity()` and trigger `try_process_next()` if true.
|
||||
|
||||
**Why trigger-based?**
|
||||
|
||||
- ✅ Lower CPU usage (no busy-waiting)
|
||||
- ✅ Instant response when capacity frees
|
||||
- ✅ Simpler reasoning (event-driven)
|
||||
|
||||
**Implementation**: [`src/purgatory/sync/throttle.rs:ThrottleManager`](../../src/purgatory/sync/throttle.rs)
|
||||
|
||||
---
|
||||
|
||||
## 30-Minute Purgatory Expiry
|
||||
|
||||
Purgatory entries **automatically expire** after 30 minutes to prevent unbounded memory growth.
|
||||
|
||||
### Why 30 Minutes?
|
||||
|
||||
From the [GRASP-01 spec](https://github.com/DanConwayDev/grasp/blob/main/01.md#purgatory):
|
||||
|
||||
> Events should be kept in purgatory and otherwise discarded after 30 minutes.
|
||||
|
||||
This balances:
|
||||
|
||||
- ⏰ **Long enough** for typical sync scenarios (git data usually arrives within minutes)
|
||||
- 🧹 **Short enough** to prevent memory leaks from abandoned events
|
||||
- 🔄 **Recoverable** events are still on other relays and can be re-submitted
|
||||
|
||||
### Implementation
|
||||
|
||||
Each purgatory entry tracks:
|
||||
|
||||
- `created_at: Instant` - When added to purgatory
|
||||
- `expires_at: Instant` - When to discard (created_at + 30min)
|
||||
|
||||
The main sync loop checks expiry before processing:
|
||||
|
||||
```rust
|
||||
if !self.has_pending_events(&identifier) {
|
||||
// No events remain (expired or released) → remove from sync queue
|
||||
self.sync_queue.remove(&identifier);
|
||||
}
|
||||
```
|
||||
|
||||
**Note**: Expiry is checked implicitly via `has_pending_events()`. If all events for an identifier have expired, the identifier is removed from the sync queue.
|
||||
|
||||
**Implementation**: [`src/purgatory/mod.rs:DEFAULT_EXPIRY`](../../src/purgatory/mod.rs)
|
||||
|
||||
---
|
||||
|
||||
## Testability: Mock-Based Architecture
|
||||
|
||||
A key design goal was **100% unit test coverage** without requiring real git servers or databases.
|
||||
|
||||
### SyncContext Trait
|
||||
|
||||
All external dependencies are abstracted behind the `SyncContext` trait:
|
||||
|
||||
```rust
|
||||
#[async_trait]
|
||||
pub trait SyncContext: Send + Sync {
|
||||
async fn fetch_repository_data(&self, identifier: &str) -> Result<RepositoryData>;
|
||||
fn collect_needed_oids(&self, identifier: &str) -> HashSet<String>;
|
||||
async fn oid_exists(&self, repo_path: &Path, oid: &str) -> bool;
|
||||
async fn fetch_oids(&self, repo_path: &Path, url: &str, oids: &[String]) -> Result<Vec<String>>;
|
||||
async fn process_newly_available_git_data(&self, ...) -> Result<ProcessResult>;
|
||||
fn has_pending_events(&self, identifier: &str) -> bool;
|
||||
fn find_target_repo(&self, data: &RepositoryData) -> Option<PathBuf>;
|
||||
fn our_domain(&self) -> Option<&str>;
|
||||
}
|
||||
```
|
||||
|
||||
**Two Implementations**:
|
||||
|
||||
1. **`RealSyncContext`** - Production implementation connecting to real systems
|
||||
2. **`MockSyncContext`** - Test implementation with configurable behavior
|
||||
|
||||
### MockSyncContext Features
|
||||
|
||||
The mock supports builder-pattern configuration:
|
||||
|
||||
```rust
|
||||
let mock = MockSyncContext::new()
|
||||
.with_repository_data("test-repo", RepositoryData {
|
||||
announcements: vec![...],
|
||||
clone_urls: vec!["https://server1.com/repo.git".to_string()],
|
||||
})
|
||||
.with_needed_oids("test-repo", hashset!["abc123", "def456"])
|
||||
.with_fetch_result("https://server1.com/repo.git", Ok(vec!["abc123"]))
|
||||
.with_fetch_result("https://server2.com/repo.git", Ok(vec!["def456"]));
|
||||
```
|
||||
|
||||
**Test Example** (from [`src/purgatory/sync/functions.rs`](../../src/purgatory/sync/functions.rs)):
|
||||
|
||||
```rust
|
||||
#[tokio::test]
|
||||
async fn test_sync_identifier_partial_success() {
|
||||
let mock = MockSyncContext::new()
|
||||
.with_repository_data("repo", RepositoryData {
|
||||
clone_urls: vec![
|
||||
"https://server1.com/repo.git".to_string(),
|
||||
"https://server2.com/repo.git".to_string(),
|
||||
],
|
||||
..Default::default()
|
||||
})
|
||||
.with_needed_oids("repo", hashset!["oid1", "oid2"])
|
||||
.with_fetch_result("https://server1.com/repo.git", Ok(vec!["oid1"]))
|
||||
.with_fetch_result("https://server2.com/repo.git", Ok(vec!["oid2"]));
|
||||
|
||||
let throttle = Arc::new(ThrottleManager::new(5, 30));
|
||||
let complete = sync_identifier(&mock, "repo", &throttle).await;
|
||||
|
||||
assert!(complete); // Both OIDs fetched
|
||||
}
|
||||
```
|
||||
|
||||
**Why this matters**:
|
||||
|
||||
- ✅ Tests run **instantly** (no network I/O)
|
||||
- ✅ Tests are **deterministic** (no flaky failures)
|
||||
- ✅ Tests cover **edge cases** easily (network errors, partial success, etc.)
|
||||
- ✅ Tests are **isolated** (no shared state between tests)
|
||||
|
||||
**Implementation**: [`src/purgatory/sync/context.rs:MockSyncContext`](../../src/purgatory/sync/context.rs)
|
||||
|
||||
---
|
||||
|
||||
## Configuration
|
||||
|
||||
Purgatory sync behavior is configurable via CLI flags or environment variables:
|
||||
|
||||
| Setting | CLI Flag | Environment Variable | Default | Description |
|
||||
| ----------------------- | -------- | -------------------- | ------- | ---------------------------------------------------- |
|
||||
| Domain concurrent limit | (future) | (future) | `5` | Max concurrent requests per domain |
|
||||
| Domain rate limit | (future) | (future) | `30` | Max requests per minute per domain |
|
||||
| Sync loop interval | N/A | N/A | `1s` | How often to check for ready identifiers (hardcoded) |
|
||||
| Default sync delay | N/A | N/A | `180s` | Delay for user-submitted events (hardcoded) |
|
||||
| Immediate sync delay | N/A | N/A | `500ms` | Delay for sync-triggered events (hardcoded) |
|
||||
| Purgatory expiry | N/A | N/A | `30min` | How long events wait before expiring (hardcoded) |
|
||||
|
||||
**Note**: Currently, throttle limits and delays are hardcoded constants. Future work may expose these as configuration options if needed.
|
||||
|
||||
---
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
### 1. Identifier-Based, Not Event-Based
|
||||
|
||||
**Decision**: Sync by repository identifier, not individual events.
|
||||
|
||||
**Rationale**: Multiple events for the same repository should trigger a single fetch operation, not N separate fetches.
|
||||
|
||||
**Impact**: Batches events efficiently, reduces server load.
|
||||
|
||||
### 2. Two Separate `tried_urls` Tracking
|
||||
|
||||
**Decision**: Main sync loop and domain queues track tried URLs independently.
|
||||
|
||||
**Main sync**: Local `HashSet<String>` for current attempt (all domains)
|
||||
**Domain queue**: Per-identifier `HashSet<String>` for this domain only
|
||||
|
||||
**Rationale**:
|
||||
|
||||
- Main sync skips throttled domains entirely (doesn't need their tried URLs)
|
||||
- Domain queue only cares about URLs from its own domain
|
||||
- No coordination needed → simpler code
|
||||
|
||||
**Impact**: Clean separation of concerns, easier to reason about.
|
||||
|
||||
### 3. Trigger-Based Domain Processing
|
||||
|
||||
**Decision**: Domain queues process on triggers (capacity freed, new enqueue), not polling.
|
||||
|
||||
**Rationale**:
|
||||
|
||||
- Polling wastes CPU cycles checking capacity every interval
|
||||
- Triggers provide instant response when capacity frees
|
||||
- Event-driven design is easier to test and debug
|
||||
|
||||
**Impact**: Lower CPU usage, faster response times.
|
||||
|
||||
### 4. Fresh Start on New Events
|
||||
|
||||
**Decision**: Reset `attempt_count` to 0 when new events arrive for an identifier.
|
||||
|
||||
**Rationale**:
|
||||
|
||||
- New events often mean fresh git data is available
|
||||
- Previous failures might have been temporary
|
||||
- Gives repositories a "second chance" without waiting for full backoff
|
||||
|
||||
**Impact**: Faster recovery from transient failures, better UX.
|
||||
|
||||
### 5. OID Copying in `process_newly_available_git_data`
|
||||
|
||||
**Decision**: Copy OIDs and release events **per successful fetch**, not at end of sync.
|
||||
|
||||
**Rationale**:
|
||||
|
||||
- Events can be released as soon as their specific OIDs are available
|
||||
- Partial success scenarios work correctly (some events release, others stay)
|
||||
- Handles multiple state events for same identifier independently
|
||||
|
||||
**Impact**: Events release faster, better handling of partial success.
|
||||
|
||||
---
|
||||
|
||||
## Observability
|
||||
|
||||
### Logging
|
||||
|
||||
Sync operations produce structured logs at different levels:
|
||||
|
||||
**INFO**: Major events
|
||||
|
||||
```
|
||||
Starting purgatory sync loop (interval: 1s)
|
||||
Sync complete - removed from sync queue (identifier=test-repo, complete=true)
|
||||
```
|
||||
|
||||
**DEBUG**: Detailed progress
|
||||
|
||||
```
|
||||
Added new sync queue entry (identifier=test-repo, delay_secs=180)
|
||||
Starting sync task for identifier (identifier=test-repo)
|
||||
Sync incomplete - applying backoff (identifier=test-repo, attempt_count=2, next_backoff_secs=40)
|
||||
```
|
||||
|
||||
**WARN**: Errors and failures
|
||||
|
||||
```
|
||||
Failed to fetch OIDs (url=https://server.com/repo.git, error=connection timeout)
|
||||
```
|
||||
|
||||
### Metrics (Future)
|
||||
|
||||
Planned Prometheus metrics for observability:
|
||||
|
||||
- `purgatory_sync_queue_size` - Number of identifiers pending sync
|
||||
- `purgatory_sync_attempts_total{identifier}` - Total sync attempts per identifier
|
||||
- `purgatory_sync_oids_fetched_total{identifier}` - OIDs successfully fetched
|
||||
- `purgatory_domain_in_flight{domain}` - Current in-flight requests per domain
|
||||
- `purgatory_domain_requests_total{domain}` - Total requests per domain
|
||||
|
||||
---
|
||||
|
||||
## Testing Strategy
|
||||
|
||||
### Unit Tests
|
||||
|
||||
Core sync functions have comprehensive unit tests using `MockSyncContext`:
|
||||
|
||||
**`sync_identifier_next_url`** (3 tests):
|
||||
|
||||
- Skips throttled domains
|
||||
- Skips tried URLs
|
||||
- Returns None when all URLs exhausted
|
||||
|
||||
**`sync_identifier_from_url`** (2 tests):
|
||||
|
||||
- Successful fetch triggers processing
|
||||
- Failed fetch doesn't trigger processing
|
||||
|
||||
**`sync_identifier`** (3 tests):
|
||||
|
||||
- Tries multiple URLs until complete
|
||||
- Enqueues throttled domains when incomplete
|
||||
- Handles partial success correctly
|
||||
|
||||
**`SyncQueueEntry`** (3 tests):
|
||||
|
||||
- Backoff calculation correct
|
||||
- Fresh start on new events
|
||||
- Ready state logic correct
|
||||
|
||||
**`DomainThrottle`** (4 tests):
|
||||
|
||||
- Concurrent limit enforced
|
||||
- Rate limit enforced
|
||||
- Round-robin fairness
|
||||
- Queue management correct
|
||||
|
||||
**Total**: 15+ unit tests covering all core logic
|
||||
|
||||
**Location**: [`src/purgatory/sync/`](../../src/purgatory/sync/) (various `#[cfg(test)]` modules)
|
||||
|
||||
### Integration Tests
|
||||
|
||||
End-to-end tests verify sync behavior with real relay instances:
|
||||
|
||||
**Planned tests**:
|
||||
|
||||
- State event syncs from remote server
|
||||
- PR event syncs from remote server
|
||||
- Partial OID aggregation across multiple servers
|
||||
- Throttling prevents overwhelming servers
|
||||
- Backoff retry after failures
|
||||
|
||||
**Location**: [`tests/purgatory_sync.rs`](../../tests/purgatory_sync.rs) (planned)
|
||||
|
||||
---
|
||||
|
||||
## Future Enhancements
|
||||
|
||||
### 1. Configurable Throttle Limits
|
||||
|
||||
**Current**: Hardcoded to 5 concurrent, 30/min per domain
|
||||
**Future**: CLI flags `--sync-domain-concurrent` and `--sync-domain-rate-limit`
|
||||
|
||||
**Use case**: Operators might want stricter limits for public servers or looser limits for trusted servers.
|
||||
|
||||
### 2. Per-Domain Throttle Configuration
|
||||
|
||||
**Current**: Same limits for all domains
|
||||
**Future**: Domain-specific overrides (e.g., `github.com:10,60` for higher limits)
|
||||
|
||||
**Use case**: Popular forges like GitHub/GitLab can handle more load than small personal servers.
|
||||
|
||||
### 3. Prometheus Metrics
|
||||
|
||||
**Current**: Structured logging only
|
||||
**Future**: Export metrics for monitoring dashboards
|
||||
|
||||
**Use case**: Operators want visibility into sync performance, throttle effectiveness, success rates.
|
||||
|
||||
### 4. Negentropy Integration
|
||||
|
||||
**Current**: Sync triggered by event arrival
|
||||
**Future**: Proactive sync discovers missing events via negentropy
|
||||
|
||||
**Use case**: Catch up with repositories after downtime without waiting for event re-submission.
|
||||
|
||||
---
|
||||
|
||||
## Related Documentation
|
||||
|
||||
- **[Purgatory Design](purgatory-design.md)** - Core purgatory concepts and event flows
|
||||
- **[GRASP-02 Proactive Sync](grasp-02-proactive-sync.md)** - Full GRASP-02 implementation (relay sync)
|
||||
- **[Unified Git Data Sync](unify-git-data-sync.md)** - Shared processing for push and sync paths
|
||||
- **[Architecture Overview](architecture.md)** - System-wide architecture
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
The purgatory sync system is a sophisticated, production-ready implementation that:
|
||||
|
||||
✅ **Batches intelligently** - Groups events by identifier for efficient fetching
|
||||
✅ **Retries smartly** - Exponential backoff with fresh start on new events
|
||||
✅ **Throttles respectfully** - 5 concurrent + 30/min per domain, round-robin fairness
|
||||
✅ **Times strategically** - 3min for user events, 500ms for synced events
|
||||
✅ **Expires responsibly** - 30min auto-cleanup prevents memory leaks
|
||||
✅ **Tests thoroughly** - Mock-based architecture enables comprehensive unit tests
|
||||
|
||||
This design ensures ngit-grasp can serve repositories reliably even when git data and Nostr events arrive out-of-order or from different sources, while respecting remote server capacity and providing excellent observability.
|
||||
@@ -4,6 +4,8 @@
|
||||
|
||||
Proactively Sync Nostr Events from other relays listed in accepted repository announcements.
|
||||
|
||||
**Note**: This document covers **relay-to-relay event sync**. For automatic git data fetching when events arrive without their data, see [GRASP-02 Purgatory Git Data Fetching](grasp-02-proactive-sync-purgatory-git-data.md).
|
||||
|
||||
Features:
|
||||
|
||||
- Fetches all repository announcements from connected relays to discover new repos listing our service
|
||||
@@ -13,6 +15,7 @@ Features:
|
||||
- Plays nicely with other relays - connection backoff and rate-limiting detection with cooldown
|
||||
- Does a full reconciliation daily
|
||||
- Prometheus metrics
|
||||
- **Triggers purgatory git data sync**: When events arrive via sync, they're enqueued for immediate git data fetching (500ms delay to batch bursts)
|
||||
|
||||
Key Architectural Points:
|
||||
|
||||
|
||||
@@ -37,16 +37,18 @@ Client Server
|
||||
```
|
||||
|
||||
**Pros:**
|
||||
|
||||
- Standard Git mechanism
|
||||
- Language-agnostic (hook can be any executable)
|
||||
- Well-documented
|
||||
|
||||
**Cons:**
|
||||
|
||||
- Hook output goes to stderr (client sees as `remote:` messages)
|
||||
- Hard to provide structured error messages
|
||||
- Requires hook installation and management
|
||||
- Difficult to test (needs Git repository setup)
|
||||
- Hook runs *after* Git has started processing
|
||||
- Hook runs _after_ Git has started processing
|
||||
|
||||
---
|
||||
|
||||
@@ -60,7 +62,7 @@ Client Server (ngit-grasp)
|
||||
|--- git push ----->|--- HTTP handler receives request
|
||||
| |
|
||||
| |--- Parse ref updates from request
|
||||
| |--- Query Nostr relay for state
|
||||
| |--- Query database + purgatory for state
|
||||
| |--- Validate push against state
|
||||
| |
|
||||
| |--- If invalid: return HTTP error
|
||||
@@ -71,13 +73,16 @@ Client Server (ngit-grasp)
|
||||
```
|
||||
|
||||
**Pros:**
|
||||
|
||||
- Full control over error messages (HTTP response)
|
||||
- Can skip spawning Git entirely for invalid pushes
|
||||
- Easier testing (pure Rust, no Git setup needed)
|
||||
- Shared state between Git and Nostr components
|
||||
- Better performance (early rejection)
|
||||
- Can check both database and purgatory for authorization
|
||||
|
||||
**Cons:**
|
||||
|
||||
- Requires parsing Git protocol ourselves
|
||||
- Less standard than hooks
|
||||
- Tighter coupling to Git HTTP protocol
|
||||
@@ -86,9 +91,41 @@ Client Server (ngit-grasp)
|
||||
|
||||
## Why Inline Authorization Is Better for GRASP
|
||||
|
||||
### 1. Better Error Messages
|
||||
### 1. Purgatory Integration
|
||||
|
||||
**Critical advantage:** Inline authorization allows checking **both database and purgatory** during authorization:
|
||||
|
||||
```rust
|
||||
// From src/git/authorization.rs
|
||||
pub async fn authorize_push(
|
||||
database: &SharedDatabase,
|
||||
identifier: &str,
|
||||
owner_pubkey: &str,
|
||||
request_body: &Bytes,
|
||||
purgatory: &Arc<Purgatory>, // Can check purgatory!
|
||||
repo_path: &std::path::Path,
|
||||
) -> anyhow::Result<AuthorizationResult>
|
||||
```
|
||||
|
||||
**Why this matters:** State events go to purgatory when git data doesn't exist yet. Without inline authorization checking purgatory, we'd have a deadlock:
|
||||
|
||||
1. State event arrives → No git data → Goes to **purgatory** (not database)
|
||||
2. Git push arrives → Hook checks **database only** → No state found → **REJECTED** ❌
|
||||
|
||||
With inline authorization:
|
||||
|
||||
1. State event arrives → No git data → Goes to purgatory
|
||||
2. Git push arrives → Checks **database + purgatory** → State found → **AUTHORIZED** ✅
|
||||
3. After push succeeds → Save event to database → Remove from purgatory
|
||||
|
||||
See [`src/git/authorization.rs:342-400`](../../src/git/authorization.rs) for implementation.
|
||||
|
||||
otherwise we'd need another way of storing purgatory events.
|
||||
|
||||
### 2. Better Error Messages
|
||||
|
||||
**With hooks:**
|
||||
|
||||
```
|
||||
$ git push
|
||||
remote: error: Push rejected - not authorized for ref refs/heads/main
|
||||
@@ -98,39 +135,37 @@ To https://gitnostr.com/alice/myrepo.git
|
||||
```
|
||||
|
||||
**With inline authorization:**
|
||||
|
||||
```
|
||||
$ git push
|
||||
error: RPC failed; HTTP 403 Forbidden
|
||||
error: {
|
||||
"error": "unauthorized",
|
||||
"ref": "refs/heads/main",
|
||||
"required_state": "event_id_abc123",
|
||||
"your_pubkey": "npub1alice...",
|
||||
"docs": "https://docs.gitnostr.com/errors/unauthorized"
|
||||
}
|
||||
error: Push rejected: No state event found in purgatory from authorized publishers
|
||||
```
|
||||
|
||||
The inline approach can return **structured JSON** with actionable information.
|
||||
The inline approach provides clear, actionable error messages directly in the HTTP response.
|
||||
|
||||
### 2. Performance Benefits
|
||||
### 3. Performance Benefits
|
||||
|
||||
**With hooks:**
|
||||
|
||||
- Git process spawns
|
||||
- Git starts receiving pack data
|
||||
- Hook runs (might query Nostr relay)
|
||||
- If rejected, Git throws away received data
|
||||
|
||||
**With inline authorization:**
|
||||
- Parse ref updates from HTTP request
|
||||
- Validate against Nostr state (cached)
|
||||
- If rejected, return HTTP 403 immediately
|
||||
|
||||
- Parse ref updates from HTTP request (pkt-line format)
|
||||
- Validate against database + purgatory state
|
||||
- If rejected, return HTTP error immediately
|
||||
- Never spawn Git for invalid pushes
|
||||
|
||||
**Result:** Faster rejection, less resource usage.
|
||||
**Result:** Faster rejection, less resource usage, no wasted pack data transfer.
|
||||
|
||||
### 3. Easier Testing
|
||||
### 4. Easier Testing
|
||||
|
||||
**With hooks:**
|
||||
|
||||
```bash
|
||||
# Test setup
|
||||
mkdir -p /tmp/test-repo
|
||||
@@ -147,6 +182,7 @@ rm -rf /tmp/test-repo
|
||||
```
|
||||
|
||||
**With inline authorization:**
|
||||
|
||||
```rust
|
||||
#[tokio::test]
|
||||
async fn test_unauthorized_push() {
|
||||
@@ -161,43 +197,55 @@ async fn test_unauthorized_push() {
|
||||
|
||||
See [`tests/push_authorization.rs`](tests/push_authorization.rs) for actual test examples.
|
||||
|
||||
### 4. Shared State and Types
|
||||
### 5. Shared State and Types
|
||||
|
||||
**With hooks:**
|
||||
|
||||
- Hook is separate process
|
||||
- Must query Nostr relay over WebSocket
|
||||
- Can't share in-memory cache
|
||||
- Can't access purgatory
|
||||
- Separate error types
|
||||
|
||||
**With inline authorization:**
|
||||
|
||||
```rust
|
||||
// From src/git/handlers.rs
|
||||
pub async fn handle_receive_pack(
|
||||
repo_path: PathBuf,
|
||||
body: Bytes,
|
||||
database: SharedDatabase, // Shared with Nostr relay!
|
||||
database: Option<SharedDatabase>, // Shared with Nostr relay!
|
||||
purgatory: Option<Arc<Purgatory>>, // Shared purgatory access!
|
||||
npub: &str,
|
||||
identifier: &str,
|
||||
) -> Result<Response<Full<Bytes>>, GitError> {
|
||||
// Direct database access for authorization
|
||||
let auth = get_authorization_for_owner(&database, pubkey, identifier).await?;
|
||||
// Direct database + purgatory access for authorization
|
||||
let auth = authorize_push(
|
||||
&database,
|
||||
identifier,
|
||||
owner_pubkey,
|
||||
&body,
|
||||
&purgatory, // Can check purgatory!
|
||||
&repo_path
|
||||
).await?;
|
||||
// ...
|
||||
}
|
||||
```
|
||||
|
||||
**Result:** Better performance, type safety, simpler architecture.
|
||||
**Result:** Better performance, type safety, simpler architecture, purgatory integration.
|
||||
|
||||
### 5. Simpler Deployment
|
||||
### 6. Simpler Deployment
|
||||
|
||||
**With hooks (ngit-relay):**
|
||||
|
||||
```
|
||||
Docker container:
|
||||
- nginx (HTTP frontend)
|
||||
- git-http-backend (C binary)
|
||||
- pre-receive hook (Go binary)
|
||||
- pre-receive hook (Go binary)
|
||||
- Khatru relay (Go binary)
|
||||
- supervisord (process manager)
|
||||
|
||||
|
||||
Setup steps:
|
||||
1. Install all components
|
||||
2. Configure nginx
|
||||
@@ -207,13 +255,14 @@ Setup steps:
|
||||
```
|
||||
|
||||
**With inline authorization (ngit-grasp):**
|
||||
|
||||
```
|
||||
Single Rust binary:
|
||||
- HTTP server (Hyper)
|
||||
- Git protocol handler
|
||||
- Nostr relay (nostr-relay-builder)
|
||||
- Authorization logic
|
||||
|
||||
|
||||
Setup steps:
|
||||
1. Run binary
|
||||
2. Configure environment variables
|
||||
@@ -227,66 +276,95 @@ Setup steps:
|
||||
|
||||
### How We Parse Ref Updates
|
||||
|
||||
The Git HTTP protocol sends ref updates in the request body:
|
||||
The Git HTTP protocol sends ref updates in pkt-line format:
|
||||
|
||||
```
|
||||
POST /alice/myrepo.git/git-receive-pack HTTP/1.1
|
||||
Content-Type: application/x-git-receive-pack-request
|
||||
|
||||
0000000000000000000000000000000000000000 abc123... refs/heads/main\0 report-status
|
||||
00a5 0000...0000 abc123...def456 refs/heads/main\0 report-status\n
|
||||
0000
|
||||
PACK...
|
||||
```
|
||||
|
||||
We parse this **before** spawning Git. See [`src/git/authorization.rs`](src/git/authorization.rs) for the implementation:
|
||||
We parse this **before** spawning Git. See [`src/git/authorization.rs:695-778`](../../src/git/authorization.rs) for the implementation:
|
||||
|
||||
```rust
|
||||
/// Parse ref updates from git-receive-pack request body
|
||||
pub fn parse_pushed_refs(body: &[u8]) -> Result<Vec<PushedRef>, AuthorizationError> {
|
||||
// Parse pkt-line format
|
||||
// Extract ref updates
|
||||
// Return structured data
|
||||
/// Parse the refs being updated from a Git pack
|
||||
///
|
||||
/// The receive-pack protocol sends ref updates in pkt-line format:
|
||||
/// - 4-byte hex length prefix (e.g., "00a5")
|
||||
/// - Payload: `<old-oid> <new-oid> <ref-name>\0<capabilities>\n`
|
||||
/// - Flush packet "0000" terminates the list
|
||||
pub fn parse_pushed_refs(data: &[u8]) -> Vec<(String, String, String)> {
|
||||
// Handles both pkt-line format (real Git clients)
|
||||
// and simple text format (for unit tests)
|
||||
}
|
||||
```
|
||||
|
||||
### How We Validate
|
||||
|
||||
Validation checks (from [`src/git/authorization.rs`](src/git/authorization.rs)):
|
||||
|
||||
1. Does pusher's pubkey have write access?
|
||||
2. Are they listed as a maintainer in the latest state event?
|
||||
3. Do the refs match the state event?
|
||||
The authorization flow (from [`src/git/authorization.rs:51-162`](../../src/git/authorization.rs)):
|
||||
|
||||
```rust
|
||||
/// Validate that pushed refs match the authorized state
|
||||
pub fn validate_push_refs(
|
||||
pushed_refs: &[PushedRef],
|
||||
state: &RepositoryState,
|
||||
) -> Result<(), AuthorizationError> {
|
||||
for pushed_ref in pushed_refs {
|
||||
if pushed_ref.ref_name.starts_with("refs/heads/") {
|
||||
// Validate branch against state
|
||||
} else if pushed_ref.ref_name.starts_with("refs/tags/") {
|
||||
// Validate tag against state
|
||||
} else if pushed_ref.ref_name.starts_with("refs/nostr/") {
|
||||
// Allow refs/nostr/<event-id> for PRs
|
||||
}
|
||||
}
|
||||
Ok(())
|
||||
pub async fn authorize_push(
|
||||
database: &SharedDatabase,
|
||||
identifier: &str,
|
||||
owner_pubkey: &str,
|
||||
request_body: &Bytes,
|
||||
purgatory: &Arc<Purgatory>,
|
||||
repo_path: &std::path::Path,
|
||||
) -> anyhow::Result<AuthorizationResult> {
|
||||
// 1. Parse refs from push request
|
||||
let pushed_refs = parse_pushed_refs(request_body);
|
||||
|
||||
// 2. Separate refs/nostr/ refs from state refs
|
||||
let (nostr_refs, state_refs) = partition_refs(&pushed_refs);
|
||||
|
||||
// 3. Handle refs/nostr/ refs (PR events)
|
||||
// - Validate event ID format
|
||||
// - Check purgatory for PR event
|
||||
// - Create placeholder if git-data-first scenario
|
||||
|
||||
// 4. Handle normal refs (state events)
|
||||
// - Check database + purgatory for state events
|
||||
// - Collect authorized maintainers
|
||||
// - Find latest authorized state
|
||||
// - Validate refs match state
|
||||
|
||||
// 5. Return authorization result with purgatory events
|
||||
}
|
||||
```
|
||||
|
||||
**Key validation checks:**
|
||||
|
||||
1. **For state refs** (`refs/heads/*`, `refs/tags/*`):
|
||||
|
||||
- Query database for announcements → collect authorized maintainers
|
||||
- Check **purgatory** for matching state events (critical for purgatory flow!)
|
||||
- Filter to events from authorized maintainers
|
||||
- Find latest state event
|
||||
- Validate pushed refs match state event refs
|
||||
|
||||
2. **For PR refs** (`refs/nostr/<event-id>`):
|
||||
- Validate event ID format
|
||||
- Check purgatory for PR event with matching commit
|
||||
- If no event found, create placeholder (git-data-first scenario)
|
||||
- Collect PR events from purgatory for post-push processing
|
||||
|
||||
---
|
||||
|
||||
## Comparison with Reference Implementation
|
||||
|
||||
| Aspect | ngit-relay (hooks) | ngit-grasp (inline) |
|
||||
|--------|-------------------|---------------------|
|
||||
| **Components** | nginx + git-http-backend + hook + Khatru | Single Rust binary |
|
||||
| **Validation** | Pre-receive hook (separate process) | Inline HTTP handler |
|
||||
| **Error messages** | Hook stderr → `remote:` | HTTP response JSON |
|
||||
| **Performance** | Spawns Git first | Validates first |
|
||||
| **Testing** | Shell scripts + Go tests | Pure Rust tests |
|
||||
| **Deployment** | Docker + supervisord | Single binary |
|
||||
| **State sharing** | WebSocket query | Direct database access |
|
||||
| Aspect | ngit-relay (hooks) | ngit-grasp (inline) |
|
||||
| ------------------ | ---------------------------------------- | ---------------------- |
|
||||
| **Components** | nginx + git-http-backend + hook + Khatru | Single Rust binary |
|
||||
| **Validation** | Pre-receive hook (separate process) | Inline HTTP handler |
|
||||
| **Error messages** | Hook stderr → `remote:` | HTTP response JSON |
|
||||
| **Performance** | Spawns Git first | Validates first |
|
||||
| **Testing** | Shell scripts + Go tests | Pure Rust tests |
|
||||
| **Deployment** | Docker + supervisord | Single binary |
|
||||
| **State sharing** | WebSocket query | Direct database access |
|
||||
|
||||
Both are GRASP-compliant, but inline authorization is simpler and more efficient.
|
||||
|
||||
@@ -295,24 +373,30 @@ Both are GRASP-compliant, but inline authorization is simpler and more efficient
|
||||
## Trade-offs and Limitations
|
||||
|
||||
### What We Gain
|
||||
|
||||
- ✅ **Purgatory integration** - Can check database + purgatory during authorization
|
||||
- ✅ **Prevents deadlock** - State events in purgatory can authorize pushes
|
||||
- ✅ Better error messages
|
||||
- ✅ Better performance
|
||||
- ✅ Easier testing
|
||||
- ✅ Simpler deployment
|
||||
- ✅ Tighter integration
|
||||
- ✅ Better performance (early rejection)
|
||||
- ✅ Easier testing (pure Rust)
|
||||
- ✅ Simpler deployment (single binary)
|
||||
- ✅ Tighter integration (shared state)
|
||||
|
||||
### What We Lose
|
||||
|
||||
- ❌ Non-standard approach (not using Git's hook system)
|
||||
- ❌ Tighter coupling to Git HTTP protocol
|
||||
- ❌ Must parse protocol ourselves
|
||||
- ❌ Must parse pkt-line protocol ourselves
|
||||
|
||||
### Is It Worth It?
|
||||
|
||||
**Yes**, because:
|
||||
1. We handle protocol parsing in [`src/git/protocol.rs`](src/git/protocol.rs)
|
||||
2. GRASP is already non-standard (Nostr authorization)
|
||||
3. Benefits far outweigh the coupling cost
|
||||
4. We can still add hook support later if needed
|
||||
**Absolutely**, because:
|
||||
|
||||
1. **Purgatory integration is essential** - Without it, we'd have a deadlock where state events in purgatory can't authorize pushes
|
||||
2. Protocol parsing is isolated in [`src/git/authorization.rs`](../../src/git/authorization.rs)
|
||||
3. GRASP is already non-standard (Nostr authorization)
|
||||
4. Benefits far outweigh the coupling cost
|
||||
5. We can still add hook support later if needed (but purgatory checking would still need inline access)
|
||||
|
||||
---
|
||||
|
||||
@@ -320,14 +404,15 @@ Both are GRASP-compliant, but inline authorization is simpler and more efficient
|
||||
|
||||
Key files in the ngit-grasp implementation:
|
||||
|
||||
| Component | Location |
|
||||
|-----------|----------|
|
||||
| HTTP routing | [`src/http/mod.rs`](src/http/mod.rs) |
|
||||
| Git handlers | [`src/git/handlers.rs`](src/git/handlers.rs) |
|
||||
| Push authorization | [`src/git/authorization.rs`](src/git/authorization.rs) |
|
||||
| Git protocol parsing | [`src/git/protocol.rs`](src/git/protocol.rs) |
|
||||
| Subprocess management | [`src/git/subprocess.rs`](src/git/subprocess.rs) |
|
||||
| Event acceptance policy | [`src/nostr/builder.rs:51`](src/nostr/builder.rs:51) - `Nip34WritePolicy` |
|
||||
| Component | Location |
|
||||
| ----------------------- | ------------------------------------------------------------------------- |
|
||||
| HTTP routing | [`src/http/mod.rs`](../../src/http/mod.rs) |
|
||||
| Git handlers | [`src/git/handlers.rs`](../../src/git/handlers.rs) |
|
||||
| Push authorization | [`src/git/authorization.rs`](../../src/git/authorization.rs) |
|
||||
| Pkt-line parsing | [`src/git/authorization.rs:695-778`](../../src/git/authorization.rs) |
|
||||
| Subprocess management | [`src/git/subprocess.rs`](../../src/git/subprocess.rs) |
|
||||
| Purgatory integration | [`src/purgatory/mod.rs`](../../src/purgatory/mod.rs) |
|
||||
| Event acceptance policy | [`src/nostr/builder.rs`](../../src/nostr/builder.rs) - `Nip34WritePolicy` |
|
||||
|
||||
---
|
||||
|
||||
@@ -345,6 +430,7 @@ pub struct GitConfig {
|
||||
```
|
||||
|
||||
This would allow:
|
||||
|
||||
- Migration path for hook-based systems
|
||||
- Extra validation for paranoid deployments
|
||||
- Compatibility with other Git tools
|
||||
@@ -352,6 +438,7 @@ This would allow:
|
||||
### If Git Protocol Changes
|
||||
|
||||
The protocol parsing is isolated in [`src/git/protocol.rs`](src/git/protocol.rs). If the Git protocol changes:
|
||||
|
||||
- Update the protocol module
|
||||
- Tests will catch any breakage
|
||||
|
||||
@@ -361,18 +448,21 @@ The protocol parsing is isolated in [`src/git/protocol.rs`](src/git/protocol.rs)
|
||||
|
||||
**Inline authorization is the right choice for ngit-grasp** because:
|
||||
|
||||
1. It provides better error messages for users
|
||||
2. It's more performant (early rejection)
|
||||
3. It's easier to test (pure Rust)
|
||||
4. It's simpler to deploy (single binary)
|
||||
5. It enables better integration (shared database)
|
||||
1. **Purgatory integration** - Without inline authorization, state events in purgatory couldn't authorize pushes, creating a deadlock
|
||||
2. **Better error messages** - Direct HTTP responses with clear rejection reasons
|
||||
3. **Better performance** - Early rejection before spawning Git
|
||||
4. **Easier testing** - Pure Rust unit tests, no Git setup needed
|
||||
5. **Simpler deployment** - Single binary with shared state
|
||||
6. **Shared database + purgatory** - Both authorization sources accessible during validation
|
||||
|
||||
The trade-off (coupling to Git HTTP protocol) is acceptable because:
|
||||
- The protocol is stable and well-specified
|
||||
- Protocol handling is isolated in one module
|
||||
|
||||
- The pkt-line protocol is stable and well-specified
|
||||
- Protocol parsing is isolated in [`src/git/authorization.rs`](../../src/git/authorization.rs)
|
||||
- Purgatory integration requires inline access anyway
|
||||
- Benefits far outweigh the cost
|
||||
|
||||
This decision aligns with our goal of creating a **developer-friendly, production-ready GRASP implementation**.
|
||||
This decision aligns with our goal of creating a **developer-friendly, production-ready GRASP implementation** that properly handles the event-git-data ordering problem via purgatory.
|
||||
|
||||
---
|
||||
|
||||
@@ -386,4 +476,4 @@ This decision aligns with our goal of creating a **developer-friendly, productio
|
||||
|
||||
---
|
||||
|
||||
*Part of the [ngit-grasp explanation docs](./)*
|
||||
_Part of the [ngit-grasp explanation docs](./)_
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -1,481 +0,0 @@
|
||||
# Unified Git Data Sync
|
||||
|
||||
## Status
|
||||
|
||||
**Proposed** - January 2026
|
||||
|
||||
## Context
|
||||
|
||||
Currently, two separate code paths handle "git data is now available" scenarios:
|
||||
|
||||
1. **`handle_receive_pack`** (src/git/handlers.rs) - After a successful `git push`
|
||||
2. **`sync_state_git_data`** (src/purgatory/mod.rs) - After purgatory sync fetches OIDs from remote servers
|
||||
|
||||
Both paths perform essentially the same post-processing:
|
||||
|
||||
| Step | `handle_receive_pack` | `sync_state_git_data` |
|
||||
|------|----------------------|----------------------|
|
||||
| Set HEAD | ✅ `try_set_head_if_available()` | ✅ (via `align_repository_with_state`) |
|
||||
| Save events to DB | ✅ `database.save_event()` | ✅ `database.save_event()` |
|
||||
| Remove from purgatory | ✅ `remove_state_event()` / `remove_pr()` | ✅ `remove_state_event()` |
|
||||
| Notify WebSocket | ✅ `relay.notify_event()` | ✅ `relay.notify_event()` |
|
||||
| Sync state to owner repos | ✅ `sync_to_owner_repos()` | ✅ `sync_to_owner_repos()` |
|
||||
| Sync PR refs to owner repos | ✅ `sync_pr_refs_to_tagged_owner_repos()` | ❌ Not implemented |
|
||||
|
||||
This duplication creates maintenance burden and inconsistent behavior (e.g., PR sync missing from purgatory path).
|
||||
|
||||
## Decision
|
||||
|
||||
Create a single unified function that handles all post-git-data-available processing:
|
||||
|
||||
```rust
|
||||
pub async fn process_newly_available_git_data(
|
||||
source_repo_path: &Path,
|
||||
new_oids: &HashSet<String>,
|
||||
database: &SharedDatabase,
|
||||
local_relay: Option<&nostr_relay_builder::LocalRelay>,
|
||||
purgatory: &Purgatory,
|
||||
git_data_path: &Path,
|
||||
) -> ProcessResult
|
||||
```
|
||||
|
||||
### Key Design Principles
|
||||
|
||||
**1. Always discover events from purgatory**
|
||||
|
||||
Rather than accepting pre-authorized events (which may have changed since authorization), the function always scans purgatory to find satisfiable events. This ensures consistency and handles race conditions where events change between authorization and processing.
|
||||
|
||||
**2. Minimal input, maximal output**
|
||||
|
||||
Callers only need to provide:
|
||||
- `source_repo_path` - Where the git data landed
|
||||
- `new_oids` - Which OIDs are now available (for efficient filtering)
|
||||
|
||||
The function handles everything else: finding events, syncing across repos, aligning refs, setting HEAD, saving to database, notifying subscribers, and cleaning up purgatory.
|
||||
|
||||
**3. Process all event types uniformly**
|
||||
|
||||
Both state events (kind 30618) and PR events (kind 1617/1618) are processed in the same flow, ensuring consistent behavior.
|
||||
|
||||
## Architecture
|
||||
|
||||
### Flow Overview
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────────────────────────┐
|
||||
│ Git Data Becomes Available │
|
||||
│ │
|
||||
│ ┌─────────────────────┐ ┌─────────────────────┐ │
|
||||
│ │ handle_receive_pack │ │ purgatory sync │ │
|
||||
│ │ (push received) │ │ (fetch completed) │ │
|
||||
│ └──────────┬──────────┘ └──────────┬──────────┘ │
|
||||
│ │ │ │
|
||||
│ │ source_repo_path │ source_repo_path │
|
||||
│ │ new_oids │ new_oids │
|
||||
│ │ │ │
|
||||
│ └────────────────┬───────────────────┘ │
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ ┌────────────────────────────────────────┐ │
|
||||
│ │ process_newly_available_git_data() │ │
|
||||
│ │ │ │
|
||||
│ │ 1. Extract identifier from path │ │
|
||||
│ │ 2. Fetch repository data from DB │ │
|
||||
│ │ 3. Find satisfiable state events │ │
|
||||
│ │ 4. Find satisfiable PR events │ │
|
||||
│ │ 5. For each event: │ │
|
||||
│ │ - Sync OIDs to owner repos │ │
|
||||
│ │ - Align refs (+ set HEAD) │ │
|
||||
│ │ - Save to database │ │
|
||||
│ │ - Notify WebSocket │ │
|
||||
│ │ - Remove from purgatory │ │
|
||||
│ └────────────────────────────────────────┘ │
|
||||
└─────────────────────────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### Event Discovery
|
||||
|
||||
The function discovers satisfiable events by scanning purgatory:
|
||||
|
||||
**For State Events:**
|
||||
1. Get all state entries for the identifier from purgatory
|
||||
2. For each entry, check if ALL required OIDs exist in source repo
|
||||
3. Quick optimization: skip if none of `new_oids` are in the state's OID set
|
||||
|
||||
**For PR Events:**
|
||||
1. Get all PR entries for the identifier from purgatory (via secondary index)
|
||||
2. For each entry with an event, check if the commit OID exists in source repo
|
||||
3. Quick optimization: skip if commit not in `new_oids`
|
||||
|
||||
### Sync to Owner Repos
|
||||
|
||||
**For State Events:**
|
||||
|
||||
For each owner whose maintainer set authorizes the state author:
|
||||
1. Skip if a newer state already exists for that owner
|
||||
2. Copy missing OIDs from source repo to target repo
|
||||
3. Align refs (create/update/delete branches and tags)
|
||||
4. Set HEAD per state announcement
|
||||
|
||||
**For PR Events:**
|
||||
|
||||
For each owner whose maintainer set includes any tagged owner (from `a` tags):
|
||||
1. Copy commit from source repo to target repo (if missing)
|
||||
2. Create `refs/nostr/<event-id>` pointing to the commit
|
||||
|
||||
## Data Structure Changes
|
||||
|
||||
### PrPurgatoryEntry
|
||||
|
||||
Add `identifier` field for secondary index lookup:
|
||||
|
||||
```rust
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct PrPurgatoryEntry {
|
||||
/// The nostr PR event, if received (None = git data arrived first)
|
||||
pub event: Option<Event>,
|
||||
|
||||
/// The expected commit SHA from 'c' tag or actual commit pushed
|
||||
pub commit: String,
|
||||
|
||||
/// Repository identifier extracted from 'a' tag (30617:<owner>:<identifier>)
|
||||
/// Used for lookup when git data arrives
|
||||
pub identifier: Option<String>,
|
||||
|
||||
/// When this entry was added to purgatory
|
||||
pub created_at: Instant,
|
||||
|
||||
/// Expiry deadline
|
||||
pub expires_at: Instant,
|
||||
}
|
||||
```
|
||||
|
||||
### Purgatory Secondary Index
|
||||
|
||||
Add index for finding PR events by identifier:
|
||||
|
||||
```rust
|
||||
pub struct Purgatory {
|
||||
/// State events indexed by repository identifier
|
||||
state_events: Arc<DashMap<String, Vec<StatePurgatoryEntry>>>,
|
||||
|
||||
/// PR events indexed by event ID (hex string)
|
||||
pr_events: Arc<DashMap<String, PrPurgatoryEntry>>,
|
||||
|
||||
/// Secondary index: identifier -> event_ids for PR events
|
||||
pr_events_by_identifier: Arc<DashMap<String, HashSet<String>>>,
|
||||
|
||||
git_data_path: PathBuf,
|
||||
}
|
||||
```
|
||||
|
||||
### New Purgatory Methods
|
||||
|
||||
```rust
|
||||
impl Purgatory {
|
||||
/// Find all PR events for an identifier
|
||||
pub fn find_prs_for_identifier(&self, identifier: &str) -> Vec<PrPurgatoryEntry>;
|
||||
|
||||
/// Add PR with automatic identifier extraction and indexing
|
||||
pub fn add_pr(&self, event: Event, event_id: String, commit: String);
|
||||
|
||||
/// Add placeholder with optional identifier
|
||||
pub fn add_pr_placeholder(&self, event_id: String, commit: String, identifier: Option<String>);
|
||||
|
||||
/// Remove PR (also cleans up secondary index)
|
||||
pub fn remove_pr(&self, event_id: &str);
|
||||
}
|
||||
```
|
||||
|
||||
## Implementation
|
||||
|
||||
### Core Function
|
||||
|
||||
```rust
|
||||
/// Unified processing of newly available git data.
|
||||
///
|
||||
/// Called whenever git data becomes available, whether from:
|
||||
/// - A successful `git push` (handle_receive_pack)
|
||||
/// - Purgatory sync fetching OIDs from remote servers
|
||||
///
|
||||
/// # What it does
|
||||
///
|
||||
/// 1. **Discover satisfiable events**: Scans purgatory for state and PR events
|
||||
/// whose required OIDs are now available in `source_repo_path`
|
||||
///
|
||||
/// 2. **For each satisfiable STATE event**:
|
||||
/// - Find all owner repos that authorize this state's author
|
||||
/// - Copy OIDs from source repo to each authorized owner repo
|
||||
/// - Align refs (create/update/delete) to match state
|
||||
/// - Set HEAD per state announcement
|
||||
/// - Save event to database
|
||||
/// - Notify WebSocket subscribers
|
||||
/// - Remove from purgatory
|
||||
///
|
||||
/// 3. **For each satisfiable PR event**:
|
||||
/// - Find all owner repos that list tagged owners as maintainers
|
||||
/// - Copy commit from source repo to each relevant owner repo
|
||||
/// - Create refs/nostr/<event-id> in each repo
|
||||
/// - Save event to database
|
||||
/// - Notify WebSocket subscribers
|
||||
/// - Remove from purgatory
|
||||
pub async fn process_newly_available_git_data(
|
||||
source_repo_path: &Path,
|
||||
new_oids: &HashSet<String>,
|
||||
database: &SharedDatabase,
|
||||
local_relay: Option<&nostr_relay_builder::LocalRelay>,
|
||||
purgatory: &Purgatory,
|
||||
git_data_path: &Path,
|
||||
) -> ProcessResult {
|
||||
let mut result = ProcessResult::default();
|
||||
|
||||
// Extract identifier from repo path
|
||||
let identifier = match extract_identifier_from_repo_path(source_repo_path, git_data_path) {
|
||||
Some(id) => id,
|
||||
None => return result,
|
||||
};
|
||||
|
||||
// Fetch repository data once for all operations
|
||||
let db_repo_data = match fetch_repository_data(database, &identifier).await {
|
||||
Ok(data) => data,
|
||||
Err(e) => {
|
||||
result.errors.push(format!("Failed to fetch repo data: {}", e));
|
||||
return result;
|
||||
}
|
||||
};
|
||||
|
||||
// Process satisfiable state events
|
||||
let state_result = process_satisfiable_state_events(
|
||||
source_repo_path,
|
||||
&identifier,
|
||||
new_oids,
|
||||
&db_repo_data,
|
||||
database,
|
||||
local_relay,
|
||||
purgatory,
|
||||
git_data_path,
|
||||
).await;
|
||||
|
||||
result.merge_state_result(state_result);
|
||||
|
||||
// Process satisfiable PR events
|
||||
let pr_result = process_satisfiable_pr_events(
|
||||
source_repo_path,
|
||||
&identifier,
|
||||
new_oids,
|
||||
&db_repo_data,
|
||||
database,
|
||||
local_relay,
|
||||
purgatory,
|
||||
git_data_path,
|
||||
).await;
|
||||
|
||||
result.merge_pr_result(pr_result);
|
||||
|
||||
result
|
||||
}
|
||||
```
|
||||
|
||||
### Result Type
|
||||
|
||||
```rust
|
||||
/// Result of processing newly available git data
|
||||
#[derive(Debug, Default)]
|
||||
pub struct ProcessResult {
|
||||
/// Number of state events released from purgatory
|
||||
pub states_released: usize,
|
||||
/// Number of PR events released from purgatory
|
||||
pub prs_released: usize,
|
||||
/// Number of owner repositories synced
|
||||
pub repos_synced: usize,
|
||||
/// Number of refs created across all repos
|
||||
pub refs_created: usize,
|
||||
/// Number of refs updated across all repos
|
||||
pub refs_updated: usize,
|
||||
/// Number of refs deleted across all repos
|
||||
pub refs_deleted: usize,
|
||||
/// Errors encountered (non-fatal)
|
||||
pub errors: Vec<String>,
|
||||
}
|
||||
```
|
||||
|
||||
### Helper: Extract Identifier from PR Event
|
||||
|
||||
```rust
|
||||
/// Extract identifier from PR event's `a` tag.
|
||||
/// Format: 30617:<owner_pubkey>:<identifier>
|
||||
fn extract_identifier_from_pr_event(event: &Event) -> Option<String> {
|
||||
event.tags.iter().find_map(|tag| {
|
||||
let tag_vec = tag.clone().to_vec();
|
||||
if tag_vec.len() >= 2 && tag_vec[0] == "a" && tag_vec[1].starts_with("30617:") {
|
||||
let parts: Vec<&str> = tag_vec[1].split(':').collect();
|
||||
if parts.len() >= 3 {
|
||||
Some(parts[2].to_string())
|
||||
} else {
|
||||
None
|
||||
}
|
||||
} else {
|
||||
None
|
||||
}
|
||||
})
|
||||
}
|
||||
```
|
||||
|
||||
### Helper: Extract Identifier from Repo Path
|
||||
|
||||
```rust
|
||||
/// Extract identifier from repository path.
|
||||
/// Path format: {git_data_path}/{npub}/{identifier}.git
|
||||
fn extract_identifier_from_repo_path(repo_path: &Path, git_data_path: &Path) -> Option<String> {
|
||||
let relative = repo_path.strip_prefix(git_data_path).ok()?;
|
||||
let components: Vec<_> = relative.components().collect();
|
||||
|
||||
if components.len() >= 2 {
|
||||
let identifier_with_git = components[1].as_os_str().to_str()?;
|
||||
Some(identifier_with_git.trim_end_matches(".git").to_string())
|
||||
} else {
|
||||
None
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Integration
|
||||
|
||||
### handle_receive_pack (Simplified)
|
||||
|
||||
```rust
|
||||
// After git receive-pack succeeds:
|
||||
|
||||
// Collect new OIDs from the push
|
||||
let new_oids: HashSet<String> = pushed_refs
|
||||
.iter()
|
||||
.filter(|(_, new_oid, _)| new_oid != "0000000000000000000000000000000000000000")
|
||||
.map(|(_, new_oid, _)| new_oid.clone())
|
||||
.collect();
|
||||
|
||||
// Single unified call handles everything
|
||||
let result = process_newly_available_git_data(
|
||||
&repo_path,
|
||||
&new_oids,
|
||||
&database,
|
||||
Some(&relay),
|
||||
&purgatory,
|
||||
Path::new(git_data_path),
|
||||
).await;
|
||||
|
||||
info!(
|
||||
"Processed push: {} states, {} PRs released, {} repos synced",
|
||||
result.states_released,
|
||||
result.prs_released,
|
||||
result.repos_synced
|
||||
);
|
||||
```
|
||||
|
||||
### Purgatory Sync (Simplified)
|
||||
|
||||
```rust
|
||||
// After fetching OIDs from remote:
|
||||
|
||||
let new_oids: HashSet<String> = fetched_oids.into_iter().collect();
|
||||
|
||||
let result = process_newly_available_git_data(
|
||||
&source_repo_path,
|
||||
&new_oids,
|
||||
&database,
|
||||
local_relay.as_ref(),
|
||||
&purgatory,
|
||||
&git_data_path,
|
||||
).await;
|
||||
```
|
||||
|
||||
### Integration with Purgatory Sync Redesign
|
||||
|
||||
The purgatory sync redesign (see `purgatory-sync-redesign.md`) uses this unified function in its `sync_identifier_from_url` implementation:
|
||||
|
||||
```rust
|
||||
pub async fn sync_identifier_from_url<C: SyncContext>(
|
||||
ctx: &C,
|
||||
identifier: &str,
|
||||
url: &str,
|
||||
throttle_manager: &Arc<ThrottleManager>,
|
||||
) -> usize {
|
||||
// ... fetch OIDs from URL ...
|
||||
|
||||
let fetched_oids = ctx.fetch_oids(&target_repo, url, &needed_oids).await?;
|
||||
|
||||
if !fetched_oids.is_empty() {
|
||||
// Use unified processing
|
||||
let new_oids: HashSet<String> = fetched_oids.into_iter().collect();
|
||||
|
||||
let result = process_newly_available_git_data(
|
||||
&target_repo,
|
||||
&new_oids,
|
||||
ctx.database(),
|
||||
ctx.local_relay(),
|
||||
ctx.purgatory(),
|
||||
ctx.git_data_path(),
|
||||
).await;
|
||||
|
||||
// Result already handled purgatory removal, DB saves, etc.
|
||||
}
|
||||
|
||||
fetched_oids.len()
|
||||
}
|
||||
```
|
||||
|
||||
The `SyncContext` trait wraps this function in its `process_newly_available_git_data` method for testability.
|
||||
|
||||
## Benefits
|
||||
|
||||
1. **Single source of truth** - One function handles all post-git-data processing
|
||||
2. **Always fresh discovery** - Events discovered from purgatory at processing time
|
||||
3. **Consistent behavior** - Push and sync paths behave identically
|
||||
4. **Simpler callers** - Just pass repo_path + new_oids
|
||||
5. **Complete processing** - Handles all event types, all repo syncing, HEAD, DB, WebSocket, purgatory
|
||||
6. **PR sync parity** - PR events now synced in purgatory path (was missing)
|
||||
|
||||
## Code to Remove/Simplify
|
||||
|
||||
After implementing the unified function:
|
||||
|
||||
1. **Remove**: Most of `sync_state_git_data` in `src/purgatory/mod.rs`
|
||||
2. **Simplify**: Event handling in `handle_receive_pack` (replace ~100 lines with single call)
|
||||
3. **Internalize**: `sync_to_owner_repos` and `sync_pr_refs_to_tagged_owner_repos` become internal helpers
|
||||
|
||||
## Testing Strategy
|
||||
|
||||
### Unit Tests
|
||||
|
||||
1. `extract_identifier_from_repo_path` - Various path formats
|
||||
2. `extract_identifier_from_pr_event` - Various tag formats
|
||||
3. Event discovery logic with mock purgatory
|
||||
|
||||
### Integration Tests
|
||||
|
||||
1. Push triggers processing and releases state event
|
||||
2. Push triggers processing and releases PR event
|
||||
3. Purgatory sync triggers processing
|
||||
4. Multiple events for same identifier processed correctly
|
||||
5. Cross-repo sync works for both state and PR events
|
||||
|
||||
## Future Considerations
|
||||
|
||||
### Batch Processing
|
||||
|
||||
Currently processes events one at a time. Could batch database saves and WebSocket notifications for efficiency with many events.
|
||||
|
||||
### Partial Failures
|
||||
|
||||
Currently continues on errors and collects them in result. Could add retry logic or transaction semantics if needed.
|
||||
|
||||
### Metrics
|
||||
|
||||
Add Prometheus metrics for:
|
||||
- Events processed by type (state/PR)
|
||||
- Repos synced per processing call
|
||||
- Processing duration
|
||||
- Errors by type
|
||||
|
||||
## Related Documents
|
||||
|
||||
- [Purgatory Sync Redesign](purgatory-sync-redesign.md) - Uses this unified function for purgatory sync operations
|
||||
Reference in New Issue
Block a user