Files
ngit-grasp/CHANGELOG.md
T
DanConwayDev 4b689a6515 fix(sync): checkpoint recovery state crash-safely
Purgatory and rejected-event snapshots were consumed on successful startup and only rewritten during graceful shutdown. A crash or power loss after restore therefore discarded all durable recovery history, allowing unresolved work and rejection backoff to restart from an empty state.

Retain restored checkpoints, replace them atomically every 60 seconds, and write a final snapshot only after background mutation tasks have stopped. Snapshot replacement syncs a complete temporary file, renames it, and syncs the parent directory so interruption leaves either the old or new complete state.

The checkpoint interval bounds mutation loss without adding synchronous disk I/O to event handling. Addressable purgatory restoration is idempotent and timestamp adjustment continues to preserve remaining lifetimes. Database durability, git repository state, and general LMDB backup policy are deliberately excluded.

Validation: cargo test --lib passed all 668 tests before commit. Purgatory and rejected-index tests verify the checkpoint survives restore and can initialize a second fresh instance; the atomic writer test verifies complete replacement. Nix validation follows after rebasing onto the passive-ownership proposal.
2026-08-08 12:10:16 +00:00

20 KiB
Raw Blame History

Changelog

All notable changes to this project will be documented in this file.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

Unreleased

Fixed

  • Preserve purgatory and rejected-event recovery across abrupt termination. Their snapshots now remain after restore and are atomically replaced every 60 seconds, bounding crash loss to one interval instead of consuming the only durable copy at startup.
  • Treat cumulative retained-subscription byte refusals as capacity signals, not temporary query-rate episodes. The sync client learns the disclosed cap, rebuilds persistent coverage within it while reserving one maximum transient REQ, serializes transient work against that reserve, and covers overflow with paced five-minute history plus overlap instead of retrying the same impossible live set.
  • Replace relay ownership when an addressable repository announcement changes, rather than retaining every relay URL ever listed. Fully expired StateOnly purgatory targets are pruned while soft-expired and promoted repositories are preserved. Healthy connections and live requests are left undisturbed; naturally closed subscriptions rebuild from current ownership, and naturally ended connections omit obsolete items when deciding whether to reconnect.
  • Raise retained subscription state per connection from 1 MiB to 5 MiB. A production 34-filter repository-sync live set reached roughly 1.2 MiB, so rust-nostr's newly introduced default repeatedly closed part of persistent coverage. The replacement remains a finite per-connection bound and matches the relay's maximum admitted WebSocket message size.

2.1.1 - 2026-08-08

ngit-grasp 2.1.1 is a patch release improving relay-to-relay sync stability under newly introduced query-rate limits. It handles explicit query refusals, paces non-urgent background work per relay, and temporarily relaxes the embedded relay's over-restrictive rust-nostr 0.45 query allowance while the upstream NIP-77 accounting is reviewed.

Fixed

  • Temporarily override rust-nostr 0.45's newly introduced LocalRelay query allowance from 120 to 1,200 messages per minute per connection. This retains a finite DoS backstop while avoiding an unusually restrictive default that also charges every SDK-managed NIP-77 NEG-MSG continuation. Re-evaluate the 10× allowance after upstream revises that NIP-77 accounting.
  • Proactively space non-urgent historic sync, pagination, hydration, retry, and dependency query starts at one per second per relay connection. Live subscriptions bypass the background gate, while the existing adaptive refusal handling remains the backstop for tighter or internally generated traffic.
  • Pace query starts after a relay reports its per-connection query-rate budget is exhausted, preventing fixed cooldown recovery from replaying the same fast historic-sync burst indefinitely. If an SDK-managed NIP-77 exchange exhausts that budget, use paced REQs for the rest of the connection session because the application cannot pace individual NEG-MSG frames. This is a reactive compatibility backstop for an unadvertised limit, layered over the proactive background pacing while leaving live coverage prioritised.

2.1.0 - 2026-08-07

ngit-grasp 2.1.0 substantially improves proactive repository-event sync reliability, compatibility, resource accounting, and recovery without adding new sync features. It also updates rust-nostr to 0.45.0 and makes the relay's defensive limits explicit and operator-configurable.

Added

  • Expanded grasp-audit with stable machine-readable results, full JSON reports, explicit audit identities, and hardened discovered-server probes.
  • Added bounded request-class labels to transient-sync watchdog logs and metrics so operators can identify which historic recovery path is stalling.

Changed

  • Upgraded the rust-nostr crates from 0.45.0-alpha.8 to the stable 0.45.0 release, including the upstream NEG-OPEN handling fix and new local-relay resource hardening. The embedded relay now imposes 500 active REQs per connection; per-minute connection quotas of 60 event writes, 120 queries, 30 authentication events, and 6,000 WebSocket messages (raised from rust-nostr's 300/minute default after it disconnected legitimate bursty clients in production); 20 filters per REQ; 500 results per filter; 250-byte subscription IDs; 1 MiB retained subscription state; 10 active negentropy sessions and 50,000 negentropy items per connection; 5 MiB WebSocket messages; and a 10-second handshake deadline. ngit-grasp deliberately:

    • makes the discoverable subscription and filter-result limits configurable and advertises them in NIP-11;
    • raises rust-nostr's new 64 KiB event bound to a configurable 192 KiB default because valid NIP-34 patch events in production reach about 149 KiB; and
    • overrides rust-nostr's new 128-connection default to preserve documented unbounded-by-default inbound admission, while keeping configured caps exact.

Security

  • Restricted event-directed sync targets to globally reachable endpoints. Relay and clone URLs from untrusted repository announcements and PR events reached the proactive WebSocket and git fetch sinks after syntax-only checks; production logs showed sync dialling ws://localhost:3334, ws://127.0.0.1:7334, and a CGNAT address, so a crafted event could point a public relay at loopback or internal infrastructure (SSRF). A fail-closed outbound target policy now runs immediately before every event-directed connection and git fetch: per-sink scheme allowlists, no credentials, no local hostnames, and only globally reachable addresses for IP literals and all DNS answers. Git fetches additionally pin the vetted DNS answers and disable redirects, proxies, credential helpers, and alternate protocols. Service admission and self-fetch filtering now compare parsed host and port instead of substrings, so gitnostr.com.attacker.example no longer satisfies a check for gitnostr.com. The operator-configured bootstrap relay remains usable even when local; event URLs never inherit that exception. The new NGIT_SYNC_ALLOW_NON_GLOBAL_TARGETS flag (default false) disables the reachability checks for tests and closed development networks. Known limitation: relay-connection DNS is re-validated before every dial but cannot be pinned through nostr-sdk's connector, leaving a narrow DNS-rebinding window that needs upstream connector support.

Fixed

  • Published the source revision in build metrics, NIP-11 metadata, and the landing page for Nix-built releases, where the filtered source archive does not contain Git metadata.
  • Stopped routine client connection resets without a WebSocket close handshake from being reported as server errors, while retaining other embedded-relay failures at their original severity.
  • Adapted historic pagination to each relay's observed page size and guarded NIP-11 default_limit hint, so omitted-limit filters continue past relays that return small pages without trusting unrelated maximum-limit metadata.
  • Batched unresolved purgatory dependency IDs into one query per relay on the existing 30-second cadence, reducing subscription and rate-limit pressure without delaying newly available repository data.
  • Prevented naughty-listed and self-referential relay targets from being rescheduled while unavailable or unnecessary.
  • Kept NIP-77 reconciliation enabled when historic hydration leaves a small recoverable residual, and stopped repeatedly requesting Git object IDs that a remote has already advertised as missing.
  • Stopped structurally malformed synced events from being downloaded and revalidated on every historic pass.
  • Kept proactive sync within each relay connection's active-subscription allowance. Live subscriptions, negentropy rounds, historic and fallback queries, pagination verification, and purgatory dependency polling now share one per-session budget derived from NIP-11 with a conservative fallback. Live coverage is consolidated first, capacity is reserved for control work, and timed-out queries are closed relay-side before their slots are reused.
  • Fixed historic sync silently abandoning events a relay identified but failed to deliver. Negentropy reconciliation finds event IDs missing locally, but a relay's exact-ID response can return only a subset; for batches without repository/root-event metadata (the Layer 1 announcements batch) no semantic REQ+EOSE fallback exists, and production logs showed such batches completing "with partial results" — dropping the missing IDs until the next daily sync up to 25 hours later. The still-missing IDs are now kept in a per-relay recovery index and refetched over the existing connection with bounded exponential backoff (30s doubling to 15min, one in-flight attempt per relay, 300 IDs per fetch). Progress clears only the recovered IDs and resets the backoff; duplicate incomplete responses merge without duplicating work; IDs satisfied by live sync are cleared without consuming attempts; 12 consecutive zero-progress attempts expire the pending IDs explicitly, and the relay stays observably degraded (ConnectedHistoricSyncFailures) until the daily sync. Startup remains non-blocking throughout, and a relay whose missing events are all recovered — with no unrelated batch failures — is now promoted back to Connected instead of reporting failures forever. Also fixed the retry-subscription-failure path confirming an incomplete batch as successful, and bounded the previously unbounded missing_ids log arrays to a five-ID sample.
  • Fixed malformed client messages tearing down the whole WebSocket connection. A single unparseable message - in production, requests carrying invalid event IDs that fail with Invalid input length 64 - closed the session, forcing clients to reconnect and re-subscribe. Invalid messages are now answered with a NOTICE and the connection stays open. Fixed upstream in rust-nostr and picked up by upgrading to the stable 0.45.0 release.
  • Fixed proactive sync losing repository events when public relays cap active subscriptions. Compatible GRASP filters now share bounded NIP-01 REQs while retaining per-filter history pagination.
  • Fixed repeated relay rate-limit notices extending the cooldown indefinitely. Notices during an active cooldown keep its original deadline, while a new rejection after recovery begins a fresh cooldown.
  • Removed superseded same-author repository states from purgatory after their replacement is promoted when the locally available Git data cannot reconstruct them. Reconstructable rollback states and other maintainers' states are retained.

2.0.0 - 2026-07-27

Breaking changes

  • Removed --relay-owner-nsec because command-line secrets are exposed through process listings and service diagnostics. Use the relay_owner_nsec systemd credential, NGIT_RELAY_OWNER_NSEC, or .relay-owner.nsec. The NixOS relayOwnerNsecFile option remains supported and now supplies a protected systemd credential.
  • Removed the hidden repair-deletion-requests command and the public ngit_grasp::repair_deletion_requests module. Deletion-request lifecycle reconciliation now runs automatically.
  • Added four deletion-request retention fields to the public Config struct. Rust consumers that construct Config with a struct literal must provide them.

Added

  • Added configurable bounded retention for NIP-09 deletion requests and NIP-62 request-to-vanish events, with cleanup telemetry for operators.

Changed

  • Deletion requests now move through a bounded served and gating lifecycle based on whether they affected accepted events. Retention catch-up runs in the background after the relay and synchronization workers start.
  • Deletion-disrespector mode now applies to both NIP-09 deletion requests and NIP-62 request-to-vanish events.

Fixed

  • Fixed maintainer invitation acceptance and synchronization across owner-only, invitee-only, shared, and temporarily unavailable GRASP servers. Acceptance now converges without an invitee state event or additional Git push, handles an invitee repository that already exists, and preserves the one-way authority granted by an invitation before reciprocal acceptance.
  • Kept invitation recovery responsive during restarts, rate limits, large relay sets, and expired dependency caches by bounding actor work, retaining exact relay hints, and keeping connection and subscription retries scheduler-owned.
  • Fixed promoted repositories remaining on state-only GRASP sync filters, which prevented remote issues, patches, and pull requests from being discovered after their Git data arrived.
  • Delayed the terminal smart-HTTP receive-pack flush until matching repository events are promoted and queryable. Sideband clients receive progress during long Git processing and post-push repository alignment.
  • Batched the one-time deletion-request lifecycle migration so large databases do not remain unavailable while historical requests are reconciled.
  • Applied NIP-01's lowest-event-ID tie-break to same-second repository state replacements in purgatory.

Security

  • Kept relay-owner private keys out of process arguments by loading NixOS secret files through a protected systemd credential. Empty or invalid configured keys now stop startup instead of rotating identity, generated fallback keys use mode 0600, and existing fallback files are restricted before being read.

1.2.0 - 2026-07-06

Fixed

  • Prevented redundant NIP-09 deletion requests from polluting relay storage and query results.

Added

  • Implemented repository lifecycle handling for removals caused by NIP-09 deletion requests, NIP-62 requests-to-vanish, repository blacklist matches, and removal from whitelist. Removed repository scopes now archive git data and move data through holding/recovery flows with a 90-day default retention period, cascade-delete related events that lose their accepted-reference path, roll back deleted state-event versions where possible, and keep served nostr state aligned with git refs. This is an enabler for moderation features.

Changed

  • Stream git smart HTTP responses instead of buffering full responses before sending them to clients, improving behavior for long-running fetch and push operations.

  • Upgraded dependencies: rust-nostr to 0.45.0-alpha.3 from a patched version to enable publishing ngit-grasp to crates.io, a Rust toolchain bump via Nix flake upgrade, and other Rust dependencies.

Fixed

  • Wrapped receive-pack error pkt-lines in sideband framing so git clients receive push rejection messages correctly during smart HTTP pushes.

1.1.0 - 2026-05-22

Added

  • GRASP-06 contributor PR submission endpoint (NGIT_GRASP06_ENABLE, default off). When enabled, the relay accepts unauthenticated git push of refs/nostr/<event-id> to /prs/<npub>/<identifier>.git from any contributor, even for repositories this relay has no accepted announcement for. The corresponding PR (kind 1618) or PR Update (kind 1619) event is accepted into purgatory when its clone tag names this relay's /prs/<signer>/<d>.git endpoint and its a tag's d-tag matches the URL identifier. When the event and the push match (signer, d-tag, c-tag commit) the event is released from purgatory and the ref is mirrored into any accepted-announcement repos on this relay. Empty /prs/ repos (probe pushes, mismatched events) are garbage-collected inline at the three runtime sites that can leave them empty (receive handler at end of push, PR-event policy when discarding a mismatched scoped placeholder, purgatory sweep when a scoped placeholder expires without a matching event) plus a one-shot startup scan that removes any zero-ref /prs/ bare repos left behind by a previous run (crash, mid-cleanup failure, or shutdown with unresolved scoped placeholders). GRASP-06 is advertised in NIP-11 supported_grasps when enabled. See how-to/enable-grasp-06.md and explanation/grasp-06-contributor-pr-submission.md.

Fixed

  • Handle HEAD requests for info/refs endpoints (previously returned 405).

1.0.2 - 2026-04-10

Fixed

  • Replacement announcements (kind 30617) for a purgatory entry were being saved to the database immediately, bypassing the purgatory gate. When a second copy of the same announcement arrived (e.g. via sync from another relay) while the original was still in purgatory awaiting git data, the policy returned Accept instead of AcceptPurgatory, causing the event to be stored without the corresponding git data or state events ever arriving. The fix returns AcceptPurgatory for replacements of purgatory entries so the updated event is held in purgatory until git data arrives.

  • Repository identifiers containing characters that require percent-encoding in URLs (e.g. spaces, emoji) are now accepted and served correctly. NIP-01 places no restriction on d tag values and NIP-34 only recommends kebab-case without mandating it, so rejecting non-kebab identifiers was overly strict. Identifiers are stored verbatim on disk and percent-encoded when used in URLs, per the nostr:// clone URL spec formalised in NIP-34 PR #2312 and the GRASP-01 HTTP path spec. The landing page clone URL now also correctly percent-encodes the identifier.

  • --git-dir is now passed as a global git option (before the subcommand) in check_repo_empty, fixing compatibility with git versions that require global options to precede the subcommand.

Changed

  • Remove arbitrary default max connections limit; when NGIT_MAX_CONNECTIONS is unset the relay imposes no connection cap, deferring to OS fd limits and infrastructure controls

  • Added cleanup-empty-repos subcommand to remove stale events for empty git repositories

1.0.1 - 2026-02-27

Fixed

  • Push authorization now correctly ignores refs/tags/<name>^{} peeled-tag entries in state events (kind 30618). These entries are git's internal notation for the dereferenced commit behind an annotated tag and are never sent as part of a push. Previously, their presence in the state event caused can_satisfy_state to reject valid annotated-tag pushes because the would-be ref state after the push did not include the spurious ^{} entry, making the exact-equality check fail.

Changed

  • Push auth rejections now send the reason to the git client via ERR pkt-line (e.g. "authorisation failed: No state events in purgatory") instead of a generic HTTP 403, so users see actionable error messages directly in their terminal

1.0.0 - 2026-02-26

Initial release of ngit-grasp, a GRASP relay implementation in Rust.