Files
amethyst/tools/search-parity
Claude 3b20aede9a test(search): pin the local search engine against a live relay's answers
Adds a recorded corpus of 64 real events from search-staging.brainstorm.world — a
vespa-relay, the same software this token language was ported from — and 18 tests
over it: 8 on the engine in :commons, 10 on `LocalCache.filter` in :amethyst.

Recorded, not fetched. `tools/search-parity/fetch_fixtures.py` drives `amy fetch`
against the relay by hand; the tests read the committed fixture, so `./gradlew
test` stays offline and the pre-push hook does not depend on somebody else's
uptime. The fixture stores each case's filter FIELDS rather than a prebuilt
filter, so the test rebuilds the Filter in view of the reader and a wrong rebuild
cannot quietly make the assertions vacuous.

What the relay can and cannot referee turned out to be the whole design, and it
was measured rather than assumed:

- **NIP-01 it can.** Given kinds, #t, since and until there is exactly one right
  answer, and across all 64 events our matcher agrees with the relay about every
  one it chose to return — 0 violations. That is now a hard assertion, with a
  converse test so it cannot pass by matching everything.
- **NIP-50 it cannot.** This relay retrieves topically: asked for `bitcoin` it
  returns a block-height summary that never says "bitcoin". Eight of 64 events
  carry no literal occurrence of the term that fetched them. Asserting our
  substring matcher reproduces that would encode someone else's semantic
  expansion as a requirement on a lexical one — a test that fails on correct
  code. So text results are deliberately not compared, and the divergence is
  pinned as a range instead: zero would mean the relay turned lexical and the
  comparison should be rewritten, a quarter would mean we regressed.

The fixture uses the `include:spam` lens, which waives the web-of-trust gate.
Also measured: it makes the corpus reproducible, where `observer:<pubkey>` ties
every answer to one account's moving trust graph — but it does not make retrieval
lexical, and in fact widens the divergence from 4 events to 8 by letting more
topical matches through.

The LocalCache tests cover what the relay knows nothing about and where the bugs
actually were: the regular/addressable split, the viewer-policy predicate
composing with rather than replacing the filter, and the result cap keeping the
newest — the ordering whose absence let `take(limit)` run before the sort.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017yKjw2WqwZpSzsqcYZMnkV
2026-09-08 15:42:35 +00:00
..