Files
amethyst/commons
Claude b79217b09c perf(commons): halve the list observables' write cost, and measure it
The copy-on-write rewrite of the two list observables was justified on
portability -- `java.util.concurrent` has no KMP equivalent, which is what
kept them in `jvmAndroid` behind an iOS stub. That says nothing about
speed, and these are hot: every consumed event is offered to each observer
whose filter could match it, from each relay's socket coroutine.

So measure it. `ObserverListBenchmark` keeps the `ConcurrentSkipListSet`
implementation verbatim as the baseline and as a differential oracle --
`bothImplementationsAgree` asserts the two emit identical lists for
identical input at limits null/50/400, which is the property the rewrite
had to preserve. Runs in ~7s, asserts on correctness only, never on wall
time.

The first run said the rewrite was a regression: slower on the limited
insert path and 4.9x slower on 8 threads, with no thread scaling. The
cause was one line, not the design. `State.plus` published `ids + added`,
and over the limit `ids + added - dropped`; each operator allocates a full
copy of the set, so the eviction path rebuilt membership twice per insert.
One `HashSet`, sized up front and mutated before publishing, is the fix.

After it, against the implementation it replaced: re-delivering an already
listed note 3.3-4.1x faster, insert into a populated unlimited list (what
nearly every observer registers) 1.3-1.5x faster and widening with n, cold
fill at n=5000 1.8x faster, concurrent inserts at parity on one thread and
2.1-2.7x slower on 4-8.

That last number is a real cost and is written on both classes rather than
left out: threads serialize on one reference and a lost CAS discards its
copy, where the skip list striped across keys and scaled. Two things bound
it -- the benchmark's threads do nothing but insert, while real ingest
spends most of its per-event budget verifying signatures and parsing
before it reaches an observer; and nearly every production observer
registers with no limit, which is the shape copy-on-write wins.

Also recorded, since it framed the earlier discussion wrong: lock-free was
never the difference between the two. The skip-list version was lock-free
too -- ConcurrentHashMap stripes per key, so `compute` holds a bin, not a
monitor. The choice was portability and speed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KkULS5SVq4GHDdoCzajKi8
2026-09-21 18:42:13 +00:00
..