Files
amethyst/marmotBench
Claude d03cc445f1 perf(marmot): head-to-head benchmark against MDK, and where our allocation goes
Adds `marmotBench` — the quartz half of a head-to-head against MDK's
`cgka-engine --bench group_lifecycle`, case for case: create_group/N,
join_welcome, send_app_message, ingest_app_message. Both sides exclude
transport crypto, run over in-memory storage, and keep setup outside the
measured window, so what is compared is the engine's own CPU cost.

Every row also reports BYTES ALLOCATED PER OPERATION, from the JDK's own
`ThreadMXBean` — no dependency added. Latency alone cannot answer "are we
avoiding GC": a JVM can win a microbenchmark and still hand the user a
dropped frame later. The counter is per-thread, so benchmark bodies run
inline via `runBlocking` rather than on a dispatcher, where the allocation
would go uncounted.

First results, same host, nothing else running:

    operation              MDK        quartz      ratio
    create_group/1      3.61 ms     16.93 ms     4.7x slower   27 MB/op
    create_group/8      9.93 ms     51.47 ms     5.2x slower   95 MB/op
    create_group/32    31.64 ms    190.08 ms     6.0x slower  333 MB/op
    join_welcome        4.77 ms      6.22 ms     1.3x slower   10 MB/op
    send_app_message    4.28 ms      1.72 ms     2.5x FASTER  2.8 MB/op

So the steady-state path a user actually exercises — sending a message — is
already faster than the reference. The gap is concentrated in key-agreement
work, and a JFR allocation profile says exactly where: 93% of all allocation
samples are `long[]` from `Curve25519Field.mul/add/sub`, which return a
freshly allocated field element on every single field operation inside
255-iteration scalar-multiplication loops. That one shape explains both the
5x latency gap and the MB-per-op allocation.

The fix (in-place field ops over caller-supplied scratch) is left as its own
change so it can be verified against the RFC vectors on its own merits.

Note: MDK's `ingest_app_message` has no number here. Their bench binary
panics in `bench_deferred_outbound_preflight_matrix` (an assertion on peeler
attempts) before reaching it, and criterion's filter does not skip that
bench's fixture construction. Reporting the gap rather than inventing a
comparison.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016kCuA6tc4JQzHPCDd39GHq
2026-09-10 13:22:43 +00:00
..

marmotBench — quartz Marmot vs MDK, head to head

Measures the same Marmot/MLS operations on both engines so "are we as fast as the reference" has an answer instead of an opinion.

The quartz half lives here. The MDK half is its own criterion suite:

# quartz
./gradlew :marmotBench:run                     # table
./gradlew :marmotBench:run --args=--json       # machine-readable

# MDK (the reference)
cd <mdk> && cargo bench -p cgka-engine --bench group_lifecycle

What is compared

this module MDK bench
create_group/N bench_create_group
join_welcome bench_join_welcome
send_app_message bench_app_message_send
ingest_app_message bench_app_message_ingest

Both sides exclude transport crypto and run over in-memory storage, so what is measured is the engine's own CPU cost. Setup is outside the measured window on both sides — criterion's iter_batched(.., PerIteration) there, an explicit setup lambda here.

One shape difference, deliberately not hidden: MDK folds invitees into the founding group (FoundingGroupCreated), while we create at epoch 0 and add in a second commit to epoch 1. create_group/N therefore includes one more commit on our side. That is a real cost, and averaging it away would be the wrong kind of favourable.

Why allocation is reported next to latency

The brief is "as fast if not faster, while avoiding GC as much as possible", and those are two different measurements. A JVM can win a microbenchmark while allocating tens of times more per operation; the bill arrives later as GC pauses on a phone, in a frame the benchmark never renders. So every row carries bytes allocated per operation from com.sun.management.ThreadMXBean (in the JDK — no dependency), alongside p50/p90/p99.

That counter is per-thread, which is why every benchmark body runs inline on the harness thread via runBlocking. Work dispatched elsewhere would allocate off-book and read as free.

The JVM runs with -XX:+UseSerialGC on a fixed 4g heap: the point is to keep allocation attributable to the benchmark thread and to keep a collection from landing inside a measured sample and corrupting the percentile it falls in.

Reading the numbers honestly

  • Rust has no GC, so alloc/op has no MDK counterpart. It is not a head-to-head column — it is our own regression signal, and the number to drive down.
  • Latency across a JVM and a Rust binary on the same host is a fair comparison of this workload on this machine. It is not a language benchmark.
  • create_group/32 builds 32 KeyPackages in setup. That cost is excluded, but it makes each iteration expensive to prepare — hence the low iteration count.