Adds `marmotBench` — the quartz half of a head-to-head against MDK's
`cgka-engine --bench group_lifecycle`, case for case: create_group/N,
join_welcome, send_app_message, ingest_app_message. Both sides exclude
transport crypto, run over in-memory storage, and keep setup outside the
measured window, so what is compared is the engine's own CPU cost.
Every row also reports BYTES ALLOCATED PER OPERATION, from the JDK's own
`ThreadMXBean` — no dependency added. Latency alone cannot answer "are we
avoiding GC": a JVM can win a microbenchmark and still hand the user a
dropped frame later. The counter is per-thread, so benchmark bodies run
inline via `runBlocking` rather than on a dispatcher, where the allocation
would go uncounted.
First results, same host, nothing else running:
operation MDK quartz ratio
create_group/1 3.61 ms 16.93 ms 4.7x slower 27 MB/op
create_group/8 9.93 ms 51.47 ms 5.2x slower 95 MB/op
create_group/32 31.64 ms 190.08 ms 6.0x slower 333 MB/op
join_welcome 4.77 ms 6.22 ms 1.3x slower 10 MB/op
send_app_message 4.28 ms 1.72 ms 2.5x FASTER 2.8 MB/op
So the steady-state path a user actually exercises — sending a message — is
already faster than the reference. The gap is concentrated in key-agreement
work, and a JFR allocation profile says exactly where: 93% of all allocation
samples are `long[]` from `Curve25519Field.mul/add/sub`, which return a
freshly allocated field element on every single field operation inside
255-iteration scalar-multiplication loops. That one shape explains both the
5x latency gap and the MB-per-op allocation.
The fix (in-place field ops over caller-supplied scratch) is left as its own
change so it can be verified against the RFC vectors on its own merits.
Note: MDK's `ingest_app_message` has no number here. Their bench binary
panics in `bench_deferred_outbound_preflight_matrix` (an assertion on peeler
attempts) before reaching it, and criterion's filter does not skip that
bench's fixture construction. Reporting the gap rather than inventing a
comparison.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016kCuA6tc4JQzHPCDd39GHq
marmotBench — quartz Marmot vs MDK, head to head
Measures the same Marmot/MLS operations on both engines so "are we as fast as the reference" has an answer instead of an opinion.
The quartz half lives here. The MDK half is its own criterion suite:
# quartz
./gradlew :marmotBench:run # table
./gradlew :marmotBench:run --args=--json # machine-readable
# MDK (the reference)
cd <mdk> && cargo bench -p cgka-engine --bench group_lifecycle
What is compared
| this module | MDK bench |
|---|---|
create_group/N |
bench_create_group |
join_welcome |
bench_join_welcome |
send_app_message |
bench_app_message_send |
ingest_app_message |
bench_app_message_ingest |
Both sides exclude transport crypto and run over in-memory storage, so what is
measured is the engine's own CPU cost. Setup is outside the measured window on
both sides — criterion's iter_batched(.., PerIteration) there, an explicit
setup lambda here.
One shape difference, deliberately not hidden: MDK folds invitees into the
founding group (FoundingGroupCreated), while we create at epoch 0 and add in
a second commit to epoch 1. create_group/N therefore includes one more commit
on our side. That is a real cost, and averaging it away would be the wrong kind
of favourable.
Why allocation is reported next to latency
The brief is "as fast if not faster, while avoiding GC as much as possible",
and those are two different measurements. A JVM can win a microbenchmark while
allocating tens of times more per operation; the bill arrives later as GC
pauses on a phone, in a frame the benchmark never renders. So every row carries
bytes allocated per operation from com.sun.management.ThreadMXBean
(in the JDK — no dependency), alongside p50/p90/p99.
That counter is per-thread, which is why every benchmark body runs inline on
the harness thread via runBlocking. Work dispatched elsewhere would allocate
off-book and read as free.
The JVM runs with -XX:+UseSerialGC on a fixed 4g heap: the point is to keep
allocation attributable to the benchmark thread and to keep a collection from
landing inside a measured sample and corrupting the percentile it falls in.
Reading the numbers honestly
- Rust has no GC, so
alloc/ophas no MDK counterpart. It is not a head-to-head column — it is our own regression signal, and the number to drive down. - Latency across a JVM and a Rust binary on the same host is a fair comparison of this workload on this machine. It is not a language benchmark.
create_group/32builds 32 KeyPackages in setup. That cost is excluded, but it makes each iteration expensive to prepare — hence the low iteration count.