mirror of
https://github.com/vitorpamplona/amethyst.git
synced 2026-10-06 03:38:23 +00:00
Adds `marmotBench` — the quartz half of a head-to-head against MDK's
`cgka-engine --bench group_lifecycle`, case for case: create_group/N,
join_welcome, send_app_message, ingest_app_message. Both sides exclude
transport crypto, run over in-memory storage, and keep setup outside the
measured window, so what is compared is the engine's own CPU cost.
Every row also reports BYTES ALLOCATED PER OPERATION, from the JDK's own
`ThreadMXBean` — no dependency added. Latency alone cannot answer "are we
avoiding GC": a JVM can win a microbenchmark and still hand the user a
dropped frame later. The counter is per-thread, so benchmark bodies run
inline via `runBlocking` rather than on a dispatcher, where the allocation
would go uncounted.
First results, same host, nothing else running:
operation MDK quartz ratio
create_group/1 3.61 ms 16.93 ms 4.7x slower 27 MB/op
create_group/8 9.93 ms 51.47 ms 5.2x slower 95 MB/op
create_group/32 31.64 ms 190.08 ms 6.0x slower 333 MB/op
join_welcome 4.77 ms 6.22 ms 1.3x slower 10 MB/op
send_app_message 4.28 ms 1.72 ms 2.5x FASTER 2.8 MB/op
So the steady-state path a user actually exercises — sending a message — is
already faster than the reference. The gap is concentrated in key-agreement
work, and a JFR allocation profile says exactly where: 93% of all allocation
samples are `long[]` from `Curve25519Field.mul/add/sub`, which return a
freshly allocated field element on every single field operation inside
255-iteration scalar-multiplication loops. That one shape explains both the
5x latency gap and the MB-per-op allocation.
The fix (in-place field ops over caller-supplied scratch) is left as its own
change so it can be verified against the RFC vectors on its own merits.
Note: MDK's `ingest_app_message` has no number here. Their bench binary
panics in `bench_deferred_outbound_preflight_matrix` (an assertion on peeler
attempts) before reaching it, and criterion's filter does not skip that
bench's fixture construction. Reporting the gap rather than inventing a
comparison.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016kCuA6tc4JQzHPCDd39GHq
63 lines
2.8 KiB
Markdown
63 lines
2.8 KiB
Markdown
# marmotBench — quartz Marmot vs MDK, head to head
|
|
|
|
Measures the same Marmot/MLS operations on both engines so "are we as fast as
|
|
the reference" has an answer instead of an opinion.
|
|
|
|
The quartz half lives here. The MDK half is its own criterion suite:
|
|
|
|
```sh
|
|
# quartz
|
|
./gradlew :marmotBench:run # table
|
|
./gradlew :marmotBench:run --args=--json # machine-readable
|
|
|
|
# MDK (the reference)
|
|
cd <mdk> && cargo bench -p cgka-engine --bench group_lifecycle
|
|
```
|
|
|
|
## What is compared
|
|
|
|
| this module | MDK bench |
|
|
|-----------------------|-----------------------------|
|
|
| `create_group/N` | `bench_create_group` |
|
|
| `join_welcome` | `bench_join_welcome` |
|
|
| `send_app_message` | `bench_app_message_send` |
|
|
| `ingest_app_message` | `bench_app_message_ingest` |
|
|
|
|
Both sides exclude transport crypto and run over in-memory storage, so what is
|
|
measured is the engine's own CPU cost. Setup is outside the measured window on
|
|
both sides — criterion's `iter_batched(.., PerIteration)` there, an explicit
|
|
`setup` lambda here.
|
|
|
|
**One shape difference, deliberately not hidden:** MDK folds invitees into the
|
|
founding group (`FoundingGroupCreated`), while we create at epoch 0 and add in
|
|
a second commit to epoch 1. `create_group/N` therefore includes one more commit
|
|
on our side. That is a real cost, and averaging it away would be the wrong kind
|
|
of favourable.
|
|
|
|
## Why allocation is reported next to latency
|
|
|
|
The brief is "as fast if not faster, while avoiding GC as much as possible",
|
|
and those are two different measurements. A JVM can win a microbenchmark while
|
|
allocating tens of times more per operation; the bill arrives later as GC
|
|
pauses on a phone, in a frame the benchmark never renders. So every row carries
|
|
**bytes allocated per operation** from `com.sun.management.ThreadMXBean`
|
|
(in the JDK — no dependency), alongside p50/p90/p99.
|
|
|
|
That counter is per-thread, which is why every benchmark body runs inline on
|
|
the harness thread via `runBlocking`. Work dispatched elsewhere would allocate
|
|
off-book and read as free.
|
|
|
|
The JVM runs with `-XX:+UseSerialGC` on a fixed 4g heap: the point is to keep
|
|
allocation attributable to the benchmark thread and to keep a collection from
|
|
landing inside a measured sample and corrupting the percentile it falls in.
|
|
|
|
## Reading the numbers honestly
|
|
|
|
- Rust has no GC, so `alloc/op` has no MDK counterpart. It is not a
|
|
head-to-head column — it is our own regression signal, and the number to
|
|
drive down.
|
|
- Latency across a JVM and a Rust binary on the same host is a fair comparison
|
|
of *this* workload on *this* machine. It is not a language benchmark.
|
|
- `create_group/32` builds 32 KeyPackages in setup. That cost is excluded, but
|
|
it makes each iteration expensive to prepare — hence the low iteration count.
|