mirror of
https://github.com/vitorpamplona/amethyst.git
synced 2026-10-05 19:28:25 +00:00
perf(marmot): benchmark commit ingest, and the message path by group size
Two gaps, both of which turned out to be hiding something. **The message path ran on groups of one and two.** Latency is genuinely flat across group sizes — an application message is sealed under the sender's own ratchet and never touches the tree — so the head-to-head claim against MDK survives being parameterised. Allocation is not flat on the send side: 68.7 KB at zero members to 253.1 KB at 32, while the receive side moves 61.5 to 67.4 KB. The asymmetry is exact rather than mysterious. MlsGroupManager.encrypt calls persistGroup unconditionally, serialising the whole group state on every message sent; decrypt calls it only when the message was a Commit that advanced the epoch, so decrypting an application message persists nothing. Sending in a 32-member group therefore spends a full state serialisation to record what amounts to a generation-counter bump. The write is necessary — a sender generation reused after a crash is a nonce-reuse-class problem — but writing all of the state for it is heavier than the invariant needs. Recorded as a finding, not changed: send-path persistence is security-sensitive. **Commit ingest was measured by nobody**, here or in MDK, despite being the operation every member performs on every membership or settings change and the only one whose cost is meant to scale with the group. It grows 1 950 -> 2 367 -> 3 009 us from 1 to 32 members: 1.5x for a 32x bigger group, which is the log2 shape MLS predicts, since the UpdatePath carries one node per LEVEL of the tree. Allocation grows 5.7x to 1.4 MB, making it the largest single allocator in the suite and the row to watch on a phone. Two measurement lessons are written into the README rather than left implicit. An isolated --only=ingest_commit run reported no trend at all and put the one-member case slowest; that was JIT warm-up on the shared group builder, and the full-suite numbers are the trustworthy ones. And create_group/32 is now reported as a range (77.5 - 195.7 ms) instead of a figure: five runs produced four within 11% and one more than twice the rest, which is what the fewest-iterations row on a shared vCPU looks like. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016kCuA6tc4JQzHPCDd39GHq
This commit is contained in:
+91
-9
@@ -16,12 +16,18 @@ cd <mdk> && cargo bench -p cgka-engine --bench group_lifecycle
|
||||
|
||||
## What is compared
|
||||
|
||||
| this module | MDK bench |
|
||||
|-----------------------|-----------------------------|
|
||||
| `create_group/N` | `bench_create_group` |
|
||||
| `join_welcome` | `bench_join_welcome` |
|
||||
| `send_app_message` | `bench_app_message_send` |
|
||||
| `ingest_app_message` | `bench_app_message_ingest` |
|
||||
| this module | MDK bench |
|
||||
|--------------------------|-----------------------------|
|
||||
| `create_group/N` | `bench_create_group` |
|
||||
| `join_welcome` | `bench_join_welcome` |
|
||||
| `send_app_message/N` | `bench_app_message_send` |
|
||||
| `ingest_app_message/N` | `bench_app_message_ingest` |
|
||||
| `ingest_commit/N` | (no counterpart) |
|
||||
|
||||
`ingest_commit` has no MDK counterpart because neither suite had one. It is the
|
||||
operation every member pays on every membership or settings change, and the
|
||||
only one whose cost is supposed to grow with the group, so leaving it
|
||||
unmeasured left the most load-bearing path in the protocol untested.
|
||||
|
||||
Both sides exclude transport crypto and run over in-memory storage, so what is
|
||||
measured is the engine's own CPU cost. Setup is outside the measured window on
|
||||
@@ -157,13 +163,20 @@ reproduces to four significant figures):
|
||||
|----------------------|------------|---------------|-----------------|--------------|
|
||||
| `create_group/1` | 3.61 ms | 16.93 ms | 5.25 - 5.85 ms | 1.5-1.6x slower |
|
||||
| `create_group/8` | 9.93 ms | 51.47 ms | 16.23 - 16.87 ms| 1.7x slower |
|
||||
| `create_group/32` | 31.64 ms | 190.08 ms | 82.49 - 86.31 ms| 2.6x slower |
|
||||
| `create_group/32` | 31.64 ms | 190.08 ms | 77.5 - 195.7 ms | 2.5x - 6x (see below) |
|
||||
| `join_welcome` | 4.77 ms | 6.22 ms | 1.89 - 1.93 ms | **2.5x faster** |
|
||||
| `send_app_message` | 4.28 ms | 1.72 ms | 0.63 - 0.68 ms | **6.5x faster** |
|
||||
| `ingest_app_message` | (n/a) | 3.11 ms | 0.90 - 0.95 ms | — |
|
||||
|
||||
`create_group` remains the weakest row, and `create_group/32` is still the
|
||||
noisiest in the suite — its two runs here disagree by nearly 2.4x at p99.
|
||||
`create_group` remains the weakest row, and `create_group/32` is not just the
|
||||
noisiest in the suite — it is the one number here that should not be quoted as
|
||||
a single figure at all. Across five post-rewrite runs on this host its p50 came
|
||||
out 77.5, 81.5, 82.5, 86.3 and 195.7 ms: four clustered within 11% of each
|
||||
other and one more than twice the rest. The row has the fewest iterations in
|
||||
the suite (each one has to build a 32-member group in setup) and this is a
|
||||
shared cloud vCPU, so a single noisy neighbour moves it in a way it cannot move
|
||||
the 300-iteration rows. Treat "roughly 2.5x MDK, occasionally much worse" as
|
||||
the honest reading, and re-run before believing any movement in it.
|
||||
|
||||
### What not publishing the founding commit was worth
|
||||
|
||||
@@ -204,3 +217,72 @@ non-canonical inputs such as `p` itself.
|
||||
|
||||
The RFC 7748 and RFC 8032 vector suites, HPKE, the MDK crypto-interop vectors
|
||||
and the full quartz + commons suites all pass unchanged.
|
||||
|
||||
## Group-size scaling: what the one-member benchmarks were hiding
|
||||
|
||||
The original `send_app_message` and `ingest_app_message` rows ran on groups of
|
||||
one and two members. That is the flattering case, and it hid a real asymmetry.
|
||||
|
||||
Numbers from one full-suite run:
|
||||
|
||||
| benchmark | p50 | alloc/op |
|
||||
|------------------------------|---------|------------|
|
||||
| `send_app_message/0 members` | 601 us | 68.7 KB |
|
||||
| `send_app_message/1 members` | 553 us | 73.8 KB |
|
||||
| `send_app_message/8 members` | 602 us | 117.1 KB |
|
||||
| `send_app_message/32 members`| 691 us | 253.1 KB |
|
||||
| `ingest_app_message/1` | 877 us | 61.5 KB |
|
||||
| `ingest_app_message/8` | 951 us | 65.1 KB |
|
||||
| `ingest_app_message/32` | 966 us | 67.4 KB |
|
||||
|
||||
**Latency is flat**, which is what MLS promises: an application message is
|
||||
sealed under the sender's own ratchet and never touches the tree, so the
|
||||
cryptography does not care how many members there are. The comparison against
|
||||
MDK's `send_app_message` therefore survives the parameterisation.
|
||||
|
||||
**Allocation is not flat on the send side** — 3.5x from 0 to 32 members, while
|
||||
the receive side barely moves. The cause is not subtle once looked at:
|
||||
|
||||
- `MlsGroupManager.encrypt` calls `persistGroup` **unconditionally**, so the
|
||||
whole group state, ratchet tree included, is serialised on every message
|
||||
sent.
|
||||
- `MlsGroupManager.decrypt` calls it **only** when the message was a Commit
|
||||
that advanced the epoch. Decrypting an application message persists nothing.
|
||||
|
||||
So a 32-member group allocates 252.8 KB to send a message and 67.1 KB to
|
||||
receive one, and the difference is a full state serialisation performed to
|
||||
record what amounts to a generation-counter bump. The write itself is
|
||||
necessary — a sender generation reused after a crash is a nonce-reuse-class
|
||||
problem — but writing the entire group state for it is heavier than the
|
||||
invariant requires. Left as a finding rather than a change: send-path
|
||||
persistence is security-sensitive and deserves its own decision, not a
|
||||
drive-by.
|
||||
|
||||
## `ingest_commit`: logarithmic in time, linear in allocation
|
||||
|
||||
Receiving someone else's Commit, from one full-suite run:
|
||||
|
||||
| benchmark | p50 | alloc/op |
|
||||
|----------------------------|----------|-------------|
|
||||
| `ingest_commit/1 members` | 1 950 us | 242.8 KB |
|
||||
| `ingest_commit/8 members` | 2 367 us | 523.0 KB |
|
||||
| `ingest_commit/32 members` | 3 009 us | 1 390.0 KB |
|
||||
|
||||
Latency grows, but far slower than the member count: 1.5x for a 32x bigger
|
||||
group. That is the shape MLS predicts. The UpdatePath a Commit carries has one
|
||||
node per LEVEL of the ratchet tree, so the receiver goes from roughly one HPKE
|
||||
open at two members to roughly five at thirty-three — a log2 curve, not a
|
||||
linear one.
|
||||
|
||||
Allocation grows 5.7x, tracking the size of the tree being parsed, rebuilt and
|
||||
persisted rather than the number of curve operations. So `ingest_commit` is
|
||||
mostly an allocation story, and it is the row to watch on a phone: every member
|
||||
performs it on every membership or settings change, and at 1.4 MB it is by some
|
||||
distance the largest single allocator in the suite.
|
||||
|
||||
A caution about isolated runs of this row. A `--only=ingest_commit` run
|
||||
reported 3 995 / 2 905 / 3 182 us — no trend at all, and the one-member case
|
||||
SLOWEST. That was JIT warm-up: `ingest_commit/1` runs first and pays for
|
||||
compiling the shared group builder. The monotonic full-suite numbers above are
|
||||
the trustworthy ones, which is the general rule here — prefer a full run, and
|
||||
distrust whichever row happens to go first.
|
||||
|
||||
@@ -94,6 +94,64 @@ fun benchCreateGroup(invitees: Int): BenchResult =
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* A group with [members] invitees already joined, at epoch 1.
|
||||
*
|
||||
* Built once per benchmark rather than per iteration where the operation under
|
||||
* test does not consume it, because assembling a 32-member group costs more
|
||||
* than everything being measured.
|
||||
*/
|
||||
private suspend fun groupWithMembers(members: Int): Triple<Client, List<Client>, HexKey> {
|
||||
val alice = Client("alice")
|
||||
val groupId = newGroupId()
|
||||
alice.manager.createCurrentProfileGroup(groupId, listOf("wss://bench.invalid"), GroupProfileV1("bench", ""))
|
||||
val invitees = (0 until members).map { Client("member-$it") }
|
||||
if (invitees.isNotEmpty()) {
|
||||
val kps = invitees.map { it.manager.generateKeyPackageEvent(relays = emptyList()) }
|
||||
val (_, welcomes) = alice.manager.addMembers(groupId, kps, emptyList())
|
||||
welcomes.forEach { delivery ->
|
||||
invitees
|
||||
.first { it.signer.pubKey == delivery.recipientPubKey }
|
||||
.manager
|
||||
.ingest(delivery.giftWrapEvent)
|
||||
}
|
||||
}
|
||||
return Triple(alice, invitees, groupId)
|
||||
}
|
||||
|
||||
/**
|
||||
* `ingest_commit/N` — receiving someone else's Commit.
|
||||
*
|
||||
* Nothing in this suite measured this, and neither does MDK's. It is the one
|
||||
* operation every member pays on every membership or settings change, and the
|
||||
* one that genuinely scales with group size: the UpdatePath it carries has a
|
||||
* node per level of the ratchet tree, so the receiver's cost grows with
|
||||
* log2(N) HPKE opens on top of the tree bookkeeping.
|
||||
*
|
||||
* Setup produces a FRESH commit per iteration — the group is built once, then
|
||||
* Alice changes the profile each time — because ingesting the same commit
|
||||
* twice is a no-op and would measure the dedup path instead.
|
||||
*/
|
||||
fun benchIngestCommit(members: Int): BenchResult {
|
||||
val (alice, invitees, groupId) =
|
||||
runBlocking { groupWithMembers(members) }
|
||||
val bob = invitees.first()
|
||||
var round = 0
|
||||
return measure(
|
||||
name = "ingest_commit/$members members",
|
||||
iterations = if (members >= 32) 40 else 100,
|
||||
warmup = if (members >= 32) 10 else 30,
|
||||
setup = {
|
||||
runBlocking {
|
||||
round++
|
||||
alice.manager.setGroupProfile(groupId, "bench-$round", "", emptyList())
|
||||
}
|
||||
},
|
||||
) { commit ->
|
||||
runBlocking { bob.manager.ingest(commit.signedEvent) }
|
||||
}
|
||||
}
|
||||
|
||||
/** `join_welcome` — MDK's `bench_join_welcome`. The invitee's side of the add. */
|
||||
fun benchJoinWelcome(): BenchResult =
|
||||
measure(
|
||||
@@ -115,17 +173,28 @@ fun benchJoinWelcome(): BenchResult =
|
||||
runBlocking { bob.manager.ingest(wrap as GiftWrapEvent) }
|
||||
}
|
||||
|
||||
/** `send_app_message` — MDK's `bench_app_message_send`. Encrypt + persist. */
|
||||
fun benchSendAppMessage(): BenchResult =
|
||||
/**
|
||||
* `send_app_message/N` — MDK's `bench_app_message_send`. Encrypt + persist.
|
||||
*
|
||||
* Parameterised by group size, which the single-member version of this
|
||||
* benchmark hid. The MLS half is O(1) in the member count — an application
|
||||
* message is sealed under the sender's own ratchet and never touches the tree
|
||||
* — but `MlsGroupManager.encrypt` calls `persistGroup`, and THAT serialises
|
||||
* the whole group state, ratchet tree included, on every send. So the cost per
|
||||
* message has a term that grows with the group while the cryptography does
|
||||
* not, and a one-member number is the flattering one.
|
||||
*
|
||||
* A fresh group per iteration, as before: sending accumulates rows in the
|
||||
* message store, and reusing one group would measure that growth instead.
|
||||
*/
|
||||
fun benchSendAppMessage(members: Int): BenchResult =
|
||||
measure(
|
||||
name = "send_app_message",
|
||||
iterations = 300,
|
||||
warmup = 100,
|
||||
name = "send_app_message/$members members",
|
||||
iterations = if (members >= 32) 50 else 300,
|
||||
warmup = if (members >= 32) 15 else 100,
|
||||
setup = {
|
||||
runBlocking {
|
||||
val alice = Client("alice")
|
||||
val groupId = newGroupId()
|
||||
alice.manager.createCurrentProfileGroup(groupId, listOf("wss://bench.invalid"), GroupProfileV1("bench", ""))
|
||||
val (alice, _, groupId) = groupWithMembers(members)
|
||||
alice to groupId
|
||||
}
|
||||
},
|
||||
@@ -133,23 +202,22 @@ fun benchSendAppMessage(): BenchResult =
|
||||
runBlocking { alice.manager.buildTextMessage(groupId, PAYLOAD) }
|
||||
}
|
||||
|
||||
/** `ingest_app_message` — MDK's `bench_app_message_ingest`. Decrypt + persist. */
|
||||
fun benchIngestAppMessage(): BenchResult =
|
||||
/**
|
||||
* `ingest_app_message/N` — MDK's `bench_app_message_ingest`. Decrypt + persist.
|
||||
*
|
||||
* Parameterised for the same reason as [benchSendAppMessage]: decryption is
|
||||
* O(1) in the member count, but the receiver persists its group state too, and
|
||||
* that is not.
|
||||
*/
|
||||
fun benchIngestAppMessage(members: Int): BenchResult =
|
||||
measure(
|
||||
name = "ingest_app_message",
|
||||
iterations = 200,
|
||||
warmup = 60,
|
||||
name = "ingest_app_message/$members members",
|
||||
iterations = if (members >= 32) 50 else 200,
|
||||
warmup = if (members >= 32) 15 else 60,
|
||||
setup = {
|
||||
runBlocking {
|
||||
val alice = Client("alice")
|
||||
val bob = Client("bob")
|
||||
val groupId = newGroupId()
|
||||
alice.manager.createCurrentProfileGroup(groupId, listOf("wss://bench.invalid"), GroupProfileV1("bench", ""))
|
||||
val kp = bob.manager.generateKeyPackageEvent(relays = emptyList())
|
||||
// A founding add publishes no commit, so there is no echo for
|
||||
// Alice to re-ingest — the Welcome is the whole delivery.
|
||||
val (_, welcome) = alice.manager.addMember(groupId, kp, emptyList())
|
||||
bob.manager.ingest(welcome!!.giftWrapEvent)
|
||||
val (alice, invitees, groupId) = groupWithMembers(members)
|
||||
val bob = invitees.first()
|
||||
val sent = alice.manager.buildTextMessage(groupId, PAYLOAD, persistOwn = false)
|
||||
bob to sent.outbound.signedEvent
|
||||
}
|
||||
@@ -173,8 +241,11 @@ private val ALL: List<Pair<String, () -> BenchResult>> =
|
||||
// the founding-only baseline, so the rows line up for comparison.
|
||||
listOf(0, 1, 8, 32).forEach { n -> add("create_group/$n" to { benchCreateGroup(n) }) }
|
||||
add("join_welcome" to { benchJoinWelcome() })
|
||||
add("send_app_message" to { benchSendAppMessage() })
|
||||
add("ingest_app_message" to { benchIngestAppMessage() })
|
||||
// Group sizes on the message path, because its persistence cost scales
|
||||
// with the member count even though its cryptography does not.
|
||||
listOf(0, 1, 8, 32).forEach { n -> add("send_app_message/$n" to { benchSendAppMessage(n) }) }
|
||||
listOf(1, 8, 32).forEach { n -> add("ingest_app_message/$n" to { benchIngestAppMessage(n) }) }
|
||||
listOf(1, 8, 32).forEach { n -> add("ingest_commit/$n" to { benchIngestCommit(n) }) }
|
||||
addAll(primitiveBenchmarks())
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user