perf(marmot): head-to-head benchmark against MDK, and where our allocation goes

Adds `marmotBench` — the quartz half of a head-to-head against MDK's
`cgka-engine --bench group_lifecycle`, case for case: create_group/N,
join_welcome, send_app_message, ingest_app_message. Both sides exclude
transport crypto, run over in-memory storage, and keep setup outside the
measured window, so what is compared is the engine's own CPU cost.

Every row also reports BYTES ALLOCATED PER OPERATION, from the JDK's own
`ThreadMXBean` — no dependency added. Latency alone cannot answer "are we
avoiding GC": a JVM can win a microbenchmark and still hand the user a
dropped frame later. The counter is per-thread, so benchmark bodies run
inline via `runBlocking` rather than on a dispatcher, where the allocation
would go uncounted.

First results, same host, nothing else running:

    operation              MDK        quartz      ratio
    create_group/1      3.61 ms     16.93 ms     4.7x slower   27 MB/op
    create_group/8      9.93 ms     51.47 ms     5.2x slower   95 MB/op
    create_group/32    31.64 ms    190.08 ms     6.0x slower  333 MB/op
    join_welcome        4.77 ms      6.22 ms     1.3x slower   10 MB/op
    send_app_message    4.28 ms      1.72 ms     2.5x FASTER  2.8 MB/op

So the steady-state path a user actually exercises — sending a message — is
already faster than the reference. The gap is concentrated in key-agreement
work, and a JFR allocation profile says exactly where: 93% of all allocation
samples are `long[]` from `Curve25519Field.mul/add/sub`, which return a
freshly allocated field element on every single field operation inside
255-iteration scalar-multiplication loops. That one shape explains both the
5x latency gap and the MB-per-op allocation.

The fix (in-place field ops over caller-supplied scratch) is left as its own
change so it can be verified against the RFC vectors on its own merits.

Note: MDK's `ingest_app_message` has no number here. Their bench binary
panics in `bench_deferred_outbound_preflight_matrix` (an assertion on peeler
attempts) before reaching it, and criterion's filter does not skip that
bench's fixture construction. Reporting the gap rather than inventing a
comparison.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016kCuA6tc4JQzHPCDd39GHq
This commit is contained in:
Claude
2026-09-10 13:22:43 +00:00
parent 5fc9aa9f58
commit d03cc445f1
8 changed files with 605 additions and 0 deletions
+62
View File
@@ -0,0 +1,62 @@
# marmotBench — quartz Marmot vs MDK, head to head
Measures the same Marmot/MLS operations on both engines so "are we as fast as
the reference" has an answer instead of an opinion.
The quartz half lives here. The MDK half is its own criterion suite:
```sh
# quartz
./gradlew :marmotBench:run # table
./gradlew :marmotBench:run --args=--json # machine-readable
# MDK (the reference)
cd <mdk> && cargo bench -p cgka-engine --bench group_lifecycle
```
## What is compared
| this module | MDK bench |
|-----------------------|-----------------------------|
| `create_group/N` | `bench_create_group` |
| `join_welcome` | `bench_join_welcome` |
| `send_app_message` | `bench_app_message_send` |
| `ingest_app_message` | `bench_app_message_ingest` |
Both sides exclude transport crypto and run over in-memory storage, so what is
measured is the engine's own CPU cost. Setup is outside the measured window on
both sides — criterion's `iter_batched(.., PerIteration)` there, an explicit
`setup` lambda here.
**One shape difference, deliberately not hidden:** MDK folds invitees into the
founding group (`FoundingGroupCreated`), while we create at epoch 0 and add in
a second commit to epoch 1. `create_group/N` therefore includes one more commit
on our side. That is a real cost, and averaging it away would be the wrong kind
of favourable.
## Why allocation is reported next to latency
The brief is "as fast if not faster, while avoiding GC as much as possible",
and those are two different measurements. A JVM can win a microbenchmark while
allocating tens of times more per operation; the bill arrives later as GC
pauses on a phone, in a frame the benchmark never renders. So every row carries
**bytes allocated per operation** from `com.sun.management.ThreadMXBean`
(in the JDK — no dependency), alongside p50/p90/p99.
That counter is per-thread, which is why every benchmark body runs inline on
the harness thread via `runBlocking`. Work dispatched elsewhere would allocate
off-book and read as free.
The JVM runs with `-XX:+UseSerialGC` on a fixed 4g heap: the point is to keep
allocation attributable to the benchmark thread and to keep a collection from
landing inside a measured sample and corrupting the percentile it falls in.
## Reading the numbers honestly
- Rust has no GC, so `alloc/op` has no MDK counterpart. It is not a
head-to-head column — it is our own regression signal, and the number to
drive down.
- Latency across a JVM and a Rust binary on the same host is a fair comparison
of *this* workload on *this* machine. It is not a language benchmark.
- `create_group/32` builds 32 KeyPackages in setup. That cost is excluded, but
it makes each iteration expensive to prepare — hence the low iteration count.
+37
View File
@@ -0,0 +1,37 @@
import org.jetbrains.kotlin.gradle.dsl.JvmTarget
plugins {
alias(libs.plugins.jetbrainsKotlinJvm)
application
}
application {
mainClass.set("com.vitorpamplona.marmotbench.MainKt")
applicationName = "marmotbench"
// `-XX:+UseSerialGC` keeps the allocation counter attributable to the
// benchmark thread instead of to background GC worker threads, and a heap
// big enough that a collection never lands mid-measurement. Both matter
// more here than raw throughput: the number we care about is bytes
// allocated per operation, and a GC pause inside a sample corrupts the
// latency percentile it lands in.
applicationDefaultJvmArgs = listOf("-Xmx4g", "-Xms4g", "-XX:+UseSerialGC", "-Dfile.encoding=UTF-8")
}
kotlin {
jvmToolchain(21)
compilerOptions {
jvmTarget.set(JvmTarget.JVM_21)
}
}
dependencies {
// The MLS engine and the Marmot codecs under test.
implementation(project(":quartz"))
// MarmotManager — the app-level entry points MDK's engine benches measure.
implementation(project(":commons"))
implementation(libs.kotlinx.coroutines.core)
// JNI secp256k1 backend quartz needs at runtime on plain JVM.
runtimeOnly(libs.secp256k1.kmp.jni.jvm)
}
@@ -0,0 +1,48 @@
/*
* Copyright (c) 2025 Vitor Pamplona
*
* Permission is hereby granted, free of charge, to any person obtaining a copy of
* this software and associated documentation files (the "Software"), to deal in
* the Software without restriction, including without limitation the rights to use,
* copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the
* Software, and to permit persons to whom the Software is furnished to do so,
* subject to the following conditions:
*
* The above copyright notice and this permission notice shall be included in all
* copies or substantial portions of the Software.
*
* THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
* IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS
* FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR
* COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN
* AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION
* WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
*/
package com.vitorpamplona.marmotbench
/**
* One measured operation.
*
* Latency percentiles AND allocation, because the two answer different
* questions and only one of them is visible in a wall-clock number. A JVM that
* allocates 40x per operation can still win a microbenchmark — the cost shows
* up later as GC pauses on a phone, which is exactly what we are trying not to
* ship.
*/
class BenchResult(
val name: String,
val samples: LongArray,
val bytesPerOp: Long,
val iterations: Int,
) {
private fun percentile(p: Double): Long {
val sorted = samples.sortedArray()
val idx = ((sorted.size - 1) * p).toInt().coerceIn(0, sorted.size - 1)
return sorted[idx]
}
val p50 get() = percentile(0.50)
val p90 get() = percentile(0.90)
val p99 get() = percentile(0.99)
val mean get() = if (samples.isEmpty()) 0L else samples.sum() / samples.size
}
@@ -0,0 +1,145 @@
/*
* Copyright (c) 2025 Vitor Pamplona
*
* Permission is hereby granted, free of charge, to any person obtaining a copy of
* this software and associated documentation files (the "Software"), to deal in
* the Software without restriction, including without limitation the rights to use,
* copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the
* Software, and to permit persons to whom the Software is furnished to do so,
* subject to the following conditions:
*
* The above copyright notice and this permission notice shall be included in all
* copies or substantial portions of the Software.
*
* THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
* IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS
* FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR
* COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN
* AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION
* WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
*/
package com.vitorpamplona.marmotbench
import com.vitorpamplona.amethyst.commons.marmot.MarmotPublisher
import com.vitorpamplona.quartz.marmot.mip00KeyPackages.KeyPackageBundleStore
import com.vitorpamplona.quartz.marmot.mls.group.MarmotMessageStore
import com.vitorpamplona.quartz.marmot.mls.group.MlsGroupStateStore
import com.vitorpamplona.quartz.nip01Core.core.Event
// In-memory stores, matching the commons test doubles byte for byte.
//
// MDK's engine benches run over in-memory SQLite for the same reason: the
// number under test is the engine's own CPU cost, not the disk under it.
// Anything slower here would add storage noise to both sides and hide the
// thing being compared.
/** Stands in for a relay that accepts every commit, so epochs actually advance. */
val ACCEPTING_RELAY = MarmotPublisher { _, _ -> true }
class MemStateStore : MlsGroupStateStore {
private val states = mutableMapOf<String, ByteArray>()
private val retained = mutableMapOf<String, List<ByteArray>>()
override suspend fun save(
nostrGroupId: String,
state: ByteArray,
) {
states[nostrGroupId] = state
}
override suspend fun load(nostrGroupId: String): ByteArray? = states[nostrGroupId]
override suspend fun delete(nostrGroupId: String) {
states.remove(nostrGroupId)
retained.remove(nostrGroupId)
}
override suspend fun listGroups(): List<String> = states.keys.toList()
override suspend fun saveRetainedEpochs(
nostrGroupId: String,
retainedSecrets: List<ByteArray>,
) {
retained[nostrGroupId] = retainedSecrets
}
override suspend fun loadRetainedEpochs(nostrGroupId: String): List<ByteArray> = retained[nostrGroupId] ?: emptyList()
}
class MemMessageStore : MarmotMessageStore {
private val messages = mutableMapOf<String, MutableList<String>>()
private val snapshots = mutableMapOf<String, String>()
private val expiries = mutableMapOf<String, MutableMap<String, Long>>()
private val epochRetentions = mutableMapOf<String, MutableMap<Long, Long>>()
override suspend fun appendMessage(
nostrGroupId: String,
innerEventJson: String,
) {
val log = messages.getOrPut(nostrGroupId) { mutableListOf() }
if (innerEventJson !in log) log.add(innerEventJson)
}
override suspend fun loadMessages(nostrGroupId: String): List<String> = messages[nostrGroupId]?.toList() ?: emptyList()
override suspend fun delete(nostrGroupId: String) {
messages.remove(nostrGroupId)
snapshots.remove(nostrGroupId)
expiries.remove(nostrGroupId)
epochRetentions.remove(nostrGroupId)
}
override suspend fun recordGroupSnapshot(
nostrGroupId: String,
snapshotJson: String,
) {
snapshots[nostrGroupId] = snapshotJson
}
override suspend fun loadGroupSnapshot(nostrGroupId: String): String? = snapshots[nostrGroupId]
// Disappearing messages. First write wins, mirroring the durable stores:
// an expiry is pinned to its message's own source epoch and a replay must
// not re-time it.
override suspend fun recordExpiry(
nostrGroupId: String,
innerEventId: String,
expiresAtSecs: Long,
) {
expiries.getOrPut(nostrGroupId) { mutableMapOf() }.putIfAbsent(innerEventId, expiresAtSecs)
}
override suspend fun loadExpiries(nostrGroupId: String): Map<String, Long> = expiries[nostrGroupId]?.toMap() ?: emptyMap()
override suspend fun removeMessages(
nostrGroupId: String,
innerEventIds: Set<String>,
) {
messages[nostrGroupId]?.removeAll { json -> Event.fromJsonOrNull(json)?.id in innerEventIds }
expiries[nostrGroupId]?.keys?.removeAll(innerEventIds)
}
override suspend fun recordEpochRetention(
nostrGroupId: String,
epoch: Long,
retentionSecs: Long,
) {
epochRetentions.getOrPut(nostrGroupId) { mutableMapOf() }.putIfAbsent(epoch, retentionSecs)
}
override suspend fun loadEpochRetentions(nostrGroupId: String): Map<Long, Long> = epochRetentions[nostrGroupId]?.toMap() ?: emptyMap()
}
class MemBundleStore : KeyPackageBundleStore {
private var snapshot: ByteArray? = null
override suspend fun save(snapshot: ByteArray) {
this.snapshot = snapshot
}
override suspend fun load(): ByteArray? = snapshot
override suspend fun delete() {
snapshot = null
}
}
@@ -0,0 +1,70 @@
/*
* Copyright (c) 2025 Vitor Pamplona
*
* Permission is hereby granted, free of charge, to any person obtaining a copy of
* this software and associated documentation files (the "Software"), to deal in
* the Software without restriction, including without limitation the rights to use,
* copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the
* Software, and to permit persons to whom the Software is furnished to do so,
* subject to the following conditions:
*
* The above copyright notice and this permission notice shall be included in all
* copies or substantial portions of the Software.
*
* THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
* IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS
* FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR
* COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN
* AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION
* WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
*/
package com.vitorpamplona.marmotbench
import java.lang.management.ManagementFactory
/**
* Bytes this thread has allocated, cumulative.
*
* `com.sun.management.ThreadMXBean` is in the JDK, so measuring GC pressure
* costs no dependency. It counts TLAB allocation for the CALLING thread only,
* which is why every benchmark body runs inline on the harness thread rather
* than on a dispatcher — work handed to a coroutine on another thread would
* allocate off-book and read as free.
*/
private val threadMx = ManagementFactory.getThreadMXBean() as com.sun.management.ThreadMXBean
private fun allocatedBytes(): Long = threadMx.getThreadAllocatedBytes(Thread.currentThread().threadId())
/**
* Run [body] until the numbers stop being about JIT.
*
* `setup` runs OUTSIDE the measured window and its cost is excluded, mirroring
* criterion's `iter_batched` with `BatchSize::PerIteration` — which is what
* MDK's engine benches use, so the two sides measure the same span of work.
*/
fun <S> measure(
name: String,
iterations: Int = 200,
warmup: Int = 50,
setup: () -> S,
body: (S) -> Unit,
): BenchResult {
repeat(warmup) { body(setup()) }
// A collection here rather than inside the measured window: the harness
// runs SerialGC on a big heap precisely so this is the last one.
System.gc()
Thread.sleep(50)
val samples = LongArray(iterations)
var allocated = 0L
repeat(iterations) { i ->
val state = setup()
val allocBefore = allocatedBytes()
val start = System.nanoTime()
body(state)
samples[i] = System.nanoTime() - start
allocated += allocatedBytes() - allocBefore
}
return BenchResult(name, samples, allocated / iterations, iterations)
}
@@ -0,0 +1,72 @@
/*
* Copyright (c) 2025 Vitor Pamplona
*
* Permission is hereby granted, free of charge, to any person obtaining a copy of
* this software and associated documentation files (the "Software"), to deal in
* the Software without restriction, including without limitation the rights to use,
* copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the
* Software, and to permit persons to whom the Software is furnished to do so,
* subject to the following conditions:
*
* The above copyright notice and this permission notice shall be included in all
* copies or substantial portions of the Software.
*
* THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
* IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS
* FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR
* COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN
* AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION
* WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
*/
package com.vitorpamplona.marmotbench
import com.vitorpamplona.quartz.utils.Log
import com.vitorpamplona.quartz.utils.LogLevel
private fun micros(nanos: Long) = nanos / 1000.0
private fun kb(bytes: Long) = bytes / 1024.0
fun main(args: Array<String>) {
val json = args.contains("--json")
// Quartz logs at DEBUG by default, and those lines land INSIDE the measured
// window: they cost time, and the string building they do is charged to the
// benchmark thread's allocation counter. Measuring the logger instead of
// the engine would make every number here fiction.
Log.minLevel = LogLevel.ERROR
val results = allBenchmarks()
if (json) {
println("[")
results.forEachIndexed { i, r ->
val comma = if (i == results.size - 1) "" else ","
println(
""" {"name":"${r.name}","p50_us":${"%.1f".format(micros(r.p50))},""" +
""""p90_us":${"%.1f".format(micros(r.p90))},"p99_us":${"%.1f".format(micros(r.p99))},""" +
""""mean_us":${"%.1f".format(micros(r.mean))},"bytes_per_op":${r.bytesPerOp},""" +
""""iterations":${r.iterations}}$comma""",
)
}
println("]")
return
}
println("quartz Marmot — latency and allocation per operation")
println("(alloc is bytes the operation allocated on the calling thread — GC pressure, not heap footprint)")
println()
println("%-28s %10s %10s %10s %12s".format("benchmark", "p50", "p90", "p99", "alloc/op"))
println("-".repeat(74))
results.forEach { r ->
println(
"%-28s %9.1fµs %9.1fµs %9.1fµs %10.1f KB".format(
r.name,
micros(r.p50),
micros(r.p90),
micros(r.p99),
kb(r.bytesPerOp),
),
)
}
}
@@ -0,0 +1,170 @@
/*
* Copyright (c) 2025 Vitor Pamplona
*
* Permission is hereby granted, free of charge, to any person obtaining a copy of
* this software and associated documentation files (the "Software"), to deal in
* the Software without restriction, including without limitation the rights to use,
* copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the
* Software, and to permit persons to whom the Software is furnished to do so,
* subject to the following conditions:
*
* The above copyright notice and this permission notice shall be included in all
* copies or substantial portions of the Software.
*
* THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
* IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS
* FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR
* COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN
* AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION
* WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
*/
package com.vitorpamplona.marmotbench
import com.vitorpamplona.amethyst.commons.marmot.MarmotManager
import com.vitorpamplona.amethyst.commons.marmot.ingest
import com.vitorpamplona.quartz.marmot.appComponents.GroupProfileV1
import com.vitorpamplona.quartz.marmot.mip00KeyPackages.KeyPackageEvent
import com.vitorpamplona.quartz.marmot.mip03GroupMessages.GroupEvent
import com.vitorpamplona.quartz.nip01Core.core.HexKey
import com.vitorpamplona.quartz.nip01Core.core.toHexKey
import com.vitorpamplona.quartz.nip01Core.crypto.KeyPair
import com.vitorpamplona.quartz.nip01Core.signers.NostrSignerInternal
import com.vitorpamplona.quartz.nip59Giftwrap.wraps.GiftWrapEvent
import com.vitorpamplona.quartz.utils.RandomInstance
import kotlinx.coroutines.runBlocking
/**
* The quartz half of the head-to-head against MDK's
* `cgka-engine --bench group_lifecycle`.
*
* Case for case, the two sides measure the same span of work: MDK's benches
* exclude transport crypto and run over in-memory storage, and so do these.
* Where MDK uses criterion's `iter_batched(.., BatchSize::PerIteration)`, the
* setup here is likewise outside the measured window — otherwise we would be
* timing group construction instead of the operation under test.
*
* Every body runs through `runBlocking` on the harness thread on purpose. The
* allocation counter is per-thread, so work handed to another dispatcher would
* not be counted and the operation would read as cheaper than it is.
*/
class Client(
val name: String,
) {
val signer = NostrSignerInternal(KeyPair())
val manager = MarmotManager(signer, MemStateStore(), MemMessageStore(), MemBundleStore(), publisher = ACCEPTING_RELAY)
}
private fun newGroupId(): HexKey = RandomInstance.bytes(32).toHexKey()
private suspend fun keyPackagesFor(count: Int): Pair<List<Client>, List<KeyPackageEvent>> {
val invitees = (0 until count).map { Client("invitee-$it") }
return invitees to invitees.map { it.manager.generateKeyPackageEvent(relays = emptyList()) }
}
/**
* `create_group` with N invitees — MDK's `bench_create_group`.
*
* N Add proposals in ONE commit on both sides, and one Welcome carrying N
* `EncryptedGroupSecrets`. The shapes are not identical and the comparison
* should not pretend otherwise: MDK folds the invitees into the founding
* group (`FoundingGroupCreated`), while we create at epoch 0 and add in a
* second commit to epoch 1. That costs us one extra commit here, which is a
* real difference worth seeing rather than hiding by measuring only the add.
*/
fun benchCreateGroup(invitees: Int): BenchResult =
measure(
name = "create_group/$invitees invitees",
iterations = 30,
warmup = 10,
setup = {
runBlocking {
val alice = Client("alice")
val (_, kps) = keyPackagesFor(invitees)
Triple(alice, kps, newGroupId())
}
},
) { (alice, kps, groupId) ->
runBlocking {
alice.manager.createCurrentProfileGroup(
nostrGroupId = groupId,
relays = listOf("wss://bench.invalid"),
profile = GroupProfileV1("bench", ""),
)
if (kps.isNotEmpty()) alice.manager.addMembers(groupId, kps, emptyList())
}
}
/** `join_welcome` — MDK's `bench_join_welcome`. The invitee's side of the add. */
fun benchJoinWelcome(): BenchResult =
measure(
name = "join_welcome",
iterations = 50,
warmup = 15,
setup = {
runBlocking {
val alice = Client("alice")
val bob = Client("bob")
val groupId = newGroupId()
alice.manager.createCurrentProfileGroup(groupId, listOf("wss://bench.invalid"), GroupProfileV1("bench", ""))
val kp = bob.manager.generateKeyPackageEvent(relays = emptyList())
val (_, welcome) = alice.manager.addMember(groupId, kp, emptyList())
bob to welcome!!.giftWrapEvent
}
},
) { (bob, wrap) ->
runBlocking { bob.manager.ingest(wrap as GiftWrapEvent) }
}
/** `send_app_message` — MDK's `bench_app_message_send`. Encrypt + persist. */
fun benchSendAppMessage(): BenchResult =
measure(
name = "send_app_message",
iterations = 300,
warmup = 100,
setup = {
runBlocking {
val alice = Client("alice")
val groupId = newGroupId()
alice.manager.createCurrentProfileGroup(groupId, listOf("wss://bench.invalid"), GroupProfileV1("bench", ""))
alice to groupId
}
},
) { (alice, groupId) ->
runBlocking { alice.manager.buildTextMessage(groupId, PAYLOAD) }
}
/** `ingest_app_message` — MDK's `bench_app_message_ingest`. Decrypt + persist. */
fun benchIngestAppMessage(): BenchResult =
measure(
name = "ingest_app_message",
iterations = 200,
warmup = 60,
setup = {
runBlocking {
val alice = Client("alice")
val bob = Client("bob")
val groupId = newGroupId()
alice.manager.createCurrentProfileGroup(groupId, listOf("wss://bench.invalid"), GroupProfileV1("bench", ""))
val kp = bob.manager.generateKeyPackageEvent(relays = emptyList())
val (commit, welcome) = alice.manager.addMember(groupId, kp, emptyList())
bob.manager.ingest(welcome!!.giftWrapEvent)
alice.manager.ingest(commit.signedEvent)
val sent = alice.manager.buildTextMessage(groupId, PAYLOAD, persistOwn = false)
bob to sent.outbound.signedEvent
}
},
) { (bob, event) ->
runBlocking { bob.manager.ingest(event as GroupEvent) }
}
private const val PAYLOAD = "marmot benchmark payload — the same 64-ish byte body both sides send"
fun allBenchmarks(): List<BenchResult> =
buildList {
// The same invitee counts MDK's `bench_create_group` uses, plus 0 as
// the founding-only baseline, so the rows line up for comparison.
listOf(0, 1, 8, 32).forEach { add(benchCreateGroup(it)) }
add(benchJoinWelcome())
add(benchSendAppMessage())
add(benchIngestAppMessage())
}
+1
View File
@@ -43,5 +43,6 @@ include(":marmotQuic")
include(":desktopApp")
include(":cli")
include(":relayBench")
include(":marmotBench")
include(":quic-interop")
project(":quic-interop").projectDir = file("quic/interop")