Files
amethyst/marmotBench/build.gradle.kts
Claude d03cc445f1 perf(marmot): head-to-head benchmark against MDK, and where our allocation goes
Adds `marmotBench` — the quartz half of a head-to-head against MDK's
`cgka-engine --bench group_lifecycle`, case for case: create_group/N,
join_welcome, send_app_message, ingest_app_message. Both sides exclude
transport crypto, run over in-memory storage, and keep setup outside the
measured window, so what is compared is the engine's own CPU cost.

Every row also reports BYTES ALLOCATED PER OPERATION, from the JDK's own
`ThreadMXBean` — no dependency added. Latency alone cannot answer "are we
avoiding GC": a JVM can win a microbenchmark and still hand the user a
dropped frame later. The counter is per-thread, so benchmark bodies run
inline via `runBlocking` rather than on a dispatcher, where the allocation
would go uncounted.

First results, same host, nothing else running:

    operation              MDK        quartz      ratio
    create_group/1      3.61 ms     16.93 ms     4.7x slower   27 MB/op
    create_group/8      9.93 ms     51.47 ms     5.2x slower   95 MB/op
    create_group/32    31.64 ms    190.08 ms     6.0x slower  333 MB/op
    join_welcome        4.77 ms      6.22 ms     1.3x slower   10 MB/op
    send_app_message    4.28 ms      1.72 ms     2.5x FASTER  2.8 MB/op

So the steady-state path a user actually exercises — sending a message — is
already faster than the reference. The gap is concentrated in key-agreement
work, and a JFR allocation profile says exactly where: 93% of all allocation
samples are `long[]` from `Curve25519Field.mul/add/sub`, which return a
freshly allocated field element on every single field operation inside
255-iteration scalar-multiplication loops. That one shape explains both the
5x latency gap and the MB-per-op allocation.

The fix (in-place field ops over caller-supplied scratch) is left as its own
change so it can be verified against the RFC vectors on its own merits.

Note: MDK's `ingest_app_message` has no number here. Their bench binary
panics in `bench_deferred_outbound_preflight_matrix` (an assertion on peeler
attempts) before reaching it, and criterion's filter does not skip that
bench's fixture construction. Reporting the gap rather than inventing a
comparison.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016kCuA6tc4JQzHPCDd39GHq
2026-09-10 13:22:43 +00:00

38 lines
1.2 KiB
Kotlin

import org.jetbrains.kotlin.gradle.dsl.JvmTarget
plugins {
alias(libs.plugins.jetbrainsKotlinJvm)
application
}
application {
mainClass.set("com.vitorpamplona.marmotbench.MainKt")
applicationName = "marmotbench"
// `-XX:+UseSerialGC` keeps the allocation counter attributable to the
// benchmark thread instead of to background GC worker threads, and a heap
// big enough that a collection never lands mid-measurement. Both matter
// more here than raw throughput: the number we care about is bytes
// allocated per operation, and a GC pause inside a sample corrupts the
// latency percentile it lands in.
applicationDefaultJvmArgs = listOf("-Xmx4g", "-Xms4g", "-XX:+UseSerialGC", "-Dfile.encoding=UTF-8")
}
kotlin {
jvmToolchain(21)
compilerOptions {
jvmTarget.set(JvmTarget.JVM_21)
}
}
dependencies {
// The MLS engine and the Marmot codecs under test.
implementation(project(":quartz"))
// MarmotManager — the app-level entry points MDK's engine benches measure.
implementation(project(":commons"))
implementation(libs.kotlinx.coroutines.core)
// JNI secp256k1 backend quartz needs at runtime on plain JVM.
runtimeOnly(libs.secp256k1.kmp.jni.jvm)
}