From 8bf1ac27d8216ce376a8be3be4951156dd7b22bc Mon Sep 17 00:00:00 2001 From: Vitor Pamplona Date: Wed, 19 Aug 2026 17:21:30 -0400 Subject: [PATCH] ci: shrink MirrorSyncThroughputTest's corpus to 100k on CI MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit test-geode has been dying at its 30-minute timeout on most runs since at least 18 Aug — six consecutive cancellations on main, both runs on #3955 — while the runs that pass finish the whole job in under five minutes. It is not a job that outgrew its budget; nothing lands in between. The cause is MirrorSyncThroughputTest, which preloads a 1,000,000-event corpus. The sink ingests slower than the in-process source serves, and MirrorWorker's intake channel is deliberately unbounded (its listener callback cannot suspend, so it trySends rather than block the shared OkHttp reader — see its kdoc). The backlog therefore grows until the runner's heap is gone. From the CI log: …116777/1000000 (10,225 ev/s inst) …121961/1000000 ( 1,507 ev/s inst) …122793/1000000 ( 247 ev/s inst) …123113/1000000 ( 86 ev/s inst) Exception: java.lang.OutOfMemoryError thrown from the UncaughtExceptionHandler in thread "kotlinx.coroutines.DefaultExecutor" That is a GC death spiral, then OOM. What turns it into a *timeout* rather than a failure is where the OOM lands: on coroutine threads, reaching the UncaughtExceptionHandler instead of the test thread. JUnit never sees a failure, the JVM never exits, and the job produces no further output until GitHub kills it — which is why the check has been red without ever saying why. The test already takes -DsyncN (default 1,000,000) and geode/build.gradle.kts already forwards it to the test JVM, so this is a workflow-only change. 100k keeps a real ev/s measurement while bounding the worst-case backlog to a tenth of what died. -DsyncN is unset everywhere else, so local and manual runs still measure the full 1M. This does not fix the underlying fragility — a bulk backfill can still exhaust the heap, and an OOM on a coroutine thread will still wedge rather than fail. Both are worth addressing separately; the MirrorWorker half is already noted in relayBench/plans/2026-07-04-sync-throughput-1m.md. Verified: :geode:test --tests "*MirrorSyncThroughputTest*" -DsyncN=100000 passes in 7s and logs "preloaded 100000"; the full suite passes locally in 1m39s. Co-Authored-By: Claude Opus 5 (1M context) --- .github/workflows/build.yml | 14 +++++++++++++- 1 file changed, 13 insertions(+), 1 deletion(-) diff --git a/.github/workflows/build.yml b/.github/workflows/build.yml index 6bf15f67fa..5eb8bbb3a3 100644 --- a/.github/workflows/build.yml +++ b/.github/workflows/build.yml @@ -136,8 +136,20 @@ jobs: with: cache-read-only: ${{ github.ref != 'refs/heads/main' }} + # -DsyncN shrinks MirrorSyncThroughputTest's corpus from its 1,000,000-event + # default. At 1M the sink cannot keep up with the in-process source, and the + # MirrorWorker's deliberately unbounded intake channel buffers the backlog until + # the runner's heap is gone: throughput collapses (13,800 -> 86 ev/s) and an + # OutOfMemoryError lands on a coroutine thread, where the UncaughtExceptionHandler + # swallows it. JUnit never sees a failure, so the JVM wedges and the job burns to + # the timeout with no signal rather than failing. 100k keeps a real ev/s number + # while bounding the worst-case backlog to a tenth of what died. + # + # Only CI is shrunk: -DsyncN is unset everywhere else, so a local or manual run + # still measures the full 1M — the number written up in + # relayBench/plans/2026-07-04-sync-throughput-1m.md. - name: Test geode (gradle) - run: ./gradlew :geode:test + run: ./gradlew :geode:test -DsyncN=100000 - name: Upload geode Test Reports uses: actions/upload-artifact@v7