mirror of
https://github.com/vitorpamplona/amethyst.git
synced 2026-10-05 11:18:24 +00:00
ci: shrink MirrorSyncThroughputTest's corpus to 100k on CI
test-geode has been dying at its 30-minute timeout on most runs since at least 18 Aug — six consecutive cancellations on main, both runs on #3955 — while the runs that pass finish the whole job in under five minutes. It is not a job that outgrew its budget; nothing lands in between. The cause is MirrorSyncThroughputTest, which preloads a 1,000,000-event corpus. The sink ingests slower than the in-process source serves, and MirrorWorker's intake channel is deliberately unbounded (its listener callback cannot suspend, so it trySends rather than block the shared OkHttp reader — see its kdoc). The backlog therefore grows until the runner's heap is gone. From the CI log: …116777/1000000 (10,225 ev/s inst) …121961/1000000 ( 1,507 ev/s inst) …122793/1000000 ( 247 ev/s inst) …123113/1000000 ( 86 ev/s inst) Exception: java.lang.OutOfMemoryError thrown from the UncaughtExceptionHandler in thread "kotlinx.coroutines.DefaultExecutor" That is a GC death spiral, then OOM. What turns it into a *timeout* rather than a failure is where the OOM lands: on coroutine threads, reaching the UncaughtExceptionHandler instead of the test thread. JUnit never sees a failure, the JVM never exits, and the job produces no further output until GitHub kills it — which is why the check has been red without ever saying why. The test already takes -DsyncN (default 1,000,000) and geode/build.gradle.kts already forwards it to the test JVM, so this is a workflow-only change. 100k keeps a real ev/s measurement while bounding the worst-case backlog to a tenth of what died. -DsyncN is unset everywhere else, so local and manual runs still measure the full 1M. This does not fix the underlying fragility — a bulk backfill can still exhaust the heap, and an OOM on a coroutine thread will still wedge rather than fail. Both are worth addressing separately; the MirrorWorker half is already noted in relayBench/plans/2026-07-04-sync-throughput-1m.md. Verified: :geode:test --tests "*MirrorSyncThroughputTest*" -DsyncN=100000 passes in 7s and logs "preloaded 100000"; the full suite passes locally in 1m39s. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
e4184a3806
commit
8bf1ac27d8
@@ -136,8 +136,20 @@ jobs:
|
||||
with:
|
||||
cache-read-only: ${{ github.ref != 'refs/heads/main' }}
|
||||
|
||||
# -DsyncN shrinks MirrorSyncThroughputTest's corpus from its 1,000,000-event
|
||||
# default. At 1M the sink cannot keep up with the in-process source, and the
|
||||
# MirrorWorker's deliberately unbounded intake channel buffers the backlog until
|
||||
# the runner's heap is gone: throughput collapses (13,800 -> 86 ev/s) and an
|
||||
# OutOfMemoryError lands on a coroutine thread, where the UncaughtExceptionHandler
|
||||
# swallows it. JUnit never sees a failure, so the JVM wedges and the job burns to
|
||||
# the timeout with no signal rather than failing. 100k keeps a real ev/s number
|
||||
# while bounding the worst-case backlog to a tenth of what died.
|
||||
#
|
||||
# Only CI is shrunk: -DsyncN is unset everywhere else, so a local or manual run
|
||||
# still measures the full 1M — the number written up in
|
||||
# relayBench/plans/2026-07-04-sync-throughput-1m.md.
|
||||
- name: Test geode (gradle)
|
||||
run: ./gradlew :geode:test
|
||||
run: ./gradlew :geode:test -DsyncN=100000
|
||||
|
||||
- name: Upload geode Test Reports
|
||||
uses: actions/upload-artifact@v7
|
||||
|
||||
Reference in New Issue
Block a user