From d7b5884000fa266a9d08422ab750ffe6cf64e0ec Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 9 Sep 2026 22:34:24 +0000 Subject: [PATCH] test(marmot): bound app-payload parsing and port two reference fuzz targets MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three pieces of hardening taken from the reference client's `:fuzz` module. **Parse bounds.** An app payload reaches a decoder only after MLS has authenticated that a group MEMBER sent it — never that it is well-intentioned, and deep nesting or a huge collection costs a parser far more than it costs whoever sent it. `MarmotJson` now pre-scans for the same three limits the reference draws, at the same values: 64 KiB, depth 16, 64 elements per container. The scan is linear and runs before any JSON library sees the string, and it deliberately does NOT double as a validity filter — malformed input inside the limits still reaches the parser, so its error paths keep being exercised. `MarmotAppEvent.decode` is the choke point, so every inner kind is covered, with kind:1210 checked again at its own entry point because `fromAppEvent` can be reached without it. The byte limit counts UTF-8 rather than UTF-16 code units, which is the difference between a 64 KiB cap and a 256 KiB one for a payload of emoji. **Two ported targets.** Neither could be a like-for-like copy, because the reference fuzzes code we do not have in that shape — their metadata walkers are deliberately Android-free byte functions, ours is `ExifInterface` over a `Uri`. What ports is the set of oracles: - Identity references, against `Nip19Parser`: never throws, deterministic, idempotent on what it canonicalises, and never emits a key that is not 32 bytes of lowercase hex. The corpus is their grammar — `nostr:`, profile links, percent-encoded separators, truncated and over-long bech32 bodies, clipboard text with several references run together. - Container sniffing, against `ShareHelper`: never throws, deterministic, always names a declared kind, a mismatched walker does not claim the container, and — the one that matters most — only the header decides. A sniffer that read past its header would let bytes deep inside a file relabel it, which is what content-type confusion needs. Seeded rather than Jazzer-driven, so no fuzzing engine joins the build and a failure reproduces from the printed seed. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_016kCuA6tc4JQzHPCDd39GHq --- .../ShareHelperContainerSniffingTest.kt | 168 +++++++++++++++ .../foundation/appEvents/MarmotAppEvent.kt | 8 + .../marmot/foundation/appEvents/MarmotJson.kt | 93 ++++++++ .../foundation/appEvents/MarmotSystemEvent.kt | 4 + .../appEvents/MarmotJsonBoundsTest.kt | 129 ++++++++++++ .../Nip19ParserAdversarialInputTest.kt | 199 ++++++++++++++++++ 6 files changed, 601 insertions(+) create mode 100644 amethyst/src/test/java/com/vitorpamplona/amethyst/ui/components/ShareHelperContainerSniffingTest.kt create mode 100644 quartz/src/commonTest/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotJsonBoundsTest.kt create mode 100644 quartz/src/commonTest/kotlin/com/vitorpamplona/quartz/nip19Bech32/Nip19ParserAdversarialInputTest.kt diff --git a/amethyst/src/test/java/com/vitorpamplona/amethyst/ui/components/ShareHelperContainerSniffingTest.kt b/amethyst/src/test/java/com/vitorpamplona/amethyst/ui/components/ShareHelperContainerSniffingTest.kt new file mode 100644 index 0000000000..bcb68cc0ff --- /dev/null +++ b/amethyst/src/test/java/com/vitorpamplona/amethyst/ui/components/ShareHelperContainerSniffingTest.kt @@ -0,0 +1,168 @@ +/* + * Copyright (c) 2025 Vitor Pamplona + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * this software and associated documentation files (the "Software"), to deal in + * the Software without restriction, including without limitation the rights to use, + * copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the + * Software, and to permit persons to whom the Software is furnished to do so, + * subject to the following conditions: + * + * The above copyright notice and this permission notice shall be included in all + * copies or substantial portions of the Software. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS + * FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR + * COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN + * AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION + * WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. + */ +package com.vitorpamplona.amethyst.ui.components + +import org.junit.Assert.assertEquals +import org.junit.Assert.assertTrue +import org.junit.Test +import java.io.File +import java.nio.file.Files +import kotlin.random.Random + +/** + * Container sniffing under adversarial bytes. + * + * Ported from the reference client's `ImageContainerBytesFuzzTest`. Their + * walkers strip metadata from a `ByteArray`; ours identifies a container from a + * file's first bytes, so the oracles carry over even though the code does not: + * arbitrary input never throws, the answer is deterministic and always a + * declared kind, a mismatched walker does not claim the container, and only the + * header can decide — trailing bytes must be irrelevant. + * + * That last one is the load-bearing invariant here. A sniffer that read past + * its header would let attacker-chosen bytes deep inside a file change how the + * file is labelled, which is exactly what content-type confusion needs. + */ +class ShareHelperContainerSniffingTest { + private val imageKinds = setOf("jpg", "png", "gif", "webp") + private val videoKinds = setOf("mp4", "mov", "webm", "avi") + + private lateinit var dir: File + + private fun file(bytes: ByteArray): File { + if (!::dir.isInitialized) dir = Files.createTempDirectory("sniffing").toFile() + val f = File.createTempFile("probe", ".bin", dir) + f.writeBytes(bytes) + return f + } + + private fun headers(): List> = + listOf( + "jpg" to byteArrayOf(0xFF.toByte(), 0xD8.toByte(), 0x00, 0x01), + "png" to byteArrayOf(0x89.toByte(), 0x50, 0x4E, 0x47), + "gif" to "GIF89a".encodeToByteArray(), + "webp" to ("RIFF".encodeToByteArray() + ByteArray(4) + "WEBP".encodeToByteArray()), + "webm" to byteArrayOf(0x1A, 0x45, 0xDF.toByte(), 0xA3.toByte()), + "avi" to ("RIFF".encodeToByteArray() + ByteArray(4) + "AVI ".encodeToByteArray()), + "mp4" to (ByteArray(4) + "ftyp".encodeToByteArray() + "isom".encodeToByteArray()), + "mov" to (ByteArray(4) + "ftyp".encodeToByteArray() + "qt ".encodeToByteArray()), + ) + + /** Random bytes, truncations, and real headers with random tails. */ + private fun corpus(seed: Int): List { + val rnd = Random(seed) + return buildList { + repeat(200) { add(ByteArray(rnd.nextInt(0, 64)) { rnd.nextInt(256).toByte() }) } + headers().forEach { (_, header) -> + add(header) + repeat(8) { add(header + ByteArray(rnd.nextInt(0, 128)) { rnd.nextInt(256).toByte() }) } + // Truncated to every prefix length: the sniffer must survive a + // header that stops in the middle of a magic number. + for (cut in 0 until header.size) add(header.copyOfRange(0, cut)) + } + add(ByteArray(0)) + } + } + + @Test + fun sniffingNeverThrowsAndAlwaysNamesADeclaredKind() { + val seed = 20260914 + corpus(seed).forEach { bytes -> + val f = file(bytes) + val image = + try { + ShareHelper.getImageExtension(f) + } catch (e: Throwable) { + throw AssertionError("image sniffing threw on " + bytes.size + " bytes (seed " + seed + ")", e) + } + val video = + try { + ShareHelper.getVideoExtension(f) + } catch (e: Throwable) { + throw AssertionError("video sniffing threw on " + bytes.size + " bytes (seed " + seed + ")", e) + } + assertTrue("unexpected image kind " + image, image in imageKinds) + assertTrue("unexpected video kind " + video, video in videoKinds) + } + } + + @Test + fun sniffingIsDeterministic() { + corpus(20260915).forEach { bytes -> + val f = file(bytes) + assertEquals(ShareHelper.getImageExtension(f), ShareHelper.getImageExtension(f)) + assertEquals(ShareHelper.getVideoExtension(f), ShareHelper.getVideoExtension(f)) + } + } + + @Test + fun onlyTheHeaderDecides() { + // Appending arbitrary bytes must not change the verdict. A sniffer that + // read further would let bytes deep inside a file relabel it. + val rnd = Random(20260916) + headers().forEach { (_, header) -> + val bare = file(header) + val image = ShareHelper.getImageExtension(bare) + val video = ShareHelper.getVideoExtension(bare) + repeat(16) { + val padded = file(header + ByteArray(rnd.nextInt(1, 512)) { rnd.nextInt(256).toByte() }) + assertEquals("a trailing byte changed the image verdict", image, ShareHelper.getImageExtension(padded)) + assertEquals("a trailing byte changed the video verdict", video, ShareHelper.getVideoExtension(padded)) + } + } + } + + @Test + fun everyRealHeaderIsIdentifiedByItsOwnWalker() { + headers().forEach { (kind, header) -> + val f = file(header + ByteArray(32)) + val sniffed = if (kind in imageKinds) ShareHelper.getImageExtension(f) else ShareHelper.getVideoExtension(f) + assertEquals("header for " + kind + " was not identified", kind, sniffed) + } + } + + @Test + fun aMismatchedWalkerDoesNotClaimTheContainer() { + // Their "a mismatched walker must reject the container", in the shape + // our API allows: asking the video sniffer about a JPEG must fall back + // to the video default rather than reporting an image kind. + headers().forEach { (kind, header) -> + val f = file(header + ByteArray(32)) + if (kind in imageKinds) { + assertTrue("an image was reported as a video kind", ShareHelper.getVideoExtension(f) in videoKinds) + } else { + assertTrue("a video was reported as an image kind", ShareHelper.getImageExtension(f) in imageKinds) + } + } + } + + @Test + fun aTruncatedHeaderFallsBackRatherThanGuessing() { + // Under four readable bytes there is nothing to decide on, and reading + // past the end is how a sniffer turns a short file into a crash. + listOf(ByteArray(0), byteArrayOf(0xFF.toByte()), byteArrayOf(0xFF.toByte(), 0xD8.toByte()), byteArrayOf(0x89.toByte(), 0x50, 0x4E)) + .forEach { bytes -> + val f = file(bytes) + assertEquals("jpg", ShareHelper.getImageExtension(f)) + assertEquals("mp4", ShareHelper.getVideoExtension(f)) + } + } +} diff --git a/quartz/src/commonMain/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotAppEvent.kt b/quartz/src/commonMain/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotAppEvent.kt index 5e63e03345..e63c2f7d2f 100644 --- a/quartz/src/commonMain/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotAppEvent.kt +++ b/quartz/src/commonMain/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotAppEvent.kt @@ -139,6 +139,14 @@ class MarmotAppEvent( * @throws IllegalArgumentException naming the reason. */ fun decode(json: String): MarmotAppEvent { + // Shape before content. MLS authenticates that a group MEMBER sent + // these bytes, never that they are well-intentioned, and a deeply + // nested or enormous payload costs a parser far more than it costs + // the sender. The pre-scan is linear and runs before any JSON + // library sees the string. + require(MarmotJson.withinResourceBounds(json)) { + "Marmot app payload exceeds the parse bounds" + } val obj = MarmotJson.parseObject(json) require(!obj.containsKey("sig")) { diff --git a/quartz/src/commonMain/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotJson.kt b/quartz/src/commonMain/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotJson.kt index 022fc4d381..6166c27554 100644 --- a/quartz/src/commonMain/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotJson.kt +++ b/quartz/src/commonMain/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotJson.kt @@ -76,6 +76,99 @@ class MarmotJsonObject( object MarmotJson { private val parser = Json { ignoreUnknownKeys = false } + /** + * Resource bounds for an app-payload `content` string. + * + * A payload reaches a decoder only after MLS has authenticated that a group + * MEMBER sent it — never that it is well-intentioned. Deep nesting and huge + * collections cost a parser far more than they cost the sender, so the + * shape is checked with a linear pre-scan BEFORE any JSON library sees the + * bytes. The reference client draws the same three limits at the same + * values, which is why they are these numbers and not rounder ones. + */ + const val MAX_INPUT_BYTES = 64 * 1024 + const val MAX_JSON_DEPTH = 16 + const val MAX_COLLECTION_ELEMENTS = 64 + + /** + * True when [json] is small enough and shallow enough to hand to a parser. + * + * Deliberately NOT a validity check: malformed input that stays inside the + * limits still goes through to the real parser, so its error paths keep + * being exercised rather than being masked by a pre-filter. + */ + fun withinResourceBounds(json: String): Boolean = !exceedsByteLimit(json) && BoundsScanner().scan(json) + + private fun exceedsByteLimit(json: String): Boolean { + var bytes = 0 + var i = 0 + while (i < json.length) { + val ch = json[i] + bytes += + when { + ch.code <= 0x7F -> 1 + ch.code <= 0x7FF -> 2 + ch.isHighSurrogate() && i + 1 < json.length && json[i + 1].isLowSurrogate() -> { + i++ + 4 + } + + else -> 3 + } + if (bytes > MAX_INPUT_BYTES) return true + i++ + } + return false + } + + /** + * One pass over the text, counting container depth and the members at each + * depth. It only has to be right about structure — string boundaries and + * escapes — so it reads nothing else. + */ + private class BoundsScanner { + private val membersByDepth = IntArray(MAX_JSON_DEPTH + 2) + private var depth = 0 + private var quote: Char? = null + private var escaped = false + + fun scan(json: String): Boolean { + for (ch in json) { + if (!consume(ch)) return false + } + return true + } + + private fun consume(ch: Char): Boolean { + val open = quote + if (open != null) { + when { + escaped -> escaped = false + ch == '\\' -> escaped = true + ch == open -> quote = null + } + return true + } + when (ch) { + '"' -> quote = ch + '{', '[' -> { + depth++ + if (depth > MAX_JSON_DEPTH) return false + membersByDepth[depth] = 0 + } + + '}', ']' -> if (depth > 0) depth-- + ',' -> { + if (depth > 0) { + membersByDepth[depth]++ + if (membersByDepth[depth] >= MAX_COLLECTION_ELEMENTS) return false + } + } + } + return true + } + } + fun parseObject(json: String): MarmotJsonObject { val element = parser.parseToJsonElement(json) val obj = element as? JsonObject ?: throw IllegalArgumentException("payload is not a JSON object") diff --git a/quartz/src/commonMain/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotSystemEvent.kt b/quartz/src/commonMain/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotSystemEvent.kt index 3f0ca9cfee..1eefadacf9 100644 --- a/quartz/src/commonMain/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotSystemEvent.kt +++ b/quartz/src/commonMain/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotSystemEvent.kt @@ -133,6 +133,10 @@ class MarmotSystemEvent( */ fun fromAppEvent(event: MarmotAppEvent): MarmotSystemEvent? { if (event.kind != MarmotAppEvent.KIND_SYSTEM) return null + // Bounded before parsed: a row's content is peer-authored, and + // deep nesting or a huge collection costs a parser far more than + // it costs whoever sent it. + if (!MarmotJson.withinResourceBounds(event.content)) return null return try { val obj = MarmotJson.parseObject(event.content) if (obj.int("v") != SCHEMA_VERSION) return null diff --git a/quartz/src/commonTest/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotJsonBoundsTest.kt b/quartz/src/commonTest/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotJsonBoundsTest.kt new file mode 100644 index 0000000000..20a60f1edb --- /dev/null +++ b/quartz/src/commonTest/kotlin/com/vitorpamplona/quartz/marmot/foundation/appEvents/MarmotJsonBoundsTest.kt @@ -0,0 +1,129 @@ +/* + * Copyright (c) 2025 Vitor Pamplona + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * this software and associated documentation files (the "Software"), to deal in + * the Software without restriction, including without limitation the rights to use, + * copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the + * Software, and to permit persons to whom the Software is furnished to do so, + * subject to the following conditions: + * + * The above copyright notice and this permission notice shall be included in all + * copies or substantial portions of the Software. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS + * FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR + * COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN + * AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION + * WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. + */ +package com.vitorpamplona.quartz.marmot.foundation.appEvents + +import kotlin.test.Test +import kotlin.test.assertEquals +import kotlin.test.assertFalse +import kotlin.test.assertNotNull +import kotlin.test.assertNull +import kotlin.test.assertTrue + +/** + * Resource bounds on an app payload's `content`, ported from the reference + * client's `GroupSystemEventFuzzTest` contract. + * + * The threat is not a malformed payload — those are dropped either way — but a + * well-formed one that is expensive. MLS authenticates that a group MEMBER sent + * these bytes and nothing more, and deep nesting or a huge collection costs a + * parser far more than it costs whoever sent it. + * + * The limits are checked with a linear pre-scan BEFORE any JSON library sees + * the string, and the pre-scan deliberately does not double as a validity + * check: malformed input that stays inside the limits still reaches the real + * parser, so its error paths keep being exercised. + */ +class MarmotJsonBoundsTest { + private fun nested(containers: Int) = + buildString { + append("{\"system_type\":\"nested\",\"data\":") + repeat(containers) { append('[') } + append('0') + repeat(containers) { append(']') } + append('}') + } + + private fun wide(members: Int) = + (0 until members).joinToString( + prefix = "{\"system_type\":\"wide\",\"data\":{", + postfix = "}}", + ) { "\"field$it\":$it" } + + @Test + fun aPayloadInsideEveryLimitIsAccepted() { + assertTrue(MarmotJson.withinResourceBounds("""{"v":1,"system_type":"member_added"}""")) + assertTrue(MarmotJson.withinResourceBounds(nested(MarmotJson.MAX_JSON_DEPTH - 2))) + assertTrue(MarmotJson.withinResourceBounds(wide(MarmotJson.MAX_COLLECTION_ELEMENTS - 2))) + } + + @Test + fun nestingBeyondTheDepthLimitIsRefused() { + assertFalse(MarmotJson.withinResourceBounds(nested(MarmotJson.MAX_JSON_DEPTH + 1))) + } + + @Test + fun aCollectionBeyondTheElementLimitIsRefused() { + assertFalse(MarmotJson.withinResourceBounds(wide(MarmotJson.MAX_COLLECTION_ELEMENTS + 2))) + } + + @Test + fun inputBeyondTheByteLimitIsRefused() { + val big = "{\"text\":\"" + "a".repeat(MarmotJson.MAX_INPUT_BYTES) + "\"}" + assertFalse(MarmotJson.withinResourceBounds(big)) + } + + @Test + fun theByteLimitCountsUtf8NotUtf16() { + // A four-byte emoji is two Kotlin chars. Counting chars would let a + // payload four times over the limit through. + val emoji = "😀" + val overshoot = "{\"text\":\"" + emoji.repeat(MarmotJson.MAX_INPUT_BYTES / 4) + "\"}" + assertFalse(MarmotJson.withinResourceBounds(overshoot)) + } + + @Test + fun bracesInsideStringsDoNotCountAsNesting() { + // The scanner has to know where strings begin and end, or a caption + // that merely mentions a bracket would be refused as too deep. + val text = "[".repeat(MarmotJson.MAX_JSON_DEPTH * 4) + assertTrue(MarmotJson.withinResourceBounds("""{"system_type":"x","text":"$text"}""")) + } + + @Test + fun anEscapedQuoteDoesNotEndTheString() { + assertTrue(MarmotJson.withinResourceBounds("""{"system_type":"x","text":"she said \"hi\" and [[["}""")) + } + + @Test + fun aBoundedButMalformedPayloadStillReachesTheParser() { + // The pre-scan is not a validity filter. If it rejected malformed input + // itself, the parser's error paths would stop being exercised. + assertTrue(MarmotJson.withinResourceBounds("{not json at all")) + } + + @Test + fun anOversizedSystemRowIsDroppedRatherThanParsed() { + val payload = nested(MarmotJson.MAX_JSON_DEPTH + 4) + val event = MarmotAppEvent("id", "a".repeat(64), 1L, MarmotAppEvent.KIND_SYSTEM, emptyArray(), payload) + assertNull(MarmotSystemEvent.fromAppEvent(event)) + } + + @Test + fun anOrdinarySystemRowStillDecodes() { + // The bound must not cost the real thing: a row this client derived has + // to survive its own round trip. + val row = MarmotSystemEvent(MarmotSystemType.GROUP_RENAMED, actor = "b".repeat(64), name = "after") + val appEvent = row.toAppEvent("b".repeat(64), 1_800_000_000L) + val decoded = assertNotNull(MarmotSystemEvent.fromAppEvent(appEvent)) + assertEquals(MarmotSystemType.GROUP_RENAMED, decoded.systemType) + assertEquals("after", decoded.name) + } +} diff --git a/quartz/src/commonTest/kotlin/com/vitorpamplona/quartz/nip19Bech32/Nip19ParserAdversarialInputTest.kt b/quartz/src/commonTest/kotlin/com/vitorpamplona/quartz/nip19Bech32/Nip19ParserAdversarialInputTest.kt new file mode 100644 index 0000000000..080b08398f --- /dev/null +++ b/quartz/src/commonTest/kotlin/com/vitorpamplona/quartz/nip19Bech32/Nip19ParserAdversarialInputTest.kt @@ -0,0 +1,199 @@ +/* + * Copyright (c) 2025 Vitor Pamplona + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * this software and associated documentation files (the "Software"), to deal in + * the Software without restriction, including without limitation the rights to use, + * copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the + * Software, and to permit persons to whom the Software is furnished to do so, + * subject to the following conditions: + * + * The above copyright notice and this permission notice shall be included in all + * copies or substantial portions of the Software. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS + * FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR + * COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN + * AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION + * WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. + */ +package com.vitorpamplona.quartz.nip19Bech32 + +import com.vitorpamplona.quartz.nip19Bech32.entities.IPubKeyEntity +import kotlin.random.Random +import kotlin.test.Test +import kotlin.test.assertEquals +import kotlin.test.assertTrue + +/** + * Identity-reference parsing under adversarial input. + * + * Ported from the reference client's `IdentityReferenceFuzzTest`: same grammar + * (`nostr:`, profile links, percent-encoded separators, truncated bech32, + * multi-token clipboard text), same invariants — the parser never throws, is + * deterministic, is idempotent on whatever it canonicalises, and never emits a + * key that is not a 32-byte lowercase hex string. + * + * Seeded rather than Jazzer-driven, so it needs no fuzzing engine on the build + * and a failure reproduces from the printed seed. That is the whole difference: + * the oracles below are theirs. + */ +class Nip19ParserAdversarialInputTest { + private val bech32Body = "qpzry9x8gf2tvdw0s3jn54khce6mua7l" + private val schemes = listOf("nostr", "http", "https", "marmot", "whitenoise", "web+nostr", "") + private val hosts = listOf("njump.me", "primal.net", "example.com", "profile", "") + private val separators = listOf("://", ":", "%3A%2F%2F", "%2F", "") + private val pathPrefixes = listOf("profile/", "profile%2F", "p/", "") + private val joiners = listOf(",", " ", "\n", "\t", " ", ";", "") + + private fun body( + rnd: Random, + length: Int, + ) = buildString { repeat(length) { append(bech32Body[rnd.nextInt(bech32Body.length)]) } } + + /** + * A shaped-but-corrupt npub: the right alphabet and length, a checksum that + * does not hold. This is what a truncated or mistyped paste actually looks + * like, and it is most of the corpus on purpose. + */ + private fun corruptNpub(rnd: Random) = "npub1" + body(rnd, 58) + + /** A real npub, checksum and all, for the cases that must SUCCEED. */ + private fun validNpub(rnd: Random) = ByteArray(32) { rnd.nextInt(256).toByte() }.toNpub() + + private fun reference(rnd: Random): String { + val key = + when (rnd.nextInt(6)) { + 0 -> validNpub(rnd) + // Truncated and over-long bodies: the 58-char rule is what + // stops a half-pasted npub from decoding to a short key. + 1 -> "npub1" + body(rnd, rnd.nextInt(1, 58)) + 2 -> "npub1" + body(rnd, rnd.nextInt(59, 90)) + 3 -> "nprofile1" + body(rnd, rnd.nextInt(1, 120)) + 4 -> corruptNpub(rnd).uppercase() + else -> body(rnd, rnd.nextInt(0, 70)) + } + val scheme = schemes[rnd.nextInt(schemes.size)] + return when (scheme) { + "" -> key + "nostr" -> "nostr:" + key + "http", "https" -> + scheme + "://" + hosts[rnd.nextInt(hosts.size)] + "/" + + pathPrefixes[rnd.nextInt(pathPrefixes.size)] + key + + else -> + scheme + separators[rnd.nextInt(separators.size)] + + pathPrefixes[rnd.nextInt(pathPrefixes.size)] + key + } + } + + private fun corpus(seed: Int): List { + val rnd = Random(seed) + return buildList { + repeat(400) { add(reference(rnd)) } + // Clipboard-shaped input: several references run together. + repeat(100) { + val count = rnd.nextInt(2, 6) + add((0 until count).joinToString(joiners[rnd.nextInt(joiners.size)]) { reference(rnd) }) + } + // Degenerate shapes the grammar above never produces. + addAll(listOf("", " ", "\n", "nostr:", "npub1", "@", "nostr:@", "://", "%", "npub1 npub1")) + } + } + + @Test + fun parsingNeverThrowsAndIsDeterministic() { + val seed = 20260909 + corpus(seed).forEach { input -> + val first = + try { + Nip19Parser.uriToRoute(input) + } catch (e: Throwable) { + throw AssertionError("uriToRoute threw on " + input.take(120) + " (seed " + seed + ")", e) + } + val second = Nip19Parser.uriToRoute(input) + assertEquals(first?.entity, second?.entity, "parsing must be deterministic for " + input.take(120)) + assertEquals(first?.nip19raw, second?.nip19raw) + } + } + + @Test + fun cleaningNeverThrowsAndIsIdempotent() { + val seed = 20260910 + corpus(seed).forEach { input -> + val cleaned = + try { + Nip19Parser.tryParseAndClean(input) + } catch (e: Throwable) { + throw AssertionError("tryParseAndClean threw on " + input.take(120) + " (seed " + seed + ")", e) + } + if (cleaned != null) { + // Re-cleaning its own output must be a fixed point, or two + // clients that clean a different number of times disagree about + // the same paste. + assertEquals( + cleaned, + Nip19Parser.tryParseAndClean(cleaned), + "cleaning is not idempotent for " + input.take(120), + ) + } + } + } + + @Test + fun aDecodedKeyIsAlwaysThirtyTwoBytesOfLowercaseHex() { + val seed = 20260911 + corpus(seed).forEach { input -> + val entity = Nip19Parser.uriToRoute(input)?.entity + if (entity is IPubKeyEntity) { + val hex = entity.hex + assertEquals( + 64, + hex.length, + "a pubkey entity must decode to 32 bytes, got " + hex.length + " from " + input.take(120), + ) + assertTrue( + hex.all { it in '0'..'9' || it in 'a'..'f' }, + "a decoded key must be lowercase hex, got " + hex, + ) + } + } + } + + @Test + fun aParsedReferenceReParsesToTheSameEntity() { + // The canonical `nip19raw` is what the app stores and re-reads. If it + // did not round-trip, a reference would decay every time it was copied + // through the UI. + val seed = 20260912 + corpus(seed).forEach { input -> + val parsed = Nip19Parser.uriToRoute(input) ?: return@forEach + val reparsed = Nip19Parser.uriToRoute(parsed.nip19raw) + assertEquals(parsed.entity, reparsed?.entity, "re-parsing " + parsed.nip19raw + " changed the entity") + } + } + + @Test + fun aWellFormedNpubIsFoundInsideEveryCarrierShape() { + // The negative cases above are only half the contract: the parser also + // has to keep finding a real reference through a link, a scheme it does + // not know, and percent-encoding. + val rnd = Random(20260913) + val key = validNpub(rnd) + val carriers = + listOf( + key, + "nostr:" + key, + "@" + key, + "https://njump.me/" + key, + "https://example.com/profile/" + key, + "web+nostr://" + key, + "text before nostr:" + key + " and after", + ) + carriers.forEach { carrier -> + val entity = Nip19Parser.uriToRoute(carrier)?.entity + assertTrue(entity is IPubKeyEntity, "no pubkey found in " + carrier) + } + } +}