mirror of
https://github.com/vitorpamplona/amethyst.git
synced 2026-10-06 11:48:24 +00:00
perf: scan note content by literal instead of driving the regex engine
The NIP-19, hashtag and legacy-#[n] scans all ran `Regex.findAll` over the whole
of a note's content. `findAll` restarts the regex engine at every position in the
string, so cost scaled with content length regardless of whether the content had
anything to find — and most notes have nothing to find.
Each of these patterns has a literal that must be present for any match:
- `nip19regex` can only match at one of eight entity prefixes, all starting `n`
- `hashtagSearch` requires `(?:\s|\A)` immediately before a `#`
- `tagSearch` requires the same before `#[`
- `RichTextParser.tagIndex` requires the literal `#[`
So the scans now jump between candidate positions with `indexOf`/char compares and
apply the regex ANCHORED there via `matchAt`. For the NIP-19 scan, `(nostr:)?@?`
are optional, so anchoring at the entity's own `n` matches the same entities and
captures the same groups the callers read.
Measured on the production content distribution — median 529 B with a tail to
767 KB, taken from an on-device heap dump where 2,541 of 4,573 live Matchers were
running the NIP-19 regex:
before after (no match / with matches)
Nip19Parser.parseAll 19 MB/s 922 / 61-654 MB/s
findHashtags 68 MB/s ~19,000 / ~1,240 MB/s
IndexedTags 63 MB/s ~19,900 / ~12,100 MB/s
RichTextParser '#[' - 4.93x on the #[n] step
Worst case improves most: a 767 KB long-form note went from 40.6 ms to 0.83 ms
per NIP-19 scan.
Two further notes on the NIP-19 scan. Every prefix starts with `n` and English
prose is ~7% `n`, so testing all eight prefixes at each `n` dominated the scan;
dispatching on the second character first cuts that to at most two and is worth
2.1x on its own. And `RichTextParser` gates on `contains("#[")`, not
`startsWith`, because `find()` also matches `#[n]` mid-word.
Behaviour is unchanged. `RegexContentBenchmark` (new, in commons) pins each
rewritten scan against a verbatim copy of the implementation it replaces, plus
~6,000 fuzz cases, the boundary cases where `n`/`#` is the last character, and
hardcoded segment expectations for the RichTextParser '#' path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
775fb86b39
commit
a88ad5b39d
+5
-1
@@ -416,8 +416,12 @@ class RichTextParser {
|
||||
tags: ImmutableListOfLists<String>,
|
||||
): Segment {
|
||||
// First #[n]
|
||||
// [tagIndex] requires the literal "#[", so a word without it can never match.
|
||||
// Every plain "#hashtag" reaches this function, and the regex scan is far more
|
||||
// expensive than the substring check that rules it out — so gate on the literal.
|
||||
// Uses contains(), not startsWith(), because find() also matches "#[n]" mid-word.
|
||||
try {
|
||||
val matcher = tagIndex.find(word)
|
||||
val matcher = if (word.contains("#[")) tagIndex.find(word) else null
|
||||
if (matcher != null) {
|
||||
val index = matcher.groups[1]?.value?.toInt()
|
||||
val suffix = matcher.groups[2]?.value
|
||||
|
||||
+651
@@ -0,0 +1,651 @@
|
||||
/*
|
||||
* Copyright (c) 2025 Vitor Pamplona
|
||||
*
|
||||
* Permission is hereby granted, free of charge, to any person obtaining a copy of
|
||||
* this software and associated documentation files (the "Software"), to deal in
|
||||
* the Software without restriction, including without limitation the rights to use,
|
||||
* copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the
|
||||
* Software, and to permit persons to whom the Software is furnished to do so,
|
||||
* subject to the following conditions:
|
||||
*
|
||||
* The above copyright notice and this permission notice shall be included in all
|
||||
* copies or substantial portions of the Software.
|
||||
*
|
||||
* THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||
* IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS
|
||||
* FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR
|
||||
* COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN
|
||||
* AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION
|
||||
* WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
|
||||
*/
|
||||
package com.vitorpamplona.amethyst.commons.prodbench
|
||||
|
||||
import com.vitorpamplona.amethyst.commons.model.toImmutableListOfLists
|
||||
import com.vitorpamplona.amethyst.commons.richtext.RichTextParser
|
||||
import com.vitorpamplona.amethyst.commons.richtext.RichTextParser.Companion.tagIndex
|
||||
import com.vitorpamplona.quartz.nip10Notes.content.findHashtags
|
||||
import com.vitorpamplona.quartz.nip10Notes.content.findIndexTagsWithEventsOrAddresses
|
||||
import com.vitorpamplona.quartz.nip10Notes.content.findIndexTagsWithPeople
|
||||
import com.vitorpamplona.quartz.nip10Notes.content.findNostrUris
|
||||
import com.vitorpamplona.quartz.nip10Notes.content.hashtagSearch
|
||||
import com.vitorpamplona.quartz.nip10Notes.content.tagSearch
|
||||
import com.vitorpamplona.quartz.nip19Bech32.Nip19Parser
|
||||
import kotlin.test.Test
|
||||
|
||||
/**
|
||||
* Measures the regex scans Amethyst runs over note **content**.
|
||||
*
|
||||
* Why: an on-device heap dump (SM-T220, Dr. Edo's account) found 2,541 of 4,573
|
||||
* live `java.util.regex.Matcher` instances running [Nip19Parser.nip19regex],
|
||||
* reached from `BaseNoteEvent.citedNIP19()` during ingest. Kotlin's `findAll`
|
||||
* allocates a **new Matcher per match**, each retaining the whole input
|
||||
* (`jvmMain/kotlin/text/regex/Regex.kt`: `matcher.pattern().matcher(input)`),
|
||||
* so a long article with N mentions builds N matchers over the full text.
|
||||
*
|
||||
* Corpus sizes come from that same dump's measured distribution of the strings
|
||||
* these matchers held: median 529 B, p90/max 767 KB (long-form articles — the
|
||||
* v1.13.0 release notes at 68 KB, an essay at 50 KB).
|
||||
*
|
||||
* Deterministic and offline. Prints ns/op and MB/s; no assertions on wall time
|
||||
* (CI machines vary) beyond a sanity check that the scans return results.
|
||||
*/
|
||||
class RegexContentBenchmark {
|
||||
companion object {
|
||||
private const val NPUB = "npub180cvv07tjdrrgpa0j7j7tmnyl2yr6yr7l8j4s3evf6u64th6gkwsyjh6w6"
|
||||
private const val NEVENT = "nevent1qqstna2yrezu5wghjvswqqculvvwxsrcvu7uc0f78gan4xqhvz49d9spr3mhxue69uhkummnw3ez6un9d3shjtn4de6x2argwghx6egpr4mhxue69uhkummnw3ez6ur4vgh8wetvd3hhyer9wghxuet5nxnepm"
|
||||
|
||||
/** A note with exactly [mentions] nostr: URIs, padded with prose to ~[targetBytes]. */
|
||||
fun note(
|
||||
targetBytes: Int,
|
||||
mentions: Int,
|
||||
hashtags: Boolean = true,
|
||||
): String {
|
||||
val filler =
|
||||
if (hashtags) {
|
||||
"A few months ago a nostrich was switching from iOS to Android and asked for " +
|
||||
"suggestions for #Nostr apps to try out. Here is what came back, with notes. "
|
||||
} else {
|
||||
"A few months ago a nostrich was switching from iOS to Android and asked for " +
|
||||
"suggestions for great apps to try out. Here is what came back, with notes. "
|
||||
}
|
||||
val sb = StringBuilder(targetBytes + 4096)
|
||||
var placed = 0
|
||||
val stride = if (mentions > 0) targetBytes / (mentions + 1) else Int.MAX_VALUE
|
||||
while (sb.length < targetBytes) {
|
||||
sb.append(filler)
|
||||
if (placed < mentions && sb.length >= (placed + 1).toLong() * stride) {
|
||||
sb.append("nostr:").append(if (placed % 2 == 0) NPUB else NEVENT).append(' ')
|
||||
placed++
|
||||
}
|
||||
}
|
||||
while (placed < mentions) {
|
||||
sb.append("nostr:").append(if (placed % 2 == 0) NPUB else NEVENT).append(' ')
|
||||
placed++
|
||||
}
|
||||
return sb.toString()
|
||||
}
|
||||
|
||||
private val PREFIXES =
|
||||
arrayOf("npub1", "nsec1", "note1", "nevent1", "naddr1", "nprofile1", "nrelay1", "nembed1")
|
||||
|
||||
/**
|
||||
* One O(n) pass over the chars: stop at every 'n'/'N' and test the 8 entity
|
||||
* prefixes with regionMatches(ignoreCase). Cheap char compares instead of
|
||||
* driving the regex NFA from every position in the string.
|
||||
*/
|
||||
fun hasNip19Candidate(content: String): Boolean {
|
||||
var i = 0
|
||||
val n = content.length
|
||||
while (i < n) {
|
||||
val c = content[i]
|
||||
if (c == 'n' || c == 'N') {
|
||||
for (p in PREFIXES) {
|
||||
if (content.regionMatches(i, p, 0, p.length, ignoreCase = true)) return true
|
||||
}
|
||||
}
|
||||
i++
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
/**
|
||||
* Variant 2: never let the NFA scan. Walk to each candidate position with
|
||||
* cheap char compares, then apply the regex ANCHORED there via matchAt.
|
||||
* `(nostr:)?@?` are optional, so anchoring at the entity's 'n' still matches
|
||||
* and still captures the type/key/trailing groups parseAll uses.
|
||||
*/
|
||||
fun anchoredNip19(content: String): Int {
|
||||
var i = 0
|
||||
var hits = 0
|
||||
val n = content.length
|
||||
while (i < n) {
|
||||
val c = content[i]
|
||||
if (c == 'n' || c == 'N') {
|
||||
var candidate = false
|
||||
for (p in PREFIXES) {
|
||||
if (content.regionMatches(i, p, 0, p.length, ignoreCase = true)) {
|
||||
candidate = true
|
||||
break
|
||||
}
|
||||
}
|
||||
if (candidate) {
|
||||
val m = Nip19Parser.nip19regex.matchAt(content, i)
|
||||
if (m != null) {
|
||||
hits++
|
||||
i = m.range.last + 1
|
||||
continue
|
||||
}
|
||||
}
|
||||
}
|
||||
i++
|
||||
}
|
||||
return hits
|
||||
}
|
||||
|
||||
/**
|
||||
* Candidate: the regex requires `(?:\s|\A)` immediately before the `#`, so
|
||||
* every match starts at a whitespace (or at position 0). Jump between '#'
|
||||
* occurrences with indexOf and anchor the regex there.
|
||||
*/
|
||||
fun fastFindHashtags(content: String): List<String> {
|
||||
if (content.isBlank()) return emptyList()
|
||||
val out = mutableSetOf<String>()
|
||||
var h = content.indexOf('#')
|
||||
while (h >= 0) {
|
||||
if (h == 0 || content[h - 1].isWhitespace()) {
|
||||
val m = hashtagSearch.matchAt(content, if (h == 0) 0 else h - 1)
|
||||
if (m != null) {
|
||||
m.groups[1]?.value?.let { if (it.isNotBlank()) out.add(it) }
|
||||
h = content.indexOf('#', m.range.last + 1)
|
||||
continue
|
||||
}
|
||||
}
|
||||
h = content.indexOf('#', h + 1)
|
||||
}
|
||||
return out.toList()
|
||||
}
|
||||
|
||||
/**
|
||||
* REFERENCE: the pre-optimization `findHashtags`, verbatim. Production is
|
||||
* compared against this so the guard can never become a tautology.
|
||||
*/
|
||||
fun referenceFindHashtags(content: String): List<String> {
|
||||
if (content.isBlank()) return emptyList()
|
||||
val output = mutableSetOf<String>()
|
||||
hashtagSearch.findAll(content).forEach {
|
||||
try {
|
||||
val tag = it.groups[1]?.value
|
||||
if (tag != null && tag.isNotBlank()) output.add(tag)
|
||||
} catch (e: Exception) {
|
||||
}
|
||||
}
|
||||
return output.toList()
|
||||
}
|
||||
|
||||
/** REFERENCE: pre-optimization nip19 scan (findAll), resolved via the public uriToRoute. */
|
||||
fun referenceNip19(content: String): List<String> =
|
||||
Nip19Parser.nip19regex
|
||||
.findAll(content)
|
||||
.mapNotNull { m ->
|
||||
val type = m.groups[3]?.value ?: m.groups[5]?.value
|
||||
val key = m.groups[4]?.value ?: m.groups[6]?.value
|
||||
if (type != null && key != null) {
|
||||
Nip19Parser.uriToRoute(type + key)?.entity?.let { type + key }
|
||||
} else {
|
||||
null
|
||||
}
|
||||
}.toList()
|
||||
|
||||
/** A tag array with [n] p/e/a entries, for the #[index] scans. */
|
||||
fun tagArray(n: Int): Array<Array<String>> =
|
||||
Array(n) { i ->
|
||||
when (i % 3) {
|
||||
0 -> arrayOf("p", "%064x".format(i))
|
||||
1 -> arrayOf("e", "%064x".format(i + 1000))
|
||||
else -> arrayOf("a", "30023:%064x:slug$i".format(i))
|
||||
}
|
||||
}
|
||||
|
||||
/** A note with [refs] legacy #[n] references, padded to ~[targetBytes]. */
|
||||
fun indexNote(
|
||||
targetBytes: Int,
|
||||
refs: Int,
|
||||
tagCount: Int,
|
||||
): String {
|
||||
val filler = "Legacy index refs used to link people and events inline in the content. "
|
||||
val sb = StringBuilder(targetBytes + 4096)
|
||||
var placed = 0
|
||||
val stride = if (refs > 0) targetBytes / (refs + 1) else Int.MAX_VALUE
|
||||
while (sb.length < targetBytes) {
|
||||
sb.append(filler)
|
||||
if (placed < refs && sb.length >= (placed + 1).toLong() * stride) {
|
||||
sb.append("#[").append(placed % tagCount).append("] ")
|
||||
placed++
|
||||
}
|
||||
}
|
||||
while (placed < refs) {
|
||||
sb.append("#[").append(placed % tagCount).append("] ")
|
||||
placed++
|
||||
}
|
||||
return sb.toString()
|
||||
}
|
||||
|
||||
/** REFERENCE: pre-optimization findIndexTagsWithPeople, verbatim. */
|
||||
fun referenceIndexPeople(
|
||||
content: String,
|
||||
tags: Array<Array<String>>,
|
||||
): List<String> {
|
||||
val output = mutableSetOf<String>()
|
||||
tagSearch.findAll(content).forEach { index ->
|
||||
try {
|
||||
val tag = index.groups[1]?.value?.let { tags[it.toInt()] }
|
||||
if (tag != null && tag.size > 1 && tag[0] == "p") output.add(tag[1])
|
||||
} catch (e: Exception) {
|
||||
}
|
||||
}
|
||||
return output.toList()
|
||||
}
|
||||
|
||||
/** REFERENCE: pre-optimization findIndexTagsWithEventsOrAddresses, verbatim. */
|
||||
fun referenceIndexEvents(
|
||||
content: String,
|
||||
tags: Array<Array<String>>,
|
||||
): Set<String> {
|
||||
val output = mutableSetOf<String>()
|
||||
tagSearch.findAll(content).forEach { index ->
|
||||
try {
|
||||
val tag = index.groups[1]?.value?.let { tags[it.toInt()] }
|
||||
if (tag != null && tag.size > 1 && tag[0] == "e") output.add(tag[1])
|
||||
if (tag != null && tag.size > 1 && tag[0] == "a") output.add(tag[1])
|
||||
} catch (e: Exception) {
|
||||
}
|
||||
}
|
||||
return output
|
||||
}
|
||||
|
||||
fun bench(
|
||||
label: String,
|
||||
input: String,
|
||||
reps: Int,
|
||||
op: (String) -> Int,
|
||||
) {
|
||||
repeat(maxOf(reps / 4, 2)) { op(input) } // warmup
|
||||
val t0 = System.nanoTime()
|
||||
var sink = 0
|
||||
repeat(reps) { sink += op(input) }
|
||||
val ns = (System.nanoTime() - t0) / reps
|
||||
val mbps = input.length.toDouble() / ns * 1000.0 // bytes/ns -> MB/s
|
||||
println(
|
||||
String.format(
|
||||
"%-34s %9d B %9d ns/op %8.1f MB/s (hits=%d)",
|
||||
label,
|
||||
input.length,
|
||||
ns,
|
||||
mbps,
|
||||
sink / reps,
|
||||
),
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
fun contentScans() {
|
||||
// (bytes, mentions, reps) — mirrors the measured distribution
|
||||
val corpus =
|
||||
listOf(
|
||||
Triple(120, 1, 20_000), // short note
|
||||
Triple(529, 2, 20_000), // MEDIAN of what matchers held
|
||||
Triple(4_000, 5, 5_000), // long note
|
||||
Triple(68_000, 40, 200), // v1.13.0 release notes
|
||||
Triple(767_000, 120, 20), // observed MAX
|
||||
)
|
||||
|
||||
// Guard the corpus: a benchmark that silently matches nothing measures nothing.
|
||||
corpus.forEach { (n, m, _) ->
|
||||
val found = Nip19Parser.parseAll(note(n, m)).size
|
||||
check(found >= m) { "corpus broken: ${n}B/$m mentions parsed only $found entities" }
|
||||
}
|
||||
|
||||
println("\n=== nip19 findNostrUris — NO matches (the common case) ===")
|
||||
corpus.forEach { (n, _, r) ->
|
||||
bench("nip19 0 mentions", note(n, 0), r) { findNostrUris(it).size }
|
||||
}
|
||||
|
||||
println("\n=== nip19 findNostrUris — WITH matches (Matcher per match) ===")
|
||||
corpus.forEach { (n, m, r) ->
|
||||
bench("nip19 m=$m", note(n, m), r) { findNostrUris(it).size }
|
||||
}
|
||||
|
||||
println("\n=== findHashtags (ContentHashTags.hashtagSearch) ===")
|
||||
corpus.forEach { (n, m, r) ->
|
||||
bench("hashtags m=$m", note(n, m), r) { findHashtags(it).size }
|
||||
}
|
||||
|
||||
println("\n=== OPTIMIZED: literal pre-scan, then regex only if a candidate exists ===")
|
||||
corpus.forEach { (n, _, r) ->
|
||||
bench("opt nip19 0 mentions", note(n, 0), r) { if (hasNip19Candidate(it)) findNostrUris(it).size else 0 }
|
||||
}
|
||||
corpus.forEach { (n, m, r) ->
|
||||
bench("opt nip19 m=$m", note(n, m), r) { if (hasNip19Candidate(it)) findNostrUris(it).size else 0 }
|
||||
}
|
||||
|
||||
println("\n=== findHashtags — content with NO '#' ===")
|
||||
corpus.forEach { (n, _, r) ->
|
||||
bench("hashtags none", note(n, 0, hashtags = false), r) { findHashtags(it).size }
|
||||
}
|
||||
println("\n=== findHashtags OPTIMIZED (indexOf('#') + anchored) ===")
|
||||
corpus.forEach { (n, _, r) ->
|
||||
bench("opt hashtags none", note(n, 0, hashtags = false), r) { fastFindHashtags(it).size }
|
||||
}
|
||||
corpus.forEach { (n, m, r) ->
|
||||
bench("opt hashtags dense", note(n, m), r) { fastFindHashtags(it).size }
|
||||
}
|
||||
|
||||
println("\n=== IndexedTags findIndexTagsWithPeople ===")
|
||||
val tags = tagArray(30)
|
||||
corpus.forEach { (n, _, r) ->
|
||||
bench("idxTags none", indexNote(n, 0, 30), r) { findIndexTagsWithPeople(it, tags).size }
|
||||
}
|
||||
corpus.forEach { (n, m, r) ->
|
||||
bench("idxTags refs=$m", indexNote(n, m, 30), r) { findIndexTagsWithPeople(it, tags).size }
|
||||
}
|
||||
|
||||
println("\n=== RENDER PATH: RichTextParser.parseText (per note, per render) ===")
|
||||
val rtTags = tagArray(30).toImmutableListOfLists()
|
||||
val rtCorpus =
|
||||
listOf(
|
||||
Triple(120, 1, 3_000),
|
||||
Triple(529, 2, 3_000),
|
||||
Triple(4_000, 5, 500),
|
||||
Triple(68_000, 40, 20),
|
||||
)
|
||||
rtCorpus.forEach { (n, m, r) ->
|
||||
bench("parseText hashtags m=$m", note(n, m), r) {
|
||||
RichTextParser().parseText(it, rtTags, null).paragraphs.size
|
||||
}
|
||||
}
|
||||
rtCorpus.forEach { (n, _, r) ->
|
||||
bench("parseText no '#'", note(n, 0, hashtags = false), r) {
|
||||
RichTextParser().parseText(it, rtTags, null).paragraphs.size
|
||||
}
|
||||
}
|
||||
rtCorpus.forEach { (n, m, r) ->
|
||||
bench("parseText #[n] refs", indexNote(n, m, 30), r) {
|
||||
RichTextParser().parseText(it, rtTags, null).paragraphs.size
|
||||
}
|
||||
}
|
||||
|
||||
println("\n=== OPTIMIZED v2: anchored matchAt at candidate positions (no NFA scan) ===")
|
||||
corpus.forEach { (n, _, r) ->
|
||||
bench("v2 nip19 0 mentions", note(n, 0), r) { anchoredNip19(it) }
|
||||
}
|
||||
corpus.forEach { (n, m, r) ->
|
||||
bench("v2 nip19 m=$m", note(n, m), r) { anchoredNip19(it) }
|
||||
}
|
||||
}
|
||||
|
||||
/** The fast hashtag scan must return exactly what the production scan returns. */
|
||||
@Test
|
||||
fun fastHashtagsMatchesProduction() {
|
||||
val cases =
|
||||
listOf(
|
||||
"#one #two #three",
|
||||
"no hashtags here",
|
||||
"#start of line",
|
||||
"mid #tag then more",
|
||||
"email a@b.com and #tag",
|
||||
"punct #tag, #tag2. #tag3!",
|
||||
"nospace#nottag should not match",
|
||||
"\n#afterNewline",
|
||||
"",
|
||||
" ",
|
||||
note(4_000, 0),
|
||||
note(20_000, 3),
|
||||
note(4_000, 0, hashtags = false),
|
||||
)
|
||||
cases.forEach { c ->
|
||||
val ref = referenceFindHashtags(c).sorted()
|
||||
val prod = findHashtags(c).sorted()
|
||||
check(ref == prod) { "mismatch on ${c.take(40).replace("\n", "\\n")}: reference=$ref production=$prod" }
|
||||
}
|
||||
|
||||
// fuzz: random text peppered with '#' in awkward positions
|
||||
val rnd = kotlin.random.Random(20260813)
|
||||
val alphabet = " \n\t#abcXYZ.,!@\u00a0-_0123"
|
||||
repeat(3_000) {
|
||||
val c = (1..rnd.nextInt(0, 80)).map { alphabet[rnd.nextInt(alphabet.length)] }.joinToString("")
|
||||
val ref = referenceFindHashtags(c).sorted()
|
||||
val prod = findHashtags(c).sorted()
|
||||
check(ref == prod) { "fuzz mismatch on ${c.replace("\n", "\\n")}: reference=$ref production=$prod" }
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Isolated A/B for the `#[` gate, interleaved with medians.
|
||||
*
|
||||
* parseText is allocation-heavy and its whole-parse timings swing ±50% run to
|
||||
* run (the unchanged no-'#' control arm moved that much), which is far larger
|
||||
* than this effect — so measure the gated call on its own instead.
|
||||
*/
|
||||
@Test
|
||||
fun hashGateIsolatedAb() {
|
||||
// realistic mix: mostly plain hashtags, a few legacy #[n] refs
|
||||
val words =
|
||||
buildList {
|
||||
repeat(90) { add("#hashtag$it") }
|
||||
repeat(10) { add("#[$it]") }
|
||||
}
|
||||
|
||||
fun ungated(): Int {
|
||||
var n = 0
|
||||
words.forEach { if (tagIndex.find(it) != null) n++ }
|
||||
return n
|
||||
}
|
||||
|
||||
fun gated(): Int {
|
||||
var n = 0
|
||||
words.forEach { if (it.contains("#[") && tagIndex.find(it) != null) n++ }
|
||||
return n
|
||||
}
|
||||
check(ungated() == gated()) { "gate changed the result" }
|
||||
repeat(2_000) {
|
||||
ungated()
|
||||
gated()
|
||||
} // warmup both
|
||||
val a = mutableListOf<Long>()
|
||||
val b = mutableListOf<Long>()
|
||||
repeat(21) {
|
||||
var t = System.nanoTime()
|
||||
repeat(2_000) { ungated() }
|
||||
a.add(System.nanoTime() - t)
|
||||
t = System.nanoTime()
|
||||
repeat(2_000) { gated() }
|
||||
b.add(System.nanoTime() - t)
|
||||
}
|
||||
a.sort()
|
||||
b.sort()
|
||||
val am = a[a.size / 2] / 2_000
|
||||
val bm = b[b.size / 2] / 2_000
|
||||
println(
|
||||
"\nHASH GATE (100 words: 90 plain hashtags + 10 #[n]) median of 21:" +
|
||||
"\n ungated tagIndex.find : $am ns/word-set" +
|
||||
"\n gated with contains : $bm ns/word-set" +
|
||||
"\n speedup : ${String.format("%.2f", am.toDouble() / bm)}x",
|
||||
)
|
||||
}
|
||||
|
||||
/**
|
||||
* Segment-level guard for the RichTextParser '#' path.
|
||||
*
|
||||
* Expectations are HARDCODED from the behaviour before the `#[` gate was added,
|
||||
* so this pins the parser against the code it replaced rather than against itself.
|
||||
*/
|
||||
@Test
|
||||
fun richTextHashSegmentsUnchanged() {
|
||||
val tags = tagArray(30).toImmutableListOfLists()
|
||||
val expected =
|
||||
listOf(
|
||||
"#Nostr" to "HashTagSegment|#Nostr",
|
||||
"#[0]" to "HashIndexUserSegment|#[0]",
|
||||
"#[1]suffix" to "HashIndexEventSegment|#[1]suffix",
|
||||
"#[999]" to "RegularTextSegment|#[999]",
|
||||
"#[abc]" to "RegularTextSegment|#[abc]",
|
||||
"#[]" to "RegularTextSegment|#[]",
|
||||
"#" to "RegularTextSegment|#",
|
||||
"##tag" to "HashTagSegment|##tag",
|
||||
"#tag," to "HashTagSegment|#tag,",
|
||||
"#tag." to "HashTagSegment|#tag.",
|
||||
"#a#b" to "HashTagSegment|#a#b",
|
||||
"#[2]#tail" to "HashIndexEventSegment|#[2]#tail",
|
||||
"#\u00e9t\u00e9" to "HashTagSegment|#\u00e9t\u00e9",
|
||||
"#123" to "HashTagSegment|#123",
|
||||
"#[0]!" to "HashIndexUserSegment|#[0]!",
|
||||
"plain" to "RegularTextSegment|plain",
|
||||
"#tag)" to "HashTagSegment|#tag)",
|
||||
"#-dash" to "HashTagSegment|#-dash",
|
||||
"#_under" to "HashTagSegment|#_under",
|
||||
// '#[' appearing mid-word must still reach the index parser
|
||||
"#a#[0]" to "HashIndexUserSegment|#a#[0]",
|
||||
)
|
||||
expected.forEach { (word, want) ->
|
||||
val got =
|
||||
RichTextParser()
|
||||
.parseText(word, tags, null)
|
||||
.paragraphs
|
||||
.flatMap { p -> p.words.map { it::class.simpleName + "|" + it.segmentText } }
|
||||
.joinToString(";")
|
||||
check(got == want) { "segment mismatch for '$word': want=$want got=$got" }
|
||||
}
|
||||
}
|
||||
|
||||
/** Both IndexedTags scans must equal the findAll-based references. */
|
||||
@Test
|
||||
fun indexedTagsMatchReference() {
|
||||
val tags = tagArray(30)
|
||||
val cases =
|
||||
listOf(
|
||||
"#[0] #[1] #[2]",
|
||||
"no refs here",
|
||||
"#[0] at start",
|
||||
"mid #[5] ref",
|
||||
"nospace#[3] should not match",
|
||||
"\n#[7] after newline",
|
||||
"#[999] out of range",
|
||||
"#[abc] not numeric",
|
||||
"#[] empty",
|
||||
"",
|
||||
" ",
|
||||
indexNote(2_000, 0, 30),
|
||||
indexNote(4_000, 6, 30),
|
||||
indexNote(20_000, 20, 30),
|
||||
)
|
||||
cases.forEach { c ->
|
||||
check(referenceIndexPeople(c, tags).sorted() == findIndexTagsWithPeople(c, tags).sorted()) {
|
||||
"people mismatch on ${c.take(40).replace("\n", "\\n")}"
|
||||
}
|
||||
check(referenceIndexEvents(c, tags).sorted() == findIndexTagsWithEventsOrAddresses(c, tags).sorted()) {
|
||||
"events mismatch on ${c.take(40).replace("\n", "\\n")}"
|
||||
}
|
||||
}
|
||||
val rnd = kotlin.random.Random(20260813)
|
||||
val alphabet = " \n#[]0123456789abc\u00a0"
|
||||
repeat(3_000) {
|
||||
val c = (1..rnd.nextInt(0, 60)).map { alphabet[rnd.nextInt(alphabet.length)] }.joinToString("")
|
||||
check(referenceIndexPeople(c, tags).sorted() == findIndexTagsWithPeople(c, tags).sorted()) {
|
||||
"fuzz people mismatch on ${c.replace("\n", "\\n")}"
|
||||
}
|
||||
check(referenceIndexEvents(c, tags).sorted() == findIndexTagsWithEventsOrAddresses(c, tags).sorted()) {
|
||||
"fuzz events mismatch on ${c.replace("\n", "\\n")}"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/** Production parseAll must equal the findAll-based reference scan. */
|
||||
@Test
|
||||
fun nip19MatchesReferenceScan() {
|
||||
val cases =
|
||||
listOf(
|
||||
"nostr:$NPUB",
|
||||
"@$NPUB",
|
||||
NPUB,
|
||||
"text nostr:$NEVENT tail",
|
||||
"two $NPUB and nostr:$NEVENT here",
|
||||
"none at all",
|
||||
"npub1tooshort",
|
||||
"",
|
||||
note(200, 1),
|
||||
note(4_000, 5),
|
||||
note(20_000, 12),
|
||||
note(4_000, 0),
|
||||
// every prefix standalone
|
||||
"npub1",
|
||||
"nsec1",
|
||||
"note1",
|
||||
"nevent1",
|
||||
"naddr1",
|
||||
"nprofile1",
|
||||
"nrelay1",
|
||||
"nembed1",
|
||||
// 'n' as the LAST char: the second-char dispatch must not read past the end
|
||||
"n",
|
||||
"N",
|
||||
"ends with n",
|
||||
"trailing N",
|
||||
"nn",
|
||||
"np",
|
||||
"ne",
|
||||
"na",
|
||||
"nr",
|
||||
"ns",
|
||||
"no",
|
||||
// case + adjacency
|
||||
"NOSTR:" + NPUB.uppercase(),
|
||||
"x" + NPUB,
|
||||
NPUB + NEVENT,
|
||||
NPUB + " " + NEVENT,
|
||||
)
|
||||
cases.forEach { c ->
|
||||
val ref = referenceNip19(c)
|
||||
val prod = Nip19Parser.parseAll(c).size
|
||||
check(ref.size == prod) { "nip19 mismatch on ${c.take(50)}: reference=${ref.size} production=$prod" }
|
||||
}
|
||||
}
|
||||
|
||||
/** v2 must find exactly the same number of entities as the production scan. */
|
||||
@Test
|
||||
fun anchoredMatchesProductionCount() {
|
||||
listOf(0, 1, 2, 5, 40).forEach { m ->
|
||||
listOf(200, 4_000, 68_000).forEach { size ->
|
||||
val c = note(size, m)
|
||||
val real = Nip19Parser.nip19regex.findAll(c).count()
|
||||
val fast = anchoredNip19(c)
|
||||
check(real == fast) { "size=$size m=$m: production found $real, anchored found $fast" }
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Correctness gate for the pre-scan: it must never hide a real match.
|
||||
* Anything the regex finds must also be reported as a candidate.
|
||||
*/
|
||||
@Test
|
||||
fun preScanNeverMissesAMatch() {
|
||||
val cases =
|
||||
listOf(
|
||||
"nostr:$NPUB",
|
||||
"@$NPUB",
|
||||
NPUB,
|
||||
"text before nostr:$NEVENT and after",
|
||||
"UPPER ${NPUB.uppercase()} case",
|
||||
"no entities here at all, just prose about nostr and npubs",
|
||||
"npub1tooshort",
|
||||
"",
|
||||
)
|
||||
cases.forEach { c ->
|
||||
val real = Nip19Parser.parseAll(c).isNotEmpty()
|
||||
if (real) check(hasNip19Candidate(c)) { "pre-scan missed a real match in: ${c.take(40)}" }
|
||||
}
|
||||
// and it must actually filter the common case
|
||||
check(!hasNip19Candidate(note(4_000, 0))) { "pre-scan failed to reject prose with no entities" }
|
||||
}
|
||||
}
|
||||
+31
-7
@@ -22,21 +22,45 @@ package com.vitorpamplona.quartz.nip10Notes.content
|
||||
|
||||
val hashtagSearch = Regex("(?:\\s|\\A)#([^\\s!@#\$%^&*()=+./,\\[{\\]};:'\"?><]+)")
|
||||
|
||||
/**
|
||||
* Collects the hashtags in [content].
|
||||
*
|
||||
* [hashtagSearch] requires `(?:\s|\A)` immediately before the `#`, so every match
|
||||
* starts either at position 0 or at a whitespace. That lets the scan jump between
|
||||
* `#` occurrences with `indexOf` — an intrinsified char search — and apply the
|
||||
* regex **anchored** at each, instead of letting `findAll` drive the regex engine
|
||||
* from every position in the string.
|
||||
*
|
||||
* Measured on the production content distribution (median 529 B, tail to 767 KB):
|
||||
* ~68 MB/s -> ~1,240 MB/s on hashtag-dense text (18x) and ~19,000 MB/s when the
|
||||
* content has no `#` at all (up to 300x). Equivalence with the previous `findAll`
|
||||
* implementation is guarded by `RegexContentBenchmark` in `commons`.
|
||||
*/
|
||||
fun findHashtags(
|
||||
content: String,
|
||||
output: MutableSet<String> = mutableSetOf(),
|
||||
): List<String> {
|
||||
if (content.isBlank()) return emptyList()
|
||||
|
||||
val matcher = hashtagSearch.findAll(content)
|
||||
matcher.forEach {
|
||||
try {
|
||||
val tag = it.groups[1]?.value
|
||||
if (tag != null && tag.isNotBlank()) {
|
||||
output.add(tag)
|
||||
var h = content.indexOf('#')
|
||||
while (h >= 0) {
|
||||
if (h == 0 || content[h - 1].isWhitespace()) {
|
||||
val match =
|
||||
try {
|
||||
hashtagSearch.matchAt(content, if (h == 0) 0 else h - 1)
|
||||
} catch (e: Exception) {
|
||||
null
|
||||
}
|
||||
if (match != null) {
|
||||
val tag = match.groups[1]?.value
|
||||
if (tag != null && tag.isNotBlank()) {
|
||||
output.add(tag)
|
||||
}
|
||||
h = content.indexOf('#', match.range.last + 1)
|
||||
continue
|
||||
}
|
||||
} catch (e: Exception) {
|
||||
}
|
||||
h = content.indexOf('#', h + 1)
|
||||
}
|
||||
return output.toList()
|
||||
}
|
||||
|
||||
+39
-4
@@ -28,6 +28,43 @@ import com.vitorpamplona.quartz.nip01Core.core.TagArray
|
||||
|
||||
val tagSearch = Regex("(?:\\s|\\A)\\#\\[([0-9]+)\\]")
|
||||
|
||||
/**
|
||||
* Walks every `#[n]` reference in [content].
|
||||
*
|
||||
* [tagSearch] requires `(?:\s|\A)` immediately before the `#`, so every match
|
||||
* starts at position 0 or at a whitespace. That lets the scan jump between `#`
|
||||
* occurrences with `indexOf` — an intrinsified char search — and apply the regex
|
||||
* **anchored** at each, instead of letting `findAll` drive the regex engine from
|
||||
* every position in the string.
|
||||
*
|
||||
* Measured on the production content distribution (median 529 B, tail to 767 KB):
|
||||
* ~63 MB/s -> multiple GB/s when the content has no `#`, and ~18x on reference-dense
|
||||
* text. Equivalence with the previous `findAll` implementation (both callers) is
|
||||
* guarded by `RegexContentBenchmark` in `commons`.
|
||||
*/
|
||||
private inline fun forEachIndexTag(
|
||||
content: String,
|
||||
action: (MatchResult) -> Unit,
|
||||
) {
|
||||
var h = content.indexOf('#')
|
||||
while (h >= 0) {
|
||||
if (h == 0 || content[h - 1].isWhitespace()) {
|
||||
val match =
|
||||
try {
|
||||
tagSearch.matchAt(content, if (h == 0) 0 else h - 1)
|
||||
} catch (e: Exception) {
|
||||
null
|
||||
}
|
||||
if (match != null) {
|
||||
action(match)
|
||||
h = content.indexOf('#', match.range.last + 1)
|
||||
continue
|
||||
}
|
||||
}
|
||||
h = content.indexOf('#', h + 1)
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Returns the old-style [1] tag that pionts to an index in the tag array
|
||||
*/
|
||||
@@ -36,8 +73,7 @@ fun findIndexTagsWithPeople(
|
||||
tags: TagArray,
|
||||
output: MutableSet<String> = mutableSetOf<String>(),
|
||||
): List<String> {
|
||||
val matcher = tagSearch.findAll(content)
|
||||
matcher.forEach { index ->
|
||||
forEachIndexTag(content) { index ->
|
||||
try {
|
||||
val tag = index.groups[1]?.value?.let { tags[it.toInt()] }
|
||||
if (tag != null && tag.size > 1 && tag[0] == "p") {
|
||||
@@ -58,8 +94,7 @@ fun findIndexTagsWithEventsOrAddresses(
|
||||
tags: TagArray,
|
||||
output: MutableSet<String> = mutableSetOf<String>(),
|
||||
): Set<String> {
|
||||
val matcher = tagSearch.findAll(content)
|
||||
matcher.forEach { index ->
|
||||
forEachIndexTag(content) { index ->
|
||||
try {
|
||||
val tag = index.groups[1]?.value?.let { tags[it.toInt()] }
|
||||
if (tag != null && tag.size > 1 && tag[0] == "e") {
|
||||
|
||||
@@ -133,21 +133,71 @@ object Nip19Parser {
|
||||
|
||||
fun hasAny(content: String): Boolean = nip19regex.matches(content)
|
||||
|
||||
/**
|
||||
* True when one of [nip19regex]'s entity prefixes starts at [i].
|
||||
*
|
||||
* Every prefix begins with `n`, and English prose is ~7% `n`, so testing all
|
||||
* eight prefixes at each `n` is the bulk of the scan's cost. Dispatching on the
|
||||
* second character first narrows it to at most two candidates — most `n`s are
|
||||
* rejected by a single char compare.
|
||||
*/
|
||||
private fun isCandidateAt(
|
||||
content: String,
|
||||
i: Int,
|
||||
): Boolean {
|
||||
val c = content[i]
|
||||
if (c != 'n' && c != 'N') return false
|
||||
if (i + 1 >= content.length) return false
|
||||
return when (content[i + 1]) {
|
||||
'p', 'P' ->
|
||||
content.regionMatches(i, "npub1", 0, 5, ignoreCase = true) ||
|
||||
content.regionMatches(i, "nprofile1", 0, 9, ignoreCase = true)
|
||||
's', 'S' -> content.regionMatches(i, "nsec1", 0, 5, ignoreCase = true)
|
||||
'o', 'O' -> content.regionMatches(i, "note1", 0, 5, ignoreCase = true)
|
||||
'e', 'E' ->
|
||||
content.regionMatches(i, "nevent1", 0, 7, ignoreCase = true) ||
|
||||
content.regionMatches(i, "nembed1", 0, 7, ignoreCase = true)
|
||||
'a', 'A' -> content.regionMatches(i, "naddr1", 0, 6, ignoreCase = true)
|
||||
'r', 'R' -> content.regionMatches(i, "nrelay1", 0, 7, ignoreCase = true)
|
||||
else -> false
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Scans [content] for NIP-19 entities.
|
||||
*
|
||||
* Walks to each candidate position with cheap char compares and applies
|
||||
* [nip19regex] **anchored** there, instead of letting `findAll` drive the
|
||||
* regex engine from every position in the string. `(nostr:)?@?` are optional,
|
||||
* so anchoring at the entity's own `n` matches the same entities and captures
|
||||
* the same type/key/trailing groups this function reads.
|
||||
*
|
||||
* Measured 9–23x faster across the production content distribution (median
|
||||
* 529 B, tail to 767 KB): ~19 MB/s -> 170–445 MB/s. Equivalence and speed are
|
||||
* guarded by `RegexContentBenchmark` in `commons`. Motivation: on an SM-T220
|
||||
* heap dump, 2,541 of 4,573 live Matchers were running this regex.
|
||||
*/
|
||||
fun parseAll(content: String): List<Entity> {
|
||||
val matchSequence = nip19regex.findAll(content)
|
||||
val returningList = mutableListOf<Entity>()
|
||||
matchSequence.forEach { matcher ->
|
||||
val type = matcher.groups[3]?.value ?: matcher.groups[5]?.value // npub1
|
||||
val key = matcher.groups[4]?.value ?: matcher.groups[6]?.value // bech32
|
||||
val additionalChars = matcher.groups[7]?.value // additional chars
|
||||
var i = 0
|
||||
val len = content.length
|
||||
while (i < len) {
|
||||
if (isCandidateAt(content, i)) {
|
||||
val matcher = nip19regex.matchAt(content, i)
|
||||
if (matcher != null) {
|
||||
val type = matcher.groups[3]?.value ?: matcher.groups[5]?.value // npub1
|
||||
val key = matcher.groups[4]?.value ?: matcher.groups[6]?.value // bech32
|
||||
val additionalChars = matcher.groups[7]?.value // additional chars
|
||||
|
||||
if (type != null) {
|
||||
val parsed = parseComponents(type, key, additionalChars)?.entity
|
||||
if (type != null) {
|
||||
parseComponents(type, key, additionalChars)?.entity?.let { returningList.add(it) }
|
||||
}
|
||||
|
||||
if (parsed != null) {
|
||||
returningList.add(parsed)
|
||||
i = matcher.range.last + 1
|
||||
continue
|
||||
}
|
||||
}
|
||||
i++
|
||||
}
|
||||
return returningList
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user