perf(quartz): one pass for list search's other tags, ASCII whitespace fast path

- The `t` pass and the natural-language pass are one walk over the tags, so
  `t` values now come in tag order among the others (golden re-recorded; the
  set of fields per event is unchanged on all 1.5M corpus events).
- The extra-field walk is private; the two text builders share one join.
- Whitespace checks take an ASCII fast path instead of the JVM's two Unicode
  table lookups per character (identical results on every corpus value).

Corpus benchmark: classifier ~400 -> ~300 ms per 3.7M values, walk ~490 ->
~365 ms per 300k events.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016XJSf4ekFVuhSj5m8eeNsP
This commit is contained in:
Claude
2026-10-05 13:37:32 +00:00
parent 9de2368337
commit 3b0120e5bf
5 changed files with 71 additions and 68 deletions
@@ -65,7 +65,7 @@ Separator legend: **NL** = `joinToString("\n")`, **SP** = `joinToString(" ")`.
| 9736 | Bolt12ZapEvent | nipB1Bolt12Zaps/zap | `content` |
| 9737 | Bolt12ZapIntentEvent | nipB1Bolt12Zaps/intent | `content` |
| 9802 | HighlightEvent | nip84Highlights | `listOfNotNull(comment(), context(), content)` NL |
| 9998 | ListHeaderEvent | experimental/decentralizedLists/header | `tags.searchableListContent()` NL — `names` (singular, plural), `titles` (singular, plural), `name`, `title`, `description`, `comments`, then every `t` value that is not `isMachineValue`, then every value of every other tag (except `alt`, `client`, `imeta`) that passes `isNaturalLanguageValue` (has whitespace or non-ASCII, or is one capitalized letters-only word; never JSON, numbers, URIs of any scheme, addresses, hex ids, UUIDs, bech32) |
| 9998 | ListHeaderEvent | experimental/decentralizedLists/header | `tags.searchableListContent()` NL — `names` (singular, plural), `titles` (singular, plural), `name`, `title`, `description`, `comments`, then in tag order every `t` value that is not `isMachineValue` and every value of every other tag (except `alt`, `client`, `imeta`) that passes `isNaturalLanguageValue` (has whitespace or non-ASCII, or is one capitalized letters-only word; never JSON, numbers, URIs of any scheme, addresses, hex ids, UUIDs, bech32) |
| 9999 | ListItemEvent | experimental/decentralizedLists/item | same as 9998 |
| 10003 | BookmarkListEvent | nip51Lists/bookmarkList | `listOfNotNull(title())` NL |
| 10100 | AgentProfileEvent | buzz/agentProfiles | `profileOrNull()?.let { listOfNotNull(it.name, it.displayName).joinToString("\n") } ?: ""` |
@@ -85,8 +85,8 @@ Filter(
All four kinds are `SearchableEvent`s sharing one walk
(`forEachSearchableListField`): `names` and `titles` (singular then plural),
`name`, `title`, `description`, `comments`, then every `t` item value, then
every value of every other tag that reads as natural language.
`name`, `title`, `description`, `comments`, then, in tag order, every `t`
item value and every value of every other tag that reads as natural language.
The tag set is open — deployments add `author`, `subject`, `artist`,
`relationshipType` and tags nobody has named yet — so that last step decides
@@ -43,11 +43,12 @@ import com.vitorpamplona.quartz.nip89AppHandlers.clientTag.ClientTag
import com.vitorpamplona.quartz.nip92IMeta.IMetaTag
/**
* The human-authored text of any kind in the family, in a fixed order: the header's `names`
* and `titles` (singular, then plural), the item's `name` and `title`, `description`,
* `comments`, then each `t` value that is not a machine value, then
* [forEachSearchableListExtraField]. Headers and items share one walk because the spec's
* nonstandard method lets an item carry header tags; each kind simply has fewer of them set.
* The human-authored text of any kind in the family: first, in a fixed order, the header's `names`
* and `titles` (singular, then plural), the item's `name` and `title`, `description` and
* `comments`; then, in tag order, each `t` value that is not a machine value and every value of
* every other tag that reads as natural language ([forEachOtherField]). Headers and items share
* one walk because the spec's nonstandard method lets an item carry header tags; each kind simply
* has fewer of them set.
*
* `content` is not part of the spec and ids/pubkeys/coordinates are served by tag filters,
* so neither is indexed.
@@ -57,7 +58,8 @@ import com.vitorpamplona.quartz.nip92IMeta.IMetaTag
fun TagArray.forEachSearchableListField(visitor: IndexableFieldVisitor): Boolean {
// The read path runs once per event per keystroke, so this is two allocation-free passes
// instead of one full scan (and one parsed object) per field. The first pass only remembers
// the first well-formed tag of each field; the fixed visiting order is applied afterwards.
// the first well-formed tag of each named field, so their fixed order can be applied before
// the second pass visits everything else.
var names: Array<String>? = null
var titles: Array<String>? = null
var name: String? = null
@@ -89,53 +91,55 @@ fun TagArray.forEachSearchableListField(visitor: IndexableFieldVisitor): Boolean
title?.let { if (!visitor.visit(it)) return false }
description?.let { if (!visitor.visit(it)) return false }
comments?.let { if (!visitor.visit(it)) return false }
fastForEach { tag ->
HashtagTag.parse(tag)?.let { if (!isMachineValue(it) && !visitor.visit(it)) return false }
}
return forEachSearchableListExtraField(visitor)
return forEachOtherField(visitor, withHashtags = true)
}
/**
* Every value of every other tag that reads as natural language ([isNaturalLanguageValue]), in
* tag order. The family's tag set is open — deployments add `author`, `subject`, `artist`,
* `relationshipType` and tags nobody has named yet — so this walk decides by what a value looks
* like rather than by the tag it sits in.
* Every tag the first pass of [forEachSearchableListField] does not own, in tag order: each `t`
* value that is not [isMachineValue] (only [withHashtags]), and every value of every other tag
* that [isNaturalLanguageValue]. The family's tag set is open — deployments add `author`,
* `subject`, `artist`, `relationshipType` and tags nobody has named yet — so this decides by what
* a value looks like rather than by the tag it sits in.
*
* Skips the tags [forEachSearchableListField] already visits, plus the NIP-defined metadata tags
* that are not the event's own text: `alt` (NIP-31 fallback text, which here restates `title`
* and `artist` behind a fixed "Song: … by …" prefix), `client` (NIP-89 app name) and `imeta`
* (NIP-92 `key value` pairs, whose space would otherwise pass every URL and hash as text).
* Skips the NIP-defined metadata tags that are not the event's own text: `alt` (NIP-31 fallback
* text, which here restates `title` and `artist` behind a fixed "Song: … by …" prefix), `client`
* (NIP-89 app name) and `imeta` (NIP-92 `key value` pairs, whose space would otherwise pass every
* URL and hash as text).
*
* @return false when the visitor stopped the walk.
*/
fun TagArray.forEachSearchableListExtraField(visitor: IndexableFieldVisitor): Boolean {
private fun TagArray.forEachOtherField(
visitor: IndexableFieldVisitor,
withHashtags: Boolean,
): Boolean {
fastForEach { tag ->
if (tag.size < 2 || !isExtraFieldTagName(tag[0])) return@fastForEach
for (i in 1 until tag.size) {
val value = tag[i]
if (isNaturalLanguageValue(value) && !visitor.visit(value)) return false
if (tag.size < 2) return@fastForEach
when (tag[0]) {
HashtagTag.TAG_NAME -> {
if (withHashtags && !isMachineValue(tag[1]) && !visitor.visit(tag[1])) return false
}
NamesTag.TAG_NAME,
TitlesTag.TAG_NAME,
NameTag.TAG_NAME,
TitleTag.TAG_NAME,
DescriptionTag.TAG_NAME,
CommentsTag.TAG_NAME,
AltTag.TAG_NAME,
ClientTag.TAG_NAME,
IMetaTag.TAG_NAME,
-> {}
else -> {
for (i in 1 until tag.size) {
if (isNaturalLanguageValue(tag[i]) && !visitor.visit(tag[i])) return false
}
}
}
}
return true
}
private fun isExtraFieldTagName(name: String) =
when (name) {
NamesTag.TAG_NAME,
TitlesTag.TAG_NAME,
NameTag.TAG_NAME,
TitleTag.TAG_NAME,
DescriptionTag.TAG_NAME,
CommentsTag.TAG_NAME,
HashtagTag.TAG_NAME,
AltTag.TAG_NAME,
ClientTag.TAG_NAME,
IMetaTag.TAG_NAME,
-> false
else -> true
}
/**
* The TITLE role of [forEachSearchableListField], for search engines that weight fields: what the
* list or item is called — the header's `names` and `titles` (singular, then plural) and the
@@ -151,25 +155,18 @@ fun TagArray.searchableListTitles(): List<String?> {
fun TagArray.searchableListDescriptions(): List<String?> = listOf(description(), comments())
/**
* The TEXT role of [forEachSearchableListField]: the [forEachSearchableListExtraField] values, one
* per line, or null when the event has none. The `t` values are not here; the search funnel
* carries them as hashtags.
* The TEXT role of [forEachSearchableListField]: the natural-language values of the tags the
* fixed-order fields do not own, one per line, or null when the event has none. The `t` values
* are not here; the search funnel carries them as hashtags.
*/
fun TagArray.searchableListExtraText(): String? =
buildString {
forEachSearchableListExtraField { field ->
if (field != null) {
if (isNotEmpty()) append('\n')
append(field)
}
true
}
}.ifEmpty { null }
fun TagArray.searchableListExtraText(): String? = joinFields { forEachOtherField(it, withHashtags = false) }.ifEmpty { null }
/** The write-path join of [forEachSearchableListField]: one field per line. */
fun TagArray.searchableListContent() =
fun TagArray.searchableListContent() = joinFields { forEachSearchableListField(it) }
private inline fun joinFields(walk: (IndexableFieldVisitor) -> Boolean) =
buildString {
forEachSearchableListField { field ->
walk { field ->
if (field != null) {
if (isNotEmpty()) append('\n')
append(field)
@@ -36,15 +36,15 @@ package com.vitorpamplona.quartz.nip50Search
fun isNaturalLanguageValue(value: String): Boolean {
var start = 0
var end = value.length
while (start < end && value[start].isWhitespace()) start++
while (end > start && value[end - 1].isWhitespace()) end--
while (start < end && value[start].isSpace()) start++
while (end > start && value[end - 1].isSpace()) end--
if (start == end || isStructuredValue(value, start, end)) return false
var nonAscii = false
for (i in start until end) {
val c = value[i]
// The token shapes in isMachineToken cannot contain whitespace.
if (c.isWhitespace()) return true
if (c.isSpace()) return true
if (c.code > 127) nonAscii = true
}
@@ -68,10 +68,10 @@ fun isNaturalLanguageValue(value: String): Boolean {
fun isMachineValue(value: String): Boolean {
var start = 0
var end = value.length
while (start < end && value[start].isWhitespace()) start++
while (end > start && value[end - 1].isWhitespace()) end--
while (start < end && value[start].isSpace()) start++
while (end > start && value[end - 1].isSpace()) end--
if (start == end || isStructuredValue(value, start, end)) return true
for (i in start until end) if (value[i].isWhitespace()) return false
for (i in start until end) if (value[i].isSpace()) return false
return isMachineToken(value, start, end)
}
@@ -99,6 +99,12 @@ private fun isMachineToken(
isUuid(v, start, end) ||
isBech32(v, start)
/**
* [isWhitespace] with an ASCII fast path: on the JVM, [isWhitespace] is two Unicode table lookups
* per character, and nearly every character this file scans is ASCII.
*/
private fun Char.isSpace() = if (code < 0x80) this == ' ' || this in '\t'..'\r' || this in '\u001C'..'\u001F' else isWhitespace()
private fun Char.isAsciiDigit() = this in '0'..'9'
private fun Char.isAsciiLetter() = this in 'a'..'z' || this in 'A'..'Z'
@@ -58,8 +58,8 @@
9736 The content body.
9737 The content body.
9802 The Comment\nThe Context\nThe content body.
9998 The Name\nThe Title\nThe Description\nhashtag1\nhashtag2\nThe Subject\nThe Summary\nThe About\nThe Comment\nThe Context\nThe Rules\nThe Text\nThe Instructions
9999 The Name\nThe Title\nThe Description\nhashtag1\nhashtag2\nThe Subject\nThe Summary\nThe About\nThe Comment\nThe Context\nThe Rules\nThe Text\nThe Instructions
9998 The Name\nThe Title\nThe Description\nThe Subject\nThe Summary\nThe About\nThe Comment\nThe Context\nThe Rules\nThe Text\nThe Instructions\nhashtag1\nhashtag2
9999 The Name\nThe Title\nThe Description\nThe Subject\nThe Summary\nThe About\nThe Comment\nThe Context\nThe Rules\nThe Text\nThe Instructions\nhashtag1\nhashtag2
10003 The Title
10100
10154 The Title\nThe Description
@@ -141,8 +141,8 @@
39092 The Title\nThe Description
39307 The content body.
39701 The Title\nThe content body.
39998 The Name\nThe Title\nThe Description\nhashtag1\nhashtag2\nThe Subject\nThe Summary\nThe About\nThe Comment\nThe Context\nThe Rules\nThe Text\nThe Instructions
39999 The Name\nThe Title\nThe Description\nhashtag1\nhashtag2\nThe Subject\nThe Summary\nThe About\nThe Comment\nThe Context\nThe Rules\nThe Text\nThe Instructions
39998 The Name\nThe Title\nThe Description\nThe Subject\nThe Summary\nThe About\nThe Comment\nThe Context\nThe Rules\nThe Text\nThe Instructions\nhashtag1\nhashtag2
39999 The Name\nThe Title\nThe Description\nThe Subject\nThe Summary\nThe About\nThe Comment\nThe Context\nThe Rules\nThe Text\nThe Instructions\nhashtag1\nhashtag2
40002 The content body.
40100 The content body.
45001 The content body.