Files
amethyst/commons
Claude 39797b2191 perf: drop the regex and the per-tag allocations from the meta scan
The tag-name check was a Regex match over a freshly cut substring, run for
every `<` in the document. Both are gone: names are compared in place against
the only four that matter (meta, head, script, style), ASCII-case-folded with
`code or 0x20`, so a non-meta tag now costs zero allocations -- no substring,
no Matcher, no RawTag. nextTag() reports a TagKind and leaves the attribute
span as two indices; only a real `<meta>` gets read, and parseAttrs() reads
that span straight out of the document instead of a copy of it.

The rest of the scan got the same treatment:

- `indexOf('<')` / `indexOf('>')` / `indexOf("-->")` instead of char-at-a-time
  predicate loops -- these are intrinsified and vectorized on the JVM.
- `Set<Char>.contains` for the attribute character classes boxed a Char per
  character of every meta tag; they are `when` branches now.
- one Pair, one Result and one lambda per attribute (`runCatching { add(Pair) }`)
  became a boolean-returning add -- a duplicate attribute no longer throws.
- `toImmutableMap()` rebuilt a persistent map for every meta tag; the Attrs
  builder is discarded at freeze(), so its own map is already private.
- the character-reference Regex only runs on values that contain an `&`.

Measured on a comment-free head, where this and the previous implementation
do identical work (same JVM, both warmed, `plainHead` corpus):

  1.1 KB head,  10 metas   12.5 us -> 5.1 us   ( 88 -> 215 MB/s)
   28 KB head, 204 metas    267 us -> 116 us   (105 -> 242 MB/s)

MetaTagsParserBenchmark joins the prodbench suite as the guard, on corpora
shaped like a Vite SPA head and a CMS head buried in analytics scripts: any
site we preview picks the input, so the scan has to stay linear in it.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FumxeDJPEgPqX8mz3xkM6b
2026-08-21 19:58:08 +00:00
..