mirror of
https://github.com/vitorpamplona/amethyst.git
synced 2026-10-05 11:18:24 +00:00
test: cover the meta-tag variants the suite never saw, and fix title/textarea
Reviewing what the tests actually reach turned up one live defect and a suite that mostly did not run. The defect is the comment bug again, in the elements whose content is text rather than markup. `<title>5 < 6, that's math</title>` is ordinary HTML: the `<` opened a phantom tag, the apostrophe opened a phantom attribute value, and every tag after it -- the whole og: block -- was swallowed. Same for `<textarea>`, and a `<meta>` written inside title text was parsed as a real one and won over the page's own. Title and textarea now skip to their end tag like script and style; the three tests for it fail without that change. The suite: MetaTagsParserTest lived in `androidDeviceTest`, so every attribute shape it covers -- unquoted values, single quotes, valueless attributes, duplicate-attribute rejection, `</head>` inside a value -- was unguarded in CI. Nothing in it is Android-specific, so it moves to commonTest. OpenGraphParser and HtmlCharsetParser had no tests at all. New coverage, all of it variants nothing exercised before: - end of scan: `</HEAD>`, `</head >`, `</head>` inside an attribute value, and a document with no `</head>` at all. - truncation: a body cut after a `/`, inside a quoted value, and right after a `<` -- the first is the crash the previous commit guarded and never pinned. - tag shapes: uppercase `<META>`, `<meta/>`, `/>` inside a quoted value, `<![CDATA[]]>`, and meta tags inside `<noscript>` (which must still be read). - attributes: a trailing valueless attribute, a value spanning lines, unknown attributes. - character references: query-string `&` left intact (og:image URLs are full of them), astral references (`😀`), unknown references left alone. - laziness: the sequence stops when the consumer does. - OpenGraphParser: property / name / itemprop sources, twitter and plain-name fallbacks, and that a field is taken in document order -- so a plain `<meta name="description">` above the og: one wins. That is why the brainstorm.world card shows the site description; pinned, not changed. - HtmlCharsetParser / HtmlParser: charset attribute, http-equiv content-type, the UTF-8 default, the 1 KB sniff window, and the response-charset > BOM > document-declaration precedence. 49 tests in the package, from 8 that ran. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FumxeDJPEgPqX8mz3xkM6b
This commit is contained in:
+24
-8
@@ -44,11 +44,18 @@ object MetaTagsParser {
|
||||
|
||||
private const val META = "meta"
|
||||
private const val HEAD = "head"
|
||||
|
||||
// Elements whose content is text rather than markup: script and style hold raw text, title and
|
||||
// textarea hold character data. A `<` inside any of them is not a tag.
|
||||
private const val SCRIPT = "script"
|
||||
private const val STYLE = "style"
|
||||
private const val TITLE = "title"
|
||||
private const val TEXTAREA = "textarea"
|
||||
|
||||
private const val SCRIPT_END = "</script"
|
||||
private const val STYLE_END = "</style"
|
||||
private const val TITLE_END = "</title"
|
||||
private const val TEXTAREA_END = "</textarea"
|
||||
private const val COMMENT_START = "!--"
|
||||
private const val COMMENT_END = "-->"
|
||||
|
||||
@@ -202,15 +209,24 @@ object MetaTagsParser {
|
||||
return if (nameIs(nameStart, nameEnd, HEAD)) TagKind.HEAD_END else TagKind.OTHER
|
||||
}
|
||||
|
||||
if (nameIs(nameStart, nameEnd, META)) return TagKind.META
|
||||
// Text-only elements are skipped whole: `for (i = 0; i < n; i++)` in a script, a quote
|
||||
// in a JS string, or the `<` and the apostrophe in `<title>5 < 6, that's math</title>`
|
||||
// are content, not markup, and scanning them as markup hides the tags that follow --
|
||||
// the same way an unbalanced quote inside a comment does. Switching on the name length
|
||||
// first keeps the common tag (a `<div>`, a `<link>`) down to one comparison.
|
||||
when (nameEnd - nameStart) {
|
||||
META.length -> if (nameIs(nameStart, nameEnd, META)) return TagKind.META
|
||||
|
||||
// Script and style bodies are raw text: a `<` in `for (i = 0; i < n; i++)` or a quote
|
||||
// in a JS string is not markup and must not be scanned as such, for the same reason
|
||||
// comments can't be.
|
||||
if (nameIs(nameStart, nameEnd, SCRIPT)) {
|
||||
skipRawText(SCRIPT_END)
|
||||
} else if (nameIs(nameStart, nameEnd, STYLE)) {
|
||||
skipRawText(STYLE_END)
|
||||
STYLE.length ->
|
||||
if (nameIs(nameStart, nameEnd, STYLE)) {
|
||||
skipRawText(STYLE_END)
|
||||
} else if (nameIs(nameStart, nameEnd, TITLE)) {
|
||||
skipRawText(TITLE_END)
|
||||
}
|
||||
|
||||
SCRIPT.length -> if (nameIs(nameStart, nameEnd, SCRIPT)) skipRawText(SCRIPT_END)
|
||||
|
||||
TEXTAREA.length -> if (nameIs(nameStart, nameEnd, TEXTAREA)) skipRawText(TEXTAREA_END)
|
||||
}
|
||||
|
||||
return TagKind.OTHER
|
||||
|
||||
+120
@@ -0,0 +1,120 @@
|
||||
/*
|
||||
* Copyright (c) 2025 Vitor Pamplona
|
||||
*
|
||||
* Permission is hereby granted, free of charge, to any person obtaining a copy of
|
||||
* this software and associated documentation files (the "Software"), to deal in
|
||||
* the Software without restriction, including without limitation the rights to use,
|
||||
* copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the
|
||||
* Software, and to permit persons to whom the Software is furnished to do so,
|
||||
* subject to the following conditions:
|
||||
*
|
||||
* The above copyright notice and this permission notice shall be included in all
|
||||
* copies or substantial portions of the Software.
|
||||
*
|
||||
* THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||
* IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS
|
||||
* FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR
|
||||
* COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN
|
||||
* AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION
|
||||
* WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
|
||||
*/
|
||||
package com.vitorpamplona.amethyst.commons.preview
|
||||
|
||||
import kotlinx.coroutines.test.runTest
|
||||
import kotlin.test.Test
|
||||
import kotlin.test.assertEquals
|
||||
|
||||
/**
|
||||
* The charset half of the meta scan: `<meta charset>` and `<meta http-equiv="content-type">` are
|
||||
* the tags that decide how the rest of the document -- including every og: value -- is decoded.
|
||||
* Get this wrong and a preview renders mojibake rather than nothing, so it fails quietly.
|
||||
*/
|
||||
class HtmlParserCharsetTest {
|
||||
private suspend fun firstContentOf(
|
||||
bytes: ByteArray,
|
||||
charsetName: String?,
|
||||
): String =
|
||||
HtmlParser()
|
||||
.parseHtml(bytes, charsetName)
|
||||
.last()
|
||||
.attr("content")
|
||||
|
||||
// -- HtmlCharsetParser --------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
fun sniffsTheCharsetAttribute() {
|
||||
assertEquals("iso-8859-1", HtmlCharsetParser.detectCharset("""<head><meta charset="iso-8859-1">""".encodeToByteArray()))
|
||||
}
|
||||
|
||||
@Test
|
||||
fun sniffsTheHttpEquivContentType() {
|
||||
val html = """<head><meta http-equiv="Content-Type" content="text/html; charset=shift_jis">"""
|
||||
|
||||
assertEquals("shift_jis", HtmlCharsetParser.detectCharset(html.encodeToByteArray()))
|
||||
}
|
||||
|
||||
@Test
|
||||
fun defaultsToUtf8WhenNothingIsDeclared() {
|
||||
assertEquals("UTF-8", HtmlCharsetParser.detectCharset("""<head><title>x</title>""".encodeToByteArray()))
|
||||
}
|
||||
|
||||
@Test
|
||||
fun onlySniffsTheFirstKilobyte() {
|
||||
// The window is deliberate -- the declaration is required to be early -- but it means a
|
||||
// charset pushed past 1 KB by a banner comment is not found, and UTF-8 is assumed.
|
||||
val pushedOut = "<head>" + "<!-- " + "x".repeat(1100) + " -->" + """<meta charset="iso-8859-1">"""
|
||||
|
||||
assertEquals("UTF-8", HtmlCharsetParser.detectCharset(pushedOut.encodeToByteArray()))
|
||||
}
|
||||
|
||||
@Test
|
||||
fun aCommentedOutCharsetIsNotSniffed() {
|
||||
val html = """<head><!-- <meta charset="iso-8859-1"> --><meta charset="utf-8">"""
|
||||
|
||||
assertEquals("utf-8", HtmlCharsetParser.detectCharset(html.encodeToByteArray()))
|
||||
}
|
||||
|
||||
// -- HtmlParser: which charset wins --------------------------------------------------------
|
||||
|
||||
@Test
|
||||
fun anExplicitCharsetWinsOverTheDocumentDeclaration() =
|
||||
runTest {
|
||||
// Content-Type said windows-1252; the document claims utf-8 and must not be believed.
|
||||
val bytes =
|
||||
byteArrayOf(0x3C) + // '<'
|
||||
"""head><meta charset="utf-8"><meta property="og:title" content="caf""".encodeToByteArray() +
|
||||
byteArrayOf(0xE9.toByte()) + // 'é' in windows-1252
|
||||
""""></head>""".encodeToByteArray()
|
||||
|
||||
assertEquals("café", firstContentOf(bytes, "windows-1252"))
|
||||
}
|
||||
|
||||
@Test
|
||||
fun aByteOrderMarkWinsOverTheDocumentDeclaration() =
|
||||
runTest {
|
||||
val utf16 = """<head><meta charset="iso-8859-1"><meta property="og:title" content="café"></head>"""
|
||||
val bytes = byteArrayOf(0xFE.toByte(), 0xFF.toByte()) + utf16.encodeToUtf16Be()
|
||||
|
||||
assertEquals("café", firstContentOf(bytes, null))
|
||||
}
|
||||
|
||||
@Test
|
||||
fun fallsBackToTheDocumentDeclarationWhenTheResponseHasNoCharset() =
|
||||
runTest {
|
||||
val bytes =
|
||||
"""<head><meta charset="windows-1252"><meta property="og:title" content="caf""".encodeToByteArray() +
|
||||
byteArrayOf(0xE9.toByte()) +
|
||||
""""></head>""".encodeToByteArray()
|
||||
|
||||
assertEquals("café", firstContentOf(bytes, null))
|
||||
}
|
||||
|
||||
private fun String.encodeToUtf16Be(): ByteArray {
|
||||
val out = ByteArray(length * 2)
|
||||
forEachIndexed { i, c ->
|
||||
out[i * 2] = (c.code shr 8).toByte()
|
||||
out[i * 2 + 1] = (c.code and 0xFF).toByte()
|
||||
}
|
||||
return out
|
||||
}
|
||||
}
|
||||
+71
-1
@@ -24,7 +24,8 @@ import kotlin.test.Test
|
||||
import kotlin.test.assertEquals
|
||||
|
||||
/**
|
||||
* Comments, declarations and script bodies are not element markup. Scanning them for
|
||||
* Comments, declarations and the text-only elements (script, style, title, textarea) are not
|
||||
* element markup. Scanning them for
|
||||
* attribute quotes lets an odd apostrophe -- "we don't" is enough -- leave the scanner
|
||||
* inside a phantom quoted value, swallowing every tag up to the next quote character.
|
||||
* That is what hid the entire og: block of https://brainstorm.world from link previews.
|
||||
@@ -100,6 +101,75 @@ class MetaTagsParserCommentTest {
|
||||
assertEquals("Real Title", metaTags[0].attr("content"))
|
||||
}
|
||||
|
||||
@Test
|
||||
fun titleTextWithALessThanDoesNotSwallowFollowingMetaTags() {
|
||||
// `<title>5 < 6, that's math</title>` is ordinary HTML: an unescaped `<` in title text,
|
||||
// then an apostrophe. Scanned as markup, the `<` opens a phantom tag and the apostrophe
|
||||
// opens a phantom attribute value that runs to the end of the document.
|
||||
val input =
|
||||
"""
|
||||
|<html><head>
|
||||
| <title>5 < 6, that's math</title>
|
||||
| <meta property="og:title" content="Real Title">
|
||||
|</head></html>
|
||||
""".trimMargin()
|
||||
|
||||
val metaTags = MetaTagsParser.parse(input).toList()
|
||||
|
||||
assertEquals(1, metaTags.size)
|
||||
assertEquals("Real Title", metaTags[0].attr("content"))
|
||||
}
|
||||
|
||||
@Test
|
||||
fun metaTagsInsideTitleTextAreNotParsed() {
|
||||
// Title content is character data, so this is a title that reads literally
|
||||
// `a <meta property="og:title" content="Fake"> b`, not a second og:title.
|
||||
val input =
|
||||
"""
|
||||
|<html><head>
|
||||
| <title>a <meta property="og:title" content="Fake"> b</title>
|
||||
| <meta property="og:title" content="Real Title">
|
||||
|</head></html>
|
||||
""".trimMargin()
|
||||
|
||||
val metaTags = MetaTagsParser.parse(input).toList()
|
||||
|
||||
assertEquals(1, metaTags.size)
|
||||
assertEquals("Real Title", metaTags[0].attr("content"))
|
||||
}
|
||||
|
||||
@Test
|
||||
fun textareaTextDoesNotSwallowFollowingMetaTags() {
|
||||
val input =
|
||||
"""
|
||||
|<html><head>
|
||||
| <textarea>x < y's z</textarea>
|
||||
| <meta property="og:title" content="Real Title">
|
||||
|</head></html>
|
||||
""".trimMargin()
|
||||
|
||||
val metaTags = MetaTagsParser.parse(input).toList()
|
||||
|
||||
assertEquals(1, metaTags.size)
|
||||
assertEquals("Real Title", metaTags[0].attr("content"))
|
||||
}
|
||||
|
||||
@Test
|
||||
fun aSelfClosedScriptStillOpensRawText() {
|
||||
// Deliberate, and what a browser does: `/` on a script start tag is ignored, so everything
|
||||
// up to `</script>` is script data. A page written this way shows nothing after it either.
|
||||
// Pinned so that "fixing" it never turns a JS string into an og: tag.
|
||||
val input =
|
||||
"""
|
||||
|<html><head>
|
||||
| <script src="a.js"/>
|
||||
| <meta property="og:title" content="Unreachable">
|
||||
|</head></html>
|
||||
""".trimMargin()
|
||||
|
||||
assertEquals(0, MetaTagsParser.parse(input).count())
|
||||
}
|
||||
|
||||
@Test
|
||||
fun doctypeAndProcessingInstructionsAreSkipped() {
|
||||
val input =
|
||||
|
||||
+208
@@ -0,0 +1,208 @@
|
||||
/*
|
||||
* Copyright (c) 2025 Vitor Pamplona
|
||||
*
|
||||
* Permission is hereby granted, free of charge, to any person obtaining a copy of
|
||||
* this software and associated documentation files (the "Software"), to deal in
|
||||
* the Software without restriction, including without limitation the rights to use,
|
||||
* copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the
|
||||
* Software, and to permit persons to whom the Software is furnished to do so,
|
||||
* subject to the following conditions:
|
||||
*
|
||||
* The above copyright notice and this permission notice shall be included in all
|
||||
* copies or substantial portions of the Software.
|
||||
*
|
||||
* THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||
* IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS
|
||||
* FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR
|
||||
* COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN
|
||||
* AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION
|
||||
* WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
|
||||
*/
|
||||
package com.vitorpamplona.amethyst.commons.preview
|
||||
|
||||
import kotlin.test.Test
|
||||
import kotlin.test.assertEquals
|
||||
import kotlin.test.assertTrue
|
||||
|
||||
/**
|
||||
* Tag and attribute shapes a preview fetch can meet in the wild, and the ones a truncated or
|
||||
* hostile response can produce. The server picks this input, so "it throws" and "it silently eats
|
||||
* the rest of the head" both have to be ruled out for every shape here.
|
||||
*/
|
||||
class MetaTagsParserEdgeCaseTest {
|
||||
private fun contents(html: String) = MetaTagsParser.parse(html).map { it.attr("content") }.toList()
|
||||
|
||||
// -- the end of the scan ------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
fun stopsAtAnUppercaseHeadEndTag() {
|
||||
val metas = contents("""<head><meta name="a" content="1"></HEAD><meta name="b" content="2">""")
|
||||
|
||||
assertEquals(listOf("1"), metas)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun stopsAtAHeadEndTagWithTrailingSpace() {
|
||||
val metas = contents("""<head><meta name="a" content="1"></head ><meta name="b" content="2">""")
|
||||
|
||||
assertEquals(listOf("1"), metas)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun aHeadEndTagInsideAnAttributeValueDoesNotEndTheScan() {
|
||||
val metas = contents("""<head><meta name="a" content="</head>"><meta name="b" content="2"></head>""")
|
||||
|
||||
assertEquals(listOf("</head>", "2"), metas)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun scansTheWholeDocumentWhenThereIsNoHeadEndTag() {
|
||||
val metas = contents("""<head><meta name="a" content="1"><body><meta name="b" content="2">""")
|
||||
|
||||
assertEquals(listOf("1", "2"), metas)
|
||||
}
|
||||
|
||||
// -- truncated responses ------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
fun aBodyTruncatedRightAfterASlashDoesNotThrow() {
|
||||
// The `/` of a `/>` as the very last byte: the self-closing check must not read past it.
|
||||
val metas = contents("""<head><meta name="a" content="1" /""")
|
||||
|
||||
assertEquals(listOf("1"), metas)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun aBodyTruncatedInsideAnAttributeValueYieldsNoValue() {
|
||||
val metas = MetaTagsParser.parse("""<head><meta property="og:title" content="Trunc""").toList()
|
||||
|
||||
assertEquals(1, metas.size)
|
||||
assertEquals("og:title", metas[0].attr("property"))
|
||||
assertEquals("", metas[0].attr("content"))
|
||||
}
|
||||
|
||||
@Test
|
||||
fun aBodyTruncatedRightAfterALessThanDoesNotThrow() {
|
||||
assertEquals(listOf("1"), contents("""<head><meta name="a" content="1"><"""))
|
||||
}
|
||||
|
||||
// -- tag shapes ---------------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
fun readsAnUppercaseMetaTagAndUppercaseAttributeNames() {
|
||||
val metas = contents("""<head><META PROPERTY="og:title" CONTENT="Real"></head>""")
|
||||
|
||||
assertEquals(listOf("Real"), metas)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun anEmptySelfClosedMetaIsSkippedWithoutDerailingTheScan() {
|
||||
// `<meta/>` has no separator before the `/`, so the name reads as `meta/` and the tag is
|
||||
// dropped. It carries nothing anyway; what matters is that the next tag still parses.
|
||||
val metas = contents("""<head><meta/><meta property="og:title" content="Real"></head>""")
|
||||
|
||||
assertEquals(listOf("Real"), metas)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun aSelfClosingSequenceInsideAQuotedValueDoesNotEndTheTag() {
|
||||
val metas =
|
||||
contents(
|
||||
"""<head><meta property="og:title" content="a /> b"><meta name="after" content="ok"></head>""",
|
||||
)
|
||||
|
||||
assertEquals(listOf("a /> b", "ok"), metas)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun readsMetaTagsInsideNoscript() {
|
||||
// noscript content is markup for a parser that isn't running scripts, and a redirect meta
|
||||
// hidden in there is exactly the kind a preview wants to see.
|
||||
val metas =
|
||||
contents(
|
||||
"""<head><noscript><meta http-equiv="refresh" content="0"></noscript><meta property="og:title" content="Real"></head>""",
|
||||
)
|
||||
|
||||
assertEquals(listOf("0", "Real"), metas)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun skipsCdataSections() {
|
||||
val metas = contents("""<head><![CDATA[ x > y ]]><meta property="og:title" content="Real"></head>""")
|
||||
|
||||
assertEquals(listOf("Real"), metas)
|
||||
}
|
||||
|
||||
// -- attribute shapes ---------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
fun keepsAttributesWhenTheTagEndsWithAValuelessOne() {
|
||||
val metas = MetaTagsParser.parse("""<head><meta property="og:title" content="T" data-foo></head>""").toList()
|
||||
|
||||
assertEquals(1, metas.size)
|
||||
assertEquals("og:title", metas[0].attr("property"))
|
||||
assertEquals("T", metas[0].attr("content"))
|
||||
}
|
||||
|
||||
@Test
|
||||
fun keepsAValueThatSpansLines() {
|
||||
val metas = contents("<head><meta property=\"og:title\" content=\"line1\nline2\"></head>")
|
||||
|
||||
assertEquals(listOf("line1\nline2"), metas)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun anUnknownAttributeIsIgnoredNotFatal() {
|
||||
val metas = contents("""<head><meta data-rh="true" property="og:title" content="Real"></head>""")
|
||||
|
||||
assertEquals(listOf("Real"), metas)
|
||||
}
|
||||
|
||||
// -- character references in values --------------------------------------------------------
|
||||
|
||||
@Test
|
||||
fun keepsQueryStringAmpersandsIntact() {
|
||||
// og:image URLs are full of `&`; only a real character reference may be resolved.
|
||||
val metas =
|
||||
contents(
|
||||
"""<head><meta property="og:image" content="https://x.com/i.png?w=1200&h=630&fit=crop&q=80"></head>""",
|
||||
)
|
||||
|
||||
assertEquals(listOf("https://x.com/i.png?w=1200&h=630&fit=crop&q=80"), metas)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun decodesCharacterReferencesOutsideTheBasicMultilingualPlane() {
|
||||
val metas = contents("""<head><meta property="og:title" content="😀 😀 hi"></head>""")
|
||||
|
||||
assertEquals(listOf("😀 😀 hi"), metas)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun leavesAnUnknownCharacterReferenceAlone() {
|
||||
val metas = contents("""<head><meta property="og:title" content="AT&T ¬areference; &#xZZ;"></head>""")
|
||||
|
||||
assertEquals(listOf("AT&T ¬areference; &#xZZ;"), metas)
|
||||
}
|
||||
|
||||
// -- laziness -------------------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
fun stopsReadingOnceTheConsumerStops() {
|
||||
// OpenGraphParser bails as soon as it has title+description+image; the sequence must not
|
||||
// have scanned the rest of the document by then. A tag after an unterminated comment is
|
||||
// unreachable, so seeing the first one proves the scan was still lazy.
|
||||
val html =
|
||||
"""
|
||||
|<head>
|
||||
| <meta property="og:title" content="T">
|
||||
| <!-- an unterminated comment swallows everything after it
|
||||
| <meta property="og:description" content="D">
|
||||
""".trimMargin()
|
||||
|
||||
val first = MetaTagsParser.parse(html).first()
|
||||
|
||||
assertEquals("T", first.attr("content"))
|
||||
assertTrue(MetaTagsParser.parse(html).count() == 1)
|
||||
}
|
||||
}
|
||||
+6
-2
@@ -20,9 +20,13 @@
|
||||
*/
|
||||
package com.vitorpamplona.amethyst.commons.preview
|
||||
|
||||
import org.junit.Assert.assertEquals
|
||||
import org.junit.Test
|
||||
import kotlin.test.Test
|
||||
import kotlin.test.assertEquals
|
||||
|
||||
/**
|
||||
* Lived in `androidDeviceTest` until it was moved here: nothing in it is Android-specific, and on
|
||||
* a device-only source set it never ran in CI -- every attribute shape below was unguarded.
|
||||
*/
|
||||
class MetaTagsParserTest {
|
||||
@Test
|
||||
fun testParse() {
|
||||
+145
@@ -0,0 +1,145 @@
|
||||
/*
|
||||
* Copyright (c) 2025 Vitor Pamplona
|
||||
*
|
||||
* Permission is hereby granted, free of charge, to any person obtaining a copy of
|
||||
* this software and associated documentation files (the "Software"), to deal in
|
||||
* the Software without restriction, including without limitation the rights to use,
|
||||
* copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the
|
||||
* Software, and to permit persons to whom the Software is furnished to do so,
|
||||
* subject to the following conditions:
|
||||
*
|
||||
* The above copyright notice and this permission notice shall be included in all
|
||||
* copies or substantial portions of the Software.
|
||||
*
|
||||
* THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||
* IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS
|
||||
* FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR
|
||||
* COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN
|
||||
* AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION
|
||||
* WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
|
||||
*/
|
||||
package com.vitorpamplona.amethyst.commons.preview
|
||||
|
||||
import kotlin.test.Test
|
||||
import kotlin.test.assertEquals
|
||||
|
||||
/**
|
||||
* The three attributes a preview can be declared under -- `property` (Open Graph), `name`
|
||||
* (Twitter cards and plain HTML) and `itemprop` (schema.org) -- and what happens when a page
|
||||
* declares the same thing under more than one.
|
||||
*/
|
||||
class OpenGraphParserTest {
|
||||
private fun extract(html: String) = OpenGraphParser().extractUrlInfo(MetaTagsParser.parse(html))
|
||||
|
||||
@Test
|
||||
fun readsOpenGraphProperties() {
|
||||
val info =
|
||||
extract(
|
||||
"""
|
||||
|<head>
|
||||
| <meta property="og:title" content="T">
|
||||
| <meta property="og:description" content="D">
|
||||
| <meta property="og:image" content="https://example.com/i.png">
|
||||
|</head>
|
||||
""".trimMargin(),
|
||||
)
|
||||
|
||||
assertEquals("T", info.title)
|
||||
assertEquals("D", info.description)
|
||||
assertEquals("https://example.com/i.png", info.image)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun fallsBackToTwitterCardNames() {
|
||||
val info =
|
||||
extract(
|
||||
"""
|
||||
|<head>
|
||||
| <meta name="twitter:title" content="T">
|
||||
| <meta name="twitter:description" content="D">
|
||||
| <meta name="twitter:image" content="https://example.com/i.png">
|
||||
|</head>
|
||||
""".trimMargin(),
|
||||
)
|
||||
|
||||
assertEquals("T", info.title)
|
||||
assertEquals("D", info.description)
|
||||
assertEquals("https://example.com/i.png", info.image)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun fallsBackToPlainNames() {
|
||||
val info =
|
||||
extract(
|
||||
"""
|
||||
|<head>
|
||||
| <meta name="title" content="T">
|
||||
| <meta name="description" content="D">
|
||||
| <meta name="image" content="https://example.com/i.png">
|
||||
|</head>
|
||||
""".trimMargin(),
|
||||
)
|
||||
|
||||
assertEquals("T", info.title)
|
||||
assertEquals("D", info.description)
|
||||
assertEquals("https://example.com/i.png", info.image)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun readsSchemaOrgItemprops() {
|
||||
val info =
|
||||
extract(
|
||||
"""
|
||||
|<head>
|
||||
| <meta itemprop="name" content="ignored, not a title key">
|
||||
| <meta itemprop="title" content="T">
|
||||
| <meta itemprop="description" content="D">
|
||||
| <meta itemprop="image" content="https://example.com/i.png">
|
||||
|</head>
|
||||
""".trimMargin(),
|
||||
)
|
||||
|
||||
assertEquals("T", info.title)
|
||||
assertEquals("D", info.description)
|
||||
assertEquals("https://example.com/i.png", info.image)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun takesEachFieldFromWhicheverTagCarriesItFirst() {
|
||||
// NOTE: this is document order, not source priority. A page that puts a plain
|
||||
// <meta name="description"> above its <meta property="og:description"> -- a very common
|
||||
// CMS layout -- has the plain one win, even though og: is the more specific declaration.
|
||||
// Pinned as the current behavior; changing it means preferring og: over name: explicitly.
|
||||
val info =
|
||||
extract(
|
||||
"""
|
||||
|<head>
|
||||
| <meta name="description" content="plain, comes first">
|
||||
| <meta property="og:description" content="og, comes second">
|
||||
| <meta property="og:title" content="T">
|
||||
| <meta property="og:image" content="https://example.com/i.png">
|
||||
|</head>
|
||||
""".trimMargin(),
|
||||
)
|
||||
|
||||
assertEquals("plain, comes first", info.description)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun missingFieldsComeBackEmptyRatherThanNull() {
|
||||
val info = extract("""<head><meta property="og:title" content="T"></head>""")
|
||||
|
||||
assertEquals("T", info.title)
|
||||
assertEquals("", info.description)
|
||||
assertEquals("", info.image)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun ignoresAMetaTagWithNoRecognizedKey() {
|
||||
val info = extract("""<head><meta name="viewport" content="width=device-width"></head>""")
|
||||
|
||||
assertEquals("", info.title)
|
||||
assertEquals("", info.description)
|
||||
assertEquals("", info.image)
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user