mirror of
https://github.com/vitorpamplona/amethyst.git
synced 2026-10-05 11:18:24 +00:00
fix(cordn): a handshake that answers is a call, and two counts of one
Three things an afternoon of driving the cordn screens on a tablet turned up. **"Ask who it is" left the coordinator card contradicting itself.** It printed "Nothing asked of it yet." directly above "Says it is: cordn-server · 0.1.0", because `health.recordSuccess` was only reached from `CordnGroupManager.call` and `serverInfo()` goes to the coordinator through `CoordinatorClient` without passing that way. The screen's own comment says health is "only ever what a call of ours already observed", and this is one — usually the very first, since the button exists so someone can find out whether a coordinator they just added answers at all. Recorded in `CordnSession.serverInfo` now, where the session already holds the health it feeds. An answer counts as a success and a throw as a failure; `null` counts as neither, because a scope that did no handshake returns `null` and a call that never left the device must not report a coordinator as reachable. **"1 single-use packages available."** — and "1 packages this device cannot open". Both are now plurals. Three tests cover the health path: an answer marks it observed, a throw counts a failure and keeps its reason, and silence leaves it unknown. --- The live migrate harness could not run on this machine at all; none of it was the product. - `stack.sh` ran the coordinator with `--network host` so it could reach a relay on the host's loopback. On Docker Desktop the daemon is inside a VM, so that is the *VM's* loopback and `127.0.0.1:$PORT` is nothing: the coordinator retried "Relay connection error" until `stack_up` gave up waiting for a pubkey that was never coming. It now uses `host.docker.internal` off Linux and keeps `--network host` on it. - `BLOB_PORT` defaulted to **877**, which is privileged, so the blob stand-in died with `PermissionError` for anyone who is not root — on Linux too. Now 8877. - `blob_up` did not check the server came up. A dead one surfaced thirty lines later as `migrate export` reporting "no server accepted the document", plus two more failures behind it, none naming the blob server. It now waits for the port and exits 2 with the real error. With those, `cli/tests/cordn/migrate.sh` passes end to end on stock defaults. Worth knowing for whoever runs it next: the image is amd64-only, so on Apple Silicon the pull needs `--platform linux/amd64` and then runs emulated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
d5d86270b9
commit
d627ba7846
+4
-2
@@ -47,6 +47,7 @@ import androidx.compose.runtime.rememberCoroutineScope
|
||||
import androidx.compose.runtime.setValue
|
||||
import androidx.compose.ui.Alignment
|
||||
import androidx.compose.ui.Modifier
|
||||
import androidx.compose.ui.platform.LocalContext
|
||||
import androidx.compose.ui.unit.dp
|
||||
import androidx.lifecycle.compose.collectAsStateWithLifecycle
|
||||
import com.vitorpamplona.amethyst.R
|
||||
@@ -57,6 +58,7 @@ import com.vitorpamplona.amethyst.commons.ui.navigation.navs.INav
|
||||
import com.vitorpamplona.amethyst.commons.ui.navigation.topbars.TopBarWithBackButton
|
||||
import com.vitorpamplona.amethyst.model.cordn.CordnKeyPackageRow
|
||||
import com.vitorpamplona.amethyst.model.cordn.CordnRuntime
|
||||
import com.vitorpamplona.amethyst.ui.pluralStringRes
|
||||
import com.vitorpamplona.amethyst.ui.screen.loggedIn.AccountViewModel
|
||||
import com.vitorpamplona.amethyst.ui.stringRes
|
||||
import kotlinx.coroutines.launch
|
||||
@@ -178,7 +180,7 @@ private fun CoordinatorKeyPackages(
|
||||
val hasLastResort = loaded.any { it.lastResort }
|
||||
|
||||
Text(
|
||||
text = stringRes(R.string.cordn_keypackages_summary, single),
|
||||
text = pluralStringRes(LocalContext.current, R.plurals.cordn_keypackages_summary, single, single),
|
||||
style = MaterialTheme.typography.bodyMedium,
|
||||
)
|
||||
Text(
|
||||
@@ -199,7 +201,7 @@ private fun CoordinatorKeyPackages(
|
||||
) {
|
||||
Column(Modifier.padding(12.dp)) {
|
||||
Text(
|
||||
text = stringRes(R.string.cordn_keypackages_orphans, orphans.size),
|
||||
text = pluralStringRes(LocalContext.current, R.plurals.cordn_keypackages_orphans, orphans.size, orphans.size),
|
||||
style = MaterialTheme.typography.titleSmall,
|
||||
)
|
||||
Text(
|
||||
|
||||
@@ -406,10 +406,16 @@
|
||||
<string name="cordn_coordinators_purge_title">Purge this coordinator?</string>
|
||||
<string name="cordn_coordinators_purge_body">Removing keeps the groups on this device, so adding the coordinator back brings them with it. Purging deletes them: the group keys, the read positions and your key packages all go, and those groups can only be re-entered by a fresh invitation.\n\nThe coordinator is not told and keeps whatever it already had.</string>
|
||||
<string name="cordn_keypackages_explainer">Others can only add you to a cordn group by taking a key package you published to that coordinator. cordn has no other place to keep one, so with none published nobody can invite you — and nothing tells them why.</string>
|
||||
<string name="cordn_keypackages_summary">%1$d single-use packages available.</string>
|
||||
<plurals name="cordn_keypackages_summary">
|
||||
<item quantity="one">%1$d single-use package available.</item>
|
||||
<item quantity="other">%1$d single-use packages available.</item>
|
||||
</plurals>
|
||||
<string name="cordn_keypackages_last_resort_yes">A last-resort package is published, so an invitation can still be made once the single-use ones run out.</string>
|
||||
<string name="cordn_keypackages_last_resort_no">No last-resort package. Once the single-use ones run out, invitations fail.</string>
|
||||
<string name="cordn_keypackages_orphans">%1$d packages this device cannot open</string>
|
||||
<plurals name="cordn_keypackages_orphans">
|
||||
<item quantity="one">%1$d package this device cannot open</item>
|
||||
<item quantity="other">%1$d packages this device cannot open</item>
|
||||
</plurals>
|
||||
<string name="cordn_keypackages_orphans_body">The coordinator will hand these to anyone inviting you, but the private half is not on this device — it belongs to another install. If that install is gone, an invitation using one of these produces a welcome nobody can ever open.</string>
|
||||
<string name="cordn_keypackages_withdraw_orphans">Withdraw those</string>
|
||||
<string name="cordn_keypackages_disclosure">Publishing signs a record under your own account key. That coordinator then knows this account exists and is invitable — permanently, and whether or not anyone invites you.</string>
|
||||
|
||||
@@ -28,7 +28,9 @@ CONTAINER="cordn-migrate"
|
||||
|
||||
export AMY_PASSPHRASE="${AMY_PASSPHRASE:-migrate}"
|
||||
|
||||
BLOB_PORT="${BLOB_PORT:-877}"
|
||||
# Above 1024: binding below it needs root, and this test has no business
|
||||
# asking for that. 877 was the default and bound for nobody.
|
||||
BLOB_PORT="${BLOB_PORT:-8877}"
|
||||
BLOB="http://127.0.0.1:$BLOB_PORT"
|
||||
BLOB_DIR="$WORK/blobs"
|
||||
BLOB_PID=""
|
||||
@@ -49,7 +51,20 @@ blob_up() {
|
||||
mkdir -p "$BLOB_DIR"
|
||||
python3 - "$BLOB_PORT" "$BLOB_DIR" >"$WORK/blob.log" 2>&1 &
|
||||
BLOB_PID=$!
|
||||
sleep 1
|
||||
# Wait for the port, and say so here if it never opens. Letting a dead blob
|
||||
# server through costs three misleading failures later — export reports "no
|
||||
# server accepted the document", and the group/blob assertions all fall over
|
||||
# behind it — none of which name the thing that is actually wrong.
|
||||
for _ in $(seq 20); do
|
||||
if python3 -c "import socket,sys; s=socket.socket(); s.settimeout(0.3); sys.exit(0 if s.connect_ex(('127.0.0.1',$BLOB_PORT))==0 else 1)"; then
|
||||
return 0
|
||||
fi
|
||||
kill -0 "$BLOB_PID" 2>/dev/null || break
|
||||
sleep 0.5
|
||||
done
|
||||
echo "the blob server never came up on $BLOB — see $WORK/blob.log"
|
||||
tail -3 "$WORK/blob.log" 2>/dev/null
|
||||
exit 2
|
||||
} <<'PYEOF'
|
||||
import hashlib, os, sys
|
||||
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||
|
||||
@@ -60,15 +60,32 @@ stack_up() {
|
||||
sleep 1
|
||||
done
|
||||
|
||||
# --network host so the container reaches a relay on the host's loopback.
|
||||
# The container has to reach geode, which is on the host's loopback.
|
||||
#
|
||||
# On Linux `--network host` puts it there. On Docker Desktop (macOS,
|
||||
# Windows) the daemon is inside a VM, so `--network host` is the *VM's*
|
||||
# loopback and 127.0.0.1:$PORT is nothing at all: the coordinator retries
|
||||
# "Relay connection error" until stack_up gives up on a pubkey that was
|
||||
# never going to arrive. There the host is reachable by name instead, over
|
||||
# the default bridge.
|
||||
#
|
||||
# A stable key so the coordinator pubkey survives a re-run against the
|
||||
# same WORK directory.
|
||||
if [ "$(uname -s)" = "Linux" ]; then
|
||||
COORD_NET="--network host"
|
||||
COORD_RELAY="$RELAY"
|
||||
else
|
||||
COORD_NET="--add-host=host.docker.internal:host-gateway"
|
||||
COORD_RELAY="ws://host.docker.internal:$PORT"
|
||||
fi
|
||||
|
||||
[ -f "$WORK/coordinator.key" ] || openssl rand -hex 32 >"$WORK/coordinator.key"
|
||||
docker rm -f "$CONTAINER" >/dev/null 2>&1
|
||||
docker run -d --name "$CONTAINER" --network host \
|
||||
# shellcheck disable=SC2086 # COORD_NET is two words on purpose
|
||||
docker run -d --name "$CONTAINER" $COORD_NET \
|
||||
-e CORDN_STORAGE_BACKEND=memory \
|
||||
-e CORDN_ANNOUNCED=false \
|
||||
-e CORDN_RELAY_URLS="$RELAY" \
|
||||
-e CORDN_RELAY_URLS="$COORD_RELAY" \
|
||||
-e CORDN_SERVER_PRIVATE_KEY="$(cat "$WORK/coordinator.key")" \
|
||||
-e CORDN_SERVER_NAME="cordn-test" \
|
||||
"$IMAGE" >/dev/null || { echo "could not start $CONTAINER"; exit 1; }
|
||||
|
||||
+25
-2
@@ -23,6 +23,7 @@ package com.vitorpamplona.amethyst.commons.cordn
|
||||
import com.vitorpamplona.quartz.cordn.spec00Coordinator.CoordinatorServerInfo
|
||||
import com.vitorpamplona.quartz.cordn.spec00Coordinator.ICoordinator
|
||||
import com.vitorpamplona.quartz.nip01Core.core.HexKey
|
||||
import com.vitorpamplona.quartz.utils.TimeUtils
|
||||
import kotlinx.coroutines.flow.MutableStateFlow
|
||||
import kotlinx.coroutines.flow.StateFlow
|
||||
import kotlinx.coroutines.flow.asStateFlow
|
||||
@@ -70,6 +71,8 @@ class CordnSession(
|
||||
val keyPackages: CordnKeyPackages,
|
||||
val health: CoordinatorHealth,
|
||||
private val scope: CordnCoordinatorScope,
|
||||
/** Seconds. Injected so tests are not at the mercy of the wall clock. */
|
||||
private val clock: () -> Long = { TimeUtils.now() },
|
||||
) {
|
||||
val coordinatorPubKey: HexKey get() = config.pubKey
|
||||
|
||||
@@ -85,8 +88,28 @@ class CordnSession(
|
||||
*/
|
||||
suspend fun exposure(gid: String): GroupExposure = manager.exposure(gid, publishedKeyPackage = keyPackages.hasPublished())
|
||||
|
||||
/** What this coordinator says about itself. Claims, never identity (§8.5). */
|
||||
suspend fun serverInfo(): CoordinatorServerInfo? = scope.serverInfo()
|
||||
/**
|
||||
* What this coordinator says about itself. Claims, never identity (§8.5).
|
||||
*
|
||||
* Counts against [health] like any other call. It is a real round trip to
|
||||
* the coordinator, and it is usually the *first* one a reader makes — the
|
||||
* settings screen offers it precisely so someone can find out whether a
|
||||
* coordinator they just added answers at all. Leaving it out had that
|
||||
* screen say "Nothing asked of it yet" directly underneath what the
|
||||
* coordinator had just answered.
|
||||
*
|
||||
* An answer counts as a success and a throw as a failure; `null` counts as
|
||||
* neither. A scope returns `null` when it did no handshake at all, so
|
||||
* treating that as a success would report a coordinator as reachable on
|
||||
* the strength of a call that never left the device.
|
||||
*/
|
||||
suspend fun serverInfo(): CoordinatorServerInfo? =
|
||||
try {
|
||||
scope.serverInfo()?.also { health.recordSuccess(clock()) }
|
||||
} catch (e: Exception) {
|
||||
health.recordFailure(clock(), e.message)
|
||||
throw e
|
||||
}
|
||||
|
||||
/** How many of this session's responses arrived reassembled over CEP-22. */
|
||||
val oversizedTransfers: Int get() = scope.coordinator.oversizedTransfers
|
||||
|
||||
+75
@@ -20,6 +20,7 @@
|
||||
*/
|
||||
package com.vitorpamplona.amethyst.commons.cordn
|
||||
|
||||
import com.vitorpamplona.quartz.cordn.spec00Coordinator.CoordinatorServerInfo
|
||||
import com.vitorpamplona.quartz.cordn.spec00Coordinator.ICoordinator
|
||||
import com.vitorpamplona.quartz.cordn.spec01GroupMetadata.CordnGroupMetadata
|
||||
import com.vitorpamplona.quartz.nip01Core.core.HexKey
|
||||
@@ -255,4 +256,78 @@ class CordnCoordinatorRegistryTest : CordnTransportHarness() {
|
||||
assertTrue(registry.coordinators.value.isEmpty())
|
||||
assertNull(registry.sessionOrNull(config.pubKey))
|
||||
}
|
||||
|
||||
/**
|
||||
* A scope whose handshake the test decides: an answer, silence, or a throw.
|
||||
*/
|
||||
private inner class InfoFactory(
|
||||
private val info: CoordinatorServerInfo?,
|
||||
private val blowUp: Boolean = false,
|
||||
) : CordnCoordinatorScopeFactory {
|
||||
override suspend fun open(
|
||||
accountPubKey: HexKey,
|
||||
config: CoordinatorConfig,
|
||||
): CordnCoordinatorScope =
|
||||
object : CordnCoordinatorScope {
|
||||
override val coordinator: ICoordinator = account.clientFor(config.pubKey)
|
||||
override val groupStore: CordnGroupStore = InMemoryCordnGroupStore()
|
||||
override val keyPackageStore: CordnKeyPackageStore = InMemoryCordnKeyPackageStore()
|
||||
|
||||
override suspend fun serverInfo(): CoordinatorServerInfo? {
|
||||
if (blowUp) throw IllegalStateException("coordinator unreachable")
|
||||
return info
|
||||
}
|
||||
|
||||
override suspend fun close() = Unit
|
||||
}
|
||||
}
|
||||
|
||||
private val anInfo = CoordinatorServerInfo(name = "cordn-server", version = "0.1.0", protocolVersion = "2025-11-25", capabilities = null)
|
||||
|
||||
/**
|
||||
* The settings screen offers "ask who it is" so someone can find out
|
||||
* whether a coordinator answers. Before this, a successful handshake left
|
||||
* the health line reading "Nothing asked of it yet" directly above the
|
||||
* answer it had just printed.
|
||||
*/
|
||||
@Test
|
||||
fun `a handshake that answers counts as a healthy call`() =
|
||||
runTest {
|
||||
val registry = registry(InfoFactory(anInfo))
|
||||
val session = driving { registry.session(config) }
|
||||
|
||||
assertTrue(session.health.state.value.isUnknown, "nothing has been asked yet")
|
||||
|
||||
assertNotNull(driving { session.serverInfo() })
|
||||
|
||||
assertTrue(!session.health.state.value.isUnknown, "the handshake should have been observed")
|
||||
assertNotNull(session.health.state.value.lastSuccessAt)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun `a handshake that throws counts against health`() =
|
||||
runTest {
|
||||
val registry = registry(InfoFactory(null, blowUp = true))
|
||||
val session = driving { registry.session(config) }
|
||||
|
||||
runCatching { driving { session.serverInfo() } }
|
||||
|
||||
assertEquals(1, session.health.state.value.consecutiveFailures)
|
||||
assertEquals("coordinator unreachable", session.health.state.value.lastFailure)
|
||||
}
|
||||
|
||||
/**
|
||||
* A scope that did no handshake returns null, and a call that never left
|
||||
* the device must not report the coordinator as reachable.
|
||||
*/
|
||||
@Test
|
||||
fun `silence is neither a success nor a failure`() =
|
||||
runTest {
|
||||
val registry = registry(InfoFactory(null))
|
||||
val session = driving { registry.session(config) }
|
||||
|
||||
assertNull(driving { session.serverInfo() })
|
||||
|
||||
assertTrue(session.health.state.value.isUnknown, "silence should leave health unknown")
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user