Files
amethyst/cli/tests/cordn/migrate.sh
T
Vitor PamplonaandClaude Opus 5 d627ba7846 fix(cordn): a handshake that answers is a call, and two counts of one
Three things an afternoon of driving the cordn screens on a tablet turned up.

**"Ask who it is" left the coordinator card contradicting itself.** It printed
"Nothing asked of it yet." directly above "Says it is: cordn-server · 0.1.0",
because `health.recordSuccess` was only reached from `CordnGroupManager.call`
and `serverInfo()` goes to the coordinator through `CoordinatorClient` without
passing that way. The screen's own comment says health is "only ever what a
call of ours already observed", and this is one — usually the very first,
since the button exists so someone can find out whether a coordinator they
just added answers at all. Recorded in `CordnSession.serverInfo` now, where
the session already holds the health it feeds.

An answer counts as a success and a throw as a failure; `null` counts as
neither, because a scope that did no handshake returns `null` and a call that
never left the device must not report a coordinator as reachable.

**"1 single-use packages available."** — and "1 packages this device cannot
open". Both are now plurals.

Three tests cover the health path: an answer marks it observed, a throw counts
a failure and keeps its reason, and silence leaves it unknown.

---

The live migrate harness could not run on this machine at all; none of it was
the product.

- `stack.sh` ran the coordinator with `--network host` so it could reach a
  relay on the host's loopback. On Docker Desktop the daemon is inside a VM,
  so that is the *VM's* loopback and `127.0.0.1:$PORT` is nothing: the
  coordinator retried "Relay connection error" until `stack_up` gave up
  waiting for a pubkey that was never coming. It now uses
  `host.docker.internal` off Linux and keeps `--network host` on it.
- `BLOB_PORT` defaulted to **877**, which is privileged, so the blob
  stand-in died with `PermissionError` for anyone who is not root — on Linux
  too. Now 8877.
- `blob_up` did not check the server came up. A dead one surfaced thirty lines
  later as `migrate export` reporting "no server accepted the document", plus
  two more failures behind it, none naming the blob server. It now waits for
  the port and exits 2 with the real error.

With those, `cli/tests/cordn/migrate.sh` passes end to end on stock defaults.
Worth knowing for whoever runs it next: the image is amd64-only, so on Apple
Silicon the pull needs `--platform linux/amd64` and then runs emulated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 18:59:09 -04:00

179 lines
7.0 KiB
Bash
Executable File

#!/usr/bin/env bash
#
# migrate.sh — one account, two devices, a handoff between them.
#
# The claim tier-b.sh does not test. There, two accounts talk to each other.
# Here ONE account moves from the phone it is on to a new one, which is the
# case `spec/applications/multi-device.md` calls device addition (§11) and we
# implement as a one-shot handoff rather than continuous sync.
#
# What this proves that the unit tests cannot:
#
# * the sealed documents survive a real HTTP round trip to a blob server,
# * the tip survives a real relay as a real replaceable event,
# * the new device, starting from an empty home and a scanned string alone,
# ends up holding the same groups.
#
# The blob server is a throwaway python stand-in for Blossom: PUT /upload
# stores by sha256, GET /<sha256> serves it back. Enough to exercise the real
# HttpCordnBlobStore, and deliberately not a Blossom implementation.
#
# Usage: cli/tests/cordn/migrate.sh
set -uo pipefail
WORK="${WORK:-$(mktemp -d)}"
CONTAINER="cordn-migrate"
# shellcheck source=stack.sh
. "$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)/stack.sh"
export AMY_PASSPHRASE="${AMY_PASSPHRASE:-migrate}"
# Above 1024: binding below it needs root, and this test has no business
# asking for that. 877 was the default and bound for nobody.
BLOB_PORT="${BLOB_PORT:-8877}"
BLOB="http://127.0.0.1:$BLOB_PORT"
BLOB_DIR="$WORK/blobs"
BLOB_PID=""
fail=0
step() { echo; echo "── $*"; }
ok() { echo " ok: $*"; }
bad() { echo " FAIL: $*"; fail=1; }
# Two HOMEs, one identity: the same nsec on an old phone and a new one. That is
# what a migration is, and it is why the new device needs no Welcome — it
# adopts the shared leaf rather than joining as a member (§9).
old() { HOME="$WORK/old" "$AMY" --account me --secret-backend ncryptsec "$@" 2>/dev/null; }
new() { HOME="$WORK/new" "$AMY" --account me --secret-backend ncryptsec "$@" 2>/dev/null; }
field() { python3 -c "import json,sys; d=json.load(sys.stdin); print(json.dumps(d$1) if not isinstance(d$1,str) else d$1)"; }
blob_up() {
mkdir -p "$BLOB_DIR"
python3 - "$BLOB_PORT" "$BLOB_DIR" >"$WORK/blob.log" 2>&1 &
BLOB_PID=$!
# Wait for the port, and say so here if it never opens. Letting a dead blob
# server through costs three misleading failures later — export reports "no
# server accepted the document", and the group/blob assertions all fall over
# behind it — none of which name the thing that is actually wrong.
for _ in $(seq 20); do
if python3 -c "import socket,sys; s=socket.socket(); s.settimeout(0.3); sys.exit(0 if s.connect_ex(('127.0.0.1',$BLOB_PORT))==0 else 1)"; then
return 0
fi
kill -0 "$BLOB_PID" 2>/dev/null || break
sleep 0.5
done
echo "the blob server never came up on $BLOB — see $WORK/blob.log"
tail -3 "$WORK/blob.log" 2>/dev/null
exit 2
} <<'PYEOF'
import hashlib, os, sys
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
port, root = int(sys.argv[1]), sys.argv[2]
class H(BaseHTTPRequestHandler):
def do_PUT(self):
body = self.rfile.read(int(self.headers.get("Content-Length", 0)))
digest = hashlib.sha256(body).hexdigest()
open(os.path.join(root, digest), "wb").write(body)
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.end_headers()
self.wfile.write(('{"sha256":"%s","size":%d}' % (digest, len(body))).encode())
def do_GET(self):
path = os.path.join(root, self.path.lstrip("/"))
if not os.path.isfile(path):
self.send_response(404); self.end_headers(); return
body = open(path, "rb").read()
self.send_response(200)
self.send_header("Content-Length", str(len(body)))
self.end_headers()
self.wfile.write(body)
def log_message(self, *a): pass
ThreadingHTTPServer(("127.0.0.1", port), H).serve_forever()
PYEOF
cleanup() {
[ -n "$BLOB_PID" ] && kill "$BLOB_PID" 2>/dev/null
stack_down
}
trap cleanup EXIT
stack_require
step "boot geode on $RELAY, the reference coordinator, and a blob server on $BLOB"
stack_up
blob_up
ok "coordinator $COORD"
step "one account, on its old phone"
mkdir -p "$WORK/old" "$WORK/new"
old create --json >/dev/null
PK=$(old whoami --json | field "['hex']")
ok "npub $PK"
# The new phone is the SAME identity. Copy the key across the way a user would
# by signing in again; the migration carries group state, never identity (§4.3).
NSEC=$(old key export --json 2>/dev/null | field "['nsec']" 2>/dev/null)
if [ -z "$NSEC" ]; then
cp -r "$WORK/old/.amy" "$WORK/new/.amy"
ok "new phone signed in (identity copied)"
else
new import --nsec "$NSEC" --json >/dev/null
ok "new phone signed in"
fi
step "the old phone has a group"
old cordn coordinator add --coordinator "$COORD" --relay "$RELAY" --label migrate --json >/dev/null
GROUP=$(old cordn group create --name "Moves with me" --about "handoff" --json)
GID=$(echo "$GROUP" | field "['gid']")
[ -n "$GID" ] && ok "gid $GID" || bad "no group to migrate"
step "the old phone exports a handoff"
EXPORT=$(old cordn migrate export --relay "$RELAY" --server "$BLOB" --json)
echo " $EXPORT"
CODE=$(echo "$EXPORT" | field "['code']")
if [ -n "$CODE" ] && [ "${CODE:0:9}" = "cordndev1" ]; then ok "code minted"; else bad "no handoff code"; fi
[ "$(echo "$EXPORT" | field "['groups']")" = "1" ] && ok "one group in the snapshot" || bad "wrong group count"
[ "$(echo "$EXPORT" | field "['uploaded']")" = "true" ] && ok "documents stored" || bad "nothing stored"
step "blobs really landed on the server"
COUNT=$(ls "$BLOB_DIR" 2>/dev/null | wc -l | tr -d ' ')
# One group document plus the meta document.
[ "$COUNT" -ge 2 ] && ok "$COUNT blobs" || bad "expected >=2 blobs, found $COUNT"
step "the new phone has nothing yet"
BEFORE=$(new cordn group list --json 2>/dev/null | field "['groups']" 2>/dev/null || echo "[]")
[ "$BEFORE" = "[]" ] && ok "empty" || echo " (starting from: $BEFORE)"
step "the new phone imports, from the code alone"
IMPORT=$(new cordn migrate import --code "$CODE" --json)
echo " $IMPORT"
[ "$(echo "$IMPORT" | field "['groups']")" = "1" ] && ok "one group adopted" || bad "group did not arrive"
[ "$(echo "$IMPORT" | field "['replaced_local_state']")" = "true" ] && ok "replaced, not merged" || bad "did not replace"
step "the group is really there, with the same gid"
AFTER=$(new cordn group list --json 2>/dev/null)
echo " $AFTER"
echo "$AFTER" | grep -q "$GID" && ok "gid $GID on the new phone" || bad "gid missing after import"
step "a group ref is not a handoff code"
# Called directly rather than through new(): the --json error contract writes
# to stderr, which the helper sends to /dev/null. The first version of this
# check compared an empty string and passed for the wrong reason.
STRANGER=$(HOME="$WORK/new" "$AMY" --account me --secret-backend ncryptsec \
cordn migrate import --code "cordn1qqqqqq" --json 2>&1 || true)
echo " $STRANGER"
echo "$STRANGER" | grep -q "bad_args" && ok "refused" || bad "accepted a non-handoff code"
echo
if [ "$fail" -eq 0 ]; then
echo "── migrate: PASS"
else
echo "── migrate: FAIL"
fi
exit "$fail"