mirror of
https://github.com/jmcorgan/fips.git
synced 2026-10-05 11:08:25 +00:00
The suite pinned three fixed /24s — 172.31.60, .61 and .62. Compose
project names are unique per run, so two concurrent runs got distinct
container and network *names*, but the address pools are constants and
both runs asked for the same ones. Whichever created a network first
won; the other died at topology start with
failed to create network ..._mc-far: invalid pool request:
Pool overlaps with other one on this address space
having tested nothing. Pushing maint, master and next within a second of
each other is enough to hit it, and the branch that loses looks broken
when it is fine.
The three `MC_*_PREFIX` overrides existed from the start but nothing
ever set them, so the defaults were the only values ever used.
This does what the nat suite already does. A free /24 per network is
claimed under 10.42.0.0/16 before the lab starts, with docker's own
`network create` as the atomic arbiter of who owns what — the run-id
derived offset is deliberately not used here for the same reason it was
rejected there: it makes a collision unlikely rather than impossible,
and a collision is the failure being removed. Both CI labels are stamped
so ci-cleanup.sh's label sweep recovers the networks when a run is
SIGKILLed, which no inline removal can cover.
Since the claim creates the networks, compose has to attach rather than
create, so `docker-compose.external-net.yml` declares the three
external, applied through a new `MC_EXTRA_COMPOSE` hook. Nothing else
sets it: the GitHub matrix runs one job per runner and invokes
`test.sh` directly, and a bare `docker compose up` is a single lab, so
both keep the fixed defaults and the addresses in the README stay
literal. The hook appends to the whole COMPOSE array rather than to the
`up` alone, so teardown addresses the same project — a `down` without
the overlay would not know the networks are external.
The claim lives in ci-local.sh rather than the suite script for the
reason the nat comment gives: the workflow invokes the script directly
and tears down with the base file only, and does not want a claim.
Release does a `compose down` before removing the networks. That order
is load-bearing rather than tidy: `docker network rm` silently no-ops on
a network that still has endpoints attached and reports success, so
without the `down` the removal fails exactly on the path it exists for.
A network left behind would not merely leak — the next invocation in the
run would hit `network with name ... already exists`, which is not a
pool overlap, so the allocator correctly refuses to advance and fails.
Also gives the suite its own compose project. It set none, so it
inherited whichever COMPOSE_PROJECT_NAME the previous suite exported —
which is why the failure named a medium-change network under the nat
project, `fipsci_<runid>_nat_mc-far`. That is not what caused the
overlap, but filing one suite's resources under another's project is a
teardown hazard: `down --remove-orphans` on either would consider the
other's containers orphans.
The /24 claim loop is now shared rather than copied a third time.
`ci_claim_nat_net` keeps its name, its `[nat]` log tag and its exported
prefix, and becomes a two-line caller.
Verified rather than assumed. The lab runs on the claimed prefixes, not
just alongside them: `node-b re-pinned to 10.42.1.10:2121`. With the
first three candidates occupied by squatter networks — the collision
path itself — the allocator advances to 10.42.3/4/5 and the suite passes
on those. Run standalone with no overlay it still renders 172.31.6x and
passes 8/8, so the workflow path is untouched. `nat-cone` passes and
still claims 10.41.0/1, so the shared loop did not disturb it. Networks
are gone after each run. `ci-local.sh --only medium-change` is 17/17.
Not exercised: the partial-claim rollback, which needs a /16 exhausted
part-way to reach. It mirrors ci_claim_nat_networks' shape.
33 lines
1.4 KiB
YAML
33 lines
1.4 KiB
YAML
# Override: attach the medium-change lab's three bridges to pre-created
|
|
# external networks instead of letting compose create them from the base
|
|
# file's fixed pins.
|
|
#
|
|
# Applied only by ci-local.sh's run_medium_change, which claims a free /24 per
|
|
# network per suite invocation (ci_claim_mc_networks) and exports
|
|
# FIPS_MC_PRIMARY_NET / FIPS_MC_SECONDARY_NET / FIPS_MC_FAR_NET alongside
|
|
# MC_PRIMARY_PREFIX / MC_SECONDARY_PREFIX / MC_FAR_PREFIX before `up`. That is
|
|
# what makes two concurrent local runs collision-safe: each claims distinct
|
|
# ranges, so neither requests address space the other holds.
|
|
#
|
|
# The GitHub matrix, the README invocation and any standalone
|
|
# `docker compose up` do NOT apply this overlay; they use the base file's
|
|
# normal networks with the `:-` defaults, which render today's exact
|
|
# addresses. This mirrors nat/docker-compose.external-net.yml,
|
|
# static/docker-compose.gateway-external-net.yml and
|
|
# sidecar/docker-compose.external-net.yml.
|
|
#
|
|
# One file carries all three: they are claimed and released together, and
|
|
# splitting them would let the lab attach to a mix of claimed and compose-made
|
|
# bridges — which is the state that produces a half-addressed topology rather
|
|
# than a clean failure.
|
|
networks:
|
|
mc-primary:
|
|
external: true
|
|
name: ${FIPS_MC_PRIMARY_NET:-fips-mc-primary}
|
|
mc-secondary:
|
|
external: true
|
|
name: ${FIPS_MC_SECONDARY_NET:-fips-mc-secondary}
|
|
mc-far:
|
|
external: true
|
|
name: ${FIPS_MC_FAR_NET:-fips-mc-far}
|