Files
2026-05-08 10:54:03 -04:00

235 lines
8.2 KiB
Markdown

# FIPS Debug Notes: `.fips` HTTP reachability incident
## Summary
We investigated why this URL failed locally but worked from the remote host:
- `http://npub1crpldvy49ef8z34wlacwujnfudy4nd7k96aqdx5wgn6ckztz7z8q9t59ud.fips/index.html`
Outcome:
1. Initial local failure was due to a stale `fd00::/8` route pointing to an unreachable next-hop.
2. After route recovery, TCP connected but larger HTTP responses stalled.
3. Root cause for the stall was path/MTU behavior (large payload transfer blackholing/retransmission).
4. Setting local UDP transport MTU to `1280` resolved the issue; full page download succeeded.
---
## What we observed
### 1) DNS was correct
- `.fips` name resolved to the expected IPv6 address.
- Name resolution was **not** the issue.
### 2) Endpoints were alive
- On remote host (`ubuntu@laantungir.net`), the same URL returned `HTTP/1.1 200 OK` with full content.
- Nginx was listening and serving properly.
### 3) Local routing was initially broken
- Local route had `fd00::/8` via a link-local next-hop that was unreachable.
- Neighbor state for that gateway showed `FAILED`.
- This produced local `No route to host` errors.
### 4) After local route fix, only large responses failed
- Local curl connected to remote `:80` and sent GET.
- Small responses worked:
- direct IPv6 host request without vhost Host header got `410` small body,
- ranged request got `206` with a small payload.
- Full `index.html` (10KB+) timed out with `0 bytes received` locally.
- Remote socket stats indicated retransmissions and congestion-window collapse, consistent with MTU/path fragmentation loss behavior.
### 5) Mitigation validated
- Local config changed:
```yaml
transports:
udp:
bind_addr: 0.0.0.0:2121
mtu: 1280
```
- Restarted local `fips.service`.
- Retested URL: `HTTP 200`, full `10156` bytes downloaded.
---
## Lessons learned
1. **Treat `.fips` failures as a pipeline problem**
- Validate in this order: DNS → route → reachability → app response size behavior.
2. **“Connect succeeds” does not prove path health**
- A successful TCP handshake can still hide payload delivery failures.
- Always test both tiny and realistic/full payload responses.
3. **Stale `fd00::/8` routes can silently break TUN startup**
- We saw TUN init warnings when route insertion conflicted with existing state.
- Route hygiene matters before/after service restarts.
4. **MTU defaults are not always safe on every real path**
- Especially over mixed internet routes and overlays.
- `1280` is often the practical “known-good” baseline for IPv6 paths.
5. **Use control-plane telemetry early**
- `fipsctl show status|peers|sessions|routing` quickly distinguishes control-plane up/down from data-plane degradation.
6. **Host-header/vhost behavior can mislead debugging**
- Explicit `Host:` header to the same IPv6 returned 200 remotely.
- Direct IP without host header returned a different vhost response (410).
- This is expected and should be factored into tests.
---
## Recommended runbook for future incidents
1. **Reproduce quickly**
- `curl -v --max-time 10 http://<npub>.fips/index.html`
2. **Check resolution and route**
- `getent ahosts <npub>.fips`
- `ip -6 route get <resolved_ipv6>`
- `ip -6 neigh show`
3. **Check FIPS health**
- `sudo /usr/local/bin/fipsctl show status`
- `sudo /usr/local/bin/fipsctl show peers`
- `sudo /usr/local/bin/fipsctl show sessions`
4. **Differentiate small vs large payload behavior**
- Small: direct IP / tiny endpoint / range request
- Large: full page/object transfer
5. **If large payloads fail but small succeed, test MTU mitigation**
- Set `transports.udp.mtu: 1280` locally
- Restart `fips.service`
- Retest full payload
6. **Persist fix and monitor**
- Keep 1280 unless measured path supports higher values reliably.
- Track retransmission symptoms and transfer success over time.
---
## Config change made during this debug
Local `/etc/fips/fips.yaml` was updated to include:
```yaml
transports:
udp:
mtu: 1280
```
A backup of the prior local file was created before modification.
---
## Final status
The target `.fips` web page is reachable from this local node after lowering local UDP MTU to `1280` and restarting `fips.service`.
---
## Additional root cause discovered later (boot-time route collision)
A subsequent incident showed another failure mode that can present as the same `curl: (7) Failed to connect` symptom:
1. A stale kernel route existed: `fd00::/8 via fe80::... dev eth0`.
2. During startup, `fips.service` attempted to install `fd00::/8 dev fips0`.
3. Kernel returned `EEXIST`, so TUN initialization failed (`tun_state: failed`, no `fips0`).
4. Without TUN, `.fips` reachability failed even if service looked partially alive.
### Key diagnostic indicators
- `sudo fipsctl show status` reports:
- `tun_state: failed`
- missing/empty `tun_name`
- `ip -6 route show fd00::/8` points to non-`fips0` device (for example `eth0`)
- `ip -o link show dev fips0` fails
### Immediate runtime recovery
```bash
sudo ip -6 route del fd00::/8
sudo systemctl restart fips.service
```
After recovery, verify:
- `tun_state: active`
- `fips0` exists and is UP
- `ip -6 route show fd00::/8` -> `dev fips0`
### Persistence/hardening
To make this robust across reboots and service restarts:
1. Keep config MTU safeguards:
- `tun.mtu: 1280`
- `transports.udp.mtu: 1280`
2. Ensure `fips.service` is enabled:
- `sudo systemctl enable fips.service`
3. Add a systemd drop-in `ExecStartPre` hook that removes stale non-`fips0` `fd00::/8` routes before daemon start.
This repository now includes deployment-script updates to install that hook automatically.
---
## Latest incident addendum (Qubes inter-qube `.fips` debug)
### New findings
1. **Mesh + DNS can be healthy while app check still fails**
- We observed successful `.fips` DNS resolution and ICMP reachability to the remote FIPS IPv6.
- Yet HTTP checks still failed due to listener/bind mismatch.
2. **Qubes resolver path must include local dnsmasq for `.fips`**
- `dig @127.0.0.1 -p 5354 <npub>.fips AAAA` worked while `getent ahosts <npub>.fips` initially failed.
- Root cause: `/etc/resolv.conf` lacked `nameserver 127.0.0.1` on the local qube.
- Adding `127.0.0.1` restored system resolver `.fips` lookups.
3. **Persistence in Qubes requires `/rw/config/rc.local` hygiene**
- Runtime fixes to `/etc/resolv.conf` are not enough.
- Persisting the resolver injection in `/rw/config/rc.local` is required for reboot durability.
- We also discovered/remediated a stale legacy block (`# === AI FIPS ROUTE START ===`) that forced `fd00::/8` via `eth0`; this must be removed to avoid breaking TUN routing on boot.
4. **Remote service bind must match transport family used by `.fips` path**
- Remote python server bound to `0.0.0.0:80` (IPv4) was not reachable via `.fips` IPv6 destination.
- Rebinding server to remote FIPS IPv6 (`fd56:...:982f`) fixed HTTP connectivity immediately.
### Practical runbook update for inter-qube checks
Add these checks before concluding “FIPS is broken”:
1. **Resolver path sanity**
- `getent ahosts <npub>.fips`
- If fail, check `/etc/resolv.conf` includes `nameserver 127.0.0.1`
- Confirm dnsmasq forwards `/fips/` to local FIPS DNS (`127.0.0.1#5354` or `::1#5354`)
2. **Control/data plane split checks**
- DNS AAAA success + `ping6 <npub>.fips` success proves overlay path exists
- If HTTP fails, test TCP port explicitly (`/dev/tcp` or `curl -v -6`)
3. **Listener-family verification on remote**
- `ss -lntp | grep ':80 '`
- Ensure remote listener is IPv6-capable for `.fips` access
### Script updates from this incident
- `test_qube_pair.sh`
- Added resolver hints when `.fips` name does not resolve
- Added IPv6 curl (`curl -6`) and configurable HTTP port via `REMOTE_HTTP_PORT`
- Added port probe diagnostics on HTTP failure
- `test_fips.sh`
- Added `/etc/resolv.conf` sanity reporting for `nameserver 127.0.0.1`
- `update_and_deploy_fips.sh`
- Persisted rc.local network overrides
- Added cleanup of legacy `# === AI FIPS ROUTE START/END ===` block when rewriting rc.local
- `start_py_http80.sh`
- Binds to detected local FIPS IPv6 by default (instead of IPv4-only bind)