Implement hop-by-hop ECN congestion signaling through the FMP layer,
transport-level congestion detection via kernel drop counters, and
chaos harness integration for end-to-end validation.
FMP/session ECN plumbing:
- Thread ce_flag parsed at link layer through dispatch_link_message,
handle_session_datagram, handle_session_payload, and
handle_encrypted_session_msg to session delivery
- Replace hardcoded false in session-layer record_recv() with actual
ce_flag, activating ecn_ce_count tracking in session MMP
ECN congestion detection and CE relay:
- Add EcnConfig (node.ecn.*) with configurable loss_threshold (5%)
and etx_threshold (3.0) for transit congestion detection
- Add send_encrypted_link_message_with_ce() that ORs FLAG_CE into FMP
header flags; original method delegates with ce_flag=false
- Compute outgoing_ce = incoming_ce || local congestion on next-hop
link, enabling hop-by-hop CE relay through transit nodes
IPv6 ECN-CE marking:
- Mark ECN-CE (0b11) in IPv6 Traffic Class on received DataPackets
before TUN delivery when FMP CE flag is set
- Only marks ECN-capable packets (ECT(0)/ECT(1)); Not-ECT packets
unchanged per RFC 3168
Transport congestion abstraction and UDP kernel drop detection:
- Add TransportCongestion struct to transport layer for transport-
agnostic local congestion indicators
- Replace tokio::UdpSocket with AsyncFd<socket2::Socket> using
libc::recvmsg() with ancillary data parsing
- Enable SO_RXQ_OVFL for kernel receive buffer drop counter on every
packet, wiring up previously-stubbed UdpStats.kernel_drops
- Add TransportDropState for per-transport delta tracking with 1s
tick sampling via sample_transport_congestion()
- Extend detect_congestion() with transport kernel drop check
alongside MMP loss metrics
Congestion monitoring and control:
- Add CongestionStats (ce_forwarded, ce_received, congestion_detected,
kernel_drop_events) to NodeStats with snapshot serialization
- Wire counters into forwarding path, session handler, and transport
drop sampling with rate-limited warn logging (5s interval)
- Expose congestion data in show_routing control query and
ecn_ce_count in show_mmp peer entries
- Add congestion counters to fipstop routing tab in two-column layout
Chaos harness integration:
- Add query_routing(), query_transports(), snapshot_all_congestion()
to chaos control module
- Add congestion/kernel-drop log analysis in logs module
- Add congestion-stress scenario: 10-node tree, 1 Mbps bandwidth,
5-10% netem loss, heavy iperf3 traffic
- Add IngressConfig for tc ingress policing with per-peer policer
filters simulating upstream bandwidth bottlenecks
- Add iperf3 JSON result capture to traffic manager for throughput
measurement across scenarios
- Add ECN A/B test scenarios (ecn-ab-on/off.yaml) with ingress
policing and comparison script
- Enable TCP ECN negotiation (tcp_ecn=1 sysctl) in container
entrypoint for end-to-end CE propagation
Tests:
- 10 ECN unit/integration tests: mark_ipv6_ecn_ce variants, CE relay
chain (3-node propagation), EcnConfig serde roundtrip
- 3 transport drop congestion detection unit tests
Documentation:
- Update fips-mesh-layer.md: replace outdated CE Echo stub with full
ECN Congestion Signaling section covering detection logic, CE relay,
IPv6 marking, session tracking, and monitoring counters
- Update fips-configuration.md: add node.ecn.* parameter table and
ecn block in complete reference YAML
- Update fips-transport-layer.md: add Congestion Reporting section
with TransportCongestion struct, congestion() trait method, and
per-transport status; document AsyncFd/recvmsg/SO_RXQ_OVFL in UDP
- Update chaos README: add congestion/ECN scenario docs, ingress
traffic control, and iperf3 JSON capture sections
- Update README.md: add ECN to features list and "What works today";
update transport and tooling entries
Track consecutive send failures in SenderState. Apply 2^n backoff
multiplier (capped at 32x) to the report interval. Suppress debug
logs after 3 consecutive failures, emit recovery summary on success.
Add a Unix domain socket interface for querying node state at runtime.
A spawned tokio task accepts connections and communicates with the main
event loop via mpsc/oneshot channels, keeping all Node access
single-threaded.
Includes:
- src/control/ module with socket lifecycle, JSON protocol, and 11
query handlers (status, peers, links, tree, sessions, bloom, mmp,
cache, connections, transports, routing)
- Separate fipsctl binary for CLI queries (fipsctl show <command>)
- ControlConfig in node configuration (enabled, socket_path)
- Integration into the main select! event loop
Rename FIPS Link Protocol (FLP) to FIPS Mesh Protocol (FMP)
The "Link Protocol" name understated the layer's scope — spanning tree
construction, bloom filter routing, greedy forwarding, and mesh-wide
coordination go well beyond link-level concerns. Rename fips-link-layer.md
to fips-mesh-layer.md, update FLP→FMP throughout docs and source code
(FLP_VERSION→FMP_VERSION, wire.rs, rx_loop.rs, spanning_tree.rs).
New SVG illustrations
- Protocol stack: color-coded layer diagram replacing ASCII art
- OSI mapping: side-by-side comparison with traditional networking layers
- Bloom filter propagation: 6-node tree with sender-colored filter boxes
showing split-horizon computation per link
- Routing decision flowchart: 5-step priority chain with candidate ranking
by tree distance and link performance
- Coordinate discovery: sequence diagram showing LookupRequest propagation,
response caching, and SessionSetup cache warming
Redesigned existing SVGs
- Architecture overview: uniform node layout, U-shaped encrypted link
connectors, separate end-to-end session line
- Node architecture: split Router Core into FSP and FMP layers, reorganize
transports into Overlay/Shared Medium/Point-to-Point categories
- Identity derivation: wider boxes, visible encode arrow, dashed npub line
fips-intro.md revisions
- Add inline references to prior work: Yggdrasil/Ironwood for coordinate
routing, Noise Protocol Framework for IK handshakes, WireGuard for
index-based session dispatch, Wikipedia for bloom filters, split-horizon,
and greedy embedding
- Add explanatory paragraphs after bloom filter diagram describing
split-horizon filter computation and candidate selection behavior
- Simplify transport abstraction language, remove I2P/LoRa references
- Fix LookupRequest wording ("propagates" not "floods"), note intermediate
node coordinate caching on lookup responses
- Rewrite architecture overview prose to match redesigned diagrams
Sends a 1-byte encrypted heartbeat (0x51) to each peer every 10s.
If no frame is received from a peer within 30s, the peer is removed
via remove_active_peer(), triggering tree reconvergence, coord cache
flush, and bloom filter recomputation.
This fixes the critical bug where UDP peers that silently died
(e.g., container stopped) were never detected or removed, leaving
the spanning tree permanently stale.
Both intervals are configurable via node.heartbeat_interval_secs
and node.link_dead_timeout_secs.
Per-packet happy-path events (UDP send/receive, TUN I/O, MMP report
processing, TreeAnnounce/FilterAnnounce sent, RTT samples) moved from
debug to trace. Periodic maintenance and retry scheduling moved from
info to debug. Session state changes (established, initiated, torn
down) and transport stop promoted from debug to info. MMP report send
failures demoted from warn to debug (normal under churn).