mirror of
https://github.com/jmcorgan/fips.git
synced 2026-10-06 03:28:24 +00:00
Every NAT rebuild is one netlink batch, sent through the default socket that rustables opens, and every message in it requested an ack. From about 105 live mappings the acks overflowed the 208 KiB receive buffer and the read failed with ENOBUFS, although the kernel had already committed the batch, so those rebuilds were logged as failed while they had taken effect. Past about 313 mappings the batch itself exceeded the send buffer, the send failed with EMSGSIZE and nothing was committed, so new .fips names got a virtual IP with no translation. Releases with the gateway through 0.5.1 sent the table delete in a batch of its own, so there a rebuild past about 313 mappings also removed the whole table. Neither errno reached the log, because the error kept only the outer message. The rebuild now encodes the batch and sends it itself. SO_SNDBUFFORCE is sized to the batch, capped at the kernel's own clamp, and a batch no send buffer can hold is refused up front. Only the last object requests an ack; an ack reader fails on any error that arrives before that ack, since the kernel still acknowledges the last message of a batch it aborted, and SO_RCVTIMEO bounds the wait. NAT errors now name the errno. The batch is still one transaction, so the table never leaves the packet path, and a failed rebuild is still only logged: the next successful rebuild installs the mapping. Unit tests cover the encoding at 2000 mappings, the send-buffer sizing and its saturation, the admission limit and the ack reader. A new gateway suite phase floods 400 names, past both old thresholds, and judges the result on the kernel's table and on the absence of NAT failure lines.