fwdelta: Semantic Diff for Firewall Policy
The problem
A firewall ruleset is reviewed as text. A firewall behaves as a set of permitted packets. Nothing in the review process connects the two. Rules are evaluated first-match, and that single property makes every edit non-local: narrowing a source range on rule 14 can expose rule 22, which was previously shadowed and therefore dead, and which now matches traffic nobody has thought about in two years. A reviewer looking at a unified diff sees one line changed. The second-order effects are not in the diff, they are in the interaction between the changed line and the two hundred lines around it.
What it does
fwdelta compiles each ruleset into the set of packets it permits, encoded as a binary decision diagram over a 120-bit header with seven match dimensions. Comparing two rulesets is then set algebra rather than string comparison, which is what makes the non-local effects visible.
The output has three parts. NEWLY BLOCKED and NEWLY ALLOWED name the packet ranges whose verdict changed, each attributed to the rule that used to decide it and the rule that decides it now. STRUCTURAL reports the rules whose role changed without their text changing: a rule that became load-bearing, a rule that became reachable because the rule shadowing it was narrowed.
Reports render as text or as a single HTML file with all CSS inline, no scripts, no remote fonts and no images fetched over the network. A report about an air-gapped network's firewall is not much use if reading it needs a CDN, and one that phones home when opened is an exfiltration path for a document describing exactly where the trust boundaries are. CI greps the generated file for every one of those patterns.
How the model is checked against reality
A verification tool that is wrong is worse than no tool, because it converts uncertainty into false confidence. The differential harness loads generated rulesets into an unprivileged network namespace, sends real packets and reads per-rule counters for the kernel's actual verdict. Disagreement is a failing test with a reproducible seed.
The harness is itself tested: a self-test mode injects five deliberately broken models and requires every one to be detected. A fault that survives means the dimension it breaks is untested, which is why oifname is rejected by the frontend rather than shipped — the harness cannot exercise the output hook and so cannot falsify it.
Separately, a sweep runs every hook against every match type with real nftables and counters. Two combinations load and are then silently ignored by the kernel; both are rejected by the frontend. That class of bug is invisible to the differential harness, which generates its rulesets from the same model it is testing.
Where it runs
One statically linked musl binary with no runtime dependencies. The dependency graph bans network-capable crates by name, and a syscall audit straces the binary and fails on a single socket call. Neither mechanism is a proof on its own and the project does not claim one: the denylist matches crate names and cannot detect capability, the strace only covers the paths the audit run reached.
The build is reproducible. The toolchain is pinned and the dependency graph committed, a script builds twice and requires identical digests, and the digest for the published binary is in the README for a third party to compare against.
Scope and limits
Every project page on this site carries this section. If a tool is not ready for something, the honest place to say so is next to the claim, not three pages into a README.
- Single base chain only. jump, goto, return and any user-defined chain are rejected at exit 2, including a helper chain nothing jumps to. Docker adds chains, so a good many real rulesets are refused outright rather than analysed.
- Connection tracking is rejected, not approximated. Any ct match of any kind refuses the whole file. The model is stateless and assumes return traffic for permitted connections is permitted.
- NAT is rejected. Address translation changes packet identity in transit, so a ruleset containing it is refused rather than analysed with NAT ignored.
- IPv4 only. ip6, inet, arp, bridge and netdev tables are rejected; the header layout is 32-bit.
- One host's filter table, not end-to-end reachability. No topology, no routing table, no BGP or OSPF. For multi-hop reachability the right tool is Batfish.
- The parser has been tested against nftables v1.0.9 only. A dump from a different nft may use syntax it has never seen. It fails loudly rather than silently, but the coverage claim is one version wide.
- Analysis cost is roughly quadratic in ruleset size. At 5,000 rules a diff takes about 73 seconds.
Building something in this territory, or hiring for it?