@push.rocks/smartnftables 2.5.2 → 3.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. package/changelog.md +31 -0
  2. package/dist_rust/smartnftables_linux_amd64_musl +0 -0
  3. package/dist_rust/smartnftables_linux_amd64_musl.tsrust-build.json +5 -5
  4. package/dist_rust/smartnftables_linux_arm64_musl +0 -0
  5. package/dist_rust/smartnftables_linux_arm64_musl.tsrust-build.json +5 -5
  6. package/dist_ts/00_commitinfo_data.js +1 -1
  7. package/dist_ts/managed.egress.types.d.ts +69 -3
  8. package/package.json +6 -6
  9. package/readme.md +342 -61
  10. package/rust/src/egress.compile.rs +157 -76
  11. package/rust/src/egress.host.rs +104 -44
  12. package/rust/src/egress.hostgrant.rs +14 -14
  13. package/rust/src/egress.keys.rs +305 -0
  14. package/rust/src/egress.leased.rs +399 -0
  15. package/rust/src/egress.poolguard.rs +45 -0
  16. package/rust/src/egress.private.rs +197 -0
  17. package/rust/src/egress.published.rs +370 -0
  18. package/rust/src/egress.router.rs +64 -295
  19. package/rust/src/egress.rs +248 -17
  20. package/rust/src/egress.workloadgrant.rs +88 -0
  21. package/rust/src/egress_hostgrant_tests.rs +196 -47
  22. package/rust/src/egress_localplatform_tests.rs +204 -0
  23. package/rust/src/egress_localport_tests.rs +266 -0
  24. package/rust/src/egress_publishedrange_tests.rs +812 -0
  25. package/rust/src/egress_routerequivalence_tests.rs +448 -0
  26. package/rust/src/egress_tests.rs +132 -230
  27. package/rust/src/egress_workloadgrant_tests.rs +313 -0
  28. package/rust/src/owner.rs +77 -15
  29. package/rust/src/owner_host_traffic_tests.rs +4 -0
  30. package/rust/src/owner_hostgrant_traffic_tests.rs +9 -4
  31. package/rust/src/owner_identity_tests.rs +228 -80
  32. package/rust/src/owner_localplatform_tests.rs +274 -0
  33. package/rust/src/owner_localport_tests.rs +241 -0
  34. package/rust/src/owner_publishedrange_traffic_tests.rs +306 -0
  35. package/rust/src/owner_router_traffic_tests.rs +2 -0
  36. package/rust/src/owner_scale_traffic_tests.rs +383 -0
  37. package/rust/src/owner_tests.rs +15 -0
  38. package/rust/src/owner_workloadgrant_tests.rs +278 -0
  39. package/rust/src/policy.rs +29 -25
  40. package/ts/00_commitinfo_data.ts +1 -1
  41. package/ts/managed.egress.types.ts +71 -3
package/readme.md CHANGED
@@ -264,9 +264,9 @@ allocation-pool guard policy kinds.
264
264
 
265
265
  | V2 scope | Required authority and behavior |
266
266
  | --- | --- |
267
- | `routerEgress` | Private `endpoints` and `rules`, one exact `links` binding per endpoint, a separate veth `handoff`, `protection`, and active `generations`. Private veth/TUN/local DNS and egress share one table so terminal private denial cannot override a separate egress table. Optional `publishedPorts` add the inbound second hop from the handoff to a workload endpoint; both directions are classified into the default conntrack zone ahead of every leased classifier, so a published endpoint port is dedicated to its publication and never becomes leased egress. Optional `hostGrants` forward exact host-origin flows from the handoff to a workload endpoint. |
268
- | `hostTransit` | Exact handoff `link`/`allocations` pairs, complete `protection`, an explicit veth or Ethernet `uplink`, and its current `snatAddress`. It checks each handoff's leased source address and protocol/port range, default conntrack zone, direction, uplink, and protected destinations before outer SNAT. Optional `publishedPorts` add inbound uplink destination NAT inside the same generation. Optional `hostGrants` let the host's own address on a handoff dial exact workload ports. |
269
- | `allocationPoolGuard` | An authenticated `authorityDigest` and complete current allocation-pool `prefixes`. Installs host-wide IPv4 destination denial before any handoff exists, without link, uplink or SNAT dependencies. Optional `hostGrants` are its only exceptions. |
267
+ | `routerEgress` | Private `endpoints` and `rules`, one exact `links` binding per endpoint, a separate veth `handoff`, `protection`, and active `generations`. Private veth/TUN/local DNS and egress share one table so terminal private denial cannot override a separate egress table. Optional `publishedPorts` add the inbound second hop from the handoff to a workload endpoint, one port or a port range; both directions are classified into the default conntrack zone ahead of every leased classifier, so a published endpoint port is dedicated to its publication and never becomes leased egress. A `symmetric` publication also lets the workload open flows from its published ports. Optional `hostGrants` forward exact host-origin flows from the handoff to a workload endpoint. Optional `workloadGrants` let one workload open one exact port of another, one way. |
268
+ | `hostTransit` | Exact handoff `link`/`allocations` pairs, complete `protection`, an explicit veth or Ethernet `uplink`, and its current `snatAddress`. It checks each handoff's leased source address and protocol/port range, default conntrack zone, direction, uplink, and protected destinations before outer SNAT. Optional `publishedPorts` add inbound uplink destination NAT inside the same generation, one port or a port range, and a `symmetric` publication also carries the workload's own flows from its published ports out through the uplink. Optional `hostGrants` let the host's own address on a handoff dial exact workload ports. Optional `localPlatformEndpoints` serve platform endpoints on the host's own addresses to leased flows. |
269
+ | `allocationPoolGuard` | An authenticated `authorityDigest` and complete current allocation-pool `prefixes`. Installs host-wide IPv4 destination denial before any handoff exists, without link, uplink or SNAT dependencies. Optional `hostGrants` are its only exceptions. Optional `localTcpPortOwners` restrict loopback TCP ports to one local user each. |
270
270
 
271
271
  `allocationPoolGuard` accepts 1–64 canonical, disjoint RFC1918 prefixes. Supply
272
272
  the actual allocation pools, not the broader protected union containing management
@@ -317,7 +317,8 @@ handoffs, unmatched handoff traffic and IPv6 forwarding are denied.
317
317
 
318
318
  `hostTransit.publishedPorts` is optional. Absent and empty are the same canonical
319
319
  policy, so an unpublished scope keeps its exact previous digest and compiled bytes.
320
- Each entry is `{ protocol, hostPort, targetPort, targetAddress, hostIp }`. `hostIp`
320
+ Each entry is `{ protocol, hostPort, targetPort, targetAddress, hostIp }`, with an
321
+ optional `hostPortEnd` and an optional `symmetric`. `hostIp`
321
322
  must be an exact current uplink address, which rtnetlink verifies at apply, recovery
322
323
  and inspection; there is no wildcard, secondary-link or default-route inference.
323
324
  `targetAddress` must be a current leased `transitSourceAddress` of one allocation,
@@ -328,21 +329,59 @@ including across uplink addresses, and a port published on the `snatAddress` may
328
329
  fall inside any leased outbound source-port range of the same protocol, because outer
329
330
  SNAT translates to that same address.
330
331
 
331
- A publication compiles inside the generation: one destination NAT rule in this
332
- table's own prerouting `nat` chain at priority -100, plus three forward admissions
333
- placed ahead of the capture barrier — original direction NEW and ESTABLISHED, with
334
- opening TCP restricted to SYN with FIN/RST/ACK clear, and the ESTABLISHED reply.
335
- Every admission checks both links, the conntrack protocol, the default zone, the
336
- direction, the translated current address/port and the original conntrack
337
- destination, so foreign NAT cannot borrow the exception for another flow. Published
338
- traffic is not source-translated, so the workload observes the actual client address.
339
- The chain, its translation and its admissions are created, replaced and deleted with
340
- the generation: a target without the entry, a failed batch, or release removes them,
341
- and a caller never needs a separate teardown.
342
-
343
- Input bounds are 64 entries, but the atomic graph reserve binds first. A publication
344
- costs four operations and about 6,000 bytes, so roughly a dozen fit beside a small
345
- host-transit graph and `prepare()` rejects an excessive set before any kernel work.
332
+ `hostPortEnd` publishes the range `hostPort..hostPortEnd`. It must be greater than
333
+ `hostPort`, and `targetPort` must equal `hostPort`: a range keeps every port, so port
334
+ `p` of the uplink reaches port `p` of `targetAddress`. One port has exactly one
335
+ canonical form, without `hostPortEnd`. Ranges follow the same rules as single ports:
336
+ no two publications of one protocol overlap anywhere in the scope, and no published
337
+ port on the `snatAddress` falls inside a leased range. A range translates the
338
+ address only, which never changes a port.
339
+
340
+ Publications are set elements, not rules. The generation compiles CONSTANT named
341
+ sets and maps of concatenated fields in which the port fields are inclusive
342
+ intervals (the kernel's pipapo backend), so one element holds one port or a whole
343
+ range: `published_port` maps a one-port publication's uplink address, protocols and
344
+ port to its target address and port, `published_address` maps a range to its target
345
+ address alone, and `published_forward` holds every publication's inbound tuple. A
346
+ constant number of rules looks them up whatever the number of publications: one
347
+ destination NAT rule per map in this table's own prerouting `nat` chain at priority
348
+ -100, and four forward admissions ahead of the capture barrier — the ESTABLISHED
349
+ original direction, a NEW opening per protocol with opening TCP restricted to SYN
350
+ with FIN/RST/ACK clear, and the ESTABLISHED reply. Every rule checks the uplink, the
351
+ default zone and the direction, and every key carries the exact handoff (index and
352
+ name), the packet and conntrack protocols, the translated current address and port
353
+ and the original conntrack destination, so foreign NAT cannot borrow the exception
354
+ for another flow. Published traffic is not source-translated, so the workload
355
+ observes the actual client address. The chain, its sets, its translations and its
356
+ admissions are created, replaced and deleted with the generation: a target without
357
+ the entry, a failed batch, or release removes them, and a caller never needs a
358
+ separate teardown. Owner verification reads every set back with its description,
359
+ every element with its interval bounds and map data, and every lookup with its
360
+ registers, so apply, inspection, adoption and recovery reject a changed element
361
+ exactly as they reject a changed rule.
362
+
363
+ Without `symmetric`, a publication carries inbound flows only: a flow the workload
364
+ opens from its target address and port toward the uplink is denied. With `symmetric:
365
+ true`, a flow from exactly `targetAddress` and the published target port(s) on the
366
+ exact handoff may leave through the uplink, NEW (a TCP opening only with SYN) or
367
+ ESTABLISHED, with its ESTABLISHED replies, and outer source NAT translates it to
368
+ `hostIp` and the published uplink port. These flows are elements of one more map,
369
+ `published_symmetric`, keyed on the handoff, protocols and the live and original
370
+ source and holding the uplink address and port span; four admissions look it up as a
371
+ set and one postrouting rule translates through it. The workload's outbound flows therefore leave
372
+ from the same address and port its clients reach, as Docker host-mode NAT kept them;
373
+ a range keeps each port unless the host itself already holds that exact tuple on
374
+ `hostIp`. The admissions follow the capture barrier, so every protected current or
375
+ original destination, the platform endpoints included, stays denied. A symmetric
376
+ target port may not fall inside a leased range of `targetAddress`, and may not overlap
377
+ the target ports of another publication of the same address and protocol. Absent and
378
+ `false` are the same canonical policy, digest and compiled graph.
379
+
380
+ Input bounds are 1024 entries. The rule graph does not grow with them: with every
381
+ kind of publication present the host compiles at most eleven rules and four sets,
382
+ about 13,400 bytes of the 100,000-byte rule reserve, and each publication costs
383
+ about 150 to 300 bytes of set elements, which share the target budget described
384
+ under host grants. 64 symmetric ranges take about 19,600 element bytes.
346
385
  The caller still owns the route to `targetAddress` through that handoff, the
347
386
  workload listener, and every other packet owner on the host. As with leased egress,
348
387
  an ACCEPT here cannot override Docker's independent FORWARD DROP; the
@@ -366,7 +405,8 @@ There is no broad ESTABLISHED/RELATED bypass.
366
405
  `routerEgress.publishedPorts` is optional and is the second hop of a published
367
406
  host port. Absent and empty are the same canonical policy, so an unpublished scope
368
407
  keeps its exact previous digest and compiled bytes. Each entry is `{ protocol,
369
- transitPort, endpointPort, endpointAddress, transitSourceAddress }`.
408
+ transitPort, endpointPort, endpointAddress, transitSourceAddress }`, with an
409
+ optional `transitPortEnd` and an optional `symmetric`.
370
410
  `transitSourceAddress` must be a current leased generation address that is present
371
411
  on `handoff`, which rtnetlink verifies at apply, recovery and inspection; the
372
412
  handoff can carry several current leased addresses, so the arrival address is
@@ -379,11 +419,24 @@ published exactly once across the scope, and a published port may not fall insid
379
419
  any leased outbound source-port range of the same protocol on that transit address,
380
420
  because source NAT translates to that same address.
381
421
 
382
- A publication compiles inside the generation: one destination NAT rule in this
383
- table's own prerouting `nat` chain at priority -100, two raw classifications, and
384
- three forward admissions placed ahead of the capture barrier — original direction
385
- NEW and ESTABLISHED, with opening TCP restricted to SYN with FIN/RST/ACK clear,
386
- and the ESTABLISHED reply. The handoff ingress classification precedes the
422
+ `transitPortEnd` publishes the range `transitPort..transitPortEnd`, the second hop of
423
+ a host range. It must be greater than `transitPort`, and `endpointPort` must equal
424
+ `transitPort`, so every port keeps its number on both hops; one port has exactly one
425
+ canonical form, without `transitPortEnd`. No two publications of one protocol overlap
426
+ anywhere in the scope, and no published port falls inside a leased range of its
427
+ transit address. The destination NAT of a range translates the address only.
428
+
429
+ Publications are set elements here too: `published_arrival` (protocol, transit
430
+ address and port), `published_endpoint` (workload link, protocol, address and port,
431
+ the union of every publication's workload ports as disjoint canonical intervals),
432
+ the `published_port` and `published_address` translation maps and
433
+ `published_forward`, with the port fields as inclusive intervals. A constant number
434
+ of rules looks them up whatever the number of publications: one destination NAT
435
+ rule per map in this table's own prerouting `nat` chain at priority -100, two raw
436
+ classifications, and four forward admissions placed ahead of the capture barrier —
437
+ the ESTABLISHED original direction, a NEW opening per protocol with opening TCP
438
+ restricted to SYN with FIN/RST/ACK clear, and the ESTABLISHED reply. The handoff
439
+ ingress classification precedes the
387
440
  barrier's protected-source denial, because the uplink or management network that
388
441
  reaches a published port is itself protected space. The endpoint ingress
389
442
  classification matches the exact workload link, address and published endpoint
@@ -393,24 +446,47 @@ would otherwise carry the workload's published packets into that generation, whe
393
446
  the publication's flow does not exist. Both classifications match the arriving
394
447
  tuple only, set no generation zone and leave the flow in the default conntrack
395
448
  zone. The private anti-spoof guards still run first, and every admission checks
396
- both links, the conntrack protocol, the default zone, the direction, the
397
- translated current address/port and the original conntrack destination.
449
+ the handoff, the default zone and the direction, and looks up the exact workload
450
+ link (index and name), both protocols, the translated current address/port and the
451
+ original conntrack destination.
398
452
 
399
- A published `endpointPort` is therefore dedicated to its publication: a new
453
+ A published endpoint port is therefore dedicated to its publication: a new
400
454
  outbound flow that the workload opens from that exact address and port on that
401
- exact link stays in the default zone too and is not translated by the leased
402
- outbound source NAT. Published traffic is not source-translated on either hop, so
403
- the workload observes the actual client address. The chain, its translation, its
404
- classifications and its admissions are created, replaced and deleted with the
405
- generation: a target without the entry, a failed batch, or release removes them.
406
-
407
- Input bounds are 64 entries, but the atomic graph reserve binds first. A router
408
- publication costs six operations and about 7,600 bytes, so a combined
409
- two-endpoint router policy leaves room for roughly five, and `prepare()` rejects an
410
- excessive set before any kernel work. The caller still owns both hops: the exact
411
- matching `hostTransit` publication, the route to the workload through this router,
412
- the workload listener and every other packet owner. A publication grants no egress;
413
- outbound flows still need their leased grants.
455
+ exact link stays in the default zone too and is never translated by the leased
456
+ outbound source NAT. Without `symmetric` such a flow is denied. Inbound published
457
+ traffic is not source-translated on either hop, so the workload observes the actual
458
+ client address. The chain, its sets, its translations, its classifications and its
459
+ admissions are created, replaced and deleted with the generation: a target without
460
+ the entry, a failed batch, or release removes them.
461
+
462
+ With `symmetric: true` the workload may open flows from exactly `endpointAddress`
463
+ and the published endpoint port(s) on its exact link toward the handoff, NEW (a TCP
464
+ opening only with SYN) or ESTABLISHED, with their ESTABLISHED replies, and source
465
+ NAT translates them to `transitSourceAddress` and the published transit port. With
466
+ the matching symmetric `hostTransit` publication the flow reaches the uplink from
467
+ `hostIp` and the published uplink port, the same address and port its clients
468
+ reach, as Docker host-mode NAT kept it; a range keeps each port unless that exact
469
+ tuple is already taken on the translated address. These flows are elements of the
470
+ `published_symmetric` map, keyed on the workload link, protocols and the live and
471
+ original source and holding the transit address and port span, behind four
472
+ admissions and one postrouting translation. Both directions are already in the
473
+ default zone through the publication's endpoint and handoff classifications, so no
474
+ leased classifier can capture them. The admissions follow the capture barrier, so
475
+ every protected current or original destination, the platform endpoints included,
476
+ stays denied; outbound flows anywhere else still need their leased grants. The
477
+ endpoint ports of a symmetric publication may not overlap the endpoint ports of
478
+ another publication of the same address and protocol, since each flow has exactly
479
+ one published source. Absent and `false` are the same canonical policy, digest and
480
+ compiled graph.
481
+
482
+ Input bounds are 1024 entries. The rule graph does not grow with them: with every
483
+ kind of publication present the router compiles at most thirteen rules and six sets,
484
+ about 15,100 bytes of the 100,000-byte rule reserve, and each publication costs
485
+ about 250 to 450 bytes of set elements, which share the target budget with the
486
+ router's endpoints, rules and grants. 64 symmetric ranges take about 28,000 element
487
+ bytes. The caller still owns both hops: the exact matching `hostTransit`
488
+ publication, the route to the workload through this router, the workload listener
489
+ and every other packet owner.
414
490
 
415
491
  #### Host grants
416
492
 
@@ -466,7 +542,7 @@ foreign set exactly as they reject a changed rule; a bound CONSTANT set also
466
542
  refuses element changes from any other socket.
467
543
 
468
544
  Absent and empty are the same canonical policy, so a scope without grants keeps its
469
- exact previous digest and compiled bytes and has no set. Withdrawing a grant is an
545
+ exact previous digest and compiled bytes and has no host grant set. Withdrawing a grant is an
470
546
  ordinary atomic replacement: the batch deletes the previous rules, chains and sets
471
547
  by handle and creates the complete target, so a target without the entry, a failed
472
548
  batch or release removes it, and an established flow stops with it. Apply grants
@@ -474,14 +550,134 @@ hop by hop from the workload side (router, host transit, guard) and withdraw the
474
550
  in the reverse order; each table is atomic on its own, and a partially applied
475
551
  grant only stays dark.
476
552
 
553
+ #### Host-local platform endpoints
554
+
555
+ A platform endpoint normally lives elsewhere and host transit forwards leased
556
+ flows to it. When the host serves an endpoint itself — a relay listener, a data
557
+ port or a VPN hub on an address the host holds — the leased flow ends in the
558
+ host's INPUT instead, where every handoff arrival is denied. The optional
559
+ `hostTransit.localPlatformEndpoints` lists the ids of `protection.platformEndpoints`
560
+ the host serves:
561
+
562
+ ```typescript
563
+ protection: { authorityDigest, prefixes: ['10.0.0.0/8', '192.0.2.0/24'], platformEndpoints: [
564
+ { id: 'relay', address: '192.0.2.2', protocol: 'tcp', port: 8443 }] },
565
+ localPlatformEndpoints: ['relay'],
566
+ ```
567
+
568
+ Each id names an existing endpoint whose address is an exact `requiredIpv4Addresses`
569
+ entry of the uplink or of a handoff link, so rtnetlink verifies that the host holds it
570
+ at apply, recovery and inspection. The refusal of a platform endpoint on a host
571
+ address is lifted only for declared ids; unknown, duplicate or unheld ids reject.
572
+ The router scope grants its workloads the endpoint like any platform endpoint
573
+ (`destination: { kind: 'platform', endpointId }`).
574
+
575
+ INPUT admits the original direction only from the exact handoff that carries a lease,
576
+ with the lease's transit source address and a source port in one of its ranges for
577
+ the endpoint's protocol, in the default conntrack zone, to the exact endpoint address
578
+ and port in both the live and the originally tracked tuple, NEW (a TCP opening only
579
+ with SYN) or ESTABLISHED. OUTPUT admits only the ESTABLISHED reply back through that
580
+ handoff. A flow translated in transit, another source address or port, another
581
+ endpoint port or protocol and any host-origin opening toward the transit address keep
582
+ meeting the handoff denials. The source port is checked against the leased range in
583
+ both tuples, not for equality between them. Each leased range adds one INPUT and one
584
+ OUTPUT jump when an endpoint of its protocol is declared, and each endpoint adds three
585
+ rules. Absent and empty are the same canonical policy, digest and compiled bytes. The
586
+ caller keeps a declared endpoint's address outside every guarded allocation pool and
587
+ owns the listener.
588
+
589
+ #### Workload grants
590
+
591
+ `routerEgress.workloadGrants` is optional: a stateful one-way flow between two
592
+ workloads behind the same router, such as an ingress workload reaching a target
593
+ workload's service port across organizations. Each entry is `{ protocol,
594
+ sourceAddress, destinationAddress, destinationPort }`: `tcp` or `udp`, the exact
595
+ current addresses of two different veth workload endpoints of the scope and one port
596
+ 1–65535. Prefixes, wildcards, ranges, router-local and platform-endpoint addresses,
597
+ TUN endpoints, both ends on one endpoint and duplicates reject. The source opens and
598
+ the destination only answers; the reverse direction is a second grant.
599
+
600
+ Four FORWARD admissions follow the private anti-spoof guards whatever the number of
601
+ grants: ESTABLISHED original direction, a NEW UDP opening, a NEW TCP opening
602
+ restricted to SYN with FIN/RST/ACK clear, and the ESTABLISHED reply, each in the
603
+ default conntrack zone. Every admission looks up the incoming link with the source
604
+ address and the outgoing link with the destination address (index and name) in the
605
+ CONSTANT set `workload_link`, and the protocols with the live and originally tracked
606
+ tuple in the CONSTANT set `workload_grant`, so a spoofed, renamed or translated flow
607
+ never matches and the destination can never open toward the source. Endpoint prefixes
608
+ are protected, so both directions stay in the default zone ahead of every leased
609
+ classifier. Owner verification reads both sets back element by element. Up to 1024
610
+ grants fit; the sets share the scope's target budget with `hostGrants` and every
611
+ other set, and 1024 of each fit together beside the reference router; a scope
612
+ beyond the budget is refused before any kernel work. Absent and empty are the same canonical policy, digest and compiled bytes, and
613
+ withdrawal is an ordinary atomic replacement that also stops established flows.
614
+
615
+ These flows need no private `rules`. A private rule is stateless: answering through
616
+ private rules needs a reverse rule, which would also let the destination open toward
617
+ the source.
618
+
619
+ #### Loopback TCP port owners
620
+
621
+ `allocationPoolGuard.localTcpPortOwners` is optional and restricts a loopback TCP
622
+ service to one local user, such as a container runtime's CRI stream server on
623
+ `127.0.0.1:10010` that only root may reach. Each entry is `{ address, port, uid }`:
624
+ one exact IPv4 loopback host address in `127.0.0.0/8` (not the network or broadcast
625
+ address), one port 1–65535 and one uid 0–4294967294. Prefixes, wildcards, IPv6,
626
+ ranges, uid lists and a second entry for the same address and port reject; the
627
+ limit is 8 entries. Absent and empty are the same canonical policy, so a guard
628
+ without owners keeps its exact previous digest and compiled bytes.
629
+
630
+ ```typescript
631
+ const guard: IManagedNftPolicyV2 = { schemaVersion: 2, revision: 1, scope: {
632
+ kind: 'allocationPoolGuard', authorityDigest, prefixes: ['10.240.0.0/16'],
633
+ localTcpPortOwners: [{ address: '127.0.0.1', port: 10010, uid: 0 }] } };
634
+ ```
635
+
636
+ Each entry compiles two rules at the head of the guard's base chains, ahead of the
637
+ host grant admissions and the pool guard, so nothing earlier in the table accepts
638
+ past them:
639
+
640
+ - OUTPUT (filter priority 0, after ordinary output destination NAT): IPv4 TCP to the
641
+ exact address and port with `meta skuid != uid` is rejected with a TCP reset, so
642
+ another user's `connect()` fails at once with `ECONNREFUSED`. This covers every
643
+ packet, not only the opening, and a flow a foreign output DNAT redirects to the
644
+ port. `meta skuid` is the socket file's file-system uid as the user namespace owning the
645
+ network namespace sees it, the initial one on a host: a process in a user namespace
646
+ that shares the host network namespace is matched by its host uid, and a user
647
+ namespace's root is not uid 0. A packet without a socket
648
+ file carries no uid and does not match: a kernel reply, the teardown of an orphaned
649
+ socket, or an in-kernel socket such as an NFS or CIFS client, which only a privileged
650
+ mount creates.
651
+ - INPUT: the exact address and port arriving on any interface but loopback (ifindex
652
+ 1 in every network namespace) are dropped, so `route_localnet` or a prerouting DNAT
653
+ cannot expose the service to another host.
654
+
655
+ The rules match the exact destination address. The service must bind exactly that
656
+ address; a wildcard listener stays reachable through the host's other addresses.
657
+ IPv6 is not covered, so the service must not listen on `::1` or `::`. A socket keeps
658
+ the uid it was created with, so a root process that hands its connected socket to
659
+ another user delegates that connection. The reject expression needs the kernel's
660
+ `nft_reject_inet` module (autoloaded on stock kernels). Owner verification reads
661
+ back every rule, including the comparison operator, the uid and the reject type, so
662
+ apply, inspection, adoption and recovery reject a changed rule. The entries are
663
+ applied, retained, recovered and released with the rest of the guard's table; after
664
+ a reboot the caller applies its retained intent again, like every other member.
665
+
477
666
  Input bounds are 1024 grants per scope, the maximum a network projection carries.
478
- Elements have their own 100,000-byte budget beside the 100,000-byte rule graph, and
479
- are sent 256 to a message. At 1024 grants the reference router needs about 94 KB
480
- of elements (two sets), host transit about 66 KB and the guard about 45 KB. A
481
- replacement batch — handle-only deletion of a previous maximum graph, a maximum
482
- rule graph and maximum elements — stays below 320,000 bytes, which the kernel
483
- accepts with any `net.core.wmem_max` of at least 160,016 (the Linux default is
484
- 212,992). The private facade IPC carries up to three complete policies in one
667
+ Elements are sent 256 to a message. At 1024 grants the reference router needs about
668
+ 98 KB of elements (two sets), host transit about 66 KB and the guard about 45 KB.
669
+
670
+ Every replacement is one nfnetlink batch in one `sendmsg`, and Linux refuses a
671
+ netlink message larger than the socket send buffer: the requested 1 MiB `SO_SNDBUF`
672
+ is clamped to `net.core.wmem_max` and doubled, 425,984 bytes with the Linux default
673
+ of 212,992. That is the hard ceiling. This package bounds a batch at 320,000 bytes,
674
+ which fits every `wmem_max` of at least 160,016, and splits it: at most 64 bytes of
675
+ batch framing, at most 107,520 bytes for the handle-only deletion of the previous
676
+ graph (fewer than 768 rules, chains and sets of at most 140 bytes each), and
677
+ 212,000 bytes for the complete target, of which rules, chains and sets may take at
678
+ most 100,000 and set elements the rest. Scopes that grow with their inputs keep
679
+ those inputs in set elements, so the element share grows where the rule share
680
+ shrinks; `prepare()` refuses a target beyond either bound before any kernel work. The private facade IPC carries up to three complete policies in one
485
681
  status and is bounded at 1,048,576 bytes per line; the facade captures at most
486
682
  1,000,000 JSON bytes and 65,536 values per request or result. The caller owns the host route
487
683
  to the workload through the handoff, selecting the transit host address as the
@@ -489,18 +685,56 @@ socket's source, the workload listener and every other packet owner.
489
685
 
490
686
  Active overlapping classifiers, duplicate zones/labels and conflicting handoff
491
687
  allocations reject. Empty router generations retain private routing while denying
492
- egress. Input bounds include 32 private endpoints, 128 private rules, 32 active
493
- generations, 128 total egress grants, 128 protected prefixes, 96 platform endpoints,
494
- 16 ranges per allocation, and 32 host handoffs/active allocations. The complete
495
- compiled graph still must fit 768 operations and 100,000 bytes; cross-products can
496
- reach that limit before individual input limits. Exhausted source-port ranges
497
- drop new flows rather than allocate outside the lease.
498
-
499
- The router compiler shares original-direction connection-state and label checks
500
- per generation and protocol. Each jump still checks the complete current/original
501
- tuple, links, zone and direction; replies and source NAT keep their exact checks.
502
- A combined two-endpoint policy with six private rules, five protected prefixes and
503
- six workload/router grants fits the existing atomic graph limit.
688
+ egress. Router input bounds are 128 private endpoints and links, 1024 private
689
+ rules, 32 active generations and 1024 total egress grants; other bounds include 128
690
+ protected prefixes, 96 platform endpoints, 16 ranges per allocation, and 32 host
691
+ handoffs/active allocations. The complete target must fit the budget above; the
692
+ router's rules grow only with its protected prefixes, so its elements bind first.
693
+ Exhausted source-port ranges drop new flows rather than allocate outside the lease.
694
+
695
+ The router compiles its private endpoints, directed rules and leased grants into
696
+ named sets and maps behind a fixed set of rules, whatever the number of workloads:
697
+ - `private_link` holds every endpoint link (index and name) and the absent
698
+ interface of router-local traffic; `private_index` and `private_name` hold every
699
+ endpoint index and name for the terminal denials, either alone enough to deny;
700
+ `private_source` holds each link index with its source prefixes. The anti-spoof
701
+ guard drops a packet on an endpoint link whose address is not one of its
702
+ prefixes, and anything but IPv4 there.
703
+ - `private_any` and `private_port` hold the directed rules: incoming and outgoing
704
+ link index (0 for the router itself), addresses, and for a transport rule the
705
+ protocol and ports.
706
+ - Per kind (platform ahead of the protected barrier, public behind it) and per
707
+ shape (whether the grant names a source port), `leased_*_zone` maps a workload
708
+ or router flow's link, protocol and tuple to its generation zone before conntrack,
709
+ `leased_*_return` maps a translated return to its zone, and `leased_*_flow` maps
710
+ the link, both protocols, the live and originally tracked tuple and the zone to
711
+ the grant's source NAT; the filter admissions look it up as a set.
712
+ `leased_label` holds each generation's zone and label, and `leased_generation`
713
+ maps a zone to the label a fresh flow receives.
714
+ A key's link index names its link because the rule also finds the interface's
715
+ index and name in `private_link`, whose indices are unique. Every element is
716
+ checked, like every rule, on apply, inspection, adoption and recovery. A decision
717
+ corpus of about 133,000 packets across five router fixtures, frozen from the
718
+ per-rule compiler of 2.6.0, is reproduced exactly by the set-backed compiler.
719
+
720
+ A router of 64 workloads, each with local DNS, public TCP and UDP egress, grants to
721
+ four platform endpoints and one to three publications, compiles about 66,000 rule
722
+ bytes and 105,000 element bytes; 89 such workloads fit the target budget:
723
+
724
+ | Workloads | Router rules | Router elements | Host rules | Host elements |
725
+ | --- | --- | --- | --- | --- |
726
+ | 1 | 65,820 | 4,968 | 42,944 | 1,128 |
727
+ | 8 | 65,820 | 16,056 | 42,944 | 2,528 |
728
+ | 32 | 65,820 | 54,072 | 42,944 | 7,328 |
729
+ | 64 | 65,820 | 104,760 | 42,944 | 13,728 |
730
+
731
+ **Upgrading from 2.x:** version 3 changes the compiled identity of every schema-v2
732
+ `routerEgress` table and of every scope with `publishedPorts`, while their policy
733
+ digests stay unchanged. Reconcile or release against such a table applied by 2.x
734
+ returns `Conflict`: the new engine cannot adopt the old graph. Recover and release
735
+ each such table with the exact previous engine and its retained intent and
736
+ receipt, under the same traffic fence as below, before upgrading; tables of other
737
+ scopes and without publications are unchanged.
504
738
 
505
739
  **Upgrading from 1.x:** version 2 changes the compiled identity of schema-v2
506
740
  `routerEgress` tables. Retained router tables and their receipts cannot be adopted
@@ -619,6 +853,27 @@ A client whose source port equals a leased grant's destination port on that gran
619
853
  public address completes the published TCP handshake and receives its UDP reply
620
854
  from the published address, so the leased classifiers of the same generation cannot
621
855
  capture the workload's published packets.
856
+ Published ranges and symmetric publications are qualified across both hops in the
857
+ same topology, with every publication a set element whose kernel dump matches the
858
+ compiled sets and maps: the first, middle and last port of a UDP range reach their listeners
859
+ with the client address preserved and answer from the published address, the ports
860
+ just outside the range stay dark against pre-policy positive controls, symmetric UDP
861
+ from a range port and from the signalling port and symmetric TCP from the signalling
862
+ port reach the uplink peer from the published address with the same port, a later
863
+ request from that peer reaches the workload on the same flow, the plain publication's
864
+ workload port opens nothing, symmetric flows into the protected union stay denied
865
+ against a pre-policy positive control, leased egress keeps its leased range, and the
866
+ range goes dark with release.
867
+ A 64-workload router is qualified in one batch, every workload in its own namespace
868
+ with local DNS, public TCP and UDP egress, four platform endpoints and a
869
+ publication, the first with the SIP shape: the kernel dumps every private, leased
870
+ and published element and the owner verifies it, and on the first, middle and last
871
+ workload local DNS answers, public UDP egress leaves from the leased range, UDP and
872
+ TCP platform endpoints answer through both hops, the publication reaches its
873
+ listener with the client address preserved, and another workload, a protected
874
+ address outside the platform endpoints and an ungranted platform port stay denied
875
+ against pre-policy positive controls, as does another workload's source address.
876
+ The router compiles 82 rules for the 64 workloads.
622
877
  Host grants are qualified across all three tables in the same four-namespace
623
878
  topology with leased public egress on every port of both protocols and 1024 grants
624
879
  in every scope: a member far from the probed tuples passes, the tuple just outside
@@ -632,6 +887,32 @@ withdrawal in every scope stops new flows and the established one. A full guard
632
887
  set persists across owner loss, is re-verified element by element before adoption,
633
888
  survives lost-ACK replay and is replaced and released like the rest of the graph,
634
889
  and another socket cannot add or delete one of its elements.
890
+ Host-local platform endpoints are qualified from a router namespace against
891
+ pre-policy positive controls: without the member the handoff denials keep both a TCP
892
+ endpoint on the uplink address and a UDP endpoint on the handoff address dark; with
893
+ it the exact leased flows reach them with the transit source preserved, while a
894
+ source port outside the lease or in the other protocol's range, an unleased source
895
+ address, another endpoint port, a flow a foreign prerouting DNAT translated onto the
896
+ endpoint and the host's own opening toward the transit address stay dark; withdrawal
897
+ stops new flows and the established one, and release reopens the path. A graph with
898
+ served endpoints persists across owner loss and survives lost-ACK replay.
899
+ Workload grants are qualified between three workload namespaces behind the router,
900
+ each with leased public egress on every port of both protocols: against pre-policy
901
+ positive controls, the grant opens exactly the source workload's TCP and UDP flows to
902
+ the granted ports with the source preserved; another port, another source workload,
903
+ another destination workload, the destination's TCP and UDP openings toward the
904
+ source stay dark; withdrawal stops new flows and the established one; release reopens
905
+ the path. A set of 1024 grants persists across owner loss, is re-verified element by
906
+ element and survives lost-ACK replay.
907
+ Loopback TCP port owners are qualified against pre-policy positive controls on the
908
+ same listeners: the owning uid (root, and uid 1000 for a second port) connects, every
909
+ other uid is reset with `ECONNREFUSED`, an unowned port and another loopback address
910
+ stay open to every user, a foreign output DNAT to the owned port stays closed to
911
+ another user, and an uplink arrival translated to loopback with `route_localnet`
912
+ enabled is dropped. The policy keeps enforcing after the owner detaches and its
913
+ process is gone, a fresh owner verifies and adopts the exact graph, and release
914
+ reopens every probe. The maximum of eight owners beside the guard persists across
915
+ owner loss and survives lost-ACK replay.
635
916
  The arm64 binary is cross-built; native packet qualification is currently x86_64.
636
917
  Kernel 6.8 is unsupported. No production activation is implied by these tests.
637
918