@push.rocks/smartnftables 2.6.0 → 3.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/changelog.md +19 -0
- package/dist_rust/smartnftables_linux_amd64_musl +0 -0
- package/dist_rust/smartnftables_linux_amd64_musl.tsrust-build.json +5 -5
- package/dist_rust/smartnftables_linux_arm64_musl +0 -0
- package/dist_rust/smartnftables_linux_arm64_musl.tsrust-build.json +5 -5
- package/dist_ts/00_commitinfo_data.js +1 -1
- package/dist_ts/managed.egress.types.d.ts +25 -3
- package/package.json +6 -6
- package/readme.md +205 -63
- package/rust/src/egress.compile.rs +155 -76
- package/rust/src/egress.host.rs +24 -51
- package/rust/src/egress.hostgrant.rs +1 -1
- package/rust/src/egress.keys.rs +305 -0
- package/rust/src/egress.leased.rs +399 -0
- package/rust/src/egress.private.rs +197 -0
- package/rust/src/egress.published.rs +370 -0
- package/rust/src/egress.router.rs +61 -294
- package/rust/src/egress.rs +126 -11
- package/rust/src/egress_hostgrant_tests.rs +149 -18
- package/rust/src/egress_localport_tests.rs +3 -1
- package/rust/src/egress_publishedrange_tests.rs +812 -0
- package/rust/src/egress_routerequivalence_tests.rs +448 -0
- package/rust/src/egress_tests.rs +132 -230
- package/rust/src/egress_workloadgrant_tests.rs +6 -7
- package/rust/src/owner.rs +69 -14
- package/rust/src/owner_host_traffic_tests.rs +4 -0
- package/rust/src/owner_identity_tests.rs +166 -80
- package/rust/src/owner_publishedrange_traffic_tests.rs +306 -0
- package/rust/src/owner_router_traffic_tests.rs +2 -0
- package/rust/src/owner_scale_traffic_tests.rs +383 -0
- package/rust/src/owner_tests.rs +6 -0
- package/rust/src/policy.rs +29 -25
- package/ts/00_commitinfo_data.ts +1 -1
- package/ts/managed.egress.types.ts +25 -3
package/changelog.md
CHANGED
|
@@ -1,5 +1,24 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 2026-09-28 - 3.0.0
|
|
4
|
+
|
|
5
|
+
### Breaking Changes
|
|
6
|
+
|
|
7
|
+
- The schema-v2 `routerEgress` compiler now keeps private endpoints, directed rules and leased egress grants in named sets and maps behind a fixed set of rules, and every scope keeps its `publishedPorts` the same way (see Features). The policy API is additive and policy digests are unchanged, but the compiled graph of every `routerEgress` table and of every scope with publications changes: reconcile, inspection or release against such a table applied by 2.x returns `Conflict`, because the new engine does not adopt the old graph. Recover and release each such table with the exact previous engine and its retained intent and receipt, under a traffic fence, before upgrading. Tables of other scopes and without publications are unchanged.
|
|
8
|
+
- Rebalance the atomic replacement budget inside the unchanged 320,000-byte batch: the complete target (rules, chains and sets at most 100,000 bytes, plus set elements) may take 212,000 bytes beside a 107,520-byte handle-only deletion of the previous graph, instead of two fixed 100,000-byte halves for rules and elements. The hard ceiling is Linux's `sendmsg` limit, the socket send buffer (`net.core.wmem_max` doubled, 425,984 bytes by default); 320,000 bytes fits any `wmem_max` of at least 160,016.
|
|
9
|
+
- Raise the router's input bounds to 128 private endpoints and links, 1024 private rules and 1024 egress grants, and the publication bound of both scopes to 1024. Schema-v1 policies keep their bounds, digests and compiled bytes.
|
|
10
|
+
|
|
11
|
+
### Features
|
|
12
|
+
|
|
13
|
+
- Router: a node of 64 workloads, each with local DNS, public TCP and UDP egress, four platform endpoints and one to three publications, compiles about 66,000 rule bytes and 105,000 element bytes (89 such workloads fit), where 2.x held three. Private links, indices, names and source prefixes, directed rules, and per-kind and per-shape leased zone, return and flow maps, with the generation labels in two exact sets, replace the per-endpoint guard chains, per-rule grants, per-grant classifiers and admissions and the per-generation state chains; every lookup keeps the operands the rules compared, a key's link index is proven with the interface name through `private_link`, and a decision corpus of about 133,000 packets frozen from the 2.6.0 compiler is reproduced exactly. Owner verification admits absence (inverted) lookups and exact maps, and interval elements without a distinct last key.
|
|
14
|
+
- Add published port ranges to both hops of a schema-v2 publication: an optional `hostPortEnd` on `hostTransit.publishedPorts` and `transitPortEnd` on `routerEgress.publishedPorts` publish `port..end` with every port kept (the inside port must equal the first port). A range compiles to one interval match per classification, translation and admission and to address-only destination NAT, so it costs the same operations as one port. Inverted and one-port ranges, a range with a differing inside port, overlaps between publications of one protocol and overlaps with a leased source-port range on the translated address are refused.
|
|
15
|
+
- Add `symmetric` publications on both hops: the workload may open flows from its published address and port(s) toward unprotected space, NEW (a SYN-only TCP opening) or ESTABLISHED with ESTABLISHED replies, source-translated to the transit address and published transit port in the router and to `hostIp` and the published uplink port on the host, so outbound flows leave from the port clients reach. The admissions follow the capture barrier, so the protected union stays denied; a symmetric publication's inside ports may not overlap another publication's on the same address or, on the host, a leased range of its target address.
|
|
16
|
+
- Compile every publication, of both hops, as elements of CONSTANT concatenated interval sets and maps (the kernel's pipapo backend) behind a constant number of rules: at most eleven rules and four sets on the host and thirteen rules and six sets in the router, whatever the number of publications, instead of four and six rules per publication. Destination NAT is a map lookup (to address and port for one port, to the address alone for a range) and symmetric source NAT a map lookup to the outside address and port span. Owner verification now admits interval and map sets, reads back their concatenation, every element's bounds and data, and map lookups with their data register, so apply, inspection, adoption and recovery reject a changed element like a changed rule. Absent and empty publications, `hostPortEnd`, `transitPortEnd` and `symmetric` (and `symmetric: false`) keep the previous canonical policy and digest, and a scope without publications keeps its compiled bytes; a scope with publications compiles a different graph for the same digest.
|
|
17
|
+
|
|
18
|
+
### Maintenance
|
|
19
|
+
|
|
20
|
+
- Release tooling: pnpm 12.5.1, `@git.zone/cli` 7.3.0, `@git.zone/tsrust` 1.15.1, `@git.zone/tstest` 6.3.2, `@git.zone/tsrun` 3.0.1, `@types/node` 26.6.2 (pnpm 12.6.0 and `@types/node` 26.6.3 stay out until they clear the seven-day rule).
|
|
21
|
+
|
|
3
22
|
## 2026-09-25 - 2.6.0
|
|
4
23
|
|
|
5
24
|
### Features
|
|
Binary file
|
|
@@ -1,13 +1,13 @@
|
|
|
1
1
|
{
|
|
2
2
|
"format": "tsrust.build-provenance.v2",
|
|
3
|
-
"binarySha256": "
|
|
3
|
+
"binarySha256": "17c4308039280f584cd51d6cf9030e3c4512746d316561925b6da95bb3c67ff8",
|
|
4
4
|
"buildInfo": {
|
|
5
5
|
"projectName": "@push.rocks/smartnftables",
|
|
6
|
-
"projectVersion": "
|
|
7
|
-
"gitCommit": "
|
|
6
|
+
"projectVersion": "3.0.0",
|
|
7
|
+
"gitCommit": "61c5a8be9c99a83cecde6509fb31c47f43e15bff",
|
|
8
8
|
"gitDirty": false,
|
|
9
|
-
"builtAt": "2026-09-
|
|
10
|
-
"tsrustVersion": "1.15.
|
|
9
|
+
"builtAt": "2026-09-28T00:27:05.031Z",
|
|
10
|
+
"tsrustVersion": "1.15.1",
|
|
11
11
|
"binary": "smartnftables",
|
|
12
12
|
"target": "linux_amd64_musl"
|
|
13
13
|
}
|
|
Binary file
|
|
@@ -1,13 +1,13 @@
|
|
|
1
1
|
{
|
|
2
2
|
"format": "tsrust.build-provenance.v2",
|
|
3
|
-
"binarySha256": "
|
|
3
|
+
"binarySha256": "3756fe460edc3bdea00c7f598a683081d8238f5172dad9776304a030508d4276",
|
|
4
4
|
"buildInfo": {
|
|
5
5
|
"projectName": "@push.rocks/smartnftables",
|
|
6
|
-
"projectVersion": "
|
|
7
|
-
"gitCommit": "
|
|
6
|
+
"projectVersion": "3.0.0",
|
|
7
|
+
"gitCommit": "61c5a8be9c99a83cecde6509fb31c47f43e15bff",
|
|
8
8
|
"gitDirty": false,
|
|
9
|
-
"builtAt": "2026-09-
|
|
10
|
-
"tsrustVersion": "1.15.
|
|
9
|
+
"builtAt": "2026-09-28T00:27:15.164Z",
|
|
10
|
+
"tsrustVersion": "1.15.1",
|
|
11
11
|
"binary": "smartnftables",
|
|
12
12
|
"target": "linux_arm64_musl"
|
|
13
13
|
}
|
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
*/
|
|
4
4
|
export const commitinfo = {
|
|
5
5
|
name: '@push.rocks/smartnftables',
|
|
6
|
-
version: '
|
|
6
|
+
version: '3.0.0',
|
|
7
7
|
description: 'A TypeScript module for managing nftables rules including NAT, firewall, and rate limiting with a high-level API.'
|
|
8
8
|
};
|
|
9
9
|
//# sourceMappingURL=data:application/json;base64,eyJ2ZXJzaW9uIjozLCJmaWxlIjoiMDBfY29tbWl0aW5mb19kYXRhLmpzIiwic291cmNlUm9vdCI6IiIsInNvdXJjZXMiOlsiLi4vdHMvMDBfY29tbWl0aW5mb19kYXRhLnRzIl0sIm5hbWVzIjpbXSwibWFwcGluZ3MiOiJBQUFBOztHQUVHO0FBQ0gsTUFBTSxDQUFDLE1BQU0sVUFBVSxHQUFHO0lBQ3hCLElBQUksRUFBRSwyQkFBMkI7SUFDakMsT0FBTyxFQUFFLE9BQU87SUFDaEIsV0FBVyxFQUFFLG1IQUFtSDtDQUNqSSxDQUFBIn0=
|
|
@@ -65,9 +65,14 @@ export interface IManagedNftEgressGenerationV2 extends IManagedNftHandoffAllocat
|
|
|
65
65
|
* `transitSourceAddress:transitPort`, this translates on to the workload. */
|
|
66
66
|
export interface IManagedNftRouterPublishedPortV2 {
|
|
67
67
|
protocol: 'tcp' | 'udp';
|
|
68
|
-
/** Published handoff port, 1–65535
|
|
69
|
-
* It is the `targetPort` of the matching host-transit publication. */
|
|
68
|
+
/** Published handoff port, 1–65535; published ports never overlap per protocol across the
|
|
69
|
+
* scope. It is the `targetPort` of the matching host-transit publication. */
|
|
70
70
|
transitPort: number;
|
|
71
|
+
/** Optional last port of a published range `transitPort..transitPortEnd`, greater than
|
|
72
|
+
* `transitPort`. A range keeps every port, so `endpointPort` must equal `transitPort`; like one
|
|
73
|
+
* port it is one interval set element per set, and its translation keeps the port. Absent for
|
|
74
|
+
* one port, which is its only canonical form. */
|
|
75
|
+
transitPortEnd?: number;
|
|
71
76
|
endpointPort: number;
|
|
72
77
|
/** Exact current workload endpoint address behind one veth endpoint of this
|
|
73
78
|
* scope. Never a router-local address, a gateway or a platform endpoint. */
|
|
@@ -75,6 +80,12 @@ export interface IManagedNftRouterPublishedPortV2 {
|
|
|
75
80
|
/** Exact leased transit source address of one current generation, present on
|
|
76
81
|
* `handoff`. No wildcard, secondary-address or route inference. */
|
|
77
82
|
transitSourceAddress: string;
|
|
83
|
+
/** Optional: the workload may open flows from `endpointAddress` and its published endpoint
|
|
84
|
+
* port(s) toward unprotected space, source-translated to `transitSourceAddress` and the
|
|
85
|
+
* published transit port (a range keeps the port). The protected union stays denied, and its
|
|
86
|
+
* endpoint ports may not overlap another publication's on the same address. Absent and
|
|
87
|
+
* `false` are the same canonical policy. */
|
|
88
|
+
symmetric?: boolean;
|
|
78
89
|
}
|
|
79
90
|
/** One exact host-origin flow: the node's own host network namespace dials one port of one local
|
|
80
91
|
* workload at its lease address, from the transit host address of the current handoff. Compile the
|
|
@@ -137,13 +148,24 @@ export interface IManagedNftRouterEgressScopeV2 {
|
|
|
137
148
|
/** One inbound uplink publication, compiled inside the same host-transit generation. */
|
|
138
149
|
export interface IManagedNftPublishedPortV2 {
|
|
139
150
|
protocol: 'tcp' | 'udp';
|
|
140
|
-
/** Published uplink port, 1–65535
|
|
151
|
+
/** Published uplink port, 1–65535; published ports never overlap per protocol across the scope. */
|
|
141
152
|
hostPort: number;
|
|
153
|
+
/** Optional last port of a published range `hostPort..hostPortEnd`, greater than `hostPort`.
|
|
154
|
+
* A range keeps every port, so `targetPort` must equal `hostPort`; like one port it is one
|
|
155
|
+
* interval set element per set, and its translation keeps the port. Absent for one port, which
|
|
156
|
+
* is its only canonical form. */
|
|
157
|
+
hostPortEnd?: number;
|
|
142
158
|
targetPort: number;
|
|
143
159
|
/** Exact leased transit source address of one current allocation; it selects the handoff. */
|
|
144
160
|
targetAddress: string;
|
|
145
161
|
/** Exact current uplink address that publishes this port. No wildcard or route inference. */
|
|
146
162
|
hostIp: string;
|
|
163
|
+
/** Optional: flows from `targetAddress` and the published target port(s) may leave through the
|
|
164
|
+
* uplink toward unprotected space, source-translated to `hostIp` and the published uplink port
|
|
165
|
+
* (a range keeps the port). The protected union stays denied, its target ports may not fall in
|
|
166
|
+
* a leased range of `targetAddress` or overlap another publication's. Absent and `false` are
|
|
167
|
+
* the same canonical policy. */
|
|
168
|
+
symmetric?: boolean;
|
|
147
169
|
}
|
|
148
170
|
/** Shared-host capture barrier and explicit outer SNAT; Docker forwarding remains a separate owner. */
|
|
149
171
|
export interface IManagedNftHostTransitScopeV2 {
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@push.rocks/smartnftables",
|
|
3
|
-
"version": "
|
|
3
|
+
"version": "3.0.0",
|
|
4
4
|
"private": false,
|
|
5
5
|
"description": "A TypeScript module for managing nftables rules including NAT, firewall, and rate limiting with a high-level API.",
|
|
6
6
|
"main": "dist_ts/index.js",
|
|
@@ -9,12 +9,12 @@
|
|
|
9
9
|
"author": "Task Venture Capital GmbH",
|
|
10
10
|
"license": "MIT",
|
|
11
11
|
"devDependencies": {
|
|
12
|
-
"@git.zone/cli": "7.
|
|
12
|
+
"@git.zone/cli": "7.3.0",
|
|
13
13
|
"@git.zone/tsbuild": "^5.0.0",
|
|
14
|
-
"@git.zone/tsrun": "^3.0.
|
|
15
|
-
"@git.zone/tsrust": "1.15.
|
|
16
|
-
"@git.zone/tstest": "6.2
|
|
17
|
-
"@types/node": "26.6.
|
|
14
|
+
"@git.zone/tsrun": "^3.0.1",
|
|
15
|
+
"@git.zone/tsrust": "1.15.1",
|
|
16
|
+
"@git.zone/tstest": "6.3.2",
|
|
17
|
+
"@types/node": "26.6.2",
|
|
18
18
|
"typescript": "^7.0.2"
|
|
19
19
|
},
|
|
20
20
|
"dependencies": {
|
package/readme.md
CHANGED
|
@@ -264,8 +264,8 @@ allocation-pool guard policy kinds.
|
|
|
264
264
|
|
|
265
265
|
| V2 scope | Required authority and behavior |
|
|
266
266
|
| --- | --- |
|
|
267
|
-
| `routerEgress` | Private `endpoints` and `rules`, one exact `links` binding per endpoint, a separate veth `handoff`, `protection`, and active `generations`. Private veth/TUN/local DNS and egress share one table so terminal private denial cannot override a separate egress table. Optional `publishedPorts` add the inbound second hop from the handoff to a workload endpoint; both directions are classified into the default conntrack zone ahead of every leased classifier, so a published endpoint port is dedicated to its publication and never becomes leased egress. Optional `hostGrants` forward exact host-origin flows from the handoff to a workload endpoint. Optional `workloadGrants` let one workload open one exact port of another, one way. |
|
|
268
|
-
| `hostTransit` | Exact handoff `link`/`allocations` pairs, complete `protection`, an explicit veth or Ethernet `uplink`, and its current `snatAddress`. It checks each handoff's leased source address and protocol/port range, default conntrack zone, direction, uplink, and protected destinations before outer SNAT. Optional `publishedPorts` add inbound uplink destination NAT inside the same generation. Optional `hostGrants` let the host's own address on a handoff dial exact workload ports. Optional `localPlatformEndpoints` serve platform endpoints on the host's own addresses to leased flows. |
|
|
267
|
+
| `routerEgress` | Private `endpoints` and `rules`, one exact `links` binding per endpoint, a separate veth `handoff`, `protection`, and active `generations`. Private veth/TUN/local DNS and egress share one table so terminal private denial cannot override a separate egress table. Optional `publishedPorts` add the inbound second hop from the handoff to a workload endpoint, one port or a port range; both directions are classified into the default conntrack zone ahead of every leased classifier, so a published endpoint port is dedicated to its publication and never becomes leased egress. A `symmetric` publication also lets the workload open flows from its published ports. Optional `hostGrants` forward exact host-origin flows from the handoff to a workload endpoint. Optional `workloadGrants` let one workload open one exact port of another, one way. |
|
|
268
|
+
| `hostTransit` | Exact handoff `link`/`allocations` pairs, complete `protection`, an explicit veth or Ethernet `uplink`, and its current `snatAddress`. It checks each handoff's leased source address and protocol/port range, default conntrack zone, direction, uplink, and protected destinations before outer SNAT. Optional `publishedPorts` add inbound uplink destination NAT inside the same generation, one port or a port range, and a `symmetric` publication also carries the workload's own flows from its published ports out through the uplink. Optional `hostGrants` let the host's own address on a handoff dial exact workload ports. Optional `localPlatformEndpoints` serve platform endpoints on the host's own addresses to leased flows. |
|
|
269
269
|
| `allocationPoolGuard` | An authenticated `authorityDigest` and complete current allocation-pool `prefixes`. Installs host-wide IPv4 destination denial before any handoff exists, without link, uplink or SNAT dependencies. Optional `hostGrants` are its only exceptions. Optional `localTcpPortOwners` restrict loopback TCP ports to one local user each. |
|
|
270
270
|
|
|
271
271
|
`allocationPoolGuard` accepts 1–64 canonical, disjoint RFC1918 prefixes. Supply
|
|
@@ -317,7 +317,8 @@ handoffs, unmatched handoff traffic and IPv6 forwarding are denied.
|
|
|
317
317
|
|
|
318
318
|
`hostTransit.publishedPorts` is optional. Absent and empty are the same canonical
|
|
319
319
|
policy, so an unpublished scope keeps its exact previous digest and compiled bytes.
|
|
320
|
-
Each entry is `{ protocol, hostPort, targetPort, targetAddress, hostIp }
|
|
320
|
+
Each entry is `{ protocol, hostPort, targetPort, targetAddress, hostIp }`, with an
|
|
321
|
+
optional `hostPortEnd` and an optional `symmetric`. `hostIp`
|
|
321
322
|
must be an exact current uplink address, which rtnetlink verifies at apply, recovery
|
|
322
323
|
and inspection; there is no wildcard, secondary-link or default-route inference.
|
|
323
324
|
`targetAddress` must be a current leased `transitSourceAddress` of one allocation,
|
|
@@ -328,21 +329,59 @@ including across uplink addresses, and a port published on the `snatAddress` may
|
|
|
328
329
|
fall inside any leased outbound source-port range of the same protocol, because outer
|
|
329
330
|
SNAT translates to that same address.
|
|
330
331
|
|
|
331
|
-
|
|
332
|
-
|
|
333
|
-
|
|
334
|
-
|
|
335
|
-
|
|
336
|
-
|
|
337
|
-
|
|
338
|
-
|
|
339
|
-
|
|
340
|
-
|
|
341
|
-
|
|
342
|
-
|
|
343
|
-
|
|
344
|
-
|
|
345
|
-
|
|
332
|
+
`hostPortEnd` publishes the range `hostPort..hostPortEnd`. It must be greater than
|
|
333
|
+
`hostPort`, and `targetPort` must equal `hostPort`: a range keeps every port, so port
|
|
334
|
+
`p` of the uplink reaches port `p` of `targetAddress`. One port has exactly one
|
|
335
|
+
canonical form, without `hostPortEnd`. Ranges follow the same rules as single ports:
|
|
336
|
+
no two publications of one protocol overlap anywhere in the scope, and no published
|
|
337
|
+
port on the `snatAddress` falls inside a leased range. A range translates the
|
|
338
|
+
address only, which never changes a port.
|
|
339
|
+
|
|
340
|
+
Publications are set elements, not rules. The generation compiles CONSTANT named
|
|
341
|
+
sets and maps of concatenated fields in which the port fields are inclusive
|
|
342
|
+
intervals (the kernel's pipapo backend), so one element holds one port or a whole
|
|
343
|
+
range: `published_port` maps a one-port publication's uplink address, protocols and
|
|
344
|
+
port to its target address and port, `published_address` maps a range to its target
|
|
345
|
+
address alone, and `published_forward` holds every publication's inbound tuple. A
|
|
346
|
+
constant number of rules looks them up whatever the number of publications: one
|
|
347
|
+
destination NAT rule per map in this table's own prerouting `nat` chain at priority
|
|
348
|
+
-100, and four forward admissions ahead of the capture barrier — the ESTABLISHED
|
|
349
|
+
original direction, a NEW opening per protocol with opening TCP restricted to SYN
|
|
350
|
+
with FIN/RST/ACK clear, and the ESTABLISHED reply. Every rule checks the uplink, the
|
|
351
|
+
default zone and the direction, and every key carries the exact handoff (index and
|
|
352
|
+
name), the packet and conntrack protocols, the translated current address and port
|
|
353
|
+
and the original conntrack destination, so foreign NAT cannot borrow the exception
|
|
354
|
+
for another flow. Published traffic is not source-translated, so the workload
|
|
355
|
+
observes the actual client address. The chain, its sets, its translations and its
|
|
356
|
+
admissions are created, replaced and deleted with the generation: a target without
|
|
357
|
+
the entry, a failed batch, or release removes them, and a caller never needs a
|
|
358
|
+
separate teardown. Owner verification reads every set back with its description,
|
|
359
|
+
every element with its interval bounds and map data, and every lookup with its
|
|
360
|
+
registers, so apply, inspection, adoption and recovery reject a changed element
|
|
361
|
+
exactly as they reject a changed rule.
|
|
362
|
+
|
|
363
|
+
Without `symmetric`, a publication carries inbound flows only: a flow the workload
|
|
364
|
+
opens from its target address and port toward the uplink is denied. With `symmetric:
|
|
365
|
+
true`, a flow from exactly `targetAddress` and the published target port(s) on the
|
|
366
|
+
exact handoff may leave through the uplink, NEW (a TCP opening only with SYN) or
|
|
367
|
+
ESTABLISHED, with its ESTABLISHED replies, and outer source NAT translates it to
|
|
368
|
+
`hostIp` and the published uplink port. These flows are elements of one more map,
|
|
369
|
+
`published_symmetric`, keyed on the handoff, protocols and the live and original
|
|
370
|
+
source and holding the uplink address and port span; four admissions look it up as a
|
|
371
|
+
set and one postrouting rule translates through it. The workload's outbound flows therefore leave
|
|
372
|
+
from the same address and port its clients reach, as Docker host-mode NAT kept them;
|
|
373
|
+
a range keeps each port unless the host itself already holds that exact tuple on
|
|
374
|
+
`hostIp`. The admissions follow the capture barrier, so every protected current or
|
|
375
|
+
original destination, the platform endpoints included, stays denied. A symmetric
|
|
376
|
+
target port may not fall inside a leased range of `targetAddress`, and may not overlap
|
|
377
|
+
the target ports of another publication of the same address and protocol. Absent and
|
|
378
|
+
`false` are the same canonical policy, digest and compiled graph.
|
|
379
|
+
|
|
380
|
+
Input bounds are 1024 entries. The rule graph does not grow with them: with every
|
|
381
|
+
kind of publication present the host compiles at most eleven rules and four sets,
|
|
382
|
+
about 13,400 bytes of the 100,000-byte rule reserve, and each publication costs
|
|
383
|
+
about 150 to 300 bytes of set elements, which share the target budget described
|
|
384
|
+
under host grants. 64 symmetric ranges take about 19,600 element bytes.
|
|
346
385
|
The caller still owns the route to `targetAddress` through that handoff, the
|
|
347
386
|
workload listener, and every other packet owner on the host. As with leased egress,
|
|
348
387
|
an ACCEPT here cannot override Docker's independent FORWARD DROP; the
|
|
@@ -366,7 +405,8 @@ There is no broad ESTABLISHED/RELATED bypass.
|
|
|
366
405
|
`routerEgress.publishedPorts` is optional and is the second hop of a published
|
|
367
406
|
host port. Absent and empty are the same canonical policy, so an unpublished scope
|
|
368
407
|
keeps its exact previous digest and compiled bytes. Each entry is `{ protocol,
|
|
369
|
-
transitPort, endpointPort, endpointAddress, transitSourceAddress }
|
|
408
|
+
transitPort, endpointPort, endpointAddress, transitSourceAddress }`, with an
|
|
409
|
+
optional `transitPortEnd` and an optional `symmetric`.
|
|
370
410
|
`transitSourceAddress` must be a current leased generation address that is present
|
|
371
411
|
on `handoff`, which rtnetlink verifies at apply, recovery and inspection; the
|
|
372
412
|
handoff can carry several current leased addresses, so the arrival address is
|
|
@@ -379,11 +419,24 @@ published exactly once across the scope, and a published port may not fall insid
|
|
|
379
419
|
any leased outbound source-port range of the same protocol on that transit address,
|
|
380
420
|
because source NAT translates to that same address.
|
|
381
421
|
|
|
382
|
-
|
|
383
|
-
|
|
384
|
-
|
|
385
|
-
|
|
386
|
-
|
|
422
|
+
`transitPortEnd` publishes the range `transitPort..transitPortEnd`, the second hop of
|
|
423
|
+
a host range. It must be greater than `transitPort`, and `endpointPort` must equal
|
|
424
|
+
`transitPort`, so every port keeps its number on both hops; one port has exactly one
|
|
425
|
+
canonical form, without `transitPortEnd`. No two publications of one protocol overlap
|
|
426
|
+
anywhere in the scope, and no published port falls inside a leased range of its
|
|
427
|
+
transit address. The destination NAT of a range translates the address only.
|
|
428
|
+
|
|
429
|
+
Publications are set elements here too: `published_arrival` (protocol, transit
|
|
430
|
+
address and port), `published_endpoint` (workload link, protocol, address and port,
|
|
431
|
+
the union of every publication's workload ports as disjoint canonical intervals),
|
|
432
|
+
the `published_port` and `published_address` translation maps and
|
|
433
|
+
`published_forward`, with the port fields as inclusive intervals. A constant number
|
|
434
|
+
of rules looks them up whatever the number of publications: one destination NAT
|
|
435
|
+
rule per map in this table's own prerouting `nat` chain at priority -100, two raw
|
|
436
|
+
classifications, and four forward admissions placed ahead of the capture barrier —
|
|
437
|
+
the ESTABLISHED original direction, a NEW opening per protocol with opening TCP
|
|
438
|
+
restricted to SYN with FIN/RST/ACK clear, and the ESTABLISHED reply. The handoff
|
|
439
|
+
ingress classification precedes the
|
|
387
440
|
barrier's protected-source denial, because the uplink or management network that
|
|
388
441
|
reaches a published port is itself protected space. The endpoint ingress
|
|
389
442
|
classification matches the exact workload link, address and published endpoint
|
|
@@ -393,24 +446,47 @@ would otherwise carry the workload's published packets into that generation, whe
|
|
|
393
446
|
the publication's flow does not exist. Both classifications match the arriving
|
|
394
447
|
tuple only, set no generation zone and leave the flow in the default conntrack
|
|
395
448
|
zone. The private anti-spoof guards still run first, and every admission checks
|
|
396
|
-
|
|
397
|
-
translated current address/port and the
|
|
449
|
+
the handoff, the default zone and the direction, and looks up the exact workload
|
|
450
|
+
link (index and name), both protocols, the translated current address/port and the
|
|
451
|
+
original conntrack destination.
|
|
398
452
|
|
|
399
|
-
A published
|
|
453
|
+
A published endpoint port is therefore dedicated to its publication: a new
|
|
400
454
|
outbound flow that the workload opens from that exact address and port on that
|
|
401
|
-
exact link stays in the default zone too and is
|
|
402
|
-
outbound source NAT.
|
|
403
|
-
|
|
404
|
-
|
|
405
|
-
|
|
406
|
-
|
|
407
|
-
|
|
408
|
-
|
|
409
|
-
|
|
410
|
-
|
|
411
|
-
|
|
412
|
-
the
|
|
413
|
-
|
|
455
|
+
exact link stays in the default zone too and is never translated by the leased
|
|
456
|
+
outbound source NAT. Without `symmetric` such a flow is denied. Inbound published
|
|
457
|
+
traffic is not source-translated on either hop, so the workload observes the actual
|
|
458
|
+
client address. The chain, its sets, its translations, its classifications and its
|
|
459
|
+
admissions are created, replaced and deleted with the generation: a target without
|
|
460
|
+
the entry, a failed batch, or release removes them.
|
|
461
|
+
|
|
462
|
+
With `symmetric: true` the workload may open flows from exactly `endpointAddress`
|
|
463
|
+
and the published endpoint port(s) on its exact link toward the handoff, NEW (a TCP
|
|
464
|
+
opening only with SYN) or ESTABLISHED, with their ESTABLISHED replies, and source
|
|
465
|
+
NAT translates them to `transitSourceAddress` and the published transit port. With
|
|
466
|
+
the matching symmetric `hostTransit` publication the flow reaches the uplink from
|
|
467
|
+
`hostIp` and the published uplink port, the same address and port its clients
|
|
468
|
+
reach, as Docker host-mode NAT kept it; a range keeps each port unless that exact
|
|
469
|
+
tuple is already taken on the translated address. These flows are elements of the
|
|
470
|
+
`published_symmetric` map, keyed on the workload link, protocols and the live and
|
|
471
|
+
original source and holding the transit address and port span, behind four
|
|
472
|
+
admissions and one postrouting translation. Both directions are already in the
|
|
473
|
+
default zone through the publication's endpoint and handoff classifications, so no
|
|
474
|
+
leased classifier can capture them. The admissions follow the capture barrier, so
|
|
475
|
+
every protected current or original destination, the platform endpoints included,
|
|
476
|
+
stays denied; outbound flows anywhere else still need their leased grants. The
|
|
477
|
+
endpoint ports of a symmetric publication may not overlap the endpoint ports of
|
|
478
|
+
another publication of the same address and protocol, since each flow has exactly
|
|
479
|
+
one published source. Absent and `false` are the same canonical policy, digest and
|
|
480
|
+
compiled graph.
|
|
481
|
+
|
|
482
|
+
Input bounds are 1024 entries. The rule graph does not grow with them: with every
|
|
483
|
+
kind of publication present the router compiles at most thirteen rules and six sets,
|
|
484
|
+
about 15,100 bytes of the 100,000-byte rule reserve, and each publication costs
|
|
485
|
+
about 250 to 450 bytes of set elements, which share the target budget with the
|
|
486
|
+
router's endpoints, rules and grants. 64 symmetric ranges take about 28,000 element
|
|
487
|
+
bytes. The caller still owns both hops: the exact matching `hostTransit`
|
|
488
|
+
publication, the route to the workload through this router, the workload listener
|
|
489
|
+
and every other packet owner.
|
|
414
490
|
|
|
415
491
|
#### Host grants
|
|
416
492
|
|
|
@@ -466,7 +542,7 @@ foreign set exactly as they reject a changed rule; a bound CONSTANT set also
|
|
|
466
542
|
refuses element changes from any other socket.
|
|
467
543
|
|
|
468
544
|
Absent and empty are the same canonical policy, so a scope without grants keeps its
|
|
469
|
-
exact previous digest and compiled bytes and has no set. Withdrawing a grant is an
|
|
545
|
+
exact previous digest and compiled bytes and has no host grant set. Withdrawing a grant is an
|
|
470
546
|
ordinary atomic replacement: the batch deletes the previous rules, chains and sets
|
|
471
547
|
by handle and creates the complete target, so a target without the entry, a failed
|
|
472
548
|
batch or release removes it, and an established flow stops with it. Apply grants
|
|
@@ -531,9 +607,9 @@ tuple in the CONSTANT set `workload_grant`, so a spoofed, renamed or translated
|
|
|
531
607
|
never matches and the destination can never open toward the source. Endpoint prefixes
|
|
532
608
|
are protected, so both directions stay in the default zone ahead of every leased
|
|
533
609
|
classifier. Owner verification reads both sets back element by element. Up to 1024
|
|
534
|
-
grants fit; the sets share the scope's
|
|
535
|
-
|
|
536
|
-
work. Absent and empty are the same canonical policy, digest and compiled bytes, and
|
|
610
|
+
grants fit; the sets share the scope's target budget with `hostGrants` and every
|
|
611
|
+
other set, and 1024 of each fit together beside the reference router; a scope
|
|
612
|
+
beyond the budget is refused before any kernel work. Absent and empty are the same canonical policy, digest and compiled bytes, and
|
|
537
613
|
withdrawal is an ordinary atomic replacement that also stops established flows.
|
|
538
614
|
|
|
539
615
|
These flows need no private `rules`. A private rule is stateless: answering through
|
|
@@ -588,13 +664,20 @@ applied, retained, recovered and released with the rest of the guard's table; af
|
|
|
588
664
|
a reboot the caller applies its retained intent again, like every other member.
|
|
589
665
|
|
|
590
666
|
Input bounds are 1024 grants per scope, the maximum a network projection carries.
|
|
591
|
-
Elements
|
|
592
|
-
|
|
593
|
-
|
|
594
|
-
replacement
|
|
595
|
-
|
|
596
|
-
|
|
597
|
-
212,992
|
|
667
|
+
Elements are sent 256 to a message. At 1024 grants the reference router needs about
|
|
668
|
+
98 KB of elements (two sets), host transit about 66 KB and the guard about 45 KB.
|
|
669
|
+
|
|
670
|
+
Every replacement is one nfnetlink batch in one `sendmsg`, and Linux refuses a
|
|
671
|
+
netlink message larger than the socket send buffer: the requested 1 MiB `SO_SNDBUF`
|
|
672
|
+
is clamped to `net.core.wmem_max` and doubled, 425,984 bytes with the Linux default
|
|
673
|
+
of 212,992. That is the hard ceiling. This package bounds a batch at 320,000 bytes,
|
|
674
|
+
which fits every `wmem_max` of at least 160,016, and splits it: at most 64 bytes of
|
|
675
|
+
batch framing, at most 107,520 bytes for the handle-only deletion of the previous
|
|
676
|
+
graph (fewer than 768 rules, chains and sets of at most 140 bytes each), and
|
|
677
|
+
212,000 bytes for the complete target, of which rules, chains and sets may take at
|
|
678
|
+
most 100,000 and set elements the rest. Scopes that grow with their inputs keep
|
|
679
|
+
those inputs in set elements, so the element share grows where the rule share
|
|
680
|
+
shrinks; `prepare()` refuses a target beyond either bound before any kernel work. The private facade IPC carries up to three complete policies in one
|
|
598
681
|
status and is bounded at 1,048,576 bytes per line; the facade captures at most
|
|
599
682
|
1,000,000 JSON bytes and 65,536 values per request or result. The caller owns the host route
|
|
600
683
|
to the workload through the handoff, selecting the transit host address as the
|
|
@@ -602,18 +685,56 @@ socket's source, the workload listener and every other packet owner.
|
|
|
602
685
|
|
|
603
686
|
Active overlapping classifiers, duplicate zones/labels and conflicting handoff
|
|
604
687
|
allocations reject. Empty router generations retain private routing while denying
|
|
605
|
-
egress.
|
|
606
|
-
|
|
607
|
-
16 ranges per allocation, and 32 host
|
|
608
|
-
|
|
609
|
-
|
|
610
|
-
drop new flows rather than allocate outside the lease.
|
|
611
|
-
|
|
612
|
-
The router
|
|
613
|
-
|
|
614
|
-
|
|
615
|
-
|
|
616
|
-
|
|
688
|
+
egress. Router input bounds are 128 private endpoints and links, 1024 private
|
|
689
|
+
rules, 32 active generations and 1024 total egress grants; other bounds include 128
|
|
690
|
+
protected prefixes, 96 platform endpoints, 16 ranges per allocation, and 32 host
|
|
691
|
+
handoffs/active allocations. The complete target must fit the budget above; the
|
|
692
|
+
router's rules grow only with its protected prefixes, so its elements bind first.
|
|
693
|
+
Exhausted source-port ranges drop new flows rather than allocate outside the lease.
|
|
694
|
+
|
|
695
|
+
The router compiles its private endpoints, directed rules and leased grants into
|
|
696
|
+
named sets and maps behind a fixed set of rules, whatever the number of workloads:
|
|
697
|
+
- `private_link` holds every endpoint link (index and name) and the absent
|
|
698
|
+
interface of router-local traffic; `private_index` and `private_name` hold every
|
|
699
|
+
endpoint index and name for the terminal denials, either alone enough to deny;
|
|
700
|
+
`private_source` holds each link index with its source prefixes. The anti-spoof
|
|
701
|
+
guard drops a packet on an endpoint link whose address is not one of its
|
|
702
|
+
prefixes, and anything but IPv4 there.
|
|
703
|
+
- `private_any` and `private_port` hold the directed rules: incoming and outgoing
|
|
704
|
+
link index (0 for the router itself), addresses, and for a transport rule the
|
|
705
|
+
protocol and ports.
|
|
706
|
+
- Per kind (platform ahead of the protected barrier, public behind it) and per
|
|
707
|
+
shape (whether the grant names a source port), `leased_*_zone` maps a workload
|
|
708
|
+
or router flow's link, protocol and tuple to its generation zone before conntrack,
|
|
709
|
+
`leased_*_return` maps a translated return to its zone, and `leased_*_flow` maps
|
|
710
|
+
the link, both protocols, the live and originally tracked tuple and the zone to
|
|
711
|
+
the grant's source NAT; the filter admissions look it up as a set.
|
|
712
|
+
`leased_label` holds each generation's zone and label, and `leased_generation`
|
|
713
|
+
maps a zone to the label a fresh flow receives.
|
|
714
|
+
A key's link index names its link because the rule also finds the interface's
|
|
715
|
+
index and name in `private_link`, whose indices are unique. Every element is
|
|
716
|
+
checked, like every rule, on apply, inspection, adoption and recovery. A decision
|
|
717
|
+
corpus of about 133,000 packets across five router fixtures, frozen from the
|
|
718
|
+
per-rule compiler of 2.6.0, is reproduced exactly by the set-backed compiler.
|
|
719
|
+
|
|
720
|
+
A router of 64 workloads, each with local DNS, public TCP and UDP egress, grants to
|
|
721
|
+
four platform endpoints and one to three publications, compiles about 66,000 rule
|
|
722
|
+
bytes and 105,000 element bytes; 89 such workloads fit the target budget:
|
|
723
|
+
|
|
724
|
+
| Workloads | Router rules | Router elements | Host rules | Host elements |
|
|
725
|
+
| --- | --- | --- | --- | --- |
|
|
726
|
+
| 1 | 65,820 | 4,968 | 42,944 | 1,128 |
|
|
727
|
+
| 8 | 65,820 | 16,056 | 42,944 | 2,528 |
|
|
728
|
+
| 32 | 65,820 | 54,072 | 42,944 | 7,328 |
|
|
729
|
+
| 64 | 65,820 | 104,760 | 42,944 | 13,728 |
|
|
730
|
+
|
|
731
|
+
**Upgrading from 2.x:** version 3 changes the compiled identity of every schema-v2
|
|
732
|
+
`routerEgress` table and of every scope with `publishedPorts`, while their policy
|
|
733
|
+
digests stay unchanged. Reconcile or release against such a table applied by 2.x
|
|
734
|
+
returns `Conflict`: the new engine cannot adopt the old graph. Recover and release
|
|
735
|
+
each such table with the exact previous engine and its retained intent and
|
|
736
|
+
receipt, under the same traffic fence as below, before upgrading; tables of other
|
|
737
|
+
scopes and without publications are unchanged.
|
|
617
738
|
|
|
618
739
|
**Upgrading from 1.x:** version 2 changes the compiled identity of schema-v2
|
|
619
740
|
`routerEgress` tables. Retained router tables and their receipts cannot be adopted
|
|
@@ -732,6 +853,27 @@ A client whose source port equals a leased grant's destination port on that gran
|
|
|
732
853
|
public address completes the published TCP handshake and receives its UDP reply
|
|
733
854
|
from the published address, so the leased classifiers of the same generation cannot
|
|
734
855
|
capture the workload's published packets.
|
|
856
|
+
Published ranges and symmetric publications are qualified across both hops in the
|
|
857
|
+
same topology, with every publication a set element whose kernel dump matches the
|
|
858
|
+
compiled sets and maps: the first, middle and last port of a UDP range reach their listeners
|
|
859
|
+
with the client address preserved and answer from the published address, the ports
|
|
860
|
+
just outside the range stay dark against pre-policy positive controls, symmetric UDP
|
|
861
|
+
from a range port and from the signalling port and symmetric TCP from the signalling
|
|
862
|
+
port reach the uplink peer from the published address with the same port, a later
|
|
863
|
+
request from that peer reaches the workload on the same flow, the plain publication's
|
|
864
|
+
workload port opens nothing, symmetric flows into the protected union stay denied
|
|
865
|
+
against a pre-policy positive control, leased egress keeps its leased range, and the
|
|
866
|
+
range goes dark with release.
|
|
867
|
+
A 64-workload router is qualified in one batch, every workload in its own namespace
|
|
868
|
+
with local DNS, public TCP and UDP egress, four platform endpoints and a
|
|
869
|
+
publication, the first with the SIP shape: the kernel dumps every private, leased
|
|
870
|
+
and published element and the owner verifies it, and on the first, middle and last
|
|
871
|
+
workload local DNS answers, public UDP egress leaves from the leased range, UDP and
|
|
872
|
+
TCP platform endpoints answer through both hops, the publication reaches its
|
|
873
|
+
listener with the client address preserved, and another workload, a protected
|
|
874
|
+
address outside the platform endpoints and an ungranted platform port stay denied
|
|
875
|
+
against pre-policy positive controls, as does another workload's source address.
|
|
876
|
+
The router compiles 82 rules for the 64 workloads.
|
|
735
877
|
Host grants are qualified across all three tables in the same four-namespace
|
|
736
878
|
topology with leased public egress on every port of both protocols and 1024 grants
|
|
737
879
|
in every scope: a member far from the probed tuples passes, the tuple just outside
|