rcekit 2.35.1__tar.gz → 2.45.3__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (36) hide show
  1. rcekit-2.45.3/CHANGELOG.md +2306 -0
  2. {rcekit-2.35.1 → rcekit-2.45.3}/CONTRIBUTING.md +65 -8
  3. {rcekit-2.35.1/rcekit.egg-info → rcekit-2.45.3}/PKG-INFO +148 -26
  4. {rcekit-2.35.1 → rcekit-2.45.3}/README.md +147 -25
  5. {rcekit-2.35.1 → rcekit-2.45.3}/docs/generation.md +10 -0
  6. {rcekit-2.35.1 → rcekit-2.45.3}/docs/guide.md +23 -6
  7. {rcekit-2.35.1 → rcekit-2.45.3}/docs/reference.md +397 -34
  8. {rcekit-2.35.1 → rcekit-2.45.3}/docs/verify-it-yourself.md +10 -0
  9. {rcekit-2.35.1 → rcekit-2.45.3/rcekit.egg-info}/PKG-INFO +148 -26
  10. {rcekit-2.35.1 → rcekit-2.45.3}/rcekit.egg-info/SOURCES.txt +1 -0
  11. {rcekit-2.35.1 → rcekit-2.45.3}/rcekit.py +2399 -119
  12. {rcekit-2.35.1 → rcekit-2.45.3}/templates/payloads.json +33 -0
  13. rcekit-2.45.3/tests/bench/README.md +225 -0
  14. rcekit-2.45.3/tests/bench/cases/log4shell-cve-2021-44228.json +69 -0
  15. rcekit-2.45.3/tests/bench/cases/struts2-s2-001.json +69 -0
  16. rcekit-2.45.3/tests/bench/cases/webmin-cve-2019-15107.json +62 -0
  17. {rcekit-2.35.1 → rcekit-2.45.3}/tests/bench/runner.py +261 -56
  18. {rcekit-2.35.1 → rcekit-2.45.3}/tests/test_bench_runner.py +314 -6
  19. rcekit-2.45.3/tests/test_docs.py +800 -0
  20. {rcekit-2.35.1 → rcekit-2.45.3}/tests/test_generator.py +3931 -204
  21. rcekit-2.35.1/CHANGELOG.md +0 -1042
  22. rcekit-2.35.1/tests/bench/README.md +0 -138
  23. rcekit-2.35.1/tests/bench/cases/struts2-s2-001.json +0 -46
  24. rcekit-2.35.1/tests/bench/cases/webmin-cve-2019-15107.json +0 -43
  25. rcekit-2.35.1/tests/test_docs.py +0 -309
  26. {rcekit-2.35.1 → rcekit-2.45.3}/LICENSE +0 -0
  27. {rcekit-2.35.1 → rcekit-2.45.3}/MANIFEST.in +0 -0
  28. {rcekit-2.35.1 → rcekit-2.45.3}/SECURITY.md +0 -0
  29. {rcekit-2.35.1 → rcekit-2.45.3}/pyproject.toml +0 -0
  30. {rcekit-2.35.1 → rcekit-2.45.3}/rcekit.egg-info/dependency_links.txt +0 -0
  31. {rcekit-2.35.1 → rcekit-2.45.3}/rcekit.egg-info/entry_points.txt +0 -0
  32. {rcekit-2.35.1 → rcekit-2.45.3}/rcekit.egg-info/top_level.txt +0 -0
  33. {rcekit-2.35.1 → rcekit-2.45.3}/setup.cfg +0 -0
  34. {rcekit-2.35.1 → rcekit-2.45.3}/tests/bench/requests/webmin.txt +0 -0
  35. {rcekit-2.35.1 → rcekit-2.45.3}/tests/test_corpus_embedding.py +0 -0
  36. {rcekit-2.35.1 → rcekit-2.45.3}/tools/embed_corpus.py +0 -0
@@ -0,0 +1,2306 @@
1
+ # Changelog
2
+
3
+ All notable changes to RCEKit are recorded here. The format follows
4
+ [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and RCEKit follows
5
+ [Semantic Versioning](https://semver.org/spec/v2.0.0.html): PATCH for fixes,
6
+ MINOR for new capabilities, MAJOR for breaking changes to the CLI, output
7
+ formats, or the template schema.
8
+
9
+ ## [Unreleased]
10
+
11
+ ### Added
12
+
13
+ - **A tier correction has to reach the prose, not only the class.** Four tests
14
+ now read the tier from `DETECTION_METHODS[name].tier` and hold the sentences
15
+ an operator actually sees against it.
16
+
17
+ `lookup` moved from `confirmed` to `lookup-sink` on the class, and the move
18
+ was made in the README's CVE table -- but `blind_sink_advice` went on
19
+ offering `--methods lookup` as a method that "confirms", in the list where
20
+ `oob` and `file` do mean confirmed execution, and the README's Log4Shell demo
21
+ heading still said `confirmed` over alt text calling the run
22
+ "auto-confirming a blind Log4Shell RCE". Three places, one correction, and
23
+ nothing compared them: the existing test pinned the one line it was written
24
+ for (`time` is marked needs-review only) rather than asking every line the
25
+ same question.
26
+
27
+ So: every `blind_sink_advice` line naming a method must state that method's
28
+ tier and no other; every CVE row must state the tier its method reports;
29
+ every demo heading must match its row in that table; and a recording below
30
+ `confirmed` may not be described as confirming, alt text included, since
31
+ that is the sentence a screen reader reads out.
32
+
33
+ The wording checks match stems rather than words, and denials such as
34
+ "never confirmed" are removed before they are applied. "confirms" alone
35
+ would have readmitted the overclaim through "confirming RCE" or "confirmed
36
+ execution" -- an assertion answering the same way for the right reason and
37
+ the broken one, which is the defect being guarded against. Each test was
38
+ checked by reintroducing the defect it exists for, and each carries a floor
39
+ on how much it inspected: a parser that quietly stopped matching would
40
+ otherwise pass exactly as a clean README does.
41
+
42
+ Test-only: no version bump, and nothing about a run changes.
43
+
44
+ - **The same check, for the two tables under `docs/`.** They were left out when
45
+ the README's were pinned, and they carry the same claim:
46
+ `docs/reference.md` names the tier each method can reach, and
47
+ `docs/guide.md` tells an operator which method to reach for next.
48
+
49
+ `reference.md` states a ceiling, so its rule is a subset rather than an
50
+ equality -- a cell may also name a weaker tier the method really emits, as
51
+ `write` and `deser` both do -- and every registered method must have a row,
52
+ so a capability cannot land unlookupable. `guide.md` has no tier column and
53
+ so is held only to the negative: a row recommending a method that cannot
54
+ confirm may not describe confirmation.
55
+
56
+ Rows are read by column heading rather than by position. Reading the whole
57
+ row made the `guide.md` check skip the one row it was written for, because
58
+ that row's prose names `oob` -- which confirms -- while recommending
59
+ `lookup`, which does not.
60
+
61
+ A reference row naming something that is not a registered method fails rather
62
+ than being passed over. Skipping it left the completeness check one-way: a
63
+ method *removed* from `DETECTION_METHODS` would leave its row behind, every
64
+ remaining row would still match, and the page would go on offering a
65
+ `--methods` value the CLI rejects.
66
+
67
+ Rows split on *unescaped* pipes. Markdown writes a literal pipe in a cell as
68
+ `\|`, which `docs/reference.md` already does in three tables, and splitting
69
+ on every pipe invents a cell: zipping against the header then drops the last
70
+ column outright, so a claim there stops being examined and every one of these
71
+ checks passes without looking at it.
72
+
73
+ A disclaimer is not read as a claim. "never confirmed" was already removed
74
+ before the `confirm` stem was looked for; "without confirmation" and
75
+ "unconfirmed" were not, and the second carries its negation inside the word
76
+ where a rule about preceding words cannot see it. Documentation saying the
77
+ honest thing would have failed the suite. One parser and one denial rule now,
78
+ where there were two of each here and a third in the generator tests.
79
+
80
+ ### Changed
81
+
82
+ - **A bench case may bring its target up once for both halves**, with
83
+ `"share_target": true`. Bringing the container up twice is the largest fixed
84
+ cost in a case and both halves usually hit the same one, so on a fast case it
85
+ is most of the wall clock.
86
+
87
+ It is opt-in, and the default is unchanged, because the teardown between the
88
+ halves is `down -v`: today's control meets a **fresh** target. A case whose
89
+ vulnerable half writes a file, plants a shell or changes a setting would hand
90
+ its control a target it had already altered, and a control measured against a
91
+ contaminated target measures nothing -- which is the one failure a benchmark
92
+ may not have. Validation rejects the key on a case whose control brings up a
93
+ different target, since there is then nothing to share and leaving it set
94
+ would read as though there were.
95
+
96
+ A shared `up` that fails falls back to per-half management rather than
97
+ carrying on, so the run reports `compose up failed` instead of two readiness
98
+ timeouts naming the wrong cause.
99
+
100
+ Both shipped cases set it, because neither half writes anything. Measured on
101
+ `struts2-s2-001` against vulhub on Docker: 33.8s to 22.9s, with both halves
102
+ reaching the same verdicts either way.
103
+
104
+ - **The benchmark reaches a callback method**, with a case for Log4Shell
105
+ (Apache Solr 8.11.0, Log4j 2.14.1). `python tests/bench/runner.py --all` is
106
+ 3/3.
107
+
108
+ Every case until now was in-band, so the listener, the token correlation and
109
+ `confirm_each` had never run against real software -- the area that produced
110
+ three P1 findings while `lookup` was being written.
111
+
112
+ Two things had to give. A JNDI lookup resolves through the system resolver,
113
+ which asks UDP 53, and a developer machine rarely has that port free; a bare
114
+ `--oob-host` IP is no way around it, because a lookup has no second channel
115
+ to carry the token. So a case may now name `run_in`, and the run happens in a
116
+ container on the target's own network at a fixed address, with an override
117
+ pointing the service's `dns:` at it. The repository is mounted read-only and
118
+ the image is a stock Python.
119
+
120
+ The other was the harness's own vocabulary. `VALID_EXPECTATIONS` and
121
+ `CONTROL_EXPECTATIONS` were written out by hand and had drifted: neither
122
+ `lookup-sink` nor `deserialization-sink` was in either, so a case for `lookup`
123
+ or `deser` could not be *loaded*, let alone run. Both are read from
124
+ `DETECTION_METHODS` now.
125
+
126
+ The control is the case: `oob` against the same Solr comes back `negative`
127
+ **although it is exploitable**, because every probe it builds is a shell
128
+ command and a `${jndi:...}` sink runs none of them. That is the gap `lookup`
129
+ was added to close, and until this case ran it rested on a fixture.
130
+
131
+ `{bench}` / `{repo}` now expand in a case's compose argv as well as its
132
+ invocation, and resolve to the mount point when the run is containerised.
133
+
134
+ The results file is created before the container starts, and it rather than
135
+ its directory is made writable. Under Docker's user-namespace remapping,
136
+ container root is a subordinate host UID, so a `mkdtemp` owned by the runner
137
+ at 0700 is not writable from inside: the file would never appear and the case
138
+ would report `nothing-tested`, as though detection had found nothing rather
139
+ than as though the channel had been shut. Granting the one file and not the
140
+ directory keeps anyone else from creating, replacing or unlinking entries
141
+ there.
142
+
143
+ Test-only: no version bump, and nothing about a run changes.
144
+
145
+ ### Fixed
146
+
147
+ - **`--methods` enumerated only some of the methods it accepts.** `write`,
148
+ `lookup` and `deser` were registered in `DETECTION_METHODS` and had never
149
+ once been named in the help, so `--help` described a whole target class as
150
+ out of reach -- an upload that stores a file, a `${jndi:...}` sink, an
151
+ endpoint that deserializes what it is handed -- while the method for it was
152
+ already shipping and documented everywhere else. Each now carries the tier
153
+ its class declares: `confirmed` for `write`, with `needs-review` for a write
154
+ that is served but not interpreted; `lookup-sink` for `lookup`;
155
+ `deserialization-sink` for `deser`; and neither of the last two ever
156
+ `confirmed`.
157
+
158
+ `oob` also called itself "the only confirmed-tier method for a fully blind
159
+ sink". The tier is right, but an exclusivity claim is the kind that goes
160
+ stale without anything failing, so the line now says what the class itself
161
+ says: it confirms a sink with no output channel and no writable web root.
162
+
163
+ A new test in `CLIDocumentationTestCase` holds the list to the registry, the
164
+ way `--eval-engines` is already held to the corpus. It reads the `--methods`
165
+ help alone rather than the whole page, because every one of these names also
166
+ occurs inside some other flag -- `file` in `--request-file`, `write` in
167
+ `--file-write-path`, `time` in `--time-base` -- and it looks for the name
168
+ followed by the parenthesis that opens its description, because the old help
169
+ carried the word `write` inside `file`'s "write+read-back" and a looser
170
+ match would have counted that as documentation.
171
+
172
+ A second pass corrected what `--oob-host` promises those two methods. Its
173
+ help offered "an IP it can reach or a domain delegated to this listener" and
174
+ was scoped `(--methods oob)`, which was true when `oob` was the only caller.
175
+ It is not true now: `LookupCallback.build_probes` returns no probes at all
176
+ for an address literal, so `lookup` reports `nothing-tested`, and `DeserSink`
177
+ builds `<token>.<IP>`, which resolves nowhere, so its DNS gadget can never
178
+ reach `deserialization-sink`. An operator who satisfied the stated
179
+ prerequisite got silence from one method and a capped tier from the other.
180
+ The flag's help, the `--methods` entries for `lookup` and `deser`, the
181
+ message printed when a callback method is selected without a host, and the
182
+ `deser` oracle table in `docs/reference.md` now all say the same thing: an IP
183
+ serves `oob` alone, because only `oob` can put its token in a URL path.
184
+
185
+ `boolean` landed on `main` while this branch was open. It was named in the
186
+ help but left out of the `--methods` row in `docs/reference.md` -- the same
187
+ drift, one table over. Both carry it now, and the new test covered it without
188
+ being told to, which is the point of reading the registry instead of a list.
189
+
190
+ No version bump -- this is documentation and tests only.
191
+
192
+ ## [Unreleased]
193
+
194
+ ## [2.45.3] — 2026-09-24
195
+
196
+ ### Fixed
197
+
198
+ - **A `deser`-only run started no listener, so its gadgets went out with
199
+ nothing to receive them.** `--methods deser --oob-host <delegated-name>`,
200
+ with neither `oob` nor `lookup` selected, built its two DNS-gadget probes and
201
+ sent them to a target while no listener was running -- with the host
202
+ configured perfectly. `deser` reached `deserialization-sink` only when it
203
+ happened to be selected beside a method that required a host; run on its own,
204
+ the way the docs present it, its strongest oracle was structurally dead and
205
+ it capped at `needs-review` without saying so.
206
+
207
+ The gate read `needs_oob_host`, which was answering two questions at once:
208
+ "should the run stop when no host is named" and "does this run need a
209
+ listener". Every other method answers both the same way. `deser` is where
210
+ they come apart -- its shape oracle proves something with no listener at all,
211
+ so a missing host must not stop the run, while its DNS gadget does call back
212
+ and needs one.
213
+
214
+ So the second question gets its own attribute, `uses_oob_host`, and the
215
+ listener starts for any selected method that calls back once a host is named.
216
+ It is declared rather than derived: `needs_oob_host or oob_needs_dns_label`
217
+ gives the right answer for every method shipping today and would miss one
218
+ that called back over HTTP without requiring a host -- which is the shape
219
+ `oob` itself would have if its host were optional.
220
+
221
+ Two notices could now reach a run with no HTTP shape in it, and said "only
222
+ the HTTP shapes are live" as though `oob`'s channel were the only one. They
223
+ say "only the shapes that need no DNS are live", which is true of `oob`'s
224
+ HTTP probes and of `deser`'s shape oracle alike. The message for an
225
+ unbindable HTTP port no longer names `lookup` as the reason a run continues.
226
+
227
+ `uses_oob_host` is the default answer rather than the answer. Whether a
228
+ method calls back can depend on its own flags: `--deser-formats` selects the
229
+ ecosystems, and three of the five -- `php`, `dotnet`, `python_pickle` --
230
+ ship no DNS gadget at all. So both callers ask `builds_callback_probes()` on
231
+ a configured instance instead of reading the class. `--deser-formats php`
232
+ with an address no longer draws a notice blaming the address, because there
233
+ was no callback probe to lose and a delegated domain could not have helped;
234
+ and it no longer starts a listener that nothing would reach.
235
+
236
+ Whether the host's *form* is one a method can use is a second question,
237
+ `can_use_oob_host()`, and only the gate asks it. A listener bound for
238
+ `--methods deser --oob-host 10.0.0.9` receives nothing -- `build_probes`
239
+ drops an address and sends no gadget -- and the run announced "the target
240
+ will open outbound connections" directly beneath the notice saying those
241
+ probes are not sent. It stays out of `builds_callback_probes()`, because the
242
+ stranded notice is *about* the address and an address test there would
243
+ silence it for the one case it exists to report.
244
+
245
+ `lookup` with an address stops getting a listener too. It built no probes
246
+ with one either way, so nothing is lost, and the stranded notice says more
247
+ about that run than the port-53 warnings it no longer reaches.
248
+
249
+ A flaky test met on the way, fixed here because it would have bitten CI
250
+ sooner or later: the fake parsing endpoint behind the `deser` shape
251
+ differential sniffed one byte (`raw[4:5] not in (b"t", b"s")`) to decide
252
+ whether a stream was well-formed. The `noise` form is the format's magic
253
+ followed by random characters, so it passed that check about once in 500
254
+ runs -- measured at 398 in 200,000 -- all three forms answered alike, and
255
+ the oracle's correct `negative` failed a test that wanted `needs-review`.
256
+ The endpoint now reads TC_STRING the way a parser does, a two-byte length
257
+ followed by exactly that many bytes: 0 collisions in 500,000 seeds. Seeding
258
+ the test's RNG would have hidden it instead, and would have gone on hiding a
259
+ real regression in how the noise form is built.
260
+
261
+ This is the second half of #94, found while fixing the first and reported as
262
+ #98 rather than folded into #97: same defect -- probes sent that no callback
263
+ could follow -- reached by a different route. The format-dependent part came
264
+ from an automated review of #97 and the host-form part from one of this
265
+ change, both verified before being accepted.
266
+
267
+ ## [2.45.2] — 2026-09-24
268
+
269
+ ### Fixed
270
+
271
+ - **`deser` fired DNS gadgets at a host that could not resolve, and the warning
272
+ built for exactly that stayed quiet.** `--oob-host` takes a delegated name or
273
+ an address. `oob` uses either: given an address it drops its DNS shapes and
274
+ carries the token in a URL path instead. `lookup` refuses an address
275
+ outright, because a lookup expression has nowhere else to put one.
276
+ `DeserSink` had neither behaviour -- it built `<token>.10.0.0.9` and sent it,
277
+ a real request the target answered from which no callback could follow by
278
+ construction, and then capped at `needs-review` without saying why.
279
+
280
+ `oob_channel_warnings` is the function that exists to say a DNS channel is
281
+ structurally dead. It returned early on an address literal, reasoning that
282
+ such a run "never builds DNS probes and has nothing to warn about" -- true of
283
+ `oob`, and written when `oob` was the only method taking the flag. So the one
284
+ mechanism that could have reported this was the one place it was ruled out.
285
+
286
+ Both halves are fixed together, because either alone still ships the problem.
287
+ `deser` treats an address as no host and sends its shape oracle alone, which
288
+ is what it can prove without a listener anyway. The notice moves into
289
+ `oob_address_strands`, which names the methods an address stranded and the
290
+ tier that puts out of reach, read off each class.
291
+
292
+ It is a separate function rather than another branch of
293
+ `oob_channel_warnings` because the two answer different questions at
294
+ different times. "Can the DNS shapes reach this listener" is only worth
295
+ asking once a listener exists, and that block runs only for a method that
296
+ *requires* a callback host. `deser` does not require one, so it never
297
+ reaches that block -- a notice left there could not have fired for the
298
+ method that prompted it. The new check runs on every verification run that
299
+ names a host.
300
+
301
+ Which methods those are is a class attribute now, `oob_needs_dns_label`,
302
+ declared beside `needs_oob_host`. Deciding it at the one place that asked is
303
+ how the two drifted apart to begin with.
304
+
305
+ The test that pinned the old silence asserted "no DNS shape is ever built" --
306
+ a claim about `oob` wearing the shape of a claim about the flag. Its scope
307
+ narrows to `oob`, and the stranded case gets tests of its own.
308
+
309
+ Found while verifying an automated review finding on #93; reported as #94 and
310
+ deliberately kept out of that documentation-only PR.
311
+
312
+ ## [2.45.1] — 2026-09-24
313
+
314
+ ### Fixed
315
+
316
+ - **`--max-payloads` is a bound on requests, and it was being checked after
317
+ they had been sent.** The cap was measured against the number of result
318
+ *rows*. For a method that answers from each probe those are the same number,
319
+ so nothing showed. An aggregate method reports one row however many probes it
320
+ cost, so the whole series went out and the cap noticed afterwards:
321
+
322
+ | `--max-payloads 1` | the cost line said | requests sent |
323
+ |---|---|---|
324
+ | `reflected`, `eval` | 1 | 1 |
325
+ | `time` | 1 | **12** |
326
+ | `deser` | 1 | **15** |
327
+ | `boolean` | 1 | **27** |
328
+
329
+ A clean split: every aggregate method overran and no per-probe method did.
330
+ That line is the only thing an operator bounding a monitored engagement has
331
+ to go on before the traffic starts, and it was wrong in the direction that
332
+ matters — under, not over.
333
+
334
+ Rows were the wrong meter in a second way, which the first fix exposed. A
335
+ series the budget declines costs requests and produces no row at all, so the
336
+ next carrier recomputed the same allowance and fired again — `time` at
337
+ `--max-payloads 5` sent 4 requests per carrier with the cap never moving. And
338
+ in a mixed run, a per-probe method never saw what an aggregate one had
339
+ already spent: `boolean,eval` at 30 sent 56. The budget is counted in
340
+ requests for the whole call now. Swept across every method this build can run
341
+ without a callback host, alone and in combination, at seven caps: 63
342
+ combinations, worst overrun 0.
343
+
344
+ - **A measurement that cannot be finished is no longer started.** What a budget
345
+ may do to a series depends on where the method's answer lives, and the class
346
+ already knows: `DetectionMethod.decides_per_probe()` reads it from whether
347
+ the class overrides `confirm_each`, rather than from a list of names that
348
+ would go stale like every other one in this repository.
349
+
350
+ Where each probe answers — `deser`, `lookup`, `oob` — running out of budget
351
+ stops the series and the probes already sent keep their verdicts. Where only
352
+ the series answers — `time`, `boolean` — a part of one is not a weaker answer
353
+ but a wrong one: `time` reports `negative` from a screen with no regression
354
+ behind it and from a regression short of four samples, and a `boolean` series
355
+ whose anchors never went out reads the same way. Both are false cleans, and
356
+ the budget would have been manufacturing them. So a wave that does not fit
357
+ abandons the measurement rather than cutting it short.
358
+
359
+ - **An abandoned measurement leaves an `inconclusive` row, not a silence.**
360
+ Dropping it quietly let the *other* carriers describe the run, and the other
361
+ carriers are the ones with nothing to find: measured at `--max-payloads 12`
362
+ against a sink that honours an injected sleep, the unix carrier's regression
363
+ was abandoned for budget and a windows carrier's honest "no separator
364
+ delayed" was the only row left — so the run reported `negative` for a target
365
+ that was vulnerable. `inconclusive` is what an abandoned measurement is in
366
+ the word the tool already uses, and it outranks `negative` in the run
367
+ verdict. Swept across 31 caps against a sink that really does delay: no cap
368
+ reports it clean, and none overruns.
369
+
370
+ - **What the budget declined is reported by name**, in a sixth run-wide tally
371
+ beside the profile drops, the safety holds, the reach notes, the settled
372
+ carriers and the refusals. A ladder that shrinks quietly is indistinguishable
373
+ from a target with nothing to find:
374
+
375
+ ```
376
+ [detect] --max-payloads held back 1 measurement(s) that could not have reached a
377
+ verdict within the budget:
378
+ [detect] 1 x boolean/raw needs 7 requests to reach a verdict and --max-payloads left 3
379
+ [!] No probes were built, so NOTHING WAS TESTED - this is not a negative result.
380
+ ```
381
+
382
+ - **The cost line and the run make the same decision.** `estimate_detection_probes`
383
+ applies the same all-or-nothing rule, so under a cap the figure printed before
384
+ the traffic is the traffic. Uncapped it stays a floor by documented design — a
385
+ wave a method picks after seeing its own timings cannot be predicted from
386
+ there.
387
+
388
+ - **`sent N probes` is a claim about what the target received**, and it was
389
+ counting result rows. Those are the same number for a method that answers
390
+ from each probe, which is why it read true for so long -- but an aggregate
391
+ method reports one row for a whole series, so a `time` run that put 20
392
+ requests on a target announced 5, and a measurement this release declines
393
+ would have announced one probe for traffic that never left. Both the run
394
+ summary and the per-point line report deliveries now, with the result count
395
+ beside them.
396
+
397
+ One figure in the 2.45.0 entry below came from that line and was mislabelled
398
+ with it. Re-measured against the same sink: `--methods reflected,eval,time`
399
+ put **2883** requests on the target, not 2426, which was the number of
400
+ results. Every one of them was still `negative`, so the claim stands; the
401
+ number did not.
402
+
403
+ - **An evasion retry is a request, and the meter could not see it.** `_escalate`
404
+ fired without incrementing `delivered_probes`, so retries were traffic
405
+ outside the bound this release makes real: measured at `--max-payloads 5
406
+ --evade high` against a filter that refuses whitespace, the target received
407
+ 12 requests and the run recorded 5. Counting them was half the fix. The other
408
+ half is that the budget is checked *per rung* rather than once before the
409
+ ladder -- one refused probe is retried at `low` and again at `high`, so it
410
+ can cost three requests, and a check at the call site alone still overshot by
411
+ one at `--max-payloads 10 --evade high`. Swept over caps 1 to 25 at all three
412
+ evasion settings: 75 combinations, no overrun, and the counter equal to the
413
+ requests the target actually received in every one.
414
+
415
+ ### Security
416
+
417
+ No authorization, deny-by-default or matrix-evaluation boundary moves, and
418
+ nothing widens `confirmed`. The effect is one-way: fewer requests than before,
419
+ and a run that could not afford to test something now says so instead of
420
+ reporting it clean.
421
+
422
+ ## [2.45.0] — 2026-09-23
423
+
424
+ ### Added
425
+
426
+ - **`--methods boolean` — a sink that evaluates a predicate and renders nothing
427
+ of it.** MongoDB `$where` is the shape, and this repository has carried the
428
+ note for two releases: a JS sandbox with no shell, no egress and no value in
429
+ the response, only a document set that a predicate narrows. Every shipped
430
+ oracle is structurally blind to it. Measured against exactly that sink:
431
+
432
+ ```
433
+ [detect] methods: reflected, eval, time
434
+ [detect] sent 2426 probes: negative=2426
435
+ ```
436
+
437
+ 2426 requests and a clean negative on a target that evaluates whatever it is
438
+ handed. The new method reads the one channel left — whether the *shape* of the
439
+ response changed between a true predicate and a false one.
440
+
441
+ **It is `needs-review` and there is no path from here to `confirmed`.** Not
442
+ because the signal is weak, but because of what it cannot distinguish: against
443
+ a sandboxed `eval` sink and against a plain SQLite comparison it produced an
444
+ identical clean differential in 40 runs each, and a query engine comparing two
445
+ numbers is not remote code execution. Extracting a locally computed product
446
+ bit by bit through the channel was tried and does not fix it — it recovers the
447
+ product through both sinks alike, for about 80 requests and a string function
448
+ a sandbox may well deny.
449
+
450
+ **The naive form of this oracle is unusable**, which is why none of it is.
451
+ `1==1` against `1==2`, with a changed response read as a finding, called a
452
+ target that only *reflected* its input vulnerable in 40 runs out of 40, and
453
+ one whose response merely wobbled in 32 of 40. Four guards, each a measured
454
+ false-finding rate rather than a precaution:
455
+
456
+ - **Compare structure, not the body and not its length.** A reflected payload
457
+ lands in the text between two tags, and the text between two tags is what
458
+ the signature throws away. A length-based signature claimed a differential
459
+ in 13 of 25 runs against a reflect-only target; comparing raw bodies was
460
+ unusable outright, reading `unstable` in 25 of 25 runs against a target that
461
+ *was* vulnerable, because one CSRF token makes every response unique.
462
+ - **Several independently randomised pairs, not one.** Against a target whose
463
+ response varies on its own, one pair claimed a differential in 46 of 200
464
+ runs; two claimed none in 200. `--probe-depth quick` trades four pairs for
465
+ two and never for one.
466
+ - **Randomised firing order.** A target that never reads the payload but
467
+ degrades part-way through a run splits an ordered true-then-false series
468
+ perfectly: at the worst point of a swept degradation, 100 false findings out
469
+ of 100.
470
+ - **An anchor before and after the series, each a different true predicate.**
471
+ Shuffling alone still left 2 in 100, which is just the chance a shuffle
472
+ lands separable. Re-measuring the channel afterwards caught it 100 times in
473
+ 100, because a target that moved during the series cannot answer the
474
+ closing anchor the way it answered the opening one. Sending *one* anchor
475
+ payload three times does not merely weaken that: a cache keyed on the query
476
+ string answers the repeats from its store, so the closing anchor agrees
477
+ with the opening one whatever the target did in between. Measured against
478
+ an input-blind target that degrades mid-series, identical anchors caught it
479
+ in 36 of 39 runs live and in 0 of 39 behind a cache. Every probe payload is
480
+ unique, so nothing else in the series is replayable.
481
+
482
+ With every guard on, a genuinely evaluating target still read as a
483
+ differential in 100 runs of 100.
484
+
485
+ - **A channel that cannot carry one bit is `inconclusive`, never `negative`.**
486
+ If the same probe draws two different shapes, or the shape moves while the
487
+ series is being fired, the run says so. `negative` asserts the probes reached
488
+ the target and found nothing; here they reached it and no answer could be read
489
+ out of them, which is the same false clean `blocked` and `nothing-tested`
490
+ exist to prevent, one oracle further in. No new verdict: there are still nine.
491
+
492
+ - **`response_shape()`** — a response reduced to its structure, with everything
493
+ it said removed. Three readings, because a response is one of three things and
494
+ the wrong reading is not a near miss: a JSON document keeps its keys, nesting
495
+ and list lengths and drops every scalar; markup keeps its tag skeleton;
496
+ anything else keeps one marker per word per line. Measured on a JSON sink, the
497
+ markup reading was unusable — `unstable` in 25 of 25 runs — and the shape tree
498
+ read the differential in 25 of 25.
499
+
500
+ ### Changed
501
+
502
+ - **`CODE_POSITION_CONTEXTS`** names the contexts that carry the injected value
503
+ as code rather than as a value, and `boolean` is offered every other one. The
504
+ first cut of this asked whether the context had a break-out prefix at all, and
505
+ got both halves wrong: it refused `attribute`, `attribute_unquoted`,
506
+ `xml_cdata` and `yaml`, whose delimiters open and close *around* the value and
507
+ leave a predicate exactly where a predicate belongs, and it offered
508
+ `unix_shell`, `windows_cmd` and `powershell`, which have no delimiters at all
509
+ and run the value as a command. The suite enumerates both sides, so a context
510
+ added to the corpus fails until somebody decides which one it is on.
511
+
512
+ - **`costly` now asks whether one probe buys an *answer*, not whether it costs
513
+ more than one response.** The two were the same question while every method's
514
+ probe was also its unit of information. `boolean` is the first where they come
515
+ apart: each of its probes is one ordinary request and none of them means
516
+ anything alone, because the answer is the partition across the whole series.
517
+ Read the old way it would have landed in the wave the enumeration driver runs
518
+ *first* — the one that exists to be answered cheaply — ahead of `reflected`
519
+ and `eval` and spending the same per-question budget, at 27 requests before it
520
+ could say a word. It sits with `time` instead, and the suite now holds every
521
+ aggregate method to that.
522
+
523
+ ### Security
524
+
525
+ - **The `OR` connectives ship behind `--verify-active-risk stateful`.** A
526
+ predicate probe breaks out of a condition the application already wrote, and
527
+ the connective is this method's command separator. `AND` differentiates only
528
+ where the application's own predicate is true and `OR` only where it is false,
529
+ so they are complements and dropping `OR` is a blind spot rather than a
530
+ saving. But a true predicate `OR`-ed into a `DELETE … WHERE` took a table from
531
+ 3 rows to 0, where the same predicate `AND`-ed into it left all 3 — so it goes
532
+ at the top rung, held back by default, with the run naming every shape it held
533
+ and the flag that sends it. `AND` and the bare form stay `safe` and change
534
+ nothing.
535
+
536
+ - **Nothing widens `confirmed`.** The new method's ceiling is one tier below it
537
+ and the suite holds that as behaviour, not as an attribute: `confirm_series`
538
+ is driven across 200 series including the perfect one, and none of them
539
+ reaches `confirmed`.
540
+
541
+ ## [2.44.0] — 2026-09-23
542
+
543
+ ### Fixed
544
+
545
+ - **`--evade low` substituted inside quoted programs and broke them.** Every
546
+ space became `${IFS}`, including the ones inside `awk 'BEGIN{print "RK" a+b
547
+ "RK"}'` — and inside single quotes `${IFS}` is literal text, not an
548
+ expansion, so awk was handed `BEGIN{print${IFS}"RK"...` and answered with a
549
+ syntax error.
550
+
551
+ Measured shape by shape against an unfiltered target: **8 probe shapes that
552
+ the canonical form executes broke at the rung, and none improved.** The
553
+ substitution stops at a quote now. Double quotes are left alone from the
554
+ other side — `${IFS}` *does* expand inside them, so substituting there would
555
+ change the string the target computes rather than the spacing around it.
556
+
557
+ ### Changed
558
+
559
+ - **A rung is a retry for a refused probe, not a posture for the run.** It was
560
+ applied to every probe regardless of whether anything was being filtered,
561
+ which is a pure loss on a target with no filter. Against a filter that blocks
562
+ whitespace it turned 1 confirmation into 5; against an unfiltered target it
563
+ now costs **zero** extra requests, because nothing was refused.
564
+
565
+ `--evade` is therefore a **ceiling**. Every probe goes out canonical, and
566
+ only a refused one is retried, up to that ceiling — at most one request per
567
+ rung. The ladder a run builds no longer depends on the rung at all: all three
568
+ settings build the same 42 probes, of which the same 32 execute. Before, the
569
+ rung built a smaller ladder and 13 fewer of its probes ran.
570
+
571
+ The run reports how many retries it made and which rung got through, whatever
572
+ it concluded — including a run that confirmed, where the retry is the reason
573
+ it did.
574
+
575
+ ### Added
576
+
577
+ - **A second evasion rung, `high`.** The measurement turned up two classes of
578
+ filter and the shipped rung only addressed one. `low` removes whitespace;
579
+ `high` also splits the command word with an expansion that vanishes
580
+ (`ec$@ho`), for a filter matching command names. It is applied before the
581
+ whitespace substitution, because afterwards the split lands inside `${IFS}`
582
+ and makes `${I$@FS}` — neither an expansion nor a command.
583
+
584
+ Four boundaries the retry has to respect, each of which was a way to turn a
585
+ vulnerable target into a `negative` — worse than the `blocked` the rung sits
586
+ beside, because `blocked` at least says the run learned nothing:
587
+
588
+ - **Unix shell probes only.** `${IFS}` and `$@` are POSIX. A cmd.exe probe
589
+ rewritten with them loses its spaces, so a whitespace filter answers 200
590
+ and the retry counts as a win while cmd.exe cannot run it.
591
+ - **A break-out context opens by *closing* a quote.** Reading that leading
592
+ quote as an opener left `'; echo …` untouched, so the rung did nothing on
593
+ exactly the contexts a filter is most likely to sit in front of.
594
+ - **A followup is read again after a retry that lands.** `file` writes its
595
+ token on the request that arrives, so a body read before the retry is a
596
+ read of a file that did not exist yet.
597
+ - **A redirect is never retried.** The build-time transform took an explicit
598
+ `evade=False` for these, and that parameter stopped doing anything when the
599
+ rung became a retry — a guard lost in the move, restored where the retry
600
+ now happens.
601
+
602
+ The aggregate methods escalate too. `time`, `oob`, `lookup` and `deser` take a
603
+ different branch, and leaving it out meant the documented ceiling did nothing
604
+ for four of the eight methods — the same branch, and the same omission, as the
605
+ refusal check one change earlier.
606
+
607
+ ## [2.43.0] — 2026-09-23
608
+
609
+ ### Added
610
+
611
+ - **A ninth verdict: `blocked`.** A run whose payloads a filter refuses has
612
+ learned nothing about the sink, and it used to say otherwise. Measured
613
+ against a real command injection behind a filter that 403s a space or a
614
+ separator:
615
+
616
+ ```
617
+ before: sent 10 probes: negative=10
618
+ after: sent 10 probes: blocked=10
619
+ ```
620
+
621
+ `negative` asserts that the probes **reached** the target — the project's own
622
+ tier rule says so. They reached a filter. That is the same false clean
623
+ `nothing-tested` exists to prevent, one level further in.
624
+
625
+ **The signal is differential**, like everything else this tool decides: the
626
+ payload-free control got through and the probe did not, so what was refused
627
+ is the payload. An endpoint answering 403 to everything — an auth wall, a
628
+ path that does not exist for this session — refuses the control too and is
629
+ not mistaken for a filter. Verified against exactly that case.
630
+
631
+ 4xx only. A 5xx is as likely to be the payload *breaking* the application,
632
+ which means it reached something, and reading that as blocked would hide the
633
+ one response saying the sink is live. No vendor list and no block-page
634
+ fingerprints: a status the control did not get is the whole signal.
635
+
636
+ A run where **some** probes got through stays a real `negative` — the sink
637
+ saw those — and the refusals are reported either way.
638
+
639
+ **A refusal never unmakes evidence.** Only a `negative` is replaced, because
640
+ it is the only verdict a refusal contradicts. An application can execute the
641
+ payload and then answer 400 with the output in its body, and the oracle has
642
+ already proven execution from a value random to that probe; overwriting that
643
+ would turn demonstrated RCE into a false negative, which is worse than the
644
+ false clean this verdict removes.
645
+
646
+ **Every method is covered, including the three that decide per probe from a
647
+ series.** `oob`, `lookup` and `deser` take a different branch, which returned
648
+ before refusal was considered — so a callback run whose every probe was
649
+ refused reported `negative` for each of them, because no callback arrived.
650
+
651
+ Refusals are counted in **requests**, like the status tally beside them: an
652
+ aggregate method fires a whole ladder and reports one row, so counting rows
653
+ described the same run with different arithmetic and hid a partly filtered
654
+ series entirely.
655
+
656
+ ### Fixed
657
+
658
+ - **The advice on a filtered run pointed at the wrong thing entirely.** "The
659
+ target may be patched" reads as a clean bill of health for a target that was
660
+ never reached; the blind-sink list names methods a filter refuses in exactly
661
+ the same way; and the second-order line said the target *accepted* an input
662
+ it had in fact rejected with a 403. None of the three fires on a refused run
663
+ now. It is replaced by what was actually observed, with the flags that change
664
+ the payload's shape.
665
+
666
+ ## [2.42.0] — 2026-09-23
667
+
668
+ ### Added
669
+
670
+ - **A run that confirms nothing now names the second-order oracle.** Measured
671
+ against a target that stores on one endpoint and renders through a shell on
672
+ another — a real RCE — every probe read `negative`, and the run answered with
673
+ four methods that are all negative there too, because the execution does not
674
+ happen on the request being measured:
675
+
676
+ ```
677
+ --methods time -> negative=4
678
+ --observe-url -> confirmed, first run
679
+ ```
680
+
681
+ The one flag that works was named nowhere. It is named now, and **not gated
682
+ on which methods have run**: the blind-sink list is, so an operator who had
683
+ already tried the expensive methods — exactly the one with nothing left but
684
+ second order — was told only that the target might be patched.
685
+
686
+ The wording follows what the run observed rather than what it assumes. Input
687
+ returned verbatim means a sink that reflects without executing; input that
688
+ never came back means this response cannot show what became of it; and a run
689
+ of aggregate methods alone, which record no per-probe observation, claims
690
+ neither. A swallowed input is equally a blind sink or a stored one — `ping
691
+ <input> >/dev/null` returns nothing either — so both are named and neither is
692
+ picked.
693
+
694
+ - **Whether the target returned the input is recorded on every probe.** It was
695
+ already computed on the confirmed path, where it becomes "target also
696
+ reflects the payload verbatim"; a negative probe never looked, and the
697
+ negative run is the one that has to say what it saw.
698
+
699
+ Measured against the payload that actually went out. A carrier that
700
+ multiplies through a filter never spells the joined `a*b` out — Liquid sends
701
+ `{{ a | times: b }}` and Django `{% widthratio a 1 b %}` — so an endpoint
702
+ echoing the whole payload recorded a measured `False`, and a target profile
703
+ filtering the `*` shapes leaves only those, at which point the run would
704
+ report "returned none of it" about a target that returned everything.
705
+
706
+ Recorded only when there was a response to look at. A delivery error is not
707
+ an observation, and a `False` for one would put an unmeasured claim exactly
708
+ where the unobserved branch belongs.
709
+
710
+ ## [2.41.0] — 2026-09-23
711
+
712
+ ### Fixed
713
+
714
+ - **A candidate that confirmed RCE was never asked whether it was also a
715
+ deserialization or lookup sink.** The enumeration driver split the methods
716
+ into a cheap wave and an expensive one and skipped the expensive wave once
717
+ execution was proven — reasonable for `time` and `oob`, which would only put
718
+ a second name on one finding, and wrong for `lookup` and `deser`, which
719
+ report *different properties* with their own remediation. They sat on the
720
+ expensive side of a hand-written set of names, so the answer was never asked
721
+ for and never reported.
722
+
723
+ The split now comes from the tier each method declares. Anything reporting
724
+ `confirmed` or `needs-review` is answering *did this target execute my
725
+ input*; anything else is a different question and is never skipped for an
726
+ answer to that one.
727
+
728
+ - **`--max-payloads` is spent per question, not per wave and not per
729
+ candidate.** Per wave it quietly doubled: a run capped at 5 sent 10 probes to
730
+ every candidate that did not confirm, while the cost line printed before any
731
+ traffic said 5. Bounding the candidate instead starves the different
732
+ question — the cheap methods eat the whole allowance and `deser` never runs,
733
+ which is the same finding lost by another route. Every method asking about
734
+ execution now shares one allowance, each different property gets its own, and
735
+ the cost line names how many questions are being asked — and **sums** the
736
+ estimate across them. Counting the cap once while the run grants it per
737
+ question advertised 44 requests for a run that sent 80, which is wrong in
738
+ the direction that matters for the operator bounding a monitored engagement.
739
+
740
+ - **A second-order confirmation settles its carrier too.** With `--observe-url`
741
+ a probe can read negative in the response it drew and `confirmed` on the
742
+ observed channel a moment later. The stop was decided from the pre-poll
743
+ verdict, so the carrier kept probing after it had in fact confirmed —
744
+ spending the budget the stop exists to hand to carriers not yet examined,
745
+ which is the coverage loss this change was written to remove, reappearing on
746
+ the one oracle that needs a second request to answer.
747
+
748
+ ### Changed
749
+
750
+ - **A carrier that has confirmed stops there** (`--confirm-depth first`, the
751
+ new default). One carrier is one method in one environment and context, and
752
+ once it has confirmed every further shape of it can only say the same thing
753
+ again.
754
+
755
+ Measured against an executing target: one candidate spent 115 of its 120
756
+ probes after the first confirmation, and printed 32 confirmations of which 29
757
+ were duplicates inside a single carrier. Those probes were not idle — they
758
+ were spent instead of reaching carriers never examined at all. At the same
759
+ budget the run went from **4 carriers examined to 23**, and from **4
760
+ environments reached to 9**. This buys coverage rather than saving requests.
761
+
762
+ The stop is per carrier and **never** per candidate: a candidate may confirm
763
+ as `unix` while a later `nodejs` carrier is the only thing a different target
764
+ would have shown. `--confirm-depth every` maps every shape a sink accepts,
765
+ which is what writing a proof of concept by hand needs, and the run reports
766
+ how many shapes it held back and which carriers stopped — a fourth tally
767
+ beside the profile drops, the safety holds and the reach notes, because it
768
+ says a fourth thing: the probe could have been sent and had nothing left to
769
+ establish.
770
+
771
+ - **A method declares what one of its probes costs.** `CHEAP_DETECTION_METHODS`
772
+ was a set literal; it is now derived from a `costly` attribute each class sets
773
+ for its own reason — a real sleep, a wait for a callback, a second fetch.
774
+ Every hand-written list naming methods in this repository has gone stale, and
775
+ this one had put `lookup` and `deser` where being skipped cost findings.
776
+
777
+ ## [2.40.0] — 2026-09-22
778
+
779
+ ### Added
780
+
781
+ - **A carrier may take the operands apart, and two engines need it.** Every
782
+ expression carrier until now substituted `__EXPR__` — the joined `a*b` —
783
+ which quietly assumed the engine has an arithmetic operator. Two widely
784
+ deployed ones do not, and both were measured as false negatives:
785
+
786
+ ```
787
+ liquid every bare form missing; {{ 45013 | times: 45989 }} -> 2070102857
788
+ django every bare form missing; {% widthratio 45013 1 45989 %} -> 2070102857
789
+ ```
790
+
791
+ Liquid multiplies with a filter and Django with a tag, so neither form can be
792
+ written as a single expression — an application that really does evaluate the
793
+ template was reported `negative`, which is the same shape as `oob` against a
794
+ `${jndi:...}` sink: probes that reach the target and cannot speak its
795
+ language. A carrier template may now use `__A__` and `__B__` as well as
796
+ `__EXPR__`.
797
+
798
+ Neither payload contains the product, so a target that merely echoes the
799
+ payload still cannot read as `confirmed`. A test pins that for every shipped
800
+ carrier.
801
+
802
+ ### Changed
803
+
804
+ - **A carrier template not parameterised by both operands is skipped, not
805
+ sent.** With no token at all it renders the same constant every probe; with
806
+ only one operand the target is never handed the other, so nothing it can
807
+ compute is the product RCEKit is looking for. Either way the product would
808
+ not be evidence the target computed anything — and a probe that cannot
809
+ confirm still counts toward the coverage a run reports, which is the part
810
+ that matters more than the wasted request.
811
+
812
+ - **`--eval-engines` names every carrier the corpus ships.** Its help listed
813
+ three engines by hand and two were added. An operator narrowing that flag
814
+ reads the list and nothing else, so a stale one says an engine needs no
815
+ carrier when it does, and they cut the only probe that could have confirmed
816
+ it. A test now holds the help text to the corpus.
817
+
818
+ - **What the survey measured and did not ship is recorded too.**
819
+ `eval_carrier_survey` in the corpus now names the engines that need no
820
+ carrier — nunjucks 3.2.4, tornado 6.5.10, mako 1.4.1, chameleon 4.6.0,
821
+ smarty 5.8.4 and Ruby's ERB all return the product from a bare form — and
822
+ those that are
823
+ out of reach. Handlebars 4.7.9 fails every bare form and, being logic-less
824
+ with no built-in arithmetic helper, has no template text that computes a
825
+ product at all. Go `text/template` 1.23 fails every bare form too, and the
826
+ only forms that do return the product — `{{printf "%d" <product>}}` and
827
+ `{{<product>}}` — hand the target the answer, so a target that merely echoed
828
+ them would read as `confirmed`. That is the rule a carrier lives under: **a
829
+ carrier may not carry its own result.**
830
+
831
+ Saying so is worth more than a carrier that cannot work, and it stops the
832
+ next person re-running the same survey — or shipping the Go form, which the
833
+ survey itself produced and which looks like a carrier until you ask what an
834
+ echoing target would return.
835
+
836
+ ## [2.39.0] — 2026-09-22
837
+
838
+ ### Added
839
+
840
+ - **Every part of a multipart body is an injection point.** None of them were
841
+ before. The form branch matched on `=` appearing anywhere in the body, so a
842
+ `multipart/form-data` capture was split on `&` and yielded exactly one
843
+ candidate, named after a `Content-Disposition` line:
844
+
845
+ ```
846
+ form | '--X\r\nContent-Disposition: form-data; name' | body param '...'
847
+ ```
848
+
849
+ Every probe for that point rewrote a *part header*, so it could confirm
850
+ nothing, while `user`, `avatar` and `note` — the fields the form actually
851
+ posts — were never reached. The run still printed a point and a probe count,
852
+ which is the part that matters: coverage reported and not delivered reads
853
+ exactly like a clean target.
854
+
855
+ Several parts may post under one name — a multi-file input and a checkbox
856
+ array both do — so a part is addressed by its **index**, as a JSON leaf is
857
+ addressed by its token path. Addressing by name alone rewrote the first part
858
+ for every candidate, which is the same failure one level further in: three
859
+ files, three points, the first file probed three times and the other two
860
+ never touched.
861
+
862
+ `multipart` is now its own kind, recognised from `Content-Type` and decided
863
+ before the form branch can see the body. A file part is a candidate too — its
864
+ content is the value under test, while its `filename` and `Content-Type` stay
865
+ as captured. Verified end to end against a target that parses with the
866
+ standard library's MIME parser: `multipart field 'note': confirmed`, with the
867
+ non-vulnerable `user` field `negative` beside it.
868
+
869
+ ### Changed
870
+
871
+ - **A GraphQL request is ordered by what can actually confirm.** Its variables
872
+ were already enumerated — they are JSON leaves — but they were tried in
873
+ document order alongside `query` and `operationName`, and those two cannot
874
+ confirm anything. A payload in `query` *replaces* the operation document, so
875
+ the server answers with a parse error before a resolver runs, and
876
+ `operationName` then names an operation that is no longer there. On the
877
+ capture this was measured against they were two points of five, each one a
878
+ full probe ladder.
879
+
880
+ They are moved behind the variables, **not dropped**. A server that logs the
881
+ query document before parsing it is reachable through exactly that field,
882
+ which is the route Log4Shell took through access logs, so a full run still
883
+ tests both and `--max-points` now cuts the least likely to pay first. A plain
884
+ `{"query": ...}` body with no `variables` is left in document order: it is as
885
+ likely to be a search API, and there the query field is the one worth testing.
886
+ Carrying both keys is not enough either — the `query` string has to open like
887
+ a GraphQL document, so `{"query": "red shoes", "variables": {...}}` keeps its
888
+ real injection point where a bounded `--max-points` run will still reach it.
889
+
890
+ ### Fixed
891
+
892
+ - **The verdict table was missing a verdict.** The README said "seven verdicts
893
+ that are never collapsed into each other" and listed seven, while the tool
894
+ reports eight: `lookup-sink` was absent from the one table whose whole job is
895
+ to enumerate them — in a README that uses the word two tables higher, in the
896
+ Log4Shell CVE row, and again in the methods table. The sentence counted the
897
+ rows the table had rather than the verdicts there are, so the omission never
898
+ contradicted itself and nothing failed.
899
+
900
+ Found by reading the README end to end for this change. A test now reads the
901
+ table against `DETECTION_METHODS`, so a verdict a method declares and the
902
+ table does not carry is a failure rather than a silence.
903
+
904
+ - **A multipart body now goes out with the line endings it needs.** `-r`
905
+ normalises the whole request to LF, so the CRLF delimiters RFC 2046 requires
906
+ were gone by the time anything was sent. Bodies rendered for a multipart point
907
+ are re-serialised canonically — probe and payload-free control alike, so the
908
+ two differ in the field under test and in nothing else. Part content is left
909
+ character for character, so a lone newline inside an uploaded text file
910
+ survives.
911
+
912
+ ## [2.38.0] — 2026-09-21
913
+
914
+ ### Added
915
+
916
+ - **A probe shape may reach past the run's tier and be sent anyway**, when its
917
+ effect is one the run undoes by saying it happened. `reaches_past` is that
918
+ declaration, beside `safety`, which stays for an effect a notice cannot take
919
+ back.
920
+
921
+ Reach wins where the two pull against each other. Detection the tool could
922
+ have done and did not is a false negative wearing a safety label, and it
923
+ costs more than the noise it saves.
924
+
925
+ `deser`'s DNS gadget is the first of these, and closes an inconsistency
926
+ recorded a version ago: it makes the target resolve a name -- the very thing
927
+ `oob` and `lookup` are refused for at `safe` -- while its only gate was
928
+ `--oob-host`. Holding it back would have sent fewer probes at the default
929
+ tier. It goes, and the run reports how far it reached.
930
+
931
+ Three tallies now, because they say three different things and one number
932
+ would state the wrong one about all of them: the profile dropped it (it could
933
+ not have reached the sink), the tier held it (it could, and was not sent), or
934
+ it reached past the tier (it was sent, further than asked).
935
+
936
+ `lookup`'s `ldap://` and `rmi://` are the edge the rule has, and stay at
937
+ `stateful`: a class fetched from an address RCEKit did not choose is not
938
+ something a notice takes back.
939
+
940
+ - **Two worked examples in the README's Quick start**, because the fullest run
941
+ RCEKit can make was not shown anywhere near the front.
942
+
943
+ The first is a captured request: most sinks worth testing sit behind a POST
944
+ with a session cookie, a content type and a body, and `--verify-url` carries
945
+ none of that. The second is everything the tool has -- injection-point
946
+ enumeration across every value in that request, every method, callbacks, the
947
+ top rung -- with a table of what each flag opens up and the cost line that
948
+ prints before it fires.
949
+
950
+ Both were run before being written down. The point the Quick start never
951
+ made: `--auto-params` needs `-r`, so the fullest run is not reachable from a
952
+ URL at all, which is worth knowing before concluding a target is clean.
953
+
954
+ ## [2.37.0] — 2026-09-20
955
+
956
+ ### Added
957
+
958
+ - **A detection method declares the risk rung it needs**, and the engine reads
959
+ it. `SAFETY_ORDER` has labelled corpus payloads `safe` / `intrusive` /
960
+ `stateful` from the start and the query-language bridges followed; detection
961
+ methods did not. Each risky one was gated by a hand-written branch in
962
+ `main()` naming it, so a new method meant remembering to add another -- and a
963
+ probe shape with nowhere to declare its rung was deleted rather than gated.
964
+
965
+ | method | rung | |
966
+ |---|---|---|
967
+ | `reflected`, `eval`, `time`, `deser` | `safe` | compute, delay, or parse |
968
+ | `oob`, `lookup` | `intrusive` | makes the target open outbound connections |
969
+ | `file`, `write` | `stateful` | writes to the target |
970
+
971
+ `file` and `write` are gated by their own configuration rather than by the
972
+ rung: neither does anything until a directory to write into and a URL to read
973
+ it back from are named, which says more than a tier would, and asking for the
974
+ flag as well would refuse a command that works today. Nothing that runs today
975
+ stops running.
976
+
977
+ - **A probe shape may need a higher rung than its method**, so coverage that
978
+ only makes sense at the top tier has somewhere to live instead of being
979
+ deleted. **`--methods lookup` sends `ldap://` and `rmi://` again**, at
980
+ `stateful`.
981
+
982
+ They were removed in 2.36.0 on the argument that `dns://` resolves wherever
983
+ `ldap://` would. That claim was too strong -- a filter catching the string
984
+ `dns:` and not `ldap:`, or a trimmed runtime without the DNS provider,
985
+ defeats it -- and a sink that takes one scheme and not the other is exactly
986
+ the sink this method is for. The reason they were removed was real: they
987
+ continue *past* resolution and connect to whatever address the answer named,
988
+ which is not an address RCEKit chose. That is a rung, not a reason to drop
989
+ coverage.
990
+
991
+ - **The run says what the rung held back**, counted apart from the target
992
+ profile's drops and with the flag that would send it. The two say different
993
+ things: a profile drop means the probe *could not have* reached the sink,
994
+ while this means it could and the operator chose not to send it. Reporting
995
+ them together would state the first about the second.
996
+
997
+ - **A method declares the weaker tiers it really reports**, not only its
998
+ ceiling. `write` reports `needs-review` for a file that is served but not
999
+ interpreted and `deser` for a shape fingerprint, and three separate places
1000
+ had to know that -- the documentation tests, the benchmark's expectation
1001
+ whitelist, and the advice printed after a clean in-band run. Each kept its
1002
+ own answer; the docs test exempted `needs-review` for *every* method, so
1003
+ `lookup` could name a verdict it never emits and pass.
1004
+
1005
+ ### Changed
1006
+
1007
+ - **`--methods file` and `--methods write` with nothing configured now say
1008
+ so by name.** They were simply not applicable before, so the run built no
1009
+ probes and reported `nothing-tested` -- which is the quietest way this tool
1010
+ can fail and reads much like a clean target.
1011
+
1012
+ - **`docs/reference.md` carries a `Rung` column**, held against the class by a
1013
+ test. The tier column already was; this is the same claim one column over.
1014
+
1015
+ ### Fixed
1016
+
1017
+ - **A config-gated method is not re-gated by the rung at runtime.** The
1018
+ pre-flight lets `--methods file --webroot ... --web-base-url ...` through
1019
+ because the configuration is the gate, but the probe filter read the run's
1020
+ default `safe` ceiling and held every probe inheriting the method's
1021
+ `stateful` rung. The CLI accepted a documented invocation and then reported
1022
+ `nothing-tested` -- the quietest way this tool can fail, and the thing the
1023
+ rung work was supposed to remove rather than add.
1024
+
1025
+ Nothing caught it because every `file` test builds the method's config
1026
+ directly, without `max_safety`, so the ceiling fell back to the method's own
1027
+ rung and the probes went out. A run through the CLI with the channel
1028
+ configured is the one thing that would have, and there is one now.
1029
+
1030
+ - **The cost estimate applies the risk tier as well as the target profile.** It
1031
+ counted every shape a method built, so once a rung could narrow a method the
1032
+ pre-flight figure over-counted -- three times over for `lookup` at the
1033
+ default tier, which is exactly the operator who narrowed the run on purpose.
1034
+ Both paths share one predicate now, and it counts nothing, so an estimate
1035
+ never moves the numbers the report prints.
1036
+
1037
+ ## [2.36.0] — 2026-09-19
1038
+
1039
+ ### Added
1040
+
1041
+ - **`--methods lookup`: confirmation for an expression-lookup sink**, the shape
1042
+ Log4Shell has, where the sink resolves a URI instead of running a command.
1043
+
1044
+ `oob` could not reach one. Its `applicable` does admit `java` -- a Java
1045
+ application can shell out, so the environment is genuinely shell-capable --
1046
+ but every probe it builds is a shell command: `nslookup`, `curl`, `certutil`,
1047
+ `iwr`. A sink that interpolates `${jndi:...}` runs none of them, so the method
1048
+ applied, sent its whole ladder, and came back `negative` on a target that is
1049
+ exploitable. The README's Log4Shell row rested on the standalone listener and
1050
+ a generated payload file, which produce no verdict row at all, so nothing in
1051
+ the engine could reproduce that claim.
1052
+
1053
+ The probes are lookups and nothing else, and they depend on the injection
1054
+ context rather than the environment, exactly as `eval`'s do. The oracle is the
1055
+ one `oob` already uses: a token the target could only have learned by
1056
+ resolving what it was handed. The expression resolves `<token>.<host>`, so
1057
+ the in-process DNS listener is the entire apparatus -- no LDAP or RMI server
1058
+ is needed, and none is started.
1059
+
1060
+ **It reports `lookup-sink`, never `confirmed`.** A callback proves the sink
1061
+ resolved a URI RCEKit chose -- that it evaluated the expression it was handed.
1062
+ It does not prove the target ran attacker code: Log4Shell becomes RCE when the
1063
+ LDAP server answers with a loadable class. So the method gets its own
1064
+ proven-sink tier beside `deserialization-sink`, with its own section in the
1065
+ report, and `confirmed` keeps meaning executed.
1066
+
1067
+ **Only `jndi:dns://` is sent, and that is the security property rather than a
1068
+ shortcut.** A name lookup can be nothing else. `ldap://` and `rmi://` continue
1069
+ *past* resolution and open a connection to whatever address the answer named
1070
+ -- by default `127.0.0.1`, which is the target's own loopback. Whatever
1071
+ replies on :389 or :1099 is not RCEKit, so a reference could come back and a
1072
+ class be instantiated: the tool would have crossed the line this method exists
1073
+ to stop short of, having promised it had not. Dropping the two schemes also
1074
+ costs no coverage -- `DnsContextFactory` ships in the JDK, so `dns://`
1075
+ resolves wherever `ldap://` would, and on 2.15.0 it still resolves where
1076
+ `ldap://` no longer does. The proof is the callback and the finding is "this
1077
+ sink resolved a URI I chose" -- the same line `deser` draws, drawn here before
1078
+ it can be crossed.
1079
+
1080
+ Behind the same two gates as `oob`, and sharing its listener: it needs
1081
+ `--oob-host`, and it is held back at the default safety tier because it makes
1082
+ the target open outbound connections. Without `--oob-host` it builds no probes
1083
+ at all, which the engine reports as `nothing-tested` -- never `negative`. An
1084
+ address literal is a literal whichever family it is from: `::1` and `[::1]`
1085
+ build nothing, the same as `10.0.0.1`, because a token can only ride in a DNS
1086
+ label. `blind_sink_advice` names the method as proving a lookup sink and not
1087
+ execution, in a list where `oob` and `file` mean confirmed execution, and the
1088
+ README's Log4Shell demo heading says `lookup-sink` rather than `confirmed` --
1089
+ it was left claiming execution beside the table row that no longer does.
1090
+
1091
+ ## [2.35.5] — 2026-09-19
1092
+
1093
+ ### Fixed
1094
+
1095
+ - **The `--insecure` notice describes the run it is in.** It was printed as
1096
+ soon as the flag was seen, so a generation-only run, `--doctor`, or a plain
1097
+ HTTP target all announced a TLS downgrade that never happened -- and it named
1098
+ both downgrade rungs even on an OpenSSL build that had refused one of them.
1099
+ Reporting a downgrade on a run that opened no TLS connection is the same
1100
+ defect as reporting a probe that was never sent, in the one line written to
1101
+ be an audit of the run.
1102
+
1103
+ Building the context and saying so are now separate questions. The context is
1104
+ built whatever the target's scheme is -- urllib follows a redirect with the
1105
+ handler it was given, so an `http://` target that lands on a self-signed or
1106
+ legacy `https://` one needs it as much as a direct HTTPS target does. The
1107
+ notice fires once, when a connection has actually reached TLS, and lists only
1108
+ the rungs that took.
1109
+
1110
+ - **A benchmark case may set `timeout` to zero.** `timeout or 900.0` replaced
1111
+ an explicit `0` with the fifteen-minute default, so a case deliberately
1112
+ bounded to no time at all ran for much longer than it asked for.
1113
+
1114
+ ## [2.35.4] — 2026-09-19
1115
+
1116
+ ### Fixed
1117
+
1118
+ - **A wave the declared profile emptied is no longer read as a finished
1119
+ method.** The probe filter drops probes between the method and the wire, and
1120
+ when it emptied an adaptive method's *first* screening wave the engine took
1121
+ the empty batch for "nothing left to send" and stopped before asking for the
1122
+ next wave at all.
1123
+
1124
+ An adaptive method holds separators back in waves precisely because a filter
1125
+ is expected: `time` screens two of them first and keeps `||`, `&&`, the
1126
+ newline and the bare command for a second wave. A sink that strips `;` and
1127
+ `|` removes exactly the first wave and leaves the rest intact -- and the rest
1128
+ were never sent. Measured against a sink reachable only through `&&`, the run
1129
+ reported **`negative`**: "the probes reached the target and found nothing",
1130
+ about probes that were never sent. Worse than `nothing-tested`, which is at
1131
+ least true.
1132
+
1133
+ The loop now tracks what the method offered separately from what the profile
1134
+ allows to be sent, and ends only when the method itself is done. The round cap
1135
+ still bounds it, so a method whose every wave is filtered cannot spin the
1136
+ engine.
1137
+
1138
+ - **The unit suite runs off Linux.** The fake vulnerable sinks in the tests are
1139
+ POSIX command-injection points and every probe built for them is POSIX, but
1140
+ they were executed through `os.popen` — which is `cmd.exe` on Windows. The
1141
+ sink the test says executes did not execute, so the oracle correctly reported
1142
+ no execution and five tests failed for a reason unrelated to the code under
1143
+ test. Worse, the tests asserting a *negative* stayed green throughout: a
1144
+ broken fixture that keeps its controls passing is the failure this project
1145
+ takes seriously everywhere else. The sinks now name a POSIX `sh` explicitly;
1146
+ on Linux and macOS that is the `/bin/sh` they always used.
1147
+
1148
+ - **A deliberately dead target no longer costs minutes.** A closed loopback port
1149
+ answers with a RST on Linux and is silently dropped on Windows, where each
1150
+ probe waits out the SYN retry instead — measured at ~2s per probe. The
1151
+ benchmark harness's own unreachable-target case paid that for the full ladder
1152
+ twice, once for the vulnerable half and once for the control: 1800s for one
1153
+ test, longer than the other 534 together. Bounded to three probes it measures
1154
+ 9.1s, and three probes prove "nothing reached the target" exactly as well as
1155
+ forty do.
1156
+
1157
+ - **A captured-request fixture reaches disk byte for byte.** The cleartext-capture
1158
+ tests write a raw HTTP request whose text already spells its own CRLF line
1159
+ endings, through `Path.write_text` — which on Windows translates the newline of
1160
+ each one again. The file on disk held a doubled carriage return, the parser
1161
+ found no headers, and two tests failed against a fixture that had stopped being
1162
+ an HTTP request at all.
1163
+
1164
+ The third test in that class passed throughout, for the wrong reason: it
1165
+ asserts that a notice is *absent*, and a request that cannot be built prints no
1166
+ notice either. A control that stays green while its fixture rots is precisely
1167
+ what this project refuses to accept from a benchmark case, so the fixture now
1168
+ has a guard of its own. The capture is written as bytes; the `newline` argument
1169
+ that would say the same thing arrived in Python 3.10, and this project
1170
+ supports 3.8.
1171
+
1172
+ ## [2.35.3] — 2026-09-18
1173
+
1174
+ The target profile an operator declares now reaches the probe ladder, not just
1175
+ the corpus.
1176
+
1177
+ ### Fixed
1178
+
1179
+ - **`--deny-chars` / `--max-length` reach the detection probes.** They were
1180
+ applied by `_filter_by_profile`, which drops *corpus records* — and stopped
1181
+ there. The probes a detection method builds from those records went out
1182
+ regardless, so a run that had been told "this target strips quotes" still paid
1183
+ for every quote-carrying rung of the ladder, on requests structurally unable
1184
+ to confirm. Those requests are not free: on a captured request with `-p all`
1185
+ they are the budget the next injection point never got. The filter now sits at
1186
+ the engine, where every probe passes through it — deliberately not inside
1187
+ `_wrap_variants`, because `_space_free_probes`, the query-language bridges,
1188
+ `eval`, `oob` and `deser` each build payloads without going through that
1189
+ helper, and a gate that reaches some methods and not others is the side path
1190
+ that once left `file`/`time`/`oob` unable to send the raw rung.
1191
+
1192
+ Denying a character narrows the ladder rather than emptying it: a target that
1193
+ strips `;` is still probed through `|`, `||`, `&&` and the newline, which is
1194
+ what the separator table has always been for.
1195
+
1196
+ Checked on the literal payload, before the delivery layer percent-encodes it
1197
+ for its injection point. That is stricter than the corpus check, which is
1198
+ applied to the encoded payload and so lets a URL-encoded quote through a quote
1199
+ filter. The layers genuinely differ: transport encoding is undone by the
1200
+ server before the value reaches the sink, so a percent-encoded quote is still
1201
+ a quote when the application's own filter sees it.
1202
+
1203
+ - **A profile strict enough to remove every probe reports `nothing-tested`.**
1204
+ Not `negative`, which would read as "not vulnerable" from a run that sent
1205
+ nothing. The message names the profile as the cause and the characters a probe
1206
+ would have to avoid, instead of the generic advice to widen `--environments` —
1207
+ which is not what emptied the run.
1208
+
1209
+ - **The cost estimate follows the profile.** `[detect] cost:` builds the probes
1210
+ and counts them, so it now counts the ones that will actually be sent. An
1211
+ estimate that ignores a filter is wrong precisely for the operator who
1212
+ narrowed the run on purpose.
1213
+
1214
+ ### Changed
1215
+
1216
+ - A run that dropped probes says so, with the reason and a count per reason. A
1217
+ ladder that shrinks quietly is the one way this filter could manufacture a
1218
+ false negative, so the removal is stated rather than left to be inferred from
1219
+ the traffic.
1220
+
1221
+ ## [2.35.2] — 2026-09-18
1222
+
1223
+ ### Fixed
1224
+
1225
+ - **`--insecure` now reaches a legacy TLS stack, not just an untrusted one.**
1226
+ Turning certificate verification off is not the same as completing a
1227
+ handshake. OpenSSL 3.x ships security level 2, which refuses the key sizes and
1228
+ signature algorithms that software of the era this tool gets pointed at still
1229
+ offers — Webmin 1.910, the build the README's `reflected` row rests on,
1230
+ answers a default client with `SSLV3_ALERT_HANDSHAKE_FAILURE` and nothing
1231
+ else. Every probe then came back `error`: correct, and useless. The run was
1232
+ honest about having measured nothing, and the sink behind that handshake was
1233
+ never tested at all. Found by running the coverage benchmark, which failed at
1234
+ its readiness gate against a container that was up and answering.
1235
+
1236
+ `--insecure` now also lowers the security level and the minimum protocol
1237
+ version. This does not widen exposure: with `check_hostname = False` and
1238
+ `CERT_NONE` the connection is already unauthenticated, so an active attacker
1239
+ is already unconstrained — accepting a 1024-bit key or a SHA-1 signature on
1240
+ top of that gives away nothing that was still being held. What it buys is the
1241
+ difference between testing the target and reporting that it could not be
1242
+ reached. A run without the flag is untouched and still verifies certificates.
1243
+
1244
+ - **The benchmark's readiness gate is as permissive as the tool it gates.**
1245
+ `wait_for_target` built its own strict context, so a case against deliberately
1246
+ old software reported "target never became ready" about a container that was
1247
+ up — a case failure with nothing wrong in it.
1248
+
1249
+ ### Changed
1250
+
1251
+ - A run that passes `--insecure` states the full extent of the downgrade on its
1252
+ first line. The flag gives up more than certificate identity now, and an
1253
+ operator on a monitored engagement should read that in the transcript rather
1254
+ than infer it from the help text.
1255
+
1256
+ ## [2.35.1] — 2026-08-21
1257
+
1258
+ A robustness pass over error handling: no new capability, four ways the tool
1259
+ could crash or mislead on input it did not choose.
1260
+
1261
+ ### Fixed
1262
+
1263
+ - **A truncated error response no longer ends the run.** Reading an
1264
+ `HTTPError`'s body happens *inside* the `except` handler, where the sibling
1265
+ `except Exception` cannot reach it — so a target that promised a
1266
+ `Content-Length` it never delivered raised `ConnectionResetError` straight out
1267
+ of `main()`, taking every probe already fired with it. The read is now
1268
+ guarded: the status still comes back, an unreadable body is reported empty.
1269
+ A body that *does* arrive is still returned in full — the 500-stack-trace
1270
+ confirmations that branch exists for are unaffected.
1271
+ - **A `--target-profile` is checked before it is used.** A profile is written by
1272
+ hand, so a typo in one is ordinary; it surfaced as a traceback. A top level
1273
+ that is not a JSON object, `deny_chars` that is not text, a `max_length` that
1274
+ is not a number, a selector field that is not a list of names — each is now an
1275
+ operator-readable `[!]` message and exit 1, the same way the sink-shape fields
1276
+ already behaved. Twelve inputs that produced an `AttributeError` or a
1277
+ `TypeError` now produce a sentence.
1278
+ - **A selector field given as one string means one name.** `"environments":
1279
+ "unix"` in a profile was iterated character by character, matched nothing, and
1280
+ the empty run that followed was reported as a success. A string is now split
1281
+ on commas: `"unix"` is `["unix"]`, `"raw, html"` is `["raw", "html"]`.
1282
+ - **Unknown `--environments` and `--encodings` are named.** Both were silent, so
1283
+ `--environments linux` — the corpus calls it `unix` — produced an empty file
1284
+ and exit 0, indistinguishable from a target with no payloads for it. Both now
1285
+ warn and list the names that exist, as unknown contexts and categories already
1286
+ did. An empty result is also no longer announced as "Successfully generated 0
1287
+ payloads"; it says the selection matched nothing and points at the filters.
1288
+ - **A callback cannot rewrite the operator's terminal.** The host and path of an
1289
+ OOB callback are chosen by the target. Printed raw, an ESC byte let that
1290
+ target colour, erase and rewrite lines — hiding a genuine `[HIT]` behind
1291
+ `\x1b[2K\r`, or forging one that never arrived. Control characters are now
1292
+ escaped for display as `\xNN`; the recorded hit and the `--listen-log` JSONL
1293
+ keep the bytes verbatim.
1294
+
1295
+ ## [2.35.0] — 2026-08-20
1296
+
1297
+ ### Added
1298
+
1299
+ - **RCEKit is installable from PyPI**: `pipx install rcekit` (or
1300
+ `pip install rcekit`) puts an `rcekit` command on PATH. Published through PyPI
1301
+ Trusted Publishing from a GitHub release — no API token, no repository secret.
1302
+
1303
+ It ships as a **single-module distribution**, not a package tree. `rcekit.py`
1304
+ stays one file at the repo root and still runs alone from a `curl` on a jump
1305
+ box or an air-gapped host; installing is a second supported shape, not a
1306
+ replacement for the first.
1307
+
1308
+ ### Changed
1309
+
1310
+ - **`import rcekit` no longer has side effects.** Logging was configured at
1311
+ module scope, and `logging.FileHandler` opens its file when it is constructed,
1312
+ so merely importing the module wrote `rcekit.log` into whatever directory the
1313
+ interpreter happened to be in. Handler setup moved into `configure_logging()`,
1314
+ called from `main()`. Running the CLI still writes `rcekit.log` exactly as
1315
+ before.
1316
+
1317
+ - **`main()` takes an optional `argv`** and returns an explicit `int`, so the
1318
+ console-script entry point is a plain zero-argument call and tests can drive
1319
+ the CLI in-process. Every exit code is unchanged.
1320
+
1321
+ - **The built-in corpus is no longer reported as a missing file.** With nothing
1322
+ but `rcekit.py` — an installed wheel, or the single-file copy — there is no
1323
+ `templates/` directory, the embedded corpus *is* the corpus, and the run is
1324
+ now silent about it; `--doctor` names it `built-in (embedded in rcekit.py)`
1325
+ and reports OK. A `templates/` directory that exists *without* its
1326
+ `payloads.json` still prints the notice, because that one is a real finding.
1327
+ A corpus that is present but corrupt, and an explicit `--template-file` that
1328
+ is missing or corrupt, still refuse to run and exit non-zero.
1329
+
1330
+ - The log file handler now runs at `DEBUG` while the console stays at `INFO`, so
1331
+ detail worth having when reconstructing a run no longer lands on the
1332
+ operator's terminal.
1333
+
1334
+ ## [2.34.1] — 2026-08-17
1335
+
1336
+ Documentation only; no behaviour change.
1337
+
1338
+ ### Changed
1339
+
1340
+ - **New tagline: "`confirmed` means the target executed the input. `negative`
1341
+ means the probes reached it."** The old one — "prove RCE, don't guess it" —
1342
+ claimed the tool always proves. It does not, and does not need to: the value
1343
+ is that each verdict has a mechanical meaning, in both directions. A promise
1344
+ can be broken; a definition cannot.
1345
+
1346
+ Both halves name their actor and object on purpose. "Executed" alone reads as
1347
+ though *RCEKit* executed something; the claim is about the target. And
1348
+ `confirmed`/`negative` are the verdict values as the code spells them, not
1349
+ looser words like "clean".
1350
+
1351
+ - **The README caught up with the engine.** It still described "two verdict
1352
+ tiers" when there are seven, and the comparison table predated the last five
1353
+ releases. Rewritten around what a verdict asserts, with the tier table as the
1354
+ centrepiece and a section on the half no other tool has: `error` and
1355
+ `nothing-tested` exist so a run that tested nothing is never reported as
1356
+ clean.
1357
+
1358
+ The comparison table gains the classes added since it was written — Windows
1359
+ `cmd.exe`/PowerShell sinks, upload → write-then-execute, second-order
1360
+ execution, query-language bridges and deserialization sinks — and the "reach
1361
+ for something else" note now says plainly that sqlmap owns the database and
1362
+ RCEKit's bridges only prove the OS is reachable from a text parameter.
1363
+
1364
+ ### Fixed
1365
+
1366
+ - Two claims in the engagement-controls table were wrong and are now accurate:
1367
+ the observed-channel fetch sends **no** credentials unless given a request with
1368
+ `--observe-request` (only the `file` read-back inherits them, same-origin), and
1369
+ an unanswered `--observe-url` is a warning about a partly blinded run rather
1370
+ than a `nothing-tested` verdict.
1371
+ - The "mechanisms that produce `inconclusive`" list said four and listed five.
1372
+
1373
+ ## [2.34.0] — 2026-08-17
1374
+
1375
+ Deserialization **sink** detection, and a verdict that is deliberately not RCE.
1376
+
1377
+ Deserialization RCE (fastjson, shiro, weblogic, jenkins) cannot be confirmed by
1378
+ the value-oracle model: the payload is a serialized object graph and gadget
1379
+ selection is classpath-specific, so whether execution is reachable depends on
1380
+ jars RCEKit cannot see. That stays out of scope. The honest middle step is
1381
+ showing the endpoint parses the data at all — a real finding, and the
1382
+ prerequisite for every gadget chain.
1383
+
1384
+ ### Added
1385
+
1386
+ - **`--methods deser`**, which **never emits `confirmed`**. Its strongest
1387
+ outcome is a new verdict, `deserialization-sink`, reported in its own section
1388
+ that states outright that reaching RCE from there depends on classpath
1389
+ gadgets. Collapsing it into `confirmed` would break the one guarantee the tool
1390
+ rests on; collapsing it into `needs-review` would throw away a proven finding.
1391
+
1392
+ `deserialization-sink` sits below both RCE tiers in the collapsed verdict: it
1393
+ is proven, but a *suspected* RCE outranks a proven non-RCE in triage.
1394
+
1395
+ - **Two oracles of deliberately different strength.** `shape` (no listener
1396
+ needed) sends a well-formed object stream, the same stream truncated, and the
1397
+ format's magic bytes plus random noise of the same length, and asks whether
1398
+ the endpoint answers the well-formed one differently from both — a
1399
+ fingerprint, so `needs-review` only, never promoted. `dns` (needs
1400
+ `--oob-host`) sends a gadget whose only side effect is a name lookup.
1401
+
1402
+ - **A URLDNS builder for Java serialization.** A `HashMap` holding one
1403
+ `java.net.URL`: `HashMap.readObject` hashes the key, `URL.hashCode` asks for
1404
+ the host address, the JVM resolves the name. It references no class outside
1405
+ `java.util`/`java.net`, so there is nothing in it that can run — the callback
1406
+ proves the object graph was reconstructed and no more.
1407
+
1408
+ Built in Python rather than declared in the corpus because the URL host is
1409
+ length-prefixed *inside* the stream and changes per probe. Its constant parts
1410
+ are the exact bytes OpenJDK's own `ObjectOutputStream` emits for that graph,
1411
+ and the result was **verified against OpenJDK 21**: it deserializes to
1412
+ `HashMap{http://<host>/=rk}` and issues a DNS query for `<host>`, with no code
1413
+ execution.
1414
+
1415
+ - **A `deser_probes` corpus section** with `java`, `php`, `dotnet`,
1416
+ `python_pickle` and `fastjson`, plus `--deser-formats` to narrow it. Only
1417
+ `java` and `fastjson` carry a DNS gadget: PHP and .NET chains all run through
1418
+ magic methods or type confusion, so there is no honest DNS-only probe for them
1419
+ and they get the shape oracle alone.
1420
+
1421
+ - Response signatures for the shape differential drop long digit and hex runs,
1422
+ so request ids and timestamps on an otherwise identical error page do not make
1423
+ every endpoint fingerprint as a parser.
1424
+
1425
+ ### Changed
1426
+
1427
+ - The README's scope note now says precisely what changed and what did not:
1428
+ deserialization **gadget chains** remain out of scope, while the **sink** is
1429
+ now reported in its own tier.
1430
+
1431
+ ## [2.33.0] — 2026-08-17
1432
+
1433
+ Query-language bridges. Several RCEs pass through a query language before
1434
+ reaching the OS — Postgres `COPY … FROM PROGRAM`, MSSQL `xp_cmdshell`, XXE
1435
+ `expect://` — and the injection point is an ordinary text value, so the oracle
1436
+ model already fitted. Only the carriers were missing.
1437
+
1438
+ ### Added
1439
+
1440
+ - **A `bridges` section in the corpus**, declared like `eval_carriers` so
1441
+ coverage grows without touching Python. Each bridge names the shell it
1442
+ reaches, its safety tier, its prerequisites and, where it creates something,
1443
+ the statement that removes it: `postgres_copy_program` (`/bin/sh`,
1444
+ `stateful`), `mssql_xp_cmdshell` (`cmd.exe`, `intrusive`) and `xxe_expect`
1445
+ (`/bin/sh`, `intrusive`).
1446
+
1447
+ - **`--bridges none|auto|NAMES`** rides the command probes through them. A
1448
+ bridge is a **carrier, not an oracle**: it wraps the command `reflected`,
1449
+ `time` and `oob` already build, so those methods prove execution through it
1450
+ and inherit every tier guarantee rather than re-deriving one. Off by default,
1451
+ because a bridge payload is SQL or XML syntax and on an ordinary shell sink it
1452
+ is a request that cannot confirm.
1453
+
1454
+ Three properties follow from that framing. A bridge only gets a core written
1455
+ in its own dialect — `xp_cmdshell` hands its argument to `cmd.exe`, so pairing
1456
+ it with a POSIX `$((a+b))` would send inert text. No separator is prepended:
1457
+ inside `COPY … FROM PROGRAM '…'` there is no running command to break out of.
1458
+ And the record's context still applies, so `--contexts sql` and a bridge
1459
+ compose instead of each reinventing the other.
1460
+
1461
+ - **The safety ordering governs bridges** exactly as it governs every corpus
1462
+ payload: a `stateful` bridge needs `--verify-active-risk stateful`, and the
1463
+ pre-flight names the tier each held-back bridge actually requires rather than
1464
+ sending the operator to raise the ceiling further than the run needs.
1465
+
1466
+ ### Changed
1467
+
1468
+ - **An aggregate method's result now carries its cleanup line.** `time` reports
1469
+ one row for a whole probe series, so a stateful bridge on the one oracle that
1470
+ reliably proves a query-language sink was the one that never said how to clean
1471
+ up after itself.
1472
+
1473
+ ### Not built, deliberately
1474
+
1475
+ - **MySQL UDF execution** is a multi-stage chain — write a shared object into the
1476
+ plugin directory, then `CREATE FUNCTION` — not something a single probe can
1477
+ carry. There is no stub for it.
1478
+ - **MongoDB `$where`** is a boolean-only channel (its JS sandbox cannot reach a
1479
+ shell), so it needs a different oracle rather than this one.
1480
+ `mongo-express/CVE-2019-10758` is a plain JS `eval` sink that `--methods eval`
1481
+ already covers.
1482
+
1483
+ The three shipped bridges are documented syntax but **not validated against live
1484
+ databases here** — this build environment has no container runtime. Each corpus
1485
+ entry says so in its `verified` field rather than implying a test that did not
1486
+ happen.
1487
+
1488
+ ## [2.32.0] — 2026-08-17
1489
+
1490
+ The second-order oracle. Execution frequently happens on a **different request**
1491
+ than injection — stored SSTI rendered on a profile page, a payload written to a
1492
+ log a template engine later renders, a queued job run asynchronously. The engine
1493
+ diffs the response it injected into, so every one of those read `negative`
1494
+ however exploitable the target was.
1495
+
1496
+ ### Added
1497
+
1498
+ - **`--observe-url URL`** names the endpoint where the execution surfaces. It is
1499
+ read after each probe and then polled after the batch, and a probe whose
1500
+ computed value turns up there is upgraded to `confirmed`.
1501
+
1502
+ It stays fully differential, which is why it reaches `confirmed` rather than
1503
+ `needs-review`: the value was computed locally from operands random to that
1504
+ probe, it must be absent from a snapshot of the endpoint taken **before any
1505
+ probe was sent**, and — the rule that carries the weight — a probe's value is
1506
+ looked for there **only when the probe's own payload does not contain it**.
1507
+
1508
+ Without that last rule the oracle would be a false-positive generator: `file`
1509
+ and `oob` expect a random token that sits verbatim in the payload, so a target
1510
+ that merely stores the payload and renders it back would hand that token
1511
+ straight to the observed page and every such probe would confirm without
1512
+ executing anything. Measured against a store-and-echo target: **0**
1513
+ confirmations. The computed-value methods pass the same rule for the opposite
1514
+ reason — reflection returns `$((a+b))`, never the sum — so it selects them
1515
+ without naming them, and a method added later inherits the right answer.
1516
+
1517
+ - **`--observe-request FILE`** takes a captured request instead, for the common
1518
+ case where the page a stored payload renders on is behind a login. It needs no
1519
+ `FUZZ` marker: the observed endpoint is read, never injected into.
1520
+
1521
+ - **`--observe-poll` / `--observe-timeout`** control the polling window
1522
+ (defaults 5s and 60s). One poll always happens, even at a zero timeout.
1523
+
1524
+ - Every probe result carries an `observe_status` in `--detect-json`:
1525
+ `confirmed`, `polled` (read, value not there), `in-control`, `not-observed`
1526
+ (not eligible) or `unreachable`. When the endpoint never answered, the run
1527
+ says so outright — negatives decided without ever reading the observed channel
1528
+ are not second-order negatives.
1529
+
1530
+ ### Changed
1531
+
1532
+ - The observed channel is read once after **each** probe as well as polled after
1533
+ the batch, so a run with `--observe-url` sends roughly twice the requests.
1534
+ Batch-then-poll alone is only correct for a channel that *accumulates* (a log,
1535
+ a comment list); where the store overwrites — a profile field, which is the
1536
+ shape this oracle most exists for — every probe but the last is gone by the
1537
+ time the batch poll runs, and the oracle confirmed nothing. The extra read is
1538
+ skipped for probes that are already confirmed in-band or not eligible, so
1539
+ `file` and `oob` add none.
1540
+
1541
+ Observing is additive throughout: the in-band verdict is computed exactly as
1542
+ before and only a non-`confirmed` one can be upgraded, so a run without the flag
1543
+ is byte-for-byte unchanged and a run with it can only gain findings.
1544
+
1545
+ ## [2.31.0] — 2026-08-17
1546
+
1547
+ The `write` method: a write primitive proven to be RCE by executing what it
1548
+ wrote. A whole family of targets was invisible — `tomcat/CVE-2017-12615` (PUT a
1549
+ JSP), `activemq/CVE-2016-3088`, `weblogic/CVE-2018-2894` — because the vulnerable
1550
+ request *stores a file* rather than evaluating anything. Nothing is computed in
1551
+ its response, so `reflected` and `eval` correctly returned `negative` on targets
1552
+ that are fully exploitable.
1553
+
1554
+ ### Added
1555
+
1556
+ - **`--methods write`** — the inverse of `file`. `file` assumes execution exists
1557
+ and uses a write as proof of it; `write` assumes a write primitive exists and
1558
+ uses execution of the written file as proof of RCE. The probe is the file's
1559
+ *content*: a one-liner computing a product on random operands, delivered
1560
+ through the ordinary injection point.
1561
+
1562
+ The fetched file is read in three tiers, and the middle one is the reason the
1563
+ method exists:
1564
+
1565
+ | fetched file contains | verdict | means |
1566
+ |---|---|---|
1567
+ | the product | `confirmed` | written **and** executed |
1568
+ | the one-liner, verbatim | `needs-review` | arbitrary file write, not interpreted |
1569
+ | neither | `negative` | no write, or not served there |
1570
+
1571
+ An upload directory that is served but not interpreted is a real finding and
1572
+ is not remote code execution, so the tiers are never merged.
1573
+
1574
+ - **`--write-url-template URL`** names where the stored file is served — the
1575
+ channel the proof comes back on, and the flag the method is gated on.
1576
+
1577
+ - **`--write-lang`** picks the file types: `auto` (default) reads the extension
1578
+ off the read-back URL, or name any of `jsp`, `jspx`, `php`, `aspx`, `erb`.
1579
+ `jsp`/`aspx`/`erb` share the `<%= %>` delimiters, so their probes are
1580
+ byte-identical and cost one request between them; with no extension to read,
1581
+ `auto` writes all five in three requests.
1582
+
1583
+ ### Changed
1584
+
1585
+ - **A `needs-review` finding now prints its cleanup line too.** It used to
1586
+ appear only under `confirmed`, which was already thin and is wrong for this
1587
+ method: a `write` reaching `needs-review` means the file *is* on the target,
1588
+ just not interpreted, so the artifact would have been left there unmentioned.
1589
+
1590
+ - The write method's operands are drawn once per run rather than once per
1591
+ carrier, so the file is written once instead of once for each of the ~13
1592
+ `(environment, context)` carriers. For a state-changing method that is not a
1593
+ request-count saving, it is a blast radius. Still fresh per run, which is what
1594
+ makes the product unforgeable.
1595
+
1596
+ - The write method declines the break-out contexts (`sql`, `javascript`,
1597
+ `shell_*`, …) and keeps the transport ones. Its payload is a whole file body:
1598
+ there is nothing to break out of, and wrapping it in `'; … -- ` would write a
1599
+ broken file. A run narrowed past `raw` and the transport contexts is told so
1600
+ rather than reporting a clean negative.
1601
+
1602
+ ## [2.30.0] — 2026-08-17
1603
+
1604
+ Per-dialect shell probes. `$((a+b))`, `sleep` and `$(echo TAG)` are POSIX
1605
+ constructs: on a cmd.exe or PowerShell sink they are inert literal text. The
1606
+ dialect was inferred from the corpus environment alone, so a run could send a
1607
+ probe no shell on the target would ever execute — including on the carrier whose
1608
+ context is literally named `powershell`.
1609
+
1610
+ ### Added
1611
+
1612
+ - **`--sink-env auto|unix|windows|powershell`** states which shell runs the
1613
+ injected command. The computed-value core, the separators and the break-out
1614
+ contexts are all chosen from it. `auto` (the default) infers it per carrier;
1615
+ pin it when the corpus environment names the *application runtime* rather than
1616
+ the OS — `--environments php --sink-env windows` is a PHP application on IIS,
1617
+ which no inference can see.
1618
+
1619
+ - **A PowerShell probe shape for every shell method**, validated against
1620
+ pwsh 7.4: `Write-Output T1$(a*b)T2` for `reflected` (an unquoted argument is an
1621
+ expandable string, so the core carries no quote and the quote-wrapping
1622
+ contexts can still carry it), `Start-Sleep -Milliseconds N` for `time`,
1623
+ `Set-Content` for `file` and `iwr -useb` for `oob`. PowerShell was previously
1624
+ reachable by no probe in any method.
1625
+
1626
+ - **cmd.exe and PowerShell carriers for the `dotnet` environment.** It is the
1627
+ one corpus environment that names a platform, and it was taking the POSIX
1628
+ shape — so .NET on Windows, the case the environment exists for, was the case
1629
+ it could not confirm on. Every other runtime keeps the POSIX shape: a language
1630
+ does not say which OS it runs on.
1631
+
1632
+ ### Fixed
1633
+
1634
+ - **The `powershell` carrier was written in cmd.exe.** Every `windows` carrier
1635
+ took the `for /f ... ('set /a a+b')` core regardless of context, so the one
1636
+ carrier explicitly shaped for PowerShell sent a payload PowerShell cannot
1637
+ execute. The dialect now follows the carrier's context first, and a carrier's
1638
+ break-out variants stay in its dialect rather than re-deriving from the
1639
+ environment.
1640
+
1641
+ - **`Set-Content`, not `>`, for the PowerShell write.** In Windows PowerShell
1642
+ 5.1 the redirect is `Out-File`, whose default encoding is UTF-16LE: the write
1643
+ lands and the read-back still does not find the token, so the probe reports
1644
+ negative on a target it owns.
1645
+
1646
+ ### Changed
1647
+
1648
+ - **cmd.exe no longer gets the `sq`, `dq` and `subshell` carriers.** It has
1649
+ neither a comment character to swallow the sink's tail nor a
1650
+ command-substitution syntax, so those four carriers per Windows run were
1651
+ requests that could only come back negative. PowerShell takes the quote
1652
+ break-outs and `$( )` — both measured — but not the backtick, which is its
1653
+ escape character rather than a substitution.
1654
+
1655
+ - **PowerShell's separator sweep carries no pipe.** `cmd | Start-Sleep
1656
+ -Milliseconds 500` is a parameter-binding error, not a fresh command with
1657
+ stdin attached the way a POSIX pipe is, and it fails that way for every cmdlet
1658
+ the probes use. `;`, a newline and (on PowerShell 7) `&&`/`||` remain.
1659
+
1660
+ - The pre-flight plan prints the sink shell alongside the sink shapes, and a
1661
+ pinned dialect narrows the printed ladder to the rungs it has syntax for.
1662
+
1663
+ ## [2.29.0] — 2026-08-17
1664
+
1665
+ Generalised read-back for the `file` method. It required a writable **web root**
1666
+ the tester already knew, which ruled out every other way a target can hand a
1667
+ file back — on exactly the internal, no-egress targets the method exists for.
1668
+
1669
+ ### Added
1670
+
1671
+ - **`--file-write-path DIR` + `--file-read-url URL`** name the two halves of the
1672
+ read-back channel directly, so an LFI endpoint, a download or export handler,
1673
+ an attachment fetcher or a `/tmp`-backed preview all work. The template takes
1674
+ `{name}` (the filename), `{path}` (the full server-side path) and `{path_enc}`
1675
+ (that path percent-encoded); only those three are substituted, so a URL that
1676
+ legitimately contains braces survives unchanged.
1677
+
1678
+ Measured against a target with a download handler and nothing serving the
1679
+ write directory: the web-root form confirms **0** — reporting an exploitable
1680
+ target clean — and the general form confirms **7**.
1681
+
1682
+ ### Fixed
1683
+
1684
+ - **The read-back fetch now carries the run's headers**, so an authenticated
1685
+ download, export, attachment or LFI handler can actually be read. It went out
1686
+ bare, which barely mattered while the channel had to be a web root — static
1687
+ file serving is rarely authenticated — and became the likely case the moment
1688
+ the channel could be an application endpoint. Measured against a handler
1689
+ behind a bearer token: the write executed on every probe and the verdict was
1690
+ `negative`, "token absent from the fetched file". Now 7 confirmations on the
1691
+ same target.
1692
+ - **Credentials are carried only to the same origin.** A read-back URL on
1693
+ another host is someone else's server, and replaying the target's session
1694
+ cookie or bearer token to it would leak the credential, so those headers are
1695
+ dropped there while the rest still go — and the run says so, because the
1696
+ symptom would otherwise look like a clean target. `Content-Type` and
1697
+ `Content-Length` are dropped from the fetch too: they describe a body the GET
1698
+ does not have.
1699
+
1700
+ ### Changed
1701
+
1702
+ - **`--webroot` / `--web-base-url` are now the web-root alias** for the general
1703
+ form: a web root is just the case where the read URL is the base plus the
1704
+ filename. Existing command lines are unaffected. Both are resolved in one
1705
+ place inside the method, so the alias and the general form cannot drift — and
1706
+ the gate, the pre-flight banner and the blind-sink advice all ask that same
1707
+ resolver instead of testing for the webroot pair.
1708
+ - `blind_sink_advice` reads its flags defensively, so an args-like object
1709
+ missing a newer field costs a line of advice rather than a traceback.
1710
+
1711
+ ## [2.28.0] — 2026-08-17
1712
+
1713
+ Injection-point enumeration. `-p NAME` needed the tester to already know which
1714
+ parameter was the sink, so a capture's other candidates — including the headers
1715
+ and nested JSON leaves that carry some of the highest-value classes — were never
1716
+ tried.
1717
+
1718
+ ### Added
1719
+
1720
+ - **`-p all` / `--auto-params KINDS`** expands one captured request into every
1721
+ candidate injection point and runs the selected `--methods` against each.
1722
+ Query values, JSON leaves addressed by path (`user.profile.name`, `tags[1]`),
1723
+ form fields, cookie crumbs and headers, each rewritten in **its own**
1724
+ serialization rather than blanket-encoded. Verified end to end: a sink
1725
+ reachable only through `User-Agent` is confirmed from `-r request.txt -p all`
1726
+ with no manual header selection.
1727
+ - **`--point-order fast|thorough`** — `fast` tries a curated high-yield header
1728
+ list (the headers real published RCEs inject through); `thorough` adds every
1729
+ remaining non-hop-by-hop header. **`--max-points N`** bounds the run and
1730
+ reports what it dropped. **`--include-path-segments`** is opt-in, because
1731
+ rewriting a path segment usually just produces a 404.
1732
+ - **The run states its cost before sending it** —
1733
+ `6 points x ~61 probes = at least 372 requests` — via a new
1734
+ `estimate_detection_probes`, which builds the probes and counts them without
1735
+ firing any. Enumeration multiplies an already-laddered probe count by the
1736
+ candidate count, and an operator on a monitored engagement has to see that
1737
+ before it happens rather than infer it from the traffic.
1738
+ - **Findings name the point they came from**: `[reflected/unix/raw] at header
1739
+ 'User-Agent' ...`.
1740
+
1741
+ ### Changed
1742
+
1743
+ - **Each candidate carries its own payload-free control.** Differencing a header
1744
+ probe against a query probe's control would compare two different responses
1745
+ and prove nothing.
1746
+ - **Cheap methods run first per candidate, and a candidate stops at its first
1747
+ confirmation.** `reflected` and `eval` cost one response each; `time` sleeps
1748
+ and `oob` waits for a callback, and on a candidate that has already proven
1749
+ execution those buy a second name for the same finding. Candidates that stay
1750
+ clean still get every method, and single-point runs are unchanged.
1751
+ - A JSON leaf is **replaced, never created**. Assigning to a missing key would
1752
+ have injected into a field the application never sends — a probe that cannot
1753
+ say anything about the parameter that does exist. Caught by its own test.
1754
+ - **JSON points are addressed by tokens, not by a joined path string.** A key may
1755
+ itself contain the separator: `{"user.name": ..., "user": {"name": ...}}`
1756
+ rendered *both* leaves as `user.name`, so the literal key was never probed and
1757
+ both candidates mutated the nested field — a false negative and a misattributed
1758
+ finding at once. Tokens remove the ambiguity, and the display form
1759
+ bracket-quotes such a key (`["user.name"]`) so the two stay distinguishable on
1760
+ screen.
1761
+ - **A deeply nested captured body no longer ends `-p all` with a traceback.**
1762
+ `json.loads` recurses in C, so `RecursionError` joins the caught exceptions in
1763
+ both the enumerator and the placer, as it already had in the response-channel
1764
+ parser. The body yields no candidates; the rest of the request still enumerates.
1765
+ - **The cost estimate honours `--max-payloads`.** It counted every probe the
1766
+ carriers could produce while the run stops at the cap, so the figure was wrong
1767
+ exactly when the operator had reached for the budget guard.
1768
+ - `Host`, `Content-Length`, `Cookie` and the hop-by-hop headers are never
1769
+ candidates: injecting into those changes the request's plumbing rather than
1770
+ testing the application, and two of them are rebuilt by the delivery layer.
1771
+
1772
+ ## [2.27.0] — 2026-08-17
1773
+
1774
+ Engine carriers for the `eval` probe. Three template engines evaluate the
1775
+ injected expression perfectly and still made RCEKit report `negative`, because
1776
+ what came back was not the bare product the oracle searches for.
1777
+
1778
+ ### Added
1779
+
1780
+ - **`eval_carriers` in the corpus**, and `--eval-engines auto|<names>` to select
1781
+ them. A carrier wraps the same random-operand arithmetic in an engine-specific
1782
+ form; it never changes the oracle, and the bare probes still run first. Each
1783
+ entry records `notes` (why it exists) and `verified` (what it was measured
1784
+ against). Declarative, so a new carrier is a JSON entry rather than a code
1785
+ change.
1786
+
1787
+ | Engine | Bare `${a*b}` returned | Carrier | Carrier returned |
1788
+ |---|---|---|---|
1789
+ | Freemarker | `2,070,761,401` (locale grouping) | `${(a*b)?c}` | `2070761401` |
1790
+ | Velocity | `${a*b}` verbatim — a *reference*, not an expression | `#set($rk=a*b)$rk` | `2070761401` |
1791
+ | Thymeleaf | `${a*b}` verbatim — needs inlining brackets | `[[${a*b}]]` | `2070761401` |
1792
+
1793
+ Measured against freemarker 2.3.32, velocity-engine-core 2.3 and thymeleaf
1794
+ 3.1.2, running RCEKit's own generated probes through each engine: bare form
1795
+ `CONFIRMS=no`, carrier `CONFIRMS=YES`, for all three.
1796
+ - **The evidence line names the carrier** — `target computed '3979016' via the
1797
+ freemarker carrier` — so a finding says which engine quirk it worked around.
1798
+ A bare confirmation reads exactly as before.
1799
+
1800
+ ### Notes
1801
+
1802
+ - **Carriers are not sandbox escapes, and no sandbox-escape carrier ships.** The
1803
+ premise that a sandboxed engine blocks the arithmetic probe did not survive
1804
+ measurement: a member-access sandbox restricts method and field access, and
1805
+ arithmetic needs neither. With OGNL member access denied for *everything*,
1806
+ `40277*51413` still returned `2070761401` while `@java.lang.Math@max(1,2)` was
1807
+ blocked; SpEL's restricted `SimpleEvaluationContext` and Jinja2's
1808
+ `SandboxedEnvironment` behaved the same way. The bare probes already cover
1809
+ those engines.
1810
+ - The frequently-cited OGNL escape `(#_memberAccess=@ognl.OgnlContext@DEFAULT_MEMBER_ACCESS)`
1811
+ additionally targets a field that **no longer exists in OGNL 3.3.4**, so on a
1812
+ current engine it is a probe that can only come back negative.
1813
+
1814
+ ## [2.26.0] — 2026-08-16
1815
+
1816
+ The sink-shape ladder. An injected value lands in a *shape* — mid-command,
1817
+ inside quotes, as the whole command — and the shape decides what can reach it.
1818
+ Two shapes had no probe that fitted, so a genuinely exploitable target reported
1819
+ clean.
1820
+
1821
+ ### Added
1822
+
1823
+ - **`--sink-shape auto|sep|raw|chain|newline|dq|sq|subshell`** (comma-separated)
1824
+ names which shapes the shell probes try. `auto` is the whole ladder and the
1825
+ default. Underneath it selects the existing separator sweep and break-out
1826
+ contexts, so naming a rung narrows a supported run rather than switching on a
1827
+ parallel path. The plan is printed before anything is sent, because the ladder
1828
+ multiplies request count and an operator on a monitored engagement needs to
1829
+ see the cost first.
1830
+ - **The `subshell` rung — `$(...)` and backticks.** Reaches a value sitting
1831
+ inside double quotes *without closing the quote*, which is the one case a
1832
+ quoted break-out loses to a filter on the quote character itself. Measured
1833
+ against `system("echo PING \"$input\"")`: with `"` stripped, `dq` is inert and
1834
+ both substitution forms execute; with `$` stripped, `dq` executes and the
1835
+ backtick form still does. Both ship because they survive different filters.
1836
+
1837
+ **Which method it helps is the counter-intuitive part.** `reflected`'s core is
1838
+ `$((a+b))`, which the shell expands inside double quotes anyway, so that
1839
+ method already confirmed there. The methods whose core must actually *run* —
1840
+ `time` (a sleep), `file` (a redirect), `oob` (a fetch) — are completely inert
1841
+ inside those quotes. On a quote-filtering sink, `--methods file` went from **0
1842
+ confirmations to 2**: it had been reporting an exploitable target as clean.
1843
+ - **The `raw` rung is now part of `auto`, for every shell method.** A
1844
+ `qx/$input/`-style sink, where the input is the whole command, previously
1845
+ needed `--sink-raw` — so it reported clean unless the operator already
1846
+ suspected the shape. One extra probe per carrier buys it. `--sink-raw` keeps
1847
+ its meaning as the narrowing alias for `--sink-shape raw`, and no existing
1848
+ command line changes behaviour. `reflected`, `file`, `time` and `oob` all
1849
+ build their candidates through one `_separator_candidates` helper, so a rung
1850
+ cannot reach some methods and not others; `time` screens it in its second
1851
+ wave, alongside the separators it holds back.
1852
+
1853
+ ### Fixed
1854
+
1855
+ - **A method that builds no probes no longer reports `negative`.** An aggregate
1856
+ method asked to judge zero samples answers honestly — "no delay was observed",
1857
+ "no callback arrived" — and that reads as "not vulnerable" from a run that
1858
+ tested nothing. The engine now emits no row for a carrier that produced no
1859
+ probes, which lets its own loud nothing-tested path fire instead. Reachable
1860
+ through any narrowing that leaves a carrier with nothing to send.
1861
+
1862
+ ### Changed
1863
+
1864
+ - **The pre-flight sink-shape plan is computed from the effective run**, not
1865
+ from the `--sink-shape` value. `--separators`, `--contexts` and `--sink-raw`
1866
+ each narrow the ladder, so printing the flag described a run that would not
1867
+ happen — and this output is presented as an audit of the traffic about to be
1868
+ sent. `effective_sink_shapes` is the single source of truth the engine and the
1869
+ plan both read.
1870
+ - **The backtick context drops probe shapes that carry their own backtick.**
1871
+ Backticks do not nest, so such a probe closes the outer substitution early and
1872
+ could only ever come back negative. `$( )` does nest and keeps every shape.
1873
+ - **Naming `--separators` now implies the sink is separator-led**, so the `raw`
1874
+ rung is dropped unless `--sink-shape` names it explicitly. A profile with
1875
+ `sink_needs_separator` drops it for the same reason. Both keep an explicitly
1876
+ narrowed run from being widened behind the operator's back.
1877
+
1878
+ ## [2.25.0] — 2026-08-16
1879
+
1880
+ A coverage benchmark, so a claim about what RCEKit confirms can be checked
1881
+ instead of asserted. The unit suite proves the tool reaches the right verdict
1882
+ from a given response; it cannot prove it confirms Webmin.
1883
+
1884
+ ### Added
1885
+
1886
+ - **`--detect-json PATH`** writes a detection run as JSON: the run's overall
1887
+ verdict, per-verdict counts, and every probe with its payload, method, context
1888
+ and evidence. Text output is unchanged. This is the supported way to consume a
1889
+ run programmatically — scraping stdout cannot be made reliable, because a
1890
+ probe payload may contain a literal newline (the newline separator is a real
1891
+ one, so line-oriented parsing splits a payload in half) and the detection path
1892
+ exits 0 whether it confirmed or came back clean.
1893
+ - **`tests/bench/` — the coverage benchmark harness.** Each case brings a real
1894
+ vulnerable build up, runs RCEKit as an operator would, checks the verdict, and
1895
+ tears it down; `--markdown` emits the coverage table. Not part of
1896
+ `python -m unittest discover -s tests` — cases need Docker and pull real
1897
+ images — so it runs by hand or in a dedicated job, and exits non-zero if any
1898
+ case fails. Two cases ship, transcribed from `docs/verify-it-yourself.md`:
1899
+ Webmin CVE-2019-15107 and Struts2 S2-001.
1900
+ - **A negative control is a required key.** A benchmark without controls measures
1901
+ nothing: a tool that shouted `confirmed` at every target would score full marks
1902
+ on the vulnerable half. Three kinds are supported — a patched build, the same
1903
+ target probed for the wrong class, and a weaker method that must stay below
1904
+ `confirmed` on a target where it happens to be right. The runner refuses four
1905
+ shapes of non-control: no control at all; one expecting `confirmed`; one that
1906
+ runs the identical invocation against an identical target (judged on what it
1907
+ would actually run, so an explicit copy of the vulnerable invocation is caught
1908
+ as well as an omitted one); and one expecting `error` or `nothing-tested`,
1909
+ since both mean the target was never exercised and such a control would stay
1910
+ green with the detection engine entirely broken. Validation and execution
1911
+ share one `control_plan` so they cannot drift.
1912
+ - **`overall_detection_verdict`** collapses a run to one verdict, ordered by what
1913
+ an operator must not miss rather than by frequency: one `confirmed` among a
1914
+ hundred negatives is the finding. `error` is reported only when *nothing*
1915
+ reached the target, and a run that built no probes is `nothing-tested` —
1916
+ never `negative`, which would read as "not vulnerable".
1917
+
1918
+ ### Changed
1919
+
1920
+ - `CONTRIBUTING.md` asks for a bench case alongside new detection coverage, and
1921
+ for the README table to state the tier the case actually reached.
1922
+
1923
+ ### Notes
1924
+
1925
+ - The two shipped cases have **not yet been executed through the harness** — it
1926
+ was written where no Docker daemon was available. Their invocations come from
1927
+ a documented, reproduced guide, but the case files themselves are unvalidated;
1928
+ `tests/bench/README.md` says so and flags the one field that is a guess. No
1929
+ README claim was changed to assert benchmark results.
1930
+
1931
+ ## [2.24.0] — 2026-08-16
1932
+
1933
+ The computed value is no longer looked for in the response body alone. A sink
1934
+ whose output surfaces anywhere else in the response was reported `negative` — a
1935
+ false negative on a class RCEKit already claims to cover, which is worse than a
1936
+ missing class. The oracle, the random operands and the control differential are
1937
+ unchanged; only the set of places searched is wider.
1938
+
1939
+ ### Added
1940
+
1941
+ - **Whole-response evidence search.** Every confirmation now sweeps the response
1942
+ body, the application response headers, individual cookie values, the redirect
1943
+ target RCEKit actually landed on, the HTTP reason phrase, and each leaf of a
1944
+ parsed JSON body. Real sinks put command output in a debug header or a
1945
+ `Set-Cookie`, and API targets surface an evaluator's result inside a nested
1946
+ error envelope — `{"error": {"detail": "cannot render 2058898001"}}` — where a
1947
+ substring search of the serialised body misses a value the encoder escaped.
1948
+ - **The evidence line names the channel that carried the value**, e.g.
1949
+ `target computed 'RK…' in header X-Cmd-Out (random operands, absent from
1950
+ control)`, so the finding stays reproducible by hand. A body-carried
1951
+ confirmation reads exactly as it did before.
1952
+
1953
+ ### Changed
1954
+
1955
+ - **The control differential now covers every channel, not just the body.** A
1956
+ value present anywhere in the payload-free control is not attributable to
1957
+ execution, so it yields `inconclusive` wherever it turned up. This is stricter
1958
+ than comparing only the channel that matched, and it is what keeps a wider
1959
+ search from becoming a looser verdict.
1960
+ - **The `file` method's control check covers every channel too**, on the same
1961
+ reasoning: its token is random, so its presence in any control channel means
1962
+ it did not get there by being written and served.
1963
+
1964
+ ### Security
1965
+
1966
+ - **A deeply nested JSON response can no longer silence detection.** Channels are
1967
+ built inside the delivery `try`/`except`, so a `RecursionError` while parsing
1968
+ or walking the body escaped as a network failure: a response that arrived
1969
+ perfectly well was reported "request never reached the target". Measured:
1970
+ every one of the 46 probes in a default `reflected` run turned into `error`,
1971
+ which a target could induce deliberately to hide a live sink behind a thousand
1972
+ nested arrays. Version-independent, though the source moves — CPython 3.12
1973
+ raised the C recursion limit its JSON scanner runs under, so on 3.12/3.13 the
1974
+ parser survives a depth that breaks it on 3.8–3.11 and the recursive leaf walk
1975
+ hit the ordinary Python limit instead. The walk is now iterative and
1976
+ depth-capped, `RecursionError` from the parser costs the JSON channels only,
1977
+ and building channels can never turn a delivered response into a delivery
1978
+ failure.
1979
+ - **Transport headers are excluded from the sweep.** `Content-Length`, `Date`,
1980
+ `Age`, `ETag` and their neighbours are generated below the application and can
1981
+ never carry a computed value, but they *are* numeric — and the `expr` probe's
1982
+ expected value is a bare boundary-fenced number. Searching them would let a
1983
+ byte count collide with an arithmetic result and read as execution. Locked in
1984
+ by a test that puts the expected value in `Content-Length` and requires
1985
+ `negative`.
1986
+
1987
+ ## [2.23.3] — 2026-08-04
1988
+
1989
+ Four items from the same review: requests and seconds spent on work that could
1990
+ not produce a result. No verdict changes — the lab still confirms 15 of 15
1991
+ vulnerable sinks with nothing on the clean five — the run just stops paying for
1992
+ probes that were structurally unable to confirm.
1993
+
1994
+ ### Changed
1995
+
1996
+ - **The `awk` probe is no longer sent into a context that wraps the payload in
1997
+ quotes.** It carries double quotes, so in `attribute` the quote closed early
1998
+ and the rest was not a command: 5 requests per carrier that could only ever
1999
+ come back negative. Measured on a verbose shell sink, that shape confirmed 8
2000
+ times in `raw` and 0 times in `attribute`. Break-out contexts such as
2001
+ `shell_double_quoted` *close* the sink's quote and comment its tail, so they
2002
+ still get it. The same guard covers the PowerShell out-of-band shape.
2003
+ - **The timing screen runs in two waves.** Every delayed screen probe costs a
2004
+ real sleep, so screening all five separators up front spent `5 × base` seconds
2005
+ on every carrier, including the ones that cannot break out at all. `; ` and
2006
+ `| ` are screened first and the rest only if neither delayed — a sink that
2007
+ filters both is still swept, it is just no longer the price everyone pays.
2008
+ - **The out-of-band callback window is no longer paid per carrier.** Callbacks
2009
+ land in a burst once the channel works, so a target that has not produced one
2010
+ across every probe fired so far is not going to. The first carrier still gets
2011
+ the full window, so a target that does call back is never cut short before its
2012
+ first hit. On a clean target with the default carriers this was 30s of pure
2013
+ waiting; it is now ~12s.
2014
+ - **`--probe-depth` documents what it does on Windows**, which is nothing:
2015
+ `cmd.exe` has no `#` comment, no `${IFS}` and no `awk`, so both depths send
2016
+ the single `set /a` probe. The docs promised three extra shapes per sink
2017
+ without that caveat.
2018
+
2019
+ ## [2.23.2] — 2026-08-04
2020
+
2021
+ Three findings from a review of the detection work in 2.22.0 and 2.23.0. All
2022
+ three are the same shape: the run said something that was not true — about what
2023
+ it had done, about what it had looked for, or about which channel was live.
2024
+
2025
+ ### Fixed
2026
+
2027
+ - **`--methods oob` ignored `--verify-active-risk`.** Detection methods build
2028
+ their own probes and so bypass every corpus-level safety filter. That was
2029
+ harmless while every method was inert, but this one makes the target open
2030
+ outbound connections — and the same run printed *"low-impact (safe) payloads
2031
+ only; pass `--verify-active-risk intrusive` to also fire … OOB"* and then fired
2032
+ OOB anyway. It now needs `--verify-active-risk intrusive`, the same tier that
2033
+ holds back the corpus OOB payloads, and refuses before the listener binds.
2034
+ - **`--probe-depth quick` silently narrowed the timing separator screen to
2035
+ `; `.** That put back the exact blind spot the screen was added to remove, so
2036
+ a sink that merely filters `;` reported negative — and only for the operator
2037
+ who chose `quick` to be gentle on a rate-limited target. Both depths now screen
2038
+ every candidate separator; `--probe-depth` governs probe *shapes* only, and
2039
+ `--separators` remains the way to narrow break-outs deliberately.
2040
+ - **The DNS out-of-band probes could not call back on the default port, and
2041
+ nothing said so.** A DNS callback travels the real resolver hierarchy, so it
2042
+ only arrives if the listener *is* the authority for the OOB domain — port 53
2043
+ plus NS delegation. On `--listen-dns-port 5335` the DNS shapes were still sent,
2044
+ never fired, and the startup line reported `DNS :5335` with no caveat. Since
2045
+ most of the shapes are DNS ones — a resolver is often the only egress a
2046
+ hardened target has — the silence was expensive. RCEKit now says which channel
2047
+ is live.
2048
+ - The blind-sink advice added in 2.23.0 suggested an `oob` command without the
2049
+ risk flag, which the gate above would refuse. Naming a command the tool then
2050
+ declines to run is a small version of the same problem, so it now spells out
2051
+ `--verify-active-risk intrusive`.
2052
+
2053
+ ## [2.23.1] — 2026-08-03
2054
+
2055
+ ### Added
2056
+
2057
+ - **[Verify it yourself](docs/verify-it-yourself.md)** — reproduce the README's
2058
+ confirmations locally against dockerised [vulhub](https://github.com/vulhub/vulhub)
2059
+ targets. Webmin CVE-2019-15107 driven from a captured request (`reflected` →
2060
+ `confirmed`, then `time` → `needs-review` on the *same* sink, which is the
2061
+ clearest demonstration that the tiers are not merged), and Struts2 S2-001
2062
+ (`eval` confirms, `reflected` does not, on a target where both were tried).
2063
+
2064
+ Log4Shell is documented as an advanced case rather than a five-minute one: its
2065
+ sink is a JNDI lookup inside a logging library, so `--methods oob` does not
2066
+ apply — that method builds shell probes for shell-capable environments. The
2067
+ `${jndi:…}` payloads come from the `oob` *category* with the listener
2068
+ correlating the callback, and the token rides in a DNS label, which needs a
2069
+ delegated domain. Saying so is cheaper than a reader discovering it mid-demo.
2070
+
2071
+ ## [2.23.0] — 2026-08-03
2072
+
2073
+ The three sinks v2.22.0 still could not reach. One was a real gap in the probe
2074
+ set; the other two were a reporting problem, not a detection one. With both
2075
+ closed, a single `--methods reflected,eval,oob` run confirms **all fifteen**
2076
+ vulnerable sinks in the lab and still reports nothing on any of the five clean
2077
+ ones.
2078
+
2079
+ ### Added
2080
+
2081
+ - **A space-free probe, sent at both probe depths.** Stripping spaces is a filter
2082
+ of the same family as stripping `;` — it looks like it disarms command
2083
+ injection and does not, because `${IFS}` is a space as far as the shell is
2084
+ concerned. Every other probe carries a space, so that one filter silenced all
2085
+ of them and the sink was only reachable if the operator thought to pass
2086
+ `--evade low`. The separator's trailing space is trimmed with it (`;echo…`, not
2087
+ `; echo…`); the newline separator is unaffected. It costs one shape, so it is
2088
+ not part of the `--probe-depth` trade-off, and it is skipped under
2089
+ `--evade low`, which already applies the same transform everywhere.
2090
+ - **Guidance when every in-band probe comes back negative.** A results-based
2091
+ method cannot confirm a sink that returns no output — there is nowhere for the
2092
+ computed value to appear — so that negative is not evidence the target is
2093
+ clean. A run of `reflected`/`eval` alone that confirms nothing now says exactly
2094
+ that and names the methods that could still reach a blind sink, with the flags
2095
+ each one needs. It is suppressed once a blind-capable method has already run,
2096
+ and the `file` line is dropped once a web root is known.
2097
+
2098
+ ## [2.22.0] — 2026-08-03
2099
+
2100
+ Detection coverage. Measured against a lab of twenty sinks — fifteen genuinely
2101
+ vulnerable, five deliberately clean — the results-based methods went from
2102
+ confirming 8 of the 15 to confirming 12, with no new false positives on any of
2103
+ the clean ones.
2104
+
2105
+ ### ⚠️ A blind-timing candidate could be pure latency drift
2106
+
2107
+ `--methods time` fired its probes in a fixed ascending delay order
2108
+ (`0,0,N,N,2N,2N`), which makes the injected delay collinear with the request
2109
+ index. A target that simply gets **slower during the run** — progressive load, a
2110
+ rate limiter backing off, a filling log — therefore produced a textbook-perfect
2111
+ linear fit while being entirely un-injectable. In the lab this reproduced on 8
2112
+ of 8 runs against a sink with no command execution anywhere in it.
2113
+
2114
+ The probe order is now randomised, and the request index enters the regression
2115
+ as a nuisance term, so drift loads onto a drift coefficient instead of
2116
+ masquerading as a sleep. The same lab sink now reports negative on 9 of 9 runs,
2117
+ with every genuine timing detection preserved. If you have a `needs-review`
2118
+ timing candidate from an earlier version against a target that was under load,
2119
+ it is worth re-running.
2120
+
2121
+ ### Added
2122
+
2123
+ - **`--methods oob`** — out-of-band detection, the first `confirmed`-tier method
2124
+ for a sink that returns nothing *and* has no writable web root. Starts the
2125
+ built-in HTTP+DNS listener in-process and asks the target to resolve or fetch
2126
+ `<token>.<oob-host>`; a callback carrying a token the target could only have
2127
+ learned by running the command is proof of execution. Each probe gets its own
2128
+ token, so the finding names the break-out that actually worked. One shape puts
2129
+ a computed value in the DNS label, so the callback proves the shell evaluated
2130
+ arithmetic rather than merely resolving a name. Requires `--oob-host`, since
2131
+ it makes the target open outbound connections.
2132
+ - **`--probe-depth quick|full`** (default `full`) — trades requests for
2133
+ coverage. `full` adds three probe shapes, each aimed at a filter that silenced
2134
+ the canonical ones: substitution-free (`awk`, bare `expr`) for sinks that strip
2135
+ `$(` and backticks; keyword-diverse (`awk`) for filters on `echo`/`expr`; and
2136
+ comment-terminated (`… #`) for applications that append a redirect, extra
2137
+ arguments or a pipe after the injection point. `quick` keeps the old probe set
2138
+ at roughly half the requests.
2139
+
2140
+ ### Fixed
2141
+
2142
+ - **A `ping '<input>'` sink could not be detected at all.** The
2143
+ `shell_single_quoted`/`shell_double_quoted` contexts exist precisely for input
2144
+ interpolated inside quotes, but they are not in `default_contexts`, so no
2145
+ record carried them and the detection engine never tried them — the one sink
2146
+ shape they exist for was the one shape that always reported clean. They are now
2147
+ probed by default, and skipped when `--contexts` names a selection explicitly.
2148
+ - **`--methods time` reported a `;`-filtering sink as negative.** A regression
2149
+ blends its probes into one measurement, so it could not sweep separators the
2150
+ way the results-based methods do and was locked to `; ` alone — while
2151
+ `| sleep 3` delayed on the same sink. It now screens every candidate separator
2152
+ with one cheap probe each, then runs the regression through whichever one
2153
+ actually delayed.
2154
+ - **A trailing redirect or pipe in the sink hid a working probe.**
2155
+ `<cmd> <input> 2>/dev/null` and `<cmd> <input> | grep …` swallow the probe's
2156
+ output, so it executed and still read as negative. The comment-terminated
2157
+ shapes comment that tail out.
2158
+
2159
+ ## [2.21.1] — 2026-08-02
2160
+
2161
+ First release since v2.15.2. The headline is not a new feature — it is that
2162
+ detection is now correct in cases where it previously was not.
2163
+
2164
+ ### ⚠️ Re-check findings from v2.15.2 and earlier
2165
+
2166
+ **A reflection could be reported as `confirmed`.** The paired same-token control
2167
+ in `run_verification` was gated on a plain `re.search`, while the verdict itself
2168
+ used the encoding-aware search. A target that only echoes input but wraps its
2169
+ output — base64, hex, URL- or HTML-encoded — skipped the control entirely and was
2170
+ reported as proven execution: precisely the case the encoding-aware search was
2171
+ added for. If you ran an earlier version against a target that encodes its
2172
+ responses, a `confirmed` verdict from that run is worth re-testing.
2173
+
2174
+ ### Fixed — false negatives on exploitable targets
2175
+
2176
+ - **Separator sweep.** Shell probes always broke out with a single hardcoded
2177
+ `; `, so a sink that strips `;` — the most common partial mitigation there is,
2178
+ and one that stops nothing on its own — defeated every probe. Measured against
2179
+ nine deliberately vulnerable local sinks, detection was correct on 5 of 9;
2180
+ three of the four misses were exploitable targets reported clean. Probes now
2181
+ sweep `; `, `| `, `|| `, `&& ` and a newline, narrowable with `--separators`.
2182
+ - **Language runtimes.** An environment names what runs the application, not what
2183
+ executes the injected command: PHP's `system()`, Python's `os.system()`,
2184
+ Node's `child_process.exec()`, Ruby's `system()`, Perl's backticks and Go's
2185
+ `os/exec` all hand the string to `/bin/sh`. Scoping a run to the language the
2186
+ application is written in — the natural thing to do — used to send no shell
2187
+ probes at all.
2188
+ - **Whole-command sinks.** `--sink-raw` sends probes as bare commands for sinks
2189
+ that execute the input as the entire command (`qx/$input/`, `sh -c "$input"`),
2190
+ where a leading `;` is a syntax error that guaranteed a false negative.
2191
+ - **Captured requests.** A trailing newline in a saved request body is no longer
2192
+ sent as part of the body.
2193
+
2194
+ ### Fixed — a failed request is not a clean result
2195
+
2196
+ - A request that never reached the target is reported `error`, not `negative`.
2197
+ - Runs that build no probes at all exit non-zero and say so, instead of ending
2198
+ in silence and exit 0 — which read exactly like a target that came back clean.
2199
+ - The OOB DNS listener no longer dies on a malformed query, and write failures
2200
+ surface instead of being swallowed.
2201
+
2202
+ ### Fixed — safety and audit
2203
+
2204
+ - **Multi-step chains now carry the same safeguards as single requests.** The
2205
+ chain path delivered to live targets without sink-shape filters, destructive
2206
+ hold-back or a pre-flight plan, so `--verify-active-risk stateful` fired
2207
+ persistence and irreversible file operations that `--verify-url` refuses to
2208
+ send without `--verify-allow-destructive`. Both paths now share one hold-back
2209
+ and print the same plan.
2210
+ - **The audit trail redacts credentials** — it records that a credential header
2211
+ was sent, never its value.
2212
+ - A capture carrying `Authorization` or `Cookie` over plain `http` is flagged
2213
+ before anything is sent.
2214
+
2215
+ ### Added
2216
+
2217
+ - **The payload corpus is embedded in `rcekit.py`**, so the single file runs on
2218
+ its own — a jump box, an air-gapped host, a bare `curl` of the raw script.
2219
+ Resolution order is `--template-file` → `templates/payloads.json` beside the
2220
+ script → the built-in copy, and falling back to the built-in copy is
2221
+ announced. A corpus that exists but does not parse still hard-fails: that
2222
+ check exists for truncated and tampered corpora. `tools/embed_corpus.py`
2223
+ regenerates the embedded copy, and the test suite fails if the two drift.
2224
+ - **`--insecure`** skips TLS verification for internal targets with self-signed
2225
+ or mismatched certificates — opt-in and explicit, like `curl -k`.
2226
+ - **`--sink-raw`** for whole-command injection sinks, also readable from a
2227
+ target profile.
2228
+ - **`--separators`** to narrow the break-out sweep once the sink's shape is
2229
+ known.
2230
+ - **Documentation split into a task-oriented tree.** The README is half its
2231
+ former length and now leads with what RCEKit is for:
2232
+ [field guide](docs/guide.md) (worked examples by situation),
2233
+ [payload generation & exports](docs/generation.md), and
2234
+ [reference](docs/reference.md) (every flag grouped by task, plus the full
2235
+ taxonomies and exit codes).
2236
+ - **A "How RCEKit compares" section** covering commix, SSTImap, Nuclei and
2237
+ interactsh, with every claim traceable to that project's own documentation.
2238
+ - Four confirmation demos against real, publicly documented CVEs (Webmin
2239
+ CVE-2019-15107, Struts2 S2-001, Log4Shell CVE-2021-44228).
2240
+
2241
+ ### Changed
2242
+
2243
+ - **Expect more requests per run.** The separator sweep and the language-runtime
2244
+ fix both widen the probe set. Narrow with `--separators`, `--contexts` and
2245
+ `--environments` once the sink's shape is known.
2246
+ - **`--doctor` output.** Its first line now names the corpus in use
2247
+ (`corpus: …`) rather than a path (`template: …`), since the corpus is no
2248
+ longer necessarily a file, and `[ok] file loaded and parsed` is now
2249
+ `[ok] corpus loaded and parsed`.
2250
+
2251
+ No breaking changes to the CLI, output formats, or the template schema.
2252
+ Standard library only, Python 3.8–3.13.
2253
+
2254
+ ## Earlier releases
2255
+
2256
+ Release notes for these live on the
2257
+ [Releases page](https://github.com/kabiri-labs/rcekit/releases); they predate
2258
+ this file and have not been restated here.
2259
+
2260
+ - **[2.15.2]** — Multi-method RCE detection &amp; confirmation
2261
+ - **[2.7.0]**
2262
+ - **[2.1.0]**
2263
+
2264
+
2265
+
2266
+ [Unreleased]: https://github.com/kabiri-labs/rcekit/compare/v2.36.0...HEAD
2267
+ [2.45.3]: https://github.com/kabiri-labs/rcekit/compare/v2.45.2...v2.45.3
2268
+ [2.45.2]: https://github.com/kabiri-labs/rcekit/compare/v2.45.1...v2.45.2
2269
+ [2.45.1]: https://github.com/kabiri-labs/rcekit/compare/v2.45.0...v2.45.1
2270
+ [2.45.0]: https://github.com/kabiri-labs/rcekit/compare/v2.44.0...v2.45.0
2271
+ [2.44.0]: https://github.com/kabiri-labs/rcekit/compare/v2.43.0...v2.44.0
2272
+ [2.43.0]: https://github.com/kabiri-labs/rcekit/compare/v2.42.0...v2.43.0
2273
+ [2.42.0]: https://github.com/kabiri-labs/rcekit/compare/v2.41.0...v2.42.0
2274
+ [2.41.0]: https://github.com/kabiri-labs/rcekit/compare/v2.40.0...v2.41.0
2275
+ [2.40.0]: https://github.com/kabiri-labs/rcekit/compare/v2.39.0...v2.40.0
2276
+ [2.39.0]: https://github.com/kabiri-labs/rcekit/compare/v2.38.0...v2.39.0
2277
+ [2.38.0]: https://github.com/kabiri-labs/rcekit/compare/v2.37.0...v2.38.0
2278
+ [2.37.0]: https://github.com/kabiri-labs/rcekit/compare/v2.36.0...v2.37.0
2279
+ [2.36.0]: https://github.com/kabiri-labs/rcekit/compare/v2.35.5...v2.36.0
2280
+ [2.35.5]: https://github.com/kabiri-labs/rcekit/compare/v2.35.4...v2.35.5
2281
+ [2.35.4]: https://github.com/kabiri-labs/rcekit/compare/v2.35.3...v2.35.4
2282
+ [2.35.3]: https://github.com/kabiri-labs/rcekit/compare/v2.35.2...v2.35.3
2283
+ [2.35.2]: https://github.com/kabiri-labs/rcekit/compare/v2.35.1...v2.35.2
2284
+ [2.35.1]: https://github.com/kabiri-labs/rcekit/compare/v2.35.0...v2.35.1
2285
+ [2.35.0]: https://github.com/kabiri-labs/rcekit/compare/v2.34.1...v2.35.0
2286
+ [2.34.1]: https://github.com/kabiri-labs/rcekit/compare/v2.34.0...v2.34.1
2287
+ [2.34.0]: https://github.com/kabiri-labs/rcekit/compare/v2.33.0...v2.34.0
2288
+ [2.33.0]: https://github.com/kabiri-labs/rcekit/compare/v2.32.0...v2.33.0
2289
+ [2.32.0]: https://github.com/kabiri-labs/rcekit/compare/v2.31.0...v2.32.0
2290
+ [2.31.0]: https://github.com/kabiri-labs/rcekit/compare/v2.30.0...v2.31.0
2291
+ [2.30.0]: https://github.com/kabiri-labs/rcekit/compare/v2.29.0...v2.30.0
2292
+ [2.29.0]: https://github.com/kabiri-labs/rcekit/compare/v2.28.0...v2.29.0
2293
+ [2.28.0]: https://github.com/kabiri-labs/rcekit/compare/v2.27.0...v2.28.0
2294
+ [2.27.0]: https://github.com/kabiri-labs/rcekit/compare/v2.26.0...v2.27.0
2295
+ [2.26.0]: https://github.com/kabiri-labs/rcekit/compare/v2.25.0...v2.26.0
2296
+ [2.25.0]: https://github.com/kabiri-labs/rcekit/compare/v2.24.0...v2.25.0
2297
+ [2.24.0]: https://github.com/kabiri-labs/rcekit/compare/v2.23.3...v2.24.0
2298
+ [2.23.3]: https://github.com/kabiri-labs/rcekit/compare/v2.23.2...v2.23.3
2299
+ [2.23.2]: https://github.com/kabiri-labs/rcekit/compare/v2.23.1...v2.23.2
2300
+ [2.23.1]: https://github.com/kabiri-labs/rcekit/compare/v2.23.0...v2.23.1
2301
+ [2.23.0]: https://github.com/kabiri-labs/rcekit/compare/v2.22.0...v2.23.0
2302
+ [2.22.0]: https://github.com/kabiri-labs/rcekit/compare/v2.21.1...v2.22.0
2303
+ [2.21.1]: https://github.com/kabiri-labs/rcekit/compare/v2.15.2...v2.21.1
2304
+ [2.15.2]: https://github.com/kabiri-labs/rcekit/releases/tag/v2.15.2
2305
+ [2.7.0]: https://github.com/kabiri-labs/rcekit/releases/tag/v2.7.0
2306
+ [2.1.0]: https://github.com/kabiri-labs/rcekit/releases/tag/v2.1.0