rcekit 2.35.1__tar.gz → 3.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. rcekit-3.0.0/CHANGELOG.md +2747 -0
  2. rcekit-3.0.0/CONTRIBUTING.md +175 -0
  3. {rcekit-2.35.1/rcekit.egg-info → rcekit-3.0.0}/PKG-INFO +243 -62
  4. {rcekit-2.35.1 → rcekit-3.0.0}/README.md +242 -61
  5. {rcekit-2.35.1 → rcekit-3.0.0}/docs/generation.md +10 -0
  6. {rcekit-2.35.1 → rcekit-3.0.0}/docs/guide.md +40 -21
  7. {rcekit-2.35.1 → rcekit-3.0.0}/docs/reference.md +457 -61
  8. {rcekit-2.35.1 → rcekit-3.0.0}/docs/verify-it-yourself.md +19 -8
  9. {rcekit-2.35.1 → rcekit-3.0.0/rcekit.egg-info}/PKG-INFO +243 -62
  10. {rcekit-2.35.1 → rcekit-3.0.0}/rcekit.egg-info/SOURCES.txt +7 -0
  11. {rcekit-2.35.1 → rcekit-3.0.0}/rcekit.py +3095 -246
  12. {rcekit-2.35.1 → rcekit-3.0.0}/templates/payloads.json +33 -0
  13. rcekit-3.0.0/tests/bench/README.md +472 -0
  14. rcekit-3.0.0/tests/bench/cases/fastjson-1.2.83.json +104 -0
  15. rcekit-3.0.0/tests/bench/cases/hugegraph-gremlin-eval.json +69 -0
  16. rcekit-3.0.0/tests/bench/cases/hugegraph-gremlin-shell.json +76 -0
  17. rcekit-3.0.0/tests/bench/cases/log4shell-cve-2021-44228.json +69 -0
  18. rcekit-3.0.0/tests/bench/cases/opentsdb-cve-2023-25826.json +93 -0
  19. rcekit-3.0.0/tests/bench/cases/struts2-s2-001.json +69 -0
  20. rcekit-3.0.0/tests/bench/cases/tomcat-cve-2017-12615.json +105 -0
  21. rcekit-3.0.0/tests/bench/cases/webmin-cve-2019-15107-file.json +83 -0
  22. rcekit-3.0.0/tests/bench/cases/webmin-cve-2019-15107.json +62 -0
  23. rcekit-3.0.0/tests/bench/runner.py +909 -0
  24. rcekit-3.0.0/tests/test_bench_runner.py +1276 -0
  25. rcekit-3.0.0/tests/test_docs.py +1458 -0
  26. {rcekit-2.35.1 → rcekit-3.0.0}/tests/test_generator.py +5200 -349
  27. rcekit-2.35.1/CHANGELOG.md +0 -1042
  28. rcekit-2.35.1/CONTRIBUTING.md +0 -104
  29. rcekit-2.35.1/tests/bench/README.md +0 -138
  30. rcekit-2.35.1/tests/bench/cases/struts2-s2-001.json +0 -46
  31. rcekit-2.35.1/tests/bench/cases/webmin-cve-2019-15107.json +0 -43
  32. rcekit-2.35.1/tests/bench/runner.py +0 -437
  33. rcekit-2.35.1/tests/test_bench_runner.py +0 -368
  34. rcekit-2.35.1/tests/test_docs.py +0 -309
  35. {rcekit-2.35.1 → rcekit-3.0.0}/LICENSE +0 -0
  36. {rcekit-2.35.1 → rcekit-3.0.0}/MANIFEST.in +0 -0
  37. {rcekit-2.35.1 → rcekit-3.0.0}/SECURITY.md +0 -0
  38. {rcekit-2.35.1 → rcekit-3.0.0}/pyproject.toml +0 -0
  39. {rcekit-2.35.1 → rcekit-3.0.0}/rcekit.egg-info/dependency_links.txt +0 -0
  40. {rcekit-2.35.1 → rcekit-3.0.0}/rcekit.egg-info/entry_points.txt +0 -0
  41. {rcekit-2.35.1 → rcekit-3.0.0}/rcekit.egg-info/top_level.txt +0 -0
  42. {rcekit-2.35.1 → rcekit-3.0.0}/setup.cfg +0 -0
  43. {rcekit-2.35.1 → rcekit-3.0.0}/tests/bench/requests/webmin.txt +0 -0
  44. {rcekit-2.35.1 → rcekit-3.0.0}/tests/test_corpus_embedding.py +0 -0
  45. {rcekit-2.35.1 → rcekit-3.0.0}/tools/embed_corpus.py +0 -0
@@ -0,0 +1,2747 @@
1
+ # Changelog
2
+
3
+ All notable changes to RCEKit are recorded here. The format follows
4
+ [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and RCEKit follows
5
+ [Semantic Versioning](https://semver.org/spec/v2.0.0.html): PATCH for fixes,
6
+ MINOR for new capabilities, MAJOR for breaking changes to the CLI, output
7
+ formats, or the template schema.
8
+
9
+ ## [Unreleased]
10
+
11
+ ### Added
12
+
13
+ - **A coverage ledger, so a measured result and a reproduced one stop looking
14
+ alike.** The README's CVE table had four rows and two columns that mattered,
15
+ and no way to say what stood behind any of them. It gained seven columns and
16
+ five more rows, and the two on the right are the point: `Bench case` says
17
+ whether `tests/bench/` reproduces the row under Docker with its negative
18
+ control, and `Recording` says whether a file in `confirmation-gifs/` shows it.
19
+
20
+ Five rows are new, measured against vulhub builds during this work: OpenTSDB
21
+ 2.4.1 (CVE-2023-25826) reaching `confirmed` through `oob` and `needs-review`
22
+ through `time`; HugeGraph 1.2.0 reaching `confirmed` through both `eval` and
23
+ `reflected`; and a Spring Boot application on fastjson 1.2.83 reaching
24
+ `deserialization-sink`. All five landed saying `not yet` under `Bench case`,
25
+ which was the honest state at the time and the reason that column exists; each
26
+ has a case now. The 4 rows that predate this work had last been executed as a
27
+ set at 2.36.0, and none of them regressed. The ledger carries the figures for
28
+ the whole set, because a running total written here rots on the next change
29
+ rather than the next release.
30
+
31
+ `Advisory` is empty on three of them, deliberately. HugeGraph 1.3.0 answers
32
+ the arithmetic exactly as 1.2.0 does -- its Gremlin API evaluates Groovy
33
+ unauthenticated by design -- and fastjson resolving an `Inet4Address` is
34
+ documented autoType behaviour. Those rows prove a *method* against real
35
+ software, which is worth recording; calling them CVE reproductions would be
36
+ the overclaim the table exists to avoid. Both facts were measured rather than
37
+ assumed: the patched HugeGraph image was pulled and run.
38
+
39
+ Three tests hold the ledger to the repository: a row claiming a bench case
40
+ must have one, a row claiming a recording must have one, and the capability
41
+ rows must stay distinguishable from the CVE rows. Each was checked by making
42
+ the false claim and watching it fail. The tier check now reads *every* row
43
+ rather than only those naming an advisory, and its parser is keyed on column
44
+ headings rather than position -- the old one was hard-coded to four columns
45
+ and found nothing at all once the table grew.
46
+
47
+ - **A bench case for `write`, the last tier-1 method without one.**
48
+ `tomcat-cve-2017-12615` -- `PUT /rcekit-probe.jsp/` with the content in the
49
+ body, which is the shape the method's own docstring names. Its control runs
50
+ `reflected` against the same point and must stay `negative`: the target is
51
+ exploitable and `write` confirms on it, but a PUT answers 204 with an empty
52
+ body, so nothing in the vulnerable response is computed. Measured at 959
53
+ probes, every one answered 204.
54
+
55
+ It is the first case here that may not share a target between its halves --
56
+ both write a file into the web root -- which turned up a harness bug: under
57
+ `--keep-up` the teardown *between* the halves was skipped too, and compose
58
+ handed the control the container the vulnerable half had just written to. That
59
+ teardown is mandatory now wherever the halves would land on the same
60
+ container, which is decided from what *starts* a target -- its `vulhub_path`
61
+ and `compose` -- and not from what stops it. A control overriding only
62
+ `compose_down`, to keep its volumes say, still meets the container the
63
+ vulnerable half brought up. A control that brings up its own target has
64
+ nothing to contaminate, so `--keep-up` leaves the vulnerable one running
65
+ there, which is the environment worth looking at after a failure. The
66
+ teardown after the last half stays the flag's to skip either way.
67
+
68
+ The case is also what caught the two `write` defects fixed in 2.45.6 and
69
+ 2.45.7, neither of which a fixture could have found.
70
+
71
+ ### Fixed
72
+
73
+ - **The comparison table named seven of the nine registered methods.** `lookup`
74
+ and `boolean` had no row in "How RCEKit compares", although `lookup` is the
75
+ Log4Shell row the CVE table directly above it already carries, and `boolean`
76
+ shipped two releases ago. The table's claim is coverage -- "every class, one
77
+ run" -- so a method missing from it is that claim understating itself.
78
+
79
+ The drift was structural, not an oversight anyone could have caught by
80
+ reading. `docs/reference.md`, the `--methods` help, the CVE rows and the demo
81
+ headings are each pinned to `DETECTION_METHODS` by a test. The comparison
82
+ table was pinned by nothing, so it was the one place a new method could fail
83
+ to appear without anything going red.
84
+
85
+ Two tests now hold it, in both directions: every registered method has a row,
86
+ and the table names no method the CLI would reject -- because a one-way check
87
+ rots. A method *removed* from the registry would otherwise leave its row
88
+ behind, advertising a `--methods` value that no longer exists, while the
89
+ completeness test went on passing.
90
+
91
+ Each row now names its method in backticks, which is what makes the check
92
+ possible and also answers the question the table used to leave hanging: a
93
+ reader who wants the row they just read has the flag in front of them. Rows
94
+ that are not methods name none on purpose -- second-order execution, the
95
+ query-language bridges and the per-dialect Windows probes are things the
96
+ methods run *through*, not entries in `--methods`.
97
+
98
+ Both checks read the capability table alone, and identify a method label
99
+ structurally: the last backtick in the leading cell, after an em-dash. The
100
+ first attempt did neither, and both shortcuts cost it the guarantee it stated.
101
+ Reading every table in the section meant a deleted capability row kept passing
102
+ on a mention elsewhere -- the engagement-controls table names `file` and
103
+ `write` under "Cleanup commands" -- which was confirmed by deleting the `file`
104
+ row and watching the check stay green. And filtering labels on lowercase
105
+ letters exempted exactly the spellings a typo or a rename produces:
106
+ `boolean2`, `second-order`, `reflected_v2` all walked past the phantom check,
107
+ while `cmd.exe` is the legitimate backtick in that column that the filter had
108
+ been reaching for. Structure separates them; a character class does not.
109
+
110
+ No version bump -- this is documentation and tests only.
111
+
112
+ ### Added
113
+
114
+ - **A tier correction has to reach the prose, not only the class.** Four tests
115
+ now read the tier from `DETECTION_METHODS[name].tier` and hold the sentences
116
+ an operator actually sees against it.
117
+
118
+ `lookup` moved from `confirmed` to `lookup-sink` on the class, and the move
119
+ was made in the README's CVE table -- but `blind_sink_advice` went on
120
+ offering `--methods lookup` as a method that "confirms", in the list where
121
+ `oob` and `file` do mean confirmed execution, and the README's Log4Shell demo
122
+ heading still said `confirmed` over alt text calling the run
123
+ "auto-confirming a blind Log4Shell RCE". Three places, one correction, and
124
+ nothing compared them: the existing test pinned the one line it was written
125
+ for (`time` is marked needs-review only) rather than asking every line the
126
+ same question.
127
+
128
+ So: every `blind_sink_advice` line naming a method must state that method's
129
+ tier and no other; every CVE row must state the tier its method reports;
130
+ every demo heading must match its row in that table; and a recording below
131
+ `confirmed` may not be described as confirming, alt text included, since
132
+ that is the sentence a screen reader reads out.
133
+
134
+ The wording checks match stems rather than words, and denials such as
135
+ "never confirmed" are removed before they are applied. "confirms" alone
136
+ would have readmitted the overclaim through "confirming RCE" or "confirmed
137
+ execution" -- an assertion answering the same way for the right reason and
138
+ the broken one, which is the defect being guarded against. Each test was
139
+ checked by reintroducing the defect it exists for, and each carries a floor
140
+ on how much it inspected: a parser that quietly stopped matching would
141
+ otherwise pass exactly as a clean README does.
142
+
143
+ Test-only: no version bump, and nothing about a run changes.
144
+
145
+ - **The same check, for the two tables under `docs/`.** They were left out when
146
+ the README's were pinned, and they carry the same claim:
147
+ `docs/reference.md` names the tier each method can reach, and
148
+ `docs/guide.md` tells an operator which method to reach for next.
149
+
150
+ `reference.md` states a ceiling, so its rule is a subset rather than an
151
+ equality -- a cell may also name a weaker tier the method really emits, as
152
+ `write` and `deser` both do -- and every registered method must have a row,
153
+ so a capability cannot land unlookupable. `guide.md` has no tier column and
154
+ so is held only to the negative: a row recommending a method that cannot
155
+ confirm may not describe confirmation.
156
+
157
+ Rows are read by column heading rather than by position. Reading the whole
158
+ row made the `guide.md` check skip the one row it was written for, because
159
+ that row's prose names `oob` -- which confirms -- while recommending
160
+ `lookup`, which does not.
161
+
162
+ A reference row naming something that is not a registered method fails rather
163
+ than being passed over. Skipping it left the completeness check one-way: a
164
+ method *removed* from `DETECTION_METHODS` would leave its row behind, every
165
+ remaining row would still match, and the page would go on offering a
166
+ `--methods` value the CLI rejects.
167
+
168
+ Rows split on *unescaped* pipes. Markdown writes a literal pipe in a cell as
169
+ `\|`, which `docs/reference.md` already does in three tables, and splitting
170
+ on every pipe invents a cell: zipping against the header then drops the last
171
+ column outright, so a claim there stops being examined and every one of these
172
+ checks passes without looking at it.
173
+
174
+ A disclaimer is not read as a claim. "never confirmed" was already removed
175
+ before the `confirm` stem was looked for; "without confirmation" and
176
+ "unconfirmed" were not, and the second carries its negation inside the word
177
+ where a rule about preceding words cannot see it. Documentation saying the
178
+ honest thing would have failed the suite. One parser and one denial rule now,
179
+ where there were two of each here and a third in the generator tests.
180
+
181
+ ### Changed
182
+
183
+ - **A bench case may bring its target up once for both halves**, with
184
+ `"share_target": true`. Bringing the container up twice is the largest fixed
185
+ cost in a case and both halves usually hit the same one, so on a fast case it
186
+ is most of the wall clock.
187
+
188
+ It is opt-in, and the default is unchanged, because the teardown between the
189
+ halves is `down -v`: today's control meets a **fresh** target. A case whose
190
+ vulnerable half writes a file, plants a shell or changes a setting would hand
191
+ its control a target it had already altered, and a control measured against a
192
+ contaminated target measures nothing -- which is the one failure a benchmark
193
+ may not have. Validation rejects the key on a case whose control brings up a
194
+ different target, since there is then nothing to share and leaving it set
195
+ would read as though there were.
196
+
197
+ A shared `up` that fails falls back to per-half management rather than
198
+ carrying on, so the run reports `compose up failed` instead of two readiness
199
+ timeouts naming the wrong cause.
200
+
201
+ Both shipped cases set it, because neither half writes anything. Measured on
202
+ `struts2-s2-001` against vulhub on Docker: 33.8s to 22.9s, with both halves
203
+ reaching the same verdicts either way.
204
+
205
+ - **The benchmark reaches a callback method**, with a case for Log4Shell
206
+ (Apache Solr 8.11.0, Log4j 2.14.1). `python tests/bench/runner.py --all` is
207
+ 3/3.
208
+
209
+ Every case until now was in-band, so the listener, the token correlation and
210
+ `confirm_each` had never run against real software -- the area that produced
211
+ three P1 findings while `lookup` was being written.
212
+
213
+ Two things had to give. A JNDI lookup resolves through the system resolver,
214
+ which asks UDP 53, and a developer machine rarely has that port free; a bare
215
+ `--oob-host` IP is no way around it, because a lookup has no second channel
216
+ to carry the token. So a case may now name `run_in`, and the run happens in a
217
+ container on the target's own network at a fixed address, with an override
218
+ pointing the service's `dns:` at it. The repository is mounted read-only and
219
+ the image is a stock Python.
220
+
221
+ The other was the harness's own vocabulary. `VALID_EXPECTATIONS` and
222
+ `CONTROL_EXPECTATIONS` were written out by hand and had drifted: neither
223
+ `lookup-sink` nor `deserialization-sink` was in either, so a case for `lookup`
224
+ or `deser` could not be *loaded*, let alone run. Both are read from
225
+ `DETECTION_METHODS` now.
226
+
227
+ The control is the case: `oob` against the same Solr comes back `negative`
228
+ **although it is exploitable**, because every probe it builds is a shell
229
+ command and a `${jndi:...}` sink runs none of them. That is the gap `lookup`
230
+ was added to close, and until this case ran it rested on a fixture.
231
+
232
+ `{bench}` / `{repo}` now expand in a case's compose argv as well as its
233
+ invocation, and resolve to the mount point when the run is containerised.
234
+
235
+ The results file is created before the container starts, and it rather than
236
+ its directory is made writable. Under Docker's user-namespace remapping,
237
+ container root is a subordinate host UID, so a `mkdtemp` owned by the runner
238
+ at 0700 is not writable from inside: the file would never appear and the case
239
+ would report `nothing-tested`, as though detection had found nothing rather
240
+ than as though the channel had been shut. Granting the one file and not the
241
+ directory keeps anyone else from creating, replacing or unlinking entries
242
+ there.
243
+
244
+ Test-only: no version bump, and nothing about a run changes.
245
+
246
+ ### Fixed
247
+
248
+ - **`--methods` enumerated only some of the methods it accepts.** `write`,
249
+ `lookup` and `deser` were registered in `DETECTION_METHODS` and had never
250
+ once been named in the help, so `--help` described a whole target class as
251
+ out of reach -- an upload that stores a file, a `${jndi:...}` sink, an
252
+ endpoint that deserializes what it is handed -- while the method for it was
253
+ already shipping and documented everywhere else. Each now carries the tier
254
+ its class declares: `confirmed` for `write`, with `needs-review` for a write
255
+ that is served but not interpreted; `lookup-sink` for `lookup`;
256
+ `deserialization-sink` for `deser`; and neither of the last two ever
257
+ `confirmed`.
258
+
259
+ `oob` also called itself "the only confirmed-tier method for a fully blind
260
+ sink". The tier is right, but an exclusivity claim is the kind that goes
261
+ stale without anything failing, so the line now says what the class itself
262
+ says: it confirms a sink with no output channel and no writable web root.
263
+
264
+ A new test in `CLIDocumentationTestCase` holds the list to the registry, the
265
+ way `--eval-engines` is already held to the corpus. It reads the `--methods`
266
+ help alone rather than the whole page, because every one of these names also
267
+ occurs inside some other flag -- `file` in `--request-file`, `write` in
268
+ `--file-write-path`, `time` in `--time-base` -- and it looks for the name
269
+ followed by the parenthesis that opens its description, because the old help
270
+ carried the word `write` inside `file`'s "write+read-back" and a looser
271
+ match would have counted that as documentation.
272
+
273
+ A second pass corrected what `--oob-host` promises those two methods. Its
274
+ help offered "an IP it can reach or a domain delegated to this listener" and
275
+ was scoped `(--methods oob)`, which was true when `oob` was the only caller.
276
+ It is not true now: `LookupCallback.build_probes` returns no probes at all
277
+ for an address literal, so `lookup` reports `nothing-tested`, and `DeserSink`
278
+ builds `<token>.<IP>`, which resolves nowhere, so its DNS gadget can never
279
+ reach `deserialization-sink`. An operator who satisfied the stated
280
+ prerequisite got silence from one method and a capped tier from the other.
281
+ The flag's help, the `--methods` entries for `lookup` and `deser`, the
282
+ message printed when a callback method is selected without a host, and the
283
+ `deser` oracle table in `docs/reference.md` now all say the same thing: an IP
284
+ serves `oob` alone, because only `oob` can put its token in a URL path.
285
+
286
+ `boolean` landed on `main` while this branch was open. It was named in the
287
+ help but left out of the `--methods` row in `docs/reference.md` -- the same
288
+ drift, one table over. Both carry it now, and the new test covered it without
289
+ being told to, which is the point of reading the registry instead of a list.
290
+
291
+ No version bump -- this is documentation and tests only.
292
+
293
+ ## [3.0.0] — 2026-09-30
294
+
295
+ ### Changed
296
+
297
+ - **BREAKING: `confirmed` is now `executed`, and three settled measurements
298
+ stopped reporting as `needs-review`.** Every verdict is now named for what the
299
+ *target* did, never for how sure RCEKit is.
300
+
301
+ Two problems, one vocabulary. `confirmed` was the only verdict that described
302
+ RCEKit's own confidence while every sibling described the target's behaviour,
303
+ and next to it a `lookup-sink` -- proven on a callback carrying a token only
304
+ that probe held -- read as something unconfirmed. And `needs-review` had four
305
+ users, of which three had measured their answer:
306
+
307
+ | Verdict | Was | What it says |
308
+ |---|---|---|
309
+ | `executed` | `confirmed` | the target executed the input |
310
+ | `timing-sink` | `needs-review` | the target honoured a delay RCEKit injected |
311
+ | `file-write` | `needs-review` | the target stored a file at a path RCEKit chose and serves it back uninterpreted |
312
+ | `evaluation-sink` | `needs-review` | an evaluator consumed the input and partitioned on it; which evaluator is not shown |
313
+ | `needs-review` | `needs-review` | unchanged, and now with one user: `deser`'s parser fingerprint, the one signal that really is a candidate |
314
+
315
+ `write`'s middle verdict is the clearest case: its evidence line opened with
316
+ the words "ARBITRARY FILE WRITE confirmed" under a status that said
317
+ `needs-review`. The tool's own prose contradicted the tool's own status,
318
+ because the status was naming a confidence the measurement did not have a
319
+ reason to lack -- the file was on the target and RCEKit had read it back.
320
+
321
+ Nothing was promoted. `executed` still rests on exactly the three channels it
322
+ did before, and none of the five weaker tiers has a path into it. `time` in
323
+ particular stays out: a delay honoured is not necessarily a *shell* honouring
324
+ it, since a sandbox that implements `sleep` answers the regression the same
325
+ way. What changed is the direction of the reporting error -- three proven
326
+ findings were being under-reported.
327
+
328
+ - **`overall_detection_verdict` ranks `needs-review` last instead of second.**
329
+ Its position was argued from "a suspected RCE outranks a proven non-RCE in
330
+ triage", which held while `time`, `boolean` and `write`'s uninterpreted file
331
+ all reported there. The one verdict left under it is a suspected
332
+ *deserialization*, so ranking it above `deserialization-sink` put the guess
333
+ above the proof of the same thing. The full order is now `executed`,
334
+ `timing-sink`, `file-write`, `evaluation-sink`, `deserialization-sink`,
335
+ `lookup-sink`, `needs-review`.
336
+
337
+ That order is stated once, in `NON_EXECUTION_SECTIONS`, and both the
338
+ collapsed verdict and the text report read it from there. They used to state
339
+ it separately -- one as a tuple, the other as the order of print statements
340
+ in `main()` -- and they disagreed: `deserialization-sink` and `lookup-sink`
341
+ were hand-written blocks above the table holding the rest, so a run
342
+ reporting both `timing-sink` and `deserialization-sink` collapsed to the
343
+ first and led the report with the second. A statement order is not data and
344
+ cannot be checked against anything.
345
+
346
+ Folding those two in unifies their heading with the others: `NOT proof of
347
+ RCE` becomes `NOT proof of execution`, which is what the other four already
348
+ said. Their rows keep the compact form they had -- truncated payload, no
349
+ `environment` -- because they carry base64 object streams and JNDI URIs.
350
+
351
+ - The classic `--verify-url` path renamed its `confirmed` verdict too, and its
352
+ section headings with it. It has no sinks and no `needs-review`, so nothing
353
+ there was mis-tiered; a tool printing `[verify] CONFIRMED` beside `[detect]
354
+ EXECUTED` would just be two vocabularies. That path has no machine-readable
355
+ output, so this is text only.
356
+
357
+ ### Benchmark
358
+
359
+ **`runner.py --all`: 9/9, in 55m50s**, against the same vulhub builds. Every
360
+ case's expected verdict changed in this release, and a rename applied to a
361
+ stored result proves nothing about what the tool now reports, so all 9 ran:
362
+
363
+ | Case | Method | Vulnerable half | Negative control |
364
+ |---|---|---|---|
365
+ | `fastjson-1.2.83` | `deser` | `deserialization-sink` | `negative` |
366
+ | `hugegraph-gremlin-eval` | `eval` | `executed` | `negative` |
367
+ | `hugegraph-gremlin-shell` | `reflected` | `executed` | `negative` |
368
+ | `log4shell-cve-2021-44228` | `lookup` | `lookup-sink` | `negative` |
369
+ | `opentsdb-cve-2023-25826` | `oob` | `executed` | `timing-sink` |
370
+ | `struts2-s2-001` | `eval` | `executed` | `negative` |
371
+ | `tomcat-cve-2017-12615` | `write` | `executed` | `negative` |
372
+ | `webmin-cve-2019-15107-file` | `file` | `executed` | `negative` |
373
+ | `webmin-cve-2019-15107` | `reflected` | `executed` | `timing-sink` |
374
+
375
+ The two `timing-sink` controls carry the weight. On a Webmin 1.910 and an
376
+ OpenTSDB 2.4.1 that are genuinely vulnerable, `time` reaches its own proven
377
+ tier and still does not reach `executed` -- the retier holding against real
378
+ software and not against a fixture.
379
+
380
+ Two tiers this release introduced are **not** in that table: `file-write` and
381
+ `evaluation-sink`. The Tomcat `write` case reaches `executed`, and `boolean`
382
+ has no case at all. Both are covered by unit tests only.
383
+
384
+ ### Migration
385
+
386
+ `--detect-json` changes values, not keys. `verdict`, the keys of `counts`, each
387
+ probe's `verdict`, and `observe_status` take the new names per the table above.
388
+ A consumer keyed on `"confirmed"` must read `"executed"`; one keyed on
389
+ `"needs-review"` should decide which of the four it actually meant. No CLI flag,
390
+ no JSON key, no template-schema field and no exit code changed.
391
+
392
+ No backward-compatible alias is shipped. Two accepted spellings for one field is
393
+ a second source of truth, and the point of this release is that there is one.
394
+
395
+ ### Fixed
396
+
397
+ - **A method's question was derived from its tier, which does not determine
398
+ it.** `detection_question` decided whether a method asks *did the target
399
+ execute my input* by looking up its tier in a set of two names. A tier says
400
+ how strong an answer a method can reach; `time` asks the execution question
401
+ and cannot reach the strongest answer to it, while `lookup` reaches a
402
+ definitive answer to a different question.
403
+
404
+ Left alone, renaming `time`'s tier in this release would have moved it out of
405
+ the execution group and out of the dedup that skips it on a candidate
406
+ `reflected` has already proven -- so the most expensive method in the tool
407
+ would have run to put a second name on a finding. `question` is now declared
408
+ on the class, `EXECUTION_TIERS` is gone, and the partition is unchanged.
409
+
410
+ The test covering this recomputed the derivation rule it was checking, so it
411
+ agreed with the rule whatever the rule became. Three tests replace it: the
412
+ partition written out, the one case where question and tier disagree, and the
413
+ `time`-leaves-the-group regression by name.
414
+
415
+ ## [2.45.7] — 2026-09-26
416
+
417
+ ### Fixed
418
+
419
+ - **The `write` product overflowed a 32-bit int, so the method reported
420
+ `negative` on more than half of the targets it was built for.** RCEKit
421
+ computes `a*b` in Python, where integers do not overflow. The interpreter on
422
+ the other side usually is not Python: JSP and ASPX are 3 of the 5 languages
423
+ this method writes, and both evaluate the expression as a signed 32-bit int.
424
+ A product past 2**31-1 therefore came back wrapped, disagreed with the value
425
+ RCEKit had computed, and the run reported `negative` against a target that is
426
+ fully exploitable -- the one failure this tool exists not to produce.
427
+
428
+ Measured against Tomcat 8.5.19 (CVE-2017-12615), writing the probe by PUT and
429
+ fetching it back. `97233*38786` was served as `-523688158` where RCEKit
430
+ expected `3771279138`; `4721*8093` came back as `38207053`, exactly. The
431
+ operand range drew a pair whose product overflows in **56%** of runs, and the
432
+ operands are drawn once per run, so there was no second probe to recover: more
433
+ than half of all runs against a JVM or .NET target lost.
434
+
435
+ The operands are now 4 digits apiece, which caps the product at 99,980,001 --
436
+ the range `eval` has always used against these same targets. Nothing about the
437
+ oracle changes: the product is still bracketed by two random tags and still
438
+ differenced against a payload-free control, which is what makes it unforgeable
439
+ rather than its magnitude.
440
+
441
+ The same run confirms now. `RKYBRDH<%=3253*8793%>RKICSMR` came back as
442
+ `RKYBRDH28603629RKICSMR` from the fetched file, and `write` reached
443
+ `confirmed` against real software for the first time.
444
+
445
+ Two tests, both written first and watched to fail: the drawn operands multiply
446
+ to something a signed 32-bit int can hold, across 400 seeds; and a store whose
447
+ evaluator wraps to 32 bits the way a JVM does reaches `confirmed` rather than
448
+ `negative`. A third test had pinned the operand width with a hard-coded
449
+ `\d{5}`; it reads the declared range now and asserts the operands fall inside
450
+ it, which the digit count had only implied.
451
+
452
+ ## [2.45.6] — 2026-09-26
453
+
454
+ ### Fixed
455
+
456
+ - **The `write` probe carried whitespace, and some sinks tokenise before they
457
+ write.** The file content this method delivers is a language one-liner, and
458
+ its templates spelled the expression with spaces around it -- `<?= a*b ?>`,
459
+ `<%= a*b %>`. The content goes through the operator's own injection point, and
460
+ a sink that splits its input on whitespace never writes a valid file.
461
+
462
+ Measured against RRDtool 1.7.2, reached through Cacti 1.2.28's
463
+ `right_axis_label` (CVE-2025-24367), which builds the file from a
464
+ `LINE1:out:<content>` argument. `RKAAA<?=7*6?>RKBBB` wrote a 47-byte file with
465
+ the content intact. `RKAAA<?= 7*6 ?>RKBBB` wrote no file at all and answered
466
+ `ERROR: '7*6' is not a valid function name in 7*6`. The whitespace was the
467
+ whole difference.
468
+
469
+ Whitespace inside `<?= ?>` and `<%= %>` is optional, so the templates drop it:
470
+ nothing is lost and every sink of that shape is gained. Verified end to end
471
+ against the real sink rather than a fixture -- RCEKit generated
472
+ `RKYOOQS<?=69294*83370?>RKGFZQP`, RRDtool wrote it whole, and PHP answered
473
+ `RKYOOQS5777040780RKGFZQP`, which is the value RCEKit had computed for that
474
+ probe.
475
+
476
+ `jspx` keeps its spaces: it is an XML document whose root element carries
477
+ namespace attributes, and a sink that tokenises on whitespace was never going
478
+ to carry one. Two tests pin the rest -- the generated probe and the templates
479
+ themselves -- both written first and watched to fail.
480
+
481
+ ## [2.45.5] — 2026-09-25
482
+
483
+ ### Fixed
484
+
485
+ - **The `deser` shape differential answered from evidence it had not
486
+ validated.** Three ways, each found by pointing the method at a live
487
+ fastjson 1.2.83 endpoint under vulhub and asking why it said what it said.
488
+
489
+ **Drift landed on whichever form went last.** The differential is read across
490
+ requests, so anything changing with the request *index* rather than with the
491
+ payload -- a rate limiter backing off, a filling log, a warming cache -- fell
492
+ on the form sent last, and noise was sent last every time. There are two
493
+ noise probes now, first and last, bracketing the structured pair. They are
494
+ one noise body sent twice, byte for byte, because the bracket asks whether
495
+ the endpoint answered *the same probe* alike and anything the target can see
496
+ a difference in it may answer a difference to: a tail beginning `n` made
497
+ fastjson say "error parse new" where every other tail said "syntax error".
498
+ A carrier whose two ends disagree is held at `inconclusive`. The engine
499
+ de-duplicates by payload on the per-probe path but not on the aggregate path
500
+ where `deser` lives, which is what makes one body twice possible; a test pins
501
+ that, because the bracket becomes a single request if it ever changes and a
502
+ drift check comparing an answer with itself passes forever.
503
+
504
+ **An endpoint echoing its input looked like a parser.** A response carrying a
505
+ fragment of the payload that begins at the format's magic, and not carrying
506
+ it in the payload-free control, is an echo and the shape channel cannot be
507
+ read through it -- `inconclusive`, not `negative`, because the channel was
508
+ unreadable rather than empty. The sentinels are anchored on the magic's
509
+ position in the payload rather than its start, since a probe is wrapped for
510
+ its injection context first and some wrappers are longer than the magic: an
511
+ `xml_cdata` pickle probe begins `<![CDATA[gASV`. They are matched raw, case
512
+ folded (because `_signature` folds case and the two halves of one oracle must
513
+ agree), and encoded -- base64 aligned on bytes rather than characters and
514
+ read from the first three-byte boundary at or past the magic, hex from the
515
+ magic itself.
516
+
517
+ **A format declaring no magic got a verdict anyway.** Without a magic there
518
+ is no sentinel, so an echoing endpoint cannot be told from a parsing one --
519
+ and the differential answers regardless, because the noise form is random
520
+ where a structured pair is not. Every shipped format declares one; a
521
+ `--template-file` need not, and the answer to a control that cannot be
522
+ validated is to say so.
523
+
524
+ - **The page an operator plans traffic from still described three payloads.**
525
+ The bracket makes the shape oracle send four per ecosystem -- 20 requests per
526
+ carrier rather than 15, with the five the corpus ships -- and
527
+ `docs/reference.md` and the method's own docstring both still said three.
528
+ README advertises an execution plan that prints the exact probe count before
529
+ anything fires, so a page that undercounts it breaks that promise in the one
530
+ place someone reads before a monitored engagement.
531
+
532
+ Both now state four, and the arithmetic, and what the repeated probe is for.
533
+ A test holds the page to `SHAPE_FORMS` and to the ecosystem count, in both
534
+ directions: adding a form without updating the page fails it, and reverting
535
+ the page without touching the code fails it too. Nothing compared the two
536
+ before, which is how the count drifted the moment the form was added.
537
+
538
+ - **The `negative` evidence line described a comparison it had not made.** "The
539
+ endpoint answers all three forms alike" was asserted rather than observed.
540
+ It now names what actually collapsed: every form alike, a truncated stream
541
+ answered as a complete one, or the noise form answered as a structured one.
542
+
543
+ ### Withdrawn before release
544
+
545
+ - **A second route through the differential, reading noise as the control.** It
546
+ was written for the shape fastjson has -- a format resolving its type name
547
+ before the parse finishes answers a truncated stream exactly as it answers a
548
+ complete one, which collapses the original comparison on the very endpoints
549
+ it exists to find -- and it did reach `needs-review` there.
550
+
551
+ It is not in this release. "Noise differs from the structured pair" has too
552
+ many explanations that are not a parser. Drift and echo are closed above. The
553
+ third is not: a filter that refuses the random payload while letting the
554
+ structured ones through produces the same differential, no repetition exposes
555
+ it because the refusal is stable, and this verdict cannot see it -- refusal
556
+ is judged against the probe the verdict is anchored to, and that is the
557
+ well-formed one.
558
+
559
+ A format of that shape is reported `negative` by the shape differential and
560
+ reaches `deserialization-sink` through the DNS gadget instead, which is this
561
+ method's real oracle and rests on a callback rather than on the shape of a
562
+ response. Measured on the same target: 14 callbacks, each carrying a token it
563
+ could only have learned by resolving the name it was handed.
564
+
565
+ ## [2.45.4] — 2026-09-25
566
+
567
+ ### Fixed
568
+
569
+ - **A compressed response body is a body the oracles cannot read, so a target
570
+ that executed the probe was reported `negative`.** Every in-band oracle --
571
+ `reflected`, `eval`, `file`, `write`, `boolean` -- confirms by finding a
572
+ value the target computed in the response. Nothing in the delivery layer
573
+ undid `Content-Encoding`, so a gzipped body arrived as bytes that *contain*
574
+ the computed value and do not spell it, and the search for a decimal number
575
+ came back empty. The verdict was `negative`: not "could not read this", but
576
+ "the probes reached the target and it is not vulnerable".
577
+
578
+ Measured against Apache HugeGraph 1.2.0 under vulhub. The probe
579
+ `echo RKPDSWY$((258675+956113))RKIEMTB$(echo RKLUQKL)RKPDSWY`, delivered
580
+ through the Gremlin API, came back HTTP 200 carrying
581
+ `RKPDSWY1214788RKIEMTBRKLUQKLRKPDSWY` -- the sum and the substitution
582
+ collapse, both of `ReflectedMath`'s proofs. Present after gunzip, absent in
583
+ the raw bytes, and the run reported `negative` across 620 probes.
584
+
585
+ What made it survive this long is that the same target confirmed on the same
586
+ run's *other* probes: HugeGraph gzips its 200s and leaves its 4xx errors
587
+ plain, so `eval` confirmed through a 400 that carried a Groovy exception
588
+ message while the 200 that carried real command execution was invisible. A
589
+ run that finds something is a run nobody rereads.
590
+
591
+ Both HTTP entry points decode now -- the verification/detection path and the
592
+ `--verify-chain` path -- through one helper, on the success body and the
593
+ `HTTPError` body alike, since an evaluator surfacing its value in a 500 is
594
+ the case that path exists for. The header is honoured whether or not RCEKit
595
+ asked for it, because HugeGraph compresses a response to a request that sent
596
+ no `Accept-Encoding` at all.
597
+
598
+ `gzip` and `deflate` only, both from the standard library, and `deflate`
599
+ tries the zlib wrapper before the raw stream because servers disagree about
600
+ which one that name means. `br` and `zstd` need a third-party module and this
601
+ tool has none to add; a body in either is returned as its raw bytes. Nothing
602
+ here raises: a body that claims an encoding it is not in comes back
603
+ undecompressed rather than ending a run that has already spent hundreds of
604
+ probes.
605
+
606
+ This does not widen what `confirmed` means. The oracle is unchanged -- a
607
+ computed value present in the response and absent from a payload-free
608
+ control -- and nothing is promoted into it. What changes is that the response
609
+ the oracle reads is the response the target sent.
610
+
611
+ - **Decompression is bounded, because compression is where the target chooses
612
+ how much memory RCEKit allocates.** Reading a body at all is new here, and
613
+ one-shot decompression handed a blank cheque to the one party in the exchange
614
+ that is untrusted by definition -- the thing being tested. Measured: 130,478
615
+ bytes of gzip expand to 134,217,728, and 203,860 bytes hold 200 MB of zeros,
616
+ a ratio of 1:1028. This path runs once per probe.
617
+
618
+ The ceiling is 32 MiB per encoding step. Whatever was produced before it is
619
+ kept rather than discarded: the value the oracle wants is short, so searching
620
+ 32 MB is a far better answer than searching nothing, and the only risk the
621
+ ceiling carries is a false negative -- never a false `confirmed`. A truncated
622
+ read is logged rather than passed off as a clean one, since a `negative`
623
+ decided on a partial body is not the same claim as a `negative` decided on
624
+ the whole of it.
625
+
626
+ Multi-member streams are still followed. `gzip.decompress` followed them, and
627
+ a ceiling that quietly became a truncation for a server that concatenates
628
+ members would be the same false negative in a new place.
629
+
630
+ ## [2.45.3] — 2026-09-24
631
+
632
+ ### Fixed
633
+
634
+ - **A `deser`-only run started no listener, so its gadgets went out with
635
+ nothing to receive them.** `--methods deser --oob-host <delegated-name>`,
636
+ with neither `oob` nor `lookup` selected, built its two DNS-gadget probes and
637
+ sent them to a target while no listener was running -- with the host
638
+ configured perfectly. `deser` reached `deserialization-sink` only when it
639
+ happened to be selected beside a method that required a host; run on its own,
640
+ the way the docs present it, its strongest oracle was structurally dead and
641
+ it capped at `needs-review` without saying so.
642
+
643
+ The gate read `needs_oob_host`, which was answering two questions at once:
644
+ "should the run stop when no host is named" and "does this run need a
645
+ listener". Every other method answers both the same way. `deser` is where
646
+ they come apart -- its shape oracle proves something with no listener at all,
647
+ so a missing host must not stop the run, while its DNS gadget does call back
648
+ and needs one.
649
+
650
+ So the second question gets its own attribute, `uses_oob_host`, and the
651
+ listener starts for any selected method that calls back once a host is named.
652
+ It is declared rather than derived: `needs_oob_host or oob_needs_dns_label`
653
+ gives the right answer for every method shipping today and would miss one
654
+ that called back over HTTP without requiring a host -- which is the shape
655
+ `oob` itself would have if its host were optional.
656
+
657
+ Two notices could now reach a run with no HTTP shape in it, and said "only
658
+ the HTTP shapes are live" as though `oob`'s channel were the only one. They
659
+ say "only the shapes that need no DNS are live", which is true of `oob`'s
660
+ HTTP probes and of `deser`'s shape oracle alike. The message for an
661
+ unbindable HTTP port no longer names `lookup` as the reason a run continues.
662
+
663
+ `uses_oob_host` is the default answer rather than the answer. Whether a
664
+ method calls back can depend on its own flags: `--deser-formats` selects the
665
+ ecosystems, and three of the five -- `php`, `dotnet`, `python_pickle` --
666
+ ship no DNS gadget at all. So both callers ask `builds_callback_probes()` on
667
+ a configured instance instead of reading the class. `--deser-formats php`
668
+ with an address no longer draws a notice blaming the address, because there
669
+ was no callback probe to lose and a delegated domain could not have helped;
670
+ and it no longer starts a listener that nothing would reach.
671
+
672
+ Whether the host's *form* is one a method can use is a second question,
673
+ `can_use_oob_host()`, and only the gate asks it. A listener bound for
674
+ `--methods deser --oob-host 10.0.0.9` receives nothing -- `build_probes`
675
+ drops an address and sends no gadget -- and the run announced "the target
676
+ will open outbound connections" directly beneath the notice saying those
677
+ probes are not sent. It stays out of `builds_callback_probes()`, because the
678
+ stranded notice is *about* the address and an address test there would
679
+ silence it for the one case it exists to report.
680
+
681
+ `lookup` with an address stops getting a listener too. It built no probes
682
+ with one either way, so nothing is lost, and the stranded notice says more
683
+ about that run than the port-53 warnings it no longer reaches.
684
+
685
+ A flaky test met on the way, fixed here because it would have bitten CI
686
+ sooner or later: the fake parsing endpoint behind the `deser` shape
687
+ differential sniffed one byte (`raw[4:5] not in (b"t", b"s")`) to decide
688
+ whether a stream was well-formed. The `noise` form is the format's magic
689
+ followed by random characters, so it passed that check about once in 500
690
+ runs -- measured at 398 in 200,000 -- all three forms answered alike, and
691
+ the oracle's correct `negative` failed a test that wanted `needs-review`.
692
+ The endpoint now reads TC_STRING the way a parser does, a two-byte length
693
+ followed by exactly that many bytes: 0 collisions in 500,000 seeds. Seeding
694
+ the test's RNG would have hidden it instead, and would have gone on hiding a
695
+ real regression in how the noise form is built.
696
+
697
+ This is the second half of #94, found while fixing the first and reported as
698
+ #98 rather than folded into #97: same defect -- probes sent that no callback
699
+ could follow -- reached by a different route. The format-dependent part came
700
+ from an automated review of #97 and the host-form part from one of this
701
+ change, both verified before being accepted.
702
+
703
+ ## [2.45.2] — 2026-09-24
704
+
705
+ ### Fixed
706
+
707
+ - **`deser` fired DNS gadgets at a host that could not resolve, and the warning
708
+ built for exactly that stayed quiet.** `--oob-host` takes a delegated name or
709
+ an address. `oob` uses either: given an address it drops its DNS shapes and
710
+ carries the token in a URL path instead. `lookup` refuses an address
711
+ outright, because a lookup expression has nowhere else to put one.
712
+ `DeserSink` had neither behaviour -- it built `<token>.10.0.0.9` and sent it,
713
+ a real request the target answered from which no callback could follow by
714
+ construction, and then capped at `needs-review` without saying why.
715
+
716
+ `oob_channel_warnings` is the function that exists to say a DNS channel is
717
+ structurally dead. It returned early on an address literal, reasoning that
718
+ such a run "never builds DNS probes and has nothing to warn about" -- true of
719
+ `oob`, and written when `oob` was the only method taking the flag. So the one
720
+ mechanism that could have reported this was the one place it was ruled out.
721
+
722
+ Both halves are fixed together, because either alone still ships the problem.
723
+ `deser` treats an address as no host and sends its shape oracle alone, which
724
+ is what it can prove without a listener anyway. The notice moves into
725
+ `oob_address_strands`, which names the methods an address stranded and the
726
+ tier that puts out of reach, read off each class.
727
+
728
+ It is a separate function rather than another branch of
729
+ `oob_channel_warnings` because the two answer different questions at
730
+ different times. "Can the DNS shapes reach this listener" is only worth
731
+ asking once a listener exists, and that block runs only for a method that
732
+ *requires* a callback host. `deser` does not require one, so it never
733
+ reaches that block -- a notice left there could not have fired for the
734
+ method that prompted it. The new check runs on every verification run that
735
+ names a host.
736
+
737
+ Which methods those are is a class attribute now, `oob_needs_dns_label`,
738
+ declared beside `needs_oob_host`. Deciding it at the one place that asked is
739
+ how the two drifted apart to begin with.
740
+
741
+ The test that pinned the old silence asserted "no DNS shape is ever built" --
742
+ a claim about `oob` wearing the shape of a claim about the flag. Its scope
743
+ narrows to `oob`, and the stranded case gets tests of its own.
744
+
745
+ Found while verifying an automated review finding on #93; reported as #94 and
746
+ deliberately kept out of that documentation-only PR.
747
+
748
+ ## [2.45.1] — 2026-09-24
749
+
750
+ ### Fixed
751
+
752
+ - **`--max-payloads` is a bound on requests, and it was being checked after
753
+ they had been sent.** The cap was measured against the number of result
754
+ *rows*. For a method that answers from each probe those are the same number,
755
+ so nothing showed. An aggregate method reports one row however many probes it
756
+ cost, so the whole series went out and the cap noticed afterwards:
757
+
758
+ | `--max-payloads 1` | the cost line said | requests sent |
759
+ |---|---|---|
760
+ | `reflected`, `eval` | 1 | 1 |
761
+ | `time` | 1 | **12** |
762
+ | `deser` | 1 | **15** |
763
+ | `boolean` | 1 | **27** |
764
+
765
+ A clean split: every aggregate method overran and no per-probe method did.
766
+ That line is the only thing an operator bounding a monitored engagement has
767
+ to go on before the traffic starts, and it was wrong in the direction that
768
+ matters — under, not over.
769
+
770
+ Rows were the wrong meter in a second way, which the first fix exposed. A
771
+ series the budget declines costs requests and produces no row at all, so the
772
+ next carrier recomputed the same allowance and fired again — `time` at
773
+ `--max-payloads 5` sent 4 requests per carrier with the cap never moving. And
774
+ in a mixed run, a per-probe method never saw what an aggregate one had
775
+ already spent: `boolean,eval` at 30 sent 56. The budget is counted in
776
+ requests for the whole call now. Swept across every method this build can run
777
+ without a callback host, alone and in combination, at seven caps: 63
778
+ combinations, worst overrun 0.
779
+
780
+ - **A measurement that cannot be finished is no longer started.** What a budget
781
+ may do to a series depends on where the method's answer lives, and the class
782
+ already knows: `DetectionMethod.decides_per_probe()` reads it from whether
783
+ the class overrides `confirm_each`, rather than from a list of names that
784
+ would go stale like every other one in this repository.
785
+
786
+ Where each probe answers — `deser`, `lookup`, `oob` — running out of budget
787
+ stops the series and the probes already sent keep their verdicts. Where only
788
+ the series answers — `time`, `boolean` — a part of one is not a weaker answer
789
+ but a wrong one: `time` reports `negative` from a screen with no regression
790
+ behind it and from a regression short of four samples, and a `boolean` series
791
+ whose anchors never went out reads the same way. Both are false cleans, and
792
+ the budget would have been manufacturing them. So a wave that does not fit
793
+ abandons the measurement rather than cutting it short.
794
+
795
+ - **An abandoned measurement leaves an `inconclusive` row, not a silence.**
796
+ Dropping it quietly let the *other* carriers describe the run, and the other
797
+ carriers are the ones with nothing to find: measured at `--max-payloads 12`
798
+ against a sink that honours an injected sleep, the unix carrier's regression
799
+ was abandoned for budget and a windows carrier's honest "no separator
800
+ delayed" was the only row left — so the run reported `negative` for a target
801
+ that was vulnerable. `inconclusive` is what an abandoned measurement is in
802
+ the word the tool already uses, and it outranks `negative` in the run
803
+ verdict. Swept across 31 caps against a sink that really does delay: no cap
804
+ reports it clean, and none overruns.
805
+
806
+ - **What the budget declined is reported by name**, in a sixth run-wide tally
807
+ beside the profile drops, the safety holds, the reach notes, the settled
808
+ carriers and the refusals. A ladder that shrinks quietly is indistinguishable
809
+ from a target with nothing to find:
810
+
811
+ ```
812
+ [detect] --max-payloads held back 1 measurement(s) that could not have reached a
813
+ verdict within the budget:
814
+ [detect] 1 x boolean/raw needs 7 requests to reach a verdict and --max-payloads left 3
815
+ [!] No probes were built, so NOTHING WAS TESTED - this is not a negative result.
816
+ ```
817
+
818
+ - **The cost line and the run make the same decision.** `estimate_detection_probes`
819
+ applies the same all-or-nothing rule, so under a cap the figure printed before
820
+ the traffic is the traffic. Uncapped it stays a floor by documented design — a
821
+ wave a method picks after seeing its own timings cannot be predicted from
822
+ there.
823
+
824
+ - **`sent N probes` is a claim about what the target received**, and it was
825
+ counting result rows. Those are the same number for a method that answers
826
+ from each probe, which is why it read true for so long -- but an aggregate
827
+ method reports one row for a whole series, so a `time` run that put 20
828
+ requests on a target announced 5, and a measurement this release declines
829
+ would have announced one probe for traffic that never left. Both the run
830
+ summary and the per-point line report deliveries now, with the result count
831
+ beside them.
832
+
833
+ One figure in the 2.45.0 entry below came from that line and was mislabelled
834
+ with it. Re-measured against the same sink: `--methods reflected,eval,time`
835
+ put **2883** requests on the target, not 2426, which was the number of
836
+ results. Every one of them was still `negative`, so the claim stands; the
837
+ number did not.
838
+
839
+ - **An evasion retry is a request, and the meter could not see it.** `_escalate`
840
+ fired without incrementing `delivered_probes`, so retries were traffic
841
+ outside the bound this release makes real: measured at `--max-payloads 5
842
+ --evade high` against a filter that refuses whitespace, the target received
843
+ 12 requests and the run recorded 5. Counting them was half the fix. The other
844
+ half is that the budget is checked *per rung* rather than once before the
845
+ ladder -- one refused probe is retried at `low` and again at `high`, so it
846
+ can cost three requests, and a check at the call site alone still overshot by
847
+ one at `--max-payloads 10 --evade high`. Swept over caps 1 to 25 at all three
848
+ evasion settings: 75 combinations, no overrun, and the counter equal to the
849
+ requests the target actually received in every one.
850
+
851
+ ### Security
852
+
853
+ No authorization, deny-by-default or matrix-evaluation boundary moves, and
854
+ nothing widens `confirmed`. The effect is one-way: fewer requests than before,
855
+ and a run that could not afford to test something now says so instead of
856
+ reporting it clean.
857
+
858
+ ## [2.45.0] — 2026-09-23
859
+
860
+ ### Added
861
+
862
+ - **`--methods boolean` — a sink that evaluates a predicate and renders nothing
863
+ of it.** MongoDB `$where` is the shape, and this repository has carried the
864
+ note for two releases: a JS sandbox with no shell, no egress and no value in
865
+ the response, only a document set that a predicate narrows. Every shipped
866
+ oracle is structurally blind to it. Measured against exactly that sink:
867
+
868
+ ```
869
+ [detect] methods: reflected, eval, time
870
+ [detect] sent 2426 probes: negative=2426
871
+ ```
872
+
873
+ 2426 requests and a clean negative on a target that evaluates whatever it is
874
+ handed. The new method reads the one channel left — whether the *shape* of the
875
+ response changed between a true predicate and a false one.
876
+
877
+ **It is `needs-review` and there is no path from here to `confirmed`.** Not
878
+ because the signal is weak, but because of what it cannot distinguish: against
879
+ a sandboxed `eval` sink and against a plain SQLite comparison it produced an
880
+ identical clean differential in 40 runs each, and a query engine comparing two
881
+ numbers is not remote code execution. Extracting a locally computed product
882
+ bit by bit through the channel was tried and does not fix it — it recovers the
883
+ product through both sinks alike, for about 80 requests and a string function
884
+ a sandbox may well deny.
885
+
886
+ **The naive form of this oracle is unusable**, which is why none of it is.
887
+ `1==1` against `1==2`, with a changed response read as a finding, called a
888
+ target that only *reflected* its input vulnerable in 40 runs out of 40, and
889
+ one whose response merely wobbled in 32 of 40. Four guards, each a measured
890
+ false-finding rate rather than a precaution:
891
+
892
+ - **Compare structure, not the body and not its length.** A reflected payload
893
+ lands in the text between two tags, and the text between two tags is what
894
+ the signature throws away. A length-based signature claimed a differential
895
+ in 13 of 25 runs against a reflect-only target; comparing raw bodies was
896
+ unusable outright, reading `unstable` in 25 of 25 runs against a target that
897
+ *was* vulnerable, because one CSRF token makes every response unique.
898
+ - **Several independently randomised pairs, not one.** Against a target whose
899
+ response varies on its own, one pair claimed a differential in 46 of 200
900
+ runs; two claimed none in 200. `--probe-depth quick` trades four pairs for
901
+ two and never for one.
902
+ - **Randomised firing order.** A target that never reads the payload but
903
+ degrades part-way through a run splits an ordered true-then-false series
904
+ perfectly: at the worst point of a swept degradation, 100 false findings out
905
+ of 100.
906
+ - **An anchor before and after the series, each a different true predicate.**
907
+ Shuffling alone still left 2 in 100, which is just the chance a shuffle
908
+ lands separable. Re-measuring the channel afterwards caught it 100 times in
909
+ 100, because a target that moved during the series cannot answer the
910
+ closing anchor the way it answered the opening one. Sending *one* anchor
911
+ payload three times does not merely weaken that: a cache keyed on the query
912
+ string answers the repeats from its store, so the closing anchor agrees
913
+ with the opening one whatever the target did in between. Measured against
914
+ an input-blind target that degrades mid-series, identical anchors caught it
915
+ in 36 of 39 runs live and in 0 of 39 behind a cache. Every probe payload is
916
+ unique, so nothing else in the series is replayable.
917
+
918
+ With every guard on, a genuinely evaluating target still read as a
919
+ differential in 100 runs of 100.
920
+
921
+ - **A channel that cannot carry one bit is `inconclusive`, never `negative`.**
922
+ If the same probe draws two different shapes, or the shape moves while the
923
+ series is being fired, the run says so. `negative` asserts the probes reached
924
+ the target and found nothing; here they reached it and no answer could be read
925
+ out of them, which is the same false clean `blocked` and `nothing-tested`
926
+ exist to prevent, one oracle further in. No new verdict: there are still nine.
927
+
928
+ - **`response_shape()`** — a response reduced to its structure, with everything
929
+ it said removed. Three readings, because a response is one of three things and
930
+ the wrong reading is not a near miss: a JSON document keeps its keys, nesting
931
+ and list lengths and drops every scalar; markup keeps its tag skeleton;
932
+ anything else keeps one marker per word per line. Measured on a JSON sink, the
933
+ markup reading was unusable — `unstable` in 25 of 25 runs — and the shape tree
934
+ read the differential in 25 of 25.
935
+
936
+ ### Changed
937
+
938
+ - **`CODE_POSITION_CONTEXTS`** names the contexts that carry the injected value
939
+ as code rather than as a value, and `boolean` is offered every other one. The
940
+ first cut of this asked whether the context had a break-out prefix at all, and
941
+ got both halves wrong: it refused `attribute`, `attribute_unquoted`,
942
+ `xml_cdata` and `yaml`, whose delimiters open and close *around* the value and
943
+ leave a predicate exactly where a predicate belongs, and it offered
944
+ `unix_shell`, `windows_cmd` and `powershell`, which have no delimiters at all
945
+ and run the value as a command. The suite enumerates both sides, so a context
946
+ added to the corpus fails until somebody decides which one it is on.
947
+
948
+ - **`costly` now asks whether one probe buys an *answer*, not whether it costs
949
+ more than one response.** The two were the same question while every method's
950
+ probe was also its unit of information. `boolean` is the first where they come
951
+ apart: each of its probes is one ordinary request and none of them means
952
+ anything alone, because the answer is the partition across the whole series.
953
+ Read the old way it would have landed in the wave the enumeration driver runs
954
+ *first* — the one that exists to be answered cheaply — ahead of `reflected`
955
+ and `eval` and spending the same per-question budget, at 27 requests before it
956
+ could say a word. It sits with `time` instead, and the suite now holds every
957
+ aggregate method to that.
958
+
959
+ ### Security
960
+
961
+ - **The `OR` connectives ship behind `--verify-active-risk stateful`.** A
962
+ predicate probe breaks out of a condition the application already wrote, and
963
+ the connective is this method's command separator. `AND` differentiates only
964
+ where the application's own predicate is true and `OR` only where it is false,
965
+ so they are complements and dropping `OR` is a blind spot rather than a
966
+ saving. But a true predicate `OR`-ed into a `DELETE … WHERE` took a table from
967
+ 3 rows to 0, where the same predicate `AND`-ed into it left all 3 — so it goes
968
+ at the top rung, held back by default, with the run naming every shape it held
969
+ and the flag that sends it. `AND` and the bare form stay `safe` and change
970
+ nothing.
971
+
972
+ - **Nothing widens `confirmed`.** The new method's ceiling is one tier below it
973
+ and the suite holds that as behaviour, not as an attribute: `confirm_series`
974
+ is driven across 200 series including the perfect one, and none of them
975
+ reaches `confirmed`.
976
+
977
+ ## [2.44.0] — 2026-09-23
978
+
979
+ ### Fixed
980
+
981
+ - **`--evade low` substituted inside quoted programs and broke them.** Every
982
+ space became `${IFS}`, including the ones inside `awk 'BEGIN{print "RK" a+b
983
+ "RK"}'` — and inside single quotes `${IFS}` is literal text, not an
984
+ expansion, so awk was handed `BEGIN{print${IFS}"RK"...` and answered with a
985
+ syntax error.
986
+
987
+ Measured shape by shape against an unfiltered target: **8 probe shapes that
988
+ the canonical form executes broke at the rung, and none improved.** The
989
+ substitution stops at a quote now. Double quotes are left alone from the
990
+ other side — `${IFS}` *does* expand inside them, so substituting there would
991
+ change the string the target computes rather than the spacing around it.
992
+
993
+ ### Changed
994
+
995
+ - **A rung is a retry for a refused probe, not a posture for the run.** It was
996
+ applied to every probe regardless of whether anything was being filtered,
997
+ which is a pure loss on a target with no filter. Against a filter that blocks
998
+ whitespace it turned 1 confirmation into 5; against an unfiltered target it
999
+ now costs **zero** extra requests, because nothing was refused.
1000
+
1001
+ `--evade` is therefore a **ceiling**. Every probe goes out canonical, and
1002
+ only a refused one is retried, up to that ceiling — at most one request per
1003
+ rung. The ladder a run builds no longer depends on the rung at all: all three
1004
+ settings build the same 42 probes, of which the same 32 execute. Before, the
1005
+ rung built a smaller ladder and 13 fewer of its probes ran.
1006
+
1007
+ The run reports how many retries it made and which rung got through, whatever
1008
+ it concluded — including a run that confirmed, where the retry is the reason
1009
+ it did.
1010
+
1011
+ ### Added
1012
+
1013
+ - **A second evasion rung, `high`.** The measurement turned up two classes of
1014
+ filter and the shipped rung only addressed one. `low` removes whitespace;
1015
+ `high` also splits the command word with an expansion that vanishes
1016
+ (`ec$@ho`), for a filter matching command names. It is applied before the
1017
+ whitespace substitution, because afterwards the split lands inside `${IFS}`
1018
+ and makes `${I$@FS}` — neither an expansion nor a command.
1019
+
1020
+ Four boundaries the retry has to respect, each of which was a way to turn a
1021
+ vulnerable target into a `negative` — worse than the `blocked` the rung sits
1022
+ beside, because `blocked` at least says the run learned nothing:
1023
+
1024
+ - **Unix shell probes only.** `${IFS}` and `$@` are POSIX. A cmd.exe probe
1025
+ rewritten with them loses its spaces, so a whitespace filter answers 200
1026
+ and the retry counts as a win while cmd.exe cannot run it.
1027
+ - **A break-out context opens by *closing* a quote.** Reading that leading
1028
+ quote as an opener left `'; echo …` untouched, so the rung did nothing on
1029
+ exactly the contexts a filter is most likely to sit in front of.
1030
+ - **A followup is read again after a retry that lands.** `file` writes its
1031
+ token on the request that arrives, so a body read before the retry is a
1032
+ read of a file that did not exist yet.
1033
+ - **A redirect is never retried.** The build-time transform took an explicit
1034
+ `evade=False` for these, and that parameter stopped doing anything when the
1035
+ rung became a retry — a guard lost in the move, restored where the retry
1036
+ now happens.
1037
+
1038
+ The aggregate methods escalate too. `time`, `oob`, `lookup` and `deser` take a
1039
+ different branch, and leaving it out meant the documented ceiling did nothing
1040
+ for four of the eight methods — the same branch, and the same omission, as the
1041
+ refusal check one change earlier.
1042
+
1043
+ ## [2.43.0] — 2026-09-23
1044
+
1045
+ ### Added
1046
+
1047
+ - **A ninth verdict: `blocked`.** A run whose payloads a filter refuses has
1048
+ learned nothing about the sink, and it used to say otherwise. Measured
1049
+ against a real command injection behind a filter that 403s a space or a
1050
+ separator:
1051
+
1052
+ ```
1053
+ before: sent 10 probes: negative=10
1054
+ after: sent 10 probes: blocked=10
1055
+ ```
1056
+
1057
+ `negative` asserts that the probes **reached** the target — the project's own
1058
+ tier rule says so. They reached a filter. That is the same false clean
1059
+ `nothing-tested` exists to prevent, one level further in.
1060
+
1061
+ **The signal is differential**, like everything else this tool decides: the
1062
+ payload-free control got through and the probe did not, so what was refused
1063
+ is the payload. An endpoint answering 403 to everything — an auth wall, a
1064
+ path that does not exist for this session — refuses the control too and is
1065
+ not mistaken for a filter. Verified against exactly that case.
1066
+
1067
+ 4xx only. A 5xx is as likely to be the payload *breaking* the application,
1068
+ which means it reached something, and reading that as blocked would hide the
1069
+ one response saying the sink is live. No vendor list and no block-page
1070
+ fingerprints: a status the control did not get is the whole signal.
1071
+
1072
+ A run where **some** probes got through stays a real `negative` — the sink
1073
+ saw those — and the refusals are reported either way.
1074
+
1075
+ **A refusal never unmakes evidence.** Only a `negative` is replaced, because
1076
+ it is the only verdict a refusal contradicts. An application can execute the
1077
+ payload and then answer 400 with the output in its body, and the oracle has
1078
+ already proven execution from a value random to that probe; overwriting that
1079
+ would turn demonstrated RCE into a false negative, which is worse than the
1080
+ false clean this verdict removes.
1081
+
1082
+ **Every method is covered, including the three that decide per probe from a
1083
+ series.** `oob`, `lookup` and `deser` take a different branch, which returned
1084
+ before refusal was considered — so a callback run whose every probe was
1085
+ refused reported `negative` for each of them, because no callback arrived.
1086
+
1087
+ Refusals are counted in **requests**, like the status tally beside them: an
1088
+ aggregate method fires a whole ladder and reports one row, so counting rows
1089
+ described the same run with different arithmetic and hid a partly filtered
1090
+ series entirely.
1091
+
1092
+ ### Fixed
1093
+
1094
+ - **The advice on a filtered run pointed at the wrong thing entirely.** "The
1095
+ target may be patched" reads as a clean bill of health for a target that was
1096
+ never reached; the blind-sink list names methods a filter refuses in exactly
1097
+ the same way; and the second-order line said the target *accepted* an input
1098
+ it had in fact rejected with a 403. None of the three fires on a refused run
1099
+ now. It is replaced by what was actually observed, with the flags that change
1100
+ the payload's shape.
1101
+
1102
+ ## [2.42.0] — 2026-09-23
1103
+
1104
+ ### Added
1105
+
1106
+ - **A run that confirms nothing now names the second-order oracle.** Measured
1107
+ against a target that stores on one endpoint and renders through a shell on
1108
+ another — a real RCE — every probe read `negative`, and the run answered with
1109
+ four methods that are all negative there too, because the execution does not
1110
+ happen on the request being measured:
1111
+
1112
+ ```
1113
+ --methods time -> negative=4
1114
+ --observe-url -> confirmed, first run
1115
+ ```
1116
+
1117
+ The one flag that works was named nowhere. It is named now, and **not gated
1118
+ on which methods have run**: the blind-sink list is, so an operator who had
1119
+ already tried the expensive methods — exactly the one with nothing left but
1120
+ second order — was told only that the target might be patched.
1121
+
1122
+ The wording follows what the run observed rather than what it assumes. Input
1123
+ returned verbatim means a sink that reflects without executing; input that
1124
+ never came back means this response cannot show what became of it; and a run
1125
+ of aggregate methods alone, which record no per-probe observation, claims
1126
+ neither. A swallowed input is equally a blind sink or a stored one — `ping
1127
+ <input> >/dev/null` returns nothing either — so both are named and neither is
1128
+ picked.
1129
+
1130
+ - **Whether the target returned the input is recorded on every probe.** It was
1131
+ already computed on the confirmed path, where it becomes "target also
1132
+ reflects the payload verbatim"; a negative probe never looked, and the
1133
+ negative run is the one that has to say what it saw.
1134
+
1135
+ Measured against the payload that actually went out. A carrier that
1136
+ multiplies through a filter never spells the joined `a*b` out — Liquid sends
1137
+ `{{ a | times: b }}` and Django `{% widthratio a 1 b %}` — so an endpoint
1138
+ echoing the whole payload recorded a measured `False`, and a target profile
1139
+ filtering the `*` shapes leaves only those, at which point the run would
1140
+ report "returned none of it" about a target that returned everything.
1141
+
1142
+ Recorded only when there was a response to look at. A delivery error is not
1143
+ an observation, and a `False` for one would put an unmeasured claim exactly
1144
+ where the unobserved branch belongs.
1145
+
1146
+ ## [2.41.0] — 2026-09-23
1147
+
1148
+ ### Fixed
1149
+
1150
+ - **A candidate that confirmed RCE was never asked whether it was also a
1151
+ deserialization or lookup sink.** The enumeration driver split the methods
1152
+ into a cheap wave and an expensive one and skipped the expensive wave once
1153
+ execution was proven — reasonable for `time` and `oob`, which would only put
1154
+ a second name on one finding, and wrong for `lookup` and `deser`, which
1155
+ report *different properties* with their own remediation. They sat on the
1156
+ expensive side of a hand-written set of names, so the answer was never asked
1157
+ for and never reported.
1158
+
1159
+ The split now comes from the tier each method declares. Anything reporting
1160
+ `confirmed` or `needs-review` is answering *did this target execute my
1161
+ input*; anything else is a different question and is never skipped for an
1162
+ answer to that one.
1163
+
1164
+ - **`--max-payloads` is spent per question, not per wave and not per
1165
+ candidate.** Per wave it quietly doubled: a run capped at 5 sent 10 probes to
1166
+ every candidate that did not confirm, while the cost line printed before any
1167
+ traffic said 5. Bounding the candidate instead starves the different
1168
+ question — the cheap methods eat the whole allowance and `deser` never runs,
1169
+ which is the same finding lost by another route. Every method asking about
1170
+ execution now shares one allowance, each different property gets its own, and
1171
+ the cost line names how many questions are being asked — and **sums** the
1172
+ estimate across them. Counting the cap once while the run grants it per
1173
+ question advertised 44 requests for a run that sent 80, which is wrong in
1174
+ the direction that matters for the operator bounding a monitored engagement.
1175
+
1176
+ - **A second-order confirmation settles its carrier too.** With `--observe-url`
1177
+ a probe can read negative in the response it drew and `confirmed` on the
1178
+ observed channel a moment later. The stop was decided from the pre-poll
1179
+ verdict, so the carrier kept probing after it had in fact confirmed —
1180
+ spending the budget the stop exists to hand to carriers not yet examined,
1181
+ which is the coverage loss this change was written to remove, reappearing on
1182
+ the one oracle that needs a second request to answer.
1183
+
1184
+ ### Changed
1185
+
1186
+ - **A carrier that has confirmed stops there** (`--confirm-depth first`, the
1187
+ new default). One carrier is one method in one environment and context, and
1188
+ once it has confirmed every further shape of it can only say the same thing
1189
+ again.
1190
+
1191
+ Measured against an executing target: one candidate spent 115 of its 120
1192
+ probes after the first confirmation, and printed 32 confirmations of which 29
1193
+ were duplicates inside a single carrier. Those probes were not idle — they
1194
+ were spent instead of reaching carriers never examined at all. At the same
1195
+ budget the run went from **4 carriers examined to 23**, and from **4
1196
+ environments reached to 9**. This buys coverage rather than saving requests.
1197
+
1198
+ The stop is per carrier and **never** per candidate: a candidate may confirm
1199
+ as `unix` while a later `nodejs` carrier is the only thing a different target
1200
+ would have shown. `--confirm-depth every` maps every shape a sink accepts,
1201
+ which is what writing a proof of concept by hand needs, and the run reports
1202
+ how many shapes it held back and which carriers stopped — a fourth tally
1203
+ beside the profile drops, the safety holds and the reach notes, because it
1204
+ says a fourth thing: the probe could have been sent and had nothing left to
1205
+ establish.
1206
+
1207
+ - **A method declares what one of its probes costs.** `CHEAP_DETECTION_METHODS`
1208
+ was a set literal; it is now derived from a `costly` attribute each class sets
1209
+ for its own reason — a real sleep, a wait for a callback, a second fetch.
1210
+ Every hand-written list naming methods in this repository has gone stale, and
1211
+ this one had put `lookup` and `deser` where being skipped cost findings.
1212
+
1213
+ ## [2.40.0] — 2026-09-22
1214
+
1215
+ ### Added
1216
+
1217
+ - **A carrier may take the operands apart, and two engines need it.** Every
1218
+ expression carrier until now substituted `__EXPR__` — the joined `a*b` —
1219
+ which quietly assumed the engine has an arithmetic operator. Two widely
1220
+ deployed ones do not, and both were measured as false negatives:
1221
+
1222
+ ```
1223
+ liquid every bare form missing; {{ 45013 | times: 45989 }} -> 2070102857
1224
+ django every bare form missing; {% widthratio 45013 1 45989 %} -> 2070102857
1225
+ ```
1226
+
1227
+ Liquid multiplies with a filter and Django with a tag, so neither form can be
1228
+ written as a single expression — an application that really does evaluate the
1229
+ template was reported `negative`, which is the same shape as `oob` against a
1230
+ `${jndi:...}` sink: probes that reach the target and cannot speak its
1231
+ language. A carrier template may now use `__A__` and `__B__` as well as
1232
+ `__EXPR__`.
1233
+
1234
+ Neither payload contains the product, so a target that merely echoes the
1235
+ payload still cannot read as `confirmed`. A test pins that for every shipped
1236
+ carrier.
1237
+
1238
+ ### Changed
1239
+
1240
+ - **A carrier template not parameterised by both operands is skipped, not
1241
+ sent.** With no token at all it renders the same constant every probe; with
1242
+ only one operand the target is never handed the other, so nothing it can
1243
+ compute is the product RCEKit is looking for. Either way the product would
1244
+ not be evidence the target computed anything — and a probe that cannot
1245
+ confirm still counts toward the coverage a run reports, which is the part
1246
+ that matters more than the wasted request.
1247
+
1248
+ - **`--eval-engines` names every carrier the corpus ships.** Its help listed
1249
+ three engines by hand and two were added. An operator narrowing that flag
1250
+ reads the list and nothing else, so a stale one says an engine needs no
1251
+ carrier when it does, and they cut the only probe that could have confirmed
1252
+ it. A test now holds the help text to the corpus.
1253
+
1254
+ - **What the survey measured and did not ship is recorded too.**
1255
+ `eval_carrier_survey` in the corpus now names the engines that need no
1256
+ carrier — nunjucks 3.2.4, tornado 6.5.10, mako 1.4.1, chameleon 4.6.0,
1257
+ smarty 5.8.4 and Ruby's ERB all return the product from a bare form — and
1258
+ those that are
1259
+ out of reach. Handlebars 4.7.9 fails every bare form and, being logic-less
1260
+ with no built-in arithmetic helper, has no template text that computes a
1261
+ product at all. Go `text/template` 1.23 fails every bare form too, and the
1262
+ only forms that do return the product — `{{printf "%d" <product>}}` and
1263
+ `{{<product>}}` — hand the target the answer, so a target that merely echoed
1264
+ them would read as `confirmed`. That is the rule a carrier lives under: **a
1265
+ carrier may not carry its own result.**
1266
+
1267
+ Saying so is worth more than a carrier that cannot work, and it stops the
1268
+ next person re-running the same survey — or shipping the Go form, which the
1269
+ survey itself produced and which looks like a carrier until you ask what an
1270
+ echoing target would return.
1271
+
1272
+ ## [2.39.0] — 2026-09-22
1273
+
1274
+ ### Added
1275
+
1276
+ - **Every part of a multipart body is an injection point.** None of them were
1277
+ before. The form branch matched on `=` appearing anywhere in the body, so a
1278
+ `multipart/form-data` capture was split on `&` and yielded exactly one
1279
+ candidate, named after a `Content-Disposition` line:
1280
+
1281
+ ```
1282
+ form | '--X\r\nContent-Disposition: form-data; name' | body param '...'
1283
+ ```
1284
+
1285
+ Every probe for that point rewrote a *part header*, so it could confirm
1286
+ nothing, while `user`, `avatar` and `note` — the fields the form actually
1287
+ posts — were never reached. The run still printed a point and a probe count,
1288
+ which is the part that matters: coverage reported and not delivered reads
1289
+ exactly like a clean target.
1290
+
1291
+ Several parts may post under one name — a multi-file input and a checkbox
1292
+ array both do — so a part is addressed by its **index**, as a JSON leaf is
1293
+ addressed by its token path. Addressing by name alone rewrote the first part
1294
+ for every candidate, which is the same failure one level further in: three
1295
+ files, three points, the first file probed three times and the other two
1296
+ never touched.
1297
+
1298
+ `multipart` is now its own kind, recognised from `Content-Type` and decided
1299
+ before the form branch can see the body. A file part is a candidate too — its
1300
+ content is the value under test, while its `filename` and `Content-Type` stay
1301
+ as captured. Verified end to end against a target that parses with the
1302
+ standard library's MIME parser: `multipart field 'note': confirmed`, with the
1303
+ non-vulnerable `user` field `negative` beside it.
1304
+
1305
+ ### Changed
1306
+
1307
+ - **A GraphQL request is ordered by what can actually confirm.** Its variables
1308
+ were already enumerated — they are JSON leaves — but they were tried in
1309
+ document order alongside `query` and `operationName`, and those two cannot
1310
+ confirm anything. A payload in `query` *replaces* the operation document, so
1311
+ the server answers with a parse error before a resolver runs, and
1312
+ `operationName` then names an operation that is no longer there. On the
1313
+ capture this was measured against they were two points of five, each one a
1314
+ full probe ladder.
1315
+
1316
+ They are moved behind the variables, **not dropped**. A server that logs the
1317
+ query document before parsing it is reachable through exactly that field,
1318
+ which is the route Log4Shell took through access logs, so a full run still
1319
+ tests both and `--max-points` now cuts the least likely to pay first. A plain
1320
+ `{"query": ...}` body with no `variables` is left in document order: it is as
1321
+ likely to be a search API, and there the query field is the one worth testing.
1322
+ Carrying both keys is not enough either — the `query` string has to open like
1323
+ a GraphQL document, so `{"query": "red shoes", "variables": {...}}` keeps its
1324
+ real injection point where a bounded `--max-points` run will still reach it.
1325
+
1326
+ ### Fixed
1327
+
1328
+ - **The verdict table was missing a verdict.** The README said "seven verdicts
1329
+ that are never collapsed into each other" and listed seven, while the tool
1330
+ reports eight: `lookup-sink` was absent from the one table whose whole job is
1331
+ to enumerate them — in a README that uses the word two tables higher, in the
1332
+ Log4Shell CVE row, and again in the methods table. The sentence counted the
1333
+ rows the table had rather than the verdicts there are, so the omission never
1334
+ contradicted itself and nothing failed.
1335
+
1336
+ Found by reading the README end to end for this change. A test now reads the
1337
+ table against `DETECTION_METHODS`, so a verdict a method declares and the
1338
+ table does not carry is a failure rather than a silence.
1339
+
1340
+ - **A multipart body now goes out with the line endings it needs.** `-r`
1341
+ normalises the whole request to LF, so the CRLF delimiters RFC 2046 requires
1342
+ were gone by the time anything was sent. Bodies rendered for a multipart point
1343
+ are re-serialised canonically — probe and payload-free control alike, so the
1344
+ two differ in the field under test and in nothing else. Part content is left
1345
+ character for character, so a lone newline inside an uploaded text file
1346
+ survives.
1347
+
1348
+ ## [2.38.0] — 2026-09-21
1349
+
1350
+ ### Added
1351
+
1352
+ - **A probe shape may reach past the run's tier and be sent anyway**, when its
1353
+ effect is one the run undoes by saying it happened. `reaches_past` is that
1354
+ declaration, beside `safety`, which stays for an effect a notice cannot take
1355
+ back.
1356
+
1357
+ Reach wins where the two pull against each other. Detection the tool could
1358
+ have done and did not is a false negative wearing a safety label, and it
1359
+ costs more than the noise it saves.
1360
+
1361
+ `deser`'s DNS gadget is the first of these, and closes an inconsistency
1362
+ recorded a version ago: it makes the target resolve a name -- the very thing
1363
+ `oob` and `lookup` are refused for at `safe` -- while its only gate was
1364
+ `--oob-host`. Holding it back would have sent fewer probes at the default
1365
+ tier. It goes, and the run reports how far it reached.
1366
+
1367
+ Three tallies now, because they say three different things and one number
1368
+ would state the wrong one about all of them: the profile dropped it (it could
1369
+ not have reached the sink), the tier held it (it could, and was not sent), or
1370
+ it reached past the tier (it was sent, further than asked).
1371
+
1372
+ `lookup`'s `ldap://` and `rmi://` are the edge the rule has, and stay at
1373
+ `stateful`: a class fetched from an address RCEKit did not choose is not
1374
+ something a notice takes back.
1375
+
1376
+ - **Two worked examples in the README's Quick start**, because the fullest run
1377
+ RCEKit can make was not shown anywhere near the front.
1378
+
1379
+ The first is a captured request: most sinks worth testing sit behind a POST
1380
+ with a session cookie, a content type and a body, and `--verify-url` carries
1381
+ none of that. The second is everything the tool has -- injection-point
1382
+ enumeration across every value in that request, every method, callbacks, the
1383
+ top rung -- with a table of what each flag opens up and the cost line that
1384
+ prints before it fires.
1385
+
1386
+ Both were run before being written down. The point the Quick start never
1387
+ made: `--auto-params` needs `-r`, so the fullest run is not reachable from a
1388
+ URL at all, which is worth knowing before concluding a target is clean.
1389
+
1390
+ ## [2.37.0] — 2026-09-20
1391
+
1392
+ ### Added
1393
+
1394
+ - **A detection method declares the risk rung it needs**, and the engine reads
1395
+ it. `SAFETY_ORDER` has labelled corpus payloads `safe` / `intrusive` /
1396
+ `stateful` from the start and the query-language bridges followed; detection
1397
+ methods did not. Each risky one was gated by a hand-written branch in
1398
+ `main()` naming it, so a new method meant remembering to add another -- and a
1399
+ probe shape with nowhere to declare its rung was deleted rather than gated.
1400
+
1401
+ | method | rung | |
1402
+ |---|---|---|
1403
+ | `reflected`, `eval`, `time`, `deser` | `safe` | compute, delay, or parse |
1404
+ | `oob`, `lookup` | `intrusive` | makes the target open outbound connections |
1405
+ | `file`, `write` | `stateful` | writes to the target |
1406
+
1407
+ `file` and `write` are gated by their own configuration rather than by the
1408
+ rung: neither does anything until a directory to write into and a URL to read
1409
+ it back from are named, which says more than a tier would, and asking for the
1410
+ flag as well would refuse a command that works today. Nothing that runs today
1411
+ stops running.
1412
+
1413
+ - **A probe shape may need a higher rung than its method**, so coverage that
1414
+ only makes sense at the top tier has somewhere to live instead of being
1415
+ deleted. **`--methods lookup` sends `ldap://` and `rmi://` again**, at
1416
+ `stateful`.
1417
+
1418
+ They were removed in 2.36.0 on the argument that `dns://` resolves wherever
1419
+ `ldap://` would. That claim was too strong -- a filter catching the string
1420
+ `dns:` and not `ldap:`, or a trimmed runtime without the DNS provider,
1421
+ defeats it -- and a sink that takes one scheme and not the other is exactly
1422
+ the sink this method is for. The reason they were removed was real: they
1423
+ continue *past* resolution and connect to whatever address the answer named,
1424
+ which is not an address RCEKit chose. That is a rung, not a reason to drop
1425
+ coverage.
1426
+
1427
+ - **The run says what the rung held back**, counted apart from the target
1428
+ profile's drops and with the flag that would send it. The two say different
1429
+ things: a profile drop means the probe *could not have* reached the sink,
1430
+ while this means it could and the operator chose not to send it. Reporting
1431
+ them together would state the first about the second.
1432
+
1433
+ - **A method declares the weaker tiers it really reports**, not only its
1434
+ ceiling. `write` reports `needs-review` for a file that is served but not
1435
+ interpreted and `deser` for a shape fingerprint, and three separate places
1436
+ had to know that -- the documentation tests, the benchmark's expectation
1437
+ whitelist, and the advice printed after a clean in-band run. Each kept its
1438
+ own answer; the docs test exempted `needs-review` for *every* method, so
1439
+ `lookup` could name a verdict it never emits and pass.
1440
+
1441
+ ### Changed
1442
+
1443
+ - **`--methods file` and `--methods write` with nothing configured now say
1444
+ so by name.** They were simply not applicable before, so the run built no
1445
+ probes and reported `nothing-tested` -- which is the quietest way this tool
1446
+ can fail and reads much like a clean target.
1447
+
1448
+ - **`docs/reference.md` carries a `Rung` column**, held against the class by a
1449
+ test. The tier column already was; this is the same claim one column over.
1450
+
1451
+ ### Fixed
1452
+
1453
+ - **A config-gated method is not re-gated by the rung at runtime.** The
1454
+ pre-flight lets `--methods file --webroot ... --web-base-url ...` through
1455
+ because the configuration is the gate, but the probe filter read the run's
1456
+ default `safe` ceiling and held every probe inheriting the method's
1457
+ `stateful` rung. The CLI accepted a documented invocation and then reported
1458
+ `nothing-tested` -- the quietest way this tool can fail, and the thing the
1459
+ rung work was supposed to remove rather than add.
1460
+
1461
+ Nothing caught it because every `file` test builds the method's config
1462
+ directly, without `max_safety`, so the ceiling fell back to the method's own
1463
+ rung and the probes went out. A run through the CLI with the channel
1464
+ configured is the one thing that would have, and there is one now.
1465
+
1466
+ - **The cost estimate applies the risk tier as well as the target profile.** It
1467
+ counted every shape a method built, so once a rung could narrow a method the
1468
+ pre-flight figure over-counted -- three times over for `lookup` at the
1469
+ default tier, which is exactly the operator who narrowed the run on purpose.
1470
+ Both paths share one predicate now, and it counts nothing, so an estimate
1471
+ never moves the numbers the report prints.
1472
+
1473
+ ## [2.36.0] — 2026-09-19
1474
+
1475
+ ### Added
1476
+
1477
+ - **`--methods lookup`: confirmation for an expression-lookup sink**, the shape
1478
+ Log4Shell has, where the sink resolves a URI instead of running a command.
1479
+
1480
+ `oob` could not reach one. Its `applicable` does admit `java` -- a Java
1481
+ application can shell out, so the environment is genuinely shell-capable --
1482
+ but every probe it builds is a shell command: `nslookup`, `curl`, `certutil`,
1483
+ `iwr`. A sink that interpolates `${jndi:...}` runs none of them, so the method
1484
+ applied, sent its whole ladder, and came back `negative` on a target that is
1485
+ exploitable. The README's Log4Shell row rested on the standalone listener and
1486
+ a generated payload file, which produce no verdict row at all, so nothing in
1487
+ the engine could reproduce that claim.
1488
+
1489
+ The probes are lookups and nothing else, and they depend on the injection
1490
+ context rather than the environment, exactly as `eval`'s do. The oracle is the
1491
+ one `oob` already uses: a token the target could only have learned by
1492
+ resolving what it was handed. The expression resolves `<token>.<host>`, so
1493
+ the in-process DNS listener is the entire apparatus -- no LDAP or RMI server
1494
+ is needed, and none is started.
1495
+
1496
+ **It reports `lookup-sink`, never `confirmed`.** A callback proves the sink
1497
+ resolved a URI RCEKit chose -- that it evaluated the expression it was handed.
1498
+ It does not prove the target ran attacker code: Log4Shell becomes RCE when the
1499
+ LDAP server answers with a loadable class. So the method gets its own
1500
+ proven-sink tier beside `deserialization-sink`, with its own section in the
1501
+ report, and `confirmed` keeps meaning executed.
1502
+
1503
+ **Only `jndi:dns://` is sent, and that is the security property rather than a
1504
+ shortcut.** A name lookup can be nothing else. `ldap://` and `rmi://` continue
1505
+ *past* resolution and open a connection to whatever address the answer named
1506
+ -- by default `127.0.0.1`, which is the target's own loopback. Whatever
1507
+ replies on :389 or :1099 is not RCEKit, so a reference could come back and a
1508
+ class be instantiated: the tool would have crossed the line this method exists
1509
+ to stop short of, having promised it had not. Dropping the two schemes also
1510
+ costs no coverage -- `DnsContextFactory` ships in the JDK, so `dns://`
1511
+ resolves wherever `ldap://` would, and on 2.15.0 it still resolves where
1512
+ `ldap://` no longer does. The proof is the callback and the finding is "this
1513
+ sink resolved a URI I chose" -- the same line `deser` draws, drawn here before
1514
+ it can be crossed.
1515
+
1516
+ Behind the same two gates as `oob`, and sharing its listener: it needs
1517
+ `--oob-host`, and it is held back at the default safety tier because it makes
1518
+ the target open outbound connections. Without `--oob-host` it builds no probes
1519
+ at all, which the engine reports as `nothing-tested` -- never `negative`. An
1520
+ address literal is a literal whichever family it is from: `::1` and `[::1]`
1521
+ build nothing, the same as `10.0.0.1`, because a token can only ride in a DNS
1522
+ label. `blind_sink_advice` names the method as proving a lookup sink and not
1523
+ execution, in a list where `oob` and `file` mean confirmed execution, and the
1524
+ README's Log4Shell demo heading says `lookup-sink` rather than `confirmed` --
1525
+ it was left claiming execution beside the table row that no longer does.
1526
+
1527
+ ## [2.35.5] — 2026-09-19
1528
+
1529
+ ### Fixed
1530
+
1531
+ - **The `--insecure` notice describes the run it is in.** It was printed as
1532
+ soon as the flag was seen, so a generation-only run, `--doctor`, or a plain
1533
+ HTTP target all announced a TLS downgrade that never happened -- and it named
1534
+ both downgrade rungs even on an OpenSSL build that had refused one of them.
1535
+ Reporting a downgrade on a run that opened no TLS connection is the same
1536
+ defect as reporting a probe that was never sent, in the one line written to
1537
+ be an audit of the run.
1538
+
1539
+ Building the context and saying so are now separate questions. The context is
1540
+ built whatever the target's scheme is -- urllib follows a redirect with the
1541
+ handler it was given, so an `http://` target that lands on a self-signed or
1542
+ legacy `https://` one needs it as much as a direct HTTPS target does. The
1543
+ notice fires once, when a connection has actually reached TLS, and lists only
1544
+ the rungs that took.
1545
+
1546
+ - **A benchmark case may set `timeout` to zero.** `timeout or 900.0` replaced
1547
+ an explicit `0` with the fifteen-minute default, so a case deliberately
1548
+ bounded to no time at all ran for much longer than it asked for.
1549
+
1550
+ ## [2.35.4] — 2026-09-19
1551
+
1552
+ ### Fixed
1553
+
1554
+ - **A wave the declared profile emptied is no longer read as a finished
1555
+ method.** The probe filter drops probes between the method and the wire, and
1556
+ when it emptied an adaptive method's *first* screening wave the engine took
1557
+ the empty batch for "nothing left to send" and stopped before asking for the
1558
+ next wave at all.
1559
+
1560
+ An adaptive method holds separators back in waves precisely because a filter
1561
+ is expected: `time` screens two of them first and keeps `||`, `&&`, the
1562
+ newline and the bare command for a second wave. A sink that strips `;` and
1563
+ `|` removes exactly the first wave and leaves the rest intact -- and the rest
1564
+ were never sent. Measured against a sink reachable only through `&&`, the run
1565
+ reported **`negative`**: "the probes reached the target and found nothing",
1566
+ about probes that were never sent. Worse than `nothing-tested`, which is at
1567
+ least true.
1568
+
1569
+ The loop now tracks what the method offered separately from what the profile
1570
+ allows to be sent, and ends only when the method itself is done. The round cap
1571
+ still bounds it, so a method whose every wave is filtered cannot spin the
1572
+ engine.
1573
+
1574
+ - **The unit suite runs off Linux.** The fake vulnerable sinks in the tests are
1575
+ POSIX command-injection points and every probe built for them is POSIX, but
1576
+ they were executed through `os.popen` — which is `cmd.exe` on Windows. The
1577
+ sink the test says executes did not execute, so the oracle correctly reported
1578
+ no execution and five tests failed for a reason unrelated to the code under
1579
+ test. Worse, the tests asserting a *negative* stayed green throughout: a
1580
+ broken fixture that keeps its controls passing is the failure this project
1581
+ takes seriously everywhere else. The sinks now name a POSIX `sh` explicitly;
1582
+ on Linux and macOS that is the `/bin/sh` they always used.
1583
+
1584
+ - **A deliberately dead target no longer costs minutes.** A closed loopback port
1585
+ answers with a RST on Linux and is silently dropped on Windows, where each
1586
+ probe waits out the SYN retry instead — measured at ~2s per probe. The
1587
+ benchmark harness's own unreachable-target case paid that for the full ladder
1588
+ twice, once for the vulnerable half and once for the control: 1800s for one
1589
+ test, longer than the other 534 together. Bounded to three probes it measures
1590
+ 9.1s, and three probes prove "nothing reached the target" exactly as well as
1591
+ forty do.
1592
+
1593
+ - **A captured-request fixture reaches disk byte for byte.** The cleartext-capture
1594
+ tests write a raw HTTP request whose text already spells its own CRLF line
1595
+ endings, through `Path.write_text` — which on Windows translates the newline of
1596
+ each one again. The file on disk held a doubled carriage return, the parser
1597
+ found no headers, and two tests failed against a fixture that had stopped being
1598
+ an HTTP request at all.
1599
+
1600
+ The third test in that class passed throughout, for the wrong reason: it
1601
+ asserts that a notice is *absent*, and a request that cannot be built prints no
1602
+ notice either. A control that stays green while its fixture rots is precisely
1603
+ what this project refuses to accept from a benchmark case, so the fixture now
1604
+ has a guard of its own. The capture is written as bytes; the `newline` argument
1605
+ that would say the same thing arrived in Python 3.10, and this project
1606
+ supports 3.8.
1607
+
1608
+ ## [2.35.3] — 2026-09-18
1609
+
1610
+ The target profile an operator declares now reaches the probe ladder, not just
1611
+ the corpus.
1612
+
1613
+ ### Fixed
1614
+
1615
+ - **`--deny-chars` / `--max-length` reach the detection probes.** They were
1616
+ applied by `_filter_by_profile`, which drops *corpus records* — and stopped
1617
+ there. The probes a detection method builds from those records went out
1618
+ regardless, so a run that had been told "this target strips quotes" still paid
1619
+ for every quote-carrying rung of the ladder, on requests structurally unable
1620
+ to confirm. Those requests are not free: on a captured request with `-p all`
1621
+ they are the budget the next injection point never got. The filter now sits at
1622
+ the engine, where every probe passes through it — deliberately not inside
1623
+ `_wrap_variants`, because `_space_free_probes`, the query-language bridges,
1624
+ `eval`, `oob` and `deser` each build payloads without going through that
1625
+ helper, and a gate that reaches some methods and not others is the side path
1626
+ that once left `file`/`time`/`oob` unable to send the raw rung.
1627
+
1628
+ Denying a character narrows the ladder rather than emptying it: a target that
1629
+ strips `;` is still probed through `|`, `||`, `&&` and the newline, which is
1630
+ what the separator table has always been for.
1631
+
1632
+ Checked on the literal payload, before the delivery layer percent-encodes it
1633
+ for its injection point. That is stricter than the corpus check, which is
1634
+ applied to the encoded payload and so lets a URL-encoded quote through a quote
1635
+ filter. The layers genuinely differ: transport encoding is undone by the
1636
+ server before the value reaches the sink, so a percent-encoded quote is still
1637
+ a quote when the application's own filter sees it.
1638
+
1639
+ - **A profile strict enough to remove every probe reports `nothing-tested`.**
1640
+ Not `negative`, which would read as "not vulnerable" from a run that sent
1641
+ nothing. The message names the profile as the cause and the characters a probe
1642
+ would have to avoid, instead of the generic advice to widen `--environments` —
1643
+ which is not what emptied the run.
1644
+
1645
+ - **The cost estimate follows the profile.** `[detect] cost:` builds the probes
1646
+ and counts them, so it now counts the ones that will actually be sent. An
1647
+ estimate that ignores a filter is wrong precisely for the operator who
1648
+ narrowed the run on purpose.
1649
+
1650
+ ### Changed
1651
+
1652
+ - A run that dropped probes says so, with the reason and a count per reason. A
1653
+ ladder that shrinks quietly is the one way this filter could manufacture a
1654
+ false negative, so the removal is stated rather than left to be inferred from
1655
+ the traffic.
1656
+
1657
+ ## [2.35.2] — 2026-09-18
1658
+
1659
+ ### Fixed
1660
+
1661
+ - **`--insecure` now reaches a legacy TLS stack, not just an untrusted one.**
1662
+ Turning certificate verification off is not the same as completing a
1663
+ handshake. OpenSSL 3.x ships security level 2, which refuses the key sizes and
1664
+ signature algorithms that software of the era this tool gets pointed at still
1665
+ offers — Webmin 1.910, the build the README's `reflected` row rests on,
1666
+ answers a default client with `SSLV3_ALERT_HANDSHAKE_FAILURE` and nothing
1667
+ else. Every probe then came back `error`: correct, and useless. The run was
1668
+ honest about having measured nothing, and the sink behind that handshake was
1669
+ never tested at all. Found by running the coverage benchmark, which failed at
1670
+ its readiness gate against a container that was up and answering.
1671
+
1672
+ `--insecure` now also lowers the security level and the minimum protocol
1673
+ version. This does not widen exposure: with `check_hostname = False` and
1674
+ `CERT_NONE` the connection is already unauthenticated, so an active attacker
1675
+ is already unconstrained — accepting a 1024-bit key or a SHA-1 signature on
1676
+ top of that gives away nothing that was still being held. What it buys is the
1677
+ difference between testing the target and reporting that it could not be
1678
+ reached. A run without the flag is untouched and still verifies certificates.
1679
+
1680
+ - **The benchmark's readiness gate is as permissive as the tool it gates.**
1681
+ `wait_for_target` built its own strict context, so a case against deliberately
1682
+ old software reported "target never became ready" about a container that was
1683
+ up — a case failure with nothing wrong in it.
1684
+
1685
+ ### Changed
1686
+
1687
+ - A run that passes `--insecure` states the full extent of the downgrade on its
1688
+ first line. The flag gives up more than certificate identity now, and an
1689
+ operator on a monitored engagement should read that in the transcript rather
1690
+ than infer it from the help text.
1691
+
1692
+ ## [2.35.1] — 2026-08-21
1693
+
1694
+ A robustness pass over error handling: no new capability, four ways the tool
1695
+ could crash or mislead on input it did not choose.
1696
+
1697
+ ### Fixed
1698
+
1699
+ - **A truncated error response no longer ends the run.** Reading an
1700
+ `HTTPError`'s body happens *inside* the `except` handler, where the sibling
1701
+ `except Exception` cannot reach it — so a target that promised a
1702
+ `Content-Length` it never delivered raised `ConnectionResetError` straight out
1703
+ of `main()`, taking every probe already fired with it. The read is now
1704
+ guarded: the status still comes back, an unreadable body is reported empty.
1705
+ A body that *does* arrive is still returned in full — the 500-stack-trace
1706
+ confirmations that branch exists for are unaffected.
1707
+ - **A `--target-profile` is checked before it is used.** A profile is written by
1708
+ hand, so a typo in one is ordinary; it surfaced as a traceback. A top level
1709
+ that is not a JSON object, `deny_chars` that is not text, a `max_length` that
1710
+ is not a number, a selector field that is not a list of names — each is now an
1711
+ operator-readable `[!]` message and exit 1, the same way the sink-shape fields
1712
+ already behaved. Twelve inputs that produced an `AttributeError` or a
1713
+ `TypeError` now produce a sentence.
1714
+ - **A selector field given as one string means one name.** `"environments":
1715
+ "unix"` in a profile was iterated character by character, matched nothing, and
1716
+ the empty run that followed was reported as a success. A string is now split
1717
+ on commas: `"unix"` is `["unix"]`, `"raw, html"` is `["raw", "html"]`.
1718
+ - **Unknown `--environments` and `--encodings` are named.** Both were silent, so
1719
+ `--environments linux` — the corpus calls it `unix` — produced an empty file
1720
+ and exit 0, indistinguishable from a target with no payloads for it. Both now
1721
+ warn and list the names that exist, as unknown contexts and categories already
1722
+ did. An empty result is also no longer announced as "Successfully generated 0
1723
+ payloads"; it says the selection matched nothing and points at the filters.
1724
+ - **A callback cannot rewrite the operator's terminal.** The host and path of an
1725
+ OOB callback are chosen by the target. Printed raw, an ESC byte let that
1726
+ target colour, erase and rewrite lines — hiding a genuine `[HIT]` behind
1727
+ `\x1b[2K\r`, or forging one that never arrived. Control characters are now
1728
+ escaped for display as `\xNN`; the recorded hit and the `--listen-log` JSONL
1729
+ keep the bytes verbatim.
1730
+
1731
+ ## [2.35.0] — 2026-08-20
1732
+
1733
+ ### Added
1734
+
1735
+ - **RCEKit is installable from PyPI**: `pipx install rcekit` (or
1736
+ `pip install rcekit`) puts an `rcekit` command on PATH. Published through PyPI
1737
+ Trusted Publishing from a GitHub release — no API token, no repository secret.
1738
+
1739
+ It ships as a **single-module distribution**, not a package tree. `rcekit.py`
1740
+ stays one file at the repo root and still runs alone from a `curl` on a jump
1741
+ box or an air-gapped host; installing is a second supported shape, not a
1742
+ replacement for the first.
1743
+
1744
+ ### Changed
1745
+
1746
+ - **`import rcekit` no longer has side effects.** Logging was configured at
1747
+ module scope, and `logging.FileHandler` opens its file when it is constructed,
1748
+ so merely importing the module wrote `rcekit.log` into whatever directory the
1749
+ interpreter happened to be in. Handler setup moved into `configure_logging()`,
1750
+ called from `main()`. Running the CLI still writes `rcekit.log` exactly as
1751
+ before.
1752
+
1753
+ - **`main()` takes an optional `argv`** and returns an explicit `int`, so the
1754
+ console-script entry point is a plain zero-argument call and tests can drive
1755
+ the CLI in-process. Every exit code is unchanged.
1756
+
1757
+ - **The built-in corpus is no longer reported as a missing file.** With nothing
1758
+ but `rcekit.py` — an installed wheel, or the single-file copy — there is no
1759
+ `templates/` directory, the embedded corpus *is* the corpus, and the run is
1760
+ now silent about it; `--doctor` names it `built-in (embedded in rcekit.py)`
1761
+ and reports OK. A `templates/` directory that exists *without* its
1762
+ `payloads.json` still prints the notice, because that one is a real finding.
1763
+ A corpus that is present but corrupt, and an explicit `--template-file` that
1764
+ is missing or corrupt, still refuse to run and exit non-zero.
1765
+
1766
+ - The log file handler now runs at `DEBUG` while the console stays at `INFO`, so
1767
+ detail worth having when reconstructing a run no longer lands on the
1768
+ operator's terminal.
1769
+
1770
+ ## [2.34.1] — 2026-08-17
1771
+
1772
+ Documentation only; no behaviour change.
1773
+
1774
+ ### Changed
1775
+
1776
+ - **New tagline: "`confirmed` means the target executed the input. `negative`
1777
+ means the probes reached it."** The old one — "prove RCE, don't guess it" —
1778
+ claimed the tool always proves. It does not, and does not need to: the value
1779
+ is that each verdict has a mechanical meaning, in both directions. A promise
1780
+ can be broken; a definition cannot.
1781
+
1782
+ Both halves name their actor and object on purpose. "Executed" alone reads as
1783
+ though *RCEKit* executed something; the claim is about the target. And
1784
+ `confirmed`/`negative` are the verdict values as the code spells them, not
1785
+ looser words like "clean".
1786
+
1787
+ - **The README caught up with the engine.** It still described "two verdict
1788
+ tiers" when there are seven, and the comparison table predated the last five
1789
+ releases. Rewritten around what a verdict asserts, with the tier table as the
1790
+ centrepiece and a section on the half no other tool has: `error` and
1791
+ `nothing-tested` exist so a run that tested nothing is never reported as
1792
+ clean.
1793
+
1794
+ The comparison table gains the classes added since it was written — Windows
1795
+ `cmd.exe`/PowerShell sinks, upload → write-then-execute, second-order
1796
+ execution, query-language bridges and deserialization sinks — and the "reach
1797
+ for something else" note now says plainly that sqlmap owns the database and
1798
+ RCEKit's bridges only prove the OS is reachable from a text parameter.
1799
+
1800
+ ### Fixed
1801
+
1802
+ - Two claims in the engagement-controls table were wrong and are now accurate:
1803
+ the observed-channel fetch sends **no** credentials unless given a request with
1804
+ `--observe-request` (only the `file` read-back inherits them, same-origin), and
1805
+ an unanswered `--observe-url` is a warning about a partly blinded run rather
1806
+ than a `nothing-tested` verdict.
1807
+ - The "mechanisms that produce `inconclusive`" list said four and listed five.
1808
+
1809
+ ## [2.34.0] — 2026-08-17
1810
+
1811
+ Deserialization **sink** detection, and a verdict that is deliberately not RCE.
1812
+
1813
+ Deserialization RCE (fastjson, shiro, weblogic, jenkins) cannot be confirmed by
1814
+ the value-oracle model: the payload is a serialized object graph and gadget
1815
+ selection is classpath-specific, so whether execution is reachable depends on
1816
+ jars RCEKit cannot see. That stays out of scope. The honest middle step is
1817
+ showing the endpoint parses the data at all — a real finding, and the
1818
+ prerequisite for every gadget chain.
1819
+
1820
+ ### Added
1821
+
1822
+ - **`--methods deser`**, which **never emits `confirmed`**. Its strongest
1823
+ outcome is a new verdict, `deserialization-sink`, reported in its own section
1824
+ that states outright that reaching RCE from there depends on classpath
1825
+ gadgets. Collapsing it into `confirmed` would break the one guarantee the tool
1826
+ rests on; collapsing it into `needs-review` would throw away a proven finding.
1827
+
1828
+ `deserialization-sink` sits below both RCE tiers in the collapsed verdict: it
1829
+ is proven, but a *suspected* RCE outranks a proven non-RCE in triage.
1830
+
1831
+ - **Two oracles of deliberately different strength.** `shape` (no listener
1832
+ needed) sends a well-formed object stream, the same stream truncated, and the
1833
+ format's magic bytes plus random noise of the same length, and asks whether
1834
+ the endpoint answers the well-formed one differently from both — a
1835
+ fingerprint, so `needs-review` only, never promoted. `dns` (needs
1836
+ `--oob-host`) sends a gadget whose only side effect is a name lookup.
1837
+
1838
+ - **A URLDNS builder for Java serialization.** A `HashMap` holding one
1839
+ `java.net.URL`: `HashMap.readObject` hashes the key, `URL.hashCode` asks for
1840
+ the host address, the JVM resolves the name. It references no class outside
1841
+ `java.util`/`java.net`, so there is nothing in it that can run — the callback
1842
+ proves the object graph was reconstructed and no more.
1843
+
1844
+ Built in Python rather than declared in the corpus because the URL host is
1845
+ length-prefixed *inside* the stream and changes per probe. Its constant parts
1846
+ are the exact bytes OpenJDK's own `ObjectOutputStream` emits for that graph,
1847
+ and the result was **verified against OpenJDK 21**: it deserializes to
1848
+ `HashMap{http://<host>/=rk}` and issues a DNS query for `<host>`, with no code
1849
+ execution.
1850
+
1851
+ - **A `deser_probes` corpus section** with `java`, `php`, `dotnet`,
1852
+ `python_pickle` and `fastjson`, plus `--deser-formats` to narrow it. Only
1853
+ `java` and `fastjson` carry a DNS gadget: PHP and .NET chains all run through
1854
+ magic methods or type confusion, so there is no honest DNS-only probe for them
1855
+ and they get the shape oracle alone.
1856
+
1857
+ - Response signatures for the shape differential drop long digit and hex runs,
1858
+ so request ids and timestamps on an otherwise identical error page do not make
1859
+ every endpoint fingerprint as a parser.
1860
+
1861
+ ### Changed
1862
+
1863
+ - The README's scope note now says precisely what changed and what did not:
1864
+ deserialization **gadget chains** remain out of scope, while the **sink** is
1865
+ now reported in its own tier.
1866
+
1867
+ ## [2.33.0] — 2026-08-17
1868
+
1869
+ Query-language bridges. Several RCEs pass through a query language before
1870
+ reaching the OS — Postgres `COPY … FROM PROGRAM`, MSSQL `xp_cmdshell`, XXE
1871
+ `expect://` — and the injection point is an ordinary text value, so the oracle
1872
+ model already fitted. Only the carriers were missing.
1873
+
1874
+ ### Added
1875
+
1876
+ - **A `bridges` section in the corpus**, declared like `eval_carriers` so
1877
+ coverage grows without touching Python. Each bridge names the shell it
1878
+ reaches, its safety tier, its prerequisites and, where it creates something,
1879
+ the statement that removes it: `postgres_copy_program` (`/bin/sh`,
1880
+ `stateful`), `mssql_xp_cmdshell` (`cmd.exe`, `intrusive`) and `xxe_expect`
1881
+ (`/bin/sh`, `intrusive`).
1882
+
1883
+ - **`--bridges none|auto|NAMES`** rides the command probes through them. A
1884
+ bridge is a **carrier, not an oracle**: it wraps the command `reflected`,
1885
+ `time` and `oob` already build, so those methods prove execution through it
1886
+ and inherit every tier guarantee rather than re-deriving one. Off by default,
1887
+ because a bridge payload is SQL or XML syntax and on an ordinary shell sink it
1888
+ is a request that cannot confirm.
1889
+
1890
+ Three properties follow from that framing. A bridge only gets a core written
1891
+ in its own dialect — `xp_cmdshell` hands its argument to `cmd.exe`, so pairing
1892
+ it with a POSIX `$((a+b))` would send inert text. No separator is prepended:
1893
+ inside `COPY … FROM PROGRAM '…'` there is no running command to break out of.
1894
+ And the record's context still applies, so `--contexts sql` and a bridge
1895
+ compose instead of each reinventing the other.
1896
+
1897
+ - **The safety ordering governs bridges** exactly as it governs every corpus
1898
+ payload: a `stateful` bridge needs `--verify-active-risk stateful`, and the
1899
+ pre-flight names the tier each held-back bridge actually requires rather than
1900
+ sending the operator to raise the ceiling further than the run needs.
1901
+
1902
+ ### Changed
1903
+
1904
+ - **An aggregate method's result now carries its cleanup line.** `time` reports
1905
+ one row for a whole probe series, so a stateful bridge on the one oracle that
1906
+ reliably proves a query-language sink was the one that never said how to clean
1907
+ up after itself.
1908
+
1909
+ ### Not built, deliberately
1910
+
1911
+ - **MySQL UDF execution** is a multi-stage chain — write a shared object into the
1912
+ plugin directory, then `CREATE FUNCTION` — not something a single probe can
1913
+ carry. There is no stub for it.
1914
+ - **MongoDB `$where`** is a boolean-only channel (its JS sandbox cannot reach a
1915
+ shell), so it needs a different oracle rather than this one.
1916
+ `mongo-express/CVE-2019-10758` is a plain JS `eval` sink that `--methods eval`
1917
+ already covers.
1918
+
1919
+ The three shipped bridges are documented syntax but **not validated against live
1920
+ databases here** — this build environment has no container runtime. Each corpus
1921
+ entry says so in its `verified` field rather than implying a test that did not
1922
+ happen.
1923
+
1924
+ ## [2.32.0] — 2026-08-17
1925
+
1926
+ The second-order oracle. Execution frequently happens on a **different request**
1927
+ than injection — stored SSTI rendered on a profile page, a payload written to a
1928
+ log a template engine later renders, a queued job run asynchronously. The engine
1929
+ diffs the response it injected into, so every one of those read `negative`
1930
+ however exploitable the target was.
1931
+
1932
+ ### Added
1933
+
1934
+ - **`--observe-url URL`** names the endpoint where the execution surfaces. It is
1935
+ read after each probe and then polled after the batch, and a probe whose
1936
+ computed value turns up there is upgraded to `confirmed`.
1937
+
1938
+ It stays fully differential, which is why it reaches `confirmed` rather than
1939
+ `needs-review`: the value was computed locally from operands random to that
1940
+ probe, it must be absent from a snapshot of the endpoint taken **before any
1941
+ probe was sent**, and — the rule that carries the weight — a probe's value is
1942
+ looked for there **only when the probe's own payload does not contain it**.
1943
+
1944
+ Without that last rule the oracle would be a false-positive generator: `file`
1945
+ and `oob` expect a random token that sits verbatim in the payload, so a target
1946
+ that merely stores the payload and renders it back would hand that token
1947
+ straight to the observed page and every such probe would confirm without
1948
+ executing anything. Measured against a store-and-echo target: **0**
1949
+ confirmations. The computed-value methods pass the same rule for the opposite
1950
+ reason — reflection returns `$((a+b))`, never the sum — so it selects them
1951
+ without naming them, and a method added later inherits the right answer.
1952
+
1953
+ - **`--observe-request FILE`** takes a captured request instead, for the common
1954
+ case where the page a stored payload renders on is behind a login. It needs no
1955
+ `FUZZ` marker: the observed endpoint is read, never injected into.
1956
+
1957
+ - **`--observe-poll` / `--observe-timeout`** control the polling window
1958
+ (defaults 5s and 60s). One poll always happens, even at a zero timeout.
1959
+
1960
+ - Every probe result carries an `observe_status` in `--detect-json`:
1961
+ `confirmed`, `polled` (read, value not there), `in-control`, `not-observed`
1962
+ (not eligible) or `unreachable`. When the endpoint never answered, the run
1963
+ says so outright — negatives decided without ever reading the observed channel
1964
+ are not second-order negatives.
1965
+
1966
+ ### Changed
1967
+
1968
+ - The observed channel is read once after **each** probe as well as polled after
1969
+ the batch, so a run with `--observe-url` sends roughly twice the requests.
1970
+ Batch-then-poll alone is only correct for a channel that *accumulates* (a log,
1971
+ a comment list); where the store overwrites — a profile field, which is the
1972
+ shape this oracle most exists for — every probe but the last is gone by the
1973
+ time the batch poll runs, and the oracle confirmed nothing. The extra read is
1974
+ skipped for probes that are already confirmed in-band or not eligible, so
1975
+ `file` and `oob` add none.
1976
+
1977
+ Observing is additive throughout: the in-band verdict is computed exactly as
1978
+ before and only a non-`confirmed` one can be upgraded, so a run without the flag
1979
+ is byte-for-byte unchanged and a run with it can only gain findings.
1980
+
1981
+ ## [2.31.0] — 2026-08-17
1982
+
1983
+ The `write` method: a write primitive proven to be RCE by executing what it
1984
+ wrote. A whole family of targets was invisible — `tomcat/CVE-2017-12615` (PUT a
1985
+ JSP), `activemq/CVE-2016-3088`, `weblogic/CVE-2018-2894` — because the vulnerable
1986
+ request *stores a file* rather than evaluating anything. Nothing is computed in
1987
+ its response, so `reflected` and `eval` correctly returned `negative` on targets
1988
+ that are fully exploitable.
1989
+
1990
+ ### Added
1991
+
1992
+ - **`--methods write`** — the inverse of `file`. `file` assumes execution exists
1993
+ and uses a write as proof of it; `write` assumes a write primitive exists and
1994
+ uses execution of the written file as proof of RCE. The probe is the file's
1995
+ *content*: a one-liner computing a product on random operands, delivered
1996
+ through the ordinary injection point.
1997
+
1998
+ The fetched file is read in three tiers, and the middle one is the reason the
1999
+ method exists:
2000
+
2001
+ | fetched file contains | verdict | means |
2002
+ |---|---|---|
2003
+ | the product | `confirmed` | written **and** executed |
2004
+ | the one-liner, verbatim | `needs-review` | arbitrary file write, not interpreted |
2005
+ | neither | `negative` | no write, or not served there |
2006
+
2007
+ An upload directory that is served but not interpreted is a real finding and
2008
+ is not remote code execution, so the tiers are never merged.
2009
+
2010
+ - **`--write-url-template URL`** names where the stored file is served — the
2011
+ channel the proof comes back on, and the flag the method is gated on.
2012
+
2013
+ - **`--write-lang`** picks the file types: `auto` (default) reads the extension
2014
+ off the read-back URL, or name any of `jsp`, `jspx`, `php`, `aspx`, `erb`.
2015
+ `jsp`/`aspx`/`erb` share the `<%= %>` delimiters, so their probes are
2016
+ byte-identical and cost one request between them; with no extension to read,
2017
+ `auto` writes all five in three requests.
2018
+
2019
+ ### Changed
2020
+
2021
+ - **A `needs-review` finding now prints its cleanup line too.** It used to
2022
+ appear only under `confirmed`, which was already thin and is wrong for this
2023
+ method: a `write` reaching `needs-review` means the file *is* on the target,
2024
+ just not interpreted, so the artifact would have been left there unmentioned.
2025
+
2026
+ - The write method's operands are drawn once per run rather than once per
2027
+ carrier, so the file is written once instead of once for each of the ~13
2028
+ `(environment, context)` carriers. For a state-changing method that is not a
2029
+ request-count saving, it is a blast radius. Still fresh per run, which is what
2030
+ makes the product unforgeable.
2031
+
2032
+ - The write method declines the break-out contexts (`sql`, `javascript`,
2033
+ `shell_*`, …) and keeps the transport ones. Its payload is a whole file body:
2034
+ there is nothing to break out of, and wrapping it in `'; … -- ` would write a
2035
+ broken file. A run narrowed past `raw` and the transport contexts is told so
2036
+ rather than reporting a clean negative.
2037
+
2038
+ ## [2.30.0] — 2026-08-17
2039
+
2040
+ Per-dialect shell probes. `$((a+b))`, `sleep` and `$(echo TAG)` are POSIX
2041
+ constructs: on a cmd.exe or PowerShell sink they are inert literal text. The
2042
+ dialect was inferred from the corpus environment alone, so a run could send a
2043
+ probe no shell on the target would ever execute — including on the carrier whose
2044
+ context is literally named `powershell`.
2045
+
2046
+ ### Added
2047
+
2048
+ - **`--sink-env auto|unix|windows|powershell`** states which shell runs the
2049
+ injected command. The computed-value core, the separators and the break-out
2050
+ contexts are all chosen from it. `auto` (the default) infers it per carrier;
2051
+ pin it when the corpus environment names the *application runtime* rather than
2052
+ the OS — `--environments php --sink-env windows` is a PHP application on IIS,
2053
+ which no inference can see.
2054
+
2055
+ - **A PowerShell probe shape for every shell method**, validated against
2056
+ pwsh 7.4: `Write-Output T1$(a*b)T2` for `reflected` (an unquoted argument is an
2057
+ expandable string, so the core carries no quote and the quote-wrapping
2058
+ contexts can still carry it), `Start-Sleep -Milliseconds N` for `time`,
2059
+ `Set-Content` for `file` and `iwr -useb` for `oob`. PowerShell was previously
2060
+ reachable by no probe in any method.
2061
+
2062
+ - **cmd.exe and PowerShell carriers for the `dotnet` environment.** It is the
2063
+ one corpus environment that names a platform, and it was taking the POSIX
2064
+ shape — so .NET on Windows, the case the environment exists for, was the case
2065
+ it could not confirm on. Every other runtime keeps the POSIX shape: a language
2066
+ does not say which OS it runs on.
2067
+
2068
+ ### Fixed
2069
+
2070
+ - **The `powershell` carrier was written in cmd.exe.** Every `windows` carrier
2071
+ took the `for /f ... ('set /a a+b')` core regardless of context, so the one
2072
+ carrier explicitly shaped for PowerShell sent a payload PowerShell cannot
2073
+ execute. The dialect now follows the carrier's context first, and a carrier's
2074
+ break-out variants stay in its dialect rather than re-deriving from the
2075
+ environment.
2076
+
2077
+ - **`Set-Content`, not `>`, for the PowerShell write.** In Windows PowerShell
2078
+ 5.1 the redirect is `Out-File`, whose default encoding is UTF-16LE: the write
2079
+ lands and the read-back still does not find the token, so the probe reports
2080
+ negative on a target it owns.
2081
+
2082
+ ### Changed
2083
+
2084
+ - **cmd.exe no longer gets the `sq`, `dq` and `subshell` carriers.** It has
2085
+ neither a comment character to swallow the sink's tail nor a
2086
+ command-substitution syntax, so those four carriers per Windows run were
2087
+ requests that could only come back negative. PowerShell takes the quote
2088
+ break-outs and `$( )` — both measured — but not the backtick, which is its
2089
+ escape character rather than a substitution.
2090
+
2091
+ - **PowerShell's separator sweep carries no pipe.** `cmd | Start-Sleep
2092
+ -Milliseconds 500` is a parameter-binding error, not a fresh command with
2093
+ stdin attached the way a POSIX pipe is, and it fails that way for every cmdlet
2094
+ the probes use. `;`, a newline and (on PowerShell 7) `&&`/`||` remain.
2095
+
2096
+ - The pre-flight plan prints the sink shell alongside the sink shapes, and a
2097
+ pinned dialect narrows the printed ladder to the rungs it has syntax for.
2098
+
2099
+ ## [2.29.0] — 2026-08-17
2100
+
2101
+ Generalised read-back for the `file` method. It required a writable **web root**
2102
+ the tester already knew, which ruled out every other way a target can hand a
2103
+ file back — on exactly the internal, no-egress targets the method exists for.
2104
+
2105
+ ### Added
2106
+
2107
+ - **`--file-write-path DIR` + `--file-read-url URL`** name the two halves of the
2108
+ read-back channel directly, so an LFI endpoint, a download or export handler,
2109
+ an attachment fetcher or a `/tmp`-backed preview all work. The template takes
2110
+ `{name}` (the filename), `{path}` (the full server-side path) and `{path_enc}`
2111
+ (that path percent-encoded); only those three are substituted, so a URL that
2112
+ legitimately contains braces survives unchanged.
2113
+
2114
+ Measured against a target with a download handler and nothing serving the
2115
+ write directory: the web-root form confirms **0** — reporting an exploitable
2116
+ target clean — and the general form confirms **7**.
2117
+
2118
+ ### Fixed
2119
+
2120
+ - **The read-back fetch now carries the run's headers**, so an authenticated
2121
+ download, export, attachment or LFI handler can actually be read. It went out
2122
+ bare, which barely mattered while the channel had to be a web root — static
2123
+ file serving is rarely authenticated — and became the likely case the moment
2124
+ the channel could be an application endpoint. Measured against a handler
2125
+ behind a bearer token: the write executed on every probe and the verdict was
2126
+ `negative`, "token absent from the fetched file". Now 7 confirmations on the
2127
+ same target.
2128
+ - **Credentials are carried only to the same origin.** A read-back URL on
2129
+ another host is someone else's server, and replaying the target's session
2130
+ cookie or bearer token to it would leak the credential, so those headers are
2131
+ dropped there while the rest still go — and the run says so, because the
2132
+ symptom would otherwise look like a clean target. `Content-Type` and
2133
+ `Content-Length` are dropped from the fetch too: they describe a body the GET
2134
+ does not have.
2135
+
2136
+ ### Changed
2137
+
2138
+ - **`--webroot` / `--web-base-url` are now the web-root alias** for the general
2139
+ form: a web root is just the case where the read URL is the base plus the
2140
+ filename. Existing command lines are unaffected. Both are resolved in one
2141
+ place inside the method, so the alias and the general form cannot drift — and
2142
+ the gate, the pre-flight banner and the blind-sink advice all ask that same
2143
+ resolver instead of testing for the webroot pair.
2144
+ - `blind_sink_advice` reads its flags defensively, so an args-like object
2145
+ missing a newer field costs a line of advice rather than a traceback.
2146
+
2147
+ ## [2.28.0] — 2026-08-17
2148
+
2149
+ Injection-point enumeration. `-p NAME` needed the tester to already know which
2150
+ parameter was the sink, so a capture's other candidates — including the headers
2151
+ and nested JSON leaves that carry some of the highest-value classes — were never
2152
+ tried.
2153
+
2154
+ ### Added
2155
+
2156
+ - **`-p all` / `--auto-params KINDS`** expands one captured request into every
2157
+ candidate injection point and runs the selected `--methods` against each.
2158
+ Query values, JSON leaves addressed by path (`user.profile.name`, `tags[1]`),
2159
+ form fields, cookie crumbs and headers, each rewritten in **its own**
2160
+ serialization rather than blanket-encoded. Verified end to end: a sink
2161
+ reachable only through `User-Agent` is confirmed from `-r request.txt -p all`
2162
+ with no manual header selection.
2163
+ - **`--point-order fast|thorough`** — `fast` tries a curated high-yield header
2164
+ list (the headers real published RCEs inject through); `thorough` adds every
2165
+ remaining non-hop-by-hop header. **`--max-points N`** bounds the run and
2166
+ reports what it dropped. **`--include-path-segments`** is opt-in, because
2167
+ rewriting a path segment usually just produces a 404.
2168
+ - **The run states its cost before sending it** —
2169
+ `6 points x ~61 probes = at least 372 requests` — via a new
2170
+ `estimate_detection_probes`, which builds the probes and counts them without
2171
+ firing any. Enumeration multiplies an already-laddered probe count by the
2172
+ candidate count, and an operator on a monitored engagement has to see that
2173
+ before it happens rather than infer it from the traffic.
2174
+ - **Findings name the point they came from**: `[reflected/unix/raw] at header
2175
+ 'User-Agent' ...`.
2176
+
2177
+ ### Changed
2178
+
2179
+ - **Each candidate carries its own payload-free control.** Differencing a header
2180
+ probe against a query probe's control would compare two different responses
2181
+ and prove nothing.
2182
+ - **Cheap methods run first per candidate, and a candidate stops at its first
2183
+ confirmation.** `reflected` and `eval` cost one response each; `time` sleeps
2184
+ and `oob` waits for a callback, and on a candidate that has already proven
2185
+ execution those buy a second name for the same finding. Candidates that stay
2186
+ clean still get every method, and single-point runs are unchanged.
2187
+ - A JSON leaf is **replaced, never created**. Assigning to a missing key would
2188
+ have injected into a field the application never sends — a probe that cannot
2189
+ say anything about the parameter that does exist. Caught by its own test.
2190
+ - **JSON points are addressed by tokens, not by a joined path string.** A key may
2191
+ itself contain the separator: `{"user.name": ..., "user": {"name": ...}}`
2192
+ rendered *both* leaves as `user.name`, so the literal key was never probed and
2193
+ both candidates mutated the nested field — a false negative and a misattributed
2194
+ finding at once. Tokens remove the ambiguity, and the display form
2195
+ bracket-quotes such a key (`["user.name"]`) so the two stay distinguishable on
2196
+ screen.
2197
+ - **A deeply nested captured body no longer ends `-p all` with a traceback.**
2198
+ `json.loads` recurses in C, so `RecursionError` joins the caught exceptions in
2199
+ both the enumerator and the placer, as it already had in the response-channel
2200
+ parser. The body yields no candidates; the rest of the request still enumerates.
2201
+ - **The cost estimate honours `--max-payloads`.** It counted every probe the
2202
+ carriers could produce while the run stops at the cap, so the figure was wrong
2203
+ exactly when the operator had reached for the budget guard.
2204
+ - `Host`, `Content-Length`, `Cookie` and the hop-by-hop headers are never
2205
+ candidates: injecting into those changes the request's plumbing rather than
2206
+ testing the application, and two of them are rebuilt by the delivery layer.
2207
+
2208
+ ## [2.27.0] — 2026-08-17
2209
+
2210
+ Engine carriers for the `eval` probe. Three template engines evaluate the
2211
+ injected expression perfectly and still made RCEKit report `negative`, because
2212
+ what came back was not the bare product the oracle searches for.
2213
+
2214
+ ### Added
2215
+
2216
+ - **`eval_carriers` in the corpus**, and `--eval-engines auto|<names>` to select
2217
+ them. A carrier wraps the same random-operand arithmetic in an engine-specific
2218
+ form; it never changes the oracle, and the bare probes still run first. Each
2219
+ entry records `notes` (why it exists) and `verified` (what it was measured
2220
+ against). Declarative, so a new carrier is a JSON entry rather than a code
2221
+ change.
2222
+
2223
+ | Engine | Bare `${a*b}` returned | Carrier | Carrier returned |
2224
+ |---|---|---|---|
2225
+ | Freemarker | `2,070,761,401` (locale grouping) | `${(a*b)?c}` | `2070761401` |
2226
+ | Velocity | `${a*b}` verbatim — a *reference*, not an expression | `#set($rk=a*b)$rk` | `2070761401` |
2227
+ | Thymeleaf | `${a*b}` verbatim — needs inlining brackets | `[[${a*b}]]` | `2070761401` |
2228
+
2229
+ Measured against freemarker 2.3.32, velocity-engine-core 2.3 and thymeleaf
2230
+ 3.1.2, running RCEKit's own generated probes through each engine: bare form
2231
+ `CONFIRMS=no`, carrier `CONFIRMS=YES`, for all three.
2232
+ - **The evidence line names the carrier** — `target computed '3979016' via the
2233
+ freemarker carrier` — so a finding says which engine quirk it worked around.
2234
+ A bare confirmation reads exactly as before.
2235
+
2236
+ ### Notes
2237
+
2238
+ - **Carriers are not sandbox escapes, and no sandbox-escape carrier ships.** The
2239
+ premise that a sandboxed engine blocks the arithmetic probe did not survive
2240
+ measurement: a member-access sandbox restricts method and field access, and
2241
+ arithmetic needs neither. With OGNL member access denied for *everything*,
2242
+ `40277*51413` still returned `2070761401` while `@java.lang.Math@max(1,2)` was
2243
+ blocked; SpEL's restricted `SimpleEvaluationContext` and Jinja2's
2244
+ `SandboxedEnvironment` behaved the same way. The bare probes already cover
2245
+ those engines.
2246
+ - The frequently-cited OGNL escape `(#_memberAccess=@ognl.OgnlContext@DEFAULT_MEMBER_ACCESS)`
2247
+ additionally targets a field that **no longer exists in OGNL 3.3.4**, so on a
2248
+ current engine it is a probe that can only come back negative.
2249
+
2250
+ ## [2.26.0] — 2026-08-16
2251
+
2252
+ The sink-shape ladder. An injected value lands in a *shape* — mid-command,
2253
+ inside quotes, as the whole command — and the shape decides what can reach it.
2254
+ Two shapes had no probe that fitted, so a genuinely exploitable target reported
2255
+ clean.
2256
+
2257
+ ### Added
2258
+
2259
+ - **`--sink-shape auto|sep|raw|chain|newline|dq|sq|subshell`** (comma-separated)
2260
+ names which shapes the shell probes try. `auto` is the whole ladder and the
2261
+ default. Underneath it selects the existing separator sweep and break-out
2262
+ contexts, so naming a rung narrows a supported run rather than switching on a
2263
+ parallel path. The plan is printed before anything is sent, because the ladder
2264
+ multiplies request count and an operator on a monitored engagement needs to
2265
+ see the cost first.
2266
+ - **The `subshell` rung — `$(...)` and backticks.** Reaches a value sitting
2267
+ inside double quotes *without closing the quote*, which is the one case a
2268
+ quoted break-out loses to a filter on the quote character itself. Measured
2269
+ against `system("echo PING \"$input\"")`: with `"` stripped, `dq` is inert and
2270
+ both substitution forms execute; with `$` stripped, `dq` executes and the
2271
+ backtick form still does. Both ship because they survive different filters.
2272
+
2273
+ **Which method it helps is the counter-intuitive part.** `reflected`'s core is
2274
+ `$((a+b))`, which the shell expands inside double quotes anyway, so that
2275
+ method already confirmed there. The methods whose core must actually *run* —
2276
+ `time` (a sleep), `file` (a redirect), `oob` (a fetch) — are completely inert
2277
+ inside those quotes. On a quote-filtering sink, `--methods file` went from **0
2278
+ confirmations to 2**: it had been reporting an exploitable target as clean.
2279
+ - **The `raw` rung is now part of `auto`, for every shell method.** A
2280
+ `qx/$input/`-style sink, where the input is the whole command, previously
2281
+ needed `--sink-raw` — so it reported clean unless the operator already
2282
+ suspected the shape. One extra probe per carrier buys it. `--sink-raw` keeps
2283
+ its meaning as the narrowing alias for `--sink-shape raw`, and no existing
2284
+ command line changes behaviour. `reflected`, `file`, `time` and `oob` all
2285
+ build their candidates through one `_separator_candidates` helper, so a rung
2286
+ cannot reach some methods and not others; `time` screens it in its second
2287
+ wave, alongside the separators it holds back.
2288
+
2289
+ ### Fixed
2290
+
2291
+ - **A method that builds no probes no longer reports `negative`.** An aggregate
2292
+ method asked to judge zero samples answers honestly — "no delay was observed",
2293
+ "no callback arrived" — and that reads as "not vulnerable" from a run that
2294
+ tested nothing. The engine now emits no row for a carrier that produced no
2295
+ probes, which lets its own loud nothing-tested path fire instead. Reachable
2296
+ through any narrowing that leaves a carrier with nothing to send.
2297
+
2298
+ ### Changed
2299
+
2300
+ - **The pre-flight sink-shape plan is computed from the effective run**, not
2301
+ from the `--sink-shape` value. `--separators`, `--contexts` and `--sink-raw`
2302
+ each narrow the ladder, so printing the flag described a run that would not
2303
+ happen — and this output is presented as an audit of the traffic about to be
2304
+ sent. `effective_sink_shapes` is the single source of truth the engine and the
2305
+ plan both read.
2306
+ - **The backtick context drops probe shapes that carry their own backtick.**
2307
+ Backticks do not nest, so such a probe closes the outer substitution early and
2308
+ could only ever come back negative. `$( )` does nest and keeps every shape.
2309
+ - **Naming `--separators` now implies the sink is separator-led**, so the `raw`
2310
+ rung is dropped unless `--sink-shape` names it explicitly. A profile with
2311
+ `sink_needs_separator` drops it for the same reason. Both keep an explicitly
2312
+ narrowed run from being widened behind the operator's back.
2313
+
2314
+ ## [2.25.0] — 2026-08-16
2315
+
2316
+ A coverage benchmark, so a claim about what RCEKit confirms can be checked
2317
+ instead of asserted. The unit suite proves the tool reaches the right verdict
2318
+ from a given response; it cannot prove it confirms Webmin.
2319
+
2320
+ ### Added
2321
+
2322
+ - **`--detect-json PATH`** writes a detection run as JSON: the run's overall
2323
+ verdict, per-verdict counts, and every probe with its payload, method, context
2324
+ and evidence. Text output is unchanged. This is the supported way to consume a
2325
+ run programmatically — scraping stdout cannot be made reliable, because a
2326
+ probe payload may contain a literal newline (the newline separator is a real
2327
+ one, so line-oriented parsing splits a payload in half) and the detection path
2328
+ exits 0 whether it confirmed or came back clean.
2329
+ - **`tests/bench/` — the coverage benchmark harness.** Each case brings a real
2330
+ vulnerable build up, runs RCEKit as an operator would, checks the verdict, and
2331
+ tears it down; `--markdown` emits the coverage table. Not part of
2332
+ `python -m unittest discover -s tests` — cases need Docker and pull real
2333
+ images — so it runs by hand or in a dedicated job, and exits non-zero if any
2334
+ case fails. Two cases ship, transcribed from `docs/verify-it-yourself.md`:
2335
+ Webmin CVE-2019-15107 and Struts2 S2-001.
2336
+ - **A negative control is a required key.** A benchmark without controls measures
2337
+ nothing: a tool that shouted `confirmed` at every target would score full marks
2338
+ on the vulnerable half. Three kinds are supported — a patched build, the same
2339
+ target probed for the wrong class, and a weaker method that must stay below
2340
+ `confirmed` on a target where it happens to be right. The runner refuses four
2341
+ shapes of non-control: no control at all; one expecting `confirmed`; one that
2342
+ runs the identical invocation against an identical target (judged on what it
2343
+ would actually run, so an explicit copy of the vulnerable invocation is caught
2344
+ as well as an omitted one); and one expecting `error` or `nothing-tested`,
2345
+ since both mean the target was never exercised and such a control would stay
2346
+ green with the detection engine entirely broken. Validation and execution
2347
+ share one `control_plan` so they cannot drift.
2348
+ - **`overall_detection_verdict`** collapses a run to one verdict, ordered by what
2349
+ an operator must not miss rather than by frequency: one `confirmed` among a
2350
+ hundred negatives is the finding. `error` is reported only when *nothing*
2351
+ reached the target, and a run that built no probes is `nothing-tested` —
2352
+ never `negative`, which would read as "not vulnerable".
2353
+
2354
+ ### Changed
2355
+
2356
+ - `CONTRIBUTING.md` asks for a bench case alongside new detection coverage, and
2357
+ for the README table to state the tier the case actually reached.
2358
+
2359
+ ### Notes
2360
+
2361
+ - The two shipped cases have **not yet been executed through the harness** — it
2362
+ was written where no Docker daemon was available. Their invocations come from
2363
+ a documented, reproduced guide, but the case files themselves are unvalidated;
2364
+ `tests/bench/README.md` says so and flags the one field that is a guess. No
2365
+ README claim was changed to assert benchmark results.
2366
+
2367
+ ## [2.24.0] — 2026-08-16
2368
+
2369
+ The computed value is no longer looked for in the response body alone. A sink
2370
+ whose output surfaces anywhere else in the response was reported `negative` — a
2371
+ false negative on a class RCEKit already claims to cover, which is worse than a
2372
+ missing class. The oracle, the random operands and the control differential are
2373
+ unchanged; only the set of places searched is wider.
2374
+
2375
+ ### Added
2376
+
2377
+ - **Whole-response evidence search.** Every confirmation now sweeps the response
2378
+ body, the application response headers, individual cookie values, the redirect
2379
+ target RCEKit actually landed on, the HTTP reason phrase, and each leaf of a
2380
+ parsed JSON body. Real sinks put command output in a debug header or a
2381
+ `Set-Cookie`, and API targets surface an evaluator's result inside a nested
2382
+ error envelope — `{"error": {"detail": "cannot render 2058898001"}}` — where a
2383
+ substring search of the serialised body misses a value the encoder escaped.
2384
+ - **The evidence line names the channel that carried the value**, e.g.
2385
+ `target computed 'RK…' in header X-Cmd-Out (random operands, absent from
2386
+ control)`, so the finding stays reproducible by hand. A body-carried
2387
+ confirmation reads exactly as it did before.
2388
+
2389
+ ### Changed
2390
+
2391
+ - **The control differential now covers every channel, not just the body.** A
2392
+ value present anywhere in the payload-free control is not attributable to
2393
+ execution, so it yields `inconclusive` wherever it turned up. This is stricter
2394
+ than comparing only the channel that matched, and it is what keeps a wider
2395
+ search from becoming a looser verdict.
2396
+ - **The `file` method's control check covers every channel too**, on the same
2397
+ reasoning: its token is random, so its presence in any control channel means
2398
+ it did not get there by being written and served.
2399
+
2400
+ ### Security
2401
+
2402
+ - **A deeply nested JSON response can no longer silence detection.** Channels are
2403
+ built inside the delivery `try`/`except`, so a `RecursionError` while parsing
2404
+ or walking the body escaped as a network failure: a response that arrived
2405
+ perfectly well was reported "request never reached the target". Measured:
2406
+ every one of the 46 probes in a default `reflected` run turned into `error`,
2407
+ which a target could induce deliberately to hide a live sink behind a thousand
2408
+ nested arrays. Version-independent, though the source moves — CPython 3.12
2409
+ raised the C recursion limit its JSON scanner runs under, so on 3.12/3.13 the
2410
+ parser survives a depth that breaks it on 3.8–3.11 and the recursive leaf walk
2411
+ hit the ordinary Python limit instead. The walk is now iterative and
2412
+ depth-capped, `RecursionError` from the parser costs the JSON channels only,
2413
+ and building channels can never turn a delivered response into a delivery
2414
+ failure.
2415
+ - **Transport headers are excluded from the sweep.** `Content-Length`, `Date`,
2416
+ `Age`, `ETag` and their neighbours are generated below the application and can
2417
+ never carry a computed value, but they *are* numeric — and the `expr` probe's
2418
+ expected value is a bare boundary-fenced number. Searching them would let a
2419
+ byte count collide with an arithmetic result and read as execution. Locked in
2420
+ by a test that puts the expected value in `Content-Length` and requires
2421
+ `negative`.
2422
+
2423
+ ## [2.23.3] — 2026-08-04
2424
+
2425
+ Four items from the same review: requests and seconds spent on work that could
2426
+ not produce a result. No verdict changes — the lab still confirms 15 of 15
2427
+ vulnerable sinks with nothing on the clean five — the run just stops paying for
2428
+ probes that were structurally unable to confirm.
2429
+
2430
+ ### Changed
2431
+
2432
+ - **The `awk` probe is no longer sent into a context that wraps the payload in
2433
+ quotes.** It carries double quotes, so in `attribute` the quote closed early
2434
+ and the rest was not a command: 5 requests per carrier that could only ever
2435
+ come back negative. Measured on a verbose shell sink, that shape confirmed 8
2436
+ times in `raw` and 0 times in `attribute`. Break-out contexts such as
2437
+ `shell_double_quoted` *close* the sink's quote and comment its tail, so they
2438
+ still get it. The same guard covers the PowerShell out-of-band shape.
2439
+ - **The timing screen runs in two waves.** Every delayed screen probe costs a
2440
+ real sleep, so screening all five separators up front spent `5 × base` seconds
2441
+ on every carrier, including the ones that cannot break out at all. `; ` and
2442
+ `| ` are screened first and the rest only if neither delayed — a sink that
2443
+ filters both is still swept, it is just no longer the price everyone pays.
2444
+ - **The out-of-band callback window is no longer paid per carrier.** Callbacks
2445
+ land in a burst once the channel works, so a target that has not produced one
2446
+ across every probe fired so far is not going to. The first carrier still gets
2447
+ the full window, so a target that does call back is never cut short before its
2448
+ first hit. On a clean target with the default carriers this was 30s of pure
2449
+ waiting; it is now ~12s.
2450
+ - **`--probe-depth` documents what it does on Windows**, which is nothing:
2451
+ `cmd.exe` has no `#` comment, no `${IFS}` and no `awk`, so both depths send
2452
+ the single `set /a` probe. The docs promised three extra shapes per sink
2453
+ without that caveat.
2454
+
2455
+ ## [2.23.2] — 2026-08-04
2456
+
2457
+ Three findings from a review of the detection work in 2.22.0 and 2.23.0. All
2458
+ three are the same shape: the run said something that was not true — about what
2459
+ it had done, about what it had looked for, or about which channel was live.
2460
+
2461
+ ### Fixed
2462
+
2463
+ - **`--methods oob` ignored `--verify-active-risk`.** Detection methods build
2464
+ their own probes and so bypass every corpus-level safety filter. That was
2465
+ harmless while every method was inert, but this one makes the target open
2466
+ outbound connections — and the same run printed *"low-impact (safe) payloads
2467
+ only; pass `--verify-active-risk intrusive` to also fire … OOB"* and then fired
2468
+ OOB anyway. It now needs `--verify-active-risk intrusive`, the same tier that
2469
+ holds back the corpus OOB payloads, and refuses before the listener binds.
2470
+ - **`--probe-depth quick` silently narrowed the timing separator screen to
2471
+ `; `.** That put back the exact blind spot the screen was added to remove, so
2472
+ a sink that merely filters `;` reported negative — and only for the operator
2473
+ who chose `quick` to be gentle on a rate-limited target. Both depths now screen
2474
+ every candidate separator; `--probe-depth` governs probe *shapes* only, and
2475
+ `--separators` remains the way to narrow break-outs deliberately.
2476
+ - **The DNS out-of-band probes could not call back on the default port, and
2477
+ nothing said so.** A DNS callback travels the real resolver hierarchy, so it
2478
+ only arrives if the listener *is* the authority for the OOB domain — port 53
2479
+ plus NS delegation. On `--listen-dns-port 5335` the DNS shapes were still sent,
2480
+ never fired, and the startup line reported `DNS :5335` with no caveat. Since
2481
+ most of the shapes are DNS ones — a resolver is often the only egress a
2482
+ hardened target has — the silence was expensive. RCEKit now says which channel
2483
+ is live.
2484
+ - The blind-sink advice added in 2.23.0 suggested an `oob` command without the
2485
+ risk flag, which the gate above would refuse. Naming a command the tool then
2486
+ declines to run is a small version of the same problem, so it now spells out
2487
+ `--verify-active-risk intrusive`.
2488
+
2489
+ ## [2.23.1] — 2026-08-03
2490
+
2491
+ ### Added
2492
+
2493
+ - **[Verify it yourself](docs/verify-it-yourself.md)** — reproduce the README's
2494
+ confirmations locally against dockerised [vulhub](https://github.com/vulhub/vulhub)
2495
+ targets. Webmin CVE-2019-15107 driven from a captured request (`reflected` →
2496
+ `confirmed`, then `time` → `needs-review` on the *same* sink, which is the
2497
+ clearest demonstration that the tiers are not merged), and Struts2 S2-001
2498
+ (`eval` confirms, `reflected` does not, on a target where both were tried).
2499
+
2500
+ Log4Shell is documented as an advanced case rather than a five-minute one: its
2501
+ sink is a JNDI lookup inside a logging library, so `--methods oob` does not
2502
+ apply — that method builds shell probes for shell-capable environments. The
2503
+ `${jndi:…}` payloads come from the `oob` *category* with the listener
2504
+ correlating the callback, and the token rides in a DNS label, which needs a
2505
+ delegated domain. Saying so is cheaper than a reader discovering it mid-demo.
2506
+
2507
+ ## [2.23.0] — 2026-08-03
2508
+
2509
+ The three sinks v2.22.0 still could not reach. One was a real gap in the probe
2510
+ set; the other two were a reporting problem, not a detection one. With both
2511
+ closed, a single `--methods reflected,eval,oob` run confirms **all fifteen**
2512
+ vulnerable sinks in the lab and still reports nothing on any of the five clean
2513
+ ones.
2514
+
2515
+ ### Added
2516
+
2517
+ - **A space-free probe, sent at both probe depths.** Stripping spaces is a filter
2518
+ of the same family as stripping `;` — it looks like it disarms command
2519
+ injection and does not, because `${IFS}` is a space as far as the shell is
2520
+ concerned. Every other probe carries a space, so that one filter silenced all
2521
+ of them and the sink was only reachable if the operator thought to pass
2522
+ `--evade low`. The separator's trailing space is trimmed with it (`;echo…`, not
2523
+ `; echo…`); the newline separator is unaffected. It costs one shape, so it is
2524
+ not part of the `--probe-depth` trade-off, and it is skipped under
2525
+ `--evade low`, which already applies the same transform everywhere.
2526
+ - **Guidance when every in-band probe comes back negative.** A results-based
2527
+ method cannot confirm a sink that returns no output — there is nowhere for the
2528
+ computed value to appear — so that negative is not evidence the target is
2529
+ clean. A run of `reflected`/`eval` alone that confirms nothing now says exactly
2530
+ that and names the methods that could still reach a blind sink, with the flags
2531
+ each one needs. It is suppressed once a blind-capable method has already run,
2532
+ and the `file` line is dropped once a web root is known.
2533
+
2534
+ ## [2.22.0] — 2026-08-03
2535
+
2536
+ Detection coverage. Measured against a lab of twenty sinks — fifteen genuinely
2537
+ vulnerable, five deliberately clean — the results-based methods went from
2538
+ confirming 8 of the 15 to confirming 12, with no new false positives on any of
2539
+ the clean ones.
2540
+
2541
+ ### ⚠️ A blind-timing candidate could be pure latency drift
2542
+
2543
+ `--methods time` fired its probes in a fixed ascending delay order
2544
+ (`0,0,N,N,2N,2N`), which makes the injected delay collinear with the request
2545
+ index. A target that simply gets **slower during the run** — progressive load, a
2546
+ rate limiter backing off, a filling log — therefore produced a textbook-perfect
2547
+ linear fit while being entirely un-injectable. In the lab this reproduced on 8
2548
+ of 8 runs against a sink with no command execution anywhere in it.
2549
+
2550
+ The probe order is now randomised, and the request index enters the regression
2551
+ as a nuisance term, so drift loads onto a drift coefficient instead of
2552
+ masquerading as a sleep. The same lab sink now reports negative on 9 of 9 runs,
2553
+ with every genuine timing detection preserved. If you have a `needs-review`
2554
+ timing candidate from an earlier version against a target that was under load,
2555
+ it is worth re-running.
2556
+
2557
+ ### Added
2558
+
2559
+ - **`--methods oob`** — out-of-band detection, the first `confirmed`-tier method
2560
+ for a sink that returns nothing *and* has no writable web root. Starts the
2561
+ built-in HTTP+DNS listener in-process and asks the target to resolve or fetch
2562
+ `<token>.<oob-host>`; a callback carrying a token the target could only have
2563
+ learned by running the command is proof of execution. Each probe gets its own
2564
+ token, so the finding names the break-out that actually worked. One shape puts
2565
+ a computed value in the DNS label, so the callback proves the shell evaluated
2566
+ arithmetic rather than merely resolving a name. Requires `--oob-host`, since
2567
+ it makes the target open outbound connections.
2568
+ - **`--probe-depth quick|full`** (default `full`) — trades requests for
2569
+ coverage. `full` adds three probe shapes, each aimed at a filter that silenced
2570
+ the canonical ones: substitution-free (`awk`, bare `expr`) for sinks that strip
2571
+ `$(` and backticks; keyword-diverse (`awk`) for filters on `echo`/`expr`; and
2572
+ comment-terminated (`… #`) for applications that append a redirect, extra
2573
+ arguments or a pipe after the injection point. `quick` keeps the old probe set
2574
+ at roughly half the requests.
2575
+
2576
+ ### Fixed
2577
+
2578
+ - **A `ping '<input>'` sink could not be detected at all.** The
2579
+ `shell_single_quoted`/`shell_double_quoted` contexts exist precisely for input
2580
+ interpolated inside quotes, but they are not in `default_contexts`, so no
2581
+ record carried them and the detection engine never tried them — the one sink
2582
+ shape they exist for was the one shape that always reported clean. They are now
2583
+ probed by default, and skipped when `--contexts` names a selection explicitly.
2584
+ - **`--methods time` reported a `;`-filtering sink as negative.** A regression
2585
+ blends its probes into one measurement, so it could not sweep separators the
2586
+ way the results-based methods do and was locked to `; ` alone — while
2587
+ `| sleep 3` delayed on the same sink. It now screens every candidate separator
2588
+ with one cheap probe each, then runs the regression through whichever one
2589
+ actually delayed.
2590
+ - **A trailing redirect or pipe in the sink hid a working probe.**
2591
+ `<cmd> <input> 2>/dev/null` and `<cmd> <input> | grep …` swallow the probe's
2592
+ output, so it executed and still read as negative. The comment-terminated
2593
+ shapes comment that tail out.
2594
+
2595
+ ## [2.21.1] — 2026-08-02
2596
+
2597
+ First release since v2.15.2. The headline is not a new feature — it is that
2598
+ detection is now correct in cases where it previously was not.
2599
+
2600
+ ### ⚠️ Re-check findings from v2.15.2 and earlier
2601
+
2602
+ **A reflection could be reported as `confirmed`.** The paired same-token control
2603
+ in `run_verification` was gated on a plain `re.search`, while the verdict itself
2604
+ used the encoding-aware search. A target that only echoes input but wraps its
2605
+ output — base64, hex, URL- or HTML-encoded — skipped the control entirely and was
2606
+ reported as proven execution: precisely the case the encoding-aware search was
2607
+ added for. If you ran an earlier version against a target that encodes its
2608
+ responses, a `confirmed` verdict from that run is worth re-testing.
2609
+
2610
+ ### Fixed — false negatives on exploitable targets
2611
+
2612
+ - **Separator sweep.** Shell probes always broke out with a single hardcoded
2613
+ `; `, so a sink that strips `;` — the most common partial mitigation there is,
2614
+ and one that stops nothing on its own — defeated every probe. Measured against
2615
+ nine deliberately vulnerable local sinks, detection was correct on 5 of 9;
2616
+ three of the four misses were exploitable targets reported clean. Probes now
2617
+ sweep `; `, `| `, `|| `, `&& ` and a newline, narrowable with `--separators`.
2618
+ - **Language runtimes.** An environment names what runs the application, not what
2619
+ executes the injected command: PHP's `system()`, Python's `os.system()`,
2620
+ Node's `child_process.exec()`, Ruby's `system()`, Perl's backticks and Go's
2621
+ `os/exec` all hand the string to `/bin/sh`. Scoping a run to the language the
2622
+ application is written in — the natural thing to do — used to send no shell
2623
+ probes at all.
2624
+ - **Whole-command sinks.** `--sink-raw` sends probes as bare commands for sinks
2625
+ that execute the input as the entire command (`qx/$input/`, `sh -c "$input"`),
2626
+ where a leading `;` is a syntax error that guaranteed a false negative.
2627
+ - **Captured requests.** A trailing newline in a saved request body is no longer
2628
+ sent as part of the body.
2629
+
2630
+ ### Fixed — a failed request is not a clean result
2631
+
2632
+ - A request that never reached the target is reported `error`, not `negative`.
2633
+ - Runs that build no probes at all exit non-zero and say so, instead of ending
2634
+ in silence and exit 0 — which read exactly like a target that came back clean.
2635
+ - The OOB DNS listener no longer dies on a malformed query, and write failures
2636
+ surface instead of being swallowed.
2637
+
2638
+ ### Fixed — safety and audit
2639
+
2640
+ - **Multi-step chains now carry the same safeguards as single requests.** The
2641
+ chain path delivered to live targets without sink-shape filters, destructive
2642
+ hold-back or a pre-flight plan, so `--verify-active-risk stateful` fired
2643
+ persistence and irreversible file operations that `--verify-url` refuses to
2644
+ send without `--verify-allow-destructive`. Both paths now share one hold-back
2645
+ and print the same plan.
2646
+ - **The audit trail redacts credentials** — it records that a credential header
2647
+ was sent, never its value.
2648
+ - A capture carrying `Authorization` or `Cookie` over plain `http` is flagged
2649
+ before anything is sent.
2650
+
2651
+ ### Added
2652
+
2653
+ - **The payload corpus is embedded in `rcekit.py`**, so the single file runs on
2654
+ its own — a jump box, an air-gapped host, a bare `curl` of the raw script.
2655
+ Resolution order is `--template-file` → `templates/payloads.json` beside the
2656
+ script → the built-in copy, and falling back to the built-in copy is
2657
+ announced. A corpus that exists but does not parse still hard-fails: that
2658
+ check exists for truncated and tampered corpora. `tools/embed_corpus.py`
2659
+ regenerates the embedded copy, and the test suite fails if the two drift.
2660
+ - **`--insecure`** skips TLS verification for internal targets with self-signed
2661
+ or mismatched certificates — opt-in and explicit, like `curl -k`.
2662
+ - **`--sink-raw`** for whole-command injection sinks, also readable from a
2663
+ target profile.
2664
+ - **`--separators`** to narrow the break-out sweep once the sink's shape is
2665
+ known.
2666
+ - **Documentation split into a task-oriented tree.** The README is half its
2667
+ former length and now leads with what RCEKit is for:
2668
+ [field guide](docs/guide.md) (worked examples by situation),
2669
+ [payload generation & exports](docs/generation.md), and
2670
+ [reference](docs/reference.md) (every flag grouped by task, plus the full
2671
+ taxonomies and exit codes).
2672
+ - **A "How RCEKit compares" section** covering commix, SSTImap, Nuclei and
2673
+ interactsh, with every claim traceable to that project's own documentation.
2674
+ - Four confirmation demos against real, publicly documented CVEs (Webmin
2675
+ CVE-2019-15107, Struts2 S2-001, Log4Shell CVE-2021-44228).
2676
+
2677
+ ### Changed
2678
+
2679
+ - **Expect more requests per run.** The separator sweep and the language-runtime
2680
+ fix both widen the probe set. Narrow with `--separators`, `--contexts` and
2681
+ `--environments` once the sink's shape is known.
2682
+ - **`--doctor` output.** Its first line now names the corpus in use
2683
+ (`corpus: …`) rather than a path (`template: …`), since the corpus is no
2684
+ longer necessarily a file, and `[ok] file loaded and parsed` is now
2685
+ `[ok] corpus loaded and parsed`.
2686
+
2687
+ No breaking changes to the CLI, output formats, or the template schema.
2688
+ Standard library only, Python 3.8–3.13.
2689
+
2690
+ ## Earlier releases
2691
+
2692
+ Release notes for these live on the
2693
+ [Releases page](https://github.com/kabiri-labs/rcekit/releases); they predate
2694
+ this file and have not been restated here.
2695
+
2696
+ - **[2.15.2]** — Multi-method RCE detection &amp; confirmation
2697
+ - **[2.7.0]**
2698
+ - **[2.1.0]**
2699
+
2700
+
2701
+
2702
+ [Unreleased]: https://github.com/kabiri-labs/rcekit/compare/v2.36.0...HEAD
2703
+ [3.0.0]: https://github.com/kabiri-labs/rcekit/compare/v2.45.7...v3.0.0
2704
+ [2.45.7]: https://github.com/kabiri-labs/rcekit/compare/v2.45.6...v2.45.7
2705
+ [2.45.6]: https://github.com/kabiri-labs/rcekit/compare/v2.45.5...v2.45.6
2706
+ [2.45.5]: https://github.com/kabiri-labs/rcekit/compare/v2.45.4...v2.45.5
2707
+ [2.45.4]: https://github.com/kabiri-labs/rcekit/compare/v2.45.3...v2.45.4
2708
+ [2.45.3]: https://github.com/kabiri-labs/rcekit/compare/v2.45.2...v2.45.3
2709
+ [2.45.2]: https://github.com/kabiri-labs/rcekit/compare/v2.45.1...v2.45.2
2710
+ [2.45.1]: https://github.com/kabiri-labs/rcekit/compare/v2.45.0...v2.45.1
2711
+ [2.45.0]: https://github.com/kabiri-labs/rcekit/compare/v2.44.0...v2.45.0
2712
+ [2.44.0]: https://github.com/kabiri-labs/rcekit/compare/v2.43.0...v2.44.0
2713
+ [2.43.0]: https://github.com/kabiri-labs/rcekit/compare/v2.42.0...v2.43.0
2714
+ [2.42.0]: https://github.com/kabiri-labs/rcekit/compare/v2.41.0...v2.42.0
2715
+ [2.41.0]: https://github.com/kabiri-labs/rcekit/compare/v2.40.0...v2.41.0
2716
+ [2.40.0]: https://github.com/kabiri-labs/rcekit/compare/v2.39.0...v2.40.0
2717
+ [2.39.0]: https://github.com/kabiri-labs/rcekit/compare/v2.38.0...v2.39.0
2718
+ [2.38.0]: https://github.com/kabiri-labs/rcekit/compare/v2.37.0...v2.38.0
2719
+ [2.37.0]: https://github.com/kabiri-labs/rcekit/compare/v2.36.0...v2.37.0
2720
+ [2.36.0]: https://github.com/kabiri-labs/rcekit/compare/v2.35.5...v2.36.0
2721
+ [2.35.5]: https://github.com/kabiri-labs/rcekit/compare/v2.35.4...v2.35.5
2722
+ [2.35.4]: https://github.com/kabiri-labs/rcekit/compare/v2.35.3...v2.35.4
2723
+ [2.35.3]: https://github.com/kabiri-labs/rcekit/compare/v2.35.2...v2.35.3
2724
+ [2.35.2]: https://github.com/kabiri-labs/rcekit/compare/v2.35.1...v2.35.2
2725
+ [2.35.1]: https://github.com/kabiri-labs/rcekit/compare/v2.35.0...v2.35.1
2726
+ [2.35.0]: https://github.com/kabiri-labs/rcekit/compare/v2.34.1...v2.35.0
2727
+ [2.34.1]: https://github.com/kabiri-labs/rcekit/compare/v2.34.0...v2.34.1
2728
+ [2.34.0]: https://github.com/kabiri-labs/rcekit/compare/v2.33.0...v2.34.0
2729
+ [2.33.0]: https://github.com/kabiri-labs/rcekit/compare/v2.32.0...v2.33.0
2730
+ [2.32.0]: https://github.com/kabiri-labs/rcekit/compare/v2.31.0...v2.32.0
2731
+ [2.31.0]: https://github.com/kabiri-labs/rcekit/compare/v2.30.0...v2.31.0
2732
+ [2.30.0]: https://github.com/kabiri-labs/rcekit/compare/v2.29.0...v2.30.0
2733
+ [2.29.0]: https://github.com/kabiri-labs/rcekit/compare/v2.28.0...v2.29.0
2734
+ [2.28.0]: https://github.com/kabiri-labs/rcekit/compare/v2.27.0...v2.28.0
2735
+ [2.27.0]: https://github.com/kabiri-labs/rcekit/compare/v2.26.0...v2.27.0
2736
+ [2.26.0]: https://github.com/kabiri-labs/rcekit/compare/v2.25.0...v2.26.0
2737
+ [2.25.0]: https://github.com/kabiri-labs/rcekit/compare/v2.24.0...v2.25.0
2738
+ [2.24.0]: https://github.com/kabiri-labs/rcekit/compare/v2.23.3...v2.24.0
2739
+ [2.23.3]: https://github.com/kabiri-labs/rcekit/compare/v2.23.2...v2.23.3
2740
+ [2.23.2]: https://github.com/kabiri-labs/rcekit/compare/v2.23.1...v2.23.2
2741
+ [2.23.1]: https://github.com/kabiri-labs/rcekit/compare/v2.23.0...v2.23.1
2742
+ [2.23.0]: https://github.com/kabiri-labs/rcekit/compare/v2.22.0...v2.23.0
2743
+ [2.22.0]: https://github.com/kabiri-labs/rcekit/compare/v2.21.1...v2.22.0
2744
+ [2.21.1]: https://github.com/kabiri-labs/rcekit/compare/v2.15.2...v2.21.1
2745
+ [2.15.2]: https://github.com/kabiri-labs/rcekit/releases/tag/v2.15.2
2746
+ [2.7.0]: https://github.com/kabiri-labs/rcekit/releases/tag/v2.7.0
2747
+ [2.1.0]: https://github.com/kabiri-labs/rcekit/releases/tag/v2.1.0