rcekit 2.35.1__tar.gz → 2.45.3__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- rcekit-2.45.3/CHANGELOG.md +2306 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/CONTRIBUTING.md +65 -8
- {rcekit-2.35.1/rcekit.egg-info → rcekit-2.45.3}/PKG-INFO +148 -26
- {rcekit-2.35.1 → rcekit-2.45.3}/README.md +147 -25
- {rcekit-2.35.1 → rcekit-2.45.3}/docs/generation.md +10 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/docs/guide.md +23 -6
- {rcekit-2.35.1 → rcekit-2.45.3}/docs/reference.md +397 -34
- {rcekit-2.35.1 → rcekit-2.45.3}/docs/verify-it-yourself.md +10 -0
- {rcekit-2.35.1 → rcekit-2.45.3/rcekit.egg-info}/PKG-INFO +148 -26
- {rcekit-2.35.1 → rcekit-2.45.3}/rcekit.egg-info/SOURCES.txt +1 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/rcekit.py +2399 -119
- {rcekit-2.35.1 → rcekit-2.45.3}/templates/payloads.json +33 -0
- rcekit-2.45.3/tests/bench/README.md +225 -0
- rcekit-2.45.3/tests/bench/cases/log4shell-cve-2021-44228.json +69 -0
- rcekit-2.45.3/tests/bench/cases/struts2-s2-001.json +69 -0
- rcekit-2.45.3/tests/bench/cases/webmin-cve-2019-15107.json +62 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/tests/bench/runner.py +261 -56
- {rcekit-2.35.1 → rcekit-2.45.3}/tests/test_bench_runner.py +314 -6
- rcekit-2.45.3/tests/test_docs.py +800 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/tests/test_generator.py +3931 -204
- rcekit-2.35.1/CHANGELOG.md +0 -1042
- rcekit-2.35.1/tests/bench/README.md +0 -138
- rcekit-2.35.1/tests/bench/cases/struts2-s2-001.json +0 -46
- rcekit-2.35.1/tests/bench/cases/webmin-cve-2019-15107.json +0 -43
- rcekit-2.35.1/tests/test_docs.py +0 -309
- {rcekit-2.35.1 → rcekit-2.45.3}/LICENSE +0 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/MANIFEST.in +0 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/SECURITY.md +0 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/pyproject.toml +0 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/rcekit.egg-info/dependency_links.txt +0 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/rcekit.egg-info/entry_points.txt +0 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/rcekit.egg-info/top_level.txt +0 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/setup.cfg +0 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/tests/bench/requests/webmin.txt +0 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/tests/test_corpus_embedding.py +0 -0
- {rcekit-2.35.1 → rcekit-2.45.3}/tools/embed_corpus.py +0 -0
|
@@ -0,0 +1,2306 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to RCEKit are recorded here. The format follows
|
|
4
|
+
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and RCEKit follows
|
|
5
|
+
[Semantic Versioning](https://semver.org/spec/v2.0.0.html): PATCH for fixes,
|
|
6
|
+
MINOR for new capabilities, MAJOR for breaking changes to the CLI, output
|
|
7
|
+
formats, or the template schema.
|
|
8
|
+
|
|
9
|
+
## [Unreleased]
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- **A tier correction has to reach the prose, not only the class.** Four tests
|
|
14
|
+
now read the tier from `DETECTION_METHODS[name].tier` and hold the sentences
|
|
15
|
+
an operator actually sees against it.
|
|
16
|
+
|
|
17
|
+
`lookup` moved from `confirmed` to `lookup-sink` on the class, and the move
|
|
18
|
+
was made in the README's CVE table -- but `blind_sink_advice` went on
|
|
19
|
+
offering `--methods lookup` as a method that "confirms", in the list where
|
|
20
|
+
`oob` and `file` do mean confirmed execution, and the README's Log4Shell demo
|
|
21
|
+
heading still said `confirmed` over alt text calling the run
|
|
22
|
+
"auto-confirming a blind Log4Shell RCE". Three places, one correction, and
|
|
23
|
+
nothing compared them: the existing test pinned the one line it was written
|
|
24
|
+
for (`time` is marked needs-review only) rather than asking every line the
|
|
25
|
+
same question.
|
|
26
|
+
|
|
27
|
+
So: every `blind_sink_advice` line naming a method must state that method's
|
|
28
|
+
tier and no other; every CVE row must state the tier its method reports;
|
|
29
|
+
every demo heading must match its row in that table; and a recording below
|
|
30
|
+
`confirmed` may not be described as confirming, alt text included, since
|
|
31
|
+
that is the sentence a screen reader reads out.
|
|
32
|
+
|
|
33
|
+
The wording checks match stems rather than words, and denials such as
|
|
34
|
+
"never confirmed" are removed before they are applied. "confirms" alone
|
|
35
|
+
would have readmitted the overclaim through "confirming RCE" or "confirmed
|
|
36
|
+
execution" -- an assertion answering the same way for the right reason and
|
|
37
|
+
the broken one, which is the defect being guarded against. Each test was
|
|
38
|
+
checked by reintroducing the defect it exists for, and each carries a floor
|
|
39
|
+
on how much it inspected: a parser that quietly stopped matching would
|
|
40
|
+
otherwise pass exactly as a clean README does.
|
|
41
|
+
|
|
42
|
+
Test-only: no version bump, and nothing about a run changes.
|
|
43
|
+
|
|
44
|
+
- **The same check, for the two tables under `docs/`.** They were left out when
|
|
45
|
+
the README's were pinned, and they carry the same claim:
|
|
46
|
+
`docs/reference.md` names the tier each method can reach, and
|
|
47
|
+
`docs/guide.md` tells an operator which method to reach for next.
|
|
48
|
+
|
|
49
|
+
`reference.md` states a ceiling, so its rule is a subset rather than an
|
|
50
|
+
equality -- a cell may also name a weaker tier the method really emits, as
|
|
51
|
+
`write` and `deser` both do -- and every registered method must have a row,
|
|
52
|
+
so a capability cannot land unlookupable. `guide.md` has no tier column and
|
|
53
|
+
so is held only to the negative: a row recommending a method that cannot
|
|
54
|
+
confirm may not describe confirmation.
|
|
55
|
+
|
|
56
|
+
Rows are read by column heading rather than by position. Reading the whole
|
|
57
|
+
row made the `guide.md` check skip the one row it was written for, because
|
|
58
|
+
that row's prose names `oob` -- which confirms -- while recommending
|
|
59
|
+
`lookup`, which does not.
|
|
60
|
+
|
|
61
|
+
A reference row naming something that is not a registered method fails rather
|
|
62
|
+
than being passed over. Skipping it left the completeness check one-way: a
|
|
63
|
+
method *removed* from `DETECTION_METHODS` would leave its row behind, every
|
|
64
|
+
remaining row would still match, and the page would go on offering a
|
|
65
|
+
`--methods` value the CLI rejects.
|
|
66
|
+
|
|
67
|
+
Rows split on *unescaped* pipes. Markdown writes a literal pipe in a cell as
|
|
68
|
+
`\|`, which `docs/reference.md` already does in three tables, and splitting
|
|
69
|
+
on every pipe invents a cell: zipping against the header then drops the last
|
|
70
|
+
column outright, so a claim there stops being examined and every one of these
|
|
71
|
+
checks passes without looking at it.
|
|
72
|
+
|
|
73
|
+
A disclaimer is not read as a claim. "never confirmed" was already removed
|
|
74
|
+
before the `confirm` stem was looked for; "without confirmation" and
|
|
75
|
+
"unconfirmed" were not, and the second carries its negation inside the word
|
|
76
|
+
where a rule about preceding words cannot see it. Documentation saying the
|
|
77
|
+
honest thing would have failed the suite. One parser and one denial rule now,
|
|
78
|
+
where there were two of each here and a third in the generator tests.
|
|
79
|
+
|
|
80
|
+
### Changed
|
|
81
|
+
|
|
82
|
+
- **A bench case may bring its target up once for both halves**, with
|
|
83
|
+
`"share_target": true`. Bringing the container up twice is the largest fixed
|
|
84
|
+
cost in a case and both halves usually hit the same one, so on a fast case it
|
|
85
|
+
is most of the wall clock.
|
|
86
|
+
|
|
87
|
+
It is opt-in, and the default is unchanged, because the teardown between the
|
|
88
|
+
halves is `down -v`: today's control meets a **fresh** target. A case whose
|
|
89
|
+
vulnerable half writes a file, plants a shell or changes a setting would hand
|
|
90
|
+
its control a target it had already altered, and a control measured against a
|
|
91
|
+
contaminated target measures nothing -- which is the one failure a benchmark
|
|
92
|
+
may not have. Validation rejects the key on a case whose control brings up a
|
|
93
|
+
different target, since there is then nothing to share and leaving it set
|
|
94
|
+
would read as though there were.
|
|
95
|
+
|
|
96
|
+
A shared `up` that fails falls back to per-half management rather than
|
|
97
|
+
carrying on, so the run reports `compose up failed` instead of two readiness
|
|
98
|
+
timeouts naming the wrong cause.
|
|
99
|
+
|
|
100
|
+
Both shipped cases set it, because neither half writes anything. Measured on
|
|
101
|
+
`struts2-s2-001` against vulhub on Docker: 33.8s to 22.9s, with both halves
|
|
102
|
+
reaching the same verdicts either way.
|
|
103
|
+
|
|
104
|
+
- **The benchmark reaches a callback method**, with a case for Log4Shell
|
|
105
|
+
(Apache Solr 8.11.0, Log4j 2.14.1). `python tests/bench/runner.py --all` is
|
|
106
|
+
3/3.
|
|
107
|
+
|
|
108
|
+
Every case until now was in-band, so the listener, the token correlation and
|
|
109
|
+
`confirm_each` had never run against real software -- the area that produced
|
|
110
|
+
three P1 findings while `lookup` was being written.
|
|
111
|
+
|
|
112
|
+
Two things had to give. A JNDI lookup resolves through the system resolver,
|
|
113
|
+
which asks UDP 53, and a developer machine rarely has that port free; a bare
|
|
114
|
+
`--oob-host` IP is no way around it, because a lookup has no second channel
|
|
115
|
+
to carry the token. So a case may now name `run_in`, and the run happens in a
|
|
116
|
+
container on the target's own network at a fixed address, with an override
|
|
117
|
+
pointing the service's `dns:` at it. The repository is mounted read-only and
|
|
118
|
+
the image is a stock Python.
|
|
119
|
+
|
|
120
|
+
The other was the harness's own vocabulary. `VALID_EXPECTATIONS` and
|
|
121
|
+
`CONTROL_EXPECTATIONS` were written out by hand and had drifted: neither
|
|
122
|
+
`lookup-sink` nor `deserialization-sink` was in either, so a case for `lookup`
|
|
123
|
+
or `deser` could not be *loaded*, let alone run. Both are read from
|
|
124
|
+
`DETECTION_METHODS` now.
|
|
125
|
+
|
|
126
|
+
The control is the case: `oob` against the same Solr comes back `negative`
|
|
127
|
+
**although it is exploitable**, because every probe it builds is a shell
|
|
128
|
+
command and a `${jndi:...}` sink runs none of them. That is the gap `lookup`
|
|
129
|
+
was added to close, and until this case ran it rested on a fixture.
|
|
130
|
+
|
|
131
|
+
`{bench}` / `{repo}` now expand in a case's compose argv as well as its
|
|
132
|
+
invocation, and resolve to the mount point when the run is containerised.
|
|
133
|
+
|
|
134
|
+
The results file is created before the container starts, and it rather than
|
|
135
|
+
its directory is made writable. Under Docker's user-namespace remapping,
|
|
136
|
+
container root is a subordinate host UID, so a `mkdtemp` owned by the runner
|
|
137
|
+
at 0700 is not writable from inside: the file would never appear and the case
|
|
138
|
+
would report `nothing-tested`, as though detection had found nothing rather
|
|
139
|
+
than as though the channel had been shut. Granting the one file and not the
|
|
140
|
+
directory keeps anyone else from creating, replacing or unlinking entries
|
|
141
|
+
there.
|
|
142
|
+
|
|
143
|
+
Test-only: no version bump, and nothing about a run changes.
|
|
144
|
+
|
|
145
|
+
### Fixed
|
|
146
|
+
|
|
147
|
+
- **`--methods` enumerated only some of the methods it accepts.** `write`,
|
|
148
|
+
`lookup` and `deser` were registered in `DETECTION_METHODS` and had never
|
|
149
|
+
once been named in the help, so `--help` described a whole target class as
|
|
150
|
+
out of reach -- an upload that stores a file, a `${jndi:...}` sink, an
|
|
151
|
+
endpoint that deserializes what it is handed -- while the method for it was
|
|
152
|
+
already shipping and documented everywhere else. Each now carries the tier
|
|
153
|
+
its class declares: `confirmed` for `write`, with `needs-review` for a write
|
|
154
|
+
that is served but not interpreted; `lookup-sink` for `lookup`;
|
|
155
|
+
`deserialization-sink` for `deser`; and neither of the last two ever
|
|
156
|
+
`confirmed`.
|
|
157
|
+
|
|
158
|
+
`oob` also called itself "the only confirmed-tier method for a fully blind
|
|
159
|
+
sink". The tier is right, but an exclusivity claim is the kind that goes
|
|
160
|
+
stale without anything failing, so the line now says what the class itself
|
|
161
|
+
says: it confirms a sink with no output channel and no writable web root.
|
|
162
|
+
|
|
163
|
+
A new test in `CLIDocumentationTestCase` holds the list to the registry, the
|
|
164
|
+
way `--eval-engines` is already held to the corpus. It reads the `--methods`
|
|
165
|
+
help alone rather than the whole page, because every one of these names also
|
|
166
|
+
occurs inside some other flag -- `file` in `--request-file`, `write` in
|
|
167
|
+
`--file-write-path`, `time` in `--time-base` -- and it looks for the name
|
|
168
|
+
followed by the parenthesis that opens its description, because the old help
|
|
169
|
+
carried the word `write` inside `file`'s "write+read-back" and a looser
|
|
170
|
+
match would have counted that as documentation.
|
|
171
|
+
|
|
172
|
+
A second pass corrected what `--oob-host` promises those two methods. Its
|
|
173
|
+
help offered "an IP it can reach or a domain delegated to this listener" and
|
|
174
|
+
was scoped `(--methods oob)`, which was true when `oob` was the only caller.
|
|
175
|
+
It is not true now: `LookupCallback.build_probes` returns no probes at all
|
|
176
|
+
for an address literal, so `lookup` reports `nothing-tested`, and `DeserSink`
|
|
177
|
+
builds `<token>.<IP>`, which resolves nowhere, so its DNS gadget can never
|
|
178
|
+
reach `deserialization-sink`. An operator who satisfied the stated
|
|
179
|
+
prerequisite got silence from one method and a capped tier from the other.
|
|
180
|
+
The flag's help, the `--methods` entries for `lookup` and `deser`, the
|
|
181
|
+
message printed when a callback method is selected without a host, and the
|
|
182
|
+
`deser` oracle table in `docs/reference.md` now all say the same thing: an IP
|
|
183
|
+
serves `oob` alone, because only `oob` can put its token in a URL path.
|
|
184
|
+
|
|
185
|
+
`boolean` landed on `main` while this branch was open. It was named in the
|
|
186
|
+
help but left out of the `--methods` row in `docs/reference.md` -- the same
|
|
187
|
+
drift, one table over. Both carry it now, and the new test covered it without
|
|
188
|
+
being told to, which is the point of reading the registry instead of a list.
|
|
189
|
+
|
|
190
|
+
No version bump -- this is documentation and tests only.
|
|
191
|
+
|
|
192
|
+
## [Unreleased]
|
|
193
|
+
|
|
194
|
+
## [2.45.3] — 2026-09-24
|
|
195
|
+
|
|
196
|
+
### Fixed
|
|
197
|
+
|
|
198
|
+
- **A `deser`-only run started no listener, so its gadgets went out with
|
|
199
|
+
nothing to receive them.** `--methods deser --oob-host <delegated-name>`,
|
|
200
|
+
with neither `oob` nor `lookup` selected, built its two DNS-gadget probes and
|
|
201
|
+
sent them to a target while no listener was running -- with the host
|
|
202
|
+
configured perfectly. `deser` reached `deserialization-sink` only when it
|
|
203
|
+
happened to be selected beside a method that required a host; run on its own,
|
|
204
|
+
the way the docs present it, its strongest oracle was structurally dead and
|
|
205
|
+
it capped at `needs-review` without saying so.
|
|
206
|
+
|
|
207
|
+
The gate read `needs_oob_host`, which was answering two questions at once:
|
|
208
|
+
"should the run stop when no host is named" and "does this run need a
|
|
209
|
+
listener". Every other method answers both the same way. `deser` is where
|
|
210
|
+
they come apart -- its shape oracle proves something with no listener at all,
|
|
211
|
+
so a missing host must not stop the run, while its DNS gadget does call back
|
|
212
|
+
and needs one.
|
|
213
|
+
|
|
214
|
+
So the second question gets its own attribute, `uses_oob_host`, and the
|
|
215
|
+
listener starts for any selected method that calls back once a host is named.
|
|
216
|
+
It is declared rather than derived: `needs_oob_host or oob_needs_dns_label`
|
|
217
|
+
gives the right answer for every method shipping today and would miss one
|
|
218
|
+
that called back over HTTP without requiring a host -- which is the shape
|
|
219
|
+
`oob` itself would have if its host were optional.
|
|
220
|
+
|
|
221
|
+
Two notices could now reach a run with no HTTP shape in it, and said "only
|
|
222
|
+
the HTTP shapes are live" as though `oob`'s channel were the only one. They
|
|
223
|
+
say "only the shapes that need no DNS are live", which is true of `oob`'s
|
|
224
|
+
HTTP probes and of `deser`'s shape oracle alike. The message for an
|
|
225
|
+
unbindable HTTP port no longer names `lookup` as the reason a run continues.
|
|
226
|
+
|
|
227
|
+
`uses_oob_host` is the default answer rather than the answer. Whether a
|
|
228
|
+
method calls back can depend on its own flags: `--deser-formats` selects the
|
|
229
|
+
ecosystems, and three of the five -- `php`, `dotnet`, `python_pickle` --
|
|
230
|
+
ship no DNS gadget at all. So both callers ask `builds_callback_probes()` on
|
|
231
|
+
a configured instance instead of reading the class. `--deser-formats php`
|
|
232
|
+
with an address no longer draws a notice blaming the address, because there
|
|
233
|
+
was no callback probe to lose and a delegated domain could not have helped;
|
|
234
|
+
and it no longer starts a listener that nothing would reach.
|
|
235
|
+
|
|
236
|
+
Whether the host's *form* is one a method can use is a second question,
|
|
237
|
+
`can_use_oob_host()`, and only the gate asks it. A listener bound for
|
|
238
|
+
`--methods deser --oob-host 10.0.0.9` receives nothing -- `build_probes`
|
|
239
|
+
drops an address and sends no gadget -- and the run announced "the target
|
|
240
|
+
will open outbound connections" directly beneath the notice saying those
|
|
241
|
+
probes are not sent. It stays out of `builds_callback_probes()`, because the
|
|
242
|
+
stranded notice is *about* the address and an address test there would
|
|
243
|
+
silence it for the one case it exists to report.
|
|
244
|
+
|
|
245
|
+
`lookup` with an address stops getting a listener too. It built no probes
|
|
246
|
+
with one either way, so nothing is lost, and the stranded notice says more
|
|
247
|
+
about that run than the port-53 warnings it no longer reaches.
|
|
248
|
+
|
|
249
|
+
A flaky test met on the way, fixed here because it would have bitten CI
|
|
250
|
+
sooner or later: the fake parsing endpoint behind the `deser` shape
|
|
251
|
+
differential sniffed one byte (`raw[4:5] not in (b"t", b"s")`) to decide
|
|
252
|
+
whether a stream was well-formed. The `noise` form is the format's magic
|
|
253
|
+
followed by random characters, so it passed that check about once in 500
|
|
254
|
+
runs -- measured at 398 in 200,000 -- all three forms answered alike, and
|
|
255
|
+
the oracle's correct `negative` failed a test that wanted `needs-review`.
|
|
256
|
+
The endpoint now reads TC_STRING the way a parser does, a two-byte length
|
|
257
|
+
followed by exactly that many bytes: 0 collisions in 500,000 seeds. Seeding
|
|
258
|
+
the test's RNG would have hidden it instead, and would have gone on hiding a
|
|
259
|
+
real regression in how the noise form is built.
|
|
260
|
+
|
|
261
|
+
This is the second half of #94, found while fixing the first and reported as
|
|
262
|
+
#98 rather than folded into #97: same defect -- probes sent that no callback
|
|
263
|
+
could follow -- reached by a different route. The format-dependent part came
|
|
264
|
+
from an automated review of #97 and the host-form part from one of this
|
|
265
|
+
change, both verified before being accepted.
|
|
266
|
+
|
|
267
|
+
## [2.45.2] — 2026-09-24
|
|
268
|
+
|
|
269
|
+
### Fixed
|
|
270
|
+
|
|
271
|
+
- **`deser` fired DNS gadgets at a host that could not resolve, and the warning
|
|
272
|
+
built for exactly that stayed quiet.** `--oob-host` takes a delegated name or
|
|
273
|
+
an address. `oob` uses either: given an address it drops its DNS shapes and
|
|
274
|
+
carries the token in a URL path instead. `lookup` refuses an address
|
|
275
|
+
outright, because a lookup expression has nowhere else to put one.
|
|
276
|
+
`DeserSink` had neither behaviour -- it built `<token>.10.0.0.9` and sent it,
|
|
277
|
+
a real request the target answered from which no callback could follow by
|
|
278
|
+
construction, and then capped at `needs-review` without saying why.
|
|
279
|
+
|
|
280
|
+
`oob_channel_warnings` is the function that exists to say a DNS channel is
|
|
281
|
+
structurally dead. It returned early on an address literal, reasoning that
|
|
282
|
+
such a run "never builds DNS probes and has nothing to warn about" -- true of
|
|
283
|
+
`oob`, and written when `oob` was the only method taking the flag. So the one
|
|
284
|
+
mechanism that could have reported this was the one place it was ruled out.
|
|
285
|
+
|
|
286
|
+
Both halves are fixed together, because either alone still ships the problem.
|
|
287
|
+
`deser` treats an address as no host and sends its shape oracle alone, which
|
|
288
|
+
is what it can prove without a listener anyway. The notice moves into
|
|
289
|
+
`oob_address_strands`, which names the methods an address stranded and the
|
|
290
|
+
tier that puts out of reach, read off each class.
|
|
291
|
+
|
|
292
|
+
It is a separate function rather than another branch of
|
|
293
|
+
`oob_channel_warnings` because the two answer different questions at
|
|
294
|
+
different times. "Can the DNS shapes reach this listener" is only worth
|
|
295
|
+
asking once a listener exists, and that block runs only for a method that
|
|
296
|
+
*requires* a callback host. `deser` does not require one, so it never
|
|
297
|
+
reaches that block -- a notice left there could not have fired for the
|
|
298
|
+
method that prompted it. The new check runs on every verification run that
|
|
299
|
+
names a host.
|
|
300
|
+
|
|
301
|
+
Which methods those are is a class attribute now, `oob_needs_dns_label`,
|
|
302
|
+
declared beside `needs_oob_host`. Deciding it at the one place that asked is
|
|
303
|
+
how the two drifted apart to begin with.
|
|
304
|
+
|
|
305
|
+
The test that pinned the old silence asserted "no DNS shape is ever built" --
|
|
306
|
+
a claim about `oob` wearing the shape of a claim about the flag. Its scope
|
|
307
|
+
narrows to `oob`, and the stranded case gets tests of its own.
|
|
308
|
+
|
|
309
|
+
Found while verifying an automated review finding on #93; reported as #94 and
|
|
310
|
+
deliberately kept out of that documentation-only PR.
|
|
311
|
+
|
|
312
|
+
## [2.45.1] — 2026-09-24
|
|
313
|
+
|
|
314
|
+
### Fixed
|
|
315
|
+
|
|
316
|
+
- **`--max-payloads` is a bound on requests, and it was being checked after
|
|
317
|
+
they had been sent.** The cap was measured against the number of result
|
|
318
|
+
*rows*. For a method that answers from each probe those are the same number,
|
|
319
|
+
so nothing showed. An aggregate method reports one row however many probes it
|
|
320
|
+
cost, so the whole series went out and the cap noticed afterwards:
|
|
321
|
+
|
|
322
|
+
| `--max-payloads 1` | the cost line said | requests sent |
|
|
323
|
+
|---|---|---|
|
|
324
|
+
| `reflected`, `eval` | 1 | 1 |
|
|
325
|
+
| `time` | 1 | **12** |
|
|
326
|
+
| `deser` | 1 | **15** |
|
|
327
|
+
| `boolean` | 1 | **27** |
|
|
328
|
+
|
|
329
|
+
A clean split: every aggregate method overran and no per-probe method did.
|
|
330
|
+
That line is the only thing an operator bounding a monitored engagement has
|
|
331
|
+
to go on before the traffic starts, and it was wrong in the direction that
|
|
332
|
+
matters — under, not over.
|
|
333
|
+
|
|
334
|
+
Rows were the wrong meter in a second way, which the first fix exposed. A
|
|
335
|
+
series the budget declines costs requests and produces no row at all, so the
|
|
336
|
+
next carrier recomputed the same allowance and fired again — `time` at
|
|
337
|
+
`--max-payloads 5` sent 4 requests per carrier with the cap never moving. And
|
|
338
|
+
in a mixed run, a per-probe method never saw what an aggregate one had
|
|
339
|
+
already spent: `boolean,eval` at 30 sent 56. The budget is counted in
|
|
340
|
+
requests for the whole call now. Swept across every method this build can run
|
|
341
|
+
without a callback host, alone and in combination, at seven caps: 63
|
|
342
|
+
combinations, worst overrun 0.
|
|
343
|
+
|
|
344
|
+
- **A measurement that cannot be finished is no longer started.** What a budget
|
|
345
|
+
may do to a series depends on where the method's answer lives, and the class
|
|
346
|
+
already knows: `DetectionMethod.decides_per_probe()` reads it from whether
|
|
347
|
+
the class overrides `confirm_each`, rather than from a list of names that
|
|
348
|
+
would go stale like every other one in this repository.
|
|
349
|
+
|
|
350
|
+
Where each probe answers — `deser`, `lookup`, `oob` — running out of budget
|
|
351
|
+
stops the series and the probes already sent keep their verdicts. Where only
|
|
352
|
+
the series answers — `time`, `boolean` — a part of one is not a weaker answer
|
|
353
|
+
but a wrong one: `time` reports `negative` from a screen with no regression
|
|
354
|
+
behind it and from a regression short of four samples, and a `boolean` series
|
|
355
|
+
whose anchors never went out reads the same way. Both are false cleans, and
|
|
356
|
+
the budget would have been manufacturing them. So a wave that does not fit
|
|
357
|
+
abandons the measurement rather than cutting it short.
|
|
358
|
+
|
|
359
|
+
- **An abandoned measurement leaves an `inconclusive` row, not a silence.**
|
|
360
|
+
Dropping it quietly let the *other* carriers describe the run, and the other
|
|
361
|
+
carriers are the ones with nothing to find: measured at `--max-payloads 12`
|
|
362
|
+
against a sink that honours an injected sleep, the unix carrier's regression
|
|
363
|
+
was abandoned for budget and a windows carrier's honest "no separator
|
|
364
|
+
delayed" was the only row left — so the run reported `negative` for a target
|
|
365
|
+
that was vulnerable. `inconclusive` is what an abandoned measurement is in
|
|
366
|
+
the word the tool already uses, and it outranks `negative` in the run
|
|
367
|
+
verdict. Swept across 31 caps against a sink that really does delay: no cap
|
|
368
|
+
reports it clean, and none overruns.
|
|
369
|
+
|
|
370
|
+
- **What the budget declined is reported by name**, in a sixth run-wide tally
|
|
371
|
+
beside the profile drops, the safety holds, the reach notes, the settled
|
|
372
|
+
carriers and the refusals. A ladder that shrinks quietly is indistinguishable
|
|
373
|
+
from a target with nothing to find:
|
|
374
|
+
|
|
375
|
+
```
|
|
376
|
+
[detect] --max-payloads held back 1 measurement(s) that could not have reached a
|
|
377
|
+
verdict within the budget:
|
|
378
|
+
[detect] 1 x boolean/raw needs 7 requests to reach a verdict and --max-payloads left 3
|
|
379
|
+
[!] No probes were built, so NOTHING WAS TESTED - this is not a negative result.
|
|
380
|
+
```
|
|
381
|
+
|
|
382
|
+
- **The cost line and the run make the same decision.** `estimate_detection_probes`
|
|
383
|
+
applies the same all-or-nothing rule, so under a cap the figure printed before
|
|
384
|
+
the traffic is the traffic. Uncapped it stays a floor by documented design — a
|
|
385
|
+
wave a method picks after seeing its own timings cannot be predicted from
|
|
386
|
+
there.
|
|
387
|
+
|
|
388
|
+
- **`sent N probes` is a claim about what the target received**, and it was
|
|
389
|
+
counting result rows. Those are the same number for a method that answers
|
|
390
|
+
from each probe, which is why it read true for so long -- but an aggregate
|
|
391
|
+
method reports one row for a whole series, so a `time` run that put 20
|
|
392
|
+
requests on a target announced 5, and a measurement this release declines
|
|
393
|
+
would have announced one probe for traffic that never left. Both the run
|
|
394
|
+
summary and the per-point line report deliveries now, with the result count
|
|
395
|
+
beside them.
|
|
396
|
+
|
|
397
|
+
One figure in the 2.45.0 entry below came from that line and was mislabelled
|
|
398
|
+
with it. Re-measured against the same sink: `--methods reflected,eval,time`
|
|
399
|
+
put **2883** requests on the target, not 2426, which was the number of
|
|
400
|
+
results. Every one of them was still `negative`, so the claim stands; the
|
|
401
|
+
number did not.
|
|
402
|
+
|
|
403
|
+
- **An evasion retry is a request, and the meter could not see it.** `_escalate`
|
|
404
|
+
fired without incrementing `delivered_probes`, so retries were traffic
|
|
405
|
+
outside the bound this release makes real: measured at `--max-payloads 5
|
|
406
|
+
--evade high` against a filter that refuses whitespace, the target received
|
|
407
|
+
12 requests and the run recorded 5. Counting them was half the fix. The other
|
|
408
|
+
half is that the budget is checked *per rung* rather than once before the
|
|
409
|
+
ladder -- one refused probe is retried at `low` and again at `high`, so it
|
|
410
|
+
can cost three requests, and a check at the call site alone still overshot by
|
|
411
|
+
one at `--max-payloads 10 --evade high`. Swept over caps 1 to 25 at all three
|
|
412
|
+
evasion settings: 75 combinations, no overrun, and the counter equal to the
|
|
413
|
+
requests the target actually received in every one.
|
|
414
|
+
|
|
415
|
+
### Security
|
|
416
|
+
|
|
417
|
+
No authorization, deny-by-default or matrix-evaluation boundary moves, and
|
|
418
|
+
nothing widens `confirmed`. The effect is one-way: fewer requests than before,
|
|
419
|
+
and a run that could not afford to test something now says so instead of
|
|
420
|
+
reporting it clean.
|
|
421
|
+
|
|
422
|
+
## [2.45.0] — 2026-09-23
|
|
423
|
+
|
|
424
|
+
### Added
|
|
425
|
+
|
|
426
|
+
- **`--methods boolean` — a sink that evaluates a predicate and renders nothing
|
|
427
|
+
of it.** MongoDB `$where` is the shape, and this repository has carried the
|
|
428
|
+
note for two releases: a JS sandbox with no shell, no egress and no value in
|
|
429
|
+
the response, only a document set that a predicate narrows. Every shipped
|
|
430
|
+
oracle is structurally blind to it. Measured against exactly that sink:
|
|
431
|
+
|
|
432
|
+
```
|
|
433
|
+
[detect] methods: reflected, eval, time
|
|
434
|
+
[detect] sent 2426 probes: negative=2426
|
|
435
|
+
```
|
|
436
|
+
|
|
437
|
+
2426 requests and a clean negative on a target that evaluates whatever it is
|
|
438
|
+
handed. The new method reads the one channel left — whether the *shape* of the
|
|
439
|
+
response changed between a true predicate and a false one.
|
|
440
|
+
|
|
441
|
+
**It is `needs-review` and there is no path from here to `confirmed`.** Not
|
|
442
|
+
because the signal is weak, but because of what it cannot distinguish: against
|
|
443
|
+
a sandboxed `eval` sink and against a plain SQLite comparison it produced an
|
|
444
|
+
identical clean differential in 40 runs each, and a query engine comparing two
|
|
445
|
+
numbers is not remote code execution. Extracting a locally computed product
|
|
446
|
+
bit by bit through the channel was tried and does not fix it — it recovers the
|
|
447
|
+
product through both sinks alike, for about 80 requests and a string function
|
|
448
|
+
a sandbox may well deny.
|
|
449
|
+
|
|
450
|
+
**The naive form of this oracle is unusable**, which is why none of it is.
|
|
451
|
+
`1==1` against `1==2`, with a changed response read as a finding, called a
|
|
452
|
+
target that only *reflected* its input vulnerable in 40 runs out of 40, and
|
|
453
|
+
one whose response merely wobbled in 32 of 40. Four guards, each a measured
|
|
454
|
+
false-finding rate rather than a precaution:
|
|
455
|
+
|
|
456
|
+
- **Compare structure, not the body and not its length.** A reflected payload
|
|
457
|
+
lands in the text between two tags, and the text between two tags is what
|
|
458
|
+
the signature throws away. A length-based signature claimed a differential
|
|
459
|
+
in 13 of 25 runs against a reflect-only target; comparing raw bodies was
|
|
460
|
+
unusable outright, reading `unstable` in 25 of 25 runs against a target that
|
|
461
|
+
*was* vulnerable, because one CSRF token makes every response unique.
|
|
462
|
+
- **Several independently randomised pairs, not one.** Against a target whose
|
|
463
|
+
response varies on its own, one pair claimed a differential in 46 of 200
|
|
464
|
+
runs; two claimed none in 200. `--probe-depth quick` trades four pairs for
|
|
465
|
+
two and never for one.
|
|
466
|
+
- **Randomised firing order.** A target that never reads the payload but
|
|
467
|
+
degrades part-way through a run splits an ordered true-then-false series
|
|
468
|
+
perfectly: at the worst point of a swept degradation, 100 false findings out
|
|
469
|
+
of 100.
|
|
470
|
+
- **An anchor before and after the series, each a different true predicate.**
|
|
471
|
+
Shuffling alone still left 2 in 100, which is just the chance a shuffle
|
|
472
|
+
lands separable. Re-measuring the channel afterwards caught it 100 times in
|
|
473
|
+
100, because a target that moved during the series cannot answer the
|
|
474
|
+
closing anchor the way it answered the opening one. Sending *one* anchor
|
|
475
|
+
payload three times does not merely weaken that: a cache keyed on the query
|
|
476
|
+
string answers the repeats from its store, so the closing anchor agrees
|
|
477
|
+
with the opening one whatever the target did in between. Measured against
|
|
478
|
+
an input-blind target that degrades mid-series, identical anchors caught it
|
|
479
|
+
in 36 of 39 runs live and in 0 of 39 behind a cache. Every probe payload is
|
|
480
|
+
unique, so nothing else in the series is replayable.
|
|
481
|
+
|
|
482
|
+
With every guard on, a genuinely evaluating target still read as a
|
|
483
|
+
differential in 100 runs of 100.
|
|
484
|
+
|
|
485
|
+
- **A channel that cannot carry one bit is `inconclusive`, never `negative`.**
|
|
486
|
+
If the same probe draws two different shapes, or the shape moves while the
|
|
487
|
+
series is being fired, the run says so. `negative` asserts the probes reached
|
|
488
|
+
the target and found nothing; here they reached it and no answer could be read
|
|
489
|
+
out of them, which is the same false clean `blocked` and `nothing-tested`
|
|
490
|
+
exist to prevent, one oracle further in. No new verdict: there are still nine.
|
|
491
|
+
|
|
492
|
+
- **`response_shape()`** — a response reduced to its structure, with everything
|
|
493
|
+
it said removed. Three readings, because a response is one of three things and
|
|
494
|
+
the wrong reading is not a near miss: a JSON document keeps its keys, nesting
|
|
495
|
+
and list lengths and drops every scalar; markup keeps its tag skeleton;
|
|
496
|
+
anything else keeps one marker per word per line. Measured on a JSON sink, the
|
|
497
|
+
markup reading was unusable — `unstable` in 25 of 25 runs — and the shape tree
|
|
498
|
+
read the differential in 25 of 25.
|
|
499
|
+
|
|
500
|
+
### Changed
|
|
501
|
+
|
|
502
|
+
- **`CODE_POSITION_CONTEXTS`** names the contexts that carry the injected value
|
|
503
|
+
as code rather than as a value, and `boolean` is offered every other one. The
|
|
504
|
+
first cut of this asked whether the context had a break-out prefix at all, and
|
|
505
|
+
got both halves wrong: it refused `attribute`, `attribute_unquoted`,
|
|
506
|
+
`xml_cdata` and `yaml`, whose delimiters open and close *around* the value and
|
|
507
|
+
leave a predicate exactly where a predicate belongs, and it offered
|
|
508
|
+
`unix_shell`, `windows_cmd` and `powershell`, which have no delimiters at all
|
|
509
|
+
and run the value as a command. The suite enumerates both sides, so a context
|
|
510
|
+
added to the corpus fails until somebody decides which one it is on.
|
|
511
|
+
|
|
512
|
+
- **`costly` now asks whether one probe buys an *answer*, not whether it costs
|
|
513
|
+
more than one response.** The two were the same question while every method's
|
|
514
|
+
probe was also its unit of information. `boolean` is the first where they come
|
|
515
|
+
apart: each of its probes is one ordinary request and none of them means
|
|
516
|
+
anything alone, because the answer is the partition across the whole series.
|
|
517
|
+
Read the old way it would have landed in the wave the enumeration driver runs
|
|
518
|
+
*first* — the one that exists to be answered cheaply — ahead of `reflected`
|
|
519
|
+
and `eval` and spending the same per-question budget, at 27 requests before it
|
|
520
|
+
could say a word. It sits with `time` instead, and the suite now holds every
|
|
521
|
+
aggregate method to that.
|
|
522
|
+
|
|
523
|
+
### Security
|
|
524
|
+
|
|
525
|
+
- **The `OR` connectives ship behind `--verify-active-risk stateful`.** A
|
|
526
|
+
predicate probe breaks out of a condition the application already wrote, and
|
|
527
|
+
the connective is this method's command separator. `AND` differentiates only
|
|
528
|
+
where the application's own predicate is true and `OR` only where it is false,
|
|
529
|
+
so they are complements and dropping `OR` is a blind spot rather than a
|
|
530
|
+
saving. But a true predicate `OR`-ed into a `DELETE … WHERE` took a table from
|
|
531
|
+
3 rows to 0, where the same predicate `AND`-ed into it left all 3 — so it goes
|
|
532
|
+
at the top rung, held back by default, with the run naming every shape it held
|
|
533
|
+
and the flag that sends it. `AND` and the bare form stay `safe` and change
|
|
534
|
+
nothing.
|
|
535
|
+
|
|
536
|
+
- **Nothing widens `confirmed`.** The new method's ceiling is one tier below it
|
|
537
|
+
and the suite holds that as behaviour, not as an attribute: `confirm_series`
|
|
538
|
+
is driven across 200 series including the perfect one, and none of them
|
|
539
|
+
reaches `confirmed`.
|
|
540
|
+
|
|
541
|
+
## [2.44.0] — 2026-09-23
|
|
542
|
+
|
|
543
|
+
### Fixed
|
|
544
|
+
|
|
545
|
+
- **`--evade low` substituted inside quoted programs and broke them.** Every
|
|
546
|
+
space became `${IFS}`, including the ones inside `awk 'BEGIN{print "RK" a+b
|
|
547
|
+
"RK"}'` — and inside single quotes `${IFS}` is literal text, not an
|
|
548
|
+
expansion, so awk was handed `BEGIN{print${IFS}"RK"...` and answered with a
|
|
549
|
+
syntax error.
|
|
550
|
+
|
|
551
|
+
Measured shape by shape against an unfiltered target: **8 probe shapes that
|
|
552
|
+
the canonical form executes broke at the rung, and none improved.** The
|
|
553
|
+
substitution stops at a quote now. Double quotes are left alone from the
|
|
554
|
+
other side — `${IFS}` *does* expand inside them, so substituting there would
|
|
555
|
+
change the string the target computes rather than the spacing around it.
|
|
556
|
+
|
|
557
|
+
### Changed
|
|
558
|
+
|
|
559
|
+
- **A rung is a retry for a refused probe, not a posture for the run.** It was
|
|
560
|
+
applied to every probe regardless of whether anything was being filtered,
|
|
561
|
+
which is a pure loss on a target with no filter. Against a filter that blocks
|
|
562
|
+
whitespace it turned 1 confirmation into 5; against an unfiltered target it
|
|
563
|
+
now costs **zero** extra requests, because nothing was refused.
|
|
564
|
+
|
|
565
|
+
`--evade` is therefore a **ceiling**. Every probe goes out canonical, and
|
|
566
|
+
only a refused one is retried, up to that ceiling — at most one request per
|
|
567
|
+
rung. The ladder a run builds no longer depends on the rung at all: all three
|
|
568
|
+
settings build the same 42 probes, of which the same 32 execute. Before, the
|
|
569
|
+
rung built a smaller ladder and 13 fewer of its probes ran.
|
|
570
|
+
|
|
571
|
+
The run reports how many retries it made and which rung got through, whatever
|
|
572
|
+
it concluded — including a run that confirmed, where the retry is the reason
|
|
573
|
+
it did.
|
|
574
|
+
|
|
575
|
+
### Added
|
|
576
|
+
|
|
577
|
+
- **A second evasion rung, `high`.** The measurement turned up two classes of
|
|
578
|
+
filter and the shipped rung only addressed one. `low` removes whitespace;
|
|
579
|
+
`high` also splits the command word with an expansion that vanishes
|
|
580
|
+
(`ec$@ho`), for a filter matching command names. It is applied before the
|
|
581
|
+
whitespace substitution, because afterwards the split lands inside `${IFS}`
|
|
582
|
+
and makes `${I$@FS}` — neither an expansion nor a command.
|
|
583
|
+
|
|
584
|
+
Four boundaries the retry has to respect, each of which was a way to turn a
|
|
585
|
+
vulnerable target into a `negative` — worse than the `blocked` the rung sits
|
|
586
|
+
beside, because `blocked` at least says the run learned nothing:
|
|
587
|
+
|
|
588
|
+
- **Unix shell probes only.** `${IFS}` and `$@` are POSIX. A cmd.exe probe
|
|
589
|
+
rewritten with them loses its spaces, so a whitespace filter answers 200
|
|
590
|
+
and the retry counts as a win while cmd.exe cannot run it.
|
|
591
|
+
- **A break-out context opens by *closing* a quote.** Reading that leading
|
|
592
|
+
quote as an opener left `'; echo …` untouched, so the rung did nothing on
|
|
593
|
+
exactly the contexts a filter is most likely to sit in front of.
|
|
594
|
+
- **A followup is read again after a retry that lands.** `file` writes its
|
|
595
|
+
token on the request that arrives, so a body read before the retry is a
|
|
596
|
+
read of a file that did not exist yet.
|
|
597
|
+
- **A redirect is never retried.** The build-time transform took an explicit
|
|
598
|
+
`evade=False` for these, and that parameter stopped doing anything when the
|
|
599
|
+
rung became a retry — a guard lost in the move, restored where the retry
|
|
600
|
+
now happens.
|
|
601
|
+
|
|
602
|
+
The aggregate methods escalate too. `time`, `oob`, `lookup` and `deser` take a
|
|
603
|
+
different branch, and leaving it out meant the documented ceiling did nothing
|
|
604
|
+
for four of the eight methods — the same branch, and the same omission, as the
|
|
605
|
+
refusal check one change earlier.
|
|
606
|
+
|
|
607
|
+
## [2.43.0] — 2026-09-23
|
|
608
|
+
|
|
609
|
+
### Added
|
|
610
|
+
|
|
611
|
+
- **A ninth verdict: `blocked`.** A run whose payloads a filter refuses has
|
|
612
|
+
learned nothing about the sink, and it used to say otherwise. Measured
|
|
613
|
+
against a real command injection behind a filter that 403s a space or a
|
|
614
|
+
separator:
|
|
615
|
+
|
|
616
|
+
```
|
|
617
|
+
before: sent 10 probes: negative=10
|
|
618
|
+
after: sent 10 probes: blocked=10
|
|
619
|
+
```
|
|
620
|
+
|
|
621
|
+
`negative` asserts that the probes **reached** the target — the project's own
|
|
622
|
+
tier rule says so. They reached a filter. That is the same false clean
|
|
623
|
+
`nothing-tested` exists to prevent, one level further in.
|
|
624
|
+
|
|
625
|
+
**The signal is differential**, like everything else this tool decides: the
|
|
626
|
+
payload-free control got through and the probe did not, so what was refused
|
|
627
|
+
is the payload. An endpoint answering 403 to everything — an auth wall, a
|
|
628
|
+
path that does not exist for this session — refuses the control too and is
|
|
629
|
+
not mistaken for a filter. Verified against exactly that case.
|
|
630
|
+
|
|
631
|
+
4xx only. A 5xx is as likely to be the payload *breaking* the application,
|
|
632
|
+
which means it reached something, and reading that as blocked would hide the
|
|
633
|
+
one response saying the sink is live. No vendor list and no block-page
|
|
634
|
+
fingerprints: a status the control did not get is the whole signal.
|
|
635
|
+
|
|
636
|
+
A run where **some** probes got through stays a real `negative` — the sink
|
|
637
|
+
saw those — and the refusals are reported either way.
|
|
638
|
+
|
|
639
|
+
**A refusal never unmakes evidence.** Only a `negative` is replaced, because
|
|
640
|
+
it is the only verdict a refusal contradicts. An application can execute the
|
|
641
|
+
payload and then answer 400 with the output in its body, and the oracle has
|
|
642
|
+
already proven execution from a value random to that probe; overwriting that
|
|
643
|
+
would turn demonstrated RCE into a false negative, which is worse than the
|
|
644
|
+
false clean this verdict removes.
|
|
645
|
+
|
|
646
|
+
**Every method is covered, including the three that decide per probe from a
|
|
647
|
+
series.** `oob`, `lookup` and `deser` take a different branch, which returned
|
|
648
|
+
before refusal was considered — so a callback run whose every probe was
|
|
649
|
+
refused reported `negative` for each of them, because no callback arrived.
|
|
650
|
+
|
|
651
|
+
Refusals are counted in **requests**, like the status tally beside them: an
|
|
652
|
+
aggregate method fires a whole ladder and reports one row, so counting rows
|
|
653
|
+
described the same run with different arithmetic and hid a partly filtered
|
|
654
|
+
series entirely.
|
|
655
|
+
|
|
656
|
+
### Fixed
|
|
657
|
+
|
|
658
|
+
- **The advice on a filtered run pointed at the wrong thing entirely.** "The
|
|
659
|
+
target may be patched" reads as a clean bill of health for a target that was
|
|
660
|
+
never reached; the blind-sink list names methods a filter refuses in exactly
|
|
661
|
+
the same way; and the second-order line said the target *accepted* an input
|
|
662
|
+
it had in fact rejected with a 403. None of the three fires on a refused run
|
|
663
|
+
now. It is replaced by what was actually observed, with the flags that change
|
|
664
|
+
the payload's shape.
|
|
665
|
+
|
|
666
|
+
## [2.42.0] — 2026-09-23
|
|
667
|
+
|
|
668
|
+
### Added
|
|
669
|
+
|
|
670
|
+
- **A run that confirms nothing now names the second-order oracle.** Measured
|
|
671
|
+
against a target that stores on one endpoint and renders through a shell on
|
|
672
|
+
another — a real RCE — every probe read `negative`, and the run answered with
|
|
673
|
+
four methods that are all negative there too, because the execution does not
|
|
674
|
+
happen on the request being measured:
|
|
675
|
+
|
|
676
|
+
```
|
|
677
|
+
--methods time -> negative=4
|
|
678
|
+
--observe-url -> confirmed, first run
|
|
679
|
+
```
|
|
680
|
+
|
|
681
|
+
The one flag that works was named nowhere. It is named now, and **not gated
|
|
682
|
+
on which methods have run**: the blind-sink list is, so an operator who had
|
|
683
|
+
already tried the expensive methods — exactly the one with nothing left but
|
|
684
|
+
second order — was told only that the target might be patched.
|
|
685
|
+
|
|
686
|
+
The wording follows what the run observed rather than what it assumes. Input
|
|
687
|
+
returned verbatim means a sink that reflects without executing; input that
|
|
688
|
+
never came back means this response cannot show what became of it; and a run
|
|
689
|
+
of aggregate methods alone, which record no per-probe observation, claims
|
|
690
|
+
neither. A swallowed input is equally a blind sink or a stored one — `ping
|
|
691
|
+
<input> >/dev/null` returns nothing either — so both are named and neither is
|
|
692
|
+
picked.
|
|
693
|
+
|
|
694
|
+
- **Whether the target returned the input is recorded on every probe.** It was
|
|
695
|
+
already computed on the confirmed path, where it becomes "target also
|
|
696
|
+
reflects the payload verbatim"; a negative probe never looked, and the
|
|
697
|
+
negative run is the one that has to say what it saw.
|
|
698
|
+
|
|
699
|
+
Measured against the payload that actually went out. A carrier that
|
|
700
|
+
multiplies through a filter never spells the joined `a*b` out — Liquid sends
|
|
701
|
+
`{{ a | times: b }}` and Django `{% widthratio a 1 b %}` — so an endpoint
|
|
702
|
+
echoing the whole payload recorded a measured `False`, and a target profile
|
|
703
|
+
filtering the `*` shapes leaves only those, at which point the run would
|
|
704
|
+
report "returned none of it" about a target that returned everything.
|
|
705
|
+
|
|
706
|
+
Recorded only when there was a response to look at. A delivery error is not
|
|
707
|
+
an observation, and a `False` for one would put an unmeasured claim exactly
|
|
708
|
+
where the unobserved branch belongs.
|
|
709
|
+
|
|
710
|
+
## [2.41.0] — 2026-09-23
|
|
711
|
+
|
|
712
|
+
### Fixed
|
|
713
|
+
|
|
714
|
+
- **A candidate that confirmed RCE was never asked whether it was also a
|
|
715
|
+
deserialization or lookup sink.** The enumeration driver split the methods
|
|
716
|
+
into a cheap wave and an expensive one and skipped the expensive wave once
|
|
717
|
+
execution was proven — reasonable for `time` and `oob`, which would only put
|
|
718
|
+
a second name on one finding, and wrong for `lookup` and `deser`, which
|
|
719
|
+
report *different properties* with their own remediation. They sat on the
|
|
720
|
+
expensive side of a hand-written set of names, so the answer was never asked
|
|
721
|
+
for and never reported.
|
|
722
|
+
|
|
723
|
+
The split now comes from the tier each method declares. Anything reporting
|
|
724
|
+
`confirmed` or `needs-review` is answering *did this target execute my
|
|
725
|
+
input*; anything else is a different question and is never skipped for an
|
|
726
|
+
answer to that one.
|
|
727
|
+
|
|
728
|
+
- **`--max-payloads` is spent per question, not per wave and not per
|
|
729
|
+
candidate.** Per wave it quietly doubled: a run capped at 5 sent 10 probes to
|
|
730
|
+
every candidate that did not confirm, while the cost line printed before any
|
|
731
|
+
traffic said 5. Bounding the candidate instead starves the different
|
|
732
|
+
question — the cheap methods eat the whole allowance and `deser` never runs,
|
|
733
|
+
which is the same finding lost by another route. Every method asking about
|
|
734
|
+
execution now shares one allowance, each different property gets its own, and
|
|
735
|
+
the cost line names how many questions are being asked — and **sums** the
|
|
736
|
+
estimate across them. Counting the cap once while the run grants it per
|
|
737
|
+
question advertised 44 requests for a run that sent 80, which is wrong in
|
|
738
|
+
the direction that matters for the operator bounding a monitored engagement.
|
|
739
|
+
|
|
740
|
+
- **A second-order confirmation settles its carrier too.** With `--observe-url`
|
|
741
|
+
a probe can read negative in the response it drew and `confirmed` on the
|
|
742
|
+
observed channel a moment later. The stop was decided from the pre-poll
|
|
743
|
+
verdict, so the carrier kept probing after it had in fact confirmed —
|
|
744
|
+
spending the budget the stop exists to hand to carriers not yet examined,
|
|
745
|
+
which is the coverage loss this change was written to remove, reappearing on
|
|
746
|
+
the one oracle that needs a second request to answer.
|
|
747
|
+
|
|
748
|
+
### Changed
|
|
749
|
+
|
|
750
|
+
- **A carrier that has confirmed stops there** (`--confirm-depth first`, the
|
|
751
|
+
new default). One carrier is one method in one environment and context, and
|
|
752
|
+
once it has confirmed every further shape of it can only say the same thing
|
|
753
|
+
again.
|
|
754
|
+
|
|
755
|
+
Measured against an executing target: one candidate spent 115 of its 120
|
|
756
|
+
probes after the first confirmation, and printed 32 confirmations of which 29
|
|
757
|
+
were duplicates inside a single carrier. Those probes were not idle — they
|
|
758
|
+
were spent instead of reaching carriers never examined at all. At the same
|
|
759
|
+
budget the run went from **4 carriers examined to 23**, and from **4
|
|
760
|
+
environments reached to 9**. This buys coverage rather than saving requests.
|
|
761
|
+
|
|
762
|
+
The stop is per carrier and **never** per candidate: a candidate may confirm
|
|
763
|
+
as `unix` while a later `nodejs` carrier is the only thing a different target
|
|
764
|
+
would have shown. `--confirm-depth every` maps every shape a sink accepts,
|
|
765
|
+
which is what writing a proof of concept by hand needs, and the run reports
|
|
766
|
+
how many shapes it held back and which carriers stopped — a fourth tally
|
|
767
|
+
beside the profile drops, the safety holds and the reach notes, because it
|
|
768
|
+
says a fourth thing: the probe could have been sent and had nothing left to
|
|
769
|
+
establish.
|
|
770
|
+
|
|
771
|
+
- **A method declares what one of its probes costs.** `CHEAP_DETECTION_METHODS`
|
|
772
|
+
was a set literal; it is now derived from a `costly` attribute each class sets
|
|
773
|
+
for its own reason — a real sleep, a wait for a callback, a second fetch.
|
|
774
|
+
Every hand-written list naming methods in this repository has gone stale, and
|
|
775
|
+
this one had put `lookup` and `deser` where being skipped cost findings.
|
|
776
|
+
|
|
777
|
+
## [2.40.0] — 2026-09-22
|
|
778
|
+
|
|
779
|
+
### Added
|
|
780
|
+
|
|
781
|
+
- **A carrier may take the operands apart, and two engines need it.** Every
|
|
782
|
+
expression carrier until now substituted `__EXPR__` — the joined `a*b` —
|
|
783
|
+
which quietly assumed the engine has an arithmetic operator. Two widely
|
|
784
|
+
deployed ones do not, and both were measured as false negatives:
|
|
785
|
+
|
|
786
|
+
```
|
|
787
|
+
liquid every bare form missing; {{ 45013 | times: 45989 }} -> 2070102857
|
|
788
|
+
django every bare form missing; {% widthratio 45013 1 45989 %} -> 2070102857
|
|
789
|
+
```
|
|
790
|
+
|
|
791
|
+
Liquid multiplies with a filter and Django with a tag, so neither form can be
|
|
792
|
+
written as a single expression — an application that really does evaluate the
|
|
793
|
+
template was reported `negative`, which is the same shape as `oob` against a
|
|
794
|
+
`${jndi:...}` sink: probes that reach the target and cannot speak its
|
|
795
|
+
language. A carrier template may now use `__A__` and `__B__` as well as
|
|
796
|
+
`__EXPR__`.
|
|
797
|
+
|
|
798
|
+
Neither payload contains the product, so a target that merely echoes the
|
|
799
|
+
payload still cannot read as `confirmed`. A test pins that for every shipped
|
|
800
|
+
carrier.
|
|
801
|
+
|
|
802
|
+
### Changed
|
|
803
|
+
|
|
804
|
+
- **A carrier template not parameterised by both operands is skipped, not
|
|
805
|
+
sent.** With no token at all it renders the same constant every probe; with
|
|
806
|
+
only one operand the target is never handed the other, so nothing it can
|
|
807
|
+
compute is the product RCEKit is looking for. Either way the product would
|
|
808
|
+
not be evidence the target computed anything — and a probe that cannot
|
|
809
|
+
confirm still counts toward the coverage a run reports, which is the part
|
|
810
|
+
that matters more than the wasted request.
|
|
811
|
+
|
|
812
|
+
- **`--eval-engines` names every carrier the corpus ships.** Its help listed
|
|
813
|
+
three engines by hand and two were added. An operator narrowing that flag
|
|
814
|
+
reads the list and nothing else, so a stale one says an engine needs no
|
|
815
|
+
carrier when it does, and they cut the only probe that could have confirmed
|
|
816
|
+
it. A test now holds the help text to the corpus.
|
|
817
|
+
|
|
818
|
+
- **What the survey measured and did not ship is recorded too.**
|
|
819
|
+
`eval_carrier_survey` in the corpus now names the engines that need no
|
|
820
|
+
carrier — nunjucks 3.2.4, tornado 6.5.10, mako 1.4.1, chameleon 4.6.0,
|
|
821
|
+
smarty 5.8.4 and Ruby's ERB all return the product from a bare form — and
|
|
822
|
+
those that are
|
|
823
|
+
out of reach. Handlebars 4.7.9 fails every bare form and, being logic-less
|
|
824
|
+
with no built-in arithmetic helper, has no template text that computes a
|
|
825
|
+
product at all. Go `text/template` 1.23 fails every bare form too, and the
|
|
826
|
+
only forms that do return the product — `{{printf "%d" <product>}}` and
|
|
827
|
+
`{{<product>}}` — hand the target the answer, so a target that merely echoed
|
|
828
|
+
them would read as `confirmed`. That is the rule a carrier lives under: **a
|
|
829
|
+
carrier may not carry its own result.**
|
|
830
|
+
|
|
831
|
+
Saying so is worth more than a carrier that cannot work, and it stops the
|
|
832
|
+
next person re-running the same survey — or shipping the Go form, which the
|
|
833
|
+
survey itself produced and which looks like a carrier until you ask what an
|
|
834
|
+
echoing target would return.
|
|
835
|
+
|
|
836
|
+
## [2.39.0] — 2026-09-22
|
|
837
|
+
|
|
838
|
+
### Added
|
|
839
|
+
|
|
840
|
+
- **Every part of a multipart body is an injection point.** None of them were
|
|
841
|
+
before. The form branch matched on `=` appearing anywhere in the body, so a
|
|
842
|
+
`multipart/form-data` capture was split on `&` and yielded exactly one
|
|
843
|
+
candidate, named after a `Content-Disposition` line:
|
|
844
|
+
|
|
845
|
+
```
|
|
846
|
+
form | '--X\r\nContent-Disposition: form-data; name' | body param '...'
|
|
847
|
+
```
|
|
848
|
+
|
|
849
|
+
Every probe for that point rewrote a *part header*, so it could confirm
|
|
850
|
+
nothing, while `user`, `avatar` and `note` — the fields the form actually
|
|
851
|
+
posts — were never reached. The run still printed a point and a probe count,
|
|
852
|
+
which is the part that matters: coverage reported and not delivered reads
|
|
853
|
+
exactly like a clean target.
|
|
854
|
+
|
|
855
|
+
Several parts may post under one name — a multi-file input and a checkbox
|
|
856
|
+
array both do — so a part is addressed by its **index**, as a JSON leaf is
|
|
857
|
+
addressed by its token path. Addressing by name alone rewrote the first part
|
|
858
|
+
for every candidate, which is the same failure one level further in: three
|
|
859
|
+
files, three points, the first file probed three times and the other two
|
|
860
|
+
never touched.
|
|
861
|
+
|
|
862
|
+
`multipart` is now its own kind, recognised from `Content-Type` and decided
|
|
863
|
+
before the form branch can see the body. A file part is a candidate too — its
|
|
864
|
+
content is the value under test, while its `filename` and `Content-Type` stay
|
|
865
|
+
as captured. Verified end to end against a target that parses with the
|
|
866
|
+
standard library's MIME parser: `multipart field 'note': confirmed`, with the
|
|
867
|
+
non-vulnerable `user` field `negative` beside it.
|
|
868
|
+
|
|
869
|
+
### Changed
|
|
870
|
+
|
|
871
|
+
- **A GraphQL request is ordered by what can actually confirm.** Its variables
|
|
872
|
+
were already enumerated — they are JSON leaves — but they were tried in
|
|
873
|
+
document order alongside `query` and `operationName`, and those two cannot
|
|
874
|
+
confirm anything. A payload in `query` *replaces* the operation document, so
|
|
875
|
+
the server answers with a parse error before a resolver runs, and
|
|
876
|
+
`operationName` then names an operation that is no longer there. On the
|
|
877
|
+
capture this was measured against they were two points of five, each one a
|
|
878
|
+
full probe ladder.
|
|
879
|
+
|
|
880
|
+
They are moved behind the variables, **not dropped**. A server that logs the
|
|
881
|
+
query document before parsing it is reachable through exactly that field,
|
|
882
|
+
which is the route Log4Shell took through access logs, so a full run still
|
|
883
|
+
tests both and `--max-points` now cuts the least likely to pay first. A plain
|
|
884
|
+
`{"query": ...}` body with no `variables` is left in document order: it is as
|
|
885
|
+
likely to be a search API, and there the query field is the one worth testing.
|
|
886
|
+
Carrying both keys is not enough either — the `query` string has to open like
|
|
887
|
+
a GraphQL document, so `{"query": "red shoes", "variables": {...}}` keeps its
|
|
888
|
+
real injection point where a bounded `--max-points` run will still reach it.
|
|
889
|
+
|
|
890
|
+
### Fixed
|
|
891
|
+
|
|
892
|
+
- **The verdict table was missing a verdict.** The README said "seven verdicts
|
|
893
|
+
that are never collapsed into each other" and listed seven, while the tool
|
|
894
|
+
reports eight: `lookup-sink` was absent from the one table whose whole job is
|
|
895
|
+
to enumerate them — in a README that uses the word two tables higher, in the
|
|
896
|
+
Log4Shell CVE row, and again in the methods table. The sentence counted the
|
|
897
|
+
rows the table had rather than the verdicts there are, so the omission never
|
|
898
|
+
contradicted itself and nothing failed.
|
|
899
|
+
|
|
900
|
+
Found by reading the README end to end for this change. A test now reads the
|
|
901
|
+
table against `DETECTION_METHODS`, so a verdict a method declares and the
|
|
902
|
+
table does not carry is a failure rather than a silence.
|
|
903
|
+
|
|
904
|
+
- **A multipart body now goes out with the line endings it needs.** `-r`
|
|
905
|
+
normalises the whole request to LF, so the CRLF delimiters RFC 2046 requires
|
|
906
|
+
were gone by the time anything was sent. Bodies rendered for a multipart point
|
|
907
|
+
are re-serialised canonically — probe and payload-free control alike, so the
|
|
908
|
+
two differ in the field under test and in nothing else. Part content is left
|
|
909
|
+
character for character, so a lone newline inside an uploaded text file
|
|
910
|
+
survives.
|
|
911
|
+
|
|
912
|
+
## [2.38.0] — 2026-09-21
|
|
913
|
+
|
|
914
|
+
### Added
|
|
915
|
+
|
|
916
|
+
- **A probe shape may reach past the run's tier and be sent anyway**, when its
|
|
917
|
+
effect is one the run undoes by saying it happened. `reaches_past` is that
|
|
918
|
+
declaration, beside `safety`, which stays for an effect a notice cannot take
|
|
919
|
+
back.
|
|
920
|
+
|
|
921
|
+
Reach wins where the two pull against each other. Detection the tool could
|
|
922
|
+
have done and did not is a false negative wearing a safety label, and it
|
|
923
|
+
costs more than the noise it saves.
|
|
924
|
+
|
|
925
|
+
`deser`'s DNS gadget is the first of these, and closes an inconsistency
|
|
926
|
+
recorded a version ago: it makes the target resolve a name -- the very thing
|
|
927
|
+
`oob` and `lookup` are refused for at `safe` -- while its only gate was
|
|
928
|
+
`--oob-host`. Holding it back would have sent fewer probes at the default
|
|
929
|
+
tier. It goes, and the run reports how far it reached.
|
|
930
|
+
|
|
931
|
+
Three tallies now, because they say three different things and one number
|
|
932
|
+
would state the wrong one about all of them: the profile dropped it (it could
|
|
933
|
+
not have reached the sink), the tier held it (it could, and was not sent), or
|
|
934
|
+
it reached past the tier (it was sent, further than asked).
|
|
935
|
+
|
|
936
|
+
`lookup`'s `ldap://` and `rmi://` are the edge the rule has, and stay at
|
|
937
|
+
`stateful`: a class fetched from an address RCEKit did not choose is not
|
|
938
|
+
something a notice takes back.
|
|
939
|
+
|
|
940
|
+
- **Two worked examples in the README's Quick start**, because the fullest run
|
|
941
|
+
RCEKit can make was not shown anywhere near the front.
|
|
942
|
+
|
|
943
|
+
The first is a captured request: most sinks worth testing sit behind a POST
|
|
944
|
+
with a session cookie, a content type and a body, and `--verify-url` carries
|
|
945
|
+
none of that. The second is everything the tool has -- injection-point
|
|
946
|
+
enumeration across every value in that request, every method, callbacks, the
|
|
947
|
+
top rung -- with a table of what each flag opens up and the cost line that
|
|
948
|
+
prints before it fires.
|
|
949
|
+
|
|
950
|
+
Both were run before being written down. The point the Quick start never
|
|
951
|
+
made: `--auto-params` needs `-r`, so the fullest run is not reachable from a
|
|
952
|
+
URL at all, which is worth knowing before concluding a target is clean.
|
|
953
|
+
|
|
954
|
+
## [2.37.0] — 2026-09-20
|
|
955
|
+
|
|
956
|
+
### Added
|
|
957
|
+
|
|
958
|
+
- **A detection method declares the risk rung it needs**, and the engine reads
|
|
959
|
+
it. `SAFETY_ORDER` has labelled corpus payloads `safe` / `intrusive` /
|
|
960
|
+
`stateful` from the start and the query-language bridges followed; detection
|
|
961
|
+
methods did not. Each risky one was gated by a hand-written branch in
|
|
962
|
+
`main()` naming it, so a new method meant remembering to add another -- and a
|
|
963
|
+
probe shape with nowhere to declare its rung was deleted rather than gated.
|
|
964
|
+
|
|
965
|
+
| method | rung | |
|
|
966
|
+
|---|---|---|
|
|
967
|
+
| `reflected`, `eval`, `time`, `deser` | `safe` | compute, delay, or parse |
|
|
968
|
+
| `oob`, `lookup` | `intrusive` | makes the target open outbound connections |
|
|
969
|
+
| `file`, `write` | `stateful` | writes to the target |
|
|
970
|
+
|
|
971
|
+
`file` and `write` are gated by their own configuration rather than by the
|
|
972
|
+
rung: neither does anything until a directory to write into and a URL to read
|
|
973
|
+
it back from are named, which says more than a tier would, and asking for the
|
|
974
|
+
flag as well would refuse a command that works today. Nothing that runs today
|
|
975
|
+
stops running.
|
|
976
|
+
|
|
977
|
+
- **A probe shape may need a higher rung than its method**, so coverage that
|
|
978
|
+
only makes sense at the top tier has somewhere to live instead of being
|
|
979
|
+
deleted. **`--methods lookup` sends `ldap://` and `rmi://` again**, at
|
|
980
|
+
`stateful`.
|
|
981
|
+
|
|
982
|
+
They were removed in 2.36.0 on the argument that `dns://` resolves wherever
|
|
983
|
+
`ldap://` would. That claim was too strong -- a filter catching the string
|
|
984
|
+
`dns:` and not `ldap:`, or a trimmed runtime without the DNS provider,
|
|
985
|
+
defeats it -- and a sink that takes one scheme and not the other is exactly
|
|
986
|
+
the sink this method is for. The reason they were removed was real: they
|
|
987
|
+
continue *past* resolution and connect to whatever address the answer named,
|
|
988
|
+
which is not an address RCEKit chose. That is a rung, not a reason to drop
|
|
989
|
+
coverage.
|
|
990
|
+
|
|
991
|
+
- **The run says what the rung held back**, counted apart from the target
|
|
992
|
+
profile's drops and with the flag that would send it. The two say different
|
|
993
|
+
things: a profile drop means the probe *could not have* reached the sink,
|
|
994
|
+
while this means it could and the operator chose not to send it. Reporting
|
|
995
|
+
them together would state the first about the second.
|
|
996
|
+
|
|
997
|
+
- **A method declares the weaker tiers it really reports**, not only its
|
|
998
|
+
ceiling. `write` reports `needs-review` for a file that is served but not
|
|
999
|
+
interpreted and `deser` for a shape fingerprint, and three separate places
|
|
1000
|
+
had to know that -- the documentation tests, the benchmark's expectation
|
|
1001
|
+
whitelist, and the advice printed after a clean in-band run. Each kept its
|
|
1002
|
+
own answer; the docs test exempted `needs-review` for *every* method, so
|
|
1003
|
+
`lookup` could name a verdict it never emits and pass.
|
|
1004
|
+
|
|
1005
|
+
### Changed
|
|
1006
|
+
|
|
1007
|
+
- **`--methods file` and `--methods write` with nothing configured now say
|
|
1008
|
+
so by name.** They were simply not applicable before, so the run built no
|
|
1009
|
+
probes and reported `nothing-tested` -- which is the quietest way this tool
|
|
1010
|
+
can fail and reads much like a clean target.
|
|
1011
|
+
|
|
1012
|
+
- **`docs/reference.md` carries a `Rung` column**, held against the class by a
|
|
1013
|
+
test. The tier column already was; this is the same claim one column over.
|
|
1014
|
+
|
|
1015
|
+
### Fixed
|
|
1016
|
+
|
|
1017
|
+
- **A config-gated method is not re-gated by the rung at runtime.** The
|
|
1018
|
+
pre-flight lets `--methods file --webroot ... --web-base-url ...` through
|
|
1019
|
+
because the configuration is the gate, but the probe filter read the run's
|
|
1020
|
+
default `safe` ceiling and held every probe inheriting the method's
|
|
1021
|
+
`stateful` rung. The CLI accepted a documented invocation and then reported
|
|
1022
|
+
`nothing-tested` -- the quietest way this tool can fail, and the thing the
|
|
1023
|
+
rung work was supposed to remove rather than add.
|
|
1024
|
+
|
|
1025
|
+
Nothing caught it because every `file` test builds the method's config
|
|
1026
|
+
directly, without `max_safety`, so the ceiling fell back to the method's own
|
|
1027
|
+
rung and the probes went out. A run through the CLI with the channel
|
|
1028
|
+
configured is the one thing that would have, and there is one now.
|
|
1029
|
+
|
|
1030
|
+
- **The cost estimate applies the risk tier as well as the target profile.** It
|
|
1031
|
+
counted every shape a method built, so once a rung could narrow a method the
|
|
1032
|
+
pre-flight figure over-counted -- three times over for `lookup` at the
|
|
1033
|
+
default tier, which is exactly the operator who narrowed the run on purpose.
|
|
1034
|
+
Both paths share one predicate now, and it counts nothing, so an estimate
|
|
1035
|
+
never moves the numbers the report prints.
|
|
1036
|
+
|
|
1037
|
+
## [2.36.0] — 2026-09-19
|
|
1038
|
+
|
|
1039
|
+
### Added
|
|
1040
|
+
|
|
1041
|
+
- **`--methods lookup`: confirmation for an expression-lookup sink**, the shape
|
|
1042
|
+
Log4Shell has, where the sink resolves a URI instead of running a command.
|
|
1043
|
+
|
|
1044
|
+
`oob` could not reach one. Its `applicable` does admit `java` -- a Java
|
|
1045
|
+
application can shell out, so the environment is genuinely shell-capable --
|
|
1046
|
+
but every probe it builds is a shell command: `nslookup`, `curl`, `certutil`,
|
|
1047
|
+
`iwr`. A sink that interpolates `${jndi:...}` runs none of them, so the method
|
|
1048
|
+
applied, sent its whole ladder, and came back `negative` on a target that is
|
|
1049
|
+
exploitable. The README's Log4Shell row rested on the standalone listener and
|
|
1050
|
+
a generated payload file, which produce no verdict row at all, so nothing in
|
|
1051
|
+
the engine could reproduce that claim.
|
|
1052
|
+
|
|
1053
|
+
The probes are lookups and nothing else, and they depend on the injection
|
|
1054
|
+
context rather than the environment, exactly as `eval`'s do. The oracle is the
|
|
1055
|
+
one `oob` already uses: a token the target could only have learned by
|
|
1056
|
+
resolving what it was handed. The expression resolves `<token>.<host>`, so
|
|
1057
|
+
the in-process DNS listener is the entire apparatus -- no LDAP or RMI server
|
|
1058
|
+
is needed, and none is started.
|
|
1059
|
+
|
|
1060
|
+
**It reports `lookup-sink`, never `confirmed`.** A callback proves the sink
|
|
1061
|
+
resolved a URI RCEKit chose -- that it evaluated the expression it was handed.
|
|
1062
|
+
It does not prove the target ran attacker code: Log4Shell becomes RCE when the
|
|
1063
|
+
LDAP server answers with a loadable class. So the method gets its own
|
|
1064
|
+
proven-sink tier beside `deserialization-sink`, with its own section in the
|
|
1065
|
+
report, and `confirmed` keeps meaning executed.
|
|
1066
|
+
|
|
1067
|
+
**Only `jndi:dns://` is sent, and that is the security property rather than a
|
|
1068
|
+
shortcut.** A name lookup can be nothing else. `ldap://` and `rmi://` continue
|
|
1069
|
+
*past* resolution and open a connection to whatever address the answer named
|
|
1070
|
+
-- by default `127.0.0.1`, which is the target's own loopback. Whatever
|
|
1071
|
+
replies on :389 or :1099 is not RCEKit, so a reference could come back and a
|
|
1072
|
+
class be instantiated: the tool would have crossed the line this method exists
|
|
1073
|
+
to stop short of, having promised it had not. Dropping the two schemes also
|
|
1074
|
+
costs no coverage -- `DnsContextFactory` ships in the JDK, so `dns://`
|
|
1075
|
+
resolves wherever `ldap://` would, and on 2.15.0 it still resolves where
|
|
1076
|
+
`ldap://` no longer does. The proof is the callback and the finding is "this
|
|
1077
|
+
sink resolved a URI I chose" -- the same line `deser` draws, drawn here before
|
|
1078
|
+
it can be crossed.
|
|
1079
|
+
|
|
1080
|
+
Behind the same two gates as `oob`, and sharing its listener: it needs
|
|
1081
|
+
`--oob-host`, and it is held back at the default safety tier because it makes
|
|
1082
|
+
the target open outbound connections. Without `--oob-host` it builds no probes
|
|
1083
|
+
at all, which the engine reports as `nothing-tested` -- never `negative`. An
|
|
1084
|
+
address literal is a literal whichever family it is from: `::1` and `[::1]`
|
|
1085
|
+
build nothing, the same as `10.0.0.1`, because a token can only ride in a DNS
|
|
1086
|
+
label. `blind_sink_advice` names the method as proving a lookup sink and not
|
|
1087
|
+
execution, in a list where `oob` and `file` mean confirmed execution, and the
|
|
1088
|
+
README's Log4Shell demo heading says `lookup-sink` rather than `confirmed` --
|
|
1089
|
+
it was left claiming execution beside the table row that no longer does.
|
|
1090
|
+
|
|
1091
|
+
## [2.35.5] — 2026-09-19
|
|
1092
|
+
|
|
1093
|
+
### Fixed
|
|
1094
|
+
|
|
1095
|
+
- **The `--insecure` notice describes the run it is in.** It was printed as
|
|
1096
|
+
soon as the flag was seen, so a generation-only run, `--doctor`, or a plain
|
|
1097
|
+
HTTP target all announced a TLS downgrade that never happened -- and it named
|
|
1098
|
+
both downgrade rungs even on an OpenSSL build that had refused one of them.
|
|
1099
|
+
Reporting a downgrade on a run that opened no TLS connection is the same
|
|
1100
|
+
defect as reporting a probe that was never sent, in the one line written to
|
|
1101
|
+
be an audit of the run.
|
|
1102
|
+
|
|
1103
|
+
Building the context and saying so are now separate questions. The context is
|
|
1104
|
+
built whatever the target's scheme is -- urllib follows a redirect with the
|
|
1105
|
+
handler it was given, so an `http://` target that lands on a self-signed or
|
|
1106
|
+
legacy `https://` one needs it as much as a direct HTTPS target does. The
|
|
1107
|
+
notice fires once, when a connection has actually reached TLS, and lists only
|
|
1108
|
+
the rungs that took.
|
|
1109
|
+
|
|
1110
|
+
- **A benchmark case may set `timeout` to zero.** `timeout or 900.0` replaced
|
|
1111
|
+
an explicit `0` with the fifteen-minute default, so a case deliberately
|
|
1112
|
+
bounded to no time at all ran for much longer than it asked for.
|
|
1113
|
+
|
|
1114
|
+
## [2.35.4] — 2026-09-19
|
|
1115
|
+
|
|
1116
|
+
### Fixed
|
|
1117
|
+
|
|
1118
|
+
- **A wave the declared profile emptied is no longer read as a finished
|
|
1119
|
+
method.** The probe filter drops probes between the method and the wire, and
|
|
1120
|
+
when it emptied an adaptive method's *first* screening wave the engine took
|
|
1121
|
+
the empty batch for "nothing left to send" and stopped before asking for the
|
|
1122
|
+
next wave at all.
|
|
1123
|
+
|
|
1124
|
+
An adaptive method holds separators back in waves precisely because a filter
|
|
1125
|
+
is expected: `time` screens two of them first and keeps `||`, `&&`, the
|
|
1126
|
+
newline and the bare command for a second wave. A sink that strips `;` and
|
|
1127
|
+
`|` removes exactly the first wave and leaves the rest intact -- and the rest
|
|
1128
|
+
were never sent. Measured against a sink reachable only through `&&`, the run
|
|
1129
|
+
reported **`negative`**: "the probes reached the target and found nothing",
|
|
1130
|
+
about probes that were never sent. Worse than `nothing-tested`, which is at
|
|
1131
|
+
least true.
|
|
1132
|
+
|
|
1133
|
+
The loop now tracks what the method offered separately from what the profile
|
|
1134
|
+
allows to be sent, and ends only when the method itself is done. The round cap
|
|
1135
|
+
still bounds it, so a method whose every wave is filtered cannot spin the
|
|
1136
|
+
engine.
|
|
1137
|
+
|
|
1138
|
+
- **The unit suite runs off Linux.** The fake vulnerable sinks in the tests are
|
|
1139
|
+
POSIX command-injection points and every probe built for them is POSIX, but
|
|
1140
|
+
they were executed through `os.popen` — which is `cmd.exe` on Windows. The
|
|
1141
|
+
sink the test says executes did not execute, so the oracle correctly reported
|
|
1142
|
+
no execution and five tests failed for a reason unrelated to the code under
|
|
1143
|
+
test. Worse, the tests asserting a *negative* stayed green throughout: a
|
|
1144
|
+
broken fixture that keeps its controls passing is the failure this project
|
|
1145
|
+
takes seriously everywhere else. The sinks now name a POSIX `sh` explicitly;
|
|
1146
|
+
on Linux and macOS that is the `/bin/sh` they always used.
|
|
1147
|
+
|
|
1148
|
+
- **A deliberately dead target no longer costs minutes.** A closed loopback port
|
|
1149
|
+
answers with a RST on Linux and is silently dropped on Windows, where each
|
|
1150
|
+
probe waits out the SYN retry instead — measured at ~2s per probe. The
|
|
1151
|
+
benchmark harness's own unreachable-target case paid that for the full ladder
|
|
1152
|
+
twice, once for the vulnerable half and once for the control: 1800s for one
|
|
1153
|
+
test, longer than the other 534 together. Bounded to three probes it measures
|
|
1154
|
+
9.1s, and three probes prove "nothing reached the target" exactly as well as
|
|
1155
|
+
forty do.
|
|
1156
|
+
|
|
1157
|
+
- **A captured-request fixture reaches disk byte for byte.** The cleartext-capture
|
|
1158
|
+
tests write a raw HTTP request whose text already spells its own CRLF line
|
|
1159
|
+
endings, through `Path.write_text` — which on Windows translates the newline of
|
|
1160
|
+
each one again. The file on disk held a doubled carriage return, the parser
|
|
1161
|
+
found no headers, and two tests failed against a fixture that had stopped being
|
|
1162
|
+
an HTTP request at all.
|
|
1163
|
+
|
|
1164
|
+
The third test in that class passed throughout, for the wrong reason: it
|
|
1165
|
+
asserts that a notice is *absent*, and a request that cannot be built prints no
|
|
1166
|
+
notice either. A control that stays green while its fixture rots is precisely
|
|
1167
|
+
what this project refuses to accept from a benchmark case, so the fixture now
|
|
1168
|
+
has a guard of its own. The capture is written as bytes; the `newline` argument
|
|
1169
|
+
that would say the same thing arrived in Python 3.10, and this project
|
|
1170
|
+
supports 3.8.
|
|
1171
|
+
|
|
1172
|
+
## [2.35.3] — 2026-09-18
|
|
1173
|
+
|
|
1174
|
+
The target profile an operator declares now reaches the probe ladder, not just
|
|
1175
|
+
the corpus.
|
|
1176
|
+
|
|
1177
|
+
### Fixed
|
|
1178
|
+
|
|
1179
|
+
- **`--deny-chars` / `--max-length` reach the detection probes.** They were
|
|
1180
|
+
applied by `_filter_by_profile`, which drops *corpus records* — and stopped
|
|
1181
|
+
there. The probes a detection method builds from those records went out
|
|
1182
|
+
regardless, so a run that had been told "this target strips quotes" still paid
|
|
1183
|
+
for every quote-carrying rung of the ladder, on requests structurally unable
|
|
1184
|
+
to confirm. Those requests are not free: on a captured request with `-p all`
|
|
1185
|
+
they are the budget the next injection point never got. The filter now sits at
|
|
1186
|
+
the engine, where every probe passes through it — deliberately not inside
|
|
1187
|
+
`_wrap_variants`, because `_space_free_probes`, the query-language bridges,
|
|
1188
|
+
`eval`, `oob` and `deser` each build payloads without going through that
|
|
1189
|
+
helper, and a gate that reaches some methods and not others is the side path
|
|
1190
|
+
that once left `file`/`time`/`oob` unable to send the raw rung.
|
|
1191
|
+
|
|
1192
|
+
Denying a character narrows the ladder rather than emptying it: a target that
|
|
1193
|
+
strips `;` is still probed through `|`, `||`, `&&` and the newline, which is
|
|
1194
|
+
what the separator table has always been for.
|
|
1195
|
+
|
|
1196
|
+
Checked on the literal payload, before the delivery layer percent-encodes it
|
|
1197
|
+
for its injection point. That is stricter than the corpus check, which is
|
|
1198
|
+
applied to the encoded payload and so lets a URL-encoded quote through a quote
|
|
1199
|
+
filter. The layers genuinely differ: transport encoding is undone by the
|
|
1200
|
+
server before the value reaches the sink, so a percent-encoded quote is still
|
|
1201
|
+
a quote when the application's own filter sees it.
|
|
1202
|
+
|
|
1203
|
+
- **A profile strict enough to remove every probe reports `nothing-tested`.**
|
|
1204
|
+
Not `negative`, which would read as "not vulnerable" from a run that sent
|
|
1205
|
+
nothing. The message names the profile as the cause and the characters a probe
|
|
1206
|
+
would have to avoid, instead of the generic advice to widen `--environments` —
|
|
1207
|
+
which is not what emptied the run.
|
|
1208
|
+
|
|
1209
|
+
- **The cost estimate follows the profile.** `[detect] cost:` builds the probes
|
|
1210
|
+
and counts them, so it now counts the ones that will actually be sent. An
|
|
1211
|
+
estimate that ignores a filter is wrong precisely for the operator who
|
|
1212
|
+
narrowed the run on purpose.
|
|
1213
|
+
|
|
1214
|
+
### Changed
|
|
1215
|
+
|
|
1216
|
+
- A run that dropped probes says so, with the reason and a count per reason. A
|
|
1217
|
+
ladder that shrinks quietly is the one way this filter could manufacture a
|
|
1218
|
+
false negative, so the removal is stated rather than left to be inferred from
|
|
1219
|
+
the traffic.
|
|
1220
|
+
|
|
1221
|
+
## [2.35.2] — 2026-09-18
|
|
1222
|
+
|
|
1223
|
+
### Fixed
|
|
1224
|
+
|
|
1225
|
+
- **`--insecure` now reaches a legacy TLS stack, not just an untrusted one.**
|
|
1226
|
+
Turning certificate verification off is not the same as completing a
|
|
1227
|
+
handshake. OpenSSL 3.x ships security level 2, which refuses the key sizes and
|
|
1228
|
+
signature algorithms that software of the era this tool gets pointed at still
|
|
1229
|
+
offers — Webmin 1.910, the build the README's `reflected` row rests on,
|
|
1230
|
+
answers a default client with `SSLV3_ALERT_HANDSHAKE_FAILURE` and nothing
|
|
1231
|
+
else. Every probe then came back `error`: correct, and useless. The run was
|
|
1232
|
+
honest about having measured nothing, and the sink behind that handshake was
|
|
1233
|
+
never tested at all. Found by running the coverage benchmark, which failed at
|
|
1234
|
+
its readiness gate against a container that was up and answering.
|
|
1235
|
+
|
|
1236
|
+
`--insecure` now also lowers the security level and the minimum protocol
|
|
1237
|
+
version. This does not widen exposure: with `check_hostname = False` and
|
|
1238
|
+
`CERT_NONE` the connection is already unauthenticated, so an active attacker
|
|
1239
|
+
is already unconstrained — accepting a 1024-bit key or a SHA-1 signature on
|
|
1240
|
+
top of that gives away nothing that was still being held. What it buys is the
|
|
1241
|
+
difference between testing the target and reporting that it could not be
|
|
1242
|
+
reached. A run without the flag is untouched and still verifies certificates.
|
|
1243
|
+
|
|
1244
|
+
- **The benchmark's readiness gate is as permissive as the tool it gates.**
|
|
1245
|
+
`wait_for_target` built its own strict context, so a case against deliberately
|
|
1246
|
+
old software reported "target never became ready" about a container that was
|
|
1247
|
+
up — a case failure with nothing wrong in it.
|
|
1248
|
+
|
|
1249
|
+
### Changed
|
|
1250
|
+
|
|
1251
|
+
- A run that passes `--insecure` states the full extent of the downgrade on its
|
|
1252
|
+
first line. The flag gives up more than certificate identity now, and an
|
|
1253
|
+
operator on a monitored engagement should read that in the transcript rather
|
|
1254
|
+
than infer it from the help text.
|
|
1255
|
+
|
|
1256
|
+
## [2.35.1] — 2026-08-21
|
|
1257
|
+
|
|
1258
|
+
A robustness pass over error handling: no new capability, four ways the tool
|
|
1259
|
+
could crash or mislead on input it did not choose.
|
|
1260
|
+
|
|
1261
|
+
### Fixed
|
|
1262
|
+
|
|
1263
|
+
- **A truncated error response no longer ends the run.** Reading an
|
|
1264
|
+
`HTTPError`'s body happens *inside* the `except` handler, where the sibling
|
|
1265
|
+
`except Exception` cannot reach it — so a target that promised a
|
|
1266
|
+
`Content-Length` it never delivered raised `ConnectionResetError` straight out
|
|
1267
|
+
of `main()`, taking every probe already fired with it. The read is now
|
|
1268
|
+
guarded: the status still comes back, an unreadable body is reported empty.
|
|
1269
|
+
A body that *does* arrive is still returned in full — the 500-stack-trace
|
|
1270
|
+
confirmations that branch exists for are unaffected.
|
|
1271
|
+
- **A `--target-profile` is checked before it is used.** A profile is written by
|
|
1272
|
+
hand, so a typo in one is ordinary; it surfaced as a traceback. A top level
|
|
1273
|
+
that is not a JSON object, `deny_chars` that is not text, a `max_length` that
|
|
1274
|
+
is not a number, a selector field that is not a list of names — each is now an
|
|
1275
|
+
operator-readable `[!]` message and exit 1, the same way the sink-shape fields
|
|
1276
|
+
already behaved. Twelve inputs that produced an `AttributeError` or a
|
|
1277
|
+
`TypeError` now produce a sentence.
|
|
1278
|
+
- **A selector field given as one string means one name.** `"environments":
|
|
1279
|
+
"unix"` in a profile was iterated character by character, matched nothing, and
|
|
1280
|
+
the empty run that followed was reported as a success. A string is now split
|
|
1281
|
+
on commas: `"unix"` is `["unix"]`, `"raw, html"` is `["raw", "html"]`.
|
|
1282
|
+
- **Unknown `--environments` and `--encodings` are named.** Both were silent, so
|
|
1283
|
+
`--environments linux` — the corpus calls it `unix` — produced an empty file
|
|
1284
|
+
and exit 0, indistinguishable from a target with no payloads for it. Both now
|
|
1285
|
+
warn and list the names that exist, as unknown contexts and categories already
|
|
1286
|
+
did. An empty result is also no longer announced as "Successfully generated 0
|
|
1287
|
+
payloads"; it says the selection matched nothing and points at the filters.
|
|
1288
|
+
- **A callback cannot rewrite the operator's terminal.** The host and path of an
|
|
1289
|
+
OOB callback are chosen by the target. Printed raw, an ESC byte let that
|
|
1290
|
+
target colour, erase and rewrite lines — hiding a genuine `[HIT]` behind
|
|
1291
|
+
`\x1b[2K\r`, or forging one that never arrived. Control characters are now
|
|
1292
|
+
escaped for display as `\xNN`; the recorded hit and the `--listen-log` JSONL
|
|
1293
|
+
keep the bytes verbatim.
|
|
1294
|
+
|
|
1295
|
+
## [2.35.0] — 2026-08-20
|
|
1296
|
+
|
|
1297
|
+
### Added
|
|
1298
|
+
|
|
1299
|
+
- **RCEKit is installable from PyPI**: `pipx install rcekit` (or
|
|
1300
|
+
`pip install rcekit`) puts an `rcekit` command on PATH. Published through PyPI
|
|
1301
|
+
Trusted Publishing from a GitHub release — no API token, no repository secret.
|
|
1302
|
+
|
|
1303
|
+
It ships as a **single-module distribution**, not a package tree. `rcekit.py`
|
|
1304
|
+
stays one file at the repo root and still runs alone from a `curl` on a jump
|
|
1305
|
+
box or an air-gapped host; installing is a second supported shape, not a
|
|
1306
|
+
replacement for the first.
|
|
1307
|
+
|
|
1308
|
+
### Changed
|
|
1309
|
+
|
|
1310
|
+
- **`import rcekit` no longer has side effects.** Logging was configured at
|
|
1311
|
+
module scope, and `logging.FileHandler` opens its file when it is constructed,
|
|
1312
|
+
so merely importing the module wrote `rcekit.log` into whatever directory the
|
|
1313
|
+
interpreter happened to be in. Handler setup moved into `configure_logging()`,
|
|
1314
|
+
called from `main()`. Running the CLI still writes `rcekit.log` exactly as
|
|
1315
|
+
before.
|
|
1316
|
+
|
|
1317
|
+
- **`main()` takes an optional `argv`** and returns an explicit `int`, so the
|
|
1318
|
+
console-script entry point is a plain zero-argument call and tests can drive
|
|
1319
|
+
the CLI in-process. Every exit code is unchanged.
|
|
1320
|
+
|
|
1321
|
+
- **The built-in corpus is no longer reported as a missing file.** With nothing
|
|
1322
|
+
but `rcekit.py` — an installed wheel, or the single-file copy — there is no
|
|
1323
|
+
`templates/` directory, the embedded corpus *is* the corpus, and the run is
|
|
1324
|
+
now silent about it; `--doctor` names it `built-in (embedded in rcekit.py)`
|
|
1325
|
+
and reports OK. A `templates/` directory that exists *without* its
|
|
1326
|
+
`payloads.json` still prints the notice, because that one is a real finding.
|
|
1327
|
+
A corpus that is present but corrupt, and an explicit `--template-file` that
|
|
1328
|
+
is missing or corrupt, still refuse to run and exit non-zero.
|
|
1329
|
+
|
|
1330
|
+
- The log file handler now runs at `DEBUG` while the console stays at `INFO`, so
|
|
1331
|
+
detail worth having when reconstructing a run no longer lands on the
|
|
1332
|
+
operator's terminal.
|
|
1333
|
+
|
|
1334
|
+
## [2.34.1] — 2026-08-17
|
|
1335
|
+
|
|
1336
|
+
Documentation only; no behaviour change.
|
|
1337
|
+
|
|
1338
|
+
### Changed
|
|
1339
|
+
|
|
1340
|
+
- **New tagline: "`confirmed` means the target executed the input. `negative`
|
|
1341
|
+
means the probes reached it."** The old one — "prove RCE, don't guess it" —
|
|
1342
|
+
claimed the tool always proves. It does not, and does not need to: the value
|
|
1343
|
+
is that each verdict has a mechanical meaning, in both directions. A promise
|
|
1344
|
+
can be broken; a definition cannot.
|
|
1345
|
+
|
|
1346
|
+
Both halves name their actor and object on purpose. "Executed" alone reads as
|
|
1347
|
+
though *RCEKit* executed something; the claim is about the target. And
|
|
1348
|
+
`confirmed`/`negative` are the verdict values as the code spells them, not
|
|
1349
|
+
looser words like "clean".
|
|
1350
|
+
|
|
1351
|
+
- **The README caught up with the engine.** It still described "two verdict
|
|
1352
|
+
tiers" when there are seven, and the comparison table predated the last five
|
|
1353
|
+
releases. Rewritten around what a verdict asserts, with the tier table as the
|
|
1354
|
+
centrepiece and a section on the half no other tool has: `error` and
|
|
1355
|
+
`nothing-tested` exist so a run that tested nothing is never reported as
|
|
1356
|
+
clean.
|
|
1357
|
+
|
|
1358
|
+
The comparison table gains the classes added since it was written — Windows
|
|
1359
|
+
`cmd.exe`/PowerShell sinks, upload → write-then-execute, second-order
|
|
1360
|
+
execution, query-language bridges and deserialization sinks — and the "reach
|
|
1361
|
+
for something else" note now says plainly that sqlmap owns the database and
|
|
1362
|
+
RCEKit's bridges only prove the OS is reachable from a text parameter.
|
|
1363
|
+
|
|
1364
|
+
### Fixed
|
|
1365
|
+
|
|
1366
|
+
- Two claims in the engagement-controls table were wrong and are now accurate:
|
|
1367
|
+
the observed-channel fetch sends **no** credentials unless given a request with
|
|
1368
|
+
`--observe-request` (only the `file` read-back inherits them, same-origin), and
|
|
1369
|
+
an unanswered `--observe-url` is a warning about a partly blinded run rather
|
|
1370
|
+
than a `nothing-tested` verdict.
|
|
1371
|
+
- The "mechanisms that produce `inconclusive`" list said four and listed five.
|
|
1372
|
+
|
|
1373
|
+
## [2.34.0] — 2026-08-17
|
|
1374
|
+
|
|
1375
|
+
Deserialization **sink** detection, and a verdict that is deliberately not RCE.
|
|
1376
|
+
|
|
1377
|
+
Deserialization RCE (fastjson, shiro, weblogic, jenkins) cannot be confirmed by
|
|
1378
|
+
the value-oracle model: the payload is a serialized object graph and gadget
|
|
1379
|
+
selection is classpath-specific, so whether execution is reachable depends on
|
|
1380
|
+
jars RCEKit cannot see. That stays out of scope. The honest middle step is
|
|
1381
|
+
showing the endpoint parses the data at all — a real finding, and the
|
|
1382
|
+
prerequisite for every gadget chain.
|
|
1383
|
+
|
|
1384
|
+
### Added
|
|
1385
|
+
|
|
1386
|
+
- **`--methods deser`**, which **never emits `confirmed`**. Its strongest
|
|
1387
|
+
outcome is a new verdict, `deserialization-sink`, reported in its own section
|
|
1388
|
+
that states outright that reaching RCE from there depends on classpath
|
|
1389
|
+
gadgets. Collapsing it into `confirmed` would break the one guarantee the tool
|
|
1390
|
+
rests on; collapsing it into `needs-review` would throw away a proven finding.
|
|
1391
|
+
|
|
1392
|
+
`deserialization-sink` sits below both RCE tiers in the collapsed verdict: it
|
|
1393
|
+
is proven, but a *suspected* RCE outranks a proven non-RCE in triage.
|
|
1394
|
+
|
|
1395
|
+
- **Two oracles of deliberately different strength.** `shape` (no listener
|
|
1396
|
+
needed) sends a well-formed object stream, the same stream truncated, and the
|
|
1397
|
+
format's magic bytes plus random noise of the same length, and asks whether
|
|
1398
|
+
the endpoint answers the well-formed one differently from both — a
|
|
1399
|
+
fingerprint, so `needs-review` only, never promoted. `dns` (needs
|
|
1400
|
+
`--oob-host`) sends a gadget whose only side effect is a name lookup.
|
|
1401
|
+
|
|
1402
|
+
- **A URLDNS builder for Java serialization.** A `HashMap` holding one
|
|
1403
|
+
`java.net.URL`: `HashMap.readObject` hashes the key, `URL.hashCode` asks for
|
|
1404
|
+
the host address, the JVM resolves the name. It references no class outside
|
|
1405
|
+
`java.util`/`java.net`, so there is nothing in it that can run — the callback
|
|
1406
|
+
proves the object graph was reconstructed and no more.
|
|
1407
|
+
|
|
1408
|
+
Built in Python rather than declared in the corpus because the URL host is
|
|
1409
|
+
length-prefixed *inside* the stream and changes per probe. Its constant parts
|
|
1410
|
+
are the exact bytes OpenJDK's own `ObjectOutputStream` emits for that graph,
|
|
1411
|
+
and the result was **verified against OpenJDK 21**: it deserializes to
|
|
1412
|
+
`HashMap{http://<host>/=rk}` and issues a DNS query for `<host>`, with no code
|
|
1413
|
+
execution.
|
|
1414
|
+
|
|
1415
|
+
- **A `deser_probes` corpus section** with `java`, `php`, `dotnet`,
|
|
1416
|
+
`python_pickle` and `fastjson`, plus `--deser-formats` to narrow it. Only
|
|
1417
|
+
`java` and `fastjson` carry a DNS gadget: PHP and .NET chains all run through
|
|
1418
|
+
magic methods or type confusion, so there is no honest DNS-only probe for them
|
|
1419
|
+
and they get the shape oracle alone.
|
|
1420
|
+
|
|
1421
|
+
- Response signatures for the shape differential drop long digit and hex runs,
|
|
1422
|
+
so request ids and timestamps on an otherwise identical error page do not make
|
|
1423
|
+
every endpoint fingerprint as a parser.
|
|
1424
|
+
|
|
1425
|
+
### Changed
|
|
1426
|
+
|
|
1427
|
+
- The README's scope note now says precisely what changed and what did not:
|
|
1428
|
+
deserialization **gadget chains** remain out of scope, while the **sink** is
|
|
1429
|
+
now reported in its own tier.
|
|
1430
|
+
|
|
1431
|
+
## [2.33.0] — 2026-08-17
|
|
1432
|
+
|
|
1433
|
+
Query-language bridges. Several RCEs pass through a query language before
|
|
1434
|
+
reaching the OS — Postgres `COPY … FROM PROGRAM`, MSSQL `xp_cmdshell`, XXE
|
|
1435
|
+
`expect://` — and the injection point is an ordinary text value, so the oracle
|
|
1436
|
+
model already fitted. Only the carriers were missing.
|
|
1437
|
+
|
|
1438
|
+
### Added
|
|
1439
|
+
|
|
1440
|
+
- **A `bridges` section in the corpus**, declared like `eval_carriers` so
|
|
1441
|
+
coverage grows without touching Python. Each bridge names the shell it
|
|
1442
|
+
reaches, its safety tier, its prerequisites and, where it creates something,
|
|
1443
|
+
the statement that removes it: `postgres_copy_program` (`/bin/sh`,
|
|
1444
|
+
`stateful`), `mssql_xp_cmdshell` (`cmd.exe`, `intrusive`) and `xxe_expect`
|
|
1445
|
+
(`/bin/sh`, `intrusive`).
|
|
1446
|
+
|
|
1447
|
+
- **`--bridges none|auto|NAMES`** rides the command probes through them. A
|
|
1448
|
+
bridge is a **carrier, not an oracle**: it wraps the command `reflected`,
|
|
1449
|
+
`time` and `oob` already build, so those methods prove execution through it
|
|
1450
|
+
and inherit every tier guarantee rather than re-deriving one. Off by default,
|
|
1451
|
+
because a bridge payload is SQL or XML syntax and on an ordinary shell sink it
|
|
1452
|
+
is a request that cannot confirm.
|
|
1453
|
+
|
|
1454
|
+
Three properties follow from that framing. A bridge only gets a core written
|
|
1455
|
+
in its own dialect — `xp_cmdshell` hands its argument to `cmd.exe`, so pairing
|
|
1456
|
+
it with a POSIX `$((a+b))` would send inert text. No separator is prepended:
|
|
1457
|
+
inside `COPY … FROM PROGRAM '…'` there is no running command to break out of.
|
|
1458
|
+
And the record's context still applies, so `--contexts sql` and a bridge
|
|
1459
|
+
compose instead of each reinventing the other.
|
|
1460
|
+
|
|
1461
|
+
- **The safety ordering governs bridges** exactly as it governs every corpus
|
|
1462
|
+
payload: a `stateful` bridge needs `--verify-active-risk stateful`, and the
|
|
1463
|
+
pre-flight names the tier each held-back bridge actually requires rather than
|
|
1464
|
+
sending the operator to raise the ceiling further than the run needs.
|
|
1465
|
+
|
|
1466
|
+
### Changed
|
|
1467
|
+
|
|
1468
|
+
- **An aggregate method's result now carries its cleanup line.** `time` reports
|
|
1469
|
+
one row for a whole probe series, so a stateful bridge on the one oracle that
|
|
1470
|
+
reliably proves a query-language sink was the one that never said how to clean
|
|
1471
|
+
up after itself.
|
|
1472
|
+
|
|
1473
|
+
### Not built, deliberately
|
|
1474
|
+
|
|
1475
|
+
- **MySQL UDF execution** is a multi-stage chain — write a shared object into the
|
|
1476
|
+
plugin directory, then `CREATE FUNCTION` — not something a single probe can
|
|
1477
|
+
carry. There is no stub for it.
|
|
1478
|
+
- **MongoDB `$where`** is a boolean-only channel (its JS sandbox cannot reach a
|
|
1479
|
+
shell), so it needs a different oracle rather than this one.
|
|
1480
|
+
`mongo-express/CVE-2019-10758` is a plain JS `eval` sink that `--methods eval`
|
|
1481
|
+
already covers.
|
|
1482
|
+
|
|
1483
|
+
The three shipped bridges are documented syntax but **not validated against live
|
|
1484
|
+
databases here** — this build environment has no container runtime. Each corpus
|
|
1485
|
+
entry says so in its `verified` field rather than implying a test that did not
|
|
1486
|
+
happen.
|
|
1487
|
+
|
|
1488
|
+
## [2.32.0] — 2026-08-17
|
|
1489
|
+
|
|
1490
|
+
The second-order oracle. Execution frequently happens on a **different request**
|
|
1491
|
+
than injection — stored SSTI rendered on a profile page, a payload written to a
|
|
1492
|
+
log a template engine later renders, a queued job run asynchronously. The engine
|
|
1493
|
+
diffs the response it injected into, so every one of those read `negative`
|
|
1494
|
+
however exploitable the target was.
|
|
1495
|
+
|
|
1496
|
+
### Added
|
|
1497
|
+
|
|
1498
|
+
- **`--observe-url URL`** names the endpoint where the execution surfaces. It is
|
|
1499
|
+
read after each probe and then polled after the batch, and a probe whose
|
|
1500
|
+
computed value turns up there is upgraded to `confirmed`.
|
|
1501
|
+
|
|
1502
|
+
It stays fully differential, which is why it reaches `confirmed` rather than
|
|
1503
|
+
`needs-review`: the value was computed locally from operands random to that
|
|
1504
|
+
probe, it must be absent from a snapshot of the endpoint taken **before any
|
|
1505
|
+
probe was sent**, and — the rule that carries the weight — a probe's value is
|
|
1506
|
+
looked for there **only when the probe's own payload does not contain it**.
|
|
1507
|
+
|
|
1508
|
+
Without that last rule the oracle would be a false-positive generator: `file`
|
|
1509
|
+
and `oob` expect a random token that sits verbatim in the payload, so a target
|
|
1510
|
+
that merely stores the payload and renders it back would hand that token
|
|
1511
|
+
straight to the observed page and every such probe would confirm without
|
|
1512
|
+
executing anything. Measured against a store-and-echo target: **0**
|
|
1513
|
+
confirmations. The computed-value methods pass the same rule for the opposite
|
|
1514
|
+
reason — reflection returns `$((a+b))`, never the sum — so it selects them
|
|
1515
|
+
without naming them, and a method added later inherits the right answer.
|
|
1516
|
+
|
|
1517
|
+
- **`--observe-request FILE`** takes a captured request instead, for the common
|
|
1518
|
+
case where the page a stored payload renders on is behind a login. It needs no
|
|
1519
|
+
`FUZZ` marker: the observed endpoint is read, never injected into.
|
|
1520
|
+
|
|
1521
|
+
- **`--observe-poll` / `--observe-timeout`** control the polling window
|
|
1522
|
+
(defaults 5s and 60s). One poll always happens, even at a zero timeout.
|
|
1523
|
+
|
|
1524
|
+
- Every probe result carries an `observe_status` in `--detect-json`:
|
|
1525
|
+
`confirmed`, `polled` (read, value not there), `in-control`, `not-observed`
|
|
1526
|
+
(not eligible) or `unreachable`. When the endpoint never answered, the run
|
|
1527
|
+
says so outright — negatives decided without ever reading the observed channel
|
|
1528
|
+
are not second-order negatives.
|
|
1529
|
+
|
|
1530
|
+
### Changed
|
|
1531
|
+
|
|
1532
|
+
- The observed channel is read once after **each** probe as well as polled after
|
|
1533
|
+
the batch, so a run with `--observe-url` sends roughly twice the requests.
|
|
1534
|
+
Batch-then-poll alone is only correct for a channel that *accumulates* (a log,
|
|
1535
|
+
a comment list); where the store overwrites — a profile field, which is the
|
|
1536
|
+
shape this oracle most exists for — every probe but the last is gone by the
|
|
1537
|
+
time the batch poll runs, and the oracle confirmed nothing. The extra read is
|
|
1538
|
+
skipped for probes that are already confirmed in-band or not eligible, so
|
|
1539
|
+
`file` and `oob` add none.
|
|
1540
|
+
|
|
1541
|
+
Observing is additive throughout: the in-band verdict is computed exactly as
|
|
1542
|
+
before and only a non-`confirmed` one can be upgraded, so a run without the flag
|
|
1543
|
+
is byte-for-byte unchanged and a run with it can only gain findings.
|
|
1544
|
+
|
|
1545
|
+
## [2.31.0] — 2026-08-17
|
|
1546
|
+
|
|
1547
|
+
The `write` method: a write primitive proven to be RCE by executing what it
|
|
1548
|
+
wrote. A whole family of targets was invisible — `tomcat/CVE-2017-12615` (PUT a
|
|
1549
|
+
JSP), `activemq/CVE-2016-3088`, `weblogic/CVE-2018-2894` — because the vulnerable
|
|
1550
|
+
request *stores a file* rather than evaluating anything. Nothing is computed in
|
|
1551
|
+
its response, so `reflected` and `eval` correctly returned `negative` on targets
|
|
1552
|
+
that are fully exploitable.
|
|
1553
|
+
|
|
1554
|
+
### Added
|
|
1555
|
+
|
|
1556
|
+
- **`--methods write`** — the inverse of `file`. `file` assumes execution exists
|
|
1557
|
+
and uses a write as proof of it; `write` assumes a write primitive exists and
|
|
1558
|
+
uses execution of the written file as proof of RCE. The probe is the file's
|
|
1559
|
+
*content*: a one-liner computing a product on random operands, delivered
|
|
1560
|
+
through the ordinary injection point.
|
|
1561
|
+
|
|
1562
|
+
The fetched file is read in three tiers, and the middle one is the reason the
|
|
1563
|
+
method exists:
|
|
1564
|
+
|
|
1565
|
+
| fetched file contains | verdict | means |
|
|
1566
|
+
|---|---|---|
|
|
1567
|
+
| the product | `confirmed` | written **and** executed |
|
|
1568
|
+
| the one-liner, verbatim | `needs-review` | arbitrary file write, not interpreted |
|
|
1569
|
+
| neither | `negative` | no write, or not served there |
|
|
1570
|
+
|
|
1571
|
+
An upload directory that is served but not interpreted is a real finding and
|
|
1572
|
+
is not remote code execution, so the tiers are never merged.
|
|
1573
|
+
|
|
1574
|
+
- **`--write-url-template URL`** names where the stored file is served — the
|
|
1575
|
+
channel the proof comes back on, and the flag the method is gated on.
|
|
1576
|
+
|
|
1577
|
+
- **`--write-lang`** picks the file types: `auto` (default) reads the extension
|
|
1578
|
+
off the read-back URL, or name any of `jsp`, `jspx`, `php`, `aspx`, `erb`.
|
|
1579
|
+
`jsp`/`aspx`/`erb` share the `<%= %>` delimiters, so their probes are
|
|
1580
|
+
byte-identical and cost one request between them; with no extension to read,
|
|
1581
|
+
`auto` writes all five in three requests.
|
|
1582
|
+
|
|
1583
|
+
### Changed
|
|
1584
|
+
|
|
1585
|
+
- **A `needs-review` finding now prints its cleanup line too.** It used to
|
|
1586
|
+
appear only under `confirmed`, which was already thin and is wrong for this
|
|
1587
|
+
method: a `write` reaching `needs-review` means the file *is* on the target,
|
|
1588
|
+
just not interpreted, so the artifact would have been left there unmentioned.
|
|
1589
|
+
|
|
1590
|
+
- The write method's operands are drawn once per run rather than once per
|
|
1591
|
+
carrier, so the file is written once instead of once for each of the ~13
|
|
1592
|
+
`(environment, context)` carriers. For a state-changing method that is not a
|
|
1593
|
+
request-count saving, it is a blast radius. Still fresh per run, which is what
|
|
1594
|
+
makes the product unforgeable.
|
|
1595
|
+
|
|
1596
|
+
- The write method declines the break-out contexts (`sql`, `javascript`,
|
|
1597
|
+
`shell_*`, …) and keeps the transport ones. Its payload is a whole file body:
|
|
1598
|
+
there is nothing to break out of, and wrapping it in `'; … -- ` would write a
|
|
1599
|
+
broken file. A run narrowed past `raw` and the transport contexts is told so
|
|
1600
|
+
rather than reporting a clean negative.
|
|
1601
|
+
|
|
1602
|
+
## [2.30.0] — 2026-08-17
|
|
1603
|
+
|
|
1604
|
+
Per-dialect shell probes. `$((a+b))`, `sleep` and `$(echo TAG)` are POSIX
|
|
1605
|
+
constructs: on a cmd.exe or PowerShell sink they are inert literal text. The
|
|
1606
|
+
dialect was inferred from the corpus environment alone, so a run could send a
|
|
1607
|
+
probe no shell on the target would ever execute — including on the carrier whose
|
|
1608
|
+
context is literally named `powershell`.
|
|
1609
|
+
|
|
1610
|
+
### Added
|
|
1611
|
+
|
|
1612
|
+
- **`--sink-env auto|unix|windows|powershell`** states which shell runs the
|
|
1613
|
+
injected command. The computed-value core, the separators and the break-out
|
|
1614
|
+
contexts are all chosen from it. `auto` (the default) infers it per carrier;
|
|
1615
|
+
pin it when the corpus environment names the *application runtime* rather than
|
|
1616
|
+
the OS — `--environments php --sink-env windows` is a PHP application on IIS,
|
|
1617
|
+
which no inference can see.
|
|
1618
|
+
|
|
1619
|
+
- **A PowerShell probe shape for every shell method**, validated against
|
|
1620
|
+
pwsh 7.4: `Write-Output T1$(a*b)T2` for `reflected` (an unquoted argument is an
|
|
1621
|
+
expandable string, so the core carries no quote and the quote-wrapping
|
|
1622
|
+
contexts can still carry it), `Start-Sleep -Milliseconds N` for `time`,
|
|
1623
|
+
`Set-Content` for `file` and `iwr -useb` for `oob`. PowerShell was previously
|
|
1624
|
+
reachable by no probe in any method.
|
|
1625
|
+
|
|
1626
|
+
- **cmd.exe and PowerShell carriers for the `dotnet` environment.** It is the
|
|
1627
|
+
one corpus environment that names a platform, and it was taking the POSIX
|
|
1628
|
+
shape — so .NET on Windows, the case the environment exists for, was the case
|
|
1629
|
+
it could not confirm on. Every other runtime keeps the POSIX shape: a language
|
|
1630
|
+
does not say which OS it runs on.
|
|
1631
|
+
|
|
1632
|
+
### Fixed
|
|
1633
|
+
|
|
1634
|
+
- **The `powershell` carrier was written in cmd.exe.** Every `windows` carrier
|
|
1635
|
+
took the `for /f ... ('set /a a+b')` core regardless of context, so the one
|
|
1636
|
+
carrier explicitly shaped for PowerShell sent a payload PowerShell cannot
|
|
1637
|
+
execute. The dialect now follows the carrier's context first, and a carrier's
|
|
1638
|
+
break-out variants stay in its dialect rather than re-deriving from the
|
|
1639
|
+
environment.
|
|
1640
|
+
|
|
1641
|
+
- **`Set-Content`, not `>`, for the PowerShell write.** In Windows PowerShell
|
|
1642
|
+
5.1 the redirect is `Out-File`, whose default encoding is UTF-16LE: the write
|
|
1643
|
+
lands and the read-back still does not find the token, so the probe reports
|
|
1644
|
+
negative on a target it owns.
|
|
1645
|
+
|
|
1646
|
+
### Changed
|
|
1647
|
+
|
|
1648
|
+
- **cmd.exe no longer gets the `sq`, `dq` and `subshell` carriers.** It has
|
|
1649
|
+
neither a comment character to swallow the sink's tail nor a
|
|
1650
|
+
command-substitution syntax, so those four carriers per Windows run were
|
|
1651
|
+
requests that could only come back negative. PowerShell takes the quote
|
|
1652
|
+
break-outs and `$( )` — both measured — but not the backtick, which is its
|
|
1653
|
+
escape character rather than a substitution.
|
|
1654
|
+
|
|
1655
|
+
- **PowerShell's separator sweep carries no pipe.** `cmd | Start-Sleep
|
|
1656
|
+
-Milliseconds 500` is a parameter-binding error, not a fresh command with
|
|
1657
|
+
stdin attached the way a POSIX pipe is, and it fails that way for every cmdlet
|
|
1658
|
+
the probes use. `;`, a newline and (on PowerShell 7) `&&`/`||` remain.
|
|
1659
|
+
|
|
1660
|
+
- The pre-flight plan prints the sink shell alongside the sink shapes, and a
|
|
1661
|
+
pinned dialect narrows the printed ladder to the rungs it has syntax for.
|
|
1662
|
+
|
|
1663
|
+
## [2.29.0] — 2026-08-17
|
|
1664
|
+
|
|
1665
|
+
Generalised read-back for the `file` method. It required a writable **web root**
|
|
1666
|
+
the tester already knew, which ruled out every other way a target can hand a
|
|
1667
|
+
file back — on exactly the internal, no-egress targets the method exists for.
|
|
1668
|
+
|
|
1669
|
+
### Added
|
|
1670
|
+
|
|
1671
|
+
- **`--file-write-path DIR` + `--file-read-url URL`** name the two halves of the
|
|
1672
|
+
read-back channel directly, so an LFI endpoint, a download or export handler,
|
|
1673
|
+
an attachment fetcher or a `/tmp`-backed preview all work. The template takes
|
|
1674
|
+
`{name}` (the filename), `{path}` (the full server-side path) and `{path_enc}`
|
|
1675
|
+
(that path percent-encoded); only those three are substituted, so a URL that
|
|
1676
|
+
legitimately contains braces survives unchanged.
|
|
1677
|
+
|
|
1678
|
+
Measured against a target with a download handler and nothing serving the
|
|
1679
|
+
write directory: the web-root form confirms **0** — reporting an exploitable
|
|
1680
|
+
target clean — and the general form confirms **7**.
|
|
1681
|
+
|
|
1682
|
+
### Fixed
|
|
1683
|
+
|
|
1684
|
+
- **The read-back fetch now carries the run's headers**, so an authenticated
|
|
1685
|
+
download, export, attachment or LFI handler can actually be read. It went out
|
|
1686
|
+
bare, which barely mattered while the channel had to be a web root — static
|
|
1687
|
+
file serving is rarely authenticated — and became the likely case the moment
|
|
1688
|
+
the channel could be an application endpoint. Measured against a handler
|
|
1689
|
+
behind a bearer token: the write executed on every probe and the verdict was
|
|
1690
|
+
`negative`, "token absent from the fetched file". Now 7 confirmations on the
|
|
1691
|
+
same target.
|
|
1692
|
+
- **Credentials are carried only to the same origin.** A read-back URL on
|
|
1693
|
+
another host is someone else's server, and replaying the target's session
|
|
1694
|
+
cookie or bearer token to it would leak the credential, so those headers are
|
|
1695
|
+
dropped there while the rest still go — and the run says so, because the
|
|
1696
|
+
symptom would otherwise look like a clean target. `Content-Type` and
|
|
1697
|
+
`Content-Length` are dropped from the fetch too: they describe a body the GET
|
|
1698
|
+
does not have.
|
|
1699
|
+
|
|
1700
|
+
### Changed
|
|
1701
|
+
|
|
1702
|
+
- **`--webroot` / `--web-base-url` are now the web-root alias** for the general
|
|
1703
|
+
form: a web root is just the case where the read URL is the base plus the
|
|
1704
|
+
filename. Existing command lines are unaffected. Both are resolved in one
|
|
1705
|
+
place inside the method, so the alias and the general form cannot drift — and
|
|
1706
|
+
the gate, the pre-flight banner and the blind-sink advice all ask that same
|
|
1707
|
+
resolver instead of testing for the webroot pair.
|
|
1708
|
+
- `blind_sink_advice` reads its flags defensively, so an args-like object
|
|
1709
|
+
missing a newer field costs a line of advice rather than a traceback.
|
|
1710
|
+
|
|
1711
|
+
## [2.28.0] — 2026-08-17
|
|
1712
|
+
|
|
1713
|
+
Injection-point enumeration. `-p NAME` needed the tester to already know which
|
|
1714
|
+
parameter was the sink, so a capture's other candidates — including the headers
|
|
1715
|
+
and nested JSON leaves that carry some of the highest-value classes — were never
|
|
1716
|
+
tried.
|
|
1717
|
+
|
|
1718
|
+
### Added
|
|
1719
|
+
|
|
1720
|
+
- **`-p all` / `--auto-params KINDS`** expands one captured request into every
|
|
1721
|
+
candidate injection point and runs the selected `--methods` against each.
|
|
1722
|
+
Query values, JSON leaves addressed by path (`user.profile.name`, `tags[1]`),
|
|
1723
|
+
form fields, cookie crumbs and headers, each rewritten in **its own**
|
|
1724
|
+
serialization rather than blanket-encoded. Verified end to end: a sink
|
|
1725
|
+
reachable only through `User-Agent` is confirmed from `-r request.txt -p all`
|
|
1726
|
+
with no manual header selection.
|
|
1727
|
+
- **`--point-order fast|thorough`** — `fast` tries a curated high-yield header
|
|
1728
|
+
list (the headers real published RCEs inject through); `thorough` adds every
|
|
1729
|
+
remaining non-hop-by-hop header. **`--max-points N`** bounds the run and
|
|
1730
|
+
reports what it dropped. **`--include-path-segments`** is opt-in, because
|
|
1731
|
+
rewriting a path segment usually just produces a 404.
|
|
1732
|
+
- **The run states its cost before sending it** —
|
|
1733
|
+
`6 points x ~61 probes = at least 372 requests` — via a new
|
|
1734
|
+
`estimate_detection_probes`, which builds the probes and counts them without
|
|
1735
|
+
firing any. Enumeration multiplies an already-laddered probe count by the
|
|
1736
|
+
candidate count, and an operator on a monitored engagement has to see that
|
|
1737
|
+
before it happens rather than infer it from the traffic.
|
|
1738
|
+
- **Findings name the point they came from**: `[reflected/unix/raw] at header
|
|
1739
|
+
'User-Agent' ...`.
|
|
1740
|
+
|
|
1741
|
+
### Changed
|
|
1742
|
+
|
|
1743
|
+
- **Each candidate carries its own payload-free control.** Differencing a header
|
|
1744
|
+
probe against a query probe's control would compare two different responses
|
|
1745
|
+
and prove nothing.
|
|
1746
|
+
- **Cheap methods run first per candidate, and a candidate stops at its first
|
|
1747
|
+
confirmation.** `reflected` and `eval` cost one response each; `time` sleeps
|
|
1748
|
+
and `oob` waits for a callback, and on a candidate that has already proven
|
|
1749
|
+
execution those buy a second name for the same finding. Candidates that stay
|
|
1750
|
+
clean still get every method, and single-point runs are unchanged.
|
|
1751
|
+
- A JSON leaf is **replaced, never created**. Assigning to a missing key would
|
|
1752
|
+
have injected into a field the application never sends — a probe that cannot
|
|
1753
|
+
say anything about the parameter that does exist. Caught by its own test.
|
|
1754
|
+
- **JSON points are addressed by tokens, not by a joined path string.** A key may
|
|
1755
|
+
itself contain the separator: `{"user.name": ..., "user": {"name": ...}}`
|
|
1756
|
+
rendered *both* leaves as `user.name`, so the literal key was never probed and
|
|
1757
|
+
both candidates mutated the nested field — a false negative and a misattributed
|
|
1758
|
+
finding at once. Tokens remove the ambiguity, and the display form
|
|
1759
|
+
bracket-quotes such a key (`["user.name"]`) so the two stay distinguishable on
|
|
1760
|
+
screen.
|
|
1761
|
+
- **A deeply nested captured body no longer ends `-p all` with a traceback.**
|
|
1762
|
+
`json.loads` recurses in C, so `RecursionError` joins the caught exceptions in
|
|
1763
|
+
both the enumerator and the placer, as it already had in the response-channel
|
|
1764
|
+
parser. The body yields no candidates; the rest of the request still enumerates.
|
|
1765
|
+
- **The cost estimate honours `--max-payloads`.** It counted every probe the
|
|
1766
|
+
carriers could produce while the run stops at the cap, so the figure was wrong
|
|
1767
|
+
exactly when the operator had reached for the budget guard.
|
|
1768
|
+
- `Host`, `Content-Length`, `Cookie` and the hop-by-hop headers are never
|
|
1769
|
+
candidates: injecting into those changes the request's plumbing rather than
|
|
1770
|
+
testing the application, and two of them are rebuilt by the delivery layer.
|
|
1771
|
+
|
|
1772
|
+
## [2.27.0] — 2026-08-17
|
|
1773
|
+
|
|
1774
|
+
Engine carriers for the `eval` probe. Three template engines evaluate the
|
|
1775
|
+
injected expression perfectly and still made RCEKit report `negative`, because
|
|
1776
|
+
what came back was not the bare product the oracle searches for.
|
|
1777
|
+
|
|
1778
|
+
### Added
|
|
1779
|
+
|
|
1780
|
+
- **`eval_carriers` in the corpus**, and `--eval-engines auto|<names>` to select
|
|
1781
|
+
them. A carrier wraps the same random-operand arithmetic in an engine-specific
|
|
1782
|
+
form; it never changes the oracle, and the bare probes still run first. Each
|
|
1783
|
+
entry records `notes` (why it exists) and `verified` (what it was measured
|
|
1784
|
+
against). Declarative, so a new carrier is a JSON entry rather than a code
|
|
1785
|
+
change.
|
|
1786
|
+
|
|
1787
|
+
| Engine | Bare `${a*b}` returned | Carrier | Carrier returned |
|
|
1788
|
+
|---|---|---|---|
|
|
1789
|
+
| Freemarker | `2,070,761,401` (locale grouping) | `${(a*b)?c}` | `2070761401` |
|
|
1790
|
+
| Velocity | `${a*b}` verbatim — a *reference*, not an expression | `#set($rk=a*b)$rk` | `2070761401` |
|
|
1791
|
+
| Thymeleaf | `${a*b}` verbatim — needs inlining brackets | `[[${a*b}]]` | `2070761401` |
|
|
1792
|
+
|
|
1793
|
+
Measured against freemarker 2.3.32, velocity-engine-core 2.3 and thymeleaf
|
|
1794
|
+
3.1.2, running RCEKit's own generated probes through each engine: bare form
|
|
1795
|
+
`CONFIRMS=no`, carrier `CONFIRMS=YES`, for all three.
|
|
1796
|
+
- **The evidence line names the carrier** — `target computed '3979016' via the
|
|
1797
|
+
freemarker carrier` — so a finding says which engine quirk it worked around.
|
|
1798
|
+
A bare confirmation reads exactly as before.
|
|
1799
|
+
|
|
1800
|
+
### Notes
|
|
1801
|
+
|
|
1802
|
+
- **Carriers are not sandbox escapes, and no sandbox-escape carrier ships.** The
|
|
1803
|
+
premise that a sandboxed engine blocks the arithmetic probe did not survive
|
|
1804
|
+
measurement: a member-access sandbox restricts method and field access, and
|
|
1805
|
+
arithmetic needs neither. With OGNL member access denied for *everything*,
|
|
1806
|
+
`40277*51413` still returned `2070761401` while `@java.lang.Math@max(1,2)` was
|
|
1807
|
+
blocked; SpEL's restricted `SimpleEvaluationContext` and Jinja2's
|
|
1808
|
+
`SandboxedEnvironment` behaved the same way. The bare probes already cover
|
|
1809
|
+
those engines.
|
|
1810
|
+
- The frequently-cited OGNL escape `(#_memberAccess=@ognl.OgnlContext@DEFAULT_MEMBER_ACCESS)`
|
|
1811
|
+
additionally targets a field that **no longer exists in OGNL 3.3.4**, so on a
|
|
1812
|
+
current engine it is a probe that can only come back negative.
|
|
1813
|
+
|
|
1814
|
+
## [2.26.0] — 2026-08-16
|
|
1815
|
+
|
|
1816
|
+
The sink-shape ladder. An injected value lands in a *shape* — mid-command,
|
|
1817
|
+
inside quotes, as the whole command — and the shape decides what can reach it.
|
|
1818
|
+
Two shapes had no probe that fitted, so a genuinely exploitable target reported
|
|
1819
|
+
clean.
|
|
1820
|
+
|
|
1821
|
+
### Added
|
|
1822
|
+
|
|
1823
|
+
- **`--sink-shape auto|sep|raw|chain|newline|dq|sq|subshell`** (comma-separated)
|
|
1824
|
+
names which shapes the shell probes try. `auto` is the whole ladder and the
|
|
1825
|
+
default. Underneath it selects the existing separator sweep and break-out
|
|
1826
|
+
contexts, so naming a rung narrows a supported run rather than switching on a
|
|
1827
|
+
parallel path. The plan is printed before anything is sent, because the ladder
|
|
1828
|
+
multiplies request count and an operator on a monitored engagement needs to
|
|
1829
|
+
see the cost first.
|
|
1830
|
+
- **The `subshell` rung — `$(...)` and backticks.** Reaches a value sitting
|
|
1831
|
+
inside double quotes *without closing the quote*, which is the one case a
|
|
1832
|
+
quoted break-out loses to a filter on the quote character itself. Measured
|
|
1833
|
+
against `system("echo PING \"$input\"")`: with `"` stripped, `dq` is inert and
|
|
1834
|
+
both substitution forms execute; with `$` stripped, `dq` executes and the
|
|
1835
|
+
backtick form still does. Both ship because they survive different filters.
|
|
1836
|
+
|
|
1837
|
+
**Which method it helps is the counter-intuitive part.** `reflected`'s core is
|
|
1838
|
+
`$((a+b))`, which the shell expands inside double quotes anyway, so that
|
|
1839
|
+
method already confirmed there. The methods whose core must actually *run* —
|
|
1840
|
+
`time` (a sleep), `file` (a redirect), `oob` (a fetch) — are completely inert
|
|
1841
|
+
inside those quotes. On a quote-filtering sink, `--methods file` went from **0
|
|
1842
|
+
confirmations to 2**: it had been reporting an exploitable target as clean.
|
|
1843
|
+
- **The `raw` rung is now part of `auto`, for every shell method.** A
|
|
1844
|
+
`qx/$input/`-style sink, where the input is the whole command, previously
|
|
1845
|
+
needed `--sink-raw` — so it reported clean unless the operator already
|
|
1846
|
+
suspected the shape. One extra probe per carrier buys it. `--sink-raw` keeps
|
|
1847
|
+
its meaning as the narrowing alias for `--sink-shape raw`, and no existing
|
|
1848
|
+
command line changes behaviour. `reflected`, `file`, `time` and `oob` all
|
|
1849
|
+
build their candidates through one `_separator_candidates` helper, so a rung
|
|
1850
|
+
cannot reach some methods and not others; `time` screens it in its second
|
|
1851
|
+
wave, alongside the separators it holds back.
|
|
1852
|
+
|
|
1853
|
+
### Fixed
|
|
1854
|
+
|
|
1855
|
+
- **A method that builds no probes no longer reports `negative`.** An aggregate
|
|
1856
|
+
method asked to judge zero samples answers honestly — "no delay was observed",
|
|
1857
|
+
"no callback arrived" — and that reads as "not vulnerable" from a run that
|
|
1858
|
+
tested nothing. The engine now emits no row for a carrier that produced no
|
|
1859
|
+
probes, which lets its own loud nothing-tested path fire instead. Reachable
|
|
1860
|
+
through any narrowing that leaves a carrier with nothing to send.
|
|
1861
|
+
|
|
1862
|
+
### Changed
|
|
1863
|
+
|
|
1864
|
+
- **The pre-flight sink-shape plan is computed from the effective run**, not
|
|
1865
|
+
from the `--sink-shape` value. `--separators`, `--contexts` and `--sink-raw`
|
|
1866
|
+
each narrow the ladder, so printing the flag described a run that would not
|
|
1867
|
+
happen — and this output is presented as an audit of the traffic about to be
|
|
1868
|
+
sent. `effective_sink_shapes` is the single source of truth the engine and the
|
|
1869
|
+
plan both read.
|
|
1870
|
+
- **The backtick context drops probe shapes that carry their own backtick.**
|
|
1871
|
+
Backticks do not nest, so such a probe closes the outer substitution early and
|
|
1872
|
+
could only ever come back negative. `$( )` does nest and keeps every shape.
|
|
1873
|
+
- **Naming `--separators` now implies the sink is separator-led**, so the `raw`
|
|
1874
|
+
rung is dropped unless `--sink-shape` names it explicitly. A profile with
|
|
1875
|
+
`sink_needs_separator` drops it for the same reason. Both keep an explicitly
|
|
1876
|
+
narrowed run from being widened behind the operator's back.
|
|
1877
|
+
|
|
1878
|
+
## [2.25.0] — 2026-08-16
|
|
1879
|
+
|
|
1880
|
+
A coverage benchmark, so a claim about what RCEKit confirms can be checked
|
|
1881
|
+
instead of asserted. The unit suite proves the tool reaches the right verdict
|
|
1882
|
+
from a given response; it cannot prove it confirms Webmin.
|
|
1883
|
+
|
|
1884
|
+
### Added
|
|
1885
|
+
|
|
1886
|
+
- **`--detect-json PATH`** writes a detection run as JSON: the run's overall
|
|
1887
|
+
verdict, per-verdict counts, and every probe with its payload, method, context
|
|
1888
|
+
and evidence. Text output is unchanged. This is the supported way to consume a
|
|
1889
|
+
run programmatically — scraping stdout cannot be made reliable, because a
|
|
1890
|
+
probe payload may contain a literal newline (the newline separator is a real
|
|
1891
|
+
one, so line-oriented parsing splits a payload in half) and the detection path
|
|
1892
|
+
exits 0 whether it confirmed or came back clean.
|
|
1893
|
+
- **`tests/bench/` — the coverage benchmark harness.** Each case brings a real
|
|
1894
|
+
vulnerable build up, runs RCEKit as an operator would, checks the verdict, and
|
|
1895
|
+
tears it down; `--markdown` emits the coverage table. Not part of
|
|
1896
|
+
`python -m unittest discover -s tests` — cases need Docker and pull real
|
|
1897
|
+
images — so it runs by hand or in a dedicated job, and exits non-zero if any
|
|
1898
|
+
case fails. Two cases ship, transcribed from `docs/verify-it-yourself.md`:
|
|
1899
|
+
Webmin CVE-2019-15107 and Struts2 S2-001.
|
|
1900
|
+
- **A negative control is a required key.** A benchmark without controls measures
|
|
1901
|
+
nothing: a tool that shouted `confirmed` at every target would score full marks
|
|
1902
|
+
on the vulnerable half. Three kinds are supported — a patched build, the same
|
|
1903
|
+
target probed for the wrong class, and a weaker method that must stay below
|
|
1904
|
+
`confirmed` on a target where it happens to be right. The runner refuses four
|
|
1905
|
+
shapes of non-control: no control at all; one expecting `confirmed`; one that
|
|
1906
|
+
runs the identical invocation against an identical target (judged on what it
|
|
1907
|
+
would actually run, so an explicit copy of the vulnerable invocation is caught
|
|
1908
|
+
as well as an omitted one); and one expecting `error` or `nothing-tested`,
|
|
1909
|
+
since both mean the target was never exercised and such a control would stay
|
|
1910
|
+
green with the detection engine entirely broken. Validation and execution
|
|
1911
|
+
share one `control_plan` so they cannot drift.
|
|
1912
|
+
- **`overall_detection_verdict`** collapses a run to one verdict, ordered by what
|
|
1913
|
+
an operator must not miss rather than by frequency: one `confirmed` among a
|
|
1914
|
+
hundred negatives is the finding. `error` is reported only when *nothing*
|
|
1915
|
+
reached the target, and a run that built no probes is `nothing-tested` —
|
|
1916
|
+
never `negative`, which would read as "not vulnerable".
|
|
1917
|
+
|
|
1918
|
+
### Changed
|
|
1919
|
+
|
|
1920
|
+
- `CONTRIBUTING.md` asks for a bench case alongside new detection coverage, and
|
|
1921
|
+
for the README table to state the tier the case actually reached.
|
|
1922
|
+
|
|
1923
|
+
### Notes
|
|
1924
|
+
|
|
1925
|
+
- The two shipped cases have **not yet been executed through the harness** — it
|
|
1926
|
+
was written where no Docker daemon was available. Their invocations come from
|
|
1927
|
+
a documented, reproduced guide, but the case files themselves are unvalidated;
|
|
1928
|
+
`tests/bench/README.md` says so and flags the one field that is a guess. No
|
|
1929
|
+
README claim was changed to assert benchmark results.
|
|
1930
|
+
|
|
1931
|
+
## [2.24.0] — 2026-08-16
|
|
1932
|
+
|
|
1933
|
+
The computed value is no longer looked for in the response body alone. A sink
|
|
1934
|
+
whose output surfaces anywhere else in the response was reported `negative` — a
|
|
1935
|
+
false negative on a class RCEKit already claims to cover, which is worse than a
|
|
1936
|
+
missing class. The oracle, the random operands and the control differential are
|
|
1937
|
+
unchanged; only the set of places searched is wider.
|
|
1938
|
+
|
|
1939
|
+
### Added
|
|
1940
|
+
|
|
1941
|
+
- **Whole-response evidence search.** Every confirmation now sweeps the response
|
|
1942
|
+
body, the application response headers, individual cookie values, the redirect
|
|
1943
|
+
target RCEKit actually landed on, the HTTP reason phrase, and each leaf of a
|
|
1944
|
+
parsed JSON body. Real sinks put command output in a debug header or a
|
|
1945
|
+
`Set-Cookie`, and API targets surface an evaluator's result inside a nested
|
|
1946
|
+
error envelope — `{"error": {"detail": "cannot render 2058898001"}}` — where a
|
|
1947
|
+
substring search of the serialised body misses a value the encoder escaped.
|
|
1948
|
+
- **The evidence line names the channel that carried the value**, e.g.
|
|
1949
|
+
`target computed 'RK…' in header X-Cmd-Out (random operands, absent from
|
|
1950
|
+
control)`, so the finding stays reproducible by hand. A body-carried
|
|
1951
|
+
confirmation reads exactly as it did before.
|
|
1952
|
+
|
|
1953
|
+
### Changed
|
|
1954
|
+
|
|
1955
|
+
- **The control differential now covers every channel, not just the body.** A
|
|
1956
|
+
value present anywhere in the payload-free control is not attributable to
|
|
1957
|
+
execution, so it yields `inconclusive` wherever it turned up. This is stricter
|
|
1958
|
+
than comparing only the channel that matched, and it is what keeps a wider
|
|
1959
|
+
search from becoming a looser verdict.
|
|
1960
|
+
- **The `file` method's control check covers every channel too**, on the same
|
|
1961
|
+
reasoning: its token is random, so its presence in any control channel means
|
|
1962
|
+
it did not get there by being written and served.
|
|
1963
|
+
|
|
1964
|
+
### Security
|
|
1965
|
+
|
|
1966
|
+
- **A deeply nested JSON response can no longer silence detection.** Channels are
|
|
1967
|
+
built inside the delivery `try`/`except`, so a `RecursionError` while parsing
|
|
1968
|
+
or walking the body escaped as a network failure: a response that arrived
|
|
1969
|
+
perfectly well was reported "request never reached the target". Measured:
|
|
1970
|
+
every one of the 46 probes in a default `reflected` run turned into `error`,
|
|
1971
|
+
which a target could induce deliberately to hide a live sink behind a thousand
|
|
1972
|
+
nested arrays. Version-independent, though the source moves — CPython 3.12
|
|
1973
|
+
raised the C recursion limit its JSON scanner runs under, so on 3.12/3.13 the
|
|
1974
|
+
parser survives a depth that breaks it on 3.8–3.11 and the recursive leaf walk
|
|
1975
|
+
hit the ordinary Python limit instead. The walk is now iterative and
|
|
1976
|
+
depth-capped, `RecursionError` from the parser costs the JSON channels only,
|
|
1977
|
+
and building channels can never turn a delivered response into a delivery
|
|
1978
|
+
failure.
|
|
1979
|
+
- **Transport headers are excluded from the sweep.** `Content-Length`, `Date`,
|
|
1980
|
+
`Age`, `ETag` and their neighbours are generated below the application and can
|
|
1981
|
+
never carry a computed value, but they *are* numeric — and the `expr` probe's
|
|
1982
|
+
expected value is a bare boundary-fenced number. Searching them would let a
|
|
1983
|
+
byte count collide with an arithmetic result and read as execution. Locked in
|
|
1984
|
+
by a test that puts the expected value in `Content-Length` and requires
|
|
1985
|
+
`negative`.
|
|
1986
|
+
|
|
1987
|
+
## [2.23.3] — 2026-08-04
|
|
1988
|
+
|
|
1989
|
+
Four items from the same review: requests and seconds spent on work that could
|
|
1990
|
+
not produce a result. No verdict changes — the lab still confirms 15 of 15
|
|
1991
|
+
vulnerable sinks with nothing on the clean five — the run just stops paying for
|
|
1992
|
+
probes that were structurally unable to confirm.
|
|
1993
|
+
|
|
1994
|
+
### Changed
|
|
1995
|
+
|
|
1996
|
+
- **The `awk` probe is no longer sent into a context that wraps the payload in
|
|
1997
|
+
quotes.** It carries double quotes, so in `attribute` the quote closed early
|
|
1998
|
+
and the rest was not a command: 5 requests per carrier that could only ever
|
|
1999
|
+
come back negative. Measured on a verbose shell sink, that shape confirmed 8
|
|
2000
|
+
times in `raw` and 0 times in `attribute`. Break-out contexts such as
|
|
2001
|
+
`shell_double_quoted` *close* the sink's quote and comment its tail, so they
|
|
2002
|
+
still get it. The same guard covers the PowerShell out-of-band shape.
|
|
2003
|
+
- **The timing screen runs in two waves.** Every delayed screen probe costs a
|
|
2004
|
+
real sleep, so screening all five separators up front spent `5 × base` seconds
|
|
2005
|
+
on every carrier, including the ones that cannot break out at all. `; ` and
|
|
2006
|
+
`| ` are screened first and the rest only if neither delayed — a sink that
|
|
2007
|
+
filters both is still swept, it is just no longer the price everyone pays.
|
|
2008
|
+
- **The out-of-band callback window is no longer paid per carrier.** Callbacks
|
|
2009
|
+
land in a burst once the channel works, so a target that has not produced one
|
|
2010
|
+
across every probe fired so far is not going to. The first carrier still gets
|
|
2011
|
+
the full window, so a target that does call back is never cut short before its
|
|
2012
|
+
first hit. On a clean target with the default carriers this was 30s of pure
|
|
2013
|
+
waiting; it is now ~12s.
|
|
2014
|
+
- **`--probe-depth` documents what it does on Windows**, which is nothing:
|
|
2015
|
+
`cmd.exe` has no `#` comment, no `${IFS}` and no `awk`, so both depths send
|
|
2016
|
+
the single `set /a` probe. The docs promised three extra shapes per sink
|
|
2017
|
+
without that caveat.
|
|
2018
|
+
|
|
2019
|
+
## [2.23.2] — 2026-08-04
|
|
2020
|
+
|
|
2021
|
+
Three findings from a review of the detection work in 2.22.0 and 2.23.0. All
|
|
2022
|
+
three are the same shape: the run said something that was not true — about what
|
|
2023
|
+
it had done, about what it had looked for, or about which channel was live.
|
|
2024
|
+
|
|
2025
|
+
### Fixed
|
|
2026
|
+
|
|
2027
|
+
- **`--methods oob` ignored `--verify-active-risk`.** Detection methods build
|
|
2028
|
+
their own probes and so bypass every corpus-level safety filter. That was
|
|
2029
|
+
harmless while every method was inert, but this one makes the target open
|
|
2030
|
+
outbound connections — and the same run printed *"low-impact (safe) payloads
|
|
2031
|
+
only; pass `--verify-active-risk intrusive` to also fire … OOB"* and then fired
|
|
2032
|
+
OOB anyway. It now needs `--verify-active-risk intrusive`, the same tier that
|
|
2033
|
+
holds back the corpus OOB payloads, and refuses before the listener binds.
|
|
2034
|
+
- **`--probe-depth quick` silently narrowed the timing separator screen to
|
|
2035
|
+
`; `.** That put back the exact blind spot the screen was added to remove, so
|
|
2036
|
+
a sink that merely filters `;` reported negative — and only for the operator
|
|
2037
|
+
who chose `quick` to be gentle on a rate-limited target. Both depths now screen
|
|
2038
|
+
every candidate separator; `--probe-depth` governs probe *shapes* only, and
|
|
2039
|
+
`--separators` remains the way to narrow break-outs deliberately.
|
|
2040
|
+
- **The DNS out-of-band probes could not call back on the default port, and
|
|
2041
|
+
nothing said so.** A DNS callback travels the real resolver hierarchy, so it
|
|
2042
|
+
only arrives if the listener *is* the authority for the OOB domain — port 53
|
|
2043
|
+
plus NS delegation. On `--listen-dns-port 5335` the DNS shapes were still sent,
|
|
2044
|
+
never fired, and the startup line reported `DNS :5335` with no caveat. Since
|
|
2045
|
+
most of the shapes are DNS ones — a resolver is often the only egress a
|
|
2046
|
+
hardened target has — the silence was expensive. RCEKit now says which channel
|
|
2047
|
+
is live.
|
|
2048
|
+
- The blind-sink advice added in 2.23.0 suggested an `oob` command without the
|
|
2049
|
+
risk flag, which the gate above would refuse. Naming a command the tool then
|
|
2050
|
+
declines to run is a small version of the same problem, so it now spells out
|
|
2051
|
+
`--verify-active-risk intrusive`.
|
|
2052
|
+
|
|
2053
|
+
## [2.23.1] — 2026-08-03
|
|
2054
|
+
|
|
2055
|
+
### Added
|
|
2056
|
+
|
|
2057
|
+
- **[Verify it yourself](docs/verify-it-yourself.md)** — reproduce the README's
|
|
2058
|
+
confirmations locally against dockerised [vulhub](https://github.com/vulhub/vulhub)
|
|
2059
|
+
targets. Webmin CVE-2019-15107 driven from a captured request (`reflected` →
|
|
2060
|
+
`confirmed`, then `time` → `needs-review` on the *same* sink, which is the
|
|
2061
|
+
clearest demonstration that the tiers are not merged), and Struts2 S2-001
|
|
2062
|
+
(`eval` confirms, `reflected` does not, on a target where both were tried).
|
|
2063
|
+
|
|
2064
|
+
Log4Shell is documented as an advanced case rather than a five-minute one: its
|
|
2065
|
+
sink is a JNDI lookup inside a logging library, so `--methods oob` does not
|
|
2066
|
+
apply — that method builds shell probes for shell-capable environments. The
|
|
2067
|
+
`${jndi:…}` payloads come from the `oob` *category* with the listener
|
|
2068
|
+
correlating the callback, and the token rides in a DNS label, which needs a
|
|
2069
|
+
delegated domain. Saying so is cheaper than a reader discovering it mid-demo.
|
|
2070
|
+
|
|
2071
|
+
## [2.23.0] — 2026-08-03
|
|
2072
|
+
|
|
2073
|
+
The three sinks v2.22.0 still could not reach. One was a real gap in the probe
|
|
2074
|
+
set; the other two were a reporting problem, not a detection one. With both
|
|
2075
|
+
closed, a single `--methods reflected,eval,oob` run confirms **all fifteen**
|
|
2076
|
+
vulnerable sinks in the lab and still reports nothing on any of the five clean
|
|
2077
|
+
ones.
|
|
2078
|
+
|
|
2079
|
+
### Added
|
|
2080
|
+
|
|
2081
|
+
- **A space-free probe, sent at both probe depths.** Stripping spaces is a filter
|
|
2082
|
+
of the same family as stripping `;` — it looks like it disarms command
|
|
2083
|
+
injection and does not, because `${IFS}` is a space as far as the shell is
|
|
2084
|
+
concerned. Every other probe carries a space, so that one filter silenced all
|
|
2085
|
+
of them and the sink was only reachable if the operator thought to pass
|
|
2086
|
+
`--evade low`. The separator's trailing space is trimmed with it (`;echo…`, not
|
|
2087
|
+
`; echo…`); the newline separator is unaffected. It costs one shape, so it is
|
|
2088
|
+
not part of the `--probe-depth` trade-off, and it is skipped under
|
|
2089
|
+
`--evade low`, which already applies the same transform everywhere.
|
|
2090
|
+
- **Guidance when every in-band probe comes back negative.** A results-based
|
|
2091
|
+
method cannot confirm a sink that returns no output — there is nowhere for the
|
|
2092
|
+
computed value to appear — so that negative is not evidence the target is
|
|
2093
|
+
clean. A run of `reflected`/`eval` alone that confirms nothing now says exactly
|
|
2094
|
+
that and names the methods that could still reach a blind sink, with the flags
|
|
2095
|
+
each one needs. It is suppressed once a blind-capable method has already run,
|
|
2096
|
+
and the `file` line is dropped once a web root is known.
|
|
2097
|
+
|
|
2098
|
+
## [2.22.0] — 2026-08-03
|
|
2099
|
+
|
|
2100
|
+
Detection coverage. Measured against a lab of twenty sinks — fifteen genuinely
|
|
2101
|
+
vulnerable, five deliberately clean — the results-based methods went from
|
|
2102
|
+
confirming 8 of the 15 to confirming 12, with no new false positives on any of
|
|
2103
|
+
the clean ones.
|
|
2104
|
+
|
|
2105
|
+
### ⚠️ A blind-timing candidate could be pure latency drift
|
|
2106
|
+
|
|
2107
|
+
`--methods time` fired its probes in a fixed ascending delay order
|
|
2108
|
+
(`0,0,N,N,2N,2N`), which makes the injected delay collinear with the request
|
|
2109
|
+
index. A target that simply gets **slower during the run** — progressive load, a
|
|
2110
|
+
rate limiter backing off, a filling log — therefore produced a textbook-perfect
|
|
2111
|
+
linear fit while being entirely un-injectable. In the lab this reproduced on 8
|
|
2112
|
+
of 8 runs against a sink with no command execution anywhere in it.
|
|
2113
|
+
|
|
2114
|
+
The probe order is now randomised, and the request index enters the regression
|
|
2115
|
+
as a nuisance term, so drift loads onto a drift coefficient instead of
|
|
2116
|
+
masquerading as a sleep. The same lab sink now reports negative on 9 of 9 runs,
|
|
2117
|
+
with every genuine timing detection preserved. If you have a `needs-review`
|
|
2118
|
+
timing candidate from an earlier version against a target that was under load,
|
|
2119
|
+
it is worth re-running.
|
|
2120
|
+
|
|
2121
|
+
### Added
|
|
2122
|
+
|
|
2123
|
+
- **`--methods oob`** — out-of-band detection, the first `confirmed`-tier method
|
|
2124
|
+
for a sink that returns nothing *and* has no writable web root. Starts the
|
|
2125
|
+
built-in HTTP+DNS listener in-process and asks the target to resolve or fetch
|
|
2126
|
+
`<token>.<oob-host>`; a callback carrying a token the target could only have
|
|
2127
|
+
learned by running the command is proof of execution. Each probe gets its own
|
|
2128
|
+
token, so the finding names the break-out that actually worked. One shape puts
|
|
2129
|
+
a computed value in the DNS label, so the callback proves the shell evaluated
|
|
2130
|
+
arithmetic rather than merely resolving a name. Requires `--oob-host`, since
|
|
2131
|
+
it makes the target open outbound connections.
|
|
2132
|
+
- **`--probe-depth quick|full`** (default `full`) — trades requests for
|
|
2133
|
+
coverage. `full` adds three probe shapes, each aimed at a filter that silenced
|
|
2134
|
+
the canonical ones: substitution-free (`awk`, bare `expr`) for sinks that strip
|
|
2135
|
+
`$(` and backticks; keyword-diverse (`awk`) for filters on `echo`/`expr`; and
|
|
2136
|
+
comment-terminated (`… #`) for applications that append a redirect, extra
|
|
2137
|
+
arguments or a pipe after the injection point. `quick` keeps the old probe set
|
|
2138
|
+
at roughly half the requests.
|
|
2139
|
+
|
|
2140
|
+
### Fixed
|
|
2141
|
+
|
|
2142
|
+
- **A `ping '<input>'` sink could not be detected at all.** The
|
|
2143
|
+
`shell_single_quoted`/`shell_double_quoted` contexts exist precisely for input
|
|
2144
|
+
interpolated inside quotes, but they are not in `default_contexts`, so no
|
|
2145
|
+
record carried them and the detection engine never tried them — the one sink
|
|
2146
|
+
shape they exist for was the one shape that always reported clean. They are now
|
|
2147
|
+
probed by default, and skipped when `--contexts` names a selection explicitly.
|
|
2148
|
+
- **`--methods time` reported a `;`-filtering sink as negative.** A regression
|
|
2149
|
+
blends its probes into one measurement, so it could not sweep separators the
|
|
2150
|
+
way the results-based methods do and was locked to `; ` alone — while
|
|
2151
|
+
`| sleep 3` delayed on the same sink. It now screens every candidate separator
|
|
2152
|
+
with one cheap probe each, then runs the regression through whichever one
|
|
2153
|
+
actually delayed.
|
|
2154
|
+
- **A trailing redirect or pipe in the sink hid a working probe.**
|
|
2155
|
+
`<cmd> <input> 2>/dev/null` and `<cmd> <input> | grep …` swallow the probe's
|
|
2156
|
+
output, so it executed and still read as negative. The comment-terminated
|
|
2157
|
+
shapes comment that tail out.
|
|
2158
|
+
|
|
2159
|
+
## [2.21.1] — 2026-08-02
|
|
2160
|
+
|
|
2161
|
+
First release since v2.15.2. The headline is not a new feature — it is that
|
|
2162
|
+
detection is now correct in cases where it previously was not.
|
|
2163
|
+
|
|
2164
|
+
### ⚠️ Re-check findings from v2.15.2 and earlier
|
|
2165
|
+
|
|
2166
|
+
**A reflection could be reported as `confirmed`.** The paired same-token control
|
|
2167
|
+
in `run_verification` was gated on a plain `re.search`, while the verdict itself
|
|
2168
|
+
used the encoding-aware search. A target that only echoes input but wraps its
|
|
2169
|
+
output — base64, hex, URL- or HTML-encoded — skipped the control entirely and was
|
|
2170
|
+
reported as proven execution: precisely the case the encoding-aware search was
|
|
2171
|
+
added for. If you ran an earlier version against a target that encodes its
|
|
2172
|
+
responses, a `confirmed` verdict from that run is worth re-testing.
|
|
2173
|
+
|
|
2174
|
+
### Fixed — false negatives on exploitable targets
|
|
2175
|
+
|
|
2176
|
+
- **Separator sweep.** Shell probes always broke out with a single hardcoded
|
|
2177
|
+
`; `, so a sink that strips `;` — the most common partial mitigation there is,
|
|
2178
|
+
and one that stops nothing on its own — defeated every probe. Measured against
|
|
2179
|
+
nine deliberately vulnerable local sinks, detection was correct on 5 of 9;
|
|
2180
|
+
three of the four misses were exploitable targets reported clean. Probes now
|
|
2181
|
+
sweep `; `, `| `, `|| `, `&& ` and a newline, narrowable with `--separators`.
|
|
2182
|
+
- **Language runtimes.** An environment names what runs the application, not what
|
|
2183
|
+
executes the injected command: PHP's `system()`, Python's `os.system()`,
|
|
2184
|
+
Node's `child_process.exec()`, Ruby's `system()`, Perl's backticks and Go's
|
|
2185
|
+
`os/exec` all hand the string to `/bin/sh`. Scoping a run to the language the
|
|
2186
|
+
application is written in — the natural thing to do — used to send no shell
|
|
2187
|
+
probes at all.
|
|
2188
|
+
- **Whole-command sinks.** `--sink-raw` sends probes as bare commands for sinks
|
|
2189
|
+
that execute the input as the entire command (`qx/$input/`, `sh -c "$input"`),
|
|
2190
|
+
where a leading `;` is a syntax error that guaranteed a false negative.
|
|
2191
|
+
- **Captured requests.** A trailing newline in a saved request body is no longer
|
|
2192
|
+
sent as part of the body.
|
|
2193
|
+
|
|
2194
|
+
### Fixed — a failed request is not a clean result
|
|
2195
|
+
|
|
2196
|
+
- A request that never reached the target is reported `error`, not `negative`.
|
|
2197
|
+
- Runs that build no probes at all exit non-zero and say so, instead of ending
|
|
2198
|
+
in silence and exit 0 — which read exactly like a target that came back clean.
|
|
2199
|
+
- The OOB DNS listener no longer dies on a malformed query, and write failures
|
|
2200
|
+
surface instead of being swallowed.
|
|
2201
|
+
|
|
2202
|
+
### Fixed — safety and audit
|
|
2203
|
+
|
|
2204
|
+
- **Multi-step chains now carry the same safeguards as single requests.** The
|
|
2205
|
+
chain path delivered to live targets without sink-shape filters, destructive
|
|
2206
|
+
hold-back or a pre-flight plan, so `--verify-active-risk stateful` fired
|
|
2207
|
+
persistence and irreversible file operations that `--verify-url` refuses to
|
|
2208
|
+
send without `--verify-allow-destructive`. Both paths now share one hold-back
|
|
2209
|
+
and print the same plan.
|
|
2210
|
+
- **The audit trail redacts credentials** — it records that a credential header
|
|
2211
|
+
was sent, never its value.
|
|
2212
|
+
- A capture carrying `Authorization` or `Cookie` over plain `http` is flagged
|
|
2213
|
+
before anything is sent.
|
|
2214
|
+
|
|
2215
|
+
### Added
|
|
2216
|
+
|
|
2217
|
+
- **The payload corpus is embedded in `rcekit.py`**, so the single file runs on
|
|
2218
|
+
its own — a jump box, an air-gapped host, a bare `curl` of the raw script.
|
|
2219
|
+
Resolution order is `--template-file` → `templates/payloads.json` beside the
|
|
2220
|
+
script → the built-in copy, and falling back to the built-in copy is
|
|
2221
|
+
announced. A corpus that exists but does not parse still hard-fails: that
|
|
2222
|
+
check exists for truncated and tampered corpora. `tools/embed_corpus.py`
|
|
2223
|
+
regenerates the embedded copy, and the test suite fails if the two drift.
|
|
2224
|
+
- **`--insecure`** skips TLS verification for internal targets with self-signed
|
|
2225
|
+
or mismatched certificates — opt-in and explicit, like `curl -k`.
|
|
2226
|
+
- **`--sink-raw`** for whole-command injection sinks, also readable from a
|
|
2227
|
+
target profile.
|
|
2228
|
+
- **`--separators`** to narrow the break-out sweep once the sink's shape is
|
|
2229
|
+
known.
|
|
2230
|
+
- **Documentation split into a task-oriented tree.** The README is half its
|
|
2231
|
+
former length and now leads with what RCEKit is for:
|
|
2232
|
+
[field guide](docs/guide.md) (worked examples by situation),
|
|
2233
|
+
[payload generation & exports](docs/generation.md), and
|
|
2234
|
+
[reference](docs/reference.md) (every flag grouped by task, plus the full
|
|
2235
|
+
taxonomies and exit codes).
|
|
2236
|
+
- **A "How RCEKit compares" section** covering commix, SSTImap, Nuclei and
|
|
2237
|
+
interactsh, with every claim traceable to that project's own documentation.
|
|
2238
|
+
- Four confirmation demos against real, publicly documented CVEs (Webmin
|
|
2239
|
+
CVE-2019-15107, Struts2 S2-001, Log4Shell CVE-2021-44228).
|
|
2240
|
+
|
|
2241
|
+
### Changed
|
|
2242
|
+
|
|
2243
|
+
- **Expect more requests per run.** The separator sweep and the language-runtime
|
|
2244
|
+
fix both widen the probe set. Narrow with `--separators`, `--contexts` and
|
|
2245
|
+
`--environments` once the sink's shape is known.
|
|
2246
|
+
- **`--doctor` output.** Its first line now names the corpus in use
|
|
2247
|
+
(`corpus: …`) rather than a path (`template: …`), since the corpus is no
|
|
2248
|
+
longer necessarily a file, and `[ok] file loaded and parsed` is now
|
|
2249
|
+
`[ok] corpus loaded and parsed`.
|
|
2250
|
+
|
|
2251
|
+
No breaking changes to the CLI, output formats, or the template schema.
|
|
2252
|
+
Standard library only, Python 3.8–3.13.
|
|
2253
|
+
|
|
2254
|
+
## Earlier releases
|
|
2255
|
+
|
|
2256
|
+
Release notes for these live on the
|
|
2257
|
+
[Releases page](https://github.com/kabiri-labs/rcekit/releases); they predate
|
|
2258
|
+
this file and have not been restated here.
|
|
2259
|
+
|
|
2260
|
+
- **[2.15.2]** — Multi-method RCE detection & confirmation
|
|
2261
|
+
- **[2.7.0]**
|
|
2262
|
+
- **[2.1.0]**
|
|
2263
|
+
|
|
2264
|
+
|
|
2265
|
+
|
|
2266
|
+
[Unreleased]: https://github.com/kabiri-labs/rcekit/compare/v2.36.0...HEAD
|
|
2267
|
+
[2.45.3]: https://github.com/kabiri-labs/rcekit/compare/v2.45.2...v2.45.3
|
|
2268
|
+
[2.45.2]: https://github.com/kabiri-labs/rcekit/compare/v2.45.1...v2.45.2
|
|
2269
|
+
[2.45.1]: https://github.com/kabiri-labs/rcekit/compare/v2.45.0...v2.45.1
|
|
2270
|
+
[2.45.0]: https://github.com/kabiri-labs/rcekit/compare/v2.44.0...v2.45.0
|
|
2271
|
+
[2.44.0]: https://github.com/kabiri-labs/rcekit/compare/v2.43.0...v2.44.0
|
|
2272
|
+
[2.43.0]: https://github.com/kabiri-labs/rcekit/compare/v2.42.0...v2.43.0
|
|
2273
|
+
[2.42.0]: https://github.com/kabiri-labs/rcekit/compare/v2.41.0...v2.42.0
|
|
2274
|
+
[2.41.0]: https://github.com/kabiri-labs/rcekit/compare/v2.40.0...v2.41.0
|
|
2275
|
+
[2.40.0]: https://github.com/kabiri-labs/rcekit/compare/v2.39.0...v2.40.0
|
|
2276
|
+
[2.39.0]: https://github.com/kabiri-labs/rcekit/compare/v2.38.0...v2.39.0
|
|
2277
|
+
[2.38.0]: https://github.com/kabiri-labs/rcekit/compare/v2.37.0...v2.38.0
|
|
2278
|
+
[2.37.0]: https://github.com/kabiri-labs/rcekit/compare/v2.36.0...v2.37.0
|
|
2279
|
+
[2.36.0]: https://github.com/kabiri-labs/rcekit/compare/v2.35.5...v2.36.0
|
|
2280
|
+
[2.35.5]: https://github.com/kabiri-labs/rcekit/compare/v2.35.4...v2.35.5
|
|
2281
|
+
[2.35.4]: https://github.com/kabiri-labs/rcekit/compare/v2.35.3...v2.35.4
|
|
2282
|
+
[2.35.3]: https://github.com/kabiri-labs/rcekit/compare/v2.35.2...v2.35.3
|
|
2283
|
+
[2.35.2]: https://github.com/kabiri-labs/rcekit/compare/v2.35.1...v2.35.2
|
|
2284
|
+
[2.35.1]: https://github.com/kabiri-labs/rcekit/compare/v2.35.0...v2.35.1
|
|
2285
|
+
[2.35.0]: https://github.com/kabiri-labs/rcekit/compare/v2.34.1...v2.35.0
|
|
2286
|
+
[2.34.1]: https://github.com/kabiri-labs/rcekit/compare/v2.34.0...v2.34.1
|
|
2287
|
+
[2.34.0]: https://github.com/kabiri-labs/rcekit/compare/v2.33.0...v2.34.0
|
|
2288
|
+
[2.33.0]: https://github.com/kabiri-labs/rcekit/compare/v2.32.0...v2.33.0
|
|
2289
|
+
[2.32.0]: https://github.com/kabiri-labs/rcekit/compare/v2.31.0...v2.32.0
|
|
2290
|
+
[2.31.0]: https://github.com/kabiri-labs/rcekit/compare/v2.30.0...v2.31.0
|
|
2291
|
+
[2.30.0]: https://github.com/kabiri-labs/rcekit/compare/v2.29.0...v2.30.0
|
|
2292
|
+
[2.29.0]: https://github.com/kabiri-labs/rcekit/compare/v2.28.0...v2.29.0
|
|
2293
|
+
[2.28.0]: https://github.com/kabiri-labs/rcekit/compare/v2.27.0...v2.28.0
|
|
2294
|
+
[2.27.0]: https://github.com/kabiri-labs/rcekit/compare/v2.26.0...v2.27.0
|
|
2295
|
+
[2.26.0]: https://github.com/kabiri-labs/rcekit/compare/v2.25.0...v2.26.0
|
|
2296
|
+
[2.25.0]: https://github.com/kabiri-labs/rcekit/compare/v2.24.0...v2.25.0
|
|
2297
|
+
[2.24.0]: https://github.com/kabiri-labs/rcekit/compare/v2.23.3...v2.24.0
|
|
2298
|
+
[2.23.3]: https://github.com/kabiri-labs/rcekit/compare/v2.23.2...v2.23.3
|
|
2299
|
+
[2.23.2]: https://github.com/kabiri-labs/rcekit/compare/v2.23.1...v2.23.2
|
|
2300
|
+
[2.23.1]: https://github.com/kabiri-labs/rcekit/compare/v2.23.0...v2.23.1
|
|
2301
|
+
[2.23.0]: https://github.com/kabiri-labs/rcekit/compare/v2.22.0...v2.23.0
|
|
2302
|
+
[2.22.0]: https://github.com/kabiri-labs/rcekit/compare/v2.21.1...v2.22.0
|
|
2303
|
+
[2.21.1]: https://github.com/kabiri-labs/rcekit/compare/v2.15.2...v2.21.1
|
|
2304
|
+
[2.15.2]: https://github.com/kabiri-labs/rcekit/releases/tag/v2.15.2
|
|
2305
|
+
[2.7.0]: https://github.com/kabiri-labs/rcekit/releases/tag/v2.7.0
|
|
2306
|
+
[2.1.0]: https://github.com/kabiri-labs/rcekit/releases/tag/v2.1.0
|