rcekit 2.35.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,1002 @@
1
+ # Changelog
2
+
3
+ All notable changes to RCEKit are recorded here. The format follows
4
+ [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and RCEKit follows
5
+ [Semantic Versioning](https://semver.org/spec/v2.0.0.html): PATCH for fixes,
6
+ MINOR for new capabilities, MAJOR for breaking changes to the CLI, output
7
+ formats, or the template schema.
8
+
9
+ ## [Unreleased]
10
+
11
+ ## [2.35.0] — 2026-08-20
12
+
13
+ ### Added
14
+
15
+ - **RCEKit is installable from PyPI**: `pipx install rcekit` (or
16
+ `pip install rcekit`) puts an `rcekit` command on PATH. Published through PyPI
17
+ Trusted Publishing from a GitHub release — no API token, no repository secret.
18
+
19
+ It ships as a **single-module distribution**, not a package tree. `rcekit.py`
20
+ stays one file at the repo root and still runs alone from a `curl` on a jump
21
+ box or an air-gapped host; installing is a second supported shape, not a
22
+ replacement for the first.
23
+
24
+ ### Changed
25
+
26
+ - **`import rcekit` no longer has side effects.** Logging was configured at
27
+ module scope, and `logging.FileHandler` opens its file when it is constructed,
28
+ so merely importing the module wrote `rcekit.log` into whatever directory the
29
+ interpreter happened to be in. Handler setup moved into `configure_logging()`,
30
+ called from `main()`. Running the CLI still writes `rcekit.log` exactly as
31
+ before.
32
+
33
+ - **`main()` takes an optional `argv`** and returns an explicit `int`, so the
34
+ console-script entry point is a plain zero-argument call and tests can drive
35
+ the CLI in-process. Every exit code is unchanged.
36
+
37
+ - **The built-in corpus is no longer reported as a missing file.** With nothing
38
+ but `rcekit.py` — an installed wheel, or the single-file copy — there is no
39
+ `templates/` directory, the embedded corpus *is* the corpus, and the run is
40
+ now silent about it; `--doctor` names it `built-in (embedded in rcekit.py)`
41
+ and reports OK. A `templates/` directory that exists *without* its
42
+ `payloads.json` still prints the notice, because that one is a real finding.
43
+ A corpus that is present but corrupt, and an explicit `--template-file` that
44
+ is missing or corrupt, still refuse to run and exit non-zero.
45
+
46
+ - The log file handler now runs at `DEBUG` while the console stays at `INFO`, so
47
+ detail worth having when reconstructing a run no longer lands on the
48
+ operator's terminal.
49
+
50
+ ## [2.34.1] — 2026-08-17
51
+
52
+ Documentation only; no behaviour change.
53
+
54
+ ### Changed
55
+
56
+ - **New tagline: "`confirmed` means the target executed the input. `negative`
57
+ means the probes reached it."** The old one — "prove RCE, don't guess it" —
58
+ claimed the tool always proves. It does not, and does not need to: the value
59
+ is that each verdict has a mechanical meaning, in both directions. A promise
60
+ can be broken; a definition cannot.
61
+
62
+ Both halves name their actor and object on purpose. "Executed" alone reads as
63
+ though *RCEKit* executed something; the claim is about the target. And
64
+ `confirmed`/`negative` are the verdict values as the code spells them, not
65
+ looser words like "clean".
66
+
67
+ - **The README caught up with the engine.** It still described "two verdict
68
+ tiers" when there are seven, and the comparison table predated the last five
69
+ releases. Rewritten around what a verdict asserts, with the tier table as the
70
+ centrepiece and a section on the half no other tool has: `error` and
71
+ `nothing-tested` exist so a run that tested nothing is never reported as
72
+ clean.
73
+
74
+ The comparison table gains the classes added since it was written — Windows
75
+ `cmd.exe`/PowerShell sinks, upload → write-then-execute, second-order
76
+ execution, query-language bridges and deserialization sinks — and the "reach
77
+ for something else" note now says plainly that sqlmap owns the database and
78
+ RCEKit's bridges only prove the OS is reachable from a text parameter.
79
+
80
+ ### Fixed
81
+
82
+ - Two claims in the engagement-controls table were wrong and are now accurate:
83
+ the observed-channel fetch sends **no** credentials unless given a request with
84
+ `--observe-request` (only the `file` read-back inherits them, same-origin), and
85
+ an unanswered `--observe-url` is a warning about a partly blinded run rather
86
+ than a `nothing-tested` verdict.
87
+ - The "mechanisms that produce `inconclusive`" list said four and listed five.
88
+
89
+ ## [2.34.0] — 2026-08-17
90
+
91
+ Deserialization **sink** detection, and a verdict that is deliberately not RCE.
92
+
93
+ Deserialization RCE (fastjson, shiro, weblogic, jenkins) cannot be confirmed by
94
+ the value-oracle model: the payload is a serialized object graph and gadget
95
+ selection is classpath-specific, so whether execution is reachable depends on
96
+ jars RCEKit cannot see. That stays out of scope. The honest middle step is
97
+ showing the endpoint parses the data at all — a real finding, and the
98
+ prerequisite for every gadget chain.
99
+
100
+ ### Added
101
+
102
+ - **`--methods deser`**, which **never emits `confirmed`**. Its strongest
103
+ outcome is a new verdict, `deserialization-sink`, reported in its own section
104
+ that states outright that reaching RCE from there depends on classpath
105
+ gadgets. Collapsing it into `confirmed` would break the one guarantee the tool
106
+ rests on; collapsing it into `needs-review` would throw away a proven finding.
107
+
108
+ `deserialization-sink` sits below both RCE tiers in the collapsed verdict: it
109
+ is proven, but a *suspected* RCE outranks a proven non-RCE in triage.
110
+
111
+ - **Two oracles of deliberately different strength.** `shape` (no listener
112
+ needed) sends a well-formed object stream, the same stream truncated, and the
113
+ format's magic bytes plus random noise of the same length, and asks whether
114
+ the endpoint answers the well-formed one differently from both — a
115
+ fingerprint, so `needs-review` only, never promoted. `dns` (needs
116
+ `--oob-host`) sends a gadget whose only side effect is a name lookup.
117
+
118
+ - **A URLDNS builder for Java serialization.** A `HashMap` holding one
119
+ `java.net.URL`: `HashMap.readObject` hashes the key, `URL.hashCode` asks for
120
+ the host address, the JVM resolves the name. It references no class outside
121
+ `java.util`/`java.net`, so there is nothing in it that can run — the callback
122
+ proves the object graph was reconstructed and no more.
123
+
124
+ Built in Python rather than declared in the corpus because the URL host is
125
+ length-prefixed *inside* the stream and changes per probe. Its constant parts
126
+ are the exact bytes OpenJDK's own `ObjectOutputStream` emits for that graph,
127
+ and the result was **verified against OpenJDK 21**: it deserializes to
128
+ `HashMap{http://<host>/=rk}` and issues a DNS query for `<host>`, with no code
129
+ execution.
130
+
131
+ - **A `deser_probes` corpus section** with `java`, `php`, `dotnet`,
132
+ `python_pickle` and `fastjson`, plus `--deser-formats` to narrow it. Only
133
+ `java` and `fastjson` carry a DNS gadget: PHP and .NET chains all run through
134
+ magic methods or type confusion, so there is no honest DNS-only probe for them
135
+ and they get the shape oracle alone.
136
+
137
+ - Response signatures for the shape differential drop long digit and hex runs,
138
+ so request ids and timestamps on an otherwise identical error page do not make
139
+ every endpoint fingerprint as a parser.
140
+
141
+ ### Changed
142
+
143
+ - The README's scope note now says precisely what changed and what did not:
144
+ deserialization **gadget chains** remain out of scope, while the **sink** is
145
+ now reported in its own tier.
146
+
147
+ ## [2.33.0] — 2026-08-17
148
+
149
+ Query-language bridges. Several RCEs pass through a query language before
150
+ reaching the OS — Postgres `COPY … FROM PROGRAM`, MSSQL `xp_cmdshell`, XXE
151
+ `expect://` — and the injection point is an ordinary text value, so the oracle
152
+ model already fitted. Only the carriers were missing.
153
+
154
+ ### Added
155
+
156
+ - **A `bridges` section in the corpus**, declared like `eval_carriers` so
157
+ coverage grows without touching Python. Each bridge names the shell it
158
+ reaches, its safety tier, its prerequisites and, where it creates something,
159
+ the statement that removes it: `postgres_copy_program` (`/bin/sh`,
160
+ `stateful`), `mssql_xp_cmdshell` (`cmd.exe`, `intrusive`) and `xxe_expect`
161
+ (`/bin/sh`, `intrusive`).
162
+
163
+ - **`--bridges none|auto|NAMES`** rides the command probes through them. A
164
+ bridge is a **carrier, not an oracle**: it wraps the command `reflected`,
165
+ `time` and `oob` already build, so those methods prove execution through it
166
+ and inherit every tier guarantee rather than re-deriving one. Off by default,
167
+ because a bridge payload is SQL or XML syntax and on an ordinary shell sink it
168
+ is a request that cannot confirm.
169
+
170
+ Three properties follow from that framing. A bridge only gets a core written
171
+ in its own dialect — `xp_cmdshell` hands its argument to `cmd.exe`, so pairing
172
+ it with a POSIX `$((a+b))` would send inert text. No separator is prepended:
173
+ inside `COPY … FROM PROGRAM '…'` there is no running command to break out of.
174
+ And the record's context still applies, so `--contexts sql` and a bridge
175
+ compose instead of each reinventing the other.
176
+
177
+ - **The safety ordering governs bridges** exactly as it governs every corpus
178
+ payload: a `stateful` bridge needs `--verify-active-risk stateful`, and the
179
+ pre-flight names the tier each held-back bridge actually requires rather than
180
+ sending the operator to raise the ceiling further than the run needs.
181
+
182
+ ### Changed
183
+
184
+ - **An aggregate method's result now carries its cleanup line.** `time` reports
185
+ one row for a whole probe series, so a stateful bridge on the one oracle that
186
+ reliably proves a query-language sink was the one that never said how to clean
187
+ up after itself.
188
+
189
+ ### Not built, deliberately
190
+
191
+ - **MySQL UDF execution** is a multi-stage chain — write a shared object into the
192
+ plugin directory, then `CREATE FUNCTION` — not something a single probe can
193
+ carry. There is no stub for it.
194
+ - **MongoDB `$where`** is a boolean-only channel (its JS sandbox cannot reach a
195
+ shell), so it needs a different oracle rather than this one.
196
+ `mongo-express/CVE-2019-10758` is a plain JS `eval` sink that `--methods eval`
197
+ already covers.
198
+
199
+ The three shipped bridges are documented syntax but **not validated against live
200
+ databases here** — this build environment has no container runtime. Each corpus
201
+ entry says so in its `verified` field rather than implying a test that did not
202
+ happen.
203
+
204
+ ## [2.32.0] — 2026-08-17
205
+
206
+ The second-order oracle. Execution frequently happens on a **different request**
207
+ than injection — stored SSTI rendered on a profile page, a payload written to a
208
+ log a template engine later renders, a queued job run asynchronously. The engine
209
+ diffs the response it injected into, so every one of those read `negative`
210
+ however exploitable the target was.
211
+
212
+ ### Added
213
+
214
+ - **`--observe-url URL`** names the endpoint where the execution surfaces. It is
215
+ read after each probe and then polled after the batch, and a probe whose
216
+ computed value turns up there is upgraded to `confirmed`.
217
+
218
+ It stays fully differential, which is why it reaches `confirmed` rather than
219
+ `needs-review`: the value was computed locally from operands random to that
220
+ probe, it must be absent from a snapshot of the endpoint taken **before any
221
+ probe was sent**, and — the rule that carries the weight — a probe's value is
222
+ looked for there **only when the probe's own payload does not contain it**.
223
+
224
+ Without that last rule the oracle would be a false-positive generator: `file`
225
+ and `oob` expect a random token that sits verbatim in the payload, so a target
226
+ that merely stores the payload and renders it back would hand that token
227
+ straight to the observed page and every such probe would confirm without
228
+ executing anything. Measured against a store-and-echo target: **0**
229
+ confirmations. The computed-value methods pass the same rule for the opposite
230
+ reason — reflection returns `$((a+b))`, never the sum — so it selects them
231
+ without naming them, and a method added later inherits the right answer.
232
+
233
+ - **`--observe-request FILE`** takes a captured request instead, for the common
234
+ case where the page a stored payload renders on is behind a login. It needs no
235
+ `FUZZ` marker: the observed endpoint is read, never injected into.
236
+
237
+ - **`--observe-poll` / `--observe-timeout`** control the polling window
238
+ (defaults 5s and 60s). One poll always happens, even at a zero timeout.
239
+
240
+ - Every probe result carries an `observe_status` in `--detect-json`:
241
+ `confirmed`, `polled` (read, value not there), `in-control`, `not-observed`
242
+ (not eligible) or `unreachable`. When the endpoint never answered, the run
243
+ says so outright — negatives decided without ever reading the observed channel
244
+ are not second-order negatives.
245
+
246
+ ### Changed
247
+
248
+ - The observed channel is read once after **each** probe as well as polled after
249
+ the batch, so a run with `--observe-url` sends roughly twice the requests.
250
+ Batch-then-poll alone is only correct for a channel that *accumulates* (a log,
251
+ a comment list); where the store overwrites — a profile field, which is the
252
+ shape this oracle most exists for — every probe but the last is gone by the
253
+ time the batch poll runs, and the oracle confirmed nothing. The extra read is
254
+ skipped for probes that are already confirmed in-band or not eligible, so
255
+ `file` and `oob` add none.
256
+
257
+ Observing is additive throughout: the in-band verdict is computed exactly as
258
+ before and only a non-`confirmed` one can be upgraded, so a run without the flag
259
+ is byte-for-byte unchanged and a run with it can only gain findings.
260
+
261
+ ## [2.31.0] — 2026-08-17
262
+
263
+ The `write` method: a write primitive proven to be RCE by executing what it
264
+ wrote. A whole family of targets was invisible — `tomcat/CVE-2017-12615` (PUT a
265
+ JSP), `activemq/CVE-2016-3088`, `weblogic/CVE-2018-2894` — because the vulnerable
266
+ request *stores a file* rather than evaluating anything. Nothing is computed in
267
+ its response, so `reflected` and `eval` correctly returned `negative` on targets
268
+ that are fully exploitable.
269
+
270
+ ### Added
271
+
272
+ - **`--methods write`** — the inverse of `file`. `file` assumes execution exists
273
+ and uses a write as proof of it; `write` assumes a write primitive exists and
274
+ uses execution of the written file as proof of RCE. The probe is the file's
275
+ *content*: a one-liner computing a product on random operands, delivered
276
+ through the ordinary injection point.
277
+
278
+ The fetched file is read in three tiers, and the middle one is the reason the
279
+ method exists:
280
+
281
+ | fetched file contains | verdict | means |
282
+ |---|---|---|
283
+ | the product | `confirmed` | written **and** executed |
284
+ | the one-liner, verbatim | `needs-review` | arbitrary file write, not interpreted |
285
+ | neither | `negative` | no write, or not served there |
286
+
287
+ An upload directory that is served but not interpreted is a real finding and
288
+ is not remote code execution, so the tiers are never merged.
289
+
290
+ - **`--write-url-template URL`** names where the stored file is served — the
291
+ channel the proof comes back on, and the flag the method is gated on.
292
+
293
+ - **`--write-lang`** picks the file types: `auto` (default) reads the extension
294
+ off the read-back URL, or name any of `jsp`, `jspx`, `php`, `aspx`, `erb`.
295
+ `jsp`/`aspx`/`erb` share the `<%= %>` delimiters, so their probes are
296
+ byte-identical and cost one request between them; with no extension to read,
297
+ `auto` writes all five in three requests.
298
+
299
+ ### Changed
300
+
301
+ - **A `needs-review` finding now prints its cleanup line too.** It used to
302
+ appear only under `confirmed`, which was already thin and is wrong for this
303
+ method: a `write` reaching `needs-review` means the file *is* on the target,
304
+ just not interpreted, so the artifact would have been left there unmentioned.
305
+
306
+ - The write method's operands are drawn once per run rather than once per
307
+ carrier, so the file is written once instead of once for each of the ~13
308
+ `(environment, context)` carriers. For a state-changing method that is not a
309
+ request-count saving, it is a blast radius. Still fresh per run, which is what
310
+ makes the product unforgeable.
311
+
312
+ - The write method declines the break-out contexts (`sql`, `javascript`,
313
+ `shell_*`, …) and keeps the transport ones. Its payload is a whole file body:
314
+ there is nothing to break out of, and wrapping it in `'; … -- ` would write a
315
+ broken file. A run narrowed past `raw` and the transport contexts is told so
316
+ rather than reporting a clean negative.
317
+
318
+ ## [2.30.0] — 2026-08-17
319
+
320
+ Per-dialect shell probes. `$((a+b))`, `sleep` and `$(echo TAG)` are POSIX
321
+ constructs: on a cmd.exe or PowerShell sink they are inert literal text. The
322
+ dialect was inferred from the corpus environment alone, so a run could send a
323
+ probe no shell on the target would ever execute — including on the carrier whose
324
+ context is literally named `powershell`.
325
+
326
+ ### Added
327
+
328
+ - **`--sink-env auto|unix|windows|powershell`** states which shell runs the
329
+ injected command. The computed-value core, the separators and the break-out
330
+ contexts are all chosen from it. `auto` (the default) infers it per carrier;
331
+ pin it when the corpus environment names the *application runtime* rather than
332
+ the OS — `--environments php --sink-env windows` is a PHP application on IIS,
333
+ which no inference can see.
334
+
335
+ - **A PowerShell probe shape for every shell method**, validated against
336
+ pwsh 7.4: `Write-Output T1$(a*b)T2` for `reflected` (an unquoted argument is an
337
+ expandable string, so the core carries no quote and the quote-wrapping
338
+ contexts can still carry it), `Start-Sleep -Milliseconds N` for `time`,
339
+ `Set-Content` for `file` and `iwr -useb` for `oob`. PowerShell was previously
340
+ reachable by no probe in any method.
341
+
342
+ - **cmd.exe and PowerShell carriers for the `dotnet` environment.** It is the
343
+ one corpus environment that names a platform, and it was taking the POSIX
344
+ shape — so .NET on Windows, the case the environment exists for, was the case
345
+ it could not confirm on. Every other runtime keeps the POSIX shape: a language
346
+ does not say which OS it runs on.
347
+
348
+ ### Fixed
349
+
350
+ - **The `powershell` carrier was written in cmd.exe.** Every `windows` carrier
351
+ took the `for /f ... ('set /a a+b')` core regardless of context, so the one
352
+ carrier explicitly shaped for PowerShell sent a payload PowerShell cannot
353
+ execute. The dialect now follows the carrier's context first, and a carrier's
354
+ break-out variants stay in its dialect rather than re-deriving from the
355
+ environment.
356
+
357
+ - **`Set-Content`, not `>`, for the PowerShell write.** In Windows PowerShell
358
+ 5.1 the redirect is `Out-File`, whose default encoding is UTF-16LE: the write
359
+ lands and the read-back still does not find the token, so the probe reports
360
+ negative on a target it owns.
361
+
362
+ ### Changed
363
+
364
+ - **cmd.exe no longer gets the `sq`, `dq` and `subshell` carriers.** It has
365
+ neither a comment character to swallow the sink's tail nor a
366
+ command-substitution syntax, so those four carriers per Windows run were
367
+ requests that could only come back negative. PowerShell takes the quote
368
+ break-outs and `$( )` — both measured — but not the backtick, which is its
369
+ escape character rather than a substitution.
370
+
371
+ - **PowerShell's separator sweep carries no pipe.** `cmd | Start-Sleep
372
+ -Milliseconds 500` is a parameter-binding error, not a fresh command with
373
+ stdin attached the way a POSIX pipe is, and it fails that way for every cmdlet
374
+ the probes use. `;`, a newline and (on PowerShell 7) `&&`/`||` remain.
375
+
376
+ - The pre-flight plan prints the sink shell alongside the sink shapes, and a
377
+ pinned dialect narrows the printed ladder to the rungs it has syntax for.
378
+
379
+ ## [2.29.0] — 2026-08-17
380
+
381
+ Generalised read-back for the `file` method. It required a writable **web root**
382
+ the tester already knew, which ruled out every other way a target can hand a
383
+ file back — on exactly the internal, no-egress targets the method exists for.
384
+
385
+ ### Added
386
+
387
+ - **`--file-write-path DIR` + `--file-read-url URL`** name the two halves of the
388
+ read-back channel directly, so an LFI endpoint, a download or export handler,
389
+ an attachment fetcher or a `/tmp`-backed preview all work. The template takes
390
+ `{name}` (the filename), `{path}` (the full server-side path) and `{path_enc}`
391
+ (that path percent-encoded); only those three are substituted, so a URL that
392
+ legitimately contains braces survives unchanged.
393
+
394
+ Measured against a target with a download handler and nothing serving the
395
+ write directory: the web-root form confirms **0** — reporting an exploitable
396
+ target clean — and the general form confirms **7**.
397
+
398
+ ### Fixed
399
+
400
+ - **The read-back fetch now carries the run's headers**, so an authenticated
401
+ download, export, attachment or LFI handler can actually be read. It went out
402
+ bare, which barely mattered while the channel had to be a web root — static
403
+ file serving is rarely authenticated — and became the likely case the moment
404
+ the channel could be an application endpoint. Measured against a handler
405
+ behind a bearer token: the write executed on every probe and the verdict was
406
+ `negative`, "token absent from the fetched file". Now 7 confirmations on the
407
+ same target.
408
+ - **Credentials are carried only to the same origin.** A read-back URL on
409
+ another host is someone else's server, and replaying the target's session
410
+ cookie or bearer token to it would leak the credential, so those headers are
411
+ dropped there while the rest still go — and the run says so, because the
412
+ symptom would otherwise look like a clean target. `Content-Type` and
413
+ `Content-Length` are dropped from the fetch too: they describe a body the GET
414
+ does not have.
415
+
416
+ ### Changed
417
+
418
+ - **`--webroot` / `--web-base-url` are now the web-root alias** for the general
419
+ form: a web root is just the case where the read URL is the base plus the
420
+ filename. Existing command lines are unaffected. Both are resolved in one
421
+ place inside the method, so the alias and the general form cannot drift — and
422
+ the gate, the pre-flight banner and the blind-sink advice all ask that same
423
+ resolver instead of testing for the webroot pair.
424
+ - `blind_sink_advice` reads its flags defensively, so an args-like object
425
+ missing a newer field costs a line of advice rather than a traceback.
426
+
427
+ ## [2.28.0] — 2026-08-17
428
+
429
+ Injection-point enumeration. `-p NAME` needed the tester to already know which
430
+ parameter was the sink, so a capture's other candidates — including the headers
431
+ and nested JSON leaves that carry some of the highest-value classes — were never
432
+ tried.
433
+
434
+ ### Added
435
+
436
+ - **`-p all` / `--auto-params KINDS`** expands one captured request into every
437
+ candidate injection point and runs the selected `--methods` against each.
438
+ Query values, JSON leaves addressed by path (`user.profile.name`, `tags[1]`),
439
+ form fields, cookie crumbs and headers, each rewritten in **its own**
440
+ serialization rather than blanket-encoded. Verified end to end: a sink
441
+ reachable only through `User-Agent` is confirmed from `-r request.txt -p all`
442
+ with no manual header selection.
443
+ - **`--point-order fast|thorough`** — `fast` tries a curated high-yield header
444
+ list (the headers real published RCEs inject through); `thorough` adds every
445
+ remaining non-hop-by-hop header. **`--max-points N`** bounds the run and
446
+ reports what it dropped. **`--include-path-segments`** is opt-in, because
447
+ rewriting a path segment usually just produces a 404.
448
+ - **The run states its cost before sending it** —
449
+ `6 points x ~61 probes = at least 372 requests` — via a new
450
+ `estimate_detection_probes`, which builds the probes and counts them without
451
+ firing any. Enumeration multiplies an already-laddered probe count by the
452
+ candidate count, and an operator on a monitored engagement has to see that
453
+ before it happens rather than infer it from the traffic.
454
+ - **Findings name the point they came from**: `[reflected/unix/raw] at header
455
+ 'User-Agent' ...`.
456
+
457
+ ### Changed
458
+
459
+ - **Each candidate carries its own payload-free control.** Differencing a header
460
+ probe against a query probe's control would compare two different responses
461
+ and prove nothing.
462
+ - **Cheap methods run first per candidate, and a candidate stops at its first
463
+ confirmation.** `reflected` and `eval` cost one response each; `time` sleeps
464
+ and `oob` waits for a callback, and on a candidate that has already proven
465
+ execution those buy a second name for the same finding. Candidates that stay
466
+ clean still get every method, and single-point runs are unchanged.
467
+ - A JSON leaf is **replaced, never created**. Assigning to a missing key would
468
+ have injected into a field the application never sends — a probe that cannot
469
+ say anything about the parameter that does exist. Caught by its own test.
470
+ - **JSON points are addressed by tokens, not by a joined path string.** A key may
471
+ itself contain the separator: `{"user.name": ..., "user": {"name": ...}}`
472
+ rendered *both* leaves as `user.name`, so the literal key was never probed and
473
+ both candidates mutated the nested field — a false negative and a misattributed
474
+ finding at once. Tokens remove the ambiguity, and the display form
475
+ bracket-quotes such a key (`["user.name"]`) so the two stay distinguishable on
476
+ screen.
477
+ - **A deeply nested captured body no longer ends `-p all` with a traceback.**
478
+ `json.loads` recurses in C, so `RecursionError` joins the caught exceptions in
479
+ both the enumerator and the placer, as it already had in the response-channel
480
+ parser. The body yields no candidates; the rest of the request still enumerates.
481
+ - **The cost estimate honours `--max-payloads`.** It counted every probe the
482
+ carriers could produce while the run stops at the cap, so the figure was wrong
483
+ exactly when the operator had reached for the budget guard.
484
+ - `Host`, `Content-Length`, `Cookie` and the hop-by-hop headers are never
485
+ candidates: injecting into those changes the request's plumbing rather than
486
+ testing the application, and two of them are rebuilt by the delivery layer.
487
+
488
+ ## [2.27.0] — 2026-08-17
489
+
490
+ Engine carriers for the `eval` probe. Three template engines evaluate the
491
+ injected expression perfectly and still made RCEKit report `negative`, because
492
+ what came back was not the bare product the oracle searches for.
493
+
494
+ ### Added
495
+
496
+ - **`eval_carriers` in the corpus**, and `--eval-engines auto|<names>` to select
497
+ them. A carrier wraps the same random-operand arithmetic in an engine-specific
498
+ form; it never changes the oracle, and the bare probes still run first. Each
499
+ entry records `notes` (why it exists) and `verified` (what it was measured
500
+ against). Declarative, so a new carrier is a JSON entry rather than a code
501
+ change.
502
+
503
+ | Engine | Bare `${a*b}` returned | Carrier | Carrier returned |
504
+ |---|---|---|---|
505
+ | Freemarker | `2,070,761,401` (locale grouping) | `${(a*b)?c}` | `2070761401` |
506
+ | Velocity | `${a*b}` verbatim — a *reference*, not an expression | `#set($rk=a*b)$rk` | `2070761401` |
507
+ | Thymeleaf | `${a*b}` verbatim — needs inlining brackets | `[[${a*b}]]` | `2070761401` |
508
+
509
+ Measured against freemarker 2.3.32, velocity-engine-core 2.3 and thymeleaf
510
+ 3.1.2, running RCEKit's own generated probes through each engine: bare form
511
+ `CONFIRMS=no`, carrier `CONFIRMS=YES`, for all three.
512
+ - **The evidence line names the carrier** — `target computed '3979016' via the
513
+ freemarker carrier` — so a finding says which engine quirk it worked around.
514
+ A bare confirmation reads exactly as before.
515
+
516
+ ### Notes
517
+
518
+ - **Carriers are not sandbox escapes, and no sandbox-escape carrier ships.** The
519
+ premise that a sandboxed engine blocks the arithmetic probe did not survive
520
+ measurement: a member-access sandbox restricts method and field access, and
521
+ arithmetic needs neither. With OGNL member access denied for *everything*,
522
+ `40277*51413` still returned `2070761401` while `@java.lang.Math@max(1,2)` was
523
+ blocked; SpEL's restricted `SimpleEvaluationContext` and Jinja2's
524
+ `SandboxedEnvironment` behaved the same way. The bare probes already cover
525
+ those engines.
526
+ - The frequently-cited OGNL escape `(#_memberAccess=@ognl.OgnlContext@DEFAULT_MEMBER_ACCESS)`
527
+ additionally targets a field that **no longer exists in OGNL 3.3.4**, so on a
528
+ current engine it is a probe that can only come back negative.
529
+
530
+ ## [2.26.0] — 2026-08-16
531
+
532
+ The sink-shape ladder. An injected value lands in a *shape* — mid-command,
533
+ inside quotes, as the whole command — and the shape decides what can reach it.
534
+ Two shapes had no probe that fitted, so a genuinely exploitable target reported
535
+ clean.
536
+
537
+ ### Added
538
+
539
+ - **`--sink-shape auto|sep|raw|chain|newline|dq|sq|subshell`** (comma-separated)
540
+ names which shapes the shell probes try. `auto` is the whole ladder and the
541
+ default. Underneath it selects the existing separator sweep and break-out
542
+ contexts, so naming a rung narrows a supported run rather than switching on a
543
+ parallel path. The plan is printed before anything is sent, because the ladder
544
+ multiplies request count and an operator on a monitored engagement needs to
545
+ see the cost first.
546
+ - **The `subshell` rung — `$(...)` and backticks.** Reaches a value sitting
547
+ inside double quotes *without closing the quote*, which is the one case a
548
+ quoted break-out loses to a filter on the quote character itself. Measured
549
+ against `system("echo PING \"$input\"")`: with `"` stripped, `dq` is inert and
550
+ both substitution forms execute; with `$` stripped, `dq` executes and the
551
+ backtick form still does. Both ship because they survive different filters.
552
+
553
+ **Which method it helps is the counter-intuitive part.** `reflected`'s core is
554
+ `$((a+b))`, which the shell expands inside double quotes anyway, so that
555
+ method already confirmed there. The methods whose core must actually *run* —
556
+ `time` (a sleep), `file` (a redirect), `oob` (a fetch) — are completely inert
557
+ inside those quotes. On a quote-filtering sink, `--methods file` went from **0
558
+ confirmations to 2**: it had been reporting an exploitable target as clean.
559
+ - **The `raw` rung is now part of `auto`, for every shell method.** A
560
+ `qx/$input/`-style sink, where the input is the whole command, previously
561
+ needed `--sink-raw` — so it reported clean unless the operator already
562
+ suspected the shape. One extra probe per carrier buys it. `--sink-raw` keeps
563
+ its meaning as the narrowing alias for `--sink-shape raw`, and no existing
564
+ command line changes behaviour. `reflected`, `file`, `time` and `oob` all
565
+ build their candidates through one `_separator_candidates` helper, so a rung
566
+ cannot reach some methods and not others; `time` screens it in its second
567
+ wave, alongside the separators it holds back.
568
+
569
+ ### Fixed
570
+
571
+ - **A method that builds no probes no longer reports `negative`.** An aggregate
572
+ method asked to judge zero samples answers honestly — "no delay was observed",
573
+ "no callback arrived" — and that reads as "not vulnerable" from a run that
574
+ tested nothing. The engine now emits no row for a carrier that produced no
575
+ probes, which lets its own loud nothing-tested path fire instead. Reachable
576
+ through any narrowing that leaves a carrier with nothing to send.
577
+
578
+ ### Changed
579
+
580
+ - **The pre-flight sink-shape plan is computed from the effective run**, not
581
+ from the `--sink-shape` value. `--separators`, `--contexts` and `--sink-raw`
582
+ each narrow the ladder, so printing the flag described a run that would not
583
+ happen — and this output is presented as an audit of the traffic about to be
584
+ sent. `effective_sink_shapes` is the single source of truth the engine and the
585
+ plan both read.
586
+ - **The backtick context drops probe shapes that carry their own backtick.**
587
+ Backticks do not nest, so such a probe closes the outer substitution early and
588
+ could only ever come back negative. `$( )` does nest and keeps every shape.
589
+ - **Naming `--separators` now implies the sink is separator-led**, so the `raw`
590
+ rung is dropped unless `--sink-shape` names it explicitly. A profile with
591
+ `sink_needs_separator` drops it for the same reason. Both keep an explicitly
592
+ narrowed run from being widened behind the operator's back.
593
+
594
+ ## [2.25.0] — 2026-08-16
595
+
596
+ A coverage benchmark, so a claim about what RCEKit confirms can be checked
597
+ instead of asserted. The unit suite proves the tool reaches the right verdict
598
+ from a given response; it cannot prove it confirms Webmin.
599
+
600
+ ### Added
601
+
602
+ - **`--detect-json PATH`** writes a detection run as JSON: the run's overall
603
+ verdict, per-verdict counts, and every probe with its payload, method, context
604
+ and evidence. Text output is unchanged. This is the supported way to consume a
605
+ run programmatically — scraping stdout cannot be made reliable, because a
606
+ probe payload may contain a literal newline (the newline separator is a real
607
+ one, so line-oriented parsing splits a payload in half) and the detection path
608
+ exits 0 whether it confirmed or came back clean.
609
+ - **`tests/bench/` — the coverage benchmark harness.** Each case brings a real
610
+ vulnerable build up, runs RCEKit as an operator would, checks the verdict, and
611
+ tears it down; `--markdown` emits the coverage table. Not part of
612
+ `python -m unittest discover -s tests` — cases need Docker and pull real
613
+ images — so it runs by hand or in a dedicated job, and exits non-zero if any
614
+ case fails. Two cases ship, transcribed from `docs/verify-it-yourself.md`:
615
+ Webmin CVE-2019-15107 and Struts2 S2-001.
616
+ - **A negative control is a required key.** A benchmark without controls measures
617
+ nothing: a tool that shouted `confirmed` at every target would score full marks
618
+ on the vulnerable half. Three kinds are supported — a patched build, the same
619
+ target probed for the wrong class, and a weaker method that must stay below
620
+ `confirmed` on a target where it happens to be right. The runner refuses four
621
+ shapes of non-control: no control at all; one expecting `confirmed`; one that
622
+ runs the identical invocation against an identical target (judged on what it
623
+ would actually run, so an explicit copy of the vulnerable invocation is caught
624
+ as well as an omitted one); and one expecting `error` or `nothing-tested`,
625
+ since both mean the target was never exercised and such a control would stay
626
+ green with the detection engine entirely broken. Validation and execution
627
+ share one `control_plan` so they cannot drift.
628
+ - **`overall_detection_verdict`** collapses a run to one verdict, ordered by what
629
+ an operator must not miss rather than by frequency: one `confirmed` among a
630
+ hundred negatives is the finding. `error` is reported only when *nothing*
631
+ reached the target, and a run that built no probes is `nothing-tested` —
632
+ never `negative`, which would read as "not vulnerable".
633
+
634
+ ### Changed
635
+
636
+ - `CONTRIBUTING.md` asks for a bench case alongside new detection coverage, and
637
+ for the README table to state the tier the case actually reached.
638
+
639
+ ### Notes
640
+
641
+ - The two shipped cases have **not yet been executed through the harness** — it
642
+ was written where no Docker daemon was available. Their invocations come from
643
+ a documented, reproduced guide, but the case files themselves are unvalidated;
644
+ `tests/bench/README.md` says so and flags the one field that is a guess. No
645
+ README claim was changed to assert benchmark results.
646
+
647
+ ## [2.24.0] — 2026-08-16
648
+
649
+ The computed value is no longer looked for in the response body alone. A sink
650
+ whose output surfaces anywhere else in the response was reported `negative` — a
651
+ false negative on a class RCEKit already claims to cover, which is worse than a
652
+ missing class. The oracle, the random operands and the control differential are
653
+ unchanged; only the set of places searched is wider.
654
+
655
+ ### Added
656
+
657
+ - **Whole-response evidence search.** Every confirmation now sweeps the response
658
+ body, the application response headers, individual cookie values, the redirect
659
+ target RCEKit actually landed on, the HTTP reason phrase, and each leaf of a
660
+ parsed JSON body. Real sinks put command output in a debug header or a
661
+ `Set-Cookie`, and API targets surface an evaluator's result inside a nested
662
+ error envelope — `{"error": {"detail": "cannot render 2058898001"}}` — where a
663
+ substring search of the serialised body misses a value the encoder escaped.
664
+ - **The evidence line names the channel that carried the value**, e.g.
665
+ `target computed 'RK…' in header X-Cmd-Out (random operands, absent from
666
+ control)`, so the finding stays reproducible by hand. A body-carried
667
+ confirmation reads exactly as it did before.
668
+
669
+ ### Changed
670
+
671
+ - **The control differential now covers every channel, not just the body.** A
672
+ value present anywhere in the payload-free control is not attributable to
673
+ execution, so it yields `inconclusive` wherever it turned up. This is stricter
674
+ than comparing only the channel that matched, and it is what keeps a wider
675
+ search from becoming a looser verdict.
676
+ - **The `file` method's control check covers every channel too**, on the same
677
+ reasoning: its token is random, so its presence in any control channel means
678
+ it did not get there by being written and served.
679
+
680
+ ### Security
681
+
682
+ - **A deeply nested JSON response can no longer silence detection.** Channels are
683
+ built inside the delivery `try`/`except`, so a `RecursionError` while parsing
684
+ or walking the body escaped as a network failure: a response that arrived
685
+ perfectly well was reported "request never reached the target". Measured:
686
+ every one of the 46 probes in a default `reflected` run turned into `error`,
687
+ which a target could induce deliberately to hide a live sink behind a thousand
688
+ nested arrays. Version-independent, though the source moves — CPython 3.12
689
+ raised the C recursion limit its JSON scanner runs under, so on 3.12/3.13 the
690
+ parser survives a depth that breaks it on 3.8–3.11 and the recursive leaf walk
691
+ hit the ordinary Python limit instead. The walk is now iterative and
692
+ depth-capped, `RecursionError` from the parser costs the JSON channels only,
693
+ and building channels can never turn a delivered response into a delivery
694
+ failure.
695
+ - **Transport headers are excluded from the sweep.** `Content-Length`, `Date`,
696
+ `Age`, `ETag` and their neighbours are generated below the application and can
697
+ never carry a computed value, but they *are* numeric — and the `expr` probe's
698
+ expected value is a bare boundary-fenced number. Searching them would let a
699
+ byte count collide with an arithmetic result and read as execution. Locked in
700
+ by a test that puts the expected value in `Content-Length` and requires
701
+ `negative`.
702
+
703
+ ## [2.23.3] — 2026-08-04
704
+
705
+ Four items from the same review: requests and seconds spent on work that could
706
+ not produce a result. No verdict changes — the lab still confirms 15 of 15
707
+ vulnerable sinks with nothing on the clean five — the run just stops paying for
708
+ probes that were structurally unable to confirm.
709
+
710
+ ### Changed
711
+
712
+ - **The `awk` probe is no longer sent into a context that wraps the payload in
713
+ quotes.** It carries double quotes, so in `attribute` the quote closed early
714
+ and the rest was not a command: 5 requests per carrier that could only ever
715
+ come back negative. Measured on a verbose shell sink, that shape confirmed 8
716
+ times in `raw` and 0 times in `attribute`. Break-out contexts such as
717
+ `shell_double_quoted` *close* the sink's quote and comment its tail, so they
718
+ still get it. The same guard covers the PowerShell out-of-band shape.
719
+ - **The timing screen runs in two waves.** Every delayed screen probe costs a
720
+ real sleep, so screening all five separators up front spent `5 × base` seconds
721
+ on every carrier, including the ones that cannot break out at all. `; ` and
722
+ `| ` are screened first and the rest only if neither delayed — a sink that
723
+ filters both is still swept, it is just no longer the price everyone pays.
724
+ - **The out-of-band callback window is no longer paid per carrier.** Callbacks
725
+ land in a burst once the channel works, so a target that has not produced one
726
+ across every probe fired so far is not going to. The first carrier still gets
727
+ the full window, so a target that does call back is never cut short before its
728
+ first hit. On a clean target with the default carriers this was 30s of pure
729
+ waiting; it is now ~12s.
730
+ - **`--probe-depth` documents what it does on Windows**, which is nothing:
731
+ `cmd.exe` has no `#` comment, no `${IFS}` and no `awk`, so both depths send
732
+ the single `set /a` probe. The docs promised three extra shapes per sink
733
+ without that caveat.
734
+
735
+ ## [2.23.2] — 2026-08-04
736
+
737
+ Three findings from a review of the detection work in 2.22.0 and 2.23.0. All
738
+ three are the same shape: the run said something that was not true — about what
739
+ it had done, about what it had looked for, or about which channel was live.
740
+
741
+ ### Fixed
742
+
743
+ - **`--methods oob` ignored `--verify-active-risk`.** Detection methods build
744
+ their own probes and so bypass every corpus-level safety filter. That was
745
+ harmless while every method was inert, but this one makes the target open
746
+ outbound connections — and the same run printed *"low-impact (safe) payloads
747
+ only; pass `--verify-active-risk intrusive` to also fire … OOB"* and then fired
748
+ OOB anyway. It now needs `--verify-active-risk intrusive`, the same tier that
749
+ holds back the corpus OOB payloads, and refuses before the listener binds.
750
+ - **`--probe-depth quick` silently narrowed the timing separator screen to
751
+ `; `.** That put back the exact blind spot the screen was added to remove, so
752
+ a sink that merely filters `;` reported negative — and only for the operator
753
+ who chose `quick` to be gentle on a rate-limited target. Both depths now screen
754
+ every candidate separator; `--probe-depth` governs probe *shapes* only, and
755
+ `--separators` remains the way to narrow break-outs deliberately.
756
+ - **The DNS out-of-band probes could not call back on the default port, and
757
+ nothing said so.** A DNS callback travels the real resolver hierarchy, so it
758
+ only arrives if the listener *is* the authority for the OOB domain — port 53
759
+ plus NS delegation. On `--listen-dns-port 5335` the DNS shapes were still sent,
760
+ never fired, and the startup line reported `DNS :5335` with no caveat. Since
761
+ most of the shapes are DNS ones — a resolver is often the only egress a
762
+ hardened target has — the silence was expensive. RCEKit now says which channel
763
+ is live.
764
+ - The blind-sink advice added in 2.23.0 suggested an `oob` command without the
765
+ risk flag, which the gate above would refuse. Naming a command the tool then
766
+ declines to run is a small version of the same problem, so it now spells out
767
+ `--verify-active-risk intrusive`.
768
+
769
+ ## [2.23.1] — 2026-08-03
770
+
771
+ ### Added
772
+
773
+ - **[Verify it yourself](docs/verify-it-yourself.md)** — reproduce the README's
774
+ confirmations locally against dockerised [vulhub](https://github.com/vulhub/vulhub)
775
+ targets. Webmin CVE-2019-15107 driven from a captured request (`reflected` →
776
+ `confirmed`, then `time` → `needs-review` on the *same* sink, which is the
777
+ clearest demonstration that the tiers are not merged), and Struts2 S2-001
778
+ (`eval` confirms, `reflected` does not, on a target where both were tried).
779
+
780
+ Log4Shell is documented as an advanced case rather than a five-minute one: its
781
+ sink is a JNDI lookup inside a logging library, so `--methods oob` does not
782
+ apply — that method builds shell probes for shell-capable environments. The
783
+ `${jndi:…}` payloads come from the `oob` *category* with the listener
784
+ correlating the callback, and the token rides in a DNS label, which needs a
785
+ delegated domain. Saying so is cheaper than a reader discovering it mid-demo.
786
+
787
+ ## [2.23.0] — 2026-08-03
788
+
789
+ The three sinks v2.22.0 still could not reach. One was a real gap in the probe
790
+ set; the other two were a reporting problem, not a detection one. With both
791
+ closed, a single `--methods reflected,eval,oob` run confirms **all fifteen**
792
+ vulnerable sinks in the lab and still reports nothing on any of the five clean
793
+ ones.
794
+
795
+ ### Added
796
+
797
+ - **A space-free probe, sent at both probe depths.** Stripping spaces is a filter
798
+ of the same family as stripping `;` — it looks like it disarms command
799
+ injection and does not, because `${IFS}` is a space as far as the shell is
800
+ concerned. Every other probe carries a space, so that one filter silenced all
801
+ of them and the sink was only reachable if the operator thought to pass
802
+ `--evade low`. The separator's trailing space is trimmed with it (`;echo…`, not
803
+ `; echo…`); the newline separator is unaffected. It costs one shape, so it is
804
+ not part of the `--probe-depth` trade-off, and it is skipped under
805
+ `--evade low`, which already applies the same transform everywhere.
806
+ - **Guidance when every in-band probe comes back negative.** A results-based
807
+ method cannot confirm a sink that returns no output — there is nowhere for the
808
+ computed value to appear — so that negative is not evidence the target is
809
+ clean. A run of `reflected`/`eval` alone that confirms nothing now says exactly
810
+ that and names the methods that could still reach a blind sink, with the flags
811
+ each one needs. It is suppressed once a blind-capable method has already run,
812
+ and the `file` line is dropped once a web root is known.
813
+
814
+ ## [2.22.0] — 2026-08-03
815
+
816
+ Detection coverage. Measured against a lab of twenty sinks — fifteen genuinely
817
+ vulnerable, five deliberately clean — the results-based methods went from
818
+ confirming 8 of the 15 to confirming 12, with no new false positives on any of
819
+ the clean ones.
820
+
821
+ ### ⚠️ A blind-timing candidate could be pure latency drift
822
+
823
+ `--methods time` fired its probes in a fixed ascending delay order
824
+ (`0,0,N,N,2N,2N`), which makes the injected delay collinear with the request
825
+ index. A target that simply gets **slower during the run** — progressive load, a
826
+ rate limiter backing off, a filling log — therefore produced a textbook-perfect
827
+ linear fit while being entirely un-injectable. In the lab this reproduced on 8
828
+ of 8 runs against a sink with no command execution anywhere in it.
829
+
830
+ The probe order is now randomised, and the request index enters the regression
831
+ as a nuisance term, so drift loads onto a drift coefficient instead of
832
+ masquerading as a sleep. The same lab sink now reports negative on 9 of 9 runs,
833
+ with every genuine timing detection preserved. If you have a `needs-review`
834
+ timing candidate from an earlier version against a target that was under load,
835
+ it is worth re-running.
836
+
837
+ ### Added
838
+
839
+ - **`--methods oob`** — out-of-band detection, the first `confirmed`-tier method
840
+ for a sink that returns nothing *and* has no writable web root. Starts the
841
+ built-in HTTP+DNS listener in-process and asks the target to resolve or fetch
842
+ `<token>.<oob-host>`; a callback carrying a token the target could only have
843
+ learned by running the command is proof of execution. Each probe gets its own
844
+ token, so the finding names the break-out that actually worked. One shape puts
845
+ a computed value in the DNS label, so the callback proves the shell evaluated
846
+ arithmetic rather than merely resolving a name. Requires `--oob-host`, since
847
+ it makes the target open outbound connections.
848
+ - **`--probe-depth quick|full`** (default `full`) — trades requests for
849
+ coverage. `full` adds three probe shapes, each aimed at a filter that silenced
850
+ the canonical ones: substitution-free (`awk`, bare `expr`) for sinks that strip
851
+ `$(` and backticks; keyword-diverse (`awk`) for filters on `echo`/`expr`; and
852
+ comment-terminated (`… #`) for applications that append a redirect, extra
853
+ arguments or a pipe after the injection point. `quick` keeps the old probe set
854
+ at roughly half the requests.
855
+
856
+ ### Fixed
857
+
858
+ - **A `ping '<input>'` sink could not be detected at all.** The
859
+ `shell_single_quoted`/`shell_double_quoted` contexts exist precisely for input
860
+ interpolated inside quotes, but they are not in `default_contexts`, so no
861
+ record carried them and the detection engine never tried them — the one sink
862
+ shape they exist for was the one shape that always reported clean. They are now
863
+ probed by default, and skipped when `--contexts` names a selection explicitly.
864
+ - **`--methods time` reported a `;`-filtering sink as negative.** A regression
865
+ blends its probes into one measurement, so it could not sweep separators the
866
+ way the results-based methods do and was locked to `; ` alone — while
867
+ `| sleep 3` delayed on the same sink. It now screens every candidate separator
868
+ with one cheap probe each, then runs the regression through whichever one
869
+ actually delayed.
870
+ - **A trailing redirect or pipe in the sink hid a working probe.**
871
+ `<cmd> <input> 2>/dev/null` and `<cmd> <input> | grep …` swallow the probe's
872
+ output, so it executed and still read as negative. The comment-terminated
873
+ shapes comment that tail out.
874
+
875
+ ## [2.21.1] — 2026-08-02
876
+
877
+ First release since v2.15.2. The headline is not a new feature — it is that
878
+ detection is now correct in cases where it previously was not.
879
+
880
+ ### ⚠️ Re-check findings from v2.15.2 and earlier
881
+
882
+ **A reflection could be reported as `confirmed`.** The paired same-token control
883
+ in `run_verification` was gated on a plain `re.search`, while the verdict itself
884
+ used the encoding-aware search. A target that only echoes input but wraps its
885
+ output — base64, hex, URL- or HTML-encoded — skipped the control entirely and was
886
+ reported as proven execution: precisely the case the encoding-aware search was
887
+ added for. If you ran an earlier version against a target that encodes its
888
+ responses, a `confirmed` verdict from that run is worth re-testing.
889
+
890
+ ### Fixed — false negatives on exploitable targets
891
+
892
+ - **Separator sweep.** Shell probes always broke out with a single hardcoded
893
+ `; `, so a sink that strips `;` — the most common partial mitigation there is,
894
+ and one that stops nothing on its own — defeated every probe. Measured against
895
+ nine deliberately vulnerable local sinks, detection was correct on 5 of 9;
896
+ three of the four misses were exploitable targets reported clean. Probes now
897
+ sweep `; `, `| `, `|| `, `&& ` and a newline, narrowable with `--separators`.
898
+ - **Language runtimes.** An environment names what runs the application, not what
899
+ executes the injected command: PHP's `system()`, Python's `os.system()`,
900
+ Node's `child_process.exec()`, Ruby's `system()`, Perl's backticks and Go's
901
+ `os/exec` all hand the string to `/bin/sh`. Scoping a run to the language the
902
+ application is written in — the natural thing to do — used to send no shell
903
+ probes at all.
904
+ - **Whole-command sinks.** `--sink-raw` sends probes as bare commands for sinks
905
+ that execute the input as the entire command (`qx/$input/`, `sh -c "$input"`),
906
+ where a leading `;` is a syntax error that guaranteed a false negative.
907
+ - **Captured requests.** A trailing newline in a saved request body is no longer
908
+ sent as part of the body.
909
+
910
+ ### Fixed — a failed request is not a clean result
911
+
912
+ - A request that never reached the target is reported `error`, not `negative`.
913
+ - Runs that build no probes at all exit non-zero and say so, instead of ending
914
+ in silence and exit 0 — which read exactly like a target that came back clean.
915
+ - The OOB DNS listener no longer dies on a malformed query, and write failures
916
+ surface instead of being swallowed.
917
+
918
+ ### Fixed — safety and audit
919
+
920
+ - **Multi-step chains now carry the same safeguards as single requests.** The
921
+ chain path delivered to live targets without sink-shape filters, destructive
922
+ hold-back or a pre-flight plan, so `--verify-active-risk stateful` fired
923
+ persistence and irreversible file operations that `--verify-url` refuses to
924
+ send without `--verify-allow-destructive`. Both paths now share one hold-back
925
+ and print the same plan.
926
+ - **The audit trail redacts credentials** — it records that a credential header
927
+ was sent, never its value.
928
+ - A capture carrying `Authorization` or `Cookie` over plain `http` is flagged
929
+ before anything is sent.
930
+
931
+ ### Added
932
+
933
+ - **The payload corpus is embedded in `rcekit.py`**, so the single file runs on
934
+ its own — a jump box, an air-gapped host, a bare `curl` of the raw script.
935
+ Resolution order is `--template-file` → `templates/payloads.json` beside the
936
+ script → the built-in copy, and falling back to the built-in copy is
937
+ announced. A corpus that exists but does not parse still hard-fails: that
938
+ check exists for truncated and tampered corpora. `tools/embed_corpus.py`
939
+ regenerates the embedded copy, and the test suite fails if the two drift.
940
+ - **`--insecure`** skips TLS verification for internal targets with self-signed
941
+ or mismatched certificates — opt-in and explicit, like `curl -k`.
942
+ - **`--sink-raw`** for whole-command injection sinks, also readable from a
943
+ target profile.
944
+ - **`--separators`** to narrow the break-out sweep once the sink's shape is
945
+ known.
946
+ - **Documentation split into a task-oriented tree.** The README is half its
947
+ former length and now leads with what RCEKit is for:
948
+ [field guide](docs/guide.md) (worked examples by situation),
949
+ [payload generation & exports](docs/generation.md), and
950
+ [reference](docs/reference.md) (every flag grouped by task, plus the full
951
+ taxonomies and exit codes).
952
+ - **A "How RCEKit compares" section** covering commix, SSTImap, Nuclei and
953
+ interactsh, with every claim traceable to that project's own documentation.
954
+ - Four confirmation demos against real, publicly documented CVEs (Webmin
955
+ CVE-2019-15107, Struts2 S2-001, Log4Shell CVE-2021-44228).
956
+
957
+ ### Changed
958
+
959
+ - **Expect more requests per run.** The separator sweep and the language-runtime
960
+ fix both widen the probe set. Narrow with `--separators`, `--contexts` and
961
+ `--environments` once the sink's shape is known.
962
+ - **`--doctor` output.** Its first line now names the corpus in use
963
+ (`corpus: …`) rather than a path (`template: …`), since the corpus is no
964
+ longer necessarily a file, and `[ok] file loaded and parsed` is now
965
+ `[ok] corpus loaded and parsed`.
966
+
967
+ No breaking changes to the CLI, output formats, or the template schema.
968
+ Standard library only, Python 3.8–3.13.
969
+
970
+ ## Earlier releases
971
+
972
+ Release notes for these live on the
973
+ [Releases page](https://github.com/kabiri-labs/rcekit/releases); they predate
974
+ this file and have not been restated here.
975
+
976
+ - **[2.15.2]** — Multi-method RCE detection &amp; confirmation
977
+ - **[2.7.0]**
978
+ - **[2.1.0]**
979
+
980
+ [Unreleased]: https://github.com/kabiri-labs/rcekit/compare/v2.35.0...HEAD
981
+ [2.35.0]: https://github.com/kabiri-labs/rcekit/compare/v2.34.1...v2.35.0
982
+ [2.34.1]: https://github.com/kabiri-labs/rcekit/compare/v2.34.0...v2.34.1
983
+ [2.34.0]: https://github.com/kabiri-labs/rcekit/compare/v2.33.0...v2.34.0
984
+ [2.33.0]: https://github.com/kabiri-labs/rcekit/compare/v2.32.0...v2.33.0
985
+ [2.32.0]: https://github.com/kabiri-labs/rcekit/compare/v2.31.0...v2.32.0
986
+ [2.31.0]: https://github.com/kabiri-labs/rcekit/compare/v2.30.0...v2.31.0
987
+ [2.30.0]: https://github.com/kabiri-labs/rcekit/compare/v2.29.0...v2.30.0
988
+ [2.29.0]: https://github.com/kabiri-labs/rcekit/compare/v2.28.0...v2.29.0
989
+ [2.28.0]: https://github.com/kabiri-labs/rcekit/compare/v2.27.0...v2.28.0
990
+ [2.27.0]: https://github.com/kabiri-labs/rcekit/compare/v2.26.0...v2.27.0
991
+ [2.26.0]: https://github.com/kabiri-labs/rcekit/compare/v2.25.0...v2.26.0
992
+ [2.25.0]: https://github.com/kabiri-labs/rcekit/compare/v2.24.0...v2.25.0
993
+ [2.24.0]: https://github.com/kabiri-labs/rcekit/compare/v2.23.3...v2.24.0
994
+ [2.23.3]: https://github.com/kabiri-labs/rcekit/compare/v2.23.2...v2.23.3
995
+ [2.23.2]: https://github.com/kabiri-labs/rcekit/compare/v2.23.1...v2.23.2
996
+ [2.23.1]: https://github.com/kabiri-labs/rcekit/compare/v2.23.0...v2.23.1
997
+ [2.23.0]: https://github.com/kabiri-labs/rcekit/compare/v2.22.0...v2.23.0
998
+ [2.22.0]: https://github.com/kabiri-labs/rcekit/compare/v2.21.1...v2.22.0
999
+ [2.21.1]: https://github.com/kabiri-labs/rcekit/compare/v2.15.2...v2.21.1
1000
+ [2.15.2]: https://github.com/kabiri-labs/rcekit/releases/tag/v2.15.2
1001
+ [2.7.0]: https://github.com/kabiri-labs/rcekit/releases/tag/v2.7.0
1002
+ [2.1.0]: https://github.com/kabiri-labs/rcekit/releases/tag/v2.1.0