rcekit 3.0.0__tar.gz → 3.0.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (38) hide show
  1. {rcekit-3.0.0 → rcekit-3.0.1}/CHANGELOG.md +52 -0
  2. {rcekit-3.0.0 → rcekit-3.0.1}/PKG-INFO +171 -43
  3. {rcekit-3.0.0 → rcekit-3.0.1}/README.md +170 -42
  4. {rcekit-3.0.0 → rcekit-3.0.1}/docs/guide.md +70 -19
  5. {rcekit-3.0.0 → rcekit-3.0.1}/docs/reference.md +13 -6
  6. {rcekit-3.0.0 → rcekit-3.0.1}/rcekit.egg-info/PKG-INFO +171 -43
  7. {rcekit-3.0.0 → rcekit-3.0.1}/rcekit.egg-info/SOURCES.txt +1 -0
  8. {rcekit-3.0.0 → rcekit-3.0.1}/rcekit.py +25 -12
  9. {rcekit-3.0.0 → rcekit-3.0.1}/tests/bench/README.md +24 -9
  10. {rcekit-3.0.0 → rcekit-3.0.1}/tests/bench/cases/hugegraph-gremlin-eval.json +9 -7
  11. rcekit-3.0.1/tests/bench/cases/hugegraph-gremlin-reflected-class.json +81 -0
  12. {rcekit-3.0.0 → rcekit-3.0.1}/tests/test_docs.py +398 -0
  13. {rcekit-3.0.0 → rcekit-3.0.1}/tests/test_generator.py +98 -9
  14. {rcekit-3.0.0 → rcekit-3.0.1}/CONTRIBUTING.md +0 -0
  15. {rcekit-3.0.0 → rcekit-3.0.1}/LICENSE +0 -0
  16. {rcekit-3.0.0 → rcekit-3.0.1}/MANIFEST.in +0 -0
  17. {rcekit-3.0.0 → rcekit-3.0.1}/SECURITY.md +0 -0
  18. {rcekit-3.0.0 → rcekit-3.0.1}/docs/generation.md +0 -0
  19. {rcekit-3.0.0 → rcekit-3.0.1}/docs/verify-it-yourself.md +0 -0
  20. {rcekit-3.0.0 → rcekit-3.0.1}/pyproject.toml +0 -0
  21. {rcekit-3.0.0 → rcekit-3.0.1}/rcekit.egg-info/dependency_links.txt +0 -0
  22. {rcekit-3.0.0 → rcekit-3.0.1}/rcekit.egg-info/entry_points.txt +0 -0
  23. {rcekit-3.0.0 → rcekit-3.0.1}/rcekit.egg-info/top_level.txt +0 -0
  24. {rcekit-3.0.0 → rcekit-3.0.1}/setup.cfg +0 -0
  25. {rcekit-3.0.0 → rcekit-3.0.1}/templates/payloads.json +0 -0
  26. {rcekit-3.0.0 → rcekit-3.0.1}/tests/bench/cases/fastjson-1.2.83.json +0 -0
  27. {rcekit-3.0.0 → rcekit-3.0.1}/tests/bench/cases/hugegraph-gremlin-shell.json +0 -0
  28. {rcekit-3.0.0 → rcekit-3.0.1}/tests/bench/cases/log4shell-cve-2021-44228.json +0 -0
  29. {rcekit-3.0.0 → rcekit-3.0.1}/tests/bench/cases/opentsdb-cve-2023-25826.json +0 -0
  30. {rcekit-3.0.0 → rcekit-3.0.1}/tests/bench/cases/struts2-s2-001.json +0 -0
  31. {rcekit-3.0.0 → rcekit-3.0.1}/tests/bench/cases/tomcat-cve-2017-12615.json +0 -0
  32. {rcekit-3.0.0 → rcekit-3.0.1}/tests/bench/cases/webmin-cve-2019-15107-file.json +0 -0
  33. {rcekit-3.0.0 → rcekit-3.0.1}/tests/bench/cases/webmin-cve-2019-15107.json +0 -0
  34. {rcekit-3.0.0 → rcekit-3.0.1}/tests/bench/requests/webmin.txt +0 -0
  35. {rcekit-3.0.0 → rcekit-3.0.1}/tests/bench/runner.py +0 -0
  36. {rcekit-3.0.0 → rcekit-3.0.1}/tests/test_bench_runner.py +0 -0
  37. {rcekit-3.0.0 → rcekit-3.0.1}/tests/test_corpus_embedding.py +0 -0
  38. {rcekit-3.0.0 → rcekit-3.0.1}/tools/embed_corpus.py +0 -0
@@ -290,6 +290,57 @@ formats, or the template schema.
290
290
 
291
291
  No version bump -- this is documentation and tests only.
292
292
 
293
+ ## [3.0.1] — 2026-10-01
294
+
295
+ ### Fixed
296
+
297
+ - **`reflected` named a shell it had not proved.** One of its probe shapes — a
298
+ bare `expr a + b` at `--probe-depth full`, whose expected value was the sum
299
+ alone, matched digit-fenced — reported `executed` under
300
+ `[reflected/<environment>/...]` against targets where nothing executed a
301
+ command. Groovy reads `expr 548286 + 675041` as the command expression
302
+ `expr(548286 + 675041)`: it computes the sum, fails to resolve the method, and
303
+ echoes the result in its error.
304
+
305
+ The tier was never wrong. The target did compute a value reflection cannot
306
+ forge, and `eval` reaches `executed` on such an endpoint honestly. What was
307
+ wrong is the **class**: a finding printed under `reflected/unix` asserts a
308
+ POSIX shell ran the input, and in a report a correct proof filed under the
309
+ wrong class is worse than a miss.
310
+
311
+ Measured against a plain Apache HugeGraph 1.2.0 Gremlin endpoint with every
312
+ shape mapped (`--confirm-depth every`): 42 probes, `executed=3`, all three
313
+ that one shape, and 0 of 36 tagged probes confirming. Measured against the
314
+ same build's ProcessBuilder body, where a real `sh -c` runs: 4 tagged shapes
315
+ confirm — the canonical `$(( ))` plus `$(echo TAG)` collapse, the backtick
316
+ `expr`, the `${IFS}` space-free form and the comment-terminated collapse — so
317
+ removing the shape costs nothing on a genuine shell. The remaining shapes all
318
+ require the target to place RCEKit's random tags *around* the value it
319
+ computed, which an expression evaluator handed the payload cannot do.
320
+
321
+ The shape is removed rather than repaired: `expr` has no portable string
322
+ concatenation, so there is no way to bracket its result in a tag without
323
+ reintroducing the `$(` or the `echo` keyword the shape existed to avoid. The
324
+ substitution-free coverage it shared is carried by `awk` alone, which is
325
+ unaffected. A sink that filters `$(`, backticks and `awk` at once is now out
326
+ of `reflected`'s reach and stays reachable through `time`, `oob` and `file`.
327
+
328
+ `--probe-depth quick` never sent the shape, so a `quick` run is unchanged.
329
+ Two guards hold the invariant over every environment, context and depth: no
330
+ probe's expected value may be the arithmetic alone, and none may match
331
+ digit-fenced. Both were run against the restored shape and both fail on it,
332
+ as does a new bench case whose control runs `reflected` against the same
333
+ Gremlin endpoint and requires `negative` — with the shape restored it reports
334
+ `executed=30`.
335
+
336
+ One second-order effect worth knowing: the retired shape was also the
337
+ shortest probe `reflected` built, so the shortest is now 41 characters rather
338
+ than about 20. A `--max-length` below 41 therefore leaves the method with no
339
+ probe at all. That fails loudly rather than quietly — the run prints
340
+ `NOTHING WAS TESTED — this is not a negative result`, names
341
+ `--max-length` as what removed all 36 probes, and says to relax it or treat
342
+ the sink as one the probe vocabulary cannot reach.
343
+
293
344
  ## [3.0.0] — 2026-09-30
294
345
 
295
346
  ### Changed
@@ -2700,6 +2751,7 @@ this file and have not been restated here.
2700
2751
 
2701
2752
 
2702
2753
  [Unreleased]: https://github.com/kabiri-labs/rcekit/compare/v2.36.0...HEAD
2754
+ [3.0.1]: https://github.com/kabiri-labs/rcekit/compare/v3.0.0...v3.0.1
2703
2755
  [3.0.0]: https://github.com/kabiri-labs/rcekit/compare/v2.45.7...v3.0.0
2704
2756
  [2.45.7]: https://github.com/kabiri-labs/rcekit/compare/v2.45.6...v2.45.7
2705
2757
  [2.45.6]: https://github.com/kabiri-labs/rcekit/compare/v2.45.5...v2.45.6
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: rcekit
3
- Version: 3.0.0
3
+ Version: 3.0.1
4
4
  Summary: RCE detection & confirmation toolkit for authorised penetration testing
5
5
  Author: Ahmad Kabiri
6
6
  License-Expression: MIT
@@ -24,7 +24,7 @@ Dynamic: license-file
24
24
 
25
25
  **`executed` means the target executed the input. `negative` means the probes reached it.**
26
26
 
27
- **Version 3.0.0** · MIT · Python 3.8+ · zero third-party dependencies
27
+ **Version 3.0.1** · MIT · Python 3.8+ · zero third-party dependencies
28
28
 
29
29
  RCEKit is an **RCE detection &amp; confirmation toolkit** for authorised penetration
30
30
  testing, red teaming and security research. Point it at a target you are allowed
@@ -73,25 +73,15 @@ method could reach in principle.
73
73
  **Bench case** says whether [`tests/bench/`](tests/bench/) reproduces the row —
74
74
  bringing the target up under Docker and checking the verdict **and** its negative
75
75
  control. All 11 rows do; `python tests/bench/runner.py --all` was last green at
76
- **3.0.0** (2026-09-30), 9/9 cases in 55m50s. That is a point-in-time claim, not
77
- a continuous one: the benchmark runs on a cadence, not on every change. A row
78
- reading *as a control* is reproduced by the control half of another row's case,
79
- which is the half that describes it.
80
-
81
- 3.0.0 renamed the verdict every case expects, so every one of the 9 was re-run
82
- under the new names and all 9 passed — the `Verdict` column above is what that
83
- run reported, not a rename applied to an older result.
84
-
85
- What each run's **control** half reported is not in this table, and two of them
86
- are the reason the run was worth doing. The full 9 rows, controls included, are
87
- in [`tests/bench/README.md`](tests/bench/README.md#status); the two rows here
88
- reading *as a control* are the `time` halves of the Webmin and OpenTSDB cases,
89
- so their `timing-sink` is a control result shown in the `Verdict` column.
90
-
91
- Those two are the ones worth having: on a Webmin 1.910 and an OpenTSDB 2.4.1
92
- that are genuinely vulnerable, `time` reaches its own proven tier and still does
93
- not reach `executed`. That is the tier ceiling holding against real software
94
- rather than against a fixture.
76
+ **3.0.1** (2026-10-01), 10/10 cases in 55m48s. That is a point-in-time claim:
77
+ the benchmark runs on a cadence, not on every change.
78
+
79
+ The controls are not in this table and two of them are the point. On a Webmin
80
+ 1.910 and an OpenTSDB 2.4.1 that are genuinely vulnerable, `time` reaches its
81
+ own proven tier and still does not reach `executed` — a ceiling holding against
82
+ real software rather than against a fixture. Those two are the rows reading *as
83
+ a control* above. The full 10, controls included, are in
84
+ [`tests/bench/README.md`](tests/bench/README.md#status).
95
85
 
96
86
  **Advisory** is empty where the verdict does not depend on the patch. Both
97
87
  HugeGraph rows and the fastjson row are `—` deliberately: HugeGraph 1.3.0 answers
@@ -195,11 +185,55 @@ rcekit --acknowledge-consent \
195
185
 
196
186
  ```
197
187
  [detect] methods: reflected, eval
198
- [detect] sent 13 probes (13 result(s)): executed=4, negative=9
188
+ [detect] sent 1430 probes (1430 result(s)): executed=17, negative=1413
199
189
 
200
- [detect] EXECUTED (4):
201
- [reflected/unix/raw] ; echo RKYZRIP$((540141+314681))RKFWVFS$(echo RKBWOOC)RKYZRIP
202
- (target computed 'RKYZRIP854822RKFWVFSRKBWOOCRKYZRIP' — random operands, absent from control)
190
+ [detect] EXECUTED (17):
191
+ [reflected/unix/raw] | awk 'BEGIN{print "RKDXCXU" 206334+786088 "RKLMVXZ"}'
192
+ (target computed 'RKDXCXU992422RKLMVXZ' (random operands, absent from control))
193
+ ```
194
+
195
+ **That is 1430 requests against this target.** With no `--environments`, RCEKit
196
+ sweeps every sink dialect it knows — POSIX, `cmd.exe`, PowerShell, and the
197
+ language runtimes that reach a shell — because which one is behind the parameter
198
+ is the thing you do not know yet. 17 of them computed the value, which is one
199
+ sink answering in seventeen dialects, not seventeen findings.
200
+
201
+ The target moves the number: at the default `--confirm-depth first`, a carrier
202
+ stops once it confirms, so a sink that answers sends **fewer** probes than one
203
+ that never does. The same command against an endpoint that only echoes sent
204
+ 1684.
205
+
206
+ Narrow it when you already know something. Against the same target, the same
207
+ command plus `--environments unix` sends 177 requests; adding `--contexts raw`
208
+ as well sends 32.
209
+
210
+ Every probe count in this quick start was measured at **3.0.0**, and like the
211
+ benchmark's they are point-in-time rather than constants. 3.0.1 retired a probe
212
+ shape, which took a measured 44 requests per point off a comparable sink —
213
+ without costing a single confirmation, which is why the `executed` counts are
214
+ the figures least likely to have moved.
215
+
216
+ **`--max-payloads` means two different things, and both will surprise you once.**
217
+
218
+ On a single point — a `--verify-url`, or `-r` without `--auto-params` — it is
219
+ one allowance **shared by every method you selected**. `--methods
220
+ reflected,deser --max-payloads 5` sends 5 probes in total, and `reflected` can
221
+ spend all five: `deser` then never runs, and the run reports nothing about
222
+ deserialization because it never asked. That is the shape to watch, because a
223
+ budget set for politeness quietly becomes a coverage gap.
224
+
225
+ Under `--auto-params` it is the opposite — a budget *per injection point, per
226
+ question*, where execution, lookup and deserialization are three questions, and
227
+ each point's payload-free control sits outside it. Measured: `--max-payloads 5`
228
+ across 3 points and 2 questions sent 30 probes, not 5.
229
+
230
+ So read the cost line. It prints before any traffic and names the multiplier it
231
+ is applying:
232
+
233
+ ```
234
+ [detect] cost: 3 points x ~10 probes (capped by --max-payloads 5 per question,
235
+ 2 question(s) asked) = at least 36 requests (each point carries its
236
+ own payload-free control)
203
237
  ```
204
238
 
205
239
  ### From a captured request — the shape most real targets have
@@ -216,13 +250,14 @@ rcekit --acknowledge-consent \
216
250
  ```
217
251
 
218
252
  ```
219
- [detect] sent 4 probes (4 result(s)): executed=3, negative=1
220
-
221
- [detect] EXECUTED (3):
222
- [reflected/unix/raw] ; echo RKHWNHK$((114157+752773))RKXGFIH$(echo RKHSEIF)RKHWNHK
223
- (target computed 'RKHWNHK866930RKXGFIHRKHSEIFRKHWNHK' — random operands, absent from control)
253
+ [detect] EXECUTED (17):
254
+ [reflected/unix/raw] | awk 'BEGIN{print "RKXJTON" 570247+922635 "RKJCVJT"}'
255
+ (target computed 'RKXJTON1492882RKJCVJT' (random operands, absent from control))
224
256
  ```
225
257
 
258
+ Same corpus, so the same scale as above — what changes is that the session
259
+ cookie, the content type and the body go with every probe.
260
+
226
261
  The method, path, headers, body and cookies are reused as captured, and each
227
262
  value is encoded for the context it lands in — a JSON leaf, a form field and a
228
263
  cookie are not escaped the same way. Drop `-p` and mark the spot with `FUZZ` or
@@ -238,12 +273,17 @@ knowing before concluding a target is clean.
238
273
  ```bash
239
274
  rcekit --acknowledge-consent \
240
275
  -r search.req --auto-params all --point-order thorough \
241
- --methods reflected,eval,time,lookup,deser \
276
+ --methods reflected,eval,time,oob,lookup,deser,boolean \
242
277
  --oob-host oob.yourdomain.example --listen-dns-port 53 \
243
278
  --verify-active-risk stateful --probe-depth full \
244
279
  --detect-json findings.json
245
280
  ```
246
281
 
282
+ That is **7 of the 9 methods**. The two left out are the ones that need a path
283
+ only you can supply: `file` wants somewhere writable the target also serves
284
+ (`--file-write-path` + `--file-read-url`), and `write` wants the URL your own
285
+ upload lands at (`--write-url-template`). Both are below.
286
+
247
287
  ```
248
288
  [verify] loaded request from search.req: enumerating 4 injection point(s)
249
289
  [detect] enumerating 4 injection point(s) x 3 method(s)
@@ -259,7 +299,7 @@ What each flag opens up:
259
299
  | `--auto-params all` | every query value, JSON leaf, form field, multipart part, cookie and header, instead of one named field |
260
300
  | `--point-order thorough` | every non-hop-by-hop header, not just the high-yield ones |
261
301
  | `--methods ...,lookup,deser` | expression-lookup and deserialization sinks, which the shell-shaped methods cannot reach |
262
- | `--oob-host` | a callback host for the blind methods. Needs a domain delegated to you; port 53 needs root |
302
+ | `--oob-host` | a callback host for the blind methods. An IP the target can reach serves `oob`; `lookup` and `deser`'s gadget can only put a token in a DNS label, so both need a domain delegated to you, and port 53 needs root |
263
303
  | `--verify-active-risk stateful` | the top rung — adds the probe shapes that make the target fetch from an address RCEKit did not choose |
264
304
  | `--probe-depth full` | every break-out shape per sink, not the cheap ones only |
265
305
  | `--detect-json` | the same verdicts as machine-readable JSON |
@@ -271,6 +311,74 @@ target, not for production.
271
311
 
272
312
  No external infrastructure, no config file.
273
313
 
314
+ ### When nothing comes back
315
+
316
+ The most common way to call a target clean when it is not: `reflected` and
317
+ `eval` need the computed value rendered somewhere, and a sink that returns no
318
+ output has nowhere to render it. A `negative` from those two is not an answer
319
+ about the sink, only about the channel.
320
+
321
+ Two methods still reach **`executed`** there, and which one you can use depends
322
+ on what you have:
323
+
324
+ ```bash
325
+ # The target can reach something of yours. An address is enough: `oob` carries
326
+ # the token in a URL path, so the HTTP channel alone can reach `executed`.
327
+ rcekit --acknowledge-consent \
328
+ --verify-url "https://target.example/render?q=FUZZ" \
329
+ --methods oob --oob-host 203.0.113.9 \
330
+ --verify-active-risk intrusive
331
+ ```
332
+
333
+ Give it a **delegated name** and `--listen-dns-port 53` instead when you have
334
+ them, and the DNS shapes come alive too — those often cross egress filtering
335
+ that blocks outbound HTTP. They cannot work off an address: a resolver reaches
336
+ the authority for a name on port 53 only, and RCEKit says so rather than sending
337
+ probes that silently never fire. The benchmark's OpenTSDB case runs the plain
338
+ address form and reaches `executed`.
339
+
340
+ ```bash
341
+ # No egress at all — but something writable you can also read back.
342
+ # Writes a random token, fetches it, and prints a cleanup line per finding.
343
+ rcekit --acknowledge-consent \
344
+ -r search.req -p q \
345
+ --methods file \
346
+ --file-write-path /var/www/html/uploads \
347
+ --file-read-url "https://target.example/uploads/{name}" \
348
+ --verify-active-risk stateful
349
+ ```
350
+
351
+ The read-back does not have to be a web root. `--file-read-url` takes `{name}`
352
+ for a handler that wants a filename, and `{path}` or `{path_enc}` for one that
353
+ takes the whole server-side path — so a writable `/tmp` plus an LFI parameter,
354
+ a download handler or an export endpoint is enough.
355
+
356
+ With neither, `--methods time` still measures a controlled delay series — but it
357
+ reports `timing-sink`, which proves the target honoured a delay and not that a
358
+ shell ran. [Blind targets](docs/guide.md#blind-targets) and
359
+ [No-egress targets](docs/guide.md#no-egress-targets) work each of these through.
360
+
361
+ ### When your own request stores the file
362
+
363
+ An upload, or a `PUT` the server accepts. Nothing in the response is computed,
364
+ so `reflected` and `eval` report `negative` against a target that is fully
365
+ exploitable. `write` inverts the question: it stores a one-liner that *computes*
366
+ a product, fetches the file back, and reads which of three things happened —
367
+ the product (`executed`), the source verbatim (`file-write`, an arbitrary write
368
+ that is served but not interpreted), or neither.
369
+
370
+ ```bash
371
+ rcekit --acknowledge-consent \
372
+ --verify-url "https://target.example/probe.jsp/" \
373
+ --verify-method PUT --verify-data FUZZ --verify-body-location raw \
374
+ --methods write --write-url-template "https://target.example/probe.jsp" \
375
+ --verify-active-risk stateful
376
+ ```
377
+
378
+ That is the CVE-2017-12615 shape, and it is one of the benchmark's cases.
379
+ [Upload and write-primitive targets](docs/guide.md#upload-and-write-primitive-targets)
380
+ covers the others.
381
+
274
382
  **Don't take the GIFs on trust** — [reproduce them yourself](docs/verify-it-yourself.md)
275
383
  against dockerised Webmin and Struts2 targets in about five minutes.
276
384
 
@@ -442,15 +550,10 @@ python rcekit.py --acknowledge-consent -r request.txt -p host --methods reflecte
442
550
  A tool reports what it found. RCEKit also reports **what it refused to believe** —
443
551
  `inconclusive` is a verdict of its own, for anything it cannot attribute to
444
552
  execution: evidence that showed up in the payload-free control too, and equally
445
- a measurement the run never finished gathering:
446
-
447
- ```
448
- [detect] methods: reflected, eval
449
- [detect] sent 13 probes (13 result(s)): executed=0, inconclusive=2, negative=11
450
- ```
553
+ a measurement the run never finished gathering. Either would have been someone
554
+ else's finding.
451
555
 
452
- Those two would have been someone else's finding. Five mechanisms produce that
453
- verdict, and they run on every confirmation:
556
+ Five mechanisms produce that verdict, and they run on every confirmation:
454
557
 
455
558
  - **A payload-free control request.** Evidence must be present *with* the payload
456
559
  and absent *without* it. Anything in both is `inconclusive`, not a finding.
@@ -508,15 +611,40 @@ exploit the database.
508
611
 
509
612
  ## Find your situation
510
613
 
511
- Each row is a worked example in the [field guide](docs/guide.md) — the command,
512
- what it sends, and how to read what comes back.
614
+ Two different questions, and they are worth keeping apart. The first is **which
615
+ method** — what the target lets you observe decides that, and it decides how
616
+ strong an answer you can get. The second is **getting the payload to land**,
617
+ which is the same method fighting a filter, a quote or an encoding.
618
+
619
+ ### Which method, and what it needs from you
620
+
621
+ | What you can observe | Method | You must already have | Reaches |
622
+ |---|---|---|---|
623
+ | Output comes back in the response | `reflected`, `eval` | nothing | **`executed`** |
624
+ | Nothing comes back; the target has egress | `oob` | a listener it can reach — a bare IP is enough. The DNS shapes also want a delegated name and port 53 | **`executed`** |
625
+ | Nothing comes back, no egress; somewhere writable you can also read back | `file` | the write path, and any URL that returns it — a web root, an LFI parameter, a download or export handler | **`executed`** |
626
+ | Your own request stores a file | `write` | the URL the file lands at | **`executed`** |
627
+ | Nothing comes back and none of the above | `time` | nothing | `timing-sink` |
628
+ | The sink interpolates `${…}` rather than shelling out | `lookup` | a delegated domain | `lookup-sink` |
629
+ | The parameter carries a serialized object | `deser` | nothing for the shape fingerprint; the DNS gadget needs a delegated name and a listener | `needs-review`, or `deserialization-sink` with the gadget |
630
+ | The sink evaluates but renders nothing of it | `boolean` | nothing; `OR` shapes need `stateful` | `evaluation-sink` |
631
+
632
+ Only the first four reach **`executed`**. `timing-sink`, `lookup-sink`,
633
+ `evaluation-sink` and `deserialization-sink` are proven findings about something
634
+ else. `needs-review` is the exception: it is the parser fingerprint `deser`
635
+ reads when it has no gadget to fire, a real candidate and **not proof** — review
636
+ it by hand rather than reporting it. [What a verdict
637
+ means](#what-a-verdict-means) spells each one out,
638
+ [Choosing methods](docs/guide.md#choosing-methods) has the costs, and the rows
639
+ below have the worked commands.
640
+
641
+ ### Getting the payload to land
513
642
 
514
643
  | Situation | Go to |
515
644
  |---|---|
516
645
  | I have a URL and a parameter | [Point at a URL](docs/guide.md#point-at-a-url) |
517
646
  | I have a request saved from Burp | [Point at a captured request](docs/guide.md#point-at-a-captured-request) |
518
647
  | The app is JSON / the payload keeps getting mangled | [Landing the payload intact](docs/guide.md#landing-the-payload-intact) |
519
- | I don't know which class it is | [Choosing methods](docs/guide.md#choosing-methods) |
520
648
  | The sink strips `;` | [When the sink filters separators](docs/guide.md#when-the-sink-filters-separators) |
521
649
  | My input lands inside `'quotes'` | [Injecting inside quotes](docs/guide.md#injecting-inside-quotes) |
522
650
  | The sink runs my input as the whole command | [Whole-command sinks](docs/guide.md#whole-command-sinks) |
@@ -2,7 +2,7 @@
2
2
 
3
3
  **`executed` means the target executed the input. `negative` means the probes reached it.**
4
4
 
5
- **Version 3.0.0** · MIT · Python 3.8+ · zero third-party dependencies
5
+ **Version 3.0.1** · MIT · Python 3.8+ · zero third-party dependencies
6
6
 
7
7
  RCEKit is an **RCE detection &amp; confirmation toolkit** for authorised penetration
8
8
  testing, red teaming and security research. Point it at a target you are allowed
@@ -51,25 +51,15 @@ method could reach in principle.
51
51
  **Bench case** says whether [`tests/bench/`](tests/bench/) reproduces the row —
52
52
  bringing the target up under Docker and checking the verdict **and** its negative
53
53
  control. All 11 rows do; `python tests/bench/runner.py --all` was last green at
54
- **3.0.0** (2026-09-30), 9/9 cases in 55m50s. That is a point-in-time claim, not
55
- a continuous one: the benchmark runs on a cadence, not on every change. A row
56
- reading *as a control* is reproduced by the control half of another row's case,
57
- which is the half that describes it.
58
-
59
- 3.0.0 renamed the verdict every case expects, so every one of the 9 was re-run
60
- under the new names and all 9 passed — the `Verdict` column above is what that
61
- run reported, not a rename applied to an older result.
62
-
63
- What each run's **control** half reported is not in this table, and two of them
64
- are the reason the run was worth doing. The full 9 rows, controls included, are
65
- in [`tests/bench/README.md`](tests/bench/README.md#status); the two rows here
66
- reading *as a control* are the `time` halves of the Webmin and OpenTSDB cases,
67
- so their `timing-sink` is a control result shown in the `Verdict` column.
68
-
69
- Those two are the ones worth having: on a Webmin 1.910 and an OpenTSDB 2.4.1
70
- that are genuinely vulnerable, `time` reaches its own proven tier and still does
71
- not reach `executed`. That is the tier ceiling holding against real software
72
- rather than against a fixture.
54
+ **3.0.1** (2026-10-01), 10/10 cases in 55m48s. That is a point-in-time claim:
55
+ the benchmark runs on a cadence, not on every change.
56
+
57
+ The controls are not in this table and two of them are the point. On a Webmin
58
+ 1.910 and an OpenTSDB 2.4.1 that are genuinely vulnerable, `time` reaches its
59
+ own proven tier and still does not reach `executed` — a ceiling holding against
60
+ real software rather than against a fixture. Those two are the rows reading *as
61
+ a control* above. The full 10, controls included, are in
62
+ [`tests/bench/README.md`](tests/bench/README.md#status).
73
63
 
74
64
  **Advisory** is empty where the verdict does not depend on the patch. Both
75
65
  HugeGraph rows and the fastjson row are `—` deliberately: HugeGraph 1.3.0 answers
@@ -173,11 +163,55 @@ rcekit --acknowledge-consent \
173
163
 
174
164
  ```
175
165
  [detect] methods: reflected, eval
176
- [detect] sent 13 probes (13 result(s)): executed=4, negative=9
166
+ [detect] sent 1430 probes (1430 result(s)): executed=17, negative=1413
177
167
 
178
- [detect] EXECUTED (4):
179
- [reflected/unix/raw] ; echo RKYZRIP$((540141+314681))RKFWVFS$(echo RKBWOOC)RKYZRIP
180
- (target computed 'RKYZRIP854822RKFWVFSRKBWOOCRKYZRIP' — random operands, absent from control)
168
+ [detect] EXECUTED (17):
169
+ [reflected/unix/raw] | awk 'BEGIN{print "RKDXCXU" 206334+786088 "RKLMVXZ"}'
170
+ (target computed 'RKDXCXU992422RKLMVXZ' (random operands, absent from control))
171
+ ```
172
+
173
+ **That is 1430 requests against this target.** With no `--environments`, RCEKit
174
+ sweeps every sink dialect it knows — POSIX, `cmd.exe`, PowerShell, and the
175
+ language runtimes that reach a shell — because which one is behind the parameter
176
+ is the thing you do not know yet. 17 of them computed the value, which is one
177
+ sink answering in seventeen dialects, not seventeen findings.
178
+
179
+ The target moves the number: at the default `--confirm-depth first`, a carrier
180
+ stops once it confirms, so a sink that answers sends **fewer** probes than one
181
+ that never does. The same command against an endpoint that only echoes sent
182
+ 1684.
183
+
184
+ Narrow it when you already know something. Against the same target, the same
185
+ command plus `--environments unix` sends 177 requests; adding `--contexts raw`
186
+ as well sends 32.
187
+
188
+ Every probe count in this quick start was measured at **3.0.0**, and like the
189
+ benchmark's they are point-in-time rather than constants. 3.0.1 retired a probe
190
+ shape, which took a measured 44 requests per point off a comparable sink —
191
+ without costing a single confirmation, which is why the `executed` counts are
192
+ the figures least likely to have moved.
193
+
194
+ **`--max-payloads` means two different things, and both will surprise you once.**
195
+
196
+ On a single point — a `--verify-url`, or `-r` without `--auto-params` — it is
197
+ one allowance **shared by every method you selected**. `--methods
198
+ reflected,deser --max-payloads 5` sends 5 probes in total, and `reflected` can
199
+ spend all five: `deser` then never runs, and the run reports nothing about
200
+ deserialization because it never asked. That is the shape to watch, because a
201
+ budget set for politeness quietly becomes a coverage gap.
202
+
203
+ Under `--auto-params` it is the opposite — a budget *per injection point, per
204
+ question*, where execution, lookup and deserialization are three questions, and
205
+ each point's payload-free control sits outside it. Measured: `--max-payloads 5`
206
+ across 3 points and 2 questions sent 30 probes, not 5.
207
+
208
+ So read the cost line. It prints before any traffic and names the multiplier it
209
+ is applying:
210
+
211
+ ```
212
+ [detect] cost: 3 points x ~10 probes (capped by --max-payloads 5 per question,
213
+ 2 question(s) asked) = at least 36 requests (each point carries its
214
+ own payload-free control)
181
215
  ```
182
216
 
183
217
  ### From a captured request — the shape most real targets have
@@ -194,13 +228,14 @@ rcekit --acknowledge-consent \
194
228
  ```
195
229
 
196
230
  ```
197
- [detect] sent 4 probes (4 result(s)): executed=3, negative=1
198
-
199
- [detect] EXECUTED (3):
200
- [reflected/unix/raw] ; echo RKHWNHK$((114157+752773))RKXGFIH$(echo RKHSEIF)RKHWNHK
201
- (target computed 'RKHWNHK866930RKXGFIHRKHSEIFRKHWNHK' — random operands, absent from control)
231
+ [detect] EXECUTED (17):
232
+ [reflected/unix/raw] | awk 'BEGIN{print "RKXJTON" 570247+922635 "RKJCVJT"}'
233
+ (target computed 'RKXJTON1492882RKJCVJT' (random operands, absent from control))
202
234
  ```
203
235
 
236
+ Same corpus, so the same scale as above — what changes is that the session
237
+ cookie, the content type and the body go with every probe.
238
+
204
239
  The method, path, headers, body and cookies are reused as captured, and each
205
240
  value is encoded for the context it lands in — a JSON leaf, a form field and a
206
241
  cookie are not escaped the same way. Drop `-p` and mark the spot with `FUZZ` or
@@ -216,12 +251,17 @@ knowing before concluding a target is clean.
216
251
  ```bash
217
252
  rcekit --acknowledge-consent \
218
253
  -r search.req --auto-params all --point-order thorough \
219
- --methods reflected,eval,time,lookup,deser \
254
+ --methods reflected,eval,time,oob,lookup,deser,boolean \
220
255
  --oob-host oob.yourdomain.example --listen-dns-port 53 \
221
256
  --verify-active-risk stateful --probe-depth full \
222
257
  --detect-json findings.json
223
258
  ```
224
259
 
260
+ That is **7 of the 9 methods**. The two left out are the ones that need a path
261
+ only you can supply: `file` wants somewhere writable the target also serves
262
+ (`--file-write-path` + `--file-read-url`), and `write` wants the URL your own
263
+ upload lands at (`--write-url-template`). Both are below.
264
+
225
265
  ```
226
266
  [verify] loaded request from search.req: enumerating 4 injection point(s)
227
267
  [detect] enumerating 4 injection point(s) x 3 method(s)
@@ -237,7 +277,7 @@ What each flag opens up:
237
277
  | `--auto-params all` | every query value, JSON leaf, form field, multipart part, cookie and header, instead of one named field |
238
278
  | `--point-order thorough` | every non-hop-by-hop header, not just the high-yield ones |
239
279
  | `--methods ...,lookup,deser` | expression-lookup and deserialization sinks, which the shell-shaped methods cannot reach |
240
- | `--oob-host` | a callback host for the blind methods. Needs a domain delegated to you; port 53 needs root |
280
+ | `--oob-host` | a callback host for the blind methods. An IP the target can reach serves `oob`; `lookup` and `deser`'s gadget can only put a token in a DNS label, so both need a domain delegated to you, and port 53 needs root |
241
281
  | `--verify-active-risk stateful` | the top rung — adds the probe shapes that make the target fetch from an address RCEKit did not choose |
242
282
  | `--probe-depth full` | every break-out shape per sink, not the cheap ones only |
243
283
  | `--detect-json` | the same verdicts as machine-readable JSON |
@@ -249,6 +289,74 @@ target, not for production.
249
289
 
250
290
  No external infrastructure, no config file.
251
291
 
292
+ ### When nothing comes back
293
+
294
+ The most common way to call a target clean when it is not: `reflected` and
295
+ `eval` need the computed value rendered somewhere, and a sink that returns no
296
+ output has nowhere to render it. A `negative` from those two is not an answer
297
+ about the sink, only about the channel.
298
+
299
+ Two methods still reach **`executed`** there, and which one you can use depends
300
+ on what you have:
301
+
302
+ ```bash
303
+ # The target can reach something of yours. An address is enough: `oob` carries
304
+ # the token in a URL path, so the HTTP channel alone can reach `executed`.
305
+ rcekit --acknowledge-consent \
306
+ --verify-url "https://target.example/render?q=FUZZ" \
307
+ --methods oob --oob-host 203.0.113.9 \
308
+ --verify-active-risk intrusive
309
+ ```
310
+
311
+ Give it a **delegated name** and `--listen-dns-port 53` instead when you have
312
+ them, and the DNS shapes come alive too — those often cross egress filtering
313
+ that blocks outbound HTTP. They cannot work off an address: a resolver reaches
314
+ the authority for a name on port 53 only, and RCEKit says so rather than sending
315
+ probes that silently never fire. The benchmark's OpenTSDB case runs the plain
316
+ address form and reaches `executed`.
317
+
318
+ ```bash
319
+ # No egress at all — but something writable you can also read back.
320
+ # Writes a random token, fetches it, and prints a cleanup line per finding.
321
+ rcekit --acknowledge-consent \
322
+ -r search.req -p q \
323
+ --methods file \
324
+ --file-write-path /var/www/html/uploads \
325
+ --file-read-url "https://target.example/uploads/{name}" \
326
+ --verify-active-risk stateful
327
+ ```
328
+
329
+ The read-back does not have to be a web root. `--file-read-url` takes `{name}`
330
+ for a handler that wants a filename, and `{path}` or `{path_enc}` for one that
331
+ takes the whole server-side path — so a writable `/tmp` plus an LFI parameter,
332
+ a download handler or an export endpoint is enough.
333
+
334
+ With neither, `--methods time` still measures a controlled delay series — but it
335
+ reports `timing-sink`, which proves the target honoured a delay and not that a
336
+ shell ran. [Blind targets](docs/guide.md#blind-targets) and
337
+ [No-egress targets](docs/guide.md#no-egress-targets) work each of these through.
338
+
339
+ ### When your own request stores the file
340
+
341
+ An upload, or a `PUT` the server accepts. Nothing in the response is computed,
342
+ so `reflected` and `eval` report `negative` against a target that is fully
343
+ exploitable. `write` inverts the question: it stores a one-liner that *computes*
344
+ a product, fetches the file back, and reads which of three things happened —
345
+ the product (`executed`), the source verbatim (`file-write`, an arbitrary write
346
+ that is served but not interpreted), or neither.
347
+
348
+ ```bash
349
+ rcekit --acknowledge-consent \
350
+ --verify-url "https://target.example/probe.jsp/" \
351
+ --verify-method PUT --verify-data FUZZ --verify-body-location raw \
352
+ --methods write --write-url-template "https://target.example/probe.jsp" \
353
+ --verify-active-risk stateful
354
+ ```
355
+
356
+ That is the CVE-2017-12615 shape, and it is one of the benchmark's cases.
357
+ [Upload and write-primitive targets](docs/guide.md#upload-and-write-primitive-targets)
358
+ covers the others.
359
+
252
360
  **Don't take the GIFs on trust** — [reproduce them yourself](docs/verify-it-yourself.md)
253
361
  against dockerised Webmin and Struts2 targets in about five minutes.
254
362
 
@@ -420,15 +528,10 @@ python rcekit.py --acknowledge-consent -r request.txt -p host --methods reflecte
420
528
  A tool reports what it found. RCEKit also reports **what it refused to believe** —
421
529
  `inconclusive` is a verdict of its own, for anything it cannot attribute to
422
530
  execution: evidence that showed up in the payload-free control too, and equally
423
- a measurement the run never finished gathering:
424
-
425
- ```
426
- [detect] methods: reflected, eval
427
- [detect] sent 13 probes (13 result(s)): executed=0, inconclusive=2, negative=11
428
- ```
531
+ a measurement the run never finished gathering. Either would have been someone
532
+ else's finding.
429
533
 
430
- Those two would have been someone else's finding. Five mechanisms produce that
431
- verdict, and they run on every confirmation:
534
+ Five mechanisms produce that verdict, and they run on every confirmation:
432
535
 
433
536
  - **A payload-free control request.** Evidence must be present *with* the payload
434
537
  and absent *without* it. Anything in both is `inconclusive`, not a finding.
@@ -486,15 +589,40 @@ exploit the database.
486
589
 
487
590
  ## Find your situation
488
591
 
489
- Each row is a worked example in the [field guide](docs/guide.md) — the command,
490
- what it sends, and how to read what comes back.
592
+ Two different questions, and they are worth keeping apart. The first is **which
593
+ method** — what the target lets you observe decides that, and it decides how
594
+ strong an answer you can get. The second is **getting the payload to land**,
595
+ which is the same method fighting a filter, a quote or an encoding.
596
+
597
+ ### Which method, and what it needs from you
598
+
599
+ | What you can observe | Method | You must already have | Reaches |
600
+ |---|---|---|---|
601
+ | Output comes back in the response | `reflected`, `eval` | nothing | **`executed`** |
602
+ | Nothing comes back; the target has egress | `oob` | a listener it can reach — a bare IP is enough. The DNS shapes also want a delegated name and port 53 | **`executed`** |
603
+ | Nothing comes back, no egress; somewhere writable you can also read back | `file` | the write path, and any URL that returns it — a web root, an LFI parameter, a download or export handler | **`executed`** |
604
+ | Your own request stores a file | `write` | the URL the file lands at | **`executed`** |
605
+ | Nothing comes back and none of the above | `time` | nothing | `timing-sink` |
606
+ | The sink interpolates `${…}` rather than shelling out | `lookup` | a delegated domain | `lookup-sink` |
607
+ | The parameter carries a serialized object | `deser` | nothing for the shape fingerprint; the DNS gadget needs a delegated name and a listener | `needs-review`, or `deserialization-sink` with the gadget |
608
+ | The sink evaluates but renders nothing of it | `boolean` | nothing; `OR` shapes need `stateful` | `evaluation-sink` |
609
+
610
+ Only the first four reach **`executed`**. `timing-sink`, `lookup-sink`,
611
+ `evaluation-sink` and `deserialization-sink` are proven findings about something
612
+ else. `needs-review` is the exception: it is the parser fingerprint `deser`
613
+ reads when it has no gadget to fire, a real candidate and **not proof** — review
614
+ it by hand rather than reporting it. [What a verdict
615
+ means](#what-a-verdict-means) spells each one out,
616
+ [Choosing methods](docs/guide.md#choosing-methods) has the costs, and the rows
617
+ below have the worked commands.
618
+
619
+ ### Getting the payload to land
491
620
 
492
621
  | Situation | Go to |
493
622
  |---|---|
494
623
  | I have a URL and a parameter | [Point at a URL](docs/guide.md#point-at-a-url) |
495
624
  | I have a request saved from Burp | [Point at a captured request](docs/guide.md#point-at-a-captured-request) |
496
625
  | The app is JSON / the payload keeps getting mangled | [Landing the payload intact](docs/guide.md#landing-the-payload-intact) |
497
- | I don't know which class it is | [Choosing methods](docs/guide.md#choosing-methods) |
498
626
  | The sink strips `;` | [When the sink filters separators](docs/guide.md#when-the-sink-filters-separators) |
499
627
  | My input lands inside `'quotes'` | [Injecting inside quotes](docs/guide.md#injecting-inside-quotes) |
500
628
  | The sink runs my input as the whole command | [Whole-command sinks](docs/guide.md#whole-command-sinks) |