rcekit 2.45.3__tar.gz → 3.0.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. {rcekit-2.45.3 → rcekit-3.0.1}/CHANGELOG.md +494 -1
  2. {rcekit-2.45.3 → rcekit-3.0.1}/CONTRIBUTING.md +23 -9
  3. {rcekit-2.45.3 → rcekit-3.0.1}/PKG-INFO +264 -77
  4. {rcekit-2.45.3 → rcekit-3.0.1}/README.md +263 -76
  5. {rcekit-2.45.3 → rcekit-3.0.1}/docs/guide.md +84 -31
  6. {rcekit-2.45.3 → rcekit-3.0.1}/docs/reference.md +80 -40
  7. {rcekit-2.45.3 → rcekit-3.0.1}/docs/verify-it-yourself.md +9 -8
  8. {rcekit-2.45.3 → rcekit-3.0.1}/rcekit.egg-info/PKG-INFO +264 -77
  9. {rcekit-2.45.3 → rcekit-3.0.1}/rcekit.egg-info/SOURCES.txt +7 -0
  10. {rcekit-2.45.3 → rcekit-3.0.1}/rcekit.py +802 -220
  11. rcekit-3.0.1/tests/bench/README.md +487 -0
  12. rcekit-3.0.1/tests/bench/cases/fastjson-1.2.83.json +104 -0
  13. rcekit-3.0.1/tests/bench/cases/hugegraph-gremlin-eval.json +71 -0
  14. rcekit-3.0.1/tests/bench/cases/hugegraph-gremlin-reflected-class.json +81 -0
  15. rcekit-3.0.1/tests/bench/cases/hugegraph-gremlin-shell.json +76 -0
  16. rcekit-3.0.1/tests/bench/cases/opentsdb-cve-2023-25826.json +93 -0
  17. {rcekit-2.45.3 → rcekit-3.0.1}/tests/bench/cases/struts2-s2-001.json +1 -1
  18. rcekit-3.0.1/tests/bench/cases/tomcat-cve-2017-12615.json +105 -0
  19. rcekit-3.0.1/tests/bench/cases/webmin-cve-2019-15107-file.json +83 -0
  20. {rcekit-2.45.3 → rcekit-3.0.1}/tests/bench/cases/webmin-cve-2019-15107.json +2 -2
  21. {rcekit-2.45.3 → rcekit-3.0.1}/tests/bench/runner.py +287 -20
  22. rcekit-3.0.1/tests/test_bench_runner.py +1276 -0
  23. rcekit-3.0.1/tests/test_docs.py +1856 -0
  24. {rcekit-2.45.3 → rcekit-3.0.1}/tests/test_generator.py +1404 -191
  25. rcekit-2.45.3/tests/bench/README.md +0 -225
  26. rcekit-2.45.3/tests/test_bench_runner.py +0 -676
  27. rcekit-2.45.3/tests/test_docs.py +0 -800
  28. {rcekit-2.45.3 → rcekit-3.0.1}/LICENSE +0 -0
  29. {rcekit-2.45.3 → rcekit-3.0.1}/MANIFEST.in +0 -0
  30. {rcekit-2.45.3 → rcekit-3.0.1}/SECURITY.md +0 -0
  31. {rcekit-2.45.3 → rcekit-3.0.1}/docs/generation.md +0 -0
  32. {rcekit-2.45.3 → rcekit-3.0.1}/pyproject.toml +0 -0
  33. {rcekit-2.45.3 → rcekit-3.0.1}/rcekit.egg-info/dependency_links.txt +0 -0
  34. {rcekit-2.45.3 → rcekit-3.0.1}/rcekit.egg-info/entry_points.txt +0 -0
  35. {rcekit-2.45.3 → rcekit-3.0.1}/rcekit.egg-info/top_level.txt +0 -0
  36. {rcekit-2.45.3 → rcekit-3.0.1}/setup.cfg +0 -0
  37. {rcekit-2.45.3 → rcekit-3.0.1}/templates/payloads.json +0 -0
  38. {rcekit-2.45.3 → rcekit-3.0.1}/tests/bench/cases/log4shell-cve-2021-44228.json +0 -0
  39. {rcekit-2.45.3 → rcekit-3.0.1}/tests/bench/requests/webmin.txt +0 -0
  40. {rcekit-2.45.3 → rcekit-3.0.1}/tests/test_corpus_embedding.py +0 -0
  41. {rcekit-2.45.3 → rcekit-3.0.1}/tools/embed_corpus.py +0 -0
@@ -10,6 +10,107 @@ formats, or the template schema.
10
10
 
11
11
  ### Added
12
12
 
13
+ - **A coverage ledger, so a measured result and a reproduced one stop looking
14
+ alike.** The README's CVE table had four rows and two columns that mattered,
15
+ and no way to say what stood behind any of them. It gained seven columns and
16
+ five more rows, and the two on the right are the point: `Bench case` says
17
+ whether `tests/bench/` reproduces the row under Docker with its negative
18
+ control, and `Recording` says whether a file in `confirmation-gifs/` shows it.
19
+
20
+ Five rows are new, measured against vulhub builds during this work: OpenTSDB
21
+ 2.4.1 (CVE-2023-25826) reaching `confirmed` through `oob` and `needs-review`
22
+ through `time`; HugeGraph 1.2.0 reaching `confirmed` through both `eval` and
23
+ `reflected`; and a Spring Boot application on fastjson 1.2.83 reaching
24
+ `deserialization-sink`. All five landed saying `not yet` under `Bench case`,
25
+ which was the honest state at the time and the reason that column exists; each
26
+ has a case now. The 4 rows that predate this work had last been executed as a
27
+ set at 2.36.0, and none of them regressed. The ledger carries the figures for
28
+ the whole set, because a running total written here rots on the next change
29
+ rather than the next release.
30
+
31
+ `Advisory` is empty on three of them, deliberately. HugeGraph 1.3.0 answers
32
+ the arithmetic exactly as 1.2.0 does -- its Gremlin API evaluates Groovy
33
+ unauthenticated by design -- and fastjson resolving an `Inet4Address` is
34
+ documented autoType behaviour. Those rows prove a *method* against real
35
+ software, which is worth recording; calling them CVE reproductions would be
36
+ the overclaim the table exists to avoid. Both facts were measured rather than
37
+ assumed: the patched HugeGraph image was pulled and run.
38
+
39
+ Three tests hold the ledger to the repository: a row claiming a bench case
40
+ must have one, a row claiming a recording must have one, and the capability
41
+ rows must stay distinguishable from the CVE rows. Each was checked by making
42
+ the false claim and watching it fail. The tier check now reads *every* row
43
+ rather than only those naming an advisory, and its parser is keyed on column
44
+ headings rather than position -- the old one was hard-coded to four columns
45
+ and found nothing at all once the table grew.
46
+
47
+ - **A bench case for `write`, the last tier-1 method without one.**
48
+ `tomcat-cve-2017-12615` -- `PUT /rcekit-probe.jsp/` with the content in the
49
+ body, which is the shape the method's own docstring names. Its control runs
50
+ `reflected` against the same point and must stay `negative`: the target is
51
+ exploitable and `write` confirms on it, but a PUT answers 204 with an empty
52
+ body, so nothing in the vulnerable response is computed. Measured at 959
53
+ probes, every one answered 204.
54
+
55
+ It is the first case here that may not share a target between its halves --
56
+ both write a file into the web root -- which turned up a harness bug: under
57
+ `--keep-up` the teardown *between* the halves was skipped too, and compose
58
+ handed the control the container the vulnerable half had just written to. That
59
+ teardown is mandatory now wherever the halves would land on the same
60
+ container, which is decided from what *starts* a target -- its `vulhub_path`
61
+ and `compose` -- and not from what stops it. A control overriding only
62
+ `compose_down`, to keep its volumes say, still meets the container the
63
+ vulnerable half brought up. A control that brings up its own target has
64
+ nothing to contaminate, so `--keep-up` leaves the vulnerable one running
65
+ there, which is the environment worth looking at after a failure. The
66
+ teardown after the last half stays the flag's to skip either way.
67
+
68
+ The case is also what caught the two `write` defects fixed in 2.45.6 and
69
+ 2.45.7, neither of which a fixture could have found.
70
+
71
+ ### Fixed
72
+
73
+ - **The comparison table named seven of the nine registered methods.** `lookup`
74
+ and `boolean` had no row in "How RCEKit compares", although `lookup` is the
75
+ Log4Shell row the CVE table directly above it already carries, and `boolean`
76
+ shipped two releases ago. The table's claim is coverage -- "every class, one
77
+ run" -- so a method missing from it is that claim understating itself.
78
+
79
+ The drift was structural, not an oversight anyone could have caught by
80
+ reading. `docs/reference.md`, the `--methods` help, the CVE rows and the demo
81
+ headings are each pinned to `DETECTION_METHODS` by a test. The comparison
82
+ table was pinned by nothing, so it was the one place a new method could fail
83
+ to appear without anything going red.
84
+
85
+ Two tests now hold it, in both directions: every registered method has a row,
86
+ and the table names no method the CLI would reject -- because a one-way check
87
+ rots. A method *removed* from the registry would otherwise leave its row
88
+ behind, advertising a `--methods` value that no longer exists, while the
89
+ completeness test went on passing.
90
+
91
+ Each row now names its method in backticks, which is what makes the check
92
+ possible and also answers the question the table used to leave hanging: a
93
+ reader who wants the row they just read has the flag in front of them. Rows
94
+ that are not methods name none on purpose -- second-order execution, the
95
+ query-language bridges and the per-dialect Windows probes are things the
96
+ methods run *through*, not entries in `--methods`.
97
+
98
+ Both checks read the capability table alone, and identify a method label
99
+ structurally: the last backtick in the leading cell, after an em-dash. The
100
+ first attempt did neither, and both shortcuts cost it the guarantee it stated.
101
+ Reading every table in the section meant a deleted capability row kept passing
102
+ on a mention elsewhere -- the engagement-controls table names `file` and
103
+ `write` under "Cleanup commands" -- which was confirmed by deleting the `file`
104
+ row and watching the check stay green. And filtering labels on lowercase
105
+ letters exempted exactly the spellings a typo or a rename produces:
106
+ `boolean2`, `second-order`, `reflected_v2` all walked past the phantom check,
107
+ while `cmd.exe` is the legitimate backtick in that column that the filter had
108
+ been reaching for. Structure separates them; a character class does not.
109
+
110
+ No version bump -- this is documentation and tests only.
111
+
112
+ ### Added
113
+
13
114
  - **A tier correction has to reach the prose, not only the class.** Four tests
14
115
  now read the tier from `DETECTION_METHODS[name].tier` and hold the sentences
15
116
  an operator actually sees against it.
@@ -189,7 +290,393 @@ formats, or the template schema.
189
290
 
190
291
  No version bump -- this is documentation and tests only.
191
292
 
192
- ## [Unreleased]
293
+ ## [3.0.1] — 2026-10-01
294
+
295
+ ### Fixed
296
+
297
+ - **`reflected` named a shell it had not proved.** One of its probe shapes — a
298
+ bare `expr a + b` at `--probe-depth full`, whose expected value was the sum
299
+ alone, matched digit-fenced — reported `executed` under
300
+ `[reflected/<environment>/...]` against targets where nothing executed a
301
+ command. Groovy reads `expr 548286 + 675041` as the command expression
302
+ `expr(548286 + 675041)`: it computes the sum, fails to resolve the method, and
303
+ echoes the result in its error.
304
+
305
+ The tier was never wrong. The target did compute a value reflection cannot
306
+ forge, and `eval` reaches `executed` on such an endpoint honestly. What was
307
+ wrong is the **class**: a finding printed under `reflected/unix` asserts a
308
+ POSIX shell ran the input, and in a report a correct proof filed under the
309
+ wrong class is worse than a miss.
310
+
311
+ Measured against a plain Apache HugeGraph 1.2.0 Gremlin endpoint with every
312
+ shape mapped (`--confirm-depth every`): 42 probes, `executed=3`, all three
313
+ that one shape, and 0 of 36 tagged probes confirming. Measured against the
314
+ same build's ProcessBuilder body, where a real `sh -c` runs: 4 tagged shapes
315
+ confirm — the canonical `$(( ))` plus `$(echo TAG)` collapse, the backtick
316
+ `expr`, the `${IFS}` space-free form and the comment-terminated collapse — so
317
+ removing the shape costs nothing on a genuine shell. The remaining shapes all
318
+ require the target to place RCEKit's random tags *around* the value it
319
+ computed, which an expression evaluator handed the payload cannot do.
320
+
321
+ The shape is removed rather than repaired: `expr` has no portable string
322
+ concatenation, so there is no way to bracket its result in a tag without
323
+ reintroducing the `$(` or the `echo` keyword the shape existed to avoid. The
324
+ substitution-free coverage it shared is carried by `awk` alone, which is
325
+ unaffected. A sink that filters `$(`, backticks and `awk` at once is now out
326
+ of `reflected`'s reach and stays reachable through `time`, `oob` and `file`.
327
+
328
+ `--probe-depth quick` never sent the shape, so a `quick` run is unchanged.
329
+ Two guards hold the invariant over every environment, context and depth: no
330
+ probe's expected value may be the arithmetic alone, and none may match
331
+ digit-fenced. Both were run against the restored shape and both fail on it,
332
+ as does a new bench case whose control runs `reflected` against the same
333
+ Gremlin endpoint and requires `negative` — with the shape restored it reports
334
+ `executed=30`.
335
+
336
+ One second-order effect worth knowing: the retired shape was also the
337
+ shortest probe `reflected` built, so the shortest is now 41 characters rather
338
+ than about 20. A `--max-length` below 41 therefore leaves the method with no
339
+ probe at all. That fails loudly rather than quietly — the run prints
340
+ `NOTHING WAS TESTED — this is not a negative result`, names
341
+ `--max-length` as what removed all 36 probes, and says to relax it or treat
342
+ the sink as one the probe vocabulary cannot reach.
343
+
344
+ ## [3.0.0] — 2026-09-30
345
+
346
+ ### Changed
347
+
348
+ - **BREAKING: `confirmed` is now `executed`, and three settled measurements
349
+ stopped reporting as `needs-review`.** Every verdict is now named for what the
350
+ *target* did, never for how sure RCEKit is.
351
+
352
+ Two problems, one vocabulary. `confirmed` was the only verdict that described
353
+ RCEKit's own confidence while every sibling described the target's behaviour,
354
+ and next to it a `lookup-sink` -- proven on a callback carrying a token only
355
+ that probe held -- read as something unconfirmed. And `needs-review` had four
356
+ users, of which three had measured their answer:
357
+
358
+ | Verdict | Was | What it says |
359
+ |---|---|---|
360
+ | `executed` | `confirmed` | the target executed the input |
361
+ | `timing-sink` | `needs-review` | the target honoured a delay RCEKit injected |
362
+ | `file-write` | `needs-review` | the target stored a file at a path RCEKit chose and serves it back uninterpreted |
363
+ | `evaluation-sink` | `needs-review` | an evaluator consumed the input and partitioned on it; which evaluator is not shown |
364
+ | `needs-review` | `needs-review` | unchanged, and now with one user: `deser`'s parser fingerprint, the one signal that really is a candidate |
365
+
366
+ `write`'s middle verdict is the clearest case: its evidence line opened with
367
+ the words "ARBITRARY FILE WRITE confirmed" under a status that said
368
+ `needs-review`. The tool's own prose contradicted the tool's own status,
369
+ because the status was naming a confidence the measurement did not have a
370
+ reason to lack -- the file was on the target and RCEKit had read it back.
371
+
372
+ Nothing was promoted. `executed` still rests on exactly the three channels it
373
+ did before, and none of the five weaker tiers has a path into it. `time` in
374
+ particular stays out: a delay honoured is not necessarily a *shell* honouring
375
+ it, since a sandbox that implements `sleep` answers the regression the same
376
+ way. What changed is the direction of the reporting error -- three proven
377
+ findings were being under-reported.
378
+
379
+ - **`overall_detection_verdict` ranks `needs-review` last instead of second.**
380
+ Its position was argued from "a suspected RCE outranks a proven non-RCE in
381
+ triage", which held while `time`, `boolean` and `write`'s uninterpreted file
382
+ all reported there. The one verdict left under it is a suspected
383
+ *deserialization*, so ranking it above `deserialization-sink` put the guess
384
+ above the proof of the same thing. The full order is now `executed`,
385
+ `timing-sink`, `file-write`, `evaluation-sink`, `deserialization-sink`,
386
+ `lookup-sink`, `needs-review`.
387
+
388
+ That order is stated once, in `NON_EXECUTION_SECTIONS`, and both the
389
+ collapsed verdict and the text report read it from there. They used to state
390
+ it separately -- one as a tuple, the other as the order of print statements
391
+ in `main()` -- and they disagreed: `deserialization-sink` and `lookup-sink`
392
+ were hand-written blocks above the table holding the rest, so a run
393
+ reporting both `timing-sink` and `deserialization-sink` collapsed to the
394
+ first and led the report with the second. A statement order is not data and
395
+ cannot be checked against anything.
396
+
397
+ Folding those two in unifies their heading with the others: `NOT proof of
398
+ RCE` becomes `NOT proof of execution`, which is what the other four already
399
+ said. Their rows keep the compact form they had -- truncated payload, no
400
+ `environment` -- because they carry base64 object streams and JNDI URIs.
401
+
402
+ - The classic `--verify-url` path renamed its `confirmed` verdict too, and its
403
+ section headings with it. It has no sinks and no `needs-review`, so nothing
404
+ there was mis-tiered; a tool printing `[verify] CONFIRMED` beside `[detect]
405
+ EXECUTED` would just be two vocabularies. That path has no machine-readable
406
+ output, so this is text only.
407
+
408
+ ### Benchmark
409
+
410
+ **`runner.py --all`: 9/9, in 55m50s**, against the same vulhub builds. Every
411
+ case's expected verdict changed in this release, and a rename applied to a
412
+ stored result proves nothing about what the tool now reports, so all 9 ran:
413
+
414
+ | Case | Method | Vulnerable half | Negative control |
415
+ |---|---|---|---|
416
+ | `fastjson-1.2.83` | `deser` | `deserialization-sink` | `negative` |
417
+ | `hugegraph-gremlin-eval` | `eval` | `executed` | `negative` |
418
+ | `hugegraph-gremlin-shell` | `reflected` | `executed` | `negative` |
419
+ | `log4shell-cve-2021-44228` | `lookup` | `lookup-sink` | `negative` |
420
+ | `opentsdb-cve-2023-25826` | `oob` | `executed` | `timing-sink` |
421
+ | `struts2-s2-001` | `eval` | `executed` | `negative` |
422
+ | `tomcat-cve-2017-12615` | `write` | `executed` | `negative` |
423
+ | `webmin-cve-2019-15107-file` | `file` | `executed` | `negative` |
424
+ | `webmin-cve-2019-15107` | `reflected` | `executed` | `timing-sink` |
425
+
426
+ The two `timing-sink` controls carry the weight. On a Webmin 1.910 and an
427
+ OpenTSDB 2.4.1 that are genuinely vulnerable, `time` reaches its own proven
428
+ tier and still does not reach `executed` -- the retier holding against real
429
+ software and not against a fixture.
430
+
431
+ Two tiers this release introduced are **not** in that table: `file-write` and
432
+ `evaluation-sink`. The Tomcat `write` case reaches `executed`, and `boolean`
433
+ has no case at all. Both are covered by unit tests only.
434
+
435
+ ### Migration
436
+
437
+ `--detect-json` changes values, not keys. `verdict`, the keys of `counts`, each
438
+ probe's `verdict`, and `observe_status` take the new names per the table above.
439
+ A consumer keyed on `"confirmed"` must read `"executed"`; one keyed on
440
+ `"needs-review"` should decide which of the four it actually meant. No CLI flag,
441
+ no JSON key, no template-schema field and no exit code changed.
442
+
443
+ No backward-compatible alias is shipped. Two accepted spellings for one field is
444
+ a second source of truth, and the point of this release is that there is one.
445
+
446
+ ### Fixed
447
+
448
+ - **A method's question was derived from its tier, which does not determine
449
+ it.** `detection_question` decided whether a method asks *did the target
450
+ execute my input* by looking up its tier in a set of two names. A tier says
451
+ how strong an answer a method can reach; `time` asks the execution question
452
+ and cannot reach the strongest answer to it, while `lookup` reaches a
453
+ definitive answer to a different question.
454
+
455
+ Left alone, renaming `time`'s tier in this release would have moved it out of
456
+ the execution group and out of the dedup that skips it on a candidate
457
+ `reflected` has already proven -- so the most expensive method in the tool
458
+ would have run to put a second name on a finding. `question` is now declared
459
+ on the class, `EXECUTION_TIERS` is gone, and the partition is unchanged.
460
+
461
+ The test covering this recomputed the derivation rule it was checking, so it
462
+ agreed with the rule whatever the rule became. Three tests replace it: the
463
+ partition written out, the one case where question and tier disagree, and the
464
+ `time`-leaves-the-group regression by name.
465
+
466
+ ## [2.45.7] — 2026-09-26
467
+
468
+ ### Fixed
469
+
470
+ - **The `write` product overflowed a 32-bit int, so the method reported
471
+ `negative` on more than half of the targets it was built for.** RCEKit
472
+ computes `a*b` in Python, where integers do not overflow. The interpreter on
473
+ the other side usually is not Python: JSP and ASPX are 3 of the 5 languages
474
+ this method writes, and both evaluate the expression as a signed 32-bit int.
475
+ A product past 2**31-1 therefore came back wrapped, disagreed with the value
476
+ RCEKit had computed, and the run reported `negative` against a target that is
477
+ fully exploitable -- the one failure this tool exists not to produce.
478
+
479
+ Measured against Tomcat 8.5.19 (CVE-2017-12615), writing the probe by PUT and
480
+ fetching it back. `97233*38786` was served as `-523688158` where RCEKit
481
+ expected `3771279138`; `4721*8093` came back as `38207053`, exactly. The
482
+ operand range drew a pair whose product overflows in **56%** of runs, and the
483
+ operands are drawn once per run, so there was no second probe to recover: more
484
+ than half of all runs against a JVM or .NET target lost.
485
+
486
+ The operands are now 4 digits apiece, which caps the product at 99,980,001 --
487
+ the range `eval` has always used against these same targets. Nothing about the
488
+ oracle changes: the product is still bracketed by two random tags and still
489
+ differenced against a payload-free control, which is what makes it unforgeable
490
+ rather than its magnitude.
491
+
492
+ The same run confirms now. `RKYBRDH<%=3253*8793%>RKICSMR` came back as
493
+ `RKYBRDH28603629RKICSMR` from the fetched file, and `write` reached
494
+ `confirmed` against real software for the first time.
495
+
496
+ Two tests, both written first and watched to fail: the drawn operands multiply
497
+ to something a signed 32-bit int can hold, across 400 seeds; and a store whose
498
+ evaluator wraps to 32 bits the way a JVM does reaches `confirmed` rather than
499
+ `negative`. A third test had pinned the operand width with a hard-coded
500
+ `\d{5}`; it reads the declared range now and asserts the operands fall inside
501
+ it, which the digit count had only implied.
502
+
503
+ ## [2.45.6] — 2026-09-26
504
+
505
+ ### Fixed
506
+
507
+ - **The `write` probe carried whitespace, and some sinks tokenise before they
508
+ write.** The file content this method delivers is a language one-liner, and
509
+ its templates spelled the expression with spaces around it -- `<?= a*b ?>`,
510
+ `<%= a*b %>`. The content goes through the operator's own injection point, and
511
+ a sink that splits its input on whitespace never writes a valid file.
512
+
513
+ Measured against RRDtool 1.7.2, reached through Cacti 1.2.28's
514
+ `right_axis_label` (CVE-2025-24367), which builds the file from a
515
+ `LINE1:out:<content>` argument. `RKAAA<?=7*6?>RKBBB` wrote a 47-byte file with
516
+ the content intact. `RKAAA<?= 7*6 ?>RKBBB` wrote no file at all and answered
517
+ `ERROR: '7*6' is not a valid function name in 7*6`. The whitespace was the
518
+ whole difference.
519
+
520
+ Whitespace inside `<?= ?>` and `<%= %>` is optional, so the templates drop it:
521
+ nothing is lost and every sink of that shape is gained. Verified end to end
522
+ against the real sink rather than a fixture -- RCEKit generated
523
+ `RKYOOQS<?=69294*83370?>RKGFZQP`, RRDtool wrote it whole, and PHP answered
524
+ `RKYOOQS5777040780RKGFZQP`, which is the value RCEKit had computed for that
525
+ probe.
526
+
527
+ `jspx` keeps its spaces: it is an XML document whose root element carries
528
+ namespace attributes, and a sink that tokenises on whitespace was never going
529
+ to carry one. Two tests pin the rest -- the generated probe and the templates
530
+ themselves -- both written first and watched to fail.
531
+
532
+ ## [2.45.5] — 2026-09-25
533
+
534
+ ### Fixed
535
+
536
+ - **The `deser` shape differential answered from evidence it had not
537
+ validated.** Three ways, each found by pointing the method at a live
538
+ fastjson 1.2.83 endpoint under vulhub and asking why it said what it said.
539
+
540
+ **Drift landed on whichever form went last.** The differential is read across
541
+ requests, so anything changing with the request *index* rather than with the
542
+ payload -- a rate limiter backing off, a filling log, a warming cache -- fell
543
+ on the form sent last, and noise was sent last every time. There are two
544
+ noise probes now, first and last, bracketing the structured pair. They are
545
+ one noise body sent twice, byte for byte, because the bracket asks whether
546
+ the endpoint answered *the same probe* alike and anything the target can see
547
+ a difference in it may answer a difference to: a tail beginning `n` made
548
+ fastjson say "error parse new" where every other tail said "syntax error".
549
+ A carrier whose two ends disagree is held at `inconclusive`. The engine
550
+ de-duplicates by payload on the per-probe path but not on the aggregate path
551
+ where `deser` lives, which is what makes one body twice possible; a test pins
552
+ that, because the bracket becomes a single request if it ever changes and a
553
+ drift check comparing an answer with itself passes forever.
554
+
555
+ **An endpoint echoing its input looked like a parser.** A response carrying a
556
+ fragment of the payload that begins at the format's magic, and not carrying
557
+ it in the payload-free control, is an echo and the shape channel cannot be
558
+ read through it -- `inconclusive`, not `negative`, because the channel was
559
+ unreadable rather than empty. The sentinels are anchored on the magic's
560
+ position in the payload rather than its start, since a probe is wrapped for
561
+ its injection context first and some wrappers are longer than the magic: an
562
+ `xml_cdata` pickle probe begins `<![CDATA[gASV`. They are matched raw, case
563
+ folded (because `_signature` folds case and the two halves of one oracle must
564
+ agree), and encoded -- base64 aligned on bytes rather than characters and
565
+ read from the first three-byte boundary at or past the magic, hex from the
566
+ magic itself.
567
+
568
+ **A format declaring no magic got a verdict anyway.** Without a magic there
569
+ is no sentinel, so an echoing endpoint cannot be told from a parsing one --
570
+ and the differential answers regardless, because the noise form is random
571
+ where a structured pair is not. Every shipped format declares one; a
572
+ `--template-file` need not, and the answer to a control that cannot be
573
+ validated is to say so.
574
+
575
+ - **The page an operator plans traffic from still described three payloads.**
576
+ The bracket makes the shape oracle send four per ecosystem -- 20 requests per
577
+ carrier rather than 15, with the five the corpus ships -- and
578
+ `docs/reference.md` and the method's own docstring both still said three.
579
+ README advertises an execution plan that prints the exact probe count before
580
+ anything fires, so a page that undercounts it breaks that promise in the one
581
+ place someone reads before a monitored engagement.
582
+
583
+ Both now state four, and the arithmetic, and what the repeated probe is for.
584
+ A test holds the page to `SHAPE_FORMS` and to the ecosystem count, in both
585
+ directions: adding a form without updating the page fails it, and reverting
586
+ the page without touching the code fails it too. Nothing compared the two
587
+ before, which is how the count drifted the moment the form was added.
588
+
589
+ - **The `negative` evidence line described a comparison it had not made.** "The
590
+ endpoint answers all three forms alike" was asserted rather than observed.
591
+ It now names what actually collapsed: every form alike, a truncated stream
592
+ answered as a complete one, or the noise form answered as a structured one.
593
+
594
+ ### Withdrawn before release
595
+
596
+ - **A second route through the differential, reading noise as the control.** It
597
+ was written for the shape fastjson has -- a format resolving its type name
598
+ before the parse finishes answers a truncated stream exactly as it answers a
599
+ complete one, which collapses the original comparison on the very endpoints
600
+ it exists to find -- and it did reach `needs-review` there.
601
+
602
+ It is not in this release. "Noise differs from the structured pair" has too
603
+ many explanations that are not a parser. Drift and echo are closed above. The
604
+ third is not: a filter that refuses the random payload while letting the
605
+ structured ones through produces the same differential, no repetition exposes
606
+ it because the refusal is stable, and this verdict cannot see it -- refusal
607
+ is judged against the probe the verdict is anchored to, and that is the
608
+ well-formed one.
609
+
610
+ A format of that shape is reported `negative` by the shape differential and
611
+ reaches `deserialization-sink` through the DNS gadget instead, which is this
612
+ method's real oracle and rests on a callback rather than on the shape of a
613
+ response. Measured on the same target: 14 callbacks, each carrying a token it
614
+ could only have learned by resolving the name it was handed.
615
+
616
+ ## [2.45.4] — 2026-09-25
617
+
618
+ ### Fixed
619
+
620
+ - **A compressed response body is a body the oracles cannot read, so a target
621
+ that executed the probe was reported `negative`.** Every in-band oracle --
622
+ `reflected`, `eval`, `file`, `write`, `boolean` -- confirms by finding a
623
+ value the target computed in the response. Nothing in the delivery layer
624
+ undid `Content-Encoding`, so a gzipped body arrived as bytes that *contain*
625
+ the computed value and do not spell it, and the search for a decimal number
626
+ came back empty. The verdict was `negative`: not "could not read this", but
627
+ "the probes reached the target and it is not vulnerable".
628
+
629
+ Measured against Apache HugeGraph 1.2.0 under vulhub. The probe
630
+ `echo RKPDSWY$((258675+956113))RKIEMTB$(echo RKLUQKL)RKPDSWY`, delivered
631
+ through the Gremlin API, came back HTTP 200 carrying
632
+ `RKPDSWY1214788RKIEMTBRKLUQKLRKPDSWY` -- the sum and the substitution
633
+ collapse, both of `ReflectedMath`'s proofs. Present after gunzip, absent in
634
+ the raw bytes, and the run reported `negative` across 620 probes.
635
+
636
+ What made it survive this long is that the same target confirmed on the same
637
+ run's *other* probes: HugeGraph gzips its 200s and leaves its 4xx errors
638
+ plain, so `eval` confirmed through a 400 that carried a Groovy exception
639
+ message while the 200 that carried real command execution was invisible. A
640
+ run that finds something is a run nobody rereads.
641
+
642
+ Both HTTP entry points decode now -- the verification/detection path and the
643
+ `--verify-chain` path -- through one helper, on the success body and the
644
+ `HTTPError` body alike, since an evaluator surfacing its value in a 500 is
645
+ the case that path exists for. The header is honoured whether or not RCEKit
646
+ asked for it, because HugeGraph compresses a response to a request that sent
647
+ no `Accept-Encoding` at all.
648
+
649
+ `gzip` and `deflate` only, both from the standard library, and `deflate`
650
+ tries the zlib wrapper before the raw stream because servers disagree about
651
+ which one that name means. `br` and `zstd` need a third-party module and this
652
+ tool has none to add; a body in either is returned as its raw bytes. Nothing
653
+ here raises: a body that claims an encoding it is not in comes back
654
+ undecompressed rather than ending a run that has already spent hundreds of
655
+ probes.
656
+
657
+ This does not widen what `confirmed` means. The oracle is unchanged -- a
658
+ computed value present in the response and absent from a payload-free
659
+ control -- and nothing is promoted into it. What changes is that the response
660
+ the oracle reads is the response the target sent.
661
+
662
+ - **Decompression is bounded, because compression is where the target chooses
663
+ how much memory RCEKit allocates.** Reading a body at all is new here, and
664
+ one-shot decompression handed a blank cheque to the one party in the exchange
665
+ that is untrusted by definition -- the thing being tested. Measured: 130,478
666
+ bytes of gzip expand to 134,217,728, and 203,860 bytes hold 200 MB of zeros,
667
+ a ratio of 1:1028. This path runs once per probe.
668
+
669
+ The ceiling is 32 MiB per encoding step. Whatever was produced before it is
670
+ kept rather than discarded: the value the oracle wants is short, so searching
671
+ 32 MB is a far better answer than searching nothing, and the only risk the
672
+ ceiling carries is a false negative -- never a false `confirmed`. A truncated
673
+ read is logged rather than passed off as a clean one, since a `negative`
674
+ decided on a partial body is not the same claim as a `negative` decided on
675
+ the whole of it.
676
+
677
+ Multi-member streams are still followed. `gzip.decompress` followed them, and
678
+ a ceiling that quietly became a truncation for a server that concatenates
679
+ members would be the same false negative in a new place.
193
680
 
194
681
  ## [2.45.3] — 2026-09-24
195
682
 
@@ -2264,6 +2751,12 @@ this file and have not been restated here.
2264
2751
 
2265
2752
 
2266
2753
  [Unreleased]: https://github.com/kabiri-labs/rcekit/compare/v2.36.0...HEAD
2754
+ [3.0.1]: https://github.com/kabiri-labs/rcekit/compare/v3.0.0...v3.0.1
2755
+ [3.0.0]: https://github.com/kabiri-labs/rcekit/compare/v2.45.7...v3.0.0
2756
+ [2.45.7]: https://github.com/kabiri-labs/rcekit/compare/v2.45.6...v2.45.7
2757
+ [2.45.6]: https://github.com/kabiri-labs/rcekit/compare/v2.45.5...v2.45.6
2758
+ [2.45.5]: https://github.com/kabiri-labs/rcekit/compare/v2.45.4...v2.45.5
2759
+ [2.45.4]: https://github.com/kabiri-labs/rcekit/compare/v2.45.3...v2.45.4
2267
2760
  [2.45.3]: https://github.com/kabiri-labs/rcekit/compare/v2.45.2...v2.45.3
2268
2761
  [2.45.2]: https://github.com/kabiri-labs/rcekit/compare/v2.45.1...v2.45.2
2269
2762
  [2.45.1]: https://github.com/kabiri-labs/rcekit/compare/v2.45.0...v2.45.1
@@ -83,12 +83,16 @@ in the clearly separated **Detection methods** section of `rcekit.py`
83
83
 
84
84
  To add one:
85
85
 
86
- 1. Subclass `DetectionMethod`: set `name` and `tier` (`confirmed` for
87
- execution-proven methods, `needs-review` for candidates — the two tiers are
88
- never merged), then implement `applicable`, `build_probes`, and `confirm`.
86
+ 1. Subclass `DetectionMethod`: set `name` and `tier` (`executed` for
87
+ execution-proven methods; otherwise a tier named for what the target did,
88
+ like `timing-sink`, and `needs-review` only when the signal really is a
89
+ candidate — tiers are never merged), then implement `applicable`,
90
+ `build_probes`, and `confirm`. Also set `question`, which says whether this
91
+ asks *did the target execute my input* or something else; it does not follow
92
+ from the tier.
89
93
  2. Compute the expected value **locally** with random inputs and compare it
90
94
  against the payload-free control (reuse `self._search`, which is
91
- encoding-aware). A verdict is `confirmed` only when the computed value is
95
+ encoding-aware). A verdict is `executed` only when the computed value is
92
96
  present and absent from the control.
93
97
  3. **Declare what the method costs and what it says.** `safety` is the
94
98
  `--verify-active-risk` rung it needs: `safe` when it only makes the target
@@ -119,10 +123,20 @@ To add one:
119
123
  bench case proves it reaches that verdict against the real software. It is
120
124
  not a per-PR gate, and a method whose case cannot run yet still ships --
121
125
  with the reason written down in `tests/bench/README.md`, not left implicit.
122
- 7. Do not add a row to the README's **CVE table** without a bench run behind it.
123
- That table is the reproduced claim, and it carries the version it was last
124
- verified at. The **methods table** describes capability, and the unit suite
125
- is what backs that.
126
+ 7. A row in the README's **coverage ledger** needs a run behind it, and its
127
+ `Bench case` column has to say which kind. `yes` means `tests/bench/`
128
+ reproduces it, target and negative control; that is the strong claim and it
129
+ carries the version it was last verified at. `not yet` means it was measured
130
+ by hand against the same build and nobody can re-run it on demand — honest,
131
+ weaker, and the reason the column is there. Never fill the column in ahead of
132
+ the case: a test reads it against `tests/bench/cases/` and
133
+ `confirmation-gifs/`.
134
+
135
+ Leave the `Advisory` column empty where the verdict does not depend on the
136
+ patch. A row proving a method against real software is worth recording;
137
+ labelling it with a CVE it does not reproduce is the overclaim the ledger
138
+ exists to prevent. The **methods table** describes capability in the
139
+ abstract, and the unit suite is what backs that.
126
140
 
127
141
  ## Adding a bench case
128
142
 
@@ -152,7 +166,7 @@ New detection coverage should arrive with a case. Every case needs a **negative
152
166
  control** — the runner refuses to load one without it — because a benchmark with
153
167
  no controls rewards aggressive probing instead of measuring accuracy. The
154
168
  control may be a patched build, the same target probed for the wrong class, or a
155
- weaker method that must stay below `confirmed` on a target where it happens to
169
+ weaker method that must stay below `executed` on a target where it happens to
156
170
  be right.
157
171
 
158
172
  Run your case against the real target before submitting, and update the README