rcekit 2.45.3__tar.gz → 3.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (40) hide show
  1. {rcekit-2.45.3 → rcekit-3.0.0}/CHANGELOG.md +442 -1
  2. {rcekit-2.45.3 → rcekit-3.0.0}/CONTRIBUTING.md +23 -9
  3. {rcekit-2.45.3 → rcekit-3.0.0}/PKG-INFO +119 -60
  4. {rcekit-2.45.3 → rcekit-3.0.0}/README.md +118 -59
  5. {rcekit-2.45.3 → rcekit-3.0.0}/docs/guide.md +17 -15
  6. {rcekit-2.45.3 → rcekit-3.0.0}/docs/reference.md +69 -36
  7. {rcekit-2.45.3 → rcekit-3.0.0}/docs/verify-it-yourself.md +9 -8
  8. {rcekit-2.45.3 → rcekit-3.0.0}/rcekit.egg-info/PKG-INFO +119 -60
  9. {rcekit-2.45.3 → rcekit-3.0.0}/rcekit.egg-info/SOURCES.txt +6 -0
  10. {rcekit-2.45.3 → rcekit-3.0.0}/rcekit.py +778 -209
  11. rcekit-3.0.0/tests/bench/README.md +472 -0
  12. rcekit-3.0.0/tests/bench/cases/fastjson-1.2.83.json +104 -0
  13. rcekit-3.0.0/tests/bench/cases/hugegraph-gremlin-eval.json +69 -0
  14. rcekit-3.0.0/tests/bench/cases/hugegraph-gremlin-shell.json +76 -0
  15. rcekit-3.0.0/tests/bench/cases/opentsdb-cve-2023-25826.json +93 -0
  16. {rcekit-2.45.3 → rcekit-3.0.0}/tests/bench/cases/struts2-s2-001.json +1 -1
  17. rcekit-3.0.0/tests/bench/cases/tomcat-cve-2017-12615.json +105 -0
  18. rcekit-3.0.0/tests/bench/cases/webmin-cve-2019-15107-file.json +83 -0
  19. {rcekit-2.45.3 → rcekit-3.0.0}/tests/bench/cases/webmin-cve-2019-15107.json +2 -2
  20. {rcekit-2.45.3 → rcekit-3.0.0}/tests/bench/runner.py +287 -20
  21. rcekit-3.0.0/tests/test_bench_runner.py +1276 -0
  22. rcekit-3.0.0/tests/test_docs.py +1458 -0
  23. {rcekit-2.45.3 → rcekit-3.0.0}/tests/test_generator.py +1306 -182
  24. rcekit-2.45.3/tests/bench/README.md +0 -225
  25. rcekit-2.45.3/tests/test_bench_runner.py +0 -676
  26. rcekit-2.45.3/tests/test_docs.py +0 -800
  27. {rcekit-2.45.3 → rcekit-3.0.0}/LICENSE +0 -0
  28. {rcekit-2.45.3 → rcekit-3.0.0}/MANIFEST.in +0 -0
  29. {rcekit-2.45.3 → rcekit-3.0.0}/SECURITY.md +0 -0
  30. {rcekit-2.45.3 → rcekit-3.0.0}/docs/generation.md +0 -0
  31. {rcekit-2.45.3 → rcekit-3.0.0}/pyproject.toml +0 -0
  32. {rcekit-2.45.3 → rcekit-3.0.0}/rcekit.egg-info/dependency_links.txt +0 -0
  33. {rcekit-2.45.3 → rcekit-3.0.0}/rcekit.egg-info/entry_points.txt +0 -0
  34. {rcekit-2.45.3 → rcekit-3.0.0}/rcekit.egg-info/top_level.txt +0 -0
  35. {rcekit-2.45.3 → rcekit-3.0.0}/setup.cfg +0 -0
  36. {rcekit-2.45.3 → rcekit-3.0.0}/templates/payloads.json +0 -0
  37. {rcekit-2.45.3 → rcekit-3.0.0}/tests/bench/cases/log4shell-cve-2021-44228.json +0 -0
  38. {rcekit-2.45.3 → rcekit-3.0.0}/tests/bench/requests/webmin.txt +0 -0
  39. {rcekit-2.45.3 → rcekit-3.0.0}/tests/test_corpus_embedding.py +0 -0
  40. {rcekit-2.45.3 → rcekit-3.0.0}/tools/embed_corpus.py +0 -0
@@ -10,6 +10,107 @@ formats, or the template schema.
10
10
 
11
11
  ### Added
12
12
 
13
+ - **A coverage ledger, so a measured result and a reproduced one stop looking
14
+ alike.** The README's CVE table had four rows and two columns that mattered,
15
+ and no way to say what stood behind any of them. It gained seven columns and
16
+ five more rows, and the two on the right are the point: `Bench case` says
17
+ whether `tests/bench/` reproduces the row under Docker with its negative
18
+ control, and `Recording` says whether a file in `confirmation-gifs/` shows it.
19
+
20
+ Five rows are new, measured against vulhub builds during this work: OpenTSDB
21
+ 2.4.1 (CVE-2023-25826) reaching `confirmed` through `oob` and `needs-review`
22
+ through `time`; HugeGraph 1.2.0 reaching `confirmed` through both `eval` and
23
+ `reflected`; and a Spring Boot application on fastjson 1.2.83 reaching
24
+ `deserialization-sink`. All five landed saying `not yet` under `Bench case`,
25
+ which was the honest state at the time and the reason that column exists; each
26
+ has a case now. The 4 rows that predate this work had last been executed as a
27
+ set at 2.36.0, and none of them regressed. The ledger carries the figures for
28
+ the whole set, because a running total written here rots on the next change
29
+ rather than the next release.
30
+
31
+ `Advisory` is empty on three of them, deliberately. HugeGraph 1.3.0 answers
32
+ the arithmetic exactly as 1.2.0 does -- its Gremlin API evaluates Groovy
33
+ unauthenticated by design -- and fastjson resolving an `Inet4Address` is
34
+ documented autoType behaviour. Those rows prove a *method* against real
35
+ software, which is worth recording; calling them CVE reproductions would be
36
+ the overclaim the table exists to avoid. Both facts were measured rather than
37
+ assumed: the patched HugeGraph image was pulled and run.
38
+
39
+ Three tests hold the ledger to the repository: a row claiming a bench case
40
+ must have one, a row claiming a recording must have one, and the capability
41
+ rows must stay distinguishable from the CVE rows. Each was checked by making
42
+ the false claim and watching it fail. The tier check now reads *every* row
43
+ rather than only those naming an advisory, and its parser is keyed on column
44
+ headings rather than position -- the old one was hard-coded to four columns
45
+ and found nothing at all once the table grew.
46
+
47
+ - **A bench case for `write`, the last tier-1 method without one.**
48
+ `tomcat-cve-2017-12615` -- `PUT /rcekit-probe.jsp/` with the content in the
49
+ body, which is the shape the method's own docstring names. Its control runs
50
+ `reflected` against the same point and must stay `negative`: the target is
51
+ exploitable and `write` confirms on it, but a PUT answers 204 with an empty
52
+ body, so nothing in the vulnerable response is computed. Measured at 959
53
+ probes, every one answered 204.
54
+
55
+ It is the first case here that may not share a target between its halves --
56
+ both write a file into the web root -- which turned up a harness bug: under
57
+ `--keep-up` the teardown *between* the halves was skipped too, and compose
58
+ handed the control the container the vulnerable half had just written to. That
59
+ teardown is mandatory now wherever the halves would land on the same
60
+ container, which is decided from what *starts* a target -- its `vulhub_path`
61
+ and `compose` -- and not from what stops it. A control overriding only
62
+ `compose_down`, to keep its volumes say, still meets the container the
63
+ vulnerable half brought up. A control that brings up its own target has
64
+ nothing to contaminate, so `--keep-up` leaves the vulnerable one running
65
+ there, which is the environment worth looking at after a failure. The
66
+ teardown after the last half stays the flag's to skip either way.
67
+
68
+ The case is also what caught the two `write` defects fixed in 2.45.6 and
69
+ 2.45.7, neither of which a fixture could have found.
70
+
71
+ ### Fixed
72
+
73
+ - **The comparison table named seven of the nine registered methods.** `lookup`
74
+ and `boolean` had no row in "How RCEKit compares", although `lookup` is the
75
+ Log4Shell row the CVE table directly above it already carries, and `boolean`
76
+ shipped two releases ago. The table's claim is coverage -- "every class, one
77
+ run" -- so a method missing from it is that claim understating itself.
78
+
79
+ The drift was structural, not an oversight anyone could have caught by
80
+ reading. `docs/reference.md`, the `--methods` help, the CVE rows and the demo
81
+ headings are each pinned to `DETECTION_METHODS` by a test. The comparison
82
+ table was pinned by nothing, so it was the one place a new method could fail
83
+ to appear without anything going red.
84
+
85
+ Two tests now hold it, in both directions: every registered method has a row,
86
+ and the table names no method the CLI would reject -- because a one-way check
87
+ rots. A method *removed* from the registry would otherwise leave its row
88
+ behind, advertising a `--methods` value that no longer exists, while the
89
+ completeness test went on passing.
90
+
91
+ Each row now names its method in backticks, which is what makes the check
92
+ possible and also answers the question the table used to leave hanging: a
93
+ reader who wants the row they just read has the flag in front of them. Rows
94
+ that are not methods name none on purpose -- second-order execution, the
95
+ query-language bridges and the per-dialect Windows probes are things the
96
+ methods run *through*, not entries in `--methods`.
97
+
98
+ Both checks read the capability table alone, and identify a method label
99
+ structurally: the last backtick in the leading cell, after an em-dash. The
100
+ first attempt did neither, and both shortcuts cost it the guarantee it stated.
101
+ Reading every table in the section meant a deleted capability row kept passing
102
+ on a mention elsewhere -- the engagement-controls table names `file` and
103
+ `write` under "Cleanup commands" -- which was confirmed by deleting the `file`
104
+ row and watching the check stay green. And filtering labels on lowercase
105
+ letters exempted exactly the spellings a typo or a rename produces:
106
+ `boolean2`, `second-order`, `reflected_v2` all walked past the phantom check,
107
+ while `cmd.exe` is the legitimate backtick in that column that the filter had
108
+ been reaching for. Structure separates them; a character class does not.
109
+
110
+ No version bump -- this is documentation and tests only.
111
+
112
+ ### Added
113
+
13
114
  - **A tier correction has to reach the prose, not only the class.** Four tests
14
115
  now read the tier from `DETECTION_METHODS[name].tier` and hold the sentences
15
116
  an operator actually sees against it.
@@ -189,7 +290,342 @@ formats, or the template schema.
189
290
 
190
291
  No version bump -- this is documentation and tests only.
191
292
 
192
- ## [Unreleased]
293
+ ## [3.0.0] — 2026-09-30
294
+
295
+ ### Changed
296
+
297
+ - **BREAKING: `confirmed` is now `executed`, and three settled measurements
298
+ stopped reporting as `needs-review`.** Every verdict is now named for what the
299
+ *target* did, never for how sure RCEKit is.
300
+
301
+ Two problems, one vocabulary. `confirmed` was the only verdict that described
302
+ RCEKit's own confidence while every sibling described the target's behaviour,
303
+ and next to it a `lookup-sink` -- proven on a callback carrying a token only
304
+ that probe held -- read as something unconfirmed. And `needs-review` had four
305
+ users, of which three had measured their answer:
306
+
307
+ | Verdict | Was | What it says |
308
+ |---|---|---|
309
+ | `executed` | `confirmed` | the target executed the input |
310
+ | `timing-sink` | `needs-review` | the target honoured a delay RCEKit injected |
311
+ | `file-write` | `needs-review` | the target stored a file at a path RCEKit chose and serves it back uninterpreted |
312
+ | `evaluation-sink` | `needs-review` | an evaluator consumed the input and partitioned on it; which evaluator is not shown |
313
+ | `needs-review` | `needs-review` | unchanged, and now with one user: `deser`'s parser fingerprint, the one signal that really is a candidate |
314
+
315
+ `write`'s middle verdict is the clearest case: its evidence line opened with
316
+ the words "ARBITRARY FILE WRITE confirmed" under a status that said
317
+ `needs-review`. The tool's own prose contradicted the tool's own status,
318
+ because the status was naming a confidence the measurement did not have a
319
+ reason to lack -- the file was on the target and RCEKit had read it back.
320
+
321
+ Nothing was promoted. `executed` still rests on exactly the three channels it
322
+ did before, and none of the five weaker tiers has a path into it. `time` in
323
+ particular stays out: a delay honoured is not necessarily a *shell* honouring
324
+ it, since a sandbox that implements `sleep` answers the regression the same
325
+ way. What changed is the direction of the reporting error -- three proven
326
+ findings were being under-reported.
327
+
328
+ - **`overall_detection_verdict` ranks `needs-review` last instead of second.**
329
+ Its position was argued from "a suspected RCE outranks a proven non-RCE in
330
+ triage", which held while `time`, `boolean` and `write`'s uninterpreted file
331
+ all reported there. The one verdict left under it is a suspected
332
+ *deserialization*, so ranking it above `deserialization-sink` put the guess
333
+ above the proof of the same thing. The full order is now `executed`,
334
+ `timing-sink`, `file-write`, `evaluation-sink`, `deserialization-sink`,
335
+ `lookup-sink`, `needs-review`.
336
+
337
+ That order is stated once, in `NON_EXECUTION_SECTIONS`, and both the
338
+ collapsed verdict and the text report read it from there. They used to state
339
+ it separately -- one as a tuple, the other as the order of print statements
340
+ in `main()` -- and they disagreed: `deserialization-sink` and `lookup-sink`
341
+ were hand-written blocks above the table holding the rest, so a run
342
+ reporting both `timing-sink` and `deserialization-sink` collapsed to the
343
+ first and led the report with the second. A statement order is not data and
344
+ cannot be checked against anything.
345
+
346
+ Folding those two in unifies their heading with the others: `NOT proof of
347
+ RCE` becomes `NOT proof of execution`, which is what the other four already
348
+ said. Their rows keep the compact form they had -- truncated payload, no
349
+ `environment` -- because they carry base64 object streams and JNDI URIs.
350
+
351
+ - The classic `--verify-url` path renamed its `confirmed` verdict too, and its
352
+ section headings with it. It has no sinks and no `needs-review`, so nothing
353
+ there was mis-tiered; a tool printing `[verify] CONFIRMED` beside `[detect]
354
+ EXECUTED` would just be two vocabularies. That path has no machine-readable
355
+ output, so this is text only.
356
+
357
+ ### Benchmark
358
+
359
+ **`runner.py --all`: 9/9, in 55m50s**, against the same vulhub builds. Every
360
+ case's expected verdict changed in this release, and a rename applied to a
361
+ stored result proves nothing about what the tool now reports, so all 9 ran:
362
+
363
+ | Case | Method | Vulnerable half | Negative control |
364
+ |---|---|---|---|
365
+ | `fastjson-1.2.83` | `deser` | `deserialization-sink` | `negative` |
366
+ | `hugegraph-gremlin-eval` | `eval` | `executed` | `negative` |
367
+ | `hugegraph-gremlin-shell` | `reflected` | `executed` | `negative` |
368
+ | `log4shell-cve-2021-44228` | `lookup` | `lookup-sink` | `negative` |
369
+ | `opentsdb-cve-2023-25826` | `oob` | `executed` | `timing-sink` |
370
+ | `struts2-s2-001` | `eval` | `executed` | `negative` |
371
+ | `tomcat-cve-2017-12615` | `write` | `executed` | `negative` |
372
+ | `webmin-cve-2019-15107-file` | `file` | `executed` | `negative` |
373
+ | `webmin-cve-2019-15107` | `reflected` | `executed` | `timing-sink` |
374
+
375
+ The two `timing-sink` controls carry the weight. On a Webmin 1.910 and an
376
+ OpenTSDB 2.4.1 that are genuinely vulnerable, `time` reaches its own proven
377
+ tier and still does not reach `executed` -- the retier holding against real
378
+ software and not against a fixture.
379
+
380
+ Two tiers this release introduced are **not** in that table: `file-write` and
381
+ `evaluation-sink`. The Tomcat `write` case reaches `executed`, and `boolean`
382
+ has no case at all. Both are covered by unit tests only.
383
+
384
+ ### Migration
385
+
386
+ `--detect-json` changes values, not keys. `verdict`, the keys of `counts`, each
387
+ probe's `verdict`, and `observe_status` take the new names per the table above.
388
+ A consumer keyed on `"confirmed"` must read `"executed"`; one keyed on
389
+ `"needs-review"` should decide which of the four it actually meant. No CLI flag,
390
+ no JSON key, no template-schema field and no exit code changed.
391
+
392
+ No backward-compatible alias is shipped. Two accepted spellings for one field is
393
+ a second source of truth, and the point of this release is that there is one.
394
+
395
+ ### Fixed
396
+
397
+ - **A method's question was derived from its tier, which does not determine
398
+ it.** `detection_question` decided whether a method asks *did the target
399
+ execute my input* by looking up its tier in a set of two names. A tier says
400
+ how strong an answer a method can reach; `time` asks the execution question
401
+ and cannot reach the strongest answer to it, while `lookup` reaches a
402
+ definitive answer to a different question.
403
+
404
+ Left alone, renaming `time`'s tier in this release would have moved it out of
405
+ the execution group and out of the dedup that skips it on a candidate
406
+ `reflected` has already proven -- so the most expensive method in the tool
407
+ would have run to put a second name on a finding. `question` is now declared
408
+ on the class, `EXECUTION_TIERS` is gone, and the partition is unchanged.
409
+
410
+ The test covering this recomputed the derivation rule it was checking, so it
411
+ agreed with the rule whatever the rule became. Three tests replace it: the
412
+ partition written out, the one case where question and tier disagree, and the
413
+ `time`-leaves-the-group regression by name.
414
+
415
+ ## [2.45.7] — 2026-09-26
416
+
417
+ ### Fixed
418
+
419
+ - **The `write` product overflowed a 32-bit int, so the method reported
420
+ `negative` on more than half of the targets it was built for.** RCEKit
421
+ computes `a*b` in Python, where integers do not overflow. The interpreter on
422
+ the other side usually is not Python: JSP and ASPX are 3 of the 5 languages
423
+ this method writes, and both evaluate the expression as a signed 32-bit int.
424
+ A product past 2**31-1 therefore came back wrapped, disagreed with the value
425
+ RCEKit had computed, and the run reported `negative` against a target that is
426
+ fully exploitable -- the one failure this tool exists not to produce.
427
+
428
+ Measured against Tomcat 8.5.19 (CVE-2017-12615), writing the probe by PUT and
429
+ fetching it back. `97233*38786` was served as `-523688158` where RCEKit
430
+ expected `3771279138`; `4721*8093` came back as `38207053`, exactly. The
431
+ operand range drew a pair whose product overflows in **56%** of runs, and the
432
+ operands are drawn once per run, so there was no second probe to recover: more
433
+ than half of all runs against a JVM or .NET target lost.
434
+
435
+ The operands are now 4 digits apiece, which caps the product at 99,980,001 --
436
+ the range `eval` has always used against these same targets. Nothing about the
437
+ oracle changes: the product is still bracketed by two random tags and still
438
+ differenced against a payload-free control, which is what makes it unforgeable
439
+ rather than its magnitude.
440
+
441
+ The same run confirms now. `RKYBRDH<%=3253*8793%>RKICSMR` came back as
442
+ `RKYBRDH28603629RKICSMR` from the fetched file, and `write` reached
443
+ `confirmed` against real software for the first time.
444
+
445
+ Two tests, both written first and watched to fail: the drawn operands multiply
446
+ to something a signed 32-bit int can hold, across 400 seeds; and a store whose
447
+ evaluator wraps to 32 bits the way a JVM does reaches `confirmed` rather than
448
+ `negative`. A third test had pinned the operand width with a hard-coded
449
+ `\d{5}`; it reads the declared range now and asserts the operands fall inside
450
+ it, which the digit count had only implied.
451
+
452
+ ## [2.45.6] — 2026-09-26
453
+
454
+ ### Fixed
455
+
456
+ - **The `write` probe carried whitespace, and some sinks tokenise before they
457
+ write.** The file content this method delivers is a language one-liner, and
458
+ its templates spelled the expression with spaces around it -- `<?= a*b ?>`,
459
+ `<%= a*b %>`. The content goes through the operator's own injection point, and
460
+ a sink that splits its input on whitespace never writes a valid file.
461
+
462
+ Measured against RRDtool 1.7.2, reached through Cacti 1.2.28's
463
+ `right_axis_label` (CVE-2025-24367), which builds the file from a
464
+ `LINE1:out:<content>` argument. `RKAAA<?=7*6?>RKBBB` wrote a 47-byte file with
465
+ the content intact. `RKAAA<?= 7*6 ?>RKBBB` wrote no file at all and answered
466
+ `ERROR: '7*6' is not a valid function name in 7*6`. The whitespace was the
467
+ whole difference.
468
+
469
+ Whitespace inside `<?= ?>` and `<%= %>` is optional, so the templates drop it:
470
+ nothing is lost and every sink of that shape is gained. Verified end to end
471
+ against the real sink rather than a fixture -- RCEKit generated
472
+ `RKYOOQS<?=69294*83370?>RKGFZQP`, RRDtool wrote it whole, and PHP answered
473
+ `RKYOOQS5777040780RKGFZQP`, which is the value RCEKit had computed for that
474
+ probe.
475
+
476
+ `jspx` keeps its spaces: it is an XML document whose root element carries
477
+ namespace attributes, and a sink that tokenises on whitespace was never going
478
+ to carry one. Two tests pin the rest -- the generated probe and the templates
479
+ themselves -- both written first and watched to fail.
480
+
481
+ ## [2.45.5] — 2026-09-25
482
+
483
+ ### Fixed
484
+
485
+ - **The `deser` shape differential answered from evidence it had not
486
+ validated.** Three ways, each found by pointing the method at a live
487
+ fastjson 1.2.83 endpoint under vulhub and asking why it said what it said.
488
+
489
+ **Drift landed on whichever form went last.** The differential is read across
490
+ requests, so anything changing with the request *index* rather than with the
491
+ payload -- a rate limiter backing off, a filling log, a warming cache -- fell
492
+ on the form sent last, and noise was sent last every time. There are two
493
+ noise probes now, first and last, bracketing the structured pair. They are
494
+ one noise body sent twice, byte for byte, because the bracket asks whether
495
+ the endpoint answered *the same probe* alike and anything the target can see
496
+ a difference in it may answer a difference to: a tail beginning `n` made
497
+ fastjson say "error parse new" where every other tail said "syntax error".
498
+ A carrier whose two ends disagree is held at `inconclusive`. The engine
499
+ de-duplicates by payload on the per-probe path but not on the aggregate path
500
+ where `deser` lives, which is what makes one body twice possible; a test pins
501
+ that, because the bracket becomes a single request if it ever changes and a
502
+ drift check comparing an answer with itself passes forever.
503
+
504
+ **An endpoint echoing its input looked like a parser.** A response carrying a
505
+ fragment of the payload that begins at the format's magic, and not carrying
506
+ it in the payload-free control, is an echo and the shape channel cannot be
507
+ read through it -- `inconclusive`, not `negative`, because the channel was
508
+ unreadable rather than empty. The sentinels are anchored on the magic's
509
+ position in the payload rather than its start, since a probe is wrapped for
510
+ its injection context first and some wrappers are longer than the magic: an
511
+ `xml_cdata` pickle probe begins `<![CDATA[gASV`. They are matched raw, case
512
+ folded (because `_signature` folds case and the two halves of one oracle must
513
+ agree), and encoded -- base64 aligned on bytes rather than characters and
514
+ read from the first three-byte boundary at or past the magic, hex from the
515
+ magic itself.
516
+
517
+ **A format declaring no magic got a verdict anyway.** Without a magic there
518
+ is no sentinel, so an echoing endpoint cannot be told from a parsing one --
519
+ and the differential answers regardless, because the noise form is random
520
+ where a structured pair is not. Every shipped format declares one; a
521
+ `--template-file` need not, and the answer to a control that cannot be
522
+ validated is to say so.
523
+
524
+ - **The page an operator plans traffic from still described three payloads.**
525
+ The bracket makes the shape oracle send four per ecosystem -- 20 requests per
526
+ carrier rather than 15, with the five the corpus ships -- and
527
+ `docs/reference.md` and the method's own docstring both still said three.
528
+ README advertises an execution plan that prints the exact probe count before
529
+ anything fires, so a page that undercounts it breaks that promise in the one
530
+ place someone reads before a monitored engagement.
531
+
532
+ Both now state four, and the arithmetic, and what the repeated probe is for.
533
+ A test holds the page to `SHAPE_FORMS` and to the ecosystem count, in both
534
+ directions: adding a form without updating the page fails it, and reverting
535
+ the page without touching the code fails it too. Nothing compared the two
536
+ before, which is how the count drifted the moment the form was added.
537
+
538
+ - **The `negative` evidence line described a comparison it had not made.** "The
539
+ endpoint answers all three forms alike" was asserted rather than observed.
540
+ It now names what actually collapsed: every form alike, a truncated stream
541
+ answered as a complete one, or the noise form answered as a structured one.
542
+
543
+ ### Withdrawn before release
544
+
545
+ - **A second route through the differential, reading noise as the control.** It
546
+ was written for the shape fastjson has -- a format resolving its type name
547
+ before the parse finishes answers a truncated stream exactly as it answers a
548
+ complete one, which collapses the original comparison on the very endpoints
549
+ it exists to find -- and it did reach `needs-review` there.
550
+
551
+ It is not in this release. "Noise differs from the structured pair" has too
552
+ many explanations that are not a parser. Drift and echo are closed above. The
553
+ third is not: a filter that refuses the random payload while letting the
554
+ structured ones through produces the same differential, no repetition exposes
555
+ it because the refusal is stable, and this verdict cannot see it -- refusal
556
+ is judged against the probe the verdict is anchored to, and that is the
557
+ well-formed one.
558
+
559
+ A format of that shape is reported `negative` by the shape differential and
560
+ reaches `deserialization-sink` through the DNS gadget instead, which is this
561
+ method's real oracle and rests on a callback rather than on the shape of a
562
+ response. Measured on the same target: 14 callbacks, each carrying a token it
563
+ could only have learned by resolving the name it was handed.
564
+
565
+ ## [2.45.4] — 2026-09-25
566
+
567
+ ### Fixed
568
+
569
+ - **A compressed response body is a body the oracles cannot read, so a target
570
+ that executed the probe was reported `negative`.** Every in-band oracle --
571
+ `reflected`, `eval`, `file`, `write`, `boolean` -- confirms by finding a
572
+ value the target computed in the response. Nothing in the delivery layer
573
+ undid `Content-Encoding`, so a gzipped body arrived as bytes that *contain*
574
+ the computed value and do not spell it, and the search for a decimal number
575
+ came back empty. The verdict was `negative`: not "could not read this", but
576
+ "the probes reached the target and it is not vulnerable".
577
+
578
+ Measured against Apache HugeGraph 1.2.0 under vulhub. The probe
579
+ `echo RKPDSWY$((258675+956113))RKIEMTB$(echo RKLUQKL)RKPDSWY`, delivered
580
+ through the Gremlin API, came back HTTP 200 carrying
581
+ `RKPDSWY1214788RKIEMTBRKLUQKLRKPDSWY` -- the sum and the substitution
582
+ collapse, both of `ReflectedMath`'s proofs. Present after gunzip, absent in
583
+ the raw bytes, and the run reported `negative` across 620 probes.
584
+
585
+ What made it survive this long is that the same target confirmed on the same
586
+ run's *other* probes: HugeGraph gzips its 200s and leaves its 4xx errors
587
+ plain, so `eval` confirmed through a 400 that carried a Groovy exception
588
+ message while the 200 that carried real command execution was invisible. A
589
+ run that finds something is a run nobody rereads.
590
+
591
+ Both HTTP entry points decode now -- the verification/detection path and the
592
+ `--verify-chain` path -- through one helper, on the success body and the
593
+ `HTTPError` body alike, since an evaluator surfacing its value in a 500 is
594
+ the case that path exists for. The header is honoured whether or not RCEKit
595
+ asked for it, because HugeGraph compresses a response to a request that sent
596
+ no `Accept-Encoding` at all.
597
+
598
+ `gzip` and `deflate` only, both from the standard library, and `deflate`
599
+ tries the zlib wrapper before the raw stream because servers disagree about
600
+ which one that name means. `br` and `zstd` need a third-party module and this
601
+ tool has none to add; a body in either is returned as its raw bytes. Nothing
602
+ here raises: a body that claims an encoding it is not in comes back
603
+ undecompressed rather than ending a run that has already spent hundreds of
604
+ probes.
605
+
606
+ This does not widen what `confirmed` means. The oracle is unchanged -- a
607
+ computed value present in the response and absent from a payload-free
608
+ control -- and nothing is promoted into it. What changes is that the response
609
+ the oracle reads is the response the target sent.
610
+
611
+ - **Decompression is bounded, because compression is where the target chooses
612
+ how much memory RCEKit allocates.** Reading a body at all is new here, and
613
+ one-shot decompression handed a blank cheque to the one party in the exchange
614
+ that is untrusted by definition -- the thing being tested. Measured: 130,478
615
+ bytes of gzip expand to 134,217,728, and 203,860 bytes hold 200 MB of zeros,
616
+ a ratio of 1:1028. This path runs once per probe.
617
+
618
+ The ceiling is 32 MiB per encoding step. Whatever was produced before it is
619
+ kept rather than discarded: the value the oracle wants is short, so searching
620
+ 32 MB is a far better answer than searching nothing, and the only risk the
621
+ ceiling carries is a false negative -- never a false `confirmed`. A truncated
622
+ read is logged rather than passed off as a clean one, since a `negative`
623
+ decided on a partial body is not the same claim as a `negative` decided on
624
+ the whole of it.
625
+
626
+ Multi-member streams are still followed. `gzip.decompress` followed them, and
627
+ a ceiling that quietly became a truncation for a server that concatenates
628
+ members would be the same false negative in a new place.
193
629
 
194
630
  ## [2.45.3] — 2026-09-24
195
631
 
@@ -2264,6 +2700,11 @@ this file and have not been restated here.
2264
2700
 
2265
2701
 
2266
2702
  [Unreleased]: https://github.com/kabiri-labs/rcekit/compare/v2.36.0...HEAD
2703
+ [3.0.0]: https://github.com/kabiri-labs/rcekit/compare/v2.45.7...v3.0.0
2704
+ [2.45.7]: https://github.com/kabiri-labs/rcekit/compare/v2.45.6...v2.45.7
2705
+ [2.45.6]: https://github.com/kabiri-labs/rcekit/compare/v2.45.5...v2.45.6
2706
+ [2.45.5]: https://github.com/kabiri-labs/rcekit/compare/v2.45.4...v2.45.5
2707
+ [2.45.4]: https://github.com/kabiri-labs/rcekit/compare/v2.45.3...v2.45.4
2267
2708
  [2.45.3]: https://github.com/kabiri-labs/rcekit/compare/v2.45.2...v2.45.3
2268
2709
  [2.45.2]: https://github.com/kabiri-labs/rcekit/compare/v2.45.1...v2.45.2
2269
2710
  [2.45.1]: https://github.com/kabiri-labs/rcekit/compare/v2.45.0...v2.45.1
@@ -83,12 +83,16 @@ in the clearly separated **Detection methods** section of `rcekit.py`
83
83
 
84
84
  To add one:
85
85
 
86
- 1. Subclass `DetectionMethod`: set `name` and `tier` (`confirmed` for
87
- execution-proven methods, `needs-review` for candidates — the two tiers are
88
- never merged), then implement `applicable`, `build_probes`, and `confirm`.
86
+ 1. Subclass `DetectionMethod`: set `name` and `tier` (`executed` for
87
+ execution-proven methods; otherwise a tier named for what the target did,
88
+ like `timing-sink`, and `needs-review` only when the signal really is a
89
+ candidate — tiers are never merged), then implement `applicable`,
90
+ `build_probes`, and `confirm`. Also set `question`, which says whether this
91
+ asks *did the target execute my input* or something else; it does not follow
92
+ from the tier.
89
93
  2. Compute the expected value **locally** with random inputs and compare it
90
94
  against the payload-free control (reuse `self._search`, which is
91
- encoding-aware). A verdict is `confirmed` only when the computed value is
95
+ encoding-aware). A verdict is `executed` only when the computed value is
92
96
  present and absent from the control.
93
97
  3. **Declare what the method costs and what it says.** `safety` is the
94
98
  `--verify-active-risk` rung it needs: `safe` when it only makes the target
@@ -119,10 +123,20 @@ To add one:
119
123
  bench case proves it reaches that verdict against the real software. It is
120
124
  not a per-PR gate, and a method whose case cannot run yet still ships --
121
125
  with the reason written down in `tests/bench/README.md`, not left implicit.
122
- 7. Do not add a row to the README's **CVE table** without a bench run behind it.
123
- That table is the reproduced claim, and it carries the version it was last
124
- verified at. The **methods table** describes capability, and the unit suite
125
- is what backs that.
126
+ 7. A row in the README's **coverage ledger** needs a run behind it, and its
127
+ `Bench case` column has to say which kind. `yes` means `tests/bench/`
128
+ reproduces it, target and negative control; that is the strong claim and it
129
+ carries the version it was last verified at. `not yet` means it was measured
130
+ by hand against the same build and nobody can re-run it on demand — honest,
131
+ weaker, and the reason the column is there. Never fill the column in ahead of
132
+ the case: a test reads it against `tests/bench/cases/` and
133
+ `confirmation-gifs/`.
134
+
135
+ Leave the `Advisory` column empty where the verdict does not depend on the
136
+ patch. A row proving a method against real software is worth recording;
137
+ labelling it with a CVE it does not reproduce is the overclaim the ledger
138
+ exists to prevent. The **methods table** describes capability in the
139
+ abstract, and the unit suite is what backs that.
126
140
 
127
141
  ## Adding a bench case
128
142
 
@@ -152,7 +166,7 @@ New detection coverage should arrive with a case. Every case needs a **negative
152
166
  control** — the runner refuses to load one without it — because a benchmark with
153
167
  no controls rewards aggressive probing instead of measuring accuracy. The
154
168
  control may be a patched build, the same target probed for the wrong class, or a
155
- weaker method that must stay below `confirmed` on a target where it happens to
169
+ weaker method that must stay below `executed` on a target where it happens to
156
170
  be right.
157
171
 
158
172
  Run your case against the real target before submitting, and update the README