supercov 1.1.1 → 2.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,8 +1,8 @@
1
- ![Code quality and coverage for coding agents](https://raw.githubusercontent.com/supercorp-ai/supercov/main/supercov.jpg)
1
+ ![Coverage, security and code quality for coding agents](https://raw.githubusercontent.com/supercorp-ai/supercov/main/supercov.jpg)
2
2
 
3
- **Code quality and coverage for coding agents**
3
+ **Coverage, security and code quality for coding agents**
4
4
 
5
- **Supercov tells your coding agent what to fix and what to test.** It scores your code quality with [Jev](https://typesafe.ai), runs the test command you already use, and turns uncovered paths into small, actionable queries. Your agent picks a target, writes a focused test or a focused refactor, proves what improved, and keeps going.
5
+ **Supercov tells your coding agent what to fix and what to test.** It scores your code quality and flags security risks with [Jev](https://typesafe.ai), runs the test command you already use, and turns uncovered paths into small, actionable queries. Your agent picks a target, writes a focused test or a focused refactor, proves what improved, and keeps going.
6
6
 
7
7
  Paste this to your coding agent to start:
8
8
 
@@ -56,6 +56,18 @@ Jev charges for what it reads and nothing for what it writes, so a megabyte of
56
56
  source costs a little over a cent. Answers are cached by content, so a second
57
57
  run pays only for what changed. See [Understanding quality](docs/quality.md).
58
58
 
59
+ ## Find security risks
60
+
61
+ ```bash
62
+ npx supercov security
63
+ npx supercov security patch # what a change introduced
64
+ ```
65
+
66
+ Twelve security checks asked of every file with the same key, from injection
67
+ and secrets in source to paths and redirects taken from a request. Nothing is
68
+ averaged: a file is clean, or it names what fired, with the line and the code on
69
+ it where Jev confirms one. See [Security surface](docs/security.md).
70
+
59
71
  ## Start with the suite you already have
60
72
 
61
73
  ```bash
@@ -102,6 +114,19 @@ npx supercov -- npm test
102
114
  npx supercov diff <previous-run-id> latest
103
115
  ```
104
116
 
117
+ Want to hand the evidence to a person? Generate one private, interactive file:
118
+
119
+ ```bash
120
+ npx supercov report
121
+ ```
122
+
123
+ The report opens in your browser with up to ten recent runs in a visual history.
124
+ Selecting a run compares it with the previous snapshot automatically. It has no
125
+ server, external assets, account, or file picker. It is written to
126
+ `.supercov/reports/`, where the store's own `.gitignore` keeps it out of git.
127
+ You can attach that one file directly to a pull request; a reviewer downloads
128
+ it and double-clicks it.
129
+
105
130
  Use whichever complete test command the repository already trusts:
106
131
 
107
132
  ```bash
@@ -233,7 +258,7 @@ Java and Kotlin have no registry of their own here, so a JVM project takes the b
233
258
  | PHP | Coming soon | — | — |
234
259
  | C | Coming soon | — | — |
235
260
 
236
- The `npx` path requires Node.js 22 or newer; the other installs need no Node. Rust support currently uses Rust 1.95; cargo-nextest 0.9.138 and 0.9.140 are supported. Python support requires CPython 3.12 or newer and measures pytest and unittest runs. Ruby support requires Ruby 3.3 or newer (3.4 or newer for full measurement) and measures RSpec, Minitest, test-unit and Cucumber runs. Go support requires Go 1.22 or newer and measures `go test`. Java and Kotlin support requires a JDK 17 or newer with Maven or Gradle, and measures every JUnit Platform engine -- JUnit 5, JUnit 4 through Vintage, Kotest and Spock -- as well as TestNG.
261
+ The `npx` path requires Node.js 22 or newer; the other installs need no Node. Rust support currently uses Rust 1.95; cargo-nextest 0.9.138 and 0.9.140 are supported. Python support requires CPython 3.9 or newer and measures pytest and unittest runs. Ruby support requires Ruby 3.3 or newer (3.4 or newer for full measurement) and measures RSpec, Minitest, test-unit and Cucumber runs. Go support requires Go 1.22 or newer and measures `go test`. Java and Kotlin support requires a JDK 17 or newer with Maven or Gradle, and measures every JUnit Platform engine -- JUnit 5, JUnit 4 through Vintage, Kotest and Spock -- as well as TestNG.
237
262
 
238
263
  ## Supported operating systems and architectures
239
264
 
@@ -245,7 +270,7 @@ The CLI is a single native binary. `npx supercov` selects the build for the mach
245
270
  | Linux | arm64, x64 | glibc 2.28 or newer (Debian 10, Ubuntu 18.10, RHEL 8 and later), or musl (Alpine); chosen automatically |
246
271
  | Windows | arm64, x64 | Windows 10 or newer |
247
272
 
248
- JavaScript, TypeScript, Python, Ruby, Rust, Go, Java and Kotlin suites are verified on Linux, macOS and Windows -- CPython 3.12, 3.13 and 3.14, Ruby 3.3, 3.4 and 4.0, and Go 1.22, 1.25, 1.26 and 1.27 on every one, and JDK 17 and 21 on every one that has them (Temurin publishes no 17 for arm64 Windows, so that pairing is 21 alone). Containers, VMs, and remote executors run the same Linux builds. Every release publishes the same binaries to npm, PyPI (`supercov-cli`, a wheel per platform) and RubyGems (`supercov`, a gem per platform except arm64 Windows, which Ruby has no platform for), and the source to crates.io (`supercov`).
273
+ JavaScript, TypeScript, Python, Ruby, Rust, Go, Java and Kotlin suites are verified on Linux, macOS and Windows -- CPython 3.9 through 3.14 (on arm64 Windows, 3.9 and 3.10 as the x64 builds, since python.org publishes no arm64 build before 3.11), Ruby 3.3, 3.4 and 4.0, and Go 1.22, 1.25, 1.26 and 1.27 on every one, and JDK 17 and 21 on every one that has them (Temurin publishes no 17 for arm64 Windows, so that pairing is 21 alone). Containers, VMs, and remote executors run the same Linux builds. Every release publishes the same binaries to npm, PyPI (`supercov-cli`, a wheel per platform) and RubyGems (`supercov`, a gem per platform except arm64 Windows, which Ruby has no platform for), and the source to crates.io (`supercov`).
249
274
 
250
275
  ## Supported test suites
251
276
 
@@ -286,6 +311,9 @@ npx supercov runs latest line app/checkout/session.ts:64
286
311
 
287
312
  # What changed between two runs
288
313
  npx supercov diff <previous-run-id> latest
314
+
315
+ # A portable interactive report for a person
316
+ npx supercov report
289
317
  ```
290
318
 
291
319
  Collections accept `--limit` and `--offset` and print a copyable next-page command. Machine-readable output is available with `--json` when an integration needs it.
@@ -299,9 +327,9 @@ The Supercov CLI does not contact a Supercov service during a coverage run. Pack
299
327
  Supercov does not rewrite your source files, tests, imports, reporter list, runner configuration, dependency tree, or normal build output. An existing user-created `supercov/` directory is never adopted.
300
328
 
301
329
  ```bash
302
- npx supercov clean --dry-run # preview cleanup
303
- npx supercov clean --keep 20 # keep the 20 newest runs
304
- npx supercov clean # remove all runs and the build cache
330
+ npx supercov runs clean --dry-run # preview cleanup
331
+ npx supercov runs clean --keep 20 # keep the 20 newest runs
332
+ npx supercov runs clean # remove all runs and the build cache
305
333
  ```
306
334
 
307
335
  ## Documentation
@@ -315,6 +343,7 @@ npx supercov clean # remove all runs and the build cache
315
343
  - [Supported languages and test suites](https://supercov.com/docs/supported-suites)
316
344
  - [Understanding coverage](https://supercov.com/docs/coverage-model)
317
345
  - [Runs and evidence](https://supercov.com/docs/evidence)
346
+ - [Portable HTML reports](https://supercov.com/docs/reports)
318
347
  - [Files, privacy, and cleanup](https://supercov.com/docs/workspace-isolation)
319
348
  - [Trusting a result](https://supercov.com/docs/verification)
320
349
  - [Speed and storage](https://supercov.com/docs/performance)
@@ -326,3 +355,6 @@ npx supercov clean # remove all runs and the build cache
326
355
  ## Contributors
327
356
 
328
357
  - [@DaVinciTachyon](https://github.com/DaVinciTachyon)
358
+ - [@Uhbif2](https://github.com/Uhbif2)
359
+ - [@maik-intellicoach](https://github.com/maik-intellicoach)
360
+ - [@untitaker](https://github.com/untitaker)
package/docs/cli.md CHANGED
@@ -21,12 +21,16 @@ npx supercov --help
21
21
  | Inspect one assertion and its flows | `npx supercov runs latest assertion <id>` |
22
22
  | Read matching current source code | `npx supercov runs latest source <path>` |
23
23
  | Compare two runs | `npx supercov diff <older> <newer>` |
24
+ | Open an interactive report | `npx supercov report` |
24
25
  | Find the tests a change affects | `npx supercov runs latest tests affected` |
25
26
  | Combine shards | `npx supercov merge <id> <id> [...]` |
26
- | Remove local data | `npx supercov clean` |
27
+ | Remove stored runs | `npx supercov runs clean` |
27
28
  | Assess code quality with Jev | `npx supercov quality` |
28
29
  | See only files with findings | `npx supercov quality gaps` |
29
30
  | Review what a change introduced | `npx supercov quality patch` |
31
+ | Remove saved assessments | `npx supercov quality clean` |
32
+ | Find security surface with Jev | `npx supercov security` |
33
+ | Review what a change introduced, security only | `npx supercov security patch` |
30
34
  | Read bundled guides | `npx supercov docs` |
31
35
 
32
36
  ## Assess source quality
@@ -39,6 +43,10 @@ supercov quality file src/a.ts # one file, every check
39
43
  supercov quality scope # which files are assessed, and why
40
44
  supercov quality snapshots # saved assessments
41
45
  supercov quality diff <older> <newer>
46
+ supercov security # twelve security checks, every file
47
+ supercov security gaps # only files something fired on
48
+ supercov security patch --base origin/main
49
+ supercov security --run latest # flagged files no test exercises
42
50
  ```
43
51
 
44
52
  With no argument the subject is the repository you are standing in. Every
@@ -56,7 +64,8 @@ the exact requests and contacts nothing.
56
64
  health, which properties appeared, and which files entered or left the scope.
57
65
 
58
66
  Assessing needs a TypeSafe AI API key in `TYPESAFE_API_KEY`; reading a saved
59
- assessment does not. The command prints a cost estimate before sending anything
67
+ assessment does not. `TYPESAFE_BASE_URL` and `TYPESAFE_DEFAULT_MODEL` use Jev
68
+ through another provider, such as OpenRouter. The command prints a cost estimate before sending anything
60
69
  and caches answers by content, so a second run pays only for what changed.
61
70
 
62
71
  See [Understanding quality](https://supercov.com/docs/quality) for what the
@@ -340,6 +349,30 @@ The same filters can focus a comparison:
340
349
  npx supercov diff <older-run> <newer-run> --kind e2e
341
350
  ```
342
351
 
352
+ ## Generate a portable HTML report
353
+
354
+ ```sh
355
+ npx supercov report
356
+ npx supercov report <run-id>
357
+ npx supercov report latest --compare <older-run-id>
358
+ ```
359
+
360
+ `report` turns stored local evidence into one self-contained interactive HTML
361
+ file and opens it in the default browser. It does not rerun tests, start a
362
+ server, load external assets, or upload anything. The default report includes
363
+ up to ten recent runs. Selecting a run compares it with the previous saved run
364
+ automatically.
365
+
366
+ ```sh
367
+ npx supercov report --runs 5
368
+ npx supercov report --output artifacts/supercov-report.html --no-open
369
+ ```
370
+
371
+ Use `--runs N` to include up to 20 stored runs. Reports are capped at 24 MB so
372
+ they remain suitable for a normal pull-request attachment; use `--runs 1` if a
373
+ large project exceeds that target. See [Portable HTML reports](reports.md) for
374
+ the privacy and source-integrity rules.
375
+
343
376
  ## Find a smaller test set
344
377
 
345
378
  ```sh supercov
@@ -386,15 +419,31 @@ incompatible merge rather than publishing a misleading aggregate.
386
419
  ## Clean local data
387
420
 
388
421
  ```sh supercov
389
- npx supercov clean --dry-run
390
- npx supercov clean --keep 20
391
- npx supercov clean
422
+ npx supercov runs clean --dry-run
423
+ npx supercov runs clean --keep 20
424
+ npx supercov runs clean
392
425
  ```
393
426
 
394
- By default, `clean` removes all stored runs and the isolated build cache.
395
- `--keep N` keeps the newest N runs. Cleanup removes only marker-owned Supercov
427
+ By default, `runs clean` removes all stored runs and the isolated build cache.
428
+ `--keep N` keeps the newest N runs, by when they started rather than by their
429
+ identifiers, which carry no order. Cleanup removes only marker-owned Supercov
396
430
  storage.
397
431
 
432
+ Saved quality assessments live in their own lane and are never removed here: an
433
+ assessment costs money and cannot be reproduced from the repository. Remove
434
+ those deliberately, with the same options:
435
+
436
+ ```sh supercov
437
+ npx supercov quality clean --dry-run
438
+ npx supercov quality clean --keep 5
439
+ npx supercov quality clean
440
+ ```
441
+ If Supercov itself fails to publish a run the tests already paid for, it keeps
442
+ that run's raw evidence in `.supercov/failed-evidence/<run id>` and says so in
443
+ the error. Report the failure with that directory: it is what diagnoses it. Only
444
+ a full `runs clean` reclaims it -- `--keep N` leaves it alone -- and the summary says
445
+ when it goes.
446
+
398
447
  ## Read bundled documentation
399
448
 
400
449
  ```sh supercov
package/docs/evidence.md CHANGED
@@ -101,15 +101,23 @@ instrumented build cache may use more space than the compressed run itself.
101
101
  Nothing is pruned in the background.
102
102
 
103
103
  ```sh supercov
104
- npx supercov clean --dry-run
105
- npx supercov clean --keep 20
106
- npx supercov clean
104
+ npx supercov runs clean --dry-run
105
+ npx supercov runs clean --keep 20
106
+ npx supercov runs clean
107
107
  ```
108
108
 
109
109
  Preview cleanup first. The final command removes all runs and the isolated build
110
110
  cache; `--keep 20` preserves the 20 newest runs. Cleanup removes only
111
111
  marker-owned Supercov data.
112
112
 
113
+ ### Evidence of a run that failed to publish
114
+
115
+ A suite pays for its measurement in wall clock. If Supercov itself cannot
116
+ publish a run the tests already finished, the run's raw evidence is kept in
117
+ `.supercov/failed-evidence/<run-id>/` rather than going with the work directory,
118
+ and the error names the path. Attach that directory to the bug report.
119
+ `runs clean` reclaims it; `--keep N` does not.
120
+
113
121
  ## Understand test kinds
114
122
 
115
123
  Reports group tests by kind, such as unit, integration or E2E. Supercov uses
@@ -21,6 +21,16 @@ Measure code quality with npx supercov.
21
21
 
22
22
  See [Understanding quality](https://supercov.com/docs/quality).
23
23
 
24
+ For security surface with the same key, ask:
25
+
26
+ ```text supercov-prompt
27
+ Find security issues with npx supercov security and show me what it found.
28
+ ```
29
+
30
+ Supercov asks twelve security checks of every file and shows each finding
31
+ with the line and the code on it. See
32
+ [Security surface](https://supercov.com/docs/security).
33
+
24
34
  When it finishes, review the test change and coverage comparison in your
25
35
  conversation. Ask separately if you want a commit or pull request.
26
36
 
@@ -88,7 +98,16 @@ npx supercov runs latest line app/checkout/session.ts:64
88
98
  outcomes and MC/DC witnesses. `line` shows the obligations and tests associated
89
99
  with one source line.
90
100
 
91
- ### 4. Add a test and prove the gain
101
+ #For a human review, create one interactive file from the same stored evidence:
102
+
103
+ ```sh
104
+ npx supercov report
105
+ ```
106
+
107
+ It opens locally with run comparison, prioritized gaps, source context, test
108
+ attribution, and scope. No test rerun or server is required.
109
+
110
+ ## 4. Add a test and prove the gain
92
111
 
93
112
  The agent writes one focused test with a meaningful assertion. It then reruns
94
113
  the same complete command and compares the two runs:
@@ -119,7 +138,7 @@ if anything is missing.
119
138
  - for Rust, the Rust 1.95 toolchain;
120
139
  - for Go, Go 1.22 or newer;
121
140
  - for Java and Kotlin, JDK 17 or newer, with Maven or Gradle;
122
- - for Python, CPython 3.12 or newer with pytest or unittest;
141
+ - for Python, CPython 3.9 or newer with pytest or unittest;
123
142
  - for Ruby, Ruby 3.4 or newer with RSpec, Minitest, test-unit or Cucumber (3.3 measures lines, methods and simple branches only).
124
143
 
125
144
  The CLI is a native binary. `npx supercov` picks the build for your operating
@@ -152,9 +171,9 @@ Supercov does not rewrite your source, tests, imports, runner configuration,
152
171
  dependencies, or ordinary build output.
153
172
 
154
173
  ```sh supercov
155
- npx supercov clean --dry-run # preview what would be removed
156
- npx supercov clean --keep 20 # keep the 20 newest runs
157
- npx supercov clean # remove all runs and the build cache
174
+ npx supercov runs clean --dry-run # preview what would be removed
175
+ npx supercov runs clean --keep 20 # keep the 20 newest runs
176
+ npx supercov runs clean # remove all runs and the build cache
158
177
  ```
159
178
 
160
179
  If the first run does not look right, go to [Troubleshooting](troubleshooting.md)
@@ -21,6 +21,11 @@ The summary separates:
21
21
  | Test command | The wrapped command, including browser, VM, or remote latency |
22
22
  | Evidence publication | Validating and storing the completed run |
23
23
 
24
+ The line Supercov prints when a run ends also reports publication: analysing
25
+ the evidence once for the stored query views, the summary and the assertion
26
+ map. That is why the first query after a run opens at once rather than
27
+ analysing the evidence itself.
28
+
24
29
  The first `npx` invocation may also download the package. That download happens
25
30
  before Supercov starts and is not coverage-engine overhead.
26
31
 
@@ -64,6 +69,48 @@ compare a cold package, browser, build, or VM cache with a warm one. Supercov
64
69
  never runs the test command a second time automatically because suites may write
65
70
  data, call paid services, or be intentionally non-repeatable.
66
71
 
72
+ ### Python subprocesses
73
+
74
+ Interpreter count matters as well as test count. Each child initializes the
75
+ Python runtime before executing user code, so a suite that launches thousands
76
+ of short-lived interpreters can have more overhead than a single-process suite
77
+ with the same number of tests. Tight child-startup deadlines can also expire
78
+ before the child produces its first output.
79
+
80
+ Supercov prepares the run's probe index once and lets later interpreters load
81
+ file and decision data as needed. The compiled-module cache is separate: it
82
+ avoids compiling unchanged measured source, while the prepared index avoids
83
+ repeatedly parsing and indexing the entire coverage plan.
84
+
85
+ Optional unittest and concurrency adapters install when their libraries are
86
+ imported, rather than importing those libraries into every helper process.
87
+ Libraries already loaded when measurement starts are patched immediately.
88
+
89
+ For development, `python3 scripts/python-startup-benchmark.py` measures cold
90
+ initialization and repeated child startup against a synthetic plan. Use
91
+ `--runtime /path/to/checkout/runtime/python` to compare implementations. This
92
+ isolates startup cost; it does not predict a complete suite's slowdown, which
93
+ also includes measured execution, concurrency and evidence publication.
94
+
95
+ For an end-to-end subprocess workload, build a release binary and run:
96
+
97
+ ```sh
98
+ cargo build --release -p supercov
99
+ python3 scripts/python-subprocess-benchmark.py \
100
+ --output /tmp/supercov-subprocess-benchmark \
101
+ --children 1000 --tests 400 \
102
+ --binary candidate="$PWD/target/release/supercov"
103
+ ```
104
+
105
+ Use a new output directory for each experiment. The fixture defaults to 321
106
+ source files and saves plain/measured timings, CLI phase timings, child latency,
107
+ test outcomes and coverage summaries. Add another `--binary label=/path` to
108
+ compare builds; measured coverage and outcomes must agree. Vary `--calls` for
109
+ hot execution, `--statements` for plan size, `--decision-width` for wider MC/DC
110
+ regions, or `--workers` for concurrent children. `--mode noop` isolates helpers
111
+ that run no measured source; `--mode script` exercises native entry scripts and
112
+ allows the corrected unmeasured-file denominator to differ between versions.
113
+
67
114
  ## Understand disk usage
68
115
 
69
116
  Each completed run stores compressed evidence and metadata under
@@ -74,9 +121,9 @@ The isolated workspace may be larger because it can contain an instrumented
74
121
  build cache. Supercov does not delete history in the background.
75
122
 
76
123
  ```sh supercov
77
- npx supercov clean --dry-run
78
- npx supercov clean --keep 20
79
- npx supercov clean
124
+ npx supercov runs clean --dry-run
125
+ npx supercov runs clean --keep 20
126
+ npx supercov runs clean
80
127
  ```
81
128
 
82
129
  Use `--dry-run` to preview cleanup. Keep enough run history for active reviews
package/docs/quality.md CHANGED
@@ -35,6 +35,27 @@ Set it the way your environment already sets secrets:
35
35
  Reading a saved assessment never needs a key, and `--dry-run` prints the exact
36
36
  requests without sending them.
37
37
 
38
+ ## Use Jev through another provider
39
+
40
+ Supercov reads the same variables as TypeSafe's own SDKs, so a provider that
41
+ serves Jev, such as [OpenRouter](https://openrouter.ai), needs only its URL, its
42
+ name for the model, and its key:
43
+
44
+ ```bash
45
+ export TYPESAFE_BASE_URL=https://openrouter.ai/api
46
+ export TYPESAFE_DEFAULT_MODEL=typesafe/jev-1.13
47
+ export TYPESAFE_API_KEY=sk-or-...
48
+ npx supercov quality
49
+ ```
50
+
51
+ | Variable | What it is | Default |
52
+ | --- | --- | --- |
53
+ | `TYPESAFE_BASE_URL` | where requests go | `https://api.typesafe.ai` |
54
+ | `TYPESAFE_DEFAULT_MODEL` | the model asked for | `jev-1.13.0` |
55
+
56
+ Snapshots record the model, and `quality diff` compares only snapshots from the
57
+ same one.
58
+
38
59
  ## Start with the repository
39
60
 
40
61
  ```bash
@@ -100,15 +121,16 @@ the merge base so commits other people landed after you branched are not
100
121
  counted as yours.
101
122
 
102
123
  It asks the twelve properties differentially, whether the new version shows
103
- something the old one did not, and adds six checks that only make sense for a
104
- change:
124
+ something the old one did not, adds three checks that only make sense for a
125
+ change, and asks the twelve [security surface](security.md) checks the same
126
+ differential way:
105
127
 
106
- - a credential written into source
107
- - untrusted input interpolated into a query
108
- - a change to how the system decides who may do what
109
128
  - a test that now checks less than it did
110
129
  - a database schema or data migration
111
130
  - debugging left behind
131
+ - a secret in source, an injection sink, unescaped output, a path or
132
+ destination taken from a caller, a handler with no visible authorisation, and
133
+ the rest of the security catalog, each only if the change introduced it
112
134
 
113
135
  Output lists only files where something appeared. A change that introduces
114
136
  nothing says so in one line.
@@ -0,0 +1,181 @@
1
+ # Security surface
2
+
3
+ `supercov security` tells you what security-relevant surface each file shows:
4
+ twelve named checks, asked file by file, each mapped to the weakness classes it
5
+ stands for. There is no score. A file is clean, or it names what fired.
6
+
7
+ Judgments come from [Jev](https://typesafe.ai), the same way `supercov quality`
8
+ gets them. Set `TYPESAFE_API_KEY`, and optionally another provider, as described
9
+ in [quality](quality.md); reading a saved assessment never needs a key.
10
+
11
+ ## Start with the repository
12
+
13
+ ```bash
14
+ npx supercov security
15
+ ```
16
+
17
+ ```
18
+ Security: 6 of 73 files flagged, 67 clean; 4 confirmed at a line.
19
+ destination_from_input 3, injection_sink 1, missing_authorization 1, unsafe_code_execution 1
20
+ Catalog security-v4, model jev-1.13.0, snapshot q_5e1c…
21
+
22
+ Flagged:
23
+ src/routes/invoices.ts
24
+ 0.91 injection_sink
25
+ line 42 0.88 db.query(`select * from invoices where id = ${req.params.id}`)
26
+ src/gateways/sseToStdio.ts
27
+ 0.77 destination_from_input (file-level only)
28
+ ```
29
+
30
+ Every line is a claim you can check against the file in seconds. The value is
31
+ the model's probability that the pattern is present; nothing is averaged,
32
+ because nothing in the evidence supported averaging security answers.
33
+
34
+ ## Two passes, two tiers
35
+
36
+ Every file is asked the twelve questions once. For the checks that fired at
37
+ all, the parser lists every call, literal, route handler and request-body
38
+ spread the file has (Python, JavaScript and TypeScript), or every template
39
+ expression and configuration value (templates and config files), and the
40
+ model says which of the fired checks, if any, each line shows, one Choice
41
+ per line. Each classified line then gets the check's own question and a
42
+ triage question: would a careful reviewer report this, or dismiss it. A finding confirmed at a line is shown with the line and
43
+ the code on it, and with the line where the outside value enters the program
44
+ when the model can name one; a finding only the first pass made is marked
45
+ file-level only. Three findings of one check within ten lines count as one.
46
+ Authorisation findings stay file-level. Whether a guard reaches a handler
47
+ and which callers it admits is a fact that lives in the wiring, not the
48
+ handler, and no per-line question has been found that decides it: every
49
+ attempt measured on the labelled corpus fired on unguarded handlers the
50
+ application leaves open by design as often as on the ones it should not.
51
+
52
+
53
+ Measured with this command on fifteen held-out repositories of a
54
+ 140-repository labelled corpus, whole repositories, at finding level: F1
55
+ 0.47 with 54% precision and 41% recall, where Semgrep scores 0.14, the
56
+ general agentic LLM scanners 0.50 to 0.60, and the two leaders 0.76 and
57
+ 0.77, at about two cents per repository against 30 cents to 4 dollars for
58
+ the agentic scanners. Injection, secrets, path, redirect and mass
59
+ assignment are found at 70 to 90% recall; authorisation and authentication
60
+ at 10%, because the guard that decides them lives in another file. A file
61
+ too large for one request is read in windows at declaration boundaries.
62
+
63
+ ## Across files
64
+
65
+ The second pass also labels every function of every file, on the same
66
+ request: does it take caller-supplied data, does a parameter reach a
67
+ dangerous operation unprotected inside it, does it sanitise or authorise what
68
+ it receives, and which kind of operation. Supercov then resolves the file's
69
+ imports itself, pairs every function that takes outside data with every
70
+ imported function that reaches a sink without sanitising, where the caller's
71
+ body names the callee, and confirms each pair with one question that carries
72
+ both function bodies. A confirmed path is printed as a call site and a sink:
73
+
74
+ ```
75
+ Cross-file paths confirmed: 2
76
+ 0.79 injection_sink app/routes/index.js:19 index -> app/routes/allocations.js:11 displayAllocations
77
+ 0.70 injection_sink app/routes/benefits.js:11 BenefitsHandler -> app/data/benefits-dao.js:2 BenefitsDAO
78
+ ```
79
+
80
+ The model never sees two files at once until a path is confirmed; the graph
81
+ is host code over resolved imports, one hop.
82
+
83
+ With `--run`, a second kind of edge joins the first: the run's own record of
84
+ which tests executed which lines. A sink function and an entry function that
85
+ took outside data under the same test are paired and confirmed the same way,
86
+ whether or not any import connects them. That is how a handler that reaches
87
+ its sink through a lookup table, a registry, a decorator or a framework is
88
+ followed; no import graph has that edge, and no trace is attempted. On a
89
+ library whose views dispatch to their sinks by name, this took recall of the
90
+ labelled weaknesses from 42% to 75%. Without a run, nothing here is asked. Measured on a third of the same
91
+ corpus, confirmed paths fired on 0.1% of unlabelled files, and on real CVEs
92
+ in mature projects this stage is where the recall has to come from, because
93
+ a real vulnerability rarely sits in one file. JavaScript, TypeScript and
94
+ Python have parsers for this; other languages get the twelve questions and
95
+ the pattern candidates only.
96
+
97
+ ## The twelve checks
98
+
99
+ | Check | Stands for |
100
+ | --- | --- |
101
+ | secret_in_source | a credential or signing secret written into the source (CWE-798, 259, 321) |
102
+ | injection_sink | a query or command assembled from a caller's value (CWE-89, 78, 77, 943, 90) |
103
+ | unsafe_code_execution | eval, exec, pickle or unsafe deserialisation of outside data (CWE-94, 95, 502, 1336) |
104
+ | unescaped_output | an outside value written into markup without escaping (CWE-79, 116) |
105
+ | path_from_input | a file operation on a caller-controlled path, name or type (CWE-22, 73, 434) |
106
+ | missing_authorization | a handler acting on a caller-chosen record with no visible check (CWE-862, 863, 639, 284, 306) |
107
+ | weak_authentication | a token, session or password handled in a way that weakens it (CWE-287, 613, 614, 347, 384, 1004) |
108
+ | weak_cryptography | a primitive weak for its purpose, or predictable randomness for a secret (CWE-327, 328, 326, 330, 338, 916) |
109
+ | sensitive_data_exposure | a secret, token or internal detail in a log or a response (CWE-200, 209, 532, 312) |
110
+ | destination_from_input | a request or redirect going wherever a caller's value says (CWE-918, 601) |
111
+ | unchecked_mass_assignment | a request body copied wholesale into a record (CWE-915, 1321) |
112
+ | insecure_configuration | a provided protection switched off or weakened, or debug mode left on (CWE-16, 295, 352, 611, 942, 1004, 489, 215) |
113
+
114
+ Each asks whether a pattern is present, never whether it is exploitable, and
115
+ each states its exception. A value is from outside the program when it arrives
116
+ at run time from a caller, request, message, upload or third-party service; the
117
+ operator's own configuration, arguments, environment and shipped files are not.
118
+
119
+ ## What is known about each check
120
+
121
+ Measured on 2026-09-20 against constructed pairs, a clean floor, and RealVuln,
122
+ a corpus of 140 deliberately vulnerable repositories with 3,381 labelled
123
+ findings in Python, TypeScript and JavaScript:
124
+
125
+ - every check caught its constructed positive at 0.97 or above and stayed quiet
126
+ on a matched safe file, both when asked of a file and of a change;
127
+ - on the labelled corpus, the check a finding's weakness class maps to fired on
128
+ 68% of files carrying a finding, from 96% for secrets in source down to 43%
129
+ for missing authorisation, and on 5 of 76 certified false-positive traps;
130
+ - on the same files and labels, Semgrep reached 10%, SonarQube 18% and Snyk 21%.
131
+
132
+ `missing_authorization` is the weak one and is shown with that number. A
133
+ per-file question cannot see a guard in another file, so it is right seven
134
+ times in ten on code people wrote and less on generated corpora that label
135
+ single unguarded routes inside otherwise guarded files.
136
+
137
+ ```bash
138
+ npx supercov security file src/routes/invoices.ts
139
+ ```
140
+
141
+ shows every check with its value and its evidence line, so a finding can be
142
+ weighed without leaving the terminal.
143
+
144
+ ## Cross it with a run
145
+
146
+ ```bash
147
+ npx supercov security --run latest
148
+ ```
149
+
150
+ adds, for every flagged file, whether the saved coverage run executed it and how
151
+ many of its measured lines no test reached. A handler that builds a query from
152
+ request input is one thing; the same handler that no test ever executes is a
153
+ stronger claim, and one that needs both halves of this product to make.
154
+
155
+ ## What a change introduced
156
+
157
+ ```bash
158
+ npx supercov security patch --base origin/main
159
+ npx supercov security patch --annotate github
160
+ ```
161
+
162
+ The same twelve, asked whether the new version shows a surface the old one did
163
+ not. `quality patch` asks them too, alongside its own catalog, so a review of a
164
+ change needs one command; `security patch` is the same answers without the
165
+ complexity findings.
166
+
167
+ ## Reference
168
+
169
+ ```bash
170
+ npx supercov security # this repository
171
+ npx supercov security gaps # only files something fired on
172
+ npx supercov security file src/server.ts # one file, every check, with evidence
173
+ npx supercov security scope # which files, and why
174
+ npx supercov security snapshots # saved assessments
175
+ npx supercov security diff <older> <newer> # what appeared
176
+ npx supercov security patch # what a change introduced
177
+ npx supercov security --run latest # flagged and untested
178
+ ```
179
+
180
+ Cache and snapshots live under `.supercov/security/`, apart from quality's, so
181
+ one can never be read as the other.
@@ -184,14 +184,16 @@ attribution.
184
184
 
185
185
  | Runner | Attribution | Current requirement |
186
186
  | --- | --- | --- |
187
- | pytest | Exact test, worker, retry, and setup/call/teardown phase identity | CPython 3.12 or newer; run with `npx supercov -- pytest` or `python -m pytest` |
187
+ | pytest | Exact test, worker, retry, and setup/call/teardown phase identity | CPython 3.9 or newer; run with `npx supercov -- pytest` or `python -m pytest` |
188
188
  | pytest-xdist | Exact per worker | Workers inherit the run through the environment |
189
189
  | pytest-rerunfailures | Exact per attempt; flaky tests are reported as such | |
190
190
  | `python -m unittest` | Exact test and setUp/test/tearDown phase identity | Serial in-process; skips and expected failures are recorded; subtest failures roll up to the parent test |
191
191
 
192
192
  Your project runs in place with its own interpreter and virtual environment.
193
- Supercov adds its monitoring and runner hooks through the process environment;
194
- you do not need to rewrite tests or configure a different build.
193
+ Supercov adds its runner hooks through the process environment and places a
194
+ probe for every obligation into each measured module as it is imported --
195
+ nothing on disk changes, and tracebacks keep their line numbers. You do not
196
+ need to rewrite tests or configure a different build.
195
197
 
196
198
  Coverage includes statements, functions, boolean decisions, loops,
197
199
  comprehensions, short-circuit operators, `match` cases and exception paths.
@@ -202,9 +204,11 @@ assertion checks.
202
204
 
203
205
  Interpreters launched with `-I`, `-E` or `-S` ignore the required startup hook
204
206
  and are not measured. Code compiled from strings at runtime has no source
205
- obligations. Completed observations can survive a hard kill, but a corrupt or
206
- exhausted evidence channel fails the run rather than reporting partial data as
207
- complete.
207
+ obligations. A measured module imported before Supercov starts, or compiled
208
+ past the import system, runs without probes; the run names it as a limitation
209
+ rather than reporting it uncovered. What a test executed survives the process
210
+ being killed -- `os._exit`, SIGTERM or SIGKILL -- but a corrupt or exhausted
211
+ evidence channel fails the run rather than reporting partial data as complete.
208
212
 
209
213
  ```sh
210
214
  npx supercov -- pytest
@@ -74,6 +74,47 @@ Compare the command with [Supported suites](supported-suites.md). If the runner
74
74
  should be supported, preserve the summary and exact command when reporting the
75
75
  problem.
76
76
 
77
+ ## A plain Python run behaves differently after using Supercov 1.2.0
78
+
79
+ Version 1.2.0 could write instrumented bytecode into ordinary Python caches
80
+ through explicit source loaders or `py_compile`. Its fallback for condition
81
+ probes could then change behavior when a later plain run loaded that bytecode.
82
+
83
+ After updating Supercov, remove the generated bytecode caches for the affected
84
+ project, or discard the dedicated `PYTHONPYCACHEPREFIX` directory used by the
85
+ old run. Rebuild those caches with plain Python before comparing test outcomes.
86
+ Updating Supercov cannot repair bytecode already written by an older version.
87
+
88
+ ## A Python file has tests but no measured lines
89
+
90
+ Python imports and `python -m package.module` pass through Supercov's probe
91
+ loader. CPython compiles a direct entry script (`python path/to/script.py`)
92
+ outside that loader. Supercov declares this boundary on the affected file,
93
+ including on Python 3.9–3.11. It does not silently rewrite the child command,
94
+ which could change imports, `sys.argv`, or process behavior.
95
+
96
+ A file can therefore have tests that touched it through an import, while a
97
+ separate direct execution makes its overall coverage unmeasured. Those lines
98
+ are excluded from both the measured denominator and the uncovered-lines list.
99
+ Inspect the file's measurement limitations before treating a zero denominator
100
+ as complete coverage.
101
+
102
+ ## Python subprocess output or timeout assertions fail
103
+
104
+ Supercov adds interpreter startup and execution time. A timeout that expires
105
+ before the child reaches its first write can return empty partial output.
106
+ When testing timeout handling, synchronize with the child's readiness first;
107
+ when testing an actual performance deadline, compare the plain run separately.
108
+
109
+ Ordinary helpers that import no application code stay silent. Set
110
+ `SUPERCOV_PYTHON_TIMING=1` only for diagnostics: it deliberately adds stderr
111
+ output. Explicit child environments that omit Supercov's plan are left alone.
112
+
113
+ Runtime evidence descriptors are close-on-exec. After an exec, Python or
114
+ Supercov may reuse a closed descriptor's number. A descriptor-leak test should
115
+ compare the original open file's identity, rather than only whether its number
116
+ is open in the new process.
117
+
77
118
  ## A Ruby file behaves differently while Supercov measures it
78
119
 
79
120
  Supercov splices probe calls into Ruby sources in memory as they load, and a
@@ -185,9 +226,9 @@ See [Speed and storage](performance.md) for practical ways to shorten a loop.
185
226
  Preview cleanup, then choose how much history to keep:
186
227
 
187
228
  ```sh supercov
188
- npx supercov clean --dry-run
189
- npx supercov clean --keep 20
190
- npx supercov clean
229
+ npx supercov runs clean --dry-run
230
+ npx supercov runs clean --keep 20
231
+ npx supercov runs clean
191
232
  ```
192
233
 
193
234
  The final command removes all stored runs and the isolated build cache. Cleanup
@@ -75,9 +75,9 @@ staging state. Completed runs remain immutable.
75
75
  Preview cleanup before removing anything:
76
76
 
77
77
  ```sh supercov
78
- npx supercov clean --dry-run
79
- npx supercov clean --keep 20
80
- npx supercov clean
78
+ npx supercov runs clean --dry-run
79
+ npx supercov runs clean --keep 20
80
+ npx supercov runs clean
81
81
  ```
82
82
 
83
83
  The final command removes all runs and the isolated build cache. `--keep 20`
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "supercov",
3
- "version": "1.1.1",
4
- "description": "Code quality and coverage for coding agents",
3
+ "version": "2.0.0",
4
+ "description": "Coverage, security and code quality for coding agents",
5
5
  "license": "MIT",
6
6
  "repository": {
7
7
  "type": "git",
@@ -19,6 +19,7 @@
19
19
  "docs/agent-loop.md",
20
20
  "docs/assertions.md",
21
21
  "docs/quality.md",
22
+ "docs/security.md",
22
23
  "docs/assertion-evidence.md",
23
24
  "docs/assertion-maps.md",
24
25
  "docs/assertion-agent.md",
@@ -63,10 +64,11 @@
63
64
  "test:isolation": "cargo build -p supercov && node scripts/isolation-integration.mjs",
64
65
  "test:filesystem": "cargo test --workspace && cargo build -p supercov && node scripts/workspace-crash-integration.mjs",
65
66
  "test:watchdog": "cargo build -p supercov && node scripts/watchdog-integration.mjs",
67
+ "test:html-report": "cargo build -p supercov && node scripts/html-report-integration.mjs",
66
68
  "test:engine-contract": "cargo build -p supercov && node scripts/engine-contract.mjs",
67
69
  "test:agent": "cargo build -p supercov && node scripts/agent-query-eval.mjs",
68
70
  "test:engine": "cargo fmt --all -- --check && cargo clippy --workspace --all-targets -- -D warnings && cargo test --workspace && npm run test:runtime && npm run test:rust-assets && npm run test:engine-integrations",
69
- "test:engine-integrations": "cargo build -p supercov && node scripts/rust-process-supervision.mjs && node scripts/rust-direct-node-integration.mjs && node scripts/rust-public-run-integration.mjs && node scripts/rust-embedded-runtime-integration.mjs && node scripts/rust-direct-vitest-integration.mjs && node scripts/rust-direct-jest-integration.mjs && node scripts/rust-direct-playwright-integration.mjs && node scripts/rust-custom-browser-playwright-integration.mjs && node scripts/rust-generic-esbuild-integration.mjs && node scripts/rust-generic-tsc-integration.mjs && node scripts/rust-generic-build-matrix.mjs && node scripts/rust-vite-playwright-integration.mjs",
71
+ "test:engine-integrations": "cargo build -p supercov && node scripts/html-report-integration.mjs && node scripts/rust-process-supervision.mjs && node scripts/rust-direct-node-integration.mjs && node scripts/rust-public-run-integration.mjs && node scripts/rust-embedded-runtime-integration.mjs && node scripts/rust-direct-vitest-integration.mjs && node scripts/rust-direct-jest-integration.mjs && node scripts/rust-direct-playwright-integration.mjs && node scripts/rust-custom-browser-playwright-integration.mjs && node scripts/rust-generic-esbuild-integration.mjs && node scripts/rust-generic-tsc-integration.mjs && node scripts/rust-generic-build-matrix.mjs && node scripts/rust-vite-playwright-integration.mjs",
70
72
  "test:platform": "cargo test --workspace && cargo build -p supercov && node scripts/rust-process-supervision.mjs && node scripts/workspace-crash-integration.mjs",
71
73
  "test:native-package": "cargo build --release -p supercov && node scripts/native-package-integration.mjs && node scripts/native-release-set-integration.mjs",
72
74
  "test:pypi-wheel": "node scripts/pypi-wheel-integration.mjs",
@@ -78,9 +80,9 @@
78
80
  "test:clang-mcdc": "node scripts/clang-mcdc-oracle.mjs",
79
81
  "test:test262": "cargo build --release -p supercov && node scripts/test262-equivalence.mjs",
80
82
  "benchmark:check": "cargo build --release -p supercov && node scripts/rust-transform-benchmark.mjs",
81
- "benchmark:python-monitoring": "cargo build -p supercov && node scripts/python-monitoring-benchmark.mjs",
83
+ "benchmark:python": "cargo build -p supercov && node scripts/python-benchmark.mjs",
82
84
  "check": "cargo fmt --all -- --check && cargo clippy --workspace --all-targets -- -D warnings && npm run test && npm run test:release-tooling && npm run test:runtime && npm run test:rust-assets && npm run test:assertion-maps && npm run test:sweep-target && npm run test:skip-guard && npm run test:windows-tls && node scripts/package-preflight.mjs && node scripts/verify-binstall-metadata.mjs --offline",
83
- "release:check": "cargo clean && node scripts/sweep-target.mjs && npm run check && npm run test:engine && npm run test:go && npm run test:jvm && npm run test:launcher && npm run test:fixture && npm run test:watchdog && npm run test:engine-contract && npm run test:agent && npm run test:child-attribution && npm run test:host-loader && npm run test:python-monitoring && npm run test:ruby-coverage && npm run test:rust-public-cargo && npm run test:packed-npx && npm run test:clang-mcdc && npm run benchmark:check",
85
+ "release:check": "cargo clean && node scripts/sweep-target.mjs && npm run check && npm run test:engine && npm run test:go && npm run test:jvm && npm run test:launcher && npm run test:fixture && npm run test:watchdog && npm run test:engine-contract && npm run test:agent && npm run test:child-attribution && npm run test:host-loader && npm run test:python && npm run test:ruby-coverage && npm run test:rust-public-cargo && npm run test:packed-npx && npm run test:clang-mcdc && npm run benchmark:check",
84
86
  "release:bump": "node scripts/bump-version.mjs",
85
87
  "sweep": "node scripts/sweep-target.mjs",
86
88
  "oracle:rust": "cargo build --release -p supercov && node scripts/rust-coverage-oracle.mjs",
@@ -88,7 +90,7 @@
88
90
  "prepublishOnly": "npm run release:check",
89
91
  "test:child-attribution": "cargo build -p supercov && node scripts/rust-child-attribution-integration.mjs",
90
92
  "test:host-loader": "cargo build -p supercov && node scripts/rust-host-loader-integration.mjs",
91
- "test:python-monitoring": "cargo build -p supercov && node scripts/python-monitoring-integration.mjs",
93
+ "test:python": "cargo build -p supercov && node scripts/python-integration.mjs",
92
94
  "test:ruby-coverage": "cargo build -p supercov && node scripts/ruby-coverage-integration.mjs",
93
95
  "test:go": "cargo test --no-fail-fast -p supercov-engine --test go_frontend --test go_run",
94
96
  "test:jvm": "cargo test --no-fail-fast -p supercov-engine --test jvm_frontend --test jvm_run -- --include-ignored --test-threads=1",
@@ -105,14 +107,14 @@
105
107
  "test:windows-tls": "node scripts/windows-tls-test.mjs"
106
108
  },
107
109
  "optionalDependencies": {
108
- "@supercov/cli-darwin-arm64": "1.1.1",
109
- "@supercov/cli-darwin-x64": "1.1.1",
110
- "@supercov/cli-linux-arm64-gnu": "1.1.1",
111
- "@supercov/cli-linux-arm64-musl": "1.1.1",
112
- "@supercov/cli-linux-x64-gnu": "1.1.1",
113
- "@supercov/cli-linux-x64-musl": "1.1.1",
114
- "@supercov/cli-win32-arm64": "1.1.1",
115
- "@supercov/cli-win32-x64": "1.1.1"
110
+ "@supercov/cli-darwin-arm64": "2.0.0",
111
+ "@supercov/cli-darwin-x64": "2.0.0",
112
+ "@supercov/cli-linux-arm64-gnu": "2.0.0",
113
+ "@supercov/cli-linux-arm64-musl": "2.0.0",
114
+ "@supercov/cli-linux-x64-gnu": "2.0.0",
115
+ "@supercov/cli-linux-x64-musl": "2.0.0",
116
+ "@supercov/cli-win32-arm64": "2.0.0",
117
+ "@supercov/cli-win32-x64": "2.0.0"
116
118
  },
117
119
  "peerDependencies": {
118
120
  "@playwright/test": ">=1.55.0",