greencheck 0.3.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- greencheck-0.3.0/CITATION.cff +22 -0
- greencheck-0.3.0/LICENSE +21 -0
- greencheck-0.3.0/MANIFEST.in +20 -0
- greencheck-0.3.0/PKG-INFO +379 -0
- greencheck-0.3.0/README.md +348 -0
- greencheck-0.3.0/case_study/results/CASE_STUDY.md +95 -0
- greencheck-0.3.0/case_study/results/tables.json +203 -0
- greencheck-0.3.0/case_study/run_case_study.py +250 -0
- greencheck-0.3.0/case_study/third_party/README.md +128 -0
- greencheck-0.3.0/case_study/third_party/gate.py +37 -0
- greencheck-0.3.0/case_study/third_party/results/mutate-report.json +715 -0
- greencheck-0.3.0/case_study/third_party/run.py +100 -0
- greencheck-0.3.0/case_study/third_party/triage.py +92 -0
- greencheck-0.3.0/data/PROVENANCE.md +84 -0
- greencheck-0.3.0/data/gate_daily.jsonl +16 -0
- greencheck-0.3.0/data/gate_summary.json +37 -0
- greencheck-0.3.0/data/ledger_metric_samples.jsonl +84 -0
- greencheck-0.3.0/docs/METHOD.md +102 -0
- greencheck-0.3.0/greencheck/__init__.py +70 -0
- greencheck-0.3.0/greencheck/cli.py +262 -0
- greencheck-0.3.0/greencheck/core.py +671 -0
- greencheck-0.3.0/greencheck/mutate.py +430 -0
- greencheck-0.3.0/greencheck/report.py +75 -0
- greencheck-0.3.0/greencheck/skills/README.md +30 -0
- greencheck-0.3.0/greencheck/skills/bare-zero/SKILL.md +80 -0
- greencheck-0.3.0/greencheck/skills/dead-check/SKILL.md +74 -0
- greencheck-0.3.0/greencheck/skills/dimension-scope/SKILL.md +63 -0
- greencheck-0.3.0/greencheck/skills/measurement-or-decoration/SKILL.md +76 -0
- greencheck-0.3.0/greencheck/skills/positive-control/SKILL.md +64 -0
- greencheck-0.3.0/greencheck.egg-info/PKG-INFO +379 -0
- greencheck-0.3.0/greencheck.egg-info/SOURCES.txt +41 -0
- greencheck-0.3.0/greencheck.egg-info/dependency_links.txt +1 -0
- greencheck-0.3.0/greencheck.egg-info/entry_points.txt +2 -0
- greencheck-0.3.0/greencheck.egg-info/requires.txt +3 -0
- greencheck-0.3.0/greencheck.egg-info/top_level.txt +1 -0
- greencheck-0.3.0/pyproject.toml +61 -0
- greencheck-0.3.0/setup.cfg +4 -0
- greencheck-0.3.0/tests/test_greencheck.py +546 -0
- greencheck-0.3.0/tools/anonymise.py +413 -0
- greencheck-0.3.0/tools/field_renames.local.example.json +5 -0
- greencheck-0.3.0/tools/forbidden.local.example.txt +13 -0
- greencheck-0.3.0/tools/leakscan.py +112 -0
- greencheck-0.3.0/tools/smoke_install.py +151 -0
|
@@ -0,0 +1,22 @@
|
|
|
1
|
+
cff-version: 1.2.0
|
|
2
|
+
message: "If you use this software or the case-study data, please cite it as below."
|
|
3
|
+
title: "greencheck: auditing self-reported instruments in long-running agents"
|
|
4
|
+
abstract: >-
|
|
5
|
+
A dependency-free checker for one rule: a number is not a measurement until it
|
|
6
|
+
has been observed to take different values on inputs that are known to differ.
|
|
7
|
+
Ships a taxonomy of six ways a self-reported metric can be correct code and
|
|
8
|
+
still measure nothing, and a case study over six instruments of an autonomous
|
|
9
|
+
agent that ran continuously for roughly two months.
|
|
10
|
+
authors:
|
|
11
|
+
- name: Simon
|
|
12
|
+
version: 0.2.0
|
|
13
|
+
date-released: 2026-09-22
|
|
14
|
+
license: MIT
|
|
15
|
+
type: software
|
|
16
|
+
keywords:
|
|
17
|
+
- autonomous-agents
|
|
18
|
+
- self-report
|
|
19
|
+
- observability
|
|
20
|
+
- falsification
|
|
21
|
+
- metrics
|
|
22
|
+
- reproducibility
|
greencheck-0.3.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Simon
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
include README.md
|
|
2
|
+
include LICENSE
|
|
3
|
+
include CITATION.cff
|
|
4
|
+
include pyproject.toml
|
|
5
|
+
|
|
6
|
+
# The skills ship with the source tree rather than inside the package: they are
|
|
7
|
+
# meant to be read as files (and browsed on GitHub), not imported. `greencheck
|
|
8
|
+
# skills` locates this directory relative to the package, so a source install
|
|
9
|
+
# finds them; a bare `pip install greencheck` from a wheel would not.
|
|
10
|
+
recursive-include skills *.md
|
|
11
|
+
recursive-include data *.json *.jsonl *.md
|
|
12
|
+
recursive-include docs *.md
|
|
13
|
+
recursive-include case_study *.py *.md *.json
|
|
14
|
+
recursive-include tests *.py
|
|
15
|
+
recursive-include tools *.py *.example.json *.example.txt
|
|
16
|
+
|
|
17
|
+
# Never ship these, whatever else changes.
|
|
18
|
+
exclude tools/*.local.json
|
|
19
|
+
exclude tools/*.local.txt
|
|
20
|
+
prune .git
|
|
@@ -0,0 +1,379 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: greencheck
|
|
3
|
+
Version: 0.3.0
|
|
4
|
+
Summary: Your validator passes inputs it should reject. greencheck finds them.
|
|
5
|
+
Author: Simon
|
|
6
|
+
License: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/simin-yuan/greencheck
|
|
8
|
+
Project-URL: Repository, https://github.com/simin-yuan/greencheck
|
|
9
|
+
Project-URL: Issues, https://github.com/simin-yuan/greencheck/issues
|
|
10
|
+
Project-URL: Changelog, https://github.com/simin-yuan/greencheck/releases
|
|
11
|
+
Keywords: mutation-testing,observability,metrics,agent,monitoring,evaluation,verification,test,testing,audit
|
|
12
|
+
Classifier: Development Status :: 4 - Beta
|
|
13
|
+
Classifier: Intended Audience :: Developers
|
|
14
|
+
Classifier: Operating System :: OS Independent
|
|
15
|
+
Classifier: Programming Language :: Python :: 3
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.9
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
20
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
21
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
22
|
+
Classifier: Topic :: Software Development :: Quality Assurance
|
|
23
|
+
Classifier: Topic :: Software Development :: Testing
|
|
24
|
+
Classifier: Topic :: System :: Monitoring
|
|
25
|
+
Requires-Python: >=3.9
|
|
26
|
+
Description-Content-Type: text/markdown
|
|
27
|
+
License-File: LICENSE
|
|
28
|
+
Provides-Extra: dev
|
|
29
|
+
Requires-Dist: pytest>=7; extra == "dev"
|
|
30
|
+
Dynamic: license-file
|
|
31
|
+
|
|
32
|
+
# greencheck
|
|
33
|
+
|
|
34
|
+
[](https://github.com/simin-yuan/greencheck/actions/workflows/tests.yml)
|
|
35
|
+
[](LICENSE)
|
|
36
|
+
[](#)
|
|
37
|
+
|
|
38
|
+
### Your validator passes inputs it should reject. This finds them.
|
|
39
|
+
|
|
40
|
+
```console
|
|
41
|
+
$ pip install https://github.com/simin-yuan/greencheck/releases/download/v0.3.0/greencheck-0.3.0-py3-none-any.whl
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Then, from a checkout of this repository (both halves run against files that
|
|
45
|
+
ship with it, so a reader can reproduce every line below):
|
|
46
|
+
|
|
47
|
+
```console
|
|
48
|
+
$ python -m greencheck.cli mutate --gate "python examples/mutate-demo/gate.py {target}/config.json" --target examples/mutate-demo/input
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
```
|
|
52
|
+
baseline rc: 0 PASS (baseline is clean, mutating)
|
|
53
|
+
mutants : 12
|
|
54
|
+
----------------------------------------------------------------------
|
|
55
|
+
caught drop-file:config.json
|
|
56
|
+
caught empty-file:config.json
|
|
57
|
+
caught drop-line:config.json:4:"replicas": 3,
|
|
58
|
+
caught blank-value:config.json:2:"service":
|
|
59
|
+
caught blank-value:config.json:3:"region":
|
|
60
|
+
* ESCAPED blank-value:config.json:4:"replicas":
|
|
61
|
+
caught blank-value:config.json:5:"owner":
|
|
62
|
+
----------------------------------------------------------------------
|
|
63
|
+
caught 11 / 12
|
|
64
|
+
|
|
65
|
+
mutants the gate let through (1):
|
|
66
|
+
* blank-value:config.json:4:"replicas":
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
The escaped mutant turns `"replicas": 3` into `"replicas": 0` — a service that
|
|
70
|
+
never starts. The gate checks that every key is present and every value is
|
|
71
|
+
non-empty. Zero is neither missing nor empty, so it says yes.
|
|
72
|
+
|
|
73
|
+
**Every structural property the gate tests for still holds.** That is exactly
|
|
74
|
+
why this class of gap survives review: it is invisible to anything that only
|
|
75
|
+
looks at shape. → [the runnable example](examples/mutate-demo/)
|
|
76
|
+
|
|
77
|
+
## It also audits metrics
|
|
78
|
+
|
|
79
|
+
Point it at a metric you already collect — a self-score, a recall rate, a health
|
|
80
|
+
probe — and it tells you whether the number varies with the thing it is named
|
|
81
|
+
after, or only with something adjacent: input existence, wall-clock freshness, a
|
|
82
|
+
single boolean.
|
|
83
|
+
|
|
84
|
+
```console
|
|
85
|
+
$ python -m greencheck.cli audit data/ledger_metric_samples.jsonl --field identity.identity_score
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
The same question, asked of numbers instead of gates: *has this ever been
|
|
89
|
+
observed to take a different value on input that differs?* If it has not, it is
|
|
90
|
+
not a measurement. It is a decoration that updates on a schedule.
|
|
91
|
+
|
|
92
|
+
---
|
|
93
|
+
|
|
94
|
+
## The one rule
|
|
95
|
+
|
|
96
|
+
> **A number is not a measurement until it has been observed to take different
|
|
97
|
+
> values on inputs that are known to differ.**
|
|
98
|
+
|
|
99
|
+
A metric that reads `1.0` forever is not telling you things are perfect. It is
|
|
100
|
+
not telling you anything. A dashboard that is green because it has never been
|
|
101
|
+
asked a question it could fail is a specific kind of silence — and this tool
|
|
102
|
+
exists to make that silence audible.
|
|
103
|
+
|
|
104
|
+
Point it at a metric you already collect: a self-score, a recall rate, a guard,
|
|
105
|
+
a health probe. It tells you whether the number varies with the thing it is
|
|
106
|
+
named after, or only with something adjacent — input existence, wall-clock
|
|
107
|
+
freshness, a single boolean.
|
|
108
|
+
|
|
109
|
+
---
|
|
110
|
+
|
|
111
|
+
## 30 seconds
|
|
112
|
+
|
|
113
|
+
```console
|
|
114
|
+
$ python -m greencheck.cli demo
|
|
115
|
+
|
|
116
|
+
[ok ] content_length (2 samples)
|
|
117
|
+
[FAIL] content_presence_score (2 samples)
|
|
118
|
+
- IDENTITY: 0/1 probe pairs separated within the claimed dimension
|
|
119
|
+
'content'. The instrument responds to something other than what it
|
|
120
|
+
claims to measure.
|
|
121
|
+
[FAIL] always_one (control) (2 samples)
|
|
122
|
+
```
|
|
123
|
+
|
|
124
|
+
The middle line is the point. `content_presence_score` **does** separate one of
|
|
125
|
+
the probe pairs — `1.0` for text, `0.0` for empty. A naive check ("did any two
|
|
126
|
+
outputs differ?") passes it. But the pair it separates differs in *existence*,
|
|
127
|
+
not in *content*, and `content` is what the instrument is named after. A
|
|
128
|
+
discriminability test that is not dimension-scoped is itself a decoration.
|
|
129
|
+
|
|
130
|
+
Audit a real ledger (this one ships in `data/`):
|
|
131
|
+
|
|
132
|
+
```console
|
|
133
|
+
$ python -m greencheck.cli audit data/ledger_metric_samples.jsonl --field identity.identity_score
|
|
134
|
+
|
|
135
|
+
[FAIL] identity.identity_score (84 samples)
|
|
136
|
+
- CONSTANT: 84/84 samples report the identical value 1.0.
|
|
137
|
+
Zero variance over the observed window.
|
|
138
|
+
stats: n_samples=84, distinct_values=1
|
|
139
|
+
```
|
|
140
|
+
|
|
141
|
+
Check whether an assertion has ever fired. This reads the per-day aggregate
|
|
142
|
+
log in `data/` — the same file the case study table is computed from, so the
|
|
143
|
+
figure is yours to re-add:
|
|
144
|
+
|
|
145
|
+
```console
|
|
146
|
+
$ python -m greencheck.cli gate data/gate_daily.jsonl --fire-event block_issued
|
|
147
|
+
|
|
148
|
+
[FAIL] data/gate_daily.jsonl — 13 full days with zero firings (27097 samples)
|
|
149
|
+
- DEAD_GATE: 27,097 events recorded across 13 days, of which 25,594 were
|
|
150
|
+
the guard reporting itself healthy, and the denial path fired 0 times.
|
|
151
|
+
The assertion has never been observed to fire. Its existence is
|
|
152
|
+
documented; its behaviour is not.
|
|
153
|
+
stats: total_events=27097, health_events=25594, firings=0, silent_days_before_first_fire=13
|
|
154
|
+
|
|
155
|
+
[ok ] data/gate_daily.jsonl — full window (16 days) (33228 samples)
|
|
156
|
+
- PASS: The path fired 5 times, the first on 2026-09-19. That first
|
|
157
|
+
firing came from a positive control written specifically to exercise
|
|
158
|
+
the path, not from ordinary traffic.
|
|
159
|
+
stats: total_events=33228, health_events=31642, firings=5
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
Both rows are the same guard. The code did not change between them; what
|
|
163
|
+
changed is that somebody finally built it an input it was known to have to
|
|
164
|
+
reject. Full window: 5 firings out of 33,228 events. Without that positive
|
|
165
|
+
control, all 33,228 would have read as health.
|
|
166
|
+
|
|
167
|
+
Zero dependencies. Python 3.9+. Nothing to install beyond the repo.
|
|
168
|
+
|
|
169
|
+
---
|
|
170
|
+
|
|
171
|
+
## What it catches
|
|
172
|
+
|
|
173
|
+
| verdict | signature | how you notice |
|
|
174
|
+
|---|---|---|
|
|
175
|
+
| `CONSTANT` | Zero variance over the whole window | 84/84 samples report the same value |
|
|
176
|
+
| `IDENTITY` | Separates inputs by existence, not by the claimed dimension | Empty vs non-empty moves it; good vs contradictory content does not |
|
|
177
|
+
| `BARE_ZERO` | `n = 0` and `n > 0, 0 hits` produce the same output | 63 of 84 "recall" samples examined nothing and reported `0.0` |
|
|
178
|
+
| `LIVENESS_BIT` | All variance traces to one boolean | A composite score whose two values are exactly a freshness flag |
|
|
179
|
+
| `DEAD_GATE` | A guard exists, logs health, has never fired | 25,594 healthy events, 0 firings, 13 full days |
|
|
180
|
+
| `NO_DATA` | Nothing matched | Reported instead of `PASS`, on purpose |
|
|
181
|
+
|
|
182
|
+
**`NO_DATA` is the failure mode of the audit tool itself.** An empty input
|
|
183
|
+
returning `PASS` is precisely the class of bug this repository exists to catch,
|
|
184
|
+
so the tool refuses to commit it. And `IDENTITY` is not "the values differ
|
|
185
|
+
somewhere" — discrimination is scoped to the dimension the instrument *claims*
|
|
186
|
+
to measure, which is a strictly harder bar and the one that actually matters.
|
|
187
|
+
|
|
188
|
+
None of these is hypothetical. Each was found on a real instrument, in
|
|
189
|
+
production, in the case study below.
|
|
190
|
+
|
|
191
|
+
---
|
|
192
|
+
|
|
193
|
+
## The case study
|
|
194
|
+
|
|
195
|
+
`data/` holds two anonymised datasets from a continuously-running autonomous
|
|
196
|
+
agent over roughly two months. Everything in this repo is reproducible from
|
|
197
|
+
them:
|
|
198
|
+
|
|
199
|
+
```bash
|
|
200
|
+
python case_study/run_case_study.py
|
|
201
|
+
```
|
|
202
|
+
|
|
203
|
+
| instrument | n | verdict |
|
|
204
|
+
|---|---:|---|
|
|
205
|
+
| `self_assessment.identity_score` | 84 | **CONSTANT** |
|
|
206
|
+
| `self_assessment.memory_recall.recall_rate` | 84 | **CONSTANT; BARE_ZERO** |
|
|
207
|
+
| `self_assessment.reflection_score` | 84 | **LIVENESS_BIT** |
|
|
208
|
+
| `self_assessment.overall (composite)` | 84 | **LIVENESS_BIT** |
|
|
209
|
+
| `guard denial path — 13 full days with zero firings` | 27,097 | **DEAD_GATE** |
|
|
210
|
+
| `guard denial path — full window` | 33,228 | **PASS** |
|
|
211
|
+
|
|
212
|
+
Six instruments. Five flagged. Full write-up:
|
|
213
|
+
[`case_study/results/CASE_STUDY.md`](case_study/results/CASE_STUDY.md).
|
|
214
|
+
|
|
215
|
+
The last two rows are **the same guard**. It reads `DEAD_GATE` for the thirteen
|
|
216
|
+
days before anyone constructed an input it was *known* to have to block, and
|
|
217
|
+
`PASS` afterwards. Nothing about the guard changed at that moment. What changed
|
|
218
|
+
is that somebody finally asked it to prove itself.
|
|
219
|
+
|
|
220
|
+
Every number above can be recomputed from the two files in `data/`. The headline
|
|
221
|
+
is 27,097 and not a larger figure, because the larger one counted events with a
|
|
222
|
+
timestamp earlier than the first firing — which needs event-level data, and only
|
|
223
|
+
aggregates are published. A reader could not have checked it. Summing the
|
|
224
|
+
zero-firing rows of [`data/gate_daily.jsonl`](data/gate_daily.jsonl) gives
|
|
225
|
+
exactly 27,097, and that is the number used everywhere in this repository.
|
|
226
|
+
|
|
227
|
+
That is the whole thesis. A control is not validated by its presence in the
|
|
228
|
+
code, nor by its green light. It is validated by having been observed to fire.
|
|
229
|
+
|
|
230
|
+
### The same loop, on a validator nobody here wrote
|
|
231
|
+
|
|
232
|
+
[`case_study/third_party/`](case_study/third_party/) points it at third-party
|
|
233
|
+
artifacts instead: a real `package.json` from a public project, checked by
|
|
234
|
+
SchemaStore's official schema. 168 mutants, 36 caught, 132 let through — and
|
|
235
|
+
then the part that decides anything: `npm` refuses two of the escapes and
|
|
236
|
+
accepts a third. The count is not the finding. The triage is.
|
|
237
|
+
|
|
238
|
+
---
|
|
239
|
+
|
|
240
|
+
## Skills
|
|
241
|
+
|
|
242
|
+
The CLI audits a ledger **after** the fact. `greencheck/skills/` is for **before** —
|
|
243
|
+
agent-readable instructions that keep the broken instrument from being built at
|
|
244
|
+
all.
|
|
245
|
+
|
|
246
|
+
```console
|
|
247
|
+
$ python -m greencheck.cli skills
|
|
248
|
+
greencheck skills — 5 available
|
|
249
|
+
|
|
250
|
+
bare-zero Use when a metric, counter, rate or score reports zero...
|
|
251
|
+
dead-check Use when writing or reviewing a monitoring rule...
|
|
252
|
+
dimension-scope Use when testing whether an instrument measures what its name claims...
|
|
253
|
+
measurement-or-decoration Use when a system reports a score, confidence, health value about itself...
|
|
254
|
+
positive-control Use when writing a guard, assertion, test, alert or validation rule...
|
|
255
|
+
|
|
256
|
+
Read one in full: python -m greencheck.cli skills <name>
|
|
257
|
+
```
|
|
258
|
+
|
|
259
|
+
Plain `SKILL.md` files with YAML frontmatter — the format used by Claude Code,
|
|
260
|
+
Codex, OpenClaw and most agent harnesses. Copy one where your harness looks for
|
|
261
|
+
skills, or let your agent read it directly:
|
|
262
|
+
|
|
263
|
+
```console
|
|
264
|
+
$ python -m greencheck.cli skills positive-control # prints the whole file
|
|
265
|
+
```
|
|
266
|
+
|
|
267
|
+
| skill | use when |
|
|
268
|
+
|---|---|
|
|
269
|
+
| [`measurement-or-decoration`](greencheck/skills/measurement-or-decoration/SKILL.md) | the system reports a score about itself. The root rule; the others are special cases. |
|
|
270
|
+
| [`positive-control`](greencheck/skills/positive-control/SKILL.md) | writing a guard, assertion or alert that is supposed to reject bad input |
|
|
271
|
+
| [`dead-check`](greencheck/skills/dead-check/SKILL.md) | writing a monitor. Covers rules that can never fire, and rules with inverted polarity. |
|
|
272
|
+
| [`dimension-scope`](greencheck/skills/dimension-scope/SKILL.md) | testing whether an instrument measures the dimension its name claims |
|
|
273
|
+
| [`bare-zero`](greencheck/skills/bare-zero/SKILL.md) | a metric reports `0` |
|
|
274
|
+
|
|
275
|
+
## Install / use
|
|
276
|
+
|
|
277
|
+
No dependencies. Python 3.9+.
|
|
278
|
+
|
|
279
|
+
From a checkout of this repository:
|
|
280
|
+
|
|
281
|
+
```console
|
|
282
|
+
$ python -m unittest discover -s tests # 45 tests, stdlib only
|
|
283
|
+
$ python -m greencheck.cli demo
|
|
284
|
+
```
|
|
285
|
+
|
|
286
|
+
Audit your own ledger:
|
|
287
|
+
|
|
288
|
+
```bash
|
|
289
|
+
python -m greencheck.cli audit my_ledger.jsonl \
|
|
290
|
+
--field identity.identity_score \
|
|
291
|
+
--count memory_recall.sampled \
|
|
292
|
+
--fresh reflection.fresh
|
|
293
|
+
```
|
|
294
|
+
|
|
295
|
+
Or drive it as a library:
|
|
296
|
+
|
|
297
|
+
```python
|
|
298
|
+
from greencheck import ProbePair, discriminate
|
|
299
|
+
|
|
300
|
+
discriminate(my_metric, [
|
|
301
|
+
ProbePair("rich vs degenerate", rich_text, "aaaa", dimension="content"),
|
|
302
|
+
ProbePair("present vs absent", "hello", "", dimension="content"),
|
|
303
|
+
], subject="my_metric", claimed_dimensions=["content"])
|
|
304
|
+
```
|
|
305
|
+
|
|
306
|
+
---
|
|
307
|
+
|
|
308
|
+
## Where the evidence comes from
|
|
309
|
+
|
|
310
|
+
The underlying logs are from production and contain filesystem paths, host
|
|
311
|
+
identifiers and operator-specific strings. **They are not published.** What is
|
|
312
|
+
published is aggregated:
|
|
313
|
+
|
|
314
|
+
- `data/ledger_metric_samples.jsonl` — 84 samples, field names neutralised; no
|
|
315
|
+
value, timestamp or count altered
|
|
316
|
+
- `data/gate_summary.json`, `data/gate_daily.jsonl` — per-day counts only
|
|
317
|
+
|
|
318
|
+
`tools/anonymise.py` produces them and **fails the run** if any forbidden pattern
|
|
319
|
+
survives. The leak gate has its own positive control (`--self-test`), because a
|
|
320
|
+
checker that has never been observed to fire is not a checker — and that rule
|
|
321
|
+
applies to this repository before it applies to yours. `tools/leakscan.py`
|
|
322
|
+
re-verifies the entire tree independently. Rules and provenance:
|
|
323
|
+
[`data/PROVENANCE.md`](data/PROVENANCE.md).
|
|
324
|
+
|
|
325
|
+
---
|
|
326
|
+
|
|
327
|
+
## Standing on
|
|
328
|
+
|
|
329
|
+
This did not start from nothing, and it would be dishonest to imply otherwise:
|
|
330
|
+
|
|
331
|
+
- **Dimension-scoped discriminability and falsification-first verification** —
|
|
332
|
+
adapted from [obra/superpowers](https://github.com/obra/superpowers) (MIT).
|
|
333
|
+
- **`n = 1` case-study discipline** — report a single system honestly instead of
|
|
334
|
+
inflating it into a general claim.
|
|
335
|
+
|
|
336
|
+
The strongest contribution in the other direction: if you have a metric that
|
|
337
|
+
passes `greencheck` and still lies to you, that is a bug here, and the issue is
|
|
338
|
+
welcome.
|
|
339
|
+
|
|
340
|
+
---
|
|
341
|
+
|
|
342
|
+
## Limits
|
|
343
|
+
|
|
344
|
+
- **This tests instruments, not systems.** A `PASS` means the instrument
|
|
345
|
+
separates inputs within its claimed dimension. It does not mean the thing
|
|
346
|
+
measured is good, or that the dimension is the right one to measure.
|
|
347
|
+
- **`n = 1` system.** One agent, roughly two months. The signatures recur across
|
|
348
|
+
its instruments, but cross-system replication is open. Treat the taxonomy as a
|
|
349
|
+
starting set, not a closed one.
|
|
350
|
+
- **The tool only sees what you record.** A metric that is never written down
|
|
351
|
+
cannot be audited. Absence of records is itself a finding — a different one.
|
|
352
|
+
- **A positive control is part of the procedure, not an optional extra.** A
|
|
353
|
+
guard's firing count is meaningless until someone has shown it fires when it
|
|
354
|
+
should. See [`docs/METHOD.md`](docs/METHOD.md).
|
|
355
|
+
- **The instruments were written by the system being measured.** That is the
|
|
356
|
+
condition under study, not a flaw in the study.
|
|
357
|
+
- **The judging standard is not set by the thing being judged.** The verdicts in
|
|
358
|
+
the case study are proposed by this tool's author and should be read as a
|
|
359
|
+
draft. Disagreeing with one of them is the useful move.
|
|
360
|
+
|
|
361
|
+
## Falsifiers
|
|
362
|
+
|
|
363
|
+
1. If any instrument in the case study is shown to separate inputs within the
|
|
364
|
+
dimension it claims to measure, the corresponding finding is false.
|
|
365
|
+
2. If the guard is shown to have fired before `2026-09-19T14:52:23Z` on an input
|
|
366
|
+
that should have been blocked, the `DEAD_GATE` finding is false.
|
|
367
|
+
3. If anonymisation changed any *value* (as opposed to a field name), every
|
|
368
|
+
number in the case study is void. `tools/anonymise.py` is the only transform
|
|
369
|
+
applied; its leak check is mandatory.
|
|
370
|
+
|
|
371
|
+
## Citation
|
|
372
|
+
|
|
373
|
+
See [`CITATION.cff`](CITATION.cff).
|
|
374
|
+
|
|
375
|
+
## License
|
|
376
|
+
|
|
377
|
+
MIT — see [`LICENSE`](LICENSE).
|
|
378
|
+
|
|
379
|
+
**Simon**, 2026.
|