token-harness 0.1.11 → 0.1.12
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +34 -55
- package/package.json +1 -1
- package/sbom.json +3 -3
- package/token-harness.mjs +3719 -523
package/README.md
CHANGED
|
@@ -3,8 +3,8 @@
|
|
|
3
3
|
**Build, verify and measure an optimization stack for Claude Code and Codex.**
|
|
4
4
|
|
|
5
5
|
Token Harness is a local **optimization stack manager**. It checks your coding agents, manages the
|
|
6
|
-
optimization components it can safely own, keeps
|
|
7
|
-
result, and reports savings only when it has evidence to support them.
|
|
6
|
+
optimization components it can safely own, keeps evaluation evidence separate from managed lifecycle,
|
|
7
|
+
verifies the result, and reports savings only when it has evidence to support them.
|
|
8
8
|
|
|
9
9
|
It is not another coding agent and it does not replace specialized projects such as RTK or
|
|
10
10
|
HarnessTrim.
|
|
@@ -34,10 +34,9 @@ of CLI commands to memorize.
|
|
|
34
34
|
3. In **Setup**, work from top to bottom:
|
|
35
35
|
- Coding agents
|
|
36
36
|
- Managed optimizers
|
|
37
|
-
- Experimental tools
|
|
38
37
|
- Optional agent tuning
|
|
39
38
|
- Checks and maintenance
|
|
40
|
-
4. For Claude Code or Codex, choose **Review
|
|
39
|
+
4. For Claude Code or Codex, choose **Review baseline** when the RTK + HarnessTrim setup is available.
|
|
41
40
|
5. Read the exact proposed changes. **Apply reviewed setup** appears only when there is a concrete
|
|
42
41
|
safe plan to apply.
|
|
43
42
|
6. Keep using Claude Code or Codex normally.
|
|
@@ -72,8 +71,10 @@ useful read-only allowance and connected-tool observations when available.
|
|
|
72
71
|
|
|
73
72
|
**2. Managed optimizers**
|
|
74
73
|
|
|
75
|
-
The
|
|
76
|
-
|
|
74
|
+
The production baseline remains **RTK + HarnessTrim** on individually reviewed combinations.
|
|
75
|
+
Token Harness also exposes **mcptoon**, **GitNexus** and **Headroom** as optional managed integrations on their exact
|
|
76
|
+
reviewed lifecycle rows; enabling any of them does not make it part of the production baseline or create a
|
|
77
|
+
savings claim. Token Harness tracks the exact combined provider set separately: if no combined-stack
|
|
77
78
|
review is recorded, Setup says so and keeps the stack incomplete rather than inferring compatibility
|
|
78
79
|
from healthy individual checks. Token Harness can prepare their integration transactionally, show the
|
|
79
80
|
exact plan, apply it only after approval, verify it, and remove only configuration it owns.
|
|
@@ -91,45 +92,17 @@ shipped combined-review registry stays empty until the captured configuration ha
|
|
|
91
92
|
verification evidence, has been benchmarked together, and has been deliberately reviewed; see
|
|
92
93
|
`docs/combined-stack-reviews.md`.
|
|
93
94
|
|
|
94
|
-
**3.
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
repositories or register MCP servers.
|
|
101
|
-
|
|
102
|
-
Choose **Run standard evaluation** to start or resume a paired candidate campaign. Campaign state is
|
|
103
|
-
scoped to both the candidate and the selected harness, so Claude Code and Codex evidence cannot be
|
|
104
|
-
mixed accidentally. The browser reads the campaign directly and shows **Progress**, the current
|
|
105
|
-
selection signal, whether the evidence is **Decision ready**, the number of evidence-bearing pairs
|
|
106
|
-
and the exact **Next** step.
|
|
107
|
-
|
|
108
|
-
For normal use, the browser can now start and finish the local benchmark capture itself. You still
|
|
109
|
-
run the actual task in Claude Code or Codex. When the task finishes, record the quality result,
|
|
110
|
-
attempt count and failed-attempt count you actually observed. Before an optimized capture, Token
|
|
111
|
-
Harness requires you to acknowledge that you enabled the candidate through its own documented
|
|
112
|
-
workflow. That acknowledgement is **not activation verification** and is never treated as promotion
|
|
113
|
-
evidence. Every browser capture action is matched against the campaign engine's current step before
|
|
114
|
-
it can write local benchmark state, so a stale tab cannot advance a different step.
|
|
115
|
-
|
|
116
|
-
`benchmark-matrix`, `benchmark-start` and `benchmark-finish` remain available as advanced terminal
|
|
117
|
-
fallbacks for debugging or automation. The browser does not run the coding task, install or activate
|
|
118
|
-
a candidate, or claim that candidate attribution proves activation.
|
|
119
|
-
|
|
120
|
-
A campaign selection assessment can report `insufficient-evidence`, `promising`, `mixed` or
|
|
121
|
-
`negative`, plus whether the evidence is decision-ready. **Decision-ready is not promotion-ready.**
|
|
122
|
-
The campaign shows the same promotion-review gate count and next gate as the candidate card, using
|
|
123
|
-
the current local observation rather than starting another environment scan. Activation
|
|
124
|
-
verification, managed lifecycle, compatibility/reversibility, project maturity and combined-stack
|
|
125
|
-
validation remain separate gates.
|
|
95
|
+
**3. Evaluation evidence (maintainers)**
|
|
96
|
+
|
|
97
|
+
Selection campaigns and historical candidate evidence remain available for maintainers, but they are
|
|
98
|
+
not a separate novice setup product. Once a tool has a reviewed managed lifecycle, Setup presents it in
|
|
99
|
+
the same **Optimization Stack** with its exact state, version, prerequisites, harness scope and safe
|
|
100
|
+
Apply/Remove controls. Evaluation evidence never turns into a savings or compatibility claim by itself.
|
|
126
101
|
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
Activation is shown separately as **Verified**, **Blocked** or **Unreviewed**, with
|
|
132
|
-
verified/blocked/unknown pair counts when campaign evidence exists.
|
|
102
|
+
Headroom is config-only managed on its exact reviewed 0.37.0 row. Token Harness can own the narrow
|
|
103
|
+
Claude Code or Codex MCP registration after the pinned CLI is present; it deliberately does not
|
|
104
|
+
bootstrap `uv`/Python, start wrappers/proxies, or claim runtime compression savings. GitNexus likewise
|
|
105
|
+
does not auto-index repositories, and its noncommercial license boundary stays visible.
|
|
133
106
|
|
|
134
107
|
**4. Optional agent tuning**
|
|
135
108
|
|
|
@@ -209,6 +182,9 @@ algorithms into this repository.
|
|
|
209
182
|
| --- | --- | --- |
|
|
210
183
|
| [RTK](https://github.com/rtk-ai/rtk) | Shell/tool output reduction | Managed on reviewed combinations |
|
|
211
184
|
| [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic output/context reduction | Managed first-party integration |
|
|
185
|
+
| mcptoon | MCP discovery / compact manifest guidance | Optional managed integration on exact reviewed 0.7.10 rows; no savings assumed |
|
|
186
|
+
| GitNexus | Repository graph / MCP context | Optional managed Claude integration for already-installed 1.6.12; license review required |
|
|
187
|
+
| Headroom | Local MCP context compression/retrieval | Optional config-only managed Claude/Codex integration for already-installed 0.37.0; package prerequisite stays user-owned |
|
|
212
188
|
| [cclimits](https://github.com/cruzanstx/cclimits) | Optional Claude allowance evidence | Read-only evidence; not an optimizer |
|
|
213
189
|
| [ccusage](https://github.com/ccusage/ccusage) | Local usage history | Read-only evidence; never subscription quota |
|
|
214
190
|
|
|
@@ -233,8 +209,8 @@ RTK has no equivalent machine-readable capability endpoint, so releases newer th
|
|
|
233
209
|
reviewed RTK set remain visible as `unknown-newer` until their consumed contract is checked. See
|
|
234
210
|
[docs/provider-version-compatibility.md](docs/provider-version-compatibility.md).
|
|
235
211
|
|
|
236
|
-
|
|
237
|
-
benchmark is not enough for promotion. Their campaign assessment is structured evidence for the
|
|
212
|
+
Historical evaluation evidence remains available for mcptoon, GitNexus and Headroom. Detection or a promising
|
|
213
|
+
benchmark is not enough for a production-stack promotion or savings claim. Their campaign assessment is structured evidence for the
|
|
238
214
|
selection gate, not an activation or promotion decision. A candidate must pass structured promotion
|
|
239
215
|
readiness across benchmark capability, category fit, selection evidence, real activation
|
|
240
216
|
verification, managed lifecycle, compatibility/reversibility, project maturity and combined-stack
|
|
@@ -307,31 +283,34 @@ The older automation contracts remain available. `ui --json` preserves its exist
|
|
|
307
283
|
report; `ui --read-only` opens the legacy read-only UI; `ui --no-open` starts the guided app without
|
|
308
284
|
launching a browser.
|
|
309
285
|
|
|
310
|
-
###
|
|
286
|
+
### Evaluation evidence (advanced / maintainers)
|
|
311
287
|
|
|
312
|
-
|
|
288
|
+
Evaluation campaigns are an advanced maintainer workflow; managed setup stays in the unified **Optimization Stack**. The app
|
|
313
289
|
keeps a resumable campaign ID for each candidate/harness pair and reads campaign progress, assessment
|
|
314
290
|
and the exact **Next** step directly in the browser. Use **Start baseline capture** or **Start optimized
|
|
315
291
|
capture**, run the requested task in the selected coding agent, then choose **Record outcome** and
|
|
316
292
|
enter the quality/attempt values you actually observed. Normal use no longer requires copying
|
|
317
293
|
`benchmark-start` or `benchmark-finish` commands into a terminal.
|
|
318
294
|
|
|
319
|
-
The equivalent advanced CLI flow starts by asking the campaign engine for its current state
|
|
295
|
+
The equivalent advanced CLI flow starts by asking the campaign engine for its current state. This
|
|
296
|
+
GitNexus example intentionally uses the only currently reviewed campaign row:
|
|
320
297
|
|
|
321
298
|
```sh
|
|
322
299
|
token-harness benchmark-matrix \
|
|
323
|
-
--benchmark-id gitnexus-
|
|
300
|
+
--benchmark-id gitnexus-claude-eval-1 \
|
|
324
301
|
--candidate gitnexus \
|
|
325
|
-
--harness
|
|
302
|
+
--harness claude
|
|
326
303
|
```
|
|
327
304
|
|
|
328
305
|
Follow only the **Next** command printed by that report, complete the task honestly, then rerun the
|
|
329
306
|
same `benchmark-matrix` command. Before an optimized run, enable the candidate through its own
|
|
330
307
|
documented workflow. Token Harness records the experiment target but does not treat attribution—or
|
|
331
|
-
the browser acknowledgement—as proof that the candidate was active. For GitNexus
|
|
332
|
-
|
|
333
|
-
server
|
|
334
|
-
|
|
308
|
+
the browser acknowledgement—as proof that the candidate was active. For GitNexus on the reviewed
|
|
309
|
+
Claude Code `2.1.269` × GitNexus `1.6.12` × native-Linux row, the harness-native MCP inventory can
|
|
310
|
+
prove only that the GitNexus server was available at both task boundaries. That is not proof Claude
|
|
311
|
+
actually called a GitNexus tool, so the activation-verification promotion gate remains blocked
|
|
312
|
+
without a separate reviewed usage witness. See
|
|
313
|
+
[`docs/candidates/gitnexus-real-campaign.md`](docs/candidates/gitnexus-real-campaign.md).
|
|
335
314
|
|
|
336
315
|
The selection assessment can become decision-ready after enough evidence across task classes, but it
|
|
337
316
|
still cannot promote a candidate by itself. The remaining lifecycle and combined-stack gates must be
|
package/package.json
CHANGED
package/sbom.json
CHANGED
|
@@ -1,14 +1,14 @@
|
|
|
1
1
|
{
|
|
2
2
|
"bomFormat": "CycloneDX",
|
|
3
3
|
"specVersion": "1.5",
|
|
4
|
-
"serialNumber": "urn:uuid:
|
|
4
|
+
"serialNumber": "urn:uuid:584dee19-a861-2f17-f82a-93efadc2de0b",
|
|
5
5
|
"version": 1,
|
|
6
6
|
"metadata": {
|
|
7
7
|
"component": {
|
|
8
8
|
"type": "application",
|
|
9
9
|
"bom-ref": "token-harness",
|
|
10
10
|
"name": "token-harness",
|
|
11
|
-
"version": "0.1.
|
|
11
|
+
"version": "0.1.12",
|
|
12
12
|
"description": "Quota-aware efficiency layer for Claude Code and Codex subscription limits.",
|
|
13
13
|
"licenses": [
|
|
14
14
|
{
|
|
@@ -20,7 +20,7 @@
|
|
|
20
20
|
"hashes": [
|
|
21
21
|
{
|
|
22
22
|
"alg": "SHA-256",
|
|
23
|
-
"content": "
|
|
23
|
+
"content": "584dee19a8612f17f82a93efadc2de0b6c8b28fc96bd7a7d44b9e5af2dbb09bc"
|
|
24
24
|
}
|
|
25
25
|
]
|
|
26
26
|
},
|