token-harness 0.1.11 → 0.1.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -3,8 +3,8 @@
3
3
  **Build, verify and measure an optimization stack for Claude Code and Codex.**
4
4
 
5
5
  Token Harness is a local **optimization stack manager**. It checks your coding agents, manages the
6
- optimization components it can safely own, keeps experimental candidates separate, verifies the
7
- result, and reports savings only when it has evidence to support them.
6
+ optimization components it can safely own, keeps evaluation evidence separate from managed lifecycle,
7
+ verifies the result, and reports savings only when it has evidence to support them.
8
8
 
9
9
  It is not another coding agent and it does not replace specialized projects such as RTK or
10
10
  HarnessTrim.
@@ -34,10 +34,9 @@ of CLI commands to memorize.
34
34
  3. In **Setup**, work from top to bottom:
35
35
  - Coding agents
36
36
  - Managed optimizers
37
- - Experimental tools
38
37
  - Optional agent tuning
39
38
  - Checks and maintenance
40
- 4. For Claude Code or Codex, choose **Review setup** when a managed setup is available.
39
+ 4. For Claude Code or Codex, choose **Review baseline** when the RTK + HarnessTrim setup is available.
41
40
  5. Read the exact proposed changes. **Apply reviewed setup** appears only when there is a concrete
42
41
  safe plan to apply.
43
42
  6. Keep using Claude Code or Codex normally.
@@ -72,8 +71,10 @@ useful read-only allowance and connected-tool observations when available.
72
71
 
73
72
  **2. Managed optimizers**
74
73
 
75
- The managed stack currently consists of **RTK + HarnessTrim** on individually reviewed
76
- combinations. Token Harness tracks the exact combined provider set separately: if no combined-stack
74
+ The production baseline remains **RTK + HarnessTrim** on individually reviewed combinations.
75
+ Token Harness also exposes **mcptoon**, **GitNexus** and **Headroom** as optional managed integrations on their exact
76
+ reviewed lifecycle rows; enabling any of them does not make it part of the production baseline or create a
77
+ savings claim. Token Harness tracks the exact combined provider set separately: if no combined-stack
77
78
  review is recorded, Setup says so and keeps the stack incomplete rather than inferring compatibility
78
79
  from healthy individual checks. Token Harness can prepare their integration transactionally, show the
79
80
  exact plan, apply it only after approval, verify it, and remove only configuration it owns.
@@ -91,45 +92,17 @@ shipped combined-review registry stays empty until the captured configuration ha
91
92
  verification evidence, has been benchmarked together, and has been deliberately reviewed; see
92
93
  `docs/combined-stack-reviews.md`.
93
94
 
94
- **3. Experimental tools**
95
-
96
- Headroom, mcptoon and GitNexus are visible as candidates, not silently promoted dependencies. Their
97
- cards distinguish CLI installation, candidate-side activation/evaluation, Token Harness benchmark
98
- evidence and promotion-readiness gates. External install or activation commands are shown for
99
- review; Token Harness does not silently execute package managers, activate wrappers, index
100
- repositories or register MCP servers.
101
-
102
- Choose **Run standard evaluation** to start or resume a paired candidate campaign. Campaign state is
103
- scoped to both the candidate and the selected harness, so Claude Code and Codex evidence cannot be
104
- mixed accidentally. The browser reads the campaign directly and shows **Progress**, the current
105
- selection signal, whether the evidence is **Decision ready**, the number of evidence-bearing pairs
106
- and the exact **Next** step.
107
-
108
- For normal use, the browser can now start and finish the local benchmark capture itself. You still
109
- run the actual task in Claude Code or Codex. When the task finishes, record the quality result,
110
- attempt count and failed-attempt count you actually observed. Before an optimized capture, Token
111
- Harness requires you to acknowledge that you enabled the candidate through its own documented
112
- workflow. That acknowledgement is **not activation verification** and is never treated as promotion
113
- evidence. Every browser capture action is matched against the campaign engine's current step before
114
- it can write local benchmark state, so a stale tab cannot advance a different step.
115
-
116
- `benchmark-matrix`, `benchmark-start` and `benchmark-finish` remain available as advanced terminal
117
- fallbacks for debugging or automation. The browser does not run the coding task, install or activate
118
- a candidate, or claim that candidate attribution proves activation.
119
-
120
- A campaign selection assessment can report `insufficient-evidence`, `promising`, `mixed` or
121
- `negative`, plus whether the evidence is decision-ready. **Decision-ready is not promotion-ready.**
122
- The campaign shows the same promotion-review gate count and next gate as the candidate card, using
123
- the current local observation rather than starting another environment scan. Activation
124
- verification, managed lifecycle, compatibility/reversibility, project maturity and combined-stack
125
- validation remain separate gates.
95
+ **3. Evaluation evidence (maintainers)**
96
+
97
+ Selection campaigns and historical candidate evidence remain available for maintainers, but they are
98
+ not a separate novice setup product. Once a tool has a reviewed managed lifecycle, Setup presents it in
99
+ the same **Optimization Stack** with its exact state, version, prerequisites, harness scope and safe
100
+ Apply/Remove controls. Evaluation evidence never turns into a savings or compatibility claim by itself.
126
101
 
127
- Choose **Compare evaluation evidence** when you want one read-only view of campaigns you already
128
- started. The comparison is loaded only on request, never creates a campaign, keeps candidate and
129
- harness ordering fixed, and reports progress, selection signal, decision readiness, runtime
130
- activation evidence and promotion gates without producing a composite score or automatic winner.
131
- Activation is shown separately as **Verified**, **Blocked** or **Unreviewed**, with
132
- verified/blocked/unknown pair counts when campaign evidence exists.
102
+ Headroom is config-only managed on its exact reviewed 0.37.0 row. Token Harness can own the narrow
103
+ Claude Code or Codex MCP registration after the pinned CLI is present; it deliberately does not
104
+ bootstrap `uv`/Python, start wrappers/proxies, or claim runtime compression savings. GitNexus likewise
105
+ does not auto-index repositories, and its noncommercial license boundary stays visible.
133
106
 
134
107
  **4. Optional agent tuning**
135
108
 
@@ -209,6 +182,9 @@ algorithms into this repository.
209
182
  | --- | --- | --- |
210
183
  | [RTK](https://github.com/rtk-ai/rtk) | Shell/tool output reduction | Managed on reviewed combinations |
211
184
  | [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic output/context reduction | Managed first-party integration |
185
+ | mcptoon | MCP discovery / compact manifest guidance | Optional managed integration on exact reviewed 0.7.10 rows; no savings assumed |
186
+ | GitNexus | Repository graph / MCP context | Optional managed Claude integration for already-installed 1.6.12; license review required |
187
+ | Headroom | Local MCP context compression/retrieval | Optional config-only managed Claude/Codex integration for already-installed 0.37.0; package prerequisite stays user-owned |
212
188
  | [cclimits](https://github.com/cruzanstx/cclimits) | Optional Claude allowance evidence | Read-only evidence; not an optimizer |
213
189
  | [ccusage](https://github.com/ccusage/ccusage) | Local usage history | Read-only evidence; never subscription quota |
214
190
 
@@ -233,8 +209,8 @@ RTK has no equivalent machine-readable capability endpoint, so releases newer th
233
209
  reviewed RTK set remain visible as `unknown-newer` until their consumed contract is checked. See
234
210
  [docs/provider-version-compatibility.md](docs/provider-version-compatibility.md).
235
211
 
236
- Current experimental candidates include Headroom, mcptoon and GitNexus. Detection or a promising
237
- benchmark is not enough for promotion. Their campaign assessment is structured evidence for the
212
+ Historical evaluation evidence remains available for mcptoon, GitNexus and Headroom. Detection or a promising
213
+ benchmark is not enough for a production-stack promotion or savings claim. Their campaign assessment is structured evidence for the
238
214
  selection gate, not an activation or promotion decision. A candidate must pass structured promotion
239
215
  readiness across benchmark capability, category fit, selection evidence, real activation
240
216
  verification, managed lifecycle, compatibility/reversibility, project maturity and combined-stack
@@ -307,31 +283,34 @@ The older automation contracts remain available. `ui --json` preserves its exist
307
283
  report; `ui --read-only` opens the legacy read-only UI; `ui --no-open` starts the guided app without
308
284
  launching a browser.
309
285
 
310
- ### Evaluating an experimental candidate
286
+ ### Evaluation evidence (advanced / maintainers)
311
287
 
312
- For normal use, open **Setup -> Experimental tools** and choose **Run standard evaluation**. The app
288
+ Evaluation campaigns are an advanced maintainer workflow; managed setup stays in the unified **Optimization Stack**. The app
313
289
  keeps a resumable campaign ID for each candidate/harness pair and reads campaign progress, assessment
314
290
  and the exact **Next** step directly in the browser. Use **Start baseline capture** or **Start optimized
315
291
  capture**, run the requested task in the selected coding agent, then choose **Record outcome** and
316
292
  enter the quality/attempt values you actually observed. Normal use no longer requires copying
317
293
  `benchmark-start` or `benchmark-finish` commands into a terminal.
318
294
 
319
- The equivalent advanced CLI flow starts by asking the campaign engine for its current state:
295
+ The equivalent advanced CLI flow starts by asking the campaign engine for its current state. This
296
+ GitNexus example intentionally uses the only currently reviewed campaign row:
320
297
 
321
298
  ```sh
322
299
  token-harness benchmark-matrix \
323
- --benchmark-id gitnexus-codex-eval-1 \
300
+ --benchmark-id gitnexus-claude-eval-1 \
324
301
  --candidate gitnexus \
325
- --harness codex
302
+ --harness claude
326
303
  ```
327
304
 
328
305
  Follow only the **Next** command printed by that report, complete the task honestly, then rerun the
329
306
  same `benchmark-matrix` command. Before an optimized run, enable the candidate through its own
330
307
  documented workflow. Token Harness records the experiment target but does not treat attribution—or
331
- the browser acknowledgement—as proof that the candidate was active. For GitNexus, new benchmark
332
- receipts can additionally verify activation from the harness-native MCP inventory when the GitNexus
333
- server is observed usable at both task boundaries; missing or ambiguous runtime evidence stays
334
- unverified.
308
+ the browser acknowledgement—as proof that the candidate was active. For GitNexus on the reviewed
309
+ Claude Code `2.1.269` × GitNexus `1.6.12` × native-Linux row, the harness-native MCP inventory can
310
+ prove only that the GitNexus server was available at both task boundaries. That is not proof Claude
311
+ actually called a GitNexus tool, so the activation-verification promotion gate remains blocked
312
+ without a separate reviewed usage witness. See
313
+ [`docs/candidates/gitnexus-real-campaign.md`](docs/candidates/gitnexus-real-campaign.md).
335
314
 
336
315
  The selection assessment can become decision-ready after enough evidence across task classes, but it
337
316
  still cannot promote a candidate by itself. The remaining lifecycle and combined-stack gates must be
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "token-harness",
3
- "version": "0.1.11",
3
+ "version": "0.1.12",
4
4
  "description": "Quota-aware efficiency layer for Claude Code and Codex subscription limits.",
5
5
  "license": "Apache-2.0",
6
6
  "type": "module",
package/sbom.json CHANGED
@@ -1,14 +1,14 @@
1
1
  {
2
2
  "bomFormat": "CycloneDX",
3
3
  "specVersion": "1.5",
4
- "serialNumber": "urn:uuid:01436fc5-44c9-56ef-29b7-e0847e19fad9",
4
+ "serialNumber": "urn:uuid:584dee19-a861-2f17-f82a-93efadc2de0b",
5
5
  "version": 1,
6
6
  "metadata": {
7
7
  "component": {
8
8
  "type": "application",
9
9
  "bom-ref": "token-harness",
10
10
  "name": "token-harness",
11
- "version": "0.1.11",
11
+ "version": "0.1.12",
12
12
  "description": "Quota-aware efficiency layer for Claude Code and Codex subscription limits.",
13
13
  "licenses": [
14
14
  {
@@ -20,7 +20,7 @@
20
20
  "hashes": [
21
21
  {
22
22
  "alg": "SHA-256",
23
- "content": "01436fc544c956ef29b7e0847e19fad9e439c4cdfccc2517e007156a2b9fb2e0"
23
+ "content": "584dee19a8612f17f82a93efadc2de0b6c8b28fc96bd7a7d44b9e5af2dbb09bc"
24
24
  }
25
25
  ]
26
26
  },