token-harness 0.1.5 → 0.1.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (4) hide show
  1. package/README.md +170 -13
  2. package/package.json +2 -2
  3. package/sbom.json +4 -4
  4. package/token-harness.mjs +12243 -8256
package/README.md CHANGED
@@ -117,6 +117,144 @@ token-harness metrics --since 7d
117
117
  `doctor` ends with a `NEXT` section. If you are unsure what to do, run the command shown
118
118
  there.
119
119
 
120
+ ### Codex: quota-aware optimization example
121
+
122
+ For a concrete Codex workflow, keep observation and mutation separate:
123
+
124
+ ```sh
125
+ # Read-only: inspect live subscription windows.
126
+ token-harness budget --harness codex
127
+
128
+ # Read-only: inspect the effective model, reasoning effort, project instructions,
129
+ # MCP servers, and visible tool inventory.
130
+ token-harness context --harness codex
131
+
132
+ # Read-only: combine quota pacing and context pressure into task-specific advice.
133
+ token-harness optimize --harness codex --task standard --profile balanced
134
+
135
+ # Dry run: preview supported Codex-native policy changes.
136
+ token-harness plan --harness codex --native-policy
137
+ ```
138
+
139
+ A typical optimization can recommend lowering reasoning effort for routine work when quota is
140
+ under-pace, while keeping the current model until empirical model-tier evidence exists. A plan may
141
+ also include reviewed provider actions such as a HarnessTrim skill install.
142
+
143
+ Apply the exact stored plan only after reviewing it:
144
+
145
+ ```sh
146
+ token-harness apply --plan <plan-id> --yes
147
+ token-harness context --harness codex
148
+ token-harness verify --harness codex
149
+ ```
150
+
151
+ On supported recent Codex builds, Token Harness uses Codex's native app-server for authoritative
152
+ rate-limit windows, effective configuration, model catalog, MCP inventory, and reviewed native
153
+ configuration writes. It does not infer quota from local token counts.
154
+
155
+ Useful variants:
156
+
157
+ ```sh
158
+ # More conservative reasoning for routine work.
159
+ token-harness optimize --harness codex --task mechanical --profile economy
160
+
161
+ # Preserve more quality headroom for difficult work.
162
+ token-harness optimize --harness codex --task hard --profile quality
163
+
164
+ # Inspect MCP exposure directly.
165
+ token-harness mcp --harness codex
166
+
167
+ # Compare recent local history when available.
168
+ token-harness history --harness codex --since 7d
169
+ ```
170
+
171
+ ### Cross-harness scheduling and compact handoff
172
+
173
+ `token-harness schedule` is a read-only recommendation, not an automatic router. In the normal
174
+ installed CLI, unknown five-hour and weekly pace fields are hydrated from the same live budget
175
+ observer used by `token-harness budget`. Candidate quality and transfer benefit are hydrated only
176
+ from attributable project-local benchmark evidence. Missing or conflicting evidence returns
177
+ `insufficient-evidence` instead of guessing.
178
+
179
+ Start with the minimal command:
180
+
181
+ ```sh
182
+ token-harness schedule \
183
+ --current claude \
184
+ --candidate codex \
185
+ --task-class hard \
186
+ --handoff-bytes 900
187
+ ```
188
+
189
+ If the task is already in progress, generate a bounded handoff rather than copying the transcript:
190
+
191
+ ```sh
192
+ token-harness handoff \
193
+ --objective "Finish the scheduler change without weakening quota evidence" \
194
+ --decision "Keep Claude and Codex quota observations provider-local" \
195
+ --changed-file apps/cli/src/schedule-main.ts \
196
+ --validation "pnpm test passes" \
197
+ --unresolved "Need one empirical Codex comparison" \
198
+ --next-action "Run the candidate benchmark in Codex" \
199
+ --max-bytes 2048 > handoff.md
200
+ ```
201
+
202
+ To teach future recommendations whether that handoff was actually worth the switch, capture one
203
+ paired experiment. The baseline and optimized variants use the same benchmark id and task class,
204
+ but different harnesses:
205
+
206
+ ```sh
207
+ # 1. Before doing the task in the current harness.
208
+ token-harness benchmark-start \
209
+ --benchmark-id scheduler-hard-01 \
210
+ --variant baseline \
211
+ --task hard \
212
+ --harness claude
213
+
214
+ # Run the baseline task in Claude Code, then record its observed outcome.
215
+ token-harness benchmark-finish \
216
+ --benchmark-id scheduler-hard-01 \
217
+ --variant baseline \
218
+ --quality passed \
219
+ --attempts 1 \
220
+ --failed-attempts 0
221
+
222
+ # 2. Before the candidate run, start the optimized variant.
223
+ token-harness benchmark-start \
224
+ --benchmark-id scheduler-hard-01 \
225
+ --variant optimized \
226
+ --task hard \
227
+ --harness codex
228
+
229
+ # Give handoff.md to Codex, run the equivalent task, then close the capture.
230
+ token-harness benchmark-finish \
231
+ --benchmark-id scheduler-hard-01 \
232
+ --variant optimized \
233
+ --quality passed \
234
+ --attempts 1 \
235
+ --failed-attempts 0
236
+
237
+ # 3. Evaluate the exact handoff used by the candidate run.
238
+ token-harness transfer \
239
+ --benchmark-id scheduler-hard-01 \
240
+ --handoff-file handoff.md
241
+
242
+ # 4. Persist the immutable transfer verdict for future schedule calls.
243
+ token-harness transfer-record \
244
+ --benchmark-id scheduler-hard-01 \
245
+ --handoff-file handoff.md
246
+ ```
247
+
248
+ `transfer-record` re-validates the project-scoped pair, hashes the exact handoff and writes one
249
+ immutable receipt. Future `schedule` calls consume receipts only for the exact route and task class.
250
+ Every attributable receipt must agree on the same non-unknown transfer verdict; one `unknown` or a
251
+ positive/non-positive conflict keeps the recommendation unresolved. Historical receipt sizes never
252
+ replace the current `--handoff-bytes` value.
253
+
254
+ Manual evidence flags remain available for controlled experiments and debugging. Explicit values,
255
+ including an explicit `unknown`, always override automatic hydration. Run
256
+ `token-harness schedule --help` for the full list.
257
+
120
258
  To try the read-only diagnosis without installing Token Harness globally:
121
259
 
122
260
  ```sh
@@ -129,7 +267,7 @@ HarnessTrim, or a coding agent.
129
267
  ### Managed compatibility rows
130
268
 
131
269
  Token Harness changes a harness configuration only when a reviewed compatibility row covers the
132
- exact provider version, harness version, platform, and configuration schema. Three rows ship, and each
270
+ exact provider version, harness version, platform, and configuration schema. Four rows ship, and each
133
271
  names the recording it stands on:
134
272
 
135
273
  | Provider | Harness | Platform | Tested versions | Tier |
@@ -137,21 +275,22 @@ names the recording it stands on:
137
275
  | RTK | Claude Code | Windows | rtk 0.44.0, Claude Code 2.1.220 | `canary` |
138
276
  | HarnessTrim | Claude Code | Windows | harnesstrim 0.1.0, Claude Code 2.1.220 | `config-only` |
139
277
  | HarnessTrim | Codex | Windows | harnesstrim 0.1.0, Codex 0.146.0 | `config-only` |
278
+ | HarnessTrim | Codex | Linux (non-WSL) | harnesstrim 0.2.1, Codex 0.152.1 | `config-only` |
140
279
 
141
280
  Everything else is refused, and that is the design rather than a gap: `doctor` detects and reports on
142
281
  every supported platform, and only the *mutation* is narrower. An uncovered combination exits 9 and
143
282
  the diagnostic names what is missing — the reviewed fixture, or the nearest row it does have.
144
283
 
145
- A live Linux check on 2026-09-02 with Codex 0.152.1 confirmed that boundary: the current row covers
146
- Codex 0.146.0 on Windows only, so `plan` and `apply --yes` refused with exit 9 and wrote nothing.
147
- That combination can still be observed and benchmarked with a provider installed through its own
148
- installer; it is not promoted to managed mutation until a real compatibility fixture is recorded.
284
+ A live Linux recording on 2026-09-02 promoted exactly one combination to managed mutation:
285
+ Codex 0.152.1 + HarnessTrim 0.2.1 on non-WSL Linux. The fixture covers an empty state, brownfield
286
+ user-owned files, skills-only apply, drift, verified rollback, and surgical uninstall. Nearby Codex
287
+ or HarnessTrim versions and WSL remain refused.
149
288
 
150
289
  What is not covered today, and why:
151
290
 
152
- - **macOS and Linux.** No row on either. The recordings a row needs are states of a real machine, and
153
- a fixture cannot be written from a machine nobody ran. On those platforms `plan` and `apply` refuse;
154
- install the provider with its own installer and Token Harness will detect, verify, and measure it.
291
+ - **macOS, WSL, and other Linux version combinations.** The recordings a row needs are states of a
292
+ real machine. Only the exact Linux combination above has been reviewed, so all other combinations
293
+ continue to refuse managed mutation while detection, verification, and measurement still work.
155
294
  - **OpenCode, and permanently rather than pending.** Both providers are detected, adopted, verified
156
295
  and measured there, and neither is written. RTK reaches OpenCode through a plugin module its own
157
296
  installer places globally, which this build has no action for. HarnessTrim's OpenCode installer
@@ -519,12 +658,25 @@ Neither command removes user-owned RTK or HarnessTrim configuration.
519
658
 
520
659
  | Command | Purpose | Changes agent/project configuration? |
521
660
  | --- | --- | --- |
522
- | `doctor` | Detect agents, providers, ownership, and problems | No |
523
- | `plan` | Resolve ownership and preview exact actions | No; stores the plan in private state |
524
- | `apply` | Apply a plan transactionally | Yes, only with `--yes` |
661
+ | `doctor` | Detect agents, providers, ownership, versions, and problems | No |
662
+ | `budget` | Read live quota/headroom windows where the harness exposes them | No |
663
+ | `context` | Inspect effective model/config, instructions, MCP exposure, and tool inventory | No |
664
+ | `mcp` | Inspect MCP servers and tool-schema exposure | No |
665
+ | `history` | Summarize attributable local usage history | No |
666
+ | `optimize` | Combine quota pacing, context pressure, and task/profile policy into advice | No |
667
+ | `schedule` | Recommend Claude Code or Codex from independently attributable evidence | No |
668
+ | `handoff` | Build a bounded compact handoff for an in-progress cross-harness move | No |
669
+ | `plan` | Resolve ownership and preview exact actions; use `--native-policy` for supported harness-native changes | No; stores the plan in private state |
670
+ | `apply` | Apply a reviewed plan transactionally | Yes, only with `--yes` |
525
671
  | `status` | Detect drift and competing hooks | No |
526
672
  | `verify` | Check the declared verification tier | No |
527
673
  | `metrics` | Import provider records and report savings | No; updates only Token Harness state |
674
+ | `benchmark` | Compare an explicit baseline/optimized receipt pair | No |
675
+ | `benchmark-start` | Start an empirical task capture | No agent/project config change; records Token Harness benchmark state |
676
+ | `benchmark-finish` | Finish an empirical task capture and record the outcome | No agent/project config change; records Token Harness benchmark state |
677
+ | `benchmark-matrix` | Aggregate complete project-scoped benchmark pairs by task class and evidence | No |
678
+ | `transfer` | Evaluate one empirical cross-harness benchmark pair and exact handoff | No |
679
+ | `transfer-record` | Persist one immutable project-scoped transfer evidence receipt | No agent/project config change; records Token Harness benchmark state |
528
680
  | `update` | Query channels and update installed providers | Yes, only with `--yes` |
529
681
  | `rollback` | Restore files from the latest committed transaction | Yes, only with `--yes` |
530
682
  | `uninstall` | Remove owned integration entries | Yes, only with `--yes` |
@@ -532,9 +684,14 @@ Neither command removes user-owned RTK or HarnessTrim configuration.
532
684
  Every command supports `--help`. Common filters are:
533
685
 
534
686
  ```text
535
- --harness claude|codex|opencode
536
- --provider rtk|harnesstrim
687
+ --harness <id>
688
+ --provider <id>
537
689
  --project <directory>
690
+ --task mechanical|standard|hard|critical
691
+ --profile economy|balanced|quality|custom
692
+ --reserve <percent>
693
+ --native-policy
694
+ --plan <plan-id>
538
695
  --json
539
696
  ```
540
697
 
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "token-harness",
3
- "version": "0.1.5",
4
- "description": "One control plane for token-efficient coding agents.",
3
+ "version": "0.1.6",
4
+ "description": "Quota-aware efficiency layer for Claude Code and Codex subscription limits.",
5
5
  "license": "Apache-2.0",
6
6
  "type": "module",
7
7
  "engines": {
package/sbom.json CHANGED
@@ -1,15 +1,15 @@
1
1
  {
2
2
  "bomFormat": "CycloneDX",
3
3
  "specVersion": "1.5",
4
- "serialNumber": "urn:uuid:f4cf234e-2ad8-b2a3-a085-7056eb8a17c4",
4
+ "serialNumber": "urn:uuid:f6a37c3c-4b64-1136-564f-ab6db227d91a",
5
5
  "version": 1,
6
6
  "metadata": {
7
7
  "component": {
8
8
  "type": "application",
9
9
  "bom-ref": "token-harness",
10
10
  "name": "token-harness",
11
- "version": "0.1.5",
12
- "description": "One control plane for token-efficient coding agents.",
11
+ "version": "0.1.6",
12
+ "description": "Quota-aware efficiency layer for Claude Code and Codex subscription limits.",
13
13
  "licenses": [
14
14
  {
15
15
  "license": {
@@ -20,7 +20,7 @@
20
20
  "hashes": [
21
21
  {
22
22
  "alg": "SHA-256",
23
- "content": "f4cf234e2ad8b2a3a0857056eb8a17c494730c3d5c11e702edc98b270f18d203"
23
+ "content": "f6a37c3c4b641136564fab6db227d91af645507cf96ea033321bc57cf567756a"
24
24
  }
25
25
  ]
26
26
  },