token-harness 0.1.5 → 0.1.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +170 -13
- package/package.json +2 -2
- package/sbom.json +4 -4
- package/token-harness.mjs +12243 -8256
package/README.md
CHANGED
|
@@ -117,6 +117,144 @@ token-harness metrics --since 7d
|
|
|
117
117
|
`doctor` ends with a `NEXT` section. If you are unsure what to do, run the command shown
|
|
118
118
|
there.
|
|
119
119
|
|
|
120
|
+
### Codex: quota-aware optimization example
|
|
121
|
+
|
|
122
|
+
For a concrete Codex workflow, keep observation and mutation separate:
|
|
123
|
+
|
|
124
|
+
```sh
|
|
125
|
+
# Read-only: inspect live subscription windows.
|
|
126
|
+
token-harness budget --harness codex
|
|
127
|
+
|
|
128
|
+
# Read-only: inspect the effective model, reasoning effort, project instructions,
|
|
129
|
+
# MCP servers, and visible tool inventory.
|
|
130
|
+
token-harness context --harness codex
|
|
131
|
+
|
|
132
|
+
# Read-only: combine quota pacing and context pressure into task-specific advice.
|
|
133
|
+
token-harness optimize --harness codex --task standard --profile balanced
|
|
134
|
+
|
|
135
|
+
# Dry run: preview supported Codex-native policy changes.
|
|
136
|
+
token-harness plan --harness codex --native-policy
|
|
137
|
+
```
|
|
138
|
+
|
|
139
|
+
A typical optimization can recommend lowering reasoning effort for routine work when quota is
|
|
140
|
+
under-pace, while keeping the current model until empirical model-tier evidence exists. A plan may
|
|
141
|
+
also include reviewed provider actions such as a HarnessTrim skill install.
|
|
142
|
+
|
|
143
|
+
Apply the exact stored plan only after reviewing it:
|
|
144
|
+
|
|
145
|
+
```sh
|
|
146
|
+
token-harness apply --plan <plan-id> --yes
|
|
147
|
+
token-harness context --harness codex
|
|
148
|
+
token-harness verify --harness codex
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
On supported recent Codex builds, Token Harness uses Codex's native app-server for authoritative
|
|
152
|
+
rate-limit windows, effective configuration, model catalog, MCP inventory, and reviewed native
|
|
153
|
+
configuration writes. It does not infer quota from local token counts.
|
|
154
|
+
|
|
155
|
+
Useful variants:
|
|
156
|
+
|
|
157
|
+
```sh
|
|
158
|
+
# More conservative reasoning for routine work.
|
|
159
|
+
token-harness optimize --harness codex --task mechanical --profile economy
|
|
160
|
+
|
|
161
|
+
# Preserve more quality headroom for difficult work.
|
|
162
|
+
token-harness optimize --harness codex --task hard --profile quality
|
|
163
|
+
|
|
164
|
+
# Inspect MCP exposure directly.
|
|
165
|
+
token-harness mcp --harness codex
|
|
166
|
+
|
|
167
|
+
# Compare recent local history when available.
|
|
168
|
+
token-harness history --harness codex --since 7d
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
### Cross-harness scheduling and compact handoff
|
|
172
|
+
|
|
173
|
+
`token-harness schedule` is a read-only recommendation, not an automatic router. In the normal
|
|
174
|
+
installed CLI, unknown five-hour and weekly pace fields are hydrated from the same live budget
|
|
175
|
+
observer used by `token-harness budget`. Candidate quality and transfer benefit are hydrated only
|
|
176
|
+
from attributable project-local benchmark evidence. Missing or conflicting evidence returns
|
|
177
|
+
`insufficient-evidence` instead of guessing.
|
|
178
|
+
|
|
179
|
+
Start with the minimal command:
|
|
180
|
+
|
|
181
|
+
```sh
|
|
182
|
+
token-harness schedule \
|
|
183
|
+
--current claude \
|
|
184
|
+
--candidate codex \
|
|
185
|
+
--task-class hard \
|
|
186
|
+
--handoff-bytes 900
|
|
187
|
+
```
|
|
188
|
+
|
|
189
|
+
If the task is already in progress, generate a bounded handoff rather than copying the transcript:
|
|
190
|
+
|
|
191
|
+
```sh
|
|
192
|
+
token-harness handoff \
|
|
193
|
+
--objective "Finish the scheduler change without weakening quota evidence" \
|
|
194
|
+
--decision "Keep Claude and Codex quota observations provider-local" \
|
|
195
|
+
--changed-file apps/cli/src/schedule-main.ts \
|
|
196
|
+
--validation "pnpm test passes" \
|
|
197
|
+
--unresolved "Need one empirical Codex comparison" \
|
|
198
|
+
--next-action "Run the candidate benchmark in Codex" \
|
|
199
|
+
--max-bytes 2048 > handoff.md
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
To teach future recommendations whether that handoff was actually worth the switch, capture one
|
|
203
|
+
paired experiment. The baseline and optimized variants use the same benchmark id and task class,
|
|
204
|
+
but different harnesses:
|
|
205
|
+
|
|
206
|
+
```sh
|
|
207
|
+
# 1. Before doing the task in the current harness.
|
|
208
|
+
token-harness benchmark-start \
|
|
209
|
+
--benchmark-id scheduler-hard-01 \
|
|
210
|
+
--variant baseline \
|
|
211
|
+
--task hard \
|
|
212
|
+
--harness claude
|
|
213
|
+
|
|
214
|
+
# Run the baseline task in Claude Code, then record its observed outcome.
|
|
215
|
+
token-harness benchmark-finish \
|
|
216
|
+
--benchmark-id scheduler-hard-01 \
|
|
217
|
+
--variant baseline \
|
|
218
|
+
--quality passed \
|
|
219
|
+
--attempts 1 \
|
|
220
|
+
--failed-attempts 0
|
|
221
|
+
|
|
222
|
+
# 2. Before the candidate run, start the optimized variant.
|
|
223
|
+
token-harness benchmark-start \
|
|
224
|
+
--benchmark-id scheduler-hard-01 \
|
|
225
|
+
--variant optimized \
|
|
226
|
+
--task hard \
|
|
227
|
+
--harness codex
|
|
228
|
+
|
|
229
|
+
# Give handoff.md to Codex, run the equivalent task, then close the capture.
|
|
230
|
+
token-harness benchmark-finish \
|
|
231
|
+
--benchmark-id scheduler-hard-01 \
|
|
232
|
+
--variant optimized \
|
|
233
|
+
--quality passed \
|
|
234
|
+
--attempts 1 \
|
|
235
|
+
--failed-attempts 0
|
|
236
|
+
|
|
237
|
+
# 3. Evaluate the exact handoff used by the candidate run.
|
|
238
|
+
token-harness transfer \
|
|
239
|
+
--benchmark-id scheduler-hard-01 \
|
|
240
|
+
--handoff-file handoff.md
|
|
241
|
+
|
|
242
|
+
# 4. Persist the immutable transfer verdict for future schedule calls.
|
|
243
|
+
token-harness transfer-record \
|
|
244
|
+
--benchmark-id scheduler-hard-01 \
|
|
245
|
+
--handoff-file handoff.md
|
|
246
|
+
```
|
|
247
|
+
|
|
248
|
+
`transfer-record` re-validates the project-scoped pair, hashes the exact handoff and writes one
|
|
249
|
+
immutable receipt. Future `schedule` calls consume receipts only for the exact route and task class.
|
|
250
|
+
Every attributable receipt must agree on the same non-unknown transfer verdict; one `unknown` or a
|
|
251
|
+
positive/non-positive conflict keeps the recommendation unresolved. Historical receipt sizes never
|
|
252
|
+
replace the current `--handoff-bytes` value.
|
|
253
|
+
|
|
254
|
+
Manual evidence flags remain available for controlled experiments and debugging. Explicit values,
|
|
255
|
+
including an explicit `unknown`, always override automatic hydration. Run
|
|
256
|
+
`token-harness schedule --help` for the full list.
|
|
257
|
+
|
|
120
258
|
To try the read-only diagnosis without installing Token Harness globally:
|
|
121
259
|
|
|
122
260
|
```sh
|
|
@@ -129,7 +267,7 @@ HarnessTrim, or a coding agent.
|
|
|
129
267
|
### Managed compatibility rows
|
|
130
268
|
|
|
131
269
|
Token Harness changes a harness configuration only when a reviewed compatibility row covers the
|
|
132
|
-
exact provider version, harness version, platform, and configuration schema.
|
|
270
|
+
exact provider version, harness version, platform, and configuration schema. Four rows ship, and each
|
|
133
271
|
names the recording it stands on:
|
|
134
272
|
|
|
135
273
|
| Provider | Harness | Platform | Tested versions | Tier |
|
|
@@ -137,21 +275,22 @@ names the recording it stands on:
|
|
|
137
275
|
| RTK | Claude Code | Windows | rtk 0.44.0, Claude Code 2.1.220 | `canary` |
|
|
138
276
|
| HarnessTrim | Claude Code | Windows | harnesstrim 0.1.0, Claude Code 2.1.220 | `config-only` |
|
|
139
277
|
| HarnessTrim | Codex | Windows | harnesstrim 0.1.0, Codex 0.146.0 | `config-only` |
|
|
278
|
+
| HarnessTrim | Codex | Linux (non-WSL) | harnesstrim 0.2.1, Codex 0.152.1 | `config-only` |
|
|
140
279
|
|
|
141
280
|
Everything else is refused, and that is the design rather than a gap: `doctor` detects and reports on
|
|
142
281
|
every supported platform, and only the *mutation* is narrower. An uncovered combination exits 9 and
|
|
143
282
|
the diagnostic names what is missing — the reviewed fixture, or the nearest row it does have.
|
|
144
283
|
|
|
145
|
-
A live Linux
|
|
146
|
-
Codex 0.
|
|
147
|
-
|
|
148
|
-
|
|
284
|
+
A live Linux recording on 2026-09-02 promoted exactly one combination to managed mutation:
|
|
285
|
+
Codex 0.152.1 + HarnessTrim 0.2.1 on non-WSL Linux. The fixture covers an empty state, brownfield
|
|
286
|
+
user-owned files, skills-only apply, drift, verified rollback, and surgical uninstall. Nearby Codex
|
|
287
|
+
or HarnessTrim versions and WSL remain refused.
|
|
149
288
|
|
|
150
289
|
What is not covered today, and why:
|
|
151
290
|
|
|
152
|
-
- **macOS and Linux
|
|
153
|
-
|
|
154
|
-
|
|
291
|
+
- **macOS, WSL, and other Linux version combinations.** The recordings a row needs are states of a
|
|
292
|
+
real machine. Only the exact Linux combination above has been reviewed, so all other combinations
|
|
293
|
+
continue to refuse managed mutation while detection, verification, and measurement still work.
|
|
155
294
|
- **OpenCode, and permanently rather than pending.** Both providers are detected, adopted, verified
|
|
156
295
|
and measured there, and neither is written. RTK reaches OpenCode through a plugin module its own
|
|
157
296
|
installer places globally, which this build has no action for. HarnessTrim's OpenCode installer
|
|
@@ -519,12 +658,25 @@ Neither command removes user-owned RTK or HarnessTrim configuration.
|
|
|
519
658
|
|
|
520
659
|
| Command | Purpose | Changes agent/project configuration? |
|
|
521
660
|
| --- | --- | --- |
|
|
522
|
-
| `doctor` | Detect agents, providers, ownership, and problems | No |
|
|
523
|
-
| `
|
|
524
|
-
| `
|
|
661
|
+
| `doctor` | Detect agents, providers, ownership, versions, and problems | No |
|
|
662
|
+
| `budget` | Read live quota/headroom windows where the harness exposes them | No |
|
|
663
|
+
| `context` | Inspect effective model/config, instructions, MCP exposure, and tool inventory | No |
|
|
664
|
+
| `mcp` | Inspect MCP servers and tool-schema exposure | No |
|
|
665
|
+
| `history` | Summarize attributable local usage history | No |
|
|
666
|
+
| `optimize` | Combine quota pacing, context pressure, and task/profile policy into advice | No |
|
|
667
|
+
| `schedule` | Recommend Claude Code or Codex from independently attributable evidence | No |
|
|
668
|
+
| `handoff` | Build a bounded compact handoff for an in-progress cross-harness move | No |
|
|
669
|
+
| `plan` | Resolve ownership and preview exact actions; use `--native-policy` for supported harness-native changes | No; stores the plan in private state |
|
|
670
|
+
| `apply` | Apply a reviewed plan transactionally | Yes, only with `--yes` |
|
|
525
671
|
| `status` | Detect drift and competing hooks | No |
|
|
526
672
|
| `verify` | Check the declared verification tier | No |
|
|
527
673
|
| `metrics` | Import provider records and report savings | No; updates only Token Harness state |
|
|
674
|
+
| `benchmark` | Compare an explicit baseline/optimized receipt pair | No |
|
|
675
|
+
| `benchmark-start` | Start an empirical task capture | No agent/project config change; records Token Harness benchmark state |
|
|
676
|
+
| `benchmark-finish` | Finish an empirical task capture and record the outcome | No agent/project config change; records Token Harness benchmark state |
|
|
677
|
+
| `benchmark-matrix` | Aggregate complete project-scoped benchmark pairs by task class and evidence | No |
|
|
678
|
+
| `transfer` | Evaluate one empirical cross-harness benchmark pair and exact handoff | No |
|
|
679
|
+
| `transfer-record` | Persist one immutable project-scoped transfer evidence receipt | No agent/project config change; records Token Harness benchmark state |
|
|
528
680
|
| `update` | Query channels and update installed providers | Yes, only with `--yes` |
|
|
529
681
|
| `rollback` | Restore files from the latest committed transaction | Yes, only with `--yes` |
|
|
530
682
|
| `uninstall` | Remove owned integration entries | Yes, only with `--yes` |
|
|
@@ -532,9 +684,14 @@ Neither command removes user-owned RTK or HarnessTrim configuration.
|
|
|
532
684
|
Every command supports `--help`. Common filters are:
|
|
533
685
|
|
|
534
686
|
```text
|
|
535
|
-
--harness
|
|
536
|
-
--provider
|
|
687
|
+
--harness <id>
|
|
688
|
+
--provider <id>
|
|
537
689
|
--project <directory>
|
|
690
|
+
--task mechanical|standard|hard|critical
|
|
691
|
+
--profile economy|balanced|quality|custom
|
|
692
|
+
--reserve <percent>
|
|
693
|
+
--native-policy
|
|
694
|
+
--plan <plan-id>
|
|
538
695
|
--json
|
|
539
696
|
```
|
|
540
697
|
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "token-harness",
|
|
3
|
-
"version": "0.1.
|
|
4
|
-
"description": "
|
|
3
|
+
"version": "0.1.6",
|
|
4
|
+
"description": "Quota-aware efficiency layer for Claude Code and Codex subscription limits.",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
6
|
"type": "module",
|
|
7
7
|
"engines": {
|
package/sbom.json
CHANGED
|
@@ -1,15 +1,15 @@
|
|
|
1
1
|
{
|
|
2
2
|
"bomFormat": "CycloneDX",
|
|
3
3
|
"specVersion": "1.5",
|
|
4
|
-
"serialNumber": "urn:uuid:
|
|
4
|
+
"serialNumber": "urn:uuid:f6a37c3c-4b64-1136-564f-ab6db227d91a",
|
|
5
5
|
"version": 1,
|
|
6
6
|
"metadata": {
|
|
7
7
|
"component": {
|
|
8
8
|
"type": "application",
|
|
9
9
|
"bom-ref": "token-harness",
|
|
10
10
|
"name": "token-harness",
|
|
11
|
-
"version": "0.1.
|
|
12
|
-
"description": "
|
|
11
|
+
"version": "0.1.6",
|
|
12
|
+
"description": "Quota-aware efficiency layer for Claude Code and Codex subscription limits.",
|
|
13
13
|
"licenses": [
|
|
14
14
|
{
|
|
15
15
|
"license": {
|
|
@@ -20,7 +20,7 @@
|
|
|
20
20
|
"hashes": [
|
|
21
21
|
{
|
|
22
22
|
"alg": "SHA-256",
|
|
23
|
-
"content": "
|
|
23
|
+
"content": "f6a37c3c4b641136564fab6db227d91af645507cf96ea033321bc57cf567756a"
|
|
24
24
|
}
|
|
25
25
|
]
|
|
26
26
|
},
|