token-harness 0.1.9 → 0.1.10
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +79 -4
- package/package.json +1 -1
- package/sbom.json +3 -3
- package/token-harness.mjs +10249 -7817
package/README.md
CHANGED
|
@@ -112,14 +112,42 @@ change is not automatically available through `token-harness@latest`.
|
|
|
112
112
|
|
|
113
113
|
### Advanced and AI-assisted use
|
|
114
114
|
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
115
|
+
The long CLI flag combinations are **not** the normal human interface. They are the controller API
|
|
116
|
+
used by the app, automation, and optionally by a coding harness. Humans can keep using the browser
|
|
117
|
+
and the two entry points above.
|
|
118
|
+
|
|
119
|
+
For in-session use, this repository now includes a portable Agent Skill at
|
|
120
|
+
[`skills/token-harness/SKILL.md`](skills/token-harness/SKILL.md). A compatible Claude Code or Codex
|
|
121
|
+
skill mechanism can load it on demand, after which you can simply ask the harness to **use Token
|
|
122
|
+
Harness for this task**. The skill classifies substantial work conservatively, calls the existing
|
|
123
|
+
local `--json` optimizer at meaningful task boundaries, and can use explicit workload scheduling
|
|
124
|
+
when you have actually supplied a backlog. It does not run Token Harness before every tool call.
|
|
125
|
+
|
|
126
|
+
The skill is deliberately thin: Token Harness remains the deterministic policy engine. No MCP
|
|
127
|
+
server, background model, persistent agent, or second quota formula is added just to make this work.
|
|
128
|
+
Local tokens are still not subscription quota, and raw Claude/Codex percentages are still not a
|
|
129
|
+
common currency.
|
|
130
|
+
|
|
131
|
+
Agent use is read-only by default. If Token Harness recommends a persistent model, reasoning, or
|
|
132
|
+
verbosity change, the harness must first build a reviewed plan, explain the exact proposed mutation,
|
|
133
|
+
and wait for your explicit approval before `apply`. A saved preference may affect future sessions;
|
|
134
|
+
it is not silently presented as a live change to the current session.
|
|
135
|
+
|
|
136
|
+
The guided app can preview **Enable in-session guidance** for each detected Claude Code or
|
|
137
|
+
Codex installation. It installs the same portable skill into the agent's documented user-level
|
|
138
|
+
Agent Skills directory through the normal transactional plan/apply path. The agent card then shows
|
|
139
|
+
the live guidance state: **Enabled** when the exact skill is currently owned by Token Harness,
|
|
140
|
+
**Enabled externally** for a byte-identical user-owned skill, or an explicit not-enabled,
|
|
141
|
+
custom/conflict, or unavailable state. Existing `token-harness` skill directories are never
|
|
142
|
+
overwritten or silently adopted, and matching bytes alone never create an ownership claim. The
|
|
143
|
+
browser remains fully usable without any skill or second AI subscription. See
|
|
144
|
+
[RFC 0023](docs/rfcs/0023-guided-agent-skill-install.md) for the install and ownership boundary.
|
|
118
145
|
|
|
119
146
|
The older automation contracts remain available: `setup`, `optimize`, `plan`, `apply`,
|
|
120
147
|
`verify`, `metrics`, `rollback`, and their JSON reports. `ui --json` preserves its existing
|
|
121
148
|
schema-1 report; `ui --read-only` opens the legacy read-only dashboard. `ui --no-open` starts
|
|
122
|
-
the guided app without launching a browser. Stop either local server with Ctrl+C.
|
|
149
|
+
the guided app without launching a browser. Stop either local server with Ctrl+C. See
|
|
150
|
+
[RFC 0022](docs/rfcs/0022-agent-native-skill.md) for the agent-facing safety boundary.
|
|
123
151
|
|
|
124
152
|
## What normal output looks like
|
|
125
153
|
|
|
@@ -239,6 +267,53 @@ Most people do not need this section. Run `token-harness <command> --help` for d
|
|
|
239
267
|
| `handoff` | Build a bounded cross-harness handoff | No |
|
|
240
268
|
| `benchmark*`, `transfer*` | Capture and compare empirical evidence | Local state only |
|
|
241
269
|
|
|
270
|
+
## Workload-aware allowance planning
|
|
271
|
+
|
|
272
|
+
If you know how many accepted tasks remain, the advanced CLI can ask whether that backlog fits the
|
|
273
|
+
**currently observed** included allowance:
|
|
274
|
+
|
|
275
|
+
```sh
|
|
276
|
+
token-harness optimize --harness codex --task standard --tasks-left 5
|
|
277
|
+
token-harness schedule --current codex --candidate claude --task-class standard --tasks-left 5
|
|
278
|
+
```
|
|
279
|
+
|
|
280
|
+
`--tasks-left` is explicit workload intent for a backlog of one task class. Token Harness does not
|
|
281
|
+
infer it from `ccusage`, local tokens, session length, or raw provider percentages. A workload-driven
|
|
282
|
+
recommendation requires complete project-local benchmark evidence for the exact model + reasoning
|
|
283
|
+
effort + verbosity policy in both the five-hour and weekly windows. If that evidence is incomplete,
|
|
284
|
+
capacity stays unknown.
|
|
285
|
+
|
|
286
|
+
When evidence proves that the current policy cannot cover the stated backlog, `optimize` protects
|
|
287
|
+
capacity instead of spending a quota-derived effort bonus and reports whether the five-hour,
|
|
288
|
+
weekly, or both windows are limiting. `schedule` can use the same target to consider the other
|
|
289
|
+
harness, but only when that candidate has enough conservative accepted-task capacity and passes the
|
|
290
|
+
existing quality, pace, availability, and transfer checks.
|
|
291
|
+
|
|
292
|
+
### Mixed task-class backlog
|
|
293
|
+
|
|
294
|
+
For queued **new tasks** spanning more than one class, give `schedule` the mix explicitly:
|
|
295
|
+
|
|
296
|
+
```sh
|
|
297
|
+
token-harness schedule --current codex --candidate claude \
|
|
298
|
+
--workload mechanical=2,standard=3,hard=1
|
|
299
|
+
```
|
|
300
|
+
|
|
301
|
+
Mixed mode does not sum per-class task capacities as if they were separate quota buckets. It charges
|
|
302
|
+
each proposed task's project-local p75 cost against the same shared five-hour and weekly allowance
|
|
303
|
+
of that harness, then returns `stay`, `split`, `switch`, `shortfall`, or
|
|
304
|
+
`insufficient-evidence`. Candidate assignments require at least three coherent quality-gated
|
|
305
|
+
observations for the exact task class plus complete five-hour and weekly capacity evidence.
|
|
306
|
+
Unproven work remains visibly unallocated.
|
|
307
|
+
|
|
308
|
+
This mode is for queued/new tasks, not an in-progress handoff. Therefore `--workload` is mutually
|
|
309
|
+
exclusive with `--task-class`, `--tasks-left`, manual pace/quality flags, and handoff/transfer
|
|
310
|
+
flags. The allocator is deterministic and conservative; it does not claim globally optimal routing,
|
|
311
|
+
launch either harness, or compare raw Claude and Codex percentages.
|
|
312
|
+
|
|
313
|
+
No capacity after a future reset is assumed. Re-run the observation after the reset rather than
|
|
314
|
+
treating a forecast as provider quota. See [RFC 0020](docs/rfcs/0020-workload-aware-allowance.md)
|
|
315
|
+
and [RFC 0021](docs/rfcs/0021-mixed-workload-allocation.md).
|
|
316
|
+
|
|
242
317
|
## Applying native recommendations
|
|
243
318
|
|
|
244
319
|
`optimize` remains read-only. Review a plan before applying a supported native change:
|
package/package.json
CHANGED
package/sbom.json
CHANGED
|
@@ -1,14 +1,14 @@
|
|
|
1
1
|
{
|
|
2
2
|
"bomFormat": "CycloneDX",
|
|
3
3
|
"specVersion": "1.5",
|
|
4
|
-
"serialNumber": "urn:uuid:
|
|
4
|
+
"serialNumber": "urn:uuid:9290051f-d6be-441a-20e1-add1f9ea0ff8",
|
|
5
5
|
"version": 1,
|
|
6
6
|
"metadata": {
|
|
7
7
|
"component": {
|
|
8
8
|
"type": "application",
|
|
9
9
|
"bom-ref": "token-harness",
|
|
10
10
|
"name": "token-harness",
|
|
11
|
-
"version": "0.1.
|
|
11
|
+
"version": "0.1.10",
|
|
12
12
|
"description": "Quota-aware efficiency layer for Claude Code and Codex subscription limits.",
|
|
13
13
|
"licenses": [
|
|
14
14
|
{
|
|
@@ -20,7 +20,7 @@
|
|
|
20
20
|
"hashes": [
|
|
21
21
|
{
|
|
22
22
|
"alg": "SHA-256",
|
|
23
|
-
"content": "
|
|
23
|
+
"content": "9290051fd6be441a20e1add1f9ea0ff88e9ae931921532ed8422d86db26f1b3a"
|
|
24
24
|
}
|
|
25
25
|
]
|
|
26
26
|
},
|