token-harness 0.1.2 → 0.1.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +118 -50
- package/package.json +1 -1
- package/sbom.json +3 -2
- package/token-harness.mjs +7059 -1062
package/README.md
CHANGED
|
@@ -1,51 +1,81 @@
|
|
|
1
1
|
# Token Harness
|
|
2
2
|
|
|
3
|
-
Token Harness has one objective: **
|
|
4
|
-
|
|
3
|
+
Token Harness has one objective: **maximize the useful coding work you can get from Claude Code
|
|
4
|
+
and Codex usage limits, without hiding quality regressions or pretending that opaque subscription
|
|
5
|
+
quota is exactly equivalent to a token count**.
|
|
5
6
|
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
7
|
+
The scarce resource is no longer just model context. Subscription users are constrained by rolling
|
|
8
|
+
usage windows, weekly limits, model-dependent burn, long-session context growth, MCP/tool schema
|
|
9
|
+
overhead, noisy tool output, and repeated work after poor model or effort choices. Token Harness is
|
|
10
|
+
evolving from a token-reduction control plane into a **quota-aware efficiency layer** for coding
|
|
11
|
+
harnesses.
|
|
10
12
|
|
|
11
|
-
|
|
12
|
-
token-saving tools on the machine, selects a compatible owner for each interception point, shows
|
|
13
|
-
every proposed change before applying it, verifies whether the integration is genuinely being
|
|
14
|
-
used, and reports how many tokens or characters were saved.
|
|
13
|
+
The product therefore optimizes two related budgets:
|
|
15
14
|
|
|
16
|
-
|
|
17
|
-
|
|
15
|
+
1. **subscription headroom** — observe the usage windows the harness itself exposes, pace work
|
|
16
|
+
against reset time, and choose native model/effort settings that fit the task and remaining
|
|
17
|
+
budget;
|
|
18
|
+
2. **context sent to the model** — keep instructions, MCP schemas, conversation history, repository
|
|
19
|
+
context, and tool output as small as possible without removing information needed to finish the
|
|
20
|
+
task correctly.
|
|
21
|
+
|
|
22
|
+
RTK and HarnessTrim remain useful providers, but reducers are now one layer of the system rather
|
|
23
|
+
than the product definition. Native harness controls come first because choosing the right model,
|
|
24
|
+
effort, context shape, and enabled tools can avoid an expensive turn entirely.
|
|
25
|
+
|
|
26
|
+
## What Token Harness should optimize
|
|
27
|
+
|
|
28
|
+
| Layer | Target behavior |
|
|
29
|
+
| --- | --- |
|
|
30
|
+
| Usage-window observability | Read five-hour, weekly, model-specific, or credit-backed limits only from surfaces the harness can actually prove; show reset time, headroom, and burn rate |
|
|
31
|
+
| Native model policy | Prefer economical models/effort for routine work, escalate only for tasks whose expected quality benefit justifies the extra quota |
|
|
32
|
+
| Session hygiene | Detect task boundaries, oversized conversations, and stale context; recommend clear/compact/new-session actions before history dominates each turn |
|
|
33
|
+
| Instruction budget | Keep `CLAUDE.md` / `AGENTS.md` concise and hierarchical instead of injecting one large global instruction file everywhere |
|
|
34
|
+
| MCP/tool budget | Disable irrelevant MCP servers, defer schemas where the harness supports it, and expose only the tools required by the current task |
|
|
35
|
+
| Tool-output budget | Reduce logs, diffs, test output, and repeated command results before they re-enter model context |
|
|
36
|
+
| Cross-harness scheduling | When Claude Code and Codex have independent headroom, recommend the harness that can do the task with the best expected quality per remaining quota |
|
|
37
|
+
| Measurement | Correlate quota delta, model/effort, context size, tool traffic, and task outcome; never convert local token savings into an unproven subscription-quota claim |
|
|
18
38
|
|
|
19
39
|
## Optimization ecosystem
|
|
20
40
|
|
|
21
|
-
The
|
|
22
|
-
|
|
23
|
-
|
|
41
|
+
The priority order below is deliberately different from the original Token Harness roadmap. A tool
|
|
42
|
+
is high priority only when it helps **included Claude Code/Codex capacity**, not merely API cost,
|
|
43
|
+
provider routing, or benchmark token counts.
|
|
24
44
|
|
|
25
|
-
| Tool | Optimization layer | Token Harness
|
|
45
|
+
| Tool or surface | Optimization layer | Token Harness direction |
|
|
26
46
|
| --- | --- | --- |
|
|
27
|
-
|
|
|
28
|
-
|
|
|
29
|
-
| [
|
|
30
|
-
| [
|
|
31
|
-
| [
|
|
32
|
-
| [
|
|
33
|
-
| [
|
|
34
|
-
| [
|
|
35
|
-
| [
|
|
36
|
-
| [
|
|
37
|
-
| [
|
|
38
|
-
|
|
|
39
|
-
| [LLMLingua](https://github.com/microsoft/LLMLingua) |
|
|
40
|
-
| [Caveman](https://github.com/JuliusBrussee/caveman) | Reduce visible model-output verbosity |
|
|
41
|
-
|
|
42
|
-
Candidate status means
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
merely because it is present on the machine. Rows marked `high priority` are the next intended
|
|
46
|
-
intake; the admission gates each one carries are recorded in
|
|
47
|
+
| **Claude Code native controls** | Usage status, model/effort choice, context inspection, clear/compact lifecycle | **P0 — build into core**. Prefer native surfaces over third-party credential scraping; discover available models dynamically |
|
|
48
|
+
| **Codex native app-server + profiles** | Structured rate-limit telemetry; model, reasoning effort, verbosity, tool-output and MCP/context controls | **P0 — build into core**. `account/rateLimits/read` is the preferred live quota source when available |
|
|
49
|
+
| [cclimits](https://github.com/cruzanstx/cclimits) | Live/local quota companion across coding tools | **P0 — optional observational companion for Claude**. Token Harness uses only cacheless Claude JSON mode, never reads its credentials, and labels the result `reported` or `cached`. Codex keeps its native app-server reader |
|
|
50
|
+
| [ccusage](https://github.com/ccusage/ccusage) | Local historical token/session/cost analytics for Claude Code and Codex | **P0 — high-priority read-only companion**. Useful for history and baselines; not a substitute for live subscription-limit telemetry |
|
|
51
|
+
| [RTK](https://github.com/rtk-ai/rtk) | Shell-command rewriting and command-output reduction | **P1 — active, keep**. Reduces context that would otherwise be resent on later turns |
|
|
52
|
+
| [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic reducers, harness adapters, skills, pipes, and MCP reduction | **P1 — active, keep**. Continue only on proven non-overlapping surfaces |
|
|
53
|
+
| [Lazy MCP](https://gitlab.com/gitlab-org/ai/lazy-mcp) | Load MCP tool schemas only when needed | **P1 — high priority**. Directly attacks schema/context overhead; benchmark against native Codex tool deferral before installing another owner |
|
|
54
|
+
| [Context Mode](https://github.com/mksglu/context-mode) | Keep raw tool/MCP results outside model context | **P2 — alternative broad context owner**. Evaluate against Headroom, not alongside overlapping reducers by default |
|
|
55
|
+
| [Headroom](https://github.com/headroomlabs-ai/headroom) | Compress tool, MCP, file, and RAG payloads | **P2 — alternative broad context owner**. Admit only after pair-specific quality and attribution fixtures |
|
|
56
|
+
| [Dejavu](https://github.com/Salnika/dejavu) | Emit only the delta when command output repeats | **P2 — useful for test/rerun loops** after the normal output-reduction path is measured |
|
|
57
|
+
| [repowise](https://github.com/repowise-dev/repowise) | Retrieve task-specific repository context | **P2 — conditional**. Its MCP overhead must be lower than the repository context it avoids |
|
|
58
|
+
| LiteLLM, Claude Code Router, RouteLLM, LLMRouter, vLLM Semantic Router | Route requests across models/providers | **P3 — overflow/API routing only**. Useful after included quota is exhausted or for explicit external-provider policy; not a core subscription-quota optimizer |
|
|
59
|
+
| [LLMLingua](https://github.com/microsoft/LLMLingua) | Generic prompt compression | **Research only** until a harness-aware lifecycle proves must-keep recall and quality |
|
|
60
|
+
| [Caveman](https://github.com/JuliusBrussee/caveman) | Reduce visible model-output verbosity | **De-prioritized**. Use native verbosity/effort controls first and measure their effect before adding another instruction owner |
|
|
61
|
+
|
|
62
|
+
Candidate status still means that Token Harness neither installs nor configures the tool until its
|
|
63
|
+
installation, conflicts, rollback behavior, verification, quality gates, and metrics attribution are
|
|
64
|
+
implemented and tested. The intake evidence lives in
|
|
47
65
|
[docs/provider-landscape.md](docs/provider-landscape.md).
|
|
48
66
|
|
|
67
|
+
### Why generic routers moved down
|
|
68
|
+
|
|
69
|
+
The original roadmap treated model routers as high priority. For the new objective that is backwards:
|
|
70
|
+
a router can reduce API cost by sending work to another provider while consuming **none of the
|
|
71
|
+
included Claude Code or Codex allowance**, or it can bypass subscription authentication entirely.
|
|
72
|
+
That can be valuable as explicit overflow, but it does not demonstrate better use of the allowance
|
|
73
|
+
the user already paid for.
|
|
74
|
+
|
|
75
|
+
Token Harness should first exploit the harness-native choices that are actually inside the plan:
|
|
76
|
+
model tier, reasoning effort, verbosity, session lifecycle, instruction size, MCP exposure, and tool
|
|
77
|
+
output. External routing becomes an opt-in second budget, never an invisible shortcut.
|
|
78
|
+
|
|
49
79
|
## Quick start
|
|
50
80
|
|
|
51
81
|
Install the CLI:
|
|
@@ -61,15 +91,24 @@ Then run the complete workflow from the project in which you use your coding age
|
|
|
61
91
|
# 1. Inspect the machine. This does not change agent configuration.
|
|
62
92
|
token-harness doctor
|
|
63
93
|
|
|
64
|
-
# 2.
|
|
94
|
+
# 2. Inspect subscription headroom and avoidable context. Still read-only.
|
|
95
|
+
token-harness budget
|
|
96
|
+
token-harness context
|
|
97
|
+
token-harness history --since 7d
|
|
98
|
+
token-harness optimize
|
|
99
|
+
|
|
100
|
+
# Optional: Claude live quota can use an already-installed cclimits build that
|
|
101
|
+
# supports --no-cache-write. Token Harness never installs it automatically.
|
|
102
|
+
|
|
103
|
+
# 3. Preview every proposed configuration change.
|
|
65
104
|
token-harness plan
|
|
66
105
|
|
|
67
|
-
#
|
|
106
|
+
# 4. Apply the reviewed plan. This is the first configuration-changing step.
|
|
68
107
|
token-harness apply --yes
|
|
69
108
|
|
|
70
|
-
#
|
|
109
|
+
# 5. Restart the coding agent, then run a normal shell command through it.
|
|
71
110
|
|
|
72
|
-
#
|
|
111
|
+
# 6. Check configuration, real interception evidence, and savings.
|
|
73
112
|
token-harness status
|
|
74
113
|
token-harness verify
|
|
75
114
|
token-harness metrics --since 7d
|
|
@@ -90,26 +129,41 @@ HarnessTrim, or a coding agent.
|
|
|
90
129
|
### Managed compatibility rows
|
|
91
130
|
|
|
92
131
|
Token Harness changes a harness configuration only when a reviewed compatibility row covers the
|
|
93
|
-
exact provider version, harness version, platform, and configuration schema.
|
|
132
|
+
exact provider version, harness version, platform, and configuration schema. Three rows ship, and each
|
|
94
133
|
names the recording it stands on:
|
|
95
134
|
|
|
96
135
|
| Provider | Harness | Platform | Tested versions | Tier |
|
|
97
136
|
| --- | --- | --- | --- | --- |
|
|
98
137
|
| RTK | Claude Code | Windows | rtk 0.44.0, Claude Code 2.1.220 | `canary` |
|
|
99
138
|
| HarnessTrim | Claude Code | Windows | harnesstrim 0.1.0, Claude Code 2.1.220 | `config-only` |
|
|
139
|
+
| HarnessTrim | Codex | Windows | harnesstrim 0.1.0, Codex 0.146.0 | `config-only` |
|
|
100
140
|
|
|
101
141
|
Everything else is refused, and that is the design rather than a gap: `doctor` detects and reports on
|
|
102
142
|
every supported platform, and only the *mutation* is narrower. An uncovered combination exits 9 and
|
|
103
143
|
the diagnostic names what is missing — the reviewed fixture, or the nearest row it does have.
|
|
104
144
|
|
|
145
|
+
A live Linux check on 2026-09-02 with Codex 0.152.1 confirmed that boundary: the current row covers
|
|
146
|
+
Codex 0.146.0 on Windows only, so `plan` and `apply --yes` refused with exit 9 and wrote nothing.
|
|
147
|
+
That combination can still be observed and benchmarked with a provider installed through its own
|
|
148
|
+
installer; it is not promoted to managed mutation until a real compatibility fixture is recorded.
|
|
149
|
+
|
|
105
150
|
What is not covered today, and why:
|
|
106
151
|
|
|
107
152
|
- **macOS and Linux.** No row on either. The recordings a row needs are states of a real machine, and
|
|
108
153
|
a fixture cannot be written from a machine nobody ran. On those platforms `plan` and `apply` refuse;
|
|
109
154
|
install the provider with its own installer and Token Harness will detect, verify, and measure it.
|
|
110
|
-
- **
|
|
111
|
-
|
|
112
|
-
|
|
155
|
+
- **OpenCode, and permanently rather than pending.** Both providers are detected, adopted, verified
|
|
156
|
+
and measured there, and neither is written. RTK reaches OpenCode through a plugin module its own
|
|
157
|
+
installer places globally, which this build has no action for. HarnessTrim's OpenCode installer
|
|
158
|
+
writes a plugin wrapper *and runs an npm install*, so a containment boundary covering what it wrote
|
|
159
|
+
would hold a `node_modules` tree — and that is not a decision deferred for want of a fixture. A
|
|
160
|
+
dependency tree is not configuration, so it cannot be a reviewed write set; snapshotting it on
|
|
161
|
+
every apply to keep the rollback honest would be slow and would be restoring upstream's install
|
|
162
|
+
rather than our change; and excluding it would leave a transaction claiming a reversibility it does
|
|
163
|
+
not have. So the assignment is not producible, and RFC 0003 is explicit about what that means: a
|
|
164
|
+
capability the provider has but cannot be asked for is not an assignable capability. OpenCode stays
|
|
165
|
+
adoption-only by decision.
|
|
166
|
+
- **RTK on Codex.** Not managed, and no row: RTK writes a Claude-shaped hook list and nothing else.
|
|
113
167
|
- **A newer Claude Code.** The range is a single observed version. `2.1.221` reads `unknown-newer` and
|
|
114
168
|
refuses rather than assuming it behaves like `2.1.220`.
|
|
115
169
|
|
|
@@ -122,21 +176,35 @@ There are three separate layers. Installing one does not automatically provide t
|
|
|
122
176
|
|
|
123
177
|
| Layer | Examples | Who installs it? |
|
|
124
178
|
| --- | --- | --- |
|
|
125
|
-
| Coding agent (harness) | Claude Code, Codex, OpenCode | You, using the agent's official installer |
|
|
179
|
+
| Coding agent (harness) | Claude Code, Codex, OpenCode, Hermes, Pi | You, using the agent's official installer |
|
|
126
180
|
| Token Harness | `token-harness` | You, from npm or this repository |
|
|
127
181
|
| Optimization provider | RTK, HarnessTrim | Both can be installed by Token Harness where a compatibility row covers the combination; otherwise install them with their own installers and Token Harness detects and measures them |
|
|
128
182
|
|
|
129
|
-
Token Harness does not install Claude Code, Codex, or
|
|
183
|
+
Token Harness does not install Claude Code, Codex, OpenCode, Hermes, or Pi. Install and run at least one of
|
|
130
184
|
them first so that `token-harness doctor` can detect it.
|
|
131
185
|
|
|
132
|
-
| Provider | Claude Code | Codex | OpenCode | Installed by Token Harness |
|
|
133
|
-
| --- | --- | --- | --- | --- |
|
|
134
|
-
| RTK | Configure, verify, and measure | Not managed | Detect, adopt, verify, and measure | **Yes**, for the supported Claude Code path |
|
|
135
|
-
| HarnessTrim | Claude skills only; no reducer hook or reduce-pipe instruction | Detect, adopt, verify, and measure | Detect, adopt, verify, and measure | **Yes**, on a covered row — see above |
|
|
186
|
+
| Provider | Claude Code | Codex | OpenCode | Hermes | Pi | Installed by Token Harness |
|
|
187
|
+
| --- | --- | --- | --- | --- | --- | --- |
|
|
188
|
+
| RTK | Configure, verify, and measure | Not managed | Detect, adopt, verify, and measure | Not managed | Not managed | **Yes**, for the supported Claude Code path |
|
|
189
|
+
| HarnessTrim | Claude skills only; no reducer hook or reduce-pipe instruction | Detect, adopt, verify, and measure | Detect, adopt, verify, and measure | Detect, verify, and measure | Detect, verify, and measure | **Yes**, on a covered row — see above |
|
|
136
190
|
|
|
137
191
|
"Not managed" does not mean the upstream tool cannot support that agent. It means this release
|
|
138
192
|
does not claim ownership of that integration and will not modify it.
|
|
139
193
|
|
|
194
|
+
Hermes is read-only in both directions: the adapter finds the HarnessTrim plugin, reads whether it
|
|
195
|
+
is enabled, and imports the telemetry it writes to `~/.hermes/harnesstrim-metrics.jsonl`, but nothing
|
|
196
|
+
here enables the plugin or restarts the gateway. Enabling it is
|
|
197
|
+
`hermes plugins enable harnesstrim`, and that stays your command to run. No compatibility row ships
|
|
198
|
+
for Hermes because a row is the precondition for a *mutation*, and none is proposed.
|
|
199
|
+
|
|
200
|
+
Pi is read-only in both directions too: the adapter finds the HarnessTrim extension module in the
|
|
201
|
+
directories Pi auto-loads (`~/.pi/agent/extensions/` and `<project>/.pi/extensions/`) and verifies
|
|
202
|
+
the configuration, but nothing here installs it, and nothing here can say which mode it runs in —
|
|
203
|
+
the extension defaults to `dryrun` and only `HARNESSTRIM_MODE=active` in Pi's environment makes it
|
|
204
|
+
reduce. Installing it is `harnesstrim install pi --apply`, and that stays your command to run. No
|
|
205
|
+
compatibility row ships for Pi because a row is the precondition for a *mutation*, and none is
|
|
206
|
+
proposed.
|
|
207
|
+
|
|
140
208
|
RTK on OpenCode is detected and verified, not written: `rtk init -g --opencode` installs a plugin
|
|
141
209
|
module at `~/.config/opencode/plugins/rtk.ts`, and Token Harness reads that file rather than
|
|
142
210
|
producing it. Note that the plugin is inert under OpenCode Desktop — see
|
package/package.json
CHANGED
package/sbom.json
CHANGED
|
@@ -1,13 +1,14 @@
|
|
|
1
1
|
{
|
|
2
2
|
"bomFormat": "CycloneDX",
|
|
3
3
|
"specVersion": "1.5",
|
|
4
|
+
"serialNumber": "urn:uuid:f4cf234e-2ad8-b2a3-a085-7056eb8a17c4",
|
|
4
5
|
"version": 1,
|
|
5
6
|
"metadata": {
|
|
6
7
|
"component": {
|
|
7
8
|
"type": "application",
|
|
8
9
|
"bom-ref": "token-harness",
|
|
9
10
|
"name": "token-harness",
|
|
10
|
-
"version": "0.1.
|
|
11
|
+
"version": "0.1.5",
|
|
11
12
|
"description": "One control plane for token-efficient coding agents.",
|
|
12
13
|
"licenses": [
|
|
13
14
|
{
|
|
@@ -19,7 +20,7 @@
|
|
|
19
20
|
"hashes": [
|
|
20
21
|
{
|
|
21
22
|
"alg": "SHA-256",
|
|
22
|
-
"content": "
|
|
23
|
+
"content": "f4cf234e2ad8b2a3a0857056eb8a17c494730c3d5c11e702edc98b270f18d203"
|
|
23
24
|
}
|
|
24
25
|
]
|
|
25
26
|
},
|