token-harness 0.1.2 → 0.1.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (4) hide show
  1. package/README.md +118 -50
  2. package/package.json +1 -1
  3. package/sbom.json +3 -2
  4. package/token-harness.mjs +7059 -1062
package/README.md CHANGED
@@ -1,51 +1,81 @@
1
1
  # Token Harness
2
2
 
3
- Token Harness has one objective: **reduce the tokens consumed by coding agents without hiding
4
- useful information or overstating the result**.
3
+ Token Harness has one objective: **maximize the useful coding work you can get from Claude Code
4
+ and Codex usage limits, without hiding quality regressions or pretending that opaque subscription
5
+ quota is exactly equivalent to a token count**.
5
6
 
6
- Coding sessions repeatedly send test logs, command output, repository context, MCP schemas,
7
- tool results, and conversation history back to the model. Specialized tools can reduce each of
8
- those sources, but installing them independently creates a second problem: overlapping hooks,
9
- double reduction, incompatible configurations, and savings counted more than once.
7
+ The scarce resource is no longer just model context. Subscription users are constrained by rolling
8
+ usage windows, weekly limits, model-dependent burn, long-session context growth, MCP/tool schema
9
+ overhead, noisy tool output, and repeated work after poor model or effort choices. Token Harness is
10
+ evolving from a token-reduction control plane into a **quota-aware efficiency layer** for coding
11
+ harnesses.
10
12
 
11
- Token Harness is the control plane for that optimization stack. It finds the coding agents and
12
- token-saving tools on the machine, selects a compatible owner for each interception point, shows
13
- every proposed change before applying it, verifies whether the integration is genuinely being
14
- used, and reports how many tokens or characters were saved.
13
+ The product therefore optimizes two related budgets:
15
14
 
16
- The reduction still happens inside specialized providers such as RTK and HarnessTrim. Token
17
- Harness makes those providers safe to combine, observable, reversible, and comparable.
15
+ 1. **subscription headroom** — observe the usage windows the harness itself exposes, pace work
16
+ against reset time, and choose native model/effort settings that fit the task and remaining
17
+ budget;
18
+ 2. **context sent to the model** — keep instructions, MCP schemas, conversation history, repository
19
+ context, and tool output as small as possible without removing information needed to finish the
20
+ task correctly.
21
+
22
+ RTK and HarnessTrim remain useful providers, but reducers are now one layer of the system rather
23
+ than the product definition. Native harness controls come first because choosing the right model,
24
+ effort, context shape, and enabled tools can avoid an expensive turn entirely.
25
+
26
+ ## What Token Harness should optimize
27
+
28
+ | Layer | Target behavior |
29
+ | --- | --- |
30
+ | Usage-window observability | Read five-hour, weekly, model-specific, or credit-backed limits only from surfaces the harness can actually prove; show reset time, headroom, and burn rate |
31
+ | Native model policy | Prefer economical models/effort for routine work, escalate only for tasks whose expected quality benefit justifies the extra quota |
32
+ | Session hygiene | Detect task boundaries, oversized conversations, and stale context; recommend clear/compact/new-session actions before history dominates each turn |
33
+ | Instruction budget | Keep `CLAUDE.md` / `AGENTS.md` concise and hierarchical instead of injecting one large global instruction file everywhere |
34
+ | MCP/tool budget | Disable irrelevant MCP servers, defer schemas where the harness supports it, and expose only the tools required by the current task |
35
+ | Tool-output budget | Reduce logs, diffs, test output, and repeated command results before they re-enter model context |
36
+ | Cross-harness scheduling | When Claude Code and Codex have independent headroom, recommend the harness that can do the task with the best expected quality per remaining quota |
37
+ | Measurement | Correlate quota delta, model/effort, context size, tool traffic, and task outcome; never convert local token savings into an unproven subscription-quota claim |
18
38
 
19
39
  ## Optimization ecosystem
20
40
 
21
- The long-term goal is to coordinate token savings across the whole coding-agent pipeline. Only
22
- tools marked **active** are integrated in this release; every other row is a candidate and is
23
- neither installed nor configured by Token Harness.
41
+ The priority order below is deliberately different from the original Token Harness roadmap. A tool
42
+ is high priority only when it helps **included Claude Code/Codex capacity**, not merely API cost,
43
+ provider routing, or benchmark token counts.
24
44
 
25
- | Tool | Optimization layer | Token Harness status |
45
+ | Tool or surface | Optimization layer | Token Harness direction |
26
46
  | --- | --- | --- |
27
- | [RTK](https://github.com/rtk-ai/rtk) | Shell-command rewriting and command-output reduction | **Active — integrated** |
28
- | [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic reducers, harness adapters, skills, pipes, and MCP reduction | **Active — integrated** |
29
- | [Dejavu](https://github.com/Salnika/dejavu) | Emit only the delta when command output repeats | Not active — candidate |
30
- | [Lazy MCP](https://github.com/voicetreelab/lazy-mcp) | Load MCP tool schemas only when needed | Not active — candidate |
31
- | [repowise](https://github.com/repowise-dev/repowise) | Retrieve task-specific repository context | Not active — candidate |
32
- | [LiteLLM](https://github.com/BerriAI/litellm) | Model routing, fallbacks, budgets, and usage telemetry | Not active — candidate |
33
- | [RouteLLM](https://github.com/lm-sys/RouteLLM) | Route simpler requests to less expensive models | Not active — candidate |
34
- | [vLLM Semantic Router](https://github.com/vllm-project/semantic-router) | Route by task, complexity, tools, and deployment locality | Not active — candidate |
35
- | [Claude Code Router](https://github.com/musistudio/claude-code-router) | Route coding-agent requests across models and providers with effort-based rules and fallback chains | Not active — candidate · high priority |
36
- | [LLMRouter](https://github.com/ulab-uiuc/LLMRouter) | Select the model by task complexity, cost, and quality across routing strategies | Not active — candidate · high priority |
37
- | [Headroom](https://github.com/headroomlabs-ai/headroom) | Compress tool, MCP, file, and RAG payloads | Not active — candidate |
38
- | [Context Mode](https://github.com/mksglu/context-mode) | Keep raw tool results outside model context | Not active — candidate |
39
- | [LLMLingua](https://github.com/microsoft/LLMLingua) | Compress long prompts and context | Not active — candidate |
40
- | [Caveman](https://github.com/JuliusBrussee/caveman) | Reduce visible model-output verbosity | Not active — candidate |
41
-
42
- Candidate status means only that the project has identified a useful optimization layer. A tool
43
- becomes active only after its installation, conflicts, rollback behavior, verification, and
44
- metrics attribution have been implemented and tested. Token Harness never installs a candidate
45
- merely because it is present on the machine. Rows marked `high priority` are the next intended
46
- intake; the admission gates each one carries are recorded in
47
+ | **Claude Code native controls** | Usage status, model/effort choice, context inspection, clear/compact lifecycle | **P0 — build into core**. Prefer native surfaces over third-party credential scraping; discover available models dynamically |
48
+ | **Codex native app-server + profiles** | Structured rate-limit telemetry; model, reasoning effort, verbosity, tool-output and MCP/context controls | **P0 — build into core**. `account/rateLimits/read` is the preferred live quota source when available |
49
+ | [cclimits](https://github.com/cruzanstx/cclimits) | Live/local quota companion across coding tools | **P0 — optional observational companion for Claude**. Token Harness uses only cacheless Claude JSON mode, never reads its credentials, and labels the result `reported` or `cached`. Codex keeps its native app-server reader |
50
+ | [ccusage](https://github.com/ccusage/ccusage) | Local historical token/session/cost analytics for Claude Code and Codex | **P0 — high-priority read-only companion**. Useful for history and baselines; not a substitute for live subscription-limit telemetry |
51
+ | [RTK](https://github.com/rtk-ai/rtk) | Shell-command rewriting and command-output reduction | **P1 — active, keep**. Reduces context that would otherwise be resent on later turns |
52
+ | [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic reducers, harness adapters, skills, pipes, and MCP reduction | **P1 — active, keep**. Continue only on proven non-overlapping surfaces |
53
+ | [Lazy MCP](https://gitlab.com/gitlab-org/ai/lazy-mcp) | Load MCP tool schemas only when needed | **P1 — high priority**. Directly attacks schema/context overhead; benchmark against native Codex tool deferral before installing another owner |
54
+ | [Context Mode](https://github.com/mksglu/context-mode) | Keep raw tool/MCP results outside model context | **P2 — alternative broad context owner**. Evaluate against Headroom, not alongside overlapping reducers by default |
55
+ | [Headroom](https://github.com/headroomlabs-ai/headroom) | Compress tool, MCP, file, and RAG payloads | **P2 — alternative broad context owner**. Admit only after pair-specific quality and attribution fixtures |
56
+ | [Dejavu](https://github.com/Salnika/dejavu) | Emit only the delta when command output repeats | **P2 — useful for test/rerun loops** after the normal output-reduction path is measured |
57
+ | [repowise](https://github.com/repowise-dev/repowise) | Retrieve task-specific repository context | **P2 — conditional**. Its MCP overhead must be lower than the repository context it avoids |
58
+ | LiteLLM, Claude Code Router, RouteLLM, LLMRouter, vLLM Semantic Router | Route requests across models/providers | **P3 — overflow/API routing only**. Useful after included quota is exhausted or for explicit external-provider policy; not a core subscription-quota optimizer |
59
+ | [LLMLingua](https://github.com/microsoft/LLMLingua) | Generic prompt compression | **Research only** until a harness-aware lifecycle proves must-keep recall and quality |
60
+ | [Caveman](https://github.com/JuliusBrussee/caveman) | Reduce visible model-output verbosity | **De-prioritized**. Use native verbosity/effort controls first and measure their effect before adding another instruction owner |
61
+
62
+ Candidate status still means that Token Harness neither installs nor configures the tool until its
63
+ installation, conflicts, rollback behavior, verification, quality gates, and metrics attribution are
64
+ implemented and tested. The intake evidence lives in
47
65
  [docs/provider-landscape.md](docs/provider-landscape.md).
48
66
 
67
+ ### Why generic routers moved down
68
+
69
+ The original roadmap treated model routers as high priority. For the new objective that is backwards:
70
+ a router can reduce API cost by sending work to another provider while consuming **none of the
71
+ included Claude Code or Codex allowance**, or it can bypass subscription authentication entirely.
72
+ That can be valuable as explicit overflow, but it does not demonstrate better use of the allowance
73
+ the user already paid for.
74
+
75
+ Token Harness should first exploit the harness-native choices that are actually inside the plan:
76
+ model tier, reasoning effort, verbosity, session lifecycle, instruction size, MCP exposure, and tool
77
+ output. External routing becomes an opt-in second budget, never an invisible shortcut.
78
+
49
79
  ## Quick start
50
80
 
51
81
  Install the CLI:
@@ -61,15 +91,24 @@ Then run the complete workflow from the project in which you use your coding age
61
91
  # 1. Inspect the machine. This does not change agent configuration.
62
92
  token-harness doctor
63
93
 
64
- # 2. Preview every proposed change.
94
+ # 2. Inspect subscription headroom and avoidable context. Still read-only.
95
+ token-harness budget
96
+ token-harness context
97
+ token-harness history --since 7d
98
+ token-harness optimize
99
+
100
+ # Optional: Claude live quota can use an already-installed cclimits build that
101
+ # supports --no-cache-write. Token Harness never installs it automatically.
102
+
103
+ # 3. Preview every proposed configuration change.
65
104
  token-harness plan
66
105
 
67
- # 3. Apply the reviewed plan. This is the first configuration-changing step.
106
+ # 4. Apply the reviewed plan. This is the first configuration-changing step.
68
107
  token-harness apply --yes
69
108
 
70
- # 4. Restart the coding agent, then run a normal shell command through it.
109
+ # 5. Restart the coding agent, then run a normal shell command through it.
71
110
 
72
- # 5. Check configuration, real interception evidence, and savings.
111
+ # 6. Check configuration, real interception evidence, and savings.
73
112
  token-harness status
74
113
  token-harness verify
75
114
  token-harness metrics --since 7d
@@ -90,26 +129,41 @@ HarnessTrim, or a coding agent.
90
129
  ### Managed compatibility rows
91
130
 
92
131
  Token Harness changes a harness configuration only when a reviewed compatibility row covers the
93
- exact provider version, harness version, platform, and configuration schema. Two rows ship, and each
132
+ exact provider version, harness version, platform, and configuration schema. Three rows ship, and each
94
133
  names the recording it stands on:
95
134
 
96
135
  | Provider | Harness | Platform | Tested versions | Tier |
97
136
  | --- | --- | --- | --- | --- |
98
137
  | RTK | Claude Code | Windows | rtk 0.44.0, Claude Code 2.1.220 | `canary` |
99
138
  | HarnessTrim | Claude Code | Windows | harnesstrim 0.1.0, Claude Code 2.1.220 | `config-only` |
139
+ | HarnessTrim | Codex | Windows | harnesstrim 0.1.0, Codex 0.146.0 | `config-only` |
100
140
 
101
141
  Everything else is refused, and that is the design rather than a gap: `doctor` detects and reports on
102
142
  every supported platform, and only the *mutation* is narrower. An uncovered combination exits 9 and
103
143
  the diagnostic names what is missing — the reviewed fixture, or the nearest row it does have.
104
144
 
145
+ A live Linux check on 2026-09-02 with Codex 0.152.1 confirmed that boundary: the current row covers
146
+ Codex 0.146.0 on Windows only, so `plan` and `apply --yes` refused with exit 9 and wrote nothing.
147
+ That combination can still be observed and benchmarked with a provider installed through its own
148
+ installer; it is not promoted to managed mutation until a real compatibility fixture is recorded.
149
+
105
150
  What is not covered today, and why:
106
151
 
107
152
  - **macOS and Linux.** No row on either. The recordings a row needs are states of a real machine, and
108
153
  a fixture cannot be written from a machine nobody ran. On those platforms `plan` and `apply` refuse;
109
154
  install the provider with its own installer and Token Harness will detect, verify, and measure it.
110
- - **Codex and OpenCode.** Both providers are detected and adopted there, and neither is written:
111
- RTK's plan builder produces a Claude-shaped hook list, and HarnessTrim's reviewed write set covers
112
- Claude only. A row would admit a mutation that nothing proposes.
155
+ - **OpenCode, and permanently rather than pending.** Both providers are detected, adopted, verified
156
+ and measured there, and neither is written. RTK reaches OpenCode through a plugin module its own
157
+ installer places globally, which this build has no action for. HarnessTrim's OpenCode installer
158
+ writes a plugin wrapper *and runs an npm install*, so a containment boundary covering what it wrote
159
+ would hold a `node_modules` tree — and that is not a decision deferred for want of a fixture. A
160
+ dependency tree is not configuration, so it cannot be a reviewed write set; snapshotting it on
161
+ every apply to keep the rollback honest would be slow and would be restoring upstream's install
162
+ rather than our change; and excluding it would leave a transaction claiming a reversibility it does
163
+ not have. So the assignment is not producible, and RFC 0003 is explicit about what that means: a
164
+ capability the provider has but cannot be asked for is not an assignable capability. OpenCode stays
165
+ adoption-only by decision.
166
+ - **RTK on Codex.** Not managed, and no row: RTK writes a Claude-shaped hook list and nothing else.
113
167
  - **A newer Claude Code.** The range is a single observed version. `2.1.221` reads `unknown-newer` and
114
168
  refuses rather than assuming it behaves like `2.1.220`.
115
169
 
@@ -122,21 +176,35 @@ There are three separate layers. Installing one does not automatically provide t
122
176
 
123
177
  | Layer | Examples | Who installs it? |
124
178
  | --- | --- | --- |
125
- | Coding agent (harness) | Claude Code, Codex, OpenCode | You, using the agent's official installer |
179
+ | Coding agent (harness) | Claude Code, Codex, OpenCode, Hermes, Pi | You, using the agent's official installer |
126
180
  | Token Harness | `token-harness` | You, from npm or this repository |
127
181
  | Optimization provider | RTK, HarnessTrim | Both can be installed by Token Harness where a compatibility row covers the combination; otherwise install them with their own installers and Token Harness detects and measures them |
128
182
 
129
- Token Harness does not install Claude Code, Codex, or OpenCode. Install and run at least one of
183
+ Token Harness does not install Claude Code, Codex, OpenCode, Hermes, or Pi. Install and run at least one of
130
184
  them first so that `token-harness doctor` can detect it.
131
185
 
132
- | Provider | Claude Code | Codex | OpenCode | Installed by Token Harness |
133
- | --- | --- | --- | --- | --- |
134
- | RTK | Configure, verify, and measure | Not managed | Detect, adopt, verify, and measure | **Yes**, for the supported Claude Code path |
135
- | HarnessTrim | Claude skills only; no reducer hook or reduce-pipe instruction | Detect, adopt, verify, and measure | Detect, adopt, verify, and measure | **Yes**, on a covered row — see above |
186
+ | Provider | Claude Code | Codex | OpenCode | Hermes | Pi | Installed by Token Harness |
187
+ | --- | --- | --- | --- | --- | --- | --- |
188
+ | RTK | Configure, verify, and measure | Not managed | Detect, adopt, verify, and measure | Not managed | Not managed | **Yes**, for the supported Claude Code path |
189
+ | HarnessTrim | Claude skills only; no reducer hook or reduce-pipe instruction | Detect, adopt, verify, and measure | Detect, adopt, verify, and measure | Detect, verify, and measure | Detect, verify, and measure | **Yes**, on a covered row — see above |
136
190
 
137
191
  "Not managed" does not mean the upstream tool cannot support that agent. It means this release
138
192
  does not claim ownership of that integration and will not modify it.
139
193
 
194
+ Hermes is read-only in both directions: the adapter finds the HarnessTrim plugin, reads whether it
195
+ is enabled, and imports the telemetry it writes to `~/.hermes/harnesstrim-metrics.jsonl`, but nothing
196
+ here enables the plugin or restarts the gateway. Enabling it is
197
+ `hermes plugins enable harnesstrim`, and that stays your command to run. No compatibility row ships
198
+ for Hermes because a row is the precondition for a *mutation*, and none is proposed.
199
+
200
+ Pi is read-only in both directions too: the adapter finds the HarnessTrim extension module in the
201
+ directories Pi auto-loads (`~/.pi/agent/extensions/` and `<project>/.pi/extensions/`) and verifies
202
+ the configuration, but nothing here installs it, and nothing here can say which mode it runs in —
203
+ the extension defaults to `dryrun` and only `HARNESSTRIM_MODE=active` in Pi's environment makes it
204
+ reduce. Installing it is `harnesstrim install pi --apply`, and that stays your command to run. No
205
+ compatibility row ships for Pi because a row is the precondition for a *mutation*, and none is
206
+ proposed.
207
+
140
208
  RTK on OpenCode is detected and verified, not written: `rtk init -g --opencode` installs a plugin
141
209
  module at `~/.config/opencode/plugins/rtk.ts`, and Token Harness reads that file rather than
142
210
  producing it. Note that the plugin is inert under OpenCode Desktop — see
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "token-harness",
3
- "version": "0.1.2",
3
+ "version": "0.1.5",
4
4
  "description": "One control plane for token-efficient coding agents.",
5
5
  "license": "Apache-2.0",
6
6
  "type": "module",
package/sbom.json CHANGED
@@ -1,13 +1,14 @@
1
1
  {
2
2
  "bomFormat": "CycloneDX",
3
3
  "specVersion": "1.5",
4
+ "serialNumber": "urn:uuid:f4cf234e-2ad8-b2a3-a085-7056eb8a17c4",
4
5
  "version": 1,
5
6
  "metadata": {
6
7
  "component": {
7
8
  "type": "application",
8
9
  "bom-ref": "token-harness",
9
10
  "name": "token-harness",
10
- "version": "0.1.2",
11
+ "version": "0.1.5",
11
12
  "description": "One control plane for token-efficient coding agents.",
12
13
  "licenses": [
13
14
  {
@@ -19,7 +20,7 @@
19
20
  "hashes": [
20
21
  {
21
22
  "alg": "SHA-256",
22
- "content": "8e2f96efbb15fd8aec72a98ae3b62d128265abd779e9f911c8c0d4bf024adfb2"
23
+ "content": "f4cf234e2ad8b2a3a0857056eb8a17c494730c3d5c11e702edc98b270f18d203"
23
24
  }
24
25
  ]
25
26
  },