token-harness 0.1.3 → 0.1.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,51 +1,81 @@
1
1
  # Token Harness
2
2
 
3
- Token Harness has one objective: **reduce the tokens consumed by coding agents without hiding
4
- useful information or overstating the result**.
3
+ Token Harness has one objective: **maximize the useful coding work you can get from Claude Code
4
+ and Codex usage limits, without hiding quality regressions or pretending that opaque subscription
5
+ quota is exactly equivalent to a token count**.
5
6
 
6
- Coding sessions repeatedly send test logs, command output, repository context, MCP schemas,
7
- tool results, and conversation history back to the model. Specialized tools can reduce each of
8
- those sources, but installing them independently creates a second problem: overlapping hooks,
9
- double reduction, incompatible configurations, and savings counted more than once.
7
+ The scarce resource is no longer just model context. Subscription users are constrained by rolling
8
+ usage windows, weekly limits, model-dependent burn, long-session context growth, MCP/tool schema
9
+ overhead, noisy tool output, and repeated work after poor model or effort choices. Token Harness is
10
+ evolving from a token-reduction control plane into a **quota-aware efficiency layer** for coding
11
+ harnesses.
10
12
 
11
- Token Harness is the control plane for that optimization stack. It finds the coding agents and
12
- token-saving tools on the machine, selects a compatible owner for each interception point, shows
13
- every proposed change before applying it, verifies whether the integration is genuinely being
14
- used, and reports how many tokens or characters were saved.
13
+ The product therefore optimizes two related budgets:
15
14
 
16
- The reduction still happens inside specialized providers such as RTK and HarnessTrim. Token
17
- Harness makes those providers safe to combine, observable, reversible, and comparable.
15
+ 1. **subscription headroom** — observe the usage windows the harness itself exposes, pace work
16
+ against reset time, and choose native model/effort settings that fit the task and remaining
17
+ budget;
18
+ 2. **context sent to the model** — keep instructions, MCP schemas, conversation history, repository
19
+ context, and tool output as small as possible without removing information needed to finish the
20
+ task correctly.
21
+
22
+ RTK and HarnessTrim remain useful providers, but reducers are now one layer of the system rather
23
+ than the product definition. Native harness controls come first because choosing the right model,
24
+ effort, context shape, and enabled tools can avoid an expensive turn entirely.
25
+
26
+ ## What Token Harness should optimize
27
+
28
+ | Layer | Target behavior |
29
+ | --- | --- |
30
+ | Usage-window observability | Read five-hour, weekly, model-specific, or credit-backed limits only from surfaces the harness can actually prove; show reset time, headroom, and burn rate |
31
+ | Native model policy | Prefer economical models/effort for routine work, escalate only for tasks whose expected quality benefit justifies the extra quota |
32
+ | Session hygiene | Detect task boundaries, oversized conversations, and stale context; recommend clear/compact/new-session actions before history dominates each turn |
33
+ | Instruction budget | Keep `CLAUDE.md` / `AGENTS.md` concise and hierarchical instead of injecting one large global instruction file everywhere |
34
+ | MCP/tool budget | Disable irrelevant MCP servers, defer schemas where the harness supports it, and expose only the tools required by the current task |
35
+ | Tool-output budget | Reduce logs, diffs, test output, and repeated command results before they re-enter model context |
36
+ | Cross-harness scheduling | When Claude Code and Codex have independent headroom, recommend the harness that can do the task with the best expected quality per remaining quota |
37
+ | Measurement | Correlate quota delta, model/effort, context size, tool traffic, and task outcome; never convert local token savings into an unproven subscription-quota claim |
18
38
 
19
39
  ## Optimization ecosystem
20
40
 
21
- The long-term goal is to coordinate token savings across the whole coding-agent pipeline. Only
22
- tools marked **active** are integrated in this release; every other row is a candidate and is
23
- neither installed nor configured by Token Harness.
41
+ The priority order below is deliberately different from the original Token Harness roadmap. A tool
42
+ is high priority only when it helps **included Claude Code/Codex capacity**, not merely API cost,
43
+ provider routing, or benchmark token counts.
24
44
 
25
- | Tool | Optimization layer | Token Harness status |
45
+ | Tool or surface | Optimization layer | Token Harness direction |
26
46
  | --- | --- | --- |
27
- | [RTK](https://github.com/rtk-ai/rtk) | Shell-command rewriting and command-output reduction | **Active — integrated** |
28
- | [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic reducers, harness adapters, skills, pipes, and MCP reduction | **Active — integrated** |
29
- | [Dejavu](https://github.com/Salnika/dejavu) | Emit only the delta when command output repeats | Not active — candidate |
30
- | [Lazy MCP](https://github.com/voicetreelab/lazy-mcp) | Load MCP tool schemas only when needed | Not active — candidate |
31
- | [repowise](https://github.com/repowise-dev/repowise) | Retrieve task-specific repository context | Not active — candidate |
32
- | [LiteLLM](https://github.com/BerriAI/litellm) | Model routing, fallbacks, budgets, and usage telemetry | Not active — candidate |
33
- | [RouteLLM](https://github.com/lm-sys/RouteLLM) | Route simpler requests to less expensive models | Not active — candidate |
34
- | [vLLM Semantic Router](https://github.com/vllm-project/semantic-router) | Route by task, complexity, tools, and deployment locality | Not active — candidate |
35
- | [Claude Code Router](https://github.com/musistudio/claude-code-router) | Route coding-agent requests across models and providers with effort-based rules and fallback chains | Not active — candidate · high priority |
36
- | [LLMRouter](https://github.com/ulab-uiuc/LLMRouter) | Select the model by task complexity, cost, and quality across routing strategies | Not active — candidate · high priority |
37
- | [Headroom](https://github.com/headroomlabs-ai/headroom) | Compress tool, MCP, file, and RAG payloads | Not active — candidate |
38
- | [Context Mode](https://github.com/mksglu/context-mode) | Keep raw tool results outside model context | Not active — candidate |
39
- | [LLMLingua](https://github.com/microsoft/LLMLingua) | Compress long prompts and context | Not active — candidate |
40
- | [Caveman](https://github.com/JuliusBrussee/caveman) | Reduce visible model-output verbosity | Not active — candidate |
41
-
42
- Candidate status means only that the project has identified a useful optimization layer. A tool
43
- becomes active only after its installation, conflicts, rollback behavior, verification, and
44
- metrics attribution have been implemented and tested. Token Harness never installs a candidate
45
- merely because it is present on the machine. Rows marked `high priority` are the next intended
46
- intake; the admission gates each one carries are recorded in
47
+ | **Claude Code native controls** | Usage status, model/effort choice, context inspection, clear/compact lifecycle | **P0 — build into core**. Prefer native surfaces over third-party credential scraping; discover available models dynamically |
48
+ | **Codex native app-server + profiles** | Structured rate-limit telemetry; model, reasoning effort, verbosity, tool-output and MCP/context controls | **P0 — build into core**. `account/rateLimits/read` is the preferred live quota source when available |
49
+ | [cclimits](https://github.com/cruzanstx/cclimits) | Live/local quota companion across coding tools | **P0 — optional observational companion for Claude**. Token Harness uses only cacheless Claude JSON mode, never reads its credentials, and labels the result `reported` or `cached`. Codex keeps its native app-server reader |
50
+ | [ccusage](https://github.com/ccusage/ccusage) | Local historical token/session/cost analytics for Claude Code and Codex | **P0 — high-priority read-only companion**. Useful for history and baselines; not a substitute for live subscription-limit telemetry |
51
+ | [RTK](https://github.com/rtk-ai/rtk) | Shell-command rewriting and command-output reduction | **P1 — active, keep**. Reduces context that would otherwise be resent on later turns |
52
+ | [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic reducers, harness adapters, skills, pipes, and MCP reduction | **P1 — active, keep**. Continue only on proven non-overlapping surfaces |
53
+ | [Lazy MCP](https://gitlab.com/gitlab-org/ai/lazy-mcp) | Load MCP tool schemas only when needed | **P1 — high priority**. Directly attacks schema/context overhead; benchmark against native Codex tool deferral before installing another owner |
54
+ | [Context Mode](https://github.com/mksglu/context-mode) | Keep raw tool/MCP results outside model context | **P2 — alternative broad context owner**. Evaluate against Headroom, not alongside overlapping reducers by default |
55
+ | [Headroom](https://github.com/headroomlabs-ai/headroom) | Compress tool, MCP, file, and RAG payloads | **P2 — alternative broad context owner**. Admit only after pair-specific quality and attribution fixtures |
56
+ | [Dejavu](https://github.com/Salnika/dejavu) | Emit only the delta when command output repeats | **P2 — useful for test/rerun loops** after the normal output-reduction path is measured |
57
+ | [repowise](https://github.com/repowise-dev/repowise) | Retrieve task-specific repository context | **P2 — conditional**. Its MCP overhead must be lower than the repository context it avoids |
58
+ | LiteLLM, Claude Code Router, RouteLLM, LLMRouter, vLLM Semantic Router | Route requests across models/providers | **P3 — overflow/API routing only**. Useful after included quota is exhausted or for explicit external-provider policy; not a core subscription-quota optimizer |
59
+ | [LLMLingua](https://github.com/microsoft/LLMLingua) | Generic prompt compression | **Research only** until a harness-aware lifecycle proves must-keep recall and quality |
60
+ | [Caveman](https://github.com/JuliusBrussee/caveman) | Reduce visible model-output verbosity | **De-prioritized**. Use native verbosity/effort controls first and measure their effect before adding another instruction owner |
61
+
62
+ Candidate status still means that Token Harness neither installs nor configures the tool until its
63
+ installation, conflicts, rollback behavior, verification, quality gates, and metrics attribution are
64
+ implemented and tested. The intake evidence lives in
47
65
  [docs/provider-landscape.md](docs/provider-landscape.md).
48
66
 
67
+ ### Why generic routers moved down
68
+
69
+ The original roadmap treated model routers as high priority. For the new objective that is backwards:
70
+ a router can reduce API cost by sending work to another provider while consuming **none of the
71
+ included Claude Code or Codex allowance**, or it can bypass subscription authentication entirely.
72
+ That can be valuable as explicit overflow, but it does not demonstrate better use of the allowance
73
+ the user already paid for.
74
+
75
+ Token Harness should first exploit the harness-native choices that are actually inside the plan:
76
+ model tier, reasoning effort, verbosity, session lifecycle, instruction size, MCP exposure, and tool
77
+ output. External routing becomes an opt-in second budget, never an invisible shortcut.
78
+
49
79
  ## Quick start
50
80
 
51
81
  Install the CLI:
@@ -61,15 +91,24 @@ Then run the complete workflow from the project in which you use your coding age
61
91
  # 1. Inspect the machine. This does not change agent configuration.
62
92
  token-harness doctor
63
93
 
64
- # 2. Preview every proposed change.
94
+ # 2. Inspect subscription headroom and avoidable context. Still read-only.
95
+ token-harness budget
96
+ token-harness context
97
+ token-harness history --since 7d
98
+ token-harness optimize
99
+
100
+ # Optional: Claude live quota can use an already-installed cclimits build that
101
+ # supports --no-cache-write. Token Harness never installs it automatically.
102
+
103
+ # 3. Preview every proposed configuration change.
65
104
  token-harness plan
66
105
 
67
- # 3. Apply the reviewed plan. This is the first configuration-changing step.
106
+ # 4. Apply the reviewed plan. This is the first configuration-changing step.
68
107
  token-harness apply --yes
69
108
 
70
- # 4. Restart the coding agent, then run a normal shell command through it.
109
+ # 5. Restart the coding agent, then run a normal shell command through it.
71
110
 
72
- # 5. Check configuration, real interception evidence, and savings.
111
+ # 6. Check configuration, real interception evidence, and savings.
73
112
  token-harness status
74
113
  token-harness verify
75
114
  token-harness metrics --since 7d
@@ -103,6 +142,11 @@ Everything else is refused, and that is the design rather than a gap: `doctor` d
103
142
  every supported platform, and only the *mutation* is narrower. An uncovered combination exits 9 and
104
143
  the diagnostic names what is missing — the reviewed fixture, or the nearest row it does have.
105
144
 
145
+ A live Linux check on 2026-09-02 with Codex 0.152.1 confirmed that boundary: the current row covers
146
+ Codex 0.146.0 on Windows only, so `plan` and `apply --yes` refused with exit 9 and wrote nothing.
147
+ That combination can still be observed and benchmarked with a provider installed through its own
148
+ installer; it is not promoted to managed mutation until a real compatibility fixture is recorded.
149
+
106
150
  What is not covered today, and why:
107
151
 
108
152
  - **macOS and Linux.** No row on either. The recordings a row needs are states of a real machine, and
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "token-harness",
3
- "version": "0.1.3",
3
+ "version": "0.1.5",
4
4
  "description": "One control plane for token-efficient coding agents.",
5
5
  "license": "Apache-2.0",
6
6
  "type": "module",
package/sbom.json CHANGED
@@ -1,13 +1,14 @@
1
1
  {
2
2
  "bomFormat": "CycloneDX",
3
3
  "specVersion": "1.5",
4
+ "serialNumber": "urn:uuid:f4cf234e-2ad8-b2a3-a085-7056eb8a17c4",
4
5
  "version": 1,
5
6
  "metadata": {
6
7
  "component": {
7
8
  "type": "application",
8
9
  "bom-ref": "token-harness",
9
10
  "name": "token-harness",
10
- "version": "0.1.3",
11
+ "version": "0.1.5",
11
12
  "description": "One control plane for token-efficient coding agents.",
12
13
  "licenses": [
13
14
  {
@@ -19,7 +20,7 @@
19
20
  "hashes": [
20
21
  {
21
22
  "alg": "SHA-256",
22
- "content": "d28d7044c0215e3471e990732b14275cca61e57a823e7a04a505748be73613c8"
23
+ "content": "f4cf234e2ad8b2a3a0857056eb8a17c494730c3d5c11e702edc98b270f18d203"
23
24
  }
24
25
  ]
25
26
  },