token-harness 0.1.3 → 0.1.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +84 -40
- package/package.json +1 -1
- package/sbom.json +3 -2
- package/token-harness.mjs +6275 -801
package/README.md
CHANGED
|
@@ -1,51 +1,81 @@
|
|
|
1
1
|
# Token Harness
|
|
2
2
|
|
|
3
|
-
Token Harness has one objective: **
|
|
4
|
-
|
|
3
|
+
Token Harness has one objective: **maximize the useful coding work you can get from Claude Code
|
|
4
|
+
and Codex usage limits, without hiding quality regressions or pretending that opaque subscription
|
|
5
|
+
quota is exactly equivalent to a token count**.
|
|
5
6
|
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
7
|
+
The scarce resource is no longer just model context. Subscription users are constrained by rolling
|
|
8
|
+
usage windows, weekly limits, model-dependent burn, long-session context growth, MCP/tool schema
|
|
9
|
+
overhead, noisy tool output, and repeated work after poor model or effort choices. Token Harness is
|
|
10
|
+
evolving from a token-reduction control plane into a **quota-aware efficiency layer** for coding
|
|
11
|
+
harnesses.
|
|
10
12
|
|
|
11
|
-
|
|
12
|
-
token-saving tools on the machine, selects a compatible owner for each interception point, shows
|
|
13
|
-
every proposed change before applying it, verifies whether the integration is genuinely being
|
|
14
|
-
used, and reports how many tokens or characters were saved.
|
|
13
|
+
The product therefore optimizes two related budgets:
|
|
15
14
|
|
|
16
|
-
|
|
17
|
-
|
|
15
|
+
1. **subscription headroom** — observe the usage windows the harness itself exposes, pace work
|
|
16
|
+
against reset time, and choose native model/effort settings that fit the task and remaining
|
|
17
|
+
budget;
|
|
18
|
+
2. **context sent to the model** — keep instructions, MCP schemas, conversation history, repository
|
|
19
|
+
context, and tool output as small as possible without removing information needed to finish the
|
|
20
|
+
task correctly.
|
|
21
|
+
|
|
22
|
+
RTK and HarnessTrim remain useful providers, but reducers are now one layer of the system rather
|
|
23
|
+
than the product definition. Native harness controls come first because choosing the right model,
|
|
24
|
+
effort, context shape, and enabled tools can avoid an expensive turn entirely.
|
|
25
|
+
|
|
26
|
+
## What Token Harness should optimize
|
|
27
|
+
|
|
28
|
+
| Layer | Target behavior |
|
|
29
|
+
| --- | --- |
|
|
30
|
+
| Usage-window observability | Read five-hour, weekly, model-specific, or credit-backed limits only from surfaces the harness can actually prove; show reset time, headroom, and burn rate |
|
|
31
|
+
| Native model policy | Prefer economical models/effort for routine work, escalate only for tasks whose expected quality benefit justifies the extra quota |
|
|
32
|
+
| Session hygiene | Detect task boundaries, oversized conversations, and stale context; recommend clear/compact/new-session actions before history dominates each turn |
|
|
33
|
+
| Instruction budget | Keep `CLAUDE.md` / `AGENTS.md` concise and hierarchical instead of injecting one large global instruction file everywhere |
|
|
34
|
+
| MCP/tool budget | Disable irrelevant MCP servers, defer schemas where the harness supports it, and expose only the tools required by the current task |
|
|
35
|
+
| Tool-output budget | Reduce logs, diffs, test output, and repeated command results before they re-enter model context |
|
|
36
|
+
| Cross-harness scheduling | When Claude Code and Codex have independent headroom, recommend the harness that can do the task with the best expected quality per remaining quota |
|
|
37
|
+
| Measurement | Correlate quota delta, model/effort, context size, tool traffic, and task outcome; never convert local token savings into an unproven subscription-quota claim |
|
|
18
38
|
|
|
19
39
|
## Optimization ecosystem
|
|
20
40
|
|
|
21
|
-
The
|
|
22
|
-
|
|
23
|
-
|
|
41
|
+
The priority order below is deliberately different from the original Token Harness roadmap. A tool
|
|
42
|
+
is high priority only when it helps **included Claude Code/Codex capacity**, not merely API cost,
|
|
43
|
+
provider routing, or benchmark token counts.
|
|
24
44
|
|
|
25
|
-
| Tool | Optimization layer | Token Harness
|
|
45
|
+
| Tool or surface | Optimization layer | Token Harness direction |
|
|
26
46
|
| --- | --- | --- |
|
|
27
|
-
|
|
|
28
|
-
|
|
|
29
|
-
| [
|
|
30
|
-
| [
|
|
31
|
-
| [
|
|
32
|
-
| [
|
|
33
|
-
| [
|
|
34
|
-
| [
|
|
35
|
-
| [
|
|
36
|
-
| [
|
|
37
|
-
| [
|
|
38
|
-
|
|
|
39
|
-
| [LLMLingua](https://github.com/microsoft/LLMLingua) |
|
|
40
|
-
| [Caveman](https://github.com/JuliusBrussee/caveman) | Reduce visible model-output verbosity |
|
|
41
|
-
|
|
42
|
-
Candidate status means
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
merely because it is present on the machine. Rows marked `high priority` are the next intended
|
|
46
|
-
intake; the admission gates each one carries are recorded in
|
|
47
|
+
| **Claude Code native controls** | Usage status, model/effort choice, context inspection, clear/compact lifecycle | **P0 — build into core**. Prefer native surfaces over third-party credential scraping; discover available models dynamically |
|
|
48
|
+
| **Codex native app-server + profiles** | Structured rate-limit telemetry; model, reasoning effort, verbosity, tool-output and MCP/context controls | **P0 — build into core**. `account/rateLimits/read` is the preferred live quota source when available |
|
|
49
|
+
| [cclimits](https://github.com/cruzanstx/cclimits) | Live/local quota companion across coding tools | **P0 — optional observational companion for Claude**. Token Harness uses only cacheless Claude JSON mode, never reads its credentials, and labels the result `reported` or `cached`. Codex keeps its native app-server reader |
|
|
50
|
+
| [ccusage](https://github.com/ccusage/ccusage) | Local historical token/session/cost analytics for Claude Code and Codex | **P0 — high-priority read-only companion**. Useful for history and baselines; not a substitute for live subscription-limit telemetry |
|
|
51
|
+
| [RTK](https://github.com/rtk-ai/rtk) | Shell-command rewriting and command-output reduction | **P1 — active, keep**. Reduces context that would otherwise be resent on later turns |
|
|
52
|
+
| [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic reducers, harness adapters, skills, pipes, and MCP reduction | **P1 — active, keep**. Continue only on proven non-overlapping surfaces |
|
|
53
|
+
| [Lazy MCP](https://gitlab.com/gitlab-org/ai/lazy-mcp) | Load MCP tool schemas only when needed | **P1 — high priority**. Directly attacks schema/context overhead; benchmark against native Codex tool deferral before installing another owner |
|
|
54
|
+
| [Context Mode](https://github.com/mksglu/context-mode) | Keep raw tool/MCP results outside model context | **P2 — alternative broad context owner**. Evaluate against Headroom, not alongside overlapping reducers by default |
|
|
55
|
+
| [Headroom](https://github.com/headroomlabs-ai/headroom) | Compress tool, MCP, file, and RAG payloads | **P2 — alternative broad context owner**. Admit only after pair-specific quality and attribution fixtures |
|
|
56
|
+
| [Dejavu](https://github.com/Salnika/dejavu) | Emit only the delta when command output repeats | **P2 — useful for test/rerun loops** after the normal output-reduction path is measured |
|
|
57
|
+
| [repowise](https://github.com/repowise-dev/repowise) | Retrieve task-specific repository context | **P2 — conditional**. Its MCP overhead must be lower than the repository context it avoids |
|
|
58
|
+
| LiteLLM, Claude Code Router, RouteLLM, LLMRouter, vLLM Semantic Router | Route requests across models/providers | **P3 — overflow/API routing only**. Useful after included quota is exhausted or for explicit external-provider policy; not a core subscription-quota optimizer |
|
|
59
|
+
| [LLMLingua](https://github.com/microsoft/LLMLingua) | Generic prompt compression | **Research only** until a harness-aware lifecycle proves must-keep recall and quality |
|
|
60
|
+
| [Caveman](https://github.com/JuliusBrussee/caveman) | Reduce visible model-output verbosity | **De-prioritized**. Use native verbosity/effort controls first and measure their effect before adding another instruction owner |
|
|
61
|
+
|
|
62
|
+
Candidate status still means that Token Harness neither installs nor configures the tool until its
|
|
63
|
+
installation, conflicts, rollback behavior, verification, quality gates, and metrics attribution are
|
|
64
|
+
implemented and tested. The intake evidence lives in
|
|
47
65
|
[docs/provider-landscape.md](docs/provider-landscape.md).
|
|
48
66
|
|
|
67
|
+
### Why generic routers moved down
|
|
68
|
+
|
|
69
|
+
The original roadmap treated model routers as high priority. For the new objective that is backwards:
|
|
70
|
+
a router can reduce API cost by sending work to another provider while consuming **none of the
|
|
71
|
+
included Claude Code or Codex allowance**, or it can bypass subscription authentication entirely.
|
|
72
|
+
That can be valuable as explicit overflow, but it does not demonstrate better use of the allowance
|
|
73
|
+
the user already paid for.
|
|
74
|
+
|
|
75
|
+
Token Harness should first exploit the harness-native choices that are actually inside the plan:
|
|
76
|
+
model tier, reasoning effort, verbosity, session lifecycle, instruction size, MCP exposure, and tool
|
|
77
|
+
output. External routing becomes an opt-in second budget, never an invisible shortcut.
|
|
78
|
+
|
|
49
79
|
## Quick start
|
|
50
80
|
|
|
51
81
|
Install the CLI:
|
|
@@ -61,15 +91,24 @@ Then run the complete workflow from the project in which you use your coding age
|
|
|
61
91
|
# 1. Inspect the machine. This does not change agent configuration.
|
|
62
92
|
token-harness doctor
|
|
63
93
|
|
|
64
|
-
# 2.
|
|
94
|
+
# 2. Inspect subscription headroom and avoidable context. Still read-only.
|
|
95
|
+
token-harness budget
|
|
96
|
+
token-harness context
|
|
97
|
+
token-harness history --since 7d
|
|
98
|
+
token-harness optimize
|
|
99
|
+
|
|
100
|
+
# Optional: Claude live quota can use an already-installed cclimits build that
|
|
101
|
+
# supports --no-cache-write. Token Harness never installs it automatically.
|
|
102
|
+
|
|
103
|
+
# 3. Preview every proposed configuration change.
|
|
65
104
|
token-harness plan
|
|
66
105
|
|
|
67
|
-
#
|
|
106
|
+
# 4. Apply the reviewed plan. This is the first configuration-changing step.
|
|
68
107
|
token-harness apply --yes
|
|
69
108
|
|
|
70
|
-
#
|
|
109
|
+
# 5. Restart the coding agent, then run a normal shell command through it.
|
|
71
110
|
|
|
72
|
-
#
|
|
111
|
+
# 6. Check configuration, real interception evidence, and savings.
|
|
73
112
|
token-harness status
|
|
74
113
|
token-harness verify
|
|
75
114
|
token-harness metrics --since 7d
|
|
@@ -103,6 +142,11 @@ Everything else is refused, and that is the design rather than a gap: `doctor` d
|
|
|
103
142
|
every supported platform, and only the *mutation* is narrower. An uncovered combination exits 9 and
|
|
104
143
|
the diagnostic names what is missing — the reviewed fixture, or the nearest row it does have.
|
|
105
144
|
|
|
145
|
+
A live Linux check on 2026-09-02 with Codex 0.152.1 confirmed that boundary: the current row covers
|
|
146
|
+
Codex 0.146.0 on Windows only, so `plan` and `apply --yes` refused with exit 9 and wrote nothing.
|
|
147
|
+
That combination can still be observed and benchmarked with a provider installed through its own
|
|
148
|
+
installer; it is not promoted to managed mutation until a real compatibility fixture is recorded.
|
|
149
|
+
|
|
106
150
|
What is not covered today, and why:
|
|
107
151
|
|
|
108
152
|
- **macOS and Linux.** No row on either. The recordings a row needs are states of a real machine, and
|
package/package.json
CHANGED
package/sbom.json
CHANGED
|
@@ -1,13 +1,14 @@
|
|
|
1
1
|
{
|
|
2
2
|
"bomFormat": "CycloneDX",
|
|
3
3
|
"specVersion": "1.5",
|
|
4
|
+
"serialNumber": "urn:uuid:f4cf234e-2ad8-b2a3-a085-7056eb8a17c4",
|
|
4
5
|
"version": 1,
|
|
5
6
|
"metadata": {
|
|
6
7
|
"component": {
|
|
7
8
|
"type": "application",
|
|
8
9
|
"bom-ref": "token-harness",
|
|
9
10
|
"name": "token-harness",
|
|
10
|
-
"version": "0.1.
|
|
11
|
+
"version": "0.1.5",
|
|
11
12
|
"description": "One control plane for token-efficient coding agents.",
|
|
12
13
|
"licenses": [
|
|
13
14
|
{
|
|
@@ -19,7 +20,7 @@
|
|
|
19
20
|
"hashes": [
|
|
20
21
|
{
|
|
21
22
|
"alg": "SHA-256",
|
|
22
|
-
"content": "
|
|
23
|
+
"content": "f4cf234e2ad8b2a3a0857056eb8a17c494730c3d5c11e702edc98b270f18d203"
|
|
23
24
|
}
|
|
24
25
|
]
|
|
25
26
|
},
|