token-harness 0.1.20 → 0.1.22

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -188,6 +188,12 @@ RTK has no equivalent machine-readable capability endpoint, so releases newer th
188
188
  reviewed RTK set remain visible as `unknown-newer` until their consumed contract is checked. See
189
189
  [docs/provider-version-compatibility.md](docs/provider-version-compatibility.md).
190
190
 
191
+ **Codex hook activation is manual.** Token Harness can write an RTK hook declaration to
192
+ `hooks.json`, but Codex separately requires the hook to be enabled and trusted in Codex. After
193
+ setup, open Codex and enable/trust the hook. Token Harness never grants trust. A declaration in
194
+ `hooks.json` is config-only evidence until Codex reports the hook enabled and trusted; run
195
+ `token-harness verify --harness codex` to inspect the available verification tier.
196
+
191
197
  Historical evaluation evidence remains available for mcptoon, GitNexus and Headroom. Detection or a promising
192
198
  benchmark is not enough for a production-stack promotion or savings claim. Their campaign assessment is structured evidence for the
193
199
  selection gate, not an activation or promotion decision. A candidate must pass structured promotion
@@ -196,8 +202,7 @@ verification, managed lifecycle, compatibility/reversibility, project maturity a
196
202
  validation. Broader context owners also require an explicit admission decision.
197
203
 
198
204
  See [docs/optimizer-priorities.md](docs/optimizer-priorities.md) and
199
- [RFC 0027](docs/rfcs/0027-optimization-stack-manager.md) and
200
- [RFC 0028](docs/rfcs/0028-smart-model-routing.md).
205
+ [RFC 0027](docs/rfcs/0027-optimization-stack-manager.md).
201
206
 
202
207
  ## Stable-stack operating model
203
208
 
@@ -248,7 +253,6 @@ browser controller itself.
248
253
  | `apply` | Apply a reviewed stored plan | Yes, only with `--yes` |
249
254
  | `verify` | Check the declared integration tier | No |
250
255
  | `metrics` | Report attributable reducer savings | No |
251
- | `routing` | Export/configure an owned CCR rule or inspect routing decisions | Yes, only after preview and `--yes` |
252
256
  | `status` | Report pipelines, drift and importer modes | No |
253
257
  | `update` | Check/update reviewed provider packages | Yes, only with `--yes` |
254
258
  | `rollback` | Restore the latest transaction snapshot | Yes, only with `--yes` |
@@ -264,79 +268,6 @@ The older automation contracts remain available. `ui --json` preserves its exist
264
268
  report; `ui --read-only` opens the legacy read-only UI; `ui --no-open` starts the guided app without
265
269
  launching a browser.
266
270
 
267
- ### Smart Model Routing (advanced)
268
-
269
- The local TypeScript classifier needs no API key or local model. On Node.js 22+, Token Harness can
270
- install and start the reviewed CCR 3.1.1 CLI in its own protected state directory. Installation and
271
- routing configuration are separate preview/apply steps. Existing authenticated CCR services can be
272
- used as-is; Token Harness does not adopt or update an external/global CCR installation. For a useful
273
- shadow report, the selected CCR profile needs an existing provider/model; when exactly one matching
274
- Claude Code or Codex provider is already configured, Token Harness adds a profile scoped to CCR CLI
275
- launches if the provider exposes one unambiguous default model. If it exposes several models, set
276
- `TOKEN_HARNESS_ROUTING_PROFILE_MODEL` to the exact `Provider/model` for the profile. To enable a
277
- conservative candidate, set `TOKEN_HARNESS_ROUTING_SIMPLE_MODEL` to an exact configured
278
- `Provider/model`; Token Harness validates it against CCR's provider catalog. It never imports OAuth
279
- credentials or edits native harness endpoints. Select provider login/import explicitly in CCR when
280
- needed. See CCR's
281
- [Agent Profiles guide](https://github.com/musistudio/claude-code-router/blob/main/docs/src/content/docs/en/configuration/profiles.md).
282
- On a first install, preview and approve CCR install/start, then preview and approve routing setup:
283
-
284
- ```sh
285
- # First preview and approve CCR install/start.
286
- token-harness routing --configure-ccr --harness codex
287
- token-harness routing --configure-ccr --harness codex --yes
288
- # Then preview and approve the routing rule and optional profile.
289
- token-harness routing --configure-ccr --harness codex
290
- token-harness routing --configure-ccr --harness codex --yes
291
- ```
292
-
293
- To update the Token Harness-owned CCR CLI to the current reviewed version pin, preview and apply:
294
-
295
- ```sh
296
- token-harness routing --update-ccr
297
- token-harness routing --update-ccr --yes
298
- ```
299
-
300
- After setup, launch the scoped profile shown by Token Harness (for Codex, typically
301
- `ccr "Token Harness Codex"`) and confirm a real request appears in CCR logs. The saved profile or
302
- gateway status alone does not prove interception. In shadow mode the rule records the proposed tier
303
- and leaves the current model unchanged.
304
-
305
- For a paired routing experiment, capture the baseline while the rule is in shadow mode, then record
306
- the task's actual quality outcome. To test conservative routing, set
307
- `TOKEN_HARNESS_ROUTING_SIMPLE_MODEL` in the Token Harness environment to an exact model already
308
- configured in CCR, roll back the owned shadow rule, and preview/apply the conservative rule. Token
309
- Harness validates the alias and embeds it in the script. Run the same task as the optimized variant,
310
- record its quality, then compare the receipt paths printed by Token Harness:
311
-
312
- ```sh
313
- token-harness benchmark-start --benchmark-id routing-codex-1 --variant baseline --task mechanical --harness codex
314
- # Run the task with CCR in shadow mode, then finish with the actual quality and attempt counts.
315
- token-harness benchmark-finish --benchmark-id routing-codex-1 --variant baseline --quality passed --attempts 1 --failed-attempts 0
316
-
317
- token-harness routing --rollback-ccr --harness codex
318
- token-harness routing --rollback-ccr --harness codex --yes
319
- token-harness routing --configure-ccr --harness codex --route-mode conservative
320
- token-harness routing --configure-ccr --harness codex --route-mode conservative --yes
321
- token-harness benchmark-start --benchmark-id routing-codex-1 --variant optimized --task mechanical --harness codex
322
- # Repeat the same task under comparable conditions, then finish with its real quality outcome.
323
- token-harness benchmark-finish --benchmark-id routing-codex-1 --variant optimized --quality passed --attempts 1 --failed-attempts 0
324
-
325
- token-harness benchmark --baseline /path/to/baseline.json --optimized /path/to/optimized.json
326
- ```
327
-
328
- To inspect decisions and local CCR request usage outside the task comparison, use
329
- `token-harness routing --route-metrics` or add `--ccr-usage`.
330
-
331
- The benchmark reads CCR session counters only when the task produced local routing events with a
332
- recognized harness identity. Receipts retain model and token aggregates, not prompts or session
333
- IDs. The comparator shows CCR token and provider-cost-estimate deltas separately and only when both
334
- observations are complete and both quality gates pass. This is not evidence of saved Codex/Claude
335
- subscription quota; quota deltas remain separately attributable, and no live savings are claimed
336
- until real paired tasks have been measured. `token-harness benchmark-matrix` also aggregates CCR
337
- usage across complete quality-passed pairs and reports withheld pairs separately. Shadow mode
338
- remains the default, and switching an owned rule's mode requires rollback before reconfiguration.
339
-
340
271
  ### Evaluation evidence (advanced / maintainers)
341
272
 
342
273
  Evaluation campaigns are an advanced maintainer workflow; managed setup stays in the unified **Optimization Stack**. The app
@@ -553,7 +484,6 @@ does not require that Corepack shim write.
553
484
 
554
485
  Before changing public behavior or architecture, read
555
486
  [RFC 0027](docs/rfcs/0027-optimization-stack-manager.md),
556
- [RFC 0028](docs/rfcs/0028-smart-model-routing.md),
557
487
  [docs/optimizer-priorities.md](docs/optimizer-priorities.md),
558
488
  [docs/release-readiness.md](docs/release-readiness.md), [PLAN.md](PLAN.md), and the accepted
559
489
  [RFCs](docs/rfcs).
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "token-harness",
3
- "version": "0.1.20",
3
+ "version": "0.1.22",
4
4
  "description": "Quota-aware efficiency layer for Claude Code and Codex subscription limits.",
5
5
  "license": "Apache-2.0",
6
6
  "type": "module",
package/sbom.json CHANGED
@@ -1,14 +1,14 @@
1
1
  {
2
2
  "bomFormat": "CycloneDX",
3
3
  "specVersion": "1.5",
4
- "serialNumber": "urn:uuid:6a907cd9-4778-ca06-b653-78caf2cf25f0",
4
+ "serialNumber": "urn:uuid:b575e831-af8f-bda5-aac1-d757d0d3687b",
5
5
  "version": 1,
6
6
  "metadata": {
7
7
  "component": {
8
8
  "type": "application",
9
9
  "bom-ref": "token-harness",
10
10
  "name": "token-harness",
11
- "version": "0.1.20",
11
+ "version": "0.1.22",
12
12
  "description": "Quota-aware efficiency layer for Claude Code and Codex subscription limits.",
13
13
  "licenses": [
14
14
  {
@@ -20,7 +20,7 @@
20
20
  "hashes": [
21
21
  {
22
22
  "alg": "SHA-256",
23
- "content": "6a907cd94778ca06b65378caf2cf25f060cc82a51168e1cd611f5f58b5e19e55"
23
+ "content": "b575e831af8fbda5aac1d757d0d3687b8206375877fa1305d014f3eb4cb92be7"
24
24
  }
25
25
  ]
26
26
  },