@velum-labs/routekit-eval-setup 1.1.1 → 1.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,163 +0,0 @@
1
- ---
2
- name: setup-eval-routing
3
- description: >-
4
- Onboard or maintain a repository's compositional RouteKit eval routing through
5
- the public routekit eval CLI. Use when the user wants to define a routing
6
- basis, review workload dimensions, author or approve evaluations, validate or
7
- estimate a billed plan, run model evidence, inspect results, activate
8
- model:auto routing, or resume an interrupted eval project.
9
- ---
10
-
11
- # Set Up Eval Routing
12
-
13
- Use the public `routekit eval` CLI as the product boundary. Do not substitute
14
- internal services, a standalone eval executable, or testkit qualification
15
- commands for missing CLI functionality.
16
-
17
- Inside the RouteKit source checkout, build first and use:
18
-
19
- ```text
20
- node packages/cli/dist/index.js
21
- ```
22
-
23
- For an installed release, use `routekit`. Refer to either form as `$ROUTEKIT`.
24
-
25
- ## Discover the available interface
26
-
27
- Run `$ROUTEKIT eval --help` and the relevant subcommand help before acting. Use
28
- only commands and flags exposed by that CLI version. Prefer `--json` for state,
29
- plans, estimates, and results.
30
-
31
- Normal model-backed commands use RouteKit's standard target resolution:
32
-
33
- 1. an explicit global `--remote <name>` or `--local` selection;
34
- 2. the active remote, when configured; otherwise
35
- 3. the local daemon.
36
-
37
- Do not ask for a gateway URL or credential when that configured target works.
38
- An explicitly supported external gateway mode may use `--gateway-url` and one
39
- private credential source, but it is for qualification only and must never
40
- publish a routing activation.
41
-
42
- ## Resume or initialize
43
-
44
- 1. Run `$ROUTEKIT --json eval status` from the repository root.
45
- 2. If no eval project exists, run `$ROUTEKIT --json eval setup`.
46
- 3. Follow `nextAction` and the returned artifact paths. Durable state lives
47
- under `.routekit/evals`; resume it rather than starting over after an
48
- interruption.
49
- 4. When setup returns a question, relay exactly that question and its context.
50
- Ask one question per turn and never answer it for the user.
51
- 5. Submit the answer unchanged with `eval answer`. Prefer a private temporary
52
- answer file for multiline text when the CLI supports one, then remove it.
53
-
54
- Candidate, classifier, author, and judge roles must use explicit
55
- `provider/model` IDs. Eval traffic must never use `model: auto`.
56
-
57
- ## Build and review the routing basis
58
-
59
- Use the CLI workflow in this order, following the current `nextAction`:
60
-
61
- 1. `eval propose dimensions`
62
- 2. review the generated routing basis and every workload dimension;
63
- 3. `eval approve dimensions`
64
- 4. `eval propose evaluations`
65
- 5. review every dimension suite, case identity, rubric, manifest, and source
66
- boundary;
67
- 6. `eval approve evaluations`
68
- 7. `eval validate`
69
-
70
- Treat proposals as review material, not activation evidence. Approval is bound
71
- to the exact artifact digest. If an artifact changes, validate and approve the
72
- new digest rather than reusing an old approval.
73
-
74
- A useful routing basis normally contains 5–10 orthogonal workload dimensions.
75
- Each definition must include positive scope, exclusions, and a contrast pair:
76
- one request that should receive majority weight on the dimension and one
77
- same-workload near-miss that should route to a sibling dimension or unknown.
78
- Request-envelope capabilities such as tools, vision, context, and maximum
79
- output are hard requirements, not semantic workload dimensions.
80
-
81
- Reject the proposal instead of approving its digest when:
82
-
83
- - any dimension lacks an exclusive in-scope request or a distinct near-miss;
84
- - product-behavior axes are mixed with repository-change/process axes; or
85
- - a dimension is an implementation, tests/docs/CI/release, eval/classifier, or
86
- other always-on layer that would receive high weight on almost every ticket.
87
-
88
- Unknown weight absorbs the remainder. Do not add catch-all axes to cover it.
89
- For a gateway product basis, protocol, selection, classification, and quota can
90
- share one classifier; repository operations belong in a second classifier or
91
- in tags, not in the same summed vector.
92
-
93
- Keep generated evaluations and sanitized structured results reviewable in the
94
- repository when the user approves committing them. Do not hand-edit immutable
95
- plans or measured run records.
96
-
97
- ## Preserve the classifier boundary
98
-
99
- The decomposition classifier receives only:
100
-
101
- - the request; and
102
- - the reviewed workload-dimension definitions.
103
-
104
- It emits one weight per dimension plus an unknown weight, normalized to sum to
105
- one. It must not receive candidate models, evidence, prices, objectives,
106
- selected models, fallbacks, or previous routing decisions.
107
-
108
- Model selection is deterministic. It combines the request decomposition, hard
109
- requirements, the approved objective and constraints, and the published
110
- model-by-dimension evidence matrix.
111
-
112
- ## Estimate and run
113
-
114
- Before every billed step:
115
-
116
- 1. explain what will call models and show the resolved target and explicit model
117
- roles;
118
- 2. run `eval estimate` for the intended scope;
119
- 3. report exact call and token limits and the CLI's pricing status—missing
120
- pricing is unknown, never zero; and
121
- 4. obtain explicit user approval for that plan.
122
-
123
- Run the immutable plan with `eval run` using the plan identifier returned by
124
- the CLI. Do not silently reduce case counts, change candidates, replace failed
125
- rows, or retry a completed paid plan. After an interruption or ambiguous
126
- result, use `eval status` and `eval results` before deciding whether work
127
- remains.
128
-
129
- Never expose credentials, prompts, responses, headers, or raw child output.
130
- Never recursively evaluate through `model: auto`.
131
-
132
- ## Review results and activate
133
-
134
- Use `eval results` to review the decomposition benchmark, every dimension
135
- suite, the composition benchmark, the complete evidence matrix, accounting,
136
- and cleanup outcome.
137
-
138
- Run `eval publish` only when all of these are true:
139
-
140
- - dimensions and evaluations were approved at their current digests;
141
- - validation passed and the immutable plan is still fresh;
142
- - every configured candidate has exactly one judged result for every expected
143
- case in every dimension;
144
- - the decomposition and composition benchmarks passed;
145
- - the run has zero active reservations and zero unknown measurements;
146
- - the user reviewed the results and explicitly approved activation; and
147
- - the run used a configured local or remote RouteKit target, not an external
148
- gateway.
149
-
150
- Publication installs already-measured evidence atomically; it must not perform
151
- another billed run. After publication, check `eval status`, then verify an
152
- ordinary headerless `model: auto` request. Use `routekit calls inspect <call-id>`
153
- to inspect sanitized routing provenance when needed.
154
-
155
- ## Safety
156
-
157
- - Never spend or publish silently.
158
- - Never send repository material before approval for the model-backed step.
159
- - Never log, echo, or commit credentials.
160
- - Never describe unknown cost as zero.
161
- - Never publish incomplete, stale, mismatched, duplicated, or cutoff evidence.
162
- - Never claim that a passing pilot or classifier-only run is production routing
163
- qualification.