@velum-labs/routekit-eval-setup 1.1.1 → 1.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +4 -2
- package/dist/effect-api.d.ts +2 -0
- package/dist/effect-api.js +1 -0
- package/dist/errors.d.ts +9 -0
- package/dist/errors.js +5 -0
- package/dist/index.d.ts +3 -1
- package/dist/index.js +2 -1
- package/dist/model-selection.d.ts +2 -0
- package/dist/model-selection.js +7 -0
- package/dist/project-artifacts.js +17 -1
- package/dist/project-authoring.d.ts +1 -1
- package/dist/project-contracts.d.ts +62 -1
- package/dist/project-contracts.js +20 -2
- package/dist/project-workflow.d.ts +7 -6
- package/dist/project-workflow.js +93 -8
- package/dist/questions.d.ts +1 -1
- package/dist/questions.js +7 -1
- package/dist/services/model-catalog/service.d.ts +10 -0
- package/dist/services/model-catalog/service.js +3 -0
- package/dist/test/model-selection.test.js +6 -1
- package/dist/test/project-workflow.test.js +81 -6
- package/dist/test/questions.test.js +7 -1
- package/package.json +4 -5
- package/dist/test/skill.test.d.ts +0 -1
- package/dist/test/skill.test.js +0 -35
- package/skills/setup-eval-routing/SKILL.md +0 -163
|
@@ -1,163 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: setup-eval-routing
|
|
3
|
-
description: >-
|
|
4
|
-
Onboard or maintain a repository's compositional RouteKit eval routing through
|
|
5
|
-
the public routekit eval CLI. Use when the user wants to define a routing
|
|
6
|
-
basis, review workload dimensions, author or approve evaluations, validate or
|
|
7
|
-
estimate a billed plan, run model evidence, inspect results, activate
|
|
8
|
-
model:auto routing, or resume an interrupted eval project.
|
|
9
|
-
---
|
|
10
|
-
|
|
11
|
-
# Set Up Eval Routing
|
|
12
|
-
|
|
13
|
-
Use the public `routekit eval` CLI as the product boundary. Do not substitute
|
|
14
|
-
internal services, a standalone eval executable, or testkit qualification
|
|
15
|
-
commands for missing CLI functionality.
|
|
16
|
-
|
|
17
|
-
Inside the RouteKit source checkout, build first and use:
|
|
18
|
-
|
|
19
|
-
```text
|
|
20
|
-
node packages/cli/dist/index.js
|
|
21
|
-
```
|
|
22
|
-
|
|
23
|
-
For an installed release, use `routekit`. Refer to either form as `$ROUTEKIT`.
|
|
24
|
-
|
|
25
|
-
## Discover the available interface
|
|
26
|
-
|
|
27
|
-
Run `$ROUTEKIT eval --help` and the relevant subcommand help before acting. Use
|
|
28
|
-
only commands and flags exposed by that CLI version. Prefer `--json` for state,
|
|
29
|
-
plans, estimates, and results.
|
|
30
|
-
|
|
31
|
-
Normal model-backed commands use RouteKit's standard target resolution:
|
|
32
|
-
|
|
33
|
-
1. an explicit global `--remote <name>` or `--local` selection;
|
|
34
|
-
2. the active remote, when configured; otherwise
|
|
35
|
-
3. the local daemon.
|
|
36
|
-
|
|
37
|
-
Do not ask for a gateway URL or credential when that configured target works.
|
|
38
|
-
An explicitly supported external gateway mode may use `--gateway-url` and one
|
|
39
|
-
private credential source, but it is for qualification only and must never
|
|
40
|
-
publish a routing activation.
|
|
41
|
-
|
|
42
|
-
## Resume or initialize
|
|
43
|
-
|
|
44
|
-
1. Run `$ROUTEKIT --json eval status` from the repository root.
|
|
45
|
-
2. If no eval project exists, run `$ROUTEKIT --json eval setup`.
|
|
46
|
-
3. Follow `nextAction` and the returned artifact paths. Durable state lives
|
|
47
|
-
under `.routekit/evals`; resume it rather than starting over after an
|
|
48
|
-
interruption.
|
|
49
|
-
4. When setup returns a question, relay exactly that question and its context.
|
|
50
|
-
Ask one question per turn and never answer it for the user.
|
|
51
|
-
5. Submit the answer unchanged with `eval answer`. Prefer a private temporary
|
|
52
|
-
answer file for multiline text when the CLI supports one, then remove it.
|
|
53
|
-
|
|
54
|
-
Candidate, classifier, author, and judge roles must use explicit
|
|
55
|
-
`provider/model` IDs. Eval traffic must never use `model: auto`.
|
|
56
|
-
|
|
57
|
-
## Build and review the routing basis
|
|
58
|
-
|
|
59
|
-
Use the CLI workflow in this order, following the current `nextAction`:
|
|
60
|
-
|
|
61
|
-
1. `eval propose dimensions`
|
|
62
|
-
2. review the generated routing basis and every workload dimension;
|
|
63
|
-
3. `eval approve dimensions`
|
|
64
|
-
4. `eval propose evaluations`
|
|
65
|
-
5. review every dimension suite, case identity, rubric, manifest, and source
|
|
66
|
-
boundary;
|
|
67
|
-
6. `eval approve evaluations`
|
|
68
|
-
7. `eval validate`
|
|
69
|
-
|
|
70
|
-
Treat proposals as review material, not activation evidence. Approval is bound
|
|
71
|
-
to the exact artifact digest. If an artifact changes, validate and approve the
|
|
72
|
-
new digest rather than reusing an old approval.
|
|
73
|
-
|
|
74
|
-
A useful routing basis normally contains 5–10 orthogonal workload dimensions.
|
|
75
|
-
Each definition must include positive scope, exclusions, and a contrast pair:
|
|
76
|
-
one request that should receive majority weight on the dimension and one
|
|
77
|
-
same-workload near-miss that should route to a sibling dimension or unknown.
|
|
78
|
-
Request-envelope capabilities such as tools, vision, context, and maximum
|
|
79
|
-
output are hard requirements, not semantic workload dimensions.
|
|
80
|
-
|
|
81
|
-
Reject the proposal instead of approving its digest when:
|
|
82
|
-
|
|
83
|
-
- any dimension lacks an exclusive in-scope request or a distinct near-miss;
|
|
84
|
-
- product-behavior axes are mixed with repository-change/process axes; or
|
|
85
|
-
- a dimension is an implementation, tests/docs/CI/release, eval/classifier, or
|
|
86
|
-
other always-on layer that would receive high weight on almost every ticket.
|
|
87
|
-
|
|
88
|
-
Unknown weight absorbs the remainder. Do not add catch-all axes to cover it.
|
|
89
|
-
For a gateway product basis, protocol, selection, classification, and quota can
|
|
90
|
-
share one classifier; repository operations belong in a second classifier or
|
|
91
|
-
in tags, not in the same summed vector.
|
|
92
|
-
|
|
93
|
-
Keep generated evaluations and sanitized structured results reviewable in the
|
|
94
|
-
repository when the user approves committing them. Do not hand-edit immutable
|
|
95
|
-
plans or measured run records.
|
|
96
|
-
|
|
97
|
-
## Preserve the classifier boundary
|
|
98
|
-
|
|
99
|
-
The decomposition classifier receives only:
|
|
100
|
-
|
|
101
|
-
- the request; and
|
|
102
|
-
- the reviewed workload-dimension definitions.
|
|
103
|
-
|
|
104
|
-
It emits one weight per dimension plus an unknown weight, normalized to sum to
|
|
105
|
-
one. It must not receive candidate models, evidence, prices, objectives,
|
|
106
|
-
selected models, fallbacks, or previous routing decisions.
|
|
107
|
-
|
|
108
|
-
Model selection is deterministic. It combines the request decomposition, hard
|
|
109
|
-
requirements, the approved objective and constraints, and the published
|
|
110
|
-
model-by-dimension evidence matrix.
|
|
111
|
-
|
|
112
|
-
## Estimate and run
|
|
113
|
-
|
|
114
|
-
Before every billed step:
|
|
115
|
-
|
|
116
|
-
1. explain what will call models and show the resolved target and explicit model
|
|
117
|
-
roles;
|
|
118
|
-
2. run `eval estimate` for the intended scope;
|
|
119
|
-
3. report exact call and token limits and the CLI's pricing status—missing
|
|
120
|
-
pricing is unknown, never zero; and
|
|
121
|
-
4. obtain explicit user approval for that plan.
|
|
122
|
-
|
|
123
|
-
Run the immutable plan with `eval run` using the plan identifier returned by
|
|
124
|
-
the CLI. Do not silently reduce case counts, change candidates, replace failed
|
|
125
|
-
rows, or retry a completed paid plan. After an interruption or ambiguous
|
|
126
|
-
result, use `eval status` and `eval results` before deciding whether work
|
|
127
|
-
remains.
|
|
128
|
-
|
|
129
|
-
Never expose credentials, prompts, responses, headers, or raw child output.
|
|
130
|
-
Never recursively evaluate through `model: auto`.
|
|
131
|
-
|
|
132
|
-
## Review results and activate
|
|
133
|
-
|
|
134
|
-
Use `eval results` to review the decomposition benchmark, every dimension
|
|
135
|
-
suite, the composition benchmark, the complete evidence matrix, accounting,
|
|
136
|
-
and cleanup outcome.
|
|
137
|
-
|
|
138
|
-
Run `eval publish` only when all of these are true:
|
|
139
|
-
|
|
140
|
-
- dimensions and evaluations were approved at their current digests;
|
|
141
|
-
- validation passed and the immutable plan is still fresh;
|
|
142
|
-
- every configured candidate has exactly one judged result for every expected
|
|
143
|
-
case in every dimension;
|
|
144
|
-
- the decomposition and composition benchmarks passed;
|
|
145
|
-
- the run has zero active reservations and zero unknown measurements;
|
|
146
|
-
- the user reviewed the results and explicitly approved activation; and
|
|
147
|
-
- the run used a configured local or remote RouteKit target, not an external
|
|
148
|
-
gateway.
|
|
149
|
-
|
|
150
|
-
Publication installs already-measured evidence atomically; it must not perform
|
|
151
|
-
another billed run. After publication, check `eval status`, then verify an
|
|
152
|
-
ordinary headerless `model: auto` request. Use `routekit calls inspect <call-id>`
|
|
153
|
-
to inspect sanitized routing provenance when needed.
|
|
154
|
-
|
|
155
|
-
## Safety
|
|
156
|
-
|
|
157
|
-
- Never spend or publish silently.
|
|
158
|
-
- Never send repository material before approval for the model-backed step.
|
|
159
|
-
- Never log, echo, or commit credentials.
|
|
160
|
-
- Never describe unknown cost as zero.
|
|
161
|
-
- Never publish incomplete, stale, mismatched, duplicated, or cutoff evidence.
|
|
162
|
-
- Never claim that a passing pilot or classifier-only run is production routing
|
|
163
|
-
qualification.
|