tiergear 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +294 -0
- package/dist/cli/bin.js +9 -0
- package/dist/cli/files.js +43 -0
- package/dist/cli/judge.js +19 -0
- package/dist/cli/launch.js +50 -0
- package/dist/cli/main.js +145 -0
- package/dist/cli/node.js +53 -0
- package/dist/cli/orca.js +105 -0
- package/dist/cli/spawn.js +70 -0
- package/dist/cli/stats.js +19 -0
- package/dist/cli/status.js +78 -0
- package/dist/core/config.js +58 -0
- package/dist/core/decide.js +253 -0
- package/dist/core/failures.js +10 -0
- package/dist/core/floor.js +44 -0
- package/dist/core/judge.js +13 -0
- package/dist/core/judges/presets.js +27 -0
- package/dist/core/judges/systemone.js +85 -0
- package/dist/core/log.js +50 -0
- package/dist/core/state.js +72 -0
- package/dist/core/status.js +65 -0
- package/dist/core/tables.js +85 -0
- package/dist/core/tiers.js +53 -0
- package/package.json +38 -0
package/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 jang2162
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
package/README.md
ADDED
|
@@ -0,0 +1,294 @@
|
|
|
1
|
+
# tiergear
|
|
2
|
+
|
|
3
|
+
A plugin and CLI that lets a decision model (the judge) pick the model and reasoning effort for Claude Code sessions and Orca workers.
|
|
4
|
+
|
|
5
|
+
- **First turn**: the first prompt goes to the judge, which returns a tier (trivial, quick, standard, deep, max). Tables A and B turn that into a model and effort. The first turn has no cache to lose, so the model changes too.
|
|
6
|
+
- **Later turns**: by default the model stays and only effort changes. Raising is easy (confidence 0.5); lowering is hard (confidence 0.85 for 2 turns in a row), unless the optional instant switch is on and the judge is sure enough.
|
|
7
|
+
- **Floor**: the tier never drops below one step under the first decision. Sessions started with `tiergear launch`/`orca-spawn` use the launch tier itself as their floor, so a `--min-tier` holds for the whole session.
|
|
8
|
+
- **Ceiling**: a session launched with `--max-tier` never rises above it, whether the judge asks for a harder tier or the session looks stuck.
|
|
9
|
+
- **Pick a tier, or off, from the band**: the band above the prompt has a button per tier and one for **off** (see [Status band](#status-band)). A picked tier is a starting point; the judge keeps moving it as usual. **off** withdraws what tiergear applies until you pick a tier again.
|
|
10
|
+
- **Manual changes pause routing**: changing the model or effort yourself with `/model` or `/effort` mid-session pauses the session, just like picking **off**, and writes `[tiergear] manual model/effort change — routing paused for this session` to the hook log once. Pick a tier to hand it back to tiergear. It compares the session values the engine reports between turns, so tiergear's own changes don't trigger it. If the engine switches to a fallback model from the first request of a turn, that can also look like a manual change.
|
|
11
|
+
|
|
12
|
+
If the judge is slow or fails, that turn proceeds untouched.
|
|
13
|
+
|
|
14
|
+
### Which prompts are judged
|
|
15
|
+
|
|
16
|
+
Only prompts the user typed are judged: terminal input (`composer`), Remote Control (`bridge`), and the SDK or `claude -p` (`sdk`). Background task notifications, scheduled tasks and `/loop`, messages from other sessions or SendMessage, observers, auto-continue, plugin-sent prompts, and unknown origins (`unclassified`) pass through unjudged, and the tier and applied values stay as they are. Prompts that enter a running turn (including ones typed and queued while a turn runs) are not judged. Slash commands starting with `/` are not judged.
|
|
17
|
+
|
|
18
|
+
## Install
|
|
19
|
+
|
|
20
|
+
Requires Claude Code 2.1.289 or later. At the prompt of a Claude Code session:
|
|
21
|
+
|
|
22
|
+
```
|
|
23
|
+
/plugin marketplace add jang2162/tiergear
|
|
24
|
+
/plugin install tiergear@tiergear
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
Pick a scope when asked (user is first). The hooks run in that session at once and in every new session under that scope. `/plugin` lists it as `tiergear@tiergear`. The default judge, jev, needs a TypeSafe API key in `TYPESAFE_API_KEY` (your shell or user settings) or in the `judgeApiKey` option; without one it answers nothing and the [ Recent ] pane shows `judge no API key for jev`. Or pick a local judge (laya, kev) in the plugin options.
|
|
28
|
+
|
|
29
|
+
To update: `claude plugin update tiergear@tiergear`, then `/reload-plugins`.
|
|
30
|
+
|
|
31
|
+
### CLI (optional)
|
|
32
|
+
|
|
33
|
+
`tiergear launch`, `tiergear orca-spawn`, `tiergear stats` and `tiergear status` are a separate command-line tool; the hooks don't need it. Requires Node 20 or later.
|
|
34
|
+
|
|
35
|
+
```bash
|
|
36
|
+
npm install -g tiergear
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
`which tiergear` should print the CLI path.
|
|
40
|
+
|
|
41
|
+
### From a clone (development)
|
|
42
|
+
|
|
43
|
+
To run the hooks and the CLI from a working copy instead, so `/reload-plugins` picks up edits without an update:
|
|
44
|
+
|
|
45
|
+
```bash
|
|
46
|
+
git clone https://github.com/jang2162/tiergear.git
|
|
47
|
+
cd tiergear
|
|
48
|
+
npm install && npm run build && npm link
|
|
49
|
+
mkdir -p ~/.claude/skills
|
|
50
|
+
ln -s "$PWD" ~/.claude/skills/tiergear
|
|
51
|
+
claude plugin list | grep -A3 tiergear
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
You should see `tiergear@skills-dir` and `Status: ✔ loaded`. If `tiergear@tiergear` is also installed, it takes precedence and the clone is not loaded (`claude plugin list` says so); uninstall it with `claude plugin uninstall tiergear@tiergear` to run from the clone.
|
|
55
|
+
|
|
56
|
+
## Judges
|
|
57
|
+
|
|
58
|
+
All three judges use TypeSafe's `<baseUrl>/v1/systemone` contract. The default is `jev`.
|
|
59
|
+
|
|
60
|
+
| Preset | Address | Model | Key env var | Timeout (first turn / later) |
|
|
61
|
+
| --- | --- | --- | --- | --- |
|
|
62
|
+
| jev | `https://api.typesafe.ai` | `jev-latest` | `TYPESAFE_API_KEY` (required) | 2000ms / 1200ms |
|
|
63
|
+
| laya | `http://localhost:11435` | `laya` | `OLLAYA_API_KEY` (optional) | 3000ms / 2500ms |
|
|
64
|
+
| kev | `http://localhost:8009` | `kev-latest` | `KEV_API_KEY` (optional) | 3000ms / 2000ms |
|
|
65
|
+
|
|
66
|
+
- jev needs a key and sends data to TypeSafe's servers.
|
|
67
|
+
- laya runs locally with `ollaya run laya`.
|
|
68
|
+
- kev points at a server you run locally or on Modal (change the address with `judgeBaseUrl`).
|
|
69
|
+
- A judge address must be `https`, or `http` to this machine (`localhost`, `127.0.0.1`, `[::1]`), since prompts travel in the request. Any other address is refused and the turn goes on untouched.
|
|
70
|
+
- A preset's key goes only to that preset's own address. Point `judgeBaseUrl` (or the CLI's `--judge-url`) elsewhere and set the key for it explicitly with `judgeApiKey` (CLI: `TIERGEAR_JUDGE_API_KEY`); `TYPESAFE_API_KEY` and the like are not sent there.
|
|
71
|
+
- A key that a repository's `.claude/settings.json` supplies (in `env`) is ignored, so a cloned repo can't route your prompts to its own account. Keys in your user settings, `.claude/settings.local.json` or your shell are used.
|
|
72
|
+
- After 3 judge failures in a row, the session skips the judge for 5 minutes and keeps its current state instead of falling back to the tables.
|
|
73
|
+
|
|
74
|
+
## Tables A and B
|
|
75
|
+
|
|
76
|
+
**Table A: the model each tier starts on**
|
|
77
|
+
|
|
78
|
+
| Tier | Claude | Codex |
|
|
79
|
+
| --- | --- | --- |
|
|
80
|
+
| trivial | sonnet | gpt-5.6-luna |
|
|
81
|
+
| quick | sonnet | gpt-5.6-terra |
|
|
82
|
+
| standard | sonnet | gpt-5.6-terra |
|
|
83
|
+
| deep | opus | gpt-5.6-terra |
|
|
84
|
+
| max | fable | gpt-5.6-terra |
|
|
85
|
+
|
|
86
|
+
Claude's trivial tier uses sonnet, not haiku: Claude Code's auto mode doesn't run on haiku, so a haiku session stops to ask for approval on commands. If you run with permission checks off, see [Haiku for trivial](#haiku-for-trivial-bypass-permissions-users).
|
|
87
|
+
|
|
88
|
+
**Table B: effort per model** (`-` means no effort is sent; the haiku row applies only if you put haiku in Table A)
|
|
89
|
+
|
|
90
|
+
| Model | trivial | quick | standard | deep | max |
|
|
91
|
+
| --- | --- | --- | --- | --- | --- |
|
|
92
|
+
| haiku | - | - | - | - | - |
|
|
93
|
+
| sonnet | low | low | medium | high | max |
|
|
94
|
+
| opus | low | low | medium | xhigh | max |
|
|
95
|
+
| fable | low | low | medium | high | xhigh |
|
|
96
|
+
| gpt-5.6-luna (Codex) | low | low | medium | high | high |
|
|
97
|
+
| gpt-5.6-terra (Codex) | low | low | medium | xhigh | max |
|
|
98
|
+
|
|
99
|
+
On Codex, deep and max both use `gpt-5.6-terra`; only the effort differs (xhigh vs max).
|
|
100
|
+
|
|
101
|
+
A session on a model without effort (haiku, if you set it) can't be raised through effort alone, so raising the tier switches to that tier's model even when `switchModelMidSession=false` (the cache breaks once).
|
|
102
|
+
|
|
103
|
+
### Customizing the tables
|
|
104
|
+
|
|
105
|
+
Put only the cells you want to change in `~/.config/tiergear/tables.json`; they are merged over the defaults.
|
|
106
|
+
|
|
107
|
+
```json
|
|
108
|
+
{ "claude": { "effort": { "opus": { "deep": "max" } } } }
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
Model cells go under `models`, as in `{"claude":{"models":{"deep":"opus"}}}`. A model must be a plain id (letters, digits, `.`, `_`, `-`, `:`, `/` and brackets as in `opus[1m]`), since it ends up in a launch command a shell reads. Effort values are low, medium, high, xhigh, max, or `null`. If any value or the JSON itself is invalid, **the whole file is ignored** and the default tables are used (this is noted in the hook log). Hooks read the file once, the first time it's needed after session start, so open a new session after editing it.
|
|
112
|
+
|
|
113
|
+
### Haiku for trivial (bypass-permissions users)
|
|
114
|
+
|
|
115
|
+
If you run Claude Code with permission checks off ("yolo" mode), auto mode doesn't matter and haiku is the cheaper choice for trivial work. Put it back with one cell:
|
|
116
|
+
|
|
117
|
+
```json
|
|
118
|
+
{ "claude": { "models": { "trivial": "haiku" } } }
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
Permission checks can be turned off in two ways:
|
|
122
|
+
|
|
123
|
+
- One session: `claude --dangerously-skip-permissions`.
|
|
124
|
+
- Every session, including the workers `orca-spawn` starts: in `~/.claude/settings.json`,
|
|
125
|
+
|
|
126
|
+
```json
|
|
127
|
+
{ "permissions": { "defaultMode": "bypassPermissions" } }
|
|
128
|
+
```
|
|
129
|
+
|
|
130
|
+
Use the settings file if you start workers with tiergear. `tiergear launch` prints a command with only `--model` and `--effort`, `orca-spawn` starts the agent with that command, and inside a Run Orca's `worker-start` builds the command itself, so none of them adds the flag. With only the flag, a haiku worker asks for approval again.
|
|
131
|
+
|
|
132
|
+
**Bypass mode runs every tool call without asking.** Use it only where you accept that, such as a sandbox or a throwaway worktree.
|
|
133
|
+
|
|
134
|
+
## Plugin options
|
|
135
|
+
|
|
136
|
+
Change these in `/config` (plugin options).
|
|
137
|
+
|
|
138
|
+
| Option | Default | Description |
|
|
139
|
+
| --- | --- | --- |
|
|
140
|
+
| `judge` | `jev` | Judge preset: one of jev, laya, kev (shown as a list in `/config`). An unknown value falls back to jev and is noted once in the hook log |
|
|
141
|
+
| `judgeBaseUrl` | preset | Leave empty for the preset's address. Must be https, or http to localhost |
|
|
142
|
+
| `judgeModel` | preset | Leave empty for the preset's model |
|
|
143
|
+
| `judgeApiKey` | preset env var | Leave empty to use TYPESAFE_API_KEY, OLLAYA_API_KEY, or KEV_API_KEY, which only go to the preset's own address (sensitive) |
|
|
144
|
+
| `switchModelMidSession` | `false` | Also change the model after the first turn. Off: only effort changes. Either change breaks the messages prompt cache; a model change also breaks the tools and system caches |
|
|
145
|
+
| `instantSwitchConfidence` | off | Confidence · instant switch: a lower tier the judge is at least this sure of applies at once, straight to that tier (never under the floor), skipping `downgradeStreak`. Leave unset or 0 to keep it off |
|
|
146
|
+
| `minDowngradeConfidence` | `0.85` | Confidence · lower: minimum confidence for a turn to count toward lowering |
|
|
147
|
+
| `minUpgradeConfidence` | `0.5` | Confidence · raise: minimum confidence to raise the tier |
|
|
148
|
+
| `stuckConfidence` | `0.6` | Confidence · stuck: raise one step when the judge's stuck probability is at or above this |
|
|
149
|
+
| `downgradeStreak` | `2` | Consecutive turns needed to lower one step |
|
|
150
|
+
| `stuckFailures` | `3` | Raise one step after this many identical tool failures in a row |
|
|
151
|
+
| `firstTurnTimeoutMs` | preset | First-turn latency budget. At most 8000ms (the whole hook budget is 10s, so larger values are cut to 8000) |
|
|
152
|
+
| `turnTimeoutMs` | preset | Later-turn latency budget. At most 8000ms |
|
|
153
|
+
| `showRecentButton` | `true` | Show the **[ Recent ]** button above the prompt. Off: `/tiergear` still opens the pane |
|
|
154
|
+
| `showPrefix` | `true` | Start the band with `tiergear` |
|
|
155
|
+
| `showStatusText` | `true` | Show the line's parts below (all four at once). Turn it off when a status line tool shows them (see [Status line tools](#status-line-tools-ccstatusline)) |
|
|
156
|
+
| `showTier` | `true` | The tier, `deep` |
|
|
157
|
+
| `showConfidence` | `true` | The judge's confidence, `0.91` (`n/d` without one) |
|
|
158
|
+
| `showModelEffort` | `true` | The model and effort in effect, `opus/xhigh` (`→ opus/xhigh` when tiergear just applied them) |
|
|
159
|
+
| `showReason` | `true` | Why nothing changed, `unchanged (same tier)` |
|
|
160
|
+
| `showTierButtons` | `true` | Show the off and tier buttons above the prompt |
|
|
161
|
+
| `showFloor` | `true` | Show the floor picker above the prompt |
|
|
162
|
+
|
|
163
|
+
## Status band
|
|
164
|
+
|
|
165
|
+
tiergear shows its state in one row just above the prompt: a line, the **Tier** picker (off and a button per tier), the **Floor** picker, and **[ Recent ]**. The choice in effect is bracketed; the others are dim.
|
|
166
|
+
|
|
167
|
+
```
|
|
168
|
+
tiergear · deep 0.91 → opus/xhigh | Tier: off ~~trivial~~ quick standard [deep] max | Floor: quick | [ Recent ]
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
It doesn't use the status line below the prompt. Before anything is decided the line reads just `tiergear` and no button is bracketed.
|
|
172
|
+
|
|
173
|
+
The line starts with `tiergear ·` and shows the model and effort the session is running on as `model/effort`; a model without effort shows `-`.
|
|
174
|
+
|
|
175
|
+
- Applied: `tiergear · deep 0.91 → opus/xhigh`.
|
|
176
|
+
- Unchanged: `tiergear · standard 0.62 · sonnet/medium · unchanged (<reason>)`. When tiergear isn't overriding anything (low confidence, off), this is the session's own model and effort.
|
|
177
|
+
- The line is set when a prompt is judged, then refreshed with the values the engine reports when the turn starts. Before the session's first turn the values may not be known yet: an unchanged line then leaves them out, and an applied line shows only the effort if tiergear isn't setting the model.
|
|
178
|
+
- `unset` appears when no tier has been decided yet, and `n/d` takes the place of the confidence when the judge gave no answer (as after a pick from the band).
|
|
179
|
+
|
|
180
|
+
Reasons for no change: `no answer` (no judge response), `low confidence`, `same tier`, `paused` (off is chosen), `manual tier` (the first prompt runs on a tier picked before it), `floor set` (the floor was changed by hand), `at floor` (can't go lower), `at ceiling` (can't go higher than the launch's `--max-tier`), `easier step N/M` (Nth lowering candidate, M needed), `stuck at max`.
|
|
181
|
+
|
|
182
|
+
### Tier buttons and floor
|
|
183
|
+
|
|
184
|
+
- **Floor** picks the lowest tier the session may go to. A tier above the current one raises the session to it at once. A tier under the floor is struck through and can't be clicked; lower the floor first. A floor above a launch's `--max-tier` stops at it. Set before the first prompt, the first decided tier starts no lower than it. The floor picker is a select: click or focus it, then use the arrow keys and Enter.
|
|
185
|
+
|
|
186
|
+
- One click picks. A tier applies from the next request, even in the middle of a running turn. Mid-session only the effort changes (from the session model's column in table B), unless `switchModelMidSession` is on. The judge goes on from the picked tier by the usual rules. A pick below the floor lowers the floor to it. A pick may go above a launch's `--max-tier`, but the judge still won't raise past it.
|
|
187
|
+
- Picked before the first prompt, a tier works like a launch tier: the first turn runs on its table A model and effort without asking the judge, and the judge takes over from the second prompt.
|
|
188
|
+
- **off** withdraws the model and effort tiergear applies, so the session runs on its own (startup flags, `/model`, `/effort`), and stops asking the judge. `[off]` stays bracketed until you pick a tier.
|
|
189
|
+
- Picking a tier while off turns tiergear back on at that tier, on the session's own model with the effort from the tier, and the judge is asked again from the next prompt. That prompt becomes the task the judge reads.
|
|
190
|
+
|
|
191
|
+
### Recent decisions
|
|
192
|
+
|
|
193
|
+
Press **[ Recent ]** to open a pane listing this session's decisions, newest first, and press it again to close it. `/tiergear` opens the pane too. To hide the button, turn off `showRecentButton` in the plugin options. Each part of the band has its own option (`showPrefix`, `showStatusText` and the line's parts, `showTierButtons`, `showFloor`, `showRecentButton`); with all of them off the band draws nothing.
|
|
194
|
+
|
|
195
|
+
```
|
|
196
|
+
12:11 judge quick 0.44 → hold deep (low confidence) · opus/xhigh
|
|
197
|
+
11:16 judge deep 0.52 → up deep (harder step) · opus/xhigh
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
Each line is: time, what the judge proposed and its confidence, what tiergear did and the resulting tier (with the reason), and the model/effort tiergear applied (`session` when it applied nothing). `judge skipped` means the judge wasn't asked (a launch floor, a picked tier, off, a paused judge); a failure shows its reason, like `judge timed out after 2000ms`. A pick from the band (a tier or off) reads `manual` instead, as in `manual → set quick (manual tier) · opus/low`. The pane reads the decision log, so it still works after a plugin reload; the line comes back with the next judged prompt.
|
|
201
|
+
|
|
202
|
+
### Status line tools (ccstatusline)
|
|
203
|
+
|
|
204
|
+
Claude Code's status JSON reports the session's own model and effort, not what tiergear applies per request, so a status line built from it misses tiergear's changes. The hook writes what the band shows to `~/.local/state/tiergear/status/<session>.json`, and `tiergear status` prints it:
|
|
205
|
+
|
|
206
|
+
```bash
|
|
207
|
+
tiergear status [tier|state|model|effort] [--session <id>] [--json] [--format <template>]
|
|
208
|
+
```
|
|
209
|
+
|
|
210
|
+
- The session is `--session`, else the `session_id` of the status JSON piped on stdin (what a status line command receives), else the session updated last.
|
|
211
|
+
- Before tiergear has a value for the session (a new session, the first turn), the model and effort come from the status JSON on stdin, the session's own; the tier stays empty. A value tiergear doesn't know yet is filled the same way.
|
|
212
|
+
- Default output: `deep · opus/xhigh`, `paused · sonnet/medium` when paused; nothing (exit 0) when neither tiergear nor the status JSON on stdin has a value, so a widget hides. The stdin fallback applies only without `--session`.
|
|
213
|
+
- A field prints that value alone, for a widget of its own: `tiergear status model` prints `opus`, `tiergear status effort` prints `xhigh`, `tiergear status tier` prints `deep`, `tiergear status state` prints `auto` or `paused`. An unknown value prints nothing. A field wins over `--json` and `--format`.
|
|
214
|
+
- `--format` fills `{tier}`, `{model}`, `{modelName}` (as Claude Code names it, `Opus 5.5`), `{effort}`, `{state}` (`auto` or `paused`) and `{line}` (the band's text); an unknown value is `-` (`unset` for the tier), and a template with no known value in it prints nothing. `--json` prints the whole record, or `null`.
|
|
215
|
+
|
|
216
|
+
With the line shown there, turn off `showStatusText` to keep the band to its buttons.
|
|
217
|
+
|
|
218
|
+
In [ccstatusline](https://github.com/sirmalloc/ccstatusline), add a **Custom Command** widget with the command `tiergear status`, or one widget per value (`tiergear status model`, `tiergear status effort`, ...) to color them apart (the CLI must be installed, see [CLI](#cli-optional)). It runs in about 50ms, well within the widget's default 1000ms timeout. The value follows a pick or a judged prompt at the next status line refresh.
|
|
219
|
+
|
|
220
|
+
## CLI
|
|
221
|
+
|
|
222
|
+
```bash
|
|
223
|
+
tiergear launch "<brief>" [--agent claude|codex] [--worktree <path>] [--min-tier <tier>] [--max-tier <tier>] [--judge jev|laya|kev] [--judge-url <url>] [--judge-model <name>]
|
|
224
|
+
tiergear orca-spawn "<brief>" --name <task> [--agent claude|codex] [--repo <dir>] [--base-branch <ref>] [--min-tier <tier>] [--max-tier <tier>] [--judge ...]
|
|
225
|
+
tiergear stats [days]
|
|
226
|
+
tiergear status [tier|state|model|effort] [--session <id>] [--json] [--format <template>]
|
|
227
|
+
```
|
|
228
|
+
|
|
229
|
+
- `launch`: judges the brief and prints the command to run (e.g. `claude --model opus --effort xhigh`). With Claude, `--worktree` writes a floor for that path. The path is stored as an absolute real path (symlinks resolved), so a relative path still works for a session opened in that folder. With `--agent codex`, floors are Claude-only, so none is written and a one-line note is printed instead.
|
|
230
|
+
- `orca-spawn`: needs Orca, the multi-agent IDE, with its `orca` CLI on your PATH (or named by `ORCA_CLI_COMMAND`). It creates an Orca worktree, writes the floor, and starts the agent there with the judged model and effort. Prints the result as JSON. `--base-branch` picks the ref the worktree starts from; without it Orca uses the repo's default base, which may be a remote branch behind your local one. A brief that starts with `!` or `/` (which the agent's prompt would run as a shell or slash command) or holds control characters other than line breaks and tabs is refused before the judge is asked.
|
|
231
|
+
- **Inside an orchestration Run** (run from the coordinator terminal after `orca orchestration run-create`): starts the agent with `orca orchestration worker-start`, so the worker gets Orca's lifecycle preamble and reports `worker_done` to the Run. The JSON includes `dispatch` (`runId`, `taskId`, `dispatchId`, `handle`). A failed `worker-start` exits non-zero with Orca's error; don't rerun it blindly, since Orca may have left resources behind.
|
|
232
|
+
- **Without a Run**: creates a terminal, waits for the agent, then types the brief as is (no preamble, `dispatch` is `null`). If Claude asks whether to trust the folder, **`orca-spawn` does not approve it for you.** Approve it yourself in Orca within 120 seconds. If the agent isn't ready, the brief is **not sent** (exit code 1) and a fallback shell may be left in the worktree.
|
|
233
|
+
- The floor is written before the agent's first prompt in both cases. Without it, the first turn would be judged on Orca's preamble, which usually gives low confidence.
|
|
234
|
+
- `stats`: number of recorded decisions, counts by change type, and response rate and average latency per judge (default 7 days).
|
|
235
|
+
- `status`: what a session runs on now, for status line tools; see [Status line tools](#status-line-tools-ccstatusline).
|
|
236
|
+
- `--min-tier <tier>` and `--max-tier <tier>` (either or both; trivial, quick, standard, deep or max) bound the launch. The judge's tier is clamped into the range, and tables A and B pick the model and effort from the clamped tier. The floor written for the session holds the clamped tier as its floor and `--max-tier` as its ceiling, so later turns stay in the range too. An unknown tier or a minimum above the maximum exits 2 before the judge is asked or a worktree is created. With `--agent codex` the range sets the starting tier only (floors are Claude-only), and a one-line note says so.
|
|
237
|
+
- Both commands report the judge's tier beside the one applied. `launch` prints `tiergear: judged <tier>, applied <tier>` on stderr (`judged none` when the judge decided nothing); stdout stays the command alone. `orca-spawn`'s JSON has `tier`, the tier applied after the range, and `judgedTier`, the judge's tier before it (`null` when the judge failed, gave no tier, or had confidence under 0.5).
|
|
238
|
+
- Without a range, a floor is written only when the judge actually decided. If the judge failed, didn't answer, or had low confidence and standard was used instead, no floor is written and a warning is printed. With `--min-tier` or `--max-tier`, standard is clamped into the range and the floor is written anyway, since the range is yours, not the judge's. So with only `--max-tier` and a failed judge, the session's floor is standard (or `--max-tier`, if lower).
|
|
239
|
+
- The CLI's judge timeout is 5000ms.
|
|
240
|
+
|
|
241
|
+
The CLI can't read plugin options, so it looks up settings from flags, then environment variables, then the preset.
|
|
242
|
+
|
|
243
|
+
| Variable | Meaning |
|
|
244
|
+
| --- | --- |
|
|
245
|
+
| `TIERGEAR_JUDGE` | Judge name (default jev) |
|
|
246
|
+
| `TIERGEAR_JUDGE_URL` | Judge address |
|
|
247
|
+
| `TIERGEAR_JUDGE_MODEL` | Judge model |
|
|
248
|
+
| `TIERGEAR_JUDGE_API_KEY` | Key (falls back to the preset's key env var, only for the preset's own address) |
|
|
249
|
+
| `ORCA_CLI_COMMAND` | The `orca` executable `orca-spawn` runs (default `orca` on your PATH) |
|
|
250
|
+
|
|
251
|
+
## Where things are stored
|
|
252
|
+
|
|
253
|
+
| What | Location |
|
|
254
|
+
| --- | --- |
|
|
255
|
+
| Tables | `~/.config/tiergear/tables.json` |
|
|
256
|
+
| Floors (valid 24 hours) | `~/.local/state/tiergear/floors/<fnv1a(worktree)>.json` |
|
|
257
|
+
| Decision log (1000 lines per file; one file per session, not removed) | `~/.local/state/tiergear/decisions/<name>.jsonl` |
|
|
258
|
+
| Status for status line tools (one small file per session, not removed) | `~/.local/state/tiergear/status/<session>.json` |
|
|
259
|
+
| Session records (kept 7 days, newest 200) | Plugin `$.store` under `session:<id>` (the first prompt is truncated to 2000 characters, and dropped while a session is paused) |
|
|
260
|
+
|
|
261
|
+
Raw prompt text is never written to the logs.
|
|
262
|
+
|
|
263
|
+
## Cost
|
|
264
|
+
|
|
265
|
+
- Changing the model and changing effort mid-session both invalidate the messages prompt cache (per Anthropic's docs). The first turn has no cache, so it costs nothing.
|
|
266
|
+
- That's why the default mid-session change is effort only, and lowering only happens after repeated high-confidence turns, to keep changes rare.
|
|
267
|
+
- Raising from a model without effort, like haiku if you set it, switches the model and breaks the cache once.
|
|
268
|
+
- The judge call itself adds cost and latency to every judged prompt (within the timeouts above).
|
|
269
|
+
|
|
270
|
+
## Data sent to the judge
|
|
271
|
+
|
|
272
|
+
- First turn: the first prompt (only the start and end if long).
|
|
273
|
+
- Later turns: the next prompt, the last 3 exchanges, the first prompt (up to 500 characters), the number of files edited, the repeated failure count, and the current tier and effort.
|
|
274
|
+
- An exchange is one prompt, the assistant's last words before the next prompt, and the names of the tools it used. Tool calls and their results are folded into it, so a busy turn doesn't push the conversation out of view.
|
|
275
|
+
- Recent text gets more room: the reply the next prompt answers is sent up to 1500 characters, keeping mostly its end, where a question usually sits. Prompts and older replies get 400 characters each. Longer text keeps its start and end.
|
|
276
|
+
- jev sends data externally (TypeSafe). laya and kev stay on your local server (if you run kev on Modal, data goes there).
|
|
277
|
+
- Prompts that aren't judged (see "Which prompts are judged" above) never go to the judge as the prompt being judged, but one can still appear among the recent exchanges: a task notification or a message from another session is a prompt in the conversation too. Nothing is masked, so with jev that text leaves your machine.
|
|
278
|
+
|
|
279
|
+
## Limitations
|
|
280
|
+
|
|
281
|
+
- Codex gets its model and effort only at launch (`launch`/`orca-spawn`). Mid-session adjustment works only in Claude Code.
|
|
282
|
+
- Hooks are an early-access feature, so the contract can change with Claude Code updates.
|
|
283
|
+
- The confidence thresholds (0.5, 0.85, 0.6) are tuned for Jev. For other judges, check response rate and latency with `tiergear stats` and adjust the options.
|
|
284
|
+
- Subagent requests are left alone; only main-loop requests are changed.
|
|
285
|
+
|
|
286
|
+
## Development
|
|
287
|
+
|
|
288
|
+
`npm test`, `npm run typecheck` and `npm run validate:plugin` check a working copy. `types/claude-code.d.ts` is the hook API declaration Claude Code 2.1.289 generated; after a Claude Code update, replace it with the `.claude-plugin/types/claude-code/index.d.ts` the engine writes next to a loaded plugin.
|
|
289
|
+
|
|
290
|
+
To release, set the same version in `package.json` and `.claude-plugin/plugin.json`, push, then publish a GitHub release tagged `v<version>` (`gh release create v<version> --generate-notes`). `.github/workflows/publish.yml` tests it and publishes it to npm through npm trusted publishing, with no token stored.
|
|
291
|
+
|
|
292
|
+
## License
|
|
293
|
+
|
|
294
|
+
MIT
|
package/dist/cli/bin.js
ADDED
|
@@ -0,0 +1,43 @@
|
|
|
1
|
+
import { mkdir, readFile, readdir, realpath, writeFile } from 'node:fs/promises';
|
|
2
|
+
import { dirname, join, resolve } from 'node:path';
|
|
3
|
+
import { floorPath, normalizePath, serializeFloor } from '../core/floor.js';
|
|
4
|
+
import { appendLogLine, decisionLogPath, decisionsDir, parseLogLines } from '../core/log.js';
|
|
5
|
+
import { DEFAULT_TABLES, parseTablesFile, tablesPath } from '../core/tables.js';
|
|
6
|
+
// The hook matches the session's cwd, which Claude Code reports as an absolute physical path.
|
|
7
|
+
export async function realWorktree(worktree) {
|
|
8
|
+
const absolute = resolve(worktree);
|
|
9
|
+
return realpath(absolute).catch(() => absolute);
|
|
10
|
+
}
|
|
11
|
+
export async function writeFloorFile(home, worktree, floor, now) {
|
|
12
|
+
const real = normalizePath(await realWorktree(worktree));
|
|
13
|
+
const path = floorPath(home, real);
|
|
14
|
+
await mkdir(dirname(path), { recursive: true });
|
|
15
|
+
await writeFile(path, serializeFloor({ worktree: real, tier: floor.tier, createdAt: now, ceiling: floor.ceiling }));
|
|
16
|
+
return path;
|
|
17
|
+
}
|
|
18
|
+
export async function appendDecision(home, name, entry) {
|
|
19
|
+
const path = decisionLogPath(home, name);
|
|
20
|
+
await mkdir(dirname(path), { recursive: true });
|
|
21
|
+
const existing = await readFile(path, 'utf8').catch(() => '');
|
|
22
|
+
await writeFile(path, appendLogLine(existing, JSON.stringify(entry)));
|
|
23
|
+
}
|
|
24
|
+
export async function readDecisions(home) {
|
|
25
|
+
const dir = decisionsDir(home);
|
|
26
|
+
const names = await readdir(dir).catch(() => []);
|
|
27
|
+
const entries = [];
|
|
28
|
+
for (const name of names) {
|
|
29
|
+
if (!name.endsWith('.jsonl'))
|
|
30
|
+
continue;
|
|
31
|
+
entries.push(...parseLogLines(await readFile(join(dir, name), 'utf8')));
|
|
32
|
+
}
|
|
33
|
+
return entries;
|
|
34
|
+
}
|
|
35
|
+
export async function readTables(home) {
|
|
36
|
+
const text = await readFile(tablesPath(home), 'utf8').catch(() => null);
|
|
37
|
+
if (text === null)
|
|
38
|
+
return DEFAULT_TABLES;
|
|
39
|
+
const parsed = parseTablesFile(text);
|
|
40
|
+
if (!parsed)
|
|
41
|
+
console.error('tiergear: tables.json ignored: invalid JSON or values; using the default tables');
|
|
42
|
+
return parsed ?? DEFAULT_TABLES;
|
|
43
|
+
}
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
import { JUDGE_PRESETS, isAllowedJudgeUrl, isJudgeName, presetKeyApplies } from '../core/judges/presets.js';
|
|
2
|
+
// The CLI cannot read the plugin's options, so it takes flags first, then TIERGEAR_* variables, then the preset.
|
|
3
|
+
export function cliJudgeOptions(flags, env) {
|
|
4
|
+
const name = flags.judge ?? env['TIERGEAR_JUDGE'] ?? 'jev';
|
|
5
|
+
if (!isJudgeName(name))
|
|
6
|
+
throw new Error(`unknown judge ${name} (use jev, laya or kev)`);
|
|
7
|
+
const preset = JUDGE_PRESETS[name];
|
|
8
|
+
const baseUrl = flags.url ?? env['TIERGEAR_JUDGE_URL'] ?? preset.baseUrl;
|
|
9
|
+
if (!isAllowedJudgeUrl(baseUrl))
|
|
10
|
+
throw new Error(`judge URL ${baseUrl} must be https, or http to localhost`);
|
|
11
|
+
const presetKey = presetKeyApplies(preset, baseUrl) ? env[preset.keyEnv] : undefined;
|
|
12
|
+
return {
|
|
13
|
+
name,
|
|
14
|
+
baseUrl,
|
|
15
|
+
model: flags.model ?? env['TIERGEAR_JUDGE_MODEL'] ?? preset.model,
|
|
16
|
+
apiKey: env['TIERGEAR_JUDGE_API_KEY'] || presetKey || undefined,
|
|
17
|
+
keyRequired: preset.keyRequired,
|
|
18
|
+
};
|
|
19
|
+
}
|
|
@@ -0,0 +1,50 @@
|
|
|
1
|
+
import { firstTurnState } from '../core/state.js';
|
|
2
|
+
import { firstTarget } from '../core/tables.js';
|
|
3
|
+
import { TIER_ORDER, clampTier, isTier, launchCommand, tierRank } from '../core/tiers.js';
|
|
4
|
+
export const CLI_TIMEOUT_MS = 5000;
|
|
5
|
+
const MIN_CONFIDENCE = 0.5;
|
|
6
|
+
export function parseTierRange(min, max) {
|
|
7
|
+
for (const [flag, value] of [['--min-tier', min], ['--max-tier', max]]) {
|
|
8
|
+
if (value !== undefined && !isTier(value))
|
|
9
|
+
return { ok: false, error: `${flag} must be one of ${TIER_ORDER.join('|')}, not "${value}"` };
|
|
10
|
+
}
|
|
11
|
+
const range = {};
|
|
12
|
+
if (isTier(min))
|
|
13
|
+
range.min = min;
|
|
14
|
+
if (isTier(max))
|
|
15
|
+
range.max = max;
|
|
16
|
+
if (range.min && range.max && tierRank(range.min) > tierRank(range.max))
|
|
17
|
+
return { ok: false, error: `--min-tier ${range.min} is above --max-tier ${range.max}` };
|
|
18
|
+
return { ok: true, range };
|
|
19
|
+
}
|
|
20
|
+
export async function planLaunch(params) {
|
|
21
|
+
const range = params.range ?? {};
|
|
22
|
+
const fallback = clampTier('standard', range);
|
|
23
|
+
const started = params.now();
|
|
24
|
+
const result = await params.judge.ask({ state: firstTurnState(params.brief), withStuck: false, timeoutMs: CLI_TIMEOUT_MS });
|
|
25
|
+
const ms = params.now() - started;
|
|
26
|
+
let judgedTier = null;
|
|
27
|
+
let confidence = null;
|
|
28
|
+
let warning = null;
|
|
29
|
+
let outcome = 'ok';
|
|
30
|
+
if (!result.ok) {
|
|
31
|
+
outcome = result.reason;
|
|
32
|
+
warning = `${params.judge.name} unavailable (${result.reason}); using ${fallback}`;
|
|
33
|
+
}
|
|
34
|
+
else if (!result.verdict.tier) {
|
|
35
|
+
warning = `${params.judge.name} gave no tier; using ${fallback}`;
|
|
36
|
+
}
|
|
37
|
+
else if (result.verdict.tier.confidence < MIN_CONFIDENCE) {
|
|
38
|
+
confidence = result.verdict.tier.confidence;
|
|
39
|
+
warning = `low confidence ${confidence.toFixed(2)} for ${result.verdict.tier.tier}; using ${fallback}`;
|
|
40
|
+
}
|
|
41
|
+
else {
|
|
42
|
+
judgedTier = result.verdict.tier.tier;
|
|
43
|
+
confidence = result.verdict.tier.confidence;
|
|
44
|
+
}
|
|
45
|
+
const tier = judgedTier === null ? fallback : clampTier(judgedTier, range);
|
|
46
|
+
// A range a person gave is not the judge's call, so it is floored even when the judge decided nothing.
|
|
47
|
+
const floor = warning === null || range.min || range.max ? { tier, ceiling: range.max } : null;
|
|
48
|
+
const target = firstTarget(params.tables, params.harness, tier);
|
|
49
|
+
return { tier, judgedTier, confidence, target, command: launchCommand(params.harness, target), warning, outcome, ms, floor };
|
|
50
|
+
}
|
package/dist/cli/main.js
ADDED
|
@@ -0,0 +1,145 @@
|
|
|
1
|
+
import { parseArgs } from 'node:util';
|
|
2
|
+
import { createSystemOneJudge } from '../core/judges/systemone.js';
|
|
3
|
+
import { appendDecision, readDecisions, readTables, writeFloorFile } from './files.js';
|
|
4
|
+
import { cliJudgeOptions } from './judge.js';
|
|
5
|
+
import { parseTierRange, planLaunch } from './launch.js';
|
|
6
|
+
import { createOrcaExec } from './orca.js';
|
|
7
|
+
import { homeDir, nodeSleep, nodeTransport, readStdin } from './node.js';
|
|
8
|
+
import { briefProblem, spawnWorker } from './spawn.js';
|
|
9
|
+
import { summarize } from './stats.js';
|
|
10
|
+
import { statusCommand } from './status.js';
|
|
11
|
+
import { isStatusField } from '../core/status.js';
|
|
12
|
+
export const USAGE = `usage:
|
|
13
|
+
tiergear launch "<brief>" [--agent claude|codex] [--worktree <path>] [--min-tier <tier>] [--max-tier <tier>] [--judge jev|laya|kev] [--judge-url <url>] [--judge-model <name>]
|
|
14
|
+
tiergear orca-spawn "<brief>" --name <task> [--agent claude|codex] [--repo <dir>] [--base-branch <ref>] [--min-tier <tier>] [--max-tier <tier>] [--judge ...]
|
|
15
|
+
tiergear stats [days]
|
|
16
|
+
tiergear status [tier|state|model|effort] [--session <id>] [--json] [--format <template>]
|
|
17
|
+
<tier> is one of trivial|quick|standard|deep|max`;
|
|
18
|
+
async function plan(command, brief, harness, home, flags, range) {
|
|
19
|
+
const options = cliJudgeOptions(flags, process.env);
|
|
20
|
+
const judge = createSystemOneJudge({ ...options, transport: nodeTransport, sleep: nodeSleep });
|
|
21
|
+
const result = await planLaunch({ brief, harness, judge, tables: await readTables(home), now: Date.now, range });
|
|
22
|
+
if (result.warning)
|
|
23
|
+
console.error(`tiergear: ${result.warning}`);
|
|
24
|
+
try {
|
|
25
|
+
await appendDecision(home, `cli-${new Date().toISOString().slice(0, 10)}`, {
|
|
26
|
+
at: Date.now(),
|
|
27
|
+
source: 'cli',
|
|
28
|
+
session: command,
|
|
29
|
+
phase: 'launch',
|
|
30
|
+
judge: options.name,
|
|
31
|
+
tier: result.tier,
|
|
32
|
+
change: 'set',
|
|
33
|
+
confidence: result.confidence,
|
|
34
|
+
stuck: null,
|
|
35
|
+
outcome: result.outcome,
|
|
36
|
+
ms: result.ms,
|
|
37
|
+
applied: { model: result.target.model, effort: result.target.effort },
|
|
38
|
+
reason: command,
|
|
39
|
+
});
|
|
40
|
+
}
|
|
41
|
+
catch (error) {
|
|
42
|
+
console.error(`tiergear: decision log not written: ${error instanceof Error ? error.message : String(error)}`);
|
|
43
|
+
}
|
|
44
|
+
return result;
|
|
45
|
+
}
|
|
46
|
+
export async function main(argv) {
|
|
47
|
+
const [command, ...rest] = argv;
|
|
48
|
+
const { values, positionals } = parseArgs({
|
|
49
|
+
args: rest,
|
|
50
|
+
allowPositionals: true,
|
|
51
|
+
options: {
|
|
52
|
+
agent: { type: 'string', default: 'claude' },
|
|
53
|
+
worktree: { type: 'string' },
|
|
54
|
+
name: { type: 'string' },
|
|
55
|
+
repo: { type: 'string' },
|
|
56
|
+
'base-branch': { type: 'string' },
|
|
57
|
+
judge: { type: 'string' },
|
|
58
|
+
'judge-url': { type: 'string' },
|
|
59
|
+
'judge-model': { type: 'string' },
|
|
60
|
+
'min-tier': { type: 'string' },
|
|
61
|
+
'max-tier': { type: 'string' },
|
|
62
|
+
session: { type: 'string' },
|
|
63
|
+
json: { type: 'boolean', default: false },
|
|
64
|
+
format: { type: 'string' },
|
|
65
|
+
},
|
|
66
|
+
});
|
|
67
|
+
const home = homeDir();
|
|
68
|
+
if (command === 'status') {
|
|
69
|
+
const field = positionals[0];
|
|
70
|
+
if (positionals.length > 1 || (field !== undefined && !isStatusField(field))) {
|
|
71
|
+
console.error(USAGE);
|
|
72
|
+
return 2;
|
|
73
|
+
}
|
|
74
|
+
const out = await statusCommand({ home, session: values.session, stdin: () => readStdin(500), json: values.json, format: values.format, field });
|
|
75
|
+
if (out)
|
|
76
|
+
console.log(out);
|
|
77
|
+
return 0;
|
|
78
|
+
}
|
|
79
|
+
if (command === 'stats') {
|
|
80
|
+
const days = Number(positionals[0] ?? 7);
|
|
81
|
+
console.log(summarize(await readDecisions(home), Date.now() - days * 86_400_000));
|
|
82
|
+
return 0;
|
|
83
|
+
}
|
|
84
|
+
const brief = positionals.join(' ').trim();
|
|
85
|
+
const harness = values.agent;
|
|
86
|
+
if ((command !== 'launch' && command !== 'orca-spawn') || !brief || (harness !== 'claude' && harness !== 'codex')) {
|
|
87
|
+
console.error(USAGE);
|
|
88
|
+
return 2;
|
|
89
|
+
}
|
|
90
|
+
const flags = { judge: values.judge, url: values['judge-url'], model: values['judge-model'] };
|
|
91
|
+
const parsedRange = parseTierRange(values['min-tier'], values['max-tier']);
|
|
92
|
+
if (!parsedRange.ok) {
|
|
93
|
+
console.error(`tiergear: ${parsedRange.error}`);
|
|
94
|
+
return 2;
|
|
95
|
+
}
|
|
96
|
+
const { range } = parsedRange;
|
|
97
|
+
const ranged = range.min !== undefined || range.max !== undefined;
|
|
98
|
+
if (command === 'launch') {
|
|
99
|
+
const launch = await plan(command, brief, harness, home, flags, range);
|
|
100
|
+
// stdout carries only the command, so the tiers go to stderr.
|
|
101
|
+
console.error(`tiergear: judged ${launch.judgedTier ?? 'none'}, applied ${launch.tier}`);
|
|
102
|
+
if (harness === 'codex' && values.worktree)
|
|
103
|
+
console.error('tiergear: --worktree ignored: floors are Claude-only');
|
|
104
|
+
else if (harness === 'codex' && ranged)
|
|
105
|
+
console.error('tiergear: the tier range sets the start only: floors are Claude-only');
|
|
106
|
+
if (values.worktree && harness === 'claude') {
|
|
107
|
+
if (launch.floor)
|
|
108
|
+
await writeFloorFile(home, values.worktree, launch.floor, Date.now());
|
|
109
|
+
else
|
|
110
|
+
console.error('tiergear: no floor written (judge fallback)');
|
|
111
|
+
}
|
|
112
|
+
console.log(launch.command);
|
|
113
|
+
return 0;
|
|
114
|
+
}
|
|
115
|
+
if (!values.name) {
|
|
116
|
+
console.error(USAGE);
|
|
117
|
+
return 2;
|
|
118
|
+
}
|
|
119
|
+
const problem = briefProblem(brief);
|
|
120
|
+
if (problem) {
|
|
121
|
+
console.error(`tiergear: ${problem}`);
|
|
122
|
+
return 2;
|
|
123
|
+
}
|
|
124
|
+
const spawnPlan = await plan(command, brief, harness, home, flags, range);
|
|
125
|
+
if (harness === 'codex' && ranged)
|
|
126
|
+
console.error('tiergear: the tier range sets the start only: floors are Claude-only');
|
|
127
|
+
const result = await spawnWorker({
|
|
128
|
+
brief,
|
|
129
|
+
name: values.name,
|
|
130
|
+
harness,
|
|
131
|
+
repoDir: values.repo ?? process.cwd(),
|
|
132
|
+
baseBranch: values['base-branch'],
|
|
133
|
+
plan: spawnPlan,
|
|
134
|
+
orca: createOrcaExec(),
|
|
135
|
+
writeFloor: async (path, floor) => {
|
|
136
|
+
await writeFloorFile(home, path, floor, Date.now());
|
|
137
|
+
},
|
|
138
|
+
log: (line) => console.error(`tiergear: ${line}`),
|
|
139
|
+
});
|
|
140
|
+
if (result.status === 'not-started') {
|
|
141
|
+
console.error('tiergear: the agent did not become ready; the brief was not sent. A fallback shell may remain in the worktree.');
|
|
142
|
+
}
|
|
143
|
+
console.log(JSON.stringify({ ...result, tier: spawnPlan.tier, judgedTier: spawnPlan.judgedTier, command: spawnPlan.command }));
|
|
144
|
+
return result.status === 'sent' ? 0 : 1;
|
|
145
|
+
}
|
package/dist/cli/node.js
ADDED
|
@@ -0,0 +1,53 @@
|
|
|
1
|
+
import { homedir } from 'node:os';
|
|
2
|
+
import { MAX_RESPONSE_CHARS } from '../core/judges/systemone.js';
|
|
3
|
+
import { CLI_TIMEOUT_MS } from './launch.js';
|
|
4
|
+
// The judge's deadline only stops waiting; aborting the request is what lets the process exit, and the
|
|
5
|
+
// byte limit keeps a runaway body from being read whole.
|
|
6
|
+
export function createNodeTransport(deadlineMs, maxBytes = MAX_RESPONSE_CHARS) {
|
|
7
|
+
return async (url, init) => {
|
|
8
|
+
const response = await fetch(url, { ...init, signal: AbortSignal.timeout(deadlineMs) });
|
|
9
|
+
const reader = response.body?.getReader();
|
|
10
|
+
const chunks = [];
|
|
11
|
+
let size = 0;
|
|
12
|
+
for (;;) {
|
|
13
|
+
if (!reader)
|
|
14
|
+
break;
|
|
15
|
+
const { done, value } = await reader.read();
|
|
16
|
+
if (done)
|
|
17
|
+
break;
|
|
18
|
+
size += value.byteLength;
|
|
19
|
+
if (size > maxBytes) {
|
|
20
|
+
await reader.cancel();
|
|
21
|
+
throw new Error('judge response too large');
|
|
22
|
+
}
|
|
23
|
+
chunks.push(value);
|
|
24
|
+
}
|
|
25
|
+
return { status: response.status, ok: response.ok, text: Buffer.concat(chunks).toString('utf8') };
|
|
26
|
+
};
|
|
27
|
+
}
|
|
28
|
+
export const nodeTransport = createNodeTransport(CLI_TIMEOUT_MS + 1000);
|
|
29
|
+
// unref so a pending deadline does not keep the CLI alive after the judge answered.
|
|
30
|
+
export const nodeSleep = (ms) => new Promise((resolve) => {
|
|
31
|
+
setTimeout(resolve, ms).unref();
|
|
32
|
+
});
|
|
33
|
+
export function homeDir() {
|
|
34
|
+
return process.env['HOME'] ?? homedir();
|
|
35
|
+
}
|
|
36
|
+
// Reads what is piped in, giving up after timeoutMs so a stdin left open never hangs the command.
|
|
37
|
+
export function readStdin(timeoutMs) {
|
|
38
|
+
if (process.stdin.isTTY)
|
|
39
|
+
return Promise.resolve(null);
|
|
40
|
+
return new Promise((resolve) => {
|
|
41
|
+
let text = '';
|
|
42
|
+
const done = (value) => {
|
|
43
|
+
clearTimeout(timer);
|
|
44
|
+
process.stdin.destroy();
|
|
45
|
+
resolve(value);
|
|
46
|
+
};
|
|
47
|
+
const timer = setTimeout(() => done(text || null), timeoutMs);
|
|
48
|
+
process.stdin.setEncoding('utf8');
|
|
49
|
+
process.stdin.on('data', (chunk) => (text += chunk));
|
|
50
|
+
process.stdin.on('end', () => done(text));
|
|
51
|
+
process.stdin.on('error', () => done(null));
|
|
52
|
+
});
|
|
53
|
+
}
|