vigiles 12.2.0 → 12.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +125 -162
- package/dist/adapters/claude-code/run-scripts.d.ts +41 -0
- package/dist/adapters/claude-code/run-scripts.js +26 -0
- package/dist/cli.js +49 -4
- package/dist/eval.d.ts +15 -0
- package/dist/eval.js +43 -0
- package/package.json +2 -2
package/README.md
CHANGED
|
@@ -1,88 +1,66 @@
|
|
|
1
1
|
<!--
|
|
2
2
|
README DIRECTION — read before editing; keep changes aligned.
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
"
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
the "bite" over a clean A 92 on 2026-06-29). No dialect-drift banner (HTML
|
|
65
|
-
report is terminal-banner-free by design). Re-render via headless Chromium on
|
|
66
|
-
the React report if the UI changes (recipe: copy a trifecta-bearing plugin to
|
|
67
|
-
my-plugin/, `node dist/cli.js audit my-plugin --no-json --no-serve`,
|
|
68
|
-
headless_shell `--window-size=820,1180 --force-device-scale-factor=2
|
|
69
|
-
--screenshot` on vigiles-report.html, then `rm -rf my-plugin
|
|
70
|
-
vigiles-report.html`). (vigiles-demo.gif was removed
|
|
71
|
-
from Proof 1 — it rendered as a frozen half-typed terminal and was redundant
|
|
72
|
-
with the code block; if a lint demo returns, it belongs in the Lint section
|
|
73
|
-
with a non-frozen asset.)
|
|
74
|
-
|
|
75
|
-
READABILITY (the 2026-06-29 pass — why this reads the way it does):
|
|
76
|
-
A. ONE bold per block, on the single phrase the eye should catch. Bold
|
|
77
|
-
everywhere = bold nowhere. Link CTAs may stay bold (they're navigation).
|
|
78
|
-
B. ONE idea per sentence. No em-dash clause-chains, no stacked parentheticals.
|
|
79
|
-
If a clause needs a paren, cut it or give it its own line.
|
|
80
|
-
C. PLAIN words in every LEAD; push jargon (rings, recall/precision,
|
|
81
|
-
interceptTools, selector, deterministic) into the linked docs. A skimmer who
|
|
82
|
-
lives in Claude Code still may not know the vocabulary.
|
|
83
|
-
D. SHOW via the proofs/code blocks; don't stack adjectives ("real, popular,
|
|
84
|
-
free, model-less") on top of what the block already proves.
|
|
85
|
-
E. SELL the outcome before the mechanism; the instruments come AFTER the proofs.
|
|
3
|
+
Front door + marketing asset for someone who lives in Claude Code / Codex.
|
|
4
|
+
Optimize for a phone-skimmer who must come away knowing WHAT IT IS and wanting
|
|
5
|
+
to run it — never scared off. Validated by a 6-persona cold-read (2026-07):
|
|
6
|
+
newcomer / power-user / plugin-author / skeptical-senior / decision-maker /
|
|
7
|
+
Codex-user. The fixes below trace to that review — don't regress them.
|
|
8
|
+
|
|
9
|
+
HOOK = FELT PAIN, then breadth (founder direction 2026-07). Bold tagline is a
|
|
10
|
+
specific second-person pain: "You have a rule your agent follows half the time
|
|
11
|
+
— and no way to know which one." It's rhetorical (the reader's own uncertainty),
|
|
12
|
+
NOT an ecosystem stat — do NOT reintroduce a bare "%"/number in the tagline (the
|
|
13
|
+
doc mocks uncited stats: "'65% fewer tokens.' Says who?", so a fake stat up top
|
|
14
|
+
reads as hypocrisy). Dev-native CRAFT bait, NOT enterprise fear (OSS dev tool,
|
|
15
|
+
not a security product). The report-card + 5 rings RIGHT BELOW show the breadth
|
|
16
|
+
so the specific hook doesn't read as narrow.
|
|
17
|
+
|
|
18
|
+
DEFINE "HARNESS" on first use (the load-bearing noun, used ~15×) — gloss it as
|
|
19
|
+
"the CLAUDE.md/AGENTS.md rules, skills, subagents, and hooks steering your
|
|
20
|
+
agent." 4 of 6 personas bounced on it being undefined. Keep the gloss.
|
|
21
|
+
|
|
22
|
+
INSTRUCTION-NEUTRAL NOUNS (Codex was the lowest score): body copy says CLAUDE.md
|
|
23
|
+
OR AGENTS.md, never CLAUDE.md alone. The tagline may keep punch, but Codex must
|
|
24
|
+
appear within the first sentence or two (the harness gloss names AGENTS.md), and
|
|
25
|
+
every "adopts your CLAUDE.md"/"verifies your CLAUDE.md" gets "or AGENTS.md".
|
|
26
|
+
|
|
27
|
+
CATEGORY = A TOOL YOU RUN (ESLint/Lighthouse/npm audit class), NOT a framework,
|
|
28
|
+
NOT a lib-collection. The FAQ says this outright — it's the #1 thing that scares
|
|
29
|
+
people off. The library/subpath exports are the automation door for the 5%.
|
|
30
|
+
|
|
31
|
+
AUDIT vs LINT — ONE CONSISTENT, ACCURATE STORY (a persona caught 3 conflicting
|
|
32
|
+
answers): `lint` = the CI gate on the deterministic checks (broken refs, tool
|
|
33
|
+
contracts, dead hooks, skill collisions — Proofs 1 & 2). `audit` = those same
|
|
34
|
+
checks + the Safety ring + two opt-in LIVE checks (MCP connects? skills fire?)
|
|
35
|
+
+ the graded report. Do NOT claim lint gates the lethal-trifecta Safety flag
|
|
36
|
+
(it's an audit ring, not a default gating rule) and do NOT say lint is
|
|
37
|
+
refs-only. The table row, the reconciliation line, and the Lint subsection must
|
|
38
|
+
all agree.
|
|
39
|
+
|
|
40
|
+
SPINE = proof/demo-led. Real, screenshotable catches on shipped plugins, THEN
|
|
41
|
+
mechanism. Every proof traces to a real dogfood run (research/dogfood/) — NEVER
|
|
42
|
+
fabricate one. Order = most-RELATABLE first (broken tool ref → skill collision →
|
|
43
|
+
secrets-exfil gotcha last as the bite). The intro triplet maps 1:1 to the 3
|
|
44
|
+
proofs. Security is ONE dev-native GOTCHA proof, never the brand. Add a repro
|
|
45
|
+
line ("run `npx vigiles audit <any-repo>` for your own") + a one-line note that
|
|
46
|
+
examples use CC subagents but the checks run on Codex too. FALSE CONFIDENCE is
|
|
47
|
+
the coined term (a guard that looks like it works and silently doesn't), defined
|
|
48
|
+
once in the Test section.
|
|
49
|
+
|
|
50
|
+
FUNNEL: "How it works" opens with the audit/lint/test/eval verb-map table — the
|
|
51
|
+
load-bearing "one tool, not four" fix. Frame it vibes → verified. Name that
|
|
52
|
+
init/compile/eject manage the spec layer (personas noticed the verb-count gap).
|
|
53
|
+
|
|
54
|
+
DON'T SHAME OSS: catches are ANONYMIZED (no obra/superpowers, madappgang,
|
|
55
|
+
claude-flow by name) — real names live only in research/dogfood/.
|
|
56
|
+
Guard/compiled-hooks + the 2/7→7/7 battery are PARKED FOR LAUNCH — not the hero.
|
|
57
|
+
|
|
58
|
+
RULES: lead with the reader's CONCRETE PAIN; ≤ ~3-line paragraphs; ONE bold per
|
|
59
|
+
block; ONE idea per sentence; NO internal vocabulary (moat/flywheel) / NO
|
|
60
|
+
research/ links / NO enterprise/national-interest framing — name the user
|
|
61
|
+
benefit; ~220-line body cap; push depth into docs/ and LINK it. Assets: the hero
|
|
62
|
+
vigiles-audit.png is a REAL current report (a community plugin as "my-plugin"),
|
|
63
|
+
C 72, five rings — re-render via headless Chromium if the UI changes.
|
|
86
64
|
-->
|
|
87
65
|
|
|
88
66
|
<p align="center">
|
|
@@ -92,11 +70,11 @@
|
|
|
92
70
|
<h1 align="center">vigiles</h1>
|
|
93
71
|
|
|
94
72
|
<p align="center">
|
|
95
|
-
<strong>
|
|
73
|
+
<strong>You have a rule your agent follows half the time — and no way to know which one.</strong>
|
|
96
74
|
</p>
|
|
97
75
|
|
|
98
76
|
<p align="center">
|
|
99
|
-
Verify your CLAUDE.md, skills, and hooks are real — and prove they actually work.
|
|
77
|
+
Verify your CLAUDE.md or AGENTS.md, skills, and hooks are real — and prove they actually work.
|
|
100
78
|
</p>
|
|
101
79
|
|
|
102
80
|
<p align="center">
|
|
@@ -107,44 +85,44 @@
|
|
|
107
85
|
|
|
108
86
|
---
|
|
109
87
|
|
|
110
|
-
**
|
|
88
|
+
**You review every PR. Nothing reviews your CLAUDE.md.**
|
|
111
89
|
|
|
112
|
-
|
|
90
|
+
Your **harness** — the CLAUDE.md or AGENTS.md rules, skills, subagents, and hooks steering your agent — is the one part nobody checks. Nobody verified it's real. Nobody tested it works. That's not a system. That's vibes.
|
|
113
91
|
|
|
114
|
-
|
|
92
|
+
And vibes break silently mid-task: a subagent wired to a tool that doesn't exist, two skills your agent can't tell apart, one helper quietly able to read your secrets and send them out.
|
|
93
|
+
|
|
94
|
+
vigiles[^name] checks your harness is _real_, not just well-formed — Claude Code and Codex alike. One command, no key, no config, safe on any repo:
|
|
115
95
|
|
|
116
96
|
```bash
|
|
117
97
|
npx vigiles audit
|
|
118
98
|
```
|
|
119
99
|
|
|
120
|
-
|
|
121
|
-
ship. ↓
|
|
100
|
+
It's free and open-source, runs entirely on your machine, and never bills per token. (`eval` is the only step that calls a model — on your own Claude subscription.) Here's what it caught on plugins people actually ship. ↓
|
|
122
101
|
|
|
123
102
|
## What it caught
|
|
124
103
|
|
|
125
104
|
<p align="center">
|
|
126
|
-
<img src="vigiles-audit.png" width="760" alt="vigiles audit report scoring my-plugin C (72/100): five categories scored A–F — Truthfulness, Triggering, Structure, Safety, Tested — with
|
|
105
|
+
<img src="vigiles-audit.png" width="760" alt="vigiles audit report scoring my-plugin C (72/100): five categories scored A–F — Truthfulness, Triggering, Structure, Safety, Tested — with an inline fix card for a subagent declaring a tool that doesn't exist" />
|
|
127
106
|
</p>
|
|
128
107
|
|
|
129
|
-
**Like Google's Lighthouse, but for your agent harness.**
|
|
130
|
-
|
|
108
|
+
**Like Google's Lighthouse, but for your agent harness.** One command grades it A–F across five categories, every fix shown inline:
|
|
109
|
+
|
|
110
|
+
- **Truthfulness** — do the references resolve?
|
|
111
|
+
- **Triggering** — do skills fire, without colliding?
|
|
112
|
+
- **Structure** — are tool contracts and configs valid?
|
|
113
|
+
- **Safety** — any way for the agent to leak your data?
|
|
114
|
+
- **Tested** — does the harness ship tests?
|
|
131
115
|
|
|
132
|
-
|
|
133
|
-
For CI gating, use `vigiles lint` instead. **[Audit a harness →](docs/for-plugin-authors.md)**
|
|
116
|
+
These are real scans of public plugins — run `npx vigiles audit <any-repo>` for your own. The examples below use Claude Code subagents; the same checks run on Codex `AGENTS.md`, skills, and hooks. ↓
|
|
134
117
|
|
|
135
|
-
## Proof 1 — your agent
|
|
118
|
+
## Proof 1 — a tool your agent thinks it has and doesn't
|
|
136
119
|
|
|
137
120
|
```text
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
reads private data (Bash, Read) · takes in untrusted web content (WebFetch)
|
|
141
|
-
· can send data out (Bash, WebFetch)
|
|
121
|
+
✗ tester — Tool "AskUserQuestion" is never available to a subagent.
|
|
122
|
+
→ remove or correct it — it's silently dropped from the contract.
|
|
142
123
|
```
|
|
143
124
|
|
|
144
|
-
|
|
145
|
-
web page can tell it to read your `.env` and POST it anywhere — no exploit code, just the
|
|
146
|
-
tools it was handed. vigiles flags it from the tool list alone, free, no model.
|
|
147
|
-
**[How the Safety check works →](docs/for-plugin-authors.md)**
|
|
125
|
+
This subagent — a helper your main agent hands work to — lists a tool that doesn't exist for it. The harness drops it without a word, so the agent quietly loses a capability it thinks it has. The markdown is perfectly valid. vigiles catches it and hands you the **one-line fix**.
|
|
148
126
|
|
|
149
127
|
## Proof 2 — two skills your agent can't tell apart
|
|
150
128
|
|
|
@@ -154,62 +132,50 @@ tools it was handed. vigiles flags it from the tool list alone, free, no model.
|
|
|
154
132
|
apart, so the wrong one fires (e.g. "agent-coder" ↔ "agent-tester", 83% alike)
|
|
155
133
|
```
|
|
156
134
|
|
|
157
|
-
One popular plugin ships **45 pairs of skills** with near-identical descriptions.
|
|
158
|
-
agent picks which skill to run by reading those descriptions, so when two match it
|
|
159
|
-
fires the wrong one. The markdown is perfectly valid.
|
|
135
|
+
One popular plugin ships **45 pairs of skills** with near-identical descriptions. Your agent picks which skill to run by _reading_ those descriptions, so when two match it fires the wrong one. Still perfectly valid markdown.
|
|
160
136
|
**[How triggering works →](docs/measuring-skills.md)**
|
|
161
137
|
|
|
162
|
-
## Proof 3 —
|
|
138
|
+
## Proof 3 — it can quietly read your secrets and send them out
|
|
163
139
|
|
|
164
140
|
```text
|
|
165
|
-
|
|
166
|
-
|
|
141
|
+
◑ Safety 80 (80/100)
|
|
142
|
+
└ subagent "tester" holds all three lethal-trifecta legs:
|
|
143
|
+
reads private data (Bash, Read) · takes in untrusted web content (WebFetch)
|
|
144
|
+
· can send data out (Bash, WebFetch)
|
|
167
145
|
```
|
|
168
146
|
|
|
169
|
-
|
|
170
|
-
doesn't exist. The harness drops it silently, so the agent loses a capability it
|
|
171
|
-
thinks it has. vigiles catches it and gives you the **one-line fix**.
|
|
147
|
+
Hand one subagent all three powers and a poisoned web page can tell it to read your `.env` and POST it anywhere — no exploit code, just the tools it was given. The **80 still looks like a B** — that's the point: a healthy-looking grade can hide a single subagent that's a data-leak waiting to happen. vigiles spots it from the tool list alone, free, no model.
|
|
172
148
|
|
|
173
|
-
That's the whole idea
|
|
174
|
-
|
|
175
|
-
where you name a linter rule, it's checked to exist _and_ be enabled (ESLint, Ruff,
|
|
176
|
-
Clippy and more).
|
|
177
|
-
**[Full guide →](docs/verifying-instruction-files.md)**
|
|
149
|
+
That's the whole idea: it checks your harness against **reality, not style**. Every tool, hook, file, script, and skill you reference is verified to actually resolve — and where you name a linter rule, it's checked to exist _and_ be enabled (ESLint, Ruff, Clippy, and more).
|
|
150
|
+
**[Everything it catches →](docs/what-vigiles-catches.md)** · point `audit` at a whole marketplace and it ranks every plugin the same way.
|
|
178
151
|
|
|
179
|
-
|
|
180
|
-
classes of bug by construction (a typed spec or compiled hook just won't compile).
|
|
181
|
-
**[Everything it catches and prevents →](docs/what-vigiles-catches.md)** · point `audit`
|
|
182
|
-
at a whole marketplace and it ranks every plugin the same way.
|
|
183
|
-
**[Audit a marketplace →](docs/for-plugin-authors.md)**
|
|
152
|
+
## How it works — vibes → verified
|
|
184
153
|
|
|
185
|
-
|
|
154
|
+
`audit` shows you where your setup is still vibes. Turning that into _verified_ is four commands over one engine — and almost none of it needs a model or a key.
|
|
186
155
|
|
|
187
|
-
|
|
188
|
-
|
|
156
|
+
| Command | Answers | Needs a model? | When to run |
|
|
157
|
+
| ------- | ------------------------------ | ------------------------ | ------------------------ |
|
|
158
|
+
| `audit` | Everything, graded A–F | No — read-only[^audit] | Anytime; it's the report |
|
|
159
|
+
| `lint` | Do the structural checks pass? | No | CI gate, every push |
|
|
160
|
+
| `test` | Does the harness behave? | No — a scripted stand-in | Every commit |
|
|
161
|
+
| `eval` | Does a skill actually help? | Yes — your subscription | On demand |
|
|
189
162
|
|
|
190
|
-
|
|
163
|
+
`audit` and `lint` share one engine. **`lint` is the CI gate** — it fails the build on broken references, bad tool contracts, dead hooks, and skill collisions (Proofs 1 and 2). **`audit`** runs those same checks, adds the Safety ring, renders the graded report, and can also run two opt-in _live_ checks (does your MCP server connect, do your skills fire). `test` and `eval` go past _does it exist_ to _does it work_. (`init` / `compile` / `eject` manage the spec layer underneath; you rarely run them by hand.)
|
|
191
164
|
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
195
|
-
edited by your agent in plain English, undone by `eject`.
|
|
165
|
+
### 🔎 Lint — your instructions stop lying
|
|
166
|
+
|
|
167
|
+
Every path, script, symbol, and rule verified against reality — plus tool contracts, skill collisions, and dead hooks (the catches above). You don't write the checks. `npx vigiles init` writes a `CLAUDE.md.spec.ts` beside your file: the same rules, each reference now wrapped so vigiles can confirm it exists. `compile` turns that back into the `CLAUDE.md` (or `AGENTS.md`) your agent already reads. Your agent edits the spec in plain English; `eject` deletes it and leaves your original untouched.
|
|
196
168
|
**[How →](docs/verifying-instruction-files.md)**
|
|
197
169
|
|
|
198
170
|
### 🧪 Test — does the harness actually do its job?
|
|
199
171
|
|
|
200
|
-
A hook that blocks nothing, a skill that hijacks unrelated prompts, context that never
|
|
201
|
-
reaches the model — each passes a naive "did it run?" check. vigiles tests the real
|
|
202
|
-
thing: hooks **block**, skills **fire**, subagents **finish what they promised**, and a
|
|
203
|
-
stray `git push` is caught before it happens. No model, no key, on every commit.
|
|
172
|
+
A hook that blocks nothing, a skill that hijacks unrelated prompts, context that never reaches the model — each passes a naive "did it run?" check. That gap is **false confidence**: a guard that looks like it works and silently doesn't. vigiles tests the real thing — hooks block, skills fire, subagents finish what they promised, a stray `git push` is caught before it happens. It drives a scripted stand-in for the model, not a live call, so it needs no key and runs on every commit.
|
|
204
173
|
**[How testing works →](docs/harness-testing.md)**
|
|
205
174
|
|
|
206
175
|
### 📊 Eval — does a skill help, or just cost more?
|
|
207
176
|
|
|
208
|
-
_"65% fewer tokens." Says who?_ vigiles
|
|
209
|
-
|
|
210
|
-
and DeepEval bill **per token, every run**; vigiles runs on your own Claude Pro/Max
|
|
211
|
-
subscription. Evals run locally — a committed lock then lets **CI catch stale results with no
|
|
212
|
-
model call**. **[Measure a skill →](docs/measuring-skills.md)**
|
|
177
|
+
_"65% fewer tokens." Says who?_ vigiles A/Bs the claim on real coding tasks and reports the token bill, whether it hit its target, and whether the code still works. promptfoo and DeepEval bill **per token, every run**; vigiles runs on your own Claude Pro/Max subscription. Evals run locally; a committed lock file — like a `package-lock` — records the result, so CI catches stale numbers without calling the model again. (On Claude Code today; Codex eval support is landing.)
|
|
178
|
+
**[Measure a skill →](docs/measuring-skills.md)**
|
|
213
179
|
|
|
214
180
|
## Quick start
|
|
215
181
|
|
|
@@ -232,21 +198,19 @@ Or run it yourself:
|
|
|
232
198
|
|
|
233
199
|
```bash
|
|
234
200
|
npx vigiles init # adopts your files (non-destructive — eject reverses), adds CI,
|
|
235
|
-
# installs
|
|
201
|
+
# installs vigiles's skills + hooks as a Claude Code plugin (in
|
|
202
|
+
# ~/.claude/, not your repo). On Codex, skills install globally too.
|
|
236
203
|
```
|
|
237
204
|
|
|
238
|
-
Interactive in a terminal, non-interactive for agents/CI (or `--yes`).
|
|
205
|
+
Interactive in a terminal, non-interactive for agents/CI (or `--yes`). **Works with Claude Code and Codex** — vigiles verifies `CLAUDE.md` and `AGENTS.md` the same way. **[Codex setup →](docs/harnesses.md)**
|
|
239
206
|
|
|
240
|
-
**Adoption is smooth: one command, then your agent does the rest.** `init` installs
|
|
241
|
-
the **skills and hooks**, so a plain-English ask does the work — no specs to
|
|
242
|
-
hand-write, no hooks to wire:
|
|
207
|
+
**Adoption is smooth: one command, then your agent does the rest.** `init` installs the **skills and hooks**, so a plain-English ask does the work — no specs to hand-write, no hooks to wire:
|
|
243
208
|
|
|
244
209
|
- _"test my skills"_ → scaffolds **and runs** a trigger/behaviour test, then commits its result so CI can check it (`test-harness`)
|
|
245
210
|
- _"harden my rules"_ → upgrades prose guidance into enforced linter rules (`strengthen`)
|
|
246
|
-
- _"add a rule to my CLAUDE.md"_ → edits the source and recompiles (`edit-spec`)
|
|
211
|
+
- _"add a rule to my CLAUDE.md or AGENTS.md"_ → edits the source and recompiles (`edit-spec`)
|
|
247
212
|
|
|
248
|
-
The **hooks** keep it honest in-loop — nudging the agent to
|
|
249
|
-
refresh a stale eval — so there are no chores to remember.
|
|
213
|
+
The **hooks** keep it honest in-loop — nudging the agent to tag a linter-rule mention so vigiles can verify it, or to re-run a test whose result just went stale — so there are no chores to remember.
|
|
250
214
|
|
|
251
215
|
<details>
|
|
252
216
|
<summary>What <code>init</code> sets up</summary>
|
|
@@ -254,35 +218,34 @@ refresh a stale eval — so there are no chores to remember.
|
|
|
254
218
|
- **Both lint and test** by default; scope with `--lint` / `--test`.
|
|
255
219
|
- **Already have a CLAUDE.md / AGENTS.md, skills, or subagents? `init` adopts them all** into specs faithfully and **non-destructively** — untouched until you `compile` (and `eject` undoes it).
|
|
256
220
|
- Adds `vigiles` to `devDependencies`; installs the Claude Code plugin (skills + hooks) via the marketplace — globally, never vendored.
|
|
257
|
-
- Wires CI as a `zernie/vigiles@v1` workflow that posts a sticky PR comment + a `valid` output.
|
|
221
|
+
- Wires CI as a `zernie/vigiles@v1` workflow (needs only read + PR-comment permissions) that posts a sticky PR comment + a `valid` output.
|
|
258
222
|
|
|
259
|
-
|
|
260
|
-
[your own harness](docs/authoring-an-adapter.md). Prefer to write tests yourself?
|
|
261
|
-
JS **or** TS (`*.harness.{mjs,ts}`) — run with `npx vigiles test`.
|
|
223
|
+
Targets Claude Code and Codex out of the box, or [your own harness](docs/authoring-an-adapter.md). Prefer to write tests yourself? JS **or** TS (`*.harness.{mjs,ts}`) — run with `npx vigiles test`.
|
|
262
224
|
|
|
263
225
|
</details>
|
|
264
226
|
|
|
265
227
|
## FAQ
|
|
266
228
|
|
|
229
|
+
- **Is this a framework I have to build around?** No. It's a tool you run — like ESLint, Lighthouse, or `npm audit`. One command, a report, an optional CI gate. There's a library API for automation, but you never touch it to get value.
|
|
267
230
|
- **Isn't this just a markdown linter?** No — it checks whether your instruction file is _true_ (every path/script/symbol/rule exists and is enabled), then tests and measures your harness. A style linter can't do any of that.
|
|
268
|
-
- **Do I have to write TypeScript?** No — your agent writes the spec (`init` adopts your CLAUDE.md into one), or plain markdown lints with zero new files. Compiler-grade guarantees are opt-in, like TS's `strict` ([why?](docs/faq.md#why-are-the-strongest-guarantees-opt-in-not-the-default)).
|
|
269
|
-
- **
|
|
231
|
+
- **Do I have to write TypeScript?** No — your agent writes the spec (`init` adopts your CLAUDE.md or AGENTS.md into one), or plain markdown lints with zero new files. Compiler-grade guarantees are opt-in, like TS's `strict` ([why?](docs/faq.md#why-are-the-strongest-guarantees-opt-in-not-the-default)).
|
|
232
|
+
- **Is it stable enough to adopt?** Yes — the CLI is stable; only the library API is still evolving ([details](STABILITY.md)).
|
|
233
|
+
- **Non-JS repo?** `npx vigiles lint` verifies your CLAUDE.md or AGENTS.md with no install (Ruff/Clippy/Pylint/… too).
|
|
270
234
|
|
|
271
235
|
**[Full FAQ →](docs/faq.md)**
|
|
272
236
|
|
|
237
|
+
**Not for you if** you want a model/capability benchmark or runtime guardrails in the request path — vigiles is build-/CI-time.
|
|
238
|
+
|
|
273
239
|
## More
|
|
274
240
|
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
|
|
278
|
-
- **[Ship plugins? The plugin-author guide →](docs/for-plugin-authors.md)** — scan a draft, make your skills fire, rank a whole marketplace — no key.
|
|
279
|
-
- **[Docs index →](docs/README.md)** · **[API reference →](https://zernie.github.io/vigiles/)** · **[Related tools →](docs/related-tools.md)**.
|
|
280
|
-
- **[Stability →](STABILITY.md)** — 0.x: the CLI is stable; the library API is still evolving.
|
|
281
|
-
- **Not for you if** you want a model/capability benchmark or runtime guardrails in the request path — vigiles is build-/CI-time.
|
|
282
|
-
- Companion to [Feedback Loop Is All You Need](https://zernie.com/blog/feedback-loop-is-all-you-need).
|
|
241
|
+
**Docs** — **[What it catches and prevents →](docs/what-vigiles-catches.md)** · **[Verifying instruction files →](docs/verifying-instruction-files.md)** ([rules matrix](docs/verifying-instruction-files.md#the-validation-rules--the-full-matrix)) · **[Harness testing →](docs/harness-testing.md)** · **[Measuring skills →](docs/measuring-skills.md)** · **[CLI →](docs/cli.md)** · **[GitHub Action →](docs/github-action.md)** · **[Skills →](docs/skills.md)** · **[Plugin-author guide →](docs/for-plugin-authors.md)** · **[Docs index →](docs/README.md)** · **[API reference →](https://zernie.github.io/vigiles/)**
|
|
242
|
+
|
|
243
|
+
**Project** — **[Stability →](STABILITY.md)** · **[Related tools →](docs/related-tools.md)** · companion to [Feedback Loop Is All You Need](https://zernie.com/blog/feedback-loop-is-all-you-need).
|
|
283
244
|
|
|
284
245
|
## License
|
|
285
246
|
|
|
286
247
|
[MIT](LICENSE)
|
|
287
248
|
|
|
288
249
|
[^name]: **vigiles** — the watchmen of ancient Rome, who guarded the city (and fought its fires) by night. _Quis custodiet ipsos custodes?_ — "who watches the watchmen?" (Juvenal, _Satire VI_).
|
|
250
|
+
|
|
251
|
+
[^audit]: `audit` reads only by default. Two deeper checks — live MCP connections and skill-firing — are opt-in and ask before they run.
|
|
@@ -46,6 +46,47 @@ export declare function discoverScripts(patterns: readonly string[], defaultGlob
|
|
|
46
46
|
export declare function runScripts(files: readonly string[], cwd: string, env?: NodeJS.ProcessEnv): ScriptRunResult[];
|
|
47
47
|
/** Whether any script FAILED (a skip is not a failure). */
|
|
48
48
|
export declare function anyFailed(results: readonly ScriptRunResult[]): boolean;
|
|
49
|
+
/**
|
|
50
|
+
* What a `test`/`eval` invocation should do about actually RUNNING the discovered
|
|
51
|
+
* scripts:
|
|
52
|
+
* - `run` — proceed.
|
|
53
|
+
* - `confirm` — interactive human, no explicit intent: ask before firing `count`.
|
|
54
|
+
* - `refuse` — headless, no explicit intent: don't silently fire the whole tree.
|
|
55
|
+
*/
|
|
56
|
+
export type RunScriptsDecision = {
|
|
57
|
+
readonly kind: "run";
|
|
58
|
+
} | {
|
|
59
|
+
readonly kind: "confirm";
|
|
60
|
+
readonly count: number;
|
|
61
|
+
} | {
|
|
62
|
+
readonly kind: "refuse";
|
|
63
|
+
readonly count: number;
|
|
64
|
+
};
|
|
65
|
+
export interface RunScriptsEnv {
|
|
66
|
+
/** `test` is free/deterministic → always runs. `eval` spends model quota. */
|
|
67
|
+
readonly kind: "test" | "eval";
|
|
68
|
+
/** The user named explicit target files/globs (positional args) — clear intent. */
|
|
69
|
+
readonly explicitTargets: boolean;
|
|
70
|
+
/** How many script files the discovery matched. */
|
|
71
|
+
readonly matchedCount: number;
|
|
72
|
+
/** A human at a terminal who can answer + wait. */
|
|
73
|
+
readonly isTTY: boolean;
|
|
74
|
+
/** `--all` — opt in to running the whole discovered set without a prompt. */
|
|
75
|
+
readonly all: boolean;
|
|
76
|
+
/** `--yes` / `--no-interactive` — agent/CI mode: never prompt. */
|
|
77
|
+
readonly yes: boolean;
|
|
78
|
+
}
|
|
79
|
+
/**
|
|
80
|
+
* Consent gate for a bare (no-target) `vigiles eval`. `eval` runs the REAL model
|
|
81
|
+
* on your subscription, and a no-target run discovers every `*.eval.*` over the
|
|
82
|
+
* whole tree — so a repo with many evals fires them all and spends quota. Mirrors
|
|
83
|
+
* `audit`'s read-vs-run consent (`decideExecute`): a paid, side-effecting verb
|
|
84
|
+
* never fans out over an unbounded glob without either an explicit target, an
|
|
85
|
+
* `--all` opt-in, or an interactive yes. `test` is free + deterministic, so it
|
|
86
|
+
* always runs. Total + pure, first match wins; the IO (prompt/refuse) lives in the
|
|
87
|
+
* CLI.
|
|
88
|
+
*/
|
|
89
|
+
export declare function decideRunScripts(o: RunScriptsEnv): RunScriptsDecision;
|
|
49
90
|
/** One line per file + an explicit pass/skip/fail tally. Skips are SHOWN, never
|
|
50
91
|
* folded into "passed" — a `⊘ SKIPPED` is loud, not a silent green. */
|
|
51
92
|
export declare function formatScriptSummary(results: readonly ScriptRunResult[]): string;
|
|
@@ -7,6 +7,7 @@ exports.detectNodeCaps = detectNodeCaps;
|
|
|
7
7
|
exports.discoverScripts = discoverScripts;
|
|
8
8
|
exports.runScripts = runScripts;
|
|
9
9
|
exports.anyFailed = anyFailed;
|
|
10
|
+
exports.decideRunScripts = decideRunScripts;
|
|
10
11
|
exports.formatScriptSummary = formatScriptSummary;
|
|
11
12
|
/**
|
|
12
13
|
* vigiles — run harness-test / eval script files via the CLI.
|
|
@@ -123,6 +124,31 @@ function runScripts(files, cwd, env = {}) {
|
|
|
123
124
|
function anyFailed(results) {
|
|
124
125
|
return results.some((r) => r.status === "fail");
|
|
125
126
|
}
|
|
127
|
+
/**
|
|
128
|
+
* Consent gate for a bare (no-target) `vigiles eval`. `eval` runs the REAL model
|
|
129
|
+
* on your subscription, and a no-target run discovers every `*.eval.*` over the
|
|
130
|
+
* whole tree — so a repo with many evals fires them all and spends quota. Mirrors
|
|
131
|
+
* `audit`'s read-vs-run consent (`decideExecute`): a paid, side-effecting verb
|
|
132
|
+
* never fans out over an unbounded glob without either an explicit target, an
|
|
133
|
+
* `--all` opt-in, or an interactive yes. `test` is free + deterministic, so it
|
|
134
|
+
* always runs. Total + pure, first match wins; the IO (prompt/refuse) lives in the
|
|
135
|
+
* CLI.
|
|
136
|
+
*/
|
|
137
|
+
function decideRunScripts(o) {
|
|
138
|
+
if (o.kind === "test")
|
|
139
|
+
return { kind: "run" };
|
|
140
|
+
if (o.explicitTargets)
|
|
141
|
+
return { kind: "run" };
|
|
142
|
+
if (o.all || o.yes)
|
|
143
|
+
return { kind: "run" };
|
|
144
|
+
// A bounded no-target run (0 = no-op, 1 = a single obviously-intended eval) is
|
|
145
|
+
// not the footgun; the footgun is fanning out over the whole tree.
|
|
146
|
+
if (o.matchedCount <= 1)
|
|
147
|
+
return { kind: "run" };
|
|
148
|
+
if (!o.isTTY)
|
|
149
|
+
return { kind: "refuse", count: o.matchedCount };
|
|
150
|
+
return { kind: "confirm", count: o.matchedCount };
|
|
151
|
+
}
|
|
126
152
|
const MARK = {
|
|
127
153
|
pass: "✓",
|
|
128
154
|
skip: "⊘",
|
package/dist/cli.js
CHANGED
|
@@ -3487,11 +3487,29 @@ async function handleGenerateHarness(args, restArgs) {
|
|
|
3487
3487
|
console.log(`\n✓ Generated ${(0, generate_harness_js_1.labelFor)(process.cwd(), fullOut)}`);
|
|
3488
3488
|
console.log(" `tsc --noEmit` over this file now checks every delegate target resolves.");
|
|
3489
3489
|
}
|
|
3490
|
+
/** Minimal TTY yes/no prompt (readline). Returns true only on an explicit y/yes. */
|
|
3491
|
+
async function promptYesNo(question) {
|
|
3492
|
+
const readline = await import("node:readline");
|
|
3493
|
+
const rl = readline.createInterface({
|
|
3494
|
+
input: process.stdin,
|
|
3495
|
+
output: process.stdout,
|
|
3496
|
+
});
|
|
3497
|
+
try {
|
|
3498
|
+
const answer = await new Promise((res) => {
|
|
3499
|
+
rl.question(question, res);
|
|
3500
|
+
});
|
|
3501
|
+
return /^y(es)?$/i.test(answer.trim());
|
|
3502
|
+
}
|
|
3503
|
+
finally {
|
|
3504
|
+
rl.close();
|
|
3505
|
+
}
|
|
3506
|
+
}
|
|
3490
3507
|
/**
|
|
3491
3508
|
* `vigiles test` / `vigiles eval` — discover and run the two-tier harness
|
|
3492
3509
|
* scripts (deterministic `*.harness.mjs` / real-model `*.eval.mjs`) as child
|
|
3493
3510
|
* `node` processes, aggregating exit codes so they work as a CI command. See
|
|
3494
|
-
* src/run-scripts.ts.
|
|
3511
|
+
* src/run-scripts.ts. A bare `vigiles eval` (no target) asks before fanning out
|
|
3512
|
+
* over the whole tree — it spends model quota (see `decideRunScripts`).
|
|
3495
3513
|
*
|
|
3496
3514
|
* `vigiles test` skips clean when the `claude` CLI is absent (the deterministic
|
|
3497
3515
|
* tier needs it, just like the node:test suite). `--trials=N` is forwarded to
|
|
@@ -3532,7 +3550,7 @@ function resolveEvalLockEnv(args) {
|
|
|
3532
3550
|
}
|
|
3533
3551
|
return env;
|
|
3534
3552
|
}
|
|
3535
|
-
function handleRunScripts(kind, args, restArgs) {
|
|
3553
|
+
async function handleRunScripts(kind, args, restArgs) {
|
|
3536
3554
|
const cwd = process.cwd();
|
|
3537
3555
|
// Harness/eval scripts may be authored in JS or TS (see run-scripts.ts).
|
|
3538
3556
|
const defaultGlob = (0, run_scripts_js_1.scriptGlob)(kind === "test" ? "harness" : "eval");
|
|
@@ -3563,6 +3581,33 @@ function handleRunScripts(kind, args, restArgs) {
|
|
|
3563
3581
|
console.log(`No ${defaultGlob} files found.`);
|
|
3564
3582
|
return;
|
|
3565
3583
|
}
|
|
3584
|
+
// Consent gate for a bare `vigiles eval`: it runs the REAL model on your
|
|
3585
|
+
// subscription, and a no-target run discovered the whole tree — so never fan out
|
|
3586
|
+
// over an unbounded glob without explicit intent. Mirrors audit's read-vs-run
|
|
3587
|
+
// consent. `test` is free → always runs (decideRunScripts returns "run").
|
|
3588
|
+
const runDecision = (0, run_scripts_js_1.decideRunScripts)({
|
|
3589
|
+
kind,
|
|
3590
|
+
explicitTargets: restArgs.length > 0,
|
|
3591
|
+
matchedCount: files.length,
|
|
3592
|
+
isTTY: (process.stdin.isTTY ?? false) && (process.stdout.isTTY ?? false),
|
|
3593
|
+
all: args.includes("--all"),
|
|
3594
|
+
yes: args.includes("--yes") || args.includes("--no-interactive"),
|
|
3595
|
+
});
|
|
3596
|
+
if (runDecision.kind === "refuse") {
|
|
3597
|
+
console.error(`✗ vigiles eval: ${String(runDecision.count)} eval file(s) matched the whole tree, and each ` +
|
|
3598
|
+
"runs the real model on your subscription. Refusing to fire them all non-interactively.\n" +
|
|
3599
|
+
" → name the eval(s): vigiles eval path/to/x.eval.mjs\n" +
|
|
3600
|
+
" → or opt in to all: vigiles eval --all");
|
|
3601
|
+
process.exit(2);
|
|
3602
|
+
}
|
|
3603
|
+
if (runDecision.kind === "confirm") {
|
|
3604
|
+
const ok = await promptYesNo(`About to run ${String(runDecision.count)} eval file(s) against the real model on your ` +
|
|
3605
|
+
"subscription (uses your Claude quota). Continue? [y/N] ");
|
|
3606
|
+
if (!ok) {
|
|
3607
|
+
console.log("Aborted. Name specific eval(s), or pass --all to run them all.");
|
|
3608
|
+
return;
|
|
3609
|
+
}
|
|
3610
|
+
}
|
|
3566
3611
|
// No blanket skip: unit-tier (runHook) tests need no `claude`, so always run.
|
|
3567
3612
|
// A script whose tier DOES need `claude` self-reports `⊘ SKIPPED` (exit 77) —
|
|
3568
3613
|
// loud, never a silent green. Just flag up front that some may skip.
|
|
@@ -4934,10 +4979,10 @@ async function main() {
|
|
|
4934
4979
|
break;
|
|
4935
4980
|
}
|
|
4936
4981
|
case "test":
|
|
4937
|
-
handleRunScripts("test", args, restArgs);
|
|
4982
|
+
await handleRunScripts("test", args, restArgs);
|
|
4938
4983
|
break;
|
|
4939
4984
|
case "eval":
|
|
4940
|
-
handleRunScripts("eval", args, restArgs);
|
|
4985
|
+
await handleRunScripts("eval", args, restArgs);
|
|
4941
4986
|
break;
|
|
4942
4987
|
case "audit": {
|
|
4943
4988
|
// The Lighthouse run: a plain `audit` is a deterministic READ — rings, each
|
package/dist/eval.d.ts
CHANGED
|
@@ -270,6 +270,21 @@ export declare function spawnAgent(a: AgentRunArgs): Promise<RunOut>;
|
|
|
270
270
|
* working model auth (e.g. `ANTHROPIC_API_KEY`). Thin wrapper over
|
|
271
271
|
* `runEvalWith` with the real agent runner.
|
|
272
272
|
*/
|
|
273
|
+
/**
|
|
274
|
+
* A `SKILL.md` written straight into a run's cwd (via an arm's `files`) is NOT
|
|
275
|
+
* registered as a skill by the harness. Claude Code — and Codex — load skills only
|
|
276
|
+
* from a plugin / `.claude/skills` layout, so a bare cwd `SKILL.md` sits unread:
|
|
277
|
+
* the arm silently measures NOTHING (baseline and "skill" become the same run). This
|
|
278
|
+
* is the exact footgun the ecosystem benchmark hit — a skill file delivered where it
|
|
279
|
+
* can never activate, with no error. Detect it so {@link runEval} / {@link measureArms}
|
|
280
|
+
* can WARN and point at `pluginDir` (a real `--plugin-dir` install) or `skillsDir`.
|
|
281
|
+
*
|
|
282
|
+
* HIGH-PRECISION (don't cry wolf): only a file whose basename is `SKILL.md` AND that
|
|
283
|
+
* carries real skill frontmatter (a `---` block naming `name`/`description`) is flagged
|
|
284
|
+
* — so an empty scratch `SKILL.md` a task is asked to AUTHOR is never flagged. Pure +
|
|
285
|
+
* exported for testing.
|
|
286
|
+
*/
|
|
287
|
+
export declare function unregisteredSkillFiles(files: Record<string, string> | undefined): string[];
|
|
273
288
|
export declare function runEval<M extends Metrics>(spec: EvalSpec<M>): Promise<EvalReport>;
|
|
274
289
|
/** A task run N times, scored against a `Trace` check vocabulary. */
|
|
275
290
|
export interface MeasureSpec {
|
package/dist/eval.js
CHANGED
|
@@ -3,6 +3,7 @@ Object.defineProperty(exports, "__esModule", { value: true });
|
|
|
3
3
|
exports.claudeEvalDriver = exports.EPHEMERAL_HOME_KEEP = void 0;
|
|
4
4
|
exports.resolveSpawnEnv = resolveSpawnEnv;
|
|
5
5
|
exports.spawnAgent = spawnAgent;
|
|
6
|
+
exports.unregisteredSkillFiles = unregisteredSkillFiles;
|
|
6
7
|
exports.runEval = runEval;
|
|
7
8
|
exports.measureWith = measureWith;
|
|
8
9
|
exports.measure = measure;
|
|
@@ -145,7 +146,48 @@ function spawnAgent(a) {
|
|
|
145
146
|
* working model auth (e.g. `ANTHROPIC_API_KEY`). Thin wrapper over
|
|
146
147
|
* `runEvalWith` with the real agent runner.
|
|
147
148
|
*/
|
|
149
|
+
/**
|
|
150
|
+
* A `SKILL.md` written straight into a run's cwd (via an arm's `files`) is NOT
|
|
151
|
+
* registered as a skill by the harness. Claude Code — and Codex — load skills only
|
|
152
|
+
* from a plugin / `.claude/skills` layout, so a bare cwd `SKILL.md` sits unread:
|
|
153
|
+
* the arm silently measures NOTHING (baseline and "skill" become the same run). This
|
|
154
|
+
* is the exact footgun the ecosystem benchmark hit — a skill file delivered where it
|
|
155
|
+
* can never activate, with no error. Detect it so {@link runEval} / {@link measureArms}
|
|
156
|
+
* can WARN and point at `pluginDir` (a real `--plugin-dir` install) or `skillsDir`.
|
|
157
|
+
*
|
|
158
|
+
* HIGH-PRECISION (don't cry wolf): only a file whose basename is `SKILL.md` AND that
|
|
159
|
+
* carries real skill frontmatter (a `---` block naming `name`/`description`) is flagged
|
|
160
|
+
* — so an empty scratch `SKILL.md` a task is asked to AUTHOR is never flagged. Pure +
|
|
161
|
+
* exported for testing.
|
|
162
|
+
*/
|
|
163
|
+
function unregisteredSkillFiles(files) {
|
|
164
|
+
if (files === undefined)
|
|
165
|
+
return [];
|
|
166
|
+
return Object.entries(files)
|
|
167
|
+
.filter(([p, c]) => skillBasename(p) && hasSkillFrontmatter(c))
|
|
168
|
+
.map(([p]) => p);
|
|
169
|
+
}
|
|
170
|
+
function skillBasename(path) {
|
|
171
|
+
return (path.split(/[\\/]/).pop() ?? path) === "SKILL.md";
|
|
172
|
+
}
|
|
173
|
+
function hasSkillFrontmatter(content) {
|
|
174
|
+
const m = /^\uFEFF?\s*---\s*\r?\n([\s\S]*?)\r?\n---/.exec(content);
|
|
175
|
+
return m !== null && /(^|\n)\s*(name|description)\s*:/.test(m[1]);
|
|
176
|
+
}
|
|
177
|
+
/** Warn (loud, non-fatal) for every arm that drops an unregistered skill file. */
|
|
178
|
+
function warnUnregisteredSkillArms(arms) {
|
|
179
|
+
for (const [name, arm] of Object.entries(arms)) {
|
|
180
|
+
for (const path of unregisteredSkillFiles(arm.files)) {
|
|
181
|
+
console.warn(`⚠ eval arm "${name}": files["${path}"] is a SKILL.md with skill ` +
|
|
182
|
+
`frontmatter, but a SKILL.md written to the run cwd is NOT registered as a ` +
|
|
183
|
+
`skill by the harness — it never activates, so this arm measures nothing. ` +
|
|
184
|
+
`Install it via \`pluginDir\` (a real --plugin-dir plugin) or \`skillsDir\`, ` +
|
|
185
|
+
`not \`files\`. See docs/harness-testing.md.`);
|
|
186
|
+
}
|
|
187
|
+
}
|
|
188
|
+
}
|
|
148
189
|
async function runEval(spec) {
|
|
190
|
+
warnUnregisteredSkillArms(spec.arms);
|
|
149
191
|
const report = await runEvalWith(spec, spawnAgent);
|
|
150
192
|
// Surface what the run spent — tokens + API-equivalent $, and a LOUD warning if
|
|
151
193
|
// it was billed to a metered API key instead of the subscription. See eval-cost.ts.
|
|
@@ -288,6 +330,7 @@ function stubArmPluginDirs(arms) {
|
|
|
288
330
|
/* v8 ignore start -- real claude subprocess; thin wrapper over measureArmsWith */
|
|
289
331
|
/** Score checks across arms against the real `claude` CLI. */
|
|
290
332
|
async function measureArms(spec) {
|
|
333
|
+
warnUnregisteredSkillArms(spec.arms);
|
|
291
334
|
const report = await measureArmsWith(spec, spawnAgent);
|
|
292
335
|
// Sum every arm's spend — an A/B run pays for both arms.
|
|
293
336
|
(0, eval_cost_js_1.emitCostSummary)((0, eval_cost_js_1.sumCosts)(Object.values(report.arms).map((a) => (0, eval_cost_js_1.costFromArm)(a.usage))));
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "vigiles",
|
|
3
|
-
"version": "12.
|
|
3
|
+
"version": "12.3.0",
|
|
4
4
|
"description": "Lint & test the harness your AI agent runs on — verify the references in your CLAUDE.md / AGENTS.md and test that your hooks and skills actually work.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"claude-code",
|
|
@@ -79,7 +79,7 @@
|
|
|
79
79
|
"test:e2e": "npm run build && vitest run --project e2e",
|
|
80
80
|
"test:cli-e2e": "bash test/e2e/run.sh",
|
|
81
81
|
"test:harness": "npm run build && node dist/cli.js test",
|
|
82
|
-
"test:eval": "npm run build && node dist/cli.js eval",
|
|
82
|
+
"test:eval": "npm run build && node dist/cli.js eval --all",
|
|
83
83
|
"test:vitest": "npm run build && vitest run --project runners",
|
|
84
84
|
"test:jest": "npm run build && jest",
|
|
85
85
|
"test:types": "npm run build && tsc --noEmit -p test/types/tsconfig.json",
|