vigiles 14.9.0 → 14.10.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +24 -51
- package/dist/audit-report.d.ts +8 -0
- package/dist/audit-report.js +3 -0
- package/dist/audit-report.template.html +30 -30
- package/dist/audit-verdict.js +17 -5
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -60,6 +60,15 @@
|
|
|
60
60
|
docs/compiled-hooks.md §disler battery + src/hook-dogfood.test.ts) — reachable
|
|
61
61
|
evidence under the Lighthouse umbrella, not the hero, not the brand.
|
|
62
62
|
|
|
63
|
+
WEBSITE (added 2026-07-22): vigiles.sh is the LIVE interactive demo (grade any
|
|
64
|
+
repo in the browser). The README is the AGENT's install door + the GitHub/npm
|
|
65
|
+
landing — it LINKS to the site prominently (hero line + badge + the catches CTA
|
|
66
|
+
+ quick-start + docs) rather than re-embedding the whole proof. The 4 detailed
|
|
67
|
+
Proof narratives were COLLAPSED to a 4-bullet list that leans on the live demo +
|
|
68
|
+
docs/what-vigiles-catches.md (the site + that doc own the depth). Don't re-expand
|
|
69
|
+
them here; deepen them on the site or in docs instead. Split of labour: site =
|
|
70
|
+
marketing/try, README = install + 60-sec what, docs/ = the shared HOW.
|
|
71
|
+
|
|
63
72
|
RULES: lead with the reader's CONCRETE PAIN; ≤ ~3-line paragraphs; ONE bold per
|
|
64
73
|
block; ONE idea per sentence; NO internal vocabulary (moat/flywheel) / NO
|
|
65
74
|
research/ links / NO enterprise/national-interest framing — name the user
|
|
@@ -88,10 +97,15 @@
|
|
|
88
97
|
Verify your CLAUDE.md or AGENTS.md, skills, and hooks are real — and prove they actually work.
|
|
89
98
|
</p>
|
|
90
99
|
|
|
100
|
+
<p align="center">
|
|
101
|
+
<a href="https://vigiles.sh"><strong>▶ Grade any public repo live at vigiles.sh</strong></a> — no install, runs in your browser.
|
|
102
|
+
</p>
|
|
103
|
+
|
|
91
104
|
<p align="center">
|
|
92
105
|
<a href="https://www.npmjs.com/package/vigiles"><img src="https://img.shields.io/npm/v/vigiles?color=orange" alt="npm version" /></a>
|
|
93
106
|
<a href="https://github.com/zernie/vigiles/actions"><img src="https://img.shields.io/github/actions/workflow/status/zernie/vigiles/ci.yml?branch=main" alt="CI" /></a>
|
|
94
107
|
<a href="https://github.com/zernie/vigiles/blob/main/LICENSE"><img src="https://img.shields.io/github/license/zernie/vigiles" alt="License" /></a>
|
|
108
|
+
<a href="https://vigiles.sh"><img src="https://img.shields.io/badge/demo-vigiles.sh-f97316" alt="live demo" /></a>
|
|
95
109
|
</p>
|
|
96
110
|
|
|
97
111
|
---
|
|
@@ -128,55 +142,14 @@ And it closes the loop from prose to enforcement: **your rules → enforced** ma
|
|
|
128
142
|
|
|
129
143
|
These are real scans of public plugins — run `npx vigiles audit <any-repo>` for your own. The examples below use Claude Code subagents; the same checks run on Codex `AGENTS.md`, skills, and hooks. ↓
|
|
130
144
|
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
```text
|
|
134
|
-
✗ tester — Tool "AskUserQuestion" is never available to a subagent.
|
|
135
|
-
→ remove or correct it — it's silently dropped from the contract.
|
|
136
|
-
```
|
|
137
|
-
|
|
138
|
-
This subagent — a helper your main agent hands work to — lists a tool that doesn't exist for it. The harness drops it without a word, so the agent quietly loses a capability it thinks it has. The markdown is perfectly valid. vigiles catches it and hands you the **one-line fix**.
|
|
139
|
-
|
|
140
|
-
## Proof 2 — two skills your agent can't tell apart
|
|
141
|
-
|
|
142
|
-
```text
|
|
143
|
-
✗ Triggering 0 (0/100)
|
|
144
|
-
└ 45 pairs of near-identical skill descriptions — the agent can't tell them
|
|
145
|
-
apart, so the wrong one fires (e.g. "agent-coder" ↔ "agent-tester", 83% alike)
|
|
146
|
-
```
|
|
147
|
-
|
|
148
|
-
One popular plugin ships **45 pairs** of near-identical skill descriptions. Your agent picks a skill by _reading_ them — so when two match, it fires the wrong one. Still perfectly valid markdown.
|
|
149
|
-
**[How triggering works →](docs/measuring-skills.md)**
|
|
150
|
-
|
|
151
|
-
## Proof 3 — a rule you wrote that nothing enforces
|
|
152
|
-
|
|
153
|
-
Point vigiles at your own repo and the rules section maps each prose rule to the lint rule that enforces it — then checks your config. Representative output:
|
|
154
|
-
|
|
155
|
-
```text
|
|
156
|
-
Your rules → enforced 1 of 4 enforced · 2 one line away · 1 contradicted by config
|
|
157
|
-
└ "always use ===" → eqeqeq is set to "off" in your ESLint config
|
|
158
|
-
your CLAUDE.md says enforce it; your config quietly turns it off
|
|
159
|
-
```
|
|
160
|
-
|
|
161
|
-
You wrote the rule. Your agent treats it as gospel and follows it — until it doesn't, and nothing tells you which time. vigiles checks each mapped rule three ways: enforced, **one line away**, or — the one people screenshot — documented but silently **turned off**. Deterministic, no model. Acting on the map is **opt-in and agent-driven**: enable a rule in one line (the `strengthen` skill does it for you), turn an action rule like `git push` into a compiled hook, leave judgment calls as prose. Nothing runs a model — or changes your config — unless you ask.
|
|
162
|
-
**[How enforcement works →](docs/verifying-instruction-files.md)**
|
|
163
|
-
|
|
164
|
-
## Proof 4 — it can quietly read your secrets and send them out
|
|
165
|
-
|
|
166
|
-
```text
|
|
167
|
-
◑ Safety 80 (80/100)
|
|
168
|
-
└ subagent "tester" holds all three lethal-trifecta legs:
|
|
169
|
-
reads private data (Bash, Read) · takes in untrusted web content (WebFetch)
|
|
170
|
-
· can send data out (Bash, WebFetch)
|
|
171
|
-
```
|
|
172
|
-
|
|
173
|
-
Hand one subagent all three powers and a poisoned web page can make it read your `.env` and POST it anywhere — no exploit code, just the tools it was given. The **80 looks like a B** — and that's the trap: a healthy grade hiding a subagent that's a data-leak waiting to happen. vigiles spots it from the tool list alone, free, no model.
|
|
145
|
+
Every one of these is **valid markdown** — parses fine, does the wrong thing. That's the gap a style linter can't see:
|
|
174
146
|
|
|
175
|
-
|
|
176
|
-
**
|
|
147
|
+
- **A tool your agent thinks it has and doesn't** — a subagent lists a tool that isn't available to it. The harness drops it silently; the agent loses a capability it believes it has.
|
|
148
|
+
- **Two skills your agent can't tell apart** — near-identical descriptions, so the agent (which picks by _reading_ them) fires the wrong one. One popular plugin ships 45 such pairs.
|
|
149
|
+
- **A rule you wrote that nothing enforces** — `"always use ==="` while `eqeqeq` is set to `"off"` in your ESLint config. Your CLAUDE.md says enforce it; your config quietly turns it off.
|
|
150
|
+
- **A subagent that can read your secrets and send them out** — one holding all three "lethal-trifecta" legs (reads data · takes untrusted web input · can send data out). A poisoned page makes it POST your `.env` anywhere — no exploit, just the tools it was given, hidden behind a healthy-looking grade.
|
|
177
151
|
|
|
178
|
-
|
|
179
|
-
**[Everything it catches →](docs/what-vigiles-catches.md)** · point `audit` at a whole marketplace and it ranks every plugin the same way.
|
|
152
|
+
**[▶ See these live at vigiles.sh](https://vigiles.sh)** — grade any repo · **[everything it catches →](docs/what-vigiles-catches.md)**. Point `audit` at a whole marketplace and it ranks every plugin the same way.
|
|
180
153
|
|
|
181
154
|
## How it works — vibes → verified
|
|
182
155
|
|
|
@@ -216,7 +189,7 @@ Point it at any harness change that claims a number — does a compression skill
|
|
|
216
189
|
|
|
217
190
|
## Quick start
|
|
218
191
|
|
|
219
|
-
**1. See what's broken** — read-only, no setup:
|
|
192
|
+
**1. See what's broken** — read-only, no setup (or try it in your browser first at **[vigiles.sh](https://vigiles.sh)**):
|
|
220
193
|
|
|
221
194
|
```bash
|
|
222
195
|
npx vigiles audit
|
|
@@ -265,8 +238,8 @@ Targets Claude Code and Codex out of the box, or [your own harness](docs/authori
|
|
|
265
238
|
|
|
266
239
|
- **Is this a framework I have to build around?** No. It's a tool you run — like ESLint, Lighthouse, or `npm audit`. One command, a report, an optional CI gate. There's a library API for automation, but you never touch it to get value.
|
|
267
240
|
- **Isn't this just a markdown linter?** No — it checks whether your instruction file is _true_ (every path/script/symbol/rule exists and is enabled), then tests and measures your harness. A style linter can't do any of that.
|
|
268
|
-
- **
|
|
269
|
-
- **
|
|
241
|
+
- **Does anything leave my machine?** No. `audit` and `lint` read your local repo — no upload, no account, no server. (`eval` is the only step that calls a model, on your own Claude subscription; the browser demo only reads a public repo you name via GitHub's API.)
|
|
242
|
+
- **What do I have to change to adopt it?** Almost nothing. Plain markdown works with zero new files, rules run in your existing linter, and the agent edits the specs for you. The typed spec is opt-in, only for structural checks a linter can't express — like TS's `strict` ([why?](docs/faq.md#why-are-the-strongest-guarantees-opt-in-not-the-default)).
|
|
270
243
|
- **Non-JS repo?** `npx vigiles lint` verifies your CLAUDE.md or AGENTS.md with no install (Ruff/Clippy/Pylint/… too).
|
|
271
244
|
|
|
272
245
|
**[Full FAQ →](docs/faq.md)**
|
|
@@ -275,7 +248,7 @@ Targets Claude Code and Codex out of the box, or [your own harness](docs/authori
|
|
|
275
248
|
|
|
276
249
|
## Docs
|
|
277
250
|
|
|
278
|
-
The **[docs index](docs/README.md)** is the full map, grouped by what you're doing:
|
|
251
|
+
**[vigiles.sh](https://vigiles.sh)** is the live demo — grade any repo in your browser. The **[docs index](docs/README.md)** is the full map, grouped by what you're doing:
|
|
279
252
|
|
|
280
253
|
- **Guides** — [verify instruction files](docs/verifying-instruction-files.md) · [test your harness](docs/harness-testing.md) · [measure a skill](docs/measuring-skills.md) · [ship a plugin](docs/for-plugin-authors.md) · [Codex & other harnesses](docs/harnesses.md)
|
|
281
254
|
- **Reference** — [CLI](docs/cli.md) · [rules matrix](docs/verifying-instruction-files.md#the-validation-rules--the-full-matrix) · [testing API](docs/testing-api.md) · [full API](https://zernie.github.io/vigiles/api/)
|
package/dist/audit-report.d.ts
CHANGED
|
@@ -98,6 +98,14 @@ export interface AuditReport {
|
|
|
98
98
|
/** The deterministic, ranked fixes (the inline recommendations). */
|
|
99
99
|
readonly recommendations: readonly Recommendation[];
|
|
100
100
|
readonly inventory: AuditInventory;
|
|
101
|
+
/**
|
|
102
|
+
* The CONCRETE intra-plugin references that don't resolve on disk — the actual
|
|
103
|
+
* paths behind the Truthfulness category's "N broken intra-plugin reference(s)"
|
|
104
|
+
* count, so the report can show WHAT is broken (a file path), not just how many.
|
|
105
|
+
* Present only when at least one dangling ref exists. Additive/optional — schema
|
|
106
|
+
* version unchanged.
|
|
107
|
+
*/
|
|
108
|
+
readonly brokenReferences?: readonly string[];
|
|
101
109
|
/**
|
|
102
110
|
* The adoption preview — "what would vigiles catch in your repo?" Present only
|
|
103
111
|
* when the model-gated tier ran (behind consent); a deterministic read omits it.
|
package/dist/audit-report.js
CHANGED
|
@@ -79,6 +79,9 @@ function buildAuditReport(report, opts) {
|
|
|
79
79
|
mcp: report.mcp,
|
|
80
80
|
untested: report.untested,
|
|
81
81
|
},
|
|
82
|
+
...(report.danglingRefs.length
|
|
83
|
+
? { brokenReferences: report.danglingRefs }
|
|
84
|
+
: {}),
|
|
82
85
|
...(adoptable ? { adoptable } : {}),
|
|
83
86
|
...(opts.observations ? { observations: opts.observations } : {}),
|
|
84
87
|
...(opts.rulesInventory && opts.rulesInventory.length
|