skilltrigger 0.0.0-stage → 0.0.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md ADDED
@@ -0,0 +1,86 @@
1
+ # Changelog
2
+
3
+ All notable changes to skilltrigger. The format is [Keep a Changelog](https://keepachangelog.com/en/1.1.0/);
4
+ versions follow [SemVer](https://semver.org/). Items reference their `ST-n` backlog id.
5
+
6
+ ## [Unreleased]
7
+
8
+ ## [0.0.2] — 2026-10-01
9
+
10
+ Tested only against the fake `claude` the tests use, never against a live CLI — the
11
+ first live run is the gate on 0.1.0 (ST-11, ST-12). The first version on npm, published
12
+ by hand so that trusted publishing can be configured (ST-13); it carries 0.0.1, which was
13
+ never released, and the entries below.
14
+
15
+ ### Added
16
+ - The report records what `compare` needs to tell two environments apart: the roster's
17
+ members beside its counts, each as the first 12 hex digits of its name's SHA-256 — a
18
+ roster names private skills, and a report is meant to be shared; the model each run's `init` event
19
+ reported, per run and counted in `runModels`; and a count of what the inherited
20
+ environment contributed — memory files for the temporary project and for the user,
21
+ hooks configured, MCP servers in the `init` event — never their contents, paths or
22
+ names (ST-20).
23
+ - A re-enable reminder. The toggles the conflict gate prints are remembered in
24
+ `.skilltrigger-toggles` under the `--out` directory, never under `~/.claude`; every
25
+ `preflight` that sees such a plugin still disabled, and the end of every `run`, prints
26
+ its `claude plugin enable …` command, until a preflight sees it enabled again or
27
+ uninstalled. `preflight` takes `--out` for it, defaulting to `./skilltrigger-results`
28
+ as `run` does. Nothing is enabled or disabled by skilltrigger itself (ST-21).
29
+
30
+ ### Changed
31
+ - The backlog check, the roadmap, the issue sync and the release-drift check are
32
+ [backlogsync](https://github.com/Allan-Nava/backlogsync), pinned by commit;
33
+ `scripts/backlog.mjs` and its test are gone, and `npm run roadmap` regenerates
34
+ `ROADMAP.md` (ST-22).
35
+ - `compare` warns when roster members differ, counting the members added and removed by hash, not
36
+ only when the counts differ; and when the models the runs reported or the inherited
37
+ environment's counts differ. Reports without the new fields compare as before (ST-20).
38
+
39
+ ### Fixed
40
+ - A verdict no longer survives losing whole queries. The 10% rule counts runs, so two
41
+ positive queries could lose every run to timeouts or errors and the report still gave
42
+ a verdict on a smaller positives total. A query that was run and measured nothing is
43
+ now no verdict, exit 3, with the lost queries named in the headline and recorded as
44
+ `lostQueries`; a query that lost some but not all of its runs keeps the verdict and is
45
+ flagged, in `partialQueries`, in its row and under the table (ST-19).
46
+ - The conflict gate no longer passes when it cannot read the plugin list. Output from
47
+ `claude plugin list --json` in a shape the parser does not know used to read as zero
48
+ plugins and the gate said `ok` with nothing checked; it now fails, quoting the first
49
+ line of what it got, truncated. A valid empty JSON array, or a human-form list that
50
+ says no plugins are installed, is still zero plugins and passes; the human-form
51
+ fallback for an older CLI without `--json` stays. `--allow-conflict` lets an
52
+ unreadable list through as a recorded warning (ST-18).
53
+ - `release.yml` closes only the milestone titled `v<version>`, alone or followed by a
54
+ space and a subtitle. It matched any title starting with `v<version>`, so a 0.0.2 tag
55
+ could have closed a `v0.0.20` milestone, and 0.1.1 one called `v0.1.10` (ST-13).
56
+
57
+ ## [0.0.1] — 2026-10-01 — not released
58
+
59
+ The first version, in the repository only: the gates, the runner, the report and their
60
+ tests against a fake `claude`. Nothing has been measured with a real model yet; the
61
+ first version on npm is 0.1.0, after that measurement (ST-11).
62
+
63
+ ### Added
64
+ - `skilltrigger preflight`: six gates — `cli` (on PATH, version recorded), `auth`
65
+ (`claude auth status`), `round-trip` (`claude -p "Reply with exactly: pong"` with the
66
+ chosen model), `conflict` (an enabled plugin or visible skill of the same name; the
67
+ disable and enable commands printed, never run; `--allow-conflict` records it),
68
+ `sleep` (`caffeinate -i -s -w <pid>` on macOS, a warning elsewhere) and `roster`
69
+ (slash commands and skills from the `init` event). Exit 2 on any failure (ST-2 to ST-5).
70
+ - `skilltrigger run`: strictly serial; one temporary project and one stub command per
71
+ run, removed afterwards; detection from the partial stream events, the process
72
+ stopped at the decision; four outcomes — `triggered`, `not-triggered`, `timeout`,
73
+ `error` — with an authentication or API error never scored as a miss, and a run whose
74
+ `init` event does not list the stub scored as an error (ST-6).
75
+ - The verdict rule: more than 10% of runs timing out or failing is no verdict, exit 3,
76
+ and the run stops as soon as that share is passed. Reports in JSON and Markdown that
77
+ carry the queries, the counts and the versions, and no description text, path or stub
78
+ name (ST-7).
79
+ - `skilltrigger compare`: per query and total differences over the shared queries, a
80
+ warning when model, CLI version or roster differ, ±1 per query at two runs labelled
81
+ noise (ST-8).
82
+ - A fake `claude` driven by environment variables, and tests for every gate, every
83
+ outcome, the no-verdict rule, compare's drift warning and the removal of the
84
+ temporary project (ST-9).
85
+ - `skilltrigger check`, CI on Node 18/20/22/24, the release workflow over npm trusted
86
+ publishing, the site, the backlog and its generated roadmap (ST-10).
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Allan Nava
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md CHANGED
@@ -1,3 +1,110 @@
1
- # Temporary Holding Version
1
+ <p align="center"><img src="https://raw.githubusercontent.com/Allan-Nava/skilltrigger/main/assets/logo.svg" width="72" height="72" alt=""></p>
2
2
 
3
- This version is a temporary placeholder for this package. An operational version to replace this has been submitted for review and is awaiting a staged release.
3
+ # skilltrigger — skill trigger rates you can trust
4
+
5
+ **skilltrigger measures how often a Claude Code skill's `description` makes the model load it** — on prompts that should trigger it and on near-misses that should not — and refuses to report a number it cannot trust. Six things have been seen to turn that measurement into a clean-looking zero that means nothing; each one is a gate it must pass first or a field recorded beside the result.
6
+
7
+ **Status:** 0.0.2, on npm — an early version that **has not yet been run against a live `claude` CLI**. Every number skilltrigger has produced so far came from the fake `claude` its tests use, and the stream shapes it reads are unconfirmed on a real run. The gate on 0.1.0 is one real measurement: the maintainer runs skilltrigger on the three skills of the [qrspi](https://github.com/Allan-Nava/qrspi) plugin, with their `evals/trigger/*.json` sets, and the result is recorded, dated, beside the 2026-09-18 numbers in qrspi's CONTRIBUTING.
8
+
9
+ ## What it measures
10
+
11
+ A skill's description sits in the model's context permanently; the model reads it and decides whether to load the skill. Whether a rewrite helps is an empirical question, and the answer is a pair of rates:
12
+
13
+ - **positives triggered** — of the runs on prompts that should load the skill, how many did;
14
+ - **negatives fired** — of the runs on near-misses that should not, how many loaded it anyway.
15
+
16
+ Each run is one `claude -p` in a fresh temporary project that holds a single stub command whose frontmatter `description` is the text under test. The stub is **loaded** when the model's first decisive message carries a tool call that names it — `Skill` or `SlashCommand` with the stub's name in its input, or a `Read` of the stub file. A first message that only fetches tool schemas (`ToolSearch`) does not decide; the next one does. Anything else first — a text answer, another tool, another skill — is a run that did **not** trigger. Detection happens on the partial stream events, and the process is stopped at the decision.
17
+
18
+ Each run ends in one of four outcomes: `triggered`, `not-triggered`, `timeout`, `error`. Only the first two are measurements. A timeout or an error is never counted as "did not trigger", and when more than 10% of the runs are either, there is **no verdict**: the report says so and the command exits 3. The share counts runs, so it is not the only rule: a query that lost *every* run to timeouts or errors is also no verdict, with the lost queries named — otherwise two positives could vanish inside the 10% and the positives total would quietly leave them out. A query that lost some of its runs but not all keeps the verdict, and the report flags it as measured on fewer runs than planned.
19
+
20
+ ## The six traps
21
+
22
+ Each was hit in practice, with skill-creator's harness, between 2026-09-09 and 2026-09-18. Each produced a number that looked like a result.
23
+
24
+ 1. **Parallel workers.** The old harness injects one stub per worker, all with the same description and names that differ only by a hash. The model invokes whichever stub it likes, and only the worker whose stub was picked counts the hit — so ten workers, the default, measure about a tenth of the true rate. *skilltrigger is strictly serial and has no workers option*; `--workers` is rejected as an unknown option rather than ignored. A detected run ends in seconds; serial is fast enough.
25
+ 2. **An installed plugin shadowing the stub.** If a plugin that carries a skill of the same name is enabled, the model loads the real skill and the stub is never seen: the rate reads zero. *The conflict gate* reads `claude plugin list --json`, looks for the skill's `name` among each enabled plugin's skills, and cross-checks the roster the model actually sees for a same-named entry. On a hit it refuses and prints the exact `claude plugin disable …` and `claude plugin enable …` commands, scope included — and runs neither. A list it cannot read is not an empty list: output in a shape it does not know — not a JSON array, not the human form an older CLI prints, not a line saying no plugins are installed — fails the gate with the first line of what it got, because "which plugins are enabled" is then unknown. `--allow-conflict` measures anyway and records in the report that it did. Because a plugin disabled as asked is no longer a conflict, the toggles the gate printed are remembered in a small file, `.skilltrigger-toggles`, under the `--out` directory — never under `~/.claude` — and every `preflight` that sees the plugin still disabled, and the end of every `run`, prints its `claude plugin enable …` command again, until a preflight sees it enabled or uninstalled. Independently, every run checks that the stub is in the roster its `init` event lists; a run where it is not is an `error`.
26
+ 3. **An outdated CLI.** An old `claude` answers 400 for a model it does not know, and the old harness scored that as "did not trigger". *The round-trip gate* runs `claude -p "Reply with exactly: pong"` with the chosen `--model` and requires the answer `pong`; the CLI version is recorded in the report, and in every run an API error line is an `error`.
27
+ 4. **An expired login.** An expired OAuth session makes every `claude -p` answer "Failed to authenticate", and a harness that discards stderr scores all forty runs as misses: 0/18 positives and 0/22 negatives, which is not a measurement. *The auth gate* requires `claude auth status` to say logged in, and the round trip catches the expired token that `auth status` can miss. In a run, an authentication failure is an `error`, never a miss.
28
+ 5. **A machine that sleeps.** Each run has a timeout, and a laptop that sleeps mid-eval turns runs into timeouts — once 5/18 for a description measured 17/18 a week earlier. *The sleep gate* starts `caffeinate -i -s -w <pid>` on macOS, which holds off idle and system sleep until skilltrigger exits; elsewhere it prints a warning and the advice. Timeouts are counted apart, and they count towards the no-verdict rule.
29
+ 6. **The skill roster changing the number.** The stub competes with every other description the model can see. A byte-identical description measured 17/18 one week and 10/18 the next because a synced folder had added 24 skills to the machine, from 59 to 83. *The roster gate* reads the `init` event of a stub-less `claude -p --output-format stream-json --verbose` and records how many slash commands and skills the model sees and their names, and `compare` warns when two reports differ in roster — by its members, not only its size — model or CLI version.
30
+
31
+ ## Install
32
+
33
+ skilltrigger needs Node 18 or later and a logged-in [Claude Code](https://code.claude.com/docs) CLI on `PATH`. It has no runtime dependencies.
34
+
35
+ From npm, without installing:
36
+
37
+ ```bash
38
+ npx skilltrigger preflight
39
+ ```
40
+
41
+ or installed, so the `skilltrigger` command is on `PATH`:
42
+
43
+ ```bash
44
+ npm install -g skilltrigger
45
+ skilltrigger preflight
46
+ ```
47
+
48
+ From a checkout:
49
+
50
+ ```bash
51
+ git clone https://github.com/Allan-Nava/skilltrigger && cd skilltrigger
52
+ node bin/skilltrigger.mjs preflight
53
+ ```
54
+
55
+ ## Usage
56
+
57
+ ```bash
58
+ skilltrigger preflight [--model M] [--skill <dir>] [--out <dir>] [--allow-conflict]
59
+ skilltrigger run --skill <dir> --eval <file> [--runs 2] [--model M] [--timeout 30] \
60
+ [--description "<override>"] [--out <dir>] [--allow-conflict]
61
+ skilltrigger compare <a.json> <b.json>
62
+ skilltrigger check
63
+ ```
64
+
65
+ `preflight` runs every gate and prints each as `ok`, `warn`, `fail` or `skip`, with the reason and, for anything not ok, the fix. It exits 2 on any failure. Without `--skill` it has no name to look for, so the conflict gate only lists the enabled plugins. `--out` is where the re-enable reminder is kept (below); it defaults to `./skilltrigger-results`, as for `run`.
66
+
67
+ `run` runs the preflight first and refuses on a failure, then runs every query `--runs` times (default 2), round-robin, serially, with a `--timeout` in seconds per run (default 30). `--description` measures an override instead of the text in `SKILL.md`, so a rewrite can be tried without editing the skill. Reports go to `--out` (default `./skilltrigger-results`). Exit codes: 0 a verdict, 1 a usage or input error, 2 a failed gate, 3 no verdict.
68
+
69
+ The eval set is skill-creator's format — a JSON array of `{ "query": string, "should_trigger": boolean }` — and any other field, a `note` for instance, is carried through to the report:
70
+
71
+ ```json
72
+ [
73
+ { "query": "summarise this session into HANDOFF.md before I clear it", "should_trigger": true },
74
+ { "query": "why is cache_read_input_tokens 0 on every call", "should_trigger": false, "note": "the neighbouring skill's territory" }
75
+ ]
76
+ ```
77
+
78
+ `compare` sets two reports side by side: per query and in total, over the queries both share. It warns first when the model, the models the runs reported, the CLI version, the roster or the inherited environment differ, because the number is a property of the text *and* the roster. Two rosters of the same size with different members are a warning that counts the members added and removed, by hash. Two runs per query resolve to ±1 per query, so a difference of one hit at two runs is labelled noise, and so is a total difference of up to two.
79
+
80
+ To judge a rewrite, measure the baseline the same day, in the same environment, then the rewrite with `--description`, and compare the two.
81
+
82
+ ## The report
83
+
84
+ Each run writes `<out>/<date>-<skill>.json` and a Markdown rendering beside it (a second report the same day gets `-2`). It holds the date, the skilltrigger and CLI versions, the model, the model each run's `init` event reported (per run, and counted in `runModels`), the roster — its counts and a hash of each member, sorted — a count of what the inherited environment contributed (memory files for the temporary project and for the user, hooks configured, MCP servers in the `init` event), runs per query, the timeout, the pass threshold (a trigger rate of 0.5) and the no-verdict threshold (10%), each query with every outcome and its hits/runs, and the totals — positives triggered, negatives fired, timeouts and errors counted apart — plus the queries that lost every run (`lostQueries`, which make it no verdict) and those measured on fewer runs than planned (`partialQueries`).
85
+
86
+ Nothing else is in it: no description text (a byte count and a SHA-256 stand in for it, so `compare` can tell two texts apart), no paths, no command, skill or stub names, no stderr, and of the environment only counts — never a memory file's contents, a hook command or a server name. The roster's members are the commands and skills installed on the machine that measured, private ones included, so each is recorded as the first 12 hex digits of its name's SHA-256: enough for `compare` to see that a member moved, not to say which. A hash hides a name from a reader, not from someone guessing it — a name anyone could guess is still guessable.
87
+
88
+ ## What it never does
89
+
90
+ - **It never touches `~/.claude`.** Every run happens in its own temporary directory, removed afterwards; nothing is written under the user's configuration, and every `claude -p` gets `--no-session-persistence`, so the runs are not saved as sessions either. The conflict gate reads a plugin's install directory and writes nothing there.
91
+ - **It never runs in parallel.** See the first trap.
92
+ - **It never reports a number through a failed gate.** A failed gate refuses the run; more than 10% of runs timing out or failing is no verdict, and so is a query that lost every run; the run stops as soon as that share is passed, because every further run would be spent on a number that will not be reported.
93
+ - **It never disables or enables a plugin itself.** It prints the commands, and reminds you of the enable command until the plugin is back.
94
+ - It never sends anything anywhere but through the `claude` CLI it is measuring with.
95
+
96
+ ## Design notes
97
+
98
+ **Why the first decisive message.** A skill is useful when the model loads it before doing the work. A model that goes exploring first and loads the skill three tool calls later has, for the purpose of a description, missed — and an open-ended wait would turn every such run into a timeout. skill-creator stops at the first tool call of any other kind; skilltrigger reads the whole first message, so a `Bash` call listed before the `Skill` call in the same message still counts as a load.
99
+
100
+ **A known limit of the method.** `claude -p` starts from nothing, so a positive that presupposes session history — "dump the state of this refactor" — measures whether the model loads the skill *before* going to look for the material. In a real session the material is already in context. Annotate such prompts in the eval set rather than dropping them, so the number stays honest about what it covers.
101
+
102
+ **Why exit 3 is not a failure of the description.** No verdict means the environment broke, not that the text is bad. The report is still written, so the outcomes can be read, and it says at the top that they are not a measurement.
103
+
104
+ ## Prior art
105
+
106
+ The mechanism — a stub command carrying the description, `claude -p --output-format stream-json --include-partial-messages`, and early detection from the stream events — comes from `scripts/run_eval.py` in the `skill-creator` skill of [Anthropic's skills repository](https://github.com/anthropics/skills), Apache-2.0. skilltrigger is a separate implementation that copies no code from it, reads the same eval-set format, and adds the gates, the four outcomes and the no-verdict rule. The six traps are documented, with the numbers above, in [qrspi's CONTRIBUTING](https://github.com/Allan-Nava/qrspi/blob/main/CONTRIBUTING.md).
107
+
108
+ ## License
109
+
110
+ MIT — see [LICENSE](LICENSE).
@@ -0,0 +1,37 @@
1
+ // A small argument parser: `util.parseArgs` arrived in Node 18.3 and the floor is 18.0.
2
+ //
3
+ // spec: { name: 'string' | 'boolean' }. Unknown options are errors, not ignored — a
4
+ // `--workers 4` that was silently dropped would look like it had been honoured.
5
+
6
+ export class UsageError extends Error {}
7
+
8
+ export function parseArgs(argv, spec) {
9
+ const opts = {}
10
+ const positional = []
11
+ for (let i = 0; i < argv.length; i++) {
12
+ const a = argv[i]
13
+ if (!a.startsWith('--')) {
14
+ positional.push(a)
15
+ continue
16
+ }
17
+ let [name, inline] = a.slice(2).split(/=(.*)/s)
18
+ const type = spec[name]
19
+ if (!type) throw new UsageError(`unknown option --${name}`)
20
+ if (type === 'boolean') {
21
+ if (inline !== undefined) throw new UsageError(`--${name} takes no value`)
22
+ opts[name] = true
23
+ continue
24
+ }
25
+ const value = inline ?? argv[++i]
26
+ if (value === undefined || (inline === undefined && value.startsWith('--'))) throw new UsageError(`--${name} needs a value`)
27
+ opts[name] = value
28
+ }
29
+ return { opts, positional }
30
+ }
31
+
32
+ export function number(opts, name, { fallback, min, integer = false }) {
33
+ if (opts[name] === undefined) return fallback
34
+ const n = Number(opts[name])
35
+ if (!Number.isFinite(n) || n < min || (integer && !Number.isInteger(n))) throw new UsageError(`--${name} must be ${integer ? 'an integer' : 'a number'} of at least ${min}, got "${opts[name]}"`)
36
+ return n
37
+ }
@@ -0,0 +1,37 @@
1
+ // The CHANGELOG, read. Two pure functions: the section a release's notes open
2
+ // with, and the rule that a Breaking entry is the first one under its heading — so the
3
+ // line an upgrader most needs is the first line they read, in the file and in the notes.
4
+ // No filesystem here; `check` and `scripts/release-notes.mjs` pass the text in.
5
+
6
+ const esc = (s) => s.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')
7
+
8
+ // The body of `## [version] — date`: everything after its heading line up to the next
9
+ // `## [`, trimmed; null when the version has no section. A prefix is not a match
10
+ // ("0.2" does not find "0.2.0").
11
+ export function changelogSection(text, version) {
12
+ const lines = String(text).split('\n')
13
+ const start = lines.findIndex((l) => new RegExp(`^## \\[${esc(version)}\\](?:\\s|$)`).test(l))
14
+ if (start < 0) return null
15
+ let end = lines.findIndex((l, i) => i > start && l.startsWith('## ['))
16
+ if (end < 0) end = lines.length
17
+ return lines.slice(start + 1, end).join('\n').trim()
18
+ }
19
+
20
+ // Every `### Heading` block, in every `## [...]` section, whose Breaking entry (a bullet
21
+ // starting `- **Breaking`) is not its first bullet: [{ section, heading }].
22
+ export function breakingOutOfPlace(text) {
23
+ const out = []
24
+ let section = null
25
+ let heading = null
26
+ let bullets = 0
27
+ for (const line of String(text).split('\n')) {
28
+ const s = line.match(/^## \[([^\]]+)\]/)
29
+ if (s) { section = s[1]; heading = null; bullets = 0; continue }
30
+ const h = line.match(/^### (.+?)\s*$/)
31
+ if (h) { heading = h[1]; bullets = 0; continue }
32
+ if (!section || !heading || !line.startsWith('- ')) continue
33
+ bullets++
34
+ if (bullets > 1 && line.startsWith('- **Breaking')) out.push({ section, heading })
35
+ }
36
+ return out
37
+ }
@@ -0,0 +1,106 @@
1
+ // `skilltrigger check` — the repository's own invariants, run by `npm test`.
2
+ //
3
+ // - package.json and CHANGELOG.md agree on the version; [Unreleased] exists; a
4
+ // Breaking entry leads its heading; release.yml builds its notes from the CHANGELOG
5
+ // - zero runtime dependencies, and package.json#files ships what the CLI needs
6
+ // - the README names each of the six traps by its phrase (bin/lib/traps.mjs)
7
+ // - no tracked file carries a home-directory path or an email address
8
+ import { execFileSync } from 'node:child_process'
9
+ import { existsSync, readFileSync, readdirSync, statSync } from 'node:fs'
10
+ import { join, relative } from 'node:path'
11
+ import { breakingOutOfPlace } from './changelog.mjs'
12
+ import { TRAPS } from './traps.mjs'
13
+
14
+ // skilltrigger:allow-private-shapes — this file defines the patterns it forbids.
15
+ // A home path: /Users/<name> or /home/<name>, the name a real one (a placeholder in
16
+ // angle brackets does not match). An email: any address but git's SSH user.
17
+ const SHAPES = [
18
+ [/(?:^|[\s"'`(=:])\/(?:Users|home)\/[A-Za-z0-9._-]+/m, 'a home directory path'],
19
+ [/\b(?!git@)[A-Za-z0-9._%+-]+@[A-Za-z0-9-]+(?:\.[A-Za-z0-9-]+)*\.[A-Za-z]{2,}\b/, 'an email address'],
20
+ ]
21
+ export const ALLOW_SHAPES = 'skilltrigger:allow-private-shapes'
22
+
23
+ // → ["path:line looks like <kind> …"] — the finding names the file and the kind, never
24
+ // the match: an error message is printed, logged by CI and pasted into issues.
25
+ export function privateFindings(files) {
26
+ const out = []
27
+ for (const { path, text } of files) {
28
+ if (text.includes(ALLOW_SHAPES)) continue
29
+ for (const [re, what] of SHAPES) {
30
+ const m = text.match(re)
31
+ if (!m) continue
32
+ const line = text.slice(0, m.index + (m[0].startsWith('/') ? 0 : 1)).split('\n').length
33
+ out.push(`${path}:${line} looks like ${what} — this repository is public`)
34
+ break
35
+ }
36
+ }
37
+ return out
38
+ }
39
+
40
+ const SKIP_DIRS = new Set(['.git', 'node_modules', 'dist'])
41
+ const BINARY = /\.(png|jpe?g|gif|ico|webp|woff2?|tgz|gz|zip|pdf)$/i
42
+
43
+ function* walk(root, dir = root) {
44
+ for (const name of readdirSync(dir).sort()) {
45
+ if (SKIP_DIRS.has(name)) continue
46
+ const p = join(dir, name)
47
+ const st = statSync(p)
48
+ if (st.isDirectory()) yield* walk(root, p)
49
+ else yield relative(root, p)
50
+ }
51
+ }
52
+
53
+ // Tracked files plus untracked ones git would add (so a check before the first commit
54
+ // sees the tree); outside a git checkout, the tree walked.
55
+ function trackedFiles(root) {
56
+ try {
57
+ const out = execFileSync('git', ['ls-files', '--cached', '--others', '--exclude-standard'], { cwd: root, encoding: 'utf8', stdio: ['ignore', 'pipe', 'ignore'] })
58
+ const top = execFileSync('git', ['rev-parse', '--show-toplevel'], { cwd: root, encoding: 'utf8', stdio: ['ignore', 'pipe', 'ignore'] }).trim()
59
+ if (relative(top, root) === '') return out.split('\n').filter(Boolean)
60
+ } catch {}
61
+ return [...walk(root)]
62
+ }
63
+
64
+ export function checkRepo(root) {
65
+ const errors = []
66
+ const fail = (m) => errors.push(m)
67
+ const read = (f) => readFileSync(join(root, f), 'utf8')
68
+ const has = (f) => existsSync(join(root, f))
69
+
70
+ let pkg = {}
71
+ try {
72
+ pkg = JSON.parse(read('package.json'))
73
+ } catch (e) {
74
+ fail(`package.json: ${e.message}`)
75
+ }
76
+ if (pkg.name !== 'skilltrigger') fail('package.json#name must be skilltrigger')
77
+ if (!/^\d+\.\d+\.\d+$/.test(pkg.version ?? '')) fail(`package.json#version is not x.y.z: ${pkg.version}`)
78
+ if (pkg.dependencies && Object.keys(pkg.dependencies).length) fail('no runtime dependencies — package.json#dependencies must be empty')
79
+ for (const f of ['bin', 'README.md', 'CHANGELOG.md', 'LICENSE']) if (!pkg.files?.includes(f)) fail(`package.json#files is missing ${f}`)
80
+ if (!/^(?:git\+)?https:\/\/github\.com\/Allan-Nava\/skilltrigger(?:\.git)?$/.test(pkg.repository?.url ?? pkg.repository ?? '')) fail('package.json#repository must be the GitHub repository URL')
81
+
82
+ if (!has('CHANGELOG.md')) fail('CHANGELOG.md is missing')
83
+ else {
84
+ const log = read('CHANGELOG.md')
85
+ if (!/^## \[Unreleased\]/m.test(log)) fail('CHANGELOG.md needs an [Unreleased] section')
86
+ if (pkg.version && !log.includes(`## [${pkg.version}]`)) fail(`CHANGELOG.md has no section for ${pkg.version}, the version in package.json`)
87
+ for (const b of breakingOutOfPlace(log)) fail(`CHANGELOG.md [${b.section}] ### ${b.heading}: a **Breaking** entry must be the first under its heading`)
88
+ }
89
+ if (has('.github/workflows/release.yml') && !read('.github/workflows/release.yml').includes('scripts/release-notes.mjs')) fail('release.yml must build the notes from the CHANGELOG with scripts/release-notes.mjs')
90
+
91
+ if (!has('README.md')) fail('README.md is missing')
92
+ else {
93
+ const readme = read('README.md').toLowerCase()
94
+ for (const t of TRAPS) if (!readme.includes(t.phrase.toLowerCase())) fail(`README.md must name the trap "${t.phrase}" (bin/lib/traps.mjs)`)
95
+ }
96
+
97
+ const files = []
98
+ for (const f of trackedFiles(root)) {
99
+ if (BINARY.test(f) || !existsSync(join(root, f))) continue
100
+ const st = statSync(join(root, f))
101
+ if (!st.isFile() || st.size > 4 * 1024 * 1024) continue
102
+ files.push({ path: f, text: read(f) })
103
+ }
104
+ errors.push(...privateFindings(files))
105
+ return errors
106
+ }
@@ -0,0 +1,106 @@
1
+ // Spawning `claude`. Two shapes: run to completion and collect the output (the
2
+ // gates), or stream newline-delimited JSON and stop the process the moment a caller
3
+ // has what it needs (the round trip, every run).
4
+ //
5
+ // Every spawn gets the environment minus CLAUDECODE, which a parent Claude Code
6
+ // session sets and which makes a nested `claude -p` refuse to start.
7
+ import { spawn } from 'node:child_process'
8
+ import { lineSplitter } from './stream.mjs'
9
+
10
+ export const claudeBin = (env = process.env) => env.SKILLTRIGGER_CLAUDE || 'claude'
11
+
12
+ export function childEnv(env = process.env) {
13
+ const out = { ...env }
14
+ delete out.CLAUDECODE
15
+ return out
16
+ }
17
+
18
+ // Kill the whole group: the CLI starts helpers (MCP servers, a shell) that would
19
+ // otherwise outlive it and hold the temporary directory open.
20
+ function killTree(child) {
21
+ if (child.exitCode !== null || child.signalCode !== null) return
22
+ try {
23
+ process.kill(-child.pid, 'SIGTERM')
24
+ } catch {
25
+ child.kill('SIGTERM')
26
+ }
27
+ const hard = setTimeout(() => {
28
+ try {
29
+ process.kill(-child.pid, 'SIGKILL')
30
+ } catch {
31
+ try {
32
+ child.kill('SIGKILL')
33
+ } catch {}
34
+ }
35
+ }, 2000)
36
+ hard.unref()
37
+ }
38
+
39
+ function start(args, { cwd, env }) {
40
+ return spawn(claudeBin(env), args, { cwd, env: childEnv(env), stdio: ['ignore', 'pipe', 'pipe'], detached: process.platform !== 'win32' })
41
+ }
42
+
43
+ // → { code, stdout, stderr, timedOut, notFound }
44
+ export function execClaude(args, { cwd = process.cwd(), env = process.env, timeoutMs = 20000 } = {}) {
45
+ return new Promise((resolve) => {
46
+ let stdout = ''
47
+ let stderr = ''
48
+ let timedOut = false
49
+ let settled = false
50
+ const child = start(args, { cwd, env })
51
+ const finish = (r) => {
52
+ if (settled) return
53
+ settled = true
54
+ clearTimeout(timer)
55
+ resolve(r)
56
+ }
57
+ const timer = setTimeout(() => {
58
+ timedOut = true
59
+ killTree(child)
60
+ }, timeoutMs)
61
+ child.stdout.on('data', (d) => (stdout += d))
62
+ child.stderr.on('data', (d) => (stderr = (stderr + d).slice(-65536)))
63
+ child.on('error', (e) => finish({ code: null, stdout, stderr: stderr || e.message, timedOut, notFound: e.code === 'ENOENT' }))
64
+ child.on('close', (code) => finish({ code, stdout, stderr, timedOut, notFound: false }))
65
+ })
66
+ }
67
+
68
+ // onEvent(event) returns a truthy value to stop: the process is killed and that value
69
+ // comes back as `stopped`. → { stopped, code, stderr, timedOut, notFound, ms }
70
+ export function streamClaude(args, { cwd, env = process.env, timeoutMs, onEvent }) {
71
+ return new Promise((resolve) => {
72
+ const t0 = Date.now()
73
+ let stderr = ''
74
+ let stopped = null
75
+ let timedOut = false
76
+ let settled = false
77
+ const child = start(args, { cwd, env })
78
+ const finish = (r) => {
79
+ if (settled) return
80
+ settled = true
81
+ clearTimeout(timer)
82
+ resolve({ ...r, ms: Date.now() - t0 })
83
+ }
84
+ const timer = setTimeout(() => {
85
+ if (stopped) return
86
+ timedOut = true
87
+ killTree(child)
88
+ }, timeoutMs)
89
+ const lines = lineSplitter((ev) => {
90
+ if (stopped || timedOut) return
91
+ const r = onEvent(ev)
92
+ if (r) {
93
+ stopped = r
94
+ killTree(child)
95
+ }
96
+ })
97
+ child.stdout.setEncoding('utf8')
98
+ child.stdout.on('data', (d) => lines.push(d))
99
+ child.stderr.on('data', (d) => (stderr = (stderr + d).slice(-65536)))
100
+ child.on('error', (e) => finish({ stopped, code: null, stderr: stderr || e.message, timedOut, notFound: e.code === 'ENOENT' }))
101
+ child.on('close', (code) => {
102
+ lines.end()
103
+ finish({ stopped, code, stderr, timedOut, notFound: false })
104
+ })
105
+ })
106
+ }
@@ -0,0 +1,102 @@
1
+ // Two reports side by side. The number is a property of the description AND of the
2
+ // environment it was measured in — the model, the CLI, and above all the roster: the
3
+ // stub competes with every other description the model can see, and one synced
4
+ // folder of 24 skills moved a byte-identical description from 18/18 to 10/18. So a
5
+ // difference in any of the three is printed as a warning before any delta is — the
6
+ // roster by its members as well as its counts, since two rosters of one size can differ
7
+ // (ST-20), and so is a difference in the models the runs reported or in what the
8
+ // inherited environment contributed.
9
+ //
10
+ // Two runs per query resolve to ±1 per query; a difference of one hit at two runs is
11
+ // labelled noise, and so is a total difference of up to two hits.
12
+
13
+ const rosterSize = (r) => (r?.roster ? r.roster.slashCommands : null)
14
+ const fmtRoster = (r) => (r?.roster ? `${r.roster.slashCommands} slash commands${r.roster.skills == null ? '' : `, ${r.roster.skills} skills`}` : 'unknown')
15
+
16
+ // Added and removed members, per kind, by hash — the report carries no names. Skipped
17
+ // when either report predates the hashes.
18
+ function memberDiff(a, b) {
19
+ const parts = []
20
+ for (const [key, label] of [['commandHashes', 'slash commands'], ['skillHashes', 'skills']]) {
21
+ const na = a.roster?.[key]
22
+ const nb = b.roster?.[key]
23
+ if (!Array.isArray(na) || !Array.isArray(nb)) continue
24
+ const sa = new Set(na)
25
+ const sb = new Set(nb)
26
+ const added = nb.filter((n) => !sa.has(n))
27
+ const removed = na.filter((n) => !sb.has(n))
28
+ if (added.length || removed.length) parts.push(`${label}: ${[added.length ? `${added.length} added (${added.join(', ')})` : '', removed.length ? `${removed.length} removed (${removed.join(', ')})` : ''].filter(Boolean).join('; ')}`)
29
+ }
30
+ return parts
31
+ }
32
+
33
+ const sameKeys = (x, y) => JSON.stringify(Object.keys(x).sort()) === JSON.stringify(Object.keys(y).sort())
34
+
35
+ function environmentDiff(ea, eb) {
36
+ if (!ea || !eb) return []
37
+ const out = []
38
+ const mem = (e) => `${e.memoryFiles?.project}+${e.memoryFiles?.user}`
39
+ if (mem(ea) !== mem(eb)) out.push(`memory files (project+user) ${mem(ea)} vs ${mem(eb)}`)
40
+ if (ea.hooks !== eb.hooks) out.push(`hooks ${ea.hooks} vs ${eb.hooks}`)
41
+ if (ea.mcpServers !== eb.mcpServers) out.push(`MCP servers ${ea.mcpServers} vs ${eb.mcpServers}`)
42
+ return out
43
+ }
44
+
45
+ export function compare(a, b) {
46
+ const warnings = []
47
+ if (a.model !== b.model) warnings.push(`model differs: ${a.model} vs ${b.model}`)
48
+ if (a.runModels && b.runModels && Object.keys(a.runModels).length && Object.keys(b.runModels).length && !sameKeys(a.runModels, b.runModels)) {
49
+ warnings.push(`the models the runs reported differ: ${Object.keys(a.runModels).sort().join(', ')} vs ${Object.keys(b.runModels).sort().join(', ')}`)
50
+ }
51
+ if (a.cliVersion !== b.cliVersion) warnings.push(`CLI version differs: ${a.cliVersion} vs ${b.cliVersion}`)
52
+ if (rosterSize(a) !== rosterSize(b) || a.roster?.skills !== b.roster?.skills) warnings.push(`roster differs: ${fmtRoster(a)} vs ${fmtRoster(b)} — the number is a property of the text and the roster; re-measure the baseline in the same environment before judging a rewrite`)
53
+ const members = memberDiff(a, b)
54
+ if (members.length) warnings.push(`roster members differ (a → b) — ${members.join(' · ')}`)
55
+ const env = environmentDiff(a.environment, b.environment)
56
+ if (env.length) warnings.push(`inherited environment differs: ${env.join(', ')} — memory files, hooks and MCP servers reach the model too`)
57
+ for (const [k, r] of [['a', a], ['b', b]]) if (r.verdict !== 'ok') warnings.push(`report ${k} (${r.date}) has no verdict — its counts are not a measurement`)
58
+ if (a.conflictAllowed || b.conflictAllowed) warnings.push('a report was measured with --allow-conflict — a same-named skill was visible')
59
+
60
+ const noiseAt = Math.min(a.runsPerQuery ?? 2, b.runsPerQuery ?? 2) <= 2
61
+ const byQuery = new Map(b.queries.map((q) => [q.query, q]))
62
+ const seen = new Set()
63
+ const rows = []
64
+ for (const qa of a.queries) {
65
+ const qb = byQuery.get(qa.query)
66
+ seen.add(qa.query)
67
+ if (!qb) {
68
+ rows.push({ query: qa.query, should_trigger: qa.should_trigger, a: qa, b: null, only: 'a' })
69
+ continue
70
+ }
71
+ const delta = qb.hits - qa.hits
72
+ rows.push({ query: qa.query, should_trigger: qa.should_trigger, a: qa, b: qb, delta, noise: delta !== 0 && Math.abs(delta) <= 1 && noiseAt })
73
+ }
74
+ for (const qb of b.queries) if (!seen.has(qb.query)) rows.push({ query: qb.query, should_trigger: qb.should_trigger, a: null, b: qb, only: 'b' })
75
+
76
+ // Totals over the queries both reports share, so a query added on one side does not
77
+ // pass for a change in the description.
78
+ const shared = rows.filter((r) => !r.only)
79
+ const total = (pred) => {
80
+ const pick = (side) => shared.filter((r) => pred(r.should_trigger)).reduce((acc, r) => ({ hits: acc.hits + r[side].hits, runs: acc.runs + r[side].runs }), { hits: 0, runs: 0 })
81
+ const ta = pick('a')
82
+ const tb = pick('b')
83
+ return { a: ta, b: tb, delta: tb.hits - ta.hits }
84
+ }
85
+ const totals = { positives: total((s) => s), negatives: total((s) => !s) }
86
+
87
+ const lines = [`skilltrigger compare — ${a.skill} ${a.date} (a) vs ${b.skill} ${b.date} (b)`, '']
88
+ for (const w of warnings) lines.push(`warning: ${w}`)
89
+ if (warnings.length) lines.push('')
90
+ if (a.description?.sha256 && b.description?.sha256) lines.push(a.description.sha256 === b.description.sha256 ? 'description: identical in both' : `description changed: ${a.description.bytes} → ${b.description.bytes} bytes`, '')
91
+ const sign = (n) => (n > 0 ? `+${n}` : String(n))
92
+ for (const r of rows) {
93
+ const tag = r.should_trigger ? 'pos' : 'neg'
94
+ if (r.only) lines.push(` ${tag} only in ${r.only}: ${r[r.only].hits}/${r[r.only].runs} ${r.query}`)
95
+ else lines.push(` ${tag} ${r.a.hits}/${r.a.runs} → ${r.b.hits}/${r.b.runs} ${r.delta === 0 ? ' 0' : sign(r.delta)}${r.noise ? ' (noise)' : ''} ${r.query}`)
96
+ }
97
+ lines.push('')
98
+ const totalNoise = (d) => (d !== 0 && Math.abs(d) <= 2 && noiseAt ? ' (within noise)' : '')
99
+ lines.push(`positives triggered ${totals.positives.a.hits}/${totals.positives.a.runs} → ${totals.positives.b.hits}/${totals.positives.b.runs} (${sign(totals.positives.delta)})${totalNoise(totals.positives.delta)}`)
100
+ lines.push(`negatives fired ${totals.negatives.a.hits}/${totals.negatives.a.runs} → ${totals.negatives.b.hits}/${totals.negatives.b.runs} (${sign(totals.negatives.delta)})${totalNoise(totals.negatives.delta)}`)
101
+ return { warnings, rows, totals, text: lines.join('\n') }
102
+ }