@rhize/skill-forge 0.14.0 → 0.16.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -2,7 +2,8 @@
2
2
 
3
3
  **The supply-chain gate for agent skills.**
4
4
 
5
- Status: pre-release (built 2026-07-11)
5
+ Status: published — `@rhize/skill-forge@0.14.0` on npm (`0.16.0` in this repo, pending release),
6
+ 0.x beta (Pro features free until 1.0)
6
7
 
7
8
  > This project is unrelated to the [`skillforge`](https://www.npmjs.com/package/skillforge)
8
9
  > package on npm, which is a Claude Skills *evaluation* framework. `skill-forge` (this package,
@@ -27,19 +28,15 @@ promote/hold/reject decision**, before it is allowed anywhere near your working
27
28
  npx @rhize/skill-forge init
28
29
  ```
29
30
 
30
- Optional, but recommended first: `init` detects which coding agents you have installed — the
31
- matrix covers 73 known agents (Claude Code, Codex CLI, Cursor, Windsurf, OpenCode, Gemini CLI,
32
- and 67 more; see `src/agents.ts`), and every entry's skill-directory paths are verified directly
33
- against the [vercel-labs/skills](https://github.com/vercel-labs/skills) CLI's own source, not
34
- just its README — none are unverified/community guesses today. (The schema carries a
35
- `verified: false` flag for any future entry that can't be confirmed that way; it's unused as of
36
- this release.) `init` lets you pick which of their skill directories should be gated, a default
37
- promotion target, and an optional agent to hand follow-up prompts off to (see
38
- [`--ingest`](#ingestion-handoff---ingest)). No configuration is
39
- required for a first run either way — skip `init` and skill-forge defaults to `<cwd>/.claude/skills`
40
- as its promotion target and `~/.skill-forge/quarantine` as its sandbox (and offers to run `init` for
41
- you the first time `add`/`scan`/`list`/`status` runs with no config present, in an interactive
42
- terminal). See [docs/configuration.md](docs/configuration.md) for the full field reference.
31
+ Optional, but recommended first: `init` detects which coding agents you have installed (73 known
32
+ agents) and lets you pick which of their skill directories to gate, a default promotion target,
33
+ and an optional agent to hand follow-up prompts off to. It also checks whether each picked root is
34
+ under Git version control yet and, if not, offers to give it a baseline commit (exact commands
35
+ shown first, default No). No configuration is required either way — skip `init` and skill-forge
36
+ defaults to `<cwd>/.claude/skills` as its promotion target (and offers to run `init` for you the
37
+ first time `add`/`scan`/`list`/`status` runs with no config present, in an interactive terminal).
38
+ See [docs/commands/init.md](docs/commands/init.md) for the full detection/menu/Git-preflight
39
+ behavior and [docs/configuration.md](docs/configuration.md) for the config field reference.
43
40
 
44
41
  ```bash
45
42
  npx @rhize/skill-forge add <owner>/<skill-name>
@@ -64,866 +61,31 @@ anything ending in `.git`), or a local filesystem path.
64
61
 
65
62
  ## Commands
66
63
 
67
- ```
68
- skill-forge init [options] Detect installed agents and set gate targets / handoff agent
69
- skill-forge add <source> [options] Quarantine-install a skill and run it through the gate
70
- skill-forge scan <source> [options] Gate a skill without installing it (always cleans up)
71
- skill-forge evolve <skill-dir> [options] Self-evolve an installed skill via SkillOpt-Sleep, re-gate, decide (Pro)
72
- skill-forge audit [options] Doctor-style health check over the configured skill/MCP set (alias: doctor)
73
- skill-forge finding accept <fp-prefix> Acknowledge a LOW/MEDIUM audit finding you've reviewed (human-only, no gate effect)
74
- skill-forge finding revoke <fp-prefix> Remove a stored acceptance so the finding reports again
75
- skill-forge finding list [options] List every accepted finding
76
- skill-forge organize [options] Set-level capability registry + dependency graph across configured skills roots (Pro)
77
- skill-forge find [query] [options] Discover skills via skills.sh and check partner security audits (free)
78
- skill-forge watch [options] Drift check across every SOURCES.md provenance ledger (Pro)
79
- skill-forge ingest [options] Hand the pending-ingestion queue off to a coding agent for the decide/absorb pass (Pro)
80
- skill-forge queue close <id> --status <s> Close a queue entry after the decide pass (Pro)
81
- skill-forge refine [options] Capture a project-scope override from real usage feedback (Pro)
82
- skill-forge refine list [options] Show refinement history
83
- skill-forge refine patterns [options] List tracked/ready/generalized/dismissed patterns
84
- skill-forge refine promote <PATTERN-ID> Merge a ready pattern into the user-scope skill (Pro)
85
- skill-forge refine rollback <backup-id> Restore a promotion backup (Pro)
86
- skill-forge refine which <skill> Print override-resolution order for a skill
87
- skill-forge routine [options] One scheduled maintenance pass: audit + drift + registry, cron-friendly (Pro)
88
- skill-forge promote <id> [options] Re-gate a held quarantine entry and promote it
89
- skill-forge reject <id> [options] Discard a held quarantine entry
90
- skill-forge config <sub> [args] Read/write open-ended preferences (list|get|set|unset|propose|review)
91
- skill-forge list List skills currently held in quarantine
92
- skill-forge status Show configuration and quarantine summary
93
- skill-forge guide [topic] Orientation: what this does, your current state, the next command
94
- ```
95
-
96
- ### `init`
97
-
98
- ```bash
99
- skill-forge init # interactive: pick targets, default target, handoff agent
100
- skill-forge init --defaults # non-interactive: agents found on PATH, first as default (CI)
101
- skill-forge init --list # print detected agent skill roots and exit — no writes
102
- skill-forge init --all-agents # don't narrow to agents whose CLI is on PATH — keep every root
103
- ```
104
-
105
- Probes the known agent matrix (`src/agents.ts`) for both project-relative (`.claude/skills`, ...)
106
- and global (`~/.codex/skills`, ...) skill directories that already exist on disk, then writes
107
- `skillsRoots`, `agents`, `defaultTarget`, and (if you pick a handoff agent) `handoffCommand` to
108
- `config.json`. Safe to re-run any time — it always starts from your existing config and only
109
- overwrites the fields it's responsible for. If no `config.json` exists yet, `add`/`scan`/`list`/
110
- `status` offer to run this for you on first use (skipped entirely for `--json`/`--yes`/non-TTY
111
- invocations, so scripted runs never block on a prompt).
112
-
113
- **Menus are keyed on the PATH, not the agent (v0.11).** Many agents share one skills directory —
114
- 18 entries in the agent matrix use `.agents/skills` as their project root — so the gate-target and
115
- default-target menus list each *distinct* root once, labelled with the agents that resolve to it
116
- (`19 agents (project): Amp, Replit, Universal, +16 more`), and `skillsRoots`/`mcpTargets` are
117
- written deduped. `config.agents` still records every agent, since it's an id→root map. Detection
118
- output says so explicitly when roots collapse (`→ 21 agents share 3 distinct skills root(s)`).
119
- Configs written by earlier versions are deduped on load, so no re-run is required to clean one up.
120
-
121
- **Narrowed to agents you actually have (v0.11).** Because a shared directory can't tell you which
122
- of its 18 agents you use, init also checks PATH: a root is pre-selected when at least one agent
123
- mapped to it has its CLI installed, and roots with none are still listed, just switched off with
124
- the reason shown.
125
-
126
- ```
127
- [x] 1. /Users/you/.agents/skills
128
- on PATH: Codex, Gemini CLI (project) (+17 other agents share this path)
129
- [ ] 3. /Users/you/.openclaw/skills
130
- OpenClaw (global) — no verified CLI name to check
131
- ```
132
-
133
- The PATH check is a **look-up only — nothing found is ever executed** (no `--version` probe), and
134
- binary names are held to the same evidence bar as the skill paths: an agent with no verified
135
- command name reports "unknown", never "not installed". `--all-agents` turns narrowing off, and it
136
- disables itself automatically if no agent CLI is found at all, so it can never reduce a working
137
- detection to an empty config. The handoff menu is only *reordered* by it — installed CLIs float to
138
- the top and everything detected stays pickable.
139
-
140
- **Selecting things (v0.11).** The multi-selects take a whole answer at once — `2`, `1 3`, `2, 4`,
141
- `2-5`, plus `all` and `none` — and Enter confirms. Anything unrecognized is named back to you and
142
- ignored, never silently swallowed. The single-choice prompts (default target, handoff agent)
143
- **re-ask** on an answer that isn't one listed number instead of falling back to option 1, so a
144
- multi-value or mistyped answer can no longer write a target you didn't choose. Ctrl+D at any
145
- prompt exits cleanly.
146
-
147
- **Init now ends by offering the audit (v0.8).** After an interactive run writes its config, it
148
- asks "Run the skills & MCP audit now? [Y/n]" (default yes) and, if accepted, runs
149
- `skill-forge audit` interactively — see [`audit`](#audit-v08) below. `init --defaults` runs it too,
150
- but non-interactively (`--yes`): a report is written, with no business-profile prompt, no
151
- foundation scaffold, and no agent handoff. `init --list` and an aborted/empty selection stay
152
- write-free, so neither writes a config nor runs the audit.
153
-
154
- ### `add`
155
-
156
- ```bash
157
- skill-forge add owner/name
158
- skill-forge add https://github.com/owner/repo.git --target ./.claude/skills
159
- skill-forge add ./local-skill-dir --yes
160
- skill-forge add owner/name --json
161
- skill-forge add owner/name --yes --ingest
162
- ```
163
-
164
- Installs the source into a quarantine sandbox, then runs it through the gate: profile → safety
165
- scan → overlap analysis against the configured skills root (Pro) → report. Nothing touches the
166
- target skills root until you decide.
167
-
168
- | Option | Effect |
169
- |---|---|
170
- | *(none)* | Prompts you to **promote**, **hold**, or **reject** the candidate. |
171
- | `-y, --yes` | Skips the prompt and honors the gate verdict: a `block` safety verdict is rejected (process exits nonzero); anything else (`pass`/`warn`) is promoted. |
172
- | `-t, --target <dir>` | Skills root to promote into. Defaults to the config's `defaultTarget`, then `skillsRoots[0]`. |
173
- | `--json` | Prints the gate result (profile, safety findings, overlap) as JSON instead of the terminal report box. Implies non-interactive: the decision is made the same way `--yes` makes it (verdict decides promote/hold/reject), never an interactive prompt. |
174
- | `--ingest` | Pro (free during the 0.x beta). After a successful promote, hands off to a coding agent — see [Ingestion handoff](#ingestion-handoff---ingest). |
175
- | `--skill-map <path>` | Path to a generated `rhize-plugins` skill map (Phase 4 of that repo's skill-map-graph-substrate plan). When given, ranks the candidate's name/description against every `skill` node in the map — a near-duplicate of an already-shipped marketplace skill is folded into the safety findings as a `HIGH`-severity finding (escalates the verdict to `block`, so it's held rather than silently promoted); a moderate overlap is `MEDIUM` (`warn`). Free (not Pro-gated) — this enforces the marketplace's own curation rule ("close the gap, don't duplicate"), not a premium ranking feature. Missing/unreadable map: printed to stderr, never fatal. **Two map variants exist, with different coverage**: the **static** map (`rhize-plugins/generated/skill-map.static.json`) is first-party-only — it only knows about that marketplace's own skills, so it cannot catch a candidate that duplicates an *installed third-party plugin's* skill. The **resolved** map (`~/.claude/context-manager/skill-map.resolved.json`, produced by `rhize-context-manager`) additionally carries the full third-party ecosystem inventory. Point `--skill-map` at the resolved map when you want ecosystem-wide duplicate detection; the static map is enough only when you're curating a single first-party marketplace against itself. |
176
-
177
- **Observability (v0.14.0):** the `--skill-map` curation check now always reports that it ran. A below-threshold PASS previously left no trace at all — no line in the report, no field in `--json` — making "checked 45 skills, nothing crossed the threshold" indistinguishable from "the check never ran". That ambiguity caused a correct PASS to be written up as a broken gate. The report now carries a `Marketplace overlap` line (`checked N skills — closest "x" (score), none over threshold` / `skipped — <reason>`) and `--json` a `mapOverlap` object with `status`, `compared`, `topMatch`, and `matches`. `topMatch` is informational only: it is never graded and never produces a finding. **Scoring and both thresholds are unchanged.** Note the check is lexical, not semantic — two skills can claim the same invocation trigger in different words and score far below threshold (measured: 0.134), so competing `description:` triggers still need a human read. See `gate/mapOverlap.ts`'s module doc.
178
-
179
- **Bugfix (v0.13.0):** `--skill-map` previously matched nothing against either real map variant — the loader read a `type` field that the real generated artifact never sets (it uses `kind`), so every node was silently invisible to the overlap check. `loadSkillMap()` now normalizes `kind` → `type` at load time; regression coverage lives in `test/mapOverlap.realmap.test.ts` against vendored real-map snapshots. If you were relying on `--skill-map` before v0.13.0, it was a no-op — re-run `add`/`scan` on anything you're unsure about.
180
-
181
- **Extends-declared overlap exemption**: a candidate can declare `metadata.rhize.extends:
182
- ["<skill-name>" | "<plugin>/<skill-name>"]` in its SKILL.md frontmatter to mark itself a
183
- deliberate specialization/layering of an existing map skill. When the top `--skill-map` match is
184
- a skill the candidate declares it extends, that finding is downgraded from a HOLD-level safety
185
- finding to an informational notice printed above the report — it never escalates the verdict.
186
- Overlap with any *other* skill (a different match, or a second candidate/skill pair) still holds
187
- at full severity: the exemption applies per declared pair, not globally, so declaring an
188
- extension of skill X never waives overlap with skill Y.
189
-
190
- A promoted skill gets a provenance entry appended to `<target>/SOURCES.md`, and every promote or
191
- hold decision is recorded to `~/.skill-forge/queue.json` (or `$SKILL_FORGE_HOME/queue.json`) — see
192
- [`docs/queue-schema.md`](docs/queue-schema.md) for the entry schema. A reject writes neither —
193
- nothing is left behind to record. These are Pro features that run free during the 0.x beta (see
194
- [docs/pro.md](docs/pro.md#beta-pricing-0x)): with no valid license, `add` still writes them, and
195
- prints a one-line notice above the report instead of skipping them.
196
-
197
- `add`/`scan` also accept `--artifact mcp` (plus `add`-only `--mcp-target <file>`/`--force`) to gate
198
- an MCP server instead of a skill — see [MCP gating](#mcp-gating-v05) below.
199
-
200
- ### `scan`
201
-
202
- ```bash
203
- skill-forge scan owner/name
204
- skill-forge scan owner/name --json
205
- ```
206
-
207
- Runs the same gate pipeline as `add` (profile → safety → overlap → report) but never promotes
208
- anything — the quarantine sandbox is always cleaned up afterward, on success or failure. Exits
209
- nonzero when the safety verdict is `block`. `--json` prints the same gate-result payload shape as
210
- `add`'s. Also accepts `--skill-map <path>` — see `add`'s option table above.
211
-
212
- ### `evolve` (v0.7)
213
-
214
- ```bash
215
- skill-forge evolve .claude/skills/my-skill
216
- skill-forge evolve .claude/skills/my-skill --yes
217
- skill-forge evolve .claude/skills/my-skill --dry-run
218
- skill-forge evolve .claude/skills/my-skill --backend claude --yes
219
- ```
220
-
221
- Pro (free during the 0.x beta). Orchestrates [microsoft/SkillOpt](https://github.com/microsoft/SkillOpt)'s
222
- `skillopt-sleep` CLI (`pip install skillopt`) to propose a self-evolution of an already-installed
223
- skill — harvest recent sessions, generate a candidate replacement `SKILL.md`/`CLAUDE.md`, and stage
224
- it — then runs the **staged proposal**, never the live skill, back through skill-forge's own static
225
- safety ruleset before you decide anything. SkillOpt-Sleep's own validation gate is score-only; it
226
- never content-vets the generated markdown. This closes that gap.
227
-
228
- | Option | Effect |
229
- |---|---|
230
- | `--project <dir>` | Project dir passed to SkillOpt-Sleep as `--project`. Defaults to cwd. |
231
- | `--dry-run` | Uses SkillOpt-Sleep's `dry-run` subcommand — report only, nothing staged. |
232
- | `--backend <name>` | SkillOpt-Sleep backend. Defaults to `mock` (offline, deterministic, no network). |
233
- | `--lookback-hours <n>` | Hours of session history for SkillOpt-Sleep to harvest. |
234
- | `-y, --yes` | Skips interactive prompts (the disclosure confirm below, and the promote/hold/reject prompt) and honors the re-gate verdict automatically. |
235
- | `--json` | Prints the re-gate result as JSON instead of the terminal report. Implies non-interactive, same as `add`'s `--json`. |
236
- | `--force` | Allows re-adopting a staging dir whose proposed content already matches the live skill (see the double-adopt guard below). |
237
-
238
- **Requires `skillopt-sleep` on PATH** — skill-forge never installs it for you. If it's missing,
239
- `evolve` prints `pip install skillopt` + a docs pointer and exits, rather than attempting an
240
- auto-install (the same "detect, don't install" discipline the rest of the gate follows).
241
-
242
- **Data boundary.** The default `mock` backend is fully offline — harvesting and proposal
243
- generation both run locally, no session data leaves the machine. Any other backend (`claude`,
244
- `codex`, `azure_openai`, ...) sends truncated excerpts from harvested sessions and derived tasks to
245
- the provider you selected; per SkillOpt-Sleep's own docs this is not currently guaranteed to be
246
- secret-free. `evolve` prints that disclosure and requires either `--yes` or an interactive `y/N`
247
- confirmation before a non-`mock` run proceeds — `--json` is non-interactive, so a non-`mock` run
248
- under `--json` without `--yes` is refused rather than silently sending data off-machine.
249
-
250
- **Re-gate.** SkillOpt-Sleep only ever *stages* a proposal (`<project>/.skillopt-sleep/staging/<timestamp>/`
251
- — full replacement files, never auto-adopted). `evolve` copies just `proposed_SKILL.md` (and
252
- `proposed_CLAUDE.md`, if present) into a fresh temporary directory and runs skill-forge's own
253
- static safety ruleset over that copy — the same one `add`/`scan` use, at your configured
254
- `strictness`. `report.json`/`diagnostics.json` (SkillOpt-Sleep's own redacted holdout evidence) are
255
- never scanned. The result renders through the same terminal report / `--json` shape as `add`/`scan`.
256
-
257
- **Decision.** Same promote/hold/reject semantics as `add`: `--yes`/`--json` honor the re-gate
258
- verdict (a `block` is rejected); otherwise you're prompted.
259
-
260
- - **promote** — runs SkillOpt-Sleep's own `adopt --staging <dir>` (which backs up the live file(s)
261
- before copying the proposal over them), then records a provenance entry to the target skill's
262
- `SOURCES.md` and a pending-ingestion queue entry (`origin: "evolve"`, see
263
- [docs/queue-schema.md](docs/queue-schema.md)) so a later ingest pass reviews the evolution rather
264
- than an external source.
265
- - **hold** — leaves the staging dir exactly as SkillOpt-Sleep produced it (its own `status` command
266
- still lists it); `report.md` has the full evidence.
267
- - **reject** — deletes the staging dir.
268
-
269
- **Double-adopt guard.** SkillOpt-Sleep's `adopt` has no confirmation of its own, and adopting the
270
- same staging dir twice overwrites *its* backup — the pre-evolution original would be lost. Before
271
- adopting, `evolve` refuses (unless `--force`) when the staged proposal is already byte-identical to
272
- the live skill it would replace, since that's the signature of a staging dir that was already
273
- adopted once.
274
-
275
- ### `audit` (v0.8)
276
-
277
- ```bash
278
- skill-forge audit
279
- skill-forge doctor # alias
280
- skill-forge audit --json --report ./audit.md
281
- skill-forge audit --yes
282
- skill-forge audit --foundation
283
- skill-forge audit --handoff
284
- ```
285
-
286
- A re-runnable, doctor-style health check over the skill/MCP set you've **already** configured —
287
- unlike `add`/`scan`, which gate a new candidate before it's installed, `audit` inventories what's
288
- already there and looks for hygiene issues and consolidation/refinement opportunities. Requires a
289
- real, saved config (`skill-forge init` first) — it never silently audits an invented default.
290
- `init` now ends by offering to run it (see [`init`](#init) above); it's equally safe to run any
291
- time on its own.
292
-
293
- | Option | Effect |
294
- |---|---|
295
- | *(none)* | Interactive: offers to capture/reuse a business profile, then runs the audit, then offers the foundation scaffold and agent handoff. |
296
- | `--json` | Prints the full audit report as JSON instead of the terminal summary. Non-interactive — skips the business-profile prompt. |
297
- | `-y, --yes` | Skips every interactive prompt (business profile, foundation, handoff). Never implies `--foundation` or `--handoff`. |
298
- | `--report <file>` | Write the report here instead of the default `~/.skill-forge/reports/audit-<ISO-timestamp>.md`. Refuses an existing path — never overwrites. |
299
- | `--foundation` | The one skills-root write this command can make: scaffold a `business-foundation` skill from the captured business profile (see below). |
300
- | `--handoff` | Hand the written report off to your configured coding agent with the bundled curation prompt, via the same handoff plumbing as `add --ingest`. |
301
-
302
- **What it inventories/checks.** Every configured `skillsRoots` entry — symlink-aware, deduped by
303
- realpath so a skill reachable via two roots or an aliased symlink is reported once, with alias
304
- locations kept rather than dropped. Per skill: strict frontmatter validation (a missing/unclosed
305
- fence or missing `name`/`description` is a finding, not silently backfilled), `SKILL.md` size and
306
- an estimated token count, and a full `scanSafety` pass — the same safety ruleset `add`/`scan` run.
307
- Every configured `mcpTargets` file: JSON targets get full server enumeration (name, command
308
- basename, package spec, arg/env **counts** — never values); TOML targets (e.g. Codex CLI's
309
- `config.toml`) get the same textual MCP safety scan `--artifact mcp` uses, since there's no
310
- structured TOML enumeration. A per-item failure (unreadable skill, broken symlink, malformed MCP
311
- config) becomes a finding — it never aborts the run.
312
-
313
- **Accepted findings (v0.12).** Every finding in the report carries a fingerprint
314
- (`(fp a1b2c3d4e5f6)`); `Findings` shows **active** findings only, and a separate
315
- **Accepted findings** section lists whatever you've reviewed and accepted via
316
- [`skill-forge finding accept`](#finding-v012) — permanently, with reason and date, even after it
317
- stops matching (see that section for the full contract). `summary.acceptedCount` and each
318
- finding's `fingerprint` are additive `--json` fields; nothing accepted is ever silently deleted.
319
-
320
- **Opportunity pass.** Beyond hygiene, the report surfaces:
321
-
322
- - **Overlap clusters** (Pro, free during the 0.x beta) — cross-root overlap scoring across every
323
- inventoried skill, grouped into connected components, each with a top pairwise score and a
324
- suggested verb code (`ABSORB`/`FORK`/`DEFER` — `opportunities.overlapLocked` in `--json` output
325
- says whether this ran or was Pro-locked, without string-matching `notices`). The fuller
326
- five-verb matrix (adding `REJECT`/`WATCH`) belongs to the agent-side curation prompt's deeper
327
- decide pass (`--handoff`), not this report. Locked → that section shows the standard upgrade
328
- notice and empty clusters; everything else in the report still runs.
329
- - **Evolve-eligible skills** — structurally valid, safety-passing skills, labeled as eligible for
330
- `skill-forge evolve` — an eligibility list, not a judgment that they need refining.
331
- - **A business-foundation opportunity** — offered whenever `config.foundationSkillPath` isn't set
332
- yet.
333
-
334
- **Business profile.** Interactive runs (never under `--yes`/`--json`/non-TTY) open with a
335
- business-name/industry/audiences/workflows/constraints Q&A, preceded by an explicit "don't enter
336
- credentials or confidential customer data" warning. An existing stored profile is offered for
337
- reuse rather than re-asked, and a fresh capture is shown back as a summary before it's saved to
338
- `config.businessProfile`. This profile grounds both the foundation scaffold and the curation
339
- handoff prompt.
340
-
341
- **Report location + privacy.** Written to `~/.skill-forge/reports/audit-<ISO-timestamp>.md` by
342
- default (`--report <file>` to override), created with `{ flag: 'wx', mode: 0o600 }` — exclusive
343
- create (refuses to overwrite an existing report) and owner-only permissions. MCP `env` values and
344
- arbitrary arg values are never written into the report or the `--json` payload — only server/
345
- command names, package specs, and counts.
346
-
347
- **Explicit opt-ins.** `--foundation` is the *only* way this command writes to a skills root: it
348
- scaffolds `<targetRoot>/business-foundation/SKILL.md` from the captured business profile
349
- (containment-checked against `config.skillsRoots`, refuses an existing or symlinked destination,
350
- writes atomically, and runs `scanSafety` on the generated content before it's placed), then
351
- records the path to `config.foundationSkillPath`. `--handoff` is the *only* way this command
352
- launches an agent — same argv-array, no-shell discipline as `--ingest` (see
353
- [Ingestion handoff](#ingestion-handoff---ingest)), using the bundled `assets/curation-prompt.md`
354
- in place of `ingest-prompt.md`. `--yes` (and `init --defaults`, which passes it through) never
355
- implies either — a non-interactive run writes only the report and, if a profile was already
356
- stored, the config; nothing else.
357
-
358
- ### `finding` (v0.12)
359
-
360
- ```bash
361
- skill-forge finding accept <fingerprint-prefix> --reason "Funnel genuinely needs root; verified"
362
- skill-forge finding revoke <fingerprint-prefix>
363
- skill-forge finding list [--stale] [--json]
364
- ```
365
-
366
- Acknowledge an `audit` finding you've reviewed and accepted, so it stops counting against
367
- `routine --fail-on` and the `add`/`scan` advisory while staying **permanently visible** in its own
368
- report section — an audit you can't triage decays into noise, and noise is how the one real
369
- finding gets skimmed. Every finding in an `audit`/`routine` report now carries a 12-char
370
- fingerprint (`(fp a1b2c3d4e5f6)`) you copy into `accept`/`revoke`.
371
-
372
- | Command | Effect |
373
- |---|---|
374
- | `finding accept <fp-prefix> --reason <text>` | Records an acceptance. Re-runs the read-only audit engine to resolve the prefix against what's **currently active** — you can only accept a finding that exists right now, never from a stale report. `--reason` is mandatory (1-200 chars, no control characters). |
375
- | `finding revoke <fp-prefix>` | Removes a stored acceptance so the finding reports again. Resolves against the **store**, not a fresh audit run — the target may no longer be producible, which is exactly when you'd want to clean up. |
376
- | `finding list [--stale] [--json]` | Lists every accepted finding (severity, rule, target, reason, accepted date, fingerprint). `--stale` filters to records whose target no longer exists on disk. |
377
-
378
- **Identity is content-derived, not a name you pick.** A finding's fingerprint is
379
- `sha256(severity + rule + target + the offending line's text)` — deliberately excluding the LINE
380
- NUMBER (drifts on unrelated edits) and deliberately including SEVERITY (so a rule re-tuned to a
381
- higher severity for the same content can't inherit an ack minted against the lower one). Any change
382
- to the offending content **fails open to re-reporting** — this is the design, not a bug: the
383
- alternative (key on rule+path only) is the classic baseline trap, where accepting one benign line
384
- silently suppresses every future line the same rule matches in that file. Full contract, the
385
- per-rule `context` table, and fail-open read semantics: see
386
- [docs/accepted-findings-schema.md](docs/accepted-findings-schema.md).
387
-
388
- **Severity-capped, human-only, no override.** `finding accept` refuses `HIGH`/`CRITICAL` findings
389
- outright — there is no `--force`. The precision problem this command exists to solve lives at
390
- `LOW`/`MEDIUM`; a `HIGH` false positive is a rule bug to fix, not something to baseline. This is a
391
- human-door CLI (`config set` model), not the agent-facing propose/review queue — both
392
- `assets/ingest-prompt.md` and `assets/curation-prompt.md` instruct agents to never run
393
- `finding accept`/`finding revoke` themselves; an agent that judges a finding a false positive says
394
- so in its summary and lets the user run the command.
395
-
396
- **Never a gate input.** Acceptances only ever affect `audit`'s own report classification (and,
397
- downstream, `routine --fail-on` and the `add`/`scan` advisory) — `add`/`scan`/`promote <id>`/
398
- `evolve`'s re-gate never reads the accepted-findings store. See
399
- [docs/gate-policy.md](docs/gate-policy.md#accepted-findings-v012).
400
-
401
- ### `organize` (v0.9, Pro)
402
-
403
- ```bash
404
- skill-forge organize
405
- skill-forge organize --json
406
- skill-forge organize --out ./registry.json
407
- skill-forge organize --usage-snapshot ./skill-monitor-snapshot.json
408
- ```
409
-
410
- Pro (free during the 0.x beta). Builds a set-level view across every configured `skillsRoots`
411
- entry: a **capability registry** (per-skill row — `tier`/`domain`/`consumes`/`provenance`/
412
- `maturity` frontmatter, plus untagged/rot flags) and a **dependency graph** (nodes, `consumes`
413
- edges, orphans, hot resources, dangling edges, cycles). This is the TS port of the plugin's
414
- `index_skills.py` + `build_dependency_graph.py`, merged into one command — graph nodes are keyed
415
- by realpath-qualified stable IDs (not bare frontmatter names), so duplicate names across roots
416
- don't collapse adjacency; duplicate-name ambiguity and unresolved `consumes` targets are reported
417
- rather than silently dropped.
418
-
419
- | Option | Effect |
420
- |---|---|
421
- | `--json` | Print the registry + graph as JSON instead of the terminal summary. |
422
- | `--out <file>` | Write the full report JSON to this path. Refuses an existing path (`wx`, `0o600` — same discipline as `audit --report`). |
423
- | `--usage-snapshot <file>` | Join per-skill usage counts from a skill-monitor snapshot JSON (`usage_joined` in the output); a zero-match snapshot is warned about, not silently ignored. |
424
-
425
- `organize` performs its own self-contained scan — it does not reuse `audit`'s inventory walker
426
- — and, like every other command, never executes anything found under a scanned skills root.
427
-
428
- ### `find` (v0.9, free)
429
-
430
- ```bash
431
- skill-forge find "pdf form filling"
432
- skill-forge find "pdf form filling" --limit 20 --json
433
- skill-forge find --audit owner/skill-name
434
- skill-forge find --get owner/skill-name
435
- skill-forge find --curated
436
- ```
437
-
438
- Free. TS port of the plugin's `skills_sh.py` — discovery only, exactly one mode per invocation
439
- (search, `--audit`, `--get`, or `--curated`; mutually exclusive):
440
-
441
- | Mode | Effect |
442
- |---|---|
443
- | `<query>` (default) | Search skills.sh; each result shows id, install count, and the install hint `skill-forge add <id>` — **not** `npx skills add`, since `add` runs the full local quarantine/gate. |
444
- | `--audit <id>` | Partner security-audit verdicts (Socket, Snyk, Gen Agent Trust Hub, …) for a specific skills.sh id — pass/warn/fail + risk. No auto-audit of search results. |
445
- | `--get <id>` | Detail + file tree for a specific skills.sh id. |
446
- | `--curated` | List skills.sh's curated skills. |
447
- | `--limit <n>` | Max search results (default 10). |
448
- | `--json` | Print the raw API payload as JSON. |
449
-
450
- Search results are labeled **unvetted** — `find` never installs or gates anything itself. Think
451
- of `find --audit` as the partner-verdict layer and `add`/`scan` as the deep local gate; they're
452
- two independent checks, not a replacement for one another.
453
-
454
- **Auth: `VERCEL_OIDC_TOKEN`.** All requests are HTTPS GETs to `https://skills.sh/api/v1`, using a
455
- bearer token read from `process.env.VERCEL_OIDC_TOKEN` — never stored in config, never printed.
456
- Without it, `find` fails loud (exit code 3) with setup guidance:
457
-
458
- ```
459
- 1. skills.sh authenticates with a short-lived Vercel OIDC token
460
- 2. npm i -g vercel && vercel link && vercel env pull (writes VERCEL_OIDC_TOKEN to .env.local)
461
- 3. export VERCEL_OIDC_TOKEN from .env.local into your shell before running this command
462
- (or set it directly: export VERCEL_OIDC_TOKEN=...) — docs: https://skills.sh/docs/api
463
- ```
464
-
465
- Note step 3: `vercel env pull` writes the token into `.env.local`, it does not export it into your
466
- shell — you (or your shell's dotenv loader) still need to export it before `find` can see it.
467
-
468
- ### `watch` (v0.9, Pro)
469
-
470
- ```bash
471
- skill-forge watch
472
- skill-forge watch --offline
473
- skill-forge watch --json
474
- skill-forge watch --skill-map generated/skill-map.static.json
475
- ```
476
-
477
- Pro (free during the 0.x beta). TS port of the plugin's `record_provenance.py --check-drift`.
478
- Scans **every** configured `skillsRoots` entry's `SOURCES.md` provenance ledger (deduped by
479
- realpath, each reported row names its ledger) and reports, per tracked entry, whether the
480
- recorded upstream ref still matches what the source currently resolves to.
481
-
482
- For entries whose `Source` parses as a git URL and whose `Upstream ref` looks like a commit/tag,
483
- `watch` runs a fixed, first-party `git ls-remote <url> [ref]` (argv array, no shell) to compare —
484
- `--offline` skips all network checks and just lists entries for manual comparison. Every other
485
- entry is listed for manual checking regardless.
486
-
487
- **`watch` NEVER executes a ledger's stored drift-check command string.** That string is
488
- attacker-influenceable data — anyone who can write to a skills root's `SOURCES.md` controls it.
489
- It is only ever printed, sanitized, as a suggestion for you (or an agent) to run yourself.
490
-
491
- **`--skill-map <path>` (Phase 4 of `rhize-plugins`' skill-map-graph-substrate plan)** additionally
492
- drift-checks every `fork-of` edge in a generated skill map: for each edge it resolves the local
493
- `skill` node's `path`/`contentHash` and the upstream (`external`) node's `url`/`path`, fetches or
494
- reads the upstream content, and compares content hashes. Same never-execute posture as the ledger
495
- check above: nothing derived from the map — including a fork-of edge's own `driftCheck`
496
- metadata — is ever executed; only fetch/read/hash/compare. A missing/unreadable map is a warning,
497
- never fatal. A node's repo-relative `path` (e.g. `rhize-context-manager/skills/x/SKILL.md`) is
498
- resolved against cwd, the map's own directory, and that directory's parent (in that order) — not
499
- cwd alone — so `watch --skill-map <path>` gives correct verdicts regardless of the directory you
500
- run it from.
501
-
502
- **Verdict**, per fork-of edge — three-way comparison, four verdict states when both hashes below
503
- are present on the map, else the older two-way fallback (`in-sync`/`drifted`, `contentHash` vs
504
- freshly-fetched upstream):
505
-
506
- | local-normalized vs baseline | upstream-now vs baseline | status | actionable |
507
- |---|---|---|---|
508
- | == | == | `in-sync` | no |
509
- | != | == | `local-only` | no (deliberate fork divergence, e.g. Rhize's added frontmatter) |
510
- | == | != | `upstream-moved` | yes |
511
- | != | != | `diverged` | yes |
512
-
513
- The three-way matrix fires only when the local `skill` node carries `contentHashNormalized` (a
514
- hash of the file with Rhize-injected frontmatter stripped, computed once by the rhize-plugins
515
- compiler) **and** the upstream `external` node carries `baselineHash` (the upstream content hash
516
- as of the last human review, recorded in `rhize-plugins`' SOURCES.md and copied onto the node by
517
- that compiler — skill-forge never computes or fetches it itself). Either field missing on a given
518
- edge falls back to the two-way compare for that edge, so older maps keep working. All four
519
- three-way verdicts — `in-sync`, `local-only`, `upstream-moved`, and `diverged` — carry
520
- `baselineHash`/`upstreamHash` in the JSON output; the two-way-fallback rows and
521
- `upstream-unreachable`/`local-missing` never do (no verdict without a successful fetch/read, so
522
- there is nothing to compare against a baseline). Every row also carries an `actionable` boolean
523
- (from `isActionable`), so callers don't have to re-derive "needs attention" from `status`/`detail`
524
- prose — it's `true` for `drifted`, `upstream-moved`, `diverged`, `upstream-unreachable`, and
525
- `local-missing`; `false` for `in-sync` and `local-only`. **Re-baselining**: after reviewing and
526
- adopting an `upstream-moved`/`diverged` change, re-run `rhize-plugins`' `scripts/baseline_upstreams.py`
527
- and commit the updated SOURCES.md — that's the "I reviewed upstream, accept its state" action that
528
- clears the row back to `in-sync`/`local-only`.
529
-
530
- ### `ingest` + `queue close` (v0.9, Pro)
531
-
532
- ```bash
533
- skill-forge ingest
534
- skill-forge ingest --list
535
- skill-forge ingest --list --json
536
- skill-forge queue close <id> --status ingested
537
- skill-forge queue close <id> --status dismissed
538
- ```
539
-
540
- Pro (free during the 0.x beta). The queue-drain UX for pending ingestions. `ingest` (no args)
541
- validates every pending `~/.skill-forge/queue.json` entry — each entry's `quarantinePath`/`installedPath` must
542
- canonicalize under the configured quarantine dir or a configured skills root/MCP target;
543
- mismatched or escaping entries are reported and excluded — then hands the surviving summary off to
544
- the configured coding agent (same handoff plumbing as `add --ingest`) with `assets/ingest-prompt.md`,
545
- which now carries the full queue-drain workflow. `--list` is read-only: it prints pending entries
546
- (`--json` for machine-readable output) without any handoff.
547
-
548
- For a single new source, use `skill-forge add <source> --ingest` instead — that runs the full
549
- quarantine/gate pipeline first; `ingest` only ever drains what's already queued.
550
-
551
- `skill-forge queue close <id> --status ingested|dismissed` lets an agent close out a queue entry
552
- after its decide pass without hand-editing `queue.json` — writes are atomic (temp file + rename).
553
-
554
- ### `refine` (v0.10, Pro)
555
-
556
- ```bash
557
- skill-forge refine --skill my-skill --category hook --override-type patch \
558
- --action insert-after --marker "Only check paths" \
559
- --content "..." --expected "..." --actual "..." --dry-run
560
- skill-forge refine list [--status <s>] [--skill <s>] [--project <p>]
561
- skill-forge refine patterns [--status tracking|ready|generalized|dismissed] [--skill <s>]
562
- skill-forge refine promote <PATTERN-ID> [--dry-run] [--force] [--user-root <dir>]
563
- skill-forge refine rollback <backup-id> [--force]
564
- skill-forge refine which <skill> [--user-root <dir>]
565
- ```
566
-
567
- Pro (free during the 0.x beta). Captures, applies, and generalizes improvements to installed
568
- skills from real usage feedback (see [CLAUDE.md's "refine (v0.10)" section](CLAUDE.md) for the
569
- full architecture notes, or `docs/refinement-schema.md` for the store shape).
570
-
571
- **Capture never mutates a base `SKILL.md`.** `refine` capture/apply writes ONLY project-scope
572
- override artifacts — `SKILL.patch.md` / `SKILL.extend.md` / `skill-config.json`, or a whole-file
573
- override copy for `full`/`hook`/`script` override types. The only command that ever touches a
574
- user-scope base skill is `refine promote`, and only against a `ready` pattern (or with `--force`).
575
-
576
- **Override precedence** (configurable roots):
577
-
578
- ```
579
- 1. PROJECT LOCAL <cwd>/.claude/skills/<skill>/ highest
580
- 2. PROJECT SHARED <cwd>/skills/<skill>/
581
- 3. USER SCOPE first configured skillsRoots entry outside cwd (fallback ~/.claude/skills)
582
- ```
583
-
584
- `refine which <skill>` prints this resolution order and which override files exist at each scope —
585
- the "why didn't my patch take effect" debugging aid, made explicit and read-only. Both `refine
586
- which` and `refine promote` accept `--user-root <dir>` to override the detected user-scope skills
587
- root (default: the first configured `skillsRoots` entry outside the current working directory,
588
- falling back to `~/.claude/skills`) — useful when the config's default root doesn't match the
589
- skill you're targeting.
590
-
591
- **Non-interactive capture contract.** Capture flags: `--skill --category --target --override-type
592
- <patch|extend|config|full|hook|script|new> --action <append|prepend|replace-section|insert-after|
593
- insert-before|delete-section> --marker --content <text> --content-file <file> --expected --actual
594
- --example --outcome --root-cause --pattern-id --scope <local|shared> --dry-run --json --yes
595
- --handoff`. `--content` takes override content inline; `--content-file <file>` reads it from a
596
- file instead (mutually exclusive with `--content`) — use it for anything multi-line rather than
597
- fighting shell quoting. All seven override types
598
- are supported: `patch`/`extend`/`config` are rendered from the flags; `full`/`hook`/`script` are
599
- verbatim override-file writes at project scope (content required); `new` creates an extension file
600
- for a capability that doesn't exist in the base skill yet. Always run `--dry-run` first and confirm
601
- the preview before writing for real — `--yes` skips re-prompting for a confirmation already given,
602
- it does not replace the dry-run preview.
603
-
604
- **The judgment step is the agent's job**, same pattern as `--ingest`/`audit --handoff`: a
605
- bundled `assets/refine-prompt.md` carries the gap-analysis rubric (category/override-type decision
606
- tables, guided-mode triggers), the patch-action syntax, the pattern-fingerprint/generalization
607
- criteria, and the verification step. `refine --handoff`
608
- (opt-in) launches your configured agent with it, same handoff plumbing as `--ingest`; `--yes` never
609
- implies `--handoff`.
610
-
611
- **Pattern tracking and promotion.** A pattern becomes `ready` only when it recurs in a **second,
612
- genuinely different project** — repeat captures in the same project never flip it (occurrence
613
- `count` is derived from unique project identities, not raw refinement counts; see
614
- `docs/refinement-schema.md`). `refine patterns` lists tracked/ready/generalized/dismissed patterns,
615
- filterable by `--status --skill --project`. `refine promote <PATTERN-ID>` merges a `ready` pattern
616
- into the user-scope base skill: it backs up affected files first (manifest with per-file sha256 +
617
- original bytes + mode, tombstones for files that didn't exist), stages the write, validates, then
618
- renames into place — `--dry-run` previews the diff without writing, `refine rollback <backup-id>`
619
- restores from the manifest (refusing if current files have drifted since promote, unless
620
- `--force`). On promote, a `SOURCES.md` provenance entry is appended (verb `DEFER`, notes
621
- `"generalized from PAT-xxxx via skill-forge refine"`).
622
-
623
- **`evolve` vs. `refine`.** Both improve an already-installed skill, but at different scopes and
624
- triggers: `evolve` (v0.7) is *automated, whole-skill* optimization — it hands the entire skill off
625
- to SkillOpt-Sleep to propose a fresh replacement, re-gates the proposal, and lets you adopt or
626
- reject it wholesale. `refine` is *targeted, human/agent-driven* — it captures one specific observed
627
- gap ("expected X, got Y") and writes the smallest override that closes it, tracked and eventually
628
- generalized only once the same gap recurs elsewhere. A `refine` patch on top of an `evolve`d skill
629
- is fine; both record their own provenance entry, so the `SOURCES.md` ledger shows which change came
630
- from which mechanism.
631
-
632
- **Legacy store — deliberate skip, not a migration.** `~/.claude/skill-refinements/` is a legacy
633
- location some users may have from earlier tooling — three flat markdown notes, no structured
634
- ledgers, no schema — and there is no migration from it. Files there stay readable in place; if
635
- anything in them still matters, re-capture it via `skill-forge refine` against the greenfield JSON
636
- store described in [`docs/refinement-schema.md`](docs/refinement-schema.md). Markdown ledgers like
637
- `refinement-history/*.md`, `aggregated-patterns.md`, or `generalization-queue.md` are not
638
- recreated — JSON is the store; `refine list`/`refine patterns` are the human-readable view over it.
639
-
640
- **Auto-trigger hooks — templates, not automation.** A CLI cannot hook a running Claude Code
641
- session, so two auto-trigger hooks ship here as documented templates instead:
642
- `assets/hooks/refinement-detector.sh` (detects refinement-shaped language in a prompt) and
643
- `assets/hooks/session-end.sh` (prompts after a substantial session). Both are optional, inert if
644
- `skill-forge` isn't on `PATH`, and never call `skill-forge` themselves — they only print a
645
- suggestion. Wire either one in by adding it to your `.claude/settings.json` (or
646
- `~/.claude/settings.json`) hooks section — the exact snippet is in each script's own header
647
- comment:
648
-
649
- ```json
650
- {
651
- "hooks": {
652
- "UserPromptSubmit": [
653
- { "hooks": [{ "type": "command", "command": "bash /path/to/refinement-detector.sh" }] }
654
- ],
655
- "SessionEnd": [
656
- { "hooks": [{ "type": "command", "command": "bash /path/to/session-end.sh" }] }
657
- ]
658
- }
659
- }
660
- ```
661
-
662
- **Security invariants** (same regime as v0.8/v0.9): nothing from a skill being refined is ever
663
- executed. Patch application writes ONLY under the resolved target scope dir for the named skill
664
- (containment-checked realpath, refuses a symlinked destination file); promotion writes ONLY under
665
- user scope, backup first. All child processes are argv arrays (git only, for context gathering).
666
- Control-char sanitization on every untrusted string that lands in human-readable output. `--yes`
667
- never implies `--handoff`. Applying a patch whose target `SKILL.md` is missing is refused; promoting
668
- a non-`ready` pattern without `--force` is refused.
669
-
670
- ### `routine` (v0.11, Pro)
671
-
672
- ```bash
673
- skill-forge routine # audit + drift + registry, one report, writes nothing else
674
- skill-forge routine --json --offline # for a scheduler: one JSON document, no network
675
- skill-forge routine --fail-on high # exit 1 when a high/critical finding exists
676
- skill-forge routine --housekeeping # opt in to bounded writes (see below)
677
- ```
678
-
679
- The whole maintenance pipeline in one command, shaped for cron. Every other command answers one
680
- question; keeping a set healthy means running four and correlating them by hand, which is exactly
681
- what nobody does on a schedule. `routine` runs the audit, drift, and registry **engines**,
682
- correlates them into one report and one summary, and exits with a code a scheduler can alert on.
683
-
684
- **Non-interactive by construction** — it never prompts, so it cannot hang a scheduled job. It is
685
- also the only schedulable command that writes, so what it may write is fenced:
686
-
687
- | | writes |
688
- |---|---|
689
- | always | the combined report + `audit-state.json`, both under `~/.skill-forge` |
690
- | `--housekeeping` | prunes quarantine sandboxes older than `--prune-days` (default 30); closes queue entries whose skill is gone from disk |
691
- | `--handoff` | launches your configured agent with the curation prompt |
692
-
693
- Housekeeping and handoff are **opt-in**: a bare `skill-forge routine` in a crontab reports and
694
- nothing else. A step that was requested but could not run — a crashed drift check, a failed
695
- prune, an unreadable queue — is listed under **Incomplete steps** and **always exits nonzero**,
696
- independent of `--fail-on`: `--fail-on` grades what the audit *found*, whereas a degraded run
697
- means the routine did not do what was scheduled, and a cron job that exits 0 on that is blind. Pruning only ever deletes inside the quarantine dir — never a skills root — is
698
- age-bounded, `--dry-run`-able, and lists every path in the report. A queue entry is only closed
699
- when the skill it points at no longer exists; a pending entry whose skill is still installed is
700
- undone work, not an orphan. `routine` never installs, promotes, or rejects: no candidate enters a
701
- skills root without a human or agent decision, and a cron job is not that.
702
-
703
- **Accepted findings (v0.12)** are honored the same way `audit` honors them: `--fail-on` grades
704
- **active** findings only, so accepting every remaining finding via `skill-forge finding accept`
705
- turns a red `--fail-on any` green, and any new or changed finding turns it red again. A
706
- corrupt/missing accepted-findings store fails open (nothing gets suppressed) and surfaces as a
707
- report **notice**, never as an "Incomplete step" — see
708
- [docs/accepted-findings-schema.md](docs/accepted-findings-schema.md).
709
-
710
- ```
711
- 0 9 * * 1 skill-forge routine --offline --fail-on high --housekeeping
712
- ```
713
-
714
- ### `promote <id>` / `reject <id>` (v0.11)
715
-
716
- Answering **hold** at the gate used to be a dead end: the sandbox sat in quarantine and no command
717
- could ever finish the decision. These close that loop.
718
-
719
- ```bash
720
- skill-forge promote <id> # re-gate the held sandbox, then promote/hold/reject
721
- skill-forge reject <id> # discard it
722
- ```
723
-
724
- Works for both artifact types: a held MCP candidate is re-gated as an MCP server and written into
725
- your MCP config (`--mcp-target`), with env values emptied and unpinned `npx` specs pinned, exactly
726
- as `add --artifact mcp` would.
727
-
728
- Promotion also re-validates the candidate against the state it was gated in: each skill dir must
729
- still be a real directory rather than a symlink, no symlink inside it may resolve out of the
730
- sandbox, and its content digest must still match the one captured at gate time. Otherwise the gap
731
- between "scanned" and "installed" — days, for a hold — is a window in which vetted content can be
732
- swapped for something that never passed the gate.
733
-
734
- `promote` **re-gates** rather than promoting blind. The original run's gate result was never
735
- persisted, so the alternative would be a queue entry with a fabricated gate record — and a hold
736
- may be days old, with the ruleset and your installed set since changed. Re-scanning is static and
737
- cheap. Ids come from `skill-forge list`; anything that resolves outside your quarantine dir
738
- (traversal, absolute path, symlink) is refused.
739
-
740
- ### `config` (v0.11)
741
-
742
- `businessProfile` is five fixed fields, so anything else learned about you has nowhere to live.
743
- Preferences are open-ended key/value facts that accumulate without a schema change:
744
-
745
- ```bash
746
- skill-forge config set prefers-typescript "strict mode, no any"
747
- skill-forge config list
748
- skill-forge config review
749
- ```
750
-
751
- **Agents get a different door.** `set` is for a human at a keyboard and takes effect immediately;
752
- `propose` only ever queues something for review:
753
-
754
- ```bash
755
- skill-forge config propose deploys-on vercel --origin "ingest agent" --note "seen in 3 skills"
756
- ```
757
-
758
- An agent draining the queue reads untrusted skill content, so anything it concludes about you is
759
- downstream of text an attacker may have written. If agents could write preferences directly, a
760
- malicious `SKILL.md` could talk one into persisting an instruction into your config, where every
761
- later run would read it back as *your stated preference* — a durable prompt-injection foothold
762
- with a laundering step in the middle. Routing agent writes through review means the worst case is
763
- a proposal you decline. `review` is interactive, or takes an explicit `--accept-all`/`--reject-all`;
764
- it refuses to fall through to a default in a non-interactive shell.
765
-
766
- Both doors run the same validation: kebab-case keys, a length cap, no control characters, and a
767
- refusal for anything credential-shaped — by key name (`api-key`, `*-token`) or by value shape
768
- (`sk-…`, `ghp_…`, `AKIA…`, PEM blocks, JWTs). `config.json` is plain text; secrets belong in your
769
- keychain. Hand-edited entries that fail those checks are dropped on load rather than rendered.
770
-
771
- ### `list` / `status`
772
-
773
- `list` shows what's currently held in quarantine (installed via `add`, answered "hold", not yet
774
- promoted or rejected). `status` shows the resolved configuration (skills roots, quarantine dir,
775
- strictness) plus a count of held entries.
776
-
777
- ### `guide` (v0.11)
778
-
779
- ```bash
780
- skill-forge guide # what this does, where you are, what to run next
781
- skill-forge guide verbs # a single topic
782
- ```
783
-
784
- Orientation for someone — or some agent — dropping into an existing session. `--help` lists flags
785
- but can't tell you what to do with them, and `status` prints configuration without saying what it
786
- means; `guide` answers "what is my current state, and what is the next command". It reads your
787
- config, quarantine, and queue and names one next step, unfinished work first: nothing configured →
788
- `init`; entries awaiting a decide-pass verb → `ingest`; candidates still held → `list`; no audit
789
- ever run → `audit`.
790
-
791
- ```
792
- Where you are:
793
- config: /Users/you/.skill-forge/config.json
794
- gate targets: 3 (default: /Users/you/.agents/skills)
795
- quarantined: 2
796
- queue: 12 pending
797
-
798
- Suggested next step:
799
- skill-forge ingest
800
- 12 promoted items still awaiting a decide-pass verb
801
- ```
802
-
803
- Topics: `pipeline`, `verbs`, `queue`, `mcp`, `refine`, `maintenance`, `pro`. **Read-only and
804
- non-interactive** — it writes nothing and never prompts, so it's safe to run inside an agent turn,
805
- and with no config it says so rather than reciting default placeholder paths as if they were
806
- settings.
807
-
808
- ## MCP gating (v0.5)
809
-
810
- `add` and `scan` can also gate an **MCP server** instead of a skill — the same quarantine →
811
- profile → safety scan → overlap analysis → report → promote/hold/reject pipeline, applied to an
812
- MCP server candidate rather than a skill:
813
-
814
- ```bash
815
- skill-forge scan ./my-mcp-server --artifact mcp
816
- skill-forge add @scope/some-mcp-server --artifact mcp --mcp-target ~/.claude.json
817
- skill-forge add https://github.com/owner/mcp-server.git --artifact mcp --yes
818
- ```
819
-
820
- Pass `--artifact mcp` explicitly — skill-forge never guesses. In particular, an npm package name
821
- (`@scope/name`, or a bare `some-mcp-server` name) is only treated as an npm source with this flag;
822
- without it, the same string resolves as a skills.sh `owner/name` slug instead (no ambiguity
823
- guessing between the two).
824
-
825
- **Source forms**
826
-
827
- | Form | Example | How it's fetched |
64
+ | Command | Description | Docs |
828
65
  |---|---|---|
829
- | Local directory | `./my-mcp-server` | Copied into quarantine, same as a local skill source. |
830
- | Git URL | `https://github.com/owner/mcp-server.git` | Shallow-cloned into quarantine, same as a git skill source. |
831
- | npm package | `@scope/name` or `some-mcp-server` | `npm pack <name> --ignore-scripts --pack-destination <quarantine>`, then tarball **extraction only** — never `npm install`, never lifecycle scripts. Every tarball member path is validated (absolute paths and `..` segments are rejected) before extraction. |
832
-
833
- **What's gated**
834
-
835
- Safety runs the same built-in deny-pattern ruleset used for skills (curl\|bash, credential-file
836
- access, etc.) plus MCP-specific rules: inline credential values in config/env (quoted or
837
- unquoted), unpinned `npx` launch commands — MEDIUM with `-y`/`--yes` (silent install), **LOW
838
- without it (v0.12)**, since npx still prompts before the first install but the tag floats once
839
- cached (a moving/dist tag like `@latest` counts as unpinned either way, and `npx` is recognized by
840
- basename so a full path or `npx.cmd` can't evade it) — `--dangerously-*`/`--no-sandbox` flags, and
841
- filesystem-root launch args — see the
842
- [MCP safety ruleset table](docs/gate-policy.md#mcp-safety-ruleset). Overlap analysis (Pro, free
843
- during the 0.x beta) ranks the candidate against the server entries already present in your
844
- configured `mcpTargets` files, instead of against a skills root.
845
-
846
- **Capability profile (v0.6).** The report also includes a **static** capability summary — tool,
847
- resource, and prompt counts, plus a `declaredConfidence` (`high`/`partial`/`none`) — parsed from
848
- the candidate's `package.json`, any shipped `.mcp.json`/manifest, and MCP SDK source-text patterns
849
- (`server.tool(...)`, `setRequestHandler(ListToolsRequestSchema, ...)`, etc.):
850
-
851
- ```
852
- Artifact type : mcp
853
- Capabilities : 2 tools, 1 resource, 0 prompts (declaredConfidence: high)
854
- ```
855
-
856
- This is free (it's profiling, not a Pro feature) and **never derived by running the candidate
857
- server** — anything that can't be determined from source text is reported as undetermined rather
858
- than discovered by executing it. `--json` includes the full `tools`/`resources`/`prompts` name
859
- lists under `profile.capabilities`.
860
-
861
- **Promote semantics**
862
-
863
- Promoting an MCP candidate writes ONE server entry into a target MCP config file's `mcpServers`
864
- map — it never touches a skills root:
865
-
866
- ```json
867
- { "mcpServers": { "<name>": { "command": "...", "args": ["..."], "env": { "SOME_KEY": "" } } } }
868
- ```
869
-
870
- - **Env values are never copied** from the candidate — every declared env var is written as an
871
- empty string, and skill-forge prints the var names you need to fill in yourself.
872
- - **Target resolution:** `--mcp-target <file>`, else `config.mcpTargets[0]`, else `<cwd>/.mcp.json`.
873
- - **Existing target file:** backed up first to `<file>.bak-<timestamp>`.
874
- - **Existing same-name server entry:** refused unless `--force` is passed.
875
- - **Missing target file/parent dirs:** created.
876
- - A candidate config listing more than one server gates/promotes the first (same "N found — using
877
- the first" convention `add`/`scan` already use for a multi-skill source), noted on stderr.
878
- - **Version pinning is enforced, including on a candidate's own documented entry (v0.7).** A
879
- candidate's `.mcp.json` server entry is used for `command`/`args` when present, but an unpinned
880
- `npx` spec there (`"args": ["-y", "pkg@latest"]`) is no longer written through as-is: it's
881
- **auto-pinned** to `<package>@<version>` when it matches the version skill-forge scanned from the
882
- candidate's own (self-declared) `package.json`, or **promotion is refused** (with an explanation
883
- and no override flag) when it can't be safely auto-pinned — a different package name, no version
884
- found, or a more complex shape (e.g. a `-p`/`--package` dependency, or more than one spec).
885
- See [docs/gate-policy.md](docs/gate-policy.md#mcp-promote-version-pin-enforcement-v07).
886
-
887
- There's no MCP equivalent of the skill provenance ledger (`SOURCES.md`) — the pending-ingestion
888
- queue (Pro) and `--ingest` handoff both apply the same way, keyed on the written config file path
889
- instead of an installed skill directory. The queued entry carries the candidate's static capability
890
- profile (`capabilities`, v0.6, above), so a `--ingest` handoff run on an MCP promote has real
891
- material to work with: `assets/ingest-prompt.md` branches on `artifactType: "mcp"` and walks the
892
- same five-verb decide (DEFER/ABSORB/FORK/REJECT/WATCH) applied to a server instead of a skill —
893
- compare declared capabilities against what's already configured, then keep/tighten/remove the
894
- promoted config entry accordingly. Same static-only rule as the CLI's own profiling: the deep pass
895
- never runs or installs the candidate server to inspect it.
896
-
897
- **TOML-format agents: detect-only.** `skill-forge init` detects MCP config files for every known
898
- agent, including TOML-format ones (Codex CLI's `config.toml`) — they show up in `init`'s MCP-target
899
- list and can be selected into `config.mcpTargets` for **overlap ranking**. But **promote only
900
- writes JSON-shaped targets** (`{ "mcpServers": { ... } }`); pointing `--mcp-target` at (or letting
901
- `config.mcpTargets[0]` resolve to) a TOML file fails when the promote step tries to parse it as
902
- JSON. Detection/overlap is agent-format-agnostic; writing is JSON-only.
66
+ | `init [options]` | Detect installed agents and set gate targets / handoff agent | [docs/commands/init.md](docs/commands/init.md) |
67
+ | `add <source> [options]` | Quarantine-install a skill and run it through the gate | [docs/commands/add.md](docs/commands/add.md) |
68
+ | `scan <source> [options]` | Gate a skill without installing it (always cleans up) | [docs/commands/scan.md](docs/commands/scan.md) |
69
+ | `evolve <skill-dir> [options]` (Pro) | Self-evolve an installed skill via SkillOpt-Sleep, re-gate, decide | [docs/commands/evolve.md](docs/commands/evolve.md) |
70
+ | `audit [options]` (alias `doctor`) | Doctor-style health check over the configured skill/MCP set | [docs/commands/audit.md](docs/commands/audit.md) |
71
+ | `finding accept\|revoke\|list` | Acknowledge, revoke, or list reviewed audit findings (human-only, no gate effect) | [docs/commands/finding.md](docs/commands/finding.md) |
72
+ | `organize [options]` (Pro) | Set-level capability registry + dependency graph across configured skills roots | [docs/commands/organize.md](docs/commands/organize.md) |
73
+ | `find [query] [options]` | Discover skills via skills.sh and check partner security audits (free) | [docs/commands/find.md](docs/commands/find.md) |
74
+ | `watch [options]` (Pro) | Drift check across every `SOURCES.md` provenance ledger | [docs/commands/watch.md](docs/commands/watch.md) |
75
+ | `ingest [options]` / `queue close <id>` (Pro) | Hand the pending-ingestion queue off to a coding agent for the decide/absorb pass | [docs/commands/ingest.md](docs/commands/ingest.md) |
76
+ | `refine [options]` (Pro) | Capture, list, review, promote, and roll back project-scope skill overrides | [docs/commands/refine.md](docs/commands/refine.md) |
77
+ | `insight <subcommand>` (Pro) | Turn sources into evidence-backed capability plans and Jira manifests | [docs/commands/insight.md](docs/commands/insight.md) |
78
+ | `routine [options]` (Pro) | One scheduled maintenance pass: audit + drift + registry, cron-friendly | [docs/commands/routine.md](docs/commands/routine.md) |
79
+ | `promote <id>` / `reject <id>` [options] | Re-gate and finish a held quarantine decision, or discard it | [docs/commands/promote-reject.md](docs/commands/promote-reject.md) |
80
+ | `config <sub> [args]` | Read/write open-ended preferences (`list`\|`get`\|`set`\|`unset`\|`propose`\|`review`) | [docs/commands/config.md](docs/commands/config.md) |
81
+ | `list` / `status` | List held quarantine entries, or show configuration and quarantine summary | [docs/commands/list-status.md](docs/commands/list-status.md) |
82
+ | `guide [topic]` | Orientation: what this does, your current state, the next command | [docs/commands/guide.md](docs/commands/guide.md) |
83
+
84
+ `add`/`scan` also gate an **MCP server** instead of a skill via `--artifact mcp` — see
85
+ [docs/mcp-gating.md](docs/mcp-gating.md).
903
86
 
904
87
  ## Free vs. Pro
905
88
 
906
- | Capability | Free | Pro |
907
- |---|:---:|:---:|
908
- | Quarantine install (skills.sh slug / git / local path) | ✓ | ✓ |
909
- | Profile (name, version, license, structure, MCP/tool deps) | ✓ | ✓ |
910
- | Safety gate — built-in ruleset + SkillSpector shell-out | ✓ | ✓ |
911
- | Terminal report + `--json` | ✓ | ✓ |
912
- | Promote / hold / reject decision | ✓ | ✓ |
913
- | `init` setup wizard (agent detection, gate targets, handoff agent) | ✓ | ✓ |
914
- | Overlap analysis against your configured skill set | | ✓ |
915
- | Provenance ledger (`SOURCES.md` audit trail) | | ✓ |
916
- | Pending-ingestion queue + `--ingest` handoff | | ✓ |
917
- | `evolve` — SkillOpt-Sleep self-evolution, re-gating, provenance, queueing (v0.7) | | ✓ |
918
- | `audit` — inventory, hygiene findings, report, business profile, foundation scaffold (v0.8) | ✓ | ✓ |
919
- | `audit`'s cross-root overlap clusters (v0.8) | | ✓ |
920
- | `organize` — set-level capability registry + dependency graph (v0.9) | | ✓ |
921
- | `find` — skills.sh discovery + partner audit verdicts (v0.9) | ✓ | ✓ |
922
- | `watch` — provenance drift check across every `SOURCES.md` ledger (v0.9) | | ✓ |
923
- | `ingest` + `queue close` — pending-queue drain handoff (v0.9) | | ✓ |
924
- | `refine` — capture/apply/generalize project-scope skill overrides (v0.10) | | ✓ |
925
- | `finding accept`/`revoke`/`list` — acknowledge audit findings, never a gate input (v0.12) | ✓ | ✓ |
926
-
927
89
  Free is the complete safety gate on its own — quarantine, profile, safety scan, and an explicit
928
90
  promote/reject decision, with nothing held back. Pro is the curation layer on top: whether a new
929
91
  candidate duplicates something you already have, and an ongoing provenance record across your
@@ -931,73 +93,29 @@ whole skill set rather than a single install-time decision.
931
93
 
932
94
  **Everything free until 1.0.** This is a 0.x beta build, and the Pro tier's runtime license check
933
95
  is intentionally asleep for the whole 0.x line: overlap analysis, the provenance ledger, and the
934
- pending-ingestion queue / `--ingest` handoff all run for everyone, licensed or not. Without a valid
935
- license (`SKILL_FORGE_LICENSE` env var or `config.json`'s `licenseKey`, verified offline), `add`
936
- prints a one-line notice — `Pro feature (...) — free during the 0.x beta; will require a license at
937
- 1.0.` — above the report (or in the `--json` payload's `notices` array) and otherwise runs exactly
938
- as a licensed run would. At 1.0 the lock re-arms and these features go back to requiring a valid
939
- key. The set-level organizer and the skills.sh partner-audit enrichment are not yet exposed by any
940
- CLI command, licensed or not, beta or not. See [docs/pro.md](docs/pro.md) for the per-feature
941
- implementation status.
96
+ pending-ingestion queue / `--ingest` handoff all run for everyone, licensed or not — `add` prints a
97
+ one-line "free during the beta" notice above the report when no valid license is set, and otherwise
98
+ runs exactly as a licensed run would.
99
+
100
+ See [docs/pro.md](docs/pro.md#free-vs-pro) for the full Free/Pro capability table and the
101
+ per-feature implementation status.
942
102
 
943
103
  ## Security model
944
104
 
945
- - **Quarantine-first.** Every source — skills.sh slug, git URL, or local path — is installed into
946
- an isolated sandbox (`~/.skill-forge/quarantine/<id>/`) before anything is inspected. Nothing
947
- reaches your working skill set without an explicit promote decision.
105
+ - **Quarantine-first.** Every source is installed into an isolated sandbox before anything is
106
+ inspected — nothing reaches your working skill set without an explicit promote decision.
948
107
  - **Built-in safety ruleset, always on, fully offline.** A deny-pattern scan (curl/wget-into-shell,
949
- base64-obfuscated exec, reverse shells, recursive force-delete, credential-file access/exfil,
950
- dynamic eval, `shell=True` subprocess, persistence via shell rc files or cron, `sudo` usage) runs
951
- against every candidate with no external dependency and no network call. See
952
- [docs/gate-policy.md](docs/gate-policy.md) for the full rule table.
108
+ base64-obfuscated exec, reverse shells, credential-file access/exfil, `shell=True` subprocess,
109
+ persistence via shell rc files or cron, `sudo` usage, and more) runs against every candidate with
110
+ no external dependency and no network call.
953
111
  - **Block on HIGH/CRITICAL.** Any finding at `HIGH` or `CRITICAL` severity blocks the candidate
954
- outright (verdict `block`); a lower-severity finding produces `warn`; a clean scan is `pass`.
955
- `--yes` honors this: `block` is rejected automatically.
956
- - **npm-sourced MCP candidates: `--ignore-scripts` + tarball member validation.** An npm-package
957
- MCP source is fetched with `npm pack --ignore-scripts`, and the tarball's member list is checked
958
- with `tar -tzf` before extraction — any absolute path or `..` path segment refuses the extract
959
- with an error instead of running `tar -xzf` (a path-traversal guard against a malicious tarball
960
- writing outside the quarantine sandbox).
961
- - **SkillSpector, when installed.** If [SkillSpector](https://github.com/NVIDIA/SkillSpector)
962
- (Apache-2.0) is on `PATH`, skill-forge shells out to it (`--no-llm` by default, so scanned skill
963
- content is never sent to an external LLM provider) and merges its findings into the same report.
964
- Purely additive — its absence never blocks the gate.
965
- - **skills.sh partner audits — implemented, not yet wired into the CLI pipeline.** skill-forge
966
- ships a client for skills.sh's documented `/api/v1/skills/audit` endpoint (partner verdicts from
967
- Socket, Snyk, Gen Agent Trust Hub, Runlayer, ZeroLeaks), gated on a user-supplied
968
- `VERCEL_OIDC_TOKEN`. The client exists (`src/gate/skillsSh.ts`) but `add`/`scan` do not call it
969
- yet in this build — see [docs/gate-policy.md](docs/gate-policy.md) for current status.
970
-
971
- ## Ingestion handoff (`--ingest`)
972
-
973
- `skill-forge` deliberately doesn't try to decide *what to extract* from a skill worth adopting —
974
- that deeper judgment (which patterns to keep, whether to absorb into an existing skill vs. fork a
975
- new one, verifying the result beats baseline) is a job for a coding agent, not the gate. `--ingest`
976
- hands a promoted skill off to one, running the bundled, agent-neutral prompt at
977
- `assets/ingest-prompt.md` — works with any agent. ABSORB extractions from an ingest pass route
978
- through `skill-forge refine` in this same package — see [`refine`](#refine-v010-pro) above and
979
- [docs/forge-workflow.md](docs/forge-workflow.md).
980
- The same flag works on an MCP server promote (`--artifact mcp --ingest`, v0.6): the bundled prompt branches on the
981
- queue entry's `artifactType` and runs the matching decide pass — see
982
- [MCP gating](#mcp-gating-v05) above.
983
-
984
- ```bash
985
- skill-forge add owner/name --yes --ingest
986
- ```
987
-
988
- The command that gets run, in order:
112
+ outright; a lower-severity finding produces `warn`; a clean scan is `pass`.
113
+ - **npm-sourced MCP candidates:** fetched with `--ignore-scripts`, and every tarball member path is
114
+ validated before extraction (a path-traversal guard).
115
+ - **SkillSpector, when installed,** is purely additive — its absence never blocks the gate.
989
116
 
990
- 1. `config.handoffCommand` — an argv-style template (`["claude", "-p", "{prompt}"]`-shaped) set by
991
- `skill-forge init`'s "handoff agent" prompt, with `{path}` (the installed skill) and `{prompt}`
992
- (the bundled prompt file) substituted in. Never shell-parsed, so it's safe even if a substituted
993
- path contains shell metacharacters.
994
- 2. Otherwise, the first known agent binary found on `PATH` (`claude`, `codex`, `cursor-agent`,
995
- `windsurf`, `opencode`, `gemini`), invoked generically with the prompt.
996
- 3. Otherwise, skill-forge prints the prompt path and skill path for you to hand off yourself.
997
-
998
- Every promote or hold is recorded to the pending queue (`~/.skill-forge/queue.json`) regardless of
999
- `--ingest` — nothing is lost if you skip the handoff. See [`docs/queue-schema.md`](docs/queue-schema.md)
1000
- for the entry schema.
117
+ See [docs/security-model.md](docs/security-model.md) for the full rule table and current status of
118
+ each check.
1001
119
 
1002
120
  ## Configuration
1003
121
 
@@ -1049,6 +167,18 @@ the provenance ledger, and the pending-ingestion queue / `--ingest` handoff all
1049
167
  with a one-line "free during the beta" notice if no valid key is set. See
1050
168
  [docs/pro.md](docs/pro.md) for details.
1051
169
 
170
+ **How do I roll back a customization?**
171
+ It depends what changed. If `skill-forge init` gave a skills root a Git baseline (or you'd already
172
+ committed it yourself), rolling back a change under that root is ordinary Git —
173
+ `git checkout -- <path>` or `git reset`, run by you; skill-forge doesn't ship its own rollback for
174
+ that. `refine rollback <backup-id>` is different and narrower: it only undoes a `refine promote`
175
+ (a base-skill update recorded by `skill-forge refine`), not a Git commit or anything else. See
176
+ [Git preflight](docs/commands/init.md#git-preflight-v016) and
177
+ [`refine`](docs/commands/refine.md) for both.
178
+
179
+ **Where's the full documentation?**
180
+ See the [docs index](docs/README.md) for every command reference and background doc.
181
+
1052
182
  ## License
1053
183
 
1054
184
  skill-forge is open-core with a split license (as of v0.2.0):
@@ -1058,7 +188,8 @@ skill-forge is open-core with a split license (as of v0.2.0):
1058
188
  - **Pro modules — Rhize Commercial License.** `src/license.ts`, `src/gate/overlap.ts`,
1059
189
  `src/provenance.ts`, `src/queue.ts` (overlap analysis, provenance ledger, pending
1060
190
  queue / `--ingest` handoff), and `src/refine/` (v0.10 — capture/apply/generalize skill
1061
- overrides). The source is available to read and audit, but production use of Pro
191
+ overrides), plus `src/insights/` (source studies, evidence validation, inventory mapping,
192
+ plans, and Jira manifests). The source is available to read and audit, but production use of Pro
1062
193
  functionality requires a license key — see [LICENSE-COMMERCIAL](LICENSE-COMMERCIAL).
1063
194
 
1064
195
  [LICENSE](LICENSE) is the authoritative map of which files fall under which license.