@rhize/skill-forge 0.6.0 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE CHANGED
@@ -18,6 +18,7 @@ Pro Modules
18
18
  - src/gate/mcpOverlap.ts
19
19
  - src/provenance.ts
20
20
  - src/queue.ts
21
+ - src/evolve.ts
21
22
 
22
23
  ...and the portions of built artifacts (e.g. dist/cli.js in the published npm
23
24
  package) generated from these files. Each Pro Module carries a header
package/README.md CHANGED
@@ -65,11 +65,12 @@ anything ending in `.git`), or a local filesystem path.
65
65
  ## Commands
66
66
 
67
67
  ```
68
- skill-forge init [options] Detect installed agents and set gate targets / handoff agent
69
- skill-forge add <source> [options] Quarantine-install a skill and run it through the gate
70
- skill-forge scan <source> [options] Gate a skill without installing it (always cleans up)
71
- skill-forge list List skills currently held in quarantine
72
- skill-forge status Show configuration and quarantine summary
68
+ skill-forge init [options] Detect installed agents and set gate targets / handoff agent
69
+ skill-forge add <source> [options] Quarantine-install a skill and run it through the gate
70
+ skill-forge scan <source> [options] Gate a skill without installing it (always cleans up)
71
+ skill-forge evolve <skill-dir> [options] Self-evolve an installed skill via SkillOpt-Sleep, re-gate, decide (Pro)
72
+ skill-forge list List skills currently held in quarantine
73
+ skill-forge status Show configuration and quarantine summary
73
74
  ```
74
75
 
75
76
  ### `init`
@@ -132,6 +133,69 @@ anything — the quarantine sandbox is always cleaned up afterward, on success o
132
133
  nonzero when the safety verdict is `block`. `--json` prints the same gate-result payload shape as
133
134
  `add`'s.
134
135
 
136
+ ### `evolve` (v0.7)
137
+
138
+ ```bash
139
+ skill-forge evolve .claude/skills/my-skill
140
+ skill-forge evolve .claude/skills/my-skill --yes
141
+ skill-forge evolve .claude/skills/my-skill --dry-run
142
+ skill-forge evolve .claude/skills/my-skill --backend claude --yes
143
+ ```
144
+
145
+ Pro (free during the 0.x beta). Orchestrates [microsoft/SkillOpt](https://github.com/microsoft/SkillOpt)'s
146
+ `skillopt-sleep` CLI (`pip install skillopt`) to propose a self-evolution of an already-installed
147
+ skill — harvest recent sessions, generate a candidate replacement `SKILL.md`/`CLAUDE.md`, and stage
148
+ it — then runs the **staged proposal**, never the live skill, back through skill-forge's own static
149
+ safety ruleset before you decide anything. SkillOpt-Sleep's own validation gate is score-only; it
150
+ never content-vets the generated markdown. This closes that gap.
151
+
152
+ | Option | Effect |
153
+ |---|---|
154
+ | `--project <dir>` | Project dir passed to SkillOpt-Sleep as `--project`. Defaults to cwd. |
155
+ | `--dry-run` | Uses SkillOpt-Sleep's `dry-run` subcommand — report only, nothing staged. |
156
+ | `--backend <name>` | SkillOpt-Sleep backend. Defaults to `mock` (offline, deterministic, no network). |
157
+ | `--lookback-hours <n>` | Hours of session history for SkillOpt-Sleep to harvest. |
158
+ | `-y, --yes` | Skips interactive prompts (the disclosure confirm below, and the promote/hold/reject prompt) and honors the re-gate verdict automatically. |
159
+ | `--json` | Prints the re-gate result as JSON instead of the terminal report. Implies non-interactive, same as `add`'s `--json`. |
160
+ | `--force` | Allows re-adopting a staging dir whose proposed content already matches the live skill (see the double-adopt guard below). |
161
+
162
+ **Requires `skillopt-sleep` on PATH** — skill-forge never installs it for you. If it's missing,
163
+ `evolve` prints `pip install skillopt` + a docs pointer and exits, rather than attempting an
164
+ auto-install (the same "detect, don't install" discipline the rest of the gate follows).
165
+
166
+ **Data boundary.** The default `mock` backend is fully offline — harvesting and proposal
167
+ generation both run locally, no session data leaves the machine. Any other backend (`claude`,
168
+ `codex`, `azure_openai`, ...) sends truncated excerpts from harvested sessions and derived tasks to
169
+ the provider you selected; per SkillOpt-Sleep's own docs this is not currently guaranteed to be
170
+ secret-free. `evolve` prints that disclosure and requires either `--yes` or an interactive `y/N`
171
+ confirmation before a non-`mock` run proceeds — `--json` is non-interactive, so a non-`mock` run
172
+ under `--json` without `--yes` is refused rather than silently sending data off-machine.
173
+
174
+ **Re-gate.** SkillOpt-Sleep only ever *stages* a proposal (`<project>/.skillopt-sleep/staging/<timestamp>/`
175
+ — full replacement files, never auto-adopted). `evolve` copies just `proposed_SKILL.md` (and
176
+ `proposed_CLAUDE.md`, if present) into a fresh temporary directory and runs skill-forge's own
177
+ static safety ruleset over that copy — the same one `add`/`scan` use, at your configured
178
+ `strictness`. `report.json`/`diagnostics.json` (SkillOpt-Sleep's own redacted holdout evidence) are
179
+ never scanned. The result renders through the same terminal report / `--json` shape as `add`/`scan`.
180
+
181
+ **Decision.** Same promote/hold/reject semantics as `add`: `--yes`/`--json` honor the re-gate
182
+ verdict (a `block` is rejected); otherwise you're prompted.
183
+
184
+ - **promote** — runs SkillOpt-Sleep's own `adopt --staging <dir>` (which backs up the live file(s)
185
+ before copying the proposal over them), then records a provenance entry to the target skill's
186
+ `SOURCES.md` and a pending-ingestion queue entry (`origin: "evolve"`, see
187
+ [docs/queue-schema.md](docs/queue-schema.md)) so a later ingest pass reviews the evolution rather
188
+ than an external source.
189
+ - **hold** — leaves the staging dir exactly as SkillOpt-Sleep produced it (its own `status` command
190
+ still lists it); `report.md` has the full evidence.
191
+ - **reject** — deletes the staging dir.
192
+
193
+ **Double-adopt guard.** SkillOpt-Sleep's `adopt` has no confirmation of its own, and adopting the
194
+ same staging dir twice overwrites *its* backup — the pre-evolution original would be lost. Before
195
+ adopting, `evolve` refuses (unless `--force`) when the staged proposal is already byte-identical to
196
+ the live skill it would replace, since that's the signature of a staging dir that was already
197
+ adopted once.
198
+
135
199
  ### `list` / `status`
136
200
 
137
201
  `list` shows what's currently held in quarantine (installed via `add`, answered "hold", not yet
@@ -161,13 +225,15 @@ guessing between the two).
161
225
  |---|---|---|
162
226
  | Local directory | `./my-mcp-server` | Copied into quarantine, same as a local skill source. |
163
227
  | Git URL | `https://github.com/owner/mcp-server.git` | Shallow-cloned into quarantine, same as a git skill source. |
164
- | npm package | `@scope/name` or `some-mcp-server` | `npm pack <name> --pack-destination <quarantine>`, then tarball **extraction only** — never `npm install`, never lifecycle scripts. |
228
+ | npm package | `@scope/name` or `some-mcp-server` | `npm pack <name> --ignore-scripts --pack-destination <quarantine>`, then tarball **extraction only** — never `npm install`, never lifecycle scripts. Every tarball member path is validated (absolute paths and `..` segments are rejected) before extraction. |
165
229
 
166
230
  **What's gated**
167
231
 
168
232
  Safety runs the same built-in deny-pattern ruleset used for skills (curl\|bash, credential-file
169
- access, etc.) plus MCP-specific rules: inline credential values in config/env, unpinned `npx -y`
170
- launch commands, `--dangerously-*`/`--no-sandbox` flags, and filesystem-root launch args — see the
233
+ access, etc.) plus MCP-specific rules: inline credential values in config/env (quoted or
234
+ unquoted), unpinned `npx -y` launch commands (a moving/dist tag like `@latest` counts as
235
+ unpinned, and `npx` is recognized by basename so a full path or `npx.cmd` can't evade it),
236
+ `--dangerously-*`/`--no-sandbox` flags, and filesystem-root launch args — see the
171
237
  [MCP safety ruleset table](docs/gate-policy.md#mcp-safety-ruleset). Overlap analysis (Pro, free
172
238
  during the 0.x beta) ranks the candidate against the server entries already present in your
173
239
  configured `mcpTargets` files, instead of against a skills root.
@@ -235,6 +301,7 @@ JSON. Detection/overlap is agent-format-agnostic; writing is JSON-only.
235
301
  | Overlap analysis against your configured skill set | | ✓ |
236
302
  | Provenance ledger (`SOURCES.md` audit trail) | | ✓ |
237
303
  | Pending-ingestion queue + `--ingest` handoff | | ✓ |
304
+ | `evolve` — SkillOpt-Sleep self-evolution, re-gating, provenance, queueing (v0.7) | | ✓ |
238
305
  | Set-level organizer (capability registry, redundancy, dependency graph) | | ✓ |
239
306
 
240
307
  Free is the complete safety gate on its own — quarantine, profile, safety scan, and an explicit
@@ -266,6 +333,11 @@ implementation status.
266
333
  - **Block on HIGH/CRITICAL.** Any finding at `HIGH` or `CRITICAL` severity blocks the candidate
267
334
  outright (verdict `block`); a lower-severity finding produces `warn`; a clean scan is `pass`.
268
335
  `--yes` honors this: `block` is rejected automatically.
336
+ - **npm-sourced MCP candidates: `--ignore-scripts` + tarball member validation.** An npm-package
337
+ MCP source is fetched with `npm pack --ignore-scripts`, and the tarball's member list is checked
338
+ with `tar -tzf` before extraction — any absolute path or `..` path segment refuses the extract
339
+ with an error instead of running `tar -xzf` (a path-traversal guard against a malicious tarball
340
+ writing outside the quarantine sandbox).
269
341
  - **SkillSpector, when installed.** If [SkillSpector](https://github.com/NVIDIA/SkillSpector)
270
342
  (Apache-2.0) is on `PATH`, skill-forge shells out to it (`--no-llm` by default, so scanned skill
271
343
  content is never sent to an external LLM provider) and merges its findings into the same report.