@rhize/skill-forge 0.11.3 → 0.13.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -70,6 +70,9 @@ skill-forge add <source> [options] Quarantine-install a skill and ru
70
70
  skill-forge scan <source> [options] Gate a skill without installing it (always cleans up)
71
71
  skill-forge evolve <skill-dir> [options] Self-evolve an installed skill via SkillOpt-Sleep, re-gate, decide (Pro)
72
72
  skill-forge audit [options] Doctor-style health check over the configured skill/MCP set (alias: doctor)
73
+ skill-forge finding accept <fp-prefix> Acknowledge a LOW/MEDIUM audit finding you've reviewed (human-only, no gate effect)
74
+ skill-forge finding revoke <fp-prefix> Remove a stored acceptance so the finding reports again
75
+ skill-forge finding list [options] List every accepted finding
73
76
  skill-forge organize [options] Set-level capability registry + dependency graph across configured skills roots (Pro)
74
77
  skill-forge find [query] [options] Discover skills via skills.sh and check partner security audits (free)
75
78
  skill-forge watch [options] Drift check across every SOURCES.md provenance ledger (Pro)
@@ -169,6 +172,18 @@ target skills root until you decide.
169
172
  | `-t, --target <dir>` | Skills root to promote into. Defaults to the config's `defaultTarget`, then `skillsRoots[0]`. |
170
173
  | `--json` | Prints the gate result (profile, safety findings, overlap) as JSON instead of the terminal report box. Implies non-interactive: the decision is made the same way `--yes` makes it (verdict decides promote/hold/reject), never an interactive prompt. |
171
174
  | `--ingest` | Pro (free during the 0.x beta). After a successful promote, hands off to a coding agent — see [Ingestion handoff](#ingestion-handoff---ingest). |
175
+ | `--skill-map <path>` | Path to a generated `rhize-plugins` skill map (Phase 4 of that repo's skill-map-graph-substrate plan). When given, ranks the candidate's name/description against every `skill` node in the map — a near-duplicate of an already-shipped marketplace skill is folded into the safety findings as a `HIGH`-severity finding (escalates the verdict to `block`, so it's held rather than silently promoted); a moderate overlap is `MEDIUM` (`warn`). Free (not Pro-gated) — this enforces the marketplace's own curation rule ("close the gap, don't duplicate"), not a premium ranking feature. Missing/unreadable map: printed to stderr, never fatal. **Two map variants exist, with different coverage**: the **static** map (`rhize-plugins/generated/skill-map.static.json`) is first-party-only — it only knows about that marketplace's own skills, so it cannot catch a candidate that duplicates an *installed third-party plugin's* skill. The **resolved** map (`~/.claude/context-manager/skill-map.resolved.json`, produced by `rhize-context-manager`) additionally carries the full third-party ecosystem inventory. Point `--skill-map` at the resolved map when you want ecosystem-wide duplicate detection; the static map is enough only when you're curating a single first-party marketplace against itself. |
176
+
177
+ **Bugfix (v0.13.0):** `--skill-map` previously matched nothing against either real map variant — the loader read a `type` field that the real generated artifact never sets (it uses `kind`), so every node was silently invisible to the overlap check. `loadSkillMap()` now normalizes `kind` → `type` at load time; regression coverage lives in `test/mapOverlap.realmap.test.ts` against vendored real-map snapshots. If you were relying on `--skill-map` before v0.13.0, it was a no-op — re-run `add`/`scan` on anything you're unsure about.
178
+
179
+ **Extends-declared overlap exemption**: a candidate can declare `metadata.rhize.extends:
180
+ ["<skill-name>" | "<plugin>/<skill-name>"]` in its SKILL.md frontmatter to mark itself a
181
+ deliberate specialization/layering of an existing map skill. When the top `--skill-map` match is
182
+ a skill the candidate declares it extends, that finding is downgraded from a HOLD-level safety
183
+ finding to an informational notice printed above the report — it never escalates the verdict.
184
+ Overlap with any *other* skill (a different match, or a second candidate/skill pair) still holds
185
+ at full severity: the exemption applies per declared pair, not globally, so declaring an
186
+ extension of skill X never waives overlap with skill Y.
172
187
 
173
188
  A promoted skill gets a provenance entry appended to `<target>/SOURCES.md`, and every promote or
174
189
  hold decision is recorded to `~/.skill-forge/queue.json` (or `$SKILL_FORGE_HOME/queue.json`) — see
@@ -190,7 +205,7 @@ skill-forge scan owner/name --json
190
205
  Runs the same gate pipeline as `add` (profile → safety → overlap → report) but never promotes
191
206
  anything — the quarantine sandbox is always cleaned up afterward, on success or failure. Exits
192
207
  nonzero when the safety verdict is `block`. `--json` prints the same gate-result payload shape as
193
- `add`'s.
208
+ `add`'s. Also accepts `--skill-map <path>` — see `add`'s option table above.
194
209
 
195
210
  ### `evolve` (v0.7)
196
211
 
@@ -293,6 +308,13 @@ basename, package spec, arg/env **counts** — never values); TOML targets (e.g.
293
308
  structured TOML enumeration. A per-item failure (unreadable skill, broken symlink, malformed MCP
294
309
  config) becomes a finding — it never aborts the run.
295
310
 
311
+ **Accepted findings (v0.12).** Every finding in the report carries a fingerprint
312
+ (`(fp a1b2c3d4e5f6)`); `Findings` shows **active** findings only, and a separate
313
+ **Accepted findings** section lists whatever you've reviewed and accepted via
314
+ [`skill-forge finding accept`](#finding-v012) — permanently, with reason and date, even after it
315
+ stops matching (see that section for the full contract). `summary.acceptedCount` and each
316
+ finding's `fingerprint` are additive `--json` fields; nothing accepted is ever silently deleted.
317
+
296
318
  **Opportunity pass.** Beyond hygiene, the report surfaces:
297
319
 
298
320
  - **Overlap clusters** (Pro, free during the 0.x beta) — cross-root overlap scoring across every
@@ -331,6 +353,49 @@ in place of `ingest-prompt.md`. `--yes` (and `init --defaults`, which passes it
331
353
  implies either — a non-interactive run writes only the report and, if a profile was already
332
354
  stored, the config; nothing else.
333
355
 
356
+ ### `finding` (v0.12)
357
+
358
+ ```bash
359
+ skill-forge finding accept <fingerprint-prefix> --reason "Funnel genuinely needs root; verified"
360
+ skill-forge finding revoke <fingerprint-prefix>
361
+ skill-forge finding list [--stale] [--json]
362
+ ```
363
+
364
+ Acknowledge an `audit` finding you've reviewed and accepted, so it stops counting against
365
+ `routine --fail-on` and the `add`/`scan` advisory while staying **permanently visible** in its own
366
+ report section — an audit you can't triage decays into noise, and noise is how the one real
367
+ finding gets skimmed. Every finding in an `audit`/`routine` report now carries a 12-char
368
+ fingerprint (`(fp a1b2c3d4e5f6)`) you copy into `accept`/`revoke`.
369
+
370
+ | Command | Effect |
371
+ |---|---|
372
+ | `finding accept <fp-prefix> --reason <text>` | Records an acceptance. Re-runs the read-only audit engine to resolve the prefix against what's **currently active** — you can only accept a finding that exists right now, never from a stale report. `--reason` is mandatory (1-200 chars, no control characters). |
373
+ | `finding revoke <fp-prefix>` | Removes a stored acceptance so the finding reports again. Resolves against the **store**, not a fresh audit run — the target may no longer be producible, which is exactly when you'd want to clean up. |
374
+ | `finding list [--stale] [--json]` | Lists every accepted finding (severity, rule, target, reason, accepted date, fingerprint). `--stale` filters to records whose target no longer exists on disk. |
375
+
376
+ **Identity is content-derived, not a name you pick.** A finding's fingerprint is
377
+ `sha256(severity + rule + target + the offending line's text)` — deliberately excluding the LINE
378
+ NUMBER (drifts on unrelated edits) and deliberately including SEVERITY (so a rule re-tuned to a
379
+ higher severity for the same content can't inherit an ack minted against the lower one). Any change
380
+ to the offending content **fails open to re-reporting** — this is the design, not a bug: the
381
+ alternative (key on rule+path only) is the classic baseline trap, where accepting one benign line
382
+ silently suppresses every future line the same rule matches in that file. Full contract, the
383
+ per-rule `context` table, and fail-open read semantics: see
384
+ [docs/accepted-findings-schema.md](docs/accepted-findings-schema.md).
385
+
386
+ **Severity-capped, human-only, no override.** `finding accept` refuses `HIGH`/`CRITICAL` findings
387
+ outright — there is no `--force`. The precision problem this command exists to solve lives at
388
+ `LOW`/`MEDIUM`; a `HIGH` false positive is a rule bug to fix, not something to baseline. This is a
389
+ human-door CLI (`config set` model), not the agent-facing propose/review queue — both
390
+ `assets/ingest-prompt.md` and `assets/curation-prompt.md` instruct agents to never run
391
+ `finding accept`/`finding revoke` themselves; an agent that judges a finding a false positive says
392
+ so in its summary and lets the user run the command.
393
+
394
+ **Never a gate input.** Acceptances only ever affect `audit`'s own report classification (and,
395
+ downstream, `routine --fail-on` and the `add`/`scan` advisory) — `add`/`scan`/`promote <id>`/
396
+ `evolve`'s re-gate never reads the accepted-findings store. See
397
+ [docs/gate-policy.md](docs/gate-policy.md#accepted-findings-v012).
398
+
334
399
  ### `organize` (v0.9, Pro)
335
400
 
336
401
  ```bash
@@ -404,6 +469,7 @@ shell — you (or your shell's dotenv loader) still need to export it before `fi
404
469
  skill-forge watch
405
470
  skill-forge watch --offline
406
471
  skill-forge watch --json
472
+ skill-forge watch --skill-map generated/skill-map.static.json
407
473
  ```
408
474
 
409
475
  Pro (free during the 0.x beta). TS port of the plugin's `record_provenance.py --check-drift`.
@@ -420,6 +486,45 @@ entry is listed for manual checking regardless.
420
486
  attacker-influenceable data — anyone who can write to a skills root's `SOURCES.md` controls it.
421
487
  It is only ever printed, sanitized, as a suggestion for you (or an agent) to run yourself.
422
488
 
489
+ **`--skill-map <path>` (Phase 4 of `rhize-plugins`' skill-map-graph-substrate plan)** additionally
490
+ drift-checks every `fork-of` edge in a generated skill map: for each edge it resolves the local
491
+ `skill` node's `path`/`contentHash` and the upstream (`external`) node's `url`/`path`, fetches or
492
+ reads the upstream content, and compares content hashes. Same never-execute posture as the ledger
493
+ check above: nothing derived from the map — including a fork-of edge's own `driftCheck`
494
+ metadata — is ever executed; only fetch/read/hash/compare. A missing/unreadable map is a warning,
495
+ never fatal. A node's repo-relative `path` (e.g. `rhize-context-manager/skills/x/SKILL.md`) is
496
+ resolved against cwd, the map's own directory, and that directory's parent (in that order) — not
497
+ cwd alone — so `watch --skill-map <path>` gives correct verdicts regardless of the directory you
498
+ run it from.
499
+
500
+ **Verdict**, per fork-of edge — three-way comparison, four verdict states when both hashes below
501
+ are present on the map, else the older two-way fallback (`in-sync`/`drifted`, `contentHash` vs
502
+ freshly-fetched upstream):
503
+
504
+ | local-normalized vs baseline | upstream-now vs baseline | status | actionable |
505
+ |---|---|---|---|
506
+ | == | == | `in-sync` | no |
507
+ | != | == | `local-only` | no (deliberate fork divergence, e.g. Rhize's added frontmatter) |
508
+ | == | != | `upstream-moved` | yes |
509
+ | != | != | `diverged` | yes |
510
+
511
+ The three-way matrix fires only when the local `skill` node carries `contentHashNormalized` (a
512
+ hash of the file with Rhize-injected frontmatter stripped, computed once by the rhize-plugins
513
+ compiler) **and** the upstream `external` node carries `baselineHash` (the upstream content hash
514
+ as of the last human review, recorded in `rhize-plugins`' SOURCES.md and copied onto the node by
515
+ that compiler — skill-forge never computes or fetches it itself). Either field missing on a given
516
+ edge falls back to the two-way compare for that edge, so older maps keep working. All four
517
+ three-way verdicts — `in-sync`, `local-only`, `upstream-moved`, and `diverged` — carry
518
+ `baselineHash`/`upstreamHash` in the JSON output; the two-way-fallback rows and
519
+ `upstream-unreachable`/`local-missing` never do (no verdict without a successful fetch/read, so
520
+ there is nothing to compare against a baseline). Every row also carries an `actionable` boolean
521
+ (from `isActionable`), so callers don't have to re-derive "needs attention" from `status`/`detail`
522
+ prose — it's `true` for `drifted`, `upstream-moved`, `diverged`, `upstream-unreachable`, and
523
+ `local-missing`; `false` for `in-sync` and `local-only`. **Re-baselining**: after reviewing and
524
+ adopting an `upstream-moved`/`diverged` change, re-run `rhize-plugins`' `scripts/baseline_upstreams.py`
525
+ and commit the updated SOURCES.md — that's the "I reviewed upstream, accept its state" action that
526
+ clears the row back to `in-sync`/`local-only`.
527
+
423
528
  ### `ingest` + `queue close` (v0.9, Pro)
424
529
 
425
530
  ```bash
@@ -593,6 +698,13 @@ when the skill it points at no longer exists; a pending entry whose skill is sti
593
698
  undone work, not an orphan. `routine` never installs, promotes, or rejects: no candidate enters a
594
699
  skills root without a human or agent decision, and a cron job is not that.
595
700
 
701
+ **Accepted findings (v0.12)** are honored the same way `audit` honors them: `--fail-on` grades
702
+ **active** findings only, so accepting every remaining finding via `skill-forge finding accept`
703
+ turns a red `--fail-on any` green, and any new or changed finding turns it red again. A
704
+ corrupt/missing accepted-findings store fails open (nothing gets suppressed) and surfaces as a
705
+ report **notice**, never as an "Incomplete step" — see
706
+ [docs/accepted-findings-schema.md](docs/accepted-findings-schema.md).
707
+
596
708
  ```
597
709
  0 9 * * 1 skill-forge routine --offline --fail-on high --housekeeping
598
710
  ```
@@ -720,9 +832,11 @@ guessing between the two).
720
832
 
721
833
  Safety runs the same built-in deny-pattern ruleset used for skills (curl\|bash, credential-file
722
834
  access, etc.) plus MCP-specific rules: inline credential values in config/env (quoted or
723
- unquoted), unpinned `npx -y` launch commands (a moving/dist tag like `@latest` counts as
724
- unpinned, and `npx` is recognized by basename so a full path or `npx.cmd` can't evade it),
725
- `--dangerously-*`/`--no-sandbox` flags, and filesystem-root launch args — see the
835
+ unquoted), unpinned `npx` launch commands — MEDIUM with `-y`/`--yes` (silent install), **LOW
836
+ without it (v0.12)**, since npx still prompts before the first install but the tag floats once
837
+ cached (a moving/dist tag like `@latest` counts as unpinned either way, and `npx` is recognized by
838
+ basename so a full path or `npx.cmd` can't evade it) — `--dangerously-*`/`--no-sandbox` flags, and
839
+ filesystem-root launch args — see the
726
840
  [MCP safety ruleset table](docs/gate-policy.md#mcp-safety-ruleset). Overlap analysis (Pro, free
727
841
  during the 0.x beta) ranks the candidate against the server entries already present in your
728
842
  configured `mcpTargets` files, instead of against a skills root.
@@ -806,6 +920,7 @@ JSON. Detection/overlap is agent-format-agnostic; writing is JSON-only.
806
920
  | `watch` — provenance drift check across every `SOURCES.md` ledger (v0.9) | | ✓ |
807
921
  | `ingest` + `queue close` — pending-queue drain handoff (v0.9) | | ✓ |
808
922
  | `refine` — capture/apply/generalize project-scope skill overrides (v0.10) | | ✓ |
923
+ | `finding accept`/`revoke`/`list` — acknowledge audit findings, never a gate input (v0.12) | ✓ | ✓ |
809
924
 
810
925
  Free is the complete safety gate on its own — quarantine, profile, safety scan, and an explicit
811
926
  promote/reject decision, with nothing held back. Pro is the curation layer on top: whether a new