@uinaf/skillcheck 0.2.0 → 0.2.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +9 -8
- package/docs/authoring.md +70 -0
- package/docs/scenarios.md +3 -0
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -49,14 +49,15 @@ CLI it documents.
|
|
|
49
49
|
|
|
50
50
|
## docs
|
|
51
51
|
|
|
52
|
-
| doc
|
|
53
|
-
|
|
|
54
|
-
| [usage](docs/usage.md)
|
|
55
|
-
| [scenarios](docs/scenarios.md)
|
|
56
|
-
| [
|
|
57
|
-
| [
|
|
58
|
-
| [
|
|
59
|
-
| [
|
|
52
|
+
| doc | when |
|
|
53
|
+
| --------------------------------------------------------------- | ------------------------------------- |
|
|
54
|
+
| [usage](docs/usage.md) | every subcommand, flag, and auth path |
|
|
55
|
+
| [scenarios](docs/scenarios.md) | writing an eval scenario |
|
|
56
|
+
| [authoring](docs/authoring.md) | writing and auditing the skill itself |
|
|
57
|
+
| [adoption](docs/adoption.md) | wiring the lint into another repo |
|
|
58
|
+
| [releasing](docs/releasing.md) | the npm pipeline |
|
|
59
|
+
| [contributing](CONTRIBUTING.md) | local setup and the verify gate |
|
|
60
|
+
| [security](https://github.com/uinaf/skillcheck/security/policy) | reporting a vulnerability |
|
|
60
61
|
|
|
61
62
|
## license
|
|
62
63
|
|
|
@@ -0,0 +1,70 @@
|
|
|
1
|
+
# authoring and auditing skills
|
|
2
|
+
|
|
3
|
+
what `lint` and evals cannot judge: whether a skill is worth routing to and
|
|
4
|
+
cheap to load. use this when writing a skill or auditing one. evidence beats
|
|
5
|
+
stylistic preference; run `skillcheck lint` first and let this cover the rest.
|
|
6
|
+
|
|
7
|
+
## metadata and discovery
|
|
8
|
+
|
|
9
|
+
- `name` is concrete and easy to say out loud. `helper`, `tools`, `utils` are
|
|
10
|
+
discovery smells.
|
|
11
|
+
- `description` is third person and says both what the skill does and when to
|
|
12
|
+
use it. it is an always-loaded retrieval pointer: front-load the concrete
|
|
13
|
+
action or domain that should activate it.
|
|
14
|
+
- one trigger per materially distinct request branch. collapse synonyms that
|
|
15
|
+
rename the same branch.
|
|
16
|
+
- state the main overlap boundary without naming another skill.
|
|
17
|
+
|
|
18
|
+
## body shape
|
|
19
|
+
|
|
20
|
+
- keep `SKILL.md` on workflow, principles, boundaries, and routing. lead with
|
|
21
|
+
the task, not a bibliography.
|
|
22
|
+
- assume the model is smart; spend tokens on repo-specific judgment. delete any
|
|
23
|
+
instruction that would not change a capable model's behavior.
|
|
24
|
+
- match freedom to risk: high for contextual judgment, medium when a preferred
|
|
25
|
+
pattern exists, low for fragile operations.
|
|
26
|
+
- say what evidence to gather and what a complete result includes. end each
|
|
27
|
+
step with an observable completion condition, not "understood" or "handled".
|
|
28
|
+
|
|
29
|
+
## progressive disclosure
|
|
30
|
+
|
|
31
|
+
- durable detail, rubrics, and long examples go in `references/`, one hop from
|
|
32
|
+
`SKILL.md`, each with a task-shaped retrieval job. material every path needs
|
|
33
|
+
stays inline.
|
|
34
|
+
- for repeated deterministic work, route to the target's existing framework,
|
|
35
|
+
schema, task graph, or library; otherwise add a tested module in the
|
|
36
|
+
project's primary language, not ad-hoc shell rendered as prose.
|
|
37
|
+
- when executable code belongs to another maintained project, link the exact
|
|
38
|
+
public artifact and state the contract it demonstrates; do not fork it into
|
|
39
|
+
prose.
|
|
40
|
+
- a package stays independently usable: state prerequisites and out-of-scope
|
|
41
|
+
next steps locally. never invoke, import, or assume a sibling skill.
|
|
42
|
+
|
|
43
|
+
## audit
|
|
44
|
+
|
|
45
|
+
grade each dimension strong, mixed, or weak:
|
|
46
|
+
|
|
47
|
+
| dimension | question |
|
|
48
|
+
| ---------------------- | --------------------------------------------------------------------- |
|
|
49
|
+
| discovery | does metadata alone route a realistic request here |
|
|
50
|
+
| workflow | does the body say how to begin, what evidence to gather, when to stop |
|
|
51
|
+
| progressive disclosure | is detail in the right file |
|
|
52
|
+
| repo fit | are links, commands, and conventions current |
|
|
53
|
+
| verification | is the strongest mechanical check named, plus a real evidence loop |
|
|
54
|
+
| boundaries | are limits and next steps stated without leaning on a sibling skill |
|
|
55
|
+
|
|
56
|
+
blockers, must-fix: invalid frontmatter; a description that fails discovery;
|
|
57
|
+
stale commands, paths, or links; a workflow with no start, evidence loop, or
|
|
58
|
+
completion; conflicts with the repo's guidance; sibling-skill dependencies.
|
|
59
|
+
|
|
60
|
+
major findings: vague name; synonym-stuffed description; bloated `SKILL.md`;
|
|
61
|
+
missing or muddy boundaries; prose re-inventing a deterministic tool; abstract
|
|
62
|
+
examples.
|
|
63
|
+
|
|
64
|
+
## improve
|
|
65
|
+
|
|
66
|
+
fix blockers first, then the highest-leverage majors. prefer the smallest
|
|
67
|
+
change that improves activation, decision quality, or proof. when pruning,
|
|
68
|
+
measure common-path context for representative requests; line count alone does
|
|
69
|
+
not reveal retrieval cost. after edits, rerun `skillcheck lint` and the repo's
|
|
70
|
+
gate, and rerun evals when behavior was the thing changed.
|
package/docs/scenarios.md
CHANGED
|
@@ -84,3 +84,6 @@ per run, under `<root>/.skillcheck/scratch/<name>/`, rebuilt from scratch each
|
|
|
84
84
|
time. the skill under test is installed at `.claude/skills/<skill>/` (and also
|
|
85
85
|
`.agents/skills/<skill>/` on the codex harness) with its `evals/` directory
|
|
86
86
|
excluded, so criteria never leak into the agent's context.
|
|
87
|
+
|
|
88
|
+
scenario quality is behavioral proof; [authoring](authoring.md) covers the
|
|
89
|
+
judgment layer lint and evals cannot grade.
|