@ngockhoale/ukit 1.6.8 → 2.0.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +43 -0
- package/manifests/platform.full.yaml +47 -0
- package/package.json +2 -1
- package/scripts/skill/audit-skill.mjs +39 -0
- package/src/cli/commands/doctor.js +22 -2
- package/src/cli/commands/memory.js +76 -1
- package/src/core/memory/store.js +125 -1
- package/src/core/skillProfile.js +45 -0
- package/src/skill/auditSkill.js +99 -0
- package/templates/.claude/agents/code-reviewer.md +51 -7
- package/templates/.claude/agents/handoff-planner.md +18 -2
- package/templates/.claude/hooks/context-hardcap-gate.sh +102 -0
- package/templates/.claude/hooks/reset-compact-pressure.sh +25 -0
- package/templates/.claude/settings.json +15 -0
- package/templates/.claude/skills/canvas-design/SKILL.md +2 -20
- package/templates/.claude/skills/canvas-design/philosophy-examples.md +23 -0
- package/templates/.claude/skills/debugging-toolkit/SKILL.md +2 -30
- package/templates/.claude/skills/debugging-toolkit/reference-tables.md +33 -0
- package/templates/.claude/skills/docs-manager/SKILL.md +7 -249
- package/templates/.claude/skills/docs-manager/conventions-and-examples.md +221 -0
- package/templates/.claude/skills/docx/SKILL.md +3 -34
- package/templates/.claude/skills/docx/redlining-reference.md +34 -0
- package/templates/.claude/skills/duraone/SKILL.md +12 -16
- package/templates/.claude/skills/executing-plans/SKILL.md +31 -19
- package/templates/.claude/skills/file-organizer/SKILL.md +2 -170
- package/templates/.claude/skills/file-organizer/examples-and-practices.md +173 -0
- package/templates/.claude/skills/pdf/SKILL.md +1 -62
- package/templates/.claude/skills/pdf/reference.md +65 -0
- package/templates/.claude/skills/pdf-processing-pro/SKILL.md +2 -73
- package/templates/.claude/skills/pdf-processing-pro/workflows-and-troubleshooting.md +80 -0
- package/templates/.claude/skills/pptx/SKILL.md +14 -286
- package/templates/.claude/skills/pptx/design-references.md +81 -0
- package/templates/.claude/skills/pptx/template-replacement-reference.md +150 -0
- package/templates/.claude/skills/pptx/utilities.md +62 -0
- package/templates/.claude/skills/project-learning/SKILL.md +32 -0
- package/templates/.claude/skills/root-cause-tracing/SKILL.md +2 -35
- package/templates/.claude/skills/root-cause-tracing/diagrams.md +44 -0
- package/templates/.claude/skills/sharing-skills/SKILL.md +1 -41
- package/templates/.claude/skills/sharing-skills/complete-example.md +41 -0
- package/templates/.claude/skills/skill-quality/SKILL.md +37 -0
- package/templates/.claude/skills/skill-quality/pressure-scenario-template.md +20 -0
- package/templates/.claude/skills/skill-quality/rationalization-table-template.md +15 -0
- package/templates/.claude/skills/skill-quality/trigger-accuracy-template.md +32 -0
- package/templates/.claude/skills/sql-optimization-patterns/SKILL.md +13 -440
- package/templates/.claude/skills/sql-optimization-patterns/references/advanced-techniques.md +128 -0
- package/templates/.claude/skills/sql-optimization-patterns/references/core-concepts.md +112 -0
- package/templates/.claude/skills/sql-optimization-patterns/references/query-patterns.md +204 -0
- package/templates/.claude/skills/subagent-driven-development/SKILL.md +4 -51
- package/templates/.claude/skills/subagent-driven-development/example-workflow.md +40 -0
- package/templates/.claude/skills/systematic-debugging/SKILL.md +2 -28
- package/templates/.claude/skills/systematic-debugging/reference-tables.md +33 -0
- package/templates/.claude/skills/test-driven-development/SKILL.md +2 -51
- package/templates/.claude/skills/test-driven-development/reference-tables.md +56 -0
- package/templates/.claude/skills/testing-anti-patterns/SKILL.md +1 -10
- package/templates/.claude/skills/testing-anti-patterns/reference-tables.md +14 -0
- package/templates/.claude/skills/verification-before-completion/SKILL.md +1 -31
- package/templates/.claude/skills/verification-before-completion/key-patterns.md +33 -0
- package/templates/.claude/ukit/runtime/compact-threshold.mjs +28 -0
- package/templates/.claude/ukit/runtime/reinject-context.mjs +14 -1
- package/templates/CLAUDE.md +4 -0
- package/templates/ukit/storage/config.json +4 -0
- package/src/core/memory/index.js +0 -2
- package/src/core/router/index.js +0 -2
- package/src/core/validation/index.js +0 -2
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
# PPTX Utility Commands
|
|
2
|
+
|
|
3
|
+
Reference commands for thumbnail generation and slide-to-image conversion. Not part of the core create/edit workflows.
|
|
4
|
+
|
|
5
|
+
## Creating Thumbnail Grids
|
|
6
|
+
|
|
7
|
+
To create visual thumbnail grids of PowerPoint slides for quick analysis and reference:
|
|
8
|
+
|
|
9
|
+
```bash
|
|
10
|
+
python scripts/thumbnail.py template.pptx [output_prefix]
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
**Features**:
|
|
14
|
+
- Creates: `thumbnails.jpg` (or `thumbnails-1.jpg`, `thumbnails-2.jpg`, etc. for large decks)
|
|
15
|
+
- Default: 5 columns, max 30 slides per grid (5×6)
|
|
16
|
+
- Custom prefix: `python scripts/thumbnail.py template.pptx my-grid`
|
|
17
|
+
- Note: The output prefix should include the path if you want output in a specific directory (e.g., `workspace/my-grid`)
|
|
18
|
+
- Adjust columns: `--cols 4` (range: 3-6, affects slides per grid)
|
|
19
|
+
- Grid limits: 3 cols = 12 slides/grid, 4 cols = 20, 5 cols = 30, 6 cols = 42
|
|
20
|
+
- Slides are zero-indexed (Slide 0, Slide 1, etc.)
|
|
21
|
+
|
|
22
|
+
**Use cases**:
|
|
23
|
+
- Template analysis: Quickly understand slide layouts and design patterns
|
|
24
|
+
- Content review: Visual overview of entire presentation
|
|
25
|
+
- Navigation reference: Find specific slides by their visual appearance
|
|
26
|
+
- Quality check: Verify all slides are properly formatted
|
|
27
|
+
|
|
28
|
+
**Examples**:
|
|
29
|
+
```bash
|
|
30
|
+
# Basic usage
|
|
31
|
+
python scripts/thumbnail.py presentation.pptx
|
|
32
|
+
|
|
33
|
+
# Combine options: custom name, columns
|
|
34
|
+
python scripts/thumbnail.py template.pptx analysis --cols 4
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
## Converting Slides to Images
|
|
38
|
+
|
|
39
|
+
To visually analyze PowerPoint slides, convert them to images using a two-step process:
|
|
40
|
+
|
|
41
|
+
1. **Convert PPTX to PDF**:
|
|
42
|
+
```bash
|
|
43
|
+
soffice --headless --convert-to pdf template.pptx
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
2. **Convert PDF pages to JPEG images**:
|
|
47
|
+
```bash
|
|
48
|
+
pdftoppm -jpeg -r 150 template.pdf slide
|
|
49
|
+
```
|
|
50
|
+
This creates files like `slide-1.jpg`, `slide-2.jpg`, etc.
|
|
51
|
+
|
|
52
|
+
Options:
|
|
53
|
+
- `-r 150`: Sets resolution to 150 DPI (adjust for quality/size balance)
|
|
54
|
+
- `-jpeg`: Output JPEG format (use `-png` for PNG if preferred)
|
|
55
|
+
- `-f N`: First page to convert (e.g., `-f 2` starts from page 2)
|
|
56
|
+
- `-l N`: Last page to convert (e.g., `-l 5` stops at page 5)
|
|
57
|
+
- `slide`: Prefix for output files
|
|
58
|
+
|
|
59
|
+
Example for specific range:
|
|
60
|
+
```bash
|
|
61
|
+
pdftoppm -jpeg -r 150 -f 2 -l 5 template.pdf slide # Converts only pages 2-5
|
|
62
|
+
```
|
|
@@ -0,0 +1,32 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: project-learning
|
|
3
|
+
description: Use after completing non-trivial work in a project, when a repeated coding pattern (naming, structure, framework usage) was observed across multiple files or multiple sessions, to propose it as a project convention for human approval. Never auto-applies a convention.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Project Learning
|
|
7
|
+
|
|
8
|
+
Detecting a project's own conventions is a judgment call, not something a deterministic script can do — it needs a live agent reading real code. This skill turns that judgment into a proposal, never a silent change.
|
|
9
|
+
|
|
10
|
+
## When to use
|
|
11
|
+
|
|
12
|
+
After finishing a task, if either is true:
|
|
13
|
+
|
|
14
|
+
- The same non-obvious pattern was seen in 2+ places in existing code you read (e.g. every form field uses `ControlInput`, never a raw `<input>`).
|
|
15
|
+
- The human corrected your output in a way that would recur (e.g. you wrote `form`, they renamed it to `workingObj` — this is exactly Example 3 in the maintainer's own global `CLAUDE.md` Surgical Changes section).
|
|
16
|
+
|
|
17
|
+
One-off patterns are not conventions. Skip this for anything seen only once.
|
|
18
|
+
|
|
19
|
+
## What NOT to do
|
|
20
|
+
|
|
21
|
+
Never write directly into `conventions`, and never mention this to the human as if it's already decided. Always propose through Task 10's storage path — call `proposePatternCandidate(projectRoot, projectId, candidate)` from `src/core/memory/store.js` (or, from a plain conversation, tell the human to check `ukit memory list --pending`). Phrase the candidate as one falsifiable sentence:
|
|
22
|
+
|
|
23
|
+
- Good: "Use `ControlInput` from `src/components/` instead of raw `<input>` for form fields."
|
|
24
|
+
- Bad: "This project seems to like custom inputs."
|
|
25
|
+
|
|
26
|
+
## Human approval loop
|
|
27
|
+
|
|
28
|
+
The human reviews pending candidates with `ukit memory list --pending` and decides with `ukit memory approve <id>` or `ukit memory reject <id>`. Until approved, a candidate has zero effect on any future session — no session reads a `pending` candidate as if it were an established convention.
|
|
29
|
+
|
|
30
|
+
## Scope
|
|
31
|
+
|
|
32
|
+
Project-local only for v2.0. Do not write to `user.preferences`/`user.rules` (global, cross-project) — that scope is deferred to a later plan.
|
|
@@ -13,20 +13,7 @@ Bugs often manifest deep in the call stack (git init in wrong directory, file cr
|
|
|
13
13
|
|
|
14
14
|
## When to Use
|
|
15
15
|
|
|
16
|
-
|
|
17
|
-
digraph when_to_use {
|
|
18
|
-
"Bug appears deep in stack?" [shape=diamond];
|
|
19
|
-
"Can trace backwards?" [shape=diamond];
|
|
20
|
-
"Fix at symptom point" [shape=box];
|
|
21
|
-
"Trace to original trigger" [shape=box];
|
|
22
|
-
"BETTER: Also add defense-in-depth" [shape=box];
|
|
23
|
-
|
|
24
|
-
"Bug appears deep in stack?" -> "Can trace backwards?" [label="yes"];
|
|
25
|
-
"Can trace backwards?" -> "Trace to original trigger" [label="yes"];
|
|
26
|
-
"Can trace backwards?" -> "Fix at symptom point" [label="no - dead end"];
|
|
27
|
-
"Trace to original trigger" -> "BETTER: Also add defense-in-depth";
|
|
28
|
-
}
|
|
29
|
-
```
|
|
16
|
+
Flow diagram: [`diagrams.md`](diagrams.md#when-to-use).
|
|
30
17
|
|
|
31
18
|
**Use when:**
|
|
32
19
|
- Error happens deep in execution (not at entry point)
|
|
@@ -134,27 +121,7 @@ Runs tests one-by-one, stops at first polluter. See script for usage.
|
|
|
134
121
|
|
|
135
122
|
## Key Principle
|
|
136
123
|
|
|
137
|
-
|
|
138
|
-
digraph principle {
|
|
139
|
-
"Found immediate cause" [shape=ellipse];
|
|
140
|
-
"Can trace one level up?" [shape=diamond];
|
|
141
|
-
"Trace backwards" [shape=box];
|
|
142
|
-
"Is this the source?" [shape=diamond];
|
|
143
|
-
"Fix at source" [shape=box];
|
|
144
|
-
"Add validation at each layer" [shape=box];
|
|
145
|
-
"Bug impossible" [shape=doublecircle];
|
|
146
|
-
"NEVER fix just the symptom" [shape=octagon, style=filled, fillcolor=red, fontcolor=white];
|
|
147
|
-
|
|
148
|
-
"Found immediate cause" -> "Can trace one level up?";
|
|
149
|
-
"Can trace one level up?" -> "Trace backwards" [label="yes"];
|
|
150
|
-
"Can trace one level up?" -> "NEVER fix just the symptom" [label="no"];
|
|
151
|
-
"Trace backwards" -> "Is this the source?";
|
|
152
|
-
"Is this the source?" -> "Trace backwards" [label="no - keeps going"];
|
|
153
|
-
"Is this the source?" -> "Fix at source" [label="yes"];
|
|
154
|
-
"Fix at source" -> "Add validation at each layer";
|
|
155
|
-
"Add validation at each layer" -> "Bug impossible";
|
|
156
|
-
}
|
|
157
|
-
```
|
|
124
|
+
Flow diagram: [`diagrams.md`](diagrams.md#key-principle).
|
|
158
125
|
|
|
159
126
|
**NEVER fix just where the error appears.** Trace back to find the original trigger.
|
|
160
127
|
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# Root Cause Tracing — Flow Diagrams
|
|
2
|
+
|
|
3
|
+
Graphviz restatements of the decision flow in `SKILL.md`. The prose in `SKILL.md` is the source of truth; these are visual aids for the same logic.
|
|
4
|
+
|
|
5
|
+
## When to Use
|
|
6
|
+
|
|
7
|
+
```dot
|
|
8
|
+
digraph when_to_use {
|
|
9
|
+
"Bug appears deep in stack?" [shape=diamond];
|
|
10
|
+
"Can trace backwards?" [shape=diamond];
|
|
11
|
+
"Fix at symptom point" [shape=box];
|
|
12
|
+
"Trace to original trigger" [shape=box];
|
|
13
|
+
"BETTER: Also add defense-in-depth" [shape=box];
|
|
14
|
+
|
|
15
|
+
"Bug appears deep in stack?" -> "Can trace backwards?" [label="yes"];
|
|
16
|
+
"Can trace backwards?" -> "Trace to original trigger" [label="yes"];
|
|
17
|
+
"Can trace backwards?" -> "Fix at symptom point" [label="no - dead end"];
|
|
18
|
+
"Trace to original trigger" -> "BETTER: Also add defense-in-depth";
|
|
19
|
+
}
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
## Key Principle
|
|
23
|
+
|
|
24
|
+
```dot
|
|
25
|
+
digraph principle {
|
|
26
|
+
"Found immediate cause" [shape=ellipse];
|
|
27
|
+
"Can trace one level up?" [shape=diamond];
|
|
28
|
+
"Trace backwards" [shape=box];
|
|
29
|
+
"Is this the source?" [shape=diamond];
|
|
30
|
+
"Fix at source" [shape=box];
|
|
31
|
+
"Add validation at each layer" [shape=box];
|
|
32
|
+
"Bug impossible" [shape=doublecircle];
|
|
33
|
+
"NEVER fix just the symptom" [shape=octagon, style=filled, fillcolor=red, fontcolor=white];
|
|
34
|
+
|
|
35
|
+
"Found immediate cause" -> "Can trace one level up?";
|
|
36
|
+
"Can trace one level up?" -> "Trace backwards" [label="yes"];
|
|
37
|
+
"Can trace one level up?" -> "NEVER fix just the symptom" [label="no"];
|
|
38
|
+
"Trace backwards" -> "Is this the source?";
|
|
39
|
+
"Is this the source?" -> "Trace backwards" [label="no - keeps going"];
|
|
40
|
+
"Is this the source?" -> "Fix at source" [label="yes"];
|
|
41
|
+
"Fix at source" -> "Add validation at each layer";
|
|
42
|
+
"Add validation at each layer" -> "Bug impossible";
|
|
43
|
+
}
|
|
44
|
+
```
|
|
@@ -99,47 +99,7 @@ EOF
|
|
|
99
99
|
)"
|
|
100
100
|
```
|
|
101
101
|
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
Here's a complete example of sharing a skill called "async-patterns":
|
|
105
|
-
|
|
106
|
-
```bash
|
|
107
|
-
# 1. Sync with upstream
|
|
108
|
-
cd ~/.config/superpowers/skills/
|
|
109
|
-
git checkout main
|
|
110
|
-
git pull upstream main
|
|
111
|
-
git push origin main
|
|
112
|
-
|
|
113
|
-
# 2. Create branch
|
|
114
|
-
git checkout -b "add-async-patterns-skill"
|
|
115
|
-
|
|
116
|
-
# 3. Create/edit the skill
|
|
117
|
-
# (Work on skills/async-patterns/SKILL.md)
|
|
118
|
-
|
|
119
|
-
# 4. Commit
|
|
120
|
-
git add skills/async-patterns/
|
|
121
|
-
git commit -m "Add async-patterns skill
|
|
122
|
-
|
|
123
|
-
Patterns for handling asynchronous operations in tests and application code.
|
|
124
|
-
|
|
125
|
-
Tested with: Multiple pressure scenarios testing agent compliance."
|
|
126
|
-
|
|
127
|
-
# 5. Push
|
|
128
|
-
git push -u origin "add-async-patterns-skill"
|
|
129
|
-
|
|
130
|
-
# 6. Create PR
|
|
131
|
-
gh pr create \
|
|
132
|
-
--repo upstream-org/upstream-repo \
|
|
133
|
-
--title "Add async-patterns skill" \
|
|
134
|
-
--body "## Summary
|
|
135
|
-
Patterns for handling asynchronous operations correctly in tests and application code.
|
|
136
|
-
|
|
137
|
-
## Testing
|
|
138
|
-
Tested with multiple application scenarios. Agents successfully apply patterns to new code.
|
|
139
|
-
|
|
140
|
-
## Context
|
|
141
|
-
Addresses common async pitfalls like race conditions, improper error handling, and timing issues."
|
|
142
|
-
```
|
|
102
|
+
Full worked example (async-patterns skill, all 6 steps filled in): [`complete-example.md`](complete-example.md).
|
|
143
103
|
|
|
144
104
|
## After PR is Merged
|
|
145
105
|
|
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
# Sharing Skills — Complete Example
|
|
2
|
+
|
|
3
|
+
Full walkthrough of the "Sharing Workflow" steps in `SKILL.md`, filled in for a skill called "async-patterns".
|
|
4
|
+
|
|
5
|
+
```bash
|
|
6
|
+
# 1. Sync with upstream
|
|
7
|
+
cd ~/.config/superpowers/skills/
|
|
8
|
+
git checkout main
|
|
9
|
+
git pull upstream main
|
|
10
|
+
git push origin main
|
|
11
|
+
|
|
12
|
+
# 2. Create branch
|
|
13
|
+
git checkout -b "add-async-patterns-skill"
|
|
14
|
+
|
|
15
|
+
# 3. Create/edit the skill
|
|
16
|
+
# (Work on skills/async-patterns/SKILL.md)
|
|
17
|
+
|
|
18
|
+
# 4. Commit
|
|
19
|
+
git add skills/async-patterns/
|
|
20
|
+
git commit -m "Add async-patterns skill
|
|
21
|
+
|
|
22
|
+
Patterns for handling asynchronous operations in tests and application code.
|
|
23
|
+
|
|
24
|
+
Tested with: Multiple pressure scenarios testing agent compliance."
|
|
25
|
+
|
|
26
|
+
# 5. Push
|
|
27
|
+
git push -u origin "add-async-patterns-skill"
|
|
28
|
+
|
|
29
|
+
# 6. Create PR
|
|
30
|
+
gh pr create \
|
|
31
|
+
--repo upstream-org/upstream-repo \
|
|
32
|
+
--title "Add async-patterns skill" \
|
|
33
|
+
--body "## Summary
|
|
34
|
+
Patterns for handling asynchronous operations correctly in tests and application code.
|
|
35
|
+
|
|
36
|
+
## Testing
|
|
37
|
+
Tested with multiple application scenarios. Agents successfully apply patterns to new code.
|
|
38
|
+
|
|
39
|
+
## Context
|
|
40
|
+
Addresses common async pitfalls like race conditions, improper error handling, and timing issues."
|
|
41
|
+
```
|
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: skill-quality
|
|
3
|
+
description: Use when creating or editing a UKit template skill/agent (templates/.claude/skills or templates/.claude/agents), before shipping the change, to pressure-test compliance and trigger accuracy. Maintainer-only.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Skill Quality
|
|
7
|
+
|
|
8
|
+
Maintainer-only. Pressure-tests a skill/agent before shipping a change, the same way a unit test pressure-tests code. UKit cannot run this itself — `scripts/skill/audit-skill.mjs` only scaffolds files and records results, it never calls a model. The actual test runs inside a live Claude Code (or Codex) session, by dispatching a subagent (via `Agent`/`Task`) to follow the pressure scenario.
|
|
9
|
+
|
|
10
|
+
## When to use
|
|
11
|
+
|
|
12
|
+
Editing or creating a **discipline-enforcing** skill/agent — one with rules meant to hold under pressure (e.g. `systematic-debugging`, `verification-before-completion`, `executing-plans`). Skip this for pure-reference skills (`pdf`, `docx`) that have no compliance rule to break.
|
|
13
|
+
|
|
14
|
+
## The loop: RED → GREEN → REFACTOR
|
|
15
|
+
|
|
16
|
+
1. **RED** — dispatch a subagent with the pressure scenario but *without* the skill loaded. Record its choice and the exact rationalization it used, verbatim.
|
|
17
|
+
2. **GREEN** — load the skill (new or patched), re-run the same scenario. Expect compliance. If it still fails, the skill's language is too weak — patch and re-run before moving on.
|
|
18
|
+
3. **REFACTOR** — if a *new* rationalization surfaces later (different scenario, same skill), don't bolt a soft exception onto the existing rule. Add an explicit counter, a row in `rationalization-table-template.md`, and a red-flag entry, then re-test the original scenario too — a new exception must not silently loosen the old one.
|
|
19
|
+
|
|
20
|
+
## Output format
|
|
21
|
+
|
|
22
|
+
```
|
|
23
|
+
Compliance: NN%
|
|
24
|
+
Weakness found: <one line>
|
|
25
|
+
Rationalization detected: "<verbatim quote>"
|
|
26
|
+
Recommended patch: <one line>
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
## Trigger accuracy check
|
|
30
|
+
|
|
31
|
+
Separate from compliance: does the skill's `description` frontmatter actually get it loaded at the right time? Dispatch N subagents with prompts that *should* trigger the skill and M that *should not*, giving them only the frontmatter `description` — the same information a real session uses to decide whether to load a skill. Fixture and report format: `trigger-accuracy-template.md`.
|
|
32
|
+
|
|
33
|
+
## Templates
|
|
34
|
+
|
|
35
|
+
- `pressure-scenario-template.md` — how to write a scenario, with a real worked example already run against `systematic-debugging`.
|
|
36
|
+
- `rationalization-table-template.md` — the closing-the-loophole table + the 3 hardening techniques.
|
|
37
|
+
- `trigger-accuracy-template.md` — fixture + report format for the check above.
|
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
# Pressure Scenario Template
|
|
2
|
+
|
|
3
|
+
## How to write a scenario
|
|
4
|
+
|
|
5
|
+
- Combine **3+ pressure types** from: time, sunk-cost, authority, economic, exhaustion, social.
|
|
6
|
+
- Force a concrete **A/B/C choice**, not an open question. "What do you do?" — not "what should you do?"
|
|
7
|
+
- Use real file paths and real numbers, not abstractions.
|
|
8
|
+
- Frame it as real: "This is a real scenario. You must choose and act. Don't ask hypothetical questions."
|
|
9
|
+
- End with: "Be honest about what you would actually do."
|
|
10
|
+
|
|
11
|
+
## Worked example — don't copy it here, read the original
|
|
12
|
+
|
|
13
|
+
Rather than duplicating a worked scenario in this file, reuse the one already written and already run in this repo: `templates/.claude/skills/systematic-debugging/test-pressure-1.md` (economic + authority + time + sunk-cost pressure, A/B/C choice, passed against `systematic-debugging`). Open it directly, swap the skill name and domain details for the skill you're testing. See its sibling `CREATION-LOG.md` → "Bulletproofing Elements" for the language choices (ALWAYS/NEVER, explicit "even if faster") that made the original pass.
|
|
14
|
+
|
|
15
|
+
## Producing your own scenario (when the worked example doesn't fit the skill's domain)
|
|
16
|
+
|
|
17
|
+
1. Identify the rule the skill is meant to enforce and the shortcut an agent would be tempted to take instead.
|
|
18
|
+
2. Pick 3+ pressure types that would plausibly justify the shortcut in the moment.
|
|
19
|
+
3. Write the A/B/C choice so the "wrong" answer is the one that feels most reasonable under pressure — that's the point of the test.
|
|
20
|
+
4. Record the subagent's raw choice and its exact reasoning, not a paraphrase.
|
|
@@ -0,0 +1,15 @@
|
|
|
1
|
+
# Rationalization Table Template
|
|
2
|
+
|
|
3
|
+
Close a found rationalization with all three, not just one:
|
|
4
|
+
|
|
5
|
+
1. **Explicit negation** — state the subagent's exact justification, then negate it directly in the skill's text. ("The deadline changes things" → "the deadline does not change this," not a vague "be careful.")
|
|
6
|
+
2. **Table row** — append below. Never delete a row; a closed loophole can reopen if wording drifts later.
|
|
7
|
+
3. **Red-flags entry** — add the excuse's *pattern* (not just its wording) to the skill's own "if you're thinking X, that's the shortcut talking" list, so a reworded version is still caught.
|
|
8
|
+
|
|
9
|
+
| Excuse | Reality |
|
|
10
|
+
|---|---|
|
|
11
|
+
| *(verbatim rationalization found in RED)* | *(one-line counter-fact)* |
|
|
12
|
+
|
|
13
|
+
## Gotcha
|
|
14
|
+
|
|
15
|
+
Don't patch a rationalization with a soft "nuance clause" (e.g. "...unless it's genuinely urgent") appended to an otherwise-working rule — that clause becomes the next rationalization. A real exception must be its own conditional keyed to an observable predicate (a specific file pattern, a specific error type), never a general escape hatch.
|
|
@@ -0,0 +1,32 @@
|
|
|
1
|
+
# Trigger Accuracy Template
|
|
2
|
+
|
|
3
|
+
Tests whether a skill's frontmatter `description` loads it at the right time — not whether its body is correct once loaded.
|
|
4
|
+
|
|
5
|
+
## Fixture format
|
|
6
|
+
|
|
7
|
+
```yaml
|
|
8
|
+
skill: <skill-name>
|
|
9
|
+
should-trigger:
|
|
10
|
+
- "prompt that clearly matches this skill's purpose"
|
|
11
|
+
- "a differently-worded prompt that should still match"
|
|
12
|
+
should-not-trigger:
|
|
13
|
+
- "a prompt from an adjacent but different skill's domain"
|
|
14
|
+
- "a prompt that mentions a keyword from the description out of context"
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
At least 3 prompts per list. `should-not-trigger` prompts should be near-misses (shared keyword, wrong intent) — obviously unrelated prompts prove nothing.
|
|
18
|
+
|
|
19
|
+
## Running it
|
|
20
|
+
|
|
21
|
+
Per prompt, dispatch a subagent given **only** the frontmatter `description` and ask: "would you load this skill for this request? yes/no" — mirrors how a real session decides.
|
|
22
|
+
|
|
23
|
+
## Report format
|
|
24
|
+
|
|
25
|
+
```
|
|
26
|
+
Should trigger: x/N
|
|
27
|
+
Should not trigger: y/M
|
|
28
|
+
False positives: <should-not-trigger prompt that fired anyway, and why>
|
|
29
|
+
False negatives: <should-trigger prompt that didn't fire, and why>
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
A false negative usually means the `description` is too narrow or uses jargon the prompt wouldn't contain. A false positive usually means it's too broad or shares vocabulary with an unrelated domain — tighten with a more specific verb/noun pairing rather than adding exclusions.
|