@meyverick/agentic 5.0.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +234 -0
- package/CHANGELOG.md +236 -0
- package/README.md +50 -0
- package/install.ts +349 -0
- package/package.json +37 -0
- package/scripts/check-deps.mjs +587 -0
- package/scripts/git-dl.mjs +100 -0
- package/skills/check/SKILL.md +108 -0
- package/skills/check/evals/benchmark.json +40 -0
- package/skills/check/evals/evals.json +38 -0
- package/skills/check/references/diagnostic-matrix.md +170 -0
- package/skills/check/references/script-anatomy.md +154 -0
- package/skills/create-skill/SKILL.md +291 -0
- package/skills/create-skill/assets/templates/SKILL.md.template +118 -0
- package/skills/create-skill/assets/templates/evals.json.template +36 -0
- package/skills/create-skill/assets/templates/grading.json.template +26 -0
- package/skills/create-skill/evals/benchmark.json +41 -0
- package/skills/create-skill/evals/evals.json +50 -0
- package/skills/create-skill/evals/grading-template.json +36 -0
- package/skills/create-skill/evals/near-misses.json +35 -0
- package/skills/create-skill/evals/trigger-queries.json +80 -0
- package/skills/create-skill/references/antipatterns.md +123 -0
- package/skills/create-skill/references/component-decomposition.md +130 -0
- package/skills/create-skill/references/content-quality-criteria.md +61 -0
- package/skills/create-skill/references/description-optimization.md +90 -0
- package/skills/create-skill/references/eval-methodology.md +100 -0
- package/skills/create-skill/references/fragility-matching.md +88 -0
- package/skills/create-skill/references/gotchas-patterns.md +80 -0
- package/skills/create-skill/references/specification.md +77 -0
- package/skills/create-skill/scripts/audit-antipatterns.mjs +164 -0
- package/skills/create-skill/scripts/compute-benchmark.mjs +111 -0
- package/skills/create-skill/scripts/run-cold-eval.mjs +118 -0
- package/skills/create-skill/scripts/scaffold-skill.mjs +86 -0
- package/skills/create-skill/scripts/validate-routing.mjs +137 -0
- package/skills/create-skill/scripts/validate-structure.mjs +223 -0
- package/skills/design-craft/SKILL.md +134 -0
- package/skills/design-craft/evals/benchmark.json +41 -0
- package/skills/design-craft/evals/evals.json +81 -0
- package/skills/design-craft/references/anti-slop-patterns.md +49 -0
- package/skills/design-craft/references/art-direction.md +89 -0
- package/skills/design-craft/references/design-engineering.md +122 -0
- package/skills/design-craft/references/motion-craft.md +124 -0
- package/skills/design-craft/references/process.md +47 -0
- package/skills/design-craft/references/review-checklist.md +121 -0
- package/skills/guardrails/SKILL.md +118 -0
- package/skills/guardrails/evals/benchmark.json +40 -0
- package/skills/guardrails/evals/evals.json +49 -0
- package/skills/guardrails/references/guardrails-patterns.md +43 -0
- package/skills/okf-docs/SKILL.md +79 -0
- package/skills/okf-docs/evals/benchmark.json +21 -0
- package/skills/okf-docs/evals/evals.json +37 -0
- package/skills/okf-docs/references/okf-spec.md +56 -0
- package/skills/okf-docs/scripts/validate-frontmatter.mjs +130 -0
- package/skills/openspec-harden/SKILL.md +138 -0
- package/skills/openspec-harden/evals/benchmark.json +40 -0
- package/skills/openspec-harden/evals/evals.json +38 -0
- package/skills/openspec-learn/SKILL.md +216 -0
- package/skills/openspec-learn/evals/benchmark.json +44 -0
- package/skills/openspec-learn/evals/evals.json +48 -0
- package/skills/openspec-learn/evals/retrieval-bench.json +27 -0
- package/skills/openspec-learn/references/conflict-handling.md +20 -0
- package/skills/openspec-learn/references/evaluation-methodology.md +126 -0
- package/skills/openspec-learn/references/examples.md +37 -0
- package/skills/openspec-learn/references/improvement-patterns.md +155 -0
- package/skills/openspec-learn/references/report-analysis.md +104 -0
- package/skills/openspec-learn/references/skill-quality.md +103 -0
- package/skills/openspec-learn/references/tool-type-detection.md +30 -0
- package/skills/openspec-report/SKILL.md +104 -0
- package/skills/openspec-report/assets/templates/assessment.md.template +84 -0
- package/skills/openspec-report/assets/templates/report.md.template +92 -0
- package/skills/openspec-report/evals/benchmark.json +44 -0
- package/skills/openspec-report/evals/evals.json +46 -0
- package/skills/qmd-research/SKILL.md +89 -0
- package/skills/qmd-research/evals/benchmark.json +40 -0
- package/skills/qmd-research/evals/evals.json +38 -0
- package/skills/qmd-research/references/index-management.md +69 -0
- package/skills/qmd-research/references/query-craft.md +82 -0
|
@@ -0,0 +1,155 @@
|
|
|
1
|
+
# Improvement Patterns
|
|
2
|
+
|
|
3
|
+
Common improvement types and templates.
|
|
4
|
+
|
|
5
|
+
## Pattern 1: Fix Validation Failures
|
|
6
|
+
|
|
7
|
+
**Trigger**: Report shows test cases failed
|
|
8
|
+
|
|
9
|
+
**Template**:
|
|
10
|
+
```
|
|
11
|
+
1. Identify failed test case
|
|
12
|
+
2. Analyze failure reason
|
|
13
|
+
3. Update SKILL.md to address failure
|
|
14
|
+
4. Re-run validation
|
|
15
|
+
5. If passes → update changelog (PATCH)
|
|
16
|
+
6. If still fails → try alternative approach
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
**Example**:
|
|
20
|
+
- Failed: "Description doesn't trigger on casual prompts"
|
|
21
|
+
- Fix: Broaden description scope
|
|
22
|
+
- Changelog: `### Fixed - Description trigger accuracy`
|
|
23
|
+
|
|
24
|
+
## Pattern 2: Address Follow-ups
|
|
25
|
+
|
|
26
|
+
**Trigger**: Report has follow-up items
|
|
27
|
+
|
|
28
|
+
**Template**:
|
|
29
|
+
```
|
|
30
|
+
1. List all follow-up items
|
|
31
|
+
2. Prioritize by impact/effort
|
|
32
|
+
3. Implement highest priority
|
|
33
|
+
4. Update changelog (MINOR for new features)
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
**Example**:
|
|
37
|
+
- Follow-up: "Add gotcha for edge case X"
|
|
38
|
+
- Fix: Add to gotchas section
|
|
39
|
+
- Changelog: `### Added - Gotcha for edge case X`
|
|
40
|
+
|
|
41
|
+
## Pattern 3: Upgrade Based on Patterns
|
|
42
|
+
|
|
43
|
+
**Trigger**: Report shows new best practices
|
|
44
|
+
|
|
45
|
+
**Template**:
|
|
46
|
+
```
|
|
47
|
+
1. Identify new pattern
|
|
48
|
+
2. Check if skill already follows it
|
|
49
|
+
3. If not, update skill to follow pattern
|
|
50
|
+
4. Update references if needed
|
|
51
|
+
5. Update changelog (MINOR)
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
**Example**:
|
|
55
|
+
- Pattern: "All skills should include fragility matching"
|
|
56
|
+
- Fix: Add fragility section to SKILL.md
|
|
57
|
+
- Changelog: `### Added - Fragility matching section`
|
|
58
|
+
|
|
59
|
+
## Pattern 4: Meta-Improvement
|
|
60
|
+
|
|
61
|
+
**Trigger**: Pattern is generalizable to all skills
|
|
62
|
+
|
|
63
|
+
**Template**:
|
|
64
|
+
```
|
|
65
|
+
1. Identify generalizable pattern
|
|
66
|
+
2. Check if create-skill already covers it
|
|
67
|
+
3. If not, update create-skill:
|
|
68
|
+
- SKILL.md workflow
|
|
69
|
+
- scripts/
|
|
70
|
+
- references/
|
|
71
|
+
- assets/templates/
|
|
72
|
+
4. Validate create-skill
|
|
73
|
+
5. Update create-skill changelog (MINOR)
|
|
74
|
+
6. Optionally propagate to existing skills
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
**Example**:
|
|
78
|
+
- Pattern: "All skills should validate near-miss negatives"
|
|
79
|
+
- Fix: Add to create-skill evaluation phase
|
|
80
|
+
- Changelog: `### Added - Near-miss negative evaluation`
|
|
81
|
+
|
|
82
|
+
## Pattern 5: Description Optimization
|
|
83
|
+
|
|
84
|
+
**Trigger**: Report shows description doesn't trigger well
|
|
85
|
+
|
|
86
|
+
**Template**:
|
|
87
|
+
```
|
|
88
|
+
1. Analyze trigger test results
|
|
89
|
+
2. Identify failing queries
|
|
90
|
+
3. Revise description:
|
|
91
|
+
- Broaden scope for should-trigger failures
|
|
92
|
+
- Add specificity for shouldn't-trigger failures
|
|
93
|
+
4. Re-test triggers
|
|
94
|
+
5. Update changelog (PATCH)
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
**Example**:
|
|
98
|
+
- Failing: "Create a skill for X" doesn't trigger
|
|
99
|
+
- Fix: Add "create" to description keywords
|
|
100
|
+
- Changelog: `### Fixed - Description trigger keywords`
|
|
101
|
+
|
|
102
|
+
## Pattern 6: Antipattern Fix
|
|
103
|
+
|
|
104
|
+
**Trigger**: Audit detects antipatterns
|
|
105
|
+
|
|
106
|
+
**Template**:
|
|
107
|
+
```
|
|
108
|
+
1. Identify antipattern violation
|
|
109
|
+
2. Analyze why it occurred
|
|
110
|
+
3. Fix the violation:
|
|
111
|
+
- A1: Remove phantom tool references
|
|
112
|
+
- A2: Consolidate duplicated invariants
|
|
113
|
+
- A3: Convert to imperative triggers
|
|
114
|
+
- A4: Single source for cheat-sheets
|
|
115
|
+
- A5: Reduce prose bloat
|
|
116
|
+
- A14: Decompose omnibus file
|
|
117
|
+
- A15: Add concrete success metrics
|
|
118
|
+
4. Re-run audit
|
|
119
|
+
5. Update changelog (PATCH)
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
## Pattern 7: Fragility Reclassification
|
|
123
|
+
|
|
124
|
+
**Trigger**: Report shows wrong fragility classification
|
|
125
|
+
|
|
126
|
+
**Template**:
|
|
127
|
+
```
|
|
128
|
+
1. Identify misclassified task
|
|
129
|
+
2. Determine correct fragility:
|
|
130
|
+
- Mutation → strict
|
|
131
|
+
- Read-only → loose
|
|
132
|
+
- Creative → low specificity
|
|
133
|
+
3. Update instructions accordingly
|
|
134
|
+
4. Update changelog (MINOR if significant)
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
## Changelog Entry Templates
|
|
138
|
+
|
|
139
|
+
### For Fixes
|
|
140
|
+
```markdown
|
|
141
|
+
### Fixed
|
|
142
|
+
- <what was fixed> (from report: <report-name>)
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
### For Additions
|
|
146
|
+
```markdown
|
|
147
|
+
### Added
|
|
148
|
+
- <what was added> (from report: <report-name>)
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
### For Changes
|
|
152
|
+
```markdown
|
|
153
|
+
### Changed
|
|
154
|
+
- <what was changed> (from report: <report-name>)
|
|
155
|
+
```
|
|
@@ -0,0 +1,104 @@
|
|
|
1
|
+
# Report Analysis Guide
|
|
2
|
+
|
|
3
|
+
How to parse and analyze reports generated by `opsx-report`.
|
|
4
|
+
|
|
5
|
+
## Report Structure
|
|
6
|
+
|
|
7
|
+
Reports follow OKF v0.2 format:
|
|
8
|
+
|
|
9
|
+
```yaml
|
|
10
|
+
---
|
|
11
|
+
type: OpenSpec Report
|
|
12
|
+
title: "Report: <change-name>"
|
|
13
|
+
description: "<summary>"
|
|
14
|
+
status: stable
|
|
15
|
+
tags: [openspec, <schema>, <capability>]
|
|
16
|
+
generated:
|
|
17
|
+
by: opsx-report/1.0
|
|
18
|
+
at: <timestamp>
|
|
19
|
+
sources:
|
|
20
|
+
- id: proposal
|
|
21
|
+
resource: artifacts/proposal.md
|
|
22
|
+
- id: design
|
|
23
|
+
resource: artifacts/design.md
|
|
24
|
+
- id: specs
|
|
25
|
+
resource: artifacts/specs/
|
|
26
|
+
- id: tasks
|
|
27
|
+
resource: artifacts/tasks.md
|
|
28
|
+
stale_after: <date>
|
|
29
|
+
---
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
## Extraction Points
|
|
33
|
+
|
|
34
|
+
### From Frontmatter
|
|
35
|
+
|
|
36
|
+
| Field | Extract | Use |
|
|
37
|
+
|-------|---------|-----|
|
|
38
|
+
| `tags` | Skill name, capability-path | Identify related skill |
|
|
39
|
+
| `generated.at` | Timestamp | Track when report was created |
|
|
40
|
+
| `sources` | Artifact paths | Locate original artifacts |
|
|
41
|
+
| `stale_after` | Date | Check if report is stale |
|
|
42
|
+
|
|
43
|
+
### From Sections
|
|
44
|
+
|
|
45
|
+
| Section | Extract | Use |
|
|
46
|
+
|---------|---------|-----|
|
|
47
|
+
| Problem Statement | Why skill exists | Understand purpose |
|
|
48
|
+
| Approach | Decisions with rationale | Understand design choices |
|
|
49
|
+
| What Changed | Requirements added/modified/removed | Understand scope |
|
|
50
|
+
| Implementation | Tasks completed, files created | Understand what was built |
|
|
51
|
+
| Validation | Test results, pass/fail | Identify failures to fix |
|
|
52
|
+
| Trade-offs | Risks and mitigations | Understand constraints |
|
|
53
|
+
| Follow-ups | Deferred work | Identify improvements to make |
|
|
54
|
+
|
|
55
|
+
## Analysis Patterns
|
|
56
|
+
|
|
57
|
+
### Pattern 1: Validation Failures
|
|
58
|
+
|
|
59
|
+
```
|
|
60
|
+
IF validation section shows failures:
|
|
61
|
+
→ Identify which tests failed
|
|
62
|
+
→ Analyze why they failed
|
|
63
|
+
→ Fix skill to address failures
|
|
64
|
+
→ Re-run validation
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
### Pattern 2: Follow-ups
|
|
68
|
+
|
|
69
|
+
```
|
|
70
|
+
IF follow-ups section has items:
|
|
71
|
+
→ List all follow-up items
|
|
72
|
+
→ Prioritize by impact
|
|
73
|
+
→ Implement highest priority
|
|
74
|
+
→ Update changelog
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
### Pattern 3: Trade-offs
|
|
78
|
+
|
|
79
|
+
```
|
|
80
|
+
IF trade-offs section has items:
|
|
81
|
+
→ Evaluate if trade-offs are still valid
|
|
82
|
+
→ Check if alternatives are now viable
|
|
83
|
+
→ Consider upgrading if better option exists
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
### Pattern 4: Generalizable Patterns
|
|
87
|
+
|
|
88
|
+
```
|
|
89
|
+
IF report shows pattern applicable to all skills:
|
|
90
|
+
→ Check if create-skill already covers it
|
|
91
|
+
→ If not, update create-skill
|
|
92
|
+
→ Propagate to existing skills
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
## Quality Signals
|
|
96
|
+
|
|
97
|
+
From report analysis, extract quality signals:
|
|
98
|
+
|
|
99
|
+
- **Validation pass rate**: Percentage of tests that passed
|
|
100
|
+
- **Antipattern count**: Number of antipatterns detected
|
|
101
|
+
- **Description trigger rate**: How well description triggers
|
|
102
|
+
- **Task completion rate**: Percentage of tasks completed
|
|
103
|
+
|
|
104
|
+
Use these to measure improvement after changes.
|
|
@@ -0,0 +1,103 @@
|
|
|
1
|
+
# Skill Quality Criteria
|
|
2
|
+
|
|
3
|
+
Quality metrics and thresholds for skill evaluation.
|
|
4
|
+
|
|
5
|
+
## Quality Dimensions
|
|
6
|
+
|
|
7
|
+
### 1. Structural Quality
|
|
8
|
+
|
|
9
|
+
- **SKILL.md exists**: Required
|
|
10
|
+
- **Frontmatter valid**: name, description present
|
|
11
|
+
- **Name format**: lowercase, hyphens, 1-64 chars
|
|
12
|
+
- **Description format**: 1-1024 chars, imperative
|
|
13
|
+
- **Directory structure**: scripts/, references/, assets/ present
|
|
14
|
+
- **File references resolve**: All markdown links valid
|
|
15
|
+
|
|
16
|
+
### 2. Content Quality
|
|
17
|
+
|
|
18
|
+
- **Description specificity**: Imperative, intent-driven, not vague
|
|
19
|
+
- **Instruction clarity**: Actionable, not ambiguous
|
|
20
|
+
- **Progressive disclosure**: SKILL.md < 500 lines
|
|
21
|
+
- **Gotchas present**: Environment-specific facts documented
|
|
22
|
+
- **Examples provided**: BAD/GOOD pairs where helpful
|
|
23
|
+
|
|
24
|
+
### 3. Fragility Matching
|
|
25
|
+
|
|
26
|
+
- **Mutation tasks**: Strict preconditions, gating
|
|
27
|
+
- **Read-only tasks**: Loose triggers, latitude
|
|
28
|
+
- **Creative tasks**: Low specificity, examples
|
|
29
|
+
|
|
30
|
+
### 4. Antipattern Compliance
|
|
31
|
+
|
|
32
|
+
No violations of:
|
|
33
|
+
- A1: Phantom tool references
|
|
34
|
+
- A2: Duplicated invariants
|
|
35
|
+
- A3: Passive-voice triggers
|
|
36
|
+
- A4: Copy-pasted cheat-sheets
|
|
37
|
+
- A5: Prose bloat in triggers
|
|
38
|
+
- A14: Single file omnibus
|
|
39
|
+
- A15: Vague success bars
|
|
40
|
+
|
|
41
|
+
### 5. Trigger Accuracy
|
|
42
|
+
|
|
43
|
+
- **Description triggers on relevant prompts**: >80% pass rate
|
|
44
|
+
- **Description doesn't trigger on irrelevant prompts**: >90% pass rate
|
|
45
|
+
- **Near-miss negatives pass**: 100% for mutation tools
|
|
46
|
+
|
|
47
|
+
## Quality Score
|
|
48
|
+
|
|
49
|
+
Calculate quality score (0-100):
|
|
50
|
+
|
|
51
|
+
```
|
|
52
|
+
Structural: 30 points
|
|
53
|
+
- Valid frontmatter: 10
|
|
54
|
+
- Valid name: 5
|
|
55
|
+
- Valid description: 5
|
|
56
|
+
- Directory structure: 5
|
|
57
|
+
- File references: 5
|
|
58
|
+
|
|
59
|
+
Content: 30 points
|
|
60
|
+
- Description specificity: 10
|
|
61
|
+
- Instruction clarity: 10
|
|
62
|
+
- Progressive disclosure: 5
|
|
63
|
+
- Gotchas present: 5
|
|
64
|
+
|
|
65
|
+
Fragility: 15 points
|
|
66
|
+
- Correct classification: 10
|
|
67
|
+
- Appropriate specificity: 5
|
|
68
|
+
|
|
69
|
+
Antipatterns: 15 points
|
|
70
|
+
- No violations: 15
|
|
71
|
+
- Minor violations: 10
|
|
72
|
+
- Major violations: 0
|
|
73
|
+
|
|
74
|
+
Trigger: 10 points
|
|
75
|
+
- Positive trigger rate: 5
|
|
76
|
+
- Negative trigger rate: 5
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
## Thresholds
|
|
80
|
+
|
|
81
|
+
| Score | Quality Level | Action |
|
|
82
|
+
|-------|---------------|--------|
|
|
83
|
+
| 90-100 | Excellent | No improvements needed |
|
|
84
|
+
| 70-89 | Good | Minor improvements |
|
|
85
|
+
| 50-69 | Fair | Significant improvements |
|
|
86
|
+
| 0-49 | Poor | Major improvements or rebuild |
|
|
87
|
+
|
|
88
|
+
## Measurement
|
|
89
|
+
|
|
90
|
+
Before improvement:
|
|
91
|
+
1. Run validate-structure.sh → structural score
|
|
92
|
+
2. Run audit-antipatterns.sh → antipattern score
|
|
93
|
+
3. Review content → content score
|
|
94
|
+
4. Calculate total quality score
|
|
95
|
+
|
|
96
|
+
After improvement:
|
|
97
|
+
1. Re-run all checks
|
|
98
|
+
2. Calculate new quality score
|
|
99
|
+
3. Calculate delta (improvement percentage)
|
|
100
|
+
|
|
101
|
+
If delta < 0 (quality degraded) → rollback
|
|
102
|
+
If delta = 0 → plateau reached
|
|
103
|
+
If delta > 0 → improvement successful
|
|
@@ -0,0 +1,30 @@
|
|
|
1
|
+
# Tool Type Detection
|
|
2
|
+
|
|
3
|
+
## Detection Matrix
|
|
4
|
+
|
|
5
|
+
| Report Content | Tool Type | Action |
|
|
6
|
+
|----------------|-----------|--------|
|
|
7
|
+
| New skill files in artifacts/ | New skill | Create proposal for new skill |
|
|
8
|
+
| Existing skill improvement | Skill update | Create proposal for skill update |
|
|
9
|
+
| New prompt files in artifacts/ | New prompt | Create proposal for new prompt |
|
|
10
|
+
| Both skill and prompt | Combo | Create proposal for both |
|
|
11
|
+
| Knowledge gaps in assessment | Reference/docs | Create proposal for reference materials |
|
|
12
|
+
| Generic workflow improvement | Prompt | Create proposal for new prompt |
|
|
13
|
+
|
|
14
|
+
## Detection Signals
|
|
15
|
+
|
|
16
|
+
### From Report Metadata
|
|
17
|
+
- `tags` field → skill/prompt name
|
|
18
|
+
- `capability-path` → skill identifier
|
|
19
|
+
- `type` field → report category
|
|
20
|
+
|
|
21
|
+
### From Artifacts
|
|
22
|
+
- `skills/` directory → skill
|
|
23
|
+
- `prompts/` directory → prompt
|
|
24
|
+
- Both → combo
|
|
25
|
+
- `references/` → reference materials
|
|
26
|
+
|
|
27
|
+
### From Assessment
|
|
28
|
+
- Knowledge gaps → reference materials needed
|
|
29
|
+
- Tool improvements → skill/prompt updates
|
|
30
|
+
- Documentation gaps → new documentation
|
|
@@ -0,0 +1,104 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: openspec-report
|
|
3
|
+
description: >
|
|
4
|
+
Generate self-reflection (meditation) on an archived OpenSpec change.
|
|
5
|
+
Use when the user wants to create a report from an archived change,
|
|
6
|
+
document lessons learned, or capture AI agent self-assessment.
|
|
7
|
+
Do NOT use when exploring ideas (use openspec-explore), implementing changes,
|
|
8
|
+
creating proposals, or analyzing reports for skill improvements.
|
|
9
|
+
allowed-tools: Bash(openspec:*), Bash(ls:*), Bash(mkdir:*), Bash(mv:*)
|
|
10
|
+
license: MIT
|
|
11
|
+
compatibility: Requires openspec CLI.
|
|
12
|
+
metadata:
|
|
13
|
+
author: agentic
|
|
14
|
+
version: "1.0.1"
|
|
15
|
+
positive_triggers:
|
|
16
|
+
- "generate a report from an archived change"
|
|
17
|
+
- "document what was learned from this change"
|
|
18
|
+
- "create self-reflection on completed work"
|
|
19
|
+
anti_triggers:
|
|
20
|
+
- "explore ideas and investigate problems"
|
|
21
|
+
- "implement a proposal or apply changes"
|
|
22
|
+
- "analyze reports to improve skills"
|
|
23
|
+
---
|
|
24
|
+
|
|
25
|
+
# Openspec Report
|
|
26
|
+
|
|
27
|
+
Generate self-reflection (meditation) on archived OpenSpec changes.
|
|
28
|
+
|
|
29
|
+
**Timestamp rule**: fill the `{{timestamp}}` frontmatter placeholder by running `date -u +"%Y-%m-%dT%H:%M:%SZ"` at generation time. Never hand-write or estimate timestamps.
|
|
30
|
+
|
|
31
|
+
## Workflow
|
|
32
|
+
|
|
33
|
+
### Phase 1: Find Archive
|
|
34
|
+
|
|
35
|
+
Archives have date prefixes. Search intelligently:
|
|
36
|
+
|
|
37
|
+
**Exact match first:**
|
|
38
|
+
```bash
|
|
39
|
+
ls openspec/changes/archive/<name>/ 2>/dev/null
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
**Fuzzy search if not found:**
|
|
43
|
+
```bash
|
|
44
|
+
ls openspec/changes/archive/ | grep -i "<name>"
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
Handle results:
|
|
48
|
+
- One match → use it
|
|
49
|
+
- Multiple matches → ask user to select
|
|
50
|
+
- No matches → error with available archives
|
|
51
|
+
|
|
52
|
+
### Phase 2: Read Archive
|
|
53
|
+
|
|
54
|
+
Read `.openspec.yaml` for metadata, then all planning artifacts: proposal.md, design.md, specs/, tasks.md.
|
|
55
|
+
|
|
56
|
+
### Phase 3: Create Report Directory
|
|
57
|
+
|
|
58
|
+
Create `openspec/reports/<name>/` (just report.md and assessment.md — no subdirectories).
|
|
59
|
+
|
|
60
|
+
### Phase 4: Generate report.md
|
|
61
|
+
|
|
62
|
+
Use template at `assets/templates/report.md.template`. Fill placeholders with self-reflection content.
|
|
63
|
+
|
|
64
|
+
**Meditation sections:**
|
|
65
|
+
- What Happened
|
|
66
|
+
- What I Learned
|
|
67
|
+
- Mental Model Shift (1–2 lines: `Before (old model):` / `After (new model):` — e.g., `Before: C# null check (if x != null)` → `After: Rust Option (if let Some(x))`)
|
|
68
|
+
- Surprise vs Expectation (Expected vs Surprise — 1–2 lines)
|
|
69
|
+
- Concrete Gotcha (single before/after code snippet with Signal — e.g., `if (user != null)` vs `if let Some(u) = user`)
|
|
70
|
+
- Time Cost (where time went — e.g., `60m borrow checker, 10m logic`)
|
|
71
|
+
- Re-use Score (`high/medium/low — one-line why` — e.g., `high — every Rust task touches Option`)
|
|
72
|
+
- What I'd Do Differently
|
|
73
|
+
- Key Decisions (table)
|
|
74
|
+
- Trade-offs Made (table)
|
|
75
|
+
- Follow-ups
|
|
76
|
+
- Archive Reference
|
|
77
|
+
|
|
78
|
+
### Phase 5: Generate assessment.md
|
|
79
|
+
|
|
80
|
+
Use template at `assets/templates/assessment.md.template`. Fill with experiential reflection.
|
|
81
|
+
|
|
82
|
+
**Assessment sections:**
|
|
83
|
+
- How It Felt (What Went Well / What Went Poorly)
|
|
84
|
+
- Difficulty Ratings (5 dimensions)
|
|
85
|
+
- What I Was Lacking (knowledge gaps, skills gaps, Before → After Model — complements report's Mental Model Shift)
|
|
86
|
+
- What Would Help Next Time (process, tools, documentation, Re-use Score)
|
|
87
|
+
- Time Cost (mirrors report)
|
|
88
|
+
- Surprise vs Expectation (mirrors report)
|
|
89
|
+
|
|
90
|
+
### Phase 6: Validate
|
|
91
|
+
|
|
92
|
+
Check report has all meditation sections and assessment has minimum quality bar.
|
|
93
|
+
|
|
94
|
+
### Phase 7: Display Summary
|
|
95
|
+
|
|
96
|
+
Report location, archive reference, files generated.
|
|
97
|
+
|
|
98
|
+
## Gotchas
|
|
99
|
+
|
|
100
|
+
- **Archive is source of truth**: Reference it, don't copy content into reports.
|
|
101
|
+
- **Reports are meditation**: Self-reflection on experience, not documentation of what happened.
|
|
102
|
+
- **Minimum quality bar**: At least 2 items per reflection section, at least 1 knowledge gap identified, and one Concrete Gotcha snippet when a mistake or discovery occurred.
|
|
103
|
+
- **1–2 line shifts only; single gotcha snippet**: Mental Model Shift Before/After are 1–2 lines each; Concrete Gotcha is one before/after pair with Signal — prevents report bloat while preserving C#→Rust transferability.
|
|
104
|
+
- **Latest version only**: If no name specified, use most recent archive.
|
|
@@ -0,0 +1,84 @@
|
|
|
1
|
+
---
|
|
2
|
+
type: Assessment
|
|
3
|
+
title: "Assessment: {{change-name}}"
|
|
4
|
+
description: "AI agent meditation on the experience"
|
|
5
|
+
generated:
|
|
6
|
+
by: openspec-report/1.0
|
|
7
|
+
at: {{timestamp}}
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# Assessment: {{change-name}}
|
|
11
|
+
|
|
12
|
+
*Meditation on the experience of this change.*
|
|
13
|
+
|
|
14
|
+
## How It Felt
|
|
15
|
+
|
|
16
|
+
### What Went Well
|
|
17
|
+
|
|
18
|
+
{{what_went_well}}
|
|
19
|
+
|
|
20
|
+
### What Went Poorly
|
|
21
|
+
|
|
22
|
+
{{what_went_poorly}}
|
|
23
|
+
|
|
24
|
+
## Difficulty Ratings
|
|
25
|
+
|
|
26
|
+
| Dimension | Rating (1-5) | Notes |
|
|
27
|
+
|-----------|--------------|-------|
|
|
28
|
+
| Overall | {{overall_rating}} | {{overall_notes}} |
|
|
29
|
+
| Technical complexity | {{technical_rating}} | {{technical_notes}} |
|
|
30
|
+
| Scope clarity | {{scope_rating}} | {{scope_notes}} |
|
|
31
|
+
| Integration | {{integration_rating}} | {{integration_notes}} |
|
|
32
|
+
| Time estimation | {{time_rating}} | {{time_notes}} |
|
|
33
|
+
|
|
34
|
+
## What I Was Lacking
|
|
35
|
+
|
|
36
|
+
### Knowledge gaps
|
|
37
|
+
|
|
38
|
+
{{knowledge_gaps}}
|
|
39
|
+
|
|
40
|
+
### Skills gaps
|
|
41
|
+
|
|
42
|
+
{{skills_gaps}}
|
|
43
|
+
|
|
44
|
+
### Before → After Model
|
|
45
|
+
|
|
46
|
+
<!-- 1-2 lines complementing report's Mental Model Shift: what model did you hold before vs after. ex: Before: C# null is runtime value → After: Rust Option is type-level -->
|
|
47
|
+
|
|
48
|
+
Before: {{before_model}}
|
|
49
|
+
|
|
50
|
+
After: {{after_model}}
|
|
51
|
+
|
|
52
|
+
## What Would Help Next Time
|
|
53
|
+
|
|
54
|
+
### Process improvements
|
|
55
|
+
|
|
56
|
+
{{process_improvements}}
|
|
57
|
+
|
|
58
|
+
### Tool improvements
|
|
59
|
+
|
|
60
|
+
{{tool_improvements}}
|
|
61
|
+
|
|
62
|
+
### Documentation gaps
|
|
63
|
+
|
|
64
|
+
{{documentation_gaps}}
|
|
65
|
+
|
|
66
|
+
### Re-use Score
|
|
67
|
+
|
|
68
|
+
<!-- high/medium/low + one-line why — feeds learn's frequency×cost prioritization. ex: high — every Rust task touches Option -->
|
|
69
|
+
|
|
70
|
+
{{reuse_score}} — {{reuse_reason}}
|
|
71
|
+
|
|
72
|
+
## Time Cost
|
|
73
|
+
|
|
74
|
+
<!-- Where time went — mirrors report's Time Cost for learn clustering. ex: 60m borrow checker, 10m logic -->
|
|
75
|
+
|
|
76
|
+
{{time_cost}}
|
|
77
|
+
|
|
78
|
+
## Surprise vs Expectation
|
|
79
|
+
|
|
80
|
+
<!-- 1-2 lines mirroring report; helps learn prioritize surprises -->
|
|
81
|
+
|
|
82
|
+
Expected: {{expected}}
|
|
83
|
+
|
|
84
|
+
Surprise: {{surprise}}
|
|
@@ -0,0 +1,92 @@
|
|
|
1
|
+
---
|
|
2
|
+
type: OpenSpec Report
|
|
3
|
+
title: "Report: {{change-name}}"
|
|
4
|
+
description: "Self-reflection on archived change"
|
|
5
|
+
generated:
|
|
6
|
+
by: openspec-report/1.0
|
|
7
|
+
at: {{timestamp}}
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# Report: {{change-name}}
|
|
11
|
+
|
|
12
|
+
*Self-reflection on archived change: {{archive-path}}*
|
|
13
|
+
|
|
14
|
+
## What Happened
|
|
15
|
+
|
|
16
|
+
{{what_happened}}
|
|
17
|
+
|
|
18
|
+
## What I Learned
|
|
19
|
+
|
|
20
|
+
{{what_i_learned}}
|
|
21
|
+
|
|
22
|
+
## Mental Model Shift
|
|
23
|
+
|
|
24
|
+
<!-- 1-2 lines: Before (old model) → After (new model). ex: Before: C# null check (if x != null) → After: Rust Option (if let Some(x) = x) — absence is type-level, not runtime value -->
|
|
25
|
+
|
|
26
|
+
Before (old model): {{mental_model_before}}
|
|
27
|
+
|
|
28
|
+
After (new model): {{mental_model_after}}
|
|
29
|
+
|
|
30
|
+
## Surprise vs Expectation
|
|
31
|
+
|
|
32
|
+
<!-- 1-2 lines: what you expected vs what actually happened -->
|
|
33
|
+
|
|
34
|
+
Expected: {{expected}}
|
|
35
|
+
|
|
36
|
+
Surprise: {{surprise}}
|
|
37
|
+
|
|
38
|
+
## Concrete Gotcha
|
|
39
|
+
|
|
40
|
+
<!-- Single before/after code snippet with signal. ex: before (wrong): if (user != null) → after (right): if let Some(u) = user — signal: compiler forces None handling -->
|
|
41
|
+
|
|
42
|
+
Before (wrong):
|
|
43
|
+
|
|
44
|
+
```{{gotcha_lang_before}}
|
|
45
|
+
{{gotcha_before}}
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
After (right):
|
|
49
|
+
|
|
50
|
+
```{{gotcha_lang_after}}
|
|
51
|
+
{{gotcha_after}}
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
Signal: {{gotcha_signal}}
|
|
55
|
+
|
|
56
|
+
## Time Cost
|
|
57
|
+
|
|
58
|
+
<!-- Where time went. ex: 60m borrow checker, 10m logic -->
|
|
59
|
+
|
|
60
|
+
{{time_cost}}
|
|
61
|
+
|
|
62
|
+
## Re-use Score
|
|
63
|
+
|
|
64
|
+
<!-- high/medium/low + one-line why. ex: high — every Rust task touches Option -->
|
|
65
|
+
|
|
66
|
+
{{reuse_score}} — {{reuse_reason}}
|
|
67
|
+
|
|
68
|
+
## What I'd Do Differently
|
|
69
|
+
|
|
70
|
+
{{what_i_would_do_differently}}
|
|
71
|
+
|
|
72
|
+
## Key Decisions
|
|
73
|
+
|
|
74
|
+
| Decision | Choice | Rationale |
|
|
75
|
+
|----------|--------|-----------|
|
|
76
|
+
{{key_decisions_table}}
|
|
77
|
+
|
|
78
|
+
## Trade-offs Made
|
|
79
|
+
|
|
80
|
+
| Risk | Mitigation |
|
|
81
|
+
|------|------------|
|
|
82
|
+
{{tradeoffs_table}}
|
|
83
|
+
|
|
84
|
+
## Follow-ups
|
|
85
|
+
|
|
86
|
+
{{followups}}
|
|
87
|
+
|
|
88
|
+
## Archive Reference
|
|
89
|
+
|
|
90
|
+
- **Location**: {{archive_path}}
|
|
91
|
+
- **Date**: {{archive_date}}
|
|
92
|
+
- **Schema**: {{schema_name}}
|