@phuc1403/musketeer 0.8.0 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (76) hide show
  1. package/README.md +49 -49
  2. package/manifest.json +333 -301
  3. package/package.json +1 -1
  4. package/template/.claude/agents/code-reviewer.md +182 -166
  5. package/template/.claude/hooks/git-skill-reminder.cjs +53 -0
  6. package/template/.claude/hooks/lib/colors.cjs +180 -122
  7. package/template/.claude/hooks/lib/transcript-parser.cjs +300 -277
  8. package/template/.claude/skills/code-review/SKILL.md +201 -54
  9. package/template/.claude/skills/code-review/references/checklist-workflow.md +96 -0
  10. package/template/.claude/skills/code-review/references/checklists/api.md +52 -52
  11. package/template/.claude/skills/code-review/references/checklists/base.md +100 -100
  12. package/template/.claude/skills/code-review/references/checklists/web-app.md +54 -54
  13. package/template/.claude/skills/code-review/references/code-review-reception.md +113 -0
  14. package/template/.claude/skills/code-review/references/codebase-scan-workflow.md +30 -0
  15. package/template/.claude/skills/code-review/references/edge-case-scouting.md +119 -0
  16. package/template/.claude/skills/code-review/references/input-mode-resolution.md +135 -0
  17. package/template/.claude/skills/code-review/references/parallel-review-workflow.md +76 -0
  18. package/template/.claude/skills/code-review/references/requesting-code-review.md +116 -0
  19. package/template/.claude/skills/code-review/references/spec-compliance-review.md +43 -0
  20. package/template/.claude/skills/code-review/references/task-management-reviews.md +140 -0
  21. package/template/.claude/skills/code-review/references/verification-before-completion.md +139 -0
  22. package/template/.claude/skills/git/SKILL.md +131 -115
  23. package/template/.claude/skills/git/references/branch-management.md +88 -88
  24. package/template/.claude/skills/git/references/commit-standards.md +46 -46
  25. package/template/.claude/skills/git/references/context-efficiency.md +54 -0
  26. package/template/.claude/skills/git/references/gh-cli-guide.md +109 -109
  27. package/template/.claude/skills/git/references/safety-protocols.md +69 -69
  28. package/template/.claude/skills/git/references/workflow-commit.md +58 -58
  29. package/template/.claude/skills/git/references/workflow-merge-pr.md +136 -0
  30. package/template/.claude/skills/git/references/workflow-merge.md +48 -48
  31. package/template/.claude/skills/git/references/workflow-pr.md +58 -58
  32. package/template/.claude/skills/git/references/workflow-push.md +52 -52
  33. package/template/.claude/skills/skill-creator/LICENSE.txt +201 -201
  34. package/template/.claude/skills/skill-creator/SKILL.md +154 -149
  35. package/template/.claude/skills/skill-creator/agents/analyzer.md +274 -274
  36. package/template/.claude/skills/skill-creator/agents/comparator.md +202 -202
  37. package/template/.claude/skills/skill-creator/agents/grader.md +223 -223
  38. package/template/.claude/skills/skill-creator/assets/eval_review.html +146 -146
  39. package/template/.claude/skills/skill-creator/eval-viewer/generate_review.py +471 -471
  40. package/template/.claude/skills/skill-creator/eval-viewer/viewer.html +1325 -1325
  41. package/template/.claude/skills/skill-creator/references/benchmark-optimization-guide.md +86 -86
  42. package/template/.claude/skills/skill-creator/references/distribution-guide.md +79 -79
  43. package/template/.claude/skills/skill-creator/references/eval-infrastructure-guide.md +129 -129
  44. package/template/.claude/skills/skill-creator/references/eval-schemas.md +121 -121
  45. package/template/.claude/skills/skill-creator/references/mcp-skills-integration.md +71 -71
  46. package/template/.claude/skills/skill-creator/references/metadata-quality-criteria.md +94 -94
  47. package/template/.claude/skills/skill-creator/references/plugin-marketplace-hosting.md +104 -104
  48. package/template/.claude/skills/skill-creator/references/plugin-marketplace-overview.md +89 -89
  49. package/template/.claude/skills/skill-creator/references/plugin-marketplace-schema.md +93 -93
  50. package/template/.claude/skills/skill-creator/references/plugin-marketplace-sources.md +103 -103
  51. package/template/.claude/skills/skill-creator/references/plugin-marketplace-troubleshooting.md +76 -76
  52. package/template/.claude/skills/skill-creator/references/script-quality-criteria.md +106 -106
  53. package/template/.claude/skills/skill-creator/references/skill-anatomy-and-requirements.md +77 -77
  54. package/template/.claude/skills/skill-creator/references/skill-creation-workflow.md +152 -151
  55. package/template/.claude/skills/skill-creator/references/skill-design-patterns.md +75 -75
  56. package/template/.claude/skills/skill-creator/references/skillmark-benchmark-criteria.md +102 -102
  57. package/template/.claude/skills/skill-creator/references/structure-organization-criteria.md +114 -114
  58. package/template/.claude/skills/skill-creator/references/testing-and-iteration.md +78 -78
  59. package/template/.claude/skills/skill-creator/references/token-efficiency-criteria.md +74 -74
  60. package/template/.claude/skills/skill-creator/references/troubleshooting-guide.md +81 -81
  61. package/template/.claude/skills/skill-creator/references/validation-checklist.md +83 -83
  62. package/template/.claude/skills/skill-creator/references/writing-effective-instructions.md +88 -88
  63. package/template/.claude/skills/skill-creator/references/yaml-frontmatter-reference.md +92 -92
  64. package/template/.claude/skills/skill-creator/scripts/aggregate_benchmark.py +401 -401
  65. package/template/.claude/skills/skill-creator/scripts/encoding_utils.py +36 -36
  66. package/template/.claude/skills/skill-creator/scripts/generate_report.py +326 -326
  67. package/template/.claude/skills/skill-creator/scripts/improve_description.py +248 -248
  68. package/template/.claude/skills/skill-creator/scripts/init_skill.py +360 -360
  69. package/template/.claude/skills/skill-creator/scripts/package_skill.py +143 -143
  70. package/template/.claude/skills/skill-creator/scripts/quick_validate.py +110 -110
  71. package/template/.claude/skills/skill-creator/scripts/run_eval.py +310 -310
  72. package/template/.claude/skills/skill-creator/scripts/run_loop.py +332 -332
  73. package/template/.claude/skills/skill-creator/scripts/utils.py +47 -47
  74. package/template/.claude/statusline.cjs +0 -0
  75. package/template/.claude/skills/code-review/references/adversarial-review.md +0 -223
  76. /package/template/.claude/hooks/{usage-context-awareness.cjs → usage-quota-cache-refresh.cjs} +0 -0
@@ -1,47 +1,47 @@
1
- """Shared utilities for skill-creator scripts."""
2
-
3
- from pathlib import Path
4
-
5
-
6
-
7
- def parse_skill_md(skill_path: Path) -> tuple[str, str, str]:
8
- """Parse a SKILL.md file, returning (name, description, full_content)."""
9
- content = (skill_path / "SKILL.md").read_text()
10
- lines = content.split("\n")
11
-
12
- if lines[0].strip() != "---":
13
- raise ValueError("SKILL.md missing frontmatter (no opening ---)")
14
-
15
- end_idx = None
16
- for i, line in enumerate(lines[1:], start=1):
17
- if line.strip() == "---":
18
- end_idx = i
19
- break
20
-
21
- if end_idx is None:
22
- raise ValueError("SKILL.md missing frontmatter (no closing ---)")
23
-
24
- name = ""
25
- description = ""
26
- frontmatter_lines = lines[1:end_idx]
27
- i = 0
28
- while i < len(frontmatter_lines):
29
- line = frontmatter_lines[i]
30
- if line.startswith("name:"):
31
- name = line[len("name:"):].strip().strip('"').strip("'")
32
- elif line.startswith("description:"):
33
- value = line[len("description:"):].strip()
34
- # Handle YAML multiline indicators (>, |, >-, |-)
35
- if value in (">", "|", ">-", "|-"):
36
- continuation_lines: list[str] = []
37
- i += 1
38
- while i < len(frontmatter_lines) and (frontmatter_lines[i].startswith(" ") or frontmatter_lines[i].startswith("\t")):
39
- continuation_lines.append(frontmatter_lines[i].strip())
40
- i += 1
41
- description = " ".join(continuation_lines)
42
- continue
43
- else:
44
- description = value.strip('"').strip("'")
45
- i += 1
46
-
47
- return name, description, content
1
+ """Shared utilities for skill-creator scripts."""
2
+
3
+ from pathlib import Path
4
+
5
+
6
+
7
+ def parse_skill_md(skill_path: Path) -> tuple[str, str, str]:
8
+ """Parse a SKILL.md file, returning (name, description, full_content)."""
9
+ content = (skill_path / "SKILL.md").read_text()
10
+ lines = content.split("\n")
11
+
12
+ if lines[0].strip() != "---":
13
+ raise ValueError("SKILL.md missing frontmatter (no opening ---)")
14
+
15
+ end_idx = None
16
+ for i, line in enumerate(lines[1:], start=1):
17
+ if line.strip() == "---":
18
+ end_idx = i
19
+ break
20
+
21
+ if end_idx is None:
22
+ raise ValueError("SKILL.md missing frontmatter (no closing ---)")
23
+
24
+ name = ""
25
+ description = ""
26
+ frontmatter_lines = lines[1:end_idx]
27
+ i = 0
28
+ while i < len(frontmatter_lines):
29
+ line = frontmatter_lines[i]
30
+ if line.startswith("name:"):
31
+ name = line[len("name:"):].strip().strip('"').strip("'")
32
+ elif line.startswith("description:"):
33
+ value = line[len("description:"):].strip()
34
+ # Handle YAML multiline indicators (>, |, >-, |-)
35
+ if value in (">", "|", ">-", "|-"):
36
+ continuation_lines: list[str] = []
37
+ i += 1
38
+ while i < len(frontmatter_lines) and (frontmatter_lines[i].startswith(" ") or frontmatter_lines[i].startswith("\t")):
39
+ continuation_lines.append(frontmatter_lines[i].strip())
40
+ i += 1
41
+ description = " ".join(continuation_lines)
42
+ continue
43
+ else:
44
+ description = value.strip('"').strip("'")
45
+ i += 1
46
+
47
+ return name, description, content
Binary file
@@ -1,223 +0,0 @@
1
- ---
2
- name: adversarial-review
3
- description: Stage 3 red-team review that actively tries to break code — finds security holes, false assumptions, failure modes, race conditions. Spawns adversarial reviewer subagent with destructive mindset. Includes scope gate for trivial changes.
4
- ---
5
-
6
- # Adversarial Review (Stage 3)
7
-
8
- Runs after every Stage 2 (Code Quality) pass. Subject to scope gate below.
9
-
10
- ## Scope Gate
11
-
12
- Skip adversarial review when ALL of these are true:
13
- - Changed files <= 2
14
- - Lines changed <= 30
15
- - No security-sensitive files touched (auth, crypto, input parsing, SQL, env)
16
- - No new dependencies added
17
-
18
- When skipped, note: `Adversarial: skipped (below threshold)` in review output.
19
-
20
- **NEVER skip when:**
21
- - Any file in: `auth/`, `middleware/`, `security/`, `crypto/`
22
- - `package.json`, `package-lock.json`, or lockfile changed
23
- - Environment variables added/changed
24
- - Database schema modified
25
- - API route added/changed
26
-
27
- ## Mindset
28
-
29
- > "You are hired to tear apart the implementer's work. Your job is to find every way this code can fail, be exploited, or produce incorrect results. Assume the implementer made mistakes. Prove it."
30
-
31
- This is NOT a standard code review. Standard reviews check if code meets requirements. Adversarial review assumes requirements are met and asks: **"How can this still break?"**
32
-
33
- ## What to Attack
34
-
35
- ### Security Holes
36
- - Injection vectors (SQL, command, XSS, template)
37
- - Auth bypass paths (missing checks, privilege escalation)
38
- - Secrets exposure (logs, error messages, stack traces)
39
- - Input trust boundaries (user input treated as safe)
40
- - SSRF, path traversal, deserialization attacks
41
-
42
- ### False Assumptions
43
- - "This will never be null" -- prove it can be
44
- - "This list always has elements" -- find the empty case
45
- - "Users always call A before B" -- find the out-of-order path
46
- - "This config value exists" -- find the missing env var
47
- - "This third-party API always returns 200" -- find the failure mode
48
- - "This API shape won't change" -- find the breaking caller
49
-
50
- ### Failure Modes & Resource Exhaustion
51
- - What happens when disk is full?
52
- - What happens when network times out mid-operation?
53
- - What happens when the database connection drops during a transaction?
54
- - Unbounded allocations from user-controlled input
55
- - Missing timeouts on external calls
56
- - Event loop blocking (sync operations in async context)
57
- - Connection/handle leaks on error paths
58
- - Regex catastrophic backtracking (ReDoS)
59
-
60
- ### Race Conditions
61
- - Shared mutable state without locks
62
- - Time-of-check-to-time-of-use (TOCTOU)
63
- - Async operations with implicit ordering assumptions
64
- - Cache invalidation during concurrent writes
65
-
66
- ### Data Corruption
67
- - Partial writes on failure (no transaction/rollback)
68
- - Type coercion surprises (string "0" as falsy)
69
- - Floating point comparison for equality
70
- - Timezone-naive datetime operations
71
-
72
- ### Supply Chain & Dependencies
73
- - New dependencies: postinstall scripts, maintainer reputation, bundle size
74
- - Lockfile changes: version drift, removed integrity hashes
75
- - Transitive deps pulling in known-vulnerable packages
76
-
77
- ### Observability Blind Spots
78
- - Swallowed errors (`catch {}` with no log)
79
- - Missing structured context in error logs
80
- - PII in log output
81
-
82
- ## Process
83
-
84
- ### 1. Spawn Adversarial Reviewer
85
-
86
- Dispatch `code-reviewer` subagent with adversarial prompt:
87
-
88
- ```
89
- You are an adversarial code reviewer. Your ONLY job is to find ways this code
90
- can fail, be exploited, or produce incorrect results.
91
-
92
- DO NOT praise the code. DO NOT note what works well.
93
- ONLY report problems. If you find nothing, say "No findings" -- but try harder first.
94
-
95
- Focus on ADDED/MODIFIED lines (+ prefix in diff). Pre-existing code is out of scope
96
- unless the change makes it newly exploitable.
97
-
98
- Context (read for understanding, DO NOT review):
99
- {CONTEXT_FILES}
100
-
101
- Runtime: {RUNTIME} (e.g., Node.js single-threaded, browser, serverless)
102
- Framework: {FRAMEWORK} (e.g., Express with global error handler at app.ts:45)
103
-
104
- Review this diff:
105
- {DIFF}
106
-
107
- Changed files: {FILES}
108
-
109
- Attack vectors to check:
110
- 1. Security holes (injection, auth bypass, secrets exposure)
111
- 2. False assumptions (null, empty, ordering, config, API contracts)
112
- 3. Failure modes + resource exhaustion (timeouts, leaks, unbounded input)
113
- 4. Race conditions (shared state, TOCTOU, async ordering)
114
- 5. Data corruption (partial writes, type coercion, encoding)
115
- 6. Supply chain (new deps, lockfile changes, transitive vulns)
116
- 7. Observability (swallowed errors, missing logs, PII in output)
117
-
118
- For each finding, report:
119
- - SEVERITY: Critical / Medium / Low
120
- - CATEGORY: Security / Assumption / Failure / Race / Data / Supply / Observability
121
- - LOCATION: file:line
122
- - ATTACK: How to trigger the problem
123
- - IMPACT: What happens when triggered
124
- - FIX: Describe the fix approach (e.g., "add null check before line 42").
125
- Do NOT write implementation code -- the implementer has full context.
126
- ```
127
-
128
- **If adversarial produces >10 findings on <100 lines changed:** likely too aggressive. Batch-reject noise, deep-review only Critical/Medium.
129
-
130
- ### 2. Adjudicate Findings
131
-
132
- Main agent reviews each adversarial finding and assigns verdict:
133
-
134
- | Verdict | Meaning | Action |
135
- |---------|---------|--------|
136
- | **Accept** | Valid flaw, reproducible or clearly reasoned | Must fix before merge |
137
- | **Reject** | False positive, already handled, or impossible path | Document why, no action |
138
- | **Defer** | Valid but low-risk, tracked for later | Create GitHub issue for tracking |
139
-
140
- **Rules:**
141
- - Every finding gets a verdict -- no silent dismissals
142
- - Critical findings: Accept unless you can PROVE false positive
143
- - Benefit of doubt goes to the adversary (safer to fix than to dismiss)
144
- - If >50% of findings are Rejected, the adversary was too aggressive -- but still report all
145
-
146
- **Calibration examples:**
147
-
148
- | Verdict | Example | Reasoning |
149
- |---------|---------|-----------|
150
- | Accept | "SQL injection via string interpolation in query builder" | Clearly exploitable, concrete path shown |
151
- | Reject | "Missing null check on config.apiUrl" | Config loaded at startup with schema validation (see config.ts:12), cannot be null at runtime |
152
- | Defer | "No rate limiting on POST /api/upload" | Valid concern but internal-only tool currently; track for public exposure |
153
-
154
- ### 3. Report Format
155
-
156
- ```
157
- ## Adversarial Review -- Stage 3
158
-
159
- ### Summary
160
- - Findings: N total (X Critical, Y Medium, Z Low)
161
- - Accepted: A (must fix)
162
- - Rejected: B (false positive)
163
- - Deferred: C (tracked via GitHub issues)
164
-
165
- ### Accepted Findings (Must Fix)
166
-
167
- #### [1] SEVERITY -- CATEGORY -- file:line
168
- **Attack:** How to trigger
169
- **Impact:** What happens
170
- **Fix:** Approach description
171
- **Verdict:** Accept -- [reason]
172
-
173
- ### Rejected Findings
174
-
175
- #### [N] SEVERITY -- CATEGORY -- file:line
176
- **Attack:** Claimed vector
177
- **Verdict:** Reject -- [reason this is a false positive]
178
-
179
- ### Deferred Findings
180
-
181
- #### [N] SEVERITY -- CATEGORY -- file:line
182
- **Attack:** How to trigger
183
- **Verdict:** Defer -- [reason] → GitHub issue #X
184
- ```
185
-
186
- ### 4. Fix Accepted Findings
187
-
188
- - Critical: Block merge. Fix immediately via `/fix` or manual edit.
189
- - Medium: Fix before merge if feasible. Defer only with explicit user approval.
190
- - Low: Track. Fix in follow-up if pattern repeats.
191
-
192
- ### Re-review Optimization
193
-
194
- On fix cycles (re-running after accepted findings were fixed):
195
- - Only pass the FIX diff to adversarial, not the full original diff
196
- - Verify accepted findings are resolved
197
- - Check for regression: did the fix introduce new issues?
198
-
199
- ## Integration with Pipeline
200
-
201
- ```
202
- Stage 1 (Spec) → PASS
203
- ↓
204
- Stage 2 (Quality) → PASS
205
- ↓
206
- Scope gate → below threshold? → skip (note in report)
207
- ↓ (above threshold)
208
- Stage 3 (Adversarial) → findings
209
- ├─ 0 Accepted → PASS → proceed
210
- ├─ Accepted Critical → BLOCK → fix → re-run Stage 3 (fix diff only)
211
- └─ Accepted Medium/Low only → fix or defer → proceed
212
- ```
213
-
214
- **Task pipeline update:** When using task-managed reviews, adversarial review gets its own task between "Review implementation" and "Fix critical issues".
215
-
216
- ## What This Is NOT
217
-
218
- - NOT a style review (Stage 2 handles that)
219
- - NOT a spec compliance check (Stage 1 handles that)
220
- - NOT dependency graph analysis or import tracing (scout handles that)
221
- - NOT a general "suggestions for improvement" pass
222
-
223
- This is a focused, hostile attempt to break the code. If the code survives, it's ready to ship.