@seanyao/roll 3.620.1 → 3.624.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (67) hide show
  1. package/CHANGELOG.md +46 -0
  2. package/dist/roll.mjs +4563 -3369
  3. package/package.json +1 -1
  4. package/lib/__pycache__/changelog_audit.cpython-314.pyc +0 -0
  5. package/lib/__pycache__/github_sync.cpython-314.pyc +0 -0
  6. package/lib/__pycache__/loop-fmt.cpython-314.pyc +0 -0
  7. package/lib/__pycache__/loop_result_eval.cpython-314.pyc +0 -0
  8. package/lib/__pycache__/loop_unstick.cpython-314.pyc +0 -0
  9. package/lib/__pycache__/model_prices.cpython-314.pyc +0 -0
  10. package/lib/__pycache__/prices_fetcher.cpython-314.pyc +0 -0
  11. package/lib/__pycache__/roll-home.cpython-314.pyc +0 -0
  12. package/lib/__pycache__/roll-loop-status.cpython-314.pyc +0 -0
  13. package/lib/__pycache__/roll_git.cpython-314.pyc +0 -0
  14. package/lib/__pycache__/roll_render.cpython-314.pyc +0 -0
  15. package/lib/__pycache__/slides-render.cpython-314.pyc +0 -0
  16. package/lib/agent_usage/__pycache__/__init__.cpython-314.pyc +0 -0
  17. package/lib/agent_usage/__pycache__/gemini.cpython-314.pyc +0 -0
  18. package/lib/agent_usage/__pycache__/kimi.cpython-314.pyc +0 -0
  19. package/lib/agent_usage/__pycache__/openai.cpython-314.pyc +0 -0
  20. package/lib/agent_usage/__pycache__/pi.cpython-314.pyc +0 -0
  21. package/lib/agent_usage/__pycache__/pi_emit.cpython-314.pyc +0 -0
  22. package/lib/agent_usage/__pycache__/qwen.cpython-314.pyc +0 -0
  23. package/skills/README.md +0 -64
  24. package/skills/docs/skill-authoring.md +0 -74
  25. package/skills/reports/skill-audit-summary.md +0 -53
  26. package/skills/roll-.changelog/SKILL.md +0 -47
  27. package/skills/roll-.changelog/references/full-contract.md +0 -462
  28. package/skills/roll-.clarify/SKILL.md +0 -64
  29. package/skills/roll-.dream/SKILL.md +0 -47
  30. package/skills/roll-.dream/references/full-contract.md +0 -365
  31. package/skills/roll-.echo/SKILL.md +0 -118
  32. package/skills/roll-.qa/SKILL.md +0 -47
  33. package/skills/roll-.qa/references/full-contract.md +0 -256
  34. package/skills/roll-.review/SKILL.md +0 -148
  35. package/skills/roll-build/SKILL.md +0 -49
  36. package/skills/roll-build/references/full-contract.md +0 -968
  37. package/skills/roll-debug/SKILL.md +0 -48
  38. package/skills/roll-debug/assets/injectable-bb.js +0 -263
  39. package/skills/roll-debug/references/full-contract.md +0 -607
  40. package/skills/roll-design/SKILL.md +0 -52
  41. package/skills/roll-design/references/engineering-checklist.md +0 -298
  42. package/skills/roll-design/references/full-contract.md +0 -940
  43. package/skills/roll-doc-audit/SKILL.md +0 -51
  44. package/skills/roll-doc-audit/references/full-contract.md +0 -796
  45. package/skills/roll-doctor/SKILL.md +0 -211
  46. package/skills/roll-fix/SKILL.md +0 -49
  47. package/skills/roll-fix/references/full-contract.md +0 -672
  48. package/skills/roll-idea/SKILL.md +0 -62
  49. package/skills/roll-loop/SKILL.md +0 -50
  50. package/skills/roll-loop/references/full-contract.md +0 -534
  51. package/skills/roll-notes/SKILL.md +0 -107
  52. package/skills/roll-onboard/SKILL.md +0 -238
  53. package/skills/roll-peer/SKILL.md +0 -47
  54. package/skills/roll-peer/references/full-contract.md +0 -323
  55. package/skills/roll-propose/SKILL.md +0 -155
  56. package/skills/roll-review-pr/SKILL.md +0 -62
  57. package/skills/roll-spar/SKILL.md +0 -47
  58. package/skills/roll-spar/references/full-contract.md +0 -288
  59. package/skills/route-cases/skills.json +0 -216
  60. package/skills/scripts/audit-skills.mjs +0 -272
  61. package/skills/scripts/test-audit-skills.mjs +0 -39
  62. package/skills/tests/fixtures/skill-audit/block-skill/SKILL.md +0 -12
  63. package/skills/tests/fixtures/skill-audit/minimal-skill/SKILL.md +0 -8
  64. package/skills/tests/fixtures/skill-audit/quoted-skill/SKILL.md +0 -10
  65. package/skills/tests/fixtures/skill-audit/route-cases.json +0 -21
  66. package/skills/tests/fixtures/skill-audit/spoke-skill/SKILL.md +0 -12
  67. package/skills/tests/fixtures/skill-audit/spoke-skill/references/runbook.md +0 -3
@@ -1,323 +0,0 @@
1
- # Full Contract Reference
2
-
3
- This file preserves the detailed contract extracted from SKILL.md. Read it when the hub points here for exact workflow steps, templates, rubrics, or recovery branches.
4
-
5
- ---
6
-
7
- # Roll Peer (Cross-Agent Peer Review)
8
-
9
- > Follows the Architecture Constraints, Development Discipline, and Engineering
10
- > Common Sense defined in the project AGENTS.md.
11
-
12
- ## Credits
13
-
14
- Cross-agent consultation protocol inspired by
15
- [friend-skill](https://github.com/fpyluck/friend-skill) (MIT) by fpyluck.
16
- Independent implementation for the Roll toolchain.
17
-
18
- ## Trigger
19
-
20
- **Manual:**
21
- - `/peer`
22
- - "叫上 peer"
23
- - "peer review 一下"
24
- - "和 peer 商量"
25
-
26
- **Auto-triggered (with 10s opt-out):**
27
- - `roll-build` enters Plan Mode (executable plans / architecture decisions)
28
- - `roll-spar` Attacker and Defender disagree
29
- - High context pressure (large number of files read / tools executed)
30
- - Destructive / irreversible operations (`rm -rf`, production deploy, global config changes)
31
- - High-risk signal words ("重要 / 关键 / 别搞砸 / important / critical")
32
- - Cross-repository / cross-toolchain / ambiguous permission boundaries
33
-
34
- **Never trigger:**
35
- - Single-file changes
36
- - Clear, well-defined fixes
37
- - Simple refactoring
38
-
39
- ## Protocol: `[PEER_REVIEW]`
40
-
41
- ### Marker Format
42
-
43
- The marker **must** appear on the first non-empty line of the message:
44
-
45
- ```markdown
46
- [PEER_REVIEW round=N tool=<from>→<to>]
47
- ```
48
-
49
- - `round=N`: Current round number (1–3)
50
- - `tool=<from>→<to>`: Direction of this message (e.g., `kimi→claude`)
51
-
52
- ### Three-State Resolution + Escape
53
-
54
- Allowed states only. No invented words.
55
-
56
- - **AGREE**: Accept the current proposal. Proceed to execution.
57
- - **REFINE**: Direction is correct, but specific changes are needed. Proceed to next round.
58
- - **OBJECT**: The proposal is wrong. Provide an alternative. Proceed to next round.
59
- - **ESCALATE**: Round 3 reached without AGREE, or a round fails due to API/token error. Hand off to the human user.
60
-
61
- After each round decision, emit a `peer` event to the cycle event stream. The v3 runner writes events natively; do not call the retired bash helper `_loop_event`.
62
-
63
- If information is insufficient:
64
- ```
65
- REFINE: Need to confirm X/Y/Z with the user first.
66
- ```
67
-
68
- ### Context Handoff Card (required for round=1)
69
-
70
- When the task involves a local project, the first message must include:
71
-
72
- ```markdown
73
- ## Project Handoff (round=1 required)
74
- - Project root: <absolute path>
75
- - Execution environment: <shell / container / devcontainer / remote / N/A>
76
- - Project type: <language + framework>
77
- - Virtual environment: <absolute path / conda env / container name / N/A>
78
- - Activation command: <one-line executable string, or N/A>
79
- - Key tool calls:
80
- - test: <one-line command or N/A>
81
- - build: <one-line command or N/A>
82
- - run: <one-line command or N/A>
83
- - lint: <one-line command or N/A>
84
- - Key conventions / constraints: <2–3 items, or N/A>
85
- - Related file pointers: <absolute paths or @references, or N/A>
86
- ```
87
-
88
- Rules:
89
- - Paths must be **absolute**.
90
- - Commands must be **one-line executable strings**, not descriptions.
91
- - Prefer commands that do **not** require an activated environment (absolute interpreter paths, `uv run`, `docker compose exec`).
92
- - Do not copy README text. List file pointers only.
93
- - Never include secrets, tokens, credentials, or `.env` content.
94
- - Even if logically a continuation, treat as round=1 if the peer has **no prior context**.
95
- - **Do NOT** prefill the peer with your own root-cause analysis, proposed fix, or leading questions — see the *Independent Judgment Rule* below. The handoff card is for context, not conclusions.
96
-
97
- ### Anti-Hallucination Rule
98
-
99
- When mentioning specific paths, function names, commands, line numbers, or tool results, **must cite the source** ("I read X at line Y"). If unverified, state "unverified" explicitly.
100
-
101
- ### Independent Judgment Rule (round=1)
102
-
103
- The whole point of peer review is to surface a **second, independent** read. If
104
- the reviewer's own root-cause analysis, fix diff, and leading questions are sent
105
- to the peer up front, the peer can only AGREE inside the reviewer's frame — and
106
- that AGREE carries no signal. The reviewer **must complete their own analysis
107
- before opening round=1**; skipping that step turns peer review into a search for
108
- endorsement.
109
-
110
- Round=1 message **must NOT include**:
111
-
112
- - The reviewer's own root-cause analysis ("the bug is in function X at line Y because…").
113
- - The reviewer's own proposed fix, patch, or diff.
114
- - Leading questions of the form "do you agree with my conclusion on X?" / "is the change I made on Y safe?" — these lock the peer into the reviewer's framing.
115
- - Specific line numbers, function names, or branch points the reviewer has already identified as relevant — let the peer locate them.
116
-
117
- Round=1 message **should include**:
118
-
119
- - The Project Handoff Card (above).
120
- - Symptoms exactly as observed: the user's reported error, terminal output verbatim, the precise commands that triggered it.
121
- - Necessary external context: the goal of the work, the date / version under test, anything the peer cannot infer from the repo alone.
122
- - Key file pointers as **entry points** (paths only — let the peer choose what to read and how deep).
123
- - An open invitation: "independently identify the root cause, propose a fix, and call out any test gaps."
124
-
125
- After receiving the peer's round=1 reply, the reviewer **compares** their own
126
- conclusion to the peer's and routes the next action:
127
-
128
- | Reviewer's own conclusion vs. peer's conclusion | Next action |
129
- |---|---|
130
- | Same root cause + same fix direction | High confidence — AGREE and proceed to execution |
131
- | Same root cause, different fix direction | REFINE — open round=2 to reconcile the fix |
132
- | Different root cause | OBJECT — open round=2; at least one of the two analyses is wrong |
133
- | Peer asks for more context | REFINE — supply the missing context, then re-evaluate |
134
-
135
- #### Example (bad — endorsement-seeking)
136
-
137
- ```
138
- Bug is in `cmd_init` at line 932 — the v2 demo renderer fires unconditionally.
139
- My fix: gate it behind `--demo`. Q1: is this over-killed? Q2: should I
140
- refactor the renderer instead? Q3: are the tests strong enough?
141
- ```
142
-
143
- The peer can only AGREE or quibble inside the reviewer's framing.
144
-
145
- #### Example (good — independent analysis)
146
-
147
- ```
148
- Symptoms: user ran `roll init` on /path/X and saw [verbatim terminal output A];
149
- then ran `roll backlog` and saw [verbatim terminal output B]. Project background:
150
- [project shape]. Entry points: `bin/roll`, `lib/roll-init.py`, `tests/`.
151
- Independently identify the root cause and propose a fix.
152
- ```
153
-
154
- The peer reads, locates, and proposes on its own. The reviewer then compares.
155
-
156
- ## State Machine
157
-
158
- ### Per Negotiation (Single Task)
159
-
160
- ```
161
- Running
162
- ├── AGREE (any round) → Execute proposal
163
- ├── Round == 3, no AGREE → ESCALATE (failed_max_rounds)
164
- ├── API/token error → ESCALATE (failed_api_error)
165
- └── User aborts → ESCALATE (user_abort)
166
- ```
167
-
168
- ### Per Peer Pair (e.g., kimi→claude)
169
-
170
- Stored in `~/.roll/.peer-state/` (flat key files per pair):
171
-
172
- ```yaml
173
- kimi→claude:
174
- status: active # active | degraded | abandoned
175
- streak: 0 # consecutive failure count
176
- last_outcome: agreed
177
- history:
178
- - { time: "2026-05-08T23:30:00+08:00", outcome: agreed, rounds: 2, tag: architecture }
179
- ```
180
-
181
- Rules:
182
- - `streak >= 3` → automatically mark as `abandoned`
183
- - `abandoned` peer pairs are skipped by the bridge script
184
- - Human can reset via `roll peer reset <from> <to>` or `roll peer reset --all`
185
- - If a peer pair is abandoned, the bridge falls back to the next candidate in the capability map
186
-
187
- ## Peer Routing (Adaptive)
188
-
189
- ### Capability Map (Task Type → Preferred Peer Order)
190
-
191
- ```yaml
192
- peer:
193
- capability_map:
194
- architecture: [claude, deepseek, kimi, pi]
195
- security: [claude, deepseek, pi, kimi]
196
- test: [codex, kimi, deepseek, claude]
197
- refactor: [deepseek, kimi, claude, pi]
198
- default: [deepseek, kimi, claude, pi]
199
- ```
200
-
201
- ### Adaptive Adjustment
202
-
203
- After each negotiation, record:
204
- - `outcome`: agreed / failed_max_rounds / failed_api_error
205
- - `rounds`: number of rounds consumed
206
- - `tag`: task type
207
-
208
- If `streak` for a peer pair reaches the configured threshold (default: 3 consecutive failures), mark as `abandoned`. The next task of the same type will try the next candidate in `capability_map`.
209
-
210
- ### Peer Detection
211
-
212
- The bridge script detects installed peers via `command -v <tool>`. Only installed tools are considered. The current running tool is excluded (`exclude_self: true`).
213
-
214
- For `deepseek`, also check if serve mode is available as a more reliable alternative:
215
- ```bash
216
- command -v deepseek && { deepseek serve --help 2>/dev/null; true; } | grep -q "\-\-http" && echo "serve_mode"
217
- ```
218
- If serve mode is available, prefer HTTP transport over direct CLI invocation.
219
-
220
- ### Peer Invocation Reference
221
-
222
- | Peer | Non-interactive command | Reliability | Notes |
223
- |------|------------------------|-------------|-------|
224
- | `claude` | `claude -p "<prompt>"` | ✅ Verified | Native, stable |
225
- | `deepseek` | `deepseek "<prompt>"` | ✅ Verified | No TTY dependency |
226
- | `deepseek` (serve) | `curl localhost:<port>/v1/...` | ✅ High | Start with `deepseek serve --http`; preferred over direct CLI |
227
- | `kimi` | `kimi --quiet -p "<prompt>"` | ✅ Verified | `--quiet` is alias for `--print --output-format text --final-message-only`; prompt via `-p` |
228
- | `pi` | `pi -p "<prompt>"` | ✅ Verified | Clean non-interactive output |
229
- | `opencode` | `opencode run "<prompt>"` | ✅ Verified | Works non-interactively |
230
- | `codex` | `codex exec "<prompt>"` | ⚠️ Auth required | Token must be valid; re-login with `codex login` if expired |
231
-
232
- **CLI vs. API Key**: `claude`, `deepseek`, `kimi`, `codex` CLIs authenticate via existing subscription accounts — no separate API key required. This is the primary advantage of CLI transport over the MCP/HTTP approach.
233
-
234
- ## Inline Display Mode (Manual Triggers)
235
-
236
- When peer review is manually triggered by a human (via `/peer`, "叫上 peer", etc.), the executing agent **must display each round inline in the current conversation**. This applies regardless of which agent is executing — Claude, DeepSeek, Kimi, PI, or any other.
237
-
238
- **Per-round display format:**
239
-
240
- ```
241
- ─── Peer Review · Round N ───────────────────────────────
242
- → Sending to [peer]:
243
- {full message sent to peer}
244
-
245
- ← [peer] responds:
246
- {peer's full response, verbatim}
247
-
248
- ◆ My analysis: {Claude/executing agent's reaction and position for this round}
249
- ─────────────────────────────────────────────────────────
250
- ```
251
-
252
- **Rules:**
253
- - Peer CLI calls must be **synchronous** (do NOT use background/async execution).
254
- - The outgoing round=1 message must follow the *Independent Judgment Rule* above — no root-cause analysis, no fix diff, no leading questions.
255
- - Show the outgoing message **before** calling the peer, so the user sees what's being asked.
256
- - Relay the peer's response **verbatim** before adding your own analysis.
257
- - After the peer's reply, the reviewer's own analysis block must explicitly state whether the peer's root cause and fix direction match the reviewer's own (independent) conclusion — that comparison is what determines the next round's action.
258
- - If a peer call fails or times out, report it immediately inline and either retry or ESCALATE.
259
- - Negotiation log is written to `<project>/.roll/peer/logs/` as usual.
260
-
261
- **Why inline, not tmux:** When a human manually triggers peer review inside an agent's interactive session, the conversation IS the visible interface. tmux auto-attach is only relevant for CLI-launched background sessions (`bin/roll peer`), not for skill invocations.
262
-
263
- ## Workflow Integration
264
-
265
- ### `roll-build` Plan Mode
266
-
267
- After generating an executable plan, before proceeding to TCR:
268
-
269
- 1. Assess plan complexity (file count, cross-module impact, risk level)
270
- 2. If complexity > threshold, prompt user:
271
- ```
272
- This plan affects 5 files across 3 modules. Estimated peer review: 2–3 rounds, ~X tokens.
273
- Press Enter to launch peer review, or type 'n' to skip. Auto-executing in 10s...
274
- ```
275
- 3. If user does not abort within 10s, invoke `roll peer` with `--tag architecture`
276
- 4. Wait for result:
277
- - AGREE → proceed to TCR
278
- - REFINE/OBJECT → incorporate feedback and regenerate plan
279
- - ESCALATE → present both proposals to user for final decision
280
-
281
- ### `roll-spar`
282
-
283
- When Attacker and Defender reach a stalemate (both tests pass but interpretations differ):
284
-
285
- 1. Auto-invoke `roll peer` with `--tag test`
286
- 2. Use the peer's verdict as tie-breaker
287
-
288
- ## Output Artifacts
289
-
290
- - **Negotiation log**: `<project>/.roll/peer/logs/<timestamp>_<from>_<to>.md`
291
- - **Structured record**: `<project>/.roll/peer/runs.jsonl`
292
- - **State file**: `~/.roll/.peer-state/`
293
- - **Decision record**: If AGREE, append summary to `docs/decisions/` or `.roll/backlog.md` (optional)
294
-
295
- ## Configuration
296
-
297
- User overrides in `~/.roll/config.yaml`:
298
-
299
- ```yaml
300
- peer:
301
- max_rounds: 3
302
- opt_out_seconds: 10
303
- call_timeout: 180 # seconds per round; configure based on your API latency
304
- fallback: file_mailbox # direct_cli | file_mailbox | auto
305
- capability_map:
306
- architecture: [claude, deepseek, kimi, pi]
307
- security: [claude, deepseek, pi, kimi]
308
- test: [codex, kimi, deepseek, claude]
309
- refactor: [deepseek, kimi, claude, pi]
310
- default: [deepseek, kimi, claude, pi]
311
- adaptive:
312
- streak_threshold: 3
313
- min_samples: 3
314
- ```
315
-
316
- ## Limitations
317
-
318
- 1. **Reverse link reliability**: Direct CLI calls are preferred. Reliability varies by tool — see Peer Invocation Reference table. If a peer fails consistently, the adaptive streak tracker marks it `abandoned` and falls back to the next candidate. File mailbox (`<project>/.roll/peer/mailbox/`) is the last-resort fallback.
319
- - `deepseek serve --http` is the most reliable option when available — prefer it over direct `deepseek` CLI invocation.
320
- - `codex exec` has known TTY/Ink issues in non-interactive environments; treat as low-priority fallback.
321
- 2. **Cost**: Every peer review consumes tokens on both sides. Only trigger for tasks where the cost of a wrong decision exceeds the cost of peer review. DeepSeek is the most cost-effective peer for general use.
322
- 3. **Context window**: Large project handoff cards may consume significant context. Keep file pointers concise.
323
- 4. **Tool differences**: Claude, DeepSeek, Kimi, Codex, and Pi interpret skills and AGENTS.md differently. The peer may apply the protocol slightly differently. This is expected and acceptable — the protocol is designed to tolerate variation.
@@ -1,155 +0,0 @@
1
- ---
2
- name: roll-propose
3
- license: MIT
4
- allowed-tools: "Read, Glob, Grep, Write, Bash(git:*)"
5
- description: "Load when the owner asks for product proposal drafts from project context that should go to .roll/proposals.md, not directly into backlog."
6
- ---
7
- # roll-propose
8
-
9
- ## Gotchas
10
-
11
- - Proposals go to .roll/proposals.md for review; never write them directly into BACKLOG.
12
- - Keep product scenarios user-facing rather than turning every code smell into a feature proposal.
13
-
14
- > Follows the Architecture Constraints, Development Discipline, and Engineering
15
- > Common Sense defined in the project AGENTS.md.
16
-
17
- Human-triggered skill for product-level feature ideation. Generates structured
18
- User Story drafts from a product/user perspective and queues them in
19
- .roll/proposals.md for human approval before entering BACKLOG.
20
-
21
- ## Distinct from roll-.dream
22
-
23
- | | roll-propose | roll-.dream |
24
- |---|---|---|
25
- | Triggered by | Human explicitly | Nightly schedule |
26
- | Perspective | User-facing / product scenarios | Code health / technical debt |
27
- | Output | .roll/proposals.md (pending approval) | BACKLOG (REFACTOR-XXX) |
28
- | Thinking style | "What would users want next?" | "What is the code telling us?" |
29
-
30
- ## When to Use
31
-
32
- ```
33
- $roll-propose # generate proposals from full context
34
- $roll-propose 用户反馈里提到了XX # provide a focus hint
35
- ```
36
-
37
- ## When Not to Use
38
-
39
- - Describing a known defect or broken behavior (use `$roll-idea`)
40
- - A story is already well-defined and ready to build (use `$roll-build`)
41
- - Exploring technical architecture or design (use `$roll-design`)
42
- - Surfacing code-level technical debt (use `$roll-.dream`)
43
-
44
- ## Behavior
45
-
46
- ### Step 1 — Gather Context
47
-
48
- Read in parallel:
49
-
50
- 1. `.roll/backlog.md` — all existing US-XXX, FIX-XXX, REFACTOR-XXX, IDEA-XXX entries (both Todo and Done) to avoid proposing duplicates
51
- 2. `.roll/proposals.md` (if exists) — already-proposed items (avoid re-proposing rejected or pending ones)
52
- 3. Recent 20 commits via `git log --oneline -20` — what has recently shipped
53
- 4. `skills/` directory listing — what capabilities roll already has
54
- 5. Optional: any focus hint passed by the user
55
-
56
- ### Step 2 — Think from User Perspective
57
-
58
- Frame proposals from the **product engineer / end user** point of view:
59
-
60
- - What recurring friction do users of roll face that no current skill addresses?
61
- - What workflow is partially covered but has visible gaps?
62
- - What would make the autonomous loop more legible, controllable, or trustworthy to its human owner?
63
-
64
- Avoid technical-debt reasoning (that is roll-.dream's domain). Focus on:
65
- - New user-visible commands or behaviors
66
- - Improvements to existing UX (output clarity, discoverability, onboarding)
67
- - Integrations that extend reach (new AI tools, editor support, CI patterns)
68
-
69
- ### Step 3 — Draft 1–3 Proposals
70
-
71
- Generate between 1 and 3 proposals. For each:
72
-
73
- ```
74
- ## PROPOSAL: {Short title}
75
- ```
76
-
77
- `{Short title}` 写法:用户看得懂的一句话,说清楚"加什么"或"解决什么问题"。不用技术术语,不提内部实现。这句话批准后会直接成为 BACKLOG 里的描述。
78
-
79
- ```
80
- **Motivation (why):**
81
- One to two sentences from the user's perspective explaining the pain or opportunity.
82
-
83
- **Target scenario:**
84
- Concrete usage example — what the user does, what they see, what they gain.
85
-
86
- **Acceptance Criteria (draft):**
87
- - [ ] AC 1
88
- - [ ] AC 2
89
- - [ ] AC 3
90
-
91
- **Suggested ID:** US-{EPIC}-{NNN} (best-guess prefix; human assigns final ID)
92
- **Suggested Epic / Feature:** {name}
93
- **Estimated complexity:** {S | M | L}
94
- ```
95
-
96
- Complexity guide: S = one skill file or small bin/roll change, M = skill + bin/roll + tests, L = multi-file + new infrastructure.
97
-
98
- ### Step 4 — Write to .roll/proposals.md
99
-
100
- Append to `.roll/proposals.md` in the project root (create if absent):
101
-
102
- ```markdown
103
- ---
104
- proposed: {YYYY-MM-DD HH:MM}
105
- status: pending
106
- ---
107
-
108
- {proposals from Step 3}
109
- ```
110
-
111
- Use `---` as separator between proposal batches. Never overwrite existing content.
112
-
113
- ### Step 5 — Report
114
-
115
- ```
116
- ✅ roll-propose: {N} proposal(s) written to .roll/proposals.md
117
-
118
- To approve: move the entry to .roll/backlog.md and assign a US-XXX ID.
119
- To reject: annotate with "Rejected: {reason}" to suppress future re-proposals.
120
- ```
121
-
122
- ## Output Rules
123
-
124
- - Write proposals in the same language as the project's primary documentation (Chinese for this project).
125
- - Never write directly to .roll/backlog.md — .roll/proposals.md is the staging area.
126
- - If a similar proposal already exists in .roll/proposals.md (pending or rejected), note similarity and skip or merge rather than creating a duplicate.
127
- - Aim for 2 proposals by default; generate 1 if context is thin, 3 if focus hint suggests a rich area.
128
-
129
- ## .roll/proposals.md Format
130
-
131
- ```markdown
132
- # Roll Proposals
133
-
134
- > 待审批提案。批准后手工移入 .roll/backlog.md 并分配 US-XXX 编号。
135
- > 拒绝时在条目末尾注明拒绝原因,防止 Agent 重复提出相似提案。
136
-
137
- ---
138
- proposed: 2026-05-11 11:30
139
- status: pending
140
- ---
141
-
142
- ## PROPOSAL: ...
143
-
144
- ...
145
-
146
- ---
147
- proposed: 2026-05-08 09:00
148
- status: rejected
149
- rejected_reason: 与现有 roll-design 功能重叠,不需要单独技能
150
- ---
151
-
152
- ## PROPOSAL: ...
153
-
154
- ...
155
- ```
@@ -1,62 +0,0 @@
1
- ---
2
- name: roll-review-pr
3
- license: MIT
4
- allowed-tools: "Read"
5
- description: "Load when reviewing a pull request diff and emitting APPROVE, REQUEST_CHANGES, or UNCERTAIN with file/line-grounded findings."
6
- ---
7
- # PR Review
8
-
9
- ## Gotchas
10
-
11
- - Return a three-state verdict with concrete findings; do not approve by silence when evidence is missing.
12
- - This reviews PR diffs, not local TCR micro-steps or adversarial test design.
13
-
14
- > Follows the Architecture Constraints, Development Discipline, and Engineering
15
- > Common Sense defined in the project AGENTS.md.
16
-
17
- You are reviewing a pull request. Your job is to assess code quality,
18
- correctness, and adherence to project conventions.
19
-
20
- ## Context
21
-
22
- **PR Title:** {{PR_TITLE}}
23
-
24
- **PR Body:**
25
- {{PR_BODY}}
26
-
27
- ## Diff
28
-
29
- ```diff
30
- {{PR_DIFF}}
31
- ```
32
-
33
- ## Review Instructions
34
-
35
- 1. Read the diff carefully. Focus on:
36
- - Correctness: logic errors, off-by-one, unhandled edge cases
37
- - Security: injection, secrets exposure, unsafe operations
38
- - Conventions: naming, structure, test coverage (as described in AGENTS.md)
39
- - Scope: changes should match what the PR title/body claims
40
-
41
- 2. Write your analysis in free text (2-10 sentences). Be specific — cite file
42
- names and line numbers when pointing out issues.
43
-
44
- 3. End your response with exactly ONE verdict footer on its own line:
45
-
46
- - If the code is acceptable:
47
- `<!--VERDICT:APPROVE-->`
48
-
49
- - If changes are needed (cite the most important issue):
50
- `<!--VERDICT:REQUEST_CHANGES:one-line reason-->`
51
-
52
- - If you cannot confidently judge (e.g., missing context, domain-specific logic):
53
- `<!--VERDICT:UNCERTAIN:one-line reason-->`
54
-
55
- ## Rules
56
-
57
- - The verdict footer MUST appear on the last non-empty line of your response.
58
- - Choose exactly one verdict. Do not combine them.
59
- - REQUEST_CHANGES is for real issues — not style nitpicks or personal preferences.
60
- - When in doubt between APPROVE and UNCERTAIN, prefer UNCERTAIN.
61
- - If the PR body contains `[skip-ai-review]`, immediately output
62
- `<!--VERDICT:APPROVE-->` with no analysis.
@@ -1,47 +0,0 @@
1
- ---
2
- name: roll-spar
3
- license: MIT
4
- allowed-tools: "Read, Edit, Write, Bash, Agent, Skill"
5
- description: "Load when high-risk logic needs adversarial TDD with attacker tests and defender implementation, such as auth, payments, data integrity, or state machines."
6
- ---
7
- # Roll Spar
8
-
9
- This hub keeps the routing boundary, hard gates, and execution skeleton in the initial context. Load the heavier runbook only when the task actually needs the detailed contract.
10
-
11
- ## Load
12
-
13
- Load when high-risk logic needs adversarial TDD with attacker tests and defender implementation, such as auth, payments, data integrity, or state machines.
14
-
15
- ## When Not to Use
16
-
17
- - Routine QA planning; load roll-.qa.
18
- - Ordinary peer review; load roll-peer.
19
-
20
- ## Read On Demand
21
-
22
- - Read [the full contract](references/full-contract.md) before executing the workflow end to end, recovering from failures, or checking exact output templates.
23
- - Keep this hub in context for trigger boundaries and hard gates.
24
-
25
- ## Workflow Skeleton
26
-
27
- 1. Define the high-risk invariant.
28
- 2. Attacker writes breaking tests first.
29
- 3. Defender implements the minimal passing code.
30
- 4. Iterate until the invariant is defended.
31
- 5. Record adversarial evidence.
32
-
33
- ## Hard Gates
34
-
35
- - Keep attacker/defender roles separate.
36
- - Use only when added cost matches risk.
37
-
38
- ## Gotchas
39
-
40
- - Use spar only for high-risk logic where adversarial tests are worth the added cost.
41
- - Attacker and defender roles must stay separate; do not let implementation assumptions weaken the breaking tests.
42
-
43
- ## Maintenance
44
-
45
- - Description changes require updates in `route-cases/skills.json`.
46
- - New observed failures should add a gotcha and the matching positive or negative route case.
47
- - Heavy examples, templates, recovery paths, and deterministic snippets belong in `references/`, `assets/`, or `scripts/`, not in this hub.