@vegastack/skills 0.9.1 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (85) hide show
  1. package/README.md +8 -3
  2. package/dist/index.js +5 -5
  3. package/package.json +1 -1
  4. package/skill/dev-architect/SKILL.md +96 -0
  5. package/skill/dev-architect/agents/openai.yaml +4 -0
  6. package/skill/dev-architect/references/ai-agents.md +89 -0
  7. package/skill/dev-architect/references/conventions.md +93 -0
  8. package/skill/{architect → dev-architect}/references/data.md +43 -44
  9. package/skill/dev-architect/references/infra.md +98 -0
  10. package/skill/dev-architect/references/mobile.md +75 -0
  11. package/skill/{architect → dev-architect}/references/pinned-facts.md +17 -16
  12. package/skill/dev-architect/references/principles.md +117 -0
  13. package/skill/{architect → dev-architect}/references/security.md +37 -44
  14. package/skill/dev-architect/references/stack.md +38 -0
  15. package/skill/dev-architect/references/web.md +102 -0
  16. package/skill/{architect → dev-architect}/refresh/REFRESH.md +8 -6
  17. package/skill/{architect → dev-architect}/refresh/sources.json +5 -10
  18. package/skill/dev-chronicle/SKILL.md +45 -0
  19. package/skill/dev-chronicle/agents/openai.yaml +4 -0
  20. package/skill/dev-chronicle/references/conventions.md +93 -0
  21. package/skill/dev-chronicle/refresh/REFRESH.md +3 -0
  22. package/skill/dev-chronicle/refresh/sources.json +6 -0
  23. package/skill/dev-debug/SKILL.md +43 -0
  24. package/skill/dev-debug/agents/openai.yaml +4 -0
  25. package/skill/dev-debug/references/conventions.md +93 -0
  26. package/skill/dev-debug/references/loop-ladder.md +20 -0
  27. package/skill/dev-debug/refresh/REFRESH.md +3 -0
  28. package/skill/dev-debug/refresh/sources.json +6 -0
  29. package/skill/dev-implement/SKILL.md +41 -36
  30. package/skill/dev-implement/references/conventions.md +93 -0
  31. package/skill/dev-implement/references/ledger-and-resume.md +27 -0
  32. package/skill/dev-implement/scripts/evidence-check.mjs +57 -0
  33. package/skill/dev-implement/scripts/lib/gh.mjs +93 -0
  34. package/skill/dev-implement/scripts/preflight.mjs +101 -0
  35. package/skill/dev-intake/SKILL.md +39 -33
  36. package/skill/dev-intake/references/brief-template.md +27 -12
  37. package/skill/dev-intake/references/conventions.md +93 -0
  38. package/skill/dev-intake/scripts/brief-lint.mjs +87 -0
  39. package/skill/dev-plan/SKILL.md +53 -0
  40. package/skill/dev-plan/agents/openai.yaml +4 -0
  41. package/skill/dev-plan/references/conventions.md +93 -0
  42. package/skill/dev-plan/references/plan-format.md +54 -0
  43. package/skill/dev-plan/refresh/REFRESH.md +3 -0
  44. package/skill/dev-plan/refresh/sources.json +6 -0
  45. package/skill/dev-plan/scripts/plan-lint.mjs +86 -0
  46. package/skill/dev-review/SKILL.md +69 -0
  47. package/skill/dev-review/agents/openai.yaml +4 -0
  48. package/skill/dev-review/assets/review-known-patterns.md.template +30 -0
  49. package/skill/dev-review/references/conventions.md +93 -0
  50. package/skill/dev-review/references/cross-agent.md +39 -0
  51. package/skill/dev-review/references/dispatch-prompts.md +104 -0
  52. package/skill/dev-review/references/security-axis.md +33 -0
  53. package/skill/dev-review/refresh/REFRESH.md +3 -0
  54. package/skill/dev-review/refresh/sources.json +6 -0
  55. package/skill/dev-setup/SKILL.md +14 -9
  56. package/skill/dev-setup/assets/agents-section.md.template +2 -2
  57. package/skill/dev-setup/assets/dev-profile.md.template +23 -5
  58. package/skill/dev-setup/references/conventions.md +93 -0
  59. package/skill/dev-setup/references/stack-playbooks.md +1 -1
  60. package/skill/dev-ship/SKILL.md +14 -7
  61. package/skill/dev-ship/references/conventions.md +93 -0
  62. package/skill/dev-ship/references/runbook.md +1 -1
  63. package/skill/dev-ship/scripts/ship-gate.mjs +213 -0
  64. package/skill/dev-status/SKILL.md +45 -0
  65. package/skill/dev-status/agents/openai.yaml +4 -0
  66. package/skill/dev-status/references/conventions.md +93 -0
  67. package/skill/dev-status/refresh/REFRESH.md +3 -0
  68. package/skill/dev-status/refresh/sources.json +6 -0
  69. package/skill/dev-status/scripts/status.mjs +152 -0
  70. package/skill/skill-maintainer/references/release-ops.md +3 -3
  71. package/skill/skillify/SKILL.md +1 -1
  72. package/skill/skillify/references/eval-playbook.md +6 -0
  73. package/skill-integrity.json +93 -33
  74. package/skill/architect/SKILL.md +0 -68
  75. package/skill/architect/agents/openai.yaml +0 -4
  76. package/skill/architect/assets/adr-template.md +0 -21
  77. package/skill/architect/assets/arch-template.md +0 -20
  78. package/skill/architect/references/advisory.md +0 -102
  79. package/skill/architect/references/ai-agents.md +0 -95
  80. package/skill/architect/references/infra.md +0 -128
  81. package/skill/architect/references/mobile.md +0 -78
  82. package/skill/architect/references/principles.md +0 -91
  83. package/skill/architect/references/project-profile.md +0 -37
  84. package/skill/architect/references/stack.md +0 -38
  85. package/skill/architect/references/web.md +0 -152
@@ -0,0 +1,45 @@
1
+ ---
2
+ name: dev-chronicle
3
+ description: The project's narrative record — what got built, why, and how it went, in plain language for the operator's future recall. Use when asked to "catch me up on this project", "what did we build here", "what happened in this repo", "tell me the story so far", "what's the project story", when a chronicle entry needs writing for finished work, or when dev-implement's hand-back cites the chronicle format. Not for the consumer-facing changelog (dev-implement writes those per the changelog knob), release notes (dev-ship), or current board state ("what needs me" is dev-status).
4
+ ---
5
+
6
+ # dev-chronicle
7
+
8
+ `.vegastack/chronicle.md` is the project's story, newest first — the answer to "what did I build here and what happened?" months later, when the operator remembers nothing. Entries are **story language for a human**, never commit-log prose: the changelog tells consumers what changed; the chronicle tells the operator what happened.
9
+
10
+ Nearest neighbors: `dev-implement` writes the entries at hand-back (the write rule lives there; the format lives here); `dev-status` answers "what needs me now" — this skill answers "how did we get here". `dev-ship`'s ship-gate checks entry presence when dev.md says `chronicle: on`.
11
+
12
+ ## The entry — one per behavior-changing branch
13
+
14
+ ```markdown
15
+ ## DD-MM-YYYY — <title: the change as a human outcome — not a mechanism, not a commit subject> (#<issue>)
16
+
17
+ **What:** <2–4 plain sentences: what exists now that didn't, from the operator's point of view>
18
+ **Why:** <the need that prompted it>
19
+ **How it went:** <the honest one-liner: smooth / what fought back / what was cut>
20
+ **Changed:** <the user-visible changes, simple words — bullets or one ·-separated line>
21
+ **Decisions:** <register lines it produced, or "none">
22
+ — approved by operator (<username>) · built by <agent> · branch <name>
23
+ ```
24
+
25
+ - Titles name the outcome ("Invoice reminders now chase late payers"), never the mechanism ("add reminderAt column").
26
+ - Prepend — newest first. File missing → create it with a two-line header naming this skill as the format home.
27
+ - **How it went** is where honesty lives: what fought back, what was cut, what surprised. "Smooth" is a fine answer; silence is not.
28
+ - Research issues get an entry only when the findings changed direction; docs/test-only merges get none (ship-gate's excuse flag covers both records at once).
29
+ - A notable ship event — rollback, failed release — becomes its own short entry on the next branch, not a rewrite of an old one. Entries are never edited except for typos; the story is append-only like the register.
30
+
31
+ ## The digest — "catch me up"
32
+
33
+ On "catch me up on this project" (or any story-so-far ask), read **only** `.vegastack/chronicle.md` and the decision register — never git archaeology — and render three parts, plain language throughout:
34
+
35
+ 1. **The story so far** — 3–5 sentences: what this project is, the arc of what's been built, where it stands.
36
+ 2. **Recent chapters** — the last 3–7 entries, one line each: date, the outcome title, and the one thing worth remembering from How-it-went.
37
+ 3. **Open threads** — pending decisions the register hasn't recorded, entries whose How-it-went named unfinished business, and (when `dev-status` is installed) a one-line pointer to run it for the live board.
38
+
39
+ Length scales with the ask: "catch me up quickly" is one paragraph; a returning-after-months operator gets all three parts. Never pad — a young project with three entries gets three honest lines.
40
+
41
+ ## Setup
42
+
43
+ The `chronicle:` knob in dev.md (`on` default | `off`) governs whether dev-implement writes entries and ship-gate checks them; `dev-setup` writes the knob. A project that turns it on mid-life starts from now — no retroactive backfill unless the operator asks, and then it's marked as reconstructed.
44
+
45
+ Close every run with the plain-language summary: what was written or rendered, and anything the story surfaced that deserves the operator's attention.
@@ -0,0 +1,4 @@
1
+ interface:
2
+ display_name: "dev-chronicle"
3
+ short_description: "The project's story — entries and the catch-me-up digest"
4
+ default_prompt: "Use $dev-chronicle to catch me up on this project."
@@ -0,0 +1,93 @@
1
+ # Workflow conventions
2
+
3
+ The single spec for the artifacts every dev-family skill reads and writes. One home per rule: skills cite this file, never restate it. Everything here is harness-neutral.
4
+
5
+ ## Comment metadata markers
6
+
7
+ Every workflow-generated issue comment opens with an invisible HTML marker followed by a human heading:
8
+
9
+ ```markdown
10
+ <!-- vsk:v1 type=<type> rev=<n> [key=value ...] -->
11
+ ## <Human title> (v<n>)
12
+ ```
13
+
14
+ | type | required keys | instances |
15
+ |---|---|---|
16
+ | `approval` | `scope=<brief\|brief+plan\|plan>` | one per approval event |
17
+ | `plan` | `rev` | one, edited in place |
18
+ | `ledger` | `branch` | one, edited in place |
19
+ | `evidence` | `rev branch sha` | one, edited in place |
20
+ | `review` | `round sha agent=<claude\|codex> verdict=<clean\|needs-fixes>` | one per review cycle, rounds appended inside |
21
+ | `decision` | — | one per decision proposal |
22
+ | `handback` | — | one per stop event |
23
+
24
+ `rev=<n>` and the matching `(v<n>)` heading suffix appear only on revisable artifacts — the brief (issue description), `plan`, and `evidence` — starting at `rev=1`/`(v1)`. Single-event comments (`approval`, `decision`, `handback`) and the `ledger` carry neither. Scripts and agents locate comments strictly by marker, never by heading text. A comment without its marker does not count as the artifact — there is no legacy fallback.
25
+
26
+ ## Operator identity
27
+
28
+ Every human reference in every artifact — approvals, revisions, decisions, changelog attributions, review adjudications — is written `operator (<github-username>)`:
29
+
30
+ - Approval: `Approved by operator (<username>) on DD-MM-YYYY: "<their words>"`
31
+ - Register line: `- DD-MM-YYYY operator (<username>) — <decision>`
32
+
33
+ ## Revision markers
34
+
35
+ Any artifact edited after its first approval: the heading gains `(v2)`, the marker gains `rev=2`, and a `Revisions:` line is appended at the bottom — `v2 — DD-MM-YYYY: <what changed>, per operator (<username>) correction`. Existing revision lines are never rewritten.
36
+
37
+ ## Scope classes
38
+
39
+ Set at intake, applied as a label, announced with its reason (operator can override):
40
+
41
+ - **`research`** — a question to answer; throwaway code allowed, never merged. No branch/PR/changelog; findings + recommendation are the evidence comment.
42
+ - **`quick-build`** — small change and the flow being changed already exists in the repo to read. Brief (description) + plan (comment) are drafted in the same conversation; **one approval covers both**; then straight to `ready`.
43
+ - **`full-plan`** — big or new ground. Brief approval → `needs-plan` → a separate, fresh-grounded planning session posts the plan → `needs-operator` → "plan approved" → `ready`. Multi-deliverable work becomes an epic; each sub-issue is classified independently.
44
+
45
+ Scope calls are revisited through the one-way ratchet, whose rules and mechanics live in the `dev-plan` skill — the one home for upgrade/downgrade behavior.
46
+
47
+ ## Labels
48
+
49
+ State — exactly one per issue (creation colors live in dev-setup's labels row, their one home):
50
+
51
+ | label | meaning |
52
+ |---|---|
53
+ | `needs-operator` | waiting on the operator: a question, a brief or plan to approve, a proposal |
54
+ | `needs-plan` | brief approved; waiting for the planning stage (full-plan only) |
55
+ | `ready` | fully approved — an agent may start |
56
+ | `working` | claimed, in progress; the ledger comment shows live progress |
57
+ | `for-operator` | done — evidence posted, awaiting operator review |
58
+
59
+ Modifiers (may coexist with the state label): `risky` · scope `research` / `quick-build` / `full-plan` · `epic` (map parents, only where the org has no native Epic issue type).
60
+
61
+ ## Titles, types, hierarchy
62
+
63
+ - **Title prefixes** on issues, branches, and PRs identically: dev.md's `branch:` knob type list (that knob stays the list's one home) plus `research:` for research issues. PR title = issue title.
64
+ - **Native issue types** where the org defines them: Feature (feat) · Bug (fix) · Task (docs/chore/refactor/research) · Epic for parents (label fallback otherwise).
65
+ - **Hierarchy:** epic parent = map only (Destination · Decisions so far as one-line gists · Not clear yet · Out of scope), children attached as native sub-issues; issues = the unit of work (brief in description, own approvals/branch/PR/evidence); tasks = checkboxes **in the plan comment only**. Blockers use native issue dependencies; phases use milestones. Only issues — never epics — get `ready`. GitHub caps issue bodies and comments at ~65,536 characters; what a plan nearing that cap means is the `dev-plan` ratchet's call.
66
+
67
+ ## The ledger
68
+
69
+ Maintained by the implement session as one comment, edited in place:
70
+
71
+ ```markdown
72
+ <!-- vsk:v1 type=ledger branch=<branch> -->
73
+ ## Ledger — <branch>
74
+ - Task <N>: complete (commits <base7>..<head7>[, review clean | K parked])
75
+ - Task <N>: fix round <R>/3 (<X> addressed, <Y> open — <one-liners>; commits <a>..<b>)
76
+ - Ruling: <what> — <why> — cost if wrong: <cost>
77
+ - Task <N>: parked — <finding> — Ruling: <why the code stands>
78
+ - Deferred minor: <one-liner>
79
+ ```
80
+
81
+ **Resume protocol:** a fresh, compacted, or (operator-handed) takeover session reads, in order: the brief → the plan comment → the ledger → `git log` on the branch — **nothing else**. Tasks with a `complete` line are DONE, never re-executed; a task whose last line is a fix round resumes at the next round. After compaction, trust the ledger and `git log` over recollection. Every `Ruling:` line surfaces in the evidence comment — a ruling that dies with the session was a decision made in secret.
82
+
83
+ ## `.vegastack/.tmp/` workspace
84
+
85
+ All transitory artifacts — subagent reports, review packages, plan drafts, extracted diffs — live at `.vegastack/.tmp/<issue-number>-<title-slug>/` (pre-issue intake drafts, which have no number yet: `.vegastack/.tmp/intake-<slug>/`), kept out of git by a self-ignoring `.gitignore` (`printf '*\n' > .vegastack/.tmp/.gitignore`, created on first use). Subagents write full reports to files there and return only short status — a dead subagent's findings survive on disk, and the primary session never holds full reports in context. The workspace lives in the working tree (never under `.git/`, which harnesses protect from writes).
86
+
87
+ ## Verification gate
88
+
89
+ Before claiming any status: **IDENTIFY** the command that proves the claim → **RUN** it fresh and complete → **READ** the full output and exit code → only then claim, with the evidence. "Should pass", a previous run, or a subagent's say-so are never evidence. Guard scripts follow the same doctrine: machine-verifiable facts **block** (exit 2 with the reason); regex or judgment heuristics only **warn** — no AI inference inside guards, and an unverifiable state fails closed.
90
+
91
+ ## Plain-language collaboration
92
+
93
+ Every skill run ends with a simple-language summary: what happened, which paths were taken — cross-agent invocations announced at trigger time AND summarized at the end — and what is worth the operator double-checking. Use mermaid or ASCII diagrams in issues wherever a picture beats prose. A vague or self-contradicting operator answer gets pushback with concrete options, never silent absorption.
@@ -0,0 +1,3 @@
1
+ # Refresh contract — dev-chronicle
2
+
3
+ Evergreen: this skill asserts no version pins, vendor mechanisms, numeric limits, or dated claims — its content is a narrative format and a reading discipline, all versionless. Revisit if a future edit introduces a volatile fact.
@@ -0,0 +1,6 @@
1
+ {
2
+ "schemaVersion": 1,
3
+ "retrievalBaseline": "2026-08-28",
4
+ "note": "Evergreen waiver recorded in REFRESH.md; sources deliberately empty.",
5
+ "sources": []
6
+ }
@@ -0,0 +1,43 @@
1
+ ---
2
+ name: dev-debug
3
+ description: Reproduce-first bug work. Use when given a bug to fix — "debug this", "this is broken and I don't know why", "users report X fails", "login intermittently 500s", implementing a fix-type issue whose brief carries a Reproduction section, or when a fix keeps not fixing the symptom. Not for writing the bug up as an issue (dev-intake), building planned features (dev-implement — this skill governs the diagnosis inside its dark mode), or reviewing a finished fix (dev-review).
4
+ ---
5
+
6
+ # dev-debug
7
+
8
+ The failure this skill prevents: reading code, forming one theory, and "fixing" something that was never the cause. The discipline is a hard order — **reproduce, shrink, suspect, test, prove, clean** — and each phase has a completion criterion you can check, not vibe. It runs inside dev-implement's dark mode: no operator questions; missing-artifact stops are one `handback` comment; every phase result is a ledger checkpoint.
9
+
10
+ Nearest neighbors: `dev-intake`'s bug variant writes the brief this skill executes; `dev-implement` owns the surrounding build ceremony; `dev-review` judges the finished fix.
11
+
12
+ ## Phase 1 — the red command. No red command, no theorizing.
13
+
14
+ Build **one named command** that demonstrates the bug: it fails right now *because of this bug's exact symptom*, and will pass once it's truly fixed. The completion criterion, all four checkable:
15
+
16
+ - **Red-capable** — it asserts the reported symptom, not "runs without erroring"; you have run it at least once and its invocation + failing output (redacted) go in the ledger.
17
+ - **Deterministic** — same verdict every run; a flaky bug substitutes a pinned, stated reproduction rate ("~1 in 12 across 50 runs").
18
+ - **Fast** — seconds, not minutes; a tight loop is the whole superpower here.
19
+ - **Agent-runnable** — no human in the loop.
20
+
21
+ Pick the cheapest rung that reaches the bug from the [loop ladder](references/loop-ladder.md). Can't build one after walking the ladder → stop: one `handback` comment listing what was tried and asking for artifacts (logs, HAR, recording, environment access), `needs-operator`. Proceeding to theories without a red command is the exact failure this skill exists to prevent.
22
+
23
+ ## Phase 2 — shrink until everything left is load-bearing
24
+
25
+ Run the loop, watch it go red on the *reported* symptom (the wrong bug means the wrong fix). Then minimise: cut inputs, callers, config, and steps **one at a time**, re-running after each cut, keeping only what the failure needs. Done when removing any remaining element turns the loop green. The minimal repro shrinks the suspect space and becomes Phase 5's regression test.
26
+
27
+ ## Phase 3 — suspects: 3–5, ranked, falsifiable, posted, then GO
28
+
29
+ List 3–5 candidate causes ranked most-likely first — a single hypothesis anchors on the first plausible idea. Each must be **falsifiable**: "if X is the cause, then changing Y makes the bug disappear / Z makes it worse." A suspect whose prediction you can't state is a vibe — discard or sharpen it. **Post the ranked list to the ledger and proceed immediately** — never pause for the operator (their async re-rank is welcome whenever it comes; dark mode holds).
30
+
31
+ ## Phase 4 — test suspects one variable at a time
32
+
33
+ Every probe maps to one suspect's prediction. Prefer a debugger/REPL breakpoint over logs; when logging, target the boundaries that separate suspects — never "log everything and grep". **Every debug log carries a `[DEBUG-<4hex>]` tag** (one random tag per session): cleanup becomes a single grep, and ship-gate blocks any tag that survives into the diff. Performance bugs: logs lie — measure a baseline first (timing harness, profiler, query plan), then bisect.
34
+
35
+ ## Phase 5 — regression test before the fix
36
+
37
+ Write the failing test **before** touching the fix, at a **correct seam**: one where the test exercises the real bug pattern as it occurs at the call site. If the only reachable seam is too shallow to replicate the trigger, **that is itself a finding** — record it in the ledger and evidence; a false-confidence test is worse than a named gap. With a correct seam: minimal repro → failing test → watch it fail → fix → watch it pass → re-run the **original un-minimised** Phase 1 loop green.
38
+
39
+ ## Phase 6 — clean up and teach
40
+
41
+ Before hand-back, all checkable: the original repro re-runs green · `git diff <base>... | grep -F '[DEBUG-'` comes back empty (fixed-string grep; ship-gate backstops the added lines) · throwaway harnesses deleted · the **winning suspect and its evidence** named in the evidence comment and the commit message — the next debugger learns what it actually was, not just that it went away.
42
+
43
+ Close with the plain-language summary: the symptom, the cause, the proof, and anything the investigation surfaced that deserves its own issue.
@@ -0,0 +1,4 @@
1
+ interface:
2
+ display_name: "dev-debug"
3
+ short_description: "Reproduce-first bug diagnosis and fix"
4
+ default_prompt: "Use $dev-debug to find and fix this bug."
@@ -0,0 +1,93 @@
1
+ # Workflow conventions
2
+
3
+ The single spec for the artifacts every dev-family skill reads and writes. One home per rule: skills cite this file, never restate it. Everything here is harness-neutral.
4
+
5
+ ## Comment metadata markers
6
+
7
+ Every workflow-generated issue comment opens with an invisible HTML marker followed by a human heading:
8
+
9
+ ```markdown
10
+ <!-- vsk:v1 type=<type> rev=<n> [key=value ...] -->
11
+ ## <Human title> (v<n>)
12
+ ```
13
+
14
+ | type | required keys | instances |
15
+ |---|---|---|
16
+ | `approval` | `scope=<brief\|brief+plan\|plan>` | one per approval event |
17
+ | `plan` | `rev` | one, edited in place |
18
+ | `ledger` | `branch` | one, edited in place |
19
+ | `evidence` | `rev branch sha` | one, edited in place |
20
+ | `review` | `round sha agent=<claude\|codex> verdict=<clean\|needs-fixes>` | one per review cycle, rounds appended inside |
21
+ | `decision` | — | one per decision proposal |
22
+ | `handback` | — | one per stop event |
23
+
24
+ `rev=<n>` and the matching `(v<n>)` heading suffix appear only on revisable artifacts — the brief (issue description), `plan`, and `evidence` — starting at `rev=1`/`(v1)`. Single-event comments (`approval`, `decision`, `handback`) and the `ledger` carry neither. Scripts and agents locate comments strictly by marker, never by heading text. A comment without its marker does not count as the artifact — there is no legacy fallback.
25
+
26
+ ## Operator identity
27
+
28
+ Every human reference in every artifact — approvals, revisions, decisions, changelog attributions, review adjudications — is written `operator (<github-username>)`:
29
+
30
+ - Approval: `Approved by operator (<username>) on DD-MM-YYYY: "<their words>"`
31
+ - Register line: `- DD-MM-YYYY operator (<username>) — <decision>`
32
+
33
+ ## Revision markers
34
+
35
+ Any artifact edited after its first approval: the heading gains `(v2)`, the marker gains `rev=2`, and a `Revisions:` line is appended at the bottom — `v2 — DD-MM-YYYY: <what changed>, per operator (<username>) correction`. Existing revision lines are never rewritten.
36
+
37
+ ## Scope classes
38
+
39
+ Set at intake, applied as a label, announced with its reason (operator can override):
40
+
41
+ - **`research`** — a question to answer; throwaway code allowed, never merged. No branch/PR/changelog; findings + recommendation are the evidence comment.
42
+ - **`quick-build`** — small change and the flow being changed already exists in the repo to read. Brief (description) + plan (comment) are drafted in the same conversation; **one approval covers both**; then straight to `ready`.
43
+ - **`full-plan`** — big or new ground. Brief approval → `needs-plan` → a separate, fresh-grounded planning session posts the plan → `needs-operator` → "plan approved" → `ready`. Multi-deliverable work becomes an epic; each sub-issue is classified independently.
44
+
45
+ Scope calls are revisited through the one-way ratchet, whose rules and mechanics live in the `dev-plan` skill — the one home for upgrade/downgrade behavior.
46
+
47
+ ## Labels
48
+
49
+ State — exactly one per issue (creation colors live in dev-setup's labels row, their one home):
50
+
51
+ | label | meaning |
52
+ |---|---|
53
+ | `needs-operator` | waiting on the operator: a question, a brief or plan to approve, a proposal |
54
+ | `needs-plan` | brief approved; waiting for the planning stage (full-plan only) |
55
+ | `ready` | fully approved — an agent may start |
56
+ | `working` | claimed, in progress; the ledger comment shows live progress |
57
+ | `for-operator` | done — evidence posted, awaiting operator review |
58
+
59
+ Modifiers (may coexist with the state label): `risky` · scope `research` / `quick-build` / `full-plan` · `epic` (map parents, only where the org has no native Epic issue type).
60
+
61
+ ## Titles, types, hierarchy
62
+
63
+ - **Title prefixes** on issues, branches, and PRs identically: dev.md's `branch:` knob type list (that knob stays the list's one home) plus `research:` for research issues. PR title = issue title.
64
+ - **Native issue types** where the org defines them: Feature (feat) · Bug (fix) · Task (docs/chore/refactor/research) · Epic for parents (label fallback otherwise).
65
+ - **Hierarchy:** epic parent = map only (Destination · Decisions so far as one-line gists · Not clear yet · Out of scope), children attached as native sub-issues; issues = the unit of work (brief in description, own approvals/branch/PR/evidence); tasks = checkboxes **in the plan comment only**. Blockers use native issue dependencies; phases use milestones. Only issues — never epics — get `ready`. GitHub caps issue bodies and comments at ~65,536 characters; what a plan nearing that cap means is the `dev-plan` ratchet's call.
66
+
67
+ ## The ledger
68
+
69
+ Maintained by the implement session as one comment, edited in place:
70
+
71
+ ```markdown
72
+ <!-- vsk:v1 type=ledger branch=<branch> -->
73
+ ## Ledger — <branch>
74
+ - Task <N>: complete (commits <base7>..<head7>[, review clean | K parked])
75
+ - Task <N>: fix round <R>/3 (<X> addressed, <Y> open — <one-liners>; commits <a>..<b>)
76
+ - Ruling: <what> — <why> — cost if wrong: <cost>
77
+ - Task <N>: parked — <finding> — Ruling: <why the code stands>
78
+ - Deferred minor: <one-liner>
79
+ ```
80
+
81
+ **Resume protocol:** a fresh, compacted, or (operator-handed) takeover session reads, in order: the brief → the plan comment → the ledger → `git log` on the branch — **nothing else**. Tasks with a `complete` line are DONE, never re-executed; a task whose last line is a fix round resumes at the next round. After compaction, trust the ledger and `git log` over recollection. Every `Ruling:` line surfaces in the evidence comment — a ruling that dies with the session was a decision made in secret.
82
+
83
+ ## `.vegastack/.tmp/` workspace
84
+
85
+ All transitory artifacts — subagent reports, review packages, plan drafts, extracted diffs — live at `.vegastack/.tmp/<issue-number>-<title-slug>/` (pre-issue intake drafts, which have no number yet: `.vegastack/.tmp/intake-<slug>/`), kept out of git by a self-ignoring `.gitignore` (`printf '*\n' > .vegastack/.tmp/.gitignore`, created on first use). Subagents write full reports to files there and return only short status — a dead subagent's findings survive on disk, and the primary session never holds full reports in context. The workspace lives in the working tree (never under `.git/`, which harnesses protect from writes).
86
+
87
+ ## Verification gate
88
+
89
+ Before claiming any status: **IDENTIFY** the command that proves the claim → **RUN** it fresh and complete → **READ** the full output and exit code → only then claim, with the evidence. "Should pass", a previous run, or a subagent's say-so are never evidence. Guard scripts follow the same doctrine: machine-verifiable facts **block** (exit 2 with the reason); regex or judgment heuristics only **warn** — no AI inference inside guards, and an unverifiable state fails closed.
90
+
91
+ ## Plain-language collaboration
92
+
93
+ Every skill run ends with a simple-language summary: what happened, which paths were taken — cross-agent invocations announced at trigger time AND summarized at the end — and what is worth the operator double-checking. Use mermaid or ASCII diagrams in issues wherever a picture beats prose. A vague or self-contradicting operator answer gets pushback with concrete options, never silent absorption.
@@ -0,0 +1,20 @@
1
+ # The loop ladder
2
+
3
+ How to build the red command, cheapest rung that reaches the bug first. The goal is always the same: symptom in, verdict out, seconds per run, no human.
4
+
5
+ 1. **Failing test** at whatever seam reaches the bug — unit, integration, or e2e. First choice when the code path is importable.
6
+ 2. **Curl / HTTP script** against a running dev server, asserting on status/body — for bugs that live in the request path.
7
+ 3. **CLI invocation with a fixture input**, diffing stdout/exit code against known-good — for tools and scripts.
8
+ 4. **Headless browser script** (Playwright or equivalent) driving the UI and asserting on DOM, console, or network — when the bug needs a real browser.
9
+ 5. **Captured-trace replay** — save a real request/payload/event log once, replay it through the code path in isolation; turns "only happens with production data" into a loop.
10
+ 6. **Throwaway harness** — boot the minimal slice of the system (one service, mocked deps) that exercises the path with a single call. Deleted in Phase 6.
11
+ 7. **Property/fuzz loop** — for "sometimes wrong output": run hundreds of random inputs and trap the failure mode; the trapped case seeds the minimised repro.
12
+ 8. **Bisection harness** — when the bug appeared between two known states (commit, dataset, version): automate "boot at state X, check, report" so `git bisect run` can drive it.
13
+
14
+ ## Tightening
15
+
16
+ A loop earns its keep on three axes — make it **faster** (mock the slow dependency, cache the boot), **sharper** (assert the exact symptom, not a proxy), **more deterministic** (pin the clock, the seed, the ordering). A 30-second flaky loop is barely better than none; a 2-second deterministic one changes what's possible.
17
+
18
+ ## When no rung works
19
+
20
+ That's a stop, not a license to guess: `handback` with the rungs tried, why each failed, and the specific artifact or access that would unlock one (a HAR of the failing request, server logs around the timestamp, a screen recording, an environment credential). The operator trades one artifact for a loop; nobody trades theories for luck.
@@ -0,0 +1,3 @@
1
+ # Refresh contract — dev-debug
2
+
3
+ Evergreen: this skill asserts no version pins, vendor mechanisms, numeric limits, or dated claims — its content is diagnosis discipline (the phase order, the red-command criterion, the ladder), all versionless. Revisit if a future edit introduces a volatile fact.
@@ -0,0 +1,6 @@
1
+ {
2
+ "schemaVersion": 1,
3
+ "retrievalBaseline": "2026-08-28",
4
+ "note": "Evergreen waiver recorded in REFRESH.md; sources deliberately empty.",
5
+ "sources": []
6
+ }
@@ -1,70 +1,75 @@
1
1
  ---
2
2
  name: dev-implement
3
- description: Implement an approved GitHub issue end to end without further user input. Use when given an issue to build - "do issue 12", "implement" plus an issue URL or number, "pick up the next ready issue", "go dark on" an issue - when returning to apply corrections the user left on a for-operator issue, or when the user directly asks in chat for a quick fix or small change. Runs preflight, claims the issue, builds on a task branch, tests, gets independent review, and posts one evidence comment in the issue. Not for writing or approving issues (dev-intake), not for creating PRs or merging (dev-ship).
3
+ description: Implement an approved GitHub issue end to end without further user input. Use when given an issue to build "do issue 12", "implement" plus an issue URL or number, "pick up the next ready issue", "go dark on" an issue when resuming a dead or compacted session's working issue the operator hands over, when returning to apply corrections the user left on a for-operator issue, or when the user directly asks in chat for a small fix. Not for writing or approving issues (dev-intake), planning them (dev-plan), reviewing finished work (dev-review), or creating PRs and merging (dev-ship).
4
4
  ---
5
5
 
6
6
  # dev-implement
7
7
 
8
- One issue, one session, end to end: preflight → claim → build dark → verify → review → evidence in the issue → stop. The user reads the result in the issue on their own time; nothing here creates a PR or merges — those are `dev-ship`, on the user's word.
8
+ One issue, one session, end to end: preflight → claim → build dark → verify → review → evidence → stop. The operator reads the result in the issue on their own time; nothing here creates a PR or merges — those are `dev-ship`, on the operator's word. Artifact formats follow the `dev-setup` skill's `references/conventions.md`; the ledger discipline lives in [ledger-and-resume](references/ledger-and-resume.md).
9
9
 
10
- Nearest neighbor: `dev-intake` writes the brief this skill executes; if the issue turns out to need decisions, that's intake work — hand it back via `needs-operator`, don't guess. `.vegastack/dev.md` missing → run `dev-setup` first. Read dev.md before anything; its knobs (review, ui-evidence, tests, branch, stop-list) govern this whole skill.
10
+ Nearest neighbors: `dev-plan` writes the plan this skill executes task by task; `dev-review` judges the result; issues that turn out to need decisions go back through `needs-operator`, never guessed. `.vegastack/dev.md` missing → run `dev-setup` first. Read dev.md before anything; its knobs govern this skill, and the `## Architecture` section governs stack-touching choices.
11
11
 
12
- ## Direct requests
12
+ ## Direct requests — trivial only, tightly bounded
13
13
 
14
- The gates exist to stop agent-invented authority, never to slow the user down. When the user directly asks in chat for a change ("fix this typo", "bump that timeout"), their words are the approval — do it, verify it, and report; no issue required. Branch as `<type>/<slug>` (no issue segment); the changelog rule below applies unchanged when the change is behavior-changing; shipping still goes through dev-ship's words, with the chat request standing in for the approval and the report for the evidence comment. Offer to record an issue when the change is material enough that its brief or evidence will matter later. Everything below is the path for issue-driven work.
14
+ When the operator directly asks in chat for a change, their words are the approval — build, verify, report; no issue needed. The bound is **trivial**: the moment the work exceeds it a behavior change beyond the asked words, a new dependency, more than 1–2 files stop and route to `dev-intake` instead of continuing. Branch `<type>/<slug>`; the changelog and chronicle rules apply unchanged when behavior changes; shipping still goes through dev-ship's words.
15
15
 
16
16
  ## Preflight — all must hold, or stop and say which failed
17
17
 
18
- - `gh auth status` works and the issue's repo matches dev.md.
19
- - The issue is open, labeled `ready`, and carries the recorded approval comment (`Approved by : "…"`). A label without the comment is not approval.
20
- - No open blockers (issue dependencies) and no other assignee an assigned or `working` issue belongs to someone else. A claim from a dead session is released only by the user: take over a `working` issue only when they explicitly hand it to you.
21
- - Read the complete brief, plus parent issue and milestone for context. If the brief leaves a material decision open — including an unresolved Assumptions entry — do not start: label `needs-operator`, comment the smallest question that unblocks it, stop.
22
- - Re-verify the brief against reality before coding: its cited touch points against the current code (things drift between approval and execution), and volatile dependency claims when stale or version-sensitive. Reality contradicting the brief is a stop — label `needs-operator` with the discrepancy; an approved brief is never a license to improvise past what's actually there.
18
+ - Run the deterministic guard first: `node <path-to-this-skill>/scripts/preflight.mjs --issue <n> --me $(gh api user -q .login) --json` (add `--repo <o/r> --dev-md <path>` when running outside the project root) — exit 2 stops you with its reasons (open + `ready` state, approval marker, scope label, plan approval on full-plan, Assumptions section, blockers, assignee, repo match); exit 1 passes with warnings — read them into the ledger. Resume and corrections runs pass `--expect working` / `--expect for-operator`.
19
+ - Then the judgment checks: read the complete brief plus parent issue and milestone for context; re-verify the brief's touch points against the current code (things drift between approval and execution) — including the version-impact line; volatile dependency claims per `dev-architect`'s verify protocol; a full-plan issue's plan still matches reality. A material decision left open — even outside a formal Assumptions section — or reality contradicting brief or plan is a stop: one `handback` comment with the smallest question, `needs-operator`.
20
+ - Resuming a dead session's issue: the operator's explicit handover is required; then follow the resume protocol in [ledger-and-resume](references/ledger-and-resume.md) brief plan ledger `git log`, nothing else.
23
21
 
24
- ## Claim and branch
22
+ ## Claim
25
23
 
26
- Assign yourself, swap `ready` → `working`. Branch from the default branch per dev.md's `branch:` knob the knob is the only source of the pattern and its type list.
24
+ Assign yourself, swap `ready` → `working`, branch from the default branch per dev.md's `branch:` knob, and **create the ledger comment as your first write**. Record each task's base sha before starting it.
27
25
 
28
- ## Build — dark
26
+ ## Build — dark, test-first, checkpointed
29
27
 
30
- No progress updates, no questions. A spike the brief flagged runs first — its result opens the evidence comment and shapes the rest of the build. Decide routine things yourself: file layout, helpers, fixtures, and root-cause fixes inside the issue's change areas. The brief's out-of-scope section and the dev.md stop-list bound you; hitting a stop condition (scope change, new dependency, spending, destructive/production action, unresolvable blocker) ends dark mode — post one `needs-operator` comment stating the smallest decision needed with your recommendation, and stop.
28
+ No progress updates, no questions. A `fix:` issue's diagnosis runs under the `dev-debug` skill its phases govern the investigation inside this dark mode, and its winning suspect feeds the evidence comment. A spike the brief flagged runs first its result opens the evidence comment and shapes the rest of the build. Then work the plan task by task:
29
+
30
+ - **Red before green.** Write the failing test first — at the seams the brief names, never elsewhere — watch it fail for the stated reason, implement the minimal code, watch it pass. One slice at a time. The tests-are-real rubric (implementation-coupled, tautological, horizontal-sliced — defined in `dev-review`'s dispatch prompts) applies to your own tests before a reviewer ever sees them.
31
+ - **Checkpoint the ledger** after every task (tick the plan checkbox in the same pass) and at every ruling, per the reference.
32
+ - **Transitory artifacts** — subagent reports, scratch diffs, drafts — live in `.vegastack/.tmp/<issue>-<slug>/`; subagents write full output to files there and return short status.
33
+ - Decide routine things yourself and ledger the rulings. A structural choice mid-build — a new dependency, table, or service — checks `dev-architect`'s trigger discipline first; a moving part with no named trigger is a stop condition.
34
+ - **The scope ratchet is a stop condition:** work revealed bigger than the issue's scope class (or plainly exceeding one session) → one `handback` comment proposing the upgrade or split (dev-plan's ratchet rules), `needs-operator`, stop.
35
+ - The brief's out-of-scope section and the dev.md stop-list bound you; hitting any stop condition ends dark mode with one `handback` comment stating the smallest decision needed, your recommendation attached.
31
36
 
32
37
  Honesty over green: a failing test gets fixed at the root or reported as failing. Weakening a test, an assertion, or acceptance to pass is a cover-up, and cover-ups surface at review with interest.
33
38
 
34
- ## Changelog — before hand-back
39
+ ## Changelog and chronicle — before hand-back
40
+
41
+ Every behavior-changing branch carries its changelog entry per dev.md's `changelog:` knob (`changesets` → write `.changeset/<slug>.md` directly, never the interactive CLI; `keep-a-changelog` / `pubspec+changelog` → one bullet under `## [Unreleased]`, creating CHANGELOG.md with the `# Changelog` + `## [Unreleased]` skeleton in the same branch when absent; `none` → skip) **and**, when dev.md says `chronicle: on`, its story entry prepended to `.vegastack/chronicle.md` (format: the `dev-chronicle` skill) — both on the branch, landing atomically with the merge. Docs the brief names as affected get updated in the same branch.
35
42
 
36
- Every behavior-changing branch carries its changelog entry per dev.md's `changelog:` knob; a ship-time guard catches misses, but don't rely on it. `changesets` → write `.changeset/<slug>.md` directly (frontmatter `"<package-name>": <bump>` from the brief's version-impact line, plus a one-paragraph summary — the changeset CLI prompt is interactive, never invoke it here). `keep-a-changelog` / `pubspec+changelog` → one bullet under `## [Unreleased]` (Added/Changed/Fixed/Removed subsection as fits); no CHANGELOG.md yet → create it in the same branch with the skeleton (`# Changelog` + `## [Unreleased]`). `none` → skip. Docs the brief names as affected get updated in the same branch.
43
+ ## Verify — the gate function
37
44
 
38
- ## Verify
45
+ Before claiming ANY status: **identify** the command that proves it → **run** it fresh and complete → **read** the full output and exit code → only then claim, with the evidence. Tests pass ⇒ a fresh run with 0 failures — never "should pass", never a previous run. Build succeeds ⇒ exit 0. Bug fixed ⇒ the original symptom re-tested. A subagent finished ⇒ you inspected its diff or report file — never its say-so.
39
46
 
40
- - Run the tests dev.md requires (`tests: required` every changed behavior has a test that runs and passes; `logic-only` content/config tweaks may skip). Record commands and results for the evidence comment.
41
- - A `risky` issue gets focused security, failure, and recovery checks on top of the required tests.
42
- - When dev.md has a `## Verify` runbook, follow it run the app and smoke-check the flows it names; that live result belongs in the evidence comment alongside the test output. Verify is pre-merge only; post-release checks live in `## Ship` and belong to dev-ship.
43
- - UI changed and `ui-evidence: playwright` → capture screenshots of the key states and flows and upload them to the shared evidence repo (dev.md `evidence-repo`) under `<this-repo-name>/<issue-number>/<timestamp>-<name>.png` — via the contents API so the repo is never cloned: `base64 < <file> | tr -d '\n' | gh api -X PUT repos/<evidence-repo>/contents/<path> -f message="evidence #<issue>" -F content=@-` (piped stdin, so large screenshots never hit argv limits; timestamped names keep re-captures from colliding; a 409 from a concurrent upload just means retry). Link them in the evidence comment — links, not embeds; private-repo images don't render inline in issues. Evidence repo missing or unreachable → name the local file paths and say so; the hand-back never blocks on it.
44
- - dev.md's Ship or Verify section is an empty TODO while release/deploy machinery visibly exists → finish this issue normally, then suggest re-running dev-setup so detection can fill it.
47
+ - Run what dev.md's `tests:` knob requires; a `risky` issue gets focused security, failure, and recovery checks on top. When dev.md has a `## Verify` runbook, run the app and smoke-check the flows it names. Verify is pre-merge only; post-release checks live in `## Ship` and belong to dev-ship.
48
+ - UI changed and `ui-evidence: playwright` → capture screenshots of the key states and upload to the shared evidence repo (dev.md `evidence-repo`) under `<this-repo-name>/<issue-number>/<timestamp>-<name>.png` via the contents API, never a clone: `base64 < <file> | tr -d '\n' | gh api -X PUT repos/<evidence-repo>/contents/<path> -f message="evidence #<issue>" -F content=@-` (piped stdin so large screenshots never hit argv limits; timestamped names avoid collisions; a 409 from a concurrent upload just means retry). Link them in the evidence comment — links, never embeds (private-repo images don't render inline). Evidence repo unreachable → name local paths and say so; the hand-back never blocks on it.
49
+ - dev.md's Ship or Verify section is an empty TODO next to visible machinery finish normally, then suggest re-running dev-setup.
45
50
 
46
- ## Independent review — per the dev.md knob
51
+ ## Independent review — invoke dev-review
47
52
 
48
- - `subagent` (default): spawn a fresh reviewer subagent that gets the diff, the brief, and dev.md and no memory of writing the code. It checks: does the change do what the brief says, does anything break, are the tests real? In a harness without subagents, do a separate fresh-eyes review pass against the brief and label it a self-review in the evidence comment; prefer cross-agent there for `risky` work.
49
- - `cross-agent` (or `cross-agent-risky` on a `risky` issue): push the branch, add to the evidence comment "awaiting cross-agent review", keep `working`, and tell the user which agent to point at the issue. The reviewing session posts findings on the issue; you apply them.
50
- - Fix real findings and rerun affected checks. Disagree with a finding → say why in the evidence comment rather than silently skipping it.
53
+ Run the `dev-review` skill per dev.md's `review:` knob — fresh subagent axes by default, cross-agent (Codex↔Claude, announced to the operator) per the knob's mapping; it owns the axes, severities, review comment, bounded fix loop, and adjudication rules. Apply its findings through its loop and re-run the affected checks. Disagree with a finding adjudicate openly per its rules, never silently skip. In a harness without subagents, run dev-review's axis briefs yourself as a labeled self-review.
51
54
 
52
55
  ## The evidence comment — exactly one, edited in place
53
56
 
54
- ```
55
- ## Result
57
+ ```markdown
58
+ <!-- vsk:v1 type=evidence rev=1 branch=<name> sha=<sha7> -->
59
+ ## Result (v1)
56
60
  **Done:** what changed, in behavior terms
57
- **Tests:** <command> → <result summary>
58
- **Review:** <mode> — <findings fixed / none / disputed with reason>
59
- **Changelog:** <entry added / none, with reason> (when the knob is not `none`)
61
+ **Tests:** <command> → <fresh result>
62
+ **Review:** <mode> — <verdict; adjudications and rulings surfaced, in order made>
63
+ **Changelog:** <entry added / none, with reason>
64
+ **Docs:** brief v<n>, plan v<n> — in sync | unchanged since approval
60
65
  **UI evidence:** <links> (when applicable)
61
- **Decision:** <one line in the register format> (only a dark-mode choice that passes dev.md's Decisions test — a proposal; dev-ship records it after naming it in the merge confirmation)
66
+ **Decision:** <register-format proposals> (only choices passing dev.md's Decisions test)
62
67
  **Not done / limits:** the honest list
63
- Branch: <name> @ <short-sha>
68
+ Branch: <name> @ <sha7>
64
69
  ```
65
70
 
66
- Post it, swap `working` → `for-operator`, unassign nothing, stop. Later corrections update this same comment a stack of stale result comments hides the current truth.
71
+ The `**Review:**` line is the one home of surfaced rulings: every ledger `Ruling:` appears there, in the order made. Run `node <path-to-this-skill>/scripts/evidence-check.mjs --file <draft> --json` before posting — exit 2 means the shape is incomplete; fix, don't post. Post it, swap `working` → `for-operator`, stop, and close with the plain-language summary (which repeats, never replaces, the evidence content): what was built, which paths were taken, the rulings, what's worth the operator double-checking.
67
72
 
68
- ## Corrections loop
73
+ ## Corrections loop — code and docs move together
69
74
 
70
- The user's comments on a `for-operator` issue are the new frontier: apply them, re-verify what they touch, update the evidence comment, back to `for-operator`. Their corrections never need re-approval ceremony unless they change scope — then it's `needs-operator` and dev-intake's recording rule.
75
+ The operator's comments on a `for-operator` issue are the new frontier. Applying a correction is **one pass**: the code change + the affected brief/plan sections edited to match (revision markers bumped, `Revisions:` line appended) + a ledger line + the evidence comment updated in place — its `sha` to the new head and its `Docs:` line to the new revisions. Re-verify what the correction touched. An operator dismissal of a review finding gets appended to `.vegastack/review-known-patterns.md` with its mandatory "Still flag if:" clause. Then back to `for-operator`. Corrections never need re-approval ceremony unless they change scope — that's `needs-operator` and intake's recording rule.
@@ -0,0 +1,93 @@
1
+ # Workflow conventions
2
+
3
+ The single spec for the artifacts every dev-family skill reads and writes. One home per rule: skills cite this file, never restate it. Everything here is harness-neutral.
4
+
5
+ ## Comment metadata markers
6
+
7
+ Every workflow-generated issue comment opens with an invisible HTML marker followed by a human heading:
8
+
9
+ ```markdown
10
+ <!-- vsk:v1 type=<type> rev=<n> [key=value ...] -->
11
+ ## <Human title> (v<n>)
12
+ ```
13
+
14
+ | type | required keys | instances |
15
+ |---|---|---|
16
+ | `approval` | `scope=<brief\|brief+plan\|plan>` | one per approval event |
17
+ | `plan` | `rev` | one, edited in place |
18
+ | `ledger` | `branch` | one, edited in place |
19
+ | `evidence` | `rev branch sha` | one, edited in place |
20
+ | `review` | `round sha agent=<claude\|codex> verdict=<clean\|needs-fixes>` | one per review cycle, rounds appended inside |
21
+ | `decision` | — | one per decision proposal |
22
+ | `handback` | — | one per stop event |
23
+
24
+ `rev=<n>` and the matching `(v<n>)` heading suffix appear only on revisable artifacts — the brief (issue description), `plan`, and `evidence` — starting at `rev=1`/`(v1)`. Single-event comments (`approval`, `decision`, `handback`) and the `ledger` carry neither. Scripts and agents locate comments strictly by marker, never by heading text. A comment without its marker does not count as the artifact — there is no legacy fallback.
25
+
26
+ ## Operator identity
27
+
28
+ Every human reference in every artifact — approvals, revisions, decisions, changelog attributions, review adjudications — is written `operator (<github-username>)`:
29
+
30
+ - Approval: `Approved by operator (<username>) on DD-MM-YYYY: "<their words>"`
31
+ - Register line: `- DD-MM-YYYY operator (<username>) — <decision>`
32
+
33
+ ## Revision markers
34
+
35
+ Any artifact edited after its first approval: the heading gains `(v2)`, the marker gains `rev=2`, and a `Revisions:` line is appended at the bottom — `v2 — DD-MM-YYYY: <what changed>, per operator (<username>) correction`. Existing revision lines are never rewritten.
36
+
37
+ ## Scope classes
38
+
39
+ Set at intake, applied as a label, announced with its reason (operator can override):
40
+
41
+ - **`research`** — a question to answer; throwaway code allowed, never merged. No branch/PR/changelog; findings + recommendation are the evidence comment.
42
+ - **`quick-build`** — small change and the flow being changed already exists in the repo to read. Brief (description) + plan (comment) are drafted in the same conversation; **one approval covers both**; then straight to `ready`.
43
+ - **`full-plan`** — big or new ground. Brief approval → `needs-plan` → a separate, fresh-grounded planning session posts the plan → `needs-operator` → "plan approved" → `ready`. Multi-deliverable work becomes an epic; each sub-issue is classified independently.
44
+
45
+ Scope calls are revisited through the one-way ratchet, whose rules and mechanics live in the `dev-plan` skill — the one home for upgrade/downgrade behavior.
46
+
47
+ ## Labels
48
+
49
+ State — exactly one per issue (creation colors live in dev-setup's labels row, their one home):
50
+
51
+ | label | meaning |
52
+ |---|---|
53
+ | `needs-operator` | waiting on the operator: a question, a brief or plan to approve, a proposal |
54
+ | `needs-plan` | brief approved; waiting for the planning stage (full-plan only) |
55
+ | `ready` | fully approved — an agent may start |
56
+ | `working` | claimed, in progress; the ledger comment shows live progress |
57
+ | `for-operator` | done — evidence posted, awaiting operator review |
58
+
59
+ Modifiers (may coexist with the state label): `risky` · scope `research` / `quick-build` / `full-plan` · `epic` (map parents, only where the org has no native Epic issue type).
60
+
61
+ ## Titles, types, hierarchy
62
+
63
+ - **Title prefixes** on issues, branches, and PRs identically: dev.md's `branch:` knob type list (that knob stays the list's one home) plus `research:` for research issues. PR title = issue title.
64
+ - **Native issue types** where the org defines them: Feature (feat) · Bug (fix) · Task (docs/chore/refactor/research) · Epic for parents (label fallback otherwise).
65
+ - **Hierarchy:** epic parent = map only (Destination · Decisions so far as one-line gists · Not clear yet · Out of scope), children attached as native sub-issues; issues = the unit of work (brief in description, own approvals/branch/PR/evidence); tasks = checkboxes **in the plan comment only**. Blockers use native issue dependencies; phases use milestones. Only issues — never epics — get `ready`. GitHub caps issue bodies and comments at ~65,536 characters; what a plan nearing that cap means is the `dev-plan` ratchet's call.
66
+
67
+ ## The ledger
68
+
69
+ Maintained by the implement session as one comment, edited in place:
70
+
71
+ ```markdown
72
+ <!-- vsk:v1 type=ledger branch=<branch> -->
73
+ ## Ledger — <branch>
74
+ - Task <N>: complete (commits <base7>..<head7>[, review clean | K parked])
75
+ - Task <N>: fix round <R>/3 (<X> addressed, <Y> open — <one-liners>; commits <a>..<b>)
76
+ - Ruling: <what> — <why> — cost if wrong: <cost>
77
+ - Task <N>: parked — <finding> — Ruling: <why the code stands>
78
+ - Deferred minor: <one-liner>
79
+ ```
80
+
81
+ **Resume protocol:** a fresh, compacted, or (operator-handed) takeover session reads, in order: the brief → the plan comment → the ledger → `git log` on the branch — **nothing else**. Tasks with a `complete` line are DONE, never re-executed; a task whose last line is a fix round resumes at the next round. After compaction, trust the ledger and `git log` over recollection. Every `Ruling:` line surfaces in the evidence comment — a ruling that dies with the session was a decision made in secret.
82
+
83
+ ## `.vegastack/.tmp/` workspace
84
+
85
+ All transitory artifacts — subagent reports, review packages, plan drafts, extracted diffs — live at `.vegastack/.tmp/<issue-number>-<title-slug>/` (pre-issue intake drafts, which have no number yet: `.vegastack/.tmp/intake-<slug>/`), kept out of git by a self-ignoring `.gitignore` (`printf '*\n' > .vegastack/.tmp/.gitignore`, created on first use). Subagents write full reports to files there and return only short status — a dead subagent's findings survive on disk, and the primary session never holds full reports in context. The workspace lives in the working tree (never under `.git/`, which harnesses protect from writes).
86
+
87
+ ## Verification gate
88
+
89
+ Before claiming any status: **IDENTIFY** the command that proves the claim → **RUN** it fresh and complete → **READ** the full output and exit code → only then claim, with the evidence. "Should pass", a previous run, or a subagent's say-so are never evidence. Guard scripts follow the same doctrine: machine-verifiable facts **block** (exit 2 with the reason); regex or judgment heuristics only **warn** — no AI inference inside guards, and an unverifiable state fails closed.
90
+
91
+ ## Plain-language collaboration
92
+
93
+ Every skill run ends with a simple-language summary: what happened, which paths were taken — cross-agent invocations announced at trigger time AND summarized at the end — and what is worth the operator double-checking. Use mermaid or ASCII diagrams in issues wherever a picture beats prose. A vague or self-contradicting operator answer gets pushback with concrete options, never silent absorption.