tech-lead-stack 1.0.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (123) hide show
  1. package/.agents/hr-workflows/hr-ad-distributor.md +18 -0
  2. package/.agents/hr-workflows/hr-candidate-sourcer.md +18 -0
  3. package/.agents/hr-workflows/hr-endorsement-synthesizer.md +18 -0
  4. package/.agents/hr-workflows/hr-intake-specifier.md +18 -0
  5. package/.agents/hr-workflows/hr-interview-auditor.md +18 -0
  6. package/.agents/hr-workflows/hr-jd-drafter.md +18 -0
  7. package/.agents/hr-workflows/hr-pipeline-translator.md +18 -0
  8. package/.agents/pm-workflows/pm-action-item-mapper.md +18 -0
  9. package/.agents/pm-workflows/pm-backlog-auditor.md +18 -0
  10. package/.agents/pm-workflows/pm-context-summarizer.md +18 -0
  11. package/.agents/pm-workflows/pm-design-system-auditor.md +18 -0
  12. package/.agents/pm-workflows/pm-effort-estimator.md +18 -0
  13. package/.agents/pm-workflows/pm-newsletter-generator.md +18 -0
  14. package/.agents/pm-workflows/pm-progress-translator.md +18 -0
  15. package/.agents/pm-workflows/pm-release-note-drafter.md +18 -0
  16. package/.agents/pm-workflows/pm-risk-detector.md +18 -0
  17. package/.agents/pm-workflows/pm-story-augmenter.md +18 -0
  18. package/.agents/pm-workflows/pm-task-specifier.md +18 -0
  19. package/.agents/workflows/accessibility-audit.md +30 -0
  20. package/.agents/workflows/ask.md +44 -0
  21. package/.agents/workflows/audit-tech-debt.md +31 -0
  22. package/.agents/workflows/changelog.md +31 -0
  23. package/.agents/workflows/clean-code-audit.md +31 -0
  24. package/.agents/workflows/code-review.md +38 -0
  25. package/.agents/workflows/competitive-analysis.md +46 -0
  26. package/.agents/workflows/design-requirements-to-architecture.md +31 -0
  27. package/.agents/workflows/design-system-review.md +113 -0
  28. package/.agents/workflows/dev-team-sub-max.md +57 -0
  29. package/.agents/workflows/dev-team-sub-pro.md +57 -0
  30. package/.agents/workflows/dev-team.md +52 -0
  31. package/.agents/workflows/feature-orchestrator.md +43 -0
  32. package/.agents/workflows/init.md +31 -0
  33. package/.agents/workflows/mission-architect.md +31 -0
  34. package/.agents/workflows/onboard-dev.md +31 -0
  35. package/.agents/workflows/plan-quick.md +33 -0
  36. package/.agents/workflows/plan.md +31 -0
  37. package/.agents/workflows/pr-automator.md +44 -0
  38. package/.agents/workflows/pr-design-review-init.md +57 -0
  39. package/.agents/workflows/qa-handover.md +40 -0
  40. package/.agents/workflows/reflexion-loop-sub-max.md +46 -0
  41. package/.agents/workflows/reflexion-loop-sub-pro.md +45 -0
  42. package/.agents/workflows/reflexion-loop.md +66 -0
  43. package/.agents/workflows/regression-bug-fix.md +31 -0
  44. package/.agents/workflows/security-audit.md +31 -0
  45. package/.agents/workflows/standup-daily-summary.md +31 -0
  46. package/.agents/workflows/strategy-target-evaluation.md +31 -0
  47. package/.agents/workflows/style-logic-exporter.md +84 -0
  48. package/.agents/workflows/ui-spec-generator.md +156 -0
  49. package/.agents/workflows/verify-changes.md +31 -0
  50. package/.agents/workflows/vertical-slice.md +52 -0
  51. package/.agents/workflows/weekly-leadership-report.md +39 -0
  52. package/.ai/agent-surfaces.json +1235 -0
  53. package/.ai/hooks/README.md +32 -0
  54. package/.ai/hooks/build-requires-approved-spec.json +10 -0
  55. package/.ai/hooks/deploy-requires-review.json +11 -0
  56. package/.ai/hooks/no-ai-approve-deploy.json +10 -0
  57. package/.ai/hooks/protected-paths.json +10 -0
  58. package/.ai/hr-skills/hr-ad-distributor.md +61 -0
  59. package/.ai/hr-skills/hr-candidate-sourcer.md +69 -0
  60. package/.ai/hr-skills/hr-endorsement-synthesizer.md +82 -0
  61. package/.ai/hr-skills/hr-intake-specifier.md +71 -0
  62. package/.ai/hr-skills/hr-interview-auditor.md +60 -0
  63. package/.ai/hr-skills/hr-jd-drafter.md +61 -0
  64. package/.ai/hr-skills/hr-pipeline-translator.md +58 -0
  65. package/.ai/pm-skills/pm-action-item-mapper.md +61 -0
  66. package/.ai/pm-skills/pm-backlog-auditor.md +57 -0
  67. package/.ai/pm-skills/pm-context-summarizer.md +61 -0
  68. package/.ai/pm-skills/pm-effort-estimator.md +79 -0
  69. package/.ai/pm-skills/pm-newsletter-generator.md +60 -0
  70. package/.ai/pm-skills/pm-progress-translator.md +59 -0
  71. package/.ai/pm-skills/pm-release-note-drafter.md +58 -0
  72. package/.ai/pm-skills/pm-risk-detector.md +58 -0
  73. package/.ai/pm-skills/pm-story-augmenter.md +70 -0
  74. package/.ai/pm-skills/pm-task-specifier.md +70 -0
  75. package/.ai/policies/diagnosis-first.md +26 -0
  76. package/.ai/policies/four-pillars.md +72 -0
  77. package/.ai/policies/user-sovereignty.md +25 -0
  78. package/.ai/skills/accessibility-auditor.md +105 -0
  79. package/.ai/skills/agent-optimizer.md +99 -0
  80. package/.ai/skills/ask.md +200 -0
  81. package/.ai/skills/capacity-planner.md +60 -0
  82. package/.ai/skills/changelog-generator.md +131 -0
  83. package/.ai/skills/clean-code.md +136 -0
  84. package/.ai/skills/code-review-checklist.md +103 -0
  85. package/.ai/skills/codebase-onboarding-intelligence.md +130 -0
  86. package/.ai/skills/competitive-analysis.md +114 -0
  87. package/.ai/skills/daily-standup.md +106 -0
  88. package/.ai/skills/design-system-review.md +308 -0
  89. package/.ai/skills/dev-team-local.md +52 -0
  90. package/.ai/skills/dev-team-orchestrator.md +289 -0
  91. package/.ai/skills/dev-team-sub-max.md +369 -0
  92. package/.ai/skills/dev-team-sub-pro.md +288 -0
  93. package/.ai/skills/dummy-skill.md +28 -0
  94. package/.ai/skills/feature-design-assistant.md +134 -0
  95. package/.ai/skills/feature-orchestrator.md +163 -0
  96. package/.ai/skills/knowledge-manager.md +103 -0
  97. package/.ai/skills/mission-architect.md +86 -0
  98. package/.ai/skills/mission-control.md +102 -0
  99. package/.ai/skills/operational-boundaries.md +94 -0
  100. package/.ai/skills/planning-expert-quick.md +164 -0
  101. package/.ai/skills/planning-expert.md +390 -0
  102. package/.ai/skills/pr-automator.md +431 -0
  103. package/.ai/skills/product-strategist.md +123 -0
  104. package/.ai/skills/qa-handover-generator.md +182 -0
  105. package/.ai/skills/reflexion-loop-local.md +39 -0
  106. package/.ai/skills/reflexion-loop-sub-max.md +214 -0
  107. package/.ai/skills/reflexion-loop-sub-pro.md +164 -0
  108. package/.ai/skills/reflexion-loop.md +119 -0
  109. package/.ai/skills/regression-bug-fix.md +95 -0
  110. package/.ai/skills/security-audit.md +97 -0
  111. package/.ai/skills/solutioning-facilitator.md +338 -0
  112. package/.ai/skills/style-logic-exporter.md +115 -0
  113. package/.ai/skills/technical-debt-auditor.md +119 -0
  114. package/.ai/skills/ui-spec-generator.md +78 -0
  115. package/.ai/skills/verification-auditor.md +101 -0
  116. package/.ai/skills/vertical-slice-decomposer.md +335 -0
  117. package/.ai/skills/visual-verifier.md +134 -0
  118. package/.ai/skills/weekly-leadership-report.md +224 -0
  119. package/.ai/skills.graph.json +1550 -0
  120. package/LICENSE +21 -0
  121. package/README.md +58 -0
  122. package/dist/mcp-server.mjs +5203 -0
  123. package/package.json +48 -0
@@ -0,0 +1,182 @@
1
+ ---
2
+ name: qa-handover-generator
3
+ description: >
4
+ Produces a QA handover + universal smoke-test criteria document for a changed
5
+ feature and delivers it to ClickUp. Splits behaviour by architecture/state
6
+ pattern, states the single source of truth per pattern (from real code), and
7
+ emits smoke-test acceptance criteria that are both agent-ingestible (for
8
+ generating formal acceptance criteria) and directly followable by a human
9
+ tester. All ClickUp output is rendered through the shared clickup-format
10
+ module (single source of truth for ClickUp formatting).
11
+ cost: ~2100 tokens
12
+ modes: [read-only, write, mcp]
13
+ surface: public
14
+ category: Ship & Communicate
15
+ how:
16
+ 'Performs Phase 0 G-Stack discovery of state architecture, maps components to
17
+ server-driven vs client-side patterns, and renders ClickUp markup via the
18
+ clickup-format module.'
19
+ useCase:
20
+ 'Generating high-fidelity QA handovers and smoke test checklists for
21
+ developers and automated testing agents.'
22
+ phase: deploy
23
+ kind: skill
24
+ domain: eng
25
+ ownership:
26
+ drive: human-ai
27
+ approve: human
28
+ targets: [local, api, subscription]
29
+ minModelClass: small
30
+ consumes: [review-report]
31
+ emits: [release]
32
+ policies:
33
+ - user-sovereignty
34
+ - diagnosis-first
35
+ - four-pillars
36
+ ---
37
+
38
+ # QA Handover Generator
39
+
40
+ ## Runtime modes
41
+
42
+ Produces a verifiable QA handover in read-only chat, and in an IDE/MCP agent
43
+ renders it via the shared ClickUp formatter and creates it in ClickUp (with a
44
+ file fallback).
45
+
46
+ **Persistence & Quality Mindset**: There is no reward for completion. The reward
47
+ comes from a handover accurate enough that a QA engineer — or an agent ingesting
48
+ it — can derive correct acceptance criteria without re-reading the source.
49
+ Persist until the architecture split and the single-source-of-truth per pattern
50
+ are stated correctly and render correctly in ClickUp.
51
+
52
+ > [!CAUTION] **ClickUp formatting is NOT hand-rolled.** All ClickUp output MUST
53
+ > be produced via the shared module `scripts/clickup-format.ts` (headings, bold,
54
+ > code, bullets, checklists, tables, document assembly). Never format ClickUp
55
+ > markdown inline in this skill — the shared module is the single source of
56
+ > truth so formatting stays consistent and testable across every ClickUp-
57
+ > producing skill. If ClickUp rendering needs a fix, fix it in that module once.
58
+
59
+ ## 🎯 Handover Gates
60
+
61
+ ### Phase 0: Skill Acquisition & Architecture Discovery (MANDATORY)
62
+
63
+ - **Skill Usage Enforcement (NON-NEGOTIABLE):**
64
+ - **FORBIDDEN:** Direct file access via `view_file` or `run_command` is
65
+ strictly prohibited for skill reading.
66
+ - **IDE / MCP-enabled Agent:** You MUST call the MCP `get_skills` tool (which
67
+ may be prefixed as `mcp_tech-lead-stack_get_skills` or
68
+ `tech-lead-stack_get_skills` depending on client prefixing).
69
+ - **Chat UI (/chat):** You MUST call the internal `get_skill` tool.
70
+
71
+ - **Scope the change:** Identify exactly which modules/screens/components the
72
+ change touches. The handover covers the feature under test, nothing else.
73
+ - **Discover the state architecture (the core of this skill):** From the real
74
+ code, determine how state is managed for the feature under test. Distinguish
75
+ the patterns that actually apply (name only what the code uses):
76
+ - **Server-driven** (e.g. URL query string is the source of truth; a hook
77
+ parses URL state and maps it to server query variables; sort/paginate/filter
78
+ round-trip to the server).
79
+ - **Client-side / in-memory / offline-first** (e.g. a static query fetches all
80
+ records once; sort/search/filter execute in-memory; URL may hold filter
81
+ state for deep-linking but no server round-trip on change).
82
+ - Any other real pattern (cursor-paginated, optimistic-update, event-driven…).
83
+ - **Per pattern, extract the mechanics** a tester needs: single source of truth,
84
+ the owning hook/query/function BY REAL NAME, where filtering/sorting executes,
85
+ reset behaviours (e.g. offset reset on tab change), browser/history/offline
86
+ integration.
87
+ - **Scoped discovery only:** exclude `node_modules`, `.next`, `.nx`, `dist`,
88
+ `build`. No unscoped recursive searches.
89
+
90
+ ### Gate 1: Architecture Overview (verified against code)
91
+
92
+ - **Positive (Pass):** The handover opens with an Architecture Overview split by
93
+ the state patterns actually found. For EACH pattern: target modules, single
94
+ source of truth, mechanism (named hooks/queries/functions), and
95
+ pattern-specific behaviours (resets, history, offline). Every claim traces to
96
+ real code.
97
+ - **Negative (Fail):** Generic overview, a pattern the code does not use, or
98
+ named symbols that do not exist. Rendered with `clickup-format` headings +
99
+ tables/lists.
100
+
101
+ ### Gate 2: Universal Smoke-Test Acceptance Criteria
102
+
103
+ - **Positive (Pass):** For EACH pattern, smoke-test criteria covering general
104
+ usage and core user flows (NOT edge cases): sort, paginate, filter, search,
105
+ tab switch, navigate — each with its expected observable result and any
106
+ pattern-specific gotcha (e.g. "server-side pagination offset is zero-indexed:
107
+ offset=0 is page 1"; "client-side table issues no server request on filter
108
+ change — manipulation is immediate/in-memory").
109
+ - **Dual-audience rule:** Each criterion MUST be (a) concrete enough for an
110
+ agent to convert into formal acceptance criteria, AND (b) followable
111
+ step-by-step by a human doing it manually. Render criteria as ClickUp
112
+ checklist items via `clickup-format.checklist(...)` so QA can tick them off.
113
+
114
+ ### Gate 3: Testability & Environment Notes
115
+
116
+ - **Positive (Pass):** States what the tester needs to run the smoke tests
117
+ locally: which modules/URLs to visit, auth/role requirements, offline/PWA
118
+ considerations, how to observe state (e.g. URL query string for server-driven
119
+ tables), and where behaviour differs by environment (live vs seeded/offline)
120
+ so QA does not report false failures.
121
+
122
+ ### Gate 4: ClickUp Delivery
123
+
124
+ - **Render:** Build the entire document through `scripts/clickup-format.ts`
125
+ (`h2`/`h3`, `bold`, `code`, `bullets`, `checklist`, `renderTable`, `section`,
126
+ `assembleDocument`). Do not concatenate raw markdown by hand.
127
+ - **Tables:** call `renderTable(table, mode)`. Default `mode` is `'list'`
128
+ (guaranteed to render correctly in ClickUp). Only pass `'pipe'` if
129
+ pipe-tables have been confirmed to render in the target ClickUp context. The
130
+ mode is the ONLY table decision — never hand-write table syntax.
131
+ - **Create in ClickUp (primary path):** When the ClickUp MCP is connected AND a
132
+ destination is provided (space/folder/list/doc id + title), create the
133
+ handover via the ClickUp MCP tools — prefer `clickup_create_document` /
134
+ `clickup_create_document_page` for a handover Doc, passing the rendered
135
+ content. Confirm the created doc's headings, table, checkboxes and code render
136
+ correctly.
137
+ - **File fallback (no destination / no MCP):** write the same rendered content
138
+ to `.ai/output/qa-handovers/<feature>-handover.md` for manual paste into
139
+ ClickUp.
140
+ - **Opt-in:** never create in ClickUp without an explicit destination.
141
+
142
+ ## Handover Structure (rendered via clickup-format)
143
+
144
+ ```md
145
+ # QA Handover & Universal Smoke Test Criteria: <Feature>
146
+
147
+ ## 1. <Feature> Architecture Overview
148
+
149
+ <framing paragraph>
150
+ ### A. <Pattern name> (e.g. Server-Side / URL-Driven)
151
+ - **Target Modules:** <real names>
152
+ - **Single Source of Truth:** <what owns state>
153
+ - **Mechanism:** `<hook/query/fn>` — <how controls map to state/server>
154
+ - <pattern-specific behaviours>
155
+ ### B. <Pattern name> (e.g. Client-Side / In-Memory / Offline-First)
156
+ - **Target Module:** <real names>
157
+ - **Mechanism:** `<query/wrapper>` — <fetch/hold data>
158
+ - **Filtering & Search:** <where filtering executes>
159
+ - **Performance / Offline:** <what to expect>
160
+
161
+ ## 2. Universal Smoke Test Acceptance Criteria
162
+
163
+ ### <Pattern A> Smoke Tests (Verify on <modules>. Note: <gotcha>.)
164
+
165
+ - [ ] <interaction> → <expected observable result>
166
+
167
+ ### <Pattern B> Smoke Tests (Verify on <module>. Important: <gotcha>.)
168
+
169
+ - [ ] <interaction> → <expected observable result>
170
+
171
+ ## 3. Testability & Environment Notes
172
+
173
+ - **Local run:** <URLs/modules>
174
+ - **Auth/role:** <requirement>
175
+ - **Data/offline:** <seeded data / PWA / env differences>
176
+ ```
177
+
178
+ ## Telemetry
179
+
180
+ When invoked via MCP skill tools, pass telemetry overrides
181
+ `{ teamRole: "qa", actorType: "AGENT", loopRunId: "<MISSION_ID>" }` so the
182
+ handover generation is attributed on the Agentic Health dashboard.
@@ -0,0 +1,39 @@
1
+ ---
2
+ name: reflexion-loop-local
3
+ description:
4
+ '[LOOP · LOCAL · SAME-MODEL] Fully offline model loop with same-model
5
+ sequential self-critique, governed by a token and wall-clock budget.'
6
+ phase: plan
7
+ kind: skill
8
+ domain: eng
9
+ ownership:
10
+ drive: ai
11
+ approve: human
12
+ targets:
13
+ - local
14
+ minModelClass: small
15
+ cost: ~250 tokens
16
+ modes: [read-only, write, mcp]
17
+ surface: public
18
+ category: Plan & Harden
19
+ policies:
20
+ - four-pillars
21
+ ---
22
+
23
+ # Reflexion Loop (Local Tier)
24
+
25
+ This workflow runs the exact same sequential self-correcting plan loop as the
26
+ standard reflexion-loop, but constrained to a single model instance. Because
27
+ local models are often smaller or slower, it uses the identical model for both
28
+ the generator and the critic steps, governed by an absolute wall-clock limit
29
+ (via `REFLEXION_MAX_WALLCLOCK_MS`) and a token limit.
30
+
31
+ - **Constraints**: No USD budget is applied (since it runs locally for free).
32
+ - **Enforcement**: Model isolation is downgraded to `same-model` via the
33
+ `TIER_POLICY`.
34
+
35
+ Usage:
36
+
37
+ ```bash
38
+ npm run reflexion -- --tier local "Your feature brief"
39
+ ```
@@ -0,0 +1,214 @@
1
+ ---
2
+ name: reflexion-loop-sub-max
3
+ description: >
4
+ [LOOP · SUB-MAX · NO API KEYS · CROSS-MODEL VERIFY] $100/mo tier
5
+ context-isolated plan hardening loop. Manages multi-vendor model isolation
6
+ (L0-L3) and exhaustion limits without losing work, delivering cross-model
7
+ verified plans without requiring API keys. (Note: The stated token cost is per
8
+ loop/run).
9
+ cost: ~2400 tokens
10
+ modes: [read-only, write, mcp]
11
+ surface: public
12
+ category: Plan & Harden
13
+ how:
14
+ 'Multi-vendor model contract, Findings Ledger, and context-firewalled critic
15
+ isolation'
16
+ useCase:
17
+ 'Plan hardening on a $100/mo subscription without requiring external API keys'
18
+ phase: plan
19
+ kind: skill
20
+ domain: eng
21
+ ownership:
22
+ drive: human-ai
23
+ approve: human
24
+ targets: [local, api, subscription]
25
+ minModelClass: small
26
+ consumes: [spec]
27
+ emits: [plan]
28
+ suggests:
29
+ [clean-code, regression-bug-fix, reflexion-loop-sub-pro, reflexion-loop]
30
+ policies:
31
+ - user-sovereignty
32
+ - diagnosis-first
33
+ - four-pillars
34
+ ---
35
+
36
+ # Reflexion Loop ($100/mo Tier - No API Keys)
37
+
38
+ > Tier siblings: reflexion-loop (API keys, dual-model) · reflexion-loop-sub-max
39
+ > ($100 tier) · reflexion-loop-sub-pro ($20 tier). See the tier table in the
40
+ > README.
41
+ >
42
+ > [!NOTE] **Tier profile ($100/mo subscription):** Hardens implementation plans
43
+ > using multi-vendor model isolation (L0–L3) without requiring API keys. Handles
44
+ > model exhaustion limits seamlessly while preserving verification integrity.
45
+
46
+ ## Pre-Flight Model Contract (MANDATORY BEFORE PHASE 0)
47
+
48
+ Read active models from the agent harness at runtime. Formulate and print the
49
+ Pre-Flight Model Contract before Phase 0:
50
+
51
+ | Role | Model class assigned | Isolation vs writer | Continuity Fallback |
52
+ | ------------------ | ----------------------------- | ------------------- | ------------------------ |
53
+ | Generator (Writer) | frontier | N/A (author) | Rung 1 -> Rung 2 |
54
+ | Critic (Auditor) | a different vendor's frontier | L0 (cross-vendor) | Rung 1 -> Rung 2 -> Park |
55
+
56
+ ### Environment Check (`CLAUDE_CODE_SUBAGENT_MODEL`)
57
+
58
+ If `CLAUDE_CODE_SUBAGENT_MODEL` is set to anything other than `inherit`, warn
59
+ plainly in chat and cap claimed isolation level at **L2** (same-model
60
+ sub-agent). Never claim L0 or L1 when overridden by environment variables.
61
+
62
+ ### Four-Level Isolation Ladder
63
+
64
+ - **L0 (Cross-Vendor)**: Generator and critic run on models from different
65
+ vendors. Default target.
66
+ - **L1 (Cross-Family, Same Vendor)**: Generator and critic run on different
67
+ model families from same vendor (shared lineage limitation apply).
68
+ - **L2 (Fresh Sub-Agent)**: Same model, fresh sub-agent context receiving ONLY
69
+ plan text + rubric + diagnosis.
70
+ - **L3 (Fresh Session)**: Same model, plan pasted into a new session cold.
71
+
72
+ **Rule:** Critic is FORBIDDEN from seeing writer's drafting conversation or
73
+ rationale.
74
+
75
+ ### Critic Resolution Ladder (MANDATORY, BEFORE PHASE 1)
76
+
77
+ Resolve the critic with the probe, never by guessing a CLI:
78
+
79
+ ```bash
80
+ ./.ai/rtk-run run resolve-critic --writer <anthropic|google|openai> --writer-model <generator-model>
81
+ # in the stack repo itself: node scripts/resolve-critic.mjs --writer <vendor> --writer-model <id>
82
+ ```
83
+
84
+ `--writer-model` is required for a Claude writer so the Claude rung picks a
85
+ different model. The probe is the single source of truth: it smoke-tests each
86
+ CLI rung in order and prints JSON for the first that answers. Every rung works
87
+ on a subscription login or a pay-as-you-go API key:
88
+
89
+ 1. **Rung G — Gemini** (L0; skipped for a Google writer): `rung: "gemini-cli"`
90
+ (standalone `gemini` with `GEMINI_API_KEY`, Vertex AI or enterprise Code
91
+ Assist), else `rung: "agy"` (Antigravity CLI, consumer Google plans).
92
+ 2. **Rung X — Codex** (`rung: "codex"`, L0; skipped for an OpenAI writer):
93
+ `codex exec` on a ChatGPT plan or an OpenAI API key.
94
+ 3. **Rung C — Claude** (L0, or L1 for a Claude writer):
95
+ `rung: "claude-subagent"` inside Claude Code — spawn a fresh sub-agent on
96
+ exactly the probe's `model`, never the writer's; otherwise
97
+ `rung: "claude-cli"` (`claude -p` on a Claude plan or `ANTHROPIC_API_KEY`).
98
+ 4. **Rung H — Another harness model** (probe returned `rung: "harness"`, or the
99
+ probe could not run): pick a different vendor (L0) or family (L1) in the
100
+ harness / sub-agent model setting, preferring Claude.
101
+ 5. **Rung S — Same model** (nothing else available): fresh sub-agent (L2) or
102
+ fresh session (L3). Verdict is forced to `PROVISIONAL` and `state.json` MUST
103
+ carry `criticAdvisory` (below). Use STATE 2, with its second line naming the
104
+ unavailable rungs instead of a usage limit.
105
+
106
+ For a CLI rung, run the probe's `command` array with the critic prompt (plan +
107
+ rubric + diagnosis only) appended as its **final argument**, e.g.
108
+ `<command...> "$(cat .loop-out/<runId>/critic-prompt-r1.txt)"`. Never pipe the
109
+ prompt on stdin — `agy` ignores stdin.
110
+
111
+ > [!WARNING] The standalone `gemini` CLI stopped serving personal Google logins
112
+ > (AI Pro/Ultra, free Code Assist) on 18 June 2026; a pay-as-you-go
113
+ > `GEMINI_API_KEY` still works. A `GOOGLE_CLOUD_PROJECT` error from it means
114
+ > **unsupported account type**, not a missing setting — the probe moves on to
115
+ > `agy`. Never ask the user to set `GOOGLE_CLOUD_PROJECT` unless they confirm an
116
+ > enterprise Code Assist licence.
117
+
118
+ Record the probe output plus the final rung as `criticResolution` in
119
+ `state.json` (on probe failure, set `probeError` and start at Rung H). On Rung S
120
+ also write:
121
+
122
+ ```json
123
+ "criticAdvisory": {
124
+ "level": "STRONG",
125
+ "code": "SAME_MODEL_CRITIC",
126
+ "message": "The critic ran on the writer's model. This audit is NOT independent; treat the score as self-assessment.",
127
+ "rungsTried": ["<criticResolution.skipped[].rung: reason>", "harness: <why no other model>"],
128
+ "remediation": "Set GEMINI_API_KEY or sign in to agy, sign in to codex, or install the claude CLI, then re-run the critic and deep-review this output before use."
129
+ }
130
+ ```
131
+
132
+ ## Quota Discipline & Findings Ledger
133
+
134
+ - **Turn Budget:** 12 agent turns per run. Print
135
+ `[turn N/12 | model: <active-model>]`.
136
+ - **Findings Ledger (`.loop-out/<runId>/findings.md`):** Write stack facts, file
137
+ paths, domain boundaries, decisions taken, and **OPTIONS REJECTED WITH
138
+ REASONS** (mandatory) to survive model swaps without re-discovery.
139
+ - **State File (`.loop-out/<runId>/state.json`):** Checkpoint after EVERY phase.
140
+ Include `generatorModel`, `criticModel`, `activeIsolationLevel`,
141
+ `criticResolution`, `criticAdvisory` (null unless Rung S), and `status`.
142
+ - **Cold Resume Protocol:** If `.loop-out/<runId>/state.json` exists, read it
143
+ and Findings Ledger, then resume from recorded phase. **Never re-run Phase 0
144
+ discovery on resume.**
145
+
146
+ ## Model Continuity, UNREVIEWED vs PROVISIONAL Verdicts
147
+
148
+ ### Exhaustion Modes A/B/C & Fallback Ladder
149
+
150
+ - **Mode A (MODEL-SCOPED):** Quota spent on one model. Try Rung 1 (cross-vendor
151
+ frontier) -> Rung 2 (same vendor lower class). Recompute isolation level. If
152
+ the critic is exhausted, continue the Critic Resolution Ladder from the NEXT
153
+ rung and update `criticResolution`.
154
+ - **Mode B (ACCOUNT-WIDE):** Consolidate into Findings Ledger, checkpoint state
155
+ as `PARKED`, report reset window. State 3 disclosure applies.
156
+ - **Mode C (SILENT DOWNGRADE):** Poll model identity at every phase boundary.
157
+ Any change is treated as a swap event.
158
+ - **PROVISIONAL Verdict:** If plan passed critique under a degraded critic (same
159
+ model or L2/L3 isolation), mark verdict as `PROVISIONAL`. PROVISIONAL plans
160
+ require explicit Tech-Lead sign-off.
161
+ - **UNREVIEWED Verdict:** If the run stopped before the critic ran or completed
162
+ (Mode B park), mark status as `UNREVIEWED`.
163
+
164
+ ## Three Mandatory End-State Disclosures (VERY FIRST LINE OF VERDICT OUTPUT)
165
+
166
+ The FIRST line of `.loop-out/<runId>/verdict.md` and final chat output MUST emit
167
+ exactly one of these three end states:
168
+
169
+ - **STATE 1 — Separation Held (Auditor finished on a different model):**
170
+
171
+ ```text
172
+ Model separation held: written by <generator-model>, audited by <critic-model>.
173
+ ```
174
+
175
+ - **STATE 2 — Separation Lost (Auditor FINISHED, but on the writer's model):**
176
+
177
+ ```text
178
+ MODEL SEPARATION LOST: <generator-model> wrote this work and also audited it.
179
+ <exhausted-model> hit its usage limit at <phase/step>, so the audit fell back to the same model that produced the work. This audit was not independent.
180
+ ```
181
+
182
+ - **STATE 3 — Audit Incomplete (Run stopped before auditor finished):**
183
+
184
+ ```text
185
+ AUDIT NOT COMPLETED: the run stopped at <phase/step> before the audit finished.
186
+ <exhausted-model> hit an account-wide usage limit, so no model was available to continue. The work below is UNREVIEWED, not approved. Quota resets <window>.
187
+ ```
188
+
189
+ ### Selection Rule
190
+
191
+ State 2 REQUIRES that an audit RAN TO COMPLETION on the writer's model. If the
192
+ audit did not complete, State 3 applies — NEVER State 2. An unfinished audit is
193
+ not a weak audit, it is an absent one.
194
+
195
+ Emit Provenance Table
196
+ (`| Phase | Role | Model | Isolation | Reason for swap | Effect on the claim |`)
197
+ beneath disclosure line.
198
+
199
+ ## Four Pillars Compliance & Anti-Rationalization
200
+
201
+ ### Anti-Rationalization Rebuttals
202
+
203
+ | Rationalization | Rebuttal |
204
+ | ----------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ |
205
+ | "The plan looks fine, let me skip the rubric." | Unscored is unhardened. Produce the rubric. |
206
+ | "I'll harden it after I start coding." | Rework after code is 10x more expensive. |
207
+ | "I have no second model so the loop is pointless." | Context isolation is degraded, not absent — declare level and proceed. |
208
+ | "The audit passed anyway, so the notice would just worry them." | A pass from the author is not a pass; the notice IS the finding. Emit disclosure line as line 1. |
209
+ | "The model swap was handled automatically, so it's an implementation detail." | Handling it seamlessly is why developer cannot see it, which is exactly why it must be stated. |
210
+ | "It is already recorded in the provenance table below." | A table row is not a disclosure; the first line is. |
211
+
212
+ ## Exit States
213
+
214
+ Concludes in one of: `PASSED`, `PROVISIONAL`, `PARKED`, `CAPPED`, `ABORTED`.
@@ -0,0 +1,164 @@
1
+ ---
2
+ name: reflexion-loop-sub-pro
3
+ description: >
4
+ [LOOP · SUB-PRO · NO API KEYS · CROSS-MODEL VERIFY] $20/mo tier
5
+ context-isolated loop. Single-pass cross-model plan check enforcing Mode B
6
+ quota handling and mandatory disclosure without requiring API keys.
7
+ cost: ~1850 tokens
8
+ modes: [read-only, write, mcp]
9
+ surface: public
10
+ category: Plan & Harden
11
+ how:
12
+ 'Single-pass Generator/Critic model contract, Mode B consolidate-and-park, and
13
+ mandatory three-state end-state disclosure'
14
+ useCase:
15
+ 'Frugal single-pass plan verification on a standard ($20/mo) subscription
16
+ without API keys'
17
+ phase: plan
18
+ kind: skill
19
+ domain: eng
20
+ ownership:
21
+ drive: human-ai
22
+ approve: human
23
+ targets: [local, api, subscription]
24
+ minModelClass: small
25
+ consumes: [spec]
26
+ emits: [plan]
27
+ suggests:
28
+ [clean-code, regression-bug-fix, reflexion-loop-sub-max, reflexion-loop]
29
+ policies:
30
+ - user-sovereignty
31
+ - diagnosis-first
32
+ - four-pillars
33
+ ---
34
+
35
+ # Reflexion Loop ($20/mo Tier - No API Keys)
36
+
37
+ > Tier siblings: reflexion-loop (API keys, dual-model) · reflexion-loop-sub-max
38
+ > ($100 tier) · reflexion-loop-sub-pro ($20 tier). See the tier table in the
39
+ > README.
40
+ >
41
+ > [!NOTE] **Tier profile ($20/mo subscription):** Honest promise: **"One good
42
+ > adversarial pass"**, not a hardened plan. Operates as a single-pass
43
+ > cross-model check on a 5-turn budget without API keys.
44
+
45
+ <!-- -->
46
+
47
+ > [!NOTE] **Frugal Compression Notice:** Multi-lane rebalancing is **OMITTED**.
48
+ > The L0–L3 ladder, fallback rungs, and Findings Ledger structure are
49
+ > **COMPRESSED**. Detailed mechanics reference `reflexion-loop-sub-max` by name.
50
+ > The Pre-Flight Model Contract, `CLAUDE_CODE_SUBAGENT_MODEL` check, Mode B
51
+ > consolidate-and-park protocol, and Mandatory Three-State Disclosure rules are
52
+ > included **FULL and uncompressed**.
53
+
54
+ ## Pre-Flight Model Contract (FULL)
55
+
56
+ Read active models from agent harness at runtime. Formulate and print:
57
+
58
+ | Role | Model class assigned | Isolation vs writer | Continuity Fallback |
59
+ | ------------------ | ----------------------------- | ------------------- | ------------------- |
60
+ | Generator (Writer) | frontier / mid | N/A (author) | Consolidate & Park |
61
+ | Critic (Auditor) | a different vendor's frontier | L0 (cross-vendor) | L1 -> L2 -> Park |
62
+
63
+ > **Sub-Pro Note:** Sub-pro is **throughput-limited** (one critique pass,
64
+ > tighter turn budget), not model-limited. Cross-vendor L0 isolation is
65
+ > reachable on the entry tier wherever the platform offers one lineup across
66
+ > paid tiers. Model availability is platform-dependent; read the harness's
67
+ > actual model list at runtime instead of assuming a tier ceiling.
68
+
69
+ ### Environment Check (`CLAUDE_CODE_SUBAGENT_MODEL` — FULL)
70
+
71
+ If `CLAUDE_CODE_SUBAGENT_MODEL` is set to anything other than `inherit`, warn
72
+ plainly in chat and cap claimed isolation level at **L2** (same-model
73
+ sub-agent). Never claim L0 or L1 when overridden by environment variables.
74
+
75
+ ### Isolation Ladder (COMPRESSED)
76
+
77
+ - **L0 (Cross-Vendor)**: Different vendor models. **L1**: Cross-family, same
78
+ vendor. **L2**: Fresh sub-agent. **L3**: Cold paste.
79
+
80
+ ### Critic Resolution Ladder (COMPRESSED — full mechanics in `reflexion-loop-sub-max`)
81
+
82
+ - Resolve the critic with
83
+ `./.ai/rtk-run run resolve-critic --writer <vendor> --writer-model <generator-model>`,
84
+ never by guessing a CLI. Order: Gemini (`gemini-cli`, then `agy`) -> `codex`
85
+ -> Claude (`claude-subagent` inside Claude Code, else `claude-cli`) -> another
86
+ harness model -> same model. Each rung works on a subscription or an API key.
87
+ - Run a CLI rung's `command` with the critic prompt as its final argument; on
88
+ `claude-subagent`, spawn a fresh sub-agent on exactly the probe's `model`.
89
+ - The standalone `gemini` CLI no longer serves personal Google logins (since 18
90
+ June 2026); its `GOOGLE_CLOUD_PROJECT` error means unsupported account type —
91
+ the probe moves on to `agy`.
92
+ - Record `criticResolution` in `.loop-out/<runId>/state.json`. On the same-model
93
+ rung, verdict is `PROVISIONAL` and `state.json` MUST carry the STRONG
94
+ `criticAdvisory` defined in `reflexion-loop-sub-max`.
95
+
96
+ ## Single-Pass Mechanics & Quota Discipline
97
+
98
+ _(Note: The limits below are generated/derived — see `TIER_POLICY['sub-pro']` in
99
+ `src/lib/ai/tier-policy.ts` for the authoritative code policy.)_
100
+
101
+ - **Deltas:** 1 critique pass, max 1 revision, pass threshold 7/10, budget 5
102
+ turns (`[turn N/5 | model: <active-model>]`), plan <= 400 words / <= 8 tasks.
103
+ - **Risk-2 Refusal:** Risk signal = 2 (auth/payments/data/infra) -> MUST refuse
104
+ single-pass hardening and escalate to `reflexion-loop-sub-max` (Enforced by
105
+ `tier-policy.ts`).
106
+
107
+ ## Mode B Consolidate-and-Park Protocol (FULL)
108
+
109
+ On encountering Mode B account-wide limit, or reaching 5 turns:
110
+
111
+ 1. Write brief, diagnosis, score, decisions, and **OPTIONS REJECTED WITH
112
+ REASONS** (mandatory) to `.loop-out/<runId>/loop.md`.
113
+ 2. Mark status as `UNREVIEWED` (or `PROVISIONAL` if degraded critic was used).
114
+ 3. Report progress and reset window in State 3 disclosure.
115
+
116
+ ## Three Mandatory End-State Disclosures (FULL — VERY FIRST LINE OF OUTPUT)
117
+
118
+ The FIRST line of `.loop-out/<runId>/loop.md` and final chat output MUST emit
119
+ exactly one of these three end states:
120
+
121
+ - **STATE 1 — Separation Held (Auditor finished on a different model):**
122
+
123
+ ```text
124
+ Model separation held: written by <generator-model>, audited by <critic-model>.
125
+ ```
126
+
127
+ - **STATE 2 — Separation Lost (Auditor FINISHED, but on the writer's model):**
128
+
129
+ ```text
130
+ MODEL SEPARATION LOST: <generator-model> wrote this work and also audited it.
131
+ <exhausted-model> hit its usage limit at <phase/step>, so the audit fell back to the same model that produced the work. This audit was not independent.
132
+ ```
133
+
134
+ - **STATE 3 — Audit Incomplete (Run stopped before auditor finished):**
135
+
136
+ ```text
137
+ AUDIT NOT COMPLETED: the run stopped at <phase/step> before the audit finished.
138
+ <exhausted-model> hit an account-wide usage limit, so no model was available to continue. The work below is UNREVIEWED, not approved. Quota resets <window>.
139
+ ```
140
+
141
+ ### Selection Rule
142
+
143
+ State 2 REQUIRES that an audit RAN TO COMPLETION on the writer's model. If the
144
+ audit did not complete, State 3 applies — NEVER State 2. An unfinished audit is
145
+ not a weak audit, it is an absent one.
146
+
147
+ Emit Provenance Table
148
+ (`| Phase | Role | Model | Isolation | Reason for swap | Effect on the claim |`)
149
+ beneath disclosure line.
150
+
151
+ ## Four Pillars & Anti-Rationalization
152
+
153
+ | Rationalization | Rebuttal |
154
+ | ----------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ |
155
+ | "The plan looks fine, let me skip the rubric." | Unscored is unhardened. Produce the rubric. |
156
+ | "Risk-2 task can be done in Sub-Pro." | Risk-2 MUST refuse single-pass hardening. Escalate. |
157
+ | "The audit passed anyway, so the notice would just worry them." | A pass from the author is not a pass; the notice IS the finding. Emit disclosure line as line 1. |
158
+ | "The model swap was handled automatically, so it's an implementation detail." | Handling it seamlessly is why developer cannot see it, which is exactly why it must be stated. |
159
+ | "It is already recorded in the provenance table below." | A table row is not a disclosure; the first line is. |
160
+
161
+ ## Exit States
162
+
163
+ Concludes in one of: `PASSED`, `PROVISIONAL`, `UNREVIEWED`, `PARKED`, `CAPPED`,
164
+ `ABORTED`.