tech-lead-stack 1.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/hr-workflows/hr-ad-distributor.md +18 -0
- package/.agents/hr-workflows/hr-candidate-sourcer.md +18 -0
- package/.agents/hr-workflows/hr-endorsement-synthesizer.md +18 -0
- package/.agents/hr-workflows/hr-intake-specifier.md +18 -0
- package/.agents/hr-workflows/hr-interview-auditor.md +18 -0
- package/.agents/hr-workflows/hr-jd-drafter.md +18 -0
- package/.agents/hr-workflows/hr-pipeline-translator.md +18 -0
- package/.agents/pm-workflows/pm-action-item-mapper.md +18 -0
- package/.agents/pm-workflows/pm-backlog-auditor.md +18 -0
- package/.agents/pm-workflows/pm-context-summarizer.md +18 -0
- package/.agents/pm-workflows/pm-design-system-auditor.md +18 -0
- package/.agents/pm-workflows/pm-effort-estimator.md +18 -0
- package/.agents/pm-workflows/pm-newsletter-generator.md +18 -0
- package/.agents/pm-workflows/pm-progress-translator.md +18 -0
- package/.agents/pm-workflows/pm-release-note-drafter.md +18 -0
- package/.agents/pm-workflows/pm-risk-detector.md +18 -0
- package/.agents/pm-workflows/pm-story-augmenter.md +18 -0
- package/.agents/pm-workflows/pm-task-specifier.md +18 -0
- package/.agents/workflows/accessibility-audit.md +30 -0
- package/.agents/workflows/ask.md +44 -0
- package/.agents/workflows/audit-tech-debt.md +31 -0
- package/.agents/workflows/changelog.md +31 -0
- package/.agents/workflows/clean-code-audit.md +31 -0
- package/.agents/workflows/code-review.md +38 -0
- package/.agents/workflows/competitive-analysis.md +46 -0
- package/.agents/workflows/design-requirements-to-architecture.md +31 -0
- package/.agents/workflows/design-system-review.md +113 -0
- package/.agents/workflows/dev-team-sub-max.md +57 -0
- package/.agents/workflows/dev-team-sub-pro.md +57 -0
- package/.agents/workflows/dev-team.md +52 -0
- package/.agents/workflows/feature-orchestrator.md +43 -0
- package/.agents/workflows/init.md +31 -0
- package/.agents/workflows/mission-architect.md +31 -0
- package/.agents/workflows/onboard-dev.md +31 -0
- package/.agents/workflows/plan-quick.md +33 -0
- package/.agents/workflows/plan.md +31 -0
- package/.agents/workflows/pr-automator.md +44 -0
- package/.agents/workflows/pr-design-review-init.md +57 -0
- package/.agents/workflows/qa-handover.md +40 -0
- package/.agents/workflows/reflexion-loop-sub-max.md +46 -0
- package/.agents/workflows/reflexion-loop-sub-pro.md +45 -0
- package/.agents/workflows/reflexion-loop.md +66 -0
- package/.agents/workflows/regression-bug-fix.md +31 -0
- package/.agents/workflows/security-audit.md +31 -0
- package/.agents/workflows/standup-daily-summary.md +31 -0
- package/.agents/workflows/strategy-target-evaluation.md +31 -0
- package/.agents/workflows/style-logic-exporter.md +84 -0
- package/.agents/workflows/ui-spec-generator.md +156 -0
- package/.agents/workflows/verify-changes.md +31 -0
- package/.agents/workflows/vertical-slice.md +52 -0
- package/.agents/workflows/weekly-leadership-report.md +39 -0
- package/.ai/agent-surfaces.json +1235 -0
- package/.ai/hooks/README.md +32 -0
- package/.ai/hooks/build-requires-approved-spec.json +10 -0
- package/.ai/hooks/deploy-requires-review.json +11 -0
- package/.ai/hooks/no-ai-approve-deploy.json +10 -0
- package/.ai/hooks/protected-paths.json +10 -0
- package/.ai/hr-skills/hr-ad-distributor.md +61 -0
- package/.ai/hr-skills/hr-candidate-sourcer.md +69 -0
- package/.ai/hr-skills/hr-endorsement-synthesizer.md +82 -0
- package/.ai/hr-skills/hr-intake-specifier.md +71 -0
- package/.ai/hr-skills/hr-interview-auditor.md +60 -0
- package/.ai/hr-skills/hr-jd-drafter.md +61 -0
- package/.ai/hr-skills/hr-pipeline-translator.md +58 -0
- package/.ai/pm-skills/pm-action-item-mapper.md +61 -0
- package/.ai/pm-skills/pm-backlog-auditor.md +57 -0
- package/.ai/pm-skills/pm-context-summarizer.md +61 -0
- package/.ai/pm-skills/pm-effort-estimator.md +79 -0
- package/.ai/pm-skills/pm-newsletter-generator.md +60 -0
- package/.ai/pm-skills/pm-progress-translator.md +59 -0
- package/.ai/pm-skills/pm-release-note-drafter.md +58 -0
- package/.ai/pm-skills/pm-risk-detector.md +58 -0
- package/.ai/pm-skills/pm-story-augmenter.md +70 -0
- package/.ai/pm-skills/pm-task-specifier.md +70 -0
- package/.ai/policies/diagnosis-first.md +26 -0
- package/.ai/policies/four-pillars.md +72 -0
- package/.ai/policies/user-sovereignty.md +25 -0
- package/.ai/skills/accessibility-auditor.md +105 -0
- package/.ai/skills/agent-optimizer.md +99 -0
- package/.ai/skills/ask.md +200 -0
- package/.ai/skills/capacity-planner.md +60 -0
- package/.ai/skills/changelog-generator.md +131 -0
- package/.ai/skills/clean-code.md +136 -0
- package/.ai/skills/code-review-checklist.md +103 -0
- package/.ai/skills/codebase-onboarding-intelligence.md +130 -0
- package/.ai/skills/competitive-analysis.md +114 -0
- package/.ai/skills/daily-standup.md +106 -0
- package/.ai/skills/design-system-review.md +308 -0
- package/.ai/skills/dev-team-local.md +52 -0
- package/.ai/skills/dev-team-orchestrator.md +289 -0
- package/.ai/skills/dev-team-sub-max.md +369 -0
- package/.ai/skills/dev-team-sub-pro.md +288 -0
- package/.ai/skills/dummy-skill.md +28 -0
- package/.ai/skills/feature-design-assistant.md +134 -0
- package/.ai/skills/feature-orchestrator.md +163 -0
- package/.ai/skills/knowledge-manager.md +103 -0
- package/.ai/skills/mission-architect.md +86 -0
- package/.ai/skills/mission-control.md +102 -0
- package/.ai/skills/operational-boundaries.md +94 -0
- package/.ai/skills/planning-expert-quick.md +164 -0
- package/.ai/skills/planning-expert.md +390 -0
- package/.ai/skills/pr-automator.md +431 -0
- package/.ai/skills/product-strategist.md +123 -0
- package/.ai/skills/qa-handover-generator.md +182 -0
- package/.ai/skills/reflexion-loop-local.md +39 -0
- package/.ai/skills/reflexion-loop-sub-max.md +214 -0
- package/.ai/skills/reflexion-loop-sub-pro.md +164 -0
- package/.ai/skills/reflexion-loop.md +119 -0
- package/.ai/skills/regression-bug-fix.md +95 -0
- package/.ai/skills/security-audit.md +97 -0
- package/.ai/skills/solutioning-facilitator.md +338 -0
- package/.ai/skills/style-logic-exporter.md +115 -0
- package/.ai/skills/technical-debt-auditor.md +119 -0
- package/.ai/skills/ui-spec-generator.md +78 -0
- package/.ai/skills/verification-auditor.md +101 -0
- package/.ai/skills/vertical-slice-decomposer.md +335 -0
- package/.ai/skills/visual-verifier.md +134 -0
- package/.ai/skills/weekly-leadership-report.md +224 -0
- package/.ai/skills.graph.json +1550 -0
- package/LICENSE +21 -0
- package/README.md +58 -0
- package/dist/mcp-server.mjs +5203 -0
- package/package.json +48 -0
|
@@ -0,0 +1,182 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: qa-handover-generator
|
|
3
|
+
description: >
|
|
4
|
+
Produces a QA handover + universal smoke-test criteria document for a changed
|
|
5
|
+
feature and delivers it to ClickUp. Splits behaviour by architecture/state
|
|
6
|
+
pattern, states the single source of truth per pattern (from real code), and
|
|
7
|
+
emits smoke-test acceptance criteria that are both agent-ingestible (for
|
|
8
|
+
generating formal acceptance criteria) and directly followable by a human
|
|
9
|
+
tester. All ClickUp output is rendered through the shared clickup-format
|
|
10
|
+
module (single source of truth for ClickUp formatting).
|
|
11
|
+
cost: ~2100 tokens
|
|
12
|
+
modes: [read-only, write, mcp]
|
|
13
|
+
surface: public
|
|
14
|
+
category: Ship & Communicate
|
|
15
|
+
how:
|
|
16
|
+
'Performs Phase 0 G-Stack discovery of state architecture, maps components to
|
|
17
|
+
server-driven vs client-side patterns, and renders ClickUp markup via the
|
|
18
|
+
clickup-format module.'
|
|
19
|
+
useCase:
|
|
20
|
+
'Generating high-fidelity QA handovers and smoke test checklists for
|
|
21
|
+
developers and automated testing agents.'
|
|
22
|
+
phase: deploy
|
|
23
|
+
kind: skill
|
|
24
|
+
domain: eng
|
|
25
|
+
ownership:
|
|
26
|
+
drive: human-ai
|
|
27
|
+
approve: human
|
|
28
|
+
targets: [local, api, subscription]
|
|
29
|
+
minModelClass: small
|
|
30
|
+
consumes: [review-report]
|
|
31
|
+
emits: [release]
|
|
32
|
+
policies:
|
|
33
|
+
- user-sovereignty
|
|
34
|
+
- diagnosis-first
|
|
35
|
+
- four-pillars
|
|
36
|
+
---
|
|
37
|
+
|
|
38
|
+
# QA Handover Generator
|
|
39
|
+
|
|
40
|
+
## Runtime modes
|
|
41
|
+
|
|
42
|
+
Produces a verifiable QA handover in read-only chat, and in an IDE/MCP agent
|
|
43
|
+
renders it via the shared ClickUp formatter and creates it in ClickUp (with a
|
|
44
|
+
file fallback).
|
|
45
|
+
|
|
46
|
+
**Persistence & Quality Mindset**: There is no reward for completion. The reward
|
|
47
|
+
comes from a handover accurate enough that a QA engineer — or an agent ingesting
|
|
48
|
+
it — can derive correct acceptance criteria without re-reading the source.
|
|
49
|
+
Persist until the architecture split and the single-source-of-truth per pattern
|
|
50
|
+
are stated correctly and render correctly in ClickUp.
|
|
51
|
+
|
|
52
|
+
> [!CAUTION] **ClickUp formatting is NOT hand-rolled.** All ClickUp output MUST
|
|
53
|
+
> be produced via the shared module `scripts/clickup-format.ts` (headings, bold,
|
|
54
|
+
> code, bullets, checklists, tables, document assembly). Never format ClickUp
|
|
55
|
+
> markdown inline in this skill — the shared module is the single source of
|
|
56
|
+
> truth so formatting stays consistent and testable across every ClickUp-
|
|
57
|
+
> producing skill. If ClickUp rendering needs a fix, fix it in that module once.
|
|
58
|
+
|
|
59
|
+
## 🎯 Handover Gates
|
|
60
|
+
|
|
61
|
+
### Phase 0: Skill Acquisition & Architecture Discovery (MANDATORY)
|
|
62
|
+
|
|
63
|
+
- **Skill Usage Enforcement (NON-NEGOTIABLE):**
|
|
64
|
+
- **FORBIDDEN:** Direct file access via `view_file` or `run_command` is
|
|
65
|
+
strictly prohibited for skill reading.
|
|
66
|
+
- **IDE / MCP-enabled Agent:** You MUST call the MCP `get_skills` tool (which
|
|
67
|
+
may be prefixed as `mcp_tech-lead-stack_get_skills` or
|
|
68
|
+
`tech-lead-stack_get_skills` depending on client prefixing).
|
|
69
|
+
- **Chat UI (/chat):** You MUST call the internal `get_skill` tool.
|
|
70
|
+
|
|
71
|
+
- **Scope the change:** Identify exactly which modules/screens/components the
|
|
72
|
+
change touches. The handover covers the feature under test, nothing else.
|
|
73
|
+
- **Discover the state architecture (the core of this skill):** From the real
|
|
74
|
+
code, determine how state is managed for the feature under test. Distinguish
|
|
75
|
+
the patterns that actually apply (name only what the code uses):
|
|
76
|
+
- **Server-driven** (e.g. URL query string is the source of truth; a hook
|
|
77
|
+
parses URL state and maps it to server query variables; sort/paginate/filter
|
|
78
|
+
round-trip to the server).
|
|
79
|
+
- **Client-side / in-memory / offline-first** (e.g. a static query fetches all
|
|
80
|
+
records once; sort/search/filter execute in-memory; URL may hold filter
|
|
81
|
+
state for deep-linking but no server round-trip on change).
|
|
82
|
+
- Any other real pattern (cursor-paginated, optimistic-update, event-driven…).
|
|
83
|
+
- **Per pattern, extract the mechanics** a tester needs: single source of truth,
|
|
84
|
+
the owning hook/query/function BY REAL NAME, where filtering/sorting executes,
|
|
85
|
+
reset behaviours (e.g. offset reset on tab change), browser/history/offline
|
|
86
|
+
integration.
|
|
87
|
+
- **Scoped discovery only:** exclude `node_modules`, `.next`, `.nx`, `dist`,
|
|
88
|
+
`build`. No unscoped recursive searches.
|
|
89
|
+
|
|
90
|
+
### Gate 1: Architecture Overview (verified against code)
|
|
91
|
+
|
|
92
|
+
- **Positive (Pass):** The handover opens with an Architecture Overview split by
|
|
93
|
+
the state patterns actually found. For EACH pattern: target modules, single
|
|
94
|
+
source of truth, mechanism (named hooks/queries/functions), and
|
|
95
|
+
pattern-specific behaviours (resets, history, offline). Every claim traces to
|
|
96
|
+
real code.
|
|
97
|
+
- **Negative (Fail):** Generic overview, a pattern the code does not use, or
|
|
98
|
+
named symbols that do not exist. Rendered with `clickup-format` headings +
|
|
99
|
+
tables/lists.
|
|
100
|
+
|
|
101
|
+
### Gate 2: Universal Smoke-Test Acceptance Criteria
|
|
102
|
+
|
|
103
|
+
- **Positive (Pass):** For EACH pattern, smoke-test criteria covering general
|
|
104
|
+
usage and core user flows (NOT edge cases): sort, paginate, filter, search,
|
|
105
|
+
tab switch, navigate — each with its expected observable result and any
|
|
106
|
+
pattern-specific gotcha (e.g. "server-side pagination offset is zero-indexed:
|
|
107
|
+
offset=0 is page 1"; "client-side table issues no server request on filter
|
|
108
|
+
change — manipulation is immediate/in-memory").
|
|
109
|
+
- **Dual-audience rule:** Each criterion MUST be (a) concrete enough for an
|
|
110
|
+
agent to convert into formal acceptance criteria, AND (b) followable
|
|
111
|
+
step-by-step by a human doing it manually. Render criteria as ClickUp
|
|
112
|
+
checklist items via `clickup-format.checklist(...)` so QA can tick them off.
|
|
113
|
+
|
|
114
|
+
### Gate 3: Testability & Environment Notes
|
|
115
|
+
|
|
116
|
+
- **Positive (Pass):** States what the tester needs to run the smoke tests
|
|
117
|
+
locally: which modules/URLs to visit, auth/role requirements, offline/PWA
|
|
118
|
+
considerations, how to observe state (e.g. URL query string for server-driven
|
|
119
|
+
tables), and where behaviour differs by environment (live vs seeded/offline)
|
|
120
|
+
so QA does not report false failures.
|
|
121
|
+
|
|
122
|
+
### Gate 4: ClickUp Delivery
|
|
123
|
+
|
|
124
|
+
- **Render:** Build the entire document through `scripts/clickup-format.ts`
|
|
125
|
+
(`h2`/`h3`, `bold`, `code`, `bullets`, `checklist`, `renderTable`, `section`,
|
|
126
|
+
`assembleDocument`). Do not concatenate raw markdown by hand.
|
|
127
|
+
- **Tables:** call `renderTable(table, mode)`. Default `mode` is `'list'`
|
|
128
|
+
(guaranteed to render correctly in ClickUp). Only pass `'pipe'` if
|
|
129
|
+
pipe-tables have been confirmed to render in the target ClickUp context. The
|
|
130
|
+
mode is the ONLY table decision — never hand-write table syntax.
|
|
131
|
+
- **Create in ClickUp (primary path):** When the ClickUp MCP is connected AND a
|
|
132
|
+
destination is provided (space/folder/list/doc id + title), create the
|
|
133
|
+
handover via the ClickUp MCP tools — prefer `clickup_create_document` /
|
|
134
|
+
`clickup_create_document_page` for a handover Doc, passing the rendered
|
|
135
|
+
content. Confirm the created doc's headings, table, checkboxes and code render
|
|
136
|
+
correctly.
|
|
137
|
+
- **File fallback (no destination / no MCP):** write the same rendered content
|
|
138
|
+
to `.ai/output/qa-handovers/<feature>-handover.md` for manual paste into
|
|
139
|
+
ClickUp.
|
|
140
|
+
- **Opt-in:** never create in ClickUp without an explicit destination.
|
|
141
|
+
|
|
142
|
+
## Handover Structure (rendered via clickup-format)
|
|
143
|
+
|
|
144
|
+
```md
|
|
145
|
+
# QA Handover & Universal Smoke Test Criteria: <Feature>
|
|
146
|
+
|
|
147
|
+
## 1. <Feature> Architecture Overview
|
|
148
|
+
|
|
149
|
+
<framing paragraph>
|
|
150
|
+
### A. <Pattern name> (e.g. Server-Side / URL-Driven)
|
|
151
|
+
- **Target Modules:** <real names>
|
|
152
|
+
- **Single Source of Truth:** <what owns state>
|
|
153
|
+
- **Mechanism:** `<hook/query/fn>` — <how controls map to state/server>
|
|
154
|
+
- <pattern-specific behaviours>
|
|
155
|
+
### B. <Pattern name> (e.g. Client-Side / In-Memory / Offline-First)
|
|
156
|
+
- **Target Module:** <real names>
|
|
157
|
+
- **Mechanism:** `<query/wrapper>` — <fetch/hold data>
|
|
158
|
+
- **Filtering & Search:** <where filtering executes>
|
|
159
|
+
- **Performance / Offline:** <what to expect>
|
|
160
|
+
|
|
161
|
+
## 2. Universal Smoke Test Acceptance Criteria
|
|
162
|
+
|
|
163
|
+
### <Pattern A> Smoke Tests (Verify on <modules>. Note: <gotcha>.)
|
|
164
|
+
|
|
165
|
+
- [ ] <interaction> → <expected observable result>
|
|
166
|
+
|
|
167
|
+
### <Pattern B> Smoke Tests (Verify on <module>. Important: <gotcha>.)
|
|
168
|
+
|
|
169
|
+
- [ ] <interaction> → <expected observable result>
|
|
170
|
+
|
|
171
|
+
## 3. Testability & Environment Notes
|
|
172
|
+
|
|
173
|
+
- **Local run:** <URLs/modules>
|
|
174
|
+
- **Auth/role:** <requirement>
|
|
175
|
+
- **Data/offline:** <seeded data / PWA / env differences>
|
|
176
|
+
```
|
|
177
|
+
|
|
178
|
+
## Telemetry
|
|
179
|
+
|
|
180
|
+
When invoked via MCP skill tools, pass telemetry overrides
|
|
181
|
+
`{ teamRole: "qa", actorType: "AGENT", loopRunId: "<MISSION_ID>" }` so the
|
|
182
|
+
handover generation is attributed on the Agentic Health dashboard.
|
|
@@ -0,0 +1,39 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: reflexion-loop-local
|
|
3
|
+
description:
|
|
4
|
+
'[LOOP · LOCAL · SAME-MODEL] Fully offline model loop with same-model
|
|
5
|
+
sequential self-critique, governed by a token and wall-clock budget.'
|
|
6
|
+
phase: plan
|
|
7
|
+
kind: skill
|
|
8
|
+
domain: eng
|
|
9
|
+
ownership:
|
|
10
|
+
drive: ai
|
|
11
|
+
approve: human
|
|
12
|
+
targets:
|
|
13
|
+
- local
|
|
14
|
+
minModelClass: small
|
|
15
|
+
cost: ~250 tokens
|
|
16
|
+
modes: [read-only, write, mcp]
|
|
17
|
+
surface: public
|
|
18
|
+
category: Plan & Harden
|
|
19
|
+
policies:
|
|
20
|
+
- four-pillars
|
|
21
|
+
---
|
|
22
|
+
|
|
23
|
+
# Reflexion Loop (Local Tier)
|
|
24
|
+
|
|
25
|
+
This workflow runs the exact same sequential self-correcting plan loop as the
|
|
26
|
+
standard reflexion-loop, but constrained to a single model instance. Because
|
|
27
|
+
local models are often smaller or slower, it uses the identical model for both
|
|
28
|
+
the generator and the critic steps, governed by an absolute wall-clock limit
|
|
29
|
+
(via `REFLEXION_MAX_WALLCLOCK_MS`) and a token limit.
|
|
30
|
+
|
|
31
|
+
- **Constraints**: No USD budget is applied (since it runs locally for free).
|
|
32
|
+
- **Enforcement**: Model isolation is downgraded to `same-model` via the
|
|
33
|
+
`TIER_POLICY`.
|
|
34
|
+
|
|
35
|
+
Usage:
|
|
36
|
+
|
|
37
|
+
```bash
|
|
38
|
+
npm run reflexion -- --tier local "Your feature brief"
|
|
39
|
+
```
|
|
@@ -0,0 +1,214 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: reflexion-loop-sub-max
|
|
3
|
+
description: >
|
|
4
|
+
[LOOP · SUB-MAX · NO API KEYS · CROSS-MODEL VERIFY] $100/mo tier
|
|
5
|
+
context-isolated plan hardening loop. Manages multi-vendor model isolation
|
|
6
|
+
(L0-L3) and exhaustion limits without losing work, delivering cross-model
|
|
7
|
+
verified plans without requiring API keys. (Note: The stated token cost is per
|
|
8
|
+
loop/run).
|
|
9
|
+
cost: ~2400 tokens
|
|
10
|
+
modes: [read-only, write, mcp]
|
|
11
|
+
surface: public
|
|
12
|
+
category: Plan & Harden
|
|
13
|
+
how:
|
|
14
|
+
'Multi-vendor model contract, Findings Ledger, and context-firewalled critic
|
|
15
|
+
isolation'
|
|
16
|
+
useCase:
|
|
17
|
+
'Plan hardening on a $100/mo subscription without requiring external API keys'
|
|
18
|
+
phase: plan
|
|
19
|
+
kind: skill
|
|
20
|
+
domain: eng
|
|
21
|
+
ownership:
|
|
22
|
+
drive: human-ai
|
|
23
|
+
approve: human
|
|
24
|
+
targets: [local, api, subscription]
|
|
25
|
+
minModelClass: small
|
|
26
|
+
consumes: [spec]
|
|
27
|
+
emits: [plan]
|
|
28
|
+
suggests:
|
|
29
|
+
[clean-code, regression-bug-fix, reflexion-loop-sub-pro, reflexion-loop]
|
|
30
|
+
policies:
|
|
31
|
+
- user-sovereignty
|
|
32
|
+
- diagnosis-first
|
|
33
|
+
- four-pillars
|
|
34
|
+
---
|
|
35
|
+
|
|
36
|
+
# Reflexion Loop ($100/mo Tier - No API Keys)
|
|
37
|
+
|
|
38
|
+
> Tier siblings: reflexion-loop (API keys, dual-model) · reflexion-loop-sub-max
|
|
39
|
+
> ($100 tier) · reflexion-loop-sub-pro ($20 tier). See the tier table in the
|
|
40
|
+
> README.
|
|
41
|
+
>
|
|
42
|
+
> [!NOTE] **Tier profile ($100/mo subscription):** Hardens implementation plans
|
|
43
|
+
> using multi-vendor model isolation (L0–L3) without requiring API keys. Handles
|
|
44
|
+
> model exhaustion limits seamlessly while preserving verification integrity.
|
|
45
|
+
|
|
46
|
+
## Pre-Flight Model Contract (MANDATORY BEFORE PHASE 0)
|
|
47
|
+
|
|
48
|
+
Read active models from the agent harness at runtime. Formulate and print the
|
|
49
|
+
Pre-Flight Model Contract before Phase 0:
|
|
50
|
+
|
|
51
|
+
| Role | Model class assigned | Isolation vs writer | Continuity Fallback |
|
|
52
|
+
| ------------------ | ----------------------------- | ------------------- | ------------------------ |
|
|
53
|
+
| Generator (Writer) | frontier | N/A (author) | Rung 1 -> Rung 2 |
|
|
54
|
+
| Critic (Auditor) | a different vendor's frontier | L0 (cross-vendor) | Rung 1 -> Rung 2 -> Park |
|
|
55
|
+
|
|
56
|
+
### Environment Check (`CLAUDE_CODE_SUBAGENT_MODEL`)
|
|
57
|
+
|
|
58
|
+
If `CLAUDE_CODE_SUBAGENT_MODEL` is set to anything other than `inherit`, warn
|
|
59
|
+
plainly in chat and cap claimed isolation level at **L2** (same-model
|
|
60
|
+
sub-agent). Never claim L0 or L1 when overridden by environment variables.
|
|
61
|
+
|
|
62
|
+
### Four-Level Isolation Ladder
|
|
63
|
+
|
|
64
|
+
- **L0 (Cross-Vendor)**: Generator and critic run on models from different
|
|
65
|
+
vendors. Default target.
|
|
66
|
+
- **L1 (Cross-Family, Same Vendor)**: Generator and critic run on different
|
|
67
|
+
model families from same vendor (shared lineage limitation apply).
|
|
68
|
+
- **L2 (Fresh Sub-Agent)**: Same model, fresh sub-agent context receiving ONLY
|
|
69
|
+
plan text + rubric + diagnosis.
|
|
70
|
+
- **L3 (Fresh Session)**: Same model, plan pasted into a new session cold.
|
|
71
|
+
|
|
72
|
+
**Rule:** Critic is FORBIDDEN from seeing writer's drafting conversation or
|
|
73
|
+
rationale.
|
|
74
|
+
|
|
75
|
+
### Critic Resolution Ladder (MANDATORY, BEFORE PHASE 1)
|
|
76
|
+
|
|
77
|
+
Resolve the critic with the probe, never by guessing a CLI:
|
|
78
|
+
|
|
79
|
+
```bash
|
|
80
|
+
./.ai/rtk-run run resolve-critic --writer <anthropic|google|openai> --writer-model <generator-model>
|
|
81
|
+
# in the stack repo itself: node scripts/resolve-critic.mjs --writer <vendor> --writer-model <id>
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
`--writer-model` is required for a Claude writer so the Claude rung picks a
|
|
85
|
+
different model. The probe is the single source of truth: it smoke-tests each
|
|
86
|
+
CLI rung in order and prints JSON for the first that answers. Every rung works
|
|
87
|
+
on a subscription login or a pay-as-you-go API key:
|
|
88
|
+
|
|
89
|
+
1. **Rung G — Gemini** (L0; skipped for a Google writer): `rung: "gemini-cli"`
|
|
90
|
+
(standalone `gemini` with `GEMINI_API_KEY`, Vertex AI or enterprise Code
|
|
91
|
+
Assist), else `rung: "agy"` (Antigravity CLI, consumer Google plans).
|
|
92
|
+
2. **Rung X — Codex** (`rung: "codex"`, L0; skipped for an OpenAI writer):
|
|
93
|
+
`codex exec` on a ChatGPT plan or an OpenAI API key.
|
|
94
|
+
3. **Rung C — Claude** (L0, or L1 for a Claude writer):
|
|
95
|
+
`rung: "claude-subagent"` inside Claude Code — spawn a fresh sub-agent on
|
|
96
|
+
exactly the probe's `model`, never the writer's; otherwise
|
|
97
|
+
`rung: "claude-cli"` (`claude -p` on a Claude plan or `ANTHROPIC_API_KEY`).
|
|
98
|
+
4. **Rung H — Another harness model** (probe returned `rung: "harness"`, or the
|
|
99
|
+
probe could not run): pick a different vendor (L0) or family (L1) in the
|
|
100
|
+
harness / sub-agent model setting, preferring Claude.
|
|
101
|
+
5. **Rung S — Same model** (nothing else available): fresh sub-agent (L2) or
|
|
102
|
+
fresh session (L3). Verdict is forced to `PROVISIONAL` and `state.json` MUST
|
|
103
|
+
carry `criticAdvisory` (below). Use STATE 2, with its second line naming the
|
|
104
|
+
unavailable rungs instead of a usage limit.
|
|
105
|
+
|
|
106
|
+
For a CLI rung, run the probe's `command` array with the critic prompt (plan +
|
|
107
|
+
rubric + diagnosis only) appended as its **final argument**, e.g.
|
|
108
|
+
`<command...> "$(cat .loop-out/<runId>/critic-prompt-r1.txt)"`. Never pipe the
|
|
109
|
+
prompt on stdin — `agy` ignores stdin.
|
|
110
|
+
|
|
111
|
+
> [!WARNING] The standalone `gemini` CLI stopped serving personal Google logins
|
|
112
|
+
> (AI Pro/Ultra, free Code Assist) on 18 June 2026; a pay-as-you-go
|
|
113
|
+
> `GEMINI_API_KEY` still works. A `GOOGLE_CLOUD_PROJECT` error from it means
|
|
114
|
+
> **unsupported account type**, not a missing setting — the probe moves on to
|
|
115
|
+
> `agy`. Never ask the user to set `GOOGLE_CLOUD_PROJECT` unless they confirm an
|
|
116
|
+
> enterprise Code Assist licence.
|
|
117
|
+
|
|
118
|
+
Record the probe output plus the final rung as `criticResolution` in
|
|
119
|
+
`state.json` (on probe failure, set `probeError` and start at Rung H). On Rung S
|
|
120
|
+
also write:
|
|
121
|
+
|
|
122
|
+
```json
|
|
123
|
+
"criticAdvisory": {
|
|
124
|
+
"level": "STRONG",
|
|
125
|
+
"code": "SAME_MODEL_CRITIC",
|
|
126
|
+
"message": "The critic ran on the writer's model. This audit is NOT independent; treat the score as self-assessment.",
|
|
127
|
+
"rungsTried": ["<criticResolution.skipped[].rung: reason>", "harness: <why no other model>"],
|
|
128
|
+
"remediation": "Set GEMINI_API_KEY or sign in to agy, sign in to codex, or install the claude CLI, then re-run the critic and deep-review this output before use."
|
|
129
|
+
}
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
## Quota Discipline & Findings Ledger
|
|
133
|
+
|
|
134
|
+
- **Turn Budget:** 12 agent turns per run. Print
|
|
135
|
+
`[turn N/12 | model: <active-model>]`.
|
|
136
|
+
- **Findings Ledger (`.loop-out/<runId>/findings.md`):** Write stack facts, file
|
|
137
|
+
paths, domain boundaries, decisions taken, and **OPTIONS REJECTED WITH
|
|
138
|
+
REASONS** (mandatory) to survive model swaps without re-discovery.
|
|
139
|
+
- **State File (`.loop-out/<runId>/state.json`):** Checkpoint after EVERY phase.
|
|
140
|
+
Include `generatorModel`, `criticModel`, `activeIsolationLevel`,
|
|
141
|
+
`criticResolution`, `criticAdvisory` (null unless Rung S), and `status`.
|
|
142
|
+
- **Cold Resume Protocol:** If `.loop-out/<runId>/state.json` exists, read it
|
|
143
|
+
and Findings Ledger, then resume from recorded phase. **Never re-run Phase 0
|
|
144
|
+
discovery on resume.**
|
|
145
|
+
|
|
146
|
+
## Model Continuity, UNREVIEWED vs PROVISIONAL Verdicts
|
|
147
|
+
|
|
148
|
+
### Exhaustion Modes A/B/C & Fallback Ladder
|
|
149
|
+
|
|
150
|
+
- **Mode A (MODEL-SCOPED):** Quota spent on one model. Try Rung 1 (cross-vendor
|
|
151
|
+
frontier) -> Rung 2 (same vendor lower class). Recompute isolation level. If
|
|
152
|
+
the critic is exhausted, continue the Critic Resolution Ladder from the NEXT
|
|
153
|
+
rung and update `criticResolution`.
|
|
154
|
+
- **Mode B (ACCOUNT-WIDE):** Consolidate into Findings Ledger, checkpoint state
|
|
155
|
+
as `PARKED`, report reset window. State 3 disclosure applies.
|
|
156
|
+
- **Mode C (SILENT DOWNGRADE):** Poll model identity at every phase boundary.
|
|
157
|
+
Any change is treated as a swap event.
|
|
158
|
+
- **PROVISIONAL Verdict:** If plan passed critique under a degraded critic (same
|
|
159
|
+
model or L2/L3 isolation), mark verdict as `PROVISIONAL`. PROVISIONAL plans
|
|
160
|
+
require explicit Tech-Lead sign-off.
|
|
161
|
+
- **UNREVIEWED Verdict:** If the run stopped before the critic ran or completed
|
|
162
|
+
(Mode B park), mark status as `UNREVIEWED`.
|
|
163
|
+
|
|
164
|
+
## Three Mandatory End-State Disclosures (VERY FIRST LINE OF VERDICT OUTPUT)
|
|
165
|
+
|
|
166
|
+
The FIRST line of `.loop-out/<runId>/verdict.md` and final chat output MUST emit
|
|
167
|
+
exactly one of these three end states:
|
|
168
|
+
|
|
169
|
+
- **STATE 1 — Separation Held (Auditor finished on a different model):**
|
|
170
|
+
|
|
171
|
+
```text
|
|
172
|
+
Model separation held: written by <generator-model>, audited by <critic-model>.
|
|
173
|
+
```
|
|
174
|
+
|
|
175
|
+
- **STATE 2 — Separation Lost (Auditor FINISHED, but on the writer's model):**
|
|
176
|
+
|
|
177
|
+
```text
|
|
178
|
+
MODEL SEPARATION LOST: <generator-model> wrote this work and also audited it.
|
|
179
|
+
<exhausted-model> hit its usage limit at <phase/step>, so the audit fell back to the same model that produced the work. This audit was not independent.
|
|
180
|
+
```
|
|
181
|
+
|
|
182
|
+
- **STATE 3 — Audit Incomplete (Run stopped before auditor finished):**
|
|
183
|
+
|
|
184
|
+
```text
|
|
185
|
+
AUDIT NOT COMPLETED: the run stopped at <phase/step> before the audit finished.
|
|
186
|
+
<exhausted-model> hit an account-wide usage limit, so no model was available to continue. The work below is UNREVIEWED, not approved. Quota resets <window>.
|
|
187
|
+
```
|
|
188
|
+
|
|
189
|
+
### Selection Rule
|
|
190
|
+
|
|
191
|
+
State 2 REQUIRES that an audit RAN TO COMPLETION on the writer's model. If the
|
|
192
|
+
audit did not complete, State 3 applies — NEVER State 2. An unfinished audit is
|
|
193
|
+
not a weak audit, it is an absent one.
|
|
194
|
+
|
|
195
|
+
Emit Provenance Table
|
|
196
|
+
(`| Phase | Role | Model | Isolation | Reason for swap | Effect on the claim |`)
|
|
197
|
+
beneath disclosure line.
|
|
198
|
+
|
|
199
|
+
## Four Pillars Compliance & Anti-Rationalization
|
|
200
|
+
|
|
201
|
+
### Anti-Rationalization Rebuttals
|
|
202
|
+
|
|
203
|
+
| Rationalization | Rebuttal |
|
|
204
|
+
| ----------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ |
|
|
205
|
+
| "The plan looks fine, let me skip the rubric." | Unscored is unhardened. Produce the rubric. |
|
|
206
|
+
| "I'll harden it after I start coding." | Rework after code is 10x more expensive. |
|
|
207
|
+
| "I have no second model so the loop is pointless." | Context isolation is degraded, not absent — declare level and proceed. |
|
|
208
|
+
| "The audit passed anyway, so the notice would just worry them." | A pass from the author is not a pass; the notice IS the finding. Emit disclosure line as line 1. |
|
|
209
|
+
| "The model swap was handled automatically, so it's an implementation detail." | Handling it seamlessly is why developer cannot see it, which is exactly why it must be stated. |
|
|
210
|
+
| "It is already recorded in the provenance table below." | A table row is not a disclosure; the first line is. |
|
|
211
|
+
|
|
212
|
+
## Exit States
|
|
213
|
+
|
|
214
|
+
Concludes in one of: `PASSED`, `PROVISIONAL`, `PARKED`, `CAPPED`, `ABORTED`.
|
|
@@ -0,0 +1,164 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: reflexion-loop-sub-pro
|
|
3
|
+
description: >
|
|
4
|
+
[LOOP · SUB-PRO · NO API KEYS · CROSS-MODEL VERIFY] $20/mo tier
|
|
5
|
+
context-isolated loop. Single-pass cross-model plan check enforcing Mode B
|
|
6
|
+
quota handling and mandatory disclosure without requiring API keys.
|
|
7
|
+
cost: ~1850 tokens
|
|
8
|
+
modes: [read-only, write, mcp]
|
|
9
|
+
surface: public
|
|
10
|
+
category: Plan & Harden
|
|
11
|
+
how:
|
|
12
|
+
'Single-pass Generator/Critic model contract, Mode B consolidate-and-park, and
|
|
13
|
+
mandatory three-state end-state disclosure'
|
|
14
|
+
useCase:
|
|
15
|
+
'Frugal single-pass plan verification on a standard ($20/mo) subscription
|
|
16
|
+
without API keys'
|
|
17
|
+
phase: plan
|
|
18
|
+
kind: skill
|
|
19
|
+
domain: eng
|
|
20
|
+
ownership:
|
|
21
|
+
drive: human-ai
|
|
22
|
+
approve: human
|
|
23
|
+
targets: [local, api, subscription]
|
|
24
|
+
minModelClass: small
|
|
25
|
+
consumes: [spec]
|
|
26
|
+
emits: [plan]
|
|
27
|
+
suggests:
|
|
28
|
+
[clean-code, regression-bug-fix, reflexion-loop-sub-max, reflexion-loop]
|
|
29
|
+
policies:
|
|
30
|
+
- user-sovereignty
|
|
31
|
+
- diagnosis-first
|
|
32
|
+
- four-pillars
|
|
33
|
+
---
|
|
34
|
+
|
|
35
|
+
# Reflexion Loop ($20/mo Tier - No API Keys)
|
|
36
|
+
|
|
37
|
+
> Tier siblings: reflexion-loop (API keys, dual-model) · reflexion-loop-sub-max
|
|
38
|
+
> ($100 tier) · reflexion-loop-sub-pro ($20 tier). See the tier table in the
|
|
39
|
+
> README.
|
|
40
|
+
>
|
|
41
|
+
> [!NOTE] **Tier profile ($20/mo subscription):** Honest promise: **"One good
|
|
42
|
+
> adversarial pass"**, not a hardened plan. Operates as a single-pass
|
|
43
|
+
> cross-model check on a 5-turn budget without API keys.
|
|
44
|
+
|
|
45
|
+
<!-- -->
|
|
46
|
+
|
|
47
|
+
> [!NOTE] **Frugal Compression Notice:** Multi-lane rebalancing is **OMITTED**.
|
|
48
|
+
> The L0–L3 ladder, fallback rungs, and Findings Ledger structure are
|
|
49
|
+
> **COMPRESSED**. Detailed mechanics reference `reflexion-loop-sub-max` by name.
|
|
50
|
+
> The Pre-Flight Model Contract, `CLAUDE_CODE_SUBAGENT_MODEL` check, Mode B
|
|
51
|
+
> consolidate-and-park protocol, and Mandatory Three-State Disclosure rules are
|
|
52
|
+
> included **FULL and uncompressed**.
|
|
53
|
+
|
|
54
|
+
## Pre-Flight Model Contract (FULL)
|
|
55
|
+
|
|
56
|
+
Read active models from agent harness at runtime. Formulate and print:
|
|
57
|
+
|
|
58
|
+
| Role | Model class assigned | Isolation vs writer | Continuity Fallback |
|
|
59
|
+
| ------------------ | ----------------------------- | ------------------- | ------------------- |
|
|
60
|
+
| Generator (Writer) | frontier / mid | N/A (author) | Consolidate & Park |
|
|
61
|
+
| Critic (Auditor) | a different vendor's frontier | L0 (cross-vendor) | L1 -> L2 -> Park |
|
|
62
|
+
|
|
63
|
+
> **Sub-Pro Note:** Sub-pro is **throughput-limited** (one critique pass,
|
|
64
|
+
> tighter turn budget), not model-limited. Cross-vendor L0 isolation is
|
|
65
|
+
> reachable on the entry tier wherever the platform offers one lineup across
|
|
66
|
+
> paid tiers. Model availability is platform-dependent; read the harness's
|
|
67
|
+
> actual model list at runtime instead of assuming a tier ceiling.
|
|
68
|
+
|
|
69
|
+
### Environment Check (`CLAUDE_CODE_SUBAGENT_MODEL` — FULL)
|
|
70
|
+
|
|
71
|
+
If `CLAUDE_CODE_SUBAGENT_MODEL` is set to anything other than `inherit`, warn
|
|
72
|
+
plainly in chat and cap claimed isolation level at **L2** (same-model
|
|
73
|
+
sub-agent). Never claim L0 or L1 when overridden by environment variables.
|
|
74
|
+
|
|
75
|
+
### Isolation Ladder (COMPRESSED)
|
|
76
|
+
|
|
77
|
+
- **L0 (Cross-Vendor)**: Different vendor models. **L1**: Cross-family, same
|
|
78
|
+
vendor. **L2**: Fresh sub-agent. **L3**: Cold paste.
|
|
79
|
+
|
|
80
|
+
### Critic Resolution Ladder (COMPRESSED — full mechanics in `reflexion-loop-sub-max`)
|
|
81
|
+
|
|
82
|
+
- Resolve the critic with
|
|
83
|
+
`./.ai/rtk-run run resolve-critic --writer <vendor> --writer-model <generator-model>`,
|
|
84
|
+
never by guessing a CLI. Order: Gemini (`gemini-cli`, then `agy`) -> `codex`
|
|
85
|
+
-> Claude (`claude-subagent` inside Claude Code, else `claude-cli`) -> another
|
|
86
|
+
harness model -> same model. Each rung works on a subscription or an API key.
|
|
87
|
+
- Run a CLI rung's `command` with the critic prompt as its final argument; on
|
|
88
|
+
`claude-subagent`, spawn a fresh sub-agent on exactly the probe's `model`.
|
|
89
|
+
- The standalone `gemini` CLI no longer serves personal Google logins (since 18
|
|
90
|
+
June 2026); its `GOOGLE_CLOUD_PROJECT` error means unsupported account type —
|
|
91
|
+
the probe moves on to `agy`.
|
|
92
|
+
- Record `criticResolution` in `.loop-out/<runId>/state.json`. On the same-model
|
|
93
|
+
rung, verdict is `PROVISIONAL` and `state.json` MUST carry the STRONG
|
|
94
|
+
`criticAdvisory` defined in `reflexion-loop-sub-max`.
|
|
95
|
+
|
|
96
|
+
## Single-Pass Mechanics & Quota Discipline
|
|
97
|
+
|
|
98
|
+
_(Note: The limits below are generated/derived — see `TIER_POLICY['sub-pro']` in
|
|
99
|
+
`src/lib/ai/tier-policy.ts` for the authoritative code policy.)_
|
|
100
|
+
|
|
101
|
+
- **Deltas:** 1 critique pass, max 1 revision, pass threshold 7/10, budget 5
|
|
102
|
+
turns (`[turn N/5 | model: <active-model>]`), plan <= 400 words / <= 8 tasks.
|
|
103
|
+
- **Risk-2 Refusal:** Risk signal = 2 (auth/payments/data/infra) -> MUST refuse
|
|
104
|
+
single-pass hardening and escalate to `reflexion-loop-sub-max` (Enforced by
|
|
105
|
+
`tier-policy.ts`).
|
|
106
|
+
|
|
107
|
+
## Mode B Consolidate-and-Park Protocol (FULL)
|
|
108
|
+
|
|
109
|
+
On encountering Mode B account-wide limit, or reaching 5 turns:
|
|
110
|
+
|
|
111
|
+
1. Write brief, diagnosis, score, decisions, and **OPTIONS REJECTED WITH
|
|
112
|
+
REASONS** (mandatory) to `.loop-out/<runId>/loop.md`.
|
|
113
|
+
2. Mark status as `UNREVIEWED` (or `PROVISIONAL` if degraded critic was used).
|
|
114
|
+
3. Report progress and reset window in State 3 disclosure.
|
|
115
|
+
|
|
116
|
+
## Three Mandatory End-State Disclosures (FULL — VERY FIRST LINE OF OUTPUT)
|
|
117
|
+
|
|
118
|
+
The FIRST line of `.loop-out/<runId>/loop.md` and final chat output MUST emit
|
|
119
|
+
exactly one of these three end states:
|
|
120
|
+
|
|
121
|
+
- **STATE 1 — Separation Held (Auditor finished on a different model):**
|
|
122
|
+
|
|
123
|
+
```text
|
|
124
|
+
Model separation held: written by <generator-model>, audited by <critic-model>.
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
- **STATE 2 — Separation Lost (Auditor FINISHED, but on the writer's model):**
|
|
128
|
+
|
|
129
|
+
```text
|
|
130
|
+
MODEL SEPARATION LOST: <generator-model> wrote this work and also audited it.
|
|
131
|
+
<exhausted-model> hit its usage limit at <phase/step>, so the audit fell back to the same model that produced the work. This audit was not independent.
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
- **STATE 3 — Audit Incomplete (Run stopped before auditor finished):**
|
|
135
|
+
|
|
136
|
+
```text
|
|
137
|
+
AUDIT NOT COMPLETED: the run stopped at <phase/step> before the audit finished.
|
|
138
|
+
<exhausted-model> hit an account-wide usage limit, so no model was available to continue. The work below is UNREVIEWED, not approved. Quota resets <window>.
|
|
139
|
+
```
|
|
140
|
+
|
|
141
|
+
### Selection Rule
|
|
142
|
+
|
|
143
|
+
State 2 REQUIRES that an audit RAN TO COMPLETION on the writer's model. If the
|
|
144
|
+
audit did not complete, State 3 applies — NEVER State 2. An unfinished audit is
|
|
145
|
+
not a weak audit, it is an absent one.
|
|
146
|
+
|
|
147
|
+
Emit Provenance Table
|
|
148
|
+
(`| Phase | Role | Model | Isolation | Reason for swap | Effect on the claim |`)
|
|
149
|
+
beneath disclosure line.
|
|
150
|
+
|
|
151
|
+
## Four Pillars & Anti-Rationalization
|
|
152
|
+
|
|
153
|
+
| Rationalization | Rebuttal |
|
|
154
|
+
| ----------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ |
|
|
155
|
+
| "The plan looks fine, let me skip the rubric." | Unscored is unhardened. Produce the rubric. |
|
|
156
|
+
| "Risk-2 task can be done in Sub-Pro." | Risk-2 MUST refuse single-pass hardening. Escalate. |
|
|
157
|
+
| "The audit passed anyway, so the notice would just worry them." | A pass from the author is not a pass; the notice IS the finding. Emit disclosure line as line 1. |
|
|
158
|
+
| "The model swap was handled automatically, so it's an implementation detail." | Handling it seamlessly is why developer cannot see it, which is exactly why it must be stated. |
|
|
159
|
+
| "It is already recorded in the provenance table below." | A table row is not a disclosure; the first line is. |
|
|
160
|
+
|
|
161
|
+
## Exit States
|
|
162
|
+
|
|
163
|
+
Concludes in one of: `PASSED`, `PROVISIONAL`, `UNREVIEWED`, `PARKED`, `CAPPED`,
|
|
164
|
+
`ABORTED`.
|