@hybridlabor-api/aos 4.12.1 → 4.13.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/vendor-manifest.json +6 -0
- package/.claude/agents/reviewer.md +1 -1
- package/.claude/agents/techlead.md +1 -1
- package/.opencode/agents/architect.md +1 -1
- package/.opencode/agents/reviewer.md +1 -1
- package/.opencode/agents/techlead.md +1 -1
- package/README.de.md +10 -9
- package/README.md +10 -9
- package/README.pt.md +10 -9
- package/THIRD_PARTY_NOTICES.md +94 -0
- package/bin/aos-store.mjs +48 -45
- package/installer.js +22 -20
- package/lib/scenario-store-index.json +1 -0
- package/lib/store-shared.mjs +81 -4
- package/lib/store-ui/index.html +14 -5
- package/lib/store-ui/server.mjs +45 -22
- package/package.json +2 -2
- package/scripts/build-plugin-manifest.mjs +33 -10
- package/scripts/build-scenario-store-index.mjs +154 -0
- package/skills/bdbrainstorm/SKILL.md +1 -0
- package/skills/global_config/agent-pipeline/SKILL.md +4 -0
- package/skills/global_config/aos-store/SKILL.md +3 -3
- package/skills/global_config/design-control-loop/SKILL.md +184 -0
- package/skills/global_config/design-control-loop/references/agent-iteration.ts +174 -0
- package/skills/global_config/design-control-loop/references/agent-runner-templates.md +157 -0
- package/skills/global_config/design-control-loop/references/control-loop-taxonomy.md +75 -0
- package/skills/global_config/design-control-loop/references/example-control-loop.md +57 -0
- package/skills/global_config/design-control-loop/references/example-skill.md +171 -0
- package/skills/global_config/design-control-loop/references/memory-template.md +7 -0
- package/skills/global_config/design-control-loop/references/prompt-template.md +58 -0
- package/skills/global_config/design-control-loop/references/response-template.md +103 -0
- package/skills/global_config/design-control-loop/references/skill-template.md +57 -0
- package/skills/global_config/design-control-loop/references/workflow-template.yml +273 -0
|
@@ -0,0 +1,157 @@
|
|
|
1
|
+
# Agent Runner Templates
|
|
2
|
+
|
|
3
|
+
Use one of these inside the workflow's agent run step. Broad permission modes are appropriate only on trusted, isolated runners.
|
|
4
|
+
|
|
5
|
+
Each agent outputs differently, so response extraction varies. The goal is to get the agent's final formatted response into `/tmp/pr-body.md` for the PR body.
|
|
6
|
+
|
|
7
|
+
**Run the actuator locally before wiring CI.** Every command below is a headless CLI invocation — export the agent's secret and set `PROMPT`, then run it in your shell against a controller-selected target to confirm the actuator works before it goes into a workflow:
|
|
8
|
+
|
|
9
|
+
```bash
|
|
10
|
+
export ANTHROPIC_API_KEY=... # or the agent's secret
|
|
11
|
+
PROMPT="$(cat /tmp/agent-prompt.md)" # your assembled actuator prompt
|
|
12
|
+
# then run the agent command for your chosen agent, below
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
## Claude Code
|
|
18
|
+
|
|
19
|
+
Secret: `ANTHROPIC_API_KEY`.
|
|
20
|
+
|
|
21
|
+
```yaml
|
|
22
|
+
- uses: actions/setup-node@v4
|
|
23
|
+
with:
|
|
24
|
+
node-version: 24
|
|
25
|
+
- run: npm install -g @anthropic-ai/claude-code
|
|
26
|
+
- name: Run Claude Code
|
|
27
|
+
env:
|
|
28
|
+
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
29
|
+
run: |
|
|
30
|
+
claude -p "$PROMPT" \
|
|
31
|
+
--permission-mode bypassPermissions \
|
|
32
|
+
--output-format stream-json \
|
|
33
|
+
--verbose \
|
|
34
|
+
2>&1 | tee /tmp/agent-output.txt
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
Use `--max-turns` or `--max-budget-usd` when the repo needs spend guards.
|
|
38
|
+
|
|
39
|
+
**Response extraction:** Claude Code with `--output-format stream-json` outputs JSON lines. Extract the final assistant message:
|
|
40
|
+
|
|
41
|
+
```yaml
|
|
42
|
+
- name: Extract PR body
|
|
43
|
+
run: |
|
|
44
|
+
# Extract the last assistant text message from stream-json output
|
|
45
|
+
cat /tmp/agent-output.txt \
|
|
46
|
+
| grep '^{' \
|
|
47
|
+
| jq -s '[.[] | select(.type == "assistant" and .message.content)] | last | .message.content[] | select(.type == "text") | .text' -r \
|
|
48
|
+
> /tmp/pr-body.md
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
---
|
|
52
|
+
|
|
53
|
+
## Codex CLI
|
|
54
|
+
|
|
55
|
+
Secret: `OPENAI_API_KEY`.
|
|
56
|
+
|
|
57
|
+
```yaml
|
|
58
|
+
- uses: actions/setup-node@v4
|
|
59
|
+
with:
|
|
60
|
+
node-version: 24
|
|
61
|
+
- run: npm install -g @openai/codex
|
|
62
|
+
- name: Login Codex
|
|
63
|
+
env:
|
|
64
|
+
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
|
65
|
+
run: printenv OPENAI_API_KEY | codex login --with-api-key
|
|
66
|
+
- name: Run Codex
|
|
67
|
+
run: |
|
|
68
|
+
codex exec "$PROMPT" \
|
|
69
|
+
--cd "$GITHUB_WORKSPACE" \
|
|
70
|
+
--ask-for-approval never \
|
|
71
|
+
--sandbox danger-full-access \
|
|
72
|
+
--json \
|
|
73
|
+
--output-last-message /tmp/pr-body.md \
|
|
74
|
+
2>&1 | tee /tmp/agent-output.txt
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
**Response extraction:** Codex has built-in support via `--output-last-message /tmp/pr-body.md`. No additional extraction needed.
|
|
78
|
+
|
|
79
|
+
---
|
|
80
|
+
|
|
81
|
+
## OpenCode
|
|
82
|
+
|
|
83
|
+
Secret: provider-specific, commonly `ANTHROPIC_API_KEY` or `OPENAI_API_KEY`.
|
|
84
|
+
|
|
85
|
+
```yaml
|
|
86
|
+
- uses: oven-sh/setup-bun@v2
|
|
87
|
+
- run: bun install -g opencode-ai
|
|
88
|
+
- name: Run OpenCode
|
|
89
|
+
env:
|
|
90
|
+
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
91
|
+
run: |
|
|
92
|
+
opencode run "$PROMPT" \
|
|
93
|
+
--dir "$GITHUB_WORKSPACE" \
|
|
94
|
+
--model anthropic/claude-sonnet-4-5 \
|
|
95
|
+
--format json \
|
|
96
|
+
--dangerously-skip-permissions \
|
|
97
|
+
2>&1 | tee /tmp/agent-output.txt
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
Change `--model` and env secrets together, for example to an `openai/...` model with `OPENAI_API_KEY`.
|
|
101
|
+
|
|
102
|
+
**Response extraction:** OpenCode with `--format json` outputs structured JSON. Extract the final message:
|
|
103
|
+
|
|
104
|
+
```yaml
|
|
105
|
+
- name: Extract PR body
|
|
106
|
+
run: |
|
|
107
|
+
# Extract the last assistant message from OpenCode JSON output
|
|
108
|
+
cat /tmp/agent-output.txt \
|
|
109
|
+
| jq -r '.messages | map(select(.role == "assistant")) | last | .content' \
|
|
110
|
+
> /tmp/pr-body.md
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
---
|
|
114
|
+
|
|
115
|
+
## CodeLayer
|
|
116
|
+
|
|
117
|
+
Secret: usually `ANTHROPIC_API_KEY` for Anthropic-backed runs.
|
|
118
|
+
|
|
119
|
+
```yaml
|
|
120
|
+
- uses: oven-sh/setup-bun@v2
|
|
121
|
+
- name: Run CodeLayer
|
|
122
|
+
env:
|
|
123
|
+
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
124
|
+
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
|
125
|
+
FORCE_COLOR: "3"
|
|
126
|
+
run: |
|
|
127
|
+
bunx @humanlayer/cli@latest codelayer \
|
|
128
|
+
--provider anthropic \
|
|
129
|
+
--model claude-opus-4-8 \
|
|
130
|
+
--thinking high \
|
|
131
|
+
--prompt "$PROMPT" \
|
|
132
|
+
2>&1 | tee /tmp/agent-output.txt
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
Use this when the repo already depends on HumanLayer/CodeLayer conventions or wants CodeLayer's PR-oriented behavior.
|
|
136
|
+
|
|
137
|
+
**Response extraction:** CodeLayer outputs plain text with ANSI colors. Use a parser script or strip formatting:
|
|
138
|
+
|
|
139
|
+
```yaml
|
|
140
|
+
- name: Extract PR body
|
|
141
|
+
run: |
|
|
142
|
+
# If you have a codelayer-output parser script:
|
|
143
|
+
# bun ci-scripts/codelayer-output.ts < /tmp/agent-output.txt > /tmp/pr-body.md
|
|
144
|
+
|
|
145
|
+
# Otherwise, strip ANSI codes and use the raw output:
|
|
146
|
+
cat /tmp/agent-output.txt | sed 's/\x1b\[[0-9;]*m//g' > /tmp/pr-body.md
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
For cleaner extraction, use a parser script like `ci-scripts/codelayer-output.ts` that extracts just the final response section.
|
|
150
|
+
|
|
151
|
+
---
|
|
152
|
+
|
|
153
|
+
## Notes
|
|
154
|
+
|
|
155
|
+
- All agents should tee output to `/tmp/agent-output.txt` for artifact upload and debugging.
|
|
156
|
+
- The extracted response goes to `/tmp/pr-body.md` which is used by the PR creation step.
|
|
157
|
+
- If extraction fails, the workflow should fall back gracefully (e.g., use raw output or a placeholder message).
|
|
@@ -0,0 +1,75 @@
|
|
|
1
|
+
# Control Loop Taxonomy
|
|
2
|
+
|
|
3
|
+
Use this to explain an agent loop as a control system before designing one, so it is observable, bounded, and reviewable instead of "an agent runs sometimes."
|
|
4
|
+
|
|
5
|
+
Every example below is an **illustration to spark discussion**, not a recommendation. The right sensor, controller, and actuator depend entirely on the user's codebase and the tooling they already use — discover that in the interview, don't assume it.
|
|
6
|
+
|
|
7
|
+
## The four components
|
|
8
|
+
|
|
9
|
+
Foreground these four (plus the set point they serve). Keep the rest of the vocabulary in your back pocket.
|
|
10
|
+
|
|
11
|
+
- **Set point** — the desired end state for some property of the codebase. An invariant ("no module imports across these boundaries"), a threshold ("coverage ≥ 80% in `core`"), or a direction ("fewer occurrences each run").
|
|
12
|
+
|
|
13
|
+
- **Sensor** — how the loop measures the current state and the gap to the set point. It can be almost anything that reports on the codebase: a static-analysis or lint tool, a structural/AST search, a type checker, a test suite, a telemetry or error query, a code-search query, a custom script — or an agent that inspects the code. Trade-offs to talk through with the user (not rules to impose): how stable and repeatable the measurement is, how much it costs to run, and whether it can be silently disabled. Aim for output a controller can act on repeatably.
|
|
14
|
+
|
|
15
|
+
- **Controller** — how the loop turns the measurement into the next change, sized to stay low-risk and reviewable. It decides *what to do now versus defer*: which target, how many, in what order. It ranges from fully deterministic (a script that sorts findings and picks one) to fully agentic (an agent choosing from natural-language criteria), with data-driven variants in between (e.g. prioritize by where production errors cluster). This is the part you **tune over time** from loop output — start simple.
|
|
16
|
+
|
|
17
|
+
- **Actuator** — what applies the change: a coding agent (Claude Code, Codex, OpenCode, CodeLayer, …) plus a repo-local skill, running in CI and opening a PR.
|
|
18
|
+
|
|
19
|
+
- **Disturbance** — anything that changes the system from outside the loop: teammates' commits, dependency updates, generated code, flaky tests, large refactors. The loop has to make progress *despite* these.
|
|
20
|
+
|
|
21
|
+
## Components can blur
|
|
22
|
+
|
|
23
|
+
In practice — especially with agents — the lines blur, and that is fine:
|
|
24
|
+
|
|
25
|
+
- **Sensor + controller fused:** a tool that both reports problems and ranks them by impact is doing both jobs.
|
|
26
|
+
- **Controller + actuator fused:** a single agent prompt that both picks the next target and changes it.
|
|
27
|
+
|
|
28
|
+
Design the loop the user actually needs; don't manufacture separation that isn't there.
|
|
29
|
+
|
|
30
|
+
## The fuller picture
|
|
31
|
+
|
|
32
|
+
You rarely need to name all of this explicitly, but it helps when reasoning about a loop:
|
|
33
|
+
|
|
34
|
+
```mermaid
|
|
35
|
+
flowchart LR
|
|
36
|
+
SetPoint[Set point] --> Compare((Compare))
|
|
37
|
+
Measured[Measured output] --> Compare
|
|
38
|
+
Compare --> Error[Measured error]
|
|
39
|
+
Error --> Controller[Controller]
|
|
40
|
+
Controller --> Actuator[Actuator]
|
|
41
|
+
Actuator --> System[System / Repository]
|
|
42
|
+
Disturbance[Disturbance] --> System
|
|
43
|
+
System --> Output[System output]
|
|
44
|
+
Output --> Sensor[Sensor]
|
|
45
|
+
Sensor --> Measured
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
- **Measured output** — the signal the sensor produces (e.g. a list of findings, a count, a score).
|
|
49
|
+
- **Measured error** — the delta between set point and measured output (e.g. findings over threshold, or new findings vs a baseline).
|
|
50
|
+
- **Controller output** — the chosen operation for this iteration (e.g. "address these N targets in these packages").
|
|
51
|
+
- **System input / output** — the patch/commit/PR the actuator produces, and the resulting repo state plus validation and review.
|
|
52
|
+
|
|
53
|
+
## Run each component locally first
|
|
54
|
+
|
|
55
|
+
Whatever the components turn out to be, make each one runnable by hand and standalone before wiring CI: run the sensor and read its output, run the controller against that output, run the actuator on a selected target. The workflow should only orchestrate pieces the user can already run locally — this keeps the loop debuggable.
|
|
56
|
+
|
|
57
|
+
## Extra loop parts to consider
|
|
58
|
+
|
|
59
|
+
- **Flow control (PR bounding)** — stop scheduled runs when an open PR for this loop already exists, so the loop doesn't outrun review.
|
|
60
|
+
- **Dampener (regression gate)** — a check (often on PRs / pushes to main) that compares the sensor's output against a baseline so the problem can't get *worse* while the loop incrementally makes it better. Offer it; not every loop needs one.
|
|
61
|
+
- **Scope gate** — restrict the actuator to safe directories; exclude generated files, vendored code, or high-risk packages unless explicitly selected.
|
|
62
|
+
- **Batch size** — cap each run by finding/file/package count to keep diffs reviewable.
|
|
63
|
+
- **Memory** — durable reviewer feedback and known false positives that steer future runs (not one-off logs).
|
|
64
|
+
|
|
65
|
+
## Design questions for the interview
|
|
66
|
+
|
|
67
|
+
1. What property are we driving, and what is the set point?
|
|
68
|
+
2. What in this repo (or its tooling) can measure the gap repeatably? What are the trade-offs of each option?
|
|
69
|
+
3. What counts as an error worth acting on this run, and how big is one reviewable increment?
|
|
70
|
+
4. How should the controller prioritize and select targets — and how will we tune that over time?
|
|
71
|
+
5. Which coding agent is the actuator, what credentials does it need, and what golden patterns should it follow?
|
|
72
|
+
6. What validation proves the actuator improved the system?
|
|
73
|
+
7. What disturbances should the loop ignore, tolerate, or dampen — and do we want a regression gate?
|
|
74
|
+
8. What memory should carry forward between runs?
|
|
75
|
+
9. How do we run each component locally before it goes into CI?
|
|
@@ -0,0 +1,57 @@
|
|
|
1
|
+
# Example Control Loop: React Doctor
|
|
2
|
+
|
|
3
|
+
One fully worked loop, to make the taxonomy concrete. **This is an illustration, not a template.** A real loop from a production monorepo, it drives React code quality in a single app. Your loop's set point, sensor, controller, and actuator will look different — copy the *shape*, not the specifics.
|
|
4
|
+
|
|
5
|
+
It is a useful example because the sensor and controller come almost entirely from an existing, configurable tool, so the whole loop is small.
|
|
6
|
+
|
|
7
|
+
## Set point
|
|
8
|
+
|
|
9
|
+
`apps/riptide-ui` stays free of high-impact React issues (lint, accessibility, correctness, architecture). A direction more than a fixed threshold: each run leaves the app a little healthier.
|
|
10
|
+
|
|
11
|
+
## Sensor — `react-doctor` + `doctor.config.ts`
|
|
12
|
+
|
|
13
|
+
The [`react-doctor`](https://github.com/millionco/react-doctor) CLI scans the app and reports prioritized issues; `doctor.config.ts` configures which rules run and which paths/rules are ignored. It runs locally exactly as it does in CI:
|
|
14
|
+
|
|
15
|
+
```bash
|
|
16
|
+
bunx react-doctor --project '@codelayer/riptide-ui' --diff false --yes
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
Chosen because it is repeatable, configurable, and lives outside the editor/lint config so a stray inline comment can't quietly switch it off — a trade-off that mattered *for this team*, not a requirement of all sensors.
|
|
20
|
+
|
|
21
|
+
## Controller — fused with the sensor
|
|
22
|
+
|
|
23
|
+
There is no separate controller: `react-doctor` returns "the top 3 rules by impact," and the loop's policy is "fix up to 5 issues from those top 3 rules this run." That selection logic lives in the actuator's prompt. This is the **sensor + controller blur** — one tool plus a small policy does both jobs.
|
|
24
|
+
|
|
25
|
+
## Actuator — CodeLayer + a repo-local skill
|
|
26
|
+
|
|
27
|
+
A CodeLayer agent runs with the repo's `react-doctor` skill. Its per-issue loop gives the agent three honest options — **fix**, **ignore** (add to `doctor.config.ts` with a reason), or **skip** (leave for a human) — and validates each change before committing it separately:
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
bun run typecheck
|
|
31
|
+
bun run quality
|
|
32
|
+
bunx react-doctor --project '@codelayer/riptide-ui' --staged # no new issues in staged files
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
## Disturbances + dampener
|
|
36
|
+
|
|
37
|
+
**Disturbance:** teammates ship React code continuously while the loop runs.
|
|
38
|
+
|
|
39
|
+
**Dampener:** a second workflow (`react-doctor.yml`) runs on every pull request and on pushes to `main`. It diffs against the merge base and comments on only the *newly introduced* issues. It is **advisory by default** (never red-Xes a teammate's PR), with a documented path to graduate to blocking once the team trusts the signal. This keeps the problem from getting worse while the scheduled loop chips away at it.
|
|
40
|
+
|
|
41
|
+
## The loop — `agent-react-doctor.yml`
|
|
42
|
+
|
|
43
|
+
A scheduled workflow (daily, plus manual dispatch and `/iterate`) runs the loop and opens a PR whose body is the agent's final message. Because the sensor, controller, and actuator are fused into one agent step here, the workflow does **not** have three separate steps — it collapses them, matching the design. A loop with a deterministic, standalone sensor and controller would instead have discrete steps.
|
|
44
|
+
|
|
45
|
+
## Human on the loop
|
|
46
|
+
|
|
47
|
+
- `.github/agent-memory/react-doctor.md` is loaded into the agent every run — standing feedback like "always use the `no-use-effect` skill for `useEffect` fixes" and "don't globally ignore a rule when only specific files need an exemption."
|
|
48
|
+
- Maintainers comment `/iterate` on the PR; a hidden marker in the PR body routes the comment to the workflow that created it, which loads the PR context and feedback and updates the memory file and the PR. Deterministic glue lives in `ci-scripts/` (`agent-iteration.ts`, `codelayer-output.ts`).
|
|
49
|
+
|
|
50
|
+
## Flow control
|
|
51
|
+
|
|
52
|
+
Every PR is labeled `agent-react-doctor`. Scheduled runs no-op when an open PR with that label already exists, so there is at most one open PR per loop; manual dispatch bypasses the gate.
|
|
53
|
+
|
|
54
|
+
## What to take from this
|
|
55
|
+
|
|
56
|
+
- **Reusable shape:** set point → sensor → controller → actuator under disturbances, plus a dampener, a memory file, `/iterate` steering, and one-PR flow control.
|
|
57
|
+
- **Not reusable:** `react-doctor`, the bun/CodeLayer commands, the "top 3 rules / 5 fixes" policy, the `riptide-ui` scope. Those are tailored to this repo — yours come from your interview.
|
|
@@ -0,0 +1,171 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: narrow-react-prop-types
|
|
3
|
+
description: narrow React component prop types to match live code paths
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Narrow React Prop Types
|
|
7
|
+
|
|
8
|
+
Use this skill when a React component's props have been widened for stories, mocks, tests, or demos and now express states the live application does not enter.
|
|
9
|
+
|
|
10
|
+
The goal is to make component types describe the real live-code-path contract, then require stories/tests/mocks to adapt to that contract instead of weakening it.
|
|
11
|
+
|
|
12
|
+
## Core Requirements
|
|
13
|
+
|
|
14
|
+
- Find the actual non-test, non-Storybook call sites before changing types.
|
|
15
|
+
- Treat live code paths as the source of truth for the prop contract.
|
|
16
|
+
- Do not preserve optional props only because they make Storybook, tests, or mock data easier.
|
|
17
|
+
- Keep props optional only when there are non-Storybook, non-test call sites which do not provide them and which have a good reason for not doing so.
|
|
18
|
+
- Types should not enable expressing states which are not observed in non-test, non-Storybook call sites.
|
|
19
|
+
- Types should be as strict as possible so code can be as simple as possible.
|
|
20
|
+
- Prefer deriving and extracting types from existing live-code-path values and APIs where possible.
|
|
21
|
+
|
|
22
|
+
## Workflow
|
|
23
|
+
|
|
24
|
+
For an example recurring GitHub Actions workflow that runs this skill through CodeLayer, see `references/agent-narrow-component-props.yml`. Its example agent memory file is `references/narrow-component-props-memory.md`.
|
|
25
|
+
|
|
26
|
+
### 1. Identify the suspect component
|
|
27
|
+
|
|
28
|
+
Look for components with these signals:
|
|
29
|
+
|
|
30
|
+
- Large props interfaces with many optional fields.
|
|
31
|
+
- Optional callback calls such as `onSelect?.(...)` or `onArchive?.(...)`.
|
|
32
|
+
- Fallback state handling such as `items ?? []`, `count ?? 0`, or `handler && ...` around values live code likely always supplies.
|
|
33
|
+
- UI affordances that always render even though their callbacks are optional.
|
|
34
|
+
- Props that look demo-oriented, such as `defaultFoo`, alternate handler shapes, or display toggles not used by live code.
|
|
35
|
+
|
|
36
|
+
Do not pick a target from a story or test alone. Use stories/tests only as supporting evidence that the type has been widened, not as evidence that a state is real.
|
|
37
|
+
|
|
38
|
+
### 2. Find every live usage
|
|
39
|
+
|
|
40
|
+
Search for all imports/usages of the component, exported prop type, and shared child primitives.
|
|
41
|
+
|
|
42
|
+
Classify call sites by whether they are live code paths or support code:
|
|
43
|
+
|
|
44
|
+
- Live code paths: app routes, wired components, providers, hooks, production package exports, and shared components used by those paths.
|
|
45
|
+
- Support code: Storybook stories, test files, fixtures, mocks, demo harnesses, and visual-only examples.
|
|
46
|
+
|
|
47
|
+
Only live code paths should determine what the component API supports.
|
|
48
|
+
|
|
49
|
+
### 3. Derive the real types from the live code paths
|
|
50
|
+
|
|
51
|
+
Read the live call sites and classify each prop:
|
|
52
|
+
|
|
53
|
+
- Required: every non-test, non-Storybook call site supplies it.
|
|
54
|
+
- Optional: at least one non-test, non-Storybook call site omits it and that omission is a meaningful runtime state.
|
|
55
|
+
- Removed: no non-test, non-Storybook call site uses it.
|
|
56
|
+
|
|
57
|
+
Nullability and optionality are different. If live code always passes a prop but the value can be empty, prefer a required nullable prop such as `focusedItem: FocusedItem | null` over `focusedItem?: FocusedItem | null`.
|
|
58
|
+
|
|
59
|
+
### 4. Tighten the public prop type
|
|
60
|
+
|
|
61
|
+
Update exported prop types to match only the states observed in live code paths.
|
|
62
|
+
|
|
63
|
+
The looser and more optional a type is, the more possible states the component has to reason about. Every optional prop creates another branch the component must handle, test, and keep correct. Prefer strict types that prevent impossible states instead of broad types that require defensive render logic.
|
|
64
|
+
|
|
65
|
+
If the component always renders an interactive affordance, require the handler that makes it work. Do not allow inert states like a visible menu item that calls `onRename?.(...)`.
|
|
66
|
+
|
|
67
|
+
### 5. Derive and extract types where possible
|
|
68
|
+
|
|
69
|
+
Prefer deriving types from the live APIs instead of restating them manually:
|
|
70
|
+
|
|
71
|
+
- `Parameters<typeof fn>[0]` for function argument types.
|
|
72
|
+
- `ReturnType<typeof fn>` for return types.
|
|
73
|
+
- `Extract<Union, Shape>` for narrowing a union to a real variant.
|
|
74
|
+
- `React.Dispatch<React.SetStateAction<T>>` for React state setters instead of approximating them as `(value: T) => void`.
|
|
75
|
+
|
|
76
|
+
Prefer explicit state type parameters when inference would widen or obscure the intended state shape:
|
|
77
|
+
|
|
78
|
+
```ts
|
|
79
|
+
const [dialogState, setDialogState] = useState<DialogState>({
|
|
80
|
+
id: null,
|
|
81
|
+
isOpen: false,
|
|
82
|
+
})
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
Avoid relying on implicit `useState(...)` inference when it produces broad nullable object shapes, string literal widening, or callback types that later need hand-written approximations.
|
|
86
|
+
|
|
87
|
+
### 6. Tighten internal child props too
|
|
88
|
+
|
|
89
|
+
Do not stop at the exported component if it passes broad props into child primitives.
|
|
90
|
+
|
|
91
|
+
If row/menu/button child components receive optional handlers only because the parent props were broad, tighten those internal props too. Replace optional calls like this:
|
|
92
|
+
|
|
93
|
+
```ts
|
|
94
|
+
onRename?.(id, name)
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
with required calls:
|
|
98
|
+
|
|
99
|
+
```ts
|
|
100
|
+
onRename(id, name)
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
### 7. Remove fallback logic for unsupported states
|
|
104
|
+
|
|
105
|
+
Once props are required, remove defensive fallbacks that only existed for widened types.
|
|
106
|
+
|
|
107
|
+
Examples:
|
|
108
|
+
|
|
109
|
+
```ts
|
|
110
|
+
new Set(expandedIds ?? defaultExpandedIds ?? [])
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
should become:
|
|
114
|
+
|
|
115
|
+
```ts
|
|
116
|
+
new Set(expandedIds)
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
```ts
|
|
120
|
+
items && items.length > 0
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
should become:
|
|
124
|
+
|
|
125
|
+
```ts
|
|
126
|
+
items.length > 0
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
### 8. Update all variants that share the prop type
|
|
130
|
+
|
|
131
|
+
If multiple components share the broad prop type, update them together so they all enforce the same live-code-path contract.
|
|
132
|
+
|
|
133
|
+
### 9. Let tests and stories adapt to live code
|
|
134
|
+
|
|
135
|
+
If a story or test breaks after narrowing props, fix it by providing realistic handlers and state. Do not make live-code-path props optional again to reduce test setup.
|
|
136
|
+
|
|
137
|
+
If the story/test setup feels verbose, create a test helper or fixture that satisfies the strict live-code-path contract. Keep the helper in support code; do not weaken the component API.
|
|
138
|
+
|
|
139
|
+
### 10. Validate the change
|
|
140
|
+
|
|
141
|
+
Run package-level typechecks for the changed package and each live app/package that consumes the changed component.
|
|
142
|
+
|
|
143
|
+
Use repository-specific validation commands when available. In this monorepo, prefer:
|
|
144
|
+
|
|
145
|
+
```bash
|
|
146
|
+
bun --bun run typecheck --filter <package>
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
### 11. Read the response template and format your final asnwer
|
|
150
|
+
Read the response template at `references/final-response-template.md` and format your answer accordingly.
|
|
151
|
+
|
|
152
|
+
## Review Checklist
|
|
153
|
+
|
|
154
|
+
- The changed prop type was derived from non-test, non-Storybook call sites.
|
|
155
|
+
- Optional callbacks are removed for always-rendered interactions.
|
|
156
|
+
- Rendered menu items and buttons cannot be inert because of missing handlers.
|
|
157
|
+
- Removed props are not used by live code paths.
|
|
158
|
+
- Nullability is preserved only for real states, such as no current focus.
|
|
159
|
+
- Types are derived or extracted where possible rather than manually duplicated.
|
|
160
|
+
- `useState<T>(...)` is used where inference would otherwise widen or obscure the intended state.
|
|
161
|
+
- Shared variants compile against the same narrowed contract.
|
|
162
|
+
- Typecheck passes for the shared package and consuming live app/package.
|
|
163
|
+
|
|
164
|
+
## Anti-Patterns to Avoid
|
|
165
|
+
|
|
166
|
+
- Making callbacks optional so stories can omit them.
|
|
167
|
+
- Rendering a menu item that calls `onAction?.(...)`.
|
|
168
|
+
- Adding `default*` props for Storybook when live code is controlled.
|
|
169
|
+
- Using `?? []` or `?? 0` to hide missing required live state.
|
|
170
|
+
- Accepting multiple API shapes when live code only uses one.
|
|
171
|
+
- Treating pure components as mock components with relaxed contracts.
|
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
# Agent Memory: <Task Title>
|
|
2
|
+
|
|
3
|
+
Standing feedback for future `Agent: <Task Title>` runs. This is the human-on-the-loop steering channel: it is loaded into the actuator on every run (after the controller), so edits here change future behavior, not just one PR. Keep durable guidance only — not one-off instructions or single-run logs.
|
|
4
|
+
|
|
5
|
+
## Guidance
|
|
6
|
+
|
|
7
|
+
- <Permanent scope constraint, false-positive area, or reviewer preference that should change future runs.>
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# Context
|
|
2
|
+
|
|
3
|
+
You are <task summary> in this repository. Begin by using the `<skill-name>` skill.
|
|
4
|
+
|
|
5
|
+
# Scope
|
|
6
|
+
|
|
7
|
+
Focus only on `<primary path or package>`.
|
|
8
|
+
|
|
9
|
+
You may inspect `<secondary path>` only when it is necessary to understand or validate the primary change.
|
|
10
|
+
|
|
11
|
+
# Instructions
|
|
12
|
+
|
|
13
|
+
1. <Work the target(s) the controller selected (see "Selected work" if present); otherwise find high-confidence targets yourself.>
|
|
14
|
+
2. <Keep this to one small, reviewable increment.>
|
|
15
|
+
3. <Use the real source of truth, not support-only examples.>
|
|
16
|
+
4. <Avoid adjacent cleanup that belongs to another workflow.>
|
|
17
|
+
5. Validate with the commands below.
|
|
18
|
+
6. Commit and push your changes.
|
|
19
|
+
|
|
20
|
+
## Validation Commands
|
|
21
|
+
|
|
22
|
+
```bash
|
|
23
|
+
<validation command 1>
|
|
24
|
+
<validation command 2>
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
## Important Rules
|
|
28
|
+
|
|
29
|
+
- <Rule that prevents the most likely wrong change.>
|
|
30
|
+
- <Rule that keeps scope narrow.>
|
|
31
|
+
- Do not run long integration tests unless explicitly requested.
|
|
32
|
+
- You are running in a sandbox or CI runner environment. Do not stop and ask for feedback from the user or request approvals.
|
|
33
|
+
|
|
34
|
+
## Finishing Up
|
|
35
|
+
|
|
36
|
+
When you are finished:
|
|
37
|
+
|
|
38
|
+
1. Make sure the validation commands pass.
|
|
39
|
+
2. Commit and push your changes.
|
|
40
|
+
3. Answer with the output format below.
|
|
41
|
+
|
|
42
|
+
## Output Format
|
|
43
|
+
|
|
44
|
+
Format your final answer as GitHub-flavored markdown:
|
|
45
|
+
|
|
46
|
+
```markdown
|
|
47
|
+
## <Task Title> Complete
|
|
48
|
+
|
|
49
|
+
### Changes Made
|
|
50
|
+
- [file/component]: [what changed and why]
|
|
51
|
+
|
|
52
|
+
### Source Of Truth Checked
|
|
53
|
+
- [path]: [why it supports the change]
|
|
54
|
+
|
|
55
|
+
### Validation
|
|
56
|
+
- [x] <validation command 1> passed
|
|
57
|
+
- [x] <validation command 2> passed
|
|
58
|
+
```
|
|
@@ -0,0 +1,103 @@
|
|
|
1
|
+
# Response Template
|
|
2
|
+
|
|
3
|
+
This template defines how the CI agent should format its final response, which becomes the body of the GitHub PR.
|
|
4
|
+
|
|
5
|
+
Customize this template based on the task type and what information reviewers need.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## Example: Fix/Migration Task
|
|
10
|
+
|
|
11
|
+
```markdown
|
|
12
|
+
Finished processing N `<tool-name>` issues:
|
|
13
|
+
- `resolved`: X
|
|
14
|
+
- `ignored`: Y
|
|
15
|
+
- `skipped`: Z
|
|
16
|
+
|
|
17
|
+
Details below:
|
|
18
|
+
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
## Issue 1
|
|
22
|
+
- **File(s)**: `path/to/file.ts`
|
|
23
|
+
- **Summary**: [concise description of the issue]
|
|
24
|
+
- **Resolution**: **FIXED**. [concise description of the fix]
|
|
25
|
+
- **Risk**: [HIGH / medium / low]
|
|
26
|
+
- **Manual verification** (for HIGH/medium risk): [steps to test the changes]
|
|
27
|
+
|
|
28
|
+
## Issue 2
|
|
29
|
+
- **File(s)**: `path/to/other.ts`
|
|
30
|
+
- **Summary**: [...]
|
|
31
|
+
- **Resolution**: **IGNORED**. [why the rule was added to ignore list]
|
|
32
|
+
|
|
33
|
+
## Issue 3
|
|
34
|
+
- **File(s)**: `path/to/another.ts`
|
|
35
|
+
- **Summary**: [...]
|
|
36
|
+
- **Resolution**: **SKIPPED**. [why this needs human attention]
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
---
|
|
40
|
+
|
|
41
|
+
## Example: Generation Task
|
|
42
|
+
|
|
43
|
+
```markdown
|
|
44
|
+
Generated N new <artifact type>:
|
|
45
|
+
|
|
46
|
+
| File | Description | Status |
|
|
47
|
+
|------|-------------|--------|
|
|
48
|
+
| `path/to/new-file.ts` | [what was generated] | Created |
|
|
49
|
+
| `path/to/updated.ts` | [what changed] | Updated |
|
|
50
|
+
|
|
51
|
+
## Validation
|
|
52
|
+
- Typecheck: PASS
|
|
53
|
+
- Tests: PASS
|
|
54
|
+
- Lint: PASS
|
|
55
|
+
|
|
56
|
+
## Notes
|
|
57
|
+
[Any context reviewers should know]
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
---
|
|
61
|
+
|
|
62
|
+
## Example: Refactor Task
|
|
63
|
+
|
|
64
|
+
```markdown
|
|
65
|
+
Refactored N components/modules:
|
|
66
|
+
|
|
67
|
+
## Changes
|
|
68
|
+
|
|
69
|
+
### `path/to/file.ts`
|
|
70
|
+
- **Before**: [brief description of old pattern]
|
|
71
|
+
- **After**: [brief description of new pattern]
|
|
72
|
+
- **Risk**: low
|
|
73
|
+
|
|
74
|
+
### `path/to/other.ts`
|
|
75
|
+
- **Before**: [...]
|
|
76
|
+
- **After**: [...]
|
|
77
|
+
- **Risk**: medium
|
|
78
|
+
- **Verify**: [how to test this change]
|
|
79
|
+
|
|
80
|
+
## Validation
|
|
81
|
+
- Typecheck: PASS
|
|
82
|
+
- Tests: PASS
|
|
83
|
+
- No behavior changes expected
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
---
|
|
87
|
+
|
|
88
|
+
## Template Variables
|
|
89
|
+
|
|
90
|
+
When writing your response template, you can use these placeholders:
|
|
91
|
+
|
|
92
|
+
- `<tool-name>` - The CLI tool or process that found issues
|
|
93
|
+
- `<artifact-type>` - What the agent creates (tests, types, docs, etc.)
|
|
94
|
+
- `<task-name>` - The name of the agent task
|
|
95
|
+
- `<date>` - Current date (MM/DD format)
|
|
96
|
+
|
|
97
|
+
## Guidelines
|
|
98
|
+
|
|
99
|
+
1. **Lead with summary stats** - Reviewers should know scope immediately
|
|
100
|
+
2. **Group by resolution type** - Fixed, ignored, skipped, or created
|
|
101
|
+
3. **Include risk levels and highlight prominently** - Help reviewers prioritize their review
|
|
102
|
+
4. **Provide verification steps** - For medium/high risk changes
|
|
103
|
+
5. **Keep it scannable and concise** - Use tables, headers, and bullet points.
|