aiblueprint-cli 1.4.103 → 1.4.105
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +2 -7
- package/agents-config/skills/audit-skills/SKILL.md +79 -0
- package/agents-config/skills/audit-skills/agents/openai.yaml +7 -0
- package/agents-config/skills/audit-skills/scripts/audit-skills.mjs +346 -0
- package/package.json +1 -1
- package/agents-config/skills/audit/SKILL.md +0 -126
- package/agents-config/skills/audit/agents/openai.yaml +0 -10
- package/agents-config/skills/audit/assets/codex-icon.svg +0 -20
- package/agents-config/skills/commit/SKILL.md +0 -44
- package/agents-config/skills/commit/agents/openai.yaml +0 -10
- package/agents-config/skills/commit/assets/codex-icon.svg +0 -17
- package/agents-config/skills/create-pr/SKILL.md +0 -55
- package/agents-config/skills/create-pr/agents/openai.yaml +0 -10
- package/agents-config/skills/create-pr/assets/codex-icon.svg +0 -17
- package/agents-config/skills/oneshot/SKILL.md +0 -44
- package/agents-config/skills/oneshot/agents/openai.yaml +0 -10
- package/agents-config/skills/oneshot/assets/codex-icon.svg +0 -18
- package/agents-config/skills/tools/SKILL.md +0 -149
- package/agents-config/skills/use-artifacts/SKILL.md +0 -211
- package/agents-config/skills/use-artifacts/agents/openai.yaml +0 -7
- package/agents-config/skills/use-artifacts/assets/codex-icon.svg +0 -18
- package/agents-config/skills/use-artifacts/assets/local-runtime.js +0 -299
- package/agents-config/skills/use-artifacts/scripts/create_artifact.py +0 -317
- package/agents-config/skills/use-delegate/SKILL.md +0 -97
- package/agents-config/skills/use-delegate/agents/openai.yaml +0 -10
- package/agents-config/skills/use-delegate/assets/codex-icon.svg +0 -20
- package/agents-config/skills/use-delegate/references/models.md +0 -32
- package/agents-config/skills/use-goal/SKILL.md +0 -121
- package/agents-config/skills/use-goal/agents/openai.yaml +0 -7
- package/agents-config/skills/use-goal/assets/codex-icon.svg +0 -18
- package/agents-config/skills/use-goal/references/claude-code-goal.md +0 -65
- package/agents-config/skills/use-goal/references/codex-goal.md +0 -70
- package/agents-config/skills/use-goal/references/verification-harnesses.md +0 -108
|
@@ -1,97 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: use-delegate
|
|
3
|
-
description: "Delegation mode: the host agent (Claude or Codex) plans and reviews while heavy work runs on cheap executors: OpenCode Kimi K3, Codex GPT-5.6 terra/sol. Use when the user invokes /use-delegate, says 'use delegate', 'delegate mode', 'orchestrator mode', or wants to save tokens/rate limits."
|
|
4
|
-
disable-model-invocation: true
|
|
5
|
-
metadata:
|
|
6
|
-
opencode/autoinvoke: "false"
|
|
7
|
-
opencode/slash: "true"
|
|
8
|
-
---
|
|
9
|
-
|
|
10
|
-
# Use Delegate
|
|
11
|
-
|
|
12
|
-
The host agent (Claude Code or Codex, whichever is running this skill) is the thinker, never the typist. Its tokens are scarce; the executors below are cheap and steerable. The host decides **what** to do and judges **whether it was done well**: everything token-hungry runs elsewhere and reports back.
|
|
13
|
-
|
|
14
|
-
## Core rule
|
|
15
|
-
|
|
16
|
-
The host does not execute. It may: read a few targeted files, search, inspect git state, think, plan, decompose, write specs and delegation prompts, review diffs and reports, judge outputs, and talk to the user.
|
|
17
|
-
|
|
18
|
-
The host must NOT directly do:
|
|
19
|
-
|
|
20
|
-
- Implementation, refactors, migrations, test writing (any multi-file or >~15-line change)
|
|
21
|
-
- Codebase-wide exploration or analysis (reading many files to "understand")
|
|
22
|
-
- Computer use, browser automation, UI/UX verification
|
|
23
|
-
- Log triage, data analysis, bulk mechanical edits
|
|
24
|
-
- Running long build/test loops and reading their full output
|
|
25
|
-
|
|
26
|
-
The only direct edits allowed: trivial single-file tweaks (a config value, a typo, a one-liner) where writing a delegation prompt would cost more than the edit itself.
|
|
27
|
-
|
|
28
|
-
## Executors
|
|
29
|
-
|
|
30
|
-
| Executor | Command | Use for |
|
|
31
|
-
|---|---|---|
|
|
32
|
-
| OpenCode · Kimi K3 | `opencode run "<prompt>" -m kimi-for-coding/k3` | Default general executor: implementation, refactors, tests, bulk edits |
|
|
33
|
-
| Codex · GPT-5.6 terra (high) | `codex exec -m gpt-5.6-terra "<prompt>"` | Low-stakes tasks: mechanical edits, scripts, quick investigations, log triage |
|
|
34
|
-
| Codex · GPT-5.6 sol (high) | `codex exec "<prompt>"` (config default = sol + high) | Compute-heavy tasks: hard bugs, migrations, architecture-sensitive changes, computer use / UI verification |
|
|
35
|
-
| Host-native subagents | Claude `Agent` tool / Codex collab threads | Exploration summaries the host plans from; taste-sensitive work (UI, copy, API design) |
|
|
36
|
-
|
|
37
|
-
Read-only investigation: `codex exec -s read-only`, `opencode run --agent plan` (built-in read-only agent).
|
|
38
|
-
|
|
39
|
-
Model rankings move fast: current DeepSWE scores, API pricing, and the refresh protocol live in `references/models.md`. Check its `Last verified` date before planning a big batch; if older than 14 days, delegate a refresh first (DeepSWE leaderboard + pricing pages), never guess rankings from memory.
|
|
40
|
-
|
|
41
|
-
## Invocation mechanics
|
|
42
|
-
|
|
43
|
-
Both CLIs: always end the command with `< /dev/null` and run in background (codex hangs forever on open stdin: full codex mechanics in `~/.claude/rules/launch-codex.md`).
|
|
44
|
-
|
|
45
|
-
**Codex** (background):
|
|
46
|
-
|
|
47
|
-
```bash
|
|
48
|
-
codex exec -C <repo-root> -m gpt-5.6-terra \
|
|
49
|
-
--output-last-message <scratchpad>/codex-<task>.md \
|
|
50
|
-
"<self-contained prompt>" < /dev/null
|
|
51
|
-
```
|
|
52
|
-
|
|
53
|
-
Effort override: `-c model_reasoning_effort=high`. Non-git dir: `--skip-git-repo-check`.
|
|
54
|
-
|
|
55
|
-
**OpenCode** (background):
|
|
56
|
-
|
|
57
|
-
```bash
|
|
58
|
-
opencode run "<self-contained prompt>" \
|
|
59
|
-
-m kimi-for-coding/k3 --title "<task>" \
|
|
60
|
-
> <scratchpad>/oc-<task>.log 2>&1 < /dev/null
|
|
61
|
-
```
|
|
62
|
-
|
|
63
|
-
- Final answer = tail of the log; `--format json` for machine-readable events.
|
|
64
|
-
- Steer or continue a session: `opencode run -s <sessionID> "<follow-up>"`.
|
|
65
|
-
- Standalone specialized agent: `--agent <name>`: the `~/.config/opencode/agent/` roster (worker, explore-fast, verifier, snipper, code-reviewer…) runs standalone with any `-m` model.
|
|
66
|
-
- Permissions are pre-allowed in the user config (build agent allows all); no interactive prompt will block a non-interactive run.
|
|
67
|
-
|
|
68
|
-
## Batch / multi-process
|
|
69
|
-
|
|
70
|
-
- **Parallel processes** (verified): launch N independent `opencode run` / `codex exec` in background; each opencode run spins up its own server + session, results stay isolated.
|
|
71
|
-
- **Shared server** for large batches: `opencode serve --port <p>` once, then N × `opencode run --attach http://localhost:<p> --dir <workdir> ...`: one server, many sessions, less startup overhead. Kill the serve process when done.
|
|
72
|
-
- **In-executor subagents**: Kimi in opencode spawns its own task-tool subagents; codex spawns collab threads (config caps 6). Prefer one executor process per independent task over one giant prompt.
|
|
73
|
-
|
|
74
|
-
## The loop
|
|
75
|
-
|
|
76
|
-
1. **Think.** Understand the request. Missing context → delegate the exploration, think on the summaries.
|
|
77
|
-
2. **Spec.** Write a self-contained delegation prompt: exact files, goal, constraints, and what "done" looks like (tests pass, lint clean, behavior X). The executor can't see this conversation: spell everything out.
|
|
78
|
-
3. **Delegate.** Fire independent tasks in parallel in background. Stay available to steer.
|
|
79
|
-
4. **Verify.** Delegate verification too: read-only pass, test run, or `verifier` agent. The host reads the report and the diff, not the whole tree.
|
|
80
|
-
5. **Judge.** Output misses the bar → refine the spec and re-delegate (better prompts beat manual fixes). Escalate terra → Kimi K3 → sol → host-native only when the cheaper tier keeps failing.
|
|
81
|
-
|
|
82
|
-
## Delegation prompt checklist
|
|
83
|
-
|
|
84
|
-
- Names exact files/paths and the repo root
|
|
85
|
-
- States the goal in one sentence, then constraints (style, deletion safety: `trash` not `rm -rf`, no scope creep)
|
|
86
|
-
- Defines done: commands to run, expected results
|
|
87
|
-
- Asks for a report: changed files, what was done, tests run, risks
|
|
88
|
-
|
|
89
|
-
## Anti-patterns
|
|
90
|
-
|
|
91
|
-
- "It's faster if I just do it": beyond a trivial tweak it isn't, and it burns the budget the whole session depends on.
|
|
92
|
-
- Reading 10 files to plan: delegate exploration, plan from the summary.
|
|
93
|
-
- Fixing an executor's output by hand: refine the prompt and rerun.
|
|
94
|
-
- Serializing independent delegations: parallelize.
|
|
95
|
-
- Escalating everything to sol or the host: terra and Kimi K3 handle most well-spec'd work.
|
|
96
|
-
|
|
97
|
-
If no delegation path works (CLIs unavailable, Bash denied), say so explicitly and ask the user before falling back to direct execution: never silently drop out of the mode.
|
|
@@ -1,10 +0,0 @@
|
|
|
1
|
-
interface:
|
|
2
|
-
display_name: "Use Delegate"
|
|
3
|
-
short_description: "Delegate heavy work to OpenCode Kimi K3 and Codex"
|
|
4
|
-
icon_small: "./assets/codex-icon.svg"
|
|
5
|
-
icon_large: "./assets/codex-icon.svg"
|
|
6
|
-
brand_color: "#802F83"
|
|
7
|
-
default_prompt: "Use $use-delegate to orchestrate this task through cheap executors."
|
|
8
|
-
|
|
9
|
-
policy:
|
|
10
|
-
allow_implicit_invocation: false
|
|
@@ -1,20 +0,0 @@
|
|
|
1
|
-
<!-- @license lucide-static v1.24.0 - ISC -->
|
|
2
|
-
<svg role="img" aria-label="use-delegate skill icon"
|
|
3
|
-
class="lucide lucide-bot"
|
|
4
|
-
xmlns="http://www.w3.org/2000/svg"
|
|
5
|
-
width="128"
|
|
6
|
-
height="128"
|
|
7
|
-
viewBox="0 0 24 24"
|
|
8
|
-
fill="none"
|
|
9
|
-
stroke="#F5F5F5"
|
|
10
|
-
stroke-width="2"
|
|
11
|
-
stroke-linecap="round"
|
|
12
|
-
stroke-linejoin="round"
|
|
13
|
-
>
|
|
14
|
-
<path d="M12 8V4H8" />
|
|
15
|
-
<rect width="16" height="12" x="4" y="8" rx="2" />
|
|
16
|
-
<path d="M2 14h2" />
|
|
17
|
-
<path d="M20 14h2" />
|
|
18
|
-
<path d="M15 13v2" />
|
|
19
|
-
<path d="M9 13v2" />
|
|
20
|
-
</svg>
|
|
@@ -1,32 +0,0 @@
|
|
|
1
|
-
# Delegation models: current rankings and pricing
|
|
2
|
-
|
|
3
|
-
Last verified: 2026-07-20
|
|
4
|
-
|
|
5
|
-
**Refresh protocol**: if the date above is older than 14 days, refresh BEFORE planning a big delegation batch. Delegate the research (exa-search skill or a read-only executor): pull the DeepSWE leaderboard (https://deepswe.datacurve.ai/) and the provider pricing pages, then update both tables and the date. DeepSWE is the reference signal: 113 original long-horizon engineering tasks, contamination-free, cost-per-task published per model.
|
|
6
|
-
|
|
7
|
-
## DeepSWE leaderboard (best config per model, snapshot 2026-07-17)
|
|
8
|
-
|
|
9
|
-
| Model | Pass@1 | Avg cost/task | Read |
|
|
10
|
-
|---|---|---|---|
|
|
11
|
-
| gpt-5.6-sol [max] | 73% | $8.39 | Top score, best cost among frontier |
|
|
12
|
-
| claude-fable-5 [max] | 70% | $21.63 | Host tier: 2.6x sol cost, never a delegation target |
|
|
13
|
-
| gpt-5.6-terra [max] | 70% | $4.95 | Sol-level score at 59% of the cost |
|
|
14
|
-
| kimi-k3 [max] | 69% | $4.65 | Within noise of terra/sol, cheapest of the top pack |
|
|
15
|
-
| gpt-5.6-luna [max] | 67% | $3.03 | Acceptable floor for trivial bulk work |
|
|
16
|
-
| gpt-5.5 [xhigh] | 67% | $7.23 | Superseded by the 5.6 family |
|
|
17
|
-
| claude-opus-4.8 [max] | 59% | $13.22 | Dominated: lower score, higher cost |
|
|
18
|
-
|
|
19
|
-
## API list pricing (per 1M tokens)
|
|
20
|
-
|
|
21
|
-
| Model | Input | Cached input | Output | Context |
|
|
22
|
-
|---|---|---|---|---|
|
|
23
|
-
| Kimi K3 (Moonshot) | $3.00 | $0.30 | $15.00 | 1M |
|
|
24
|
-
| GPT-5.6 Sol | $5.00 | $0.50 | $30.00 | 1.05M |
|
|
25
|
-
| GPT-5.6 Terra | $2.50 | $0.25 | $15.00 | 1.05M |
|
|
26
|
-
| GPT-5.6 Luna | $1.00 | ~$0.10 | $6.00 | 1.05M |
|
|
27
|
-
|
|
28
|
-
## Access notes (this machine)
|
|
29
|
-
|
|
30
|
-
- Kimi K3: flat-rate through the kimi-for-coding subscription in opencode (`-m kimi-for-coding/k3`), so marginal cost per delegation is ~zero. The opencode-go gateway is unfunded (insufficient balance): do not route through it. Open weights due 2026-07-27; 1M context; native vision; Terminal-Bench 88.3.
|
|
31
|
-
- GPT-5.6 sol/terra/luna: covered by the Codex subscription (`codex exec -m gpt-5.6-<tier>`); config default is sol + effort high.
|
|
32
|
-
- Current call (2026-07-20): Kimi K3 is the default executor (top-pack score, subscription-covered). Terra for low-stakes tasks. Sol for compute-heavy work. Luna only for trivial bulk edits.
|
|
@@ -1,121 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: use-goal
|
|
3
|
-
description: Use when the user asks to create, draft, set, start, or refine a Codex or Claude Code `/goal` objective for persistent multi-turn work.
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# Use Goal
|
|
7
|
-
|
|
8
|
-
Create or draft a Codex or Claude Code Goal that follows the official `/goal` contract: one persistent objective with evidence-based completion criteria.
|
|
9
|
-
|
|
10
|
-
## When To Use
|
|
11
|
-
|
|
12
|
-
Use this skill when the user explicitly asks to:
|
|
13
|
-
|
|
14
|
-
- create, set, start, or use a Goal
|
|
15
|
-
- turn a task into a strong `/goal`
|
|
16
|
-
- make Codex or Claude Code continue until an outcome is actually done
|
|
17
|
-
- define success criteria for longer debugging, optimization, migration, refactor, benchmark, flaky-test, or research work
|
|
18
|
-
|
|
19
|
-
Do not introduce a Goal for a one-off edit, short explanation, simple code review, or single answer unless the user explicitly asks for Goal mode.
|
|
20
|
-
Do not use a Goal for a loose backlog or unrelated task list. A good Goal is bigger than one prompt but smaller than an open-ended project.
|
|
21
|
-
|
|
22
|
-
## Pick The Platform
|
|
23
|
-
|
|
24
|
-
Before drafting or creating a Goal, identify the active platform from the runtime and available tools:
|
|
25
|
-
|
|
26
|
-
- **Codex**: use `references/codex-goal.md`.
|
|
27
|
-
- **Claude Code**: use `references/claude-code-goal.md`.
|
|
28
|
-
- **Unknown platform**: draft a plain `/goal ...` command and state that the user should run it in the target agent.
|
|
29
|
-
|
|
30
|
-
If Goal tools are available in Codex, use them rather than only printing a slash command. If the runtime only exposes slash commands, return the exact `/goal ...` command unless the harness can dispatch it directly.
|
|
31
|
-
|
|
32
|
-
## Goal Shape
|
|
33
|
-
|
|
34
|
-
Before writing or creating the Goal, think through the verification strategy. Do a short discovery pass when the evidence surface is not already obvious:
|
|
35
|
-
|
|
36
|
-
- Inspect repository docs, package scripts, test commands, CI config, benchmark scripts, failing logs, linked issue text, plans, or referenced files.
|
|
37
|
-
- Identify which command, artifact, report, screenshot, benchmark, source document, or manual check can prove completion.
|
|
38
|
-
- For external libraries, APIs, or current product behavior, use the appropriate docs/research skill before relying on memory.
|
|
39
|
-
- Prefer existing project commands and documented workflows over inventing new validation.
|
|
40
|
-
- If no reliable verification surface exists, ask one concise question or make the Goal explicitly require creating one.
|
|
41
|
-
|
|
42
|
-
For broad refactors, deletions, migrations, moving files, or "remove all X" goals, read `references/verification-harnesses.md` before creating the Goal. Default to a measurable harness: establish a baseline count/list first, then make the Goal continue until the validation command exits successfully at the target condition, such as count `0`.
|
|
43
|
-
|
|
44
|
-
Write the Goal as a compact, well-formatted contract with these fields embedded in natural language:
|
|
45
|
-
|
|
46
|
-
1. Outcome: what must be true when the work is done.
|
|
47
|
-
2. Verification surface: the tests, commands, benchmarks, artifacts, reports, logs, source material, or other concrete evidence that proves it.
|
|
48
|
-
3. Constraints: what must not regress or be violated.
|
|
49
|
-
4. Boundaries: allowed files, tools, repositories, data, and resources when relevant.
|
|
50
|
-
5. Iteration policy: how to choose the next best action after each attempt.
|
|
51
|
-
6. Blocked stop condition: when to stop and what to report if no defensible path remains.
|
|
52
|
-
|
|
53
|
-
For long-running implementation work, also include:
|
|
54
|
-
|
|
55
|
-
- one objective and one stopping condition
|
|
56
|
-
- the files, docs, issue, logs, or plan the agent should inspect first
|
|
57
|
-
- the commands or artifacts that prove progress
|
|
58
|
-
- checkpoint behavior and a short progress log requirement
|
|
59
|
-
|
|
60
|
-
Prefer this pattern:
|
|
61
|
-
|
|
62
|
-
```text
|
|
63
|
-
<desired end state>, verified by <specific evidence>, while preserving <constraints>. Use <allowed inputs, tools, or boundaries>. Between iterations, <how to choose and record the next best action>. If blocked or no valid paths remain, stop with <attempted paths, evidence gathered, blocker, and next input needed>.
|
|
64
|
-
```
|
|
65
|
-
|
|
66
|
-
For implementation Goals, include exact command names when known:
|
|
67
|
-
|
|
68
|
-
```text
|
|
69
|
-
<desired end state>, verified by `<test or build command>` and <artifact/manual check>, while preserving <constraints>. First inspect <files/docs/logs>. Work in checkpoints: after each change, run the narrowest relevant verification, record the result, and choose the next smallest defensible step. Stop only when the verification passes, or stop blocked with the failed command output, attempted paths, and the missing input needed.
|
|
70
|
-
```
|
|
71
|
-
|
|
72
|
-
Keep the objective non-empty and at most 4,000 characters. If the needed instructions are longer, create or point to a file and make the Goal refer to that file.
|
|
73
|
-
|
|
74
|
-
## Create Or Draft
|
|
75
|
-
|
|
76
|
-
When goal tools are available, use this order:
|
|
77
|
-
|
|
78
|
-
1. Call the status tool first to check whether a Goal already exists.
|
|
79
|
-
2. If the user explicitly asked to create, set, start, or use a new Goal and no Goal exists, call the create tool with the refined objective.
|
|
80
|
-
3. Set a token budget only when the user explicitly provided one.
|
|
81
|
-
4. If a Goal already exists, do not overwrite, clear, pause, or resume it unless the user explicitly asks for that lifecycle action.
|
|
82
|
-
|
|
83
|
-
For Codex, this means calling `get_goal` first, then `create_goal` with the refined objective when creation is requested and no active Goal blocks it. Read `references/codex-goal.md` before doing so.
|
|
84
|
-
|
|
85
|
-
For Claude Code, `/goal <condition>` is a real slash command, but the model cannot launch it unless the harness exposes a slash-command dispatch tool. If no dispatch tool is available, return the exact manual `/goal ...` command, ask the user to paste/run it, and wait for confirmation before continuing Goal-driven work. Do not say Claude Code lacks `/goal`, do not call it Codex-only, and do not substitute a task list or goal file as equivalent unless the user explicitly asks for that fallback. Read `references/claude-code-goal.md` before drafting or instructing a Claude Code Goal.
|
|
86
|
-
|
|
87
|
-
If the user asks only to draft, rewrite, explain, or refine a Goal, return the final `/goal ...` text instead of activating it.
|
|
88
|
-
|
|
89
|
-
Ask a clarifying question only when a missing detail would make the Goal unverifiable or unsafe. Prefer one concise question. Otherwise infer conservative defaults from the repository, task, and available evidence.
|
|
90
|
-
|
|
91
|
-
## Evidence Rules
|
|
92
|
-
|
|
93
|
-
Completion must be evidence-based. Do not mark a Goal complete because the work seems likely done. First compare the objective to concrete evidence such as changed files, command output, tests, benchmarks, generated artifacts, logs, or source-backed research findings.
|
|
94
|
-
|
|
95
|
-
If a budget limit is reached, stop substantive work, summarize progress and blockers, and identify the next useful step. Do not treat budget exhaustion as completion.
|
|
96
|
-
|
|
97
|
-
If blocked, report the attempted paths, evidence gathered, blocker, and exact input or external change that would unlock progress.
|
|
98
|
-
|
|
99
|
-
If status reports become vague, tighten the Goal instead of adding more one-off instructions. Name the current checkpoint, what was verified, what remains, and what should cause a pause.
|
|
100
|
-
|
|
101
|
-
Only mark a Goal complete after verifying the stated stopping condition. Only mark it blocked when the same blocking condition has repeated enough that no meaningful progress is possible without user input or an external change.
|
|
102
|
-
|
|
103
|
-
## References
|
|
104
|
-
|
|
105
|
-
- `references/codex-goal.md`: Use for Codex Goal mode, `get_goal` / `create_goal`, CLI/app `/goal`, feature setup, and completion/blocking rules.
|
|
106
|
-
- `references/claude-code-goal.md`: Use for Claude Code manual `/goal`, evaluator behavior, requirements, status, clear/resume behavior, and non-interactive usage.
|
|
107
|
-
- `references/verification-harnesses.md`: Use for measurable refactor, deletion, migration, move, rename, dependency-removal, and "remove all X" Goals.
|
|
108
|
-
|
|
109
|
-
## Good Examples
|
|
110
|
-
|
|
111
|
-
```text
|
|
112
|
-
/goal Reduce p95 checkout latency below 120 ms, verified by the checkout benchmark, while keeping the correctness suite green. Use only the checkout service, benchmark fixtures, and related tests. Between iterations, record what changed, what the benchmark showed, and the next best experiment to try. If the benchmark cannot run or no valid paths remain, stop with the attempted paths, the evidence gathered, the blocker, and the next input needed.
|
|
113
|
-
```
|
|
114
|
-
|
|
115
|
-
```text
|
|
116
|
-
/goal Make the checkout test suite pass on the current branch, verified by the repository's documented test command, while preserving public API behavior. Use the failing tests, adjacent implementation files, and existing test helpers. Between iterations, inspect the latest failure, make the smallest defensible change, and rerun the relevant test surface. If no valid path remains, stop with the failures, changes tried, and the missing decision or dependency.
|
|
117
|
-
```
|
|
118
|
-
|
|
119
|
-
```text
|
|
120
|
-
/goal Produce the strongest evidence-backed reproduction report for the provided paper using available materials and local resources. Attempt the headline claims where feasible, verify outputs where possible, and end with a report that separates confirmed findings, approximate reconstructions, blocked claims, and remaining uncertainty.
|
|
121
|
-
```
|
|
@@ -1,7 +0,0 @@
|
|
|
1
|
-
interface:
|
|
2
|
-
display_name: "Use Goal"
|
|
3
|
-
short_description: "Use when the user asks to create, draft, set, start, or..."
|
|
4
|
-
icon_small: "./assets/codex-icon.svg"
|
|
5
|
-
icon_large: "./assets/codex-icon.svg"
|
|
6
|
-
brand_color: "#C70A64"
|
|
7
|
-
default_prompt: "Use $use-goal to help with this task."
|
|
@@ -1,18 +0,0 @@
|
|
|
1
|
-
<!-- @license lucide-static v1.24.0 - ISC -->
|
|
2
|
-
<svg role="img" aria-label="use-goal skill icon"
|
|
3
|
-
class="lucide lucide-sparkles"
|
|
4
|
-
xmlns="http://www.w3.org/2000/svg"
|
|
5
|
-
width="128"
|
|
6
|
-
height="128"
|
|
7
|
-
viewBox="0 0 24 24"
|
|
8
|
-
fill="none"
|
|
9
|
-
stroke="#F5F5F5"
|
|
10
|
-
stroke-width="2"
|
|
11
|
-
stroke-linecap="round"
|
|
12
|
-
stroke-linejoin="round"
|
|
13
|
-
>
|
|
14
|
-
<path d="M11.017 2.814a1 1 0 0 1 1.966 0l1.051 5.558a2 2 0 0 0 1.594 1.594l5.558 1.051a1 1 0 0 1 0 1.966l-5.558 1.051a2 2 0 0 0-1.594 1.594l-1.051 5.558a1 1 0 0 1-1.966 0l-1.051-5.558a2 2 0 0 0-1.594-1.594l-5.558-1.051a1 1 0 0 1 0-1.966l5.558-1.051a2 2 0 0 0 1.594-1.594z" />
|
|
15
|
-
<path d="M20 2v4" />
|
|
16
|
-
<path d="M22 4h-4" />
|
|
17
|
-
<circle cx="4" cy="20" r="2" />
|
|
18
|
-
</svg>
|
|
@@ -1,65 +0,0 @@
|
|
|
1
|
-
# Claude Code Goal Reference
|
|
2
|
-
|
|
3
|
-
Use this reference when the active agent is Claude Code.
|
|
4
|
-
|
|
5
|
-
Official reference:
|
|
6
|
-
|
|
7
|
-
- https://code.claude.com/docs/en/goal
|
|
8
|
-
|
|
9
|
-
## Manual Activation
|
|
10
|
-
|
|
11
|
-
Claude Code has `/goal`, but the model cannot launch it unless the harness exposes a slash-command dispatch tool. If no dispatch tool is available, output the exact `/goal ...` command, ask the user to paste/run it manually, and wait for confirmation before continuing Goal-driven work.
|
|
12
|
-
|
|
13
|
-
Do not say Claude Code lacks `/goal`. Do not call `/goal` Codex-only. Do not replace `/goal` with a task list, TODO list, or goal file and describe it as equivalent. Those can be supporting artifacts only when the user asks for them or when they are useful after the manual `/goal` command has been provided.
|
|
14
|
-
|
|
15
|
-
If the user asks to continue without manually pasting the command, continue normal work only after acknowledging that no Claude Code Goal is active.
|
|
16
|
-
|
|
17
|
-
## Command Surface
|
|
18
|
-
|
|
19
|
-
Claude Code uses `/goal` to set a completion condition for the current session.
|
|
20
|
-
|
|
21
|
-
- `/goal <condition>` sets the Goal and immediately starts a turn using the condition as the directive.
|
|
22
|
-
- `/goal` with no argument shows the current state, evaluated turns, token spend, and latest evaluator reason.
|
|
23
|
-
- `/goal clear` removes the active Goal before it is met.
|
|
24
|
-
- `stop`, `off`, `reset`, `none`, and `cancel` are aliases for `clear`.
|
|
25
|
-
- `/clear` starts a new conversation and removes any active Goal.
|
|
26
|
-
- `claude -p "/goal <condition>"` can run a Goal non-interactively until the condition is met or the process is interrupted.
|
|
27
|
-
|
|
28
|
-
Only one Goal can be active per Claude Code session. Setting a new `/goal <condition>` replaces the active Goal. Do not overwrite an existing Goal unless the user explicitly asks to replace it.
|
|
29
|
-
|
|
30
|
-
If an active Goal existed when a Claude Code session ended, it is restored on `--resume` or `--continue`; the condition carries over, but the timer, turn count, and token-spend baseline reset.
|
|
31
|
-
|
|
32
|
-
## Requirements
|
|
33
|
-
|
|
34
|
-
Claude Code `/goal` requires Claude Code v2.1.139 or later.
|
|
35
|
-
|
|
36
|
-
It only runs in trusted workspaces because it uses the hooks system. It is unavailable when hooks are disabled through `disableAllHooks` or restricted through `allowManagedHooksOnly`; Claude Code should explain that condition when the command fails.
|
|
37
|
-
|
|
38
|
-
## Evaluator Behavior
|
|
39
|
-
|
|
40
|
-
Claude Code evaluates the Goal after each turn with a separate small fast model. A "no" result starts another turn and passes the evaluator reason as guidance. A "yes" result clears the Goal and records the achieved condition in the transcript.
|
|
41
|
-
|
|
42
|
-
The evaluator does not call tools, read files, or run commands independently. It judges only the condition and what Claude has surfaced in the conversation so far. Therefore the Goal must require Claude to put the proof in the transcript.
|
|
43
|
-
|
|
44
|
-
Good Claude Code Goal conditions include:
|
|
45
|
-
|
|
46
|
-
- one measurable end state, such as a passing test, clean build, target count, or empty queue
|
|
47
|
-
- a stated check, such as a command exiting `0`, a generated report, or a reviewed artifact
|
|
48
|
-
- constraints that matter, such as files not to modify or behavior not to regress
|
|
49
|
-
- a bounded stop clause when useful, such as "or stop after 20 turns with the remaining blocker"
|
|
50
|
-
|
|
51
|
-
Goal conditions can be up to 4,000 characters. If the instructions are longer, put details in a file and make the Goal point to that file.
|
|
52
|
-
|
|
53
|
-
## Claude Code Draft Pattern
|
|
54
|
-
|
|
55
|
-
Prefer this pattern:
|
|
56
|
-
|
|
57
|
-
```text
|
|
58
|
-
/goal <desired end state>, verified by <proof Claude must surface in the transcript>, while preserving <constraints>. First inspect <files/docs/logs>. After each turn, report the current checkpoint, command/artifact result, remaining gap, and next smallest step. Stop when the proof is in the transcript, or after <bound> with attempted paths, evidence, blocker, and needed input.
|
|
59
|
-
```
|
|
60
|
-
|
|
61
|
-
For test or build work:
|
|
62
|
-
|
|
63
|
-
```text
|
|
64
|
-
/goal <desired end state>, verified by `<test/build command>` exiting 0 with the relevant output included in the transcript, while preserving <constraints>. First inspect <files/logs>. After each turn, rerun the narrowest relevant check and summarize the result. Stop only when the command output proves success, or stop after <bound> with the failing output, attempted fixes, and missing input.
|
|
65
|
-
```
|
|
@@ -1,70 +0,0 @@
|
|
|
1
|
-
# Codex Goal Reference
|
|
2
|
-
|
|
3
|
-
Use this reference when the active agent is OpenAI Codex, including the Codex app, IDE extension, or CLI.
|
|
4
|
-
|
|
5
|
-
Official references:
|
|
6
|
-
|
|
7
|
-
- https://developers.openai.com/codex/use-cases/follow-goals
|
|
8
|
-
- https://developers.openai.com/codex/app/commands
|
|
9
|
-
- https://developers.openai.com/codex/cli/slash-commands
|
|
10
|
-
|
|
11
|
-
## Command Surface
|
|
12
|
-
|
|
13
|
-
`/goal <objective>` starts Goal mode. `/goal` views the current Goal. `/goal pause`, `/goal resume`, and `/goal clear` manage lifecycle.
|
|
14
|
-
|
|
15
|
-
Goal objectives must be non-empty and at most 4,000 characters. For longer instructions, create or point to a file and make the Goal refer to that file.
|
|
16
|
-
|
|
17
|
-
If `/goal` is missing, tell the user to enable Goals with:
|
|
18
|
-
|
|
19
|
-
```toml
|
|
20
|
-
[features]
|
|
21
|
-
goals = true
|
|
22
|
-
```
|
|
23
|
-
|
|
24
|
-
They can also run:
|
|
25
|
-
|
|
26
|
-
```bash
|
|
27
|
-
codex features enable goals
|
|
28
|
-
```
|
|
29
|
-
|
|
30
|
-
## Tool Contract
|
|
31
|
-
|
|
32
|
-
When Codex Goal tools are available, use the tools instead of printing a slash command for activation:
|
|
33
|
-
|
|
34
|
-
1. Call `get_goal` before any lifecycle action.
|
|
35
|
-
2. If the user asked to create, set, start, activate, or use a new Goal and no active Goal exists, call `create_goal` with the refined objective.
|
|
36
|
-
3. Pass `token_budget` only when the user explicitly provided a budget.
|
|
37
|
-
4. If a Goal already exists, do not overwrite, clear, pause, resume, mark complete, or mark blocked unless the user explicitly asked for that lifecycle action or the active Goal's stated status condition is actually met.
|
|
38
|
-
|
|
39
|
-
Use slash-command text only when the user asks for a draft, the tool surface is unavailable, or the target is a separate Codex session.
|
|
40
|
-
|
|
41
|
-
## Codex Goal Shape
|
|
42
|
-
|
|
43
|
-
A strong Codex Goal should define:
|
|
44
|
-
|
|
45
|
-
- one objective and one stopping condition
|
|
46
|
-
- the files, docs, issue, logs, or plan Codex should inspect first
|
|
47
|
-
- the commands, artifacts, screenshots, benchmarks, reports, or manual checks that prove progress
|
|
48
|
-
- constraints that must not regress
|
|
49
|
-
- checkpoint behavior and compact progress logging
|
|
50
|
-
- the exact blocked stop condition and what evidence to report
|
|
51
|
-
|
|
52
|
-
Prefer this pattern:
|
|
53
|
-
|
|
54
|
-
```text
|
|
55
|
-
<desired end state>, verified by <specific evidence>, while preserving <constraints>. Use <allowed inputs, tools, or boundaries>. Between iterations, <how to choose and record the next best action>. If blocked or no valid paths remain, stop with <attempted paths, evidence gathered, blocker, and next input needed>.
|
|
56
|
-
```
|
|
57
|
-
|
|
58
|
-
For implementation Goals, include exact commands when known:
|
|
59
|
-
|
|
60
|
-
```text
|
|
61
|
-
<desired end state>, verified by `<test or build command>` and <artifact/manual check>, while preserving <constraints>. First inspect <files/docs/logs>. Work in checkpoints: after each change, run the narrowest relevant verification, record the result, and choose the next smallest defensible step. Stop only when the verification passes, or stop blocked with the failed command output, attempted paths, and the missing input needed.
|
|
62
|
-
```
|
|
63
|
-
|
|
64
|
-
## Completion And Blocking
|
|
65
|
-
|
|
66
|
-
Completion must be evidence-based. Compare the active Goal to concrete evidence in the thread: changed files, command output, tests, benchmarks, generated artifacts, logs, screenshots, or source-backed research findings.
|
|
67
|
-
|
|
68
|
-
Do not mark a Goal complete because the work seems likely done, because a budget is exhausted, or because no more work is planned. Only mark it complete after verifying the stated stopping condition.
|
|
69
|
-
|
|
70
|
-
Only mark a Goal blocked when the stated blocker has repeated enough that no meaningful progress is possible without user input or an external change. Report the attempted paths, gathered evidence, exact blocker, and input needed.
|
|
@@ -1,108 +0,0 @@
|
|
|
1
|
-
# Verification Harnesses For Refactors
|
|
2
|
-
|
|
3
|
-
Use this reference when a Goal involves refactoring, deletion, migration, moving files, eliminating a pattern, or reducing a code smell. The core tactic is to create a small measurable harness before the main work, then make the Goal continue until the harness reaches the target.
|
|
4
|
-
|
|
5
|
-
## Principle
|
|
6
|
-
|
|
7
|
-
For broad changes, do not rely only on subjective review. Convert the desired end state into a number, list, or deterministic command result.
|
|
8
|
-
|
|
9
|
-
Examples:
|
|
10
|
-
|
|
11
|
-
- Remove all explicit TypeScript `any` -> count explicit `any` occurrences and require `0`.
|
|
12
|
-
- Delete dead files -> scan imports/references and require no references to removed paths.
|
|
13
|
-
- Move a module -> scan imports and require all imports use the new path.
|
|
14
|
-
- Rename an API -> count old symbol references and require `0`, then run tests/typecheck.
|
|
15
|
-
- Remove a dependency -> scan package manifests and lockfiles, then run install/typecheck/test.
|
|
16
|
-
- Split a large file -> check file size or exported symbol boundaries, then run typecheck/test.
|
|
17
|
-
|
|
18
|
-
The harness should make progress visible after every iteration.
|
|
19
|
-
|
|
20
|
-
## Harness Rules
|
|
21
|
-
|
|
22
|
-
1. First establish the baseline count or failure list before editing.
|
|
23
|
-
2. Prefer existing repo tooling: tests, lint rules, typecheck, dependency analyzers, codemods, static analyzers.
|
|
24
|
-
3. If no existing command measures the target, create a narrow validation script.
|
|
25
|
-
4. Make the script deterministic and fast enough to run repeatedly.
|
|
26
|
-
5. Exclude generated, vendored, build output, lockfiles, snapshots, and irrelevant binary assets unless the task explicitly includes them.
|
|
27
|
-
6. Print actionable output: total count, grouped files, and the top remaining offenders.
|
|
28
|
-
7. Exit with code `0` only when the target condition is met. Exit non-zero while work remains.
|
|
29
|
-
8. Keep the harness scoped to the Goal. Remove temporary harnesses before completion unless they are useful project validation and the user or repo conventions support keeping them.
|
|
30
|
-
|
|
31
|
-
## Goal Pattern
|
|
32
|
-
|
|
33
|
-
Use this shape for count-based refactor Goals:
|
|
34
|
-
|
|
35
|
-
```text
|
|
36
|
-
<desired refactor>, verified by `<validation command>` returning success with <target count/list condition>, while preserving <tests/typecheck/public behavior>. First establish the baseline with `<validation command>` and inspect the highest-signal offenders. Work in checkpoints: after each batch, rerun `<validation command>`, record the count/list delta, run the narrowest relevant tests, and continue until the validation command exits 0. If the target cannot be reached safely, stop with the remaining offenders, attempted paths, failing output, and the decision needed.
|
|
37
|
-
```
|
|
38
|
-
|
|
39
|
-
Example:
|
|
40
|
-
|
|
41
|
-
```text
|
|
42
|
-
Remove all explicit TypeScript `any` from the codebase, verified by `node scripts/check-explicit-any.mjs` exiting 0 with count 0, while keeping `pnpm typecheck` and relevant tests green. First run the checker to record the baseline and prioritize files with the most occurrences. Work in checkpoints: after each batch, rerun the checker, record the remaining count, and run the narrowest relevant typecheck/tests. Continue until the checker exits 0. If some `any` cannot be removed safely, stop with the remaining locations, attempted replacements, compiler/test output, and the type information needed.
|
|
43
|
-
```
|
|
44
|
-
|
|
45
|
-
## Script Patterns
|
|
46
|
-
|
|
47
|
-
Use structured parsers when reasonable. For TypeScript, prefer the TypeScript compiler API, `ts-morph`, ESLint, or an existing lint rule over raw text search when false positives matter.
|
|
48
|
-
|
|
49
|
-
For a quick bootstrap harness, a text scanner is acceptable if the Goal explicitly treats it as an approximate first pass and follows up with typecheck/lint.
|
|
50
|
-
|
|
51
|
-
Example quick scanner:
|
|
52
|
-
|
|
53
|
-
```js
|
|
54
|
-
#!/usr/bin/env node
|
|
55
|
-
import { readdirSync, readFileSync, statSync } from "node:fs";
|
|
56
|
-
import { join } from "node:path";
|
|
57
|
-
|
|
58
|
-
const root = process.cwd();
|
|
59
|
-
const ignoredDirs = new Set([".git", "node_modules", "dist", "build", ".next", "coverage"]);
|
|
60
|
-
const extensions = new Set([".ts", ".tsx"]);
|
|
61
|
-
const matches = [];
|
|
62
|
-
|
|
63
|
-
function walk(dir) {
|
|
64
|
-
for (const entry of readdirSync(dir)) {
|
|
65
|
-
const path = join(dir, entry);
|
|
66
|
-
const stat = statSync(path);
|
|
67
|
-
if (stat.isDirectory()) {
|
|
68
|
-
if (!ignoredDirs.has(entry)) walk(path);
|
|
69
|
-
continue;
|
|
70
|
-
}
|
|
71
|
-
if (![...extensions].some((ext) => path.endsWith(ext))) continue;
|
|
72
|
-
const text = readFileSync(path, "utf8");
|
|
73
|
-
const lines = text.split("\n");
|
|
74
|
-
lines.forEach((line, index) => {
|
|
75
|
-
if (/\bany\b/.test(line)) matches.push(`${path.replace(`${root}/`, "")}:${index + 1}: ${line.trim()}`);
|
|
76
|
-
});
|
|
77
|
-
}
|
|
78
|
-
}
|
|
79
|
-
|
|
80
|
-
walk(root);
|
|
81
|
-
|
|
82
|
-
console.log(`explicit_any_count=${matches.length}`);
|
|
83
|
-
for (const match of matches.slice(0, 50)) console.log(match);
|
|
84
|
-
if (matches.length > 50) console.log(`...and ${matches.length - 50} more`);
|
|
85
|
-
process.exit(matches.length === 0 ? 0 : 1);
|
|
86
|
-
```
|
|
87
|
-
|
|
88
|
-
## Refactor Targets
|
|
89
|
-
|
|
90
|
-
For deletion Goals, verify both absence and behavior:
|
|
91
|
-
|
|
92
|
-
- target files or symbols are gone
|
|
93
|
-
- no imports, string references, routes, config entries, docs links, or tests point to them
|
|
94
|
-
- typecheck/build/test still passes
|
|
95
|
-
|
|
96
|
-
For moving Goals, verify all call sites:
|
|
97
|
-
|
|
98
|
-
- old import path count is `0`
|
|
99
|
-
- new import path exists where expected
|
|
100
|
-
- public exports remain compatible unless changing them is part of the Goal
|
|
101
|
-
- typecheck/build/test still passes
|
|
102
|
-
|
|
103
|
-
For migration Goals, verify old surface removal and new surface behavior:
|
|
104
|
-
|
|
105
|
-
- old package/API/pattern count reaches `0` or the explicitly allowed exception list
|
|
106
|
-
- new package/API/pattern is used consistently
|
|
107
|
-
- tests/typecheck/build pass
|
|
108
|
-
- manual or browser verification is included when behavior is visual or interactive
|