@akinet/akidevrule 3.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +835 -0
- package/LICENSE +21 -0
- package/README.md +356 -0
- package/claude/CLAUDE.md +40 -0
- package/claude/agents/aki-challenger.md +38 -0
- package/claude/agents/aki-conduct.md +54 -0
- package/claude/agents/aki-hands.md +59 -0
- package/claude/agents/aki-judge.md +37 -0
- package/claude/agents/aki-maker.md +36 -0
- package/claude/fragments/settings.akidoc.fragment.json +15 -0
- package/claude/hooks/aki-update-check.mjs +160 -0
- package/claude/hooks/aki_version_check.mjs +83 -0
- package/docs/ref/macos-codesign-tcc.md +59 -0
- package/install.mjs +1067 -0
- package/install.ps1 +11 -0
- package/install.sh +12 -0
- package/package.json +52 -0
- package/payload/GEMINI.md +147 -0
- package/payload/METHOD-audit-flow.md +147 -0
- package/payload/METHOD-audit-subtraction.md +67 -0
- package/payload/METHOD-audit-zero-trust.md +49 -0
- package/payload/METHOD-deep-think.md +172 -0
- package/payload/METHOD-proportionality.md +62 -0
- package/payload/METHOD-ux-psych.md +60 -0
- package/payload/RULE-agent-behavior.md +138 -0
- package/payload/RULE-biz.md +51 -0
- package/payload/RULE-coding.md +130 -0
- package/payload/RULE-content-write.md +54 -0
- package/payload/RULE-db-design.md +26 -0
- package/payload/RULE-docs.md +144 -0
- package/payload/RULE-pattern-core.md +80 -0
- package/payload/RULE-release.md +215 -0
- package/payload/RULE-seo.md +173 -0
- package/payload/RULE-stack-akiNuxtCf.md +179 -0
- package/payload/RULE-stack-tauri.md +59 -0
- package/payload/RULE-ui-pattern.md +167 -0
- package/payload/index.md +91 -0
- package/skills/aki-article-writer/SKILL.md +50 -0
- package/skills/aki-article-writer/references/article-workflow.md +377 -0
- package/skills/akidevsync-notes/SKILL.md +48 -0
- package/skills/akidevsync-notes/scripts/notes_cli.py +212 -0
- package/skills/akiflow/SKILL.md +221 -0
- package/skills/akiflow/references/harness-facts.md +215 -0
- package/skills/akiflow/scripts/council-cost.sh +4 -0
- package/skills/akiflow/scripts/council-open.sh +4 -0
- package/skills/akiflow/scripts/council-read.sh +4 -0
- package/skills/akiflow/scripts/council-verify.sh +4 -0
- package/skills/akiflow/scripts/council_cost.py +149 -0
- package/skills/akiflow/scripts/council_open.py +323 -0
- package/skills/akiflow/scripts/council_read.py +148 -0
- package/skills/akiflow/scripts/council_verify.py +315 -0
- package/skills/akiflow/scripts/scythe.py +307 -0
- package/skills/akiflow/scripts/scythe.sh +4 -0
- package/skills/akigitcommit/SKILL.md +85 -0
- package/skills/akihelp/SKILL.md +47 -0
- package/skills/akihtmlreport/SKILL.md +59 -0
- package/skills/akilint/SKILL.md +29 -0
- package/skills/akirule/SKILL.md +155 -0
- package/skills/akiship/SKILL.md +55 -0
- package/skills/akithink/SKILL.md +59 -0
|
@@ -0,0 +1,221 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: akiflow
|
|
3
|
+
description: Lead-coordinated agent council for work that needs more than one kind of judgment. The lead pins the owner's verbatim words as the run's immutable anchor, decomposes the request into owned work items each quoting a requirement from it, checks a three-condition activation gate, and convenes seats from the five installed agent definitions in ~/.claude/agents/ — one batch, each seat traced to a requirement, never picked from a menu. Every mechanism is off by default and turns on only when this run produces a reason. The lead does no menial work and settles what doctrine answers, escalating only a one-way door, a contradiction with documented design, or scope expansion — then writes the owner's answer back into doctrine so the same question never escalates twice. Two shapes discriminated by whether anything is being arbitrated — a council of items with adversaries, or a dispatch of lanes with exclusive file ownership for fan-out work whose answer is already knowable; and three modes discriminated by what changes outside the room — discuss, audit (read-only), execute. A mechanical gate script refuses closure on a missing anchor, a REQ that quotes nothing the owner wrote, a declared seat that left neither a turn nor a file, a seat with no rule receipt, an untagged seat, an unanswered reminder, or a requirement no item covers. Explicit invoke only.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# akiflow — lead-coordinated agent council
|
|
7
|
+
|
|
8
|
+
Invoke with `/akiflow <request>`. The session agent becomes the **lead**: it anchors, decomposes, convenes, arbitrates, and decides.
|
|
9
|
+
|
|
10
|
+
## The lead's job is two laws
|
|
11
|
+
|
|
12
|
+
Everything below is one of these two made operational. They are the lead's job description, not corpus-wide law.
|
|
13
|
+
|
|
14
|
+
| Law | Statement | What it forbids |
|
|
15
|
+
|---|---|---|
|
|
16
|
+
| **R1 — ANCHOR** | The owner's words are immutable and are the final test: the content, the mechanism they named, and the shape of the answer. | Paraphrasing the request into a pinned problem statement. The moment the lead restates, the anchor is gone and every seat downstream inherits the restatement — the failure mode that cost two full council runs. |
|
|
17
|
+
| **R2 — JUSTIFICATION** | Every mechanism — seat, check, step, script, consult — is **OFF by default** and turns on only when *this run* produces a reason. | "It is in the skill", "it is a standing seat", "we always run it". Being documented is not a reason. A gate that forces a seat to exist manufactures work; a roster derived from a tier instead of from a requirement is over-staffing with a procedure attached. |
|
|
18
|
+
|
|
19
|
+
The council is worth its cost only against three structural failures of a single thread, none of which a stronger model fixes: **context flooding** (one context holds request, code, plan, diff and review), **role collapse** (one agent applying one standard of "correct" to problems judged by different ones), and **self-approval** (the context that produced a decision cannot judge it). With none of the three present, this skill costs more than it returns.
|
|
20
|
+
|
|
21
|
+
## Step 0 — anchor, then cut
|
|
22
|
+
|
|
23
|
+
```bash
|
|
24
|
+
python3 ~/.claude/skills/akiflow/scripts/council_open.py <slug> "<the owner's message, verbatim>"
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
It refuses to open a room without the message, and writes it as chat.md's immutable `## anchor` block. Pinning used to be a discipline; two consecutive runs skipped it, so it is now a mechanism whose absence is impossible.
|
|
28
|
+
|
|
29
|
+
**The requirement ledger.** Extraction is bulk work, so it goes to `aki-hands`, not the lead: one numbered line per distinct requirement, `REQ-1 … REQ-n`, **each carrying a "quoted fragment" copied from the anchor**. The quote is what makes a line a requirement rather than an interpretation, and the closure gate checks it. The lead ratifies the draft against the anchor before cutting anything — only a ratified ledger counts, and the lead owns coverage even though a worker did the labor.
|
|
30
|
+
|
|
31
|
+
A **work item** is the atomic unit:
|
|
32
|
+
|
|
33
|
+
```
|
|
34
|
+
ITEM <id> · <one-line statement of what must be decided or built>
|
|
35
|
+
covers: <REQ-n, REQ-m, ...>
|
|
36
|
+
owner: <seat name>
|
|
37
|
+
challenger: <a different seat name>
|
|
38
|
+
closes when:<a criterion someone else can check>
|
|
39
|
+
rationale: <filled in at closure, <=3 lines>
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
Every REQ must be covered by ≥1 item; an orphan REQ is a decomposition bug, not a footnote. Cut along real boundaries — each item having its own definition of "correct" — never into phases of one undivided question. The checklist is a precondition for opening the room, never a product of it.
|
|
43
|
+
|
|
44
|
+
**`aki-challenger`'s first assignment is always the decomposition itself**, before any content attack: the REQ with no item, the missing item, the item with two owners, the item whose closing criterion nobody can check. A bad cut means the room debates the wrong squares thoroughly — the one failure no other mechanism catches, because every other mechanism operates *within* the item structure.
|
|
45
|
+
|
|
46
|
+
## Step 1 — the gate, and the mode
|
|
47
|
+
|
|
48
|
+
Declare both before any other output:
|
|
49
|
+
|
|
50
|
+
```
|
|
51
|
+
[akiflow] mode=execute · REQ 1-6 → 4 items · trigger: schema + API shape + migration ordering
|
|
52
|
+
roster: judge-schema(mid) · challenger(mid) · hands-callsites(cheap, ro:--tools Read,Grep,Glob) · maker-api(mid)
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
**Shape comes before the gate: is there anything to arbitrate?** If two competent seats could reach different defensible answers, this is a **council** and the rest of this step applies. If the answer is knowable and the work is merely large, it is a **dispatch** (Step 1b) — the same anchor, ledger, receipts and closure gate, partitioned into lanes with exclusive file ownership instead of items with an adversary. Shape and mode are independent: a dispatch can be `audit` or `execute` just as a council can.
|
|
56
|
+
|
|
57
|
+
All three conditions must hold, or this is not a council: **decomposable** into ≥2 items with real boundaries · **≥2 kinds of "correct"** (schema-correct ≠ UX-correct ≠ price-correct; if one standard covers everything, one good head suffices) · **cost of error exceeds cost of coordination**. Ambiguity resolves downward. Never ask the user which tier they want — the gate is auditable through these two lines.
|
|
58
|
+
|
|
59
|
+
**Mode is decided by one question: what changes outside the room?**
|
|
60
|
+
|
|
61
|
+
| Mode | Changes outside the room | Produces | Notes |
|
|
62
|
+
|---|---|---|---|
|
|
63
|
+
| `discuss` | nothing | a decision plus its record | no `aki-maker` is convened; a room that writes files is not in this mode |
|
|
64
|
+
| `audit` | nothing — read-only by construction (`agent.B5`) | findings, and a plan that schedules fixes | one item per domain, each owned by a `judge` seated on that domain's standard: `docs.C` · `ui.C` · `flow` · `release.B` · `ux.C` · `biz` · `subtract`. Fixes are a separate run through this gate |
|
|
65
|
+
| `execute` | files | a diff, verified | `aki-maker` is the only seat permitted to write |
|
|
66
|
+
|
|
67
|
+
**Bulk mechanical work is not a council — but it no longer has to leave the skill.** The same transform across many files, or a sweep whose paths are known up front, has nothing for a roster to arbitrate and grows the lead's context with the item count. It still wants the anchor, the REQ ledger, the `[RULES]` receipts, the durable record and the closure gate, and that combination is a dispatch (Step 1b), not a reason to fall back to bare spawns. Claude Code's native `Workflow` tool is still the better fit where the loop itself must be held outside any model's context — the owner must invoke it, this skill cannot. A subtraction audit is the clearest split: the scanning is dispatch lanes, and the council convenes only at classification, where *dead* versus *load-bearing but ugly* is the judgment the owner acts on.
|
|
68
|
+
|
|
69
|
+
## Step 1b — dispatch
|
|
70
|
+
|
|
71
|
+
A fan-out with a paper trail. Reuses this skill's workspace, its three file kinds and its closure gate unchanged; replaces the debate with a partition. Declare it the same way:
|
|
72
|
+
|
|
73
|
+
```
|
|
74
|
+
[akiflow] shape=dispatch · mode=execute · REQ 1-4 → 3 lanes
|
|
75
|
+
lanes: scripts(maker mid, writes skills/akiflow/scripts/*.py) · docs(maker mid, writes docs/** + CHANGELOG.md) · sweep(hands cheap, ro:--tools Read,Grep,Glob, writes none)
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
**A lane is an item whose adversary is replaced by an exclusive file set.** Every lane carries `covers` (the REQs it satisfies), `worker`, `writes`, `reads` and `returns`. `writes` is exclusive: a path claimed by two lanes is a gate failure, refused by `council_open.py --convene` before a token is spent, because two workers editing one file is the failure a fan-out actually dies of and it is invisible until the second one clobbers the first. `returns` exists because the lead merges the lanes and cannot merge a shape that was never specified.
|
|
79
|
+
|
|
80
|
+
**Every lane leaves its own trace, and the lead does not hold them all.** The trace name for a lane IS the lane's own short name, never its `worker:` value — a lane whose worker can write puts its report in `<lane>.md` in the session directory and returns only its conclusion; a read-only lane returns to the lead, who posts it as one turn under the lane name. Either satisfies `council_verify.py`, and the first keeps a long report out of the lead's context until something makes it worth reading. `worker:` stays a required field — it is roster/cost metadata, the same as a council item's `model` — but it is never the trace identity: two lanes may declare the same `worker` type and must still leave two separately-traceable reports.
|
|
81
|
+
|
|
82
|
+
**What dispatch drops, and why that is safe:** no challenger and no judge, because nothing is being arbitrated · no turn-numbered debate and no peer-to-peer laws, because lanes do not talk to each other · no three-condition gate, because the condition that justifies a dispatch is different — work partitionable into ≥2 independent lanes, each lane's paths and question nameable up front, and a result that must outlive the session. Fail the third and a bare spawn is enough; fail the second and it was a council question all along.
|
|
83
|
+
|
|
84
|
+
**What dispatch keeps is the whole point:** the anchor immutable (R1 is unconditional), every REQ quoting the owner's own words, every worker's `[RULES]` receipt, the durable on-disk record, and all seven checks of `council_verify.py` — a lane's `worker` is a declared seat exactly as a council's `owner` is, and an untraced one is a ghost.
|
|
85
|
+
|
|
86
|
+
**A lane declared `worker: agy` carries the failure-report clause in its prompt, not just in the roster line** — `references/harness-facts.md` § Worker invocation quick-facts has the verbatim text. Gemini's helpful-bias fabricates a result exactly where a tool call was silently denied; the clause is what turns that into an honest `BLOCKED:` instead.
|
|
87
|
+
|
|
88
|
+
## Step 2 — convene from the definitions, never from a menu
|
|
89
|
+
|
|
90
|
+
The five agents live in `~/.claude/agents/` and carry their own tools, model tier, rule manifest and output contract. Read the definition rather than re-describing it here.
|
|
91
|
+
|
|
92
|
+
| Definition | Seat is for |
|
|
93
|
+
|---|---|
|
|
94
|
+
| `aki-hands` | retrieval with `file:line`; judgment forbidden. Also the file that names every worker substrate (Claude subagent · agy · kiro-cli · `cl-9rt`) and the recorded harness facts that make re-probing them unnecessary |
|
|
95
|
+
| `aki-judge` | one standard, named at spawn — `pattern`, `proportion`, `ux`, `db`, `docs`, `release`, `biz`, whichever the item is judged by |
|
|
96
|
+
| `aki-conduct` | the process: whether rules arrived (LOAD-fail) and whether they were followed (COMPLY-fail); `scythe.py` is its tool |
|
|
97
|
+
| `aki-challenger` | attacks the result from a clean context; closes on *"what can be cut?"* and *"does this answer the anchored words?"* |
|
|
98
|
+
| `aki-maker` | turns a decision into a diff; `execute` mode only |
|
|
99
|
+
|
|
100
|
+
**The convening rule: a seat exists only when it traces to a requirement in the anchor.** Five definitions on disk is a catalog, not a roster, and picking seats from a catalog is exactly the over-staffing this skill was rebuilt to end. There are no standing seats — `conduct` is convened when the run writes durable artifacts, `judge -proportion` when an item adds, sizes, keeps or removes a guard, limit or accepted risk, and so on. A seat with nothing to act on is R2 violated with a procedure attached.
|
|
101
|
+
|
|
102
|
+
**Name each seat `<definition>-<scope>`** — `judge-schema`, `hands-callsites`, `challenger` — because the name is the address `SendMessage` routes to and the key `chat.md` turns are grouped by. **Convene the whole roster in one call**: each subagent's sibling list is captured at its own startup, so an agent named later is invisible to those named earlier — a silent one-way channel with no error. Assign each a distinct turn-number block at the same time (`judge-schema` 10–19, `challenger` 20–29), since parallel writers cannot see each other's latest number and would collide on `#1`. Check `disallowedTools` does not strip `SendMessage`, or the roster is decorative.
|
|
103
|
+
|
|
104
|
+
**Before the spawn batch, the checklist must pass its own gate** — the room is not convened on an uncut question:
|
|
105
|
+
|
|
106
|
+
```bash
|
|
107
|
+
python3 ~/.claude/skills/akiflow/scripts/council_open.py --convene <session-dir>
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
Exit 1 unless ≥1 `ITEM` carries all of `owner` / `challenger` / `closes when`. It gates *convening*, not file creation: the anchor has to be pinned before the ledger can quote it (R1), so `chat.md` necessarily exists first — the cost this prevents is N agents circling an undecomposed question, and that cost is paid at spawn.
|
|
111
|
+
|
|
112
|
+
**A seat is declared by tier — `top` / `mid` / `cheap` — never by a model name; the host running this skill resolves the tier through its own row in `references/harness-facts.md` § Model tiers › Host resolution.** On Claude Code the tier lives in each agent definition's frontmatter — that file is the Claude row, and a spawn passes `model` only to override it, not as a ritual on every call. On any other host (Cursor reads the same `~/.claude/agents/` files) the frontmatter value is a Claude alias the host cannot resolve, so the lead declares each seat's model explicitly at spawn from the host's row — a Claude alias written into a Cursor or Codex spawn buys an expensive API-pool model where a native cheap tier existed. The real remaining hazard is a generic subagent such as `general-purpose`, which carries no tier of its own and therefore inherits the lead's expensive default — this matters concretely at Step 6, where the cost seat must hold `Bash` and is spawned generically. The in-session Agent tool has **no `effort` parameter**; only headless calls take `--effort`, so do not declare per-seat effort for in-session spawns. A read-only seat names its enforcing mechanism (`--tools`, `--mode plan`, `--trust-tools=`) on the roster line; read-only by wording is not read-only.
|
|
113
|
+
|
|
114
|
+
**The lead does no menial work — ever.** Arbitration quality is its only product and it degrades with every unrelated token. Bulk reads, greps, inventory scans go to `aki-hands` even when doing it directly feels faster, because "faster" spends the one context the run cannot replace. The lead reads at orientation depth: the anchor, the checklist, `--stats`, and the specific excerpt a decision turns on.
|
|
115
|
+
|
|
116
|
+
## Step 3 — the room
|
|
117
|
+
|
|
118
|
+
`council_open.py` creates `~/.aki/agent-council/<project>/<YYYY.MM.DD-HHMM>-<slug>/` and prunes sessions older than 30 days. Three files, three jobs, never merged:
|
|
119
|
+
|
|
120
|
+
| File | Writer | Holds |
|
|
121
|
+
|---|---|---|
|
|
122
|
+
| `<seat-name>.md` | that seat | its pinned mandate and working notes |
|
|
123
|
+
| `chat.md` | everyone | the anchor, the pinned block, and the meeting in time order |
|
|
124
|
+
| `checklist.md` | **the lead only** | the REQ ledger, the items, their closures and rationale; the durable copy goes to `docs/plan/` (`docs.B1`) |
|
|
125
|
+
|
|
126
|
+
Turns are `### <HH:MM> <seat-name> #<n>`, under 200 words, never hard-wrapped, everyone appends and nobody edits another's turn. **Every claim in a turn is tagged `FACT` / `CONSTRAINT` / `ASSUMPTION`** (`METHOD-deep-think.md` B2) — the gate checks each posting agent used at least one, because an untagged room is where a guess closes an item wearing a fact's clothes.
|
|
127
|
+
|
|
128
|
+
```bash
|
|
129
|
+
R="python3 ~/.claude/skills/akiflow/scripts/council_read.py"
|
|
130
|
+
$R <chat.md> --index | --pinned | --stats | --agent challenger --tail 5 | --from 12
|
|
131
|
+
$R <chat.md> --grep "quota|pricing" --agent red-team # locate: matching lines, tagged with the turn each came from
|
|
132
|
+
$R <chat.md> --turn 14 # then read only that turn
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
**Locate, then read — never scan.** `--grep` answers "did anyone raise X?" for a few hundred bytes; reading the room to answer the same question costs the whole file. That is the only pair of commands that removes a whole-file read, because every other flag (`--agent`, `--from`, `--tail`) needs you to already know where to look. A seat arriving reads `--pinned` plus `--index`, then only the turns its own mandate names — it has no `--from` to resume from, and pulling the whole file is exactly the context the spawn existed to avoid. A seat rejoining reads `--pinned` plus `--from <its last turn>`. The lead watches `--stats` and `--index` and reads full turns only where something looks wrong; a lead that reads the whole room has rebuilt the flooded main thread with the least independent context now holding the arbitration seat.
|
|
136
|
+
|
|
137
|
+
**A read is a subscription, not a purchase.** Every tool call re-sends the whole history, so anything pulled into context is charged again on every later turn: a read of size `S` at turn `t` of a `T`-turn run costs about `S × (T − t)`, not `S`. Reading a 50k-token room at turn 50 of 200 costs ~7.5M cache-read tokens — a fifth of a real measured run's entire lead spend, from one call. `--stats` prints bytes per agent precisely so the lead can price a read before making it.
|
|
138
|
+
|
|
139
|
+
**Peer-to-peer laws.** Direct challenge is the point of the room, but it removes the lead's view of how a conclusion was reached: every peer exchange ends in a `DECISION:` or `CONFLICT:` turn posted by whoever closes it · peer agreement is not a decision, only the lead closes an item · three rounds per pair then escalate, and cyclic chains (A→B→C→A) are forbidden · every message costs a full turn of the receiving agent.
|
|
140
|
+
|
|
141
|
+
**Domain consults are mandatory once a seat exists.** An item whose closure touches a domain with a seated judge closes only after that judge's recorded turn. "Nobody asked UX" is a closure defect. This does not create seats — it binds the ones the anchor already justified.
|
|
142
|
+
|
|
143
|
+
**Recurring conflict is a design smell, not a refereeing job.** The same ground contested across two or more items is the signature of a missing pattern underneath (`pattern.A8`). Open a root item, name the pattern, re-close the conflicted items against it.
|
|
144
|
+
|
|
145
|
+
**Steering is judgment, not a counter.** Depth is why this skill exists; never cut a productive argument short for being long. Intervene on a real signal — ground re-covered with no new evidence, a closing criterion that has stopped getting closer, scope drifting outside mandates, cost visibly outrunning the decision's worth — and then do the minimum: one pinned `CHECKPOINT` line, messaged only to the seats that are drifting.
|
|
146
|
+
|
|
147
|
+
## Step 4 — closure, and what reaches the owner
|
|
148
|
+
|
|
149
|
+
Before items close, and again before the Step 6 tally, the lead runs the gate and pastes its output into the room:
|
|
150
|
+
|
|
151
|
+
```bash
|
|
152
|
+
python3 ~/.claude/skills/akiflow/scripts/council_verify.py <session-dir>
|
|
153
|
+
```
|
|
154
|
+
|
|
155
|
+
It fails on a missing anchor, a REQ quoting nothing the owner wrote, a declared owner/challenger that never posted, a posting agent with no `[RULES]` receipt, a posting agent that never tagged evidence, an unanswered `REMIND-<n>`, and a ledger `REQ-<n>` no item's `covers` names. That last one is the only check aimed at the lead itself: `aki-challenger` sees the items the lead cut and therefore cannot see a requirement that never became one, so the gate diffs the ratified ledger against `covers` rather than asking the lead to declare its own omissions — a declaration the omitting party is the worst-placed to make (`agent.B2`). A FAIL is a closure blocker, not a note. It proves presence, never quality — and it deliberately does not require any named seat, because a gate that manufactures a seat gets gamed rather than questioned.
|
|
156
|
+
|
|
157
|
+
**A reminder from `conduct` blocks what it targets**: the target answers `ACK REMIND-<n>` plus the fix applied, or the lead posts `OVERRULE REMIND-<n> <reason>` — a logged lead judgment, never a default and never the target's own call.
|
|
158
|
+
|
|
159
|
+
**The lead decides and reports.** Exactly three things escalate: a genuine one-way door (`think.A1`) · anything contradicting `docs/biz/` or documented project design · scope expansion beyond what was asked. A fourth is possible but rare — the room deadlocked on something important *and* the lead cannot break the tie on the merits; present it as a decision with positions, tradeoff and a recommendation, never as an open question handed back. A seat-raised `CONFLICT` is a candidate, not an escalation: it reaches the owner only if it survives `agent.A3`'s kill-tests.
|
|
160
|
+
|
|
161
|
+
**Doctrine first, always.** Before anything reaches the owner the lead verifies the question is not already answered by `docs/biz/`, the project `CLAUDE.md`, or the relevant `docs/arch|feat` — and the escalation cites that search: which files were read and where exactly they fall silent. **Every owner answer becomes doctrine** in the same turn it is applied, so the identical question can never escalate again. An answer left in chat evaporates with the session; asking twice is failing twice.
|
|
162
|
+
|
|
163
|
+
**The closure rationale answers two one-liners**: *did this achieve what the owner actually asked for* (compared against the anchor, not against the lead's restatement of it), and *why is this the smallest shape that does*. An item that cannot name what was cut — or state that nothing needed cutting — has not faced the subtraction pass. Verification asks "did I do what I said"; adversarial review asks "should this have been done"; neither asks the first question, which is why it is written explicitly here.
|
|
164
|
+
|
|
165
|
+
## Step 5 — `execute` mode
|
|
166
|
+
|
|
167
|
+
| Job | Mechanism |
|
|
168
|
+
|---|---|
|
|
169
|
+
| Implement from the plan | `aki-maker`, given the plan doc path and the checklist items in its prompt — the plan doc *is* the continuity mechanism, since no subagent inherits the lead's context |
|
|
170
|
+
| Mechanical fan-out | `aki-hands` on the cheapest capable tier |
|
|
171
|
+
| **Verify** — "did I do what I said?" plus the drift sweep | a subagent given the diff and the closing criteria. The drift sweep is part of verifying, not an extra: every doc, comment, i18n string and CHANGELOG line referencing the changed behavior, reported if it still describes the old one |
|
|
172
|
+
| **Adversarial review** — "should this have been done?" | `aki-challenger`, at the tier its frontmatter declares. Contamination is disqualifying — that is what makes the seat work, not the tier |
|
|
173
|
+
| **Author user-facing prose** | a separate writer worker on the softest capable writing tier, never the implementer as a side effect, under an anti-fabrication brief: it phrases facts supplied in its prompt and tags anything else ASSUMPTION |
|
|
174
|
+
|
|
175
|
+
**The one boundary that must never blur.** A reviewer briefed with the lead's self-justifying chain will agree with it — sycophancy given structure, worse than no review because it emits a stamp. `aki-challenger` receives **only the diff, the closing criteria, and the anchor**. What is withheld is the lead's *reasoning*, never the behavior floor: rules are constraints, not justification, so they cannot contaminate a review.
|
|
176
|
+
|
|
177
|
+
**Parallel writers need isolation** (`isolation: "worktree"`) — it costs setup time and disk per agent, so a lone implementer or a read-only sweep does not get one. **Do not stall the roster on its slowest member**: wait for a whole stage only when the next stage genuinely needs every result together. **A truncated scope must be stated** — capping at top-N or sampling is allowed, doing it silently is not. **Unknown-size discovery loops until it runs dry** (two consecutive rounds surfacing nothing new), never until a fixed count. **Adversarial verifiers get distinct lenses**, not the same question repeated.
|
|
178
|
+
|
|
179
|
+
**The roster stays convened**, so a blocker goes back into the room rather than onto the owner's desk. A Phase B blocker that invalidates an assumption behind a closed item **reopens that item** — message its owner and challenger, re-close with a new rationale. It is never patched quietly; that is how a plan doc becomes fiction while everyone still cites it.
|
|
180
|
+
|
|
181
|
+
Even with no room at all, close direct work with a verifier subagent against what was promised, running the same drift sweep. It catches what dominates small tasks: a missed call site, an unupdated CHANGELOG, a "tested" claim that was never run.
|
|
182
|
+
|
|
183
|
+
## Step 6 — close-out accounting
|
|
184
|
+
|
|
185
|
+
The roster line declared `model` before a token was spent; this closes that loop with what was actually spent. Mandatory, not an extra the owner requests: a run that cost ten times its worth must not be indistinguishable from one that didn't. `council_verify.py` must already have passed.
|
|
186
|
+
|
|
187
|
+
```bash
|
|
188
|
+
python3 ~/.claude/skills/akiflow/scripts/council_cost.py <session-dir> # session id read from the room's chat.md stamp
|
|
189
|
+
python3 ~/.claude/skills/akiflow/scripts/council_cost.py --session <uuid> # for a room opened before the stamp existed
|
|
190
|
+
```
|
|
191
|
+
|
|
192
|
+
**One `haiku` subagent runs it and reports the table — never the lead**; reading the raw transcript is the largest bulk-read in the run. It must be a seat that actually holds `Bash` — `aki-conduct`, or a generic subagent — since the retrieval seat the lead reaches for by reflex (`aki-hands`) is `Read, Grep, Glob` only and will return a refusal rather than a table. The script aggregates in-shell and prints tokens only, because per-model prices drift and a hardcoded table in a distributed script would rot: look the price up, bill `input + cache_creation` as input, price `cache_read` and `output` separately. **It measures the Claude meter completely and stops there** — the main session plus every subagent under it, which is exactly the budget Step 1's roster declared. A headless lane (`agy`, or a separate `claude -p`) writes no turn into this transcript and bills another quota: report its own `usage` beside the table if it matters, never summed into it, or the total is in no single currency. The close-out line goes into the run's `docs/plan/` record, beside the roster declaration it reconciles.
|
|
193
|
+
|
|
194
|
+
## Anti-patterns — all of these are default behaviour unless forbidden by name
|
|
195
|
+
|
|
196
|
+
1. **A pinned problem statement that paraphrases the owner** → every seat inherits the paraphrase, and the room answers a question nobody asked. R1, and the most expensive failure this skill has actually produced.
|
|
197
|
+
2. **Convening a seat because it exists** → a seat with no surface to act on burns a full agent's budget on artifacts nobody will read again. R2.
|
|
198
|
+
3. **Opening the room before the checklist exists** → agents circling an uncut question at many times the cost of solving it alone.
|
|
199
|
+
4. **Briefing the adversarial reviewer with the lead's reasoning chain** → a rubber stamp wearing a review's clothes.
|
|
200
|
+
5. **Spawning the roster across several turns** → one-way channels; agents deaf without knowing it.
|
|
201
|
+
6. **Agreement with no falsifier** → manufactured consensus. Three prior agents agreeing is not evidence.
|
|
202
|
+
7. **Handing the owner an open question instead of a decision, or escalating past doctrine** → the council did not do its job, and the owner pays attention twice for one question.
|
|
203
|
+
8. **Nesting a subagent for context-dependent work** → the grandchild invents and the parent cannot tell. One level deep, mechanical only: if the task cannot be written in under 200 words with no project context, it is not a nested-spawn task.
|
|
204
|
+
9. **Merging `chat.md` into `checklist.md`** → the argument buries the conclusion, and execution inherits conclusions stripped of their reasons.
|
|
205
|
+
10. **Re-probing a CLI whose flags are already recorded** → one run spent three calls re-learning what `references/harness-facts.md` already stated. Liveness and quota are probeable; capability is not.
|
|
206
|
+
11. **Lead writing `chat.md` content on behalf of a seat** → the lead ran the council scripts correctly but then authored the turn content that only a real subagent should produce. `council_verify.py` catches this as `ghost-seats`, but the budget is already spent. The lead opens the room, decomposes, spawns — it never *is* a seat. Same severity as #1: a council whose seats are the lead in costume answers nothing the lead alone could not, at many times the cost.
|
|
207
|
+
|
|
208
|
+
## Harness notes
|
|
209
|
+
|
|
210
|
+
**Before choosing sequential or parallel execution, the lead checks its own tool set.** If `invoke_subagent` (or `Agent` on Claude Code) is available, the roster is spawned as real concurrent subagents — the fallback sequential path does not apply. If no spawn mechanism exists (headless `claude -p`, a read-only subagent), the sequential single-session path applies. This check is mandatory; defaulting to sequential when a spawn tool is present is anti-pattern #11.
|
|
211
|
+
|
|
212
|
+
- **Claude Code:** roster in one batch; `SendMessage` for peer challenge and for resuming completed seats; continuity travels as the plan doc or diff named in the prompt, since the one `subagent_type` that inherits session history (`fork`) is gated off by default; `isolation: "worktree"` for concurrent writers. An agent the *user* stopped refuses to resume via message and must be resumed from its own transcript panel — do not respawn a duplicate.
|
|
213
|
+
- **Headless (`claude -p`):** nobody can answer an escalation or a permission prompt. Record it as `BLOCKED: needs owner` in `checklist.md` and continue the other items — never guess what the owner would have wanted.
|
|
214
|
+
- **Antigravity / AGY:** supports native subagents via `invoke_subagent`. When `/akiflow` is invoked with multiple experts or dispatch lanes, the lead MUST spawn the roster via `invoke_subagent` concurrently in one batch. Simulating multiple seats sequentially in a single session context without spawning real subagents is strictly forbidden (role collapse / self-approval violation). Where AGY is reachable from a Claude Code lead, it may also serve as a wide-context worker substrate (`aki-hands`).
|
|
215
|
+
- **Script paths in this skill's literal commands are written for Claude Code** (`~/.claude/skills/akiflow/scripts/...`). This file is deployed byte-identical to `~/.gemini/config/skills/` too (`docs/ref/agent-skills-standard.md`), so either root's path runs the same script under Antigravity/agy. **Run the command exactly as written above** — the installer pre-allows both roots in both renderings (expanded and tilde-literal), so the form you copy is never what gets denied. Background: agy's matcher compares command strings literally, with no glob or tilde expansion, so a rule and a command that render the same path differently do not match — which is why the pre-allow covers every rendering instead of asking you to normalize one (`docs/ref/cli-permission-allowlist-standard.md` §1.2).
|
|
216
|
+
|
|
217
|
+
Verified harness facts behind every flag named here: `references/harness-facts.md` — its § Worker invocation quick-facts is the lookup table (literal command, read-only mechanism, silent failure per lane); the rest of the file is why. Design record: `docs/arch/akiflow.md` in the akidevrule repo.
|
|
218
|
+
|
|
219
|
+
## Invocation scope
|
|
220
|
+
|
|
221
|
+
Explicit invoke only — akirule never auto-triggers this skill. When ordinary work makes the three activation conditions obvious, suggest `/akiflow` in a single line; do not self-invoke.
|
|
@@ -0,0 +1,215 @@
|
|
|
1
|
+
# akiflow — harness facts and cost model
|
|
2
|
+
|
|
3
|
+
The skill's rules are consequences of these facts. If a fact changes, the rule it supports must be revisited rather than patched. Each entry is marked:
|
|
4
|
+
|
|
5
|
+
- **[doc]** — stated in Anthropic's (or Antigravity's) published documentation, linked at the bottom.
|
|
6
|
+
- **[obs]** — observed runtime behaviour of the tooling (Claude Code or Antigravity/agy), not found in published docs. Treat as true-until-contradicted, and re-verify before relying on a detail.
|
|
7
|
+
- **[owner]** — supplied by the owner from a machine this repo has never run on. Not verified here and not verifiable here. Weakest tier: never let a rule depend on one without a verification step attached.
|
|
8
|
+
|
|
9
|
+
Every entry carries the date it was checked, because all of it is version-bound and expected to rot.
|
|
10
|
+
|
|
11
|
+
## Worker invocation quick-facts
|
|
12
|
+
|
|
13
|
+
The lookup table: literal command, read-only mechanism, and the one silent failure each lane hides. Every section below this one is the *why* — a caller assigning a lane needs none of it.
|
|
14
|
+
|
|
15
|
+
**Cross-file lock.** Sections in this file are addressed **by name** from `claude/agents/aki-hands.md` (§ Substrates) and `skills/akiflow/SKILL.md` (the closing pointer under § Harness notes). Renaming or removing a section means updating both in the same change; a stale section name reads as a missing fact and sends the reader back to probing.
|
|
16
|
+
|
|
17
|
+
| Lane | Literal command | Read-only by | Silent failure to check |
|
|
18
|
+
|---|---|---|---|
|
|
19
|
+
| **agy flash** — discovery default | `agy --model gemini-3.7-flash-high --mode plan --output-format json -p "<prompt>"` | `--mode plan` (mechanism, not wording) | a denied call still returns `status: "SUCCESS"` with empty `response`; `-p` takes the next token as its value, so any flag written after it is sent as prompt text |
|
|
20
|
+
| **kiro-cli** | `kiro-cli chat --no-interactive --trust-tools=fs_read --model claude-sonnet-4.5 --effort high "<prompt>"` | `--trust-tools=fs_read` | none recorded; `--effort` is operative on every Kiro model, unlike `claude` + haiku |
|
|
21
|
+
| **claude via proxy gateway** | `CLAUDE_CONFIG_DIR=~/.claude-9rt claude -p --tools "Read,Grep" --model <alias> --effort low "<prompt>"` | `--tools` allowlist | `cl-9rt` is a shell alias and does not exist in a non-interactive shell — run the expanded literal; the gateway may route an alias to a non-Anthropic core |
|
|
22
|
+
| **claude in-harness subagent** | Agent tool, `model` passed explicitly | the agent file's `tools:` frontmatter | an omitted `model` inherits the caller's top tier; the Agent tool has no `effort` parameter at all, so a declared effort is decorative |
|
|
23
|
+
| **claude persistent worker** | `claude -p --session-id <uuid> …`, then `--resume <uuid>` | `--tools` allowlist | the id is cwd-scoped — resuming from another directory fails with `No conversation found` (exit 5) |
|
|
24
|
+
|
|
25
|
+
Two cheapness dials, set independently: agy carries the thinking tier inside the model name (`-low`/`-medium`/`-high`) and has no separate dial; Claude-family lanes need `--model` **and** `--effort`, except haiku, which has no extended thinking and ignores the flag.
|
|
26
|
+
|
|
27
|
+
**Mandatory brief clause for any agy lane (L3, `docs/plan/done/agy-helpful-bias-containment.md`).** Gemini's helpful/shortcut bias is aligned-in, not a prompting accident (Google's own guidance: the model "may over-analyze verbose or overly complex prompt engineering," and the industry-default "if a tool fails, try a different approach" pattern trains the exact fabricate-past-failure reflex this bias produces). The agy flash row above already names the silent-`SUCCESS` failure mode; a silently denied tool is the vacuum the bias fills with an invented answer. Every prompt string sent to an agy lane — dispatch or council — carries this verbatim clause, not a paraphrase, and not a growth of `payload/GEMINI.md` itself (that file's size is now treated as a cost, per the containment plan's L4):
|
|
28
|
+
|
|
29
|
+
> If any tool call or permission fails: do NOT retry, do NOT work around it, do NOT substitute content. Output `BLOCKED:` followed by the verbatim error, then stop. Always prefer actual tool output over memory; never report a result no tool returned.
|
|
30
|
+
|
|
31
|
+
Verification scope, stated precisely: the clause was carried in three test prompts, but only **two of them actually hit a denial** (T2, T8) — T9 was an allow and evidences nothing about failure reporting. So the **failure-report sentence** (first sentence) is 2 for 2 on real denial events: both times, flash-tier agy emitted the `BLOCKED:` line with the verbatim error instead of a confident substitute. Two events is a signal, not a pass rate; no control run with default phrasing was recorded, so "better than the default" is inference from the documented industry pattern, not a measured A/B. Its wording here is normalized from the tested contract, and this block is the canonical form workers copy. The **tool-output-primacy sentence** (second) is an untested ToolFailBench-derived mitigation riding along, covered by nothing. A dispatch lane declared with `worker: agy` (`SKILL.md` Step 1b) is out of contract without this block in its prompt; the roster declaration alone does not carry it.
|
|
32
|
+
|
|
33
|
+
**Scope limit — in headless `-p`, a denied file write never reaches the model, so the clause cannot cover it (measured 2026-08-21, agy 1.1.17, Linux; every worker lane here is headless).** A denied `write_file` makes the CLI end the turn at the permission check: the caller gets a non-`SUCCESS` status and an empty `response`, and no model turn remains in which to emit `BLOCKED:`. The non-`SUCCESS` envelope varies and the reason is not always in the JSON — `ERROR` populates `error`, `CANCELED` leaves it empty and puts the verbatim reason on **stderr** (both seen on the same trap minutes apart, 2026-08-21), so **key on `status`, never on the `error` field being present**. The clause still earns its place for failures that surface *inside* a model turn (a failing shell command, a missing file, a tool returning an error the model then reacts to). **The caller's duty is unconditional: check `status != "SUCCESS"` OR a `SUCCESS` with an empty body** — those are two different gates with two different signatures (vendor doc: an approval-needed tool is *soft*-denied, run continues, exit 0, stderr notice; a boundary refusal is a hard error), and checking for only one misreads the other. Nothing here describes interactive or IDE agy: the vendor states there is no interactive prompt in headless mode, i.e. it is a different path, unmeasured. Evidence: `docs/research/handoff-vs-self-verification-aug21.md` §5.7 + [Antigravity headless docs](https://antigravity.google/docs/cli/headless).
|
|
34
|
+
|
|
35
|
+
Probe exactly two things, once, at the moment of assigning a lane: liveness/quota (a one-token `… -p "ok"`) and `test -d ~/.claude-9rt`. Capability is recorded — never re-run `--help` or `agy models`, unless a caller names a model absent from § agy headless.
|
|
36
|
+
|
|
37
|
+
## Subagents
|
|
38
|
+
|
|
39
|
+
| Fact | Design consequence |
|
|
40
|
+
|---|---|
|
|
41
|
+
| **[doc]** A subagent runs in its own context window with its own tool set and model; it does not see the parent's conversation. | Independence is available by construction. It also means a plain subagent inherits **no** akirule routing — every plain-subagent prompt must name the exact `~/.aki/akidevrule/*.md` files to Read. |
|
|
42
|
+
| **[obs]** A subagent that has a **name** and the `SendMessage` tool receives a *sibling roster* listing the other named agents, captured **at its own startup**. | Agents named later are invisible to agents named earlier — a silent one-way channel with no error. The roster is therefore convened in **one batch**, and mid-run escalation reconvenes rather than appends. |
|
|
43
|
+
| **[doc]** `/fork` and `/subtask` are **interactive slash commands**, not an Agent-tool `subagent_type`; both require `CLAUDE_CODE_FORK_SUBAGENT=1`. A real run on 2026-07-30 failed every `subagent_type: fork` spawn with `Agent type 'fork' not found`, matching official docs at the time (`code.claude.com/docs/en/agents`, `.../sub-agents`) — which is why a prior version of this row said flatly that no `subagent_type: fork` value existed. **[obs]** That claim is corrected for Claude Code 2.1.220: the `claude` binary contains a real fork agent type (telemetry field `is_fork`; error strings `"Fork is not available inside a forked worker"` and `"Fork cannot use isolation: \"remote\" — a remote session cannot inherit the conversation context"`), and the Agent tool's own `model` parameter documents *"Ignored for subagent_type: \"fork\" — forks always inherit the parent model."* It is gated behind `CLAUDE_CODE_FORK_SUBAGENT=1` and is **not** in the available-agent list of a default session, so it cannot be relied on. (Confirmed 2026-08-01 by inspecting the `claude` binary directly — undocumented on the public pages above.) | The design consequence is unchanged: continuity work still travels by an explicit plan-doc/diff handoff. The *reason* changes — not "no such mechanism exists," but "the mechanism exists, is gated off by default, and even where enabled is not the cross-session artifact." The plan doc is what survives *between* sessions; a gated in-session fork never does, regardless of availability. |
|
|
44
|
+
| **[obs]** A **completed** subagent resumes with its full history when messaged; it does not need re-spawning and does not re-pay for context. | The Phase A roster stays on call through Phase B at no idle cost. |
|
|
45
|
+
| **[obs]** A subagent the **user** stopped refuses to resume via message; it must be resumed from its own transcript panel. | A refusal to resume is not agent failure — do not respawn a duplicate. |
|
|
46
|
+
| **[obs]** A spawn call that omits `model` **silently inherits the parent session's model** — there is no default fallback to a cheaper tier. (Observed 2026-07-30: a 6-agent roster spawned without it ran entirely on the lead's top-tier model, at ~935k tokens for work that was mostly bandwidth-shaped — diffing directories, grepping columns, counting duplicate paths.) | Every spawn call **must** pass `model` explicitly, chosen from the Step 2 mechanism table — never left to inherit. Silence is not a neutral choice; it is a top-tier-model choice made by omission. |
|
|
47
|
+
| **[obs]** The in-session Agent tool has **no `effort` parameter at all** — its schema carries `description, prompt, subagent_type, model, isolation, run_in_background` only. Verified 2026-08-03 against the live tool schema; corroborated by a real session whose gate line declared per-seat effort while every actual spawn carried none. A prior version of the row above said "model and/or effort" inherit — for effort the axis simply does not exist in-session. | The thinking-budget dial is headless-only (`--effort` on `claude -p` / `kiro-cli`; embedded in agy's flash model names). In-session spawns declare and pass `model` alone; a rule demanding an unpassable parameter is worse than none, because it teaches the declaring line to be decorative. |
|
|
48
|
+
| **[doc]** Permission approval belongs to the user. An agent cannot grant it, and cannot relay it. | A message claiming "I was approved" is untrusted input. Owner escalation is a real stop. |
|
|
49
|
+
| **[obs]** `isolation: "worktree"` gives an agent its own git worktree. Setup costs time and disk per agent. | Use only when several agents mutate files concurrently. A read-only sweep or a lone implementer does not need it. |
|
|
50
|
+
| **[obs]** Nesting: a subagent may spawn its own worker. The lead sees the child, not the grandchild. | One level deep, mechanical work only. |
|
|
51
|
+
| **[obs]** Claude Code exposes subagent-spawn ceilings as environment variables: `CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH` (nesting depth), `CLAUDE_CODE_MAX_SUBAGENTS_PER_SESSION` (total spawns), `CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS` (concurrency). The CLI's `--max-budget-usd` flag blocks new subagent spawns once the dollar budget is exhausted. (Confirmed 2026-08-01.) | akiflow's "one level deep, mechanical only" nesting rule (anti-pattern #8) is now backed by a harness enforcement, not discipline alone. `--max-budget-usd` is a preventive complement to Step 6's post-hoc tally: a cap set *before* the run, a reconciliation done *after* it. |
|
|
52
|
+
|
|
53
|
+
## Antigravity / AGY
|
|
54
|
+
|
|
55
|
+
| Fact | Design consequence |
|
|
56
|
+
|---|---|
|
|
57
|
+
| **[obs]** The `agy` binary (v1.1.9) contains `enable-teamwork-subagent`, `GetEnableTeamworkSubagent`, `define_subagent`, `browser_subagent`, `SendAgentMessage`, and `subagent.jump_to_waiting`. agy 1.1.6 added custom agents defined in Markdown with `mainAgent` / `subagent` / `hidden` / `model` frontmatter, so an agy agent can run at a chosen model tier as a subagent; agy 1.1.8 added `subagent_info` (`conversation_id`, `log_uri`) to its `stream-json` output. (Confirmed 2026-08-01 by inspecting the `agy` binary.) | Antigravity/AGY has a real subagent mechanism. Any prior claim that it has none is wrong and must be corrected everywhere it appears in `SKILL.md`. |
|
|
58
|
+
| **[obs]** agy ships a built-in agent council, `/teamwork-preview`, with a **fixed** roster read from the binary: `orchestrator_pure`, `explorer`, `spec_miner`, `armed_worker`, `armed_critic`, `empirical_challenger`, `forensic_auditor`, `reviewer_critic`, `sentinel`, `test_writer`, `victory_auditor`, `challenger`. (Confirmed 2026-08-01.) | Two consequences. **Convergent validation**: an independently designed council landed on nearly the same role split, including a *pure* orchestrator that only orchestrates — akiflow's "the lead does no menial work" was arrived at independently, not merely fashionable. **A gap**: akiflow has no **victory audit** role — *"did this achieve the goal that was asked for?"*, distinct from verification's *"did I do what I said?"* and adversarial review's *"should this have been done?"* **What not to copy**: the roster is fixed; akiflow derives its roster from the items it cut, and a fixed roster is exactly what R2 ("a seat is convened only when it traces to a requirement in the anchor") forbids. |
|
|
59
|
+
|
|
60
|
+
| **[owner]** `gemini-3.1-pro` (agy `gemini-3.1-pro-{low,high}`) writes noticeably softer, more natural user-facing prose than Sonnet-class implementers (owner observation across projects, 2026-08-03); its agy quota headroom is unmeasured here. | Default engine for akiflow's writer role (`SKILL.md` Step 5): a role split from the implementer, under an anti-fabrication brief — the writer phrases facts supplied in its brief, never sources its own. Probe quota before a long batch; fallback is a Claude-family writer plus a softening pass. |
|
|
61
|
+
|
|
62
|
+
## Cross-CLI worker (Claude Code lead → agy headless)
|
|
63
|
+
|
|
64
|
+
Verified by real runs, 2026-08-01; model re-probed 2026-08-15. Invocation, flag order load-bearing:
|
|
65
|
+
|
|
66
|
+
```
|
|
67
|
+
agy --model gemini-3.7-flash-high --mode plan --output-format json -p "<prompt>"
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
| Fact | Design consequence |
|
|
71
|
+
|---|---|
|
|
72
|
+
| **[obs]** `-p`/`--print` takes the prompt as **its value**, so it must come last. `agy -p --model X "prompt"` silently sends the wrong thing and returns a confident, unrelated answer with no error. (Hit in testing.) | The flag-order trap is the single most likely way this mechanism fails silently — state it explicitly in every prompt/script that invokes agy headless. |
|
|
73
|
+
| **[obs]** `--mode plan` enforces read-only **by mechanism**, not by prompt wording. | Strictly stronger than Claude Code's own subagent read-only story: the Agent tool's `mode` parameter is deprecated and ignored, and a Claude subagent inherits the parent session's permission mode. For read-only sweeps and audit mode, agy is the better substrate. akiflow's "restate read-only in every audit prompt" rule stays necessary on the Claude side and becomes belt-and-braces on the agy side. |
|
|
74
|
+
| **[obs]** `~/.gemini/GEMINI.md` auto-loads into every agy call, including headless `-p` — 15 sections of the akirule behavior baseline (scope discipline, no unrequested artifacts, no model-credit trailers, absolute factuality, communication-is-read-only, skill-token dispatch). | An agy worker gets the behavior floor for free, where a Claude subagent must be handed the file list explicitly. This does not exempt the prompt from naming the item's domain rule files on top — GEMINI.md carries the baseline, not the domain-specific layer. |
|
|
75
|
+
| **[obs]** agy headless cannot prompt for permission; a denied action auto-fails and the call still returns `status: "SUCCESS"` with `response: ""`. `~/.gemini/antigravity-cli/settings.json` already allows the common read tools (grep, rg, find, cat, head, ls, sed, awk, wc), so read-only sweeps work today without `--dangerously-skip-permissions`. | A caller that does not test for an empty `response` reads a failed call as a clean sweep. Every cross-CLI call must check for this before trusting its result — see the new anti-pattern in `SKILL.md`. |
|
|
76
|
+
| **[obs]** agy has `allowNonWorkspaceAccess: true` and a global workspace index; in testing it resolved file paths back to the real repository even when run inside a copied directory. | `cwd` is not a reliable scope boundary for an agy worker — the prompt must name the paths explicitly. |
|
|
77
|
+
| **[obs]** Measured on a real read-only repo sweep: 8.2s wall / 3.4s model time, correct answer; ~20–26k tokens of fixed input overhead per call (agy's system prompt plus `~/.gemini/GEMINI.md`). *A prior version of this row cited one observed `cache_read_tokens: 32621` as evidence that repeat calls hit a warm cache. That reading was too generous — see § Stateful workers, where a controlled three-turn test shows the cache is unreliable and the latency curve is the real constraint.* | The fixed overhead means this mechanism pays for itself on a non-trivial sweep, not on a one-line lookup — the same shape as the "self-contained question" cutoff already in the Step 2 mechanism table. |
|
|
78
|
+
| **[obs]** The `json`/`stream-json` output carries `usage`: `input_tokens`, `output_tokens`, `thinking_tokens`, `cache_read_tokens`, plus `conversation_id`. | Cost is measurable per call — but it is invisible to `council_cost.py`, which only parses the Claude Code session transcript. The lead must add it to the close-out tally by hand (`SKILL.md` Step 6). |
|
|
79
|
+
| **[obs]** agy 1.1.9 expands skills in print mode, so `agy -p "/akiflow …"` resolves the skill; `akiflow` is already deployed to agy at `~/.gemini/config/skills/akiflow`. | A cross-CLI call can invoke the skill itself, not just an ad hoc prompt — relevant if a future revision routes part of a run through agy directly. |
|
|
80
|
+
| **[owner]** + **[obs]** Model choice inside agy is not free-form. **`gemini-3.7-flash-high` is the default discovery tier** (owner directive, 2026-08-15, superseding the prior `gemini-3.6-flash-medium` default once `gemini-3.7-flash-*` shipped — see § agy headless below for the full re-probed model list): ~1M context, generous quota. Its weakness is carelessness, not capacity — it skims. The counter is prompt precision, not a bigger model: name the exact paths, the exact question, and the exact output shape, leaving it nothing to improvise. **`claude-sonnet-4-6` / `claude-opus-4-6-thinking` inside agy are quota-scarce even on a Pro plan** (owner-reported) and additionally sit on the no-cache resume curve above. | Route discovery to `gemini-3.7-flash-high` by default and hand it a fully-specified task. Reach for agy's Claude tiers only for a single-shot, self-contained, high-value call where context and cache are demonstrably under control — never for a conversation, never as a habit. When strong-model judgment is needed *and* stateful, that is a Claude session id, not agy. |
|
|
81
|
+
| **[obs]** A flash-tier worker (`gemini-*-flash-*`, any generation) is for **retrieval, never for judgment**. | akiflow's thinking floor turns on the FACT/CONSTRAINT/ASSUMPTION distinction, which the skill already names as the one unrecoverable error to mislabel — exactly what a cheap model does worst. Hard rule wherever this mechanism is used, in the same voice as the existing "never downgrade implementation to save cost": retrieval only. |
|
|
82
|
+
|
|
83
|
+
## Cost model
|
|
84
|
+
|
|
85
|
+
| Fact | Design consequence |
|
|
86
|
+
|---|---|
|
|
87
|
+
| **[doc]** Prompt caching has a limited TTL — 5 minutes by default, 1 hour on the extended option. Sessions differ. | Plain subagents spawned in one batch, sharing the same rule-file prefix, hit a warm cache; a subagent spawned late (after a long Phase A) pays a colder read instead. Reason enough, on top of the sibling-roster fact above, to convene the roster in one batch (Step 2). |
|
|
88
|
+
| **[obs]** Every `SendMessage` costs a full turn of the receiving agent. | Peer-to-peer is not free; it merely skips the lead. Budget it in the roster brief. |
|
|
89
|
+
| **[obs]** A skill's `SKILL.md` loads only when the skill is invoked (progressive disclosure); `references/` and `scripts/` load only when read or run. | Keep the runnable contract in `SKILL.md`; park the reasoning here so it costs nothing until someone needs it. |
|
|
90
|
+
| **[obs]** Claude Code writes the session transcript as JSONL under `~/.claude/projects/<cwd-path-with-slashes-as-dashes>/<session-id>.jsonl`. Every assistant turn is one line carrying `message.model` and `message.usage` (`input_tokens`, `output_tokens`, `cache_creation_input_tokens`, `cache_read_input_tokens`). **A subagent's turns are in a separate file, not the parent's** — `<session-id>/subagents/agent-<agent-id>.jsonl`, each with a sidecar `agent-<agent-id>.meta.json` carrying `agentType`, `description`, `toolUseId`, `spawnDepth` and `model`. That directory is flat even for nested spawns: a `spawnDepth: 2` seat sits beside its parent and names it in `parentAgentId`. `"isSidechain": true` marks subagent rows and appears **only** in those files — measured 2026-08-15 across 1094 transcripts on this machine, 71,839 occurrences (in 876 files), none in a main-session file. (First observed 2026-07-31, when both lived in one file; re-verified and corrected 2026-08-15.) | Per-agent token accounting needs no new bookkeeping during the run — the numbers already exist. A cheap subagent parses them at close-out (`scripts/council_cost.py`, Step 6) and aggregates in-shell. Labels come from the sidecar, and must be `agentType` **plus** `description`: real rooms spawn almost every seat as `general-purpose`, so `agentType` alone silently merges unrelated seats into one row. **Dollar prices are not in the transcript and drift — never bake a price table into the script; multiply tokens by the current per-model price at report time.** |
|
|
91
|
+
|
|
92
|
+
## Model tiers
|
|
93
|
+
|
|
94
|
+
Family names, not version ids — the mapping outlives any single release.
|
|
95
|
+
|
|
96
|
+
| Work | Tier | Why |
|
|
97
|
+
|---|---|---|
|
|
98
|
+
| Decomposition, arbitration, the final call (the lead) | top tier | every downstream cut inherits this |
|
|
99
|
+
| Adversarial review, business/UX judgment | top tier, high effort | judgment, and the output the owner acts on |
|
|
100
|
+
| Implementation | mid tier or above, never downgraded to save cost | code quality is created at the keyboard, not recovered in review |
|
|
101
|
+
| Verification against a written promise | mid tier, plain subagent handed the diff + the promise | mechanical comparison, but must know what was promised |
|
|
102
|
+
| Mechanical sweeps: inventory, grep, call-site lists, stats | cheapest capable tier, low effort | the task describes itself; instruct it to aggregate in-shell rather than pulling raw data into context |
|
|
103
|
+
| Read-only retrieval via the cross-CLI worker (agy headless, flash tier, `--mode plan`) | cheapest available tier, mechanical retrieval only, never judgment | `--mode plan` enforces read-only by mechanism rather than by prompt; the FACT/CONSTRAINT/ASSUMPTION judgment stays with the lead — see § Cross-CLI worker |
|
|
104
|
+
|
|
105
|
+
### Host resolution — a tier is a word, the host supplies the model
|
|
106
|
+
|
|
107
|
+
Skills are deployed unmodified to five hosts, and **[doc]** Cursor additionally reads `~/.claude/skills/`, `~/.agents/skills/`, `~/.claude/agents/` and `~/.codex/agents/` for compatibility (cursor.com/docs/skills, cursor.com/docs/subagents, checked 2026-09-08). So any model name written into a skill or agent file is read by every host, and a Claude alias in a shared file becomes a Claude call on a host that has cheaper native models. The rule: **a roster, lane, or brief names a tier — `top` / `mid` / `cheap` — and the host running the skill resolves it through its own row below. A skill never names another vendor's model.** The one deliberate exception is `claude/agents/*.md` frontmatter: `model: sonnet` / `model: haiku` are the Claude Code row written in the only syntax Claude Code reads, and Claude Code is the primary host, so they stay exact rather than degrading to `inherit`. On a non-Claude host that value is not a native model id: declare the seat's model explicitly at spawn from the host's row, and treat the frontmatter as Claude-only.
|
|
108
|
+
|
|
109
|
+
| Host | `cheap` — retrieval, never judgment | `mid` — implementation, verification | `top` — lead, adversarial judgment | Where the tier is set | Status |
|
|
110
|
+
|---|---|---|---|---|---|
|
|
111
|
+
| Claude Code | `haiku` (no `--effort`) | `sonnet` | `opus`, or `inherit` from the lead | agent frontmatter `model:`; Agent tool `model`; `claude -p --model <alias> --effort <e>` | **[obs]** verified across this file |
|
|
112
|
+
| Cursor (IDE + `agent` CLI) | Composer family — the current id from Cursor's model picker (`composer-2`-style), never an API-pool Claude/GPT model | `inherit` (the session's model) | `inherit`, or the session's top API-pool model | `.cursor/agents/*.md` or `~/.claude/agents/*.md` frontmatter `model: inherit \| <id>[effort=…]`; `agent -p --model <id>` | **[doc]** field and syntax; **UNCONFIRMED** how Cursor treats a Claude alias (`haiku`) it cannot resolve — reopen trigger: one measured Cursor run of an `aki-hands` spawn |
|
|
113
|
+
| Antigravity `agy` | `gemini-3.7-flash-high` (§ Cross-CLI worker) | `gemini-3.1-pro-low` | `gemini-3.1-pro-high` (`claude-opus-4-6-thinking` is quota-scarce, § agy headless) | `agy --model <slug>` — effort is inside the slug; agy 1.1.6+ agent markdown carries `model` | **[obs]** 2026-08-15 |
|
|
114
|
+
| Codex CLI | UNCONFIRMED low-cost alias | `[agents] default_subagent_model` in `config.toml`; `codex exec -c model=<id> -c model_reasoning_effort=medium` | same model, `model_reasoning_effort=xhigh` | `config.toml [agents]`, per-agent `model`; `codex exec -c …` | **[doc, secondary]** 2026-09-08, model ids drift monthly — read them from `codex` itself |
|
|
115
|
+
| Kiro CLI | `qwen3-coder-next` (0.05×) or `claude-haiku-4.5` (0.4×) | `auto` (1×) or `claude-sonnet-4.5` (1.3×) | Opus-class, ~22× — rarely worth it on this host | `kiro-cli chat --no-interactive --model <id> --effort <e>`; custom agent JSON `model` | **[obs]** 2026-08-02 list; multipliers re-read with `--list-models` |
|
|
116
|
+
| Grok CLI | UNCONFIRMED | `grok-build-0.1` default | UNCONFIRMED | `grok -p` (`--model` flag unconfirmed) | **[doc, secondary]** 2026-09-08 |
|
|
117
|
+
|
|
118
|
+
The `model` key in `SKILL.md` frontmatter is a Claude Code extension, rejected by the open-standard validator (github.com/anthropics/claude-code/issues/25380) — never put a tier or a model there.
|
|
119
|
+
|
|
120
|
+
## Headless — the cost levers, per CLI
|
|
121
|
+
|
|
122
|
+
Full narrative and the measurements behind these rows: `docs/research/headless-cli-workers-aug1.md`.
|
|
123
|
+
|
|
124
|
+
### Claude Code (`claude -p`), verified 2026-08-01 against 2.1.220
|
|
125
|
+
|
|
126
|
+
| Flag | Fact | Design consequence |
|
|
127
|
+
|---|---|---|
|
|
128
|
+
| `--bare` | **[obs]** Skips hooks, LSP, plugin sync, attribution, auto-memory, background prefetches, keychain reads, and CLAUDE.md auto-discovery. **It also refuses OAuth — auth is strictly `ANTHROPIC_API_KEY` or an `apiKeyHelper`.** A live call on an OAuth-only machine returned `is_error: true`, `terminal_reason: "api_error"`, zero tokens. | The largest available input-token cut on the Claude side, and unusable without a separate API key. Do not write it into a mechanism that must work on the owner's normal login — but on an API-key proxy-gateway lane (§ claude via a proxy gateway below) the objection vanishes and the cut applies in full. |
|
|
129
|
+
| `--disallowedTools "Workflow DesignSync"` | **[obs]+[doc]** Session-level tool block, no auth change (unlike `--bare`, still OAuth). Verified 2026-08-03: `Workflow` is Claude Code's native autonomous multi-step orchestrator (code.claude.com/docs/en/common-workflows) — akiflow's own `SKILL.md` (Step 1) already states it only *tells the owner* to invoke Workflow, never calls it itself, so blocking it removes nothing akiflow uses. `DesignSync` is the `/design-sync` bridge to claude.ai/design (design-token/component import-export, support.claude.com) — a different product surface with zero overlap with akiflow. | The OAuth-compatible alternative to `--bare` for cutting tool/context surface: same intent (fewer tools loaded at session start), none of `--bare`'s auth cost, and zero functional loss for akiflow specifically. Owner alias: `cl-9rt-min='CLAUDE_CONFIG_DIR="$HOME/.claude-9rt" claude --disallowedTools "Workflow DesignSync"'`. |
|
|
130
|
+
| `--tools "Read,Grep"` | **[obs]** Restricts the session to a named subset of built-in tools; `""` disables all. | Read-only **by mechanism** on the Claude side — the missing counterpart to agy's `--mode plan`. Prefer it over telling a subagent to behave. |
|
|
131
|
+
| `--json-schema` | **[obs]** Enforces structured output. `agy --json-schema` is the same capability. | The one Workflow feature worth having is available headless on both CLIs, so it is not a reason to adopt Workflow. |
|
|
132
|
+
| `--max-budget-usd` | **[obs]** Hard dollar cap; `-p` only. | Preventive counterpart to Step 6's post-hoc tally, and the second Workflow-only feature that turns out not to be Workflow-only. |
|
|
133
|
+
| `--effort low` | **[obs]** `low\|medium\|high\|xhigh\|max`. Present on `claude`, `agy`, and `kiro-cli`. **`claude` + haiku**: no `--effort` option — haiku has no extended thinking at API level; flag silently ignored. | Two cheapness axes: model tier **and** thinking budget. Cut both on sweeps. On `claude` CLI with haiku, only the model axis is available. |
|
|
134
|
+
| `--exclude-dynamic-system-prompt-sections` | **[doc, in `--help`]** Moves cwd/env/memory/git-status out of the system prompt into the first user message, improving cross-call prompt-cache reuse. | Matters when the same worker shape is called many times in a run — cache hits, not flag count, are what make fan-out cheap. |
|
|
135
|
+
| `--agents <json>` | **[obs]** Defines custom agents inline, as JSON, per call. | A worker roster can be declared at the call site without installing anything. |
|
|
136
|
+
| `--permission-mode plan`, `--add-dir`, `--no-session-persistence`, `--fallback-model`, `--disable-slash-commands`, `--setting-sources` | **[obs]** Present. | Scope, durability, and resilience are all per-call settable; a headless worker need not inherit the caller's environment. |
|
|
137
|
+
|
|
138
|
+
### claude via a proxy gateway (9router) — a parallel worker lane every machine should have
|
|
139
|
+
|
|
140
|
+
**[obs]** A gateway like 9router provisions a separate Claude config dir (`CLAUDE_CONFIG_DIR=~/.claude-9rt`) pointing the same `claude` binary at its own endpoint + API key — same harness, separate metering, and the gateway decides which core actually serves each model alias. Owner's `cl-9rt-min` alias pairs it with `--disallowedTools "Workflow DesignSync"` (row above) for a minimal-surface session. **`cl-9rt`/`cl-9rt-min` is a shell alias, not a binary** — it is defined in the owner's interactive shell rc and does not exist in a spawned/non-interactive process. A worker or script must run the expanded literal command (`CLAUDE_CONFIG_DIR=~/.claude-9rt claude ...`, `aki-hands.md` § Substrates), never the alias name.
|
|
141
|
+
|
|
142
|
+
| Fact | Design consequence |
|
|
143
|
+
|---|---|
|
|
144
|
+
| Separate config dir = separate quota, metered by the gateway, concurrent with the lead's session | A genuine **parallel** lane, not a spillover for "primary quota exhausted" — a second real worker runs alongside the lead without competing for the same tokens. |
|
|
145
|
+
| **[owner]** A gateway alias may route to a non-Anthropic core: the owner's lane's "Sn" alias is a deepseek-v4-pro-class model (2026-08-03). Which core answers is the gateway's routing table — per machine, re-verify before relying on model-specific behavior. | Treat the lane as an **explore/bandwidth seat** beside agy flash, never a second seat of the lead's model: strong enough for wide parallel exploration and synthesis, while judgment and arbitration stay with the lead. |
|
|
146
|
+
| Auth on this lane is the gateway's API key, not OAuth — so `--bare` (rejected for the main login, table above) **works here**. | The largest token cut becomes usable: `--bare --tools "Read,Grep,Bash"` is the minimal fresh-context worker shape for this lane. |
|
|
147
|
+
| Depends on a config dir the owner provisions (`~/.claude-9rt`); not present by default on every machine | Probe `test -d ~/.claude-9rt` like any cross-CLI lane. Where absent, **recommend the one-time setup** (CONFIG_DIR + gateway endpoint — every machine benefits from having this lane) instead of silently substituting another mechanism. |
|
|
148
|
+
|
|
149
|
+
*Consequence:* the quota-payer axis doubles as a parallelism axis — the lead plus one or more gateway workers run at once, each metered on its own account.
|
|
150
|
+
|
|
151
|
+
### agy headless — see § Cross-CLI worker above
|
|
152
|
+
|
|
153
|
+
**[obs]** Re-probed 2026-08-15 (prior check 2026-08-02 predates the `gemini-3.7-flash-*` release — do not cite the old list), `agy models`: `gemini-3.7-flash-{low,medium,high}`, `gemini-3.6-flash-{low,medium,high}`, `gemini-3.5-flash-{low,medium,high}`, `gemini-3.1-pro-{low,high}`, **`claude-sonnet-4-6`**, **`claude-opus-4-6-thinking`**, `gpt-oss-120b-medium`. Also present: `--json-schema`, `--effort`, `--agent`, `--add-dir`, `--print-timeout`, `--disable-slash-commands`, and an `agents` subcommand (empty on this machine — no custom agy agents defined).
|
|
154
|
+
|
|
155
|
+
*Consequence:* a Claude-family model can be reached **on the Antigravity quota**. The vendor paying and the model reasoning are independent choices, which is a second axis the Step 2 mechanism table did not previously have.
|
|
156
|
+
|
|
157
|
+
### Kiro CLI (`kiro-cli` 2.16.0) — **[obs]**, verified 2026-08-02
|
|
158
|
+
|
|
159
|
+
| Flag / Feature | Fact | Design consequence |
|
|
160
|
+
|---|---|---|
|
|
161
|
+
| Batch | `kiro-cli chat --no-interactive "<prompt>"` — positional arg also accepted | Drop-in headless substrate, same invocation shape as `claude -p` |
|
|
162
|
+
| Read-only | `--trust-tools=` blocks all; `--trust-tools=fs_read,fs_write` restricts to named set; `-a`/`--trust-all-tools` approves all | Read-only **by mechanism** — equivalent of `--mode plan` / `--tools ""`. Prefer over prompt wording. |
|
|
163
|
+
| Fail-loud | `--require-mcp-startup` — exit code 3 if any MCP server fails to start | A broken worker fails loudly instead of silently degrading |
|
|
164
|
+
| Effort | `--effort low\|medium\|high\|xhigh\|max` — present and operative on all model tiers | Unlike `claude` CLI + haiku, `--effort` is not a no-op on any Kiro model |
|
|
165
|
+
| Session | `--resume-id <ID>`, `-r` (most recent), `--resume-picker`; `--list-sessions`, `--delete-session` | Same persistent-worker pattern as `claude -p --session-id`; cwd-scoped |
|
|
166
|
+
| Engine | `--agent-engine v1\|v2\|v3` (v2 default); `--mode default\|spec` with v3; `--agent <name>`; built-ins: `kiro_default`, `kiro_planner`, `kiro_help`; custom in `.kiro/agents` or `~/.kiro/agents` | v3 spec mode available when needed |
|
|
167
|
+
| ACP | `kiro-cli acp` — exposes Kiro as an Agent Client Protocol server | Only machine protocol in this stack; an external orchestrator can drive Kiro directly |
|
|
168
|
+
| Models | **[obs]** (`--list-models`, 2026-08-02): `auto`(1×), `claude-sonnet-4.5`(1.3×), `claude-haiku-4.5`(0.4×), `minimax-m2.5`(0.25×), `glm-5`(0.5×), `qwen3-coder-next`(0.05×, 256k) | Cheapest bandwidth tier in this stack; only one exposing a machine protocol (ACP) |
|
|
169
|
+
|
|
170
|
+
### Stateful workers — both CLIs offer one, and they behave oppositely
|
|
171
|
+
|
|
172
|
+
**[obs]** Controlled test, 2026-08-01. Store a token in turn 1, ask for it back in turn 2, send a trivial turn 3. Both CLIs recalled it correctly, so *functionally* both resume. Economically they are not comparable:
|
|
173
|
+
|
|
174
|
+
| | `agy --conversation <id>` | `claude -p --session-id <uuid>` / `--resume <uuid>` |
|
|
175
|
+
|---|---|---|
|
|
176
|
+
| Turn 1 | 26.5k input, 2.6s | $0.0358 — 16.9k cache-creation, 17.5k cache-read |
|
|
177
|
+
| Turn 2 | 53.6k input, **`cache_read_tokens: 0`**, 8.8s | **$0.0043** — 34.4k cache-read, 219 new. ~8× cheaper than turn 1 |
|
|
178
|
+
| Turn 3 (trivial prompt) | 56.4k input, 24.5k cache-read, **57.6s** | $0.0038 — 34.6k cache-read, 90 new |
|
|
179
|
+
| Shape of the curve | input grows every turn, cache intermittent, **latency 2.6s → 8.8s → 57.6s** | flat: cost and latency stable from turn 2 onward |
|
|
180
|
+
| Lookup scope | conversation id resolves globally | **cwd-scoped** — resuming the same id from another directory fails with `No conversation found` (exit 5) |
|
|
181
|
+
|
|
182
|
+
Design consequences, and they are the sharpest in this file:
|
|
183
|
+
|
|
184
|
+
1. **`claude -p --session-id` is the persistent-worker mechanism this skill was missing.** A named uuid, called repeatedly, keeps its full history and gets *cheaper* after the first turn because the prefix is cached. It is not a fork — it is better for this purpose, because it survives between top-level sessions. Continuity no longer has to travel exclusively by plan-doc handoff; a long-lived worker can simply be re-addressed.
|
|
185
|
+
2. **Its cwd-scoping is a feature, not a limitation.** A worker id is bound to the project directory it was created in, so a worker cannot be accidentally re-addressed from the wrong repo.
|
|
186
|
+
3. **agy conversations are a trap. Use agy one-shot only.** It resumes correctly and looks fine at turn 2, then the latency curve makes it unusable — 57.6 seconds to answer "reply ok" on turn 3. Anything needing more than one exchange belongs on a Claude session id; agy's value is a fast, wide, self-contained single call.
|
|
187
|
+
4. **The cheap axis and the stateful axis are different mechanisms.** agy flash is cheap *per call* and stateless in practice; a Claude session id is cheap *per turn after the first* and stateful. Choosing "cheap" without saying which of the two is meant is how a run ends up paying full price for both.
|
|
188
|
+
|
|
189
|
+
### What breaks in headless mode, on every CLI
|
|
190
|
+
|
|
191
|
+
Two things, and both are structural:
|
|
192
|
+
|
|
193
|
+
1. **Nobody can answer an owner escalation.** The lead must not guess what the owner would have wanted. Record the escalation in the checklist as `BLOCKED: needs owner` and stop that item — the rest of the run continues.
|
|
194
|
+
2. **Nobody can answer a permission prompt.** Anything requiring approval fails rather than waits. Headless work must be scoped to what the current permissions already allow.
|
|
195
|
+
|
|
196
|
+
A bare cheap-model headless call is the right shape for a self-contained mechanical question — one that fits in a couple of hundred words with no project context and returns a short answer. It is the wrong shape for anything that would need to ask a follow-up question.
|
|
197
|
+
|
|
198
|
+
## Sources
|
|
199
|
+
|
|
200
|
+
Claude Code and Antigravity documentation both move; verify a detail before relying on it.
|
|
201
|
+
|
|
202
|
+
- Subagents (fork/subtask, `CLAUDE_CODE_FORK_SUBAGENT`) — <https://code.claude.com/docs/en/sub-agents>
|
|
203
|
+
- Run agents in parallel (subagents vs agent view vs agent teams vs workflows) — <https://code.claude.com/docs/en/agents>
|
|
204
|
+
- Changelog (version-dated behavior changes, e.g. the `/fork` → `/subtask` rename) — <https://code.claude.com/docs/en/changelog>
|
|
205
|
+
- Agent Skills — <https://code.claude.com/docs/en/skills>
|
|
206
|
+
- CLI reference (`-p` / non-interactive) — <https://code.claude.com/docs/en/cli-reference>
|
|
207
|
+
- Settings and permissions — <https://code.claude.com/docs/en/settings>
|
|
208
|
+
- Prompt caching and TTL — <https://docs.claude.com/en/docs/build-with-claude/prompt-caching>
|
|
209
|
+
- Cursor skills discovery paths (incl. `~/.claude/skills/`, `~/.agents/skills/`) — <https://cursor.com/docs/skills>
|
|
210
|
+
- Cursor subagent frontmatter (`model: inherit | <id>[effort=…]`, reads `~/.claude/agents/`) — <https://cursor.com/docs/subagents>
|
|
211
|
+
- Cursor headless CLI (`agent -p --model`) — <https://cursor.com/docs/cli/headless>
|
|
212
|
+
- Antigravity headless model slugs — <https://antigravity.google/docs/cli/headless/>
|
|
213
|
+
- `model` in SKILL.md is a Claude Code extension, rejected by the open-standard validator — <https://github.com/anthropics/claude-code/issues/25380>
|
|
214
|
+
|
|
215
|
+
Entries marked **[obs]** are not in these pages, including all of the Antigravity/AGY and cross-CLI rows above — no equivalent published reference for `agy`'s internal mechanisms was found; they were observed directly against the `claude` and `agy` binaries and a live transcript, and are recorded here so a future reader can tell the difference between what is documented and what is merely believed.
|
|
@@ -0,0 +1,4 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# Transitional Unix wrapper — the Python file is the SSOT (cross-platform, incl. Windows).
|
|
3
|
+
# Kept so existing prompts/docs that name "council-verify.sh" keep working on Unix during the transition.
|
|
4
|
+
exec python3 "$(dirname "$0")/council_verify.py" "$@"
|