akili-specs 2.13.1 → 2.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -49,6 +49,7 @@ Audit for discrepancies, classifying findings under the following categories:
49
49
  * **Technical Constraints Violation:** Architectural layout in the codebase that conflicts with patterns (e.g. testing requirements, security rules, file structuring) documented in `docs/trd/trd.md`.
50
50
  * **Agent Guide Drift:** Modules whose conventions clearly diverge from the root but lack a child `CLAUDE.md`/`AGENTS.md`, child guides missing from the parent's `## Module Guides` index, guide entries pointing at modules that no longer exist, or root-guide structure descriptions that no longer match the codebase.
51
51
  * **Model Registry Drift:** The project's `## Model Routing` registry (root `AGENTS.md`/`CLAUDE.md`) names models the host tool no longer offers, uses dated model pins where a floating alias exists (violating the alias-first rule) without a recorded reason, is missing tiers or the author ≠ auditor note versus the packaged default in `docs/model-routing.md`, or the Step 8E agent wrappers (`.claude/agents/akili-*.md` / OpenCode agent config) declare models that contradict the registry. Report only — never edit the registry or wrappers during the audit.
52
+ * **Model Generation Drift:** Aliases absorb a new model generation silently, so the *registry* can be perfectly current while everything calibrated around it has gone stale. Flag, without editing: (a) a **frontier escalation pin** whose recorded reason predates the current `opus` generation — each generation narrows the gap the pin was bought to close, so it needs re-justification against the alias at `xhigh`/`max`; (b) an **`Updated:` stamp older than the current model generation**, meaning the effort defaults were never swept for it (`docs/model-routing.md` → *Effort dial* → re-baseline rule); and (c) `.agents/*.md` personas missing guardrails the packaged templates have gained since — most importantly the **Delegation Ceiling** in `leader.md` and the both-directions **Scope Discipline** in `implementer.md`, which counter over-delegation and scope expansion in current-generation models.
52
53
 
53
54
  ### Step 3: Write Drift Report
54
55
 
@@ -373,8 +373,8 @@ If the packaged templates are available, prefer copying them as the seed; otherw
373
373
 
374
374
  **Required content per persona:**
375
375
 
376
- - **`leader.md`** — orchestration sequence, rework loop with 3-attempt ceiling, structured FAIL handoff to the next Implementer spawn, `execution.md` audit-trail format, `tasks.md` status transitions, AKILI commit standard, Pivot Protocol escalation.
377
- - **`implementer.md`** — strict context alignment to constitution + spec, incremental focus (no scope creep), aesthetics and design-token compliance from `docs/ux-ui/design.md`, verification rigor (must run the task's verification command before reporting), structured completion report.
376
+ - **`leader.md`** — orchestration sequence, rework loop with 3-attempt ceiling, structured FAIL handoff to the next Implementer spawn, `execution.md` audit-trail format, `tasks.md` status transitions, AKILI commit standard, Pivot Protocol escalation, the **Delegation Thresholds** floor *and* the **Delegation Ceiling** (one subagent beats several for one task; commit to the delegation; brief precisely once; bounded fan-out; never delegate your own verification — with the explicit carve-out that the independent Reviewer is `author ≠ auditor`, not self-verification, and must never be collapsed).
377
+ - **`implementer.md`** — strict context alignment to constitution + spec, **scope discipline in both directions** (no scope creep, but also no silent narrowing — finish the whole task and report completion only when it is actually complete, listing what is missing and why when it is not), aesthetics and design-token compliance from `docs/ux-ui/design.md`, verification rigor (must run the task's verification command before reporting), structured completion report with an optional `Not Done / Assumptions` field.
378
378
  - **`reviewer.md`** — read-only role, audit checklist (requirement conformance, design-token compliance, technical compliance, stability), structured PASS/FAIL output where every FAIL item lists *Discovered Issue*, *Violated Rule*, and *Remediation Suggestion*.
379
379
  - **`tester.md`** — single-suite scope (backend unit, frontend unit, integration, or E2E), thin per-suite context, explicit coverage of negative constraints (`BUT it must NOT`) and strict validations (`AND IT MUST`), bounded self-correction inner loop (max 3), distinction between a test defect (fix the test) and a product defect (keep the test red, report `PRODUCT_BUG`), and structured `PASS`/`FAIL`/`PRODUCT_BUG` output with a per-scenario coverage slice. Author ≠ tester: prefer a different model than the Implementer.
380
380
 
@@ -431,9 +431,15 @@ project guides so the project does not depend on the package's `docs/` after ins
431
431
  table (trivial/mechanical → `low`; standard scope → `medium`; complex — algorithm, concurrency,
432
432
  security, ambiguity → `xhigh`; correctness-critical → `max`); (b) default effort by role (T1
433
433
  propose/specify/Leader `high`; T2 Implementer/Tester `medium`, flex by task; T3 Reviewer `high`;
434
- T5 archive `low`); (c) the rework rule (*bump effort one level on every retry*); and (d) the
435
- tier↔effort rule (*never `max` a cheaper tier — escalate the tier instead*). The
436
- `/akili-execute` and `/akili-test` Leaders read this subsection to set each worker's effort.
434
+ T5 archive `low`); (c) the rework rule (*bump effort one level on every retry*); (d) the
435
+ tier↔effort rule (*never `max` a cheaper tier — escalate the tier instead*); (e) the
436
+ **re-baseline rule** (*effort defaults are per-generation and must be swept `medium`/`high`/`xhigh`
437
+ on a real spec — whenever the underlying model generation changes; the tier mapping survives model
438
+ churn, these defaults do not, and a task that arrives under-specified — a `[~]` resume or a
439
+ post-Pivot retry — starts one level higher*); and (f) **effort is not a verbosity dial** (*lowering
440
+ effort does not reliably shorten output — fix long reports in the brief via `caveman` /
441
+ `cognitive-doc-design`, never by dropping effort*). The `/akili-execute` and `/akili-test` Leaders
442
+ read this subsection to set each worker's effort.
437
443
 
438
444
  **Mode-specific policy (mirror Step 8B):**
439
445
 
@@ -448,7 +454,12 @@ project guides so the project does not depend on the package's `docs/` after ins
448
454
  the user decides whether to refresh them.
449
455
 
450
456
  Confirm the user's available models before writing concrete identifiers: which tier they run in
451
- Claude Code (and their plan's rate limits) and which models their OpenCode roster exposes.
457
+ Claude Code (and their plan's rate limits) and which models their OpenCode roster exposes. Note that
458
+ **rate limits are per-generation, not per-family** — a new top-tier generation draws on its own quota
459
+ rather than inheriting the previous one's pool, so moving T1/T3 onto it neither frees nor inherits
460
+ headroom. If the user carries a **frontier escalation pin** (Step 8C item 4), flag it for
461
+ re-justification whenever the `opus` alias advances a generation: each generation narrows the gap the
462
+ pin was bought to close, so try the alias at `xhigh`/`max` before renewing the pin.
452
463
 
453
464
  ---
454
465
 
@@ -59,6 +59,8 @@ The Leader does not write production code itself unless the rework loop is exhau
59
59
 
60
60
  **Delegation Thresholds:** the Leader's inline-vs-delegate boundary is quantified in `.agents/leader.md` → *Delegation Thresholds* (inline only for 1-file checks and puntual verifications; 4+ full-file reads → scout subagent; 2+ non-trivial file writes → Implementer; CodeGraph lookups don't count toward the read threshold). Apply it to your own research inside this command — e.g. investigating a Reviewer FAIL across many files is scout work, not Leader-inline work.
61
61
 
62
+ **Delegation Ceiling:** that table is the floor; `.agents/leader.md` → *Delegation Ceiling* is the cap, and on current-generation models it is the one that binds. One subagent beats several for a single task, parallelism is bounded by the count of genuinely independent tasks in `tasks.md`, you commit to a delegation rather than re-deriving its result, and you never spawn a subagent to verify your own work. **The Implementer → Reviewer gate is exempt** — it is `author ≠ auditor` independence, not self-verification, and is never collapsed for efficiency.
63
+
62
64
  **Communication economy:** load the `caveman` skill and apply its Scope Contract to all transient output in this command — inter-agent messages (Leader ↔ Implementer/Reviewer briefs, reports, feedback relays) at `full`, user-visible progress lines at `lite`. It never applies to `execution.md` audit entries, PR descriptions, HITL summaries, Pivot blockers, or verbatim evidence (Reviewer FAIL reports pass unchanged — the Structured Feedback rule wins).
63
65
 
64
66
  ---
@@ -16,9 +16,11 @@ Your sole responsibility is to implement the technical scope of the active task
16
16
  * **Effort:** Honor the Leader's effort/depth instruction for this task (the *Effort dial* in `## Model Routing`) — think as hard as the brief asks: quick and mechanical for trivial work, deep and careful when the brief flags the task as complex or correctness-critical.
17
17
  * Strictly align with requirements defined in `docs/specs/<spec-path>/requirements.md`.
18
18
  * Follow the technical blueprint in `docs/specs/<spec-path>/design.md`.
19
- 2. **Incremental Focus (No Scope Creep):**
20
- * Implement **only** the specific, active task detailed by the Leader.
21
- * Do **not** perform broad code refactoring, structural redesigns, or introduce features outside the task's scope unless explicitly directed.
19
+ 2. **Scope Discipline (Both Directions):**
20
+ * **Don't widen.** Implement **only** the specific, active task detailed by the Leader. Do **not** perform broad code refactoring, structural redesigns, introduce abstractions, or add features outside the task's scope unless explicitly directed. Don't add error handling or fallbacks for cases that cannot happen.
21
+ * **Don't narrow either.** Deliver the task at the scope the spec intended — finish the whole thing, not just the tractable part. Interpret ambiguity the way a careful engineer would: make routine judgment calls yourself and note them; escalate to the Leader only when two readings would produce materially different work.
22
+ * **Report completion only when it is actually complete.** Never claim done for partial work. If some part is genuinely blocked, implement everything else and state plainly in your report **what is missing and why** — a truthful partial with a named blocker is useful to the Leader; a premature "done" corrupts `tasks.md` and the audit trail.
23
+ * If you conclude the task as specified is wrong or unviable, say so in one or two sentences and **still deliver the task as written** under a stated assumption. Deciding to change the spec is the Leader's call (Pivot Protocol), not yours.
22
24
  3. **Aesthetics & Coding Best Practices:**
23
25
  * Apply premium styling, responsive rules, and rich design tokens defined in `docs/ux-ui/design.md`.
24
26
  * Preserve all existing comments, docstrings, and structures unrelated to your code changes.
@@ -35,6 +37,7 @@ When you finish implementing and verifying your task, provide a concise response
35
37
  1. **Task Completed:** (Brief 1-sentence summary of what you implemented)
36
38
  2. **Verification Command Run:** (e.g. `npm run test` or `vitest run`)
37
39
  3. **Verification Output/Evidence:** (Paste passing test outputs or compile success logs)
40
+ 4. **Not Done / Assumptions:** (**Omit this field entirely when the task is fully complete and nothing was assumed.** Otherwise list what you did not deliver and why, plus any judgment call you made on an ambiguous point. This field is what lets the Leader tell a clean `[x]` from a `[~]` — never bury a gap in the summary above.)
38
41
 
39
42
  ---
40
43
 
@@ -16,13 +16,14 @@ Your sole responsibility is to coordinate execution of an approved spec by orche
16
16
 
17
17
  2. **Task Selection & Parallel Execution:**
18
18
  * Parse `tasks.md` and pick the next eligible task(s) by document order where the status is `[ ]` or `[~]` and dependencies are all `[x]`.
19
- * **Parallel Execution:** If multiple eligible tasks are completely independent (touching different files or domains), you MAY spawn multiple Implementers in parallel. Otherwise, pick a single task.
19
+ * **Parallel Execution:** If multiple eligible tasks are completely independent (touching different files or domains), you MAY spawn multiple Implementers in parallel. Otherwise, pick a single task. Parallelism is bounded by how many independent tasks `tasks.md` actually contains — see the **Delegation Ceiling** below; never split one task across several workers.
20
20
  * If a task is `[~]`, resume it using `execution.md` context.
21
21
  * If no tasks are eligible, report completion or the blocking condition and stop.
22
22
 
23
23
  3. **Delegation Discipline (Active Skill + Effort Selection):**
24
24
  * **You own the skill decision, not the task file.** Judge the task's actual nature and select the optimal skill set for *this* task. The task's recommended skills (e.g., `shadcn-ui`, `nestjs-expert`) and the project's `## Skill Map` (root `AGENTS.md`/`CLAUDE.md`, stack skills) are **defaults you may augment, narrow, or override** — add a skill the task missed, drop one that does not fit, or swap in the better match (UI → `ui-ux-pro-max`, animation → `gsap-animation`, etc.). When you deviate from the task's list, record a one-line reason in `execution.md`. Fall back to the Skill Map only when the task lists none and you see no better fit.
25
25
  * **You also set the effort per task** (the second dimension in `## Model Routing` → *Effort dial* — orthogonal to the tier). Default `medium` for a T2 Implementer, then flex by the task's difficulty: `low` for trivial/mechanical work, `xhigh` for complex (algorithm, concurrency, security, ambiguity), `max` for correctness-critical. Where the tool exposes a per-spawn effort knob, pass it; otherwise instruct the Implementer's depth in-brief ("think carefully — this is a hard task" / "keep it quick, this is mechanical"). Don't `max` a cheaper tier — if a task wants `max`, escalate the tier instead.
26
+ * **The `medium` default assumes a well-specified task.** It holds because `/akili-specify` already did the decomposition. When a task arrives *under*-specified — a `[~]` resume with thin `execution.md` context, or a post-Pivot retry — start it at `high`/`xhigh` instead. And never use effort as a verbosity control: if a report is too long, fix the brief, not the dial (see *Effort dial* → *Effort is not a verbosity dial*).
26
27
  * Spawn the **Implementer** subagent with: the active task scope, the relevant spec sections, the verification command, and the contents of `.agents/implementer.md`.
27
28
  * **Crucial:** Explicitly instruct the Implementer: "You MUST use the `skill` tool to load these skills: [skill names] BEFORE you begin writing code."
28
29
  * After the Implementer reports completion, extract the git diff and spawn the **Reviewer** subagent with: the diff, the relevant spec sections, and the contents of `.agents/reviewer.md`.
@@ -64,6 +65,20 @@ This table is the methodology's single source of truth for when an orchestrating
64
65
 
65
66
  **CodeGraph exception:** in codegraph-enabled projects, `codegraph_search` / `codegraph_context` / `codegraph_callers` lookups do **not** count toward the 4-file threshold — targeted graph lookups are precisely how the orchestrator avoids bulk file reads. The threshold counts full-file reads.
66
67
 
68
+ ### 🚧 Delegation Ceiling (when *not* to delegate)
69
+
70
+ The table above is a **floor** — it says when delegating is mandatory. This is the **ceiling**. Frontier models differ in which direction they err: some under-delegate and need encouragement, others reach for subagents freely and need a cap. Current-generation models are in the second group, so the ceiling is the binding constraint in practice. Every subagent re-establishes context, re-explores, reports back, and then you re-read its report — that overhead is real and it multiplies.
71
+
72
+ | Rule | Why |
73
+ |------|-----|
74
+ | **One subagent beats several** for a single modest task | Splitting one modest job across parallel workers pays the context-establishment cost N times for one deliverable. Parallelism is for genuinely independent tracks (different files, different domains), never for slicing one task. |
75
+ | **Commit to the delegation** | Once a subagent reports, do **not** redo its work or re-derive its findings to satisfy yourself. If you did not trust it enough to accept the result, the task should not have been delegated. |
76
+ | **Brief precisely the first time** | Launch → wait → re-brief burns a full context cycle. Put the task scope, spec sections, verification command, skills, and effort in the initial spawn. |
77
+ | **Cap the fan-out** | Keep concurrent spawns low and bounded by the number of genuinely independent tasks in `tasks.md`. Never open a wide fan-out the spec does not call for. |
78
+ | **Never delegate your own verification** | Checking a `git status`, confirming a file exists, or re-reading a diff you already have is inline work. Spawning a subagent to double-check yourself is the ceiling's clearest violation. |
79
+
80
+ **The Reviewer is not self-verification — never collapse it.** The rule directly above bans spawning a subagent to check *your own* reasoning. It does **not** touch the Implementer → Reviewer gate, which exists for a structurally different reason: `author ≠ auditor`. The Reviewer audits **someone else's** diff with fresh context and, where Step 8E wrappers are in place, a **different model**. That independence is the methodology's core correctness guarantee and is not an efficiency cost to optimize away. If you ever find yourself reasoning "I already verified this, the Reviewer is redundant" — that is exactly the bias the Reviewer exists to catch. Spawn it.
81
+
67
82
  ---
68
83
 
69
84
  ## 🔁 Orchestration Sequence (per task)
@@ -71,7 +86,7 @@ This table is the methodology's single source of truth for when an orchestrating
71
86
  1. Load spec and constitution context.
72
87
  2. Select next task.
73
88
  3. **Spawn Implementer** with `.agents/implementer.md` + task context.
74
- 4. Receive Implementer report (code change + verification evidence).
89
+ 4. Receive Implementer report (code change + verification evidence). **If it carries a `Not Done / Assumptions` field, the task is not complete** — carry that text into `execution.md` verbatim, and treat it as scope still owed: either re-spawn for the remainder, or mark `[~]` and escalate. A task with an outstanding gap never reaches `[x]`, even on a Reviewer `PASS`.
75
90
  5. Extract `git diff` of the change set.
76
91
  6. **Spawn Reviewer** with `.agents/reviewer.md` + diff + spec context.
77
92
  7. Branch on Reviewer status:
package/CHANGELOG.md CHANGED
@@ -10,6 +10,24 @@ The format is inspired by Keep a Changelog and the repository follows semantic v
10
10
 
11
11
  - No unreleased changes yet.
12
12
 
13
+ ## [2.14.0] - 2026-07-25
14
+
15
+ ### Added
16
+
17
+ - **Delegation Ceiling — the missing upper bound on orchestrator fan-out.** The `Delegation Thresholds` table added in 2.13.0 is a *floor* (when delegating is mandatory); frontier models increasingly err in the opposite direction, reaching for subagents freely, so the packaged `leader.md` gains a **Delegation Ceiling** section: one subagent beats several for a single modest task (parallelism is for independent tracks, never for slicing one task), **commit to the delegation** (never re-derive a subagent's result), brief precisely on the first spawn (launch → wait → re-brief burns a context cycle), bounded fan-out, and **never delegate your own verification**. The section carries an explicit carve-out: *the Implementer → Reviewer gate is `author ≠ auditor` independence, not self-verification, and is never collapsed for efficiency* — without it, a ceiling rule banning "subagents to double-check yourself" reads as license to drop the Reviewer. Referenced from `/akili-execute` and scaffolded via `/akili-constitution` Step 8B. Leader Instruction #2 now bounds parallelism by the count of genuinely independent tasks in `tasks.md`.
18
+ - **Implementer scope discipline is now bidirectional, with a `Not Done / Assumptions` report field.** Instruction #2 only guarded against scope *creep*; current-generation models also expand scope and can report completion on partial work. It becomes **Scope Discipline (Both Directions)**: don't widen (no unrequested refactors, abstractions, or impossible-case error handling), **don't narrow either** (finish the whole task at the intended scope), **report completion only when actually complete**, and when the task as specified looks wrong, say so in a sentence but still deliver it as written — changing the spec is the Leader's Pivot Protocol call. The completion report gains an optional field 4 (**omitted entirely when the task is clean**) naming what was not delivered and any judgment call made; `leader.md`'s orchestration sequence consumes it — a task carrying an outstanding gap never reaches `[x]`, even on a Reviewer `PASS`.
19
+ - **Model Generation Drift check in `/akili-audit`.** Aliases absorb a new model generation silently, so the registry can be current while everything calibrated around it has gone stale. The new check flags (report-only, as with all drift checks): a **frontier escalation pin** whose recorded reason predates the current `opus` generation; an `Updated:` stamp older than the current generation (effort defaults never swept); and `.agents/*.md` personas missing guardrails the packaged templates have since gained (Delegation Ceiling, bidirectional Scope Discipline).
20
+
21
+ ### Changed
22
+
23
+ - **Effort dial re-baselined as a per-generation calibration, not a constant.** `docs/model-routing.md` now states the split explicitly: the tier mapping survives model churn, **effort defaults do not**. AKILI's defaults (T2 Implementer `medium`) are calibrated for its shape — tasks arrive *already decomposed and spec-bounded* from `/akili-specify` — whereas vendor guidance for a frontier model is written for one open-ended agentic request with no spec, and therefore starts higher (Claude Opus 5's published starting points are `xhigh` for coding/agentic, `high` elsewhere, then sweep down). The reconciliation is three rules: **sweep don't assume** (run a real spec at `medium`/`high`/`xhigh` on every generation and keep the cheapest level whose Reviewer outcome holds), newer generations get *more* out of the low end so the sweep usually confirms `medium`, and an **under-specified task** (a `[~]` resume with thin `execution.md` context, or a post-Pivot retry) starts one level higher because it resembles the open-ended case. New **Opus specifics** paragraph (start high and iterate *down* — more effort up front reduces total turns on agentic work; the rework bump is the cheapest place to spend effort since it only fires after a failure). Mirrored into the `/akili-constitution` Step 8C scaffold as items (e) and (f).
24
+ - **Effort is not a verbosity dial.** New explicit rule in `docs/model-routing.md` and the Step 8C scaffold: on current-generation models, lowering effort does **not** reliably shorten user-facing output — it changes how much the model *thinks*, not how much it *writes*. Long Implementer reports and Leader narration are fixed in the brief (`caveman` for transient agent output, `cognitive-doc-design` for artifacts), never by dropping effort, which otherwise buys verbose output *and* shallower reasoning. The `tier ↔ effort` illustration is now labelled as generation-specific data supporting a generation-stable rule.
25
+ - **Frontier escalation tier corrected and repositioned.** The section claimed Fable carries *"a refusal classifier Opus does not"* — no longer true: current-generation Opus ships elevated bio/cyber safeguards of its own and can decline a request outright. That claim is removed from the Fable-only constraints list (which keeps no-ZDR, minutes-long turns, and less-prescriptive prompting) and replaced by a methodology-level rule: a declined security-adjacent task is a **Pivot Protocol** case recorded in `execution.md`, **never a Reviewer `FAIL` to be reworked** — rework cannot fix a refusal. The tier is also repositioned as a **decision to re-justify on every generation** rather than a standing configuration, since each Opus generation narrows the gap the pin was bought to close: try the alias at `xhigh`/`max` first.
26
+ - **Rate limits documented as per-generation, not per-family.** `docs/model-routing.md` and the `/akili-constitution` Step 8C model-confirmation step now record that a new top-tier generation draws on its **own** quota rather than inheriting the previous generation's pool, so moving T1/T3 onto it neither frees headroom on the old pool nor inherits it. The "reserve the top tier for T1/T3" rule is unchanged — it is about where the budget earns its cost, not which pool it comes from.
27
+
28
+ ### Notes
29
+
30
+ - **The model registry itself required zero edits for the Opus 5 release** — the `opus` alias resolved to it automatically and T1/T3 followed, which is the first row of the *Replacing a model* runbook (*do nothing*). This is recorded in `docs/model-routing.md` as a worked example of the alias-first rule, alongside the corollary this release makes explicit: what a new generation *does* invalidate is the per-task effort calibration and the prompt-level guardrails, not the tier mapping.
13
31
  ## [2.13.1] - 2026-07-24
14
32
 
15
33
  ### Added
@@ -99,6 +99,13 @@ tier (to the deeper reasoner) to preserve independence.
99
99
  This is the single editable source of truth. Phases reference **tiers**; only this table names
100
100
  models. When models change, edit only this table. *Registry updated: 2026-07.*
101
101
 
102
+ > **Worked example — the Opus 5 release required zero edits to this table.** When Anthropic shipped
103
+ > Claude Opus 5, the `opus` alias moved to it on its own; T1/T3 followed automatically. That is the
104
+ > alias-first rule paying off (first row of the *Replacing a model* runbook: **do nothing**). What
105
+ > a new generation *does* require is re-reading the **Effort dial** and the behavioural notes below
106
+ > — the tier mapping survives model churn, but the per-task effort defaults and the prompt-level
107
+ > guardrails do not.
108
+
102
109
  **Alias-first rule: never pin a dated model name where a floating alias exists.** Claude Code's
103
110
  `opus` / `sonnet` / `haiku` aliases always resolve to the latest version of each family — when
104
111
  Anthropic ships a new generation, an alias-based registry needs **zero edits**. Pin a dated ID only
@@ -131,6 +138,13 @@ unchanged (e.g. when the top family moves from one generation to the next, `opus
131
138
  Users on plans that expose a frontier model above Opus (Fable 5 / Mythos 5) can pin it for T1/T3 —
132
139
  see **Frontier escalation tier** below.
133
140
 
141
+ **Rate limits are per-generation, not per-family.** A new Opus generation draws on its **own**
142
+ quota rather than inheriting the previous generation's pool — Opus 5, for instance, does not share
143
+ the combined Opus 4.x bucket. So moving T1/T3 onto a newer Opus neither frees headroom on the old
144
+ pool nor inherits it: check the new generation's limits before shifting volume onto it. The
145
+ "reserve the top tier for T1/T3" rule holds regardless — it is about *where the budget earns its
146
+ cost*, not about which pool the budget comes from.
147
+
134
148
  **OpenCode Go.** The strongest open models anchor the highest-leverage tiers:
135
149
 
136
150
  - **GLM-5.2 → T2 Coder.** A-tier open coder (753B, GPQA 91.2), successor to GLM-5.1, purpose-built
@@ -165,6 +179,13 @@ high-volume fan-out): at ~2× Opus pricing (**$10 / $50** per 1M in/out vs `opus
165
179
  only earns its cost where token volume is low and the decision gates the whole run — T1 (Architect +
166
180
  the execute/test Leader) and T3 (Auditor).
167
181
 
182
+ **Each Opus generation raises the bar for escalating.** Opus 5 closed much of the gap on exactly the
183
+ work this tier existed for — deep reasoning, long-horizon agentic runs, and bug-finding — at half
184
+ Fable's price. Treat escalation as a decision to **re-justify on every generation**, not a standing
185
+ configuration: try the current `opus` at `xhigh` or `max` first, and pin the frontier model only if
186
+ a concrete spec demonstrably fails there. A pin inherited from a previous generation is a pin worth
187
+ re-testing.
188
+
168
189
  | Slot | Default (alias) | Frontier escalation pin | Fallback |
169
190
  |---|---|---|---|
170
191
  | **T1 Architect / Leader** | `opus` | `claude-fable-5` *(pin — record the reason)* | `opus` |
@@ -175,12 +196,18 @@ Two things follow from Fable having **no floating alias**:
175
196
  - **It is a dated pin — record the reason** (per the alias-first rule), and it re-inherits the
176
197
  concrete-slug maintenance the aliases spare you. `/akili-audit`'s drift check watches it; the
177
198
  *Replacing a model* runbook below is how you swap it when a newer frontier model ships.
178
- - **Always give it a Fallback to `opus`**, because Fable carries operational constraints Opus does
179
- not: a **refusal classifier** on bio/cyber that can false-positive on legitimate security-adjacent
180
- specs (a Fable Implementer/Reviewer may decline valid work); **no zero-data-retention** (requires
181
- 30-day retention unavailable to ZDR projects); **minutes-long turns** (budget HITL gates and
182
- progress narration accordingly); and **less prescriptive prompting** (over-prescribed `.agents/*`
183
- personas can *reduce* Fable output quality — loosen them when routing a role to Fable).
199
+ - **Always give it a Fallback to `opus`**, because Fable carries operational constraints the Opus
200
+ alias does not: **no zero-data-retention** (requires 30-day retention unavailable to ZDR
201
+ projects); **minutes-long turns** (budget HITL gates and progress narration accordingly); and
202
+ **less prescriptive prompting** (over-prescribed `.agents/*` personas can *reduce* Fable output
203
+ quality loosen them when routing a role to Fable).
204
+
205
+ **Refusal classifiers are no longer a Fable-only concern.** Current-generation Opus ships elevated
206
+ bio/cyber safeguards of its own and can decline a request outright, so a security-adjacent spec may
207
+ stall the Implementer or Reviewer on **either** tier. Plan for it at the methodology level rather
208
+ than treating it as an escalation-only risk: when a spec is security-adjacent, say so in the task
209
+ brief, and treat a declined task as a **Pivot Protocol** case (record it in `execution.md` and
210
+ escalate to the user) — never as a Reviewer `FAIL` to be reworked, since rework cannot fix a refusal.
184
211
 
185
212
  **Cross-family author ≠ auditor.** Fable pairs well with a *different-lab* auditor — pin the
186
213
  Implementer/Leader on `claude-fable-5` and the Reviewer on a Kimi/DeepSeek OpenCode slug (or
@@ -227,21 +254,53 @@ representative model (GPT-5.6 Sol, Artificial Analysis Intelligence Index):
227
254
  | T3 **Reviewer** / `/akili-validate` | `high` (auditor thoroughness) |
228
255
  | T5 `/akili-archive` / setup steps | `low` |
229
256
 
257
+ **Re-baseline these defaults on every model generation — the tier mapping survives model churn, the
258
+ effort defaults do not.** The table above is calibrated for AKILI's shape: tasks arrive *already
259
+ decomposed and spec-bounded* from `/akili-specify`, which is exactly the case where a mid-range
260
+ effort holds up. Vendor guidance for a frontier model is written for the opposite case — one
261
+ open-ended agentic request with no spec — and therefore starts higher (for Claude Opus 5, the
262
+ published starting points are **`xhigh` for coding and agentic work, `high` elsewhere, then sweep
263
+ down**). Both are right for their context. The reconciliation:
264
+
265
+ - **Sweep, don't assume.** On a new generation, run the same spec at `medium`, `high`, and `xhigh`
266
+ and keep the cheapest level whose Reviewer outcome holds. Effort defaults inherited from a
267
+ previous generation are a guess, not a measurement.
268
+ - **Newer generations get more out of the low end, not less.** Each Opus generation has made
269
+ `low`/`medium` stronger relative to its own ceiling — so the sweep usually confirms the T2
270
+ `medium` default rather than pushing it up. Raise the Implementer to `xhigh` for the task
271
+ signals in the policy table above, not by default.
272
+ - **Where a task arrives under-specified** — a `[~]` resume with thin `execution.md` context, or a
273
+ Pivot Protocol retry — it is closer to the vendor's open-ended case. Start it at `high`/`xhigh`.
274
+
275
+ **Effort is not a verbosity dial.** On current-generation models, lowering effort does **not**
276
+ reliably shorten user-facing output — it changes how much the model *thinks*, not how much it
277
+ *writes*. If an Implementer's report or a Leader's narration is too long, fix it in the brief
278
+ (`caveman` for transient agent output, `cognitive-doc-design` for artifacts), never by dropping
279
+ effort — that buys verbose output *and* shallower reasoning. The same applies in reverse: raising
280
+ effort to get a more thorough **document** is the wrong lever; ask for the depth explicitly.
281
+
230
282
  **Escalate effort on rework.** In the `/akili-execute` rework loop, a Reviewer `FAIL` bumps effort
231
283
  one level on the retry (attempt 1 `medium` → attempt 2 `high` → attempt 3 `xhigh`) — cheap (only
232
284
  when it already failed) and it targets the usual cause (under-thinking, not missing instructions).
233
285
 
234
286
  **Tier ↔ effort interaction — don't `max` a cheaper tier.** Maxing a lower tier erodes its cost
235
- advantage: Sonnet 5 at `max` (53 intel / $1.53) approaches Opus 4.8 at `max` (56 intel / $1.80) —
236
- near-Opus price for below-Opus intelligence. If you find yourself wanting a cheaper tier at `max`,
237
- escalate the **tier** (to `opus` at `high`/`medium`) instead of the effort.
287
+ advantage: on the 4.8-generation ladder, Sonnet 5 at `max` (53 intel / $1.53) approached Opus 4.8 at
288
+ `max` (56 intel / $1.80) — near-Opus price for below-Opus intelligence. Those figures are a
289
+ generation-specific illustration, but the rule they demonstrate is stable across generations: if you
290
+ find yourself wanting a cheaper tier at `max`, escalate the **tier** (to `opus` at `high`/`medium`)
291
+ instead of the effort.
238
292
 
239
293
  **Sonnet specifics.** Sonnet respects effort strictly, especially at the low end — at `low`/`medium`
240
294
  it scopes work to exactly what was asked. If you see shallow reasoning on a hard problem, **raise
241
295
  the effort, don't prompt around it**. `high` is the default sweet spot; give `max_tokens` headroom
242
296
  at `xhigh`/`max` (thinking consumes the budget — too tight truncates with `stop_reason: max_tokens`).
243
- On Opus/frontier the nuance inverts slightly: start at `high` and iterate — more effort up front
244
- often *reduces* total turns and cost on agentic work.
297
+
298
+ **Opus specifics.** On Opus the nuance inverts: start high and iterate **down** more effort up
299
+ front often *reduces* total turns and total cost on agentic work, because the model plans better and
300
+ re-does less. Two consequences for the AKILI loop: (1) the **rework bump** (above) is the cheapest
301
+ place to spend effort, since it only fires after a failure; and (2) `max` is for the
302
+ correctness-critical and latency-insensitive case only — it can overthink a routine task and is
303
+ where diminishing returns bite hardest.
245
304
 
246
305
  ## Enforced routing (tool-native agent bindings)
247
306
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "akili-specs",
3
- "version": "2.13.1",
3
+ "version": "2.14.0",
4
4
  "description": "Portable AKILI-SPECS methodology commands and skills for AI-assisted development.",
5
5
  "homepage": "https://github.com/JuankCadavid/akili-specs#readme",
6
6
  "repository": {