@orkestrel/scaffold 0.0.76 → 0.0.77

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -24,14 +24,14 @@ follows it. This file adds only what Claude Code does differently, and cannot we
24
24
  - Every Workflow `agent()` node names its model alias explicitly. The Workflow custom-agent path
25
25
  does not apply a role file's `model:` pin, so a node that omits the alias runs on the session
26
26
  model and the lane reads normal on the wrong engine.
27
- - Run the main session on `opus` at high effort, set by `/model opus` or `"model": "opus"`. Opus 5
27
+ - Run the main session on `opus` at high effort, set by `/model opus` or `"model": "opus"`. Opus 5.5
28
28
  is the Orchestrator in this harness. Its Orchestrator duties are unchanged if it is configured
29
29
  otherwise.
30
- - The Orchestrator shares its engine with `planner`, `reviewer`, and `opus`. Run the Sol `analyst`
30
+ - The Orchestrator shares its engine with `planner`, `reviewer`, and `opus`. Run the Astra `analyst`
31
31
  in every design round and every audit round so the judgment is not single-engine, and confirm from
32
- its journal that it reached Sol. A bridge driver that answers from its own engine
32
+ its journal that it reached Astra. A bridge driver that answers from its own engine
33
33
  collapses the round to one engine and its ruling still reads as normal.
34
- - Claude role frontmatter accepts Claude models only. Reach Grok through `grok`, and Sol through
34
+ - Claude role frontmatter accepts Claude models only. Reach Grok through `grok`, and Astra through
35
35
  `analyst` and `sol`. Never put an external model in `model:`. Treat
36
36
  `.agents/transports/codex.md` as the transport contract for those bridges, never as a
37
37
  dispatchable role.
@@ -39,11 +39,11 @@ follows it. This file adds only what Claude Code does differently, and cannot we
39
39
 
40
40
  ## Bench wiring
41
41
 
42
- - `.mcp.json` registers `codex mcp-server` for short interactive exchanges with Sol. Project MCP
42
+ - `.mcp.json` registers `codex mcp-server` for short interactive exchanges with Astra. Project MCP
43
43
  servers are enabled without prompting, so the wiring works headless.
44
44
  - `.claude/skills/<name>/SKILL.md` is a bridge that loads the canonical skill from
45
45
  `.agents/skills/<name>/SKILL.md`. It adds no independent process.
46
- - Claude Code exposes `claude mcp serve`, which is how a Codex-primary session reaches Opus 5.
46
+ - Claude Code exposes `claude mcp serve`, which is how a Codex-primary session reaches Opus 5.5.
47
47
 
48
48
  ## Claude Code Cloud
49
49
 
@@ -29,11 +29,11 @@ One workflow runs across all providers. Each engine has one job and never takes
29
29
  | Engine | Job | Posture |
30
30
  | --------------- | ----------------------------------------------------- | -------------------------------------------- |
31
31
  | **Cursor Grok** | Absorption, distillation, scouting, bounded research | Read-only; returns evidence, never decisions |
32
- | **Opus 5** | Subjective design, design-fit review, implementation | Proposes, audits, implements; never accepts |
33
- | **GPT-5.6 Sol** | Objective analysis, correctness audit, implementation | Proposes, audits, implements; never accepts |
32
+ | **Opus 5.5** | Subjective design, design-fit review, implementation | Proposes, audits, implements; never accepts |
33
+ | **GPT-6 Astra** | Objective analysis, correctness audit, implementation | Proposes, audits, implements; never accepts |
34
34
 
35
- - Route each nontrivial implementation unit to Opus or Sol. Objective, constraint-heavy,
36
- mechanical-precision work goes to Sol. API-shape, naming, and documentation-voice work goes to
35
+ - Route each nontrivial implementation unit to Opus or Astra. Objective, constraint-heavy,
36
+ mechanical-precision work goes to Astra. API-shape, naming, and documentation-voice work goes to
37
37
  Opus. Cursor Composer is not an implementation route, and no `composer` role exists.
38
38
  - Design runs the adversarial pass. § Execution loop's audit step fixes which lanes an audit runs.
39
39
 
@@ -44,12 +44,12 @@ reasoning effort.
44
44
 
45
45
  | Harness | Orchestrator engine |
46
46
  | ----------- | ------------------- |
47
- | Claude Code | Opus 5 |
48
- | Codex | GPT-5.6 Sol |
47
+ | Claude Code | Opus 5.5 |
48
+ | Codex | GPT-6 Astra |
49
49
  | Cursor | Cursor Grok |
50
50
 
51
51
  - The Orchestrator reconciles. No engine reconciles itself or accepts its own work.
52
- - The Orchestrator shares its engine with one lane: Opus in Claude Code, Sol in Codex. That lane is
52
+ - The Orchestrator shares its engine with one lane: Opus in Claude Code, Astra in Codex. That lane is
53
53
  still dispatched as a separate subagent with a clean context, never run inline.
54
54
  - In a fix round the auditor is an engine that did not write it. When the writer's engine is the
55
55
  Orchestrator's engine, the auditor is the other lane.
@@ -84,12 +84,12 @@ names nothing else, so never write it of a lane. A verdict file's recorded reaso
84
84
 
85
85
  ### Engine assignment
86
86
 
87
- By default Opus 5 holds the subjective lane and Sol holds the objective lane.
87
+ By default Opus 5.5 holds the subjective lane and Astra holds the objective lane.
88
88
 
89
89
  Swap the lanes whenever the round needs an engine that is not the one running that lane. Bench
90
90
  darkness is one trigger and the writer's engine is another: § Execution loop's audit step requires
91
- an auditor that did not write the work, so where Sol wrote the work under audit, give the objective
92
- lane to Opus 5 and the subjective lane to Sol, and reverse that where Opus 5 wrote it. Both engines
91
+ an auditor that did not write the work, so where Astra wrote the work under audit, give the objective
92
+ lane to Opus 5.5 and the subjective lane to Astra, and reverse that where Opus 5.5 wrote it. Both engines
93
93
  still run, so this is a lane swap rather than a substitution. Record which engine held which lane in
94
94
  the routing ledger.
95
95
 
@@ -97,15 +97,15 @@ When one engine is unavailable, the remaining engine runs **every** lane — sti
97
97
  subagents, still clean contexts, still blind to each other, each told which perspective it holds.
98
98
  Record the substitution.
99
99
 
100
- | Harness | Engine unavailable | Runs every lane |
101
- | ----------- | --------------------------------- | --------------- |
102
- | Claude Code | Sol (Codex bench dark) | Opus 5 |
103
- | Codex | Opus 5 (Claude CLI dark) | GPT-5.6 Sol |
104
- | Cursor | Opus 5 and Sol (MCP servers dark) | Cursor Grok |
100
+ | Harness | Engine unavailable | Runs every lane |
101
+ | ----------- | ------------------------------------- | --------------- |
102
+ | Claude Code | Astra (Codex bench dark) | Opus 5.5 |
103
+ | Codex | Opus 5.5 (Claude CLI dark) | GPT-6 Astra |
104
+ | Cursor | Opus 5.5 and Astra (MCP servers dark) | Cursor Grok |
105
105
 
106
106
  - Never assign Grok to either lane in Claude Code or Codex. If the remaining native engine is also
107
107
  unavailable there, the pass cannot run: stop and report rather than substituting Grok.
108
- - Grok takes every lane only in Cursor, and only when Opus 5 and Sol are both unavailable.
108
+ - Grok takes every lane only in Cursor, and only when Opus 5.5 and Astra are both unavailable.
109
109
  - Treat a lane that returns no verdicts as a lane that did not run. Re-probe the bench before ruling
110
110
  on why, per Bench laws rule "One lane at a time per bench". A bench lane reporting that its driver
111
111
  executed and its engine was never reached is a dark bench, not a result. Record the bench dark
@@ -133,7 +133,7 @@ Fall back in this order and record the substitution:
133
133
  reading and `researcher` excludes repository-scale absorption, so neither takes that step.
134
134
  - Never route absorption to the Orchestrator itself, even when the Orchestrator is Grok. Keep the
135
135
  main context at decision level; in Cursor that means a Grok executor session, not this one.
136
- - Never spend Opus 5 or Sol on it.
136
+ - Never spend Opus 5.5 or Astra on it.
137
137
  - Grok is read-only, so a writing unit never routes there. Fully specified mechanical writing goes
138
138
  to `builder` or `application` on the harness's cheap native tier.
139
139
  - `verifier` runs commands and reports exit codes, so it stays on the native tier too.
@@ -162,11 +162,11 @@ when the role file already pins it.
162
162
  | Job | Claude role (`.claude/agents/`) | Codex role (`.codex/agents/`) | Engine |
163
163
  | ---------------------------------------- | ------------------------------- | ----------------------------- | ----------------------------- |
164
164
  | Absorption, distillation, scouting | `grok` | `grok` | Cursor Grok (bridge) |
165
- | Creative design and alternatives | `planner` | `planner` | Opus 5 (native / bridge) |
166
- | Design-fit review and audit | `reviewer` | `reviewer` | Opus 5 (native / bridge) |
167
- | Objective analysis and correctness audit | `analyst` | `analyst` | GPT-5.6 Sol (bridge / native) |
168
- | Nontrivial implementation (objective) | `sol` | `sol` | GPT-5.6 Sol (bridge / native) |
169
- | Nontrivial implementation (subjective) | `opus` | `opus` | Opus 5 (native / bridge) |
165
+ | Creative design and alternatives | `planner` | `planner` | Opus 5.5 (native / bridge) |
166
+ | Design-fit review and audit | `reviewer` | `reviewer` | Opus 5.5 (native / bridge) |
167
+ | Objective analysis and correctness audit | `analyst` | `analyst` | GPT-6 Astra (bridge / native) |
168
+ | Nontrivial implementation (objective) | `sol` | `sol` | GPT-6 Astra (bridge / native) |
169
+ | Nontrivial implementation (subjective) | `opus` | `opus` | Opus 5.5 (native / bridge) |
170
170
  | Bulk reading and evidence distillation | `distiller` | `distiller` | Grok → Luna → Sonnet |
171
171
  | Bounded primary-source research | `researcher` | `researcher` | Grok → Luna → Sonnet |
172
172
  | Repository reconnaissance | `scout` | `scout` | Grok → Luna → Sonnet |
@@ -197,18 +197,18 @@ when the role file already pins it.
197
197
  job to a bench means shipping that catalog across, which costs more than the bench saves.
198
198
  - A transport contract lives in `.agents/transports/`, not in an agents directory. A harness lists
199
199
  its dispatchable agents from that directory, so a contract that is never dispatched sits outside
200
- it. `.agents/transports/codex.md` is the shared Sol transport contract and
200
+ it. `.agents/transports/codex.md` is the shared Astra transport contract and
201
201
  `.agents/transports/claude.md` the shared Opus transport contract. Neither is a route: `analyst`
202
- and `sol` are the named Sol bridges, `planner`, `reviewer`, and `opus` the named Opus bridges, and
202
+ and `sol` are the named Astra bridges, `planner`, `reviewer`, and `opus` the named Opus bridges, and
203
203
  each binds its own contract by reference and pins only its route and sandbox.
204
204
  `.agents/transports/cursor.md` is the shared Cursor transport contract, and both harnesses' `grok`
205
205
  bridges bind it, because Cursor is native to neither. A contract's home is the provider it carries,
206
206
  never the harness that reaches it.
207
207
  - Mirroring is by work class, not filename. A transport contract is provider-specific: the Codex
208
- contract carries the Sol transport the Claude-side bridges follow, the Claude contract carries the
208
+ contract carries the Astra transport the Claude-side bridges follow, the Claude contract carries the
209
209
  Opus transport the Codex-side bridges follow, and each bridge binds the contract of the provider it
210
210
  reaches.
211
- - Opus and Sol roles use high effort. Native cheap-tier roles use low or medium. Bridge drivers use
211
+ - Opus and Astra roles use high effort. Native cheap-tier roles use low or medium. Bridge drivers use
212
212
  the cheapest tier that can run a CLI.
213
213
  - Never route orchestration or acceptance across a bridge.
214
214
 
@@ -345,7 +345,7 @@ longer holds.
345
345
  units, dependencies, ownership, parallel and serial order, acceptance criteria, risks.
346
346
  - Surface the plan before dispatch, including a routing ledger naming each unit's role **and**
347
347
  engine. Routing a unit to a Claude-native agent when its work class belongs to a bench —
348
- reading-heavy to Grok, objective audit or objective implementation to Sol — without a recorded
348
+ reading-heavy to Grok, objective audit or objective implementation to Astra — without a recorded
349
349
  bench-dark deviation is a dispatch deviation.
350
350
  - State the goal's exit criterion beside the units: the enumerated capabilities whose closure
351
351
  ends the campaign, each to end implemented, repaired, retained, or intentionally excluded on
@@ -369,11 +369,23 @@ longer holds.
369
369
  - State the audit's subject as numbered falsifiable claims and require per-claim verdicts with
370
370
  evidence, per the Falsification law in `.claude/rules/quality.md` and the `orkestrel-falsify`
371
371
  value set, unless the dispatch names a different skill that fixes another.
372
+ - Require every lane, before confirming a claim **about a proof**, to name the mutation that would
373
+ make that proof fail and to say whether its assertions distinguish that mutation from the
374
+ passing case. Put the answer in the evidence. Describing what a proof does is not ruling on it:
375
+ a proof that cannot fail reads exactly like one that can, and structural reading confirms both.
376
+ This is the discipline that separates a lane which runs the disputed thing from one which reads
377
+ it, and the reading lane is the one that confirms a hole.
372
378
  - In a fix round, give the unit to an auditor engine that did not write it.
373
379
  - Run the `orkestrel-falsify` skill for multi-round audits. It owns the brief anatomy, the
374
380
  successor-brief rule, the verdict shape and its single terminal line, and the reconciliation
375
381
  discipline.
376
382
  - Reconcile the lanes that ran. Drop, on the record, any finding no lane can substantiate.
383
+ - Where two lanes contradict each other, check first that each one's citations resolve in the
384
+ file it names, before weighing the arguments. A line number past a file's end settles the
385
+ disagreement in one command, and a lane that confirmed a claim on evidence that cannot exist
386
+ has its verdict on that claim discarded rather than balanced against the other lane's. Discard
387
+ the claim, not the lane: verify a sample of that lane's other citations, and keep the readings
388
+ that hold.
377
389
  6. **Verify.** Have one independent `verifier` run the authoritative gates.
378
390
  7. **Re-baseline.** Reconcile the remaining plan against what the phase revealed, before dispatching
379
391
  the next one.
@@ -461,6 +473,14 @@ The harness bridge names the concrete mechanism for each of these.
461
473
  - Name a corrected unit's effective brief and report and the pair they supersede before that unit
462
474
  integrates. A unit whose correction landed as a serial patch or a direct reconciliation, with no
463
475
  successor pair on disk, cannot be re-run from what the campaign kept.
476
+ - When you copy a previous round's or unit's launch artifact to derive this one's, rewrite every
477
+ field that names the subject, not only the ones the edit was about. That covers each brief's lane
478
+ focus, its evidence and report paths, its subject files, and the path it gives for the law; the
479
+ claims file's subject line; each driver and watch script's header, journal, and output paths; and
480
+ the workflow's `meta.description` and every node's label. Read the derived copy start to finish
481
+ against the round it launches before launching it. A field carried over from the source names a
482
+ different subject and still reads as deliberate, so the lane works around it or rules on the wrong
483
+ file, and the retained record attributes that to the round.
464
484
  - Write an audit round's numbered claims to `tmp/audit/<unit>-audit-claims.md` and point every lane
465
485
  of the round at that one file, so "both lanes ran the same brief" stays checkable after the round.
466
486
  Retain it as `.orkestrel/<package>/<unit>-audit-claims.md`, beside the round's verdict.
@@ -475,6 +495,11 @@ The harness bridge names the concrete mechanism for each of these.
475
495
  ignores rather than into the checkout root.
476
496
  - Before dispatching a successor or accepting a round, open every file the effective brief names
477
497
  and confirm it resolves from the executor's root. Refuse the transition when one does not.
498
+ - Check the launch argument's own file list the same way. A driver prompt names records beside the
499
+ brief it points at, and those names are part of the effective instruction even though they sit
500
+ outside the brief file. Staging the brief and its terrain is not enough: list every path the prompt
501
+ mentions and confirm each one resolves, because the unit stops on the one that does not and the
502
+ stop is charged to the round rather than to the launch.
478
503
  - Send a decision taken mid-campaign to every unit already in flight whose brief it invalidates. An
479
504
  executor cannot see a change made after it was dispatched, so it writes the state its brief
480
505
  described and the defect surfaces as its own.
@@ -575,6 +600,23 @@ filled.
575
600
  - Paste the command and its output behind every factual claim — paths, counts, registrations, file
576
601
  existence. Name the scope any search covered, and check a fact against the code rather than
577
602
  against another artifact that states it.
603
+ - Search the campaign's own record for every fact a brief asserts, and reconcile any disagreement
604
+ before dispatch. A brief that contradicts a measurement the campaign already took is the most
605
+ expensive kind of wrong fact, because the scope read cannot save you: it checks the brief against
606
+ the tree with the brief's framing in hand, so it reproduces the error as often as it catches it.
607
+ The unit reading the record cold is the only reader positioned to refuse, and by then the round is
608
+ spent. Grep the plan and the retained reports for the subject, and not the tree alone.
609
+ - Better, give each measurement one home and keep the brief out of it. Put the measurements in a
610
+ terrain record the brief names and stages beside itself, and write the brief as rulings and
611
+ obligations that point at it. A measurement restated in a brief is a second copy that can drift
612
+ from the first, and the drift is invisible because both artifacts look authoritative. Tell the unit
613
+ which artifact wins when they disagree, and to stop rather than resolve it.
614
+ - Cite a site by its symbol in any artifact meant to outlive a landing, and name a line only as
615
+ approximate ("around line 40"). A record written before one unit lands and read after another does is the
616
+ normal case, not the exception: a line citation that was exact when measured goes stale the moment
617
+ something is inserted above it, and the reader cannot tell a stale number from a wrong one. Name
618
+ the function, the mixin, the case title, or the surrounding construct, and tell the reader to
619
+ locate it by that.
578
620
  - Take every measurement under the conditions the unit runs in, or have the unit take it before
579
621
  doing anything else.
580
622
  - Ask what the change does to every fact you measured, and fix each criterion to the state the unit
@@ -584,8 +626,32 @@ filled.
584
626
  golden digest over generated output, the consumer script naming the removed union member.
585
627
  - Derive that set by running the suite. Where you cannot run it, name the search's bound in the
586
628
  brief so the unit re-derives the set.
629
+ - Find the enumerating assertions by searching for the population's EXISTING members, never by
630
+ reasoning about which files look relevant. A literal set or list naming every member of a growing
631
+ population goes false the moment a unit adds one, and it can sit in a file whose name suggests it
632
+ holds only machinery. Grep the tree for a member the population already has; every file that comes
633
+ back is a file the unit must own. Reasoning finds the obvious site and misses the rest, and the
634
+ miss surfaces as the unit's stop.
635
+ - Answer any question of the form "can this tree do X" the same way: search for X already being done,
636
+ never by inspecting the interface that would do it. Grep for the behaviour, the export, the staged
637
+ preference, the driven surface. A sibling already doing it is the answer, and its call site is the
638
+ pattern the unit copies. An interface's surface is evidence about that interface and never about the
639
+ tree, so a missing control there proves nothing — the mechanism is as often a function an installed
640
+ package exports as a method on the object in hand. Record the answer as unreachable only after a
641
+ search for the thing itself came back empty, and name the pattern that search used. A brief that
642
+ asserts a capability is absent tells the unit to stop looking, so this error costs the round even
643
+ when the tree is correct.
644
+ - Diff the previous unit's actual status against this brief's owned set whenever a series of units
645
+ adds members to one population. A grant the previous sibling needed and this brief omits is the
646
+ likeliest gap, because each brief is written from the plan rather than from what the last unit
647
+ touched.
587
648
  - Grant a behaviour with the tests that pin it, a constant with every fixture and expectation
588
649
  derived from it, and a template change with the materialized copy the package generates from it.
650
+ - Grant the file a brief tells the unit to copy a pattern from, wherever a standing gate forbids
651
+ duplicating that pattern. Naming a pattern to follow and withholding the file it lives in
652
+ instructs the unit into a gate failure it cannot fix inside its scope: the honest fix touches both
653
+ copies, and extracting one side alone breaks the same rule from the other direction. The unit then
654
+ stops, correctly, and the round is spent on a contradiction the brief carried.
589
655
  - Read each criterion against the off-limits list, line by line. Grant the file a criterion needs or
590
656
  strike that criterion. A file the change will break that appears in neither list is unscoped.
591
657
  - Where a scope line names the `tests/**` or `src/**` glob, name the paths the `scaffold repair`
@@ -614,6 +680,11 @@ filled.
614
680
  After reconciling findings into briefs, walk the retained finding list once. Every finding names
615
681
  the brief item that carries it. A finding with no carrier is a dropped finding.
616
682
 
683
+ Name a carrier by the unit, never by a condition. "Whichever unit next touches this file" is a
684
+ description that stops being true the moment the next unit is chosen, and nobody re-reads the verdict
685
+ at that moment — so the finding reads as assigned and is not. Where the carrying unit is not yet
686
+ known, say that plainly and re-check the assignment at the next dispatch.
687
+
617
688
  Every finding names exactly one carrier. A further brief claiming the same finding is not redundancy
618
689
  that costs a little duplicated work — it is a conflict the executor discovers mid-unit, between
619
690
  documents you told it to obey, with no way to tell which you meant. It will either implement the row
@@ -927,8 +998,8 @@ dependency takes the same shape when its consumers' gates read its unpublished t
927
998
  ## Acceptance laws
928
999
 
929
1000
  - No writer's and no external engine's self-assessment is authoritative.
930
- - Never spend Opus 5 or Sol on absorption, distillation, scouting, or mechanical edits. Never route
931
- judgment-bearing implementation away from Opus 5 or Sol.
1001
+ - Never spend Opus 5.5 or Astra on absorption, distillation, scouting, or mechanical edits. Never route
1002
+ judgment-bearing implementation away from Opus 5.5 or Astra.
932
1003
  - Substitute an engine only when the same session records the bench dark — CLI missing, auth
933
1004
  expired, model unavailable. Name the fallback in the plan; never improvise it silently. The
934
1005
  tedious-work ladder is the only pre-approved substitution, and each step down it is still recorded.
@@ -1,13 +1,13 @@
1
1
  # Claude transport contract
2
2
 
3
3
  The transport contract every Codex-side driver follows when it carries a brief to the
4
- Claude Opus 5 bench: invocation, journalling, session ids, availability, and recovery.
4
+ Claude Opus 5.5 bench: invocation, journalling, session ids, availability, and recovery.
5
5
  Reach a route by its own name — `planner`, `reviewer`, `opus`. This file is a contract,
6
6
  not a role: it is never dispatched, and the drivers that bind it pin their own model,
7
7
  effort, and sandbox mode.
8
8
 
9
9
  Read `.agents/orchestration.md` first. It owns the role set, the routing, and the
10
- dispatch contract. You dispatch the external Claude Opus 5 bench. Every Codex-side
10
+ dispatch contract. You dispatch the external Claude Opus 5.5 bench. Every Codex-side
11
11
  driver — planner, reviewer, opus — binds this contract by reference and pins only its
12
12
  own route, permission mode, and brief shape.
13
13
 
@@ -20,7 +20,7 @@ claude -p "<brief or pointer>" --model opus --effort high
20
20
  with the permission mode the route pins. Never substitute a fixed Claude model id.
21
21
 
22
22
  Verify that the `claude` CLI resolves and is authenticated before first use. On either
23
- failure return it immediately with the fallback named, so the Sol main session records
23
+ failure return it immediately with the fallback named, so the Astra main session records
24
24
  Opus unavailable for the round. Where the binary is absent, that report names the install
25
25
  command for the `claude` CLI, so the Orchestrator can put it to the user in the same turn
26
26
  it records the bench dark and re-probe when the user answers. Never install, authenticate,
@@ -1,26 +1,26 @@
1
1
  # Codex transport contract
2
2
 
3
3
  The transport contract every Claude-side driver follows when it carries a brief to the
4
- GPT-5.6 Sol bench: work class to transport, the exact exec form, journalling, session ids,
4
+ GPT-6 Astra bench: work class to transport, the exact exec form, journalling, session ids,
5
5
  and recovery. Reach a route by its own name — `analyst` for audit, `sol` for
6
6
  implementation. This file is a contract, not a role: it is never dispatched, and the
7
7
  drivers that bind it pin their own tools, model, effort, and permission mode.
8
8
 
9
- You dispatch the external Codex Sol bench.
9
+ You dispatch the external Codex Astra bench.
10
10
 
11
11
  Read `.agents/orchestration.md` first. It owns the role set, the routing, and the dispatch
12
12
  contract.
13
13
 
14
14
  The dispatch names exactly one route and includes the objective, evidence slice, rules,
15
15
  skill, guide or spec, scope, output contract, and acceptance criteria. Spawn no Claude
16
- agent, never implement directly, and never treat Sol's response as authoritative.
16
+ agent, never implement directly, and never treat Astra's response as authoritative.
17
17
 
18
18
  ## Models and effort
19
19
 
20
20
  ```text
21
- CODEX_ANALYST_MODEL=gpt-5.6-sol
21
+ CODEX_ANALYST_MODEL=gpt-6-astra
22
22
  CODEX_ANALYST_EFFORT=high
23
- CODEX_IMPLEMENTER_MODEL=gpt-5.6-sol
23
+ CODEX_IMPLEMENTER_MODEL=gpt-6-astra
24
24
  CODEX_IMPLEMENTER_EFFORT=high
25
25
  ```
26
26
 
@@ -52,7 +52,7 @@ background command under a hard cap.
52
52
  Create `tmp/codex/`, then write the full brief to `tmp/codex/<unit>-brief.md`. Briefs
53
53
  never travel as shell arguments. Return the exact resolved command with a pointer prompt:
54
54
 
55
- `timeout <cap> codex exec --json -C <working-directory> --sandbox <route-sandbox> --model gpt-5.6-sol -c "model_reasoning_effort=\"high\"" --output-last-message tmp/codex/<unit>-last.md "Read and execute the brief at tmp/codex/<unit>-brief.md exactly. Your final message must be the report it specifies." < /dev/null > tmp/codex/<unit>.jsonl`
55
+ `timeout <cap> codex exec --json -C <working-directory> --sandbox <route-sandbox> --model gpt-6-astra -c "model_reasoning_effort=\"high\"" --output-last-message tmp/codex/<unit>-last.md "Read and execute the brief at tmp/codex/<unit>-brief.md exactly. Your final message must be the report it specifies." < /dev/null > tmp/codex/<unit>.jsonl`
56
56
 
57
57
  - Return the brief path, that resolved command, and the journal path. Leave
58
58
  `<cap>` unresolved — the Orchestrator owns it, per **Long-running commands → Launching**
@@ -64,7 +64,7 @@ never travel as shell arguments. Return the exact resolved command with a pointe
64
64
  repository, and `--output-schema <file>` when the Orchestrator supplies one.
65
65
  - The journal at `tmp/codex/<unit>.jsonl` is the live progress record and its mtime is
66
66
  the liveness signal the Orchestrator watches. Never re-print the stream into your report.
67
- - The Orchestrator reads Sol's answer from the `--output-last-message` file rather than
67
+ - The Orchestrator reads Astra's answer from the `--output-last-message` file rather than
68
68
  stdout, and records the session id (`thread_id` in the journal's opening events)
69
69
  beside the result; a follow-up on a finished exec is a fresh dispatch.
70
70
 
@@ -73,7 +73,7 @@ never travel as shell arguments. Return the exact resolved command with a pointe
73
73
  `codex exec` runs with `--unshare-net`. Any unit needing the registry or another remote
74
74
  endpoint — lockfile generation, real installs, live fetches — belongs to the
75
75
  Orchestrator's own tracked commands or a network-capable native agent. Never put it in a
76
- brief. A Sol exec hanging on `npm` until its cap fires is this misroute, not a slow bench.
76
+ brief. A Astra exec hanging on `npm` until its cap fires is this misroute, not a slow bench.
77
77
 
78
78
  The namespace has its own loopback, so a host daemon on `127.0.0.1` is unreachable and a bind can
79
79
  fail `EPERM`. It has no IPv6, so `::1` fails `EAFNOSUPPORT`. Any proof that must reach a daemon,
@@ -93,7 +93,7 @@ mutated restores it by rewriting the original text, and proves it with
93
93
 
94
94
  On any interruption or missing result, in order:
95
95
 
96
- 1. Interrupted MCP call with a persisted thread id → `mcp__codex__codex-reply` asking Sol
96
+ 1. Interrupted MCP call with a persisted thread id → `mcp__codex__codex-reply` asking Astra
97
97
  to re-emit the complete final report. The reasoning may have finished server-side.
98
98
  2. No persisted id, or the reply fails → prepare a fresh journaled CLI launch with the
99
99
  same brief file and return it.
@@ -112,12 +112,12 @@ correctness and security audit, and constraint review. Capture repository status
112
112
  and after. Require evidence for every claim and return unsupported claims as dropped.
113
113
 
114
114
  An audit brief states its subject as a numbered list of falsifiable claims rather than a
115
- diff to read, and requires Sol to attempt refutation. The Falsification section of
115
+ diff to read, and requires Astra to attempt refutation. The Falsification section of
116
116
  `.claude/rules/quality.md` owns the method and the evidence each verdict carries. The verdict shape
117
117
  defaults to `orkestrel-falsify`; a dispatch may name a different skill that fixes another. That
118
118
  skill owns the value set and the terminal line. Point the brief at both; restate neither.
119
119
 
120
- ## Sol route
120
+ ## Astra route
121
121
 
122
122
  Sandbox `workspace-write`, the checkout the route writes in, its sole serial writer from a clean committed
123
123
  baseline, with owned files, off-limits files, and a deviation contract. The brief forbids
@@ -147,7 +147,7 @@ on one unit, at the same point in the work, with nothing written to disk either
147
147
  Route such a unit to `opus` from the start and record the Codex
148
148
  bench dark for that unit with this reason. Do not soften or obscure a brief to slip past
149
149
  the filter; a bench that declines work is a routing fact, not an obstacle. The exclusion is
150
- per unit — everything else still routes to Sol, and an audit that merely reads existing
150
+ per unit — everything else still routes to Astra, and an audit that merely reads existing
151
151
  negative tests is unaffected.
152
152
 
153
153
  ## Availability
@@ -162,7 +162,7 @@ negative tests is unaffected.
162
162
  output captured to `tmp/codex/login.log`, surfaces the verification URL and one-time code
163
163
  from that file, and re-probes `codex login status` on completion.
164
164
  - Recovery impossible — device login unavailable, declined, or expired: the Codex bench is
165
- dark. Name the fallback explicitly: `planner` and `reviewer` (Opus 5) for judgment, and
165
+ dark. Name the fallback explicitly: `planner` and `reviewer` (Opus 5.5) for judgment, and
166
166
  `builder` for fully specified mechanics.
167
167
  - Never authenticate, log out, inspect auth files, or substitute an API key, access token,
168
168
  or copied `auth.json`.
@@ -13,7 +13,7 @@ bridges bind this file.
13
13
  ## Model
14
14
 
15
15
  ```text
16
- CURSOR_GROK_MODEL=cursor-grok-4.6-high
16
+ CURSOR_GROK_MODEL=cursor-grok-4.7-high
17
17
  ```
18
18
 
19
19
  That id was read from `agent models` on 2026-08-13. Resolve the model from the variable at
@@ -1,14 +1,14 @@
1
1
  ---
2
2
  name: analyst
3
- description: 'Claude-side driver for the GPT-5.6 Sol `analyst` route — the adversarial objective design argument, diagnosis, and correctness and constraint audit. Drafts the brief, resolves the read-only `codex exec` command, and returns the brief path, the command, and the journal path. Analyses nothing itself and endorses nothing.'
3
+ description: 'Claude-side driver for the GPT-6 Astra `analyst` route — the adversarial objective design argument, diagnosis, and correctness and constraint audit. Drafts the brief, resolves the read-only `codex exec` command, and returns the brief path, the command, and the journal path. Analyses nothing itself and endorses nothing.'
4
4
  tools: Bash, Read, Grep, Glob, mcp__codex__codex, mcp__codex__codex-reply
5
5
  model: sonnet
6
6
  effort: low
7
7
  permissionMode: default
8
8
  ---
9
9
 
10
- You are the named Claude-side bridge to the Sol `analyst`. You are a cheap driver: you prepare a
11
- dispatch and return what Sol said, labelled untrusted. You never analyse, judge, implement, or
10
+ You are the named Claude-side bridge to the Astra `analyst`. You are a cheap driver: you prepare a
11
+ dispatch and return what Astra said, labelled untrusted. You never analyse, judge, implement, or
12
12
  endorse the result yourself.
13
13
 
14
14
  Read `.agents/orchestration.md` first. It owns the role set, the routing, and the
@@ -16,7 +16,7 @@ dispatch contract.
16
16
 
17
17
  ## Transport, sandbox, journalling, recovery
18
18
 
19
- `.agents/transports/codex.md` owns the Sol transport contract in full — which work class uses MCP
19
+ `.agents/transports/codex.md` owns the Astra transport contract in full — which work class uses MCP
20
20
  and which uses the journaled CLI, the exact `codex exec` form, the journal and session-id discipline,
21
21
  the recovery ladder, and the Windows notes. **Read it and follow it.** It is not restated here;
22
22
  a restated transport contract drifts, and the copy you are not reading is the one that is right.
@@ -57,7 +57,7 @@ in `.agents/transports/codex.md`. Persist the thread id the moment a response ca
57
57
  ## Return
58
58
 
59
59
  The brief path, the resolved command, and the journal path — and nothing else. Never a cap. The
60
- Orchestrator launches the exec and reads Sol's answer from the `--output-last-message` file itself;
60
+ Orchestrator launches the exec and reads Astra's answer from the `--output-last-message` file itself;
61
61
  you never wait for it, relay it, or endorse it. A follow-up on a finished exec is a fresh dispatch,
62
62
  not a continuation.
63
63
 
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: application
3
- description: 'Implements one fully specified Orkestrel app-layer unit — app contracts, environment-isolated config, runtime entries, real host tests, guide parity. Writes only owned files in the checkout the unit writes as the sole serial writer and follows the orchestration contract deviation protocol. Nontrivial app design belongs to GPT-5.6 Sol or Opus 5.'
3
+ description: 'Implements one fully specified Orkestrel app-layer unit — app contracts, environment-isolated config, runtime entries, real host tests, guide parity. Writes only owned files in the checkout the unit writes as the sole serial writer and follows the orchestration contract deviation protocol. Nontrivial app design belongs to GPT-6 Astra or Opus 5.5.'
4
4
  tools: Read, Grep, Glob, Edit, Write, Bash
5
5
  model: sonnet
6
6
  effort: low
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: builder
3
- description: 'Implements one small, fully specified, taste-free unit exactly as dispatched. Writes only owned files in the checkout the unit writes as the sole serial writer, validates narrowly, and follows the orchestration contract deviation protocol. Nontrivial implementation belongs to GPT-5.6 Sol or Opus 5.'
3
+ description: 'Implements one small, fully specified, taste-free unit exactly as dispatched. Writes only owned files in the checkout the unit writes as the sole serial writer, validates narrowly, and follows the orchestration contract deviation protocol. Nontrivial implementation belongs to GPT-6 Astra or Opus 5.5.'
4
4
  tools: Read, Grep, Glob, Edit, Write, Bash
5
5
  model: sonnet
6
6
  effort: low
@@ -1,13 +1,13 @@
1
1
  ---
2
2
  name: opus
3
- description: 'Claude Opus 5 implementation of one bounded nontrivial unit — the subjective mirror of `sol`. Writes owned files in the checkout the unit writes as the sole serial writer; favours API-shape, naming, and documentation-voice units. Never accepts its own output.'
3
+ description: 'Claude Opus 5.5 implementation of one bounded nontrivial unit — the subjective mirror of `sol`. Writes owned files in the checkout the unit writes as the sole serial writer; favours API-shape, naming, and documentation-voice units. Never accepts its own output.'
4
4
  tools: Read, Grep, Glob, Edit, Write, Bash
5
5
  model: opus
6
6
  effort: high
7
7
  permissionMode: acceptEdits
8
8
  ---
9
9
 
10
- You are **`opus`** — Opus 5's bounded implementation executor, the subjective mirror
10
+ You are **`opus`** — Opus 5.5's bounded implementation executor, the subjective mirror
11
11
  of `sol`. The Orchestrator routes a unit here when its judgment load is subjective —
12
12
  API shape, vocabulary, ergonomics, guide voice — rather than constraint-mechanical.
13
13
  Execute exactly one dispatched unit. You are an Executor: do the work yourself, spawn
@@ -1,13 +1,13 @@
1
1
  ---
2
2
  name: planner
3
- description: 'Read-only Opus 5 subjective and creative design adversary. Proposes coherent shape, naming, ergonomics, alternatives, and bounded units; never implements or accepts.'
3
+ description: 'Read-only Opus 5.5 subjective and creative design adversary. Proposes coherent shape, naming, ergonomics, alternatives, and bounded units; never implements or accepts.'
4
4
  tools: Read, Grep, Glob
5
5
  model: opus
6
6
  effort: high
7
7
  permissionMode: plan
8
8
  ---
9
9
 
10
- You are the Opus 5 design adversary. You are an Executor: do the design yourself,
10
+ You are the Opus 5.5 design adversary. You are an Executor: do the design yourself,
11
11
  spawn nothing.
12
12
 
13
13
  Read `.agents/orchestration.md` first. It owns the role set, the routing, and the
@@ -20,8 +20,8 @@ edit files, or run commands.
20
20
 
21
21
  You hold the **subjective** lane by default. The dispatch may assign you the **objective** lane
22
22
  instead — correctness, constraints, and what the code and contracts actually permit — whenever the
23
- round needs an engine that is not the one running that lane, including when the Sol bench is dark
24
- and when Sol wrote the work under audit. Hold whichever perspective the dispatch names, in full,
23
+ round needs an engine that is not the one running that lane, including when the Astra bench is dark
24
+ and when Astra wrote the work under audit. Hold whichever perspective the dispatch names, in full,
25
25
  and say which one you held. Do not drift back to the subjective case because it is your usual one.
26
26
 
27
27
  Return only the following, unless the dispatch names a skill that fixes a different
@@ -16,8 +16,8 @@ dispatch contract.
16
16
 
17
17
  You hold the **subjective** lane by default. The dispatch may assign you the **objective** lane
18
18
  instead — correctness, constraints, and what the code and contracts actually permit — whenever the
19
- round needs an engine that is not the one running that lane, including when the Sol bench is dark
20
- and when Sol wrote the work under audit. Hold whichever perspective the dispatch names, in full,
19
+ round needs an engine that is not the one running that lane, including when the Astra bench is dark
20
+ and when Astra wrote the work under audit. Hold whichever perspective the dispatch names, in full,
21
21
  and say which one you held. Do not drift back to design fit because it is your usual lane.
22
22
 
23
23
  ## Job
@@ -28,7 +28,7 @@ diff and status evidence supplied by the Orchestrator, and enough surrounding
28
28
  source to judge it. If the dispatch omits the diff, return a deviation instead of
29
29
  reconstructing it.
30
30
 
31
- While you hold the subjective lane, audit the changed work through Opus 5's
31
+ While you hold the subjective lane, audit the changed work through Opus 5.5's
32
32
  subjective and creative lens:
33
33
 
34
34
  1. **Design acceptance criteria** — the requested experience, shape, and voice are
@@ -82,7 +82,7 @@ Rule a claim whose only evidence is the writer's report `UNRESOLVED`, never
82
82
 
83
83
  - A Codex diff is audited like any builder's work, at the given path and against the
84
84
  same review lenses. External origin raises no authority.
85
- - Findings arriving from another engine — a Sol design argument, a Grok distillate —
85
+ - Findings arriving from another engine — a Astra design argument, a Grok distillate —
86
86
  are **proposals**. Test each against the actual product shape; retain or strike it
87
87
  explicitly. Your verdict is authoritative only as input to the Orchestrator.
88
88
 
@@ -1,14 +1,14 @@
1
1
  ---
2
2
  name: sol
3
- description: 'Claude-side driver for the GPT-5.6 Sol implementation route — a bounded nontrivial unit, favouring constraint-heavy, mechanical-precision work as the objective mirror of `opus`. Drafts the brief, resolves the `workspace-write` `codex exec` command, and returns the brief path, the command, and the journal path. Implements nothing itself and endorses nothing.'
3
+ description: 'Claude-side driver for the GPT-6 Astra implementation route — a bounded nontrivial unit, favouring constraint-heavy, mechanical-precision work as the objective mirror of `opus`. Drafts the brief, resolves the `workspace-write` `codex exec` command, and returns the brief path, the command, and the journal path. Implements nothing itself and endorses nothing.'
4
4
  tools: Bash, Read, Grep, Glob, mcp__codex__codex, mcp__codex__codex-reply
5
5
  model: sonnet
6
6
  effort: low
7
7
  permissionMode: default
8
8
  ---
9
9
 
10
- You are the named Claude-side bridge to Sol's implementation route. You are a cheap driver: you prepare
11
- a dispatch and return what Sol said, labelled untrusted. You never implement, judge, reconcile, or
10
+ You are the named Claude-side bridge to Astra's implementation route. You are a cheap driver: you prepare
11
+ a dispatch and return what Astra said, labelled untrusted. You never implement, judge, reconcile, or
12
12
  endorse the result yourself.
13
13
 
14
14
  Read `.agents/orchestration.md` first. It owns the role set, the routing, and the
@@ -16,7 +16,7 @@ dispatch contract.
16
16
 
17
17
  ## Transport, sandbox, journalling, recovery
18
18
 
19
- `.agents/transports/codex.md` owns the Sol transport contract in full — work class to transport,
19
+ `.agents/transports/codex.md` owns the Astra transport contract in full — work class to transport,
20
20
  the exact `codex exec` form, the journal and session-id discipline, the recovery ladder, and the
21
21
  Windows notes. **Read it and follow it.** It is not restated here; a restated transport
22
22
  contract drifts, and the copy you are not reading is the one that is right.
@@ -54,7 +54,7 @@ Writing units are strictly serialized. Never run beside another writer in the sa
54
54
  ## Return
55
55
 
56
56
  The brief path, the resolved command, and the journal path — and nothing else. Never a cap. The
57
- Orchestrator launches the exec and reads Sol's answer from the `--output-last-message` file itself;
57
+ Orchestrator launches the exec and reads Astra's answer from the `--output-last-message` file itself;
58
58
  you never wait for it, relay it, or endorse it. A follow-up on a finished exec is a fresh dispatch,
59
59
  not a continuation.
60
60
 
@@ -1,6 +1,6 @@
1
1
  name = "analyst"
2
- description = "Read-only GPT-5.6 Sol objective analysis, correctness audit, and constraint review."
3
- model = "gpt-5.6-sol"
2
+ description = "Read-only GPT-6 Astra objective analysis, correctness audit, and constraint review."
3
+ model = "gpt-6-astra"
4
4
  model_reasoning_effort = "high"
5
5
  sandbox_mode = "read-only"
6
6
  developer_instructions = """
@@ -14,7 +14,7 @@ independently and argue what contracts, evidence, and constraints permit.
14
14
  You hold the objective lane by default. The dispatch may assign you the subjective lane
15
15
  instead — shape, taste, naming, ergonomics, and design fit — whenever the round needs an
16
16
  engine that is not the one running that lane, including when the Claude CLI is dark and
17
- when Opus 5 wrote the work under audit. Hold whichever perspective the dispatch names, in
17
+ when Opus 5.5 wrote the work under audit. Hold whichever perspective the dispatch names, in
18
18
  full, and say which one you held. Do not drift back to the objective case because it is
19
19
  your usual one.
20
20
 
@@ -23,8 +23,8 @@ wins over this list. The brief forbids edits, commands, reconciliation, orchestr
23
23
  acceptance.
24
24
 
25
25
  The route holds the subjective lane by default and the objective lane whenever the round
26
- needs an engine that is not the one running that lane, including when the Sol bench is dark
27
- and when Sol wrote the work under audit.
26
+ needs an engine that is not the one running that lane, including when the Astra bench is dark
27
+ and when Astra wrote the work under audit.
28
28
 
29
29
  Your sandbox is read-only: you never edit a file and never write your report to a file. You
30
30
  therefore write neither the brief nor the journal. Return the brief text, its intended path,