@yemi33/minions 0.1.2453 → 0.1.2455

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (48) hide show
  1. package/bin/install-internal-minions.js +136 -6
  2. package/bin/minions.js +27 -13
  3. package/dashboard/js/agent-identity.js +121 -0
  4. package/dashboard/js/modal.js +4 -0
  5. package/dashboard/js/refresh.js +38 -3
  6. package/dashboard/js/render-agents.js +14 -2
  7. package/dashboard/js/render-dispatch.js +2 -2
  8. package/dashboard/js/render-other.js +147 -13
  9. package/dashboard/js/render-prd.js +6 -5
  10. package/dashboard/js/render-prs.js +44 -25
  11. package/dashboard/js/render-work-items.js +442 -22
  12. package/dashboard/js/settings.js +8 -0
  13. package/dashboard/js/utils.js +40 -0
  14. package/dashboard/pages/tools.html +1 -0
  15. package/dashboard/pages/work.html +13 -0
  16. package/dashboard/shared/pr-author.js +75 -0
  17. package/dashboard/shared/pr-filters.js +28 -5
  18. package/dashboard/styles.css +85 -0
  19. package/dashboard-build.js +3 -3
  20. package/dashboard.js +93 -7
  21. package/docs/README.md +1 -0
  22. package/docs/copilot-cli-schema.md +1 -0
  23. package/docs/engine-restart.md +4 -2
  24. package/docs/internal-install.md +44 -5
  25. package/docs/named-agents.md +48 -0
  26. package/docs/pr-author-identity.md +63 -10
  27. package/docs/runtime-adapters.md +39 -0
  28. package/docs/temporary-agents.md +172 -0
  29. package/engine/ado/comment.js +261 -4
  30. package/engine/agents/llm.js +26 -0
  31. package/engine/agents/playbook.js +2 -1
  32. package/engine/api/settings-validation.js +25 -0
  33. package/engine/core/operator-identity.js +23 -1
  34. package/engine/core/queries.js +53 -2
  35. package/engine/core/shared.js +286 -9
  36. package/engine/db/migrations/032-review-enrolled-pr-context-only.js +95 -0
  37. package/engine/operations/cli.js +102 -1
  38. package/engine/orchestration/lifecycle.js +7 -0
  39. package/engine/orchestration/routing.js +4 -1
  40. package/engine/providers/gh-comment.js +159 -0
  41. package/engine/recovery/stop-stack.js +16 -2
  42. package/engine/runtimes/claude.js +3 -0
  43. package/engine/runtimes/codex.js +4 -0
  44. package/engine/runtimes/copilot.js +23 -2
  45. package/engine.js +2 -2
  46. package/package.json +1 -1
  47. package/playbooks/fix.md +23 -2
  48. package/playbooks/shared-rules.md +24 -2
@@ -29,6 +29,48 @@ The `name`, `emoji`, `role`, and `expertise` fields are **human-facing metadata*
29
29
  label the agent in the dashboard and document intent. They do **not** drive dispatch (see
30
30
  §3) and the `expertise` array is not executable (see §4).
31
31
 
32
+ ### 1.1 Temporary agents and friendly call signs
33
+
34
+ When every permanent agent is busy and `engine.allowTempAgents` is set, the engine
35
+ spawns an **ephemeral agent** with a machine id of the form `temp-<uid>` (see
36
+ [`engine/orchestration/routing.js`](../engine/orchestration/routing.js)). That canonical
37
+ `temp-…` id is the **only** machine identity — it is unchanged for routing, SQL keys,
38
+ filesystem paths, live-output logs, process ownership, APIs, per-agent memory rules
39
+ (§8: `temp-*` ids are skipped for personal memory), and PR/reconciliation records.
40
+
41
+ For display only, each temp agent also gets a stable, memorable **call sign** such as
42
+ `Cosmic Comet`, rendered next to a visible `Temp` badge (`CallSign` + badge). This is
43
+ an *additive display identity*, never a rename of the machine id.
44
+
45
+ **Deterministic-from-id invariant (no persisted state).** The call sign is derived
46
+ purely from the immutable canonical id by
47
+ [`shared.tempAgentCallSign(id)`](../engine/core/shared.js): a single FNV-1a 32-bit hash of
48
+ the lower-cased id selects an adjective from its low bits (`h % ADJ`) and a name from its
49
+ high bits (`⌊h / ADJ⌋ % NAMES`), each drawn from a curated, roster-disjoint vocabulary.
50
+ Because the derivation depends only on the id and nothing else:
51
+
52
+ - **Restart- and history-stable** — the same id always yields the same call sign across
53
+ engine/dashboard restarts and for long-exited agents in archived/history views, so no
54
+ SQL migration or JSON runtime state is needed (and none is added). An already-assigned
55
+ name is preserved by construction — it cannot drift when the active set changes.
56
+ - **Collision behavior is deterministic** — the ~24×24 adjective/name space plus the
57
+ single-hash bit-slice keeps concurrently-visible temp agents well-spread; the curated
58
+ `name` list is disjoint from the shipped roster (Ripley/Dallas/Lambert/Rebecca/Ralph)
59
+ so a call sign never equals a configured display name, and the always-visible `Temp`
60
+ badge disambiguates even a rare full collision.
61
+ - **Case-insensitive** — legacy mixed-case `Temp-*` ids derive the same call sign as
62
+ their canonical lower-case form.
63
+
64
+ The client mirrors this vocabulary + algorithm in
65
+ [`dashboard/js/agent-identity.js`](../dashboard/js/agent-identity.js) (`MinionsAgentIdentity`),
66
+ the single formatter every user-facing renderer calls, so a legacy record missing the
67
+ server-stamped `agentName` still renders the identical stable call sign. Parity between
68
+ the two implementations is asserted by
69
+ [`test/unit/temp-agent-identity.test.js`](../test/unit/temp-agent-identity.test.js).
70
+ API payloads expose structured fields (`id`, friendly `name`/`callSign`, `isTemp`) rather
71
+ than replacing ids with formatted strings, and the canonical `temp-…` id stays discoverable
72
+ in detail views/tooltips for debugging.
73
+
32
74
  ## 2. How an agent is configured
33
75
 
34
76
  The full per-agent config shape supported under `config.agents.<id>`:
@@ -123,6 +165,12 @@ context; `_any_` routes to any available idle agent (lowest error rate first).
123
165
  `implement:large` is selected for items with `estimated_complexity: "large"`. If both
124
166
  preferred and fallback are busy, the engine falls back to any idle agent.
125
167
 
168
+ **Overflow → temp agents.** If no configured agent is idle and the operator has enabled
169
+ `engine.allowTempAgents` (default off), the engine may spawn an ephemeral, personaless
170
+ `temp-<uid>` agent as a last-resort fallback. Temp agents inherit fleet-default runtime/model/
171
+ tools and shared/project memory but carry no charter and no personal memory. See
172
+ [`temporary-agents.md`](temporary-agents.md).
173
+
126
174
  **Dispatch rules** (from `routing.md`):
127
175
 
128
176
  - **Eager by default** — the engine spawns every agent that can start work, not one at a time.
@@ -62,8 +62,15 @@ agent id. `pr.agent` is now **never** a human login.
62
62
  non-empty incoming value wins, otherwise the existing value is preserved, so a
63
63
  **partial provider refresh never erases a richer author already stored**. A
64
64
  provider change trusts the newer identity wholesale.
65
- - `prAuthorLabel(author)` — human-readable label (displayName → login → id) for
66
- the typed SQL column and filter/search text.
65
+ - `prAuthorDisplayName(author)` — the **email-free** human-readable name every
66
+ surface should show. Prefers a real `displayName` → `login`, and only when the
67
+ sole usable identity is an email does it fall back to the **local-part** (domain
68
+ dropped). Any `@` anywhere forces local-part extraction, so a malformed value
69
+ (`"Name <a@b>"`, `weird@x@y`, a bare address) can never leak an `@`. Returns
70
+ `''` (→ **Unknown**) when nothing usable remains. Presentation only — the
71
+ structured `login`/`id`/`descriptor` fields are left untouched.
72
+ - `prAuthorLabel(author)` — thin delegate to `prAuthorDisplayName` (kept for
73
+ callers wanting a "label"); therefore also email-free.
67
74
 
68
75
  `upsertPullRequestRecord` normalizes `author` on the create path and merges it
69
76
  via `mergePrAuthorIdentity` on the update path (and in `_mergeDuplicatePrInto`),
@@ -76,7 +83,11 @@ manual link, dashboard enrich, dedup — converge on one merge behavior.
76
83
  projection column (plus `idx_pr_author`). The canonical structured identity
77
84
  always lives in the row's `data` JSON; the column is a label projection for
78
85
  filters/sort. `engine/persistence/pull-requests-store.js` writes it via a local
79
- `_authorLabel(pr.author)` (displayName → login → id).
86
+ `_authorLabel(pr.author)` (displayName → login → id). This column feeds only the
87
+ `idx_pr_author` index for server-side filter/sort — it is **never selected back
88
+ or surfaced to any client** (reads return the `data` JSON), so it may still
89
+ contain an email and is deliberately left as machine-facing projection, not a
90
+ display label.
80
91
 
81
92
  Backfill is **fail-closed**. A legacy row whose `agent` is a human login (i.e.
82
93
  not `'human'`, not a configured Minions agent id from `config.json`, not
@@ -96,14 +107,26 @@ alongside the existing `agent` field.
96
107
 
97
108
  ## Dashboard
98
109
 
110
+ Every dashboard surface that shows a PR author routes through **one** shared
111
+ browser formatter, `window.MinionsPrAuthor` (`dashboard/shared/pr-author.js`),
112
+ which mirrors the engine `prAuthorDisplayName` so the operator-visible name is
113
+ always **email-free** — an ADO author stored as `Name <email>` or a bare email
114
+ `uniqueName` renders as the name / email local-part, never the address. It is
115
+ listed in `DASHBOARD_SHARED_JS` before `pr-filters`, so classic and slim share it.
116
+
99
117
  `dashboard/js/render-prs.js` replaces the Pull Requests table's **Agent** column
100
- with **Author** (display name + `@login`, safely escaped, linked to the provider
101
- profile only when a validated `http(s)` URL is present). The Minions agent stays
102
- visible in the PR detail panel under **Minions agent**, next to a new **Author**
103
- line, so operators can still answer both "who is this for?" and "which agent
104
- handled it?". The shared author filter (`dashboard/shared/pr-filters.js`) now
105
- reads structured `pr.author` only and no longer falls back to `pr.agent`, so the
106
- filter and the displayed column stay aligned (legacy rows group as Unknown).
118
+ with **Author** (`MinionsPrAuthor.display` → name + optional `@handle` — a handle
119
+ is offered only for a real non-email login, safely escaped, linked to the
120
+ provider profile only when a validated `http(s)` URL is present). The Minions
121
+ agent stays visible in the PR detail panel under **Minions agent**, next to the
122
+ **Author** line, so operators can still answer both "who is this for?" and "which
123
+ agent handled it?". The shared author filter (`dashboard/shared/pr-filters.js`)
124
+ shows the same email-free **label** while grouping on a stable machine **key**
125
+ (`id` → `descriptor` → `login` → … via `_prAuthorMachineKey`), so two people who
126
+ share a display name never collapse and the filter stays aligned with the column
127
+ (legacy rows group as Unknown). The work-item follow-up chip
128
+ (`dashboard/js/render-work-items.js`) strips the email from a
129
+ `pr_followup.parent_comment_author` the same way.
107
130
 
108
131
  ## Backward-compatible reads
109
132
 
@@ -112,3 +135,33 @@ filter and the displayed column stay aligned (legacy rows group as Unknown).
112
135
  author string keeps rendering. This is a tolerant read, not a retained
113
136
  compatibility shim — the overloaded `agent = login` write behavior was removed
114
137
  outright (no `docs/deprecated.json` entry required).
138
+
139
+ ## AutoFix foreign-author guardrail (W-msd8pr2300ev95e1)
140
+
141
+ When AutoFix is enabled on a PR whose author is **not** the current authenticated
142
+ identity for that provider/repo scope, the Pull Requests table's **AutoFix** cell
143
+ renders a prominent warning icon beside the toggle: a red `⚠`
144
+ ("Auto-fix enabled for another author's PR", naming the author when known) or, if
145
+ ownership cannot be established, a yellow `⚠` ("author ownership unknown"). This
146
+ never changes behavior — enabling auto-fix on another author's PR stays an
147
+ explicit operator action — it only surfaces the risk.
148
+
149
+ Ownership is resolved by `shared.resolvePrAuthorOwnership(pr, selfByProvider)`
150
+ (`self` | `foreign` | `unknown`) using **stable provider identities only** —
151
+ GitHub canonical `login` / numeric `id`; ADO identity `id` / `descriptor`. Display
152
+ names and emails are never compared, and the determination **fails closed** to
153
+ `unknown` (shown as the neutral warning, never silently treated as self) whenever
154
+ either side's stable identity is missing.
155
+
156
+ `engine/core/queries.js#getPullRequests` stamps `pr._authorOwnership` once per
157
+ build, resolving the current identity per provider:
158
+
159
+ - **GitHub** — `operator-identity.resolveGithubViewerLoginStrict(config)`, which
160
+ accepts only `engine.operatorLogin` or `gh api user` (never a git-email / OS
161
+ fallback).
162
+ - **ADO** — the optional operator-configured `engine.operatorAdoIdentity`
163
+ (`{ id?, descriptor? }`). There is no automatic ADO current-identity seam, so
164
+ when it is unset every ADO PR reports `unknown` (fail-closed). Set it under
165
+ **Settings → Operator & Comments → ADO identity id / descriptor**. It is a
166
+ read-only identity value (like `operatorLogin`), never an email — the ADO
167
+ `uniqueName`/login is deliberately excluded from both storage and comparison.
@@ -108,6 +108,44 @@ marker and visible sign-off, and lifecycle persists the same value as
108
108
  `minionsReview.model`. Model catalogs are never used to guess an implicit
109
109
  runtime default.
110
110
 
111
+ ### Billable-unit usage accounting
112
+
113
+ Some runtimes bill a native, non-USD unit — the GitHub Copilot CLI reports
114
+ **premium requests** per session/turn. `parseOutput().usage` therefore carries a
115
+ generic typed descriptor alongside the token/cost fields so the engine accounts
116
+ for credits without ever branching on a runtime name:
117
+
118
+ ```js
119
+ usage.billable = { unit: 'premiumRequests', value: <number|null>, reported: <bool> }
120
+ ```
121
+
122
+ - **Authoritative, never derived.** `value` is the count the CLI reported; the
123
+ engine never estimates credits from tokens, model, elapsed time, or pricing.
124
+ - **Unavailable ≠ 0.** When the installed CLI did not expose the field for a run,
125
+ the adapter emits `value: null, reported: false` (and Copilot leaves
126
+ `premiumRequests` / `costUsd` `null`, not `0`). Surfaces render unreported as
127
+ `N/A`, distinct from a genuine reported `0`.
128
+ - **Already scoped — no delta diffing.** Copilot `parseOutput` reassigns
129
+ `usage` on each terminal `result` event, so the value is the finalized count
130
+ for THIS invocation / pooled turn, not a running session total. Downstream code
131
+ therefore **sums** per-invocation values; it does not diff cumulative counters.
132
+ This per-invocation scoping is what makes streamed, resumed, retried, and
133
+ pooled dispatches safe from double-counting.
134
+
135
+ `shared.accumulateBillableUnits(target, usage.billable)` folds a descriptor into
136
+ `target.billableUnits[unit] = { total, reported, unavailable }`, used by
137
+ `engine/orchestration/lifecycle.js#updateMetrics` (per-agent / `_engine` /
138
+ `_daily`) and `engine/agents/llm.js` (Command Center / doc-chat). USD is
139
+ deliberately **excluded** from this path — dollar cost already flows through the
140
+ richer `costUsd` + token accounting, so routing it here too would double-count.
141
+ The map persists inside the existing metrics JSON blob (no migration; older rows
142
+ simply lack the field), and `engine/core/queries.js#getMetrics` resolves
143
+ `m.billableUnit` from the runtime capability so the dashboard keys off typed
144
+ metadata, never the runtime name. A future adapter can surface its own unit by
145
+ setting `capabilities.billableUnit` and emitting `usage.billable` — no call-site
146
+ changes required.
147
+
148
+
111
149
  ## Capability Flags
112
150
 
113
151
  Engine code branches on consumed flags, never on runtime names. A capability
@@ -122,6 +160,7 @@ to act on that behavioral split.
122
160
  | `systemPromptFile` | ✓ | ✗ | ✗ | sysprompt via `--system-prompt-file` (else inlined into stdin by `buildPrompt`). |
123
161
  | `effortLevels` | ✓ | ✓ | ✓ | Runtime reasoning-effort controls. Codex maps through `--config model_reasoning_effort=...`; Copilot maps `'max'` to `'xhigh'`; Claude leaves `'max'` alone. |
124
162
  | `costTracking` | ✓ | ✗ | ✗ | USD + token counts in the result event. Copilot only emits `premiumRequests`; Codex accounting is not assumed. |
163
+ | `billableUnit` | `usd` | `premiumRequests` | `usd` | Typed native billable unit for the display layer. The dashboard renders THIS unit in the same position a USD runtime shows dollar cost, switching labels/units/tooltips off the metadata — never a runtime-name branch. Adapters with no native credit unit default to `'usd'`. |
125
164
  | `modelShorthands` | ✓ | ✗ | ✗ | Bare `sonnet`/`opus`/`haiku` accepted by Claude only. |
126
165
  | `modelDiscovery` | ✗ | ✓ | ✓ | `listModels()` returns a real catalog when supported (Copilot API, Codex `debug models --bundled`). |
127
166
  | `strictModelCatalog` | — | — | ✗ | An explicit false lets the runtime validate IDs missing from its cached catalog; Codex uses this because bundled catalogs can lag staged/custom models. |
@@ -0,0 +1,172 @@
1
+ # Temporary Agents — ephemeral `temp-<uid>` fallback dispatch
2
+
3
+ Minions normally dispatches work to a fixed roster of **named agents** (Ripley, Dallas, …;
4
+ see [`named-agents.md`](named-agents.md)). When every eligible named agent is busy and the
5
+ operator has opted in, the engine can spawn a **temporary agent** — a one-shot, unnamed
6
+ `temp-<uid>` identity that exists only to absorb overflow work so the queue does not stall.
7
+
8
+ This document describes the complete temp-agent behavior, verified against the current
9
+ implementation. Temp agents are deliberately minimal: they have **no charter, no persona,
10
+ no personal memory, and no persistent config record**. They are a scheduling escape valve,
11
+ not a sixth teammate.
12
+
13
+ > Orientation only — the load-bearing contracts live in code. Cross-references:
14
+ > [`named-agents.md`](named-agents.md), [`team-memory.md`](team-memory.md),
15
+ > [`workspace-manifests.md`](workspace-manifests.md), [`runtime-adapters.md`](runtime-adapters.md),
16
+ > and [`../CLAUDE.md`](../CLAUDE.md).
17
+
18
+ ## 1. When and why a temp agent is created
19
+
20
+ Temp agents are an **opt-in, last-resort** step of agent resolution in
21
+ `resolveAgent(workType, config, opts)` (`engine/orchestration/routing.js`). Resolution tries,
22
+ in order: agent hints → routing `preferred` → routing `fallback` → any idle configured agent
23
+ (lowest error rate first). Only if **all** of those miss does it consider a temp agent, and
24
+ only when every one of these gates passes:
25
+
26
+ 1. **Config gate — `engine.allowTempAgents`.** Defaults to **`false`**
27
+ (`ENGINE_DEFAULTS.allowTempAgents`, `engine/core/shared.js`). With it off, `resolveAgent`
28
+ returns `null` and the work item simply stays `pending` until a named agent frees up — no
29
+ temp is ever created (`engine/orchestration/routing.js`, the `if (config.engine?.allowTempAgents)`
30
+ branch). It is exposed as a Settings toggle ("Allow Temp Agents", `set-allowTempAgents`) in
31
+ `dashboard/js/settings.js` and surfaced on `GET /api/*` status via `dashboard.js`.
32
+ 2. **Concurrency gate — per-tick temp budget.** The engine calls `setTempBudget(maxConcurrent - activeCount)`
33
+ once per tick (`engine.js`), so temps count against `engine.maxConcurrent` exactly like
34
+ named agents. When the budget is exhausted, `resolveAgent` logs
35
+ `Temp agent refused … per-tick budget exhausted` and returns `null`. This guard stops a
36
+ mass-discovery pass (e.g. a PR-poll sweep over 20 failing PRs) from registering one temp per
37
+ pending item and spawning far more processes than `maxConcurrent` allows (closes #1209).
38
+
39
+ Because a `temp-<uid>` id is by construction never equal to a named author agent, a temp is
40
+ always a valid **non-author** reviewer under the self-review ban, and a valid reassignment
41
+ target under `excludeAgent`. Temp creation is also the documented fallback for the
42
+ **decompose** flow: when `engine.autoDecompose` splits an `implement:large` item and all
43
+ permanent agents are busy, the engine can spawn a `temp-<uid>` to run a sub-task
44
+ (see [`../CLAUDE.md`](../CLAUDE.md) → "Decomposition").
45
+
46
+ ## 2. Naming and runtime/state representation
47
+
48
+ - **Id:** `temp-${shared.uid()}` — e.g. `temp-mscilxge004b1de2`. The `temp-` prefix is a
49
+ reserved sentinel checked throughout the engine; a configured agent id may not collide with
50
+ it (`docs/specs/agent-configurability.md`).
51
+ - **Display name / role:** `{ name: "Temp-<4chars>", role: "Temporary Agent" }`, derived from
52
+ characters 5–8 of the id (`engine/orchestration/routing.js`).
53
+ - **In-memory only:** the record lives in a module-level `Map` `tempAgents` in
54
+ `engine/orchestration/routing.js` (`tempAgentId → { name, role, createdAt }`). It is **not**
55
+ written to `config.json`, and **not** persisted to SQLite. The authoritative agent identity
56
+ for the work travels on the work item / dispatch record (`item.agent`, `item.agentName`,
57
+ `item.agentRole`), which persist normally.
58
+ - **Dashboard/API:** an *active* temp agent is added to the agent roster snapshot by
59
+ `engine/core/queries.js#getAgents` with `emoji: 💨`, `expertise: []`, and `_temp: true`, so
60
+ it shows up as a tile while it runs. Temp agents are deliberately **excluded** from per-agent
61
+ aggregate metrics — PR counts and runtime-by-agent skip `temp-*` ids
62
+ (`engine/core/queries.js`).
63
+
64
+ ## 3. Charter / persona — not inherited
65
+
66
+ A temp agent does **not** inherit, copy, or reference the charter or persona of any permanent
67
+ agent. It has no `charter.md`: `buildSystemPrompt` reads
68
+ `agents/<id>/charter.md`, which returns empty for a temp id, and falls back to a bare identity
69
+ of `{ name, role: "Temporary Agent", expertise: [] }` (`engine/agents/playbook.js`). The system
70
+ prompt therefore carries only the generic role label and the shared team rules — no expertise
71
+ tags, no voice, no boundaries section.
72
+
73
+ ## 4. Memory, notes, and pinned context — retrieved, not inherited
74
+
75
+ Temp agents get the **shared, non-personal** context every dispatch gets, but accrue and read
76
+ **no personal memory**. Exact behavior (`engine/agents/playbook.js`):
77
+
78
+ | Context source | Temp agent behavior |
79
+ |----------------|---------------------|
80
+ | `pinned.md` (operator pinned context) | **Injected** — unconditional, same as any agent. |
81
+ | Task-selected **Relevant Memory** pack (`memoryRetrieval` on) | **Injected** — the retrieval query is keyed on the **work item / project**, not the agent persona, so temps get the same project-scoped pack. |
82
+ | `notes.md` (team notes) | **Injected** (when the Relevant Memory pack did not already replace legacy injection). Not gated on agent id. |
83
+ | Applicable **review-learning** lessons | **Injected** — project/file-scoped recall, agent-independent. |
84
+ | Per-agent memory `knowledge/agents/<id>.md` | **NOT injected** — the injection is explicitly skipped for ids starting with `temp-` (`engine/agents/playbook.js`), and no such file exists for a temp anyway. |
85
+
86
+ It does **not** inherit a *specific* permanent agent's `knowledge/agents/<agentId>.md`
87
+ notebook — there is no "borrow Ripley's memory" path. A temp reads shared/project memory only.
88
+
89
+ **Write side (consolidation).** A temp agent's findings never become durable personal memory.
90
+ The consolidation pipeline (`engine/memory/consolidation.js`) skips `temp-*` authors in
91
+ `appendToAgentMemory`, `reconcileAndAppendToAgentMemory`, and `maybeSummarizeAgentMemory`, and
92
+ memory routing treats `temp-*` as an ineligible route. Team-wide consolidation into `notes.md`
93
+ and the categorized `knowledge/` KB still happens from any inbox note the temp writes (those
94
+ are team-scoped, not per-agent), but nothing is filed under a `temp-*` identity.
95
+
96
+ ## 5. Runtime, CLI, model, tools, and workspace manifest
97
+
98
+ Because a temp agent has **no `config.agents.<id>` entry**, every per-agent override resolves to
99
+ `null` and falls through to the **engine fleet defaults** (`engine.js#spawnAgent`, with
100
+ `agentConfig = config.agents?.[agentId] || null`):
101
+
102
+ - **CLI runtime:** `resolveAgentCli(null, engine)` → `engine.defaultCli` → `'copilot'`.
103
+ - **Model:** `resolveAgentModel(null, engine)` → `engine.defaultModel` → runtime default.
104
+ - **Budget:** `resolveAgentMaxBudget(null, engine)` → `engine.maxBudgetUsd`.
105
+ - **Bare mode:** `resolveAgentBareMode(null, engine)` → `engine.claudeBareMode` → `false`.
106
+ - **Workspace manifest:** `resolveAgentManifest(null)` returns the **permissive default** — no
107
+ `allowed_repos` gate, no `allowed_tools` narrowing (`engine/core/shared.js`,
108
+ [`workspace-manifests.md`](workspace-manifests.md)). A temp therefore inherits the fleet
109
+ `--allowedTools` set unmodified and is never blocked by the repo gate.
110
+ - **Session resume:** temp ids are excluded from session persistence in every runtime adapter
111
+ (`getResumeSessionId` / session save short-circuit on `temp-*` in
112
+ `engine/runtimes/{claude,copilot,codex}.js`), so a temp never resumes a prior session.
113
+
114
+ Spawn otherwise follows the normal path — worktree allocation, prompt build, and
115
+ `engine/agents/spawn-agent.js` — identical to a named agent.
116
+
117
+ ## 6. Routing, retries, and permanent-agent availability
118
+
119
+ - **Overflow only.** A temp is only ever chosen after hints, preferred, fallback, and any idle
120
+ configured agent all miss. The moment a named agent is idle, routing prefers it — temps are
121
+ never chosen ahead of an available permanent.
122
+ - **API validation.** `POST /api/work-items` accepts an explicit `temp-*` agent id **only when
123
+ `allowTempAgents` is true**; otherwise it is rejected as an unknown agent
124
+ (`engine/planning/work-item-validation.js`; the `tempAllowed` check).
125
+ - **Retry / reassignment.** The per-agent retry-reassignment machinery
126
+ ([`named-agents.md` §3](named-agents.md)) treats a temp id like any other agent for
127
+ `excludeAgent` purposes. If `allowTempAgents` is off (or the budget is exhausted) and no
128
+ alternate named agent exists, the engine writes a deduped inbox note advising the operator to
129
+ enable `allowTempAgents` or add another routing target, then re-dispatches to the same agent to
130
+ avoid deadlock (`engine.js`).
131
+ - **Unknown-agent guard.** A pending dispatch whose `item.agent` is a `temp-*` id is **not**
132
+ treated as an "unknown configured agent" and is not re-routed on that basis (`engine.js`), so a
133
+ temp assignment survives to spawn even though it is absent from `config.agents`.
134
+
135
+ ## 7. Lifecycle and cleanup
136
+
137
+ - **Ephemeral by design.** `cleanupTempAgent(agentId)` (`engine.js`) runs on **every**
138
+ dispatch-completion path (success, failure, timeout, orphan recovery, …). It deletes the id
139
+ from the in-memory `tempAgents` map and removes the `agents/<temp-id>/` directory (live-output
140
+ log etc.), wrapped in `shared._retryFsOp` so Windows `EBUSY`/`EPERM` from AV/file-indexer locks
141
+ are retried (regression test: `test/unit/cleanup-temp-agent-retry.test.js`).
142
+ - **Orphan sweep.** As a backstop, the periodic cleanup pass
143
+ (`engine/orchestration/cleanup.js`) reaps `agents/temp-*` directories that are no longer
144
+ referenced by any dispatch record, using a **1-hour mtime gate** so a still-spawning temp is
145
+ never reaped mid-launch.
146
+ - **Restart behavior.** The `tempAgents` map is in-memory, so an engine restart discards the
147
+ `{ name, role }` record. The persisted work-item/dispatch fields (`item.agent`,
148
+ `item.agentName`, `item.agentRole`) survive, and if a pending temp dispatch is resumed after
149
+ restart, `buildSystemPrompt` degrades gracefully to the default `Temporary Agent` identity
150
+ (`engine/agents/playbook.js`). No temp identity is re-registered or "revived" — it simply runs
151
+ out its one dispatch and is cleaned up.
152
+ - **Retained artifacts.** Only team-scoped artifacts a temp produced during its run persist:
153
+ inbox notes it wrote (subject to normal team consolidation), tracked PRs it opened (attributed
154
+ to the `temp-*` id but excluded from per-agent PR metrics), and the work item's own history.
155
+ No personal memory, charter, or config record is retained.
156
+
157
+ ## 8. Limitations and isolation implications
158
+
159
+ - **Off by default; overflow only.** With `allowTempAgents: false` (the default) temps never
160
+ exist; enabling them is an explicit operator decision (config + Settings toggle).
161
+ - **Fleet-default permissions.** A temp runs with the fleet-default runtime, model, budget, and
162
+ tool set, and a **permissive** workspace manifest (no repo gate, no tool narrowing). If your
163
+ isolation model depends on per-agent `workspace_manifest` restrictions, note that a temp agent
164
+ bypasses them by construction — it has no manifest to enforce. Constrain temps through
165
+ fleet-level `engine.*` defaults (e.g. `engine.maxBudgetUsd`, `engine.allowedTools`) rather than
166
+ per-agent config.
167
+ - **No accountable persona / no learning.** Temps carry no charter and accrue no personal
168
+ memory, so their work is not shaped by, and does not feed back into, any named agent's memory.
169
+ They are appropriate for burst capacity, not for work that should build durable per-agent
170
+ context.
171
+ - **Concurrency-bounded.** The per-tick temp budget ties temp creation to `engine.maxConcurrent`;
172
+ temps cannot exceed the global concurrency cap.