@hybridlabor-api/aos 4.13.1 → 4.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (151) hide show
  1. package/.agents/AGENTS.md +8 -0
  2. package/.agents/nodes.json +5 -2
  3. package/.claude/hooks/conventional-commits.mjs +14 -15
  4. package/.claude/hooks/env-file-protection.mjs +14 -15
  5. package/.claude/hooks/go-gate.mjs +157 -13
  6. package/.claude/hooks/go-token.mjs +55 -0
  7. package/.claude/hooks/memb-inject.mjs +75 -62
  8. package/.claude/hooks/trail-autostart.mjs +27 -0
  9. package/.claude/settings.json +18 -0
  10. package/.claude/workflows/startcycle-dispatch.mjs +11 -4
  11. package/.opencode/plugins/bdb-aos.js +98 -121
  12. package/.opencode/plugins/lib/trail-autostart.js +38 -0
  13. package/CLAUDE.md +1 -1
  14. package/README.de.md +6 -10
  15. package/README.md +6 -10
  16. package/README.pt.md +6 -10
  17. package/THIRD_PARTY_NOTICES.md +19 -3
  18. package/assets/header-v5.png +0 -0
  19. package/bin/aos-acp.mjs +211 -0
  20. package/bin/aos-doctor.mjs +1 -1
  21. package/bin/aos-uninstall.mjs +2 -2
  22. package/docs/master-session-acp.md +51 -0
  23. package/installer.js +398 -65
  24. package/mcps/mcsc/packages/mcp/server.js +6 -7
  25. package/package.json +4 -3
  26. package/scripts/validate-skills.mjs +76 -0
  27. package/skills/basic/bdbmediastorm/SKILL.md +1 -1
  28. package/skills/basic/godmode-shipping/SKILL.md +3 -0
  29. package/skills/basic/master-session/SKILL.md +89 -0
  30. package/skills/basic/startcycle/SKILL.md +1 -1
  31. package/skills/basic/startcycle-graph/SKILL.md +2 -2
  32. package/skills/basic/startcycle-graph-user/SKILL.md +1 -1
  33. package/skills/basic/teamwork-preview/SKILL.md +1 -1
  34. package/skills/bdbrainstorm/SKILL.md +7 -1
  35. package/skills/global_config/agentic-harness-patterns/SKILL.md +257 -0
  36. package/skills/global_config/agentic-harness-patterns/metadata.json +10 -0
  37. package/skills/global_config/agentic-harness-patterns/references/agent-orchestration-pattern.md +97 -0
  38. package/skills/global_config/agentic-harness-patterns/references/bootstrap-sequence-pattern.md +106 -0
  39. package/skills/global_config/agentic-harness-patterns/references/context-engineering/compress-pattern.md +78 -0
  40. package/skills/global_config/agentic-harness-patterns/references/context-engineering/isolate-pattern.md +82 -0
  41. package/skills/global_config/agentic-harness-patterns/references/context-engineering/select-pattern.md +86 -0
  42. package/skills/global_config/agentic-harness-patterns/references/context-engineering-pattern.md +29 -0
  43. package/skills/global_config/agentic-harness-patterns/references/hook-lifecycle-pattern.md +111 -0
  44. package/skills/global_config/agentic-harness-patterns/references/memory-persistence-pattern.md +109 -0
  45. package/skills/global_config/agentic-harness-patterns/references/permission-gate-pattern.md +111 -0
  46. package/skills/global_config/agentic-harness-patterns/references/skill-runtime-pattern.md +104 -0
  47. package/skills/global_config/agentic-harness-patterns/references/task-decomposition-pattern.md +92 -0
  48. package/skills/global_config/agentic-harness-patterns/references/tool-registry-pattern.md +101 -0
  49. package/skills/global_config/agenttrail/SKILL.md +8 -0
  50. package/skills/global_config/agenttrail/bin/agenttrail.mjs +14 -0
  51. package/skills/global_config/agenttrail/bin/ensure.mjs +118 -0
  52. package/skills/global_config/aos-setup/scripts/aos-doctor.mjs +1 -1
  53. package/skills/global_config/bdb-memb-mcp/SKILL.md +4 -5
  54. package/skills/global_config/bdb-visual-edit/SKILL.md +51 -0
  55. package/skills/global_config/bdb-visual-edit/references/vite-react-source-attr.md +59 -0
  56. package/skills/global_config/bdb-visual-edit/scripts/pick-snippet.js +27 -0
  57. package/skills/global_config/bdb-visual-edit/scripts/sanitize-element.mjs +123 -0
  58. package/skills/global_config/factory-collect/SKILL.md +74 -0
  59. package/skills/global_config/factory-human-digest/SKILL.md +92 -0
  60. package/skills/global_config/factory-lookback/SKILL.md +95 -0
  61. package/skills/global_config/factory-review-prs/SKILL.md +63 -0
  62. package/skills/global_config/git-pr-review/SKILL.md +3 -0
  63. package/skills/global_config/grilling/SKILL.md +2 -0
  64. package/skills/global_config/mcsc/SKILL.md +1 -1
  65. package/skills/global_config/plan-arbiter/SKILL.md +125 -0
  66. package/skills/global_config/plan-canvas/SKILL.md +63 -9
  67. package/skills/global_config/plan-canvas/scripts/lib/loopback-guard.js +19 -3
  68. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/README.md +285 -0
  69. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/agent-trail.js +129 -0
  70. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/board-client.js +124 -0
  71. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/demo-plan/canvas.mdx +19 -0
  72. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/demo-plan/plan.mdx +18 -0
  73. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/recap-demo/plan.mdx +72 -0
  74. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/README.md +29 -0
  75. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/architecture.json +30 -0
  76. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/builder/00_architecture.html +14950 -0
  77. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/builder/canvas.mdx +511 -0
  78. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/builder/plan.mdx +208 -0
  79. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/recap/plan.mdx +102 -0
  80. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/standard/plan.md +136 -0
  81. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/signup-storyboard/canvas.mdx +124 -0
  82. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/signup-storyboard/plan.mdx +37 -0
  83. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/index.js +188 -0
  84. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/kit.js +123 -0
  85. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/mdx.js +411 -0
  86. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/render.js +1291 -0
  87. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/architecture/meta.json +1 -0
  88. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/architecture/plan.mdx +195 -0
  89. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/architecture/standard.md +95 -0
  90. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/bugfix/meta.json +1 -0
  91. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/bugfix/plan.mdx +105 -0
  92. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/bugfix/standard.md +76 -0
  93. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/feature/canvas.mdx +81 -0
  94. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/feature/meta.json +1 -0
  95. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/feature/plan.mdx +145 -0
  96. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/feature/standard.md +76 -0
  97. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/migration/meta.json +1 -0
  98. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/migration/plan.mdx +172 -0
  99. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/migration/standard.md +100 -0
  100. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap/meta.json +1 -0
  101. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap/plan.mdx +67 -0
  102. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap/standard.md +49 -0
  103. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-board/canvas.mdx +63 -0
  104. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-board/meta.json +1 -0
  105. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-board/plan.mdx +49 -0
  106. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-board/standard.md +39 -0
  107. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-review/meta.json +1 -0
  108. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-review/plan.mdx +118 -0
  109. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-review/standard.md +57 -0
  110. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/release/meta.json +1 -0
  111. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/release/plan.mdx +173 -0
  112. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/release/standard.md +96 -0
  113. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/research/meta.json +1 -0
  114. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/research/plan.mdx +91 -0
  115. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/research/standard.md +54 -0
  116. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/show-control/canvas.mdx +53 -0
  117. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/show-control/meta.json +1 -0
  118. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/show-control/plan.mdx +225 -0
  119. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/show-control/standard.md +111 -0
  120. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/theme.css +472 -0
  121. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/trail.js +216 -0
  122. package/skills/global_config/plan-canvas/scripts/lib/plan-canvas/markdown.js +1 -1
  123. package/skills/global_config/plan-canvas/scripts/lib/plan-canvas/server.js +37 -4
  124. package/skills/global_config/plan-canvas/scripts/lib/plan-canvas/ui.js +125 -29
  125. package/skills/global_config/plan-canvas/scripts/plan-canvas.js +196 -8
  126. package/skills/global_config/pr-recap/SKILL.md +47 -0
  127. package/skills/global_config/pr-recap/scripts/pr-recap.mjs +200 -0
  128. package/skills/global_config/quick-recap/SKILL.md +55 -0
  129. package/skills/global_config/stay-within-limits/SKILL.md +85 -0
  130. package/skills/global_config/triage/SKILL.md +3 -0
  131. package/skills/global_config/visual-edit/README.md +96 -0
  132. package/skills/global_config/visual-edit/SKILL.md +615 -0
  133. package/skills/global_config/visual-plan/README.md +93 -0
  134. package/skills/global_config/visual-plan/SKILL.md +544 -0
  135. package/skills/global_config/visual-plan/references/canvas.md +139 -0
  136. package/skills/global_config/visual-plan/references/connection.md +51 -0
  137. package/skills/global_config/visual-plan/references/document-quality.md +186 -0
  138. package/skills/global_config/visual-plan/references/exemplar.md +62 -0
  139. package/skills/global_config/visual-plan/references/local-files.md +99 -0
  140. package/skills/global_config/visual-plan/references/wireframe.md +319 -0
  141. package/skills/global_config/visual-recap/README.md +103 -0
  142. package/skills/global_config/visual-recap/SKILL.md +560 -0
  143. package/skills/global_config/visual-recap/references/connection.md +51 -0
  144. package/skills/global_config/visual-recap/references/local-files.md +99 -0
  145. package/skills/global_config/visual-recap/references/wireframe.md +319 -0
  146. package/skills/playbooks/pb-ci-fix/SKILL.md +49 -0
  147. package/skills/playbooks/pb-event-tracker/SKILL.md +45 -0
  148. package/skills/playbooks/pb-meeting-actions/SKILL.md +42 -0
  149. package/skills/playbooks/pb-project-new/SKILL.md +48 -0
  150. package/skills/playbooks/pb-week-plan/SKILL.md +45 -0
  151. package/assets/header-v4.jpg +0 -0
@@ -0,0 +1,97 @@
1
+ # Agent Orchestration Pattern
2
+
3
+ ## Problem
4
+
5
+ Without deliberate orchestration structure, multi-agent systems collapse into one of two failure modes: a single overloaded agent that serializes everything and runs out of context, or an uncontrolled fan-out where every sub-agent spawns more sub-agents, producing recursive depth that is impossible to track or cancel. A third failure is "lazy delegation" — passing raw research findings to implementation workers instead of synthesizing them into precise specifications, which degrades output quality proportionally to task complexity.
6
+
7
+ These problems emerge in any agent runtime that supports delegation to sub-agents — they are not specific to a single implementation.
8
+
9
+ ## Golden Rules
10
+
11
+ ### Synthesize, don't delegate understanding
12
+
13
+ The orchestrator's job is not to forward raw findings from one worker to the next. After a research phase completes, the orchestrator must read the results, extract the relevant facts, and compose a self-contained specification for the next worker. A worker prompt that says "based on your findings" delegates understanding to a process that has no findings — it started from a blank context. Every worker prompt must be interpretable by someone who has never seen the prior conversation.
14
+
15
+ ### Choose delegation patterns deliberately; do not mix conflicting assumptions
16
+
17
+ The delegation patterns (coordinator, fork, swarm/peer) solve fundamentally different problems and make different assumptions about context sharing, depth, and lifecycle. Before mixing patterns in a single session, verify that their assumptions are compatible. Coordinator workers that start from blank context are incompatible with the context-sharing assumption of forking. Swarm peers that coordinate through shared state are incompatible with coordinator-style phased sequencing. If your runtime enforces mutual exclusion between modes, that is a valid simplification — it eliminates an entire class of ownership ambiguity at the cost of flexibility.
18
+
19
+ ### Depth must be bounded by design
20
+
21
+ Recursive delegation is the single most dangerous failure mode in agent orchestration. Every pattern must enforce a hard depth limit: coordinator workers may spawn sub-workers but the overall tree must remain tractable; fork children cannot fork again; swarm peers cannot spawn other peers. These constraints exist because unbounded depth produces exponential fan-out that is impossible to monitor, cancel, or reason about.
22
+
23
+ ### Workers get only the tools they need
24
+
25
+ Sub-agents that inherit the full tool set of the parent can take actions the orchestrator never intended — spawning their own background workers, modifying orchestration state, or accessing resources outside their task scope. Filter each worker's available tools down to what their specific task requires. Asynchronous workers need a more restricted set than synchronous ones, because they run without direct supervision.
26
+
27
+ ## When To Use
28
+
29
+ - Your agent needs to split work across multiple concurrent sub-agents.
30
+ - A task has distinct phases that must sequence (research, then synthesis, then implementation, then verification).
31
+ - You need parallel execution of independent subtasks while maintaining a single coordination point.
32
+ - A parent agent has accumulated significant context and needs to share it with workers without re-deriving it.
33
+ - Long-running peer agents need to collaborate on shared state over many turns.
34
+ - You need independent verification of implementation quality by a worker that does not share the implementer's assumptions.
35
+
36
+ ## Tradeoffs
37
+
38
+ | Decision | Benefit | Cost |
39
+ |---|---|---|
40
+ | Three distinct patterns instead of one | Each pattern is simple and well-suited to its use case | Developers must choose correctly; wrong choice degrades performance |
41
+ | Coordinator starts workers from blank context | Clean separation of concerns; no stale assumptions leak through | Orchestrator must do real synthesis work to produce self-contained prompts |
42
+ | Fork inherits full parent context | Workers immediately have all relevant background; no re-research needed | Context size is doubled per child; parent changes after fork are invisible to children |
43
+ | Single-level fork constraint | Prevents exponential fan-out; keeps the process tree shallow and cancellable | Cannot decompose fork children's work further; parent must plan granularity up front |
44
+ | Flat swarm roster | Every peer is visible and addressable; no hidden sub-teams | Coordination complexity grows linearly with team size; no hierarchical delegation |
45
+ | Mutual exclusion between modes (if enforced) | No ambiguity about who owns orchestration | Cannot combine coordinator's phased workflow with fork's context sharing in one session |
46
+ | Worker tool filtering | Workers cannot take unintended actions outside their task scope | Filtering logic must be maintained as the tool set evolves |
47
+ | Continue-vs-spawn decision | Reusing a worker preserves loaded context; spawning fresh avoids assumption leakage | Continuing with stale context is worse than the cost of re-loading from scratch |
48
+
49
+ ## Implementation Patterns
50
+
51
+ - Decide on the delegation pattern before dispatching work. If your runtime enforces mutual exclusion between modes, gate checks should reject conflicting activations explicitly rather than silently degrading.
52
+ - When using the coordinator pattern, structure work in explicit phases: research workers gather information, the coordinator synthesizes findings into specifications, implementation workers execute against those specifications, and verification workers independently confirm results.
53
+ - Write every worker prompt as a self-contained document. Include the specific file paths, exact changes required, and success criteria. Never reference "prior findings" or "the context above" — the worker has no prior context.
54
+ - After research workers report back, read their full results before dispatching implementation work. The synthesis step is where the orchestrator adds value — skipping it produces the "lazy delegation" failure.
55
+ - When forking, construct the child's initial messages so that the divergence point is as late as possible and the shared prefix is as long as possible. This maximizes cache hits across sibling workers. Only the per-child task directive should differ.
56
+ - Enforce the single-level fork constraint at invocation time, not at tool-definition time. Removing the delegation tool from a fork child's schema would change the tool set, breaking cache alignment with siblings. Instead, keep the tool present but reject the call with a clear error.
57
+ - For swarm peers, coordinate through a shared artifact (task list, file, or message bus) rather than through peer-to-peer spawning. The roster is defined at swarm creation and does not grow during execution.
58
+ - Filter each worker's tool set based on its role. Read-only research workers do not need write tools. Implementation workers do not need the ability to spawn their own background agents. Asynchronous workers get a more restrictive allow-list than synchronous ones.
59
+ - Decide whether to continue an existing worker or spawn a fresh one based on context relevance. If the worker's loaded context directly overlaps the next task, continue it. If the next task requires a different perspective (especially for verification), spawn fresh to avoid assumption leakage.
60
+ - For workers running in an isolated filesystem copy, inject a notice explaining that paths from the parent context refer to a different root. The worker must translate inherited paths to its own working directory.
61
+ - Include a purpose statement in every worker prompt so the worker can calibrate depth and scope. "This research will inform a PR description" produces different output than "this research will inform a security audit."
62
+
63
+ ## Gotchas
64
+
65
+ **Do not skip the synthesis step.** The most common coordinator failure is passing a research worker's raw output directly to an implementation worker. The research worker optimized for breadth; the implementation worker needs precise, actionable specifications. The orchestrator must bridge that gap.
66
+
67
+ **Verification workers must start fresh.** Never continue a verification task from the implementation worker's context. The implementer's loaded context carries assumptions about correctness that will blind the verifier to the exact bugs it should catch.
68
+
69
+ **Fork children cannot fork.** Do not design prompts that instruct fork children to further decompose their work via forking. The recursive guard rejects the attempt at call time, wasting a turn. Plan the decomposition granularity at the parent level.
70
+
71
+ **Identical shared prefixes are load-bearing for cache efficiency.** When forking, the placeholder content in shared message slots must be byte-identical across all siblings. Customizing per-child content in the shared prefix region destroys the cache benefit that justifies forking over coordinator-style delegation.
72
+
73
+ **Swarm peers cannot spawn other peers.** The team roster is fixed at creation. Design coordination around shared state, not dynamic team expansion. If you need a new peer, the session orchestrator adds it — peers do not recruit.
74
+
75
+ **In-process peers have lifecycle constraints.** A peer agent running inside the leader's process cannot manage independent background work, because its lifecycle is tied to the leader's. If a peer needs autonomous long-running tasks, it must run in its own process (e.g., a separate terminal session).
76
+
77
+ **Tool filtering has layers.** Different agent types may receive different tool subsets. Built-in, custom, and asynchronous agents each face their own filtering rules. Custom or less-trusted agents face additional restrictions beyond the base disallow-list. A tool's trust source — built-in, user-installed, or dynamically loaded — may determine whether it passes through all filtering layers or bypasses some.
78
+
79
+ **Be cautious when assigning weaker models to sub-tasks.** Orchestrators cannot reliably predict sub-task complexity. Routing "simple" tasks to a cheaper model is a false economy when the complexity assessment itself requires the judgment of a capable model.
80
+
81
+ **Mode checks happen at call time, not at definition time.** The orchestration mode gates do not remove tools from the schema; they reject calls when invoked under the wrong mode. This means the tool appears available but will fail. Design error handling around this — do not retry the same call expecting a different result.
82
+
83
+ ## Claude Code Evidence
84
+
85
+ Claude Code implements all three orchestration patterns as mutually exclusive modes within a single agent runtime. Only one mode can be active per session — if coordinator mode is active, forking is disabled, and vice versa. This is a deliberate simplification that eliminates ownership ambiguity at the cost of flexibility.
86
+
87
+ **Coordinator mode as a distinct operational persona.** When coordinator mode is activated, the agent receives an entirely different system prompt that replaces the default. This prompt encodes the phased workflow (research, synthesis, implementation, verification) and the "always synthesize" rule as first-class instructions. The coordinator dispatches workers as async tool calls and receives their results as structured notifications injected into the conversation. The coordinator distinguishes these notifications from real user messages by their markup structure, preventing confusion between human input and worker output.
88
+
89
+ **Fork as context-sharing with cache optimization.** Claude Code's fork implementation clones the parent's full message history into each child, then appends a synthetic turn where all tool-result slots contain identical placeholder text. Only the final directive block differs per child. This design was chosen specifically to maximize prompt cache hits — the entire shared prefix (system prompt, message history, tool definitions, and placeholder results) is byte-identical across siblings, so the inference provider can serve all children from the same cached prefix. The single-level guard keeps the delegation tool in fork children's tool schemas (preserving byte-identical tool definitions) but rejects any fork attempt at invocation time with a clear error message.
90
+
91
+ **Swarm as flat peer topology.** Claude Code's team system assigns each peer a name and runs it in a dedicated process. Peers coordinate through a shared task list rather than through message-passing or peer spawning. The runtime enforces the flat roster constraint with an explicit rejection message when a teammate attempts to create another teammate. In-process teammates face additional lifecycle restrictions — they cannot spawn background agents because their execution is coupled to the leader process — while process-isolated teammates manage their own background work independently.
92
+
93
+ **Tool filtering as defense in depth.** Claude Code applies multiple layers of tool filtering depending on agent type. All sub-agents face a base disallow-list. Custom agents (those defined by users rather than built into the runtime) face an additional restriction layer. Asynchronous agents are restricted to an explicit allow-list rather than the full filtered pool. Extension-provided tools bypass all filtering — a deliberate design choice reflecting that extensions are user-installed and trusted. This layered approach means each worker operates with the minimum tool surface required for its role.
94
+
95
+ **Model selection as a coordinator responsibility.** Claude Code's coordinator design guidance recommends against overriding the default model for individual workers, based on the observation that orchestrators cannot reliably predict sub-task complexity. Specifying a weaker model for "simple" tasks is treated as a false economy.
96
+
97
+ **Continue-vs-spawn as an explicit decision point.** The coordinator pattern surfaces the continue-or-spawn choice as a first-class decision in the orchestration flow. The coordinator can send a follow-up message to an existing worker (preserving its accumulated context) or spawn a fresh worker (starting from a clean slate). Claude Code's design guidance is explicit: continue when the existing context directly overlaps the next task, spawn fresh when it does not — especially for verification, where the implementation worker's context carries assumptions that would compromise independent review.
@@ -0,0 +1,106 @@
1
+ # Bootstrap Sequence Pattern
2
+
3
+ ## Problem
4
+
5
+ Without a disciplined bootstrap sequence, security-critical initialization steps execute out of order: TLS certificates load after the first network connection has already completed its handshake, proxy agents are not configured before the first outbound TCP connection, and trust-gated subsystems — telemetry, full environment variable sets — activate before the user has granted permission. Concurrent callers that trigger initialization simultaneously cause expensive subsystems to double-initialize, wasting hundreds of kilobytes of module loading and creating subtle state corruption. Meanwhile, trivial commands that need no subsystems at all still pay the cost of loading the entire module graph.
6
+
7
+ These problems emerge in any agent runtime with ordered dependencies between configuration, networking, trust, and observability — they are not specific to a single implementation.
8
+
9
+ ## Golden Rules
10
+
11
+ ### Initialize by dependency order, with trust as the critical inflection point
12
+
13
+ Initialization must follow the dependency graph: no subsystem should activate before the subsystems it depends on. The two universal ordering constraints are: (1) configuration parsing must complete before any subsystem reads config values, and (2) trust establishment — the point where the user grants or withholds consent — must complete before activating subsystems that depend on consent (telemetry, secret environment variables, analytics). Between config and trust, the ordering depends on your runtime's specific dependency graph. Transport setup (TLS, proxy, certificates) typically comes early because most subsystems depend on correct network state. Observability typically comes last because it depends on config, transport, and trust. The exact layer count and sequence will vary by runtime — the principle is that you must identify and respect the dependency edges, not that every runtime has exactly the same layers.
14
+
15
+ ### Memoize the top-level init to collapse concurrent callers
16
+
17
+ Wrap the entire initialization function in a memoization boundary so that all concurrent callers share a single in-flight promise. Without this, two callers invoking init simultaneously will each execute the full sequence, potentially double-configuring global HTTP agents, double-loading telemetry modules, or racing on state that the first caller has not finished writing. The memoization boundary is at the outermost async function, not at individual sub-steps.
18
+
19
+ ### Split environment variables across the trust boundary
20
+
21
+ Environment variables fall into two categories: those with no security implications (editor preferences, locale, display settings) and those that carry secrets or alter trust posture (API keys, proxy credentials, auth tokens). Apply the safe set before the trust dialog so that early subsystems can read harmless configuration. Apply the full set only after the user has granted trust. Reversing this order means env-injected credentials are active before consent is obtained.
22
+
23
+ ### Dispatch trivial commands before loading anything
24
+
25
+ Commands that require zero subsystem state — version queries, schema dumps, diagnostic flags — should be dispatched from the entry point before any dynamic imports or initialization runs. This keeps their latency at the floor of process startup and prevents the module graph from loading hundreds of kilobytes of code that will never execute.
26
+
27
+ ### Register cleanup during init, not at usage sites
28
+
29
+ Every cleanup handler — graceful shutdown, transport teardown, child-process reaping, session cleanup — must be registered during initialization, not scattered across usage sites. Registering at init guarantees that cleanup runs unconditionally on exit, regardless of which code path was taken. Cleanup handlers registered at usage sites are silently skipped when an unexpected exit path bypasses that site.
30
+
31
+ ## When To Use
32
+
33
+ - Your agent runtime has security-critical ordering between TLS, proxy, and network subsystems.
34
+ - You support multiple entry modes (CLI, server, SDK, headless) that must share the same initialization path.
35
+ - You need a trust boundary that separates pre-consent and post-consent configuration.
36
+ - Expensive subsystems (telemetry, tracing, observability) should load only when permitted and needed.
37
+ - Trivial commands should respond instantly without paying the cost of full initialization.
38
+ - Multiple callers or code paths may trigger initialization concurrently.
39
+
40
+ ## Tradeoffs
41
+
42
+ | Decision | Benefit | Cost |
43
+ |---|---|---|
44
+ | Dependency-ordered sequential init | Security-critical steps cannot execute out of order | Slower cold start than parallel init; each layer blocks on its dependencies |
45
+ | Memoized top-level init | Concurrent callers share one promise; no double-init | Memoize caches rejections in some libraries, blocking retry on transient errors |
46
+ | Trust-split environment variables | Secrets are never active before user consent | Two separate application passes to maintain; easy to put a variable in the wrong set |
47
+ | Fast-path dispatch for trivial commands | Sub-millisecond response for version/help/diagnostics | Each fast path is a special case that must be maintained outside the normal init flow |
48
+ | Lazy dynamic imports for heavy modules | Hundreds of kilobytes deferred until actually needed | Import latency shifts to first use; harder to reason about load order |
49
+ | Parallel fire-and-forget for non-critical work | Background caches populate without blocking the critical path | Failures are silent; callers must tolerate cache misses gracefully |
50
+ | Build-time feature-flag gating | Dead code for disabled modes is eliminated from the shipped artifact | Feature checks must appear inline at call sites; wrapping in helpers defeats the bundler |
51
+ | Transport as a narrow interface | Wire protocol versions are swappable without touching callers | Behavioral differences between protocol versions can hide behind the same interface |
52
+
53
+ ## Implementation Patterns
54
+
55
+ - Parse and validate all configuration files as the very first step. Surface parse errors as user-visible diagnostics before any network contact occurs.
56
+ - Apply TLS certificate configuration immediately after config validation and before instantiating any HTTP client or agent. Many runtimes cache the TLS certificate store at process boot; certificates applied after the first handshake may have no effect on connections that reuse the cached pool.
57
+ - Configure mutual TLS before proxy agents, and proxy agents before any outbound connection including preconnect or warm-up calls. The dependency chain is: certificates, then mTLS, then agents, then connections.
58
+ - Apply the safe subset of environment variables (no secrets, no auth tokens) before the trust dialog. Defer the full set until after trust is established.
59
+ - Wrap the entire top-level async initialization function in a memoization boundary. Concurrent callers must receive the same promise, not trigger parallel execution of the same sequence.
60
+ - Fire non-critical background work (repository detection, IDE detection, cache population, analytics preamble) after the transport layer is configured but before blocking on trust. These tasks should not be awaited in the critical path.
61
+ - Gate telemetry initialization behind trust establishment. Use a separate boolean guard distinct from the init memoization so that telemetry can be retried independently if the first attempt fails.
62
+ - When the runtime supports managed remote settings, wait for those settings to load before finalizing telemetry configuration, so that organization-configured endpoints are honored.
63
+ - Use dynamic imports for heavy dependency trees (telemetry, tracing, protocol buffers, gRPC) so that their module weight is deferred until the subsystem is actually initialized.
64
+ - At the entry point, check for trivial commands (version, help, schema dump) and dispatch them immediately with zero dynamic imports — do not load the full module graph for commands that will never use it.
65
+ - Support multiple entry modes (interactive CLI, server/daemon, SDK embedding, headless) through the same initialization path, branching only at the mode-dispatch layer above init. For headless modes, skip interactive dialogs and write errors to stderr with a non-zero exit code.
66
+ - Use build-time feature flags for mode-specific branches. The flag check must appear inline at the call site so the bundler can perform dead-code elimination; wrapping it in a helper function defeats this optimization.
67
+ - Define transport as a narrow interface — write, batch-write, connect, close, state-report — so that different wire protocol versions are interchangeable behind a single surface.
68
+ - Register all cleanup handlers (graceful shutdown, transport teardown, child-process cleanup, session finalization) during initialization, not at individual usage sites.
69
+
70
+ ## Gotchas
71
+
72
+ **Memoization hides retry failures.** Some memoize implementations cache the first returned promise, including rejected ones. If a transient config error causes init to reject, every subsequent caller receives the cached rejection forever. Either use a memoize variant that clears on rejection, or implement a manual once-guard that resets its flag on error.
73
+
74
+ **Trust boundary leakage through environment variables.** Applying the full environment variable set before the trust dialog means env-injected API keys or proxy credentials are active before the user consents. This is a security violation, not just an ordering preference. Always apply the safe pass first and the full pass only after explicit trust grant.
75
+
76
+ **TLS certificate store caching defeats late application.** On many runtimes, the TLS certificate store is read once at process boot. Applying custom CA certificates after even one TLS handshake has completed may have no effect on connections that reuse the cached certificate pool. CA certificate configuration must be the very first network-adjacent step, before any connection — including preconnect warm-ups.
77
+
78
+ **Telemetry double-init across async boundaries.** When multiple code paths can trigger telemetry initialization (eager synchronous init for non-interactive sessions, deferred async init after remote settings load), the guard flag must be set before the async initializer resolves, not after. Setting it after the resolution leaves a window where a second caller can slip through and double-initialize.
79
+
80
+ **Feature flags must be inline for dead-code elimination.** Build-time feature flags only enable dead-code elimination when the flag check appears directly at the call site. Extracting the check into a helper function, even a trivially inlined one, defeats the bundler's static analysis and ships the gated code in all builds.
81
+
82
+ **Transport protocol version differences hide behind the interface.** When a transport interface abstracts multiple wire protocol versions, behavioral differences (such as one version supporting dropped-batch counts and another always returning zero) are invisible to callers. Callers that depend on version-specific behavior must not rely on the abstracted interface for those checks.
83
+
84
+ **Global state modules must be leaves in the import graph.** Any module that holds global singleton state must import only pure leaf modules with no transitive dependencies on the rest of the system. Adding a non-leaf import to a global state module creates circular-dependency risks that cascade across the entire bootstrap graph.
85
+
86
+ **Headless modes must not block on interactive dialogs.** When the runtime detects a non-interactive session (no TTY, piped input, CI environment), initialization must skip any interactive trust dialog or prompt. Blocking on user input in a headless context hangs the process indefinitely.
87
+
88
+ ## Claude Code Evidence
89
+
90
+ Claude Code's bootstrap sequence is a production implementation of these principles, supporting seven distinct entry modes through a single ordered initialization path.
91
+
92
+ **Four-layer ordering with explicit dependency comments.** The initialization function enforces strict ordering: config parsing first, then safe environment variable application, then TLS CA certificate configuration, then graceful-shutdown registration, then mTLS configuration, then global HTTP agent setup, then API preconnection. Each step includes a comment explaining why it must precede the next. The TLS step carries a specific note about the runtime caching its certificate store at boot, making late application ineffective.
93
+
94
+ **Memoized init collapsing concurrent callers.** The top-level initialization function is wrapped in a single memoization boundary, ensuring that regardless of how many code paths trigger init simultaneously, the sequence executes exactly once and all callers share the same promise.
95
+
96
+ **Trust-split environment variables.** Environment variable application is split into two distinct passes. The "safe" pass runs before any trust check, applying only variables with no security implications. The "full" pass runs only after the trust dialog completes, applying variables that may carry API keys, proxy credentials, or other sensitive configuration.
97
+
98
+ **Fast-path dispatch for trivial commands.** The CLI entry point checks for version flags, schema dumps, and diagnostic commands before loading any modules. These paths respond with zero dynamic imports and exit immediately, keeping their latency at the floor of process startup.
99
+
100
+ **Lazy loading of heavy subsystems.** Telemetry initialization uses dynamic imports to defer approximately 400KB of observability and protocol-buffer modules until telemetry is actually needed. A further layer of lazy loading defers approximately 700KB of gRPC transport modules until tracing instrumentation is activated. This two-tier lazy loading keeps the critical path fast for sessions that never activate tracing.
101
+
102
+ **Parallel fire-and-forget for background caches.** After the transport layer is configured but before blocking on trust, the bootstrap fires off non-critical background tasks — IDE detection, repository detection, OAuth cache population, analytics preamble — as unawaited promises. These populate shared caches that later synchronous callers read, without blocking the critical initialization path.
103
+
104
+ **Multi-mode entry through a single init path.** Seven entry modes — interactive REPL, server/daemon, SDK embedding, bridge/remote-control, background sessions, template runner, and self-hosted runner — all flow through the same memoized initialization function. Mode-specific behavior branches above the init call, not within it. Feature-gated modes use build-time flags with inline checks so that unused modes are eliminated from the shipped artifact.
105
+
106
+ **Transport abstraction for swappable wire protocols.** The bridge transport is defined as a narrow interface with write, batch-write, connect, close, and state-reporting methods. Two protocol versions — one using WebSocket plus POST, the other using server-sent events for reads and a separate write channel — are interchangeable behind this surface. The interface documents where behavioral differences exist between versions.
@@ -0,0 +1,78 @@
1
+ # Context Compression and Snapshot Management
2
+
3
+ ## Problem
4
+
5
+ Long-running agent sessions inevitably exhaust the context window. Each turn adds tool calls, results, and reasoning, and none of it shrinks on its own. Without compression, the agent either hits the hard token limit and fails, or the model's attention degrades as useful information is buried under pages of stale output. Variable-length context blocks — directory listings, git status, search results — make the problem worse because their size is unpredictable and can spike by orders of magnitude on a single turn.
6
+
7
+ A second, subtler problem: context that was accurate when captured becomes misleading as the session progresses. A git status from turn 3 is treated as current truth at turn 30 unless something marks it as a point-in-time snapshot.
8
+
9
+ These problems affect any agent runtime that persists across turns. They are not specific to any single implementation.
10
+
11
+ ## Golden Rules
12
+
13
+ ### Truncate with recovery, never truncate blind
14
+
15
+ When a context block exceeds a character or token cap, truncate it — but always append a recovery pointer: a specific instruction telling the model which tool to call, with which arguments, to retrieve the full output. Truncation without a recovery pointer is a dead end. The model knows the data exists, knows it was cut off, and has no path back. This produces hallucinated completions or silent failures.
16
+
17
+ ### Compact reactively, not on a schedule
18
+
19
+ Fixed-window compaction (e.g., "summarize every 20 turns") either fires too early and loses useful context or too late and hits the token ceiling. Reactive compaction monitors the actual fill ratio and fires when the window is genuinely full. The compaction pass summarizes older turns while preserving recent ones and active tool results, recovering budget mid-session without losing the thread of work.
20
+
21
+ ### Label every snapshot
22
+
23
+ Any context block that captures point-in-time state — git status, directory listings, process output, file contents — must carry an explicit label: the time of capture and a note that the data will not auto-update. Without this label, the model treats stale data as current and produces reasoning that contradicts the actual state of the world.
24
+
25
+ ### Cap every variable-length block
26
+
27
+ Every context block whose length depends on external state must have a hard character or token ceiling. Directory listings, search results, git diffs, log output — all of these can spike unpredictably. Without a cap, a single large result can consume the majority of the context budget in one turn, crowding out everything else.
28
+
29
+ ## When To Use
30
+
31
+ - Agent performance degrades noticeably as sessions get longer.
32
+ - The agent repeatedly acts on stale data (old git status, outdated file contents, stale search results).
33
+ - Variable-length context blocks occasionally spike and crowd out other information.
34
+ - Sessions are long enough that the context window approaches its token limit.
35
+ - The agent hallucinates completions of truncated output instead of re-fetching.
36
+
37
+ ## Tradeoffs
38
+
39
+ | Decision | Benefit | Cost |
40
+ |---|---|---|
41
+ | Truncation with recovery pointers | Model can always retrieve full data | Extra round-trip when full data is needed |
42
+ | Reactive compaction | Extends effective session length dynamically | Older context is lossy-compressed, not losslessly preserved |
43
+ | Snapshot labeling | Prevents stale-data reasoning | Every context injection site must add the label |
44
+ | Hard character caps on variable blocks | No single block can starve the budget | Useful data may be cut; recovery pointer quality matters |
45
+ | Preserving recent turns during compaction | Model retains working memory of current task | Compaction logic must distinguish "recent" from "old" |
46
+
47
+ ## Implementation Patterns
48
+
49
+ - Enforce a character or token cap on every variable-length context block. The cap should be set conservatively — it is easier to raise a cap that is too tight than to debug a session where one block consumed the entire budget.
50
+ - Append a truncation notice that names the specific tool and invocation needed to fetch the full data. The notice should be concrete: "Run [tool] with [these arguments] to see the complete output," not a vague "output was truncated."
51
+ - Tag every injected snapshot with a capture timestamp and a note that the data is static. Place the label at the top of the block, not the bottom — the model processes tokens sequentially and needs the staleness warning before it starts reasoning about the content.
52
+ - Implement a fill-ratio monitor that triggers compaction when the context window reaches a threshold (e.g., 80% full). The compaction pass should summarize older turns into a condensed narrative while preserving the most recent turns and any active tool results verbatim.
53
+ - During compaction, preserve any recovery instructions that were part of truncated blocks. If a truncation notice is compacted away, the model loses its path back to the full data.
54
+ - Include explicit instructions in the compaction output telling the model what was summarized and how to re-fetch details if needed. The model should know that compaction happened and what was lost.
55
+
56
+ ## Gotchas
57
+
58
+ **Truncation without a recovery pointer is a dead end.** If you cap a context block and do not tell the model which tool to call for the full output, the model has no path back to the data. It will either hallucinate the missing content or silently drop the task that depended on it.
59
+
60
+ **Snapshot labels prevent stale reasoning.** A git status injected without a "this is a snapshot taken at time T" label causes the model to treat the data as live. It will produce branch, diff, and merge reasoning that contradicts the actual current state.
61
+
62
+ **Compaction must preserve recovery pointers.** If a truncation notice from an earlier turn is itself summarized away during compaction, the model loses the ability to re-fetch the full data. Recovery pointers should survive compaction passes.
63
+
64
+ **Variable-length blocks can spike by orders of magnitude.** A directory listing in a small project is 20 lines; in a monorepo it can be 20,000. A git diff is usually small; after a large refactor it can be enormous. Caps must be set for the worst case, not the common case.
65
+
66
+ **Reactive compaction must not fire during active tool use.** If compaction triggers while a multi-step tool sequence is in progress, it may summarize away intermediate results that the model needs for the next step. Gate compaction on a quiescent state — between turns, not mid-turn.
67
+
68
+ ## Claude Code Evidence
69
+
70
+ Claude Code's compression strategy demonstrates these principles in a long-running interactive session environment:
71
+
72
+ **Git status as a labeled, capped snapshot.** The git status string injected into context carries an explicit note: "this status is a snapshot in time, and will not update during the conversation." When the output exceeds a 2000-character threshold, it is truncated with a message directing the model to run a specific tool to see the full output. This pairing of label plus recovery pointer means the model always knows the data is stale and always has a path to fresh data.
73
+
74
+ **Reactive compaction for extended sessions.** When the context window fills during a long session, the runtime triggers a compaction pass that summarizes older conversation turns while preserving the most recent exchanges and any active tool results. The model is given explicit context about what was summarized. This approach extends effective session length without a hard turn limit — sessions can run for hours as long as compaction can recover enough budget on each pass.
75
+
76
+ **Character caps on all variable-length injections.** Every context block whose size depends on external state — git status, directory listings, file contents — has a hard character ceiling. The design philosophy is that no single context source should be able to starve the budget for all others. When a cap is hit, the truncation notice always includes the specific tool invocation needed to retrieve the complete output, ensuring the model is never left at a dead end.
77
+
78
+ **Recovery instructions as first-class content.** Truncation notices in Claude Code are not afterthoughts or generic warnings. Each one names the exact tool and arguments the model should use. This specificity matters: a vague "output was truncated" message forces the model to guess how to recover, while a concrete instruction lets it act immediately.
@@ -0,0 +1,82 @@
1
+ # Context Isolation for Delegated Work
2
+
3
+ ## Problem
4
+
5
+ When an agent delegates work to sub-agents, every piece of shared state becomes a potential collision point. A sub-agent that inherits the parent's full context may act on stale assumptions; a sub-agent that shares the parent's working directory may overwrite files the parent is reading; a sub-agent that can recursively spawn its own children creates exponential context multiplication. Without explicit isolation boundaries, the blast radius of a single sub-agent mistake extends to the entire session.
6
+
7
+ These problems emerge in any agent runtime that supports delegation — coordinator patterns, fork patterns, or any form of parallel sub-agent work. They are inherent to multi-agent architectures, not specific to any single implementation.
8
+
9
+ ## Golden Rules
10
+
11
+ ### Zero-inheritance is the safest default
12
+
13
+ When delegating work to a sub-agent in a coordinator pattern, pass only the explicit prompt. The sub-agent receives no conversation history, no accumulated context, no session state. This is the narrowest possible boundary: the sub-agent can only act on what it was told, and nothing from the parent's session leaks accidentally. The cost is that the prompt must be self-sufficient — everything the sub-agent needs to do its job must be spelled out.
14
+
15
+ ### Full-inheritance forks must be single-level
16
+
17
+ When a sub-agent needs the parent's full context (the fork pattern), copy the context but enforce a single-level boundary: the forked agent cannot itself fork. Without this constraint, recursive forking multiplies context cost exponentially — each level copies everything from the level above, and the total cost is O(2^n) in the depth of the fork tree. A single-level boundary keeps the cost linear.
18
+
19
+ ### Filesystem isolation prevents path collisions
20
+
21
+ When a sub-agent modifies files, it needs its own working copy of the repository. Shared working directories create race conditions: the parent reads a file, the sub-agent modifies it, and the parent's next read returns unexpected content. Filesystem isolation (via worktrees, temp directories, or copy-on-write clones) gives each agent its own view. Path translations must be injected so the agent operates on its isolated copy, not the shared original.
22
+
23
+ ### Choose the narrowest boundary that works
24
+
25
+ The isolation boundary determines the blast radius of a sub-agent's mistakes. A zero-inheritance worker can corrupt only its own output. A full-inheritance fork can produce results inconsistent with the parent's state. A shared-filesystem agent can corrupt the parent's working directory. Always start with the narrowest boundary and widen only when the sub-agent genuinely needs more context.
26
+
27
+ ## When To Use
28
+
29
+ - Your agent delegates work to sub-agents or spawns concurrent workers.
30
+ - Sub-agents modify files in the same repository the parent is working in.
31
+ - Delegated sub-agents produce results that contradict the parent's state or each other.
32
+ - You need to prevent a sub-agent failure from corrupting the parent session.
33
+ - Parallel agents need to work on the same codebase without stepping on each other's changes.
34
+
35
+ ## Tradeoffs
36
+
37
+ | Decision | Benefit | Cost |
38
+ |---|---|---|
39
+ | Zero-inheritance delegation | Clean context boundaries, no accidental leakage | Sub-agent must be self-sufficient from its prompt alone |
40
+ | Full-inheritance fork (single-level) | Sub-agent has full context for complex tasks | Doubles context cost; must enforce no-recursive-fork rule |
41
+ | Filesystem isolation via worktrees | No file-level race conditions between agents | Worktree creation and teardown overhead; disk cost |
42
+ | Path translation injection | Agent operates on its own copy transparently | Translation logic must cover all file-operation tools |
43
+ | Immutable shared state | Safe concurrent reads across agents | State updates require new snapshots, not in-place mutation |
44
+ | Single-level fork boundary | Linear cost, not exponential | Forked agents that need further delegation must use coordinator pattern instead |
45
+
46
+ ## Implementation Patterns
47
+
48
+ - For coordinator-pattern delegation, construct the sub-agent prompt as a self-contained document. Include all necessary context, constraints, and output format requirements in the prompt itself — do not rely on inherited session state.
49
+ - For fork-pattern delegation, copy the parent's context at fork time and enforce a single-level boundary. The forked agent should be unable to invoke the fork mechanism itself. If a forked agent needs further delegation, it must use the coordinator pattern (zero-inheritance) for its own sub-agents.
50
+ - When sub-agents modify files, create an isolated filesystem context (worktree, temp directory, or copy-on-write clone) before the agent starts. Inject path translations so every file-operation tool targets the isolated copy.
51
+ - On sub-agent completion, merge results back to the parent's working directory through a controlled integration point — not by letting the sub-agent write directly to the shared filesystem.
52
+ - Mark all state shared between parent and sub-agents as immutable. If a sub-agent needs to communicate state changes, it should return them as a result, not mutate shared objects in place.
53
+ - Register cleanup handlers that tear down isolated filesystems when the sub-agent completes or fails. Leaked worktrees or temp directories accumulate and waste disk.
54
+ - When multiple sub-agents work in parallel on the same codebase, assign non-overlapping file sets where possible. Filesystem isolation handles the general case, but non-overlapping assignments prevent merge conflicts at the integration point.
55
+
56
+ ## Gotchas
57
+
58
+ **Recursive forks create exponential cost.** If a forked agent can itself fork, the context cost doubles at each level. Three levels of recursive forking mean eight copies of the original context. Enforce the single-level boundary at the delegation mechanism, not as a convention.
59
+
60
+ **Path translations must cover all file-operation tools.** If the isolation mechanism injects a worktree path but one file-operation tool bypasses the translation, that tool writes to the shared original directory. Every tool that touches the filesystem must go through the translation layer.
61
+
62
+ **Zero-inheritance prompts must be truly self-contained.** A common failure mode is constructing a sub-agent prompt that implicitly relies on context the parent has but the sub-agent does not. Review zero-inheritance prompts for hidden dependencies: file paths that assume a specific working directory, references to earlier conversation turns, or assumptions about available tools.
63
+
64
+ **Merging results from isolated agents can conflict.** When two parallel agents modify the same file in their respective worktrees, merging both back produces a conflict. Design task assignments to minimize overlap, and have a conflict resolution strategy for when overlap is unavoidable.
65
+
66
+ **Worktree cleanup must happen on all exit paths.** If cleanup only runs on successful completion, a failed or killed sub-agent leaks its worktree. Register cleanup on the process-level shutdown handler, not just in the success path.
67
+
68
+ **Immutable shared state does not mean no communication.** Sub-agents can still return results to the parent — the constraint is that they return values rather than mutating shared objects. The parent integrates returned results into its own state through a controlled update path.
69
+
70
+ ## Claude Code Evidence
71
+
72
+ Claude Code's delegation system implements these isolation principles across multiple agent coordination patterns:
73
+
74
+ **Zero-inheritance as the default for coordinated work.** When the runtime dispatches a sub-agent in the coordinator pattern, the sub-agent receives only its explicit task prompt. No conversation history, no accumulated memory, no parent session state is inherited. This is a deliberate design choice: the sub-agent's blast radius is limited to its own output. The cost — that every sub-agent prompt must be fully self-contained — is accepted as worthwhile because it eliminates an entire class of state-leakage bugs.
75
+
76
+ **Single-level fork boundary.** The fork mechanism copies the parent's full conversation context to the forked agent but enforces that the forked agent cannot itself invoke the fork mechanism. If a forked agent needs to delegate further, it must use the coordinator pattern (zero-inheritance) for its own sub-tasks. This keeps the total context cost at 2x the parent's context, not 2^n. The design explicitly rejects recursive forking because the exponential cost makes it impractical for any real workload.
77
+
78
+ **Worktree-based filesystem isolation.** When a sub-agent needs to modify files, the runtime creates a git worktree that gives the agent its own working copy of the repository. Path translations are injected so that every file-operation tool targets the worktree, not the parent's working directory. This eliminates file-level race conditions between concurrent agents. Worktrees are cleaned up on all exit paths, including failure and kill, through process-level cleanup handlers.
79
+
80
+ **Immutable shared state across concurrent agents.** The application state shared between the parent and its sub-agents is wrapped in a deep-immutability type. Sub-agents cannot mutate this state in place. When a sub-agent produces results, they are returned as values and integrated by the parent through a controlled update path. This design prevents the most insidious class of concurrency bugs: silent in-place mutation of shared objects that causes different agents to see inconsistent state.
81
+
82
+ **Blast radius as the primary design criterion.** The isolation level for each delegation type was chosen by asking "what is the worst thing this sub-agent can break?" Zero-inheritance workers can only produce bad output. Full-inheritance forks can produce output inconsistent with the parent's context. Shared-filesystem agents can corrupt the parent's working directory. The runtime defaults to the narrowest boundary and only widens when the use case demands it.
@@ -0,0 +1,86 @@
1
+ # Context Selection and Progressive Disclosure
2
+
3
+ ## Problem
4
+
5
+ Agent runtimes that eagerly load all available context on every turn pay a compounding tax: token costs grow with catalog size, startup latency scales with the number of information sources, and concurrent agents duplicate the same expensive I/O. The model receives thousands of tokens it never reads while the data it actually needs arrives late or not at all. Without deliberate selection, the context window becomes a landfill — technically full, practically empty.
6
+
7
+ These problems appear in any agent that maintains a growing skill catalog, reads from multiple persistent stores, or serves concurrent requests. They are not specific to any single runtime.
8
+
9
+ ## Golden Rules
10
+
11
+ ### Lazy beats eager
12
+
13
+ Load context at the moment it becomes relevant, not at the moment the session starts. Eagerly injecting everything "just in case" means every API call pays for tokens the model may never use. Deferring the load means you pay only when the model actually activates a capability. The one-time latency of a just-in-time fetch is almost always cheaper than the per-turn cost of a bloated prompt.
14
+
15
+ ### Three-tier progressive disclosure
16
+
17
+ Structure every context source into three tiers of increasing cost:
18
+
19
+ 1. **Metadata** (~100 tokens per item): names, descriptions, and trigger hints. Always present. This is the "is this relevant?" signal.
20
+ 2. **Instructions** (bounded, e.g. < 5000 tokens): the full body of the activated capability. Loaded only on activation.
21
+ 3. **Resources** (unbounded): reference files, scripts, data assets. Loaded only when the agent explicitly requests them.
22
+
23
+ The discovery cost (tier 1) scales with catalog size, but the execution cost (tiers 2-3) is constant per activated item. Keeping discovery cheap lets you maintain a large catalog without paying for it on every turn.
24
+
25
+ ### Memoize the promise, not the result
26
+
27
+ When multiple callers request the same expensive context simultaneously, memoize the in-flight promise so all callers share a single I/O operation. Memoizing only the resolved value creates a window where concurrent callers each see the cache as empty and launch duplicate work. The promise itself is the deduplication key.
28
+
29
+ ### Invalidate at the mutation site, not on a timer
30
+
31
+ Cache expiration based on time or reactive subscriptions either serves stale data or rebuilds too often. Instead, place explicit invalidation calls at each known mutation point. This is more work to maintain — every new write path must add its own invalidation — but it guarantees that the cache is exactly as fresh as the most recent deliberate change.
32
+
33
+ ## When To Use
34
+
35
+ - The skill or capability catalog is growing and the always-on prompt is becoming expensive.
36
+ - Startup latency is high because context is built eagerly on every turn.
37
+ - Multiple concurrent agents or invocations duplicate the same expensive I/O.
38
+ - You need to support a large catalog without linearly increasing per-turn token cost.
39
+ - Context sources span multiple directories, stores, or services that may overlap.
40
+
41
+ ## Tradeoffs
42
+
43
+ | Decision | Benefit | Cost |
44
+ |---|---|---|
45
+ | Lazy loading | Low idle token cost | One round-trip latency on first activation |
46
+ | Progressive disclosure (3 tiers) | Large catalogs stay cheap | Model cannot reason about unactivated capabilities |
47
+ | Promise memoization | No duplicate I/O under concurrency | Cache logic is more subtle than simple value caching |
48
+ | Manual invalidation at mutation sites | Only rebuilds when necessary | Every new write path must include its own invalidation call |
49
+ | Token budget for discovery tier | Catalog cost is bounded and predictable | Each entry must fit within a tight character cap |
50
+ | Path-conditional activation | Domain-specific skills stay dormant until relevant | Activation state must survive process restarts or be re-evaluated |
51
+
52
+ ## Implementation Patterns
53
+
54
+ - Wrap every expensive context builder in a memoized async function. Store the promise itself so concurrent callers share the same in-flight request rather than spawning duplicate I/O.
55
+ - Expose a named cache-clear call for each memoized builder. Invoke it only from known mutation handlers — never from a general observer or polling loop.
56
+ - For capability catalogs, compute an estimated token count from metadata fields only (name, description, trigger hint). Store the full body in a deferred closure, fetched only on activation.
57
+ - Cap the total discovery-tier budget at a fixed fraction of the context window (e.g., ~1%), with each entry limited to a tight character ceiling.
58
+ - Use path-condition metadata to gate domain-specific capabilities: a capability stays dormant until a matching file or directory is touched in the session.
59
+ - Deduplicate multi-source loads by canonical path. Resolve symlinks and normalize before inserting into the active set — string comparison of original paths produces false negatives, and inode comparison is unreliable on virtual or network filesystems.
60
+ - Mark shared state as deeply immutable. Document any fields excluded from the immutability wrapper (e.g., function-typed fields) at the point of exclusion, not silently.
61
+
62
+ ## Gotchas
63
+
64
+ **Stale context after mutation.** Because the cache never expires automatically, any external change (new file, toggled flag, updated store) is invisible until the cache is explicitly cleared. Every mutation point must call the corresponding invalidation. Missing one means the model operates on stale data for the rest of the session.
65
+
66
+ **Concurrent initialization races without promise memoization.** If two invocations trigger the same extraction simultaneously and both see the cache as empty, both will attempt the same I/O. Only one will succeed cleanly; the other may fail or produce corrupt output. Memoizing the promise eliminates the race entirely.
67
+
68
+ **Premature disclosure inflates every turn.** Injecting full capability bodies into the system prompt means every API call pays for tokens the model may never use. With 20+ capabilities, this can consume a significant share of the context budget before the first user message arrives.
69
+
70
+ **Conditional capabilities must track activation state.** Capabilities gated by path conditions are dormant until a matching file is touched. If activation state is lost (e.g., process restart), the capability must be re-evaluated against current context — otherwise it silently disappears from the session.
71
+
72
+ **Deduplication must use canonical paths, not original paths.** The same file can appear under different paths when symlinks, mounts, or multiple search directories are involved. String comparison of the original paths will miss duplicates; resolve to canonical paths before comparing.
73
+
74
+ ## Claude Code Evidence
75
+
76
+ Claude Code's context system is a production implementation of these principles, managing a growing catalog of skills across multiple source directories:
77
+
78
+ **Memoize-plus-manual-invalidate, not reactive subscriptions.** Both the system context builder and the user context builder are memoized async functions. The cache persists for the entire process lifetime. When something must change — for instance, a system prompt injection is updated — the runtime calls cache-clear explicitly at that specific mutation point. This is a deliberate choice: context is cheap to serve from cache and expensive to rebuild. A reactive model that rebuilt on every state update would defeat the purpose of caching entirely.
79
+
80
+ **Promise memoization for concurrent initialization.** The bundled skills extraction stores the extraction promise itself rather than the resolved value. If two skill invocations race during startup, both await the same promise. Without this design, both would attempt to write the same extracted files concurrently, and the second would fail or corrupt the output.
81
+
82
+ **Token count as the gating metric for progressive disclosure.** The skill loader computes an estimated token count from only the name, description, and trigger hint fields — never from the full markdown body. The body is stored in a deferred closure and fetched only at invocation time. This cleanly separates the "is this relevant?" signal (cheap, always present) from the "what does this skill do?" content (expensive, deferred). The total listing budget is capped at roughly one percent of the context window, with each entry limited to approximately 250 characters.
83
+
84
+ **Realpath deduplication across multiple source directories.** When loading skills from managed, user, project, and additional directories, the same skill file can appear under symlinked paths. The loader resolves every file to its canonical path before inserting into a seen-set. The team chose canonical-path comparison over inode comparison because inodes are unreliable on virtual and network filesystems.
85
+
86
+ **Deeply immutable shared state.** The main application state is wrapped in a deep-immutability type to prevent accidental in-place mutation across concurrent agents. Task state is excluded from the immutability wrapper because it contains function types that the wrapper cannot handle — but the exclusion is documented at the point of exclusion, not silently dropped.
@@ -0,0 +1,29 @@
1
+ # Context Engineering Pattern
2
+
3
+ ## The Problem
4
+
5
+ Without deliberate management, context windows fill with redundant or stale data on every turn. Rebuilding expensive context on each invocation adds latency. Exposing the full capability catalog up front bloats every prompt. Delegated sub-agents pollute their parent's context with intermediate state. These are inherent to any agent that persists across turns and delegates to sub-agents.
6
+
7
+ ## The Four-Axis Framework
8
+
9
+ Context engineering decomposes into four concerns:
10
+
11
+ 1. **Select** -- decide what enters the context window, when, and at what granularity.
12
+ 2. **Write** -- the agent does not just consume context; it writes back to persistent storage, creating the learning loop.
13
+ 3. **Compress** -- long sessions exhaust the window; truncation, compaction, and snapshot labeling recover budget.
14
+ 4. **Isolate** -- delegated work must not pollute or corrupt the parent's context or filesystem.
15
+
16
+ ### A note on "Write"
17
+
18
+ The write axis is a cross-cutting principle rather than a standalone pattern. Every axis involves write-back: selections are cached and invalidated, compression produces summaries that are stored, isolation boundaries return results that the parent integrates. The write-back loop -- auto-memory, session extraction, persisted permissions, task state updates -- is what transforms a stateless tool-caller into a learning system. It is covered as a design principle in the main skill definition rather than a separate sub-file.
19
+
20
+ ## Sub-File Index
21
+
22
+ ### [Select: Context Selection and Progressive Disclosure](context-engineering/select-pattern.md)
23
+ Lazy loading, three-tier progressive disclosure, promise memoization for concurrency, manual cache invalidation, token budgeting for capability discovery, canonical-path deduplication, and path-conditional activation.
24
+
25
+ ### [Compress: Context Compression and Snapshot Management](context-engineering/compress-pattern.md)
26
+ Truncation with recovery pointers, reactive compaction triggered by fill ratio, snapshot labeling with staleness warnings, character and token caps on variable-length blocks, and recovery-pointer preservation across compaction passes.
27
+
28
+ ### [Isolate: Context Isolation for Delegated Work](context-engineering/isolate-pattern.md)
29
+ Zero-inheritance workers, full-inheritance forks with single-level boundaries, filesystem isolation via worktrees, path translation injection, blast radius as the primary design criterion, and immutable shared state across concurrent agents.