@hybridlabor-api/aos 4.13.2 → 4.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (150) hide show
  1. package/.agents/AGENTS.md +8 -0
  2. package/.agents/nodes.json +5 -2
  3. package/.claude/hooks/conventional-commits.mjs +14 -15
  4. package/.claude/hooks/env-file-protection.mjs +14 -15
  5. package/.claude/hooks/go-gate.mjs +152 -10
  6. package/.claude/hooks/go-token.mjs +55 -0
  7. package/.claude/hooks/memb-inject.mjs +75 -62
  8. package/.claude/hooks/trail-autostart.mjs +27 -0
  9. package/.claude/settings.json +18 -0
  10. package/.claude/workflows/startcycle-dispatch.mjs +11 -4
  11. package/.opencode/plugins/bdb-aos.js +98 -121
  12. package/.opencode/plugins/lib/trail-autostart.js +38 -0
  13. package/CLAUDE.md +1 -1
  14. package/README.de.md +6 -6
  15. package/README.md +6 -6
  16. package/README.pt.md +6 -6
  17. package/THIRD_PARTY_NOTICES.md +19 -3
  18. package/assets/header-v5.png +0 -0
  19. package/bin/aos-acp.mjs +211 -0
  20. package/bin/aos-doctor.mjs +1 -1
  21. package/bin/aos-uninstall.mjs +2 -2
  22. package/docs/master-session-acp.md +51 -0
  23. package/installer.js +252 -35
  24. package/mcps/mcsc/packages/mcp/server.js +6 -7
  25. package/package.json +4 -3
  26. package/scripts/validate-skills.mjs +76 -0
  27. package/skills/basic/bdbmediastorm/SKILL.md +1 -1
  28. package/skills/basic/godmode-shipping/SKILL.md +3 -0
  29. package/skills/basic/master-session/SKILL.md +89 -0
  30. package/skills/basic/startcycle/SKILL.md +1 -1
  31. package/skills/basic/startcycle-graph/SKILL.md +2 -2
  32. package/skills/basic/startcycle-graph-user/SKILL.md +1 -1
  33. package/skills/basic/teamwork-preview/SKILL.md +1 -1
  34. package/skills/bdbrainstorm/SKILL.md +7 -1
  35. package/skills/global_config/agentic-harness-patterns/SKILL.md +257 -0
  36. package/skills/global_config/agentic-harness-patterns/metadata.json +10 -0
  37. package/skills/global_config/agentic-harness-patterns/references/agent-orchestration-pattern.md +97 -0
  38. package/skills/global_config/agentic-harness-patterns/references/bootstrap-sequence-pattern.md +106 -0
  39. package/skills/global_config/agentic-harness-patterns/references/context-engineering/compress-pattern.md +78 -0
  40. package/skills/global_config/agentic-harness-patterns/references/context-engineering/isolate-pattern.md +82 -0
  41. package/skills/global_config/agentic-harness-patterns/references/context-engineering/select-pattern.md +86 -0
  42. package/skills/global_config/agentic-harness-patterns/references/context-engineering-pattern.md +29 -0
  43. package/skills/global_config/agentic-harness-patterns/references/hook-lifecycle-pattern.md +111 -0
  44. package/skills/global_config/agentic-harness-patterns/references/memory-persistence-pattern.md +109 -0
  45. package/skills/global_config/agentic-harness-patterns/references/permission-gate-pattern.md +111 -0
  46. package/skills/global_config/agentic-harness-patterns/references/skill-runtime-pattern.md +104 -0
  47. package/skills/global_config/agentic-harness-patterns/references/task-decomposition-pattern.md +92 -0
  48. package/skills/global_config/agentic-harness-patterns/references/tool-registry-pattern.md +101 -0
  49. package/skills/global_config/agenttrail/SKILL.md +8 -0
  50. package/skills/global_config/agenttrail/bin/agenttrail.mjs +14 -0
  51. package/skills/global_config/agenttrail/bin/ensure.mjs +118 -0
  52. package/skills/global_config/aos-setup/scripts/aos-doctor.mjs +1 -1
  53. package/skills/global_config/bdb-visual-edit/SKILL.md +51 -0
  54. package/skills/global_config/bdb-visual-edit/references/vite-react-source-attr.md +59 -0
  55. package/skills/global_config/bdb-visual-edit/scripts/pick-snippet.js +27 -0
  56. package/skills/global_config/bdb-visual-edit/scripts/sanitize-element.mjs +123 -0
  57. package/skills/global_config/factory-collect/SKILL.md +74 -0
  58. package/skills/global_config/factory-human-digest/SKILL.md +92 -0
  59. package/skills/global_config/factory-lookback/SKILL.md +95 -0
  60. package/skills/global_config/factory-review-prs/SKILL.md +63 -0
  61. package/skills/global_config/git-pr-review/SKILL.md +3 -0
  62. package/skills/global_config/grilling/SKILL.md +2 -0
  63. package/skills/global_config/mcsc/SKILL.md +1 -1
  64. package/skills/global_config/plan-arbiter/SKILL.md +125 -0
  65. package/skills/global_config/plan-canvas/SKILL.md +62 -5
  66. package/skills/global_config/plan-canvas/scripts/lib/loopback-guard.js +19 -3
  67. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/README.md +285 -0
  68. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/agent-trail.js +129 -0
  69. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/board-client.js +124 -0
  70. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/demo-plan/canvas.mdx +19 -0
  71. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/demo-plan/plan.mdx +18 -0
  72. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/recap-demo/plan.mdx +72 -0
  73. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/README.md +29 -0
  74. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/architecture.json +30 -0
  75. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/builder/00_architecture.html +14950 -0
  76. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/builder/canvas.mdx +511 -0
  77. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/builder/plan.mdx +208 -0
  78. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/recap/plan.mdx +102 -0
  79. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/standard/plan.md +136 -0
  80. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/signup-storyboard/canvas.mdx +124 -0
  81. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/signup-storyboard/plan.mdx +37 -0
  82. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/index.js +188 -0
  83. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/kit.js +123 -0
  84. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/mdx.js +411 -0
  85. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/render.js +1291 -0
  86. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/architecture/meta.json +1 -0
  87. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/architecture/plan.mdx +195 -0
  88. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/architecture/standard.md +95 -0
  89. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/bugfix/meta.json +1 -0
  90. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/bugfix/plan.mdx +105 -0
  91. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/bugfix/standard.md +76 -0
  92. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/feature/canvas.mdx +81 -0
  93. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/feature/meta.json +1 -0
  94. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/feature/plan.mdx +145 -0
  95. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/feature/standard.md +76 -0
  96. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/migration/meta.json +1 -0
  97. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/migration/plan.mdx +172 -0
  98. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/migration/standard.md +100 -0
  99. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap/meta.json +1 -0
  100. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap/plan.mdx +67 -0
  101. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap/standard.md +49 -0
  102. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-board/canvas.mdx +63 -0
  103. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-board/meta.json +1 -0
  104. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-board/plan.mdx +49 -0
  105. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-board/standard.md +39 -0
  106. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-review/meta.json +1 -0
  107. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-review/plan.mdx +118 -0
  108. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-review/standard.md +57 -0
  109. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/release/meta.json +1 -0
  110. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/release/plan.mdx +173 -0
  111. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/release/standard.md +96 -0
  112. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/research/meta.json +1 -0
  113. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/research/plan.mdx +91 -0
  114. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/research/standard.md +54 -0
  115. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/show-control/canvas.mdx +53 -0
  116. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/show-control/meta.json +1 -0
  117. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/show-control/plan.mdx +225 -0
  118. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/show-control/standard.md +111 -0
  119. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/theme.css +472 -0
  120. package/skills/global_config/plan-canvas/scripts/lib/plan-builder/trail.js +216 -0
  121. package/skills/global_config/plan-canvas/scripts/lib/plan-canvas/markdown.js +1 -1
  122. package/skills/global_config/plan-canvas/scripts/lib/plan-canvas/server.js +37 -4
  123. package/skills/global_config/plan-canvas/scripts/lib/plan-canvas/ui.js +125 -29
  124. package/skills/global_config/plan-canvas/scripts/plan-canvas.js +196 -8
  125. package/skills/global_config/pr-recap/SKILL.md +47 -0
  126. package/skills/global_config/pr-recap/scripts/pr-recap.mjs +200 -0
  127. package/skills/global_config/quick-recap/SKILL.md +55 -0
  128. package/skills/global_config/stay-within-limits/SKILL.md +85 -0
  129. package/skills/global_config/triage/SKILL.md +3 -0
  130. package/skills/global_config/visual-edit/README.md +96 -0
  131. package/skills/global_config/visual-edit/SKILL.md +615 -0
  132. package/skills/global_config/visual-plan/README.md +93 -0
  133. package/skills/global_config/visual-plan/SKILL.md +544 -0
  134. package/skills/global_config/visual-plan/references/canvas.md +139 -0
  135. package/skills/global_config/visual-plan/references/connection.md +51 -0
  136. package/skills/global_config/visual-plan/references/document-quality.md +186 -0
  137. package/skills/global_config/visual-plan/references/exemplar.md +62 -0
  138. package/skills/global_config/visual-plan/references/local-files.md +99 -0
  139. package/skills/global_config/visual-plan/references/wireframe.md +319 -0
  140. package/skills/global_config/visual-recap/README.md +103 -0
  141. package/skills/global_config/visual-recap/SKILL.md +560 -0
  142. package/skills/global_config/visual-recap/references/connection.md +51 -0
  143. package/skills/global_config/visual-recap/references/local-files.md +99 -0
  144. package/skills/global_config/visual-recap/references/wireframe.md +319 -0
  145. package/skills/playbooks/pb-ci-fix/SKILL.md +49 -0
  146. package/skills/playbooks/pb-event-tracker/SKILL.md +45 -0
  147. package/skills/playbooks/pb-meeting-actions/SKILL.md +42 -0
  148. package/skills/playbooks/pb-project-new/SKILL.md +48 -0
  149. package/skills/playbooks/pb-week-plan/SKILL.md +45 -0
  150. package/assets/header-v4.jpg +0 -0
@@ -0,0 +1,106 @@
1
+ # Bootstrap Sequence Pattern
2
+
3
+ ## Problem
4
+
5
+ Without a disciplined bootstrap sequence, security-critical initialization steps execute out of order: TLS certificates load after the first network connection has already completed its handshake, proxy agents are not configured before the first outbound TCP connection, and trust-gated subsystems — telemetry, full environment variable sets — activate before the user has granted permission. Concurrent callers that trigger initialization simultaneously cause expensive subsystems to double-initialize, wasting hundreds of kilobytes of module loading and creating subtle state corruption. Meanwhile, trivial commands that need no subsystems at all still pay the cost of loading the entire module graph.
6
+
7
+ These problems emerge in any agent runtime with ordered dependencies between configuration, networking, trust, and observability — they are not specific to a single implementation.
8
+
9
+ ## Golden Rules
10
+
11
+ ### Initialize by dependency order, with trust as the critical inflection point
12
+
13
+ Initialization must follow the dependency graph: no subsystem should activate before the subsystems it depends on. The two universal ordering constraints are: (1) configuration parsing must complete before any subsystem reads config values, and (2) trust establishment — the point where the user grants or withholds consent — must complete before activating subsystems that depend on consent (telemetry, secret environment variables, analytics). Between config and trust, the ordering depends on your runtime's specific dependency graph. Transport setup (TLS, proxy, certificates) typically comes early because most subsystems depend on correct network state. Observability typically comes last because it depends on config, transport, and trust. The exact layer count and sequence will vary by runtime — the principle is that you must identify and respect the dependency edges, not that every runtime has exactly the same layers.
14
+
15
+ ### Memoize the top-level init to collapse concurrent callers
16
+
17
+ Wrap the entire initialization function in a memoization boundary so that all concurrent callers share a single in-flight promise. Without this, two callers invoking init simultaneously will each execute the full sequence, potentially double-configuring global HTTP agents, double-loading telemetry modules, or racing on state that the first caller has not finished writing. The memoization boundary is at the outermost async function, not at individual sub-steps.
18
+
19
+ ### Split environment variables across the trust boundary
20
+
21
+ Environment variables fall into two categories: those with no security implications (editor preferences, locale, display settings) and those that carry secrets or alter trust posture (API keys, proxy credentials, auth tokens). Apply the safe set before the trust dialog so that early subsystems can read harmless configuration. Apply the full set only after the user has granted trust. Reversing this order means env-injected credentials are active before consent is obtained.
22
+
23
+ ### Dispatch trivial commands before loading anything
24
+
25
+ Commands that require zero subsystem state — version queries, schema dumps, diagnostic flags — should be dispatched from the entry point before any dynamic imports or initialization runs. This keeps their latency at the floor of process startup and prevents the module graph from loading hundreds of kilobytes of code that will never execute.
26
+
27
+ ### Register cleanup during init, not at usage sites
28
+
29
+ Every cleanup handler — graceful shutdown, transport teardown, child-process reaping, session cleanup — must be registered during initialization, not scattered across usage sites. Registering at init guarantees that cleanup runs unconditionally on exit, regardless of which code path was taken. Cleanup handlers registered at usage sites are silently skipped when an unexpected exit path bypasses that site.
30
+
31
+ ## When To Use
32
+
33
+ - Your agent runtime has security-critical ordering between TLS, proxy, and network subsystems.
34
+ - You support multiple entry modes (CLI, server, SDK, headless) that must share the same initialization path.
35
+ - You need a trust boundary that separates pre-consent and post-consent configuration.
36
+ - Expensive subsystems (telemetry, tracing, observability) should load only when permitted and needed.
37
+ - Trivial commands should respond instantly without paying the cost of full initialization.
38
+ - Multiple callers or code paths may trigger initialization concurrently.
39
+
40
+ ## Tradeoffs
41
+
42
+ | Decision | Benefit | Cost |
43
+ |---|---|---|
44
+ | Dependency-ordered sequential init | Security-critical steps cannot execute out of order | Slower cold start than parallel init; each layer blocks on its dependencies |
45
+ | Memoized top-level init | Concurrent callers share one promise; no double-init | Memoize caches rejections in some libraries, blocking retry on transient errors |
46
+ | Trust-split environment variables | Secrets are never active before user consent | Two separate application passes to maintain; easy to put a variable in the wrong set |
47
+ | Fast-path dispatch for trivial commands | Sub-millisecond response for version/help/diagnostics | Each fast path is a special case that must be maintained outside the normal init flow |
48
+ | Lazy dynamic imports for heavy modules | Hundreds of kilobytes deferred until actually needed | Import latency shifts to first use; harder to reason about load order |
49
+ | Parallel fire-and-forget for non-critical work | Background caches populate without blocking the critical path | Failures are silent; callers must tolerate cache misses gracefully |
50
+ | Build-time feature-flag gating | Dead code for disabled modes is eliminated from the shipped artifact | Feature checks must appear inline at call sites; wrapping in helpers defeats the bundler |
51
+ | Transport as a narrow interface | Wire protocol versions are swappable without touching callers | Behavioral differences between protocol versions can hide behind the same interface |
52
+
53
+ ## Implementation Patterns
54
+
55
+ - Parse and validate all configuration files as the very first step. Surface parse errors as user-visible diagnostics before any network contact occurs.
56
+ - Apply TLS certificate configuration immediately after config validation and before instantiating any HTTP client or agent. Many runtimes cache the TLS certificate store at process boot; certificates applied after the first handshake may have no effect on connections that reuse the cached pool.
57
+ - Configure mutual TLS before proxy agents, and proxy agents before any outbound connection including preconnect or warm-up calls. The dependency chain is: certificates, then mTLS, then agents, then connections.
58
+ - Apply the safe subset of environment variables (no secrets, no auth tokens) before the trust dialog. Defer the full set until after trust is established.
59
+ - Wrap the entire top-level async initialization function in a memoization boundary. Concurrent callers must receive the same promise, not trigger parallel execution of the same sequence.
60
+ - Fire non-critical background work (repository detection, IDE detection, cache population, analytics preamble) after the transport layer is configured but before blocking on trust. These tasks should not be awaited in the critical path.
61
+ - Gate telemetry initialization behind trust establishment. Use a separate boolean guard distinct from the init memoization so that telemetry can be retried independently if the first attempt fails.
62
+ - When the runtime supports managed remote settings, wait for those settings to load before finalizing telemetry configuration, so that organization-configured endpoints are honored.
63
+ - Use dynamic imports for heavy dependency trees (telemetry, tracing, protocol buffers, gRPC) so that their module weight is deferred until the subsystem is actually initialized.
64
+ - At the entry point, check for trivial commands (version, help, schema dump) and dispatch them immediately with zero dynamic imports — do not load the full module graph for commands that will never use it.
65
+ - Support multiple entry modes (interactive CLI, server/daemon, SDK embedding, headless) through the same initialization path, branching only at the mode-dispatch layer above init. For headless modes, skip interactive dialogs and write errors to stderr with a non-zero exit code.
66
+ - Use build-time feature flags for mode-specific branches. The flag check must appear inline at the call site so the bundler can perform dead-code elimination; wrapping it in a helper function defeats this optimization.
67
+ - Define transport as a narrow interface — write, batch-write, connect, close, state-report — so that different wire protocol versions are interchangeable behind a single surface.
68
+ - Register all cleanup handlers (graceful shutdown, transport teardown, child-process cleanup, session finalization) during initialization, not at individual usage sites.
69
+
70
+ ## Gotchas
71
+
72
+ **Memoization hides retry failures.** Some memoize implementations cache the first returned promise, including rejected ones. If a transient config error causes init to reject, every subsequent caller receives the cached rejection forever. Either use a memoize variant that clears on rejection, or implement a manual once-guard that resets its flag on error.
73
+
74
+ **Trust boundary leakage through environment variables.** Applying the full environment variable set before the trust dialog means env-injected API keys or proxy credentials are active before the user consents. This is a security violation, not just an ordering preference. Always apply the safe pass first and the full pass only after explicit trust grant.
75
+
76
+ **TLS certificate store caching defeats late application.** On many runtimes, the TLS certificate store is read once at process boot. Applying custom CA certificates after even one TLS handshake has completed may have no effect on connections that reuse the cached certificate pool. CA certificate configuration must be the very first network-adjacent step, before any connection — including preconnect warm-ups.
77
+
78
+ **Telemetry double-init across async boundaries.** When multiple code paths can trigger telemetry initialization (eager synchronous init for non-interactive sessions, deferred async init after remote settings load), the guard flag must be set before the async initializer resolves, not after. Setting it after the resolution leaves a window where a second caller can slip through and double-initialize.
79
+
80
+ **Feature flags must be inline for dead-code elimination.** Build-time feature flags only enable dead-code elimination when the flag check appears directly at the call site. Extracting the check into a helper function, even a trivially inlined one, defeats the bundler's static analysis and ships the gated code in all builds.
81
+
82
+ **Transport protocol version differences hide behind the interface.** When a transport interface abstracts multiple wire protocol versions, behavioral differences (such as one version supporting dropped-batch counts and another always returning zero) are invisible to callers. Callers that depend on version-specific behavior must not rely on the abstracted interface for those checks.
83
+
84
+ **Global state modules must be leaves in the import graph.** Any module that holds global singleton state must import only pure leaf modules with no transitive dependencies on the rest of the system. Adding a non-leaf import to a global state module creates circular-dependency risks that cascade across the entire bootstrap graph.
85
+
86
+ **Headless modes must not block on interactive dialogs.** When the runtime detects a non-interactive session (no TTY, piped input, CI environment), initialization must skip any interactive trust dialog or prompt. Blocking on user input in a headless context hangs the process indefinitely.
87
+
88
+ ## Claude Code Evidence
89
+
90
+ Claude Code's bootstrap sequence is a production implementation of these principles, supporting seven distinct entry modes through a single ordered initialization path.
91
+
92
+ **Four-layer ordering with explicit dependency comments.** The initialization function enforces strict ordering: config parsing first, then safe environment variable application, then TLS CA certificate configuration, then graceful-shutdown registration, then mTLS configuration, then global HTTP agent setup, then API preconnection. Each step includes a comment explaining why it must precede the next. The TLS step carries a specific note about the runtime caching its certificate store at boot, making late application ineffective.
93
+
94
+ **Memoized init collapsing concurrent callers.** The top-level initialization function is wrapped in a single memoization boundary, ensuring that regardless of how many code paths trigger init simultaneously, the sequence executes exactly once and all callers share the same promise.
95
+
96
+ **Trust-split environment variables.** Environment variable application is split into two distinct passes. The "safe" pass runs before any trust check, applying only variables with no security implications. The "full" pass runs only after the trust dialog completes, applying variables that may carry API keys, proxy credentials, or other sensitive configuration.
97
+
98
+ **Fast-path dispatch for trivial commands.** The CLI entry point checks for version flags, schema dumps, and diagnostic commands before loading any modules. These paths respond with zero dynamic imports and exit immediately, keeping their latency at the floor of process startup.
99
+
100
+ **Lazy loading of heavy subsystems.** Telemetry initialization uses dynamic imports to defer approximately 400KB of observability and protocol-buffer modules until telemetry is actually needed. A further layer of lazy loading defers approximately 700KB of gRPC transport modules until tracing instrumentation is activated. This two-tier lazy loading keeps the critical path fast for sessions that never activate tracing.
101
+
102
+ **Parallel fire-and-forget for background caches.** After the transport layer is configured but before blocking on trust, the bootstrap fires off non-critical background tasks — IDE detection, repository detection, OAuth cache population, analytics preamble — as unawaited promises. These populate shared caches that later synchronous callers read, without blocking the critical initialization path.
103
+
104
+ **Multi-mode entry through a single init path.** Seven entry modes — interactive REPL, server/daemon, SDK embedding, bridge/remote-control, background sessions, template runner, and self-hosted runner — all flow through the same memoized initialization function. Mode-specific behavior branches above the init call, not within it. Feature-gated modes use build-time flags with inline checks so that unused modes are eliminated from the shipped artifact.
105
+
106
+ **Transport abstraction for swappable wire protocols.** The bridge transport is defined as a narrow interface with write, batch-write, connect, close, and state-reporting methods. Two protocol versions — one using WebSocket plus POST, the other using server-sent events for reads and a separate write channel — are interchangeable behind this surface. The interface documents where behavioral differences exist between versions.
@@ -0,0 +1,78 @@
1
+ # Context Compression and Snapshot Management
2
+
3
+ ## Problem
4
+
5
+ Long-running agent sessions inevitably exhaust the context window. Each turn adds tool calls, results, and reasoning, and none of it shrinks on its own. Without compression, the agent either hits the hard token limit and fails, or the model's attention degrades as useful information is buried under pages of stale output. Variable-length context blocks — directory listings, git status, search results — make the problem worse because their size is unpredictable and can spike by orders of magnitude on a single turn.
6
+
7
+ A second, subtler problem: context that was accurate when captured becomes misleading as the session progresses. A git status from turn 3 is treated as current truth at turn 30 unless something marks it as a point-in-time snapshot.
8
+
9
+ These problems affect any agent runtime that persists across turns. They are not specific to any single implementation.
10
+
11
+ ## Golden Rules
12
+
13
+ ### Truncate with recovery, never truncate blind
14
+
15
+ When a context block exceeds a character or token cap, truncate it — but always append a recovery pointer: a specific instruction telling the model which tool to call, with which arguments, to retrieve the full output. Truncation without a recovery pointer is a dead end. The model knows the data exists, knows it was cut off, and has no path back. This produces hallucinated completions or silent failures.
16
+
17
+ ### Compact reactively, not on a schedule
18
+
19
+ Fixed-window compaction (e.g., "summarize every 20 turns") either fires too early and loses useful context or too late and hits the token ceiling. Reactive compaction monitors the actual fill ratio and fires when the window is genuinely full. The compaction pass summarizes older turns while preserving recent ones and active tool results, recovering budget mid-session without losing the thread of work.
20
+
21
+ ### Label every snapshot
22
+
23
+ Any context block that captures point-in-time state — git status, directory listings, process output, file contents — must carry an explicit label: the time of capture and a note that the data will not auto-update. Without this label, the model treats stale data as current and produces reasoning that contradicts the actual state of the world.
24
+
25
+ ### Cap every variable-length block
26
+
27
+ Every context block whose length depends on external state must have a hard character or token ceiling. Directory listings, search results, git diffs, log output — all of these can spike unpredictably. Without a cap, a single large result can consume the majority of the context budget in one turn, crowding out everything else.
28
+
29
+ ## When To Use
30
+
31
+ - Agent performance degrades noticeably as sessions get longer.
32
+ - The agent repeatedly acts on stale data (old git status, outdated file contents, stale search results).
33
+ - Variable-length context blocks occasionally spike and crowd out other information.
34
+ - Sessions are long enough that the context window approaches its token limit.
35
+ - The agent hallucinates completions of truncated output instead of re-fetching.
36
+
37
+ ## Tradeoffs
38
+
39
+ | Decision | Benefit | Cost |
40
+ |---|---|---|
41
+ | Truncation with recovery pointers | Model can always retrieve full data | Extra round-trip when full data is needed |
42
+ | Reactive compaction | Extends effective session length dynamically | Older context is lossy-compressed, not losslessly preserved |
43
+ | Snapshot labeling | Prevents stale-data reasoning | Every context injection site must add the label |
44
+ | Hard character caps on variable blocks | No single block can starve the budget | Useful data may be cut; recovery pointer quality matters |
45
+ | Preserving recent turns during compaction | Model retains working memory of current task | Compaction logic must distinguish "recent" from "old" |
46
+
47
+ ## Implementation Patterns
48
+
49
+ - Enforce a character or token cap on every variable-length context block. The cap should be set conservatively — it is easier to raise a cap that is too tight than to debug a session where one block consumed the entire budget.
50
+ - Append a truncation notice that names the specific tool and invocation needed to fetch the full data. The notice should be concrete: "Run [tool] with [these arguments] to see the complete output," not a vague "output was truncated."
51
+ - Tag every injected snapshot with a capture timestamp and a note that the data is static. Place the label at the top of the block, not the bottom — the model processes tokens sequentially and needs the staleness warning before it starts reasoning about the content.
52
+ - Implement a fill-ratio monitor that triggers compaction when the context window reaches a threshold (e.g., 80% full). The compaction pass should summarize older turns into a condensed narrative while preserving the most recent turns and any active tool results verbatim.
53
+ - During compaction, preserve any recovery instructions that were part of truncated blocks. If a truncation notice is compacted away, the model loses its path back to the full data.
54
+ - Include explicit instructions in the compaction output telling the model what was summarized and how to re-fetch details if needed. The model should know that compaction happened and what was lost.
55
+
56
+ ## Gotchas
57
+
58
+ **Truncation without a recovery pointer is a dead end.** If you cap a context block and do not tell the model which tool to call for the full output, the model has no path back to the data. It will either hallucinate the missing content or silently drop the task that depended on it.
59
+
60
+ **Snapshot labels prevent stale reasoning.** A git status injected without a "this is a snapshot taken at time T" label causes the model to treat the data as live. It will produce branch, diff, and merge reasoning that contradicts the actual current state.
61
+
62
+ **Compaction must preserve recovery pointers.** If a truncation notice from an earlier turn is itself summarized away during compaction, the model loses the ability to re-fetch the full data. Recovery pointers should survive compaction passes.
63
+
64
+ **Variable-length blocks can spike by orders of magnitude.** A directory listing in a small project is 20 lines; in a monorepo it can be 20,000. A git diff is usually small; after a large refactor it can be enormous. Caps must be set for the worst case, not the common case.
65
+
66
+ **Reactive compaction must not fire during active tool use.** If compaction triggers while a multi-step tool sequence is in progress, it may summarize away intermediate results that the model needs for the next step. Gate compaction on a quiescent state — between turns, not mid-turn.
67
+
68
+ ## Claude Code Evidence
69
+
70
+ Claude Code's compression strategy demonstrates these principles in a long-running interactive session environment:
71
+
72
+ **Git status as a labeled, capped snapshot.** The git status string injected into context carries an explicit note: "this status is a snapshot in time, and will not update during the conversation." When the output exceeds a 2000-character threshold, it is truncated with a message directing the model to run a specific tool to see the full output. This pairing of label plus recovery pointer means the model always knows the data is stale and always has a path to fresh data.
73
+
74
+ **Reactive compaction for extended sessions.** When the context window fills during a long session, the runtime triggers a compaction pass that summarizes older conversation turns while preserving the most recent exchanges and any active tool results. The model is given explicit context about what was summarized. This approach extends effective session length without a hard turn limit — sessions can run for hours as long as compaction can recover enough budget on each pass.
75
+
76
+ **Character caps on all variable-length injections.** Every context block whose size depends on external state — git status, directory listings, file contents — has a hard character ceiling. The design philosophy is that no single context source should be able to starve the budget for all others. When a cap is hit, the truncation notice always includes the specific tool invocation needed to retrieve the complete output, ensuring the model is never left at a dead end.
77
+
78
+ **Recovery instructions as first-class content.** Truncation notices in Claude Code are not afterthoughts or generic warnings. Each one names the exact tool and arguments the model should use. This specificity matters: a vague "output was truncated" message forces the model to guess how to recover, while a concrete instruction lets it act immediately.
@@ -0,0 +1,82 @@
1
+ # Context Isolation for Delegated Work
2
+
3
+ ## Problem
4
+
5
+ When an agent delegates work to sub-agents, every piece of shared state becomes a potential collision point. A sub-agent that inherits the parent's full context may act on stale assumptions; a sub-agent that shares the parent's working directory may overwrite files the parent is reading; a sub-agent that can recursively spawn its own children creates exponential context multiplication. Without explicit isolation boundaries, the blast radius of a single sub-agent mistake extends to the entire session.
6
+
7
+ These problems emerge in any agent runtime that supports delegation — coordinator patterns, fork patterns, or any form of parallel sub-agent work. They are inherent to multi-agent architectures, not specific to any single implementation.
8
+
9
+ ## Golden Rules
10
+
11
+ ### Zero-inheritance is the safest default
12
+
13
+ When delegating work to a sub-agent in a coordinator pattern, pass only the explicit prompt. The sub-agent receives no conversation history, no accumulated context, no session state. This is the narrowest possible boundary: the sub-agent can only act on what it was told, and nothing from the parent's session leaks accidentally. The cost is that the prompt must be self-sufficient — everything the sub-agent needs to do its job must be spelled out.
14
+
15
+ ### Full-inheritance forks must be single-level
16
+
17
+ When a sub-agent needs the parent's full context (the fork pattern), copy the context but enforce a single-level boundary: the forked agent cannot itself fork. Without this constraint, recursive forking multiplies context cost exponentially — each level copies everything from the level above, and the total cost is O(2^n) in the depth of the fork tree. A single-level boundary keeps the cost linear.
18
+
19
+ ### Filesystem isolation prevents path collisions
20
+
21
+ When a sub-agent modifies files, it needs its own working copy of the repository. Shared working directories create race conditions: the parent reads a file, the sub-agent modifies it, and the parent's next read returns unexpected content. Filesystem isolation (via worktrees, temp directories, or copy-on-write clones) gives each agent its own view. Path translations must be injected so the agent operates on its isolated copy, not the shared original.
22
+
23
+ ### Choose the narrowest boundary that works
24
+
25
+ The isolation boundary determines the blast radius of a sub-agent's mistakes. A zero-inheritance worker can corrupt only its own output. A full-inheritance fork can produce results inconsistent with the parent's state. A shared-filesystem agent can corrupt the parent's working directory. Always start with the narrowest boundary and widen only when the sub-agent genuinely needs more context.
26
+
27
+ ## When To Use
28
+
29
+ - Your agent delegates work to sub-agents or spawns concurrent workers.
30
+ - Sub-agents modify files in the same repository the parent is working in.
31
+ - Delegated sub-agents produce results that contradict the parent's state or each other.
32
+ - You need to prevent a sub-agent failure from corrupting the parent session.
33
+ - Parallel agents need to work on the same codebase without stepping on each other's changes.
34
+
35
+ ## Tradeoffs
36
+
37
+ | Decision | Benefit | Cost |
38
+ |---|---|---|
39
+ | Zero-inheritance delegation | Clean context boundaries, no accidental leakage | Sub-agent must be self-sufficient from its prompt alone |
40
+ | Full-inheritance fork (single-level) | Sub-agent has full context for complex tasks | Doubles context cost; must enforce no-recursive-fork rule |
41
+ | Filesystem isolation via worktrees | No file-level race conditions between agents | Worktree creation and teardown overhead; disk cost |
42
+ | Path translation injection | Agent operates on its own copy transparently | Translation logic must cover all file-operation tools |
43
+ | Immutable shared state | Safe concurrent reads across agents | State updates require new snapshots, not in-place mutation |
44
+ | Single-level fork boundary | Linear cost, not exponential | Forked agents that need further delegation must use coordinator pattern instead |
45
+
46
+ ## Implementation Patterns
47
+
48
+ - For coordinator-pattern delegation, construct the sub-agent prompt as a self-contained document. Include all necessary context, constraints, and output format requirements in the prompt itself — do not rely on inherited session state.
49
+ - For fork-pattern delegation, copy the parent's context at fork time and enforce a single-level boundary. The forked agent should be unable to invoke the fork mechanism itself. If a forked agent needs further delegation, it must use the coordinator pattern (zero-inheritance) for its own sub-agents.
50
+ - When sub-agents modify files, create an isolated filesystem context (worktree, temp directory, or copy-on-write clone) before the agent starts. Inject path translations so every file-operation tool targets the isolated copy.
51
+ - On sub-agent completion, merge results back to the parent's working directory through a controlled integration point — not by letting the sub-agent write directly to the shared filesystem.
52
+ - Mark all state shared between parent and sub-agents as immutable. If a sub-agent needs to communicate state changes, it should return them as a result, not mutate shared objects in place.
53
+ - Register cleanup handlers that tear down isolated filesystems when the sub-agent completes or fails. Leaked worktrees or temp directories accumulate and waste disk.
54
+ - When multiple sub-agents work in parallel on the same codebase, assign non-overlapping file sets where possible. Filesystem isolation handles the general case, but non-overlapping assignments prevent merge conflicts at the integration point.
55
+
56
+ ## Gotchas
57
+
58
+ **Recursive forks create exponential cost.** If a forked agent can itself fork, the context cost doubles at each level. Three levels of recursive forking mean eight copies of the original context. Enforce the single-level boundary at the delegation mechanism, not as a convention.
59
+
60
+ **Path translations must cover all file-operation tools.** If the isolation mechanism injects a worktree path but one file-operation tool bypasses the translation, that tool writes to the shared original directory. Every tool that touches the filesystem must go through the translation layer.
61
+
62
+ **Zero-inheritance prompts must be truly self-contained.** A common failure mode is constructing a sub-agent prompt that implicitly relies on context the parent has but the sub-agent does not. Review zero-inheritance prompts for hidden dependencies: file paths that assume a specific working directory, references to earlier conversation turns, or assumptions about available tools.
63
+
64
+ **Merging results from isolated agents can conflict.** When two parallel agents modify the same file in their respective worktrees, merging both back produces a conflict. Design task assignments to minimize overlap, and have a conflict resolution strategy for when overlap is unavoidable.
65
+
66
+ **Worktree cleanup must happen on all exit paths.** If cleanup only runs on successful completion, a failed or killed sub-agent leaks its worktree. Register cleanup on the process-level shutdown handler, not just in the success path.
67
+
68
+ **Immutable shared state does not mean no communication.** Sub-agents can still return results to the parent — the constraint is that they return values rather than mutating shared objects. The parent integrates returned results into its own state through a controlled update path.
69
+
70
+ ## Claude Code Evidence
71
+
72
+ Claude Code's delegation system implements these isolation principles across multiple agent coordination patterns:
73
+
74
+ **Zero-inheritance as the default for coordinated work.** When the runtime dispatches a sub-agent in the coordinator pattern, the sub-agent receives only its explicit task prompt. No conversation history, no accumulated memory, no parent session state is inherited. This is a deliberate design choice: the sub-agent's blast radius is limited to its own output. The cost — that every sub-agent prompt must be fully self-contained — is accepted as worthwhile because it eliminates an entire class of state-leakage bugs.
75
+
76
+ **Single-level fork boundary.** The fork mechanism copies the parent's full conversation context to the forked agent but enforces that the forked agent cannot itself invoke the fork mechanism. If a forked agent needs to delegate further, it must use the coordinator pattern (zero-inheritance) for its own sub-tasks. This keeps the total context cost at 2x the parent's context, not 2^n. The design explicitly rejects recursive forking because the exponential cost makes it impractical for any real workload.
77
+
78
+ **Worktree-based filesystem isolation.** When a sub-agent needs to modify files, the runtime creates a git worktree that gives the agent its own working copy of the repository. Path translations are injected so that every file-operation tool targets the worktree, not the parent's working directory. This eliminates file-level race conditions between concurrent agents. Worktrees are cleaned up on all exit paths, including failure and kill, through process-level cleanup handlers.
79
+
80
+ **Immutable shared state across concurrent agents.** The application state shared between the parent and its sub-agents is wrapped in a deep-immutability type. Sub-agents cannot mutate this state in place. When a sub-agent produces results, they are returned as values and integrated by the parent through a controlled update path. This design prevents the most insidious class of concurrency bugs: silent in-place mutation of shared objects that causes different agents to see inconsistent state.
81
+
82
+ **Blast radius as the primary design criterion.** The isolation level for each delegation type was chosen by asking "what is the worst thing this sub-agent can break?" Zero-inheritance workers can only produce bad output. Full-inheritance forks can produce output inconsistent with the parent's context. Shared-filesystem agents can corrupt the parent's working directory. The runtime defaults to the narrowest boundary and only widens when the use case demands it.
@@ -0,0 +1,86 @@
1
+ # Context Selection and Progressive Disclosure
2
+
3
+ ## Problem
4
+
5
+ Agent runtimes that eagerly load all available context on every turn pay a compounding tax: token costs grow with catalog size, startup latency scales with the number of information sources, and concurrent agents duplicate the same expensive I/O. The model receives thousands of tokens it never reads while the data it actually needs arrives late or not at all. Without deliberate selection, the context window becomes a landfill — technically full, practically empty.
6
+
7
+ These problems appear in any agent that maintains a growing skill catalog, reads from multiple persistent stores, or serves concurrent requests. They are not specific to any single runtime.
8
+
9
+ ## Golden Rules
10
+
11
+ ### Lazy beats eager
12
+
13
+ Load context at the moment it becomes relevant, not at the moment the session starts. Eagerly injecting everything "just in case" means every API call pays for tokens the model may never use. Deferring the load means you pay only when the model actually activates a capability. The one-time latency of a just-in-time fetch is almost always cheaper than the per-turn cost of a bloated prompt.
14
+
15
+ ### Three-tier progressive disclosure
16
+
17
+ Structure every context source into three tiers of increasing cost:
18
+
19
+ 1. **Metadata** (~100 tokens per item): names, descriptions, and trigger hints. Always present. This is the "is this relevant?" signal.
20
+ 2. **Instructions** (bounded, e.g. < 5000 tokens): the full body of the activated capability. Loaded only on activation.
21
+ 3. **Resources** (unbounded): reference files, scripts, data assets. Loaded only when the agent explicitly requests them.
22
+
23
+ The discovery cost (tier 1) scales with catalog size, but the execution cost (tiers 2-3) is constant per activated item. Keeping discovery cheap lets you maintain a large catalog without paying for it on every turn.
24
+
25
+ ### Memoize the promise, not the result
26
+
27
+ When multiple callers request the same expensive context simultaneously, memoize the in-flight promise so all callers share a single I/O operation. Memoizing only the resolved value creates a window where concurrent callers each see the cache as empty and launch duplicate work. The promise itself is the deduplication key.
28
+
29
+ ### Invalidate at the mutation site, not on a timer
30
+
31
+ Cache expiration based on time or reactive subscriptions either serves stale data or rebuilds too often. Instead, place explicit invalidation calls at each known mutation point. This is more work to maintain — every new write path must add its own invalidation — but it guarantees that the cache is exactly as fresh as the most recent deliberate change.
32
+
33
+ ## When To Use
34
+
35
+ - The skill or capability catalog is growing and the always-on prompt is becoming expensive.
36
+ - Startup latency is high because context is built eagerly on every turn.
37
+ - Multiple concurrent agents or invocations duplicate the same expensive I/O.
38
+ - You need to support a large catalog without linearly increasing per-turn token cost.
39
+ - Context sources span multiple directories, stores, or services that may overlap.
40
+
41
+ ## Tradeoffs
42
+
43
+ | Decision | Benefit | Cost |
44
+ |---|---|---|
45
+ | Lazy loading | Low idle token cost | One round-trip latency on first activation |
46
+ | Progressive disclosure (3 tiers) | Large catalogs stay cheap | Model cannot reason about unactivated capabilities |
47
+ | Promise memoization | No duplicate I/O under concurrency | Cache logic is more subtle than simple value caching |
48
+ | Manual invalidation at mutation sites | Only rebuilds when necessary | Every new write path must include its own invalidation call |
49
+ | Token budget for discovery tier | Catalog cost is bounded and predictable | Each entry must fit within a tight character cap |
50
+ | Path-conditional activation | Domain-specific skills stay dormant until relevant | Activation state must survive process restarts or be re-evaluated |
51
+
52
+ ## Implementation Patterns
53
+
54
+ - Wrap every expensive context builder in a memoized async function. Store the promise itself so concurrent callers share the same in-flight request rather than spawning duplicate I/O.
55
+ - Expose a named cache-clear call for each memoized builder. Invoke it only from known mutation handlers — never from a general observer or polling loop.
56
+ - For capability catalogs, compute an estimated token count from metadata fields only (name, description, trigger hint). Store the full body in a deferred closure, fetched only on activation.
57
+ - Cap the total discovery-tier budget at a fixed fraction of the context window (e.g., ~1%), with each entry limited to a tight character ceiling.
58
+ - Use path-condition metadata to gate domain-specific capabilities: a capability stays dormant until a matching file or directory is touched in the session.
59
+ - Deduplicate multi-source loads by canonical path. Resolve symlinks and normalize before inserting into the active set — string comparison of original paths produces false negatives, and inode comparison is unreliable on virtual or network filesystems.
60
+ - Mark shared state as deeply immutable. Document any fields excluded from the immutability wrapper (e.g., function-typed fields) at the point of exclusion, not silently.
61
+
62
+ ## Gotchas
63
+
64
+ **Stale context after mutation.** Because the cache never expires automatically, any external change (new file, toggled flag, updated store) is invisible until the cache is explicitly cleared. Every mutation point must call the corresponding invalidation. Missing one means the model operates on stale data for the rest of the session.
65
+
66
+ **Concurrent initialization races without promise memoization.** If two invocations trigger the same extraction simultaneously and both see the cache as empty, both will attempt the same I/O. Only one will succeed cleanly; the other may fail or produce corrupt output. Memoizing the promise eliminates the race entirely.
67
+
68
+ **Premature disclosure inflates every turn.** Injecting full capability bodies into the system prompt means every API call pays for tokens the model may never use. With 20+ capabilities, this can consume a significant share of the context budget before the first user message arrives.
69
+
70
+ **Conditional capabilities must track activation state.** Capabilities gated by path conditions are dormant until a matching file is touched. If activation state is lost (e.g., process restart), the capability must be re-evaluated against current context — otherwise it silently disappears from the session.
71
+
72
+ **Deduplication must use canonical paths, not original paths.** The same file can appear under different paths when symlinks, mounts, or multiple search directories are involved. String comparison of the original paths will miss duplicates; resolve to canonical paths before comparing.
73
+
74
+ ## Claude Code Evidence
75
+
76
+ Claude Code's context system is a production implementation of these principles, managing a growing catalog of skills across multiple source directories:
77
+
78
+ **Memoize-plus-manual-invalidate, not reactive subscriptions.** Both the system context builder and the user context builder are memoized async functions. The cache persists for the entire process lifetime. When something must change — for instance, a system prompt injection is updated — the runtime calls cache-clear explicitly at that specific mutation point. This is a deliberate choice: context is cheap to serve from cache and expensive to rebuild. A reactive model that rebuilt on every state update would defeat the purpose of caching entirely.
79
+
80
+ **Promise memoization for concurrent initialization.** The bundled skills extraction stores the extraction promise itself rather than the resolved value. If two skill invocations race during startup, both await the same promise. Without this design, both would attempt to write the same extracted files concurrently, and the second would fail or corrupt the output.
81
+
82
+ **Token count as the gating metric for progressive disclosure.** The skill loader computes an estimated token count from only the name, description, and trigger hint fields — never from the full markdown body. The body is stored in a deferred closure and fetched only at invocation time. This cleanly separates the "is this relevant?" signal (cheap, always present) from the "what does this skill do?" content (expensive, deferred). The total listing budget is capped at roughly one percent of the context window, with each entry limited to approximately 250 characters.
83
+
84
+ **Realpath deduplication across multiple source directories.** When loading skills from managed, user, project, and additional directories, the same skill file can appear under symlinked paths. The loader resolves every file to its canonical path before inserting into a seen-set. The team chose canonical-path comparison over inode comparison because inodes are unreliable on virtual and network filesystems.
85
+
86
+ **Deeply immutable shared state.** The main application state is wrapped in a deep-immutability type to prevent accidental in-place mutation across concurrent agents. Task state is excluded from the immutability wrapper because it contains function types that the wrapper cannot handle — but the exclusion is documented at the point of exclusion, not silently dropped.
@@ -0,0 +1,29 @@
1
+ # Context Engineering Pattern
2
+
3
+ ## The Problem
4
+
5
+ Without deliberate management, context windows fill with redundant or stale data on every turn. Rebuilding expensive context on each invocation adds latency. Exposing the full capability catalog up front bloats every prompt. Delegated sub-agents pollute their parent's context with intermediate state. These are inherent to any agent that persists across turns and delegates to sub-agents.
6
+
7
+ ## The Four-Axis Framework
8
+
9
+ Context engineering decomposes into four concerns:
10
+
11
+ 1. **Select** -- decide what enters the context window, when, and at what granularity.
12
+ 2. **Write** -- the agent does not just consume context; it writes back to persistent storage, creating the learning loop.
13
+ 3. **Compress** -- long sessions exhaust the window; truncation, compaction, and snapshot labeling recover budget.
14
+ 4. **Isolate** -- delegated work must not pollute or corrupt the parent's context or filesystem.
15
+
16
+ ### A note on "Write"
17
+
18
+ The write axis is a cross-cutting principle rather than a standalone pattern. Every axis involves write-back: selections are cached and invalidated, compression produces summaries that are stored, isolation boundaries return results that the parent integrates. The write-back loop -- auto-memory, session extraction, persisted permissions, task state updates -- is what transforms a stateless tool-caller into a learning system. It is covered as a design principle in the main skill definition rather than a separate sub-file.
19
+
20
+ ## Sub-File Index
21
+
22
+ ### [Select: Context Selection and Progressive Disclosure](context-engineering/select-pattern.md)
23
+ Lazy loading, three-tier progressive disclosure, promise memoization for concurrency, manual cache invalidation, token budgeting for capability discovery, canonical-path deduplication, and path-conditional activation.
24
+
25
+ ### [Compress: Context Compression and Snapshot Management](context-engineering/compress-pattern.md)
26
+ Truncation with recovery pointers, reactive compaction triggered by fill ratio, snapshot labeling with staleness warnings, character and token caps on variable-length blocks, and recovery-pointer preservation across compaction passes.
27
+
28
+ ### [Isolate: Context Isolation for Delegated Work](context-engineering/isolate-pattern.md)
29
+ Zero-inheritance workers, full-inheritance forks with single-level boundaries, filesystem isolation via worktrees, path translation injection, blast radius as the primary design criterion, and immutable shared state across concurrent agents.
@@ -0,0 +1,111 @@
1
+ # Hook Lifecycle Pattern
2
+
3
+ ## Problem
4
+
5
+ Without a centralized hook lifecycle, extensibility degrades into chaos: hook authors attach side effects at arbitrary points in agent execution with no consistent trust enforcement, no deterministic ordering across concurrent hooks, and no structured way to propagate blocking decisions back to the caller. External hooks may execute before workspace trust has been established, enabling remote code execution from untrusted configuration files. Long-running hooks that background themselves without proper registry handoff leave orphaned processes that deliver stale or contradictory signals at unexpected times.
6
+
7
+ These problems emerge in any agent runtime that supports extensible hook points — they are not specific to a single implementation.
8
+
9
+ ## Golden Rules
10
+
11
+ ### All hooks flow through a single dispatch point
12
+
13
+ Never allow hook execution to scatter across multiple call sites. A single dispatch function is the chokepoint for trust checks, source merging, type routing, and result aggregation. If a new hook type is added, the dispatch point is the only place that needs to change. Scattered execution sites inevitably diverge on trust enforcement or result handling.
14
+
15
+ ### Trust is all-or-nothing
16
+
17
+ Before any hook fires, check workspace trust exactly once. If trust has not been established, skip every hook — not just external ones. In-process hooks included. This prevents a half-trusted state where some hooks execute and others do not, which creates unpredictable behavior that is harder to debug than a clean skip-all.
18
+
19
+ ### Multiple sources merge with explicit priority
20
+
21
+ Hooks arrive from multiple sources: persisted configuration, programmatic SDK registration, and ephemeral session-scoped registrations. Merge them into a single ordered list with a defined priority layering. When enterprise policy restricts hooks, entire source layers are excluded cleanly rather than individual hooks being filtered ad hoc.
22
+
23
+ ### Exit codes carry semantic weight
24
+
25
+ For external process hooks, reserve specific exit codes for specific meanings. One code means success. One specific code means "block this action and inject my error message into the agent's context." All other non-zero codes mean "warn the user but do not block." This three-tier convention lets external scripts control agent behavior without needing structured JSON output.
26
+
27
+ ### Deny beats ask beats allow
28
+
29
+ When multiple hooks in the same batch return conflicting permission decisions, apply a strict precedence: deny overrides ask, ask overrides allow, and passthrough never sets the aggregated decision. This deterministic resolution eliminates ambiguity when hooks disagree.
30
+
31
+ ### Session hooks are ephemeral by design
32
+
33
+ Hooks registered at runtime for a specific agent session must be scoped to that session's ID and automatically cleaned up when the session ends. They should never be written to persisted configuration. If a sub-agent registers a hook, that hook must not fire in the parent agent's session or in a sibling sub-agent's session. This isolation prevents cross-session side effects during parallel fan-out.
34
+
35
+ ## When To Use
36
+
37
+ - Your agent runtime exposes lifecycle events that external code can hook into.
38
+ - You need trust enforcement that cannot be bypassed by any hook type.
39
+ - Hooks arrive from multiple configuration sources that must be merged deterministically.
40
+ - You support both synchronous blocking hooks and long-running background hooks.
41
+ - Multiple hooks can return conflicting permission decisions for the same action.
42
+ - You need session-scoped hooks that are automatically cleaned up when a session ends.
43
+
44
+ ## Tradeoffs
45
+
46
+ | Decision | Benefit | Cost |
47
+ |---|---|---|
48
+ | Single dispatch point | One place for trust, routing, and aggregation | Every hook type must conform to the same dispatch interface |
49
+ | All-or-nothing trust gate | No half-trusted states, simple mental model | In-process hooks that pose no security risk are still blocked |
50
+ | Multi-source merge with priority | Clean policy enforcement, predictable ordering | Hook authors must understand which source layer they belong to |
51
+ | Multiple hook types | Right tool for each job (process, LLM, HTTP, in-process) | More type surface to document and maintain |
52
+ | Async generator composition | Streaming partial results, early exit on blocking errors | Callers must handle incremental accumulation |
53
+ | Backgrounding with registry handoff | Long-running hooks do not block the main loop | Orphan risk if the registry handoff is skipped |
54
+ | Session-scoped hooks | Ephemeral hooks auto-clean on session end | Hooks scoped to one session do not fire in sub-agent sessions |
55
+ | Deduplication across sources | No double-firing when the same hook appears in multiple configs | Dedup key design must account for source-specific prefixes |
56
+
57
+ ## Implementation Patterns
58
+
59
+ - Route every hook invocation through a single dispatch function. This function is the sole location for the trust check, source merge, type dispatch, and result aggregation.
60
+ - Implement the trust check as the first operation in dispatch. If trust is absent, return immediately with an empty result — do not evaluate which hooks would have matched.
61
+ - Merge hook sources in a defined priority order (e.g., persisted config, then SDK-registered, then session-scoped). When a policy flag restricts to managed hooks only, exclude entire source layers rather than filtering individual entries.
62
+ - Deduplicate across sources using a composite key that includes source-specific context. Two hooks with the same command from different plugins are distinct; the same command from user config and project config collapses to the last-wins entry.
63
+ - Support at least these hook type categories: external process (shell command), LLM sub-call, sub-agent delegation, HTTP endpoint, in-process callback with full output control, and lightweight boolean gate for simple allow/deny decisions.
64
+ - Run all matched hooks in parallel using async generator composition. Each hook yields typed result fragments as they become available. The dispatch function merges all generators and yields aggregated fragments to the caller.
65
+ - For external process hooks, map exit codes to semantic outcomes: one code for success, one reserved code for blocking errors (with the error message on stderr), and all other non-zero codes for non-blocking warnings.
66
+ - Support two backgrounding modes: static (declared in hook configuration before execution) and dynamic (the hook signals backgrounding as its first output line). Backgrounded hooks are handed off to a pending-hook registry so they are tracked across turns.
67
+ - Scope session hooks to a specific agent or session ID. Hooks registered under the main session do not fire in sub-agent sessions. Provide explicit removal functions, but also clear all session hooks automatically when the session ends.
68
+ - When aggregating permission decisions from multiple hooks, apply strict precedence: deny wins over ask, ask wins over allow, passthrough is ignored.
69
+ - Choose hook types based on the capability needed. Use external process hooks for shell-level side effects. Use LLM sub-call hooks when the hook itself needs reasoning. Use HTTP hooks for external service integration. Use in-process callbacks when the hook needs full control over structured output fields (permission decisions, injected context, input rewriting). Use lightweight boolean gates when you only need a yes/no decision against conversation history without JSON serialization overhead.
70
+ - Register a process-level cleanup handler that terminates all running hooks on shutdown. Without this, backgrounded hooks survive the parent process and become orphans.
71
+ - Guard against event-type restrictions: some events may be incompatible with certain hook types. For example, HTTP hooks during early session setup can deadlock if the event fires before the response consumer is ready. Document these restrictions in the dispatch layer rather than relying on hook authors to discover them.
72
+
73
+ ## Gotchas
74
+
75
+ **The trust gate blocks all hook types, not just external ones.** Every hook — including in-process callbacks and boolean gates — is skipped when trust is absent. Never assume a hook fired just because it runs in-process. If a hook did not fire, check trust status first.
76
+
77
+ **A missing external script can produce the same exit code as an intentional block.** If a hook script is deleted between registration and execution, the shell itself may exit with the reserved blocking code. Validate that hook scripts exist at registration time, not only at execution time.
78
+
79
+ **Policy-restricted modes silently drop session hooks.** When enterprise policy limits hooks to managed sources only, session-scoped hooks registered by skills or sub-agents receive zero matches with no error. Tested behavior may differ between managed and unmanaged deployments.
80
+
81
+ **Background hooks with "rewake" semantics bypass the pending registry.** Normal backgrounded hooks appear in the registry and are resolved before the next turn. Rewake-style hooks survive across turns by design — they inject a notification when they complete. Do not use rewake for hooks that must finish before the next tool call.
82
+
83
+ **Boolean gate hooks cannot be serialized.** Lightweight in-process gates that return only true/false are inherently ephemeral. They are excluded from configuration snapshots, analytics, and telemetry. Use the full callback type if you need persistence or observability.
84
+
85
+ **Deduplication uses last-wins, which can surprise multi-scope authors.** When the same hook command appears in both user-level and project-level config, the project-level entry wins. Source-specific prefixes in the dedup key prevent cross-source collisions, but same-source collisions are resolved silently.
86
+
87
+ **HTTP hooks can deadlock during early lifecycle events.** If an HTTP hook is registered for a session-start or setup event, the outbound request may fire before the response consumer is ready, creating a deadlock. The dispatch layer should explicitly filter incompatible hook-type/event-type combinations rather than leaving this to the hook author.
88
+
89
+ **Session hooks are invisible after session cleanup.** Once a session ends and its hooks are cleared, there is no record they ever existed. If debugging requires understanding which hooks were active during a past session, you need separate logging — the session store itself is ephemeral.
90
+
91
+ ## Claude Code Evidence
92
+
93
+ Claude Code's hook system is a production implementation of these principles, supporting six hook types across dozens of lifecycle events.
94
+
95
+ **Single dispatch with trust gate.** All hook execution flows through one async generator function. The first thing it does is check workspace trust. In interactive mode, this requires the user to have accepted a trust dialog. In SDK (non-interactive) mode, trust is implicit. If trust is absent, the generator returns immediately — no hooks of any type fire.
96
+
97
+ **Three-layer source merge with policy enforcement.** Hooks are drawn from three sources in priority order: snapshot hooks from settings files, SDK-registered callback hooks, and session-scoped hooks stored in a Map-based session state. When enterprise policy sets a managed-only flag, session and plugin hooks are excluded at the source-merge layer, not filtered individually. This design means policy enforcement is a single conditional that drops entire source lists, rather than per-hook filtering that would need to be maintained for every new hook type.
98
+
99
+ **Async generator composition for streaming results.** Each hook runs as an independent async generator yielding typed result fragments — blocking errors, permission decisions, additional context, modified inputs. A parallel merge utility combines all generators so the caller receives fragments as they arrive. This is critical for blocking errors: a slow hook in the batch does not delay the blocking signal from a fast hook.
100
+
101
+ **Map-based session state for concurrency.** Session hooks use a Map rather than a plain object for the session store. Map mutation returns the same container reference, so the application store's equality check sees no change and none of the store's listeners fire. This avoids quadratic listener-notification costs during parallel sub-agent fan-out, where many hooks may be registered in rapid succession. The design lesson is general: when a store-backed data structure receives frequent writes during concurrent operations, choose a container whose mutation does not trigger change-detection cascades.
102
+
103
+ **Backgrounding with two modes.** Hooks can be backgrounded statically (declared in configuration) or dynamically (the hook emits a backgrounding signal as its first output line before doing real work). A further variant, "rewake" backgrounding, is used for hooks that must survive across multiple agent turns — on completion with a blocking exit code, they inject a task notification that wakes the model rather than appearing in the pending-hook registry. This three-tier backgrounding design reflects the real-world range of hook lifetimes: same-turn, next-turn, and multi-turn.
104
+
105
+ **Permission precedence.** When multiple hooks return conflicting permission decisions for the same tool-use event, the aggregation follows strict precedence: deny overrides ask, ask overrides allow, passthrough is ignored. This ensures that a single security-minded hook can always veto, regardless of how many permissive hooks also respond.
106
+
107
+ **Six hook types with distinct roles.** The system supports command hooks (shell processes), prompt hooks (LLM sub-calls), agent hooks (full sub-agent delegation), HTTP hooks (external service calls with mandatory JSON responses), callback hooks (in-process functions with full structured output control), and function hooks (session-only boolean gates against conversation history). The callback/function split is deliberate: callbacks follow the same protocol as external hooks and can influence permissions, inject context, and rewrite inputs. Function hooks receive only the message history and return a boolean — they exist for lightweight validation without JSON serialization overhead and are never persisted or included in telemetry.
108
+
109
+ **Exit-code discipline.** External command hooks follow a strict three-tier convention: exit 0 is success, exit 2 is a blocking error whose stderr message is injected into the model's context, and any other non-zero exit is a non-blocking warning shown to the user without halting execution. This convention gives shell script authors fine-grained control over agent behavior using only the process exit code.
110
+
111
+ **Deduplication with source-aware keys.** When the same hook command appears in multiple settings scopes (user, project, workspace-local), the merge layer deduplicates using a composite key. For hooks from the same source, last-wins applies. For hooks from different plugin roots, source-specific prefixes in the dedup key prevent cross-plugin collisions from being silently collapsed.