@hybridlabor-api/aos 4.13.1 → 4.14.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/AGENTS.md +8 -0
- package/.agents/nodes.json +5 -2
- package/.claude/hooks/conventional-commits.mjs +14 -15
- package/.claude/hooks/env-file-protection.mjs +14 -15
- package/.claude/hooks/go-gate.mjs +157 -13
- package/.claude/hooks/go-token.mjs +55 -0
- package/.claude/hooks/memb-inject.mjs +75 -62
- package/.claude/hooks/trail-autostart.mjs +27 -0
- package/.claude/settings.json +18 -0
- package/.claude/workflows/startcycle-dispatch.mjs +11 -4
- package/.opencode/plugins/bdb-aos.js +98 -121
- package/.opencode/plugins/lib/trail-autostart.js +38 -0
- package/CLAUDE.md +1 -1
- package/README.de.md +6 -10
- package/README.md +6 -10
- package/README.pt.md +6 -10
- package/THIRD_PARTY_NOTICES.md +19 -3
- package/assets/header-v5.png +0 -0
- package/bin/aos-acp.mjs +211 -0
- package/bin/aos-doctor.mjs +1 -1
- package/bin/aos-uninstall.mjs +2 -2
- package/docs/master-session-acp.md +51 -0
- package/installer.js +398 -65
- package/mcps/mcsc/packages/mcp/server.js +6 -7
- package/package.json +4 -3
- package/scripts/validate-skills.mjs +76 -0
- package/skills/basic/bdbmediastorm/SKILL.md +1 -1
- package/skills/basic/godmode-shipping/SKILL.md +3 -0
- package/skills/basic/master-session/SKILL.md +89 -0
- package/skills/basic/startcycle/SKILL.md +1 -1
- package/skills/basic/startcycle-graph/SKILL.md +2 -2
- package/skills/basic/startcycle-graph-user/SKILL.md +1 -1
- package/skills/basic/teamwork-preview/SKILL.md +1 -1
- package/skills/bdbrainstorm/SKILL.md +7 -1
- package/skills/global_config/agentic-harness-patterns/SKILL.md +257 -0
- package/skills/global_config/agentic-harness-patterns/metadata.json +10 -0
- package/skills/global_config/agentic-harness-patterns/references/agent-orchestration-pattern.md +97 -0
- package/skills/global_config/agentic-harness-patterns/references/bootstrap-sequence-pattern.md +106 -0
- package/skills/global_config/agentic-harness-patterns/references/context-engineering/compress-pattern.md +78 -0
- package/skills/global_config/agentic-harness-patterns/references/context-engineering/isolate-pattern.md +82 -0
- package/skills/global_config/agentic-harness-patterns/references/context-engineering/select-pattern.md +86 -0
- package/skills/global_config/agentic-harness-patterns/references/context-engineering-pattern.md +29 -0
- package/skills/global_config/agentic-harness-patterns/references/hook-lifecycle-pattern.md +111 -0
- package/skills/global_config/agentic-harness-patterns/references/memory-persistence-pattern.md +109 -0
- package/skills/global_config/agentic-harness-patterns/references/permission-gate-pattern.md +111 -0
- package/skills/global_config/agentic-harness-patterns/references/skill-runtime-pattern.md +104 -0
- package/skills/global_config/agentic-harness-patterns/references/task-decomposition-pattern.md +92 -0
- package/skills/global_config/agentic-harness-patterns/references/tool-registry-pattern.md +101 -0
- package/skills/global_config/agenttrail/SKILL.md +8 -0
- package/skills/global_config/agenttrail/bin/agenttrail.mjs +14 -0
- package/skills/global_config/agenttrail/bin/ensure.mjs +118 -0
- package/skills/global_config/aos-setup/scripts/aos-doctor.mjs +1 -1
- package/skills/global_config/bdb-memb-mcp/SKILL.md +4 -5
- package/skills/global_config/bdb-visual-edit/SKILL.md +51 -0
- package/skills/global_config/bdb-visual-edit/references/vite-react-source-attr.md +59 -0
- package/skills/global_config/bdb-visual-edit/scripts/pick-snippet.js +27 -0
- package/skills/global_config/bdb-visual-edit/scripts/sanitize-element.mjs +123 -0
- package/skills/global_config/factory-collect/SKILL.md +74 -0
- package/skills/global_config/factory-human-digest/SKILL.md +92 -0
- package/skills/global_config/factory-lookback/SKILL.md +95 -0
- package/skills/global_config/factory-review-prs/SKILL.md +63 -0
- package/skills/global_config/git-pr-review/SKILL.md +3 -0
- package/skills/global_config/grilling/SKILL.md +2 -0
- package/skills/global_config/mcsc/SKILL.md +1 -1
- package/skills/global_config/plan-arbiter/SKILL.md +125 -0
- package/skills/global_config/plan-canvas/SKILL.md +63 -9
- package/skills/global_config/plan-canvas/scripts/lib/loopback-guard.js +19 -3
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/README.md +285 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/agent-trail.js +129 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/board-client.js +124 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/demo-plan/canvas.mdx +19 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/demo-plan/plan.mdx +18 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/recap-demo/plan.mdx +72 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/README.md +29 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/architecture.json +30 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/builder/00_architecture.html +14950 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/builder/canvas.mdx +511 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/builder/plan.mdx +208 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/recap/plan.mdx +102 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/showcase/standard/plan.md +136 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/signup-storyboard/canvas.mdx +124 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/examples/signup-storyboard/plan.mdx +37 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/index.js +188 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/kit.js +123 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/mdx.js +411 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/render.js +1291 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/architecture/meta.json +1 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/architecture/plan.mdx +195 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/architecture/standard.md +95 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/bugfix/meta.json +1 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/bugfix/plan.mdx +105 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/bugfix/standard.md +76 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/feature/canvas.mdx +81 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/feature/meta.json +1 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/feature/plan.mdx +145 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/feature/standard.md +76 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/migration/meta.json +1 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/migration/plan.mdx +172 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/migration/standard.md +100 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap/meta.json +1 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap/plan.mdx +67 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap/standard.md +49 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-board/canvas.mdx +63 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-board/meta.json +1 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-board/plan.mdx +49 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-board/standard.md +39 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-review/meta.json +1 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-review/plan.mdx +118 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/recap-review/standard.md +57 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/release/meta.json +1 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/release/plan.mdx +173 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/release/standard.md +96 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/research/meta.json +1 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/research/plan.mdx +91 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/research/standard.md +54 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/show-control/canvas.mdx +53 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/show-control/meta.json +1 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/show-control/plan.mdx +225 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/templates/show-control/standard.md +111 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/theme.css +472 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-builder/trail.js +216 -0
- package/skills/global_config/plan-canvas/scripts/lib/plan-canvas/markdown.js +1 -1
- package/skills/global_config/plan-canvas/scripts/lib/plan-canvas/server.js +37 -4
- package/skills/global_config/plan-canvas/scripts/lib/plan-canvas/ui.js +125 -29
- package/skills/global_config/plan-canvas/scripts/plan-canvas.js +196 -8
- package/skills/global_config/pr-recap/SKILL.md +47 -0
- package/skills/global_config/pr-recap/scripts/pr-recap.mjs +200 -0
- package/skills/global_config/quick-recap/SKILL.md +55 -0
- package/skills/global_config/stay-within-limits/SKILL.md +85 -0
- package/skills/global_config/triage/SKILL.md +3 -0
- package/skills/global_config/visual-edit/README.md +96 -0
- package/skills/global_config/visual-edit/SKILL.md +615 -0
- package/skills/global_config/visual-plan/README.md +93 -0
- package/skills/global_config/visual-plan/SKILL.md +544 -0
- package/skills/global_config/visual-plan/references/canvas.md +139 -0
- package/skills/global_config/visual-plan/references/connection.md +51 -0
- package/skills/global_config/visual-plan/references/document-quality.md +186 -0
- package/skills/global_config/visual-plan/references/exemplar.md +62 -0
- package/skills/global_config/visual-plan/references/local-files.md +99 -0
- package/skills/global_config/visual-plan/references/wireframe.md +319 -0
- package/skills/global_config/visual-recap/README.md +103 -0
- package/skills/global_config/visual-recap/SKILL.md +560 -0
- package/skills/global_config/visual-recap/references/connection.md +51 -0
- package/skills/global_config/visual-recap/references/local-files.md +99 -0
- package/skills/global_config/visual-recap/references/wireframe.md +319 -0
- package/skills/playbooks/pb-ci-fix/SKILL.md +49 -0
- package/skills/playbooks/pb-event-tracker/SKILL.md +45 -0
- package/skills/playbooks/pb-meeting-actions/SKILL.md +42 -0
- package/skills/playbooks/pb-project-new/SKILL.md +48 -0
- package/skills/playbooks/pb-week-plan/SKILL.md +45 -0
- package/assets/header-v4.jpg +0 -0
|
@@ -0,0 +1,111 @@
|
|
|
1
|
+
# Hook Lifecycle Pattern
|
|
2
|
+
|
|
3
|
+
## Problem
|
|
4
|
+
|
|
5
|
+
Without a centralized hook lifecycle, extensibility degrades into chaos: hook authors attach side effects at arbitrary points in agent execution with no consistent trust enforcement, no deterministic ordering across concurrent hooks, and no structured way to propagate blocking decisions back to the caller. External hooks may execute before workspace trust has been established, enabling remote code execution from untrusted configuration files. Long-running hooks that background themselves without proper registry handoff leave orphaned processes that deliver stale or contradictory signals at unexpected times.
|
|
6
|
+
|
|
7
|
+
These problems emerge in any agent runtime that supports extensible hook points — they are not specific to a single implementation.
|
|
8
|
+
|
|
9
|
+
## Golden Rules
|
|
10
|
+
|
|
11
|
+
### All hooks flow through a single dispatch point
|
|
12
|
+
|
|
13
|
+
Never allow hook execution to scatter across multiple call sites. A single dispatch function is the chokepoint for trust checks, source merging, type routing, and result aggregation. If a new hook type is added, the dispatch point is the only place that needs to change. Scattered execution sites inevitably diverge on trust enforcement or result handling.
|
|
14
|
+
|
|
15
|
+
### Trust is all-or-nothing
|
|
16
|
+
|
|
17
|
+
Before any hook fires, check workspace trust exactly once. If trust has not been established, skip every hook — not just external ones. In-process hooks included. This prevents a half-trusted state where some hooks execute and others do not, which creates unpredictable behavior that is harder to debug than a clean skip-all.
|
|
18
|
+
|
|
19
|
+
### Multiple sources merge with explicit priority
|
|
20
|
+
|
|
21
|
+
Hooks arrive from multiple sources: persisted configuration, programmatic SDK registration, and ephemeral session-scoped registrations. Merge them into a single ordered list with a defined priority layering. When enterprise policy restricts hooks, entire source layers are excluded cleanly rather than individual hooks being filtered ad hoc.
|
|
22
|
+
|
|
23
|
+
### Exit codes carry semantic weight
|
|
24
|
+
|
|
25
|
+
For external process hooks, reserve specific exit codes for specific meanings. One code means success. One specific code means "block this action and inject my error message into the agent's context." All other non-zero codes mean "warn the user but do not block." This three-tier convention lets external scripts control agent behavior without needing structured JSON output.
|
|
26
|
+
|
|
27
|
+
### Deny beats ask beats allow
|
|
28
|
+
|
|
29
|
+
When multiple hooks in the same batch return conflicting permission decisions, apply a strict precedence: deny overrides ask, ask overrides allow, and passthrough never sets the aggregated decision. This deterministic resolution eliminates ambiguity when hooks disagree.
|
|
30
|
+
|
|
31
|
+
### Session hooks are ephemeral by design
|
|
32
|
+
|
|
33
|
+
Hooks registered at runtime for a specific agent session must be scoped to that session's ID and automatically cleaned up when the session ends. They should never be written to persisted configuration. If a sub-agent registers a hook, that hook must not fire in the parent agent's session or in a sibling sub-agent's session. This isolation prevents cross-session side effects during parallel fan-out.
|
|
34
|
+
|
|
35
|
+
## When To Use
|
|
36
|
+
|
|
37
|
+
- Your agent runtime exposes lifecycle events that external code can hook into.
|
|
38
|
+
- You need trust enforcement that cannot be bypassed by any hook type.
|
|
39
|
+
- Hooks arrive from multiple configuration sources that must be merged deterministically.
|
|
40
|
+
- You support both synchronous blocking hooks and long-running background hooks.
|
|
41
|
+
- Multiple hooks can return conflicting permission decisions for the same action.
|
|
42
|
+
- You need session-scoped hooks that are automatically cleaned up when a session ends.
|
|
43
|
+
|
|
44
|
+
## Tradeoffs
|
|
45
|
+
|
|
46
|
+
| Decision | Benefit | Cost |
|
|
47
|
+
|---|---|---|
|
|
48
|
+
| Single dispatch point | One place for trust, routing, and aggregation | Every hook type must conform to the same dispatch interface |
|
|
49
|
+
| All-or-nothing trust gate | No half-trusted states, simple mental model | In-process hooks that pose no security risk are still blocked |
|
|
50
|
+
| Multi-source merge with priority | Clean policy enforcement, predictable ordering | Hook authors must understand which source layer they belong to |
|
|
51
|
+
| Multiple hook types | Right tool for each job (process, LLM, HTTP, in-process) | More type surface to document and maintain |
|
|
52
|
+
| Async generator composition | Streaming partial results, early exit on blocking errors | Callers must handle incremental accumulation |
|
|
53
|
+
| Backgrounding with registry handoff | Long-running hooks do not block the main loop | Orphan risk if the registry handoff is skipped |
|
|
54
|
+
| Session-scoped hooks | Ephemeral hooks auto-clean on session end | Hooks scoped to one session do not fire in sub-agent sessions |
|
|
55
|
+
| Deduplication across sources | No double-firing when the same hook appears in multiple configs | Dedup key design must account for source-specific prefixes |
|
|
56
|
+
|
|
57
|
+
## Implementation Patterns
|
|
58
|
+
|
|
59
|
+
- Route every hook invocation through a single dispatch function. This function is the sole location for the trust check, source merge, type dispatch, and result aggregation.
|
|
60
|
+
- Implement the trust check as the first operation in dispatch. If trust is absent, return immediately with an empty result — do not evaluate which hooks would have matched.
|
|
61
|
+
- Merge hook sources in a defined priority order (e.g., persisted config, then SDK-registered, then session-scoped). When a policy flag restricts to managed hooks only, exclude entire source layers rather than filtering individual entries.
|
|
62
|
+
- Deduplicate across sources using a composite key that includes source-specific context. Two hooks with the same command from different plugins are distinct; the same command from user config and project config collapses to the last-wins entry.
|
|
63
|
+
- Support at least these hook type categories: external process (shell command), LLM sub-call, sub-agent delegation, HTTP endpoint, in-process callback with full output control, and lightweight boolean gate for simple allow/deny decisions.
|
|
64
|
+
- Run all matched hooks in parallel using async generator composition. Each hook yields typed result fragments as they become available. The dispatch function merges all generators and yields aggregated fragments to the caller.
|
|
65
|
+
- For external process hooks, map exit codes to semantic outcomes: one code for success, one reserved code for blocking errors (with the error message on stderr), and all other non-zero codes for non-blocking warnings.
|
|
66
|
+
- Support two backgrounding modes: static (declared in hook configuration before execution) and dynamic (the hook signals backgrounding as its first output line). Backgrounded hooks are handed off to a pending-hook registry so they are tracked across turns.
|
|
67
|
+
- Scope session hooks to a specific agent or session ID. Hooks registered under the main session do not fire in sub-agent sessions. Provide explicit removal functions, but also clear all session hooks automatically when the session ends.
|
|
68
|
+
- When aggregating permission decisions from multiple hooks, apply strict precedence: deny wins over ask, ask wins over allow, passthrough is ignored.
|
|
69
|
+
- Choose hook types based on the capability needed. Use external process hooks for shell-level side effects. Use LLM sub-call hooks when the hook itself needs reasoning. Use HTTP hooks for external service integration. Use in-process callbacks when the hook needs full control over structured output fields (permission decisions, injected context, input rewriting). Use lightweight boolean gates when you only need a yes/no decision against conversation history without JSON serialization overhead.
|
|
70
|
+
- Register a process-level cleanup handler that terminates all running hooks on shutdown. Without this, backgrounded hooks survive the parent process and become orphans.
|
|
71
|
+
- Guard against event-type restrictions: some events may be incompatible with certain hook types. For example, HTTP hooks during early session setup can deadlock if the event fires before the response consumer is ready. Document these restrictions in the dispatch layer rather than relying on hook authors to discover them.
|
|
72
|
+
|
|
73
|
+
## Gotchas
|
|
74
|
+
|
|
75
|
+
**The trust gate blocks all hook types, not just external ones.** Every hook — including in-process callbacks and boolean gates — is skipped when trust is absent. Never assume a hook fired just because it runs in-process. If a hook did not fire, check trust status first.
|
|
76
|
+
|
|
77
|
+
**A missing external script can produce the same exit code as an intentional block.** If a hook script is deleted between registration and execution, the shell itself may exit with the reserved blocking code. Validate that hook scripts exist at registration time, not only at execution time.
|
|
78
|
+
|
|
79
|
+
**Policy-restricted modes silently drop session hooks.** When enterprise policy limits hooks to managed sources only, session-scoped hooks registered by skills or sub-agents receive zero matches with no error. Tested behavior may differ between managed and unmanaged deployments.
|
|
80
|
+
|
|
81
|
+
**Background hooks with "rewake" semantics bypass the pending registry.** Normal backgrounded hooks appear in the registry and are resolved before the next turn. Rewake-style hooks survive across turns by design — they inject a notification when they complete. Do not use rewake for hooks that must finish before the next tool call.
|
|
82
|
+
|
|
83
|
+
**Boolean gate hooks cannot be serialized.** Lightweight in-process gates that return only true/false are inherently ephemeral. They are excluded from configuration snapshots, analytics, and telemetry. Use the full callback type if you need persistence or observability.
|
|
84
|
+
|
|
85
|
+
**Deduplication uses last-wins, which can surprise multi-scope authors.** When the same hook command appears in both user-level and project-level config, the project-level entry wins. Source-specific prefixes in the dedup key prevent cross-source collisions, but same-source collisions are resolved silently.
|
|
86
|
+
|
|
87
|
+
**HTTP hooks can deadlock during early lifecycle events.** If an HTTP hook is registered for a session-start or setup event, the outbound request may fire before the response consumer is ready, creating a deadlock. The dispatch layer should explicitly filter incompatible hook-type/event-type combinations rather than leaving this to the hook author.
|
|
88
|
+
|
|
89
|
+
**Session hooks are invisible after session cleanup.** Once a session ends and its hooks are cleared, there is no record they ever existed. If debugging requires understanding which hooks were active during a past session, you need separate logging — the session store itself is ephemeral.
|
|
90
|
+
|
|
91
|
+
## Claude Code Evidence
|
|
92
|
+
|
|
93
|
+
Claude Code's hook system is a production implementation of these principles, supporting six hook types across dozens of lifecycle events.
|
|
94
|
+
|
|
95
|
+
**Single dispatch with trust gate.** All hook execution flows through one async generator function. The first thing it does is check workspace trust. In interactive mode, this requires the user to have accepted a trust dialog. In SDK (non-interactive) mode, trust is implicit. If trust is absent, the generator returns immediately — no hooks of any type fire.
|
|
96
|
+
|
|
97
|
+
**Three-layer source merge with policy enforcement.** Hooks are drawn from three sources in priority order: snapshot hooks from settings files, SDK-registered callback hooks, and session-scoped hooks stored in a Map-based session state. When enterprise policy sets a managed-only flag, session and plugin hooks are excluded at the source-merge layer, not filtered individually. This design means policy enforcement is a single conditional that drops entire source lists, rather than per-hook filtering that would need to be maintained for every new hook type.
|
|
98
|
+
|
|
99
|
+
**Async generator composition for streaming results.** Each hook runs as an independent async generator yielding typed result fragments — blocking errors, permission decisions, additional context, modified inputs. A parallel merge utility combines all generators so the caller receives fragments as they arrive. This is critical for blocking errors: a slow hook in the batch does not delay the blocking signal from a fast hook.
|
|
100
|
+
|
|
101
|
+
**Map-based session state for concurrency.** Session hooks use a Map rather than a plain object for the session store. Map mutation returns the same container reference, so the application store's equality check sees no change and none of the store's listeners fire. This avoids quadratic listener-notification costs during parallel sub-agent fan-out, where many hooks may be registered in rapid succession. The design lesson is general: when a store-backed data structure receives frequent writes during concurrent operations, choose a container whose mutation does not trigger change-detection cascades.
|
|
102
|
+
|
|
103
|
+
**Backgrounding with two modes.** Hooks can be backgrounded statically (declared in configuration) or dynamically (the hook emits a backgrounding signal as its first output line before doing real work). A further variant, "rewake" backgrounding, is used for hooks that must survive across multiple agent turns — on completion with a blocking exit code, they inject a task notification that wakes the model rather than appearing in the pending-hook registry. This three-tier backgrounding design reflects the real-world range of hook lifetimes: same-turn, next-turn, and multi-turn.
|
|
104
|
+
|
|
105
|
+
**Permission precedence.** When multiple hooks return conflicting permission decisions for the same tool-use event, the aggregation follows strict precedence: deny overrides ask, ask overrides allow, passthrough is ignored. This ensures that a single security-minded hook can always veto, regardless of how many permissive hooks also respond.
|
|
106
|
+
|
|
107
|
+
**Six hook types with distinct roles.** The system supports command hooks (shell processes), prompt hooks (LLM sub-calls), agent hooks (full sub-agent delegation), HTTP hooks (external service calls with mandatory JSON responses), callback hooks (in-process functions with full structured output control), and function hooks (session-only boolean gates against conversation history). The callback/function split is deliberate: callbacks follow the same protocol as external hooks and can influence permissions, inject context, and rewrite inputs. Function hooks receive only the message history and return a boolean — they exist for lightweight validation without JSON serialization overhead and are never persisted or included in telemetry.
|
|
108
|
+
|
|
109
|
+
**Exit-code discipline.** External command hooks follow a strict three-tier convention: exit 0 is success, exit 2 is a blocking error whose stderr message is injected into the model's context, and any other non-zero exit is a non-blocking warning shown to the user without halting execution. This convention gives shell script authors fine-grained control over agent behavior using only the process exit code.
|
|
110
|
+
|
|
111
|
+
**Deduplication with source-aware keys.** When the same hook command appears in multiple settings scopes (user, project, workspace-local), the merge layer deduplicates using a composite key. For hooks from the same source, last-wins applies. For hooks from different plugin roots, source-specific prefixes in the dedup key prevent cross-plugin collisions from being silently collapsed.
|
package/skills/global_config/agentic-harness-patterns/references/memory-persistence-pattern.md
ADDED
|
@@ -0,0 +1,109 @@
|
|
|
1
|
+
# Memory and Persistence Pattern
|
|
2
|
+
|
|
3
|
+
## Problem
|
|
4
|
+
|
|
5
|
+
Without persistent memory, an agent loses all user preferences, project context, and behavioral feedback the moment a session ends. Users must repeat corrections every session ("use bun, not npm"), and the agent cannot accumulate the working knowledge that makes it genuinely useful over time. A single flat memory store also fails: instruction-level rules need different durability than transient working notes, and organization-wide conventions must coexist with per-project and per-user overrides without colliding.
|
|
6
|
+
|
|
7
|
+
These problems emerge in any agent runtime that persists across sessions — they are not specific to a single implementation.
|
|
8
|
+
|
|
9
|
+
## Golden Rules
|
|
10
|
+
|
|
11
|
+
### Separate layers by scope and durability
|
|
12
|
+
|
|
13
|
+
Not all memory is equal. Instruction-level rules ("always use tabs") belong in version-controlled files that travel with the project. Personal workflow preferences ("open diffs in VS Code") belong in a user-scoped store that spans all projects. Transient working notes belong in an auto-memory layer that the agent manages itself. Merging these into a single store destroys the ability to share, audit, or override each independently. Design at least three tiers: organization/team, user, and project — with a local override that is never shared.
|
|
14
|
+
|
|
15
|
+
### Local overrides win — always
|
|
16
|
+
|
|
17
|
+
When the same topic is addressed at multiple scopes, the most-local instruction takes priority. Organization-wide policy sets the floor, user-level preferences narrow it, project-level conventions narrow further, and a local override file has final say. Concatenate layers in ascending priority order so the most-local content appears last and receives the most model attention. Builders must understand this hierarchy or they will set a global rule expecting it to dominate, only to have a local file silently override it.
|
|
18
|
+
|
|
19
|
+
### Taxonomize auto-memory by type, not by recency
|
|
20
|
+
|
|
21
|
+
Auto-memory entries should each carry a type drawn from a small, fixed taxonomy — who the user is, behavioral corrections, project context not derivable from code, and stable reference facts. A type taxonomy prevents the memory store from devolving into a chronological dump and makes it possible to audit, prune, and promote entries systematically. Exclude anything derivable from the codebase or version history — saving it wastes index space and goes stale.
|
|
22
|
+
|
|
23
|
+
### Two-step save: write the topic file, then update the index
|
|
24
|
+
|
|
25
|
+
Every memory write is a two-step operation: first write the full content to a dedicated topic file, then append a one-line pointer to the index. If the process crashes between steps, the worst outcome is an orphaned topic file — the index remains consistent. Never write memory content directly into the index; the index is a table of contents, not a document.
|
|
26
|
+
|
|
27
|
+
### The index is bounded always-on context; topic files are on-demand detail
|
|
28
|
+
|
|
29
|
+
The memory index is loaded into every conversation — its cost is paid on every turn. Keep it small: one line per entry, under a hard line-count and byte-count cap with a truncation warning when the cap fires. Full detail lives in topic files that the agent recalls on demand. This separation keeps context cost constant regardless of how much the agent has learned.
|
|
30
|
+
|
|
31
|
+
### Background extraction writes directly to auto-memory
|
|
32
|
+
|
|
33
|
+
Session memory extraction should run as a background process that reads the session transcript after the agent produces a final response. The extractor directly writes auto-memory entries — topic files and index updates — following the same two-step save invariant used by the main agent. This keeps extraction off the critical path (no added latency to user-visible responses) while producing durable memories autonomously. Human control is preserved at the cross-layer boundary: promoting auto-memory entries to project conventions, personal instructions, or team-shared memory is a separate propose-and-review workflow that never applies changes without explicit approval.
|
|
34
|
+
|
|
35
|
+
### The main agent and the extraction agent are mutually exclusive
|
|
36
|
+
|
|
37
|
+
If the main agent wrote to memory during a turn (because the user explicitly asked it to remember something), the background extractor skips that turn entirely and advances its cursor. Two writers targeting the same memory store in the same turn will produce conflicts. Enforce mutual exclusion per turn, not per file.
|
|
38
|
+
|
|
39
|
+
### Sandbox the extraction agent
|
|
40
|
+
|
|
41
|
+
The background extractor runs with restricted permissions: it may read freely but may only write to the auto-memory directory, and it must not execute arbitrary commands. Cap its turn count (a small number like five) to prevent verification rabbit-holes. It shares the parent's prompt cache to keep cost low, but it must not inherit the parent's full write surface.
|
|
42
|
+
|
|
43
|
+
## When To Use
|
|
44
|
+
|
|
45
|
+
- Your agent persists across sessions and must recall user preferences, project context, or behavioral corrections.
|
|
46
|
+
- Multiple scopes of instruction coexist (organization, user, project, local) and need a clear priority ordering.
|
|
47
|
+
- The agent should learn from sessions without the user manually curating a knowledge base.
|
|
48
|
+
- You need a cross-layer review process that proposes promotions between memory tiers.
|
|
49
|
+
- Your runtime supports background work and you want extraction to happen without blocking the user.
|
|
50
|
+
|
|
51
|
+
## Tradeoffs
|
|
52
|
+
|
|
53
|
+
| Decision | Benefit | Cost |
|
|
54
|
+
|---|---|---|
|
|
55
|
+
| Layered memory with four+ scopes | Each scope can be shared, audited, and overridden independently | More files to discover and concatenate at startup |
|
|
56
|
+
| Local-wins priority ordering | Users can always override without touching shared files | A global rule can be silently overridden — surprising if untested with the full stack |
|
|
57
|
+
| Type taxonomy for auto-memory | Structured pruning, promotion, and audit | Taxonomy must be defined upfront; miscategorized entries degrade quality |
|
|
58
|
+
| Two-step save (topic file then index) | Crash-safe — index never points at missing content, orphaned files are harmless | Two writes per save; cleanup of orphaned topic files is a separate concern |
|
|
59
|
+
| Bounded index with on-demand topic files | Constant context cost regardless of total memory volume | Agent must perform an extra retrieval step to access full detail |
|
|
60
|
+
| Background extraction via forked agent | No latency added to user-visible responses; shares prompt cache | Race window between extraction and next user turn; failed extractions may retry stale content |
|
|
61
|
+
| Mutual exclusion per turn | No write conflicts between main agent and extractor | Extraction is skipped entirely when the main agent writes — some turns produce no extraction |
|
|
62
|
+
| Sandboxed extractor with turn cap | Prevents runaway loops and unintended side effects | May miss nuanced memories that require deeper reasoning |
|
|
63
|
+
| Propose-not-auto-write for cross-layer promotion | Human stays in control of what enters shared or version-controlled memory | Adds a manual review step; unreviewed proposals accumulate |
|
|
64
|
+
|
|
65
|
+
## Implementation Patterns
|
|
66
|
+
|
|
67
|
+
- Define a small, fixed taxonomy of auto-memory types (e.g., user identity, behavioral feedback, project context, stable reference) and document what belongs in each. Exclude anything derivable from the codebase or version history.
|
|
68
|
+
- Ensure the memory directory exists before the agent's first write via an idempotent creation step. Do not ask the agent to check for existence at runtime.
|
|
69
|
+
- Treat the memory index as a table of contents: one line per entry with a title, a link to the topic file, and a one-line summary hook. Enforce a per-entry character budget (around 150 characters) to stay within the index's hard caps.
|
|
70
|
+
- Give each topic file structured frontmatter with at minimum a name, description, and type field drawn from the taxonomy.
|
|
71
|
+
- Implement the two-step save invariant: write the topic file first, then append to the index. Never reverse the order.
|
|
72
|
+
- Discover instruction files at startup by walking a known hierarchy of locations (organization-managed, user home, project root and ancestors, local override). Concatenate in ascending priority order.
|
|
73
|
+
- Support an include directive in instruction files so that large configurations can be composed from smaller fragments.
|
|
74
|
+
- Fire background extraction only after the agent produces a final response with no pending tool calls. If the main agent wrote to memory during that turn, skip extraction and advance the cursor.
|
|
75
|
+
- Restrict the extraction agent's write surface to the auto-memory directory. Allow reads broadly but deny arbitrary command execution.
|
|
76
|
+
- Cap extraction turns at a small number to prevent verification loops. Coalesce concurrent extraction requests so only one runs at a time.
|
|
77
|
+
- Drain in-flight extractions before process shutdown, after the response is flushed but before any shutdown failsafe timer fires.
|
|
78
|
+
- Build a review mechanism that audits across all layers and proposes promotions (to project conventions, personal instructions, team-shared memory, or remain in auto-memory) — but never applies changes without explicit user approval.
|
|
79
|
+
- Provide an environment variable or setting to disable auto-memory entirely for contexts where persistence is undesirable (ephemeral runs, CI, remote sandboxes).
|
|
80
|
+
|
|
81
|
+
## Gotchas
|
|
82
|
+
|
|
83
|
+
**Index truncation is silent until it fires.** Hard caps on line count and byte count are enforced at read time. Long-line entries (multi-sentence summaries instead of one-line hooks) can hit the byte cap while staying under the line cap. Keep entries short and put detail in topic files.
|
|
84
|
+
|
|
85
|
+
**Priority ordering is counterintuitive.** Local overrides beat project rules, which beat user rules, which beat organization rules. If you inject a rule at the user level expecting it to dominate, a local override file in the project root will silently win. Always test with the full instruction-file stack present.
|
|
86
|
+
|
|
87
|
+
**Extraction timing creates a race window.** The background extractor fires at the end of a response, but the user can start the next turn before extraction completes. The overlap guard coalesces concurrent calls and the cursor only advances after a successful run — but a failed extraction means those messages are reconsidered next time, potentially producing duplicate proposals.
|
|
88
|
+
|
|
89
|
+
**Derivable content does not belong in memory.** Architecture, code patterns, and version history are re-derivable from the project itself. Saving them wastes index space and goes stale. The type taxonomy should exclude derivable content by design, but agents will still attempt to save it unless the prompt explicitly forbids it.
|
|
90
|
+
|
|
91
|
+
**Team-shared memory depends on auto-memory being enabled.** If auto-memory is disabled (via environment variable or settings), team memory is also disabled because the shared layer builds on the same directory and index infrastructure. Disabling one silently disables both.
|
|
92
|
+
|
|
93
|
+
**Do not use memory for within-session state.** Plans, task lists, and scratchpads are better handled by dedicated in-session primitives. Memory is for cross-session recall only. Mixing the two inflates the index with ephemeral content that is stale by the next session.
|
|
94
|
+
|
|
95
|
+
**Orphaned topic files are harmless but accumulate.** If the process crashes after writing a topic file but before updating the index, the topic file becomes an orphan. Orphans do not corrupt the index, but they consume disk space. A periodic sweep that deletes topic files not referenced by the index is a reasonable maintenance step.
|
|
96
|
+
|
|
97
|
+
## Claude Code Evidence
|
|
98
|
+
|
|
99
|
+
Claude Code's memory system is a production implementation of these principles across five cooperating subsystems:
|
|
100
|
+
|
|
101
|
+
**Four-level instruction hierarchy.** Claude Code discovers CLAUDE.md files at four scopes: a managed location for organization-wide policy, a user-home location for personal instructions, project-root and ancestor-directory locations for project conventions (including a rules subdirectory), and a CLAUDE.local.md file for private per-project overrides that is never checked into version control. Files are concatenated in ascending priority order so local overrides appear last in the prompt. An @include directive allows composition from fragments. A recommended per-file character cap prevents any single instruction file from dominating context.
|
|
102
|
+
|
|
103
|
+
**Four-type auto-memory with bounded index.** Auto-memory lives in a per-project directory with a MEMORY.md index and individual topic files. Each topic file carries YAML frontmatter with name, description, and type fields. The four types — user, feedback, project, and reference — are documented in the prompt so the agent categorizes entries consistently. The index is capped at 200 lines and 25,000 bytes; a truncation warning is appended when either cap fires. The directory is created idempotently before prompt construction so the agent never encounters a missing path.
|
|
104
|
+
|
|
105
|
+
**Team memory as a shared extension.** When a feature gate and an enablement check both pass, a second directory within the same project slug is added for team-shared memory. It uses the same four-type taxonomy and index structure. Team memory requires auto-memory to be enabled — disabling auto-memory via environment variable also disables team memory.
|
|
106
|
+
|
|
107
|
+
**Background session extraction with mutual exclusion.** At the end of each query loop, a forked sub-agent examines the session transcript and writes memories directly to the auto-memory directory. The fork shares the parent's prompt cache to keep cost low. If the main agent already wrote to a memory path during that turn, the extractor skips entirely and advances its read cursor. The extractor's tool permissions are restricted: it may read, search, and glob freely, but may only write to the auto-memory directory and may not execute arbitrary shell commands. A hard turn cap of five prevents verification rabbit-holes. Concurrent extraction requests are coalesced so only one runs at a time, with a trailing run picking up any messages that arrived during the active run. In-flight extractions are drained after the response is flushed but before the shutdown failsafe timer fires.
|
|
108
|
+
|
|
109
|
+
**Propose-not-auto-write review skill.** A bundled "remember" skill audits across all memory layers and proposes promotions grouped by action: promote to project conventions, promote to personal instructions, promote to team memory, clean up, ambiguous, or no action. It never applies changes autonomously — it presents a structured report and waits for explicit user approval before touching any files. The four promotion destinations mirror the four instruction-file scopes.
|
|
@@ -0,0 +1,111 @@
|
|
|
1
|
+
# Permission Gate Pattern
|
|
2
|
+
|
|
3
|
+
## Problem
|
|
4
|
+
|
|
5
|
+
Without a single permission gate, every tool invocation either silently proceeds (creating security risks) or blindly blocks (creating friction). The core difficulty is that permission decisions depend on context that varies across execution environments: interactive sessions need a human in the loop, headless agents must auto-decide, and swarm workers must delegate to a coordinator. Scattering these checks inside individual tools produces duplicated logic, inconsistent behavior when abort signals arrive mid-check, and no clean way to record audit decisions across all paths.
|
|
6
|
+
|
|
7
|
+
These problems emerge in any agent runtime that supports tool use — they are not specific to a single implementation.
|
|
8
|
+
|
|
9
|
+
## Golden Rules
|
|
10
|
+
|
|
11
|
+
### Every tool invocation passes through one gate
|
|
12
|
+
|
|
13
|
+
Route every tool call through a single permission checkpoint before execution begins. The gate evaluates rules and returns exactly one of three behaviors: allow, deny, or ask. No tool should perform its own authorization bypass; the gate is the sole authority. This makes auditing trivial — every decision flows through one chokepoint — and prevents individual tools from silently skipping checks.
|
|
14
|
+
|
|
15
|
+
### Three behaviors, no more
|
|
16
|
+
|
|
17
|
+
The permission vocabulary is exactly three words. "Allow" means the tool runs immediately. "Deny" means the tool is refused and the agent receives a rejection reason. "Ask" means the runtime cannot decide alone and must consult an external authority — a human, a coordinator, or an escalation hook. Carrying a discriminated reason on each behavior enables downstream logging and display without re-parsing the decision.
|
|
18
|
+
|
|
19
|
+
### Layered rules evaluate in strict priority order
|
|
20
|
+
|
|
21
|
+
Rule evaluation follows a fixed priority chain. Explicit deny rules fire first — they are the fastest and most restrictive. Explicit ask rules fire next. Then tool-specific content checks run (for tools that need logic beyond the general system, such as path-based access control). After tool-specific checks, bypass-immune safety checks fire. Only after all of these does the system consult mode-level overrides (like "auto-approve everything") or broad allow rules. The default, if nothing matches, is ask. This ordering guarantees that safety-critical constraints are never silently overridden by a permissive mode.
|
|
22
|
+
|
|
23
|
+
### Some "ask" results are bypass-immune
|
|
24
|
+
|
|
25
|
+
Certain conditions must always surface as "ask" regardless of how permissive the runtime mode is. Tools that inherently require a human present, content-level rules that explicitly demand confirmation, and safety checks for protected paths (such as version-control internals or agent configuration directories) all produce ask results that skip the auto-approve shortcut entirely. Without this immunity, a misconfigured bypass mode could silently mutate protected state.
|
|
26
|
+
|
|
27
|
+
### Rule evaluation is separated from UX dispatch
|
|
28
|
+
|
|
29
|
+
The decision about what to do (allow, deny, or ask) lives in a dedicated rule evaluator, separate from the UX dispatch layer. The decision about how to present an "ask" result — show a dialog, send a message to a coordinator, forward to a swarm leader — lives in environment-specific handlers. The evaluator itself may carry state (denial tracking, mode transforms) as side effects of evaluation, but the key invariant is that adding a new execution environment requires only a new handler and one dispatch branch, not changes to the rule logic.
|
|
30
|
+
|
|
31
|
+
### Resolution is race-safe through atomic claim
|
|
32
|
+
|
|
33
|
+
When multiple concurrent paths can resolve an "ask" — user input, background hooks, and an automated classifier all racing — the first to finish must win cleanly. An atomic claim mechanism marks the permission as resolved before any asynchronous work proceeds, preventing a second async winner from calling the resolver after the first has already committed. Without this, two paths can both believe they won, producing duplicate or contradictory resolutions.
|
|
34
|
+
|
|
35
|
+
### The permission context is a frozen service object
|
|
36
|
+
|
|
37
|
+
Bundle all stateful operations — abort-signal checking, queue management, audit logging, and rule persistence — into a single context object, then freeze it. Handlers receive this object as their only interface to permission state. Freezing prevents accidental mutation across handler boundaries, while consolidating operations onto the object prevents handlers from duplicating wiring code. This is not a mutable options bag that callers can patch; it is a sealed service contract.
|
|
38
|
+
|
|
39
|
+
### Rule sources are ordered and composable
|
|
40
|
+
|
|
41
|
+
Permission rules come from multiple origins: user-level settings, project-level settings, local overrides, feature flags, policy directives, CLI arguments, command-level scopes, and session-level grants. Each source contributes its own deny, ask, and allow lists independently. At evaluation time, all sources are merged and the strict priority chain applies across the combined set. Persistence targets the appropriate source — "remember for this project" writes to project settings, "remember for this session" writes to session state — so that rules compose without overwriting each other.
|
|
42
|
+
|
|
43
|
+
## When To Use
|
|
44
|
+
|
|
45
|
+
- Your agent runtime executes tools that can modify files, run commands, or access external services.
|
|
46
|
+
- You need consistent permission enforcement across interactive, headless, and multi-agent environments.
|
|
47
|
+
- You need an audit trail showing who authorized each tool invocation and why.
|
|
48
|
+
- Your tools have heterogeneous risk profiles — some are safe to auto-approve, others must always require confirmation.
|
|
49
|
+
- You support multiple configuration scopes (user, project, session) that need to compose without conflict.
|
|
50
|
+
- You need to add new execution environments without rewriting your permission logic.
|
|
51
|
+
|
|
52
|
+
## Tradeoffs
|
|
53
|
+
|
|
54
|
+
| Decision | Benefit | Cost |
|
|
55
|
+
|---|---|---|
|
|
56
|
+
| Single gate for all tools | One audit chokepoint, consistent enforcement | Every tool call pays the gate's latency |
|
|
57
|
+
| Three-behavior vocabulary | Simple mental model, exhaustive case handling | "Ask" is a broad bucket — different environments handle it differently |
|
|
58
|
+
| Strict priority ordering | Deny always wins, safety checks cannot be bypassed | Rule authors must understand the evaluation order to predict outcomes |
|
|
59
|
+
| Bypass-immune ask results | Protected paths stay protected regardless of mode | Some tools feel slower than expected even in auto-approve mode |
|
|
60
|
+
| Separated evaluation and dispatch | New environments need only a new handler | Two layers to understand instead of one |
|
|
61
|
+
| Atomic claim for resolution | No duplicate or contradictory resolutions | Slightly more complex handler authoring |
|
|
62
|
+
| Frozen context object | Handlers cannot corrupt shared state | No ad-hoc extensions at call sites — all operations must be defined upfront |
|
|
63
|
+
| Multi-source rule composition | Fine-grained scoping (user, project, session) | Merge logic is more complex; debugging "which source won?" requires tracing |
|
|
64
|
+
|
|
65
|
+
## Implementation Patterns
|
|
66
|
+
|
|
67
|
+
- Define a permission decision type with exactly three behaviors: allow, deny, and ask. Attach a discriminated reason to each so that logging and display messages never need to re-derive the cause.
|
|
68
|
+
- Implement tool-specific permission checks only when a tool needs logic beyond the general rule system (path-based access control, quota enforcement, etc.). The default should delegate entirely to the rule-based system. Return deny only for clear violations; return ask for uncertain cases.
|
|
69
|
+
- Mark any tool that cannot function without a human present as requiring user interaction. The gate always surfaces these as ask regardless of bypass mode.
|
|
70
|
+
- Tag safety-critical ask results with a distinct reason type. The gate treats these as bypass-immune and skips any auto-approve shortcut.
|
|
71
|
+
- Build the permission context as a sealed service object that bundles abort checking, queue operations, audit logging, and rule persistence. Freeze it at creation time.
|
|
72
|
+
- Use an atomic claim guard on any promise that can be resolved by multiple concurrent paths. Claim before any asynchronous work inside an async callback.
|
|
73
|
+
- Register remote or swarm callbacks before sending the outbound request to eliminate the race where a responder answers before the callback exists.
|
|
74
|
+
- Log every decision — accept, reject, source, whether it was persisted durably — at the point of resolution, not at individual call sites. Centralizing this on the context object prevents missed logging when new resolution paths are added.
|
|
75
|
+
- Persist permission updates through both in-memory state and durable storage in a single operation. Return whether any update was durable so the audit log can distinguish ephemeral from permanent grants.
|
|
76
|
+
- Re-read runtime mode at the mode-check step, not at function entry. If the mode changes between rule evaluation and the mode shortcut, a cached value will apply stale permissions.
|
|
77
|
+
- When converting ask to deny in a "don't ask" mode, perform the conversion after the inner rule chain completes. This ensures bypass-immune checks always produce ask first, which is then converted to deny at the outer layer. Merging these steps risks suppressing bypass-immune paths.
|
|
78
|
+
|
|
79
|
+
## Gotchas
|
|
80
|
+
|
|
81
|
+
**Do not cache runtime mode across the entire evaluation.** The mode-check step must re-read the current mode at evaluation time. If you snapshot the mode at function entry, a mode switch that occurs between rule evaluation and the bypass shortcut will be invisible, and you will apply stale permissions.
|
|
82
|
+
|
|
83
|
+
**Do not run hooks inside the rule evaluator.** Hook execution belongs in environment-specific handlers. If you move it into the evaluator, headless agents that never reach a handler will silently skip hooks. Headless environments need a separate hook execution path precisely because they do not go through the interactive handler.
|
|
84
|
+
|
|
85
|
+
**Allow results from rules still need input normalization.** Even when the gate resolves early via a bypass shortcut or a broad allow rule, the tool's input may need scrubbing. Content-specific allow rules (such as prefix-match rules for shell commands) can return a normalized input that differs from the raw invocation. Always apply input normalization before passing through to execution.
|
|
86
|
+
|
|
87
|
+
**Do not merge the inner rule chain with the outer mode conversion.** The inner chain produces bypass-immune ask results. The outer layer converts ask to deny in restrictive modes. If you flatten these into one function, bypass-immune results may not trigger correctly because the deny conversion can intercept them before they are tagged as immune.
|
|
88
|
+
|
|
89
|
+
**Add a grace period before honoring interaction signals in the interactive handler.** Without a short delay (on the order of a few hundred milliseconds), incidental keystrokes — such as the Enter that submitted the original prompt — can cancel a running classifier before it has time to auto-approve. The grace period absorbs input that arrived before the permission dialog was actually visible.
|
|
90
|
+
|
|
91
|
+
**Freezing the context object is not a substitute for atomic claim.** Freezing prevents external mutation of the context fields but does nothing for async race conditions on the resolution promise. Both mechanisms are required: freeze for encapsulation, atomic claim for resolution safety.
|
|
92
|
+
|
|
93
|
+
**Register callbacks before sending outbound requests in distributed environments.** In swarm or coordinator topologies, the responder may reply before the local handler has registered its callback. Registering first eliminates this window.
|
|
94
|
+
|
|
95
|
+
## Claude Code Evidence
|
|
96
|
+
|
|
97
|
+
Claude Code's permission system is a production implementation of these principles, handling tool authorization across interactive terminals, headless CI agents, and multi-agent swarm topologies.
|
|
98
|
+
|
|
99
|
+
**Single gate, three behaviors.** Every tool invocation in Claude Code passes through one top-level permission function before execution. The function returns allow, deny, or ask, each carrying a discriminated reason that drives both the audit log and the user-facing explanation message. No tool bypasses this gate.
|
|
100
|
+
|
|
101
|
+
**Strict layered evaluation.** The rule evaluator follows a fixed sequence: explicit deny rules, explicit ask rules, tool-specific content checks, bypass-immune safety checks (for paths like `.git/` and `.claude/`), mode-level shortcuts, broad allow rules, and a default fallback to ask. The ordering is deliberate — safety checks fire before mode shortcuts, so even full bypass mode cannot silently approve writes to protected directories.
|
|
102
|
+
|
|
103
|
+
**Bypass-immune safety checks.** Three categories of ask results skip the auto-approve shortcut entirely: tools that declare they require user interaction, content-level ask rules that explicitly demand confirmation, and a safety check for protected paths. This design was motivated by the observation that a single misconfigured "approve everything" flag could otherwise allow silent mutation of version-control state or agent configuration files.
|
|
104
|
+
|
|
105
|
+
**Frozen permission context.** The context object is frozen at creation time. All handler operations — queue push and remove, permission persistence, abort-signal checking, audit logging — are methods on this sealed object. This prevents handlers in different execution environments from accidentally patching each other's state, which was a recurring source of bugs during early development of the multi-handler architecture.
|
|
106
|
+
|
|
107
|
+
**Atomic claim for resolution races.** The interactive handler races three concurrent paths: user input from the confirmation dialog, background permission hooks, and an automated classifier. A claim mechanism atomically marks the permission as resolved before any of these paths proceeds with asynchronous work. The design was introduced after observing that without it, both the classifier and the user could "win" the race, producing duplicate resolution callbacks.
|
|
108
|
+
|
|
109
|
+
**Three-way handler dispatch.** After rule evaluation, the system dispatches ask results to one of three handlers based on execution environment: a coordinator handler that runs hooks and classifiers sequentially before showing a dialog, a swarm-worker handler that forwards the decision to a leader via mailbox, or a default interactive handler that pushes a confirmation dialog and races user input against automated checks. Adding a new execution environment requires only a new handler and one dispatch branch — the rule evaluator does not change.
|
|
110
|
+
|
|
111
|
+
**Composable multi-source rules.** Permission rules come from eight distinct sources spanning user settings, project settings, local overrides, feature flags, policy directives, CLI arguments, command-level scopes, and session grants. Each source contributes independent deny, ask, and allow lists. When a user says "always allow this tool for this project," the grant is persisted to the project-settings source, leaving user-level and session-level rules untouched. This composability allows teams to ship restrictive project-level defaults that individual developers can relax in their user settings without conflicting.
|
|
@@ -0,0 +1,104 @@
|
|
|
1
|
+
# Skill Runtime and Packaging Pattern
|
|
2
|
+
|
|
3
|
+
## Problem
|
|
4
|
+
|
|
5
|
+
Without a skill system, every reusable workflow must be re-explained in each conversation. The agent cannot self-select behaviors from a library of instructions, and there is no structured way to package permissions, model overrides, or lifecycle hooks alongside a prompt. The agent becomes a blank slate on each invocation, requiring the user to manually reconstruct context for repeated tasks.
|
|
6
|
+
|
|
7
|
+
A skill system introduces its own scaling problem: if the runtime loads every skill's full body into the system prompt, the context window fills before the first user message arrives. The runtime needs a discovery layer that is cheap enough to run at startup and a loading layer that defers the expensive content until the moment the skill is actually invoked.
|
|
8
|
+
|
|
9
|
+
These problems emerge in any agent runtime that supports reusable instruction sets — they are not specific to a single implementation.
|
|
10
|
+
|
|
11
|
+
## Golden Rules
|
|
12
|
+
|
|
13
|
+
### Metadata must have a single source of truth
|
|
14
|
+
|
|
15
|
+
All skill metadata — name, description, trigger hints, tool permissions, execution mode, hooks — should be co-located with the skill's instruction content rather than split into a separate sidecar file. A separate metadata file (JSON manifest, YAML config, etc.) creates a synchronization problem: the two sources can drift, and the runtime must decide which wins. A single source of truth eliminates this class of bug entirely. Common implementations use frontmatter blocks, header annotations, or inline structured comments — the specific mechanism matters less than the guarantee that there is exactly one authoritative location for each metadata field.
|
|
16
|
+
|
|
17
|
+
### Discovery is budget-constrained; loading is lazy
|
|
18
|
+
|
|
19
|
+
At startup, the runtime lists all available skills in the system prompt so the model knows what it can invoke. This listing includes only metadata — a short description and trigger hints per skill — never the full body. Each entry is hard-capped at a fixed character limit, and the total listing is capped at a small fraction of the context window (on the order of one percent). The full skill body is loaded only when the model or user actually invokes the skill. This two-phase approach — cheap discovery, deferred loading — keeps the first-turn prompt small and cacheable while preserving the model's ability to self-select the right skill.
|
|
20
|
+
|
|
21
|
+
### Trigger language is a design surface, not an afterthought
|
|
22
|
+
|
|
23
|
+
A skill has two distinct text fields for discovery: a short label (the description) and a detailed trigger guide (the "when to use" field). The description is what appears in listings. The trigger guide carries example phrases, situational cues, and invocation conditions that help the model decide when to activate the skill automatically. Both fields are concatenated for the discovery listing, but separating them lets authors write a concise label and a verbose trigger guide without conflating the two. The quality of trigger language directly determines whether the model selects the right skill without user intervention.
|
|
24
|
+
|
|
25
|
+
### Bundled skills are privileged but not exempt from all limits
|
|
26
|
+
|
|
27
|
+
Skills compiled into the agent binary occupy a privileged trust tier. When budget pressure forces the listing to shrink, bundled entries retain their full descriptions while external entries are progressively trimmed. However, all entries — including bundled — face a per-entry character cap on the concatenated description. Bundled skills are also registered programmatically rather than discovered from the filesystem, and their assets are extracted securely at runtime. This privilege reflects a trust boundary: bundled skills are authored by the runtime maintainers and can be granted capabilities that external skills cannot.
|
|
28
|
+
|
|
29
|
+
### Inline is the default; isolation is opt-in
|
|
30
|
+
|
|
31
|
+
When a skill executes inline, its content is injected into the current conversation as a user message and the model processes it within the existing context. The user can steer the skill mid-execution, and the skill's work is visible in the conversation history. Isolated (forked) execution spawns a separate sub-agent with its own token budget and its own conversation context. The parent sees only the final result. Isolation is opt-in — the skill author must explicitly declare it — because inline execution is almost always preferable: it is simpler, more transparent, and allows human-in-the-loop steering.
|
|
32
|
+
|
|
33
|
+
### Deduplication follows the filesystem, not the name
|
|
34
|
+
|
|
35
|
+
When skills are loaded from multiple overlapping directories (user-wide, project-local, plugin), the same physical file can appear under different search paths. The runtime resolves each skill file to its canonical filesystem path and deduplicates on that identity. This prevents the same skill from appearing twice in the listing and avoids double-registration bugs without requiring unique naming conventions across independent skill sources.
|
|
36
|
+
|
|
37
|
+
## When To Use
|
|
38
|
+
|
|
39
|
+
- Your agent needs a library of reusable instruction sets that persist across conversations.
|
|
40
|
+
- You need to list available capabilities in the system prompt without exhausting the context window.
|
|
41
|
+
- Skills come from multiple trust levels — built-in, user-installed, project-local, and remote/plugin sources.
|
|
42
|
+
- Some skills must run in isolation to avoid contaminating the parent conversation's context.
|
|
43
|
+
- You want skill authors to package permissions, model preferences, and hooks alongside the instructions themselves.
|
|
44
|
+
- The skill catalog is large enough that naive full-body inclusion would crowd out the user's actual work.
|
|
45
|
+
|
|
46
|
+
## Tradeoffs
|
|
47
|
+
|
|
48
|
+
| Decision | Benefit | Cost |
|
|
49
|
+
|---|---|---|
|
|
50
|
+
| Co-located metadata as sole source | Single source of truth, no sync bugs | All metadata must fit the chosen co-location format |
|
|
51
|
+
| Budget-capped discovery listing | First-turn prompt stays small and cacheable | Long descriptions are silently truncated; model may miss poorly-described skills |
|
|
52
|
+
| Lazy body loading on invocation | Context window is not wasted on unused skills | One extra round-trip to load content at invocation time |
|
|
53
|
+
| Inline execution by default | User can steer mid-process; work is visible | Skill output consumes parent context window tokens |
|
|
54
|
+
| Opt-in isolated (forked) execution | Clean context boundary; parent stays uncluttered | Forked skill cannot read or write parent conversation state |
|
|
55
|
+
| Bundled skills exempt from budget degradation | Core capabilities retain descriptions under pressure | Non-bundled skills get less budget headroom as bundled catalog grows |
|
|
56
|
+
| Realpath-based deduplication | Same file is never listed twice regardless of symlinks | Depends on filesystem returning consistent canonical paths |
|
|
57
|
+
| Secure exclusive-create for bundled asset extraction | Prevents symlink-following attacks on temp directories | Slightly more complex file creation; requires exclusive-open semantics |
|
|
58
|
+
|
|
59
|
+
## Implementation Patterns
|
|
60
|
+
|
|
61
|
+
- Discover skills from four source categories at startup: bundled (compiled-in), user-installed (user home directory), project-local (workspace directory), and plugin/remote (MCP or equivalent). Load them in this order so that precedence is predictable.
|
|
62
|
+
- Parse only the co-located metadata block of each skill file (e.g., YAML frontmatter). Extract at minimum: name, description, trigger hints, tool permissions, execution mode, and any hook definitions. Ignore sidecar files entirely.
|
|
63
|
+
- Register each discovered skill as a command object with a deferred body loader — a closure or callback that reads the full Markdown content only when the skill is invoked. Do not read the body at discovery time.
|
|
64
|
+
- Resolve each skill file to its canonical filesystem path before registration. If two search paths yield the same canonical path, register the skill only once.
|
|
65
|
+
- Build the discovery listing by concatenating each skill's description and trigger hints, then hard-truncating each entry at a fixed character limit. Sum all entries and cap the total at a small percentage of the context window.
|
|
66
|
+
- When budget pressure forces truncation, apply a graceful degradation sequence: first drop trigger hints from non-bundled skills, then drop descriptions entirely, then fall back to names only. Bundled skills are exempt from degradation but still face the per-entry character cap.
|
|
67
|
+
- On invocation, check the skill's declared execution mode. If inline (the default), inject the skill's full Markdown body as a user message in the current conversation. If isolated, spawn a sub-agent with its own token budget, run the skill body there, and return only the result text to the parent.
|
|
68
|
+
- For bundled skills that include reference files, extract those files to a per-process temporary directory at first invocation. Use exclusive-create file semantics (fail if the file already exists) and refuse to follow symlinks during extraction to prevent symlink-based attacks.
|
|
69
|
+
- Block inline shell execution for skills sourced from untrusted remote origins. Only bundled and local disk skills should be able to embed executable commands.
|
|
70
|
+
- Validate that skills declaring restricted invocation modes (such as "model cannot invoke this skill") are enforced at the tool validation layer, not at the listing layer. The skill should still appear in listings if appropriate; the restriction applies only to programmatic invocation by the model.
|
|
71
|
+
|
|
72
|
+
## Gotchas
|
|
73
|
+
|
|
74
|
+
**There is no sidecar metadata file.** All metadata must live co-located with the skill's instruction content (e.g., in a frontmatter block). If a runtime consumer places a separate JSON or YAML metadata file next to the skill document, it should be silently ignored. This is the most common integration mistake.
|
|
75
|
+
|
|
76
|
+
**Description truncation is silent.** If the combined description and trigger-hint text exceeds the per-entry character cap, the listing entry is hard-truncated with an ellipsis. The model sees the truncated version and may fail to trigger the skill. Keep both fields concise, or accept reduced auto-invocation reliability.
|
|
77
|
+
|
|
78
|
+
**Forked execution discards parent conversation state.** A skill running in isolated mode cannot read or modify the parent's conversation messages. It receives a fresh context. Use inline execution for any skill that needs to observe or update the ongoing conversation.
|
|
79
|
+
|
|
80
|
+
**Budget overflow degrades non-bundled skills first.** If non-bundled skills collectively exceed the remaining character budget after bundled entries are reserved, non-bundled entries progressively lose their descriptions, then fall to names-only. The model can still invoke them by name but cannot auto-select them based on purpose. Authors who depend on auto-invocation must keep their metadata compact.
|
|
81
|
+
|
|
82
|
+
**Realpath deduplication depends on filesystem fidelity.** On virtual, container, or network filesystems that return degenerate inode values, realpath-based deduplication may behave unexpectedly. On filesystems with inode precision loss, false deduplication can occur. The runtime should use canonical path resolution consistently rather than mixing inode and path strategies.
|
|
83
|
+
|
|
84
|
+
**Restricted invocation modes are not the same as hidden skills.** A flag that prevents the model from invoking a skill programmatically is distinct from a flag that hides the skill from user-facing listings. A skill can be visible to users but blocked from model invocation, or vice versa. Confusing the two creates either a security gap or a usability gap.
|
|
85
|
+
|
|
86
|
+
**Invalid hook schemas are silently dropped.** If a skill declares hooks in its frontmatter but the schema does not match what the runtime expects, the hooks are ignored without an error. The skill loads and runs, but the hooks never fire. Validate hook schemas during development, not in production.
|
|
87
|
+
|
|
88
|
+
## Claude Code Evidence
|
|
89
|
+
|
|
90
|
+
Claude Code's skill system is a production implementation of these principles, supporting four distinct skill sources and managing discovery under tight budget constraints.
|
|
91
|
+
|
|
92
|
+
**Four-source discovery architecture.** Skills are loaded from bundled (compiled directly into the binary), user-installed (under the user's home configuration directory), project-local (under the workspace configuration directory), and MCP-sourced (remote plugin) origins. Each source is scanned independently at startup and merged into a single command registry. Symlink-aware deduplication ensures that overlapping directories — common when a user's home config and project config share symlinked skill directories — never produce duplicate listings.
|
|
93
|
+
|
|
94
|
+
**YAML frontmatter as the metadata contract.** Claude Code uses YAML frontmatter at the top of each skill's Markdown file as the single source of truth for all metadata. There is no sidecar metadata file — if one exists alongside a skill document, it is silently ignored. This eliminates the synchronization bugs that arise when metadata lives in two places.
|
|
95
|
+
|
|
96
|
+
**Budget-constrained listing with graceful degradation.** The listing budget is set at one percent of the context window, with each individual entry capped at 250 characters — a cap that applies to all entries, including bundled. Each entry concatenates the skill's short description, a separator, and its trigger-hint text. When the non-bundled catalog exceeds the remaining budget after bundled entries are reserved, the runtime progressively strips descriptions from non-bundled skills, eventually reducing them to bare names. Bundled skills are not subject to this degradation — they retain their full (but character-capped) descriptions regardless of budget pressure.
|
|
97
|
+
|
|
98
|
+
**Lazy body loading via deferred closures.** Each skill registers a deferred body loader at discovery time rather than reading its full Markdown content upfront. The body is materialized only when the model or user invokes the skill. This keeps startup fast and the first-turn system prompt compact — critical for turn-one cache efficiency in production workloads with large skill catalogs.
|
|
99
|
+
|
|
100
|
+
**Inline versus forked execution.** The default execution path injects the skill body directly into the parent conversation. When a skill declares isolated context, the runtime spawns a forked sub-agent with its own token budget, runs the skill body there, and returns only the result text. The fork path is explicitly opt-in because inline execution is almost always preferable — it allows the user to steer mid-process and keeps the skill's intermediate reasoning visible.
|
|
101
|
+
|
|
102
|
+
**Secure extraction of bundled skill assets.** Bundled skills that ship with reference files extract those files to a per-process temporary directory using exclusive-create semantics and symlink-following prevention. The directory name includes a cryptographic nonce to resist path prediction attacks. This design reflects the principle that even trusted, compiled-in assets should be extracted defensively — the temporary filesystem is a shared attack surface.
|
|
103
|
+
|
|
104
|
+
**Trigger language as a first-class design surface.** The separation of description and trigger-hint fields is deliberate. Skill authors are guided to write a short, human-readable description and a separate, detailed trigger guide starting with phrases like "Use when..." followed by example user utterances. The concatenated form is what the model sees in the discovery listing, but the structural separation encourages authors to optimize each field for its purpose rather than cramming everything into a single line.
|
package/skills/global_config/agentic-harness-patterns/references/task-decomposition-pattern.md
ADDED
|
@@ -0,0 +1,92 @@
|
|
|
1
|
+
# Long-running Work Management
|
|
2
|
+
|
|
3
|
+
## Problem
|
|
4
|
+
|
|
5
|
+
Without a structured work management layer, concurrent agent work collides in shared state: two sub-agents writing to the same in-memory buffer corrupt each other's output, there is no canonical signal for "this work is done and will not change," and the parent coordinator has no safe window to evict completed work from memory. Untyped work identifiers make it impossible to route kill signals or status queries to the right handler at scale.
|
|
6
|
+
|
|
7
|
+
These problems emerge in any agent runtime that supports concurrent or background work — they are not specific to a single implementation.
|
|
8
|
+
|
|
9
|
+
## Golden Rules
|
|
10
|
+
|
|
11
|
+
### Every work unit gets a typed identity
|
|
12
|
+
|
|
13
|
+
Assign each unit of async work a prefixed, typed ID at creation time. The type prefix encodes what kind of work it is (shell command, agent, remote worker, teammate, etc.), making log filtering, routing, and kill dispatch unambiguous without parsing additional fields. Use a collision-resistant random suffix — the ID space should be large enough that brute-force collisions are infeasible even across hostile environments.
|
|
14
|
+
|
|
15
|
+
### Strict state machine with permanent terminal states
|
|
16
|
+
|
|
17
|
+
Every work unit follows the same lifecycle: it starts (most register directly as "running," skipping any "pending" phase), it runs, and it ends in exactly one terminal state — completed, failed, or killed. Terminal states are permanent. Encode this invariant in a single canonical check function, and use that function everywhere instead of inlining status comparisons. If a new terminal state is ever added, inline copies will silently diverge.
|
|
18
|
+
|
|
19
|
+
### Output goes to disk; memory holds only an offset
|
|
20
|
+
|
|
21
|
+
Keeping full transcripts in memory for every concurrent sub-agent is unbounded. Instead, write output to a per-work-unit file on disk. The in-memory state only stores a read offset. On each poll cycle, read the delta since the last offset and advance atomically. This keeps the in-memory footprint constant regardless of how long a work unit runs.
|
|
22
|
+
|
|
23
|
+
### Eviction is two-phase, gated by notification
|
|
24
|
+
|
|
25
|
+
When work reaches a terminal state:
|
|
26
|
+
|
|
27
|
+
1. **Disk cleanup** happens eagerly — output files are removed at the terminal transition.
|
|
28
|
+
2. **Memory cleanup** happens lazily — the in-memory record is only removed after the parent has been notified of the result.
|
|
29
|
+
|
|
30
|
+
The notification gate is critical. Without it, the framework would delete the work record before the parent could read the result, creating a race between eviction and result retrieval.
|
|
31
|
+
|
|
32
|
+
## When To Use
|
|
33
|
+
|
|
34
|
+
- Your agent spawns concurrent sub-agents or background tasks.
|
|
35
|
+
- You need to track the lifecycle of work that outlives a single turn.
|
|
36
|
+
- You need to display task status in a UI or route kill signals to specific work units.
|
|
37
|
+
- Long-running agents produce output that would be too large to keep in memory.
|
|
38
|
+
- You need a clean GC strategy that doesn't race with result consumption.
|
|
39
|
+
|
|
40
|
+
## Tradeoffs
|
|
41
|
+
|
|
42
|
+
| Decision | Benefit | Cost |
|
|
43
|
+
|---|---|---|
|
|
44
|
+
| Typed prefixed IDs | Unambiguous routing, easy log grep | One more field to generate and propagate |
|
|
45
|
+
| Skip "pending" in practice | Simpler, faster registration | UIs that assume "pending" as a required phase will break |
|
|
46
|
+
| Disk-backed output | Constant memory, survives interruption | I/O latency per poll, disk cleanup obligations |
|
|
47
|
+
| Two-phase eviction | No race between GC and result retrieval | More complex lifecycle — both phases must happen |
|
|
48
|
+
| Notification-gated GC | Parent always sees the result | Un-notified terminal tasks leak memory indefinitely |
|
|
49
|
+
| Retain flag for UI | Users can view completed work | Must be explicitly cleared or the task leaks |
|
|
50
|
+
|
|
51
|
+
## Implementation Patterns
|
|
52
|
+
|
|
53
|
+
- Mint a typed, prefixed ID before allocating any state. The ID is the work unit's identity for its entire lifecycle.
|
|
54
|
+
- Initialize shared base fields (status, output file path, read offset) via a factory function. The factory sets a safe default status; the concrete constructor overrides to "running" if the work starts immediately.
|
|
55
|
+
- Register work through a single entry point into shared state. Never write directly to the task store — the registration function is the chokepoint for validation and deduplication.
|
|
56
|
+
- For agent-type work, initialize the output file as a symlink to the agent's existing transcript. This avoids copying and lets the output file resolve immediately.
|
|
57
|
+
- Transition to "running" at or before registration. The "pending" state from the base factory is a safe default, not a required phase — most work starts immediately.
|
|
58
|
+
- Use a generic, type-parameterized update function for all mutations. Skip the state spread when the updater returns the same reference to prevent spurious re-renders.
|
|
59
|
+
- On every terminal transition: set the end timestamp, clean up the disk output, and set an eviction deadline (unless the work is being actively viewed).
|
|
60
|
+
- Enqueue a parent notification exactly once per terminal transition. Use a "notified" flag inside the update function to make the enqueue idempotent.
|
|
61
|
+
- Register a cleanup handler at process level that kills all running work on shutdown.
|
|
62
|
+
- Guard offset patches against stale state: after an async disk read, re-check the work unit's status before applying the new offset. The unit may have completed during the read.
|
|
63
|
+
|
|
64
|
+
## Gotchas
|
|
65
|
+
|
|
66
|
+
**Do not evict before the parent is notified.** The eviction function silently no-ops when the notification flag is false. If you call it early, the work unit stays in memory indefinitely and the parent never learns the result.
|
|
67
|
+
|
|
68
|
+
**Retained work units are never auto-evicted.** When a work unit is being actively viewed in a UI, its eviction deadline is set to infinity. The UI must explicitly clear the retain flag when the user navigates away — otherwise the terminal work unit leaks forever.
|
|
69
|
+
|
|
70
|
+
**Update functions must not mutate the existing state object.** Return a new object or the original reference. In-place mutations are invisible to the immutable-update pattern and will not propagate to subscribers.
|
|
71
|
+
|
|
72
|
+
**Do not hold a full state snapshot across an async boundary.** If you store the full work-unit state before an async disk read, a concurrent terminal transition during the read will be clobbered when you spread the stale snapshot back. Store only the fields you need (e.g., the new offset).
|
|
73
|
+
|
|
74
|
+
**Use the canonical terminal-status check, not inline comparisons.** If a new terminal status is ever added, inline `status === 'completed' || ...` copies will silently diverge from the authoritative check.
|
|
75
|
+
|
|
76
|
+
**Eviction is two-phase — skipping either phase leaks resources.** Skipping disk cleanup leaks files. Skipping memory cleanup leaks state-store entries. Both must happen for clean GC.
|
|
77
|
+
|
|
78
|
+
**Kill is a no-op on non-running work.** The kill guard only acts on "running" work units. Double-killing is safe, but you cannot kill a "pending" unit — wait for it to become "running" first.
|
|
79
|
+
|
|
80
|
+
## Claude Code Evidence
|
|
81
|
+
|
|
82
|
+
Claude Code's task system is a production implementation of these principles, managing seven distinct work types concurrently:
|
|
83
|
+
|
|
84
|
+
**Typed prefixed IDs.** Each task type has a single-character prefix (agent, bash, remote, teammate, workflow, monitor, dream). The remaining characters are drawn from a case-insensitive-safe alphabet (digits + lowercase), producing ~2.8 trillion combinations per prefix. The design comment notes this is intentionally large enough to resist brute-force symlink attacks on the output file paths.
|
|
85
|
+
|
|
86
|
+
**Practical skip of "pending."** The base factory initializes status as "pending," but in practice, agent tasks, remote tasks, and dream tasks all override to "running" at registration time. The "pending" state exists as a safe default for the type system, not as a phase that real tasks pass through.
|
|
87
|
+
|
|
88
|
+
**Disk-backed output with offset-based polling.** Agent task output is written to a per-task file on disk. For local agent tasks, the output file is initialized as a symlink to the agent's existing JSONL transcript — no data is copied. The framework polls this file during "running" state, reading only the delta since the last offset. The offset is advanced atomically after the async read, with an explicit comment about avoiding clobbering a status transition that may have occurred during the `await`.
|
|
89
|
+
|
|
90
|
+
**Two-phase eviction with notification gate.** Terminal tasks set an eviction deadline (30 seconds by default) to give the UI time to display the completion state. Tasks being actively viewed set a "retain" flag that blocks eviction entirely. The eviction function checks the "notified" flag and no-ops if the parent hasn't been signaled yet. Memory records are cleaned lazily on a separate sweep cycle.
|
|
91
|
+
|
|
92
|
+
**Session-level task IDs.** The core task ID prefix map has seven entries. A separate session-management task type generates its own prefix outside the core map, using the same underlying local-agent-task machinery but with a different identity namespace. This separation keeps the generic task system clean while allowing specialized reuse.
|