ahead-pi 0.2.1 → 0.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +26 -41
- package/dist/ahead_wasm.wasm +0 -0
- package/generated/corrective-debugging/ai-audit.md +39 -0
- package/generated/corrective-debugging/ai-review.md +46 -0
- package/generated/corrective-debugging/characterize.md +53 -0
- package/generated/corrective-debugging/conclude.md +59 -0
- package/generated/corrective-debugging/correction.md +53 -0
- package/generated/corrective-debugging/deploy.md +38 -0
- package/generated/corrective-debugging/human-review.md +45 -0
- package/generated/corrective-debugging/implement.md +42 -0
- package/generated/corrective-debugging/investigate.md +59 -0
- package/generated/corrective-debugging/manifest.json +30 -0
- package/generated/corrective-debugging/model.md +53 -0
- package/generated/corrective-debugging/outcome.md +38 -0
- package/generated/corrective-debugging/plan.md +53 -0
- package/generated/corrective-debugging/verify.md +47 -0
- package/generated/decision/compare.md +45 -0
- package/generated/decision/criteria.md +45 -0
- package/generated/decision/decide.md +45 -0
- package/generated/decision/frame.md +45 -0
- package/generated/decision/manifest.json +21 -0
- package/generated/decision/options.md +47 -0
- package/generated/decision/publish.md +38 -0
- package/generated/decision/research.md +45 -0
- package/generated/internal-improvement/ai-audit.md +39 -0
- package/generated/internal-improvement/ai-review.md +46 -0
- package/generated/internal-improvement/baseline.md +46 -0
- package/generated/internal-improvement/decision.md +45 -0
- package/generated/internal-improvement/deploy.md +38 -0
- package/generated/internal-improvement/human-review.md +45 -0
- package/generated/internal-improvement/implement.md +42 -0
- package/generated/internal-improvement/invariants.md +38 -0
- package/generated/internal-improvement/manifest.json +29 -0
- package/generated/internal-improvement/options.md +47 -0
- package/generated/internal-improvement/outcome.md +38 -0
- package/generated/internal-improvement/plan.md +53 -0
- package/generated/internal-improvement/target.md +45 -0
- package/generated/internal-improvement/verify.md +45 -0
- package/generated/investigation/bound.md +45 -0
- package/generated/investigation/conclude.md +45 -0
- package/generated/investigation/explore.md +60 -0
- package/generated/investigation/frame.md +45 -0
- package/generated/investigation/gather.md +45 -0
- package/generated/investigation/manifest.json +21 -0
- package/generated/investigation/synthesize.md +51 -0
- package/generated/operational-stabilization/assess.md +46 -0
- package/generated/operational-stabilization/execute-observe.md +45 -0
- package/generated/operational-stabilization/manifest.json +19 -0
- package/generated/operational-stabilization/monitor.md +45 -0
- package/generated/operational-stabilization/outcome.md +38 -0
- package/generated/operational-stabilization/respond.md +40 -0
- package/generated/operational-stabilization/verify-recovery.md +45 -0
- package/generated/product-change/ai-audit.md +7 -4
- package/generated/product-change/ai-review.md +15 -5
- package/generated/product-change/decision.md +11 -2
- package/generated/product-change/define.md +4 -2
- package/generated/product-change/deploy.md +4 -2
- package/generated/product-change/human-review.md +11 -2
- package/generated/product-change/implement.md +4 -2
- package/generated/product-change/manifest.json +8 -3
- package/generated/product-change/options.md +11 -2
- package/generated/product-change/outcome.md +4 -2
- package/generated/product-change/plan.md +17 -2
- package/generated/product-change/questions.md +17 -2
- package/generated/product-change/research.md +11 -2
- package/generated/product-change/verify.md +4 -2
- package/generated/recommended-skills.json +24 -0
- package/generated/reference/CONSTITUTION.md +2 -0
- package/generated/reference/docs/evidence/README.md +17 -0
- package/generated/reference/docs/evidence/evidence-standard.md +2 -0
- package/generated/reference/docs/evidence/research-map.md +2 -0
- package/generated/reference/docs/{references → evidence/sources}/pragmatic-programmer-page-index.md +3 -1
- package/generated/reference/docs/{references → evidence/sources}/submitted-engineering-notes.md +3 -1
- package/generated/reference/docs/guide/README.md +28 -0
- package/generated/reference/docs/{acceptable-ai-use.md → guide/acceptable-ai-use.md} +4 -2
- package/generated/reference/docs/{engineering-practice.md → guide/engineering-practice.md} +5 -3
- package/generated/reference/docs/{rationale.md → guide/rationale.md} +3 -1
- package/generated/reference/docs/guide/recommended-skills.md +21 -0
- package/generated/reference/docs/{workflows → guide/workflows}/README.md +6 -4
- package/generated/reference/docs/{workflows → guide/workflows}/corrective-debugging.md +39 -19
- package/generated/reference/docs/{workflows → guide/workflows}/decision.md +4 -2
- package/generated/reference/docs/{workflows → guide/workflows}/internal-improvement.md +37 -23
- package/generated/reference/docs/{workflows → guide/workflows}/investigation.md +5 -1
- package/generated/reference/docs/{workflows → guide/workflows}/operational-stabilization.md +16 -12
- package/generated/reference/docs/{workflows → guide/workflows}/product-change.md +16 -3
- package/generated/reference/index.json +200 -87
- package/package.json +34 -25
- package/src/engine.ts +27 -8
- package/src/flow-guides.ts +168 -0
- package/src/guidance.ts +220 -78
- package/src/index.ts +852 -189
- package/src/reference-viewer.ts +20 -18
- package/src/reference.ts +76 -14
- package/src/review.ts +360 -0
- package/src/skills.ts +133 -0
- package/src/storage.ts +139 -15
- package/src/types.ts +1 -0
- package/generated/reference/docs/design/debugging-and-operations.md +0 -119
- package/generated/reference/docs/design/executable-workflows.md +0 -110
- package/generated/reference/docs/design/process-taxonomy.md +0 -144
- package/generated/reference/docs/releasing-pi.md +0 -89
package/src/storage.ts
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
import { execFileSync } from "node:child_process";
|
|
2
2
|
import { randomUUID } from "node:crypto";
|
|
3
|
-
import { mkdir, readFile, rename, writeFile } from "node:fs/promises";
|
|
3
|
+
import { mkdir, readFile, readdir, rename, rm, unlink, writeFile } from "node:fs/promises";
|
|
4
4
|
import { dirname, join, relative, resolve } from "node:path";
|
|
5
5
|
import type { Actor, Run } from "./types.js";
|
|
6
6
|
|
|
@@ -11,9 +11,11 @@ interface CurrentRunPointer {
|
|
|
11
11
|
|
|
12
12
|
export class RunStore {
|
|
13
13
|
readonly aheadDirectory: string;
|
|
14
|
+
readonly projectRoot: string;
|
|
14
15
|
|
|
15
|
-
constructor(
|
|
16
|
-
this.
|
|
16
|
+
constructor(rootPath: string) {
|
|
17
|
+
this.projectRoot = rootPath;
|
|
18
|
+
this.aheadDirectory = join(rootPath, ".ahead");
|
|
17
19
|
}
|
|
18
20
|
|
|
19
21
|
newRunId(): string {
|
|
@@ -23,19 +25,20 @@ export class RunStore {
|
|
|
23
25
|
|
|
24
26
|
async loadCurrent(): Promise<Run | undefined> {
|
|
25
27
|
try {
|
|
26
|
-
const pointer =
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
}
|
|
28
|
+
const pointer = parseCurrentRunPointer(
|
|
29
|
+
await readFile(join(this.aheadDirectory, "current.json"), "utf8"),
|
|
30
|
+
);
|
|
30
31
|
return this.load(pointer.run_id);
|
|
31
32
|
} catch (error) {
|
|
32
|
-
if (isMissing(error))
|
|
33
|
+
if (isMissing(error)) {
|
|
34
|
+
return undefined;
|
|
35
|
+
}
|
|
33
36
|
throw error;
|
|
34
37
|
}
|
|
35
38
|
}
|
|
36
39
|
|
|
37
40
|
async load(runId: string): Promise<Run> {
|
|
38
|
-
return
|
|
41
|
+
return parseRun(await readFile(this.runPath(runId), "utf8"));
|
|
39
42
|
}
|
|
40
43
|
|
|
41
44
|
async save(run: Run, makeCurrent = true): Promise<void> {
|
|
@@ -46,9 +49,48 @@ export class RunStore {
|
|
|
46
49
|
}
|
|
47
50
|
}
|
|
48
51
|
|
|
52
|
+
async listRunIds(): Promise<string[]> {
|
|
53
|
+
try {
|
|
54
|
+
const entries = await readdir(join(this.aheadDirectory, "runs"), { withFileTypes: true });
|
|
55
|
+
return entries
|
|
56
|
+
.filter((entry) => entry.isDirectory() && isSafeRunId(entry.name))
|
|
57
|
+
.map((entry) => entry.name)
|
|
58
|
+
.toSorted((left, right) => right.localeCompare(left));
|
|
59
|
+
} catch (error) {
|
|
60
|
+
if (isMissing(error)) {
|
|
61
|
+
return [];
|
|
62
|
+
}
|
|
63
|
+
throw error;
|
|
64
|
+
}
|
|
65
|
+
}
|
|
66
|
+
|
|
67
|
+
async saveCurrentForResume(runId: string): Promise<void> {
|
|
68
|
+
await this.load(runId);
|
|
69
|
+
await this.clearCurrent(runId);
|
|
70
|
+
}
|
|
71
|
+
|
|
72
|
+
async resume(runId: string): Promise<Run> {
|
|
73
|
+
const run = await this.load(runId);
|
|
74
|
+
const pointer: CurrentRunPointer = { api_version: "ahead.current/v0", run_id: runId };
|
|
75
|
+
await atomicJson(join(this.aheadDirectory, "current.json"), pointer);
|
|
76
|
+
return run;
|
|
77
|
+
}
|
|
78
|
+
|
|
79
|
+
async discardCurrent(runId: string): Promise<void> {
|
|
80
|
+
const directory = this.runDirectory(runId);
|
|
81
|
+
await this.clearCurrent(runId);
|
|
82
|
+
await rm(directory, { recursive: true, force: false });
|
|
83
|
+
}
|
|
84
|
+
|
|
49
85
|
artifactPath(run: Run, phase: string, kind: string): { absolute: string; relative: string } {
|
|
50
86
|
const sequence = String(run.events.length + 1).padStart(4, "0");
|
|
51
|
-
const absolute = join(
|
|
87
|
+
const absolute = join(
|
|
88
|
+
this.aheadDirectory,
|
|
89
|
+
"runs",
|
|
90
|
+
run.id,
|
|
91
|
+
"artifacts",
|
|
92
|
+
`${sequence}-${phase}-${kind}.md`,
|
|
93
|
+
);
|
|
52
94
|
return { absolute, relative: relative(this.projectRoot, absolute) };
|
|
53
95
|
}
|
|
54
96
|
|
|
@@ -62,15 +104,51 @@ export class RunStore {
|
|
|
62
104
|
await writeFile(resolved, `${content.trim()}\n`, { encoding: "utf8", flag: "wx" });
|
|
63
105
|
}
|
|
64
106
|
|
|
107
|
+
async readArtifact(path: string): Promise<string> {
|
|
108
|
+
const resolved = resolve(this.projectRoot, path);
|
|
109
|
+
const artifactsRoot = resolve(this.aheadDirectory, "runs");
|
|
110
|
+
if (!resolved.startsWith(`${artifactsRoot}/`)) {
|
|
111
|
+
throw new Error("artifact path escaped .ahead/runs");
|
|
112
|
+
}
|
|
113
|
+
return readFile(resolved, "utf8");
|
|
114
|
+
}
|
|
115
|
+
|
|
65
116
|
private runPath(runId: string): string {
|
|
66
|
-
|
|
67
|
-
|
|
117
|
+
return join(this.runDirectory(runId), "run.json");
|
|
118
|
+
}
|
|
119
|
+
|
|
120
|
+
private runDirectory(runId: string): string {
|
|
121
|
+
if (!isSafeRunId(runId)) {
|
|
122
|
+
throw new Error("unsafe AHEAD run id");
|
|
123
|
+
}
|
|
124
|
+
return join(this.aheadDirectory, "runs", runId);
|
|
125
|
+
}
|
|
126
|
+
|
|
127
|
+
private async clearCurrent(expectedRunId: string): Promise<void> {
|
|
128
|
+
const path = join(this.aheadDirectory, "current.json");
|
|
129
|
+
let pointer: CurrentRunPointer;
|
|
130
|
+
try {
|
|
131
|
+
pointer = parseCurrentRunPointer(await readFile(path, "utf8"));
|
|
132
|
+
} catch (error) {
|
|
133
|
+
if (isMissing(error)) {
|
|
134
|
+
return;
|
|
135
|
+
}
|
|
136
|
+
throw error;
|
|
137
|
+
}
|
|
138
|
+
if (pointer.run_id !== expectedRunId) {
|
|
139
|
+
throw new Error(
|
|
140
|
+
`active AHEAD run changed from ${expectedRunId} to ${pointer.run_id}; stop or resume again`,
|
|
141
|
+
);
|
|
142
|
+
}
|
|
143
|
+
await unlink(path);
|
|
68
144
|
}
|
|
69
145
|
}
|
|
70
146
|
|
|
71
147
|
export function humanActor(cwd: string): Actor {
|
|
72
148
|
const explicit = process.env.AHEAD_HUMAN_IDENTITY?.trim();
|
|
73
|
-
if (explicit)
|
|
149
|
+
if (explicit) {
|
|
150
|
+
return { kind: "human", identity: explicit };
|
|
151
|
+
}
|
|
74
152
|
for (const key of ["user.email", "user.name"]) {
|
|
75
153
|
try {
|
|
76
154
|
const value = execFileSync("git", ["config", key], {
|
|
@@ -78,7 +156,9 @@ export function humanActor(cwd: string): Actor {
|
|
|
78
156
|
encoding: "utf8",
|
|
79
157
|
stdio: ["ignore", "pipe", "ignore"],
|
|
80
158
|
}).trim();
|
|
81
|
-
if (value)
|
|
159
|
+
if (value) {
|
|
160
|
+
return { kind: "human", identity: value };
|
|
161
|
+
}
|
|
82
162
|
} catch {
|
|
83
163
|
// Fall through to the next local identity source.
|
|
84
164
|
}
|
|
@@ -93,7 +173,9 @@ export function projectRoot(cwd: string): string {
|
|
|
93
173
|
encoding: "utf8",
|
|
94
174
|
stdio: ["ignore", "pipe", "ignore"],
|
|
95
175
|
}).trim();
|
|
96
|
-
if (root)
|
|
176
|
+
if (root) {
|
|
177
|
+
return root;
|
|
178
|
+
}
|
|
97
179
|
} catch {
|
|
98
180
|
// AHEAD can also persist beside work that is not yet in Git.
|
|
99
181
|
}
|
|
@@ -110,3 +192,45 @@ async function atomicJson(path: string, value: unknown): Promise<void> {
|
|
|
110
192
|
function isMissing(error: unknown): boolean {
|
|
111
193
|
return !!error && typeof error === "object" && "code" in error && error.code === "ENOENT";
|
|
112
194
|
}
|
|
195
|
+
|
|
196
|
+
function isSafeRunId(runId: string): boolean {
|
|
197
|
+
return runId !== "." && runId !== ".." && /^[A-Za-z0-9._-]+$/.test(runId);
|
|
198
|
+
}
|
|
199
|
+
|
|
200
|
+
function parseCurrentRunPointer(content: string): CurrentRunPointer {
|
|
201
|
+
const value: unknown = JSON.parse(content);
|
|
202
|
+
if (
|
|
203
|
+
!isRecord(value) ||
|
|
204
|
+
value.api_version !== "ahead.current/v0" ||
|
|
205
|
+
typeof value.run_id !== "string" ||
|
|
206
|
+
value.run_id.length === 0
|
|
207
|
+
) {
|
|
208
|
+
throw new Error("invalid .ahead/current.json");
|
|
209
|
+
}
|
|
210
|
+
return { api_version: value.api_version, run_id: value.run_id };
|
|
211
|
+
}
|
|
212
|
+
|
|
213
|
+
function parseRun(content: string): Run {
|
|
214
|
+
const value: unknown = JSON.parse(content);
|
|
215
|
+
if (!isRun(value)) {
|
|
216
|
+
throw new Error("invalid AHEAD run record");
|
|
217
|
+
}
|
|
218
|
+
return value;
|
|
219
|
+
}
|
|
220
|
+
|
|
221
|
+
function isRun(value: unknown): value is Run {
|
|
222
|
+
return (
|
|
223
|
+
isRecord(value) &&
|
|
224
|
+
typeof value.api_version === "string" &&
|
|
225
|
+
typeof value.id === "string" &&
|
|
226
|
+
typeof value.title === "string" &&
|
|
227
|
+
typeof value.workflow_id === "string" &&
|
|
228
|
+
typeof value.workflow_version === "string" &&
|
|
229
|
+
typeof value.owner === "string" &&
|
|
230
|
+
Array.isArray(value.events)
|
|
231
|
+
);
|
|
232
|
+
}
|
|
233
|
+
|
|
234
|
+
function isRecord(value: unknown): value is Record<string, unknown> {
|
|
235
|
+
return typeof value === "object" && value !== null;
|
|
236
|
+
}
|
package/src/types.ts
CHANGED
|
@@ -1,119 +0,0 @@
|
|
|
1
|
-
# Debugging and Operational Investigation
|
|
2
|
-
|
|
3
|
-
Status: design discussion, not an approved workflow specification
|
|
4
|
-
|
|
5
|
-
The minimal [corrective-debugging](../workflows/corrective-debugging.md) and [operational-stabilization](../workflows/operational-stabilization.md) profiles translate this discussion into pilotable flows. This document retains the reasoning and unresolved questions behind them.
|
|
6
|
-
|
|
7
|
-
## Human ownership
|
|
8
|
-
|
|
9
|
-
Debugging is human-owned. AI may help collect and organize evidence, suggest hypotheses and tests, identify contradictions, explain systems, and challenge conclusions. The human chooses what to investigate, performs or authorizes tests, interprets the evidence, selects interventions, and accepts the conclusion or remaining uncertainty.
|
|
10
|
-
|
|
11
|
-
There is no “AI investigation” phase. Investigation is a human engineering activity in which AI may participate.
|
|
12
|
-
|
|
13
|
-
## Shared reasoning loop
|
|
14
|
-
|
|
15
|
-
```text
|
|
16
|
-
Observation
|
|
17
|
-
→ Characterize
|
|
18
|
-
→ Build or update the mental model
|
|
19
|
-
→ Generate hypotheses
|
|
20
|
-
→ Human selects a discriminating test
|
|
21
|
-
→ Predict expected results
|
|
22
|
-
→ Run the test
|
|
23
|
-
→ Record evidence
|
|
24
|
-
→ Update or refute hypotheses
|
|
25
|
-
→ Repeat
|
|
26
|
-
```
|
|
27
|
-
|
|
28
|
-
The process distinguishes:
|
|
29
|
-
|
|
30
|
-
- **Fact:** directly observed and linked to evidence.
|
|
31
|
-
- **Inference:** an interpretation derived from facts.
|
|
32
|
-
- **Hypothesis:** a falsifiable proposed explanation.
|
|
33
|
-
- **Test:** an experiment or observation capable of changing confidence in a hypothesis.
|
|
34
|
-
- **Result:** what the test actually produced.
|
|
35
|
-
- **Conclusion:** a human-accepted explanation with confidence, limits, and remaining uncertainty.
|
|
36
|
-
|
|
37
|
-
Predictions should be recorded before a test when practical. This reduces hindsight interpretation of ambiguous results.
|
|
38
|
-
|
|
39
|
-
## Bug debugging
|
|
40
|
-
|
|
41
|
-
A bug is a defect where observed software behavior conflicts with intended behavior.
|
|
42
|
-
|
|
43
|
-
```text
|
|
44
|
-
Observe → Reproduce or establish
|
|
45
|
-
→ Characterize
|
|
46
|
-
→ Human-led investigation loop
|
|
47
|
-
→ Human accepts diagnosis or uncertainty
|
|
48
|
-
→ Choose fix → Plan → Implement → Validate locally
|
|
49
|
-
→ AI review → Independent human review
|
|
50
|
-
→ Human-authorized deploy or release when applicable
|
|
51
|
-
→ Verify the original failure and observe the outcome
|
|
52
|
-
→ Audit assumptions → Human outcome
|
|
53
|
-
```
|
|
54
|
-
|
|
55
|
-
Reproduction is valuable but not universally required. A failure may be intermittent, historical, production-only, environment-specific, or already mitigated.
|
|
56
|
-
|
|
57
|
-
## Operational investigation
|
|
58
|
-
|
|
59
|
-
An operational issue is undesirable system behavior that may not be a software defect. Examples include reconciliation storms, configuration drift, capacity exhaustion, cloud-provider behavior, dependency failures, identity or certificate failures, resource contention, bad rollout sequencing, and emergent controller interactions.
|
|
60
|
-
|
|
61
|
-
```text
|
|
62
|
-
Observed condition
|
|
63
|
-
→ Desired state versus actual state
|
|
64
|
-
→ Impact and scope
|
|
65
|
-
→ Timeline
|
|
66
|
-
→ System and control-loop model
|
|
67
|
-
→ Recent changes and external events
|
|
68
|
-
→ Evidence/hypothesis/test loop
|
|
69
|
-
→ Intervention decision
|
|
70
|
-
→ Verify convergence and user-visible behavior
|
|
71
|
-
→ Monitor recurrence
|
|
72
|
-
→ Corrective actions
|
|
73
|
-
```
|
|
74
|
-
|
|
75
|
-
The initial cause classification must remain tentative. Labeling an issue a code bug, configuration error, or provider failure before collecting evidence can bias the investigation.
|
|
76
|
-
|
|
77
|
-
## Incident mode
|
|
78
|
-
|
|
79
|
-
“Incident” describes impact and urgency rather than cause. An incident may be caused by a bug, configuration, capacity, provider behavior, a security event, data, or an interaction that remains partially unexplained.
|
|
80
|
-
|
|
81
|
-
Incident mode adds parallel concerns:
|
|
82
|
-
|
|
83
|
-
```text
|
|
84
|
-
Response: assess impact → contain → recover → monitor
|
|
85
|
-
Investigation: evidence → model → hypotheses → tests → conclusion
|
|
86
|
-
Coordination: ownership → decisions → communications → timeline
|
|
87
|
-
```
|
|
88
|
-
|
|
89
|
-
Mitigation and recovery must not be blocked on completing a full diagnosis. Risky interventions still require explicit human authorization. After recovery, unresolved causal analysis and prevention work may continue as a linked bug, operational investigation, security issue, or technical-debt item.
|
|
90
|
-
|
|
91
|
-
## Current design direction
|
|
92
|
-
|
|
93
|
-
- Treat bug debugging and operational investigation as separate work types.
|
|
94
|
-
- Treat incident mode as an overlay that may apply to several work types.
|
|
95
|
-
- Keep the evidence and hypothesis loop flexible rather than gating every iteration.
|
|
96
|
-
- Reserve hard gates for human accountability, risky tests, consequential interventions, accepted conclusions, implementation plans, reviews, and verified outcomes.
|
|
97
|
-
- Distinguish hypothesis testing, fix validation, and post-deployment outcome verification.
|
|
98
|
-
- Allow causal conclusions to include a failure mechanism, trigger, enabling conditions, and detection or containment gaps instead of insisting on one root cause.
|
|
99
|
-
|
|
100
|
-
## Open questions
|
|
101
|
-
|
|
102
|
-
- What minimum evidence is needed before a human may accept a diagnosis?
|
|
103
|
-
- When may a team remediate while explicitly accepting that the cause is unknown?
|
|
104
|
-
- Which tests require approval based on environment, reversibility, or blast radius?
|
|
105
|
-
- How should AHEAD represent multiple interacting causes and confidence changes?
|
|
106
|
-
- When does an operational anomaly become incident mode?
|
|
107
|
-
- Which incident records must be produced during response, and which may be reconstructed afterward?
|
|
108
|
-
- How should follow-up work remain linked without keeping the incident itself permanently open?
|
|
109
|
-
|
|
110
|
-
## Evidence basis
|
|
111
|
-
|
|
112
|
-
The current reasoning loop is supported by direct empirical software-engineering research, though the exact AHEAD recording requirements are not yet validated:
|
|
113
|
-
|
|
114
|
-
- [Li and Coblenz, *A Grounded Theory of Debugging in Professional Software Engineering Practice*](https://arxiv.org/abs/2602.11435) observed professional developers and describes debugging as iterative mental-model construction that guides information gathering.
|
|
115
|
-
- [Alaboudi and LaToza, *Using Hypotheses as a Debugging Aid*](https://doi.org/10.1109/VL/HCC50065.2020.9127273) found that early correct hypotheses predicted success and that supplying potential hypotheses helped more than supplying fault locations in their controlled experiment.
|
|
116
|
-
- [Sillito and Kutomi, *Failures and Fixes*](https://doi.org/10.1109/ICSME46990.2020.00027) analyzed 30 incidents and identified distinct investigative and mitigative strategies.
|
|
117
|
-
- [Ghosh et al., *How to Fight Production Incidents?*](https://doi.org/10.1145/3542929.3563482) studied hundreds of high-severity cloud incidents, including non-code causes, across detection, diagnosis, and mitigation.
|
|
118
|
-
|
|
119
|
-
These studies support the shape of the process. They do not prove that requiring engineers to record every fact, hypothesis, or test improves results. AHEAD must test the minimum useful structure and remove requirements that interrupt investigation without improving reasoning, handoff, or learning.
|
|
@@ -1,110 +0,0 @@
|
|
|
1
|
-
# Executable AHEAD Workflows
|
|
2
|
-
|
|
3
|
-
Status: initial vertical slice v0.1
|
|
4
|
-
|
|
5
|
-
## Purpose
|
|
6
|
-
|
|
7
|
-
The executable layer makes AHEAD workflow state durable and makes selected human/AI boundaries enforceable across integrations. It does not turn judgment into a checklist or make workflow artifacts proof of understanding.
|
|
8
|
-
|
|
9
|
-
The first vertical slice implements the Product Change workflow. The other five pilot workflows remain manual profiles until dogfooding provides evidence about the reusable state model.
|
|
10
|
-
|
|
11
|
-
## Architecture
|
|
12
|
-
|
|
13
|
-
```text
|
|
14
|
-
CONSTITUTION / ACCEPTABLE-AI-USE
|
|
15
|
-
│
|
|
16
|
-
▼
|
|
17
|
-
CANONICAL WORKFLOW SPEC + POLICY FRAGMENTS
|
|
18
|
-
│
|
|
19
|
-
┌────────┴────────┐
|
|
20
|
-
▼ ▼
|
|
21
|
-
RUST WORKFLOW CORE GENERATED INSTRUCTIONS
|
|
22
|
-
│ │
|
|
23
|
-
└────────┬────────┘
|
|
24
|
-
▼
|
|
25
|
-
INTEGRATION ADAPTER
|
|
26
|
-
(Pi first; others later)
|
|
27
|
-
│
|
|
28
|
-
┌────────┴─────────┐
|
|
29
|
-
▼ ▼
|
|
30
|
-
HOST TOOLS / MODEL DURABLE RUN FILES
|
|
31
|
-
```
|
|
32
|
-
|
|
33
|
-
The Rust core owns workflow semantics. It is deterministic and has no filesystem, network, clock, model-provider, or editor dependency. Hosts supply identity and timestamps, persist returned runs, and map their tools to canonical capabilities.
|
|
34
|
-
|
|
35
|
-
The initial WebAssembly boundary is a small versioned JSON ABI. This avoids coupling the core to one JavaScript binding generator and lets a later editor or CI integration load the same state machine.
|
|
36
|
-
|
|
37
|
-
## Sources of truth
|
|
38
|
-
|
|
39
|
-
| Concern | Canonical source |
|
|
40
|
-
|---|---|
|
|
41
|
-
| Durable principles | `CONSTITUTION.md` |
|
|
42
|
-
| AI authority | `docs/acceptable-ai-use.md` |
|
|
43
|
-
| Human-readable Product Change flow | `docs/workflows/product-change.md` |
|
|
44
|
-
| Executable phases, artifacts, gates, transitions, and capabilities | `spec/workflows/product-change-v0.1.json` |
|
|
45
|
-
| Compact binding agent profile and shared AI behavior | `policy/common.md` |
|
|
46
|
-
| Phase AI behavior | `policy/product-change/*.md` |
|
|
47
|
-
| State transition enforcement | `crates/ahead-core` |
|
|
48
|
-
| Host mapping, storage, and UI | `integrations/pi` |
|
|
49
|
-
|
|
50
|
-
Generated integration instructions are build artifacts. They include the compact agent profile, active phase policy, enforced contract, workflow version, and a source hash and must not be edited directly.
|
|
51
|
-
|
|
52
|
-
The Pi package also copies the canonical Constitution and `docs/**/*.md` into a generated reference catalog. These full documents are not injected into every prompt. The adapter recommends references applicable to the active phase, lets humans read them through `/ahead-guide`, and lets AI retrieve a specific source through `ahead_get_reference`. This keeps the binding prompt small while making the framework, rationale, evidence, and original page-level provenance available on demand.
|
|
53
|
-
|
|
54
|
-
## State and evidence
|
|
55
|
-
|
|
56
|
-
A run is an append-only event log. Events record an actor kind and identity, host-supplied timestamp, sequence, and one action:
|
|
57
|
-
|
|
58
|
-
- start the run;
|
|
59
|
-
- record an artifact;
|
|
60
|
-
- accept a human gate;
|
|
61
|
-
- advance or return between phases;
|
|
62
|
-
- close the run.
|
|
63
|
-
|
|
64
|
-
State is derived by replay. The core rejects invalid history rather than trusting a cached phase field. A return transition creates a new visit to the target phase. Earlier evidence remains in history, but only evidence recorded during the current visit satisfies its gate.
|
|
65
|
-
|
|
66
|
-
Pi stores state under the work's Git root:
|
|
67
|
-
|
|
68
|
-
```text
|
|
69
|
-
.ahead/
|
|
70
|
-
├── current.json
|
|
71
|
-
└── runs/
|
|
72
|
-
└── <run-id>/
|
|
73
|
-
├── run.json
|
|
74
|
-
└── artifacts/
|
|
75
|
-
└── <sequence>-<phase>-<kind>.md
|
|
76
|
-
```
|
|
77
|
-
|
|
78
|
-
These are intended to be inspectable, diffable repository artifacts. A team can decide which records belong in Git, while CI and GitHub enforcement are later adapters over the same run contract.
|
|
79
|
-
|
|
80
|
-
## Enforced boundaries in v0.1
|
|
81
|
-
|
|
82
|
-
- Only a human actor can start a run, accept a gate, transition a phase, return work, or close a run.
|
|
83
|
-
- Artifact definitions state whether a human, AI, or either may record them.
|
|
84
|
-
- Human-first option and plan artifacts unlock AI assistance in those phases.
|
|
85
|
-
- Required current-visit artifacts must exist before gate acceptance.
|
|
86
|
-
- Advancement requires the current human gate.
|
|
87
|
-
- Independent human review must be recorded by an identity other than the latest changeset implementer, and that reviewer must accept the review gate.
|
|
88
|
-
- The Product Change flow distinguishes implementation, deployment, observation, audit, and human outcome.
|
|
89
|
-
- Model-invoked host tools require an explicit adapter mapping to an allowed canonical capability.
|
|
90
|
-
|
|
91
|
-
Instructions explain these boundaries to the model. The Rust core enforces the transition, actor, artifact, identity, and capability decisions even if instructions are ignored.
|
|
92
|
-
|
|
93
|
-
During implementation, instructions explicitly permit questions, explanation, debugging help, and bounded suggestions while keeping the engineer first. If the engineer has not supplied a current model, attempted approach, or intended behavior, AI asks for it before proposing a solution. A request for help does not authorize AI to take over the implementation.
|
|
94
|
-
|
|
95
|
-
## Trust boundaries and limits
|
|
96
|
-
|
|
97
|
-
The Pi adapter is an engineering workflow control, not a security sandbox.
|
|
98
|
-
|
|
99
|
-
- Local human identity comes from `AHEAD_HUMAN_IDENTITY`, then Git email/name, then the local user. It is self-attested. GitHub review identity and protected-branch rules will provide a stronger boundary later.
|
|
100
|
-
- Pi's direct `!` shell is a human action and is not intercepted. Model-invoked `bash` is intercepted.
|
|
101
|
-
- Unknown model tools are denied until the adapter classifies them. This prevents a newly installed effectful tool from silently acquiring authority.
|
|
102
|
-
- Artifact and run writes are atomic, but v0.1 has no multi-process lock. One writer should operate a run at a time.
|
|
103
|
-
- Workflow files can prove that a named action was recorded, not that a person genuinely understood it. Human review and organizational accountability remain necessary.
|
|
104
|
-
- No GitHub checks, PR gates, migration engine, signature scheme, or backwards-compatible workflow upgrade exists yet.
|
|
105
|
-
|
|
106
|
-
## Reuse path
|
|
107
|
-
|
|
108
|
-
The reusable boundary is the engine API, not a CLI. Pi is the first adapter. A VS Code extension, GitHub check, or future WASM-capable editor can reuse the same compiled core and canonical fragments while providing its own UI, storage transport, identity strength, and tool-capability map.
|
|
109
|
-
|
|
110
|
-
Dogfooding should test whether phase visits, artifacts, gates, returns, capability vocabulary, and identity rules generalize before the remaining five workflow profiles are encoded.
|
|
@@ -1,144 +0,0 @@
|
|
|
1
|
-
# AHEAD Process Taxonomy
|
|
2
|
-
|
|
3
|
-
Status: proposed design
|
|
4
|
-
Last reviewed: 2026-08-12
|
|
5
|
-
|
|
6
|
-
## Why classify by outcome
|
|
7
|
-
|
|
8
|
-
AHEAD should not create a workflow for every issue label. “Security,” “performance,” “data,” “incident,” and “technical debt” often describe risk, domain, urgency, or cause—not the kind of reasoning needed to complete the work.
|
|
9
|
-
|
|
10
|
-
The primary workflow should be selected by the **dominant outcome** the human is trying to produce. Variants and overlays then adapt that workflow to context.
|
|
11
|
-
|
|
12
|
-
This gives AHEAD six proposed process families. All six now have minimal [pilot workflow profiles](../workflows/README.md) for use and evaluation; that does not yet validate the taxonomy or justify automated enforcement.
|
|
13
|
-
|
|
14
|
-
## The six process families
|
|
15
|
-
|
|
16
|
-
| Process family | Dominant question | Terminal outcome | Examples | Status |
|
|
17
|
-
|---|---|---|---|---|
|
|
18
|
-
| 1. [Product change](../workflows/product-change.md) | What behavior or capability should exist, and how should we deliver it? | Verified intended behavior | Feature, API change, integration, migration, dependency adaptation, decommission | Pilot v0.1 |
|
|
19
|
-
| 2. [Corrective debugging](../workflows/corrective-debugging.md) | Why does observed behavior differ from intended behavior, and how should we correct it? | Verified correction or explicitly accepted uncertainty | Deterministic bug, flaky failure, regression, incorrect data processing | Pilot v0.1 |
|
|
20
|
-
| 3. [Operational stabilization](../workflows/operational-stabilization.md) | Why is a live system outside an acceptable operating state, and how do we restore and stabilize it? | Demonstrated recovery/convergence and follow-up disposition | Reconciliation storm, capacity exhaustion, configuration drift, dependency outage | Pilot v0.1 |
|
|
21
|
-
| 4. [Decision](../workflows/decision.md) | Which course should humans choose, given goals, evidence, constraints, and tradeoffs? | Approved decision and rationale | Architecture decision, buy/build, technology selection, policy or platform choice | Pilot v0.1 |
|
|
22
|
-
| 5. [Investigation](../workflows/investigation.md) | What is true, feasible, or likely when no intervention has yet been selected? | Bounded conclusion, confidence, evidence, and remaining unknowns | Technical spike, feasibility study, causal follow-up, capacity study, vendor evaluation | Pilot v0.1 |
|
|
23
|
-
| 6. [Internal improvement](../workflows/internal-improvement.md) | How can we improve system qualities while preserving an explicit behavioral contract? | Verified invariants plus improved target qualities | Refactor, preventive maintenance, maintainability debt, performance optimization without semantic change | Pilot v0.1 |
|
|
24
|
-
|
|
25
|
-
Six is a working taxonomy, not a sacred number. The threshold for adding a seventh family is deliberately high.
|
|
26
|
-
|
|
27
|
-
## Selection test
|
|
28
|
-
|
|
29
|
-
```text
|
|
30
|
-
Is the primary outcome new or changed externally meaningful behavior?
|
|
31
|
-
→ Product change
|
|
32
|
-
|
|
33
|
-
Is an observed behavior wrong and the main work is causal diagnosis plus correction?
|
|
34
|
-
→ Corrective debugging
|
|
35
|
-
|
|
36
|
-
Is a live system unhealthy, unstable, or failing to converge, with restoration as the immediate outcome?
|
|
37
|
-
→ Operational stabilization
|
|
38
|
-
|
|
39
|
-
Is the deliverable an accountable choice among alternatives?
|
|
40
|
-
→ Decision
|
|
41
|
-
|
|
42
|
-
Is the deliverable knowledge or reduced uncertainty, without a predetermined change?
|
|
43
|
-
→ Investigation
|
|
44
|
-
|
|
45
|
-
Must behavior remain invariant while internal qualities improve?
|
|
46
|
-
→ Internal improvement
|
|
47
|
-
```
|
|
48
|
-
|
|
49
|
-
A large effort may link several runs. An architecture decision can lead to a product change. An incident can create an operational investigation, a corrective bug, and an internal-improvement follow-up. A technical spike can end in a decision without pretending that knowledge production and option selection are the same activity.
|
|
50
|
-
|
|
51
|
-
## Why the additional three differ
|
|
52
|
-
|
|
53
|
-
### Decision
|
|
54
|
-
|
|
55
|
-
A feature includes decisions, but some engineering work ends with a decision rather than code. Its quality depends on framing, option coverage, evidence, tradeoffs, consequences, reversibility, and accountable approval. Forcing it through implementation and deployment creates meaningless states.
|
|
56
|
-
|
|
57
|
-
### Investigation
|
|
58
|
-
|
|
59
|
-
An investigation begins with a question, not an assumed defect or desired change. It may conclude that no action is needed, evidence is insufficient, a vendor owns the behavior, or several interventions remain viable. Its terminal quality is epistemic: evidence, confidence, limitations, and unknowns.
|
|
60
|
-
|
|
61
|
-
### Internal improvement
|
|
62
|
-
|
|
63
|
-
Refactoring and preventive work are judged differently from feature work. They begin by specifying invariants and target qualities. Success means that required behavior was preserved while maintainability, performance, safety, comprehensibility, or another quality improved. Treating this as a feature encourages invented product outcomes; treating it as a bug assumes a failure that may not exist.
|
|
64
|
-
|
|
65
|
-
## Overlays, not primary processes
|
|
66
|
-
|
|
67
|
-
### Incident mode
|
|
68
|
-
|
|
69
|
-
Incident mode represents urgency, impact, coordination, containment, communication, and recovery. It can overlay corrective debugging, operational stabilization, a security event, or a data issue. It relaxes nonessential documentation during response but strengthens action authorization and decision logging.
|
|
70
|
-
|
|
71
|
-
### Security
|
|
72
|
-
|
|
73
|
-
Security adds confidentiality, evidence preservation, threat modeling, restricted AI access, disclosure, and security approval. A vulnerability may use corrective debugging; proactive hardening may use internal improvement; an active compromise may use incident-mode operational stabilization; a threat assessment may use investigation.
|
|
74
|
-
|
|
75
|
-
### Safety, regulatory, and compliance
|
|
76
|
-
|
|
77
|
-
These overlays strengthen traceability, independence, evidence retention, required reviewers, and non-waivable gates. They do not change whether the underlying work is a change, correction, operation, decision, investigation, or improvement.
|
|
78
|
-
|
|
79
|
-
### Emergency
|
|
80
|
-
|
|
81
|
-
Emergency handling changes sequencing and permits explicitly governed deferrals. It does not erase human accountability or evidence requirements; it moves some reconstruction and learning after stabilization.
|
|
82
|
-
|
|
83
|
-
## Labels and modifiers
|
|
84
|
-
|
|
85
|
-
Context belongs in typed modifiers rather than new workflow definitions:
|
|
86
|
-
|
|
87
|
-
```yaml
|
|
88
|
-
process: operational-stabilization
|
|
89
|
-
urgency: incident
|
|
90
|
-
domain: infrastructure
|
|
91
|
-
assurance: standard
|
|
92
|
-
failure_character: intermittent
|
|
93
|
-
environment: production
|
|
94
|
-
data_classification: internal
|
|
95
|
-
```
|
|
96
|
-
|
|
97
|
-
Useful modifiers may include:
|
|
98
|
-
|
|
99
|
-
- urgency: normal, expedited, incident, emergency;
|
|
100
|
-
- assurance: standard, security, safety-critical, regulated;
|
|
101
|
-
- environment: local, test, staging, production, external provider;
|
|
102
|
-
- failure character: deterministic, intermittent, performance, data, distributed, unknown;
|
|
103
|
-
- change character: additive, adaptive, migration, retirement;
|
|
104
|
-
- reversibility and blast radius;
|
|
105
|
-
- evidence sensitivity.
|
|
106
|
-
|
|
107
|
-
## Where common work maps
|
|
108
|
-
|
|
109
|
-
| Work label | Primary process or routing rule |
|
|
110
|
-
|---|---|
|
|
111
|
-
| Feature | Product change |
|
|
112
|
-
| Bug | Corrective debugging |
|
|
113
|
-
| Production reconciliation storm | Operational stabilization; add incident mode when impact/urgency warrants it |
|
|
114
|
-
| Architecture decision | Decision; link resulting implementation separately |
|
|
115
|
-
| Technical debt | Internal improvement when preserving behavior; product change when behavior changes; corrective debugging when it represents a known defect |
|
|
116
|
-
| Refactor | Internal improvement |
|
|
117
|
-
| Performance regression | Corrective debugging |
|
|
118
|
-
| Proactive performance optimization | Internal improvement or product change, depending on whether performance is a new product outcome |
|
|
119
|
-
| Security vulnerability | Corrective debugging plus security overlay |
|
|
120
|
-
| Active security compromise | Operational stabilization plus incident and security overlays |
|
|
121
|
-
| Security hardening | Internal improvement or product change plus security overlay |
|
|
122
|
-
| Research spike | Investigation |
|
|
123
|
-
| Compliance audit | Investigation plus compliance overlay; corrective or improvement runs handle findings |
|
|
124
|
-
| Dependency or platform upgrade | Product change with adaptive-change modifier |
|
|
125
|
-
| Service retirement | Product change with retirement and risk modifiers |
|
|
126
|
-
|
|
127
|
-
## Test for adding another family
|
|
128
|
-
|
|
129
|
-
A new primary process family should be added only when all of these are true:
|
|
130
|
-
|
|
131
|
-
1. It has a distinct terminal outcome.
|
|
132
|
-
2. It has a distinct central reasoning loop.
|
|
133
|
-
3. It requires materially different human decisions or gates.
|
|
134
|
-
4. It cannot be represented clearly as a variant, overlay, modifier, or linked combination of existing families.
|
|
135
|
-
5. Evidence or repeated practice shows that using an existing family creates confusion, unsafe behavior, or process theater.
|
|
136
|
-
6. The additional cognitive and tooling cost is justified.
|
|
137
|
-
|
|
138
|
-
## Evidence basis and limits
|
|
139
|
-
|
|
140
|
-
[ISO/IEC/IEEE 12207:2026](https://standards.ieee.org/ieee/12207/11416/) covers development, operation, maintenance, support, and retirement and allows processes to operate concurrently, iteratively, and recursively. [ISO/IEC/IEEE 14764:2022](https://www.iso.org/standard/80710.html) separately establishes software-maintenance types. These standards support broad coverage and composition, but they do not validate this six-family taxonomy.
|
|
141
|
-
|
|
142
|
-
Empirical debugging research supports a mental-model and hypothesis-testing process distinct from planned change. Empirical production-incident research distinguishes code and non-code causes and separates detection, investigation, and mitigation. Those findings support keeping corrective debugging and operational stabilization separate.
|
|
143
|
-
|
|
144
|
-
The proposed six-family classification itself remains an AHEAD design hypothesis. The pilot profiles should be tested against a diverse sample of real engineering work by asking whether teams can route work consistently, whether important states or gates differ, which records improve reasoning or handoff, and whether any family is rarely used or routinely misclassified.
|