@tea-agent/loop-agent 0.13.0 → 0.15.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +157 -157
- package/CHANGELOG.md +116 -305
- package/README.md +357 -334
- package/bin/agent-worker.js +22 -22
- package/bin/loop-agent.js +21 -21
- package/dist/commands/cursor-prompt.js +6 -6
- package/dist/commands/init.js +505 -505
- package/dist/commands/loop-benchmark.js +11 -11
- package/dist/commands/pi-reuse-benchmark.js +16 -16
- package/dist/executors/pi-event-serializer.js +33 -11
- package/dist/sidecars/cursor-prompt/executor.js +1 -1
- package/dist/task/runtime.js +27 -27
- package/dist/worker/observe/spec-evidence.js +19 -10
- package/dist/worker/observe/static/api.js +46 -46
- package/dist/worker/observe/static/app.js +151 -150
- package/dist/worker/observe/static/constants.js +156 -148
- package/dist/worker/observe/static/copy.js +67 -67
- package/dist/worker/observe/static/dag-helpers.js +201 -172
- package/dist/worker/observe/static/dag-layout.d.ts +31 -31
- package/dist/worker/observe/static/dag-layout.js +83 -83
- package/dist/worker/observe/static/dag-model.js +72 -72
- package/dist/worker/observe/static/dom.js +122 -122
- package/dist/worker/observe/static/format-pool.d.ts +71 -0
- package/dist/worker/observe/static/format-pool.js +134 -67
- package/dist/worker/observe/static/format.js +317 -292
- package/dist/worker/observe/static/index.html +350 -308
- package/dist/worker/observe/static/kpi.js +100 -94
- package/dist/worker/observe/static/markdown-render.js +124 -0
- package/dist/worker/observe/static/relations.js +133 -133
- package/dist/worker/observe/static/router.js +93 -93
- package/dist/worker/observe/static/run-processing.js +148 -148
- package/dist/worker/observe/static/shell-chrome.js +74 -68
- package/dist/worker/observe/static/state.js +273 -267
- package/dist/worker/observe/static/styles.css +2504 -1902
- package/dist/worker/observe/static/views/batch.js +227 -227
- package/dist/worker/observe/static/views/dag-graph.js +172 -172
- package/dist/worker/observe/static/views/dag-inspector.js +530 -627
- package/dist/worker/observe/static/views/dag.js +371 -371
- package/dist/worker/observe/static/views/dashboard.js +86 -100
- package/dist/worker/observe/static/views/failures.js +143 -143
- package/dist/worker/observe/static/views/feature.js +492 -492
- package/dist/worker/observe/static/views/pool.js +708 -350
- package/dist/worker/observe/static/views/run.js +453 -453
- package/dist/worker/observe/static/views/session-timeline.js +771 -219
- package/dist/worker/observe/static/views/shell.js +7 -7
- package/dist/worker/observe/static/views/task.js +314 -314
- package/dist/worker/observe/static/views/timeline.js +163 -163
- package/dist/workflows/dag/canvas-observer.js +275 -275
- package/dist/workflows/dag/init-hybrid.js +27 -11
- package/docs/README.md +106 -104
- package/docs/architecture/README.md +26 -26
- package/docs/architecture/dag-execution.md +140 -140
- package/docs/architecture/evolution.md +54 -54
- package/docs/architecture/facts-and-state.md +71 -71
- package/docs/architecture/runtime-boundaries.md +191 -191
- package/docs/architecture/system-overview.md +93 -93
- package/docs/architecture/worker-and-feature.md +85 -85
- package/docs/harness-methodology-debugging.md +153 -153
- package/docs/harness-methodology-tdd.md +130 -130
- package/docs/harness-methodology-verification.md +27 -27
- package/docs/init-surface.manifest.json +304 -307
- package/docs/skills/README.md +7 -7
- package/docs/skills/vetted-skill-registry.md +29 -29
- package/docs/templates/adr.md +60 -60
- package/docs/templates/agent-dag-authority-surface-audit.prompt.md +94 -94
- package/docs/templates/agent-dag-decision-envelope.schema.json +213 -213
- package/docs/templates/agent-dag-decision-gate-dogfood-report.md +117 -117
- package/docs/templates/agent-dag-decision-gate.prompt.md +246 -246
- package/docs/templates/agent-dag-process-supervisor.prompt.md +98 -98
- package/docs/templates/agent-dag-report.schema.json +473 -473
- package/docs/templates/agent-dag-review-verdict.prompt.md +68 -68
- package/docs/templates/agent-dag.base.json +190 -190
- package/docs/templates/agent-dag.final-verification.json +185 -185
- package/docs/templates/agent-dag.schema.json +411 -411
- package/docs/templates/agent-dag.supervised-implementation.json +620 -620
- package/docs/templates/backend-test-analysis.schema.json +44 -44
- package/docs/templates/backend-test-case-manifest.schema.json +190 -190
- package/docs/templates/backend-test-dag.classify.prompt.md +75 -75
- package/docs/templates/backend-test-dag.generate-pytest.prompt.md +204 -204
- package/docs/templates/backend-test-dag.json +559 -559
- package/docs/templates/backend-test-dag.retrospect.prompt.md +139 -139
- package/docs/templates/backend-test-dag.review-cases.prompt.md +83 -83
- package/docs/templates/backend-test-execution.schema.json +133 -133
- package/docs/templates/backend-test-result.schema.json +99 -99
- package/docs/templates/branch-merge-report.md +0 -1
- package/docs/templates/exec-plan.md +64 -64
- package/docs/templates/feature-spec.md +53 -53
- package/docs/templates/frontend-design-contract.md +42 -42
- package/docs/templates/frontend-eval/fixtures/failures/01-type-build-error.md +17 -17
- package/docs/templates/frontend-eval/fixtures/failures/02-unit-component-test-fail.md +16 -16
- package/docs/templates/frontend-eval/fixtures/failures/03-fixture-schema-drift.md +16 -16
- package/docs/templates/frontend-eval/fixtures/failures/04-missing-loading-empty-error-state.md +16 -16
- package/docs/templates/frontend-eval/fixtures/failures/05-forbidden-write-writeset-expansion.md +16 -16
- package/docs/templates/frontend-eval/fixtures/failures/06-unapproved-dependency-add.md +16 -16
- package/docs/templates/frontend-eval/fixtures/failures/07-mock-production-on.md +21 -21
- package/docs/templates/frontend-eval/fixtures/functional/01-simple-component-style.md +29 -29
- package/docs/templates/frontend-eval/fixtures/functional/02-form-validation.md +28 -28
- package/docs/templates/frontend-eval/fixtures/functional/03-list-detail-page.md +28 -28
- package/docs/templates/frontend-eval/fixtures/functional/04-api-mock.md +29 -29
- package/docs/templates/frontend-eval/fixtures/functional/05-permission-auth-gated-ui.md +27 -27
- package/docs/templates/frontend-eval/fixtures/functional/06-ssr-server-client-boundary.md +28 -28
- package/docs/templates/frontend-eval/fixtures/functional/07-shared-public-component-api.md +28 -28
- package/docs/templates/frontend-eval/fixtures/functional/08-pure-local-no-remote.md +27 -27
- package/docs/templates/frontend-eval/metrics.md +138 -138
- package/docs/templates/frontend-eval/smoke-targets.md +53 -53
- package/docs/templates/frontend-implementation-contract.schema.json +27 -27
- package/docs/templates/frontend-task-constraints.md +35 -35
- package/docs/templates/frontend-task-requirement.md +70 -70
- package/docs/templates/frontend-test-dag.generate-cases.prompt.md +5 -5
- package/docs/templates/frontend-test-dag.json +23 -23
- package/docs/templates/frontend-test-dag.retrieve-context.prompt.md +3 -3
- package/docs/templates/frontend-test-dag.retrospect.prompt.md +3 -3
- package/docs/templates/frontend-test-dag.review-cases.prompt.md +3 -3
- package/docs/templates/frontend-test-dag.review-execution.prompt.md +3 -3
- package/docs/templates/harness.schema.json +221 -221
- package/docs/templates/hybrid-dag.json +188 -188
- package/docs/templates/init-evolution-review.md +35 -35
- package/docs/templates/interactive-ui-round2-experiment.md +66 -66
- package/docs/templates/knowledge-graph-bootstrap-dag.json +118 -118
- package/docs/templates/knowledge-sync-dag.json +178 -178
- package/docs/templates/knowledge-sync-draft.schema.json +71 -71
- package/docs/templates/product-line/AGENTS.md +8 -8
- package/docs/templates/product-line/README.md +9 -9
- package/docs/templates/product-line/acceptance.yaml +14 -14
- package/docs/templates/product-line/closeout.yaml +9 -9
- package/docs/templates/product-line/design.md +13 -13
- package/docs/templates/product-line/links.md +10 -10
- package/docs/templates/product-line/requirement.md +17 -17
- package/docs/templates/product-line/task-graph.yaml +15 -15
- package/docs/templates/product-line/task.yaml +64 -64
- package/docs/templates/product-line/test-plan.md +7 -7
- package/docs/templates/production-readiness-checklist.md +57 -57
- package/docs/templates/progress-log.md +17 -17
- package/docs/templates/project-start-checklist.md +9 -9
- package/docs/templates/qa-report.md +48 -48
- package/docs/templates/sprint-contract.md +29 -29
- package/docs/templates/worker-dogfood-evidence.md +80 -80
- package/docs/templates/worker-dogfood-setup.md +68 -68
- package/examples/decision-gate-agent-dag.json +173 -173
- package/examples/example-dag.json +46 -46
- package/examples/hybrid-loop-agent-dag.json +188 -188
- package/harness.json +66 -66
- package/package.json +78 -52
- package/scripts/kb-bootstrap-init-skeleton.sh +240 -240
- package/scripts/kb-graph-incremental-prepare.mjs +386 -386
- package/scripts/kb-graph-materialize.mjs +105 -105
- package/scripts/kb-graph-promote.mjs +164 -164
- package/scripts/kb-query.mjs +554 -554
- package/skills/agent-worker/SKILL.md +39 -39
- package/skills/agent-worker/references/agent-worker-operator.md +60 -60
- package/skills/ai-engineering-context/SKILL.md +48 -48
- package/skills/analyze-product-dependencies/SKILL.md +67 -67
- package/skills/analyze-product-dependencies/agents/openai.yaml +4 -4
- package/skills/analyze-product-dependencies/references/api-documentation-schema.md +30 -30
- package/skills/analyze-product-dependencies/references/dependency-analysis-schema.md +28 -28
- package/skills/analyze-product-dependencies/references/example.md +76 -76
- package/skills/analyze-product-dependencies/references/forward-test-cases.md +35 -35
- package/skills/analyze-product-dependencies/references/input-contract.md +11 -11
- package/skills/analyze-product-dependencies/references/scouting-rules.md +61 -61
- package/skills/analyze-product-dependencies/scripts/test-validators.mjs +267 -267
- package/skills/analyze-product-dependencies/scripts/validate-api-documentation.mjs +101 -101
- package/skills/analyze-product-dependencies/scripts/validate-dependency-analysis.mjs +142 -142
- package/skills/analyze-product-dependencies/scripts/validate-product-requirement-input.mjs +76 -76
- package/skills/analyze-product-dependencies/scripts/validation-helpers.mjs +146 -146
- package/skills/analyze-product-requirements/SKILL.md +90 -90
- package/skills/analyze-product-requirements/agents/openai.yaml +4 -4
- package/skills/analyze-product-requirements/references/acceptance-criteria.md +91 -91
- package/skills/analyze-product-requirements/references/clarification-and-knowledge.md +56 -56
- package/skills/analyze-product-requirements/references/example.md +86 -86
- package/skills/analyze-product-requirements/references/forward-test-cases.md +66 -66
- package/skills/analyze-product-requirements/references/product-analysis-schema.md +32 -32
- package/skills/analyze-product-requirements/references/product-requirement-schema.md +33 -33
- package/skills/analyze-product-requirements/references/requirement-clarification-schema.md +35 -35
- package/skills/analyze-product-requirements/scripts/test-validators.mjs +193 -193
- package/skills/analyze-product-requirements/scripts/validate-product-analysis.mjs +69 -69
- package/skills/analyze-product-requirements/scripts/validate-product-requirement.mjs +97 -97
- package/skills/analyze-product-requirements/scripts/validate-requirement-clarification.mjs +98 -98
- package/skills/analyze-product-requirements/scripts/validation-helpers.mjs +156 -156
- package/skills/browser-tools/SKILL.md +196 -196
- package/skills/browser-tools/browser-content.js +103 -103
- package/skills/browser-tools/browser-cookies.js +35 -35
- package/skills/browser-tools/browser-eval.js +53 -53
- package/skills/browser-tools/browser-hn-scraper.js +108 -108
- package/skills/browser-tools/browser-nav.js +44 -44
- package/skills/browser-tools/browser-pick.js +162 -162
- package/skills/browser-tools/browser-screenshot.js +34 -34
- package/skills/browser-tools/browser-start.js +86 -86
- package/skills/browser-tools/package-lock.json +2556 -2556
- package/skills/browser-tools/package.json +19 -19
- package/skills/code-review-core/SKILL.md +20 -20
- package/skills/codebase-scout/SKILL.md +19 -19
- package/skills/frontend-design-review/SKILL.md +66 -66
- package/skills/frontend-design-review/references/review-checklist.md +40 -58
- package/skills/frontend-implementation/SKILL.md +49 -49
- package/skills/frontend-implementation/references/code-standards.md +32 -32
- package/skills/frontend-implementation/references/design-spec.md +46 -46
- package/skills/frontend-implementation/references/node-contracts.md +27 -27
- package/skills/frontend-review/SKILL.md +61 -59
- package/skills/frontend-review/references/review-findings.md +48 -47
- package/skills/frontend-verification/SKILL.md +55 -53
- package/skills/frontend-verification/references/verification-checklist.md +59 -68
- package/skills/grill-me/SKILL.md +10 -10
- package/skills/grill-with-docs/SKILL.md +88 -88
- package/skills/grill-with-docs/adr-format.md +47 -47
- package/skills/grill-with-docs/context-format.md +60 -60
- package/skills/init-capability-evolution/SKILL.md +70 -70
- package/skills/loop-agent/SKILL.md +151 -151
- package/skills/loop-agent/references/README.md +67 -67
- package/skills/loop-agent/references/command-reference.md +527 -527
- package/skills/loop-agent/references/docs-converge.md +126 -126
- package/skills/loop-agent/references/harness-policy.md +263 -263
- package/skills/loop-agent/references/hybrid-dag.md +243 -243
- package/skills/loop-agent/references/learned/README.md +21 -21
- package/skills/loop-agent/references/long-running-loop.md +57 -57
- package/skills/loop-agent/references/model-routing.md +36 -36
- package/skills/loop-agent/references/multi-worktree.md +54 -54
- package/skills/loop-agent/references/one-shot-runs.md +85 -85
- package/skills/loop-agent/references/orchestrator-and-interventions.md +169 -169
- package/skills/loop-agent/references/pi-prompt.md +23 -23
- package/skills/loop-agent/references/pi-subagent-assisted-mode.md +84 -84
- package/skills/loop-agent/references/post-implementation-and-patterns.md +44 -44
- package/skills/loop-agent/references/task-workflow.md +89 -89
- package/skills/loop-agent/references/verification-and-failure-handling.md +141 -141
- package/skills/playwright-cli/SKILL.md +420 -420
- package/skills/playwright-cli/references/element-attributes.md +23 -23
- package/skills/playwright-cli/references/playwright-tests.md +39 -39
- package/skills/playwright-cli/references/request-mocking.md +87 -87
- package/skills/playwright-cli/references/running-code.md +241 -241
- package/skills/playwright-cli/references/session-management.md +225 -225
- package/skills/playwright-cli/references/storage-state.md +275 -275
- package/skills/playwright-cli/references/test-generation.md +433 -433
- package/skills/playwright-cli/references/tracing.md +139 -139
- package/skills/playwright-cli/references/video-recording.md +143 -143
- package/skills/playwright-cli-case-generator/SKILL.md +74 -74
- package/skills/requesting-code-review/SKILL.md +101 -101
- package/skills/requesting-code-review/code-reviewer.md +168 -168
- package/skills/systematic-debugging/CREATION-LOG.md +119 -119
- package/skills/systematic-debugging/SKILL.md +296 -296
- package/skills/systematic-debugging/condition-based-waiting-example.ts +158 -158
- package/skills/systematic-debugging/condition-based-waiting.md +115 -115
- package/skills/systematic-debugging/defense-in-depth.md +122 -122
- package/skills/systematic-debugging/find-polluter.sh +63 -63
- package/skills/systematic-debugging/root-cause-tracing.md +169 -169
- package/skills/systematic-debugging/test-academic.md +14 -14
- package/skills/systematic-debugging/test-pressure-1.md +58 -58
- package/skills/systematic-debugging/test-pressure-2.md +68 -68
- package/skills/systematic-debugging/test-pressure-3.md +69 -69
- package/skills/test-driven-development/SKILL.md +20 -20
- package/skills/using-git-worktrees/SKILL.md +215 -215
- package/skills/verification-before-completion/SKILL.md +154 -154
- package/skills/webapp-testing/SKILL.md +19 -19
- package/docs/agent-dag-recovery-playbook.md +0 -195
- package/docs/agent-dag-runner.md +0 -67
- package/docs/cursor-prompt-sidecar.md +0 -36
- package/docs/decisions/README.md +0 -18
- package/docs/design/README.md +0 -167
- package/docs/development-principles.md +0 -73
- package/docs/exec-plans/README.md +0 -6
- package/docs/exec-plans/active/README.md +0 -12
- package/docs/exec-plans/completed/README.md +0 -107
- package/docs/feature-workflow.md +0 -414
- package/docs/loop-agent-harness.md +0 -142
- package/docs/production-readiness.md +0 -96
- package/docs/progress/README.md +0 -80
- package/docs/reports/README.md +0 -159
- package/docs/verification-matrix.md +0 -70
- package/scripts/check-product-line-docs.sh +0 -29
- package/scripts/check-task-pool-root.sh +0 -32
- package/scripts/kb-graph-incremental-prepare.sh +0 -5
- package/scripts/kb-graph-materialize.sh +0 -4
- package/scripts/kb-graph-promote.sh +0 -4
- package/scripts/kb-query.sh +0 -5
|
@@ -1,154 +1,154 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: verification-before-completion
|
|
3
|
-
description: 在宣称 work complete、fixed 或 passing,或在 commit / 创建 PR 之前使用——须先运行 verification commands 并确认 output,再作任何 success claims;始终 evidence before assertions
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# Verification Before Completion
|
|
7
|
-
|
|
8
|
-
## Overview
|
|
9
|
-
|
|
10
|
-
未经验证就宣称 work complete 是不诚实,不是效率。
|
|
11
|
-
|
|
12
|
-
**Core principle:** 始终 evidence before claims。
|
|
13
|
-
|
|
14
|
-
**违反本条字面即违反其精神。**
|
|
15
|
-
|
|
16
|
-
## The Iron Law
|
|
17
|
-
|
|
18
|
-
```
|
|
19
|
-
NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE
|
|
20
|
-
```
|
|
21
|
-
|
|
22
|
-
若本 message 中尚未运行 verification command,不得宣称 passes。
|
|
23
|
-
|
|
24
|
-
## The Gate Function
|
|
25
|
-
|
|
26
|
-
```
|
|
27
|
-
BEFORE claiming any status or expressing satisfaction:
|
|
28
|
-
|
|
29
|
-
1. IDENTIFY: What command proves this claim?
|
|
30
|
-
2. RUN: Execute the FULL command (fresh, complete)
|
|
31
|
-
3. READ: Full output, check exit code, count failures
|
|
32
|
-
4. VERIFY: Does output confirm the claim?
|
|
33
|
-
- If NO: State actual status with evidence
|
|
34
|
-
- If YES: State claim WITH evidence
|
|
35
|
-
5. ONLY THEN: Make the claim
|
|
36
|
-
|
|
37
|
-
Skip any step = lying, not verifying
|
|
38
|
-
```
|
|
39
|
-
|
|
40
|
-
## Common Failures
|
|
41
|
-
|
|
42
|
-
| Claim | Requires | Not Sufficient |
|
|
43
|
-
|-------|----------|----------------|
|
|
44
|
-
| Tests pass | Test command output: 0 failures | Previous run, "should pass" |
|
|
45
|
-
| Linter clean | Linter output: 0 errors | Partial check, extrapolation |
|
|
46
|
-
| Build succeeds | Build command: exit 0 | Linter passing, logs look good |
|
|
47
|
-
| Bug fixed | Test original symptom: passes | Code changed, assumed fixed |
|
|
48
|
-
| Regression test works | Red-green cycle verified | Test passes once |
|
|
49
|
-
| Agent completed | VCS diff shows changes | Agent reports "success" |
|
|
50
|
-
| Requirements met | Line-by-line checklist | Tests passing |
|
|
51
|
-
| Harness integrity | `scripts/check-repo.sh` exit 0 | Files look correct |
|
|
52
|
-
| Harness CI | `scripts/ci.sh` exit 0 | Individual checks pass |
|
|
53
|
-
| Contract handoff | Handoff checklist completed + progress/report updated | "Should be fine"
|
|
54
|
-
|
|
55
|
-
## Red Flags - STOP
|
|
56
|
-
|
|
57
|
-
- 使用 "should"、"probably"、"seems to"
|
|
58
|
-
- 验证前表达满意("Great!"、"Perfect!"、"Done!" 等)
|
|
59
|
-
- 未验证就要 commit/push/PR
|
|
60
|
-
- 信任 agent success reports
|
|
61
|
-
- 依赖 partial verification
|
|
62
|
-
- 认为 "just this once"
|
|
63
|
-
- 疲惫想结束工作
|
|
64
|
-
- **任何未运行 verification 却暗示 success 的措辞**
|
|
65
|
-
|
|
66
|
-
## Rationalization Prevention
|
|
67
|
-
|
|
68
|
-
| Excuse | Reality |
|
|
69
|
-
|--------|---------|
|
|
70
|
-
| "Should work now" | RUN the verification |
|
|
71
|
-
| "I'm confident" | Confidence ≠ evidence |
|
|
72
|
-
| "Just this once" | No exceptions |
|
|
73
|
-
| "Linter passed" | Linter ≠ compiler |
|
|
74
|
-
| "Agent said success" | Verify independently |
|
|
75
|
-
| "I'm tired" | Exhaustion ≠ excuse |
|
|
76
|
-
| "Partial check is enough" | Partial proves nothing |
|
|
77
|
-
| "Different words so rule doesn't apply" | Spirit over letter |
|
|
78
|
-
|
|
79
|
-
## Key Patterns
|
|
80
|
-
|
|
81
|
-
**Tests:**
|
|
82
|
-
```
|
|
83
|
-
✅ [Run test command] [See: 34/34 pass] "All tests pass"
|
|
84
|
-
❌ "Should pass now" / "Looks correct"
|
|
85
|
-
```
|
|
86
|
-
|
|
87
|
-
**Regression tests (TDD Red-Green):**
|
|
88
|
-
```
|
|
89
|
-
✅ Write → Run (pass) → Revert fix → Run (MUST FAIL) → Restore → Run (pass)
|
|
90
|
-
❌ "I've written a regression test" (without red-green verification)
|
|
91
|
-
```
|
|
92
|
-
|
|
93
|
-
**Build:**
|
|
94
|
-
```
|
|
95
|
-
✅ [Run build] [See: exit 0] "Build passes"
|
|
96
|
-
❌ "Linter passed" (linter doesn't check compilation)
|
|
97
|
-
```
|
|
98
|
-
|
|
99
|
-
**Requirements:**
|
|
100
|
-
```
|
|
101
|
-
✅ Re-read plan → Create checklist → Verify each → Report gaps or completion
|
|
102
|
-
❌ "Tests pass, phase complete"
|
|
103
|
-
```
|
|
104
|
-
|
|
105
|
-
**Agent delegation:**
|
|
106
|
-
```
|
|
107
|
-
✅ Agent reports success → Check VCS diff → Verify changes → Report actual state
|
|
108
|
-
❌ Trust agent report
|
|
109
|
-
```
|
|
110
|
-
|
|
111
|
-
## Harness-Specific Verification
|
|
112
|
-
|
|
113
|
-
在 harness-governed repo 中工作(存在 `harness.json`)时:
|
|
114
|
-
|
|
115
|
-
- **Docs/structure changes** → `bash scripts/check-repo.sh`
|
|
116
|
-
- **Full-repo delivery** → `bash scripts/ci.sh`
|
|
117
|
-
- **Cross-platform changes** → 验证 OpenCode 与 Pi-Agent 两条路径
|
|
118
|
-
- **Contract changes** → 验证 contract docs 已更新 + tests 对齐
|
|
119
|
-
- **Handoff** → 宣称 complete 前运行 `handoff check`
|
|
120
|
-
|
|
121
|
-
完整 command 选择见项目 `ai_workspace/loop-agent/verification-matrix.md`。
|
|
122
|
-
|
|
123
|
-
## Why This Matters
|
|
124
|
-
|
|
125
|
-
来自 24 条 failure memories:
|
|
126
|
-
- human partner 说 "I don't believe you" — trust 已破裂
|
|
127
|
-
- Undefined functions 已 ship — 会 crash
|
|
128
|
-
- Missing requirements 已 ship — 功能不完整
|
|
129
|
-
- 虚假完成浪费时间 → redirect → rework
|
|
130
|
-
- 违反:"Honesty is a core value. If you lie, you'll be replaced."
|
|
131
|
-
|
|
132
|
-
## When To Apply
|
|
133
|
-
|
|
134
|
-
**在以下情况之前 ALWAYS:**
|
|
135
|
-
- 任何 success/completion claims 的变体
|
|
136
|
-
- 任何表达满意
|
|
137
|
-
- 任何关于 work state 的正面陈述
|
|
138
|
-
- Commit、PR creation、task completion
|
|
139
|
-
- 进入 next task
|
|
140
|
-
- 委派给 agents
|
|
141
|
-
|
|
142
|
-
**规则适用于:**
|
|
143
|
-
- 精确短语
|
|
144
|
-
- paraphrases 与同义词
|
|
145
|
-
- success 的暗示
|
|
146
|
-
- 任何暗示 completion/correctness 的沟通
|
|
147
|
-
|
|
148
|
-
## The Bottom Line
|
|
149
|
-
|
|
150
|
-
**Verification 无捷径。**
|
|
151
|
-
|
|
152
|
-
Run the command. Read the output. THEN claim the result.
|
|
153
|
-
|
|
154
|
-
This is non-negotiable.
|
|
1
|
+
---
|
|
2
|
+
name: verification-before-completion
|
|
3
|
+
description: 在宣称 work complete、fixed 或 passing,或在 commit / 创建 PR 之前使用——须先运行 verification commands 并确认 output,再作任何 success claims;始终 evidence before assertions
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Verification Before Completion
|
|
7
|
+
|
|
8
|
+
## Overview
|
|
9
|
+
|
|
10
|
+
未经验证就宣称 work complete 是不诚实,不是效率。
|
|
11
|
+
|
|
12
|
+
**Core principle:** 始终 evidence before claims。
|
|
13
|
+
|
|
14
|
+
**违反本条字面即违反其精神。**
|
|
15
|
+
|
|
16
|
+
## The Iron Law
|
|
17
|
+
|
|
18
|
+
```
|
|
19
|
+
NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
若本 message 中尚未运行 verification command,不得宣称 passes。
|
|
23
|
+
|
|
24
|
+
## The Gate Function
|
|
25
|
+
|
|
26
|
+
```
|
|
27
|
+
BEFORE claiming any status or expressing satisfaction:
|
|
28
|
+
|
|
29
|
+
1. IDENTIFY: What command proves this claim?
|
|
30
|
+
2. RUN: Execute the FULL command (fresh, complete)
|
|
31
|
+
3. READ: Full output, check exit code, count failures
|
|
32
|
+
4. VERIFY: Does output confirm the claim?
|
|
33
|
+
- If NO: State actual status with evidence
|
|
34
|
+
- If YES: State claim WITH evidence
|
|
35
|
+
5. ONLY THEN: Make the claim
|
|
36
|
+
|
|
37
|
+
Skip any step = lying, not verifying
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
## Common Failures
|
|
41
|
+
|
|
42
|
+
| Claim | Requires | Not Sufficient |
|
|
43
|
+
|-------|----------|----------------|
|
|
44
|
+
| Tests pass | Test command output: 0 failures | Previous run, "should pass" |
|
|
45
|
+
| Linter clean | Linter output: 0 errors | Partial check, extrapolation |
|
|
46
|
+
| Build succeeds | Build command: exit 0 | Linter passing, logs look good |
|
|
47
|
+
| Bug fixed | Test original symptom: passes | Code changed, assumed fixed |
|
|
48
|
+
| Regression test works | Red-green cycle verified | Test passes once |
|
|
49
|
+
| Agent completed | VCS diff shows changes | Agent reports "success" |
|
|
50
|
+
| Requirements met | Line-by-line checklist | Tests passing |
|
|
51
|
+
| Harness integrity | `scripts/check-repo.sh` exit 0 | Files look correct |
|
|
52
|
+
| Harness CI | `scripts/ci.sh` exit 0 | Individual checks pass |
|
|
53
|
+
| Contract handoff | Handoff checklist completed + progress/report updated | "Should be fine"
|
|
54
|
+
|
|
55
|
+
## Red Flags - STOP
|
|
56
|
+
|
|
57
|
+
- 使用 "should"、"probably"、"seems to"
|
|
58
|
+
- 验证前表达满意("Great!"、"Perfect!"、"Done!" 等)
|
|
59
|
+
- 未验证就要 commit/push/PR
|
|
60
|
+
- 信任 agent success reports
|
|
61
|
+
- 依赖 partial verification
|
|
62
|
+
- 认为 "just this once"
|
|
63
|
+
- 疲惫想结束工作
|
|
64
|
+
- **任何未运行 verification 却暗示 success 的措辞**
|
|
65
|
+
|
|
66
|
+
## Rationalization Prevention
|
|
67
|
+
|
|
68
|
+
| Excuse | Reality |
|
|
69
|
+
|--------|---------|
|
|
70
|
+
| "Should work now" | RUN the verification |
|
|
71
|
+
| "I'm confident" | Confidence ≠ evidence |
|
|
72
|
+
| "Just this once" | No exceptions |
|
|
73
|
+
| "Linter passed" | Linter ≠ compiler |
|
|
74
|
+
| "Agent said success" | Verify independently |
|
|
75
|
+
| "I'm tired" | Exhaustion ≠ excuse |
|
|
76
|
+
| "Partial check is enough" | Partial proves nothing |
|
|
77
|
+
| "Different words so rule doesn't apply" | Spirit over letter |
|
|
78
|
+
|
|
79
|
+
## Key Patterns
|
|
80
|
+
|
|
81
|
+
**Tests:**
|
|
82
|
+
```
|
|
83
|
+
✅ [Run test command] [See: 34/34 pass] "All tests pass"
|
|
84
|
+
❌ "Should pass now" / "Looks correct"
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
**Regression tests (TDD Red-Green):**
|
|
88
|
+
```
|
|
89
|
+
✅ Write → Run (pass) → Revert fix → Run (MUST FAIL) → Restore → Run (pass)
|
|
90
|
+
❌ "I've written a regression test" (without red-green verification)
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
**Build:**
|
|
94
|
+
```
|
|
95
|
+
✅ [Run build] [See: exit 0] "Build passes"
|
|
96
|
+
❌ "Linter passed" (linter doesn't check compilation)
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
**Requirements:**
|
|
100
|
+
```
|
|
101
|
+
✅ Re-read plan → Create checklist → Verify each → Report gaps or completion
|
|
102
|
+
❌ "Tests pass, phase complete"
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
**Agent delegation:**
|
|
106
|
+
```
|
|
107
|
+
✅ Agent reports success → Check VCS diff → Verify changes → Report actual state
|
|
108
|
+
❌ Trust agent report
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
## Harness-Specific Verification
|
|
112
|
+
|
|
113
|
+
在 harness-governed repo 中工作(存在 `harness.json`)时:
|
|
114
|
+
|
|
115
|
+
- **Docs/structure changes** → `bash scripts/check-repo.sh`
|
|
116
|
+
- **Full-repo delivery** → `bash scripts/ci.sh`
|
|
117
|
+
- **Cross-platform changes** → 验证 OpenCode 与 Pi-Agent 两条路径
|
|
118
|
+
- **Contract changes** → 验证 contract docs 已更新 + tests 对齐
|
|
119
|
+
- **Handoff** → 宣称 complete 前运行 `handoff check`
|
|
120
|
+
|
|
121
|
+
完整 command 选择见项目 `ai_workspace/loop-agent/verification-matrix.md`。
|
|
122
|
+
|
|
123
|
+
## Why This Matters
|
|
124
|
+
|
|
125
|
+
来自 24 条 failure memories:
|
|
126
|
+
- human partner 说 "I don't believe you" — trust 已破裂
|
|
127
|
+
- Undefined functions 已 ship — 会 crash
|
|
128
|
+
- Missing requirements 已 ship — 功能不完整
|
|
129
|
+
- 虚假完成浪费时间 → redirect → rework
|
|
130
|
+
- 违反:"Honesty is a core value. If you lie, you'll be replaced."
|
|
131
|
+
|
|
132
|
+
## When To Apply
|
|
133
|
+
|
|
134
|
+
**在以下情况之前 ALWAYS:**
|
|
135
|
+
- 任何 success/completion claims 的变体
|
|
136
|
+
- 任何表达满意
|
|
137
|
+
- 任何关于 work state 的正面陈述
|
|
138
|
+
- Commit、PR creation、task completion
|
|
139
|
+
- 进入 next task
|
|
140
|
+
- 委派给 agents
|
|
141
|
+
|
|
142
|
+
**规则适用于:**
|
|
143
|
+
- 精确短语
|
|
144
|
+
- paraphrases 与同义词
|
|
145
|
+
- success 的暗示
|
|
146
|
+
- 任何暗示 completion/correctness 的沟通
|
|
147
|
+
|
|
148
|
+
## The Bottom Line
|
|
149
|
+
|
|
150
|
+
**Verification 无捷径。**
|
|
151
|
+
|
|
152
|
+
Run the command. Read the output. THEN claim the result.
|
|
153
|
+
|
|
154
|
+
This is non-negotiable.
|
|
@@ -1,19 +1,19 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: webapp-testing
|
|
3
|
-
description: 任务明确涉及 browser 渲染行为时,用于前端或本地 web UI 验证。
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# Webapp Testing
|
|
7
|
-
|
|
8
|
-
仅当任务包含 browser UI 或本地 web app 时使用本 skill。
|
|
9
|
-
|
|
10
|
-
## 规则
|
|
11
|
-
|
|
12
|
-
- 优先使用项目现有的 dev server 与 test tooling。
|
|
13
|
-
- 当 visual 或 interaction 行为重要时,用 browser 或文档化的 UI test 命令验证渲染行为。
|
|
14
|
-
- UI 有变更时,检查 desktop 与 mobile 布局的 overlap、clipping、blank state。
|
|
15
|
-
- 默认不添加 networked services 或第三方 scan。
|
|
16
|
-
|
|
17
|
-
## Output
|
|
18
|
-
|
|
19
|
-
报告确切的 server 命令、URL、browser/test 命令与观察结果。
|
|
1
|
+
---
|
|
2
|
+
name: webapp-testing
|
|
3
|
+
description: 任务明确涉及 browser 渲染行为时,用于前端或本地 web UI 验证。
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Webapp Testing
|
|
7
|
+
|
|
8
|
+
仅当任务包含 browser UI 或本地 web app 时使用本 skill。
|
|
9
|
+
|
|
10
|
+
## 规则
|
|
11
|
+
|
|
12
|
+
- 优先使用项目现有的 dev server 与 test tooling。
|
|
13
|
+
- 当 visual 或 interaction 行为重要时,用 browser 或文档化的 UI test 命令验证渲染行为。
|
|
14
|
+
- UI 有变更时,检查 desktop 与 mobile 布局的 overlap、clipping、blank state。
|
|
15
|
+
- 默认不添加 networked services 或第三方 scan。
|
|
16
|
+
|
|
17
|
+
## Output
|
|
18
|
+
|
|
19
|
+
报告确切的 server 命令、URL、browser/test 命令与观察结果。
|
|
@@ -1,195 +0,0 @@
|
|
|
1
|
-
# Agent DAG Recovery Playbook(恢复手册)
|
|
2
|
-
|
|
3
|
-
> **关联**:[`agent-dag-runner.md`](agent-dag-runner.md)(CLI 与 run 语义)· [`templates/agent-dag-decision-gate.prompt.md`](templates/agent-dag-decision-gate.prompt.md)(Decision Gate 消费 recovery 证据)
|
|
4
|
-
|
|
5
|
-
## 定位
|
|
6
|
-
|
|
7
|
-
Agent DAG **recovery planning 是只读、派生、advisory** 的。`dag report` 与 `buildDagDecisionGateEvidence()` 从 `.harness/dag-runs/` 的 canonical facts 聚合 `normalizedFailureCategory` → `recoveryRecommendation`,供人工或 Decision Gate prompt 消费。
|
|
8
|
-
|
|
9
|
-
中断后不要从上游摘要手工生成 impl-only DAG。先修复 `.harness/tasks/<task-id>/source/` 或计划,再对同一 task 重新执行 `dag run-task`、严格 `dag validate` 和新的 `run-dag`。新生成的完整 DAG 会重新冻结 `sourceBinding` 并经过 contract/scout/plan/gate;v3 孤立 writer 如果既无来源绑定、也无只读 planner 上游,会被 strict governance 拒绝。完整规则见 [`design/dag-source-binding-and-recovery.md`](design/dag-source-binding-and-recovery.md)。
|
|
10
|
-
|
|
11
|
-
Production Readiness v0.1 在 normalized DAG category 之上增加 product-line routing。Report 与 doctor 输出应保留 raw DAG fact 并派生,不重写已完成 facts:
|
|
12
|
-
|
|
13
|
-
```text
|
|
14
|
-
raw_failure_category
|
|
15
|
-
dag_normalized_failure_category
|
|
16
|
-
product_line_failure_category
|
|
17
|
-
recommended_follow_up
|
|
18
|
-
```
|
|
19
|
-
|
|
20
|
-
Product-line taxonomy 定义见 `ai_workspace/loop-agent/design/state-and-failure-taxonomy.md`。
|
|
21
|
-
|
|
22
|
-
### 前端设计门禁专用恢复路径
|
|
23
|
-
|
|
24
|
-
前端 DAG 的 design gate shell 失败(`frontend-first-design-gate-shell`、`frontend-final-design-gate-shell`、`frontend-design-gate-shell`)**不路由为 `ProductBug` / `dev-fix`**。此类失败固定路由为:
|
|
25
|
-
|
|
26
|
-
- `productLineFailureCategory`: `ContractMismatch`
|
|
27
|
-
- `recommendedFollowUp`: `frontend-plan-revision-and-rerun`
|
|
28
|
-
|
|
29
|
-
恢复动作由 `planDagRecovery` 根据实际的 `normalizedFailureCategory` 和 run status 决定(通常为 `rerun-after-fix` 或 `manual-review`),但 product-line 维度的分类确保 Task Pool 和 morning report 不会将其混入普通 bug backlog。
|
|
30
|
-
|
|
31
|
-
**非目标(本 playbook 不覆盖、runner 不实现):**
|
|
32
|
-
|
|
33
|
-
- 自动 retry / resume 节点执行
|
|
34
|
-
- 修改 `completed/` 或 `paused/` 下的历史 run facts
|
|
35
|
-
- 把 `autoRetryEligible` 当作 runtime 触发器
|
|
36
|
-
- 仅凭 recovery 派生字段自动 approve Decision Gate
|
|
37
|
-
|
|
38
|
-
## 快速命令
|
|
39
|
-
|
|
40
|
-
```bash
|
|
41
|
-
cd .
|
|
42
|
-
|
|
43
|
-
# 全局 runtime 健康(active/paused/completed 摘要 + healthIssues;advisoryOnly)
|
|
44
|
-
npm run dev -- dag doctor
|
|
45
|
-
|
|
46
|
-
# 单 run 生命周期(approvalFlow、hasHumanApproval、nextRecommendedAction)
|
|
47
|
-
npm run dev -- dag status --run-id <run-id>
|
|
48
|
-
|
|
49
|
-
# 聚焦最新 paused run(--paused-latest ≡ --lifecycle paused --latest)
|
|
50
|
-
npm run dev -- dag report --paused-latest [--json|--markdown]
|
|
51
|
-
|
|
52
|
-
# 默认 compact Markdown 表格
|
|
53
|
-
npm run dev -- dag report --run-id <run-id>
|
|
54
|
-
|
|
55
|
-
# 机器可读 JSON(含 primaryFailure / primaryRecovery / downstreamSkippedNodes)
|
|
56
|
-
npm run dev -- dag report --run-id <run-id> --json
|
|
57
|
-
|
|
58
|
-
# 人类交接 Recovery Plan(四段结构化 Markdown)
|
|
59
|
-
npm run dev -- dag report --run-id <run-id> --markdown
|
|
60
|
-
|
|
61
|
-
# 过滤器
|
|
62
|
-
npm run dev -- dag report --failed-only # 仅失败/需恢复
|
|
63
|
-
npm run dev -- dag report --latest --failed-only # 最新一条需恢复 run
|
|
64
|
-
npm run dev -- dag report --action retry-node # 按 primaryRecovery.action 筛选
|
|
65
|
-
npm run dev -- dag report --lifecycle paused --action resume-or-reject
|
|
66
|
-
|
|
67
|
-
# Decision Gate envelope dry-run(不 resume/retry;validate 无效时 exit 1)
|
|
68
|
-
npm run dev -- dag decision inspect --run-id <run-id> [--node-id <node-id>]
|
|
69
|
-
npm run dev -- dag decision validate --run-id <run-id> [--node-id <node-id>]
|
|
70
|
-
```
|
|
71
|
-
|
|
72
|
-
### Paused run operator 路径
|
|
73
|
-
|
|
74
|
-
1. `dag report --paused-latest --json` 或 `dag doctor` — 定位最新 paused run 与 `primaryRecovery`
|
|
75
|
-
2. `dag status --run-id <id>` — 读 `approvalFlow`、`escalationArtifactPath`、`pendingNodes`
|
|
76
|
-
3. (可选)`dag decision validate --run-id <id>` — envelope preflight
|
|
77
|
-
4. `dag approve --run-id <id> --option <option-id>` → `dag resume --run-id <id>`;或 `dag reject --run-id <id> --reason "..."`
|
|
78
|
-
|
|
79
|
-
精确 approval 顺序见 [`agent-dag-runner.md`](agent-dag-runner.md) §Paused lifecycle。
|
|
80
|
-
|
|
81
|
-
Decision Gate prompt 侧:`buildDagDecisionGateEvidence()`(`./src/core/dag-decision-evidence.ts`)从 `DagRunReportEntry` 生成 prompt-friendly 摘要,字段与 JSON report 对齐,**不**写回 run state。
|
|
82
|
-
|
|
83
|
-
## `dag report --json` schema 锁定
|
|
84
|
-
|
|
85
|
-
- **Schema 文件**:`ai_workspace/loop-agent/templates/agent-dag-report.schema.json`
|
|
86
|
-
- **Envelope**:`{ schemaVersion: 1, runs: DagRunReportEntry[] }`
|
|
87
|
-
- **稳定消费字段**(Decision Gate / tooling 应依赖):`primaryFailure`、`primaryRecovery`、`downstreamSkippedNodes`、`recoveryRecommendation`、`normalizedFailureCategory`;node 级 `decisionEnvelope`、`artifacts`;paused 级 `pausedByNodeId`、`pauseReason`
|
|
88
|
-
- **测试**:`./test/dag-report.test.ts` §`dag report JSON schema contract` 对 fixture run 做 schema 校验
|
|
89
|
-
- **变更策略**:breaking 字段变更须 bump `schemaVersion` 并同步 schema 文件与测试
|
|
90
|
-
|
|
91
|
-
## Recovery Action 枚举
|
|
92
|
-
|
|
93
|
-
| Action | 含义 | 典型触发 |
|
|
94
|
-
|--------|------|----------|
|
|
95
|
-
| `none` | 无需恢复 | 成功完成 |
|
|
96
|
-
| `monitor` | 进行中,等待结束 | `PENDING` / `RUNNING` |
|
|
97
|
-
| `retry-node` | 修复瞬态条件后可重跑节点 | timeout;executor 瞬态(network/quota/rate-limit/unavailable) |
|
|
98
|
-
| `rerun-after-fix` | 先修根因再重跑 | auth、validation、shell-command、static-error、非瞬态 executor |
|
|
99
|
-
| `resume-or-reject` | 人工审批后继续或拒绝 | paused + decision-envelope / human-required |
|
|
100
|
-
| `manual-review` | 人工审查后再定路径 | write-guard、human-rejected、unknown、非 paused 的 decision-envelope |
|
|
101
|
-
| `inspect-upstream` | 先查上游失败 | SKIPPED 下游节点 |
|
|
102
|
-
| `unknown` | 未映射类别(不应出现在正常派生路径) | 内部兜底 |
|
|
103
|
-
|
|
104
|
-
## Product-Line Routing v0.1
|
|
105
|
-
|
|
106
|
-
| Product-line category | Default follow-up |
|
|
107
|
-
|---|---|
|
|
108
|
-
| `SpecUnclear` | `spec-clarification` |
|
|
109
|
-
| `ContractMismatch` | `architecture-contract-fix` |
|
|
110
|
-
| `ProductBug` | `dev-fix` |
|
|
111
|
-
| `TestBug` | `qa-fix-test` |
|
|
112
|
-
| `EnvFailure` | `env-fix` 或 retry verify |
|
|
113
|
-
| `FlakyTest` | `flaky-test-analysis` |
|
|
114
|
-
| `RiskyChange` | `human-review` / `architecture-review` |
|
|
115
|
-
| `DependencyFailure` | unblock dependency |
|
|
116
|
-
| `NeedsHuman` | `human-review` |
|
|
117
|
-
| `Unknown` | human triage |
|
|
118
|
-
|
|
119
|
-
## 类别 → 动作 → operator 指引
|
|
120
|
-
|
|
121
|
-
| Normalized category | Recovery action | Operator guidance | Anti-patterns |
|
|
122
|
-
|---------------------|-----------------|-------------------|---------------|
|
|
123
|
-
| `success` | `none` | 归档验收;按需 review artifacts | 对成功 run 发起 retry |
|
|
124
|
-
| `timeout` | `retry-node` | 查日志/artifacts 确认瞬态;人工重跑节点 | 未查根因就循环重试;指望 runner 自动 retry |
|
|
125
|
-
| `executor`(network/quota/rate-limit/unavailable) | `retry-node` | 等后端/配额恢复后重跑 | 把 auth/validation 误判为瞬态 executor |
|
|
126
|
-
| `executor`(其他 raw) | `rerun-after-fix` | 查 executor.jsonl、node result | 盲目 retry 非瞬态 backend 错误 |
|
|
127
|
-
| `auth` | `rerun-after-fix` | 更新 API key/凭证后重跑 | 在凭证未修复时 retry |
|
|
128
|
-
| `write-guard` | `manual-review` | 审 writeSet/writePolicy、prompt、result.summary | read-only 节点写根 `artifacts/`;扩大 writeSet 掩盖违规 |
|
|
129
|
-
| `validation` | `rerun-after-fix` | 修 schema/output/test 后再跑 | 跳过验证直接 approve |
|
|
130
|
-
| `shell-command` | `rerun-after-fix` | 读 stdout/stderr、修命令或 repo 状态 | 只重跑 shell 不改命令 |
|
|
131
|
-
| `static-error` | `rerun-after-fix` | 查 static config 与 emitted markdown | 当 LLM 节点 retry |
|
|
132
|
-
| `decision-envelope`(paused) | `resume-or-reject` | `dag approve --run-id <id> --option <option-id>` / `dag reject --run-id <id> --reason "..."` → `dag resume --run-id <id>` | 未读 envelope 就 approve;用 recovery 字段单独 auto-approve |
|
|
133
|
-
| `decision-envelope`(非 paused) | `manual-review` | 读 decision.envelope.json / validation artifact | 绕过 Decision Gate schema |
|
|
134
|
-
| `human-required`(paused) | `resume-or-reject` | 提供人工输入 → approve/resume | 在 escalation 未解决时 resume |
|
|
135
|
-
| `human-required`(非 paused) | `manual-review` | 读 human-escalation artifacts | 忽略 `requiresHuman` |
|
|
136
|
-
| `human-rejected` | `manual-review` | 修订 contract/source;**新 run** | 对同一 contract 自动 retry |
|
|
137
|
-
| `skipped` | `inspect-upstream` | 修上游 ERROR/SKIPPED 再考虑下游 | 直接 retry SKIPPED 节点 |
|
|
138
|
-
| `unknown` | `manual-review` | 读 state.json、executor.jsonl、node artifacts | 假设 `autoRetryEligible` 会触发执行 |
|
|
139
|
-
|
|
140
|
-
## Handoff Recovery Plan 结构
|
|
141
|
-
|
|
142
|
-
`dag report --markdown` 的 **Recovery Plan** 含四段(与 JSON 稳定字段一一对应):
|
|
143
|
-
|
|
144
|
-
1. **Primary Failure** — `primaryFailure`(node 或 run scope)
|
|
145
|
-
2. **Recovery Action** — `primaryRecovery`(action、summary、reason、flags、commandHint)
|
|
146
|
-
3. **Blocked Downstream / Skipped Nodes** — `downstreamSkippedNodes`
|
|
147
|
-
4. **Recommended Operator Action** — 面向 operator 的步骤摘要
|
|
148
|
-
|
|
149
|
-
保存 handoff 时重定向到平台临时目录或 `ai_workspace/loop-agent/reports/`,不要写入 `.harness/dag-runs/`。
|
|
150
|
-
|
|
151
|
-
## Decision Gate 消费约定
|
|
152
|
-
|
|
153
|
-
1. 优先 `dag report --json` 或 `buildDagDecisionGateEvidence()` 的 **verified** 派生摘要。
|
|
154
|
-
2. 映射到 `decision` / `nextAction` 须保守;recovery 证据是 **advisory only, not an execution directive**。
|
|
155
|
-
3. `autoRetryEligible: true` 仅表示「规划上可人工重试」,**不**触发 runner。
|
|
156
|
-
4. paused run 的人类路径仍是 M5 CLI:`dag approve --run-id <id> --option <option-id>` / `dag reject --run-id <id> --reason "..."` / `dag resume --run-id <id>`(见 [`agent-dag-runner.md`](agent-dag-runner.md) §Decision Gate)。
|
|
157
|
-
5. Envelope 干跑:`dag decision inspect|validate` 重解析 run facts;`validate` 无效时 exit 1;**不**写 artifact、**不** resume。
|
|
158
|
-
|
|
159
|
-
## Active stale run recovery(advisory detection)
|
|
160
|
-
|
|
161
|
-
`dag doctor` 与 `dag status` 通过 `detectDagRunHealthIssues()` 检测 lifecycle 不一致,**不** mutate run facts。
|
|
162
|
-
|
|
163
|
-
| Code | 典型场景 | operator 指引 |
|
|
164
|
-
|------|----------|------------|
|
|
165
|
-
| `terminal-in-active` | run 已完成但 `active/<run-id>/` 残留 | 对照 `completed/` canonical facts;手动 archive 或删除 stale 目录 |
|
|
166
|
-
| `paused-in-active` | pause 后目录未迁至 `paused/` | `dag doctor` 诊断;修复 facts 后再 approve/resume |
|
|
167
|
-
| `lifecycle-status-mismatch` | `paused/` 下 status 非 paused | 同上 |
|
|
168
|
-
| `missing-approval-artifact` | approve 后 artifact 缺失 | 勿 resume;re-approve 或 restore artifact |
|
|
169
|
-
| `non-terminal-in-completed` | completed 目录 status 异常 | manual-review only |
|
|
170
|
-
| `run-id-mismatch` / `missing-state-json` | 目录损坏或命名错误 | Inspect;勿 auto-mutate completed facts |
|
|
171
|
-
|
|
172
|
-
**Deferred runtime**:无 `dag recover apply` 或自动 cleanup;未来可能增加只读 `dag recover plan`(设计占位,未实现)。
|
|
173
|
-
|
|
174
|
-
## 事实源与边界
|
|
175
|
-
|
|
176
|
-
| 类型 | 位置 | 规则 |
|
|
177
|
-
|------|------|------|
|
|
178
|
-
| Canonical run facts | `.harness/dag-runs/{active\|paused\|completed}/<run-id>/` | **只读**;report 不写回 |
|
|
179
|
-
| 派生 report | stdout / 重定向文件 | 可随时再生 |
|
|
180
|
-
| 工作块摘要 | 根 `artifacts/` | 非 per-run 历史;read-only DAG 节点不得写 |
|
|
181
|
-
|
|
182
|
-
## 验证
|
|
183
|
-
|
|
184
|
-
```bash
|
|
185
|
-
cd . && npx vitest run \
|
|
186
|
-
test/dag-report.test.ts \
|
|
187
|
-
test/dag-recovery-recommendation.test.ts \
|
|
188
|
-
test/dag-decision-gate-recovery-dogfood.test.ts \
|
|
189
|
-
test/dag-decision-evidence.test.ts \
|
|
190
|
-
test/dag-decision-envelope.test.ts \
|
|
191
|
-
test/dag-approve-resume.test.ts \
|
|
192
|
-
test/cli-contract.test.ts
|
|
193
|
-
```
|
|
194
|
-
|
|
195
|
-
实现细节与映射逻辑:`./src/core/dag-recovery-recommendation.ts`、`dag-report.ts`、`dag-decision-evidence.ts`。
|
package/docs/agent-dag-runner.md
DELETED
|
@@ -1,67 +0,0 @@
|
|
|
1
|
-
# Agent DAG Runner
|
|
2
|
-
|
|
3
|
-
Agent DAG 是 loop-agent 的声明式编排 runtime。DAG 将工作拆为节点、按序执行 eligible ranks、记录 artifacts,并用 gate 做 review 与验证。
|
|
4
|
-
|
|
5
|
-
## 基本用法
|
|
6
|
-
|
|
7
|
-
```bash
|
|
8
|
-
loop-agent dag run-task <task-id> --profile auto --strict-models --output <temp-dir>/<task-id>-dag.json
|
|
9
|
-
loop-agent dag validate --dag <temp-dir>/<task-id>-dag.json --strict-models --strict-governance
|
|
10
|
-
loop-agent run-dag --dag <temp-dir>/<task-id>-dag.json --cwd .
|
|
11
|
-
```
|
|
12
|
-
|
|
13
|
-
`<temp-dir>` 为平台原生临时目录。Windows 上 `--output`、`--dag`、`--cwd` 的实际值用原生路径。
|
|
14
|
-
|
|
15
|
-
## Executors
|
|
16
|
-
|
|
17
|
-
- `static`:确定性生成的 artifacts 或 notes
|
|
18
|
-
- `shell`:验证与文件系统检查
|
|
19
|
-
- `pi`:规划、review、诊断;节点设 `toolProfile: "write"` 时有界写入
|
|
20
|
-
|
|
21
|
-
## Retry (read-only Pi nodes)
|
|
22
|
-
|
|
23
|
-
planner/scout/reviewer/verifier/closeout 角色的只读 Pi 节点可声明 opt-in `retryPolicy`,用于在同一 run 内有界重试模型连接中断、provider 限流、临时不可用或请求 timeout。生成器会为这些安全节点自动声明默认策略:总尝试次数 3(手工配置上限 5),指数退避,单次等待上限 30s。
|
|
24
|
-
|
|
25
|
-
- 仅以下原始失败分类默认可重试:`timeout`、`network`、`rate-limit`、`unavailable`。
|
|
26
|
-
- `quota`、`auth`、`invalid-output`、`write-guard`、`decision-envelope` 与未知失败不重试。`quota` 不是 rate limit,不会被自动重试。
|
|
27
|
-
- 资格由确定性 helper 判断:仅 `writePolicy=read-only|none`(或 Pi 默认只读)的 planner/scout/reviewer/verifier/closeout 可用。supervisor、implementer、writer(`toolProfile=write` 或 `writePolicy=exclusive`)、docs-only、dynamic、shell、static 与 decision-gate 节点一律不重试,DAG validation 会拒绝其策略。
|
|
28
|
-
- 每次 attempt 写入独立不可变证据(`<node-id>/attempt-<n>.json`,run-relative path),最终 node record 的 `attempts` 字段引用完整 attempt 历史;后一次成功不会覆盖前一次失败证据。
|
|
29
|
-
- 重试期间复用同一 run、controller identity、skill snapshot、prompt、model 与上游输入。节点终态的 `durationMs`、`tokensUsed`、`parsedEvents` 聚合全部 attempts;退避等待会刷新 `lastActivityAt`,避免被误判为 node-quiet。当前退避会占用该节点所在的并发槽。
|
|
30
|
-
|
|
31
|
-
示例:
|
|
32
|
-
|
|
33
|
-
```json
|
|
34
|
-
{
|
|
35
|
-
"retryPolicy": {
|
|
36
|
-
"maxAttempts": 3,
|
|
37
|
-
"backoff": "exponential",
|
|
38
|
-
"initialDelayMs": 2000,
|
|
39
|
-
"maxDelayMs": 30000,
|
|
40
|
-
"retryCategories": ["timeout", "network", "rate-limit", "unavailable"]
|
|
41
|
-
}
|
|
42
|
-
}
|
|
43
|
-
```
|
|
44
|
-
|
|
45
|
-
未声明 `retryPolicy` 的历史 DAG 行为不变(单次执行、无 `attempts` 字段,也不新增 attempt artifact)。
|
|
46
|
-
|
|
47
|
-
## Skills
|
|
48
|
-
|
|
49
|
-
DAG spec 可声明 `defaults.skills`、`skillsByRole` 与节点级 `skills`。Runner 优先从目标项目 `.agents/skills/<skill-name>/SKILL.md` 解析本地指令,再回退到包内 `.agents/skills/`,并在各节点 `skills.json` artifact 中记录解析元数据。
|
|
50
|
-
|
|
51
|
-
执行前可用 `dag validate --strict-skills` 做 opt-in skill audit;该门禁会在 missing/error/truncated skill 或 unresolved reference 出现时失败。默认 role skill 应来自 `ai_workspace/loop-agent/.agents/skills/vetted-skill-registry.md` 中记录的 repo-local wrapper。
|
|
52
|
-
|
|
53
|
-
目标项目的 `loop-agent` skill 位于 `.agents/skills/loop-agent/SKILL.md`。loop-agent 源仓库和 npm 包内置版本仍位于 `.agents/skills/loop-agent/SKILL.md`;遗留根路径 `skill/SKILL.md` 仅为旧 worktree 保留兼容 fallback。
|
|
54
|
-
|
|
55
|
-
## Artifacts
|
|
56
|
-
|
|
57
|
-
DAG artifacts 位于:
|
|
58
|
-
|
|
59
|
-
```text
|
|
60
|
-
.harness/dag-runs/<state>/<run-id>/artifacts/<node-id>/
|
|
61
|
-
```
|
|
62
|
-
|
|
63
|
-
根目录 `artifacts/` 不是有效的默认 DAG artifact 位置。
|
|
64
|
-
|
|
65
|
-
## Shell Gates
|
|
66
|
-
|
|
67
|
-
- `shell.verdictGate` 从注入的当前 run 目录读取 `$HARNESS_DAG_RUN_DIR/<fromNodeId>.json`;不应自行发现 active run paths。
|
|
@@ -1,36 +0,0 @@
|
|
|
1
|
-
# cursor-prompt Sidecar
|
|
2
|
-
|
|
3
|
-
`cursor-prompt` 是显式、手工触发的 one-shot sidecar。它不是受治理 Agent runtime,也不参与 DAG、Loop 自动写入、Delegate `--auto-run` 或 task writer 选择。
|
|
4
|
-
|
|
5
|
-
## 产品定位
|
|
6
|
-
|
|
7
|
-
| 路径 | 角色 |
|
|
8
|
-
|---|---|
|
|
9
|
-
| Pi DAG (`implement-pi` / `repair-pi`) | 唯一受治理 Agent writer |
|
|
10
|
-
| shell / static | 确定性验证与静态输出 |
|
|
11
|
-
| `cursor-prompt` | 人工 one-shot 干预;成功不等于任务完成 |
|
|
12
|
-
|
|
13
|
-
## 用法
|
|
14
|
-
|
|
15
|
-
```bash
|
|
16
|
-
loop-agent cursor-prompt --cwd . "bounded task prompt"
|
|
17
|
-
loop-agent cursor-prompt --file <path>
|
|
18
|
-
loop-agent cursor-prompt --stdin
|
|
19
|
-
loop-agent cursor-prompt --model <id>
|
|
20
|
-
loop-agent cursor-prompt --timeout <ms>
|
|
21
|
-
loop-agent cursor-prompt --stream
|
|
22
|
-
loop-agent cursor-prompt --list-models
|
|
23
|
-
```
|
|
24
|
-
|
|
25
|
-
调用时才加载 `@cursor/sdk`。缺少 SDK 或 `CURSOR_API_KEY` 时,只有这条命令失败;普通 Agent DAG / doctor / init 不要求 Cursor。
|
|
26
|
-
|
|
27
|
-
## 约束
|
|
28
|
-
|
|
29
|
-
- 不读取 `harness.json` task config / DAG facts 作为授权来源。
|
|
30
|
-
- 不复制 DAG `writeSet`、repair、resume 或 Loop auto-execute 能力。
|
|
31
|
-
- 返回后由主会话检查 diff,并显式运行 shell verification。
|
|
32
|
-
- one-shot evidence 写入 `.harness/runs/{active,completed,failed}`。
|
|
33
|
-
|
|
34
|
-
## 迁移说明
|
|
35
|
-
|
|
36
|
-
旧 `executor: "cursor"` DAG、`executors.cursor`、`loopAutoWritePolicy` 与 `cursor-fix` 已硬切删除。需要写入时请重新生成 Pi-only DAG,或仅在人工干预场景使用本 sidecar。
|