@tea-agent/loop-agent 0.42.0-next.1 → 0.42.0-next.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +18 -0
- package/dist/build-stamp.json +3 -3
- package/dist/task/config-types.js +71 -0
- package/dist/workflows/dag/backend-test-case-coverage-analysis.js +99 -4
- package/dist/workflows/dag/backend-test-markdown-workflow.js +13 -1
- package/dist/workflows/dag/backend-test-writer-completeness.js +45 -3
- package/dist/workflows/dag/init-hybrid.js +101 -34
- package/dist/workflows/dag/node-execution.js +20 -0
- package/dist/workflows/dag/retry-policy.js +5 -0
- package/docs/templates/backend-test-dag.json +18 -18
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -3,6 +3,10 @@
|
|
|
3
3
|
## [Unreleased]
|
|
4
4
|
|
|
5
5
|
- Console「Pi 插件」新增升级功能:进入页面自动检查已安装插件的可用更新(从 npm 配置的仓库拉取每个已装包的 latest,与本地实际版本做语义化比较),检测到新版本时该插件行展示「升级 v<新版本>」按钮,一键升级到目标版本;工具栏提供「检查更新」手动重查入口,摘要行显示「已安装 N 个 · M 个可升级」;单包拉取失败不影响其它插件的判定,检查结果在服务端缓存 60 秒避免频繁刷新打爆 registry。
|
|
6
|
+
- backend-test 新增可选 strict `backendTest.moduleLayout`:可结构化冻结模块、业务资源、operation ownership 和不同 stem 的 Markdown/Python 精确路径,归一化合同 SHA 进入 Run-owned manifest;`output-budget` 拆分必须携带版本化数值 proof,纯理由字符串不再通过。
|
|
7
|
+
- 修复 backend-test Markdown 分片把同一接口的参数轴、组合、响应字段和回归类型误拆成多个物理模块的问题:Plan 现在优先保留主需求显式列出的 Markdown/Python 文件对,否则按最小稳定业务资源分组;Module Index 增加业务资源、operation/rule/Case ownership、拆分原因和精确路径,manifest 在 map 展开前拒绝测试目的型资源及无依据的同接口多模块,并可在权威合同无冲突时最多一次修复 Run-owned Plan 后再展开 map。
|
|
8
|
+
- 修复 backend-test Markdown writer 已明确写出 TP 分类语义但因圆括号、方括号或 pipe 标注不符合扫描格式而连续三次高 token 重试的问题:Completeness Gate 会在当前 child writeSet 内确定性补写 canonical 变体/断言/横切标签并复验;仍为单文件纯 binding-format 缺陷时最多一次 lightweight repair,不再重放第三次 full-context writer;否定、模糊、重复和未声明分类继续 fail-closed。
|
|
9
|
+
- 修复 backend-test 显式 Markdown/Python 路径使用不同 stem 时 Markdown validation 仍按 `test_<markdown-stem>.py` 推导默认脚本并误报的问题:validator 现在与 manifest/correspondence 一样读取冻结 Module Index 的 exact pytest path。
|
|
6
10
|
- 修复 DAG 节点「执行过程」导出 HTML 在内网反代/子路径部署下整页样式丢失的问题:导出不再写死从主机根抓 /inspect/styles.css(反代下 404 后静默回退兜底样式,按钮与卡片样式全丢),改为直接复用页面实际加载的样式表地址(相对当前页面解析,任何部署形态都拿到同一份 CSS,多张样式表合并);导出文档 body 补 data-shell 属性让内联页面 CSS 的设计令牌与前缀规则原生命中;兜底样式补充最小按钮样式。
|
|
7
11
|
- 修复 DAG 检查器节点「执行过程」导出 HTML 在系统深色模式下黑底黑字不可读的问题:导出文档自带完整的 `:root` 颜色变量定义(页面样式变量挂在 `body[data-shell]` 选择器下、独立导出文档不生效)并强制 light 配色,深色系统打开导出文件正文对比度恢复正常(15.4:1)。
|
|
8
12
|
- 修复「Pi 插件」安装时内网 Nexus 拉 scoped 包(如 `@0xkobold/pi-ollama`)版本列表失败的问题:包名整段编码不再把 `%2F` 还原成字面 `/`,与 npm 客户端请求形式一致,npmjs.org 与内网 Nexus 均正常返回。
|
|
@@ -104,6 +108,20 @@
|
|
|
104
108
|
- 右侧辅助栏统一承载文件、上下文、调用记录、运行活动和当前 DAG;点击消息流中的运行状态即可在同一侧栏查看运行、门禁与产物,并可自由调整侧栏宽度。
|
|
105
109
|
- 新增任务清单与结构化提问卡片,让长任务的进度和需要用户选择的内容更容易理解。
|
|
106
110
|
- 「对话 / 操作 / 观测」使用统一导航;操作和观测页面可以切换工作区查看对应数据。
|
|
111
|
+
## [0.41.0] - 2026-08-31
|
|
112
|
+
|
|
113
|
+
### 新增
|
|
114
|
+
|
|
115
|
+
- 新增了本发布周期内已经完成并通过验证的功能。
|
|
116
|
+
|
|
117
|
+
### 改进
|
|
118
|
+
|
|
119
|
+
- 完善了日常使用体验和运行稳定性。
|
|
120
|
+
|
|
121
|
+
### 修复
|
|
122
|
+
|
|
123
|
+
- 修复了影响使用体验或运行稳定性的问题。
|
|
124
|
+
|
|
107
125
|
## [0.40.0] - 2026-08-31
|
|
108
126
|
|
|
109
127
|
### 重点更新
|
package/dist/build-stamp.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"schemaVersion": 1,
|
|
3
|
-
"version": "0.42.0-next.
|
|
4
|
-
"gitSha": "
|
|
5
|
-
"builtAt": "2026-08-
|
|
3
|
+
"version": "0.42.0-next.2",
|
|
4
|
+
"gitSha": "464c462ece3604d8b6c7b44a7a915e5465bbc84f",
|
|
5
|
+
"builtAt": "2026-08-31T12:53:12.759Z"
|
|
6
6
|
}
|
|
@@ -131,6 +131,75 @@ export const frontendTestConfigSchema = z.object({
|
|
|
131
131
|
.strict()
|
|
132
132
|
.optional(),
|
|
133
133
|
});
|
|
134
|
+
const backendTestBudgetProofSchema = z
|
|
135
|
+
.object({
|
|
136
|
+
groupId: z.string().min(1),
|
|
137
|
+
estimatedCaseCount: z.number().int().positive(),
|
|
138
|
+
estimatedTestPointCount: z.number().int().positive(),
|
|
139
|
+
estimatedOutputChars: z.number().int().positive(),
|
|
140
|
+
maxOutputCharsPerWriter: z.number().int().positive(),
|
|
141
|
+
estimatorVersion: z.literal("backend-test-output-estimator-v1"),
|
|
142
|
+
inputSha256: z.string().regex(/^[a-f0-9]{64}$/),
|
|
143
|
+
})
|
|
144
|
+
.strict();
|
|
145
|
+
const backendTestModuleLayoutEntrySchema = z
|
|
146
|
+
.object({
|
|
147
|
+
stem: z.string().regex(/^[a-z][a-z0-9_]*$/),
|
|
148
|
+
businessResource: z.string().min(1),
|
|
149
|
+
markdownPath: z.string().min(1),
|
|
150
|
+
pytestPath: z.string().min(1),
|
|
151
|
+
ownedOperations: z.array(z.string().regex(/^(?:GET|POST|PUT|PATCH|DELETE|HEAD|OPTIONS) \/\S+$/)).min(1),
|
|
152
|
+
splitReason: z.enum([
|
|
153
|
+
"explicit-user-layout",
|
|
154
|
+
"primary-business-resource",
|
|
155
|
+
"independent-business-resource",
|
|
156
|
+
"output-budget",
|
|
157
|
+
]),
|
|
158
|
+
budgetProof: backendTestBudgetProofSchema.optional(),
|
|
159
|
+
})
|
|
160
|
+
.strict()
|
|
161
|
+
.superRefine((entry, ctx) => {
|
|
162
|
+
const safe = (value) => !value.startsWith("/") &&
|
|
163
|
+
!value.startsWith("\\") &&
|
|
164
|
+
!value.includes("..") &&
|
|
165
|
+
!value.includes("\\");
|
|
166
|
+
if (!safe(entry.markdownPath) || !entry.markdownPath.endsWith(".md")) {
|
|
167
|
+
ctx.addIssue({ code: z.ZodIssueCode.custom, path: ["markdownPath"], message: "must be a safe repository-relative .md path" });
|
|
168
|
+
}
|
|
169
|
+
if (!safe(entry.pytestPath) || !entry.pytestPath.endsWith(".py")) {
|
|
170
|
+
ctx.addIssue({ code: z.ZodIssueCode.custom, path: ["pytestPath"], message: "must be a safe repository-relative .py path" });
|
|
171
|
+
}
|
|
172
|
+
if (entry.splitReason === "output-budget" && !entry.budgetProof) {
|
|
173
|
+
ctx.addIssue({ code: z.ZodIssueCode.custom, path: ["budgetProof"], message: "is required for output-budget split" });
|
|
174
|
+
}
|
|
175
|
+
if (entry.splitReason !== "output-budget" && entry.budgetProof) {
|
|
176
|
+
ctx.addIssue({ code: z.ZodIssueCode.custom, path: ["budgetProof"], message: "is only valid for output-budget split" });
|
|
177
|
+
}
|
|
178
|
+
});
|
|
179
|
+
const backendTestModuleLayoutSchema = z
|
|
180
|
+
.object({
|
|
181
|
+
schemaId: z.literal("backend-test-module-layout-v1"),
|
|
182
|
+
mode: z.enum(["explicit-file-layout", "business-resource-layout"]),
|
|
183
|
+
modules: z.array(backendTestModuleLayoutEntrySchema).min(1).max(8),
|
|
184
|
+
})
|
|
185
|
+
.strict()
|
|
186
|
+
.superRefine((layout, ctx) => {
|
|
187
|
+
for (const key of ["stem", "markdownPath", "pytestPath"]) {
|
|
188
|
+
const seen = new Set();
|
|
189
|
+
for (let index = 0; index < layout.modules.length; index += 1) {
|
|
190
|
+
const value = layout.modules[index][key];
|
|
191
|
+
if (seen.has(value))
|
|
192
|
+
ctx.addIssue({ code: z.ZodIssueCode.custom, path: ["modules", index, key], message: `duplicate ${key}` });
|
|
193
|
+
seen.add(value);
|
|
194
|
+
}
|
|
195
|
+
}
|
|
196
|
+
if (layout.mode === "explicit-file-layout") {
|
|
197
|
+
layout.modules.forEach((entry, index) => {
|
|
198
|
+
if (entry.splitReason !== "explicit-user-layout")
|
|
199
|
+
ctx.addIssue({ code: z.ZodIssueCode.custom, path: ["modules", index, "splitReason"], message: "explicit-file-layout requires explicit-user-layout" });
|
|
200
|
+
});
|
|
201
|
+
}
|
|
202
|
+
});
|
|
134
203
|
/** Backend-test artifact layout and pipeline mode configuration (plan A/B/D). */
|
|
135
204
|
export const backendTestConfigSchema = z
|
|
136
205
|
.object({
|
|
@@ -152,6 +221,8 @@ export const backendTestConfigSchema = z
|
|
|
152
221
|
workspaceControl: z.enum(["git", "filesystem-only"]).optional(),
|
|
153
222
|
/** Per-command timeout for the readiness-authorized pytest/HTML execution node. */
|
|
154
223
|
executeTimeoutMs: z.number().int().min(60_000).max(7_200_000).optional(),
|
|
224
|
+
/** Optional strict physical/business module layout; authoritative over model-derived layout. */
|
|
225
|
+
moduleLayout: backendTestModuleLayoutSchema.optional(),
|
|
155
226
|
})
|
|
156
227
|
.strict()
|
|
157
228
|
.optional();
|
|
@@ -668,6 +668,92 @@ function automationLabelBody(body, labels) {
|
|
|
668
668
|
const match = new RegExp(`^\\s*[-*+]\\s*(?:${escaped})\\s*[::]\\s*(.*)$`, "mi").exec(automation);
|
|
669
669
|
return match?.[1]?.replaceAll("`", "").trim() ?? "";
|
|
670
670
|
}
|
|
671
|
+
const BINDING_NEGATION_OR_AMBIGUITY = /(?:不是|并非|非变体|非断言|非横切|可能|待定|或者|\b(?:TBD|unknown)\b)/i;
|
|
672
|
+
function explicitBindingModeFromLine(line) {
|
|
673
|
+
if (BINDING_NEGATION_OR_AMBIGUITY.test(line))
|
|
674
|
+
return undefined;
|
|
675
|
+
const modes = TEST_POINT_BINDING_MODES.filter((mode) => AUTOMATION_BINDING_LABELS[mode].some((label) => line.includes(label)));
|
|
676
|
+
return modes.length === 1 ? modes[0] : undefined;
|
|
677
|
+
}
|
|
678
|
+
/**
|
|
679
|
+
* Canonicalize only explicit, unambiguous Test Point classification marks.
|
|
680
|
+
* The model occasionally writes `TP-X(变体测试点)`, `[场景断言测试点] TP-X`
|
|
681
|
+
* or a pipe-delimited equivalent. Those forms carry deterministic semantics,
|
|
682
|
+
* but are not the machine contract consumed by automationLabelBody(). Keep the
|
|
683
|
+
* original prose and materialize the same binding under `### 自动化映射` as the
|
|
684
|
+
* canonical single-line labels. Never infer a mode from ordinary prose.
|
|
685
|
+
*/
|
|
686
|
+
export function canonicalizeBackendTestMarkdownTestPointBindings(markdown) {
|
|
687
|
+
const headings = [...markdown.matchAll(CASE_HEADING)];
|
|
688
|
+
let output = "";
|
|
689
|
+
let cursor = 0;
|
|
690
|
+
const unsafeTestPoints = [];
|
|
691
|
+
const canonicalizedTestPoints = [];
|
|
692
|
+
for (let index = 0; index < headings.length; index += 1) {
|
|
693
|
+
const heading = headings[index];
|
|
694
|
+
const start = heading.index;
|
|
695
|
+
const end = headings[index + 1]?.index ?? markdown.length;
|
|
696
|
+
const body = markdown.slice(start, end);
|
|
697
|
+
output += markdown.slice(cursor, start);
|
|
698
|
+
cursor = end;
|
|
699
|
+
const declared = listTokens(body, CASE_SECTION_ALIASES.testPoints, TEST_POINT);
|
|
700
|
+
const current = testPointBindingFacts(body, declared);
|
|
701
|
+
const currentByMode = new Map(TEST_POINT_BINDING_MODES.map((mode) => [
|
|
702
|
+
mode,
|
|
703
|
+
current.testPointBindings.filter((item) => item.mode === mode).map((item) => item.testPoint),
|
|
704
|
+
]));
|
|
705
|
+
const unresolved = new Set(current.unclassifiedTestPoints);
|
|
706
|
+
for (const line of body.split(/\r?\n/)) {
|
|
707
|
+
const mode = explicitBindingModeFromLine(line);
|
|
708
|
+
if (!mode)
|
|
709
|
+
continue;
|
|
710
|
+
const tokens = orderedUnique(line.match(/\bTP-[A-Z0-9]+(?:-[A-Z0-9]+)*\b/g) ?? []);
|
|
711
|
+
for (const token of tokens) {
|
|
712
|
+
if (!unresolved.has(token))
|
|
713
|
+
continue;
|
|
714
|
+
const existingModes = TEST_POINT_BINDING_MODES.filter((candidate) => (currentByMode.get(candidate) ?? []).includes(token));
|
|
715
|
+
if (existingModes.length > 0)
|
|
716
|
+
continue;
|
|
717
|
+
currentByMode.set(mode, orderedUnique([...(currentByMode.get(mode) ?? []), token]));
|
|
718
|
+
unresolved.delete(token);
|
|
719
|
+
canonicalizedTestPoints.push(token);
|
|
720
|
+
}
|
|
721
|
+
}
|
|
722
|
+
unsafeTestPoints.push(...unresolved);
|
|
723
|
+
if (unresolved.size === current.unclassifiedTestPoints.length) {
|
|
724
|
+
output += body;
|
|
725
|
+
continue;
|
|
726
|
+
}
|
|
727
|
+
const automationMarker = /^###\s+(?:自动化映射|Automation Mapping)\s*$/mi.exec(body);
|
|
728
|
+
if (!automationMarker) {
|
|
729
|
+
output += body;
|
|
730
|
+
continue;
|
|
731
|
+
}
|
|
732
|
+
const automationStart = automationMarker.index + automationMarker[0].length;
|
|
733
|
+
const rest = body.slice(automationStart);
|
|
734
|
+
const nextSection = /^###\s+/m.exec(rest);
|
|
735
|
+
const automationEnd = automationStart + (nextSection?.index ?? rest.length);
|
|
736
|
+
let automation = body.slice(automationStart, automationEnd);
|
|
737
|
+
for (const mode of TEST_POINT_BINDING_MODES) {
|
|
738
|
+
const labels = AUTOMATION_BINDING_LABELS[mode];
|
|
739
|
+
const escaped = labels.map((label) => label.replace(/[.*+?^${}()|[\]\\]/g, "\\$&")).join("|");
|
|
740
|
+
const canonicalLine = `- ${AUTOMATION_BINDING_LABELS[mode][0]}:${(currentByMode.get(mode) ?? []).join("; ") || "无"}`;
|
|
741
|
+
const linePattern = new RegExp(`^\\s*[-*+]\\s*(?:${escaped})\\s*[::].*$`, "mi");
|
|
742
|
+
if (linePattern.test(automation))
|
|
743
|
+
automation = automation.replace(linePattern, canonicalLine);
|
|
744
|
+
else
|
|
745
|
+
automation = `${automation.replace(/\s*$/, "")}\n${canonicalLine}\n`;
|
|
746
|
+
}
|
|
747
|
+
output += body.slice(0, automationStart) + automation + body.slice(automationEnd);
|
|
748
|
+
}
|
|
749
|
+
output += markdown.slice(cursor);
|
|
750
|
+
return {
|
|
751
|
+
markdown: output,
|
|
752
|
+
changed: output !== markdown,
|
|
753
|
+
unsafeTestPoints: orderedUnique(unsafeTestPoints),
|
|
754
|
+
canonicalizedTestPoints: orderedUnique(canonicalizedTestPoints),
|
|
755
|
+
};
|
|
756
|
+
}
|
|
671
757
|
export function inspectBackendTestMarkdownTestPointBindings(markdown) {
|
|
672
758
|
const headings = [...markdown.matchAll(CASE_HEADING)];
|
|
673
759
|
return headings.map((heading, index) => {
|
|
@@ -1312,8 +1398,16 @@ export async function analyzeBackendTestCaseCoverage(input) {
|
|
|
1312
1398
|
function normalizeModuleStem(markdownPath) {
|
|
1313
1399
|
return path.posix.basename(markdownPath).replace(/\.md$/i, "").toLowerCase().replace(/[^a-z0-9]+/g, "_").replace(/^_+|_+$/g, "").replace(/_+/g, "_");
|
|
1314
1400
|
}
|
|
1315
|
-
function expectedScript(markdownPath, layout) {
|
|
1401
|
+
function expectedScript(markdownPath, layout, declaredScripts = []) {
|
|
1316
1402
|
const resolved = layout ?? resolveBackendTestLayout(undefined);
|
|
1403
|
+
if (declaredScripts.length === 1) {
|
|
1404
|
+
const candidate = declaredScripts[0].replaceAll("\\", "/").replace(/^\.\//, "");
|
|
1405
|
+
const prefix = resolved.scriptDir === "." ? "" : `${resolved.scriptDir}/`;
|
|
1406
|
+
const relative = prefix && candidate.startsWith(prefix) ? candidate.slice(prefix.length) : candidate;
|
|
1407
|
+
if ((!prefix || candidate.startsWith(prefix)) && !relative.includes("..") && /^test_[A-Za-z0-9_.-]+\.py$/.test(path.posix.basename(candidate))) {
|
|
1408
|
+
return candidate;
|
|
1409
|
+
}
|
|
1410
|
+
}
|
|
1317
1411
|
return `${resolved.scriptDir}/test_${normalizeModuleStem(markdownPath)}.py`;
|
|
1318
1412
|
}
|
|
1319
1413
|
function symbolCaseId(symbol) {
|
|
@@ -1863,7 +1957,7 @@ export async function analyzeBackendTestMarkdownPytestCorrespondence(input) {
|
|
|
1863
1957
|
const layout = input.layout ?? resolveBackendTestLayout(undefined);
|
|
1864
1958
|
const cases = await parseMarkdownCases(input.workspaceRoot, layout);
|
|
1865
1959
|
const modulePaths = orderedUnique(cases.map((item) => item.markdownPath));
|
|
1866
|
-
const scripts = orderedUnique(cases.flatMap((item) => [
|
|
1960
|
+
const scripts = orderedUnique(cases.flatMap((item) => [expectedScript(item.markdownPath, layout, item.declaredScripts)]));
|
|
1867
1961
|
const allSymbols = [];
|
|
1868
1962
|
const scriptFiles = [];
|
|
1869
1963
|
for (const script of scripts) {
|
|
@@ -1876,7 +1970,7 @@ export async function analyzeBackendTestMarkdownPytestCorrespondence(input) {
|
|
|
1876
1970
|
const findings = [];
|
|
1877
1971
|
const entries = [];
|
|
1878
1972
|
for (const testCase of cases) {
|
|
1879
|
-
const expected = expectedScript(testCase.markdownPath, layout);
|
|
1973
|
+
const expected = expectedScript(testCase.markdownPath, layout, testCase.declaredScripts);
|
|
1880
1974
|
const declared = testCase.declaredScripts[0];
|
|
1881
1975
|
const refs = allSymbols.filter((item) => item.caseIds.includes(testCase.caseId));
|
|
1882
1976
|
const actualScripts = orderedUnique(refs.map((item) => item.script));
|
|
@@ -2034,8 +2128,9 @@ export async function analyzeBackendTestMarkdownPytestCorrespondence(input) {
|
|
|
2034
2128
|
for (const [signature, owners] of duplicateSignatureGroups)
|
|
2035
2129
|
findings.push(`duplicate execution signature ${signature}: ${orderedUnique(owners).join(", ")}`);
|
|
2036
2130
|
const exactModuleCount = modulePaths.filter((modulePath) => {
|
|
2037
|
-
const expected = expectedScript(modulePath, layout);
|
|
2038
2131
|
const moduleCases = cases.filter((item) => item.markdownPath === modulePath);
|
|
2132
|
+
const moduleDeclaredScripts = orderedUnique(moduleCases.flatMap((item) => item.declaredScripts));
|
|
2133
|
+
const expected = expectedScript(modulePath, layout, moduleDeclaredScripts);
|
|
2039
2134
|
return moduleCases.every((item) => item.declaredScripts.length === 1 && item.declaredScripts[0] === expected) && scriptFiles.some((file) => path.relative(input.workspaceRoot, file).replaceAll(path.sep, "/") === expected);
|
|
2040
2135
|
}).length;
|
|
2041
2136
|
const caseEntries = entries.filter((entry) => entry.cardinality !== "0:1");
|
|
@@ -285,6 +285,7 @@ export async function resolveBackendTestManifestModules(workspaceRoot, layout) {
|
|
|
285
285
|
moduleIndexHeadingCount,
|
|
286
286
|
declaredStems: [],
|
|
287
287
|
declaredMdFiles: [],
|
|
288
|
+
declaredScriptByMarkdown: {},
|
|
288
289
|
declaredMissing: [],
|
|
289
290
|
unmanifested: [],
|
|
290
291
|
});
|
|
@@ -306,6 +307,15 @@ export async function resolveBackendTestManifestModules(workspaceRoot, layout) {
|
|
|
306
307
|
return invalid("duplicate", moduleIndexHeadingCount);
|
|
307
308
|
}
|
|
308
309
|
const declaredStems = extractDeclaredModuleIndexFilenames(readme);
|
|
310
|
+
const declaredScriptByMarkdown = {};
|
|
311
|
+
const mdPrefix = escapeRegExp(`${resolved.markdownDir}/`);
|
|
312
|
+
const scriptPrefix = escapeRegExp(`${resolved.scriptDir}/`);
|
|
313
|
+
for (const line of readme.split(/\r?\n/)) {
|
|
314
|
+
const markdown = new RegExp(`(${mdPrefix}[A-Za-z0-9_.-]+\\.md)`).exec(line)?.[1];
|
|
315
|
+
const pytest = new RegExp(`(${scriptPrefix}[A-Za-z0-9_./-]+\\.py)`).exec(line)?.[1];
|
|
316
|
+
if (markdown && pytest)
|
|
317
|
+
declaredScriptByMarkdown[markdown] = pytest;
|
|
318
|
+
}
|
|
309
319
|
if (declaredStems.length === 0) {
|
|
310
320
|
return invalid("empty-unparsable", moduleIndexHeadingCount);
|
|
311
321
|
}
|
|
@@ -345,6 +355,7 @@ export async function resolveBackendTestManifestModules(workspaceRoot, layout) {
|
|
|
345
355
|
moduleIndexHeadingCount,
|
|
346
356
|
declaredStems,
|
|
347
357
|
declaredMdFiles: declaredMdFiles.sort(),
|
|
358
|
+
declaredScriptByMarkdown,
|
|
348
359
|
declaredMissing: declaredMissing.sort(),
|
|
349
360
|
unmanifested: unmanifested.sort(),
|
|
350
361
|
};
|
|
@@ -694,7 +705,8 @@ export async function validateBackendMarkdownCases(input) {
|
|
|
694
705
|
if (!hasAssertableExpectedResult(expected)) {
|
|
695
706
|
findings.push(`${testCase.id} has no structured expected result`);
|
|
696
707
|
}
|
|
697
|
-
const expectedScript =
|
|
708
|
+
const expectedScript = manifest.declaredScriptByMarkdown[relativeFile] ??
|
|
709
|
+
expectedBackendTestPytestScriptForMarkdownModule(relativeFile, layout);
|
|
698
710
|
const mappedScripts = extractMappedPytestScripts(testCase.body);
|
|
699
711
|
if (isEvidenceOnlyMarkdownCaseBody(testCase.body)) {
|
|
700
712
|
// Evidence-only meta cases intentionally declare 脚本/primary symbol as 无 and keep only assertion/cross-cutting TPs.
|
|
@@ -4,7 +4,7 @@ import path from "node:path";
|
|
|
4
4
|
import { z } from "zod";
|
|
5
5
|
import { BACKEND_TEST_MAX_MODULE_COUNT, isOpaqueHashBackendTestModuleStem, isPriorityOnlyBackendTestModuleStem, } from "./backend-test-module-stem.js";
|
|
6
6
|
import { resolveBackendTestLayout, } from "./backend-test-layout.js";
|
|
7
|
-
import { inspectBackendTestMarkdownTestPointBindings } from "./backend-test-case-coverage-analysis.js";
|
|
7
|
+
import { canonicalizeBackendTestMarkdownTestPointBindings, inspectBackendTestMarkdownTestPointBindings, } from "./backend-test-case-coverage-analysis.js";
|
|
8
8
|
/**
|
|
9
9
|
* Maps a backend-test generation writer task id to its writer-progress role.
|
|
10
10
|
* Returns undefined for nodes that are not completeness-gated generators.
|
|
@@ -47,8 +47,16 @@ const progressSchema = z
|
|
|
47
47
|
path: z.string().optional(),
|
|
48
48
|
detail: z.string(),
|
|
49
49
|
recoverable: z.boolean(),
|
|
50
|
+
failureSubreason: z.literal("markdown-test-point-binding-format").optional(),
|
|
51
|
+
affectedCases: z.array(z.string()).optional(),
|
|
52
|
+
testPointIds: z.array(z.string()).optional(),
|
|
53
|
+
canonicalizationEligible: z.boolean().optional(),
|
|
54
|
+
canonicalizationOutcome: z.enum(["not-attempted", "changed", "already-canonical", "unsafe", "revalidation-failed"]).optional(),
|
|
50
55
|
})
|
|
51
56
|
.strict()),
|
|
57
|
+
canonicalizationOutcome: z.enum(["not-attempted", "changed", "already-canonical", "unsafe", "revalidation-failed"]).optional(),
|
|
58
|
+
canonicalizedTestPointCount: z.number().int().min(0).optional(),
|
|
59
|
+
lightweightRepairEligible: z.boolean().optional(),
|
|
52
60
|
})
|
|
53
61
|
.strict();
|
|
54
62
|
async function exists(filePath) {
|
|
@@ -520,6 +528,9 @@ export async function assessBackendTestShardChildCompleteness(input) {
|
|
|
520
528
|
const actualPaths = [];
|
|
521
529
|
const missingPaths = [];
|
|
522
530
|
const brokenPaths = [];
|
|
531
|
+
let canonicalizationOutcome = role === "md-generate" ? "already-canonical" : undefined;
|
|
532
|
+
let canonicalizedTestPointCount = 0;
|
|
533
|
+
let lightweightRepairEligible = false;
|
|
523
534
|
for (const target of targetPaths) {
|
|
524
535
|
// Rendered map-child writeSet entries are concrete (e.g.
|
|
525
536
|
// testcase/md/health.md / testcase/test_orders.py); glob-only entries are
|
|
@@ -538,16 +549,44 @@ export async function assessBackendTestShardChildCompleteness(input) {
|
|
|
538
549
|
continue;
|
|
539
550
|
}
|
|
540
551
|
actualPaths.push(target);
|
|
541
|
-
|
|
552
|
+
let body = await readFile(absolute, "utf8");
|
|
542
553
|
if (role === "md-generate") {
|
|
554
|
+
const beforeBindings = inspectBackendTestMarkdownTestPointBindings(body);
|
|
555
|
+
const canonicalization = canonicalizeBackendTestMarkdownTestPointBindings(body);
|
|
556
|
+
if (canonicalization.changed) {
|
|
557
|
+
await writeFile(absolute, canonicalization.markdown, "utf8");
|
|
558
|
+
body = canonicalization.markdown;
|
|
559
|
+
canonicalizationOutcome = "changed";
|
|
560
|
+
canonicalizedTestPointCount += canonicalization.canonicalizedTestPoints.length;
|
|
561
|
+
}
|
|
562
|
+
const afterBindings = inspectBackendTestMarkdownTestPointBindings(body);
|
|
563
|
+
const unresolved = afterBindings.flatMap((entry) => entry.unclassifiedTestPoints);
|
|
564
|
+
const bindingCases = afterBindings
|
|
565
|
+
.filter((entry) => entry.unclassifiedTestPoints.length > 0 || entry.duplicateBindingTestPoints.length > 0)
|
|
566
|
+
.map((entry) => entry.caseId);
|
|
543
567
|
const structural = moduleStructurallyComplete(body);
|
|
544
568
|
if (!structural.ok) {
|
|
569
|
+
lightweightRepairEligible = bindingCases.length > 0;
|
|
570
|
+
if (bindingCases.length > 0)
|
|
571
|
+
canonicalizationOutcome = canonicalization.changed ? "revalidation-failed" : "unsafe";
|
|
545
572
|
brokenPaths.push(target);
|
|
573
|
+
const bindingFormatIssue = bindingCases.length > 0;
|
|
546
574
|
issues.push({
|
|
547
575
|
code: "T5",
|
|
548
576
|
path: target,
|
|
549
577
|
detail: moduleStructuralDetail(structural),
|
|
550
578
|
recoverable: true,
|
|
579
|
+
...(bindingFormatIssue
|
|
580
|
+
? {
|
|
581
|
+
failureSubreason: "markdown-test-point-binding-format",
|
|
582
|
+
affectedCases: orderedUnique(bindingCases),
|
|
583
|
+
testPointIds: orderedUnique(unresolved),
|
|
584
|
+
canonicalizationEligible: canonicalization.canonicalizedTestPoints.length > 0,
|
|
585
|
+
canonicalizationOutcome: canonicalization.changed
|
|
586
|
+
? (unresolved.length === 0 ? "changed" : "revalidation-failed")
|
|
587
|
+
: (beforeBindings.some((entry) => entry.unclassifiedTestPoints.length > 0) ? "unsafe" : "already-canonical"),
|
|
588
|
+
}
|
|
589
|
+
: {}),
|
|
551
590
|
});
|
|
552
591
|
}
|
|
553
592
|
}
|
|
@@ -586,6 +625,9 @@ export async function assessBackendTestShardChildCompleteness(input) {
|
|
|
586
625
|
brokenPaths: orderedUnique(brokenPaths),
|
|
587
626
|
targetPaths: recoverableTargets,
|
|
588
627
|
issues,
|
|
628
|
+
...(role === "md-generate"
|
|
629
|
+
? { canonicalizationOutcome, canonicalizedTestPointCount, lightweightRepairEligible }
|
|
630
|
+
: {}),
|
|
589
631
|
};
|
|
590
632
|
}
|
|
591
633
|
export async function assessBackendTestMdWriterCompleteness(workspaceRoot, layout) {
|
|
@@ -853,7 +895,7 @@ export async function writeBackendTestWriterProgressArtifacts(input) {
|
|
|
853
895
|
"## Issues",
|
|
854
896
|
"",
|
|
855
897
|
...(input.progress.issues.length
|
|
856
|
-
? input.progress.issues.map((issue) => `- [${issue.code}] ${issue.path ?? "(workspace)"}: ${issue.detail}`)
|
|
898
|
+
? input.progress.issues.map((issue) => `- [${issue.code}] ${issue.path ?? "(workspace)"}: ${issue.detail}${issue.failureSubreason ? `; failureSubreason=${issue.failureSubreason}; canonicalizationOutcome=${issue.canonicalizationOutcome ?? "not-attempted"}; canonicalizationEligible=${issue.canonicalizationEligible ?? false}; affectedCases=${(issue.affectedCases ?? []).join(",") || "none"}; testPointIds=${(issue.testPointIds ?? []).join(",") || "none"}` : ""}`)
|
|
857
899
|
: ["- None"]),
|
|
858
900
|
"",
|
|
859
901
|
`<!-- sha256:${createHash("sha256").update(JSON.stringify(parsed)).digest("hex")} -->`,
|
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
import { createHash } from "node:crypto";
|
|
2
2
|
import { access, readdir, readFile, realpath } from "node:fs/promises";
|
|
3
3
|
import { existsSync, readFileSync } from "node:fs";
|
|
4
|
+
import { deflateRawSync } from "node:zlib";
|
|
4
5
|
import path from "node:path";
|
|
5
6
|
import { writeJsonAtomic } from "../../infrastructure/harness/atomic-write.js";
|
|
6
7
|
import { assertValidDagSpec } from "./validate.js";
|
|
@@ -11,7 +12,7 @@ import { pathMatchesPattern } from "../../shared/git-progress.js";
|
|
|
11
12
|
import { extractTaskSourceOpenspecPaths } from "../../shared/openspec-spec.js";
|
|
12
13
|
import { BASELINE_FORBIDDEN_PATHS } from "./governance-constants.js";
|
|
13
14
|
import { buildDecisionEnvelopePromptContract } from "./decision-envelope.js";
|
|
14
|
-
import { DEFAULT_READ_ONLY_PI_RETRY_POLICY, PLANNER_OUTPUT_LIMIT_RETRY_POLICY, PROTOCOL_AWARE_PI_RETRY_POLICY, STRUCTURED_REQUIRED_PI_RETRY_POLICY, BACKEND_TEST_WRITER_COMPLETENESS_RETRY_POLICY, WRITER_TRANSPORT_RETRY_POLICY, TARGET_TEMPLATE_TRANSIENT_RETRY_PROFILE, TARGET_TEMPLATE_WRITER_TRANSPORT_RETRY_POLICY, isCanonicalFinalVerifyShellRetryCandidate, isSafeReadOnlyPiRetryCandidate, isTargetTemplateImplementPi, isWriterTransportRetryCandidate, } from "./retry-policy.js";
|
|
15
|
+
import { DEFAULT_READ_ONLY_PI_RETRY_POLICY, PLANNER_OUTPUT_LIMIT_RETRY_POLICY, PROTOCOL_AWARE_PI_RETRY_POLICY, STRUCTURED_REQUIRED_PI_RETRY_POLICY, BACKEND_TEST_WRITER_COMPLETENESS_RETRY_POLICY, BACKEND_TEST_MARKDOWN_BINDING_RETRY_POLICY, WRITER_TRANSPORT_RETRY_POLICY, TARGET_TEMPLATE_TRANSIENT_RETRY_PROFILE, TARGET_TEMPLATE_WRITER_TRANSPORT_RETRY_POLICY, isCanonicalFinalVerifyShellRetryCandidate, isSafeReadOnlyPiRetryCandidate, isTargetTemplateImplementPi, isWriterTransportRetryCandidate, } from "./retry-policy.js";
|
|
15
16
|
import { REVIEW_JSON_VERDICT_OUTPUT_PROTOCOL, REVIEW_VERDICT_OUTPUT_PROTOCOL, } from "./output-protocol.js";
|
|
16
17
|
import { resolveAdapter } from "../../adapters/index.js";
|
|
17
18
|
import { loadHarnessManifest } from "../../governance/harness.js";
|
|
@@ -4017,7 +4018,7 @@ export function applyBackendTestLayoutToText(text, layout) {
|
|
|
4017
4018
|
* normalization. Output is exactly one trailing JSON line `{modules:[{stem}]}`
|
|
4018
4019
|
* that `parseJsonFromText` accepts after shell command echoes.
|
|
4019
4020
|
*/
|
|
4020
|
-
function buildBackendTestModuleManifestShellCommand(layout) {
|
|
4021
|
+
function buildBackendTestModuleManifestShellCommand(layout, moduleLayout) {
|
|
4021
4022
|
// The extractor is base64-encoded so the shell command is fully opaque to
|
|
4022
4023
|
// bash: no backticks (command substitution), no regex \/ escaping, no
|
|
4023
4024
|
// backslash-counting through TS-string -> JSON.stringify -> bash -c -> node -e.
|
|
@@ -4026,18 +4027,31 @@ function buildBackendTestModuleManifestShellCommand(layout) {
|
|
|
4026
4027
|
// mdDir/testPrefix are injected as JSON literals so the same extractor
|
|
4027
4028
|
// works for any configured backendTest layout (plan A).
|
|
4028
4029
|
const mdDirLiteral = JSON.stringify(layout.markdownDir);
|
|
4030
|
+
const scriptDirLiteral = JSON.stringify(layout.scriptDir);
|
|
4031
|
+
const strictLayoutCompressed = moduleLayout
|
|
4032
|
+
? deflateRawSync(Buffer.from(JSON.stringify(moduleLayout), "utf8")).toString("base64")
|
|
4033
|
+
: null;
|
|
4034
|
+
const strictLayoutExpression = strictLayoutCompressed
|
|
4035
|
+
? `JSON.parse(zlib.inflateRawSync(Buffer.from(${JSON.stringify(strictLayoutCompressed)},'base64')).toString('utf8'))`
|
|
4036
|
+
: "null";
|
|
4037
|
+
const strictLayoutShaLiteral = JSON.stringify(moduleLayout
|
|
4038
|
+
? createHash("sha256").update(JSON.stringify(moduleLayout)).digest("hex")
|
|
4039
|
+
: null);
|
|
4029
4040
|
const escOpen = String.fromCharCode(92, 91); // \[
|
|
4030
4041
|
const escClose = String.fromCharCode(92, 93); // \]
|
|
4031
4042
|
const escBslash = String.fromCharCode(92, 92); // \\
|
|
4032
|
-
const script = `const fs=require('fs'),path=require('path'),crypto=require('crypto');
|
|
4043
|
+
const script = `const fs=require('fs'),path=require('path'),crypto=require('crypto'),zlib=require('zlib');
|
|
4033
4044
|
const mdDir=${mdDirLiteral};
|
|
4045
|
+
const scriptDir=${scriptDirLiteral};
|
|
4046
|
+
const strictLayout=${strictLayoutExpression};
|
|
4047
|
+
const strictLayoutSha256=${strictLayoutShaLiteral};
|
|
4034
4048
|
const esc=s=>s.replace(/[${escOpen}${escClose}{}()*+?^$|${escBslash}]/g,'${escBslash}$&');
|
|
4035
4049
|
const rxMdPath=new RegExp(esc(mdDir)+'${escBslash}/([A-Za-z0-9_.-]+)${escBslash}.md','g');
|
|
4036
4050
|
const runDir=process.env.HARNESS_DAG_RUN_DIR||'';
|
|
4037
4051
|
if(!runDir){process.stderr.write('missing HARNESS_DAG_RUN_DIR for backend-test Markdown plan artifact\\n');process.exit(2);}
|
|
4038
4052
|
const planPath=path.join(runDir,'generate-backend-md-plan-pi','plan.md');
|
|
4039
4053
|
if(!fs.existsSync(planPath)){process.stderr.write('missing backend-test Markdown plan artifact: '+planPath+'\\n');process.exit(2);}
|
|
4040
|
-
|
|
4054
|
+
let readme=fs.readFileSync(planPath,'utf8');
|
|
4041
4055
|
const norm=s=>String(s).toLowerCase().replace(/[^a-z0-9]+/g,'_').replace(/^_+|_+$/g,'').replace(/_+/g,'_');
|
|
4042
4056
|
const bt=String.fromCharCode(96);
|
|
4043
4057
|
const stripBackticks=s=>s.split(bt).join('');
|
|
@@ -4051,23 +4065,68 @@ const start=headings[0]+1;let end=allLines.length;for(let i=start;i<allLines.len
|
|
|
4051
4065
|
const section=allLines.slice(start,end).join('\\n');
|
|
4052
4066
|
const raw=[];
|
|
4053
4067
|
const lines=section.split('\\n').filter(l=>l.includes('|'));
|
|
4054
|
-
|
|
4055
|
-
|
|
4056
|
-
|
|
4068
|
+
const cells=line=>line.split('|').slice(1,-1).map(value=>stripBackticks(value).trim());
|
|
4069
|
+
const expectedHeader=['Module Stem','Business Resource','Owned Operations','Owned Rule Keys','Case IDs','Split Reason','Markdown Path','Pytest Path'];
|
|
4070
|
+
const headerIndex=lines.findIndex(line=>{const row=cells(line);return expectedHeader.every((value,index)=>row[index]===value);});
|
|
4071
|
+
if(headerIndex<0){process.stderr.write('invalid-module-index-header: require exact business ownership and path columns\\n');process.exit(2);}
|
|
4072
|
+
const dataRows=lines.slice(headerIndex+2).map(cells).filter(row=>row.length>=8&&row[0]&&row[0]!=='Module Stem');
|
|
4073
|
+
const allowedSplit=new Set(['explicit-user-layout','primary-business-resource','independent-business-resource','output-budget']);
|
|
4074
|
+
const strictExpectedStems=new Set((strictLayout&&strictLayout.modules||[]).map(item=>item.stem));
|
|
4075
|
+
const operationOwners=new Map();const declaredModules=[];let planRepairApplied=false;
|
|
4076
|
+
for(const row of dataRows){
|
|
4077
|
+
const stem=norm(row[0]),resource=String(row[1]||'').trim(),operations=String(row[2]||'').split(';').map(value=>value.trim()).filter(Boolean),split=String(row[5]||'').trim();
|
|
4078
|
+
if(strictLayout&&!strictExpectedStems.has(stem)){planRepairApplied=true;continue;}
|
|
4079
|
+
if(!resource){process.stderr.write('missing-business-resource: '+stem+'\\n');process.exit(2);}
|
|
4080
|
+
if(/^(?:response|resp|regression|positive|negative|boundary|error|combo|filter)(?:[_ -]|$)/i.test(resource)){process.stderr.write('test-purpose-business-resource: '+stem+' -> '+resource+'\\n');process.exit(2);}
|
|
4081
|
+
if(!allowedSplit.has(split)){process.stderr.write('invalid-module-split-reason: '+stem+' -> '+split+'\\n');process.exit(2);}
|
|
4082
|
+
if(operations.length===0){process.stderr.write('missing-owned-operation: '+stem+'\\n');process.exit(2);}
|
|
4083
|
+
const mdMatches=[...String(row[6]||'').matchAll(rxMdPath)];let markdownPath=mdMatches[0]&&mdMatches[0][0];
|
|
4084
|
+
let pytestPath=String(row[7]||'').replace(/^\\.\\//,'').replace(/\\\\/g,'/');
|
|
4085
|
+
if(mdMatches.length!==1||norm(mdMatches[0][1])!==stem){process.stderr.write('module-markdown-path-mismatch: '+stem+'\\n');process.exit(2);}
|
|
4086
|
+
if(!/^[A-Za-z0-9_./-]+\\.py$/.test(pytestPath)||pytestPath.includes('..')){process.stderr.write('invalid-module-pytest-path: '+stem+' -> '+pytestPath+'\\n');process.exit(2);}
|
|
4087
|
+
const required=strictLayout&&strictLayout.modules.find(item=>item.stem===stem);
|
|
4088
|
+
if(required){
|
|
4089
|
+
if(!required.markdownPath.startsWith(mdDir+'/')||!required.pytestPath.startsWith(scriptDir+'/')||required.markdownPath.includes('..')||required.pytestPath.includes('..')){process.stderr.write('MODULE_LAYOUT_CONFLICT: strict paths outside frozen roots for '+stem+'\\n');process.exit(2);}
|
|
4090
|
+
if(markdownPath!==required.markdownPath||pytestPath!==required.pytestPath||resource!==required.businessResource||JSON.stringify(operations)!==JSON.stringify(required.ownedOperations)||split!==required.splitReason){planRepairApplied=true;}
|
|
4091
|
+
markdownPath=required.markdownPath;pytestPath=required.pytestPath;
|
|
4092
|
+
if(required.splitReason==='output-budget'){
|
|
4093
|
+
const proof=required.budgetProof;const peers=strictLayout.modules.filter(item=>item.budgetProof&&item.budgetProof.groupId===proof.groupId);
|
|
4094
|
+
if(!proof||proof.estimatedOutputChars<=proof.maxOutputCharsPerWriter||Math.ceil(proof.estimatedOutputChars/Math.max(1,peers.length))>proof.maxOutputCharsPerWriter||peers.some(item=>item.budgetProof.inputSha256!==proof.inputSha256||item.budgetProof.estimatorVersion!==proof.estimatorVersion)){process.stderr.write('INVALID_OUTPUT_BUDGET_PROOF: '+stem+'\\n');process.exit(2);}
|
|
4095
|
+
}
|
|
4096
|
+
}
|
|
4097
|
+
const item={stem,markdownPath,pytestPath,businessResource:required?required.businessResource:resource,ownedOperations:required?required.ownedOperations:operations,ownedRuleKeys:String(row[3]||'').split(';').map(value=>value.trim()).filter(Boolean),caseIds:String(row[4]||'').split(';').map(value=>value.trim()).filter(Boolean),splitReason:required?required.splitReason:split,...(required&&required.budgetProof?{budgetProof:required.budgetProof}:{})};
|
|
4098
|
+
declaredModules.push(item);
|
|
4099
|
+
for(const operation of item.ownedOperations){const owners=operationOwners.get(operation)||[];owners.push({stem,split:item.splitReason});operationOwners.set(operation,owners);}
|
|
4100
|
+
}
|
|
4101
|
+
if(strictLayout){
|
|
4102
|
+
const missing=strictLayout.modules.filter(item=>!declaredModules.some(actual=>actual.stem===item.stem));
|
|
4103
|
+
if(missing.length){process.stderr.write('MODULE_LAYOUT_CONFLICT: plan missing required modules '+missing.map(item=>item.stem).join(',')+'; deterministic Plan-only repair cannot invent Rule/Case ownership\\n');process.exit(2);}
|
|
4104
|
+
if(planRepairApplied){
|
|
4105
|
+
const table=['| '+expectedHeader.join(' | ')+' |','|'+expectedHeader.map(()=>'---').join('|')+'|',...declaredModules.map(item=>'| '+[item.stem,item.businessResource,item.ownedOperations.join('; '),item.ownedRuleKeys.join('; '),item.caseIds.join('; '),item.splitReason,'['+item.stem+'](./'+item.stem+'.md) '+item.markdownPath,item.pytestPath].join(' | ')+' |')].join('\\n');
|
|
4106
|
+
readme=[...allLines.slice(0,headings[0]+1),table,...allLines.slice(end)].join('\\n');
|
|
4107
|
+
fs.writeFileSync(planPath,readme,'utf8');
|
|
4108
|
+
}
|
|
4109
|
+
}
|
|
4110
|
+
for(const [operation,owners] of operationOwners){if(owners.length>1&&!owners.every(owner=>owner.split==='explicit-user-layout'||owner.split==='output-budget')){process.stderr.write('overlapping-operation-modules: '+operation+' -> '+owners.map(owner=>owner.stem).join(',')+'; merge by business resource or use an authoritative explicit-user-layout\\n');process.exit(2);}}
|
|
4111
|
+
if(!strictLayout){
|
|
4112
|
+
for(const line of lines){
|
|
4113
|
+
const bare=stripBackticks(line);
|
|
4114
|
+
for(const m of bare.matchAll(rxMdPath)){raw.push(m[1]);}
|
|
4115
|
+
}
|
|
4116
|
+
for(const m of section.matchAll(rxRelLink)){raw.push(m[1]);}
|
|
4057
4117
|
}
|
|
4058
|
-
for(const m of section.matchAll(rxRelLink)){raw.push(m[1]);}
|
|
4059
4118
|
const invalid=[];for(const r of raw){const reason=invalidReason(r);if(reason)invalid.push({stem:norm(r),reason});}
|
|
4060
4119
|
if(invalid.length){for(const item of invalid)process.stderr.write(item.reason+': '+item.stem+'; use a stable business resource/domain stem\\n');process.exit(2);}
|
|
4061
4120
|
const seen=new Set();const modules=[];
|
|
4062
|
-
for(const
|
|
4121
|
+
for(const item of declaredModules){if(valid(item.stem)&&!seen.has(item.stem)){seen.add(item.stem);modules.push(item);}}
|
|
4063
4122
|
if(modules.length===0){process.stderr.write('empty-module-index: require at least one stable business module\\n');process.exit(2);}
|
|
4064
4123
|
if(modules.length>8){process.stderr.write('excessive-module-count: '+modules.length+' > 8; merge by the smallest stable business resource/domain set\\n');process.exit(2);}
|
|
4065
4124
|
const planReadPath=path.relative(process.cwd(),planPath).split(path.sep).join('/');
|
|
4066
4125
|
const planSha256=crypto.createHash('sha256').update(readme).digest('hex');
|
|
4067
|
-
process.stdout.write(JSON.stringify({modules:modules.map(item=>({...item,planReadPath})),planReadPath,planSha256}));
|
|
4126
|
+
process.stdout.write(JSON.stringify({modules:modules.map(item=>({...item,planReadPath})),planReadPath,planSha256,moduleLayout:{schemaId:'backend-test-module-layout-facts-v1',source:strictLayout?'task-contract':'plan-derived',contractSha256:strictLayoutSha256,mode:strictLayout&&strictLayout.mode||'business-resource-layout',planRepairAttemptCount:planRepairApplied?1:0,planRepairOutcome:planRepairApplied?'changed':'not-required'}}));
|
|
4068
4127
|
`;
|
|
4069
|
-
const encoded = Buffer.from(script, "utf8").toString("base64");
|
|
4070
|
-
return `node -e "eval(Buffer.from('${encoded}','base64').toString('utf8'))"`;
|
|
4128
|
+
const encoded = deflateRawSync(Buffer.from(script, "utf8")).toString("base64");
|
|
4129
|
+
return `node -e "eval(require('zlib').inflateRawSync(Buffer.from('${encoded}','base64')).toString('utf8'))"`;
|
|
4071
4130
|
}
|
|
4072
4131
|
const BACKEND_TEST_SKILLS_BY_ROLE = {
|
|
4073
4132
|
planner: ["loop-agent"],
|
|
@@ -4655,7 +4714,14 @@ async function buildBackendTestHybridDag(sources) {
|
|
|
4655
4714
|
"Each Rule Key must appear in exactly one Matrix row. Preserve each AC/REQ/BR Rule Key as one row; if one product rule spans multiple dimensions, use a concise composite Dimension in that single row instead of duplicating the key. Derive OpenAPI Rule Keys exactly as the deterministic analyzer does: operation token is `<HTTP-METHOD>-<PATH>` with braces removed and every non-alphanumeric run replaced by a hyphen, uppercase (for example POST `/api/resource-notes` → `POST-API-RESOURCE-NOTES`); response statuses use `API-<OPERATION>-RESPONSE-STATUS`; body/parameter fields use `API-<OPERATION>-<FIELD>-REQUIRED|ENUM|MIN-LENGTH|MAX-LENGTH|MINIMUM|MAXIMUM|PATTERN|FORMAT`. Do not invent aliases such as API-CREATE-FIELDS when a deterministic key applies.",
|
|
4656
4715
|
"Coverage priority is strict inside the declared scope: P0 product requirements/task hard constraints always remain in scope; P1 exhaustively supplements documented operations, fields, business rules, statuses and errors only for Affected Operations; P2 adds bounded protocol robustness only when it is relevant to the change and does not invent product behavior. Coverage percentages describe the declared affected scope, never whole-API completeness unless every operation is explicitly listed. Conflicts or undefined expectations must stay visible as GAP/CONFLICT with precise source pointers, never guessed.",
|
|
4657
4716
|
"For uniqueness/lifecycle rules cover absent, active-existing, deleted-existing, create-delete-recreate, restore-then-recreate and documented scope/case-normalization states. For every enum cover every valid value plus bounded invalid equivalence classes (unknown, case variant, whitespace, empty, null/missing and wrong types as applicable). For every length/number rule cover min-1, min, nominal, max and max+1. For format rules cover each allowed class separately plus a valid mixed value, and representative forbidden classes including uppercase, internal/leading/trailing whitespace, tab/newline, unsupported punctuation, slash, emoji or control characters when the source contract supports that expectation.",
|
|
4658
|
-
"Mandatory module
|
|
4717
|
+
"Mandatory module layout contract: first inspect the PRIMARY requirement for an explicit list of required Markdown/Python output path pairs. When explicit paths are present, they are authoritative `explicit-user-layout`: reproduce their exact filenames, count and one-to-one pairs in Module Index; do not rename, merge, split, omit or add a module from reference/example scripts. Only when the primary requirement has no explicit file layout may you derive the smallest `business-resource-layout`. Existing examples, historical regression functions and shared setup may add evidence/assertions to an existing required module, but never create an extra physical module by themselves. A reference-only `resp_regression`, positive/negative/boundary/error/response module is forbidden.",
|
|
4718
|
+
...(taskConfig.backendTest?.moduleLayout
|
|
4719
|
+
? [
|
|
4720
|
+
"A strict task-contract `backendTest.moduleLayout` is bound and is authoritative over model-derived layout. Reproduce every stem, businessResource, ownedOperations, splitReason, markdownPath and pytestPath exactly; do not add, omit, rename or reorder physical modules. The downstream preflight compares exact path sets and may perform at most one deterministic Plan-only pruning/path normalization; it cannot invent missing Rule/Case ownership.",
|
|
4721
|
+
`STRICT_BACKEND_TEST_MODULE_LAYOUT=${JSON.stringify(taskConfig.backendTest.moduleLayout)}`,
|
|
4722
|
+
]
|
|
4723
|
+
: []),
|
|
4724
|
+
"Include exactly one `## Module Index` table with this exact header: `| Module Stem | Business Resource | Owned Operations | Owned Rule Keys | Case IDs | Split Reason | Markdown Path | Pytest Path |`. Use canonical relative links `[label](./<stem>.md)` inside the Markdown Path cell followed by the resolved `${layout.markdownDir}/<stem>.md` path. Split Reason is exactly one of `explicit-user-layout`, `primary-business-resource`, `independent-business-resource`, or `output-budget`. Group by stable business resource/domain, not by CRUD operation, AC, parameter/field axis, scenario type or regression purpose: one resource's list/detail/create/update/delete and its filters/response assertions/regression floor belong in one module. Multiple modules owning the same exact `METHOD /path` are forbidden unless every such row is `explicit-user-layout` from primary-requirement path pairs or has a documented `output-budget` proof. Keep the total module count at the smallest safe value and never exceed 8 modules. Name model-derived modules with stable lowercase business stems such as `health` or `resource_notes`; explicit-user-layout preserves the primary requirement filename stem even when it is more specific. Do not use priority-only stems `p0`, `p1` or `p2`; Priority belongs only in the Coverage Matrix. Pure hexadecimal/hash-like opaque stems and test-purpose-only stems are forbidden. Do not use Case-ID-like module filenames. The relative link target, Markdown Path, Pytest Path and downstream automation mapping must be one-to-one and exact; for model-derived modules the default pair remains `testcase/md/<module>.md` and `testcase/test_<module>.py`, while explicit-user-layout preserves the primary requirement paths. Do not hand-write a conflicting module count in prose; the Module Index row count is the only count truth.",
|
|
4659
4725
|
"Scenario Partitions (query/filter axes): for every affected GET/list operation, declare one row per enum or classification axis used for filtering (query/path parameters such as type/status/category). Add a mandatory machine-readable `## Scenario Partitions` section after the Coverage Matrix using exactly `| Partition ID | Operation | Axis | Domain | Required Slots | Expected by Slot | Bind Rule |` with the separator row. Partition ID is a stable `SP-<OPERATION>-<AXIS>` token; Domain must copy the legal values verbatim from the bound OpenAPI enum or requirement sentence (never guess), using bare semicolon-separated identifier values inside the single table cell (for example `ACTIVE; ARCHIVED`, with no Markdown backticks or prose); Required Slots must contain `each-value` and exactly one `not-in-set`, plus `omitted` only when the parameter is optional. Before returning, expand every declared partition into its complete deterministic exact slot ID set: one `TP-<Partition ID>-<VALUE-TOKEN>` per Domain value, `TP-<Partition ID>-OMITTED` only for an optional axis, and exactly one `TP-<Partition ID>-NOT-IN-SET`. Every expanded slot ID must appear verbatim in the binding Rule's `Required Test Points` cell and be assigned to concrete Case IDs in that same Coverage Matrix row; ordinary alias/family Test Points do not replace this inventory. Expected by Slot states the documented expectation per slot kind (`domain-value`, `default-behavior`, `empty-result`/`excluded-result` when documented, or `GAP` when the source does not document the complement expectation — never invent 空列表/400). POST/PUT body field-validation enums stay in the Coverage Matrix as `TP-<FIELD>-ENUM-*` and MUST NOT get a Scenario Partition row. Do not create partitions for axes without a documented legal-value domain. Only GET/list query or path parameters whose bound source documents a finite enum or classification set may become a Scenario Partition. Do not create partitions for free-form strings, primary keys, required-or-optional-only parameters, or boundary/format-only axes. If an axis has no finite legal-value domain, do not declare a Partition row and do not invent NOT-IN-SET cases. Cross-axis combinations stay as ONE nominal Case; never declare a cross-axis cartesian partition.",
|
|
4660
4726
|
"Before finalizing README, calculate the predicted collected-item count as `sum(max(1, number of variant Test Points in each Case))`. If the task declares an item budget, the prediction must not exceed it. Reduce excess only by removing duplicate execution and converting same-request checkpoints to assertions; never drop required rules, boundaries, enums, operation-specific inputs, or business states. Record the prediction in README. Use only environment-supported fixtures/targets/isolation, record evidence gaps in Chinese, and do not emit JSON, pytest, or execute commands.",
|
|
4661
4727
|
...(sharedSetupPrompt ? [sharedSetupPrompt] : []),
|
|
@@ -4675,10 +4741,10 @@ async function buildBackendTestHybridDag(sources) {
|
|
|
4675
4741
|
writePolicy: "read-only",
|
|
4676
4742
|
allowedPaths: ro,
|
|
4677
4743
|
forbiddenPaths: forbidden,
|
|
4678
|
-
outputContract: "Stdout JSON {modules:[{stem,planReadPath}],planReadPath,planSha256} parsed from the run-owned generate-backend-md-plan-pi/plan.md artifact
|
|
4679
|
-
subtask_prompt: "Parse only $HARNESS_DAG_RUN_DIR/generate-backend-md-plan-pi/plan.md and
|
|
4744
|
+
outputContract: "Stdout JSON {modules:[{stem,planReadPath}],planReadPath,planSha256,moduleLayout} parsed from the run-owned generate-backend-md-plan-pi/plan.md artifact after strict-layout validation and at most one deterministic Plan-only repair.",
|
|
4745
|
+
subtask_prompt: "Parse only $HARNESS_DAG_RUN_DIR/generate-backend-md-plan-pi/plan.md, validate the optional strict module layout and output-budget proof, apply at most one deterministic Plan-only Module Index repair without project writes, then emit one JSON line. No testcase/md/README.md fallback.",
|
|
4680
4746
|
shell: {
|
|
4681
|
-
commands: [buildBackendTestModuleManifestShellCommand(layout)],
|
|
4747
|
+
commands: [buildBackendTestModuleManifestShellCommand(layout, taskConfig.backendTest?.moduleLayout)],
|
|
4682
4748
|
cwd: ".",
|
|
4683
4749
|
timeoutMs: 60000,
|
|
4684
4750
|
},
|
|
@@ -4714,9 +4780,9 @@ async function buildBackendTestHybridDag(sources) {
|
|
|
4714
4780
|
toolProfile: "write",
|
|
4715
4781
|
complexity: "MED",
|
|
4716
4782
|
writePolicy: "exclusive",
|
|
4717
|
-
allowedPaths: ["
|
|
4783
|
+
allowedPaths: ["{{item.markdownPath}}"],
|
|
4718
4784
|
forbiddenPaths: forbidden,
|
|
4719
|
-
writeSet: ["
|
|
4785
|
+
writeSet: ["{{item.markdownPath}}"],
|
|
4720
4786
|
readSet: [
|
|
4721
4787
|
"{{item.planReadPath}}",
|
|
4722
4788
|
toDagSourcePath(sources, sources.requirementPath),
|
|
@@ -4724,21 +4790,22 @@ async function buildBackendTestHybridDag(sources) {
|
|
|
4724
4790
|
? [toDagSourcePath(sources, sources.constraintPath)]
|
|
4725
4791
|
: []),
|
|
4726
4792
|
...intake.referenceIndex.map((entry) => entry.readPath),
|
|
4727
|
-
|
|
4793
|
+
"{{item.markdownPath}}",
|
|
4728
4794
|
],
|
|
4729
4795
|
writerOutcomePolicy: {
|
|
4730
4796
|
type: "implementation-outcome-v1",
|
|
4731
4797
|
requireChangedFiles: true,
|
|
4732
4798
|
},
|
|
4733
|
-
retryPolicy:
|
|
4734
|
-
outputContract: "Write exactly
|
|
4799
|
+
retryPolicy: BACKEND_TEST_MARKDOWN_BINDING_RETRY_POLICY,
|
|
4800
|
+
outputContract: "Write exactly the frozen `{{item.markdownPath}}` Chinese module Markdown case-card file with BE-<MODULE>-<NNN> cases and the seven required h3 sections; keep machine IDs/literals exact and do not execute pytest or modify production code/config or the README.",
|
|
4735
4801
|
subtaskPromptTemplate: [
|
|
4736
|
-
"This is a required file-generation node for exactly one Markdown module. Read the upstream run-owned Markdown plan artifact at `{{item.planReadPath}}` (Coverage Scope + Coverage Matrix + Module Index) and the bounded references, then immediately use write tools to create the single file
|
|
4802
|
+
"This is a required file-generation node for exactly one Markdown module. Read the upstream run-owned Markdown plan artifact at `{{item.planReadPath}}` (Coverage Scope + Coverage Matrix + Module Index) and the bounded references, then immediately use write tools to create the single frozen file `{{item.markdownPath}}`. Do not read or recreate testcase/md/README.md. Do not end after analysis or planning, and do not return before a non-empty bounded diff exists. Do not modify any other module file.",
|
|
4737
4803
|
"Output budget protocol (hard, max output <=16K per turn): Never paste full Matrix, other modules' case bodies, or source text into assistant chat. Each write/edit tool call touches at most one file (this module). Compact tables/lists are required; omitting required sections or in-scope variants is forbidden. If a Completeness Gate / OUTPUT_LIMIT_RECOVERY retry is injected, continue only listed target paths.",
|
|
4738
4804
|
"The first non-empty response line must be exactly IMPLEMENTATION_OUTCOME: changed after the module file has been written, or IMPLEMENTATION_OUTCOME: blocked when precise missing evidence prevents safe generation. already-satisfied is not valid for this node.",
|
|
4739
4805
|
"Write human-readable content in Simplified Chinese by default. Keep English only for machine-readable IDs and technical literals such as Case/AC/REQ/BR IDs, HTTP methods, paths, field names, enum values, commands, filenames, code symbols and exact source citations.",
|
|
4740
4806
|
'Write the module {{item.stem}} as readable case cards covering every in-scope rule/Test Point the README Coverage Matrix assigns to this module. Every case starts with `## BE-<MODULE>-<NNN>|<中文用例名称>`. `<NNN>` is exactly three zero-padded digits (`001`, `002`, ...), never two digits (`01`), a bare number, or an alphabetic suffix such as `011A`. Every case must include `### 覆盖规则`, `### 测试点`, `### 场景类型`, `### 前置条件`, `### 操作步骤`, `### 预期结果`, and `### 自动化映射` Do not group cases under "## 测试类 ..." (or any h2 grouping) headings that force Cases down to h3; each Case must be a direct h2 (`##`), and its seven sections must be h3 (`###`) children of that Case. If you need to convey a pytest class, state it inside the Case\'s `### 自动化映射` instead. Forbidden: `## 测试类 X` then `### BE-PD-001` and `### 覆盖规则` at the same h3 level. Required: `## BE-PD-001` then `### 覆盖规则`.; `覆盖规则` and `测试点` must reference exact Matrix Rule Keys/Test Points. Add `测试目的`, `验收标准`, `需求依据`, and `测试数据` for readable evidence. The `验收标准` section must list the exact applicable `AC-...` IDs, and every explicit task AC must appear in at least one Case. Every automatable case explicitly names its target pytest script and exactly one primary symbol so traceability scans only that script/symbol. Evidence-only meta cases that exist solely for non-executable assertion/cross-cutting process evidence may declare `脚本:无` and `primary symbol:无` with empty `变体测试点`, and must not invent a business pytest item.',
|
|
4741
|
-
"Name this module file with the
|
|
4807
|
+
"Name this module file with the exact frozen Module Index stem `{{item.stem}}` (filename `{{item.markdownPath}}`). Never reinterpret or rename an explicit-user-layout stem. Priority-only stems `p0`, `p1` and `p2` are forbidden and must never produce `p0.md` or `test_p0.py`. Pure hexadecimal/hash-like opaque stems such as `a401606` and `deadbeef` are also forbidden. Do not use Case-ID-like module filenames. For every automatable case, `自动化映射` must name exactly `{{item.pytestPath}}`, where the module stem is this Markdown filename without `.md`, lowercased, with non-alphanumeric characters replaced by underscores. Example: `health` → `testcase/test_health.py`; `resource_notes` → `testcase/test_resource_notes.py`. Never invent a different pytest path in Markdown than the module stem implies.",
|
|
4808
|
+
"AUTOMATION_BINDING_FORMAT_V1 is a literal machine contract. Under every Case's `### 自动化映射`, write these independent lines exactly: `- 脚本:<path|无>`, `- primary symbol:<symbol|无>`, `- 变体测试点:<semicolon-separated TP IDs|无>`, `- 场景断言测试点:<semicolon-separated TP IDs|无>`, `- 横切证据测试点:<semicolon-separated TP IDs|无>`. TP IDs must be on the same line after the colon. Forbidden classification forms include `TP-X(变体测试点)`, `[变体测试点] TP-X`, `【变体测试点】:TP-X`, pipe-delimited annotations, tables, or nested TP lists. Before returning, verify that the Case `### 测试点` exact set equals the pairwise-disjoint union of the three canonical binding lines; do not add, remove, rename or duplicate a TP to make the format pass.",
|
|
4742
4809
|
"Scenario Partition slots: when README declares `## Scenario Partitions`, every slot of each declared partition MUST appear in this module's Cases as exactly one variant Test Point with the deterministic ID `TP-<Partition ID>-<VALUE-TOKEN>` (each-value), `TP-<Partition ID>-OMITTED` (optional axis only) and exactly one `TP-<Partition ID>-NOT-IN-SET` complement slot with `intent=enum-invalid`. Example: Partition ID `SP-GET-API-RESOURCE-NOTES-STATUS` → `TP-SP-GET-API-RESOURCE-NOTES-STATUS-ACTIVE`. Slot IDs copy the declared Partition ID exactly; never drop the HTTP method, invent, merge, renumber or split slot IDs. Before returning, derive the complete exact slot set from every applicable Scenario Partitions row and verify that the module Cases declare and bind every slot assigned by the Coverage Matrix; ordinary alias Test Points do not satisfy a partition slot, and aggregate aliases such as `SINGLE`/`MULTIPLE` are forbidden. Prefer ONE Case per partition with a parameter table over duplicated Cases per value. The not-in-set slot value must be a concrete literal absent from the Domain (e.g. `UNKNOWN_TYPE`) and its expected result must come from the bound source — when Expected by Slot is GAP, the Case states the expectation as GAP evidence, never a guessed 空列表/400. Never create cross-axis combination variants beyond the single documented nominal.",
|
|
4743
4810
|
"For every variant Test Point, write its machine-checkable `场景意图: <TP-ID>; operation=...; target=...; intent=...` line inside that same Case body/自动化映射. Never collect Scenario Intent lines in a file-level appendix, implementation-details block, or another Case; local TP ownership is mandatory.",
|
|
4744
4811
|
"Every Case must keep at least one numbered executable line under `### 操作步骤`; a compact variant/result table may follow but must not replace the numbered action anchor. Keep numbered/bulleted independently assertable results under `### 预期结果`. The exact `### 操作步骤` and `### 预期结果` headings must remain present for every Case, including compact/table-based Cases; never compress later Cases by dropping required headings. Every result must name the observable HTTP status, response field/value, state transition or membership condition, never vague wording such as ‘符合预期’.",
|
|
@@ -4781,7 +4848,7 @@ async function buildBackendTestHybridDag(sources) {
|
|
|
4781
4848
|
"For every variant Test Point, ensure the Markdown scenario intent is machine-checkable and located inside that same Case body/自动化映射, never in a file-level appendix, implementation-details block, or another Case. Use an exact transport target: `场景意图: <TP-ID>; operation=<METHOD /path>; target=<body.field|query.field|path.field|header.field|request>; intent=<empty|missing|null|min-1|min|max|max+1|pattern-invalid|enum-invalid|wrong-type|nominal-operation|custom-literal:V>; bound=<n optional>; example=<optional>; expectedCode=<optional>`. Never use vague targets such as field=resource/health. Keep pytest params aligned to the exact target. For intent=missing/empty/default-omit, pytest may use `_OMIT` or delete the key; for intent=enum-invalid use a concrete invalid enum literal (for example `UNKNOWN_STATUS`), never `_OMIT`/missing-key; for trim/padded samples use `custom-literal:trim` or a real padded string, not a bare token like `filter-active` when the intent is `custom-literal:ACTIVE`.",
|
|
4782
4849
|
"Treat the requirement document as the coverage baseline; scope is limited to operations/rules it (or its referenced API contract) describes, and API contract evidence supplements scenario dimensions. For every in-scope operation, check applicable lifecycle/uniqueness states (including deleted-existing when in scope), valid enum values, bounded invalid classes, min-1/min/nominal/max/max+1, allowed/forbidden format classes, required/null/missing/wrong-type semantics, status/error codes, auth and state transitions. Inspect shared validator/helper/DTO/query builder evidence and expand Affected Operations when the same affected path can affect them; unresolved impact stays visible as GAP/CONFLICT. Directly add in-scope omissions; reject scope expansion to operations absent from the requirement document; undefined impact remains GAP/CONFLICT rather than invented behavior.",
|
|
4783
4850
|
"Check AC completeness/meaning, endpoint, fields/shape, status/error codes, rules, states, documented boundaries/auth, positive/negative coverage, executable steps and assertable results. Require the exact `## Coverage Scope` Field/Value table with the `|---|---|` separator row, a valid classification-policy pair, non-empty Affected Operations/Rule Keys/Scope Evidence, and the classification-specific Regression Floor. Require the exact unnumbered `## Coverage Matrix` heading in the immutable run-owned plan artifact, exact headers, exactly 9 cells in every data row (including a non-empty Dimension), deterministic OpenAPI Rule Keys for every in-scope affected operation, exactly one Matrix row per Rule Key (merge multi-dimension product rows), and bidirectional Matrix Rule/Test Point ↔ Case bindings. Never describe affected-scope coverage as whole-API completeness. Every explicit AC ID must appear in at least one Case `验收标准`; every explicit in-scope AC/REQ/BR Rule Key cited by a Case must have exactly one Coverage Matrix row, and no Case may cite a source Rule Key omitted from the Matrix. Every Matrix Case ID must share at least one of that row's Required Test Points and the Case must cite that Rule Key. Perform an explicit execution-redundancy review: merge checkpoint-only parameter rows, repeated default/read-back assertions, DELETE status/body/follow-up-read checks, response schema/Content-Type checks, PUT full-update/timestamp checks, repeated list setup and identical null/empty inputs when endpoint, input partition, precondition state and expected outcome are the same. Preserve separate POST/PUT, boundary, enum, wrong-type, role/tenant and distinct business-state variants. Directly repair malformed headings/rows/keys and binding modes rather than merely commenting on them. Reject avoidable English prose, duplicated bilingual wording, repeated boilerplate, oversized unstructured sections, a `### 操作步骤` section that contains only a table without any numbered executable line, vague results such as ‘符合预期’, Case-ID-like module filenames (for example `BE-HEALTH.md`), dropped exact `### 操作步骤`/`### 预期结果` headings, and missing or drifted script/function mapping where it can be derived.",
|
|
4784
|
-
"Correct testcase/md/** directly: add documented omissions, remove unsupported cases, rename
|
|
4851
|
+
"Correct testcase/md/** directly: add documented omissions, remove unsupported cases, preserve every frozen Module Index filename exactly (never rename an explicit-user-layout module; model-derived invalid stems must have been rejected before map expansion), normalize every Case ID to hyphen-separated module segments plus exactly three zero-padded digits (`BE-RESOURCE_NOTES-01` → `BE-RESOURCE-NOTES-001`; `BE-RN-011A` must be renumbered or merged) consistently across headings/index/mappings, fix automation mappings so each automatable case points at `testcase/test_<module>.py` derived from that module filename and declares exactly one primary symbol (evidence-only meta cases may keep `脚本/primary symbol=无` with empty variants), assign every Test Point exactly one of `变体测试点`/`场景断言测试点`/`横切证据测试点`, then perform an exact-set check: each Case's `### 测试点` set must equal (not merely contain) the union of those three binding lists; delete stale/legacy aliases and ensure every binding-list Test Point is present, expand every variant parameter row into its own atomic TP ID, make every non-cross-cutting TP Case-specific and owned by exactly one Case, require every primary symbol to start with the canonical Case prefix, ensure every explicit AC ID appears in an applicable Case `验收标准`, merge execution duplicates, improve navigation/tables/Chinese wording, or record gaps in Chinese. Remove every credential/header value, placeholder, fake token and anti-example from Markdown. Sensitive key names may remain only as a plain list; values must be described as runtime-only and omitted, with no colon/value pair or literal example anywhere, including details blocks and explanatory text. Keep Case IDs, AC/REQ/BR IDs, HTTP methods, paths, fields, enum values, filenames, code symbols and source citations as exact machine-readable identifiers; only normalize Case ID separator/sequence formatting as specified above. Recalculate predicted collected items as `sum(max(1, variant count per Case))`; when the task declares a budget, directly merge redundant journeys/reclassify same-request checkpoints until the prediction is within budget, while preserving all required coverage. The validator accepts Chinese and legacy English section aliases; retain or converge to the Chinese human-readable headings without losing structure.",
|
|
4785
4852
|
"This is the single Markdown incremental synchronization round. Read every authoritative reference index entry whose role hints include acceptance-criteria, api-contract, data-contract or business-rule; do not rely on the derived PRD as a complete inventory. Preserve every explicit AC/REQ/BR ID, every documented HTTP/business error code, every DTO/JSON field, enum value, boundary, format, nested shape, transaction/state/idempotency/uniqueness/auth/tenant/cross-field rule. For each natural-language normative business rule preserved as required scope, include its exact source sentence without paraphrase together with source path and line/heading anchor so the deterministic ledger can verify quote/hash provenance. Ensure every Case declares exactly `Payload Contract: none` or the three labels `Payload Required Paths`, `Payload Allowed Paths`, and `Payload Enum`; every label must occupy its own machine-readable list line, and a Case must never concatenate target/setup operations or multiple `Payload Contract` tokens onto one line, and explanatory prose/details must not repeat any `Payload Contract:` token; never infer missing keys or enum values. A target GET/DELETE operation with no request body must remain `Payload Contract: none` even when its setup journey performs POST/PUT with a DTO; setup payloads never redefine the target Case payload contract. Add only missing Matrix rows/Test Points/Cases/assertions or repair exact drift; do not rewrite already-valid unrelated modules. Work gap-targeted: inspect source anchors and affected modules first, leave unrelated valid modules byte-stable, and return `already-satisfied` without restating the full suite when no gap exists.",
|
|
4786
4853
|
"For affected API fields, use one valid nominal payload plus atomic required/missing/null/empty/wrong-type, every documented enum value plus bounded invalid classes, documented min-1/min/nominal/max/max+1, formats and nested object/array constraints. Do not generate a Cartesian product or invent undocumented constraints. Do not invent a concrete identifier type when the source only requires presence; for a missing-resource 404 path with unspecified identifier syntax/type, synchronize the Case to a create-delete-derived valid identifier journey rather than an arbitrary UUID/text placeholder.",
|
|
4787
4854
|
"Scenario Partitions synchronization: when the run-owned plan declares `## Scenario Partitions`, verify each declared partition's slots are fully materialized as variant Test Points with exact `TP-<Partition ID>-...` IDs (each-value per Domain value, OMITTED only for optional axes, exactly one NOT-IN-SET with intent=enum-invalid). Before returning, derive the complete exact slot set from every legal Scenario Partitions row and compare it with both the binding Coverage Matrix Rule's Required Test Points and the final Case `### 测试点`/`变体测试点` sets; directly add every missing exact slot to the already-assigned Case IDs; ordinary alias Test Points do not satisfy a partition slot, and aggregate aliases such as `SINGLE`/`MULTIPLE` are forbidden. Directly add missing slot rows/Cases. Record an illegal plan Partition row that has no source-backed finite domain as GAP/CONFLICT and remove only its derived `TP-SP-*` slots/Cases from target modules; never modify the immutable plan artifact. Never delete a legal source-backed partition or drop its complement slot to force coverage green. When the bound source does not document the complement expectation, keep the slot with GAP expected instead of guessing. Body-field validation enums (`TP-<FIELD>-ENUM-*`) are NOT partitions — do not add partition rows for them.",
|
|
@@ -4834,10 +4901,10 @@ async function buildBackendTestHybridDag(sources) {
|
|
|
4834
4901
|
writePolicy: "read-only",
|
|
4835
4902
|
allowedPaths: ro,
|
|
4836
4903
|
forbiddenPaths: forbidden,
|
|
4837
|
-
outputContract: "Stdout JSON {modules:[{stem,planReadPath}],planReadPath,planSha256} parsed from the run-owned Markdown plan artifact,
|
|
4838
|
-
subtask_prompt: "Parse only the run-owned generate-backend-md-plan-pi/plan.md artifact
|
|
4904
|
+
outputContract: "Stdout JSON {modules:[{stem,planReadPath}],planReadPath,planSha256,moduleLayout} parsed from the final strict-layout-validated run-owned Markdown plan artifact, matching the Markdown map manifest.",
|
|
4905
|
+
subtask_prompt: "Parse only the final run-owned generate-backend-md-plan-pi/plan.md artifact with the same strict module-layout and output-budget contract; do not search for or fall back to testcase/**/README.md. No project file writes.",
|
|
4839
4906
|
shell: {
|
|
4840
|
-
commands: [buildBackendTestModuleManifestShellCommand(layout)],
|
|
4907
|
+
commands: [buildBackendTestModuleManifestShellCommand(layout, taskConfig.backendTest?.moduleLayout)],
|
|
4841
4908
|
cwd: ".",
|
|
4842
4909
|
timeoutMs: 60000,
|
|
4843
4910
|
},
|
|
@@ -4873,7 +4940,7 @@ async function buildBackendTestHybridDag(sources) {
|
|
|
4873
4940
|
toolProfile: "write",
|
|
4874
4941
|
complexity: "MED",
|
|
4875
4942
|
writePolicy: "exclusive",
|
|
4876
|
-
allowedPaths: ["
|
|
4943
|
+
allowedPaths: ["{{item.pytestPath}}"],
|
|
4877
4944
|
forbiddenPaths: Array.from(new Set([
|
|
4878
4945
|
...forbidden,
|
|
4879
4946
|
"testcase/md/**",
|
|
@@ -4882,22 +4949,22 @@ async function buildBackendTestHybridDag(sources) {
|
|
|
4882
4949
|
"pyproject.toml",
|
|
4883
4950
|
"setup.cfg",
|
|
4884
4951
|
])),
|
|
4885
|
-
writeSet: ["
|
|
4952
|
+
writeSet: ["{{item.pytestPath}}"],
|
|
4886
4953
|
writerOutcomePolicy: {
|
|
4887
4954
|
type: "implementation-outcome-v1",
|
|
4888
4955
|
requireChangedFiles: true,
|
|
4889
4956
|
},
|
|
4890
4957
|
retryPolicy: BACKEND_TEST_WRITER_COMPLETENESS_RETRY_POLICY,
|
|
4891
|
-
outputContract: "Write exactly
|
|
4958
|
+
outputContract: "Write exactly the frozen pytest module file `{{item.pytestPath}}` whose actual test function region contains the exact Case ID, preferably in the function name or docstring. the frozen `{{item.markdownPath}}` maps one-to-one to `{{item.pytestPath}}`; never merge or split modules. No JSON and no pytest execution.",
|
|
4892
4959
|
subtaskPromptTemplate: [
|
|
4893
|
-
"Convert the single Markdown module
|
|
4894
|
-
"Output budget protocol (hard, max output <=16K per turn): Write exactly
|
|
4960
|
+
"Convert the single frozen Markdown module `{{item.markdownPath}}` into one self-contained pytest module. Before writing, also read the run-owned Markdown plan artifact at `{{item.planReadPath}}` and use its explicit API target/environment table as the authoritative fallback base URL for every module. A task/Markdown `API_BASE_URL` target takes precedence over project README dev-server URLs; never infer a backend API fallback from a frontend/Vite port such as localhost:3000. After reading the module Markdown, the run-owned plan artifact, and the bounded pytest config/conftest, immediately use write tools to create the single frozen file `{{item.pytestPath}}`. Define any bounded HTTP client fixture, request logging/redaction/truncation helper and payload builders needed by this module inside that same file; do not import generated testcase/**/helpers/** or testcase/**/factories/** assets. Do not end after analysis or planning. Do not modify Markdown, conftest, helpers/factories, or any other module's pytest script.",
|
|
4961
|
+
"Output budget protocol (hard, max output <=16K per turn): Write exactly the frozen `{{item.pytestPath}}`. Never paste full Python modules into assistant chat. Do not merge or split modules. Do not reduce params/assertions/skips to fit. If OUTPUT_LIMIT_RECOVERY is injected, continue only listed missing/broken scripts.",
|
|
4895
4962
|
"Align every variant pytest.param payload with the Markdown scenario intent (empty/missing/null/length/pattern/enum/wrong-type/nominal). Prefer literal payloads over Faker for intent-critical fields so pre-execution scenario-param checks can verify them. Hard contract: intent=enum-invalid MUST pass a concrete invalid value literal (string/number/boolean), never `_OMIT`/None/missing key; intent=missing/empty may use `_OMIT` or delete the key; intent=custom-literal:trim|whitespace-padded requires a leading/trailing whitespace string with non-empty trimmed content (all-whitespace belongs to empty/whitespace-only, not trim); intent=custom-literal:ACTIVE|ARCHIVED requires the exact enum string, never descriptive tokens like filter-active; intent=max/min/max+1 should pass a repeated-string length expression, a bare length number N, or a helper named _*_LEN{N} / _*_MAX_LENGTH / _*_OVER_LENGTH — never a bare 1 for oversize. Hard contract: request payload dicts may only contain DTO field keys from Payload Allowed Paths; never put expect/expected/echo_* helper keys inside the JSON body dict. Path/query/header identifiers and scenario-control metadata (including `id`, expected codes, and selector labels) must stay in separate pytest parameters and helper arguments; never merge them into a DTO patch or JSON body unless that exact path is allowed by the Markdown payload contract. Normalize the configured API base URL with `rstrip(\"/\")` (or equivalently join exactly one slash) before appending endpoint paths; generated requests must never contain a `//api/...` path. When the bound source documents a concrete non-secret local API URL, generated clients must use it as the fallback in `os.environ.get(\"API_BASE_URL\", \"<documented-url>\")`; do not require an otherwise-uninjected environment variable or fail setup solely because it is absent. Missing-field helpers must remove keys idempotently with `payload.pop(field, None)`, never `del payload[field]`, because optional fields may already be absent.",
|
|
4896
4963
|
"For every response contract that requires an object or pagination envelope, first assert that each envelope/data value is a dict and that required keys exist, then index fields and assert values. Never let an incidental KeyError or list/string TypeError stand in for the explicit response-shape contract failure.",
|
|
4897
4964
|
'Ensure every automatable final Markdown Case ID in this module appears in exactly one primary pytest test function or pytest test class method region, using the exact `primary symbol` declared by Markdown. Skip evidence-only meta Cases that declare `脚本/primary symbol=无` with empty variants; do not invent a business pytest symbol for them. The symbol must start with `test_BE_<MODULE>_<NNN>_` so every parameterized collected item remains associated with its Case. Module-level functions and class-based pytest methods are both supported. Only `变体测试点` may use stable `pytest.param(..., id="TP-...")` IDs, and every atomic variant ID must appear exactly once with a genuine input/state/outcome change. Use a literal direct `pytest.param(..., id=...)` expression for every row; never hide or wrap it behind `_post_case`, `_put_case`, row-factory functions, comprehensions, generators, or dynamically returned parameter lists; do not use decorator-level `ids=[...]`, generated suffixes, or IDs that extend/shorten the exact Markdown TP. Do not parameterize `场景断言测试点` or `横切证据测试点`; execute all assertion checkpoints within the same business journey/item and use shared helpers for cross-cutting evidence. The primary symbol docstring must contain exact metadata lines `Case-ID: BE-...`, `Assertion-Test-Points: TP-...;TP-...` and `Cross-Cutting-Test-Points: TP-...;TP-...` (use `none` when empty). Implement request dictionaries so their direct and nested key paths and enum literals exactly satisfy the Case `Payload Required Paths`, `Payload Allowed Paths`, and `Payload Enum`; for `Payload Contract: none`, do not invent a JSON/body DTO. GET/DELETE setup journeys may create resources, but their setup DTO must not change the target operation\'s no-body payload contract. No Test Point may be invented, renamed, omitted or bound in two modes. The generated pytest collection shape must equal the Markdown prediction `sum(max(1, variant count per Case))`; keep it at or below the task\'s explicit budget by removing duplicate execution, never by collapsing multiple parameter rows under a coarse family TP. Assertions come only from 预期结果 and setup comes only from 前置条件/测试数据/自动化映射.',
|
|
4898
|
-
"Name the generated pytest file so it corresponds one-to-one with its source Markdown module file: this module stem `{{item.stem}}` maps to exactly
|
|
4965
|
+
"Name the generated pytest file so it corresponds one-to-one with its source Markdown module file: this module stem `{{item.stem}}` maps to exactly the frozen `{{item.pytestPath}}`. The <module> stem is the Markdown filename without the `.md` extension, lowercased and with non-alphanumeric characters replaced by underscores. For example, `resource_notes` → `testcase/test_resource_notes.py`, `health` → `testcase/test_health.py`. If Markdown automation mapping names a different path than this module stem path, still write the frozen manifest pytest path and do not invent prefixes. Never merge multiple Markdown modules into one pytest file, never split one module across several files, and never invent pytest filenames unrelated to the Markdown modules.",
|
|
4899
4966
|
"Scenario Partition slots: every `TP-<Partition ID>-...` variant Test Point declared by this module's Markdown MUST become exactly one literal direct `pytest.param(..., id=\"TP-<Partition ID>-...\")` row with the exact slot ID; the not-in-set slot passes a concrete literal absent from the documented Domain (e.g. `UNKNOWN_TYPE`) — never `_OMIT`, never a descriptive token. Never split one slot into multiple params or merge several slots under a family TP id. Slot filtering requests hit the documented list endpoint with the slot value as the query/path filter.",
|
|
4900
|
-
"Keep this module self-contained: define module-local fixtures and helpers directly in
|
|
4967
|
+
"Keep this module self-contained: define module-local fixtures and helpers directly in `{{item.pytestPath}}`, so pytest discovers every fixture dependency without external plugin registration. The request log must include method, URL/path, and request parameters (query plus JSON/body/payload summary). The response log must include status code and response result (JSON/text/body summary), and both records must be visible in pytest stdout/stderr without changing assertions. Recursively redact sensitive values and apply bounded truncation before logging.",
|
|
4901
4968
|
"Materialize every automatable Markdown Case exactly once as one canonical primary pytest symbol. Preserve every explicit variant Test Point as a stable pytest.param id and every assertion/cross-cutting binding as declared. Build request payloads from the effective Markdown test data literally: keep all declared DTO keys, nested shapes, enum values, missing/null/boundary variants and business-state preconditions; never substitute guessed convenience fields or rename contract fields. Never assert an identifier's concrete Python/JSON type unless the Markdown or bound contract explicitly declares that type; when only presence is required, accept any non-null scalar identifier and serialize it safely into the path. For a nonexistent-resource 404 Case whose identifier syntax/type is not declared, obtain a syntactically valid identifier from a live create response and delete it before the 404 request; never invent an arbitrary UUID/text identifier that may fail path conversion with 400. Respect every local helper's actual return signature: never tuple-unpack a scalar status/id/helper result, and never treat a tuple response as a scalar.",
|
|
4902
4969
|
"Do not read source/**, add cases, reassign ACs, modify conftest/config/production code, use skip/xfail, swallow assertions, execute pytest, or emit JSON. For best-effort cleanup, catch only the narrow transport exception actually raised by the selected HTTP client (for example `requests.RequestException` or `urllib.error.URLError`); never use bare `except`, `Exception`, or `BaseException` with `pass`.",
|
|
4903
4970
|
...(sharedSetupPytestPrompt ? [sharedSetupPytestPrompt] : []),
|
|
@@ -305,6 +305,26 @@ function buildAttemptPrompt(task, basePrompt, attemptNumber, previousFailureCate
|
|
|
305
305
|
}
|
|
306
306
|
if (previousFailureCategory === "incomplete-write-set") {
|
|
307
307
|
const maxAttempts = task.retryPolicy?.maxAttempts ?? 3;
|
|
308
|
+
const bindingOnly = task.id?.startsWith("generate-backend-md-case-") === true &&
|
|
309
|
+
(recoveryTargetPaths?.length ?? 0) === 1 &&
|
|
310
|
+
Object.values(recoveryDiagnostics ?? {}).flat().some((detail) => /(?:unclassified Test Points|duplicate Test Point bindings)/i.test(detail));
|
|
311
|
+
if (bindingOnly) {
|
|
312
|
+
const target = recoveryTargetPaths[0];
|
|
313
|
+
const diagnostics = (recoveryDiagnostics?.[target] ?? []).slice(0, 4);
|
|
314
|
+
return [
|
|
315
|
+
"<retry_instruction>",
|
|
316
|
+
"BACKEND_TEST_LIGHTWEIGHT_BINDING_REPAIR attempt=1/1",
|
|
317
|
+
`target_path=${target}`,
|
|
318
|
+
"Read and edit only that single Markdown file. Do not read the PRD, references, plan, other modules, pytest, source code, or reports.",
|
|
319
|
+
"Modify only the affected Case `### 自动化映射` binding lines. Preserve every Case ID, Rule, Test Point, scenario intent, payload contract, script path, primary symbol, step and expected result.",
|
|
320
|
+
"Materialize exactly these canonical lines in every affected Case: `- 变体测试点:...`, `- 场景断言测试点:...`, `- 横切证据测试点:...`; use `无` for an empty set.",
|
|
321
|
+
"Do not guess a classification from negation, alternatives or vague prose. If exact classification cannot be proven from the same Case, return IMPLEMENTATION_OUTCOME: blocked without changing the file.",
|
|
322
|
+
"Diagnostics:",
|
|
323
|
+
...diagnostics.map((detail) => `- ${detail}`),
|
|
324
|
+
"After editing, verify the declared Test Point set exactly equals the disjoint union of the three binding lists. Return a short IMPLEMENTATION_OUTCOME only.",
|
|
325
|
+
"</retry_instruction>",
|
|
326
|
+
].join("\n");
|
|
327
|
+
}
|
|
308
328
|
// Embed concrete missing/broken paths from the run-owned progress facts
|
|
309
329
|
// so the continuation attempt is fully self-contained and never reads
|
|
310
330
|
// a forbidden `.harness/**` evidence file. When the loader finds no
|
|
@@ -209,6 +209,11 @@ export const BACKEND_TEST_WRITER_COMPLETENESS_RETRY_POLICY = {
|
|
|
209
209
|
INCOMPLETE_WRITE_SET_RETRY_CATEGORY,
|
|
210
210
|
],
|
|
211
211
|
};
|
|
212
|
+
/** Markdown shard writers use one full attempt plus at most one bounded binding repair. */
|
|
213
|
+
export const BACKEND_TEST_MARKDOWN_BINDING_RETRY_POLICY = {
|
|
214
|
+
...BACKEND_TEST_WRITER_COMPLETENESS_RETRY_POLICY,
|
|
215
|
+
maxAttempts: 2,
|
|
216
|
+
};
|
|
212
217
|
/**
|
|
213
218
|
* Deterministic helper: is this raw failure category eligible for retry under
|
|
214
219
|
* the given policy? Pure function; executor never decides retry eligibility.
|
|
@@ -148,7 +148,7 @@
|
|
|
148
148
|
"validate-backend-test-environment-shell"
|
|
149
149
|
],
|
|
150
150
|
"complexity": "MED",
|
|
151
|
-
"subtask_prompt": "This is a required plan-generation node. Read only the strict read set and return the complete Markdown plan in the assistant response. The runtime persists the response as a run-owned Harness artifact named generate-backend-md-plan-pi/plan.md. Do not write testcase/md/README.md or any project file; module case cards are written by downstream sharded nodes.\n\nOutput budget protocol (hard, max output <=16K per turn): Never paste full Matrix, case bodies, or source text into assistant chat. README holds only Scope+Matrix+module index; never inline full case bodies. If a Completeness Gate / OUTPUT_LIMIT_RECOVERY retry is injected, continue only listed target paths.\n\nReturn the complete plan as plain Markdown. Do not emit JSON or code fences. The plan must contain the exact ## Coverage Scope, ## Coverage Matrix and ## Module Index sections required by the downstream manifest.\n\nRead the upstream environment report only through the strict read set. Generate the Markdown-first backend test plan; it will be persisted under the current DAG run's Harness artifacts, not under testcase/md/.\n\nWrite human-readable content in Simplified Chinese by default. Keep English only for machine-readable IDs and technical literals such as Case/AC/REQ/BR IDs, HTTP methods, paths, field names, enum values, commands, filenames, code symbols and exact source citations.\n\nCreate the concise plan entry page: test objective, target/environment, isolation/cleanup, module summary and a linked case index table with Case ID, Chinese case name, scenario type, endpoint and expected status/result. Avoid repeating every case body in the plan artifact.\n\nBefore the Coverage Matrix, write a mandatory machine-readable `## Coverage Scope` section in the plan artifact using exactly `| Field | Value |`, immediately followed by the separator row `|---|---|`, and these six unique rows: `Change Classification`, `Coverage Policy`, `Affected Operations`, `Affected Rule Keys`, `Regression Floor`, `Scope Evidence`. Always set `Change Classification` to `new-operation` and `Coverage Policy` to `full-contract`; do NOT reason about whether operations are new or existing. Cover all in-scope rules from the requirement document at full depth; treat the product requirement as the coverage baseline and use API contract evidence (fields/status/enum/boundary/format) to supplement scenario dimensions. Scope is limited to operations/rules the requirement document (or its referenced API contract) explicitly describes; do not expand to unrelated operations that the requirement does not mention. List affected operations exactly as `METHOD /path`, stable rule keys separated by semicolons, and precise source pointers as Scope Evidence.\n\nCoverage depth is full over the in-scope rules: fully cover every documented status, request/response field rule, requiredness, enum, boundary, format, auth and business state of each affected operation the requirement describes, but do not re-test unrelated operations the requirement does not mention. Inspect shared validator/helper/DTO/query builder evidence and expand Affected Operations when the same affected path can affect them; unresolved impact stays visible as GAP/CONFLICT.\n\nBefore writing cases, build the mandatory machine-readable Coverage Matrix inside the plan artifact itself. Its section heading line must be exactly `## Coverage Matrix` with no numeric prefix/suffix; never place the canonical Matrix only in a module file. Use this exact header: `| Rule Key | Priority | Source | Endpoint/Field | Dimension | Rule | Required Test Points | Case IDs | Status |`. Every data row must contain exactly 9 pipe-delimited cells and must never omit `Dimension`; use concise dimensions such as requirement, operation, response-status, requiredness, enum, boundary, format, business-state or error. Use only P0/P1/P2 and COVERED/PARTIAL/GAP/CONFLICT. Use stable `TP-<UPPERCASE-HYPHENATED-ID>` test points separated by semicolons.\n\nEach Rule Key must appear in exactly one Matrix row. Preserve each AC/REQ/BR Rule Key as one row; if one product rule spans multiple dimensions, use a concise composite Dimension in that single row instead of duplicating the key. Derive OpenAPI Rule Keys exactly as the deterministic analyzer does: operation token is `<HTTP-METHOD>-<PATH>` with braces removed and every non-alphanumeric run replaced by a hyphen, uppercase (for example POST `/api/resource-notes` → `POST-API-RESOURCE-NOTES`); response statuses use `API-<OPERATION>-RESPONSE-STATUS`; body/parameter fields use `API-<OPERATION>-<FIELD>-REQUIRED|ENUM|MIN-LENGTH|MAX-LENGTH|MINIMUM|MAXIMUM|PATTERN|FORMAT`. Do not invent aliases such as API-CREATE-FIELDS when a deterministic key applies.\n\nCoverage priority is strict inside the declared scope: P0 product requirements/task hard constraints always remain in scope; P1 exhaustively supplements documented operations, fields, business rules, statuses and errors only for Affected Operations; P2 adds bounded protocol robustness only when it is relevant to the change and does not invent product behavior. Coverage percentages describe the declared affected scope, never whole-API completeness unless every operation is explicitly listed. Conflicts or undefined expectations must stay visible as GAP/CONFLICT with precise source pointers, never guessed.\n\nFor uniqueness/lifecycle rules cover absent, active-existing, deleted-existing, create-delete-recreate, restore-then-recreate and documented scope/case-normalization states. For every enum cover every valid value plus bounded invalid equivalence classes (unknown, case variant, whitespace, empty, null/missing and wrong types as applicable). For every length/number rule cover min-1, min, nominal, max and max+1. For format rules cover each allowed class separately plus a valid mixed value, and representative forbidden classes including uppercase, internal/leading/trailing whitespace, tab/newline, unsupported punctuation, slash, emoji or control characters when the source contract supports that expectation.\n\nMandatory module index: include a `## Module Index` table in the plan artifact that lists every planned module as a canonical relative link of the exact form `[label](./<stem>.md)` plus a `testcase/md/<stem>.md` path cell, so a downstream deterministic manifest can parse the module list. Group by stable business resource/domain, not by CRUD operation: one resource's list/detail/create/update/delete cases belong in one module such as `resource_notes`; split only when a single module would exceed the per-child 16K output protocol, keep the total module count at the smallest safe value, and never exceed 8 modules. Name each module file with a stable lowercase business stem such as `health` or `resource_notes`. Pure hexadecimal/hash-like opaque stems such as `a401606` or `deadbeef` are forbidden. Do not use priority-only stems `p0`, `p1` or `p2`; Priority belongs only in the Coverage Matrix and never defines module files. Do not use Case-ID-like module filenames such as `BE-HEALTH.md` or `BE-NOTES.md`. The relative link target MUST equal the on-disk filename stem the sharded writer will create. For every automatable case, `自动化映射` must name exactly `testcase/test_<module>.py`, where <module> is that Markdown filename without `.md`, lowercased, with non-alphanumeric characters replaced by underscores. Example: `testcase/md/health.md` → `testcase/test_health.py`; `testcase/md/resource_notes.md` → `testcase/test_resource_notes.py`. Never invent a different pytest path in Markdown than the module stem implies.\n\nScenario Partitions (query/filter axes): for every affected GET/list operation, declare one row per enum or classification axis used for filtering (query/path parameters such as type/status/category). Add a mandatory machine-readable `## Scenario Partitions` section after the Coverage Matrix using exactly `| Partition ID | Operation | Axis | Domain | Required Slots | Expected by Slot | Bind Rule |` with the separator row. Partition ID is a stable `SP-<OPERATION>-<AXIS>` token; Domain must copy the legal values verbatim from the bound OpenAPI enum or requirement sentence (never guess), using bare semicolon-separated identifier values inside the single table cell (for example `ACTIVE; ARCHIVED`, with no Markdown backticks or prose); Required Slots must contain `each-value` and exactly one `not-in-set`, plus `omitted` only when the parameter is optional. Before returning, expand every declared partition into its complete deterministic exact slot ID set: one `TP-<Partition ID>-<VALUE-TOKEN>` per Domain value, `TP-<Partition ID>-OMITTED` only for an optional axis, and exactly one `TP-<Partition ID>-NOT-IN-SET`. Every expanded slot ID must appear verbatim in the binding Rule's `Required Test Points` cell and be assigned to concrete Case IDs in that same Coverage Matrix row; ordinary alias/family Test Points do not replace this inventory. Expected by Slot states the documented expectation per slot kind (`domain-value`, `default-behavior`, `empty-result`/`excluded-result` when documented, or `GAP` when the source does not document the complement expectation — never invent 空列表/400). POST/PUT body field-validation enums stay in the Coverage Matrix as `TP-<FIELD>-ENUM-*` and MUST NOT get a Scenario Partition row. Do not create partitions for axes without a documented legal-value domain. Only GET/list query or path parameters whose bound source documents a finite enum or classification set may become a Scenario Partition. Do not create partitions for free-form strings, primary keys, required-or-optional-only parameters, or boundary/format-only axes. If an axis has no finite legal-value domain, do not declare a Partition row and do not invent NOT-IN-SET cases. Cross-axis combinations stay as ONE nominal Case; never declare a cross-axis cartesian partition.\n\nBefore finalizing README, calculate the predicted collected-item count as `sum(max(1, number of variant Test Points in each Case))`. If the task declares an item budget, the prediction must not exceed it. Reduce excess only by removing duplicate execution and converting same-request checkpoints to assertions; never drop required rules, boundaries, enums, operation-specific inputs, or business states. Record the prediction in README. Use only environment-supported fixtures/targets/isolation, record evidence gaps in Chinese, and do not emit JSON, pytest, or execute commands.\n\n## Derived task contract: 需求.md\n\n# Backend test\n- AC-001 proof\n\n## Authoritative reference index\n\n[]\n\nFor each index entry, use `readPath` for Pi read-tool calls and copy `path` exactly into Markdown Source References. Bound files under .harness/tasks/<taskId>/source/** are read-only inputs: reading them is allowed even though writing .harness/** is forbidden. Never resolve `path` relative to the repository root, search for substitutes, or fall back to docs/** when a bound read fails.\n\nRead only precise indexed references needed for AC/API/field/rule evidence; references remain authoritative over derived text.",
|
|
151
|
+
"subtask_prompt": "This is a required plan-generation node. Read only the strict read set and return the complete Markdown plan in the assistant response. The runtime persists the response as a run-owned Harness artifact named generate-backend-md-plan-pi/plan.md. Do not write testcase/md/README.md or any project file; module case cards are written by downstream sharded nodes.\n\nOutput budget protocol (hard, max output <=16K per turn): Never paste full Matrix, case bodies, or source text into assistant chat. README holds only Scope+Matrix+module index; never inline full case bodies. If a Completeness Gate / OUTPUT_LIMIT_RECOVERY retry is injected, continue only listed target paths.\n\nReturn the complete plan as plain Markdown. Do not emit JSON or code fences. The plan must contain the exact ## Coverage Scope, ## Coverage Matrix and ## Module Index sections required by the downstream manifest.\n\nRead the upstream environment report only through the strict read set. Generate the Markdown-first backend test plan; it will be persisted under the current DAG run's Harness artifacts, not under testcase/md/.\n\nWrite human-readable content in Simplified Chinese by default. Keep English only for machine-readable IDs and technical literals such as Case/AC/REQ/BR IDs, HTTP methods, paths, field names, enum values, commands, filenames, code symbols and exact source citations.\n\nCreate the concise plan entry page: test objective, target/environment, isolation/cleanup, module summary and a linked case index table with Case ID, Chinese case name, scenario type, endpoint and expected status/result. Avoid repeating every case body in the plan artifact.\n\nBefore the Coverage Matrix, write a mandatory machine-readable `## Coverage Scope` section in the plan artifact using exactly `| Field | Value |`, immediately followed by the separator row `|---|---|`, and these six unique rows: `Change Classification`, `Coverage Policy`, `Affected Operations`, `Affected Rule Keys`, `Regression Floor`, `Scope Evidence`. Always set `Change Classification` to `new-operation` and `Coverage Policy` to `full-contract`; do NOT reason about whether operations are new or existing. Cover all in-scope rules from the requirement document at full depth; treat the product requirement as the coverage baseline and use API contract evidence (fields/status/enum/boundary/format) to supplement scenario dimensions. Scope is limited to operations/rules the requirement document (or its referenced API contract) explicitly describes; do not expand to unrelated operations that the requirement does not mention. List affected operations exactly as `METHOD /path`, stable rule keys separated by semicolons, and precise source pointers as Scope Evidence.\n\nCoverage depth is full over the in-scope rules: fully cover every documented status, request/response field rule, requiredness, enum, boundary, format, auth and business state of each affected operation the requirement describes, but do not re-test unrelated operations the requirement does not mention. Inspect shared validator/helper/DTO/query builder evidence and expand Affected Operations when the same affected path can affect them; unresolved impact stays visible as GAP/CONFLICT.\n\nBefore writing cases, build the mandatory machine-readable Coverage Matrix inside the plan artifact itself. Its section heading line must be exactly `## Coverage Matrix` with no numeric prefix/suffix; never place the canonical Matrix only in a module file. Use this exact header: `| Rule Key | Priority | Source | Endpoint/Field | Dimension | Rule | Required Test Points | Case IDs | Status |`. Every data row must contain exactly 9 pipe-delimited cells and must never omit `Dimension`; use concise dimensions such as requirement, operation, response-status, requiredness, enum, boundary, format, business-state or error. Use only P0/P1/P2 and COVERED/PARTIAL/GAP/CONFLICT. Use stable `TP-<UPPERCASE-HYPHENATED-ID>` test points separated by semicolons.\n\nEach Rule Key must appear in exactly one Matrix row. Preserve each AC/REQ/BR Rule Key as one row; if one product rule spans multiple dimensions, use a concise composite Dimension in that single row instead of duplicating the key. Derive OpenAPI Rule Keys exactly as the deterministic analyzer does: operation token is `<HTTP-METHOD>-<PATH>` with braces removed and every non-alphanumeric run replaced by a hyphen, uppercase (for example POST `/api/resource-notes` → `POST-API-RESOURCE-NOTES`); response statuses use `API-<OPERATION>-RESPONSE-STATUS`; body/parameter fields use `API-<OPERATION>-<FIELD>-REQUIRED|ENUM|MIN-LENGTH|MAX-LENGTH|MINIMUM|MAXIMUM|PATTERN|FORMAT`. Do not invent aliases such as API-CREATE-FIELDS when a deterministic key applies.\n\nCoverage priority is strict inside the declared scope: P0 product requirements/task hard constraints always remain in scope; P1 exhaustively supplements documented operations, fields, business rules, statuses and errors only for Affected Operations; P2 adds bounded protocol robustness only when it is relevant to the change and does not invent product behavior. Coverage percentages describe the declared affected scope, never whole-API completeness unless every operation is explicitly listed. Conflicts or undefined expectations must stay visible as GAP/CONFLICT with precise source pointers, never guessed.\n\nFor uniqueness/lifecycle rules cover absent, active-existing, deleted-existing, create-delete-recreate, restore-then-recreate and documented scope/case-normalization states. For every enum cover every valid value plus bounded invalid equivalence classes (unknown, case variant, whitespace, empty, null/missing and wrong types as applicable). For every length/number rule cover min-1, min, nominal, max and max+1. For format rules cover each allowed class separately plus a valid mixed value, and representative forbidden classes including uppercase, internal/leading/trailing whitespace, tab/newline, unsupported punctuation, slash, emoji or control characters when the source contract supports that expectation.\n\nMandatory module layout contract: first inspect the PRIMARY requirement for an explicit list of required Markdown/Python output path pairs. When explicit paths are present, they are authoritative `explicit-user-layout`: reproduce their exact filenames, count and one-to-one pairs in Module Index; do not rename, merge, split, omit or add a module from reference/example scripts. Only when the primary requirement has no explicit file layout may you derive the smallest `business-resource-layout`. Existing examples, historical regression functions and shared setup may add evidence/assertions to an existing required module, but never create an extra physical module by themselves. A reference-only `resp_regression`, positive/negative/boundary/error/response module is forbidden.\n\nInclude exactly one `## Module Index` table with this exact header: `| Module Stem | Business Resource | Owned Operations | Owned Rule Keys | Case IDs | Split Reason | Markdown Path | Pytest Path |`. Use canonical relative links `[label](./<stem>.md)` inside the Markdown Path cell followed by the resolved `${layout.markdownDir}/<stem>.md` path. Split Reason is exactly one of `explicit-user-layout`, `primary-business-resource`, `independent-business-resource`, or `output-budget`. Group by stable business resource/domain, not by CRUD operation, AC, parameter/field axis, scenario type or regression purpose: one resource's list/detail/create/update/delete and its filters/response assertions/regression floor belong in one module. Multiple modules owning the same exact `METHOD /path` are forbidden unless every such row is `explicit-user-layout` from primary-requirement path pairs or has a documented `output-budget` proof. Keep the total module count at the smallest safe value and never exceed 8 modules. Name model-derived modules with stable lowercase business stems such as `health` or `resource_notes`; explicit-user-layout preserves the primary requirement filename stem even when it is more specific. Do not use priority-only stems `p0`, `p1` or `p2`; Priority belongs only in the Coverage Matrix. Pure hexadecimal/hash-like opaque stems and test-purpose-only stems are forbidden. Do not use Case-ID-like module filenames. The relative link target, Markdown Path, Pytest Path and downstream automation mapping must be one-to-one and exact; for model-derived modules the default pair remains `testcase/md/<module>.md` and `testcase/test_<module>.py`, while explicit-user-layout preserves the primary requirement paths. Do not hand-write a conflicting module count in prose; the Module Index row count is the only count truth.\n\nScenario Partitions (query/filter axes): for every affected GET/list operation, declare one row per enum or classification axis used for filtering (query/path parameters such as type/status/category). Add a mandatory machine-readable `## Scenario Partitions` section after the Coverage Matrix using exactly `| Partition ID | Operation | Axis | Domain | Required Slots | Expected by Slot | Bind Rule |` with the separator row. Partition ID is a stable `SP-<OPERATION>-<AXIS>` token; Domain must copy the legal values verbatim from the bound OpenAPI enum or requirement sentence (never guess), using bare semicolon-separated identifier values inside the single table cell (for example `ACTIVE; ARCHIVED`, with no Markdown backticks or prose); Required Slots must contain `each-value` and exactly one `not-in-set`, plus `omitted` only when the parameter is optional. Before returning, expand every declared partition into its complete deterministic exact slot ID set: one `TP-<Partition ID>-<VALUE-TOKEN>` per Domain value, `TP-<Partition ID>-OMITTED` only for an optional axis, and exactly one `TP-<Partition ID>-NOT-IN-SET`. Every expanded slot ID must appear verbatim in the binding Rule's `Required Test Points` cell and be assigned to concrete Case IDs in that same Coverage Matrix row; ordinary alias/family Test Points do not replace this inventory. Expected by Slot states the documented expectation per slot kind (`domain-value`, `default-behavior`, `empty-result`/`excluded-result` when documented, or `GAP` when the source does not document the complement expectation — never invent 空列表/400). POST/PUT body field-validation enums stay in the Coverage Matrix as `TP-<FIELD>-ENUM-*` and MUST NOT get a Scenario Partition row. Do not create partitions for axes without a documented legal-value domain. Only GET/list query or path parameters whose bound source documents a finite enum or classification set may become a Scenario Partition. Do not create partitions for free-form strings, primary keys, required-or-optional-only parameters, or boundary/format-only axes. If an axis has no finite legal-value domain, do not declare a Partition row and do not invent NOT-IN-SET cases. Cross-axis combinations stay as ONE nominal Case; never declare a cross-axis cartesian partition.\n\nBefore finalizing README, calculate the predicted collected-item count as `sum(max(1, number of variant Test Points in each Case))`. If the task declares an item budget, the prediction must not exceed it. Reduce excess only by removing duplicate execution and converting same-request checkpoints to assertions; never drop required rules, boundaries, enums, operation-specific inputs, or business states. Record the prediction in README. Use only environment-supported fixtures/targets/isolation, record evidence gaps in Chinese, and do not emit JSON, pytest, or execute commands.\n\n## Derived task contract: 需求.md\n\n# Backend test\n- AC-001 proof\n\n## Authoritative reference index\n\n[]\n\nFor each index entry, use `readPath` for Pi read-tool calls and copy `path` exactly into Markdown Source References. Bound files under .harness/tasks/<taskId>/source/** are read-only inputs: reading them is allowed even though writing .harness/** is forbidden. Never resolve `path` relative to the repository root, search for substitutes, or fall back to docs/** when a bound read fails.\n\nRead only precise indexed references needed for AC/API/field/rule evidence; references remain authoritative over derived text.",
|
|
152
152
|
"executor": "pi",
|
|
153
153
|
"role": "planner",
|
|
154
154
|
"toolProfile": "read-only",
|
|
@@ -185,18 +185,18 @@
|
|
|
185
185
|
"generate-backend-md-plan-pi"
|
|
186
186
|
],
|
|
187
187
|
"complexity": "LOW",
|
|
188
|
-
"subtask_prompt": "Parse only $HARNESS_DAG_RUN_DIR/generate-backend-md-plan-pi/plan.md and
|
|
188
|
+
"subtask_prompt": "Parse only $HARNESS_DAG_RUN_DIR/generate-backend-md-plan-pi/plan.md, validate the optional strict module layout and output-budget proof, apply at most one deterministic Plan-only Module Index repair without project writes, then emit one JSON line. No testcase/md/README.md fallback.",
|
|
189
189
|
"executor": "shell",
|
|
190
190
|
"role": "verifier",
|
|
191
191
|
"writePolicy": "read-only",
|
|
192
192
|
"shell": {
|
|
193
193
|
"commands": [
|
|
194
|
-
"node -e \"eval(Buffer.from('
|
|
194
|
+
"node -e \"eval(require('zlib').inflateRawSync(Buffer.from('pRprc9s28nt+Bd1kDLIiKTu95lKptCevtr5L4oydtHMnKR6YXEm48FUAsq2I+u83CxAkqIedTr/YBLBY7HsXC8VFLqQzFRGHPxeMg0umgnh+SeW8ncIR8fyYL0tZtNN6TDz/a8qu22kcEW/4KFaos+Q149F3EoSMqYB+lnxnlkTMWSk7y+2a5CyWb+myWMgoX6TproXLOX3647POMog4EtGJCDmUKY3B7Y/Go/FktXa973unn59U4/GkP/PJePzksCWS371LPiDLOdw6FzB7c1e6IGJXEe/1yHjcd0cvgv/S4OtR8NNVGEx63ngcZgnxycxCs8iRm5IXMQgRQn4T/vbi4v2by8ur1y9+vbr49P7q9dlFVREyfMSm7oGG91Zmg5AJcB7ecibBJRkTguUzZwcKZ1pw55rGXyBPApSd847yL0lxmztlSnOHcsmmNJbjnHjDhpw7Jt2n3nBdU4uQimlUb/i/guWuJsgnM8iBUwmBOSNLAgQPSkZ8gl/Iu6e5mCrUQorLZR67Bqv3EFvfQP7AIT2Dr0f2MJOCdDjQJINoimqnyS8shQ4tPlnI6fNWT3nBM7SSS8lZPnOFF8ribXEL/BUV4HqW7XzWGp/00GauiLX0+apXXfWe4Lw9fWUgzVnXMtLHhFNeZK/mlL8qEnB/eubZFl2+pPEXyeIvQluvKFMm3Wvpab2QFh/Lb2jKkgugosgjTm+jk5XBEyFjLqe33pBN3f7ncnQUPJ086YcoY1dIz+MgFzx3SMlZwZlcBkWeLoOsSBYpBEJCRvTOEQ2mE/yDzK+e+c9+WO9CU5T0zwUEcyrmW0iEjKKIaM2QZgcHAfwGki3wA3XoVzxU+djk+10H1twHYplLelcT654OrqGSZUXjisOf1TX3RlfBpM+292OMCVL2Bbrn18sqkKyNpNVJWsIdoSsBR1En7PC7C0jfsvxL1B+PRp/Hk0lvPBm743BX4AizZOz1Z2YvTXEnYAhGWbW2NObjXMWq3LawMW/mtJVox6iRzYEmLJ+JaDQZTgvuonew6GjIfjbHhCnkMzkfsl7PW7Gpa+ZHbBJKzjIXeSOPHzvvlIScszyBO+IZxGG5EHOXoeexqdvMaqQHUXS8x+83QaMoOjo1wcBog6mzBiRZlCmLMQB1FrweGTp1nnHgjsYyXTpFXn87GzQ7AmLJivzeICgk5TIytI2OJr3jIcoM8iTalFgrT7XrHpn2Pz9+PBa98WVtwtsy9rwVHsGG1xzol+G6oUfT3J4tUhaDqw70IU9MQLB1jiY6mphRqmypxmObSDhlqQTuptFJGrI8ThcJCJdUxGswxZCmIkIM0Qn+NdsrNDZFyLEfHHthRkv3hqYLiE660UvPeobJJivflRBLSH4DmgCPRqTW0yV6n09eLgQSLZwLEMWCx0B8cn6bQ+Kcl5iHWJGLZuoCN/4bljiD8do5e42fl0ipoz2U+KRJKBj/iU8+LFWeUaOJ7SzAlbEopkU4ZXmixq4WQh1aeXEbKdmoac/Eiy5bIdwAX7paBL6yWC864cXtSH1PoijS0hmudea0jv/5aF+6NBHPdoRA7xx0fcG5NmIsbnPgYs5Kh+aJg+ndiYt0keXiXl9IqKQXxa2oZaEVbhHZe6o1ryTRmBPKBrmsfeAken54iEwfTcz/gyjqKNyzIl9xC4lSnSq9LkG6IwJ36P1MBgsBPEhVqYdlB2cZ5cvAsBnw1lpQLCXkCeRy53qxkOUCl5IZSDLxutXkm1qPSJ1oCHHtUvPw0B6FWhuiqkYTLRMmIYtO8G+IKaU1/cJY8LlSisL+jpauN6yFDnFKOSRaQCpuY4zB0uUCSsr4i7JMGSTRlKYCho8wCDVG6RTTRmve6pHjGK4gq2sBpQDPN6KoaxE1fzzBStT4qt8QKmygpzVQHQiGpOP86p9xdmMQL4siBZp7vtpjI/uxc+LwkeOoOsGW8sEOjYRzKlwt1dW2XCRfAMpSsnwBw7VGemAYfqAK3TYWrDnxrL31psKvig4OoixyARV+VBxmHIRgRV6VhWCS3UCVw4yqj+tikSeULyvgvOBVXGTXRaXl5bmng9GVE0yqJ54pWRrq95GPUEG54GUh4D4enODEIT0zfz9PB7Y7apnj114aNiKTAg64CsCbBKi1+09vzc8qEB7SHga6JGi2foPuzKX0HZXxHN0tDEPLQp/VFprh8os0dc3N0NNumdVpRd2bGiwq1tmj0dEErVs5sko8Ct4655/1Oe11AkvFfvcuMR5jpdcnxlOaE6xqq6qUn3cOP554B1GkHGaf/LTKDDcBpoggY0Kx/W0egPW6Vdj2g0lvHJZLU7S3bHtV1Q6suiMMybealt6vqNy0rBb1g0qvk2USPRzXVR2wGdOjWqa1Ngw6FXdN0NFToW0moardxB9MznU/oYcqraoW2hKPBdt0R2r43ci78rSg/rLM352//vT2zdXbF/85//Tx6tX5+1/enr36OKhTpKoihFMspGAJOFNefIXc4UUhhWpGPGgy2oIt0g+iaCdLtrnYMO0scqkDmr1uoqCpIavqX5fn70OhnI5Nl1aEQffYWGzQqJDSlp1eVangZZ+kJnSluTchaZY74WInt0MrPuxgdWgkt+twvKd1C5vaFo3Fl7wopraEEOgDTtbFRwlYkuzxAJXMLR+wth8ebs6EM14syrMkiqLSHnvDmiL0D7VSVRoAhGQZlZCcKxawMyJ+rjdn9M6a/QD8D7RRXlXv0J5jYKm7H0lfAWX0zj32FYN1uPS8kwewa2hRZLCH75DlpWk7HhhOrbmq2tpRU1jw34FjcdBs21zY65hn739/8fbs9dX5p48fPn28evnp9a9vPl59uDg//+UbQrXj4N82CirGVrjJt+3Qb23O3/SkgbGg073ONjAe6W/4z469mxCtX+rNeMHD+93Aypc//I0yFHs+Z0kH3T/+blWrHXAHc/aq+vbDMGy89/BwlzOerqzBYBfEerBae2t0pY3rQt2NMWmpvRw0QsUrgjLKzcBWX231XTHauKaEM5BtwPTwnjPUkPpEbUCKwYFOkC3fa2+4iU3Y2HyNCM1T9ZDs8GNdYuoa7xui08GmUJQH01guaBqd6P8mg3duaDon6XNMkPiL2VH1q01L26jOqcl0SM9g33VHrFs5vu5rJSCBZyxnQrLY+YDddmzOOlwlFyemeV6o9i/kUjVB+qr70dz17yvWtpJUnSi0pCW9TiEakcohvY2WhqbQqRyk0amIT6otGGTN9aITEgQBMUxVuKEiaP2b2mlFoU4cNQLx6+DZDS3+LuutTxk6xLPWTeTYWq1DwNa8ZbU+GZFeQ0qPTNyw35kIs8Rz6plO5FQzbficbAnNm3Sadij4+sECbx8bnb4jv9OM9HylHX8bEJuB24inQlvs9huIPrJ9CsGksLbaCaNNB51g6NjwZNXdrANB3e85Pjw8qGd0A0wNohP1T0sYC5WdTZ2q2oDaKGf2eWNxAzylZanugIbC+rYgMCk2k+aeUFOIttelb9sRM+AzcK6XbT/NJDen4M5CgINPVAs5xxcUdb13djG32x9VxDvYCnmtGrD3hoJXPbiOm15TDtFGu1W3I7XmGxQZ7kfoHddYb8XprY7hGV4Uh6Y62Nhu+scWhvqFYxeKdfdxyjw+1PcuxIdvJqaZqovXjTcV9WSll7x6yUo1A93NwkZWnWKUIA2gid3toeiUKvVpCG+nHSnP1Rh7ZOB03H2oNY0PBNcpbNtCPykyynLVa7v/hQEgb3qKpuuXtd2+R9tEb0RM5XSKDbdNHtgqA8hVn8ZKais1SRMbdJhtFwvaEM3Cwx0XyEq57DSh2+4zlU4KFGuJHLbkpbfskdAWCSfP9xJwh5PspnkSiotFrl6Iuxh6xDlxnlteLOfgiIymKejnnvu1CQ+/mV8ATdp3cw6pigGu2RLfJq7nN8/gdY2pYAWUJtb02y44gta/ZNC/qAhjDlTCb1TMXSLUCvHCRZlQCa4O416YsBn2Wcgc38WGjyyhYZ2khbZxx12ZAJltZWJ3FYYhfvo2h2vP64z9llRf49AhbLAS8RwyepYMiP2sb1SlA2KAT/oiuDkmfn1psMPgKZFUfAmwkctpLMlA/cwgSICzG0iIbxb08YPtn4EgRV2U270dqCqy1S1tXhisKklKNPhXysS2iqfT48GRBXy+kHGRwQ44Es9pPoOEDEheYG9Ul4dkvcbq8/8=','base64')).toString('utf8'))\""
|
|
195
195
|
],
|
|
196
196
|
"timeoutMs": 60000,
|
|
197
197
|
"cwd": "."
|
|
198
198
|
},
|
|
199
|
-
"outputContract": "Stdout JSON {modules:[{stem,planReadPath}],planReadPath,planSha256} parsed from the run-owned generate-backend-md-plan-pi/plan.md artifact
|
|
199
|
+
"outputContract": "Stdout JSON {modules:[{stem,planReadPath}],planReadPath,planSha256,moduleLayout} parsed from the run-owned generate-backend-md-plan-pi/plan.md artifact after strict-layout validation and at most one deterministic Plan-only repair.",
|
|
200
200
|
"allowedPaths": [
|
|
201
201
|
"testcase/**",
|
|
202
202
|
"docs/test-reports/**"
|
|
@@ -243,12 +243,12 @@
|
|
|
243
243
|
"verification-before-completion"
|
|
244
244
|
],
|
|
245
245
|
"complexity": "MED",
|
|
246
|
-
"subtaskPromptTemplate": "This is a required file-generation node for exactly one Markdown module. Read the upstream run-owned Markdown plan artifact at `{{item.planReadPath}}` (Coverage Scope + Coverage Matrix + Module Index) and the bounded references, then immediately use write tools to create the single file testcase/md/{{item.stem}}.md. Do not read or recreate testcase/md/README.md. Do not end after analysis or planning, and do not return before a non-empty bounded diff exists. Do not modify any other module file.\n\nOutput budget protocol (hard, max output <=16K per turn): Never paste full Matrix, other modules' case bodies, or source text into assistant chat. Each write/edit tool call touches at most one file (this module). Compact tables/lists are required; omitting required sections or in-scope variants is forbidden. If a Completeness Gate / OUTPUT_LIMIT_RECOVERY retry is injected, continue only listed target paths.\n\nThe first non-empty response line must be exactly IMPLEMENTATION_OUTCOME: changed after the module file has been written, or IMPLEMENTATION_OUTCOME: blocked when precise missing evidence prevents safe generation. already-satisfied is not valid for this node.\n\nWrite human-readable content in Simplified Chinese by default. Keep English only for machine-readable IDs and technical literals such as Case/AC/REQ/BR IDs, HTTP methods, paths, field names, enum values, commands, filenames, code symbols and exact source citations.\n\nWrite the module {{item.stem}} as readable case cards covering every in-scope rule/Test Point the README Coverage Matrix assigns to this module. Every case starts with `## BE-<MODULE>-<NNN>|<中文用例名称>`. `<NNN>` is exactly three zero-padded digits (`001`, `002`, ...), never two digits (`01`), a bare number, or an alphabetic suffix such as `011A`. Every case must include `### 覆盖规则`, `### 测试点`, `### 场景类型`, `### 前置条件`, `### 操作步骤`, `### 预期结果`, and `### 自动化映射` Do not group cases under \"## 测试类 ...\" (or any h2 grouping) headings that force Cases down to h3; each Case must be a direct h2 (`##`), and its seven sections must be h3 (`###`) children of that Case. If you need to convey a pytest class, state it inside the Case's `### 自动化映射` instead. Forbidden: `## 测试类 X` then `### BE-PD-001` and `### 覆盖规则` at the same h3 level. Required: `## BE-PD-001` then `### 覆盖规则`.; `覆盖规则` and `测试点` must reference exact Matrix Rule Keys/Test Points. Add `测试目的`, `验收标准`, `需求依据`, and `测试数据` for readable evidence. The `验收标准` section must list the exact applicable `AC-...` IDs, and every explicit task AC must appear in at least one Case. Every automatable case explicitly names its target pytest script and exactly one primary symbol so traceability scans only that script/symbol. Evidence-only meta cases that exist solely for non-executable assertion/cross-cutting process evidence may declare `脚本:无` and `primary symbol:无` with empty `变体测试点`, and must not invent a business pytest item.\n\nName this module file with the stable lowercase business stem `{{item.stem}}` (filename `testcase/md/{{item.stem}}.md`). Priority-only stems `p0`, `p1` and `p2` are forbidden and must never produce `p0.md` or `test_p0.py`. Pure hexadecimal/hash-like opaque stems such as `a401606` and `deadbeef` are also forbidden. Do not use Case-ID-like module filenames. For every automatable case, `自动化映射` must name exactly `testcase/test_{{item.stem}}.py`, where the module stem is this Markdown filename without `.md`, lowercased, with non-alphanumeric characters replaced by underscores. Example: `health` → `testcase/test_health.py`; `resource_notes` → `testcase/test_resource_notes.py`. Never invent a different pytest path in Markdown than the module stem implies.\n\nScenario Partition slots: when README declares `## Scenario Partitions`, every slot of each declared partition MUST appear in this module's Cases as exactly one variant Test Point with the deterministic ID `TP-<Partition ID>-<VALUE-TOKEN>` (each-value), `TP-<Partition ID>-OMITTED` (optional axis only) and exactly one `TP-<Partition ID>-NOT-IN-SET` complement slot with `intent=enum-invalid`. Example: Partition ID `SP-GET-API-RESOURCE-NOTES-STATUS` → `TP-SP-GET-API-RESOURCE-NOTES-STATUS-ACTIVE`. Slot IDs copy the declared Partition ID exactly; never drop the HTTP method, invent, merge, renumber or split slot IDs. Before returning, derive the complete exact slot set from every applicable Scenario Partitions row and verify that the module Cases declare and bind every slot assigned by the Coverage Matrix; ordinary alias Test Points do not satisfy a partition slot, and aggregate aliases such as `SINGLE`/`MULTIPLE` are forbidden. Prefer ONE Case per partition with a parameter table over duplicated Cases per value. The not-in-set slot value must be a concrete literal absent from the Domain (e.g. `UNKNOWN_TYPE`) and its expected result must come from the bound source — when Expected by Slot is GAP, the Case states the expectation as GAP evidence, never a guessed 空列表/400. Never create cross-axis combination variants beyond the single documented nominal.\n\nFor every variant Test Point, write its machine-checkable `场景意图: <TP-ID>; operation=...; target=...; intent=...` line inside that same Case body/自动化映射. Never collect Scenario Intent lines in a file-level appendix, implementation-details block, or another Case; local TP ownership is mandatory.\n\nEvery Case must keep at least one numbered executable line under `### 操作步骤`; a compact variant/result table may follow but must not replace the numbered action anchor. Keep numbered/bulleted independently assertable results under `### 预期结果`. The exact `### 操作步骤` and `### 预期结果` headings must remain present for every Case, including compact/table-based Cases; never compress later Cases by dropping required headings. Every result must name the observable HTTP status, response field/value, state transition or membership condition, never vague wording such as ‘符合预期’.\n\nIn every `自动化映射`, use exactly these machine-readable list labels: `脚本`, `primary symbol`, `变体测试点`, `场景断言测试点`, `横切证据测试点`, plus a deterministic payload contract. For operations without a request body write `Payload Contract: none`. Otherwise write `Payload Required Paths`, `Payload Allowed Paths`, and `Payload Enum` (write `none` when there is no enum); nested fields use dot paths such as `approver.name`. Each Case describes exactly one target request payload contract: put every payload label on its own list line, never concatenate multiple operations or setup POST/PUT contracts into one label line, and never repeat a `Payload Contract:` token inside explanatory prose/details after the machine-readable line. Values must come only from bound API/DTO evidence, never guesses. Each Test Point from `### 测试点` must appear in exactly one binding list, and every Test Point named in any binding list must also be declared in that Case's `### 测试点`; write `无` for an empty list. A variant Test Point is atomic: one exact endpoint/input/precondition/outcome row equals one exact pytest item and one exact TP ID. If a parameter table has five rows, declare five distinct variant TP IDs in Markdown; never declare one family TP and append row suffixes only in pytest. Classify as `variant` only when endpoint, request input, precondition business state, or expected outcome genuinely changes and therefore needs an independent pytest parameter item. Classify CRUD checkpoints, status/body/header/schema assertions and multiple checks over the same response/journey as `assertion`; classify shared HTTP logging/redaction/truncation evidence as `cross-cutting`. Never create a Test Point merely to parameterize a checkpoint. Every non-cross-cutting TP ID is owned by exactly one Case; when the same response/schema/error assertion is needed in different Cases, use distinct Case-specific TP IDs instead of reusing one assertion TP across Cases. Keep the script path identical to the module one-to-one path and declare exactly one primary symbol named with the canonical Case prefix, for example `BE-RN-003` → `test_BE_RN_003_<description>`; non-Case-prefixed primary symbols are forbidden because parameterized item association must remain deterministic. For evidence-only meta Cases with no executable business journey, write `脚本:无` and `primary symbol:无`, keep `变体测试点:无`, and place process evidence only in assertion/cross-cutting lists. If the bound contract only says an identifier is returned/present, do not declare a concrete identifier type. If a 404 Case needs a nonexistent path identifier but its syntax/type is unspecified, define a create-delete-derived valid identifier journey instead of an arbitrary UUID/text placeholder. For redaction scenarios, list sensitive header/field key names only. Never write any header-name-and-value pair, credential placeholder, fake token, anti-example, or other secret-shaped literal in Markdown; state only that a test-only value is supplied at runtime and omitted. Put implementation-only restrictions in a concise `<details>` block rather than dominating the main case flow. Use only environment-supported fixtures/targets/isolation, record evidence gaps in Chinese, and do not emit JSON, pytest, or execute commands.\n\n## Derived task contract: 需求.md\n\n# Backend test\n- AC-001 proof\n\n## Authoritative reference index\n\n[]\n\nFor each index entry, use `readPath` for Pi read-tool calls and copy `path` exactly into Markdown Source References. Bound files under .harness/tasks/<taskId>/source/** are read-only inputs: reading them is allowed even though writing .harness/** is forbidden. Never resolve `path` relative to the repository root, search for substitutes, or fall back to docs/** when a bound read fails.\n\nRead only precise indexed references needed for AC/API/field/rule evidence; references remain authoritative over derived text.",
|
|
247
|
-
"outputContract": "Write exactly
|
|
246
|
+
"subtaskPromptTemplate": "This is a required file-generation node for exactly one Markdown module. Read the upstream run-owned Markdown plan artifact at `{{item.planReadPath}}` (Coverage Scope + Coverage Matrix + Module Index) and the bounded references, then immediately use write tools to create the single frozen file `{{item.markdownPath}}`. Do not read or recreate testcase/md/README.md. Do not end after analysis or planning, and do not return before a non-empty bounded diff exists. Do not modify any other module file.\n\nOutput budget protocol (hard, max output <=16K per turn): Never paste full Matrix, other modules' case bodies, or source text into assistant chat. Each write/edit tool call touches at most one file (this module). Compact tables/lists are required; omitting required sections or in-scope variants is forbidden. If a Completeness Gate / OUTPUT_LIMIT_RECOVERY retry is injected, continue only listed target paths.\n\nThe first non-empty response line must be exactly IMPLEMENTATION_OUTCOME: changed after the module file has been written, or IMPLEMENTATION_OUTCOME: blocked when precise missing evidence prevents safe generation. already-satisfied is not valid for this node.\n\nWrite human-readable content in Simplified Chinese by default. Keep English only for machine-readable IDs and technical literals such as Case/AC/REQ/BR IDs, HTTP methods, paths, field names, enum values, commands, filenames, code symbols and exact source citations.\n\nWrite the module {{item.stem}} as readable case cards covering every in-scope rule/Test Point the README Coverage Matrix assigns to this module. Every case starts with `## BE-<MODULE>-<NNN>|<中文用例名称>`. `<NNN>` is exactly three zero-padded digits (`001`, `002`, ...), never two digits (`01`), a bare number, or an alphabetic suffix such as `011A`. Every case must include `### 覆盖规则`, `### 测试点`, `### 场景类型`, `### 前置条件`, `### 操作步骤`, `### 预期结果`, and `### 自动化映射` Do not group cases under \"## 测试类 ...\" (or any h2 grouping) headings that force Cases down to h3; each Case must be a direct h2 (`##`), and its seven sections must be h3 (`###`) children of that Case. If you need to convey a pytest class, state it inside the Case's `### 自动化映射` instead. Forbidden: `## 测试类 X` then `### BE-PD-001` and `### 覆盖规则` at the same h3 level. Required: `## BE-PD-001` then `### 覆盖规则`.; `覆盖规则` and `测试点` must reference exact Matrix Rule Keys/Test Points. Add `测试目的`, `验收标准`, `需求依据`, and `测试数据` for readable evidence. The `验收标准` section must list the exact applicable `AC-...` IDs, and every explicit task AC must appear in at least one Case. Every automatable case explicitly names its target pytest script and exactly one primary symbol so traceability scans only that script/symbol. Evidence-only meta cases that exist solely for non-executable assertion/cross-cutting process evidence may declare `脚本:无` and `primary symbol:无` with empty `变体测试点`, and must not invent a business pytest item.\n\nName this module file with the exact frozen Module Index stem `{{item.stem}}` (filename `{{item.markdownPath}}`). Never reinterpret or rename an explicit-user-layout stem. Priority-only stems `p0`, `p1` and `p2` are forbidden and must never produce `p0.md` or `test_p0.py`. Pure hexadecimal/hash-like opaque stems such as `a401606` and `deadbeef` are also forbidden. Do not use Case-ID-like module filenames. For every automatable case, `自动化映射` must name exactly `{{item.pytestPath}}`, where the module stem is this Markdown filename without `.md`, lowercased, with non-alphanumeric characters replaced by underscores. Example: `health` → `testcase/test_health.py`; `resource_notes` → `testcase/test_resource_notes.py`. Never invent a different pytest path in Markdown than the module stem implies.\n\nAUTOMATION_BINDING_FORMAT_V1 is a literal machine contract. Under every Case's `### 自动化映射`, write these independent lines exactly: `- 脚本:<path|无>`, `- primary symbol:<symbol|无>`, `- 变体测试点:<semicolon-separated TP IDs|无>`, `- 场景断言测试点:<semicolon-separated TP IDs|无>`, `- 横切证据测试点:<semicolon-separated TP IDs|无>`. TP IDs must be on the same line after the colon. Forbidden classification forms include `TP-X(变体测试点)`, `[变体测试点] TP-X`, `【变体测试点】:TP-X`, pipe-delimited annotations, tables, or nested TP lists. Before returning, verify that the Case `### 测试点` exact set equals the pairwise-disjoint union of the three canonical binding lines; do not add, remove, rename or duplicate a TP to make the format pass.\n\nScenario Partition slots: when README declares `## Scenario Partitions`, every slot of each declared partition MUST appear in this module's Cases as exactly one variant Test Point with the deterministic ID `TP-<Partition ID>-<VALUE-TOKEN>` (each-value), `TP-<Partition ID>-OMITTED` (optional axis only) and exactly one `TP-<Partition ID>-NOT-IN-SET` complement slot with `intent=enum-invalid`. Example: Partition ID `SP-GET-API-RESOURCE-NOTES-STATUS` → `TP-SP-GET-API-RESOURCE-NOTES-STATUS-ACTIVE`. Slot IDs copy the declared Partition ID exactly; never drop the HTTP method, invent, merge, renumber or split slot IDs. Before returning, derive the complete exact slot set from every applicable Scenario Partitions row and verify that the module Cases declare and bind every slot assigned by the Coverage Matrix; ordinary alias Test Points do not satisfy a partition slot, and aggregate aliases such as `SINGLE`/`MULTIPLE` are forbidden. Prefer ONE Case per partition with a parameter table over duplicated Cases per value. The not-in-set slot value must be a concrete literal absent from the Domain (e.g. `UNKNOWN_TYPE`) and its expected result must come from the bound source — when Expected by Slot is GAP, the Case states the expectation as GAP evidence, never a guessed 空列表/400. Never create cross-axis combination variants beyond the single documented nominal.\n\nFor every variant Test Point, write its machine-checkable `场景意图: <TP-ID>; operation=...; target=...; intent=...` line inside that same Case body/自动化映射. Never collect Scenario Intent lines in a file-level appendix, implementation-details block, or another Case; local TP ownership is mandatory.\n\nEvery Case must keep at least one numbered executable line under `### 操作步骤`; a compact variant/result table may follow but must not replace the numbered action anchor. Keep numbered/bulleted independently assertable results under `### 预期结果`. The exact `### 操作步骤` and `### 预期结果` headings must remain present for every Case, including compact/table-based Cases; never compress later Cases by dropping required headings. Every result must name the observable HTTP status, response field/value, state transition or membership condition, never vague wording such as ‘符合预期’.\n\nIn every `自动化映射`, use exactly these machine-readable list labels: `脚本`, `primary symbol`, `变体测试点`, `场景断言测试点`, `横切证据测试点`, plus a deterministic payload contract. For operations without a request body write `Payload Contract: none`. Otherwise write `Payload Required Paths`, `Payload Allowed Paths`, and `Payload Enum` (write `none` when there is no enum); nested fields use dot paths such as `approver.name`. Each Case describes exactly one target request payload contract: put every payload label on its own list line, never concatenate multiple operations or setup POST/PUT contracts into one label line, and never repeat a `Payload Contract:` token inside explanatory prose/details after the machine-readable line. Values must come only from bound API/DTO evidence, never guesses. Each Test Point from `### 测试点` must appear in exactly one binding list, and every Test Point named in any binding list must also be declared in that Case's `### 测试点`; write `无` for an empty list. A variant Test Point is atomic: one exact endpoint/input/precondition/outcome row equals one exact pytest item and one exact TP ID. If a parameter table has five rows, declare five distinct variant TP IDs in Markdown; never declare one family TP and append row suffixes only in pytest. Classify as `variant` only when endpoint, request input, precondition business state, or expected outcome genuinely changes and therefore needs an independent pytest parameter item. Classify CRUD checkpoints, status/body/header/schema assertions and multiple checks over the same response/journey as `assertion`; classify shared HTTP logging/redaction/truncation evidence as `cross-cutting`. Never create a Test Point merely to parameterize a checkpoint. Every non-cross-cutting TP ID is owned by exactly one Case; when the same response/schema/error assertion is needed in different Cases, use distinct Case-specific TP IDs instead of reusing one assertion TP across Cases. Keep the script path identical to the module one-to-one path and declare exactly one primary symbol named with the canonical Case prefix, for example `BE-RN-003` → `test_BE_RN_003_<description>`; non-Case-prefixed primary symbols are forbidden because parameterized item association must remain deterministic. For evidence-only meta Cases with no executable business journey, write `脚本:无` and `primary symbol:无`, keep `变体测试点:无`, and place process evidence only in assertion/cross-cutting lists. If the bound contract only says an identifier is returned/present, do not declare a concrete identifier type. If a 404 Case needs a nonexistent path identifier but its syntax/type is unspecified, define a create-delete-derived valid identifier journey instead of an arbitrary UUID/text placeholder. For redaction scenarios, list sensitive header/field key names only. Never write any header-name-and-value pair, credential placeholder, fake token, anti-example, or other secret-shaped literal in Markdown; state only that a test-only value is supplied at runtime and omitted. Put implementation-only restrictions in a concise `<details>` block rather than dominating the main case flow. Use only environment-supported fixtures/targets/isolation, record evidence gaps in Chinese, and do not emit JSON, pytest, or execute commands.\n\n## Derived task contract: 需求.md\n\n# Backend test\n- AC-001 proof\n\n## Authoritative reference index\n\n[]\n\nFor each index entry, use `readPath` for Pi read-tool calls and copy `path` exactly into Markdown Source References. Bound files under .harness/tasks/<taskId>/source/** are read-only inputs: reading them is allowed even though writing .harness/** is forbidden. Never resolve `path` relative to the repository root, search for substitutes, or fall back to docs/** when a bound read fails.\n\nRead only precise indexed references needed for AC/API/field/rule evidence; references remain authoritative over derived text.",
|
|
247
|
+
"outputContract": "Write exactly the frozen `{{item.markdownPath}}` Chinese module Markdown case-card file with BE-<MODULE>-<NNN> cases and the seven required h3 sections; keep machine IDs/literals exact and do not execute pytest or modify production code/config or the README.",
|
|
248
248
|
"toolProfile": "write",
|
|
249
249
|
"writePolicy": "exclusive",
|
|
250
250
|
"allowedPaths": [
|
|
251
|
-
"
|
|
251
|
+
"{{item.markdownPath}}"
|
|
252
252
|
],
|
|
253
253
|
"forbiddenPaths": [
|
|
254
254
|
".harness/**",
|
|
@@ -256,19 +256,19 @@
|
|
|
256
256
|
"artifacts/**"
|
|
257
257
|
],
|
|
258
258
|
"writeSet": [
|
|
259
|
-
"
|
|
259
|
+
"{{item.markdownPath}}"
|
|
260
260
|
],
|
|
261
261
|
"readSet": [
|
|
262
262
|
"{{item.planReadPath}}",
|
|
263
263
|
"source/需求.md",
|
|
264
|
-
"
|
|
264
|
+
"{{item.markdownPath}}"
|
|
265
265
|
],
|
|
266
266
|
"writerOutcomePolicy": {
|
|
267
267
|
"type": "implementation-outcome-v1",
|
|
268
268
|
"requireChangedFiles": true
|
|
269
269
|
},
|
|
270
270
|
"retryPolicy": {
|
|
271
|
-
"maxAttempts":
|
|
271
|
+
"maxAttempts": 2,
|
|
272
272
|
"backoff": "exponential",
|
|
273
273
|
"initialDelayMs": 2000,
|
|
274
274
|
"maxDelayMs": 30000,
|
|
@@ -291,7 +291,7 @@
|
|
|
291
291
|
"materialize-backend-md-module-manifest-shell"
|
|
292
292
|
],
|
|
293
293
|
"complexity": "MED",
|
|
294
|
-
"subtask_prompt": "Perform one gap-targeted synchronization, not a full-suite rewrite or stylistic review. Read the immutable run-owned Markdown plan from the direct upstream manifest's planReadPath. Start from explicit bound source IDs/error codes/DTO fields/normative quoted rules and the plan Coverage Matrix; open and edit only modules that own a missing or conflicting rule. Never create or edit testcase/md/README.md and never modify the run-owned plan artifact. Preserve unrelated valid modules byte-for-byte and avoid optional wording cleanup.\n\nOutput budget protocol: never dump full Matrix/case bodies into assistant chat. Inspect the immutable run-owned plan first, build a concise target list from its Matrix and Module Index, then read/write only target modules one file per tool call. Do not traverse every module when the Matrix and source token inventory show no gap; return `already-satisfied`. When adding omitted in-scope cases, keep every required section. Do not bulk-delete in-scope cases to save tokens.\n\nFor every variant Test Point, ensure the Markdown scenario intent is machine-checkable and located inside that same Case body/自动化映射, never in a file-level appendix, implementation-details block, or another Case. Use an exact transport target: `场景意图: <TP-ID>; operation=<METHOD /path>; target=<body.field|query.field|path.field|header.field|request>; intent=<empty|missing|null|min-1|min|max|max+1|pattern-invalid|enum-invalid|wrong-type|nominal-operation|custom-literal:V>; bound=<n optional>; example=<optional>; expectedCode=<optional>`. Never use vague targets such as field=resource/health. Keep pytest params aligned to the exact target. For intent=missing/empty/default-omit, pytest may use `_OMIT` or delete the key; for intent=enum-invalid use a concrete invalid enum literal (for example `UNKNOWN_STATUS`), never `_OMIT`/missing-key; for trim/padded samples use `custom-literal:trim` or a real padded string, not a bare token like `filter-active` when the intent is `custom-literal:ACTIVE`.\n\nTreat the requirement document as the coverage baseline; scope is limited to operations/rules it (or its referenced API contract) describes, and API contract evidence supplements scenario dimensions. For every in-scope operation, check applicable lifecycle/uniqueness states (including deleted-existing when in scope), valid enum values, bounded invalid classes, min-1/min/nominal/max/max+1, allowed/forbidden format classes, required/null/missing/wrong-type semantics, status/error codes, auth and state transitions. Inspect shared validator/helper/DTO/query builder evidence and expand Affected Operations when the same affected path can affect them; unresolved impact stays visible as GAP/CONFLICT. Directly add in-scope omissions; reject scope expansion to operations absent from the requirement document; undefined impact remains GAP/CONFLICT rather than invented behavior.\n\nCheck AC completeness/meaning, endpoint, fields/shape, status/error codes, rules, states, documented boundaries/auth, positive/negative coverage, executable steps and assertable results. Require the exact `## Coverage Scope` Field/Value table with the `|---|---|` separator row, a valid classification-policy pair, non-empty Affected Operations/Rule Keys/Scope Evidence, and the classification-specific Regression Floor. Require the exact unnumbered `## Coverage Matrix` heading in the immutable run-owned plan artifact, exact headers, exactly 9 cells in every data row (including a non-empty Dimension), deterministic OpenAPI Rule Keys for every in-scope affected operation, exactly one Matrix row per Rule Key (merge multi-dimension product rows), and bidirectional Matrix Rule/Test Point ↔ Case bindings. Never describe affected-scope coverage as whole-API completeness. Every explicit AC ID must appear in at least one Case `验收标准`; every explicit in-scope AC/REQ/BR Rule Key cited by a Case must have exactly one Coverage Matrix row, and no Case may cite a source Rule Key omitted from the Matrix. Every Matrix Case ID must share at least one of that row's Required Test Points and the Case must cite that Rule Key. Perform an explicit execution-redundancy review: merge checkpoint-only parameter rows, repeated default/read-back assertions, DELETE status/body/follow-up-read checks, response schema/Content-Type checks, PUT full-update/timestamp checks, repeated list setup and identical null/empty inputs when endpoint, input partition, precondition state and expected outcome are the same. Preserve separate POST/PUT, boundary, enum, wrong-type, role/tenant and distinct business-state variants. Directly repair malformed headings/rows/keys and binding modes rather than merely commenting on them. Reject avoidable English prose, duplicated bilingual wording, repeated boilerplate, oversized unstructured sections, a `### 操作步骤` section that contains only a table without any numbered executable line, vague results such as ‘符合预期’, Case-ID-like module filenames (for example `BE-HEALTH.md`), dropped exact `### 操作步骤`/`### 预期结果` headings, and missing or drifted script/function mapping where it can be derived.\n\nCorrect testcase/md/** directly: add documented omissions, remove unsupported cases, rename module files to stable lowercase stems when needed, normalize every Case ID to hyphen-separated module segments plus exactly three zero-padded digits (`BE-RESOURCE_NOTES-01` → `BE-RESOURCE-NOTES-001`; `BE-RN-011A` must be renumbered or merged) consistently across headings/index/mappings, fix automation mappings so each automatable case points at `testcase/test_<module>.py` derived from that module filename and declares exactly one primary symbol (evidence-only meta cases may keep `脚本/primary symbol=无` with empty variants), assign every Test Point exactly one of `变体测试点`/`场景断言测试点`/`横切证据测试点`, then perform an exact-set check: each Case's `### 测试点` set must equal (not merely contain) the union of those three binding lists; delete stale/legacy aliases and ensure every binding-list Test Point is present, expand every variant parameter row into its own atomic TP ID, make every non-cross-cutting TP Case-specific and owned by exactly one Case, require every primary symbol to start with the canonical Case prefix, ensure every explicit AC ID appears in an applicable Case `验收标准`, merge execution duplicates, improve navigation/tables/Chinese wording, or record gaps in Chinese. Remove every credential/header value, placeholder, fake token and anti-example from Markdown. Sensitive key names may remain only as a plain list; values must be described as runtime-only and omitted, with no colon/value pair or literal example anywhere, including details blocks and explanatory text. Keep Case IDs, AC/REQ/BR IDs, HTTP methods, paths, fields, enum values, filenames, code symbols and source citations as exact machine-readable identifiers; only normalize Case ID separator/sequence formatting as specified above. Recalculate predicted collected items as `sum(max(1, variant count per Case))`; when the task declares a budget, directly merge redundant journeys/reclassify same-request checkpoints until the prediction is within budget, while preserving all required coverage. The validator accepts Chinese and legacy English section aliases; retain or converge to the Chinese human-readable headings without losing structure.\n\nThis is the single Markdown incremental synchronization round. Read every authoritative reference index entry whose role hints include acceptance-criteria, api-contract, data-contract or business-rule; do not rely on the derived PRD as a complete inventory. Preserve every explicit AC/REQ/BR ID, every documented HTTP/business error code, every DTO/JSON field, enum value, boundary, format, nested shape, transaction/state/idempotency/uniqueness/auth/tenant/cross-field rule. For each natural-language normative business rule preserved as required scope, include its exact source sentence without paraphrase together with source path and line/heading anchor so the deterministic ledger can verify quote/hash provenance. Ensure every Case declares exactly `Payload Contract: none` or the three labels `Payload Required Paths`, `Payload Allowed Paths`, and `Payload Enum`; every label must occupy its own machine-readable list line, and a Case must never concatenate target/setup operations or multiple `Payload Contract` tokens onto one line, and explanatory prose/details must not repeat any `Payload Contract:` token; never infer missing keys or enum values. A target GET/DELETE operation with no request body must remain `Payload Contract: none` even when its setup journey performs POST/PUT with a DTO; setup payloads never redefine the target Case payload contract. Add only missing Matrix rows/Test Points/Cases/assertions or repair exact drift; do not rewrite already-valid unrelated modules. Work gap-targeted: inspect source anchors and affected modules first, leave unrelated valid modules byte-stable, and return `already-satisfied` without restating the full suite when no gap exists.\n\nFor affected API fields, use one valid nominal payload plus atomic required/missing/null/empty/wrong-type, every documented enum value plus bounded invalid classes, documented min-1/min/nominal/max/max+1, formats and nested object/array constraints. Do not generate a Cartesian product or invent undocumented constraints. Do not invent a concrete identifier type when the source only requires presence; for a missing-resource 404 path with unspecified identifier syntax/type, synchronize the Case to a create-delete-derived valid identifier journey rather than an arbitrary UUID/text placeholder.\n\nScenario Partitions synchronization: when the run-owned plan declares `## Scenario Partitions`, verify each declared partition's slots are fully materialized as variant Test Points with exact `TP-<Partition ID>-...` IDs (each-value per Domain value, OMITTED only for optional axes, exactly one NOT-IN-SET with intent=enum-invalid). Before returning, derive the complete exact slot set from every legal Scenario Partitions row and compare it with both the binding Coverage Matrix Rule's Required Test Points and the final Case `### 测试点`/`变体测试点` sets; directly add every missing exact slot to the already-assigned Case IDs; ordinary alias Test Points do not satisfy a partition slot, and aggregate aliases such as `SINGLE`/`MULTIPLE` are forbidden. Directly add missing slot rows/Cases. Record an illegal plan Partition row that has no source-backed finite domain as GAP/CONFLICT and remove only its derived `TP-SP-*` slots/Cases from target modules; never modify the immutable plan artifact. Never delete a legal source-backed partition or drop its complement slot to force coverage green. When the bound source does not document the complement expectation, keep the slot with GAP expected instead of guessing. Body-field validation enums (`TP-<FIELD>-ENUM-*`) are NOT partitions — do not add partition rows for them.\n\nBefore returning, verify that every explicit source AC/REQ/BR, error code and strong DTO field token appears in the run-owned plan or an applicable module Case. If a fact cannot be safely automated, retain it as GAP/CONFLICT with its exact source pointer instead of dropping it. Return already-satisfied only when no target file needs an incremental edit.\n\nRead only indexed source paths. Do not scan the repository, modify source/**, generate pytest, execute tests, or emit JSON.\n\n## Derived task contract: 需求.md\n\n# Backend test\n- AC-001 proof\n\n## Authoritative reference index\n\n[]\n\nFor each index entry, use `readPath` for Pi read-tool calls and keep `path` as the exact Markdown Source References citation. Bound files under .harness/tasks/<taskId>/source/** are read-only inputs: reading them is allowed even though writing .harness/** is forbidden. Never resolve `path` relative to the repository root, search for substitutes, or fall back to docs/** when a bound read fails.",
|
|
294
|
+
"subtask_prompt": "Perform one gap-targeted synchronization, not a full-suite rewrite or stylistic review. Read the immutable run-owned Markdown plan from the direct upstream manifest's planReadPath. Start from explicit bound source IDs/error codes/DTO fields/normative quoted rules and the plan Coverage Matrix; open and edit only modules that own a missing or conflicting rule. Never create or edit testcase/md/README.md and never modify the run-owned plan artifact. Preserve unrelated valid modules byte-for-byte and avoid optional wording cleanup.\n\nOutput budget protocol: never dump full Matrix/case bodies into assistant chat. Inspect the immutable run-owned plan first, build a concise target list from its Matrix and Module Index, then read/write only target modules one file per tool call. Do not traverse every module when the Matrix and source token inventory show no gap; return `already-satisfied`. When adding omitted in-scope cases, keep every required section. Do not bulk-delete in-scope cases to save tokens.\n\nFor every variant Test Point, ensure the Markdown scenario intent is machine-checkable and located inside that same Case body/自动化映射, never in a file-level appendix, implementation-details block, or another Case. Use an exact transport target: `场景意图: <TP-ID>; operation=<METHOD /path>; target=<body.field|query.field|path.field|header.field|request>; intent=<empty|missing|null|min-1|min|max|max+1|pattern-invalid|enum-invalid|wrong-type|nominal-operation|custom-literal:V>; bound=<n optional>; example=<optional>; expectedCode=<optional>`. Never use vague targets such as field=resource/health. Keep pytest params aligned to the exact target. For intent=missing/empty/default-omit, pytest may use `_OMIT` or delete the key; for intent=enum-invalid use a concrete invalid enum literal (for example `UNKNOWN_STATUS`), never `_OMIT`/missing-key; for trim/padded samples use `custom-literal:trim` or a real padded string, not a bare token like `filter-active` when the intent is `custom-literal:ACTIVE`.\n\nTreat the requirement document as the coverage baseline; scope is limited to operations/rules it (or its referenced API contract) describes, and API contract evidence supplements scenario dimensions. For every in-scope operation, check applicable lifecycle/uniqueness states (including deleted-existing when in scope), valid enum values, bounded invalid classes, min-1/min/nominal/max/max+1, allowed/forbidden format classes, required/null/missing/wrong-type semantics, status/error codes, auth and state transitions. Inspect shared validator/helper/DTO/query builder evidence and expand Affected Operations when the same affected path can affect them; unresolved impact stays visible as GAP/CONFLICT. Directly add in-scope omissions; reject scope expansion to operations absent from the requirement document; undefined impact remains GAP/CONFLICT rather than invented behavior.\n\nCheck AC completeness/meaning, endpoint, fields/shape, status/error codes, rules, states, documented boundaries/auth, positive/negative coverage, executable steps and assertable results. Require the exact `## Coverage Scope` Field/Value table with the `|---|---|` separator row, a valid classification-policy pair, non-empty Affected Operations/Rule Keys/Scope Evidence, and the classification-specific Regression Floor. Require the exact unnumbered `## Coverage Matrix` heading in the immutable run-owned plan artifact, exact headers, exactly 9 cells in every data row (including a non-empty Dimension), deterministic OpenAPI Rule Keys for every in-scope affected operation, exactly one Matrix row per Rule Key (merge multi-dimension product rows), and bidirectional Matrix Rule/Test Point ↔ Case bindings. Never describe affected-scope coverage as whole-API completeness. Every explicit AC ID must appear in at least one Case `验收标准`; every explicit in-scope AC/REQ/BR Rule Key cited by a Case must have exactly one Coverage Matrix row, and no Case may cite a source Rule Key omitted from the Matrix. Every Matrix Case ID must share at least one of that row's Required Test Points and the Case must cite that Rule Key. Perform an explicit execution-redundancy review: merge checkpoint-only parameter rows, repeated default/read-back assertions, DELETE status/body/follow-up-read checks, response schema/Content-Type checks, PUT full-update/timestamp checks, repeated list setup and identical null/empty inputs when endpoint, input partition, precondition state and expected outcome are the same. Preserve separate POST/PUT, boundary, enum, wrong-type, role/tenant and distinct business-state variants. Directly repair malformed headings/rows/keys and binding modes rather than merely commenting on them. Reject avoidable English prose, duplicated bilingual wording, repeated boilerplate, oversized unstructured sections, a `### 操作步骤` section that contains only a table without any numbered executable line, vague results such as ‘符合预期’, Case-ID-like module filenames (for example `BE-HEALTH.md`), dropped exact `### 操作步骤`/`### 预期结果` headings, and missing or drifted script/function mapping where it can be derived.\n\nCorrect testcase/md/** directly: add documented omissions, remove unsupported cases, preserve every frozen Module Index filename exactly (never rename an explicit-user-layout module; model-derived invalid stems must have been rejected before map expansion), normalize every Case ID to hyphen-separated module segments plus exactly three zero-padded digits (`BE-RESOURCE_NOTES-01` → `BE-RESOURCE-NOTES-001`; `BE-RN-011A` must be renumbered or merged) consistently across headings/index/mappings, fix automation mappings so each automatable case points at `testcase/test_<module>.py` derived from that module filename and declares exactly one primary symbol (evidence-only meta cases may keep `脚本/primary symbol=无` with empty variants), assign every Test Point exactly one of `变体测试点`/`场景断言测试点`/`横切证据测试点`, then perform an exact-set check: each Case's `### 测试点` set must equal (not merely contain) the union of those three binding lists; delete stale/legacy aliases and ensure every binding-list Test Point is present, expand every variant parameter row into its own atomic TP ID, make every non-cross-cutting TP Case-specific and owned by exactly one Case, require every primary symbol to start with the canonical Case prefix, ensure every explicit AC ID appears in an applicable Case `验收标准`, merge execution duplicates, improve navigation/tables/Chinese wording, or record gaps in Chinese. Remove every credential/header value, placeholder, fake token and anti-example from Markdown. Sensitive key names may remain only as a plain list; values must be described as runtime-only and omitted, with no colon/value pair or literal example anywhere, including details blocks and explanatory text. Keep Case IDs, AC/REQ/BR IDs, HTTP methods, paths, fields, enum values, filenames, code symbols and source citations as exact machine-readable identifiers; only normalize Case ID separator/sequence formatting as specified above. Recalculate predicted collected items as `sum(max(1, variant count per Case))`; when the task declares a budget, directly merge redundant journeys/reclassify same-request checkpoints until the prediction is within budget, while preserving all required coverage. The validator accepts Chinese and legacy English section aliases; retain or converge to the Chinese human-readable headings without losing structure.\n\nThis is the single Markdown incremental synchronization round. Read every authoritative reference index entry whose role hints include acceptance-criteria, api-contract, data-contract or business-rule; do not rely on the derived PRD as a complete inventory. Preserve every explicit AC/REQ/BR ID, every documented HTTP/business error code, every DTO/JSON field, enum value, boundary, format, nested shape, transaction/state/idempotency/uniqueness/auth/tenant/cross-field rule. For each natural-language normative business rule preserved as required scope, include its exact source sentence without paraphrase together with source path and line/heading anchor so the deterministic ledger can verify quote/hash provenance. Ensure every Case declares exactly `Payload Contract: none` or the three labels `Payload Required Paths`, `Payload Allowed Paths`, and `Payload Enum`; every label must occupy its own machine-readable list line, and a Case must never concatenate target/setup operations or multiple `Payload Contract` tokens onto one line, and explanatory prose/details must not repeat any `Payload Contract:` token; never infer missing keys or enum values. A target GET/DELETE operation with no request body must remain `Payload Contract: none` even when its setup journey performs POST/PUT with a DTO; setup payloads never redefine the target Case payload contract. Add only missing Matrix rows/Test Points/Cases/assertions or repair exact drift; do not rewrite already-valid unrelated modules. Work gap-targeted: inspect source anchors and affected modules first, leave unrelated valid modules byte-stable, and return `already-satisfied` without restating the full suite when no gap exists.\n\nFor affected API fields, use one valid nominal payload plus atomic required/missing/null/empty/wrong-type, every documented enum value plus bounded invalid classes, documented min-1/min/nominal/max/max+1, formats and nested object/array constraints. Do not generate a Cartesian product or invent undocumented constraints. Do not invent a concrete identifier type when the source only requires presence; for a missing-resource 404 path with unspecified identifier syntax/type, synchronize the Case to a create-delete-derived valid identifier journey rather than an arbitrary UUID/text placeholder.\n\nScenario Partitions synchronization: when the run-owned plan declares `## Scenario Partitions`, verify each declared partition's slots are fully materialized as variant Test Points with exact `TP-<Partition ID>-...` IDs (each-value per Domain value, OMITTED only for optional axes, exactly one NOT-IN-SET with intent=enum-invalid). Before returning, derive the complete exact slot set from every legal Scenario Partitions row and compare it with both the binding Coverage Matrix Rule's Required Test Points and the final Case `### 测试点`/`变体测试点` sets; directly add every missing exact slot to the already-assigned Case IDs; ordinary alias Test Points do not satisfy a partition slot, and aggregate aliases such as `SINGLE`/`MULTIPLE` are forbidden. Directly add missing slot rows/Cases. Record an illegal plan Partition row that has no source-backed finite domain as GAP/CONFLICT and remove only its derived `TP-SP-*` slots/Cases from target modules; never modify the immutable plan artifact. Never delete a legal source-backed partition or drop its complement slot to force coverage green. When the bound source does not document the complement expectation, keep the slot with GAP expected instead of guessing. Body-field validation enums (`TP-<FIELD>-ENUM-*`) are NOT partitions — do not add partition rows for them.\n\nBefore returning, verify that every explicit source AC/REQ/BR, error code and strong DTO field token appears in the run-owned plan or an applicable module Case. If a fact cannot be safely automated, retain it as GAP/CONFLICT with its exact source pointer instead of dropping it. Return already-satisfied only when no target file needs an incremental edit.\n\nRead only indexed source paths. Do not scan the repository, modify source/**, generate pytest, execute tests, or emit JSON.\n\n## Derived task contract: 需求.md\n\n# Backend test\n- AC-001 proof\n\n## Authoritative reference index\n\n[]\n\nFor each index entry, use `readPath` for Pi read-tool calls and keep `path` as the exact Markdown Source References citation. Bound files under .harness/tasks/<taskId>/source/** are read-only inputs: reading them is allowed even though writing .harness/** is forbidden. Never resolve `path` relative to the repository root, search for substitutes, or fall back to docs/** when a bound read fails.",
|
|
295
295
|
"executor": "pi",
|
|
296
296
|
"role": "implementer",
|
|
297
297
|
"toolProfile": "write",
|
|
@@ -383,18 +383,18 @@
|
|
|
383
383
|
"generate-backend-pytest-plan-pi"
|
|
384
384
|
],
|
|
385
385
|
"complexity": "LOW",
|
|
386
|
-
"subtask_prompt": "Parse only the run-owned generate-backend-md-plan-pi/plan.md artifact
|
|
386
|
+
"subtask_prompt": "Parse only the final run-owned generate-backend-md-plan-pi/plan.md artifact with the same strict module-layout and output-budget contract; do not search for or fall back to testcase/**/README.md. No project file writes.",
|
|
387
387
|
"executor": "shell",
|
|
388
388
|
"role": "verifier",
|
|
389
389
|
"writePolicy": "read-only",
|
|
390
390
|
"shell": {
|
|
391
391
|
"commands": [
|
|
392
|
-
"node -e \"eval(Buffer.from('
|
|
392
|
+
"node -e \"eval(require('zlib').inflateRawSync(Buffer.from('pRprc9s28nt+Bd1kDLIiKTu95lKptCevtr5L4oydtHMnKR6YXEm48FUAsq2I+u83CxAkqIedTr/YBLBY7HsXC8VFLqQzFRGHPxeMg0umgnh+SeW8ncIR8fyYL0tZtNN6TDz/a8qu22kcEW/4KFaos+Q149F3EoSMqYB+lnxnlkTMWSk7y+2a5CyWb+myWMgoX6TproXLOX3647POMog4EtGJCDmUKY3B7Y/Go/FktXa973unn59U4/GkP/PJePzksCWS371LPiDLOdw6FzB7c1e6IGJXEe/1yHjcd0cvgv/S4OtR8NNVGEx63ngcZgnxycxCs8iRm5IXMQgRQn4T/vbi4v2by8ur1y9+vbr49P7q9dlFVREyfMSm7oGG91Zmg5AJcB7ecibBJRkTguUzZwcKZ1pw55rGXyBPApSd847yL0lxmztlSnOHcsmmNJbjnHjDhpw7Jt2n3nBdU4uQimlUb/i/guWuJsgnM8iBUwmBOSNLAgQPSkZ8gl/Iu6e5mCrUQorLZR67Bqv3EFvfQP7AIT2Dr0f2MJOCdDjQJINoimqnyS8shQ4tPlnI6fNWT3nBM7SSS8lZPnOFF8ribXEL/BUV4HqW7XzWGp/00GauiLX0+apXXfWe4Lw9fWUgzVnXMtLHhFNeZK/mlL8qEnB/eubZFl2+pPEXyeIvQluvKFMm3Wvpab2QFh/Lb2jKkgugosgjTm+jk5XBEyFjLqe33pBN3f7ncnQUPJ086YcoY1dIz+MgFzx3SMlZwZlcBkWeLoOsSBYpBEJCRvTOEQ2mE/yDzK+e+c9+WO9CU5T0zwUEcyrmW0iEjKKIaM2QZgcHAfwGki3wA3XoVzxU+djk+10H1twHYplLelcT654OrqGSZUXjisOf1TX3RlfBpM+292OMCVL2Bbrn18sqkKyNpNVJWsIdoSsBR1En7PC7C0jfsvxL1B+PRp/Hk0lvPBm743BX4AizZOz1Z2YvTXEnYAhGWbW2NObjXMWq3LawMW/mtJVox6iRzYEmLJ+JaDQZTgvuonew6GjIfjbHhCnkMzkfsl7PW7Gpa+ZHbBJKzjIXeSOPHzvvlIScszyBO+IZxGG5EHOXoeexqdvMaqQHUXS8x+83QaMoOjo1wcBog6mzBiRZlCmLMQB1FrweGTp1nnHgjsYyXTpFXn87GzQ7AmLJivzeICgk5TIytI2OJr3jIcoM8iTalFgrT7XrHpn2Pz9+PBa98WVtwtsy9rwVHsGG1xzol+G6oUfT3J4tUhaDqw70IU9MQLB1jiY6mphRqmypxmObSDhlqQTuptFJGrI8ThcJCJdUxGswxZCmIkIM0Qn+NdsrNDZFyLEfHHthRkv3hqYLiE660UvPeobJJivflRBLSH4DmgCPRqTW0yV6n09eLgQSLZwLEMWCx0B8cn6bQ+Kcl5iHWJGLZuoCN/4bljiD8do5e42fl0ipoz2U+KRJKBj/iU8+LFWeUaOJ7SzAlbEopkU4ZXmixq4WQh1aeXEbKdmoac/Eiy5bIdwAX7paBL6yWC864cXtSH1PoijS0hmudea0jv/5aF+6NBHPdoRA7xx0fcG5NmIsbnPgYs5Kh+aJg+ndiYt0keXiXl9IqKQXxa2oZaEVbhHZe6o1ryTRmBPKBrmsfeAken54iEwfTcz/gyjqKNyzIl9xC4lSnSq9LkG6IwJ36P1MBgsBPEhVqYdlB2cZ5cvAsBnw1lpQLCXkCeRy53qxkOUCl5IZSDLxutXkm1qPSJ1oCHHtUvPw0B6FWhuiqkYTLRMmIYtO8G+IKaU1/cJY8LlSisL+jpauN6yFDnFKOSRaQCpuY4zB0uUCSsr4i7JMGSTRlKYCho8wCDVG6RTTRmve6pHjGK4gq2sBpQDPN6KoaxE1fzzBStT4qt8QKmygpzVQHQiGpOP86p9xdmMQL4siBZp7vtpjI/uxc+LwkeOoOsGW8sEOjYRzKlwt1dW2XCRfAMpSsnwBw7VGemAYfqAK3TYWrDnxrL31psKvig4OoixyARV+VBxmHIRgRV6VhWCS3UCVw4yqj+tikSeULyvgvOBVXGTXRaXl5bmng9GVE0yqJ54pWRrq95GPUEG54GUh4D4enODEIT0zfz9PB7Y7apnj114aNiKTAg64CsCbBKi1+09vzc8qEB7SHga6JGi2foPuzKX0HZXxHN0tDEPLQp/VFprh8os0dc3N0NNumdVpRd2bGiwq1tmj0dEErVs5sko8Ct4655/1Oe11AkvFfvcuMR5jpdcnxlOaE6xqq6qUn3cOP554B1GkHGaf/LTKDDcBpoggY0Kx/W0egPW6Vdj2g0lvHJZLU7S3bHtV1Q6suiMMybealt6vqNy0rBb1g0qvk2USPRzXVR2wGdOjWqa1Ngw6FXdN0NFToW0moardxB9MznU/oYcqraoW2hKPBdt0R2r43ci78rSg/rLM352//vT2zdXbF/85//Tx6tX5+1/enr36OKhTpKoihFMspGAJOFNefIXc4UUhhWpGPGgy2oIt0g+iaCdLtrnYMO0scqkDmr1uoqCpIavqX5fn70OhnI5Nl1aEQffYWGzQqJDSlp1eVangZZ+kJnSluTchaZY74WInt0MrPuxgdWgkt+twvKd1C5vaFo3Fl7wopraEEOgDTtbFRwlYkuzxAJXMLR+wth8ebs6EM14syrMkiqLSHnvDmiL0D7VSVRoAhGQZlZCcKxawMyJ+rjdn9M6a/QD8D7RRXlXv0J5jYKm7H0lfAWX0zj32FYN1uPS8kwewa2hRZLCH75DlpWk7HhhOrbmq2tpRU1jw34FjcdBs21zY65hn739/8fbs9dX5p48fPn28evnp9a9vPl59uDg//+UbQrXj4N82CirGVrjJt+3Qb23O3/SkgbGg073ONjAe6W/4z469mxCtX+rNeMHD+93Aypc//I0yFHs+Z0kH3T/+blWrHXAHc/aq+vbDMGy89/BwlzOerqzBYBfEerBae2t0pY3rQt2NMWmpvRw0QsUrgjLKzcBWX231XTHauKaEM5BtwPTwnjPUkPpEbUCKwYFOkC3fa2+4iU3Y2HyNCM1T9ZDs8GNdYuoa7xui08GmUJQH01guaBqd6P8mg3duaDon6XNMkPiL2VH1q01L26jOqcl0SM9g33VHrFs5vu5rJSCBZyxnQrLY+YDddmzOOlwlFyemeV6o9i/kUjVB+qr70dz17yvWtpJUnSi0pCW9TiEakcohvY2WhqbQqRyk0amIT6otGGTN9aITEgQBMUxVuKEiaP2b2mlFoU4cNQLx6+DZDS3+LuutTxk6xLPWTeTYWq1DwNa8ZbU+GZFeQ0qPTNyw35kIs8Rz6plO5FQzbficbAnNm3Sadij4+sECbx8bnb4jv9OM9HylHX8bEJuB24inQlvs9huIPrJ9CsGksLbaCaNNB51g6NjwZNXdrANB3e85Pjw8qGd0A0wNohP1T0sYC5WdTZ2q2oDaKGf2eWNxAzylZanugIbC+rYgMCk2k+aeUFOIttelb9sRM+AzcK6XbT/NJDen4M5CgINPVAs5xxcUdb13djG32x9VxDvYCnmtGrD3hoJXPbiOm15TDtFGu1W3I7XmGxQZ7kfoHddYb8XprY7hGV4Uh6Y62Nhu+scWhvqFYxeKdfdxyjw+1PcuxIdvJqaZqovXjTcV9WSll7x6yUo1A93NwkZWnWKUIA2gid3toeiUKvVpCG+nHSnP1Rh7ZOB03H2oNY0PBNcpbNtCPykyynLVa7v/hQEgb3qKpuuXtd2+R9tEb0RM5XSKDbdNHtgqA8hVn8ZKais1SRMbdJhtFwvaEM3Cwx0XyEq57DSh2+4zlU4KFGuJHLbkpbfskdAWCSfP9xJwh5PspnkSiotFrl6Iuxh6xDlxnlteLOfgiIymKejnnvu1CQ+/mV8ATdp3cw6pigGu2RLfJq7nN8/gdY2pYAWUJtb02y44gta/ZNC/qAhjDlTCb1TMXSLUCvHCRZlQCa4O416YsBn2Wcgc38WGjyyhYZ2khbZxx12ZAJltZWJ3FYYhfvo2h2vP64z9llRf49AhbLAS8RwyepYMiP2sb1SlA2KAT/oiuDkmfn1psMPgKZFUfAmwkctpLMlA/cwgSICzG0iIbxb08YPtn4EgRV2U270dqCqy1S1tXhisKklKNPhXysS2iqfT48GRBXy+kHGRwQ44Es9pPoOEDEheYG9Ul4dkvcbq8/8=','base64')).toString('utf8'))\""
|
|
393
393
|
],
|
|
394
394
|
"timeoutMs": 60000,
|
|
395
395
|
"cwd": "."
|
|
396
396
|
},
|
|
397
|
-
"outputContract": "Stdout JSON {modules:[{stem,planReadPath}],planReadPath,planSha256} parsed from the run-owned Markdown plan artifact,
|
|
397
|
+
"outputContract": "Stdout JSON {modules:[{stem,planReadPath}],planReadPath,planSha256,moduleLayout} parsed from the final strict-layout-validated run-owned Markdown plan artifact, matching the Markdown map manifest.",
|
|
398
398
|
"allowedPaths": [
|
|
399
399
|
"testcase/**",
|
|
400
400
|
"docs/test-reports/**"
|
|
@@ -441,12 +441,12 @@
|
|
|
441
441
|
"verification-before-completion"
|
|
442
442
|
],
|
|
443
443
|
"complexity": "MED",
|
|
444
|
-
"subtaskPromptTemplate": "Convert the single Markdown module
|
|
445
|
-
"outputContract": "Write exactly
|
|
444
|
+
"subtaskPromptTemplate": "Convert the single frozen Markdown module `{{item.markdownPath}}` into one self-contained pytest module. Before writing, also read the run-owned Markdown plan artifact at `{{item.planReadPath}}` and use its explicit API target/environment table as the authoritative fallback base URL for every module. A task/Markdown `API_BASE_URL` target takes precedence over project README dev-server URLs; never infer a backend API fallback from a frontend/Vite port such as localhost:3000. After reading the module Markdown, the run-owned plan artifact, and the bounded pytest config/conftest, immediately use write tools to create the single frozen file `{{item.pytestPath}}`. Define any bounded HTTP client fixture, request logging/redaction/truncation helper and payload builders needed by this module inside that same file; do not import generated testcase/**/helpers/** or testcase/**/factories/** assets. Do not end after analysis or planning. Do not modify Markdown, conftest, helpers/factories, or any other module's pytest script.\n\nOutput budget protocol (hard, max output <=16K per turn): Write exactly the frozen `{{item.pytestPath}}`. Never paste full Python modules into assistant chat. Do not merge or split modules. Do not reduce params/assertions/skips to fit. If OUTPUT_LIMIT_RECOVERY is injected, continue only listed missing/broken scripts.\n\nAlign every variant pytest.param payload with the Markdown scenario intent (empty/missing/null/length/pattern/enum/wrong-type/nominal). Prefer literal payloads over Faker for intent-critical fields so pre-execution scenario-param checks can verify them. Hard contract: intent=enum-invalid MUST pass a concrete invalid value literal (string/number/boolean), never `_OMIT`/None/missing key; intent=missing/empty may use `_OMIT` or delete the key; intent=custom-literal:trim|whitespace-padded requires a leading/trailing whitespace string with non-empty trimmed content (all-whitespace belongs to empty/whitespace-only, not trim); intent=custom-literal:ACTIVE|ARCHIVED requires the exact enum string, never descriptive tokens like filter-active; intent=max/min/max+1 should pass a repeated-string length expression, a bare length number N, or a helper named _*_LEN{N} / _*_MAX_LENGTH / _*_OVER_LENGTH — never a bare 1 for oversize. Hard contract: request payload dicts may only contain DTO field keys from Payload Allowed Paths; never put expect/expected/echo_* helper keys inside the JSON body dict. Path/query/header identifiers and scenario-control metadata (including `id`, expected codes, and selector labels) must stay in separate pytest parameters and helper arguments; never merge them into a DTO patch or JSON body unless that exact path is allowed by the Markdown payload contract. Normalize the configured API base URL with `rstrip(\"/\")` (or equivalently join exactly one slash) before appending endpoint paths; generated requests must never contain a `//api/...` path. When the bound source documents a concrete non-secret local API URL, generated clients must use it as the fallback in `os.environ.get(\"API_BASE_URL\", \"<documented-url>\")`; do not require an otherwise-uninjected environment variable or fail setup solely because it is absent. Missing-field helpers must remove keys idempotently with `payload.pop(field, None)`, never `del payload[field]`, because optional fields may already be absent.\n\nFor every response contract that requires an object or pagination envelope, first assert that each envelope/data value is a dict and that required keys exist, then index fields and assert values. Never let an incidental KeyError or list/string TypeError stand in for the explicit response-shape contract failure.\n\nEnsure every automatable final Markdown Case ID in this module appears in exactly one primary pytest test function or pytest test class method region, using the exact `primary symbol` declared by Markdown. Skip evidence-only meta Cases that declare `脚本/primary symbol=无` with empty variants; do not invent a business pytest symbol for them. The symbol must start with `test_BE_<MODULE>_<NNN>_` so every parameterized collected item remains associated with its Case. Module-level functions and class-based pytest methods are both supported. Only `变体测试点` may use stable `pytest.param(..., id=\"TP-...\")` IDs, and every atomic variant ID must appear exactly once with a genuine input/state/outcome change. Use a literal direct `pytest.param(..., id=...)` expression for every row; never hide or wrap it behind `_post_case`, `_put_case`, row-factory functions, comprehensions, generators, or dynamically returned parameter lists; do not use decorator-level `ids=[...]`, generated suffixes, or IDs that extend/shorten the exact Markdown TP. Do not parameterize `场景断言测试点` or `横切证据测试点`; execute all assertion checkpoints within the same business journey/item and use shared helpers for cross-cutting evidence. The primary symbol docstring must contain exact metadata lines `Case-ID: BE-...`, `Assertion-Test-Points: TP-...;TP-...` and `Cross-Cutting-Test-Points: TP-...;TP-...` (use `none` when empty). Implement request dictionaries so their direct and nested key paths and enum literals exactly satisfy the Case `Payload Required Paths`, `Payload Allowed Paths`, and `Payload Enum`; for `Payload Contract: none`, do not invent a JSON/body DTO. GET/DELETE setup journeys may create resources, but their setup DTO must not change the target operation's no-body payload contract. No Test Point may be invented, renamed, omitted or bound in two modes. The generated pytest collection shape must equal the Markdown prediction `sum(max(1, variant count per Case))`; keep it at or below the task's explicit budget by removing duplicate execution, never by collapsing multiple parameter rows under a coarse family TP. Assertions come only from 预期结果 and setup comes only from 前置条件/测试数据/自动化映射.\n\nName the generated pytest file so it corresponds one-to-one with its source Markdown module file: this module stem `{{item.stem}}` maps to exactly the frozen `{{item.pytestPath}}`. The <module> stem is the Markdown filename without the `.md` extension, lowercased and with non-alphanumeric characters replaced by underscores. For example, `resource_notes` → `testcase/test_resource_notes.py`, `health` → `testcase/test_health.py`. If Markdown automation mapping names a different path than this module stem path, still write the frozen manifest pytest path and do not invent prefixes. Never merge multiple Markdown modules into one pytest file, never split one module across several files, and never invent pytest filenames unrelated to the Markdown modules.\n\nScenario Partition slots: every `TP-<Partition ID>-...` variant Test Point declared by this module's Markdown MUST become exactly one literal direct `pytest.param(..., id=\"TP-<Partition ID>-...\")` row with the exact slot ID; the not-in-set slot passes a concrete literal absent from the documented Domain (e.g. `UNKNOWN_TYPE`) — never `_OMIT`, never a descriptive token. Never split one slot into multiple params or merge several slots under a family TP id. Slot filtering requests hit the documented list endpoint with the slot value as the query/path filter.\n\nKeep this module self-contained: define module-local fixtures and helpers directly in `{{item.pytestPath}}`, so pytest discovers every fixture dependency without external plugin registration. The request log must include method, URL/path, and request parameters (query plus JSON/body/payload summary). The response log must include status code and response result (JSON/text/body summary), and both records must be visible in pytest stdout/stderr without changing assertions. Recursively redact sensitive values and apply bounded truncation before logging.\n\nMaterialize every automatable Markdown Case exactly once as one canonical primary pytest symbol. Preserve every explicit variant Test Point as a stable pytest.param id and every assertion/cross-cutting binding as declared. Build request payloads from the effective Markdown test data literally: keep all declared DTO keys, nested shapes, enum values, missing/null/boundary variants and business-state preconditions; never substitute guessed convenience fields or rename contract fields. Never assert an identifier's concrete Python/JSON type unless the Markdown or bound contract explicitly declares that type; when only presence is required, accept any non-null scalar identifier and serialize it safely into the path. For a nonexistent-resource 404 Case whose identifier syntax/type is not declared, obtain a syntactically valid identifier from a live create response and delete it before the 404 request; never invent an arbitrary UUID/text identifier that may fail path conversion with 400. Respect every local helper's actual return signature: never tuple-unpack a scalar status/id/helper result, and never treat a tuple response as a scalar.\n\nDo not read source/**, add cases, reassign ACs, modify conftest/config/production code, use skip/xfail, swallow assertions, execute pytest, or emit JSON. For best-effort cleanup, catch only the narrow transport exception actually raised by the selected HTTP client (for example `requests.RequestException` or `urllib.error.URLError`); never use bare `except`, `Exception`, or `BaseException` with `pass`.",
|
|
445
|
+
"outputContract": "Write exactly the frozen pytest module file `{{item.pytestPath}}` whose actual test function region contains the exact Case ID, preferably in the function name or docstring. the frozen `{{item.markdownPath}}` maps one-to-one to `{{item.pytestPath}}`; never merge or split modules. No JSON and no pytest execution.",
|
|
446
446
|
"toolProfile": "write",
|
|
447
447
|
"writePolicy": "exclusive",
|
|
448
448
|
"allowedPaths": [
|
|
449
|
-
"
|
|
449
|
+
"{{item.pytestPath}}"
|
|
450
450
|
],
|
|
451
451
|
"forbiddenPaths": [
|
|
452
452
|
".harness/**",
|
|
@@ -459,7 +459,7 @@
|
|
|
459
459
|
"setup.cfg"
|
|
460
460
|
],
|
|
461
461
|
"writeSet": [
|
|
462
|
-
"
|
|
462
|
+
"{{item.pytestPath}}"
|
|
463
463
|
],
|
|
464
464
|
"writerOutcomePolicy": {
|
|
465
465
|
"type": "implementation-outcome-v1",
|