@tea-agent/loop-agent 0.44.0-next.11 → 0.44.0-next.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,10 @@
4
4
 
5
5
  ### 新增
6
6
 
7
+ - 前端 DAG 的行为验证不再只认 Vitest:测试框架改为可插拔适配层,首批支持 Vitest、Jest 与 Playwright Test。三类项目都能用**真实报告**证明契约绑定的测试文件确实执行且通过,`npm`/`pnpm`/`yarn` 的常用测试入口走同一套识别逻辑,直接命令与 `package.json` 脚本不会再得出不同结论。
8
+ - 未适配的测试框架(Cypress、`node --test`、Mocha、自定义脚本等)现在明确报「缺少适配能力」,并附上具体命令与原因;此前这类命令会被判为不支持,错误信息还暗示你「必须改用 Vitest」,把工具的能力缺口显示成了用户配置错误。
9
+ - 行为验证保留「测试确实执行过」的校验:零测试、全部跳过、报告缺失、报告过期、契约绑定文件未执行都判失败,不会退回只看退出码(全部跳过的运行退出码是 0)。
10
+
7
11
  - main 上的 npm `next` 预发布改为跑默认测试(fast + integration),不再在每次发布时重复跑约 6 分钟的 I/O-heavy 池;完整三池仍由 PR、正式 release 线、本地推送前检查与发布前钩子承担。同时把该预发布任务的 I/O 并发从 1 调到 2,并把 npm 包可见性确认的等待窗口从约 3 分钟拉长到约 10 分钟,避免包已发出却因仓库异步处理未完成而被判失败。
8
12
 
9
13
  - 修复 main 分支的 next 发布链在类型检查阶段内存溢出失败的问题:完整源码与测试的类型图已超过 Node 默认约 2 GiB 堆上限,类型检查进程因此中止(exit 134),导致 `next` 发布无法完成。此前该上限只在部分验证入口生效,其余入口直接运行类型检查就会因默认堆不足而失败;现把上限固化到类型检查命令本身,各验证入口与本地开发统一生效,不再依赖各入口自行声明。
@@ -137,6 +141,7 @@
137
141
 
138
142
  ### 改进
139
143
 
144
+ - DAG 运行时内部按职责拆开:模型配置、领域工具、分段会话、模板族、终态判定与检查点各自独立维护,测试也按行为分成可单独调度的文件。命令、任务流程与用户可见结果不变。本机默认测试并行下整池耗时未证实变快。
140
145
  - Operator Chat 输入框下方的会话统计对齐 DSH:长串静态文字收敛为两个可点击摘要,可分别查看模型/工具用时、TTFT、TPS 与 Token 分项;详情卡互斥并支持外部点击或 Escape 关闭,窄屏打开时不会推动输入框或消息流。
141
146
  - Operator Chat 长会话新增完整轮次轨道与按需历史定位,可从任意用户消息精确新建会话;历史图片刷新后仍可预览,仓库文件链接可直接定位到行,最终回答显示单轮 token/耗时,并支持可持久化的紧凑/完整阅读模式。图片文件与消息记录分离保存,删除、分叉和失败回滚会同步清理,附件读取重新校验类型与引用。
142
147
  - 本地分支合并默认只执行与影响面匹配的定向测试、类型检查、构建和治理检查;完整测试推迟到 push 或完整交付前,并通过 receipt、`verify:tree` 或 pre-push 在最终 tree 上只执行一次。发布与推送门禁仍包含 I/O-heavy 测试。
@@ -149,6 +154,12 @@
149
154
 
150
155
  ### 修复
151
156
 
157
+ - 前端 DAG 的验证失败不再把配置问题算到业务代码头上:报告缺失或过期、报告协议错误、命令或脚本哈希漂移、框架未适配都归为验证配置问题,不再触发一轮无意义的实现重写;测试断言失败仍归实现问题。
158
+ - 已经自带 `--reporter`/`--outputFile` 的测试命令或脚本不会被悄悄覆盖;无法安全注入报告参数时在生成阶段直接报错,要求改用可注入的入口。
159
+ - 修复 Playwright 项目在 Review 阶段被误判失败的问题:审查重读报告时没有带上协议,仍按 Vitest 解析,导致 Playwright 报告因缺少 `testResults` 报错,即使执行与 trace 已经通过。report 引用现在记录协议,review 按同一协议解析并校验一致性。
160
+ - 修复 Playwright 报告路径按 `testDir` 解析错误的问题:报告中的文件路径相对 Playwright 自身的 `config.rootDir`,此前按进程 cwd 解析,配了 `testDir: "./tests"` 的项目会得到 `orders.spec.js` 而非 `tests/orders.spec.js`,契约目标永远无法匹配。
161
+ - 修复新建 Jest/Playwright 项目被误报“绑定漂移”的问题:`deferred` 观测的协议只是占位值,此前无条件参与比较,writer 一补上合法测试脚本就触发 `FRONTEND_TEST_BINDING_DRIFT`;现在只有已解析出 adapter 的观测才比较协议,真实脚本改动仍会被拦截。
162
+
152
163
  - 修复 macOS/Linux 开发路径在 Windows 暴露的三类兼容问题:前端组合验证夹具不再依赖 POSIX 单引号;无 `.gitmodules` 的仓库不再把 Git 空结果误判为启动失败;定时任务用 Git 规范化内容生成稳定树指纹,CRLF checkout 不再造成虚假 evidence drift。托管任务路径与嵌套 submodule 测试也统一使用跨平台路径语义。
153
164
  - 切换会话后,旧会话尚未确认的输入不会再出现在新会话中。
154
165
  - 修复 Windows 上 Pi 节点在没有提交改动时误报“文件在打开期间被替换”的问题,同时保留原有文件安全检查。
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "schemaVersion": 1,
3
- "version": "0.44.0-next.11",
4
- "gitSha": "332db6a894b640eef5c495e3a3ba92b383fcf259",
5
- "builtAt": "2026-09-17T05:51:06.283Z"
3
+ "version": "0.44.0-next.12",
4
+ "gitSha": "91181001a8cbed75a00a8a6d8e76ae5fab889b36",
5
+ "builtAt": "2026-09-17T10:04:11.141Z"
6
6
  }
@@ -1,7 +1,7 @@
1
1
  import { collectExistingFrontendVerificationPaths } from "../workflows/dag/frontend-implementation-contract.js";
2
2
  import { spawn } from "node:child_process";
3
3
  import { createHash, randomUUID } from "node:crypto";
4
- import { inspectFrontendTestCommand, compareFrontendTestObservation, parseFrontendTestExecutionReport } from "../workflows/dag/frontend-test-execution-evidence.js";
4
+ import { FRONTEND_TEST_OBSERVATION_FAILURES, inspectFrontendTestCommand, compareFrontendTestObservation, parseFrontendTestExecutionReport } from "../workflows/dag/frontend-test-execution-evidence.js";
5
5
  import { appendFileSync, existsSync, mkdirSync, writeFileSync } from "node:fs";
6
6
  import { mkdir, readFile, realpath, stat, writeFile } from "node:fs/promises";
7
7
  import path from "node:path";
@@ -135,6 +135,13 @@ function shouldReuseSuccessfulShellCommands(shell) {
135
135
  return shell.verifyEvidence?.phase === "intermediate" || shell.verifyEvidence?.phase === "final";
136
136
  }
137
137
  const DEFAULT_SHELL_TIMEOUT_MS = 300_000;
138
+ /**
139
+ * Tolerance for filesystem mtime granularity when rejecting a test report as
140
+ * stale. Reports are written by the command we just ran, so the only legitimate
141
+ * skew is coarse mtime resolution (1s on some filesystems) — not a real clock
142
+ * difference.
143
+ */
144
+ const REPORT_FRESHNESS_SKEW_MS = 2_000;
138
145
  const SUMMARY_STDOUT_MAX = 4_000;
139
146
  const SUMMARY_STDERR_MAX = 2_000;
140
147
  const TEST_TIMEOUT_CLASSIFY_MAX_CHARS = 8_000;
@@ -312,7 +319,10 @@ export async function executeShellCommand(input) {
312
319
  : undefined;
313
320
  const child = spawn(resolveBashExecutable(), ["-c", input.command], {
314
321
  cwd: input.cwd,
315
- env: buildShellProcessEnv(input.envAllowlist, injectedEnv),
322
+ env: buildShellProcessEnv(input.envAllowlist, {
323
+ ...injectedEnv,
324
+ ...input.extraEnv,
325
+ }),
316
326
  ...processTreeSpawnOptions(),
317
327
  stdio: ["ignore", "pipe", "pipe"],
318
328
  });
@@ -2140,16 +2150,20 @@ async function executeFrontendVerificationBundle(input, meta) {
2140
2150
  let result = cached ? reuseCommandResult(cached, command) : undefined;
2141
2151
  if (!result) {
2142
2152
  const commandNumber = results.length + 1;
2143
- const reportPath = path.join(meta.runDir, input.task.id, "commands", `${commandNumber}-${randomUUID()}.vitest.json`);
2144
2153
  const observedTest = testPreflight.inspected.get(command);
2145
- const observedCommand = observedTest?.withReporter?.(reportPath);
2146
- if (observedCommand)
2154
+ // Name the report after the resolved protocol so an operator can
2155
+ // tell at a glance which runner produced it.
2156
+ const reportScheme = observedTest?.observation.scheme ?? "vitest-json-v1";
2157
+ const reportPath = path.join(meta.runDir, input.task.id, "commands", `${commandNumber}-${randomUUID()}.${reportScheme}.json`);
2158
+ const observedInvocation = observedTest?.withReporter?.(reportPath);
2159
+ if (observedInvocation)
2147
2160
  await mkdir(path.dirname(reportPath), { recursive: true });
2148
2161
  result = await executeShellCommand({
2149
- command: observedCommand ?? command,
2162
+ command: observedInvocation?.command ?? command,
2150
2163
  cwd,
2151
2164
  timeoutMs: shell.timeoutMs ?? DEFAULT_SHELL_TIMEOUT_MS,
2152
2165
  envAllowlist: shell.envAllowlist,
2166
+ extraEnv: observedInvocation?.env,
2153
2167
  dagRunMeta: { runDir: meta.runDir, runId: meta.runId },
2154
2168
  reportActivity: input.reportActivity,
2155
2169
  outputArtifacts: {
@@ -2158,21 +2172,35 @@ async function executeFrontendVerificationBundle(input, meta) {
2158
2172
  },
2159
2173
  });
2160
2174
  let reportError;
2161
- if (observedCommand && result.ok) {
2175
+ if (observedInvocation && result.ok) {
2162
2176
  try {
2163
- // Vitest reports physical paths (e.g. /private/var on macOS).
2177
+ // Runners report physical paths (e.g. /private/var on macOS).
2164
2178
  // Match those in workspace coordinates even when shell.cwd is
2165
2179
  // a subdirectory or the caller used a directory alias.
2180
+ const reportStat = await stat(reportPath);
2181
+ // A stale report from an earlier run would otherwise let a
2182
+ // command that wrote nothing at all look verified.
2183
+ if (reportStat.mtimeMs < started - REPORT_FRESHNESS_SKEW_MS) {
2184
+ throw new Error(`stale test execution report (${FRONTEND_TEST_OBSERVATION_FAILURES.reportProtocol}): report predates this command`);
2185
+ }
2166
2186
  const rawReport = await readFile(reportPath, "utf8");
2167
2187
  const reportSha256 = createHash("sha256").update(rawReport).digest("hex");
2168
- commandExecutionEvidence.set(key, parseFrontendTestExecutionReport(JSON.parse(rawReport), await realpath(observedTest?.effectiveCwd ?? cwd), await realpath(input.cwd)));
2188
+ const parsedExecutions = parseFrontendTestExecutionReport(JSON.parse(rawReport), await realpath(observedTest?.effectiveCwd ?? cwd), await realpath(input.cwd), observedTest.observation.scheme);
2189
+ // Zero tests is not evidence: the runner may have matched no
2190
+ // file at all. The trace gate also rejects a target file with
2191
+ // no matching record, but failing here keeps the report itself
2192
+ // from being recorded as a successful observation.
2193
+ if (parsedExecutions.length === 0) {
2194
+ throw new Error(`test execution report contains no test cases (${FRONTEND_TEST_OBSERVATION_FAILURES.reportProtocol})`);
2195
+ }
2196
+ commandExecutionEvidence.set(key, parsedExecutions);
2169
2197
  const actual = commandExecutionEvidence.get(key);
2170
2198
  for (let offset = 0; offset < actual.length; offset += 24) {
2171
2199
  const relative = `${input.task.id}/commands/${commandNumber}-execution-${offset / 24}.json`;
2172
- const excerpt = { schemaVersion: 1, commandSha256: observedTest.observation.commandSha256, cwd: observedTest.observation.cwd, reportSha256, offset, total: actual.length, behaviorCoverage: "unconfirmed", tests: actual.slice(offset, offset + 24).map(test => ({ file: test.file, status: test.status, title: test.title.slice(0, 500), ...(test.title.length > 500 ? { titleTruncated: true } : {}) })) };
2200
+ const excerpt = { schemaVersion: 1, scheme: observedTest.observation.scheme, commandSha256: observedTest.observation.commandSha256, cwd: observedTest.observation.cwd, reportSha256, offset, total: actual.length, behaviorCoverage: "unconfirmed", tests: actual.slice(offset, offset + 24).map(test => ({ file: test.file, status: test.status, title: test.title.slice(0, 500), ...(test.title.length > 500 ? { titleTruncated: true } : {}) })) };
2173
2201
  await writeDagRunJsonArtifact(meta.runDir, relative, excerpt);
2174
2202
  const bytes = await readFile(path.join(meta.runDir, relative));
2175
- testExecutionEvidence.push({ path: relative, sha256: createHash("sha256").update(bytes).digest("hex"), reportPath: path.relative(meta.runDir, reportPath).replace(/\\/g, "/"), reportSha256 });
2203
+ testExecutionEvidence.push({ path: relative, sha256: createHash("sha256").update(bytes).digest("hex"), reportPath: path.relative(meta.runDir, reportPath).replace(/\\/g, "/"), reportSha256, scheme: observedTest.observation.scheme });
2176
2204
  }
2177
2205
  }
2178
2206
  catch (error) {
@@ -38,6 +38,11 @@ function contentTypeFor(pathname) {
38
38
  return "text/css; charset=utf-8";
39
39
  if (lower.endsWith(".json"))
40
40
  return "application/json; charset=utf-8";
41
+ // A web app manifest must carry this exact type: browsers ignore the
42
+ // document (and therefore offer no install entry) when it arrives as
43
+ // application/octet-stream.
44
+ if (lower.endsWith(".webmanifest"))
45
+ return "application/manifest+json";
41
46
  if (lower.endsWith(".svg"))
42
47
  return "image/svg+xml";
43
48
  if (lower.endsWith(".png"))
@@ -0,0 +1,39 @@
1
+ <svg xmlns="http://www.w3.org/2000/svg" width="64" height="64" viewBox="0 0 64 64">
2
+ <title>Loop Agent App Icon</title>
3
+ <style>
4
+ .loop-ink { stroke: #26251e; }
5
+ @media (prefers-color-scheme: dark) {
6
+ .loop-ink { stroke: #f7f7f4; }
7
+ }
8
+ </style>
9
+ <defs>
10
+ <mask id="square-handoff" maskUnits="userSpaceOnUse" x="0" y="6" width="64" height="52">
11
+ <rect x="0" y="6" width="64" height="52" fill="#ffffff" />
12
+ <path
13
+ d="m31.15 32.28 2.83 3.11"
14
+ fill="none"
15
+ stroke="#000000"
16
+ stroke-linecap="round"
17
+ stroke-width="6.23"
18
+ />
19
+ </mask>
20
+ </defs>
21
+ <g transform="translate(4.8485 13.899) scale(0.56566)">
22
+ <path
23
+ class="loop-ink"
24
+ d="M12 32c0-16 16-20 28-6l17 18c11 11 27 5 27-12S68 9 57 20L39 39C28 51 12 48 12 32Z"
25
+ fill="none"
26
+ stroke-linecap="round"
27
+ stroke-linejoin="round"
28
+ stroke-width="11"
29
+ mask="url(#square-handoff)"
30
+ />
31
+ </g>
32
+ <path
33
+ d="m31.15 32.28 2.83 3.11"
34
+ fill="none"
35
+ stroke="#f54e00"
36
+ stroke-linecap="round"
37
+ stroke-width="5.37"
38
+ />
39
+ </svg>
@@ -3,7 +3,10 @@
3
3
  <head>
4
4
  <meta charset="UTF-8" />
5
5
  <meta name="viewport" content="width=device-width, initial-scale=1.0" />
6
+ <meta name="theme-color" content="#f4f2ee" />
7
+ <link rel="manifest" href="/manifest.webmanifest" />
6
8
  <link rel="icon" type="image/svg+xml" href="/favicon.svg" />
9
+ <link rel="apple-touch-icon" href="/app-icon.svg" />
7
10
  <link rel="preconnect" href="https://cdn.jsdelivr.net" />
8
11
  <link
9
12
  rel="stylesheet"
@@ -0,0 +1,19 @@
1
+ {
2
+ "id": "/",
3
+ "name": "Loop Operator Console",
4
+ "short_name": "Loop Console",
5
+ "description": "Loop Agent 操作台:任务、Agent DAG 运行与交付证据的只读观察与受控操作界面",
6
+ "start_url": "/",
7
+ "scope": "/",
8
+ "display": "standalone",
9
+ "background_color": "#f4f2ee",
10
+ "theme_color": "#f4f2ee",
11
+ "icons": [
12
+ {
13
+ "src": "/app-icon.svg",
14
+ "sizes": "any",
15
+ "type": "image/svg+xml",
16
+ "purpose": "any"
17
+ }
18
+ ]
19
+ }
@@ -7,10 +7,46 @@ export const CONSOLE_THEMES = Object.freeze([
7
7
 
8
8
  const STORAGE_KEY = "loop-console.theme.v1";
9
9
 
10
+ /**
11
+ * Titlebar colour per theme, matching each palette's --canvas. An installed PWA
12
+ * paints the OS window frame with the theme-color meta value, so it must track
13
+ * the active theme: a fixed brand colour leaves an orange titlebar above the
14
+ * dark midnight palette. Values mirror console-theme.css --canvas.
15
+ */
16
+ const TITLEBAR_COLOR = Object.freeze({
17
+ default: "#f4f2ee",
18
+ midnight: "#1c222a",
19
+ daylight: "#eef0f2",
20
+ });
21
+
10
22
  function normalizeTheme(value) {
11
23
  return CONSOLE_THEMES.some((theme) => theme.id === value) ? value : "default";
12
24
  }
13
25
 
26
+ /**
27
+ * Project the theme onto the document's theme-color meta tag. Browsers read
28
+ * this tag live (the manifest theme_color only applies at install time), and
29
+ * an installed window repaints its frame when it changes.
30
+ *
31
+ * Strictly best-effort: the titlebar is cosmetic, so a missing or partial DOM
32
+ * (SSR, a test stub, an exotic embedder) must never break theme switching.
33
+ */
34
+ function syncTitlebarColor(theme) {
35
+ try {
36
+ const head = document?.head;
37
+ if (!head || typeof document.createElement !== "function") return;
38
+ let meta = document.querySelector?.('meta[name="theme-color"]');
39
+ if (!meta) {
40
+ meta = document.createElement("meta");
41
+ meta.setAttribute("name", "theme-color");
42
+ head.appendChild(meta);
43
+ }
44
+ meta.setAttribute("content", TITLEBAR_COLOR[theme] ?? TITLEBAR_COLOR.default);
45
+ } catch {
46
+ // The palette switch above already succeeded; the titlebar is optional.
47
+ }
48
+ }
49
+
14
50
  export function applyStoredConsoleTheme() {
15
51
  let value;
16
52
  try {
@@ -20,6 +56,7 @@ export function applyStoredConsoleTheme() {
20
56
  }
21
57
  const theme = normalizeTheme(value);
22
58
  document.documentElement.dataset.consoleTheme = theme;
59
+ syncTitlebarColor(theme);
23
60
  return theme;
24
61
  }
25
62
 
@@ -27,6 +64,7 @@ export function setConsoleTheme(value) {
27
64
  const theme = normalizeTheme(value);
28
65
  // Apply before persistence: a storage failure must not disable the picker.
29
66
  document.documentElement.dataset.consoleTheme = theme;
67
+ syncTitlebarColor(theme);
30
68
  try {
31
69
  window.localStorage.setItem(STORAGE_KEY, theme);
32
70
  } catch {
@@ -1,5 +1,19 @@
1
1
  import { z } from "zod";
2
2
  import { FRONTEND_NO_VERIFICATION_MARKER_TEXT, } from "./frontend-verification-trace.js";
3
+ import { FRONTEND_TEST_OBSERVATION_FAILURES } from "./frontend-test-execution-evidence.js";
4
+ /**
5
+ * Substrings that prove a behavior-verification failure came from the
6
+ * verification configuration rather than the candidate's code. Kept next to the
7
+ * taxonomy so a new failure code is routed deliberately instead of falling
8
+ * through to the string-blob heuristics below.
9
+ */
10
+ const FRONTEND_TEST_VERIFICATION_CONFIG_MARKERS = [
11
+ FRONTEND_TEST_OBSERVATION_FAILURES.bindingDrift,
12
+ FRONTEND_TEST_OBSERVATION_FAILURES.reportProtocol,
13
+ FRONTEND_TEST_OBSERVATION_FAILURES.capabilityGap,
14
+ "FRONTEND_TEST_OBSERVATION_INVALID",
15
+ "frontend test execution report unavailable",
16
+ ];
3
17
  export const frontendRepairFailureClassSchema = z.enum([
4
18
  "typecheck",
5
19
  "build",
@@ -73,6 +87,16 @@ export function classifyFrontendFailure(input) {
73
87
  if (input.failureCategory === "write-guard") {
74
88
  return "path";
75
89
  }
90
+ // Verification-config faults must never be mistaken for product defects.
91
+ // A missing/unreadable report, a stale report, an unsupported runner or a
92
+ // drifted command binding are all tooling/contract problems: rewriting the
93
+ // candidate's business code cannot fix them and burns a repair round. These
94
+ // are detected structurally, before any blob heuristic can see a runner name
95
+ // (the `unit-test` rule matches the literal word "jest") and misroute the
96
+ // failure to the implementation owner.
97
+ if (FRONTEND_TEST_VERIFICATION_CONFIG_MARKERS.some((marker) => `${input.stderr ?? ""}\n${input.stdout ?? ""}`.includes(marker))) {
98
+ return "contract";
99
+ }
76
100
  if (`${input.stderr ?? ""}\n${input.stdout ?? ""}`.includes(FRONTEND_NO_VERIFICATION_MARKER_TEXT)) {
77
101
  // Verification unavailability is a generation-time contract gap: no
78
102
  // authorized verification command exists. Bounded repair cannot fix it;
@@ -1,4 +1,5 @@
1
1
  import { parseFrontendTestExecutionReport } from "./frontend-test-execution-evidence.js";
2
+ import { isFrontendTestReportScheme } from "./frontend-test-framework-adapters.js";
2
3
  import { readFileBoundarySafe } from "../../shared/path-safety.js";
3
4
  import { createHash, randomUUID } from "node:crypto";
4
5
  import { readFile, realpath, lstat } from "node:fs/promises";
@@ -109,12 +110,16 @@ export async function loadFrontendReviewScopes(runDir, phase, targetBytes, works
109
110
  const reportCoverage = new Map();
110
111
  const reportedTests = [];
111
112
  for (const ref of payload.verificationTrace?.testExecutionEvidence ?? []) {
112
- if (!isArtifactReference(ref) || typeof ref.reportPath !== "string" || !/^[a-f0-9]{64}$/.test(ref.reportSha256))
113
+ if (!isArtifactReference(ref) || typeof ref.reportPath !== "string" || !/^[a-f0-9]{64}$/.test(ref.reportSha256) || !isFrontendTestReportScheme(ref.scheme))
113
114
  throw invalidScope("invalid test execution reference");
114
115
  const excerpt = await read(ref.path, ref.sha256);
115
116
  const data = JSON.parse(excerpt.text);
116
117
  if (data.reportSha256 !== ref.reportSha256)
117
118
  throw invalidScope("test execution report binding mismatch");
119
+ // The excerpt records the protocol that produced it. A mismatch means the
120
+ // trace and its report disagree, so refuse rather than guess an adapter.
121
+ if (data.scheme !== ref.scheme)
122
+ throw invalidScope("test execution report protocol binding mismatch");
118
123
  const coverage = reportCoverage.get(ref.reportPath) ?? { count: 0, total: data.total };
119
124
  if (!Number.isInteger(data.total) || data.total < 1 || data.total !== coverage.total || data.offset !== coverage.count || !Array.isArray(data.tests) || !data.tests.length || data.tests.length > 24)
120
125
  throw invalidScope("test report excerpt coverage invalid");
@@ -128,7 +133,10 @@ export async function loadFrontendReviewScopes(runDir, phase, targetBytes, works
128
133
  const cwdPath = await realpath(path.resolve(workspaceRoot, data.cwd));
129
134
  if (cwdPath !== rootPath && !cwdPath.startsWith(rootPath + path.sep))
130
135
  throw invalidScope("test report cwd escaped workspace");
131
- const actual = parseFrontendTestExecutionReport(JSON.parse(report.text), cwdPath, rootPath);
136
+ // Parse with the protocol the runner actually used. Defaulting to
137
+ // Vitest here failed every Playwright/Jest report at review time even
138
+ // though execution and trace had accepted it.
139
+ const actual = parseFrontendTestExecutionReport(JSON.parse(report.text), cwdPath, rootPath, ref.scheme);
132
140
  const expected = actual.slice(data.offset, data.offset + 24).map(test => ({ file: test.file, status: test.status, title: test.title.slice(0, 500), ...(test.title.length > 500 ? { titleTruncated: true } : {}) }));
133
141
  if (data.total !== actual.length || !expected.length || sha256OfCanonicalJson(expected) !== sha256OfCanonicalJson(data.tests))
134
142
  throw invalidScope("test report excerpt differs from actual runner report");
@@ -3,14 +3,32 @@ import { createHash } from "node:crypto";
3
3
  import { pathMatchesPattern } from "../../shared/git-progress.js";
4
4
  import { readFile, realpath } from "node:fs/promises";
5
5
  import { z } from "zod";
6
+ import { FRONTEND_TEST_ADAPTERS, detectUnadaptedRunner, frontendTestReportSchemeSchema, resolveFrontendTestAdapter, resolveRunnerArgv, } from "./frontend-test-framework-adapters.js";
7
+ /** Scheme → adapter lookup, built once from the adapter registry. */
8
+ const FRONTEND_TEST_ADAPTERS_BY_SCHEME = Object.fromEntries(FRONTEND_TEST_ADAPTERS.map((adapter) => [adapter.scheme, adapter]));
9
+ /**
10
+ * Failure taxonomy for behavior verification. Kept separate from the writer's
11
+ * `failureCategory` so a broken report protocol never reads as broken product
12
+ * code and trigger a pointless implementation rewrite.
13
+ */
14
+ export const FRONTEND_TEST_OBSERVATION_FAILURES = {
15
+ /** The frozen command drift from what the contract bound. */
16
+ bindingDrift: "FRONTEND_TEST_BINDING_DRIFT",
17
+ /** The runner produced no readable report, or the report was stale. */
18
+ reportProtocol: "FRONTEND_TEST_REPORT_PROTOCOL",
19
+ /** The runner itself is not adapted (Cypress, `node --test`, …). */
20
+ capabilityGap: "FRONTEND_TEST_CAPABILITY_GAP",
21
+ };
6
22
  export const frontendTestObservationSchema = z.object({
7
- scheme: z.literal("vitest-json-v1"),
23
+ scheme: frontendTestReportSchemeSchema,
8
24
  commandId: z.string().optional(),
9
25
  status: z.enum(["supported", "deferred", "unsupported"]),
10
26
  commandSha256: z.string().regex(/^[a-f0-9]{64}$/),
11
27
  cwd: z.string(),
12
28
  packageScript: z.object({ file: z.string(), name: z.string(), sha256: z.string().nullable() }).optional(),
13
29
  reason: z.string().optional(),
30
+ /** Set when `status: "unsupported"` because no adapter exists yet. */
31
+ capabilityGap: z.string().optional(),
14
32
  }).strict();
15
33
  const hash = (text) => createHash("sha256").update(text).digest("hex");
16
34
  const quote = (text) => `'${text.replace(/'/g, `'"'"'`)}'`;
@@ -124,20 +142,68 @@ function literalCommand(command, cwd, depth = 0) {
124
142
  }
125
143
  return { argv, cwd, append: suffix => command.trimEnd() + suffix };
126
144
  }
127
- function directVitest(argv) {
128
- const args = [...argv];
129
- if (["npx"].includes(args[0] ?? ""))
130
- args.shift();
131
- else if (["pnpm", "yarn"].includes(args[0] ?? "") && args[1] === "exec")
132
- args.splice(0, 2);
133
- return args[0] === "vitest" && args[1] === "run";
134
- }
135
145
  function hasReporter(argv) { return argv.some(arg => /^--(?:reporter|outputFile)(?:[.=]|$)/.test(arg)); }
136
- /** Shared read-only capability resolver, used at generation, admission and
137
- * immediately before execution. It never executes or rewrites an npm script. */
146
+ /** Locate the `npm|pnpm|yarn` script invocation in a literal argv. */
147
+ function scriptInvocation(argv) {
148
+ const manager = argv[0];
149
+ if (manager !== "npm" && manager !== "pnpm" && manager !== "yarn")
150
+ return undefined;
151
+ // `exec`/`dlx`/`create` run a binary rather than a package script, so they
152
+ // are not script invocations — `yarn exec vitest run` must resolve as a
153
+ // direct runner command, not as a script literally named "exec".
154
+ const subcommands = new Set(["exec", "dlx", "create", "init", "add", "install", "why", "config"]);
155
+ // npm test | npm run <name> | pnpm test | pnpm run <name> | yarn <name>
156
+ let name;
157
+ let rest = [];
158
+ if (argv[1] === "test") {
159
+ name = "test";
160
+ rest = argv.slice(2);
161
+ }
162
+ else if (argv[1] === "run") {
163
+ name = argv[2];
164
+ rest = argv.slice(3);
165
+ }
166
+ else if (manager === "yarn" && argv[1] && !argv[1].startsWith("-") && !subcommands.has(argv[1])) {
167
+ name = argv[1];
168
+ rest = argv.slice(2);
169
+ }
170
+ if (!name || !/^[\w:-]+$/.test(name))
171
+ return undefined;
172
+ if (rest.length && rest[0] !== "--")
173
+ throw Error("package manager script arguments require --");
174
+ return { name, rest };
175
+ }
176
+ /**
177
+ * Build the report invocation for an adapter.
178
+ *
179
+ * When the adapter names a `reportPathEnv`, the path travels in the environment
180
+ * and the suffix only selects the reporter — appending the path to the command
181
+ * would hand it to the runner as a positional argument (Playwright treats a
182
+ * stray argument as a test-name filter and then finds no tests at all).
183
+ *
184
+ * Otherwise the path is joined to the suffix: a suffix ending in `=` (Vitest's
185
+ * `--outputFile=`) binds it directly, and any other suffix names a flag
186
+ * (`--outputFile`, `--json`) that takes the path as the next argument.
187
+ */
188
+ function reportInvocation(adapter, append, suffix, reportPath) {
189
+ if (adapter.reportPathEnv) {
190
+ return {
191
+ command: append(suffix),
192
+ env: { [adapter.reportPathEnv]: reportPath },
193
+ };
194
+ }
195
+ const joiner = suffix.endsWith("=") ? "" : " ";
196
+ return { command: append(`${suffix}${joiner}${quote(reportPath)}`) };
197
+ }
198
+ /**
199
+ * Shared read-only capability resolver, used at generation, admission and
200
+ * immediately before execution. It never executes or rewrites an npm script.
201
+ */
138
202
  export async function inspectFrontendTestCommand(input) {
139
203
  const root = path.resolve(input.workspaceRoot ?? input.cwd);
140
204
  let cwd = path.resolve(input.cwd);
205
+ // `scheme` is unknowable until an adapter resolves; the placeholder is only
206
+ // ever surfaced on the unsupported path, where no report is produced.
141
207
  const base = { scheme: "vitest-json-v1", status: "unsupported", commandSha256: hash(input.command), cwd: path.relative(root, cwd).replace(/\\/g, "/") || "." };
142
208
  try {
143
209
  const parsed = literalCommand(input.command, cwd);
@@ -154,79 +220,109 @@ export async function inspectFrontendTestCommand(input) {
154
220
  throw Error("test cwd resolves outside the workspace");
155
221
  if (hasReporter(parsed.argv))
156
222
  throw Error("frozen command already owns reporter/outputFile flags");
157
- let separator = "";
158
- if (!directVitest(parsed.argv)) {
159
- const argv = parsed.argv;
160
- const name = argv[0] === "npm" ? argv[1] === "test" ? "test" : argv[1] === "run" ? argv[2] : undefined : undefined;
161
- if (!name || !/^[\w:-]+$/.test(name))
162
- throw Error("no supported test reporter adapter for this command");
163
- const rest = argv.slice(argv[1] === "test" ? 2 : 3);
164
- if (rest.length && rest[0] !== "--")
165
- throw Error("npm script arguments require --");
166
- const file = path.relative(root, path.join(cwd, "package.json")).replace(/\\/g, "/");
167
- let manifest = {};
168
- try {
169
- manifest = JSON.parse(await readFile(path.join(cwd, "package.json"), "utf8"));
170
- }
171
- catch (error) {
172
- if (error.code !== "ENOENT")
173
- throw Error("package.json is unreadable or invalid");
174
- }
175
- const script = manifest?.scripts?.[name];
176
- base.packageScript = { file, name, sha256: typeof script === "string" ? hash(JSON.stringify({ script, pre: manifest.scripts?.[`pre${name}`] ?? null, post: manifest.scripts?.[`post${name}`] ?? null })) : null };
177
- if (script === undefined) {
178
- const authorized = input.allowedPaths?.some(pattern => pathMatchesPattern(file, pattern)) && !input.forbiddenPaths?.some(pattern => pathMatchesPattern(file, pattern));
179
- if (!authorized)
180
- throw Error(`missing npm script ${name}; ${file} is not an authorized delivery`);
181
- return { observation: { ...base, status: "deferred", reason: `Awaiting authorized ${file} script ${name} with vitest-json-v1 capability` }, effectiveCwd: cwd };
182
- }
183
- if (typeof script !== "string")
184
- throw Error("npm script must be a string");
185
- const scriptWords = words(script);
186
- if (scriptWords.some(word => word.operator) || !directVitest(scriptWords.map(w => w.value)) || hasReporter(scriptWords.map(w => w.value)))
187
- throw Error("npm script has no supported Vitest JSON reporter capability");
188
- separator = rest.length ? "" : " --";
223
+ const runner = resolveRunnerArgv(parsed.argv);
224
+ const invocation = scriptInvocation(parsed.argv);
225
+ if (!invocation) {
226
+ const adapter = resolveFrontendTestAdapter({ argv: parsed.argv, runner });
227
+ if (!adapter)
228
+ throw capabilityGapError(runner);
229
+ const suffix = adapter.reportSuffix({ argv: parsed.argv, runner });
230
+ if (suffix === undefined)
231
+ throw Error(`${adapter.runner} project configuration owns report output; declare an explicit regression command instead of overriding it`);
232
+ return { observation: { ...base, scheme: adapter.scheme, status: "supported" }, effectiveCwd: cwd,
233
+ withReporter: reportPath => reportInvocation(adapter, inner => parsed.append(` ${inner}`), suffix, reportPath) };
234
+ }
235
+ const file = path.relative(root, path.join(cwd, "package.json")).replace(/\\/g, "/");
236
+ let manifest = {};
237
+ try {
238
+ manifest = JSON.parse(await readFile(path.join(cwd, "package.json"), "utf8"));
189
239
  }
190
- return { observation: { ...base, status: "supported" }, effectiveCwd: cwd,
191
- withReporter: reportPath => parsed.append(`${separator} --reporter=json --outputFile=${quote(reportPath)}`) };
240
+ catch (error) {
241
+ if (error.code !== "ENOENT")
242
+ throw Error("package.json is unreadable or invalid");
243
+ }
244
+ const script = manifest?.scripts?.[invocation.name];
245
+ base.packageScript = { file, name: invocation.name, sha256: typeof script === "string" ? hash(JSON.stringify({ script, pre: manifest.scripts?.[`pre${invocation.name}`] ?? null, post: manifest.scripts?.[`post${invocation.name}`] ?? null })) : null };
246
+ if (script === undefined) {
247
+ const authorized = input.allowedPaths?.some(pattern => pathMatchesPattern(file, pattern)) && !input.forbiddenPaths?.some(pattern => pathMatchesPattern(file, pattern));
248
+ if (!authorized)
249
+ throw Error(`missing npm script ${invocation.name}; ${file} is not an authorized delivery`);
250
+ return { observation: { ...base, status: "deferred", reason: `Awaiting authorized ${file} script ${invocation.name} with a supported test report capability` }, effectiveCwd: cwd };
251
+ }
252
+ if (typeof script !== "string")
253
+ throw Error("npm script must be a string");
254
+ const scriptWords = words(script);
255
+ if (scriptWords.some(word => word.operator))
256
+ throw Error("npm script must be a single literal command");
257
+ const scriptArgv = scriptWords.map(w => w.value);
258
+ if (hasReporter(scriptArgv))
259
+ throw Error("npm script already owns reporter/outputFile flags");
260
+ const scriptRunner = resolveRunnerArgv(scriptArgv);
261
+ const adapter = resolveFrontendTestAdapter({ argv: parsed.argv, runner: scriptRunner, scriptName: invocation.name, scriptArgv });
262
+ if (!adapter)
263
+ throw capabilityGapError(scriptRunner);
264
+ const suffix = adapter.reportSuffix({ argv: scriptArgv, runner: scriptRunner });
265
+ if (suffix === undefined)
266
+ throw Error(`${adapter.runner} project configuration owns report output; declare an explicit regression command instead of overriding it`);
267
+ // `npm test -- --json` forwards the flags to the script; a bare script
268
+ // uses the same `--` separator so npm never eats the report flags.
269
+ // Space-prefixed because `append` concatenates onto the trimmed
270
+ // command; a bare script needs `--` so npm forwards the report flags.
271
+ const separator = invocation.rest.length ? "" : " -- ";
272
+ return { observation: { ...base, scheme: adapter.scheme, status: "supported" }, effectiveCwd: cwd,
273
+ withReporter: reportPath => reportInvocation(adapter, inner => parsed.append(invocation.rest.length ? ` ${inner}` : `${separator}${inner}`), suffix, reportPath) };
192
274
  }
193
275
  catch (error) {
194
- return { observation: { ...base, status: "unsupported", reason: error instanceof Error ? error.message : String(error) }, effectiveCwd: cwd };
276
+ const gap = error.capabilityGap;
277
+ return { observation: { ...base, status: "unsupported", reason: error instanceof Error ? error.message : String(error), ...(gap ? { capabilityGap: gap } : {}) }, effectiveCwd: cwd };
195
278
  }
196
279
  }
280
+ /**
281
+ * An unadapted runner is a capability gap in this tool, not a user mistake: the
282
+ * message must never suggest switching frameworks.
283
+ */
284
+ function capabilityGapError(runner) {
285
+ const known = detectUnadaptedRunner(runner);
286
+ const binary = known ?? (runner[0] ?? "unknown");
287
+ const error = new Error(`${FRONTEND_TEST_OBSERVATION_FAILURES.capabilityGap}: no test report adapter for ${binary}; ` +
288
+ `frontend behavior verification currently adapts vitest, jest and playwright test, ` +
289
+ `so this command cannot produce execution evidence`);
290
+ error.capabilityGap = binary;
291
+ return error;
292
+ }
197
293
  export function compareFrontendTestObservation(frozen, observed, phase, writeSet = []) {
198
294
  if (frozen.status === "unsupported")
199
295
  return frozen.reason ?? "unsupported test observation scheme";
200
296
  if (observed.status === "unsupported")
201
297
  return observed.reason ?? "unsupported test observation scheme";
202
- if (frozen.scheme !== observed.scheme || frozen.commandSha256 !== observed.commandSha256 || frozen.cwd !== observed.cwd || frozen.packageScript?.file !== observed.packageScript?.file || frozen.packageScript?.name !== observed.packageScript?.name)
203
- return "frozen test command/cwd/script binding changed";
298
+ // A `deferred` observation has no resolved adapter, so its `scheme` is the
299
+ // placeholder rather than a binding. Comparing it would reject every new
300
+ // Jest/Playwright project the moment the writer adds the very script the
301
+ // contract asked for. The scheme only becomes a binding once the frozen
302
+ // observation is `supported`; script identity is still checked below.
303
+ const schemeBound = frozen.status !== "deferred";
304
+ if ((schemeBound && frozen.scheme !== observed.scheme) || frozen.commandSha256 !== observed.commandSha256 || frozen.cwd !== observed.cwd || frozen.packageScript?.file !== observed.packageScript?.file || frozen.packageScript?.name !== observed.packageScript?.name)
305
+ return `${FRONTEND_TEST_OBSERVATION_FAILURES.bindingDrift}: frozen test command/cwd/script binding changed`;
204
306
  if (frozen.status === "supported" && (observed.status !== "supported" || frozen.packageScript?.sha256 !== observed.packageScript?.sha256))
205
- return "frozen npm test script changed";
307
+ return `${FRONTEND_TEST_OBSERVATION_FAILURES.bindingDrift}: frozen npm test script changed`;
206
308
  if (frozen.status === "deferred" && phase === "admission" && (!frozen.packageScript || !writeSet.includes(frozen.packageScript.file)))
207
309
  return "future test script is outside the admitted concrete writeSet";
208
310
  if (phase === "execution" && observed.status !== "supported")
209
311
  return "future test reporter capability is still unavailable";
210
312
  return undefined;
211
313
  }
212
- /** Observe supported frozen Vitest commands without changing their selection. */
314
+ /** Observe supported frozen commands without changing their selection. */
213
315
  export async function withFrontendTestReporter(input) {
214
- return (await inspectFrontendTestCommand(input)).withReporter?.(input.reportPath);
316
+ return (await inspectFrontendTestCommand(input)).withReporter?.(input.reportPath)?.command;
215
317
  }
216
- const resultSchema = z.object({
217
- testResults: z.array(z.object({
218
- name: z.string().min(1),
219
- assertionResults: z.array(z.object({
220
- title: z.string(), fullName: z.string().optional(),
221
- ancestorTitles: z.array(z.string()).optional(), status: z.string(),
222
- })),
223
- })),
224
- });
225
- /** Resolve runner paths against its cwd, then use the contract workspace coordinates. */
226
- export function parseFrontendTestExecutionReport(raw, cwd, workspaceRoot = cwd) {
227
- return resultSchema.parse(raw).testResults.flatMap((suite) => suite.assertionResults.map((test) => ({
228
- file: path.relative(workspaceRoot, path.resolve(cwd, suite.name)).replace(/\\/g, "/"),
229
- title: test.fullName ?? [...(test.ancestorTitles ?? []), test.title].join(" "),
230
- status: test.status,
231
- })));
318
+ /**
319
+ * Parse a runner report into uniform execution records using the adapter for
320
+ * `scheme`. A malformed report throws; callers classify that as a report
321
+ * protocol failure, never as a product defect.
322
+ */
323
+ export function parseFrontendTestExecutionReport(raw, cwd, workspaceRoot = cwd, scheme = "vitest-json-v1") {
324
+ const adapter = FRONTEND_TEST_ADAPTERS_BY_SCHEME[scheme];
325
+ if (!adapter)
326
+ throw Error(`no test report adapter for scheme ${scheme}`);
327
+ return adapter.parse(raw, cwd, workspaceRoot);
232
328
  }
@@ -0,0 +1,309 @@
1
+ import { realpathSync } from "node:fs";
2
+ import path from "node:path";
3
+ import { z } from "zod";
4
+ /**
5
+ * Framework adapters decouple the frontend behavior-verification contract from
6
+ * any single test runner.
7
+ *
8
+ * The contract, writer admission and execution all need three things from a
9
+ * frozen verify command:
10
+ * 1. whether the runner can emit a machine-readable execution report at all
11
+ * (capability), so an unsupported runner fails closed *before* implement;
12
+ * 2. how to ask that runner for a report without clobbering the project's
13
+ * own reporter configuration;
14
+ * 3. how to read that report back as uniform file/title/status triples.
15
+ *
16
+ * Historically all three were hardcoded to Vitest (`vitest-json-v1`), while the
17
+ * command classifier was already framework-agnostic. Jest and Playwright
18
+ * commands were therefore classified into the behavior lane correctly and then
19
+ * rejected as "unsupported", surfacing a capability gap to the user as if they
20
+ * had misconfigured their project. Adapters make the two sources of truth agree.
21
+ */
22
+ /** Report protocols a frozen behavior command may be observed through. */
23
+ export const frontendTestReportSchemeSchema = z.enum([
24
+ "vitest-json-v1",
25
+ "jest-json-v1",
26
+ "playwright-json-v1",
27
+ ]);
28
+ /** Runtime guard for a persisted scheme, e.g. read back from a trace artifact. */
29
+ export const isFrontendTestReportScheme = (value) => typeof value === "string" &&
30
+ frontendTestReportSchemeSchema.options.includes(value);
31
+ const asRecord = (value) => typeof value === "object" && value !== null
32
+ ? value
33
+ : undefined;
34
+ /**
35
+ * Resolve a runner-relative path into contract workspace coordinates.
36
+ *
37
+ * A runner may report an absolute path (`config.rootDir` is absolute in a
38
+ * Playwright report) or one relative to its own root. Either way the result is
39
+ * expressed against `workspaceRoot`, which is resolved physically first: a
40
+ * workspace reached through a symlink (macOS `/var` -> `/private/var`, i.e. the
41
+ * default `os.tmpdir()`) is never a lexical prefix of the runner's canonical
42
+ * path, so comparing raw strings would yield a `../..` escape instead of a
43
+ * contract-relative file.
44
+ */
45
+ const toWorkspacePath = (raw, cwd, workspaceRoot) => {
46
+ const absolute = path.resolve(cwd, raw);
47
+ const physical = (target) => {
48
+ try {
49
+ return realpathSync(target);
50
+ }
51
+ catch {
52
+ return target;
53
+ }
54
+ };
55
+ const root = physical(path.resolve(workspaceRoot));
56
+ // The path may itself sit under the unresolved root; re-anchor it physically
57
+ // by resolving its root-relative tail against the physical root.
58
+ const anchored = absolute.startsWith(path.resolve(workspaceRoot) + path.sep)
59
+ ? path.join(root, path.relative(path.resolve(workspaceRoot), absolute))
60
+ : absolute;
61
+ return path.relative(root, physical(anchored)).replace(/\\/g, "/");
62
+ };
63
+ /**
64
+ * Strip package-manager indirection so `npx vitest run`, `pnpm exec vitest
65
+ * run`, `yarn exec vitest run` and a bare `vitest run` all reduce to the same
66
+ * runner invocation. Returns the remaining words.
67
+ */
68
+ export function resolveRunnerArgv(argv) {
69
+ const words = [...argv];
70
+ // Skip leading env assignments (`CI=1 vitest run`).
71
+ while (words.length > 0 && /^[A-Za-z_][A-Za-z0-9_]*=/.test(words[0])) {
72
+ words.shift();
73
+ }
74
+ const head = words[0];
75
+ if (head === "npx" || head === "pnpx" || head === "bunx") {
76
+ words.shift();
77
+ // `npx --no-install vitest run` / `npx -y vitest run`
78
+ while (words[0]?.startsWith("-"))
79
+ words.shift();
80
+ return words;
81
+ }
82
+ if ((head === "pnpm" || head === "yarn" || head === "npm") && words[1] === "exec") {
83
+ words.splice(0, 2);
84
+ while (words[0]?.startsWith("-"))
85
+ words.shift();
86
+ return words;
87
+ }
88
+ return words;
89
+ }
90
+ /** Runner binaries each adapter claims, keyed by adapter scheme. */
91
+ const VITEST_BINARIES = new Set(["vitest"]);
92
+ const JEST_BINARIES = new Set(["jest"]);
93
+ const PLAYWRIGHT_BINARIES = new Set(["playwright"]);
94
+ const basenameOf = (value) => value.split(/[\\/]/).pop() ?? "";
95
+ /** True when `runner` invokes `binary` directly (allowing an explicit path). */
96
+ const invokesBinary = (runner, binaries) => binaries.has(basenameOf(runner[0] ?? ""));
97
+ const vitestAdapter = {
98
+ scheme: "vitest-json-v1",
99
+ runner: "vitest",
100
+ matches: ({ runner }) => invokesBinary(runner, VITEST_BINARIES),
101
+ // Vitest's json reporter writes to stdout unless an output file is given,
102
+ // and the executor reads the report from disk, so both flags are required.
103
+ reportSuffix: () => `--reporter=json --outputFile=`,
104
+ parse: (raw, cwd, workspaceRoot) => {
105
+ const schema = z.object({
106
+ testResults: z.array(z.object({
107
+ name: z.string().min(1),
108
+ assertionResults: z.array(z.object({
109
+ title: z.string(),
110
+ fullName: z.string().optional(),
111
+ ancestorTitles: z.array(z.string()).optional(),
112
+ status: z.string(),
113
+ })),
114
+ })),
115
+ });
116
+ return schema
117
+ .parse(raw)
118
+ .testResults.flatMap((suite) => suite.assertionResults.map((test) => ({
119
+ file: toWorkspacePath(suite.name, cwd, workspaceRoot),
120
+ title: test.fullName ??
121
+ [...(test.ancestorTitles ?? []), test.title].join(" "),
122
+ status: test.status,
123
+ })));
124
+ },
125
+ };
126
+ const jestAdapter = {
127
+ scheme: "jest-json-v1",
128
+ runner: "jest",
129
+ matches: ({ runner }) => invokesBinary(runner, JEST_BINARIES),
130
+ reportSuffix: () => `--json --outputFile`,
131
+ parse: (raw, cwd, workspaceRoot) => {
132
+ // Jest's `--json` shape: testResults[] carry a per-file `name` plus
133
+ // assertionResults[] that mirror Vitest's field names.
134
+ const schema = z.object({
135
+ testResults: z.array(z.object({
136
+ name: z.string().min(1),
137
+ assertionResults: z.array(z.object({
138
+ title: z.string(),
139
+ fullName: z.string().optional(),
140
+ ancestorTitles: z.array(z.string()).optional(),
141
+ status: z.string(),
142
+ })),
143
+ })),
144
+ });
145
+ return schema
146
+ .parse(raw)
147
+ .testResults.flatMap((suite) => suite.assertionResults.map((test) => ({
148
+ file: toWorkspacePath(suite.name, cwd, workspaceRoot),
149
+ title: test.fullName ??
150
+ [...(test.ancestorTitles ?? []), test.title].join(" "),
151
+ status: test.status,
152
+ })));
153
+ },
154
+ };
155
+ /**
156
+ * Playwright's JSON reporter emits a nested `suites` tree whose leaf `specs`
157
+ * carry `tests[].results[].status`. File paths live on the enclosing `file`
158
+ * field of each suite node, so the tree is walked while tracking it.
159
+ */
160
+ const playwrightAdapter = {
161
+ scheme: "playwright-json-v1",
162
+ runner: "playwright",
163
+ matches: ({ runner }) => {
164
+ if (!invokesBinary(runner, PLAYWRIGHT_BINARIES))
165
+ return false;
166
+ // `playwright test` runs tests; `playwright install`/`codegen` do not.
167
+ return runner[1] === "test" || runner[1] === undefined;
168
+ },
169
+ reportSuffix: () => `--reporter=json`,
170
+ // Playwright ignores an output filename on the CLI (`--output` is its
171
+ // artifacts directory); its JSON reporter goes to stdout unless this
172
+ // variable names a file.
173
+ reportPathEnv: "PLAYWRIGHT_JSON_OUTPUT_NAME",
174
+ parse: (raw, cwd, workspaceRoot) => {
175
+ const suiteSchema = z.lazy(() => z.object({
176
+ file: z.string().optional(),
177
+ title: z.string().optional(),
178
+ specs: z
179
+ .array(z.object({
180
+ title: z.string(),
181
+ ok: z.boolean(),
182
+ tests: z.array(z.object({
183
+ status: z.string().optional(),
184
+ results: z
185
+ .array(z.object({
186
+ status: z.string(),
187
+ }))
188
+ .optional(),
189
+ })),
190
+ }))
191
+ .optional(),
192
+ suites: z.array(suiteSchema).optional(),
193
+ }));
194
+ const root = z
195
+ .object({
196
+ suites: z.array(suiteSchema).optional(),
197
+ // Playwright reports each suite's `file` relative to the resolved
198
+ // `testDir` (`config.rootDir`), not to the process cwd. A project
199
+ // with `testDir: "./tests"` reports `orders.spec.js` for
200
+ // `tests/orders.spec.js`, so resolving against cwd would produce a
201
+ // path that can never match a contract target.
202
+ config: z
203
+ .object({ rootDir: z.string().optional() })
204
+ .optional(),
205
+ })
206
+ .parse(raw);
207
+ // Prefer the runner's own reported root; fall back to cwd only when the
208
+ // report omitted it (older report shapes).
209
+ const reportedRoot = root.config?.rootDir;
210
+ const baseCwd = typeof reportedRoot === "string" && reportedRoot.length > 0
211
+ ? reportedRoot
212
+ : cwd;
213
+ const out = [];
214
+ const walk = (node, inheritedFile) => {
215
+ const record = asRecord(node);
216
+ if (!record)
217
+ return;
218
+ const file = typeof record.file === "string" ? record.file : inheritedFile;
219
+ for (const spec of record.specs ?? []) {
220
+ const specRecord = asRecord(spec);
221
+ if (!specRecord)
222
+ continue;
223
+ const title = String(specRecord.title ?? "");
224
+ const tests = specRecord.tests ?? [];
225
+ // A spec with several projects/retries yields several tests; the
226
+ // spec only passes when every attempt passed.
227
+ const statuses = tests.flatMap((test) => {
228
+ const testRecord = asRecord(test);
229
+ const results = testRecord?.results ?? [];
230
+ const attempts = results
231
+ .map((result) => asRecord(result)?.status)
232
+ .filter((value) => typeof value === "string");
233
+ return attempts.length > 0
234
+ ? attempts
235
+ : [String(testRecord?.status ?? "unknown")];
236
+ });
237
+ const status = statuses.every((value) => value === "passed")
238
+ ? "passed"
239
+ : statuses.includes("passed")
240
+ ? "failed"
241
+ : (statuses[0] ?? "unknown");
242
+ out.push({
243
+ file: toWorkspacePath(file ?? "", baseCwd, workspaceRoot),
244
+ title,
245
+ status,
246
+ });
247
+ }
248
+ for (const child of record.suites ?? []) {
249
+ walk(child, file);
250
+ }
251
+ };
252
+ for (const suite of root.suites ?? [])
253
+ walk(suite, undefined);
254
+ return out;
255
+ },
256
+ };
257
+ export const FRONTEND_TEST_ADAPTERS = [
258
+ vitestAdapter,
259
+ jestAdapter,
260
+ playwrightAdapter,
261
+ ];
262
+ /**
263
+ * Runners the container classifier recognizes as behavior verification but
264
+ * that have no adapter yet. Naming them lets the capability gap read as
265
+ * "framework not adapted yet" instead of a user configuration error.
266
+ */
267
+ const KNOWN_UNADAPTED_RUNNERS = new Set([
268
+ "mocha",
269
+ "ava",
270
+ "tap",
271
+ "cypress",
272
+ "testcafe",
273
+ "nightwatch",
274
+ "wdio",
275
+ "webdriverio",
276
+ ]);
277
+ /** Resolve the adapter for a command, or `undefined` when none matches. */
278
+ export function resolveFrontendTestAdapter(context) {
279
+ // A package script owns the detail: `npm test` must resolve through the
280
+ // script body, not the wrapper, or `npm test` and `vitest run` would
281
+ // disagree about capability.
282
+ if (context.scriptArgv) {
283
+ const scriptRunner = resolveRunnerArgv(context.scriptArgv);
284
+ const fromScript = FRONTEND_TEST_ADAPTERS.find((adapter) => adapter.matches({ ...context, runner: scriptRunner }));
285
+ if (fromScript)
286
+ return fromScript;
287
+ }
288
+ return FRONTEND_TEST_ADAPTERS.find((adapter) => adapter.matches(context));
289
+ }
290
+ /**
291
+ * Detect a recognized-but-unadapted runner so callers can report a capability
292
+ * gap. Returns the runner name when the command is a test runner we know of
293
+ * but cannot observe.
294
+ */
295
+ export function detectUnadaptedRunner(runner) {
296
+ const binary = basenameOf(runner[0] ?? "").replace(/\.(?:js|cjs|mjs|cmd|exe)$/u, "");
297
+ if (KNOWN_UNADAPTED_RUNNERS.has(binary))
298
+ return binary;
299
+ // `node --test`, `node --test-only`, plus `deno test`/`bun test`.
300
+ if (binary === "node" && runner.includes("--test"))
301
+ return "node --test";
302
+ if (binary === "deno" && runner[1] === "test")
303
+ return "deno test";
304
+ if (binary === "bun" && runner[1] === "test")
305
+ return "bun test";
306
+ return undefined;
307
+ }
308
+ /** Empty `file` means the runner reported a spec outside the workspace. */
309
+ export const isOutsideWorkspacePath = (file) => file === ".." || file.startsWith("../") || path.isAbsolute(file);
@@ -9,7 +9,7 @@
9
9
  ```text
10
10
  src/commands/dag-run-task.ts runDagRunTask
11
11
  → src/application/dag/generate-task-dag.ts generateTaskDagUseCase
12
- └─ src/workflows/dag/init-hybrid.ts initHybridDagFromTask (生成 DagSpec)
12
+ └─ src/workflows/dag/init-hybrid.ts initHybridDagFromTask (生成 DagSpec;模板族在 hybrid/templates/,输入准备在 hybrid/types.ts + hybrid/sources.ts)
13
13
  └─ src/application/dag/validate-dag.ts validateDagUseCase (候选 + 最终校验)
14
14
  └─ assertSafeForExecution (执行前安全检查)
15
15
  └─ src/application/dag/run-dag.ts runDagUseCase (执行)
@@ -43,7 +43,7 @@ src/commands/dag-validate.ts runDagValidate
43
43
 
44
44
  ### 专用 taskKind 与模板选择(架构摘要)
45
45
 
46
- - `initHybridDagFromTask` 按 `task.json.taskKind`(及前端需求自动分类)选择专用拓扑:`frontend-implementation`、`frontend-test`、`backend-test`、`knowledge-sync`、`knowledge-graph-bootstrap` 等;默认 `standard` 走 governance profile 通用实现链。
46
+ - `initHybridDagFromTask` 按 `task.json.taskKind`(及前端需求自动分类)选择专用拓扑:`frontend-implementation`、`frontend-test`、`backend-test`、`knowledge-sync`、`knowledge-graph-bootstrap` 等;默认 `standard` 走 governance profile 通用实现链。入口只做分发与治理装饰,六个模板 builder 在 `src/workflows/dag/hybrid/templates/`。
47
47
  - **backend-test(0.17.x)**:Markdown-first 固定短链(环境硬门 → Markdown 用例/Review → 单次 pytest + HTML → 报告与 L-5);不在本页展开节点清单,见 `docs/runtime/backend-test-workflow.md` 与 completed markdown-first plan。
48
48
  - 专用模板不是新的 governance profile:风险等级仍由 profile 规则判断;写边界仍受 `allowedPaths` / `forbiddenPaths` / `writeSet` 约束。
49
49
 
@@ -73,7 +73,7 @@ src/commands/dag-validate.ts runDagValidate
73
73
  - schema:`src/workflows/dag/types.ts`
74
74
  `dagNodeExecutorSchema = z.enum(["pi","shell","static"])`;DagNodeExecutor 默认 `"pi"`(`executor: dagNodeExecutorSchema.default("pi")`)。
75
75
  - `executor: "cursor"` 在 schema refine 阶段抛 `CURSOR_DAG_EXECUTOR_REMOVED_ERROR`(Pi-only = ADR 0001)。
76
- - Pi handler 先 `requireModel(input)` 校验已解析 model,再委托 `executeDagPiNode`(`src/executors/dag-pi-executor.ts`)。
76
+ - Pi handler 先 `requireModel(input)` 校验已解析 model,再委托 `executeDagPiNode`(`src/executors/dag-pi-executor.ts`;领域工具、plan ledger、分段会话在 `src/executors/dag-pi/`)。
77
77
  - shell handler = `executeDagShellNode`(`src/executors/shell-executor.ts`);static handler = `executeDagStaticNode`(`src/executors/dag-static-executor.ts`)。
78
78
 
79
79
  Executor 不得依赖 commands 或 CLI formatting;Cursor 不在受治理路径(`runtime-boundaries.md` §Executors)。
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@tea-agent/loop-agent",
3
- "version": "0.44.0-next.11",
3
+ "version": "0.44.0-next.12",
4
4
  "type": "module",
5
5
  "bin": {
6
6
  "loop-agent": "bin/loop-agent.js",