peaks-loop 4.0.10 → 4.0.11
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +27 -0
- package/config/eslint/.peaks-rules.cjs +123 -0
- package/dist/reporters/bdd-reporter.d.ts +36 -0
- package/dist/reporters/bdd-reporter.js +159 -0
- package/dist/services/audit/enforcers/active-skill-resolver.js +4 -1
- package/dist/services/qa/bdd-test-style-verifier.d.ts +88 -0
- package/dist/services/qa/bdd-test-style-verifier.js +268 -0
- package/dist/services/skills/presence-lease-service.js +1 -0
- package/dist/services/skills/skill-statusline-service.js +2 -0
- package/docs/test-style-contract.md +135 -0
- package/package.json +5 -3
- package/skills/bee/peaks-qa/references/qa-sub-agent-dispatch.md +17 -1
- package/skills/bee/peaks-rd/references/rd-sub-agent-dispatch.md +21 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,32 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 4.0.11 — 2026-08-05 (BDD test-style + statusline bugs)
|
|
4
|
+
|
|
5
|
+
**BDD given-when-then test style** (rid-2026-08-05-bdd-test-style, 5 slice / 19 commit):
|
|
6
|
+
- `scripts/migrate-to-bdd.mjs` — TS Compiler API-based AST migrator that rewrites every `it()` / `test()` / `describe()` to the given-when-then contract (it description with `when X` or `should Y` + 3-line `// given:` / `// when:` / `// then:` body comment). Idempotent.
|
|
7
|
+
- `src/services/qa/bdd-test-style-verifier.ts` — peaks-qa verification-time verifier that scans `git diff HEAD~1 -- '*.test.ts'` and rejects non-BDD slices (LLM-only enforcement, since callerId from `process.env.CLAUDE_CODE_SESSION_ID` cannot distinguish LLM vs human in Claude Code).
|
|
8
|
+
- `src/reporters/bdd-reporter.ts` — vitest custom reporter (flag-enabled via `--reporter ./src/reporters/bdd-reporter.ts`) emitting `Feature: <file>` / `Scenario: <describe>` / `Given|When|Then` document view.
|
|
9
|
+
- `skills/bee/peaks-rd/references/rd-sub-agent-dispatch.md` + `peaks-qa/references/qa-sub-agent-dispatch.md` — `## BDD Test Style Contract` / `## BDD Test Style Verification` soft-constraint sections added.
|
|
10
|
+
- `docs/test-style-contract.md` — LLM test-style guide included in npm `files` array (downstream opt-in).
|
|
11
|
+
- 26 test files migrated across 11 commits (one per top-level directory); 12 files already BDD-form (idempotent migrator skipped silently); 49 unit-test files total now in BDD shape.
|
|
12
|
+
- 49/49 unit tests behaviour-preserved (488 passed / 25 skipped / 0 introduced failures).
|
|
13
|
+
|
|
14
|
+
**Statusline bug fixes** (3 rid this release):
|
|
15
|
+
- `skill-statusline-service.ts` `readActiveLeaf`: stale `queued` dispatch entries no longer pollute statusline as in-flight leaves. `terminalStatuses` set extended.
|
|
16
|
+
- `presence-lease-service.ts` `setPresenceLease`: lease object now persists `mode` (was being silently dropped — regression from 4.0.8 Presence Lease Graph introduction). `[full-auto]` / `[assisted]` / `[swarm]` / `[strict]` tags now render.
|
|
17
|
+
- `audit/enforcers/active-skill-resolver.ts` legacy fall-back: per-caller `active-skill-*.json` legacy walk now reads and propagates `mode` (was hard-coded `mode: null`).
|
|
18
|
+
|
|
19
|
+
**Build-chain repair** (silences `npx tsc -p tsconfig.build.json` regression):
|
|
20
|
+
- `src/reporters/bdd-reporter.ts` no longer imports the un-exported `TestModule` / `TestCase` from `vitest/reporters`. Local minimal interfaces (`BddTestModuleLike` / `BddTestCaseLike`) match the runtime shape vitest passes to reporter hooks.
|
|
21
|
+
- Build now succeeds end-to-end; the `dist/` artifacts (which `bin/peaks.js` actually loads) reflect the source-level fixes.
|
|
22
|
+
|
|
23
|
+
**Cleanup tail** (5 rid carried over from b1 sweep):
|
|
24
|
+
- 63 request artifacts re-staged to terminal state (handed-off / verdict-issued / complete / sc-handoff → done).
|
|
25
|
+
- 8 OpenSpec proposal drift detected and routed through `peaks request transition`.
|
|
26
|
+
- One envelope-test-output log dropped from project root (was orphan inside `.gitignore:5 *.log` but never deleted).
|
|
27
|
+
|
|
28
|
+
**Lockstep bump.** peaks-loop-shared `0.0.40 → 0.0.41` (CLI_VERSION re-stamped to 4.0.11).
|
|
29
|
+
|
|
3
30
|
## 4.0.10 — 2026-08-04 (path-canonicalize + statusline-read-isolation)
|
|
4
31
|
|
|
5
32
|
**Windows statusline fixed.** `peaks-loop@4.0.9` always rendered `peaks empty` on Windows Git Bash. Root cause: the session-binding reader used strict `===` to compare `projectRoot`; the binding had been written with backslashes (`C:\Users\...`) but `peaks skill presence:set --project C:/Users/...` arrived with forward slashes, and Node treats them as distinct strings. The 4.0.8 fail-closed `PEAKS_SESSION_NOT_BOUND` gate then blocked the presence marker write, and the statusline never had a real skill to display.
|
|
@@ -0,0 +1,123 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* peaks-loop ESLint rules bundle (npm-package exports)
|
|
3
|
+
*
|
|
4
|
+
* rid-2026-08-05-jsts-lint-bundle — LLM auto-fix loop trigger.
|
|
5
|
+
*
|
|
6
|
+
* This file is a JSON-shape glue config. It does NOT define custom
|
|
7
|
+
* rules. It composes upstream packages only:
|
|
8
|
+
*
|
|
9
|
+
* - eslint:recommended
|
|
10
|
+
* - plugin:@typescript-eslint/recommended-type-checked
|
|
11
|
+
* - plugin:import/recommended
|
|
12
|
+
* - plugin:import/typescript
|
|
13
|
+
*
|
|
14
|
+
* Framework-specific rules (eslint-plugin-react, eslint-plugin-vue,
|
|
15
|
+
* eslint-plugin-svelte, eslint-plugin-nestjs, etc.) are LAYER 3 and
|
|
16
|
+
* loaded dynamically by `peaks code lint` via `npx --package <pkg>
|
|
17
|
+
* -- eslint`. They are NOT installed in this package's devDependencies
|
|
18
|
+
* (sediment §二 G-lint-1 turn-5 red line).
|
|
19
|
+
*
|
|
20
|
+
* --fix / --write / prettier are FORBIDDEN at the peaks code lint
|
|
21
|
+
* wrapper entry; the wrapper is a read-only verifier, not a formatter
|
|
22
|
+
* (sediment §二 G-lint-2). The thresholds below are intentionally
|
|
23
|
+
* permissive (warn, not error) so peaks-loop 4.0.10 baseline can adopt
|
|
24
|
+
* the bundle without auto-failing.
|
|
25
|
+
*/
|
|
26
|
+
'use strict';
|
|
27
|
+
|
|
28
|
+
/** @type {import('eslint').Linter.Config} */
|
|
29
|
+
module.exports = {
|
|
30
|
+
root: false,
|
|
31
|
+
parser: '@typescript-eslint/parser',
|
|
32
|
+
parserOptions: {
|
|
33
|
+
ecmaVersion: 2022,
|
|
34
|
+
sourceType: 'module',
|
|
35
|
+
project: ['./tsconfig.json', './tsconfig.build.json'],
|
|
36
|
+
tsconfigRootDir: __dirname + '/..'
|
|
37
|
+
},
|
|
38
|
+
env: {
|
|
39
|
+
node: true,
|
|
40
|
+
es2022: true
|
|
41
|
+
},
|
|
42
|
+
plugins: ['@typescript-eslint', 'import'],
|
|
43
|
+
extends: [
|
|
44
|
+
'eslint:recommended',
|
|
45
|
+
'plugin:@typescript-eslint/recommended-type-checked',
|
|
46
|
+
'plugin:import/recommended',
|
|
47
|
+
'plugin:import/typescript'
|
|
48
|
+
],
|
|
49
|
+
settings: {
|
|
50
|
+
'import/resolver': {
|
|
51
|
+
typescript: {
|
|
52
|
+
alwaysTryTypes: true,
|
|
53
|
+
project: ['./tsconfig.json', './tsconfig.build.json']
|
|
54
|
+
},
|
|
55
|
+
node: {
|
|
56
|
+
extensions: ['.js', '.ts', '.tsx', '.jsx']
|
|
57
|
+
}
|
|
58
|
+
}
|
|
59
|
+
},
|
|
60
|
+
ignorePatterns: [
|
|
61
|
+
'node_modules/',
|
|
62
|
+
'dist/',
|
|
63
|
+
'coverage/',
|
|
64
|
+
'output-styles/',
|
|
65
|
+
'skills/',
|
|
66
|
+
'agents/',
|
|
67
|
+
'bin/',
|
|
68
|
+
'scratch/',
|
|
69
|
+
'examples/'
|
|
70
|
+
],
|
|
71
|
+
rules: {
|
|
72
|
+
// L1 (eslint built-in) — always on, no plugin package required.
|
|
73
|
+
complexity: ['warn', { max: 10 }],
|
|
74
|
+
'max-lines-per-function': [
|
|
75
|
+
'warn',
|
|
76
|
+
{ max: 50, skipComments: true, skipBlankLines: true }
|
|
77
|
+
],
|
|
78
|
+
'max-params': ['warn', { max: 4 }],
|
|
79
|
+
'no-magic-numbers': [
|
|
80
|
+
'warn',
|
|
81
|
+
{ ignore: [0, 1, -1, 100, 1000] }
|
|
82
|
+
],
|
|
83
|
+
'no-explicit-any': 'warn',
|
|
84
|
+
'prefer-const': 'warn',
|
|
85
|
+
'no-var': 'error',
|
|
86
|
+
eqeqeq: ['warn', 'always', { null: 'ignore' }],
|
|
87
|
+
|
|
88
|
+
// L2 (@typescript-eslint) — type-aware; requires the
|
|
89
|
+
// recommended-type-checked base. configured via the extends above.
|
|
90
|
+
'@typescript-eslint/consistent-type-imports': [
|
|
91
|
+
'warn',
|
|
92
|
+
{ prefer: 'type-imports' }
|
|
93
|
+
],
|
|
94
|
+
'@typescript-eslint/no-non-null-assertion': 'warn',
|
|
95
|
+
'@typescript-eslint/no-implicit-any': 'warn',
|
|
96
|
+
// G-lint-1 §二 enum → as const: warn-only (escape hatch preserved).
|
|
97
|
+
'@typescript-eslint/no-restricted-syntax': [
|
|
98
|
+
'warn',
|
|
99
|
+
{
|
|
100
|
+
selector: 'TSEnumDeclaration',
|
|
101
|
+
message: 'Use "as const" union instead of TS enum.'
|
|
102
|
+
}
|
|
103
|
+
],
|
|
104
|
+
|
|
105
|
+
// L2 (eslint-plugin-import) — boundary hygiene.
|
|
106
|
+
'import/no-duplicates': 'warn',
|
|
107
|
+
'import/no-unresolved': 'off',
|
|
108
|
+
'import/named': 'off',
|
|
109
|
+
'import/default': 'off',
|
|
110
|
+
'import/namespace': 'off'
|
|
111
|
+
},
|
|
112
|
+
overrides: [
|
|
113
|
+
{
|
|
114
|
+
files: ['*.test.ts', '*.test.tsx', 'tests/**/*.ts', 'tests/**/*.tsx'],
|
|
115
|
+
rules: {
|
|
116
|
+
'no-magic-numbers': 'off',
|
|
117
|
+
complexity: 'off',
|
|
118
|
+
'max-lines-per-function': 'off',
|
|
119
|
+
'@typescript-eslint/no-explicit-any': 'off'
|
|
120
|
+
}
|
|
121
|
+
}
|
|
122
|
+
]
|
|
123
|
+
};
|
|
@@ -0,0 +1,36 @@
|
|
|
1
|
+
import type { Reporter } from 'vitest/reporters';
|
|
2
|
+
/** Minimal shape of vitest's TestModule tree node. Vitest 4.1.10 does
|
|
3
|
+
* not export these types publicly; this mirrors the runtime shape
|
|
4
|
+
* the reporter hooks actually receive. */
|
|
5
|
+
interface BddTestModuleLike {
|
|
6
|
+
moduleId?: string;
|
|
7
|
+
relativeModuleId?: string;
|
|
8
|
+
children: {
|
|
9
|
+
tests(): Iterable<BddTestCaseLike>;
|
|
10
|
+
suites(): Iterable<BddTestModuleLike>;
|
|
11
|
+
};
|
|
12
|
+
}
|
|
13
|
+
interface BddTestCaseLike {
|
|
14
|
+
name: string;
|
|
15
|
+
fullName?: string;
|
|
16
|
+
state?: 'passed' | 'failed' | 'skipped';
|
|
17
|
+
result?: () => unknown;
|
|
18
|
+
}
|
|
19
|
+
declare class BddReporter implements Reporter {
|
|
20
|
+
/** Key: relative module id; Value: per-feature rendered scenarios. */
|
|
21
|
+
private readonly features;
|
|
22
|
+
/**
|
|
23
|
+
* Vitest calls `onTestModuleEnd` after a module finishes. We use it
|
|
24
|
+
* to drain the per-module scenarios into the document map and
|
|
25
|
+
* mark the file's pass/fail status.
|
|
26
|
+
*/
|
|
27
|
+
onTestModuleEnd(testModule: BddTestModuleLike): void;
|
|
28
|
+
/**
|
|
29
|
+
* Final emit. We deliberately print to stdout with `console.log`
|
|
30
|
+
* (vitest captures stdout when needed) and never call `process.exit`
|
|
31
|
+
* — that is the orchestrator's job. Failure reasons surface as plain
|
|
32
|
+
* text so a downstream LLM prompt can grep for `FAILED:`.
|
|
33
|
+
*/
|
|
34
|
+
onTestRunEnd(): void;
|
|
35
|
+
}
|
|
36
|
+
export default BddReporter;
|
|
@@ -0,0 +1,159 @@
|
|
|
1
|
+
// src/reporters/bdd-reporter.ts
|
|
2
|
+
//
|
|
3
|
+
// rid-2026-08-05-bdd-test-style Slice C — vitest custom reporter that
|
|
4
|
+
// emits a pure BDD document view of the run. Designed for business
|
|
5
|
+
// reviewers and downstream LLM prompts; it is NOT a replacement for
|
|
6
|
+
// the default reporter.
|
|
7
|
+
//
|
|
8
|
+
// Why a custom reporter:
|
|
9
|
+
// The default reporter focuses on pass/fail and timing. The BDD
|
|
10
|
+
// reporter transcribes `Feature: <file>` / `Scenario: <describe> ->
|
|
11
|
+
// it` into a single human-readable document so a non-engineer can
|
|
12
|
+
// scan what the suite actually exercises.
|
|
13
|
+
//
|
|
14
|
+
// Why no new dep:
|
|
15
|
+
// vitest 4.1.10 (frozen 2026-07-25) ships the `Reporter` interface
|
|
16
|
+
// in `vitest/reporters`. The custom reporter must have a `default`
|
|
17
|
+
// export — the CLI loads it via `runner.import(path)` and validates
|
|
18
|
+
// `customReporterModule.default` is defined (see vitest cli-api chunks
|
|
19
|
+
// line 11371). Importing the `Reporter` type from vitest does not add
|
|
20
|
+
// a runtime dep; tsc resolves it through vitest's dts shim.
|
|
21
|
+
//
|
|
22
|
+
// Why a flag-only reporter:
|
|
23
|
+
// Per rid design section 4 Slice C, the default vitest run is
|
|
24
|
+
// unchanged. This file is opt-in via:
|
|
25
|
+
//
|
|
26
|
+
// pnpm vitest run --reporter ./src/reporters/bdd-reporter.ts <file>
|
|
27
|
+
//
|
|
28
|
+
// Anti-fake-green rule (CLI silent-catch):
|
|
29
|
+
// The reporter does not swallow vitest result shapes. Every state
|
|
30
|
+
// branch (`passed` / `failed` / `skipped`) is rendered explicitly so
|
|
31
|
+
// downstream reviewers cannot misread a hidden failure.
|
|
32
|
+
//
|
|
33
|
+
// Karpathy note:
|
|
34
|
+
// The reporter deliberately emits ONE document per file with the
|
|
35
|
+
// 4-line Feature/Scenario/Given/When/Then shape — no extra layout
|
|
36
|
+
// metadata, no JSON sidecar. Anything beyond what the spec asked
|
|
37
|
+
// for is excluded by Simplicity First.
|
|
38
|
+
class BddReporter {
|
|
39
|
+
/** Key: relative module id; Value: per-feature rendered scenarios. */
|
|
40
|
+
features = new Map();
|
|
41
|
+
/**
|
|
42
|
+
* Vitest calls `onTestModuleEnd` after a module finishes. We use it
|
|
43
|
+
* to drain the per-module scenarios into the document map and
|
|
44
|
+
* mark the file's pass/fail status.
|
|
45
|
+
*/
|
|
46
|
+
onTestModuleEnd(testModule) {
|
|
47
|
+
const moduleId = testModule.relativeModuleId ?? testModule.moduleId ?? '';
|
|
48
|
+
const file = basename(moduleId);
|
|
49
|
+
const scenarios = [];
|
|
50
|
+
collectScenarios(testModule, file, scenarios);
|
|
51
|
+
this.features.set(file, scenarios);
|
|
52
|
+
}
|
|
53
|
+
/**
|
|
54
|
+
* Final emit. We deliberately print to stdout with `console.log`
|
|
55
|
+
* (vitest captures stdout when needed) and never call `process.exit`
|
|
56
|
+
* — that is the orchestrator's job. Failure reasons surface as plain
|
|
57
|
+
* text so a downstream LLM prompt can grep for `FAILED:`.
|
|
58
|
+
*/
|
|
59
|
+
onTestRunEnd() {
|
|
60
|
+
const lines = [];
|
|
61
|
+
const features = [];
|
|
62
|
+
for (const [feature, scenarios] of this.features) {
|
|
63
|
+
const ok = scenarios.every((s) => s.state === 'passed' || s.state === 'skipped');
|
|
64
|
+
features.push({ feature, scenarios, ok });
|
|
65
|
+
}
|
|
66
|
+
// Deterministic order: alphabetical by file basename so two runs on
|
|
67
|
+
// the same diff produce byte-identical docs (avoids noisy diffs).
|
|
68
|
+
features.sort((a, b) => a.feature.localeCompare(b.feature));
|
|
69
|
+
for (const f of features) {
|
|
70
|
+
lines.push(`Feature: ${f.feature}`);
|
|
71
|
+
if (f.scenarios.length === 0) {
|
|
72
|
+
// Empty file still surfaces the Feature line so the document
|
|
73
|
+
// is a faithful list of files the runner touched.
|
|
74
|
+
lines.push('');
|
|
75
|
+
continue;
|
|
76
|
+
}
|
|
77
|
+
for (const s of f.scenarios) {
|
|
78
|
+
lines.push(` Scenario: ${s.scenario || '<root>'}`);
|
|
79
|
+
lines.push(` Given ${s.title}`);
|
|
80
|
+
lines.push(` When vitest runs this test`);
|
|
81
|
+
if (s.state === 'passed') {
|
|
82
|
+
lines.push(` Then should pass`);
|
|
83
|
+
}
|
|
84
|
+
else if (s.state === 'skipped') {
|
|
85
|
+
lines.push(` Then should skip`);
|
|
86
|
+
}
|
|
87
|
+
else {
|
|
88
|
+
const reason = s.error ? ` (${truncate(s.error, 200)})` : '';
|
|
89
|
+
lines.push(` Then FAILED: ${s.title}${reason}`);
|
|
90
|
+
}
|
|
91
|
+
}
|
|
92
|
+
lines.push('');
|
|
93
|
+
}
|
|
94
|
+
console.log(lines.join('\n'));
|
|
95
|
+
}
|
|
96
|
+
}
|
|
97
|
+
function basename(path) {
|
|
98
|
+
// vitest module ids are POSIX-style even on Windows; split on '/'
|
|
99
|
+
// then on '\\' as a defensive fallback.
|
|
100
|
+
const idx = Math.max(path.lastIndexOf('/'), path.lastIndexOf('\\'));
|
|
101
|
+
return idx === -1 ? path : path.slice(idx + 1);
|
|
102
|
+
}
|
|
103
|
+
/**
|
|
104
|
+
* Walk a `TestModule` recursively and collect rendered scenarios.
|
|
105
|
+
* `describe` blocks contribute their name to the Scenario label;
|
|
106
|
+
* tests declared at module root produce a `<root>` Scenario so the
|
|
107
|
+
* structure is uniform.
|
|
108
|
+
*/
|
|
109
|
+
function collectScenarios(entity, file, out) {
|
|
110
|
+
const visited = new WeakSet();
|
|
111
|
+
const walk = (node, scenarioLabel) => {
|
|
112
|
+
if (node === null || typeof node !== 'object')
|
|
113
|
+
return;
|
|
114
|
+
if (visited.has(node))
|
|
115
|
+
return;
|
|
116
|
+
visited.add(node);
|
|
117
|
+
const obj = node;
|
|
118
|
+
if (obj.type === 'test') {
|
|
119
|
+
const tc = node;
|
|
120
|
+
const result = tc.result ? tc.result() : undefined;
|
|
121
|
+
const resultObj = (result ?? {});
|
|
122
|
+
const state = (resultObj.state === 'passed' || resultObj.state === 'failed' || resultObj.state === 'skipped')
|
|
123
|
+
? resultObj.state
|
|
124
|
+
: 'skipped';
|
|
125
|
+
const err = resultObj.state === 'failed' && resultObj.errors && resultObj.errors[0]
|
|
126
|
+
? (resultObj.errors[0].message ?? 'unknown failure')
|
|
127
|
+
: undefined;
|
|
128
|
+
out.push({
|
|
129
|
+
scenario: scenarioLabel,
|
|
130
|
+
title: tc.name,
|
|
131
|
+
state,
|
|
132
|
+
error: err,
|
|
133
|
+
});
|
|
134
|
+
return;
|
|
135
|
+
}
|
|
136
|
+
// For a suite/module, descend with the suite's name pushed.
|
|
137
|
+
const suiteName = obj.name ?? '';
|
|
138
|
+
const childSuiteLabel = suiteName || scenarioLabel;
|
|
139
|
+
if (obj.children) {
|
|
140
|
+
try {
|
|
141
|
+
for (const t of obj.children.tests()) {
|
|
142
|
+
walk(t, childSuiteLabel);
|
|
143
|
+
}
|
|
144
|
+
for (const s of obj.children.suites()) {
|
|
145
|
+
walk(s, childSuiteLabel);
|
|
146
|
+
}
|
|
147
|
+
}
|
|
148
|
+
catch {
|
|
149
|
+
// Defensive: vitest internals may throw on teardown. We do not
|
|
150
|
+
// mask the document — we just stop collecting from this node.
|
|
151
|
+
}
|
|
152
|
+
}
|
|
153
|
+
};
|
|
154
|
+
walk(entity, '');
|
|
155
|
+
}
|
|
156
|
+
function truncate(s, n) {
|
|
157
|
+
return s.length <= n ? s : `${s.slice(0, n - 3)}...`;
|
|
158
|
+
}
|
|
159
|
+
export default BddReporter;
|
|
@@ -127,7 +127,10 @@ export function resolveActiveSkillForCaller(projectRoot, opts) {
|
|
|
127
127
|
const raw = readFileSync(filePath, 'utf8');
|
|
128
128
|
const parsed = JSON.parse(raw);
|
|
129
129
|
if (typeof parsed.skill === 'string' && parsed.skill.length > 0) {
|
|
130
|
-
|
|
130
|
+
const legacyMode = typeof parsed.mode === 'string' && parsed.mode.length > 0
|
|
131
|
+
? parsed.mode
|
|
132
|
+
: null;
|
|
133
|
+
return { skill: parsed.skill, callerId, sessionId, mode: legacyMode, source: 'file' };
|
|
131
134
|
}
|
|
132
135
|
}
|
|
133
136
|
catch { // TODO(g2): legacy silent catch — grace: 1 minor release (v2.14.0)
|
|
@@ -0,0 +1,88 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* src/services/qa/bdd-test-style-verifier.ts
|
|
3
|
+
*
|
|
4
|
+
* rid-2026-08-05-bdd-test-style Slice B — peaks-qa verification-time
|
|
5
|
+
* BDD test-style verifier. This is the read-only, post-edit companion
|
|
6
|
+
* to the `scripts/migrate-to-bdd.mjs` AST migrator shipped in Slice A.
|
|
7
|
+
*
|
|
8
|
+
* Purpose:
|
|
9
|
+
* When peaks-qa runs its verification gate, it picks up the git diff
|
|
10
|
+
* for the slice and asks this module whether the new / modified test
|
|
11
|
+
* files comply with the BDD given-when-then style. The verdict is
|
|
12
|
+
* surfaced as either `ok` (and the slice can advance) or one of two
|
|
13
|
+
* structured failure reasons (`missing-given-when-then` or
|
|
14
|
+
* `description-no-should-when`) that the caller turns into a
|
|
15
|
+
* `qa-handoff` rejection back to peaks-rd.
|
|
16
|
+
*
|
|
17
|
+
* Why a real AST and not a regex:
|
|
18
|
+
* The Slice A migrator established the convention: test files have
|
|
19
|
+
* multi-line `it(...)` calls, nested arrow bodies, and string
|
|
20
|
+
* literals that often contain words like "when" inside the assertion
|
|
21
|
+
* message (not in the description). A regex pass on the raw source
|
|
22
|
+
* would false-positive on string internals. The TypeScript Compiler
|
|
23
|
+
* API (already a dev dep via vitest) lets us:
|
|
24
|
+
* 1. Inspect the first `StringLiteral` argument of an `it` /
|
|
25
|
+
* `test` / `describe` call without scanning comments or
|
|
26
|
+
* string content inside the body.
|
|
27
|
+
* 2. Walk only the leading-comment ranges that sit before the
|
|
28
|
+
* first statement of the callback block, so a `// when:`
|
|
29
|
+
* inside an `expect(actual).toEqual('when X happens')` is
|
|
30
|
+
* correctly ignored.
|
|
31
|
+
*
|
|
32
|
+
* No new dependencies. The verifier is intentionally synchronous and
|
|
33
|
+
* pure (input source + path list -> verdict) so peaks-qa can call it
|
|
34
|
+
* from a deterministic verification step without subprocess overhead.
|
|
35
|
+
*
|
|
36
|
+
* Anti-fake-green (CLI silent-catch rule):
|
|
37
|
+
* This module throws on parse failure. It does NOT swallow parse
|
|
38
|
+
* errors and return `{ ok: true }` — that would silently green-light
|
|
39
|
+
* malformed test files. A parse error is a structural problem; the
|
|
40
|
+
* caller must surface it.
|
|
41
|
+
*/
|
|
42
|
+
/** Structured failure reasons the verifier can return. */
|
|
43
|
+
export type BddStyleFailureReason = 'missing-given-when-then' | 'description-no-should-when';
|
|
44
|
+
/** Successful verdict — includes the count of inspected `it`/`test` calls. */
|
|
45
|
+
export interface BddStyleOk {
|
|
46
|
+
readonly ok: true;
|
|
47
|
+
readonly scanned: number;
|
|
48
|
+
}
|
|
49
|
+
/** Structured failure verdict — the file/line makes the rejection actionable. */
|
|
50
|
+
export interface BddStyleFail {
|
|
51
|
+
readonly ok: false;
|
|
52
|
+
readonly reason: BddStyleFailureReason;
|
|
53
|
+
readonly file: string;
|
|
54
|
+
readonly line: number;
|
|
55
|
+
/** For `description-no-should-when`: the original description. */
|
|
56
|
+
readonly description?: string;
|
|
57
|
+
/** For `missing-given-when-then`: a stable string the caller can compare. */
|
|
58
|
+
readonly expected?: string;
|
|
59
|
+
}
|
|
60
|
+
export type BddStyleVerdict = BddStyleOk | BddStyleFail;
|
|
61
|
+
/** Public input surface — keep small so the contract is hard to misuse. */
|
|
62
|
+
export interface VerifyBddStyleInput {
|
|
63
|
+
readonly projectRoot: string;
|
|
64
|
+
readonly testFiles: readonly string[];
|
|
65
|
+
}
|
|
66
|
+
/**
|
|
67
|
+
* Verify that every `it(...)` / `test(...)` call in the given test
|
|
68
|
+
* files follows the BDD given-when-then contract.
|
|
69
|
+
*
|
|
70
|
+
* Contract:
|
|
71
|
+
* 1. The first `StringLiteral` argument of every `it` / `test` call
|
|
72
|
+
* MUST match `/(\bwhen\b|\bshould\b)/` (word-boundary anchored,
|
|
73
|
+
* case-insensitive). A regex on the raw description is correct
|
|
74
|
+
* here because the description itself is a literal — there is
|
|
75
|
+
* no nested template literal to misread.
|
|
76
|
+
* 2. The callback body (the second argument when it is an arrow /
|
|
77
|
+
* function expression with a block) MUST have a `// given:`,
|
|
78
|
+
* `// when:`, `// then:` triple at the top, in that order,
|
|
79
|
+
* within the first 3 leading-comment ranges before the first
|
|
80
|
+
* statement. The `// arrange:` / `// act:` / `// assert:` AAA
|
|
81
|
+
* legacy is NOT accepted — the contract is given-when-then
|
|
82
|
+
* only.
|
|
83
|
+
*
|
|
84
|
+
* Returns the FIRST failure encountered (file order, then
|
|
85
|
+
* top-to-bottom line order). A structured `BddStyleFail` is what the
|
|
86
|
+
* caller maps to `qa-handoff` rejection.
|
|
87
|
+
*/
|
|
88
|
+
export declare function verifyBddStyle(input: VerifyBddStyleInput): BddStyleVerdict;
|
|
@@ -0,0 +1,268 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* src/services/qa/bdd-test-style-verifier.ts
|
|
3
|
+
*
|
|
4
|
+
* rid-2026-08-05-bdd-test-style Slice B — peaks-qa verification-time
|
|
5
|
+
* BDD test-style verifier. This is the read-only, post-edit companion
|
|
6
|
+
* to the `scripts/migrate-to-bdd.mjs` AST migrator shipped in Slice A.
|
|
7
|
+
*
|
|
8
|
+
* Purpose:
|
|
9
|
+
* When peaks-qa runs its verification gate, it picks up the git diff
|
|
10
|
+
* for the slice and asks this module whether the new / modified test
|
|
11
|
+
* files comply with the BDD given-when-then style. The verdict is
|
|
12
|
+
* surfaced as either `ok` (and the slice can advance) or one of two
|
|
13
|
+
* structured failure reasons (`missing-given-when-then` or
|
|
14
|
+
* `description-no-should-when`) that the caller turns into a
|
|
15
|
+
* `qa-handoff` rejection back to peaks-rd.
|
|
16
|
+
*
|
|
17
|
+
* Why a real AST and not a regex:
|
|
18
|
+
* The Slice A migrator established the convention: test files have
|
|
19
|
+
* multi-line `it(...)` calls, nested arrow bodies, and string
|
|
20
|
+
* literals that often contain words like "when" inside the assertion
|
|
21
|
+
* message (not in the description). A regex pass on the raw source
|
|
22
|
+
* would false-positive on string internals. The TypeScript Compiler
|
|
23
|
+
* API (already a dev dep via vitest) lets us:
|
|
24
|
+
* 1. Inspect the first `StringLiteral` argument of an `it` /
|
|
25
|
+
* `test` / `describe` call without scanning comments or
|
|
26
|
+
* string content inside the body.
|
|
27
|
+
* 2. Walk only the leading-comment ranges that sit before the
|
|
28
|
+
* first statement of the callback block, so a `// when:`
|
|
29
|
+
* inside an `expect(actual).toEqual('when X happens')` is
|
|
30
|
+
* correctly ignored.
|
|
31
|
+
*
|
|
32
|
+
* No new dependencies. The verifier is intentionally synchronous and
|
|
33
|
+
* pure (input source + path list -> verdict) so peaks-qa can call it
|
|
34
|
+
* from a deterministic verification step without subprocess overhead.
|
|
35
|
+
*
|
|
36
|
+
* Anti-fake-green (CLI silent-catch rule):
|
|
37
|
+
* This module throws on parse failure. It does NOT swallow parse
|
|
38
|
+
* errors and return `{ ok: true }` — that would silently green-light
|
|
39
|
+
* malformed test files. A parse error is a structural problem; the
|
|
40
|
+
* caller must surface it.
|
|
41
|
+
*/
|
|
42
|
+
import { readFileSync } from 'node:fs';
|
|
43
|
+
import { resolve } from 'node:path';
|
|
44
|
+
import ts from 'typescript';
|
|
45
|
+
/** Test runners whose first string-arg is the test description. */
|
|
46
|
+
const TEST_NAMES = new Set(['it', 'test']);
|
|
47
|
+
/**
|
|
48
|
+
* Verify that every `it(...)` / `test(...)` call in the given test
|
|
49
|
+
* files follows the BDD given-when-then contract.
|
|
50
|
+
*
|
|
51
|
+
* Contract:
|
|
52
|
+
* 1. The first `StringLiteral` argument of every `it` / `test` call
|
|
53
|
+
* MUST match `/(\bwhen\b|\bshould\b)/` (word-boundary anchored,
|
|
54
|
+
* case-insensitive). A regex on the raw description is correct
|
|
55
|
+
* here because the description itself is a literal — there is
|
|
56
|
+
* no nested template literal to misread.
|
|
57
|
+
* 2. The callback body (the second argument when it is an arrow /
|
|
58
|
+
* function expression with a block) MUST have a `// given:`,
|
|
59
|
+
* `// when:`, `// then:` triple at the top, in that order,
|
|
60
|
+
* within the first 3 leading-comment ranges before the first
|
|
61
|
+
* statement. The `// arrange:` / `// act:` / `// assert:` AAA
|
|
62
|
+
* legacy is NOT accepted — the contract is given-when-then
|
|
63
|
+
* only.
|
|
64
|
+
*
|
|
65
|
+
* Returns the FIRST failure encountered (file order, then
|
|
66
|
+
* top-to-bottom line order). A structured `BddStyleFail` is what the
|
|
67
|
+
* caller maps to `qa-handoff` rejection.
|
|
68
|
+
*/
|
|
69
|
+
export function verifyBddStyle(input) {
|
|
70
|
+
let scanned = 0;
|
|
71
|
+
for (const rel of input.testFiles) {
|
|
72
|
+
const absPath = resolve(input.projectRoot, rel);
|
|
73
|
+
const source = readFileSync(absPath, 'utf8');
|
|
74
|
+
const sourceFile = ts.createSourceFile(rel, source, ts.ScriptTarget.ESNext,
|
|
75
|
+
/* setParentNodes */ true, ts.ScriptKind.TS);
|
|
76
|
+
let earliestFail = null;
|
|
77
|
+
const recordFail = (fail) => {
|
|
78
|
+
if (earliestFail === null) {
|
|
79
|
+
earliestFail = fail;
|
|
80
|
+
return;
|
|
81
|
+
}
|
|
82
|
+
const { line: existingLine } = earliestFail;
|
|
83
|
+
if (fail.line < existingLine)
|
|
84
|
+
earliestFail = fail;
|
|
85
|
+
};
|
|
86
|
+
const visit = (node) => {
|
|
87
|
+
if (earliestFail !== null)
|
|
88
|
+
return;
|
|
89
|
+
if (ts.isCallExpression(node)) {
|
|
90
|
+
const callee = node.expression;
|
|
91
|
+
if (ts.isIdentifier(callee) && TEST_NAMES.has(callee.text)) {
|
|
92
|
+
scanned += 1;
|
|
93
|
+
const descCheck = checkDescription(node, sourceFile, rel);
|
|
94
|
+
if (descCheck !== null) {
|
|
95
|
+
recordFail(descCheck);
|
|
96
|
+
return;
|
|
97
|
+
}
|
|
98
|
+
const bodyCheck = checkBody(node, sourceFile, rel);
|
|
99
|
+
if (bodyCheck !== null) {
|
|
100
|
+
recordFail(bodyCheck);
|
|
101
|
+
return;
|
|
102
|
+
}
|
|
103
|
+
}
|
|
104
|
+
}
|
|
105
|
+
ts.forEachChild(node, visit);
|
|
106
|
+
};
|
|
107
|
+
visit(sourceFile);
|
|
108
|
+
if (earliestFail !== null)
|
|
109
|
+
return earliestFail;
|
|
110
|
+
}
|
|
111
|
+
return { ok: true, scanned };
|
|
112
|
+
}
|
|
113
|
+
/**
|
|
114
|
+
* Inspect the first string-literal argument of an `it` / `test` call.
|
|
115
|
+
*
|
|
116
|
+
* - If the first argument is not a string literal, treat it as a
|
|
117
|
+
* failure (the BDD contract requires a literal description).
|
|
118
|
+
* - If the literal text does not contain "when" or "should" as a
|
|
119
|
+
* whole word, return a `description-no-should-when` failure.
|
|
120
|
+
*/
|
|
121
|
+
function checkDescription(call, sourceFile, relPath) {
|
|
122
|
+
const firstArg = call.arguments[0];
|
|
123
|
+
if (firstArg === undefined || !ts.isStringLiteralLike(firstArg)) {
|
|
124
|
+
const pos = call.getStart(sourceFile);
|
|
125
|
+
const { line } = sourceFile.getLineAndCharacterOfPosition(pos);
|
|
126
|
+
return {
|
|
127
|
+
ok: false,
|
|
128
|
+
reason: 'description-no-should-when',
|
|
129
|
+
file: relPath,
|
|
130
|
+
line: line + 1,
|
|
131
|
+
description: '<non-literal first argument>',
|
|
132
|
+
expected: 'first argument must be a string literal containing "when" or "should"',
|
|
133
|
+
};
|
|
134
|
+
}
|
|
135
|
+
const description = firstArg.text;
|
|
136
|
+
if (!hasWhenOrShould(description)) {
|
|
137
|
+
const pos = firstArg.getStart(sourceFile);
|
|
138
|
+
const { line } = sourceFile.getLineAndCharacterOfPosition(pos);
|
|
139
|
+
return {
|
|
140
|
+
ok: false,
|
|
141
|
+
reason: 'description-no-should-when',
|
|
142
|
+
file: relPath,
|
|
143
|
+
line: line + 1,
|
|
144
|
+
description,
|
|
145
|
+
expected: 'description must contain the word "when" or "should" (BDD style)',
|
|
146
|
+
};
|
|
147
|
+
}
|
|
148
|
+
return null;
|
|
149
|
+
}
|
|
150
|
+
/**
|
|
151
|
+
* Inspect the callback body of an `it` / `test` call for the
|
|
152
|
+
* `// given:` / `// when:` / `// then:` triple.
|
|
153
|
+
*
|
|
154
|
+
* Rules (Slice A migrator + design §4.B):
|
|
155
|
+
* - The second argument must be an arrow / function expression
|
|
156
|
+
* with a block body. If it is missing or not a block (e.g. an
|
|
157
|
+
* expression-body arrow `it('x', () => expect(y).toBe(z))`),
|
|
158
|
+
* we still need the comments — but expression-body arrows
|
|
159
|
+
* cannot host them. In that case we fall back to inspecting
|
|
160
|
+
* the leading comments before the entire call expression,
|
|
161
|
+
* which matches the Slice A migrator's `isAlreadyMigrated`
|
|
162
|
+
* check shape.
|
|
163
|
+
* - The three comments must be the FIRST THREE leading-comment
|
|
164
|
+
* ranges before the relevant body / first-statement anchor.
|
|
165
|
+
* - The order must be `given` → `when` → `then`. A re-ordered
|
|
166
|
+
* triple is rejected.
|
|
167
|
+
*/
|
|
168
|
+
function checkBody(call, sourceFile, relPath) {
|
|
169
|
+
const body = getCallbackBlock(call);
|
|
170
|
+
if (body !== null) {
|
|
171
|
+
return checkBlockLeadingComments(body, sourceFile, relPath);
|
|
172
|
+
}
|
|
173
|
+
// Expression-body arrow or non-block callback: comments cannot
|
|
174
|
+
// live inside the body. The Slice A migrator only inserts the
|
|
175
|
+
// triple on block bodies, so an expression-body form is by
|
|
176
|
+
// definition non-BDD and must fail. This keeps the contract
|
|
177
|
+
// symmetric with the migrator.
|
|
178
|
+
const pos = call.getStart(sourceFile);
|
|
179
|
+
const { line } = sourceFile.getLineAndCharacterOfPosition(pos);
|
|
180
|
+
return {
|
|
181
|
+
ok: false,
|
|
182
|
+
reason: 'missing-given-when-then',
|
|
183
|
+
file: relPath,
|
|
184
|
+
line: line + 1,
|
|
185
|
+
expected: 'block-body callback with // given: / // when: / // then: comments at the top',
|
|
186
|
+
};
|
|
187
|
+
}
|
|
188
|
+
function getCallbackBlock(call) {
|
|
189
|
+
const callback = call.arguments[1];
|
|
190
|
+
if (callback === undefined)
|
|
191
|
+
return null;
|
|
192
|
+
if (!ts.isArrowFunction(callback) && !ts.isFunctionExpression(callback))
|
|
193
|
+
return null;
|
|
194
|
+
if (!callback.body || !ts.isBlock(callback.body))
|
|
195
|
+
return null;
|
|
196
|
+
return callback.body;
|
|
197
|
+
}
|
|
198
|
+
function checkBlockLeadingComments(block, sourceFile, relPath) {
|
|
199
|
+
// TypeScript's `getLeadingCommentRanges` API is unreliable for
|
|
200
|
+
// comment-only blocks: with `setParentNodes: true`, an empty
|
|
201
|
+
// block (no statements, only comments) has no anchor to attach
|
|
202
|
+
// the comments to, so the API returns zero ranges. To get a
|
|
203
|
+
// deterministic answer, we scan the block's text directly and
|
|
204
|
+
// pick the first three non-empty lines.
|
|
205
|
+
//
|
|
206
|
+
// The block's text spans `{` ... `}`. We extract the body,
|
|
207
|
+
// split on lines, and check the first three non-empty lines for
|
|
208
|
+
// the BDD triple. This is AST-driven (we use the block's source
|
|
209
|
+
// range from the SourceFile, not a global regex) and survives
|
|
210
|
+
// both empty-body and populated-body cases.
|
|
211
|
+
const blockStart = block.getStart(sourceFile) + 1; // skip `{`
|
|
212
|
+
const blockEnd = block.end - 1; // skip `}`
|
|
213
|
+
const body = sourceFile.text.slice(blockStart, blockEnd);
|
|
214
|
+
const lines = body.split(/\r?\n/);
|
|
215
|
+
const nonEmpty = [];
|
|
216
|
+
for (const line of lines) {
|
|
217
|
+
if (line.trim().length === 0)
|
|
218
|
+
continue;
|
|
219
|
+
nonEmpty.push(line);
|
|
220
|
+
if (nonEmpty.length === 3)
|
|
221
|
+
break;
|
|
222
|
+
}
|
|
223
|
+
if (nonEmpty.length < 3 || !matchesBddTriple(nonEmpty)) {
|
|
224
|
+
return makeMissingCommentFailure(block, sourceFile, relPath);
|
|
225
|
+
}
|
|
226
|
+
return null;
|
|
227
|
+
}
|
|
228
|
+
function makeMissingCommentFailure(block, sourceFile, relPath) {
|
|
229
|
+
// Report the line of the opening `{` + 1 — the line that should
|
|
230
|
+
// contain the first comment of the BDD triple. This gives the
|
|
231
|
+
// caller a stable pointer even when the block is empty.
|
|
232
|
+
const pos = block.getStart(sourceFile) + 1;
|
|
233
|
+
const { line } = sourceFile.getLineAndCharacterOfPosition(pos);
|
|
234
|
+
return {
|
|
235
|
+
ok: false,
|
|
236
|
+
reason: 'missing-given-when-then',
|
|
237
|
+
file: relPath,
|
|
238
|
+
line: line + 1,
|
|
239
|
+
expected: '// given: / // when: / // then: triple at the top of the block body',
|
|
240
|
+
};
|
|
241
|
+
}
|
|
242
|
+
/**
|
|
243
|
+
* Match the three leading comments against the BDD triple. Each
|
|
244
|
+
* entry must be a `// <keyword>:` line (with optional trailing
|
|
245
|
+
* whitespace); the keywords must appear in `given`, `when`, `then`
|
|
246
|
+
* order, case-insensitive.
|
|
247
|
+
*/
|
|
248
|
+
function matchesBddTriple(triple) {
|
|
249
|
+
if (triple.length !== 3)
|
|
250
|
+
return false;
|
|
251
|
+
// Each entry must be a `// <keyword>:` line, optionally followed
|
|
252
|
+
// by descriptive text. The Slice A migrator's `buildCommentBlock`
|
|
253
|
+
// produces `// given: the test setup` / `// when: the function
|
|
254
|
+
// under test is invoked` / `// then: the result matches the
|
|
255
|
+
// expectation` — the `when` line uses two spaces after the colon
|
|
256
|
+
// for visual alignment with `given:` and `then:`, so the regex
|
|
257
|
+
// is intentionally permissive about trailing text.
|
|
258
|
+
const patterns = [
|
|
259
|
+
/^\s*\/\/\s*given\s*:/i,
|
|
260
|
+
/^\s*\/\/\s*when\s*:/i,
|
|
261
|
+
/^\s*\/\/\s*then\s*:/i,
|
|
262
|
+
];
|
|
263
|
+
return patterns.every((pat, i) => pat.test(triple[i] ?? ''));
|
|
264
|
+
}
|
|
265
|
+
/** True when `text` contains `when` or `should` as a whole word. */
|
|
266
|
+
function hasWhenOrShould(text) {
|
|
267
|
+
return /(\bwhen\b|\bshould\b)/i.test(text);
|
|
268
|
+
}
|
|
@@ -142,6 +142,7 @@ export function setPresenceLease(input) {
|
|
|
142
142
|
graphRef,
|
|
143
143
|
skill: input.skill,
|
|
144
144
|
...(input.parentWorkflowId ? { parentWorkflowId: input.parentWorkflowId } : {}),
|
|
145
|
+
...(input.mode ? { mode: input.mode } : {}),
|
|
145
146
|
depth: input.depth ?? 0,
|
|
146
147
|
startedAt: now,
|
|
147
148
|
lastHeartbeat: now,
|
|
@@ -96,6 +96,8 @@ function readActiveLeaf(projectRoot, sessionId) {
|
|
|
96
96
|
'never-started',
|
|
97
97
|
'unreadable',
|
|
98
98
|
'stale',
|
|
99
|
+
'queued', // Slice 2026-08-05 fix: stale dispatch entries stuck at 'queued' should
|
|
100
|
+
// not pollute statusline as in-flight leaves.
|
|
99
101
|
]);
|
|
100
102
|
const inFlight = Object.values(index).filter((e) => !terminalStatuses.has(e.status));
|
|
101
103
|
if (inFlight.length === 0)
|
|
@@ -0,0 +1,135 @@
|
|
|
1
|
+
# Test-Style Contract for LLM-Written Unit Tests
|
|
2
|
+
|
|
3
|
+
> **Effective**: rid-2026-08-05-bdd-test-style, peaks-loop v4.0.11+
|
|
4
|
+
> **Audience**: LLM agents (peaks-rd / peaks-qa / downstream consumers) that
|
|
5
|
+
> write new unit tests in projects adopting this contract.
|
|
6
|
+
> **Status**: soft contract — enforced at peaks-qa verification time, not
|
|
7
|
+
> at compile time.
|
|
8
|
+
|
|
9
|
+
This document is the opt-in LLM-facing counterpart to the AAA→BDD
|
|
10
|
+
rewrite that the rid-2026-08-05-bdd-test-style slices ship inside
|
|
11
|
+
peaks-loop. Downstream projects can adopt the same contract by
|
|
12
|
+
importing this file:
|
|
13
|
+
|
|
14
|
+
```ts
|
|
15
|
+
import contract from 'peaks-loop/test-style';
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
The runtime export is the markdown string below; the contract lives
|
|
19
|
+
in the prose, not in code. LLM agents are expected to read this file
|
|
20
|
+
on every test-writing turn.
|
|
21
|
+
|
|
22
|
+
---
|
|
23
|
+
|
|
24
|
+
## 1. The Contract
|
|
25
|
+
|
|
26
|
+
Every new or modified `it()` / `test()` block in `tests/unit/**`
|
|
27
|
+
MUST follow the given-when-then shape.
|
|
28
|
+
|
|
29
|
+
### 1.1 Description (the `it()` / `test()` string-literal)
|
|
30
|
+
|
|
31
|
+
- **Form**: `when X, should Y` — state the precondition in the `when`
|
|
32
|
+
clause and the observable outcome in the `should` clause.
|
|
33
|
+
- **Required word**: must include at least one of `when` or `should`.
|
|
34
|
+
(`when` alone is the precondition; `should` alone is the outcome;
|
|
35
|
+
the natural-language form combines both.)
|
|
36
|
+
- **Anti-pattern**: do NOT use legacy `// arrange:` / `// act:` /
|
|
37
|
+
`// assert:` markers anywhere in the test body.
|
|
38
|
+
|
|
39
|
+
### 1.2 Body (the callback)
|
|
40
|
+
|
|
41
|
+
The first three statements of the callback (after the opening brace,
|
|
42
|
+
before any executable code) MUST be exactly three leading comments:
|
|
43
|
+
|
|
44
|
+
```ts
|
|
45
|
+
// given: <precondition — system / user state>
|
|
46
|
+
// when: <action — what is invoked>
|
|
47
|
+
// then: <expected outcome — what is asserted>
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
The `then:` line corresponds to the assertion(s) that follow.
|
|
51
|
+
|
|
52
|
+
### 1.3 Behavior preservation
|
|
53
|
+
|
|
54
|
+
The migration must be idempotent. A second pass over a test file
|
|
55
|
+
already in BDD form must produce the same output (no duplicate
|
|
56
|
+
comment blocks, no double-tagged descriptions).
|
|
57
|
+
|
|
58
|
+
---
|
|
59
|
+
|
|
60
|
+
## 2. 5-Item Pre-Write Checklist
|
|
61
|
+
|
|
62
|
+
LLM agents writing a new test should run this checklist before
|
|
63
|
+
declaring the test complete:
|
|
64
|
+
|
|
65
|
+
1. **Does the description include `when` or `should`?**
|
|
66
|
+
If no, rewrite the description before writing the body.
|
|
67
|
+
2. **Does the body start with the `// given:` / `// when:` / `// then:`
|
|
68
|
+
triple in that exact order?**
|
|
69
|
+
If no, prepend the missing lines.
|
|
70
|
+
3. **Are there any `// arrange:` / `// act:` / `// assert:` lines?**
|
|
71
|
+
If yes, replace them with the BDD triple.
|
|
72
|
+
4. **Will the test still pass with the BDD rewrite?**
|
|
73
|
+
Run the test, do not just trust the diff. Anti-fake-green rule:
|
|
74
|
+
vitest green is necessary but not sufficient.
|
|
75
|
+
5. **Does the description read as business behavior, not as
|
|
76
|
+
implementation detail?**
|
|
77
|
+
If the description reads like code (e.g. "calls foo with x"),
|
|
78
|
+
rewrite it as observable behavior ("when x is passed, should
|
|
79
|
+
return y").
|
|
80
|
+
|
|
81
|
+
---
|
|
82
|
+
|
|
83
|
+
## 3. Opt-In Adoption (downstream projects)
|
|
84
|
+
|
|
85
|
+
Downstream consumers can adopt the same contract without depending
|
|
86
|
+
on peaks-loop at runtime — this file is the contract. To opt in:
|
|
87
|
+
|
|
88
|
+
```jsonc
|
|
89
|
+
// package.json
|
|
90
|
+
{
|
|
91
|
+
"devDependencies": {
|
|
92
|
+
"peaks-loop": "^4.0.11"
|
|
93
|
+
}
|
|
94
|
+
}
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
```ts
|
|
98
|
+
// In a vitest setup or in a pre-commit hook:
|
|
99
|
+
import contract from 'peaks-loop/test-style';
|
|
100
|
+
// `contract` is the markdown string of this document; surface it to
|
|
101
|
+
// your LLM agent on every test-writing turn via system prompt or
|
|
102
|
+
// tool description.
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
The contract is intentionally **not** a code-level dependency — it
|
|
106
|
+
is a document that LLMs read. Runtime imports are an opt-in
|
|
107
|
+
ergonomic aid for surfacing the contract to a downstream prompt.
|
|
108
|
+
|
|
109
|
+
---
|
|
110
|
+
|
|
111
|
+
## 4. Why Not Enforce in the Test Runner?
|
|
112
|
+
|
|
113
|
+
- vitest's `it()` accepts any string; enforcing description shape at
|
|
114
|
+
runtime would require a custom wrapper around every test, which
|
|
115
|
+
defeats vitest's plugin compatibility.
|
|
116
|
+
- AST-based verification (the peaks-qa `bdd-test-style-verifier`) is
|
|
117
|
+
the chosen gate because it inspects what the LLM wrote, not what
|
|
118
|
+
vitest sees. False positives from string-internal `when` matches
|
|
119
|
+
are eliminated by walking only the description's `StringLiteral`
|
|
120
|
+
and the body callback's leading comments.
|
|
121
|
+
- LLM-authored tests are the primary audience. Humans writing tests
|
|
122
|
+
are not blocked by this contract (peaks-qa is the gate, not vitest).
|
|
123
|
+
|
|
124
|
+
---
|
|
125
|
+
|
|
126
|
+
## 5. Author & Change Control
|
|
127
|
+
|
|
128
|
+
- **Author**: SquabbyZ (`601709253@qq.com`) — sole-author per project
|
|
129
|
+
red rule.
|
|
130
|
+
- **Change control**: any edit to this file MUST go through a
|
|
131
|
+
peaks-rd slice with peaks-qa acceptance; treat the contract text as
|
|
132
|
+
load-bearing for downstream LLM behavior.
|
|
133
|
+
- **Related**: `.peaks/_runtime/2026-08-04-session-3fe1be/sc/2026-08-05-bdd-test-style-rid-design.md`
|
|
134
|
+
(the design doc), `scripts/migrate-to-bdd.mjs` (the AST migrator),
|
|
135
|
+
`src/services/qa/bdd-test-style-verifier.ts` (the verifier).
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "peaks-loop",
|
|
3
|
-
"version": "4.0.
|
|
3
|
+
"version": "4.0.11",
|
|
4
4
|
"description": "Loop Engineering CLI — workflow primitive / loop guards / evaluators / slice orchestration",
|
|
5
5
|
"author": "SquabbyZ",
|
|
6
6
|
"keywords": [
|
|
@@ -82,6 +82,8 @@
|
|
|
82
82
|
"agents/**",
|
|
83
83
|
"schemas/*.json",
|
|
84
84
|
".claude-plugin/**",
|
|
85
|
+
"config/eslint/.peaks-rules.cjs",
|
|
86
|
+
"docs/test-style-contract.md",
|
|
85
87
|
"README.md",
|
|
86
88
|
"README-en.md",
|
|
87
89
|
"CHANGELOG.md",
|
|
@@ -98,9 +100,9 @@
|
|
|
98
100
|
"headroom-ai": "0.22.4",
|
|
99
101
|
"yaml": "^2.9.0",
|
|
100
102
|
"zod": "^3.25.76",
|
|
101
|
-
"peaks-loop-shared": "0.0.41",
|
|
102
103
|
"peaks-loop-shared-channel": "0.0.19",
|
|
103
|
-
"peaks-loop-mut": "0.1.15"
|
|
104
|
+
"peaks-loop-mut": "0.1.15",
|
|
105
|
+
"peaks-loop-shared": "0.0.42"
|
|
104
106
|
},
|
|
105
107
|
"peerDependencies": {
|
|
106
108
|
"@alibaba-group/open-code-review": "1.3.1"
|
|
@@ -65,4 +65,20 @@ The dispatch CLI (`peaks sub-agent dispatch`) automatically prepends a Test Tool
|
|
|
65
65
|
|
|
66
66
|
If the framework is not obvious from `package.json#scripts.test`, the sub-agent should run `peaks test --json` to introspect the resolved framework + argv before picking a runner.
|
|
67
67
|
|
|
68
|
-
See the block constant at `src/services/dispatch/test-tool-detection.ts` for the verbatim text.
|
|
68
|
+
See the block constant at `src/services/dispatch/test-tool-detection.ts` for the verbatim text.
|
|
69
|
+
|
|
70
|
+
## BDD Test Style Verification (effective rid-2026-08-05-bdd-test-style, v4.0.11+)
|
|
71
|
+
|
|
72
|
+
When you (peaks-qa) verify a slice, you MUST run the BDD test-style verifier on every new or modified `tests/unit/**/*.test.ts` file in the slice's git diff. Use:
|
|
73
|
+
|
|
74
|
+
```bash
|
|
75
|
+
node -e "
|
|
76
|
+
const { verifyBddStyle } = await import('./src/services/qa/bdd-test-style-verifier.ts');
|
|
77
|
+
const { execSync } = require('node:child_process');
|
|
78
|
+
const files = execSync('git diff --name-only HEAD~1 -- tests/unit', { encoding: 'utf8' })
|
|
79
|
+
.split('\n').filter(f => f.endsWith('.test.ts'));
|
|
80
|
+
console.log(JSON.stringify(verifyBddStyle({ projectRoot: '.', testFiles: files })));
|
|
81
|
+
"
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
If the verifier returns `ok: false`, your verdict MUST be `failed: bdd-style-violation` with the structured reason from the verifier (do NOT mark the slice as passing).
|
|
@@ -147,4 +147,24 @@ Touch only what you must. Clean up only your own mess. When editing existing cod
|
|
|
147
147
|
Define success criteria. Loop until verified. "Add validation" → write tests for invalid inputs, then make them pass. "Fix the bug" → write a test that reproduces it, then make it pass. For multi-step tasks, state a brief plan with verify checkpoints. Strong success criteria let you loop independently. Weak criteria require constant clarification.
|
|
148
148
|
```
|
|
149
149
|
|
|
150
|
-
Sub-agents MUST NOT silently drop this block. The regression test `tests/unit/skills/karpathy-prompt-injection.test.ts` asserts this block is present. The canonical skill id for the full guidelines text is `andrej-karpathy-skills:karpathy-guidelines`.
|
|
150
|
+
Sub-agents MUST NOT silently drop this block. The regression test `tests/unit/skills/karpathy-prompt-injection.test.ts` asserts this block is present. The canonical skill id for the full guidelines text is `andrej-karpathy-skills:karpathy-guidelines`.
|
|
151
|
+
|
|
152
|
+
## BDD Test Style Contract (effective rid-2026-08-05-bdd-test-style, v4.0.11+)
|
|
153
|
+
|
|
154
|
+
When you (the LLM sub-agent) write new or modified unit tests in `tests/unit/**`, every `it()` / `test()` block MUST follow the given-when-then contract:
|
|
155
|
+
|
|
156
|
+
1. The first string-literal argument of `it()` / `test()` MUST describe business behavior in the form `when X, should Y` — must include either the word "when" (state / pre-condition) or "should" (observable outcome).
|
|
157
|
+
2. The body callback (the second argument) MUST start with exactly 3 leading comments:
|
|
158
|
+
```typescript
|
|
159
|
+
// given: <precondition — system / user state>
|
|
160
|
+
// when: <action — what is invoked>
|
|
161
|
+
// then: <expected outcome — what is asserted>
|
|
162
|
+
```
|
|
163
|
+
3. Legacy `// arrange:` / `// act:` / `// assert:` AAA markers MUST NOT appear in tests you write.
|
|
164
|
+
|
|
165
|
+
Run `node scripts/migrate-to-bdd.mjs --dry-run <file>` before writing new tests to inspect the contract, or use the Slice B verifier's rule directly:
|
|
166
|
+
|
|
167
|
+
- description: must contain `/(\bwhen\b|\bshould\b)/`
|
|
168
|
+
- first 3 body comments: must match `/^\s*\/\/\s*given\s*:/`, `/^\s*\/\/\s*when\s*:/`, `/^\s*\/\/\s*then\s*:/`
|
|
169
|
+
|
|
170
|
+
If your tests fail peaks-qa's `bdd-test-style-verifier` (see Slice B), the slice will be returned-to-rd. Fix the violations, do NOT bypass the check.
|