jev-agent-tools 0.1.4 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +106 -1
- package/CONTRIBUTING.md +43 -0
- package/README.md +58 -17
- package/SECURITY.md +43 -0
- package/dist/adapters/analysis-context.js +75 -0
- package/dist/adapters/ask-files.js +198 -0
- package/dist/adapters/ask-proof.js +200 -0
- package/dist/adapters/ask-syntax.js +385 -0
- package/dist/adapters/canonical-path.js +17 -0
- package/dist/adapters/command.js +234 -0
- package/dist/adapters/docs.js +192 -0
- package/dist/adapters/evidence-context.js +119 -0
- package/dist/adapters/exec.js +207 -0
- package/dist/adapters/files.js +418 -0
- package/dist/adapters/find.js +150 -0
- package/dist/adapters/git-base.js +32 -0
- package/dist/adapters/git-inventory.js +71 -0
- package/dist/adapters/git.js +483 -0
- package/dist/adapters/locate-file.js +197 -0
- package/dist/adapters/output-lines.js +46 -0
- package/dist/adapters/private-storage.js +106 -0
- package/dist/adapters/risk-callers.js +429 -0
- package/dist/adapters/runner-version.js +78 -0
- package/dist/adapters/shell.js +92 -0
- package/dist/adapters/syntax.js +187 -0
- package/dist/adapters/test-inventory.js +139 -0
- package/dist/adapters/usage.js +20 -0
- package/dist/adapters/utf8.js +47 -0
- package/dist/configuration.js +267 -0
- package/dist/constants.js +140 -0
- package/dist/core/ask-closure.js +282 -0
- package/dist/core/ask-proof.js +1 -0
- package/dist/core/ask-references.js +278 -0
- package/dist/core/asks.js +507 -0
- package/dist/core/batches.js +65 -0
- package/dist/core/command-output.js +224 -0
- package/dist/core/diff.js +178 -0
- package/dist/core/docs.js +302 -0
- package/dist/core/find.js +108 -0
- package/dist/core/git.js +1 -0
- package/dist/core/imports.js +550 -0
- package/dist/core/integrity.js +45 -0
- package/dist/core/lexical.js +132 -0
- package/dist/core/locate.js +169 -0
- package/dist/core/output.js +137 -0
- package/dist/core/pointer.js +29 -0
- package/dist/core/result-report.js +302 -0
- package/dist/core/risk-callers.js +851 -0
- package/dist/core/runner-version.js +45 -0
- package/dist/core/secret-path.js +34 -0
- package/dist/core/sections.js +230 -0
- package/dist/core/state.js +51 -0
- package/dist/core/syntax.js +1 -0
- package/dist/core/test-commands.js +334 -0
- package/dist/core/test-coverage.js +74 -0
- package/dist/core/test-discovery.js +1382 -0
- package/dist/core/test-evidence.js +527 -0
- package/dist/core/test-state.js +81 -0
- package/dist/core/truncate.js +12 -0
- package/dist/core/units.js +349 -0
- package/dist/describe.js +23 -0
- package/dist/guide.js +33 -0
- package/dist/host.js +24 -0
- package/dist/jev/client.js +456 -0
- package/dist/jev/pool.js +54 -0
- package/dist/jev/types.js +1 -0
- package/dist/mcp/main.js +124 -0
- package/dist/mcp/protocol.js +210 -0
- package/dist/mcp/tools.js +129 -0
- package/dist/presets/docs.js +62 -0
- package/dist/presets/risk.js +179 -0
- package/dist/presets/spec.js +81 -0
- package/dist/presets/witnesses.js +249 -0
- package/dist/render.js +114 -0
- package/dist/report-schema.js +1356 -0
- package/dist/result-types.js +1 -0
- package/dist/result.js +3 -0
- package/dist/runtime.js +1 -0
- package/dist/session.js +147 -0
- package/dist/texts/ask-files.js +3 -0
- package/dist/texts/ask.js +4 -0
- package/dist/texts/check-diff.js +20 -0
- package/dist/texts/configuration.js +1 -0
- package/dist/texts/find.js +19 -0
- package/dist/texts/guide.js +3 -0
- package/dist/texts/instructions.js +72 -0
- package/dist/texts/locate.js +15 -0
- package/dist/texts/select-tests.js +4 -0
- package/dist/tools/ask-files.js +450 -0
- package/dist/tools/ask-schema.js +70 -0
- package/dist/tools/ask.js +1147 -0
- package/dist/tools/check-diff.js +594 -0
- package/dist/tools/docs-check.js +408 -0
- package/dist/tools/find.js +682 -0
- package/dist/tools/locate.js +602 -0
- package/dist/tools/review-report.js +230 -0
- package/dist/tools/select-tests.js +821 -0
- package/dist/tools/spec-check.js +263 -0
- package/docs/adr/0001-strict-typescript-pure-core-offline-tests.md +31 -0
- package/docs/adr/0002-one-http-protocol-across-hosts.md +17 -0
- package/docs/adr/0003-explicit-scope-conservative-automation.md +19 -0
- package/docs/adr/0004-compiled-typed-intents.md +19 -0
- package/docs/adr/0005-evidence-construction-before-judgment.md +19 -0
- package/docs/adr/0006-visible-uncertainty-constrained-controls.md +21 -0
- package/docs/adr/0007-bounded-evidence-visible-limits.md +21 -0
- package/docs/adr/0008-static-test-discovery-conservative-plans.md +19 -0
- package/docs/adr/0009-session-cache-requested-model-identity.md +17 -0
- package/docs/adr/0010-mcp-server-thin-host.md +23 -0
- package/docs/agent-instructions.md +120 -0
- package/docs/design.md +16 -4
- package/docs/mcp.md +233 -0
- package/docs/tools/jev_ask.md +8 -5
- package/docs/tools/jev_ask_files.md +2 -1
- package/docs/tools/jev_check_diff.md +4 -1
- package/docs/tools/jev_find_files.md +2 -1
- package/docs/tools/jev_locate_in_file.md +5 -0
- package/docs/tools/jev_select_tests.md +4 -1
- package/package.json +19 -4
- package/rules/jev-ask.md +22 -1
- package/server.json +57 -0
- package/src/adapters/ask-files.ts +11 -3
- package/src/adapters/ask-proof.ts +69 -11
- package/src/adapters/canonical-path.ts +18 -0
- package/src/adapters/command.ts +102 -36
- package/src/adapters/docs.ts +33 -14
- package/src/adapters/evidence-context.ts +169 -0
- package/src/adapters/exec.ts +226 -0
- package/src/adapters/files.ts +146 -16
- package/src/adapters/find.ts +37 -7
- package/src/adapters/git-base.ts +7 -1
- package/src/adapters/git.ts +61 -8
- package/src/adapters/locate-file.ts +51 -9
- package/src/adapters/private-storage.ts +155 -0
- package/src/adapters/risk-callers.ts +7 -2
- package/src/adapters/shell.ts +113 -0
- package/src/adapters/test-inventory.ts +12 -4
- package/src/configuration.ts +55 -14
- package/src/constants.ts +37 -5
- package/src/core/ask-references.ts +262 -146
- package/src/core/asks.ts +79 -7
- package/src/core/command-output.ts +17 -1
- package/src/core/import-boundaries.ts +8 -3
- package/src/core/locate.ts +8 -5
- package/src/core/output.ts +34 -0
- package/src/core/result-report.ts +410 -0
- package/src/core/secret-path.ts +37 -0
- package/src/core/state.ts +8 -1
- package/src/core/units.ts +3 -2
- package/src/host.ts +11 -0
- package/src/index.ts +3 -0
- package/src/jev/client.ts +66 -16
- package/src/jev/types.ts +24 -3
- package/src/mcp/main.ts +135 -0
- package/src/mcp/protocol.ts +332 -0
- package/src/mcp/tools.ts +179 -0
- package/src/render.ts +109 -0
- package/src/report-schema.ts +1380 -0
- package/src/result-types.ts +234 -0
- package/src/result.ts +4 -1
- package/src/runtime.ts +6 -0
- package/src/session.ts +59 -0
- package/src/setup.ts +13 -5
- package/src/texts/ask-files.ts +4 -1
- package/src/texts/ask.ts +8 -1
- package/src/texts/check-diff.ts +7 -4
- package/src/texts/find.ts +8 -2
- package/src/texts/guide.ts +8 -16
- package/src/texts/instructions.ts +98 -0
- package/src/texts/locate.ts +8 -2
- package/src/texts/run-end.ts +2 -2
- package/src/texts/select-tests.ts +4 -1
- package/src/tools/ask-files.ts +311 -18
- package/src/tools/ask.ts +722 -95
- package/src/tools/check-diff.ts +337 -31
- package/src/tools/docs-check.ts +241 -38
- package/src/tools/find.ts +389 -29
- package/src/tools/locate.ts +387 -25
- package/src/tools/review-report.ts +308 -0
- package/src/tools/select-tests.ts +484 -23
- package/src/tools/spec-check.ts +194 -19
|
@@ -0,0 +1,263 @@
|
|
|
1
|
+
import { resolveEvidenceContext, withEvidenceContext, } from "../adapters/evidence-context.js";
|
|
2
|
+
import { collectFiles } from "../adapters/files.js";
|
|
3
|
+
import { collectUnits } from "../adapters/git.js";
|
|
4
|
+
import { resolveBase } from "../adapters/git-base.js";
|
|
5
|
+
import { CHOICE_MAX_OPTIONS, STATE_MAX_CHARS, TIMEOUT_MS, } from "../constants.js";
|
|
6
|
+
import { buildEnvelope, } from "../core/output.js";
|
|
7
|
+
import { known, } from "../core/result-report.js";
|
|
8
|
+
import { prepareSpecCheck, readSpecJudgment, } from "../presets/spec.js";
|
|
9
|
+
import { NOT_CONFIGURED } from "../texts/configuration.js";
|
|
10
|
+
import { ReviewReport, reportMetrics } from "./review-report.js";
|
|
11
|
+
export async function runSpecCheck(deps, input) {
|
|
12
|
+
const started = performance.now();
|
|
13
|
+
const report = new ReviewReport();
|
|
14
|
+
const admission = input.evidenceContext
|
|
15
|
+
? undefined
|
|
16
|
+
: await resolveEvidenceContext(input.cwd, undefined, {
|
|
17
|
+
exec: deps.exec,
|
|
18
|
+
signal: input.signal,
|
|
19
|
+
origin: deps.evidenceOrigin,
|
|
20
|
+
});
|
|
21
|
+
const evidenceContext = input.evidenceContext ?? admission?.context;
|
|
22
|
+
if (!evidenceContext)
|
|
23
|
+
throw new Error("Evidence context was not established");
|
|
24
|
+
evidenceContext.requestedBase = input.base ?? "HEAD";
|
|
25
|
+
const judgments = [];
|
|
26
|
+
const findings = [];
|
|
27
|
+
const limitations = [];
|
|
28
|
+
const unchecked = [];
|
|
29
|
+
let budget;
|
|
30
|
+
let sent = 0;
|
|
31
|
+
let incomplete = false;
|
|
32
|
+
let emptyBase;
|
|
33
|
+
const finish = (refusal, cause = "internal_error") => {
|
|
34
|
+
const answers = findings.map((finding) => ({
|
|
35
|
+
label: finding.kind === "requirement"
|
|
36
|
+
? `${input.specPath}:${finding.requirement?.start}-${finding.requirement?.end} ${finding.label} — requirement violated`
|
|
37
|
+
: `${finding.unit?.file}:${finding.unit?.afterRange?.start ?? finding.unit?.beforeRange?.start}-${finding.unit?.afterRange?.end ?? finding.unit?.beforeRange?.end} ${finding.label} — external behavior absent from the specification`,
|
|
38
|
+
value: {
|
|
39
|
+
head: finding.kind === "requirement" ? "violates" : finding.label,
|
|
40
|
+
p: finding.probability,
|
|
41
|
+
},
|
|
42
|
+
band: incomplete ? "unsure" : "verdict",
|
|
43
|
+
...(incomplete
|
|
44
|
+
? { reason: "incomplete change evidence: read the omitted pieces" }
|
|
45
|
+
: {}),
|
|
46
|
+
}));
|
|
47
|
+
const envelope = buildEnvelope({
|
|
48
|
+
answers,
|
|
49
|
+
limitations: emptyBase !== undefined && !refusal
|
|
50
|
+
? [
|
|
51
|
+
...limitations,
|
|
52
|
+
{
|
|
53
|
+
fact: `no changed units against ${emptyBase}`,
|
|
54
|
+
next: "nothing judged; pass base= or check the working directory",
|
|
55
|
+
},
|
|
56
|
+
]
|
|
57
|
+
: limitations,
|
|
58
|
+
unchecked,
|
|
59
|
+
refusal,
|
|
60
|
+
budget,
|
|
61
|
+
...(!refusal &&
|
|
62
|
+
!budget &&
|
|
63
|
+
!unchecked.length &&
|
|
64
|
+
!findings.length &&
|
|
65
|
+
emptyBase === undefined &&
|
|
66
|
+
report.items.size > 0 &&
|
|
67
|
+
[...report.items.values()].every((item) => item.treatment === "judged")
|
|
68
|
+
? {
|
|
69
|
+
lines: [
|
|
70
|
+
{
|
|
71
|
+
type: "list",
|
|
72
|
+
title: "spec: no violations or drift reported",
|
|
73
|
+
items: [],
|
|
74
|
+
},
|
|
75
|
+
],
|
|
76
|
+
}
|
|
77
|
+
: {}),
|
|
78
|
+
yield: {
|
|
79
|
+
calls: judgments.reduce((n, j) => n + (j.calls ?? 0), 0),
|
|
80
|
+
questions: judgments.reduce((n, j) => n + (j.questions ?? 0), 0),
|
|
81
|
+
costUsd: judgments.length && judgments.every((j) => j.usage !== undefined)
|
|
82
|
+
? judgments.reduce((n, j) => n + (j.usage?.costUsd ?? 0), 0)
|
|
83
|
+
: undefined,
|
|
84
|
+
cacheHits: judgments.reduce((n, j) => n + (j.cacheHits ?? 0), 0),
|
|
85
|
+
cacheRequests: judgments.reduce((n, j) => n + (j.cacheRequests ?? 0), 0),
|
|
86
|
+
elapsedMs: performance.now() - started,
|
|
87
|
+
},
|
|
88
|
+
});
|
|
89
|
+
if (emptyBase !== undefined)
|
|
90
|
+
report.diagnose("no_changed_units", `No changed units against ${emptyBase}`, emptyBase, [], false);
|
|
91
|
+
if (refusal) {
|
|
92
|
+
report.refusal = cause !== "not_configured";
|
|
93
|
+
report.diagnose(cause, refusal, input.specPath);
|
|
94
|
+
}
|
|
95
|
+
const result = report.build("jev_check_diff", evidenceContext, reportMetrics(envelope));
|
|
96
|
+
return { ok: !refusal, envelope, judgments, findings, result };
|
|
97
|
+
};
|
|
98
|
+
if (!input.specPath)
|
|
99
|
+
return finish("spec_path required: no specification to check against.", "missing_required");
|
|
100
|
+
if (!deps.client)
|
|
101
|
+
return finish(NOT_CONFIGURED, "not_configured");
|
|
102
|
+
const comparison = await resolveBase(deps.exec, input.cwd, input.base, input.signal);
|
|
103
|
+
if (!comparison.ok)
|
|
104
|
+
return finish(comparison.error, comparison.cause ?? "invalid_base");
|
|
105
|
+
const base = comparison.base;
|
|
106
|
+
evidenceContext.resolvedBase = base;
|
|
107
|
+
const root = await deps.exec("git", ["rev-parse", "--show-toplevel"], {
|
|
108
|
+
cwd: input.cwd,
|
|
109
|
+
timeout: TIMEOUT_MS,
|
|
110
|
+
signal: input.signal,
|
|
111
|
+
});
|
|
112
|
+
if (root.code || root.killed)
|
|
113
|
+
return finish("Repository root not found.", root.killed ? "cancelled" : "git_failure");
|
|
114
|
+
const cwd = root.stdout.trim();
|
|
115
|
+
if (evidenceContext.effectiveRoot)
|
|
116
|
+
evidenceContext.effectiveRoot.path = cwd;
|
|
117
|
+
const [collected, specification] = await Promise.all([
|
|
118
|
+
collectUnits(deps.exec, { cwd, base, signal: input.signal }),
|
|
119
|
+
collectFiles(cwd, [input.specPath], input.signal, { exec: deps.exec }),
|
|
120
|
+
]);
|
|
121
|
+
if (!collected.ok)
|
|
122
|
+
return finish(collected.error, collected.cause ?? "git_failure");
|
|
123
|
+
if (!specification.ok)
|
|
124
|
+
return finish(specification.error, specification.cause ?? "file_unavailable");
|
|
125
|
+
const text = Object.values(specification.files)[0];
|
|
126
|
+
if (!text)
|
|
127
|
+
return finish("The specification is empty: nothing to check.", "empty_required");
|
|
128
|
+
for (const limit of collected.limits)
|
|
129
|
+
limitations.push({
|
|
130
|
+
fact: `${limit.file} : ${limit.kind}`,
|
|
131
|
+
next: "Read the complete change before concluding.",
|
|
132
|
+
});
|
|
133
|
+
const unavailableUnits = collected.units.filter((unit) => unit.before === null && unit.after === null);
|
|
134
|
+
const units = collected.units.filter((unit) => unit.before !== null || unit.after !== null);
|
|
135
|
+
for (const unit of unavailableUnits)
|
|
136
|
+
unchecked.push(`${unit.file} ${unit.name} (changed source unavailable)`);
|
|
137
|
+
const prepared = prepareSpecCheck(text, units);
|
|
138
|
+
for (const requirement of prepared.requirements)
|
|
139
|
+
report.expect(`spec:${requirement.id}`, requirement.label, "requirement");
|
|
140
|
+
report.expect("spec:drift", "Specification drift pointer", "pointer");
|
|
141
|
+
report.inventories.push({
|
|
142
|
+
id: "spec-requirements",
|
|
143
|
+
kind: "sections",
|
|
144
|
+
rules: ["### REQ- headings in the supplied specification"],
|
|
145
|
+
restrictions: [input.specPath],
|
|
146
|
+
discovered: known(prepared.requirements.length),
|
|
147
|
+
considered: known(prepared.requirements.length),
|
|
148
|
+
scopeRestricted: true,
|
|
149
|
+
criteria: [],
|
|
150
|
+
});
|
|
151
|
+
report.inventories.push({
|
|
152
|
+
id: "changed-units",
|
|
153
|
+
kind: "units",
|
|
154
|
+
rules: ["Changed source units against resolved base"],
|
|
155
|
+
restrictions: [],
|
|
156
|
+
discovered: known(collected.units.length),
|
|
157
|
+
considered: known(units.length),
|
|
158
|
+
scopeRestricted: false,
|
|
159
|
+
criteria: [],
|
|
160
|
+
});
|
|
161
|
+
if (!collected.units.length && prepared.requirements.length)
|
|
162
|
+
report.items.clear();
|
|
163
|
+
const reportIds = [...report.items.keys()];
|
|
164
|
+
for (const unit of unavailableUnits)
|
|
165
|
+
report.diagnose("binary_or_non_utf8", "Changed source unavailable", unit.file, reportIds, units.length === 0, [`${unit.file} ${unit.name}`]);
|
|
166
|
+
for (const limit of collected.limits)
|
|
167
|
+
report.diagnose(limit.kind === "secret_pattern" ? "secret_pattern" : "collection_omitted", limit.kind, limit.file, reportIds, units.length === 0, ["unit" in limit ? `${limit.file} ${limit.unit}` : limit.file]);
|
|
168
|
+
report.missingWork =
|
|
169
|
+
unavailableUnits.length > 0 || collected.limits.length > 0;
|
|
170
|
+
if (!prepared.requirements.length)
|
|
171
|
+
return finish("No ### REQ-… requirements in the specification: nothing to check.", "missing_required");
|
|
172
|
+
if (!collected.units.length) {
|
|
173
|
+
emptyBase = base;
|
|
174
|
+
return finish();
|
|
175
|
+
}
|
|
176
|
+
if (unchecked.length) {
|
|
177
|
+
unchecked.push(...prepared.requirements.map((requirement) => `${requirement.label} (changed source unavailable)`), "drift (changed source unavailable)");
|
|
178
|
+
}
|
|
179
|
+
if (prepared.tableWarning) {
|
|
180
|
+
report.diagnose("unsupported_syntax", "Specification Markdown table interpretation is uncalibrated", input.specPath, [], false);
|
|
181
|
+
limitations.push({
|
|
182
|
+
fact: "The specification contains a Markdown table.",
|
|
183
|
+
next: "Read the table requirements: their interpretation is uncalibrated.",
|
|
184
|
+
});
|
|
185
|
+
}
|
|
186
|
+
incomplete = report.missingWork;
|
|
187
|
+
const state = withEvidenceContext(prepared.state, evidenceContext);
|
|
188
|
+
if (units.length + 1 > CHOICE_MAX_OPTIONS ||
|
|
189
|
+
JSON.stringify(state).length > STATE_MAX_CHARS)
|
|
190
|
+
return finish(`State or spec pointer exceeds the limits; compare against a closer base (STATE_MAX_CHARS=${STATE_MAX_CHARS}).`, "evidence_too_large");
|
|
191
|
+
if (!units.length)
|
|
192
|
+
return finish();
|
|
193
|
+
const judgment = await deps.client.judge(state, prepared.questions, {
|
|
194
|
+
signal: input.signal,
|
|
195
|
+
...deps.runtime.session.requestGate(),
|
|
196
|
+
admissionCause: () => budget?.kind === "session" ? "session_budget" : "call_budget",
|
|
197
|
+
beforeRequest(questionCount) {
|
|
198
|
+
if (input.maxCalls !== undefined && sent >= input.maxCalls) {
|
|
199
|
+
budget = {
|
|
200
|
+
kind: "max_calls",
|
|
201
|
+
message: `max_calls=${input.maxCalls} reached`,
|
|
202
|
+
};
|
|
203
|
+
return { ok: false, error: budget.message };
|
|
204
|
+
}
|
|
205
|
+
const admitted = deps.runtime.session.admit(questionCount);
|
|
206
|
+
if (!admitted.ok)
|
|
207
|
+
budget = { kind: "session", message: admitted.error };
|
|
208
|
+
else
|
|
209
|
+
sent++;
|
|
210
|
+
return admitted;
|
|
211
|
+
},
|
|
212
|
+
onUsage: (usage) => deps.runtime.session.recordUsage(usage),
|
|
213
|
+
});
|
|
214
|
+
judgments.push(judgment);
|
|
215
|
+
if (!judgment.ok)
|
|
216
|
+
report.failure(judgment, reportIds, budget
|
|
217
|
+
? budget.kind === "session"
|
|
218
|
+
? "session_budget"
|
|
219
|
+
: "call_budget"
|
|
220
|
+
: undefined);
|
|
221
|
+
else {
|
|
222
|
+
for (const requirement of prepared.requirements)
|
|
223
|
+
report.answer(`spec:${requirement.id}`, judgment.answers[requirement.id], { band: incomplete ? "unsure" : "verdict" });
|
|
224
|
+
const driftAnswer = judgment.answers.drift;
|
|
225
|
+
const selectedUnit = driftAnswer?.type === "choice"
|
|
226
|
+
? units.find((unit) => unit.id === driftAnswer.choice)
|
|
227
|
+
: undefined;
|
|
228
|
+
const pointer = report.items.get("spec:drift");
|
|
229
|
+
if (pointer && selectedUnit)
|
|
230
|
+
report.items.set("spec:drift", {
|
|
231
|
+
...pointer,
|
|
232
|
+
label: `${selectedUnit.file} ${selectedUnit.name} specification drift`,
|
|
233
|
+
});
|
|
234
|
+
const selectedProbability = driftAnswer?.type === "choice"
|
|
235
|
+
? driftAnswer.probabilities[driftAnswer.choice]
|
|
236
|
+
: undefined;
|
|
237
|
+
report.answer("spec:drift", driftAnswer, {
|
|
238
|
+
band: incomplete ? "unsure" : "verdict",
|
|
239
|
+
...(driftAnswer?.type === "choice" && selectedProbability !== undefined
|
|
240
|
+
? {
|
|
241
|
+
value: {
|
|
242
|
+
head: driftAnswer.choice,
|
|
243
|
+
p: selectedProbability,
|
|
244
|
+
},
|
|
245
|
+
}
|
|
246
|
+
: {}),
|
|
247
|
+
});
|
|
248
|
+
}
|
|
249
|
+
if (!judgment.ok) {
|
|
250
|
+
unchecked.push(...prepared.requirements.map((req) => `${req.label} (${judgment.error})`), `drift (${judgment.error})`);
|
|
251
|
+
return finish();
|
|
252
|
+
}
|
|
253
|
+
for (const requirement of prepared.requirements) {
|
|
254
|
+
const answer = judgment.answers[requirement.id];
|
|
255
|
+
if (!answer || answer.type === "unjudged")
|
|
256
|
+
unchecked.push(requirement.label);
|
|
257
|
+
}
|
|
258
|
+
const drift = judgment.answers.drift;
|
|
259
|
+
if (!drift || drift.type === "unjudged")
|
|
260
|
+
unchecked.push("drift");
|
|
261
|
+
findings.push(...readSpecJudgment(prepared, judgment.answers));
|
|
262
|
+
return finish();
|
|
263
|
+
}
|
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
# 0001 — Strict TypeScript, a pure core and offline conformance tests
|
|
2
|
+
|
|
3
|
+
Status: Accepted
|
|
4
|
+
|
|
5
|
+
## Context
|
|
6
|
+
|
|
7
|
+
The same extension runs inside pi and omp. Host-specific implementations would duplicate judgment policy and make deterministic checks harder to maintain. Optional native capabilities must not become mandatory runtime dependencies.
|
|
8
|
+
|
|
9
|
+
## Decision
|
|
10
|
+
|
|
11
|
+
Use one strict TypeScript package with erasable syntax, Node APIs and fetch rather than Bun-only APIs. Pure core transformations and question compilation consume and return values; thin adapters own host, filesystem, Git, network and clock effects. Tools compose collection, evidence construction, requests and result rendering. Discriminated unions represent expected refusals, unjudged work and abstention as values rather than exceptional control flow.
|
|
12
|
+
|
|
13
|
+
Keep limits and thresholds in src/constants.ts. Batch questions for the same evidence state, run independent work concurrently under the shared limiter, and bound reads before loading content into memory.
|
|
14
|
+
|
|
15
|
+
Enforce downward import layers, including erased type dependencies:
|
|
16
|
+
|
|
17
|
+
- core/ imports core/, constants.ts, result.ts, result-types.ts and type-only contracts from jev/types.ts.
|
|
18
|
+
- presets/ and adapters/ each import their own layer, core/ and those neutral modules; adapters/ does not import presets/ or tools/.
|
|
19
|
+
- jev/ imports its own layer, core/, constants.ts, result.ts and result-types.ts.
|
|
20
|
+
- texts/ imports its own layer, constants.ts and presets/; it has no external dependencies.
|
|
21
|
+
- constants.ts and result-types.ts import nothing. result.ts may import the dependency-free result-types.ts contract. Root integration modules and tools/ compose the layers.
|
|
22
|
+
|
|
23
|
+
Reject all cycles, including type-only cycles, and nonliteral module loading. Shared contracts belong below their consumers. Pure layers do not import external modules, with explicit core exceptions for pure node:path functions and erased import type contracts from @ast-grep/napi. Inline type specifiers that retain runtime loading do not qualify. Filesystem, process and network modules, direct fetch calls and Node builtin-module loaders are forbidden in the core; import checks are not an exhaustive proof against indirect global effects.
|
|
24
|
+
|
|
25
|
+
Keep @ast-grep/napi, @ast-grep/lang-python and @ff-labs/fff-node optional. Load native parsing and search at the edges, use the available fallback and report missing capabilities rather than claiming equivalent syntax precision.
|
|
26
|
+
|
|
27
|
+
Run deterministic offline conformance checks with tsc --noEmit, the import checker, Biome and node --test. Use constructed behavior and boundary cases and local HTTP fixtures for transport failures, retries and timeouts; no Jev credentials are required.
|
|
28
|
+
|
|
29
|
+
## Consequences
|
|
30
|
+
|
|
31
|
+
This separates effects from policy and makes host-independent behavior reproducible. Missing optional acceleration can reduce parsing or search capability, so limits stay visible. Native optimization is not a reason to introduce a separate judgment path or an extension-specific daemon.
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
# 0002 — One HTTP protocol across hosts
|
|
2
|
+
|
|
3
|
+
Status: Accepted
|
|
4
|
+
|
|
5
|
+
## Context
|
|
6
|
+
|
|
7
|
+
pi and omp provide different integration APIs. Delegating judgment to a host's chat model would change the protocol and behavior according to the host.
|
|
8
|
+
|
|
9
|
+
## Decision
|
|
10
|
+
|
|
11
|
+
Use one HTTP client compatible with the Jev API format in both hosts. Read the user-supplied endpoint and Bearer credential from JEV_TOOLS_URL and JEV_TOOLS_API_KEY and the requested model from JEV_TOOLS_MODEL, defaulting to openjev.
|
|
12
|
+
|
|
13
|
+
Without the endpoint or key, keep tools registered and explain the missing configuration; disable the automatic documentation check. Never fall back to a chat model or a host-internal judgment service.
|
|
14
|
+
|
|
15
|
+
## Consequences
|
|
16
|
+
|
|
17
|
+
Users choose the service and must review which evidence it receives. Host adapters manage registration and lifecycle, not a second judgment implementation. Configuration failures remain explicit rather than producing superficially equivalent answers.
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# 0003 — Explicit scope and conservative automation
|
|
2
|
+
|
|
3
|
+
Status: Accepted
|
|
4
|
+
|
|
5
|
+
## Context
|
|
6
|
+
|
|
7
|
+
Repository navigation, change review and test selection involve different evidence and next actions. Broad automatic suggestions could be mistaken for complete knowledge of the repository.
|
|
8
|
+
|
|
9
|
+
## Decision
|
|
10
|
+
|
|
11
|
+
Expose six explicit tools: jev_ask for one combined situation, jev_ask_files for independent judgments of files, jev_find_files for discovery, jev_locate_in_file for ranges, jev_check_diff for fixed reviews and jev_select_tests for existing test plans. Hide jev_find_files in omp when its native find tool is active.
|
|
12
|
+
|
|
13
|
+
Allow only one automatic run-end documentation check of a dirty tree. A flagged stale sentence can request another turn; errors do not block the host or cause an automatic review loop. JEV_TOOLS_AUTO_DOCS=0 disables the hook.
|
|
14
|
+
|
|
15
|
+
Do not infer required new test scenarios, missing documentation or arbitrary companion edits from incomplete collection. Test selection and residual coverage checks describe only the discovered inventory.
|
|
16
|
+
|
|
17
|
+
## Consequences
|
|
18
|
+
|
|
19
|
+
Callers retain control of consequential actions and must read or execute the evidence indicated by results. A clean review is not a completeness guarantee. Automatic documentation checking targets existing sentences, not an inferred obligation to write new documentation.
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# 0004 — Compile typed intents instead of accepting raw question maps
|
|
2
|
+
|
|
3
|
+
Status: Accepted
|
|
4
|
+
|
|
5
|
+
## Context
|
|
6
|
+
|
|
7
|
+
Caller-written question wording and option sets can drift even when the intended judgment is the same. A typed declaration can enforce a consistent shape without establishing the truth of its answer.
|
|
8
|
+
|
|
9
|
+
## Decision
|
|
10
|
+
|
|
11
|
+
Accept typed asks and compile verify, classify, rate, decide and locate intents into canonical questions and options. Accept an array, a single intent object or its JSON representation; reject invalid intent contracts with guidance. Fix substantive option order by template and add applicable other and cannot_tell outcomes.
|
|
12
|
+
|
|
13
|
+
Preserve the exact supplied claim text in verification results. Warn about negation or compound claims without silently rewriting or reversing their polarity. jev_ask verification combines issue outcomes with an exact-statement boolean cross-check.
|
|
14
|
+
|
|
15
|
+
Retain free for custom bool, choice or score questions, visibly marked uncalibrated. The compiler supplies required structural safeguards but does not turn arbitrary wording into a calibrated preset.
|
|
16
|
+
|
|
17
|
+
## Consequences
|
|
18
|
+
|
|
19
|
+
Consumers get consistent result structure and explicit escape-hatch limits. Canonical compilation is a form and policy guarantee, not proof of better accuracy. Typed asks remain distinct from the fixed review questions described in [Evidence construction before judgment](0005-evidence-construction-before-judgment.md).
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# 0005 — Evidence construction before judgment
|
|
2
|
+
|
|
3
|
+
Status: Accepted
|
|
4
|
+
|
|
5
|
+
## Context
|
|
6
|
+
|
|
7
|
+
A confident answer can still be wrong when a required dependency or the object named by a question is absent. Waiting for uncertainty before collecting context cannot repair that failure reliably.
|
|
8
|
+
|
|
9
|
+
## Decision
|
|
10
|
+
|
|
11
|
+
Prepare bounded discriminating evidence before judgment. Include the failing test and the code it exercises when distinguishing a bug from a wrong test; recover a named test from command output when identifiable. Add statically resolved declarations through a depth-one import closure under its own budget and report what was added or omitted. An absent or empty named part prevents its question; a uniquely resolvable orphan reference can be added before asking.
|
|
12
|
+
|
|
13
|
+
For diff checks, construct before-and-after evidence units from declarations, slices, files or hunks. Use fixed outcome or pointer choices when one answer is expected and boolean matrices when multiple units can be relevant. Keep test evidence separate from risk units. Keep policy thresholds in code rather than caller-written questions.
|
|
14
|
+
|
|
15
|
+
For eligible replaced member accesses, prepare isolated local-caller judgments alongside the risk matrix, using statically resolved callers and providers and including coordinated migrations. Apply the same evidence construction policy to supported Python syntax, imports and test fixtures, with missing grammar capabilities visible.
|
|
16
|
+
|
|
17
|
+
## Consequences
|
|
18
|
+
|
|
19
|
+
Static source evidence is never reported as execution. Import closure, local-caller collection and an empty finding list cannot prove completeness of dynamic dependencies or repository-wide safety. Required evidence that cannot fit together remains explicitly unjudged; [bounded evidence](0007-bounded-evidence-visible-limits.md) takes precedence over a plausible verdict.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
# 0006 — Visible uncertainty and constrained controls
|
|
2
|
+
|
|
3
|
+
Status: Accepted
|
|
4
|
+
|
|
5
|
+
## Context
|
|
6
|
+
|
|
7
|
+
Probability alone does not establish correctness, and a control can reveal missing evidence or sensitivity without identifying the right answer. Hiding those outcomes would encourage stronger conclusions than the evidence supports.
|
|
8
|
+
|
|
9
|
+
## Decision
|
|
10
|
+
|
|
11
|
+
Use visible verdict, unsure and abstain outcomes. Apply policy bands to leading-option probabilities, not the response's separate confidence field. Preserve missing-evidence outcomes before ordinary verdict bands and identify the piece to add.
|
|
12
|
+
|
|
13
|
+
Reverse substantive choice order only for applicable ambiguous choices, keeping canonical template order otherwise. Use decoy and known-reference witnesses on supported preset matrices, not on every arbitrary caller ask. Keep controls specialized: rival hypotheses belong to decide; attribution removes a selected pointer's evidence when that pointer drives an action.
|
|
14
|
+
|
|
15
|
+
In jev_ask verify, compare the issue outcome and exact-claim boolean in the same request group. Substantive disagreement, or an unknown issue outcome opposed by a strong boolean yes, becomes unsure with both values and the exact claim. A failed order check, witness or cross-check never promotes a verdict.
|
|
16
|
+
|
|
17
|
+
Share result marks, visible limits and calls/cost/cache/time reporting across tools. Custom questions remain marked uncalibrated.
|
|
18
|
+
|
|
19
|
+
## Consequences
|
|
20
|
+
|
|
21
|
+
Controls detect a limited class of evidence or setup problems; passing them does not prove correctness. An exact boolean cannot override an issue outcome into a contradiction. Callers must distinguish unconfirmed conclusions from their own subsequent verification.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
# 0007 — Bounded evidence without silent truncation
|
|
2
|
+
|
|
3
|
+
Status: Accepted
|
|
4
|
+
|
|
5
|
+
## Context
|
|
6
|
+
|
|
7
|
+
Evidence collection, transport and display each require bounds. Silently trimming required evidence or splitting a control away from its judgment would change the question while appearing to answer it.
|
|
8
|
+
|
|
9
|
+
## Decision
|
|
10
|
+
|
|
11
|
+
Keep character admission limits separate from heuristic token estimates: admission does not guarantee provider acceptance. Subdivide requests only between indivisible judgment groups, keeping exact-claim cross-checks and contrastive controls together and repeating applicable witnesses. Refuse a group that cannot fit, or required pieces that cannot be admitted together, with explicit unjudged work; do not retry an identical rejected request indefinitely.
|
|
12
|
+
|
|
13
|
+
Apply shared throughput and concurrency limits, bounded transport retries, per-tool call caps and session call/cost budgets. Report exhausted budgets and work not judged.
|
|
14
|
+
|
|
15
|
+
Capture optional command output privately with bounded execution time and stream bookkeeping. Compress repetitive line shapes by rarity while preserving failure evidence; if it still cannot fit, use bounded passage selection before the caller's judgment. Report overlong lines, shape limits and oversized streams explicitly. The stream-size refusal occurs after execution and is not a disk cap or command sandbox. Head/tail context is not a substitute for rarity compression.
|
|
16
|
+
|
|
17
|
+
Keep collection omissions and display limits visible. Distinguish a result rendered concisely from evidence that was unavailable to judgment.
|
|
18
|
+
|
|
19
|
+
## Consequences
|
|
20
|
+
|
|
21
|
+
A refusal is preferable to an answer about silently altered evidence. Budgeting cannot promise zero provider refusals without a provider token-counting contract. Command permissions remain those of the host, and JEV_TOOLS_ALLOW_COMMAND=0 disables the optional command path.
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# 0008 — Static test discovery and conservative execution plans
|
|
2
|
+
|
|
3
|
+
Status: Accepted
|
|
4
|
+
|
|
5
|
+
## Context
|
|
6
|
+
|
|
7
|
+
Finding tests related to a changed unit is not the same as determining how a runner executes them. Evaluating third-party configuration during discovery would introduce unrequested execution.
|
|
8
|
+
|
|
9
|
+
## Decision
|
|
10
|
+
|
|
11
|
+
Discover existing tests from tracked sources, literal configuration, scripts and static imports without evaluating configuration or automatically collecting tests. Preserve runner, working directory, configuration, project and framework boundaries in separate command plans, including project options. Keep runtime and type-test evidence and commands separate.
|
|
12
|
+
|
|
13
|
+
Select touched tests directly, follow statically resolved import closure through unchanged intermediates and applicable pytest fixtures, and judge remaining scenarios against changed-unit evidence. Retain tests when judgment is missing or unavailable. Use the selection threshold conservatively; when at least 80% of a file is selected, names are uncertain or counts are unknown, run the whole file.
|
|
14
|
+
|
|
15
|
+
For unsupported, dynamic or unresolved runners, keep the plan unsure and identify the next manual action rather than inventing executable arguments. Residual exported-unit coverage checks apply only within the discovered inventory.
|
|
16
|
+
|
|
17
|
+
## Consequences
|
|
18
|
+
|
|
19
|
+
The tool returns plans and never runs them. The caller must execute the commands and inspect results. Static discovery cannot establish all repository tests or a requirement to add a new scenario; these scope limits follow [Explicit scope and conservative automation](0003-explicit-scope-conservative-automation.md).
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
# 0009 — Session-local caching and requested-model identity
|
|
2
|
+
|
|
3
|
+
Status: Accepted
|
|
4
|
+
|
|
5
|
+
## Context
|
|
6
|
+
|
|
7
|
+
Repeated identical judgments can reuse a result within a session, but optional commands can observe or mutate changing state. A model name in a response need not identify the implementation that actually served the request.
|
|
8
|
+
|
|
9
|
+
## Decision
|
|
10
|
+
|
|
11
|
+
Cache only successful non-command judgments in session memory. Include the requested model string in the cache identity along with the judgment inputs; do not persist the cache across sessions or cache command-bearing requests.
|
|
12
|
+
|
|
13
|
+
Treat JEV_TOOLS_MODEL as the requested model identity. The default openjev is a moving alias. A version-shaped request string or an echoed model field does not prove that the served model is pinned. Do not claim runtime model-drift detection from that string.
|
|
14
|
+
|
|
15
|
+
## Consequences
|
|
16
|
+
|
|
17
|
+
Session caching avoids repeated identical work without presenting command output as immutable. Changing the requested model separates cached judgments, but cannot establish the identity or behavior of the model actually served. Cache accounting remains visible in result output.
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
# 0010 - MCP server as a thin host over the same tools
|
|
2
|
+
|
|
3
|
+
Status: Accepted
|
|
4
|
+
|
|
5
|
+
## Context
|
|
6
|
+
|
|
7
|
+
The tools were usable only inside pi and omp. Other agents (Claude Code, Claude Desktop, Kiro, Cursor, VS Code, Codex) integrate external tools through the Model Context Protocol. A separate implementation per client would duplicate evidence collection, judgment and budgets, and drift from the pi and omp behavior.
|
|
8
|
+
|
|
9
|
+
## Decision
|
|
10
|
+
|
|
11
|
+
Add `src/mcp/` as one more host, alongside the pi/omp entry point in `src/index.ts`:
|
|
12
|
+
|
|
13
|
+
- `protocol.ts` implements JSON-RPC 2.0 dispatch for the tools capability only, with no I/O.
|
|
14
|
+
- `tools.ts` calls the six existing tool factories unchanged and adapts their results, with the same `ToolDependencies`, session limits, HTTP client and configuration precedence ([ADR 0002](0002-one-http-protocol-across-hosts.md)).
|
|
15
|
+
- `main.ts` is the stdio transport and the `jev-agent-tools-mcp` binary.
|
|
16
|
+
|
|
17
|
+
Write the protocol by hand instead of depending on an MCP SDK, so the package keeps its single runtime dependency. Ship the binary as JavaScript compiled from `src/mcp/main.ts` into `dist/`, because Node refuses to strip TypeScript types under `node_modules`. pi and omp keep loading `src/` directly.
|
|
18
|
+
|
|
19
|
+
Host-specific behavior stays out of the tools: MCP uses a generic host (`mcpHost()`), a host-neutral process runner and annotations that mark `jev_ask` as not read-only while commands are enabled. The run-end documentation check stays a pi/omp hook; MCP users call `jev_check_diff` with `check: "docs"`.
|
|
20
|
+
|
|
21
|
+
## Consequences
|
|
22
|
+
|
|
23
|
+
Every tool change reaches all hosts at once, and MCP tests exercise the real factories. The server must follow MCP protocol revisions itself; it supports the initialize-based versions and `server/discover`, and offers no resources or prompts. Clients decide approval and may ignore server `instructions`, so projects add [agent instructions](../agent-instructions.md) to their own instruction files. The build step is required before packing, enforced by `prepack`.
|
|
@@ -0,0 +1,120 @@
|
|
|
1
|
+
# Agent instructions for MCP clients
|
|
2
|
+
|
|
3
|
+
pi and OMP deliver the reading guide and discretionary Jev policy automatically. The MCP server sends the same semantic policy as its `instructions`, with MCP's explicit absence of an automatic docs hook, but clients may ignore that field. Add the generated block below to the project's instruction file so the agent can choose useful evidence tools and report their results accurately.
|
|
4
|
+
|
|
5
|
+
Copy the block unchanged, or trim the optional tool table to the tools you enable. It contains no secrets and is safe to commit. Copies are not updated automatically: replace them from this page when upgrading a release.
|
|
6
|
+
|
|
7
|
+
## The block
|
|
8
|
+
|
|
9
|
+
<!-- BEGIN GENERATED JEV INSTRUCTIONS -->
|
|
10
|
+
|
|
11
|
+
```markdown
|
|
12
|
+
## Jev evidence tools (jev_* via MCP)
|
|
13
|
+
|
|
14
|
+
<!-- Generated policy 2026-10-03.1 from src/texts/instructions.ts. Update this copied block from docs/agent-instructions.md on each release; pasted copies are not updated automatically. -->
|
|
15
|
+
|
|
16
|
+
jev_* tools: choosing evidence and reading results.
|
|
17
|
+
|
|
18
|
+
Use Jev for a bounded semantic judgment when its answer could change an open decision or focus the next inspection, and the relevant evidence is available. Use decisive reading, search or authorized execution directly when it settles the question. Exact lookup, counting and runtime causality belong to native tools. A Jev call is not a prerequisite for a conclusion, a review or task completion; no explanation is needed for choosing native tools.
|
|
19
|
+
|
|
20
|
+
Before a chosen call, identify the open decision and supply only the evidence needed for a bounded question; a complete prior analysis is not required. Ask about one positive, self-contained observable fact per claim, with its sources, and show both sides of a comparison.
|
|
21
|
+
- Failure: relevant output, the failing test and the implementation it exercises; output alone is a lead, not a bug-versus-wrong-test diagnosis. Include configuration or runtime evidence when the hypothesis depends on it. A static judgment does not establish runtime causality.
|
|
22
|
+
- Change: current evidence and the earlier reference via base. Use the checkout/root corresponding to the work and admitted by the tool; another clone is not the same context.
|
|
23
|
+
- Plan/documentation: precise observable commitments and the passages that constrain them. Local confirmation does not establish that omitted obligations were searched.
|
|
24
|
+
- Selection/review: compatible inventory, references and configuration. Suggested commands are not executed commands; existing tests need not cover a new scenario. Read the diff natively for its contents, not as proof of safety or global coverage.
|
|
25
|
+
|
|
26
|
+
The report begins with execution state: complete means the requested admitted work was processed; partial means some requested work remains unjudged or limited; not_judged means no admissible judgment; refused means a call-wide refusal. These states describe execution, not correctness, safety, documentation completeness or global test coverage. Static selection or conservative fallback can be useful without a Jev judgment; inspect the selection reason and remaining limits.
|
|
27
|
+
|
|
28
|
+
Context identifies canonical host/server authority, requested and effective root, base and resolved reference, inventory restrictions and command execution/cwd when collected. Compare that context to your actual work before using a result. Unknown, not_collected and not_applicable are distinct, not guessed values or a substitute checkout.
|
|
29
|
+
|
|
30
|
+
Each item distinguishes a judged result sourced from fresh or cache, static treatment without a Jev judgment, or unjudged work. A cached judgment is not a fresh verification. Only judged items carry a judgment measure; unjudged and static items have no invented probability or confidence. Source and treatment are independent of the probability band.
|
|
31
|
+
|
|
32
|
+
Jev reads what you pass at a glance. A verdict is a lead to check before an irreversible action, not a proof. About 1 in 100 clear verdicts is wrong, and far more when the evidence is cut, from the wrong file, or depends on a file that was not passed; no mark can see that, so check the evidence yourself before you edit, delete or report done.
|
|
33
|
+
|
|
34
|
+
Marks:
|
|
35
|
+
- unsure: Jev does not see it clearly (yes/no between 0.20 and 0.80; a category or level under 0.85 confidence; findings between their thresholds), or a control of the call failed and the line says which. Inspect the passage or obtain the decisive file; unchanged rewording does not settle it.
|
|
36
|
+
- abstain: a necessary piece is missing from what you passed; the line names it. Obtain that piece or leave the conclusion open. A further Jev call is optional when the changed evidence makes it useful.
|
|
37
|
+
- "no (not shown)" or "not addressed": the file or state does not show it. That is not "false".
|
|
38
|
+
- uncalibrated: no error rate has been measured for this kind of ask; read it as a hint.
|
|
39
|
+
- Diagnostics: typed cause, origin, affected scope and materiality explain limits or blocking work, with referenced next actions. An uncertain judged item and work never judged are different. Inspect the named evidence or action without treating a diagnostic as a probability or a clean bill of health.
|
|
40
|
+
|
|
41
|
+
Accounting separates tool invocation, HTTP attempts, questions actually sent, requested results (fresh/cache/not judged/static), cache probes (hits/requests), auxiliary controls and passage work, current reported cost and elapsed time. These counts are not interchangeable: an HTTP attempt is not a result, a cache probe is not a requested judgment, and a cached result does not imply a request. Unknown current cost means unreported, not zero or reconstructed historical cached cost. Use max_calls to bound a chosen wide ask.
|
|
42
|
+
|
|
43
|
+
After an unavailable, refused or out-of-scope result, continue natively within the existing permissions; do not widen sharing or confinement to obtain a judgment. For uncertainty or missing evidence, inspect or obtain the decisive piece, or leave the conclusion open. Revisit Jev only when new evidence, a material context change or a new useful question makes the judgment useful; rewording unchanged evidence is not a reason to retry. There is no retry quota for genuinely changed evidence, and no certification is needed once decisive evidence settles the question. A reached cap or an unjudged result is not evidence of safety or zero affected tests.
|
|
44
|
+
|
|
45
|
+
Report current conclusions first, then decisive evidence, origin and scope, then material reservations. Established means supported by relevant decisive evidence; a reported check remains explicitly reported, not observed execution. Distinguish native reading/execution, Jev's static judgment and testimony. A call count or global status is not proof.
|
|
46
|
+
If native evidence settles the same scope and context after a Jev uncertainty, attribute the current conclusion to that evidence; old unsure or abstain results need not be recited as current reservations. A later Jev confirmation is attributed to Jev and retains its static limits. Independent reservations survive either resolution.
|
|
47
|
+
Where an uncertainty actually returned by Jev remains material, explicitly say “Jev did not confirm X”, identify the missing evidence or guarantee, its known or indeterminate impact and the evidence to obtain. Preserve that reservation in the main text, summary and recommendation. Name native reservations and unjudged work as such, not as Jev uncertainty. Unresolved contradiction, stale evidence or a different checkout/base/scenario remains a reservation; the latest favorable answer does not win by default.
|
|
48
|
+
Keep material limits in the main text, not only behind a link. Do not generalize a checked scenario into a universal guarantee. Separate relevant unestablished hypotheses from observations; omit irrelevant hypotheses. Explain unavailable historical causes only when material to action or requested, without inventing causality.
|
|
49
|
+
Use existing history references when useful or requested; no exhaustive historical relay, persistent register or History block is required. A requested audit details available steps separately from the current state and does not reopen settled conclusions. Preserve existing traces; identify unavailable traces when they limit evidence or audit, without reconstructing evidence, executions or call counts. pi/OMP may reference an existing trace; MCP references must be client-accessible or explicitly unavailable.
|
|
50
|
+
|
|
51
|
+
Ask about facts the files show, in positive sentences, with the evidence attached; do counting and searching with your text search tool or code.
|
|
52
|
+
|
|
53
|
+
Evidence passed to jev_* tools leaves the machine for the configured endpoint. Review its data handling; share only authorized evidence, never secrets or credentials. Host/client approvals still apply. jev_ask commands run with normal shell permissions and no sandbox; prefer read-only commands. Tool choice does not relax quality, proof, approval or confidentiality obligations.
|
|
54
|
+
|
|
55
|
+
jev_check_diff: Review a stable diff for semantic risks, stale documentation or specification drift when that review can inform an open decision. Read the diff natively when you need its contents. Choose risk, docs or spec for the question at hand; neither risk followed by docs nor a Jev review before done is required. Findings are leads within the inspected scope, not proof of global safety or completeness.
|
|
56
|
+
|
|
57
|
+
MCP has no automatic run-end documentation hook. Its absence does not require manual replacement calls, including risk followed by docs. pi and OMP retain an opt-out host hook; that is an explicit automatic exception, not an agent call requirement.
|
|
58
|
+
|
|
59
|
+
### Optional tool choices
|
|
60
|
+
|
|
61
|
+
| Tool | Useful open decision |
|
|
62
|
+
|---|---|
|
|
63
|
+
| jev_ask | One bounded judgment combines a note, files, before/after evidence or command output. |
|
|
64
|
+
| jev_ask_files | The same questions apply independently to many candidate files. |
|
|
65
|
+
| jev_find_files | Find an entry point by behavior when the filename is unknown. |
|
|
66
|
+
| jev_locate_in_file | Find a useful range in one file over 19 KB. |
|
|
67
|
+
| jev_check_diff | Review a stable diff for the chosen risk, docs or spec question. |
|
|
68
|
+
| jev_select_tests | Suggest commands for existing affected tests when selection can change the execution plan; it never runs them. |
|
|
69
|
+
|
|
70
|
+
Use your text search tool for exact strings and known symbols, your file-name search tool for known filenames, and native reading or authorized execution for decisive evidence. Tool descriptions retain their full evidence recipes and limitations.
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
<!-- END GENERATED JEV INSTRUCTIONS -->
|
|
74
|
+
|
|
75
|
+
## Where to put it
|
|
76
|
+
|
|
77
|
+
| Client | File | Notes |
|
|
78
|
+
|---|---|---|
|
|
79
|
+
| Claude Code | `CLAUDE.md` at the project root | See [CLAUDE.md](#claudemd). `CLAUDE.local.md` for a personal, uncommitted copy. |
|
|
80
|
+
| Codex CLI, and agents that read `AGENTS.md` | `AGENTS.md` at the project root | See [AGENTS.md](#agentsmd). |
|
|
81
|
+
| Kiro | `.kiro/steering/jev.md` | See [Kiro steering](#kiro-steering). |
|
|
82
|
+
| Cursor, VS Code, Windsurf, Claude Desktop | The client's project rules or custom instructions | Paste the block; Claude Desktop has no project file, so add it to the project's instructions in the app. |
|
|
83
|
+
|
|
84
|
+
### CLAUDE.md
|
|
85
|
+
|
|
86
|
+
Either paste the block into `CLAUDE.md`, or keep it in its own file and import it. Claude Code expands `@path` references in `CLAUDE.md` when it starts:
|
|
87
|
+
|
|
88
|
+
```markdown
|
|
89
|
+
# Project instructions
|
|
90
|
+
|
|
91
|
+
@docs/jev-agent-instructions.md
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
Save the block as `docs/jev-agent-instructions.md` for that form.
|
|
95
|
+
|
|
96
|
+
### AGENTS.md
|
|
97
|
+
|
|
98
|
+
Paste the block under its own heading. Codex reads `AGENTS.md`; Claude Code reads it too.
|
|
99
|
+
|
|
100
|
+
### Kiro steering
|
|
101
|
+
|
|
102
|
+
Save as `.kiro/steering/jev.md` with front matter that includes it in every session:
|
|
103
|
+
|
|
104
|
+
```markdown
|
|
105
|
+
---
|
|
106
|
+
inclusion: always
|
|
107
|
+
---
|
|
108
|
+
|
|
109
|
+
<the block>
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
## Keeping it current
|
|
113
|
+
|
|
114
|
+
The versioned canonical fragments and host renderers live in [`src/texts/instructions.ts`](../src/texts/instructions.ts). Runtime guides, descriptions/guidelines, the OMP rule and this MCP block reuse them; the MCP block includes policy version, source and copy-update instructions. Thresholds continue to come from [policy constants](design.md#policy-thresholds).
|
|
115
|
+
|
|
116
|
+
After changing the canonical policy, run `node scripts/generate-instructions.ts` to regenerate this block and `rules/jev-ask.md`. Run `node scripts/generate-instructions.ts --check` to reject stale generated outputs. During release verification, inspect delivery across pi/OMP/MCP for semantic parity, including the automatic pi/OMP hook versus no MCP hook; shared wording is not evidence of agent behavior. A release updates runtime/server instructions, not instruction blocks already pasted into external projects. Update those copies from this page; no synchronization with external copies is claimed.
|
|
117
|
+
|
|
118
|
+
## Automatic hook and evaluation costs
|
|
119
|
+
|
|
120
|
+
pi and OMP preserve the existing automatic run-end docs hook and its conditions/budget; `JEV_TOOLS_AUTO_DOCS=0` disables it. This is an explicit exception to agent-discretionary calls, not a requirement to invoke Jev, and does not certify complete documentation. MCP has no hook and no required manual replacement. In usage evaluation, record hook cost separately from discretionary calls and include both in full-task cost; no product price, budget setting or new runtime counter is introduced by this instruction policy.
|