jev-agent-tools 0.1.3 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (135) hide show
  1. package/CHANGELOG.md +88 -3
  2. package/CONTRIBUTING.md +40 -0
  3. package/README.md +64 -9
  4. package/SECURITY.md +27 -0
  5. package/dist/adapters/analysis-context.js +75 -0
  6. package/dist/adapters/ask-files.js +189 -0
  7. package/dist/adapters/ask-proof.js +144 -0
  8. package/dist/adapters/ask-syntax.js +385 -0
  9. package/dist/adapters/canonical-path.js +17 -0
  10. package/dist/adapters/command.js +181 -0
  11. package/dist/adapters/docs.js +172 -0
  12. package/dist/adapters/exec.js +207 -0
  13. package/dist/adapters/files.js +293 -0
  14. package/dist/adapters/find.js +122 -0
  15. package/dist/adapters/git-base.js +26 -0
  16. package/dist/adapters/git-inventory.js +71 -0
  17. package/dist/adapters/git.js +439 -0
  18. package/dist/adapters/locate-file.js +159 -0
  19. package/dist/adapters/output-lines.js +46 -0
  20. package/dist/adapters/private-storage.js +98 -0
  21. package/dist/adapters/risk-callers.js +426 -0
  22. package/dist/adapters/runner-version.js +78 -0
  23. package/dist/adapters/shell.js +76 -0
  24. package/dist/adapters/syntax.js +187 -0
  25. package/dist/adapters/test-inventory.js +131 -0
  26. package/dist/adapters/usage.js +20 -0
  27. package/dist/adapters/utf8.js +47 -0
  28. package/dist/configuration.js +257 -0
  29. package/dist/constants.js +119 -0
  30. package/dist/core/ask-closure.js +282 -0
  31. package/dist/core/ask-proof.js +1 -0
  32. package/dist/core/ask-references.js +194 -0
  33. package/dist/core/asks.js +436 -0
  34. package/dist/core/batches.js +65 -0
  35. package/dist/core/command-output.js +224 -0
  36. package/dist/core/diff.js +178 -0
  37. package/dist/core/docs.js +302 -0
  38. package/dist/core/find.js +108 -0
  39. package/dist/core/git.js +1 -0
  40. package/dist/core/imports.js +550 -0
  41. package/dist/core/integrity.js +45 -0
  42. package/dist/core/lexical.js +132 -0
  43. package/dist/core/locate.js +169 -0
  44. package/dist/core/output.js +120 -0
  45. package/dist/core/pointer.js +29 -0
  46. package/dist/core/risk-callers.js +851 -0
  47. package/dist/core/runner-version.js +45 -0
  48. package/dist/core/sections.js +230 -0
  49. package/dist/core/state.js +44 -0
  50. package/dist/core/syntax.js +1 -0
  51. package/dist/core/test-commands.js +334 -0
  52. package/dist/core/test-coverage.js +74 -0
  53. package/dist/core/test-discovery.js +1382 -0
  54. package/dist/core/test-evidence.js +527 -0
  55. package/dist/core/test-state.js +81 -0
  56. package/dist/core/truncate.js +12 -0
  57. package/dist/core/units.js +349 -0
  58. package/dist/describe.js +23 -0
  59. package/dist/guide.js +33 -0
  60. package/dist/host.js +24 -0
  61. package/dist/jev/client.js +434 -0
  62. package/dist/jev/pool.js +54 -0
  63. package/dist/jev/types.js +1 -0
  64. package/dist/mcp/main.js +124 -0
  65. package/dist/mcp/protocol.js +187 -0
  66. package/dist/mcp/tools.js +116 -0
  67. package/dist/presets/docs.js +62 -0
  68. package/dist/presets/risk.js +179 -0
  69. package/dist/presets/spec.js +81 -0
  70. package/dist/presets/witnesses.js +249 -0
  71. package/dist/render.js +42 -0
  72. package/dist/result.js +3 -0
  73. package/dist/runtime.js +1 -0
  74. package/dist/session.js +147 -0
  75. package/dist/texts/ask-files.js +1 -0
  76. package/dist/texts/ask.js +2 -0
  77. package/dist/texts/check-diff.js +17 -0
  78. package/dist/texts/configuration.js +1 -0
  79. package/dist/texts/find.js +14 -0
  80. package/dist/texts/guide.js +16 -0
  81. package/dist/texts/locate.js +10 -0
  82. package/dist/texts/select-tests.js +2 -0
  83. package/dist/tools/ask-files.js +217 -0
  84. package/dist/tools/ask-schema.js +70 -0
  85. package/dist/tools/ask.js +686 -0
  86. package/dist/tools/check-diff.js +402 -0
  87. package/dist/tools/docs-check.js +299 -0
  88. package/dist/tools/find.js +389 -0
  89. package/dist/tools/locate.js +303 -0
  90. package/dist/tools/select-tests.js +567 -0
  91. package/dist/tools/spec-check.js +166 -0
  92. package/docs/adr/0001-strict-typescript-pure-core-offline-tests.md +31 -0
  93. package/docs/adr/0002-one-http-protocol-across-hosts.md +17 -0
  94. package/docs/adr/0003-explicit-scope-conservative-automation.md +19 -0
  95. package/docs/adr/0004-compiled-typed-intents.md +19 -0
  96. package/docs/adr/0005-evidence-construction-before-judgment.md +19 -0
  97. package/docs/adr/0006-visible-uncertainty-constrained-controls.md +21 -0
  98. package/docs/adr/0007-bounded-evidence-visible-limits.md +21 -0
  99. package/docs/adr/0008-static-test-discovery-conservative-plans.md +19 -0
  100. package/docs/adr/0009-session-cache-requested-model-identity.md +17 -0
  101. package/docs/adr/0010-mcp-server-thin-host.md +23 -0
  102. package/docs/agent-instructions.md +91 -0
  103. package/docs/design.md +3 -3
  104. package/docs/mcp.md +231 -0
  105. package/package.json +19 -4
  106. package/server.json +57 -0
  107. package/src/adapters/canonical-path.ts +18 -0
  108. package/src/adapters/command.ts +7 -4
  109. package/src/adapters/exec.ts +226 -0
  110. package/src/adapters/private-storage.ts +143 -0
  111. package/src/adapters/risk-callers.ts +4 -2
  112. package/src/adapters/shell.ts +97 -0
  113. package/src/configuration.ts +294 -0
  114. package/src/constants.ts +11 -0
  115. package/src/core/command-output.ts +17 -1
  116. package/src/host-tui.d.ts +14 -0
  117. package/src/host.ts +11 -0
  118. package/src/index.ts +13 -5
  119. package/src/jev/client.ts +12 -0
  120. package/src/jev/types.ts +6 -0
  121. package/src/mcp/main.ts +135 -0
  122. package/src/mcp/protocol.ts +282 -0
  123. package/src/mcp/tools.ts +166 -0
  124. package/src/secret-input.ts +222 -0
  125. package/src/session.ts +59 -0
  126. package/src/setup.ts +170 -0
  127. package/src/texts/configuration.ts +1 -1
  128. package/src/tools/ask-files.ts +8 -13
  129. package/src/tools/ask.ts +29 -28
  130. package/src/tools/check-diff.ts +11 -11
  131. package/src/tools/docs-check.ts +1 -0
  132. package/src/tools/find.ts +8 -8
  133. package/src/tools/locate.ts +8 -9
  134. package/src/tools/select-tests.ts +10 -10
  135. package/src/tools/spec-check.ts +1 -0
@@ -0,0 +1,166 @@
1
+ import { collectFiles } from "../adapters/files.js";
2
+ import { collectUnits } from "../adapters/git.js";
3
+ import { resolveBase } from "../adapters/git-base.js";
4
+ import { CHOICE_MAX_OPTIONS, STATE_MAX_CHARS, TIMEOUT_MS, } from "../constants.js";
5
+ import { buildEnvelope, } from "../core/output.js";
6
+ import { prepareSpecCheck, readSpecJudgment, } from "../presets/spec.js";
7
+ import { NOT_CONFIGURED } from "../texts/configuration.js";
8
+ export async function runSpecCheck(deps, input) {
9
+ const started = performance.now();
10
+ const judgments = [];
11
+ const findings = [];
12
+ const limitations = [];
13
+ const unchecked = [];
14
+ let budget;
15
+ let sent = 0;
16
+ let incomplete = false;
17
+ let emptyBase;
18
+ const finish = (refusal) => {
19
+ const answers = findings.map((finding) => ({
20
+ label: finding.kind === "requirement"
21
+ ? `${input.specPath}:${finding.requirement?.start}-${finding.requirement?.end} ${finding.label} — requirement violated`
22
+ : `${finding.unit?.file}:${finding.unit?.afterRange?.start ?? finding.unit?.beforeRange?.start}-${finding.unit?.afterRange?.end ?? finding.unit?.beforeRange?.end} ${finding.label} — external behavior absent from the specification`,
23
+ value: {
24
+ head: finding.kind === "requirement" ? "violates" : finding.label,
25
+ p: finding.probability,
26
+ },
27
+ band: incomplete ? "unsure" : "verdict",
28
+ ...(incomplete
29
+ ? { reason: "incomplete change evidence: read the omitted pieces" }
30
+ : {}),
31
+ }));
32
+ const envelope = buildEnvelope({
33
+ answers,
34
+ limitations: emptyBase !== undefined && !refusal
35
+ ? [
36
+ ...limitations,
37
+ {
38
+ fact: `no changed units against ${emptyBase}`,
39
+ next: "nothing judged; pass base= or check the working directory",
40
+ },
41
+ ]
42
+ : limitations,
43
+ unchecked,
44
+ refusal,
45
+ budget,
46
+ ...(!refusal &&
47
+ !budget &&
48
+ !unchecked.length &&
49
+ !findings.length &&
50
+ emptyBase === undefined
51
+ ? {
52
+ lines: [
53
+ {
54
+ type: "list",
55
+ title: "spec: no violations or drift reported",
56
+ items: [],
57
+ },
58
+ ],
59
+ }
60
+ : {}),
61
+ yield: {
62
+ calls: judgments.reduce((n, j) => n + (j.calls ?? 0), 0),
63
+ questions: judgments.reduce((n, j) => n + (j.questions ?? 0), 0),
64
+ costUsd: judgments.reduce((n, j) => n + (j.usage?.costUsd ?? 0), 0),
65
+ cacheHits: judgments.reduce((n, j) => n + (j.cacheHits ?? 0), 0),
66
+ cacheRequests: judgments.reduce((n, j) => n + (j.cacheRequests ?? 0), 0),
67
+ elapsedMs: performance.now() - started,
68
+ },
69
+ });
70
+ return { ok: !refusal, envelope, judgments, findings };
71
+ };
72
+ if (!input.specPath)
73
+ return finish("spec_path required: no specification to check against.");
74
+ if (!deps.client)
75
+ return finish(NOT_CONFIGURED);
76
+ const comparison = await resolveBase(deps.exec, input.cwd, input.base, input.signal);
77
+ if (!comparison.ok)
78
+ return finish(comparison.error);
79
+ const base = comparison.base;
80
+ const root = await deps.exec("git", ["rev-parse", "--show-toplevel"], {
81
+ cwd: input.cwd,
82
+ timeout: TIMEOUT_MS,
83
+ signal: input.signal,
84
+ });
85
+ if (root.code || root.killed)
86
+ return finish("Repository root not found.");
87
+ const cwd = root.stdout.trim();
88
+ const [collected, specification] = await Promise.all([
89
+ collectUnits(deps.exec, { cwd, base, signal: input.signal }),
90
+ collectFiles(cwd, [input.specPath], input.signal, { exec: deps.exec }),
91
+ ]);
92
+ if (!collected.ok)
93
+ return finish(collected.error);
94
+ if (!specification.ok)
95
+ return finish(specification.error);
96
+ const text = Object.values(specification.files)[0];
97
+ if (!text)
98
+ return finish("The specification is empty: nothing to check.");
99
+ for (const limit of collected.limits)
100
+ limitations.push({
101
+ fact: `${limit.file} : ${limit.kind}`,
102
+ next: "Read the complete change before concluding.",
103
+ });
104
+ const units = collected.units.filter((unit) => {
105
+ if (unit.before !== null || unit.after !== null)
106
+ return true;
107
+ unchecked.push(`${unit.file} ${unit.name} (changed source unavailable)`);
108
+ return false;
109
+ });
110
+ const prepared = prepareSpecCheck(text, units);
111
+ if (!prepared.requirements.length)
112
+ return finish("No ### REQ-… requirements in the specification: nothing to check.");
113
+ if (!collected.units.length) {
114
+ emptyBase = base;
115
+ return finish();
116
+ }
117
+ if (unchecked.length) {
118
+ unchecked.push(...prepared.requirements.map((requirement) => `${requirement.label} (changed source unavailable)`), "drift (changed source unavailable)");
119
+ }
120
+ if (prepared.tableWarning)
121
+ limitations.push({
122
+ fact: "The specification contains a Markdown table.",
123
+ next: "Read the table requirements: their interpretation is uncalibrated.",
124
+ });
125
+ incomplete = collected.limits.length > 0;
126
+ if (units.length + 1 > CHOICE_MAX_OPTIONS ||
127
+ JSON.stringify(prepared.state).length > STATE_MAX_CHARS)
128
+ return finish(`State or spec pointer exceeds the limits; compare against a closer base (STATE_MAX_CHARS=${STATE_MAX_CHARS}).`);
129
+ if (!units.length)
130
+ return finish();
131
+ const judgment = await deps.client.judge(prepared.state, prepared.questions, {
132
+ signal: input.signal,
133
+ ...deps.runtime.session.requestGate(),
134
+ beforeRequest(questionCount) {
135
+ if (input.maxCalls !== undefined && sent >= input.maxCalls) {
136
+ budget = {
137
+ kind: "max_calls",
138
+ message: `max_calls=${input.maxCalls} reached`,
139
+ };
140
+ return { ok: false, error: budget.message };
141
+ }
142
+ const admitted = deps.runtime.session.admit(questionCount);
143
+ if (!admitted.ok)
144
+ budget = { kind: "session", message: admitted.error };
145
+ else
146
+ sent++;
147
+ return admitted;
148
+ },
149
+ onUsage: (usage) => deps.runtime.session.recordUsage(usage),
150
+ });
151
+ judgments.push(judgment);
152
+ if (!judgment.ok) {
153
+ unchecked.push(...prepared.requirements.map((req) => `${req.label} (${judgment.error})`), `drift (${judgment.error})`);
154
+ return finish();
155
+ }
156
+ for (const requirement of prepared.requirements) {
157
+ const answer = judgment.answers[requirement.id];
158
+ if (!answer || answer.type === "unjudged")
159
+ unchecked.push(requirement.label);
160
+ }
161
+ const drift = judgment.answers.drift;
162
+ if (!drift || drift.type === "unjudged")
163
+ unchecked.push("drift");
164
+ findings.push(...readSpecJudgment(prepared, judgment.answers));
165
+ return finish();
166
+ }
@@ -0,0 +1,31 @@
1
+ # 0001 — Strict TypeScript, a pure core and offline conformance tests
2
+
3
+ Status: Accepted
4
+
5
+ ## Context
6
+
7
+ The same extension runs inside pi and omp. Host-specific implementations would duplicate judgment policy and make deterministic checks harder to maintain. Optional native capabilities must not become mandatory runtime dependencies.
8
+
9
+ ## Decision
10
+
11
+ Use one strict TypeScript package with erasable syntax, Node APIs and fetch rather than Bun-only APIs. Pure core transformations and question compilation consume and return values; thin adapters own host, filesystem, Git, network and clock effects. Tools compose collection, evidence construction, requests and result rendering. Discriminated unions represent expected refusals, unjudged work and abstention as values rather than exceptional control flow.
12
+
13
+ Keep limits and thresholds in src/constants.ts. Batch questions for the same evidence state, run independent work concurrently under the shared limiter, and bound reads before loading content into memory.
14
+
15
+ Enforce downward import layers, including erased type dependencies:
16
+
17
+ - core/ imports core/, constants.ts, result.ts and type-only contracts from jev/types.ts.
18
+ - presets/ and adapters/ each import their own layer, core/ and those neutral modules; adapters/ does not import presets/ or tools/.
19
+ - jev/ imports its own layer, core/, constants.ts and result.ts.
20
+ - texts/ imports its own layer, constants.ts and presets/; it has no external dependencies.
21
+ - constants.ts and result.ts import nothing. Root integration modules and tools/ compose the layers.
22
+
23
+ Reject all cycles, including type-only cycles, and nonliteral module loading. Shared contracts belong below their consumers. Pure layers do not import external modules, with explicit core exceptions for pure node:path functions and erased import type contracts from @ast-grep/napi. Inline type specifiers that retain runtime loading do not qualify. Filesystem, process and network modules, direct fetch calls and Node builtin-module loaders are forbidden in the core; import checks are not an exhaustive proof against indirect global effects.
24
+
25
+ Keep @ast-grep/napi, @ast-grep/lang-python and @ff-labs/fff-node optional. Load native parsing and search at the edges, use the available fallback and report missing capabilities rather than claiming equivalent syntax precision.
26
+
27
+ Run deterministic offline conformance checks with tsc --noEmit, the import checker, Biome and node --test. Use constructed behavior and boundary cases and local HTTP fixtures for transport failures, retries and timeouts; no Jev credentials are required.
28
+
29
+ ## Consequences
30
+
31
+ This separates effects from policy and makes host-independent behavior reproducible. Missing optional acceleration can reduce parsing or search capability, so limits stay visible. Native optimization is not a reason to introduce a separate judgment path or an extension-specific daemon.
@@ -0,0 +1,17 @@
1
+ # 0002 — One HTTP protocol across hosts
2
+
3
+ Status: Accepted
4
+
5
+ ## Context
6
+
7
+ pi and omp provide different integration APIs. Delegating judgment to a host's chat model would change the protocol and behavior according to the host.
8
+
9
+ ## Decision
10
+
11
+ Use one HTTP client compatible with the Jev API format in both hosts. Read the user-supplied endpoint and Bearer credential from JEV_TOOLS_URL and JEV_TOOLS_API_KEY and the requested model from JEV_TOOLS_MODEL, defaulting to openjev.
12
+
13
+ Without the endpoint or key, keep tools registered and explain the missing configuration; disable the automatic documentation check. Never fall back to a chat model or a host-internal judgment service.
14
+
15
+ ## Consequences
16
+
17
+ Users choose the service and must review which evidence it receives. Host adapters manage registration and lifecycle, not a second judgment implementation. Configuration failures remain explicit rather than producing superficially equivalent answers.
@@ -0,0 +1,19 @@
1
+ # 0003 — Explicit scope and conservative automation
2
+
3
+ Status: Accepted
4
+
5
+ ## Context
6
+
7
+ Repository navigation, change review and test selection involve different evidence and next actions. Broad automatic suggestions could be mistaken for complete knowledge of the repository.
8
+
9
+ ## Decision
10
+
11
+ Expose six explicit tools: jev_ask for one combined situation, jev_ask_files for independent judgments of files, jev_find_files for discovery, jev_locate_in_file for ranges, jev_check_diff for fixed reviews and jev_select_tests for existing test plans. Hide jev_find_files in omp when its native find tool is active.
12
+
13
+ Allow only one automatic run-end documentation check of a dirty tree. A flagged stale sentence can request another turn; errors do not block the host or cause an automatic review loop. JEV_TOOLS_AUTO_DOCS=0 disables the hook.
14
+
15
+ Do not infer required new test scenarios, missing documentation or arbitrary companion edits from incomplete collection. Test selection and residual coverage checks describe only the discovered inventory.
16
+
17
+ ## Consequences
18
+
19
+ Callers retain control of consequential actions and must read or execute the evidence indicated by results. A clean review is not a completeness guarantee. Automatic documentation checking targets existing sentences, not an inferred obligation to write new documentation.
@@ -0,0 +1,19 @@
1
+ # 0004 — Compile typed intents instead of accepting raw question maps
2
+
3
+ Status: Accepted
4
+
5
+ ## Context
6
+
7
+ Caller-written question wording and option sets can drift even when the intended judgment is the same. A typed declaration can enforce a consistent shape without establishing the truth of its answer.
8
+
9
+ ## Decision
10
+
11
+ Accept typed asks and compile verify, classify, rate, decide and locate intents into canonical questions and options. Accept an array, a single intent object or its JSON representation; reject invalid intent contracts with guidance. Fix substantive option order by template and add applicable other and cannot_tell outcomes.
12
+
13
+ Preserve the exact supplied claim text in verification results. Warn about negation or compound claims without silently rewriting or reversing their polarity. jev_ask verification combines issue outcomes with an exact-statement boolean cross-check.
14
+
15
+ Retain free for custom bool, choice or score questions, visibly marked uncalibrated. The compiler supplies required structural safeguards but does not turn arbitrary wording into a calibrated preset.
16
+
17
+ ## Consequences
18
+
19
+ Consumers get consistent result structure and explicit escape-hatch limits. Canonical compilation is a form and policy guarantee, not proof of better accuracy. Typed asks remain distinct from the fixed review questions described in [Evidence construction before judgment](0005-evidence-construction-before-judgment.md).
@@ -0,0 +1,19 @@
1
+ # 0005 — Evidence construction before judgment
2
+
3
+ Status: Accepted
4
+
5
+ ## Context
6
+
7
+ A confident answer can still be wrong when a required dependency or the object named by a question is absent. Waiting for uncertainty before collecting context cannot repair that failure reliably.
8
+
9
+ ## Decision
10
+
11
+ Prepare bounded discriminating evidence before judgment. Include the failing test and the code it exercises when distinguishing a bug from a wrong test; recover a named test from command output when identifiable. Add statically resolved declarations through a depth-one import closure under its own budget and report what was added or omitted. An absent or empty named part prevents its question; a uniquely resolvable orphan reference can be added before asking.
12
+
13
+ For diff checks, construct before-and-after evidence units from declarations, slices, files or hunks. Use fixed outcome or pointer choices when one answer is expected and boolean matrices when multiple units can be relevant. Keep test evidence separate from risk units. Keep policy thresholds in code rather than caller-written questions.
14
+
15
+ For eligible replaced member accesses, prepare isolated local-caller judgments alongside the risk matrix, using statically resolved callers and providers and including coordinated migrations. Apply the same evidence construction policy to supported Python syntax, imports and test fixtures, with missing grammar capabilities visible.
16
+
17
+ ## Consequences
18
+
19
+ Static source evidence is never reported as execution. Import closure, local-caller collection and an empty finding list cannot prove completeness of dynamic dependencies or repository-wide safety. Required evidence that cannot fit together remains explicitly unjudged; [bounded evidence](0007-bounded-evidence-visible-limits.md) takes precedence over a plausible verdict.
@@ -0,0 +1,21 @@
1
+ # 0006 — Visible uncertainty and constrained controls
2
+
3
+ Status: Accepted
4
+
5
+ ## Context
6
+
7
+ Probability alone does not establish correctness, and a control can reveal missing evidence or sensitivity without identifying the right answer. Hiding those outcomes would encourage stronger conclusions than the evidence supports.
8
+
9
+ ## Decision
10
+
11
+ Use visible verdict, unsure and abstain outcomes. Apply policy bands to leading-option probabilities, not the response's separate confidence field. Preserve missing-evidence outcomes before ordinary verdict bands and identify the piece to add.
12
+
13
+ Reverse substantive choice order only for applicable ambiguous choices, keeping canonical template order otherwise. Use decoy and known-reference witnesses on supported preset matrices, not on every arbitrary caller ask. Keep controls specialized: rival hypotheses belong to decide; attribution removes a selected pointer's evidence when that pointer drives an action.
14
+
15
+ In jev_ask verify, compare the issue outcome and exact-claim boolean in the same request group. Substantive disagreement, or an unknown issue outcome opposed by a strong boolean yes, becomes unsure with both values and the exact claim. A failed order check, witness or cross-check never promotes a verdict.
16
+
17
+ Share result marks, visible limits and calls/cost/cache/time reporting across tools. Custom questions remain marked uncalibrated.
18
+
19
+ ## Consequences
20
+
21
+ Controls detect a limited class of evidence or setup problems; passing them does not prove correctness. An exact boolean cannot override an issue outcome into a contradiction. Callers must distinguish unconfirmed conclusions from their own subsequent verification.
@@ -0,0 +1,21 @@
1
+ # 0007 — Bounded evidence without silent truncation
2
+
3
+ Status: Accepted
4
+
5
+ ## Context
6
+
7
+ Evidence collection, transport and display each require bounds. Silently trimming required evidence or splitting a control away from its judgment would change the question while appearing to answer it.
8
+
9
+ ## Decision
10
+
11
+ Keep character admission limits separate from heuristic token estimates: admission does not guarantee provider acceptance. Subdivide requests only between indivisible judgment groups, keeping exact-claim cross-checks and contrastive controls together and repeating applicable witnesses. Refuse a group that cannot fit, or required pieces that cannot be admitted together, with explicit unjudged work; do not retry an identical rejected request indefinitely.
12
+
13
+ Apply shared throughput and concurrency limits, bounded transport retries, per-tool call caps and session call/cost budgets. Report exhausted budgets and work not judged.
14
+
15
+ Capture optional command output privately with bounded execution time and stream bookkeeping. Compress repetitive line shapes by rarity while preserving failure evidence; if it still cannot fit, use bounded passage selection before the caller's judgment. Report overlong lines, shape limits and oversized streams explicitly. The stream-size refusal occurs after execution and is not a disk cap or command sandbox. Head/tail context is not a substitute for rarity compression.
16
+
17
+ Keep collection omissions and display limits visible. Distinguish a result rendered concisely from evidence that was unavailable to judgment.
18
+
19
+ ## Consequences
20
+
21
+ A refusal is preferable to an answer about silently altered evidence. Budgeting cannot promise zero provider refusals without a provider token-counting contract. Command permissions remain those of the host, and JEV_TOOLS_ALLOW_COMMAND=0 disables the optional command path.
@@ -0,0 +1,19 @@
1
+ # 0008 — Static test discovery and conservative execution plans
2
+
3
+ Status: Accepted
4
+
5
+ ## Context
6
+
7
+ Finding tests related to a changed unit is not the same as determining how a runner executes them. Evaluating third-party configuration during discovery would introduce unrequested execution.
8
+
9
+ ## Decision
10
+
11
+ Discover existing tests from tracked sources, literal configuration, scripts and static imports without evaluating configuration or automatically collecting tests. Preserve runner, working directory, configuration, project and framework boundaries in separate command plans, including project options. Keep runtime and type-test evidence and commands separate.
12
+
13
+ Select touched tests directly, follow statically resolved import closure through unchanged intermediates and applicable pytest fixtures, and judge remaining scenarios against changed-unit evidence. Retain tests when judgment is missing or unavailable. Use the selection threshold conservatively; when at least 80% of a file is selected, names are uncertain or counts are unknown, run the whole file.
14
+
15
+ For unsupported, dynamic or unresolved runners, keep the plan unsure and identify the next manual action rather than inventing executable arguments. Residual exported-unit coverage checks apply only within the discovered inventory.
16
+
17
+ ## Consequences
18
+
19
+ The tool returns plans and never runs them. The caller must execute the commands and inspect results. Static discovery cannot establish all repository tests or a requirement to add a new scenario; these scope limits follow [Explicit scope and conservative automation](0003-explicit-scope-conservative-automation.md).
@@ -0,0 +1,17 @@
1
+ # 0009 — Session-local caching and requested-model identity
2
+
3
+ Status: Accepted
4
+
5
+ ## Context
6
+
7
+ Repeated identical judgments can reuse a result within a session, but optional commands can observe or mutate changing state. A model name in a response need not identify the implementation that actually served the request.
8
+
9
+ ## Decision
10
+
11
+ Cache only successful non-command judgments in session memory. Include the requested model string in the cache identity along with the judgment inputs; do not persist the cache across sessions or cache command-bearing requests.
12
+
13
+ Treat JEV_TOOLS_MODEL as the requested model identity. The default openjev is a moving alias. A version-shaped request string or an echoed model field does not prove that the served model is pinned. Do not claim runtime model-drift detection from that string.
14
+
15
+ ## Consequences
16
+
17
+ Session caching avoids repeated identical work without presenting command output as immutable. Changing the requested model separates cached judgments, but cannot establish the identity or behavior of the model actually served. Cache accounting remains visible in result output.
@@ -0,0 +1,23 @@
1
+ # 0010 - MCP server as a thin host over the same tools
2
+
3
+ Status: Accepted
4
+
5
+ ## Context
6
+
7
+ The tools were usable only inside pi and omp. Other agents (Claude Code, Claude Desktop, Kiro, Cursor, VS Code, Codex) integrate external tools through the Model Context Protocol. A separate implementation per client would duplicate evidence collection, judgment and budgets, and drift from the pi and omp behavior.
8
+
9
+ ## Decision
10
+
11
+ Add `src/mcp/` as one more host, alongside the pi/omp entry point in `src/index.ts`:
12
+
13
+ - `protocol.ts` implements JSON-RPC 2.0 dispatch for the tools capability only, with no I/O.
14
+ - `tools.ts` calls the six existing tool factories unchanged and adapts their results, with the same `ToolDependencies`, session limits, HTTP client and configuration precedence ([ADR 0002](0002-one-http-protocol-across-hosts.md)).
15
+ - `main.ts` is the stdio transport and the `jev-agent-tools-mcp` binary.
16
+
17
+ Write the protocol by hand instead of depending on an MCP SDK, so the package keeps its single runtime dependency. Ship the binary as JavaScript compiled from `src/mcp/main.ts` into `dist/`, because Node refuses to strip TypeScript types under `node_modules`. pi and omp keep loading `src/` directly.
18
+
19
+ Host-specific behavior stays out of the tools: MCP uses a generic host (`mcpHost()`), a host-neutral process runner and annotations that mark `jev_ask` as not read-only while commands are enabled. The run-end documentation check stays a pi/omp hook; MCP users call `jev_check_diff` with `check: "docs"`.
20
+
21
+ ## Consequences
22
+
23
+ Every tool change reaches all hosts at once, and MCP tests exercise the real factories. The server must follow MCP protocol revisions itself; it supports the initialize-based versions and `server/discover`, and offers no resources or prompts. Clients decide approval and may ignore server `instructions`, so projects add [agent instructions](../agent-instructions.md) to their own instruction files. The build step is required before packing, enforced by `prepack`.
@@ -0,0 +1,91 @@
1
+ # Agent instructions for MCP clients
2
+
3
+ pi and omp inject the jev reading guide and the `jev_ask` policy automatically. The MCP server sends the same text as its `instructions`, but clients may ignore that field. Add the block below to the project's instruction file so the agent knows when to use the tools and how to read their results.
4
+
5
+ Copy it unchanged, or trim the tool table to the tools you enable. It contains no secrets and is safe to commit.
6
+
7
+ ## The block
8
+
9
+ ```markdown
10
+ ## Jev evidence tools (jev_* via MCP)
11
+
12
+ The jev_* tools send repository evidence to a judgment endpoint and return typed answers with probabilities. They complement reading, searching and running code; they do not replace them.
13
+
14
+ ### When to use which
15
+
16
+ | Tool | Use it when |
17
+ |---|---|
18
+ | jev_ask | One judgment combines a note, files, an earlier version (base) or a command's output. |
19
+ | jev_ask_files | The same questions apply independently to each of many candidate files. |
20
+ | jev_find_files | You need the entry point for a behaviour and do not know the filename. |
21
+ | jev_locate_in_file | You need the relevant line range in one file of at least 19 KB. |
22
+ | jev_check_diff | Changes are complete and need risk, documentation or specification review. |
23
+ | jev_select_tests | You need the commands for existing tests affected by the changes (it does not run them). |
24
+
25
+ Use ordinary search and reading for exact strings, known symbols, filenames, line numbers, counts and arithmetic. Run commands yourself when you need their full output.
26
+
27
+ ### Required habits
28
+
29
+ - Before concluding that a failure is a bug in the code, a wrong test or an environment problem, or that a plan matches the documentation, pass the relevant files to jev_ask and weigh its answer against your own reading. Include both the failing test and the code it exercises.
30
+ - Ask about facts the files show, in positive sentences, with the evidence attached.
31
+ - Before reporting a change as done, call jev_check_diff with check "risk", then with check "docs" (MCP has no automatic documentation check). Update or justify every flagged sentence.
32
+ - Bound wide asks with max_calls.
33
+
34
+ ### Reading results
35
+
36
+ - A line with no mark is a verdict: a lead to check before editing, deleting or reporting done, not a proof.
37
+ - unsure: the answer is ambiguous or a control failed. Read the passage or file the line points to, or add the file that settles it. Rewording the question does not help.
38
+ - abstain: a necessary piece is missing; the line names it. Add that file or command and ask once.
39
+ - "no (not shown)" or "not addressed": the evidence does not show it. That is not "false".
40
+ - uncalibrated: no measured error rate; treat it as a hint.
41
+ - Lines in brackets: what limited the call and the next action or parameter to use.
42
+
43
+ ### Final answers
44
+
45
+ Identify every conclusion a jev_* tool marked unsure or abstain, say Jev did not confirm it, and name the evidence still needed. If you settled it by reading the decisive evidence yourself, say so and cite that evidence, separately from Jev's result.
46
+
47
+ ### Data
48
+
49
+ Evidence passed to jev_* tools leaves the machine. Do not pass secrets, credentials or files the user has not agreed to share. jev_ask commands run with normal shell permissions and no sandbox; prefer read-only commands.
50
+ ```
51
+
52
+ ## Where to put it
53
+
54
+ | Client | File | Notes |
55
+ |---|---|---|
56
+ | Claude Code | `CLAUDE.md` at the project root | See [CLAUDE.md](#claudemd). `CLAUDE.local.md` for a personal, uncommitted copy. |
57
+ | Codex CLI, and agents that read `AGENTS.md` | `AGENTS.md` at the project root | See [AGENTS.md](#agentsmd). |
58
+ | Kiro | `.kiro/steering/jev.md` | See [Kiro steering](#kiro-steering). |
59
+ | Cursor, VS Code, Windsurf, Claude Desktop | The client's project rules or custom instructions | Paste the block; Claude Desktop has no project file, so add it to the project's instructions in the app. |
60
+
61
+ ### CLAUDE.md
62
+
63
+ Either paste the block into `CLAUDE.md`, or keep it in its own file and import it. Claude Code expands `@path` references in `CLAUDE.md` when it starts:
64
+
65
+ ```markdown
66
+ # Project instructions
67
+
68
+ @docs/jev-agent-instructions.md
69
+ ```
70
+
71
+ Save the block as `docs/jev-agent-instructions.md` for that form.
72
+
73
+ ### AGENTS.md
74
+
75
+ Paste the block under its own heading. Codex reads `AGENTS.md`; Claude Code reads it too.
76
+
77
+ ### Kiro steering
78
+
79
+ Save as `.kiro/steering/jev.md` with front matter that includes it in every session:
80
+
81
+ ```markdown
82
+ ---
83
+ inclusion: always
84
+ ---
85
+
86
+ <the block>
87
+ ```
88
+
89
+ ## Keeping it current
90
+
91
+ The block restates the guide in [`src/texts/guide.ts`](../src/texts/guide.ts) and the policy in [`rules/jev-ask.md`](../rules/jev-ask.md). Thresholds are in [design](design.md#policy-thresholds). If a release changes them, the server's `instructions` change automatically; update the copied block from this page.
package/docs/design.md CHANGED
@@ -8,7 +8,7 @@ Supply the discriminating evidence, not an argument about it. For comparisons, s
8
8
 
9
9
  ## Pure core, thin adapters
10
10
 
11
- Pure core transformations construct states, units, questions and display envelopes. Adapters own filesystem, Git, parsing and host effects; the HTTP client owns transport. Presets define fixed review questions and texts define host-facing guidance. Dependency checks keep shared contracts below consumers, reject cycles and account for erased type imports. Pure path operations and erased parser types are explicit architectural exceptions. Both hosts use the same HTTP judgment protocol.
11
+ Pure core transformations construct states, units, questions and display envelopes. Adapters own filesystem, Git, parsing and host effects; the HTTP client owns transport. Presets define fixed review questions and texts define host-facing guidance. Hosts sit on top: `src/index.ts` registers the tools with pi and omp, and `src/mcp/` serves the same tool factories to any MCP client over stdio ([ADR 0010](adr/0010-mcp-server-thin-host.md), [setup](mcp.md)). Dependency checks keep shared contracts below consumers, reject cycles and account for erased type imports. Pure path operations and erased parser types are explicit architectural exceptions. Both hosts use the same HTTP judgment protocol.
12
12
 
13
13
  ## Typed intents and fixed checks
14
14
 
@@ -20,7 +20,7 @@ Leading-option probability (p_max) is the probability of the most likely choice,
20
20
 
21
21
  ## Bounded work, visible limits
22
22
 
23
- Admission limits, request planning, rate limiting and retry bounds are distinct. Oversized evidence is refused or visibly omitted, never silently converted into an ordinary verdict. Per-invocation max_calls and session call/cost limits are separate; control and severity requests count too. Unjudged tests stay selected. Automatic documentation review is the only run-end automation and can request at most one extra turn.
23
+ Admission limits, request planning, rate limiting and retry bounds are distinct. Oversized evidence is refused or visibly omitted, never silently converted into an ordinary verdict. Per-invocation max_calls and session call/cost limits are separate; control and severity requests count too. Under a session USD limit, requests are admitted one at a time so each admission sees all cost reported so far. Unjudged tests stay selected. Automatic documentation review is the only run-end automation and can request at most one extra turn.
24
24
 
25
25
  Successful non-command judgments are cached only in session memory by canonical evidence, question and requested model string. Errors are not cached; command judgments bypass caching. The default openjev alias can move, and an echoed model name does not prove served-model identity. Fractional budgets bound admission, not an exact prediction of the final request's cost.
26
26
 
@@ -184,4 +184,4 @@ Optional native syntax parsing and file-search acceleration can be absent. Tools
184
184
 
185
185
  ## Architecture decisions
186
186
 
187
- See the [architecture decision records](adr/) for durable trade-offs. Start with the [README](../README.md) for installation and follow its six tool references for complete parameter contracts.
187
+ See the [architecture decision records](adr/) for durable trade-offs. Start with the [README](../README.md) for installation and follow its six tool references for complete parameter contracts. MCP clients: [setup guide](mcp.md) and [agent instructions](agent-instructions.md).