jev-agent-tools 0.1.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (95) hide show
  1. package/CHANGELOG.md +66 -0
  2. package/LICENSE +21 -0
  3. package/README.md +119 -0
  4. package/docs/design.md +187 -0
  5. package/docs/tools/jev_ask.md +92 -0
  6. package/docs/tools/jev_ask_files.md +80 -0
  7. package/docs/tools/jev_check_diff.md +77 -0
  8. package/docs/tools/jev_find_files.md +59 -0
  9. package/docs/tools/jev_locate_in_file.md +51 -0
  10. package/docs/tools/jev_select_tests.md +59 -0
  11. package/package.json +85 -0
  12. package/rules/jev-ask.md +4 -0
  13. package/src/adapters/analysis-context.ts +100 -0
  14. package/src/adapters/ask-files.ts +236 -0
  15. package/src/adapters/ask-proof.ts +202 -0
  16. package/src/adapters/ask-syntax.ts +463 -0
  17. package/src/adapters/command.ts +240 -0
  18. package/src/adapters/docs.ts +222 -0
  19. package/src/adapters/files.ts +411 -0
  20. package/src/adapters/find.ts +151 -0
  21. package/src/adapters/git-base.ts +32 -0
  22. package/src/adapters/git-inventory.ts +94 -0
  23. package/src/adapters/git.ts +525 -0
  24. package/src/adapters/locate-file.ts +197 -0
  25. package/src/adapters/output-lines.ts +50 -0
  26. package/src/adapters/risk-callers.ts +525 -0
  27. package/src/adapters/runner-version.ts +102 -0
  28. package/src/adapters/syntax.ts +229 -0
  29. package/src/adapters/test-inventory.ts +168 -0
  30. package/src/adapters/usage.ts +21 -0
  31. package/src/adapters/utf8.ts +57 -0
  32. package/src/constants.ts +109 -0
  33. package/src/core/ask-closure.ts +419 -0
  34. package/src/core/ask-proof.ts +32 -0
  35. package/src/core/ask-references.ts +249 -0
  36. package/src/core/asks.ts +616 -0
  37. package/src/core/batches.ts +83 -0
  38. package/src/core/command-output.ts +249 -0
  39. package/src/core/diff.ts +226 -0
  40. package/src/core/docs.ts +399 -0
  41. package/src/core/find.ts +157 -0
  42. package/src/core/git.ts +5 -0
  43. package/src/core/import-boundaries.ts +102 -0
  44. package/src/core/imports.ts +691 -0
  45. package/src/core/integrity.ts +64 -0
  46. package/src/core/lexical.ts +130 -0
  47. package/src/core/locate.ts +213 -0
  48. package/src/core/output.ts +264 -0
  49. package/src/core/pointer.ts +51 -0
  50. package/src/core/risk-callers.ts +1270 -0
  51. package/src/core/runner-version.ts +66 -0
  52. package/src/core/sections.ts +269 -0
  53. package/src/core/state.ts +53 -0
  54. package/src/core/syntax.ts +8 -0
  55. package/src/core/test-commands.ts +430 -0
  56. package/src/core/test-coverage.ts +103 -0
  57. package/src/core/test-discovery.ts +1695 -0
  58. package/src/core/test-evidence.ts +649 -0
  59. package/src/core/test-state.ts +99 -0
  60. package/src/core/truncate.ts +14 -0
  61. package/src/core/units.ts +531 -0
  62. package/src/describe.ts +26 -0
  63. package/src/guide.ts +42 -0
  64. package/src/host.ts +22 -0
  65. package/src/index.ts +40 -0
  66. package/src/jev/client.ts +505 -0
  67. package/src/jev/pool.ts +60 -0
  68. package/src/jev/types.ts +60 -0
  69. package/src/presets/docs.ts +85 -0
  70. package/src/presets/risk.ts +263 -0
  71. package/src/presets/spec.ts +111 -0
  72. package/src/presets/witnesses.ts +313 -0
  73. package/src/render.ts +45 -0
  74. package/src/result.ts +4 -0
  75. package/src/run-end.ts +157 -0
  76. package/src/runtime.ts +16 -0
  77. package/src/session.ts +106 -0
  78. package/src/texts/ask-files.ts +2 -0
  79. package/src/texts/ask.ts +4 -0
  80. package/src/texts/check-diff.ts +24 -0
  81. package/src/texts/configuration.ts +2 -0
  82. package/src/texts/find.ts +14 -0
  83. package/src/texts/guide.ts +16 -0
  84. package/src/texts/locate.ts +10 -0
  85. package/src/texts/run-end.ts +13 -0
  86. package/src/texts/select-tests.ts +3 -0
  87. package/src/tools/ask-files.ts +263 -0
  88. package/src/tools/ask-schema.ts +116 -0
  89. package/src/tools/ask.ts +925 -0
  90. package/src/tools/check-diff.ts +510 -0
  91. package/src/tools/docs-check.ts +399 -0
  92. package/src/tools/find.ts +529 -0
  93. package/src/tools/locate.ts +369 -0
  94. package/src/tools/select-tests.ts +746 -0
  95. package/src/tools/spec-check.ts +210 -0
package/CHANGELOG.md ADDED
@@ -0,0 +1,66 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project will be documented in this file.
4
+
5
+ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
6
+ and this project follows [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
+ Before 1.0.0, incompatible changes will be called out explicitly. Internal source
8
+ modules are not a stable library API.
9
+
10
+ ## [Unreleased]
11
+
12
+ ## [0.1.3] - 2026-10-02
13
+
14
+ Initial npm release candidate, under the name `jev-agent-tools`. Versions 0.1.0,
15
+ 0.1.1 and 0.1.2 were tagged but never published; their outcomes are recorded below.
16
+
17
+ ### Added
18
+
19
+ - Six evidence-oriented tools for pi and omp: `jev_ask`, `jev_ask_files`,
20
+ `jev_find_files`, `jev_locate_in_file`, `jev_check_diff` and `jev_select_tests`.
21
+ - Typed judgment intents, fixed diff checks, static test discovery and conservative
22
+ execution plans, with visible uncertainty, missing evidence and collection limits.
23
+ - A single non-blocking run-end check for stale existing documentation, with an
24
+ option to disable it.
25
+ - Repository-confined file evidence collection, explicit command controls and
26
+ session call/cost budgets. Optional commands use host permissions, not a sandbox.
27
+ - npm extension installation for both hosts, shared result-reading guidance,
28
+ tool reference documentation and architecture decisions.
29
+ - In pi, the decision policy ships with the `jev_ask` tool guidelines and is present
30
+ whenever `jev_ask` is active; omp retains native rule discovery.
31
+
32
+ ### Compatibility
33
+
34
+ Requires Node.js 24 or later. Supported host baseline: pi 0.87.1 or later, or
35
+ omp 18.4.10 or later. Install with `pi install npm:jev-agent-tools@0.1.3` or
36
+ `omp plugin install jev-agent-tools@0.1.3`. Omp uses its Bun runtime; Node.js is also
37
+ required for Node-based project checks. These are support baselines, not the first
38
+ host releases to implement package installation. Later host versions are expected
39
+ to remain compatible but are not all individually validated. Optional native
40
+ dependencies may be unavailable on some platforms; affected tools report their
41
+ limitations. The HTTP client is compatible with the Jev API format; configure your
42
+ own endpoint and credentials.
43
+
44
+ ### Fixed
45
+
46
+ - The publish workflow passes the approved tarball to npm as a local path.
47
+ - In pi, the `jev_ask` policy appears once in the system prompt instead of twice.
48
+
49
+ ## [0.1.2]
50
+
51
+ Tagged but never published to npm: the scoped candidate was cancelled before
52
+ publication after the package name changed to `jev-agent-tools`.
53
+
54
+ ## [0.1.1]
55
+
56
+ Tagged but never published to npm: the unscoped package name was rejected.
57
+
58
+ ## [0.1.0]
59
+
60
+ Tagged but never published to npm: the publish workflow failed before upload.
61
+
62
+ [Unreleased]: https://github.com/NomenAK/jev-tools/compare/v0.1.3...HEAD
63
+ [0.1.3]: https://github.com/NomenAK/jev-tools/compare/v0.1.0...v0.1.3
64
+ [0.1.2]: https://github.com/NomenAK/jev-tools/tree/v0.1.2
65
+ [0.1.1]: https://github.com/NomenAK/jev-tools/tree/v0.1.1
66
+ [0.1.0]: https://github.com/NomenAK/jev-tools/tree/v0.1.0
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 NomenAK
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,119 @@
1
+ # jev-agent-tools
2
+
3
+ Six evidence-oriented tools for pi and omp, compatible with the Jev API format. Use them to navigate unfamiliar code, ask typed questions about repository evidence, review completed changes and select existing tests. They complement reading, searching and execution; they do not replace them.
4
+
5
+ ## Install
6
+
7
+ Install through your host's package manager. The npm package is `jev-agent-tools`. It ships TypeScript sources; no separate compilation is required.
8
+
9
+ ### pi
10
+
11
+ ```sh
12
+ pi install npm:jev-agent-tools@0.1.3
13
+ # Project-local installation:
14
+ pi install -l npm:jev-agent-tools@0.1.3
15
+ ```
16
+
17
+ ### omp
18
+
19
+ ```sh
20
+ omp plugin install jev-agent-tools@0.1.3
21
+ ```
22
+
23
+ ### Requirements and compatibility
24
+
25
+ Requires Node.js 24 or later. Supported host baselines are pi 0.87.1 and omp 18.4.10. These are support baselines, not claims that every later version has been individually validated. omp uses its Bun runtime; Node.js is also required for Node-based project checks. Git and, for optional command evidence, Bash must be available. Optional native parsers and file-search acceleration may be unavailable on some platforms; affected tools report their limitations.
26
+
27
+ ## Configure
28
+
29
+ Set these before starting the host, using your own endpoint and credentials:
30
+
31
+ ```sh
32
+ export JEV_TOOLS_URL="${YOUR_JEV_ENDPOINT}"
33
+ export JEV_TOOLS_API_KEY="${YOUR_JEV_API_KEY}"
34
+ # Optional; already the default:
35
+ export JEV_TOOLS_MODEL="openjev"
36
+ ```
37
+
38
+ `YOUR_*` placeholders are inputs you supply, not additional product settings. Configuration is read when the extension loads; restart the host after changing it.
39
+
40
+ | Variable | Meaning |
41
+ |---|---|
42
+ | `JEV_TOOLS_URL` | Required complete endpoint URL compatible with the Jev API format. |
43
+ | `JEV_TOOLS_API_KEY` | Required Bearer credential; configuration values are not printed in tool output. |
44
+ | `JEV_TOOLS_MODEL` | Requested model string, default `openjev`; a moving alias, not a guarantee of served-model identity. |
45
+ | `JEV_TOOLS_MAX_CALLS` | Session-wide non-negative safe-integer call limit; absent or empty means unlimited. Invalid values refuse requests. |
46
+ | `JEV_TOOLS_MAX_USD` | Session-wide finite non-negative cost limit, including fractions; absent or empty means unlimited. Invalid values refuse requests. |
47
+ | `JEV_TOOLS_ALLOW_COMMAND` | `0` disables `command` in `jev_ask`; otherwise commands run with ordinary shell permissions, without an additional sandbox. |
48
+ | `JEV_TOOLS_AUTO_DOCS` | `0` disables the automatic run-end documentation check. |
49
+
50
+ Without the endpoint or key, tools remain registered and explain the missing configuration; the automatic documentation check is disabled. There is no fallback to a chat model. Per-tool `max_calls` is separate from session limits. A model name echoed by the response does not establish which model was actually served.
51
+
52
+ ## Choose a tool
53
+
54
+ | Tool | Choose it when |
55
+ |---|---|
56
+ | [jev_ask](docs/tools/jev_ask.md) | One judgment combines a note, files, earlier versions or command output. |
57
+ | [jev_ask_files](docs/tools/jev_ask_files.md) | The same questions apply independently to each candidate file. |
58
+ | [jev_find_files](docs/tools/jev_find_files.md) | You need an entry point for a behavioral goal and do not know its filename. In omp, use native `find` when active. |
59
+ | [jev_locate_in_file](docs/tools/jev_locate_in_file.md) | You need the relevant range in one file of at least 19 KB. |
60
+ | [jev_check_diff](docs/tools/jev_check_diff.md) | Completed changes need risk, existing-documentation or specification review. |
61
+ | [jev_select_tests](docs/tools/jev_select_tests.md) | You need commands for affected existing tests, without running or collecting them. |
62
+
63
+ Use native read/search tools or code for exact source text, known symbols, filenames, line numbers, counts and arithmetic. Run commands yourself when you need their full output. Tool reference examples use fictional repository paths and are illustrative calls, not recorded executions.
64
+
65
+ ## Read the results
66
+
67
+ A line without a mark is a **verdict**: a lead to check before editing, deleting or reporting completion, not a proof. Probabilities concern the evidence shown, not everything in your repository.
68
+
69
+ | Mark | Meaning and next action |
70
+ |---|---|
71
+ | `unsure` | The answer is ambiguous or a control failed. Read the indicated passage or add the specific evidence that would settle it. Do not merely reword the question. |
72
+ | `abstain` | A necessary piece is missing. Add the named file or command evidence and ask once. |
73
+ | `no (not shown)` / `not addressed` | The supplied evidence does not show the statement; that does not make it false. |
74
+ | `uncalibrated` | No established error-rate calibration applies to this ask; treat it as a hint even if its probability is high. |
75
+ | Bracketed lines | Collection, parsing, budget or display limitations, with the next manual action. |
76
+
77
+ Ordinary boolean verdict bands are at or below 0.20 and at or above 0.80; category/level verdicts require a leading-option probability of at least 0.85 after applicable controls. Fixed checks and navigation tools have their own thresholds, described in their references and [design](docs/design.md).
78
+
79
+ The footer reports **calls · questions · cost · cache · time**: request count, questions judged, reported USD cost, cache hits/requests and elapsed time. A tool invocation may require several requests for batching or controls. Missing cost reporting is not evidence of a free request.
80
+
81
+ In a final report, explicitly identify conclusions marked `unsure` or `abstain` as unconfirmed by Jev. If subsequent reading settles them, distinguish that verification from the tool's result and cite the decisive evidence. Otherwise retain the uncertainty in your summary and recommendation.
82
+
83
+ ## Usage guidance for pi and omp
84
+
85
+ The extension supplies the shared reading guide in both hosts. omp discovers enabled npm plugin rules during normal startup; pi does not automatically discover the package's `rules/` directory. In pi, the same decision policy is part of the `jev_ask` tool guidelines, so it is present whenever `jev_ask` is active. A forced opaque prompt override may bypass this integration; disabled tools or disabled omp rules are not covered. This README block is recommended usage guidance, not itself an installed instruction:
86
+
87
+ > Before concluding that a failure is a code bug, an incorrect test or an environment problem, or that a plan matches documentation, pass the relevant files to [jev_ask](docs/tools/jev_ask.md) and weigh its answer against your own reading. Include both the failing test and the code it exercises; identify any conclusion that remains unconfirmed.
88
+
89
+ ## Data and command safety
90
+
91
+ Repository evidence, notes and optional command output are sent to your configured endpoint. Review its data-handling policy before using confidential repositories. See [security guidance](SECURITY.md).
92
+
93
+ File collection is confined to the repository: absolute paths, parent traversal, escaping symlinks, Git metadata and internal URLs are not file inputs. Build output, binaries, lockfiles and oversized files are skipped or refused with visible limits; evidence is not silently truncated into a verdict. This confinement does **not** sandbox a command. `jev_ask` commands can read, write or access the network with the host's shell permissions. omp uses execution approval for commands; pi does not supply an additional per-tool command approval. Set `JEV_TOOLS_ALLOW_COMMAND=0` to disable them.
94
+
95
+ ## Automatic documentation check
96
+
97
+ On a dirty tree, the extension can check existing Markdown documentation once at run end against changes from `HEAD`, including untracked files. A flagged existing sentence can request one additional turn to update it or explain why it remains correct. Merely unsure sections do not trigger another turn. Missing configuration, disabled automation, invalid/exhausted session budgets or a clean tree skip the check. Errors and timeout do not block the host. This is not a check for every missing documentation obligation.
98
+
99
+ ## Known limits
100
+
101
+ Static evidence and probability do not prove execution, safety or completeness. Import closure cannot discover every relationship; dynamic code, unsupported syntax and optional parser failures leave explicit gaps. No findings is not proof that unseen callers or documentation are correct. Test selection considers existing discovered scenarios, not whether a new scenario must be added. Session caching cannot establish the identity or stability of a moving model alias.
102
+
103
+ ## Development and contributions
104
+
105
+ See [CONTRIBUTING.md](CONTRIBUTING.md), [CHANGELOG.md](CHANGELOG.md), [design](docs/design.md) and [architecture decisions](docs/adr/). Development checks run locally without contacting a judgment endpoint:
106
+
107
+ ```sh
108
+ npm ci --include=optional
109
+ npm run typecheck
110
+ npm run check:imports
111
+ npm run lint
112
+ npm test
113
+ ```
114
+
115
+ Internal source modules are implementation details, not a stable library interface.
116
+
117
+ ## License
118
+
119
+ [MIT](LICENSE).
package/docs/design.md ADDED
@@ -0,0 +1,187 @@
1
+ # Design
2
+
3
+ The tools construct bounded evidence before asking for judgment. They preserve uncertainty and omissions instead of treating probability as proof.
4
+
5
+ ## Evidence before judgment
6
+
7
+ Supply the discriminating evidence, not an argument about it. For comparisons, show both sides: a failing test and its implementation, or a file before and after. A bounded static import closure adds declarations before judgment where supported; it cannot establish completeness or discover relationships without imports. Commands are evidence sources, not a sandbox or a substitute for reading the output directly.
8
+
9
+ ## Pure core, thin adapters
10
+
11
+ Pure core transformations construct states, units, questions and display envelopes. Adapters own filesystem, Git, parsing and host effects; the HTTP client owns transport. Presets define fixed review questions and texts define host-facing guidance. Dependency checks keep shared contracts below consumers, reject cycles and account for erased type imports. Pure path operations and erased parser types are explicit architectural exceptions. Both hosts use the same HTTP judgment protocol.
12
+
13
+ ## Typed intents and fixed checks
14
+
15
+ Caller intents compile to typed questions with canonical options and exact statement text. Verification of one combined situation preserves the distinction between holds, contradicted, not addressed and cannot tell, with an exact-statement cross-check that cannot override missing evidence. Custom questions remain uncalibrated. Diff checks use reviewed questions over before/after evidence units; tests are evidence, not changed units to judge. A matrix asks boolean cells across units; a pointer chooses one candidate or none. Neither executes the code.
16
+
17
+ ## Probability is not proof
18
+
19
+ Leading-option probability (p_max) is the probability of the most likely choice, not the response's separate confidence field. Gray bands and applicable option-order checks expose ambiguity. Missing evidence remains abstention. Preset witnesses use a decoy and a known positive reference to detect a biased setup; failed controls do not promote findings. No mark establishes that unseen evidence is complete. See [result reading](../README.md#read-the-results).
20
+
21
+ ## Bounded work, visible limits
22
+
23
+ Admission limits, request planning, rate limiting and retry bounds are distinct. Oversized evidence is refused or visibly omitted, never silently converted into an ordinary verdict. Per-invocation max_calls and session call/cost limits are separate; control and severity requests count too. Unjudged tests stay selected. Automatic documentation review is the only run-end automation and can request at most one extra turn.
24
+
25
+ Successful non-command judgments are cached only in session memory by canonical evidence, question and requested model string. Errors are not cached; command judgments bypass caching. The default openjev alias can move, and an echoed model name does not prove served-model identity. Fractional budgets bound admission, not an exact prediction of the final request's cost.
26
+
27
+ ## Policy thresholds
28
+
29
+ Values below are implementation policy, not measurement results, universal accuracy guarantees or settings to edit during a session. Boolean bands preserve weak answers as unsure; choice bands require a clear leader. Fixed findings use a separate threshold. Test selection favors retaining potentially affected tests and whole-file commands when name filtering is fragile.
30
+
31
+ ### Judgment policy
32
+
33
+ | Constant and value | Purpose |
34
+ |---|---|
35
+ | `BAND_BOOL_GRAY_A = 0.2` | Leave intermediate yes/no probabilities visibly unsure rather than reporting a weak verdict. |
36
+ | `BAND_BOOL_YES_MIN = 0.8` | Leave intermediate yes/no probabilities visibly unsure rather than reporting a weak verdict. |
37
+ | `BAND_CHOICE_VERDICT_MIN = 0.85` | Require a clear leading option for a choice or level verdict; this does not calibrate arbitrary caller asks. |
38
+ | `ORDER_REVERSE_BELOW = 0.85` | Recheck ambiguous substantive option order and expose order sensitivity instead of promoting it. |
39
+ | `ORDER_DISAGREE_MIN = 0.1` | Recheck ambiguous substantive option order and expose order sensitivity instead of promoting it. |
40
+ | `CANNOT_TELL_MIN = 0.3` | Preserve an explicit missing-evidence outcome before applying ordinary verdict bands. |
41
+ | `FLAG_MIN = 0.7` | Use one finding threshold across fixed review checks, distinct from caller-ask bands. |
42
+ | `SELECT_MIN = 0.5` | Prefer running an extra test over dropping a potentially affected test. |
43
+ | `SELECT_FILE_SHARE = 0.8` | Run the whole file when filtering would save little or introduce fragile name selection. |
44
+ | `DOCS_CHECK_MIN = 0.2` | Surface potentially stale wording for manual checking without making it a finding. |
45
+ | `WITNESS_LURE_MAX = 0.15` | Use size-aware decoy limits to detect a biased preset setup. |
46
+ | `WITNESS_LARGE_LURE_MAX = 0.25` | Use size-aware decoy limits to detect a biased preset setup. |
47
+ | `WITNESS_LARGE_CHARS = 13000` | Use size-aware decoy limits to detect a biased preset setup. |
48
+ | `COVERAGE_WITNESS_LURE_MAX = 0.2` | Apply the coverage-specific small-decoy policy separately from risk decoys. |
49
+ | `WITNESS_ETALON_MIN = 0.7` | Check a known positive reference and enable automatic witnesses on sufficiently large matrices. |
50
+ | `WITNESS_AUTO_MIN_CELLS = 8` | Check a known positive reference and enable automatic witnesses on sufficiently large matrices. |
51
+
52
+ ## Engineering bounds
53
+
54
+ Collection, transport and display limits have separate purposes. The state limit counts serialized characters; byte admission and token planning are different checks. Token estimates plan batching, not guaranteed tokenizer acceptance. Command stream limits are checked after execution and do not cap disk use or stop a running command.
55
+
56
+ <details>
57
+ <summary>Collection, transport and display constants</summary>
58
+
59
+ ### Evidence and request bounds
60
+
61
+ | Constant and value | Purpose |
62
+ |---|---|
63
+ | `STATE_MAX_CHARS = 80000` | Bound admitted evidence and describe the file limit consistently; character admission is not guaranteed token acceptance. |
64
+ | `FILE_MAX_KB = 80` | Bound admitted evidence and describe the file limit consistently; character admission is not guaranteed token acceptance. |
65
+ | `TEST_SOURCE_MAX_BYTES = 1280000` | Bound test-source reads before discovery while keeping constructed states within their separate limit. |
66
+ | `ASK_NOTE_MAX_CHARS = 8000` | Keep a combined situation's caller note and directly requested files bounded. |
67
+ | `ASK_MAX_FILES = 20` | Keep a combined situation's caller note and directly requested files bounded. |
68
+ | `MAX_FILES = 255` | Bound bulk file triage and choice cardinality; pointer choices must leave room for `none`. |
69
+ | `CHOICE_MAX_OPTIONS = 255` | Bound bulk file triage and choice cardinality; pointer choices must leave room for `none`. |
70
+ | `SCORE_MIN_LEVELS = 2` | Require a small, concrete set of ordered scoring situations. |
71
+ | `SCORE_MAX_LEVELS = 10` | Require a small, concrete set of ordered scoring situations. |
72
+ | `UNIT_MAX_CHARS = 6000` | Keep changed evidence focused through slices/hunks with local context. |
73
+ | `UNIT_CONTEXT_LINES = 8` | Keep changed evidence focused through slices/hunks with local context. |
74
+ | `CLOSURE_MAX_CHARS = 20000` | Bound automatically added static declarations; the retained budget is not a demonstrated optimum or a completeness guarantee. |
75
+ | `REQUEST_MAX_TOKENS = 60000` | Leave planning headroom when forming question batches, without rejecting admissible states on a heuristic alone. |
76
+ | `QUESTION_TOKENS = 1 / 3.3` | Estimate batches only; these coefficients are not exact costs or universal tokenizer bounds. |
77
+ | `STATE_TOKENS = 0.283` | Estimate batches only; these coefficients are not exact costs or universal tokenizer bounds. |
78
+ | `REQUEST_BASE_TOKENS = 456` | Estimate batches only; these coefficients are not exact costs or universal tokenizer bounds. |
79
+ | `RATE_PER_SECOND = 8` | Share throughput and in-flight limits across requests rather than mistaking concurrency for rate control. |
80
+ | `CONCURRENCY = 8` | Share throughput and in-flight limits across requests rather than mistaking concurrency for rate control. |
81
+ | `TIMEOUT_MS = 10000` | Bound transport waiting and retry attempts. |
82
+ | `REQUEST_ATTEMPTS = 3` | Bound transport waiting and retry attempts. |
83
+ | `RETRY_BASE_MS = 500` | Back off transient failures with a bounded delay and handle longer server delays explicitly. |
84
+ | `RETRY_MAX_MS = 5000` | Back off transient failures with a bounded delay and handle longer server delays explicitly. |
85
+ | `HTTP_ERROR_MAX_CHARS = 300` | Keep error diagnostics concise and redact credentials. |
86
+ | `GIT_BLOB_BATCH_SIZE = 400` | Bound historical blob batches so object IDs and options fit command-line limits. |
87
+ | `RUNNER_PACKAGE_MAX_BYTES = 65536` | Bound local runner-version metadata reads without sending package contents as judgment evidence. |
88
+
89
+ ### Command evidence
90
+
91
+ | Constant and value | Purpose |
92
+ |---|---|
93
+ | `ASK_TIMEOUT_S = 60` | Give optional commands a bounded default and caller-selectable maximum. |
94
+ | `ASK_TIMEOUT_MAX_S = 300` | Give optional commands a bounded default and caller-selectable maximum. |
95
+ | `OUTPUT_REPEAT_MIN = 5` | Group recurring line shapes while preserving rarer failure evidence. |
96
+ | `OUTPUT_CHUNK_CHARS = 2500` | Bound the passage-finding stage when compressed output still cannot fit. |
97
+ | `OUTPUT_FIND_MAX_CALLS = 40` | Bound the passage-finding stage when compressed output still cannot fit. |
98
+ | `OUTPUT_HEAD_SHARE = 0.05` | Reserve beginning/end context within the final evidence stage, not as a substitute for rarity compression. |
99
+ | `OUTPUT_TAIL_SHARE = 0.15` | Reserve beginning/end context within the final evidence stage, not as a substitute for rarity compression. |
100
+ | `OUTPUT_FILE_MAX_BYTES = 67108864` | Refuse oversized captured streams after execution; this is not a disk or command-execution cap. |
101
+ | `OUTPUT_LINE_MAX_CHARS = 8192` | Bound streaming line/shape bookkeeping and report the resulting limits explicitly. |
102
+ | `OUTPUT_SHAPE_MAX_COUNT = 2048` | Bound streaming line/shape bookkeeping and report the resulting limits explicitly. |
103
+ | `OUTPUT_FAILURE_WINDOW_LINES = 20` | Preserve local failure neighborhoods when identifying the failing test. |
104
+
105
+ ### Documentation collection and result display
106
+
107
+ | Constant and value | Purpose |
108
+ |---|---|
109
+ | `DOCS_MAX_SECTIONS = 40` | Bound the set of existing Markdown sections judged by the documentation preset. |
110
+ | `HOOK_BUDGET_MS = 15000` | Limit automatic run-end work without turning failures into a blocking host loop. |
111
+ | `DOCS_COLLECT_BUDGET_MS = 750` | Stop progressive collection with omissions visible; this is not a strict wall-time or coverage guarantee. |
112
+ | `DOCS_NAME_MAX_FILES = 32` | Bound name-owner bookkeeping and per-anchor traversal rather than claiming complete repository reachability. |
113
+ | `DOCS_ANCHOR_MAX_FILES = 32` | Bound name-owner bookkeeping and per-anchor traversal rather than claiming complete repository reachability. |
114
+ | `CALLER_DISPLAY_MAX_SPANS = 8` | Keep local-caller evidence lines readable while preserving full details. |
115
+ | `LIMIT_DISPLAY_MAX_PATHS = 4` | Prioritize actionable limits and documentation locations while reporting remaining counts. |
116
+ | `DOCS_DISPLAY_MAX_SECTIONS = 5` | Prioritize actionable limits and documentation locations while reporting remaining counts. |
117
+
118
+ ### File finding
119
+
120
+ | Constant and value | Purpose |
121
+ |---|---|
122
+ | `FIND_NAME_CANDIDATES = 128` | Bound the default name-ranking and content-reading stages separately. |
123
+ | `FIND_NAME_BATCH = 64` | Bound the default name-ranking and content-reading stages separately. |
124
+ | `FIND_FILES_READ = 20` | Bound the default name-ranking and content-reading stages separately. |
125
+ | `FIND_READ_PER_CALL = 5` | Judge small groups with keyword-selected evidence rather than bare paths. |
126
+ | `FIND_EXCERPT_CHARS = 3000` | Judge small groups with keyword-selected evidence rather than bare paths. |
127
+ | `FIND_POINTER_EXCERPT_CHARS = 1500` | Judge small groups with keyword-selected evidence rather than bare paths. |
128
+ | `FIND_POINTER_MAX = 8` | Keep the entry-point choice focused on a bounded shortlist. |
129
+ | `FIND_NAME_GUARD_MIN = 0.9` | Preserve a strongly named leading candidate when its excerpt is less conclusive. |
130
+ | `FIND_NAME_GUARD_TOP = 3` | Preserve a strongly named leading candidate when its excerpt is less conclusive. |
131
+ | `FIND_CONFIRM_MIN = 0.8` | Separate entry confirmation from admission to the content shortlist. |
132
+ | `FIND_CONTENT_MIN = 0.5` | Separate entry confirmation from admission to the content shortlist. |
133
+ | `FIND_GOAL_MIN_WORDS = 5` | Require a behavioral sentence before presenting an unqualified entry verdict. |
134
+ | `FIND_QUICK_CANDIDATES = 32` | Offer a smaller bounded first-pass search. |
135
+ | `FIND_QUICK_READ = 8` | Offer a smaller bounded first-pass search. |
136
+ | `FIND_THOROUGH_CANDIDATES = 256` | Broaden related-file exploration without promising a better entry point. |
137
+ | `FIND_THOROUGH_READ = 40` | Broaden related-file exploration without promising a better entry point. |
138
+ | `FIND_EXCERPT_HEAD_CHARS = 1000` | Preserve leading file context when building excerpts. |
139
+ | `FIND_PATH_WEIGHT = 3` | Prefer lexical goal matches in paths during deterministic pre-ranking. |
140
+ | `FIND_GREP_PAGE_SIZE = 1024` | Bound lexical-result pages and streamed excerpt reads. |
141
+ | `FIND_READ_CHUNK_BYTES = 16384` | Bound lexical-result pages and streamed excerpt reads. |
142
+
143
+ ### Locating within a file
144
+
145
+ | Constant and value | Purpose |
146
+ |---|---|
147
+ | `LOCATE_MIN_KB = 19` | Route small files to direct reading; eligibility uses file bytes, not tokenization. |
148
+ | `LOCATE_VERDICT_MIN = 0.7` | Distinguish one clear range, two plausible ranges and a weak choice requiring narrowing. |
149
+ | `LOCATE_GRAY_MIN = 0.4` | Distinguish one clear range, two plausible ranges and a weak choice requiring narrowing. |
150
+ | `LOCATE_SHRINK_TOP = 3` | Retry only within a narrowed shortlist plus `none`, rather than appending arbitrary context. |
151
+ | `LOCATE_SECTION_MAX_LINES = 150` | Keep fallback sections bounded while avoiding tiny fragments. |
152
+ | `LOCATE_SECTION_MIN_LINES = 8` | Keep fallback sections bounded while avoiding tiny fragments. |
153
+ | `LOCATE_WINDOW_LINES = 80` | Bound streamed windows and their display labels. |
154
+ | `LOCATE_LABEL_MAX_CHARS = 120` | Bound streamed windows and their display labels. |
155
+ | `LOCATE_WHOLE_MAX_CHARS = 320000` | Bound whole-file admission before using the large-file path. |
156
+ | `LOCATE_WHOLE_MAX_BYTES = 1280000` | Bound whole-file admission before using the large-file path. |
157
+ | `LOCATE_READ_BUFFER_BYTES = 65536` | Bound the streaming read buffer. |
158
+ | `LOCATE_WINDOW_MAX_SERIALIZED_CHARS = 20000` | Leave serialization headroom for multi-window refinement and metadata; this is an engineering bound, not calibrated accuracy. |
159
+
160
+ ### Architecture policy
161
+
162
+ | Constant | Purpose |
163
+ |---|---|
164
+ | `IMPORT_LAYERS` | Enforce allowed downward local dependencies among core, presets, adapters, HTTP client and texts, including erased type imports, while rejecting cycles. |
165
+
166
+ </details>
167
+
168
+ ## Compatibility and limitations
169
+
170
+ Optional native syntax parsing and file-search acceleration can be absent. Tools report unsupported languages, incomplete parsing and heuristic sections rather than implying precise coverage. Static test discovery reads literal configuration without evaluating third-party code, keeps runtime and type tests separate, and preserves runner/project options. It cannot infer missing test obligations. File finding ranks names then reads bounded excerpts, not every file in the repository. Large-file location uses sections or streamed window outlines with visible limits.
171
+
172
+ ## Glossary
173
+
174
+ | Term | Definition |
175
+ |---|---|
176
+ | Evidence state | The JSON evidence assembled for a judgment request. |
177
+ | Intent | The typed judgment the caller declares and code compiles into questions. |
178
+ | Claim | One self-contained factual statement supplied for verification. |
179
+ | Evidence unit | A changed declaration, slice, file or hunk shown with its before/after evidence. |
180
+ | Preset | A fixed, reviewed set of questions for a specific diff check. |
181
+ | Leading-option probability | The probability of the most likely choice after applicable controls; not the response's separate confidence field. |
182
+ | Witness | Known-reference evidence used to check a preset setup, not a runtime guarantee. |
183
+ | Result mark | unsure or abstain identifying a conclusion the tool did not confirm. |
184
+
185
+ ## Architecture decisions
186
+
187
+ See the [architecture decision records](adr/) for durable trade-offs. Start with the [README](../README.md) for installation and follow its six tool references for complete parameter contracts.
@@ -0,0 +1,92 @@
1
+ # jev_ask
2
+
3
+ Ask typed questions about one situation assembled from a note, files, earlier versions and optional command output.
4
+
5
+ ## Use for
6
+
7
+ Compare a failing test with the code it calls, check a plan against documentation, or judge behavior that depends on several pieces of evidence. Provide both sides of a comparison and the evidence that distinguishes rival explanations.
8
+
9
+ ## Use another tool when
10
+
11
+ Use [jev_ask_files](jev_ask_files.md) for independent answers per file. Use native read or shell execution when you need source text or full output, and native search for exact matches. Use [jev_check_diff](jev_check_diff.md) for fixed review checks over a finished diff.
12
+
13
+ ## Evidence model
14
+
15
+ The tool assembles one JSON situation. A note appears as `state`; repository files as `files["path"]`; historical versions as `files_before["path"]`; and command evidence as `output` with `command`, `exit_code`, `timed_out`, `stdout` and `stderr`. You receive answers, not the assembled file contents or captured output. Jev reads the evidence at a glance; it does not count, calculate or infer missing dependencies. Evidence outweighs an explanation of its role.
16
+
17
+ Before judgment, bounded depth-one static import closure adds supported used declarations and data files, including before/after versions when applicable. A `closure:` line records additions and gaps. This is not completeness: configuration, documentation and other flow relationships without imports must be passed explicitly. Named missing file references can be added if uniquely resolvable; unresolved or empty referenced evidence prevents the affected ask from being posed.
18
+
19
+ For recognizable assertion failures, the tool attempts to attach the failing test named by the log. An unidentifiable target produces a warning; any bug-versus-wrong-test conclusion under that warning remains unproven. Pass the failing test and implementation yourself whenever possible.
20
+
21
+ ## Parameters
22
+
23
+ At least one of `state`, `paths` or `command` must provide evidence.
24
+
25
+ | Field | Type / default | Meaning |
26
+ |---|---|---|
27
+ | `state` | Optional string, at most 8,000 characters | Short request, plan or observation not shown by the files; not a place to paste files or logs. JSON object text is accepted as structured state; `files` is reserved and rejected in such a note. |
28
+ | `paths` | Optional array of nonempty strings, at most 20 | Repository-relative files to read; not recursive directory or glob triage. |
29
+ | `base` | Optional nonempty Git ref | Add each requested file's earlier version; absent earlier files are `null`. Use `HEAD` for uncommitted before/after questions. Historical evidence counts toward admission. |
30
+ | `command` | Optional nonempty string | Execute `bash -c` at the repository root with `CI=1`, normal shell permissions and no additional sandbox. |
31
+ | `timeout_s` | Optional number, default 60, range 1–300 | Command timeout in seconds. |
32
+ | `asks` | Required intent, nonempty intent array, or JSON string encoding them | Typed questions; one intent per judgment. |
33
+ | `max_calls` | Optional non-negative integer | Limit requests including controls and output passage finding; unjudged work is reported. |
34
+
35
+ All intents may add `about: string` naming the relevant evidence, such as `state`, `output`, `files` or `files["src/billing.ts"]`.
36
+
37
+ | Intent | Fields | Meaning |
38
+ |---|---|---|
39
+ | `verify` | `claims: {id: statement}` | Verify one positive self-contained factual statement per claim, preserving its exact text. |
40
+ | `classify` | `categories: {name: description}`, optional `pick: "one"` (default) or `"many"`, optional `by: string` | Pick a category or independently judge each; `by` describes the dimension. |
41
+ | `decide` | `hypotheses: {name: statement}` | Choose among mutually exclusive rivals, including fallback outcomes. |
42
+ | `rate` | `dimension: string`, `levels: string[]` | Choose among 2–10 concrete ordered situations on one dimension. |
43
+ | `locate` | `target: string`, nonempty `among: string[]`, optional `count: "one"` (default) or `"many"`, optional `attribution: boolean` | Identify candidates serving the target. Attribution is supported only for `count: "one"`. |
44
+ | `free` | `question: {type, instructions, criteria?}` | Custom uncalibrated bool, choice or score. |
45
+
46
+ Nonempty maps are limited to 255 entries, including room for generated choice fallbacks. `free.instructions` must be nonempty; bool criteria optionally describe `true`/`false`, choice criteria map options to descriptions, and score criteria contain 2–10 levels. Unknown fields and invalid intent shapes are rejected. Claims should not be questions, counts, dates or compound assertions; warnings preserve your original wording. Put every ask you need in one request. Attribution makes an additional judgment without the pointed file: `attributed` means the pick moved away; `does not rest on it` means it did not. This is an evidence-dependence check, not runtime causality.
47
+
48
+ ## Example
49
+
50
+ Illustrative call with fictional repository paths, not a recorded execution:
51
+
52
+ ```json
53
+ {
54
+ "paths": ["src/billing.ts", "test/billing.test.ts"],
55
+ "base": "HEAD",
56
+ "asks": [
57
+ {
58
+ "intent": "verify",
59
+ "about": "files[\"src/billing.ts\"]",
60
+ "claims": {"rounding": "prorate rounds down to the whole cent"}
61
+ },
62
+ {
63
+ "intent": "classify",
64
+ "about": "files",
65
+ "categories": {
66
+ "bug_in_code": "the implementation violates behavior required by the test and contract",
67
+ "wrong_test": "the assertion contradicts the documented contract"
68
+ },
69
+ "pick": "one"
70
+ }
71
+ ],
72
+ "max_calls": 4
73
+ }
74
+ ```
75
+
76
+ ## Results and next action
77
+
78
+ Verification distinguishes `holds`, `contradicted`, `not addressed by the state` and `cannot tell`. An exact-statement boolean cross-check cannot override evidence issues or prove contradiction alone; scope/evidence disagreement produces unsure with raw values. Choices under 0.85 or failed controls produce unsure; missing evidence produces abstain and names the needed piece. Read surprising files or output directly. Add the decisive evidence rather than rewording. See [shared result reading](../../README.md#read-the-results).
79
+
80
+ ## Limits and failure behavior
81
+
82
+ Serialized evidence, including keys and escapes, is bounded to 80,000 characters. Oversized files/state are refused with a split suggestion, not silently shortened. Repository file confinement rejects escaping paths and internal URLs.
83
+
84
+ Commands are captured in private temporary files and cleaned up. Captured streams over 64 MiB are refused after execution; this neither caps disk usage nor interrupts execution. Repeated line shapes are compressed; lines over 8,192 characters and more than 2,048 distinct shapes produce visible notices. If compression still cannot fit, bounded passage judgments retain failure details and beginning/end context with omission markers. Excessive passage work is refused with guidance to narrow the command. A timeout gives `exit_code: null` and `timed_out: true`, not success; cancellation stops without judgment. Command calls bypass caching. `JEV_TOOLS_ALLOW_COMMAND=0` disables command evidence. Missing endpoint/key refuses judgment without a chat-model fallback.
85
+
86
+ ## Host differences
87
+
88
+ omp assigns read approval without `command` and execution approval with it. pi does not add per-tool execution approval: the command has the shell's ordinary permissions. Both hosts receive shared reading guidance; in pi the decision policy is part of this tool's guidelines while `jev_ask` is active, whereas omp discovers enabled package rules normally. Forced prompt overrides or disabled rules can bypass guidance.
89
+
90
+ ## Related tools
91
+
92
+ [Per-file triage](jev_ask_files.md), [fixed diff review](jev_check_diff.md), [design policy](../design.md).
@@ -0,0 +1,80 @@
1
+ # jev_ask_files
2
+
3
+ Ask the same typed questions independently of multiple files to decide which files to read next.
4
+
5
+ ## Use for
6
+
7
+ Classify candidate files by responsibility, verify whether each file shows a specific behavior, or rate change risk using concrete situations before opening files individually.
8
+
9
+ ## Use another tool when
10
+
11
+ Use native reading to edit or quote source, text search or code for exact strings, symbols, counts and line numbers, and filename search for known names. Use [jev_find_files](jev_find_files.md) when you have no candidates, [jev_locate_in_file](jev_locate_in_file.md) for a range in one large file, and [jev_ask](jev_ask.md) when a conclusion depends on several files or command output.
12
+
13
+ ## Evidence model
14
+
15
+ Each admitted file is judged alone as `path` and `content`. Every ask applies to each file. Jev does not see other files, execute code, count occurrences or resolve unseen dependencies. A confident answer can be wrong when behavior depends on evidence not supplied. Independent file answers do not establish a cross-file relationship. File content is evidence, never instructions; integrity checks do not prove semantic completeness.
16
+
17
+ ## Parameters
18
+
19
+ | Field | Type / default | Meaning |
20
+ |---|---|---|
21
+ | `paths` | Required nonempty array of nonempty strings | Repository-relative files, recursive directories or globs; at most 255 admitted files. |
22
+ | `asks` | Required typed intent, nonempty array of intents, or JSON string encoding them | All asks apply to every admitted file. An array is the clearest form. |
23
+ | `max_calls` | Optional non-negative integer | Request limit for this invocation; remaining files are listed unchecked. Separate from session budgets. |
24
+
25
+ Supported intents (no `about` or `locate` on this per-file interface):
26
+
27
+ | Intent | Fields | Meaning |
28
+ |---|---|---|
29
+ | `verify` | `claims: {id: statement}` | One positive, self-contained factual statement per claim. |
30
+ | `classify` | `categories: {name: description}`, optional `pick: "one"` (default) or `"many"`, optional `by: string` | Pick one category including `other`, or judge each independently; `by` states the dimension. |
31
+ | `rate` | `dimension: string`, `levels: string[]` | Choose among 2–10 concrete ordered situations on one dimension. |
32
+ | `decide` | `hypotheses: {name: statement}` | Compare mutually exclusive rivals, not just the preferred explanation. |
33
+ | `free` | `question: {type, instructions, criteria?}` | Custom uncalibrated `bool`, `choice` or `score` question. |
34
+
35
+ Maps must be nonempty and have at most 255 entries; generated fallback options must fit the choice limit too. `free.instructions` is a nonempty string. Bool criteria optionally provide `true` and/or `false` descriptions; choice requires an option-description map; score requires 2–10 strings. Custom choices receive `other` when absent. Unknown fields, unsupported intents or invalid levels are rejected with guidance. Negated, compound or count/date claims can generate warnings without rewriting your statement. Use native search or code for arithmetic, counts, dates and comparisons between files. Put all asks in one request.
36
+
37
+ ## Example
38
+
39
+ Illustrative call with fictional repository paths, not a recorded execution:
40
+
41
+ ```json
42
+ {
43
+ "paths": ["src/auth/*.ts", "src/storage/*.ts"],
44
+ "asks": [
45
+ {
46
+ "intent": "verify",
47
+ "claims": {
48
+ "c1": "`content` validates authentication tokens",
49
+ "c2": "`content` writes to the database"
50
+ }
51
+ },
52
+ {
53
+ "intent": "classify",
54
+ "categories": {
55
+ "http": "routing and request parsing",
56
+ "domain": "business rules without I/O",
57
+ "storage": "queries and persistence"
58
+ },
59
+ "pick": "one"
60
+ }
61
+ ],
62
+ "max_calls": 4
63
+ }
64
+ ```
65
+
66
+ ## Results and next action
67
+
68
+ Results preserve the exact statement next to each file's answer. Boolean `yes` or `no (not shown)` includes the probability of yes; intermediate probabilities between 0.20 and 0.80 are unsure. Categories and levels need a leading-option probability of at least 0.85 after applicable controls. `classify`, `rate`, `decide` and `free` are uncalibrated; sharing a display band does not establish an error rate. Failed integrity or option-order controls make clear-looking results unsure. Read the candidate file before acting. Skipped and unchecked files and bracketed limits explain missing work. See [shared result reading](../../README.md#read-the-results).
69
+
70
+ ## Limits and failure behavior
71
+
72
+ Build output, binaries, lockfiles, ignored untracked files, empty files and oversized files are excluded or reported. The advertised 80 KB bound derives from 80,000-character evidence admission; readers also check bytes and valid UTF-8. Kept file evidence is not silently truncated. Paths cannot escape through absolute paths, parent segments, symlinks or Git metadata; internal URLs are not files. Missing configuration refuses judgment, never falls back to a chat model. Batching and controls mean one file judgment is not a promise of exactly one HTTP request.
73
+
74
+ ## Host differences
75
+
76
+ Judgment behavior is shared. Native search names adapt to the host. omp marks this tool read-only; pi does not automatically discover npm package rules, but the extension delivers its shared reading guide.
77
+
78
+ ## Related tools
79
+
80
+ [Combined evidence](jev_ask.md), [file finding](jev_find_files.md), [large-file location](jev_locate_in_file.md), [design policy](../design.md).