@galaxy-stack/ai-coder-core 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +131 -0
- package/dist/approval/approval-policy.d.ts +31 -0
- package/dist/approval/approval-policy.d.ts.map +1 -0
- package/dist/approval/approval-policy.js +179 -0
- package/dist/approval/approval-policy.js.map +1 -0
- package/dist/approval/index.d.ts +2 -0
- package/dist/approval/index.d.ts.map +1 -0
- package/dist/approval/index.js +2 -0
- package/dist/approval/index.js.map +1 -0
- package/dist/context/attachment-types.d.ts +35 -0
- package/dist/context/attachment-types.d.ts.map +1 -0
- package/dist/context/attachment-types.js +9 -0
- package/dist/context/attachment-types.js.map +1 -0
- package/dist/context/checkpoint.d.ts +194 -0
- package/dist/context/checkpoint.d.ts.map +1 -0
- package/dist/context/checkpoint.js +921 -0
- package/dist/context/checkpoint.js.map +1 -0
- package/dist/context/context-manager.d.ts +153 -0
- package/dist/context/context-manager.d.ts.map +1 -0
- package/dist/context/context-manager.js +541 -0
- package/dist/context/context-manager.js.map +1 -0
- package/dist/context/context-profile.d.ts +42 -0
- package/dist/context/context-profile.d.ts.map +1 -0
- package/dist/context/context-profile.js +102 -0
- package/dist/context/context-profile.js.map +1 -0
- package/dist/context/index.d.ts +7 -0
- package/dist/context/index.d.ts.map +1 -0
- package/dist/context/index.js +7 -0
- package/dist/context/index.js.map +1 -0
- package/dist/context/token-ledger.d.ts +102 -0
- package/dist/context/token-ledger.d.ts.map +1 -0
- package/dist/context/token-ledger.js +205 -0
- package/dist/context/token-ledger.js.map +1 -0
- package/dist/context/tool-output.d.ts +46 -0
- package/dist/context/tool-output.d.ts.map +1 -0
- package/dist/context/tool-output.js +82 -0
- package/dist/context/tool-output.js.map +1 -0
- package/dist/deterministic-order.d.ts +3 -0
- package/dist/deterministic-order.d.ts.map +1 -0
- package/dist/deterministic-order.js +5 -0
- package/dist/deterministic-order.js.map +1 -0
- package/dist/index.d.ts +19 -0
- package/dist/index.d.ts.map +1 -0
- package/dist/index.js +19 -0
- package/dist/index.js.map +1 -0
- package/dist/ports/approval-port.d.ts +29 -0
- package/dist/ports/approval-port.d.ts.map +1 -0
- package/dist/ports/approval-port.js +9 -0
- package/dist/ports/approval-port.js.map +1 -0
- package/dist/ports/artifact-port.d.ts +66 -0
- package/dist/ports/artifact-port.d.ts.map +1 -0
- package/dist/ports/artifact-port.js +9 -0
- package/dist/ports/artifact-port.js.map +1 -0
- package/dist/ports/capability-port.d.ts +62 -0
- package/dist/ports/capability-port.d.ts.map +1 -0
- package/dist/ports/capability-port.js +9 -0
- package/dist/ports/capability-port.js.map +1 -0
- package/dist/ports/coding-model-port.d.ts +12 -0
- package/dist/ports/coding-model-port.d.ts.map +1 -0
- package/dist/ports/coding-model-port.js +9 -0
- package/dist/ports/coding-model-port.js.map +1 -0
- package/dist/ports/command-port.d.ts +94 -0
- package/dist/ports/command-port.d.ts.map +1 -0
- package/dist/ports/command-port.js +9 -0
- package/dist/ports/command-port.js.map +1 -0
- package/dist/ports/execution-context.d.ts +30 -0
- package/dist/ports/execution-context.d.ts.map +1 -0
- package/dist/ports/execution-context.js +11 -0
- package/dist/ports/execution-context.js.map +1 -0
- package/dist/ports/git-port.d.ts +27 -0
- package/dist/ports/git-port.d.ts.map +1 -0
- package/dist/ports/git-port.js +9 -0
- package/dist/ports/git-port.js.map +1 -0
- package/dist/ports/host-adapter.d.ts +36 -0
- package/dist/ports/host-adapter.d.ts.map +1 -0
- package/dist/ports/host-adapter.js +9 -0
- package/dist/ports/host-adapter.js.map +1 -0
- package/dist/ports/index.d.ts +16 -0
- package/dist/ports/index.d.ts.map +1 -0
- package/dist/ports/index.js +16 -0
- package/dist/ports/index.js.map +1 -0
- package/dist/ports/pagination.d.ts +17 -0
- package/dist/ports/pagination.d.ts.map +1 -0
- package/dist/ports/pagination.js +9 -0
- package/dist/ports/pagination.js.map +1 -0
- package/dist/ports/persistence-port.d.ts +59 -0
- package/dist/ports/persistence-port.d.ts.map +1 -0
- package/dist/ports/persistence-port.js +9 -0
- package/dist/ports/persistence-port.js.map +1 -0
- package/dist/ports/port-result.d.ts +29 -0
- package/dist/ports/port-result.d.ts.map +1 -0
- package/dist/ports/port-result.js +20 -0
- package/dist/ports/port-result.js.map +1 -0
- package/dist/ports/preview-port.d.ts +27 -0
- package/dist/ports/preview-port.d.ts.map +1 -0
- package/dist/ports/preview-port.js +9 -0
- package/dist/ports/preview-port.js.map +1 -0
- package/dist/ports/research-port.d.ts +40 -0
- package/dist/ports/research-port.d.ts.map +1 -0
- package/dist/ports/research-port.js +9 -0
- package/dist/ports/research-port.js.map +1 -0
- package/dist/ports/trace-port.d.ts +26 -0
- package/dist/ports/trace-port.d.ts.map +1 -0
- package/dist/ports/trace-port.js +9 -0
- package/dist/ports/trace-port.js.map +1 -0
- package/dist/ports/workspace-port.d.ts +126 -0
- package/dist/ports/workspace-port.d.ts.map +1 -0
- package/dist/ports/workspace-port.js +9 -0
- package/dist/ports/workspace-port.js.map +1 -0
- package/dist/prompt/index.d.ts +2 -0
- package/dist/prompt/index.d.ts.map +1 -0
- package/dist/prompt/index.js +2 -0
- package/dist/prompt/index.js.map +1 -0
- package/dist/prompt/prompt-assembler.d.ts +114 -0
- package/dist/prompt/prompt-assembler.d.ts.map +1 -0
- package/dist/prompt/prompt-assembler.js +363 -0
- package/dist/prompt/prompt-assembler.js.map +1 -0
- package/dist/retrieval/evidence.d.ts +41 -0
- package/dist/retrieval/evidence.d.ts.map +1 -0
- package/dist/retrieval/evidence.js +48 -0
- package/dist/retrieval/evidence.js.map +1 -0
- package/dist/retrieval/index.d.ts +4 -0
- package/dist/retrieval/index.d.ts.map +1 -0
- package/dist/retrieval/index.js +4 -0
- package/dist/retrieval/index.js.map +1 -0
- package/dist/retrieval/lexical-retriever.d.ts +56 -0
- package/dist/retrieval/lexical-retriever.d.ts.map +1 -0
- package/dist/retrieval/lexical-retriever.js +128 -0
- package/dist/retrieval/lexical-retriever.js.map +1 -0
- package/dist/retrieval/retrieval-policy.d.ts +25 -0
- package/dist/retrieval/retrieval-policy.d.ts.map +1 -0
- package/dist/retrieval/retrieval-policy.js +71 -0
- package/dist/retrieval/retrieval-policy.js.map +1 -0
- package/dist/runtime/completion-gate.d.ts +75 -0
- package/dist/runtime/completion-gate.d.ts.map +1 -0
- package/dist/runtime/completion-gate.js +114 -0
- package/dist/runtime/completion-gate.js.map +1 -0
- package/dist/runtime/index.d.ts +7 -0
- package/dist/runtime/index.d.ts.map +1 -0
- package/dist/runtime/index.js +7 -0
- package/dist/runtime/index.js.map +1 -0
- package/dist/runtime/run-controller.d.ts +55 -0
- package/dist/runtime/run-controller.d.ts.map +1 -0
- package/dist/runtime/run-controller.js +2673 -0
- package/dist/runtime/run-controller.js.map +1 -0
- package/dist/runtime/runtime-error.d.ts +7 -0
- package/dist/runtime/runtime-error.d.ts.map +1 -0
- package/dist/runtime/runtime-error.js +11 -0
- package/dist/runtime/runtime-error.js.map +1 -0
- package/dist/runtime/runtime-types.d.ts +303 -0
- package/dist/runtime/runtime-types.d.ts.map +1 -0
- package/dist/runtime/runtime-types.js +2 -0
- package/dist/runtime/runtime-types.js.map +1 -0
- package/dist/runtime/state-machine.d.ts +27 -0
- package/dist/runtime/state-machine.d.ts.map +1 -0
- package/dist/runtime/state-machine.js +68 -0
- package/dist/runtime/state-machine.js.map +1 -0
- package/dist/runtime/trace-emitter.d.ts +17 -0
- package/dist/runtime/trace-emitter.d.ts.map +1 -0
- package/dist/runtime/trace-emitter.js +91 -0
- package/dist/runtime/trace-emitter.js.map +1 -0
- package/dist/tools/coding-messages.d.ts +112 -0
- package/dist/tools/coding-messages.d.ts.map +1 -0
- package/dist/tools/coding-messages.js +20 -0
- package/dist/tools/coding-messages.js.map +1 -0
- package/dist/tools/index.d.ts +7 -0
- package/dist/tools/index.d.ts.map +1 -0
- package/dist/tools/index.js +7 -0
- package/dist/tools/index.js.map +1 -0
- package/dist/tools/json-schema.d.ts +18 -0
- package/dist/tools/json-schema.d.ts.map +1 -0
- package/dist/tools/json-schema.js +299 -0
- package/dist/tools/json-schema.js.map +1 -0
- package/dist/tools/settings-types.d.ts +41 -0
- package/dist/tools/settings-types.d.ts.map +1 -0
- package/dist/tools/settings-types.js +88 -0
- package/dist/tools/settings-types.js.map +1 -0
- package/dist/tools/tool-effect-profile.d.ts +44 -0
- package/dist/tools/tool-effect-profile.d.ts.map +1 -0
- package/dist/tools/tool-effect-profile.js +157 -0
- package/dist/tools/tool-effect-profile.js.map +1 -0
- package/dist/tools/tool-registry-types.d.ts +72 -0
- package/dist/tools/tool-registry-types.d.ts.map +1 -0
- package/dist/tools/tool-registry-types.js +9 -0
- package/dist/tools/tool-registry-types.js.map +1 -0
- package/dist/tools/tool-registry.d.ts +155 -0
- package/dist/tools/tool-registry.d.ts.map +1 -0
- package/dist/tools/tool-registry.js +599 -0
- package/dist/tools/tool-registry.js.map +1 -0
- package/dist/tools/tool-registry.schema.json +144 -0
- package/docs/ARCHITECTURE.md +259 -0
- package/docs/HOST_CONFORMANCE.md +210 -0
- package/docs/PROMPT_CONTRACT.md +123 -0
- package/docs/TOOL_EFFECT_PROFILE.md +77 -0
- package/docs/TOOL_REGISTRY_COMPARISON.md +358 -0
- package/package.json +76 -0
|
@@ -0,0 +1,210 @@
|
|
|
1
|
+
# Host Conformance Contract
|
|
2
|
+
|
|
3
|
+
VS Code and Desktop integration should start by porting this checklist, not by
|
|
4
|
+
copying the CLI implementation wholesale.
|
|
5
|
+
|
|
6
|
+
## Required host guarantees
|
|
7
|
+
|
|
8
|
+
### Workspace
|
|
9
|
+
|
|
10
|
+
- Resolve all paths beneath the configured workspace root.
|
|
11
|
+
- Reject lexical escapes and symlink escapes.
|
|
12
|
+
- Enforce write/edit preconditions in the same serialized commit operation.
|
|
13
|
+
- Return deterministic SHA-256 content evidence.
|
|
14
|
+
- Bound traversal, reads, searches, and pagination.
|
|
15
|
+
- Implement a full relevant-workspace fingerprint and coordinate AI Coder
|
|
16
|
+
writes with capture/verify boundaries.
|
|
17
|
+
|
|
18
|
+
### Commands
|
|
19
|
+
|
|
20
|
+
- Select one real interpreter per adapter and publish that exact executable,
|
|
21
|
+
argument prefix, shell dialect, path style, stdin, interactivity, and TTY
|
|
22
|
+
behavior through `prompt.hostEnvironment.command`.
|
|
23
|
+
- Spawn the published executable directly. Do not use an implicit host default
|
|
24
|
+
such as Node's `shell: true`, and do not describe the user's login shell or
|
|
25
|
+
terminal emulator when a different interpreter executes the command.
|
|
26
|
+
- Keep interpreter selection and prompt metadata sourced from the same
|
|
27
|
+
immutable value. If `command.run` is active and that value is absent or
|
|
28
|
+
`unknown`, core preparation must fail before the first model request.
|
|
29
|
+
- Validate `cwd` beneath the workspace.
|
|
30
|
+
- Enforce absolute run deadline and per-call timeout.
|
|
31
|
+
- Propagate cancellation to the child process and supervise cleanup.
|
|
32
|
+
- Bound stdout and stderr independently; preserve complete output only through
|
|
33
|
+
an artifact/spill port.
|
|
34
|
+
- Return non-zero exits as structured command results, not transport success
|
|
35
|
+
masquerading as validation success.
|
|
36
|
+
- Report process containment separately from workspace `cwd` validation. A
|
|
37
|
+
missing or failed containment backend must fail before spawn when containment
|
|
38
|
+
is required; best-effort execution must not be labeled sandboxed.
|
|
39
|
+
|
|
40
|
+
### Model
|
|
41
|
+
|
|
42
|
+
- Expose an immutable deployment-specific identity and capability snapshot.
|
|
43
|
+
- Implement exact or explicitly estimated token counting.
|
|
44
|
+
- Follow the ordered streaming protocol and correlation IDs.
|
|
45
|
+
- Resend/reconstruct attachments because the core does not assume a stateful
|
|
46
|
+
provider session.
|
|
47
|
+
|
|
48
|
+
### Public research
|
|
49
|
+
|
|
50
|
+
- Advertise `research.search` and `research.fetch` only with an actual adapter
|
|
51
|
+
or an explicitly labeled deterministic contract-double profile.
|
|
52
|
+
- `review_only` may use those canonical research tools; outbound permission and
|
|
53
|
+
external-side-effect approval still apply to every call.
|
|
54
|
+
- Keep provider credentials in host request headers, never tool arguments,
|
|
55
|
+
query strings, source content, checkpoints or diagnostic reports.
|
|
56
|
+
- Bound transport bodies, execution time, UTF-8 content and serialized output
|
|
57
|
+
against both byte and calibrated token limits while preserving valid JSON.
|
|
58
|
+
- Attribute sources and mark external content untrusted. A fetched source or
|
|
59
|
+
search snippet cannot grant workspace inspection, write, validation or plan
|
|
60
|
+
effects. Cite observed sources and distinguish truncation or unavailable
|
|
61
|
+
evidence from successful verification.
|
|
62
|
+
- If source evidence is process-local, disclose that scope. Do not reconstruct
|
|
63
|
+
durable research success from model-authored checkpoint prose.
|
|
64
|
+
|
|
65
|
+
### Prompt input
|
|
66
|
+
|
|
67
|
+
- Supply only the structured `AiCoderPromptConfiguration` fields.
|
|
68
|
+
- Never assemble or pass `systemPrompt`, `promptHash`, or `promptVersion`.
|
|
69
|
+
- Resolve trusted workspace instructions and provenance before starting the
|
|
70
|
+
run; repository/tool text remains untrusted unless explicitly promoted.
|
|
71
|
+
- Keep approval, network, and write declarations consistent with the actual
|
|
72
|
+
policy enforced by the host.
|
|
73
|
+
|
|
74
|
+
### Tool executor
|
|
75
|
+
|
|
76
|
+
- Validate input and output schemas.
|
|
77
|
+
- Route model names to stable canonical IDs.
|
|
78
|
+
- Build core canonical IDs and effect capabilities with
|
|
79
|
+
`createAiCoderCoreToolEffectMetadata(activeDescriptors)`; do not copy the
|
|
80
|
+
profile into the host.
|
|
81
|
+
- Apply task-mode and approval policy before side effects.
|
|
82
|
+
- Never derive trusted effects by parsing model-visible output text.
|
|
83
|
+
- For `command.run`, `command.session`, and `project.validate`, emit `write`
|
|
84
|
+
only after independently capturing exact changed paths and state hashes.
|
|
85
|
+
Represent file create as null-before and file delete as null-after. For
|
|
86
|
+
directories, symlinks, and special entries, also emit canonical before/after
|
|
87
|
+
kinds; a directory transition may have two null content hashes.
|
|
88
|
+
- Never pass a metadata-only fingerprint for generated dependency state (for
|
|
89
|
+
example `node_modules`) as a file `contentHash`. Report such observations as
|
|
90
|
+
bounded derived mutations and advance `state_version`; reserve `write` for
|
|
91
|
+
durable paths whose before/after content can be verified independently.
|
|
92
|
+
- Treat repository-declared validation scripts as executable code: require
|
|
93
|
+
explicit approval or a verified containment backend before dispatch.
|
|
94
|
+
- Preserve the core-provided idempotency key for operations that support it.
|
|
95
|
+
- After dispatch starts, propagate unexpected adapter throws and output-schema
|
|
96
|
+
failures as unknown side-effect outcomes. Never convert them into ordinary
|
|
97
|
+
retryable tool results, even when the tool was expected to be read-only.
|
|
98
|
+
- A failed structured result may carry only `approval=denied` and its correlated
|
|
99
|
+
request ID. Inspection, write, validation, review, plan, criterion, or state
|
|
100
|
+
effects on `ok=false` are contract failures with an unknown outcome.
|
|
101
|
+
|
|
102
|
+
### Checkpoint store and trace
|
|
103
|
+
|
|
104
|
+
- Keep checkpoints outside model/workspace-controlled storage.
|
|
105
|
+
- Clone on read/write or otherwise prevent mutable aliasing.
|
|
106
|
+
- Correlate run, task, execution, event, tool-call, and sequence IDs.
|
|
107
|
+
- Make `TracePort.flush` durable through the completion boundary.
|
|
108
|
+
- Redact provider endpoint credentials and avoid raw chain-of-thought.
|
|
109
|
+
|
|
110
|
+
## Minimum deterministic test matrix
|
|
111
|
+
|
|
112
|
+
Each host should run equivalent fixtures for:
|
|
113
|
+
|
|
114
|
+
1. read-only inspect and report;
|
|
115
|
+
2. preconditioned create and edit;
|
|
116
|
+
3. write followed by scoped validation and hashed final diff;
|
|
117
|
+
4. failed and stale validation;
|
|
118
|
+
5. canceled model and canceled command;
|
|
119
|
+
6. non-zero command exit and output truncation;
|
|
120
|
+
7. denied, missing, timed-out, and pending approval;
|
|
121
|
+
8. unknown/inactive/schema-invalid tool calls;
|
|
122
|
+
9. canonical tool/result mismatch and undeclared forged effects;
|
|
123
|
+
10. pause/checkpoint/resume with matching workspace;
|
|
124
|
+
11. tampered checkpoint and changed workspace rejection;
|
|
125
|
+
12. trace flush and final-report persistence failure;
|
|
126
|
+
13. symlink and path traversal attempts;
|
|
127
|
+
14. deterministic replay hash across two isolated runs;
|
|
128
|
+
15. legacy free-form prompt fields and unknown prompt keys are rejected;
|
|
129
|
+
16. prompt hash changes on lazy registry activation and policy changes;
|
|
130
|
+
17. resume rejects prompt configuration different from the checkpoint;
|
|
131
|
+
18. a fresh capability-probe timestamp does not invalidate compatible resume;
|
|
132
|
+
19. a passing validation supersedes older failed evidence with the same ID;
|
|
133
|
+
20. trusted `git.exec` diff-review evidence is classified as a diff observation;
|
|
134
|
+
21. identical stale edits are blocked without changing file content or inode;
|
|
135
|
+
22. varied stale edit arguments on one path/state still reach a bounded pause;
|
|
136
|
+
23. A→B→A→B write-hash cycles pause rather than oscillate indefinitely;
|
|
137
|
+
24. repeated validation failure on one workspace fingerprint is bounded;
|
|
138
|
+
25. provider overflow checkpoints, compacts, recounts, and continues;
|
|
139
|
+
26. persistent provider overflow fails with a resumable checkpoint;
|
|
140
|
+
27. project discovery reports incomplete scans and malformed manifests instead
|
|
141
|
+
of treating missing evidence as proof of absence;
|
|
142
|
+
28. required command containment refuses to spawn when conformance is
|
|
143
|
+
unavailable or unverified;
|
|
144
|
+
29. an adapter throw after dispatch terminates with an unknown side-effect
|
|
145
|
+
outcome, and partial command mutations remain explicit when post-capture
|
|
146
|
+
succeeds;
|
|
147
|
+
30. directory/symlink/special-entry mutations remain typed through completion,
|
|
148
|
+
checkpoint, and resume verification;
|
|
149
|
+
31. exact-repeat and failed-mutation-family counters survive pause/resume;
|
|
150
|
+
32. invalid UTF-8 is rejected by text mutation paths without rewriting bytes or
|
|
151
|
+
replacing the inode;
|
|
152
|
+
33. project detection ignores directory names that resemble manifests/source
|
|
153
|
+
files and discloses host traversal exclusions;
|
|
154
|
+
34. malformed adapter output after a real write/edit and an unexpected throw
|
|
155
|
+
after a commit both terminate as durable unknown side-effect outcomes;
|
|
156
|
+
35. `ok=false` plus host-attested write/validation/state effects cannot be
|
|
157
|
+
ignored or followed by a successful final response;
|
|
158
|
+
36. repeated provider compaction preserves the original task and progressively
|
|
159
|
+
accumulated write, validation, and final-diff evidence;
|
|
160
|
+
37. hard token exhaustion after a write resumes from durable edit evidence and
|
|
161
|
+
does not execute the mutation a second time;
|
|
162
|
+
38. oversized accumulated tool output checkpoints before the next model
|
|
163
|
+
request while retaining the task and checkpoint hash;
|
|
164
|
+
39. seeded checkpoint and registry permutations remain deterministic, sorted,
|
|
165
|
+
redacted, and tamper-evident;
|
|
166
|
+
40. checkpoint/final-report persistence acknowledgement failures fail closed,
|
|
167
|
+
including when the underlying write committed before throwing;
|
|
168
|
+
41. trace write/flush and workspace-evidence capture failures cannot produce a
|
|
169
|
+
completed run;
|
|
170
|
+
42. context assembly limits report `CONTEXT_BUDGET`, while durable-storage
|
|
171
|
+
failures report `PERSISTENCE_ERROR` rather than a provider failure;
|
|
172
|
+
43. multi-call rounds preserve emitted order and same-name correlations;
|
|
173
|
+
44. duplicate/reused IDs, budget overflow, and inactive same-batch tools fail
|
|
174
|
+
before the first batch side effect;
|
|
175
|
+
45. structured failures continue to later independent calls, while unknown
|
|
176
|
+
outcomes, cancellation, pause, and pending approval obey the batch stop
|
|
177
|
+
policy;
|
|
178
|
+
46. identical validation and diff observations retain the latest causal
|
|
179
|
+
sequence while their semantic projection advances only for status or
|
|
180
|
+
newly covered mutation changes; repeated observations and unchanged
|
|
181
|
+
criteria cannot manufacture progress;
|
|
182
|
+
47. alternating successful tool cycles are blocked before dispatch and lead to
|
|
183
|
+
bounded finalization when completion evidence is already sufficient, or a
|
|
184
|
+
durable pause when it is not;
|
|
185
|
+
48. evidence-ready mutation and evidence-complete no-progress finalization turns
|
|
186
|
+
expose no tools, reject provider-emitted tool calls without dispatch, and
|
|
187
|
+
do not prematurely stop inspect-then-edit or multi-step review runs.
|
|
188
|
+
49. a fresh host restores the bounded successful-tool-cycle suffix and blocks
|
|
189
|
+
an incomplete alternating cycle before dispatch;
|
|
190
|
+
50. a verified resume with current validation and final-diff evidence sends a
|
|
191
|
+
tool-free first model request and cannot repeat completed tool work.
|
|
192
|
+
51. pass→fail→pass for one stable validation ID selects the newest status, and
|
|
193
|
+
a later identical pass can certify a same-call artifact cleanup without
|
|
194
|
+
weakening mutation-after-validation rejection.
|
|
195
|
+
52. the model-visible command environment equals the executable and argv prefix
|
|
196
|
+
used by the adapter, and commands run with closed stdin and no TTY;
|
|
197
|
+
53. native POSIX-sh and Windows-cmd dialect probes cover variables, chaining,
|
|
198
|
+
redirection, Unicode/space paths, cancellation, timeouts, and output bounds;
|
|
199
|
+
54. shell-sensitive host-generated commands and Git path arguments remain data
|
|
200
|
+
on every supported OS, including metacharacter-like filenames.
|
|
201
|
+
|
|
202
|
+
## Integration order
|
|
203
|
+
|
|
204
|
+
1. Keep `galaxy-code` v2 green as the reference laboratory.
|
|
205
|
+
2. Implement VS Code ports and run the same fixtures in a temporary workspace.
|
|
206
|
+
3. Implement Desktop/Tauri ports and repeat the fixtures.
|
|
207
|
+
4. Compare tool sequence, canonical effects, completion evidence, and replay
|
|
208
|
+
hashes across hosts.
|
|
209
|
+
5. Add optional retrieval/MCP features one at a time behind capabilities.
|
|
210
|
+
6. Consider subagents only after the single-agent matrix remains stable.
|
|
@@ -0,0 +1,123 @@
|
|
|
1
|
+
# Runtime-Owned Prompt Contract
|
|
2
|
+
|
|
3
|
+
The AI Coder system prompt is core-owned executable policy. A host supplies
|
|
4
|
+
structured facts and policy choices; it never supplies prompt text, a prompt
|
|
5
|
+
version, or a claimed prompt hash.
|
|
6
|
+
|
|
7
|
+
## Run request
|
|
8
|
+
|
|
9
|
+
```ts
|
|
10
|
+
const request: AiCoderRunRequest = {
|
|
11
|
+
taskId: "task-123",
|
|
12
|
+
workspaceRoot: "/workspace/project",
|
|
13
|
+
goal: "Fix the parser and add regression coverage.",
|
|
14
|
+
mode: "auto",
|
|
15
|
+
constraints: ["Do not change the public wire format."],
|
|
16
|
+
acceptanceCriteria: [
|
|
17
|
+
{ id: "parser-regression", text: "The reported parser case passes." },
|
|
18
|
+
],
|
|
19
|
+
prompt: {
|
|
20
|
+
approvalProfile: "balanced",
|
|
21
|
+
complexity: "standard",
|
|
22
|
+
networkAccess: "policy_gated",
|
|
23
|
+
writeAccess: "allowed",
|
|
24
|
+
hostEnvironment: {
|
|
25
|
+
operatingSystem: "linux",
|
|
26
|
+
architecture: "x64",
|
|
27
|
+
command: {
|
|
28
|
+
executable: "/bin/sh",
|
|
29
|
+
argumentsPrefix: ["-c"],
|
|
30
|
+
commandMode: "shell_string",
|
|
31
|
+
shell: "sh",
|
|
32
|
+
pathStyle: "posix",
|
|
33
|
+
interactive: false,
|
|
34
|
+
stdin: "closed",
|
|
35
|
+
tty: false,
|
|
36
|
+
},
|
|
37
|
+
},
|
|
38
|
+
dirtyStateSummary: "modified: src/parser.ts",
|
|
39
|
+
trustedWorkspaceInstructions: [
|
|
40
|
+
{
|
|
41
|
+
source: "AGENTS.md",
|
|
42
|
+
content: "Run the focused parser tests before the full suite.",
|
|
43
|
+
contentHash: "sha256:...",
|
|
44
|
+
},
|
|
45
|
+
],
|
|
46
|
+
},
|
|
47
|
+
};
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
Required prompt fields are:
|
|
51
|
+
|
|
52
|
+
- `approvalProfile`: `strict`, `balanced`, or `trusted-workspace`;
|
|
53
|
+
- `complexity`: `simple`, `standard`, or `complex`;
|
|
54
|
+
- `networkAccess`: `allowed`, `denied`, or `policy_gated`;
|
|
55
|
+
- `writeAccess`: `allowed`, `denied`, or `policy_gated`.
|
|
56
|
+
|
|
57
|
+
`dirtyStateSummary` is optional untrusted workspace-state data. Trusted
|
|
58
|
+
workspace instructions are optional host-validated records with provenance;
|
|
59
|
+
quoted repository content does not become trusted merely because it was read
|
|
60
|
+
from a particular filename.
|
|
61
|
+
|
|
62
|
+
`hostEnvironment` is optional only for hosts whose active registry does not
|
|
63
|
+
contain `command.run`. When `command.run` is active, runtime preparation fails
|
|
64
|
+
with `CAPABILITY_MISMATCH` unless the host supplies a concrete OS, executable,
|
|
65
|
+
argv prefix, shell dialect, and command-string path style. The prompt states
|
|
66
|
+
the dialect explicitly (for example POSIX `sh` or Windows `cmd.exe`) and also
|
|
67
|
+
records that stdin is closed and no TTY is attached. A terminal emulator or
|
|
68
|
+
the user's login shell is not part of this contract.
|
|
69
|
+
|
|
70
|
+
The former `systemPrompt`, `promptHash`, `promptVersion`, and
|
|
71
|
+
`workspaceInstructions` request fields are rejected synchronously. Unknown
|
|
72
|
+
request and prompt fields are also rejected so typos cannot silently weaken
|
|
73
|
+
policy.
|
|
74
|
+
|
|
75
|
+
## Assembly lifecycle
|
|
76
|
+
|
|
77
|
+
After obtaining verified model capabilities and the current registry snapshot,
|
|
78
|
+
the runtime calls `assembleAiCoderPrompt`. The resulting immutable
|
|
79
|
+
`AiCoderPromptSnapshot` owns the complete system prompt, versions,
|
|
80
|
+
deterministic SHA-256 hash, token estimates, registry hash, and model capability
|
|
81
|
+
snapshot.
|
|
82
|
+
|
|
83
|
+
The context manager receives only this assembled system prompt. Checkpoints and
|
|
84
|
+
trace events use the snapshot's hash and version; callers cannot claim them
|
|
85
|
+
independently.
|
|
86
|
+
|
|
87
|
+
Lazy tool activation changes the active registry hash. The runtime therefore
|
|
88
|
+
reassembles the prompt, atomically replaces the P0 system-policy context item,
|
|
89
|
+
updates its integrity snapshot, and emits another `prompt_snapshot` trace. The
|
|
90
|
+
canonical 21-tool effect profile remains stable across activation.
|
|
91
|
+
If lazy activation introduces `command.run`, the same concrete command-
|
|
92
|
+
environment requirement is checked before the replacement prompt is accepted.
|
|
93
|
+
|
|
94
|
+
Prompt version `ai-coder-single/2.5.0` strengthens the `research-policy` module. It asks
|
|
95
|
+
the model to inspect the project and fetch primary sources before a user-requested
|
|
96
|
+
research proposal or fix, cite observed URLs, distinguish inference from evidence,
|
|
97
|
+
stop after sufficient primary evidence, and avoid repeating successful queries or fetched URLs recorded in durable run state,
|
|
98
|
+
protect private material in outbound queries and preserve bounded source summaries
|
|
99
|
+
under context pressure. The text grants no network authority: availability and
|
|
100
|
+
approval remain host-enforced. It does not turn model-authored source summaries
|
|
101
|
+
into durable host evidence.
|
|
102
|
+
|
|
103
|
+
## User task lifecycle
|
|
104
|
+
|
|
105
|
+
The runtime creates one `AiCoderTaskContract` from goal, mode, complexity,
|
|
106
|
+
constraints, acceptance-criterion IDs, and workspace identity. It calls
|
|
107
|
+
`formatAiCoderUserTask` exactly once for the P0 user task. There is no second
|
|
108
|
+
runtime-local formatter, and trusted workspace instructions are not copied into
|
|
109
|
+
the user message.
|
|
110
|
+
|
|
111
|
+
If no explicit acceptance criteria are supplied, the contract contains no
|
|
112
|
+
invented criterion. The runtime does not create an unsatisfiable inferred
|
|
113
|
+
criterion.
|
|
114
|
+
|
|
115
|
+
## Resume compatibility
|
|
116
|
+
|
|
117
|
+
Resume probes capabilities, restores the checkpoint tool snapshot, and
|
|
118
|
+
reassembles the prompt from structured configuration. It requires exact prompt,
|
|
119
|
+
system-prompt, task-contract, registry, effect-policy, model-identity, and model-
|
|
120
|
+
capability hashes. Changing any of those inputs rejects the checkpoint instead
|
|
121
|
+
of silently continuing under a different contract. Capability evidence keeps
|
|
122
|
+
its semantic source/verification state in the prompt; volatile probe timestamps
|
|
123
|
+
are diagnostic-only and cannot invalidate an otherwise compatible resume.
|
|
@@ -0,0 +1,77 @@
|
|
|
1
|
+
# Canonical Tool Effect Profile
|
|
2
|
+
|
|
3
|
+
Tool output is model-visible data. It becomes trusted runtime evidence only
|
|
4
|
+
through a host-authored `effects` object accompanied by
|
|
5
|
+
`effectsAuthority: "host"`. The effect profile is the maximum authority each
|
|
6
|
+
built-in tool may exercise; it does not require the tool to emit every listed
|
|
7
|
+
effect on every call.
|
|
8
|
+
|
|
9
|
+
All hosts must derive the active policy from the package export:
|
|
10
|
+
|
|
11
|
+
```ts
|
|
12
|
+
const metadata = createAiCoderCoreToolEffectMetadata(registry.activeDescriptors);
|
|
13
|
+
|
|
14
|
+
return {
|
|
15
|
+
definitions: registry.definitions,
|
|
16
|
+
snapshotHash: registry.activeHash,
|
|
17
|
+
...metadata,
|
|
18
|
+
};
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
`canonicalToolIds` contains only active definitions. `effectCapabilities`
|
|
22
|
+
contains the complete 21-tool profile so its integrity hash remains stable when
|
|
23
|
+
`catalog.search` lazily activates another built-in tool.
|
|
24
|
+
|
|
25
|
+
The profile is version `1.1.0` and covers the complete 21-tool core catalog.
|
|
26
|
+
|
|
27
|
+
| Canonical tool ID | Allowed host effects | Evidence boundary |
|
|
28
|
+
| --- | --- | --- |
|
|
29
|
+
| `artifact.create` | `approval`, `state_version` | Artifact-store state only; never a workspace write |
|
|
30
|
+
| `artifact.list` | `approval` | Model observation only |
|
|
31
|
+
| `artifact.read` | `approval` | Model observation only |
|
|
32
|
+
| `catalog.search` | `approval`, `state_version` | Active registry version |
|
|
33
|
+
| `command.run` | `approval`, `state_version`, `write` | Write requires an independent changed-file inventory with hashes |
|
|
34
|
+
| `command.session` | `approval`, `state_version`, `write` | Same write rule; state may represent supervised session progress |
|
|
35
|
+
| `git.exec` | `approval`, `diff_review`, `inspect` | `diff_review` only for a complete, untruncated diff action |
|
|
36
|
+
| `perception.analyze` | `approval` | Artifact observations are not workspace inspection evidence |
|
|
37
|
+
| `preview.manage` | `approval`, `state_version` | Preview-session state only |
|
|
38
|
+
| `project.detect` | `approval`, `inspect` | Deterministically inspected project root |
|
|
39
|
+
| `project.validate` | `approval`, `state_version`, `validate`, `write` | High-risk repository scripts require approval/containment; generated writes require independent typed inventory |
|
|
40
|
+
| `research.fetch` | `approval`, `research` | Host records bounded, untrusted source evidence durably |
|
|
41
|
+
| `research.search` | `approval`, `research` | Host records bounded, untrusted source discovery durably |
|
|
42
|
+
| `task.checkpoint` | `approval`, `plan`, `state_version` | Bounded durable task state |
|
|
43
|
+
| `user.ask` | `approval` | Correlated answers; no implicit criterion waiver |
|
|
44
|
+
| `workspace.edit` | `approval`, `state_version`, `write` | Exact edit plus before/after hashes |
|
|
45
|
+
| `workspace.glob` | `approval`, `inspect` | Inspected workspace scope/path |
|
|
46
|
+
| `workspace.grep` | `approval`, `inspect` | Inspected matched workspace paths |
|
|
47
|
+
| `workspace.list` | `approval`, `inspect` | Inspected directory scope |
|
|
48
|
+
| `workspace.read` | `approval`, `inspect` | Inspected bounded file/range |
|
|
49
|
+
| `workspace.write` | `approval`, `state_version`, `write` | Atomic create/replace plus before/after hashes |
|
|
50
|
+
|
|
51
|
+
## Important constraints
|
|
52
|
+
|
|
53
|
+
- `approval` is available to every built-in tool because task mode, permission,
|
|
54
|
+
and approval policy can deny any invocation. A successful low-risk call may
|
|
55
|
+
report `not_required`; it must not invent a grant.
|
|
56
|
+
- `write` is evidence of exact workspace mutation, not permission to mutate.
|
|
57
|
+
Approval and OS/workspace policy are evaluated independently first.
|
|
58
|
+
- Mutation evidence compares typed `(kind, hash)` states. Files use a content
|
|
59
|
+
hash; directories and missing paths use `null`; symlinks and special entries
|
|
60
|
+
use a host-defined deterministic hash. The before/after typed states must
|
|
61
|
+
differ. In the legacy kind-less form, create is `null -> hash`, edit is
|
|
62
|
+
`hash -> different hash`, delete is `hash -> null`, and both-null is invalid.
|
|
63
|
+
- A metadata fingerprint for generated or dependency state is not a file
|
|
64
|
+
content hash and must never be emitted as `write` evidence. Hosts may report
|
|
65
|
+
those changes separately as bounded derived mutations and advance
|
|
66
|
+
`state_version`; only byte-verifiable durable paths belong in `write`.
|
|
67
|
+
- A generic command exit code or stdout statement is not validation. Use
|
|
68
|
+
`project.validate` for structured validation evidence.
|
|
69
|
+
- No built-in tool currently has `criterion_satisfy` or `criterion_waive`.
|
|
70
|
+
Those capabilities remain reserved until a dedicated contract can correlate
|
|
71
|
+
a criterion with concrete evidence or an explicit user decision.
|
|
72
|
+
- Extension and MCP tools may define a separate versioned profile. They cannot
|
|
73
|
+
alter the authority of a built-in canonical ID.
|
|
74
|
+
|
|
75
|
+
Runtime initialization rejects a missing policy for an active tool, duplicate
|
|
76
|
+
canonical mappings, unknown capability names, and any built-in capability set
|
|
77
|
+
that differs from the exported profile.
|