clearai-dsh 0.1.3 → 0.1.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +54 -0
- package/README.md +2 -1
- package/README.zh-CN.md +2 -1
- package/lib/client.js +95 -95
- package/lib/fold.js +151 -66
- package/lib/host.js +20 -18
- package/package.json +4 -2
- package/presets/clearai/agent.cordis.yml +107 -36
- package/presets/clearai/plugins/brain.js +2 -2
- package/presets/clearai/plugins/clearai-kernel.js +552 -284
- package/presets/clearai/plugins/commands.js +199 -0
- package/presets/clearai/plugins/ontology.js +3 -3
- package/presets/clearai/plugins/prompts.js +51 -25
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,60 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to this project are recorded here. The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
4
4
|
|
|
5
|
+
## [0.1.4] — 2026-09-16
|
|
6
|
+
|
|
7
|
+
**本版重点:子任务的交付链修好了。** 侦察与世界线执行者的结论此前只进账本、模型读不到
|
|
8
|
+
(账本里也有过「派出去就再也没人收」的挂空)。现在四类子任务(侦察 / 世界线执行者 /
|
|
9
|
+
评估者 / 横评仲裁)统一走原生 `subagents.start()` 的一次性句柄:账本只认本进程攥着的
|
|
10
|
+
`run.result`,结论正文由**收集那一刻的工具返回**交给模型,全文另落
|
|
11
|
+
`clear/knowledge/materials/<id>.md` 供模型、独立评估者与人共读。
|
|
12
|
+
试过的另一条路(拿运行时的结算通知当账本信号)已撤回——它是 best-effort,当不了承重结构。
|
|
13
|
+
|
|
14
|
+
### Added
|
|
15
|
+
|
|
16
|
+
- **ClearAI's own `/` command menu.** `/goal` `/plan` `/evidence` `/worldline` are read-only state windows computed from the ledger on the spot; `/plan-review` re-presents the active plan through the native review card instead of stamping anything itself (commands carry no mutation channel — the authority boundary test pins this).
|
|
17
|
+
- **Native working tools return.** todo, subagent (with model selection), workflow and ralph mount from the standard preset's own rows; the composition suite pins both directions — present: these four; absent: `tool-goal`, `command-goal`, `plan-mode` (the second ledger stays off).
|
|
18
|
+
- **`test/prompt-sections.test.mjs`.** All 23 prompt sections carry a `hard` / `native` / `advisory` class tag, and the suite pins that the classification matches the content (hard sections name a mechanism anchor; native sections name no kernel tool; advisory sections make no mechanism promises).
|
|
19
|
+
|
|
20
|
+
### Changed
|
|
21
|
+
|
|
22
|
+
- **Hypothesis floor is now a hard boundary.** `SetGoal` rejects zero or one hypotheses when `minHypotheses > 0` (kernel default 0 stays neutral; the preset sets 2). Revisions of an existing goal are exempt.
|
|
23
|
+
- **Authorization wording unified to one sentence everywhere.** Kernel messages, the runtime card and the prompts all say: an unapproved plan does not auto-continue; when you advance it explicitly, the first delivery records attribution as it happened (behaviour is authorization). The card says it in human words — ledger field names no longer appear.
|
|
24
|
+
- **Stale native-tool contracts rewritten.** `edit` is literal replacement, not unified diff; `web_search`/`web_fetch` parameter references that no longer exist were removed.
|
|
25
|
+
- **Comment debt cleared to zero.** ~310 comments rewritten to the style rule (why / what breaks / boundary — no dates, no internal section numbers, no incident narratives); the ratchet quotas are now {0, 0, 0}.
|
|
26
|
+
|
|
27
|
+
### Changed
|
|
28
|
+
|
|
29
|
+
- **Async sub-runs now deliver their conclusions through the runtime's own settlement notice.** `SpawnScout`, `MapScouts` and the worldline executors are started with `subagents.startContinuable()`, whose Activation delivers the child's closing message to the parent as a durable user message; the kernel keeps owning only what it must (the `scout/dispatched` / `scout/settled` ledger, the full-text material file under `clear/knowledge/materials/`, and the pointer on the runtime card). Evaluators and the arbitration reviewer stay on the one-shot path: the native durable-child descriptor deliberately omits `outputSchema`, which belongs to a one-shot activation's result contract.
|
|
30
|
+
- **Scout conclusions are persisted in full** to `clear/knowledge/materials/<id>.md` so the model, the independent evaluator and the human read the same copy; the ledger and the runtime card carry a pointer plus a bounded excerpt, and over-long text is marked `…truncated (N chars total, see <path>)` instead of being silently cut. The scout persona now caps its answer at 3000 characters.
|
|
31
|
+
- **A goal's closure records the hypotheses nobody touched.** `goal/closed` carries `unjudged`, the card writes `(untouched)` for a hypothesis with no evidence at all (distinct from "judged inconclusive"), and the loop contract asks for either one touch of evidence or an explicit note about why there was none — no verdict is ever forced.
|
|
32
|
+
|
|
33
|
+
### Fixed
|
|
34
|
+
|
|
35
|
+
- **`install.sh` aborted on macOS (bash 3.2).** It expanded an empty array as `"${OLD_PANEL_PKGS[@]}"` under `set -u`; bash only tolerates that from 4.4 on, while macOS ships 3.2 — so the documented developer install died at step ② for every macOS contributor. CI runs on Linux (bash 5), which is why it never caught it. Both expansions now use the portable `${arr[@]+"${arr[@]}"}` form.
|
|
36
|
+
- **Long-run evidence was overwritten or lost.** Every run now gets its own timestamped archive: the light half (stdout, structured result, provenance, decoded one-line-per-event trajectory, append-only index) is committed, while the heavy half (workspace, raw session log) stays on disk under `~/.dsh/e2e-archive/` so it can be re-judged offline with `tools/e2e-replay.mjs`. `--workspace` pointing inside any git repository is now refused outright: ClearAI commits each delivery into the workspace's own repository, so an in-repo workspace had the kernel commit its delivery snapshots — and the author's uncommitted work — into the host project.
|
|
37
|
+
- **Session-directory name derivation dropped dots.** DSH keeps `.` (and `_`) when it slugs a workspace path; the old rule folded both away, so a workspace under `~/.dsh/…` was reported as "no session log" (38 assertions red in one run). The rule is now taken from a real directory comparison.
|
|
38
|
+
- **Scout conclusions never came back in a scouts-only run.** `AwaitWorldlines` decided whether to keep waiting from the wait-lines produced by `sweepWorldlineExecutors()`, and `sweepScouts()` never produced one — so with only scouts in flight the loop exited on its first tick, even though `SpawnScout`'s own reply tells the model to "wait for it this turn with `AwaitWorldlines`". The one-shot form added a second layer: with no next turn, a conclusion that settled after the last tool call never met another collection point. `sweepScouts()` now reports how many scouts are still unsettled, `collectExecutors()` passes it through, and `AwaitWorldlines` counts it. Verified end to end: the same scenario that stalled twice (goal left open, evaluator refusing `inconclusive`) now finishes 36/36 with the conclusion in the material surface and the goal `achieved`.
|
|
39
|
+
- **Scout conclusions were invisible to the model even after they settled.** They landed only in the mutation record while the section the model reads every step is the runtime card — which had no material surface. Now delivery rides the native settlement notice, the card lists foreign observations (pointer + excerpt) plus the scouts still in flight, and a long-run invariant asserts that an async sub-run's conclusion appears in model-visible text rather than only in the ledger.
|
|
40
|
+
- **Scouts had no "lost" ending.** Worldline executors and evaluators already wrote one; a scout whose child was gone stayed "not yet reported" forever. The judgement now mirrors the evaluator's: in-process dispatches are alive, the native child catalog decides what is still running, and an unavailable read surface means *no judgement* rather than a fabricated one.
|
|
41
|
+
- **The prompt described `SpawnScout` as synchronous.** It said the conclusion comes straight back as the return value and that the tool waits; the kernel is deliberately asynchronous (the blocking wait used to lose the "dispatched" fact when a run was interrupted). The delegation table now teaches the real contract (fire-and-forget, conclusion replays into the material surface, wait with `AwaitWorldlines`), and a drift check pins it.
|
|
42
|
+
- **`AwaitWorldlines` counted mutations, not conclusions.** One scout writes two mutations (`scout/settled` plus the observation), so a single scout was reported as "回灌 2 条". The count and wording now speak of conclusions.
|
|
43
|
+
- **`clearai-commands` cross-plane import.** It imported `ui/lib/fold.js` from the preset plane; in the installed package the relative layout differs, so switching to the preset in a browser failed at import. The command renderers now use the host-provided `clearai` facade for `derive` as well, and the boundary suite pins that no preset plugin imports across planes.
|
|
44
|
+
- **macOS temp-path realpath mismatches.** Session-log lookup and clean-install workspace registration now resolve realpaths (`/var` is a symlink to `/private/var`), which had made e2e logs unfindable and browser session attach fail.
|
|
45
|
+
|
|
46
|
+
### Changed (sub-run lifecycle, unified)
|
|
47
|
+
|
|
48
|
+
- **All four sub-run kinds now share one native lifecycle.** Scout, worldline executor, evaluator and the arbitration reviewer all go through one-shot `subagents.start()` handles; the ledger settles from the `run.result` this process holds, and the conclusion text reaches the model in the tool return of the collecting call. The attempt to use the runtime's settlement notice as a ledger signal is withdrawn: a notice is best-effort, and a live kernel could not reliably see it through either the projection or its own session log. Role differences are now only persona, tool face, and how the result is interpreted — permission and authority boundaries are unchanged.
|
|
49
|
+
- **A scout's conclusion is recorded even when its in-memory entry exists.** The old collection guards skipped exactly the sub-runs the table was holding, so a scout could sit "dispatched, never collected" forever. The sweep now walks every unsettled sub-run in the projection, and de-duplication moved to a session+epoch map that also works after a restart.
|
|
50
|
+
|
|
51
|
+
### Tooling
|
|
52
|
+
|
|
53
|
+
- **Long-run end-to-end scenarios with offline re-judging.** Five scenarios (`worldline-arbitration`, `falsification`, `long-plan`, `scout-first`, `goal-chain`) plus a set of cross-mechanism invariants (no advance without admission, no dangling evaluator, no orphaned fork, promotion level consistency, no dangling scout, evidence bound to real steps, declared artifacts on disk). `tools/e2e-parallel.mjs` runs them concurrently (cap 3, because each run spawns its own worldline executors and evaluators), `tools/e2e-replay.mjs` re-judges a saved session log without spending tokens, and `test/e2e-scenarios.test.mjs` pins every invariant with a negative case so a mis-written judge cannot report a false green.
|
|
54
|
+
|
|
55
|
+
### Validated
|
|
56
|
+
|
|
57
|
+
- **deepseek-flash end-to-end, two headless scenarios** (24/24 plan-and-stop; 25/25 full completion including independent-evaluator settlement) and **one real-browser session** (clean install + Chrome): preset switching, native review-card approval landing `by='user'`, the full thirteen-beat chain, `/goal` rendering, and all four panels drawing — screenshots in `docs/shots/browser-e2e-*.png`.
|
|
58
|
+
|
|
5
59
|
## [0.1.3] — 2026-09-15
|
|
6
60
|
|
|
7
61
|
### Added
|
package/README.md
CHANGED
|
@@ -101,7 +101,7 @@ The plugin contributes three surfaces on top of stock DSH: a **deliverables** vi
|
|
|
101
101
|
|
|
102
102
|
## Where it lands in DSH
|
|
103
103
|
|
|
104
|
-
ClearAI adds an epistemic layer on the DSH **composition surface** — one host package, one agent preset, one client module. The DSH engine is not modified.
|
|
104
|
+
ClearAI adds an epistemic layer on the DSH **composition surface** — one host package, one agent preset, one client module. The DSH engine is not modified. `/goal` `/plan` `/evidence` `/worldline` `/plan-review` are the human's read-only state windows in the `/` menu (computed from the ledger on the spot); todo, subagents, workflows and model switching are DSH-native — working style is unbounded, but none of it can write the authoritative ledger (the authority boundary is pinned by tests).
|
|
105
105
|
|
|
106
106
|

|
|
107
107
|
|
|
@@ -123,6 +123,7 @@ They are illustrations of the mechanism, not shipped run records.
|
|
|
123
123
|
- [Glossary](docs/glossary.md)
|
|
124
124
|
- [Loop philosophy](docs/loop-philosophy.md) · [Verification ontology](docs/verification-loop.md)
|
|
125
125
|
- [Known gaps](docs/known-gaps.md) · [Release verification](docs/release-verification.md)
|
|
126
|
+
- [Convergence and slimming plan](docs/optimization/plan.md) · [Full-coverage design](docs/optimization/epistemic-coverage.md) · [Execution progress](docs/optimization/progress.zh-CN.md)
|
|
126
127
|
|
|
127
128
|
## Work attribution
|
|
128
129
|
|
package/README.zh-CN.md
CHANGED
|
@@ -101,7 +101,7 @@ ClearAI **不**声称实现递归自我改进。它提供的是自我改进系
|
|
|
101
101
|
|
|
102
102
|
## 它落在 DSH 的哪一层
|
|
103
103
|
|
|
104
|
-
ClearAI 把认识论层加在 DSH 的**组合面**上——一个宿主包、一个 agent 预设、一个客户端模块,**DSH
|
|
104
|
+
ClearAI 把认识论层加在 DSH 的**组合面**上——一个宿主包、一个 agent 预设、一个客户端模块,**DSH 引擎一行都没改**。`/goal` `/plan` `/evidence` `/worldline` `/plan-review` 是人在 `/` 菜单里的状态窗(只读,从账本现算);todo、子代理、workflow、模型切换用 DSH 原生的——工作方式不设限,但它们写不进权威账本(权威边界由测试钉死)。
|
|
105
105
|
|
|
106
106
|

|
|
107
107
|
|
|
@@ -123,6 +123,7 @@ ClearAI 把认识论层加在 DSH 的**组合面**上——一个宿主包、一
|
|
|
123
123
|
- [术语表](docs/glossary.zh-CN.md)
|
|
124
124
|
- [循环哲学](docs/loop-philosophy.zh-CN.md) · [验证本体](docs/verification-loop.zh-CN.md)
|
|
125
125
|
- [已知缺口](docs/known-gaps.zh-CN.md) · [发布验收](docs/release-verification.zh-CN.md)
|
|
126
|
+
- [收敛与瘦身计划](docs/optimization/plan.zh-CN.md) · [认识论循环全覆盖设计](docs/optimization/epistemic-coverage.zh-CN.md) · [执行进度](docs/optimization/progress.zh-CN.md)
|
|
126
127
|
|
|
127
128
|
## 工作署名
|
|
128
129
|
|