better-dsh 0.0.0 → 0.2.2-a
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +24 -0
- package/README.md +294 -4
- package/control-prompt.md +37 -0
- package/cordis.patch.yml +53 -0
- package/docs/00_adr/0001-bridge-tool-layer-not-service-layer.md +14 -0
- package/docs/00_adr/0002-masking-is-presentation-only.md +15 -0
- package/docs/10_plans/A2A-messaging-channel-test-archive.md +256 -0
- package/docs/10_plans/code-mode-vs-rlm-ipython-comparison.md +137 -0
- package/docs/10_plans/dashr-blueprint-review.md +201 -0
- package/docs/10_plans/dashr-blueprint.md +561 -0
- package/docs/10_plans/dashr-compaction-window-and-archive.md +307 -0
- package/docs/10_plans/dashr-profile-layer-feasibility.md +367 -0
- package/docs/10_plans/dashr-sandbox-escalation-semantics-gap.md +171 -0
- package/docs/10_plans/dashr-security-sandbox-analysis.md +187 -0
- package/docs/10_plans/dashr-surface-invariant-and-omp-imports.md +97 -0
- package/docs/10_plans/ipython-kernel-interactive-interface-test-report.md +152 -0
- package/docs/10_plans/kernel-refactoring/Dash-IPython-Control-Prompt-draft.md +146 -0
- package/docs/10_plans/kernel-refactoring/Dash-IPython-Control-Prompt-draft_v3.md +50 -0
- package/docs/10_plans/kernel-refactoring/Dash-IPython-Control-Prompt-draft_v4.md +79 -0
- package/docs/10_plans/kernel-refactoring/Dash-vs-PrimeAgent-systemprompt-toolcatalog-comparison.md +138 -0
- package/docs/10_plans/kernel-refactoring/RLM-system-prompt-injection-gap-report.md +161 -0
- package/docs/10_plans/kernel-refactoring/V0.1.5-development-plan.md +109 -0
- package/docs/10_plans/kernel-refactoring/actinoable-surface-to-llm-in-agent-runtime_dsh.md +50 -0
- package/docs/10_plans/kernel-refactoring/actinoable-surface-to-llm-in-agent-runtime_prime.md +113 -0
- package/docs/10_plans/recallable-compaction.md +147 -0
- package/docs/10_plans/spike-tag-repro.mjs +102 -0
- package/docs/10_plans/upstream-analysis.md +128 -0
- package/docs/50_test-reports/REPL-/345/267/245/345/205/267/350/260/203/347/224/250-/346/210/252/346/226/255/350/257/212/346/226/255.md +110 -0
- package/docs/50_test-reports/kernel-provisioning.md +44 -0
- package/docs/50_test-reports/repl-kernel-provisioning-test-report.md +87 -0
- package/docs/50_test-reports/upstream-dsh-0.1.2-alpha.5-local-test-report.md +81 -0
- package/docs/50_test-reports/upstream-dsh-0.1.2-alpha.5-report.md +93 -0
- package/docs/50_test-reports/v0.1.8-improved-/345/256/236/346/265/213/346/212/245/345/221/212.md +142 -0
- package/docs/50_test-reports/v0.1.8-/345/256/236/346/265/213/346/212/245/345/221/212.md +193 -0
- package/docs/50_test-reports/v0.1.8b-/345/256/236/346/265/213/346/212/245/345/221/212.md +96 -0
- package/docs/50_test-reports/v0.1.8c-/345/256/236/346/265/213/346/212/245/345/221/212.md +127 -0
- package/docs/50_test-reports/v0.1.8d-/345/256/236/346/265/213/346/212/245/345/221/212.md +150 -0
- package/docs/50_test-reports/v0.1.8d_artifacts/README.md +138 -0
- package/docs/50_test-reports/v0.1.8d_artifacts/code-mode-repl-only.observation.md +74 -0
- package/docs/50_test-reports/v0.1.8d_artifacts/dsh-session-session-4a293388-9ae1-474b-87a0-9e17bb556d94.jsonl +3890 -0
- package/docs/50_test-reports/v0.1.8d_artifacts/dsh-session-session-4a293388-9ae1-474b-87a0-9e17bb556d94.w-sample-0435.jsonl +544 -0
- package/docs/50_test-reports/v0.1.8d_artifacts/functions.json +592 -0
- package/docs/50_test-reports/v0.1.8d_artifacts/skills-catalog.snapshot.md +30 -0
- package/docs/50_test-reports/v0.1.8d_artifacts/tools-sdk.output-schemas.json +1236 -0
- package/docs/50_test-reports/v0.1.8d_artifacts/tools-sdk.python.txt +592 -0
- package/docs/50_test-reports/v0.1.8d_artifacts/tools-sdk.typescript.txt +516 -0
- package/docs/50_test-reports/v0.1.8d_artifacts/wire-vs-transcription.diff.md +54 -0
- package/docs/50_test-reports/v0.1.8e-/345/256/236/346/265/213/346/212/245/345/221/212.md +224 -0
- package/docs/50_test-reports/v0.1.9a-/345/256/236/346/265/213/346/212/245/345/221/212.md +168 -0
- package/docs/50_test-reports/v0.2.0b-/345/256/236/346/265/213/346/212/245/345/221/212.md +123 -0
- package/docs/50_test-reports/v0.2.0b_artifacts/f2probe/Cargo.lock +7 -0
- package/docs/50_test-reports/v0.2.0b_artifacts/f2probe/Cargo.toml +6 -0
- package/docs/50_test-reports/v0.2.0b_artifacts/f2probe/src/bin/messy.rs +8 -0
- package/docs/50_test-reports/v0.2.0b_artifacts/f2probe/src/main.rs +4 -0
- package/docs/50_test-reports/v0.2.0b_artifacts/hashline-probe.md +5 -0
- package/docs/50_test-reports/v0.2.0b_artifacts/slowprobe/Cargo.lock +7 -0
- package/docs/50_test-reports/v0.2.0b_artifacts/slowprobe/Cargo.toml +7 -0
- package/docs/50_test-reports/v0.2.0b_artifacts/slowprobe/build.rs +4 -0
- package/docs/50_test-reports/v0.2.0b_artifacts/slowprobe/src/main.rs +13 -0
- package/docs/50_test-reports/v0.2.1-/345/256/236/346/265/213/346/212/245/345/221/212.md +110 -0
- package/docs/50_test-reports/v0.2.1b-/345/256/236/346/265/213/346/212/245/345/221/212.md +86 -0
- package/docs/50_test-reports/v0.2.1c-/345/256/236/346/265/213/346/212/245/345/221/212.md +66 -0
- package/docs/50_test-reports/v0.2.1d-/345/256/236/346/265/213/346/212/245/345/221/212.md +67 -0
- package/docs/50_test-reports/v0.2.1e-P1-/345/256/236/346/265/213/346/212/245/345/221/212.md +136 -0
- package/docs/50_test-reports/v0.2.1ef-dev-audit-report.md +73 -0
- package/docs/50_test-reports/v0.2.1f-plugin-shipped-ui-patches/345/256/236/346/265/213/346/212/245/345/221/212.md +102 -0
- package/docs/60_exploration-and-research/cordis-research.md +350 -0
- package/docs/60_exploration-and-research/dsh-web-profile-package-map.md +186 -0
- package/docs/60_exploration-and-research/dsh-web-ui-slot-system-research.md +310 -0
- package/docs/60_exploration-and-research/dsh-webui-strip-boundary-research.md +300 -0
- package/docs/60_exploration-and-research/ios-chat-app-bridge-research.md +324 -0
- package/docs/60_exploration-and-research/web-frontend-composability-research.md +191 -0
- package/docs/REPL-/345/267/245/345/205/267/350/260/203/347/224/250-/346/210/252/346/226/255/350/257/212/346/226/255.md +110 -0
- package/docs/adr/0001-bridge-tool-layer-not-service-layer.md +14 -0
- package/docs/adr/0002-masking-is-presentation-only.md +15 -0
- package/docs/distro-blueprint.md +81 -0
- package/docs/dsh-webUI-with-rlm-mode.png +0 -0
- package/docs/plans/A2A-messaging-channel-test-archive.md +256 -0
- package/docs/plans/code-mode-vs-rlm-ipython-comparison.md +137 -0
- package/docs/plans/dashr-blueprint-review.md +201 -0
- package/docs/plans/dashr-blueprint.md +561 -0
- package/docs/plans/dashr-compaction-window-and-archive.md +307 -0
- package/docs/plans/dashr-profile-layer-feasibility.md +367 -0
- package/docs/plans/dashr-sandbox-escalation-semantics-gap.md +171 -0
- package/docs/plans/dashr-security-sandbox-analysis.md +187 -0
- package/docs/plans/dashr-surface-invariant-and-omp-imports.md +97 -0
- package/docs/plans/ipython-kernel-interactive-interface-test-report.md +152 -0
- package/docs/plans/kernel-refactoring/Dash-IPython-Control-Prompt-draft.md +146 -0
- package/docs/plans/kernel-refactoring/Dash-IPython-Control-Prompt-draft_v3.md +50 -0
- package/docs/plans/kernel-refactoring/Dash-IPython-Control-Prompt-draft_v4.md +79 -0
- package/docs/plans/kernel-refactoring/Dash-vs-PrimeAgent-systemprompt-toolcatalog-comparison.md +138 -0
- package/docs/plans/kernel-refactoring/RLM-system-prompt-injection-gap-report.md +161 -0
- package/docs/plans/kernel-refactoring/V0.1.5-development-plan.md +109 -0
- package/docs/plans/kernel-refactoring/actinoable-surface-to-llm-in-agent-runtime_dsh.md +50 -0
- package/docs/plans/kernel-refactoring/actinoable-surface-to-llm-in-agent-runtime_prime.md +113 -0
- package/docs/plans/recallable-compaction.md +147 -0
- package/docs/plans/spike-tag-repro.mjs +102 -0
- package/docs/plans/upstream-analysis.md +128 -0
- package/docs/repositioning-and-rebranding.md +102 -0
- package/docs/v0.1.8-improved-/345/256/236/346/265/213/346/212/245/345/221/212.md +142 -0
- package/docs/v0.1.8-/345/256/236/346/265/213/346/212/245/345/221/212.md +193 -0
- package/docs/v0.1.8b-/345/256/236/346/265/213/346/212/245/345/221/212.md +96 -0
- package/docs/v0.1.8c-/345/256/236/346/265/213/346/212/245/345/221/212.md +127 -0
- package/docs/v0.1.8d-/345/256/236/346/265/213/346/212/245/345/221/212.md +150 -0
- package/docs/v0.1.8d_artifacts/README.md +138 -0
- package/docs/v0.1.8d_artifacts/code-mode-repl-only.observation.md +74 -0
- package/docs/v0.1.8d_artifacts/dsh-session-session-4a293388-9ae1-474b-87a0-9e17bb556d94.jsonl +3890 -0
- package/docs/v0.1.8d_artifacts/dsh-session-session-4a293388-9ae1-474b-87a0-9e17bb556d94.w-sample-0435.jsonl +544 -0
- package/docs/v0.1.8d_artifacts/functions.json +592 -0
- package/docs/v0.1.8d_artifacts/skills-catalog.snapshot.md +30 -0
- package/docs/v0.1.8d_artifacts/tools-sdk.output-schemas.json +1236 -0
- package/docs/v0.1.8d_artifacts/tools-sdk.python.txt +592 -0
- package/docs/v0.1.8d_artifacts/tools-sdk.typescript.txt +516 -0
- package/docs/v0.1.8d_artifacts/wire-vs-transcription.diff.md +54 -0
- package/docs/v0.1.8e-/345/256/236/346/265/213/346/212/245/345/221/212.md +224 -0
- package/docs/v0.1.9a-/345/256/236/346/265/213/346/212/245/345/221/212.md +168 -0
- package/docs/v0.2.0b-/345/256/236/346/265/213/346/212/245/345/221/212.md +123 -0
- package/docs/v0.2.0b_artifacts/f2probe/Cargo.lock +7 -0
- package/docs/v0.2.0b_artifacts/f2probe/Cargo.toml +6 -0
- package/docs/v0.2.0b_artifacts/f2probe/src/bin/messy.rs +8 -0
- package/docs/v0.2.0b_artifacts/f2probe/src/main.rs +4 -0
- package/docs/v0.2.0b_artifacts/hashline-probe.md +5 -0
- package/docs/v0.2.0b_artifacts/slowprobe/Cargo.lock +7 -0
- package/docs/v0.2.0b_artifacts/slowprobe/Cargo.toml +7 -0
- package/docs/v0.2.0b_artifacts/slowprobe/build.rs +4 -0
- package/docs/v0.2.0b_artifacts/slowprobe/src/main.rs +13 -0
- package/docs/v0.2.1-/345/256/236/346/265/213/346/212/245/345/221/212.md +110 -0
- package/docs/v0.2.1b-/345/256/236/346/265/213/346/212/245/345/221/212.md +86 -0
- package/docs/v0.2.1c-/345/256/236/346/265/213/346/212/245/345/221/212.md +66 -0
- package/lib/client/index.js +473 -0
- package/lib/index.d.ts +736 -0
- package/lib/index.js +11518 -0
- package/lib/kernel-env-hxaihi9C.js +195 -0
- package/lib/kernel-env.d.ts +80 -0
- package/lib/kernel-env.js +3 -0
- package/lib/py-sdk-BCaOGYz7.d.ts +125 -0
- package/lib/py-sdk-CbgYiX8O.js +691 -0
- package/lib/py-sdk.d.ts +2 -0
- package/lib/py-sdk.js +3 -0
- package/package.json +325 -4
- package/scripts/kernel-provision.mjs +35 -0
- package/index.js +0 -3
|
@@ -0,0 +1,79 @@
|
|
|
1
|
+
## The Dashr IPython interface
|
|
2
|
+
|
|
3
|
+
The IPython kernel is the one actionable interface of your agent runtime — a
|
|
4
|
+
persistent control environment for reasoning, context management, state, tool
|
|
5
|
+
orchestration, and recursive subcalls. The ONLY function call the runtime
|
|
6
|
+
accepts directly is `ipython`, whose schema is:
|
|
7
|
+
|
|
8
|
+
{ "name": "ipython", "arguments": { "cell": "<one IPython cell>", "description": "<description>" } }
|
|
9
|
+
|
|
10
|
+
Every action you take is one program you write inside `cell` — one call, one
|
|
11
|
+
cell, executed top-to-bottom. Every other tool is called from inside a cell,
|
|
12
|
+
as `await name(args)` (see below).
|
|
13
|
+
|
|
14
|
+
Use the kernel to keep intermediate variables, inspect and transform outputs,
|
|
15
|
+
write small helper functions, and preserve useful state across turns or
|
|
16
|
+
compaction.
|
|
17
|
+
|
|
18
|
+
## Tools as Functions in REPL
|
|
19
|
+
|
|
20
|
+
Every tool in the Tool Catalog is a flat top-level function in IPython kernel: await name(args) with ONE positional arguments object — `await read({"file_path": "x"})`, never `read(file_path="x")`.
|
|
21
|
+
|
|
22
|
+
One program in one cell — write it, Shift+Enter (one `ipython` call), get the result back (print or final expression), like a REPL scratchpad.
|
|
23
|
+
|
|
24
|
+
```
|
|
25
|
+
┌── cell ─ "Read the project README" ───────────────────-──┐
|
|
26
|
+
│ │
|
|
27
|
+
│ # One step cell │
|
|
28
|
+
│ await read({"file_path": "docs/README.md"}) │
|
|
29
|
+
│ │
|
|
30
|
+
└──────────────────────────────────────────────────────────┘
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
To execute it via the only `ipython` entrance:
|
|
34
|
+
|
|
35
|
+
{ "name": "ipython", "arguments": { "cell": "await read({\"file_path\": \"docs/README.md\"})", "description": "Read the project README" } }
|
|
36
|
+
|
|
37
|
+
Never call tools directly as a function-call like:
|
|
38
|
+
|
|
39
|
+
❌ ~~{ "name": "read", "arguments": { "file_path": "docs/README.md", "description": "Read the project README" } }~~
|
|
40
|
+
|
|
41
|
+
|
|
42
|
+
## Examples
|
|
43
|
+
|
|
44
|
+
```python
|
|
45
|
+
# One step cell
|
|
46
|
+
print(await read({"file_path": "docs/README.md"}))
|
|
47
|
+
|
|
48
|
+
# shell is just another typed callable
|
|
49
|
+
r = await bash({"command": "ls -la src/", "description": "List source directory"})
|
|
50
|
+
print(r["stdout"]["text"])
|
|
51
|
+
|
|
52
|
+
# typed tools in script
|
|
53
|
+
for old in ("DEBUG = False", "DEBUG=False"):
|
|
54
|
+
try:
|
|
55
|
+
await edit({"file_path": "src/config.py", "old_string": old, "new_string": "DEBUG = True"})
|
|
56
|
+
break
|
|
57
|
+
except ToolCallError as e:
|
|
58
|
+
print(f"retrying ({e})")
|
|
59
|
+
|
|
60
|
+
# import, fan-out with gather
|
|
61
|
+
import asyncio
|
|
62
|
+
matches, files = await asyncio.gather(
|
|
63
|
+
grep({"pattern": "TODO", "path": "src"}),
|
|
64
|
+
file_glob({"pattern": "**/*.ts", "path": "src"}),
|
|
65
|
+
)
|
|
66
|
+
|
|
67
|
+
# variables persist across cells and turns — the kernel is your working memory
|
|
68
|
+
cfg = await read({"file_path": "config.yaml"}) # cfg stays alive in later cells
|
|
69
|
+
child = await rlm({"mode": "spawn", "prompt": "summarize the failing tests", "label": "summarizer"})
|
|
70
|
+
print(child["subagentId"]) # background admission — keep working
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
## Rules
|
|
74
|
+
|
|
75
|
+
- Do not assume IPython is the native runtime of the external thing being investigated. Evaluate external systems through their own interface, then use IPython to coordinate the process and analyze what comes back.
|
|
76
|
+
- Only print or return what you need next; everything else stays in the kernel.
|
|
77
|
+
- Variables persist across cells and turns, but they live in the kernel
|
|
78
|
+
subprocess: keep durable state in files or the Continual Harness (refine
|
|
79
|
+
writes it).
|
package/docs/10_plans/kernel-refactoring/Dash-vs-PrimeAgent-systemprompt-toolcatalog-comparison.md
ADDED
|
@@ -0,0 +1,138 @@
|
|
|
1
|
+
# Dash RLM vs Prime Agent:System Prompt 与 Toolset Catalog 对比
|
|
2
|
+
|
|
3
|
+
> 整理日期(UTC):2026-08-18T06:11:40Z
|
|
4
|
+
> 目的:把两边「LLM 实际看到的文本」摆在一起,供拍板调整方向。
|
|
5
|
+
> 口径:**System Prompt** = LLM 收到的关于「界面/范式」的说明文本;**Toolset Catalog** = LLM 运行时能看到的工具描述文本(≠ schema,schema 是代码级合约,catalog 是描述文本)。
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## 1. System Prompt 对比
|
|
10
|
+
|
|
11
|
+
### 1.1 Dash 现有(RLM 相关的全部文本,共三段)
|
|
12
|
+
|
|
13
|
+
**① 引言(唯一一句「范式」说明)**
|
|
14
|
+
|
|
15
|
+
> You are a DASHR agent powered by the deepseek-v4-pro model. Your working directory is /home/u1/workspaces/dashr. You have a persistent Python kernel as your **primary** tool interface: each `run_cell` call executes one Python program on it, and everything you import, define, or assign in one cell stays available in the next. The full tool catalog is callable from inside the kernel as `tools.<name>(...)` bindings. Prefer the kernel for multi-step work.
|
|
16
|
+
|
|
17
|
+
**② `run_cell` 工具的 description(model 直接看到的唯一工具)**
|
|
18
|
+
|
|
19
|
+
> Execute one Python cell on the persistent kernel. Takes two required arguments: `code`, the cell (top-level `await` and `return` work; variables, imports, and definitions from earlier cells are still alive), and `description`, a short summary of what the cell does. Call tools as `await tools.name(args)` per the declarations in the system prompt. Only what you print or return comes back — curate it.
|
|
20
|
+
|
|
21
|
+
**③ `## Writing cells for run_cell`(run_cell 用法 prose)**
|
|
22
|
+
|
|
23
|
+
> `run_cell` takes two required arguments: `code` — one Python cell — and `description` … The cell runs on a PERSISTENT IPython kernel: variables, imports, and definitions created in any earlier `run_cell` call of this session are still alive in later ones … treat the kernel's namespace as your working memory. Top-level `await` and `return` both work. At run time exactly two of the names declared below are bound: `tools` and `ToolCallError`. Everything else is a STATIC STUB … Inside a cell:
|
|
24
|
+
> - Call tools as `await tools.name(args)` …
|
|
25
|
+
> - A FAILED tool call raises `ToolCallError` …
|
|
26
|
+
> - Independent calls may overlap under `asyncio.gather` …
|
|
27
|
+
> - Emit the answer with `print(...)` and/or a top-level `return <value>`. … ONLY what you print and the returned value come back …
|
|
28
|
+
|
|
29
|
+
> 小结:三段是「接口如何使用」的说明(工具化口吻),**没有一句「这是唯一界面 / 你是用代码解决问题的 agent / code-first 范式」**,也没有示例。
|
|
30
|
+
|
|
31
|
+
### 1.2 Prime Agent(`buildRlmPrompt` 渲染,root 会话)
|
|
32
|
+
|
|
33
|
+
**开篇(范式,3 句)**
|
|
34
|
+
|
|
35
|
+
> You are a general purpose agent that uses code to solve tasks.
|
|
36
|
+
> You solve tasks by breaking down problems into sub-tasks, writing and executing code, observing results, and iterating one step at a time.
|
|
37
|
+
> When you are done, stop calling tools and state your final answer.
|
|
38
|
+
>
|
|
39
|
+
> Working directory: <cwd>
|
|
40
|
+
> Conversation log: <messagesPath>
|
|
41
|
+
> Recursive agent depth: 0
|
|
42
|
+
> Pre-installed Python packages: …
|
|
43
|
+
|
|
44
|
+
**`IPYTHON_CONTROL_PROMPT`(整段,核心)**
|
|
45
|
+
|
|
46
|
+
> IPython is the agent's long-lived notebook: a persistent control environment for reasoning, context management, state, tool orchestration, and recursive subcalls. Use it to keep intermediate variables, inspect and transform outputs, write small helper functions, and preserve useful state across turns or compaction.
|
|
47
|
+
>
|
|
48
|
+
> Do not assume IPython is the native runtime of the external thing being investigated. … Evaluate external systems through their own interface, then use IPython to coordinate the process and analyze what comes back.
|
|
49
|
+
>
|
|
50
|
+
> When running shell commands from IPython, use `%%bash` cells. If you use `%%bash`, it must be the first line of the code cell … Avoid `!cmd` shell escapes …
|
|
51
|
+
>
|
|
52
|
+
> Important: do not install dependencies into the IPython kernel just to make an external project import or run there. … run it through that project's own environment …
|
|
53
|
+
>
|
|
54
|
+
> Use Python for reading, searching, and editing files … Always assign read/search results to named variables …
|
|
55
|
+
>
|
|
56
|
+
> Each `%%bash` cell runs in a throw-away subshell … use `%cd <dir>` … or `os.environ['VAR'] = '...'` …
|
|
57
|
+
>
|
|
58
|
+
> Python state in the kernel, by contrast, persists across cells … Tool calls are themselves Python `await` expressions, so their return values can be bound to variables and composed into program logic …
|
|
59
|
+
>
|
|
60
|
+
> Continual harness state is available as `rlm.harness` and `rlm.get_harness_state()` …
|
|
61
|
+
>
|
|
62
|
+
> RLM-native call contract: installed Python skills are pre-imported modules. Read the matching SKILL.md and call its documented function, such as `await <skill_import>.<function>(...)` … **Do not invent non-native wrappers such as `call_skill(...)` or `run_subagent(...)`.**
|
|
63
|
+
|
|
64
|
+
**递归 + skill + 消息(节选)**
|
|
65
|
+
|
|
66
|
+
> A callable `rlm` is already in your global namespace. `await rlm('sub-task')` spawns a child and returns immediately after task admission …
|
|
67
|
+
> Children reply explicitly with `await agent_message.send(message, receiver_role='parent')` …
|
|
68
|
+
> Agent messaging is restricted to your parent, siblings, and direct children …
|
|
69
|
+
|
|
70
|
+
> 小结:Prime Agent 是**范式驱动**——先立「code-first」人设,再用一整段 prose 讲清 IPython 是唯一长驻环境、`%%bash`、状态持久化、`await` 工具调用,并给示例 + 明令「不要自造非原生 wrapper」。
|
|
71
|
+
|
|
72
|
+
### 1.3 差异小结
|
|
73
|
+
|
|
74
|
+
| 维度 | Dash 现有 | Prime Agent |
|
|
75
|
+
|---|---|---|
|
|
76
|
+
| 范式定位 | 无(只说 kernel 是 *primary* interface) | 明确「code-first agent」 |
|
|
77
|
+
| 「唯一界面」强调 | 无 | IPython 是长驻控制环境,反复强调 |
|
|
78
|
+
| 示例 | 无 | 有(`await rlm(...)` / `agent_message.send(...)` / `edit(...)`) |
|
|
79
|
+
| 反模式告诫 | 无 | 「Do not invent non-native wrappers」 |
|
|
80
|
+
| 口吻 | 工具化(接口说明) | 范式化(人设 + 工作法) |
|
|
81
|
+
|
|
82
|
+
---
|
|
83
|
+
|
|
84
|
+
## 2. Toolset Catalog 对比
|
|
85
|
+
|
|
86
|
+
### 2.1 DASH 原生(约 29 个工具,DSH 原生语义)
|
|
87
|
+
|
|
88
|
+
模型在 system prompt 的 `The available tools:`(即 `tools:dashr-sdk` 段)里看到的是:把 DSH registry 的原生工具**渲染成 Python SDK**(`class Tools(Protocol)` + 每个工具的 TypedDict 参数 + docstring 描述)。工具名与描述是 **DSH 原生词汇**,未为 kernel 范式重写:
|
|
89
|
+
|
|
90
|
+
- 文件:`read` `write` `edit` `glob` `grep` `read_image`
|
|
91
|
+
- Shell/任务:`bash` `job_output` `job_list` `job_kill`
|
|
92
|
+
- Agent 编排:`subagent` `subagent_fork` `send_message` `list_agents` `interrupt_agent` `workflow` `ralph`
|
|
93
|
+
- 目标/记忆:`create_goal` `get_goal` `update_goal` `memory_add` `memory_search` `memory_list` `memory_flush`
|
|
94
|
+
- 其他:`todo_write` `web_search` `ask_user_question` `skill` `exit_plan_mode`
|
|
95
|
+
|
|
96
|
+
关键点:**这 29 个工具里,模型能直接调用的只有 `run_cell` 一个**(`system-prompt/assemble` 把 `assembly.tools` filter 到只剩 `run_cell`;直接调其他工具会被 guard 拒绝并提示「only run_cell is callable directly」)。于是出现「catalog 列了 29 个工具 → 实际只能调 run_cell」的错位。
|
|
97
|
+
|
|
98
|
+
### 2.2 Prime Agent(1 个 core tool + 13 个 skills)
|
|
99
|
+
|
|
100
|
+
- **core tool 只有 `ipython` 一个**(`allToolNames = Set(["ipython"])`):
|
|
101
|
+
> Python scratchpad code or `%%bash` shell cells to execute in the agent kernel. Use the target project's own environment for project imports, tests, scripts, CLIs, and dependency checks instead of direct kernel imports.
|
|
102
|
+
|
|
103
|
+
- **skills(在 cell 内以 `await <skill>.<fn>(...)` 调用)**,每个 skill 的 SKILL.md description 就是它的 catalog 文本:
|
|
104
|
+
|
|
105
|
+
| skill | description(首行) |
|
|
106
|
+
|---|---|
|
|
107
|
+
| agent-message | Message an agent's parent, siblings, or direct children through the daemon |
|
|
108
|
+
| agent-observe | Read-only observation of an agent's parent, siblings, and direct children |
|
|
109
|
+
| edit | Replace an exact, unique string in an existing file |
|
|
110
|
+
| attach-image | Load an on-disk image into the model's context |
|
|
111
|
+
| compact | Check context usage and compact the conversation from IPython |
|
|
112
|
+
| goal | Manage the persistent thread goal from IPython |
|
|
113
|
+
| refine | Trigger continual harness refinement from IPython |
|
|
114
|
+
| websearch | Search Google via the Serper API |
|
|
115
|
+
| rlm-heartbeat | Manage agent-owned RLM heartbeats from IPython |
|
|
116
|
+
| skill-creator | Create, validate, and install Prime Agent skills |
|
|
117
|
+
| linear / notion / prime-intellect | MCP/CLI 集成(外部服务) |
|
|
118
|
+
|
|
119
|
+
关键点:catalog 里**没有把「bash/read/write/edit/glob/grep」等底层能力各自做成独立原生工具**——shell 走 `%%bash`、文件编辑走 `edit` skill、读取/搜索走「用 Python」的 prose 指导。工具集是**为 kernel 范式量身定做的**,且每个 skill 的 description 都写明「from IPython」。
|
|
120
|
+
|
|
121
|
+
### 2.3 差异小结
|
|
122
|
+
|
|
123
|
+
| 维度 | DASH 原生 | Prime Agent |
|
|
124
|
+
|---|---|---|
|
|
125
|
+
| 工具数量 | ~29 个原生工具 | 1 core(ipython)+ ~13 skills |
|
|
126
|
+
| 工具词汇来源 | DSH 原生(bash/read/write/subagent/…) | 为 kernel 范式重写(ipython + skills) |
|
|
127
|
+
| 呈现位置 | 塞在 system prompt 的 Python SDK 文本块 | tool schema 在 prompt body 之外;skill 用 SKILL.md |
|
|
128
|
+
| 与「唯一界面」的一致性 | 不一致(列 29 个却只能调 run_cell) | 一致(core=ipython,其余都是 cell 内调用) |
|
|
129
|
+
|
|
130
|
+
---
|
|
131
|
+
|
|
132
|
+
## 3. 附:源码位置
|
|
133
|
+
|
|
134
|
+
- Dash 插件:`dashr/dashr/src/index.ts`(`run_cell` 定义 `RUN_CELL_DESCRIPTION`、`tools:dashr-sdk` section、schema collapse、guard)
|
|
135
|
+
- Dash SDK 渲染:`dashr/dashr/src/py-sdk.ts`(`renderToolsSdkPy`)
|
|
136
|
+
- Prime Agent prompt:`upstream/prime-agent/packages/coding-agent/src/core/prompts/rlm.ts`(`buildRlmPrompt`、`IPYTHON_CONTROL_PROMPT`)
|
|
137
|
+
- Prime Agent tools:`upstream/prime-agent/packages/coding-agent/src/core/tools/index.ts`(`allToolNames`)
|
|
138
|
+
- Prime Agent skills:`upstream/prime-agent/packages/coding-agent/skills/*/SKILL.md`
|
|
@@ -0,0 +1,161 @@
|
|
|
1
|
+
# Dash RLM 插件「system prompt 注入缺口」实测报告
|
|
2
|
+
|
|
3
|
+
> 实测日期(UTC):2026-08-18T06:02:38Z
|
|
4
|
+
> 观测主体:Dash Agent 运行时自身(`agent-preset/selected = rlm-mode`)
|
|
5
|
+
> 对照对象:上游 Prime Agent(`upstream/prime-agent` + 已安装实例 `~/.prime/agent`)
|
|
6
|
+
> 方式:读自身 session 日志 + 读上游源码,所有结论均来自真实文件与真实工具返回。
|
|
7
|
+
>
|
|
8
|
+
> 版本说明:v2 —— 新增「三大缺口」明确清单;A2A 消息缺口按用户判定降级为「非致命,暂缓」。
|
|
9
|
+
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
## 1. 问题缘起
|
|
13
|
+
|
|
14
|
+
用户观察到:本 agent 会话刚开始时,对 `run_cell`(Python kernel cell)机制显得陌生,第一反应是把 `bash` / `todo_write` / `list_agents` 当普通函数工具直接调用,而不是放进 `run_cell` 执行。
|
|
15
|
+
|
|
16
|
+
核心判断:**插件已经把「工具调用界面」和「互动界面」全部归一化成了一个 `run_cell`,却没有提供足够的上下文告诉运行时该如何使用。** 这才是致命问题(相较之下,A2A 消息缺口顶多像 one-shot calling,非致命)。
|
|
17
|
+
|
|
18
|
+
---
|
|
19
|
+
|
|
20
|
+
## 2. 观测一:我自身的行为(两次直接调用工具失败)
|
|
21
|
+
|
|
22
|
+
会话历史里,本 agent 至少有两次把工具当"原生函数调用"直接发起,均被运行时拒绝:
|
|
23
|
+
|
|
24
|
+
1. **第 1 轮**:直接调用 `todo_write` / `bash` / `list_agents` → 报错
|
|
25
|
+
`Error: only run_cell is callable directly — call <tool> from inside a run_cell program instead`
|
|
26
|
+
2. **本轮(形成本报告时)**:再次直接调用 `todo_write` → 同样报错。
|
|
27
|
+
|
|
28
|
+
结论:模型默认心智是"工具 = 原生函数调用",而运行时唯一的原生工具是 `run_cell`。二者不一致,且没有 system prompt 在纠偏。
|
|
29
|
+
|
|
30
|
+
---
|
|
31
|
+
|
|
32
|
+
## 3. 观测二:Dash 侧 system prompt 的缺口
|
|
33
|
+
|
|
34
|
+
### 3.1 system prompt 从哪注入
|
|
35
|
+
|
|
36
|
+
本 session 日志 `session-d11e193f-....jsonl.zstd` 首部:
|
|
37
|
+
|
|
38
|
+
- `type=agent-preset/selected` → `{"agentPreset": "rlm-mode"}`
|
|
39
|
+
- `type=request/header` → `data.header.system` 为完整 system prompt
|
|
40
|
+
|
|
41
|
+
### 3.2 规模与结构
|
|
42
|
+
|
|
43
|
+
- system prompt 全长 **45,578 字符 / 684 行**。
|
|
44
|
+
- 结构(字符偏移):
|
|
45
|
+
- `0` — "You are an AI agent powered by DeepSeek Harness"(DSH 基础层)
|
|
46
|
+
- `1393` — "You are a DASHR agent powered by the deepseek-v4-pro model"(DASHR/RLM 层)
|
|
47
|
+
- `8817` — `## Writing cells for run_cell`(run_cell 用法说明)
|
|
48
|
+
- `10748` — `The available tools:`(约 20KB 的 Python 代码块,`class Tools(Protocol)` + 30 个工具 TypedDict)
|
|
49
|
+
- `29589` — `class Tools(Protocol)`
|
|
50
|
+
- `29918` — `async def bash`(第一个工具方法)
|
|
51
|
+
|
|
52
|
+
### 3.3 缺失的指令(检索证据)
|
|
53
|
+
|
|
54
|
+
全文关键词检索,以下纠偏型指令 **全部 0 命中**:
|
|
55
|
+
|
|
56
|
+
| 检索短语 | 命中 |
|
|
57
|
+
|---|---|
|
|
58
|
+
| `run_cell is the only` | 0 |
|
|
59
|
+
| `cannot call ... directly` | 0 |
|
|
60
|
+
| `callable directly` | 0 |
|
|
61
|
+
| `must be called from inside` | 0 |
|
|
62
|
+
| `do not call` | 0 |
|
|
63
|
+
| `must go inside` | 0 |
|
|
64
|
+
|
|
65
|
+
唯一接近的一句(`1393` 处):
|
|
66
|
+
|
|
67
|
+
> "You have a persistent Python kernel as your **primary** tool interface … The full tool catalog is callable from inside the kernel as `tools.<name>(...)` bindings."
|
|
68
|
+
|
|
69
|
+
措辞是 **"primary"(主要)而非 "only"(唯一)**,且紧接着用 `The available tools:` 把 30 个工具列得像**可直接调用的原生函数工具**,与"只有 run_cell 能直接调"的运行时事实相矛盾。
|
|
70
|
+
|
|
71
|
+
---
|
|
72
|
+
|
|
73
|
+
## 4. 三大缺口(明确清单)
|
|
74
|
+
|
|
75
|
+
### 缺口 1:框架描述的缺失 —— 缺少「唯一界面」的框架定位
|
|
76
|
+
|
|
77
|
+
- **现状**:Dash 插件只有极简短的一句提醒(`persistent Python kernel as your primary tool interface`),其余工具仍沿用 DSH(DeepSeek Harness)原生 schema 描述。整个 prompt 没有把「这是一个 Python kernel、且是**唯一**界面」讲清楚,也没有给出清晰的示例。
|
|
78
|
+
- **对照 Prime Agent**:`buildRlmPrompt()` 明确 `You are a general purpose agent that uses code to solve tasks`,并用一整段 `IPYTHON_CONTROL_PROMPT` 讲清 IPython 是 long-lived notebook、`%%bash` cell、状态跨 cell 持久化、工具调用是 cell 内的 `await` 表达式,同时给出具体示例(`await rlm('sub-task')`、`await agent_message.send(...)`、`await edit(...)`)。
|
|
79
|
+
- **建议**:明确指出「这是 Python kernel,作为唯一界面」,并给出清晰示例。
|
|
80
|
+
|
|
81
|
+
### 缺口 2:工具 schema 的呈现 —— 工具目录 vs 范式
|
|
82
|
+
|
|
83
|
+
- **现状**:Dash 的 system prompt 用 `The available tools:` 列了一份约 20KB 的 `Tools(Protocol)`,30 个工具以 DSH 原生 schema 呈现(TypedDict 参数 + `async def` 方法),读起来像可直接调用的原生函数工具——与实际「唯一工具是 run_cell」的事实冲突。
|
|
84
|
+
- **对照 Prime Agent**:整个框架按 Python kernel 编写,tool calling 的 catalog 按「在 cell 内直接可用的 schema」呈现(`await <module>.<fn>(...)` 形态),且工具 schema 放在 prompt body 之外(`system-prompt.ts` 第 12 行注释:*Tool schemas carry tool descriptions outside the prompt body*)。
|
|
85
|
+
- **待确认**:是否也应采用 Prime Agent 的标准——工具目录按「cell 内可直接调用」的形态呈现,而非 DSH 原生函数 schema。
|
|
86
|
+
|
|
87
|
+
### 缺口 3:原生 tool schema 的转译 —— 无法硬编码归一化
|
|
88
|
+
|
|
89
|
+
- **现实约束**:DASH 核心会按自身机制加载工具,从而产生**原生可执行的 schema**。这些工具描述是语义层 + 文本层的,很难做硬编码的转译或归一化。
|
|
90
|
+
- **倾向方案(用户)**:不做插件侧硬编码转译,而是在 System Prompt 中强调——
|
|
91
|
+
1. 你现在面对的是**唯一界面**;
|
|
92
|
+
2. 工具**已经过转译处理**;
|
|
93
|
+
3. 如果看到原生工具的 schema,运行时需**自行将其转化为 run_cell 可接受的格式**;
|
|
94
|
+
4. **只要格式不差太远,放进 run_cell 里执行即可**。
|
|
95
|
+
|
|
96
|
+
---
|
|
97
|
+
|
|
98
|
+
## 5. 对比:上游 Prime Agent 怎么做
|
|
99
|
+
|
|
100
|
+
### 5.1 system prompt 是"范式驱动"的散文,而非"工具目录"
|
|
101
|
+
|
|
102
|
+
`upstream/prime-agent/packages/coding-agent/src/core/prompts/rlm.ts` 的 `buildRlmPrompt()` 开篇:
|
|
103
|
+
|
|
104
|
+
> "You are a general purpose agent that uses code to solve tasks."
|
|
105
|
+
> "You solve tasks by breaking down problems into sub-tasks, writing and executing code, observing results, and iterating one step at a time."
|
|
106
|
+
> "When you are done, stop calling tools and state your final answer."
|
|
107
|
+
|
|
108
|
+
随后用一长段 `IPYTHON_CONTROL_PROMPT` 明确 IPython 是"long-lived notebook / persistent control environment",并显式约定:
|
|
109
|
+
|
|
110
|
+
- shell 用 `%%bash` cell(必须是 cell 第一行);
|
|
111
|
+
- Python 状态跨 cell 持久化;工具调用是 cell 里的 `await` 表达式,返回值可绑定变量;
|
|
112
|
+
- "**Do not invent non-native wrappers such as `call_skill(...)` or `run_subagent(...)`**"。
|
|
113
|
+
|
|
114
|
+
### 5.2 工具 schema 与 prompt 分离
|
|
115
|
+
|
|
116
|
+
`core/system-prompt.ts` 第 12 行注释:
|
|
117
|
+
|
|
118
|
+
> "Active tools. **Tool schemas carry tool descriptions outside the prompt body.**"
|
|
119
|
+
|
|
120
|
+
工具能力描述放在 function-calling 的 schema 里,**不在 prompt 正文里重复列一份工具目录**;prompt 正文只讲范式。
|
|
121
|
+
|
|
122
|
+
### 5.3 system prompt 不落盘、按请求动态构建
|
|
123
|
+
|
|
124
|
+
Prime Agent 的 session 文件(`~/.prime/agent/sessions/*.jsonl`)里**没有 system message**,只存 `user` / `assistant` / `toolResult`;system prompt 由 `buildSystemPrompt()` 每次请求动态拼装。反观 Dash 是把 45KB system prompt 固化进 `request/header`。
|
|
125
|
+
|
|
126
|
+
---
|
|
127
|
+
|
|
128
|
+
## 6. A2A 消息拓扑(非致命,暂缓)
|
|
129
|
+
|
|
130
|
+
> 用户判定:A2A 消息缺口不是致命的,顶多像 one-shot calling。本节仅存档,不阻塞主线。
|
|
131
|
+
|
|
132
|
+
1. 上游明确写了消息拓扑(`rlm.ts`):
|
|
133
|
+
> "Agent messaging is restricted to your **parent, siblings, and direct children**; roots are siblings, and deeper communication relays through the intermediate child."
|
|
134
|
+
2. 上游 `agent_message` 支持三向(`skills/agent-message/SKILL.md`):`receiver_role="parent" | "sibling" | "child"`。
|
|
135
|
+
3. Dash 实测更严格:`send_message` 只允许「父 → 直接子」,sibling/向上/跨层被拒(`subagent "<id>" belongs to another parent session`)。
|
|
136
|
+
|
|
137
|
+
---
|
|
138
|
+
|
|
139
|
+
## 7. 结论与建议
|
|
140
|
+
|
|
141
|
+
1. **致命问题确认**:插件把工具调用界面与互动界面归一化成单一 `run_cell`,却未注入足够上下文。模型开局两次直接调用工具失败即是证据。
|
|
142
|
+
2. **三大缺口**对应处理:
|
|
143
|
+
- **框架描述**:明确「Python kernel = 唯一界面」+ 清晰示例;
|
|
144
|
+
- **工具 schema 呈现**:待确认是否改按「cell 内可直接调用」形态呈现(对齐 Prime Agent);
|
|
145
|
+
- **原生 schema 转译**:不硬编码转译,在 System Prompt 中强调「唯一界面 + 已转译 + 运行时自行把原生 schema 转成 run_cell 格式,格式接近即可放进 cell 执行」。
|
|
146
|
+
3. **A2A 消息**:非致命,暂缓。
|
|
147
|
+
4. **建议注入的最小补丁**(置于 system prompt 靠前):
|
|
148
|
+
- 「本 session 运行在 RLM 模式:`run_cell` 是**唯一**能直接调用的工具。」
|
|
149
|
+
- 「不要直接调用 `bash`/`write`/`read`/`edit`/`glob`/`grep`/`todo_write`/`list_agents`/`send_message`/`subagent` 等——它们会报 `only run_cell is callable directly`。一律写成 `await tools.<name>(...)` 放进 cell。」
|
|
150
|
+
- 「你是用代码解决问题的 agent:工具调用 = cell 内的 `await` 表达式,返回值可绑定变量、可组合。」
|
|
151
|
+
- 「(转译)若看到原生工具 schema,请自行转化为 run_cell 可接受的格式;格式接近即可放入 cell 执行。」
|
|
152
|
+
|
|
153
|
+
---
|
|
154
|
+
|
|
155
|
+
## 8. 证据清单
|
|
156
|
+
|
|
157
|
+
- 自身 session 日志:`/home/u1/.dsh/sessions/--home-u1-workspaces-dashr--/session-d11e193f-a813-4200-96b1-0465ae8e8063/session.jsonl.zstd`(`request/header.header.system`,45,578 字符)
|
|
158
|
+
- 上游 prompt:`upstream/prime-agent/packages/coding-agent/src/core/prompts/rlm.ts`(`buildRlmPrompt`、`IPYTHON_CONTROL_PROMPT`、`buildChildAgentDoctrine`)
|
|
159
|
+
- 上游装配:`upstream/prime-agent/packages/coding-agent/src/core/system-prompt.ts`(第 12 行注释:tool schemas 在 prompt body 之外)
|
|
160
|
+
- 上游消息 skill:`upstream/prime-agent/packages/coding-agent/skills/agent-message/SKILL.md`
|
|
161
|
+
- Prime Agent 实例:`~/.prime/agent/sessions/*.jsonl`(无 system message,仅 user/assistant/toolResult)
|
|
@@ -0,0 +1,109 @@
|
|
|
1
|
+
# DASHR V0.1.5 开发计划
|
|
2
|
+
|
|
3
|
+
> 产出方式:grilling interview(2026-08-18/19),全部决策经用户逐条确认。
|
|
4
|
+
> 基线版本:0.1.4 → 目标 0.1.5。
|
|
5
|
+
> 版本说明:原计划为 V0.2(目标 0.2.0);实现完成复核后发现本批改动未达到架构/范式级更新,
|
|
6
|
+
> 故降级为补丁版本 V0.1.5 发布。
|
|
7
|
+
> 配套术语表:仓库根 CONTEXT.md。相关 ADR:docs/adr/0001(桥接工具层)、docs/adr/0002(masking 纯呈现层)。
|
|
8
|
+
|
|
9
|
+
---
|
|
10
|
+
|
|
11
|
+
## 1. 决策记录(interview 定案)
|
|
12
|
+
|
|
13
|
+
| # | 主题 | 定案 |
|
|
14
|
+
|---|------|------|
|
|
15
|
+
| Q1 | 实现顺序 | Architecture → A2A+RLM → System Prompt(最后写 prompt,因为它依赖 A2A/RLM 的实际形态) |
|
|
16
|
+
| Q2/Q12 | 架构分层 | DashrDaemon = profile 层概念,v0.1.5 为空壳(不写代码,YAML 保持 `[]`)。IPythonCodeRuntime 改名 **DashrRuntime**,待在 mount 层(standing mount),cross-session 管理 kernel Map——它才是实际上的 daemon |
|
|
17
|
+
| Q3 | Control Prompt | 独立 systemPrompt.section(dashr:control-prompt),与 Persona 解耦。基于 draft v2 内容 |
|
|
18
|
+
| Q4/Q7 | Tool Catalog | 保留 block,改名 "Tool Catalog"。保留硬编码转换机制(JSON Schema → code 形态呈现),信息量不减、只改呈现。去掉 `tools.` 前缀。DASHR 独有函数加 example |
|
|
19
|
+
| Q5 | 递归深度 | 可配置,默认 10,上限 10 |
|
|
20
|
+
| Q6 | 函数重映射 | subagent→rlm("spawn")、subagent_fork→rlm("fork")、interrupt_agent→rlm("interrupt")、send_message→agent_message、list_agents→agent_list、workflow→rlm_workflow()、ralph→rlm_ralph() |
|
|
21
|
+
| Q8 | 内核绑定 | 平级 top-level globals,无 holder、无前缀。glob→**file_glob**(catalog 注释从模型视角说明 stdlib shadow,不提"上游") |
|
|
22
|
+
| Q13 | file_glob | 唯一重命名的上游工具,其余平级原名 |
|
|
23
|
+
| Q14 | 桥接层级 | rlm() 桥接**工具层**(registry 工具执行管线),继承上游全部 policy(审批/沙箱/maxDepth/config)。one-shot vs continuable 跟随工具层 flag(run_in_background 参数 + backgroundMode 配置) |
|
|
24
|
+
| Q15 | rlm_await | **删除**。阻塞语义由工具层 foreground 模式覆盖(run_in_background=False)。不写自造轮询逻辑 |
|
|
25
|
+
| Q16 | rlm("remove") | **不做**。child 生命周期交给 harness(parent 结束时 drainContinuableDescendants 自动回收 continuable 后代) |
|
|
26
|
+
| Q17 | 调用级 depth override | 不可能——maxDepth 是部署级配置,不进调用 schema |
|
|
27
|
+
| Q18 | 命名 | 单数:agent_message / agent_list |
|
|
28
|
+
| Q19 | rlm 签名 | rlm(mode, ...) mode 为第一位置参数 |
|
|
29
|
+
| Q20 | rlm_workflow/rlm_ralph 展示 | 与 rlm 家族合并展示,共用一段说明 |
|
|
30
|
+
| Q21 | Masking 机制 | **纯呈现层排除**:只在 (1) Tool Catalog 文本 (2) kernel 绑定名两处排除 7 个名字。上游 registry 一指不碰,工具照常注册可执行。不用 restrict(),不用 disable patch |
|
|
31
|
+
| Q22/Q23 | 深度 10 实现 | include patch 改标准 preset 的 tool-subagent / tool-subagent-fork 两行 config(全量复述 + maxDepth: 10)。不自挂实例(避免被上游生态 exclude 的风险) |
|
|
32
|
+
| Q24 | model 选择 | 丢弃 per-call model kwarg。subagentModel 降级为 patch 里 agentOptions.model(静态)。默认 parent 继承 |
|
|
33
|
+
| Q25 | agent_message 上行 | `agent_message(receiver, message, *, subagent_id=None)` 双用:receiver='child' → 桥接 send_message 工具(下行,需 subagent_id);receiver='parent' → 桥接服务层 `ctx.subagents.reportFrom(exec.agent, ...)`(上行,零 ID:父节点由服务层从 child.session.header.parentSession 自举,authorizeReporter 免费校验 live continuable)。delivery 固定 wakeup(上游默认),不暴露给模型。仅 continuable child 可上行(root 调得到结构化 UNAUTHORIZED) |
|
|
34
|
+
|
|
35
|
+
被 masking 的 7 个名字:`subagent`, `subagent_fork`, `send_message`, `list_agents`, `interrupt_agent`, `workflow`, `ralph`。
|
|
36
|
+
|
|
37
|
+
---
|
|
38
|
+
|
|
39
|
+
## 2. 架构定案(v0.1.5 后形态)
|
|
40
|
+
|
|
41
|
+
```
|
|
42
|
+
profile 层(进程级) DashrDaemon —— v0.1.5 空壳(概念 + YAML [],无代码)
|
|
43
|
+
mount 层(standing) DashrRuntime —— 改名自 IPythonCodeRuntime;cross-session kernel Map;
|
|
44
|
+
实际 daemon:监听 agent/disposed、键控 kernels、快照、回环分发
|
|
45
|
+
session 层 ipykernel 子进程 —— 纯 Python 解释器
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
child 继承:spawn/fork 的 child **加入 parent 的 standing mount**(composeFrom),所以每个子 agent 也是 DASHR session——有 run_cell、有 rlm(),递归自调用成立,depth 上限有意义。
|
|
49
|
+
|
|
50
|
+
---
|
|
51
|
+
|
|
52
|
+
## 3. 实现阶段
|
|
53
|
+
|
|
54
|
+
### Phase 1 — 架构(改名,零行为变化)
|
|
55
|
+
|
|
56
|
+
| 文件 | 改动 |
|
|
57
|
+
|------|------|
|
|
58
|
+
| `dashr/src/runtime.ts` | class IPythonCodeRuntime → **DashrRuntime**;module doc 更新层级描述 |
|
|
59
|
+
| `dashr/src/index.ts` | import / export 同步改名;顶部注释更新 |
|
|
60
|
+
| `dashr/test/*.spec.ts` | 引用改名的用例同步 |
|
|
61
|
+
| `dashr/cordis.patch.yml` | 注释更新:DashrDaemon 为 v0.1.5 空壳、DashrRuntime 在 mount 层 |
|
|
62
|
+
| `dashr/preset/rlm-mode/agent.cordis.yml` | 注释更新同层 |
|
|
63
|
+
| `README.md` | 命名与架构段更新 |
|
|
64
|
+
|
|
65
|
+
验收:`npm test` + `npm run typecheck` 全绿;行为与 0.1.4 完全一致。
|
|
66
|
+
|
|
67
|
+
### Phase 2 — A2A + RLM(桥接重构)
|
|
68
|
+
|
|
69
|
+
| 文件 | 改动 |
|
|
70
|
+
|------|------|
|
|
71
|
+
| `dashr/src/index.ts` | ① collectSdkSchemas 排除 7 个 masked 名;② 绑定循环(约 :857)跳过 7 名;③ runtime.run bindings 从单一 tools holder 改为逐工具 top-level namespace(glob→file_glob 映射在此);④ rlm callable 重构为 mode 分发(spawn/fork→subagent/subagent_fork 工具、interrupt→interrupt_agent 工具,nested sub-dispatch);⑤ 新增 agent_message(receiver='child'→send_message 工具 / 'parent'→ctx.subagents.reportFrom(exec.agent),delivery 固定 wakeup)、agent_list、rlm_workflow、rlm_ralph callables;⑥ 删除 rlm_await callable;⑦ RlmRunRegistry 若失去唯一消费者则移除(child 生命周期归 harness) |
|
|
72
|
+
| `dashr/src/py-sdk.ts` | SDK 渲染改为平级形态:无 `class Tools(Protocol)`、无 `tools: Tools` 单例;每个工具渲染为顶层 `async def <name>(args: XArgs) -> XOutput` 签名 + TypedDict 参数/返回类;SDK_INSTRUCTIONS 文本去掉 `tools.` 前缀表述,section 改名 Tool Catalog |
|
|
73
|
+
| `dashr/preset/rlm-mode/agent.cordis.yml` | include patches 新增两行:id=tool-subagent / tool-subagent-fork,config 全量复述 + `maxDepth: 10`(两行现状已是 continuable,复述时保持) |
|
|
74
|
+
| `dashr/test/` | 更新/新增:catalog 不含 7 名、file_glob 出现且 glob 不出现、rlm 四 mode 分发、agent_* 桥接、无 rlm_await、深度 patch 生效(preset.spec) |
|
|
75
|
+
|
|
76
|
+
实现期校验点(不进决策,实现时验证):
|
|
77
|
+
- include patch 的 config 值是否支持 `!!js`(决定 maxDepth/subagentModel 能否 env 化);不支持则硬编码 10 并注明配置点。
|
|
78
|
+
- rlm("spawn") 的 label 缺省处理:上游工具 description 必填,缺省给 'subagent'。
|
|
79
|
+
|
|
80
|
+
验收:单元测试全绿;test harness 挂载 preset 后模型可见 surface 正确(只有 run_cell + 平级 catalog);nested sub-dispatch 通过 guard。
|
|
81
|
+
|
|
82
|
+
### Phase 3 — System Prompt(最后)
|
|
83
|
+
|
|
84
|
+
| 文件 | 改动 |
|
|
85
|
+
|------|------|
|
|
86
|
+
| `dashr/src/` 新增 control-prompt 模块 | 注册 `dashr:control-prompt` section(order ~100,早于 Tool Catalog 的 150)。内容基于 draft v2 适配:平级绑定(无 tools. 前缀)、rlm 家族、agent_message/agent_list、report 工具、foreground/background 双模式、guard 契约("only run_cell is callable directly" 写进说明书不靠失败学习) |
|
|
87
|
+
| `dashr/preset/rlm-mode/agent.cordis.yml` | persona patch 保持极简身份("You are a DASHR agent..."),不承载操作指引 |
|
|
88
|
+
| Tool Catalog 文本 | DASHR 独有函数(rlm/agent_message/agent_list/rlm_workflow/rlm_ralph)带 1-2 个代码 example;通用工具(read/write/bash 等)不加 example |
|
|
89
|
+
|
|
90
|
+
验收:live session 实测——新 session 第一轮不再出现"把工具当原生函数直接调用"的错误;rlm spawn → child report → parent 收到消息全链路走通。
|
|
91
|
+
|
|
92
|
+
---
|
|
93
|
+
|
|
94
|
+
## 4. 风险与坑(来自调查)
|
|
95
|
+
|
|
96
|
+
1. **restrict() 时序炸弹**:不用 restrict(),已规避。
|
|
97
|
+
2. **patch config 全量替换**:上游日后改 delegation 行 config 时我们复述的会静默漂移。缓解:agent.cordis.yml 注释 + 上游升级时对照检查。用户已知情接受。
|
|
98
|
+
3. **messaging 只在 continuable child 上存在**:标准 preset 两行已是 continuable,满足。若上游改回 one-shot,agent_message/report 会失效——升级检查点。
|
|
99
|
+
4. **report 工具只注册在 continuable child scope**:root 的 kernel 里没有 report。catalog 文案需体现(仅子 agent 可用)——Phase 3 处理。
|
|
100
|
+
5. **删除 rlm_await 的同步 fan-in 缺口**:多子任务汇总场景靠 foreground 模式 + asyncio.gather,或事件驱动 report 收集(agent_message(receiver='parent') 上行经 wakeup 唤醒 parent)。用户已确认接受。
|
|
101
|
+
6. **上行通道仅限 continuable child**:root 调 agent_message(receiver='parent') 得到结构化 UNAUTHORIZED。authorizeReporter 免费校验,无需桥接层自造检查。
|
|
102
|
+
|
|
103
|
+
---
|
|
104
|
+
|
|
105
|
+
## 5. 版本与发布
|
|
106
|
+
|
|
107
|
+
- package.json version → 0.1.5
|
|
108
|
+
- README 更新功能列表(A2A、递归深度 10、平级 catalog)
|
|
109
|
+
- 单 commit 一个 feature batch(仓库惯例:public repo single clean commit per feature batch)
|
|
@@ -0,0 +1,50 @@
|
|
|
1
|
+
# Self Survey of `dsh` (DeepSeek Harness Agent) and Prime Agent - Actionable Surface to LLM in Agent runtime
|
|
2
|
+
|
|
3
|
+
|
|
4
|
+
- Actionables: tools, toolset(s), skills, bash, executable code scripts (e.g. python ...)
|
|
5
|
+
- Surface to LLM: the text/prompts exposed to LLM during agent runtime
|
|
6
|
+
|
|
7
|
+
|
|
8
|
+
|
|
9
|
+
|
|
10
|
+
|
|
11
|
+
## `dsh` Native
|
|
12
|
+
|
|
13
|
+
### 地基 — 我"调用工具"的 wire 形态(function calling)
|
|
14
|
+
|
|
15
|
+
- 我要执行动作时,模型输出一个 **JSON function**,里面只有两个 element:
|
|
16
|
+
- `name` —— 我暴露的工具枚举清单里的某个工具名(**平级、flat**,没有 `tools.` 之类的前缀命名空间)
|
|
17
|
+
- `arguments` —— 这个工具自己声明的 tool-call 参数 schema 的 JSON 对象(键值对)
|
|
18
|
+
- 客户端拿到这个 JSON,用 `name` 定位工具、用 `arguments` 做校验,然后硬编码执行。**除此之外没有别的字段。**
|
|
19
|
+
|
|
20
|
+
示例 —— `Function: { name, arguments }`,直接写字段值:
|
|
21
|
+
|
|
22
|
+
1. 网络请求:`Function: { name: "bash", arguments: {"command": "curl -s http://example.com/api"} }`
|
|
23
|
+
2. 代码执行(脚本同样经 `bash` 跑):`Function: { name: "bash", arguments: {"command": "python3 /path/to/script.py"} }`
|
|
24
|
+
|
|
25
|
+
### Actionables — 平级工具枚举清单(29 个,全同级,无前缀)
|
|
26
|
+
|
|
27
|
+
分类只是归类,运行时它们全在同一个 flat 清单里:
|
|
28
|
+
|
|
29
|
+
- shell / 后台任务:`bash` `job_output` `job_list` `job_kill`
|
|
30
|
+
- 文件系统:`read` `write` `edit` `glob` `grep` `read_image`
|
|
31
|
+
- 委派与编排:`subagent` `subagent_fork` `send_message` `list_agents` `interrupt_agent` `workflow` `ralph`
|
|
32
|
+
- goal:`create_goal` `get_goal` `update_goal`
|
|
33
|
+
- memory:`memory_add` `memory_search` `memory_list` `memory_flush`
|
|
34
|
+
- 其他:`todo_write` `web_search` `ask_user_question` `skill` `exit_plan_mode`
|
|
35
|
+
|
|
36
|
+
- `skill` 就是上面清单里的一个平级工具:`arguments.name` 传技能名,把 SKILL.md 指令文本注入上下文(不是独立命名空间)。
|
|
37
|
+
- 可执行代码脚本:走 `bash`(一次性子进程,python / node / shell / CLI 皆可)。
|
|
38
|
+
|
|
39
|
+
### Surface to LLM — 我暴露给模型的文本
|
|
40
|
+
|
|
41
|
+
- 系统提示(每 step 组装,排序分段):harness identity → persona → 各工具包的 cross-call 指引 → plan-mode 段(仅计划模式)→ 变量 `{{model}}` `{{cwd}}`
|
|
42
|
+
- 工具 schema:上面 29 个平级工具各自的 JSON schema(`name` + `parameters`),模型可见子集,按配置或字典序排列
|
|
43
|
+
- skills catalog:`available_skills` 块(技能名 + 一句话摘要)
|
|
44
|
+
- 运行时上下文快照(sourced user-role 消息):沙箱策略、审批策略、workspace、时间、委派范围
|
|
45
|
+
- AGENTS.md 指令链:`~/.dsh/AGENTS.md` + 项目 `AGENTS.md` + 嵌套作用域,以 `<system-reminder>` 注入
|
|
46
|
+
- memory(Corti)召回:每次回复前注入
|
|
47
|
+
- 对话历史 + 压缩(tool-result pruner:threshold 8192 / head 4096 / tail 1024 字符)
|
|
48
|
+
|
|
49
|
+
|
|
50
|
+
|