aiterm-mcp 0.27.9 → 0.28.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.ja.md +43 -37
- package/README.md +44 -40
- package/dist/agent-resolver.js +14 -5
- package/dist/agent-shared.js +18 -5
- package/dist/aiterm-wait-cli.js +1 -1
- package/dist/core.js +171 -22
- package/dist/index.js +110 -58
- package/dist/vendors/codex.js +2 -1
- package/dist/vendors/cursor.js +393 -0
- package/dist/vendors/grok.js +2 -1
- package/package.json +2 -2
package/README.ja.md
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
> **任意のMCPクライアントから、Claude・Codex・Grok・
|
|
1
|
+
> **任意のMCPクライアントから、Claude Code・Codex CLI・Grok CLI・Cursor Agent CLIを単一のharness APIで永続対話TUIへ起動する。**
|
|
2
2
|
|
|
3
3
|
<p align="center">
|
|
4
4
|
<img src=".github/og.png" alt="Aiterm — 異なる知性が一つの持続する実行現場を共有する森の観測拠点" width="100%">
|
|
@@ -16,7 +16,7 @@
|
|
|
16
16
|
|
|
17
17
|
> *(English: [README.md](README.md))*
|
|
18
18
|
|
|
19
|
-
> **あなたの AI に、ほかの AI を操らせる。**
|
|
19
|
+
> **あなたの AI に、ほかの AI を操らせる。** `agent_launch`の1回の呼び出しで、実行基盤harnessとmodelを別々に選び、永続sessionを受け取る。CursorでGPT/Claude/Grokを選んでも、session・hook・transcriptはCursorが所有する。
|
|
20
20
|
>
|
|
21
21
|
> **これは何か:** AI が握る 1 本の永続 MCP 端末——その中に他のコーディングエージェントも起動できる。`ssh`・`docker exec`・REPL・別エージェントの TUI は、すべてその 1 本の端末の中へ「送るだけのテキスト」として入れ子になる。仕組みはあえて素朴——MCP クライアントが相手エージェントの端末を 1 ターンずつ操作するだけ。隠れたプロトコルも・aiterm独自の共有メモリ層も・自律的な交渉も無い。起動したagentは、直接CLIと同じproject/vendorの通常memory・設定を読む。
|
|
22
22
|
>
|
|
@@ -94,7 +94,9 @@ host統合は、kitepon.devの製品開発を支える内部基盤
|
|
|
94
94
|
|
|
95
95
|
**言葉でなく実測で:** 記録済み203テストのベンチマークでは、`pty_read` はコンテキストに載るトークンを生ログの **約 7.1 分の 1** に減らす。しかも pass/fail の判定は畳んでも残る。→ [組み込みシェルツールとの使い分け](#組み込みシェルツールとの使い分け)
|
|
96
96
|
|
|
97
|
-
|
|
97
|
+
15ツール: 6つのPTYツール、正規のagent起動入口`agent_launch`、移行用の旧4alias、`agent_configure`、`claude_turn`、`claude_approval`、`diagnostics`。バックエンドはtmuxなので、MCPサーバやAIクライアントが再起動してもsessionは生き残る。
|
|
98
|
+
|
|
99
|
+
**v0.28.0では実行基盤harnessとmodelを分離した。** harnessはagent loop・認証・hook・session・transcriptを所有し、modelはその上で選ぶ。Cursor Agent CLIでGPT/Claude/Grokを選んでも完了契約はCursor方式のまま。Composerは別harnessではなく、`harness:"grok-cli", model:"grok-composer-2.5-fast"`で表す。旧4起動ツールは同じ実装へ流れる互換alias。
|
|
98
100
|
|
|
99
101
|
**v0.25.2ではGrok 4.6を含む同一sessionの連続設定変更を安定化。** Grok Build 1.0.3で
|
|
100
102
|
`/model`の成功通知が再描画により消えても、変更前には無かった要求model/effortが常駐footerへ現れた
|
|
@@ -151,7 +153,7 @@ runtime-error store は canonical dotagents config の `collection.enabled: true
|
|
|
151
153
|
場合だけ収集し、既定OFF、network送信は行いません。tag起点CIのnpm provenance(OIDC Trusted
|
|
152
154
|
Publishing)で公開し、GitHub Release が Official MCP Registry を再登録します。
|
|
153
155
|
|
|
154
|
-
**状態:** 開発継続中 · この分野では新参で、別の形に賭けている([既存手段との比較](#既存手段との比較)参照)· 動作対象は Linux · WSL2 · macOS · Windows
|
|
156
|
+
**状態:** 開発継続中 · この分野では新参で、別の形に賭けている([既存手段との比較](#既存手段との比較)参照)· 動作対象は Linux · WSL2 · macOS · Windows ネイティブ · MIT · [変更履歴](CHANGELOG.md)。
|
|
155
157
|
|
|
156
158
|
## なぜ今
|
|
157
159
|
|
|
@@ -174,14 +176,14 @@ pty_read(id, { wait: true }) → 削減済みの出力を読む(完了
|
|
|
174
176
|
|
|
175
177
|
### 2. その端末の中に他のコーディングエージェントを起動する — オーケストレーションの旗艦
|
|
176
178
|
|
|
177
|
-
同じ
|
|
179
|
+
同じprimitiveが別エージェントのTUIを宿す。`agent_launch`の`harness`はagent loop・認証・hook・session・transcriptを所有する実行基盤、`model`は独立した選択。起動processは直接CLIと同じproject/user環境を使い、通常config、MCP、plugin、skill、permission、trust、memory、historyをcopy・filter・置換しない。
|
|
178
180
|
|
|
179
|
-
|
|
181
|
+
`aiterm.agent-launch-result.v1`は正規`harness`を返し、旧`provider`は互換fieldとして残す。同じ`harness`はagent dispatch、`aiterm-wait`、`agent_configure`、`pty_list`のagent行にも載り、旧vendor/provider/agent fieldは互換用に残る。Codexは通常rollout、Grok CLIは通常session event、Claudeはlaunch固有Stop hook、Cursorは通常agent transcript末尾の`turn_ended`を完了正本に使う。`pty_send`は非ブロックdispatchで、vendor別完了境界を表すopaqueな整数`event_cursor`を返し、完了通知は`aiterm-wait`を親のターンを塞がない別processで受ける。
|
|
180
182
|
|
|
181
|
-
`
|
|
183
|
+
`agent_launch`は任意の`write_scope`も受ける。Codex/Grokのread-onlyは`--sandbox read-only`、Cursorは公式`--mode ask`で実効化する。path説明は同等CLI引数がないためdeclaration-only。
|
|
182
184
|
|
|
183
185
|
```text
|
|
184
|
-
|
|
186
|
+
agent_launch({ harness: "codex-cli", session_name: "codex1", cwd: "/repo",
|
|
185
187
|
prompt: "port test/legacy.py to vitest",
|
|
186
188
|
model: "gpt-5.6-sol", reasoning_effort: "high",
|
|
187
189
|
write_scope: "test/ only; no commit" })
|
|
@@ -192,14 +194,16 @@ $ aiterm-wait --session codex1 --cursor <event_cursor> # exit 0=done / 3=timeo
|
|
|
192
194
|
→ 操舵し、Codex の次の入力境界で返る
|
|
193
195
|
```
|
|
194
196
|
|
|
195
|
-
|
|
197
|
+
正規のharness選択肢:
|
|
196
198
|
|
|
197
|
-
|
|
|
199
|
+
| `harness` | 起動するもの | modelの扱い |
|
|
198
200
|
| --- | --- | --- |
|
|
199
|
-
| `
|
|
200
|
-
| `
|
|
201
|
-
| `
|
|
202
|
-
| `
|
|
201
|
+
| `claude-code` | Claude Code CLI | Claude model/effort |
|
|
202
|
+
| `codex-cli` | Codex CLI | OpenAI model/effort |
|
|
203
|
+
| `grok-cli` | Grok Build CLI | Grok/Composer model、live catalog照合 |
|
|
204
|
+
| `cursor-cli` | Cursor Agent CLI | Cursor catalog上のGPT/Claude/Grok等 |
|
|
205
|
+
|
|
206
|
+
Cursorの`model`は`gpt-5.6-luna`のようなbase model、`reasoning_effort`は`high`のように別指定する。adapterは現行`model-effort` IDを`cursor-agent models`へ照合し、起動中変更はCursor標準model pickerのparameter editorを使う。不在時は別modelへfallbackしない。
|
|
203
207
|
|
|
204
208
|
`env_vars`は環境変数の**名前**だけを並べるallowlistであり、name/value mapではない。aitermは
|
|
205
209
|
launcher起動時に現在のMCP processから各名前を読み、存在する値をshell quoteして、その1回のvendor
|
|
@@ -208,7 +212,7 @@ launcher起動時に現在のMCP processから各名前を読み、存在する
|
|
|
208
212
|
PTYの起動コマンドとして送られ、sessionの`.lastcmd`にも保持されるため、起動先vendorと同じOS userへ
|
|
209
213
|
到達する。秘密転送路ではなく、席identityやworkflow用の非secret変数だけに使う。
|
|
210
214
|
|
|
211
|
-
|
|
215
|
+
選んだharnessのCLIを公式経路で導入・認証しておく。Cursorは`curl https://cursor.com/install -fsS | bash`、`agent login`、更新は`agent update`が公式経路で、Aitermは曖昧な`agent`でなく`cursor-agent`を起動する。CLI不在・未認証・不正引数はsession作成前に明示失敗し、別経路へfallbackしない。
|
|
212
216
|
|
|
213
217
|
portable forkは任意である。`throughline_source_session`を使う場合、`prompt`は必須の新ミッションとなり、
|
|
214
218
|
`launch_operation_id`とは併用できない。aitermは`THROUGHLINE_BIN`、次に`PATH`からThroughlineを解決し、
|
|
@@ -216,7 +220,7 @@ portable forkは任意である。`throughline_source_session`を使う場合、
|
|
|
216
220
|
この経路だけ`throughline >= 0.9.0`が必要で、元sessionのDB所属は変わらない。引数省略時には
|
|
217
221
|
Throughline自体が不要である。
|
|
218
222
|
|
|
219
|
-
エージェント間の隠れたプロトコルは無い。起動した
|
|
223
|
+
エージェント間の隠れたプロトコルは無い。起動したharnessは利用者がattachできるもう1本の永続sessionであり、MCPクライアントが通常のPTY操作で駆動する。
|
|
220
224
|
|
|
221
225
|
## デモ
|
|
222
226
|
|
|
@@ -271,7 +275,7 @@ Throughline自体が不要である。
|
|
|
271
275
|
Claude Code を再起動して、接続を確認:
|
|
272
276
|
|
|
273
277
|
```bash
|
|
274
|
-
/mcp # aiterm が connected・
|
|
278
|
+
/mcp # aiterm が connected・15 ツール公開、と出る
|
|
275
279
|
```
|
|
276
280
|
|
|
277
281
|
最初のセッション——4 回の呼び出しで、1 個の永続端末:
|
|
@@ -286,7 +290,7 @@ pty_close("t1") → 端末を解放
|
|
|
286
290
|
`pty_close` は冪等で、`closed` / `already_closed` のstructured receiptを返す。
|
|
287
291
|
MCP応答を失ったdurable callerも同じ`session_id`への再試行だけでclose結果を確定できる。
|
|
288
292
|
|
|
289
|
-
これだけ。`t1` の端末は本物で永続——`ssh`・`docker exec`・REPL・起動したエージェントの
|
|
293
|
+
これだけ。`t1` の端末は本物で永続——`ssh`・`docker exec`・REPL・起動したエージェントのTUIは、そこに住む「もの」に過ぎない。ワーカー起動も1コールで、`agent_launch({ harness: "codex-cli" })`が返す`session_id`を同じ`pty_read`/`pty_send`で操作する。
|
|
290
294
|
|
|
291
295
|
**グローバル導入や別クライアントが良い場合は:**
|
|
292
296
|
|
|
@@ -312,11 +316,11 @@ MCP クライアントが aiterm を stdio 越しにプログラムから駆動
|
|
|
312
316
|
|
|
313
317
|
```mermaid
|
|
314
318
|
flowchart LR
|
|
315
|
-
AI["AI / MCP client<br/>(the orchestrator)"] -->|"pty_send ·
|
|
319
|
+
AI["AI / MCP client<br/>(the orchestrator)"] -->|"pty_send · agent_launch · agent_configure · claude_turn · claude_approval<br/>旧launcher alias · diagnostics"| S["aiterm-mcp<br/>stdio MCP · 15 tools"]
|
|
316
320
|
S -->|"pty_read<br/>token-reduced"| AI
|
|
317
321
|
S -->|"tmux send-keys<br/>capture-pane"| P["persistent PTYs<br/>tmux · survive restarts"]
|
|
318
322
|
P -->|"ssh · docker · repl"| R["nested<br/>remote · container · REPL"]
|
|
319
|
-
P -->|"launches a fresh PTY per agent"| A["another coding-agent TUI<br/>Claude · Codex · Grok ·
|
|
323
|
+
P -->|"launches a fresh PTY per agent"| A["another coding-agent TUI<br/>Claude · Codex · Grok · Cursor"]
|
|
320
324
|
```
|
|
321
325
|
|
|
322
326
|
primitive は「PTY を 1 個握る」ことだけ。それ以外——SSH・コンテナ・REPL・起動したエージェント TUI——は、永続端末の中で動く「対話的な何か」に過ぎず、同じ `pty_send` / `pty_read` で操作する。各起動ツールは自分専用の新しい PTY を開く。PTY は tmux 上にあるので、MCP サーバや AI クライアントが再起動してもセッションは生き残る。
|
|
@@ -363,7 +367,7 @@ aiterm は 2 つの系譜の交点にいる——端末を操作する MCP サ
|
|
|
363
367
|
| --- | --- | --- | --- | --- |
|
|
364
368
|
| 永続セッション | ✅ tmux・再起動を跨ぐ | ❌ 毎回新シェル | ⚠️ まちまち | ✅ tmux |
|
|
365
369
|
| SSH / コンテナ / REPL | `pty_send` 1 回でネスト | 毎コマンド接続し直し | ⚠️ ツールが分かれがち | ✅ tmux(人が操作) |
|
|
366
|
-
| 1 コールで別エージェント起動 | ✅ `
|
|
370
|
+
| 1 コールで別エージェント起動 | ✅ `agent_launch(harness=…)` | ❌ | ❌ | ⚠️ 人が動かす tmux に CLI + skills で参加 |
|
|
367
371
|
| ヘッドレス(人が tmux に居ない) | ✅ MCP 駆動・プログラム的 | ✅ | ⚠️ まちまち | ❌ 人が tmux に居る前提 |
|
|
368
372
|
| MCP ネイティブ(任意の MCP クライアント) | ✅ `claude mcp add` 1 行 | ✅ | ✅(MCP なので) | ❌ tmux 設定 + CLI + Agent Skills |
|
|
369
373
|
| トークン削減読取 | ✅ コマンド別 reducer | ❌ 生出力 | ⚠️ ほぼ無し | ❌ 生 tmux |
|
|
@@ -392,8 +396,10 @@ aiterm は同じ核心の洞察——端末を出会いの場にする——を
|
|
|
392
396
|
| `pty_read` | 出力を削減して読む(既定は増分) | `session_id`, `wait`, `until`, `until_regex`, `timeout`, `screen`, `full`, `lines`, `line_range`, `raw`, `rtk`, `agent_transcript`, `operation_id` |
|
|
393
397
|
| `pty_key` | 制御キーを送る | `session_id`, `key`(`C-c`/`Enter`/`Up`…) |
|
|
394
398
|
| `pty_close` | 冪等に閉じ、`closed` / `already_closed`を返す | `session_id` |
|
|
395
|
-
| `pty_list` |
|
|
396
|
-
| `
|
|
399
|
+
| `pty_list` | セッション一覧(agent行は正規`harness=<id>`と互換`agent=<kind>`を含む) | (なし) |
|
|
400
|
+
| `agent_launch` | harnessとmodelを別軸で選ぶ正規agent起動入口 | `harness`, `prompt?`, `model?`, `reasoning_effort?`, `cwd?`, `write_scope?` |
|
|
401
|
+
| `claude_agent` / `codex_agent` / `grok_agent` / `composer_agent` | deprecated互換alias | 旧launcher引数 |
|
|
402
|
+
| `agent_configure` | 起動中のClaude/Codex/Grok/Composer/Cursorを再起動せずmodel/effort変更 | `session_id`, `model?`, `reasoning_effort?` |
|
|
397
403
|
| `claude_turn` | 相関済みClaude operationをdispatch(issue)または回収(recover) | `action`, `session_id`, `operation_id`, `text?` |
|
|
398
404
|
| `claude_approval` | 現在表示中の相関済みClaude承認UIを検査または応答 | `action`, `session_id`, `operation_id?`, `approval_choice?`, `observed_prompt_digest?` |
|
|
399
405
|
| `diagnostics` | 機械可読 JSON による read-only factory readiness | (なし) |
|
|
@@ -406,31 +412,31 @@ aiterm は同じ核心の洞察——端末を出会いの場にする——を
|
|
|
406
412
|
|
|
407
413
|
consumer は `aiterm-runtime-errors snapshot` を読み、durable ingestion 後に `aiterm-runtime-errors ack --cursor N` を呼ぶ。運用上の明示操作は `resolve|reopen --fingerprint SHA256`。MCP からの収集・diagnostic read は timeout 付き child process に隔離し、FIFOや停止 filesystem が端末本体を止めない。store mutation は期限付き bakery ticket queue で直列化する。各waiterは PID+process start identity+owner token を持つ再利用されない固有ticketを所有するため、死んだownerだけを固有名で除去でき、固定path回収のABAを作らない。queueの期限は正常な前任者を含む総待ち時間ではなく、同じ先頭ownerが進まない時間を測る。通常pollはprocessの生存確認だけを行い、process start identityはblockerがstallした時に照合する。POSIX state は `$XDG_STATE_HOME/aiterm-mcp/`(既定 `~/.local/state/aiterm-mcp/`)へ atomic replacement で置き、every read で owner/mode を再検証する。Windows native は `%LOCALAPPDATA%\aiterm-mcp\` で current SID の非継承 FullControl ACE 1件だけへ DACL を再構築し readback する。今回 Windows は path/DACL/timeout の純粋テストだけであり、新しい実機統合成功は主張しない。
|
|
408
414
|
|
|
409
|
-
###
|
|
415
|
+
### 対話エージェントharness
|
|
410
416
|
|
|
411
|
-
|
|
417
|
+
`agent_launch`は選んだharnessの対話TUIを新しい永続PTYに起動し、`session_id`を返す。harnessはagent loop・認証・hook・session・transcriptを所有し、modelは独立。以後は他sessionと同じ`pty_read`/`pty_send`で操作する。
|
|
412
418
|
|
|
413
|
-
`agent_configure({ session_id, model?, reasoning_effort? })`は
|
|
419
|
+
`agent_configure({ session_id, model?, reasoning_effort? })`はharness標準操作で起動中のClaude/Codex/Grok/Composer/Cursorを変更し、PTYと会話contextを維持する。
|
|
414
420
|
|
|
415
|
-
|
|
|
421
|
+
| `harness` | 起動するもの | modelの扱い |
|
|
416
422
|
| --- | --- | --- |
|
|
417
|
-
| `
|
|
418
|
-
| `
|
|
419
|
-
| `
|
|
420
|
-
| `
|
|
423
|
+
| `claude-code` | Claude Code CLI | Claude model/effort |
|
|
424
|
+
| `codex-cli` | Codex CLI | OpenAI model/effort |
|
|
425
|
+
| `grok-cli` | Grok Build CLI | Grok/Composer model、live catalog照合 |
|
|
426
|
+
| `cursor-cli` | Cursor Agent CLI | Cursor catalog上のGPT/Claude/Grok等 |
|
|
421
427
|
|
|
422
|
-
対応するCLI(`claude
|
|
428
|
+
対応するCLI(`claude`/`codex`/`grok`/`cursor-agent`)の公式導入・認証が必要。前提違反はsession作成前に明示失敗する。全harnessが通常project/user環境と同じ非ブロックdispatch契約を使う。
|
|
423
429
|
|
|
424
430
|
`throughline_source_session`と空でない新ミッション`prompt`を指定すると、Throughlineの読み取り専用
|
|
425
431
|
handoff contextを前置きできる。この任意経路は`throughline >= 0.9.0`を必要とし、
|
|
426
432
|
`launch_operation_id`とは併用不可で、元sessionのDB所属を変更しない。Throughlineは
|
|
427
433
|
`THROUGHLINE_BIN`、次に`PATH`から解決し、不在・不正・空のexportはPTY作成前に明示失敗する。
|
|
428
434
|
|
|
429
|
-
エージェントの回答が画面
|
|
435
|
+
エージェントの回答が画面tailより長ければ、`pty_read({ agent_transcript:true })`で再promptなしに全文回収する。Claudeはlaunch相関Stop hook、Codexは通常rollout、Grokは通常session history、Cursorはlaunch IDでbindした通常agent transcriptから同じturnを回収する。
|
|
430
436
|
|
|
431
437
|
### 完了検出(5 層)
|
|
432
438
|
|
|
433
|
-
`pty_read({ wait: true })`は通常PTYを、process終了/`mark:true` sentinel/`until`一致/shell復帰を伴う出力静止/timeoutの5
|
|
439
|
+
`pty_read({ wait: true })`は通常PTYを、process終了/`mark:true` sentinel/`until`一致/shell復帰を伴う出力静止/timeoutの5層で判定する。agent sessionは第6の正確な層を使い、Codexは通常rollout、Grokは通常session event、Claudeはlaunch相関Stop event、Cursorは通常agent transcriptの`turn_ended`を`aiterm-wait --cursor`が観測する。親はブロックもポーリングもしない。
|
|
434
440
|
|
|
435
441
|
### トークン削減
|
|
436
442
|
|
|
@@ -446,7 +452,7 @@ handoff contextを前置きできる。この任意経路は`throughline >= 0.9.
|
|
|
446
452
|
|
|
447
453
|
## 人が覗く
|
|
448
454
|
|
|
449
|
-
セッションは共有
|
|
455
|
+
セッションは共有tmux socket(Windows nativeはpsmux namespace)上にある。`pty_open`/`agent_launch`の戻り値に表示されるattachコマンドで、人間がClaude/Codex/Grok/Cursor harnessの同じ端末へ入り、途中でキーボードを引き取れる。
|
|
450
456
|
|
|
451
457
|
## 要件
|
|
452
458
|
|
|
@@ -454,7 +460,7 @@ handoff contextを前置きできる。この任意経路は`throughline >= 0.9.
|
|
|
454
460
|
- **tmux または psmux**(実行時の前提)
|
|
455
461
|
- **macOS / Linux / WSL2** は tmux を直接使う。macOS は同梱されないので `brew install tmux` で導入する。MCP クライアントがターミナルでなく **GUI から起動**された場合、Homebrew の bin(Apple Silicon: `/opt/homebrew/bin`、Intel: `/usr/local/bin`)が `PATH` に入らないことがある。その場合 aiterm が自動で探索するか、**`AITERM_TMUX=/path/to/tmux`** で明示指定する。
|
|
456
462
|
- **Windows ネイティブ**は WSL を使わず、tmux CLI互換の [psmux](https://github.com/psmux/psmux) **3.3.8以上**を直接使う(`winget install marlocarlo.psmux`)。pane shell用に Git for Windows も必要。解決先は **`AITERM_PSMUX`**/**`AITERM_BASH`** で上書きできる。Windows toolはSSHと同じく入れ子で握れ、`pty_send "powershell.exe"`でPowerShellへ入れる。
|
|
457
|
-
-
|
|
463
|
+
- **agent harness**を使う場合: 対応CLIを製品所有者の公式経路で導入・認証する。CursorはmacOS/Linux/WSLで`curl https://cursor.com/install -fsS | bash`、Windows nativeで`irm 'https://cursor.com/install?win32=true' | iex`を使い、`agent login`で認証、`agent update`で更新する。Aitermは`cursor-agent`を起動する。portable forkだけは追加で`throughline >= 0.9.0`が必要。
|
|
458
464
|
- 任意: [`rtk`](https://github.com/rtk-ai/rtk) バイナリ(`pty_send` の `rtk: true` 委譲で使う。無くても動く)
|
|
459
465
|
|
|
460
466
|
## 既知の制約(バグではなく仕様)
|
|
@@ -462,7 +468,7 @@ handoff contextを前置きできる。この任意経路は`throughline >= 0.9.
|
|
|
462
468
|
- **ネスト中(ssh / docker / REPL / 起動したエージェント TUI)は quiescence が原理的に効かない。** 前面コマンドがシェル集合(bash/sh/zsh/fish/dash)の外になるため。ネスト中で `until` も `mark` も無いときは、待っても完了を確定できる信号が無いので、`pty_read({ wait: true })` はフル `timeout` を空費せず出力静止時点で `is_complete=False via nested` と早期に返し、`until`(既定リテラル部分一致・`until_regex: true` で正規表現)か `mark: true`(終了コード付き sentinel・自動検出)の指定を促す。全画面のエージェント TUI なら、出力が落ち着いた時点で `{ screen: true }` を読む。
|
|
463
469
|
- **`is_complete=False` は失敗ではない。** 「timeout 内に完了を観測できなかった」という意味。長時間コマンドでは `timeout` を伸ばすか `until`/`mark` を使う。
|
|
464
470
|
- **破壊ゲートはサンドボックスではなく tripwire。** よくある破壊形だけを弾く。相対パスの `rm`、`$VAR` 展開後に危険化するもの、ssh 先で実行されるコマンドは捕捉しない——起動したコーディングエージェントが自分のセッション内で何をするかも取り締まらない。
|
|
465
|
-
-
|
|
471
|
+
- **agent harnessは実物TUIを起動し、model APIを代理しない。** model・認証・挙動は選んだharnessのもの。隠れたagent間protocolはなく、MCPクライアントがClaude/Codex/Grok/Cursor TUIへ入力を送り出力を読む。
|
|
466
472
|
- **`pty_send({ rtk: true })` は単行コマンドのみ+外部 `rtk` バイナリが必要**(無ければ素通し)。一方 `pty_read({ rtk: true })` の reducer は自前実装で rtk 非依存。
|
|
467
473
|
- **`pytest` reducer は件数・罫線・`FAILURES` ブロック整形が rtk 0.42.0 と byte 一致**(回帰テストで固定)。ただし `-ra`/`-rf` 時の `FAILED` 要約行の理由は**全文を保持する**(rtk 0.42.0 は最初の `" - "` 区切りで切るが、本実装は可読性優先で情報を残すため、この行は意図的に rtk と完全一致させない)。rtk が大出力時に付ける `[full output: …]`(tee ポインタ)行は read 側では再現しない。
|
|
468
474
|
- **tmux は `-f /dev/null` 起動**なので `~/.tmux.conf` を読まない(環境差を排除するため)。
|
package/README.md
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
> **From any MCP client, launch Claude, Codex, Grok, or
|
|
1
|
+
> **From any MCP client, launch Claude Code, Codex CLI, Grok CLI, or Cursor Agent CLI through one harness API inside a persistent interactive TUI.**
|
|
2
2
|
|
|
3
3
|
<p align="center">
|
|
4
4
|
<img src=".github/og.png" alt="Aiterm — a shared forest observatory where different intelligences work in one persistent execution space" width="100%">
|
|
@@ -16,7 +16,7 @@
|
|
|
16
16
|
|
|
17
17
|
> *(日本語: [README.ja.md](README.ja.md))*
|
|
18
18
|
|
|
19
|
-
> **Let your AI orchestrate other AIs.**
|
|
19
|
+
> **Let your AI orchestrate other AIs.** One `agent_launch` call selects the execution harness separately from its model and hands you a persistent session to drive. Cursor can run GPT, Claude, or Grok while Cursor still owns the session, hooks, and transcript.
|
|
20
20
|
>
|
|
21
21
|
> **What it is:** one persistent MCP terminal your AI drives — and can launch other coding agents into. `ssh`, `docker exec`, a REPL, or another agent's TUI all nest inside that one terminal as just text you send in. The mechanism is deliberately plain — your MCP client drives the other agent's terminal turn by turn: no hidden protocol, no separate aiterm-owned shared-memory layer, no autonomous negotiation. Launched agents still read the normal project and vendor memory/configuration that a direct CLI launch would use.
|
|
22
22
|
>
|
|
@@ -94,7 +94,9 @@ toolchain behind kitepon.dev's products.
|
|
|
94
94
|
|
|
95
95
|
**Measured, not claimed:** in the recorded 203-test benchmark, a `pty_read` puts **~7.1× fewer tokens** in your context than the raw log — and the pass/fail verdict survives the fold. → [When to reach for it vs. the built-in shell](#when-to-reach-for-it-vs-the-built-in-shell)
|
|
96
96
|
|
|
97
|
-
|
|
97
|
+
Fifteen tools: six **PTY tools** — `pty_open` / `pty_send` / `pty_read` / `pty_key` / `pty_close` / `pty_list` — to open, drive, and read one persistent terminal; one canonical **agent launcher**, `agent_launch`, which selects `claude-code`, `codex-cli`, `grok-cli`, or `cursor-cli` as the execution harness; four deprecated launcher aliases kept for migration; `agent_configure`; `claude_turn`; `claude_approval`; and `diagnostics`. The backend is **tmux**, so sessions survive even if the MCP server or the AI client restarts.
|
|
98
|
+
|
|
99
|
+
**v0.28.0 separates the execution harness from the model.** The harness owns the agent loop, authentication, hooks, session, and transcript; `model` is what that harness runs. Cursor Agent CLI can therefore select GPT, Claude, or Grok without changing the completion contract from Cursor hooks to another vendor's. Grok Composer is a Grok CLI model preset, not another harness: use `harness: "grok-cli", model: "grok-composer-2.5-fast"`. The old four launcher tools are thin compatibility aliases over the same implementation.
|
|
98
100
|
|
|
99
101
|
**v0.25.2 stabilizes repeated in-place configuration changes, including Grok 4.6.** If Grok Build
|
|
100
102
|
1.0.3 redraws before its `/model` success notice can be observed, aiterm confirms the requested model/effort
|
|
@@ -165,7 +167,7 @@ collection is off by default and performs no network I/O. It ships via
|
|
|
165
167
|
tag-triggered CI with npm provenance (OIDC Trusted Publishing); the GitHub
|
|
166
168
|
Release re-registers the Official MCP Registry entry.
|
|
167
169
|
|
|
168
|
-
**Status:** actively maintained · the newcomer here, betting on a different shape (see [vs. the alternatives](#vs-the-alternatives)) · runs on Linux · WSL2 · macOS · native Windows (tmux on POSIX, the tmux-CLI-compatible [psmux](https://github.com/psmux/psmux) on native Windows —
|
|
170
|
+
**Status:** actively maintained · the newcomer here, betting on a different shape (see [vs. the alternatives](#vs-the-alternatives)) · runs on Linux · WSL2 · macOS · native Windows (tmux on POSIX, the tmux-CLI-compatible [psmux](https://github.com/psmux/psmux) on native Windows — no WSL required) · MIT · see the [CHANGELOG](CHANGELOG.md).
|
|
169
171
|
|
|
170
172
|
## Why now
|
|
171
173
|
|
|
@@ -194,16 +196,16 @@ pty_read(id, { wait: true }) → read the token-reduced output, completion
|
|
|
194
196
|
|
|
195
197
|
### 2. Launch other coding agents into that terminal — the orchestration flagship
|
|
196
198
|
|
|
197
|
-
The same primitive hosts another agent's TUI.
|
|
199
|
+
The same primitive hosts another agent's TUI. `agent_launch` starts a selected execution harness inside a fresh persistent terminal and returns a `session_id`. `harness` names the component that owns the agent loop, authentication, hooks, session, and transcript; `model` remains an independent choice. The launched process sees the same project and user environment as a direct CLI invocation: normal configuration, MCPs, plugins, skills, permissions, trust decisions, memory, and history are not copied, filtered, or replaced. Aiterm adds only completion correlation and a non-user sub-agent context containing `role=subagent`, the parent session, delegation depth, lineage, and `delegation_allowed=true`.
|
|
198
200
|
|
|
199
|
-
The human-readable launch text is accompanied by an `aiterm.agent-launch-result.v1` structured receipt
|
|
201
|
+
The human-readable launch text is accompanied by an `aiterm.agent-launch-result.v1` structured receipt containing the canonical `harness`; the old `provider` field remains for compatibility. The same `harness` is carried by agent dispatch, `aiterm-wait`, `agent_configure`, and agent rows in `pty_list`, while their old vendor/provider/agent fields remain compatibility fields. Codex completion comes from its normal durable rollout transcript, Grok CLI from its normal session events, Claude Code from a launch-specific Stop hook settings addition, and Cursor from its normal agent transcript's terminal `turn_ended` record. Sending to any agent session is a non-blocking **dispatch** — the call returns immediately with an opaque, harness-specific integer `event_cursor`, and completion arrives via [`aiterm-wait`](#completion-push-for-parent-agents-aiterm-wait).
|
|
200
202
|
|
|
201
|
-
`
|
|
203
|
+
`agent_launch` accepts an optional `write_scope`: either `"read-only"` or a human-readable description of writable paths. Codex/Grok use `--sandbox read-only`; Cursor uses its official read-only `--mode ask`. A path description remains declaration-only because these CLI launch surfaces provide no equivalent path allowlist flag.
|
|
202
204
|
|
|
203
205
|
For a correlated Claude turn stopped at `Do you want to proceed?`, use `claude_approval(action: "inspect", ...)` to capture the active operation and SHA-256 screen digest, review the displayed command, then call `respond` with that exact digest and either `approve_once` or `deny`. The relay rechecks the operation and screen under the send lock, never exposes arbitrary input or permanent approval, keeps the active marker intact, and records a prompt-free owner-only receipt. `pty_send(force: true)` does not bypass this boundary.
|
|
204
206
|
|
|
205
207
|
```text
|
|
206
|
-
|
|
208
|
+
agent_launch({ harness: "codex-cli", session_name: "codex1", cwd: "/repo",
|
|
207
209
|
prompt: "port test/legacy.py to vitest",
|
|
208
210
|
model: "gpt-5.6-sol", reasoning_effort: "high",
|
|
209
211
|
write_scope: "test/ only; no commit" })
|
|
@@ -215,14 +217,14 @@ $ aiterm-wait --session codex1 --cursor <event_cursor> # never in the parent's
|
|
|
215
217
|
pty_read("codex1", { agent_transcript: true }) → collect the full answer
|
|
216
218
|
```
|
|
217
219
|
|
|
218
|
-
|
|
220
|
+
The canonical harness choices are:
|
|
219
221
|
|
|
220
|
-
|
|
|
222
|
+
| `harness` | Launches | Notes |
|
|
221
223
|
| --- | --- | --- |
|
|
222
|
-
| `
|
|
223
|
-
| `
|
|
224
|
-
| `
|
|
225
|
-
| `
|
|
224
|
+
| `claude-code` | Claude Code CLI | Claude model and effort controls; correlated Stop hook |
|
|
225
|
+
| `codex-cli` | Codex CLI | OpenAI model and effort controls; durable rollout completion |
|
|
226
|
+
| `grok-cli` | Grok Build CLI | Grok or Composer model selected with `model`; live catalog check |
|
|
227
|
+
| `cursor-cli` | Cursor Agent CLI | GPT, Claude, Grok, or another Cursor catalog model; normal transcript completion |
|
|
226
228
|
|
|
227
229
|
`env_vars` is an allowlist of environment-variable **names**, not a name/value map. At launch,
|
|
228
230
|
aiterm reads each valid name from its current MCP process, shell-quotes present values, and places
|
|
@@ -233,7 +235,7 @@ PTY launch command and retained in aiterm's per-session `.lastcmd`; the launched
|
|
|
233
235
|
processes with access to the same OS user may read them. Use this for non-secret seat identity and
|
|
234
236
|
workflow variables, not as a secret transport.
|
|
235
237
|
|
|
236
|
-
The
|
|
238
|
+
The selected harness CLI must be installed and authenticated. Aiterm resolves `CLAUDE_BIN` / `CODEX_BIN` / `GROK_BIN` / `CURSOR_AGENT_BIN`, then the documented default binary, then `PATH`. Cursor resolution deliberately uses `cursor-agent`, never the ambiguous `agent` name. Claude and Cursor authentication are checked before a PTY exists, so a failed preflight leaves no session. All harnesses use their normal vendor-owned credential and configuration stores in place.
|
|
237
239
|
|
|
238
240
|
Portable fork is optional. When `throughline_source_session` is present, `prompt` is the required
|
|
239
241
|
new mission and `launch_operation_id` cannot be combined with it. aiterm resolves Throughline via
|
|
@@ -242,9 +244,9 @@ places its returned context before a fixed separator and the mission. This route
|
|
|
242
244
|
`throughline >= 0.9.0`; it reads the source memory without changing that database's session
|
|
243
245
|
ownership. No Throughline dependency is needed when the field is omitted.
|
|
244
246
|
|
|
245
|
-
|
|
247
|
+
Harness adapters translate `model` and `reasoning_effort` into each CLI's public controls. Explicit Grok models are checked against `grok models`; Cursor combines a base model such as `gpt-5.6-luna` with a separate effort such as `high`, checks the resulting current catalog ID, and uses Cursor's standard model picker for in-session changes. Missing models are errors, with no cache, retry, or fallback. Claude adds only launch-local Stop-hook settings, Codex reads its normal rollout store, Grok reads its normal session event/history, and Cursor binds its normal agent transcript with the launch ID. Pass an absolute `cwd`; `~` is not expanded.
|
|
246
248
|
|
|
247
|
-
There is no hidden protocol between agents:
|
|
249
|
+
There is no hidden protocol between agents: every launched harness is another user-visible persistent terminal session. The MCP client drives that TUI with ordinary PTY operations, and a human can attach to watch or take over.
|
|
248
250
|
|
|
249
251
|
## Demo
|
|
250
252
|
|
|
@@ -299,7 +301,7 @@ The only edits to the captures above are the two `⋮` lines (a long head/tail r
|
|
|
299
301
|
Restart Claude Code, then verify the connection:
|
|
300
302
|
|
|
301
303
|
```bash
|
|
302
|
-
/mcp # aiterm should show as connected, exposing
|
|
304
|
+
/mcp # aiterm should show as connected, exposing 15 tools
|
|
303
305
|
```
|
|
304
306
|
|
|
305
307
|
Your first session — four calls, one persistent terminal:
|
|
@@ -314,7 +316,7 @@ pty_close("t1") → terminal released
|
|
|
314
316
|
`pty_close` is idempotent and returns a structured `closed` / `already_closed`
|
|
315
317
|
receipt, so durable callers can retry the same `session_id` after losing the MCP response.
|
|
316
318
|
|
|
317
|
-
That's it. The terminal in `t1` is real and persistent — `ssh`, `docker exec`, a REPL, or a launched agent's TUI are just things that live inside it. To launch a worker agent instead, one call does it: `
|
|
319
|
+
That's it. The terminal in `t1` is real and persistent — `ssh`, `docker exec`, a REPL, or a launched agent's TUI are just things that live inside it. To launch a worker agent instead, one call does it: `agent_launch({ harness: "codex-cli" })` returns a `session_id` you drive with the same `pty_read` / `pty_send`.
|
|
318
320
|
|
|
319
321
|
**Prefer a global install, or a different client?**
|
|
320
322
|
|
|
@@ -340,11 +342,11 @@ The terminal is real and shared, so a human *can* jump in ([A human can watch](#
|
|
|
340
342
|
|
|
341
343
|
```mermaid
|
|
342
344
|
flowchart LR
|
|
343
|
-
AI["AI / MCP client<br/>(the orchestrator)"] -->|"pty_send ·
|
|
345
|
+
AI["AI / MCP client<br/>(the orchestrator)"] -->|"pty_send · agent_launch · agent_configure · claude_turn · claude_approval<br/>legacy launcher aliases · diagnostics"| S["aiterm-mcp<br/>stdio MCP · 15 tools"]
|
|
344
346
|
S -->|"pty_read<br/>token-reduced"| AI
|
|
345
347
|
S -->|"tmux send-keys<br/>capture-pane"| P["persistent PTYs<br/>tmux · survive restarts"]
|
|
346
348
|
P -->|"ssh · docker · repl"| R["nested<br/>remote · container · REPL"]
|
|
347
|
-
P -->|"launches a fresh PTY per agent"| A["another coding-agent
|
|
349
|
+
P -->|"launches a fresh PTY per agent"| A["another coding-agent harness<br/>Claude Code · Codex CLI · Grok CLI · Cursor CLI"]
|
|
348
350
|
```
|
|
349
351
|
|
|
350
352
|
One PTY is the only primitive. Everything else — SSH, containers, REPLs, and the launched agent TUIs — is just something interactive running inside a persistent terminal, driven with the same `pty_send` / `pty_read`. Each launcher opens its own fresh PTY. Because the PTYs live in tmux, sessions outlive the MCP server and the AI client.
|
|
@@ -393,7 +395,7 @@ aiterm sits at the intersection of two families: terminal-driving MCP servers, a
|
|
|
393
395
|
| --- | --- | --- | --- | --- |
|
|
394
396
|
| Persistent session | ✅ tmux, survives restarts | ❌ new shell every call | ⚠️ varies | ✅ tmux |
|
|
395
397
|
| SSH / containers / REPLs | nest with one `pty_send` | reconnect every command | ⚠️ often separate tools | ✅ tmux (human drives) |
|
|
396
|
-
| Launch another agent in one call | ✅ `
|
|
398
|
+
| Launch another agent in one call | ✅ `agent_launch(harness=…)` | ❌ | ❌ | ⚠️ agents join a human-run tmux via a CLI + skills |
|
|
397
399
|
| Headless (no human at a tmux) | ✅ MCP-driven, programmatic | ✅ | ⚠️ varies | ❌ built around a human in the tmux |
|
|
398
400
|
| MCP-native (any MCP client) | ✅ one `claude mcp add` | ✅ | ✅ (they are MCPs) | ❌ tmux config + CLI + Agent Skills |
|
|
399
401
|
| Token-reduced reads | ✅ per-command reducers | ❌ raw output | ⚠️ rarely | ❌ raw tmux |
|
|
@@ -409,7 +411,7 @@ aiterm takes the same core insight — the terminal as the meeting point — and
|
|
|
409
411
|
|
|
410
412
|
1. **Headless by construction.** Because aiterm is driven programmatically over MCP, an AI can launch and drive another agent with *no human sitting in the tmux* — from an orchestration loop, a CI step, or a cron job. The shared-tmux tools lead with a human at the keyboard (their docs center on interactive pane navigation), so unattended operation isn't their native mode; aiterm's is.
|
|
411
413
|
2. **MCP-native, not a workflow you adopt.** aiterm is a stdio MCP server: one `claude mcp add` line and it works as structured tools in any MCP client that speaks stdio (tested in Claude Code; Cursor, Cline, and Claude Desktop speak the same protocol and should work the same way). It doesn't ask you to adopt a tmux config, learn pane navigation, or install skills into your setup — the client already knows how to call tools.
|
|
412
|
-
3. **Launching an agent is one tool call — an orchestration primitive.** `
|
|
414
|
+
3. **Launching an agent is one tool call — an orchestration primitive.** `agent_launch({ harness: "codex-cli" })` spawns Codex in a persistent terminal and returns a session you drive immediately. You don't arrange panes or paste between them by hand; the launch, the steering, and the reads are all tool calls the orchestrating model can make on its own.
|
|
413
415
|
|
|
414
416
|
On top of that sits a productized layer a raw tmux bridge doesn't have: **token-reduced reads**, **5-layer completion detection**, and a **destructive-command tripwire**. None of this makes the human-in-the-tmux model wrong — it's a different, complementary bet on where the human is standing.
|
|
415
417
|
|
|
@@ -422,8 +424,10 @@ On top of that sits a productized layer a raw tmux bridge doesn't have: **token-
|
|
|
422
424
|
| `pty_read` | Read output, token-reduced (incremental by default) | `session_id`, `wait`, `until`, `until_regex`, `timeout`, `screen`, `full`, `lines`, `line_range`, `raw`, `rtk`, `agent_transcript`, `operation_id` |
|
|
423
425
|
| `pty_key` | Send a control key | `session_id`, `key` (`C-c`/`Enter`/`Up`…) |
|
|
424
426
|
| `pty_close` | Close idempotently; return `closed` / `already_closed` | `session_id` |
|
|
425
|
-
| `pty_list` | List sessions (agent rows carry `agent=<kind>`
|
|
426
|
-
| `
|
|
427
|
+
| `pty_list` | List sessions (agent rows carry canonical `harness=<id>` plus compatibility `agent=<kind>`) | (none) |
|
|
428
|
+
| `agent_launch` | Canonical agent launch; harness and model are independent | `harness`, `prompt?`, `model?`, `reasoning_effort?`, `cwd?`, `write_scope?` |
|
|
429
|
+
| `claude_agent` / `codex_agent` / `grok_agent` / `composer_agent` | Deprecated compatibility aliases | legacy launcher arguments |
|
|
430
|
+
| `agent_configure` | Change model/effort in a running Claude, Codex, Grok, Composer, or Cursor session without restarting it | `session_id`, `model?`, `reasoning_effort?` |
|
|
427
431
|
| `claude_turn` | Issue (dispatch-only) or recover one correlated Claude operation | `action`, `session_id`, `operation_id`, `text?` |
|
|
428
432
|
| `claude_approval` | Inspect or answer the current correlated Claude approval prompt | `action`, `session_id`, `operation_id?`, `approval_choice?`, `observed_prompt_digest?` |
|
|
429
433
|
| `diagnostics` | Read-only factory readiness as machine-readable JSON | (none) |
|
|
@@ -436,20 +440,20 @@ On top of that sits a productized layer a raw tmux bridge doesn't have: **token-
|
|
|
436
440
|
|
|
437
441
|
Consumer flow is `aiterm-runtime-errors snapshot`, then `aiterm-runtime-errors ack --cursor N` after durable ingestion. Operators can use `resolve|reopen --fingerprint SHA256`. MCP collection and diagnostic reads run in timeout-bounded child processes, so a FIFO or stalled filesystem cannot block terminal work; child failure emits only the fixed store diagnostic. Store mutation uses a bounded bakery ticket queue: every waiter owns a never-reused ticket containing PID, process-start identity, and an owner token, so dead owners are removed by unique filename without fixed-path reclaim ABA. The queue deadline measures lack of progress by the same head owner, not total wait behind healthy predecessors; normal polling uses the native process-liveness check and validates process-start identity only when a blocker stalls. Worker deadlines use forced termination so a SIGTERM-ignoring child cannot mutate state after timeout. POSIX state is atomically replaced under `$XDG_STATE_HOME/aiterm-mcp/` (default `~/.local/state/aiterm-mcp/`) with owner/mode rechecked on every read. Windows native uses `%LOCALAPPDATA%\aiterm-mcp\`; each DACL is rebuilt and read back as one non-inherited FullControl ACE for the current SID. Windows path/DACL/timeout behavior is covered by pure tests in this change; no new Windows integration success is claimed.
|
|
438
442
|
|
|
439
|
-
### Interactive agent
|
|
443
|
+
### Interactive agent harnesses
|
|
440
444
|
|
|
441
|
-
|
|
445
|
+
`agent_launch` starts a selected harness's interactive coding-agent TUI inside a fresh persistent PTY and returns its `session_id`. The harness owns the agent loop, authentication, hooks, session, and transcript; `model` is independent. The TUI is a full-screen app, so read it with `pty_read({ screen: true })` for the rendered view.
|
|
442
446
|
|
|
443
|
-
`agent_configure({ session_id, model?, reasoning_effort? })` changes a running Claude, Codex, Grok, or
|
|
447
|
+
`agent_configure({ session_id, model?, reasoning_effort? })` changes a running Claude, Codex, Grok, Composer, or Cursor TUI through the harness's standard controls, preserving the PTY and conversation context.
|
|
444
448
|
|
|
445
|
-
|
|
|
449
|
+
| `harness` | Launches | Model behavior |
|
|
446
450
|
| --- | --- | --- |
|
|
447
|
-
| `
|
|
448
|
-
| `
|
|
449
|
-
| `
|
|
450
|
-
| `
|
|
451
|
+
| `claude-code` | Claude Code CLI | Claude catalog model; native effort controls |
|
|
452
|
+
| `codex-cli` | Codex CLI | OpenAI catalog model; native effort controls |
|
|
453
|
+
| `grok-cli` | Grok Build CLI | Grok/Composer catalog model; Composer is `model: "grok-composer-2.5-fast"` |
|
|
454
|
+
| `cursor-cli` | Cursor Agent CLI | Cursor catalog model, including GPT/Claude/Grok; effort uses model parameter override |
|
|
451
455
|
|
|
452
|
-
The
|
|
456
|
+
The selected harness CLI must be installed and authenticated. Use each product owner's official installer and updater; Aiterm does not distribute alternate CLI tarballs. For Cursor Agent CLI, use `curl https://cursor.com/install -fsS | bash` on macOS/Linux/WSL or `irm 'https://cursor.com/install?win32=true' | iex` on native Windows, authenticate once with `agent login`, and update with `agent update`; Aiterm invokes the unambiguous `cursor-agent` binary. Missing binaries, invalid model/effort values, unavailable Grok catalog models, and nonexistent `cwd` fail before a session exists.
|
|
453
457
|
|
|
454
458
|
Set `throughline_source_session` together with a non-empty mission in `prompt` to prepend
|
|
455
459
|
Throughline's read-only handoff context. This optional route requires `throughline >= 0.9.0`,
|
|
@@ -467,7 +471,7 @@ When an agent's answer is longer than the on-screen tail (pane height ≈ 24 lin
|
|
|
467
471
|
|
|
468
472
|
As of v0.16 a parent agent **never blocks** on aiterm — there is no wait parameter anywhere (v0.17 makes the waiter's exit codes mirror its outcome). The whole flow is dispatch + one universal waiter:
|
|
469
473
|
|
|
470
|
-
1. Launch the child (
|
|
474
|
+
1. Launch the child with `agent_launch({ harness: ... })`; every launch shares the normal project/user environment and adds only completion correlation plus lineage. Send a turn with plain `pty_send` (or `claude_turn issue` for durable Claude operations). The call returns immediately with an `event_cursor` in its structured receipt.
|
|
471
475
|
2. Run `aiterm-wait --session <id> --cursor <event_cursor> [--operation sha256:<64hex>] [--timeout <sec>]`. It observes the vendor-native completion source, plus Claude's additive launch hook, as a **pure reader** and exits with a one-line `aiterm.agent-wait-result.v1` receipt. **Exit ≠ done**: the receipt's `outcome` is authoritative (`0` = `done`, `3` = `timeout`, `4` = `closed`, `1` = error).
|
|
472
476
|
3. **The parent never runs the waiter in its own foreground.** Waiting is correct — but the waiter is a separate process, not the parent's turn. A harness that re-invokes its agent when a background task exits (Claude Code) runs the waiter **in the background** and gets woken with zero polling. So that this is not left to interpretation, aiterm reads `clientInfo.name` from the MCP `initialize` handshake and its receipts name the concrete invocation for the detected host — for Claude Code, literally `Bash(command: "aiterm-wait …", run_in_background: true)`. Unknown or undeclared hosts get the generic "start it as a process that does not block the parent's turn" wording; nothing else about the contract changes. Every receipt leads with the same rule: dispatch and let go, then go do something else or end the turn.
|
|
473
477
|
4. Collect the result exactly as before: `pty_read(agent_transcript: true)`, or `claude_turn recover` for durable Claude operations. The waiter carries the signal, never the payload.
|
|
@@ -490,7 +494,7 @@ Each `pty_send` accepts at most 64 KiB of UTF-8 text. Sends to the same session
|
|
|
490
494
|
|
|
491
495
|
## A human can watch
|
|
492
496
|
|
|
493
|
-
Sessions live on a shared tmux socket. The `tmux -S … attach -t <id>` line printed by `pty_open`
|
|
497
|
+
Sessions live on a shared tmux socket. The `tmux -S … attach -t <id>` line printed by `pty_open` and `agent_launch` lets a human attach to the same terminal and intervene (`Ctrl-b d` to detach), including a Claude/Codex/Grok/Cursor harness session. On native Windows the printed line is the psmux form — `psmux -L <namespace> attach -t <id>` — pointing at the same Windows-native session.
|
|
494
498
|
|
|
495
499
|
## Requirements
|
|
496
500
|
|
|
@@ -498,7 +502,7 @@ Sessions live on a shared tmux socket. The `tmux -S … attach -t <id>` line pri
|
|
|
498
502
|
- **tmux** (runtime prerequisite; check with `tmux -V`. Install with `apt install tmux` / `brew install tmux`)
|
|
499
503
|
- **macOS / Linux / WSL2** run tmux directly. On macOS install it with `brew install tmux` (stock macOS ships none). If your MCP client is launched from the **GUI** rather than a terminal, Homebrew's bin (`/opt/homebrew/bin` on Apple Silicon, `/usr/local/bin` on Intel) may be off its `PATH`; aiterm auto-searches those locations, or set **`AITERM_TMUX=/path/to/tmux`** to point at it explicitly.
|
|
500
504
|
- **Native Windows** has no tmux, so aiterm drives [psmux](https://github.com/psmux/psmux) — a tmux-CLI-compatible native multiplexer — with a per-install `-L` namespace. **No WSL is required.** Install psmux **3.3.8 or newer** (`winget install marlocarlo.psmux`; 3.3.8 is the first release whose `pipe-pane` file sink, byte-exact `paste-buffer` wire, and foreground `#{pane_current_command}` behave the way aiterm's capture/dispatch paths rely on), plus [Git for Windows](https://gitforwindows.org/) whose `bash.exe` becomes the pane shell (System32's `bash.exe` is the WSL launcher and is deliberately not used). Override resolution with **`AITERM_PSMUX`** / **`AITERM_BASH`** when the binaries live elsewhere. You reach Windows tools the same way you reach SSH: `pty_send "powershell.exe …"` nests into PowerShell. `grok_agent`/`composer_agent` launch the **Windows-native** Grok CLI (`%USERPROFILE%\.grok\bin\grok.exe`, or `GROK_BIN` pointing at a `.exe`) as a Windows process, and a WSL-side grok is rejected before a session is created so vendor auth and session records never split across an OS boundary.
|
|
501
|
-
- For
|
|
505
|
+
- For **agent harnesses**: the selected CLI, installed and authenticated through its product owner's official path — `claude`, `codex`, `grok`, or Cursor's `cursor-agent`. Portable fork additionally needs `throughline >= 0.9.0`; ordinary clean launch does not. (Not needed if you only use the PTY tools.)
|
|
502
506
|
- Optional: the [`rtk`](https://github.com/rtk-ai/rtk) binary (used by `pty_send`'s `rtk: true` delegation; works fine without it)
|
|
503
507
|
|
|
504
508
|
## Known constraints (by design, not bugs)
|
|
@@ -506,7 +510,7 @@ Sessions live on a shared tmux socket. The `tmux -S … attach -t <id>` line pri
|
|
|
506
510
|
- **While nested (ssh / docker / REPL / a launched agent TUI), quiescence cannot fire by design**, because the foreground command is no longer in the shell set (bash/sh/zsh/fish/dash). When nested with no `until` and no `mark`, `pty_read({ wait: true })` returns early as `is_complete=False via nested` (rather than burning the full `timeout`, since no signal can confirm completion there) with a note to pass `until` (a literal substring by default; `until_regex: true` for a regex) or `mark: true` (an exit-code sentinel, auto-detected) for a confirmed completion. For a full-screen agent TUI, read `{ screen: true }` once its output settles.
|
|
507
511
|
- **`is_complete=False` is not a failure.** It means "completion was not observed within `timeout`." For long commands, raise `timeout` or use `until`/`mark`.
|
|
508
512
|
- **The destructive gate is a tripwire, not a sandbox.** It blocks common destructive forms only. It does **not** catch relative-path `rm`, things that become dangerous after `$VAR` expansion, or commands run on the far side of an SSH session — and it does not police what a launched coding agent does inside its own session.
|
|
509
|
-
- **
|
|
513
|
+
- **Agent harnesses run their real TUI; aiterm doesn't proxy the model API.** The selected harness owns model choice, authentication, and behavior. There is no hidden inter-agent protocol; the MCP client drives the Claude/Codex/Grok/Cursor TUI with ordinary send/read operations.
|
|
510
514
|
- **`pty_send({ rtk: true })` is single-line only and needs the external `rtk` binary** (passthrough without it). The `pty_read({ rtk: true })` reducer, by contrast, is self-contained and rtk-independent.
|
|
511
515
|
- **The `pytest` reducer matches rtk 0.42.0** on test counts, the rule line, and `FAILURES`-block formatting (locked by regression tests). It **deliberately preserves the full failure reason** on the `FAILED` summary lines (emitted under `-ra`/`-rf`), whereas rtk 0.42.0 truncates the reason at the first `" - "` — a readability choice, so those lines are intentionally not byte-identical to rtk. The `[full output: …]` tee-pointer line rtk appends on large output is not reproduced on the read side.
|
|
512
516
|
- **tmux is started with `-f /dev/null`**, so it does not read `~/.tmux.conf` (to keep behavior reproducible across machines).
|
|
@@ -547,13 +551,13 @@ If aiterm let your AI hand a task to another agent — or saved you a round-trip
|
|
|
547
551
|
|
|
548
552
|
## Shared agent environment
|
|
549
553
|
|
|
550
|
-
All
|
|
554
|
+
All harnesses use the caller's normal project and user environment. Aiterm does not copy,
|
|
551
555
|
symlink, filter, or replace vendor configuration, authentication, MCP, plugin, skill, permission,
|
|
552
556
|
trust, memory, or history stores. Cleanup removes only aiterm-owned launch metadata and completion
|
|
553
557
|
correlation files.
|
|
554
558
|
|
|
555
559
|
The ordinary environment still comes from the shell/tmux session. When a caller needs a value that
|
|
556
|
-
belongs to the current MCP process rather than the older persistent tmux server, every
|
|
560
|
+
belongs to the current MCP process rather than the older persistent tmux server, every harness
|
|
557
561
|
accepts `env_vars: ["NAME", ...]`. Only those names are refreshed at launch; this is a narrow
|
|
558
562
|
per-launch overlay, not a replacement environment or configuration snapshot.
|
|
559
563
|
|
package/dist/agent-resolver.js
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
// エージェント CLI(claude / codex / grok)・Throughline・pane shell(Windows は Git Bash)の
|
|
1
|
+
// エージェント CLI(claude / codex / grok / cursor-agent)・Throughline・pane shell(Windows は Git Bash)の
|
|
2
2
|
// 実行ファイルをどう見つけ、どう起動するかの所有者(OS 分岐の所有者。tmux とは独立)。
|
|
3
3
|
import { spawnSync } from "node:child_process";
|
|
4
4
|
import * as fs from "node:fs";
|
|
@@ -115,7 +115,9 @@ export function resolveAgentBin(kind) {
|
|
|
115
115
|
? ["CLAUDE_BIN", [".local", "bin", "claude"], "claude"]
|
|
116
116
|
: kind === "codex"
|
|
117
117
|
? ["CODEX_BIN", [".local", "bin", "codex"], "codex"]
|
|
118
|
-
:
|
|
118
|
+
: kind === "cursor"
|
|
119
|
+
? ["CURSOR_AGENT_BIN", [".local", "bin", "cursor-agent"], "cursor-agent"]
|
|
120
|
+
: ["GROK_BIN", [".grok", "bin", isWin ? "grok.exe" : "grok"], "grok"];
|
|
119
121
|
const fromEnv = process.env[envVar];
|
|
120
122
|
if (fromEnv) {
|
|
121
123
|
// 明示指定 env は実在を検証する。存在しないパスを黙って返すと、session を作って
|
|
@@ -125,9 +127,16 @@ export function resolveAgentBin(kind) {
|
|
|
125
127
|
return resolveWindowsCodexShim(kind, fromEnv);
|
|
126
128
|
throw new AitermError(`${envVar} に指定された ${name} が存在しません: ${fromEnv}`, 2);
|
|
127
129
|
}
|
|
128
|
-
const
|
|
129
|
-
if (
|
|
130
|
-
|
|
130
|
+
const defaultCandidates = [path.join(home, ...rel)];
|
|
131
|
+
if (kind === "cursor" && isWin && process.env.LOCALAPPDATA) {
|
|
132
|
+
// Cursor公式Windows installerは %LOCALAPPDATA%\cursor-agent をPATHへ追加し、
|
|
133
|
+
// cursor-agent.exe/cmdを置く。GUI hostの古いPATHでも公式配置を直接解決する。
|
|
134
|
+
defaultCandidates.unshift(path.join(process.env.LOCALAPPDATA, "cursor-agent", "cursor-agent.exe"), path.join(process.env.LOCALAPPDATA, "cursor-agent", "cursor-agent.cmd"));
|
|
135
|
+
}
|
|
136
|
+
for (const cand of defaultCandidates) {
|
|
137
|
+
if (isUsableAgentExecutableFile(cand))
|
|
138
|
+
return resolveWindowsCodexShim(kind, cand);
|
|
139
|
+
}
|
|
131
140
|
const w = spawnSync(isWin ? "where" : "which", [name], {
|
|
132
141
|
encoding: "utf8",
|
|
133
142
|
timeout: 5000,
|