aiterm-mcp 0.27.9 → 0.28.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.ja.md CHANGED
@@ -1,4 +1,4 @@
1
- > **任意のMCPクライアントから、Claude・Codex・Grok・Composerをクロスベンダーでも同一ベンダーでも永続対話TUIへ起動する。Codexのスラッシュコマンドや[`$imagegen`](https://learn.chatgpt.com/docs/image-generation#generate-or-edit-an-image)のような固有機能もそのまま使える。**
1
+ > **任意のMCPクライアントから、Claude Code・Codex CLI・Grok CLI・Cursor Agent CLIを単一のharness APIで永続対話TUIへ起動する。**
2
2
 
3
3
  <p align="center">
4
4
  <img src=".github/og.png" alt="Aiterm — 異なる知性が一つの持続する実行現場を共有する森の観測拠点" width="100%">
@@ -16,7 +16,7 @@
16
16
 
17
17
  > *(English: [README.md](README.md))*
18
18
 
19
- > **あなたの AI に、ほかの AI を操らせる。** 任意の MCP クライアントから 1 回の呼び出しで、コーディングエージェント(Claude・Codex・Grok・Composer)を永続端末の中に起動し、操作用のセッションを手渡す。何をしているかをトークン削減して読み、次の指示を送る。呼び出し元と起動先のベンダーは独立しており、ClaudeからClaude/Codexを、CodexからClaude/Codexを起動できる。
19
+ > **あなたの AI に、ほかの AI を操らせる。** `agent_launch`の1回の呼び出しで、実行基盤harnessとmodelを別々に選び、永続sessionを受け取る。CursorでGPT/Claude/Grokを選んでも、session・hook・transcriptはCursorが所有する。
20
20
  >
21
21
  > **これは何か:** AI が握る 1 本の永続 MCP 端末——その中に他のコーディングエージェントも起動できる。`ssh`・`docker exec`・REPL・別エージェントの TUI は、すべてその 1 本の端末の中へ「送るだけのテキスト」として入れ子になる。仕組みはあえて素朴——MCP クライアントが相手エージェントの端末を 1 ターンずつ操作するだけ。隠れたプロトコルも・aiterm独自の共有メモリ層も・自律的な交渉も無い。起動したagentは、直接CLIと同じproject/vendorの通常memory・設定を読む。
22
22
  >
@@ -94,7 +94,9 @@ host統合は、kitepon.devの製品開発を支える内部基盤
94
94
 
95
95
  **言葉でなく実測で:** 記録済み203テストのベンチマークでは、`pty_read` はコンテキストに載るトークンを生ログの **約 7.1 分の 1** に減らす。しかも pass/fail の判定は畳んでも残る。→ [組み込みシェルツールとの使い分け](#組み込みシェルツールとの使い分け)
96
96
 
97
- 14 ツール: 6 つの **PTY ツール**(`pty_open` / `pty_send` / `pty_read` / `pty_key` / `pty_close` / `pty_list`)で 1 本の永続端末を開き・操作し・読む。加えて 4 つの **エージェント起動ツール**(`claude_agent` / `codex_agent` / `grok_agent` / `composer_agent`)が別のコーディングエージェントの TUI を新しい端末の中に起動し、`agent_configure`が起動中のClaude/Codex/Grok/Composerのmodel・effortを再起動なしで変更し、`claude_turn`がdurable caller向けの構造化issue/recoveryを、`claude_approval`が相関済みClaude承認UI中継を、`diagnostics`が安全なfactory readinessを返す。バックエンドは **tmux** なので、MCP サーバや AI クライアントが再起動してもセッションは生き残る。
97
+ 15ツール: 6つのPTYツール、正規のagent起動入口`agent_launch`、移行用の旧4alias、`agent_configure`、`claude_turn`、`claude_approval`、`diagnostics`。バックエンドはtmuxなので、MCPサーバやAIクライアントが再起動してもsessionは生き残る。
98
+
99
+ **v0.28.0では実行基盤harnessとmodelを分離した。** harnessはagent loop・認証・hook・session・transcriptを所有し、modelはその上で選ぶ。Cursor Agent CLIでGPT/Claude/Grokを選んでも完了契約はCursor方式のまま。Composerは別harnessではなく、`harness:"grok-cli", model:"grok-composer-2.5-fast"`で表す。旧4起動ツールは同じ実装へ流れる互換alias。
98
100
 
99
101
  **v0.25.2ではGrok 4.6を含む同一sessionの連続設定変更を安定化。** Grok Build 1.0.3で
100
102
  `/model`の成功通知が再描画により消えても、変更前には無かった要求model/effortが常駐footerへ現れた
@@ -151,7 +153,7 @@ runtime-error store は canonical dotagents config の `collection.enabled: true
151
153
  場合だけ収集し、既定OFF、network送信は行いません。tag起点CIのnpm provenance(OIDC Trusted
152
154
  Publishing)で公開し、GitHub Release が Official MCP Registry を再登録します。
153
155
 
154
- **状態:** 開発継続中 · この分野では新参で、別の形に賭けている([既存手段との比較](#既存手段との比較)参照)· 動作対象は Linux · WSL2 · macOS · Windows ネイティブ(4 launcher と相関付き完了を含む)· MIT · [変更履歴](CHANGELOG.md)。
156
+ **状態:** 開発継続中 · この分野では新参で、別の形に賭けている([既存手段との比較](#既存手段との比較)参照)· 動作対象は Linux · WSL2 · macOS · Windows ネイティブ · MIT · [変更履歴](CHANGELOG.md)。
155
157
 
156
158
  ## なぜ今
157
159
 
@@ -174,14 +176,14 @@ pty_read(id, { wait: true }) → 削減済みの出力を読む(完了
174
176
 
175
177
  ### 2. その端末の中に他のコーディングエージェントを起動する — オーケストレーションの旗艦
176
178
 
177
- 同じ primitive が別エージェントの TUI を宿す。4 つの起動ツールが、Claude/Codex/Grok/Composer の対話 TUI を新しい永続端末の中に起動し、`session_id` を返す。起動processは直接CLIと同じproject/user環境を使い、通常config、MCP、plugin、skill、permission、trust、memory、historyをcopy・filter・置換しない。aitermが加えるのは完了相関と、`role=subagent`、親session、delegation depth、lineage、`delegation_allowed=true`を持つ非user instructionだけ。孫以降の委譲も許可され、固定depth capはない。
179
+ 同じprimitiveが別エージェントのTUIを宿す。`agent_launch`の`harness`はagent loop・認証・hook・session・transcriptを所有する実行基盤、`model`は独立した選択。起動processは直接CLIと同じproject/user環境を使い、通常config、MCP、plugin、skill、permission、trust、memory、historyをcopy・filter・置換しない。
178
180
 
179
- 既存の人間向けtextに加えて`aiterm.agent-launch-result.v1` structured receiptも返すため、durable callerは表示文字列を解析せずsession handleを取得できる。Codexは通常rollout transcriptの`task_complete`、Grok/Composerは通常session event、Claudeは通常settingsへ加算したlaunch固有Stop hookを完了正本に使う。agent sessionへの`pty_send`は非ブロックの **dispatch** になり`event_cursor`入りreceiptを即返す。完了通知は`aiterm-wait --session <id> --cursor <event_cursor>`を親のターンを塞がない別processで受ける。durable machine callerは`claude_turn`を使い、recoveryは再送せず、検証済み完了だけがexact `raw_output`を持つ。
181
+ `aiterm.agent-launch-result.v1`は正規`harness`を返し、旧`provider`は互換fieldとして残す。同じ`harness`はagent dispatch、`aiterm-wait`、`agent_configure`、`pty_list`のagent行にも載り、旧vendor/provider/agent fieldは互換用に残る。Codexは通常rollout、Grok CLIは通常session event、Claudeはlaunch固有Stop hook、Cursorは通常agent transcript末尾の`turn_ended`を完了正本に使う。`pty_send`は非ブロックdispatchで、vendor別完了境界を表すopaqueな整数`event_cursor`を返し、完了通知は`aiterm-wait`を親のターンを塞がない別processで受ける。
180
182
 
181
- `codex_agent`・`grok_agent`・`composer_agent`は任意の`write_scope`(`"read-only"`または書込み許可パスの説明)も受ける。指定値はlaunch receipt・session metadata・`pty_list`へ保存する。3 launcherすべてで`write_scope:"read-only"`は実効能力壁となり、aitermがCLIの`--sandbox read-only`を付ける。パス説明をallowlistへ変換する同等CLI引数はないため、そちらだけは`write_scope_enforcement:"declaration_only_unsupported"`を返す。`write_scope`を省略した起動は従来どおりである。
183
+ `agent_launch`は任意の`write_scope`も受ける。Codex/Grokのread-onlyは`--sandbox read-only`、Cursorは公式`--mode ask`で実効化する。path説明は同等CLI引数がないためdeclaration-only。
182
184
 
183
185
  ```text
184
- codex_agent({ session_name: "codex1", cwd: "/repo",
186
+ agent_launch({ harness: "codex-cli", session_name: "codex1", cwd: "/repo",
185
187
  prompt: "port test/legacy.py to vitest",
186
188
  model: "gpt-5.6-sol", reasoning_effort: "high",
187
189
  write_scope: "test/ only; no commit" })
@@ -192,14 +194,16 @@ $ aiterm-wait --session codex1 --cursor <event_cursor> # exit 0=done / 3=timeo
192
194
  → 操舵し、Codex の次の入力境界で返る
193
195
  ```
194
196
 
195
- モデルごとに 1 ツール=ツール名を見ればどのモデルか分かる:
197
+ 正規のharness選択肢:
196
198
 
197
- | ツール | 起動するもの | 主な引数 |
199
+ | `harness` | 起動するもの | modelの扱い |
198
200
  | --- | --- | --- |
199
- | `claude_agent` | Claude Code CLI(Anthropic) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`), `env_vars?`, `cwd?`, `session_name?` |
200
- | `codex_agent` | Codex CLI(OpenAI・端末設定/CLI既定、`model?`で上書き) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`/`ultra`), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
201
- | `grok_agent` | Grok Build(xAI、既定`grok-4.5`、`model?`で上書き) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(vendor対応値), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
202
- | `composer_agent` | Grok Build(xAI、既定`grok-composer-2.5-fast`、`model?`で上書き。全modelをlive catalog照合) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(vendor対応値), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
201
+ | `claude-code` | Claude Code CLI | Claude model/effort |
202
+ | `codex-cli` | Codex CLI | OpenAI model/effort |
203
+ | `grok-cli` | Grok Build CLI | Grok/Composer model、live catalog照合 |
204
+ | `cursor-cli` | Cursor Agent CLI | Cursor catalog上のGPT/Claude/Grok等 |
205
+
206
+ Cursorの`model`は`gpt-5.6-luna`のようなbase model、`reasoning_effort`は`high`のように別指定する。adapterは現行`model-effort` IDを`cursor-agent models`へ照合し、起動中変更はCursor標準model pickerのparameter editorを使う。不在時は別modelへfallbackしない。
203
207
 
204
208
  `env_vars`は環境変数の**名前**だけを並べるallowlistであり、name/value mapではない。aitermは
205
209
  launcher起動時に現在のMCP processから各名前を読み、存在する値をshell quoteして、その1回のvendor
@@ -208,7 +212,7 @@ launcher起動時に現在のMCP processから各名前を読み、存在する
208
212
  PTYの起動コマンドとして送られ、sessionの`.lastcmd`にも保持されるため、起動先vendorと同じOS userへ
209
213
  到達する。秘密転送路ではなく、席identityやworkflow用の非secret変数だけに使う。
210
214
 
211
- 各ベンダーのCLIが導入・認証済みであること。CLI不在・不正なmodel/effort・実在しない`cwd`はsession作成前に失敗し、残骸を残さない。ClaudeはさらにPTY作成前に同じCLIの`auth status --json`が`loggedIn:true`を返すことを要求する。4 launcherは通常のvendor credential/config storeをその場で使い、fake `HOME`、private `CODEX_HOME`/`GROK_HOME`、project/user config snapshotを作らない。Claudeだけは完了相関用Stop hook settingsを通常の`user,project,local` settingsへ加算する。Grok/Composerは画面入力欄だけでなく通常sessionの`mcp_init_completed` eventも確認してから送信し、共有MCP初期化中の早送信を防ぐ。相関付きClaudeのactive turn中はC-c以外の`pty_key`と素送信を拒否し、承認UIは`claude_approval`で単発Yes/Noだけを相関付きで中継する。
215
+ 選んだharnessのCLIを公式経路で導入・認証しておく。Cursorは`curl https://cursor.com/install -fsS | bash`、`agent login`、更新は`agent update`が公式経路で、Aitermは曖昧な`agent`でなく`cursor-agent`を起動する。CLI不在・未認証・不正引数はsession作成前に明示失敗し、別経路へfallbackしない。
212
216
 
213
217
  portable forkは任意である。`throughline_source_session`を使う場合、`prompt`は必須の新ミッションとなり、
214
218
  `launch_operation_id`とは併用できない。aitermは`THROUGHLINE_BIN`、次に`PATH`からThroughlineを解決し、
@@ -216,7 +220,7 @@ portable forkは任意である。`throughline_source_session`を使う場合、
216
220
  この経路だけ`throughline >= 0.9.0`が必要で、元sessionのDB所属は変わらない。引数省略時には
217
221
  Throughline自体が不要である。
218
222
 
219
- エージェント間の隠れたプロトコルは無い。起動したClaude/Codex/Grok/Composerは利用者がattachできるもう1本の永続sessionであり、MCPクライアントが通常のPTY操作で駆動する。
223
+ エージェント間の隠れたプロトコルは無い。起動したharnessは利用者がattachできるもう1本の永続sessionであり、MCPクライアントが通常のPTY操作で駆動する。
220
224
 
221
225
  ## デモ
222
226
 
@@ -271,7 +275,7 @@ Throughline自体が不要である。
271
275
  Claude Code を再起動して、接続を確認:
272
276
 
273
277
  ```bash
274
- /mcp # aiterm が connected・14 ツール公開、と出る
278
+ /mcp # aiterm が connected・15 ツール公開、と出る
275
279
  ```
276
280
 
277
281
  最初のセッション——4 回の呼び出しで、1 個の永続端末:
@@ -286,7 +290,7 @@ pty_close("t1") → 端末を解放
286
290
  `pty_close` は冪等で、`closed` / `already_closed` のstructured receiptを返す。
287
291
  MCP応答を失ったdurable callerも同じ`session_id`への再試行だけでclose結果を確定できる。
288
292
 
289
- これだけ。`t1` の端末は本物で永続——`ssh`・`docker exec`・REPL・起動したエージェントの TUI は、そこに住む「もの」に過ぎない。代わりにワーカーのエージェントを起動するのも 1 コール: `codex_agent()` が返す `session_id` を、同じ `pty_read` / `pty_send` で操作する。
293
+ これだけ。`t1` の端末は本物で永続——`ssh`・`docker exec`・REPL・起動したエージェントのTUIは、そこに住む「もの」に過ぎない。ワーカー起動も1コールで、`agent_launch({ harness: "codex-cli" })`が返す`session_id`を同じ`pty_read`/`pty_send`で操作する。
290
294
 
291
295
  **グローバル導入や別クライアントが良い場合は:**
292
296
 
@@ -312,11 +316,11 @@ MCP クライアントが aiterm を stdio 越しにプログラムから駆動
312
316
 
313
317
  ```mermaid
314
318
  flowchart LR
315
- AI["AI / MCP client<br/>(the orchestrator)"] -->|"pty_send · agent_configure · claude_agent · claude_turn · claude_approval · codex_agent<br/>grok_agent · composer_agent · diagnostics"| S["aiterm-mcp<br/>stdio MCP · 14 tools"]
319
+ AI["AI / MCP client<br/>(the orchestrator)"] -->|"pty_send · agent_launch · agent_configure · claude_turn · claude_approval<br/>旧launcher alias · diagnostics"| S["aiterm-mcp<br/>stdio MCP · 15 tools"]
316
320
  S -->|"pty_read<br/>token-reduced"| AI
317
321
  S -->|"tmux send-keys<br/>capture-pane"| P["persistent PTYs<br/>tmux · survive restarts"]
318
322
  P -->|"ssh · docker · repl"| R["nested<br/>remote · container · REPL"]
319
- P -->|"launches a fresh PTY per agent"| A["another coding-agent TUI<br/>Claude · Codex · Grok · Composer"]
323
+ P -->|"launches a fresh PTY per agent"| A["another coding-agent TUI<br/>Claude · Codex · Grok · Cursor"]
320
324
  ```
321
325
 
322
326
  primitive は「PTY を 1 個握る」ことだけ。それ以外——SSH・コンテナ・REPL・起動したエージェント TUI——は、永続端末の中で動く「対話的な何か」に過ぎず、同じ `pty_send` / `pty_read` で操作する。各起動ツールは自分専用の新しい PTY を開く。PTY は tmux 上にあるので、MCP サーバや AI クライアントが再起動してもセッションは生き残る。
@@ -363,7 +367,7 @@ aiterm は 2 つの系譜の交点にいる——端末を操作する MCP サ
363
367
  | --- | --- | --- | --- | --- |
364
368
  | 永続セッション | ✅ tmux・再起動を跨ぐ | ❌ 毎回新シェル | ⚠️ まちまち | ✅ tmux |
365
369
  | SSH / コンテナ / REPL | `pty_send` 1 回でネスト | 毎コマンド接続し直し | ⚠️ ツールが分かれがち | ✅ tmux(人が操作) |
366
- | 1 コールで別エージェント起動 | ✅ `codex_agent` / `grok_agent` / `composer_agent` | ❌ | ❌ | ⚠️ 人が動かす tmux に CLI + skills で参加 |
370
+ | 1 コールで別エージェント起動 | ✅ `agent_launch(harness=…)` | ❌ | ❌ | ⚠️ 人が動かす tmux に CLI + skills で参加 |
367
371
  | ヘッドレス(人が tmux に居ない) | ✅ MCP 駆動・プログラム的 | ✅ | ⚠️ まちまち | ❌ 人が tmux に居る前提 |
368
372
  | MCP ネイティブ(任意の MCP クライアント) | ✅ `claude mcp add` 1 行 | ✅ | ✅(MCP なので) | ❌ tmux 設定 + CLI + Agent Skills |
369
373
  | トークン削減読取 | ✅ コマンド別 reducer | ❌ 生出力 | ⚠️ ほぼ無し | ❌ 生 tmux |
@@ -392,8 +396,10 @@ aiterm は同じ核心の洞察——端末を出会いの場にする——を
392
396
  | `pty_read` | 出力を削減して読む(既定は増分) | `session_id`, `wait`, `until`, `until_regex`, `timeout`, `screen`, `full`, `lines`, `line_range`, `raw`, `rtk`, `agent_transcript`, `operation_id` |
393
397
  | `pty_key` | 制御キーを送る | `session_id`, `key`(`C-c`/`Enter`/`Up`…) |
394
398
  | `pty_close` | 冪等に閉じ、`closed` / `already_closed`を返す | `session_id` |
395
- | `pty_list` | セッション一覧 | (なし) |
396
- | `agent_configure` | 起動中のClaude/Codex/Grok/Composerを再起動せずmodel/effort変更 | `session_id`, `model?`, `reasoning_effort?` |
399
+ | `pty_list` | セッション一覧(agent行は正規`harness=<id>`と互換`agent=<kind>`を含む) | (なし) |
400
+ | `agent_launch` | harnessとmodelを別軸で選ぶ正規agent起動入口 | `harness`, `prompt?`, `model?`, `reasoning_effort?`, `cwd?`, `write_scope?` |
401
+ | `claude_agent` / `codex_agent` / `grok_agent` / `composer_agent` | deprecated互換alias | 旧launcher引数 |
402
+ | `agent_configure` | 起動中のClaude/Codex/Grok/Composer/Cursorを再起動せずmodel/effort変更 | `session_id`, `model?`, `reasoning_effort?` |
397
403
  | `claude_turn` | 相関済みClaude operationをdispatch(issue)または回収(recover) | `action`, `session_id`, `operation_id`, `text?` |
398
404
  | `claude_approval` | 現在表示中の相関済みClaude承認UIを検査または応答 | `action`, `session_id`, `operation_id?`, `approval_choice?`, `observed_prompt_digest?` |
399
405
  | `diagnostics` | 機械可読 JSON による read-only factory readiness | (なし) |
@@ -406,31 +412,31 @@ aiterm は同じ核心の洞察——端末を出会いの場にする——を
406
412
 
407
413
  consumer は `aiterm-runtime-errors snapshot` を読み、durable ingestion 後に `aiterm-runtime-errors ack --cursor N` を呼ぶ。運用上の明示操作は `resolve|reopen --fingerprint SHA256`。MCP からの収集・diagnostic read は timeout 付き child process に隔離し、FIFOや停止 filesystem が端末本体を止めない。store mutation は期限付き bakery ticket queue で直列化する。各waiterは PID+process start identity+owner token を持つ再利用されない固有ticketを所有するため、死んだownerだけを固有名で除去でき、固定path回収のABAを作らない。queueの期限は正常な前任者を含む総待ち時間ではなく、同じ先頭ownerが進まない時間を測る。通常pollはprocessの生存確認だけを行い、process start identityはblockerがstallした時に照合する。POSIX state は `$XDG_STATE_HOME/aiterm-mcp/`(既定 `~/.local/state/aiterm-mcp/`)へ atomic replacement で置き、every read で owner/mode を再検証する。Windows native は `%LOCALAPPDATA%\aiterm-mcp\` で current SID の非継承 FullControl ACE 1件だけへ DACL を再構築し readback する。今回 Windows は path/DACL/timeout の純粋テストだけであり、新しい実機統合成功は主張しない。
408
414
 
409
- ### 対話エージェント起動ツール
415
+ ### 対話エージェントharness
410
416
 
411
- 各ツールは特定ベンダーの対話型コーディングエージェント TUI を新しい永続 PTY の中に起動し、`session_id` を返す。以後は他のセッションと同様に `pty_read` / `pty_send` で操作する。モデルごとに 1 ツール=ツール名を見ればどのモデルか分かる。TUI は全画面アプリなので、`pty_read({ screen: true })` で描画済みの画面を読む。
417
+ `agent_launch`は選んだharnessの対話TUIを新しい永続PTYに起動し、`session_id`を返す。harnessはagent loop・認証・hook・session・transcriptを所有し、modelは独立。以後は他sessionと同じ`pty_read`/`pty_send`で操作する。
412
418
 
413
- `agent_configure({ session_id, model?, reasoning_effort? })`はvendor標準操作で起動中のClaude/Codex/Grok/Composerを変更し、PTYと会話contextを維持する。Grok/Composerは同じlive model catalog照合後に`/model <model> [effort]`または`/effort <effort>`を使う。Claude Code標準の`/model`・`/effort`は、新しいClaude sessionの既定値も同時に保存する。
419
+ `agent_configure({ session_id, model?, reasoning_effort? })`はharness標準操作で起動中のClaude/Codex/Grok/Composer/Cursorを変更し、PTYと会話contextを維持する。
414
420
 
415
- | ツール | 起動するもの | 主な引数 |
421
+ | `harness` | 起動するもの | modelの扱い |
416
422
  | --- | --- | --- |
417
- | `claude_agent` | Claude Code CLI(Anthropic) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`), `env_vars?`, `cwd?`, `session_name?` |
418
- | `codex_agent` | Codex CLI(OpenAI・端末設定/CLI既定、`model?`で上書き) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`/`ultra`), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
419
- | `grok_agent` | Grok Build(xAI、既定`grok-4.5`、`model?`で上書き) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(vendor対応値), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
420
- | `composer_agent` | Grok Build(xAI、既定`grok-composer-2.5-fast`、`model?`で上書き。全modelをlive catalog照合) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(vendor対応値), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
423
+ | `claude-code` | Claude Code CLI | Claude model/effort |
424
+ | `codex-cli` | Codex CLI | OpenAI model/effort |
425
+ | `grok-cli` | Grok Build CLI | Grok/Composer model、live catalog照合 |
426
+ | `cursor-cli` | Cursor Agent CLI | Cursor catalog上のGPT/Claude/Grok等 |
421
427
 
422
- 対応するCLI(`claude` / `codex` / `grok`)の導入・認証が必要。前提違反はsession作成前に明示失敗する。4 launcherすべてが通常project/user環境と同じ非ブロックdispatch契約を使う。Claude/Codex/Grok/Composerのdepth 1 live smokeと、Claude親→Claude孫のdepth 2 nested delegation smokeはgreenであり、fixtureによる検証とは区別して記録する。
428
+ 対応するCLI(`claude`/`codex`/`grok`/`cursor-agent`)の公式導入・認証が必要。前提違反はsession作成前に明示失敗する。全harnessが通常project/user環境と同じ非ブロックdispatch契約を使う。
423
429
 
424
430
  `throughline_source_session`と空でない新ミッション`prompt`を指定すると、Throughlineの読み取り専用
425
431
  handoff contextを前置きできる。この任意経路は`throughline >= 0.9.0`を必要とし、
426
432
  `launch_operation_id`とは併用不可で、元sessionのDB所属を変更しない。Throughlineは
427
433
  `THROUGHLINE_BIN`、次に`PATH`から解決し、不在・不正・空のexportはPTY作成前に明示失敗する。
428
434
 
429
- エージェントの回答が画面 tailより長ければ、対話callerは`pty_read({ agent_transcript:true })`で再promptなしに全文回収する。Claudeはlaunch相関付きStop hookのowner-only resultを検証し、private transcriptを読まない。Codexは通常rollout transcript、Grok/Composerは通常session historyから同じturnを回収する。不在・非agent・抽出不能は明示エラー。
435
+ エージェントの回答が画面tailより長ければ、`pty_read({ agent_transcript:true })`で再promptなしに全文回収する。Claudeはlaunch相関Stop hook、Codexは通常rollout、Grokは通常session history、Cursorはlaunch IDでbindした通常agent transcriptから同じturnを回収する。
430
436
 
431
437
  ### 完了検出(5 層)
432
438
 
433
- `pty_read({ wait: true })`は通常PTYを、process終了/`mark:true` sentinel/`until`一致/shell復帰を伴う出力静止/timeoutの5層で判定する。`mark`はPOSIX shellでは終了コード、PowerShellでは成功`0`/失敗`1`を出力する。fish/csh/tcshはどちらの状態取得構文にも従わないため送信前に拒否する。agent sessionは第6の正確な層を使う。Codexは通常rolloutの`task_complete`、Grok/Composerは通常sessionの`turn_ended`、Claudeは通常settingsへ加算したlaunch相関Stop eventを`aiterm-wait --cursor`が観測する。親はブロックもポーリングもしない。Grok/Composerは`mcp_init_completed`確認前に入力欄が見えても送信しない。
439
+ `pty_read({ wait: true })`は通常PTYを、process終了/`mark:true` sentinel/`until`一致/shell復帰を伴う出力静止/timeoutの5層で判定する。agent sessionは第6の正確な層を使い、Codexは通常rollout、Grokは通常session event、Claudeはlaunch相関Stop event、Cursorは通常agent transcriptの`turn_ended`を`aiterm-wait --cursor`が観測する。親はブロックもポーリングもしない。
434
440
 
435
441
  ### トークン削減
436
442
 
@@ -446,7 +452,7 @@ handoff contextを前置きできる。この任意経路は`throughline >= 0.9.
446
452
 
447
453
  ## 人が覗く
448
454
 
449
- セッションは共有 tmux socket(Windows nativeはpsmux namespace)上にある。`pty_open`(および各エージェント起動ツール)の戻り値に表示される `tmux -S … attach -t <id>`、Windowsでは`psmux -L … attach -t <id>`で人間が同じ端末に入って介入できる(抜けるのは `Ctrl-b d`)——起動した Claude/Codex/Grok/Composer のセッションを見たり、途中でキーボードを引き取ったりもできる。
455
+ セッションは共有tmux socket(Windows nativeはpsmux namespace)上にある。`pty_open`/`agent_launch`の戻り値に表示されるattachコマンドで、人間がClaude/Codex/Grok/Cursor harnessの同じ端末へ入り、途中でキーボードを引き取れる。
450
456
 
451
457
  ## 要件
452
458
 
@@ -454,7 +460,7 @@ handoff contextを前置きできる。この任意経路は`throughline >= 0.9.
454
460
  - **tmux または psmux**(実行時の前提)
455
461
  - **macOS / Linux / WSL2** は tmux を直接使う。macOS は同梱されないので `brew install tmux` で導入する。MCP クライアントがターミナルでなく **GUI から起動**された場合、Homebrew の bin(Apple Silicon: `/opt/homebrew/bin`、Intel: `/usr/local/bin`)が `PATH` に入らないことがある。その場合 aiterm が自動で探索するか、**`AITERM_TMUX=/path/to/tmux`** で明示指定する。
456
462
  - **Windows ネイティブ**は WSL を使わず、tmux CLI互換の [psmux](https://github.com/psmux/psmux) **3.3.8以上**を直接使う(`winget install marlocarlo.psmux`)。pane shell用に Git for Windows も必要。解決先は **`AITERM_PSMUX`**/**`AITERM_BASH`** で上書きできる。Windows toolはSSHと同じく入れ子で握れ、`pty_send "powershell.exe"`でPowerShellへ入れる。
457
- - **エージェント起動ツール**を使う場合: 対応するベンダー CLI が導入・認証済みであること——`claude_agent` は `claude`、`codex_agent` は `codex`、`grok_agent` / `composer_agent` は `grok`。portable forkだけは追加で`throughline >= 0.9.0`が必要だが、通常のclean launchには不要。(PTY ツールだけ使うなら不要。)
463
+ - **agent harness**を使う場合: 対応CLIを製品所有者の公式経路で導入・認証する。CursorはmacOS/Linux/WSLで`curl https://cursor.com/install -fsS | bash`、Windows nativeで`irm 'https://cursor.com/install?win32=true' | iex`を使い、`agent login`で認証、`agent update`で更新する。Aitermは`cursor-agent`を起動する。portable forkだけは追加で`throughline >= 0.9.0`が必要。
458
464
  - 任意: [`rtk`](https://github.com/rtk-ai/rtk) バイナリ(`pty_send` の `rtk: true` 委譲で使う。無くても動く)
459
465
 
460
466
  ## 既知の制約(バグではなく仕様)
@@ -462,7 +468,7 @@ handoff contextを前置きできる。この任意経路は`throughline >= 0.9.
462
468
  - **ネスト中(ssh / docker / REPL / 起動したエージェント TUI)は quiescence が原理的に効かない。** 前面コマンドがシェル集合(bash/sh/zsh/fish/dash)の外になるため。ネスト中で `until` も `mark` も無いときは、待っても完了を確定できる信号が無いので、`pty_read({ wait: true })` はフル `timeout` を空費せず出力静止時点で `is_complete=False via nested` と早期に返し、`until`(既定リテラル部分一致・`until_regex: true` で正規表現)か `mark: true`(終了コード付き sentinel・自動検出)の指定を促す。全画面のエージェント TUI なら、出力が落ち着いた時点で `{ screen: true }` を読む。
463
469
  - **`is_complete=False` は失敗ではない。** 「timeout 内に完了を観測できなかった」という意味。長時間コマンドでは `timeout` を伸ばすか `until`/`mark` を使う。
464
470
  - **破壊ゲートはサンドボックスではなく tripwire。** よくある破壊形だけを弾く。相対パスの `rm`、`$VAR` 展開後に危険化するもの、ssh 先で実行されるコマンドは捕捉しない——起動したコーディングエージェントが自分のセッション内で何をするかも取り締まらない。
465
- - **エージェント起動ツールはベンダー TUI を起動するだけで、包んだり代理したりしない。** aiterm は前提を検証して CLI を永続 PTY で起動する——モデル・認証・挙動はベンダー CLI のもの。エージェント間の隠れたプロトコルはなく、「会話」とはMCPクライアントがClaude/Codex/Grok/Composer TUIへ入力を送り出力を読むことだ。
471
+ - **agent harnessは実物TUIを起動し、model APIを代理しない。** model・認証・挙動は選んだharnessのもの。隠れたagent間protocolはなく、MCPクライアントがClaude/Codex/Grok/Cursor TUIへ入力を送り出力を読む。
466
472
  - **`pty_send({ rtk: true })` は単行コマンドのみ+外部 `rtk` バイナリが必要**(無ければ素通し)。一方 `pty_read({ rtk: true })` の reducer は自前実装で rtk 非依存。
467
473
  - **`pytest` reducer は件数・罫線・`FAILURES` ブロック整形が rtk 0.42.0 と byte 一致**(回帰テストで固定)。ただし `-ra`/`-rf` 時の `FAILED` 要約行の理由は**全文を保持する**(rtk 0.42.0 は最初の `" - "` 区切りで切るが、本実装は可読性優先で情報を残すため、この行は意図的に rtk と完全一致させない)。rtk が大出力時に付ける `[full output: …]`(tee ポインタ)行は read 側では再現しない。
468
474
  - **tmux は `-f /dev/null` 起動**なので `~/.tmux.conf` を読まない(環境差を排除するため)。
package/README.md CHANGED
@@ -1,4 +1,4 @@
1
- > **From any MCP client, launch Claude, Codex, Grok, or Composer — cross-vendor or same-vendor — inside a persistent interactive TUI, with native features such as Codex slash commands and [`$imagegen`](https://learn.chatgpt.com/docs/image-generation#generate-or-edit-an-image) available.**
1
+ > **From any MCP client, launch Claude Code, Codex CLI, Grok CLI, or Cursor Agent CLI through one harness API inside a persistent interactive TUI.**
2
2
 
3
3
  <p align="center">
4
4
  <img src=".github/og.png" alt="Aiterm — a shared forest observatory where different intelligences work in one persistent execution space" width="100%">
@@ -16,7 +16,7 @@
16
16
 
17
17
  > *(日本語: [README.ja.md](README.ja.md))*
18
18
 
19
- > **Let your AI orchestrate other AIs.** From any MCP client, one call spawns a coding agent (Claude, Codex, Grok, or Composer) inside a persistent terminal and hands you a session to drive: read what it's doing token-reduced, send it the next instruction. The caller and launched vendor are independent: Claude can launch Claude or Codex, and Codex can launch Claude or Codex.
19
+ > **Let your AI orchestrate other AIs.** One `agent_launch` call selects the execution harness separately from its model and hands you a persistent session to drive. Cursor can run GPT, Claude, or Grok while Cursor still owns the session, hooks, and transcript.
20
20
  >
21
21
  > **What it is:** one persistent MCP terminal your AI drives — and can launch other coding agents into. `ssh`, `docker exec`, a REPL, or another agent's TUI all nest inside that one terminal as just text you send in. The mechanism is deliberately plain — your MCP client drives the other agent's terminal turn by turn: no hidden protocol, no separate aiterm-owned shared-memory layer, no autonomous negotiation. Launched agents still read the normal project and vendor memory/configuration that a direct CLI launch would use.
22
22
  >
@@ -94,7 +94,9 @@ toolchain behind kitepon.dev's products.
94
94
 
95
95
  **Measured, not claimed:** in the recorded 203-test benchmark, a `pty_read` puts **~7.1× fewer tokens** in your context than the raw log — and the pass/fail verdict survives the fold. → [When to reach for it vs. the built-in shell](#when-to-reach-for-it-vs-the-built-in-shell)
96
96
 
97
- Fourteen tools: six **PTY tools** — `pty_open` / `pty_send` / `pty_read` / `pty_key` / `pty_close` / `pty_list` — to open, drive, and read one persistent terminal, four **agent launchers** — `claude_agent` / `codex_agent` / `grok_agent` / `composer_agent` — that each start another coding agent's TUI inside a fresh one, `agent_configure` to change a running Claude/Codex/Grok/Composer session's model and effort without restarting it, `claude_turn` for durable structured issue/recovery, `claude_approval` for correlated Claude approval prompts, and `diagnostics` for safe factory readiness. The backend is **tmux**, so sessions survive even if the MCP server or the AI client restarts.
97
+ Fifteen tools: six **PTY tools** — `pty_open` / `pty_send` / `pty_read` / `pty_key` / `pty_close` / `pty_list` — to open, drive, and read one persistent terminal; one canonical **agent launcher**, `agent_launch`, which selects `claude-code`, `codex-cli`, `grok-cli`, or `cursor-cli` as the execution harness; four deprecated launcher aliases kept for migration; `agent_configure`; `claude_turn`; `claude_approval`; and `diagnostics`. The backend is **tmux**, so sessions survive even if the MCP server or the AI client restarts.
98
+
99
+ **v0.28.0 separates the execution harness from the model.** The harness owns the agent loop, authentication, hooks, session, and transcript; `model` is what that harness runs. Cursor Agent CLI can therefore select GPT, Claude, or Grok without changing the completion contract from Cursor hooks to another vendor's. Grok Composer is a Grok CLI model preset, not another harness: use `harness: "grok-cli", model: "grok-composer-2.5-fast"`. The old four launcher tools are thin compatibility aliases over the same implementation.
98
100
 
99
101
  **v0.25.2 stabilizes repeated in-place configuration changes, including Grok 4.6.** If Grok Build
100
102
  1.0.3 redraws before its `/model` success notice can be observed, aiterm confirms the requested model/effort
@@ -165,7 +167,7 @@ collection is off by default and performs no network I/O. It ships via
165
167
  tag-triggered CI with npm provenance (OIDC Trusted Publishing); the GitHub
166
168
  Release re-registers the Official MCP Registry entry.
167
169
 
168
- **Status:** actively maintained · the newcomer here, betting on a different shape (see [vs. the alternatives](#vs-the-alternatives)) · runs on Linux · WSL2 · macOS · native Windows (tmux on POSIX, the tmux-CLI-compatible [psmux](https://github.com/psmux/psmux) on native Windows — all four launchers and correlated completion included, no WSL required) · MIT · see the [CHANGELOG](CHANGELOG.md).
170
+ **Status:** actively maintained · the newcomer here, betting on a different shape (see [vs. the alternatives](#vs-the-alternatives)) · runs on Linux · WSL2 · macOS · native Windows (tmux on POSIX, the tmux-CLI-compatible [psmux](https://github.com/psmux/psmux) on native Windows — no WSL required) · MIT · see the [CHANGELOG](CHANGELOG.md).
169
171
 
170
172
  ## Why now
171
173
 
@@ -194,16 +196,16 @@ pty_read(id, { wait: true }) → read the token-reduced output, completion
194
196
 
195
197
  ### 2. Launch other coding agents into that terminal — the orchestration flagship
196
198
 
197
- The same primitive hosts another agent's TUI. Four launchers each start one vendor's interactive coding-agent TUI inside a fresh persistent terminal and return a `session_id`. The launched process sees the same project and user environment as a direct CLI invocation: normal configuration, MCPs, plugins, skills, permissions, trust decisions, memory, and history are not copied, filtered, or replaced. Aiterm adds only completion correlation and a non-user sub-agent context containing `role=subagent`, the parent session, delegation depth, lineage, and `delegation_allowed=true`. Nested delegation is supported: a grandchild receives depth 2 and the extended lineage rather than being forbidden from launching another agent.
199
+ The same primitive hosts another agent's TUI. `agent_launch` starts a selected execution harness inside a fresh persistent terminal and returns a `session_id`. `harness` names the component that owns the agent loop, authentication, hooks, session, and transcript; `model` remains an independent choice. The launched process sees the same project and user environment as a direct CLI invocation: normal configuration, MCPs, plugins, skills, permissions, trust decisions, memory, and history are not copied, filtered, or replaced. Aiterm adds only completion correlation and a non-user sub-agent context containing `role=subagent`, the parent session, delegation depth, lineage, and `delegation_allowed=true`.
198
200
 
199
- The human-readable launch text is accompanied by an `aiterm.agent-launch-result.v1` structured receipt, so durable callers never parse display text for the session handle; when the launch carries an initial `prompt`, the receipt also includes the `event_cursor`, a ready-made `wait_command` for the completion waiter, and a `submit_residue` observation (`true` = the prompt is likely still sitting unsubmitted in the composer — the hint explains recovery; `false` = no residue observed, not a proof of submission; `null` = not applicable). From there you drive it with the same `pty_read` / `pty_send` you'd use on any shell. Codex completion comes from its normal durable rollout transcript's `task_complete`; Grok/Composer use their normal session events; Claude receives a launch-specific Stop hook settings addition while still loading normal user/project/local settings. Sending to an agent session is a non-blocking **dispatch** — the call returns immediately with an `event_cursor`, and completion arrives via [`aiterm-wait`](#completion-push-for-parent-agents-aiterm-wait). Durable machine callers use `claude_turn({ action: "issue" | "recover", session_id, operation_id, ... })`: it returns fixed `accepted` / `pending` / `completed` / `unknown` states without parsing human-facing errors, never resends during recovery, and includes exact `raw_output` only for a verified completion. An initial `prompt` on `claude_agent`/`codex_agent` is submitted through the same ready gate and the launcher returns without waiting; on Grok/Composer it is passed on the CLI's argv. This needs the vendor's own CLI installed and authenticated — see [Requirements](#requirements).
201
+ The human-readable launch text is accompanied by an `aiterm.agent-launch-result.v1` structured receipt containing the canonical `harness`; the old `provider` field remains for compatibility. The same `harness` is carried by agent dispatch, `aiterm-wait`, `agent_configure`, and agent rows in `pty_list`, while their old vendor/provider/agent fields remain compatibility fields. Codex completion comes from its normal durable rollout transcript, Grok CLI from its normal session events, Claude Code from a launch-specific Stop hook settings addition, and Cursor from its normal agent transcript's terminal `turn_ended` record. Sending to any agent session is a non-blocking **dispatch** — the call returns immediately with an opaque, harness-specific integer `event_cursor`, and completion arrives via [`aiterm-wait`](#completion-push-for-parent-agents-aiterm-wait).
200
202
 
201
- `codex_agent`, `grok_agent`, and `composer_agent` also accept an optional `write_scope`: either `"read-only"` or a human-readable description of writable paths. A supplied value is retained in the launch receipt, session metadata, and `pty_list`. For all three launchers, `write_scope: "read-only"` is an effective boundary: aiterm adds the CLI's `--sandbox read-only` flag. A path description remains declaration-only because none of these CLI launch surfaces provides an equivalent path allowlist flag; that case returns `write_scope_enforcement: "declaration_only_unsupported"`. Omitting `write_scope` preserves prior behavior.
203
+ `agent_launch` accepts an optional `write_scope`: either `"read-only"` or a human-readable description of writable paths. Codex/Grok use `--sandbox read-only`; Cursor uses its official read-only `--mode ask`. A path description remains declaration-only because these CLI launch surfaces provide no equivalent path allowlist flag.
202
204
 
203
205
  For a correlated Claude turn stopped at `Do you want to proceed?`, use `claude_approval(action: "inspect", ...)` to capture the active operation and SHA-256 screen digest, review the displayed command, then call `respond` with that exact digest and either `approve_once` or `deny`. The relay rechecks the operation and screen under the send lock, never exposes arbitrary input or permanent approval, keeps the active marker intact, and records a prompt-free owner-only receipt. `pty_send(force: true)` does not bypass this boundary.
204
206
 
205
207
  ```text
206
- codex_agent({ session_name: "codex1", cwd: "/repo",
208
+ agent_launch({ harness: "codex-cli", session_name: "codex1", cwd: "/repo",
207
209
  prompt: "port test/legacy.py to vitest",
208
210
  model: "gpt-5.6-sol", reasoning_effort: "high",
209
211
  write_scope: "test/ only; no commit" })
@@ -215,14 +217,14 @@ $ aiterm-wait --session codex1 --cursor <event_cursor> # never in the parent's
215
217
  pty_read("codex1", { agent_transcript: true }) → collect the full answer
216
218
  ```
217
219
 
218
- One call per model, so the tool name itself tells you which model you get:
220
+ The canonical harness choices are:
219
221
 
220
- | Tool | Launches | Key args |
222
+ | `harness` | Launches | Notes |
221
223
  | --- | --- | --- |
222
- | `claude_agent` | Claude Code CLI (Anthropic) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`), `env_vars?`, `cwd?`, `session_name?`, `launch_operation_id?` |
223
- | `codex_agent` | Codex CLI (OpenAI; terminal config/CLI default unless overridden) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`/`ultra`; ultra enables proactive automatic delegation), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
224
- | `grok_agent` | Grok Build, model `grok-4.6` by default (`model?` overrides) (xAI) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (vendor-supported value), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
225
- | `composer_agent` | Grok Build, model `grok-composer-2.5-fast` by default (`model?` overrides); every model is live-catalog checked (xAI) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (vendor-supported value), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
224
+ | `claude-code` | Claude Code CLI | Claude model and effort controls; correlated Stop hook |
225
+ | `codex-cli` | Codex CLI | OpenAI model and effort controls; durable rollout completion |
226
+ | `grok-cli` | Grok Build CLI | Grok or Composer model selected with `model`; live catalog check |
227
+ | `cursor-cli` | Cursor Agent CLI | GPT, Claude, Grok, or another Cursor catalog model; normal transcript completion |
226
228
 
227
229
  `env_vars` is an allowlist of environment-variable **names**, not a name/value map. At launch,
228
230
  aiterm reads each valid name from its current MCP process, shell-quotes present values, and places
@@ -233,7 +235,7 @@ PTY launch command and retained in aiterm's per-session `.lastcmd`; the launched
233
235
  processes with access to the same OS user may read them. Use this for non-secret seat identity and
234
236
  workflow variables, not as a secret transport.
235
237
 
236
- The vendor CLI must be installed and authenticated (`claude` for `claude_agent`; `codex` for `codex_agent`; `grok` for both Grok tools). aiterm resolves the binary via `CLAUDE_BIN` / `CODEX_BIN` / `GROK_BIN`, then each documented default location, then `PATH`. Prerequisites are checked **before** a session exists: empty `model` values and unsupported effort values are rejected up front; a missing CLI binary or a nonexistent `cwd` fails for all four. Before creating a Claude session, aiterm also requires a successful structured `claude auth status --json` result with `loggedIn: true`; unavailable, malformed, or failed authentication leaves **zero leftover session**. All launchers use the normal vendor-owned credential and configuration stores in place. No launcher creates a fake `HOME`, a private `CODEX_HOME`/`GROK_HOME`, or a snapshot of project/user configuration.
238
+ The selected harness CLI must be installed and authenticated. Aiterm resolves `CLAUDE_BIN` / `CODEX_BIN` / `GROK_BIN` / `CURSOR_AGENT_BIN`, then the documented default binary, then `PATH`. Cursor resolution deliberately uses `cursor-agent`, never the ambiguous `agent` name. Claude and Cursor authentication are checked before a PTY exists, so a failed preflight leaves no session. All harnesses use their normal vendor-owned credential and configuration stores in place.
237
239
 
238
240
  Portable fork is optional. When `throughline_source_session` is present, `prompt` is the required
239
241
  new mission and `launch_operation_id` cannot be combined with it. aiterm resolves Throughline via
@@ -242,9 +244,9 @@ places its returned context before a fixed separator and the mission. This route
242
244
  `throughline >= 0.9.0`; it reads the source memory without changing that database's session
243
245
  ownership. No Throughline dependency is needed when the field is omitted.
244
246
 
245
- All four launchers forward `model` and `reasoning_effort` through public CLI flags. Explicit Grok/Composer models and Composer's default model are checked against the current `grok models` catalog before a PTY exists; missing models are errors, with no cache, retry, or fallback to another model. Pass an absolute path for `cwd` — `~` is not expanded. Durable callers can make a promptless Claude launch exactly replayable by passing an explicit `session_name` and `launch_operation_id`. Claude adds a launch-local settings file only for the correlated Stop hook and loads it together with normal `user,project,local` setting sources; it does not replace normal hooks, MCPs, plugins, permissions, or trust state. The hook event contains no answer body, and the bounded owner-only result is returned by `pty_read({ agent_transcript:true })` without reading Claude's private transcript. While a correlated Claude turn is active, raw sends and non-interrupt keys are rejected. Exact `/login` and `/logout` dispatches are rejected so shared authentication is repaired once in a normal terminal. Codex reads its normal rollout store; Grok/Composer read their normal session event/history files. Before dispatch, Codex waits for an idle TUI, while Grok/Composer additionally require the vendor's structured `mcp_init_completed` event so a visible input box cannot accept a prompt too early. Correlated completion works for all four vendors on Linux, WSL2, macOS, and native Windows. On native Windows the Grok/Composer launchers start only the Windows-native `grok.exe` (a WSL-side grok is rejected before a session is created, so vendor auth and session records never split across an OS boundary), and Claude/Codex completion reads the same Windows-side hook/rollout records the vendors write.
247
+ Harness adapters translate `model` and `reasoning_effort` into each CLI's public controls. Explicit Grok models are checked against `grok models`; Cursor combines a base model such as `gpt-5.6-luna` with a separate effort such as `high`, checks the resulting current catalog ID, and uses Cursor's standard model picker for in-session changes. Missing models are errors, with no cache, retry, or fallback. Claude adds only launch-local Stop-hook settings, Codex reads its normal rollout store, Grok reads its normal session event/history, and Cursor binds its normal agent transcript with the launch ID. Pass an absolute `cwd`; `~` is not expanded.
246
248
 
247
- There is no hidden protocol between agents: a launched Claude, Codex, Grok, or Composer is another user-visible persistent terminal session. The MCP client drives that TUI with ordinary PTY operations, and a human can attach to watch or take over.
249
+ There is no hidden protocol between agents: every launched harness is another user-visible persistent terminal session. The MCP client drives that TUI with ordinary PTY operations, and a human can attach to watch or take over.
248
250
 
249
251
  ## Demo
250
252
 
@@ -299,7 +301,7 @@ The only edits to the captures above are the two `⋮` lines (a long head/tail r
299
301
  Restart Claude Code, then verify the connection:
300
302
 
301
303
  ```bash
302
- /mcp # aiterm should show as connected, exposing 14 tools
304
+ /mcp # aiterm should show as connected, exposing 15 tools
303
305
  ```
304
306
 
305
307
  Your first session — four calls, one persistent terminal:
@@ -314,7 +316,7 @@ pty_close("t1") → terminal released
314
316
  `pty_close` is idempotent and returns a structured `closed` / `already_closed`
315
317
  receipt, so durable callers can retry the same `session_id` after losing the MCP response.
316
318
 
317
- That's it. The terminal in `t1` is real and persistent — `ssh`, `docker exec`, a REPL, or a launched agent's TUI are just things that live inside it. To launch a worker agent instead, one call does it: `codex_agent()` returns a `session_id` you drive with the same `pty_read` / `pty_send`.
319
+ That's it. The terminal in `t1` is real and persistent — `ssh`, `docker exec`, a REPL, or a launched agent's TUI are just things that live inside it. To launch a worker agent instead, one call does it: `agent_launch({ harness: "codex-cli" })` returns a `session_id` you drive with the same `pty_read` / `pty_send`.
318
320
 
319
321
  **Prefer a global install, or a different client?**
320
322
 
@@ -340,11 +342,11 @@ The terminal is real and shared, so a human *can* jump in ([A human can watch](#
340
342
 
341
343
  ```mermaid
342
344
  flowchart LR
343
- AI["AI / MCP client<br/>(the orchestrator)"] -->|"pty_send · agent_configure · claude_agent · claude_turn · claude_approval · codex_agent<br/>grok_agent · composer_agent · diagnostics"| S["aiterm-mcp<br/>stdio MCP · 14 tools"]
345
+ AI["AI / MCP client<br/>(the orchestrator)"] -->|"pty_send · agent_launch · agent_configure · claude_turn · claude_approval<br/>legacy launcher aliases · diagnostics"| S["aiterm-mcp<br/>stdio MCP · 15 tools"]
344
346
  S -->|"pty_read<br/>token-reduced"| AI
345
347
  S -->|"tmux send-keys<br/>capture-pane"| P["persistent PTYs<br/>tmux · survive restarts"]
346
348
  P -->|"ssh · docker · repl"| R["nested<br/>remote · container · REPL"]
347
- P -->|"launches a fresh PTY per agent"| A["another coding-agent TUI<br/>Claude · Codex · Grok · Composer"]
349
+ P -->|"launches a fresh PTY per agent"| A["another coding-agent harness<br/>Claude Code · Codex CLI · Grok CLI · Cursor CLI"]
348
350
  ```
349
351
 
350
352
  One PTY is the only primitive. Everything else — SSH, containers, REPLs, and the launched agent TUIs — is just something interactive running inside a persistent terminal, driven with the same `pty_send` / `pty_read`. Each launcher opens its own fresh PTY. Because the PTYs live in tmux, sessions outlive the MCP server and the AI client.
@@ -393,7 +395,7 @@ aiterm sits at the intersection of two families: terminal-driving MCP servers, a
393
395
  | --- | --- | --- | --- | --- |
394
396
  | Persistent session | ✅ tmux, survives restarts | ❌ new shell every call | ⚠️ varies | ✅ tmux |
395
397
  | SSH / containers / REPLs | nest with one `pty_send` | reconnect every command | ⚠️ often separate tools | ✅ tmux (human drives) |
396
- | Launch another agent in one call | ✅ `claude_agent` / `codex_agent` / `grok_agent` / `composer_agent` | ❌ | ❌ | ⚠️ agents join a human-run tmux via a CLI + skills |
398
+ | Launch another agent in one call | ✅ `agent_launch(harness=…)` | ❌ | ❌ | ⚠️ agents join a human-run tmux via a CLI + skills |
397
399
  | Headless (no human at a tmux) | ✅ MCP-driven, programmatic | ✅ | ⚠️ varies | ❌ built around a human in the tmux |
398
400
  | MCP-native (any MCP client) | ✅ one `claude mcp add` | ✅ | ✅ (they are MCPs) | ❌ tmux config + CLI + Agent Skills |
399
401
  | Token-reduced reads | ✅ per-command reducers | ❌ raw output | ⚠️ rarely | ❌ raw tmux |
@@ -409,7 +411,7 @@ aiterm takes the same core insight — the terminal as the meeting point — and
409
411
 
410
412
  1. **Headless by construction.** Because aiterm is driven programmatically over MCP, an AI can launch and drive another agent with *no human sitting in the tmux* — from an orchestration loop, a CI step, or a cron job. The shared-tmux tools lead with a human at the keyboard (their docs center on interactive pane navigation), so unattended operation isn't their native mode; aiterm's is.
411
413
  2. **MCP-native, not a workflow you adopt.** aiterm is a stdio MCP server: one `claude mcp add` line and it works as structured tools in any MCP client that speaks stdio (tested in Claude Code; Cursor, Cline, and Claude Desktop speak the same protocol and should work the same way). It doesn't ask you to adopt a tmux config, learn pane navigation, or install skills into your setup — the client already knows how to call tools.
412
- 3. **Launching an agent is one tool call — an orchestration primitive.** `codex_agent()` spawns Codex in a persistent terminal and returns a session you drive immediately. You don't arrange panes or paste between them by hand; the launch, the steering, and the reads are all tool calls the orchestrating model can make on its own.
414
+ 3. **Launching an agent is one tool call — an orchestration primitive.** `agent_launch({ harness: "codex-cli" })` spawns Codex in a persistent terminal and returns a session you drive immediately. You don't arrange panes or paste between them by hand; the launch, the steering, and the reads are all tool calls the orchestrating model can make on its own.
413
415
 
414
416
  On top of that sits a productized layer a raw tmux bridge doesn't have: **token-reduced reads**, **5-layer completion detection**, and a **destructive-command tripwire**. None of this makes the human-in-the-tmux model wrong — it's a different, complementary bet on where the human is standing.
415
417
 
@@ -422,8 +424,10 @@ On top of that sits a productized layer a raw tmux bridge doesn't have: **token-
422
424
  | `pty_read` | Read output, token-reduced (incremental by default) | `session_id`, `wait`, `until`, `until_regex`, `timeout`, `screen`, `full`, `lines`, `line_range`, `raw`, `rtk`, `agent_transcript`, `operation_id` |
423
425
  | `pty_key` | Send a control key | `session_id`, `key` (`C-c`/`Enter`/`Up`…) |
424
426
  | `pty_close` | Close idempotently; return `closed` / `already_closed` | `session_id` |
425
- | `pty_list` | List sessions (agent rows carry `agent=<kind>` metadata) | (none) |
426
- | `agent_configure` | Change model/effort in a running Claude, Codex, Grok, or Composer session without restarting it | `session_id`, `model?`, `reasoning_effort?` |
427
+ | `pty_list` | List sessions (agent rows carry canonical `harness=<id>` plus compatibility `agent=<kind>`) | (none) |
428
+ | `agent_launch` | Canonical agent launch; harness and model are independent | `harness`, `prompt?`, `model?`, `reasoning_effort?`, `cwd?`, `write_scope?` |
429
+ | `claude_agent` / `codex_agent` / `grok_agent` / `composer_agent` | Deprecated compatibility aliases | legacy launcher arguments |
430
+ | `agent_configure` | Change model/effort in a running Claude, Codex, Grok, Composer, or Cursor session without restarting it | `session_id`, `model?`, `reasoning_effort?` |
427
431
  | `claude_turn` | Issue (dispatch-only) or recover one correlated Claude operation | `action`, `session_id`, `operation_id`, `text?` |
428
432
  | `claude_approval` | Inspect or answer the current correlated Claude approval prompt | `action`, `session_id`, `operation_id?`, `approval_choice?`, `observed_prompt_digest?` |
429
433
  | `diagnostics` | Read-only factory readiness as machine-readable JSON | (none) |
@@ -436,20 +440,20 @@ On top of that sits a productized layer a raw tmux bridge doesn't have: **token-
436
440
 
437
441
  Consumer flow is `aiterm-runtime-errors snapshot`, then `aiterm-runtime-errors ack --cursor N` after durable ingestion. Operators can use `resolve|reopen --fingerprint SHA256`. MCP collection and diagnostic reads run in timeout-bounded child processes, so a FIFO or stalled filesystem cannot block terminal work; child failure emits only the fixed store diagnostic. Store mutation uses a bounded bakery ticket queue: every waiter owns a never-reused ticket containing PID, process-start identity, and an owner token, so dead owners are removed by unique filename without fixed-path reclaim ABA. The queue deadline measures lack of progress by the same head owner, not total wait behind healthy predecessors; normal polling uses the native process-liveness check and validates process-start identity only when a blocker stalls. Worker deadlines use forced termination so a SIGTERM-ignoring child cannot mutate state after timeout. POSIX state is atomically replaced under `$XDG_STATE_HOME/aiterm-mcp/` (default `~/.local/state/aiterm-mcp/`) with owner/mode rechecked on every read. Windows native uses `%LOCALAPPDATA%\aiterm-mcp\`; each DACL is rebuilt and read back as one non-inherited FullControl ACE for the current SID. Windows path/DACL/timeout behavior is covered by pure tests in this change; no new Windows integration success is claimed.
438
442
 
439
- ### Interactive agent launchers
443
+ ### Interactive agent harnesses
440
444
 
441
- Each launcher starts a specific vendor's interactive coding-agent TUI inside a fresh persistent PTY and returns its `session_id` — from there you drive it with plain `pty_read` / `pty_send`, exactly like any other session. One tool per model, so the tool name itself tells you which model you get. The TUI is a full-screen app, so read it with `pty_read({ screen: true })` for the rendered view.
445
+ `agent_launch` starts a selected harness's interactive coding-agent TUI inside a fresh persistent PTY and returns its `session_id`. The harness owns the agent loop, authentication, hooks, session, and transcript; `model` is independent. The TUI is a full-screen app, so read it with `pty_read({ screen: true })` for the rendered view.
442
446
 
443
- `agent_configure({ session_id, model?, reasoning_effort? })` changes a running Claude, Codex, Grok, or Composer TUI through the vendor's standard controls, preserving the PTY and conversation context. Grok/Composer use `/model <model> [effort]` or `/effort <effort>` after the same live model-catalog check. Claude Code's native `/model` and `/effort` commands also save those choices as defaults for new Claude sessions.
447
+ `agent_configure({ session_id, model?, reasoning_effort? })` changes a running Claude, Codex, Grok, Composer, or Cursor TUI through the harness's standard controls, preserving the PTY and conversation context.
444
448
 
445
- | Tool | Launches | Key args |
449
+ | `harness` | Launches | Model behavior |
446
450
  | --- | --- | --- |
447
- | `claude_agent` | Claude Code CLI (Anthropic) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`), `env_vars?`, `cwd?`, `session_name?`, `launch_operation_id?` |
448
- | `codex_agent` | Codex CLI (OpenAI; terminal config/CLI default unless overridden) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`/`ultra`; ultra enables proactive automatic delegation), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
449
- | `grok_agent` | Grok Build, model `grok-4.6` by default (`model?` overrides) (xAI) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (vendor-supported value), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
450
- | `composer_agent` | Grok Build, model `grok-composer-2.5-fast` by default (`model?` overrides); every model is live-catalog checked (xAI) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (vendor-supported value), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
451
+ | `claude-code` | Claude Code CLI | Claude catalog model; native effort controls |
452
+ | `codex-cli` | Codex CLI | OpenAI catalog model; native effort controls |
453
+ | `grok-cli` | Grok Build CLI | Grok/Composer catalog model; Composer is `model: "grok-composer-2.5-fast"` |
454
+ | `cursor-cli` | Cursor Agent CLI | Cursor catalog model, including GPT/Claude/Grok; effort uses model parameter override |
451
455
 
452
- The vendor CLI must be installed and authenticated (`claude` for `claude_agent`; `codex` for `codex_agent`; `grok` for both Grok tools). Binary resolution uses `CLAUDE_BIN` / `CODEX_BIN` / `GROK_BIN`, then each documented default location, then `PATH`. Missing binaries, invalid model/effort values, unavailable Grok/Composer catalog models, and nonexistent `cwd` fail before a session is created. Claude additionally requires a structured healthy authentication status before any PTY exists, and correlated Claude sessions reject `/login` and `/logout`; repair authentication once in a normal terminal. All four launchers share the normal project/user environment and the same non-blocking dispatch contract. Claude, Codex, Grok, and Composer depth-1 live smokes and a Claude depth-2 nested-delegation smoke are green; fixture coverage remains a separate claim. Native Windows supports all four launchers with correlated completion (Claude verified with a live end-to-end launch on psmux ≥ 3.3.8; Grok launches only the Windows-native `grok.exe`).
456
+ The selected harness CLI must be installed and authenticated. Use each product owner's official installer and updater; Aiterm does not distribute alternate CLI tarballs. For Cursor Agent CLI, use `curl https://cursor.com/install -fsS | bash` on macOS/Linux/WSL or `irm 'https://cursor.com/install?win32=true' | iex` on native Windows, authenticate once with `agent login`, and update with `agent update`; Aiterm invokes the unambiguous `cursor-agent` binary. Missing binaries, invalid model/effort values, unavailable Grok catalog models, and nonexistent `cwd` fail before a session exists.
453
457
 
454
458
  Set `throughline_source_session` together with a non-empty mission in `prompt` to prepend
455
459
  Throughline's read-only handoff context. This optional route requires `throughline >= 0.9.0`,
@@ -467,7 +471,7 @@ When an agent's answer is longer than the on-screen tail (pane height ≈ 24 lin
467
471
 
468
472
  As of v0.16 a parent agent **never blocks** on aiterm — there is no wait parameter anywhere (v0.17 makes the waiter's exit codes mirror its outcome). The whole flow is dispatch + one universal waiter:
469
473
 
470
- 1. Launch the child (`claude_agent` / `codex_agent` / ...; every launch shares the normal project/user environment and adds only completion correlation plus lineage). Send a turn with plain `pty_send` (or `claude_turn issue` for durable Claude operations). The call passes the TUI ready gate, submits, and returns immediately with an `event_cursor` in its structured receipt.
474
+ 1. Launch the child with `agent_launch({ harness: ... })`; every launch shares the normal project/user environment and adds only completion correlation plus lineage. Send a turn with plain `pty_send` (or `claude_turn issue` for durable Claude operations). The call returns immediately with an `event_cursor` in its structured receipt.
471
475
  2. Run `aiterm-wait --session <id> --cursor <event_cursor> [--operation sha256:<64hex>] [--timeout <sec>]`. It observes the vendor-native completion source, plus Claude's additive launch hook, as a **pure reader** and exits with a one-line `aiterm.agent-wait-result.v1` receipt. **Exit ≠ done**: the receipt's `outcome` is authoritative (`0` = `done`, `3` = `timeout`, `4` = `closed`, `1` = error).
472
476
  3. **The parent never runs the waiter in its own foreground.** Waiting is correct — but the waiter is a separate process, not the parent's turn. A harness that re-invokes its agent when a background task exits (Claude Code) runs the waiter **in the background** and gets woken with zero polling. So that this is not left to interpretation, aiterm reads `clientInfo.name` from the MCP `initialize` handshake and its receipts name the concrete invocation for the detected host — for Claude Code, literally `Bash(command: "aiterm-wait …", run_in_background: true)`. Unknown or undeclared hosts get the generic "start it as a process that does not block the parent's turn" wording; nothing else about the contract changes. Every receipt leads with the same rule: dispatch and let go, then go do something else or end the turn.
473
477
  4. Collect the result exactly as before: `pty_read(agent_transcript: true)`, or `claude_turn recover` for durable Claude operations. The waiter carries the signal, never the payload.
@@ -490,7 +494,7 @@ Each `pty_send` accepts at most 64 KiB of UTF-8 text. Sends to the same session
490
494
 
491
495
  ## A human can watch
492
496
 
493
- Sessions live on a shared tmux socket. The `tmux -S … attach -t <id>` line printed by `pty_open` (and by each agent launcher) lets a human attach to the same terminal and intervene (`Ctrl-b d` to detach) — including watching a launched Claude/Codex/Grok/Composer session run and taking the keyboard from your AI mid-task. On native Windows the printed line is the psmux form — `psmux -L <namespace> attach -t <id>` — pointing at the same Windows-native session.
497
+ Sessions live on a shared tmux socket. The `tmux -S … attach -t <id>` line printed by `pty_open` and `agent_launch` lets a human attach to the same terminal and intervene (`Ctrl-b d` to detach), including a Claude/Codex/Grok/Cursor harness session. On native Windows the printed line is the psmux form — `psmux -L <namespace> attach -t <id>` — pointing at the same Windows-native session.
494
498
 
495
499
  ## Requirements
496
500
 
@@ -498,7 +502,7 @@ Sessions live on a shared tmux socket. The `tmux -S … attach -t <id>` line pri
498
502
  - **tmux** (runtime prerequisite; check with `tmux -V`. Install with `apt install tmux` / `brew install tmux`)
499
503
  - **macOS / Linux / WSL2** run tmux directly. On macOS install it with `brew install tmux` (stock macOS ships none). If your MCP client is launched from the **GUI** rather than a terminal, Homebrew's bin (`/opt/homebrew/bin` on Apple Silicon, `/usr/local/bin` on Intel) may be off its `PATH`; aiterm auto-searches those locations, or set **`AITERM_TMUX=/path/to/tmux`** to point at it explicitly.
500
504
  - **Native Windows** has no tmux, so aiterm drives [psmux](https://github.com/psmux/psmux) — a tmux-CLI-compatible native multiplexer — with a per-install `-L` namespace. **No WSL is required.** Install psmux **3.3.8 or newer** (`winget install marlocarlo.psmux`; 3.3.8 is the first release whose `pipe-pane` file sink, byte-exact `paste-buffer` wire, and foreground `#{pane_current_command}` behave the way aiterm's capture/dispatch paths rely on), plus [Git for Windows](https://gitforwindows.org/) whose `bash.exe` becomes the pane shell (System32's `bash.exe` is the WSL launcher and is deliberately not used). Override resolution with **`AITERM_PSMUX`** / **`AITERM_BASH`** when the binaries live elsewhere. You reach Windows tools the same way you reach SSH: `pty_send "powershell.exe …"` nests into PowerShell. `grok_agent`/`composer_agent` launch the **Windows-native** Grok CLI (`%USERPROFILE%\.grok\bin\grok.exe`, or `GROK_BIN` pointing at a `.exe`) as a Windows process, and a WSL-side grok is rejected before a session is created so vendor auth and session records never split across an OS boundary.
501
- - For the **agent launchers**: the corresponding vendor CLI, installed and authenticated — `claude` for `claude_agent`, `codex` for `codex_agent`, `grok` for `grok_agent` / `composer_agent`. Portable fork additionally needs `throughline >= 0.9.0`; ordinary clean launch does not. (Not needed if you only use the PTY tools.)
505
+ - For **agent harnesses**: the selected CLI, installed and authenticated through its product owner's official path — `claude`, `codex`, `grok`, or Cursor's `cursor-agent`. Portable fork additionally needs `throughline >= 0.9.0`; ordinary clean launch does not. (Not needed if you only use the PTY tools.)
502
506
  - Optional: the [`rtk`](https://github.com/rtk-ai/rtk) binary (used by `pty_send`'s `rtk: true` delegation; works fine without it)
503
507
 
504
508
  ## Known constraints (by design, not bugs)
@@ -506,7 +510,7 @@ Sessions live on a shared tmux socket. The `tmux -S … attach -t <id>` line pri
506
510
  - **While nested (ssh / docker / REPL / a launched agent TUI), quiescence cannot fire by design**, because the foreground command is no longer in the shell set (bash/sh/zsh/fish/dash). When nested with no `until` and no `mark`, `pty_read({ wait: true })` returns early as `is_complete=False via nested` (rather than burning the full `timeout`, since no signal can confirm completion there) with a note to pass `until` (a literal substring by default; `until_regex: true` for a regex) or `mark: true` (an exit-code sentinel, auto-detected) for a confirmed completion. For a full-screen agent TUI, read `{ screen: true }` once its output settles.
507
511
  - **`is_complete=False` is not a failure.** It means "completion was not observed within `timeout`." For long commands, raise `timeout` or use `until`/`mark`.
508
512
  - **The destructive gate is a tripwire, not a sandbox.** It blocks common destructive forms only. It does **not** catch relative-path `rm`, things that become dangerous after `$VAR` expansion, or commands run on the far side of an SSH session — and it does not police what a launched coding agent does inside its own session.
509
- - **The agent launchers spawn a vendor TUI; they don't wrap or proxy it.** aiterm validates prerequisites and starts the CLI in a persistent PTY — the model, auth, and behavior are the vendor CLI's. There is no hidden inter-agent protocol; "conversation" is your MCP client driving the Claude/Codex/Grok/Composer TUI (send input, read output).
513
+ - **Agent harnesses run their real TUI; aiterm doesn't proxy the model API.** The selected harness owns model choice, authentication, and behavior. There is no hidden inter-agent protocol; the MCP client drives the Claude/Codex/Grok/Cursor TUI with ordinary send/read operations.
510
514
  - **`pty_send({ rtk: true })` is single-line only and needs the external `rtk` binary** (passthrough without it). The `pty_read({ rtk: true })` reducer, by contrast, is self-contained and rtk-independent.
511
515
  - **The `pytest` reducer matches rtk 0.42.0** on test counts, the rule line, and `FAILURES`-block formatting (locked by regression tests). It **deliberately preserves the full failure reason** on the `FAILED` summary lines (emitted under `-ra`/`-rf`), whereas rtk 0.42.0 truncates the reason at the first `" - "` — a readability choice, so those lines are intentionally not byte-identical to rtk. The `[full output: …]` tee-pointer line rtk appends on large output is not reproduced on the read side.
512
516
  - **tmux is started with `-f /dev/null`**, so it does not read `~/.tmux.conf` (to keep behavior reproducible across machines).
@@ -547,13 +551,13 @@ If aiterm let your AI hand a task to another agent — or saved you a round-trip
547
551
 
548
552
  ## Shared agent environment
549
553
 
550
- All four launchers use the caller's normal project and user environment. Aiterm does not copy,
554
+ All harnesses use the caller's normal project and user environment. Aiterm does not copy,
551
555
  symlink, filter, or replace vendor configuration, authentication, MCP, plugin, skill, permission,
552
556
  trust, memory, or history stores. Cleanup removes only aiterm-owned launch metadata and completion
553
557
  correlation files.
554
558
 
555
559
  The ordinary environment still comes from the shell/tmux session. When a caller needs a value that
556
- belongs to the current MCP process rather than the older persistent tmux server, every launcher
560
+ belongs to the current MCP process rather than the older persistent tmux server, every harness
557
561
  accepts `env_vars: ["NAME", ...]`. Only those names are refreshed at launch; this is a narrow
558
562
  per-launch overlay, not a replacement environment or configuration snapshot.
559
563
 
@@ -1,4 +1,4 @@
1
- // エージェント CLI(claude / codex / grok)・Throughline・pane shell(Windows は Git Bash)の
1
+ // エージェント CLI(claude / codex / grok / cursor-agent)・Throughline・pane shell(Windows は Git Bash)の
2
2
  // 実行ファイルをどう見つけ、どう起動するかの所有者(OS 分岐の所有者。tmux とは独立)。
3
3
  import { spawnSync } from "node:child_process";
4
4
  import * as fs from "node:fs";
@@ -115,7 +115,9 @@ export function resolveAgentBin(kind) {
115
115
  ? ["CLAUDE_BIN", [".local", "bin", "claude"], "claude"]
116
116
  : kind === "codex"
117
117
  ? ["CODEX_BIN", [".local", "bin", "codex"], "codex"]
118
- : ["GROK_BIN", [".grok", "bin", isWin ? "grok.exe" : "grok"], "grok"];
118
+ : kind === "cursor"
119
+ ? ["CURSOR_AGENT_BIN", [".local", "bin", "cursor-agent"], "cursor-agent"]
120
+ : ["GROK_BIN", [".grok", "bin", isWin ? "grok.exe" : "grok"], "grok"];
119
121
  const fromEnv = process.env[envVar];
120
122
  if (fromEnv) {
121
123
  // 明示指定 env は実在を検証する。存在しないパスを黙って返すと、session を作って
@@ -125,9 +127,16 @@ export function resolveAgentBin(kind) {
125
127
  return resolveWindowsCodexShim(kind, fromEnv);
126
128
  throw new AitermError(`${envVar} に指定された ${name} が存在しません: ${fromEnv}`, 2);
127
129
  }
128
- const cand = path.join(home, ...rel);
129
- if (isUsableAgentExecutableFile(cand))
130
- return resolveWindowsCodexShim(kind, cand);
130
+ const defaultCandidates = [path.join(home, ...rel)];
131
+ if (kind === "cursor" && isWin && process.env.LOCALAPPDATA) {
132
+ // Cursor公式Windows installerは %LOCALAPPDATA%\cursor-agent をPATHへ追加し、
133
+ // cursor-agent.exe/cmdを置く。GUI hostの古いPATHでも公式配置を直接解決する。
134
+ defaultCandidates.unshift(path.join(process.env.LOCALAPPDATA, "cursor-agent", "cursor-agent.exe"), path.join(process.env.LOCALAPPDATA, "cursor-agent", "cursor-agent.cmd"));
135
+ }
136
+ for (const cand of defaultCandidates) {
137
+ if (isUsableAgentExecutableFile(cand))
138
+ return resolveWindowsCodexShim(kind, cand);
139
+ }
131
140
  const w = spawnSync(isWin ? "where" : "which", [name], {
132
141
  encoding: "utf8",
133
142
  timeout: 5000,