aiterm-mcp 0.24.2 → 0.25.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.ja.md CHANGED
@@ -1,4 +1,4 @@
1
- > **Claude Code から Codex CLI の対話 TUI を操作する——スラッシュコマンドや [`$imagegen`](https://learn.chatgpt.com/docs/image-generation#generate-or-edit-an-image) のようなスキルまで、MCP越しに使える。**
1
+ > **任意のMCPクライアントから、Claude・Codex・Grok・Composerをクロスベンダーでも同一ベンダーでも永続対話TUIへ起動する。Codexのスラッシュコマンドや[`$imagegen`](https://learn.chatgpt.com/docs/image-generation#generate-or-edit-an-image)のような固有機能もそのまま使える。**
2
2
 
3
3
  <p align="center">
4
4
  <img src=".github/og.png" alt="Aiterm — 異なる知性が一つの持続する実行現場を共有する森の観測拠点" width="100%">
@@ -8,7 +8,7 @@
8
8
 
9
9
  # Aiterm
10
10
 
11
- [![CI](https://github.com/kitepon-rgb/aiterm-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/kitepon-rgb/aiterm-mcp/actions/workflows/ci.yml)
11
+ [![CI](https://github.com/kitepon/aiterm-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/kitepon/aiterm-mcp/actions/workflows/ci.yml)
12
12
  [![npm](https://img.shields.io/npm/v/aiterm-mcp.svg)](https://www.npmjs.com/package/aiterm-mcp)
13
13
  [![週間ダウンロード](https://img.shields.io/npm/dw/aiterm-mcp.svg)](https://www.npmjs.com/package/aiterm-mcp)
14
14
  [![node](https://img.shields.io/node/v/aiterm-mcp)](https://nodejs.org)
@@ -16,7 +16,7 @@
16
16
 
17
17
  > *(English: [README.md](README.md))*
18
18
 
19
- > **あなたの AI に、ほかの AI を操らせる。** 任意の MCP クライアントから 1 回の呼び出しで、コーディングエージェント(Claude・Codex・Grok・Composer)を永続端末の中に起動し、操作用のセッションを手渡す。何をしているかをトークン削減して読み、次の指示を送る。
19
+ > **あなたの AI に、ほかの AI を操らせる。** 任意の MCP クライアントから 1 回の呼び出しで、コーディングエージェント(Claude・Codex・Grok・Composer)を永続端末の中に起動し、操作用のセッションを手渡す。何をしているかをトークン削減して読み、次の指示を送る。呼び出し元と起動先のベンダーは独立しており、ClaudeからClaude/Codexを、CodexからClaude/Codexを起動できる。
20
20
  >
21
21
  > **これは何か:** AI が握る 1 本の永続 MCP 端末——その中に他のコーディングエージェントも起動できる。`ssh`・`docker exec`・REPL・別エージェントの TUI は、すべてその 1 本の端末の中へ「送るだけのテキスト」として入れ子になる。仕組みはあえて素朴——MCP クライアントが相手エージェントの端末を 1 ターンずつ操作するだけ。隠れたプロトコルも・aiterm独自の共有メモリ層も・自律的な交渉も無い。起動したagentは、直接CLIと同じproject/vendorの通常memory・設定を読む。
22
22
  >
@@ -90,11 +90,22 @@ claude mcp add --scope user --transport stdio aiterm -- npx -y aiterm-mcp
90
90
 
91
91
  **所有境界:** 本repositoryは永続PTYと外部agent実行レーンを所有します。製品横断の導入と
92
92
  host統合は、kitepon.devの製品開発を支える内部基盤
93
- [dotagents](https://github.com/kitepon-rgb/dotagents)が担当します。
93
+ [dotagents](https://github.com/kitepon/dotagents)が担当します。
94
94
 
95
95
  **言葉でなく実測で:** 記録済み203テストのベンチマークでは、`pty_read` はコンテキストに載るトークンを生ログの **約 7.1 分の 1** に減らす。しかも pass/fail の判定は畳んでも残る。→ [組み込みシェルツールとの使い分け](#組み込みシェルツールとの使い分け)
96
96
 
97
- 14 ツール: 6 つの **PTY ツール**(`pty_open` / `pty_send` / `pty_read` / `pty_key` / `pty_close` / `pty_list`)で 1 本の永続端末を開き・操作し・読む。加えて 4 つの **エージェント起動ツール**(`claude_agent` / `codex_agent` / `grok_agent` / `composer_agent`)が別のコーディングエージェントの TUI を新しい端末の中に起動し、`agent_configure`が起動中のCodex/Claudeのmodel・effortを再起動なしで変更し、`claude_turn`がdurable caller向けの構造化issue/recoveryを、`claude_approval`が相関済みClaude承認UI中継を、`diagnostics`が安全なfactory readinessを返す。バックエンドは **tmux** なので、MCP サーバや AI クライアントが再起動してもセッションは生き残る。
97
+ 14 ツール: 6 つの **PTY ツール**(`pty_open` / `pty_send` / `pty_read` / `pty_key` / `pty_close` / `pty_list`)で 1 本の永続端末を開き・操作し・読む。加えて 4 つの **エージェント起動ツール**(`claude_agent` / `codex_agent` / `grok_agent` / `composer_agent`)が別のコーディングエージェントの TUI を新しい端末の中に起動し、`agent_configure`が起動中のClaude/Codex/Grok/Composerのmodel・effortを再起動なしで変更し、`claude_turn`がdurable caller向けの構造化issue/recoveryを、`claude_approval`が相関済みClaude承認UI中継を、`diagnostics`が安全なfactory readinessを返す。バックエンドは **tmux** なので、MCP サーバや AI クライアントが再起動してもセッションは生き残る。
98
+
99
+ **v0.25.0ではGrok/Composerへ共通launcher制御を同等実装。** 起動時`reasoning_effort`、
100
+ `write_scope:"read-only"`の`--sandbox read-only`強制、`agent_configure`による同一session内の
101
+ model/effort変更に対応した。明示したGrok/Composer modelとComposer既定modelはPTY作成前に
102
+ 現在の`grok models` catalogへ照合し、不在時は別modelへ黙ってfallbackせず明示失敗する。
103
+
104
+ **v0.24.3ではlauncherへ渡す環境変数を現在のMCP processから明示選択できる。** `env_vars`へ
105
+ 変数名だけを指定すると、aitermは起動時の現在値を読み、存在する値だけをそのagentへ渡す。永続tmux
106
+ serverがMCP processより先に起動していても、古いserver環境に席identityやworkflow変数を消されない。
107
+ あわせてCodex v0.147が長寿命footerへ加える任意`fast`を認識し、idleな`medium fast ·` sessionでも
108
+ 再描画・再試行・再起動なしに`agent_configure`できる。
98
109
 
99
110
  **v0.24.2では長寿命Codexでも設定変更を維持。** 起動時headerがcapture範囲外へ流れた後は、
100
111
  常駐するmodel/effort footerと入力欄でCodexを識別する。idle sessionをそのまま変更でき、
@@ -162,7 +173,7 @@ pty_read(id, { wait: true }) → 削減済みの出力を読む(完了
162
173
 
163
174
  既存の人間向けtextに加えて`aiterm.agent-launch-result.v1` structured receiptも返すため、durable callerは表示文字列を解析せずsession handleを取得できる。Codexは通常rollout transcriptの`task_complete`、Grok/Composerは通常session event、Claudeは通常settingsへ加算したlaunch固有Stop hookを完了正本に使う。agent sessionへの`pty_send`は非ブロックの **dispatch** になり`event_cursor`入りreceiptを即返す。完了通知は`aiterm-wait --session <id> --cursor <event_cursor>`を親のターンを塞がない別processで受ける。durable machine callerは`claude_turn`を使い、recoveryは再送せず、検証済み完了だけがexact `raw_output`を持つ。
164
175
 
165
- `codex_agent`・`grok_agent`・`composer_agent`は任意の`write_scope`(`"read-only"`または書込み許可パスの説明)も受ける。指定値はlaunch receipt・session metadata・`pty_list`へ保存する。Codexの`write_scope:"read-only"`だけは実効能力壁であり、aitermがCLIの`--sandbox read-only`を付ける。Grok/Composerには対応する対話起動sandboxがなく、Codexにもパス説明をallowlistへ変換するフラグがないため、それらは強制済みと偽らず`write_scope_enforcement:"declaration_only_unsupported"`を返す。`write_scope`を省略した起動は従来どおりである。
176
+ `codex_agent`・`grok_agent`・`composer_agent`は任意の`write_scope`(`"read-only"`または書込み許可パスの説明)も受ける。指定値はlaunch receipt・session metadata・`pty_list`へ保存する。3 launcherすべてで`write_scope:"read-only"`は実効能力壁となり、aitermがCLIの`--sandbox read-only`を付ける。パス説明をallowlistへ変換する同等CLI引数はないため、そちらだけは`write_scope_enforcement:"declaration_only_unsupported"`を返す。`write_scope`を省略した起動は従来どおりである。
166
177
 
167
178
  ```text
168
179
  codex_agent({ session_name: "codex1", cwd: "/repo",
@@ -180,10 +191,17 @@ $ aiterm-wait --session codex1 --cursor <event_cursor> # exit 0=done / 3=timeo
180
191
 
181
192
  | ツール | 起動するもの | 主な引数 |
182
193
  | --- | --- | --- |
183
- | `claude_agent` | Claude Code CLI(Anthropic) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`), `cwd?`, `session_name?` |
184
- | `codex_agent` | Codex CLI(OpenAI・端末設定/CLI既定、`model?`で上書き) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`/`ultra`), `cwd?`, `session_name?`, `write_scope?` |
185
- | `grok_agent` | Grok Build(xAI、既定`grok-4.5`、`model?`で上書き) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`は非対応(指定時は明示エラー), `cwd?`, `session_name?`, `write_scope?` |
186
- | `composer_agent` | Grok Build(xAI、既定`grok-composer-2.5-fast`、`model?`で上書き) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`は非対応(指定時は明示エラー), `cwd?`, `session_name?`, `write_scope?` |
194
+ | `claude_agent` | Claude Code CLI(Anthropic) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`), `env_vars?`, `cwd?`, `session_name?` |
195
+ | `codex_agent` | Codex CLI(OpenAI・端末設定/CLI既定、`model?`で上書き) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`/`ultra`), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
196
+ | `grok_agent` | Grok Build(xAI、既定`grok-4.5`、`model?`で上書き) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(vendor対応値), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
197
+ | `composer_agent` | Grok Build(xAI、既定`grok-composer-2.5-fast`、`model?`で上書き。全modelをlive catalog照合) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(vendor対応値), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
198
+
199
+ `env_vars`は環境変数の**名前**だけを並べるallowlistであり、name/value mapではない。aitermは
200
+ launcher起動時に現在のMCP processから各名前を読み、存在する値をshell quoteして、その1回のvendor
201
+ 起動コマンドへ入れる。未設定名は省略し、shell変数名として不正な名前はsession作成前に失敗する。
202
+ 全環境の暗黙copy、tmux server再起動、retry、fallbackは行わない。値はMCP tool引数には入らないが、
203
+ PTYの起動コマンドとして送られ、sessionの`.lastcmd`にも保持されるため、起動先vendorと同じOS userへ
204
+ 到達する。秘密転送路ではなく、席identityやworkflow用の非secret変数だけに使う。
187
205
 
188
206
  各ベンダーのCLIが導入・認証済みであること。CLI不在・不正なmodel/effort・実在しない`cwd`はsession作成前に失敗し、残骸を残さない。ClaudeはさらにPTY作成前に同じCLIの`auth status --json`が`loggedIn:true`を返すことを要求する。4 launcherは通常のvendor credential/config storeをその場で使い、fake `HOME`、private `CODEX_HOME`/`GROK_HOME`、project/user config snapshotを作らない。Claudeだけは完了相関用Stop hook settingsを通常の`user,project,local` settingsへ加算する。Grok/Composerは画面入力欄だけでなく通常sessionの`mcp_init_completed` eventも確認してから送信し、共有MCP初期化中の早送信を防ぐ。相関付きClaudeのactive turn中はC-c以外の`pty_key`と素送信を拒否し、承認UIは`claude_approval`で単発Yes/Noだけを相関付きで中継する。
189
207
 
@@ -277,9 +295,9 @@ claude mcp add --scope user --transport stdio aiterm -- aiterm-mcp
277
295
 
278
296
  ## ヘッドレス: 端末に人が居ない
279
297
 
280
- MCP クライアントが aiterm を stdio 越しにプログラムから駆動するので、上のすべては **tmux に誰も座らないまま**動く。あなたの Claude Code セッションは、`codex_agent()` でタスクを起こし、`pty_read` で結果を読み、それを使って動ける——無人で。これは、人が操作する端末が向かない場所にこそ aiterm が合うということ:
298
+ MCP クライアントが aiterm を stdio 越しにプログラムから駆動するので、上のすべては **tmux に誰も座らないまま**動く。Claude/Codexのどちらからでも、自分自身を含む任意のlauncherを呼び、`pty_read`で結果を読んで次へ進める——無人で。これは、人が操作する端末が向かない場所にこそ aiterm が合うということ:
281
299
 
282
- - **複数エージェントのオーケストレーション** — 統括役がサブタスクを Codex / Grok / Composer に渡し、各々を専用の永続セッションに置き、全部を読み戻す。
300
+ - **複数エージェントのオーケストレーション** — 統括役がサブタスクを Claude / Codex / Grok / Composer に渡し、各々を専用の永続セッションに置き、全部を読み戻す。
283
301
  - **CI** — ジョブのステップがエージェントを起こし、操作し、片付けられる。
284
302
  - **cron** — スケジュール実行がエージェントを起動して出力を回収できる。
285
303
 
@@ -370,7 +388,7 @@ aiterm は同じ核心の洞察——端末を出会いの場にする——を
370
388
  | `pty_key` | 制御キーを送る | `session_id`, `key`(`C-c`/`Enter`/`Up`…) |
371
389
  | `pty_close` | 冪等に閉じ、`closed` / `already_closed`を返す | `session_id` |
372
390
  | `pty_list` | セッション一覧 | (なし) |
373
- | `agent_configure` | 起動中のCodex/Claudeを再起動せずmodel/effort変更 | `session_id`, `model?`, `reasoning_effort?` |
391
+ | `agent_configure` | 起動中のClaude/Codex/Grok/Composerを再起動せずmodel/effort変更 | `session_id`, `model?`, `reasoning_effort?` |
374
392
  | `claude_turn` | 相関済みClaude operationをdispatch(issue)または回収(recover) | `action`, `session_id`, `operation_id`, `text?` |
375
393
  | `claude_approval` | 現在表示中の相関済みClaude承認UIを検査または応答 | `action`, `session_id`, `operation_id?`, `approval_choice?`, `observed_prompt_digest?` |
376
394
  | `diagnostics` | 機械可読 JSON による read-only factory readiness | (なし) |
@@ -387,14 +405,14 @@ consumer は `aiterm-runtime-errors snapshot` を読み、durable ingestion 後
387
405
 
388
406
  各ツールは特定ベンダーの対話型コーディングエージェント TUI を新しい永続 PTY の中に起動し、`session_id` を返す。以後は他のセッションと同様に `pty_read` / `pty_send` で操作する。モデルごとに 1 ツール=ツール名を見ればどのモデルか分かる。TUI は全画面アプリなので、`pty_read({ screen: true })` で描画済みの画面を読む。
389
407
 
390
- `agent_configure({ session_id, model?, reasoning_effort? })`はvendor標準操作で起動中のCodex/Claudeを変更し、PTYと会話contextを維持する。Claude Code標準の`/model`・`/effort`は、新しいClaude sessionの既定値も同時に保存する。
408
+ `agent_configure({ session_id, model?, reasoning_effort? })`はvendor標準操作で起動中のClaude/Codex/Grok/Composerを変更し、PTYと会話contextを維持する。Grok/Composerは同じlive model catalog照合後に`/model <model> [effort]`または`/effort <effort>`を使う。Claude Code標準の`/model`・`/effort`は、新しいClaude sessionの既定値も同時に保存する。
391
409
 
392
410
  | ツール | 起動するもの | 主な引数 |
393
411
  | --- | --- | --- |
394
- | `claude_agent` | Claude Code CLI(Anthropic) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`), `cwd?`, `session_name?` |
395
- | `codex_agent` | Codex CLI(OpenAI・端末設定/CLI既定、`model?`で上書き) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`/`ultra`), `cwd?`, `session_name?`, `write_scope?` |
396
- | `grok_agent` | Grok Build(xAI、既定`grok-4.5`、`model?`で上書き) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`は非対応(指定時は明示エラー), `cwd?`, `session_name?`, `write_scope?` |
397
- | `composer_agent` | Grok Build(xAI、既定`grok-composer-2.5-fast`、`model?`で上書き) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`は非対応(指定時は明示エラー), `cwd?`, `session_name?`, `write_scope?` |
412
+ | `claude_agent` | Claude Code CLI(Anthropic) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`), `env_vars?`, `cwd?`, `session_name?` |
413
+ | `codex_agent` | Codex CLI(OpenAI・端末設定/CLI既定、`model?`で上書き) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`/`ultra`), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
414
+ | `grok_agent` | Grok Build(xAI、既定`grok-4.5`、`model?`で上書き) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(vendor対応値), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
415
+ | `composer_agent` | Grok Build(xAI、既定`grok-composer-2.5-fast`、`model?`で上書き。全modelをlive catalog照合) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?`(vendor対応値), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
398
416
 
399
417
  対応するCLI(`claude` / `codex` / `grok`)の導入・認証が必要。前提違反はsession作成前に明示失敗する。4 launcherすべてが通常project/user環境と同じ非ブロックdispatch契約を使う。Claude/Codex/Grok/Composerのdepth 1 live smokeと、Claude親→Claude孫のdepth 2 nested delegation smokeはgreenであり、fixtureによる検証とは区別して記録する。
400
418
 
@@ -454,6 +472,11 @@ npm test # build してから node:test 回帰スイート(tmux 必
454
472
  npm link # ローカルで `aiterm-mcp` を PATH に
455
473
  ```
456
474
 
475
+ 開発中は変更に直結するfocused testを先にローカルで実行します。GitHub Actionsの最終gateは
476
+ self-hostedのmacOS native・Linux native・Windows native・WSL2で同じ`npm test`を同時実行し、
477
+ OS別の縮小suiteで代用しません。tag起点のnpm公開は4環境greenとtagged commitの`origin/main`
478
+ 祖先確認を通過した後だけ実行します。
479
+
457
480
  ロジックは `src/core.ts`(tmux 制御・削減・完了検出・安全・エージェント起動)と `src/rtk.ts`(コマンド別 reducer)、公開は `src/index.ts`。設計の出発点と reducer の移植元(pytest reducer は本家 rtk 0.42.0 と一致するよう移植・ただし上記の `FAILED` 行の差異は意図的・回帰テストで固定)は `prototype/python/` を参照。
458
481
 
459
482
  ## 試す
@@ -464,10 +487,10 @@ npm link # ローカルで `aiterm-mcp` を PATH に
464
487
  claude mcp add --scope user --transport stdio aiterm -- npx -y aiterm-mcp
465
488
  ```
466
489
 
467
- aiterm が、あなたの AI に別のエージェントへ仕事を渡させたなら——あるいはトークンの往復を 1 回でも省けたなら——**[リポジトリに star](https://github.com/kitepon-rgb/aiterm-mcp)** を。他の人に見つけてもらう一番安い方法です。
490
+ aiterm が、あなたの AI に別のエージェントへ仕事を渡させたなら——あるいはトークンの往復を 1 回でも省けたなら——**[リポジトリに star](https://github.com/kitepon/aiterm-mcp)** を。他の人に見つけてもらう一番安い方法です。
468
491
 
469
492
  - **npm:** https://www.npmjs.com/package/aiterm-mcp
470
- - **Issue / バグ報告:** https://github.com/kitepon-rgb/aiterm-mcp/issues
493
+ - **Issue / バグ報告:** https://github.com/kitepon/aiterm-mcp/issues
471
494
 
472
495
  ## ライセンス
473
496
 
package/README.md CHANGED
@@ -1,4 +1,4 @@
1
- > **Drive Codex CLI's interactive TUI from Claude Code — including slash commands and skills such as [`$imagegen`](https://learn.chatgpt.com/docs/image-generation#generate-or-edit-an-image) — through MCP.**
1
+ > **From any MCP client, launch Claude, Codex, Grok, or Composer — cross-vendor or same-vendor — inside a persistent interactive TUI, with native features such as Codex slash commands and [`$imagegen`](https://learn.chatgpt.com/docs/image-generation#generate-or-edit-an-image) available.**
2
2
 
3
3
  <p align="center">
4
4
  <img src=".github/og.png" alt="Aiterm — a shared forest observatory where different intelligences work in one persistent execution space" width="100%">
@@ -8,7 +8,7 @@
8
8
 
9
9
  # Aiterm
10
10
 
11
- [![CI](https://github.com/kitepon-rgb/aiterm-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/kitepon-rgb/aiterm-mcp/actions/workflows/ci.yml)
11
+ [![CI](https://github.com/kitepon/aiterm-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/kitepon/aiterm-mcp/actions/workflows/ci.yml)
12
12
  [![npm](https://img.shields.io/npm/v/aiterm-mcp.svg)](https://www.npmjs.com/package/aiterm-mcp)
13
13
  [![weekly downloads](https://img.shields.io/npm/dw/aiterm-mcp.svg)](https://www.npmjs.com/package/aiterm-mcp)
14
14
  [![node](https://img.shields.io/node/v/aiterm-mcp)](https://nodejs.org)
@@ -16,7 +16,7 @@
16
16
 
17
17
  > *(日本語: [README.ja.md](README.ja.md))*
18
18
 
19
- > **Let your AI orchestrate other AIs.** From any MCP client, one call spawns a coding agent (Claude, Codex, Grok, or Composer) inside a persistent terminal and hands you a session to drive: read what it's doing token-reduced, send it the next instruction.
19
+ > **Let your AI orchestrate other AIs.** From any MCP client, one call spawns a coding agent (Claude, Codex, Grok, or Composer) inside a persistent terminal and hands you a session to drive: read what it's doing token-reduced, send it the next instruction. The caller and launched vendor are independent: Claude can launch Claude or Codex, and Codex can launch Claude or Codex.
20
20
  >
21
21
  > **What it is:** one persistent MCP terminal your AI drives — and can launch other coding agents into. `ssh`, `docker exec`, a REPL, or another agent's TUI all nest inside that one terminal as just text you send in. The mechanism is deliberately plain — your MCP client drives the other agent's terminal turn by turn: no hidden protocol, no separate aiterm-owned shared-memory layer, no autonomous negotiation. Launched agents still read the normal project and vendor memory/configuration that a direct CLI launch would use.
22
22
  >
@@ -89,12 +89,25 @@ Save this as `.cursor/mcp.json` for the project, or `~/.cursor/mcp.json` globall
89
89
 
90
90
  **Ownership boundary:** this repository owns the persistent PTY and external-agent
91
91
  execution lane. Cross-product installation and host integration are handled by
92
- [dotagents](https://github.com/kitepon-rgb/dotagents), the internal development
92
+ [dotagents](https://github.com/kitepon/dotagents), the internal development
93
93
  toolchain behind kitepon.dev's products.
94
94
 
95
95
  **Measured, not claimed:** in the recorded 203-test benchmark, a `pty_read` puts **~7.1× fewer tokens** in your context than the raw log — and the pass/fail verdict survives the fold. → [When to reach for it vs. the built-in shell](#when-to-reach-for-it-vs-the-built-in-shell)
96
96
 
97
- Fourteen tools: six **PTY tools** — `pty_open` / `pty_send` / `pty_read` / `pty_key` / `pty_close` / `pty_list` — to open, drive, and read one persistent terminal, four **agent launchers** — `claude_agent` / `codex_agent` / `grok_agent` / `composer_agent` — that each start another coding agent's TUI inside a fresh one, `agent_configure` to change a running Codex/Claude session's model and effort without restarting it, `claude_turn` for durable structured issue/recovery, `claude_approval` for correlated Claude approval prompts, and `diagnostics` for safe factory readiness. The backend is **tmux**, so sessions survive even if the MCP server or the AI client restarts.
97
+ Fourteen tools: six **PTY tools** — `pty_open` / `pty_send` / `pty_read` / `pty_key` / `pty_close` / `pty_list` — to open, drive, and read one persistent terminal, four **agent launchers** — `claude_agent` / `codex_agent` / `grok_agent` / `composer_agent` — that each start another coding agent's TUI inside a fresh one, `agent_configure` to change a running Claude/Codex/Grok/Composer session's model and effort without restarting it, `claude_turn` for durable structured issue/recovery, `claude_approval` for correlated Claude approval prompts, and `diagnostics` for safe factory readiness. The backend is **tmux**, so sessions survive even if the MCP server or the AI client restarts.
98
+
99
+ **v0.25.0 gives Grok and Composer the same shared launcher controls.** Their launchers now pass
100
+ `reasoning_effort`, enforce `write_scope: "read-only"` with `--sandbox read-only`, and support
101
+ in-place model/effort changes through `agent_configure`. Before creating a PTY, aiterm checks an
102
+ explicit Grok/Composer model—and Composer's default model—against the live `grok models` catalog.
103
+ An unavailable model fails visibly instead of letting the vendor CLI fall back to another model.
104
+
105
+ **v0.24.3 forwards explicitly selected launcher environment variables from the current MCP process.**
106
+ Pass variable names in `env_vars`; aiterm reads their current values at launch and injects only the
107
+ present ones into that agent. This works even when the persistent tmux server predates the MCP
108
+ process, so a stale tmux-server environment cannot erase per-seat identity or workflow variables.
109
+ It also recognizes Codex v0.147's optional `fast` token in long-lived model/effort footers, keeping
110
+ `agent_configure` available on an idle `medium fast ·` session without redraw, retry, or restart.
98
111
 
99
112
  **v0.24.2 keeps in-place configuration working in long-lived Codex sessions.** Once the
100
113
  startup header has scrolled out of the captured pane, aiterm recognizes Codex by its persistent
@@ -155,7 +168,7 @@ A lot of 2026's agent tooling is converging on orchestration: a lead model deleg
155
168
 
156
169
  ## Built with Codex and GPT-5.6 for OpenAI Build Week 2026
157
170
 
158
- aiterm predates Build Week, so the event work is kept visible in dated commits. During the submission window (July 14–16, 2026), I extended it with safe serialized delivery for long PTY input, correlated operation IDs and bounded result recovery, machine-readable launch and idempotent close receipts, and a hardened readiness gate that prevents prompts from disappearing during TUI startup redraws. The public comparison from the pre-event release is [`v0.12.2...main`](https://github.com/kitepon-rgb/aiterm-mcp/compare/v0.12.2...main).
171
+ aiterm predates Build Week, so the event work is kept visible in dated commits. During the submission window (July 14–16, 2026), I extended it with safe serialized delivery for long PTY input, correlated operation IDs and bounded result recovery, machine-readable launch and idempotent close receipts, and a hardened readiness gate that prevents prompts from disappearing during TUI startup redraws. The public comparison from the pre-event release is [`v0.12.2...main`](https://github.com/kitepon/aiterm-mcp/compare/v0.12.2...main).
159
172
 
160
173
  I used **Codex with GPT-5.6** as an engineering collaborator: it inspected the implementation, challenged the API and recovery contracts, generated focused regression cases, and helped verify race, security, timeout, and malformed-event paths. I reviewed the diffs and test evidence and retained the final product and architecture decisions. At that Build Week checkpoint, the regression suite contained 262 tests covering normal operation as well as failure and recovery behavior; current release receipts live in the [CHANGELOG](CHANGELOG.md) and release ADRs.
161
174
 
@@ -180,7 +193,7 @@ The same primitive hosts another agent's TUI. Four launchers each start one vend
180
193
 
181
194
  The human-readable launch text is accompanied by an `aiterm.agent-launch-result.v1` structured receipt, so durable callers never parse display text for the session handle; when the launch carries an initial `prompt`, the receipt also includes the `event_cursor`, a ready-made `wait_command` for the completion waiter, and a `submit_residue` observation (`true` = the prompt is likely still sitting unsubmitted in the composer — the hint explains recovery; `false` = no residue observed, not a proof of submission; `null` = not applicable). From there you drive it with the same `pty_read` / `pty_send` you'd use on any shell. Codex completion comes from its normal durable rollout transcript's `task_complete`; Grok/Composer use their normal session events; Claude receives a launch-specific Stop hook settings addition while still loading normal user/project/local settings. Sending to an agent session is a non-blocking **dispatch** — the call returns immediately with an `event_cursor`, and completion arrives via [`aiterm-wait`](#completion-push-for-parent-agents-aiterm-wait). Durable machine callers use `claude_turn({ action: "issue" | "recover", session_id, operation_id, ... })`: it returns fixed `accepted` / `pending` / `completed` / `unknown` states without parsing human-facing errors, never resends during recovery, and includes exact `raw_output` only for a verified completion. An initial `prompt` on `claude_agent`/`codex_agent` is submitted through the same ready gate and the launcher returns without waiting; on Grok/Composer it is passed on the CLI's argv. This needs the vendor's own CLI installed and authenticated — see [Requirements](#requirements).
182
195
 
183
- `codex_agent`, `grok_agent`, and `composer_agent` also accept an optional `write_scope`: either `"read-only"` or a human-readable description of writable paths. A supplied value is retained in the launch receipt, session metadata, and `pty_list`. For Codex, `write_scope: "read-only"` is an effective boundary: aiterm adds the CLI's `--sandbox read-only` flag. Grok/Composer have no corresponding interactive-launch sandbox, and Codex has no path-description allowlist flag; those cases return `write_scope_enforcement: "declaration_only_unsupported"` rather than claiming enforcement. Omitting `write_scope` preserves prior behavior.
196
+ `codex_agent`, `grok_agent`, and `composer_agent` also accept an optional `write_scope`: either `"read-only"` or a human-readable description of writable paths. A supplied value is retained in the launch receipt, session metadata, and `pty_list`. For all three launchers, `write_scope: "read-only"` is an effective boundary: aiterm adds the CLI's `--sandbox read-only` flag. A path description remains declaration-only because none of these CLI launch surfaces provides an equivalent path allowlist flag; that case returns `write_scope_enforcement: "declaration_only_unsupported"`. Omitting `write_scope` preserves prior behavior.
184
197
 
185
198
  For a correlated Claude turn stopped at `Do you want to proceed?`, use `claude_approval(action: "inspect", ...)` to capture the active operation and SHA-256 screen digest, review the displayed command, then call `respond` with that exact digest and either `approve_once` or `deny`. The relay rechecks the operation and screen under the send lock, never exposes arbitrary input or permanent approval, keeps the active marker intact, and records a prompt-free owner-only receipt. `pty_send(force: true)` does not bypass this boundary.
186
199
 
@@ -201,10 +214,19 @@ One call per model, so the tool name itself tells you which model you get:
201
214
 
202
215
  | Tool | Launches | Key args |
203
216
  | --- | --- | --- |
204
- | `claude_agent` | Claude Code CLI (Anthropic) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`), `cwd?`, `session_name?`, `launch_operation_id?` |
205
- | `codex_agent` | Codex CLI (OpenAI; terminal config/CLI default unless overridden) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`/`ultra`; ultra enables proactive automatic delegation), `cwd?`, `session_name?`, `write_scope?` |
206
- | `grok_agent` | Grok Build, model `grok-4.5` by default (`model?` overrides) (xAI) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` unsupported (an explicit value is an error; Grok CLI `--effort` is headless-only), `cwd?`, `session_name?`, `write_scope?` |
207
- | `composer_agent` | Grok Build, model `grok-composer-2.5-fast` by default (`model?` overrides) (xAI) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` unsupported (an explicit value is an error), `cwd?`, `session_name?`, `write_scope?` |
217
+ | `claude_agent` | Claude Code CLI (Anthropic) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`), `env_vars?`, `cwd?`, `session_name?`, `launch_operation_id?` |
218
+ | `codex_agent` | Codex CLI (OpenAI; terminal config/CLI default unless overridden) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`/`ultra`; ultra enables proactive automatic delegation), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
219
+ | `grok_agent` | Grok Build, model `grok-4.5` by default (`model?` overrides) (xAI) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (vendor-supported value), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
220
+ | `composer_agent` | Grok Build, model `grok-composer-2.5-fast` by default (`model?` overrides); every model is live-catalog checked (xAI) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (vendor-supported value), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
221
+
222
+ `env_vars` is an allowlist of environment-variable **names**, not a name/value map. At launch,
223
+ aiterm reads each valid name from its current MCP process, shell-quotes present values, and places
224
+ them on that one vendor launch command. Missing names are omitted; invalid shell variable names
225
+ fail before session creation. There is no implicit whole-environment copy, tmux-server restart,
226
+ retry, or fallback. Values do not enter the MCP tool arguments, but they are delivered through the
227
+ PTY launch command and retained in aiterm's per-session `.lastcmd`; the launched vendor and other
228
+ processes with access to the same OS user may read them. Use this for non-secret seat identity and
229
+ workflow variables, not as a secret transport.
208
230
 
209
231
  The vendor CLI must be installed and authenticated (`claude` for `claude_agent`; `codex` for `codex_agent`; `grok` for both Grok tools). aiterm resolves the binary via `CLAUDE_BIN` / `CODEX_BIN` / `GROK_BIN`, then each documented default location, then `PATH`. Prerequisites are checked **before** a session exists: empty `model` values and unsupported effort values are rejected up front; a missing CLI binary or a nonexistent `cwd` fails for all four. Before creating a Claude session, aiterm also requires a successful structured `claude auth status --json` result with `loggedIn: true`; unavailable, malformed, or failed authentication leaves **zero leftover session**. All launchers use the normal vendor-owned credential and configuration stores in place. No launcher creates a fake `HOME`, a private `CODEX_HOME`/`GROK_HOME`, or a snapshot of project/user configuration.
210
232
 
@@ -215,7 +237,7 @@ places its returned context before a fixed separator and the mission. This route
215
237
  `throughline >= 0.9.0`; it reads the source memory without changing that database's session
216
238
  ownership. No Throughline dependency is needed when the field is omitted.
217
239
 
218
- Claude and Codex launchers forward `model` and `reasoning_effort` through public CLI flags; Grok/Composer reject `reasoning_effort` because it is headless-only. Pass an absolute path for `cwd` — `~` is not expanded. Durable callers can make a promptless Claude launch exactly replayable by passing an explicit `session_name` and `launch_operation_id`. Claude adds a launch-local settings file only for the correlated Stop hook and loads it together with normal `user,project,local` setting sources; it does not replace normal hooks, MCPs, plugins, permissions, or trust state. The hook event contains no answer body, and the bounded owner-only result is returned by `pty_read({ agent_transcript:true })` without reading Claude's private transcript. While a correlated Claude turn is active, raw sends and non-interrupt keys are rejected. Exact `/login` and `/logout` dispatches are rejected so shared authentication is repaired once in a normal terminal. Codex reads its normal rollout store; Grok/Composer read their normal session event/history files. Before dispatch, Codex waits for an idle TUI, while Grok/Composer additionally require the vendor's structured `mcp_init_completed` event so a visible input box cannot accept a prompt too early. Correlated completion requires POSIX filesystem semantics (Linux, WSL2, macOS).
240
+ All four launchers forward `model` and `reasoning_effort` through public CLI flags. Explicit Grok/Composer models and Composer's default model are checked against the current `grok models` catalog before a PTY exists; missing models are errors, with no cache, retry, or fallback to another model. Pass an absolute path for `cwd` — `~` is not expanded. Durable callers can make a promptless Claude launch exactly replayable by passing an explicit `session_name` and `launch_operation_id`. Claude adds a launch-local settings file only for the correlated Stop hook and loads it together with normal `user,project,local` setting sources; it does not replace normal hooks, MCPs, plugins, permissions, or trust state. The hook event contains no answer body, and the bounded owner-only result is returned by `pty_read({ agent_transcript:true })` without reading Claude's private transcript. While a correlated Claude turn is active, raw sends and non-interrupt keys are rejected. Exact `/login` and `/logout` dispatches are rejected so shared authentication is repaired once in a normal terminal. Codex reads its normal rollout store; Grok/Composer read their normal session event/history files. Before dispatch, Codex waits for an idle TUI, while Grok/Composer additionally require the vendor's structured `mcp_init_completed` event so a visible input box cannot accept a prompt too early. Correlated completion requires POSIX filesystem semantics (Linux, WSL2, macOS).
219
241
 
220
242
  There is no hidden protocol between agents: a launched Claude, Codex, Grok, or Composer is another user-visible persistent terminal session. The MCP client drives that TUI with ordinary PTY operations, and a human can attach to watch or take over.
221
243
 
@@ -301,9 +323,9 @@ This registers it in `~/.claude.json`; you'll get an approval prompt the first t
301
323
 
302
324
  ## Headless: no human at the terminal
303
325
 
304
- Because an MCP client drives aiterm programmatically over stdio, everything above can run with **nobody sitting at a tmux**. Your Claude Code session can `codex_agent()` a task, `pty_read` the result, and act on it — unattended. That makes aiterm a fit for exactly the places a human-driven terminal isn't:
326
+ Because an MCP client drives aiterm programmatically over stdio, everything above can run with **nobody sitting at a tmux**. A Claude or Codex session can call any launcher — including another instance of itself — then `pty_read` the result and act on it unattended. That makes aiterm a fit for exactly the places a human-driven terminal isn't:
305
327
 
306
- - **Multi-agent orchestration** — an orchestrator hands sub-tasks to Codex / Grok / Composer, each in its own persistent session, and reads them all back.
328
+ - **Multi-agent orchestration** — an orchestrator hands sub-tasks to Claude / Codex / Grok / Composer, each in its own persistent session, and reads them all back.
307
329
  - **CI** — a job step can spin up an agent, drive it, and tear it down.
308
330
  - **cron** — a scheduled run can launch an agent and collect its output.
309
331
 
@@ -396,7 +418,7 @@ On top of that sits a productized layer a raw tmux bridge doesn't have: **token-
396
418
  | `pty_key` | Send a control key | `session_id`, `key` (`C-c`/`Enter`/`Up`…) |
397
419
  | `pty_close` | Close idempotently; return `closed` / `already_closed` | `session_id` |
398
420
  | `pty_list` | List sessions (agent rows carry `agent=<kind>` metadata) | (none) |
399
- | `agent_configure` | Change model/effort in a running Codex or Claude session without restarting it | `session_id`, `model?`, `reasoning_effort?` |
421
+ | `agent_configure` | Change model/effort in a running Claude, Codex, Grok, or Composer session without restarting it | `session_id`, `model?`, `reasoning_effort?` |
400
422
  | `claude_turn` | Issue (dispatch-only) or recover one correlated Claude operation | `action`, `session_id`, `operation_id`, `text?` |
401
423
  | `claude_approval` | Inspect or answer the current correlated Claude approval prompt | `action`, `session_id`, `operation_id?`, `approval_choice?`, `observed_prompt_digest?` |
402
424
  | `diagnostics` | Read-only factory readiness as machine-readable JSON | (none) |
@@ -413,16 +435,16 @@ Consumer flow is `aiterm-runtime-errors snapshot`, then `aiterm-runtime-errors a
413
435
 
414
436
  Each launcher starts a specific vendor's interactive coding-agent TUI inside a fresh persistent PTY and returns its `session_id` — from there you drive it with plain `pty_read` / `pty_send`, exactly like any other session. One tool per model, so the tool name itself tells you which model you get. The TUI is a full-screen app, so read it with `pty_read({ screen: true })` for the rendered view.
415
437
 
416
- `agent_configure({ session_id, model?, reasoning_effort? })` changes a running Codex or Claude TUI through the vendor's standard controls, preserving the PTY and conversation context. Claude Code's native `/model` and `/effort` commands also save those choices as defaults for new Claude sessions.
438
+ `agent_configure({ session_id, model?, reasoning_effort? })` changes a running Claude, Codex, Grok, or Composer TUI through the vendor's standard controls, preserving the PTY and conversation context. Grok/Composer use `/model <model> [effort]` or `/effort <effort>` after the same live model-catalog check. Claude Code's native `/model` and `/effort` commands also save those choices as defaults for new Claude sessions.
417
439
 
418
440
  | Tool | Launches | Key args |
419
441
  | --- | --- | --- |
420
- | `claude_agent` | Claude Code CLI (Anthropic) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`), `cwd?`, `session_name?`, `launch_operation_id?` |
421
- | `codex_agent` | Codex CLI (OpenAI; terminal config/CLI default unless overridden) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`/`ultra`; ultra enables proactive automatic delegation), `cwd?`, `session_name?`, `write_scope?` |
422
- | `grok_agent` | Grok Build, model `grok-4.5` by default (`model?` overrides) (xAI) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` unsupported (an explicit value is an error; Grok CLI `--effort` is headless-only), `cwd?`, `session_name?`, `write_scope?` |
423
- | `composer_agent` | Grok Build, model `grok-composer-2.5-fast` by default (`model?` overrides) (xAI) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` unsupported (an explicit value is an error), `cwd?`, `session_name?`, `write_scope?` |
442
+ | `claude_agent` | Claude Code CLI (Anthropic) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`), `env_vars?`, `cwd?`, `session_name?`, `launch_operation_id?` |
443
+ | `codex_agent` | Codex CLI (OpenAI; terminal config/CLI default unless overridden) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`/`ultra`; ultra enables proactive automatic delegation), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
444
+ | `grok_agent` | Grok Build, model `grok-4.5` by default (`model?` overrides) (xAI) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (vendor-supported value), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
445
+ | `composer_agent` | Grok Build, model `grok-composer-2.5-fast` by default (`model?` overrides); every model is live-catalog checked (xAI) | `prompt?`, `throughline_source_session?`, `model?`, `reasoning_effort?` (vendor-supported value), `env_vars?`, `cwd?`, `session_name?`, `write_scope?` |
424
446
 
425
- The vendor CLI must be installed and authenticated (`claude` for `claude_agent`; `codex` for `codex_agent`; `grok` for both Grok tools). Binary resolution uses `CLAUDE_BIN` / `CODEX_BIN` / `GROK_BIN`, then each documented default location, then `PATH`. Missing binaries, invalid model/effort values, and nonexistent `cwd` fail before a session is created. Claude additionally requires a structured healthy authentication status before any PTY exists, and correlated Claude sessions reject `/login` and `/logout`; repair authentication once in a normal terminal. All four launchers share the normal project/user environment and the same non-blocking dispatch contract. Claude, Codex, Grok, and Composer depth-1 live smokes and a Claude depth-2 nested-delegation smoke are green; fixture coverage remains a separate claim. Native Windows can launch agents but correlated completion is not supported yet.
447
+ The vendor CLI must be installed and authenticated (`claude` for `claude_agent`; `codex` for `codex_agent`; `grok` for both Grok tools). Binary resolution uses `CLAUDE_BIN` / `CODEX_BIN` / `GROK_BIN`, then each documented default location, then `PATH`. Missing binaries, invalid model/effort values, unavailable Grok/Composer catalog models, and nonexistent `cwd` fail before a session is created. Claude additionally requires a structured healthy authentication status before any PTY exists, and correlated Claude sessions reject `/login` and `/logout`; repair authentication once in a normal terminal. All four launchers share the normal project/user environment and the same non-blocking dispatch contract. Claude, Codex, Grok, and Composer depth-1 live smokes and a Claude depth-2 nested-delegation smoke are green; fixture coverage remains a separate claim. Native Windows can launch agents but correlated completion is not supported yet.
426
448
 
427
449
  Set `throughline_source_session` together with a non-empty mission in `prompt` to prepend
428
450
  Throughline's read-only handoff context. This optional route requires `throughline >= 0.9.0`,
@@ -434,7 +456,7 @@ When an agent's answer is longer than the on-screen tail (pane height ≈ 24 lin
434
456
 
435
457
  ### Completion detection (5 layers)
436
458
 
437
- `pty_read({ wait: true })` decides "is the command done?" via five layers: process exit / a `mark:true` sentinel / an `until` match / output quiescence with shell return / timeout. Agent sessions add a sixth exact layer: Codex observes normal rollout `task_complete`; Grok/Composer observe normal session `turn_ended`; Claude observes its additive launch-correlated Stop event. `aiterm-wait --cursor` performs that vendor-specific observation without the parent blocking or polling. Pre-send readiness failures are MCP errors, and late completion remains recoverable without resending.
459
+ `pty_read({ wait: true })` decides "is the command done?" via five layers: process exit / a `mark:true` sentinel / an `until` match / output quiescence with shell return / timeout. When `mark` or `until` is active, that requested evidence takes precedence and a momentarily quiet shell cannot complete the read as quiescent. Agent sessions add a sixth exact layer: Codex observes normal rollout `task_complete`; Grok/Composer observe normal session `turn_ended`; Claude observes its additive launch-correlated Stop event. `aiterm-wait --cursor` performs that vendor-specific observation without the parent blocking or polling. Pre-send readiness failures are MCP errors, and late completion remains recoverable without resending.
438
460
 
439
461
  ### Completion push for parent agents (`aiterm-wait`)
440
462
 
@@ -459,7 +481,7 @@ As of v0.16 a parent agent **never blocks** on aiterm — there is no wait param
459
481
 
460
482
  Before sending, `pty_send` blocks destructive commands (`rm -rf /`, `mkfs`, `dd of=/dev/…`, `DROP TABLE`, …) — pass `force: true` to override — and sanitizes ESC / bracketed-paste terminators. `pty_read` neutralizes control characters in what it returns by default (`raw: true` returns the bytes verbatim). This is a **tripwire, not a sandbox** (see [Known constraints](#known-constraints-by-design-not-bugs)).
461
483
 
462
- Each `pty_send` accepts at most 64 KiB of UTF-8 text. Sends to the same session are serialized across aiterm processes so chunks cannot interleave. On macOS, text is pasted through tmux in UTF-8-safe 256-byte chunks to avoid the platform PTY truncation observed with long input; Linux and WSL use one bounded paste. Sanitized multiline text sent while a POSIX shell is in the foreground is encoded as one newline-free `eval` input: the shell receives the complete script before it runs the first line, so a pager or REPL started mid-script cannot consume later lines as interactive keystrokes. Single-line input, `raw:true`, and non-shell frontends remain direct PTY pastes. Agent dispatches additionally paste with tmux bracketed paste (`paste-buffer -p`): panes that requested bracketed-paste mode (the vendor TUIs) receive each chunk wrapped in `ESC[200~/201~`, hardening prompt injection against mid-word key-interpretation corruption and dropped submits. If a later chunk fails, aiterm reports the partial-send state and does not press Enter automatically. A lock left by a terminated sender fails closed before sending; close and recreate that session (or use `pty_kill_all` when every session is disposable) to clean it up safely.
484
+ Each `pty_send` accepts at most 64 KiB of UTF-8 text. Sends to the same session are serialized across aiterm processes so chunks cannot interleave. Every OS pastes through tmux in UTF-8-safe 256-byte chunks with a 10 ms drain interval; macOS, Linux, and WSL2 have all demonstrated silent middle/trailing loss when a long input is pushed without that boundary. Sanitized multiline text sent while a POSIX shell is in the foreground is encoded as one newline-free `eval` input: the shell receives the complete script before it runs the first line, so a pager or REPL started mid-script cannot consume later lines as interactive keystrokes. Single-line input, `raw:true`, and non-shell frontends remain direct PTY pastes. Agent dispatches additionally paste with tmux bracketed paste (`paste-buffer -p`): panes that requested bracketed-paste mode (the vendor TUIs) receive each chunk wrapped in `ESC[200~/201~`, hardening prompt injection against mid-word key-interpretation corruption and dropped submits. If a later chunk fails, aiterm reports the partial-send state and does not press Enter automatically. A lock left by a terminated sender fails closed before sending; close and recreate that session (or use `pty_kill_all` when every session is disposable) to clean it up safely.
463
485
 
464
486
  ## A human can watch
465
487
 
@@ -494,6 +516,13 @@ npm test # build, then the node:test regression suite (requires tmux)
494
516
  npm link # put `aiterm-mcp` on PATH locally
495
517
  ```
496
518
 
519
+ Development uses focused local tests first. The final GitHub Actions gate starts the same full
520
+ `npm test` concurrently on self-hosted macOS native, Linux native, Windows native, and WSL2
521
+ runners; it does not replace any OS with a reduced suite. Tag-triggered npm publishing runs only
522
+ after all four environments pass and the tagged commit is confirmed on `origin/main`. The native
523
+ Windows runner must run as the interactive Windows user that owns the initialized WSL distro;
524
+ `NETWORK SERVICE` cannot see that user's WSL/tmux environment and is not a valid runner identity.
525
+
497
526
  Logic lives in `src/core.ts` (tmux control, reduction, completion detection, safety, agent launch) and `src/rtk.ts` (per-command reducers); `src/index.ts` is the MCP surface. The design origin and the reducer's porting source (the pytest reducer is ported to match upstream rtk 0.42.0, except the deliberate `FAILED`-line difference noted above, and is locked by regression tests) are in `prototype/python/`.
498
527
 
499
528
  ## Try it
@@ -504,10 +533,10 @@ One command, no clone, no build:
504
533
  claude mcp add --scope user --transport stdio aiterm -- npx -y aiterm-mcp
505
534
  ```
506
535
 
507
- If aiterm let your AI hand a task to another agent — or saved you a round-trip of tokens — **[star the repo](https://github.com/kitepon-rgb/aiterm-mcp)**. It's the cheapest way to help others find it.
536
+ If aiterm let your AI hand a task to another agent — or saved you a round-trip of tokens — **[star the repo](https://github.com/kitepon/aiterm-mcp)**. It's the cheapest way to help others find it.
508
537
 
509
538
  - **npm:** https://www.npmjs.com/package/aiterm-mcp
510
- - **Issues / bug reports:** https://github.com/kitepon-rgb/aiterm-mcp/issues
539
+ - **Issues / bug reports:** https://github.com/kitepon/aiterm-mcp/issues
511
540
 
512
541
  ## Shared agent environment
513
542
 
@@ -516,6 +545,11 @@ symlink, filter, or replace vendor configuration, authentication, MCP, plugin, s
516
545
  trust, memory, or history stores. Cleanup removes only aiterm-owned launch metadata and completion
517
546
  correlation files.
518
547
 
548
+ The ordinary environment still comes from the shell/tmux session. When a caller needs a value that
549
+ belongs to the current MCP process rather than the older persistent tmux server, every launcher
550
+ accepts `env_vars: ["NAME", ...]`. Only those names are refreshed at launch; this is a narrow
551
+ per-launch overlay, not a replacement environment or configuration snapshot.
552
+
519
553
  ## License
520
554
 
521
555
  MIT
package/dist/core.js CHANGED
@@ -73,6 +73,8 @@ const AGENT_SUBMIT_RESIDUE_MAX_SAMPLES = 5;
73
73
  const AGENT_SUBMIT_RESIDUE_TAIL_CHARS = 32;
74
74
  const AGENT_SUBMIT_RESIDUE_MIN_TAIL_CHARS = 8;
75
75
  const GROK_AUTH_MAX_BYTES = 64 * 1024;
76
+ const GROK_MODELS_MAX_BYTES = 1024 * 1024;
77
+ const GROK_MODELS_TIMEOUT_MS = 15_000;
76
78
  const CLAUDE_RESULT_MAX_BYTES = 4 * 1024 * 1024;
77
79
  // 出力削減(RTK の CAP 思想を移植)
78
80
  const MAX_LINES_BEFORE_ELIDE = 60;
@@ -84,9 +86,11 @@ const LINE_TAIL_CHARS = 600;
84
86
  const DEDUP_MIN_RUN = 3; // 同一行がこれ以上連続したら 1 行+件数に畳む
85
87
  const MAX_FULL_BYTES = 8 * 1024 * 1024; // full/range 読取で一度にメモリへ載せる上限(B7)
86
88
  const MAX_SEND_BYTES = 64 * 1024;
87
- // macOSのPTY入力queueは、tmuxが長文を1回で流すと後半を落とすことがある。
89
+ // PTY入力queueはOSを問わず、tmuxが長文を1回で流すと後半を落とすことがある。
88
90
  // UTF-8境界を守って小さいtmux client roundtripに分け、各回にserver event loopがPTYへdrainできる境界を作る。
89
- const PTY_PASTE_CHUNK_BYTES = process.platform === "darwin" ? 256 : MAX_SEND_BYTES;
91
+ const PTY_PASTE_CHUNK_BYTES = 256;
92
+ const PTY_PASTE_CHUNK_PAUSE_MS = 10;
93
+ const PTY_PASTE_PAUSE_BUFFER = new Int32Array(new SharedArrayBuffer(4));
90
94
  const SESSION_SEND_LOCK_WAIT_MS = 10_000;
91
95
  const SESSION_SEND_LOCK_POLL_MS = 25;
92
96
  // 安全: send 前に弾く破壊的コマンド(外部システム境界の防御)
@@ -785,7 +789,7 @@ async function waitCompletion(name, untilStr, untilRegex, timeout) {
785
789
  if (safeStatSize(logpath(name)) !== size) {
786
790
  stable = 0;
787
791
  }
788
- else if (SHELLS.has(fg)) {
792
+ else if (!until && !markActive && SHELLS.has(fg)) {
789
793
  if (isWin)
790
794
  await settleWinLog(name);
791
795
  return [true, "quiescent"]; // 出力静止 ∧ シェル復帰 = 確証つき完了
@@ -794,7 +798,8 @@ async function waitCompletion(name, untilStr, untilRegex, timeout) {
794
798
  // ネスト中(前面が ssh/docker/REPL 等でシェル集合外)は quiescence の「シェル復帰」条件を
795
799
  // 原理的に満たせない。until も mark も無ければこれ以上待っても確証は増えない(until/dead/
796
800
  // quiescent/mark のいずれも発火し得ない)ので、出力静止時点で「未確定」のまま早期返却する。
797
- // markActive のときは sentinel を待つべく早期返却せず、非シェル前面(sleep 等)でも待ち続ける。
801
+ // until/markActive のときは指定した証拠を待つべく早期返却せず、shell builtinや
802
+ // 非シェル前面(sleep 等)でも待ち続ける。
798
803
  // fg==="" は前面コマンド取得失敗=ネスト断定不可なので早期返却せず従来どおり timeout まで待つ。
799
804
  if (isWin)
800
805
  await settleWinLog(name);
@@ -1002,9 +1007,9 @@ export function send(name, text, o = {}) {
1002
1007
  /* noop */
1003
1008
  }
1004
1009
  }
1005
- // `send-keys -l`と単発`paste-buffer`は、長文をPTY入力queueへ一度に流し、macOS CIで
1006
- // 途中以降が欠落しても tmux 自体は code=0 を返した。macOSだけUTF-8を壊さない256byte以下に分け、
1007
- // Linux/WSLは一括のままとする。全chunk+Enterはsession単位のcross-process lock内で直列化する。
1010
+ // `send-keys -l`と単発`paste-buffer`は、長文をPTY入力queueへ一度に流すと、OSを問わず
1011
+ // 途中以降が欠落しても tmux 自体は code=0 を返す。UTF-8を壊さない256byte以下に分け、
1012
+ // 全chunk+Enterをsession単位のcross-process lock内で直列化する。
1008
1013
  const pasteSupportsNoSanitize = pasteBufferSupportsNoSanitizeFlag();
1009
1014
  const chunks = splitPtyText(text);
1010
1015
  const bufferBase = `aiterm-${process.pid}-${randomBytes(8).toString("hex")}`;
@@ -1034,6 +1039,11 @@ export function send(name, text, o = {}) {
1034
1039
  throw new AitermError(`tmux bufferのPTY送信に失敗しました` +
1035
1040
  `(chunk ${i + 1}/${chunks.length}): ${pasted.stderr.trim() || `code=${pasted.code}`}.${partial}`, 2);
1036
1041
  }
1042
+ if (i + 1 < chunks.length) {
1043
+ // tmux clientの終了だけでは、serverが前chunkをPTYへdrainしたことを保証しない。
1044
+ // 次chunkまで短く間を空け、外部PTY入力queueの飽和による黙った末尾欠落を防ぐ。
1045
+ Atomics.wait(PTY_PASTE_PAUSE_BUFFER, 0, 0, PTY_PASTE_CHUNK_PAUSE_MS);
1046
+ }
1037
1047
  }
1038
1048
  if (enter) {
1039
1049
  const entered = tmux("send-keys", "-t", name, "Enter");
@@ -1659,6 +1669,41 @@ function resolveAndValidateGrokAuth(srcHome) {
1659
1669
  fs.closeSync(fd);
1660
1670
  }
1661
1671
  }
1672
+ function grokModelCatalog(bin, cwd) {
1673
+ const result = spawnAgentControlCommand(bin, ["models"], cwd, {
1674
+ cwd,
1675
+ encoding: "utf8",
1676
+ env: process.env,
1677
+ timeout: GROK_MODELS_TIMEOUT_MS,
1678
+ maxBuffer: GROK_MODELS_MAX_BYTES,
1679
+ windowsHide: true,
1680
+ });
1681
+ if (result.error || result.status !== 0) {
1682
+ const detail = result.error?.message || result.stderr?.trim() || `exit=${result.status ?? "unknown"}`;
1683
+ throw new AitermError(`Grok model catalog を取得できません: ${detail}`, 2);
1684
+ }
1685
+ const text = result.stdout.replace(/\x1b\[[0-?]*[ -/]*[@-~]/g, "");
1686
+ const marker = "Available models:";
1687
+ const start = text.indexOf(marker);
1688
+ if (start < 0)
1689
+ throw new AitermError("Grok model catalog の出力形式が不正です(Available models がありません)", 2);
1690
+ const models = [];
1691
+ for (const line of text.slice(start + marker.length).split(/\r?\n/)) {
1692
+ const match = line.match(/^\s{2}[-*]\s+(.+?)(?:\s+\(default\))?\s*$/);
1693
+ if (match?.[1])
1694
+ models.push(match[1]);
1695
+ }
1696
+ if (models.length === 0)
1697
+ throw new AitermError("Grok model catalog に利用可能なmodelがありません", 2);
1698
+ return models;
1699
+ }
1700
+ function assertGrokModelAvailable(bin, cwd, model) {
1701
+ const models = grokModelCatalog(bin, cwd);
1702
+ if (!models.includes(model)) {
1703
+ throw new AitermError(`Grok model catalog に ${JSON.stringify(model)} がありません。利用可能: ${models.join(", ")}。` +
1704
+ "別modelへfallbackせず起動を中止しました", 2);
1705
+ }
1706
+ }
1662
1707
  function writeAgentMetadata(meta) {
1663
1708
  writeJson0600(agentMetadataPath(meta.aiterm_session, meta.launch_id), meta);
1664
1709
  }
@@ -3326,7 +3371,7 @@ function isAgentTuiReady(kind, screen) {
3326
3371
  // 起動直後は製品header、長寿命sessionでは常駐footerがCodex TUIの識別子になる。
3327
3372
  // capture-paneは直近45行だけなので、会話が進むとheaderは正常に画面外へ流れる。
3328
3373
  const codexFrontend = screen.includes("OpenAI Codex")
3329
- || /(^|\n)\s*\S+\s+(?:low|medium|high|xhigh|max|ultra)\s+·\s+\S.*$/m.test(screen);
3374
+ || /(^|\n)\s*\S+\s+(?:low|medium|high|xhigh|max|ultra)(?:\s+fast)?\s+·\s+\S.*$/m.test(screen);
3330
3375
  return codexFrontend && /(^|\n)\s*[›>]/.test(screen);
3331
3376
  }
3332
3377
  // Grok Build 0.2.117 は起動完了後に製品名を消し、model footerだけを残す。
@@ -3577,6 +3622,35 @@ async function waitForScreenText(name, text, timeoutMs = 3_000) {
3577
3622
  } while (performance.now() < deadline);
3578
3623
  throw new AitermError(`${agentLabel(loadAgentMetadata(name).kind)} の ${text} 画面を確認できません`, 2);
3579
3624
  }
3625
+ function lineCounts(screen) {
3626
+ const counts = new Map();
3627
+ for (const line of screen.split("\n").map((value) => value.trim()).filter(Boolean)) {
3628
+ counts.set(line, (counts.get(line) ?? 0) + 1);
3629
+ }
3630
+ return counts;
3631
+ }
3632
+ async function waitForGrokConfigurationResult(name, before, timeoutMs = 3_000) {
3633
+ const beforeCounts = lineCounts(before);
3634
+ const deadline = performance.now() + timeoutMs;
3635
+ do {
3636
+ const screen = captureScreen(name, AGENT_TUI_READY_LINES);
3637
+ const seen = new Map();
3638
+ const added = screen.split("\n").map((value) => value.trim()).filter((line) => {
3639
+ if (!line)
3640
+ return false;
3641
+ const count = (seen.get(line) ?? 0) + 1;
3642
+ seen.set(line, count);
3643
+ return count > (beforeCounts.get(line) ?? 0);
3644
+ });
3645
+ const error = added.find((line) => /^(?:Unknown model:|unknown effort level|Usage: \/(?:model|effort)\b|Invalid (?:model|reasoning effort)|.*does not support reasoning effort)/i.test(line));
3646
+ if (error)
3647
+ throw new AitermError(`Grokの設定変更に失敗しました: ${error}`, 2);
3648
+ if (added.some((line) => /^(?:Switched to |✓?\s*Default model:)/.test(line)))
3649
+ return;
3650
+ await sleep(100);
3651
+ } while (performance.now() < deadline);
3652
+ throw new AitermError(`${agentLabel(loadAgentMetadata(name).kind)} の設定変更完了を確認できません`, 2);
3653
+ }
3580
3654
  function sendMenuChoice(name, choice) {
3581
3655
  const sent = tmux("send-keys", "-t", name, choice);
3582
3656
  if (sent.code !== 0) {
@@ -3625,10 +3699,10 @@ export async function configureAgent(name, opts) {
3625
3699
  const effort = opts.reasoning_effort?.trim() || null;
3626
3700
  if (!model && !effort)
3627
3701
  throw new AitermError("model または reasoning_effort を指定してください", 2);
3628
- const meta = loadAgentMetadata(name);
3629
- if (meta.kind !== "claude" && meta.kind !== "codex") {
3630
- throw new AitermError("agent_configure はClaudeとCodexのsessionだけに対応します", 2);
3702
+ if ((model && /[\r\n]/.test(model)) || (effort && /[\r\n]/.test(effort))) {
3703
+ throw new AitermError("model/reasoning_effort に改行を含められません", 2);
3631
3704
  }
3705
+ const meta = loadAgentMetadata(name);
3632
3706
  bindCompletedInitialPrompt(meta);
3633
3707
  const ready = await waitAgentTuiReady(name, meta, AGENT_TUI_READY_TIMEOUT_MS);
3634
3708
  if (!ready.ready)
@@ -3654,6 +3728,27 @@ export async function configureAgent(name, opts) {
3654
3728
  reasoning_effort: effort,
3655
3729
  };
3656
3730
  }
3731
+ if (meta.kind === "grok" || meta.kind === "composer") {
3732
+ if (model) {
3733
+ const bin = resolveAgentBin(meta.kind);
3734
+ if (!bin)
3735
+ throw new AitermError(`${agentLabel(meta.kind)} の CLI が見つかりません`, 2);
3736
+ assertGrokModelAvailable(bin, meta.cwd ?? process.cwd(), model);
3737
+ }
3738
+ const command = model
3739
+ ? `/model ${model}${effort ? ` ${effort}` : ""}`
3740
+ : `/effort ${effort}`;
3741
+ const before = captureScreen(name, AGENT_TUI_READY_LINES);
3742
+ await sendAgentPromptText(name, command);
3743
+ await waitForGrokConfigurationResult(name, before);
3744
+ return {
3745
+ schema: "aiterm.agent-configure-result.v1",
3746
+ session_id: name,
3747
+ provider: meta.kind,
3748
+ model,
3749
+ reasoning_effort: effort,
3750
+ };
3751
+ }
3657
3752
  await sendAgentPromptText(name, "/model");
3658
3753
  const modelScreen = await waitForScreenText(name, "Select Model and Effort");
3659
3754
  if (model) {
@@ -3828,27 +3923,32 @@ function resolveAgentBin(kind) {
3828
3923
  // 明示指定 env は実在を検証する。存在しないパスを黙って返すと、session を作って
3829
3924
  // `'/typo' ...` を送信し bash が command not found を出すだけで openAgent は「起動した」と
3830
3925
  // 偽成功を返す(既定パス/PATH 経路は検証するのに env だけ無検証だった非対称の解消・A3)。
3831
- if (isUsableExecutableFile(fromEnv))
3832
- return fromEnv;
3926
+ if (isUsableAgentExecutableFile(fromEnv))
3927
+ return resolveWindowsCodexShim(kind, fromEnv);
3833
3928
  throw new AitermError(`${envVar} に指定された ${name} が存在しません: ${fromEnv}`, 2);
3834
3929
  }
3835
3930
  const cand = path.join(home, ...rel);
3836
- if (isUsableExecutableFile(cand))
3837
- return cand;
3931
+ if (isUsableAgentExecutableFile(cand))
3932
+ return resolveWindowsCodexShim(kind, cand);
3838
3933
  const w = spawnSync(isWin ? "where" : "which", [name], {
3839
3934
  encoding: "utf8",
3840
3935
  timeout: 5000,
3841
3936
  });
3842
3937
  if (w.status === 0 && (w.stdout ?? "").trim()) {
3843
- const resolved = w.stdout.trim().split(/\r?\n/)[0];
3844
- if (isUsableExecutableFile(resolved))
3845
- return resolved;
3938
+ const found = w.stdout.trim().split(/\r?\n/).filter(Boolean);
3939
+ const ordered = isWin
3940
+ ? [...found.filter((p) => /\.(?:exe|com|cmd|bat)$/i.test(p)), ...found.filter((p) => !/\.(?:exe|com|cmd|bat)$/i.test(p))]
3941
+ : found;
3942
+ for (const resolved of ordered) {
3943
+ if (isUsableAgentExecutableFile(resolved))
3944
+ return resolveWindowsCodexShim(kind, resolved);
3945
+ }
3846
3946
  }
3847
3947
  return null;
3848
3948
  }
3849
3949
  const CLAUDE_AUTH_STATUS_TIMEOUT_MS = 5_000;
3850
3950
  function assertClaudeAuthenticationReady(bin) {
3851
- const result = spawnSync(bin, ["auth", "status", "--json"], {
3951
+ const result = spawnAgentControlCommand(bin, ["auth", "status", "--json"], process.cwd(), {
3852
3952
  encoding: "utf8",
3853
3953
  timeout: CLAUDE_AUTH_STATUS_TIMEOUT_MS,
3854
3954
  maxBuffer: 64 * 1024,
@@ -3893,6 +3993,71 @@ function isUsableExecutableFile(candidate) {
3893
3993
  return false;
3894
3994
  }
3895
3995
  }
3996
+ function isWindowsDrivePath(candidate) {
3997
+ return /^[A-Za-z]:[\\/]/.test(candidate);
3998
+ }
3999
+ function isWindowsNativeExecutable(candidate) {
4000
+ return isWindowsDrivePath(candidate) && /\.(?:exe|com|cmd|bat)$/i.test(candidate);
4001
+ }
4002
+ function isUsableWslExecutable(candidate) {
4003
+ if (!isWin)
4004
+ return false;
4005
+ let wslPath = candidate;
4006
+ if (isWindowsDrivePath(candidate)) {
4007
+ try {
4008
+ if (!fs.statSync(candidate).isFile())
4009
+ return false;
4010
+ wslPath = toWslPath(candidate);
4011
+ }
4012
+ catch {
4013
+ return false;
4014
+ }
4015
+ }
4016
+ else if (!candidate.startsWith("/")) {
4017
+ return false;
4018
+ }
4019
+ const checked = spawnSync("wsl.exe", ["-e", "test", "-f", wslPath, "-a", "-x", wslPath], {
4020
+ encoding: "utf8",
4021
+ timeout: 5000,
4022
+ });
4023
+ return checked.status === 0;
4024
+ }
4025
+ function isUsableAgentExecutableFile(candidate) {
4026
+ if (!isWin)
4027
+ return isUsableExecutableFile(candidate);
4028
+ if (isWindowsNativeExecutable(candidate))
4029
+ return isUsableExecutableFile(candidate);
4030
+ return isUsableWslExecutable(candidate);
4031
+ }
4032
+ function resolveWindowsCodexShim(kind, candidate) {
4033
+ if (!isWin || kind !== "codex" || !/\.(?:cmd|bat)$/i.test(candidate))
4034
+ return candidate;
4035
+ const packageRoot = path.join(path.dirname(candidate), "node_modules", "@openai", "codex", "node_modules", "@openai");
4036
+ try {
4037
+ for (const platformPackage of fs.readdirSync(packageRoot).filter((name) => name.startsWith("codex-win32-"))) {
4038
+ const vendorRoot = path.join(packageRoot, platformPackage, "vendor");
4039
+ for (const target of fs.readdirSync(vendorRoot)) {
4040
+ const executable = path.join(vendorRoot, target, "bin", "codex.exe");
4041
+ if (isUsableExecutableFile(executable))
4042
+ return executable;
4043
+ }
4044
+ }
4045
+ }
4046
+ catch {
4047
+ /* 下の明示エラーへ */
4048
+ }
4049
+ throw new AitermError(`CODEX_BIN のnpm shimからWindows native codex.exeを解決できません: ${candidate}。` +
4050
+ "@openai/codexを再インストールするか、CODEX_BINへcodex.exeを指定してください", 2);
4051
+ }
4052
+ function agentBinForWslShell(bin) {
4053
+ return isWin && isWindowsDrivePath(bin) ? toWslPath(bin) : bin;
4054
+ }
4055
+ function spawnAgentControlCommand(bin, args, cwd, options) {
4056
+ if (!isWin || isWindowsNativeExecutable(bin))
4057
+ return spawnSync(bin, args, options);
4058
+ const wslCwd = isWindowsDrivePath(cwd) ? toWslPath(cwd) : cwd;
4059
+ return spawnSync("wsl.exe", ["--cd", wslCwd, "-e", agentBinForWslShell(bin), ...args], options);
4060
+ }
3896
4061
  const THROUGHLINE_HANDOFF_CONTEXT_SCHEMA = "throughline.handoff_context.v1";
3897
4062
  const PORTABLE_FORK_MISSION_SEPARATOR = "\n\n---\n\n## Portable fork mission\n\n";
3898
4063
  function resolveThroughlineBin() {
@@ -4004,12 +4169,16 @@ function buildAgentCmd(kind, bin, model, effort, prompt, meta = null) {
4004
4169
  }
4005
4170
  }
4006
4171
  else {
4007
- // grok / composer は同じ grok CLI をモデル違いで起動。--effort は headless(grok -p)専用で
4008
- // 対話 TUI では警告の上無視されるため渡さない(openAgent が指定を事前拒否する)。
4172
+ // grok / composer は同じ grok CLI をモデル違いで起動する。
4009
4173
  parts.push("--no-auto-update");
4010
4174
  if (meta?.kind === "grok" || meta?.kind === "composer")
4011
4175
  parts.push("--no-alt-screen");
4012
4176
  parts.push("--model", shq(model ?? GROK_MODEL_DEFAULTS[kind]));
4177
+ if (effort)
4178
+ parts.push("--reasoning-effort", shq(effort));
4179
+ if ((meta?.kind === "grok" || meta?.kind === "composer") && meta.write_scope === "read-only") {
4180
+ parts.push("--sandbox", "read-only");
4181
+ }
4013
4182
  if ((meta?.kind === "grok" || meta?.kind === "composer") && meta.hook_route === "shared_grok_home") {
4014
4183
  parts.push("--session-id", shq(meta.vendor_session_id ?? ""), "--rules", shq(subagentInstruction(meta)));
4015
4184
  }
@@ -4020,10 +4189,15 @@ function buildAgentCmd(kind, bin, model, effort, prompt, meta = null) {
4020
4189
  parts.push(shq(prompt)); // 初手プロンプト(任意)
4021
4190
  return parts.join(" ");
4022
4191
  }
4023
- function agentEnvPrefix(meta, sid) {
4192
+ function agentEnvPrefix(meta, sid, envVars = []) {
4193
+ const inherited = envVars.flatMap((name) => {
4194
+ const value = process.env[name];
4195
+ return value === undefined ? [] : [`${name}=${shq(value)}`];
4196
+ });
4024
4197
  if (!meta)
4025
- return "";
4198
+ return inherited.length ? inherited.join(" ") + " " : "";
4026
4199
  const common = [
4200
+ ...inherited,
4027
4201
  `AITERM_AGENT_KIND=${shq(meta.kind)}`,
4028
4202
  `AITERM_SESSION_ID=${shq(sid)}`,
4029
4203
  `AITERM_AGENT_SESSION_ID=${shq(sid)}`,
@@ -4061,15 +4235,15 @@ function agentLabel(kind) {
4061
4235
  function buildAgentLaunchNote(kind, model, effort, meta) {
4062
4236
  const writeScopeNote = meta?.write_scope === undefined
4063
4237
  ? ""
4064
- : kind === "codex" && meta.write_scope === "read-only"
4065
- ? `\n能力宣言: write_scope=${JSON.stringify(meta.write_scope)}。Codex CLIへ --sandbox read-only を付与し、書込みを実効禁止。`
4066
- : `\n能力宣言: write_scope=${JSON.stringify(meta.write_scope)}。${kind === "grok" || kind === "composer" ? "このCLIには起動sandbox機構がないため" : "パス単位のsandbox allowlistに対応するCLI引数がないため"}宣言の記録のみ(構造的unsupported)。`;
4238
+ : (kind === "codex" || kind === "grok" || kind === "composer") && meta.write_scope === "read-only"
4239
+ ? `\n能力宣言: write_scope=${JSON.stringify(meta.write_scope)}。${agentLabel(kind)} CLIへ --sandbox read-only を付与し、書込みを実効禁止。`
4240
+ : `\n能力宣言: write_scope=${JSON.stringify(meta.write_scope)}。パス単位のsandbox allowlistに対応するCLI引数がないため宣言の記録のみ(構造的unsupported)。`;
4067
4241
  if (kind === "claude") {
4068
4242
  return `起動設定: model=${model ?? "CLI既定"} effort=${effort ?? "CLI既定"}。${writeScopeNote}`;
4069
4243
  }
4070
4244
  if (kind !== "codex") {
4071
4245
  return (`起動設定: model=${model ?? GROK_MODEL_DEFAULTS[kind]}(${model ? "引数" : "ツール既定"})。` +
4072
- "reasoning effort は対話 TUI 非対応=未指定で起動。" + writeScopeNote);
4246
+ `effort=${effort ?? "CLI/model既定"}。` + writeScopeNote);
4073
4247
  }
4074
4248
  const configPath = meta?.kind === "codex" && meta.codex_home
4075
4249
  ? path.join(meta.codex_home, "config.toml")
@@ -4131,21 +4305,15 @@ export function openAgent(kind, opts = {}) {
4131
4305
  }
4132
4306
  const effort = opts.reasoning_effort ?? null;
4133
4307
  const writeScope = opts.write_scope;
4308
+ const envVars = opts.env_vars ?? [];
4309
+ for (const name of envVars) {
4310
+ if (!/^[A-Za-z_][A-Za-z0-9_]*$/.test(name)) {
4311
+ throw new AitermError(`env_vars に無効な環境変数名があります: ${JSON.stringify(name)}`, 2);
4312
+ }
4313
+ }
4134
4314
  if (effort && kind === "claude" && !CLAUDE_EFFORTS.has(effort)) {
4135
4315
  throw new AitermError("Claude Code の reasoning_effort は low/medium/high/xhigh/max のいずれかです", 2);
4136
4316
  }
4137
- // grok CLI の --effort は headless(grok -p)専用で、対話 TUI では警告の上無視される。
4138
- // 黙って no-op の引数を受けない=起動前に明示エラーで拒否する(codex は CLI 側の値集合が
4139
- // 版で変わるため縛らず送信まで通す)。
4140
- if (effort && kind === "grok") {
4141
- throw new AitermError(`${label} は reasoning_effort を指定できません。grok CLI の --effort は headless(grok -p)専用で、` +
4142
- "対話 TUI では警告の上無視されます(grok-4.5 の TUI 既定 effort は high)。" +
4143
- "effort 制御が必要なら通常 PTY で `grok -p --effort low|medium|high ...` を使ってください", 2);
4144
- }
4145
- if (effort && kind === "composer") {
4146
- throw new AitermError(`${label} は reasoning_effort を指定できません。grok-composer-2.5-fast は reasoning effort 非対応です` +
4147
- "(モデルカタログ supports_reasoning_effort=false)", 2);
4148
- }
4149
4317
  const agentDone = !!opts.agent_done;
4150
4318
  const launchOperationId = opts.launch_operation_id == null
4151
4319
  ? null
@@ -4220,9 +4388,14 @@ export function openAgent(kind, opts = {}) {
4220
4388
  // パスで解決)。toWslPath は session を作る前に呼ぶ=変換失敗(非ドライブパス)で残骸 session を残さない。
4221
4389
  // 未検証リスク: npm グローバル導入の codex.cmd/.bat シムや WSL interop 上の対話 TUI 描画は実 Windows
4222
4390
  // でしか確認できない(CI 非対象。docs/03_audit-sweep-2026-07.md 参照)。
4223
- const binForCmd = isWin ? toWslPath(bin) : bin;
4391
+ const binForCmd = agentBinForWslShell(bin);
4224
4392
  const cwdForCmd = cwd && isWin ? toWslPath(cwd) : cwd;
4225
4393
  const grokAuthPath = agentDone && (kind === "grok" || kind === "composer") ? resolveAndValidateGrokAuth(realGrokHome()) : null;
4394
+ if (kind === "grok" || kind === "composer") {
4395
+ const requestedModel = model ?? (kind === "composer" ? GROK_MODEL_DEFAULTS.composer : null);
4396
+ if (requestedModel)
4397
+ assertGrokModelAvailable(bin, cwd ?? process.cwd(), requestedModel);
4398
+ }
4226
4399
  let sid;
4227
4400
  let hint;
4228
4401
  try {
@@ -4249,7 +4422,7 @@ export function openAgent(kind, opts = {}) {
4249
4422
  agentMetadataNegativeCache.delete(sid);
4250
4423
  launchNote = buildAgentLaunchNote(kind, model, effort, meta);
4251
4424
  const cmd = buildAgentCmd(kind, binForCmd, model, effort, opts.prompt ?? null, meta);
4252
- const envPrefix = agentEnvPrefix(meta, sid);
4425
+ const envPrefix = agentEnvPrefix(meta, sid, envVars);
4253
4426
  const full = cwdForCmd ? `cd ${shq(cwdForCmd)} && ${envPrefix}${cmd}` : `${envPrefix}${cmd}`;
4254
4427
  // force:true で送る。起動骨格は `bin '...'` の固定形で、prompt/cwd/effort は shq でクオート済みの
4255
4428
  // 引数=シェルは決して破壊コマンドとして実行しない。破壊ゲート(生シェルコマンド想定)を prompt に
@@ -4320,6 +4493,7 @@ export async function openAgentWithInitialPrompt(kind, opts = {}) {
4320
4493
  agent_done: true,
4321
4494
  launch_operation_id: opts.launch_operation_id ?? null,
4322
4495
  write_scope: opts.write_scope,
4496
+ env_vars: opts.env_vars,
4323
4497
  });
4324
4498
  // argv prompt(grok/composer)は composer を経由しないため submit 座礁観測の対象外。
4325
4499
  return [sid, hint, prompt ? 0 : null, null];
@@ -4333,6 +4507,7 @@ export async function openAgentWithInitialPrompt(kind, opts = {}) {
4333
4507
  agent_done: true,
4334
4508
  launch_operation_id: opts.launch_operation_id ?? null,
4335
4509
  write_scope: opts.write_scope,
4510
+ env_vars: opts.env_vars,
4336
4511
  });
4337
4512
  try {
4338
4513
  const initial = await sendInitialAgentPrompt(sid, prompt, {
package/dist/index.js CHANGED
@@ -397,8 +397,8 @@ server.registerTool("claude_approval", {
397
397
  }
398
398
  });
399
399
  server.registerTool("agent_configure", {
400
- description: "起動済みのCodex/Claude agent sessionを再起動せず、会話contextを保ったままmodel/reasoning effortを変更する。" +
401
- "ClaudeはCLI標準の/model・/effort、CodexはCLI標準の/model選択画面を使う。",
400
+ description: "起動済みのClaude/Codex/Grok/Composer agent sessionを再起動せず、会話contextを保ったままmodel/reasoning effortを変更する。" +
401
+ "Claude/Grok/ComposerはCLI標準の/model・/effort、CodexはCLI標準の/model選択画面を使う。",
402
402
  inputSchema: {
403
403
  session_id: z.string().regex(/^[A-Za-z0-9_-]{1,64}$/),
404
404
  model: z.string().min(1).nullish().describe("変更後のmodel。省略時はmodelを変更しない"),
@@ -407,7 +407,7 @@ server.registerTool("agent_configure", {
407
407
  outputSchema: {
408
408
  schema: z.literal("aiterm.agent-configure-result.v1"),
409
409
  session_id: z.string().regex(/^[A-Za-z0-9_-]{1,64}$/),
410
- provider: z.enum(["claude", "codex"]),
410
+ provider: z.enum(["claude", "codex", "grok", "composer"]),
411
411
  model: z.string().nullable(),
412
412
  reasoning_effort: z.string().nullable(),
413
413
  },
@@ -430,14 +430,16 @@ const agentModelDesc = (kind) => kind === "claude"
430
430
  : kind === "codex"
431
431
  ? "起動モデル(例: gpt-5.6-sol / gpt-5.6-terra / gpt-5.6-luna)。省略時は端末 config/CLI 既定を継承" +
432
432
  "(端末側のピンがそのまま効く。実効値は起動応答に明示される)"
433
- : `起動モデル。省略時は ${kind === "grok" ? "grok-4.5" : "grok-composer-2.5-fast"}`;
433
+ : `起動モデル。省略時は ${kind === "grok" ? "grok-4.5" : "grok-composer-2.5-fast"}。` +
434
+ (kind === "composer"
435
+ ? "既定/explicit modelを起動前にlive catalogへ照合し、不在ならfallbackせずエラー"
436
+ : "explicit modelを起動前にlive catalogへ照合し、不在ならfallbackせずエラー");
434
437
  const agentEffortDesc = (kind) => kind === "claude"
435
438
  ? "Claude Code reasoning effort。low/medium/high/xhigh/max。省略時はCLI既定"
436
- : kind === "grok" || kind === "composer"
437
- ? "指定不可(grok CLI の --effort は headless 専用で、対話 TUI では警告の上無視される。" +
438
- "composer は effort 自体非対応)。指定すると起動前にエラーを返す"
439
- : "reasoning effort(思考レベル)。low/medium/high/xhigh/max/ultra(CLI 版依存)。" +
440
- "ultra は max 推論+proactive 自動委譲 ON=使用量急増注意(明示要求時のみ)。省略時は端末 config/CLI 既定。";
439
+ : kind === "codex"
440
+ ? "reasoning effort(思考レベル)。low/medium/high/xhigh/max/ultra(CLI/model 版依存)。" +
441
+ "ultra は max 推論+proactive 自動委譲 ON=使用量急増注意(明示要求時のみ)。省略時は端末 config/CLI 既定。"
442
+ : "Grok Build reasoning effort。利用可能値はCLI/modelのlive catalogに従う。省略時はCLI/model既定。";
441
443
  // 全launcher共通の完了受信ガイド。待ちコマンドは起動応答の wait_command(初回prompt時)または
442
444
  // pty_send dispatch の event_cursor から組む。文型は NON_BLOCKING_RULE と同じく「待たない」が先。
443
445
  const agentCompletionDesc = `起動して投げたら投げっぱなしでよい=親はここで待たない。` +
@@ -453,7 +455,7 @@ function registerAgentTool(toolName, kind, desc) {
453
455
  const supportsWriteScope = kind === "codex" || kind === "grok" || kind === "composer";
454
456
  const writeScopeInputSchema = supportsWriteScope
455
457
  ? {
456
- write_scope: z.string().min(1).optional().describe("能力宣言。read-only、または書込みを許可するパスの説明文字列。Codexのread-onlyだけはCLI sandboxで実効禁止する"),
458
+ write_scope: z.string().min(1).optional().describe("能力宣言。read-only、または書込みを許可するパスの説明文字列。Codex/Grok/Composerのread-onlyはCLI sandboxで実効禁止する"),
457
459
  }
458
460
  : {};
459
461
  const writeScopeOutputSchema = supportsWriteScope
@@ -480,9 +482,9 @@ function registerAgentTool(toolName, kind, desc) {
480
482
  .optional()
481
483
  .describe("同一端末のThroughline sessionから所有権を変えずに記憶を読み、promptのmissionより前へ注入する"),
482
484
  model: z.string().nullish().describe(agentModelDesc(kind)),
483
- // grok/composer の effort は対話 TUI で無効(headless 専用)=core 側が起動前に明示エラーで拒否。
484
- // codex は CLI 側の値集合が版で変わるため縛らない(core 側も同方針)。
485
+ // CLI/model側の値集合が版で変わるため公開enumでは縛らない(core側も同方針)。
485
486
  reasoning_effort: z.string().nullish().describe(agentEffortDesc(kind)),
487
+ env_vars: z.array(z.string()).optional().describe("起動したagentへ現在のMCP processから継承する環境変数名。値はtool引数へ渡さない"),
486
488
  cwd: z.string().nullish().describe("作業ディレクトリ(対象リポのルート等・任意)"),
487
489
  session_name: z.string().nullish().describe("セッション名(省略で自動採番)"),
488
490
  ...writeScopeInputSchema,
@@ -502,13 +504,14 @@ function registerAgentTool(toolName, kind, desc) {
502
504
  submit_residue: z.boolean().nullable(),
503
505
  ...writeScopeOutputSchema,
504
506
  },
505
- }, async ({ prompt, throughline_source_session, model, reasoning_effort, cwd, session_name, launch_operation_id, write_scope }) => {
507
+ }, async ({ prompt, throughline_source_session, model, reasoning_effort, env_vars, cwd, session_name, launch_operation_id, write_scope }) => {
506
508
  try {
507
509
  const [sid, hint, eventCursor, submitResidue] = await core.openAgentWithInitialPrompt(kind, {
508
510
  prompt: prompt ?? undefined,
509
511
  throughline_source_session,
510
512
  model: model ?? undefined,
511
513
  reasoning_effort: reasoning_effort ?? undefined,
514
+ env_vars,
512
515
  cwd: cwd ?? undefined,
513
516
  session_name: session_name ?? undefined,
514
517
  launch_operation_id: launch_operation_id ?? undefined,
@@ -525,7 +528,7 @@ function registerAgentTool(toolName, kind, desc) {
525
528
  ...(supportsWriteScope && write_scope !== undefined
526
529
  ? {
527
530
  write_scope,
528
- write_scope_enforcement: kind === "codex" && write_scope === "read-only"
531
+ write_scope_enforcement: (kind === "codex" || kind === "grok" || kind === "composer") && write_scope === "read-only"
529
532
  ? "enforced_read_only"
530
533
  : "declaration_only_unsupported",
531
534
  }
@@ -561,12 +564,13 @@ registerAgentTool("grok_agent", "grok", "【Grok Build の Grok モデル (既
561
564
  agentEnvironmentDesc +
562
565
  "turn は pty_send で送る(自動で非ブロック dispatch になる)。" +
563
566
  agentCompletionDesc +
564
- "model を引数で指定可。reasoning_effort は対話 TUI 非対応(指定はエラー)。");
567
+ "model/reasoning_effortを引数で指定可。read-only sandboxとagent_configureに対応。");
565
568
  registerAgentTool("composer_agent", "composer", "【Grok Build の Composer モデル (既定 grok-composer-2.5-fast)】の対話エージェント TUI を永続端末に起動する。" +
566
569
  agentEnvironmentDesc +
567
570
  "turn は pty_send で送る(自動で非ブロック dispatch になる)。" +
568
571
  agentCompletionDesc +
569
- "model を引数で指定可。reasoning_effort は非対応(指定はエラー)。");
572
+ "model/reasoning_effortを引数で指定可。live catalogにComposer modelがなければGrokへfallbackせず明示エラー。" +
573
+ "read-only sandboxとagent_configureに対応。");
570
574
  async function main() {
571
575
  // 親ホストを initialize の clientInfo.name から確定させ、receipt の完了待ちコマンドを
572
576
  // そのホストの実際の起動形で名指しする(実測: claude-code は initialize → notifications/initialized
package/package.json CHANGED
@@ -1,8 +1,8 @@
1
1
  {
2
2
  "name": "aiterm-mcp",
3
- "version": "0.24.2",
4
- "mcpName": "io.github.kitepon-rgb/aiterm-mcp",
5
- "description": "Persistent tmux terminal MCP that lets Claude Code drive Codex CLI's interactive TUI, including slash commands and $imagegen. Also runs durable PTY sessions for SSH, containers, REPLs, and coding agents.",
3
+ "version": "0.25.1",
4
+ "mcpName": "io.github.kitepon/aiterm-mcp",
5
+ "description": "Persistent tmux terminal MCP for launching and driving Claude, Codex, Grok, or Composer from any MCP client, cross-vendor or same-vendor. Also runs durable PTY sessions for SSH, containers, and REPLs.",
6
6
  "keywords": [
7
7
  "mcp",
8
8
  "mcp-server",
@@ -31,12 +31,12 @@
31
31
  },
32
32
  "repository": {
33
33
  "type": "git",
34
- "url": "git+https://github.com/kitepon-rgb/aiterm-mcp.git"
34
+ "url": "git+https://github.com/kitepon/aiterm-mcp.git"
35
35
  },
36
36
  "bugs": {
37
- "url": "https://github.com/kitepon-rgb/aiterm-mcp/issues"
37
+ "url": "https://github.com/kitepon/aiterm-mcp/issues"
38
38
  },
39
- "homepage": "https://github.com/kitepon-rgb/aiterm-mcp#readme",
39
+ "homepage": "https://github.com/kitepon/aiterm-mcp#readme",
40
40
  "type": "module",
41
41
  "bin": {
42
42
  "aiterm-mcp": "dist/index.js",