aiterm-mcp 0.12.2 → 0.15.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.ja.md +31 -23
- package/README.md +63 -22
- package/dist/aiterm-wait-cli.js +94 -0
- package/dist/claude-stop-hook.js +237 -0
- package/dist/core.js +1027 -149
- package/dist/index.js +133 -28
- package/dist/runtime-error-store.js +15 -7
- package/package.json +4 -3
package/README.ja.md
CHANGED
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
<p align="center">
|
|
2
|
-
<img src=".github/og.svg" alt="aiterm-mcp — AI が握る 1 本の永続 MCP 端末。その中へ他のコーディングエージェント(Codex/Grok/Composer)を起動する(tmux ベースの stdio MCP サーバ)" width="100%">
|
|
2
|
+
<img src=".github/og.svg" alt="aiterm-mcp — AI が握る 1 本の永続 MCP 端末。その中へ他のコーディングエージェント(Claude/Codex/Grok/Composer)を起動する(tmux ベースの stdio MCP サーバ)" width="100%">
|
|
3
3
|
</p>
|
|
4
4
|
|
|
5
5
|
# aiterm-mcp
|
|
@@ -12,7 +12,7 @@
|
|
|
12
12
|
|
|
13
13
|
> *(English: [README.md](README.md))*
|
|
14
14
|
|
|
15
|
-
> **あなたの AI に、ほかの AI を操らせる。**
|
|
15
|
+
> **あなたの AI に、ほかの AI を操らせる。** 任意の MCP クライアントから 1 回の呼び出しで、コーディングエージェント(Claude・Codex・Grok・Composer)を永続端末の中に起動し、操作用のセッションを手渡す。何をしているかをトークン削減して読み、次の指示を送る。
|
|
16
16
|
>
|
|
17
17
|
> **これは何か:** AI が握る 1 本の永続 MCP 端末——その中に他のコーディングエージェントも起動できる。`ssh`・`docker exec`・REPL・別エージェントの TUI は、すべてその 1 本の端末の中へ「送るだけのテキスト」として入れ子になる。仕組みはあえて素朴——MCP クライアントが相手エージェントの端末を 1 ターンずつ操作するだけ。隠れたプロトコルも・共有メモリも・自律的な交渉も無い。
|
|
18
18
|
>
|
|
@@ -22,7 +22,7 @@
|
|
|
22
22
|
|
|
23
23
|
**言葉でなく実測で:** このリポジトリ自身の 203 テストで、`pty_read` はコンテキストに載るトークンを生ログの **約 7.1 分の 1** に減らす。しかも pass/fail の判定は畳んでも残る。→ [組み込みシェルツールとの使い分け](#組み込みシェルツールとの使い分け)
|
|
24
24
|
|
|
25
|
-
|
|
25
|
+
12 ツール: 6 つの **PTY ツール**(`pty_open` / `pty_send` / `pty_read` / `pty_key` / `pty_close` / `pty_list`)で 1 本の永続端末を開き・操作し・読む。加えて 4 つの **エージェント起動ツール**(`claude_agent` / `codex_agent` / `grok_agent` / `composer_agent`)が別のコーディングエージェントの TUI を新しい端末の中に起動し、`claude_turn`がdurable caller向けの構造化issue/recoveryを、`diagnostics`が安全なfactory readinessを返す。バックエンドは **tmux** なので、MCP サーバや AI クライアントが再起動してもセッションは生き残る。
|
|
26
26
|
|
|
27
27
|
**v0.12.2 は release candidate(公開待ち)です。** factory diagnostics と local
|
|
28
28
|
runtime-error store は canonical dotagents config の `collection.enabled: true` が明示された
|
|
@@ -52,7 +52,7 @@ pty_read(id, { wait: true }) → 削減済みの出力を読む(完了
|
|
|
52
52
|
|
|
53
53
|
### 2. その端末の中に他のコーディングエージェントを起動する — オーケストレーションの旗艦
|
|
54
54
|
|
|
55
|
-
同じ primitive が別エージェントの TUI を宿す。
|
|
55
|
+
同じ primitive が別エージェントの TUI を宿す。4 つの起動ツールが、Claude/Codex/Grok/Composer の対話 TUI を新しい永続端末の中に起動し、`session_id` を返す。既存の人間向けtextに加えて`aiterm.agent-launch-result.v1` structured receiptも返すため、durable callerは表示文字列を解析せずsession handleを取得できる。以後は同じ `pty_read` / `pty_send` で継続操作する。`agent_done:true` なら Stop hook でターン完了を待てる。managed Claudeでは全turnに`wait:"agent_done"`を必須とする。durable machine callerは`claude_turn({ action:"issue"|"recover", session_id, operation_id, ... })`を使い、人間向けerror文字列を解析せず`accepted`/`pending`/`completed`/`unknown`を判定できる。recoveryは再送せず、検証済み完了だけがexact `raw_output`を持つ。通常の`pty_send`/`pty_read`は対話callerと人間向けに維持する。`C-c`後もmarkerを保持し、Stopが来なければsessionをcloseする。`claude_agent` と `codex_agent` は、TUI ready 後に初回 `prompt` を送り `wait:"agent_done"` で待つ入口も公開する。Codex/Grok/Composer の live smoke は成功済み。Claude実モデルの初回/follow-up smokeは明示承認待ちであり、まだ成功扱いしない。
|
|
56
56
|
|
|
57
57
|
```text
|
|
58
58
|
codex_agent({ session_name: "codex1", cwd: "/repo", agent_done: true,
|
|
@@ -67,13 +67,14 @@ pty_send("codex1", "also fix the imports it broke", { wait: "agent_done" })
|
|
|
67
67
|
|
|
68
68
|
| ツール | 起動するもの | 主な引数 |
|
|
69
69
|
| --- | --- | --- |
|
|
70
|
-
| `
|
|
71
|
-
| `
|
|
72
|
-
| `
|
|
70
|
+
| `claude_agent` | Claude Code CLI(Anthropic) | `prompt?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`), `cwd?`, `session_name?`, `agent_done?`, `wait?`, `timeout?`, `screen?`, `lines?` |
|
|
71
|
+
| `codex_agent` | Codex CLI(OpenAI・端末設定/CLI既定、`model?`で上書き) | `prompt?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`/`ultra`), `cwd?`, `session_name?`, `agent_done?`, `wait?`, `timeout?`, `screen?`, `lines?` |
|
|
72
|
+
| `grok_agent` | Grok Build(xAI、既定`grok-4.5`、`model?`で上書き) | `prompt?`, `model?`, `reasoning_effort?`は非対応(指定時は明示エラー), `cwd?`, `session_name?`, `agent_done?` |
|
|
73
|
+
| `composer_agent` | Grok Build(xAI、既定`grok-composer-2.5-fast`、`model?`で上書き) | `prompt?`, `model?`, `reasoning_effort?`は非対応(指定時は明示エラー), `cwd?`, `session_name?`, `agent_done?` |
|
|
73
74
|
|
|
74
|
-
各ベンダーの CLI が導入・認証済みであること(`codex_agent` は `codex`、Grok
|
|
75
|
+
各ベンダーの CLI が導入・認証済みであること(`claude_agent` は `claude`、`codex_agent` は `codex`、Grok 系は `grok`)。バイナリは `CLAUDE_BIN` / `CODEX_BIN` / `GROK_BIN`、各既定path、`PATH` の順で解決する。CLI不在・不正なmodel/effort・実在しない`cwd`はsession作成前に失敗し、残骸を残さない。Claudeの`agent_done:true`は通常settingsを継承しないlaunch専用settingsとStop hookを使い、本文なしeventとowner-only bounded resultを分離する。`pty_read({ agent_transcript:true })`はdigestとbyte数を検証したresultだけを返し、Claude private transcriptを読まない。timeoutは成功扱いせず同じsessionを残し、後着resultをprompt再送なしで回収できる。managed Claudeでは通常`pty_send(wait:"none")`とC-c以外の`pty_key`を拒否する。自由なkey操作が必要なら`agent_done:false`で起動し、managed routeでは中断に`C-c`、解除に`pty_close`を使う。
|
|
75
76
|
|
|
76
|
-
|
|
77
|
+
エージェント間の隠れたプロトコルは無い。起動したClaude/Codex/Grok/Composerは利用者がattachできるもう1本の永続sessionであり、MCPクライアントが通常のPTY操作で駆動する。
|
|
77
78
|
|
|
78
79
|
## デモ
|
|
79
80
|
|
|
@@ -134,7 +135,7 @@ claude mcp add --scope user --transport stdio aiterm -- npx -y aiterm-mcp
|
|
|
134
135
|
Claude Code を再起動して、接続を確認:
|
|
135
136
|
|
|
136
137
|
```bash
|
|
137
|
-
/mcp # aiterm が connected・
|
|
138
|
+
/mcp # aiterm が connected・12 ツール公開、と出る
|
|
138
139
|
```
|
|
139
140
|
|
|
140
141
|
最初のセッション——4 回の呼び出しで、1 個の永続端末:
|
|
@@ -146,6 +147,9 @@ pty_read("t1", { wait: true }) → "hello" (トークン削減・完了
|
|
|
146
147
|
pty_close("t1") → 端末を解放
|
|
147
148
|
```
|
|
148
149
|
|
|
150
|
+
`pty_close` は冪等で、`closed` / `already_closed` のstructured receiptを返す。
|
|
151
|
+
MCP応答を失ったdurable callerも同じ`session_id`への再試行だけでclose結果を確定できる。
|
|
152
|
+
|
|
149
153
|
これだけ。`t1` の端末は本物で永続——`ssh`・`docker exec`・REPL・起動したエージェントの TUI は、そこに住む「もの」に過ぎない。代わりにワーカーのエージェントを起動するのも 1 コール: `codex_agent()` が返す `session_id` を、同じ `pty_read` / `pty_send` で操作する。
|
|
150
154
|
|
|
151
155
|
**グローバル導入や別クライアントが良い場合は:**
|
|
@@ -172,11 +176,11 @@ MCP クライアントが aiterm を stdio 越しにプログラムから駆動
|
|
|
172
176
|
|
|
173
177
|
```mermaid
|
|
174
178
|
flowchart LR
|
|
175
|
-
AI["AI / MCP client<br/>(the orchestrator)"] -->|"pty_send · codex_agent<br/>grok_agent · composer_agent · diagnostics"| S["aiterm-mcp<br/>stdio MCP ·
|
|
179
|
+
AI["AI / MCP client<br/>(the orchestrator)"] -->|"pty_send · claude_agent · claude_turn · codex_agent<br/>grok_agent · composer_agent · diagnostics"| S["aiterm-mcp<br/>stdio MCP · 12 tools"]
|
|
176
180
|
S -->|"pty_read<br/>token-reduced"| AI
|
|
177
181
|
S -->|"tmux send-keys<br/>capture-pane"| P["persistent PTYs<br/>tmux · survive restarts"]
|
|
178
182
|
P -->|"ssh · docker · repl"| R["nested<br/>remote · container · REPL"]
|
|
179
|
-
P -->|"launches a fresh PTY per agent"| A["another coding-agent TUI<br/>Codex · Grok · Composer"]
|
|
183
|
+
P -->|"launches a fresh PTY per agent"| A["another coding-agent TUI<br/>Claude · Codex · Grok · Composer"]
|
|
180
184
|
```
|
|
181
185
|
|
|
182
186
|
primitive は「PTY を 1 個握る」ことだけ。それ以外——SSH・コンテナ・REPL・起動したエージェント TUI——は、永続端末の中で動く「対話的な何か」に過ぎず、同じ `pty_send` / `pty_read` で操作する。各起動ツールは自分専用の新しい PTY を開く。PTY は tmux 上にあるので、MCP サーバや AI クライアントが再起動してもセッションは生き残る。
|
|
@@ -248,11 +252,12 @@ aiterm は同じ核心の洞察——端末を出会いの場にする——を
|
|
|
248
252
|
| ツール | 役割 | 主な引数 |
|
|
249
253
|
| --- | --- | --- |
|
|
250
254
|
| `pty_open` | 端末を 1 個握り `session_id` を返す | `name?`, `shell="bash"` |
|
|
251
|
-
| `pty_send` | テキスト(コマンド)を送る | `session_id`, `text`, `enter=true`, `wait`, `timeout`, `screen`, `lines`, `mark`, `force`, `rtk`, `raw` |
|
|
252
|
-
| `pty_read` | 出力を削減して読む(既定は増分) | `session_id`, `wait`, `until`, `until_regex`, `timeout`, `screen`, `full`, `lines`, `line_range`, `raw`, `rtk`, `agent_transcript` |
|
|
255
|
+
| `pty_send` | テキスト(コマンド)を送る | `session_id`, `text`, `enter=true`, `wait`, `timeout`, `screen`, `lines`, `operation_id`, `mark`, `force`, `rtk`, `raw` |
|
|
256
|
+
| `pty_read` | 出力を削減して読む(既定は増分) | `session_id`, `wait`, `until`, `until_regex`, `timeout`, `screen`, `full`, `lines`, `line_range`, `raw`, `rtk`, `agent_transcript`, `operation_id` |
|
|
253
257
|
| `pty_key` | 制御キーを送る | `session_id`, `key`(`C-c`/`Enter`/`Up`…) |
|
|
254
|
-
| `pty_close` |
|
|
258
|
+
| `pty_close` | 冪等に閉じ、`closed` / `already_closed`を返す | `session_id` |
|
|
255
259
|
| `pty_list` | セッション一覧 | (なし) |
|
|
260
|
+
| `claude_turn` | 相関済みmanaged Claude operationを送信または回収 | `action`, `session_id`, `operation_id`, `text?`, `timeout?` |
|
|
256
261
|
| `diagnostics` | 機械可読 JSON による read-only factory readiness | (なし) |
|
|
257
262
|
|
|
258
263
|
`diagnostics` は PTY やエージェントを起動しない。パッケージ版、MCP 呼出 readiness、read-only な PTY 一覧要約、bounded runtime-error-store status、任意 vendor launcher の可用性だけを返す。path・環境値・認証情報・コマンド本文・PTY 出力・raw log は意図的に返さない。通常未設定の任意依存は `not_applicable`、安全に確定できない状態は `unverified` と表す。
|
|
@@ -269,13 +274,14 @@ consumer は `aiterm-runtime-errors snapshot` を読み、durable ingestion 後
|
|
|
269
274
|
|
|
270
275
|
| ツール | 起動するもの | 主な引数 |
|
|
271
276
|
| --- | --- | --- |
|
|
272
|
-
| `
|
|
273
|
-
| `
|
|
274
|
-
| `
|
|
277
|
+
| `claude_agent` | Claude Code CLI(Anthropic) | `prompt?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`), `cwd?`, `session_name?`, `agent_done?`, `wait?`, `timeout?`, `screen?`, `lines?` |
|
|
278
|
+
| `codex_agent` | Codex CLI(OpenAI・端末設定/CLI既定、`model?`で上書き) | `prompt?`, `model?`, `reasoning_effort?`(`low`/`medium`/`high`/`xhigh`/`max`/`ultra`), `cwd?`, `session_name?`, `agent_done?`, `wait?`, `timeout?`, `screen?`, `lines?` |
|
|
279
|
+
| `grok_agent` | Grok Build(xAI、既定`grok-4.5`、`model?`で上書き) | `prompt?`, `model?`, `reasoning_effort?`は非対応(指定時は明示エラー), `cwd?`, `session_name?`, `agent_done?` |
|
|
280
|
+
| `composer_agent` | Grok Build(xAI、既定`grok-composer-2.5-fast`、`model?`で上書き) | `prompt?`, `model?`, `reasoning_effort?`は非対応(指定時は明示エラー), `cwd?`, `session_name?`, `agent_done?` |
|
|
275
281
|
|
|
276
|
-
|
|
282
|
+
対応するCLI(`claude` / `codex` / `grok`)の導入・認証が必要。解決順は`CLAUDE_BIN` / `CODEX_BIN` / `GROK_BIN`、既定path、`PATH`。前提違反はsession作成前に明示失敗する。Claude/Codexは初回prompt完了待ちを公開し、4 launcherすべてが対応環境で同じfollow-up `pty_send(wait:"agent_done")`契約を使う。Claudeはisolated managed settingsとhook-captured resultを使い、private transcriptへ依存しない。既存3 vendorのlive smokeはgreen、Claude実モデルsmokeは承認待ちでありfixture成功と混同しない。
|
|
277
283
|
|
|
278
|
-
エージェントの回答が `wait:"agent_done"` の画面 tail
|
|
284
|
+
エージェントの回答が `wait:"agent_done"` の画面 tailより長ければ、対話callerは`pty_read({ agent_transcript:true })`で再promptなしに全文回収する。Claudeはmanaged Stop hookがowner-only resultへ保存した本文をdigest/byte数で検証して返し、private transcriptを読まない。durable machine callerは`claude_turn`を使う。`issue`は一度だけ送信し、`recover`は決して再送せず、`pending`を破損やidentity不一致と区別する。検証済みの`completed`だけがexact `raw_output`を持ち、`unknown`は未dispatchと帰属不能を区別する。不一致・破損は成功statusへ丸めずtool errorのままにする。IDなしの対話turnも匿名markerで直列化するため、現在Stop待ちの間に古い回答を返さない。CodexはStop hookの`turn_id`で構造化transcriptへjoinし、Grok/Composerは最後の実user行より後ろのassistant行を採る。不在・非agent・抽出不能は明示エラー。
|
|
279
285
|
|
|
280
286
|
### 完了検出(5 層)
|
|
281
287
|
|
|
@@ -291,9 +297,11 @@ consumer は `aiterm-runtime-errors snapshot` を読み、durable ingestion 後
|
|
|
291
297
|
|
|
292
298
|
`pty_send` は送信前に破壊的コマンド(`rm -rf /`, `mkfs`, `dd of=/dev/…`, `DROP TABLE` 等)を遮断し(`force: true` で越える)、ESC・ブラケットペースト終端などをサニタイズする。`pty_read` は既定で制御文字を無害化して返す(`raw: true` はバイトをそのまま返す)。これは**サンドボックスではなく tripwire**([既知の制約](#既知の制約バグではなく仕様)参照)。
|
|
293
299
|
|
|
300
|
+
1回の `pty_send` が受理する本文はUTF-8で最大64KiB。同一sessionへの送信はaiterm processをまたいで直列化し、chunk同士の混線を防ぐ。macOSでは長いPTY入力の欠落を避けるためUTF-8境界を壊さない256-byte単位でtmux pasteし、Linux/WSLでは上限内を1回でpasteする。途中chunkが失敗した場合は部分送信済みであることを明示し、自動でEnterを押さない。送信processの異常終了でlockが残った場合は送信前にfail-closedする。そのsessionを `pty_close` して作り直すか、全sessionを破棄できる場合だけ `pty_kill_all` で安全に掃除する。
|
|
301
|
+
|
|
294
302
|
## 人が覗く
|
|
295
303
|
|
|
296
|
-
セッションは共有 tmux ソケット上にある。`pty_open`(および各エージェント起動ツール)の戻り値に表示される `tmux -S … attach -t <id>` で人間が同じ端末に入って介入できる(抜けるのは `Ctrl-b d`)——起動した Codex/Grok/Composer
|
|
304
|
+
セッションは共有 tmux ソケット上にある。`pty_open`(および各エージェント起動ツール)の戻り値に表示される `tmux -S … attach -t <id>` で人間が同じ端末に入って介入できる(抜けるのは `Ctrl-b d`)——起動した Claude/Codex/Grok/Composer のセッションを見たり、途中でキーボードを引き取ったりもできる。
|
|
297
305
|
|
|
298
306
|
## 要件
|
|
299
307
|
|
|
@@ -301,7 +309,7 @@ consumer は `aiterm-runtime-errors snapshot` を読み、durable ingestion 後
|
|
|
301
309
|
- **tmux**(実行時の前提。`tmux -V` で確認。未導入なら `apt install tmux` / `brew install tmux`)
|
|
302
310
|
- **macOS / Linux / WSL2** は tmux を直接使う。macOS は同梱されないので `brew install tmux` で導入する。MCP クライアントがターミナルでなく **GUI から起動**された場合、Homebrew の bin(Apple Silicon: `/opt/homebrew/bin`、Intel: `/usr/local/bin`)が `PATH` に入らないことがある。その場合 aiterm が自動で探索するか、**`AITERM_TMUX=/path/to/tmux`** で明示指定する。
|
|
303
311
|
- **Windows ネイティブ**には tmux が無いため、aiterm は裏で **WSL の中の tmux** を透過的に使う。[WSL](https://learn.microsoft.com/ja-jp/windows/wsl/) を導入・初期化し、**WSL のディストリ内に tmux を入れる**こと(`sudo apt install tmux`)。`wsl tmux -V` で確認できる。セッション・ソケット・人の `attach` はすべて WSL 側にあり、AI は Windows 側のコマンドから操作するだけ。(Windows のツールは SSH と同じく入れ子で握る: `pty_send "powershell.exe …"` で PowerShell に入る。)
|
|
304
|
-
- **エージェント起動ツール**を使う場合: 対応するベンダー CLI が導入・認証済みであること——`codex_agent` は `codex`、`grok_agent` / `composer_agent` は `grok`。(PTY ツールだけ使うなら不要。)
|
|
312
|
+
- **エージェント起動ツール**を使う場合: 対応するベンダー CLI が導入・認証済みであること——`claude_agent` は `claude`、`codex_agent` は `codex`、`grok_agent` / `composer_agent` は `grok`。(PTY ツールだけ使うなら不要。)
|
|
305
313
|
- 任意: [`rtk`](https://github.com/rtk-ai/rtk) バイナリ(`pty_send` の `rtk: true` 委譲で使う。無くても動く)
|
|
306
314
|
|
|
307
315
|
## 既知の制約(バグではなく仕様)
|
|
@@ -309,7 +317,7 @@ consumer は `aiterm-runtime-errors snapshot` を読み、durable ingestion 後
|
|
|
309
317
|
- **ネスト中(ssh / docker / REPL / 起動したエージェント TUI)は quiescence が原理的に効かない。** 前面コマンドがシェル集合(bash/sh/zsh/fish/dash)の外になるため。ネスト中で `until` も `mark` も無いときは、待っても完了を確定できる信号が無いので、`pty_read({ wait: true })` はフル `timeout` を空費せず出力静止時点で `is_complete=False via nested` と早期に返し、`until`(既定リテラル部分一致・`until_regex: true` で正規表現)か `mark: true`(終了コード付き sentinel・自動検出)の指定を促す。全画面のエージェント TUI なら、出力が落ち着いた時点で `{ screen: true }` を読む。
|
|
310
318
|
- **`is_complete=False` は失敗ではない。** 「timeout 内に完了を観測できなかった」という意味。長時間コマンドでは `timeout` を伸ばすか `until`/`mark` を使う。
|
|
311
319
|
- **破壊ゲートはサンドボックスではなく tripwire。** よくある破壊形だけを弾く。相対パスの `rm`、`$VAR` 展開後に危険化するもの、ssh 先で実行されるコマンドは捕捉しない——起動したコーディングエージェントが自分のセッション内で何をするかも取り締まらない。
|
|
312
|
-
- **エージェント起動ツールはベンダー TUI を起動するだけで、包んだり代理したりしない。** aiterm は前提を検証して CLI を永続 PTY で起動する——モデル・認証・挙動はベンダー CLI
|
|
320
|
+
- **エージェント起動ツールはベンダー TUI を起動するだけで、包んだり代理したりしない。** aiterm は前提を検証して CLI を永続 PTY で起動する——モデル・認証・挙動はベンダー CLI のもの。エージェント間の隠れたプロトコルはなく、「会話」とはMCPクライアントがClaude/Codex/Grok/Composer TUIへ入力を送り出力を読むことだ。
|
|
313
321
|
- **`pty_send({ rtk: true })` は単行コマンドのみ+外部 `rtk` バイナリが必要**(無ければ素通し)。一方 `pty_read({ rtk: true })` の reducer は自前実装で rtk 非依存。
|
|
314
322
|
- **`pytest` reducer は件数・罫線・`FAILURES` ブロック整形が rtk 0.42.0 と byte 一致**(回帰テストで固定)。ただし `-ra`/`-rf` 時の `FAILED` 要約行の理由は**全文を保持する**(rtk 0.42.0 は最初の `" - "` 区切りで切るが、本実装は可読性優先で情報を残すため、この行は意図的に rtk と完全一致させない)。rtk が大出力時に付ける `[full output: …]`(tee ポインタ)行は read 側では再現しない。
|
|
315
323
|
- **tmux は `-f /dev/null` 起動**なので `~/.tmux.conf` を読まない(環境差を排除するため)。
|
package/README.md
CHANGED
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
<p align="center">
|
|
2
|
-
<img src=".github/og.svg" alt="aiterm-mcp — one persistent MCP terminal your AI drives, and launches other coding agents (Codex/Grok/Composer) into (tmux-backed stdio MCP server)" width="100%">
|
|
2
|
+
<img src=".github/og.svg" alt="aiterm-mcp — one persistent MCP terminal your AI drives, and launches other coding agents (Claude/Codex/Grok/Composer) into (tmux-backed stdio MCP server)" width="100%">
|
|
3
3
|
</p>
|
|
4
4
|
|
|
5
5
|
# aiterm-mcp
|
|
@@ -12,7 +12,7 @@
|
|
|
12
12
|
|
|
13
13
|
> *(日本語: [README.ja.md](README.ja.md))*
|
|
14
14
|
|
|
15
|
-
> **Let your AI orchestrate other AIs.** From
|
|
15
|
+
> **Let your AI orchestrate other AIs.** From any MCP client, one call spawns a coding agent (Claude, Codex, Grok, or Composer) inside a persistent terminal and hands you a session to drive: read what it's doing token-reduced, send it the next instruction.
|
|
16
16
|
>
|
|
17
17
|
> **What it is:** one persistent MCP terminal your AI drives — and can launch other coding agents into. `ssh`, `docker exec`, a REPL, or another agent's TUI all nest inside that one terminal as just text you send in. The mechanism is deliberately plain — your MCP client drives the other agent's terminal turn by turn: no hidden protocol, no shared memory, no autonomous negotiation.
|
|
18
18
|
>
|
|
@@ -22,13 +22,16 @@
|
|
|
22
22
|
|
|
23
23
|
**Measured, not claimed:** on this repo's own 203-test suite, a `pty_read` puts **~7.1× fewer tokens** in your context than the raw log — and the pass/fail verdict survives the fold. → [When to reach for it vs. the built-in shell](#when-to-reach-for-it-vs-the-built-in-shell)
|
|
24
24
|
|
|
25
|
-
|
|
25
|
+
Twelve tools: six **PTY tools** — `pty_open` / `pty_send` / `pty_read` / `pty_key` / `pty_close` / `pty_list` — to open, drive, and read one persistent terminal, four **agent launchers** — `claude_agent` / `codex_agent` / `grok_agent` / `composer_agent` — that each start another coding agent's TUI inside a fresh one, `claude_turn` for durable structured issue/recovery, and `diagnostics` for safe factory readiness. The backend is **tmux**, so sessions survive even if the MCP server or the AI client restarts.
|
|
26
26
|
|
|
27
|
-
**v0.
|
|
27
|
+
**v0.15.0 was published on 2026-07-18.** It brings the interactive agent
|
|
28
|
+
launchers, durable `claude_turn` issue/recovery, machine-readable launch and
|
|
29
|
+
idempotent close receipts, the hardened TUI readiness gate, and the new
|
|
30
|
+
`aiterm-wait` completion-push binary to npm. Factory diagnostics and the local
|
|
28
31
|
runtime-error store collect only when canonical dotagents config explicitly sets
|
|
29
32
|
`collection.enabled: true`; collection is off by default and performs no network
|
|
30
|
-
I/O.
|
|
31
|
-
|
|
33
|
+
I/O. It ships via tag-triggered CI with npm provenance (OIDC Trusted Publishing);
|
|
34
|
+
the GitHub Release re-registers the Official MCP Registry entry.
|
|
32
35
|
|
|
33
36
|
**Status:** actively maintained · the newcomer here, betting on a different shape (see [vs. the alternatives](#vs-the-alternatives)) · runs on Linux · WSL2 · macOS · native Windows for the core PTY tools (`agent_done` is POSIX/WSL/macOS only for now) · MIT · see the [CHANGELOG](CHANGELOG.md).
|
|
34
37
|
|
|
@@ -36,6 +39,12 @@ are pending.
|
|
|
36
39
|
|
|
37
40
|
A lot of 2026's agent tooling is converging on orchestration: a lead model delegating a mechanical refactor to Codex, running Composer on a bulk edit while it reviews the diff, fanning one task across several agents to spare its own context window. All of those agents already live in a terminal. aiterm makes that terminal a first-class, MCP-native tool — so the model doing the orchestrating can **spawn and steer the others without a human wiring up panes.**
|
|
38
41
|
|
|
42
|
+
## Built with Codex and GPT-5.6 for OpenAI Build Week 2026
|
|
43
|
+
|
|
44
|
+
aiterm predates Build Week, so the event work is kept visible in dated commits. During the submission window (July 14–16, 2026), I extended it with safe serialized delivery for long PTY input, correlated operation IDs and bounded result recovery, machine-readable launch and idempotent close receipts, and a hardened readiness gate that prevents prompts from disappearing during TUI startup redraws. The public comparison from the pre-event release is [`v0.12.2...main`](https://github.com/kitepon-rgb/aiterm-mcp/compare/v0.12.2...main).
|
|
45
|
+
|
|
46
|
+
I used **Codex with GPT-5.6** as an engineering collaborator: it inspected the implementation, challenged the API and recovery contracts, generated focused regression cases, and helped verify race, security, timeout, and malformed-event paths. I reviewed the diffs and test evidence and retained the final product and architecture decisions. The result is a 262-test regression suite covering normal operation as well as failure and recovery behavior.
|
|
47
|
+
|
|
39
48
|
## Two ways to use it
|
|
40
49
|
|
|
41
50
|
### 1. Drive SSH, containers, and REPLs in one persistent terminal — the primitive
|
|
@@ -53,7 +62,7 @@ pty_read(id, { wait: true }) → read the token-reduced output, completion
|
|
|
53
62
|
|
|
54
63
|
### 2. Launch other coding agents into that terminal — the orchestration flagship
|
|
55
64
|
|
|
56
|
-
The same primitive hosts another agent's TUI.
|
|
65
|
+
The same primitive hosts another agent's TUI. Four launchers each start one vendor's interactive coding-agent TUI inside a fresh persistent terminal and return a `session_id`. Their existing human-readable text is accompanied by an `aiterm.agent-launch-result.v1` structured receipt, so durable callers never parse display text for the session handle. From there you drive it with the same `pty_read` / `pty_send` you'd use on any shell: read its output token-reduced, send it the next step. (The TUIs are full-screen apps, so `pty_read({ screen: true })` gives you the rendered view.) Agent launchers can also opt into hook-backed turn completion with `agent_done: true`, letting `pty_send({ wait: "agent_done" })` return after the agent turn ends. A managed Claude session requires that wait mode for every turn. Durable machine callers use `claude_turn({ action: "issue" | "recover", session_id, operation_id, ... })`: it returns fixed `accepted` / `pending` / `completed` / `unknown` states without parsing human-facing errors, never resends during recovery, and includes exact `raw_output` only for a verified completion. The same operation ID is carried through the dispatch receipt, active marker, Stop event, and result. The ordinary `pty_send` / `pty_read` surface remains available for interactive callers and humans. `C-c` keeps the marker for a delayed Stop; if no Stop arrives, close the session. `claude_agent` and `codex_agent` can additionally wait for their initial `prompt`: pass `prompt`, `agent_done: true`, and `wait: "agent_done"`; aiterm starts the TUI first, waits until its input area is ready, submits the prompt, then returns after that first turn's Stop hook. Codex/Grok/Composer have passing live smokes for follow-up completion, and Codex initial-prompt wait is also live-smoked. Claude's real-model initial/follow-up smoke remains an explicit approval gate. Grok/Composer initial-prompt wait is intentionally not exposed until the post-OAuth smoke passes. This needs the vendor's own CLI installed and authenticated — see [Requirements](#requirements).
|
|
57
66
|
|
|
58
67
|
```text
|
|
59
68
|
codex_agent({ session_name: "codex1", cwd: "/repo", agent_done: true,
|
|
@@ -68,13 +77,14 @@ One call per model, so the tool name itself tells you which model you get:
|
|
|
68
77
|
|
|
69
78
|
| Tool | Launches | Key args |
|
|
70
79
|
| --- | --- | --- |
|
|
80
|
+
| `claude_agent` | Claude Code CLI (Anthropic) | `prompt?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`), `cwd?`, `session_name?`, `agent_done?`, `launch_operation_id?`, `wait?`, `timeout?`, `screen?`, `lines?` |
|
|
71
81
|
| `codex_agent` | Codex CLI (OpenAI; terminal config/CLI default unless overridden) | `prompt?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`/`ultra`; ultra enables proactive automatic delegation), `cwd?`, `session_name?`, `agent_done?`, `wait?`, `timeout?`, `screen?`, `lines?` |
|
|
72
82
|
| `grok_agent` | Grok Build, model `grok-4.5` by default (`model?` overrides) (xAI) | `prompt?`, `model?`, `reasoning_effort?` unsupported (an explicit value is an error; Grok CLI `--effort` is headless-only), `cwd?`, `session_name?`, `agent_done?` |
|
|
73
83
|
| `composer_agent` | Grok Build, model `grok-composer-2.5-fast` by default (`model?` overrides) (xAI) | `prompt?`, `model?`, `reasoning_effort?` unsupported (an explicit value is an error), `cwd?`, `session_name?`, `agent_done?` |
|
|
74
84
|
|
|
75
|
-
The vendor CLI must be installed and authenticated (`codex` for `codex_agent`; `grok` for both Grok tools). aiterm resolves the binary via `CODEX_BIN` / `GROK_BIN`, then `~/.local/bin/codex` / `~/.grok/bin/grok`, then `PATH`. Prerequisites are checked **before** a session exists: empty `model` values and
|
|
85
|
+
The vendor CLI must be installed and authenticated (`claude` for `claude_agent`; `codex` for `codex_agent`; `grok` for both Grok tools). aiterm resolves the binary via `CLAUDE_BIN` / `CODEX_BIN` / `GROK_BIN`, then `~/.local/bin/claude` / `~/.local/bin/codex` / `~/.grok/bin/grok`, then `PATH`. Prerequisites are checked **before** a session exists: empty `model` values and unsupported effort values are rejected up front; a missing CLI binary or a nonexistent `cwd` fails for all four. A rejected launch leaves **zero leftover session** behind. Claude and Codex launchers forward `model` and `reasoning_effort` through their vendor CLI's public flags; Grok/Composer reject `reasoning_effort` because it is headless-only. Pass an absolute path for `cwd` — `~` is not expanded. Durable callers can make a promptless managed Claude launch exactly replayable by passing an explicit `session_name`, `agent_done:true`, and a `launch_operation_id` formatted as `sha256:<64 lowercase hex>`. Repeating the identical launch returns the same structured session receipt without starting the CLI twice; a different correlation ID or launch argument for that session fails explicitly. With `agent_done:true`, Claude uses launch-local managed settings containing only aiterm's Stop hook: normal user/project/local hooks are not inherited, the hook event contains no answer body, and the bounded owner-only result is returned by `pty_read({ agent_transcript:true })` without reading Claude's private transcript. A timeout remains `is_complete=False`; the same session and a late result remain recoverable without re-sending the prompt. Ordinary `pty_send(wait:"none")` and non-interrupt keys are rejected on this managed Claude route; use `wait:"agent_done"`, `pty_key("C-c")`, and `pty_close`. Start with `agent_done:false` when unconstrained manual key-by-key driving is desired. Codex uses a managed `CODEX_HOME`; Grok/Composer isolate their managed homes and pass validated OAuth state through `GROK_AUTH_PATH`. Before the first unbound completion-waiting send, aiterm waits for the vendor TUI's input prompt and fails before sending if it is not ready. `agent_done` requires POSIX filesystem semantics, so it is supported on Linux, WSL2, and macOS; native Windows can still use launchers without `agent_done`.
|
|
76
86
|
|
|
77
|
-
There is
|
|
87
|
+
There is no hidden protocol between agents: a launched Claude, Codex, Grok, or Composer is another user-visible persistent terminal session. The MCP client drives that TUI with ordinary PTY operations, and a human can attach to watch or take over.
|
|
78
88
|
|
|
79
89
|
## Demo
|
|
80
90
|
|
|
@@ -135,7 +145,7 @@ claude mcp add --scope user --transport stdio aiterm -- npx -y aiterm-mcp
|
|
|
135
145
|
Restart Claude Code, then verify the connection:
|
|
136
146
|
|
|
137
147
|
```bash
|
|
138
|
-
/mcp # aiterm should show as connected, exposing
|
|
148
|
+
/mcp # aiterm should show as connected, exposing 12 tools
|
|
139
149
|
```
|
|
140
150
|
|
|
141
151
|
Your first session — four calls, one persistent terminal:
|
|
@@ -147,6 +157,9 @@ pty_read("t1", { wait: true }) → "hello" (token-reduced, completion det
|
|
|
147
157
|
pty_close("t1") → terminal released
|
|
148
158
|
```
|
|
149
159
|
|
|
160
|
+
`pty_close` is idempotent and returns a structured `closed` / `already_closed`
|
|
161
|
+
receipt, so durable callers can retry the same `session_id` after losing the MCP response.
|
|
162
|
+
|
|
150
163
|
That's it. The terminal in `t1` is real and persistent — `ssh`, `docker exec`, a REPL, or a launched agent's TUI are just things that live inside it. To launch a worker agent instead, one call does it: `codex_agent()` returns a `session_id` you drive with the same `pty_read` / `pty_send`.
|
|
151
164
|
|
|
152
165
|
**Prefer a global install, or a different client?**
|
|
@@ -173,11 +186,11 @@ The terminal is real and shared, so a human *can* jump in ([A human can watch](#
|
|
|
173
186
|
|
|
174
187
|
```mermaid
|
|
175
188
|
flowchart LR
|
|
176
|
-
AI["AI / MCP client<br/>(the orchestrator)"] -->|"pty_send · codex_agent<br/>grok_agent · composer_agent · diagnostics"| S["aiterm-mcp<br/>stdio MCP ·
|
|
189
|
+
AI["AI / MCP client<br/>(the orchestrator)"] -->|"pty_send · claude_agent · claude_turn · codex_agent<br/>grok_agent · composer_agent · diagnostics"| S["aiterm-mcp<br/>stdio MCP · 12 tools"]
|
|
177
190
|
S -->|"pty_read<br/>token-reduced"| AI
|
|
178
191
|
S -->|"tmux send-keys<br/>capture-pane"| P["persistent PTYs<br/>tmux · survive restarts"]
|
|
179
192
|
P -->|"ssh · docker · repl"| R["nested<br/>remote · container · REPL"]
|
|
180
|
-
P -->|"launches a fresh PTY per agent"| A["another coding-agent TUI<br/>Codex · Grok · Composer"]
|
|
193
|
+
P -->|"launches a fresh PTY per agent"| A["another coding-agent TUI<br/>Claude · Codex · Grok · Composer"]
|
|
181
194
|
```
|
|
182
195
|
|
|
183
196
|
One PTY is the only primitive. Everything else — SSH, containers, REPLs, and the launched agent TUIs — is just something interactive running inside a persistent terminal, driven with the same `pty_send` / `pty_read`. Each launcher opens its own fresh PTY. Because the PTYs live in tmux, sessions outlive the MCP server and the AI client.
|
|
@@ -226,7 +239,7 @@ aiterm sits at the intersection of two families: terminal-driving MCP servers, a
|
|
|
226
239
|
| --- | --- | --- | --- | --- |
|
|
227
240
|
| Persistent session | ✅ tmux, survives restarts | ❌ new shell every call | ⚠️ varies | ✅ tmux |
|
|
228
241
|
| SSH / containers / REPLs | nest with one `pty_send` | reconnect every command | ⚠️ often separate tools | ✅ tmux (human drives) |
|
|
229
|
-
| Launch another agent in one call | ✅ `codex_agent` / `grok_agent` / `composer_agent` | ❌ | ❌ | ⚠️ agents join a human-run tmux via a CLI + skills |
|
|
242
|
+
| Launch another agent in one call | ✅ `claude_agent` / `codex_agent` / `grok_agent` / `composer_agent` | ❌ | ❌ | ⚠️ agents join a human-run tmux via a CLI + skills |
|
|
230
243
|
| Headless (no human at a tmux) | ✅ MCP-driven, programmatic | ✅ | ⚠️ varies | ❌ built around a human in the tmux |
|
|
231
244
|
| MCP-native (any MCP client) | ✅ one `claude mcp add` | ✅ | ✅ (they are MCPs) | ❌ tmux config + CLI + Agent Skills |
|
|
232
245
|
| Token-reduced reads | ✅ per-command reducers | ❌ raw output | ⚠️ rarely | ❌ raw tmux |
|
|
@@ -251,11 +264,12 @@ On top of that sits a productized layer a raw tmux bridge doesn't have: **token-
|
|
|
251
264
|
| Tool | Role | Key args |
|
|
252
265
|
| --- | --- | --- |
|
|
253
266
|
| `pty_open` | Grab one terminal, return a `session_id` | `name?`, `shell="bash"` |
|
|
254
|
-
| `pty_send` | Send text (a command) | `session_id`, `text`, `enter=true`, `wait`, `timeout`, `screen`, `lines`, `mark`, `force`, `rtk`, `raw` |
|
|
255
|
-
| `pty_read` | Read output, token-reduced (incremental by default) | `session_id`, `wait`, `until`, `until_regex`, `timeout`, `screen`, `full`, `lines`, `line_range`, `raw`, `rtk`, `agent_transcript` |
|
|
267
|
+
| `pty_send` | Send text (a command) | `session_id`, `text`, `enter=true`, `wait`, `timeout`, `screen`, `lines`, `operation_id`, `mark`, `force`, `rtk`, `raw` |
|
|
268
|
+
| `pty_read` | Read output, token-reduced (incremental by default) | `session_id`, `wait`, `until`, `until_regex`, `timeout`, `screen`, `full`, `lines`, `line_range`, `raw`, `rtk`, `agent_transcript`, `operation_id` |
|
|
256
269
|
| `pty_key` | Send a control key | `session_id`, `key` (`C-c`/`Enter`/`Up`…) |
|
|
257
|
-
| `pty_close` | Close
|
|
270
|
+
| `pty_close` | Close idempotently; return `closed` / `already_closed` | `session_id` |
|
|
258
271
|
| `pty_list` | List sessions (agent rows carry `agent=<kind>` metadata) | (none) |
|
|
272
|
+
| `claude_turn` | Issue or recover one correlated managed-Claude operation | `action`, `session_id`, `operation_id`, `text?`, `timeout?` |
|
|
259
273
|
| `diagnostics` | Read-only factory readiness as machine-readable JSON | (none) |
|
|
260
274
|
|
|
261
275
|
`diagnostics` never starts a PTY or agent. It reports package version, MCP call readiness, a read-only PTY-list summary, bounded runtime-error-store status, and optional vendor-launcher availability. It deliberately excludes paths, environment values, credentials, command text, PTY output, and raw logs; normal unset optional dependencies are `not_applicable`, while an indeterminate probe is `unverified`.
|
|
@@ -272,17 +286,31 @@ Each launcher starts a specific vendor's interactive coding-agent TUI inside a f
|
|
|
272
286
|
|
|
273
287
|
| Tool | Launches | Key args |
|
|
274
288
|
| --- | --- | --- |
|
|
289
|
+
| `claude_agent` | Claude Code CLI (Anthropic) | `prompt?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`), `cwd?`, `session_name?`, `agent_done?`, `launch_operation_id?`, `wait?`, `timeout?`, `screen?`, `lines?` |
|
|
275
290
|
| `codex_agent` | Codex CLI (OpenAI; terminal config/CLI default unless overridden) | `prompt?`, `model?`, `reasoning_effort?` (`low`/`medium`/`high`/`xhigh`/`max`/`ultra`; ultra enables proactive automatic delegation), `cwd?`, `session_name?`, `agent_done?`, `wait?`, `timeout?`, `screen?`, `lines?` |
|
|
276
291
|
| `grok_agent` | Grok Build, model `grok-4.5` by default (`model?` overrides) (xAI) | `prompt?`, `model?`, `reasoning_effort?` unsupported (an explicit value is an error; Grok CLI `--effort` is headless-only), `cwd?`, `session_name?`, `agent_done?` |
|
|
277
292
|
| `composer_agent` | Grok Build, model `grok-composer-2.5-fast` by default (`model?` overrides) (xAI) | `prompt?`, `model?`, `reasoning_effort?` unsupported (an explicit value is an error), `cwd?`, `session_name?`, `agent_done?` |
|
|
278
293
|
|
|
279
|
-
The vendor CLI must be installed and authenticated (`codex` for `codex_agent`; `grok` for both Grok tools).
|
|
294
|
+
The vendor CLI must be installed and authenticated (`claude` for `claude_agent`; `codex` for `codex_agent`; `grok` for both Grok tools). Binary resolution uses `CLAUDE_BIN` / `CODEX_BIN` / `GROK_BIN`, then each documented default location, then `PATH`. Missing binaries, invalid model/effort values, and nonexistent `cwd` fail before a session is created. Claude and Codex expose initial-prompt completion waiting; all four use the same follow-up `pty_send(wait:"agent_done")` contract where supported. Claude uses isolated managed settings and a hook-captured bounded result rather than private transcript access. Codex/Grok/Composer live smokes are green; Claude real-model smoke remains approval-gated and is not claimed from fixtures. Native Windows can launch agents but `agent_done` is not supported yet.
|
|
280
295
|
|
|
281
|
-
When an agent's answer is longer than the `wait:"agent_done"` screen tail (pane height ≈ 24 lines), recover it in full with `pty_read({ agent_transcript: true })`. It returns the most recently completed turn's final assistant message in plain text
|
|
296
|
+
When an agent's answer is longer than the `wait:"agent_done"` screen tail (pane height ≈ 24 lines), interactive callers can recover it in full with `pty_read({ agent_transcript: true })`. It returns the most recently completed turn's final assistant message in plain text with no re-prompting. Claude reads the bounded owner-only result captured by the managed Stop hook and verifies its digest/byte count; it never reads Claude's private transcript. Durable machine callers should use `claude_turn`: `issue` sends once, `recover` never sends, `pending` is distinct from unsafe or malformed state, and only `completed` carries the exact verified `raw_output`. `unknown` distinguishes `operation_not_found` from a receipt whose result can no longer be attributed. Mismatch and corruption remain tool errors rather than being folded into a successful status. ID-less interactive Claude turns are still serialized by an anonymous marker, so an older answer is not returned while the current Stop is pending. Codex joins its structured transcript on the Stop hook `turn_id`; Grok/Composer take the assistant rows after the last real user row. A missing result/transcript, a non-agent session, or an unextractable message is an explicit error, never a silent empty.
|
|
282
297
|
|
|
283
298
|
### Completion detection (5 layers)
|
|
284
299
|
|
|
285
|
-
`pty_read({ wait: true })` decides "is the command done?" via five layers: process exit / a `mark:true` sentinel (auto-detected — see below) / an `until` match (a literal substring by default; pass `until_regex: true` for a regex) / output is quiescent ∧ the shell is back (quiescence) / timeout. While nested (inside SSH, a container, a REPL, or a launched agent's TUI), the "shell is back" check cannot fire, so pass `until` with the inner prompt — or send with `mark: true` and `pty_read({ wait: true })` auto-detects the completion sentinel (no `until` needed, works nested too) — or, for a full-screen agent TUI, read `{ screen: true }` once its output settles. Sessions launched with `agent_done:true` can instead use `pty_send({ wait:"agent_done" })`, which first waits for the agent TUI input prompt when needed, then waits for the vendor Stop hook and returns the screen after the turn boundary; `codex_agent` launcher `wait:"agent_done"`
|
|
300
|
+
`pty_read({ wait: true })` decides "is the command done?" via five layers: process exit / a `mark:true` sentinel (auto-detected — see below) / an `until` match (a literal substring by default; pass `until_regex: true` for a regex) / output is quiescent ∧ the shell is back (quiescence) / timeout. While nested (inside SSH, a container, a REPL, or a launched agent's TUI), the "shell is back" check cannot fire, so pass `until` with the inner prompt — or send with `mark: true` and `pty_read({ wait: true })` auto-detects the completion sentinel (no `until` needed, works nested too) — or, for a full-screen agent TUI, read `{ screen: true }` once its output settles. Sessions launched with `agent_done:true` can instead use `pty_send({ wait:"agent_done" })`, which first waits for the agent TUI input prompt when needed, then waits for the vendor Stop hook and returns the screen after the turn boundary; `claude_agent` and `codex_agent` launcher `wait:"agent_done"` do the same for the initial `prompt`. Pre-send readiness failures are MCP errors for `pty_send`, while launch-time initial prompt readiness failures return the session with `initial_prompt=not_sent`; timeouts after sending are reported as `is_complete=False via agent_timeout`, not as success. A late Claude completion remains recoverable from the same session with `pty_read({ agent_transcript:true })`, without resending. Normal `pty_read` on an agent session can append auxiliary metadata such as `agent_event_seen=true completion_attribution=none`, but a stale hook event is not promoted to `is_complete=True`. If a complete hook JSONL line is malformed, the timeout suffix includes `malformed_events=N` for diagnosis. If the turn is done but the terminal screen/log does not settle within the flush window, aiterm appends `agent_done_but_screen_unstable`.
|
|
301
|
+
|
|
302
|
+
### Completion push for parent agents (`aiterm-wait`)
|
|
303
|
+
|
|
304
|
+
The `wait:"agent_done"` route blocks the caller's tool call. For fire-and-forget orchestration ("B-style": dispatch now, collect later, never poll), the `aiterm-wait` binary turns the Stop-hook completion event into a process exit that a host harness can treat as a push notification:
|
|
305
|
+
|
|
306
|
+
1. Launch the child agent (`agent_done: true`) and start `aiterm-wait --session <id> [--operation sha256:<64hex>] [--timeout <sec>]` as a **background task of the parent's harness**.
|
|
307
|
+
2. Dispatch the prompt with `claude_turn issue` (`timeout: 0`) or `pty_send(wait:"agent_done", timeout: 0)` — the TUI ready gate and submit sequencing still run; the call returns immediately.
|
|
308
|
+
3. The parent is free. When the child's turn ends, the vendor Stop hook appends the event, `aiterm-wait` exits with a one-line receipt (`aiterm.agent-wait-result.v1`, outcome `done` / `timeout` / `closed`), and a harness that re-invokes its agent on background-task exit (Claude Code does) wakes the parent with zero polling.
|
|
309
|
+
4. The parent collects the result exactly as before: `claude_turn recover` (Claude) or `pty_read` (other vendors). The waiter carries the signal, never the payload.
|
|
310
|
+
|
|
311
|
+
`aiterm-wait` is a **pure reader**: it takes no locks, never writes session state, and never dispatches — so any number of waiters can run beside the MCP server and beside each other (one per launch), and `pty_close`/concurrent sends are unaffected. Start the waiter *before* dispatching (or pass `--operation`, which is start-order-independent) so no completion can slip past the observation boundary.
|
|
312
|
+
|
|
313
|
+
Codex CLI as the *parent* currently has no equivalent "wake on background completion" hook (its `notify` config is human-facing and MCP notifications are not surfaced to the model — see openai/codex#17543 / #18056), so a Codex parent still uses the blocking wait or manual recovery until upstream support lands.
|
|
286
314
|
|
|
287
315
|
### Token reduction
|
|
288
316
|
|
|
@@ -294,9 +322,11 @@ When an agent's answer is longer than the `wait:"agent_done"` screen tail (pane
|
|
|
294
322
|
|
|
295
323
|
Before sending, `pty_send` blocks destructive commands (`rm -rf /`, `mkfs`, `dd of=/dev/…`, `DROP TABLE`, …) — pass `force: true` to override — and sanitizes ESC / bracketed-paste terminators. `pty_read` neutralizes control characters in what it returns by default (`raw: true` returns the bytes verbatim). This is a **tripwire, not a sandbox** (see [Known constraints](#known-constraints-by-design-not-bugs)).
|
|
296
324
|
|
|
325
|
+
Each `pty_send` accepts at most 64 KiB of UTF-8 text. Sends to the same session are serialized across aiterm processes so chunks cannot interleave. On macOS, text is pasted through tmux in UTF-8-safe 256-byte chunks to avoid the platform PTY truncation observed with long input; Linux and WSL use one bounded paste. If a later chunk fails, aiterm reports the partial-send state and does not press Enter automatically. A lock left by a terminated sender fails closed before sending; close and recreate that session (or use `pty_kill_all` when every session is disposable) to clean it up safely.
|
|
326
|
+
|
|
297
327
|
## A human can watch
|
|
298
328
|
|
|
299
|
-
Sessions live on a shared tmux socket. The `tmux -S … attach -t <id>` line printed by `pty_open` (and by each agent launcher) lets a human attach to the same terminal and intervene (`Ctrl-b d` to detach) — including watching a launched Codex/Grok/Composer session run and taking the keyboard from your AI mid-task. On native Windows the printed line is the WSL form — `wsl tmux -S … attach -t <id>` — since the session lives inside WSL.
|
|
329
|
+
Sessions live on a shared tmux socket. The `tmux -S … attach -t <id>` line printed by `pty_open` (and by each agent launcher) lets a human attach to the same terminal and intervene (`Ctrl-b d` to detach) — including watching a launched Claude/Codex/Grok/Composer session run and taking the keyboard from your AI mid-task. On native Windows the printed line is the WSL form — `wsl tmux -S … attach -t <id>` — since the session lives inside WSL.
|
|
300
330
|
|
|
301
331
|
## Requirements
|
|
302
332
|
|
|
@@ -304,7 +334,7 @@ Sessions live on a shared tmux socket. The `tmux -S … attach -t <id>` line pri
|
|
|
304
334
|
- **tmux** (runtime prerequisite; check with `tmux -V`. Install with `apt install tmux` / `brew install tmux`)
|
|
305
335
|
- **macOS / Linux / WSL2** run tmux directly. On macOS install it with `brew install tmux` (stock macOS ships none). If your MCP client is launched from the **GUI** rather than a terminal, Homebrew's bin (`/opt/homebrew/bin` on Apple Silicon, `/usr/local/bin` on Intel) may be off its `PATH`; aiterm auto-searches those locations, or set **`AITERM_TMUX=/path/to/tmux`** to point at it explicitly.
|
|
306
336
|
- **Native Windows** has no tmux, so aiterm transparently runs tmux **inside WSL**. It needs [WSL](https://learn.microsoft.com/windows/wsl/) installed and initialized, with **tmux installed inside your WSL distro** (`sudo apt install tmux`); verify with `wsl tmux -V`. Sessions, the socket, and human `attach` all live on the WSL side — the AI just drives them from the Windows-side command. (You reach Windows tools the same way you reach SSH: `pty_send "powershell.exe …"` nests into PowerShell.)
|
|
307
|
-
- For the **agent launchers**: the corresponding vendor CLI, installed and authenticated — `codex` for `codex_agent`, `grok` for `grok_agent` / `composer_agent`. (Not needed if you only use the PTY tools.)
|
|
337
|
+
- For the **agent launchers**: the corresponding vendor CLI, installed and authenticated — `claude` for `claude_agent`, `codex` for `codex_agent`, `grok` for `grok_agent` / `composer_agent`. (Not needed if you only use the PTY tools.)
|
|
308
338
|
- Optional: the [`rtk`](https://github.com/rtk-ai/rtk) binary (used by `pty_send`'s `rtk: true` delegation; works fine without it)
|
|
309
339
|
|
|
310
340
|
## Known constraints (by design, not bugs)
|
|
@@ -312,7 +342,7 @@ Sessions live on a shared tmux socket. The `tmux -S … attach -t <id>` line pri
|
|
|
312
342
|
- **While nested (ssh / docker / REPL / a launched agent TUI), quiescence cannot fire by design**, because the foreground command is no longer in the shell set (bash/sh/zsh/fish/dash). When nested with no `until` and no `mark`, `pty_read({ wait: true })` returns early as `is_complete=False via nested` (rather than burning the full `timeout`, since no signal can confirm completion there) with a note to pass `until` (a literal substring by default; `until_regex: true` for a regex) or `mark: true` (an exit-code sentinel, auto-detected) for a confirmed completion. For a full-screen agent TUI, read `{ screen: true }` once its output settles.
|
|
313
343
|
- **`is_complete=False` is not a failure.** It means "completion was not observed within `timeout`." For long commands, raise `timeout` or use `until`/`mark`.
|
|
314
344
|
- **The destructive gate is a tripwire, not a sandbox.** It blocks common destructive forms only. It does **not** catch relative-path `rm`, things that become dangerous after `$VAR` expansion, or commands run on the far side of an SSH session — and it does not police what a launched coding agent does inside its own session.
|
|
315
|
-
- **The agent launchers spawn a vendor TUI; they don't wrap or proxy it.** aiterm validates prerequisites and starts the CLI in a persistent PTY — the model, auth, and behavior are the vendor CLI's. There is no
|
|
345
|
+
- **The agent launchers spawn a vendor TUI; they don't wrap or proxy it.** aiterm validates prerequisites and starts the CLI in a persistent PTY — the model, auth, and behavior are the vendor CLI's. There is no hidden inter-agent protocol; "conversation" is your MCP client driving the Claude/Codex/Grok/Composer TUI (send input, read output).
|
|
316
346
|
- **`pty_send({ rtk: true })` is single-line only and needs the external `rtk` binary** (passthrough without it). The `pty_read({ rtk: true })` reducer, by contrast, is self-contained and rtk-independent.
|
|
317
347
|
- **The `pytest` reducer matches rtk 0.42.0** on test counts, the rule line, and `FAILURES`-block formatting (locked by regression tests). It **deliberately preserves the full failure reason** on the `FAILED` summary lines (emitted under `-ra`/`-rf`), whereas rtk 0.42.0 truncates the reason at the first `" - "` — a readability choice, so those lines are intentionally not byte-identical to rtk. The `[full output: …]` tee-pointer line rtk appends on large output is not reproduced on the read side.
|
|
318
348
|
- **tmux is started with `-f /dev/null`**, so it does not read `~/.tmux.conf` (to keep behavior reproducible across machines).
|
|
@@ -344,4 +374,15 @@ If aiterm let your AI hand a task to another agent — or saved you a round-trip
|
|
|
344
374
|
|
|
345
375
|
## License
|
|
346
376
|
|
|
377
|
+
> Grok OAuthのauth/lock共有は0.9.1当時の契約で、2026-07-14に廃止した。現行はmanaged隔離を維持し、検証済み通常auth正本を`GROK_AUTH_PATH`でvendorへ渡す。aitermはlock・atomic replace・copy-backを所有しない。
|
|
378
|
+
|
|
347
379
|
MIT
|
|
380
|
+
|
|
381
|
+
## Grok OAuth isolation
|
|
382
|
+
|
|
383
|
+
For `agent_done`, Grok/Composer keep launch-local `GROK_HOME` and fake `HOME`.
|
|
384
|
+
The child receives `GROK_AUTH_PATH` pointing at the validated normal auth
|
|
385
|
+
canonical file; managed homes never contain auth or lock symlinks/copies.
|
|
386
|
+
aiterm does not create locks or copy credentials back. An inherited
|
|
387
|
+
`GROK_AUTH_PATH` must be absolute and safe; only an absent default auth file is
|
|
388
|
+
allowed when `XAI_API_KEY` is set.
|
|
@@ -0,0 +1,94 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// aiterm-wait — agent turn 完了eventの純リーダー観測CLI。
|
|
3
|
+
// 完了/timeout/close を1行のJSON receiptで返してexitする。lock・PTY・dispatch状態には一切触れない。
|
|
4
|
+
// 親AIホストのバックグラウンドタスクとして起動し、exitを「完了通知」として使う。
|
|
5
|
+
import { fileURLToPath } from "node:url";
|
|
6
|
+
import * as fs from "node:fs";
|
|
7
|
+
import { AitermError, observeAgentDone } from "./core.js";
|
|
8
|
+
const SESSION_RE = /^[A-Za-z0-9_-]{1,64}$/;
|
|
9
|
+
const OPERATION_RE = /^sha256:[0-9a-f]{64}$/;
|
|
10
|
+
const USAGE = "usage: aiterm-wait --session <name> [--operation sha256:<64hex>] [--timeout <sec>]";
|
|
11
|
+
export function parseArgs(argv) {
|
|
12
|
+
let session = null;
|
|
13
|
+
let operationId = null;
|
|
14
|
+
let timeout = null;
|
|
15
|
+
for (let i = 0; i < argv.length; i++) {
|
|
16
|
+
const a = argv[i];
|
|
17
|
+
if (a === "--session" || a === "--operation" || a === "--timeout") {
|
|
18
|
+
const v = argv[i + 1];
|
|
19
|
+
if (v === undefined)
|
|
20
|
+
throw new Error(`${a} に値がありません。${USAGE}`);
|
|
21
|
+
i++;
|
|
22
|
+
if (a === "--session") {
|
|
23
|
+
if (session !== null)
|
|
24
|
+
throw new Error(`--session が重複しています。${USAGE}`);
|
|
25
|
+
if (!SESSION_RE.test(v))
|
|
26
|
+
throw new Error(`--session が不正です。${USAGE}`);
|
|
27
|
+
session = v;
|
|
28
|
+
}
|
|
29
|
+
else if (a === "--operation") {
|
|
30
|
+
if (operationId !== null)
|
|
31
|
+
throw new Error(`--operation が重複しています。${USAGE}`);
|
|
32
|
+
if (!OPERATION_RE.test(v))
|
|
33
|
+
throw new Error(`--operation が不正です。${USAGE}`);
|
|
34
|
+
operationId = v;
|
|
35
|
+
}
|
|
36
|
+
else {
|
|
37
|
+
if (timeout !== null)
|
|
38
|
+
throw new Error(`--timeout が重複しています。${USAGE}`);
|
|
39
|
+
if (!/^\d+$/.test(v))
|
|
40
|
+
throw new Error(`--timeout は0以上の整数秒だけを受理します。${USAGE}`);
|
|
41
|
+
timeout = Number(v);
|
|
42
|
+
if (timeout > 86400)
|
|
43
|
+
throw new Error(`--timeout は86400秒以下だけを受理します。${USAGE}`);
|
|
44
|
+
}
|
|
45
|
+
}
|
|
46
|
+
else {
|
|
47
|
+
throw new Error(`不明な引数です: ${a}。${USAGE}`);
|
|
48
|
+
}
|
|
49
|
+
}
|
|
50
|
+
if (session === null)
|
|
51
|
+
throw new Error(`--session は必須です。${USAGE}`);
|
|
52
|
+
return { session, operationId, timeout: timeout ?? 600 };
|
|
53
|
+
}
|
|
54
|
+
function emit(value) {
|
|
55
|
+
process.stdout.write(JSON.stringify(value) + "\n");
|
|
56
|
+
}
|
|
57
|
+
export async function main(argv) {
|
|
58
|
+
const cmd = parseArgs(argv);
|
|
59
|
+
const result = await observeAgentDone(cmd.session, {
|
|
60
|
+
operation_id: cmd.operationId,
|
|
61
|
+
timeout: cmd.timeout,
|
|
62
|
+
});
|
|
63
|
+
emit(result);
|
|
64
|
+
}
|
|
65
|
+
function isDirectExecution() {
|
|
66
|
+
const entry = process.argv[1];
|
|
67
|
+
if (!entry)
|
|
68
|
+
return false;
|
|
69
|
+
try {
|
|
70
|
+
const self = fileURLToPath(import.meta.url);
|
|
71
|
+
const a = fs.realpathSync(entry);
|
|
72
|
+
const b = fs.realpathSync(self);
|
|
73
|
+
if (a === b)
|
|
74
|
+
return true;
|
|
75
|
+
return process.platform === "win32" && a.toLowerCase() === b.toLowerCase();
|
|
76
|
+
}
|
|
77
|
+
catch {
|
|
78
|
+
return false;
|
|
79
|
+
}
|
|
80
|
+
}
|
|
81
|
+
if (isDirectExecution()) {
|
|
82
|
+
main(process.argv.slice(2)).catch((e) => {
|
|
83
|
+
// 引数不正・相関エラーの文言は親AI向けに設計済みのため envelope に載せる。
|
|
84
|
+
// 想定外の例外は raw を反映せず固定文言に落とす(privacy-safe)。
|
|
85
|
+
const known = e instanceof AitermError || (e instanceof Error && e.message.includes(USAGE));
|
|
86
|
+
emit({
|
|
87
|
+
ok: false,
|
|
88
|
+
code: "AITERM_WAIT_FAILED",
|
|
89
|
+
message: known ? e.message : "aiterm-wait: operation failed",
|
|
90
|
+
});
|
|
91
|
+
process.stderr.write("aiterm-wait: operation failed\n");
|
|
92
|
+
process.exitCode = 1;
|
|
93
|
+
});
|
|
94
|
+
}
|