conductor-remote 1.111.0 → 1.113.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## PWA fleet call over WebRTC
4
4
 
5
- The primary foreground voice mode is one fleet-wide control room. In the
5
+ The workspace-list phone opens a fleet-wide control room. In the
6
6
  workspace-list header, tap the phone immediately left of **+**, then start the
7
7
  call. The orchestrator surveys every workspace, presents one bounded decision
8
8
  at a time, creates new workspaces, and queues exact prompts. Both writes happen
@@ -10,9 +10,23 @@ only after it reads the target and text back and you confirm. It is not owned by
10
10
  the chat currently on screen: hide the sheet or move between workspaces and the
11
11
  same call continues.
12
12
 
13
+ To call the current chat, tap the phone on the right side of the chat composer,
14
+ immediately left of the context control, then **Start workspace call**.
15
+ **Call this workspace** for the active chat is also available in the command menu.
16
+ The call starts with that chat's recent
17
+ conversation already loaded, so you can discuss the task or ask for an update.
18
+ The sheet names both the workspace and chat. Browsing another tab keeps the call
19
+ on its original conversation; end it before starting a call for another chat.
20
+
21
+ The relay reads up to 24 recent user and assistant messages, capped at 16,000
22
+ characters, keeping the latest user request even after a long run. Queued prompts,
23
+ reasoning, tool output, and native child-agent messages are excluded. These messages
24
+ are sent to OpenAI with the call; an update request reads the chat again. Sending a
25
+ prompt back still requires spoken confirmation.
26
+
13
27
  Live captions keep both sides readable, and the text box in the call sheet is a
14
- fallback when speaking is inconvenient. The eight available actions are roll
15
- call, fresh workspace overview, next decision, repository list, workspace-create
28
+ fallback when speaking is inconvenient. The nine available actions are roll
29
+ call, fresh workspace overview, chat context, next decision, repository list, workspace-create
16
30
  preview and confirmed creation, plus send preview and confirmed send. The two
17
31
  writes use the same persisted, one-use preview/confirmation gate as the dial-in
18
32
  orchestrator.
@@ -26,7 +40,7 @@ how recently its selected chat changed.
26
40
  This path needs the managed relay, its usual private phone URL, and an OpenAI API
27
41
  key. It does **not** need a phone number, SIP, a webhook, Funnel, or any other
28
42
  public endpoint. The permanent key stays on the Mac: the relay creates the
29
- WebRTC call, then executes those eight tools over its private sideband connection.
43
+ WebRTC call, then executes those tools over its private sideband connection.
30
44
 
31
45
  Store the key without leaving it in shell history:
32
46
 
@@ -57,9 +71,10 @@ screen locked, so use the optional dial-in transport for a pocketed commute.
57
71
 
58
72
  The optional voice listener turns a phone call into a small Conductor control room: hear a bounded fleet tally, walk one decision at a time, create a workspace, and dispatch an exact prompt after a spoken read-back and explicit confirmation. It does not expose the PWA or the relay API publicly.
59
73
 
60
- The scoped endpoint has eight tools: roll call, a fresh paged workspace overview,
74
+ The scoped endpoint has nine tools: roll call, a fresh paged workspace overview, chat context,
61
75
  next decision, repository list, create preview, confirmed create, send preview,
62
- and confirmed send. Forward-to-owner answers, artifact pushes, and voice grooming
76
+ and confirmed send. Dial-in calls open with the fleet roll call.
77
+ Forward-to-owner answers, artifact pushes, and voice grooming
63
78
  remain later milestones.
64
79
 
65
80
  ## What you need
@@ -154,7 +169,7 @@ conductor-remote config set voice.webhook-secret "$OPENAI_WEBHOOK_SECRET"
154
169
  unset OPENAI_WEBHOOK_SECRET
155
170
  ```
156
171
 
157
- The relay accepts a valid incoming call through `POST /v1/realtime/calls/{call_id}/accept`, attaches an authenticated sideband WebSocket, and gives the session only the eight scoped remote MCP tools. Remote MCP follow-up responses are driven by the broker only after both the response and every tool call in it have finished, as required by OpenAI's [Realtime MCP guide](https://developers.openai.com/api/docs/guides/realtime-mcp).
172
+ The relay accepts a valid incoming call through `POST /v1/realtime/calls/{call_id}/accept`, attaches an authenticated sideband WebSocket, and gives the session only the nine scoped remote MCP tools. Remote MCP follow-up responses are driven by the broker only after both the response and every tool call in it have finished, as required by OpenAI's [Realtime MCP guide](https://developers.openai.com/api/docs/guides/realtime-mcp).
158
173
 
159
174
  ## 5. Configure the Twilio number
160
175
 
@@ -194,6 +209,30 @@ conductor-remote service logs
194
209
 
195
210
  Each call starts with the Mac's lock state. A confirmed send returns to the voice session immediately and reuses the relay's existing transcript receipt, retry, idempotency, and parked-prompt path. A landed send stays silent; a locked or failed send is announced. Merely hearing a decision does not clear it—dispatching it or explicitly skipping it advances the read mark.
196
211
 
212
+ ### Saved transcripts
213
+
214
+ New browser and dial-in calls are saved automatically on the Mac. Open **Control room → Call history** to read a past call, copy its text, or export a `.txt` file. Hanging up a browser call opens its saved transcript. History is available to every device authenticated to this relay, including when voice calling is no longer configured.
215
+
216
+ The archive lives at `~/Library/Application Support/conductor-remote/voice-history.db`, with owner-only file permissions. It is a separate SQLite database owned by the relay; Conductor's database remains read-only. Back it up using SQLite's backup facility, or stop the relay before copying it, since an active database can have a `-wal` file alongside it. Calls are kept indefinitely.
217
+
218
+ The relay saves caller transcriptions, typed messages, assistant text, tool names, and call timestamps through its existing OpenAI sideband connection. Completed utterances are committed immediately; partial captions are checkpointed every half-second and flushed on shutdown. Audio recordings, raw event payloads, tool arguments and authentication headers are not archived. Saving works while the call panel is hidden and does not depend on a final upload from the phone.
219
+
220
+ Transcription can finish out of order, so the archive follows conversation item IDs and predecessor links rather than the arrival order of captions. See OpenAI's [transcription events](https://developers.openai.com/api/docs/guides/realtime-transcription) and [sideband controls](https://developers.openai.com/api/docs/guides/realtime-server-controls).
221
+
222
+ An interrupted reply is labeled because generated text can include words that were never played. Failed transcription is shown explicitly. A lost observer connection or relay restart marks possible gaps; reconnecting does not promise to recover events missed while disconnected. Previously completed calls that were never captured cannot be backfilled. Storage failures appear in the call panel and history, with details in the relay logs.
223
+
224
+ Authenticated reads are `GET /api/voice/history?limit=30&offset=0` for call summaries and `GET /api/voice/history/:callId` for one complete transcript. `?summary=1` reads its recording status without transferring the conversation.
225
+
226
+ The standard conductor-remote MCP server exposes the same archive over both stdio and HTTP:
227
+
228
+ - `list_voice_calls` lists recent calls with previews and call IDs. Use `limit` and `offset` to page through older calls.
229
+ - `search_voice_calls` searches caller and assistant text, with the same words/quoted-phrase grammar as `search_chats`. Each hit includes a call ID and an item ID; `call_id` can restrict the search to one call. Partial captions are searchable and are replaced when their corrected final text arrives.
230
+ - `read_voice_call` reads a call's latest entries, or a window around `near` (an item ID from search). `before`, `after`, `limit`, and `max_chars` bound the result. `older_item` and `newer_item` let the agent continue through the conversation. Gaps, failed transcription and interrupted replies remain explicit in MCP output.
231
+
232
+ For example, an agent can call `search_voice_calls({"query":"\"release Friday\""})`, then `read_voice_call({"call_id":"rtc_…","near":"item_…","before":4,"after":4})` with IDs from the result. These tools only read the archive through authenticated relay routes and never drive Conductor's UI or start a call. They are part of the main MCP server; the live voice session's scoped action tools remain separate. After installing the release, reconnect an existing MCP client so it discovers the new tools.
233
+
234
+ Search uses `GET /api/voice/search?q=…&limit=12&offset=0`, optionally with `callId=…`. Its local full-text index contains only caller and assistant text; tool payloads and internal relay nudges are excluded. Existing archives are indexed automatically without replacing their saved transcripts.
235
+
197
236
  ### Cost
198
237
 
199
238
  The default is `gpt-realtime-2.1-mini`. OpenAI currently publishes these per-million-token prices:
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "conductor-remote",
3
- "version": "1.111.0",
3
+ "version": "1.113.0",
4
4
  "type": "module",
5
5
  "packageManager": "yarn@4.15.0",
6
6
  "description": "Phone control panel for local Conductor agents. Reads ride SQLite + git; prompts ride Conductor's own dispatch path.",