osborn 0.9.88 → 0.9.90

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -22,18 +22,35 @@ Two trigger patterns:
22
22
 
23
23
  **Do NOT use this skill** for normal user voice-native messages that don't fit those patterns — those get spoken responses as usual.
24
24
 
25
+ ## CRITICAL: delegate the file + transcript work to the `writer` sub-agent
26
+
27
+ The main orchestrator agent has a **hard limit of 3 direct tool calls per turn** (enforced in PreToolUse — Read/Write/Bash/Glob are DENIED after the 3rd call). Writing `meeting-todos.md` and pulling transcripts (curl + jq + Write) is far more than 3 calls, so **doing it directly gets you blocked** ("all tools blocked").
28
+
29
+ **Sub-agents are exempt from this budget and have full permissions.** So for ALL meeting file/transcript work, **delegate to the `writer` sub-agent in ONE `Task` call** and let it do the whole job (fetch transcript, parse, write `meeting-todos.md`). That's a single tool call for you, and the writer has no budget cap.
30
+
31
+ ```
32
+ Task(
33
+ subagent_type: 'writer',
34
+ run_in_background: true, // silent — don't block voice
35
+ description: 'update meeting-todos.md',
36
+ prompt: '<the full instructions below: workspace path, bot ID, what to fetch/parse/write>'
37
+ )
38
+ ```
39
+
40
+ Give the writer everything it needs in the prompt: the session-workspace path, the bot ID, the `us-west-2.recall.ai` endpoint rule, and the `meeting-todos.md` structure. The writer runs the curl/jq/Write steps itself. For research, delegate to the `researcher` sub-agent the same way.
41
+
25
42
  ## How to behave (auto-tagged chunks)
26
43
 
27
44
  For every `[MEETING — *]:` message:
28
45
 
29
46
  1. **Do NOT speak.** No TTS output. No conversational reply.
30
- 2. **Update `meeting-todos.md`** in the session workspace. Append new action items, decisions, open questions. One file, evolving.
31
- 3. **Optionally trigger background research silently** via Task tool.
47
+ 2. **Delegate to the `writer` sub-agent** (see above) to append new action items, decisions, and open questions to `meeting-todos.md`. Do NOT write the file yourself — you'll hit the 3-call budget and get blocked. Batch chunks if they arrive faster than the writer finishes; one evolving file.
48
+ 3. **Delegate research to the `researcher` sub-agent** via `Task` (background, silent) when a chunk warrants it.
32
49
  4. **Don't consume voice-native attention.** The user can interrupt with a voice-native message at any time — that's the only kind that gets spoken responses.
33
50
 
34
51
  ## How to pull transcripts on demand (Bash + curl)
35
52
 
36
- When the user explicitly asks (see triggers above), run these commands. Speak briefly first ("On it"), do the work, then speak the result.
53
+ When the user explicitly asks (see triggers above): speak briefly first ("On it"), then **delegate the fetch+parse+write to the `writer` sub-agent** in one `Task` call (the steps below are what you put in the writer's prompt — they're 4+ Bash/Write calls, over your 3-call budget). When the writer finishes, speak the result. The commands below are the recipe the writer runs, not calls you make directly.
37
54
 
38
55
  ### Step 1: Get the bot ID
39
56
 
package/dist/index.js CHANGED
@@ -515,6 +515,41 @@ function startApiServer(workingDir, port) {
515
515
  console.log(`[canvas] client connected (${canvasClients.size} total)`);
516
516
  return;
517
517
  }
518
+ // ── Meeting canvas: TTS audio for speaking INTO the meeting ──────────────
519
+ // GET /tts?text=... → mp3 (OpenAI TTS). The canvas plays this as a real
520
+ // <audio> element so Recall's webpage output pipes it into the meeting —
521
+ // speechSynthesis is NOT captured by Recall, a media element IS.
522
+ if (req.method === 'GET' && url.pathname === '/tts') {
523
+ const text = (url.searchParams.get('text') || '').slice(0, 4000);
524
+ const voice = url.searchParams.get('voice') || 'alloy';
525
+ const key = process.env.OPENAI_API_KEY;
526
+ if (!text || !key) {
527
+ res.writeHead(400, { 'Content-Type': 'application/json' });
528
+ res.end(JSON.stringify({ error: !key ? 'no OPENAI_API_KEY' : 'no text' }));
529
+ return;
530
+ }
531
+ try {
532
+ const tts = await fetch('https://api.openai.com/v1/audio/speech', {
533
+ method: 'POST',
534
+ headers: { 'Authorization': `Bearer ${key}`, 'Content-Type': 'application/json' },
535
+ body: JSON.stringify({ model: 'gpt-4o-mini-tts', voice, input: text, response_format: 'mp3' }),
536
+ });
537
+ if (!tts.ok) {
538
+ const e = await tts.text().catch(() => '');
539
+ res.writeHead(502, { 'Content-Type': 'application/json' });
540
+ res.end(JSON.stringify({ error: `tts ${tts.status}`, detail: e.slice(0, 200) }));
541
+ return;
542
+ }
543
+ const buf = Buffer.from(await tts.arrayBuffer());
544
+ res.writeHead(200, { 'Content-Type': 'audio/mpeg', 'Cache-Control': 'no-store', 'Content-Length': buf.length });
545
+ res.end(buf);
546
+ }
547
+ catch (e) {
548
+ res.writeHead(500, { 'Content-Type': 'application/json' });
549
+ res.end(JSON.stringify({ error: e.message }));
550
+ }
551
+ return;
552
+ }
518
553
  // ── Meeting canvas: control endpoint (director / agent tool) ─────────────
519
554
  // POST /canvas { kind:'say', text } | { kind:'show', mode, title?, items?, url?, text? }
520
555
  if (req.method === 'POST' && url.pathname === '/canvas') {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "osborn",
3
- "version": "0.9.88",
3
+ "version": "0.9.90",
4
4
  "description": "Voice AI coding assistant - local agent that connects to Osborn frontend",
5
5
  "type": "module",
6
6
  "bin": {