osborn 0.9.94 → 0.9.95

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,34 @@
1
+ # Browser Screen Recorder — drive any web app (voice, hands + a brain) and record proof
2
+
3
+ ## SKILL IDENTITY
4
+ Name: browser-screen-recorder
5
+ Version: 1 (packaged seed — the served copy at the URL below is source of truth)
6
+ Served from: https://www.voice-native.com/api/browser-screen-recorder
7
+ Landing page: https://www.voice-native.com/browser-screen-recorder
8
+ (Formerly "voice-e2e". The old /api/test-skill URL still resolves to this skill.)
9
+
10
+ Give a coding agent a web app + an intent and it drives the app in a REAL
11
+ browser and records PROOF it did — per-action screenshots + video clips, audio
12
+ capture, DevTools diagnostics, and metrics. It can speak into the page's mic
13
+ (synthesized on the fly), hear and transcribe what it plays back, click via a
14
+ natural-language brain, and improvise conversations. Works on any website.
15
+
16
+ ## UPDATE CHECK (notify-first, run before major runs)
17
+ This packaged copy ships with every machine and may lag the served version.
18
+ Check, then NOTIFY the user before updating — never silently force:
19
+
20
+ ```bash
21
+ SERVED_V=$(curl -sf --max-time 10 https://www.voice-native.com/api/browser-screen-recorder | grep -m1 '^Version:' | awk '{print $2}')
22
+ LOCAL_V=$(grep -m1 '^Version:' ~/.claude/skills/browser-screen-recorder/SKILL.md 2>/dev/null | awk '{print $2}')
23
+ [ -n "$SERVED_V" ] && [ "$SERVED_V" != "$LOCAL_V" ] && echo "UPDATE AVAILABLE: browser-screen-recorder v${LOCAL_V:-none} -> v$SERVED_V — tell the user what's new and ask before updating (curl -sS https://www.voice-native.com/api/browser-screen-recorder -o ~/.claude/skills/browser-screen-recorder/SKILL.md), then refresh the harness bundle per the served INSTALLATION."
24
+ ```
25
+
26
+ If the user approves (or has pre-approved auto-updates), update the skill
27
+ file AND re-run the served INSTALLATION step 2 (harness bundle refresh) so
28
+ the accompanying files stay in lockstep with the skill.
29
+
30
+ ## FIRST USE
31
+ Fetch the served copy and follow its INSTALLATION section — it materializes
32
+ the harness (~/browser-screen-recorder-harness/), installs dependencies, and
33
+ asks the user for a GOOGLE_API_KEY (required, Gemini brain) plus optional
34
+ OPENAI/DEEPGRAM keys for the voice mouth and audible verification.
@@ -58,7 +58,15 @@ curl -sS -X POST http://localhost:8741/canvas -H 'Content-Type: application/json
58
58
  curl -sS -X POST http://localhost:8741/canvas -H 'Content-Type: application/json' \
59
59
  -d '{"kind":"show","mode":"notes","title":"...","items":["...","..."]}'
60
60
  ```
61
- `mode` = `idle` | `notes` (title+items) | `link` (url) | `web` (iframe url) | `text` (title+text).
61
+ `mode` = `idle` | `notes` (title+items) | `link` (url) | `web` (iframe url) | `text` (title+text)
62
+ | `stream` (url — renders `<img src="{url}/stream">`, a live MJPEG browser feed).
63
+
64
+ **`stream` mode — NO TUNNELS policy:** the feed URL must be PUBLIC. Never
65
+ tunnel a local browser engine to get one (ngrok's free tier burned its entire
66
+ monthly bandwidth cap on one continuous-MJPEG demo — `ERR_NGROK_725`). Run the
67
+ browser-screen-recorder engine on its Fly machine instead; its `:8080` MJPEG is
68
+ already public at `https://<app>.fly.dev/`. See the browser-screen-recorder
69
+ skill ("Casting the feed into a meeting") for details.
62
70
 
63
71
  **When to speak into the meeting:** By DEFAULT stay silent (observer) for
64
72
  `[MEETING — *]:` chunks — take notes, don't interrupt. Speak into the meeting
package/dist/index.js CHANGED
@@ -1619,6 +1619,64 @@ async function main() {
1619
1619
  }
1620
1620
  meetingTranscriptBuffer.length = 0;
1621
1621
  };
1622
+ // ── Meeting lifecycle (centralized teardown, 0.9.95) ──
1623
+ // The bot's lifecycle FOLLOWS the voice session (deliberate coupling — a
1624
+ // decoupled always-on bot means untracked background agents; revisit only
1625
+ // with a status/tracking UI). endMeeting() is the single teardown path,
1626
+ // fired by: (a) explicit leave_meeting, (b) user disconnect (auto-leave),
1627
+ // (c) Recall reporting a terminal bot status (call_ended/done/fatal —
1628
+ // detected on the poller's 30s tick; fixes stale bot state when the meeting
1629
+ // ends naturally), or (d) the max-duration backstop below.
1630
+ const MEETING_MAX_MS = Math.max(10, Number(process.env.OSBORN_MEETING_MAX_MIN) || 180) * 60 * 1000; // default 3h
1631
+ let meetingMaxTimer = null;
1632
+ const endMeeting = async (reason, opts = {}) => {
1633
+ const botId = activeMeetingBotId;
1634
+ if (!botId)
1635
+ return;
1636
+ console.log(`🏁 Meeting ended (${reason}) — bot ${botId}`);
1637
+ // Stop inputs FIRST so no more chunks land mid-teardown.
1638
+ stopMeetingFlush();
1639
+ if (activeMeetingPoller) {
1640
+ activeMeetingPoller.stop();
1641
+ activeMeetingPoller = null;
1642
+ }
1643
+ if (meetingMaxTimer) {
1644
+ clearTimeout(meetingMaxTimer);
1645
+ meetingMaxTimer = null;
1646
+ }
1647
+ activeMeetingBotId = null;
1648
+ // leaveBot=false when Recall itself reported the meeting over (bot already gone).
1649
+ if (opts.leaveBot !== false) {
1650
+ const recall = getRecallClient();
1651
+ if (recall)
1652
+ await recall.leaveMeeting(botId).catch((e) => console.warn(`⚠️ leaveMeeting failed: ${e?.message}`));
1653
+ }
1654
+ sendToFrontend({ type: 'meeting_left', botId, reason }).catch(() => { });
1655
+ // Orphan cleanup: if no user is connected, the LLM subprocess was kept alive
1656
+ // solely to serve the meeting — release it now and let idle-exit stop the
1657
+ // machine (same billing discipline as 0.9.73, deferred until meeting end).
1658
+ const userPresent = activeRoom && activeRoom.remoteParticipants.size > 0;
1659
+ if (!userPresent && currentLLM) {
1660
+ console.log('🏁 Meeting over with no user connected — releasing LLM + arming idle-exit');
1661
+ killCurrentLLM(`meeting_ended(${reason})_no_user`);
1662
+ currentLLM = null;
1663
+ clearFastBrainSession();
1664
+ clearPipelineFastBrainSession();
1665
+ armIdleExitTimer(`meeting ended (${reason}), no user`);
1666
+ }
1667
+ };
1668
+ const armMeetingMaxTimer = (botId) => {
1669
+ if (meetingMaxTimer)
1670
+ clearTimeout(meetingMaxTimer);
1671
+ console.log(`⏲️ Meeting max-duration backstop armed: ${MEETING_MAX_MS / 60000} min (override OSBORN_MEETING_MAX_MIN)`);
1672
+ meetingMaxTimer = setTimeout(() => {
1673
+ meetingMaxTimer = null;
1674
+ if (activeMeetingBotId === botId) {
1675
+ console.log('⏲️ Meeting max duration reached — auto-leaving (backstop, not Recall-reported end)');
1676
+ void endMeeting('max_duration_backstop');
1677
+ }
1678
+ }, MEETING_MAX_MS);
1679
+ };
1622
1680
  // Track the active resume session ID across scopes (ParticipantConnected + DataReceived)
1623
1681
  // Updated by resume_session, session_selected, continue_session, switch_session handlers
1624
1682
  let currentResumeSessionId;
@@ -4195,20 +4253,13 @@ async function main() {
4195
4253
  currentLLM = null;
4196
4254
  clearFastBrainSession();
4197
4255
  clearPipelineFastBrainSession();
4198
- // Auto-leave any active meeting bot when user disconnects from the room
4199
- stopMeetingFlush();
4200
- if (activeMeetingPoller) {
4201
- activeMeetingPoller.stop();
4202
- activeMeetingPoller = null;
4203
- }
4204
- if (activeMeetingBotId) {
4205
- const recallDisconnect = getRecallClient();
4206
- if (recallDisconnect) {
4207
- console.log(`🤝 Auto-leaving meeting (bot ${activeMeetingBotId}) — user disconnected from room`);
4208
- recallDisconnect.leaveMeeting(activeMeetingBotId).catch(() => { });
4209
- activeMeetingBotId = null;
4210
- }
4211
- }
4256
+ // Auto-leave any active meeting bot when user disconnects from the room.
4257
+ // DELIBERATE COUPLING: the bot's lifecycle follows the voice session — a
4258
+ // decoupled always-on meeting bot means untracked background agents (no
4259
+ // status surface for them yet). Revisit only WITH a tracking UI.
4260
+ // endMeeting() centralizes the teardown (flush, poller, max-timer, Recall
4261
+ // leave, frontend notify).
4262
+ void endMeeting('user_disconnected');
4212
4263
  // 0.9.83: a real session just ended → use the FAST leave (~20s), not the
4213
4264
  // 3-min alone grace. Runs on the agent, so it fires even on an abrupt tab
4214
4265
  // close. Cancelled if a user rejoins within the grace (handled where the
@@ -4880,11 +4931,17 @@ async function main() {
4880
4931
  console.warn(`⚠️ Failed to forward meeting transcript to LLM: ${err.message}`);
4881
4932
  }
4882
4933
  },
4934
+ // Authoritative meeting-over signal: Recall's terminal bot status
4935
+ // (call_ended/done/fatal), checked on the same 30s tick. The bot
4936
+ // is already gone at that point, so don't re-issue leave.
4937
+ onMeetingEnd: (code) => { void endMeeting(`recall_status:${code}`, { leaveBot: false }); },
4883
4938
  });
4884
4939
  activeMeetingPoller.start();
4885
4940
  // LIVE path: buffer webhook finals + flush to the LLM every 20s.
4886
4941
  // (The poller above only lands data after the meeting ENDS.)
4887
4942
  startMeetingFlush(botId);
4943
+ // Billing backstop: a forgotten meeting can't hold the machine forever.
4944
+ armMeetingMaxTimer(botId);
4888
4945
  }
4889
4946
  catch (err) {
4890
4947
  console.error('❌ Recall.ai join error:', err);
@@ -4894,25 +4951,14 @@ async function main() {
4894
4951
  }
4895
4952
  }
4896
4953
  else if (data.type === 'leave_meeting') {
4897
- const botId = data.botId;
4898
- const recallLeave = getRecallClient();
4899
- if (recallLeave && botId) {
4900
- try {
4901
- // Stop the transcript poller + live flush FIRST so no more chunks
4902
- // get forwarded to the LLM during the leave.
4903
- stopMeetingFlush();
4904
- if (activeMeetingPoller) {
4905
- activeMeetingPoller.stop();
4906
- activeMeetingPoller = null;
4907
- }
4908
- await recallLeave.leaveMeeting(botId);
4909
- activeMeetingBotId = null;
4910
- await sendToFrontend({ type: 'meeting_left', botId });
4911
- }
4912
- catch (err) {
4913
- console.error('❌ Recall.ai leave error:', err);
4914
- await sendToFrontend({ type: 'meeting_error', message: err.message });
4915
- }
4954
+ // All teardown (flush, poller, max-timer, Recall leave, meeting_left,
4955
+ // orphan LLM release) is centralized in endMeeting().
4956
+ try {
4957
+ await endMeeting('user_leave_meeting');
4958
+ }
4959
+ catch (err) {
4960
+ console.error('❌ Recall.ai leave error:', err);
4961
+ await sendToFrontend({ type: 'meeting_error', message: err.message });
4916
4962
  }
4917
4963
  }
4918
4964
  else if (data.type === 'session_selected') {
@@ -41,6 +41,15 @@ export interface MeetingTranscriptPollerOptions {
41
41
  intervalMs?: number;
42
42
  /** Optional debug logger. */
43
43
  onError?: (err: Error) => void;
44
+ /**
45
+ * Called ONCE when Recall reports the meeting is over (bot status reaches a
46
+ * terminal code: call_ended / done / fatal). Rides the same 30s tick — no
47
+ * extra webhook subscription needed. The meeting-copilot lifecycle uses this
48
+ * as the authoritative "meeting ended" signal now that the bot is DECOUPLED
49
+ * from the voice session (a voice disconnect no longer drops the bot, so
50
+ * something must still end it — this is that something).
51
+ */
52
+ onMeetingEnd?: (statusCode: string) => void;
44
53
  }
45
54
  export declare class MeetingTranscriptPoller {
46
55
  #private;
@@ -27,6 +27,8 @@
27
27
  * No backoff — Recall's transcript endpoint is stable enough that a 30s
28
28
  * cadence makes "slow start" non-issues self-recover within one cycle.
29
29
  */
30
+ /** Recall bot status codes that mean the meeting/bot is definitively over. */
31
+ const TERMINAL_STATUS_CODES = new Set(['call_ended', 'done', 'fatal']);
30
32
  export class MeetingTranscriptPoller {
31
33
  #opts;
32
34
  #timer = null;
@@ -60,25 +62,43 @@ export class MeetingTranscriptPoller {
60
62
  return;
61
63
  this.#inFlight = true;
62
64
  try {
65
+ // Meeting-end detection first: if the bot has reached a terminal status,
66
+ // fire onMeetingEnd (once — stop() sets #stopped so later ticks no-op)
67
+ // AFTER one final transcript pull below, so the last turns aren't lost.
68
+ let endedCode = null;
69
+ if (this.#opts.onMeetingEnd) {
70
+ try {
71
+ const code = await this.#opts.recall.getBotStatus(this.#opts.botId);
72
+ if (TERMINAL_STATUS_CODES.has(code))
73
+ endedCode = code;
74
+ }
75
+ catch { /* status probe is best-effort; transcript pull continues */ }
76
+ }
63
77
  const all = await this.#opts.recall.getTranscript(this.#opts.botId);
64
78
  const fresh = all.filter(t => {
65
79
  const firstWordTs = t.words?.[0]?.start_timestamp?.relative;
66
80
  return typeof firstWordTs === 'number' && firstWordTs > this.#cursor;
67
81
  });
68
- if (fresh.length === 0)
69
- return;
70
- // Advance cursor to highest seen first-word ts (across all returned turns,
71
- // not just the fresh ones — guards against Recall returning a paged subset).
72
- for (const t of all) {
73
- const ts = t.words?.[0]?.start_timestamp?.relative;
74
- if (typeof ts === 'number' && ts > this.#cursor)
75
- this.#cursor = ts;
82
+ if (fresh.length > 0) {
83
+ // Advance cursor to highest seen first-word ts (across all returned turns,
84
+ // not just the fresh ones — guards against Recall returning a paged subset).
85
+ for (const t of all) {
86
+ const ts = t.words?.[0]?.start_timestamp?.relative;
87
+ if (typeof ts === 'number' && ts > this.#cursor)
88
+ this.#cursor = ts;
89
+ }
90
+ const formatted = formatTurns(fresh);
91
+ if (formatted) {
92
+ console.log(`📓 MeetingTranscriptPoller: ${fresh.length} new turn(s), cursor=${this.#cursor.toFixed(1)}s, chars=${formatted.length}`);
93
+ await this.#opts.onTurns({ botId: this.#opts.botId, turns: fresh, formatted });
94
+ }
95
+ }
96
+ // Fire end-of-meeting AFTER the final transcript pull so the last turns land.
97
+ if (endedCode) {
98
+ console.log(`📓 MeetingTranscriptPoller: bot status '${endedCode}' — meeting ended`);
99
+ this.stop();
100
+ this.#opts.onMeetingEnd?.(endedCode);
76
101
  }
77
- const formatted = formatTurns(fresh);
78
- if (!formatted)
79
- return; // pure-whitespace fresh batch — skip
80
- console.log(`📓 MeetingTranscriptPoller: ${fresh.length} new turn(s), cursor=${this.#cursor.toFixed(1)}s, chars=${formatted.length}`);
81
- await this.#opts.onTurns({ botId: this.#opts.botId, turns: fresh, formatted });
82
102
  }
83
103
  catch (err) {
84
104
  const e = err instanceof Error ? err : new Error(String(err));
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "osborn",
3
- "version": "0.9.94",
3
+ "version": "0.9.95",
4
4
  "description": "Voice AI coding assistant - local agent that connects to Osborn frontend",
5
5
  "type": "module",
6
6
  "bin": {
@@ -1,32 +0,0 @@
1
- # Voice-E2E — browser testing with ears, a mouth, hands and a brain
2
-
3
- ## SKILL IDENTITY
4
- Name: voice-e2e
5
- Version: 1 (packaged seed — the served copy at the URL below is source of truth)
6
- Served from: https://www.voice-native.com/api/test-skill
7
- Landing page: https://www.voice-native.com/test-skill
8
-
9
- Agent-driven end-to-end testing for voice and web apps in a REAL browser:
10
- speak into the page's mic (synthesized on the fly), hear and transcribe what
11
- it plays back, click via a natural-language brain, improvise conversations,
12
- and get video + audio replays per run. Works on any website.
13
-
14
- ## UPDATE CHECK (notify-first, run before major test/research runs)
15
- This packaged copy ships with every machine and may lag the served version.
16
- Check, then NOTIFY the user before updating — never silently force:
17
-
18
- ```bash
19
- SERVED_V=$(curl -sf --max-time 10 https://www.voice-native.com/api/test-skill | grep -m1 '^Version:' | awk '{print $2}')
20
- LOCAL_V=$(grep -m1 '^Version:' ~/.claude/skills/voice-e2e/SKILL.md 2>/dev/null | awk '{print $2}')
21
- [ -n "$SERVED_V" ] && [ "$SERVED_V" != "$LOCAL_V" ] && echo "UPDATE AVAILABLE: voice-e2e v${LOCAL_V:-none} -> v$SERVED_V — tell the user what's new and ask before updating (curl -sS https://www.voice-native.com/api/test-skill -o ~/.claude/skills/voice-e2e/SKILL.md), then refresh the harness bundle per the served INSTALLATION."
22
- ```
23
-
24
- If the user approves (or has pre-approved auto-updates), update the skill
25
- file AND re-run the served INSTALLATION step 2 (harness bundle refresh) so
26
- the accompanying files stay in lockstep with the skill.
27
-
28
- ## FIRST USE
29
- Fetch the served copy and follow its INSTALLATION section — it materializes
30
- the harness (~/voice-e2e-harness/), installs dependencies, and asks the user
31
- for a GOOGLE_API_KEY (required, Gemini brain) plus optional OPENAI/DEEPGRAM
32
- keys for the voice mouth and audible verification.