osborn 0.9.94 → 0.9.95
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/browser-screen-recorder/SKILL.md +34 -0
- package/.claude/skills/meetings/SKILL.md +9 -1
- package/dist/index.js +79 -33
- package/dist/meeting-transcript-poller.d.ts +9 -0
- package/dist/meeting-transcript-poller.js +33 -13
- package/package.json +1 -1
- package/.claude/skills/voice-e2e/SKILL.md +0 -32
|
@@ -0,0 +1,34 @@
|
|
|
1
|
+
# Browser Screen Recorder — drive any web app (voice, hands + a brain) and record proof
|
|
2
|
+
|
|
3
|
+
## SKILL IDENTITY
|
|
4
|
+
Name: browser-screen-recorder
|
|
5
|
+
Version: 1 (packaged seed — the served copy at the URL below is source of truth)
|
|
6
|
+
Served from: https://www.voice-native.com/api/browser-screen-recorder
|
|
7
|
+
Landing page: https://www.voice-native.com/browser-screen-recorder
|
|
8
|
+
(Formerly "voice-e2e". The old /api/test-skill URL still resolves to this skill.)
|
|
9
|
+
|
|
10
|
+
Give a coding agent a web app + an intent and it drives the app in a REAL
|
|
11
|
+
browser and records PROOF it did — per-action screenshots + video clips, audio
|
|
12
|
+
capture, DevTools diagnostics, and metrics. It can speak into the page's mic
|
|
13
|
+
(synthesized on the fly), hear and transcribe what it plays back, click via a
|
|
14
|
+
natural-language brain, and improvise conversations. Works on any website.
|
|
15
|
+
|
|
16
|
+
## UPDATE CHECK (notify-first, run before major runs)
|
|
17
|
+
This packaged copy ships with every machine and may lag the served version.
|
|
18
|
+
Check, then NOTIFY the user before updating — never silently force:
|
|
19
|
+
|
|
20
|
+
```bash
|
|
21
|
+
SERVED_V=$(curl -sf --max-time 10 https://www.voice-native.com/api/browser-screen-recorder | grep -m1 '^Version:' | awk '{print $2}')
|
|
22
|
+
LOCAL_V=$(grep -m1 '^Version:' ~/.claude/skills/browser-screen-recorder/SKILL.md 2>/dev/null | awk '{print $2}')
|
|
23
|
+
[ -n "$SERVED_V" ] && [ "$SERVED_V" != "$LOCAL_V" ] && echo "UPDATE AVAILABLE: browser-screen-recorder v${LOCAL_V:-none} -> v$SERVED_V — tell the user what's new and ask before updating (curl -sS https://www.voice-native.com/api/browser-screen-recorder -o ~/.claude/skills/browser-screen-recorder/SKILL.md), then refresh the harness bundle per the served INSTALLATION."
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
If the user approves (or has pre-approved auto-updates), update the skill
|
|
27
|
+
file AND re-run the served INSTALLATION step 2 (harness bundle refresh) so
|
|
28
|
+
the accompanying files stay in lockstep with the skill.
|
|
29
|
+
|
|
30
|
+
## FIRST USE
|
|
31
|
+
Fetch the served copy and follow its INSTALLATION section — it materializes
|
|
32
|
+
the harness (~/browser-screen-recorder-harness/), installs dependencies, and
|
|
33
|
+
asks the user for a GOOGLE_API_KEY (required, Gemini brain) plus optional
|
|
34
|
+
OPENAI/DEEPGRAM keys for the voice mouth and audible verification.
|
|
@@ -58,7 +58,15 @@ curl -sS -X POST http://localhost:8741/canvas -H 'Content-Type: application/json
|
|
|
58
58
|
curl -sS -X POST http://localhost:8741/canvas -H 'Content-Type: application/json' \
|
|
59
59
|
-d '{"kind":"show","mode":"notes","title":"...","items":["...","..."]}'
|
|
60
60
|
```
|
|
61
|
-
`mode` = `idle` | `notes` (title+items) | `link` (url) | `web` (iframe url) | `text` (title+text)
|
|
61
|
+
`mode` = `idle` | `notes` (title+items) | `link` (url) | `web` (iframe url) | `text` (title+text)
|
|
62
|
+
| `stream` (url — renders `<img src="{url}/stream">`, a live MJPEG browser feed).
|
|
63
|
+
|
|
64
|
+
**`stream` mode — NO TUNNELS policy:** the feed URL must be PUBLIC. Never
|
|
65
|
+
tunnel a local browser engine to get one (ngrok's free tier burned its entire
|
|
66
|
+
monthly bandwidth cap on one continuous-MJPEG demo — `ERR_NGROK_725`). Run the
|
|
67
|
+
browser-screen-recorder engine on its Fly machine instead; its `:8080` MJPEG is
|
|
68
|
+
already public at `https://<app>.fly.dev/`. See the browser-screen-recorder
|
|
69
|
+
skill ("Casting the feed into a meeting") for details.
|
|
62
70
|
|
|
63
71
|
**When to speak into the meeting:** By DEFAULT stay silent (observer) for
|
|
64
72
|
`[MEETING — *]:` chunks — take notes, don't interrupt. Speak into the meeting
|
package/dist/index.js
CHANGED
|
@@ -1619,6 +1619,64 @@ async function main() {
|
|
|
1619
1619
|
}
|
|
1620
1620
|
meetingTranscriptBuffer.length = 0;
|
|
1621
1621
|
};
|
|
1622
|
+
// ── Meeting lifecycle (centralized teardown, 0.9.95) ──
|
|
1623
|
+
// The bot's lifecycle FOLLOWS the voice session (deliberate coupling — a
|
|
1624
|
+
// decoupled always-on bot means untracked background agents; revisit only
|
|
1625
|
+
// with a status/tracking UI). endMeeting() is the single teardown path,
|
|
1626
|
+
// fired by: (a) explicit leave_meeting, (b) user disconnect (auto-leave),
|
|
1627
|
+
// (c) Recall reporting a terminal bot status (call_ended/done/fatal —
|
|
1628
|
+
// detected on the poller's 30s tick; fixes stale bot state when the meeting
|
|
1629
|
+
// ends naturally), or (d) the max-duration backstop below.
|
|
1630
|
+
const MEETING_MAX_MS = Math.max(10, Number(process.env.OSBORN_MEETING_MAX_MIN) || 180) * 60 * 1000; // default 3h
|
|
1631
|
+
let meetingMaxTimer = null;
|
|
1632
|
+
const endMeeting = async (reason, opts = {}) => {
|
|
1633
|
+
const botId = activeMeetingBotId;
|
|
1634
|
+
if (!botId)
|
|
1635
|
+
return;
|
|
1636
|
+
console.log(`🏁 Meeting ended (${reason}) — bot ${botId}`);
|
|
1637
|
+
// Stop inputs FIRST so no more chunks land mid-teardown.
|
|
1638
|
+
stopMeetingFlush();
|
|
1639
|
+
if (activeMeetingPoller) {
|
|
1640
|
+
activeMeetingPoller.stop();
|
|
1641
|
+
activeMeetingPoller = null;
|
|
1642
|
+
}
|
|
1643
|
+
if (meetingMaxTimer) {
|
|
1644
|
+
clearTimeout(meetingMaxTimer);
|
|
1645
|
+
meetingMaxTimer = null;
|
|
1646
|
+
}
|
|
1647
|
+
activeMeetingBotId = null;
|
|
1648
|
+
// leaveBot=false when Recall itself reported the meeting over (bot already gone).
|
|
1649
|
+
if (opts.leaveBot !== false) {
|
|
1650
|
+
const recall = getRecallClient();
|
|
1651
|
+
if (recall)
|
|
1652
|
+
await recall.leaveMeeting(botId).catch((e) => console.warn(`⚠️ leaveMeeting failed: ${e?.message}`));
|
|
1653
|
+
}
|
|
1654
|
+
sendToFrontend({ type: 'meeting_left', botId, reason }).catch(() => { });
|
|
1655
|
+
// Orphan cleanup: if no user is connected, the LLM subprocess was kept alive
|
|
1656
|
+
// solely to serve the meeting — release it now and let idle-exit stop the
|
|
1657
|
+
// machine (same billing discipline as 0.9.73, deferred until meeting end).
|
|
1658
|
+
const userPresent = activeRoom && activeRoom.remoteParticipants.size > 0;
|
|
1659
|
+
if (!userPresent && currentLLM) {
|
|
1660
|
+
console.log('🏁 Meeting over with no user connected — releasing LLM + arming idle-exit');
|
|
1661
|
+
killCurrentLLM(`meeting_ended(${reason})_no_user`);
|
|
1662
|
+
currentLLM = null;
|
|
1663
|
+
clearFastBrainSession();
|
|
1664
|
+
clearPipelineFastBrainSession();
|
|
1665
|
+
armIdleExitTimer(`meeting ended (${reason}), no user`);
|
|
1666
|
+
}
|
|
1667
|
+
};
|
|
1668
|
+
const armMeetingMaxTimer = (botId) => {
|
|
1669
|
+
if (meetingMaxTimer)
|
|
1670
|
+
clearTimeout(meetingMaxTimer);
|
|
1671
|
+
console.log(`⏲️ Meeting max-duration backstop armed: ${MEETING_MAX_MS / 60000} min (override OSBORN_MEETING_MAX_MIN)`);
|
|
1672
|
+
meetingMaxTimer = setTimeout(() => {
|
|
1673
|
+
meetingMaxTimer = null;
|
|
1674
|
+
if (activeMeetingBotId === botId) {
|
|
1675
|
+
console.log('⏲️ Meeting max duration reached — auto-leaving (backstop, not Recall-reported end)');
|
|
1676
|
+
void endMeeting('max_duration_backstop');
|
|
1677
|
+
}
|
|
1678
|
+
}, MEETING_MAX_MS);
|
|
1679
|
+
};
|
|
1622
1680
|
// Track the active resume session ID across scopes (ParticipantConnected + DataReceived)
|
|
1623
1681
|
// Updated by resume_session, session_selected, continue_session, switch_session handlers
|
|
1624
1682
|
let currentResumeSessionId;
|
|
@@ -4195,20 +4253,13 @@ async function main() {
|
|
|
4195
4253
|
currentLLM = null;
|
|
4196
4254
|
clearFastBrainSession();
|
|
4197
4255
|
clearPipelineFastBrainSession();
|
|
4198
|
-
// Auto-leave any active meeting bot when user disconnects from the room
|
|
4199
|
-
|
|
4200
|
-
|
|
4201
|
-
|
|
4202
|
-
|
|
4203
|
-
|
|
4204
|
-
|
|
4205
|
-
const recallDisconnect = getRecallClient();
|
|
4206
|
-
if (recallDisconnect) {
|
|
4207
|
-
console.log(`🤝 Auto-leaving meeting (bot ${activeMeetingBotId}) — user disconnected from room`);
|
|
4208
|
-
recallDisconnect.leaveMeeting(activeMeetingBotId).catch(() => { });
|
|
4209
|
-
activeMeetingBotId = null;
|
|
4210
|
-
}
|
|
4211
|
-
}
|
|
4256
|
+
// Auto-leave any active meeting bot when user disconnects from the room.
|
|
4257
|
+
// DELIBERATE COUPLING: the bot's lifecycle follows the voice session — a
|
|
4258
|
+
// decoupled always-on meeting bot means untracked background agents (no
|
|
4259
|
+
// status surface for them yet). Revisit only WITH a tracking UI.
|
|
4260
|
+
// endMeeting() centralizes the teardown (flush, poller, max-timer, Recall
|
|
4261
|
+
// leave, frontend notify).
|
|
4262
|
+
void endMeeting('user_disconnected');
|
|
4212
4263
|
// 0.9.83: a real session just ended → use the FAST leave (~20s), not the
|
|
4213
4264
|
// 3-min alone grace. Runs on the agent, so it fires even on an abrupt tab
|
|
4214
4265
|
// close. Cancelled if a user rejoins within the grace (handled where the
|
|
@@ -4880,11 +4931,17 @@ async function main() {
|
|
|
4880
4931
|
console.warn(`⚠️ Failed to forward meeting transcript to LLM: ${err.message}`);
|
|
4881
4932
|
}
|
|
4882
4933
|
},
|
|
4934
|
+
// Authoritative meeting-over signal: Recall's terminal bot status
|
|
4935
|
+
// (call_ended/done/fatal), checked on the same 30s tick. The bot
|
|
4936
|
+
// is already gone at that point, so don't re-issue leave.
|
|
4937
|
+
onMeetingEnd: (code) => { void endMeeting(`recall_status:${code}`, { leaveBot: false }); },
|
|
4883
4938
|
});
|
|
4884
4939
|
activeMeetingPoller.start();
|
|
4885
4940
|
// LIVE path: buffer webhook finals + flush to the LLM every 20s.
|
|
4886
4941
|
// (The poller above only lands data after the meeting ENDS.)
|
|
4887
4942
|
startMeetingFlush(botId);
|
|
4943
|
+
// Billing backstop: a forgotten meeting can't hold the machine forever.
|
|
4944
|
+
armMeetingMaxTimer(botId);
|
|
4888
4945
|
}
|
|
4889
4946
|
catch (err) {
|
|
4890
4947
|
console.error('❌ Recall.ai join error:', err);
|
|
@@ -4894,25 +4951,14 @@ async function main() {
|
|
|
4894
4951
|
}
|
|
4895
4952
|
}
|
|
4896
4953
|
else if (data.type === 'leave_meeting') {
|
|
4897
|
-
|
|
4898
|
-
|
|
4899
|
-
|
|
4900
|
-
|
|
4901
|
-
|
|
4902
|
-
|
|
4903
|
-
|
|
4904
|
-
|
|
4905
|
-
activeMeetingPoller.stop();
|
|
4906
|
-
activeMeetingPoller = null;
|
|
4907
|
-
}
|
|
4908
|
-
await recallLeave.leaveMeeting(botId);
|
|
4909
|
-
activeMeetingBotId = null;
|
|
4910
|
-
await sendToFrontend({ type: 'meeting_left', botId });
|
|
4911
|
-
}
|
|
4912
|
-
catch (err) {
|
|
4913
|
-
console.error('❌ Recall.ai leave error:', err);
|
|
4914
|
-
await sendToFrontend({ type: 'meeting_error', message: err.message });
|
|
4915
|
-
}
|
|
4954
|
+
// All teardown (flush, poller, max-timer, Recall leave, meeting_left,
|
|
4955
|
+
// orphan LLM release) is centralized in endMeeting().
|
|
4956
|
+
try {
|
|
4957
|
+
await endMeeting('user_leave_meeting');
|
|
4958
|
+
}
|
|
4959
|
+
catch (err) {
|
|
4960
|
+
console.error('❌ Recall.ai leave error:', err);
|
|
4961
|
+
await sendToFrontend({ type: 'meeting_error', message: err.message });
|
|
4916
4962
|
}
|
|
4917
4963
|
}
|
|
4918
4964
|
else if (data.type === 'session_selected') {
|
|
@@ -41,6 +41,15 @@ export interface MeetingTranscriptPollerOptions {
|
|
|
41
41
|
intervalMs?: number;
|
|
42
42
|
/** Optional debug logger. */
|
|
43
43
|
onError?: (err: Error) => void;
|
|
44
|
+
/**
|
|
45
|
+
* Called ONCE when Recall reports the meeting is over (bot status reaches a
|
|
46
|
+
* terminal code: call_ended / done / fatal). Rides the same 30s tick — no
|
|
47
|
+
* extra webhook subscription needed. The meeting-copilot lifecycle uses this
|
|
48
|
+
* as the authoritative "meeting ended" signal now that the bot is DECOUPLED
|
|
49
|
+
* from the voice session (a voice disconnect no longer drops the bot, so
|
|
50
|
+
* something must still end it — this is that something).
|
|
51
|
+
*/
|
|
52
|
+
onMeetingEnd?: (statusCode: string) => void;
|
|
44
53
|
}
|
|
45
54
|
export declare class MeetingTranscriptPoller {
|
|
46
55
|
#private;
|
|
@@ -27,6 +27,8 @@
|
|
|
27
27
|
* No backoff — Recall's transcript endpoint is stable enough that a 30s
|
|
28
28
|
* cadence makes "slow start" non-issues self-recover within one cycle.
|
|
29
29
|
*/
|
|
30
|
+
/** Recall bot status codes that mean the meeting/bot is definitively over. */
|
|
31
|
+
const TERMINAL_STATUS_CODES = new Set(['call_ended', 'done', 'fatal']);
|
|
30
32
|
export class MeetingTranscriptPoller {
|
|
31
33
|
#opts;
|
|
32
34
|
#timer = null;
|
|
@@ -60,25 +62,43 @@ export class MeetingTranscriptPoller {
|
|
|
60
62
|
return;
|
|
61
63
|
this.#inFlight = true;
|
|
62
64
|
try {
|
|
65
|
+
// Meeting-end detection first: if the bot has reached a terminal status,
|
|
66
|
+
// fire onMeetingEnd (once — stop() sets #stopped so later ticks no-op)
|
|
67
|
+
// AFTER one final transcript pull below, so the last turns aren't lost.
|
|
68
|
+
let endedCode = null;
|
|
69
|
+
if (this.#opts.onMeetingEnd) {
|
|
70
|
+
try {
|
|
71
|
+
const code = await this.#opts.recall.getBotStatus(this.#opts.botId);
|
|
72
|
+
if (TERMINAL_STATUS_CODES.has(code))
|
|
73
|
+
endedCode = code;
|
|
74
|
+
}
|
|
75
|
+
catch { /* status probe is best-effort; transcript pull continues */ }
|
|
76
|
+
}
|
|
63
77
|
const all = await this.#opts.recall.getTranscript(this.#opts.botId);
|
|
64
78
|
const fresh = all.filter(t => {
|
|
65
79
|
const firstWordTs = t.words?.[0]?.start_timestamp?.relative;
|
|
66
80
|
return typeof firstWordTs === 'number' && firstWordTs > this.#cursor;
|
|
67
81
|
});
|
|
68
|
-
if (fresh.length
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
82
|
+
if (fresh.length > 0) {
|
|
83
|
+
// Advance cursor to highest seen first-word ts (across all returned turns,
|
|
84
|
+
// not just the fresh ones — guards against Recall returning a paged subset).
|
|
85
|
+
for (const t of all) {
|
|
86
|
+
const ts = t.words?.[0]?.start_timestamp?.relative;
|
|
87
|
+
if (typeof ts === 'number' && ts > this.#cursor)
|
|
88
|
+
this.#cursor = ts;
|
|
89
|
+
}
|
|
90
|
+
const formatted = formatTurns(fresh);
|
|
91
|
+
if (formatted) {
|
|
92
|
+
console.log(`📓 MeetingTranscriptPoller: ${fresh.length} new turn(s), cursor=${this.#cursor.toFixed(1)}s, chars=${formatted.length}`);
|
|
93
|
+
await this.#opts.onTurns({ botId: this.#opts.botId, turns: fresh, formatted });
|
|
94
|
+
}
|
|
95
|
+
}
|
|
96
|
+
// Fire end-of-meeting AFTER the final transcript pull so the last turns land.
|
|
97
|
+
if (endedCode) {
|
|
98
|
+
console.log(`📓 MeetingTranscriptPoller: bot status '${endedCode}' — meeting ended`);
|
|
99
|
+
this.stop();
|
|
100
|
+
this.#opts.onMeetingEnd?.(endedCode);
|
|
76
101
|
}
|
|
77
|
-
const formatted = formatTurns(fresh);
|
|
78
|
-
if (!formatted)
|
|
79
|
-
return; // pure-whitespace fresh batch — skip
|
|
80
|
-
console.log(`📓 MeetingTranscriptPoller: ${fresh.length} new turn(s), cursor=${this.#cursor.toFixed(1)}s, chars=${formatted.length}`);
|
|
81
|
-
await this.#opts.onTurns({ botId: this.#opts.botId, turns: fresh, formatted });
|
|
82
102
|
}
|
|
83
103
|
catch (err) {
|
|
84
104
|
const e = err instanceof Error ? err : new Error(String(err));
|
package/package.json
CHANGED
|
@@ -1,32 +0,0 @@
|
|
|
1
|
-
# Voice-E2E — browser testing with ears, a mouth, hands and a brain
|
|
2
|
-
|
|
3
|
-
## SKILL IDENTITY
|
|
4
|
-
Name: voice-e2e
|
|
5
|
-
Version: 1 (packaged seed — the served copy at the URL below is source of truth)
|
|
6
|
-
Served from: https://www.voice-native.com/api/test-skill
|
|
7
|
-
Landing page: https://www.voice-native.com/test-skill
|
|
8
|
-
|
|
9
|
-
Agent-driven end-to-end testing for voice and web apps in a REAL browser:
|
|
10
|
-
speak into the page's mic (synthesized on the fly), hear and transcribe what
|
|
11
|
-
it plays back, click via a natural-language brain, improvise conversations,
|
|
12
|
-
and get video + audio replays per run. Works on any website.
|
|
13
|
-
|
|
14
|
-
## UPDATE CHECK (notify-first, run before major test/research runs)
|
|
15
|
-
This packaged copy ships with every machine and may lag the served version.
|
|
16
|
-
Check, then NOTIFY the user before updating — never silently force:
|
|
17
|
-
|
|
18
|
-
```bash
|
|
19
|
-
SERVED_V=$(curl -sf --max-time 10 https://www.voice-native.com/api/test-skill | grep -m1 '^Version:' | awk '{print $2}')
|
|
20
|
-
LOCAL_V=$(grep -m1 '^Version:' ~/.claude/skills/voice-e2e/SKILL.md 2>/dev/null | awk '{print $2}')
|
|
21
|
-
[ -n "$SERVED_V" ] && [ "$SERVED_V" != "$LOCAL_V" ] && echo "UPDATE AVAILABLE: voice-e2e v${LOCAL_V:-none} -> v$SERVED_V — tell the user what's new and ask before updating (curl -sS https://www.voice-native.com/api/test-skill -o ~/.claude/skills/voice-e2e/SKILL.md), then refresh the harness bundle per the served INSTALLATION."
|
|
22
|
-
```
|
|
23
|
-
|
|
24
|
-
If the user approves (or has pre-approved auto-updates), update the skill
|
|
25
|
-
file AND re-run the served INSTALLATION step 2 (harness bundle refresh) so
|
|
26
|
-
the accompanying files stay in lockstep with the skill.
|
|
27
|
-
|
|
28
|
-
## FIRST USE
|
|
29
|
-
Fetch the served copy and follow its INSTALLATION section — it materializes
|
|
30
|
-
the harness (~/voice-e2e-harness/), installs dependencies, and asks the user
|
|
31
|
-
for a GOOGLE_API_KEY (required, Gemini brain) plus optional OPENAI/DEEPGRAM
|
|
32
|
-
keys for the voice mouth and audible verification.
|