osborn 0.9.94 → 0.9.96
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/browser-screen-recorder/SKILL.md +41 -0
- package/.claude/skills/meetings/SKILL.md +9 -1
- package/dist/claude-llm.d.ts +20 -0
- package/dist/claude-llm.js +126 -120
- package/dist/index.js +117 -34
- package/dist/meeting-transcript-poller.d.ts +9 -0
- package/dist/meeting-transcript-poller.js +33 -13
- package/package.json +1 -1
- package/.claude/skills/voice-e2e/SKILL.md +0 -32
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
# Browser Screen Recorder — drive any web app (voice, hands + a brain) and record proof
|
|
2
|
+
|
|
3
|
+
## SKILL IDENTITY
|
|
4
|
+
Name: browser-screen-recorder
|
|
5
|
+
Version: 1 (packaged seed — the served copy at the URL below is source of truth)
|
|
6
|
+
Served from: https://www.voice-native.com/api/browser-screen-recorder
|
|
7
|
+
Landing page: https://www.voice-native.com/browser-screen-recorder
|
|
8
|
+
(Formerly "voice-e2e". The old /api/test-skill URL still resolves to this skill.)
|
|
9
|
+
|
|
10
|
+
Give a coding agent a web app + an intent and it drives the app in a REAL
|
|
11
|
+
browser and records PROOF it did — per-action screenshots + video clips, audio
|
|
12
|
+
capture, DevTools diagnostics, and metrics. It can speak into the page's mic
|
|
13
|
+
(synthesized on the fly), hear and transcribe what it plays back, click via a
|
|
14
|
+
natural-language brain, and improvise conversations. Works on any website.
|
|
15
|
+
|
|
16
|
+
USE IT BROADLY — not just for tests: MANDATORY for debugging anything
|
|
17
|
+
web-reachable (real browser + DevTools on camera + /logs, never debug a
|
|
18
|
+
frontend blind) and for verifying deployed web changes (journey + clip, or it
|
|
19
|
+
is a claim, not a verification). Also: reproducing user issues (mobile view),
|
|
20
|
+
the meeting-canvas cast source, site research, voice testing. Every use
|
|
21
|
+
returns media — review it and deliver it to the user, no silent runs.
|
|
22
|
+
|
|
23
|
+
## UPDATE CHECK (notify-first, run before major runs)
|
|
24
|
+
This packaged copy ships with every machine and may lag the served version.
|
|
25
|
+
Check, then NOTIFY the user before updating — never silently force:
|
|
26
|
+
|
|
27
|
+
```bash
|
|
28
|
+
SERVED_V=$(curl -sf --max-time 10 https://www.voice-native.com/api/browser-screen-recorder | grep -m1 '^Version:' | awk '{print $2}')
|
|
29
|
+
LOCAL_V=$(grep -m1 '^Version:' ~/.claude/skills/browser-screen-recorder/SKILL.md 2>/dev/null | awk '{print $2}')
|
|
30
|
+
[ -n "$SERVED_V" ] && [ "$SERVED_V" != "$LOCAL_V" ] && echo "UPDATE AVAILABLE: browser-screen-recorder v${LOCAL_V:-none} -> v$SERVED_V — tell the user what's new and ask before updating (curl -sS https://www.voice-native.com/api/browser-screen-recorder -o ~/.claude/skills/browser-screen-recorder/SKILL.md), then refresh the harness bundle per the served INSTALLATION."
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
If the user approves (or has pre-approved auto-updates), update the skill
|
|
34
|
+
file AND re-run the served INSTALLATION step 2 (harness bundle refresh) so
|
|
35
|
+
the accompanying files stay in lockstep with the skill.
|
|
36
|
+
|
|
37
|
+
## FIRST USE
|
|
38
|
+
Fetch the served copy and follow its INSTALLATION section — it materializes
|
|
39
|
+
the harness (~/browser-screen-recorder-harness/), installs dependencies, and
|
|
40
|
+
asks the user for a GOOGLE_API_KEY (required, Gemini brain) plus optional
|
|
41
|
+
OPENAI/DEEPGRAM keys for the voice mouth and audible verification.
|
|
@@ -58,7 +58,15 @@ curl -sS -X POST http://localhost:8741/canvas -H 'Content-Type: application/json
|
|
|
58
58
|
curl -sS -X POST http://localhost:8741/canvas -H 'Content-Type: application/json' \
|
|
59
59
|
-d '{"kind":"show","mode":"notes","title":"...","items":["...","..."]}'
|
|
60
60
|
```
|
|
61
|
-
`mode` = `idle` | `notes` (title+items) | `link` (url) | `web` (iframe url) | `text` (title+text)
|
|
61
|
+
`mode` = `idle` | `notes` (title+items) | `link` (url) | `web` (iframe url) | `text` (title+text)
|
|
62
|
+
| `stream` (url — renders `<img src="{url}/stream">`, a live MJPEG browser feed).
|
|
63
|
+
|
|
64
|
+
**`stream` mode — NO TUNNELS policy:** the feed URL must be PUBLIC. Never
|
|
65
|
+
tunnel a local browser engine to get one (ngrok's free tier burned its entire
|
|
66
|
+
monthly bandwidth cap on one continuous-MJPEG demo — `ERR_NGROK_725`). Run the
|
|
67
|
+
browser-screen-recorder engine on its Fly machine instead; its `:8080` MJPEG is
|
|
68
|
+
already public at `https://<app>.fly.dev/`. See the browser-screen-recorder
|
|
69
|
+
skill ("Casting the feed into a meeting") for details.
|
|
62
70
|
|
|
63
71
|
**When to speak into the meeting:** By DEFAULT stay silent (observer) for
|
|
64
72
|
`[MEETING — *]:` chunks — take notes, don't interrupt. Speak into the meeting
|
package/dist/claude-llm.d.ts
CHANGED
|
@@ -35,6 +35,26 @@ export interface ClaudeLLMOptions {
|
|
|
35
35
|
trigger?: string;
|
|
36
36
|
}) => void;
|
|
37
37
|
}
|
|
38
|
+
export declare const NAMED_AGENTS: {
|
|
39
|
+
researcher: {
|
|
40
|
+
description: string;
|
|
41
|
+
tools: string[];
|
|
42
|
+
model: string;
|
|
43
|
+
prompt: string;
|
|
44
|
+
};
|
|
45
|
+
reasoner: {
|
|
46
|
+
description: string;
|
|
47
|
+
tools: string[];
|
|
48
|
+
model: string;
|
|
49
|
+
prompt: string;
|
|
50
|
+
};
|
|
51
|
+
writer: {
|
|
52
|
+
description: string;
|
|
53
|
+
tools: string[];
|
|
54
|
+
model: string;
|
|
55
|
+
prompt: string;
|
|
56
|
+
};
|
|
57
|
+
};
|
|
38
58
|
/**
|
|
39
59
|
* Claude LLM - Wraps Claude Agent SDK for LiveKit
|
|
40
60
|
* Research mode: reads anything, writes only to session workspace
|
package/dist/claude-llm.js
CHANGED
|
@@ -121,6 +121,128 @@ function loadAllSkills(_workingDir) {
|
|
|
121
121
|
return `<available-skills>\n${[...skillMap.values()].join('\n\n---\n\n')}\n</available-skills>`;
|
|
122
122
|
}
|
|
123
123
|
// Research mode tools — full research capabilities
|
|
124
|
+
// Named sub-agents — the orchestrator delegates to these specialists. Each has
|
|
125
|
+
// a specific role, model, and tool set. Module-level + exported so the HTTP
|
|
126
|
+
// API (/agents) and the frontend agents manager can list them without
|
|
127
|
+
// duplicating the definitions. The query options reference this same object.
|
|
128
|
+
export const NAMED_AGENTS = {
|
|
129
|
+
researcher: {
|
|
130
|
+
description: [
|
|
131
|
+
'Information gathering agent (Sonnet). Use for: codebase exploration, web research,',
|
|
132
|
+
'finding patterns, reading multiple files, searching for examples.',
|
|
133
|
+
'Returns structured findings — does NOT make decisions or edit files.',
|
|
134
|
+
'Use this for ANY task that needs more than 2 tool calls to gather information.',
|
|
135
|
+
].join(' '),
|
|
136
|
+
tools: ['Read', 'Glob', 'Grep', 'Bash', 'WebSearch', 'WebFetch', 'Task'],
|
|
137
|
+
model: 'sonnet',
|
|
138
|
+
prompt: [
|
|
139
|
+
'You are Osborn\'s research agent. Your job is information gathering — thorough, structured, factual.',
|
|
140
|
+
'',
|
|
141
|
+
'## Your role',
|
|
142
|
+
'Gather information the main agent needs to answer the user\'s question or make a decision.',
|
|
143
|
+
'You are a scout — go find things, read them carefully, and report back.',
|
|
144
|
+
'',
|
|
145
|
+
'## How to work',
|
|
146
|
+
'1. Understand what information is needed and why.',
|
|
147
|
+
'2. Search broadly first (Glob, Grep, WebSearch), then read deeply (Read specific files).',
|
|
148
|
+
'3. For large investigations, use the Task tool to run parallel searches.',
|
|
149
|
+
'4. Cap yourself at 5-8 tool calls unless the task clearly requires more.',
|
|
150
|
+
'',
|
|
151
|
+
'## What to return',
|
|
152
|
+
'Structured findings with specifics:',
|
|
153
|
+
'- File paths and line numbers where you found relevant code',
|
|
154
|
+
'- Exact values, configs, versions — not paraphrases',
|
|
155
|
+
'- Direct quotes from documentation or web sources',
|
|
156
|
+
'- What you looked for but did NOT find (negative results matter)',
|
|
157
|
+
'',
|
|
158
|
+
'## What NOT to do',
|
|
159
|
+
'- Do NOT make recommendations or decisions — just surface facts',
|
|
160
|
+
'- Do NOT edit or write any files',
|
|
161
|
+
'- Do NOT run destructive commands (no rm, no git push, no npm publish)',
|
|
162
|
+
'- If you need clarification, ask the main agent — it will relay to the user if needed',
|
|
163
|
+
].join('\n'),
|
|
164
|
+
},
|
|
165
|
+
reasoner: {
|
|
166
|
+
description: [
|
|
167
|
+
'Deep reasoning agent (Opus). Use for: architecture decisions, complex problem analysis,',
|
|
168
|
+
'tradeoff evaluation, generating implementation plans, understanding hard problems.',
|
|
169
|
+
'Slow but thorough — only use for genuinely complex problems that need careful thought.',
|
|
170
|
+
'Does NOT edit files — returns a clear plan for the writer agent to execute.',
|
|
171
|
+
].join(' '),
|
|
172
|
+
tools: ['Read', 'Glob', 'Grep', 'WebSearch', 'WebFetch'],
|
|
173
|
+
model: 'opus',
|
|
174
|
+
prompt: [
|
|
175
|
+
'You are Osborn\'s reasoning agent. Your job is deep analysis, architectural thinking, and decision-making.',
|
|
176
|
+
'',
|
|
177
|
+
'## Your role',
|
|
178
|
+
'Think hard about complex problems. Consider multiple approaches. Identify risks and edge cases.',
|
|
179
|
+
'Return a clear, opinionated recommendation with reasoning — not just a list of options.',
|
|
180
|
+
'',
|
|
181
|
+
'## How to work',
|
|
182
|
+
'1. Read and understand the full context before forming an opinion.',
|
|
183
|
+
'2. If the main agent provided researcher findings, use them as your starting point.',
|
|
184
|
+
'3. Consider at least 2-3 alternative approaches before recommending one.',
|
|
185
|
+
'4. Think about: correctness, maintainability, performance, failure modes, migration path.',
|
|
186
|
+
'5. Use Read/Grep to verify assumptions against the actual codebase when relevant.',
|
|
187
|
+
'',
|
|
188
|
+
'## What to return',
|
|
189
|
+
'- RECOMMENDATION: what to do (one clear answer, not "it depends")',
|
|
190
|
+
'- REASONING: why this approach wins over alternatives (2-3 sentences)',
|
|
191
|
+
'- PLAN: step-by-step implementation instructions specific enough for the writer agent',
|
|
192
|
+
'- RISKS: what could go wrong and how to mitigate',
|
|
193
|
+
'- If the problem is genuinely ambiguous, say what additional information would resolve it',
|
|
194
|
+
'',
|
|
195
|
+
'## What NOT to do',
|
|
196
|
+
'- Do NOT edit or write files — return a plan for the writer agent',
|
|
197
|
+
'- Do NOT give wishy-washy "both options are valid" non-answers — commit to a recommendation',
|
|
198
|
+
'- If you need more information, ask the main agent to delegate to the researcher',
|
|
199
|
+
].join('\n'),
|
|
200
|
+
},
|
|
201
|
+
writer: {
|
|
202
|
+
description: [
|
|
203
|
+
'Execution agent with file write/edit permissions (Sonnet).',
|
|
204
|
+
'Handles ALL file operations: code, config, docs, scripts, data files.',
|
|
205
|
+
'VERIFY-FIRST workflow: checks assumptions before making changes, runs tests after.',
|
|
206
|
+
'If anything is unclear, asks the main agent for clarification before touching files.',
|
|
207
|
+
].join(' '),
|
|
208
|
+
tools: ['Read', 'Write', 'Edit', 'MultiEdit', 'Bash', 'Glob', 'Grep', 'NotebookRead', 'NotebookEdit'],
|
|
209
|
+
model: 'sonnet',
|
|
210
|
+
prompt: [
|
|
211
|
+
'You are Osborn\'s writer agent. You execute file changes with a verify-first approach.',
|
|
212
|
+
'',
|
|
213
|
+
'## Your role',
|
|
214
|
+
'Handle ALL file operations — code, config, documentation, scripts, data files.',
|
|
215
|
+
'You are the only agent that writes. The main agent and reasoner produce plans; you execute them.',
|
|
216
|
+
'',
|
|
217
|
+
'## VERIFY-FIRST workflow (mandatory)',
|
|
218
|
+
'',
|
|
219
|
+
'### Step 1: Verify assumptions',
|
|
220
|
+
'1. Read the files you\'re about to modify. Confirm they match what the plan expects.',
|
|
221
|
+
'2. If the plan references specific code patterns, grep to confirm they exist.',
|
|
222
|
+
'3. If applicable, run the current test suite or build to confirm the starting state works.',
|
|
223
|
+
'4. If ANYTHING has drifted from the plan (file moved, code refactored, dependency changed):',
|
|
224
|
+
' STOP and report back to the main agent. Do NOT improvise.',
|
|
225
|
+
'',
|
|
226
|
+
'### Step 2: Clarify unknowns',
|
|
227
|
+
'1. If the plan is vague or ambiguous — ask the main agent a specific clarifying question.',
|
|
228
|
+
' Examples: "Which config format — YAML or JSON?", "New file or extend existing auth.ts?"',
|
|
229
|
+
'2. The main agent will answer from context or relay to the user.',
|
|
230
|
+
'3. Do NOT guess. One clear question is better than a wrong assumption.',
|
|
231
|
+
'4. Restate what you will do before doing it: which files, what changes, in what order.',
|
|
232
|
+
'',
|
|
233
|
+
'### Step 3: Execute changes',
|
|
234
|
+
'- Make ONLY the changes described in the plan.',
|
|
235
|
+
'- Do NOT refactor adjacent code, fix unrelated issues, add unrequested comments/docs.',
|
|
236
|
+
'- If you hit an unexpected issue, STOP and report to the main agent.',
|
|
237
|
+
'',
|
|
238
|
+
'### Step 4: Verify results',
|
|
239
|
+
'1. Run tests if available (npm test, pytest, cargo test, etc.).',
|
|
240
|
+
'2. Run the build if applicable (npm run build, tsc --noEmit, etc.).',
|
|
241
|
+
'3. If tests or build fail: attempt to fix the issue you introduced. Re-run.',
|
|
242
|
+
'4. Report: files changed, what changed in each, test results, any failures.',
|
|
243
|
+
].join('\n'),
|
|
244
|
+
},
|
|
245
|
+
};
|
|
124
246
|
const RESEARCH_TOOLS = [
|
|
125
247
|
'Read', 'Write', 'Edit', 'Glob', 'Grep',
|
|
126
248
|
'Bash', 'WebSearch', 'WebFetch',
|
|
@@ -1169,126 +1291,10 @@ class ClaudeLLMStream extends llm.LLMStream {
|
|
|
1169
1291
|
}]
|
|
1170
1292
|
}]
|
|
1171
1293
|
},
|
|
1172
|
-
// Named sub-agents —
|
|
1173
|
-
//
|
|
1174
|
-
agents
|
|
1175
|
-
|
|
1176
|
-
description: [
|
|
1177
|
-
'Information gathering agent (Sonnet). Use for: codebase exploration, web research,',
|
|
1178
|
-
'finding patterns, reading multiple files, searching for examples.',
|
|
1179
|
-
'Returns structured findings — does NOT make decisions or edit files.',
|
|
1180
|
-
'Use this for ANY task that needs more than 2 tool calls to gather information.',
|
|
1181
|
-
].join(' '),
|
|
1182
|
-
tools: ['Read', 'Glob', 'Grep', 'Bash', 'WebSearch', 'WebFetch', 'Task'],
|
|
1183
|
-
model: 'sonnet',
|
|
1184
|
-
prompt: [
|
|
1185
|
-
'You are Osborn\'s research agent. Your job is information gathering — thorough, structured, factual.',
|
|
1186
|
-
'',
|
|
1187
|
-
'## Your role',
|
|
1188
|
-
'Gather information the main agent needs to answer the user\'s question or make a decision.',
|
|
1189
|
-
'You are a scout — go find things, read them carefully, and report back.',
|
|
1190
|
-
'',
|
|
1191
|
-
'## How to work',
|
|
1192
|
-
'1. Understand what information is needed and why.',
|
|
1193
|
-
'2. Search broadly first (Glob, Grep, WebSearch), then read deeply (Read specific files).',
|
|
1194
|
-
'3. For large investigations, use the Task tool to run parallel searches.',
|
|
1195
|
-
'4. Cap yourself at 5-8 tool calls unless the task clearly requires more.',
|
|
1196
|
-
'',
|
|
1197
|
-
'## What to return',
|
|
1198
|
-
'Structured findings with specifics:',
|
|
1199
|
-
'- File paths and line numbers where you found relevant code',
|
|
1200
|
-
'- Exact values, configs, versions — not paraphrases',
|
|
1201
|
-
'- Direct quotes from documentation or web sources',
|
|
1202
|
-
'- What you looked for but did NOT find (negative results matter)',
|
|
1203
|
-
'',
|
|
1204
|
-
'## What NOT to do',
|
|
1205
|
-
'- Do NOT make recommendations or decisions — just surface facts',
|
|
1206
|
-
'- Do NOT edit or write any files',
|
|
1207
|
-
'- Do NOT run destructive commands (no rm, no git push, no npm publish)',
|
|
1208
|
-
'- If you need clarification, ask the main agent — it will relay to the user if needed',
|
|
1209
|
-
].join('\n'),
|
|
1210
|
-
},
|
|
1211
|
-
reasoner: {
|
|
1212
|
-
description: [
|
|
1213
|
-
'Deep reasoning agent (Opus). Use for: architecture decisions, complex problem analysis,',
|
|
1214
|
-
'tradeoff evaluation, generating implementation plans, understanding hard problems.',
|
|
1215
|
-
'Slow but thorough — only use for genuinely complex problems that need careful thought.',
|
|
1216
|
-
'Does NOT edit files — returns a clear plan for the writer agent to execute.',
|
|
1217
|
-
].join(' '),
|
|
1218
|
-
tools: ['Read', 'Glob', 'Grep', 'WebSearch', 'WebFetch'],
|
|
1219
|
-
model: 'opus',
|
|
1220
|
-
prompt: [
|
|
1221
|
-
'You are Osborn\'s reasoning agent. Your job is deep analysis, architectural thinking, and decision-making.',
|
|
1222
|
-
'',
|
|
1223
|
-
'## Your role',
|
|
1224
|
-
'Think hard about complex problems. Consider multiple approaches. Identify risks and edge cases.',
|
|
1225
|
-
'Return a clear, opinionated recommendation with reasoning — not just a list of options.',
|
|
1226
|
-
'',
|
|
1227
|
-
'## How to work',
|
|
1228
|
-
'1. Read and understand the full context before forming an opinion.',
|
|
1229
|
-
'2. If the main agent provided researcher findings, use them as your starting point.',
|
|
1230
|
-
'3. Consider at least 2-3 alternative approaches before recommending one.',
|
|
1231
|
-
'4. Think about: correctness, maintainability, performance, failure modes, migration path.',
|
|
1232
|
-
'5. Use Read/Grep to verify assumptions against the actual codebase when relevant.',
|
|
1233
|
-
'',
|
|
1234
|
-
'## What to return',
|
|
1235
|
-
'- RECOMMENDATION: what to do (one clear answer, not "it depends")',
|
|
1236
|
-
'- REASONING: why this approach wins over alternatives (2-3 sentences)',
|
|
1237
|
-
'- PLAN: step-by-step implementation instructions specific enough for the writer agent',
|
|
1238
|
-
'- RISKS: what could go wrong and how to mitigate',
|
|
1239
|
-
'- If the problem is genuinely ambiguous, say what additional information would resolve it',
|
|
1240
|
-
'',
|
|
1241
|
-
'## What NOT to do',
|
|
1242
|
-
'- Do NOT edit or write files — return a plan for the writer agent',
|
|
1243
|
-
'- Do NOT give wishy-washy "both options are valid" non-answers — commit to a recommendation',
|
|
1244
|
-
'- If you need more information, ask the main agent to delegate to the researcher',
|
|
1245
|
-
].join('\n'),
|
|
1246
|
-
},
|
|
1247
|
-
writer: {
|
|
1248
|
-
description: [
|
|
1249
|
-
'Execution agent with file write/edit permissions (Sonnet).',
|
|
1250
|
-
'Handles ALL file operations: code, config, docs, scripts, data files.',
|
|
1251
|
-
'VERIFY-FIRST workflow: checks assumptions before making changes, runs tests after.',
|
|
1252
|
-
'If anything is unclear, asks the main agent for clarification before touching files.',
|
|
1253
|
-
].join(' '),
|
|
1254
|
-
tools: ['Read', 'Write', 'Edit', 'MultiEdit', 'Bash', 'Glob', 'Grep', 'NotebookRead', 'NotebookEdit'],
|
|
1255
|
-
model: 'sonnet',
|
|
1256
|
-
prompt: [
|
|
1257
|
-
'You are Osborn\'s writer agent. You execute file changes with a verify-first approach.',
|
|
1258
|
-
'',
|
|
1259
|
-
'## Your role',
|
|
1260
|
-
'Handle ALL file operations — code, config, documentation, scripts, data files.',
|
|
1261
|
-
'You are the only agent that writes. The main agent and reasoner produce plans; you execute them.',
|
|
1262
|
-
'',
|
|
1263
|
-
'## VERIFY-FIRST workflow (mandatory)',
|
|
1264
|
-
'',
|
|
1265
|
-
'### Step 1: Verify assumptions',
|
|
1266
|
-
'1. Read the files you\'re about to modify. Confirm they match what the plan expects.',
|
|
1267
|
-
'2. If the plan references specific code patterns, grep to confirm they exist.',
|
|
1268
|
-
'3. If applicable, run the current test suite or build to confirm the starting state works.',
|
|
1269
|
-
'4. If ANYTHING has drifted from the plan (file moved, code refactored, dependency changed):',
|
|
1270
|
-
' STOP and report back to the main agent. Do NOT improvise.',
|
|
1271
|
-
'',
|
|
1272
|
-
'### Step 2: Clarify unknowns',
|
|
1273
|
-
'1. If the plan is vague or ambiguous — ask the main agent a specific clarifying question.',
|
|
1274
|
-
' Examples: "Which config format — YAML or JSON?", "New file or extend existing auth.ts?"',
|
|
1275
|
-
'2. The main agent will answer from context or relay to the user.',
|
|
1276
|
-
'3. Do NOT guess. One clear question is better than a wrong assumption.',
|
|
1277
|
-
'4. Restate what you will do before doing it: which files, what changes, in what order.',
|
|
1278
|
-
'',
|
|
1279
|
-
'### Step 3: Execute changes',
|
|
1280
|
-
'- Make ONLY the changes described in the plan.',
|
|
1281
|
-
'- Do NOT refactor adjacent code, fix unrelated issues, add unrequested comments/docs.',
|
|
1282
|
-
'- If you hit an unexpected issue, STOP and report to the main agent.',
|
|
1283
|
-
'',
|
|
1284
|
-
'### Step 4: Verify results',
|
|
1285
|
-
'1. Run tests if available (npm test, pytest, cargo test, etc.).',
|
|
1286
|
-
'2. Run the build if applicable (npm run build, tsc --noEmit, etc.).',
|
|
1287
|
-
'3. If tests or build fail: attempt to fix the issue you introduced. Re-run.',
|
|
1288
|
-
'4. Report: files changed, what changed in each, test results, any failures.',
|
|
1289
|
-
].join('\n'),
|
|
1290
|
-
},
|
|
1291
|
-
},
|
|
1294
|
+
// Named sub-agents — the orchestrator delegates to these specialists.
|
|
1295
|
+
// Definitions live at module level (NAMED_AGENTS) so the HTTP API and
|
|
1296
|
+
// frontend agents manager can list them. See top of file.
|
|
1297
|
+
agents: NAMED_AGENTS,
|
|
1292
1298
|
};
|
|
1293
1299
|
// Run Claude Agent SDK query() and stream results
|
|
1294
1300
|
let hasOutput = false;
|
package/dist/index.js
CHANGED
|
@@ -29,7 +29,7 @@ const __dirname = dirname(__filename);
|
|
|
29
29
|
import { createPatch } from 'diff';
|
|
30
30
|
import { loadConfig, getMcpServers, getEnabledMcpServerNames, getVoiceMode, getRealtimeConfig, getDirectConfig, listAllClaudeSessions, getMostRecentSessionId, sessionExists, getSessionSummary, getConversationHistory, ensureSessionWorkspace, getSessionWorkspace, getMcpServerStatusList, buildMcpServersForKeys, listWorkspaceArtifacts } from './config.js';
|
|
31
31
|
import { createSTT, createTTS, createRealtimeModelFromConfig, DIRECT_MODE_STT, DIRECT_MODE_TTS } from './voice-io.js';
|
|
32
|
-
import { createClaudeLLM } from './claude-llm.js';
|
|
32
|
+
import { createClaudeLLM, NAMED_AGENTS } from './claude-llm.js';
|
|
33
33
|
import { clearPipelineFastBrainSession, prewarmBM25Index } from './pipeline-fastbrain.js';
|
|
34
34
|
import { ensureClaudeAuth } from './claude-auth.js';
|
|
35
35
|
import { createSmitheryProxy, destroySmitheryProxy, parseSmitheryUrl, isSmitheryUrl, SmitheryAuthorizationError } from './smithery-proxy.js';
|
|
@@ -286,6 +286,30 @@ function startApiServer(workingDir, port) {
|
|
|
286
286
|
}
|
|
287
287
|
return;
|
|
288
288
|
}
|
|
289
|
+
if (req.method === 'GET' && url.pathname === '/skills') {
|
|
290
|
+
// Installed skills — same list the chat's get_skills data-channel message
|
|
291
|
+
// returns, exposed over HTTP so the DASHBOARD (no LiveKit connection) can
|
|
292
|
+
// render the skills manager too. process.cwd() === sessionBaseDir (the
|
|
293
|
+
// osborn install dir where .claude/skills lives — see main()).
|
|
294
|
+
res.writeHead(200, { 'Content-Type': 'application/json' });
|
|
295
|
+
res.end(JSON.stringify({ skills: loadSkillsList(process.cwd()) }));
|
|
296
|
+
return;
|
|
297
|
+
}
|
|
298
|
+
if (req.method === 'GET' && url.pathname === '/agents') {
|
|
299
|
+
// Named sub-agents (researcher/reasoner/writer) — definitions come from
|
|
300
|
+
// claude-llm.ts NAMED_AGENTS (single source of truth, same object the SDK
|
|
301
|
+
// query uses). Prompts are omitted: the UI manager needs role/model/tools,
|
|
302
|
+
// not the full instruction text.
|
|
303
|
+
const agents = Object.entries(NAMED_AGENTS).map(([name, a]) => ({
|
|
304
|
+
name,
|
|
305
|
+
description: a.description,
|
|
306
|
+
model: a.model,
|
|
307
|
+
tools: a.tools,
|
|
308
|
+
}));
|
|
309
|
+
res.writeHead(200, { 'Content-Type': 'application/json', 'Access-Control-Allow-Origin': '*' });
|
|
310
|
+
res.end(JSON.stringify({ agents }));
|
|
311
|
+
return;
|
|
312
|
+
}
|
|
289
313
|
if (req.method === 'GET' && url.pathname === '/health') {
|
|
290
314
|
// Include osborn version — primary signal used by machines.readInstalledOsbornVersion()
|
|
291
315
|
// to detect which agent version is running. Without this the consumer falls back to
|
|
@@ -1619,6 +1643,64 @@ async function main() {
|
|
|
1619
1643
|
}
|
|
1620
1644
|
meetingTranscriptBuffer.length = 0;
|
|
1621
1645
|
};
|
|
1646
|
+
// ── Meeting lifecycle (centralized teardown, 0.9.95) ──
|
|
1647
|
+
// The bot's lifecycle FOLLOWS the voice session (deliberate coupling — a
|
|
1648
|
+
// decoupled always-on bot means untracked background agents; revisit only
|
|
1649
|
+
// with a status/tracking UI). endMeeting() is the single teardown path,
|
|
1650
|
+
// fired by: (a) explicit leave_meeting, (b) user disconnect (auto-leave),
|
|
1651
|
+
// (c) Recall reporting a terminal bot status (call_ended/done/fatal —
|
|
1652
|
+
// detected on the poller's 30s tick; fixes stale bot state when the meeting
|
|
1653
|
+
// ends naturally), or (d) the max-duration backstop below.
|
|
1654
|
+
const MEETING_MAX_MS = Math.max(10, Number(process.env.OSBORN_MEETING_MAX_MIN) || 180) * 60 * 1000; // default 3h
|
|
1655
|
+
let meetingMaxTimer = null;
|
|
1656
|
+
const endMeeting = async (reason, opts = {}) => {
|
|
1657
|
+
const botId = activeMeetingBotId;
|
|
1658
|
+
if (!botId)
|
|
1659
|
+
return;
|
|
1660
|
+
console.log(`🏁 Meeting ended (${reason}) — bot ${botId}`);
|
|
1661
|
+
// Stop inputs FIRST so no more chunks land mid-teardown.
|
|
1662
|
+
stopMeetingFlush();
|
|
1663
|
+
if (activeMeetingPoller) {
|
|
1664
|
+
activeMeetingPoller.stop();
|
|
1665
|
+
activeMeetingPoller = null;
|
|
1666
|
+
}
|
|
1667
|
+
if (meetingMaxTimer) {
|
|
1668
|
+
clearTimeout(meetingMaxTimer);
|
|
1669
|
+
meetingMaxTimer = null;
|
|
1670
|
+
}
|
|
1671
|
+
activeMeetingBotId = null;
|
|
1672
|
+
// leaveBot=false when Recall itself reported the meeting over (bot already gone).
|
|
1673
|
+
if (opts.leaveBot !== false) {
|
|
1674
|
+
const recall = getRecallClient();
|
|
1675
|
+
if (recall)
|
|
1676
|
+
await recall.leaveMeeting(botId).catch((e) => console.warn(`⚠️ leaveMeeting failed: ${e?.message}`));
|
|
1677
|
+
}
|
|
1678
|
+
sendToFrontend({ type: 'meeting_left', botId, reason }).catch(() => { });
|
|
1679
|
+
// Orphan cleanup: if no user is connected, the LLM subprocess was kept alive
|
|
1680
|
+
// solely to serve the meeting — release it now and let idle-exit stop the
|
|
1681
|
+
// machine (same billing discipline as 0.9.73, deferred until meeting end).
|
|
1682
|
+
const userPresent = activeRoom && activeRoom.remoteParticipants.size > 0;
|
|
1683
|
+
if (!userPresent && currentLLM) {
|
|
1684
|
+
console.log('🏁 Meeting over with no user connected — releasing LLM + arming idle-exit');
|
|
1685
|
+
killCurrentLLM(`meeting_ended(${reason})_no_user`);
|
|
1686
|
+
currentLLM = null;
|
|
1687
|
+
clearFastBrainSession();
|
|
1688
|
+
clearPipelineFastBrainSession();
|
|
1689
|
+
armIdleExitTimer(`meeting ended (${reason}), no user`);
|
|
1690
|
+
}
|
|
1691
|
+
};
|
|
1692
|
+
const armMeetingMaxTimer = (botId) => {
|
|
1693
|
+
if (meetingMaxTimer)
|
|
1694
|
+
clearTimeout(meetingMaxTimer);
|
|
1695
|
+
console.log(`⏲️ Meeting max-duration backstop armed: ${MEETING_MAX_MS / 60000} min (override OSBORN_MEETING_MAX_MIN)`);
|
|
1696
|
+
meetingMaxTimer = setTimeout(() => {
|
|
1697
|
+
meetingMaxTimer = null;
|
|
1698
|
+
if (activeMeetingBotId === botId) {
|
|
1699
|
+
console.log('⏲️ Meeting max duration reached — auto-leaving (backstop, not Recall-reported end)');
|
|
1700
|
+
void endMeeting('max_duration_backstop');
|
|
1701
|
+
}
|
|
1702
|
+
}, MEETING_MAX_MS);
|
|
1703
|
+
};
|
|
1622
1704
|
// Track the active resume session ID across scopes (ParticipantConnected + DataReceived)
|
|
1623
1705
|
// Updated by resume_session, session_selected, continue_session, switch_session handlers
|
|
1624
1706
|
let currentResumeSessionId;
|
|
@@ -4060,6 +4142,9 @@ async function main() {
|
|
|
4060
4142
|
enabledMcpServers: enabledMcpNames,
|
|
4061
4143
|
workingDirectory: workingDir,
|
|
4062
4144
|
skills: loadSkillsList(sessionBaseDir),
|
|
4145
|
+
namedAgents: Object.entries(NAMED_AGENTS).map(([name, a]) => ({
|
|
4146
|
+
name, description: a.description, model: a.model, tools: a.tools,
|
|
4147
|
+
})),
|
|
4063
4148
|
});
|
|
4064
4149
|
};
|
|
4065
4150
|
const readyInterval = setInterval(sendReady, 2000);
|
|
@@ -4195,20 +4280,13 @@ async function main() {
|
|
|
4195
4280
|
currentLLM = null;
|
|
4196
4281
|
clearFastBrainSession();
|
|
4197
4282
|
clearPipelineFastBrainSession();
|
|
4198
|
-
// Auto-leave any active meeting bot when user disconnects from the room
|
|
4199
|
-
|
|
4200
|
-
|
|
4201
|
-
|
|
4202
|
-
|
|
4203
|
-
|
|
4204
|
-
|
|
4205
|
-
const recallDisconnect = getRecallClient();
|
|
4206
|
-
if (recallDisconnect) {
|
|
4207
|
-
console.log(`🤝 Auto-leaving meeting (bot ${activeMeetingBotId}) — user disconnected from room`);
|
|
4208
|
-
recallDisconnect.leaveMeeting(activeMeetingBotId).catch(() => { });
|
|
4209
|
-
activeMeetingBotId = null;
|
|
4210
|
-
}
|
|
4211
|
-
}
|
|
4283
|
+
// Auto-leave any active meeting bot when user disconnects from the room.
|
|
4284
|
+
// DELIBERATE COUPLING: the bot's lifecycle follows the voice session — a
|
|
4285
|
+
// decoupled always-on meeting bot means untracked background agents (no
|
|
4286
|
+
// status surface for them yet). Revisit only WITH a tracking UI.
|
|
4287
|
+
// endMeeting() centralizes the teardown (flush, poller, max-timer, Recall
|
|
4288
|
+
// leave, frontend notify).
|
|
4289
|
+
void endMeeting('user_disconnected');
|
|
4212
4290
|
// 0.9.83: a real session just ended → use the FAST leave (~20s), not the
|
|
4213
4291
|
// 3-min alone grace. Runs on the agent, so it fires even on an abrupt tab
|
|
4214
4292
|
// close. Cancelled if a user rejoins within the grace (handled where the
|
|
@@ -4755,6 +4833,16 @@ async function main() {
|
|
|
4755
4833
|
skills: loadSkillsList(sessionBaseDir),
|
|
4756
4834
|
});
|
|
4757
4835
|
}
|
|
4836
|
+
else if (data.type === 'get_agents') {
|
|
4837
|
+
// Named sub-agents for the chat agents manager — same NAMED_AGENTS
|
|
4838
|
+
// source as GET /agents (prompts omitted).
|
|
4839
|
+
await sendToFrontend({
|
|
4840
|
+
type: 'agents_status',
|
|
4841
|
+
agents: Object.entries(NAMED_AGENTS).map(([name, a]) => ({
|
|
4842
|
+
name, description: a.description, model: a.model, tools: a.tools,
|
|
4843
|
+
})),
|
|
4844
|
+
});
|
|
4845
|
+
}
|
|
4758
4846
|
else if (data.type === 'skill_add') {
|
|
4759
4847
|
const skillName = (data.name || '').trim().toLowerCase().replace(/[^a-z0-9-]/g, '-');
|
|
4760
4848
|
const skillContent = (data.content || '').trim();
|
|
@@ -4880,11 +4968,17 @@ async function main() {
|
|
|
4880
4968
|
console.warn(`⚠️ Failed to forward meeting transcript to LLM: ${err.message}`);
|
|
4881
4969
|
}
|
|
4882
4970
|
},
|
|
4971
|
+
// Authoritative meeting-over signal: Recall's terminal bot status
|
|
4972
|
+
// (call_ended/done/fatal), checked on the same 30s tick. The bot
|
|
4973
|
+
// is already gone at that point, so don't re-issue leave.
|
|
4974
|
+
onMeetingEnd: (code) => { void endMeeting(`recall_status:${code}`, { leaveBot: false }); },
|
|
4883
4975
|
});
|
|
4884
4976
|
activeMeetingPoller.start();
|
|
4885
4977
|
// LIVE path: buffer webhook finals + flush to the LLM every 20s.
|
|
4886
4978
|
// (The poller above only lands data after the meeting ENDS.)
|
|
4887
4979
|
startMeetingFlush(botId);
|
|
4980
|
+
// Billing backstop: a forgotten meeting can't hold the machine forever.
|
|
4981
|
+
armMeetingMaxTimer(botId);
|
|
4888
4982
|
}
|
|
4889
4983
|
catch (err) {
|
|
4890
4984
|
console.error('❌ Recall.ai join error:', err);
|
|
@@ -4894,25 +4988,14 @@ async function main() {
|
|
|
4894
4988
|
}
|
|
4895
4989
|
}
|
|
4896
4990
|
else if (data.type === 'leave_meeting') {
|
|
4897
|
-
|
|
4898
|
-
|
|
4899
|
-
|
|
4900
|
-
|
|
4901
|
-
|
|
4902
|
-
|
|
4903
|
-
|
|
4904
|
-
|
|
4905
|
-
activeMeetingPoller.stop();
|
|
4906
|
-
activeMeetingPoller = null;
|
|
4907
|
-
}
|
|
4908
|
-
await recallLeave.leaveMeeting(botId);
|
|
4909
|
-
activeMeetingBotId = null;
|
|
4910
|
-
await sendToFrontend({ type: 'meeting_left', botId });
|
|
4911
|
-
}
|
|
4912
|
-
catch (err) {
|
|
4913
|
-
console.error('❌ Recall.ai leave error:', err);
|
|
4914
|
-
await sendToFrontend({ type: 'meeting_error', message: err.message });
|
|
4915
|
-
}
|
|
4991
|
+
// All teardown (flush, poller, max-timer, Recall leave, meeting_left,
|
|
4992
|
+
// orphan LLM release) is centralized in endMeeting().
|
|
4993
|
+
try {
|
|
4994
|
+
await endMeeting('user_leave_meeting');
|
|
4995
|
+
}
|
|
4996
|
+
catch (err) {
|
|
4997
|
+
console.error('❌ Recall.ai leave error:', err);
|
|
4998
|
+
await sendToFrontend({ type: 'meeting_error', message: err.message });
|
|
4916
4999
|
}
|
|
4917
5000
|
}
|
|
4918
5001
|
else if (data.type === 'session_selected') {
|
|
@@ -41,6 +41,15 @@ export interface MeetingTranscriptPollerOptions {
|
|
|
41
41
|
intervalMs?: number;
|
|
42
42
|
/** Optional debug logger. */
|
|
43
43
|
onError?: (err: Error) => void;
|
|
44
|
+
/**
|
|
45
|
+
* Called ONCE when Recall reports the meeting is over (bot status reaches a
|
|
46
|
+
* terminal code: call_ended / done / fatal). Rides the same 30s tick — no
|
|
47
|
+
* extra webhook subscription needed. The meeting-copilot lifecycle uses this
|
|
48
|
+
* as the authoritative "meeting ended" signal now that the bot is DECOUPLED
|
|
49
|
+
* from the voice session (a voice disconnect no longer drops the bot, so
|
|
50
|
+
* something must still end it — this is that something).
|
|
51
|
+
*/
|
|
52
|
+
onMeetingEnd?: (statusCode: string) => void;
|
|
44
53
|
}
|
|
45
54
|
export declare class MeetingTranscriptPoller {
|
|
46
55
|
#private;
|
|
@@ -27,6 +27,8 @@
|
|
|
27
27
|
* No backoff — Recall's transcript endpoint is stable enough that a 30s
|
|
28
28
|
* cadence makes "slow start" non-issues self-recover within one cycle.
|
|
29
29
|
*/
|
|
30
|
+
/** Recall bot status codes that mean the meeting/bot is definitively over. */
|
|
31
|
+
const TERMINAL_STATUS_CODES = new Set(['call_ended', 'done', 'fatal']);
|
|
30
32
|
export class MeetingTranscriptPoller {
|
|
31
33
|
#opts;
|
|
32
34
|
#timer = null;
|
|
@@ -60,25 +62,43 @@ export class MeetingTranscriptPoller {
|
|
|
60
62
|
return;
|
|
61
63
|
this.#inFlight = true;
|
|
62
64
|
try {
|
|
65
|
+
// Meeting-end detection first: if the bot has reached a terminal status,
|
|
66
|
+
// fire onMeetingEnd (once — stop() sets #stopped so later ticks no-op)
|
|
67
|
+
// AFTER one final transcript pull below, so the last turns aren't lost.
|
|
68
|
+
let endedCode = null;
|
|
69
|
+
if (this.#opts.onMeetingEnd) {
|
|
70
|
+
try {
|
|
71
|
+
const code = await this.#opts.recall.getBotStatus(this.#opts.botId);
|
|
72
|
+
if (TERMINAL_STATUS_CODES.has(code))
|
|
73
|
+
endedCode = code;
|
|
74
|
+
}
|
|
75
|
+
catch { /* status probe is best-effort; transcript pull continues */ }
|
|
76
|
+
}
|
|
63
77
|
const all = await this.#opts.recall.getTranscript(this.#opts.botId);
|
|
64
78
|
const fresh = all.filter(t => {
|
|
65
79
|
const firstWordTs = t.words?.[0]?.start_timestamp?.relative;
|
|
66
80
|
return typeof firstWordTs === 'number' && firstWordTs > this.#cursor;
|
|
67
81
|
});
|
|
68
|
-
if (fresh.length
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
82
|
+
if (fresh.length > 0) {
|
|
83
|
+
// Advance cursor to highest seen first-word ts (across all returned turns,
|
|
84
|
+
// not just the fresh ones — guards against Recall returning a paged subset).
|
|
85
|
+
for (const t of all) {
|
|
86
|
+
const ts = t.words?.[0]?.start_timestamp?.relative;
|
|
87
|
+
if (typeof ts === 'number' && ts > this.#cursor)
|
|
88
|
+
this.#cursor = ts;
|
|
89
|
+
}
|
|
90
|
+
const formatted = formatTurns(fresh);
|
|
91
|
+
if (formatted) {
|
|
92
|
+
console.log(`📓 MeetingTranscriptPoller: ${fresh.length} new turn(s), cursor=${this.#cursor.toFixed(1)}s, chars=${formatted.length}`);
|
|
93
|
+
await this.#opts.onTurns({ botId: this.#opts.botId, turns: fresh, formatted });
|
|
94
|
+
}
|
|
95
|
+
}
|
|
96
|
+
// Fire end-of-meeting AFTER the final transcript pull so the last turns land.
|
|
97
|
+
if (endedCode) {
|
|
98
|
+
console.log(`📓 MeetingTranscriptPoller: bot status '${endedCode}' — meeting ended`);
|
|
99
|
+
this.stop();
|
|
100
|
+
this.#opts.onMeetingEnd?.(endedCode);
|
|
76
101
|
}
|
|
77
|
-
const formatted = formatTurns(fresh);
|
|
78
|
-
if (!formatted)
|
|
79
|
-
return; // pure-whitespace fresh batch — skip
|
|
80
|
-
console.log(`📓 MeetingTranscriptPoller: ${fresh.length} new turn(s), cursor=${this.#cursor.toFixed(1)}s, chars=${formatted.length}`);
|
|
81
|
-
await this.#opts.onTurns({ botId: this.#opts.botId, turns: fresh, formatted });
|
|
82
102
|
}
|
|
83
103
|
catch (err) {
|
|
84
104
|
const e = err instanceof Error ? err : new Error(String(err));
|
package/package.json
CHANGED
|
@@ -1,32 +0,0 @@
|
|
|
1
|
-
# Voice-E2E — browser testing with ears, a mouth, hands and a brain
|
|
2
|
-
|
|
3
|
-
## SKILL IDENTITY
|
|
4
|
-
Name: voice-e2e
|
|
5
|
-
Version: 1 (packaged seed — the served copy at the URL below is source of truth)
|
|
6
|
-
Served from: https://www.voice-native.com/api/test-skill
|
|
7
|
-
Landing page: https://www.voice-native.com/test-skill
|
|
8
|
-
|
|
9
|
-
Agent-driven end-to-end testing for voice and web apps in a REAL browser:
|
|
10
|
-
speak into the page's mic (synthesized on the fly), hear and transcribe what
|
|
11
|
-
it plays back, click via a natural-language brain, improvise conversations,
|
|
12
|
-
and get video + audio replays per run. Works on any website.
|
|
13
|
-
|
|
14
|
-
## UPDATE CHECK (notify-first, run before major test/research runs)
|
|
15
|
-
This packaged copy ships with every machine and may lag the served version.
|
|
16
|
-
Check, then NOTIFY the user before updating — never silently force:
|
|
17
|
-
|
|
18
|
-
```bash
|
|
19
|
-
SERVED_V=$(curl -sf --max-time 10 https://www.voice-native.com/api/test-skill | grep -m1 '^Version:' | awk '{print $2}')
|
|
20
|
-
LOCAL_V=$(grep -m1 '^Version:' ~/.claude/skills/voice-e2e/SKILL.md 2>/dev/null | awk '{print $2}')
|
|
21
|
-
[ -n "$SERVED_V" ] && [ "$SERVED_V" != "$LOCAL_V" ] && echo "UPDATE AVAILABLE: voice-e2e v${LOCAL_V:-none} -> v$SERVED_V — tell the user what's new and ask before updating (curl -sS https://www.voice-native.com/api/test-skill -o ~/.claude/skills/voice-e2e/SKILL.md), then refresh the harness bundle per the served INSTALLATION."
|
|
22
|
-
```
|
|
23
|
-
|
|
24
|
-
If the user approves (or has pre-approved auto-updates), update the skill
|
|
25
|
-
file AND re-run the served INSTALLATION step 2 (harness bundle refresh) so
|
|
26
|
-
the accompanying files stay in lockstep with the skill.
|
|
27
|
-
|
|
28
|
-
## FIRST USE
|
|
29
|
-
Fetch the served copy and follow its INSTALLATION section — it materializes
|
|
30
|
-
the harness (~/voice-e2e-harness/), installs dependencies, and asks the user
|
|
31
|
-
for a GOOGLE_API_KEY (required, Gemini brain) plus optional OPENAI/DEEPGRAM
|
|
32
|
-
keys for the voice mouth and audible verification.
|