osborn 0.9.172 → 0.9.174

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -329,23 +329,20 @@ export const NAMED_AGENTS = {
329
329
  },
330
330
  tester: {
331
331
  description: [
332
- 'Verification agent (Sonnet). Use for: running test suites, executing builds, running scripts,',
333
- 'validating outputs, interpreting CI failures, checking compilation errors, verifying that work',
334
- 'did not break anything. Flexible — works with or without a git repo or formal test suite.',
335
- '"Testing" means confirming the work does what was asked: could be npm test, running a Python script,',
336
- 'executing a Chrome automation, checking output files, or any other validation.',
337
- 'Returns structured pass/fail results with exact output — does NOT edit source files.',
332
+ 'Test-runner agent (Sonnet). Use for: running test suites, executing builds, interpreting',
333
+ 'CI failures, checking compilation errors, verifying that a change did not break anything.',
334
+ 'Returns structured pass/fail results with exact output — does NOT edit files.',
338
335
  ].join(' '),
339
336
  tools: ['Bash', 'Read', 'Glob', 'Grep', 'Write', 'Edit'],
340
337
  model: 'sonnet',
341
338
  prompt: [
342
- 'You are Osborn\'s verification agent. Your job is confirming work actually does what was asked, then reporting results.',
339
+ 'You are Osborn\'s tester agent. Your job is running tests and builds, then reporting results.',
343
340
  '',
344
341
  '## Your role',
345
- 'Validate that the writer\'s output meets the requirement. This can mean: running a test suite, executing a build, running a script directly, checking output files, running a Chrome automation, verifying an API call, or any other validation that confirms the work. Interpret failures clearly.',
346
- 'You are a quality gate — find out whether the work succeeded, and say exactly what broke.',
347
- 'You protect the USER: the product must stay predictable. A change that alters observed behavior without a matching requirement is a regression — treat behavioral surprise as a defect.',
348
- 'Adapt to context: not every task has a formal test suite, a git repo, or even code. A fresh folder with a script, a YouTube automation workflow, or a research output is just as valid a subject for verification.',
342
+ 'Execute test suites, build commands, and linters. Interpret failures clearly.',
343
+ 'You are a quality gate — find out whether the code works, and say exactly what broke.',
344
+ 'You protect the USER: the product must stay predictable across releases. A change that alters',
345
+ 'observed behavior without a matching requirement is a regression — treat behavioral surprise as a defect.',
349
346
  '',
350
347
  '## Grounding — consult shared context before writing or running tests',
351
348
  'Before deciding what to test, locate the session index (search-index.txt under .claude/projects/<slug>/osb/<session>/; newest if several) and Grep it for the changes/work under test. Also check project docs and known-issues files. Key doc locations to consult: `/workspace/osborn/CLAUDE.md`, `/workspace/osborn/docs/critical-patterns.md`, the `docs/` directory, `README.md`, and `CHANGELOG.md`. Check these for: (a) KNOWN ISSUES and gotchas already recorded, and (b) what behavior is ALREADY covered by existing tests.',
@@ -372,13 +369,13 @@ export const NAMED_AGENTS = {
372
369
  'test-WRITING phase only.)',
373
370
  '',
374
371
  '## How to work',
375
- '0. **Identify what changed (adapt to context):**',
376
- ' - If the working directory is a git repo: run `git diff HEAD~1 HEAD --name-only` then `git diff HEAD~1 HEAD` and build your plan around the specific changed files.',
377
- ' - If there is NO git repo or the folder is fresh/empty: list the files the writer created or modified (Glob/Read), and use those as your test scope. Many valid tasks — automation scripts, research outputs, YouTube workflows, Chrome experiments — produce real deliverables with no git history. Adapt accordingly.',
378
- '1. Identify the correct validation command from the task context: package.json test script, Makefile, a direct script run (`python script.py`, `node index.js`, `bash run.sh`), a build command, or a manual verification step. "Testing" means confirming the work actually does what was asked — not just running npm test.',
379
- '2. Run the FULL existing test suite first to establish the regression baseline (skip if none exists).',
380
- '3. Generate and run tests targeted at the SPECIFIC change — at both unit and integration levels where applicable.',
381
- ' Focus on: the changed or created files/functions, their callers, and any behavior the work modifies.',
372
+ '0. **Get the diff first (MANDATORY):** Run `git diff HEAD~1 HEAD --name-only` to get the list of',
373
+ ' changed files, then `git diff HEAD~1 HEAD` for the full diff. Build your entire test plan around',
374
+ ' the SPECIFIC files and functions that changed — not a generic sweep.',
375
+ '1. Identify the correct test / build command from package.json, Makefile, or the task brief.',
376
+ '2. Run the FULL existing test suite first to establish the regression baseline.',
377
+ '3. Generate and run tests targeted at the SPECIFIC diff/change — at both unit and integration levels.',
378
+ ' Focus on: the changed functions/components, their callers, and any behavior the diff modifies.',
382
379
  '4. Exercise edge cases: boundary values, empty inputs, error/exception paths, null/undefined.',
383
380
  '5. Execution loop: write test → run it → read failure output → fix the test OR flag as a real bug in the code. Do NOT silently paper over a real defect.',
384
381
  '6. If a command fails, read the relevant source files to locate the root cause.',
@@ -477,12 +474,10 @@ export const NAMED_AGENTS = {
477
474
  },
478
475
  reviewer: {
479
476
  description: [
480
- 'Review agent (Opus). Use for: the VERIFY step in a generator-verifier loop — after the',
481
- 'writer completes work, the reviewer reads what was produced, checks correctness, spec/requirement',
482
- 'adherence, obvious issues, and quality, then tags each finding BLOCKER/MAJOR/MINOR/NIT',
483
- 'and returns an ACCEPT or REJECT verdict with specific, actionable feedback.',
484
- 'Works with or without a git repo — reviews files directly when there is no git history.',
485
- 'May write documentation files (.md etc.) only.',
477
+ 'Code-review agent (Opus). Use for: the VERIFY step in a generator-verifier loop — after the',
478
+ 'writer completes a change, the reviewer reads the diff, checks correctness, spec/requirement',
479
+ 'adherence, obvious bugs, and security issues, then tags each finding BLOCKER/MAJOR/MINOR/NIT',
480
+ 'and returns an ACCEPT or REJECT verdict with specific, actionable feedback. May write documentation files (.md etc.) only.',
486
481
  ].join(' '),
487
482
  tools: ['Read', 'Glob', 'Grep', 'Bash', 'Write', 'Edit'],
488
483
  model: 'opus',
@@ -541,12 +536,11 @@ export const NAMED_AGENTS = {
541
536
  'of whether the diff looks clean.',
542
537
  '',
543
538
  '## How to work',
544
- '0. **Identify what changed (adapt to context):**',
545
- ' - If the working directory is a git repo with commits: run `git diff HEAD~1 HEAD --stat` then `git diff HEAD~1 HEAD`. Build your review around what ACTUALLY changed.',
546
- ' - If there is NO git repo, the folder is fresh, or git history is empty: do NOT treat this as an error. Many valid tasks — automation scripts, YouTube workflows, Chrome experiments, research outputs, content creation — produce real deliverables with no git history. Instead, Glob/Read the files the writer produced and review them directly against the task spec and any project standards.',
547
- ' - Either way: your job is to verify the work meets the requirement, not to enforce a git workflow.',
548
- '1. Read the files involved (whether from git diff or direct Glob/Read) to see exactly what was produced or changed.',
549
- '2. Read any file that needs context to evaluate the work (interfaces, callers, specs, existing code).',
539
+ '0. **Get the diff first (MANDATORY):** Run `git diff HEAD~1 HEAD --stat` then `git diff HEAD~1 HEAD`.',
540
+ ' Build your entire review around what ACTUALLY changed — not the writer\'s narrative alone.',
541
+ ' If the task provides a diff, still verify it matches git history.',
542
+ '1. Run `git diff` (or read the files listed in the task) to see exactly what changed.',
543
+ '2. Read any file that needs context to evaluate the diff (interfaces, callers, tests).',
550
544
  '3. Run the build or test suite if available to catch compile/runtime regressions.',
551
545
  '4. Check against the spec or requirement provided in the task brief AND any project standards found above.',
552
546
  '5. Look for: logic errors, missing edge cases, security issues (injection, path traversal,',
package/dist/index.js CHANGED
@@ -6088,7 +6088,10 @@ async function main() {
6088
6088
  }
6089
6089
  // Feature B — per-flow stop (does NOT affect other running flows)
6090
6090
  else if (data.type === 'stop_dispatch' && currentLLM) {
6091
- currentLLM.stopAgent?.(String(data.agentId));
6091
+ const agentId = String(data.agentId);
6092
+ currentLLM.stopAgent?.(agentId);
6093
+ // Always confirm so frontend clears the card (covers ghost/stale flows too)
6094
+ await sendToFrontend({ type: 'agent_stopped', agent_id: agentId });
6092
6095
  }
6093
6096
  else if (data.type === 'join_meeting') {
6094
6097
  const meetingUrl = data.url;
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "osborn",
3
- "version": "0.9.172",
3
+ "version": "0.9.174",
4
4
  "description": "Voice AI coding assistant - local agent that connects to Osborn frontend",
5
5
  "type": "module",
6
6
  "bin": {