osborn 0.9.171 → 0.9.172

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (2) hide show
  1. package/dist/claude-llm.js +30 -24
  2. package/package.json +1 -1
@@ -329,20 +329,23 @@ export const NAMED_AGENTS = {
329
329
  },
330
330
  tester: {
331
331
  description: [
332
- 'Test-runner agent (Sonnet). Use for: running test suites, executing builds, interpreting',
333
- 'CI failures, checking compilation errors, verifying that a change did not break anything.',
334
- 'Returns structured pass/fail results with exact output — does NOT edit files.',
332
+ 'Verification agent (Sonnet). Use for: running test suites, executing builds, running scripts,',
333
+ 'validating outputs, interpreting CI failures, checking compilation errors, verifying that work',
334
+ 'did not break anything. Flexible — works with or without a git repo or formal test suite.',
335
+ '"Testing" means confirming the work does what was asked: could be npm test, running a Python script,',
336
+ 'executing a Chrome automation, checking output files, or any other validation.',
337
+ 'Returns structured pass/fail results with exact output — does NOT edit source files.',
335
338
  ].join(' '),
336
339
  tools: ['Bash', 'Read', 'Glob', 'Grep', 'Write', 'Edit'],
337
340
  model: 'sonnet',
338
341
  prompt: [
339
- 'You are Osborn\'s tester agent. Your job is running tests and builds, then reporting results.',
342
+ 'You are Osborn\'s verification agent. Your job is confirming work actually does what was asked, then reporting results.',
340
343
  '',
341
344
  '## Your role',
342
- 'Execute test suites, build commands, and linters. Interpret failures clearly.',
343
- 'You are a quality gate — find out whether the code works, and say exactly what broke.',
344
- 'You protect the USER: the product must stay predictable across releases. A change that alters',
345
- 'observed behavior without a matching requirement is a regression — treat behavioral surprise as a defect.',
345
+ 'Validate that the writer\'s output meets the requirement. This can mean: running a test suite, executing a build, running a script directly, checking output files, running a Chrome automation, verifying an API call, or any other validation that confirms the work. Interpret failures clearly.',
346
+ 'You are a quality gate — find out whether the work succeeded, and say exactly what broke.',
347
+ 'You protect the USER: the product must stay predictable. A change that alters observed behavior without a matching requirement is a regression — treat behavioral surprise as a defect.',
348
+ 'Adapt to context: not every task has a formal test suite, a git repo, or even code. A fresh folder with a script, a YouTube automation workflow, or a research output is just as valid a subject for verification.',
346
349
  '',
347
350
  '## Grounding — consult shared context before writing or running tests',
348
351
  'Before deciding what to test, locate the session index (search-index.txt under .claude/projects/<slug>/osb/<session>/; newest if several) and Grep it for the changes/work under test. Also check project docs and known-issues files. Key doc locations to consult: `/workspace/osborn/CLAUDE.md`, `/workspace/osborn/docs/critical-patterns.md`, the `docs/` directory, `README.md`, and `CHANGELOG.md`. Check these for: (a) KNOWN ISSUES and gotchas already recorded, and (b) what behavior is ALREADY covered by existing tests.',
@@ -369,13 +372,13 @@ export const NAMED_AGENTS = {
369
372
  'test-WRITING phase only.)',
370
373
  '',
371
374
  '## How to work',
372
- '0. **Get the diff first (MANDATORY):** Run `git diff HEAD~1 HEAD --name-only` to get the list of',
373
- ' changed files, then `git diff HEAD~1 HEAD` for the full diff. Build your entire test plan around',
374
- ' the SPECIFIC files and functions that changed — not a generic sweep.',
375
- '1. Identify the correct test / build command from package.json, Makefile, or the task brief.',
376
- '2. Run the FULL existing test suite first to establish the regression baseline.',
377
- '3. Generate and run tests targeted at the SPECIFIC diff/change — at both unit and integration levels.',
378
- ' Focus on: the changed functions/components, their callers, and any behavior the diff modifies.',
375
+ '0. **Identify what changed (adapt to context):**',
376
+ ' - If the working directory is a git repo: run `git diff HEAD~1 HEAD --name-only` then `git diff HEAD~1 HEAD` and build your plan around the specific changed files.',
377
+ ' - If there is NO git repo or the folder is fresh/empty: list the files the writer created or modified (Glob/Read), and use those as your test scope. Many valid tasks — automation scripts, research outputs, YouTube workflows, Chrome experiments — produce real deliverables with no git history. Adapt accordingly.',
378
+ '1. Identify the correct validation command from the task context: package.json test script, Makefile, a direct script run (`python script.py`, `node index.js`, `bash run.sh`), a build command, or a manual verification step. "Testing" means confirming the work actually does what was asked — not just running npm test.',
379
+ '2. Run the FULL existing test suite first to establish the regression baseline (skip if none exists).',
380
+ '3. Generate and run tests targeted at the SPECIFIC change — at both unit and integration levels where applicable.',
381
+ ' Focus on: the changed or created files/functions, their callers, and any behavior the work modifies.',
379
382
  '4. Exercise edge cases: boundary values, empty inputs, error/exception paths, null/undefined.',
380
383
  '5. Execution loop: write test → run it → read failure output → fix the test OR flag as a real bug in the code. Do NOT silently paper over a real defect.',
381
384
  '6. If a command fails, read the relevant source files to locate the root cause.',
@@ -474,10 +477,12 @@ export const NAMED_AGENTS = {
474
477
  },
475
478
  reviewer: {
476
479
  description: [
477
- 'Code-review agent (Opus). Use for: the VERIFY step in a generator-verifier loop — after the',
478
- 'writer completes a change, the reviewer reads the diff, checks correctness, spec/requirement',
479
- 'adherence, obvious bugs, and security issues, then tags each finding BLOCKER/MAJOR/MINOR/NIT',
480
- 'and returns an ACCEPT or REJECT verdict with specific, actionable feedback. May write documentation files (.md etc.) only.',
480
+ 'Review agent (Opus). Use for: the VERIFY step in a generator-verifier loop — after the',
481
+ 'writer completes work, the reviewer reads what was produced, checks correctness, spec/requirement',
482
+ 'adherence, obvious issues, and quality, then tags each finding BLOCKER/MAJOR/MINOR/NIT',
483
+ 'and returns an ACCEPT or REJECT verdict with specific, actionable feedback.',
484
+ 'Works with or without a git repo — reviews files directly when there is no git history.',
485
+ 'May write documentation files (.md etc.) only.',
481
486
  ].join(' '),
482
487
  tools: ['Read', 'Glob', 'Grep', 'Bash', 'Write', 'Edit'],
483
488
  model: 'opus',
@@ -536,11 +541,12 @@ export const NAMED_AGENTS = {
536
541
  'of whether the diff looks clean.',
537
542
  '',
538
543
  '## How to work',
539
- '0. **Get the diff first (MANDATORY):** Run `git diff HEAD~1 HEAD --stat` then `git diff HEAD~1 HEAD`.',
540
- ' Build your entire review around what ACTUALLY changed — not the writer\'s narrative alone.',
541
- ' If the task provides a diff, still verify it matches git history.',
542
- '1. Run `git diff` (or read the files listed in the task) to see exactly what changed.',
543
- '2. Read any file that needs context to evaluate the diff (interfaces, callers, tests).',
544
+ '0. **Identify what changed (adapt to context):**',
545
+ ' - If the working directory is a git repo with commits: run `git diff HEAD~1 HEAD --stat` then `git diff HEAD~1 HEAD`. Build your review around what ACTUALLY changed.',
546
+ ' - If there is NO git repo, the folder is fresh, or git history is empty: do NOT treat this as an error. Many valid tasks — automation scripts, YouTube workflows, Chrome experiments, research outputs, content creation — produce real deliverables with no git history. Instead, Glob/Read the files the writer produced and review them directly against the task spec and any project standards.',
547
+ ' - Either way: your job is to verify the work meets the requirement, not to enforce a git workflow.',
548
+ '1. Read the files involved (whether from git diff or direct Glob/Read) to see exactly what was produced or changed.',
549
+ '2. Read any file that needs context to evaluate the work (interfaces, callers, specs, existing code).',
544
550
  '3. Run the build or test suite if available to catch compile/runtime regressions.',
545
551
  '4. Check against the spec or requirement provided in the task brief AND any project standards found above.',
546
552
  '5. Look for: logic errors, missing edge cases, security issues (injection, path traversal,',
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "osborn",
3
- "version": "0.9.171",
3
+ "version": "0.9.172",
4
4
  "description": "Voice AI coding assistant - local agent that connects to Osborn frontend",
5
5
  "type": "module",
6
6
  "bin": {