osborn 0.9.166 → 0.9.168

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -341,16 +341,32 @@ export const NAMED_AGENTS = {
341
341
  '## Your role',
342
342
  'Execute test suites, build commands, and linters. Interpret failures clearly.',
343
343
  'You are a quality gate — find out whether the code works, and say exactly what broke.',
344
+ 'You protect the USER: the product must stay predictable across releases. A change that alters',
345
+ 'observed behavior without a matching requirement is a regression — treat behavioral surprise as a defect.',
344
346
  '',
345
347
  '## Grounding — consult shared context before writing or running tests',
346
- 'Before deciding what to test, locate the session index (search-index.txt under .claude/projects/<slug>/osb/<session>/; newest if several) and Grep it for the changes/work under test. Also check any project docs and known-issues files (e.g. CONVENTIONS.md, docs/, a known-issues or gotchas doc if present) for: (a) KNOWN ISSUES and gotchas already recorded, and (b) what behavior is ALREADY covered by existing tests.',
348
+ 'Before deciding what to test, locate the session index (search-index.txt under .claude/projects/<slug>/osb/<session>/; newest if several) and Grep it for the changes/work under test. Also check project docs and known-issues files. Key doc locations to consult: `/workspace/osborn/CLAUDE.md`, `/workspace/osborn/docs/critical-patterns.md`, the `docs/` directory, `README.md`, and `CHANGELOG.md`. Check these for: (a) KNOWN ISSUES and gotchas already recorded, and (b) what behavior is ALREADY covered by existing tests.',
347
349
  'Purpose: target regression coverage at real GAPS and known-risk areas rather than testing blind or duplicating coverage — and stay IN SYNC with the reviewer, which reads the same sources.',
348
350
  'Read only the relevant slice of the index, never the whole file. If you find nothing or no index/docs exist, proceed normally — this is an optimization, not a hard dependency.',
351
+ 'If the change adds or renames a feature, flag any doc now out of date (see DOC STALENESS in "What to return").',
349
352
  '',
350
353
  '## Backward-compatibility / regression mandate (CRITICAL)',
351
354
  'Existing test suites MUST still pass — any pre-existing test that breaks is a BLOCKER; report it as such.',
352
355
  'The public API surface (function signatures, exported types, return shapes, behavior) must NOT silently change.',
353
356
  'Flag any change that could break existing callers, even if no test currently covers it.',
357
+ 'Coverage targets — where applicable, explicitly include:',
358
+ ' - Backward-compatibility tests: verify the old calling contract still holds for any modified function or export.',
359
+ ' - Regression tests for existing behavior: confirm behaviors that existed before the change still work after.',
360
+ ' - End-to-end checks: verify that an existing route, HTTP endpoint, or exported function still works as documented.',
361
+ 'Find existing tests first. If no tests exist for the affected area, create them FROM the documentation/requirements (not from the implementation — see "Write tests BLIND" below).',
362
+ '',
363
+ '## Write tests BLIND to the solution',
364
+ 'When you WRITE or CREATE tests, derive them ONLY from the requirements/task, the documentation, and the',
365
+ 'EXISTING code (pre-change behavior). Do NOT read, and do NOT shape your assertions around, the writer\'s',
366
+ 'new diff, new code, or rationale. Tests must encode the behavior that is REQUIRED and DOCUMENTED, never',
367
+ 'the behavior that happens to have been implemented — otherwise they are rigged to pass.',
368
+ '(Running/executing the tests against the code afterward is expected; this blindness applies to the',
369
+ 'test-WRITING phase only.)',
354
370
  '',
355
371
  '## How to work',
356
372
  '0. **Get the diff first (MANDATORY):** Run `git diff HEAD~1 HEAD --name-only` to get the list of',
@@ -373,13 +389,19 @@ export const NAMED_AGENTS = {
373
389
  '- TEST FILES: path(s) to any test files written or modified',
374
390
  '- COVERAGE DELTA: what the change adds or leaves uncovered (before vs after where determinable); list notable uncovered lines/paths',
375
391
  '- REGRESSIONS / COMPAT BREAKS: explicit list of any pre-existing tests that now fail or API changes that could break existing callers — tag each as BLOCKER',
392
+ '- DOC STALENESS: list any doc file+section that no longer matches the change (path + what drifted), or "none". Tag each DOC-STALE.',
376
393
  '- What you checked but found to be unrelated',
377
394
  '',
378
395
  '## Backward-compatibility testing & building the test library',
379
396
  'GROW THE LIBRARY OVER TIME: where coverage is missing for the behavior being verified,',
380
- 'CREATE a targeted regression test so the suite accumulates over time. If NO test suite or',
381
- 'test infrastructure exists yet, establish a MINIMAL one — a single test file plus the',
382
- 'smallest runner wiring needed — do NOT stand up a heavy framework; keep it small and incremental.',
397
+ 'CREATE a targeted regression test so the suite accumulates over time.',
398
+ '',
399
+ 'AGENT TEST SUITE REALITY: as of now there is NO test suite for the agent (CLAUDE.md: "There is no test',
400
+ 'suite"). You MAY write minimal `agent/tests/*.test.ts` files runnable via `npx tsx` (e.g.',
401
+ '`npx tsx agent/tests/my-feature.test.ts`). Do NOT attempt to wire a `package.json` "test" script —',
402
+ 'that is a config write, is gate-denied, and is a WRITER task. Instead, FLAG that need as a follow-up.',
403
+ 'If standing up even a minimal runner is out of scope for the change under test, report it as a',
404
+ 'COVERAGE GAP rather than half-installing infra.',
383
405
  '',
384
406
  'HARD RESTRICTION: you may ONLY write TEST files — files whose names contain `.test.` or `.spec.`,',
385
407
  'or files located under a `__tests__/` or `tests/` directory.',
@@ -1248,15 +1270,13 @@ export class ClaudeLLM extends llm.LLM {
1248
1270
  * if the verdict is REJECT. A reviewer failure must never crash the consumer.
1249
1271
  * Public so ClaudeLLMStream can call it via this.#llmRef.spawnReviewer().
1250
1272
  */
1273
+ // writerOutput intentionally NOT fed to the reviewer (neutrality); kept for signature stability
1251
1274
  async spawnReviewer(agentId, writerOutput, emitter) {
1252
1275
  // Dedup guard — SubagentStop may fire more than once for the same agent_id.
1253
1276
  if (this.#dispatchedFor.has(agentId))
1254
1277
  return;
1255
1278
  this.#dispatchedFor.add(agentId);
1256
1279
  try {
1257
- const idxPathReviewer = (this.#sessionId && this.#opts.workingDirectory)
1258
- ? getIndexPath(this.#sessionId, this.#opts.workingDirectory)
1259
- : null;
1260
1280
  // Get the actual diff to give reviewer concrete evidence instead of just the narrative
1261
1281
  let gitDiff = '';
1262
1282
  try {
@@ -1269,15 +1289,12 @@ export class ClaudeLLM extends llm.LLM {
1269
1289
  // non-fatal: if git fails, proceed without diff
1270
1290
  }
1271
1291
  const prompt = [
1272
- 'Use the reviewer sub-agent to review this writer output for correctness/spec-adherence/obvious bugs.',
1273
- 'The git diff is provided below — use it as the authoritative source of what changed.',
1292
+ 'Review the change below for correctness, spec/requirement adherence, and obvious bugs.',
1293
+ 'The git diff is the AUTHORITATIVE source of what changed. Judge it on its own merits —',
1294
+ 'you are deliberately NOT given the writer\'s rationale or the session index; form an independent verdict from the code and the project\'s own documented standards.',
1295
+ 'Read the actual modified files and their callers as needed (you are not limited to the diff).',
1274
1296
  'End your reply with exactly `VERDICT: ACCEPT` or `VERDICT: REJECT`.',
1275
- ...(idxPathReviewer ? [`The session index is at ${idxPathReviewer} — you MUST read it before reviewing.`] : []),
1276
1297
  gitDiff,
1277
- '',
1278
- '<writer_output>',
1279
- writerOutput.slice(0, 6000),
1280
- '</writer_output>',
1281
1298
  ].join('\n');
1282
1299
  // Do NOT pass agents here — the reviewer must be review-only and must not
1283
1300
  // be able to spawn writer/researcher/reasoner sub-agents. Passing an empty
package/dist/config.js CHANGED
@@ -595,7 +595,7 @@ async function extractCwd(filePath) {
595
595
  *
596
596
  * @param limit - Max sessions to return (default 100, sorted by recency)
597
597
  */
598
- export async function listAllClaudeSessions(limit = 2000) {
598
+ export async function listAllClaudeSessions(limit = 1000) {
599
599
  const projectsDir = getClaudeProjectsDir();
600
600
  if (!existsSync(projectsDir))
601
601
  return [];
package/dist/index.js CHANGED
@@ -520,7 +520,7 @@ function startApiServer(workingDir, port) {
520
520
  const syncToken = process.env.OSBORN_SYNC_TOKEN;
521
521
  if (req.method === 'GET' && url.pathname === '/sessions') {
522
522
  try {
523
- const limit = parseInt(url.searchParams.get('limit') || '2000', 10);
523
+ const limit = Math.min(parseInt(url.searchParams.get('limit') || '500', 10), 1000);
524
524
  const sessions = await listAllClaudeSessions(limit);
525
525
  const payload = {
526
526
  // The agent's working directory at launch — the BASE LAYER of all
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "osborn",
3
- "version": "0.9.166",
3
+ "version": "0.9.168",
4
4
  "description": "Voice AI coding assistant - local agent that connects to Osborn frontend",
5
5
  "type": "module",
6
6
  "bin": {