@ngockhoale/ukit 2.3.13 → 2.3.15

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,64 @@
2
2
 
3
3
  All notable changes to UKit are documented here.
4
4
 
5
+ ## 2.3.15 - 2026-09-11
6
+
7
+ Vision-lane root cause, part 4 — native-first dispatch order. Found by mining the user's own
8
+ failing sessions across four real projects (UnicDB, VSDB, BAGuide, AI-Gateway): the API `model`
9
+ field in the transcripts proved failed specialist dispatches landed on `glm-5-turbo` / lite lanes
10
+ (self-reported `MODEL: unic-lite`, `STATUS: WRONG_MODEL`) while the same sessions' main Claude
11
+ lanes read images fine (`STATUS: OK` ×14 in one transcript). Two stacked defects:
12
+
13
+ - **No fallback when the specialist fails.** The honored-lane hint ordered the parent to dispatch
14
+ `ukit-vision-analyst` and simply continue; when the gateway routed the alias to a backend that
15
+ cannot receive the image through the Read tool_result, the analyst returned `WRONG_MODEL` and
16
+ the image ended up unread even though the parent could have read it natively. The hint now
17
+ verifies the parent's own vision FIRST (step 2: Read the materialized files yourself, write the
18
+ receipts, do NOT dispatch), dispatches the specialist only when the parent's own Read yields no
19
+ image (step 3), and defines the fallback: on `WRONG_MODEL`/`NO_IMAGE` re-read natively, analyse
20
+ from your own view for pasted images, and never fabricate when no reader can see it (step 4).
21
+ - **Refusal by name instead of by probe.** The omp analyst template mandated refusing *before
22
+ touching any image* when the runtime model could not be confirmed vision-capable — field
23
+ transcripts show exactly this refusal (`I must refuse to open, describe, or guess`). Both agent
24
+ templates now decide ONLY by the first image Read (the probe): a backend that self-identifies as
25
+ `glm-5-turbo`/`MiniMax-M3`/a lite lane does not license refusal when the probe actually shows
26
+ the image; a vision-sounding name with a failed probe is still `WRONG_MODEL`.
27
+
28
+ Unchanged doctrine: never guess at image contents — the fix removes the two paths that turned
29
+ "don't guess" into "don't read". omp mirrors synced. TDD: `vision-router-hint` tests 13-14 and
30
+ `vision-agent` test 13 written RED first, then GREEN; all vision suites pass.
31
+
32
+ ## 2.3.14 - 2026-09-11
33
+
34
+ Vision-lane root cause, part 3 — the decisive one: the extractor could not see images in REAL
35
+ Claude Code transcripts at all. Found by reproducing the failure against the user's own real
36
+ session transcripts after they insisted the bug lived "right at the image" — grep proved the
37
+ transcripts contained `"type":"image"` while every extraction returned `imageCount: 0`. The
38
+ 2.3.12/2.3.13 suites were green because the fixtures modeled an invented transcript shape.
39
+
40
+ **P1 — real transcript image envelopes were invisible to the extractor.** Verified against
41
+ real transcripts from four projects, Claude Code stores images in THREE shapes:
42
+
43
+ - flat: `message.content[i] = { type: 'image', source: { type: 'base64', … } }` — the only
44
+ shape the pre-fix parser handled;
45
+ - pasted: a top-level `attachment` envelope (`type: 'attachment'`,
46
+ `attachment.prompt[]` = content blocks) — every real pasted image lives here, NOT in
47
+ `message.content`; the pre-fix parser extracted nothing from any pasted image;
48
+ - tool-returned: `message.content[i] = { type: 'tool_result', content: [ { type: 'image',
49
+ … } ] }` — screenshots coming back from tools nest one level down; also invisible.
50
+
51
+ Fix: `extractImageBlocks()` now walks a bounded collector (depth ≤ 3) across
52
+ `message.content[]`, nested `content[]` arrays inside blocks, and the top-level
53
+ `attachment.prompt[]` / `attachment.content` envelopes. Same-base64 payloads across
54
+ envelopes dedupe by sha as before. Real-data validation after the fix: the same four
55
+ transcripts now extract valid images (PNG 2278×2154 from the paste envelope, JPEG 1400×884
56
+ from a tool_result, PNG/JPEG from the flat shape), and a live analyst dispatch read the
57
+ real pasted screenshot.
58
+
59
+ Tests: `tests/handoff/cycle10/extract-image-contract.test.mjs` gains three cases whose
60
+ fixtures copy the real envelope structures verbatim (RED pre-fix: `imageCount 0`), plus a
61
+ cross-envelope dedupe pin. All cycle10 + cycle4 vision suites green.
62
+
5
63
  ## 2.3.13 - 2026-09-11
6
64
 
7
65
  Vision-lane root cause, part 2: the Codex model catalog could still deny the vision lane on
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ngockhoale/ukit",
3
- "version": "2.3.13",
3
+ "version": "2.3.15",
4
4
  "description": "Install/update an index-first AI workspace for Claude Code, OpenAI Codex, OpenCode, and omp (Oh My Pi).",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -27,6 +27,12 @@ demonstrated, never assumed from a name:
27
27
  actually give you the image (error, empty result, or you can only see text you were told about),
28
28
  you are not vision-capable right now: emit `STATUS: WRONG_MODEL` and stop immediately. Do not
29
29
  describe, summarise, or guess at any image content.
30
+ - Decide ONLY by the probe result — never by the lane's name. A mapping that sounds non-visual
31
+ (e.g. your environment says you are a lite lane) does NOT license refusal when the probe
32
+ actually shows you the image: if you can see it, you are vision-capable right now — report the
33
+ mapping you actually ran on in `MODEL:` and continue. Conversely, a vision-sounding name with a
34
+ failed probe is still `STATUS: WRONG_MODEL`. Refusing by name without probing was the exact
35
+ failure that left images unread in the field.
30
36
  - Never guess at image contents. A non-visual "analysis" is worse than no analysis at all, because
31
37
  it looks authoritative while being fabricated. Refusing loudly is always safer than guessing
32
38
  quietly.
@@ -283,31 +283,34 @@ const { pathToFileURL } = require('url');
283
283
  }
284
284
 
285
285
  // Lane honored (or unverified — missing aliasAvailable defaults to today's
286
- // unic-vision dispatch). Capability-framed, never provider identity: the
287
- // remedy is "don't guess — use a verified reader", which holds whether the
288
- // active model reads images natively or hands off to the specialist.
286
+ // dispatch). NATIVE-FIRST ordering (2026-09-11): field evidence showed
287
+ // specialist dispatches landing on gateways/backends that could not return
288
+ // the image through the Read tool_result (STATUS: WRONG_MODEL, self-reported
289
+ // unic-lite / glm-5-turbo) while the parent — usually a Claude lane — could
290
+ // have read the image itself. So the hint verifies the parent's own vision
291
+ // FIRST by reading the materialized file, dispatches the specialist only as
292
+ // a fallback, and defines a fallback when the specialist fails anyway —
293
+ // an image must never end up unread just because one lane failed.
289
294
  const unicNote = (unicMode === true && gatewayResult?.visionModel)
290
295
  ? ` (UNIC gateway active — ${gatewayResult.visionModel} routes through it.)`
291
296
  : '';
292
297
 
293
- const reasonLines = [
294
- 'Advisory: never guess at image contents. If the active model has VERIFIED native',
295
- 'vision for these images it may read them directly; otherwise dispatch the specialist',
296
- 'before relying on them:',
297
- ];
298
-
299
298
  const lines = [
300
299
  `UKIT VISION ROUTE — new image input detected (${cases.join(', ')}).`,
301
- ...reasonLines,
300
+ 'Advisory: never guess at image contents. Verify a real reader before relying on them:',
302
301
  ` 1. ${materializeCmd}`,
303
- ' Materialize the images to disk. With --session the extractor targets the',
304
- ' PARENT transcript explicitly, so a subagent spawn cannot drop the image.',
305
- ' 2. Agent(subagent_type: "ukit-vision-analyst") [model: unic-vision]',
306
- ' Send the ABSOLUTE paths from images[].path as TEXT (subagents do NOT inherit',
307
- ' image blocks; they can only Read files). Include the task envelope: the ORIGINAL',
308
- ' user prompt verbatim, the visual question, and the task goal — the analyst must',
309
- ' know what the images are FOR.',
310
- ' 3. Continue the real task using the analyst\'s OBSERVATIONS + INFERENCES.',
302
+ ' Materialize to disk (--session targets the PARENT transcript, so a subagent',
303
+ ' spawn cannot drop the image).',
304
+ ' 2. Read each ABSOLUTE path in images[].path YOURSELF, in the main session. If',
305
+ ' the image actually arrives you have VERIFIED native vision: analyse it and',
306
+ ' write receipts (analyzed-<sha>.json; sha/sessionId from the extractor output;',
307
+ ' model = your actual model). No dispatch.',
308
+ ' 3. Only if your own Read shows no image: Agent(subagent_type: "ukit-vision-analyst")',
309
+ ' [model: unic-vision] — send paths as TEXT (subagents do NOT inherit image',
310
+ ' blocks) plus the task envelope: original prompt, visual question, goal.',
311
+ ' 4. Specialist answers WRONG_MODEL or NO_IMAGE → do not leave the image unread:',
312
+ ' re-Read yourself; for a pasted image you can see directly, analyse from your',
313
+ ' own view. No reader can see it → say so plainly, never fabricate.',
311
314
  ];
312
315
  if (sessionId) {
313
316
  lines.push(` Markers armed under sessionId: ${sessionId} (receipts: analyzed-<sha>.json).`);
@@ -381,25 +381,64 @@ function readLines(filePath) {
381
381
  return raw.split('\n').filter((line) => line.trim().length > 0);
382
382
  }
383
383
 
384
+ /**
385
+ * Deepest content-blocks nesting the collector walks. 1 covers the flat
386
+ * message.content[] / attachment.prompt[] arrays; 2 covers the real-world
387
+ * tool_result nesting (message.content[i].content[]); 3 is a cheap safety
388
+ * bound — deeper structures are not a Claude Code shape and are ignored.
389
+ */
390
+ const MAX_BLOCK_WALK_DEPTH = 3;
391
+
392
+ /**
393
+ * Collects base64 image blocks from one content-blocks array. Real Claude
394
+ * Code transcripts nest images in more than one place (verified against real
395
+ * session transcripts 2026-09-11 — the pre-fix parser only saw the flat
396
+ * shape and extracted NOTHING from real paste/tool-image traffic):
397
+ * A. message.content[i] = { type: 'image', source: { type: 'base64', … } }
398
+ * C. message.content[i] = { type: 'tool_result', content: [ { type: 'image', … } ] }
399
+ * Blocks that carry their own `content` array (tool_result etc.) are walked
400
+ * one level deeper, bounded by MAX_BLOCK_WALK_DEPTH.
401
+ */
402
+ function collectImageBlocksFromContent(content, lineIndex, blocks, depth) {
403
+ if (!Array.isArray(content) || depth > MAX_BLOCK_WALK_DEPTH) return;
404
+ for (const block of content) {
405
+ if (!block || typeof block !== 'object') continue;
406
+ if (block.type === 'image') {
407
+ const source = block.source;
408
+ if (!source || source.type !== 'base64') continue;
409
+ const mediaType = source.media_type;
410
+ const data = source.data;
411
+ if (typeof mediaType !== 'string' || typeof data !== 'string' || !data) continue;
412
+ blocks.push({ mediaType, data, lineIndex });
413
+ } else if (Array.isArray(block.content)) {
414
+ collectImageBlocksFromContent(block.content, lineIndex, blocks, depth + 1);
415
+ }
416
+ }
417
+ }
418
+
384
419
  /**
385
420
  * Parses each JSONL line independently. A malformed line is skipped via its
386
421
  * own try/catch and never throws out to the caller.
422
+ *
423
+ * Envelope shapes collected per line (real-transcript evidence 2026-09-11):
424
+ * A/C. the API message envelope — flat image blocks, plus tool_result
425
+ * blocks that nest theirs one level down (see the collector above).
426
+ * B. pasted images arrive in a TOP-LEVEL attachment envelope, NOT in
427
+ * message.content: attachment.prompt[] is the content-blocks array
428
+ * (attachment.content accepted defensively for the same reason).
429
+ * The same payload appearing in two envelopes (attachment + the later user
430
+ * message) dedupes by sha in selectImages().
387
431
  */
388
432
  function extractImageBlocks(lines) {
389
433
  const blocks = [];
390
434
  for (let i = 0; i < lines.length; i += 1) {
391
435
  try {
392
436
  const obj = JSON.parse(lines[i]);
393
- const content = obj?.message?.content;
394
- if (!Array.isArray(content)) continue;
395
- for (const block of content) {
396
- if (block?.type !== 'image') continue;
397
- const source = block.source;
398
- if (!source || source.type !== 'base64') continue;
399
- const mediaType = source.media_type;
400
- const data = source.data;
401
- if (typeof mediaType !== 'string' || typeof data !== 'string' || !data) continue;
402
- blocks.push({ mediaType, data, lineIndex: i });
437
+ collectImageBlocksFromContent(obj?.message?.content, i, blocks, 1);
438
+ const attachment = obj?.attachment;
439
+ if (attachment && typeof attachment === 'object') {
440
+ collectImageBlocksFromContent(attachment.prompt, i, blocks, 1);
441
+ collectImageBlocksFromContent(attachment.content, i, blocks, 1);
403
442
  }
404
443
  } catch {
405
444
  // malformed line: skip and continue
@@ -17,19 +17,23 @@ actually in them. You never write, edit, or refactor product code — you analys
17
17
  - Stay end-user-invisible: this lane is internal UKit orchestration, not something end users invoke
18
18
  by name.
19
19
 
20
- ## 2. Model self-check — first action, before touching any image
21
-
22
- Before reading any image, determine the model you are actually running on right now.
23
-
24
- - Vision-capable means: `unic-vision`, or — when `unicMode` is off — whatever
25
- `modelTiers.vision.fallbackModel` resolves to per
26
- `node .claude/ukit/index/unic-gateway.mjs --json`.
27
- - If you cannot confirm you are running on a vision-capable model, you must **refuse**: emit
28
- `STATUS: WRONG_MODEL` in the output block below and stop immediately. Do not open, describe, or
29
- guess at any image content.
30
- - Never guess at image contents. A wrong-model "analysis" is worse than no analysis at all,
31
- because it looks authoritative while being fabricated. Refusing loudly is always safer than
32
- guessing quietly.
20
+ ## 2. Capability self-check — the first image read is the probe
21
+
22
+ Model names are gateway mappings whose backend can change at any time, so capability is
23
+ demonstrated, never assumed from a name:
24
+
25
+ - Your first `read` of an image file doubles as the capability probe. If the tool result does not
26
+ actually give you the image (error, empty result, or you can only see text you were told about),
27
+ you are not vision-capable right now: emit `STATUS: WRONG_MODEL` in the output block below and
28
+ stop immediately. Do not describe, summarise, or guess at any image content.
29
+ - Decide ONLY by the probe result — never by the lane's name. A backend that self-identifies as
30
+ something else (e.g. `glm-5-turbo`, `MiniMax-M3`, a lite lane) does NOT license refusal when the
31
+ probe actually shows you the image: if you can see it, you are vision-capable right now — report
32
+ what you actually ran on in `MODEL:` and continue. Refusing by name without probing was the
33
+ exact failure that left images unread in the field.
34
+ - Never guess at image contents. A non-visual "analysis" is worse than no analysis at all, because
35
+ it looks authoritative while being fabricated. Refusing loudly is always safer than guessing
36
+ quietly.
33
37
 
34
38
  ## 3. Input protocol (priority order)
35
39