openmerit 0.1.4 → 0.1.6-preview.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (147) hide show
  1. package/CHANGELOG.md +40 -0
  2. package/README.md +121 -386
  3. package/dist/core/src/index.d.ts +101 -0
  4. package/dist/core/src/index.js +1649 -0
  5. package/dist/core/src/store.d.ts +35 -0
  6. package/dist/core/src/store.js +102 -0
  7. package/dist/pi/src/index.d.ts +32 -0
  8. package/dist/pi/src/index.js +794 -0
  9. package/dist/pi/src/scheduler.d.ts +11 -0
  10. package/dist/pi/src/scheduler.js +137 -0
  11. package/dist/pi/src/wakeup.d.ts +2 -0
  12. package/dist/pi/src/wakeup.js +108 -0
  13. package/dist/protocol/src/index.d.ts +484 -0
  14. package/dist/protocol/src/index.js +47 -0
  15. package/dist/protocol/src/schemas.d.ts +576 -0
  16. package/dist/protocol/src/schemas.js +280 -0
  17. package/dist/terminal/public/app.js +297 -0
  18. package/dist/terminal/public/brands/anthropic.png +0 -0
  19. package/dist/terminal/public/brands/baai.png +0 -0
  20. package/dist/terminal/public/brands/baseten.png +0 -0
  21. package/dist/terminal/public/brands/cerebras.png +0 -0
  22. package/dist/terminal/public/brands/cohere.png +0 -0
  23. package/dist/terminal/public/brands/deepseek.ico +0 -0
  24. package/dist/terminal/public/brands/google.png +0 -0
  25. package/dist/terminal/public/brands/groq.ico +0 -0
  26. package/dist/terminal/public/brands/lm-studio.png +0 -0
  27. package/dist/terminal/public/brands/meta.ico +0 -0
  28. package/dist/terminal/public/brands/mistral.png +0 -0
  29. package/dist/terminal/public/brands/nomic.png +0 -0
  30. package/dist/terminal/public/brands/ollama.png +0 -0
  31. package/dist/terminal/public/brands/openai.png +0 -0
  32. package/dist/terminal/public/brands/openrouter.png +0 -0
  33. package/dist/terminal/public/brands/qwen.png +0 -0
  34. package/dist/terminal/public/brands/vllm.ico +0 -0
  35. package/dist/terminal/public/brands/vllm.png +0 -0
  36. package/dist/terminal/public/favicon.svg +1 -0
  37. package/dist/terminal/public/flow.css +1 -0
  38. package/dist/terminal/public/flow.js +770 -0
  39. package/dist/terminal/public/index.html +21 -0
  40. package/dist/terminal/public/styles.css +779 -0
  41. package/dist/terminal/src/activity-merge.mjs +64 -0
  42. package/dist/terminal/src/browser.mjs +29 -0
  43. package/dist/terminal/src/cli.mjs +60 -0
  44. package/dist/terminal/src/collect.mjs +311 -0
  45. package/dist/terminal/src/discovery.mjs +93 -0
  46. package/dist/terminal/src/hardware.mjs +57 -0
  47. package/dist/terminal/src/project-activity.mjs +156 -0
  48. package/dist/terminal/src/sample.mjs +171 -0
  49. package/dist/terminal/src/server.mjs +56 -0
  50. package/dist/terminal/src/services.mjs +62 -0
  51. package/dist/terminal/src/topology.mjs +30 -0
  52. package/docs/adapter-guide.md +189 -0
  53. package/docs/architecture.md +59 -0
  54. package/docs/automation.md +74 -0
  55. package/docs/budgets.md +37 -0
  56. package/docs/commands.md +85 -0
  57. package/docs/demo-backfill.md +29 -0
  58. package/docs/demo-fieldkit.md +47 -0
  59. package/docs/demo-placement.md +30 -0
  60. package/docs/demo-spam.md +15 -0
  61. package/docs/demo-support.md +42 -0
  62. package/docs/demo.md +57 -0
  63. package/docs/first-trial.md +60 -0
  64. package/docs/getting-started.md +65 -0
  65. package/docs/index.md +40 -0
  66. package/docs/inference-terminal.md +439 -0
  67. package/docs/lifecycle.md +30 -0
  68. package/docs/memo.md +126 -0
  69. package/docs/metrics-and-evidence.md +48 -0
  70. package/docs/operations.md +40 -0
  71. package/docs/pareto-spec.md +76 -0
  72. package/docs/pi-extension.md +54 -0
  73. package/docs/roadmap.md +28 -0
  74. package/docs/security.md +37 -0
  75. package/docs/site-artwork-linocut.md +23 -0
  76. package/docs/site-artwork-miniature-diverse.md +28 -0
  77. package/docs/site-artwork-miniature.md +26 -0
  78. package/docs/site-demo.md +177 -0
  79. package/docs/site-design.md +94 -0
  80. package/docs/site-documentation.md +83 -0
  81. package/docs/site-dynamic-og.md +35 -0
  82. package/docs/site-faq-maintenance.md +115 -0
  83. package/docs/site-hero-resolution.md +60 -0
  84. package/docs/site-illustration-sequences.md +227 -0
  85. package/docs/site-inference-terminal.md +203 -0
  86. package/docs/site-memo.md +39 -0
  87. package/docs/site-og-image.md +38 -0
  88. package/docs/site-og-workshop.md +21 -0
  89. package/docs/site-section-artwork.md +56 -0
  90. package/docs/site-skill-review.md +57 -0
  91. package/docs/site-terminal-preview.md +85 -0
  92. package/docs/testing.md +118 -0
  93. package/docs/troubleshooting.md +55 -0
  94. package/docs/ux-reference.md +32 -0
  95. package/package.json +74 -42
  96. package/benchmark/invoice_ocr/data/invoice_01_ground_truth.json +0 -38
  97. package/benchmark/invoice_ocr/data/invoice_01_row_2.jpg +0 -0
  98. package/benchmark/invoice_ocr/data/invoice_02_ground_truth.json +0 -32
  99. package/benchmark/invoice_ocr/data/invoice_02_row_5.jpg +0 -0
  100. package/benchmark/invoice_ocr/data/invoice_03_ground_truth.json +0 -26
  101. package/benchmark/invoice_ocr/data/invoice_03_row_6.jpg +0 -0
  102. package/benchmark/invoice_ocr/data/invoice_04_ground_truth.json +0 -26
  103. package/benchmark/invoice_ocr/data/invoice_04_row_7.jpg +0 -0
  104. package/benchmark/invoice_ocr/data/invoice_05_ground_truth.json +0 -38
  105. package/benchmark/invoice_ocr/data/invoice_05_row_947.jpg +0 -0
  106. package/benchmark/invoice_ocr/data/invoice_06_ground_truth.json +0 -38
  107. package/benchmark/invoice_ocr/data/invoice_06_row_948.jpg +0 -0
  108. package/benchmark/invoice_ocr/data/invoice_07_ground_truth.json +0 -20
  109. package/benchmark/invoice_ocr/data/invoice_07_row_949.jpg +0 -0
  110. package/benchmark/invoice_ocr/data/invoice_08_ground_truth.json +0 -38
  111. package/benchmark/invoice_ocr/data/invoice_08_row_1888.jpg +0 -0
  112. package/benchmark/invoice_ocr/data/invoice_09_ground_truth.json +0 -26
  113. package/benchmark/invoice_ocr/data/invoice_09_row_1890.jpg +0 -0
  114. package/benchmark/invoice_ocr/data/invoice_10_ground_truth.json +0 -20
  115. package/benchmark/invoice_ocr/data/invoice_10_row_1892.jpg +0 -0
  116. package/benchmark/invoice_ocr/data/manifest.json +0 -97
  117. package/dist/benchmarks.js +0 -98
  118. package/dist/catalog.js +0 -61
  119. package/dist/cli.js +0 -188
  120. package/dist/daemon.js +0 -407
  121. package/dist/diagnostics.js +0 -227
  122. package/dist/frontier.js +0 -56
  123. package/dist/harness.js +0 -1
  124. package/dist/integrations.js +0 -19
  125. package/dist/invoice-eval.js +0 -33
  126. package/dist/invoice-score.js +0 -124
  127. package/dist/judge.js +0 -43
  128. package/dist/llm.js +0 -207
  129. package/dist/pi-config.js +0 -46
  130. package/dist/pi-trials.js +0 -373
  131. package/dist/policy.js +0 -185
  132. package/dist/providers.js +0 -1
  133. package/dist/recommend.js +0 -76
  134. package/dist/routes.js +0 -74
  135. package/dist/standalone.js +0 -224
  136. package/dist/store.js +0 -89
  137. package/dist/strategist.js +0 -68
  138. package/dist/task-input.js +0 -54
  139. package/dist/traces.js +0 -127
  140. package/dist/trials.js +0 -140
  141. package/dist/types.js +0 -2
  142. package/examples/invoice-prompt.txt +0 -19
  143. package/examples/task.example.json +0 -7
  144. package/extension/openmerit.ts +0 -947
  145. package/instructions/OPENMERIT.md +0 -63
  146. package/instructions/openmerit.policy.json +0 -37
  147. package/rules.md +0 -43
@@ -0,0 +1,21 @@
1
+ # Workshop artwork for generated OG images
2
+
3
+ This asset is retained for possible future use. Separate page OG images are currently disabled; documentation uses the default homepage image.
4
+
5
+ The reusable background is `site/assets/openmerit-og-workshop-v1.png` (1200 × 630 PNG). It is derived from the default homepage OG image, `site/assets/openmerit-og-v1.png`, using the built-in imagegen tool to remove the existing headline while retaining the miniature makers, workshop, tools, and spaniel.
6
+
7
+ Generated source: `/Users/lazim/.codex/generated_images/01a0cd09-8dd8-73e0-806b-38ca331e4443/exec-86ca3e66-89ed-4866-aa68-2c1ff508a8cc.png` (1731 × 909). The source remains unchanged. Exported with `sips -z 630 1200` to the project asset. Background retouching used the built-in tool; the build adds the page heading separately.
8
+
9
+ The dynamic template places the maintained page heading over this background. The final generated card is `site/docs/og.png`. The separate OpenMerit logo, document icon, subtitle, and footer details have been removed to match the default card's composition. See [generation and build behavior](site-dynamic-og.md).
10
+
11
+ ## Exact built-in imagegen prompt
12
+
13
+ Use case: precise-object-edit.
14
+ Asset type: reusable background artwork for dynamically generated OpenMerit social sharing cards.
15
+ Input image 1 is the edit target: the existing default OpenMerit Open Graph image.
16
+
17
+ Remove the headline text "Good models prove themselves in the work." completely and reconstruct the clean warm paper background behind it. Produce a text-free background plate for the website's code to place different page titles on later.
18
+
19
+ Preserve everything else as closely as possible: the same three realistic miniature adult makers, their identities and clothing and poses, the East Asian man measuring on the left, the older South Asian man comparing hand planes near the center, the Black woman writing in a notebook on the right, the entire wooden bench and tools, brass balance, resting spaniel, all feet and bench legs, soft shadows, scale, framing and placement. Keep the existing photographic handcrafted miniature aesthetic and fine material detail. Keep the full scene in the lower portion of the image and preserve the large blank upper headline area.
20
+
21
+ The final canvas should have the same 1200:630 aspect ratio as the source. Seamless warm paper background #f8f7f4. No text anywhere, no OpenMerit logo or wordmark, no URL, no version, no icons, no page diagrams, no extra objects, no border, no footer line. Do not crop any person, dog, furniture, or tools. Change only the removal of text.
@@ -0,0 +1,56 @@
1
+ # Supporting miniature illustrations
2
+
3
+ Created with the built-in imagegen tool, using the current diverse miniature hero as the visual reference. Each final asset is a 768 × 512 WebP exported from a 1536 × 1024 source at quality 86. No new sections or interactions were added.
4
+
5
+ The illustrations use opaque white backgrounds and CSS multiply blending into the existing section colors. Initial transparency requests produced baked-in checkerboards; those drafts were corrected with imagegen and are not used by the website. The section assets are decorative (empty alt text), reserve their aspect ratio, and load lazily.
6
+
7
+ ## setup
8
+
9
+ Final asset: `site/assets/workshop-setup.webp`.
10
+
11
+ Final generated source: `/Users/lazim/.codex/generated_images/01a0c9d4-a3f3-7e12-aa03-2b46f42b5e83/exec-ba95a12e-f79d-4cdd-ba1d-5b98ac06129a.png`.
12
+
13
+ ### Initial generation prompt
14
+
15
+ > Use case: illustration-story / photographic miniature editorial vignette.
16
+ > Input image 1 is a STYLE AND MATERIAL REFERENCE ONLY: the existing OpenMerit miniature-workshop hero. Create ONE NEW SMALL SUPPORTING VIGNETTE, not a copy of the wide scene. Match its convincing physical miniature photography, tactile worn wood, brushed brass and blackened metal, linen and cotton, slight handmade imperfections, natural soft light, quiet muted color. The website will show this image only about 160–220 CSS pixels wide. Use a compact, clearly legible silhouette with a few purposeful objects; it must stay secondary to text.
17
+ > Format: 1536 × 1024 pixels, landscape 3:2, with a genuinely TRANSPARENT alpha-channel background. No white or colored rectangle, no ground plane, no vignette or scenery. A faint semi-transparent contact shadow directly beneath the subjects is okay. Isolated cutout. All objects entirely in frame, centered horizontally, modest padding at sides and above, baseline near the bottom. No words, text, numbers, marks resembling letters, labels, logos, watermark, decorative symbols, vector art, ink drawing, 3D plastic gloss, exaggerated cartoon proportions, or generic app-icon presentation.
18
+ > Lighting and palette: soft broad window light, low contrast, natural oak, linen, faded indigo and brass. No bright accents. This is a generated photographic miniature, with tangible material grain and realistic small construction details.
19
+ > SECTION: Getting started / preparing your tools.
20
+ > SCENE: A small open, weathered oak carpenter's toolbox with its hinged lid propped up. A well-used miniature hand plane rests partly inside; a little pair of brass calipers and a neatly folded oatmeal linen cloth are laid out beside the box, ready to begin work. Three objects plus the cloth, no clutter. Simple low three-quarter viewpoint, object arrangement wider than tall. No people, no dog, no plant, no extra scene. The specific feeling is preparation and careful selection, not a product display. A restrained companion to source-installation instructions.
21
+
22
+ ## faq
23
+
24
+ Final asset: `site/assets/workshop-questions.webp`.
25
+
26
+ Final generated source: `/Users/lazim/.codex/generated_images/01a0c9d4-a3f3-7e12-aa03-2b46f42b5e83/exec-683d1e35-8caf-4e28-9690-f43a6aaf57c2.png`.
27
+
28
+ ### Initial generation prompt
29
+
30
+ > Use case: illustration-story / photographic miniature editorial vignette.
31
+ > Input image 1 is a STYLE AND MATERIAL REFERENCE ONLY: the existing OpenMerit miniature-workshop hero. Create ONE NEW SMALL SUPPORTING VIGNETTE, not a copy of the wide scene. Match its convincing physical miniature photography, tactile worn wood, brushed brass and blackened metal, linen and cotton, slight handmade imperfections, natural soft light, quiet muted color. The website will show this image only about 160–220 CSS pixels wide. Use a compact, clearly legible silhouette with a few purposeful objects; it must stay secondary to text.
32
+ > Format: 1536 × 1024 pixels, landscape 3:2, with a genuinely TRANSPARENT alpha-channel background. No white or colored rectangle, no ground plane, no vignette or scenery. A faint semi-transparent contact shadow directly beneath the subjects is okay. Isolated cutout. All objects entirely in frame, centered horizontally, modest padding at sides and above, baseline near the bottom. No words, text, numbers, marks resembling letters, labels, logos, watermark, decorative symbols, vector art, ink drawing, 3D plastic gloss, exaggerated cartoon proportions, or generic app-icon presentation.
33
+ > Lighting and palette: soft broad window light, low contrast, natural oak, linen, faded indigo and brass. No bright accents. This is a generated photographic miniature, with tangible material grain and realistic small construction details.
34
+ > SECTION: Frequently asked questions / understanding together.
35
+ > SCENE: Two small adult-proportioned artisan figures from the reference's cast having a quiet practical conversation over an open cream-paper notebook on a very small plain oak worktable. A Black woman with dark brown skin and coiled hair in a bun, wearing faded indigo and an oatmeal linen apron, points to a blank notebook page; an older South Asian man with medium-brown skin, gray beard, round spectacles and an ochre knit vest sits on a short stool and listens thoughtfully. Both look at the notebook. One pencil lies on the table. Exactly two people, one small table, one stool and the notebook. Full bodies, feet and heads visible, hands anatomically believable. Small carved faces, visible stitched cloth, adult natural proportions, no cute large heads. Quiet colleague-to-colleague exchange, no lecture gestures. Compact grouping, not a long bench or panorama. No other people, tools or props.
36
+
37
+ ## footer
38
+
39
+ Final asset: `site/assets/workshop-rest.webp`.
40
+
41
+ Final generated source: `/Users/lazim/.codex/generated_images/01a0c9d4-a3f3-7e12-aa03-2b46f42b5e83/exec-f04adc05-0565-4469-9235-2429dcc6e901.png`.
42
+
43
+ ### Initial generation prompt
44
+
45
+ > Use case: illustration-story / photographic miniature editorial vignette.
46
+ > Input image 1 is a STYLE AND MATERIAL REFERENCE ONLY: the existing OpenMerit miniature-workshop hero. Create ONE NEW SMALL SUPPORTING VIGNETTE, not a copy of the wide scene. Match its convincing physical miniature photography, tactile worn wood, brushed brass and blackened metal, linen and cotton, slight handmade imperfections, natural soft light, quiet muted color. The website will show this image only about 160–220 CSS pixels wide. Use a compact, clearly legible silhouette with a few purposeful objects; it must stay secondary to text.
47
+ > Format: 1536 × 1024 pixels, landscape 3:2, with a genuinely TRANSPARENT alpha-channel background. No white or colored rectangle, no ground plane, no vignette or scenery. A faint semi-transparent contact shadow directly beneath the subjects is okay. Isolated cutout. All objects entirely in frame, centered horizontally, modest padding at sides and above, baseline near the bottom. No words, text, numbers, marks resembling letters, labels, logos, watermark, decorative symbols, vector art, ink drawing, 3D plastic gloss, exaggerated cartoon proportions, or generic app-icon presentation.
48
+ > Lighting and palette: soft broad window light, low contrast, natural oak, linen, faded indigo and brass. No bright accents. This is a generated photographic miniature, with tangible material grain and realistic small construction details.
49
+ > SECTION: Footer / the workshop at rest.
50
+ > SCENE: The same small brown-and-white spaniel-like workshop dog from the reference sleeps curled up, chin on its paws, beside a small CLOSED weathered oak toolbox. A soft folded linen work apron rests on the box lid. Exactly one sleeping dog, one small box, one folded apron. Low horizontal composition, dog on the left, box on the right, dog and box of comparable height. Softly observed miniature fur and quiet wood grain, subtly crafted rather than hyperreal animal advertising. The dog is fully asleep and relaxed. No people, no text, no extra props, no tall handle sticking upward, no floor or backdrop. An intimate, unobtrusive closing detail that still reads at 130px wide.
51
+
52
+ ## Background correction prompt applied to all three drafts
53
+
54
+ > Use case: precise-object-edit / background correction.
55
+ > Replace ALL of the checkerboard pattern in this image with a perfectly flat pure WHITE #FFFFFF background. This is an OPAQUE WHITE RGB export, not a transparency request. There must be absolutely no checkerboard squares, gray pattern, grunge, paper texture, borders or backdrop gradient anywhere, including gaps between and underneath the subjects. Edge pixels of the canvas must be pure #FFFFFF. Keep ONLY a very faint natural contact shadow immediately under the objects.
56
+ > Preserve the foreground miniature subjects EXACTLY: their placement, size, proportions, faces where present, clothing, every object, material texture, color and lighting. Do not change or invent any foreground details. Same 1536 × 1024 dimensions and framing. The finished image will be placed on a warm website background using multiply blending, so the blank regions must be clean #FFFFFF. No text or watermark.
@@ -0,0 +1,57 @@
1
+ # Focused design and motion review
2
+
3
+ ## Current illustration implementation
4
+
5
+ On 2026-09-23 the user rejected the displacement effect because it visibly stretched a single illustration. It has been removed and replaced with [separately generated pose sequences](site-illustration-sequences.md). The earlier filter reviews below are historical and do not describe the current player.
6
+
7
+ | Before | After | Why |
8
+ | --- | --- | --- |
9
+ | SVG displacement stretched the same pixels to simulate gestures. | Canvas playback selects complete, separately illustrated poses from a sprite atlas. | Hands, tools, and bodies change their actual pose and silhouette. |
10
+ | A rendering loop ran at approximately 30Hz. | Timers wake only at illustrated frame boundaries, with holds at each end. | Deliberate stop-motion pacing with fewer redraws. |
11
+ | Pausing reset the illustration to the source image. | Pausing freezes the current pose; reduced motion uses a matching static poster. | Pause behaves as expected while preserving the static accessibility fallback. |
12
+
13
+ ## Earlier implementation
14
+
15
+ Reviewed 2026-09-22 using three selected skills from [emilkowalski/skills](https://github.com/emilkowalski/skills), installed with Codex's skill-installer:
16
+
17
+ - [emil-design-eng](https://github.com/emilkowalski/skills/blob/main/skills/emil-design-eng/SKILL.md): interaction decisions, restraint, and press feedback.
18
+ - [review-animations](https://github.com/emilkowalski/skills/blob/main/skills/review-animations/SKILL.md): review of existing motion, including frequency, easing, and rendering cost.
19
+ - [mobile-native](https://github.com/emilkowalski/skills/blob/main/skills/mobile-native/SKILL.md): capability-based hover styles and touch feedback.
20
+
21
+ Installed under `/Users/lazim/.codex/skills/`. The rest of the collection was not needed for this static marketing page. No UI or animation library was added.
22
+
23
+ ## Findings and applied changes
24
+
25
+ | Before | After | Why |
26
+ | --- | --- | --- |
27
+ | Link and button hover effects applied on every input type. | All hover effects use `(hover: hover) and (pointer: fine)` in [styles.css](/Users/lazim/Documents/GitHub/openmerit/site/styles.css:97). | Touch should not inherit a persistent hover treatment. |
28
+ | No custom press feedback after browser tap highlighting was suppressed. | The primary CTA scales to 0.97 over 160ms and returns over 100ms; other controls use immediate color/opacity feedback. See [styles.css](/Users/lazim/Documents/GitHub/openmerit/site/styles.css:24). | Make pointer-down visible while keeping the page quiet. The curve comes directly from the skill's standards. |
29
+ | Opening an FAQ replayed a decorative scene, including when using the keyboard. | FAQ activation only reveals the answer; keyboard activation changes the indicator instantly. Illustration playback is independent of FAQ interaction. See [illustrations.js](/Users/lazim/Documents/GitHub/openmerit/site/illustrations.js) and [styles.css](/Users/lazim/Documents/GitHub/openmerit/site/styles.css:84). | Repeated reading actions do not need a fresh decorative performance. |
30
+ | Hover/indicator movement used the generic `ease` curve. | Short 160ms transitions use `cubic-bezier(.23, 1, .32, 1)`. | Responsive, explicit timing for small UI changes. |
31
+ | SVG displacement attributes were written on every display frame, including unchanged values. | Writes are limited to about 30 updates/second and unchanged values are skipped in [illustrations.js](/Users/lazim/Documents/GitHub/openmerit/site/illustrations.js:155). | Reduce paint work on high-refresh displays without altering the slow gestures. |
32
+
33
+ ## Verification
34
+
35
+ - Desktop Chromium at 1440px: CTA press reached scale 0.97 and returned to its initial state; keyboard FAQ activation worked with no indicator transition or illustration replay; no runtime errors.
36
+ - Touch emulation at 390px: hover and fine-pointer queries were false; hovering caused no arrow movement; taps opened and closed the menu and FAQ; no horizontal overflow; body copy remained selectable.
37
+ - Reduced motion on initial load produced no illustration layers or entrance animation.
38
+ - A sampled hero gesture produced 26 attribute updates over one second, versus writing every display frame previously.
39
+ - JavaScript syntax and whitespace checks passed.
40
+
41
+ ## Verdict
42
+
43
+ **Performance:** The photographic illustrations still use painted SVG displacement rather than compositor-only transforms. There is no equivalent transform-only change that preserves local breathing and hand/head movement in the flattened artwork. These effects remain finite, localized, throttled, and disabled offscreen or for reduced motion; this is a deliberate tradeoff, not a claim of GPU-only rendering.
44
+
45
+ **Accessibility and mobile:** Keyboard behavior, capability queries, and reduced-motion behavior were checked in the browser. Physical-phone touch feel, Safari tap behavior, and performance on older hardware have not been verified.
46
+
47
+ **Approve the targeted code changes based on the code and browser checks above.** Physical-device validation remains a separate, unverified check.
48
+
49
+ Published to the existing Cloudflare Worker as version `59b9ebf7-092e-4902-a55f-c7ebb71da8ea`. The updated CSS and illustration script on both public domains match the local files; the production page loads and starts the hero animation without runtime errors.
50
+
51
+ ## Playback follow-up
52
+
53
+ The user found the motion difficult to see. The initial one-pass behavior also permanently stopped a scene if it left the viewport mid-gesture. Playback now alternates finite gestures with scene-specific rest periods, restarts on re-entry, and has a persistent footer pause control. The hero gestures are slightly stronger. Reduced motion remains respected, and FAQ activation remains independent of the decorative scene.
54
+
55
+ Actual before/after frame comparisons confirmed visible changes in every scene, rather than checking animated attributes alone. Repeat cycles, interrupted re-entry, pause/resume persistence, reduced motion, and 320px/390px overflow checks passed.
56
+
57
+ Follow-up deployment: `2ec895d2-f206-4d30-aa63-6a673fd602b8`. The live HTML, CSS, and illustration script on both domains match the local files. A complete play–rest–replay cycle and the footer dog playback were verified on the production page with no runtime errors.
@@ -0,0 +1,85 @@
1
+ # Hosted Inference Terminal sample
2
+
3
+ This separate preview is configured for `https://dash.openmerit.site/`. It uses the isolated terminal branch and its own Cloudflare Worker, `openmerit-terminal-preview`. It does not deploy `wrangler.jsonc`, the main website, or the interactive inference demo.
4
+
5
+ ## Build and deploy
6
+
7
+ ```sh
8
+ npm ci
9
+ npm run verify
10
+ npm run deploy:dash-preview
11
+ ```
12
+
13
+ The deploy command pins Wrangler 4.136.3 and uses `wrangler.terminal.jsonc`. Its build hook rebuilds the terminal preview, runs documentation checks, and checks/tests the terminal workspace before upload. The exact `dash.openmerit.site` custom domain uses the existing Cloudflare account. No storage, container, AI, secret, or external service binding is configured. The only binding serves bundled public assets from `dist/terminal-preview/`.
14
+
15
+ `scripts/build-terminal-preview.mjs` creates this asset directory from the same browser bundle and icons as the local terminal. It adds preview metadata, noindex, and a sample-only bootstrap marker; it never copies a project snapshot. `packages/terminal/hosted/worker.mjs` imports only the fictional sample generator, never the collector or local CLI. The Worker reuses each generated sample for its 15-second time bucket. Browser polling and reduced-motion behavior remain the same as the local interface.
16
+
17
+ `/api/info` selects sample mode and identifies the hosted preview. `/api/sample` exposes the fictional environment; other API paths return 404, and mutations return 405. The visible sample disclosure remains present, and the workspace picker offers no real environment. Static assets and API responses carry a same-origin content policy and indexing restrictions. HTML responses use `no-transform` so the zone’s automatic analytics script injection cannot alter the preview HTML. Script and stylesheet responses allow Cloudflare compression to reduce transfer time; main-site analytics settings are unchanged. `/robots.txt` also requests exclusion. This is public link sharing, not password protection.
18
+
19
+ ## Package delivery
20
+
21
+ The terminal ships within the existing `openmerit` npm package. `npm pack` and `npm publish` use the root manifest and compiled core, protocol, Pi adapter, and terminal under `dist/`. The file allowlist excludes the hosted sample, website source, repository tooling, dependencies, and local evidence. Browser assets are prebuilt; registry installs do not require a repository checkout. Node engine requirements, runtime dependencies, exports, and the legal license file remain intact.
22
+
23
+ For friends' testing, publish a prerelease with the `preview` dist-tag while `latest` remains the stable version. Run `npm run verify`, inspect `npm pack --dry-run`, install the actual tarball in an isolated npm prefix, and validate existing application activity before `npm publish --tag preview`. Confirm the registry version and both tags, then repeat installation with `npm install -g openmerit@preview`. The first registry preview is `0.1.6-preview.0`; confirm publication and install evidence in the deployment record before announcing availability.
24
+
25
+ The earlier domain-hosted package was an interim delivery method. Its deployment record below is historical. The source build no longer produces installer downloads; deploy the sample only after npm delivery is verified so the already-shared interim URL remains usable during the transition.
26
+
27
+ ## Sharing image
28
+
29
+ The page title and social preview title are `Inference Terminal by OpenMerit`. `packages/terminal/hosted/social-image.svg` is the editable 1200 × 630 artwork. `scripts/terminal-social-image.mjs` renders it with the bundled Inter font; the preview build writes a content-hashed PNG and its absolute HTTPS URL into Open Graph and Twitter large-card metadata. Rendering needs no network or system fonts. The card and its metadata deploy with the preview assets. The preview robots file permits named link-preview crawlers while keeping the general exclusion and noindex headers; its API remains disallowed to those crawlers.
30
+
31
+ ## Validation
32
+
33
+ The terminal test suite covers hosted API boundaries, sample identity/counts/freshness, asset headers, and HEAD requests. Before declaring a deployment complete, verify the live custom domain over HTTPS, all six tabs on desktop/mobile, inspector and selector interactions, the sample disclosure, absence of a live-environment switch, read-only API behavior, local asset loading, and no runtime errors. Check the live page title, Open Graph/Twitter image URLs, PNG content type and dimensions when changing the sharing card. A successful local build alone is not publication.
34
+
35
+ ## Deployment record
36
+
37
+ Published on 7 October 2026 as Worker version `f9c37164-aa44-4a84-9c09-e61272bb87c3`, at [dash.openmerit.site](https://dash.openmerit.site/). The earlier first upload was `fa5114d4-fa78-47c8-98fc-085d8b0e58e9`; the final version adds the scoped `no-transform` response header after live browser testing detected automatic zone-level beacon injection.
38
+
39
+ `npm run verify` passed 158 tests plus syntax, typecheck, build, demo, and documentation checks. Both deployments repeated the preview build, documentation checks, and all 48 terminal tests. The public HTTPS page returned 200 and matched the built HTML; its sample API returned 40 agents, 140 tasks, 21 models, and 9,600 fictional requests. The live snapshot API returned 404, and POST to the sample API returned 405.
40
+
41
+ All six tabs were checked at 1440 px and 390 px with no page overflow. Keyboard agent selection, agent and model inspectors, the fictional-data disclosure, and absence of a live-environment switch passed. The final browser check recorded no JavaScript errors, console errors, or external requests. Browser TLS verification used the published Cloudflare addresses while this machine's resolver retained its earlier negative DNS result; Cloudflare and Google public resolvers returned the new domain's A records. Physical phones were not tested.
42
+
43
+ Only the separate terminal Worker and its 19 public assets were published. No main-site deployment, npm release, provider call, upload, or local environment exposure was introduced. Deployment and verification logs are retained locally as `/private/tmp/terminal-hosted-deploy.log` and `/private/tmp/terminal-hosted-verify.log`; the final screenshot is `/private/tmp/terminal-hosted-final.png`.
44
+
45
+ ### Sharing card update — 7 October 2026
46
+
47
+ Published Worker version `09a8e6ad-1185-4e27-9252-78e368755949`. The 1200 × 630 PNG is 41,977 bytes and is served as `/inference-terminal-og-2330cc415994.png`. Live page title, Open Graph, and Twitter card metadata use `Inference Terminal by OpenMerit`. Requests with browser, Twitterbot, Facebook, WhatsApp, and Slack link-preview user agents returned 200 for both HTML and image; the PNG matched the local build byte for byte. Named preview crawler rules permit the page and image, but disallow API crawling. These checks verify origin responses, not the state of third-party preview caches.
48
+
49
+ Full verification passed before deployment. The deployment hook repeated documentation checks and all 48 terminal tests after the crawler-rule adjustment. A fresh Chrome session loaded the public domain using normal DNS, rendered the overview, and reported no JavaScript errors. Logs: `/private/tmp/terminal-og-verify.log`, `/private/tmp/terminal-og-deploy.log`, and `/private/tmp/terminal-og-live.json`.
50
+
51
+ ### Service identity update — 8 October 2026
52
+
53
+ The sample now includes fictional Baseten, Groq, and Cerebras connections with separately declared model creators. Its catalog remains 21 models and 27 deployments, with 14 shared pools (12 recent, one stale, one missing). It illustrates environment-variable evidence, shared and dedicated serving modes, and unreported cloud hardware. Official service icons are bundled locally. The hosted Worker still imports only the fictional sample generator; it does not run environment detection or collect visitor data.
54
+
55
+ Before publication, the real collector was exercised against an isolated synthetic project with fake inherited credentials, an OpenAI-compatible endpoint override, activity metadata, deployment configuration, and an accelerator export. Groq, Cerebras, and Baseten were identified without labeling the overridden OpenAI client as OpenAI. Credential values and raw URLs were absent from the snapshot. Cloud deployments retained unreported hardware. Browser checks covered all six sample tabs at 1440 px and 390 px, separate creator/service details, configuration evidence, and locally served icons, with no overflow or JavaScript errors.
56
+
57
+ Published at `dash.openmerit.site` as Worker version `27295601-f3a5-4c32-a6f9-68b71911b40f`. Full verification passed 163 tests; deployment repeated all 53 terminal tests and documentation checks. Live HTTPS verification confirmed the three service icons, configuration evidence, creator/service separation, and sample counts. The snapshot API remains 404. The browser reported no JavaScript errors or external requests. Logs are `/private/tmp/terminal-services-final-verify.log` and `/private/tmp/terminal-services-deploy.log`; the live inspector screenshot is `/private/tmp/terminal-services-live.png`. The isolated collector test used synthetic credentials and metadata, not customer cloud accounts.
58
+
59
+ ### Services inventory update — 8 October 2026
60
+
61
+ Services replaces Connections while keeping six top-level tabs. Its default inventory groups the fictional environment's 18 services into five Model APIs, three Inference clouds, four Routers, three Gateways, and three Runtimes. Unclassified entries remain available when a source does not establish a type. The environment map is a secondary view; old `#connections` links still open it. Service details link models, deployments, reported paths, explicitly associated limits, and identification evidence. Creators alone never create API connections.
62
+
63
+ Local verification passed 171 tests, including eight new service-aggregation cases, plus syntax, typecheck, build, demo, and documentation checks. Browser verification covered type filters, search/reset, service-to-model and source details, command search, the overview map link, legacy hashes, keyboard focus across passive refreshes, and all six tabs at 1440 px and 390 px. A synthetic browser fixture verified configured-only and unclassified entries with no activity, redacted endpoint identity, and an inspector without horizontal overflow on an emulated touch screen. No JavaScript errors or external requests were observed. Physical phones and customer cloud accounts were not tested. Logs: `/private/tmp/terminal-service-inventory-verify.log`; screenshots: `/private/tmp/terminal-service-inventory-desktop.png` and `/private/tmp/terminal-service-inventory-detail.png`.
64
+
65
+ Published at `dash.openmerit.site` as Worker version `be6cbecd-0094-42e9-943d-58d63158dabd`. The deployment hook repeated documentation checks and all 61 terminal tests. Live HTTPS checks confirmed the Services tab, 18 service rows, all three inference clouds, service/source inspectors, the secondary map, and no page overflow at 390 px. The title and fictional-data boundary remain intact: `/api/snapshot` returns 404, sample mutations return 405, and there is no local-environment switch. No browser or console errors or external requests were observed. Deployment log: `/private/tmp/terminal-service-inventory-deploy.log`; live screenshot: `/private/tmp/terminal-service-inventory-live.png`. Only the separate terminal preview was deployed.
66
+
67
+
68
+ ### Application scope and complete maps — 9 October 2026
69
+
70
+ Published source commit `43ad7c5` at `dash.openmerit.site` as Worker version `a16b01be-b280-448b-9dbe-503b20a1dfdd`. This includes the complete grouped environment map, the separately bundled `?map=flow` experiment, and application-only source guidance. The sample still contains 9,600 fictional requests, 40 agents, 140 tasks, and 21 models. Local application collectors and VM data are not included in the Worker.
71
+
72
+ The first upload (`1cd39eb1-a394-43e6-9a1e-573910400108`) exposed a public-network loading problem: a compressed sample transfer took about 13 seconds while the browser used a 5-second localhost deadline. The final build allows 30 seconds for hosted sample refreshes, preserving 5 seconds for local snapshots and the existing single-request guard. This is a response deadline, not a change to the 15-second polling cadence or data retention. A follow-up fixes script/style headers so Cloudflare can compress the larger optional map bundle; HTML retains `no-transform` to prevent zone-level analytics injection.
73
+
74
+ The deployment hook passed all 92 terminal tests, 11 documentation checks, and syntax checks. Live HTML, app JavaScript, styles, and both map assets matched the built files byte for byte. The sample-only info and API counts were verified over HTTPS; `/api/snapshot` returned 404 and sample mutations returned 405. All six default tabs passed browser checks at desktop and 390-pixel widths without page overflow. Agent selection represented the complete focused request count, the agent inspector opened, and the environment selector offered only the fictional sample. The optional map represented all 9,600 requests, workload search returned its recorded member, and its 390-pixel view had no page overflow. A fresh browser load and subsequent refresh succeeded with no console errors. The compressed map script transferred 585,941 bytes instead of 1,934,755 bytes and matched the built file after decoding. One Node HTTP check timed out; the HTTP/2 transfer and fresh browser load succeeded.
75
+
76
+ Deployment log: `/private/tmp/terminal-application-scope-compressed-deploy.log`. Asset and API evidence: `/private/tmp/terminal-application-scope-live.json`. Live screenshot: `/private/tmp/terminal-application-scope-live.png`. The published npm package and main website were not changed.
77
+
78
+
79
+ ### Friends’ install download — 9 October 2026
80
+
81
+ Published source commit `9c62859` as Worker version `9453d410-b7bc-4719-a0e7-9dc4139049c4`. The public test download is `openmerit@0.1.5-dash.d1288465e28b`, 1,026,220 bytes, with SHA-256 `0fd079d4ab241d8e83010ad147cbe6559c39aa37fcd803fff57afbb06b60b13d`. The stable npm release and main website were not changed. Only two new static assets were uploaded: the package and its build information; the sample's browser assets and Worker collection boundaries are unchanged.
82
+
83
+ Full verification passed 205 tests plus syntax, typecheck, build, and documentation checks. The deployment hook repeated all 95 terminal tests and 11 documentation checks. A disposable VM npm prefix installed directly from `https://dash.openmerit.site/install/openmerit-preview.tgz` in seven seconds. Its installed `openmerit dash`, run from the existing application directory, retained all 240 requests, 40 worker IDs, and 120 task IDs across a passive refresh, with 111 evidence files unchanged and zero inference calls. The public package metadata matched the deployed build. Local macOS acceptance verified actual default-browser opening; headless Linux acceptance verified the printed URL. Windows desktop opening remains unit-tested only.
84
+
85
+ Logs: `/private/tmp/terminal-friends-verify.log`, `/private/tmp/terminal-friends-deploy.log`, and `/private/tmp/terminal-friends-public-install.log`. Default-browser proof: `/private/tmp/terminal-friends-browser-open.png`.
@@ -0,0 +1,118 @@
1
+ # Testing
2
+
3
+ ## Standard verification
4
+
5
+ ```sh
6
+ npm install
7
+ npm test
8
+ npm run check
9
+ npm run verify
10
+ npm pack --dry-run
11
+ ```
12
+
13
+ Tests cover protocol contracts, lifecycle decisions, post-build target detection, incumbent identity, baseline sufficiency, durable candidate handoffs, candidate/frontier/swap sequencing, conservative Pareto verification, persistence, audit events, safe evidence identifiers, Pi commands, private intent delivery, evidence requirements, and mutation boundaries.
14
+
15
+ The routine `packages/pi/test/live-wakeup.test.ts` test uses a disposable project to verify that a due, insufficient baseline is checked and recorded without starting Pi. A prior opt-in live OpenRouter run proved the former harness-driven assessment could issue an intent and call tools, but Pi did not complete it within the bounded test. Candidate discovery and model changes still require a representative headless end-to-end trial before claiming unattended operation.
16
+
17
+ Automation tests additionally cover task and elapsed cadence, idempotent target-bound signal delivery, durable pending work, catalogue changes, metric regression thresholds, native wakeup requests, unsupported-scheduler fallbacks, delayed post-swap verification, and Pi dispatch from explicit application-task signals. Pi's own settled-turn telemetry is not a product-observation test path.
18
+
19
+ ## Inference Terminal verification (unreleased)
20
+
21
+ `npm test --workspace=@openmerit/terminal` tests metadata projection, unknown/redacted values, incremental UTF-8 imports, stable-ID updates, rotation, truncation, project boundaries, source precedence, local runtime contracts, aggregate counter scope, CLI arguments, sample isolation, chart sample gaps and inspection, and the loopback read-only API. Local HTTP fixtures prove which endpoints the collectors call; they do not prove every external runtime version is compatible.
22
+
23
+ The source UI was also reviewed in a browser at desktop and 390-pixel widths, including all six views, model search and location filters, time windows, node inspection, the command palette, Escape dismissal, and a real empty project. The full `npm run verify` command includes terminal tests and syntax checks through the workspace scripts. See the [terminal guide](inference-terminal.md) for current collection coverage and limitations.
24
+
25
+ The enterprise sample tests assert 40 reported agents, 140 tasks, 21 used models, 9,600 unique requests, consistent task-to-agent assignments, all path classifications, and unknown measurements on running requests. Browser checks cover the larger inventories, agent/task lists and nested details, command search, map coverage, and desktop/mobile layouts. This fixture exercises interface density; it does not establish production throughput or hardware performance.
26
+
27
+ The sidebar and animation revision was additionally checked for persistent chart/DOM identity across polls, CPU/memory selection, exact-sample keyboard inspection, compact navigation, Torph transitions, immediate keyboard updates, and static charts under reduced motion. These are browser checks, not a committed browser automation suite. Browser source lives in `packages/terminal/client/`; `scripts/build-terminal.mjs` bundles it into the ignored `public/app.js` before `npm run dash` and during the package build.
28
+
29
+ Path-category tests cover explicit metadata precedence, unknown/invalid types, case-insensitive inventory/source matching, conflicting declarations, and mixed paths without double counting. They verify that endpoint location alone never establishes a request path. Sample data exercises direct, local, router, and gateway paths, including one model reached through multiple types.
30
+
31
+ Deployment and capacity tests cover strict identity matching, duplicate-ID rejection, ordered path projection and entry conflicts, shared-pool deduplication, stale/future/missing/zero readings, positive-limit ratios, safe project-local capacity files, export failure clearing, and highest-series vLLM cache usage. The enterprise fixture checks 27 deployments, 12 pools, 9,479 identified requests, and explicit stale/missing capacity coverage.
32
+
33
+ ## Pi load smoke test
34
+
35
+ ```sh
36
+ ./node_modules/.bin/pi --no-session --no-extensions -e ./packages/pi/src/index.ts --list-models
37
+ ```
38
+
39
+ This proves the real Pi runtime loads the extension and enumerates its available catalogue. It does not prove a provider can complete an end-to-end product intent.
40
+
41
+ CI also packs `openmerit`, installs the tarball in a temporary directory,
42
+ imports its public core and protocol entry points, and loads the installed Pi
43
+ extension. This checks the published artifact rather than only the workspace.
44
+
45
+ ## Evidence levels
46
+
47
+ - **Unit verified** — deterministic behavior has automated coverage.
48
+ - **Runtime smoke verified** — the actual Pi runtime loads the extension.
49
+ - **Product trial verified** — Pi creates evals and observability in a real product and completes an intent with evidence.
50
+ - **Production verified** — representative repeated runs support reliability and distribution claims.
51
+
52
+ The current repository is unit- and runtime-smoke-verified. Product and production claims require target-product evidence.
53
+
54
+ ## Experimental Pi/OpenMerit sandbox demo
55
+
56
+ This repository includes a disposable, live-provider demo for showing the complete experimental loop. It is not the shipping product workflow and its small evaluation set is not evidence for reliability claims.
57
+
58
+ Requirements: Node.js 22.19+, Pi 0.87.x with OpenRouter coding-harness authentication, and an OpenRouter application credential available as `OPENROUTER_API_KEY` or in Pi's local `auth.json`. The coding-harness credential is not passed to the generated application; the demo runner supplies the application credential only while making provider calls.
59
+
60
+ ```sh
61
+ npm run demo
62
+ ```
63
+
64
+ Pi builds a small sentiment-classification app and infers a task/evaluation plan in a temporary directory. The runner validates the plan against available collectors, shows the selected metrics and rationales, smoke-tests the app's configurable route through a live application call, and passes the inferred profile to OpenMerit's coordinator. Task quality is scored only by exact sentiment-label match; rationale plausibility is not graded. The runner collects a baseline, randomly interleaves incumbent/challenger calls over Pi's same three to five representative cases with two repetitions per case, records provider-reported token use and request cost plus measured latency and task outcomes, asks OpenMerit to calculate and verify the Pareto frontier, applies a selected frontier route only in that temporary app, and runs a new post-swap check for every case/repetition. It stops on missing model availability, missing provider usage/cost, invalid output, or a $1 default application-provider spend threshold checked after each response. The saved summary includes its randomization seed.
65
+
66
+ The default baseline is `openai/gpt-4.1-mini`, with `openai/gpt-4o-mini` as the challenger. Override the selection with `OPENMERIT_DEMO_BASELINE_MODEL` and comma-separated `OPENMERIT_DEMO_CANDIDATE_MODELS`; candidates must be available in the current OpenRouter catalogue. `OPENMERIT_DEMO_PI_MODEL` selects the Pi coding model. `OPENMERIT_DEMO_REPETITIONS` accepts 2–5, `OPENMERIT_DEMO_MAX_SPEND_USD` accepts a cap up to $5, and `OPENMERIT_DEMO_SEED` supplies a repeatable ordering seed.
67
+
68
+ The temporary directory path is printed at startup and completion. It contains the generated app and inferred plan, `.openmerit/` lifecycle state and redacted event log, per-call `artifacts/application-runs.jsonl`, comparison manifest, and summary. The runner executes application calls serially. Its spend threshold measures application calls only; Pi's coding-model spend is not measured or included. Its tiny task set and sample count are a mechanics demo only: do not present its observed rates, average latency, or frontier as representative reliability, quality, or performance claims.
69
+
70
+ ### Inference Terminal hardware checks (unreleased)
71
+
72
+ `npm test --workspace=@openmerit/terminal` covers bounded and project-contained hardware exports, duplicate/reserved host IDs, future and stale timestamps, missing versus empty residency, per-host history deduplication/reset, sample placement, and the hashes of locally bundled official favicons. Browser verification covers path-series selection and keyboard bin inspection, enlarged table headers, host and location selection, separate host chart histories, stale/missing hardware, exact-name connection icons, and narrow-screen layouts. These checks exercise passive metadata; they do not run inference or discover remote hardware.
73
+
74
+ The terminal identity checks additionally cover all 21 sample models across inventory, requests, deployments and residency; official versus arbitrary model namespaces; unknown-name fallbacks; runtime/maker separation; and consistent HTML/SVG asset selection. Browser checks cover model and service marks on all six views, search results, deployment/path/source/model inspectors, and loaded-model rows, plus image loading without external requests, mobile overflow, and nonduplicated accessible names.
75
+
76
+
77
+ The terminal naming tests cover source-supplied IDs without display names, absent workload labels, configured-address fallback for unnamed runtimes, full system hostnames, exported host name/hostname/ID precedence, and missing site/rack metadata. Browser checks cover colored all-path stacks and isolated paths, readable activity subtitles, compact Sources rows, name provenance inspectors, and host filters that appear only for reported sites. Sample naming is explicitly fictional.
78
+
79
+
80
+ Top-navigation browser checks cover all six tabs, manual arrow/Home/End activation, panel labels, direct hash views, source coverage, environment switching, command search, sticky positioning, responsive tab overflow, and focus/scroll preservation through a passive poll. The old sidebar and collapse controls are absent. Desktop and narrow layouts retain readable content with no page-level horizontal overflow.
81
+
82
+ ### Installed terminal smoke test
83
+
84
+ The package-smoke CI job installs the npm tarball into a fresh directory, then runs `scripts/terminal-package-smoke.mjs` against its installed executable. It checks the CLI, packaged browser assets, cached passive collection, request-update deduplication, source loss/recovery, sample separation, read-only HTTP boundary, and shutdown. Its default runtime is a deterministic fixture; that test makes no inference calls.
85
+
86
+ For a separate real-runtime check, supply a disposable Ollama server with `smollm2:135m` already installed and unloaded:
87
+
88
+ ```sh
89
+ node scripts/terminal-package-smoke.mjs /path/to/install/node_modules/.bin/openmerit --ollama-url http://127.0.0.1:11434 --output /tmp/terminal-evidence
90
+ ```
91
+
92
+ This opt-in test application makes four real local generation calls. The terminal itself still reads only inventory endpoints and a temporary metadata export. Two requests are exported with Ollama's actual token and timing measurements. A separate, explicitly synthetic running-to-completed record tests update semantics. A generation without an export must produce no invented request or agent history. The test removes its temporary project and stops its terminal process; the caller owns the disposable runtime's lifecycle. The JSON report excludes prompts and generated text. No cloud credentials are needed.
93
+
94
+ ### Automatic application discovery (unreleased)
95
+
96
+ Terminal tests cover bounded literal project declarations, credential omission, SDK hints without invented usage, canonical project containment, malformed/oversized inputs, safe loopback endpoints, runtime discovery without configuration, outage handling, and no added polling between discovery passes for never-seen endpoints. Application request logs and explicit exports reconcile without duplicating requests. Application agent identities remain intact. Records explicitly marked as development usage are rejected.
97
+
98
+ A normal CLI fixture starts with application logs, valid legacy Pi history grants, native Pi history, and other coding-tool history directories. It verifies that only application requests contribute to the snapshot and that OpenMerit `product_task_completed` summaries still work when request-level evidence is absent. Retired access flags give an explanatory error; sample JSON bypasses collectors. The installed-package smoke covers runtime/request lifecycle and read-only HTTP boundaries.
99
+
100
+ The earlier history-based checks below document a superseded preview. Pi history is no longer a supported terminal source.
101
+
102
+ ### Historical VM check with Pi history (superseded)
103
+
104
+ On 8 October 2026, the preview was built and installed in a dedicated Ubuntu 24.04.5 ARM64 VM with four virtual CPUs and 4 GiB of RAM, using Node 22.23.1, Ollama 0.40.1, and Pi 0.87.0. Eighteen actual Pi requests ran across `smollm2:135m`, `qwen2.5:0.5b`, and `smollm2:360m`. After explicit project-scoped history consent, the installed dashboard collected all 18 completed requests and their reported 1,402 input and 587 output tokens directly from native Pi sessions. No activity export supplied those requests.
105
+
106
+ This was a controlled local-model workload, not a production-traffic test. Pi's configured local-model rates were zero; displayed cost does not measure hardware or electricity. In that initial build, route names remained unknown because the adapter omitted Pi’s recorded provider connection ID. Deployments, agent/task identities, and request durations were not reported by the history. The VM had no host-directory mounts or forwarded host credentials, and its dashboard was forwarded only to host loopback.
107
+
108
+ The real browser check exposed an overview coverage check that omitted native history. A regression test then covered Pi-only usage, connected empty sources, retained partial history, and inventory without request evidence; it now covers application logs instead. The same count check is used by the overview, path summary, and model table.
109
+
110
+
111
+ ### Historical Pi connection identity (superseded, 9 October 2026)
112
+
113
+ That preview adapter preserved Pi's recorded provider connection ID as the route name. Its regression coverage included Ollama, an arbitrarily named local runtime, a declared gateway, unknown connection types, conflicting declarations, and an end-to-end approved-history collector check. Generic activity records still do not acquire a route from their serving-provider name. These checks do not infer physical deployments or reconstruct historical endpoints from mutable credentials or current Pi model configuration.
114
+
115
+
116
+ ### Application-only VM acceptance (unreleased, 9 October 2026)
117
+
118
+ The updated installed CLI was run as `openmerit dash` from the VM application directory with no terminal configuration or custom activity feed. All 240 existing application requests, 40 worker identities, and 120 task identities remained unchanged; the 36 earlier Pi-history requests were excluded. The default server, JSON CLI, next passive poll, and private dashboard agreed. Application log hashes and the legacy grant file were unchanged, and runtime inventory remained available. No new inference was performed. This supersedes the Pi-history acceptance scope above and verifies supported existing application logs, not universal discovery.
@@ -0,0 +1,55 @@
1
+ # Troubleshooting
2
+
3
+ Start with the current status and the local audit trail. These show what OpenMerit requested, what the harness reported, and which evidence references were recorded.
4
+
5
+ ```text
6
+ /openmerit status
7
+ /openmerit doctor
8
+ /openmerit logs
9
+ ```
10
+
11
+ ## The command is missing
12
+
13
+ Check that OpenMerit is installed and enabled in Pi:
14
+
15
+ ```sh
16
+ pi list
17
+ ```
18
+
19
+ The published extension targets Pi 0.87 and Node.js 22.19 or newer. Follow [getting started](getting-started.md) to install it, then start Pi in your product repository.
20
+
21
+ ## Setup did not begin
22
+
23
+ Automatic setup requires an unconfigured project, an interactive Pi UI, and a bounded source-detection match for an application LLM call. Non-interactive runs do not start the confirmation flow, and a project without a detected target remains idle until `/openmerit setup` is requested. Existing project state or active work may also prevent a fresh setup request.
24
+
25
+ Inspect status first. If no intent is active, use `/openmerit setup` to request setup explicitly; this is also the supported correction path when source detection or a prior setup recorded the wrong incumbent. A new setup preserves audit history but replaces the active task profile, candidate/assessment handoffs, frontier, and baseline state after the new result is accepted. Do not delete `.openmerit/` as a routine reset: it contains your policy, lifecycle state, evidence references, and audit history.
26
+
27
+ ## Evidence is missing or incomplete
28
+
29
+ A task profile is not the same as working telemetry. Ask Pi which required metrics are uncovered, where their measurements should come from, and which evaluation artifacts have been run. Normal coding checkpoints do not replace those measurements.
30
+
31
+ The [metrics guide](metrics-and-evidence.md) explains missing and insufficient evidence. More samples or instrumentation may be needed before a candidate can be compared.
32
+
33
+ ## A result was rejected
34
+
35
+ Read the validation error and audit event. Successful work needs evidence references and the output structure required by that intent. Frontier results must match the task and product revisions, identify their assessments, and satisfy the comparison rules.
36
+
37
+ Have the harness correct the measurement or output problem. Do not mark an intent successful merely to advance the workflow.
38
+
39
+ ## Work appears stuck after a restart
40
+
41
+ Pending or active intent state is durable. Inspect `state.json` and `events.jsonl` using the paths from `/openmerit logs`, and ask the harness to reconcile that work with the current session before requesting another intent. The current adapter may need supervised recovery; automatic completion after a restart is not guaranteed.
42
+
43
+ ## Checks are not running while Pi is closed
44
+
45
+ This is a current adapter limitation. Pi reports session and task signals while running, but does not provision its own background wakeups. A persistent policy reports an external-scheduler requirement. See [automatic checks](automation.md).
46
+
47
+ ## Status says paused but a check starts
48
+
49
+ This is a known 0.1.5 limitation: the automation-signal dispatch path does not check the pause flag. The flag is also reset by a process restart. Pause does not cancel existing work. See [current limitations](roadmap.md).
50
+
51
+ ## Report a reproducible problem
52
+
53
+ Include the OpenMerit, Pi, and Node.js versions, the triggering command or signal, the lifecycle stage, and a sanitized error. Remove credentials, private traces, and proprietary evaluation data before sharing anything in a public issue.
54
+
55
+ [Open an issue on GitHub](https://github.com/laz-aslam/openmerit/issues)
@@ -0,0 +1,32 @@
1
+ # Pi UX behavioral reference
2
+
3
+ This document records user-visible behaviors from the earlier OpenMerit implementation. It is not an implementation specification; the current implementation carries forward only the behaviors listed below, rather than its source code, storage schema, or internal contracts.
4
+
5
+ ## Behaviors worth preserving
6
+
7
+ - A single `/openmerit` command is the primary in-harness entry point.
8
+ - Running the command without arguments gives a concise operational summary.
9
+ - Long-running work produces short progress notifications instead of appearing silent.
10
+ - Users can pause and resume proactive work without uninstalling the extension.
11
+ - Users can request work explicitly even when proactive behavior is paused.
12
+ - Diagnostics are available from the same command surface.
13
+ - Skipped work explains why it was skipped and how the user can proceed.
14
+ - Failures are visible, actionable, and do not masquerade as successful evidence.
15
+ - The extension behaves sensibly in TUI, RPC, JSON, and print modes; prompts are not attempted when no interactive UI exists.
16
+
17
+ ## Staged behavior
18
+
19
+ The command shell can carry a proposal through an authorized model change, but only after baseline evidence, budgeted challenger trials, a harness-calculated and OpenMerit-verified frontier, policy approval, and post-swap verification.
20
+
21
+ ## Current direction
22
+
23
+ The command surface centers on evidence readiness. Its status view should answer:
24
+
25
+ 1. What task profile is active?
26
+ 2. Which required metrics are observable?
27
+ 3. Which evaluations exist and have been verified?
28
+ 4. What intelligent work is Pi currently performing for OpenMerit?
29
+ 5. What evidence is still missing or insufficient?
30
+ 6. When is the next assessment due?
31
+
32
+ The approved command surface is documented in [Pi extension](pi-extension.md).