openmerit 0.1.4 → 0.1.6-preview.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +40 -0
- package/README.md +121 -386
- package/dist/core/src/index.d.ts +101 -0
- package/dist/core/src/index.js +1649 -0
- package/dist/core/src/store.d.ts +35 -0
- package/dist/core/src/store.js +102 -0
- package/dist/pi/src/index.d.ts +32 -0
- package/dist/pi/src/index.js +794 -0
- package/dist/pi/src/scheduler.d.ts +11 -0
- package/dist/pi/src/scheduler.js +137 -0
- package/dist/pi/src/wakeup.d.ts +2 -0
- package/dist/pi/src/wakeup.js +108 -0
- package/dist/protocol/src/index.d.ts +484 -0
- package/dist/protocol/src/index.js +47 -0
- package/dist/protocol/src/schemas.d.ts +576 -0
- package/dist/protocol/src/schemas.js +280 -0
- package/dist/terminal/public/app.js +297 -0
- package/dist/terminal/public/brands/anthropic.png +0 -0
- package/dist/terminal/public/brands/baai.png +0 -0
- package/dist/terminal/public/brands/baseten.png +0 -0
- package/dist/terminal/public/brands/cerebras.png +0 -0
- package/dist/terminal/public/brands/cohere.png +0 -0
- package/dist/terminal/public/brands/deepseek.ico +0 -0
- package/dist/terminal/public/brands/google.png +0 -0
- package/dist/terminal/public/brands/groq.ico +0 -0
- package/dist/terminal/public/brands/lm-studio.png +0 -0
- package/dist/terminal/public/brands/meta.ico +0 -0
- package/dist/terminal/public/brands/mistral.png +0 -0
- package/dist/terminal/public/brands/nomic.png +0 -0
- package/dist/terminal/public/brands/ollama.png +0 -0
- package/dist/terminal/public/brands/openai.png +0 -0
- package/dist/terminal/public/brands/openrouter.png +0 -0
- package/dist/terminal/public/brands/qwen.png +0 -0
- package/dist/terminal/public/brands/vllm.ico +0 -0
- package/dist/terminal/public/brands/vllm.png +0 -0
- package/dist/terminal/public/favicon.svg +1 -0
- package/dist/terminal/public/flow.css +1 -0
- package/dist/terminal/public/flow.js +770 -0
- package/dist/terminal/public/index.html +21 -0
- package/dist/terminal/public/styles.css +779 -0
- package/dist/terminal/src/activity-merge.mjs +64 -0
- package/dist/terminal/src/browser.mjs +29 -0
- package/dist/terminal/src/cli.mjs +60 -0
- package/dist/terminal/src/collect.mjs +311 -0
- package/dist/terminal/src/discovery.mjs +93 -0
- package/dist/terminal/src/hardware.mjs +57 -0
- package/dist/terminal/src/project-activity.mjs +156 -0
- package/dist/terminal/src/sample.mjs +171 -0
- package/dist/terminal/src/server.mjs +56 -0
- package/dist/terminal/src/services.mjs +62 -0
- package/dist/terminal/src/topology.mjs +30 -0
- package/docs/adapter-guide.md +189 -0
- package/docs/architecture.md +59 -0
- package/docs/automation.md +74 -0
- package/docs/budgets.md +37 -0
- package/docs/commands.md +85 -0
- package/docs/demo-backfill.md +29 -0
- package/docs/demo-fieldkit.md +47 -0
- package/docs/demo-placement.md +30 -0
- package/docs/demo-spam.md +15 -0
- package/docs/demo-support.md +42 -0
- package/docs/demo.md +57 -0
- package/docs/first-trial.md +60 -0
- package/docs/getting-started.md +65 -0
- package/docs/index.md +40 -0
- package/docs/inference-terminal.md +439 -0
- package/docs/lifecycle.md +30 -0
- package/docs/memo.md +126 -0
- package/docs/metrics-and-evidence.md +48 -0
- package/docs/operations.md +40 -0
- package/docs/pareto-spec.md +76 -0
- package/docs/pi-extension.md +54 -0
- package/docs/roadmap.md +28 -0
- package/docs/security.md +37 -0
- package/docs/site-artwork-linocut.md +23 -0
- package/docs/site-artwork-miniature-diverse.md +28 -0
- package/docs/site-artwork-miniature.md +26 -0
- package/docs/site-demo.md +177 -0
- package/docs/site-design.md +94 -0
- package/docs/site-documentation.md +83 -0
- package/docs/site-dynamic-og.md +35 -0
- package/docs/site-faq-maintenance.md +115 -0
- package/docs/site-hero-resolution.md +60 -0
- package/docs/site-illustration-sequences.md +227 -0
- package/docs/site-inference-terminal.md +203 -0
- package/docs/site-memo.md +39 -0
- package/docs/site-og-image.md +38 -0
- package/docs/site-og-workshop.md +21 -0
- package/docs/site-section-artwork.md +56 -0
- package/docs/site-skill-review.md +57 -0
- package/docs/site-terminal-preview.md +85 -0
- package/docs/testing.md +118 -0
- package/docs/troubleshooting.md +55 -0
- package/docs/ux-reference.md +32 -0
- package/package.json +74 -42
- package/benchmark/invoice_ocr/data/invoice_01_ground_truth.json +0 -38
- package/benchmark/invoice_ocr/data/invoice_01_row_2.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_02_ground_truth.json +0 -32
- package/benchmark/invoice_ocr/data/invoice_02_row_5.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_03_ground_truth.json +0 -26
- package/benchmark/invoice_ocr/data/invoice_03_row_6.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_04_ground_truth.json +0 -26
- package/benchmark/invoice_ocr/data/invoice_04_row_7.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_05_ground_truth.json +0 -38
- package/benchmark/invoice_ocr/data/invoice_05_row_947.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_06_ground_truth.json +0 -38
- package/benchmark/invoice_ocr/data/invoice_06_row_948.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_07_ground_truth.json +0 -20
- package/benchmark/invoice_ocr/data/invoice_07_row_949.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_08_ground_truth.json +0 -38
- package/benchmark/invoice_ocr/data/invoice_08_row_1888.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_09_ground_truth.json +0 -26
- package/benchmark/invoice_ocr/data/invoice_09_row_1890.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_10_ground_truth.json +0 -20
- package/benchmark/invoice_ocr/data/invoice_10_row_1892.jpg +0 -0
- package/benchmark/invoice_ocr/data/manifest.json +0 -97
- package/dist/benchmarks.js +0 -98
- package/dist/catalog.js +0 -61
- package/dist/cli.js +0 -188
- package/dist/daemon.js +0 -407
- package/dist/diagnostics.js +0 -227
- package/dist/frontier.js +0 -56
- package/dist/harness.js +0 -1
- package/dist/integrations.js +0 -19
- package/dist/invoice-eval.js +0 -33
- package/dist/invoice-score.js +0 -124
- package/dist/judge.js +0 -43
- package/dist/llm.js +0 -207
- package/dist/pi-config.js +0 -46
- package/dist/pi-trials.js +0 -373
- package/dist/policy.js +0 -185
- package/dist/providers.js +0 -1
- package/dist/recommend.js +0 -76
- package/dist/routes.js +0 -74
- package/dist/standalone.js +0 -224
- package/dist/store.js +0 -89
- package/dist/strategist.js +0 -68
- package/dist/task-input.js +0 -54
- package/dist/traces.js +0 -127
- package/dist/trials.js +0 -140
- package/dist/types.js +0 -2
- package/examples/invoice-prompt.txt +0 -19
- package/examples/task.example.json +0 -7
- package/extension/openmerit.ts +0 -947
- package/instructions/OPENMERIT.md +0 -63
- package/instructions/openmerit.policy.json +0 -37
- package/rules.md +0 -43
|
@@ -0,0 +1,177 @@
|
|
|
1
|
+
# Hosted demo maintenance
|
|
2
|
+
|
|
3
|
+
The demo UI is in `site/demo/`. Its maintained public guides are `docs/demo.md` and the five `docs/demo-*.md` walkthroughs, registered in a dedicated Demo group in `scripts/docs.config.mjs`. The original homepage and generated documentation remain static Cloudflare assets.
|
|
4
|
+
|
|
5
|
+
## Playback-only revision
|
|
6
|
+
|
|
7
|
+
The public player now contains only recording controls and static evidence downloads. Live start, approval, custom-input, stop, and mode-switch controls and their browser API calls are removed. The task-specific “How it works” link opens the corresponding walkthrough. `DEMO_ENABLED` is false in the deployment configuration; the runner, API implementation, admission registry, and limits are retained for operator-led recording work. Published with the six-task Field kit recording as Worker `f72b3b7b-2b0b-46ec-9bb6-8cf4b726d629`, build `demo-20260928-fieldkit-v1`.
|
|
8
|
+
|
|
9
|
+
The docs explain each example’s structure, model sequence, grading, constraints, and captured outcomes. Field kit now includes the verified local six-task recording described below. Homepage FAQ answers were checked against `docs/site-faq-maintenance.md`: they describe published package behavior and require no change for this player revision.
|
|
10
|
+
|
|
11
|
+
The September 28 description update is published as Worker `31c181c9-c083-4eb4-85a4-f47516eba877`. Every example now has a two-sentence use-case description below its tabs, outside the changing playback story. All five descriptions were checked at the completed playback state, plus reset behavior and the longest description at 320px without horizontal overflow. `npm run verify` passed; the four player files and updated overview HTML/Markdown match production on both domains (12 byte comparisons). The recorded evidence, runtime, and disabled public inference setting are unchanged.
|
|
12
|
+
|
|
13
|
+
The sections below retain the previous deployments and live validation history; descriptions of live controls refer to those versions.
|
|
14
|
+
|
|
15
|
+
## Recording and controls
|
|
16
|
+
|
|
17
|
+
The initial terminal redesign was published as Worker version `3cb3f84f-2396-4666-91c4-173415f9af33`, build `demo-20260928-tui-v2`.
|
|
18
|
+
|
|
19
|
+
The focused presentation was published as Worker version `913b9a07-0d09-4f7c-8a8b-48805b34b2c9`, build `demo-20260928-tui-v3`. It reduces the default view to one current-step headline and the two model costs with fixture matches. A thin styled timeline retains native range keyboard controls. Policy, latency, complete playback trace, sample outputs, provenance, and evidence downloads live inside a collapsed **Run details** disclosure. Costs are the measured per-input means scaled to 1,000 inputs and rounded to three decimal places; original evidence remains unrounded. Live start still displays its policy before consent, and a custom input opens its returned outputs.
|
|
20
|
+
|
|
21
|
+
The presentation revision passed `npm run verify`. Browser checks covered autoplay, pause, replay, reset, seeking backward through approval, the expanded evidence view, and the hosted live start screen. Desktop, 390px, and 320px browser layouts had no horizontal overflow, and the playback controls fit one row at 320px. No additional paid live run was started for this presentation change; the prior end-to-end hosted validation below still covers the unchanged runtime. Physical-phone testing has not been performed.
|
|
22
|
+
|
|
23
|
+
The terminal UI autoplays `site/demo/recording.json`, a curated export of the verified 2026-09-28 hosted run. It contains actual aggregate measurements and one fixture's ticket outputs, with the original evidence SHA-256 and build label. Presentation timestamps are compressed to 36 seconds and are explicitly labeled; they are not claimed to be execution timings. The playback module has no inference or session API access.
|
|
24
|
+
|
|
25
|
+
Pause and timeline seeking stop playback. Replay starts it again; reset returns it to the start paused. Hidden tabs do not advance. “Try it live” discards the displayed recording state, checks availability, and offers a separate explicit start. Live changes still require approval. Live stop, reset, and returning to replay await container termination; failed termination leaves the live session visible. Controls are disabled during a session mutation so a late response cannot silently leave an orphaned sandbox. The admission registry and limits remain unchanged.
|
|
26
|
+
|
|
27
|
+
The recording does not include credentials, user-entered emails, or invented measurements. To replace it, export a successfully verified run and curate only its synthetic fixture data; update its provenance and tests together. Do not relabel recorded approval or measurements as current live activity. Homepage product FAQ claims remain unchanged: replay controls describe the website player, not the published adapter's pause behavior.
|
|
28
|
+
|
|
29
|
+
The redesign passed repository verification, nine demo tests, and eleven documentation checks. Browser checks covered autoplay, pointer and keyboard pause, seeking, replay, reset, and recovery when the live backend is unavailable. The layout was inspected at desktop, 390px, and 320px without horizontal overflow. A fresh hosted run through these controls reached manual approval, passed all six post-swap verification emails, and returned matching `technical` / `high` / `null` tickets for a custom synthetic outage email. This does not constitute physical-phone testing.
|
|
30
|
+
|
|
31
|
+
`demo/worker.mjs` extends the existing `openmerit-site` Worker with `/demo/api/*` and a token-protected `/demo/internal/*` inference proxy. `DemoGate` is a SQLite Durable Object that serializes admission and inference allowances. `DemoContainer` owns one Linux container per random session ID, starts the prepared Node runner, and schedules destruction at the absolute ten-minute deadline. `max_instances` additionally caps the platform at two containers. Provider credentials are never put into the container image or startup environment.
|
|
32
|
+
|
|
33
|
+
The image uses Node 22, Pi 0.87.0 from the root lockfile, and compiled OpenMerit from this checkout. `demo/runtime/adapter.mjs` wraps only the demo's result transport: a measured artifact name expands into the original adapter's full completion or signal input. The original adapter and core remain unchanged. Helpers are serialized per operation to avoid duplicated paid calls when a harness issues parallel tools.
|
|
34
|
+
|
|
35
|
+
## Jev task integration
|
|
36
|
+
|
|
37
|
+
The two-task demo was introduced as Worker version `2c9a5ccf-11f9-4629-bb12-f55845fc1479`, build `demo-20260928-jev-v1`. Both domains served the verified task assets and recordings. The original support recording is unchanged. The hosted Jev live start screen and task-specific evidence download were checked after publishing; the downloaded JSON matched `site/demo/recording-spam.json`.
|
|
38
|
+
|
|
39
|
+
The `spam` task is a prepared community-comment spam filter. It compares GPT-4.1 mini with TypeSafe's `typesafe/jev-1.13` through OpenRouter's Decisions API. `demo/runtime/moderation.mjs` defines the shared policy, 24 balanced synthetic fixtures, six disjoint verification fixtures, and Jev's two-choice question. The Decisions adapter is prepared by the demo author, not generated by OpenMerit. Both paths normalize to `{decision: "publish" | "hold"}`; the generative baseline additionally parses JSON. Confidence and probabilities remain evidence, not quality scores.
|
|
40
|
+
|
|
41
|
+
Sessions are bound to `support` or `spam` at admission and container startup. Both tasks share the existing daily, concurrent, request, expiry, and reservation limits. Only a spam session may use the Decisions proxy, which reconstructs one fixed question from an input of at most 1,200 characters. It pins Jev 1.13 and maximum token prices of $0.05/M input and $0/M output. Jev availability and pricing come from the model catalogue with `output_modalities=decisions`. Missing provider usage cost fails the comparison.
|
|
42
|
+
|
|
43
|
+
The hosted backend was validated on 2026-09-28 as Worker version `4e9086b5-1b7c-44db-9865-2dcab1f58a51`, build `demo-20260928-jev-v1`, using container image `sha256:3c7ac32b23b15dbcc4af3d99c79e2a29d7848a06629899051b0a32428dd30ac4`. The real Pi/OpenMerit flow reached a verified Jev swap after operator approval. GPT-4.1 mini matched 22/24 at $0.00007383333333333333 per comment and 780.75ms mean latency; Jev matched 24/24 at $0.0000184415 and 387ms. All six separate verification comments passed. A custom promotional comment returned `hold` through both paths, and the sandbox was ended after evidence export.
|
|
44
|
+
|
|
45
|
+
`site/demo/recording-spam.json` was curated with `node demo/recording.mjs evidence.json site/demo/recording-spam.json`. Its source SHA-256 is `4d6b4c85e65604633d8affc292758fbb206e2dbe80618e470bb2c21cf9e25855`. The curated sample is fixture `spam-08`, the first fixed-path error, which Jev answered correctly. Aggregates retain every fixture; the curation does not remove failures. Jev's native decision probabilities are retained in that sample's exported evidence. The recorder rejects missing provider cost, incomplete batches, absent approval, rollback, or failed verification.
|
|
46
|
+
|
|
47
|
+
The two-tab UI uses separate recordings and resets playback when selecting a task. It ends an active sandbox before a task change, awaits termination, and never relabels an existing cookie's task. A task-specific URL selects the recording directly. Browser checks covered keyboard tab navigation with retained focus, independent recorded models, the sample output difference, and 390px/320px layouts without horizontal overflow. Repository verification passed, including fourteen demo checks and eleven documentation checks. The existing homepage FAQ still describes the published adapter; these are hosted-demo changes.
|
|
48
|
+
|
|
49
|
+
## Latency presentation
|
|
50
|
+
|
|
51
|
+
Published as Worker version `5338f32a-c101-4ad9-bfdd-209c6d8f1c24`, retaining build `demo-20260928-jev-v1` and the validated container image.
|
|
52
|
+
|
|
53
|
+
The main comparison now pairs cost with mean end-to-end application request latency for both recorded and live runs. The verified headline computes both percentage changes from the original unrounded means. A slower candidate is labeled as higher latency; missing timings do not produce a latency claim. Reset and task changes clear the timing values and improvement highlight. The support recording shows 26% lower latency (1.64s to 1.21s); Jev shows 50% lower latency (0.78s to 0.39s). Both recordings retain their original evidence.
|
|
54
|
+
|
|
55
|
+
Repository verification passed. Browser checks covered both task results, reset, and switching tasks, plus desktop, 390px, and 320px layouts without horizontal overflow or JavaScript errors. This is a presentation change; the runtime and prepared evaluation remain unchanged. No additional paid run was needed. The published-package homepage FAQ remains unchanged.
|
|
56
|
+
|
|
57
|
+
## Visible run trace
|
|
58
|
+
|
|
59
|
+
Worker version `617895c7-4061-4443-bee4-1b56900f085d` moves the existing run trace below the playback controls, outside the collapsed **Run details** section. Both recorded tasks and live mode use this same visible trace. Policy, measurement notes, sample outputs, provenance, and downloads remain in **Run details**, whose hint now reads “Policy & evidence.” Trace rendering, recordings, runtime, and container image are unchanged.
|
|
60
|
+
|
|
61
|
+
Repository verification passed. Browser checks covered both recorded task traces, switching tasks, the remaining disclosure contents, and wrapping at 320px without horizontal overflow. No JavaScript warnings or errors were observed. The homepage FAQ does not require a change for this presentation update.
|
|
62
|
+
|
|
63
|
+
## Agent workflow scope
|
|
64
|
+
|
|
65
|
+
Published as Worker `ea26ce87-f5eb-4731-a340-d2c7795ae6c8`, build `demo-20260928-agent-v2`, with image `sha256:8bd042beeb33d948a527e3427135bf0ae7619be0cf61e41cc7076267a7281fb6`. Repository verification passed, including nineteen demo tests and eleven documentation checks. Desktop, 390px, and 320px browser checks covered scope selection, keyboard focus, reset, and trace isolation without overflow or JavaScript errors.
|
|
66
|
+
|
|
67
|
+
The top-level scope tabs separate the two original single-task recordings from a prepared backfill-recovery agent. The `backfill` task uses eight synthetic goals and six disjoint verification goals, with three dependent model turns and two read-only JSON-dispatched tools per goal. The fixed model is GPT-4.1 mini; the challenger is Gemini 2.5 Flash Lite. OpenMerit changes one shared application-model route across the three call sites. It does not generate the workflow, choose a separate model per step, or execute recovery writes.
|
|
68
|
+
|
|
69
|
+
The required profile checks complete manifests, tool use, JSON structure, checkpoint safety, evidence citations, mean whole-goal latency ≤8s, and mean whole-goal cost ≤$0.003. Cost metrics retain actual window spend for core budget accounting; the cost constraint scales by the eight-goal comparison window and the six-goal verification window. Tool-use measurements reference actual provider calls and tool traces with `harness_trace` evidence. The first hosted attempt was rejected because that reference initially used the evaluation-artifact type; a regression check now covers this boundary.
|
|
70
|
+
|
|
71
|
+
The original hosted validation used Worker `c5b4f8a2-e4de-44ce-8f49-f983cdc200ee`. The evidence-type correction was deployed as `62dc0c47-075f-4821-a693-fcb713da3f97` while retaining the existing public UI assets. The public sandbox allowance then blocked another engineering start. It was not reset or bypassed. Subsequent validation used the same runner locally with real OpenRouter calls through a loopback-only proxy enforcing the same 120-request and $1.25 reservation bounds. No credential was written to source files.
|
|
72
|
+
|
|
73
|
+
The selected recording comes from the local `demo-20260928-agent-local-v2` run, verified at `2026-09-28T11:51:32.222Z`. The source export SHA-256 is `cdd60cea8b450219b63c37829c4b2c53c6fb9366e7f0afff749a4d1f14e19b04`. GPT-4.1 mini completed 7/8 goals at mean $0.0007626 and 3674.375ms; Flash Lite completed 8/8 at $0.00020525 and 2065ms. OpenMerit proposed the actual candidate, operator approval was supplied, the configuration changed, and all six separate goals passed. The sample preserves the fixed agent incorrectly marking a checksum conflict complete versus the managed agent returning investigate. Every fixture remains in the aggregates.
|
|
74
|
+
|
|
75
|
+
After verification, the harness closing response reached the earlier output reservation, so no custom request was exercised in that run. This boundary is disclosed in the recording's `sourceNotes` and the public guide. The live agent task now caps output at 2,048 tokens; application turns request 350. A later run under the lower cap completed a valid no-change comparison: both routes scored 7/8, with an incomplete provider response in the challenger, and no model was proposed. That run used 67 inference requests and about $0.526 of conservative reservation. The runner and proxy were shut down afterward. An earlier local run also returned no proposal; its evidence was retained. These outcomes are not overwritten or presented as successful swaps.
|
|
76
|
+
|
|
77
|
+
The replay uses the verified application outcome and is labeled as a local recording. It retains the original source hash, actual sample tool requests/results, and the post-verification reservation note. The corrected hosted runtime has not received another full engineering run because of the public allowance. Its controls and API configuration are checked separately after publishing; do not claim a new hosted or custom-goal success from those read-only checks.
|
|
78
|
+
|
|
79
|
+
The scope tabs reset playback, end an active sandbox before switching, preserve the last single-task choice, and keep keyboard movement within each tab row. The workflow steps and requirements appear in the default view; the run trace remains visible; policy, sample manifests, provenance, and evidence downloads remain in Run details. The homepage FAQ still describes the published adapter and is unchanged.
|
|
80
|
+
|
|
81
|
+
## Secrets
|
|
82
|
+
|
|
83
|
+
The existing Worker needs an `OPENROUTER_API_KEY` secret. Use an inference key with a provider-side credit limit and expiry. Do not store it in the repository, image, browser, or build args.
|
|
84
|
+
|
|
85
|
+
```sh
|
|
86
|
+
npx --yes wrangler@4.136.3 secret put OPENROUTER_API_KEY
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
`DEMO_ENABLED` controls new public work and should stay false until end-to-end validation has passed. Key exhaustion or expiration requires operator action. The initial provided key expires on 2026-10-05; replace it securely if the demo should remain available beyond that date.
|
|
90
|
+
|
|
91
|
+
## Development and deployment
|
|
92
|
+
|
|
93
|
+
Install root dependencies and `npm ci --ignore-scripts --prefix demo`. Build OpenMerit with `npm run build`. For a local run, configure `OPENROUTER_API_KEY` in the process environment, set `DEMO_LOCAL=1` and an absolute temporary `DEMO_WORKSPACE`, then run `node demo/runtime/server.mjs`. The local page is `http://127.0.0.1:45832/demo`. Local mode is a single supervised session; public admission and proxy controls are enforced by the Worker.
|
|
94
|
+
|
|
95
|
+
Run `npm run verify` before deployment. Docker must be running for `npx --yes wrangler@4.136.3 deploy`, which runs the build, demo tests, documentation checks, and image build. The Docker build context uses an allowlist in `.dockerignore`; no local operational state or credentials are included. Static homepage/docs routes must be smoke-tested after publishing alongside the demo.
|
|
96
|
+
|
|
97
|
+
## Release validation
|
|
98
|
+
|
|
99
|
+
Published to `openmerit.site/demo` and `www.openmerit.site/demo` as Worker version `db7f39aa-aca8-47c6-b4a3-5681d91b1bfd`, build label `demo-20260928-cloudflare-v1`. The validated container image is `sha256:7b60b4d54e61a7970bfb698c9e310896d68a8dc30466c11a4282efe432119172`.
|
|
100
|
+
|
|
101
|
+
Validate a fresh session through setup, all comparison calls (72 for support, 48 for spam), accepted frontier, explicit approval, six verification calls, custom input, evidence download, and session termination. Verify the no-change and ineligible-candidate branches with deterministic tests. Inspect both desktop and narrow browser layouts. Do not substitute mock results for live checks.
|
|
102
|
+
|
|
103
|
+
On 2026-09-28, the hosted container completed a live comparison and approved swap. GPT-4.1 mini matched 24/24 fixtures at $0.000106 per email and 1.64 seconds mean latency. GPT-4.1 nano matched 17/24 and was ineligible. Gemini 2.5 Flash Lite matched 24/24 at $0.000027 per email and 1.21 seconds, was accepted on the frontier, and passed all six separate post-swap emails. A custom synthetic billing email returned the expected ticket through both paths. The browser downloaded 244,331 bytes of JSON evidence, including all measured batches and 18 OpenMerit audit events. These are one run's observations, not fixed demo results or promised savings.
|
|
104
|
+
|
|
105
|
+
The hosted homepage, documentation index, and demo guide returned HTTP 200. Approval without a session returned 410; cross-origin session creation and an invalid internal proxy token returned 403. Ending the completed sandbox returned the page to “Session ended” with its inference controls disabled. The full repository verification passed, including six demo checks and eleven documentation checks. Desktop and 390px/320px browser layouts were checked without page overflow. Physical phone testing has not been performed.
|
|
106
|
+
|
|
107
|
+
Pi 0.87 requires `"$DEMO_PROXY_TOKEN"`, including the dollar sign, for the `models.json` environment reference. A local dummy-token proxy verified the actual outgoing Authorization header. Cloudflare's image rollout can finish after `wrangler deploy` returns; confirm the application's current image version before starting a validation session. The public admission registry is `demo-admission-public-v1`, separate from the initial engineering trials. Keep that name stable in future releases to preserve limits.
|
|
108
|
+
|
|
109
|
+
Known boundaries: the demo uses observed point values on a finite synthetic suite, fixed helpers and candidate shortlist, and an application configuration restoration on failed verification. It does not establish unattended production optimization, broad statistical confidence, or independent evidence authenticity. Physical phone testing is separate from browser viewport checks.
|
|
110
|
+
|
|
111
|
+
|
|
112
|
+
## Models per agent task
|
|
113
|
+
|
|
114
|
+
Published as Worker `d0c48936-6a6e-478e-988b-db751cb72120`, build `demo-20260928-agent-routes-v1`, with container image `sha256:2b70880671bd09677ad5830c40d28f59cd3dcfb6bc4cb58db1a4f05815ae85bb`.
|
|
115
|
+
|
|
116
|
+
The mixed-route revision replaces the shared-model backfill challenger with `backfill/mixed-v1`: GPT-4.1 nano for incident inspection, GPT-4.1 mini for checkpoint lookup, and Gemini 2.5 Flash Lite for the manifest. The fixed route uses GPT-4.1 mini at every step. Each call receives prior context; the runtime dispatches the provider model for that step and records it in calls and traces. The application config diff includes all three assignments. OpenMerit compares the two prepared route configurations using the same whole-goal constraints and approval lifecycle. This is a bounded supplied route comparison, not open-ended per-step search. The core and published package contracts are unchanged; the candidate model ID names a demo application configuration that the prepared adapter resolves.
|
|
117
|
+
|
|
118
|
+
The proxy allows only the constituent provider models and rejects route IDs as provider IDs. Catalogue refresh checks every constituent. Curation rejects a recording if its displayed route differs from the actual provider call models, including verification. The main comparison replaces each single model label with three task/model rows. Playback reveals the managed assignments after application, and backward seeking or reset restores the fixed assignment. Custom requests dispatch through the selected route. Single-task demos keep their original model presentation.
|
|
119
|
+
|
|
120
|
+
A fresh bounded local session `demo-20260928-agent-routes-local-v1` completed the actual Pi/OpenMerit lifecycle, explicit approval, all six post-swap checks, a custom request, evidence export, and termination. Verification completed at `2026-09-28T12:19:57.721Z`. Source SHA-256: `092a958790e6399682244666240141dfc60356e0a06f368919c6a4336ea69700`. Baseline: 6/8, mean $0.0007742 and 4119.125ms. Mixed route: 8/8, mean $0.0004018625 and 3802.125ms. Verification: 6/6, mean $0.00039938333333333333 and 3186.5ms. This gives 48% lower observed cost and 8% lower latency. The sample is BF-4102: the fixed route incorrectly includes quarantined partition 2 in replay; the mixed route excludes it. All fixtures remain in the aggregate. The later custom request for that job returned the expected plan through both paths; it does not replace the measured baseline error.
|
|
121
|
+
|
|
122
|
+
This session used 99 inference requests and about $0.883 of conservative reservation under unchanged 120-request/$1.25 bounds. The public admission registry was not reset. The new recording supersedes the shared-route recording and its closing-response limitation; the historical outcomes above remain documented. A fresh hosted inference run remains unverified because the daily public allowance was exhausted. The hosted code uses the same validated runtime; deployed assets and configuration are checked separately. Homepage FAQ claims remain about published 0.1.5 and require no change for this demo-only behavior.
|
|
123
|
+
|
|
124
|
+
Validation for this revision: `npm run verify` passed, including 21 demo tests and 11 documentation checks. Browser checks covered the three model rows, actual per-step trace labels, backward seeking, reset, switching to the existing single-task scope, and 390px/320px layouts without horizontal overflow or JavaScript warnings/errors. Physical phone testing was not performed.
|
|
125
|
+
|
|
126
|
+
|
|
127
|
+
## Second agent workflow: GPU placement
|
|
128
|
+
|
|
129
|
+
The `placement` example adds a second tab within Agent workflow, with its own deep link, recording, strict policy, custom-job help, sample formatter, and remembered selection when switching scopes. Backfill and both single-task recordings remain unchanged. The UI displays each example's actual model assignments, whole-goal costs and latency, and visible tool trace.
|
|
130
|
+
|
|
131
|
+
The initial unpublished route used `mistralai/mistral-small-3.2-24b-instruct`, `meta-llama/llama-4-scout`, and `qwen/qwen3-235b-a22b-2507`, respectively, for job lookup, capacity lookup, and placement, with GPT-4.1 mini as baseline. These models do not satisfy the later request for releases in the preceding six months and have been replaced before publication. All records are synthetic. Tools never reserve GPUs or launch jobs. Six constraints determine pool eligibility; ties use quote, completion minute, and alphabetical pool ID. Exact plan matching also verifies the tie break. All eight goals and four auxiliary checks must pass; mean inference latency must be ≤12s and cost ≤$0.003 per goal. Six disjoint goals verify a swap. GPU price quotes are synthetic input data and are never reported as measured inference savings.
|
|
132
|
+
|
|
133
|
+
Initial integration findings are retained: the first local run stopped on a 404 because Llama Scout's JSON-object endpoint was unavailable (Google required a separate key; the available DeepInfra endpoint supported JSON Schema). One minimal JSON Schema request confirmed the supported mode. Both routes now use the same schema per turn. A second full local run produced 7/8 for each route and correctly proposed no swap: the challenger was influenced by a raw queue note. The adapter now passes only trusted structured job requirements and pool inventory to later stages, with the raw note retained only at initial job lookup and in the synthetic source fixture. This handoff is identical for both paths; the fixtures, expected answers, and acceptance thresholds were not relaxed. Model/schema/constraint checks remain scored from actual responses. Public admission counters and all provider/session spending limits are unchanged.
|
|
134
|
+
|
|
135
|
+
The structured-handoff run rejected the Qwen3 30B route as well: both paths selected a pool that missed one deadline, yielding 7/8. That full comparison is retained in `openmerit-placement-validation3`. The prepared planner was consequently changed to Qwen3 235B A22B Instruct 2507 for a new comparison; eligibility and fixtures remain unchanged.
|
|
136
|
+
|
|
137
|
+
The larger-planner run also produced no proposal (7/8 fixed, 6/8 mixed): deadline/queue-time mistakes remained. The final capacity adapter now includes a deterministic `finish_minute` for each pool, calculated from the same original start/runtime fields, and the shared prompt explicitly filters deadline failures before ranking. No pool is filtered or selected by the tool; expected answers and thresholds remain unchanged. This is prepared application engineering, not a capability attributed to automatic OpenMerit discovery.
|
|
138
|
+
|
|
139
|
+
|
|
140
|
+
### Recent-model lineup
|
|
141
|
+
|
|
142
|
+
The final lineup uses Gemma 4 26B A4B (April 3, 2026) for job lookup, MiMo-V2.5 (April 22) for capacity lookup, and DeepSeek V4.1 Flash (September 10) for planning. Qwen3.6 Plus (April 2) is the fixed model for all three baseline tasks. OpenRouter's model pages explicitly provide those release dates; the public guide links each source. The checked window is March 28–September 28, 2026. The route ID is `placement/recent-v1`; no prior trials using older models are relabeled as this route. Optional reasoning is disabled identically for both paths; the final planner returns explicit pool constraint checks plus its plan within 900 tokens, while lookups request 350. A six-call smoke preflight exercised both routes successfully and counts within the local proxy's original request and reservation limits.
|
|
143
|
+
|
|
144
|
+
The UI derives the incumbent from the selected task/configuration/recording instead of assuming GPT-4.1 mini. This matters before measurement, when rewinding, and when deciding whether a live route actually changed. Release provenance stays in Run details and the evidence download. Tests check the six-month release window, distinct provider assignments, bounded JSON Schema requests, task-specific allowlists, and structured context handoffs.
|
|
145
|
+
|
|
146
|
+
|
|
147
|
+
The recent-model run completed setup, eight baseline goals, eight mixed-route goals, actual OpenMerit comparison, explicit approval, configuration change, six disjoint verification goals, a custom blocked-job request, evidence export, and termination. Build: `demo-20260928-placement-recent-local-v1`. Verified at `2026-09-28T13:00:40.486Z`. Export SHA-256: `bd228d890df7ff335cbc8177201ecd102f119f856d652956812168cd9b9a8bed`. Baseline: 8/8, $0.001053771875 and 8187.25ms mean. Challenger: 8/8, $0.000409093125 and 7091.25ms mean. Verification: 6/6, $0.0003396550708333333 and 10956.5ms mean. All constraints remained unchanged throughout this run. The sample is GPU-6101 and both outputs agree; savings do not imply an accuracy improvement. The custom GPU-6105 request returned blocked from both routes. The full record is curated into `site/demo/recording-placement.json`; the original three recordings are unchanged. Hosted inference has not been rerun because the shared public daily allowance is exhausted; local calls enforce the same request/reservation controls.
|
|
148
|
+
|
|
149
|
+
The local session and smoke preflight consumed 105 inference requests with $0.963 of conservative reservation, within the unchanged 120-request/$1.25 bounds. The sandbox was ended and its temporary local runner stopped after validation.
|
|
150
|
+
|
|
151
|
+
Published as Worker `4a68a965-108a-4154-9f44-2fa229202988`, build `demo-20260928-placement-recent-v1`, with container image `sha256:b0ed8723a37c651fe1463421af6944d3d462ad7427729d9edaa5d12a58c01699`. Cloudflare reports the container application ready. `npm run verify` passed all package checks, 27 demo tests, and 11 documentation checks. Browser checks covered the new deep link, per-task model rows, completed playback, reset to Qwen3.6 Plus, keyboard tab selection, remembered scope selection, release links, and 390px/320px layouts without horizontal overflow or JavaScript warnings/errors. The published recording was also checked in the browser. Twenty-four asset responses across both domains matched the local files, including all four recordings, demo code, homepage, and generated guide; both public configuration responses expose the new build and correct routes. Hosted inference remains a separate unverified check because the public daily allowance is exhausted. Homepage FAQ claims remain unchanged and refer to released 0.1.5 behavior.
|
|
152
|
+
|
|
153
|
+
## Third agent workflow: six-task field kit
|
|
154
|
+
|
|
155
|
+
The `fieldkit` example starts with a heterogeneous six-model incumbent and compares it with a second heterogeneous six-model route. `fieldkit-models.mjs` maintains the assignments, release dates, role descriptions, and per-model API requirements. The steps exercise Spanish translation, image perception, HTML extraction, static code analysis, constraint reasoning, and structured writing. Original provider IDs are Qwen3.6 Plus, MiMo-V2.5, Gemini 3.1 Flash Lite, MiMo-V2.5 Pro, Qwen3.7 Plus, and Gemma 4 31B. The challenger uses Hy-MT2 1.8B, Perceptron Mk1.5, Schematron V2 Turbo, Laguna XS 2.1, DeepSeek V4.1 Flash, and Gemma 4 26B A4B. The two route IDs are `fieldkit/original-v1` and `fieldkit/specialists-v1`; neither is a provider model ID.
|
|
156
|
+
|
|
157
|
+
The runner dispatches read-only source adapters in a fixed six-turn sequence. Translation and code output use no `response_format`; all other structured stages use JSON Schema. The extraction instructions live in the schema to match Schematron's API. Vision receives an embedded PNG, never a text transcription. `demo/build-fieldkit-labels.mjs` regenerates the nine synthetic labels using the existing Resvg development dependency; the checked PNGs ship inside the runtime image, with no new production dependency. The same inputs, formats, token ceilings, and reasoning settings apply to both routes. Tasks one through four request 180 tokens each; feasibility requests 2,048 including low-effort reasoning; final writing requests 350. Optional reasoning is disabled for other stages where supported.
|
|
158
|
+
|
|
159
|
+
Every structured intermediate answer is checked separately; a correct final answer cannot hide an extraction, code, or reasoning mistake. The prepared source-adapter completion check uses the existing `tool_use_performance` metric, with this meaning disclosed in the public guide. There is no model-directed arbitrary tool use. Four evaluation goals and four disjoint verification goals cover offline behavior, sample rate, battery, storage, deadline, connector, and exact boundaries. Both lists are fixed before the complete comparison. One custom comparison is permitted. Session request, reservation, duration, and public daily limits remain unchanged. The original four recordings are retained unchanged.
|
|
160
|
+
|
|
161
|
+
Recording curation now derives step and verification counts from each task instead of assuming three tasks and six fresh goals. It validates every provider assignment and trace in both evaluation and verification. Task-specific verification counts also reach the actual OpenMerit policy, local API, live proposal text, and evidence. Plain-text translation output is explicitly selected by its prepared adapter; invalid JSON from structured stages remains a failure. Provider finish reason and reasoning-token count are retained for diagnosis. Catalogue discovery checks the parameters and modalities actually required by each constituent, including models that intentionally lack JSON output or tool-calling support.
|
|
162
|
+
|
|
163
|
+
Unpublished engineering checks are retained in `/private/tmp/openmerit-fieldkit-preflight-v*.json` and runner logs. The first smoke test found empty Qwen translation and Cohere extraction responses with a small output budget; optional reasoning was corrected. Cohere then rejected the bounded request with HTTP 422 and was replaced with Gemini 3.1 Flash Lite before evaluation. The specialist's first complete smoke test incorrectly rejected adequate battery capacity while reasoning was disabled. Feasibility now gets the same bounded low reasoning effort on both routes. A later MiMo vision smoke call returned HTTP 502; no result was invented and that partial check was not curated. The final specialist smoke test passed all six tasks. These integration checks are separate from the fresh supervised comparison and from the shared public allowance.
|
|
164
|
+
|
|
165
|
+
The first complete comparison (`openmerit-fieldkit-validation/evidence-final.json`) reached the frontier and correctly proposed no change. The original mix passed 4/4 at $0.00172305907 and 34949ms mean; the challenger passed 3/4 at $0.000653305625 and 42400.75ms. Its last feasibility call consumed all 1,536 tokens in reasoning and returned no decision. The downstream writer's unsupported answer was retained and failed the grader. This is not a successful recording. The final output allowance was increased to 2,048 for feasibility on both routes, and the shared instruction now asks for one pass through the six comparisons without exploring alternative plans. The first three HTML extraction calls also queued for 28–38s when goals ran concurrently; the six-task example now uses concurrency one identically on both routes. Other tasks retain concurrency three. The four fixtures, four verification goals, six constraints, exact graders, 45s/$0.006 policy, and session limits are unchanged. A second run uses a fresh workspace and build `demo-20260928-fieldkit-local-v2`; no first-run measurements are relabeled or reused as fresh evidence.
|
|
166
|
+
|
|
167
|
+
The second and third comparisons stopped during baseline collection after HTTP 504 responses from MiMo-V2.5 image inference. Their incomplete evidence is retained in the v2/v3 validation directories; no savings are claimed from either. The incumbent vision stage was replaced with Qwen3.6 Flash, released April 27, 2026 with image and JSON Schema support verified against its OpenRouter page. This creates `fieldkit/original-v2`; the challenger is unchanged. A fresh fourth comparison uses build `demo-20260928-fieldkit-local-v4`. Public playback remains recorded-only.
|
|
168
|
+
|
|
169
|
+
The fourth run measured 4/4 for the revised incumbent ($0.00161658683 and 26141.25ms mean), but only 3/4 for the challenger ($0.00038325581875 and 12703.25ms). DeepSeek incorrectly rejected sufficient storage on KIT-8101; its output and the downstream blocked handoff remain in the failed evidence. OpenMerit proposed no change. The next prepared challenger, `fieldkit/specialists-v2`, retains the incumbent Qwen3.7 Plus for feasibility while changing the other five assignments. Both routes still contain six distinct models. No fixtures, graders, or constraints were relaxed.
|
|
170
|
+
|
|
171
|
+
The final local run used build `demo-20260928-fieldkit-local-v5` and verified at `2026-09-28T15:00:27.321Z`. The source export SHA-256 is `9f0986917f41354bb8ce80421bb9a413825fb457eb37c78af119747880a1514b`. The original six-model mix passed 4/4 at $0.00186733932 and 29486ms mean; the approved six-model mix passed 4/4 at $0.0014490875 and 25628ms. All stage checks passed. The four separate post-swap goals passed at $0.00120398075 and 22501ms mean. The actual OpenMerit proposal was approved, the configuration changed, evidence was exported, and the local sandbox ended. The recording is curated from this single complete run; no earlier observations are spliced into its aggregates. It demonstrates 22% lower observed cost and 13% lower mean latency, with the same complete goal success. Both paths return the same correct KIT-8101 handoff. Qwen3.7 Plus is retained for reasoning; five other assignments change, and the release list deduplicates the shared model. The public player remains recording-only, with no inference or session creation.
|
|
172
|
+
|
|
173
|
+
Final six-task validation passed `npm run verify`, including 33 demo tests and 11 documentation checks. Browser checks covered autoplay, completed measurements, all six assignments, reset, scope memory, keyboard task navigation, the collapsed policy/evidence section, and the evidence download (parsed JSON matched the curated file). Desktop, 390px, and 320px checks showed no horizontal overflow or JavaScript warnings/errors. Physical phone testing was not performed. The successful local session used 102 inference requests and $0.991 of conservative reservation under the existing limits.
|
|
174
|
+
|
|
175
|
+
Published as Worker `f72b3b7b-2b0b-46ec-9bb6-8cf4b726d629`, build `demo-20260928-fieldkit-v1`, with container image `sha256:db1c69574b9777b7e4eaf13a4ef282b1a00377625fff280832ea03b5e557c090`. Cloudflare reports the container application ready. Thirty-six asset responses across both domains match the checked local files, including all five recordings and all six demo guide pages. Both configuration responses identify the new build, expose the six-stage routes and four-goal verification policy, and report `enabled: false`. Unconfirmed session-creation probes return HTTP 503, so no public inference session was started. The published demo and its Field kit guide were checked in the browser; the completed view shows the six distinct models on each side and the recorded metrics with no JavaScript warnings/errors. Temporary recording runners ended after evidence capture.
|
|
176
|
+
|
|
177
|
+
The full-site publication was refreshed as Worker `6b9dfbd1-d9a0-4f13-9120-4e754774998a` on September 28, 2026. All 190 public-file responses across both domains match the local build, including all demo assets and guides. The container image and measured recordings are unchanged, and public inference remains disabled. No new paid run was started for this publication.
|
|
@@ -0,0 +1,94 @@
|
|
|
1
|
+
# OpenMerit website
|
|
2
|
+
|
|
3
|
+
The website lives in `site/` and uses the existing Cloudflare asset configuration in `wrangler.jsonc`. The homepage is plain HTML, CSS, and JavaScript. The [documentation section](site-documentation.md) is generated from curated Markdown at build time; no browser framework is required. The [guided demo](site-demo.md) adds a Worker API and temporary Cloudflare Containers; Docker must be running for deployment.
|
|
4
|
+
|
|
5
|
+
The playback-only revision is published as Worker `31c181c9-c083-4eb4-85a4-f47516eba877`, build `demo-20260928-fieldkit-v1`. It removes live sandbox controls, disables public session creation, adds a dedicated Demo documentation group, and adds the six-task Field kit recording. Each of the five examples has a persistent two-sentence description below its tabs. Both Field kit paths already use six distinct models; five assignments change while the feasibility model is retained. The recorded comparison shows 22% lower observed cost and 13% lower mean latency, with 4/4 evaluation and 4/4 separate verification goals passing. The deployment history below describes earlier published behavior.
|
|
6
|
+
|
|
7
|
+
## Local preview
|
|
8
|
+
|
|
9
|
+
From the repository root:
|
|
10
|
+
|
|
11
|
+
```sh
|
|
12
|
+
npm run check:docs
|
|
13
|
+
python3 -m http.server 45831 --bind 127.0.0.1 --directory site
|
|
14
|
+
```
|
|
15
|
+
|
|
16
|
+
Open http://127.0.0.1:45831/. This port avoids another local project's cached service worker on ports 4173 and 4174.
|
|
17
|
+
|
|
18
|
+
## Cloudflare deployment
|
|
19
|
+
|
|
20
|
+
The 2026-09-28 guided demo was introduced as Worker version `db7f39aa-aca8-47c6-b4a3-5681d91b1bfd`. The preceding version `4a68a965-108a-4154-9f44-2fa229202988` includes top-level Single task and Agent workflow tabs. Single task retains support tickets and the Jev spam filter; Agent workflow contains Backfill recovery and GPU placement examples, each with three dependent tasks and stricter constraints. GPU placement uses application models released within March 28–September 28, 2026, with dated source links in Run details. The mixed-route revision displays the model assigned to each task in both columns, with aggregate whole-goal metrics. Both show mean latency alongside cost, with percentage changes after a verified swap. The run trace is visible below the controls; policy, outputs, and evidence stay in the collapsed details section. It retains the homepage and generated documentation assets; `/demo` plays real recorded runs by default and starts bounded Cloudflare Containers on demand. See [demo architecture and live validation](site-demo.md). The following record describes the preceding static-site release.
|
|
21
|
+
|
|
22
|
+
Published on 2026-09-23 to https://openmerit.site/ and https://www.openmerit.site/ using the existing `openmerit-site` Worker and routes in `wrangler.jsonc`:
|
|
23
|
+
|
|
24
|
+
```sh
|
|
25
|
+
npx --yes wrangler@4.136.3 deploy
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
Deployment version: `90bd0c29-5d32-4386-93df-a53e680d9cc5` (all documentation pages share the default OG image; separate cards disabled). Previous version: `921eb118-0539-4959-ad5f-e9b928e76d70` (documentation OG image aligned with the default workshop card).
|
|
29
|
+
|
|
30
|
+
Production verification: 54 URLs across both domains match the checked local files, including all 16 documentation pages, the homepage, downloadable Markdown samples, docs CSS and JavaScript, the search loader and manifest, both search assets, the shared default OG image, and the legacy docs-image path. The documentation build revision is `e205a278c84e`. Facebook and Twitter crawler user agents received metadata pointing to the default image and matching PNG bytes on both domains; the previous docs-image URL now serves the same default artwork. All 11 documentation checks passed. The homepage retains the “Autonomously.” headline ending and all 13 questions. Documentation behavior is recorded in [the documentation notes](site-documentation.md), and image configuration is recorded in [the sharing-image notes](site-dynamic-og.md). Publication preserves the previously deployed full-resolution hero frames and illustration player.
|
|
31
|
+
|
|
32
|
+
## Design and behavior
|
|
33
|
+
|
|
34
|
+
- Reference: https://kobbe.io/ — quiet typography, warm neutral surfaces, a wide illustration, and generous whitespace. Copy and artwork are original to OpenMerit.
|
|
35
|
+
- The page contains the hero, getting-started section, FAQ, and footer.
|
|
36
|
+
- The standalone [memo](site-memo.md) is generated from `docs/memo.md` at `/memo/`, with a homepage navigation link and a reading layout using the site's typography and colors. It describes the company's direction separately from released product documentation. The memo, header artwork, and navigation link are published through the existing Cloudflare workflow.
|
|
37
|
+
- The [documentation section](site-documentation.md) publishes 22 guides with grouped navigation, section outlines, lazy-loaded browser search, copyable examples, and mobile disclosures. It builds from maintained Markdown and validates local links before deployment.
|
|
38
|
+
- The FAQ has 13 questions covering the harness boundary, supported adapters, evaluations, model tradeoffs, evidence, approvals, regressions, background checks, pause, data, cost, budget enforcement, and readiness. Claims are reviewed against the published release using the [FAQ maintenance procedure](site-faq-maintenance.md). The 0.1.5 pause limitation is explicit; budget and evidence checks are described without implying independent enforcement or evidence authentication.
|
|
39
|
+
- Small miniature vignettes accompany setup (prepared tools), the FAQ (two makers comparing notes), and the footer (the workshop dog at rest). See [assets and generation prompts](site-section-artwork.md). They are decorative, lazy-loaded, and deliberately much smaller than the hero.
|
|
40
|
+
- `site/illustrations.js` plays sequences of separately generated poses: the seated maker raises a hand plane, the toolbox lid closes and opens, the FAQ makers point across their notebook, and the dog lifts its head. The hero uses four individual 2172 × 724 images; eight poses for each supporting scene are packed into WebP atlases. A canvas copies complete frames with no displacement filters, morphing, crossfades, or procedural deformation. See [the sequence assets, exact prompts, and packaging notes](site-illustration-sequences.md) and [the hero resolution correction](site-hero-resolution.md).
|
|
41
|
+
- Playback steps forward and back through the poses, with a short hold at each end. The first pose rests for 1.1s in the hero, 1.4s in setup and FAQ, and 2.4s in the footer. The final pose holds for 0.7s, 0.6s, 1s, and 1.6s respectively. Intermediate frames last 120–180ms. FAQ interactions do not trigger the scene. A footer control freezes/resumes the current pose and remembers the choice locally.
|
|
42
|
+
- Atlases load only when the illustration is visible and motion is allowed. Timers run only when a different frame is due and stop offscreen or in a hidden tab. Reduced motion displays the static first-frame poster. Posters also work without JavaScript or if an atlas fails to load. There are no animation dependencies.
|
|
43
|
+
- Responsive navigation, native FAQ disclosures, and a copyable `pi install npm:openmerit` command work without dependencies. The setup section explains automatic onboarding in the user's product repository, evaluation budget and permission confirmation, and evidence gathering during ordinary work. Requirements match the published package: Node.js 22.19+ and Pi 0.87 with a configured provider. Source installation and diagnostics remain in the linked guide. The cost FAQ reflects npm availability. The static content, navigation, and disclosures remain available without JavaScript.
|
|
44
|
+
- The self-hosted Inter Latin variable font comes from Google Fonts. Its SIL Open Font License is included at `site/assets/inter-LICENSE.txt`.
|
|
45
|
+
- The “Works with pi” link uses Pi’s unmodified [official primary SVG](https://pi.dev/logo.svg), downloaded from its [press kit](https://pi.dev/press-kit) to `site/assets/pi-logo.svg`, and links to https://pi.dev/.
|
|
46
|
+
- The 1200 × 630 PNG social card contains only “Good models prove themselves in the work.” and a matching miniature workshop illustration. Open Graph and Twitter metadata reference `site/assets/openmerit-og-v1.png`; see [the generation prompt and asset provenance](site-og-image.md).
|
|
47
|
+
- A focused pass with Emil Kowalski's design, animation-review, and mobile skills adds capability-gated hover effects, small pointer press feedback, and instant FAQ indicator updates for keyboard use. See [the review and verification limits](site-skill-review.md).
|
|
48
|
+
|
|
49
|
+
## Verification
|
|
50
|
+
|
|
51
|
+
- The 2026-09-23 hero resolution correction was checked in Chromium and WebKit at 1440 CSS pixels / DPR 2 and 390 CSS pixels / DPR 3. All four source frames and the canvas retain 2172 × 724 pixels, all poses render distinct pixels, and no low-resolution v1 hero assets are loaded. Supporting animations, pause/resume persistence, reduced motion, sharp fallback for missing or undersized frames, no-JavaScript rendering, and overflow checks passed. See [the sources, prompts, and verification details](site-hero-resolution.md).
|
|
52
|
+
- The 2026-09-23 FAQ update was checked locally in Chrome at 1440, 390, and 320 CSS pixels. All 13 answers opened, Enter and Space toggled disclosures, and there were no page or answer-content overflows or runtime errors. Desktop and narrow-screen screenshots were visually reviewed. All 13 native disclosures also opened with JavaScript disabled at 390px. FAQ documentation links resolve to existing repository files, IDs are unique, and whitespace checks pass. Published to both domains as version `34defac0-60c3-49a2-9501-9d40ff840ea4`; served HTML matches the checked local file.
|
|
53
|
+
|
|
54
|
+
- The 2026-09-23 frame-sequence replacement was checked locally and on the deployed site in Chromium and WebKit at 1440px and 390px. All four hero poses and all eight poses in each supporting scene were observed through a full repeat; canvas draws use whole atlas cells with equal source and destination dimensions. Pause freezes the current pose and persists across reload. Resume, runtime reduced motion, offscreen stopping/re-entry, and overflow checks passed without runtime errors. Initial reduced motion avoids atlas downloads; JavaScript-disabled and failed-atlas checks preserve the posters. The HTML, CSS, player, and all eight sequence/poster assets on both domains match the local files. Physical phone hardware was not tested.
|
|
55
|
+
- The 2026-09-23 motion visibility fix was checked locally and on the deployed site in Chromium and WebKit at 1440px and 390px. Rendered frame comparisons confirmed movement in all four scenes through the actual playback function. Natural repeat cycles, pause/resume persistence across reload, runtime reduced-motion changes, offscreen stopping and re-entry, no horizontal overflow, and no runtime errors passed in all four browser/viewport combinations. The HTML and versioned illustration script on both domains match the local files. Syntax and whitespace checks passed. Physical phone hardware was not tested.
|
|
56
|
+
- The 2026-09-23 install-copy update was checked at 320, 1024, and 1263 CSS pixels, with no page or command overflow. Verified copy success feedback and the revised cost FAQ. The command stays at 14px on narrow layouts; the changed stylesheet and script use versioned URLs to refresh cached previews. These checks used a browser, not physical phone hardware.
|
|
57
|
+
- Browser checked at 320, 390, 580, 768, 1024, 1440, and 1920 CSS pixels; no page-level horizontal overflow.
|
|
58
|
+
- Checked keyboard operation, mobile navigation and Escape, FAQ expansion, clipboard success and denied-permission fallback.
|
|
59
|
+
- Axe WCAG 2 A/AA and 2.1 AA checks found no violations at desktop and mobile sizes. This is an automated check, not a claim of exhaustive accessibility certification.
|
|
60
|
+
- Confirmed readable static content, navigation, and FAQ interaction without JavaScript.
|
|
61
|
+
- No browser console or runtime errors; local font and artwork load.
|
|
62
|
+
- Illustration motion checked at 1440px and 390px: no overflow or visible patch seams; gestures end cleanly, offscreen motion stops, FAQ interaction does not replay the scene, and reduced motion works both on initial load and when changed at runtime. Static artwork and FAQ disclosures also verified with JavaScript disabled.
|
|
63
|
+
- After correcting the one-pass behavior, rendered frame pairs confirmed visible pixel changes inside all four scenes. Verified natural repeat/rest cycles, interrupted playback on re-entry, persistent pause/resume, and no horizontal overflow at 320px and 390px. This is browser verification, not physical-phone testing.
|
|
64
|
+
- Repository verification: all 35 tests, workspace syntax checks, and TypeScript checks pass.
|
|
65
|
+
|
|
66
|
+
## Illustration asset
|
|
67
|
+
|
|
68
|
+
Current project assets: four full-size hero frames, three supporting frame atlases, and the supporting static posters in `site/assets/motion/`. Every hero frame and its poster are 2172 × 724; the supporting posters are 543 × 362. See [sequence provenance and prompts](site-illustration-sequences.md).
|
|
69
|
+
|
|
70
|
+
The original reference remains at `site/assets/workshop-miniature-diverse.webp` (2172 × 724; compressed WebP). The photographic miniature workshop has a racially diverse cast of six makers, retaining their activities and the tool-evaluation story. Generated with the built-in imagegen tool; see [the original edit prompt and provenance](site-artwork-miniature-diverse.md).
|
|
71
|
+
|
|
72
|
+
The previous miniature artwork remains at `site/assets/workshop-miniature.webp`, with its [original generation prompt](site-artwork-miniature.md).
|
|
73
|
+
|
|
74
|
+
The earlier linocut is retained at `site/assets/workshop-linocut.webp`; its [prompt and provenance](site-artwork-linocut.md) are preserved.
|
|
75
|
+
|
|
76
|
+
### Original illustration, retained
|
|
77
|
+
|
|
78
|
+
Original project asset: `site/assets/workshop.webp` (2172 × 724; compressed WebP).
|
|
79
|
+
|
|
80
|
+
Generated with the built-in imagegen tool. The uncompressed original is retained at `/Users/lazim/.codex/generated_images/01a0c9d4-a3f3-7e12-aa03-2b46f42b5e83/exec-08953716-7583-4a26-ad48-5262ab6f9ca7.png`. The website uses only the optimized local asset.
|
|
81
|
+
|
|
82
|
+
### Exact generation prompt
|
|
83
|
+
|
|
84
|
+
> Use case: illustration-story
|
|
85
|
+
> Asset type: original wide editorial illustration for the OpenMerit software website.
|
|
86
|
+
> Primary request: an understated hand-drawn illustrated frieze about testing, measuring, and choosing tools by their merit. A small open-air makers' workshop rendered as fine graphite and dark ink crosshatching, like a thoughtful 1950s scientific magazine illustration with very restrained watercolor.
|
|
87
|
+
> Scene: six small human craftspeople spaced naturally along one continuous rough wooden workbench and a low stone ground line. From left to right: a woman arriving with a small box of tools; a person carefully measuring a wooden object with calipers; an older seated craftsperson examining three different small mechanical tools on a table with a traditional balance scale; a woman recording observations in a notebook; a person adjusting a modest bench-mounted mechanism; and a person carrying the chosen tool away. A small dog sleeps underneath. Lived-in, patient, humane, charming but not whimsical. Ordinary human proportions. Objects are simply hand tools, no futuristic technology.
|
|
88
|
+
> Composition: ultra-wide landscape 3:1 aspect ratio. The entire scene lives in the bottom 55% of the frame. Full figures, heads and feet in frame. Broad empty paper above and very generous empty paper at left and right edges. Delicate fine lines, lots of negative space, no enclosing scenery or buildings. Thin continuous ground under the figures. The art should remain readable as a shallow panoramic strip at the bottom of a web hero.
|
|
89
|
+
> Color palette: background flat warm almost-white paper #f8f7f4, ink warm charcoal, very sparse washes of muted olive, faded terracotta, parchment. Mostly monochrome. No paper edge, no vignette, no noisy texture.
|
|
90
|
+
> Text: none. No lettering, logo, typography, watermark, UI, charts, robots, glowing effects, bright colors, gradients, 3D, flat vector blobs, or corporate cartoon style. This should look commissioned and drawn by an editorial illustrator, with subtle imperfections and varied poses.
|
|
91
|
+
|
|
92
|
+
## Full-site publication, September 28, 2026
|
|
93
|
+
|
|
94
|
+
All current website changes were deployed as Worker `6b9dfbd1-d9a0-4f13-9120-4e754774998a`. `npm run verify` passed before publishing, and Cloudflare repeated its build, demo, and documentation checks. All 95 public files matched the local build on both domains (190 byte-for-byte checks), covering the homepage, memo and artwork, 22 documentation guides and their source downloads, all five demo recordings, scripts, styles, fonts, and image assets. Both demo configuration responses report `enabled: false`; no inference session was started. The existing container image was unchanged and reused.
|
|
@@ -0,0 +1,83 @@
|
|
|
1
|
+
# Website documentation
|
|
2
|
+
|
|
3
|
+
The public documentation lives at https://openmerit.site/docs/. It uses the homepage's typography, colors, and Cloudflare hosting. Twenty-two curated guides, including a dedicated Demo group with an overview and five example walkthroughs, are generated from Markdown in `docs/`; internal website design and maintenance notes are excluded.
|
|
4
|
+
|
|
5
|
+
Article content has one maintained source: `docs/*.md`. Update those files alongside product changes, then rebuild and deploy. Do not edit the generated `site/docs/` files. The repository's `AGENTS.md` records this rule for future coding work.
|
|
6
|
+
|
|
7
|
+
## Reference and implementation
|
|
8
|
+
|
|
9
|
+
The structure draws on [Pi's documentation](https://pi.dev/docs/latest): a grouped left navigation, central article, right page outline, mobile disclosures, and browser search. OpenMerit's presentation and content are its own. The public Pi website repository is an older archived implementation, so it is not a reliable description of Pi's current deployment stack.
|
|
10
|
+
|
|
11
|
+
- `scripts/docs.config.mjs` controls navigation, routes, titles, and descriptions.
|
|
12
|
+
- `scripts/build-docs.mjs` renders Markdown with MarkdownIt and builds a MiniSearch index containing page and section destinations.
|
|
13
|
+
- `site/docs.css` supplies the responsive layout; `site/docs.js` adds search, copy feedback, and the active article outline.
|
|
14
|
+
- `site/docs-search.js` loads the search assets and recovers from an outdated page after deployment.
|
|
15
|
+
- All documentation pages currently share the homepage's default Open Graph/Twitter image. Separate card generation is disabled for now; the optional template is retained in `scripts/social-image.mjs`. See [sharing-image configuration](site-dynamic-og.md).
|
|
16
|
+
- `site/docs/` contains generated HTML, downloadable Markdown, the search index, the self-hosted search library and its license, and a manifest. Generated files are ignored by Git.
|
|
17
|
+
- Search loads its library and index on the first query of at least two characters. No search query is sent to a search service. Hashed filenames and versioned CSS/JavaScript URLs update with the source or supporting assets. If a tab still references assets removed by a newer deployment, search fetches the current manifest without using its cache and retries with the current library and index. Recovery uses a fresh module URL because browsers can cache failed imports. If recovery also fails, the error message lets the reader retry by typing again; retries do not loop in the background.
|
|
18
|
+
- Each page displays the package version and compact actions for viewing its source on GitHub, editing that source, and reading the generated Markdown. This is documentation for the current release, with no historical version selector yet.
|
|
19
|
+
|
|
20
|
+
## Pi component structure
|
|
21
|
+
|
|
22
|
+
The 2026-09-23 component pass follows the rendered [Pi documentation layout](https://pi.dev/docs/latest/quickstart), preserving OpenMerit's font, colors, logo, and article content. Search remains above the central article. Navigation and the article outline are independent panels, with the article in its own bordered reading panel below search. A shared documentation introduction sits above all three columns.
|
|
23
|
+
|
|
24
|
+
| Before | After | Why |
|
|
25
|
+
| --- | --- | --- |
|
|
26
|
+
| Navigation, article, and outline shared an undivided surface. | Each has a distinct panel with an appropriate header. | The page's reading and navigation components have clear boundaries. |
|
|
27
|
+
| Source and correction links occupied a separate metadata row. | Version, source, edit, and Markdown actions sit beside the article title and wrap below it on narrow screens. | Page-level tools are collected in one compact place. |
|
|
28
|
+
| Mobile navigation and the article outline appeared separately. | One native Navigation disclosure contains the current outline and an expandable Documentation list. | Both navigation scopes are available in one component. |
|
|
29
|
+
| Heading links navigated to sections. | They also copy the section URL with success feedback when clipboard access works. | Readers can share a precise reference while retaining native anchor behavior. |
|
|
30
|
+
|
|
31
|
+
At intermediate widths, the documentation rail remains visible and the article outline moves above the article. Narrow layouts use the combined navigation component; its disclosures and section links work without JavaScript. Search and code copying remain available. Previous/next page links are omitted to match Pi's component structure. The component pass introduces no runtime framework or separate article content.
|
|
32
|
+
|
|
33
|
+
The header contains the OpenMerit brand and Home/GitHub links. The sidebar begins directly with its navigation groups and ends with the package version. The repeated “Documentation” labels in the header and sidebar, and the sidebar's “Early development” link, were removed on 2026-09-23.
|
|
34
|
+
|
|
35
|
+
## Editing and previewing
|
|
36
|
+
|
|
37
|
+
Edit the Markdown source, then run:
|
|
38
|
+
|
|
39
|
+
```sh
|
|
40
|
+
npm ci
|
|
41
|
+
npm run check:docs
|
|
42
|
+
python3 -m http.server 45831 --bind 127.0.0.1 --directory site
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
Open http://127.0.0.1:45831/docs/. Rebuild after changing Markdown, navigation, CSS, or JavaScript. To add a page, create a Markdown document beginning with one H1 and add it to the appropriate group in `scripts/docs.config.mjs`. Relative links to published guides are rewritten to their website routes. Other relative Markdown links point to the repository.
|
|
46
|
+
|
|
47
|
+
`npm run verify` includes the documentation checks. Cloudflare runs `npm run check:docs` before each deployment, so an outdated generated folder or a broken internal link cannot silently pass through the normal deployment command. Publish using the command in [site design and deployment](site-design.md).
|
|
48
|
+
|
|
49
|
+
The GitHub CI workflow runs verification for pull requests and pushes to `main`; it does not deploy the website. A local edit or successful CI run appears on the public site after a Cloudflare deployment. This publishing step does not require editing the documentation a second time.
|
|
50
|
+
|
|
51
|
+
## Content review
|
|
52
|
+
|
|
53
|
+
Before a release or website update, review affected FAQ answers and guides against the published implementation using the [FAQ source map and maintenance procedure](site-faq-maintenance.md). Verify approval, budget, pause, scheduling, data-location, and readiness claims against their implementation. Keep known limitations visible until a released change and relevant verification justify updating them.
|
|
54
|
+
|
|
55
|
+
For a package release, update source guides and `package.json` together, then rebuild. The build keeps the website synchronized with the maintained Markdown; it does not independently discover changes in product behavior.
|
|
56
|
+
|
|
57
|
+
## Interaction and layout review
|
|
58
|
+
|
|
59
|
+
| Before | After | Why |
|
|
60
|
+
| --- | --- | --- |
|
|
61
|
+
| Documentation links led into the repository. | A consistent reading layout links between product guides. | Readers can follow setup, use, and reference material without losing their place. |
|
|
62
|
+
| Repository search was the only search route. | Search links directly to matching documentation sections and supports keyboard browsing. | Commands and operational topics are easier to find. |
|
|
63
|
+
| An open tab could lose search after a deployment removed its old search files. | Search recovers using the current deployment manifest. | Readers can keep searching across website updates. |
|
|
64
|
+
| Long navigation would crowd a narrow viewport. | Native disclosure controls expose navigation and the article outline. | Reading and navigation remain usable without JavaScript. |
|
|
65
|
+
| Code examples required manual text selection. | Copy buttons report success and select text if clipboard access fails. | Examples remain usable when clipboard permission is unavailable. |
|
|
66
|
+
|
|
67
|
+
Search uses a labeled 16px input, visible focus, native result links, and live status messages. Keyboard feedback is immediate. Hover styling is restricted to hover-capable fine pointers. Code and tables scroll within their own regions; the page keeps native vertical scrolling and permits zoom. Safe-area padding is included. Physical phone hardware has not been tested.
|
|
68
|
+
|
|
69
|
+
## Verification
|
|
70
|
+
|
|
71
|
+
The header/sidebar label removal was checked across all 16 generated pages and in Chrome at 1263px and 390px. The main documentation title and package version remain visible, the desktop screenshot was reviewed, and mobile navigation and budget search passed without overflow or runtime errors. All eight documentation checks passed, and 50 deployed URLs across both domains matched the local output.
|
|
72
|
+
|
|
73
|
+
The 2026-09-23 search review reproduced a production failure with the previous deployment's library and index URLs returning 404. The recovery fix passed locally and on the deployed site in Chrome and WebKit with those outdated URLs, plus separate interrupted index, library, and loader downloads followed by a successful retry. Mouse and emulated-touch result selection, command queries, lazy loading, the two-character minimum, arrow keys, Escape, empty results, and clearing a query while the index was loading passed. Four additional automated tests cover normal loading, recovery when either or both assets are missing, bounded failure, and invalid recovery URLs. `npm run check:docs` now runs eight checks.
|
|
74
|
+
|
|
75
|
+
The Pi component pass was checked in Chrome and WebKit across all 16 pages at 1440px and 320px, with no page overflow or runtime errors. Screenshots were reviewed at desktop, tablet, and phone-sized widths. Combined mobile navigation, nested documentation lists, outline links, Escape focus restoration, source/edit destinations, and native navigation without JavaScript passed. Section URL copying and code copying passed in Chrome; denied clipboard access preserves section navigation in both engines. Mouse and emulated-touch search selection passed in both engines. Axe WCAG 2 A/AA and 2.1 AA checks passed on the overview, budget tables, adapter examples, and open search at 1440px and 390px. Physical phone hardware was not tested.
|
|
76
|
+
|
|
77
|
+
The documentation checks validate every local homepage/docs link, asset reference, and section destination; one H1 and unique IDs; exact downloadable Markdown; exclusion of internal site notes; and useful search results for budgets, pause, and rollback.
|
|
78
|
+
|
|
79
|
+
Chrome browser checks covered all 16 pages at 1440px and 320px with HTTP 200 responses, one H1, and no page overflow. Screenshots covered 1440, 1263, 1024, 768, 390, and 320 CSS pixels. Search lazy loading, the keyboard shortcut, arrow-key navigation, Escape, no results, deep links, history navigation, clipboard success and denied-access fallback, mobile navigation, and article outlines passed. Search also recovered after a failed index request. Reading and mobile navigation passed with JavaScript disabled. No runtime errors were observed.
|
|
80
|
+
|
|
81
|
+
WebKit also loaded all 16 pages without overflow at 1440px and 390px. Search, keyboard result navigation, Escape, mobile navigation, section links, and the clipboard fallback passed without runtime errors. Axe WCAG 2 A/AA and 2.1 AA checks passed for the overview, budget table, adapter code examples, and open search results at 1440px and 390px after increasing code-label contrast. These are browser and automated accessibility checks, not exhaustive accessibility certification or physical-device validation.
|
|
82
|
+
|
|
83
|
+
Production smoke testing caught an event-ordering bug in pointer selection: a `focusout` microtask could close the result panel before the clicked link received focus. Closing the panel on a completed `focusin` outside search fixes that race. Mouse and touch result navigation, keyboard browsing, Escape, and moving focus outside search subsequently passed in both Chrome and WebKit. Include actual mouse and touch result selection in future browser reviews alongside keyboard tests.
|
|
@@ -0,0 +1,35 @@
|
|
|
1
|
+
# Page sharing images
|
|
2
|
+
|
|
3
|
+
Separate page images are disabled for now at the user's request. The homepage and all documentation pages use `https://openmerit.site/assets/openmerit-og-v1.png` for Open Graph and Twitter cards. No entries in `scripts/docs.config.mjs` currently enable `socialImage`.
|
|
4
|
+
|
|
5
|
+
The optional template and workshop artwork remain available for later use. Normal docs builds do not render a separate page card. The previously published `/docs/og.png` path serves a copy of the default image so cached sharing metadata still resolves.
|
|
6
|
+
|
|
7
|
+
The shared-image configuration is live as Cloudflare version `90bd0c29-5d32-4386-93df-a53e680d9cc5` (2026-09-23). All 11 documentation checks passed. Live Facebook/Twitter crawler checks confirmed default-image metadata and valid image responses on both domains; 54 deployed URLs matched local output.
|
|
8
|
+
|
|
9
|
+
## Optional template and build
|
|
10
|
+
|
|
11
|
+
`scripts/social-image.mjs` renders a reusable SVG template to PNG with the pinned `@resvg/resvg-js` development dependency. It combines the text-free workshop asset at `site/assets/openmerit-og-workshop-v1.png` with a heading in the existing `site/assets/inter-latin.woff2` font. The font's SIL Open Font License is retained. There are no network requests, system-font dependencies, or image-generation API calls during builds. The source artwork and exact edit prompt are recorded in [the workshop background notes](site-og-workshop.md).
|
|
12
|
+
|
|
13
|
+
For an entry with `socialImage: true`, `scripts/build-docs.mjs` passes the maintained page metadata into the renderer, which draws the page title over the workshop. The page description remains in the Open Graph and Twitter description tags. `npm run build:docs` and `npm run check:docs` generate enabled cards, and the existing Cloudflare deployment runs the latter before publishing.
|
|
14
|
+
|
|
15
|
+
Enabled cards are generated during the build and served as PNGs. They are not rendered separately for each visitor or sharing request. Both default and custom images have an absolute `og:image` URL, PNG type, dimensions, alt text, and a Twitter large-image card. Page titles and descriptions stay specific to each documentation page.
|
|
16
|
+
|
|
17
|
+
## URLs and future pages
|
|
18
|
+
|
|
19
|
+
When enabled for the overview, a custom image is written to `site/docs/og.png`, alongside the generated page, and assigned through `https://openmerit.site/docs/og.png?v=<content-hash>`. The hash comes from the rendered PNG, so changing the title, workshop artwork, font, or template changes the URL when it changes the pixels. Rebuilding the same card produces the same URL. Description or package-version updates do not invalidate an unchanged custom image; those details are not printed on the card.
|
|
20
|
+
|
|
21
|
+
The physical PNG path stays stable. Previously shared image URLs with an older query parameter still resolve after a deployment instead of pointing to a deleted filename. Social platforms control their own preview caches and may retain a previously fetched card.
|
|
22
|
+
|
|
23
|
+
To re-enable the optional template after a future user request, add `socialImage: true` to a page's entry in `scripts/docs.config.mjs` and rebuild; the generator assigns an image under that page's route. Do not edit `site/docs/og.png` or the generated HTML.
|
|
24
|
+
|
|
25
|
+
The template measures and wraps the heading into at most two lines in the upper-left area. If a future title cannot fit, the build fails rather than silently cropping the text or covering the makers. Review new cards at full size and as thumbnails when expanding the rollout.
|
|
26
|
+
|
|
27
|
+
## Verification and references
|
|
28
|
+
|
|
29
|
+
The documentation checks validate the assigned Open Graph/Twitter URLs, PNG signature and 1200 × 630 dimensions, image hash, alt text, deterministic regeneration, heading updates, stable artwork when unprinted metadata changes, escaping, and rejection of text that cannot fit. The rendered `/docs/` card is visually reviewed before publishing.
|
|
30
|
+
|
|
31
|
+
The initial generic card shipped on 2026-09-23 as Cloudflare version `e7791e53-4f0c-4666-89d6-f7d9bc72ae4d`; it was then restyled to match the default workshop OG image at the user's request. Current deployment and live verification are recorded in [site design notes](site-design.md).
|
|
32
|
+
|
|
33
|
+
The workshop revision shipped as `921eb118-0539-4959-ad5f-e9b928e76d70`, with assigned image `https://openmerit.site/docs/og.png?v=b8404648c9bd`. Full-size and 600 × 315 previews were reviewed against the default image. All 11 documentation checks and script syntax checks passed. Live metadata and image bytes matched for Facebook and Twitter crawler user agents on both domains, older query URLs still resolved, and 54 deployed URLs matched the local build.
|
|
34
|
+
|
|
35
|
+
Implementation references: [Open Graph metadata and image properties](https://ogp.me/) and [resvg-js](https://github.com/thx/resvg-js). Font provenance and license are recorded in the site assets and [site design notes](site-design.md).
|
|
@@ -0,0 +1,115 @@
|
|
|
1
|
+
# Website FAQ maintenance
|
|
2
|
+
|
|
3
|
+
The public FAQ lives in `site/index.html`. Its answers describe the published package, including known limitations, rather than planned capabilities or unreleased fixes.
|
|
4
|
+
|
|
5
|
+
Last factual review: 2026-09-23, against npm `openmerit@0.1.5`, the Pi 0.87 adapter, and the current source. The FAQ contains 13 questions. The earlier audit reproduced a task-triggered assessment while `/openmerit status` reported paused and confirmed the same dispatch path in the published package.
|
|
6
|
+
|
|
7
|
+
Source review on 2026-09-27: unreleased post-build target detection, harness-inferred user-confirmed metric sampling plans, incumbent and comparison handoffs, setup-time model leads, deterministic core baseline-window checks, macOS Pi LaunchAgent provisioning, and durable pause are under development. The homepage intentionally continues to describe published npm `openmerit@0.1.5`; update it only when a release containing those changes is published and verified.
|
|
8
|
+
|
|
9
|
+
The experimental `npm run demo` sandbox is a repository-local live-provider demonstration, not a released package workflow. Its end-to-end result and tiny pilot sample do not change the homepage's shipping-behavior or reliability claims.
|
|
10
|
+
|
|
11
|
+
## Sources to review when behavior changes
|
|
12
|
+
|
|
13
|
+
| FAQ anchor | Claims to verify | Primary sources |
|
|
14
|
+
| --- | --- | --- |
|
|
15
|
+
| `faq-agent` | Harness responsibility; limits of result validation | `packages/pi/src/index.ts` (`privateIntentMessage`), `packages/core/src/index.ts` (`interpretIntentResult`), `docs/architecture.md` |
|
|
16
|
+
| `faq-harnesses` | Implemented adapters; supported Pi versions | `packages/`, root `package.json` peer dependencies, `docs/adapter-guide.md` |
|
|
17
|
+
| `faq-evaluations` | Interactive setup; task-specific instrumentation | Pi `session_start`, core `establish_evals` outcome and observability validation, `docs/getting-started.md` |
|
|
18
|
+
| `faq-comparisons` | User objectives; candidate tradeoffs; no promised savings | Core `verifyFrontierSnapshot`, `docs/pareto-spec.md` |
|
|
19
|
+
| `faq-evidence` | Required samples; unresolved candidates; evidence trust | Core feasibility and frontier verification, `docs/metrics-and-evidence.md`, `docs/security.md` |
|
|
20
|
+
| `faq-approval` | Opt-in automation; zero-capable graduation threshold; post-swap evaluation | Core `nextImprovementAction`, `packages/protocol/src/schemas.ts` (`SupervisedGraduationPolicySchema`), Pi `approve` command |
|
|
21
|
+
| `faq-regressions` | Post-swap regression handling; conditional rollback | Core `interpretIntentResult` and `nextImprovementAction`, `docs/lifecycle.md` |
|
|
22
|
+
| `faq-background` | Active-session signals; no built-in persistent scheduler | Pi `PI_HARNESS_DESCRIPTOR`, core `buildAutomationPlan`, `docs/automation.md` |
|
|
23
|
+
| `faq-pause` | Which automatic dispatches stop; restart behavior; active work | Pi `reportAutomationSignal`, `createInitialUiState`, completion handler and `pause` command; `docs/operations.md` |
|
|
24
|
+
| `faq-data` | Project files; Pi session entries; harness/provider data flows | `packages/core/src/store.ts`, Pi `appendEntry` and `privateIntentMessage`, `docs/security.md` |
|
|
25
|
+
| `faq-cost` | License; npm availability; provider and infrastructure costs | `LICENSE`, published npm metadata, harness execution path |
|
|
26
|
+
| `faq-budget` | Budget passed to harness; absence or presence of metering/cutoffs | Core `createOpenMeritIntent`, budget consumers across `packages/` |
|
|
27
|
+
| `faq-readiness` | Implemented, tested, and product-validated behavior | `docs/testing.md`, `docs/roadmap.md`, actual product-trial evidence |
|
|
28
|
+
|
|
29
|
+
## Review during releases and website updates
|
|
30
|
+
|
|
31
|
+
1. Check the current published npm version and compare its behavior with the answers above. Keep unreleased behavior clearly labeled; update version-specific caveats only when the relevant release is available.
|
|
32
|
+
2. Review affected answers whenever adapters, setup, evidence checks, approval, rollback, scheduling, pause, storage, budget handling, or readiness change. Read implementation and meaningful test evidence before strengthening a claim.
|
|
33
|
+
3. Keep the linked operational documentation consistent. In particular, a pause fix needs coverage for task signals, measurement signals, session-start signals, lifecycle follow-ups, and restart behavior before the FAQ can promise a complete pause.
|
|
34
|
+
4. Check the expanded disclosures at desktop and narrow mobile widths, including keyboard operation, links, and JavaScript-disabled reading. Verify the served FAQ after publishing.
|
|
35
|
+
5. Record the reviewed version, date, and any remaining limits here.
|
|
36
|
+
|
|
37
|
+
## Wording boundaries
|
|
38
|
+
|
|
39
|
+
- Separate issuing a request from the harness successfully executing it.
|
|
40
|
+
- Distinguish structured-result and comparison validation from independent proof that evidence is authentic.
|
|
41
|
+
- Describe a budget as a harness constraint until metering and spending cutoffs exist.
|
|
42
|
+
- Describe verification as post-swap evaluation, not a guarantee that a change is safe before it is applied.
|
|
43
|
+
- Claim local storage only for the relevant OpenMerit files; account for Pi session history and configured external systems.
|
|
44
|
+
- Keep ordinary coding checkpoints distinct from task-specific product measurements.
|
|
45
|
+
|
|
46
|
+
## Inference Terminal — unreleased source, 7 October 2026
|
|
47
|
+
|
|
48
|
+
Reviewed the homepage evidence, background, data, cost, and readiness claims against the new terminal. It reads existing project metadata and loopback runtime inventory on a bounded interval and samples whole-machine CPU/memory. It makes no inference calls, routes no traffic, runs no evaluations, and stops collecting when its CLI exits. Sample data is labeled and separate. Credential-variable detection reveals provider names only; it never uses credentials for provider calls.
|
|
49
|
+
|
|
50
|
+
Deployments, reported path hops, shared pools, additional hosts, sites, and racks come from explicit configuration or bounded exports. Source timestamps determine freshness. Agent/task names come only from recorded metadata; source code alone does not establish execution. Full local hostnames are system-reported. Missing identities and readings remain explicit. No network discovery, model verification, routing enforcement, available-headroom guarantee, or per-agent resource attribution is implied.
|
|
51
|
+
|
|
52
|
+
Top navigation, shared dropdowns, model/service images, path colors, chart interactions, and denser source rows change presentation only. Images and animation libraries are bundled locally. This remains an unreleased command; published package behavior and homepage FAQ claims are unchanged. No homepage wording changes are needed before release.
|
|
53
|
+
|
|
54
|
+
## Hosted terminal sample — 7 October 2026
|
|
55
|
+
|
|
56
|
+
The separate `dash.openmerit.site` preview serves generated fictional data and terminal assets only. It has no collector, filesystem access, provider credentials, inference endpoint, upload path, or connection to a visitor's local environment. Public access is by link; noindex is a search-engine request, not authentication. The root site's FAQ and published npm claims remain unchanged. The deployment uses its own Worker and exact custom domain, and does not redeploy the main site or its demo/memo services.
|
|
57
|
+
|
|
58
|
+
The hosted preview’s dedicated sharing card and `Inference Terminal by OpenMerit` title change presentation only. Its social description still identifies the environment as fictional; homepage FAQ and package claims are unchanged.
|
|
59
|
+
|
|
60
|
+
## Inference service identity — unreleased source, 8 October 2026
|
|
61
|
+
|
|
62
|
+
Baseten and Cerebras credential-name detection joins Groq and the existing services. Known HTTPS endpoint overrides identify configuration independently of SDK vendor; unknown overrides suppress credential-name attribution. Only inherited variables are checked, and no credential values or raw URLs reach the snapshot. No cloud calls, account access, source-code scanning, `.env` loading, or physical hardware inference is introduced. Model creators, deployment modes, and accelerator vendor/kind come from explicit metadata. Configured services remain distinct from observed requests. The hosted sample uses fictional examples; package and homepage claims remain unchanged.
|
|
63
|
+
|
|
64
|
+
## Services inventory — unreleased source, 8 October 2026
|
|
65
|
+
|
|
66
|
+
Reviewed the new Services view against evidence, background, data, and readiness claims. It reorganizes existing metadata into a service inventory and a secondary environment map. Model creators alone never establish service connections; registry roles are subordinate to explicit roles. Selected-window service counts are nonadditive and configured-only entries remain distinct from activity. Recognized endpoint hostnames can now appear as evidence; deployment identifiers, URL paths, queries, and credential values remain excluded. Collection stays passive. This does not release the CLI or change the homepage FAQ's published-package claims.
|
|
67
|
+
|
|
68
|
+
## Automatic terminal discovery — unreleased source, 8 October 2026
|
|
69
|
+
|
|
70
|
+
Reviewed data, evidence, approval, background, and readiness claims. The terminal now checks four fixed project declaration files and bounded known loopback API candidates automatically. SDK presence never establishes usage. Optional Pi session-format-3 metadata needs explicit CLI consent saved outside the repository and tied to canonical project/source paths; session files containing conversation text are parsed only after consent, and only usage metadata is retained. Revocation clears it on the next poll. Activity exports, Pi history, and audit summaries use precedence rather than additive totals. There are no cloud usage calls or general filesystem/process scans. This supersedes the earlier source-only note that `.env` and home history were never read. Published-package behavior, hosted-sample isolation, and homepage claims remain unchanged; no FAQ wording change is needed before release.
|
|
71
|
+
|
|
72
|
+
|
|
73
|
+
## Native-history request counts — unreleased source, 8 October 2026
|
|
74
|
+
|
|
75
|
+
Reviewed evidence and readiness claims after the dedicated Linux VM run. Overview and model request counts now recognize approved Pi history and retained requests from partially covered sources. Collection and permissions are unchanged. The real run is bounded local-model test evidence, not a claim of universal discovery or production readiness. Published-package and homepage FAQ claims remain unchanged.
|
|
76
|
+
|
|
77
|
+
|
|
78
|
+
## Pi connection names — unreleased source, 9 October 2026
|
|
79
|
+
|
|
80
|
+
Reviewed evidence and data claims. Pi's existing assistant-message provider ID now supplies its recorded connection name; classification uses the existing unambiguous inventory/runtime matcher. Generic serving-provider fields still do not establish routes, and physical deployments, hidden hops, timing, and agent names remain uninferred. No additional files, credentials, inference calls, or permissions are introduced. Published-package and homepage FAQ claims remain unchanged.
|
|
81
|
+
|
|
82
|
+
|
|
83
|
+
## Environment map coverage — unreleased source, 9 October 2026
|
|
84
|
+
|
|
85
|
+
Reviewed evidence and readiness claims. The map now selects representatives around every recorded entry connection in the selected window or agent scope. It no longer drops connections behind workload or model layout caps. Larger maps scroll, and represented request/workload coverage remains explicit. This changes presentation only; no inference, discovery, permission, billing, or published-package behavior changes. Homepage FAQ wording is unchanged.
|
|
86
|
+
|
|
87
|
+
|
|
88
|
+
## Complete environment map — unreleased source, 9 October 2026
|
|
89
|
+
|
|
90
|
+
The map no longer selects representative workloads or models. It groups all retained activity by identical recorded entry connections and lets users expand every member. Request counts and relationships stay complete in collapsed and expanded views. Compact stacks preserve the original map scale and show a real member name with a count of remaining members. Model nodes display reported services and explicitly matched deployment information; absent deployment identities are explained in inspectors. No source scanning, inferred topology, new permissions, or inference calls are introduced. Homepage FAQ claims about the published package remain unchanged.
|
|
91
|
+
|
|
92
|
+
|
|
93
|
+
## Optional graph component experiment — unreleased source, 9 October 2026
|
|
94
|
+
|
|
95
|
+
The `map=flow` URL flag enables a separately bundled React Flow + ELK renderer over the same retained metadata. The existing map remains the default, with immediate return through Original map. Search and member highlighting change presentation only; grouping, request totals, source collection, permissions, and sample/live separation are unchanged. This is an opt-in experiment, not a claim of production scale or released package behavior. Homepage FAQ wording remains unchanged.
|
|
96
|
+
|
|
97
|
+
|
|
98
|
+
## Existing activity audit — unreleased source, 9 October 2026
|
|
99
|
+
|
|
100
|
+
Reviewed homepage evidence, data, background, approval, and readiness claims against automatic bounded project JSONL readers and multi-source reconciliation. The terminal now reads supported existing request logs inside the selected project and combines their metadata with approved native Pi usage and explicit feeds. Shared recorded identities prevent double counting; unresolved overlaps and conflicting measurements are disclosed. Pi's bounded file selection increases to 128 sessions; permission scope and revocation are unchanged. No inference calls, provider account ingestion, arbitrary process/database scanning, or framework-wide automatic support is implied. Published package and homepage FAQ claims remain unchanged. VM acceptance evidence is recorded in the internal terminal notes only after the run succeeds.
|
|
101
|
+
|
|
102
|
+
|
|
103
|
+
## Application-only terminal scope — unreleased source, 9 October 2026
|
|
104
|
+
|
|
105
|
+
Reviewed evidence, data, approval, and readiness claims. The terminal now excludes coding-harness usage: the Pi history reader, saved-grant lookup, and access prompts are removed. Prior history-related entries in this audit are historical and superseded. Application agents, jobs, evaluations, supported request logs, and OpenMerit product-task evidence remain in scope. Explicit development-scoped records are rejected, but unmarked request data cannot be classified by purpose automatically. Runtime inventories and whole-machine measurements remain shared infrastructure evidence, not project-exclusive usage. Published package and homepage FAQ wording remain unchanged; the terminal is still unreleased.
|
|
106
|
+
|
|
107
|
+
|
|
108
|
+
## Terminal test installation — preview, 9 October 2026
|
|
109
|
+
|
|
110
|
+
Reviewed setup, background, data, and readiness claims. A prebuilt test package on the existing sample domain supports a single global npm install, followed by `openmerit dash` in the selected application directory. The CLI opens the default desktop browser, keeps a printed URL for headless sessions, and falls back to a free loopback port only when no explicit port was requested. Collection scope, file access, and consent boundaries are unchanged. This does not publish or change the stable npm release; homepage FAQ claims remain about the published package. Platform URL-handler code is tested separately from actual desktop acceptance, and preview verification must not be presented as broad production readiness.
|
|
111
|
+
|
|
112
|
+
|
|
113
|
+
## npm terminal preview delivery — prepared, 9 October 2026
|
|
114
|
+
|
|
115
|
+
The terminal prerelease is prepared in the existing `openmerit` npm package as `0.1.6-preview.0`, targeting the `preview` tag. `latest` remains `0.1.5`; successful registry publication and install verification must be recorded before claiming the new command is available through npm. The compiled package retains core/protocol/Pi exports and uses an explicit dist-directory allowlist so a sample-site build cannot accidentally add hosted assets or a nested installer archive. The standalone domain-hosted packaging path is removed from the source build. This changes delivery only: collection behavior, browser launch, consent, and application scope remain unchanged. Homepage stable-package claims are unchanged.
|