openmerit 0.1.4 → 0.1.6-preview.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +40 -0
- package/README.md +121 -386
- package/dist/core/src/index.d.ts +101 -0
- package/dist/core/src/index.js +1649 -0
- package/dist/core/src/store.d.ts +35 -0
- package/dist/core/src/store.js +102 -0
- package/dist/pi/src/index.d.ts +32 -0
- package/dist/pi/src/index.js +794 -0
- package/dist/pi/src/scheduler.d.ts +11 -0
- package/dist/pi/src/scheduler.js +137 -0
- package/dist/pi/src/wakeup.d.ts +2 -0
- package/dist/pi/src/wakeup.js +108 -0
- package/dist/protocol/src/index.d.ts +484 -0
- package/dist/protocol/src/index.js +47 -0
- package/dist/protocol/src/schemas.d.ts +576 -0
- package/dist/protocol/src/schemas.js +280 -0
- package/dist/terminal/public/app.js +297 -0
- package/dist/terminal/public/brands/anthropic.png +0 -0
- package/dist/terminal/public/brands/baai.png +0 -0
- package/dist/terminal/public/brands/baseten.png +0 -0
- package/dist/terminal/public/brands/cerebras.png +0 -0
- package/dist/terminal/public/brands/cohere.png +0 -0
- package/dist/terminal/public/brands/deepseek.ico +0 -0
- package/dist/terminal/public/brands/google.png +0 -0
- package/dist/terminal/public/brands/groq.ico +0 -0
- package/dist/terminal/public/brands/lm-studio.png +0 -0
- package/dist/terminal/public/brands/meta.ico +0 -0
- package/dist/terminal/public/brands/mistral.png +0 -0
- package/dist/terminal/public/brands/nomic.png +0 -0
- package/dist/terminal/public/brands/ollama.png +0 -0
- package/dist/terminal/public/brands/openai.png +0 -0
- package/dist/terminal/public/brands/openrouter.png +0 -0
- package/dist/terminal/public/brands/qwen.png +0 -0
- package/dist/terminal/public/brands/vllm.ico +0 -0
- package/dist/terminal/public/brands/vllm.png +0 -0
- package/dist/terminal/public/favicon.svg +1 -0
- package/dist/terminal/public/flow.css +1 -0
- package/dist/terminal/public/flow.js +770 -0
- package/dist/terminal/public/index.html +21 -0
- package/dist/terminal/public/styles.css +779 -0
- package/dist/terminal/src/activity-merge.mjs +64 -0
- package/dist/terminal/src/browser.mjs +29 -0
- package/dist/terminal/src/cli.mjs +60 -0
- package/dist/terminal/src/collect.mjs +311 -0
- package/dist/terminal/src/discovery.mjs +93 -0
- package/dist/terminal/src/hardware.mjs +57 -0
- package/dist/terminal/src/project-activity.mjs +156 -0
- package/dist/terminal/src/sample.mjs +171 -0
- package/dist/terminal/src/server.mjs +56 -0
- package/dist/terminal/src/services.mjs +62 -0
- package/dist/terminal/src/topology.mjs +30 -0
- package/docs/adapter-guide.md +189 -0
- package/docs/architecture.md +59 -0
- package/docs/automation.md +74 -0
- package/docs/budgets.md +37 -0
- package/docs/commands.md +85 -0
- package/docs/demo-backfill.md +29 -0
- package/docs/demo-fieldkit.md +47 -0
- package/docs/demo-placement.md +30 -0
- package/docs/demo-spam.md +15 -0
- package/docs/demo-support.md +42 -0
- package/docs/demo.md +57 -0
- package/docs/first-trial.md +60 -0
- package/docs/getting-started.md +65 -0
- package/docs/index.md +40 -0
- package/docs/inference-terminal.md +439 -0
- package/docs/lifecycle.md +30 -0
- package/docs/memo.md +126 -0
- package/docs/metrics-and-evidence.md +48 -0
- package/docs/operations.md +40 -0
- package/docs/pareto-spec.md +76 -0
- package/docs/pi-extension.md +54 -0
- package/docs/roadmap.md +28 -0
- package/docs/security.md +37 -0
- package/docs/site-artwork-linocut.md +23 -0
- package/docs/site-artwork-miniature-diverse.md +28 -0
- package/docs/site-artwork-miniature.md +26 -0
- package/docs/site-demo.md +177 -0
- package/docs/site-design.md +94 -0
- package/docs/site-documentation.md +83 -0
- package/docs/site-dynamic-og.md +35 -0
- package/docs/site-faq-maintenance.md +115 -0
- package/docs/site-hero-resolution.md +60 -0
- package/docs/site-illustration-sequences.md +227 -0
- package/docs/site-inference-terminal.md +203 -0
- package/docs/site-memo.md +39 -0
- package/docs/site-og-image.md +38 -0
- package/docs/site-og-workshop.md +21 -0
- package/docs/site-section-artwork.md +56 -0
- package/docs/site-skill-review.md +57 -0
- package/docs/site-terminal-preview.md +85 -0
- package/docs/testing.md +118 -0
- package/docs/troubleshooting.md +55 -0
- package/docs/ux-reference.md +32 -0
- package/package.json +74 -42
- package/benchmark/invoice_ocr/data/invoice_01_ground_truth.json +0 -38
- package/benchmark/invoice_ocr/data/invoice_01_row_2.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_02_ground_truth.json +0 -32
- package/benchmark/invoice_ocr/data/invoice_02_row_5.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_03_ground_truth.json +0 -26
- package/benchmark/invoice_ocr/data/invoice_03_row_6.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_04_ground_truth.json +0 -26
- package/benchmark/invoice_ocr/data/invoice_04_row_7.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_05_ground_truth.json +0 -38
- package/benchmark/invoice_ocr/data/invoice_05_row_947.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_06_ground_truth.json +0 -38
- package/benchmark/invoice_ocr/data/invoice_06_row_948.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_07_ground_truth.json +0 -20
- package/benchmark/invoice_ocr/data/invoice_07_row_949.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_08_ground_truth.json +0 -38
- package/benchmark/invoice_ocr/data/invoice_08_row_1888.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_09_ground_truth.json +0 -26
- package/benchmark/invoice_ocr/data/invoice_09_row_1890.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_10_ground_truth.json +0 -20
- package/benchmark/invoice_ocr/data/invoice_10_row_1892.jpg +0 -0
- package/benchmark/invoice_ocr/data/manifest.json +0 -97
- package/dist/benchmarks.js +0 -98
- package/dist/catalog.js +0 -61
- package/dist/cli.js +0 -188
- package/dist/daemon.js +0 -407
- package/dist/diagnostics.js +0 -227
- package/dist/frontier.js +0 -56
- package/dist/harness.js +0 -1
- package/dist/integrations.js +0 -19
- package/dist/invoice-eval.js +0 -33
- package/dist/invoice-score.js +0 -124
- package/dist/judge.js +0 -43
- package/dist/llm.js +0 -207
- package/dist/pi-config.js +0 -46
- package/dist/pi-trials.js +0 -373
- package/dist/policy.js +0 -185
- package/dist/providers.js +0 -1
- package/dist/recommend.js +0 -76
- package/dist/routes.js +0 -74
- package/dist/standalone.js +0 -224
- package/dist/store.js +0 -89
- package/dist/strategist.js +0 -68
- package/dist/task-input.js +0 -54
- package/dist/traces.js +0 -127
- package/dist/trials.js +0 -140
- package/dist/types.js +0 -2
- package/examples/invoice-prompt.txt +0 -19
- package/examples/task.example.json +0 -7
- package/extension/openmerit.ts +0 -947
- package/instructions/OPENMERIT.md +0 -63
- package/instructions/openmerit.policy.json +0 -37
- package/rules.md +0 -43
package/docs/demo.md
ADDED
|
@@ -0,0 +1,57 @@
|
|
|
1
|
+
# Demo
|
|
2
|
+
|
|
3
|
+
The [demo player](https://openmerit.site/demo) shows recorded, supervised OpenMerit comparisons. The guides in this section explain each example's inputs, model assignments, evaluation policy, and outcome. These are early development demonstrations from this checkout; they do not establish general product reliability or behavior in the published npm release.
|
|
4
|
+
|
|
5
|
+
## Choose a walkthrough
|
|
6
|
+
|
|
7
|
+
| Scope | Demo | How it is structured |
|
|
8
|
+
| --- | --- | --- |
|
|
9
|
+
| Tasks | [Support tickets](demo-support.md) | One model call turns an email into a ticket; three models face the same exact-match checks. |
|
|
10
|
+
| Tasks | [Spam filter · Jev](demo-spam.md) | A chat model and a typed-decision specialist classify comments under the same promotional-spam policy. |
|
|
11
|
+
| Agents | [Backfill recovery](demo-backfill.md) | Three dependent turns inspect an incident, read a checkpoint, and prepare a recovery manifest. |
|
|
12
|
+
| Agents | [GPU placement](demo-placement.md) | Three dependent turns read requirements, inspect capacity, and choose a pool under six constraints. |
|
|
13
|
+
| Agents | [Field kit · Six tasks](demo-fieldkit.md) | A prepared six-model workflow translates, reads an image, extracts HTML, inspects code, checks feasibility, and writes a handoff. Both paths use six distinct models. |
|
|
14
|
+
|
|
15
|
+
## How a comparison works
|
|
16
|
+
|
|
17
|
+
Every recorded comparison uses a prepared application, synthetic fixtures, a grading rubric, instrumentation, API adapters, and a bounded model shortlist. Both application paths receive the same inputs and task rules. The **without OpenMerit** path keeps its original configuration. The managed path follows this sequence:
|
|
18
|
+
|
|
19
|
+
1. **Measure a baseline.** Pi runs the incumbent on the fixed evaluation suite and saves outputs, costs, latency, and grading results.
|
|
20
|
+
2. **Compare challengers.** The same suite measures alternative models or complete model routes. OpenMerit checks the returned evidence references, result contracts, and frontier conformance. A cheap candidate that fails a requirement cannot qualify.
|
|
21
|
+
3. **Request approval.** The prepared demo policy proposes the lowest observed cost among eligible frontier candidates only when it is below the incumbent cost. Keeping the incumbent is a valid outcome.
|
|
22
|
+
4. **Apply and verify.** After operator approval, Pi changes the application's configuration and checks it on separate inputs. The confirmed policy requests restoration of the incumbent if verification fails.
|
|
23
|
+
|
|
24
|
+
The agent examples compare supplied route configurations as whole workflows. They do not search all possible per-task model assignments. Pi's orchestration model stays fixed. The adapters, fixtures, and helpers are prepared in advance; OpenMerit does not generate those integrations in these demos.
|
|
25
|
+
|
|
26
|
+
Pi invokes bounded helpers and submits saved artifacts by name through a demo-specific transport bridge. The bridge expands those references into the adapter's canonical results and signals. The actual OpenMerit lifecycle then validates and advances the comparison. Missing provider cost stops a comparison instead of being treated as zero; a partial run is not a successful outcome.
|
|
27
|
+
|
|
28
|
+
The unreleased source runner now supplies explicit sampling plans, metric aggregations, and an incumbent descriptor to match the current protocol. Candidate limits count the incumbent and challengers together. Its comparison manifest records the fixed fixtures, rubric, model routes, and actual run order; the incumbent's baseline is reused in the comparison and included in evaluation-spend accounting. Runs are not randomized. These source changes do not alter the existing recordings or establish a new live validation result.
|
|
29
|
+
|
|
30
|
+
## Playback controls
|
|
31
|
+
|
|
32
|
+
The player is for recorded playback. It makes no model calls, starts no sandbox, and accepts no custom inputs. The **How it works** link opens the guide for the selected example.
|
|
33
|
+
|
|
34
|
+
- **Pause / Play** stops or resumes playback. Seeking on the timeline pauses at the selected position.
|
|
35
|
+
- **Replay** starts the recording again from the beginning.
|
|
36
|
+
- **Reset** returns to the beginning paused and closes Run details.
|
|
37
|
+
- **Run details** shows the evaluation policy, sample input and outputs, provenance, and **Download evidence**, which saves the selected recording as JSON.
|
|
38
|
+
|
|
39
|
+
Switching task or scope starts the selected recording from the beginning and closes Run details. Returning to a scope restores its last selected example. Left/Right arrows and Home/End move within each tab row. Task-specific URLs open the matching scope and example. Playback stops at the end and does not advance while the browser tab is hidden.
|
|
40
|
+
|
|
41
|
+
The available recordings were captured on September 28, 2026. Support and spam ran in Cloudflare; backfill, GPU placement, and field kit ran locally with real OpenRouter calls, as labeled in their provenance. The 36-second presentation compresses the original timing. Operator approval, outputs, and measurements are captured events, not new work performed while watching.
|
|
42
|
+
|
|
43
|
+
## Reading the results
|
|
44
|
+
|
|
45
|
+
Each example has a short description below its tabs explaining the task and intended result. It remains visible throughout playback, including when paused, completed, or reset.
|
|
46
|
+
|
|
47
|
+
The main view shows the decision, both application paths, and the run trace. A single-task demo shows one model per path; an agent recording shows each task's model. The managed assignments stay on the baseline until the recorded change is applied. Resetting or seeking backward restores the corresponding earlier state.
|
|
48
|
+
|
|
49
|
+
Quality is the number of matching fixtures or completed goals in the finite suite. Cost is the observed mean application cost multiplied by 1,000 and rounded to $0.001; it is not a separate 1,000-input test. Latency is mean end-to-end application request time, including the complete sequential workflow for agents, shown to two decimal places in seconds. Percentage changes use the original unrounded measurements, and slower outcomes are labeled as higher latency.
|
|
50
|
+
|
|
51
|
+
Application costs exclude Pi orchestration and infrastructure. GPU quotes inside a placement plan are synthetic job data, separate from measured inference costs. Evidence downloads retain original unrounded measurements and provenance; the player presents a curated recording rather than the full temporary runner filesystem.
|
|
52
|
+
|
|
53
|
+
## Scope and limitations
|
|
54
|
+
|
|
55
|
+
These examples use small synthetic suites, supplied prompts, fixed helpers, and preselected candidates. Their observed results are not production success-rate estimates, confidence bounds, or promised savings. Recovery manifests are never executed, GPU capacity is never reserved, and field-kit scripts are never run. Watching downloads static assets and recorded synthetic evidence; it sends no inputs to model providers.
|
|
56
|
+
|
|
57
|
+
For your own workload, follow [your first trial](first-trial.md). The [architecture guide](architecture.md) explains the boundary between OpenMerit and the coding harness.
|
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
# Your first supervised trial
|
|
2
|
+
|
|
3
|
+
Use one representative product task to learn whether OpenMerit and your harness can collect useful evidence. This guide describes a trial to perform; it does not claim that a trial has already succeeded for your product.
|
|
4
|
+
|
|
5
|
+
## Choose a bounded task
|
|
6
|
+
|
|
7
|
+
Pick a task with outcomes you can inspect, such as extracting fields from a document, answering questions over a known collection, or completing a repeatable coding task. Identify the current model configuration so it can be restored if a challenger regresses.
|
|
8
|
+
|
|
9
|
+
Choose examples that resemble actual use, including difficult cases. A single successful example cannot establish a success rate or reliability claim.
|
|
10
|
+
|
|
11
|
+
## Confirm the setup
|
|
12
|
+
|
|
13
|
+
Complete [installation](getting-started.md) and open the product repository in an interactive Pi session. During setup, confirm:
|
|
14
|
+
|
|
15
|
+
- the required quality, cost, latency, and other metrics;
|
|
16
|
+
- the sample requirements and unacceptable outcomes;
|
|
17
|
+
- the evaluation budget and any candidate or run limits;
|
|
18
|
+
- when reassessments and post-swap checks should happen;
|
|
19
|
+
- whether model changes or rollback may happen automatically.
|
|
20
|
+
|
|
21
|
+
For a supervised first trial, leave automatic swaps disabled. A zero graduation threshold can allow an immediate automatic swap if automation is enabled. Review [budgets and permissions](budgets.md) before confirming.
|
|
22
|
+
|
|
23
|
+
## Inspect the evaluations
|
|
24
|
+
|
|
25
|
+
Ask Pi to show the evaluation artifacts and instrumentation it created, explain what they measure, and demonstrate a representative run. Check that all required metrics have a source and that evaluation data is appropriate to send to the configured providers.
|
|
26
|
+
|
|
27
|
+
```text
|
|
28
|
+
/openmerit status
|
|
29
|
+
/openmerit logs
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
These commands help locate readiness and evidence. They do not independently certify that an evaluation is correct. Inspect the artifacts and measurements themselves.
|
|
33
|
+
|
|
34
|
+
## Collect and compare
|
|
35
|
+
|
|
36
|
+
Use the product normally while its evaluations collect representative measurements. OpenMerit also records completed-work checkpoints; these checkpoints alone do not establish product quality.
|
|
37
|
+
|
|
38
|
+
In the unreleased source checkout, OpenMerit core checks baseline sufficiency from reported metric windows; when they pass, it asks Pi to discover candidates, run challenger trials, and calculate a comparison. Published `openmerit@0.1.5` still asks Pi to assess the baseline. You can request a check explicitly:
|
|
39
|
+
|
|
40
|
+
```text
|
|
41
|
+
/openmerit assess
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Review the candidate results, sample counts, costs, and unresolved measurements. A missing metric is a reason to gather evidence, not to assume a model passed or failed.
|
|
45
|
+
|
|
46
|
+
## Approve and verify a change
|
|
47
|
+
|
|
48
|
+
If a proposal passes the comparison checks and you want to try it, approve the specific proposal:
|
|
49
|
+
|
|
50
|
+
```text
|
|
51
|
+
/openmerit approve
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
The harness applies the change. Allow the configured post-swap evaluation window to complete and inspect its results. If a regression is reported, the workflow can request rollback when your policy permits it. Ask Pi to demonstrate that the previous configuration can be restored.
|
|
55
|
+
|
|
56
|
+
## Decide what to trust next
|
|
57
|
+
|
|
58
|
+
Keep the trial’s artifacts and audit trail. Before enabling automatic changes, confirm that your evaluations detect meaningful failures, provider costs are visible, and post-swap verification and rollback work for this task.
|
|
59
|
+
|
|
60
|
+
Pi does not wake itself while closed. Version 0.1.5 also has a [partial pause limitation](roadmap.md#known-limitations). Plan supervision around those behaviors.
|
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
# Getting started
|
|
2
|
+
|
|
3
|
+
## Prerequisites
|
|
4
|
+
|
|
5
|
+
- Node.js 22.19 or newer.
|
|
6
|
+
- npm.
|
|
7
|
+
- Pi 0.87-compatible provider authentication.
|
|
8
|
+
- A product repository in which Pi can create and verify evaluation and observability artifacts.
|
|
9
|
+
|
|
10
|
+
OpenMerit does not require or accept an LLM provider credential. Provider access belongs to the coding harness.
|
|
11
|
+
|
|
12
|
+
## Install from npm
|
|
13
|
+
|
|
14
|
+
```sh
|
|
15
|
+
pi install npm:openmerit
|
|
16
|
+
pi list
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
Run `/openmerit doctor` inside Pi to confirm that the extension is available.
|
|
20
|
+
|
|
21
|
+
## Install from source
|
|
22
|
+
|
|
23
|
+
```sh
|
|
24
|
+
git clone https://github.com/laz-aslam/openmerit.git
|
|
25
|
+
cd openmerit
|
|
26
|
+
npm ci
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
Start Pi with the extension:
|
|
30
|
+
|
|
31
|
+
```sh
|
|
32
|
+
./node_modules/.bin/pi --no-extensions -e ./packages/pi/src/index.ts
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
`--no-extensions` prevents a globally installed npm copy from conflicting with the source extension. Run `/openmerit doctor` inside Pi to confirm that the adapter and intent executor are available.
|
|
36
|
+
|
|
37
|
+
## Establish a product profile
|
|
38
|
+
|
|
39
|
+
Open the product repository in interactive Pi with the extension loaded. If the project has no OpenMerit state, the Pi adapter performs a bounded source preflight. In the unreleased source checkout it also scans after a settled Pi build turn, so an application created from an initially empty directory can enter setup without a slash command. A repository without such a target remains idle; use `/openmerit setup` only to recover or configure a custom integration explicitly.
|
|
40
|
+
|
|
41
|
+
OpenMerit sends Pi a private, typed `establish_evals` intent. Pi must identify the application LLM call or shared route under evaluation and its exact incumbent application model. It inspects the application's outputs, tools, and failure modes to infer task-relevant metrics instead of asking you to invent them. Before running setup, Pi shows the proposed metrics, why each matters, the representative case and repetition counts, aggregation, any actual performance threshold, and estimated evaluation cost. You confirm or edit that plan together with the stable target ID, call sites, route key, budget, check cadence, scheduling mode, and automation permissions. Pi's own model is never an application target.
|
|
42
|
+
|
|
43
|
+
For an active-session demo, Pi defaults the baseline and post-swap verification cadence to one completed evaluation task unless you change it. OpenMerit rejects empty baseline or post-swap cadences.
|
|
44
|
+
|
|
45
|
+
In this source checkout (not yet published in npm `openmerit@0.1.5`), Pi is also asked to research preliminary model leads during setup. Each reported lead has a current price estimate and reputation summary with sources and check times. OpenMerit stores these leads with the project profile. The field is optional for compatibility with existing protocol-0.1 harnesses, and an empty list means no credible leads were found. It does not change the application model or count as a measurement.
|
|
46
|
+
|
|
47
|
+
OpenMerit rejects successful setup without a confirmed task profile, at least one valid objective with an explicit sampling plan, a supervised-graduation policy, non-empty baseline and post-swap cadences, a budget, and evidence.
|
|
48
|
+
|
|
49
|
+
The source checkout's setup instruction asks Pi to run the application path with working provider authentication, execute each required collector and grader, inspect the resulting metric values, and retain verification artifacts. A metric that cannot be produced must be marked missing. This smoke run verifies the pipeline; it is not counted as a production baseline sample. OpenMerit still checks the reported coverage and evidence references, not the authenticity of a grader run itself.
|
|
50
|
+
|
|
51
|
+
## Collect the baseline
|
|
52
|
+
|
|
53
|
+
After setup, instrument the confirmed application target and use the product normally. Report application-task completions and cumulative metric windows through Pi's typed signal tool, including the application model where applicable, sample counts, and durable evidence. Pi's own model, tokens, cost, and latency are orchestration telemetry and do not satisfy product metrics. In the unreleased source checkout, OpenMerit core checks the stored baseline windows at the approved cadence without calling Pi; missing evidence stays insufficient. Published `openmerit@0.1.5` still nudges Pi to assess.
|
|
54
|
+
|
|
55
|
+
## Find the frontier
|
|
56
|
+
|
|
57
|
+
When the baseline meets every required metric threshold, OpenMerit issues the next approved nudges: discover candidates, run controlled challenger trials, and calculate the Pareto frontier. Pi performs that work. OpenMerit checks the returned result against the [Pareto specification](pareto-spec.md).
|
|
58
|
+
|
|
59
|
+
Candidate discovery uses current price, capability, availability, reputation, and benchmark signals to shortlist models worth testing. These are starting clues, not evidence that a model performs well on your application. In npm `openmerit@0.1.5`, the shortlist comes after the application's baseline has enough evidence. The unreleased source checkout collects preliminary leads at setup, then refreshes them, probes availability, and freezes the incumbent plus challengers. Pi must use the same cases, application revision, prompt, tools, grader, and run count, with a recorded shuffle seed. It preserves one experiment manifest and reports measured quality, cost, latency, tokens, and other required metrics. OpenMerit durably carries the accepted candidate set and assessments into later intents rather than relying on Pi conversation history. A model proposal must come from the verified frontier of those measured candidate results, not from price or reputation alone.
|
|
60
|
+
|
|
61
|
+
## Approve an early swap
|
|
62
|
+
|
|
63
|
+
With automatic swaps disabled, a model proposal that passes frontier validation waits for `/openmerit approve`. That authorizes Pi to apply only the proposed change. OpenMerit then requires post-swap evaluation; the approved policy determines rollback behavior. Automatic swaps may bypass per-proposal approval when enabled and the configured verified-swap threshold is reached, including when that threshold is zero.
|
|
64
|
+
|
|
65
|
+
Continue with [your first supervised trial](first-trial.md) for a practical walkthrough.
|
package/docs/index.md
ADDED
|
@@ -0,0 +1,40 @@
|
|
|
1
|
+
# OpenMerit documentation
|
|
2
|
+
|
|
3
|
+
OpenMerit helps your coding harness compare models against the work your product needs to do. You define what matters; the harness gathers measurements and tests candidates; OpenMerit coordinates the process and checks the returned results.
|
|
4
|
+
|
|
5
|
+
The current release is intended for supervised trials with Pi. Start with a representative task and working evaluations before relying on automatic model changes.
|
|
6
|
+
|
|
7
|
+
## Start with your product
|
|
8
|
+
|
|
9
|
+
Install the extension in Pi, then open your product repository:
|
|
10
|
+
|
|
11
|
+
```sh
|
|
12
|
+
pi install npm:openmerit
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
Interactive setup asks Pi to establish your task requirements, evaluation budget, automation permissions, and measurements. Follow the [getting started guide](getting-started.md), then work through a [first supervised trial](first-trial.md).
|
|
16
|
+
|
|
17
|
+
## Explore the demos
|
|
18
|
+
|
|
19
|
+
The [Demo section](demo.md) explains how each prepared example works, from a single ticket-classification call to a multi-model agent workflow. Watch the recorded comparisons and inspect their inputs, grading rules, model routes, and evidence.
|
|
20
|
+
|
|
21
|
+
## Understand the decisions
|
|
22
|
+
|
|
23
|
+
OpenMerit compares candidates against the objectives you confirm, such as quality, cost, and latency. Several models may offer useful tradeoffs. A cheaper model is not automatically a better fit.
|
|
24
|
+
|
|
25
|
+
- [Metrics and evidence](metrics-and-evidence.md) explains which measurements and samples a comparison needs.
|
|
26
|
+
- [Comparisons and swaps](lifecycle.md) follows a baseline through challenger trials, approval, and post-swap checks.
|
|
27
|
+
- [Budgets and permissions](budgets.md) describes the limits you give the harness and where enforcement depends on it.
|
|
28
|
+
- [Automatic checks](automation.md) explains when work becomes due and what needs an external scheduler.
|
|
29
|
+
|
|
30
|
+
## Operate with visibility
|
|
31
|
+
|
|
32
|
+
Use [commands](commands.md) to inspect status and request work. [Project files and logs](operations.md) explains what is stored, and [troubleshooting](troubleshooting.md) covers common setup and evidence problems.
|
|
33
|
+
|
|
34
|
+
OpenMerit checks structured results and model comparisons. It relies on the harness to execute the work and provide honest evidence. Read the [security notes](security.md) and [current limitations](roadmap.md) before enabling automation.
|
|
35
|
+
|
|
36
|
+
## Build an integration
|
|
37
|
+
|
|
38
|
+
Pi 0.87 is the only implemented adapter today. The core and protocol can support other harnesses through an adapter.
|
|
39
|
+
|
|
40
|
+
Start with the [architecture](architecture.md), then the [adapter guide](adapter-guide.md). The [Pareto specification](pareto-spec.md) defines the comparison rules, and [testing and validation](testing.md) explains the evidence behind implementation claims.
|