openmerit 0.1.4 → 0.1.6-preview.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +40 -0
- package/README.md +121 -386
- package/dist/core/src/index.d.ts +101 -0
- package/dist/core/src/index.js +1649 -0
- package/dist/core/src/store.d.ts +35 -0
- package/dist/core/src/store.js +102 -0
- package/dist/pi/src/index.d.ts +32 -0
- package/dist/pi/src/index.js +794 -0
- package/dist/pi/src/scheduler.d.ts +11 -0
- package/dist/pi/src/scheduler.js +137 -0
- package/dist/pi/src/wakeup.d.ts +2 -0
- package/dist/pi/src/wakeup.js +108 -0
- package/dist/protocol/src/index.d.ts +484 -0
- package/dist/protocol/src/index.js +47 -0
- package/dist/protocol/src/schemas.d.ts +576 -0
- package/dist/protocol/src/schemas.js +280 -0
- package/dist/terminal/public/app.js +297 -0
- package/dist/terminal/public/brands/anthropic.png +0 -0
- package/dist/terminal/public/brands/baai.png +0 -0
- package/dist/terminal/public/brands/baseten.png +0 -0
- package/dist/terminal/public/brands/cerebras.png +0 -0
- package/dist/terminal/public/brands/cohere.png +0 -0
- package/dist/terminal/public/brands/deepseek.ico +0 -0
- package/dist/terminal/public/brands/google.png +0 -0
- package/dist/terminal/public/brands/groq.ico +0 -0
- package/dist/terminal/public/brands/lm-studio.png +0 -0
- package/dist/terminal/public/brands/meta.ico +0 -0
- package/dist/terminal/public/brands/mistral.png +0 -0
- package/dist/terminal/public/brands/nomic.png +0 -0
- package/dist/terminal/public/brands/ollama.png +0 -0
- package/dist/terminal/public/brands/openai.png +0 -0
- package/dist/terminal/public/brands/openrouter.png +0 -0
- package/dist/terminal/public/brands/qwen.png +0 -0
- package/dist/terminal/public/brands/vllm.ico +0 -0
- package/dist/terminal/public/brands/vllm.png +0 -0
- package/dist/terminal/public/favicon.svg +1 -0
- package/dist/terminal/public/flow.css +1 -0
- package/dist/terminal/public/flow.js +770 -0
- package/dist/terminal/public/index.html +21 -0
- package/dist/terminal/public/styles.css +779 -0
- package/dist/terminal/src/activity-merge.mjs +64 -0
- package/dist/terminal/src/browser.mjs +29 -0
- package/dist/terminal/src/cli.mjs +60 -0
- package/dist/terminal/src/collect.mjs +311 -0
- package/dist/terminal/src/discovery.mjs +93 -0
- package/dist/terminal/src/hardware.mjs +57 -0
- package/dist/terminal/src/project-activity.mjs +156 -0
- package/dist/terminal/src/sample.mjs +171 -0
- package/dist/terminal/src/server.mjs +56 -0
- package/dist/terminal/src/services.mjs +62 -0
- package/dist/terminal/src/topology.mjs +30 -0
- package/docs/adapter-guide.md +189 -0
- package/docs/architecture.md +59 -0
- package/docs/automation.md +74 -0
- package/docs/budgets.md +37 -0
- package/docs/commands.md +85 -0
- package/docs/demo-backfill.md +29 -0
- package/docs/demo-fieldkit.md +47 -0
- package/docs/demo-placement.md +30 -0
- package/docs/demo-spam.md +15 -0
- package/docs/demo-support.md +42 -0
- package/docs/demo.md +57 -0
- package/docs/first-trial.md +60 -0
- package/docs/getting-started.md +65 -0
- package/docs/index.md +40 -0
- package/docs/inference-terminal.md +439 -0
- package/docs/lifecycle.md +30 -0
- package/docs/memo.md +126 -0
- package/docs/metrics-and-evidence.md +48 -0
- package/docs/operations.md +40 -0
- package/docs/pareto-spec.md +76 -0
- package/docs/pi-extension.md +54 -0
- package/docs/roadmap.md +28 -0
- package/docs/security.md +37 -0
- package/docs/site-artwork-linocut.md +23 -0
- package/docs/site-artwork-miniature-diverse.md +28 -0
- package/docs/site-artwork-miniature.md +26 -0
- package/docs/site-demo.md +177 -0
- package/docs/site-design.md +94 -0
- package/docs/site-documentation.md +83 -0
- package/docs/site-dynamic-og.md +35 -0
- package/docs/site-faq-maintenance.md +115 -0
- package/docs/site-hero-resolution.md +60 -0
- package/docs/site-illustration-sequences.md +227 -0
- package/docs/site-inference-terminal.md +203 -0
- package/docs/site-memo.md +39 -0
- package/docs/site-og-image.md +38 -0
- package/docs/site-og-workshop.md +21 -0
- package/docs/site-section-artwork.md +56 -0
- package/docs/site-skill-review.md +57 -0
- package/docs/site-terminal-preview.md +85 -0
- package/docs/testing.md +118 -0
- package/docs/troubleshooting.md +55 -0
- package/docs/ux-reference.md +32 -0
- package/package.json +74 -42
- package/benchmark/invoice_ocr/data/invoice_01_ground_truth.json +0 -38
- package/benchmark/invoice_ocr/data/invoice_01_row_2.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_02_ground_truth.json +0 -32
- package/benchmark/invoice_ocr/data/invoice_02_row_5.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_03_ground_truth.json +0 -26
- package/benchmark/invoice_ocr/data/invoice_03_row_6.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_04_ground_truth.json +0 -26
- package/benchmark/invoice_ocr/data/invoice_04_row_7.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_05_ground_truth.json +0 -38
- package/benchmark/invoice_ocr/data/invoice_05_row_947.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_06_ground_truth.json +0 -38
- package/benchmark/invoice_ocr/data/invoice_06_row_948.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_07_ground_truth.json +0 -20
- package/benchmark/invoice_ocr/data/invoice_07_row_949.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_08_ground_truth.json +0 -38
- package/benchmark/invoice_ocr/data/invoice_08_row_1888.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_09_ground_truth.json +0 -26
- package/benchmark/invoice_ocr/data/invoice_09_row_1890.jpg +0 -0
- package/benchmark/invoice_ocr/data/invoice_10_ground_truth.json +0 -20
- package/benchmark/invoice_ocr/data/invoice_10_row_1892.jpg +0 -0
- package/benchmark/invoice_ocr/data/manifest.json +0 -97
- package/dist/benchmarks.js +0 -98
- package/dist/catalog.js +0 -61
- package/dist/cli.js +0 -188
- package/dist/daemon.js +0 -407
- package/dist/diagnostics.js +0 -227
- package/dist/frontier.js +0 -56
- package/dist/harness.js +0 -1
- package/dist/integrations.js +0 -19
- package/dist/invoice-eval.js +0 -33
- package/dist/invoice-score.js +0 -124
- package/dist/judge.js +0 -43
- package/dist/llm.js +0 -207
- package/dist/pi-config.js +0 -46
- package/dist/pi-trials.js +0 -373
- package/dist/policy.js +0 -185
- package/dist/providers.js +0 -1
- package/dist/recommend.js +0 -76
- package/dist/routes.js +0 -74
- package/dist/standalone.js +0 -224
- package/dist/store.js +0 -89
- package/dist/strategist.js +0 -68
- package/dist/task-input.js +0 -54
- package/dist/traces.js +0 -127
- package/dist/trials.js +0 -140
- package/dist/types.js +0 -2
- package/examples/invoice-prompt.txt +0 -19
- package/examples/task.example.json +0 -7
- package/extension/openmerit.ts +0 -947
- package/instructions/OPENMERIT.md +0 -63
- package/instructions/openmerit.policy.json +0 -37
- package/rules.md +0 -43
|
@@ -0,0 +1,189 @@
|
|
|
1
|
+
# Building a harness adapter
|
|
2
|
+
|
|
3
|
+
OpenMerit adapters are transport bridges. They do not reproduce OpenMerit's lifecycle, validation, persistence, or Pareto rules, and OpenMerit does not reproduce the harness's intelligence.
|
|
4
|
+
|
|
5
|
+
## Required components
|
|
6
|
+
|
|
7
|
+
An adapter provides a `HarnessAdapter` with:
|
|
8
|
+
|
|
9
|
+
- a stable harness descriptor;
|
|
10
|
+
- the OpenMerit protocol versions it supports;
|
|
11
|
+
- declared capabilities;
|
|
12
|
+
- supported execution modes;
|
|
13
|
+
- a `dispatch` method that accepts a canonical `OpenMeritIntent` and returns an acceptance receipt.
|
|
14
|
+
|
|
15
|
+
```ts
|
|
16
|
+
import {
|
|
17
|
+
OpenMeritCoordinator,
|
|
18
|
+
ProjectStore,
|
|
19
|
+
type HarnessAdapter,
|
|
20
|
+
} from "openmerit/core";
|
|
21
|
+
import { PROTOCOL_VERSION, OPENMERIT_SCHEMA_DIALECT } from "openmerit/protocol";
|
|
22
|
+
|
|
23
|
+
const adapter: HarnessAdapter = {
|
|
24
|
+
descriptor: {
|
|
25
|
+
id: "my-harness",
|
|
26
|
+
name: "My coding harness",
|
|
27
|
+
version: "1.0.0",
|
|
28
|
+
protocolVersions: [PROTOCOL_VERSION],
|
|
29
|
+
capabilities: [
|
|
30
|
+
"user_confirmation",
|
|
31
|
+
"evaluation_authoring",
|
|
32
|
+
"observability_instrumentation",
|
|
33
|
+
"production_observation",
|
|
34
|
+
"candidate_discovery",
|
|
35
|
+
"challenger_execution",
|
|
36
|
+
"frontier_calculation",
|
|
37
|
+
"model_mutation",
|
|
38
|
+
"post_swap_verification",
|
|
39
|
+
"rollback",
|
|
40
|
+
],
|
|
41
|
+
executionModes: ["interactive"],
|
|
42
|
+
structuredOutput: {
|
|
43
|
+
schemaDialect: OPENMERIT_SCHEMA_DIALECT,
|
|
44
|
+
presentation: "adapter_validation",
|
|
45
|
+
enforcement: "adapter",
|
|
46
|
+
},
|
|
47
|
+
automation: {
|
|
48
|
+
supportedSignals: ["product_task_completed", "metric_window_available", "scheduled_tick", "model_catalog_changed", "verification_window_completed"],
|
|
49
|
+
persistentScheduling: false,
|
|
50
|
+
backgroundExecution: false,
|
|
51
|
+
wakeupProvisioning: "none",
|
|
52
|
+
},
|
|
53
|
+
},
|
|
54
|
+
async dispatch(intent) {
|
|
55
|
+
await myHarness.enqueue(intent);
|
|
56
|
+
return { accepted: true, harnessJobId: intent.id };
|
|
57
|
+
},
|
|
58
|
+
};
|
|
59
|
+
|
|
60
|
+
const coordinator = new OpenMeritCoordinator(
|
|
61
|
+
new ProjectStore(process.cwd()),
|
|
62
|
+
adapter,
|
|
63
|
+
);
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
`ProjectStore` is the included local-filesystem backend. A remote or hosted integration can implement the `OpenMeritStore` interface instead.
|
|
67
|
+
|
|
68
|
+
The confirmed task profile identifies the application LLM call under evaluation. Every metric, assessment, candidate, frontier artifact, and automation signal must carry that target ID; the harness's own model usage is orchestration data and must not be reported as product evidence. Application-provider cost, token, and latency telemetry belongs on the application signal or metric record.
|
|
69
|
+
|
|
70
|
+
In the unreleased source checkout, setup may return `preliminaryModelLeads`: model IDs with sourced price estimates and reputation summaries. The field is optional under protocol 0.1 for older adapters. Core stores the leads and includes both the confirmed task profile and these leads in later discovery intent constraints, so a headless harness run does not depend on prior chat context. Leads are research priors, never measured candidate assessments.
|
|
71
|
+
|
|
72
|
+
## External telemetry export
|
|
73
|
+
|
|
74
|
+
OpenMerit always writes its canonical, redacted audit stream locally. A host integration may pass `eventExporters` to `ProjectStore` to mirror those same safe events to OpenTelemetry, Langfuse, LangSmith, or another observability backend:
|
|
75
|
+
|
|
76
|
+
```ts
|
|
77
|
+
const store = new ProjectStore(projectRoot, {
|
|
78
|
+
eventExporters: [{
|
|
79
|
+
id: "my-otel-bridge",
|
|
80
|
+
async export(event) {
|
|
81
|
+
await telemetry.emit("openmerit.audit", event);
|
|
82
|
+
},
|
|
83
|
+
}],
|
|
84
|
+
});
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
Exporters are not the durable source of truth. They are best effort, and failures are recorded locally without blocking OpenMerit decisions.
|
|
88
|
+
|
|
89
|
+
## Dispatching work
|
|
90
|
+
|
|
91
|
+
```ts
|
|
92
|
+
const issued = await coordinator.issue("establish_evals");
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
The coordinator builds the canonical intent, checks the descriptor's protocol and capabilities, creates durable state, records audit events, and calls the adapter. The harness decides how to execute the job.
|
|
96
|
+
|
|
97
|
+
If a required capability is absent, dispatch is rejected before harness work starts.
|
|
98
|
+
|
|
99
|
+
## Reporting progress
|
|
100
|
+
|
|
101
|
+
A harness may send durable lifecycle progress through `recordLifecycle`:
|
|
102
|
+
|
|
103
|
+
```ts
|
|
104
|
+
await coordinator.recordLifecycle({
|
|
105
|
+
protocolVersion: PROTOCOL_VERSION,
|
|
106
|
+
intentId: issued.intent.id,
|
|
107
|
+
status: "running",
|
|
108
|
+
occurredAt: new Date().toISOString(),
|
|
109
|
+
message: "Generating representative evaluation cases",
|
|
110
|
+
});
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
Use `waiting_for_user` when an interactive confirmation is required but cannot be completed in the current execution mode.
|
|
114
|
+
|
|
115
|
+
## Completing work
|
|
116
|
+
|
|
117
|
+
Each intent kind has a typed output in `IntentOutputMap`. Return an `IntentResult<K>` with durable evidence references, then submit it to the coordinator:
|
|
118
|
+
|
|
119
|
+
```ts
|
|
120
|
+
await coordinator.acceptResult(result);
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
The coordinator validates the active intent, protocol version, evidence references, and intent-specific output, including applicable frontier and candidate-identity checks, before updating lifecycle state. Adapters must not duplicate this logic. The harness remains responsible for performing only the authorized work; result validation does not sandbox its tools or authenticate external evidence.
|
|
124
|
+
|
|
125
|
+
### Present the canonical result schema
|
|
126
|
+
|
|
127
|
+
Import `completionToolSchemaFor(intent.kind)` from `openmerit/protocol` and present it through the harness's strongest structured-output mechanism. Do not recreate result schemas inside an adapter. A harness that supports changing tool definitions should specialize the completion tool whenever an intent becomes active; other harnesses may expose intent-specific tools or perform adapter-side validation.
|
|
128
|
+
|
|
129
|
+
Declare the behavior in `HarnessDescriptor.structuredOutput`:
|
|
130
|
+
|
|
131
|
+
- `schemaDialect` identifies the canonical OpenMerit JSON Schema dialect;
|
|
132
|
+
- `presentation` states whether the adapter uses a dynamic tool, a static union, or adapter validation;
|
|
133
|
+
- `enforcement` states whether validation occurs at the provider and harness, harness only, or adapter only.
|
|
134
|
+
|
|
135
|
+
Schema validation catches malformed structure. The coordinator still owns semantic checks such as authorized candidate identity, metric readiness for frontier comparisons, lifecycle consistency, and Pareto conformance.
|
|
136
|
+
|
|
137
|
+
## Reporting normal work and catalogue changes
|
|
138
|
+
|
|
139
|
+
After a real product task completes:
|
|
140
|
+
|
|
141
|
+
```ts
|
|
142
|
+
await coordinator.recordTaskObservation({
|
|
143
|
+
modelId: "vendor/model",
|
|
144
|
+
});
|
|
145
|
+
|
|
146
|
+
// In the unreleased source checkout, core has already checked baseline
|
|
147
|
+
// sufficiency when the confirmed cadence became due. Inspect durable state
|
|
148
|
+
// and dispatch only a pending downstream intent, if any.
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
`recordTaskObservation` resolves the confirmed application target from project state. New adapters should prefer `recordAutomationSignal` so the task signal can carry the explicit `targetId`, durable evidence, and application-provider telemetry.
|
|
152
|
+
|
|
153
|
+
When the harness's available model catalogue changes, compute a stable fingerprint and call `recordModelCatalogFingerprint`. The returned `reassessmentDue` flag tells the adapter whether the lifecycle justifies candidate discovery.
|
|
154
|
+
|
|
155
|
+
New adapters should use `recordAutomationSignal` for product tasks, metric windows, scheduled ticks, catalogue observations, and verification windows. Signal IDs must remain stable across delivery retries. If the returned decision issues an intent, dispatch it through the coordinator. `recordTaskObservation` and `recordModelCatalogFingerprint` remain compatibility helpers.
|
|
156
|
+
|
|
157
|
+
Call `reconcileAutomation` after configuration changes or when restoring a project. It returns an `AutomationPlan`. Implement `scheduleWakeup` only when the harness can genuinely arrange the requested wakeup. Declare persistent scheduling, background execution, supported signal sources, and provisioning requirements accurately in `HarnessDescriptor.automation`.
|
|
158
|
+
|
|
159
|
+
If persistent scheduling is unavailable, expose the plan's active-session or external-scheduler fallback. Do not report a seven-day schedule as configured merely because the adapter checks policy when an interactive session starts.
|
|
160
|
+
|
|
161
|
+
## Capability mapping
|
|
162
|
+
|
|
163
|
+
| Intent | Required capability |
|
|
164
|
+
| --- | --- |
|
|
165
|
+
| `establish_evals` | confirmation, evaluation authoring, observability instrumentation |
|
|
166
|
+
| `instrument_observability` | observability instrumentation |
|
|
167
|
+
| `run_assessment` | production observation |
|
|
168
|
+
| `discover_candidates` | candidate discovery |
|
|
169
|
+
| `run_challenger_trials` | challenger execution |
|
|
170
|
+
| `calculate_frontier` | frontier calculation |
|
|
171
|
+
| `investigate_regression` | production observation |
|
|
172
|
+
| `apply_model_swap` | model mutation |
|
|
173
|
+
| `verify_model_swap` | post-swap verification |
|
|
174
|
+
| `rollback_model_swap` | rollback |
|
|
175
|
+
|
|
176
|
+
## Conformance checklist
|
|
177
|
+
|
|
178
|
+
An adapter is ready when it:
|
|
179
|
+
|
|
180
|
+
1. advertises only capabilities it can actually execute;
|
|
181
|
+
2. preserves intent IDs and protocol versions;
|
|
182
|
+
3. never reports success without durable evidence;
|
|
183
|
+
4. returns the typed output for every supported intent;
|
|
184
|
+
5. performs model mutation only when the intent explicitly authorizes it;
|
|
185
|
+
6. reports ordinary product tasks without counting OpenMerit orchestration turns;
|
|
186
|
+
7. passes a durable issue-and-complete cycle using `OpenMeritCoordinator`;
|
|
187
|
+
8. proves the requested side effect, not merely a generated description.
|
|
188
|
+
|
|
189
|
+
The fake non-Pi adapter in `packages/core/test/adapter-conformance.test.ts` is the executable reference.
|
|
@@ -0,0 +1,59 @@
|
|
|
1
|
+
# Architecture
|
|
2
|
+
|
|
3
|
+
## Design principle
|
|
4
|
+
|
|
5
|
+
OpenMerit is an orchestrator, not an intelligence layer. The product application LLM call is the system under test; the coding harness is the evaluator and execution environment.
|
|
6
|
+
|
|
7
|
+
```text
|
|
8
|
+
User or lifecycle event
|
|
9
|
+
|
|
|
10
|
+
v
|
|
11
|
+
OpenMerit deterministic orchestration
|
|
12
|
+
|
|
|
13
|
+
| typed intent: outcome, constraints, evidence
|
|
14
|
+
v
|
|
15
|
+
Coding harness such as Pi
|
|
16
|
+
|
|
|
17
|
+
| intelligent work, tools, evals, calculation
|
|
18
|
+
v
|
|
19
|
+
Typed result and durable evidence references
|
|
20
|
+
|
|
|
21
|
+
v
|
|
22
|
+
OpenMerit validation, policy, persistence, next nudge
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
An automatic OpenMerit action means automatic issuance of an approved intent. OpenMerit does not perform candidate research, evaluate quality, calculate a frontier, or mutate model configuration.
|
|
26
|
+
|
|
27
|
+
## Packages
|
|
28
|
+
|
|
29
|
+
### `openmerit/protocol`
|
|
30
|
+
|
|
31
|
+
Defines the JSON-safe boundary: durable intents, evidence references, metrics, task profiles, candidate assessments, frontier snapshots, budgets, policies, state, and audit events.
|
|
32
|
+
|
|
33
|
+
It also owns the versioned runtime schema registry for all intent results. Harness adapters consume these schemas rather than defining their own result shapes, report their structured-output capabilities, and may present the active schema using harness-native tools or adapter validation. A confirmed application LLM target is carried through profiles, metrics, candidates, frontier results, automation signals, and swaps; the harness model is not a valid substitute target.
|
|
34
|
+
|
|
35
|
+
### `openmerit/core`
|
|
36
|
+
|
|
37
|
+
Contains deterministic behavior: choosing the next lifecycle intent according to confirmation and rollback policy, persisting state, and checking harness-produced frontiers against the normative definition. The verifier rejects invalid harness results; it is not a model-selection engine.
|
|
38
|
+
|
|
39
|
+
`OpenMeritCoordinator` is the reusable adapter boundary. It constructs canonical intents, negotiates capabilities, records lifecycle progress and observations, validates results, and advances durable state. `OpenMeritStore` makes persistence replaceable; `ProjectStore` is the bundled filesystem backend.
|
|
40
|
+
|
|
41
|
+
The coordinator also accepts idempotent application-target automation signals and evaluates the user-confirmed `AutomationCheckPolicy`. `buildAutomationPlan` tells an adapter which signals and wakeups are required and reports unsupported persistent scheduling honestly. OpenMerit never installs a scheduler or creates a paid cloud worker itself.
|
|
42
|
+
|
|
43
|
+
### `openmerit/pi`
|
|
44
|
+
|
|
45
|
+
Exposes the command surface, supplies private typed intent context, triggers Pi, accepts one structured completion for the active intent, records evidence, advances the lifecycle, and detects model-catalog changes.
|
|
46
|
+
|
|
47
|
+
The Pi package implements `HarnessAdapter`. It does not define canonical intents, per-intent result contracts, stage transitions, or validation rules.
|
|
48
|
+
|
|
49
|
+
### Inference Terminal (unreleased)
|
|
50
|
+
|
|
51
|
+
`packages/terminal` supplies a standalone `openmerit dash` executable, passive collectors, a loopback read-only HTTP API, and bundled dark-only browser assets. It does not invoke the coordinator or harness. One collector cycle reads existing metadata and runtime inventory; browser clients receive the same cached snapshot. Imported activity, runtime counters, and sampled whole-machine resources retain their separate scopes. Missing measurements remain unknown, and fictional sample data is isolated from the real snapshot. See the [terminal guide](inference-terminal.md) for the supported sources and JSON contract.
|
|
52
|
+
|
|
53
|
+
## Intelligence resources
|
|
54
|
+
|
|
55
|
+
Models, web research, tools, evaluators, and fast decision systems such as Jev System One are harness resources. They are configured and invoked by Pi. This keeps OpenMerit provider-neutral and lets future adapters choose their own implementation.
|
|
56
|
+
|
|
57
|
+
## Trust boundary
|
|
58
|
+
|
|
59
|
+
The harness is trusted to perform the work and return honest evidence. OpenMerit reduces unsupported claims through structured outputs, evidence references, compatible profile revisions, complete comparisons, and conservative uncertainty handling. It cannot prove that an arbitrary external evidence URI contains truthful data.
|
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
# Harness-neutral automation
|
|
2
|
+
|
|
3
|
+
The user defines cadence, thresholds, budget, scheduling mode, and automation permission. OpenMerit stores that policy, evaluates durable signals, and decides whether a bounded intent is due. A harness adapter arranges native wakeups and reports events. The coding harness performs the resulting intelligent work.
|
|
4
|
+
|
|
5
|
+
```text
|
|
6
|
+
User-confirmed policy
|
|
7
|
+
↓
|
|
8
|
+
OpenMerit due-check engine
|
|
9
|
+
↓
|
|
10
|
+
Harness-native wakeup or lifecycle hook
|
|
11
|
+
↓
|
|
12
|
+
Standard automation signal
|
|
13
|
+
↓
|
|
14
|
+
No work due, or one typed OpenMerit intent
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
## Check policy
|
|
18
|
+
|
|
19
|
+
`AutomationCheckPolicy` independently configures:
|
|
20
|
+
|
|
21
|
+
- baseline assessment cadence;
|
|
22
|
+
- frontier reassessment cadence and catalogue-change behavior;
|
|
23
|
+
- regression-check cadence and per-metric relative deterioration thresholds;
|
|
24
|
+
- post-swap verification windows;
|
|
25
|
+
- a cooldown between mature-system checks;
|
|
26
|
+
- active-session or persistent execution;
|
|
27
|
+
- whether wakeup provisioning may make infrastructure changes.
|
|
28
|
+
|
|
29
|
+
Task-count and elapsed-time triggers use OR semantics. A policy with `afterCompletedTasks: 100` and `afterElapsedSeconds: 604800` becomes due at the first threshold reached. A scheduled wakeup does not directly launch expensive work: it submits `scheduled_tick`, and OpenMerit evaluates policy and durable state before issuing an intent.
|
|
30
|
+
|
|
31
|
+
In the unreleased source checkout, the baseline elapsed-time clock starts when setup or required instrumentation finishes. Ordinary task and metric signals do not restart that clock. OpenMerit retains the latest reported aggregate window for each application metric, including sample count and evidence references. It never sums possibly overlapping windows. At a due baseline check, core requires ready instrumentation and a finite measured value, enough samples, and an evidence reference for every required objective. An insufficient check records the gaps and waits for more data; sufficient evidence advances to candidate discovery.
|
|
32
|
+
|
|
33
|
+
## Signals
|
|
34
|
+
|
|
35
|
+
Adapters may report:
|
|
36
|
+
|
|
37
|
+
- `product_task_completed`;
|
|
38
|
+
- `metric_window_available`;
|
|
39
|
+
- `scheduled_tick`;
|
|
40
|
+
- `model_catalog_changed`;
|
|
41
|
+
- `verification_window_completed`.
|
|
42
|
+
|
|
43
|
+
Signal IDs are idempotency keys. OpenMerit persists a bounded history, does not count duplicate task completions twice, and keeps due work pending until an adapter successfully dispatches it.
|
|
44
|
+
|
|
45
|
+
## Capability negotiation
|
|
46
|
+
|
|
47
|
+
Every `HarnessDescriptor` declares supported signals, persistent scheduling, background execution, and whether wakeup provisioning is native, unavailable, or requires an infrastructure change. OpenMerit produces an `AutomationPlan` containing required signals, the next wakeup, scheduling strategy, and explicit gaps.
|
|
48
|
+
|
|
49
|
+
If persistent scheduling is requested but unsupported, the plan says `external_scheduler_required`. The adapter must not claim that a schedule exists. Active-session adapters can submit `scheduled_tick` when they start. An external scheduler may also wake a non-interactive harness and submit the same signal.
|
|
50
|
+
|
|
51
|
+
OpenMerit calls `scheduleWakeup` only when the confirmed policy requests persistent execution, the adapter declares persistent scheduling and background execution, and any infrastructure-changing provisioning has been authorized. Paid agents, system scheduler installation, and cloud infrastructure remain harness implementation details and authorization boundaries.
|
|
52
|
+
|
|
53
|
+
## Intent mapping
|
|
54
|
+
|
|
55
|
+
OpenMerit owns deterministic mapping from due conditions to intents:
|
|
56
|
+
|
|
57
|
+
| Due condition | Intent |
|
|
58
|
+
| --- | --- |
|
|
59
|
+
| Baseline cadence | Core checks stored baseline-window sufficiency; if ready, `discover_candidates` |
|
|
60
|
+
| Mature frontier cadence or changed catalogue | `discover_candidates` |
|
|
61
|
+
| Regression cadence or crossed metric threshold | `investigate_regression` |
|
|
62
|
+
| Completed post-swap window | `verify_model_swap` |
|
|
63
|
+
|
|
64
|
+
Subsequent lifecycle transitions continue through candidate trials, frontier calculation, supervised swap, verification, or rollback. The harness executes those intents; OpenMerit validates their structured result and evidence.
|
|
65
|
+
|
|
66
|
+
## Pi behavior
|
|
67
|
+
|
|
68
|
+
Pi reports catalogue and scheduled-tick signals on session start. Application instrumentation submits product-task, metric-window, and post-swap verification signals through its typed `openmerit_report_signal` tool; Pi's own settled-turn telemetry is not treated as product evidence. Published `openmerit@0.1.5` cannot provision persistent background wakeups and reports an external-scheduler gap.
|
|
69
|
+
|
|
70
|
+
In the unreleased source checkout on macOS, a confirmed policy with `execution.mode: "persistent"` and `infrastructureChangesAllowed: true` allows the Pi adapter to install a per-project user LaunchAgent. It uses the model selected in the interactive Pi session and checks that Pi authentication works without shell-only credentials. The agent checks local state every five minutes without calling a model. When a baseline check is due, the runner submits `scheduled_tick` directly to core: an insufficient result exits without starting Pi; a sufficient result starts Pi for candidate discovery. The Mac must be logged in; this is not a cloud service, an exact-time alarm, or a source of product metrics. Task-count and metric-window checks still depend on application signals.
|
|
71
|
+
|
|
72
|
+
This background path is experimental in the source checkout. A local runner test verifies that an insufficient baseline check needs no Pi process. An earlier live OpenRouter/Pi test reached the former assessment intent and made tool calls, but did not complete it within the bounded window. Candidate discovery, comparison, and swapping from a background Pi run are not yet product-validated. A background Pi run is capped at 15 minutes and writes its exit code and any still-active intent ID to `.openmerit/scheduler-status.json`; an unfinished intent requires inspection and explicit recovery in Pi.
|
|
73
|
+
|
|
74
|
+
`/openmerit pause` in this source checkout persists the pause, unloads the project's LaunchAgent, and prevents automatic due decisions while still recording incoming signals. `/openmerit resume` clears the pause and restores an authorized schedule. Neither command cancels already-active work. A declined or failed scheduler installation is reported as a gap, not success. On other platforms, persistent policies still require an external scheduler.
|
package/docs/budgets.md
ADDED
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
# Budgets and permissions
|
|
2
|
+
|
|
3
|
+
Setup asks you to confirm the limits and permissions for evaluation work. OpenMerit stores the policy and passes its constraints to the harness.
|
|
4
|
+
|
|
5
|
+
## Evaluation costs
|
|
6
|
+
|
|
7
|
+
OpenMerit has no subscription fee. Model-provider charges still apply to setup, assessments, challenger trials, and verification. Paid tools, telemetry, or infrastructure used by your harness may add costs. Application-provider cost for ordinary product tasks can be reported separately as production evidence; Pi's own orchestration usage is not application cost and is not charged to the challenger-evaluation budget.
|
|
8
|
+
|
|
9
|
+
## Budget fields
|
|
10
|
+
|
|
11
|
+
| Field | Meaning |
|
|
12
|
+
| --- | --- |
|
|
13
|
+
| `currency` | Currency for the evaluation budget. |
|
|
14
|
+
| `maximumSpend` | Spending limit the harness is instructed to respect. |
|
|
15
|
+
| `maximumCandidateCount` | Optional limit on candidate count. |
|
|
16
|
+
| `maximumRunsPerCandidate` | Optional limit on runs per candidate. |
|
|
17
|
+
| `expiresAt` | Optional expiration for the evaluation budget. |
|
|
18
|
+
|
|
19
|
+
In the unreleased development version, OpenMerit keeps a separate ledger for controlled challenger trials. Each challenger intent receives the amount already spent and the remaining allowance. Challenger results that exceed the approved limit are rejected, and another challenger intent cannot start after the limit is exhausted. Provider-side spending controls remain the strongest protection against a request that exceeds its estimate while it is running.
|
|
20
|
+
|
|
21
|
+
Setup, baseline product work, assessments, and swap verification remain visible costs, but do not consume `evaluationBudget.maximumSpend`. This separation prevents normal product usage from silently exhausting the allowance reserved for comparing challengers.
|
|
22
|
+
|
|
23
|
+
## Model changes
|
|
24
|
+
|
|
25
|
+
OpenMerit normally waits for `/openmerit approve` before issuing a swap request. It can request an automatic swap when `automaticSwapsEnabled` is true and `confirmationRequiredUntilVerifiedSwaps` has been reached. That threshold may be zero.
|
|
26
|
+
|
|
27
|
+
Post-swap evaluation remains part of the workflow. The harness performs the change and reports whether it passed evaluation. `rollbackOnRegression` determines whether a reported post-swap regression can trigger a rollback request.
|
|
28
|
+
|
|
29
|
+
## Infrastructure changes
|
|
30
|
+
|
|
31
|
+
Persistent scheduling is a separate permission. An adapter that provisions infrastructure must declare that capability, and the confirmed check policy must allow infrastructure changes before OpenMerit asks it to provision a wakeup.
|
|
32
|
+
|
|
33
|
+
The current Pi adapter does not provide persistent scheduling or background execution. An external scheduler is needed to start Pi for unattended checks. See [automatic checks](automation.md).
|
|
34
|
+
|
|
35
|
+
## What a policy can establish
|
|
36
|
+
|
|
37
|
+
OpenMerit uses policy and lifecycle state to decide which requests to issue, then checks the results returned by the harness. It is not a sandbox around the harness’s tools or provider access. See [security](security.md) for the responsibility boundary.
|
package/docs/commands.md
ADDED
|
@@ -0,0 +1,85 @@
|
|
|
1
|
+
# Commands
|
|
2
|
+
|
|
3
|
+
Run these slash commands inside Pi with the OpenMerit extension loaded. Shell commands such as `pi install` run in your terminal instead.
|
|
4
|
+
|
|
5
|
+
## Inference Terminal (preview)
|
|
6
|
+
|
|
7
|
+
The `openmerit@preview` npm package adds a shell command, `openmerit dash [path]`, for a passive local view of models, connections, usage, and compute. From this checkout, use `npm run dash -- [path]`. It is independent of Pi setup and does not run model comparisons or inference. Published `openmerit@0.1.5` does not include it.
|
|
8
|
+
|
|
9
|
+
See the [Inference Terminal guide](inference-terminal.md) for local installation, automatic discovery, `--no-discovery`, `--json`, `--demo`, collection intervals, and data limitations. Discovery is part of `dash` and is scoped to application evidence. Coding-harness histories are not collected.
|
|
10
|
+
|
|
11
|
+
## Inspect status
|
|
12
|
+
|
|
13
|
+
```text
|
|
14
|
+
/openmerit
|
|
15
|
+
/openmerit status
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
Both show the task profile, metric readiness, evaluation status, active work, next assessment, lifecycle stage, incumbent candidate, candidate and assessment counts, challenger spend, proposed candidate, and pause flag. A readiness label is based on the harness’s structured reports; inspect the evidence for the underlying measurements.
|
|
19
|
+
|
|
20
|
+
## Start or recover setup
|
|
21
|
+
|
|
22
|
+
```text
|
|
23
|
+
/openmerit setup
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
Requests setup explicitly or reopens setup to correct stale target/incumbent state. Accepted replacement setup clears prior candidate, assessment, frontier, and baseline handoffs while retaining the audit log.
|
|
27
|
+
|
|
28
|
+
Request setup explicitly only when recovering, rerunning, or configuring a custom integration. In the unreleased source checkout, an unconfigured project starts setup when bounded source detection finds an application LLM call at session start or after a settled Pi build turn. OpenMerit refuses another intent while work is active.
|
|
29
|
+
|
|
30
|
+
## Cancel interrupted work
|
|
31
|
+
|
|
32
|
+
```text
|
|
33
|
+
/openmerit cancel
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
Record the active intent as cancelled after an interrupted or abandoned harness run. This clears durable active-work state without claiming success and releases any interactive request that was queued behind setup.
|
|
37
|
+
|
|
38
|
+
## Request an assessment
|
|
39
|
+
|
|
40
|
+
```text
|
|
41
|
+
/openmerit assess
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
In the unreleased source checkout, ask OpenMerit core to check stored baseline metric windows now. It reports missing or immature required metrics without starting Pi, or starts candidate discovery when they pass. Published `openmerit@0.1.5` still asks Pi to assess. Neither path creates missing measurements or overrides sample requirements.
|
|
45
|
+
|
|
46
|
+
## Request a frontier
|
|
47
|
+
|
|
48
|
+
```text
|
|
49
|
+
/openmerit frontier
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
Ask the harness to calculate a model comparison now. OpenMerit validates the returned frontier against the confirmed profile and supplied assessments. The command does not itself authorize a model change.
|
|
53
|
+
|
|
54
|
+
## Approve a proposal
|
|
55
|
+
|
|
56
|
+
```text
|
|
57
|
+
/openmerit approve
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
Approve the current proposal after its frontier result has passed validation. If there is no proposal in the expected lifecycle stage, the command reports that no verified proposal is awaiting approval. See [comparisons and swaps](lifecycle.md) for automatic-swap policy.
|
|
61
|
+
|
|
62
|
+
## Pause and resume
|
|
63
|
+
|
|
64
|
+
```text
|
|
65
|
+
/openmerit pause
|
|
66
|
+
/openmerit resume
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
In 0.1.5, pause suppresses some lifecycle follow-ups, but task, measurement, and session-start signals can still start due checks. The flag resets on process restart and does not cancel active work. Explicit commands still work. Do not treat “paused” as a guarantee that all OpenMerit work has stopped.
|
|
70
|
+
|
|
71
|
+
## Locate logs and evidence
|
|
72
|
+
|
|
73
|
+
```text
|
|
74
|
+
/openmerit logs
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
Show the local audit, exporter-error, evidence-manifest, and state paths. See [project files and logs](operations.md) for their contents.
|
|
78
|
+
|
|
79
|
+
## Check the installation
|
|
80
|
+
|
|
81
|
+
```text
|
|
82
|
+
/openmerit doctor
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
Show adapter diagnostics, UI availability, and scheduling support. This checks extension wiring; it does not prove provider authentication, successful evaluations, or a completed model swap.
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
# Backfill recovery
|
|
2
|
+
|
|
3
|
+
The [agent workflow](https://openmerit.site/demo?task=backfill) prepares a recovery manifest for an interrupted data backfill. Its goal is to resume recoverable partitions without duplicating committed data, quarantine malformed input, and stop for investigation when the failure is unknown or involves a checksum conflict. It never executes the manifest or writes to a database.
|
|
4
|
+
|
|
5
|
+
## Workflow structure
|
|
6
|
+
|
|
7
|
+
Three dependent model turns perform one goal: request the incident through `inspect_job`, request its authoritative ledger through `read_checkpoint`, then combine those results with the recovery policy into a manifest. The tools are read-only functions over fixed synthetic records. Their requests are JSON commands dispatched by the prepared adapter, not arbitrary code. The runner fixes the three-task sequence; the model supplies the job identifiers and final plan. Invalid requests become scored failures and are not silently repaired.
|
|
8
|
+
|
|
9
|
+
The fixed agent uses GPT-4.1 mini for all three tasks. Its challenger is one prepared mixed-model route:
|
|
10
|
+
|
|
11
|
+
| Task | Without OpenMerit | With the proposed route |
|
|
12
|
+
| --- | --- | --- |
|
|
13
|
+
| Inspect incident | GPT-4.1 mini | GPT-4.1 nano |
|
|
14
|
+
| Read checkpoint | GPT-4.1 mini | GPT-4.1 mini |
|
|
15
|
+
| Plan recovery | GPT-4.1 mini | Gemini 2.5 Flash Lite |
|
|
16
|
+
|
|
17
|
+
Each model receives the accumulated context and tool results from preceding tasks. OpenMerit evaluates the **whole route** and can change the application's three-step model configuration after approval. The baseline remains fixed. The route is a supplied candidate, not an unrestricted search over all model combinations or evidence that each individual assignment is optimal. The workflow itself remains prepared in advance.
|
|
18
|
+
|
|
19
|
+
## Evaluation policy
|
|
20
|
+
|
|
21
|
+
Eight synthetic recovery goals cover stale runbooks, non-contiguous commits, malformed rows, completed jobs, checksum conflicts, unknown failures, and instructions embedded in incident notes. Eligibility requires all eight complete manifests to match, correct tool use and output structure on every goal, no committed or nonexistent partition scheduled for replay/quarantine, and all required evidence citations. Mean whole-goal latency must be at most eight seconds and mean whole-goal cost at most $0.003. These are constraints on the observed suite, not per-request timeouts or guarantees about other jobs. The lowest-cost eligible frontier candidate must also cost less than the incumbent. Six disjoint goals verify an approved change.
|
|
22
|
+
|
|
23
|
+
Cost adds every model turn in each goal. Latency measures the complete sequential workflow, including tool dispatch. The `total_task_cost` assessment retains actual spend over the sample window, including every agent call. The per-goal budget is multiplied by the window size for comparison and by the six-goal verification size after a swap; the display divides actual spend by goal count. Provider-level call records, tool requests/results, checks, and manifests are retained in the run evidence. The recording includes the sample's actual tool trace. The evaluation records are `BF-4101` through `BF-4108`; their incidents and ledgers are fixed.
|
|
24
|
+
|
|
25
|
+
## Recorded outcome
|
|
26
|
+
|
|
27
|
+
In the recorded run, the fixed agent completed 6/8 goals and the mixed-model route completed 8/8. Mean whole-goal costs were $0.0007742 and $0.0004018625; latencies were 4.12s and 3.80s. This is about 48% lower cost and 8% lower latency in that run. All six post-swap goals passed, and a custom request returned the expected quarantine plan through both paths. The displayed sample preserves a baseline mistake: it scheduled the malformed partition for both replay and quarantine; the mixed route excluded it from replay. The recording's step models come from the actual provider calls and tool traces, not labels added to the earlier single-model recording. These are one run's observations, not promised outcomes.
|
|
28
|
+
|
|
29
|
+
See [the demo overview](demo.md) for playback controls, the shared comparison lifecycle, and how to read the measurements.
|
|
@@ -0,0 +1,47 @@
|
|
|
1
|
+
# Field kit · Six tasks
|
|
2
|
+
|
|
3
|
+
[Watch the six-task field-kit recording](https://openmerit.site/demo/?task=fieldkit). The captured run compares two heterogeneous model routes using real OpenRouter calls through Pi and OpenMerit. Each route uses six distinct models; five assignments change and the feasibility model is retained.
|
|
4
|
+
|
|
5
|
+
The field-kit workflow prepares an offline recording handoff from a Spanish technician brief, a synthetic equipment-label image, an HTML operating manual, and a Python capture script. Both paths already use **six different models**. The approved route pairs narrow specialists with broader coding, reasoning, and writing models:
|
|
6
|
+
|
|
7
|
+
## Workflow structure
|
|
8
|
+
|
|
9
|
+
| Task | Original mix | Approved mix | Capability being exercised |
|
|
10
|
+
| --- | --- | --- | --- |
|
|
11
|
+
| Translate brief | [Qwen3.6 Plus](https://openrouter.ai/qwen/qwen3.6-plus) | [Hy-MT2 1.8B](https://openrouter.ai/tencent/hy-mt2-1.8b) | Spanish → English translation; a compact translation specialist |
|
|
12
|
+
| Read label | [Qwen3.6 Flash](https://openrouter.ai/qwen/qwen3.6-flash) | [Perceptron Mk1.5](https://openrouter.ai/perceptron/perceptron-mk1.5) | Visual perception from an actual PNG input |
|
|
13
|
+
| Extract limits | [Gemini 3.1 Flash Lite](https://openrouter.ai/google/gemini-3.1-flash-lite) | [Schematron V2 Turbo](https://openrouter.ai/inference-net/schematron-v2-turbo) | Active-profile extraction from HTML; a 3B extraction model |
|
|
14
|
+
| Inspect code | [MiMo-V2.5 Pro](https://openrouter.ai/xiaomi/mimo-v2.5-pro) | [Laguna XS 2.1](https://openrouter.ai/poolside/laguna-xs-2.1) | Static Python analysis; a coding model with 33B total / 3B active parameters |
|
|
15
|
+
| Check feasibility | [Qwen3.7 Plus](https://openrouter.ai/qwen/qwen3.7-plus) | Qwen3.7 Plus (retained) | Reconcile the translated brief with equipment, operating limits, and code |
|
|
16
|
+
| Write handoff | [Gemma 4 31B](https://openrouter.ai/google/gemma-4-31b-it) | [Gemma 4 26B A4B](https://openrouter.ai/google/gemma-4-26b-a4b-it) | Copy the decision into the final schema with ordered evidence |
|
|
17
|
+
|
|
18
|
+
These are the supplied model assignments for this development workflow. Each route uses six distinct models; five assignments change and the reasoning model is retained. They describe the capabilities being exercised; the small suite does not establish broad capability rankings.
|
|
19
|
+
|
|
20
|
+
## Handoff rules
|
|
21
|
+
|
|
22
|
+
Six constraints govern the handoff: no network uploads, a compatible capture sample rate, enough battery energy, enough storage, completion by the same-day deadline, and a matching connector. Equality at a limit passes. The result must list every failing constraint in policy order, or return `ready` when none fail. No hardware is changed or dispatched, scripts are never executed, and no recording is uploaded.
|
|
23
|
+
|
|
24
|
+
The runner fixes the six-stage sequence and dispatches read-only source adapters. Translation feeds feasibility; label and manual extraction provide equipment limits; code analysis reports the final sample rate and upload flag; feasibility combines those actual outputs; writing copies that decision and its evidence. Intermediate errors are retained and scored even if a later model happens to produce a correct final answer. Raw scripts and source documents are data, not instructions to change the workflow. The vision models receive the same image pixels, without a text transcription of the label. The HTML extractor receives its extraction instructions in the JSON Schema. Translation uses plain text; code analysis requests JSON without enforced response formatting because Laguna does not support that API parameter. Other structured tasks use JSON Schema. Both routes receive the same task inputs, encodings, and output budgets: 180 tokens for each of the first four tasks, then 2,048 for the feasibility model (including a low reasoning effort) and 350 for the handoff. Optional reasoning is disabled for the other tasks where supported. No failed answer is repaired or substituted.
|
|
25
|
+
|
|
26
|
+
## Evaluation policy
|
|
27
|
+
|
|
28
|
+
Four evaluation goals cover ready and blocked kits across all six constraints. Every structured intermediate output, the final manifest, and the evidence list must match. The translation is checked through the resulting duration, deadline, and offline decision rather than a verbatim wording match. The source-adapter completion check is reported as `tool_use_performance`; it does not imply free-form tool selection by the models. Mean whole-goal latency must be ≤45 seconds and inference cost ≤$0.006. Four disjoint goals verify the approved route, including exact resource and deadline boundaries. This smaller suite keeps six-turn goals within the existing 120-request, ten-minute, and $1.25 session bounds. The evaluation records are `KIT-8101` through `KIT-8104`.
|
|
29
|
+
|
|
30
|
+
## Captured result
|
|
31
|
+
|
|
32
|
+
The supervised local run was verified on September 28, 2026.
|
|
33
|
+
|
|
34
|
+
| Path | Complete evaluation goals | Mean model cost per goal | Mean whole-goal latency |
|
|
35
|
+
| --- | --- | --- | --- |
|
|
36
|
+
| Original six-model mix | 4/4 | $0.001867339 | 29.49s |
|
|
37
|
+
| Approved six-model mix | 4/4 | $0.001449088 | 25.63s |
|
|
38
|
+
|
|
39
|
+
The approved route had **22% lower measured cost** and **13% lower mean latency**, with all intermediate and final checks passing. After operator approval, the application configuration changed and all 4/4 disjoint verification goals passed. The displayed sample is fieldkit-01; both paths return the same correct handoff. This demonstrates an observed efficiency improvement on the prepared suite, not an accuracy gain or a production reliability estimate.
|
|
40
|
+
|
|
41
|
+
The recording is a compressed replay of local real-provider calls, not a fresh Cloudflare run. All application models were released within March 28–September 28, 2026; dated OpenRouter sources are included in Run details and the downloaded evidence. Build: `demo-20260928-fieldkit-local-v5`. Source export SHA-256: `9f0986917f41354bb8ce80421bb9a413825fb457eb37c78af119747880a1514b`. Costs exclude Pi orchestration and infrastructure.
|
|
42
|
+
|
|
43
|
+
## What OpenMerit compares
|
|
44
|
+
|
|
45
|
+
The supplied original route and challenger are compared as complete configurations, with goals executed one at a time on both paths. OpenMerit measures them, validates the comparison, proposes an eligible improvement, and changes the application configuration after approval. It does not generate these API adapters or exhaustively search model assignments. Earlier integration comparisons were rejected after a feasibility answer exhausted its output budget and after a cheaper reasoning model misread sufficient storage. The revised challenger retains Qwen3.7 Plus for feasibility. Provider timeouts also stopped two incomplete comparisons. Those attempts are not presented as successful swaps.
|
|
46
|
+
|
|
47
|
+
See [the demo overview](demo.md) for playback controls, the shared comparison lifecycle, and how to read the measurements.
|
|
@@ -0,0 +1,30 @@
|
|
|
1
|
+
# GPU placement
|
|
2
|
+
|
|
3
|
+
The [GPU placement workflow](https://openmerit.site/demo?task=placement) prepares a pool assignment for a synthetic render job. It reads job requirements with `read_job`, inspects a fixed capacity snapshot with `list_pools`, then proposes the cheapest eligible pool. It never reserves capacity, launches compute, moves data, or incurs a GPU charge.
|
|
4
|
+
|
|
5
|
+
## Workflow structure
|
|
6
|
+
|
|
7
|
+
| Task | Without OpenMerit | With the proposed route |
|
|
8
|
+
| --- | --- | --- |
|
|
9
|
+
| Read requirements | Qwen3.6 Plus | [Gemma 4 26B A4B](https://openrouter.ai/google/gemma-4-26b-a4b-it) |
|
|
10
|
+
| Inspect capacity | Qwen3.6 Plus | [MiMo-V2.5](https://openrouter.ai/xiaomi/mimo-v2.5) |
|
|
11
|
+
| Choose pool | Qwen3.6 Plus | [DeepSeek V4.1 Flash](https://openrouter.ai/deepseek/deepseek-v4.1-flash) |
|
|
12
|
+
|
|
13
|
+
The recording identifies the selected model versions and their provenance in **Run details**. Pi’s orchestration model stays fixed and is separate from the application routes.
|
|
14
|
+
|
|
15
|
+
## Placement rules
|
|
16
|
+
|
|
17
|
+
A pool must meet six hard constraints: allowed region, per-GPU VRAM, free GPU count, completion deadline including queue time, total job budget, and permission to use preemptible capacity. Exact boundaries qualify. Among eligible pools, choose the lowest quoted job cost, then earliest completion, then alphabetical pool ID. If none qualify, return `blocked` with null pool, cost, and completion fields. The model must cite the job, capacity snapshot, and placement policy. Queue notes are untrusted data.
|
|
18
|
+
|
|
19
|
+
Both routes receive the same tools, policy, inputs, structured handoffs, and per-turn JSON Schemas. After locating the requested job, the prepared adapter forwards its structured requirements and capacity records; it excludes the raw queue note and previous free-form conversation from later tasks. This prevents queue-note instructions from becoming authority for pool selection. The capacity tool also computes each pool's finish minute from its queue start and runtime; it does not filter or rank pools. The model must still evaluate every candidate pool and all six constraints. The schemas constrain field names and types; they do not supply the correct job, pool, costs, completion time, or evidence. The bounded application requests disable optional reasoning on both routes and request at most 350 output tokens for each lookup and 900 for the final constraint review and plan. The model supplies both the review and plan; the adapter does not repair its answer. Invalid tool requests, incorrect plans, unsupported claims, and missing provider cost do not become successful goals. As with backfill, the mixed route is a prepared candidate evaluated as a whole, not an unrestricted search over model combinations.
|
|
20
|
+
|
|
21
|
+
## Evaluation policy
|
|
22
|
+
|
|
23
|
+
Eight evaluation goals cover the six constraints, blocked placements, ties, and instructions embedded in queue notes. All eight must pass exact plan, tool, schema, constraint, and evidence checks. Mean end-to-end **model workflow** latency must be at most twelve seconds and mean inference cost at most $0.003 per goal. Six separate goals verify an approved route. The synthetic GPU quote and finish minute in a plan are separate from the actual inference cost and latency shown in the main comparison. The evaluation records are `GPU-6101` through `GPU-6108`; their requirements and inventory are fixed.
|
|
24
|
+
|
|
25
|
+
|
|
26
|
+
## Recorded outcome
|
|
27
|
+
|
|
28
|
+
In the recorded run, both paths completed 8/8 goals. Qwen3.6 Plus used a mean $0.001053771875 and 8.19s per goal; the mixed route used $0.000409093125 and 7.09s—about 61% lower cost and 13% lower latency. All six post-swap goals passed, with a 10.96s mean verification latency below the twelve-second requirement. A custom budget-constrained request correctly returned `blocked` through both paths. These are real local OpenRouter calls from one prepared run, not promised savings or production reliability estimates.
|
|
29
|
+
|
|
30
|
+
See [the demo overview](demo.md) for playback controls, the shared comparison lifecycle, and how to read the measurements.
|
|
@@ -0,0 +1,15 @@
|
|
|
1
|
+
# Spam filter · Jev
|
|
2
|
+
|
|
3
|
+
The [Jev task](https://openmerit.site/demo?task=spam) classifies a community comment as `publish` or `hold` under a narrow promotional-spam policy. Ordinary discussion, criticism, and reports of spam are allowed. Unsolicited promotional solicitations are held. The fixtures contain twelve examples of each class, with six separate verification comments.
|
|
4
|
+
|
|
5
|
+
## Models and integration
|
|
6
|
+
|
|
7
|
+
The challenger is [TypeSafe's Jev 1.13](https://openrouter.ai/docs/guides/community/jev), a model that returns typed decisions through OpenRouter's Decisions API. The baseline generates a JSON decision; Jev answers a prepared Choice question. Both receive the same comment and decision criteria, and the application normalizes both results to the same one-field schema. The API adapter is prepared in advance. This demo shows OpenMerit evaluating and switching the application configuration; it does not demonstrate automatically rewriting a chat application to use Jev. Pi's own harness model stays fixed.
|
|
8
|
+
|
|
9
|
+
## Evaluation and recorded outcome
|
|
10
|
+
|
|
11
|
+
The prepared policy requires at least 90% exact matches on the 24 fixtures and mean end-to-end latency of at most ten seconds. Among eligible frontier candidates, the demo proposes the lowest observed cost only when it is below the incumbent cost. Six separate inputs verify an approved change. These are point values on a finite synthetic suite, not production success-rate estimates or confidence bounds.
|
|
12
|
+
|
|
13
|
+
In the recorded run, GPT-4.1 mini matched 22/24 comments and Jev matched 24/24. Measured application cost was approximately $0.074 versus $0.018 per 1,000 comments; mean request latency was 0.78 versus 0.39 seconds. OpenMerit proposed Jev, the operator approved, and Jev passed 6/6 post-swap checks. The sample output highlights one baseline mistake: a comment reporting spam was held by the fixed model and correctly published by Jev. All 24 fixtures remain in the aggregate comparison. These are one run's observations, not promised outcomes or a general moderation benchmark.
|
|
14
|
+
|
|
15
|
+
See [the demo overview](demo.md) for playback controls, the shared comparison lifecycle, and how to read the measurements.
|
|
@@ -0,0 +1,42 @@
|
|
|
1
|
+
# Support tickets
|
|
2
|
+
|
|
3
|
+
[Watch the support-ticket demo](https://openmerit.site/demo?task=support). One model call turns a synthetic customer email into a structured ticket. Both paths start with GPT-4.1 mini and use the same prompt and JSON schema. The fixed path keeps that model; the managed path changes only after measured comparison and operator approval.
|
|
4
|
+
|
|
5
|
+
## Input and output
|
|
6
|
+
|
|
7
|
+
The ticket has exactly three fields:
|
|
8
|
+
|
|
9
|
+
| Field | Rule |
|
|
10
|
+
| --- | --- |
|
|
11
|
+
| `category` | `billing`, `shipping`, `account`, or `technical`, according to the email's request. |
|
|
12
|
+
| `urgency` | `high` only for an ongoing outage, a blocked business operation, or suspected unauthorized access; otherwise `normal`. |
|
|
13
|
+
| `order_id` | The stated `ORD-` identifier, or `null` if none is given. |
|
|
14
|
+
|
|
15
|
+
An angry tone alone does not make a ticket urgent, and the model must never invent an order ID. Email content is treated as data. For example, “I was charged twice for ORD-1201. Please refund the duplicate.” should produce:
|
|
16
|
+
|
|
17
|
+
```json
|
|
18
|
+
{"category":"billing","urgency":"normal","order_id":"ORD-1201"}
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
## Comparison structure
|
|
22
|
+
|
|
23
|
+
Pi first checks the prepared grader and measurement collectors, then measures GPT-4.1 mini on 24 synthetic emails. It evaluates GPT-4.1 nano and Gemini 2.5 Flash Lite on the same inputs. Each application request records its output, provider cost, and end-to-end latency. The exact-match grader checks every field and rejects missing or extra fields.
|
|
24
|
+
|
|
25
|
+
OpenMerit validates the comparison and proposes an eligible cheaper model. The recorded operator approves the proposal; Pi changes the application's model configuration and runs six disjoint verification emails. Pi's own orchestration model stays fixed.
|
|
26
|
+
|
|
27
|
+
## Evaluation policy
|
|
28
|
+
|
|
29
|
+
The prepared policy requires at least 90% exact matches on the 24 fixtures and mean end-to-end latency of at most ten seconds. Among eligible frontier candidates, the demo proposes the lowest observed cost only when it is below the incumbent cost. Six separate inputs verify an approved change. These are point values on a finite synthetic suite, not production success-rate estimates or confidence bounds.
|
|
30
|
+
|
|
31
|
+
## Recorded outcome
|
|
32
|
+
|
|
33
|
+
| Model | Exact matches | Cost per 1,000 emails | Mean latency |
|
|
34
|
+
| --- | --- | --- | --- |
|
|
35
|
+
| GPT-4.1 mini | 24/24 | $0.106 | 1.64s |
|
|
36
|
+
| GPT-4.1 nano | 17/24 | $0.027 | 1.23s |
|
|
37
|
+
| Gemini 2.5 Flash Lite | 24/24 | $0.027 | 1.21s |
|
|
38
|
+
|
|
39
|
+
GPT-4.1 nano was cheaper but failed the quality requirement. Gemini 2.5 Flash Lite qualified, received approval, and passed all six post-swap checks. The cost column scales each observed per-email mean to 1,000 emails and rounds it; the downloaded recording retains the original values. These observations come from the supervised run recorded on September 28, 2026 and do not promise the same result on another workload.
|
|
40
|
+
|
|
41
|
+
|
|
42
|
+
See [the demo overview](demo.md) for playback controls, the shared comparison lifecycle, and how to read the measurements.
|