openmerit 0.1.3 → 0.1.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (77) hide show
  1. package/CHANGELOG.md +26 -0
  2. package/README.md +83 -312
  3. package/dist/core/src/index.d.ts +90 -0
  4. package/dist/core/src/index.js +1137 -0
  5. package/dist/core/src/store.d.ts +35 -0
  6. package/dist/core/src/store.js +102 -0
  7. package/dist/pi/src/index.d.ts +15 -0
  8. package/dist/pi/src/index.js +423 -0
  9. package/dist/protocol/src/index.d.ts +402 -0
  10. package/dist/protocol/src/index.js +47 -0
  11. package/dist/protocol/src/schemas.d.ts +450 -0
  12. package/dist/protocol/src/schemas.js +224 -0
  13. package/docs/adapter-guide.md +172 -0
  14. package/docs/architecture.md +55 -0
  15. package/docs/automation.md +66 -0
  16. package/docs/getting-started.md +55 -0
  17. package/docs/lifecycle.md +30 -0
  18. package/docs/metrics-and-evidence.md +40 -0
  19. package/docs/operations.md +31 -0
  20. package/docs/pareto-spec.md +76 -0
  21. package/docs/pi-extension.md +44 -0
  22. package/docs/roadmap.md +26 -0
  23. package/docs/security.md +23 -0
  24. package/docs/testing.md +36 -0
  25. package/docs/ux-reference.md +32 -0
  26. package/package.json +45 -54
  27. package/benchmark/invoice_ocr/data/invoice_01_ground_truth.json +0 -38
  28. package/benchmark/invoice_ocr/data/invoice_01_row_2.jpg +0 -0
  29. package/benchmark/invoice_ocr/data/invoice_02_ground_truth.json +0 -32
  30. package/benchmark/invoice_ocr/data/invoice_02_row_5.jpg +0 -0
  31. package/benchmark/invoice_ocr/data/invoice_03_ground_truth.json +0 -26
  32. package/benchmark/invoice_ocr/data/invoice_03_row_6.jpg +0 -0
  33. package/benchmark/invoice_ocr/data/invoice_04_ground_truth.json +0 -26
  34. package/benchmark/invoice_ocr/data/invoice_04_row_7.jpg +0 -0
  35. package/benchmark/invoice_ocr/data/invoice_05_ground_truth.json +0 -38
  36. package/benchmark/invoice_ocr/data/invoice_05_row_947.jpg +0 -0
  37. package/benchmark/invoice_ocr/data/invoice_06_ground_truth.json +0 -38
  38. package/benchmark/invoice_ocr/data/invoice_06_row_948.jpg +0 -0
  39. package/benchmark/invoice_ocr/data/invoice_07_ground_truth.json +0 -20
  40. package/benchmark/invoice_ocr/data/invoice_07_row_949.jpg +0 -0
  41. package/benchmark/invoice_ocr/data/invoice_08_ground_truth.json +0 -38
  42. package/benchmark/invoice_ocr/data/invoice_08_row_1888.jpg +0 -0
  43. package/benchmark/invoice_ocr/data/invoice_09_ground_truth.json +0 -26
  44. package/benchmark/invoice_ocr/data/invoice_09_row_1890.jpg +0 -0
  45. package/benchmark/invoice_ocr/data/invoice_10_ground_truth.json +0 -20
  46. package/benchmark/invoice_ocr/data/invoice_10_row_1892.jpg +0 -0
  47. package/benchmark/invoice_ocr/data/manifest.json +0 -97
  48. package/dist/benchmarks.js +0 -98
  49. package/dist/catalog.js +0 -61
  50. package/dist/cli.js +0 -188
  51. package/dist/daemon.js +0 -388
  52. package/dist/diagnostics.js +0 -194
  53. package/dist/frontier.js +0 -56
  54. package/dist/harness.js +0 -1
  55. package/dist/integrations.js +0 -19
  56. package/dist/invoice-eval.js +0 -33
  57. package/dist/invoice-score.js +0 -124
  58. package/dist/judge.js +0 -43
  59. package/dist/llm.js +0 -203
  60. package/dist/pi-trials.js +0 -366
  61. package/dist/policy.js +0 -115
  62. package/dist/providers.js +0 -1
  63. package/dist/recommend.js +0 -76
  64. package/dist/routes.js +0 -59
  65. package/dist/standalone.js +0 -220
  66. package/dist/store.js +0 -89
  67. package/dist/strategist.js +0 -68
  68. package/dist/task-input.js +0 -54
  69. package/dist/traces.js +0 -127
  70. package/dist/trials.js +0 -140
  71. package/dist/types.js +0 -2
  72. package/examples/invoice-prompt.txt +0 -19
  73. package/examples/task.example.json +0 -7
  74. package/extension/openmerit.ts +0 -820
  75. package/instructions/OPENMERIT.md +0 -54
  76. package/instructions/openmerit.policy.json +0 -33
  77. package/rules.md +0 -39
@@ -1,54 +0,0 @@
1
- # OpenMerit — Model Merit Harness
2
-
3
- You are running under a main agent harness that is observed by **OpenMerit**, an
4
- external merit harness. OpenMerit's job is to make sure you are always running on
5
- the best model for the task at hand, at the best price, with a vetted fallback.
6
-
7
- ## What OpenMerit does
8
-
9
- 1. **Observes** this harness's session traces (models used, tokens, cost,
10
- latency, errors) without intercepting or slowing down your work.
11
- 2. **Evaluates** candidate models for each completed task sequentially through
12
- pi and scores their outputs with a judge model. Candidate tools are disabled
13
- by default.
14
- 3. **Maintains a pareto frontier** per task (quality vs. cost vs. latency) and
15
- an aggregate frontier across all of your tasks.
16
- 4. **Watches for new model releases** (provider catalogs + public benchmarks)
17
- and queues promising releases for trial against your existing frontier.
18
- 5. **Recommends or applies model swaps**, with reasoning, and keeps your
19
- fallback model up to date.
20
-
21
- ## What you should do
22
-
23
- - **Treat model changes as explicit routing decisions.** If you notice your
24
- model identity change, it was approved by the user or passed their opt-in
25
- auto-apply policy. Continue the task; your instructions and context are
26
- unchanged.
27
- - **Respect the fallback.** If your current model errors, rate-limits, or is
28
- retired, OpenMerit can offer the current fallback or apply it under an
29
- opt-in automatic policy. Treat an approved fallback as normal operation.
30
- - **State your task clearly in your first message** of a session when possible.
31
- OpenMerit keys its pareto frontier to task signatures; clear task statements
32
- produce better model choices for you.
33
- - **Do not edit OpenMerit state files** (`~/.openmerit/`). If something looks
34
- wrong, tell the user.
35
-
36
- ## What you can ask the user for
37
-
38
- - `/openmerit` — show current model, fallback, frontier position, and pending
39
- recommendations.
40
- - `/openmerit apply` / `/openmerit dismiss` — act on a pending recommendation
41
- when automatic application is declined by the policy gate.
42
- - Policy changes (auto-apply thresholds, budgets, provider allow-lists) are
43
- made by the user in `~/.openmerit/policy.json`, not by you.
44
-
45
- ## Guarantees
46
-
47
- - Per-task comparisons send the task text, uploaded files or images (when
48
- present), and candidate answers through the model routes configured in Pi
49
- for model runs and judging. Candidate Pi runs default to no tools; any tool access is an
50
- explicit user opt-in. Use non-sensitive examples while
51
- evaluating this alpha.
52
- - The shipped policy is supervised. Swaps happen automatically only after the
53
- user opts in and the configured quality/cost guardrails pass; otherwise they
54
- remain recommendations for a human to approve.
@@ -1,33 +0,0 @@
1
- {
2
- "version": 1,
3
- "mode": "recommend",
4
- "auto_apply": {
5
- "enabled": false,
6
- "min_score_gain": 0.1,
7
- "max_price_ratio": 1.5,
8
- "require_frontier": true
9
- },
10
- "budgets": {
11
- "max_usd_per_trial": 0.25,
12
- "max_trials_per_day": 20,
13
- "max_usd_per_day": 5.0
14
- },
15
- "providers": {
16
- "allow": ["*"],
17
- "deny": []
18
- },
19
- "watch": {
20
- "catalog_interval_min": 360,
21
- "traces_interval_sec": 5,
22
- "trial_interval_min": 30,
23
- "models_per_task": 3
24
- },
25
- "fallback": {
26
- "auto_update": true,
27
- "min_score": 0.6,
28
- "apply_on_error": true
29
- },
30
- "judge_model": null,
31
- "strategist_model": null,
32
- "max_usd_per_m": 20.0
33
- }
package/rules.md DELETED
@@ -1,39 +0,0 @@
1
- # OpenMerit product and architecture rules
2
-
3
- - **Optimize for useful model choices.** OpenMerit exists to find better task-specific tradeoffs between quality, cost, and latency—not to become a general-purpose agent or telemetry platform.
4
-
5
- - **Act as a control plane, not a gateway.** Normal traffic stays between the harness and provider; OpenMerit observes evidence, runs explicit trials, and returns recommendations without intercepting ordinary work.
6
-
7
- - **Make Pi plug-and-play first.** Pi and Pi-based harnesses are the first supported user experience because that is where early users already work; broader harness support can follow without weakening this path.
8
-
9
- - **Keep the core harness-neutral.** Pi is the first `HarnessAdapter`, not a permanent assumption in scoring, recommendations, policy, or stored data, so other harnesses can be added without rewriting the merit loop.
10
-
11
- - **Keep model providers interchangeable.** OpenRouter is a first-class provider and discovery source, not the foundation of the domain model; provider-specific behavior belongs behind a `ModelProviderAdapter`.
12
-
13
- - **Require end-to-end provider neutrality.** Discovery, authentication, execution, scoring, and swapping must preserve the selected route; a feature is not provider-neutral if only its API client is abstracted.
14
-
15
- - **Separate model identity from execution route.** Keep logical model IDs in stable `vendor/model` form, while recording the actual provider or route separately, so the same model can be compared through OpenRouter, a native provider, or a local provider.
16
-
17
- - **Respect the harness's eligible model pool.** Prefer models already available and configured in the active harness; external catalogs may enrich or expand discovery but must not silently override harness scope or credentials.
18
-
19
- - **Treat public benchmarks as priors, not proof.** Benchmarks help shortlist candidates, but merit comes from trials on the user's actual task.
20
-
21
- - **Keep four integration boundaries distinct.** Model execution (`ModelProviderAdapter`), agent execution (`HarnessAdapter`), incoming traces (`ObservationSource`), and outgoing telemetry (`EventSink`) solve different problems and must not be coupled.
22
-
23
- - **Use provider-neutral core records.** Normalize integrations into stable concepts such as `TaskObservation`, `CandidateRun`, `TrialScore`, `Frontier`, `Recommendation`, and `MeritEvent`, so integrations do not leak their schemas into the decision engine.
24
-
25
- - **Optimize per task before aggregating per agent.** Agent-level conclusions must be built from measured task evidence rather than assumed from global model rankings.
26
-
27
- - **Version persisted events and evolve them additively.** Existing 0.1.x state must remain readable, and append-only recommendation history must stay intact; migrations should normalize old records rather than invalidate them.
28
-
29
- - **Keep policy as the sole auto-swap authority.** Trials and strategists may recommend changes, but only the policy gate may approve automatic application, and every recommendation must retain its gate reasons.
30
-
31
- - **Default to safe, explicit trials.** Candidate tool access stays off unless deliberately allowed, trials remain isolated, and temporary workspaces are cleaned up because trying a model must not expose or damage a user's project by surprise.
32
-
33
- - **Keep the shadow track isolated, not necessarily concurrent.** Candidate trials may run sequentially to respect cost, rate, and safety limits while remaining separate from the user's live task.
34
-
35
- - **Treat local JSONL as the default integration, not a lock-in.** Local traces and events should work without an external service; systems such as Langfuse can later plug in as observation sources or event sinks.
36
-
37
- - **Keep integrations optional and the core lightweight.** New providers, harnesses, and observability services should not impose credentials, network calls, or heavy dependencies on users who do not enable them.
38
-
39
- - **Keep 0.1.x focused.** The near-term bar is reliable, provider-neutral use for Pi tinkerers and solo hackers; additional harnesses and hosted observability integrations belong in later releases unless required to prove the boundaries work.