openmerit 0.1.4 → 0.1.6-preview.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (147) hide show
  1. package/CHANGELOG.md +40 -0
  2. package/README.md +121 -386
  3. package/dist/core/src/index.d.ts +101 -0
  4. package/dist/core/src/index.js +1649 -0
  5. package/dist/core/src/store.d.ts +35 -0
  6. package/dist/core/src/store.js +102 -0
  7. package/dist/pi/src/index.d.ts +32 -0
  8. package/dist/pi/src/index.js +794 -0
  9. package/dist/pi/src/scheduler.d.ts +11 -0
  10. package/dist/pi/src/scheduler.js +137 -0
  11. package/dist/pi/src/wakeup.d.ts +2 -0
  12. package/dist/pi/src/wakeup.js +108 -0
  13. package/dist/protocol/src/index.d.ts +484 -0
  14. package/dist/protocol/src/index.js +47 -0
  15. package/dist/protocol/src/schemas.d.ts +576 -0
  16. package/dist/protocol/src/schemas.js +280 -0
  17. package/dist/terminal/public/app.js +297 -0
  18. package/dist/terminal/public/brands/anthropic.png +0 -0
  19. package/dist/terminal/public/brands/baai.png +0 -0
  20. package/dist/terminal/public/brands/baseten.png +0 -0
  21. package/dist/terminal/public/brands/cerebras.png +0 -0
  22. package/dist/terminal/public/brands/cohere.png +0 -0
  23. package/dist/terminal/public/brands/deepseek.ico +0 -0
  24. package/dist/terminal/public/brands/google.png +0 -0
  25. package/dist/terminal/public/brands/groq.ico +0 -0
  26. package/dist/terminal/public/brands/lm-studio.png +0 -0
  27. package/dist/terminal/public/brands/meta.ico +0 -0
  28. package/dist/terminal/public/brands/mistral.png +0 -0
  29. package/dist/terminal/public/brands/nomic.png +0 -0
  30. package/dist/terminal/public/brands/ollama.png +0 -0
  31. package/dist/terminal/public/brands/openai.png +0 -0
  32. package/dist/terminal/public/brands/openrouter.png +0 -0
  33. package/dist/terminal/public/brands/qwen.png +0 -0
  34. package/dist/terminal/public/brands/vllm.ico +0 -0
  35. package/dist/terminal/public/brands/vllm.png +0 -0
  36. package/dist/terminal/public/favicon.svg +1 -0
  37. package/dist/terminal/public/flow.css +1 -0
  38. package/dist/terminal/public/flow.js +770 -0
  39. package/dist/terminal/public/index.html +21 -0
  40. package/dist/terminal/public/styles.css +779 -0
  41. package/dist/terminal/src/activity-merge.mjs +64 -0
  42. package/dist/terminal/src/browser.mjs +29 -0
  43. package/dist/terminal/src/cli.mjs +60 -0
  44. package/dist/terminal/src/collect.mjs +311 -0
  45. package/dist/terminal/src/discovery.mjs +93 -0
  46. package/dist/terminal/src/hardware.mjs +57 -0
  47. package/dist/terminal/src/project-activity.mjs +156 -0
  48. package/dist/terminal/src/sample.mjs +171 -0
  49. package/dist/terminal/src/server.mjs +56 -0
  50. package/dist/terminal/src/services.mjs +62 -0
  51. package/dist/terminal/src/topology.mjs +30 -0
  52. package/docs/adapter-guide.md +189 -0
  53. package/docs/architecture.md +59 -0
  54. package/docs/automation.md +74 -0
  55. package/docs/budgets.md +37 -0
  56. package/docs/commands.md +85 -0
  57. package/docs/demo-backfill.md +29 -0
  58. package/docs/demo-fieldkit.md +47 -0
  59. package/docs/demo-placement.md +30 -0
  60. package/docs/demo-spam.md +15 -0
  61. package/docs/demo-support.md +42 -0
  62. package/docs/demo.md +57 -0
  63. package/docs/first-trial.md +60 -0
  64. package/docs/getting-started.md +65 -0
  65. package/docs/index.md +40 -0
  66. package/docs/inference-terminal.md +439 -0
  67. package/docs/lifecycle.md +30 -0
  68. package/docs/memo.md +126 -0
  69. package/docs/metrics-and-evidence.md +48 -0
  70. package/docs/operations.md +40 -0
  71. package/docs/pareto-spec.md +76 -0
  72. package/docs/pi-extension.md +54 -0
  73. package/docs/roadmap.md +28 -0
  74. package/docs/security.md +37 -0
  75. package/docs/site-artwork-linocut.md +23 -0
  76. package/docs/site-artwork-miniature-diverse.md +28 -0
  77. package/docs/site-artwork-miniature.md +26 -0
  78. package/docs/site-demo.md +177 -0
  79. package/docs/site-design.md +94 -0
  80. package/docs/site-documentation.md +83 -0
  81. package/docs/site-dynamic-og.md +35 -0
  82. package/docs/site-faq-maintenance.md +115 -0
  83. package/docs/site-hero-resolution.md +60 -0
  84. package/docs/site-illustration-sequences.md +227 -0
  85. package/docs/site-inference-terminal.md +203 -0
  86. package/docs/site-memo.md +39 -0
  87. package/docs/site-og-image.md +38 -0
  88. package/docs/site-og-workshop.md +21 -0
  89. package/docs/site-section-artwork.md +56 -0
  90. package/docs/site-skill-review.md +57 -0
  91. package/docs/site-terminal-preview.md +85 -0
  92. package/docs/testing.md +118 -0
  93. package/docs/troubleshooting.md +55 -0
  94. package/docs/ux-reference.md +32 -0
  95. package/package.json +74 -42
  96. package/benchmark/invoice_ocr/data/invoice_01_ground_truth.json +0 -38
  97. package/benchmark/invoice_ocr/data/invoice_01_row_2.jpg +0 -0
  98. package/benchmark/invoice_ocr/data/invoice_02_ground_truth.json +0 -32
  99. package/benchmark/invoice_ocr/data/invoice_02_row_5.jpg +0 -0
  100. package/benchmark/invoice_ocr/data/invoice_03_ground_truth.json +0 -26
  101. package/benchmark/invoice_ocr/data/invoice_03_row_6.jpg +0 -0
  102. package/benchmark/invoice_ocr/data/invoice_04_ground_truth.json +0 -26
  103. package/benchmark/invoice_ocr/data/invoice_04_row_7.jpg +0 -0
  104. package/benchmark/invoice_ocr/data/invoice_05_ground_truth.json +0 -38
  105. package/benchmark/invoice_ocr/data/invoice_05_row_947.jpg +0 -0
  106. package/benchmark/invoice_ocr/data/invoice_06_ground_truth.json +0 -38
  107. package/benchmark/invoice_ocr/data/invoice_06_row_948.jpg +0 -0
  108. package/benchmark/invoice_ocr/data/invoice_07_ground_truth.json +0 -20
  109. package/benchmark/invoice_ocr/data/invoice_07_row_949.jpg +0 -0
  110. package/benchmark/invoice_ocr/data/invoice_08_ground_truth.json +0 -38
  111. package/benchmark/invoice_ocr/data/invoice_08_row_1888.jpg +0 -0
  112. package/benchmark/invoice_ocr/data/invoice_09_ground_truth.json +0 -26
  113. package/benchmark/invoice_ocr/data/invoice_09_row_1890.jpg +0 -0
  114. package/benchmark/invoice_ocr/data/invoice_10_ground_truth.json +0 -20
  115. package/benchmark/invoice_ocr/data/invoice_10_row_1892.jpg +0 -0
  116. package/benchmark/invoice_ocr/data/manifest.json +0 -97
  117. package/dist/benchmarks.js +0 -98
  118. package/dist/catalog.js +0 -61
  119. package/dist/cli.js +0 -188
  120. package/dist/daemon.js +0 -407
  121. package/dist/diagnostics.js +0 -227
  122. package/dist/frontier.js +0 -56
  123. package/dist/harness.js +0 -1
  124. package/dist/integrations.js +0 -19
  125. package/dist/invoice-eval.js +0 -33
  126. package/dist/invoice-score.js +0 -124
  127. package/dist/judge.js +0 -43
  128. package/dist/llm.js +0 -207
  129. package/dist/pi-config.js +0 -46
  130. package/dist/pi-trials.js +0 -373
  131. package/dist/policy.js +0 -185
  132. package/dist/providers.js +0 -1
  133. package/dist/recommend.js +0 -76
  134. package/dist/routes.js +0 -74
  135. package/dist/standalone.js +0 -224
  136. package/dist/store.js +0 -89
  137. package/dist/strategist.js +0 -68
  138. package/dist/task-input.js +0 -54
  139. package/dist/traces.js +0 -127
  140. package/dist/trials.js +0 -140
  141. package/dist/types.js +0 -2
  142. package/examples/invoice-prompt.txt +0 -19
  143. package/examples/task.example.json +0 -7
  144. package/extension/openmerit.ts +0 -947
  145. package/instructions/OPENMERIT.md +0 -63
  146. package/instructions/openmerit.policy.json +0 -37
  147. package/rules.md +0 -43
package/docs/memo.md ADDED
@@ -0,0 +1,126 @@
1
+ # Optimum Inference For All
2
+
3
+ *Lazim*
4
+
5
+ How can we accelerate the distribution of intelligence?
6
+
7
+ Building and deploying more models and improving the compute landscape is where most of the energy being spent right now. I have another, less explored direction in my mind, which I call, the optimum inference problem.
8
+
9
+ I see a substantial part of what limits autonomous systems is how the models are selected and assessed throughout their lifecycle. It's a harness-level infrastructure concern.
10
+
11
+ #### Does this need to get solved?
12
+
13
+ In companies that I had the privilege of working with, I have built or maintained or used a rudimentary version of this. When budget constraints strangled the roadmap, teams scrambled their way out by building something enough to justify why they need to pay y amount of money to solve x task. This was never a maintained system nor it was ever prioritised because it was seen as an optional optimisation issue and overlooked the power it had over the dissemination of intelligence inside the org and the products they ship.
14
+
15
+ If it was ever solved, it was always at a task level and if it had room and engineers to pressure, maybe at the agent level but rarely. even rarer at the harness level. As the system develops, its inference choices should develop with it. I don't think every team should have to build and maintain that process and stack themselves. We are building OpenMerit to solve this. In the Mission of optimum inference for all.
16
+
17
+ #### What optimum means
18
+
19
+ Optimum depends on the work. We need to know what counts as success, how reliably it must be delivered and the customer's constraints on cost, latency and operation.
20
+
21
+ A cheaper model that needs repeated attempts may make the completed job more expensive. A stronger model may add cost without improving the result. We need to account for the whole workflow, including anything a person has to fix.
22
+
23
+ The range of models matters here. There are frontier models from leading labs, general-purpose open-weight models, fine-tuned and distilled variants on Hugging Face, and specialists for coding, retrieval, vision and speech. Jev, by TypeSafe, is designed for fast structured decisions. Its recent release shows there is demand for a different breed of system-one models. Smaller models can run task-specific programs optimised through tools such as DSPy. An agent may have a place for several of these. We want OpenMerit to assess all of them on the work they can actually do, to allocate as it required.
24
+
25
+ The right deployment could use a direct API from the lab, a router, an inference cloud or a customer-operated model. It could combine them. A company using only commercial APIs should get a complete product. So should one operating entirely within its own environment.
26
+
27
+ Hardware is equally part of the decision. For customer-operated deployments, we need to account for available compute and memory, concurrent workloads and the cost of loading or swapping models. With managed APIs, we work with measurable outcomes and the constraints the provider exposes.
28
+
29
+ Local models, open weights and fine-tunes extend the possibilities. We have no preference for moving customers towards them. If the industry the customer operates in demands sovereignty then it should be able to, without sacrificing the intelligence spread.
30
+
31
+ Better allocation also matters beyond the bill. Where a smaller configuration meets the requirements and the freed capacity can be reused, the same resources can support more work. That is part of what I mean by distributing intelligence.
32
+
33
+ #### What we are building
34
+
35
+ There's no end-to-end product or platform clarity at the moment, but on a broader spectrum, you could place OpenMerit as infrastructure for assessing, benchmarking, selecting, assigning and maintaining all model choices within your harness.
36
+
37
+ It operates at task, agent and harness levels. An assignment should suit the individual task while accounting for the agent's overall result. Across a harness, those choices have to work within shared budgets and capacity.
38
+
39
+ To get there, OpenMerit first needs to establish which options the customer can use. It should then coordinate comparisons on representative work using their success criteria. Production evidence tells us how the current configuration behaves. It does not establish how an untested alternative would perform.
40
+
41
+ After deployment, OpenMerit should check whether the expected improvement happened and decide where further investigation is worthwhile. Keeping the current assignment is a valid decision.
42
+
43
+ The interface should be agent-native. A harness should be able to request a comparison and retrieve an assignment with its evidence programmatically. Within the customer's permissions, OpenMerit should coordinate validated changes through the existing infrastructure, then retain or reverse them based on the results.
44
+
45
+ We will remain harness-agnostic. The customer's harness owns execution and coordination. Its providers and serving systems run inference. We integrate with its observability and evaluation tools. We will not build those products ourselves as much as we can, and production traffic will not have to pass through us to fulfil the task, but sure can control the traffic based on the evidence.
46
+
47
+ #### A continuing research problem
48
+
49
+ Most of the components are still very nascent and will increasingly need newer ways of operating as models become increasingly intelligent and the compute becomes increasingly efficient. Hence we see OpenMerit as a research lab as much as a product company.
50
+
51
+ To zoom out, we see optimum inference as a long-term research and deployment effort. There is no final configuration we can find, hand over and consider the problem solved.
52
+
53
+ We need better methods for deciding how much evidence an assignment requires and when that evidence needs updating. Testing every option is expensive. Testing too little can leave us maintaining a poor decision. OpenMerit should spend its testing budget where additional evidence is likely to be useful.
54
+
55
+ We also need to understand how individual choices affect the complete outcome. A model that performs well on one step may leave the next step with harder work. Several assignments that look sensible separately may compete for the same compute.
56
+
57
+ These are some of the questions we will have to study in production. Deployment will expose cases our experiments missed and show us which improvements were worth pursuing. That experience should inform the next investigation and the methods we build.
58
+
59
+ OpenMerit should support this work throughout a system's creation, iteration and ongoing operation.
60
+
61
+ #### Where we start
62
+
63
+ Our first customers are engineering and AI teams running repeatable production workflows with measurable outcomes and meaningful inference costs or capacity constraints.
64
+
65
+ One way to think about it goes like this. A developer or their coding agent connects one workflow. OpenMerit establishes a baseline, compares eligible alternatives and returns a decision they can act on. They keep using it because maintaining that assignment is useful, then connect more work.
66
+
67
+ The first line of product combines L1, organisation-wide understanding, with L2, repeatable workflows. Individual workflows provide the evidence for a wider view of which models are assigned where, why those choices were made and which ones deserve reassessment and swapping.
68
+
69
+ L3, adaptive agents, is the longer-term research direction. Agents may revise their plans or create subagents while working. We want to support inference decisions during that execution without taking over the harness's responsibility.
70
+
71
+ L2 should support a business on its own. It also gives us production experience for the harder questions ahead.
72
+
73
+ #### A separate company, not a feature
74
+
75
+ We want OpenMerit to follow the customer's work across whichever models and infrastructure are appropriate.
76
+
77
+ A customer should be able to change providers, introduce a fine-tune or replace a harness without rebuilding how they make model decisions. Some evidence will need to be revalidated. The history of what they tried and why they changed should remain useful.
78
+
79
+ That continuing responsibility is substantial enough to build a company around. We intend to earn it through the quality of our decisions and the maintenance we remove from the customer's team.
80
+
81
+ We are thinking of building an open-core, with hosted and customer-operated options, and charge for continuous assessment and decision management. We will not mark up inference.
82
+
83
+ #### We need to prove
84
+
85
+ Our evaluations have to reflect the customer's actual work. We should begin with representative examples and usable success criteria, and limit autonomous changes to decisions the evidence supports. Where the criteria are inadequate, we need to identify that before automating the wrong objective.
86
+
87
+ Our own costs belong in the calculation. Comparisons consume resources. Switching can introduce failures and maintenance. OpenMerit needs to recognise when an apparent improvement is too small to pursue.
88
+
89
+ We also need dependable integrations. We should start with a narrow set and make them work well before extending our coverage. Supporting different infrastructure does not require supporting everything immediately.
90
+
91
+ The standard is a measurable improvement after accounting for testing, switching and running OpenMerit. Depending on the workload, that means better outcomes within the same budget, lower total cost at the required quality or more completed work with available capacity.
92
+
93
+ #### What is already here in some shape or form
94
+
95
+ There are useful tools here, including ones that already make model decisions. OpenRouter's Auto Router selects models for incoming requests. Portkey's gateway supports conditional routing, fallbacks and controlled testing of new models in production. Not Diamond lets teams train routers on their own data and evaluation criteria, with quality, cost and latency preferences.
96
+
97
+ Langfuse helps teams compare changes through evaluation scores, cost, latency and individual failures. Braintrust supports experiments and continuous production evaluation, including bringing production examples back into testing. These tools already help teams investigate and improve their systems.
98
+
99
+ Many of these capabilities overlap with parts of what we want to explore. They are also components we expect to work with.
100
+
101
+ We consider optimum inference far from solved. A system needs to determine when the evidence behind an assignment is no longer enough, which alternatives deserve testing and whether a change will improve the completed work after its costs are included. It needs to maintain that process across changes to its tasks, models and infrastructure.
102
+
103
+ OpenMerit exists to take on that continuing responsibility. The tools for an investigation should be available to the harness, along with the means to act on the result. The customer should not inherit a permanent engineering project to keep it all working.
104
+
105
+ #### Path ahead
106
+
107
+ We want teams to know when the strongest model is worth paying for, when an inexpensive alternative is sufficient and when a different deployment would let them do more.
108
+
109
+ We want those decisions to become part of what a system can maintain for itself. Its engineers should be able to define the work and its constraints without inheriting a permanent job of reconsidering every model assignment.
110
+
111
+ How much more useful work can we make possible with the intelligence and compute available to us?
112
+
113
+ ---
114
+
115
+ **Additional Reading**
116
+
117
+ 1. [**"The Shift from Models to Compound AI Systems"**](https://bair.berkeley.edu/blog/2024/02/18/compound-ai-systems/) — Matei Zaharia et al., BAIR. (February 18, 2024)
118
+ 2. [**"DSPy: Program, Don't Prompt"**](https://dspy.ai/current/) — DSPy / Stanford NLP.
119
+ 3. [**"Optimas: Optimizing Compound AI Systems with Globally Aligned Local Rewards"**](https://optimas.stanford.edu/) — Shirley Wu et al., arXiv. (February 9, 2026)
120
+ 4. [**"AI Agents That Matter"**](https://arxiv.org/abs/2407.01502) — Sayash Kapoor, Benedikt Stroebl, Zachary S. Siegel, Nitya Nadgir, and Arvind Narayanan, arXiv. (July 1, 2024)
121
+ 5. [**"Training a Custom Router"**](https://docs.notdiamond.ai/docs/router-training-quickstart) — Not Diamond.
122
+ 6. [**"Auto Router — Intelligent Model Selection"**](https://openrouter.ai/docs/guides/routing/routers/auto-router) — OpenRouter.
123
+ 7. [**"AI Gateway"**](https://portkey.ai/docs/product/ai-gateway) — Portkey.
124
+ 8. [**"Compare Experiments"**](https://langfuse.com/docs/evaluation/experiments/compare-experiments) — Langfuse.
125
+ 9. [**"Evaluate Systematically"**](https://www.braintrust.dev/docs/evaluate) — Braintrust.
126
+ 10. [**"Introducing System One Models & Jev"**](https://typesafe.ai/blog/introducing-system-one-models-and-jev) — Diogo Almeida, TypeSafe AI. (September 15, 2026)
@@ -0,0 +1,48 @@
1
+ # Metrics and evidence
2
+
3
+ ## Task-specific profiles
4
+
5
+ The full catalogue is available to every product, but the harness infers a task-specific proposal from the application's outputs, tools, and likely failure modes. Before collection, the user sees and confirms the required metrics, direction, sampling design, minimum sample count, aggregation, tolerance, optional hard constraints, and estimated cost.
6
+
7
+ ## Initial metric catalogue
8
+
9
+ Quality and behavior: task quality, task success, consistency, instruction following, tool-use performance, structured-output reliability, hallucination rate, reasoning efficiency, recovery ability, context handling, retrieval-use quality, long-horizon performance, and preference fit.
10
+
11
+ Cost and resources: input tokens, output tokens, total task cost, cost per successful task, and agent steps.
12
+
13
+ Latency and operations: time to first useful output, end-to-end latency, throughput, retry rate, and human-intervention rate.
14
+
15
+ ## Metric states
16
+
17
+ - `measured` — numeric value and evidence exist.
18
+ - `not_applicable` — the metric does not apply.
19
+ - `not_configured` — collection or evaluation is missing.
20
+ - `insufficient_evidence` — relevant, but without enough representative samples.
21
+
22
+ Missing measurements are never converted to zero.
23
+
24
+ ## Single and repeated runs
25
+
26
+ A single execution can report its cost, latency, token use, steps, success, and evaluator score. Distribution and reliability claims require repeated runs: rates, consistency, variance, averages, percentiles, throughput, prompt sensitivity, and regression stability.
27
+
28
+ Two forms of repetition may be needed: repeat the same input to measure stochastic variance, and run representative task instances to measure general task performance.
29
+
30
+ Each objective records one sampling strategy: `single_run`, `repeated_same_input`, `representative_task_set`, `repeated_representative_task_set`, or `load_test`. It also records representative-case count, repetitions per case, aggregation, and rationale. Every measured metric record repeats that aggregation explicitly (`single_value`, `sum`, `mean`, `rate`, `distribution`, or `percentile`). Rate, mean, distribution, percentile, and catalogue-defined reliability metrics are rejected when their plan contains fewer than two runs or uses `single_run`; `total_task_cost` sample windows must use `sum` because that value is used for budget accounting.
31
+
32
+ The task profile sets the minimum sample count for each required metric; there is no universal number of runs. `minimumSamples` is evidence quantity. An optional `constraint` is a real performance threshold in the metric's unit; OpenMerit never interprets a three-sample requirement as a metric value of at least three. Ratio thresholds must remain between zero and one. A completed-task signal can advance an assessment cadence, but it does not by itself supply a quality score. Quality requires a task-specific evaluator, known outcome, or other justified measurement. Cost, token use, and latency must come from the application call or its observability layer. The baseline and controlled challenger runs supply the measured evidence used for comparison. Catalogue price and model reputation help the harness choose candidates to test; they are not substitutes for these measurements. In the unreleased source checkout, setup records the exact incumbent, preliminary leads are stored separately from metric records, and discovery freezes the accepted candidate set. Challenger trials preserve one experiment manifest containing the application revision, recorded shuffle seed, comparable cases, exact model identities, and per-run evidence. Accepted assessments are carried into frontier calculation rather than reconstructed from chat history. Published `openmerit@0.1.5` does not provide these handoffs.
33
+
34
+ In the unreleased source checkout, each `metric_window_available` signal replaces the stored aggregate window for that metric. Send a cumulative, application-run window with its sample count, explicit aggregation, and evidence reference; OpenMerit does not add separate or overlapping window counts. Challenger metric windows must also fall inside their assessment execution window. Core marks the baseline ready only when instrumentation is ready and every required metric has a finite measured value, the confirmed minimum sample count, and an evidence reference. The harness remains responsible for measuring the real application call and authenticating those references; a reported window is not independently re-run or verified by core.
35
+
36
+ ## Evidence references
37
+
38
+ Results point to durable evidence rather than embedding arbitrary traces. Sources include harness traces, evaluation artifacts, observability records, external sources, and user feedback. References may contain a URI, media type, and digest. Successful intents require evidence; frontier certification also requires calculation evidence and exact assessment IDs.
39
+
40
+ ## Readiness contract
41
+
42
+ Setup must return an explicit observability coverage report containing every required metric, which metrics are covered, which remain missing, and when the check occurred. OpenMerit derives readiness from those sets; a harness cannot label coverage `ready` while a required metric is absent.
43
+
44
+ If setup reports a required gap, OpenMerit issues `instrument_observability`. A successful instrumentation result is accepted only when every required metric is reported as covered. The unreleased source checkout explicitly asks Pi to exercise each collector and grader through the application path, inspect output values, and retain artifacts. OpenMerit validates the coverage report but does not independently run the collector or grader, so `ready` is still a harness claim rather than independently verified end-to-end proof. Request a reported production metric window before relying on it for model decisions.
45
+
46
+ Product instrumentation and OpenMerit auditing are separate. The confirmed application LLM target emits measurements such as quality, success, tokens, cost, latency, retries, tool calls, and intervention. Every metric record carries that target ID. OpenMerit records the signals, policy decision, intent lifecycle, evidence identifiers, and readiness or frontier decision that followed.
47
+
48
+ Pi's own assistant messages are orchestration data, not application evidence. Application provider cost, input and output tokens, and end-to-end task latency must come from the confirmed application call or its observability layer. Tool-use performance requires harness-trace evidence; code coverage alone is not evidence of tool-use quality. Quality and task-specific behavior can still come from verified project evaluators.
@@ -0,0 +1,40 @@
1
+ # Project state and operations
2
+
3
+ ## State location
4
+
5
+ Operational data lives under `.openmerit/` in the target product repository. It should be ignored by version control because it can contain local lifecycle state and evidence locations.
6
+
7
+ - `config.json` — task profiles and policy.
8
+ - `state.json` — current cycle, active or pending intent, last result, explicit observability coverage, processed signal IDs, task counters, latest aggregate baseline metric windows with evidence references, catalogue fingerprint, and assessment/swap checkpoints.
9
+ - `events.jsonl` — append-only audit events.
10
+ - `audit-export-errors.jsonl` — failures from optional external event exporters; these failures never replace or erase the canonical local event.
11
+ - `evidence/*.json` — evidence-reference manifests.
12
+ - `scheduler-status.json` — last headless wakeup time and exit code in the unreleased macOS source implementation, when a wakeup actually ran.
13
+ - `terminal.json` — optional source configuration and declared inventory for the unreleased Inference Terminal.
14
+ - `inference.jsonl` — optional request-metadata export supplied by your existing tools for the terminal. The terminal reads this file; it does not instrument your application or create the feed.
15
+
16
+ Snapshot writes are atomic. Files use owner-only permissions where POSIX modes are supported.
17
+
18
+ The [Inference Terminal](inference-terminal.md) runs independently of the orchestration lifecycle. Its collector caches are memory-only and stop when the command exits. Permissions and retention for imported source files remain the responsibility of the tool that writes them.
19
+
20
+ ## Diagnostics
21
+
22
+ Start with `/openmerit status`, which reports profile, metric and evaluation readiness, active work, next assessment, lifecycle, and pause state. `/openmerit doctor` confirms adapter loading without printing secrets.
23
+
24
+ In published version 0.1.5, `/openmerit pause` suppresses lifecycle follow-up nudges but does not gate the automation-signal dispatch path. Task, measurement, and session-start signals can still issue due work while status says paused. The setting is process-local, resets on restart, and does not cancel active work. Explicit commands still work. Do not rely on that version's pause command to stop all OpenMerit work.
25
+
26
+ In the unreleased source checkout, pause is stored in `.openmerit/state.json`, blocks automatic due decisions, and unloads an OpenMerit-managed macOS LaunchAgent for that project. Resume restores scheduling only if the confirmed policy still permits persistent execution and infrastructure changes. Incoming signals can still be recorded while paused, and an already-active intent is not cancelled. `/openmerit doctor` shows the LaunchAgent and status-file paths. A LaunchAgent runs only while the Mac user is logged in; a headless run requires Pi credentials available without a shell-only environment variable.
27
+
28
+ Inspect `scheduler-status.json` after a headless wakeup. A successful baseline-only check records a decision without starting Pi. A nonzero `exitCode` or `activeIntentId` means model-driven work did not complete; open the project in interactive Pi, inspect `/openmerit status` and `.openmerit/events.jsonl`, then use `/openmerit cancel` only if you intend to abandon that work. Model-driven runs have a 15-minute limit. The runner does not silently mark incomplete work successful.
29
+
30
+ `/openmerit logs` prints the exact local paths to the audit log, exporter failures, evidence manifests, and current state. During an end-to-end run, `tail -f .openmerit/events.jsonl | jq .` shows signals and decisions as they occur.
31
+
32
+ ## Audit and recovery
33
+
34
+ OpenMerit records initialization, intent requests and completions, automation signals and duplicates, due checks, task checkpoints, catalogue changes, frontier validation, confirmation boundaries, and lifecycle advancement.
35
+
36
+ Automation-signal records include safe application measurement summaries: application model ID and application-reported token/cost/latency fields for completed tasks; or metric ID, state, scope, direction, method, value, unit, sample count, interval/window, and evidence IDs for metric windows. Pi's own model usage is not recorded as product-task evidence. Intent completions record their kind, status, next stage, harness, summary, and evidence IDs/sources. Secret-shaped keys and common API-key patterns are redacted before local persistence or export.
37
+
38
+ The local JSONL file is the canonical source of truth. `ProjectStore` can also receive optional event exporters for OpenTelemetry or vendor-specific adapters. Export is best effort: a failed exporter is written to `audit-export-errors.jsonl` and does not interrupt lifecycle state or canonical logging.
39
+
40
+ If Pi stops during an intent, durable state retains it. Inspect `state.json` and `events.jsonl` before manual changes. Deleting `.openmerit/` discards the audit trail, budget, profile, and evidence references.
@@ -0,0 +1,76 @@
1
+ # OpenMerit Pareto conformance specification
2
+
3
+ Status: initial normative specification
4
+
5
+ OpenMerit owns this definition. A connected harness performs the calculation, returns a `FrontierSnapshot`, and supplies calculation evidence. OpenMerit verifies the returned result against these rules.
6
+
7
+ ## Comparable assessment set
8
+
9
+ Every assessment in one frontier calculation must share:
10
+
11
+ - task-profile ID and revision;
12
+ - product revision;
13
+ - harness identity and materially equivalent tool configuration;
14
+ - objective definitions, units, and directions;
15
+ - a representative evaluation workload and compatible evidence window.
16
+
17
+ The harness must not silently compare measurements from incompatible conditions.
18
+
19
+ ## Metric readiness
20
+
21
+ A required objective is ready only when:
22
+
23
+ - a metric record exists;
24
+ - its state is `measured`;
25
+ - it has a numeric value;
26
+ - its sample count meets the objective's minimum;
27
+ - its unit and direction are compatible across candidates.
28
+
29
+ Missing or immature evidence is `unresolved`. It is never converted to zero and never treated as success.
30
+
31
+ ## Feasibility
32
+
33
+ Hard constraints are evaluated conservatively against uncertainty intervals. If an interval is absent, the point value is both bounds.
34
+
35
+ - `at_least`: eligible only when the lower bound meets the threshold; ineligible when the upper bound is below it; otherwise unresolved.
36
+ - `at_most`: eligible only when the upper bound meets the threshold; ineligible when the lower bound exceeds it; otherwise unresolved.
37
+
38
+ Only eligible candidates may be certified as members of the frontier.
39
+
40
+ ## Dominance
41
+
42
+ For candidate A to dominate candidate B, A must be proven no worse than B on every objective and proven materially better on at least one objective.
43
+
44
+ For a maximize objective with tolerance `t`:
45
+
46
+ - A is proven no worse when `A.lower + t >= B.upper`.
47
+ - A is proven materially better when `A.lower > B.upper + t`.
48
+
49
+ For a minimize objective with tolerance `t`:
50
+
51
+ - A is proven no worse when `A.upper <= B.lower + t`.
52
+ - A is proven materially better when `A.upper + t < B.lower`.
53
+
54
+ If complete evidence proves A worse on any objective, A does not dominate B. If the evidence cannot establish no-worse or worse because intervals overlap, the comparison is unresolved.
55
+
56
+ ## Frontier
57
+
58
+ An eligible candidate is on the certified frontier when:
59
+
60
+ - no eligible candidate is proven to dominate it; and
61
+ - none of its required pairwise comparisons is unresolved.
62
+
63
+ Candidates with incomplete evidence, unresolved constraints, or unresolved comparisons are reported separately. A dominated, ineligible, or unresolved candidate must not appear on the certified frontier.
64
+
65
+ ## Required result evidence
66
+
67
+ Every frontier result must include:
68
+
69
+ - the assessment IDs used;
70
+ - all directed pairwise comparisons between eligible candidates;
71
+ - reasons for every comparison;
72
+ - separate frontier, dominated, ineligible, and unresolved sets;
73
+ - at least one durable calculation-evidence reference.
74
+
75
+ The calculation evidence should identify the harness execution or artifact that produced the result.
76
+
@@ -0,0 +1,54 @@
1
+ # Pi extension
2
+
3
+ ## Supported version
4
+
5
+ The adapter targets Pi 0.87 and declares `>=0.87.0 <0.88.0`. Development uses 0.87.0 without modifying a globally installed Pi.
6
+
7
+ ## Loading
8
+
9
+ Install the published package with `pi install npm:openmerit`. For development
10
+ from a checkout, load the source extension with:
11
+
12
+ ```sh
13
+ ./node_modules/.bin/pi --no-extensions -e ./packages/pi/src/index.ts
14
+ ```
15
+
16
+ The isolation flag avoids duplicate `openmerit_complete_intent` and `openmerit_report_signal` registrations when npm `openmerit` is also installed in Pi's user settings.
17
+
18
+ Use `/openmerit doctor` to confirm that the adapter and intent executor are available.
19
+
20
+ Use `/openmerit logs` to locate the project-local audit stream, exporter failures, evidence manifests, and state while testing an end-to-end run.
21
+
22
+ On `session_start`, an interactive unconfigured project receives a bounded source preflight. In the unreleased source checkout, the adapter also repeats that bounded detection after a settled Pi turn. This lets Pi build an application from an empty directory and then receive an automatic `establish_evals` intent without requiring `/openmerit setup`. Durable active-intent state prevents duplicate setup nudges. Non-interactive runs do not attempt a confirmation flow.
23
+
24
+ If Pi exits while an intent is active, the unreleased adapter restores its completion tool and resends that same persisted intent on the next interactive session so the harness can continue it.
25
+
26
+ For configured projects with a confirmed application target, session start reports the available model catalogue and a scheduled tick. Application instrumentation reports idempotent completed-task and metric-window signals. OpenMerit evaluates the confirmed policy before issuing work. Published `openmerit@0.1.5` has no persistent Pi scheduler. In the unreleased source checkout on macOS, Pi can install a per-project user LaunchAgent after you confirm persistent execution and authorize infrastructure changes in the policy. It uses the selected authenticated Pi model, checks for due work locally every five minutes, and starts a headless Pi only when the due core check finds downstream harness work. If installation or background authentication fails, the adapter reports a scheduling gap. Other platforms still need an external scheduler.
27
+
28
+ The macOS background runner is not yet product-validated for candidate work. A local runner test verifies that an insufficient baseline is checked and recorded without starting Pi. An earlier live OpenRouter test reached the former due `run_assessment` intent and called Pi tools, but did not obtain a terminal result within the test window. A model-driven runner is capped at 15 minutes and records a nonzero exit and active intent ID on failure. Do not infer that unattended candidate discovery or swapping has passed end-to-end validation.
29
+
30
+ The adapter does not convert Pi's own provider usage into application metrics. Application-task signals must carry the application model, target ID, and durable evidence; application telemetry may include provider-reported USD cost, token counts, and wall-clock latency. If automatic setup is active when interactive input arrives, the input is queued and released after setup completes instead of racing the setup turn.
31
+
32
+ Pi also exposes `openmerit_report_signal` for verified `product_task_completed`, `metric_window_available`, and `verification_window_completed` events. Evals or observability work performed by Pi can submit measured application-target metric records through this typed ingress using stable signal IDs. For the unreleased core check, send cumulative metric windows with sample counts and evidence references. The background wakeup only checks whether a time-based action is due; it does not collect application results by itself. `/openmerit assess` checks those stored windows locally and starts candidate discovery only when all required metrics are ready.
33
+
34
+ ## Intent delivery
35
+
36
+ The extension persists each intent as a custom Pi session entry, sends private extension-authored context with `display: false`, and triggers a follow-up Pi turn. The intent describes the outcome, constraints, evidence, task profile, and authorization boundary without prescribing Pi's method.
37
+
38
+ Pi must finish by calling `openmerit_complete_intent` exactly once. Success without evidence is rejected. Setup, assessment, and frontier intents receive additional structured-output checks. During setup, Pi first presents its inferred metrics and sampling plan to the user. The accepted task profile records how many representative cases and repetitions each metric needs and how results are aggregated; sample counts are kept separate from optional performance thresholds.
39
+
40
+ If Pi exits while an intent is active, the next session reports the recovered intent. Use `/openmerit cancel` to record an explicit cancellation before retrying; OpenMerit does not silently reinterpret interrupted work as success.
41
+
42
+ The extension replaces the completion tool definition whenever an intent becomes active. Its `outputs` property therefore exposes the canonical schema for that exact intent instead of untyped JSON. Pi validates the arguments before execution and requests provider-side JSON Schema constrained sampling when supported. OpenMerit validates the canonical schema and semantic rules again when accepting the result. The source checkout records the exact incumbent during setup, freezes accepted candidates after discovery, and carries accepted assessments and the experiment manifest into frontier calculation so these transitions do not rely on conversational memory.
43
+
44
+ ## Model changes
45
+
46
+ The extension never treats Pi's model as the application target. After policy authorization, it sends a bounded, target-bound `apply_model_swap` intent. Pi performs the application-route change and returns configuration proof; the target ID must match the confirmed application LLM target, and Pi's own model must remain unchanged.
47
+
48
+ ## Jev System One
49
+
50
+ Jev is not integrated into OpenMerit. If configured in Pi, the intent asks Pi to use harness-configured fast decision resources aggressively for bounded ranking, triage, and judgment. Credentials remain outside OpenMerit.
51
+
52
+ ## Non-interactive modes
53
+
54
+ A harness must not attempt interactive confirmation in a mode without dialog-capable UI. It should return a waiting-for-user or failed result with an actionable reason.
@@ -0,0 +1,28 @@
1
+ # Current limitations and roadmap
2
+
3
+ ## Current status
4
+
5
+ OpenMerit is ready for an initial supervised product trial with Pi. Its protocol, lifecycle, persistence, metric catalogue, frontier verifier, and Pi adapter are implemented and tested. That is not evidence that it optimizes every product; each target workload requires its own profile and representative evaluations.
6
+
7
+ ## Known limitations
8
+
9
+ - No second harness adapter exists.
10
+ - Jev must be configured and invoked by the harness; OpenMerit intentionally provides no client.
11
+ - Generic Pi task checkpoints do not replace product telemetry. Setup must create the infrastructure required by the profile.
12
+ - Published 0.1.5 reports active-session signals and typed metric windows but cannot provision persistent background wakeups. The unreleased source checkout adds an opt-in macOS user LaunchAgent; non-macOS platforms still need an external scheduler.
13
+ - The source LaunchAgent now performs a deterministic core baseline check without Pi when evidence is insufficient. This has a local runner test, not a representative product trial. Earlier headless Pi assessment attempts timed out; unattended candidate discovery and model decisions are not yet validated.
14
+ - The source checkout can detect an application call created during a settled interactive Pi build turn, records the incumbent, and durably hands candidates and assessments between comparison stages. These paths have deterministic adapter tests but still need the representative live sandbox trial below.
15
+ - In published 0.1.5, pause does not block automatic dispatch from task, measurement, or session-start signals and resets on process restart. The unreleased source checkout persists pause and gates automatic work; neither version cancels already-active work.
16
+ - External evidence authenticity depends on the referenced system.
17
+ - The core and protocol workspaces are private; the single `openmerit` npm
18
+ package exposes their public entry points.
19
+
20
+ ## Next validation milestones
21
+
22
+ 1. Run setup in one representative product.
23
+ 2. Verify generated eval and observability artifacts and their side effects.
24
+ 3. Collect repeated evidence for the confirmed profile.
25
+ 4. Complete a budgeted A/B/C challenger cycle.
26
+ 5. Validate the harness-calculated frontier and recommendation.
27
+ 6. Perform a supervised swap, post-swap verification, and rollback drill.
28
+ 7. Repeat with a materially different product task before broad readiness claims.
@@ -0,0 +1,37 @@
1
+ # Security
2
+
3
+ ## Credential ownership
4
+
5
+ OpenMerit accepts no model-provider or Jev credential. Credentials belong to the coding harness, its environment, or an external secret store.
6
+
7
+ Never place secrets in profiles, policies, intent constraints, summaries, evidence URIs, audit events, fixtures, traces, or committed environment files.
8
+
9
+ ## Mutation authorization
10
+
11
+ Most intents carry `modelMutationAllowed: false`. Only application-target apply-swap and rollback intents can authorize a mutation, limited to the relevant proposal, prior configuration, and confirmed target ID. Pi's model is explicitly excluded. Automatic swaps require explicit policy opt-in and meeting the configured verified-swap threshold, which can be zero. Otherwise a proposal requires approval. Post-swap evaluation depends on the harness performing the checks and reporting their results. OpenMerit is not a sandbox around the harness's tools.
12
+
13
+ Persistent wakeup provisioning is a separate authorization boundary. An adapter that needs to install a system scheduler, create a paid cloud agent, or change infrastructure must declare `wakeupProvisioning: "infrastructure_change"`; OpenMerit will not request that provisioning unless the confirmed check policy sets `infrastructureChangesAllowed: true`.
14
+
15
+ The unreleased macOS Pi adapter installs a user LaunchAgent only under that authorization and persistent execution mode. The plist stores the project path, selected Pi model identifier, and executable paths, but no API key. It verifies that Pi authentication works without shell-only credentials before installing. The agent launches Pi headlessly only for due elapsed-time checks, so the configured model provider may receive project context and incur charges under the confirmed policy. Pause unloads the agent; already-running work is not cancelled.
16
+
17
+ The source runner imposes a 15-minute wall-clock limit and records a nonzero status if Pi fails to complete a due intent. This is a safety bound, not a provider-spend cap. Headless completion remains unverified in the current live test.
18
+
19
+ ## Local state
20
+
21
+ `.openmerit/` may reveal task names, candidate identifiers, evidence locations, budgets, or model choices. Treat it as sensitive project metadata even though it must not contain credentials.
22
+
23
+ Pi also records OpenMerit requests and results in its session history. Project context and evaluation data may be sent to the model providers, tools, and evidence systems configured in the harness. Local OpenMerit storage does not imply that all processing stays on your machine.
24
+
25
+ ## Evidence integrity
26
+
27
+ OpenMerit validates structure and deterministic Pareto conformance, not the truth of arbitrary external evidence. Use authenticated observability, immutable artifacts, and digests where stronger provenance is required.
28
+
29
+ Do not include credentials, private traces, or proprietary evaluation data in public issues.
30
+
31
+ ## Inference Terminal (preview)
32
+
33
+ The source-only terminal reads selected project metadata, bounded loopback runtime inventory/metrics endpoints, and whole-machine CPU and memory. It calls no inference endpoint and exposes no mutation API. Default discovery reads only the project's `.env`, `.env.local`, `package.json`, and `requirements.txt`, with per-file bounds and project containment after symlink resolution. It never executes those files or expands expressions. Known credential-variable presence can identify a configured provider; credential contents are neither retained, returned, nor used for provider calls. Only safe loopback origins and recognized service identity metadata survive projection.
34
+
35
+ Coding-harness histories are outside the terminal’s application scope and are not scanned or imported. The retired Pi history reader and access prompts are removed; legacy grants in `~/.openmerit/terminal-access.json` are never consulted. Application request logs and explicit exports retain only allowlisted metadata. Sources must belong to the selected application; the terminal cannot infer the purpose of an unlabeled API response. Records explicitly marked `scope: "development"` are rejected.
36
+
37
+ The local web service binds to `127.0.0.1` and rejects unexpected Host, Origin, and cross-site browser requests. It only serves bundled assets and specific read-only JSON routes. It is not authenticated against other local processes or machine users. Names and identifiers in imported metadata must still be kept free of secrets. Usage caches disappear when the process exits. Legacy history permission files are left untouched but have no effect. No cloud usage APIs, credential stores, arbitrary source-tree indexing, process inspection, or general home-directory scans are performed. See [terminal data handling and limits](inference-terminal.md) for supported sources, precedence, retention bounds, and sample-data separation.
@@ -0,0 +1,23 @@
1
+ # OpenMerit workshop — two-ink linocut
2
+
3
+ Final asset: `site/assets/workshop-linocut.webp`.
4
+
5
+ Dimensions: exactly 2172 × 724 pixels (3:1), unchanged from the previous artwork. The existing hero size and responsive crop are unchanged.
6
+
7
+ Created using the built-in imagegen tool, with `site/assets/workshop.webp` as the edit target and composition reference. Exported to WebP with cwebp at quality 88. The previous illustration is retained at its original path.
8
+
9
+ Generated source: `/Users/lazim/.codex/generated_images/01a0c9d4-a3f3-7e12-aa03-2b46f42b5e83/exec-208f9a64-a1d4-4ae9-95e9-065bbd509298.png`.
10
+
11
+ ## Exact edit prompt
12
+
13
+ > Use case: style-transfer
14
+ > Asset type: replacement hero illustration for the OpenMerit website.
15
+ > Input image 1: edit target and strict composition reference. Preserve its context, scene, framing, aspect ratio, and dimensions. Output exactly 2172 × 724 pixels, a 3:1 horizontal panorama.
16
+ >
17
+ > Reinterpret this entire illustration as an original two-ink linocut from a small independent workshop's technical field guide. This is a substantial change in drawing language, not a color filter on the existing watercolor. Use carved charcoal-black shapes with deliberate white gouge marks, sharply observed but simplified human faces and hands, angular planes of clothing, and coarse directional cut marks that describe the material. One muted iron-oxide brick-red spot ink (#a65037) appears selectively on aprons, the tool box, and small mechanical parts. Black ink is warm charcoal (#35372f), never glossy. The black printing plate carries the drawing; the red plate occasionally sits a hair out of register. Restrained ink bite and broken edges ONLY inside the printed objects, never an all-over grunge filter. Forms should feel specifically cut from a block, with crisp cut silhouettes and varied negative-space marks. An idiosyncratic craftsman's manual, not a decorative retro stock illustration.
18
+ >
19
+ > STRICT INVARIANTS: Same six adult craftspeople, same left-to-right order, same varied poses and practical activity, same central long workbench, same hanging balance scale, same notebook, measuring calipers, bench mechanism, plants and tools, same curled sleeping dog under the bench. Far-left person carries a tool box; next person measures an object; middle seated older craftsperson examines a tool; next person writes observations; next person adjusts the mechanism; far-right person walks away with a tool. Keep all six entire figures and all feet in frame. Preserve normal human anatomy and believable tool handling. Keep the current relative scale and placement of the figures, with the scene occupying the same lower part of the composition and extensive clean negative space above. Keep the ground line at the same height. No extra people or objects.
20
+ >
21
+ > BACKGROUND: flat clean warm off-white #f8f7f4 to blend seamlessly with the webpage. No texture in blank areas, no visible paper rectangle, no frame, no border, no vignette. The ground remains a thin, irregular print line and a few material marks, not a heavy black plinth.
22
+ > Mood: attentive, patient, utilitarian, observant. A real workshop with individual people, not heroic workers, caricatures, or decorative mascots.
23
+ > No text, no labels, no lettering, no logos, no watermark. No watercolor, soft pencil shading, etching-like fine gray crosshatching, gradients, soft airbrush, 3D, cute rounded corporate characters, vector blobs, generic Bauhaus geometry, or photorealism. The result must read as a confident, spare, distinctly carved two-color print at website size.
@@ -0,0 +1,28 @@
1
+ # OpenMerit miniature workshop — diverse cast
2
+
3
+ Final asset: `site/assets/workshop-miniature-diverse.webp`.
4
+
5
+ Dimensions: 2172 × 724 pixels (3:1), unchanged. Exported with cwebp at quality 89 without resizing or cropping.
6
+
7
+ Created with the built-in imagegen tool as a targeted edit of `site/assets/workshop-miniature.webp`. The six fictional makers have varied racial appearances; their roles, composition, workwear, workshop objects, and photographic miniature aesthetic are preserved.
8
+
9
+ Generated source: `/Users/lazim/.codex/generated_images/01a0c9d4-a3f3-7e12-aa03-2b46f42b5e83/exec-c120bc1a-6cea-4824-a27a-81b382be4063.png`.
10
+
11
+ ## Exact edit prompt
12
+
13
+ > Use case: precise-object-edit.
14
+ > Input image 1 is the edit target: a photographic miniature workshop.
15
+ > User request: make the six fictional adult makers racially diverse.
16
+ >
17
+ > Change ONLY the six people's racial appearances, their facial features, skin tones, and hair where needed. Preserve the same miniature figures' genders, approximate ages, positions, poses, activities, outfits and proportions. Render facial features respectfully and individually, never as exaggerated racial caricatures.
18
+ > From left to right:
19
+ > 1. The woman carrying the tool box: a Black woman with dark brown skin and natural tightly curled hair gathered into a practical bun.
20
+ > 2. The young man using calipers: an East Asian man with light-medium skin and short, slightly tousled straight black hair.
21
+ > 3. The seated older man inspecting tools: a South Asian man with medium-brown skin, gray hair and beard, retaining his glasses and ochre vest.
22
+ > 4. The woman writing in the notebook: a white woman with fair skin and auburn hair tied in a bun.
23
+ > 5. The man using the bench vise: a Latino man with warm medium-brown skin, keeping his cap and work clothes.
24
+ > 6. The bearded man carrying the hand plane at the far right: a Black man with deep brown skin, short natural coiled hair and a neatly shaped beard.
25
+ >
26
+ > Match hands, necks and all exposed skin to each person's face. Keep realistic adult anatomy and the tactile miniature-set materials. All are ordinary skilled colleagues equally engaged in the work; do not add identity symbols or stereotyped costumes.
27
+ >
28
+ > STRICTLY preserve EVERYTHING else: exact 2172 × 724 output dimensions, 3:1 aspect ratio, all six whole figures, their positions and actions, the long wooden workbenches, every tool and container, calipers, balance scale, notebook, vise, houseplant, sleeping dog, linen clothing textures, photographic lighting, contact shadows, warm off-white background, empty space above, the camera angle and framing. No scene redesign, new objects, labels, lettering, text or watermark. This is a targeted cast-diversity edit of the current photographic miniature artwork, not a new illustration style.
@@ -0,0 +1,26 @@
1
+ # OpenMerit workshop — photographic miniatures
2
+
3
+ Final asset: `site/assets/workshop-miniature.webp`.
4
+
5
+ Export dimensions: exactly 2172 × 724 pixels (3:1). The imagegen source was 2171 × 724; cwebp normalized it to the requested dimensions during WebP export at quality 89. No scene cropping or site layout change was applied.
6
+
7
+ Generated with the built-in imagegen tool. This is a generated image with a photographic miniature-set aesthetic, not a photograph of an actual fabricated set. The linocut was supplied only to retain the workshop narrative. Prior illustrations remain available in `site/assets/`.
8
+
9
+ Generated source: `/Users/lazim/.codex/generated_images/01a0c9d4-a3f3-7e12-aa03-2b46f42b5e83/exec-c6b76303-e93e-4e82-b255-d982c78207ab.png`.
10
+
11
+ ## Exact generation prompt
12
+
13
+ > Use case: style-transfer with complete medium and composition redesign.
14
+ > Asset: OpenMerit website panoramic hero. Output exactly 2172 × 724 pixels, 3:1 aspect ratio.
15
+ > Input image 1 is ONLY a reference for the narrative: craftspeople evaluate, measure, compare, and select tools. It is NOT a style reference. The previous attempt was rejected for remaining too similar. Discard ALL line illustration, outlines, etching, woodcut, gray crosshatching, drawing marks, and the flat same-pose lineup. Build a radically different image.
16
+ >
17
+ > Create a convincing editorial studio PHOTOGRAPH of a real hand-built miniature workshop, an intricate small-scale mixed-media assemblage. Physical set, miniature wooden worktables, articulated human figures about 10 cm tall, photographed with a 70 mm macro lens and enough depth of field to see every activity. The aesthetic is the considered, slightly eccentric model-making found in an independent industrial-design exhibition: tactile, intellectual, warm, observant. Not a toy advertisement or a rendered startup illustration.
18
+ >
19
+ > Six adult-proportioned miniature makers with distinctly carved pale-wood faces and hands, visible wood grain and plane marks, tiny wire-jointed limbs, and carefully stitched linen work clothes. Very small carved noses and painted facial details; ordinary adult proportions, natural working poses, no oversized heads, no cute faces, no smiling at the viewer. Wardrobe: washed indigo jackets, raw oatmeal linen aprons, one ochre knit vest. Tiny brass measuring instruments, red-oxide painted vise, a blackened-steel hand plane. Show real fibers, stitching, tool marks, miniature fasteners and occasional delicate wire joints. Materials should look touched and handled. A believable photograph of something assembled by hand, not smooth plastic or glossy CGI.
20
+ >
21
+ > NEW COMPOSITION: a gently staggered workbench arrangement with shallow three-dimensional depth, viewed from a low three-quarter angle slightly above bench height. The work flows from left to right, but people stand on BOTH sides of the tables and face different directions, with varied standing, leaning and seated poses, unlike the source's flat frieze. Keep it as one cohesive narrow panorama, not separate islands or panels. Left: a figure setting a tool tray onto a short bench and another leaning close to measure a small part with calipers. Center: a seated older maker inspecting several alternate tools while another person writes observations in a folded paper notebook; a tiny brass balance scale stands on the central bench. Right: one figure works at the vise while another carries the selected hand plane away. Small wooden tool boxes, narrow rulers and material offcuts provide sparse, precise details. A small sleeping dog is made from folded felt under the central bench. Human craft and careful evaluation remain the story.
22
+ >
23
+ > LAYOUT: the miniature set spans most of the width with roughly 6% side margins, all six whole figures and every head and foot fully visible. Scene sits in the LOWER 58% of the image. The UPPER 40% is completely empty off-white. All objects rest directly on a seamless matte warm-white studio surface (#f8f7f4). No plinth, stage, wooden groundboard, enclosing wall, skyline, building, colored background field or hard horizon. Subtle real contact shadows under feet and tables; gentle raking window light from upper left reveals physical material texture. The background stays clean and almost perfectly uniform to blend into a pale website. No dark vignette, dramatic cinematic lighting or artificial bokeh.
24
+ >
25
+ > This must be immediately unmistakable as a PHOTOGRAPHIC miniature material assemblage rather than any form of drawing. Preserve the workshop meaning and the exact dimensions, but do NOT preserve old poses, outlines, colors, or illustration style.
26
+ > No words, labels, captions, lettering, logo, watermark, robots, digital interfaces, charts, gradients, floating geometric shapes, cartoon clay, toy-like rounded bodies, isometric SaaS art, vector art, generically cheerful stock illustration, or line drawing.