openmerit 0.1.4 → 0.1.6-preview.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (147) hide show
  1. package/CHANGELOG.md +40 -0
  2. package/README.md +121 -386
  3. package/dist/core/src/index.d.ts +101 -0
  4. package/dist/core/src/index.js +1649 -0
  5. package/dist/core/src/store.d.ts +35 -0
  6. package/dist/core/src/store.js +102 -0
  7. package/dist/pi/src/index.d.ts +32 -0
  8. package/dist/pi/src/index.js +794 -0
  9. package/dist/pi/src/scheduler.d.ts +11 -0
  10. package/dist/pi/src/scheduler.js +137 -0
  11. package/dist/pi/src/wakeup.d.ts +2 -0
  12. package/dist/pi/src/wakeup.js +108 -0
  13. package/dist/protocol/src/index.d.ts +484 -0
  14. package/dist/protocol/src/index.js +47 -0
  15. package/dist/protocol/src/schemas.d.ts +576 -0
  16. package/dist/protocol/src/schemas.js +280 -0
  17. package/dist/terminal/public/app.js +297 -0
  18. package/dist/terminal/public/brands/anthropic.png +0 -0
  19. package/dist/terminal/public/brands/baai.png +0 -0
  20. package/dist/terminal/public/brands/baseten.png +0 -0
  21. package/dist/terminal/public/brands/cerebras.png +0 -0
  22. package/dist/terminal/public/brands/cohere.png +0 -0
  23. package/dist/terminal/public/brands/deepseek.ico +0 -0
  24. package/dist/terminal/public/brands/google.png +0 -0
  25. package/dist/terminal/public/brands/groq.ico +0 -0
  26. package/dist/terminal/public/brands/lm-studio.png +0 -0
  27. package/dist/terminal/public/brands/meta.ico +0 -0
  28. package/dist/terminal/public/brands/mistral.png +0 -0
  29. package/dist/terminal/public/brands/nomic.png +0 -0
  30. package/dist/terminal/public/brands/ollama.png +0 -0
  31. package/dist/terminal/public/brands/openai.png +0 -0
  32. package/dist/terminal/public/brands/openrouter.png +0 -0
  33. package/dist/terminal/public/brands/qwen.png +0 -0
  34. package/dist/terminal/public/brands/vllm.ico +0 -0
  35. package/dist/terminal/public/brands/vllm.png +0 -0
  36. package/dist/terminal/public/favicon.svg +1 -0
  37. package/dist/terminal/public/flow.css +1 -0
  38. package/dist/terminal/public/flow.js +770 -0
  39. package/dist/terminal/public/index.html +21 -0
  40. package/dist/terminal/public/styles.css +779 -0
  41. package/dist/terminal/src/activity-merge.mjs +64 -0
  42. package/dist/terminal/src/browser.mjs +29 -0
  43. package/dist/terminal/src/cli.mjs +60 -0
  44. package/dist/terminal/src/collect.mjs +311 -0
  45. package/dist/terminal/src/discovery.mjs +93 -0
  46. package/dist/terminal/src/hardware.mjs +57 -0
  47. package/dist/terminal/src/project-activity.mjs +156 -0
  48. package/dist/terminal/src/sample.mjs +171 -0
  49. package/dist/terminal/src/server.mjs +56 -0
  50. package/dist/terminal/src/services.mjs +62 -0
  51. package/dist/terminal/src/topology.mjs +30 -0
  52. package/docs/adapter-guide.md +189 -0
  53. package/docs/architecture.md +59 -0
  54. package/docs/automation.md +74 -0
  55. package/docs/budgets.md +37 -0
  56. package/docs/commands.md +85 -0
  57. package/docs/demo-backfill.md +29 -0
  58. package/docs/demo-fieldkit.md +47 -0
  59. package/docs/demo-placement.md +30 -0
  60. package/docs/demo-spam.md +15 -0
  61. package/docs/demo-support.md +42 -0
  62. package/docs/demo.md +57 -0
  63. package/docs/first-trial.md +60 -0
  64. package/docs/getting-started.md +65 -0
  65. package/docs/index.md +40 -0
  66. package/docs/inference-terminal.md +439 -0
  67. package/docs/lifecycle.md +30 -0
  68. package/docs/memo.md +126 -0
  69. package/docs/metrics-and-evidence.md +48 -0
  70. package/docs/operations.md +40 -0
  71. package/docs/pareto-spec.md +76 -0
  72. package/docs/pi-extension.md +54 -0
  73. package/docs/roadmap.md +28 -0
  74. package/docs/security.md +37 -0
  75. package/docs/site-artwork-linocut.md +23 -0
  76. package/docs/site-artwork-miniature-diverse.md +28 -0
  77. package/docs/site-artwork-miniature.md +26 -0
  78. package/docs/site-demo.md +177 -0
  79. package/docs/site-design.md +94 -0
  80. package/docs/site-documentation.md +83 -0
  81. package/docs/site-dynamic-og.md +35 -0
  82. package/docs/site-faq-maintenance.md +115 -0
  83. package/docs/site-hero-resolution.md +60 -0
  84. package/docs/site-illustration-sequences.md +227 -0
  85. package/docs/site-inference-terminal.md +203 -0
  86. package/docs/site-memo.md +39 -0
  87. package/docs/site-og-image.md +38 -0
  88. package/docs/site-og-workshop.md +21 -0
  89. package/docs/site-section-artwork.md +56 -0
  90. package/docs/site-skill-review.md +57 -0
  91. package/docs/site-terminal-preview.md +85 -0
  92. package/docs/testing.md +118 -0
  93. package/docs/troubleshooting.md +55 -0
  94. package/docs/ux-reference.md +32 -0
  95. package/package.json +74 -42
  96. package/benchmark/invoice_ocr/data/invoice_01_ground_truth.json +0 -38
  97. package/benchmark/invoice_ocr/data/invoice_01_row_2.jpg +0 -0
  98. package/benchmark/invoice_ocr/data/invoice_02_ground_truth.json +0 -32
  99. package/benchmark/invoice_ocr/data/invoice_02_row_5.jpg +0 -0
  100. package/benchmark/invoice_ocr/data/invoice_03_ground_truth.json +0 -26
  101. package/benchmark/invoice_ocr/data/invoice_03_row_6.jpg +0 -0
  102. package/benchmark/invoice_ocr/data/invoice_04_ground_truth.json +0 -26
  103. package/benchmark/invoice_ocr/data/invoice_04_row_7.jpg +0 -0
  104. package/benchmark/invoice_ocr/data/invoice_05_ground_truth.json +0 -38
  105. package/benchmark/invoice_ocr/data/invoice_05_row_947.jpg +0 -0
  106. package/benchmark/invoice_ocr/data/invoice_06_ground_truth.json +0 -38
  107. package/benchmark/invoice_ocr/data/invoice_06_row_948.jpg +0 -0
  108. package/benchmark/invoice_ocr/data/invoice_07_ground_truth.json +0 -20
  109. package/benchmark/invoice_ocr/data/invoice_07_row_949.jpg +0 -0
  110. package/benchmark/invoice_ocr/data/invoice_08_ground_truth.json +0 -38
  111. package/benchmark/invoice_ocr/data/invoice_08_row_1888.jpg +0 -0
  112. package/benchmark/invoice_ocr/data/invoice_09_ground_truth.json +0 -26
  113. package/benchmark/invoice_ocr/data/invoice_09_row_1890.jpg +0 -0
  114. package/benchmark/invoice_ocr/data/invoice_10_ground_truth.json +0 -20
  115. package/benchmark/invoice_ocr/data/invoice_10_row_1892.jpg +0 -0
  116. package/benchmark/invoice_ocr/data/manifest.json +0 -97
  117. package/dist/benchmarks.js +0 -98
  118. package/dist/catalog.js +0 -61
  119. package/dist/cli.js +0 -188
  120. package/dist/daemon.js +0 -407
  121. package/dist/diagnostics.js +0 -227
  122. package/dist/frontier.js +0 -56
  123. package/dist/harness.js +0 -1
  124. package/dist/integrations.js +0 -19
  125. package/dist/invoice-eval.js +0 -33
  126. package/dist/invoice-score.js +0 -124
  127. package/dist/judge.js +0 -43
  128. package/dist/llm.js +0 -207
  129. package/dist/pi-config.js +0 -46
  130. package/dist/pi-trials.js +0 -373
  131. package/dist/policy.js +0 -185
  132. package/dist/providers.js +0 -1
  133. package/dist/recommend.js +0 -76
  134. package/dist/routes.js +0 -74
  135. package/dist/standalone.js +0 -224
  136. package/dist/store.js +0 -89
  137. package/dist/strategist.js +0 -68
  138. package/dist/task-input.js +0 -54
  139. package/dist/traces.js +0 -127
  140. package/dist/trials.js +0 -140
  141. package/dist/types.js +0 -2
  142. package/examples/invoice-prompt.txt +0 -19
  143. package/examples/task.example.json +0 -7
  144. package/extension/openmerit.ts +0 -947
  145. package/instructions/OPENMERIT.md +0 -63
  146. package/instructions/openmerit.policy.json +0 -37
  147. package/rules.md +0 -43
@@ -0,0 +1,439 @@
1
+ # Inference Terminal
2
+
3
+ The Inference Terminal is available as a test preview and in the source checkout; it is not in the stable npm release. It gives people and agents a local overview of an application’s models, providers, connections, inference activity, and compute. The selected repository or application path defines its scope. A coding agent can launch or read the terminal, but the coding harness’s own requests, tokens, and costs are outside that scope. Published `openmerit@0.1.5` does not include this command.
4
+
5
+ The terminal reads existing metadata independently of inference. It does not call a model, change a route, run evaluations, load or unload models, or store prompts and responses. Closing it does not stop your workloads. Collection consumes a small amount of local CPU, memory, and runtime API capacity; it is not zero-overhead monitoring.
6
+
7
+ ## Share the sample
8
+
9
+ The [hosted sample preview](https://dash.openmerit.site/) lets you explore the fictional enterprise environment in a browser without installing the terminal. It includes 40 agents, 140 tasks, and 21 models, with the same tabs and inspectors as the local interface. All identities, requests, costs, and hardware readings are illustrative. Shared links use the title **Inference Terminal by OpenMerit** and a dedicated preview image.
10
+
11
+ The preview is accessible to anyone with the link and asks search engines not to index it. It does not connect to your machine, accept project uploads, make inference calls, or expose the local snapshot API. Its read-only `/api/sample` endpoint returns fictional data. Real environment collection runs locally through the test install or source-build commands below; the hosted page never runs collectors. Hosted sample refreshes allow up to 30 seconds for the large fictional response on slower connections; local snapshot requests retain their 5-second limit. Only one refresh request runs at a time, even when a transfer takes longer than the normal polling interval.
12
+
13
+ ## Install the test preview
14
+
15
+ The terminal is included in `openmerit@0.1.6-preview.0`, under npm’s `preview` tag. Requires **Node.js 22.19 or newer**, with npm available in your terminal. From your application's directory:
16
+
17
+ ```sh
18
+ cd /path/to/your/project
19
+ npm install -g openmerit@preview
20
+ openmerit dash
21
+ ```
22
+
23
+ The install uses the existing `openmerit` npm package, with the terminal bundled alongside its coordinator, protocol, and Pi adapter. It downloads a prebuilt test version. No Git checkout, build step, API key, or separate audit command is needed. Installation is global for your Node environment: it does not add a dependency, lockfile, or configuration to the application directory. It replaces any globally installed `openmerit` in the same npm prefix. This preview is separate from the stable package on npm; `npm install -g openmerit` still installs the stable release without `dash`.
24
+
25
+ `openmerit dash` reads the current directory, starts a server bound to `127.0.0.1`, and opens its URL in your default browser after the first snapshot is ready. Port 4310 is preferred; when occupied, the operating system chooses a free local port. The actual URL is always printed. macOS, Windows, and Linux desktop opening use the platform's default URL handler. SSH and headless Linux sessions print the URL instead; use a private SSH port forward to view a remote instance. A browser-launch failure does not stop collection.
26
+
27
+ Run the same install command to update the preview. Remove it with `npm uninstall -g openmerit`, or restore the stable package with `npm install -g openmerit@0.1.5`. Registry preview builds have a `-preview.` version suffix; `npm ls -g openmerit --depth=0` shows which one is installed. `npm view openmerit dist-tags` shows the published preview and stable versions. The earlier hosted `-dash.` test build can be replaced with the same npm install command with the registry preview.
28
+
29
+ If npm reports `EACCES`, use a user-owned Node installation/npm prefix rather than installing with `sudo`. If `openmerit` is not found after a successful install, reopen the terminal and check that your npm global executable directory is on `PATH`.
30
+
31
+ The page uses dark mode only. Keep the terminal command running while using the dashboard. No extra operating-system permission prompt is expected for readable project files and loopback runtime APIs. Files that your user cannot read are reported as unavailable. Supported request logs populate activity automatically; a repository with only model configuration can show its services but cannot supply missing request history.
32
+
33
+ ## Run from a source checkout
34
+
35
+ With dependencies installed in the OpenMerit checkout:
36
+
37
+ ```sh
38
+ npm run dash
39
+ npm run dash -- /path/to/your/project
40
+ ```
41
+
42
+ `npm run dash` bundles the interface before starting. `npm run build` includes the executable and local web assets under `dist/terminal/`. Torph and Liveline are bundled locally. To install this source build for shell use, run `npm install -g .` from the checkout.
43
+
44
+ The shell command also accepts a path and explicit settings:
45
+
46
+ ```sh
47
+ openmerit dash /path/to/your/project --port 4311 --interval 20
48
+ openmerit dash --no-open
49
+ ```
50
+
51
+ | Option | Behavior |
52
+ | --- | --- |
53
+ | `[path]` | Project directory; defaults to the current working directory |
54
+ | `--port <number>` | Require this loopback port; fail if it is occupied. Without this option, use 4310 or a free port |
55
+ | `--interval <seconds>` | Collection interval from 10 to 30 seconds; defaults to 15 |
56
+ | `--no-discovery` | Skip project-file and request-log discovery and additional runtime discovery for this run; existing OS, Ollama, declared sources, and process-environment metadata remain available |
57
+ | `--no-open` | Start collection and print the URL without launching a browser |
58
+ | `--json` | Print one snapshot and exit without starting an HTTP server |
59
+ | `--demo` | Use an explicitly fictional environment; no local collectors run |
60
+ | `--help` | Show usage |
61
+
62
+ Use Ctrl+C to close the terminal. There is no persistent daemon or background scheduler. Collection runs only while the command is running. The terminal is separate from Pi's `/openmerit pause` and other orchestration commands.
63
+
64
+ ## Automatic application discovery
65
+
66
+ Discovery runs as part of `openmerit dash`; no separate audit command or exporter is needed for supported sources. The first snapshot checks the selected project and known loopback runtime addresses. Project declarations and previously unresponsive addresses are checked again every 60 seconds. Responding runtimes, OS readings, project request logs, and existing exports refresh at the normal 10–30-second collection interval.
67
+
68
+ Configuration discovery reads only four fixed files: `.env`, `.env.local`, `package.json`, and `requirements.txt`, each limited to 128 KiB and required to resolve to a regular file inside the project. It does not recursively index source code, execute project code, source shell files, or expand environment expressions. Supported SDK declarations are shown as **Detected** source hints; an installed SDK is not proof of a service connection or request. Supported literal `.env` declarations provide the same service identity evidence as inherited environment variables. Each file is independent evidence; discovery does not pretend to know an application's loading order. Credential contents and unrelated settings are discarded before metadata is retained or returned.
69
+
70
+ Alongside Ollama's existing default address, discovery checks `http://127.0.0.1:1234` and `http://127.0.0.1:8000`, plus safe HTTP loopback origins in the supported endpoint variables listed below. Endpoint paths must be empty or `/v1`; credentials, queries, fragments, and non-loopback hosts are rejected. At most eight additional origins are checked with read-only `GET /v1/models` and `GET /metrics`, without authentication or redirects. Unknown or nonresponding candidates do not become service rows. Configured origins are not probed twice. A compatible catalog supplies model availability; recognized vLLM counters can also identify vLLM. A port number alone never establishes LM Studio or vLLM identity, so an otherwise unidentified compatible runtime is named by its loopback origin. This is bounded discovery on this machine, not network or remote-host scanning.
71
+
72
+ **Sources** shows automatic discovery status and per-source coverage. Configuration, responding inventory APIs, and recorded activity remain distinct. Missing usage, identities, paths, and measurements stay unknown. Overview request totals and path counts include supported application request logs even when no activity export exists; retained requests still count when source coverage is partial.
73
+
74
+ Set `"discovery": false` in `.openmerit/terminal.json`, or pass `--no-discovery`, to skip these additional scans. Existing explicit exports, declared runtimes, inherited service configuration, machine readings, and Ollama's existing collector remain available. Set `"ollama": false` separately to disable that collector. Restart after changing `terminal.json`; supported declaration files are reread automatically.
75
+
76
+ ### Existing project request logs
77
+
78
+ Preview behavior: `openmerit dash /path/to/project` also discovers existing `.jsonl` logs (including numeric rotations such as `.jsonl.1`) under `logs/`, `runs/`, and `results/`, descending at most two subdirectories beneath each root. No exporter, project configuration, credential, or inference call is required when the application already writes a supported format. Selecting the project authorizes these bounded project-local reads. Log files may contain conversation text; the parser retains only allowlisted request metadata, discarding prompts, outputs, headers, and credentials. Symbolic-link files and directories are skipped. Discovery does not execute source code, read arbitrary databases, or scan dependency directories.
79
+
80
+ Supported records are:
81
+
82
+ - Structured request results with stable `id`, ISO `occurredAt`, `model`, explicit `status` (`running`, `completed`, or `failed`), and a `usage` object containing numeric `input_tokens`/`output_tokens` or `prompt_tokens`/`completion_tokens`. Explicit agent, task, connection, timing, and cost fields use the same names as the activity export. Failure/running records without usage need a numeric `httpStatus` or `event: "inference.request"`.
83
+ - Saved response objects with `object: "response"`, stable `id`, `model`, `created_at`, recognized status, and numeric usage, either directly in each line or under `response` in a metadata envelope.
84
+ - Saved chat completion objects with `object: "chat.completion"`, stable `id`, `model`, `created`, completed choices, and numeric usage, directly or under `response`.
85
+
86
+ These are format adapters, not automatic OpenAI/OpenRouter account ingestion. An API-compatible object does not establish its serving company or connection. A structured result must explicitly supply those identities; otherwise they stay unknown. Without an application/workload label, the relative log filename is shown. Agent/task names are never inferred from prompts or source code. Costs are copied when supplied; no missing price is synthesized. A recorded running event establishes running status only until a later event updates it; the reader cannot determine whether an abandoned request is still executing.
87
+
88
+ The reader checks at most 2,048 directory entries and the newest 128 candidate files each poll. It reads up to 512 KiB per file and 4 MiB overall, caches up to 1,000 records per file and 10,000 recent records overall. Initial oversized files import a bounded tail. Appends, split UTF-8 lines, same-file updates, truncation, and rotation are handled; deleted/unreadable sources lose their cached records. Malformed or unsupported lines and scan limits appear as partial coverage in **Sources**. Ordinary unstructured text logs and standalone `.json` response files are not supported.
89
+
90
+ ### Application scope
91
+
92
+ The terminal does not discover or read Pi, Codex, Claude Code, OpenCode, or other coding-harness histories. Older preview builds offered Pi history access; that reader and the `--access` / `--revoke-access` options have been removed. Existing `~/.openmerit/terminal-access.json` grants are ignored, left untouched, and cannot enable history collection. Restart an older terminal process after upgrading to apply the new boundary.
93
+
94
+ Application agents, background jobs, and evaluations remain in scope when they produce supported request evidence. The `agent` scope means an agent in the application, not the coding assistant developing it. Explicit exports and discovered request logs must belong to the application; records explicitly marked `scope: "development"` are rejected. An unlabeled API response cannot establish whether it came from an application or a coding tool, so the terminal does not guess its purpose from a model or agent name. Keep development logs separate from application request sources.
95
+
96
+ Configuration identifies declared services; it cannot establish past executions or costs. Local runtime inventory and whole-machine counters can describe shared infrastructure and are not attributed exclusively to the selected application. The collector does not retrieve cloud account usage, inspect arbitrary processes, or connect to a remote deployment from its source code alone. Applications without a supported request source can use the optional activity export below.
97
+
98
+ ## Explore the environment
99
+
100
+ Pages, detail panels, and inventory cards lead directly with their title. Supporting context appears below the title; there are no small uppercase labels above headings.
101
+
102
+ - **Overview:** inventory counts, observed usage, a map of recorded workload-to-deployment connections, shared capacity, and recent activity.
103
+ - **Models:** filter by self-hosted (local or private) or cloud location, search models, and inspect their available measurements. Rows lead with the model name, followed by the serving service, location, and observed request paths. Select the name to open its details.
104
+ - **Services:** a searchable inventory of model APIs, inference clouds, routers, gateways, and runtimes identified in the environment. Each service opens its models, deployments, recorded paths, explicitly linked capacity pools, and identity evidence. The secondary **Environment map** view contains the deployment map, agent focus selector, deployment list, and shared pools. Selecting an agent opens Deployments, Paths, and Capacity views. Selecting an entry connection reveals its reported hop sequence; deployment and capacity nodes open their evidence and shared dependencies. A path is a source-reported connection, not a network inspection or an enforced routing policy.
105
+ - **Activity:** a request-volume chart stacked in the five request-path colors, with selectable path totals and the most recent 100 request summaries in the selected 1-, 6-, or 24-hour window. Aggregate totals include all retained records in that window. Selecting a path isolates its chart series while the activity table keeps the full window. Chart bins support pointer inspection and Left/Right or Home/End keyboard navigation; map nodes also support keyboard inspection.
106
+ - **Local compute:** this machine and exported self-hosted servers, grouped by supplied site and rack labels, a reported-site filter when site labels exist, and per-host CPU, system RAM, accelerator readings, loaded models, and optional vLLM counters. Select a host to inspect its resources. Missing measurements and stale readings remain explicit. This does not attribute machine usage to an individual request.
107
+ - **Sources:** successful connections, unavailable sources, collection gaps, and setup instructions.
108
+
109
+ Use Command+K or Control+K to find a view, model, source, or reported agent/task. Escape closes details. Keyboard actions remain immediate, and reduced-motion preferences disable movement transitions.
110
+
111
+ Dropdowns share a dark menu, inset chevron, and neutral focus and selection states. For the agent and host-location selectors, use Enter, Space, or an arrow key to open; arrows, Home/End, or typing a name prefix move through options. Enter or Space confirms, Escape cancels, and Tab confirms and continues through the page. Clicking outside dismisses the menu. An open menu preserves its options and position during passive polling; underlying collection continues.
112
+
113
+ When request metadata includes stable agent or task IDs, the overview leads directly with their counts for the selected window, without an “Observed” prefix. Select a count to browse those agents or tasks and inspect their request totals, models, connections, and available measurements. Records without these IDs do not contribute to those counts. These are activity summaries, not an agent registry or a task lifecycle tracker; a completed inference request does not establish that its task is complete. The overview previews up to twelve models and ten recent requests. The Models view contains the full retained model inventory; Activity lists the latest 100 requests.
114
+
115
+ ### Where names come from
116
+
117
+ The terminal does not infer agent roles, task descriptions, service identities, or physical locations from arbitrary source code, prompts, or model output. Being present in a repository does not establish that an agent ran. This build reads supported metadata exports, explicit configuration, runtime inventory, and the operating system. A codebase integration must export its existing identifiers; it should not generate nicer descriptions to fill gaps.
118
+
119
+ | Identity | Source and fallback |
120
+ | --- | --- |
121
+ | Project | The selected directory’s basename |
122
+ | Agent / task | `agent` / `task` from activity metadata; otherwise the corresponding `agentId` / `taskId`. Without an ID, the record does not establish a counted agent or task. |
123
+ | Workload | The exported `workload` or the target ID in an OpenMerit application audit summary. Missing values display “Unattributed workload.” |
124
+ | Connection | The exported `route` or first reported path hop; inventory names are explicitly declared in `.openmerit/terminal.json`. No name means an unattributed route. |
125
+ | Runtime | Its configured name; an unnamed compatible runtime displays its validated loopback address. The built-in Ollama collector names the protocol it queries. |
126
+ | This machine | The full hostname reported by the operating system, without stripping its domain suffix |
127
+ | Exported host | Explicit `name`, then `hostname`, then `id` from the hardware export |
128
+ | Site / rack | Only the export’s `site` / `rack` fields. Missing placement stays unspecified; neither hostname nor IP address establishes a physical location. |
129
+
130
+ Agent and task inspectors include “Where this name comes from,” with the ID, supplied display name or fallback, metadata field, and source. Activity details keep IDs and display names separate. Local compute shows each host’s name source. “Agents recorded on this deployment” is based on agent IDs joined to matching deployment metadata in the selected window; it is not an assertion about currently running processes.
131
+
132
+ A sticky top header contains the environment selector, search, and six text tabs: Overview, Models, Services, Activity, Local compute, and Sources. A thin underline marks the selected view. Source coverage is a compact link at the right on desktop. There is no sidebar or collapse setting. On narrow screens, tabs scroll horizontally and the selected tab is brought into view when navigating. The environment selector and search remain available.
133
+
134
+ The tab row is one keyboard stop. Left/Right arrows move focus between tabs, Home/End reach the first/last tab, and Enter/Space selects a view. Keyboard selection is immediate. Focus and horizontal tab position survive passive refreshes. Direct hash URLs still open the corresponding view; Command+K or Control+K continues to open search.
135
+
136
+ Request paths use the same colors, custom icons, and text labels across the overview counts, map nodes and legend, model badges, connection cards, and activity chart legend. Map threads retain the matching path color. Icons accompany labels rather than replacing them:
137
+
138
+ | Path | Color | Icon | Meaning |
139
+ | --- | --- | --- | --- |
140
+ | Direct API | Royal blue | Two endpoints joined by a straight line | A reported direct connection to a provider's API |
141
+ | Local runtime | Amber | Processor chip | A reported connection to a local or private inference runtime |
142
+ | Router | Violet | One path branching into two | A reported routing service |
143
+ | Gateway | Cyan | An arrow passing through a frame | A reported gateway connection |
144
+ | Unclassified | Gray | Broken ring | The metadata does not establish the path type |
145
+
146
+ Path counts use the selected activity window, counting each retained request once. The short bars show each category's share of those requests. Models reached through multiple kinds of connection display each observed type. A loaded or configured model without recorded activity says “No path observed.” Model location remains a separate field: a local model can be reached through a gateway, and a cloud provider does not by itself establish a direct path. With **All agents & workloads** selected, every retained request, workload, entry connection, and model/service pair in the selected time window is represented. There are no top-N workload or model limits. Workloads and models with identical sets of recorded entry connections form compact card stacks with complete counts. Each stack shows an actual member name, a `+N` count for the other members, and totals for the whole stack. Select a group to reveal every member, or use **Expand all** and **Collapse all**. Expanded maps scroll at the same card and text scale, and keyboard users can expand groups with Enter or Space. Group expansion survives passive refreshes; newly observed members join their matching group.
147
+
148
+ The three columns show workloads, entry connections, and models. Each thread aggregates only actual recorded relationships. The footer reports all represented requests in the selected scope; expanding or collapsing never changes that total. Focusing an agent applies the same complete representation to that agent’s activity. Model identity includes its serving service, so the same model served by different services stays distinct. Model captions show the reported service and, when available, explicitly matched deployment information. Missing deployment identities are explained in the inspector rather than repeated on map nodes. The map does not infer a physical instance from a model name. Select a connection to inspect its reported hop sequences; the Deployments control opens declared and runtime-discovered inventory, which can include instances without attributable request activity.
149
+
150
+ An explicit `routeKind` determines the category. Otherwise, the terminal matches the route name against declared inventory or local runtime source names, ignoring case and surrounding whitespace. A matching provider is direct; a matching router or gateway retains that type; a runtime must be declared local or private. The “Local runtime” path category includes self-hosted private runtimes; the model location and host placement distinguish them from this machine. Conflicting inventory types remain unclassified. The literal route names `Direct` and `Direct API` are also treated as direct connections. It does not promote a generic activity record’s serving `provider` into a route or guess from model location or arbitrary substrings. The category describes the entry connection once per request. An optional ordered `path` supplies additional reported hops; the terminal never infers hidden intermediaries. When `route` or `routeKind` is omitted, a valid path supplies its first-hop value. A path contradicting an explicit entry name or type is discarded, preserving that explicit entry metadata.
151
+
152
+ Local compute uses Liveline for CPU and memory history. These are passive readings on the configured collection interval, not a live inference stream. The curve interpolates between recorded points and holds the latest value between polls. Hovering provides chart inspection; the readout below it identifies the nearest **actual sample**, with its timestamp. Focus that readout and use Left/Right, Home, or End to inspect recorded samples with the keyboard. Missing readings or missed collection windows break the history; the chart resumes from the latest uninterrupted run. Reduced motion uses a static chart. Whole-machine usage is not attributed to individual inference requests.
153
+
154
+ Torph animates changed totals, table measurements, status labels, and source counts while preserving the surrounding interface between polls. Neither animation library changes the collection interval or makes inference calls.
155
+
156
+ The environment selector opens `inference-sample`, a fictional enterprise workspace with 40 agents, 140 tasks, 21 models, and 9,600 requests over 24 hours. It includes all five path classifications, traffic bursts, running and failed requests, missing measurements, partial and unavailable sources, three self-hosted runtimes, 27 distinct deployments, 14 shared capacity pools, and an hour of host history. Its six hosts include a workstation and five rack servers across two racks, with eleven accelerators listed. Four hosts have recent readings, one is stale, and one has no measurement. Seven loaded models are placed explicitly across three hosts. Six cloud models have separate `us-1` and `eu-1` example deployments. The capacity sample includes twelve recent readings, one stale reading, and one missing reading. Shared gateway pools connect consumers across reported agent IDs. Names resemble exported configuration keys (for example, `support.review` and `llm-gateway`); some agent/task display names are absent and the UI uses their IDs. These identifiers are fictional, not discovered in the current repository. The sample workstation has no site label; server sites and racks are fictional exporter fields. Baseten, Groq, and Cerebras illustrate serving services distinct from model creators; their sample configuration evidence, deployment modes, and quotas are fictional. Model rates, latency, allocation, and utilization values are illustrative; the sample is for inspecting the interface at scale, not a performance benchmark. Sample data never enters the real snapshot or project files. A command started with `--demo` only exposes a sample workspace; restart without that flag to collect real data.
157
+
158
+ ### Optional map preview
159
+
160
+ The test preview includes an opt-in React Flow + ELK map experiment. Add `?map=flow` before the page hash in your local terminal URL (for example, `http://127.0.0.1:4310/?map=flow#overview`). To try fictional enterprise data instead, use `?sample=1&map=flow#overview`. The existing map remains the default; **Original map** removes the preview flag immediately. No data migration or server restart is needed to switch views.
161
+
162
+ The preview keeps the existing card dimensions and uses ELK to arrange all collapsed groups. Layout is recalculated only when the set of nodes or relationships changes, not when request counts refresh. The graph scrolls at a fixed scale rather than automatically zooming to fit.
163
+
164
+ **Browse all** opens a searchable inventory of every workload, entry connection, and model/service pair in the selected scope. Selecting a stack filters that inventory to its exact members; the visible result count always shows the filter coverage. Selecting a member spotlights its actual recorded paths and brings a matching member to the front of each related stack without expanding the diagram. This does not change the total request counts. **Details** opens the existing inspector; **Clear spotlight** or Escape clears the selection. The browser sits beside the map on wide screens and below it on narrower screens, with independently scrolling lists. Keyboard users can navigate its tabs with arrows, Home, and End.
165
+
166
+ This is a presentation experiment within the preview, not a new data source. Its separately bundled scripts and styles are loaded from the local server only when the preview is selected. It makes no inference calls and adds no collection permissions. There is additional download and initial layout cost; very large and densely connected environments still need user evaluation. The original map remains available for comparison and rollback.
167
+
168
+ ## Collection sources
169
+
170
+ ### OpenMerit evidence
171
+
172
+ The terminal reads `.openmerit/events.jsonl` for recorded `product_task_completed` automation signals. It imports the model and target identifiers and available cost, token, and duration measurements. It does not treat generic Pi activity as application evidence.
173
+
174
+ The existing audit redactor may have replaced token-shaped fields with `[REDACTED]`. These remain unknown in the terminal. The terminal does not reconstruct missing counts from cost or other fields.
175
+
176
+ ### Existing activity exports
177
+
178
+ Export request metadata from an application, agent, router, gateway, or existing evidence system to `.openmerit/inference.jsonl`. Each record must occupy one newline-terminated JSON line. The file is an optional import contract, not an instrumentation SDK or an inline proxy.
179
+
180
+ ```json
181
+ {"id":"request-123","occurredAt":"2026-10-07T08:00:00Z","workload":"your-service-id","scope":"application","model":"your-model","provider":"your-provider","route":"your-gateway","routeKind":"gateway","location":"cloud","status":"completed","inputTokens":120,"outputTokens":40,"cacheReadTokens":0,"durationMs":850,"cost":0.002,"costKind":"reported","currency":"USD"}
182
+ ```
183
+
184
+ Required fields are `id`, `occurredAt` (a valid timestamp), and `model`. Optional fields have the following meanings:
185
+
186
+ | Field | Meaning |
187
+ | --- | --- |
188
+ | `workload` | Application, agent, or evaluation name |
189
+ | `agentId`, `agent` | Optional stable agent ID and display name; the ID enables grouping and observed-agent counts. Without a name, the ID is displayed. |
190
+ | `taskId`, `task` | Optional stable task ID and display name; the ID enables grouping and observed-task counts. These fields carry no task lifecycle or execution trace. |
191
+ | `scope` | `application`, `agent`, or `evaluation`; defaults to `application` |
192
+ | `provider` | Serving service or runtime name, shown as “Served by”; otherwise unattributed. This field does not establish who created the model. |
193
+ | `modelCreator` | Optional source-reported creator name, capped at 160 characters. Missing values stay unknown; the terminal does not derive this field from the serving company, SDK, model name, or API catalog ownership. |
194
+ | `route` | Reported gateway, router, or direct connection name; otherwise unattributed |
195
+ | `routeKind` | Optional `direct`, `local`, `router`, `gateway`, or `unknown`; explicit values take precedence over name matching. Omitted/invalid values become `null` in the snapshot and permit the fallback classification above. |
196
+ | `deploymentId` | Optional stable deployment identity. Joins only when the declared/discovered deployment ID, model, and provider all match. No join is inferred from a model name alone. |
197
+ | `path` | Optional ordered array of 1–8 `{ "name": "connection", "kind": "gateway" }` hops. Kinds use the same values as `routeKind`. Invalid/oversized paths are discarded as a whole; unsupported kinds become `unknown`. The first hop must agree with explicit entry metadata. |
198
+ | `location` | `local`, `private`, `cloud`, or `unknown`; defaults to `unknown` |
199
+ | `status` | `running`, `completed`, or `failed`; defaults to `completed` |
200
+ | `inputTokens`, `outputTokens`, `cacheReadTokens` | Nonnegative measured counts, when available |
201
+ | `durationMs` | Nonnegative measured duration in milliseconds |
202
+ | `cost` | Nonnegative source-supplied cost |
203
+ | `costKind` | `reported` or `estimated`; defaults to `reported` |
204
+ | `currency` | `USD` is supported; other currencies are excluded from dollar totals |
205
+ | `requestId` | Optional correlation metadata; not used for automatic cross-source joining |
206
+
207
+ Use a stable `id` to append updates to an ongoing request. A later complete record with that ID replaces the prior record, so updates must include all fields you want retained. Use different IDs for separately charged attempts or retries. IDs must be unique within the activity feed. Correlate and deduplicate upstream when exporting overlapping systems.
208
+
209
+ Missing, invalid, negative, or redacted measurements remain unknown. Numeric zero is preserved. Only allowlisted metadata fields are projected into the terminal; prompts, completions, headers, and other arbitrary payload fields are omitted. Do not put secrets in metadata names or identifiers.
210
+
211
+ Explicit application activity and supported project request logs are reconciled together. A record merges with a copy only through either (a) the same recorded ID, timestamp, model, provider, and connection or (b) the same API request ID, model, and recorded connection (provider if no connection is supplied). Equal timestamps or token counts alone never establish identity. Distinct attempts must have distinct IDs. A completed record replaces a running copy; otherwise explicit activity takes precedence, then project logs in scan order. All contributing source IDs are retained. Conflicting measurements are disclosed rather than added together.
212
+
213
+ When sources cover the same connection/model and unmatched records lack a comparable API request ID, lower-priority unmatched records are excluded from combined totals; **Sources** retains each source's separate count and explains the unresolved overlap. This conservative rule can omit genuinely distinct activity rather than claim an unverified sum. Sources using different aliases for the same connection cannot always be correlated; this is not universal cross-system matching. OpenMerit application summaries remain a fallback only when no request-level evidence is available. The snapshot reports `discovery.requestSource` (`combined` for multiple contributing sources), `requestSources`, and reconciliation counts for included, matching, excluded, and conflicting records.
214
+
215
+ ### Local model runtimes
216
+
217
+ Ollama is read at `http://127.0.0.1:11434` using `GET /api/tags` and `GET /api/ps`. This provides installed models, loaded models, and available model-allocation metadata. It does not supply request history. Loaded means resident in memory, not necessarily generating.
218
+
219
+ Other local servers can be declared in `.openmerit/terminal.json`:
220
+
221
+ ```json
222
+ {
223
+ "activityFile": ".openmerit/inference.jsonl",
224
+ "ollama": "http://127.0.0.1:11434",
225
+ "runtimes": [
226
+ { "name": "LM Studio", "type": "openai-compatible", "url": "http://127.0.0.1:1234" },
227
+ { "name": "vLLM", "type": "vllm", "url": "http://127.0.0.1:8000" }
228
+ ],
229
+ "inventory": [
230
+ { "name": "Your gateway", "kind": "gateway", "location": "cloud" },
231
+ { "name": "Your model", "kind": "model", "provider": "Your provider", "location": "cloud" }
232
+ ]
233
+ }
234
+ ```
235
+
236
+ Restart the terminal after configuration changes. Set `ollama` to `false` to disable its discovery. Runtime URLs must be HTTP loopback origins with no credentials, path, query, or fragment. Remote and authenticated runtime endpoints are not supported in this initial version. The runtime list is limited to eight entries.
237
+
238
+ `openai-compatible` reads `GET /v1/models`. Models are labeled available; the list does not prove they are loaded. `vllm` also reads `GET /metrics` for running/waiting requests, inference cache usage, and cumulative input-token, output-token, and finished-request counters. Cache usage is the highest valid `vllm:kv_cache_usage_perc` series, expressed as a percentage; percentages are not summed across workers. These runtime aggregates are displayed separately from imported request totals and are not attributed to projects. Counters can reset when a runtime restarts.
239
+
240
+ The optional `inventory` accepts `model`, `provider`, `gateway`, `router`, and `runtime` entries. Entries are declared configuration until supported by observations. Known process-environment variables can also identify configured services, as described below. Their credential values are not returned or used to call those services. Automatic discovery additionally reads the supported project declarations and request logs described above. Credential stores and arbitrary processes are not scanned.
241
+
242
+ ### Service inventory
243
+
244
+ The **Services** tab replaces Connections. It opens on a compact inventory with type filters and service/model search; the **Environment map** switch preserves the topology view. Existing `#connections` links open that map, while `#services` opens Services. The overview keeps its compact map.
245
+
246
+ Service rows come from declared non-model inventory, configured runtime sources, serving-provider metadata, deployments, and recorded path names. A model's `modelCreator` does not create a service row or establish an API connection to that lab. Names remain as reported. Exact curated aliases can share a row only when there is one unambiguous existing service and the alias is not separately declared. Explicitly declared service names remain separate.
247
+
248
+ Types are **Model API**, **Inference cloud**, **Router**, **Gateway**, **Runtime**, or **Unclassified**. Explicit declared/reported router, gateway, and runtime roles take precedence over the known-service registry. A `provider` inventory entry can set `serviceType` to `model-api` or `inference-cloud`; this also takes precedence over the registry. Multiple explicit roles remain visible. Known API services use Model API; Baseten, Groq, and Cerebras use Inference cloud. Custom providers without a supplied subtype remain unclassified. These roles describe services; the existing five request-path categories are unchanged. No new type implies discovered hardware or a verified network path.
249
+
250
+ Requests, paths, median latency, and supplied USD spend cover the selected window. A request counts once per service, even if a hop repeats, but can touch several services: row totals are not additive. Latency and spend describe the whole recorded request, not service-added latency or the service's bill. Last activity comes from all retained records; a service can therefore show activity outside the selected window. Models/deployments include inventory plus relationships reported in that window. Shared pools are linked only through explicit service or serving-deployment associations, not assigned to every intermediary in a path. A displayed capacity reading is one reported pool measurement, with its description directly beneath the value on the same right edge; full pool details remain in the inspector.
251
+
252
+ Configuration, inventory, and activity evidence are distinct. A configured service with no requests says **Configured only**; inventory does not imply connectivity, current health, available capacity, or usage. Every inspector includes the fields and sources behind the identity. Filtering and inspecting services remain passive and make no external requests.
253
+
254
+ ### Inference services and model creators
255
+
256
+ The terminal recognizes configuration evidence for Baseten, Groq, and Cerebras alongside the existing services. Inherited variables and supported literal declarations in the selected project’s `.env` and `.env.local` files provide evidence. Shell histories, arbitrary source files, and credential stores are not scanned. A configured service does not establish an active connection, model inventory, usage, available capacity, or a hardware allocation.
257
+
258
+ | Service | Credential-variable names | Optional endpoint-variable names |
259
+ | --- | --- | --- |
260
+ | OpenAI | `OPENAI_API_KEY` | `OPENAI_BASE_URL`, `OPENAI_API_BASE` |
261
+ | Anthropic | `ANTHROPIC_API_KEY` | `ANTHROPIC_BASE_URL` |
262
+ | OpenRouter | `OPENROUTER_API_KEY` | — |
263
+ | Google | `GOOGLE_GENERATIVE_AI_API_KEY`, `GEMINI_API_KEY` | — |
264
+ | Groq | `GROQ_API_KEY` | `GROQ_BASE_URL` |
265
+ | Cerebras | `CEREBRAS_API_KEY` | `CEREBRAS_BASE_URL` |
266
+ | Baseten | `BASETEN_API_KEY` | `BASETEN_BASE_URL` |
267
+
268
+ Nonempty credential variables identify a configured service by variable name. Endpoint overrides take precedence for that SDK's group: for example, `OPENAI_BASE_URL=https://api.groq.com/openai/v1` identifies Groq rather than OpenAI. Recognized HTTPS hosts are `api.openai.com`, `api.anthropic.com`, `openrouter.ai`, `api.groq.com`, `api.cerebras.ai`, `inference.baseten.co`, and Baseten's `model-<id>.api.baseten.co` hosts. User information and nonstandard ports are rejected. Unknown, local, or malformed overrides suppress that group's credential-name fallback; the source reports only the unrecognized variable name. Multiple configured endpoints do not establish which one an application uses.
269
+
270
+ Service details show the identifying variable and whether the evidence was a credential name or a recognized endpoint. **Sources → Service configuration** explains inherited variables; project-file sources identify declarations from each file and any unknown endpoints. Credential values and raw endpoint URLs, including their paths and queries, never enter the snapshot. Recognized endpoint hostnames may appear in service evidence; Baseten deployment identifiers are redacted as `model-….api.baseten.co`. Configured local runtimes can display their validated loopback host and port. No arbitrary endpoint URL becomes a link. Detection makes no cloud requests and does not assign a request path or rewrite recorded providers from ambient environment configuration.
271
+
272
+ Supply `modelCreator` in activity metadata, model inventory, deployments, or exported loaded-model entries when that identity is available. The model inspector shows the reported creator separately from **Served by**, including its source. Conflicting creator reports for the same model/service leave the inventory creator unresolved. Existing family icons are visual recognition only; they do not populate creator metadata. An explicit creator takes precedence over family icon matching. Baseten, Groq, and Cerebras service icons are bundled locally.
273
+
274
+ For example, an export may report `model: "gpt-oss-120b"`, `modelCreator: "OpenAI"`, and `provider: "Groq"`. An application using a gateway still reports its gateway path; the serving company does not change that category. No account API, inference call, remote hardware discovery, or model availability check is performed by this detection.
275
+
276
+ ### Deployments and shared capacity
277
+
278
+ A model identity is separate from a deployment of that model. Declare deployment identities and shared resource pools in `.openmerit/terminal.json`. These are metadata declarations, not credentials or endpoints to call:
279
+
280
+ ```json
281
+ {
282
+ "deployments": [
283
+ {
284
+ "id": "assistant-primary",
285
+ "name": "US East / primary",
286
+ "model": "your-model",
287
+ "provider": "your-serving-service",
288
+ "modelCreator": "your-model-creator",
289
+ "servingMode": "dedicated",
290
+ "location": "cloud",
291
+ "poolIds": ["provider-account"]
292
+ }
293
+ ],
294
+ "capacityPools": [
295
+ { "id": "provider-account", "name": "Provider / shared account", "kind": "quota", "windowSeconds": 60 },
296
+ { "id": "llm-gateway-pool", "name": "llm-gateway", "kind": "gateway" }
297
+ ],
298
+ "inventory": [
299
+ { "name": "llm-gateway", "kind": "gateway", "location": "cloud", "poolIds": ["llm-gateway-pool"] }
300
+ ],
301
+ "capacityFile": ".openmerit/capacity.json"
302
+ }
303
+ ```
304
+
305
+ Deployments may supply `modelCreator` and `servingMode` (`shared`, `dedicated`, `self-hosted`, or `unknown`). These are declarations; no mode is inferred from a company name or credential. Declared deployments also contribute configured models to the inventory when no runtime or activity source has listed them. Hardware stays “Not reported” until a matching host export supplies it.
306
+
307
+ Deployment locations accept `local`, `private`, `cloud`, or `unknown`. Pool kinds accept `quota`, `runtime`, `gateway`, or `unknown`. Limits are 200 deployments, 100 pools, and 16 pool references per entry; labels and IDs are capped at 160 characters. Duplicate IDs are excluded rather than resolved arbitrarily. IDs beginning with `runtime:` are reserved for discovery. An optional deployment `hostId` links to the exact host ID in Local compute; an absent or unmatched ID stays unplaced. Configuration changes require restarting the terminal.
308
+
309
+ Activity exports can add `"deploymentId":"assistant-primary"` to a record with the same model and provider. An optional `path` can report ordered router/gateway/runtime connections. Connection-to-pool links use exact, unambiguous name and kind matches in inventory; deployment-to-pool links use explicit IDs. These dependencies are declarations, not independently verified network paths. One request may touch several pools; repeated links to the same pool count it only once. Pool request counts must not be summed as a total of unique requests.
310
+
311
+ Local runtime inventories also create deployment identities and a shared runtime pool. Discovered deployments use `hostId: "this-machine"`. Discovered IDs are `runtime:<model inventory id>` for deployments and `runtime:<source id>` for pools; inspect the snapshot to obtain them. Existing activity must explicitly reference those IDs to join to discovered deployments. No association is guessed from an equal model name. A loaded model means resident memory, not available concurrency. Declared and discovered instances are not automatically merged.
312
+
313
+ Export existing capacity measurements to the project-local JSON file (defaults to `.openmerit/capacity.json`):
314
+
315
+ ```json
316
+ {
317
+ "pools": [
318
+ {
319
+ "id": "provider-account",
320
+ "observedAt": "2026-10-07T08:00:00Z",
321
+ "resetAt": "2026-10-07T08:01:00Z",
322
+ "metrics": { "requestsUsed": 842, "requestsLimit": 1000, "tokensUsed": 630000, "tokensLimit": 1000000 }
323
+ },
324
+ {
325
+ "id": "llm-gateway-pool",
326
+ "observedAt": "2026-10-07T08:00:00Z",
327
+ "metrics": { "running": 46, "concurrencyLimit": 50, "waiting": 4 }
328
+ }
329
+ ]
330
+ }
331
+ ```
332
+
333
+ The file is read on the normal passive interval. Write a complete JSON object, preferably by atomic replacement. It is limited to 128 KiB and must resolve to a regular file within the project. Only declared pool IDs and allowlisted fields are used. Invalid or unavailable files clear exported readings; no old measurement is silently marked current. Only the first 100 entries are considered. Duplicate IDs and invalid/future timestamps are excluded. The terminal does not call cloud account APIs to populate these measurements.
334
+
335
+ Supported nonnegative numeric metrics are `running`, `waiting`, `concurrencyLimit`, `requestsUsed`, `requestsLimit`, `tokensUsed`, `tokensLimit`, `cachePercent` (0–100), `memoryUsedBytes`, and `memoryTotalBytes`. Missing or invalid values stay unknown; zero stays zero. Limit bars require both a usage value and a positive limit. Counts may exceed a limit; text retains the measurement while bar width caps at 100%. A pool's declared `windowSeconds` describes the exported quota window; it is independent of the selected activity window. `resetAt` is source-reported metadata and does not trigger a reset or action.
336
+
337
+ The inspector shows the source timestamp and marks readings older than three collection intervals stale. Checking a file does not refresh the measurement's age. vLLM running/waiting and cache readings populate its discovered runtime pool; Ollama and compatible inventory alone provide no capacity measurement. Pool totals cover all consumers of that resource, including other projects. Shared agents are identified from retained activity in the selected window, not from pool counters. These readings do not establish agent-level resource attribution, bottleneck causality, remaining headroom, or guaranteed readiness.
338
+
339
+ The overview previews three pools, prioritizing recent measured utilization fractions. The full list includes stale and missing readings. Agent inspectors offer separate Deployments, Paths, and Capacity views. Capacity inspectors show shared agents, and selecting one returns to that agent's environment. Inspector contents refresh on the normal poll without resetting scroll or keyboard focus. Unknown deployment identities, missing pool links, and unavailable readings remain visible gaps.
340
+
341
+ ### This machine
342
+
343
+ CPU utilization is sampled from operating-system CPU time deltas and is unknown until two samples exist. Memory is total physical memory minus reported free memory. These measurements describe the whole machine. macOS cache and reclaimable-memory behavior can differ from Activity Monitor's presentation.
344
+
345
+ The terminal retains up to 240 samples per host in memory. The operating-system collector does not collect GPU utilization; additional accelerator readings can come from the hardware export below. Energy and hardware costs are not inferred. Ollama's reported VRAM allocation can overlap system memory on unified-memory machines; these values must not be added together.
346
+
347
+ ### Additional hosts and racks
348
+
349
+ The **Local compute** inventory includes this machine automatically, with the reserved ID `this-machine`. Its display name is the full operating-system hostname. Additional hosts are supplied by an existing exporter through `.openmerit/hosts.json`. An optional `hardwareFile` in `.openmerit/terminal.json` changes that project-local path. The terminal reads this file on its normal 10–30-second polling interval; it does not connect to remote hosts, scan networks, or install collectors. Each host may supply `name` and/or `hostname`; display precedence is `name`, `hostname`, then `id`. Snapshots expose the selected field as `nameField`. Site and rack are optional export labels, not discovered locations. The site filter is hidden when no host reports a site.
350
+
351
+ For example, an exporter can atomically replace this JSON file with its latest measurements:
352
+
353
+ ```json
354
+ {
355
+ "hosts": [{
356
+ "id": "gpu-01",
357
+ "name": "Inference 01",
358
+ "site": "Primary",
359
+ "rack": "Rack A",
360
+ "observedAt": "2026-10-07T12:00:00Z",
361
+ "chip": "AMD EPYC",
362
+ "cores": 64,
363
+ "platform": "linux",
364
+ "arch": "x64",
365
+ "cpuPercent": 32.4,
366
+ "memoryTotal": 274877906944,
367
+ "memoryUsed": 137438953472,
368
+ "accelerators": [{
369
+ "id": "gpu-0",
370
+ "name": "NVIDIA H100",
371
+ "vendor": "NVIDIA",
372
+ "kind": "gpu",
373
+ "utilization": 72,
374
+ "memoryTotal": 85899345920,
375
+ "memoryUsed": 42949672960
376
+ }],
377
+ "loadedModels": [{
378
+ "id": "model-instance-1",
379
+ "name": "example-model",
380
+ "provider": "Private runtime",
381
+ "runtime": "vLLM",
382
+ "deploymentId": "private-model-primary",
383
+ "memoryBytes": 38654705664,
384
+ "memoryKind": "accelerator"
385
+ }]
386
+ }]
387
+ }
388
+ ```
389
+
390
+ Replace the example timestamp with the actual measurement time. Each host retains its own timestamp; reading the file again does not refresh it or create a new history sample. A reading older than three collection intervals is marked stale. Future or absent timestamps cannot establish utilization or model residency. Static inventory such as processor name and memory capacity can still be listed. A removed host, invalid file, oversized file, or failed read clears imported hosts and their histories; this machine remains available.
391
+
392
+ The export is bounded to 256 KiB, 64 hosts, 16 accelerators per host, and 100 loaded-model entries per host. Labels and IDs are capped at 160 characters. Duplicate IDs within each list are excluded, and exported `this-machine` entries are rejected. Files must be regular files inside the project after symlink resolution. Unsupported fields are discarded. Only supplied metadata is retained; endpoints, credentials, prompts, and responses are not imported.
393
+
394
+ CPU and accelerator utilization are percentages from 0 to 100; memory fields are nonnegative bytes, displayed as GiB. Used memory exceeding a supplied total is discarded. Missing metrics remain unknown, and zero remains zero. System RAM, accelerator memory, and model allocations are displayed separately and are never added into a misleading fleet total. Hosts can contain mixed accelerator types. Optional accelerator `vendor` and `kind` fields identify reported hardware; supported kinds are `gpu`, `lpu`, `wse`, `npu`, `tpu`, `other`, and `unknown`. They are never inferred from an inference service, and do not supply missing utilization or memory readings. Fleet counts describe listed inventory, not usable capacity.
395
+
396
+ **Loaded models** means source-reported residency, not an active request or guaranteed readiness. `memoryKind` accepts `accelerator`, `system`, `unified`, or `unknown`; it identifies the allocation being reported. An omitted `loadedModels` field means residency is unknown; an empty array means the source explicitly reported no loaded models. Stale hosts show their last reported models. Exported model names also enter the model inventory as private models, with unknown current state when their host reading is stale. Replicas are kept separate in each host's residency list. Model identity matches do not establish deployment placement: declare the matching `hostId` on a deployment to make that link.
397
+
398
+ The operating-system view currently receives loaded-model evidence from Ollama's `/api/ps`. A compatible `/v1/models` catalog establishes availability, not residency. Runtime counters remain associated with their host and retain their separate lifetime scope.
399
+
400
+ ### Model and service identities
401
+
402
+ Small, locally bundled official marks identify recognized services and model families throughout the terminal: the overview and Models table, environment-map nodes, deployment lists, activity rows, loaded models, runtime counters, source rows, shared-capacity panels, search results, and detail inspectors. Marks are unframed and scaled to the surrounding text. The custom colored path-type icons remain separate: a service's identity and its routing role convey different information.
403
+
404
+ Service names use an exact curated registry, including supported “API” aliases. It covers OpenAI, Anthropic, Google AI, DeepSeek, Mistral, Cohere, Qwen, Meta, Nomic, BAAI, Ollama, vLLM, LM Studio, OpenRouter, and OpenMerit's own mark. Custom services remain text-only beside their path-type icons. An OpenAI-compatible endpoint does not inherit OpenAI branding merely because it supports that protocol.
405
+
406
+ Model marks identify the recognized family independently of where the model runs. For example, Qwen models carry Qwen's mark while their Ollama runtime carries Ollama's; Llama carries Meta's while vLLM retains its own identity. The registry recognizes GPT and selected OpenAI reasoning/embedding families, Claude, Gemini/Gemma, DeepSeek, Mistral families, Cohere Command/Embed, Qwen, Llama, Nomic Embed, and BGE. Family matching is anchored at the start of a model name. Namespaced IDs must use a recognized official namespace for that family; arbitrary namespaces and unfamiliar names receive no model mark or placeholder. Display names and the reported provider/runtime remain unchanged. A family mark is a visual aid, not verification of model weights or provenance.
407
+
408
+ All 21 models in the enterprise sample have a consistent identity across inventory, requests, deployments, and host residency. Shared pools receive a service mark only when their declared links resolve unambiguously to one recognized service, or their own exact name is recognized. Decorative images are hidden from assistive technology so a model or service name is announced once; text labels and keyboard behavior remain available.
409
+
410
+ There are no favicon lookups, third-party image services, or remote asset requests while using the terminal. Asset origins and hashes are recorded in `packages/terminal/brand-assets.json` in the source checkout. The additional model-family references are the publishers' model cards for [Qwen3](https://huggingface.co/Qwen/Qwen3-8B), [Llama 3.3](https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct), [Nomic Embed](https://huggingface.co/nomic-ai/nomic-embed-text-v1.5), and [BGE](https://huggingface.co/BAAI/bge-reranker-v2-m3). Marks identify their respective makers or services and do not imply affiliation.
411
+
412
+
413
+ ## Freshness, retention, and limits
414
+
415
+ Collectors run once per interval with no overlapping collection cycles. Each runtime request has a timeout and a response-size bound. Browser refreshes read the cached snapshot, so additional tabs do not multiply runtime polling. The interface shows a delayed-collection notice when its snapshot is older than three intervals. A successful source check does not mean that source has recent inference activity.
416
+
417
+ File readers handle partial lines, split UTF-8 writes, and ordinary file rotation. Activity and audit file reads are bounded to 4 MiB, and each reader retains up to 10,000 unique records. Discovered project logs have the stricter per-file and aggregate limits described above. A large initial file imports its tail. Truncation and malformed records are disclosed in source details. Results are limited to the most recent 24 hours and exclude future timestamps. An unavailable or removed file is not kept looking connected.
418
+
419
+ Collected usage is not persisted to a new database. Optional access choices are stored separately in the user-level permissions file. Machine history and imported-record caches are process-local and disappear when the terminal closes. Existing source files remain under their owner's retention policy.
420
+
421
+ The map represents every retained request, workload, entry connection, and model/service pair in its selected scope. Collapsible groups reduce visual density without removing records. Its footer states the complete represented request count, while deployment evidence and gaps remain available in inspectors. Full retained request metadata is available through the JSON API. There is no inferred quality score, provider billing verification, automatic cloud account connection, or universal inference discovery.
422
+
423
+ ## Agent access
424
+
425
+ ```sh
426
+ openmerit dash --json
427
+ ```
428
+
429
+ From the checkout:
430
+
431
+ ```sh
432
+ node packages/terminal/src/cli.mjs dash /path/to/project --json
433
+ ```
434
+
435
+ While the server is running, `GET /api/snapshot` returns the cached JSON object. Schema version 1 contains `mode`, `observedAt`, `interval`, `project`, `models`, `inventory`, `records`, `sources`, `runtimeStats`, `deployments`, `capacityPools`, `hosts`, `machine`, and `discovery`. The optional `discovery` object reports whether discovery is enabled, its last check time, the 60-second configuration discovery interval, candidate count, contributing request sources, and reconciliation coverage. Request records can include `sourceIds`, a supported log `evidenceFormat`, and `conflictingFields`. Colliding local IDs from demonstrably distinct records receive a stable namespaced ID while retaining `originalId`. The host, deployment, and pool fields extend the unreleased version-1 snapshot; older snapshots without them show unknown coverage. `machine` remains the current host for compatibility; `hosts` is the full hardware inventory. Fields containing unsupported measurements use `null` or are absent. `GET /api/sample` returns only fictional data. `GET /api/info` describes the process's mode and interval.
436
+
437
+ The server binds to `127.0.0.1`, accepts read-only GET requests, checks Host and Origin, rejects cross-site browser requests, and serves only its bundled assets and explicit API routes. It provides no general file-serving endpoint. Local processes and other users with access to this machine may still reach the loopback service; this is not a multi-user authentication boundary.
438
+
439
+ See [security and data](security.md) for the separate OpenMerit orchestration and harness data boundaries.
@@ -0,0 +1,30 @@
1
+ # Improvement lifecycle
2
+
3
+ ## Stages
4
+
5
+ 1. **Unconfigured** — no confirmed profile exists.
6
+ 2. **Establishing evidence** — Pi confirms the exact incumbent application model, infers a task-specific metric and sampling plan, explains that plan and estimated cost for user confirmation, then creates and verifies evals and observability. In the unreleased source checkout, setup can also record sourced preliminary price and reputation leads; published `openmerit@0.1.5` does not yet do this.
7
+ 3. **Collecting baseline** — normal product work accumulates evidence.
8
+ 4. **Baseline ready** — in the unreleased source checkout, OpenMerit core checks stored application metric windows against every required objective and its sample threshold. Pi still authors graders and collectors and reports the measurements.
9
+ 5. **Discovering candidates** — Pi researches credible candidates within budget, refreshing any preliminary setup leads and considering new ones. This stage still follows baseline sufficiency.
10
+ 6. **Running challenger trials** — Pi runs the frozen incumbent and challenger set over the same cases under one recorded experiment manifest.
11
+ 7. **Calculating frontier** — Pi calculates and explains the Pareto frontier.
12
+ 8. **Frontier ready** — OpenMerit verifies the returned snapshot.
13
+ 9. **Approval or swap** — policy decides whether confirmation is required.
14
+ 10. **Verifying swap** — the harness is asked to evaluate the applied change.
15
+ 11. **Monitoring** — production evidence continues to accumulate.
16
+ 12. **Rollback** — the harness is asked to restore the prior configuration when regression policy permits it.
17
+
18
+ ## Event-driven nudges
19
+
20
+ OpenMerit remains quiet when no action is justified. On first interactive entry to an unconfigured project, the Pi adapter performs a bounded source preflight. The unreleased source checkout repeats detection after a settled build turn, allowing Pi to build an application in an empty directory before OpenMerit automatically establishes the profile. Later nudges become due when durable application-target signals satisfy the configured task-count or elapsed-time cadence, a metric crosses a regression threshold, the model catalogue changes under an enabled policy, a post-swap window completes, an intent enables the next stage, rollback policy covers a regression, or the user invokes a command.
21
+
22
+ Every signal has an idempotency key. A due intent is persisted before dispatch so duplicate webhooks, restarts, and retrying adapters do not double-count tasks or silently lose pending work.
23
+
24
+ ## Evaluation budget
25
+
26
+ Candidate work can cost money. The contract passes the confirmed currency, spend, candidate count, runs per candidate, and expiration constraints to Pi. The harness is responsible for staying within those limits; OpenMerit does not independently meter spending or cut off provider requests. See [budgets and permissions](budgets.md).
27
+
28
+ ## Supervised graduation
29
+
30
+ Proposals wait for `/openmerit approve` unless automatic swaps are enabled and the configured number of verified swaps has been reached. That number can be zero. Post-swap evaluation remains part of the workflow and depends on the harness reporting its results. A reported post-swap regression can trigger a rollback request when policy permits it.