ephemeris-cli 0.0.0-stage → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,62 @@
1
+ # Contracts and ownership
2
+
3
+ Ephemeris CLI is the customer interface. It calls the existing public application API; Paracast owns model execution and routing. The CLI owns local preparation, portable receipts, deterministic presentation, and PNG/SVG rendering. It does not launch a separately installed Python toolkit.
4
+
5
+ ## Inspected upstream implementations
6
+
7
+ - `paracast-web` commit `7ae625aa69c7f751f26e87e6c92a56691aa4f543`: `public/openapi-m1.json`, forecast validation, model/account routes, and hosted MCP contracts.
8
+ - `Gnomon` commit `03b7635a0672644eb41c0a36c3187c8ead8df534`: data/time semantics, frozen data references, `agent_summary.py`, forecasting request/result contracts, Ephemeris connector, HTTP transport, and decision-memory/evaluation docs.
9
+ - `ephemeris-mcp`: npm distribution and environment bearer-key conventions.
10
+
11
+ At inspection time the Gnomon connector has gateway/direct mapping but no explicit forecast idempotency-key argument/header. Its transport deliberately does not retry POST. Its evaluation docs still mark authenticated live Ephemeris inference as unverified. This release does not alter or depend on that connector.
12
+
13
+ ## Reuse decision
14
+
15
+ There is no language-neutral Gnomon runtime package in the inspected implementation. Installing Python or launching its CLI by default would violate the single-product installation requirement. Running its logic remotely would require new hosted operations, authorization, deployment, and contracts; no such endpoint is assumed here.
16
+
17
+ This release reuses **contract semantics**, not a copied Python engine: Gnomon's summary field structure (`scope`, `result`, `basis`, `limitations`, `references`, `followups`), explicit unknowns, time/shape boundaries, and observed-statistic definitions. A developer script executes the actual Gnomon summary function and Python statistics to generate pinned conformance fixtures consumed by Node tests. That checks these narrow shared contracts and arithmetic, not full Gnomon equivalence.
18
+
19
+ The Node implementation is deliberately narrower: no repairs, historical replay, leakage checks, revision tracking, evidence memory, scores, or scheduler. Next-phase outcome/review work should integrate Gnomon rather than build a second ledger. A shared language-neutral schema package or versioned hosted API is a future integration decision, not an implemented capability.
20
+
21
+ ## Version 1 local records
22
+
23
+ | Kind | Purpose | Contains history |
24
+ | --- | --- | --- |
25
+ | `ephemeris.data` | Frozen selected rows plus inspection evidence | Yes, selected values/timestamps; invalid/missing cells are null with diagnostic counts |
26
+ | `ephemeris.request` | Effective API request plus local context | Yes, all values to submit |
27
+ | `ephemeris.forecast` | Paid-result receipt reusable without inference | Yes, exact submitted request and complete saved response |
28
+
29
+ Each record has `schema_version: "1"`. Machine schemas are bundled and available from `schema forecast --kind data|prepared|artifact`; no credentials/network are needed. The default `--kind request` prints the public API request schema. Schemas are generated by `scripts/generate-schemas.js` and checked in CI.
30
+
31
+ A forecast record contains:
32
+
33
+ - `artifact_id`: local UUID, not a Gnomon execution ID.
34
+ - `created_at`: local receipt creation time, not source availability or an observation timestamp.
35
+ - `request`: effective public wire request; local context never enters it.
36
+ - `request_sha256`: SHA-256 of recursively key-sorted compact JSON. This detects request mutation, not response authenticity or authorship.
37
+ - `response`: API JSON including quantiles, metadata, IDs, billing. It is saved before shape interpretation so a malformed paid response remains recoverable. Credential echoes are redacted and disclosed in transport metadata.
38
+ - `context.series`: metadata aligned to request series, including original history timestamps, a derived/explicit future axis, caller units/measurement/timezone/target, and known-future/scenario role labels.
39
+ - `context.scenario_assumptions`: caller assertions, not model input or validated causes.
40
+ - `transport`: chosen service origin, idempotency key, request/replay headers, HTTP status. No bearer header or API key.
41
+ - `disclosure`: explicitly says submitted history is included and model/service context may be truncated.
42
+ - `interoperability`: explicitly marks Gnomon import and ledger recording as absent.
43
+
44
+ Response shape validation happens on describe/plot. Null numeric results stay null. Extra quantiles are allowed; duplicate numeric keys, invalid levels, inconsistent variate counts, and horizon mismatch are rejected. Crossed quantiles are reported rather than silently rearranged. Unknown metadata is null, never inferred from price-like values or filenames.
45
+
46
+ For Gnomon interoperability later, the artifact preserves values, order, labels, units, timestamps when known, frequency, declared availability cutoffs, local receipt time, source snapshot identity, request fingerprint, service IDs, and role distinctions. Date-only inputs and step-only forecasts may need additional time semantics before a Gnomon import could be valid. `source_available_at`/`recorded_as_of` are caller metadata, not proof of leakage safety. No import function, compatible ledger identity, automatic outcome ingestion, or scheduling is claimed.
47
+
48
+ ## Updating contracts
49
+
50
+ 1. Compare the bundled OpenAPI with a current checked-out backend using `npm run check:api -- PATH`.
51
+ 2. Review changes against backend semantic validation, including limits/defaults and billing/retry behaviour. Schema equality alone cannot detect all implementation drift.
52
+ 3. Update runtime validation and help together. Regenerate schemas and reference docs.
53
+ 4. If Gnomon changes, run `python3 scripts/gnomon-fixtures.py /path/to/Gnomon --check`, review differences, and explicitly regenerate the pinned fixture.
54
+ 5. Run tests, artifact checks, packaging, and actual agent acceptance trials before release claims.
55
+
56
+ Artifact hashes are local integrity checks, not security boundaries. Inspecting a file does not prove its provenance or that its values measure the target the user intended.
57
+
58
+ ## Submission receipts and image export (0.3)
59
+
60
+ The `ephemeris.submission` v1 contract stores the prepared request, request SHA-256, original API origin, idempotency key, and creation time before submission. Its immutable `completion_unknown` status does not claim server completion. `forecast run --resume` reads this record, validates it, and resends the same request/key to the saved origin without automatic retries. It has no status-query or Gnomon import endpoint. Receipt creation is an explicit documented side effect of paid submission, including when result output is stdout. Receipts contain history and local context, never credentials, and remain until explicitly removed. A write failure prevents submission.
61
+
62
+ The request hash checks accidental request modification, not receipt authenticity; only trusted receipts should be resumed. Server retention rules still govern replay. PNG and SVG are presentation outputs of the saved artifact, never provider calls or tracking records. The default recent-history view discloses its window and does not mutate the source history. PNG uses a packaged renderer and licensed bundled font; no separately installed plotting tool is needed.
@@ -0,0 +1,23 @@
1
+ # Release operations
2
+
3
+ The hosted MCP lives in paracast-web and deploys through the existing Vercel Git integration. The npm ephemeris-mcp package is a separate stdio bridge; the dedicated ephemeris-cli package is another independent distribution. A hosted MCP version change does not require either npm package to be published.
4
+
5
+ ## Hosted MCP
6
+
7
+ Open a PR, require the Application checks / test check, and merge only after preview deployment and checks pass. CI covers tests, TypeScript, shared CLI/API contract drift, tool documentation, production build and PNG renderer/font tracing. Vercel deploys main. The Production MCP smoke workflow runs after successful production deployments, daily, and manually. It checks the production custom domain using the official SDK: version, six tool schemas, and HTTP 401 for a keyless saved-artifact call. It never submits a paid forecast and does not verify authenticated rendering or forecast quality.
8
+
9
+ A failed smoke check alerts through GitHub Actions; it does not automatically roll back. Investigate logs without exposing credentials. If the release is faulty, promote the last known good deployment through Vercel or revert the PR through the same checks. Re-run Production MCP smoke after recovery. Do not use new paid forecasts as a health probe.
10
+
11
+ ## CLI and bridge packages
12
+
13
+ Package CI installs locked dependencies, runs its checks, and saves npm tarballs as workflow artifacts. CLI checks exercise the installed tarball, including PNG rendering against a synthetic local service, on Linux and macOS with Node 22 and 24. Windows is not yet release-certified. Bridge CI checks supported Node versions and live keyless tool discovery through the stdio bridge.
14
+
15
+ npm publication remains separate. No publish token is stored or publication workflow enabled by these changes. Before first CLI publication, confirm npm name ownership, licensing and distribution visibility, configure the registry's trusted publisher for an explicitly approved release workflow, and test installation of the exact candidate tarball. For subsequent releases bump package and lockfile versions, run all checks, publish the verified candidate under the intended tag, then smoke its installed executable. Do not reuse a published version. Roll package consumers back to a previous known-good version if needed; a hosted rollback is independent of npm versions.
16
+
17
+ The CLI repository is initially private; source access does not imply public npm availability. Web documentation must not advertise an npm installation until it exists. Current CLI source installation uses `npm install --global .` from an authorized checkout.
18
+
19
+ ## Documentation and client refresh
20
+
21
+ Keep web docs, llms.txt, llms-full.txt, the bridge README and both bundled forecasting skills aligned. The CLI reference and schemas are generated and checked in its own CI. Archived npm READMEs and published skill catalog entries do not change solely because GitHub changes; refresh them through their own release process when needed.
22
+
23
+ Customers using hosted MCP should reconnect or start a fresh conversation to refresh cached tools. In HeyDitto, Settings → MCP Servers exposes tool discovery. Its published guide currently says SSE-only, so do not promise compatibility with Streamable HTTP without testing the actual connection. PNG/SVG attachment display also depends on the host.
@@ -0,0 +1,33 @@
1
+ # Agent usability audit and corrections — 0.3
2
+
3
+ The audit found that passing scripted workflows did not establish unfamiliar-agent usability. This release addresses the concrete failures and retains actual agent results separately from integration checks.
4
+
5
+ | Finding | Change | Evidence |
6
+ | --- | --- | --- |
7
+ | Retry key existed only in terminal output before completion | Private, fsynced `ephemeris.submission` v1 receipt before the paid POST; `forecast run --resume` reuses the exact origin/request/key | Dropped-response test discards diagnostics, changes original input, and recovers with one execution; receipt write failure prevents any paid call |
8
+ | Long history squeezed the forecast into about 1.3 pixels | Default view shows recent max(24, 2*horizon) observations with visible shown/total counts; `--history N` and `--history all` are explicit | 16,384-observation/24-step regression keeps the forecast wider than 250 pixels and leaves artifact history unchanged |
9
+ | Only SVG export | Built-in PNG export through a packaged renderer and bundled font; SVG retained; format/extension conflicts rejected | PNG signature/dimensions, binary stdout, no overwrite, package installation, and visual review |
10
+ | Help examples depended on absent files or mismatched context | Complete synthetic examples for run, describe, and plot; matching request/context grids | Every command example executes in an empty temporary directory against a local fixture |
11
+ | Argument errors did not identify the problem | Safe option names, typo suggestions, expected values, and next actions; raw values never echoed | Misspelled help, missing value, out-of-range values, and credential-like input checks |
12
+ | Scripted tests were the only evidence of discovery | Actual fresh-session runner, tool traces, final explanations, model/version/token metrics, source fingerprint and qualitative review | See evaluation evidence below |
13
+
14
+ The first development agent trial completed seven mechanical tasks but failed qualitative review on billing units and absent-future-input interpretation. Help now explicitly distinguishes millicredits from credits and “not supplied” from “assumed absent.” Description includes exact decimal billing conversions and the saved-response source pointer. It also states that discontinuities or constant-width bands alone do not establish forecast quality. The failing development evidence remains in `evals/results/development-trial.json` and `development-review.json`.
15
+
16
+ ## Verification boundaries
17
+
18
+ Unit/integration tests use synthetic observations and local fixture services. They do not establish live forecasting quality. Renderer/package checks were performed on this Linux environment; CI also declares Node 22/24 checks, but this audit does not claim other operating systems were exercised.
19
+
20
+ Receipts preserve history and remain on disk until explicitly removed. They are integrity-checked records, not authenticity signatures or proof of server completion. Server idempotency retention still governs replay; indefinite replay without another charge is not promised.
21
+
22
+ The actual agent runner starts fresh sessions with shell access and an instruction to learn through help/schema, without preloading source/docs or skills. It is not a security sandbox. Review traces and explanations as well as the mechanical grade. Passing sampled tasks on one runner/model does not establish universal usability.
23
+
24
+ ## Evaluation evidence
25
+
26
+ - **40 unit/integration tests passed**, including durable receipt recovery, no-charge receipt failures, exact billing conversion, private files, PNG export, long-history layout, and execution of every command's help examples.
27
+ - **9/9 scripted scenarios passed.** These exercise the harness and are not LLM evaluations.
28
+ - **9/9 actual fresh-agent tasks completed mechanically** on Claude Code 2.1.295, model identifier `claude-opus-5-5` as reported by the runner. Tasks took 3–6 tool calls in this sample. Runtime sources stayed unchanged during the trial.
29
+ - Qualitative review accepted eight task explanations within their tested scope and found one incorrect claim that missing labels require paid inference. Help was corrected; **2/2 fresh plotting followups passed**, and both agents explicitly explained that caller-supplied labels can be added to a copy and redrawn offline without a new forecast.
30
+
31
+ See [the nine-task report](../evals/results/acceptance-trial.json), [targeted followup](../evals/results/metadata-followup-trial.json), and [qualitative review](../evals/results/acceptance-review.json). Reports retain source hashes, actual model/runner metadata, tool traces, explanations, and metrics. The review was performed by the implementation assistant, not an independent human evaluator. The full suite and followup have different recorded hashes; the full suite was not repeated after the final help-only clarification.
32
+
33
+ This is sampled acceptance evidence for one runner/model, not a universal success rate. Some natural-language imprecision remains (for example, describing the boundary as “at step 0” rather than between the last observed and first forecast steps). The saved artifact and deterministic description remain the authoritative numeric record. No package publication or live paid forecast was performed.
@@ -0,0 +1,15 @@
1
+ date,sales
2
+ 2026-01-01,100
3
+ 2026-01-02,102
4
+ 2026-01-03,97
5
+ 2026-01-04,106
6
+ 2026-01-05,108
7
+ 2026-01-06,109
8
+ 2026-01-07,112
9
+ 2026-01-08,107
10
+ 2026-01-09,114
11
+ 2026-01-10,116
12
+ 2026-01-11,113
13
+ 2026-01-12,119
14
+ 2026-01-13,121
15
+ 2026-01-14,120
@@ -0,0 +1,8 @@
1
+ {
2
+ "mode": "route",
3
+ "series": [
4
+ { "values": [100, 101, 99, 103, 104, 102, 105, 107, 106, 108, 110, 109], "freq": "H" }
5
+ ],
6
+ "horizon": 24,
7
+ "quantiles": [0.1, 0.5, 0.9]
8
+ }
package/package.json CHANGED
@@ -1,6 +1,49 @@
1
1
  {
2
2
  "name": "ephemeris-cli",
3
- "version": "0.0.0-stage",
4
- "stub": true,
5
- "description": "Temporary package placeholder for staged publishing"
6
- }
3
+ "version": "0.4.0",
4
+ "description": "Agent-friendly CLI for Ephemeris probabilistic time-series forecasts",
5
+ "type": "module",
6
+ "bin": {
7
+ "ephemeris": "bin/ephemeris.js"
8
+ },
9
+ "engines": {
10
+ "node": ">=22"
11
+ },
12
+ "files": [
13
+ "bin/",
14
+ "src/",
15
+ "schema/",
16
+ "examples/",
17
+ "docs/",
18
+ "assets/",
19
+ "build-info.json"
20
+ ],
21
+ "scripts": {
22
+ "test": "node --test",
23
+ "docs": "node scripts/generate-docs.js",
24
+ "docs:check": "node scripts/generate-docs.js --check",
25
+ "demo": "node scripts/demo.js",
26
+ "eval:smoke": "node scripts/eval-smoke.js",
27
+ "schemas": "node scripts/generate-schemas.js",
28
+ "schemas:check": "node scripts/generate-schemas.js --check",
29
+ "check:api": "node scripts/check-api.js",
30
+ "package:check": "node scripts/check-package.js",
31
+ "prepack": "node scripts/stamp-build.js"
32
+ },
33
+ "keywords": [
34
+ "forecasting",
35
+ "time-series",
36
+ "cli",
37
+ "agents",
38
+ "ephemeris"
39
+ ],
40
+ "homepage": "https://ephemeris.cascade.industries",
41
+ "dependencies": {
42
+ "@resvg/resvg-js": "2.6.2"
43
+ },
44
+ "repository": {
45
+ "type": "git",
46
+ "url": "git+https://github.com/TensorLink-AI/ephemeris-cli.git"
47
+ },
48
+ "license": "UNLICENSED"
49
+ }