@mercury-fw/cli 0.30.0 → 0.31.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,27 @@
1
1
  # @mercury-fw/cli
2
2
 
3
+ ## 0.31.1
4
+
5
+ ### Patch Changes
6
+
7
+ - Updated dependencies [7926aa2]
8
+ - @mercury-fw/core@0.31.1
9
+
10
+ ## 0.31.0
11
+
12
+ ### Minor Changes
13
+
14
+ - 6ae8133: - `mfw local-packages <folder>` makes an app install `@mercury-fw/*` packages from local tarballs (`bun pm pack`) instead of the registry, transitive dependencies included; `--off` goes back to the registry.
15
+ - `mfw e2e` runs end-to-end tests against an app's real model through its REPL: cases of turns with checks on the tool calls and the answer, repeated runs, results kept in `e2e/results/`.
16
+ - Test files declare their cases with `e2e()` from `@mercury-fw/cli/e2e`.
17
+ - A new app's Dockerfile copies `.packs/` when present, and its `.gitignore` leaves out `.packs/` and `e2e/results/`.
18
+ - In the REPL, `/dump` after a confirmation no longer writes the turn before it.
19
+
20
+ ### Patch Changes
21
+
22
+ - Updated dependencies [6ae8133]
23
+ - @mercury-fw/core@0.31.0
24
+
3
25
  ## 0.30.0
4
26
 
5
27
  ### Minor Changes
package/README.md CHANGED
@@ -20,6 +20,8 @@
20
20
  - [`mfw credentials set <plugin> [--from <dir>] [--print]`](#mfw-credentials-set-plugin---from-dir---print)
21
21
  - [`mfw credentials reset <plugin>`](#mfw-credentials-reset-plugin)
22
22
  - [`mfw google-chat set-key <key-file> [--subscription <name>]`](#mfw-google-chat-set-key-key-file---subscription-name)
23
+ - [`mfw local-packages <folder>` / `mfw local-packages --off`](#mfw-local-packages-folder--mfw-local-packages---off)
24
+ - [`mfw e2e [tests...] [--repeat N]`](#mfw-e2e-tests---repeat-n)
23
25
  - [`mfw upgrade`](#mfw-upgrade)
24
26
  - [Help](#help)
25
27
 
@@ -131,7 +133,7 @@ Maintains the wiki vault, which lives on a Docker volume and not in the app's fo
131
133
  |---|---|
132
134
  | `list` | Every note. |
133
135
  | `read <path>` | One note (a path that isn't one says so, exit 1). |
134
- | `grep <pattern>` | Every line matching `<pattern>`, a regular expression, as `path:line:text`. A pattern starting with `-` goes after `--` (`mfw vault grep -- -h`). |
136
+ | `grep <pattern>` | Every line matching `<pattern>`, a regular expression (case ignored), as `path:line:text`. A pattern starting with `-` goes after `--` (`mfw vault grep -- -h`). |
135
137
  | `write-curated <path> [--author NAME]` | Writes a curated note, the body read from stdin. |
136
138
  | `write-raw <path>` | Writes raw material for the nightly review to triage, the body read from stdin. |
137
139
 
@@ -212,6 +214,88 @@ mfw google-chat set-key key.json --subscription projects/my-project/subscription
212
214
  mfw google-chat set-key new-key.json
213
215
  ```
214
216
 
217
+ ### `mfw local-packages <folder>` / `mfw local-packages --off`
218
+
219
+ Makes the app install `@mercury-fw/*` packages from local tarballs instead of the registry: the way to try a framework or plugin change in a real app before it's published. `<folder>` holds the `.tgz` files `bun pm pack` writes, one per package; the command copies them into the app's `.packs/`, points each package at its tarball with `overrides` in `package.json`, and runs `bun install`.
220
+
221
+ Overrides, and not the dependencies themselves, because a packed package names the packages it depends on by version, and the registry has those versions too: only an override sends them to the tarballs as well (a packed `@mercury-fw/core` depends on `@mercury-fw/plugin-types`, for example). The Dockerfile `mfw create` writes copies `.packs/` before installing, so the image gets the same packages (an app created before 0.31.0 needs `COPY --chown=mercury:mercury .pack[s] ./.packs/` added after the line copying `package.json`); `.gitignore` leaves it out of the repository.
222
+
223
+ Run it again after packing anew: it replaces the tarballs and the overrides of the previous run, and leaves alone the overrides you wrote yourself. `--off` takes the app back to the registry: the overrides it wrote and `.packs/` go, then `bun install`. After either, `mfw start` rebuilds the image with the packages now installed.
224
+
225
+ ```bash
226
+ mfw local-packages ../mercury-fw/apps/testbed/.packs
227
+ mfw local-packages --off
228
+ ```
229
+
230
+ ### `mfw e2e [tests...] [--repeat N]`
231
+
232
+ Runs end-to-end tests against the app's real model: each test case sends its turns to the app's REPL in the container, as `mfw repl` would, and checks what each turn did, the tool calls with their inputs and results, and the answer. It's how you find out whether the model actually uses a plugin the way its skill says, at the first try, with the app's own model, configuration, wiki and credentials. That's also why it doesn't belong in CI: the results depend on the model, on what's in the vault and in Qdrant, and on accounts a CI runner shouldn't have.
233
+
234
+ Without arguments it runs every `e2e/*.e2e.ts` in the app; with files, those (they can live anywhere). It prints every check of every run, keeps everything (each turn's calls and answer, each check) in `e2e/results/<time>/`, which `.gitignore` leaves out, and exits 1 when a case didn't pass. The app must be built and its `.env` filled in, as for `mfw repl`. `--repeat N` runs each case N times, overriding its own `repeat`.
235
+
236
+ #### Writing a test
237
+
238
+ A test is a TypeScript file whose default export is `e2e({ … })`, from `@mercury-fw/cli/e2e` (every app has the CLI among its devDependencies, so the types are there):
239
+
240
+ ```ts
241
+ import { e2e } from "@mercury-fw/cli/e2e";
242
+
243
+ export default e2e({
244
+ // What the app must have, by catalog id: checked before anything runs.
245
+ plugins: ["jira"],
246
+ channels: [],
247
+ cases: [
248
+ {
249
+ name: "project key, at the first try",
250
+ turns: ["On Jira, what is the project key of Customer Support?"],
251
+ repeat: 3, // the model isn't deterministic
252
+ minPasses: 3, // how many runs must pass; default every one
253
+ check: (run, expect) => {
254
+ expect.everyCall("jiraCommand", (c) => String((c.input as { command: string }).command).includes("--select "));
255
+ expect.noFailedCalls();
256
+ expect.callCount({ max: 3 });
257
+ expect.answer(/\bCS\b/);
258
+ },
259
+ },
260
+ ],
261
+ });
262
+ ```
263
+
264
+ Each run of a case gets a fresh REPL session. `check` receives the run, `run.turns` in order and `run.last`, each turn with:
265
+
266
+ - `calls`: the tool calls, each with `tool`, `input`, `output`, `ok` (false for a failed call, or one that never got a result) and `pending` (an irreversible command staged for confirmation: it worked, and its output holds the token);
267
+ - `answer`: the turn's final text (what the model wrote; a list `present` adds afterwards isn't part of it);
268
+ - `seconds`: how long it took, for the report.
269
+
270
+ The `expect` helpers record a named check each and never stop the others, so a failed run shows every check that failed:
271
+
272
+ | Helper | Passes when |
273
+ |---|---|
274
+ | `call(tool, match?)` | at least one call to `tool` (matching `match`, when given) |
275
+ | `everyCall(tool, match)` | there are calls to `tool`, and every one matches |
276
+ | `noFailedCalls()` | no call failed |
277
+ | `callCount({ min?, max? }, tool?)` | the number of calls, to any tool or to `tool`, is within the bounds |
278
+ | `answer(text \| regex)` | the last answer contains the text, or matches |
279
+ | `answerNot(text \| regex)` | it doesn't |
280
+ | `that(label, condition)` | `condition` is true |
281
+
282
+ The call helpers look at every turn of the run, the answer helpers at the last one; each takes an optional label as its last argument. A case that records no check fails: it proves nothing.
283
+
284
+ A turn can be a function of the one before, for a follow-up or a confirmation: `(previous) => …` returns the next message from `previous.calls` and `previous.answer` (an irreversible command comes back as a call whose output holds the pending confirmation's token, and sending the token as the next turn confirms it). Each turn is one line, as the REPL reads them.
285
+
286
+ `before`, `after` and `check` also get `cli(command)`, which runs `sh -c command` in the app's container, outside the model, and resolves to its exit code and output: to prepare data (an issue to work on, a note in the wiki with the vault CLI), to check what a turn changed (`jira issue get KEY --select fields.assignee.displayName` after an assignment), to clean up. `after` runs even when the run or the check failed. A test that changes an external system has to clean up after itself, so start from read-only ones.
287
+
288
+ #### What it can and can't test
289
+
290
+ It can test how the model uses a plugin through its skill (the commands and flags it picks, rejected or failed calls, how many calls it takes), what the answer says or must not say, what a turn changed in an external system, the confirmation of an irreversible command, conversations over several turns, several plugins working together, and what the wiki and memory make of a conversation.
291
+
292
+ It can't test how a channel shows a turn (Google Chat's cards, the HTTP channel's events: the REPL runs the same turn, without them), the background jobs (the nightly wiki review, idle-session capture), or the exact wording of an answer: checks look for properties, and `repeat` with `minPasses` says how steady the behaviour is. Durations are in the report and in `e2e/results/`, never a check.
293
+
294
+ ```bash
295
+ mfw e2e
296
+ mfw e2e e2e/jira.e2e.ts --repeat 3
297
+ ```
298
+
215
299
  ### `mfw upgrade`
216
300
 
217
301
  Installs the registry's latest `mfw` globally (`bun add -g @mercury-fw/cli@<latest>`) when it's newer than the one running, and says it's already the latest otherwise; a registry that doesn't answer exits 1. It's about the global `mfw` only: an app's framework, its own CLI included, moves with `bun update` in the app. `mfw create` run by a stale global `mfw` creates the app with the latest one anyway, and says to run this.
@@ -74,6 +74,19 @@ export declare function appCommands(app: App, deps: AppDeps): {
74
74
  googleChatSetKey: (keyFile: string, { subscription }: {
75
75
  subscription?: string;
76
76
  }) => Promise<number>;
77
+ /** Makes the app install the packages in `from`'s tarballs (`bun pm
78
+ * pack`) instead of the registry's: copied into `.packs/` (which the
79
+ * image copies too), overridden in the manifest, then installed.
80
+ * Everything is checked before anything is written. */
81
+ localPackages: (from: string) => Promise<number>;
82
+ /** Runs e2e tests (`tests`, or the app's `e2e/*.e2e.ts`) against the app's
83
+ * REPL in its container, keeping the turns and checks in
84
+ * `e2e/results/<time>/`; returns 1 when a case didn't pass. */
85
+ e2e: (tests: string[], { repeat }: {
86
+ repeat?: number;
87
+ }) => Promise<number>;
88
+ /** Undoes `localPackages`: the app installs from the registry again. */
89
+ localPackagesOff: () => Promise<number>;
77
90
  };
78
91
  /** The real deps: docker on the user's terminal, questions on `input`
79
92
  * (stdin by default). */
@@ -0,0 +1,26 @@
1
+ /**
2
+ * What `mfw local-packages` writes: the app's `package.json` with `overrides`
3
+ * pointing packages at local tarballs (`bun pm pack`) copied into `.packs/`,
4
+ * which the image copies before `bun install`. Overrides, and not the
5
+ * dependencies themselves, because a packed package names its siblings by
6
+ * version, which the registry has too: only an override sends those
7
+ * transitive dependencies to the tarballs as well.
8
+ */
9
+ /** The tarballs' folder inside the app, as the Dockerfile copies it. */
10
+ export declare const LOCAL_PACKS_DIR = ".packs";
11
+ type Manifest = Record<string, unknown> & {
12
+ overrides?: Record<string, string>;
13
+ };
14
+ /** `pkg` with each of `packs` (a package name and its tarball's file name in
15
+ * `.packs/`) as an override, in place of any left by an earlier run. */
16
+ export declare function withLocalOverrides(pkg: Manifest, packs: Array<{
17
+ name: string;
18
+ file: string;
19
+ }>): Manifest;
20
+ /** `pkg` without the overrides this command writes; without `overrides` at
21
+ * all when none of the user's own is left. */
22
+ export declare function withoutLocalOverrides(pkg: Manifest): Manifest;
23
+ /** The package name in a `bun pm pack` tarball, read from its
24
+ * `package/package.json`. */
25
+ export declare function packageNameOf(tarball: string): Promise<string>;
26
+ export {};
@@ -0,0 +1,73 @@
1
+ /**
2
+ * The shape of an e2e test, as a test file imports it
3
+ * (`import { e2e } from "@mercury-fw/cli/e2e"`): the plugins and channels the
4
+ * app must have, and cases of turns sent to the app's real model through its
5
+ * REPL, each with checks on the calls it made and the answer it gave.
6
+ * `mfw e2e` runs them (`runner.ts`).
7
+ */
8
+ import type { Call, TurnData } from "./dump.ts";
9
+ export type { Call } from "./dump.ts";
10
+ /** One turn as a check sees it: its calls, its answer, how long it took. */
11
+ export type Turn = TurnData & {
12
+ seconds: number;
13
+ };
14
+ /** A case's run: every turn in order, and the last one. */
15
+ export type Run = {
16
+ turns: Turn[];
17
+ last: Turn;
18
+ };
19
+ /** Runs `command` in the app's container (`sh -c`), outside the model:
20
+ * preparing data, reading what a turn changed, cleaning up. */
21
+ export type Cli = (command: string) => Promise<{
22
+ code: number;
23
+ output: string;
24
+ }>;
25
+ /** What `before`, `after` and `check` can reach besides the run. */
26
+ export type Context = {
27
+ cli: Cli;
28
+ };
29
+ /** The checks a case makes. Call helpers look at every turn's calls, answer
30
+ * helpers at the last answer; each records a named check and never throws,
31
+ * so a case reports every failure at once. */
32
+ export type Expect = {
33
+ /** At least one call to `tool`, matching `match` when given. */
34
+ call(tool: string, match?: (call: Call) => boolean, label?: string): void;
35
+ /** Every call to `tool` matches `match` (and there is at least one). */
36
+ everyCall(tool: string, match: (call: Call) => boolean, label?: string): void;
37
+ /** No call failed or went without a result. */
38
+ noFailedCalls(label?: string): void;
39
+ /** The number of calls, to every tool or to `tool`, within the bounds. */
40
+ callCount(bounds: {
41
+ min?: number;
42
+ max?: number;
43
+ }, tool?: string, label?: string): void;
44
+ /** The answer contains `pattern`, or matches it. */
45
+ answer(pattern: string | RegExp, label?: string): void;
46
+ /** The answer doesn't contain `pattern`, or doesn't match it. */
47
+ answerNot(pattern: string | RegExp, label?: string): void;
48
+ /** Anything else. */
49
+ that(label: string, condition: boolean): void;
50
+ };
51
+ /** One case: the turns sent, in one REPL session, and the checks on them. */
52
+ export type E2eCase = {
53
+ name: string;
54
+ /** A message, or a function of the turn before (a follow-up, a confirmation token). */
55
+ turns: Array<string | ((previous: Turn) => string)>;
56
+ /** How many times to run it (the model isn't deterministic); default 1. */
57
+ repeat?: number;
58
+ /** How many runs must pass; default every one. */
59
+ minPasses?: number;
60
+ before?: (ctx: Context) => unknown;
61
+ after?: (ctx: Context) => unknown;
62
+ check: (run: Run, expect: Expect, ctx: Context) => unknown;
63
+ };
64
+ /** An e2e test: what the app must have, and its cases. */
65
+ export type E2eTest = {
66
+ /** Catalog ids of the tool plugins the app must have (`jira`, …). */
67
+ plugins?: string[];
68
+ /** Catalog ids of the channels the app must have (`http`, …). */
69
+ channels?: string[];
70
+ cases: E2eCase[];
71
+ };
72
+ /** Declares an e2e test: the identity, there for the types. */
73
+ export declare function e2e(test: E2eTest): E2eTest;
@@ -0,0 +1,24 @@
1
+ /**
2
+ * A turn read out of the file the REPL's `/dump` writes: the AI SDK's step
3
+ * results for the last turn, whose `content` parts are the tool calls, their
4
+ * results (or errors) and the model's text. Read as data, not through the
5
+ * SDK's types: only the parts listed here matter.
6
+ */
7
+ /** One tool call as a check sees it. `ok` is false when the call failed or
8
+ * never got a result; a result without an `ok` of its own counts as worked.
9
+ * `pending` is an irreversible command staged for confirmation, which worked
10
+ * (its output has the token) though it reports `ok: false`. */
11
+ export type Call = {
12
+ tool: string;
13
+ input: unknown;
14
+ output: unknown;
15
+ ok: boolean;
16
+ pending: boolean;
17
+ };
18
+ /** What a turn did: its tool calls in order, and its final text. */
19
+ export type TurnData = {
20
+ calls: Call[];
21
+ answer: string;
22
+ };
23
+ /** The calls and the answer in `dump`, the parsed content of a `/dump` file. */
24
+ export declare function turnFromDump(dump: unknown): TurnData;
@@ -0,0 +1,12 @@
1
+ import type { Expect, Run } from "./define.ts";
2
+ /** One check's outcome; `detail` says what was found when it failed. */
3
+ export type Check = {
4
+ label: string;
5
+ ok: boolean;
6
+ detail?: string;
7
+ };
8
+ /** An `expect` bound to `run`, and the checks it recorded so far. */
9
+ export declare function createExpect(run: Run): {
10
+ expect: Expect;
11
+ checks: () => Check[];
12
+ };
@@ -0,0 +1,9 @@
1
+ import type { E2eTest } from "./define.ts";
2
+ /** The files to run: `named` resolved from `cwd`, or, when none is named,
3
+ * the app's `e2e/*.e2e.ts` in name order. */
4
+ export declare function findTests(named: string[], { appDir, cwd }: {
5
+ appDir: string;
6
+ cwd: string;
7
+ }): string[];
8
+ /** The test `file` exports as default; throws naming the file when it isn't one. */
9
+ export declare function loadTest(file: string): Promise<E2eTest>;
@@ -0,0 +1,42 @@
1
+ import type { Context, E2eTest, Turn } from "./define.ts";
2
+ import { type Check } from "./expect.ts";
3
+ /** A REPL session in the app: `turn` sends one line and resolves with what
4
+ * `/dump` wrote for it and what the REPL printed meanwhile. */
5
+ export type Session = {
6
+ turn: (line: string) => Promise<{
7
+ dump: unknown;
8
+ output: string;
9
+ }>;
10
+ close: () => Promise<void>;
11
+ };
12
+ export type RunnerDeps = {
13
+ openSession: () => Promise<Session>;
14
+ cli: Context["cli"];
15
+ /** The app's dependencies, to check the test's plugins and channels against. */
16
+ appPackages: Record<string, string>;
17
+ print: (line: string) => void;
18
+ /** Milliseconds, for each turn's duration. */
19
+ now: () => number;
20
+ writeReport: (report: CaseReport[]) => Promise<void>;
21
+ };
22
+ /** One run of a case, as the report keeps it. */
23
+ export type RunReport = {
24
+ ok: boolean;
25
+ turns: Turn[];
26
+ checks: Check[];
27
+ };
28
+ export type CaseReport = {
29
+ file: string;
30
+ case: string;
31
+ passed: boolean;
32
+ runs: RunReport[];
33
+ };
34
+ /** Runs the tests in `tests` (each with the file it came from); returns the
35
+ * exit code: 0 when every case passed enough runs. `repeat` overrides each
36
+ * case's own. */
37
+ export declare function runE2e(tests: Array<{
38
+ file: string;
39
+ test: E2eTest;
40
+ }>, opts: {
41
+ repeat?: number;
42
+ }, deps: RunnerDeps): Promise<number>;
@@ -0,0 +1,17 @@
1
+ import type { Session } from "./runner.ts";
2
+ export type ReplSessionOptions = {
3
+ /** The command that starts the REPL. */
4
+ argv: string[];
5
+ cwd: string;
6
+ /** Where the dumps are, on this side, and as the REPL sees the same folder. */
7
+ hostDir: string;
8
+ replDir: string;
9
+ /** Prefix of this session's dump files, unique among sessions sharing the folder. */
10
+ name: string;
11
+ /** How long one turn may take. */
12
+ timeoutMs: number;
13
+ };
14
+ /** Starts the REPL and returns the session over it once the REPL is ready
15
+ * (its first prompt is out), so a turn's time is only the turn's; throws
16
+ * when the REPL exits or doesn't get there in time. */
17
+ export declare function openReplSession(opts: ReplSessionOptions): Promise<Session>;
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@mercury-fw/cli",
3
- "version": "0.30.0",
3
+ "version": "0.31.1",
4
4
  "type": "module",
5
5
  "license": "MIT",
6
6
  "repository": {
@@ -27,6 +27,11 @@
27
27
  "mercury-fw-source": "./src/main.ts",
28
28
  "types": "./dist/src/main.d.ts",
29
29
  "default": "./src/main.ts"
30
+ },
31
+ "./e2e": {
32
+ "mercury-fw-source": "./src/e2e/define.ts",
33
+ "types": "./dist/src/e2e/define.d.ts",
34
+ "default": "./src/e2e/define.ts"
30
35
  }
31
36
  },
32
37
  "scripts": {
@@ -36,16 +41,16 @@
36
41
  "comment:shape": "The Mercury CLI (`mfw`). `mfw create <dir>` writes a new Mercury app from the template. The framework packages move in lockstep with it, so a new app gets them at the CLI's own version; plugins and channels at the registry's latest. The plugin and channel packages are devDependencies only for the catalog test. `main(argv)` is exported for create-mercury-agent.",
37
42
  "dependencies": {
38
43
  "@clack/prompts": "^1.8.1",
39
- "@mercury-fw/core": "0.30.0",
44
+ "@mercury-fw/core": "0.31.1",
40
45
  "commander": "^15.0.0"
41
46
  },
42
47
  "devDependencies": {
43
48
  "@mercury-fw/channel-google-chat": "0.1.3",
44
- "@mercury-fw/channel-http": "0.1.0",
45
- "@mercury-fw/formatter": "0.30.0",
49
+ "@mercury-fw/channel-http": "0.1.1",
50
+ "@mercury-fw/formatter": "0.31.1",
46
51
  "@mercury-fw/plugin-atlassian-admin": "0.1.0",
47
52
  "@mercury-fw/plugin-bitbucket": "0.1.0",
48
- "@mercury-fw/plugin-jira": "0.3.0",
53
+ "@mercury-fw/plugin-jira": "0.4.1",
49
54
  "@mercury-fw/typescript-config": "*",
50
55
  "@types/bun": "^1.4.2",
51
56
  "typescript": "^6.0.3"
@@ -5,12 +5,16 @@
5
5
  * and validated before any of this runs (`program.ts`); what runs a command is
6
6
  * injected (`AppDeps`), which is how the tests see the exact calls.
7
7
  */
8
- import { readFileSync } from "node:fs";
8
+ import { chmodSync, cpSync, existsSync, mkdirSync, readdirSync, readFileSync, rmSync, writeFileSync } from "node:fs";
9
9
  import { homedir } from "node:os";
10
- import { join, resolve } from "node:path";
10
+ import { join, relative, resolve } from "node:path";
11
11
  import { CATALOG, type CliCredentials } from "../catalog.ts";
12
12
  import { packCredentials, readServiceAccountKey, setEnvVar } from "./credentials.ts";
13
13
  import type { App } from "./find-app.ts";
14
+ import { LOCAL_PACKS_DIR, packageNameOf, withLocalOverrides, withoutLocalOverrides } from "./local-packages.ts";
15
+ import { findTests, loadTest } from "../e2e/load.ts";
16
+ import { runE2e } from "../e2e/runner.ts";
17
+ import { openReplSession } from "../e2e/session.ts";
14
18
 
15
19
  export type AppDeps = {
16
20
  /** Runs `argv` in `cwd` on the user's terminal (stdin, stdout, stderr) and returns its exit code. */
@@ -48,6 +52,16 @@ export const RESET_TARGETS = {
48
52
  } as const;
49
53
  export type ResetTarget = keyof typeof RESET_TARGETS;
50
54
 
55
+ /** How long one e2e turn may take: a local model on a long, tool-heavy turn is slow. */
56
+ const E2E_TURN_TIMEOUT_MS = 10 * 60_000;
57
+
58
+ /** Runs `argv` in `cwd`, returning its exit code and its output (stdout and stderr together). */
59
+ async function captureCode(argv: string[], cwd: string): Promise<{ code: number; output: string }> {
60
+ const proc = Bun.spawn(argv, { cwd, stdin: "ignore", stdout: "pipe", stderr: "pipe" });
61
+ const [stdout, stderr, code] = await Promise.all([new Response(proc.stdout).text(), new Response(proc.stderr).text(), proc.exited]);
62
+ return { code, output: stdout + stderr };
63
+ }
64
+
51
65
  /** The compose calls that build and start the app; `recreate` restarts containers even when nothing changed. */
52
66
  function startCalls(noCache: boolean, recreate: boolean): string[][] {
53
67
  const up = [...COMPOSE, "up", "-d", ...(noCache ? [] : ["--build"]), ...(recreate ? ["--force-recreate"] : [])];
@@ -174,8 +188,77 @@ export function appCommands(app: App, deps: AppDeps) {
174
188
  deps.print(`Delete ${source} now, the env file holds the key. mfw start applies it to a running app.`);
175
189
  return 0;
176
190
  },
191
+ /** Makes the app install the packages in `from`'s tarballs (`bun pm
192
+ * pack`) instead of the registry's: copied into `.packs/` (which the
193
+ * image copies too), overridden in the manifest, then installed.
194
+ * Everything is checked before anything is written. */
195
+ localPackages: async (from: string) => {
196
+ const source = resolve(from);
197
+ const target = join(app.dir, LOCAL_PACKS_DIR);
198
+ if (source === target) throw new Error(`${source} is the app's own .packs/: give the folder the tarballs were packed into.`);
199
+ if (!existsSync(source)) throw new Error(`${source} doesn't exist.`);
200
+ const files = readdirSync(source).filter((f) => f.endsWith(".tgz")).sort();
201
+ if (files.length === 0) throw new Error(`No .tgz in ${source}: pack the packages there first (bun pm pack).`);
202
+ const packs = await Promise.all(files.map(async (file) => ({ name: await packageNameOf(join(source, file)), file })));
203
+ const manifest = withLocalOverrides(readManifest(), packs);
204
+ rmSync(target, { recursive: true, force: true });
205
+ mkdirSync(target);
206
+ for (const { file } of packs) cpSync(join(source, file), join(target, file));
207
+ writeManifest(manifest);
208
+ deps.print(`${packs.length} local packages in ${target}: ${packs.map((p) => p.name).join(", ")}.`);
209
+ return deps.run(["bun", "install"], { cwd: app.dir });
210
+ },
211
+ /** Runs e2e tests (`tests`, or the app's `e2e/*.e2e.ts`) against the app's
212
+ * REPL in its container, keeping the turns and checks in
213
+ * `e2e/results/<time>/`; returns 1 when a case didn't pass. */
214
+ e2e: async (tests: string[], { repeat }: { repeat?: number }) => {
215
+ const files = findTests(tests, { appDir: app.dir, cwd: process.cwd() });
216
+ const loaded = await Promise.all(files.map(async (file) => ({ file: relative(process.cwd(), file) || file, test: await loadTest(file) })));
217
+ const results = join(app.dir, "e2e", "results", new Date().toISOString().replace(/[:.]/g, "-"));
218
+ mkdirSync(results, { recursive: true });
219
+ // The container's user writes the dumps here; on a Linux host it isn't
220
+ // the folder's owner.
221
+ chmodSync(results, 0o777);
222
+ let sessions = 0;
223
+ return runE2e(loaded, repeat === undefined ? {} : { repeat }, {
224
+ openSession: async () =>
225
+ openReplSession({
226
+ argv: [...COMPOSE, "run", "--rm", "-T", "-v", `${results}:/e2e`, SERVICE, "bun", "run", "repl"],
227
+ cwd: app.dir,
228
+ hostDir: results,
229
+ replDir: "/e2e",
230
+ name: `run-${++sessions}`,
231
+ timeoutMs: E2E_TURN_TIMEOUT_MS,
232
+ }),
233
+ cli: (command) => captureCode([...COMPOSE, "run", "--rm", "-T", "--no-deps", SERVICE, "sh", "-c", command], app.dir),
234
+ appPackages: (readManifest() as { dependencies?: Record<string, string> }).dependencies ?? {},
235
+ print: deps.print,
236
+ now: Date.now,
237
+ writeReport: async (report) => {
238
+ writeFileSync(join(results, "report.json"), `${JSON.stringify(report, null, 2)}\n`);
239
+ deps.print(`Turns and checks in ${results}`);
240
+ },
241
+ });
242
+ },
243
+ /** Undoes `localPackages`: the app installs from the registry again. */
244
+ localPackagesOff: async () => {
245
+ writeManifest(withoutLocalOverrides(readManifest()));
246
+ rmSync(join(app.dir, LOCAL_PACKS_DIR), { recursive: true, force: true });
247
+ deps.print("Local packages removed: installing from the registry.");
248
+ return deps.run(["bun", "install"], { cwd: app.dir });
249
+ },
177
250
  };
178
251
 
252
+ /** The app's manifest, as an object. */
253
+ function readManifest(): Record<string, unknown> {
254
+ return JSON.parse(readFileSync(join(app.dir, "package.json"), "utf-8")) as Record<string, unknown>;
255
+ }
256
+
257
+ /** Writes the app's manifest, two-space indented like the one `mfw create` writes. */
258
+ function writeManifest(manifest: Record<string, unknown>): void {
259
+ writeFileSync(join(app.dir, "package.json"), `${JSON.stringify(manifest, null, 2)}\n`);
260
+ }
261
+
179
262
  /** Stops `service`, runs `steps`, starts it again; stops at the first
180
263
  * failure and, once the service is stopped, says it's down and how to bring
181
264
  * it back. Returns the exit code. */
@@ -0,0 +1,51 @@
1
+ /**
2
+ * What `mfw local-packages` writes: the app's `package.json` with `overrides`
3
+ * pointing packages at local tarballs (`bun pm pack`) copied into `.packs/`,
4
+ * which the image copies before `bun install`. Overrides, and not the
5
+ * dependencies themselves, because a packed package names its siblings by
6
+ * version, which the registry has too: only an override sends those
7
+ * transitive dependencies to the tarballs as well.
8
+ */
9
+
10
+ /** The tarballs' folder inside the app, as the Dockerfile copies it. */
11
+ export const LOCAL_PACKS_DIR = ".packs";
12
+
13
+ /** The prefix of every override this command writes, how it tells them from the user's own. */
14
+ const LOCAL_OVERRIDE = `file:./${LOCAL_PACKS_DIR}/`;
15
+
16
+ type Manifest = Record<string, unknown> & { overrides?: Record<string, string> };
17
+
18
+ /** The user's own overrides in `pkg`: every one this command didn't write. */
19
+ function ownOverrides(pkg: Manifest): Record<string, string> {
20
+ return Object.fromEntries(Object.entries(pkg.overrides ?? {}).filter(([, spec]) => !spec.startsWith(LOCAL_OVERRIDE)));
21
+ }
22
+
23
+ /** `pkg` with each of `packs` (a package name and its tarball's file name in
24
+ * `.packs/`) as an override, in place of any left by an earlier run. */
25
+ export function withLocalOverrides(pkg: Manifest, packs: Array<{ name: string; file: string }>): Manifest {
26
+ const local = Object.fromEntries(packs.map((p) => [p.name, `${LOCAL_OVERRIDE}${p.file}`]));
27
+ return { ...pkg, overrides: { ...ownOverrides(pkg), ...local } };
28
+ }
29
+
30
+ /** `pkg` without the overrides this command writes; without `overrides` at
31
+ * all when none of the user's own is left. */
32
+ export function withoutLocalOverrides(pkg: Manifest): Manifest {
33
+ const { overrides: _, ...rest } = pkg;
34
+ const own = ownOverrides(pkg);
35
+ return Object.keys(own).length > 0 ? { ...rest, overrides: own } : rest;
36
+ }
37
+
38
+ /** The package name in a `bun pm pack` tarball, read from its
39
+ * `package/package.json`. */
40
+ export async function packageNameOf(tarball: string): Promise<string> {
41
+ const proc = Bun.spawn(["tar", "-xOzf", tarball, "package/package.json"], { stdout: "pipe", stderr: "pipe" });
42
+ const [manifest, code] = await Promise.all([new Response(proc.stdout).text(), proc.exited]);
43
+ try {
44
+ if (code !== 0) throw new Error(`tar exited with ${code}`);
45
+ const name = (JSON.parse(manifest) as { name?: unknown }).name;
46
+ if (typeof name !== "string") throw new Error("no name in its package.json");
47
+ return name;
48
+ } catch (err) {
49
+ throw new Error(`${tarball} isn't a package tarball: ${err instanceof Error ? err.message : String(err)}`);
50
+ }
51
+ }
@@ -0,0 +1,71 @@
1
+ /**
2
+ * The shape of an e2e test, as a test file imports it
3
+ * (`import { e2e } from "@mercury-fw/cli/e2e"`): the plugins and channels the
4
+ * app must have, and cases of turns sent to the app's real model through its
5
+ * REPL, each with checks on the calls it made and the answer it gave.
6
+ * `mfw e2e` runs them (`runner.ts`).
7
+ */
8
+ import type { Call, TurnData } from "./dump.ts";
9
+
10
+ export type { Call } from "./dump.ts";
11
+
12
+ /** One turn as a check sees it: its calls, its answer, how long it took. */
13
+ export type Turn = TurnData & { seconds: number };
14
+
15
+ /** A case's run: every turn in order, and the last one. */
16
+ export type Run = { turns: Turn[]; last: Turn };
17
+
18
+ /** Runs `command` in the app's container (`sh -c`), outside the model:
19
+ * preparing data, reading what a turn changed, cleaning up. */
20
+ export type Cli = (command: string) => Promise<{ code: number; output: string }>;
21
+
22
+ /** What `before`, `after` and `check` can reach besides the run. */
23
+ export type Context = { cli: Cli };
24
+
25
+ /** The checks a case makes. Call helpers look at every turn's calls, answer
26
+ * helpers at the last answer; each records a named check and never throws,
27
+ * so a case reports every failure at once. */
28
+ export type Expect = {
29
+ /** At least one call to `tool`, matching `match` when given. */
30
+ call(tool: string, match?: (call: Call) => boolean, label?: string): void;
31
+ /** Every call to `tool` matches `match` (and there is at least one). */
32
+ everyCall(tool: string, match: (call: Call) => boolean, label?: string): void;
33
+ /** No call failed or went without a result. */
34
+ noFailedCalls(label?: string): void;
35
+ /** The number of calls, to every tool or to `tool`, within the bounds. */
36
+ callCount(bounds: { min?: number; max?: number }, tool?: string, label?: string): void;
37
+ /** The answer contains `pattern`, or matches it. */
38
+ answer(pattern: string | RegExp, label?: string): void;
39
+ /** The answer doesn't contain `pattern`, or doesn't match it. */
40
+ answerNot(pattern: string | RegExp, label?: string): void;
41
+ /** Anything else. */
42
+ that(label: string, condition: boolean): void;
43
+ };
44
+
45
+ /** One case: the turns sent, in one REPL session, and the checks on them. */
46
+ export type E2eCase = {
47
+ name: string;
48
+ /** A message, or a function of the turn before (a follow-up, a confirmation token). */
49
+ turns: Array<string | ((previous: Turn) => string)>;
50
+ /** How many times to run it (the model isn't deterministic); default 1. */
51
+ repeat?: number;
52
+ /** How many runs must pass; default every one. */
53
+ minPasses?: number;
54
+ before?: (ctx: Context) => unknown;
55
+ after?: (ctx: Context) => unknown;
56
+ check: (run: Run, expect: Expect, ctx: Context) => unknown;
57
+ };
58
+
59
+ /** An e2e test: what the app must have, and its cases. */
60
+ export type E2eTest = {
61
+ /** Catalog ids of the tool plugins the app must have (`jira`, …). */
62
+ plugins?: string[];
63
+ /** Catalog ids of the channels the app must have (`http`, …). */
64
+ channels?: string[];
65
+ cases: E2eCase[];
66
+ };
67
+
68
+ /** Declares an e2e test: the identity, there for the types. */
69
+ export function e2e(test: E2eTest): E2eTest {
70
+ return test;
71
+ }
@@ -0,0 +1,48 @@
1
+ /**
2
+ * A turn read out of the file the REPL's `/dump` writes: the AI SDK's step
3
+ * results for the last turn, whose `content` parts are the tool calls, their
4
+ * results (or errors) and the model's text. Read as data, not through the
5
+ * SDK's types: only the parts listed here matter.
6
+ */
7
+
8
+ /** One tool call as a check sees it. `ok` is false when the call failed or
9
+ * never got a result; a result without an `ok` of its own counts as worked.
10
+ * `pending` is an irreversible command staged for confirmation, which worked
11
+ * (its output has the token) though it reports `ok: false`. */
12
+ export type Call = { tool: string; input: unknown; output: unknown; ok: boolean; pending: boolean };
13
+
14
+ /** What a turn did: its tool calls in order, and its final text. */
15
+ export type TurnData = { calls: Call[]; answer: string };
16
+
17
+ type Part = { type?: unknown; toolCallId?: unknown; toolName?: unknown; input?: unknown; output?: unknown; error?: unknown; text?: unknown };
18
+
19
+ /** The calls and the answer in `dump`, the parsed content of a `/dump` file. */
20
+ export function turnFromDump(dump: unknown): TurnData {
21
+ if (!Array.isArray(dump)) throw new Error("the file isn't a /dump of steps");
22
+ const parts: Part[][] = dump.map((step) =>
23
+ Array.isArray((step as { content?: unknown } | null)?.content) ? ((step as { content: Part[] }).content) : [],
24
+ );
25
+
26
+ const calls: Array<Call & { id: unknown }> = [];
27
+ for (const part of parts.flat()) {
28
+ if (part.type === "tool-call") {
29
+ calls.push({ id: part.toolCallId, tool: String(part.toolName), input: part.input, output: undefined, ok: false, pending: false });
30
+ continue;
31
+ }
32
+ const call = calls.find((c) => c.id === part.toolCallId);
33
+ if (call === undefined) continue;
34
+ if (part.type === "tool-result") {
35
+ call.output = part.output;
36
+ const result = part.output as { ok?: unknown; pendingConfirmation?: unknown } | null;
37
+ call.pending = result?.pendingConfirmation === true;
38
+ call.ok = call.pending || (typeof result?.ok === "boolean" ? result.ok : true);
39
+ } else if (part.type === "tool-error") {
40
+ call.output = { error: part.error };
41
+ call.ok = false;
42
+ }
43
+ }
44
+
45
+ const last = parts.at(-1) ?? [];
46
+ const answer = last.filter((p) => p.type === "text" && typeof p.text === "string").map((p) => p.text as string).join("");
47
+ return { calls: calls.map(({ id: _, ...call }) => call), answer };
48
+ }
@@ -0,0 +1,59 @@
1
+ /**
2
+ * The `expect` a case's `check` receives: helpers that record named checks
3
+ * against a run, passed or failed with a detail, and never throw.
4
+ */
5
+ import type { Call } from "./dump.ts";
6
+ import type { Expect, Run } from "./define.ts";
7
+
8
+ /** One check's outcome; `detail` says what was found when it failed. */
9
+ export type Check = { label: string; ok: boolean; detail?: string };
10
+
11
+ /** A call, one line, for a failure's detail. */
12
+ const describe = (c: Call) => `${c.tool} ${JSON.stringify(c.input)}`;
13
+
14
+ /** `pattern`, as a label says it. */
15
+ const said = (pattern: string | RegExp) => (typeof pattern === "string" ? `contains "${pattern}"` : `matches ${pattern}`);
16
+ const saidNot = (pattern: string | RegExp) => (typeof pattern === "string" ? `doesn't contain "${pattern}"` : `doesn't match ${pattern}`);
17
+ const found = (text: string, pattern: string | RegExp) => (typeof pattern === "string" ? text.includes(pattern) : pattern.test(text));
18
+
19
+ /** The checks' bounds, as a label says them. */
20
+ function boundsLabel({ min, max }: { min?: number; max?: number }, what: string): string {
21
+ if (min !== undefined && max !== undefined) return `${min} to ${max} ${what}`;
22
+ if (max !== undefined) return `at most ${max} ${what}`;
23
+ return `at least ${min ?? 0} ${what}`;
24
+ }
25
+
26
+ /** An `expect` bound to `run`, and the checks it recorded so far. */
27
+ export function createExpect(run: Run): { expect: Expect; checks: () => Check[] } {
28
+ const checks: Check[] = [];
29
+ const record = (label: string, ok: boolean, detail?: string) =>
30
+ void checks.push(ok || detail === undefined ? { label, ok } : { label, ok, detail });
31
+ const calls = run.turns.flatMap((t) => t.calls);
32
+ const of = (tool: string) => calls.filter((c) => c.tool === tool);
33
+
34
+ const expect: Expect = {
35
+ call: (tool, match, label) => {
36
+ const ok = of(tool).some((c) => match?.(c) ?? true);
37
+ record(label ?? `a ${tool} call`, ok, `calls: ${calls.map((c) => c.tool).join(", ") || "none"}`);
38
+ },
39
+ everyCall: (tool, match, label) => {
40
+ const mine = of(tool);
41
+ const bad = mine.find((c) => !match(c));
42
+ record(label ?? `every ${tool} call matches`, mine.length > 0 && bad === undefined, bad ? describe(bad) : `no ${tool} call`);
43
+ },
44
+ noFailedCalls: (label) => {
45
+ const failed = calls.filter((c) => !c.ok);
46
+ record(label ?? "no failed calls", failed.length === 0, failed.map(describe).join("; "));
47
+ },
48
+ callCount: (bounds, tool, label) => {
49
+ const n = (tool === undefined ? calls : of(tool)).length;
50
+ const ok = n >= (bounds.min ?? 0) && n <= (bounds.max ?? Infinity);
51
+ const what = tool === undefined ? "calls" : `${tool} call${bounds.max === 1 || bounds.min === 1 ? "" : "s"}`;
52
+ record(label ?? boundsLabel(bounds, what), ok, String(n));
53
+ },
54
+ answer: (pattern, label) => record(label ?? `answer ${said(pattern)}`, found(run.last.answer, pattern), run.last.answer),
55
+ answerNot: (pattern, label) => record(label ?? `answer ${saidNot(pattern)}`, !found(run.last.answer, pattern), run.last.answer),
56
+ that: (label, condition) => record(label, condition),
57
+ };
58
+ return { expect, checks: () => [...checks] };
59
+ }
@@ -0,0 +1,40 @@
1
+ /**
2
+ * Which e2e test files `mfw e2e` runs, and loading one: Bun imports the
3
+ * TypeScript file as it is, and its default export is checked to look like
4
+ * a test before anything starts.
5
+ */
6
+ import { existsSync, readdirSync } from "node:fs";
7
+ import { join, resolve } from "node:path";
8
+ import { pathToFileURL } from "node:url";
9
+ import type { E2eTest } from "./define.ts";
10
+
11
+ /** The files to run: `named` resolved from `cwd`, or, when none is named,
12
+ * the app's `e2e/*.e2e.ts` in name order. */
13
+ export function findTests(named: string[], { appDir, cwd }: { appDir: string; cwd: string }): string[] {
14
+ if (named.length > 0) return named.map((f) => resolve(cwd, f));
15
+ const folder = join(appDir, "e2e");
16
+ const found = existsSync(folder) ? readdirSync(folder).filter((f) => f.endsWith(".e2e.ts")).sort() : [];
17
+ if (found.length === 0) throw new Error(`No tests: none named, and no *.e2e.ts in ${folder}`);
18
+ return found.map((f) => join(folder, f));
19
+ }
20
+
21
+ /** The test `file` exports as default; throws naming the file when it isn't one. */
22
+ export async function loadTest(file: string): Promise<E2eTest> {
23
+ const loaded = ((await import(pathToFileURL(file).href)) as { default?: unknown }).default as Partial<E2eTest> | undefined;
24
+ if (loaded === undefined || loaded === null || !Array.isArray(loaded.cases)) {
25
+ throw new Error(`${file} doesn't export a test as default (export default e2e({ … }))`);
26
+ }
27
+ for (const [i, c] of loaded.cases.entries()) {
28
+ const which = `${file}: case ${i + 1} ("${c?.name ?? "?"}")`;
29
+ if (typeof c?.name !== "string") throw new Error(`${which} has no name`);
30
+ if (!Array.isArray(c.turns) || c.turns.length === 0) throw new Error(`${which} has no turns`);
31
+ if (typeof c.check !== "function") throw new Error(`${which} has no check`);
32
+ const whole = (n: unknown) => typeof n === "number" && Number.isInteger(n) && n >= 1;
33
+ if (c.repeat !== undefined && !whole(c.repeat)) throw new Error(`${which} repeat takes a positive whole number`);
34
+ if (c.minPasses !== undefined && !whole(c.minPasses)) throw new Error(`${which} minPasses takes a positive whole number`);
35
+ if (c.minPasses !== undefined && c.minPasses > (c.repeat ?? 1)) {
36
+ throw new Error(`${which} minPasses (${c.minPasses}) is more than repeat (${c.repeat ?? 1})`);
37
+ }
38
+ }
39
+ return loaded as E2eTest;
40
+ }
@@ -0,0 +1,153 @@
1
+ /**
2
+ * Runs e2e tests against an app: for each case (as many times as it
3
+ * repeats), a fresh REPL session gets the case's turns, the checks run on
4
+ * what each turn did, and every check is printed; the exit code says whether
5
+ * every case passed enough runs. The session, the container commands and the
6
+ * report's destination are injected (`session.ts` and `commands.ts` hold the
7
+ * real ones), so the tests drive the runner with scripted turns.
8
+ */
9
+ import { CATALOG } from "../catalog.ts";
10
+ import type { Context, E2eCase, E2eTest, Run, Turn } from "./define.ts";
11
+ import { turnFromDump } from "./dump.ts";
12
+ import { createExpect, type Check } from "./expect.ts";
13
+
14
+ /** A REPL session in the app: `turn` sends one line and resolves with what
15
+ * `/dump` wrote for it and what the REPL printed meanwhile. */
16
+ export type Session = {
17
+ turn: (line: string) => Promise<{ dump: unknown; output: string }>;
18
+ close: () => Promise<void>;
19
+ };
20
+
21
+ export type RunnerDeps = {
22
+ openSession: () => Promise<Session>;
23
+ cli: Context["cli"];
24
+ /** The app's dependencies, to check the test's plugins and channels against. */
25
+ appPackages: Record<string, string>;
26
+ print: (line: string) => void;
27
+ /** Milliseconds, for each turn's duration. */
28
+ now: () => number;
29
+ writeReport: (report: CaseReport[]) => Promise<void>;
30
+ };
31
+
32
+ /** One run of a case, as the report keeps it. */
33
+ export type RunReport = { ok: boolean; turns: Turn[]; checks: Check[] };
34
+ export type CaseReport = { file: string; case: string; passed: boolean; runs: RunReport[] };
35
+
36
+ /** The prompt the REPL prints after a reply, with its context-usage suffix. */
37
+ const PROMPT = /(\[[^\]\n]*\] )?> $/;
38
+
39
+ /** A turn's printed output as an answer: no styling, no trailing prompt. */
40
+ function answerFromOutput(output: string): string {
41
+ return output.replace(/\u001b\[[0-9;]*m/g, "").replace(PROMPT, "").trim();
42
+ }
43
+
44
+ /** The test's plugins and channels the app doesn't depend on, as a message, or undefined. */
45
+ function missing(test: E2eTest, packages: Record<string, string>): string | undefined {
46
+ const wanted = [
47
+ ...(test.plugins ?? []).map((id) => ({ id, kind: "tool" as const, word: "plugin" })),
48
+ ...(test.channels ?? []).map((id) => ({ id, kind: "channel" as const, word: "channel" })),
49
+ ];
50
+ const lacking = wanted.flatMap(({ id, kind, word }) => {
51
+ const pkg = CATALOG.find((e) => e.kind === kind && e.id === id)?.package;
52
+ if (pkg === undefined) return [`${word} ${id} (not in the catalog)`];
53
+ return packages[pkg] === undefined ? [`${word} ${id} (${pkg})`] : [];
54
+ });
55
+ return lacking.length > 0 ? `the app lacks what the test needs: ${lacking.join(", ")}` : undefined;
56
+ }
57
+
58
+ /** Runs the tests in `tests` (each with the file it came from); returns the
59
+ * exit code: 0 when every case passed enough runs. `repeat` overrides each
60
+ * case's own. */
61
+ export async function runE2e(
62
+ tests: Array<{ file: string; test: E2eTest }>,
63
+ opts: { repeat?: number },
64
+ deps: RunnerDeps,
65
+ ): Promise<number> {
66
+ const report: CaseReport[] = [];
67
+ let code = 0;
68
+ for (const { file, test } of tests) {
69
+ deps.print(file);
70
+ const lacking = missing(test, deps.appPackages);
71
+ if (lacking !== undefined) {
72
+ deps.print(` ${lacking}`);
73
+ code = 1;
74
+ continue;
75
+ }
76
+ for (const c of test.cases) {
77
+ const result = await runCase(c, opts.repeat, deps);
78
+ report.push({ file, case: c.name, ...result });
79
+ if (!result.passed) code = 1;
80
+ }
81
+ }
82
+ await deps.writeReport(report);
83
+ return code;
84
+ }
85
+
86
+ /** Runs one case as many times as it repeats, printing each run's checks and the verdict. */
87
+ async function runCase(c: E2eCase, repeatOverride: number | undefined, deps: RunnerDeps): Promise<{ passed: boolean; runs: RunReport[] }> {
88
+ const repeat = repeatOverride ?? c.repeat ?? 1;
89
+ const minPasses = Math.min(c.minPasses ?? repeat, repeat);
90
+ deps.print(` ${c.name}`);
91
+ const runs: RunReport[] = [];
92
+ for (let i = 1; i <= repeat; i++) {
93
+ if (repeat > 1) deps.print(` run ${i}/${repeat}`);
94
+ const run = await runOnce(c, deps);
95
+ const indent = repeat > 1 ? " " : " ";
96
+ for (const check of run.checks) {
97
+ deps.print(`${indent}${check.ok ? "✓" : "✗"} ${check.label}${check.ok || check.detail === undefined ? "" : `: ${check.detail}`}`);
98
+ }
99
+ runs.push(run);
100
+ }
101
+ const passes = runs.filter((r) => r.ok).length;
102
+ const passed = passes >= minPasses;
103
+ deps.print(
104
+ repeat > 1
105
+ ? ` ${c.name}: ${passes}/${repeat} runs passed (needs ${minPasses})${passed ? "" : " FAILED"}`
106
+ : ` ${c.name}: ${passed ? "passed" : "FAILED"}`,
107
+ );
108
+ return { passed, runs };
109
+ }
110
+
111
+ /** One run of a case in a fresh session: `before`, the turns, the checks, `after`. */
112
+ async function runOnce(c: E2eCase, deps: RunnerDeps): Promise<RunReport> {
113
+ const ctx: Context = { cli: deps.cli };
114
+ const turns: Turn[] = [];
115
+ let checks: Check[] = [];
116
+ try {
117
+ await c.before?.(ctx);
118
+ const session = await deps.openSession();
119
+ try {
120
+ for (const next of c.turns) {
121
+ const line = typeof next === "string" ? next : next(turns.at(-1)!);
122
+ if (line.includes("\n")) throw new Error("a turn must be one line (the REPL reads one line per turn)");
123
+ const start = deps.now();
124
+ const { dump, output } = await session.turn(line);
125
+ const seconds = (deps.now() - start) / 1000;
126
+ const data = turnFromDump(dump);
127
+ const answer = data.calls.length === 0 && data.answer === "" ? answerFromOutput(output) : data.answer;
128
+ turns.push({ calls: data.calls, answer, seconds });
129
+ }
130
+ } finally {
131
+ await session.close();
132
+ }
133
+ const run: Run = { turns, last: turns.at(-1)! };
134
+ const recorder = createExpect(run);
135
+ try {
136
+ await c.check(run, recorder.expect, ctx);
137
+ checks = recorder.checks();
138
+ // A case that checks nothing proves nothing.
139
+ if (checks.length === 0) checks = [{ label: "the case made no checks", ok: false }];
140
+ } catch (err) {
141
+ checks = [...recorder.checks(), { label: `check threw: ${err instanceof Error ? err.message : String(err)}`, ok: false }];
142
+ }
143
+ } catch (err) {
144
+ checks = [{ label: `run failed: ${err instanceof Error ? err.message : String(err)}`, ok: false }];
145
+ } finally {
146
+ try {
147
+ await c.after?.(ctx);
148
+ } catch (err) {
149
+ checks = [...checks, { label: `after threw: ${err instanceof Error ? err.message : String(err)}`, ok: false }];
150
+ }
151
+ }
152
+ return { ok: checks.every((ch) => ch.ok), turns, checks };
153
+ }
@@ -0,0 +1,104 @@
1
+ /**
2
+ * The REPL session `mfw e2e` drives: the app's REPL as a child process (in
3
+ * the app's container, through `docker compose run`), fed one turn at a time.
4
+ * A turn writes its line and then `/dump <file>` together: the REPL handles
5
+ * one line after the other, so the dump runs once the turn is over, and its
6
+ * "wrote … to <file>" line is the turn's end, unlike the prompt, which an
7
+ * answer's own text could look like. The dump is read from the host side of
8
+ * the folder the REPL writes it into.
9
+ */
10
+ import { readFileSync } from "node:fs";
11
+ import { join } from "node:path";
12
+ import type { Session } from "./runner.ts";
13
+
14
+ export type ReplSessionOptions = {
15
+ /** The command that starts the REPL. */
16
+ argv: string[];
17
+ cwd: string;
18
+ /** Where the dumps are, on this side, and as the REPL sees the same folder. */
19
+ hostDir: string;
20
+ replDir: string;
21
+ /** Prefix of this session's dump files, unique among sessions sharing the folder. */
22
+ name: string;
23
+ /** How long one turn may take. */
24
+ timeoutMs: number;
25
+ };
26
+
27
+ /** Starts the REPL and returns the session over it once the REPL is ready
28
+ * (its first prompt is out), so a turn's time is only the turn's; throws
29
+ * when the REPL exits or doesn't get there in time. */
30
+ export async function openReplSession(opts: ReplSessionOptions): Promise<Session> {
31
+ const proc = Bun.spawn(opts.argv, { cwd: opts.cwd, stdin: "pipe", stdout: "pipe", stderr: "pipe" });
32
+ let stdout = "";
33
+ let stderr = "";
34
+ let exited: number | undefined;
35
+ /** Wakes whoever waits for more output or for the exit. */
36
+ let notify = () => {};
37
+ const pump = async (stream: ReadableStream<Uint8Array>, add: (text: string) => void) => {
38
+ const decoder = new TextDecoder();
39
+ for await (const chunk of stream) {
40
+ add(decoder.decode(chunk, { stream: true }));
41
+ notify();
42
+ }
43
+ };
44
+ void pump(proc.stdout, (t) => (stdout += t));
45
+ void pump(proc.stderr, (t) => (stderr += t));
46
+ void proc.exited.then((code) => {
47
+ exited = code;
48
+ notify();
49
+ });
50
+
51
+ /** Waits until `done` holds on the output so far; throws when the REPL exits or time runs out. */
52
+ const waitFor = async <T>(done: () => T | undefined): Promise<T> => {
53
+ const deadline = Date.now() + opts.timeoutMs;
54
+ for (;;) {
55
+ const result = done();
56
+ if (result !== undefined) return result;
57
+ if (exited !== undefined) throw new Error(`the REPL exited with code ${exited}: ${stderr.trim().split("\n").slice(-3).join(" ")}`);
58
+ const left = deadline - Date.now();
59
+ if (left <= 0) throw new Error(`no reply within ${opts.timeoutMs / 1000} s`);
60
+ await new Promise<void>((resolve) => {
61
+ const timer = setTimeout(resolve, left);
62
+ notify = () => {
63
+ clearTimeout(timer);
64
+ resolve();
65
+ };
66
+ });
67
+ }
68
+ };
69
+
70
+ // What the REPL prints while starting isn't part of any turn. One that
71
+ // doesn't get there is stopped: a docker compose run container otherwise
72
+ // keeps running.
73
+ try {
74
+ await waitFor(() => (/> $/.test(stdout) ? true : undefined));
75
+ } catch (err) {
76
+ proc.kill();
77
+ throw err;
78
+ }
79
+
80
+ let turns = 0;
81
+ return {
82
+ turn: async (line) => {
83
+ turns++;
84
+ const file = `${opts.name}-turn-${turns}.json`;
85
+ const marker = `to ${opts.replDir}/${file}`;
86
+ const from = stdout.length;
87
+ proc.stdin.write(`${line}\n/dump ${opts.replDir}/${file}\n`);
88
+ proc.stdin.flush();
89
+ return waitFor(() => {
90
+ const at = stdout.indexOf(marker, from);
91
+ if (at === -1) return undefined;
92
+ // What the turn printed: everything before the dump's own line.
93
+ const output = stdout.slice(from, stdout.lastIndexOf("wrote ", at));
94
+ return { dump: JSON.parse(readFileSync(join(opts.hostDir, file), "utf-8")) as unknown, output };
95
+ });
96
+ },
97
+ close: async () => {
98
+ if (exited !== undefined) return;
99
+ proc.stdin.end();
100
+ const done = await Promise.race([proc.exited, new Promise((r) => setTimeout(() => r("timeout"), 10_000))]);
101
+ if (done === "timeout") proc.kill();
102
+ },
103
+ };
104
+ }
package/src/program.ts CHANGED
@@ -24,12 +24,14 @@ export type ProgramHandlers = {
24
24
  app: () => AppCommands;
25
25
  };
26
26
 
27
- /** `--limit`'s value: a positive whole number. */
28
- function positiveInt(value: string): string {
29
- if (!/^\d+$/.test(value) || Number(value) < 1) {
30
- throw new InvalidArgumentError(`--limit takes a positive whole number (got "${value}").`);
31
- }
32
- return value;
27
+ /** A parser for `flag`'s value: a positive whole number, kept as typed. */
28
+ function positiveInt(flag: string): (value: string) => string {
29
+ return (value) => {
30
+ if (!/^\d+$/.test(value) || Number(value) < 1) {
31
+ throw new InvalidArgumentError(`${flag} takes a positive whole number (got "${value}").`);
32
+ }
33
+ return value;
34
+ };
33
35
  }
34
36
 
35
37
  /** The help's closing paragraph for the commands that run inside an app. */
@@ -199,7 +201,7 @@ Examples:
199
201
  "Prints a collection's points, each as its id and one line per payload field: newest first where the collection has a timestamp index, in Qdrant's own order otherwise.",
200
202
  )
201
203
  .argument("<collection>", "as mfw memory list prints it")
202
- .option("--limit <n>", "how many points (default: 20)", positiveInt)
204
+ .option("--limit <n>", "how many points (default: 20)", positiveInt("--limit"))
203
205
  .action(async (collection: string, opts: { limit?: string }) =>
204
206
  inApp((app) => app.memory(["read", collection, ...(opts.limit === undefined ? [] : ["--limit", opts.limit])]))(),
205
207
  );
@@ -267,6 +269,33 @@ Examples:
267
269
  )(),
268
270
  );
269
271
 
272
+ program
273
+ .command("local-packages")
274
+ .summary("installs @mercury-fw packages from local tarballs")
275
+ .description(
276
+ "Makes the app install the packages in <folder>'s tarballs (bun pm pack) instead of the registry's: copies them into .packs/ (the image copies it too), points each package at its tarball with overrides in package.json, so packages that depend on it get it as well, and runs bun install. Run it again after packing anew. --off removes them and installs from the registry.",
277
+ )
278
+ .argument("[folder]", "the folder holding the .tgz files")
279
+ .option("--off", "go back to the registry's packages")
280
+ .addHelpText("after", `\nExamples:\n mfw local-packages ../mercury-fw/apps/testbed/.packs\n mfw local-packages --off${INSIDE_AN_APP}`)
281
+ .action(async (folder: string | undefined, opts: { off?: boolean }) => {
282
+ if ((folder === undefined) === (opts.off !== true)) throw new Error("local-packages takes a folder of tarballs, or --off");
283
+ await inApp((app) => (folder === undefined ? app.localPackagesOff() : app.localPackages(folder)))();
284
+ });
285
+
286
+ program
287
+ .command("e2e")
288
+ .summary("runs the app's end-to-end tests against its model")
289
+ .description(
290
+ "Runs e2e tests: each case's turns go to the app's real model through its REPL (in the container, like mfw repl), and its checks run on the tool calls each turn made and the answer it gave. Without files, every e2e/*.e2e.ts in the app. Prints every check, exits 1 when a case didn't pass enough runs, and keeps the turns and checks in e2e/results/<time>/. The app must be built and its env file filled in, as for mfw repl.",
291
+ )
292
+ .argument("[tests...]", "test files (default: e2e/*.e2e.ts)")
293
+ .option("--repeat <n>", "runs of each case, overriding its own repeat", positiveInt("--repeat"))
294
+ .addHelpText("after", `\nExamples:\n mfw e2e\n mfw e2e e2e/jira.e2e.ts --repeat 3${INSIDE_AN_APP}`)
295
+ .action(async (tests: string[], opts: { repeat?: string }) =>
296
+ inApp((app) => app.e2e(tests, opts.repeat === undefined ? {} : { repeat: Number(opts.repeat) }))(),
297
+ );
298
+
270
299
  return program;
271
300
  }
272
301
 
@@ -12,8 +12,11 @@ RUN groupadd -r mercury && useradd -r -g mercury mercury
12
12
  WORKDIR /app
13
13
 
14
14
  # Each tool plugin downloads its pinned CLI in its postinstall (allowed by
15
- # package.json's trustedDependencies); the binaries go on PATH.
15
+ # package.json's trustedDependencies); the binaries go on PATH. .packs/ holds
16
+ # local tarballs when the app installs them (mfw local-packages), and is
17
+ # skipped when it isn't there.
16
18
  COPY --chown=mercury:mercury package.json bun.lock* ./
19
+ COPY --chown=mercury:mercury .pack[s] ./.packs/
17
20
  RUN bun install --production && chown -R mercury:mercury node_modules
18
21
  RUN find /app/node_modules -path '*/@mercury-fw/*/bin/*' -type f -exec ln -sf {} /usr/local/bin/ \;
19
22
 
@@ -4,3 +4,5 @@ node_modules
4
4
  !.env.example
5
5
  *.log
6
6
  .DS_Store
7
+ .packs/
8
+ e2e/results/