@trawlme/cli 3.12.5 → 3.12.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +8 -6
- package/dist/commands/create.js +65 -26
- package/dist/commands/scraps.js +38 -35
- package/dist/lib/api.d.ts +15 -3
- package/dist/lib/api.js +198 -35
- package/dist/lib/errors.d.ts +2 -0
- package/dist/lib/errors.js +4 -2
- package/dist/lib/format.d.ts +6 -1
- package/dist/lib/format.js +30 -10
- package/dist/lib/json.js +19 -14
- package/dist/lib/pinchAnimation.d.ts +2 -1
- package/dist/lib/pinchAnimation.js +6 -1
- package/docs/agent-quickstart.md +23 -13
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -68,18 +68,20 @@ trawl spec [--json] Print a versioned, machine-readab
|
|
|
68
68
|
|
|
69
69
|
> **No breaking change:** every verb above is also still reachable under its pre-reorg path, `trawl scraps <verb>` (e.g. `trawl scraps list`, `trawl scraps run <id>`) — kept as a hidden alias so scripts written before the surface reorg keep working. `trawl --help` only shows the top-level form above; `trawl scraps --help` only shows the remaining scrap-management commands below.
|
|
70
70
|
|
|
71
|
-
`create`/`whoami`/`ping` are fully non-interactive — all three read auth only from `TRAWL_API_KEY`/`TRAWL_TOKEN`/the stored login token, never prompt (`whoami` is JWT-only — see [Authentication](#authentication) — so it still refuses honestly under a key instead of prompting). `trawl create` runs the AI wizard server-side (`POST /api/ai/wizard`): generate scrap code from `--prompt` via LLM, persist the scrap, trigger its FIRST run, and auto-fix on failure (default on — `--no-autofix` disables it, sending `autoFix:false`). On a successful first run (human mode) it prints a small **data sample** (item count + first-item fields + one truncated value) as proof of value — best-effort, silent if the sample can't be fetched — and points `Next step` at `trawl data <id>` (the data), with `trawl get <id>` as the secondary detail view. `--json` skips the sample fetch and prints the raw wizard payload verbatim. `success` is an honest outcome of that first run, not "did the HTTP call succeed" — a failed first run is still a 200 response (the scrap was still created; auto-fix, when enabled, retries in the background), and the CLI exits 1 in that case (both human and `--json` modes) even though `--json` always prints the raw payload verbatim. The call legitimately takes many minutes server-side (AI generation + a proxy-tier probe ladder + an iteration loop + a real run), so it arms a
|
|
71
|
+
`create`/`whoami`/`ping` are fully non-interactive — all three read auth only from `TRAWL_API_KEY`/`TRAWL_TOKEN`/the stored login token, never prompt (`whoami` is JWT-only — see [Authentication](#authentication) — so it still refuses honestly under a key instead of prompting). `trawl create` runs the AI wizard server-side (`POST /api/ai/wizard`): generate scrap code from `--prompt` via LLM, persist the scrap, trigger its FIRST run, and auto-fix on failure (default on — `--no-autofix` disables it, sending `autoFix:false`). On a successful first run (human mode) it prints a small **data sample** (item count + first-item fields + one truncated value) as proof of value — best-effort, silent if the sample can't be fetched — and points `Next step` at `trawl data <id>` (the data), with `trawl get <id>` as the secondary detail view. `--json` skips the sample fetch and prints the raw wizard payload verbatim. `success` is an honest outcome of that first run, not "did the HTTP call succeed" — a failed first run is still a 200 response (the scrap was still created; auto-fix, when enabled, retries in the background), and the CLI exits 1 in that case (both human and `--json` modes) even though `--json` always prints the raw payload verbatim. The call legitimately takes many minutes server-side (AI generation + a proxy-tier probe ladder + an iteration loop + a real run), so it reads the wizard's NDJSON progress stream (`Accept: application/x-ndjson`) — human mode prints one short line per stage on stderr, `--json` prints only the final payload, same shape as always — and arms a 3600s safety ceiling of its own, longer than the 660s `run`/`data --fresh`/`trigger --wait` use below, over a transport of its own (Node's global `fetch` cannot wait past ~300s whatever ceiling is asked for). `trawl whoami`/`trawl ping` mirror the MCP `trawl_whoami`/`trawl_health_ping` tools as closely as the REST surface allows (`GET /api/users/me` / `GET /api/health`) — `ping`'s `--json` payload is admin-enriched (version/uptime/db) and just `{"status":"ok"}` for anyone else.
|
|
72
72
|
|
|
73
|
-
> **`create` is NOT idempotent, and every wizard-created scrap runs on a
|
|
73
|
+
> **`create` is NOT idempotent, and every wizard-created scrap runs on a WEEKLY cron by default.** A transport failure — a client-side timeout, a progress stream cut or ended without a result, a connection dropped after the API already answered (exit `5`, a `NetworkError`), or a gateway `502`/`504` — does not mean the wizard failed server-side: scrap creation + the first run keep going after the CLI gives up waiting, so the scrap may already exist. In human mode the CLI prints a stderr warning whenever the request reached the API without returning a complete answer; under `--json` there is no warning and no extra envelope field — a machine caller sees exit `5` (or, for a gateway `502`/`504`, exit `1` with `status` 502/504) and must check `trawl list` itself. Run `trawl list` and look for a matching URL/title **before** retrying — a blind retry creates a DUPLICATE scrap and burns AI-generation quota a second time for the same goal. Separately, the scrap the wizard creates is scheduled to re-run every week, on Monday at 07:00 UTC (`cron: "0 7 * * 1"`, hardcoded server-side, unrelated to `--no-autofix`) — each of those recurring runs consumes execute quota like any other run. Review the generated scrap, then change or disable the schedule with `trawl scraps update <id> --cron <expr>` (or `--no-cron` to disable it). Because the call can legitimately run many minutes, also confirm `TRAWL_TIMEOUT` isn't set to something tighter than `create` needs — the env var always wins over `create`'s own 3600s default (see [Environment variables](#environment-variables)), so a value set for another purpose (e.g. a tight CI smoke-test budget) silently clamps `create` too; unset it or raise it before running `create`.
|
|
74
|
+
|
|
75
|
+
> **A failure the wizard reports itself** — the stream's terminal `error` event — arrives after the HTTP `200` was already sent, so there is no status to report: `create` exits `1` as `kind:"refused"` with **no `status` field** (`{"error":{"message":"…","kind":"refused","retryable":false}}` under `--json`, the server's message verbatim; `✗ <message>` on stderr otherwise, stripped of control characters and capped at 200 characters). Branch on `kind`, not on a status.
|
|
74
76
|
|
|
75
77
|
- `list` has a short alias, `ls` (matches `trawl --help`'s `list|ls`).
|
|
76
78
|
- `history` lists past runs (newest first); `run-info <hid>` shows details of a single run from that history.
|
|
77
79
|
- `data` returns the last persisted run payload (no execute quota); `--fresh` runs the scrap live instead (consumes execute quota); `--errors` shows the last run's error detail (`--json` on a never-run scrap returns `{"status":"no_runs"}`, exit 0, matching `scraps doctor --json`). `[]` on stdout means a genuine zero-item successful run — a scrap that has never run, whose last run failed, or whose payload aged out of retention returns a `--json` error envelope (exit 4/1/4 respectively) instead. Two more honest states: a run still **in flight** (`status: null` server-side) returns a `kind:"in_progress"` error envelope (exit 1, "retry shortly" — never suggests `--fresh`, which would just 429 against the run already holding the lock); a run whose item count **regressed** vs baseline (`statusDetail: "regression"`) still returns the real, non-empty items on stdout (exit 0) plus a stderr warning pointing at `scraps doctor <id>` — the data itself is genuine even though the run is flagged.
|
|
78
80
|
- `get` (and anything reading through it, like `data`'s default path) embeds only the newest 100 history rows on the returned scrap object — `run-info` and `scraps doctor` fetch a single run directly and are unaffected by that cap. `list`/`get` show a distinct amber `▼` "regression" badge, never the red `✗` a genuine failure gets (matches `scraps doctor`'s own badge).
|
|
79
81
|
- `list`/`get` also show a **health** badge next to the status badge. By default it reads the scrap's own `lastCronOutcome` field (always present, no extra cost): `⚠ cron paused` when the last scheduled tick was skipped for being unhealthy, `—` otherwise — a breadcrumb of the last tick, not a live read (never set on a no-cron scrap, can lag by one tick). `list --unhealthy` asks the API to filter to scraps whose *current* consecutive-failure streak is 3+ (the server's own threshold) and annotates each with the exact `consecutiveFailedRuns`/`unhealthySince` — shown instead of the badge whenever present, and never re-derived client-side from history. `--json` carries whichever fields the request produced: `lastCronOutcome` always, `consecutiveFailedRuns`/`unhealthySince` only when `--unhealthy` was passed. **Old-server safety:** an older server version that predates this filter silently ignores the unknown `--unhealthy` query param and returns everything, unfiltered — the CLI detects that (no returned item carries `consecutiveFailedRuns`) and prints a stderr warning instead of presenting the full list as "your unhealthy scraps".
|
|
80
|
-
- **Long-running calls (`run`, `data --fresh`, `trigger --wait`):** these hit server-side paths that can legitimately take
|
|
81
|
-
- **`create` is longer still:** the wizard adds a proxy-tier probe ladder and an iteration loop on top of generation + creation + a first run, so it arms **
|
|
82
|
-
- **`--watch` is poll-based, not a live stream:** the activities SSE endpoint has no backlog and, for the default async `trigger` (no `--wait`), runs in a separate cron-consumer pod whose events never reach the API pod holding the SSE connection — a naive "await the run, then open SSE" shows nothing. `run --watch` and `trigger --watch` instead poll `GET /api/scraps/:id` (terminal status) and the activities REST list until the run finishes, printing each new activity line as it appears. The watched run's outcome drives the exit code too, in BOTH human and `--json` mode: a genuinely failed terminal run, a poll timeout, or a persistently unreachable API all exit non-zero — a clean successful run is the only exit `0`. A run that never reaches a terminal status within
|
|
82
|
+
- **Long-running calls (`run`, `data --fresh`, `trigger --wait`):** these hit server-side paths that can legitimately take minutes (proxy tier escalation + AI-fix retries, up to the server's 600s per-run budget) — the CLI arms a 660s timeout (600s + 60s) for exactly these three call sites instead of the generic 30s default, over the same raw `node:http(s)` transport `create` uses: Node's global `fetch` kills any call at ~300s whatever ceiling is asked for.
|
|
83
|
+
- **`create` is longer still:** the wizard adds a proxy-tier probe ladder and an iteration loop on top of generation + creation + a first run, so it arms **3600s** for that one call site (600s probe ladder + 900s iteration loop + 600s for a dry-run still in flight + 600s first run + margin) — a safety ceiling only, since `create` reads the wizard's live NDJSON progress stream. Raising the 660s above to cover it would make the three calls above wait an hour on an unreachable API, so the two ceilings are deliberately separate. `TRAWL_TIMEOUT` (see below) still overrides ALL requests, including both — set it if you need a tighter or looser ceiling, but note a global override sized for a quick call also clamps `create`.
|
|
84
|
+
- **`--watch` is poll-based, not a live stream:** the activities SSE endpoint has no backlog and, for the default async `trigger` (no `--wait`), runs in a separate cron-consumer pod whose events never reach the API pod holding the SSE connection — a naive "await the run, then open SSE" shows nothing. `run --watch` and `trigger --watch` instead poll `GET /api/scraps/:id` (terminal status) and the activities REST list until the run finishes, printing each new activity line as it appears. The watched run's outcome drives the exit code too, in BOTH human and `--json` mode: a genuinely failed terminal run, a poll timeout, or a persistently unreachable API all exit non-zero — a clean successful run is the only exit `0`. A run that never reaches a terminal status within 660s prints an honest timeout notice pointing at `scraps doctor <id>` (human mode) — see the `--json` shape below. `scraps watch <id>` (the standalone command, no trigger) is unchanged — it still opens the live SSE stream directly.
|
|
83
85
|
- `spec --json` prints the CLI's own command tree — `{specVersion, cliVersion, commands[], exitCodes, errorKinds, kindExitCodes, docsUrl?, llmsUrl?}` — DERIVED at runtime by walking the live commander tree (never a hand-maintained file, which would silently drift from reality). Each entry in `commands[]` carries its full path (e.g. `"scraps account session set"`), description, `hidden` (the legacy `scraps <verb>` aliases above), `leaf` (false for a pure namespace/group node like `scraps`/`skills`/`telemetry` — invoking one directly is a guaranteed-failing tool, not a real command), `aliases`, `arguments`, `options`, and an optional per-command `docs` deep link (e.g. every `scraps account *` command points at the account-sessions guide); `exitCodes`/`errorKinds`/`kindExitCodes` are read from the exact same source `classifyError` uses (see [Exit codes](#exit-codes)) — never a second copy. `kindExitCodes` is the inverse of `exitCodes`: `kind -> exitCode`, since exit code `1` alone is a shared bucket (`api`/`refused`/`unknown`/`in_progress`/`run_failed`/`upgrade_failed`) that `exitCodes`' flat label can't disambiguate. `docsUrl` is resolved server-first (`externalDocs.url` on the configured API base's own OpenAPI document, when declared), else derived from a KNOWN first-party `trawl.me` API host, else **omitted entirely** — never a guessed URL pointed at the wrong docs host for a self-hosted install. `llmsUrl` and every per-command `docs` link are narrower: they only ever come from that same known-host derivation, NEVER from a server-declared `docsUrl` — this CLI's own guide slugs have no reason to exist on a third party's own docs root, so a self-hosted server that declares `externalDocs.url` gets a correct top-level `docsUrl` but no `llmsUrl` and no per-command `docs` links. A `failureKind:"auth"` run's JSON payload (`doctor`/`data --errors`/`run-info`) carries the same `docs` field, by the same known-host-only rule. Without `--json`, `spec` prints one short human line (version + visible command count) pointing at `--json`.
|
|
84
86
|
- `run --json`/`trigger --json` bypass the spinner and print the raw launch/trigger payload on stdout; combined with `--watch`, every intermediate progress line stays suppressed (stdout stays pure JSON) and, once the watch reaches its outcome, exactly ONE final NDJSON line is emitted: `{"runId","status"}` (the honest terminal status — `success`/`error`/`empty`/`regression`/…), `{"runId","status":"timeout"}` on a poll timeout, or `{"runId","status":"poll_error","error"}` if the API stays unreachable for several consecutive polls — `process.exitCode` is non-zero for all three except a genuine success. `scraps watch --json` emits one raw JSON object per activity line (NDJSON) instead of the formatted `[time] message` text — there's no single final payload to wait for on a live stream.
|
|
85
87
|
|
|
@@ -238,7 +240,7 @@ Under `--json`, a failing command emits a single error envelope on stdout — `{
|
|
|
238
240
|
| `TRAWL_TELEMETRY` | Set to `0` to disable telemetry for the current session |
|
|
239
241
|
| `DO_NOT_TRACK` | Set to `1` to disable telemetry (cross-vendor convention, https://consoledonottrack.com) — same effect as `TRAWL_TELEMETRY=0` |
|
|
240
242
|
| `TRAWL_CONFIG_DIR` | Override where the config file (token, API URL, telemetry state) is stored — useful for hermetic CI runs or concurrent `trawl login`s that must not share one on-disk file |
|
|
241
|
-
| `TRAWL_TIMEOUT` | Override the per-request fetch timeout in milliseconds (default `30000`; `
|
|
243
|
+
| `TRAWL_TIMEOUT` | Override the per-request fetch timeout in milliseconds (default `30000`; `660000` for `run`/`data --fresh`/`trigger --wait`, `3600000` for `create` — this env var always wins over those longer defaults too) |
|
|
242
244
|
| `TRAWL_SKILLS_SYNC` | Set to `0` to disable the startup skills auto-sync entirely |
|
|
243
245
|
|
|
244
246
|
Session override (no `trawl login` mutation, ideal for CI/QA against another env):
|
package/dist/commands/create.js
CHANGED
|
@@ -1,11 +1,11 @@
|
|
|
1
1
|
import { Command } from 'commander';
|
|
2
2
|
import chalk from 'chalk';
|
|
3
3
|
import { spin } from '../lib/spinner.js';
|
|
4
|
-
import { api, WIZARD_TIMEOUT_MS, NetworkError } from '../lib/api.js';
|
|
5
|
-
import { json } from '../lib/format.js';
|
|
4
|
+
import { api, WIZARD_TIMEOUT_MS, ApiError, NetworkError, StreamError } from '../lib/api.js';
|
|
5
|
+
import { json, printable } from '../lib/format.js';
|
|
6
6
|
import { parseServerJson } from '../lib/json.js';
|
|
7
7
|
import { requireUrl, requireString } from '../lib/validate.js';
|
|
8
|
-
import { UsageError } from '../lib/errors.js';
|
|
8
|
+
import { UsageError, RefusalError, HUMAN_ERROR_MAX } from '../lib/errors.js';
|
|
9
9
|
import { renderPinch, pinchEnabled } from '../lib/pinch.js';
|
|
10
10
|
import { startPinchAnimation } from '../lib/pinchAnimation.js';
|
|
11
11
|
function firstRunLabel(data, autoFixEnabled) {
|
|
@@ -17,21 +17,24 @@ function firstRunLabel(data, autoFixEnabled) {
|
|
|
17
17
|
return 'failed (auto-fix retrying in the background)';
|
|
18
18
|
return 'failed';
|
|
19
19
|
}
|
|
20
|
+
const WEEKDAYS = ['Sunday', 'Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday', 'Saturday'];
|
|
20
21
|
function describeCron(cron) {
|
|
21
|
-
const match = /^(\d{1,2})\s+(\d{1,2})\s+\*\s+\*\s
|
|
22
|
+
const match = /^(\d{1,2})\s+(\d{1,2})\s+\*\s+\*\s+(\*|[0-7])$/.exec(cron.trim());
|
|
22
23
|
if (!match)
|
|
23
24
|
return cron;
|
|
24
|
-
const [, min, hour] = match;
|
|
25
|
-
|
|
25
|
+
const [, min, hour, dow] = match;
|
|
26
|
+
const time = `${hour.padStart(2, '0')}:${min.padStart(2, '0')}`;
|
|
27
|
+
return dow === '*' ? `daily ${time}` : `weekly ${WEEKDAYS[Number(dow) % 7]} ${time}`;
|
|
26
28
|
}
|
|
27
29
|
function scheduleLabel(scrap) {
|
|
28
30
|
if (!scrap)
|
|
29
31
|
return null;
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
32
|
+
const cron = typeof scrap.cron === 'string' ? printable(scrap.cron) : '';
|
|
33
|
+
if (cron !== '') {
|
|
34
|
+
const tz = (typeof scrap.cronTimezone === 'string' ? printable(scrap.cronTimezone) : '') || 'UTC';
|
|
35
|
+
return `${describeCron(cron)} ${tz} (cron ${cron})`;
|
|
33
36
|
}
|
|
34
|
-
return 'scheduled
|
|
37
|
+
return 'scheduled weekly by default (Monday 07:00 UTC)';
|
|
35
38
|
}
|
|
36
39
|
async function printDataSample(historyId) {
|
|
37
40
|
try {
|
|
@@ -45,18 +48,36 @@ async function printDataSample(historyId) {
|
|
|
45
48
|
const first = items[0];
|
|
46
49
|
if (first && typeof first === 'object') {
|
|
47
50
|
const keys = Object.keys(first);
|
|
48
|
-
console.log(chalk.dim(` Fields: `) + keys.join(', '));
|
|
51
|
+
console.log(chalk.dim(` Fields: `) + printable(keys.join(', '), HUMAN_ERROR_MAX));
|
|
49
52
|
const firstKey = keys[0];
|
|
50
53
|
if (firstKey) {
|
|
51
|
-
const
|
|
52
|
-
|
|
53
|
-
console.log(chalk.dim(` ${firstKey.slice(0, 10).padEnd(10)}: `) + val);
|
|
54
|
+
const val = printable(first[firstKey], 61);
|
|
55
|
+
console.log(chalk.dim(` ${printable(firstKey).slice(0, 10).padEnd(10)}: `) + val);
|
|
54
56
|
}
|
|
55
57
|
}
|
|
56
58
|
}
|
|
57
59
|
catch {
|
|
58
60
|
}
|
|
59
61
|
}
|
|
62
|
+
function stageLine(event) {
|
|
63
|
+
if (event.type !== 'stage')
|
|
64
|
+
return null;
|
|
65
|
+
const label = typeof event['label'] === 'string' ? printable(event['label']) : '';
|
|
66
|
+
if (label !== '')
|
|
67
|
+
return label;
|
|
68
|
+
const stage = typeof event['stage'] === 'string' ? printable(event['stage']) : '';
|
|
69
|
+
if (stage === '')
|
|
70
|
+
return null;
|
|
71
|
+
if (stage === 'probing') {
|
|
72
|
+
const tier = event['tier'];
|
|
73
|
+
const shown = typeof tier === 'string' || typeof tier === 'number' ? printable(String(tier)) : '';
|
|
74
|
+
return shown !== '' ? `Probing the site (tier ${shown})` : 'Probing the site';
|
|
75
|
+
}
|
|
76
|
+
return stage.charAt(0).toUpperCase() + stage.slice(1);
|
|
77
|
+
}
|
|
78
|
+
function mayHaveCompleted(err) {
|
|
79
|
+
return (err instanceof NetworkError || err instanceof ApiError) && err.mayHaveCompleted;
|
|
80
|
+
}
|
|
60
81
|
export const create = new Command('create')
|
|
61
82
|
.description('Create a persistent, self-healing scrap from a URL + a goal (AI-generated) — distinct from `trawl scraps create`, which is raw manual entry')
|
|
62
83
|
.argument('[url]', 'Target public URL (http/https)')
|
|
@@ -78,24 +99,41 @@ export const create = new Command('create')
|
|
|
78
99
|
goal,
|
|
79
100
|
...(opts.autofix === false && { autoFix: false }),
|
|
80
101
|
};
|
|
81
|
-
const call = () => api
|
|
102
|
+
const call = (onEvent) => api
|
|
103
|
+
.postLongRunning('/api/ai/wizard', body, { timeoutMs: WIZARD_TIMEOUT_MS, onEvent })
|
|
104
|
+
.catch((err) => {
|
|
105
|
+
if (!(err instanceof StreamError))
|
|
106
|
+
throw err;
|
|
107
|
+
throw new RefusalError(opts.json ? err.message : printable(err.message) || 'the server reported an error');
|
|
108
|
+
});
|
|
82
109
|
let data;
|
|
83
110
|
if (opts.json) {
|
|
84
|
-
data = await call();
|
|
111
|
+
data = await call(() => { });
|
|
85
112
|
}
|
|
86
113
|
else {
|
|
87
114
|
const anim = pinchEnabled() ? startPinchAnimation(`Creating a scrap from ${url}…`) : null;
|
|
115
|
+
let lastStage = '';
|
|
116
|
+
const showStage = (event) => {
|
|
117
|
+
const line = stageLine(event);
|
|
118
|
+
if (line === null || line === lastStage)
|
|
119
|
+
return;
|
|
120
|
+
lastStage = line;
|
|
121
|
+
if (anim)
|
|
122
|
+
anim.setStatus(line);
|
|
123
|
+
else
|
|
124
|
+
console.error(`${process.stderr.isTTY ? '\r\x1b[2K' : ''}${chalk.dim(` ${line}`)}`);
|
|
125
|
+
};
|
|
88
126
|
try {
|
|
89
127
|
data = anim
|
|
90
|
-
? await call()
|
|
91
|
-
: await spin(call, {
|
|
128
|
+
? await call(showStage)
|
|
129
|
+
: await spin(() => call(showStage), {
|
|
92
130
|
text: `Creating a scrap from ${url}…`,
|
|
93
131
|
successText: 'Request complete',
|
|
94
132
|
});
|
|
95
133
|
}
|
|
96
134
|
catch (err) {
|
|
97
135
|
anim?.stop();
|
|
98
|
-
if (err
|
|
136
|
+
if (mayHaveCompleted(err)) {
|
|
99
137
|
console.error(chalk.yellow('⚠ The request did not complete client-side, but the scrap may STILL have been created server-side — run `trawl list` before retrying (a retry creates a DUPLICATE scrap + burns quota).'));
|
|
100
138
|
}
|
|
101
139
|
throw err;
|
|
@@ -106,14 +144,15 @@ export const create = new Command('create')
|
|
|
106
144
|
json(data);
|
|
107
145
|
}
|
|
108
146
|
else {
|
|
109
|
-
const scrapId = data.scrap?._id;
|
|
110
|
-
const
|
|
147
|
+
const scrapId = printable(data.scrap?._id);
|
|
148
|
+
const historyId = printable(data.historyId);
|
|
149
|
+
const title = printable(data.scrap?.title ?? url);
|
|
111
150
|
const autoFixEnabled = opts.autofix !== false;
|
|
112
151
|
const schedule = scheduleLabel(data.scrap);
|
|
113
152
|
console.log(`${data.success ? chalk.green('✓') : chalk.red('✗')} ${chalk.bold(title)}`);
|
|
114
|
-
console.log(chalk.dim(` Scrap ID: `) + (scrapId
|
|
115
|
-
if (
|
|
116
|
-
console.log(chalk.dim(` History ID: `) +
|
|
153
|
+
console.log(chalk.dim(` Scrap ID: `) + (scrapId || '—'));
|
|
154
|
+
if (historyId)
|
|
155
|
+
console.log(chalk.dim(` History ID: `) + historyId);
|
|
117
156
|
console.log(chalk.dim(` First run: `) + firstRunLabel(data, autoFixEnabled));
|
|
118
157
|
if (schedule)
|
|
119
158
|
console.log(chalk.dim(` Schedule: `) + schedule);
|
|
@@ -123,8 +162,8 @@ export const create = new Command('create')
|
|
|
123
162
|
console.log(chalk.dim(` Next step: `) + `trawl data ${scrapId}` + chalk.dim(` (details: trawl get ${scrapId})`));
|
|
124
163
|
}
|
|
125
164
|
if (!data.success && data.scrap && autoFixEnabled) {
|
|
126
|
-
const pollTarget =
|
|
127
|
-
? `\`trawl data ${scrapId}\` or \`trawl run-info ${
|
|
165
|
+
const pollTarget = historyId
|
|
166
|
+
? `\`trawl data ${scrapId}\` or \`trawl run-info ${historyId}\``
|
|
128
167
|
: `\`trawl data ${scrapId}\``;
|
|
129
168
|
console.log(chalk.yellow(` Note: `) +
|
|
130
169
|
`Auto-fix is retrying in the background — do NOT re-run create; poll ${pollTarget}.`);
|
package/dist/commands/scraps.js
CHANGED
|
@@ -2,11 +2,11 @@ import { Command } from 'commander';
|
|
|
2
2
|
import chalk from 'chalk';
|
|
3
3
|
import { spin, confirmNonTTY } from '../lib/spinner.js';
|
|
4
4
|
import { api, LONG_RUN_TIMEOUT_MS } from '../lib/api.js';
|
|
5
|
-
import { table, json, formatDate } from '../lib/format.js';
|
|
5
|
+
import { table, json, jsonText, formatDate, printable } from '../lib/format.js';
|
|
6
6
|
import { parseServerJson } from '../lib/json.js';
|
|
7
7
|
import { promptPassword } from '../lib/prompt.js';
|
|
8
8
|
import { validateObjectId, requireUrl } from '../lib/validate.js';
|
|
9
|
-
import { classifyError, reportError, retryFieldsFor, UsageError, RefusalError } from '../lib/errors.js';
|
|
9
|
+
import { classifyError, reportError, retryFieldsFor, UsageError, RefusalError, HUMAN_ERROR_MAX } from '../lib/errors.js';
|
|
10
10
|
import { confirmDestructive, isInteractive } from '../lib/confirm.js';
|
|
11
11
|
import { formatDoctor, formatAutofix, fetchRunAndFix, pickRun, pickFix, isAuthWall } from './doctor.js';
|
|
12
12
|
import { renderPinch, pinchEnabled } from '../lib/pinch.js';
|
|
@@ -127,11 +127,11 @@ async function watchActivities(id, asJson) {
|
|
|
127
127
|
try {
|
|
128
128
|
const activity = parseServerJson(event);
|
|
129
129
|
if (asJson) {
|
|
130
|
-
console.log(
|
|
130
|
+
console.log(jsonText(activity));
|
|
131
131
|
continue;
|
|
132
132
|
}
|
|
133
133
|
const time = new Date(activity.createdAt).toLocaleTimeString();
|
|
134
|
-
console.log(`${chalk.dim(`[${time}]`)} ${activity.message}`);
|
|
134
|
+
console.log(`${chalk.dim(`[${time}]`)} ${printable(activity.message, HUMAN_ERROR_MAX)}`);
|
|
135
135
|
}
|
|
136
136
|
catch {
|
|
137
137
|
if (process.env['DEBUG'] && event.trim()) {
|
|
@@ -185,10 +185,10 @@ export async function pollRunProgress(id, before, opts = {}) {
|
|
|
185
185
|
const message = err instanceof Error ? err.message : String(err);
|
|
186
186
|
if (asJson) {
|
|
187
187
|
const outcome = { runId: beforeId, status: 'poll_error', error: message };
|
|
188
|
-
console.log(
|
|
188
|
+
console.log(jsonText(outcome));
|
|
189
189
|
}
|
|
190
190
|
else {
|
|
191
|
-
console.log(chalk.red(`✗ Polling failed repeatedly (${message}) — check status with: trawl scraps doctor ${id}`));
|
|
191
|
+
console.log(chalk.red(`✗ Polling failed repeatedly (${printable(message, HUMAN_ERROR_MAX)}) — check status with: trawl scraps doctor ${id}`));
|
|
192
192
|
}
|
|
193
193
|
process.exitCode = 1;
|
|
194
194
|
return;
|
|
@@ -217,7 +217,7 @@ export async function pollRunProgress(id, before, opts = {}) {
|
|
|
217
217
|
if (asJson)
|
|
218
218
|
continue;
|
|
219
219
|
const time = new Date(a.createdAt).toLocaleTimeString();
|
|
220
|
-
console.log(`${chalk.dim(`[${time}]`)} ${a.message}`);
|
|
220
|
+
console.log(`${chalk.dim(`[${time}]`)} ${printable(a.message, HUMAN_ERROR_MAX)}`);
|
|
221
221
|
}
|
|
222
222
|
}
|
|
223
223
|
catch {
|
|
@@ -227,10 +227,10 @@ export async function pollRunProgress(id, before, opts = {}) {
|
|
|
227
227
|
const isGenuineFailure = last.status === false && outcome !== 'empty' && outcome !== 'regression';
|
|
228
228
|
if (asJson) {
|
|
229
229
|
const payload = { runId: last._id, status: outcome };
|
|
230
|
-
console.log(
|
|
230
|
+
console.log(jsonText(payload));
|
|
231
231
|
}
|
|
232
232
|
else {
|
|
233
|
-
console.log(chalk.dim(`Run finished: ${outcome}`));
|
|
233
|
+
console.log(chalk.dim(`Run finished: ${printable(outcome)}`));
|
|
234
234
|
if (last.status === true && pinchEnabled()) {
|
|
235
235
|
console.log(renderPinch('celebrating'));
|
|
236
236
|
}
|
|
@@ -241,7 +241,7 @@ export async function pollRunProgress(id, before, opts = {}) {
|
|
|
241
241
|
}
|
|
242
242
|
if (asJson) {
|
|
243
243
|
const outcome = { runId: beforeId, status: 'timeout' };
|
|
244
|
-
console.log(
|
|
244
|
+
console.log(jsonText(outcome));
|
|
245
245
|
}
|
|
246
246
|
else {
|
|
247
247
|
console.log(chalk.yellow(`⚠ Timed out waiting for the run to finish — check status with: trawl scraps doctor ${id}`));
|
|
@@ -308,14 +308,14 @@ export function attachListCommand(parent, attachOpts = {}) {
|
|
|
308
308
|
catch (err) {
|
|
309
309
|
const isDebug = Boolean(cmd.optsWithGlobals().debug || process.env['DEBUG']);
|
|
310
310
|
const { exitCode, envelope } = classifyError(err);
|
|
311
|
-
const
|
|
311
|
+
const prefix = 'Failed to fetch scraps: ';
|
|
312
312
|
if (isDebug)
|
|
313
313
|
console.error(err);
|
|
314
314
|
if (opts.json) {
|
|
315
|
-
console.log(
|
|
315
|
+
console.log(jsonText({ error: { ...envelope, message: prefix + envelope.message } }));
|
|
316
316
|
}
|
|
317
317
|
else if (!isDebug) {
|
|
318
|
-
console.error(chalk.red(`✗ ${message}`));
|
|
318
|
+
console.error(chalk.red(`✗ ${prefix}${printable(envelope.message, HUMAN_ERROR_MAX)}`));
|
|
319
319
|
}
|
|
320
320
|
process.exitCode = exitCode;
|
|
321
321
|
return;
|
|
@@ -336,7 +336,9 @@ export function attachListCommand(parent, attachOpts = {}) {
|
|
|
336
336
|
'last run': lastRun(s),
|
|
337
337
|
updated: formatDate(s.updatedAt),
|
|
338
338
|
}));
|
|
339
|
-
table(tableRows, ['id', 'title', 'cron', 'status', 'health', 'last run', 'updated']
|
|
339
|
+
table(tableRows, ['id', 'title', 'cron', 'status', 'health', 'last run', 'updated'], {
|
|
340
|
+
styled: ['status', 'health'],
|
|
341
|
+
});
|
|
340
342
|
if (limit !== undefined && rows.length < totalMatched) {
|
|
341
343
|
console.log(chalk.dim(`Showing ${rows.length} of ${totalMatched} — omit --limit to see all`));
|
|
342
344
|
}
|
|
@@ -353,9 +355,9 @@ export function attachGetCommand(parent, attachOpts = {}) {
|
|
|
353
355
|
const data = await api.get(`/api/scraps/${id}`);
|
|
354
356
|
if (opts.json)
|
|
355
357
|
return json(data);
|
|
356
|
-
console.log(chalk.bold(data.title || '(untitled)'));
|
|
357
|
-
console.log(chalk.dim(` ID: `) + data._id);
|
|
358
|
-
console.log(chalk.dim(` Cron: `) + (data.cron || '—'));
|
|
358
|
+
console.log(chalk.bold(printable(data.title) || '(untitled)'));
|
|
359
|
+
console.log(chalk.dim(` ID: `) + printable(data._id));
|
|
360
|
+
console.log(chalk.dim(` Cron: `) + (printable(data.cron) || '—'));
|
|
359
361
|
console.log(chalk.dim(` Status: `) + statusIcon(lastStatus(data)));
|
|
360
362
|
console.log(chalk.dim(` Health: `) + healthBadge(data));
|
|
361
363
|
console.log(chalk.dim(` Last run: `) + lastRun(data));
|
|
@@ -369,22 +371,23 @@ function renderTierOverrideHuman(data) {
|
|
|
369
371
|
if (!ov)
|
|
370
372
|
return;
|
|
371
373
|
if (ov.refused) {
|
|
372
|
-
console.error(chalk.red(` ✗ tier ceiling override refused: ${ov.reason
|
|
373
|
-
+ chalk.dim(` (requested ${ov.requestedMaxTier
|
|
374
|
+
console.error(chalk.red(` ✗ tier ceiling override refused: ${printable(ov.reason, HUMAN_ERROR_MAX) || 'unknown'}`)
|
|
375
|
+
+ chalk.dim(` (requested ${printable(ov.requestedMaxTier) || '—'}; kept the registry cap)`));
|
|
374
376
|
}
|
|
375
377
|
else if (ov.effectiveMaxTier) {
|
|
376
|
-
|
|
377
|
-
|
|
378
|
+
const provider = printable(ov.provider);
|
|
379
|
+
console.log(chalk.green(` ✓ tier ceiling: ${printable(ov.effectiveMaxTier)}`)
|
|
380
|
+
+ chalk.dim(` (${printable(ov.reason, HUMAN_ERROR_MAX)}${provider ? `, ${provider}` : ''})`));
|
|
378
381
|
if (ov.warning)
|
|
379
|
-
console.log(chalk.yellow(` ⚠ ${ov.warning}`));
|
|
382
|
+
console.log(chalk.yellow(` ⚠ ${printable(ov.warning, HUMAN_ERROR_MAX)}`));
|
|
380
383
|
}
|
|
381
384
|
if (ov.proxyTier) {
|
|
382
385
|
if (ov.proxyTier.clamped) {
|
|
383
|
-
console.log(chalk.yellow(` ⚠ proxyTier requested ${ov.proxyTier.requested} → applied ${ov.proxyTier.effective}`)
|
|
384
|
-
+ chalk.dim(` (${ov.proxyTier.reason
|
|
386
|
+
console.log(chalk.yellow(` ⚠ proxyTier requested ${printable(ov.proxyTier.requested)} → applied ${printable(ov.proxyTier.effective)}`)
|
|
387
|
+
+ chalk.dim(` (${printable(ov.proxyTier.reason, HUMAN_ERROR_MAX) || 'capped'})`));
|
|
385
388
|
}
|
|
386
389
|
else {
|
|
387
|
-
console.log(chalk.dim(` proxyTier: `) + ov.proxyTier.effective);
|
|
390
|
+
console.log(chalk.dim(` proxyTier: `) + printable(ov.proxyTier.effective));
|
|
388
391
|
}
|
|
389
392
|
}
|
|
390
393
|
}
|
|
@@ -426,7 +429,7 @@ scraps
|
|
|
426
429
|
request: opts.request || '',
|
|
427
430
|
...(opts.description !== undefined && { description: opts.description }),
|
|
428
431
|
...(opts.tier !== undefined && { proxyTier: opts.tier }),
|
|
429
|
-
}), { text: 'Creating scrap…', successText: (d) => `Scrap created: ${chalk.bold(d._id)}` });
|
|
432
|
+
}), { text: 'Creating scrap…', successText: (d) => `Scrap created: ${chalk.bold(printable(d._id))}` });
|
|
430
433
|
const tierWasRequested = opts.tier !== undefined;
|
|
431
434
|
warnIfUnconfirmedTier(data, tierWasRequested, data._id);
|
|
432
435
|
const refused = Boolean(data._tierOverride?.refused);
|
|
@@ -438,7 +441,7 @@ scraps
|
|
|
438
441
|
json(withTierUnconfirmed(data, tierWasRequested));
|
|
439
442
|
return;
|
|
440
443
|
}
|
|
441
|
-
console.log(chalk.dim(` Title: ${data.title}`));
|
|
444
|
+
console.log(chalk.dim(` Title: ${printable(data.title)}`));
|
|
442
445
|
renderTierOverrideHuman(data);
|
|
443
446
|
if (refused)
|
|
444
447
|
process.exitCode = 1;
|
|
@@ -536,7 +539,7 @@ scraps
|
|
|
536
539
|
}
|
|
537
540
|
const data = await spin(() => api.put(`/api/scraps/${id}`, body), {
|
|
538
541
|
text: 'Updating scrap…',
|
|
539
|
-
successText: (d) => `Scrap updated: ${chalk.bold(d._id)}`,
|
|
542
|
+
successText: (d) => `Scrap updated: ${chalk.bold(printable(d._id))}`,
|
|
540
543
|
});
|
|
541
544
|
const tierWasRequested = opts.tier !== undefined || opts.forceTier !== undefined;
|
|
542
545
|
warnIfUnconfirmedTier(data, tierWasRequested, id);
|
|
@@ -554,7 +557,7 @@ scraps
|
|
|
554
557
|
if (key === 'proxyTier' || key === 'proxyMaxTier')
|
|
555
558
|
continue;
|
|
556
559
|
const src = key in shown ? shown[key] : body[key];
|
|
557
|
-
console.log(chalk.dim(` ${key}: `) +
|
|
560
|
+
console.log(chalk.dim(` ${key}: `) + (printable(src, HUMAN_ERROR_MAX) || '—'));
|
|
558
561
|
}
|
|
559
562
|
renderTierOverrideHuman(data);
|
|
560
563
|
if (refused)
|
|
@@ -569,7 +572,7 @@ export function attachRunCommand(parent, attachOpts = {}) {
|
|
|
569
572
|
.action(async (id, opts) => {
|
|
570
573
|
validateObjectId(id);
|
|
571
574
|
const beforeRun = opts.watch ? await captureBeforeRunState(id) : undefined;
|
|
572
|
-
const call = () => api.
|
|
575
|
+
const call = () => api.getLongRunning(`/api/scraps/load/${id}`, { timeoutMs: LONG_RUN_TIMEOUT_MS });
|
|
573
576
|
const data = opts.json
|
|
574
577
|
? await call()
|
|
575
578
|
: await spin(call, { text: 'Launching scrap…', successText: 'Scrap launched' });
|
|
@@ -595,7 +598,7 @@ function renderScrapItems(items, asJson) {
|
|
|
595
598
|
console.log(chalk.bold('Last run data:'));
|
|
596
599
|
console.log(chalk.dim(` Items: ${items.length}`));
|
|
597
600
|
if (items.length > 0 && typeof items[0] === 'object' && items[0] !== null) {
|
|
598
|
-
console.log(chalk.dim(` First item keys: ${Object.keys(items[0]).join(', ')}`));
|
|
601
|
+
console.log(chalk.dim(` First item keys: ${printable(Object.keys(items[0]).join(', '), HUMAN_ERROR_MAX)}`));
|
|
599
602
|
}
|
|
600
603
|
console.log(chalk.dim(' Use --json for full output.'));
|
|
601
604
|
}
|
|
@@ -604,7 +607,7 @@ function reportDataState(message, exitCode, kind, wantsJson, retryOverride) {
|
|
|
604
607
|
if (wantsJson) {
|
|
605
608
|
const { retryable, next } = retryOverride ?? retryFieldsFor(kind);
|
|
606
609
|
const envelope = { message, kind, retryable, ...(next ? { next } : {}) };
|
|
607
|
-
console.log(
|
|
610
|
+
console.log(jsonText({ error: envelope }));
|
|
608
611
|
}
|
|
609
612
|
process.exitCode = exitCode;
|
|
610
613
|
}
|
|
@@ -640,7 +643,7 @@ export function attachDataCommand(parent, attachOpts = {}) {
|
|
|
640
643
|
return;
|
|
641
644
|
}
|
|
642
645
|
if (opts.fresh) {
|
|
643
|
-
const loaded = await spin(() => api.
|
|
646
|
+
const loaded = await spin(() => api.getLongRunning(`/api/scraps/load/${id}`, { timeoutMs: LONG_RUN_TIMEOUT_MS }), {
|
|
644
647
|
text: 'Launching a fresh scrap run (consumes execute quota)…',
|
|
645
648
|
successText: 'Fresh run complete',
|
|
646
649
|
});
|
|
@@ -857,7 +860,7 @@ export function attachTriggerCommand(parent, attachOpts = {}) {
|
|
|
857
860
|
validateObjectId(id);
|
|
858
861
|
const beforeRun = opts.watch ? await captureBeforeRunState(id) : undefined;
|
|
859
862
|
const path = opts.wait ? `/api/scraps/worker/${id}` : `/api/scraps/worker/${id}?wait=false`;
|
|
860
|
-
const call = () => (opts.wait ? api.
|
|
863
|
+
const call = () => (opts.wait ? api.postLongRunning(path, undefined, { timeoutMs: LONG_RUN_TIMEOUT_MS }) : api.post(path));
|
|
861
864
|
const data = opts.json
|
|
862
865
|
? await call()
|
|
863
866
|
: await spin(call, {
|
|
@@ -1121,7 +1124,7 @@ accountSession
|
|
|
1121
1124
|
reportCaptureFailure({ ok: false, reason: 'non_interactive', message: nonInteractiveReason }, opts);
|
|
1122
1125
|
return;
|
|
1123
1126
|
}
|
|
1124
|
-
console.error(chalk.dim(`Chrome will open at ${scrap.url} for you to log in there (2FA included); capture completes when you press Enter back in this terminal.`));
|
|
1127
|
+
console.error(chalk.dim(`Chrome will open at ${printable(scrap.url)} for you to log in there (2FA included); capture completes when you press Enter back in this terminal.`));
|
|
1125
1128
|
const result = await captureSession(scrap.url, opts.chrome ? { findChrome: () => opts.chrome ?? null } : {});
|
|
1126
1129
|
if (!result.ok) {
|
|
1127
1130
|
reportCaptureFailure(result, opts);
|
package/dist/lib/api.d.ts
CHANGED
|
@@ -2,6 +2,10 @@ export declare class ApiError extends Error {
|
|
|
2
2
|
status: number;
|
|
3
3
|
next?: string[] | undefined;
|
|
4
4
|
constructor(status: number, message: string, next?: string[] | undefined);
|
|
5
|
+
mayHaveCompleted: boolean;
|
|
6
|
+
}
|
|
7
|
+
export declare class StreamError extends Error {
|
|
8
|
+
constructor(message: string);
|
|
5
9
|
}
|
|
6
10
|
export declare class NetworkError extends Error {
|
|
7
11
|
readonly mayHaveCompleted: boolean;
|
|
@@ -13,13 +17,20 @@ export declare class AuthError extends Error {
|
|
|
13
17
|
}
|
|
14
18
|
export declare function notLoggedInError(): AuthError;
|
|
15
19
|
export declare function apiKeyUnsupportedError(command: string): AuthError;
|
|
16
|
-
export declare const LONG_RUN_TIMEOUT_MS =
|
|
17
|
-
export declare const WIZARD_TIMEOUT_MS =
|
|
20
|
+
export declare const LONG_RUN_TIMEOUT_MS = 660000;
|
|
21
|
+
export declare const WIZARD_TIMEOUT_MS = 3600000;
|
|
18
22
|
export interface RequestOptions {
|
|
19
23
|
timeoutMs?: number;
|
|
20
24
|
}
|
|
21
25
|
export declare function effectivePort(u: URL): string;
|
|
22
26
|
export declare function sameEffectivePort(from: URL, next: URL): boolean;
|
|
27
|
+
export interface StreamEvent {
|
|
28
|
+
type: string;
|
|
29
|
+
[key: string]: unknown;
|
|
30
|
+
}
|
|
31
|
+
export interface LongRunningOptions extends RequestOptions {
|
|
32
|
+
onEvent?: (event: StreamEvent) => void;
|
|
33
|
+
}
|
|
23
34
|
export declare const api: {
|
|
24
35
|
get: <T>(path: string, opts?: RequestOptions) => Promise<T>;
|
|
25
36
|
publicGet: <T>(path: string, opts?: RequestOptions) => Promise<T>;
|
|
@@ -28,7 +39,8 @@ export declare const api: {
|
|
|
28
39
|
}) => Promise<T | null>;
|
|
29
40
|
getText: (path: string, opts?: RequestOptions) => Promise<string>;
|
|
30
41
|
post: <T>(path: string, body?: unknown, opts?: RequestOptions) => Promise<T>;
|
|
31
|
-
postLongRunning: <T>(path: string, body?: unknown, opts?:
|
|
42
|
+
postLongRunning: <T>(path: string, body?: unknown, opts?: LongRunningOptions) => Promise<T>;
|
|
43
|
+
getLongRunning: <T>(path: string, opts?: RequestOptions) => Promise<T>;
|
|
32
44
|
put: <T>(path: string, body?: unknown, opts?: RequestOptions) => Promise<T>;
|
|
33
45
|
delete: <T>(path: string, opts?: RequestOptions) => Promise<T>;
|
|
34
46
|
upload: <T>(path: string, formData: FormData, opts?: RequestOptions) => Promise<T>;
|
package/dist/lib/api.js
CHANGED
|
@@ -4,7 +4,9 @@ import { dirname, resolve } from 'node:path';
|
|
|
4
4
|
import { request as httpRequest, } from 'node:http';
|
|
5
5
|
import { request as httpsRequest } from 'node:https';
|
|
6
6
|
import { brotliDecompressSync, gunzipSync, inflateRawSync, inflateSync } from 'node:zlib';
|
|
7
|
+
import { StringDecoder } from 'node:string_decoder';
|
|
7
8
|
import { getApiUrl, getToken, getAuthMode, getLiveAuthEnvVar } from './config.js';
|
|
9
|
+
import { printable } from './format.js';
|
|
8
10
|
import { parseServerJson } from './json.js';
|
|
9
11
|
const __dirname = dirname(fileURLToPath(import.meta.url));
|
|
10
12
|
const pkg = JSON.parse(readFileSync(resolve(__dirname, '../../package.json'), 'utf8'));
|
|
@@ -18,6 +20,13 @@ export class ApiError extends Error {
|
|
|
18
20
|
this.next = next;
|
|
19
21
|
this.name = 'ApiError';
|
|
20
22
|
}
|
|
23
|
+
mayHaveCompleted = false;
|
|
24
|
+
}
|
|
25
|
+
export class StreamError extends Error {
|
|
26
|
+
constructor(message) {
|
|
27
|
+
super(message);
|
|
28
|
+
this.name = 'StreamError';
|
|
29
|
+
}
|
|
21
30
|
}
|
|
22
31
|
export class NetworkError extends Error {
|
|
23
32
|
mayHaveCompleted;
|
|
@@ -64,8 +73,8 @@ function authHeaders(token) {
|
|
|
64
73
|
return getAuthMode(token) === 'apiKey' ? { Authorization: `Bearer ${token}` } : { Cookie: `TOKEN=${token}` };
|
|
65
74
|
}
|
|
66
75
|
const DEFAULT_TIMEOUT_MS = 30_000;
|
|
67
|
-
export const LONG_RUN_TIMEOUT_MS =
|
|
68
|
-
export const WIZARD_TIMEOUT_MS =
|
|
76
|
+
export const LONG_RUN_TIMEOUT_MS = 660_000;
|
|
77
|
+
export const WIZARD_TIMEOUT_MS = 3_600_000;
|
|
69
78
|
function getTimeoutMs(overrideMs) {
|
|
70
79
|
const raw = process.env['TRAWL_TIMEOUT']?.trim();
|
|
71
80
|
if (raw) {
|
|
@@ -373,6 +382,8 @@ const MAX_REDIRECTS = 3;
|
|
|
373
382
|
const REDIRECT_KEEPS_METHOD = new Set([307, 308]);
|
|
374
383
|
const REDIRECT_DROPS_BODY = new Set([301, 302, 303]);
|
|
375
384
|
const MAX_DECODED_BYTES = 64 * 1024 * 1024;
|
|
385
|
+
const MAX_BODY_BYTES = MAX_DECODED_BYTES;
|
|
386
|
+
const MAX_STREAM_LINE_CHARS = 32 * 1024 * 1024;
|
|
376
387
|
const DECODE_LIMIT = { maxOutputLength: MAX_DECODED_BYTES };
|
|
377
388
|
function isOutputOverflow(err) {
|
|
378
389
|
return err?.code === 'ERR_BUFFER_TOO_LARGE';
|
|
@@ -393,6 +404,7 @@ const CONTENT_DECODERS = Object.freeze({
|
|
|
393
404
|
deflate: inflateEitherForm,
|
|
394
405
|
br: (body) => brotliDecompressSync(body, DECODE_LIMIT),
|
|
395
406
|
});
|
|
407
|
+
const ACCEPT_ENCODING = 'gzip, deflate, br';
|
|
396
408
|
function decodeBody(body, contentEncoding) {
|
|
397
409
|
const encoding = (contentEncoding ?? '').trim().toLowerCase();
|
|
398
410
|
if (encoding === '' || encoding === 'identity')
|
|
@@ -412,7 +424,35 @@ function decodeBody(body, contentEncoding) {
|
|
|
412
424
|
return { reason: `Content-Encoding "${encoding}" failed to decode (${err.message})` };
|
|
413
425
|
}
|
|
414
426
|
}
|
|
415
|
-
function
|
|
427
|
+
function createLineReader(onLine) {
|
|
428
|
+
let pending = '';
|
|
429
|
+
const emit = (line) => {
|
|
430
|
+
if (line.length > MAX_STREAM_LINE_CHARS)
|
|
431
|
+
return 'overflow';
|
|
432
|
+
return line.trim() !== '' && onLine(line) ? 'stop' : 'more';
|
|
433
|
+
};
|
|
434
|
+
return {
|
|
435
|
+
push(text) {
|
|
436
|
+
let start = 0;
|
|
437
|
+
for (let newline = text.indexOf('\n'); newline !== -1; newline = text.indexOf('\n', start)) {
|
|
438
|
+
const line = pending + text.slice(start, newline);
|
|
439
|
+
pending = '';
|
|
440
|
+
start = newline + 1;
|
|
441
|
+
const verdict = emit(line);
|
|
442
|
+
if (verdict !== 'more')
|
|
443
|
+
return verdict;
|
|
444
|
+
}
|
|
445
|
+
pending += text.slice(start);
|
|
446
|
+
return pending.length > MAX_STREAM_LINE_CHARS ? 'overflow' : 'more';
|
|
447
|
+
},
|
|
448
|
+
end() {
|
|
449
|
+
const line = pending;
|
|
450
|
+
pending = '';
|
|
451
|
+
return emit(line);
|
|
452
|
+
},
|
|
453
|
+
};
|
|
454
|
+
}
|
|
455
|
+
function rawRequest(method, url, payload, headers, timeoutMs, onStreamLine) {
|
|
416
456
|
return new Promise((resolve, reject) => {
|
|
417
457
|
let target;
|
|
418
458
|
try {
|
|
@@ -426,17 +466,23 @@ function postRaw(url, payload, headers, timeoutMs) {
|
|
|
426
466
|
reject(connectionNetworkError(url, Object.assign(new Error(`unsupported protocol "${target.protocol}"`), { code: 'ERR_INVALID_PROTOCOL' })));
|
|
427
467
|
return;
|
|
428
468
|
}
|
|
469
|
+
const streaming = onStreamLine !== undefined;
|
|
429
470
|
let settled = false;
|
|
430
471
|
let timedOut = false;
|
|
431
472
|
let mayHaveCompleted = false;
|
|
432
473
|
let timer;
|
|
433
474
|
let active;
|
|
475
|
+
const hops = [];
|
|
476
|
+
const release = () => {
|
|
477
|
+
for (const hop of hops)
|
|
478
|
+
hop.destroy();
|
|
479
|
+
};
|
|
434
480
|
const settleWith = (err) => {
|
|
435
481
|
if (settled)
|
|
436
482
|
return;
|
|
437
483
|
settled = true;
|
|
438
484
|
clearTimeout(timer);
|
|
439
|
-
|
|
485
|
+
release();
|
|
440
486
|
reject(err);
|
|
441
487
|
};
|
|
442
488
|
const fail = (err) => {
|
|
@@ -447,58 +493,116 @@ function postRaw(url, payload, headers, timeoutMs) {
|
|
|
447
493
|
return;
|
|
448
494
|
settled = true;
|
|
449
495
|
clearTimeout(timer);
|
|
496
|
+
release();
|
|
450
497
|
resolve(value);
|
|
451
498
|
};
|
|
452
|
-
const
|
|
499
|
+
const lineTooLong = () => new NetworkError(`The response stream from ${url} sent a line longer than ${MAX_STREAM_LINE_CHARS / (1024 * 1024)} MB — refusing to buffer it`, mayHaveCompleted);
|
|
500
|
+
const onResponse = (res, from, hopsLeft, failThisHop, hop) => {
|
|
453
501
|
const status = res.statusCode ?? 0;
|
|
454
502
|
const statusText = res.statusMessage ?? '';
|
|
455
503
|
res.on('error', failThisHop);
|
|
456
|
-
if (REDIRECT_KEEPS_METHOD.has(status)) {
|
|
457
|
-
res.resume();
|
|
504
|
+
if (REDIRECT_KEEPS_METHOD.has(status) || (method === 'GET' && REDIRECT_DROPS_BODY.has(status))) {
|
|
458
505
|
const location = res.headers.location;
|
|
459
506
|
const next = location ? safeUrl(location, from) : null;
|
|
460
507
|
const refusal = !next
|
|
461
508
|
? 'it carries no usable Location header'
|
|
462
509
|
: next.protocol !== 'http:' && next.protocol !== 'https:'
|
|
463
510
|
? `its target uses an unsupported protocol ("${next.protocol}")`
|
|
464
|
-
: next
|
|
465
|
-
?
|
|
466
|
-
:
|
|
467
|
-
? `its target is a different
|
|
468
|
-
: from
|
|
469
|
-
?
|
|
470
|
-
:
|
|
471
|
-
?
|
|
472
|
-
:
|
|
511
|
+
: introducesUserinfo(from, next)
|
|
512
|
+
? 'its target embeds credentials in the URL (user:password@) — this request carries your credential'
|
|
513
|
+
: next.hostname !== from.hostname
|
|
514
|
+
? `its target is a different host (${next.hostname}) — this request carries your credential`
|
|
515
|
+
: !sameEffectivePort(from, next)
|
|
516
|
+
? `its target is a different port (${effectivePort(next)}) — this request carries your credential`
|
|
517
|
+
: from.protocol === 'https:' && next.protocol === 'http:'
|
|
518
|
+
? 'it downgrades https to http — this request carries your credential'
|
|
519
|
+
: hopsLeft <= 0
|
|
520
|
+
? `it exceeds ${MAX_REDIRECTS} redirects`
|
|
521
|
+
: null;
|
|
473
522
|
if (refusal || !next) {
|
|
474
|
-
settleWith(new ApiError(status, `${status} ${statusText}: refused to follow this redirect because ${refusal}. Point TRAWL_API_URL at the final URL.`));
|
|
523
|
+
settleWith(new ApiError(status, `${status} ${printable(statusText)}: refused to follow this redirect because ${refusal}. Point TRAWL_API_URL at the final URL.`));
|
|
475
524
|
return;
|
|
476
525
|
}
|
|
477
526
|
send(next, hopsLeft - 1);
|
|
527
|
+
hop.destroy();
|
|
478
528
|
return;
|
|
479
529
|
}
|
|
480
530
|
if (REDIRECT_DROPS_BODY.has(status)) {
|
|
481
|
-
res.
|
|
482
|
-
settleWith(new ApiError(status, `${status} ${statusText}: this endpoint redirected the POST to ${res.headers.location ?? 'an unspecified location'}, and that status rewrites it to a bodyless GET — only 307/308 are followed. Point TRAWL_API_URL at the final URL.`));
|
|
531
|
+
settleWith(new ApiError(status, `${status} ${printable(statusText)}: this endpoint redirected the POST to ${printable(res.headers.location ?? 'an unspecified location')}, and that status rewrites it to a bodyless GET — only 307/308 are followed. Point TRAWL_API_URL at the final URL.`));
|
|
483
532
|
return;
|
|
484
533
|
}
|
|
485
534
|
mayHaveCompleted = status >= 200 && status < 300;
|
|
486
|
-
const
|
|
487
|
-
res.
|
|
488
|
-
|
|
489
|
-
const
|
|
490
|
-
|
|
491
|
-
|
|
492
|
-
|
|
493
|
-
|
|
494
|
-
|
|
495
|
-
|
|
535
|
+
const lineSink = mayHaveCompleted && /ndjson/i.test(String(res.headers['content-type'] ?? '')) ? onStreamLine : undefined;
|
|
536
|
+
const encoding = String(res.headers['content-encoding'] ?? '').trim().toLowerCase();
|
|
537
|
+
if (lineSink && (encoding === '' || encoding === 'identity')) {
|
|
538
|
+
const decoder = new StringDecoder('utf8');
|
|
539
|
+
const reader = createLineReader(lineSink);
|
|
540
|
+
const finish = (verdict) => {
|
|
541
|
+
if (verdict === 'overflow')
|
|
542
|
+
settleWith(lineTooLong());
|
|
543
|
+
else
|
|
544
|
+
done({ status, statusText, headers: toHeaders(res.headers), text: '', streamed: true });
|
|
545
|
+
};
|
|
546
|
+
res.on('data', (chunk) => {
|
|
547
|
+
if (settled)
|
|
548
|
+
return;
|
|
549
|
+
const verdict = reader.push(decoder.write(chunk));
|
|
550
|
+
if (verdict !== 'more')
|
|
551
|
+
finish(verdict);
|
|
552
|
+
});
|
|
553
|
+
res.on('end', () => {
|
|
554
|
+
if (settled)
|
|
555
|
+
return;
|
|
556
|
+
let verdict = reader.push(decoder.end());
|
|
557
|
+
if (verdict === 'more')
|
|
558
|
+
verdict = reader.end();
|
|
559
|
+
finish(verdict);
|
|
560
|
+
});
|
|
561
|
+
}
|
|
562
|
+
else {
|
|
563
|
+
const chunks = [];
|
|
564
|
+
let received = 0;
|
|
565
|
+
res.on('data', (chunk) => {
|
|
566
|
+
if (settled)
|
|
567
|
+
return;
|
|
568
|
+
received += chunk.length;
|
|
569
|
+
if (received > MAX_BODY_BYTES) {
|
|
570
|
+
settleWith(new NetworkError(`Response from ${url} (HTTP ${status}) is larger than the ${MAX_BODY_BYTES / (1024 * 1024)} MB limit — refusing to buffer it`, mayHaveCompleted));
|
|
571
|
+
return;
|
|
572
|
+
}
|
|
573
|
+
chunks.push(chunk);
|
|
574
|
+
});
|
|
575
|
+
res.on('end', () => {
|
|
576
|
+
if (settled)
|
|
577
|
+
return;
|
|
578
|
+
const decoded = decodeBody(Buffer.concat(chunks), res.headers['content-encoding']);
|
|
579
|
+
if (!Buffer.isBuffer(decoded)) {
|
|
580
|
+
settleWith(new NetworkError(`Unreadable response from ${url} (HTTP ${status}): ${printable(decoded.reason)}${streaming ? ' — the server ignored "Accept-Encoding: identity"' : ''}`, mayHaveCompleted));
|
|
581
|
+
return;
|
|
582
|
+
}
|
|
583
|
+
const text = decoded.toString('utf8');
|
|
584
|
+
if (lineSink) {
|
|
585
|
+
const reader = createLineReader(lineSink);
|
|
586
|
+
let verdict = reader.push(text);
|
|
587
|
+
if (verdict === 'more')
|
|
588
|
+
verdict = reader.end();
|
|
589
|
+
if (verdict === 'overflow')
|
|
590
|
+
settleWith(lineTooLong());
|
|
591
|
+
else
|
|
592
|
+
done({ status, statusText, headers: toHeaders(res.headers), text: '', streamed: true });
|
|
593
|
+
return;
|
|
594
|
+
}
|
|
595
|
+
done({ status, statusText, headers: toHeaders(res.headers), text });
|
|
596
|
+
});
|
|
597
|
+
}
|
|
496
598
|
res.on('close', () => {
|
|
497
599
|
if (!res.complete)
|
|
498
600
|
failThisHop(new Error('response aborted before it finished'));
|
|
499
601
|
});
|
|
500
602
|
};
|
|
501
603
|
const send = (to, hopsLeft) => {
|
|
604
|
+
if (settled)
|
|
605
|
+
return;
|
|
502
606
|
mayHaveCompleted = false;
|
|
503
607
|
const transport = to.protocol === 'http:' ? httpRequest : httpsRequest;
|
|
504
608
|
let req;
|
|
@@ -508,13 +612,14 @@ function postRaw(url, payload, headers, timeoutMs) {
|
|
|
508
612
|
fail(err);
|
|
509
613
|
};
|
|
510
614
|
try {
|
|
511
|
-
req = transport(to, { method
|
|
615
|
+
req = transport(to, { method, headers: { 'Accept-Encoding': streaming ? 'identity' : ACCEPT_ENCODING, ...headers } }, (res) => onResponse(res, to, hopsLeft, failThisHop, req));
|
|
512
616
|
}
|
|
513
617
|
catch (err) {
|
|
514
618
|
fail(err);
|
|
515
619
|
return;
|
|
516
620
|
}
|
|
517
621
|
active = req;
|
|
622
|
+
hops.push(req);
|
|
518
623
|
req.on('error', failThisHop);
|
|
519
624
|
if (payload !== undefined)
|
|
520
625
|
req.write(payload);
|
|
@@ -535,6 +640,10 @@ export function sameEffectivePort(from, next) {
|
|
|
535
640
|
return true;
|
|
536
641
|
return effectivePort(from) === effectivePort(next);
|
|
537
642
|
}
|
|
643
|
+
function introducesUserinfo(from, next) {
|
|
644
|
+
return ((next.username !== '' || next.password !== '') &&
|
|
645
|
+
(next.username !== from.username || next.password !== from.password));
|
|
646
|
+
}
|
|
538
647
|
function safeUrl(raw, base) {
|
|
539
648
|
try {
|
|
540
649
|
return new URL(raw, base);
|
|
@@ -543,20 +652,71 @@ function safeUrl(raw, base) {
|
|
|
543
652
|
return null;
|
|
544
653
|
}
|
|
545
654
|
}
|
|
546
|
-
|
|
655
|
+
const GATEWAY_CUT_STATUSES = new Set([502, 504]);
|
|
656
|
+
async function longRunning(method, path, body, reqOpts = {}) {
|
|
547
657
|
const token = getToken();
|
|
548
658
|
if (!token)
|
|
549
659
|
throw notLoggedInError();
|
|
550
660
|
const url = `${getApiUrl()}${path}`;
|
|
551
661
|
const timeoutMs = getTimeoutMs(reqOpts.timeoutMs);
|
|
552
662
|
const payload = body ? JSON.stringify(body) : undefined;
|
|
553
|
-
const
|
|
663
|
+
const { onEvent } = reqOpts;
|
|
664
|
+
let terminal;
|
|
665
|
+
const onStreamLine = onEvent
|
|
666
|
+
? (line) => {
|
|
667
|
+
if (terminal)
|
|
668
|
+
return true;
|
|
669
|
+
if (!line.trimStart().startsWith('{'))
|
|
670
|
+
return false;
|
|
671
|
+
let event;
|
|
672
|
+
try {
|
|
673
|
+
event = parseServerJson(line);
|
|
674
|
+
}
|
|
675
|
+
catch {
|
|
676
|
+
return false;
|
|
677
|
+
}
|
|
678
|
+
if (event === null || typeof event !== 'object' || typeof event.type !== 'string')
|
|
679
|
+
return false;
|
|
680
|
+
const { type, ...rest } = event;
|
|
681
|
+
if (type === 'done') {
|
|
682
|
+
terminal = { event: 'done', payload: rest };
|
|
683
|
+
return true;
|
|
684
|
+
}
|
|
685
|
+
if (type === 'error') {
|
|
686
|
+
const message = rest['message'];
|
|
687
|
+
terminal = { event: 'error', message: typeof message === 'string' && message ? message : 'the server reported an error' };
|
|
688
|
+
return true;
|
|
689
|
+
}
|
|
690
|
+
try {
|
|
691
|
+
onEvent({ type, ...rest });
|
|
692
|
+
}
|
|
693
|
+
catch {
|
|
694
|
+
}
|
|
695
|
+
return false;
|
|
696
|
+
}
|
|
697
|
+
: undefined;
|
|
698
|
+
const raw = await rawRequest(method, url, payload, {
|
|
554
699
|
'Content-Type': 'application/json',
|
|
555
700
|
'User-Agent': USER_AGENT,
|
|
701
|
+
...(onEvent ? { Accept: 'application/x-ndjson' } : {}),
|
|
556
702
|
...(payload !== undefined ? { 'Content-Length': String(Buffer.byteLength(payload)) } : {}),
|
|
557
703
|
...authHeaders(token),
|
|
558
|
-
}, timeoutMs);
|
|
559
|
-
|
|
704
|
+
}, timeoutMs, onStreamLine);
|
|
705
|
+
try {
|
|
706
|
+
await throwIfError(toResponse(raw), false, getAuthMode(token));
|
|
707
|
+
}
|
|
708
|
+
catch (err) {
|
|
709
|
+
if (err instanceof ApiError && GATEWAY_CUT_STATUSES.has(err.status))
|
|
710
|
+
err.mayHaveCompleted = true;
|
|
711
|
+
throw err;
|
|
712
|
+
}
|
|
713
|
+
if (raw.streamed) {
|
|
714
|
+
if (terminal?.event === 'done')
|
|
715
|
+
return terminal.payload;
|
|
716
|
+
if (terminal?.event === 'error')
|
|
717
|
+
throw new StreamError(terminal.message);
|
|
718
|
+
throw new NetworkError(`The response stream from ${url} ended before a result arrived (no "done" or "error" event)`, true);
|
|
719
|
+
}
|
|
560
720
|
try {
|
|
561
721
|
if (!raw.text)
|
|
562
722
|
return {};
|
|
@@ -567,6 +727,8 @@ async function postLongRunning(path, body, reqOpts = {}) {
|
|
|
567
727
|
return parsed;
|
|
568
728
|
}
|
|
569
729
|
catch {
|
|
730
|
+
if (onEvent === undefined)
|
|
731
|
+
throw new Error('Invalid JSON in server response');
|
|
570
732
|
throw new NetworkError('Invalid JSON in server response', true);
|
|
571
733
|
}
|
|
572
734
|
}
|
|
@@ -595,7 +757,8 @@ export const api = {
|
|
|
595
757
|
method: 'POST',
|
|
596
758
|
body: body ? JSON.stringify(body) : undefined,
|
|
597
759
|
}, opts),
|
|
598
|
-
postLongRunning: (path, body, opts) =>
|
|
760
|
+
postLongRunning: (path, body, opts) => longRunning('POST', path, body, opts),
|
|
761
|
+
getLongRunning: (path, opts) => longRunning('GET', path, undefined, opts),
|
|
599
762
|
put: (path, body, opts) => request(path, {
|
|
600
763
|
method: 'PUT',
|
|
601
764
|
body: body ? JSON.stringify(body) : undefined,
|
package/dist/lib/errors.d.ts
CHANGED
|
@@ -1,3 +1,4 @@
|
|
|
1
|
+
import { HUMAN_ERROR_MAX } from './format.js';
|
|
1
2
|
export declare class UsageError extends Error {
|
|
2
3
|
constructor(message: string);
|
|
3
4
|
}
|
|
@@ -37,6 +38,7 @@ export declare function retryFieldsFor(kind: string): {
|
|
|
37
38
|
next?: string[];
|
|
38
39
|
};
|
|
39
40
|
export declare function classifyError(err: unknown): ClassifiedError;
|
|
41
|
+
export { HUMAN_ERROR_MAX };
|
|
40
42
|
export declare function reportError(err: unknown, opts?: {
|
|
41
43
|
json?: boolean;
|
|
42
44
|
quiet?: boolean;
|
package/dist/lib/errors.js
CHANGED
|
@@ -1,5 +1,6 @@
|
|
|
1
1
|
import chalk from 'chalk';
|
|
2
2
|
import { ApiError, AuthError, NetworkError } from './api.js';
|
|
3
|
+
import { printable, jsonText, HUMAN_ERROR_MAX } from './format.js';
|
|
3
4
|
export class UsageError extends Error {
|
|
4
5
|
constructor(message) {
|
|
5
6
|
super(message);
|
|
@@ -103,13 +104,14 @@ export function classifyError(err) {
|
|
|
103
104
|
}
|
|
104
105
|
return { exitCode: EXIT_CODES.UNKNOWN, envelope: { message, kind: 'unknown', ...retryFieldsFor('unknown') } };
|
|
105
106
|
}
|
|
107
|
+
export { HUMAN_ERROR_MAX };
|
|
106
108
|
export function reportError(err, opts = {}) {
|
|
107
109
|
const { exitCode, envelope } = classifyError(err);
|
|
108
110
|
if (opts.json) {
|
|
109
|
-
console.log(
|
|
111
|
+
console.log(jsonText({ error: envelope }));
|
|
110
112
|
}
|
|
111
113
|
else if (!opts.quiet) {
|
|
112
|
-
console.error(chalk.red('✗ ' + envelope.message));
|
|
114
|
+
console.error(chalk.red('✗ ' + printable(envelope.message, HUMAN_ERROR_MAX)));
|
|
113
115
|
}
|
|
114
116
|
return exitCode;
|
|
115
117
|
}
|
package/dist/lib/format.d.ts
CHANGED
|
@@ -1,3 +1,8 @@
|
|
|
1
|
-
export declare
|
|
1
|
+
export declare const HUMAN_ERROR_MAX = 1000;
|
|
2
|
+
export declare function table(rows: Record<string, unknown>[], columns: string[], opts?: {
|
|
3
|
+
styled?: readonly string[];
|
|
4
|
+
}): void;
|
|
5
|
+
export declare function jsonText(data: unknown, space?: number): string;
|
|
2
6
|
export declare function json(data: unknown): void;
|
|
7
|
+
export declare function printable(text: unknown, max?: number): string;
|
|
3
8
|
export declare function formatDate(value: string | number | Date | null | undefined): string;
|
package/dist/lib/format.js
CHANGED
|
@@ -1,27 +1,47 @@
|
|
|
1
1
|
import chalk from 'chalk';
|
|
2
2
|
import stripAnsi from 'strip-ansi';
|
|
3
|
-
export
|
|
3
|
+
export const HUMAN_ERROR_MAX = 1000;
|
|
4
|
+
export function table(rows, columns, opts = {}) {
|
|
4
5
|
if (!rows.length) {
|
|
5
6
|
console.log(chalk.dim('No results.'));
|
|
6
7
|
return;
|
|
7
8
|
}
|
|
8
|
-
const
|
|
9
|
+
const styled = new Set(opts.styled);
|
|
10
|
+
const cells = rows.map((row) => columns.map((col) => (styled.has(col) ? String(row[col] ?? '') : printable(row[col], HUMAN_ERROR_MAX))));
|
|
11
|
+
const visibleLength = (cell) => stripAnsi(cell).length;
|
|
12
|
+
const widths = columns.map((col, i) => Math.max(col.length, ...cells.map((row) => visibleLength(row[i]))));
|
|
9
13
|
const header = columns.map((col, i) => chalk.bold(col.padEnd(widths[i]))).join(' ');
|
|
10
14
|
console.log(header);
|
|
11
15
|
console.log(chalk.dim('─'.repeat(header.replace(/\x1b\[[0-9;]*m/g, '').length)));
|
|
12
|
-
for (const row of
|
|
13
|
-
const line =
|
|
14
|
-
.map((
|
|
15
|
-
const raw = String(row[col] ?? '');
|
|
16
|
-
const visLen = stripAnsi(raw).length;
|
|
17
|
-
return raw + ' '.repeat(Math.max(0, widths[i] - visLen));
|
|
18
|
-
})
|
|
16
|
+
for (const row of cells) {
|
|
17
|
+
const line = row
|
|
18
|
+
.map((cell, i) => cell + ' '.repeat(Math.max(0, widths[i] - visibleLength(cell))))
|
|
19
19
|
.join(' ');
|
|
20
20
|
console.log(line);
|
|
21
21
|
}
|
|
22
22
|
}
|
|
23
|
+
const RAW_CONTROL_CHARS = /[\u007f-\u009f]/g;
|
|
24
|
+
export function jsonText(data, space) {
|
|
25
|
+
const text = JSON.stringify(data, null, space);
|
|
26
|
+
if (typeof text !== 'string')
|
|
27
|
+
return text;
|
|
28
|
+
return text.replace(RAW_CONTROL_CHARS, (c) => `\\u${c.charCodeAt(0).toString(16).padStart(4, '0')}`);
|
|
29
|
+
}
|
|
23
30
|
export function json(data) {
|
|
24
|
-
console.log(
|
|
31
|
+
console.log(jsonText(data, 2));
|
|
32
|
+
}
|
|
33
|
+
const INPUT_SPAN = 8;
|
|
34
|
+
export function printable(text, max = 200) {
|
|
35
|
+
const value = typeof text === 'string' ? text : text == null ? '' : String(text);
|
|
36
|
+
const clean = stripAnsi(value.slice(0, max * INPUT_SPAN))
|
|
37
|
+
.replace(/[\t\n\v\f\r\u0085]+/g, ' ')
|
|
38
|
+
.replace(/[\u0000-\u001f\u007f-\u009f]/g, '')
|
|
39
|
+
.trim();
|
|
40
|
+
if (clean.length <= max)
|
|
41
|
+
return clean;
|
|
42
|
+
const cut = clean.slice(0, Math.max(0, max - 1));
|
|
43
|
+
const last = cut.charCodeAt(cut.length - 1);
|
|
44
|
+
return `${last >= 0xd800 && last <= 0xdbff ? cut.slice(0, -1) : cut}…`;
|
|
25
45
|
}
|
|
26
46
|
export function formatDate(value) {
|
|
27
47
|
if (value === null || value === undefined)
|
package/dist/lib/json.js
CHANGED
|
@@ -5,49 +5,54 @@ const NAMED_CONTROL_CHAR_ESCAPES = {
|
|
|
5
5
|
'\b': '\\b',
|
|
6
6
|
'\f': '\\f',
|
|
7
7
|
};
|
|
8
|
+
const CONTROL_CHAR_ESCAPES = Array.from({ length: 0x20 }, (_, code) => NAMED_CONTROL_CHAR_ESCAPES[String.fromCharCode(code)] ?? `\\u${code.toString(16).padStart(4, '0')}`);
|
|
8
9
|
export function escapeRawControlCharsInJsonStrings(text) {
|
|
9
|
-
|
|
10
|
+
const parts = [];
|
|
11
|
+
let copied = 0;
|
|
10
12
|
let inString = false;
|
|
11
13
|
let escaped = false;
|
|
12
14
|
for (let i = 0; i < text.length; i += 1) {
|
|
13
|
-
const ch = text[i];
|
|
14
15
|
const code = text.charCodeAt(i);
|
|
15
16
|
if (!inString) {
|
|
16
|
-
if (
|
|
17
|
+
if (code === 0x22)
|
|
17
18
|
inString = true;
|
|
18
|
-
out += ch;
|
|
19
19
|
continue;
|
|
20
20
|
}
|
|
21
21
|
if (escaped) {
|
|
22
|
-
out += ch;
|
|
23
22
|
escaped = false;
|
|
24
23
|
continue;
|
|
25
24
|
}
|
|
26
|
-
if (
|
|
27
|
-
out += ch;
|
|
25
|
+
if (code === 0x5c) {
|
|
28
26
|
escaped = true;
|
|
29
27
|
continue;
|
|
30
28
|
}
|
|
31
|
-
if (
|
|
29
|
+
if (code === 0x22) {
|
|
32
30
|
inString = false;
|
|
33
|
-
out += ch;
|
|
34
31
|
continue;
|
|
35
32
|
}
|
|
36
33
|
if (code <= 0x1f) {
|
|
37
|
-
|
|
38
|
-
|
|
34
|
+
if (i > copied)
|
|
35
|
+
parts.push(text.slice(copied, i));
|
|
36
|
+
parts.push(CONTROL_CHAR_ESCAPES[code]);
|
|
37
|
+
copied = i + 1;
|
|
39
38
|
}
|
|
40
|
-
out += ch;
|
|
41
39
|
}
|
|
42
|
-
|
|
40
|
+
if (parts.length === 0)
|
|
41
|
+
return text;
|
|
42
|
+
if (copied < text.length)
|
|
43
|
+
parts.push(text.slice(copied));
|
|
44
|
+
return parts.join('');
|
|
43
45
|
}
|
|
44
46
|
export function parseServerJson(text) {
|
|
45
47
|
try {
|
|
46
48
|
return JSON.parse(text);
|
|
47
49
|
}
|
|
48
50
|
catch (firstErr) {
|
|
51
|
+
const repaired = escapeRawControlCharsInJsonStrings(text);
|
|
52
|
+
if (repaired === text)
|
|
53
|
+
throw firstErr;
|
|
49
54
|
try {
|
|
50
|
-
return JSON.parse(
|
|
55
|
+
return JSON.parse(repaired);
|
|
51
56
|
}
|
|
52
57
|
catch {
|
|
53
58
|
throw firstErr;
|
|
@@ -1,7 +1,9 @@
|
|
|
1
1
|
import chalk from 'chalk';
|
|
2
|
+
import { printable } from './format.js';
|
|
2
3
|
import { renderPinchArt } from './pinch.js';
|
|
3
4
|
const FRAME_MS = 600;
|
|
4
|
-
export function startPinchAnimation(
|
|
5
|
+
export function startPinchAnimation(initialStatus) {
|
|
6
|
+
let statusText = initialStatus;
|
|
5
7
|
const frames = [
|
|
6
8
|
renderPinchArt('working', { compact: true, frame: 0 }),
|
|
7
9
|
renderPinchArt('working', { compact: true, frame: 1 }),
|
|
@@ -30,6 +32,9 @@ export function startPinchAnimation(statusText) {
|
|
|
30
32
|
timer.unref();
|
|
31
33
|
let stopped = false;
|
|
32
34
|
return {
|
|
35
|
+
setStatus(text) {
|
|
36
|
+
statusText = printable(text);
|
|
37
|
+
},
|
|
33
38
|
stop() {
|
|
34
39
|
if (stopped)
|
|
35
40
|
return;
|
package/docs/agent-quickstart.md
CHANGED
|
@@ -103,30 +103,40 @@ succeed" — a failed first run is still a 200 response (the scrap was still
|
|
|
103
103
|
created), and the CLI exits `1` in that case even though `--json` always
|
|
104
104
|
prints the raw payload verbatim. The call legitimately runs many minutes
|
|
105
105
|
server-side (AI generation + a proxy-tier probe ladder + an iteration loop +
|
|
106
|
-
a real run)
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
106
|
+
a real run), so the CLI reads the wizard's NDJSON progress stream: human mode
|
|
107
|
+
prints one short line per stage on stderr, `--json` prints only the final
|
|
108
|
+
payload (same shape as always). It arms a 3600s safety ceiling for this
|
|
109
|
+
endpoint alone, above the 660s `run`/`data --fresh`/`trigger --wait` use and
|
|
110
|
+
well above the generic 30s default. Both use a transport of their own: Node's
|
|
111
|
+
global `fetch` cannot wait longer than ~300s no matter what ceiling is asked
|
|
112
|
+
for.
|
|
113
|
+
|
|
114
|
+
> **Not idempotent, and not a one-off run.** A transport failure — a
|
|
115
|
+
> client-side timeout, a progress stream cut or ended without a result, a
|
|
116
|
+
> connection dropped after the API already answered (exit `5`, a
|
|
117
|
+
> `NetworkError`), or a gateway `502`/`504` — does not mean the wizard failed server-side: scrap
|
|
114
118
|
> creation + the first run keep going after the CLI gives up waiting, so the
|
|
115
119
|
> scrap may already exist. ⚠ The stderr warning about this is HUMAN-mode
|
|
116
120
|
> only: under `--json` you get exit `5` and nothing else — no warning line,
|
|
117
|
-
> no extra envelope field — so treat ANY exit `5` from `create
|
|
118
|
-
> created". Check `trawl list --json` for a matching URL/title
|
|
121
|
+
> no extra envelope field — so treat ANY exit `5` from `create`, and an
|
|
122
|
+
> exit `1` with `status` 502/504, as "maybe created". Check `trawl list --json` for a matching URL/title
|
|
119
123
|
> **before** retrying — a blind retry creates a duplicate scrap and burns
|
|
120
124
|
> quota a second time for the same goal. Separately, the created scrap is
|
|
121
|
-
> scheduled to re-run every
|
|
122
|
-
> (`cron: "0 7 * *
|
|
125
|
+
> scheduled to re-run every week, Monday at 07:00 UTC, by default
|
|
126
|
+
> (`cron: "0 7 * * 1"`, unrelated to `--no-autofix`) — each recurring run
|
|
123
127
|
> consumes execute quota. Disable or change it once you've reviewed the
|
|
124
128
|
> scrap: `trawl scraps update <id> --no-cron` (or `--cron <expr>`). Finally,
|
|
125
|
-
> if `TRAWL_TIMEOUT` is set globally for a tighter budget than
|
|
129
|
+
> if `TRAWL_TIMEOUT` is set globally for a tighter budget than 3600s, it
|
|
126
130
|
> clamps `create`'s ceiling too (env always wins) — so a value picked for a
|
|
127
131
|
> quick call, e.g. 600s, aborts `create` mid-wizard. Unset it or raise it
|
|
128
132
|
> before calling `create`.
|
|
129
133
|
|
|
134
|
+
> **A failure the wizard reports itself** — the stream's terminal `error`
|
|
135
|
+
> event — arrives after the HTTP `200` was already sent, so it has no
|
|
136
|
+
> `status`: `create` exits `1` with
|
|
137
|
+
> `{"error":{"message":"…","kind":"refused","retryable":false}}` (no `status`
|
|
138
|
+
> field). Branch on `kind`, not on a status.
|
|
139
|
+
|
|
130
140
|
> **No breaking change:** every verb above is also still reachable under its
|
|
131
141
|
> pre-reorg path, `trawl scraps <verb>` (e.g. `trawl scraps list`) — kept as
|
|
132
142
|
> a hidden alias. Prefer the bare top-level form above; it's what
|