@trawlme/cli 3.12.3 → 3.12.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +5 -4
- package/dist/commands/create.js +4 -4
- package/dist/lib/api.d.ts +6 -1
- package/dist/lib/api.js +246 -6
- package/dist/lib/config.js +12 -0
- package/docs/agent-quickstart.md +17 -10
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -68,16 +68,17 @@ trawl spec [--json] Print a versioned, machine-readab
|
|
|
68
68
|
|
|
69
69
|
> **No breaking change:** every verb above is also still reachable under its pre-reorg path, `trawl scraps <verb>` (e.g. `trawl scraps list`, `trawl scraps run <id>`) — kept as a hidden alias so scripts written before the surface reorg keep working. `trawl --help` only shows the top-level form above; `trawl scraps --help` only shows the remaining scrap-management commands below.
|
|
70
70
|
|
|
71
|
-
`create`/`whoami`/`ping` are fully non-interactive — all three read auth only from `TRAWL_API_KEY`/`TRAWL_TOKEN`/the stored login token, never prompt (`whoami` is JWT-only — see [Authentication](#authentication) — so it still refuses honestly under a key instead of prompting). `trawl create` runs the AI wizard server-side (`POST /api/ai/wizard`): generate scrap code from `--prompt` via LLM, persist the scrap, trigger its FIRST run, and auto-fix on failure (default on — `--no-autofix` disables it, sending `autoFix:false`). On a successful first run (human mode) it prints a small **data sample** (item count + first-item fields + one truncated value) as proof of value — best-effort, silent if the sample can't be fetched — and points `Next step` at `trawl data <id>` (the data), with `trawl get <id>` as the secondary detail view. `--json` skips the sample fetch and prints the raw wizard payload verbatim. `success` is an honest outcome of that first run, not "did the HTTP call succeed" — a failed first run is still a 200 response (the scrap was still created; auto-fix, when enabled, retries in the background), and the CLI exits 1 in that case (both human and `--json` modes) even though `--json` always prints the raw payload verbatim. The call legitimately takes
|
|
71
|
+
`create`/`whoami`/`ping` are fully non-interactive — all three read auth only from `TRAWL_API_KEY`/`TRAWL_TOKEN`/the stored login token, never prompt (`whoami` is JWT-only — see [Authentication](#authentication) — so it still refuses honestly under a key instead of prompting). `trawl create` runs the AI wizard server-side (`POST /api/ai/wizard`): generate scrap code from `--prompt` via LLM, persist the scrap, trigger its FIRST run, and auto-fix on failure (default on — `--no-autofix` disables it, sending `autoFix:false`). On a successful first run (human mode) it prints a small **data sample** (item count + first-item fields + one truncated value) as proof of value — best-effort, silent if the sample can't be fetched — and points `Next step` at `trawl data <id>` (the data), with `trawl get <id>` as the secondary detail view. `--json` skips the sample fetch and prints the raw wizard payload verbatim. `success` is an honest outcome of that first run, not "did the HTTP call succeed" — a failed first run is still a 200 response (the scrap was still created; auto-fix, when enabled, retries in the background), and the CLI exits 1 in that case (both human and `--json` modes) even though `--json` always prints the raw payload verbatim. The call legitimately takes many minutes server-side (AI generation + a proxy-tier probe ladder + an iteration loop + a real run), so it arms a 1500s ceiling of its own — longer than the 300s `run`/`data --fresh`/`trigger --wait` use below, and over a transport of its own (Node's global `fetch` cannot wait past ~300s whatever ceiling is asked for). `trawl whoami`/`trawl ping` mirror the MCP `trawl_whoami`/`trawl_health_ping` tools as closely as the REST surface allows (`GET /api/users/me` / `GET /api/health`) — `ping`'s `--json` payload is admin-enriched (version/uptime/db) and just `{"status":"ok"}` for anyone else.
|
|
72
72
|
|
|
73
|
-
> **`create` is NOT idempotent, and every wizard-created scrap runs on a DAILY cron by default.** A client-side
|
|
73
|
+
> **`create` is NOT idempotent, and every wizard-created scrap runs on a DAILY cron by default.** A client-side transport failure (exit `5`, a `NetworkError`) — a timeout, or a connection dropped after the API already answered — does not mean the wizard failed server-side: scrap creation + the first run keep going after the CLI gives up waiting, so the scrap may already exist. In human mode the CLI prints a stderr warning whenever the request reached the API without returning a complete answer; under `--json` there is no warning and no extra envelope field — a machine caller sees exit `5` and must check `trawl list` itself. Run `trawl list` and look for a matching URL/title **before** retrying — a blind retry creates a DUPLICATE scrap and burns AI-generation quota a second time for the same goal. Separately, the scrap the wizard creates is scheduled to re-run every day at 07:00 UTC (`cron: "0 7 * * *"`, hardcoded server-side, unrelated to `--no-autofix`) — each of those recurring runs consumes execute quota like any other run. Review the generated scrap, then change or disable the schedule with `trawl scraps update <id> --cron <expr>` (or `--no-cron` to disable it). Because the call can legitimately run many minutes, also confirm `TRAWL_TIMEOUT` isn't set to something tighter than `create` needs — the env var always wins over `create`'s own 1500s default (see [Environment variables](#environment-variables)), so a value set for another purpose (e.g. a tight CI smoke-test budget) silently clamps `create` too; unset it or raise it before running `create`.
|
|
74
74
|
|
|
75
75
|
- `list` has a short alias, `ls` (matches `trawl --help`'s `list|ls`).
|
|
76
76
|
- `history` lists past runs (newest first); `run-info <hid>` shows details of a single run from that history.
|
|
77
77
|
- `data` returns the last persisted run payload (no execute quota); `--fresh` runs the scrap live instead (consumes execute quota); `--errors` shows the last run's error detail (`--json` on a never-run scrap returns `{"status":"no_runs"}`, exit 0, matching `scraps doctor --json`). `[]` on stdout means a genuine zero-item successful run — a scrap that has never run, whose last run failed, or whose payload aged out of retention returns a `--json` error envelope (exit 4/1/4 respectively) instead. Two more honest states: a run still **in flight** (`status: null` server-side) returns a `kind:"in_progress"` error envelope (exit 1, "retry shortly" — never suggests `--fresh`, which would just 429 against the run already holding the lock); a run whose item count **regressed** vs baseline (`statusDetail: "regression"`) still returns the real, non-empty items on stdout (exit 0) plus a stderr warning pointing at `scraps doctor <id>` — the data itself is genuine even though the run is flagged.
|
|
78
78
|
- `get` (and anything reading through it, like `data`'s default path) embeds only the newest 100 history rows on the returned scrap object — `run-info` and `scraps doctor` fetch a single run directly and are unaffected by that cap. `list`/`get` show a distinct amber `▼` "regression" badge, never the red `✗` a genuine failure gets (matches `scraps doctor`'s own badge).
|
|
79
79
|
- `list`/`get` also show a **health** badge next to the status badge. By default it reads the scrap's own `lastCronOutcome` field (always present, no extra cost): `⚠ cron paused` when the last scheduled tick was skipped for being unhealthy, `—` otherwise — a breadcrumb of the last tick, not a live read (never set on a no-cron scrap, can lag by one tick). `list --unhealthy` asks the API to filter to scraps whose *current* consecutive-failure streak is 3+ (the server's own threshold) and annotates each with the exact `consecutiveFailedRuns`/`unhealthySince` — shown instead of the badge whenever present, and never re-derived client-side from history. `--json` carries whichever fields the request produced: `lastCronOutcome` always, `consecutiveFailedRuns`/`unhealthySince` only when `--unhealthy` was passed. **Old-server safety:** an older server version that predates this filter silently ignores the unknown `--unhealthy` query param and returns everything, unfiltered — the CLI detects that (no returned item carries `consecutiveFailedRuns`) and prints a stderr warning instead of presenting the full list as "your unhealthy scraps".
|
|
80
|
-
- **Long-running calls (`
|
|
80
|
+
- **Long-running calls (`run`, `data --fresh`, `trigger --wait`):** these hit server-side paths that can legitimately take 30–250s+ (proxy tier escalation + AI-fix retries) — the CLI arms a 300s timeout for exactly these three call sites instead of the generic 30s default.
|
|
81
|
+
- **`create` is longer still:** the wizard adds a proxy-tier probe ladder and an iteration loop on top of generation + creation + a first run, so it arms **1500s** for that one call site (600s probe ladder + 480s iteration loop + 250s worst-case first run + margin). Raising the 300s above to cover it would make the three calls above wait 25 minutes on an unreachable API, so the two ceilings are deliberately separate. `create` is also the one call that does not go through Node's global `fetch`, which caps any wait for response headers at ~300s regardless of the ceiling asked for. `TRAWL_TIMEOUT` (see below) still overrides ALL requests, including both — set it if you need a tighter or looser ceiling, but note a global override sized for a quick call also clamps `create`.
|
|
81
82
|
- **`--watch` is poll-based, not a live stream:** the activities SSE endpoint has no backlog and, for the default async `trigger` (no `--wait`), runs in a separate cron-consumer pod whose events never reach the API pod holding the SSE connection — a naive "await the run, then open SSE" shows nothing. `run --watch` and `trigger --watch` instead poll `GET /api/scraps/:id` (terminal status) and the activities REST list until the run finishes, printing each new activity line as it appears. The watched run's outcome drives the exit code too, in BOTH human and `--json` mode: a genuinely failed terminal run, a poll timeout, or a persistently unreachable API all exit non-zero — a clean successful run is the only exit `0`. A run that never reaches a terminal status within 300s prints an honest timeout notice pointing at `scraps doctor <id>` (human mode) — see the `--json` shape below. `scraps watch <id>` (the standalone command, no trigger) is unchanged — it still opens the live SSE stream directly.
|
|
82
83
|
- `spec --json` prints the CLI's own command tree — `{specVersion, cliVersion, commands[], exitCodes, errorKinds, kindExitCodes, docsUrl?, llmsUrl?}` — DERIVED at runtime by walking the live commander tree (never a hand-maintained file, which would silently drift from reality). Each entry in `commands[]` carries its full path (e.g. `"scraps account session set"`), description, `hidden` (the legacy `scraps <verb>` aliases above), `leaf` (false for a pure namespace/group node like `scraps`/`skills`/`telemetry` — invoking one directly is a guaranteed-failing tool, not a real command), `aliases`, `arguments`, `options`, and an optional per-command `docs` deep link (e.g. every `scraps account *` command points at the account-sessions guide); `exitCodes`/`errorKinds`/`kindExitCodes` are read from the exact same source `classifyError` uses (see [Exit codes](#exit-codes)) — never a second copy. `kindExitCodes` is the inverse of `exitCodes`: `kind -> exitCode`, since exit code `1` alone is a shared bucket (`api`/`refused`/`unknown`/`in_progress`/`run_failed`/`upgrade_failed`) that `exitCodes`' flat label can't disambiguate. `docsUrl` is resolved server-first (`externalDocs.url` on the configured API base's own OpenAPI document, when declared), else derived from a KNOWN first-party `trawl.me` API host, else **omitted entirely** — never a guessed URL pointed at the wrong docs host for a self-hosted install. `llmsUrl` and every per-command `docs` link are narrower: they only ever come from that same known-host derivation, NEVER from a server-declared `docsUrl` — this CLI's own guide slugs have no reason to exist on a third party's own docs root, so a self-hosted server that declares `externalDocs.url` gets a correct top-level `docsUrl` but no `llmsUrl` and no per-command `docs` links. A `failureKind:"auth"` run's JSON payload (`doctor`/`data --errors`/`run-info`) carries the same `docs` field, by the same known-host-only rule. Without `--json`, `spec` prints one short human line (version + visible command count) pointing at `--json`.
|
|
83
84
|
- `run --json`/`trigger --json` bypass the spinner and print the raw launch/trigger payload on stdout; combined with `--watch`, every intermediate progress line stays suppressed (stdout stays pure JSON) and, once the watch reaches its outcome, exactly ONE final NDJSON line is emitted: `{"runId","status"}` (the honest terminal status — `success`/`error`/`empty`/`regression`/…), `{"runId","status":"timeout"}` on a poll timeout, or `{"runId","status":"poll_error","error"}` if the API stays unreachable for several consecutive polls — `process.exitCode` is non-zero for all three except a genuine success. `scraps watch --json` emits one raw JSON object per activity line (NDJSON) instead of the formatted `[time] message` text — there's no single final payload to wait for on a live stream.
|
|
@@ -237,7 +238,7 @@ Under `--json`, a failing command emits a single error envelope on stdout — `{
|
|
|
237
238
|
| `TRAWL_TELEMETRY` | Set to `0` to disable telemetry for the current session |
|
|
238
239
|
| `DO_NOT_TRACK` | Set to `1` to disable telemetry (cross-vendor convention, https://consoledonottrack.com) — same effect as `TRAWL_TELEMETRY=0` |
|
|
239
240
|
| `TRAWL_CONFIG_DIR` | Override where the config file (token, API URL, telemetry state) is stored — useful for hermetic CI runs or concurrent `trawl login`s that must not share one on-disk file |
|
|
240
|
-
| `TRAWL_TIMEOUT` | Override the per-request fetch timeout in milliseconds (default `30000`; `300000` for `
|
|
241
|
+
| `TRAWL_TIMEOUT` | Override the per-request fetch timeout in milliseconds (default `30000`; `300000` for `run`/`data --fresh`/`trigger --wait`, `1500000` for `create` — this env var always wins over those longer defaults too) |
|
|
241
242
|
| `TRAWL_SKILLS_SYNC` | Set to `0` to disable the startup skills auto-sync entirely |
|
|
242
243
|
|
|
243
244
|
Session override (no `trawl login` mutation, ideal for CI/QA against another env):
|
package/dist/commands/create.js
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
import { Command } from 'commander';
|
|
2
2
|
import chalk from 'chalk';
|
|
3
3
|
import { spin } from '../lib/spinner.js';
|
|
4
|
-
import { api,
|
|
4
|
+
import { api, WIZARD_TIMEOUT_MS, NetworkError } from '../lib/api.js';
|
|
5
5
|
import { json } from '../lib/format.js';
|
|
6
6
|
import { parseServerJson } from '../lib/json.js';
|
|
7
7
|
import { requireUrl, requireString } from '../lib/validate.js';
|
|
@@ -78,7 +78,7 @@ export const create = new Command('create')
|
|
|
78
78
|
goal,
|
|
79
79
|
...(opts.autofix === false && { autoFix: false }),
|
|
80
80
|
};
|
|
81
|
-
const call = () => api.
|
|
81
|
+
const call = () => api.postLongRunning('/api/ai/wizard', body, { timeoutMs: WIZARD_TIMEOUT_MS });
|
|
82
82
|
let data;
|
|
83
83
|
if (opts.json) {
|
|
84
84
|
data = await call();
|
|
@@ -95,8 +95,8 @@ export const create = new Command('create')
|
|
|
95
95
|
}
|
|
96
96
|
catch (err) {
|
|
97
97
|
anim?.stop();
|
|
98
|
-
if (err instanceof NetworkError &&
|
|
99
|
-
console.error(chalk.yellow('⚠ The request
|
|
98
|
+
if (err instanceof NetworkError && err.mayHaveCompleted) {
|
|
99
|
+
console.error(chalk.yellow('⚠ The request did not complete client-side, but the scrap may STILL have been created server-side — run `trawl list` before retrying (a retry creates a DUPLICATE scrap + burns quota).'));
|
|
100
100
|
}
|
|
101
101
|
throw err;
|
|
102
102
|
}
|
package/dist/lib/api.d.ts
CHANGED
|
@@ -4,7 +4,8 @@ export declare class ApiError extends Error {
|
|
|
4
4
|
constructor(status: number, message: string, next?: string[] | undefined);
|
|
5
5
|
}
|
|
6
6
|
export declare class NetworkError extends Error {
|
|
7
|
-
|
|
7
|
+
readonly mayHaveCompleted: boolean;
|
|
8
|
+
constructor(message: string, mayHaveCompleted?: boolean);
|
|
8
9
|
}
|
|
9
10
|
export declare class AuthError extends Error {
|
|
10
11
|
next?: string[] | undefined;
|
|
@@ -13,9 +14,12 @@ export declare class AuthError extends Error {
|
|
|
13
14
|
export declare function notLoggedInError(): AuthError;
|
|
14
15
|
export declare function apiKeyUnsupportedError(command: string): AuthError;
|
|
15
16
|
export declare const LONG_RUN_TIMEOUT_MS = 300000;
|
|
17
|
+
export declare const WIZARD_TIMEOUT_MS = 1500000;
|
|
16
18
|
export interface RequestOptions {
|
|
17
19
|
timeoutMs?: number;
|
|
18
20
|
}
|
|
21
|
+
export declare function effectivePort(u: URL): string;
|
|
22
|
+
export declare function sameEffectivePort(from: URL, next: URL): boolean;
|
|
19
23
|
export declare const api: {
|
|
20
24
|
get: <T>(path: string, opts?: RequestOptions) => Promise<T>;
|
|
21
25
|
publicGet: <T>(path: string, opts?: RequestOptions) => Promise<T>;
|
|
@@ -24,6 +28,7 @@ export declare const api: {
|
|
|
24
28
|
}) => Promise<T | null>;
|
|
25
29
|
getText: (path: string, opts?: RequestOptions) => Promise<string>;
|
|
26
30
|
post: <T>(path: string, body?: unknown, opts?: RequestOptions) => Promise<T>;
|
|
31
|
+
postLongRunning: <T>(path: string, body?: unknown, opts?: RequestOptions) => Promise<T>;
|
|
27
32
|
put: <T>(path: string, body?: unknown, opts?: RequestOptions) => Promise<T>;
|
|
28
33
|
delete: <T>(path: string, opts?: RequestOptions) => Promise<T>;
|
|
29
34
|
upload: <T>(path: string, formData: FormData, opts?: RequestOptions) => Promise<T>;
|
package/dist/lib/api.js
CHANGED
|
@@ -1,8 +1,9 @@
|
|
|
1
1
|
import { readFileSync } from 'node:fs';
|
|
2
2
|
import { fileURLToPath } from 'node:url';
|
|
3
3
|
import { dirname, resolve } from 'node:path';
|
|
4
|
-
import { request as httpRequest } from 'node:http';
|
|
4
|
+
import { request as httpRequest, } from 'node:http';
|
|
5
5
|
import { request as httpsRequest } from 'node:https';
|
|
6
|
+
import { brotliDecompressSync, gunzipSync, inflateRawSync, inflateSync } from 'node:zlib';
|
|
6
7
|
import { getApiUrl, getToken, getAuthMode, getLiveAuthEnvVar } from './config.js';
|
|
7
8
|
import { parseServerJson } from './json.js';
|
|
8
9
|
const __dirname = dirname(fileURLToPath(import.meta.url));
|
|
@@ -19,8 +20,10 @@ export class ApiError extends Error {
|
|
|
19
20
|
}
|
|
20
21
|
}
|
|
21
22
|
export class NetworkError extends Error {
|
|
22
|
-
|
|
23
|
+
mayHaveCompleted;
|
|
24
|
+
constructor(message, mayHaveCompleted = false) {
|
|
23
25
|
super(message);
|
|
26
|
+
this.mayHaveCompleted = mayHaveCompleted;
|
|
24
27
|
this.name = 'NetworkError';
|
|
25
28
|
}
|
|
26
29
|
}
|
|
@@ -62,6 +65,7 @@ function authHeaders(token) {
|
|
|
62
65
|
}
|
|
63
66
|
const DEFAULT_TIMEOUT_MS = 30_000;
|
|
64
67
|
export const LONG_RUN_TIMEOUT_MS = 300_000;
|
|
68
|
+
export const WIZARD_TIMEOUT_MS = 1_500_000;
|
|
65
69
|
function getTimeoutMs(overrideMs) {
|
|
66
70
|
const raw = process.env['TRAWL_TIMEOUT']?.trim();
|
|
67
71
|
if (raw) {
|
|
@@ -71,6 +75,16 @@ function getTimeoutMs(overrideMs) {
|
|
|
71
75
|
}
|
|
72
76
|
return overrideMs !== undefined && overrideMs > 0 ? overrideMs : DEFAULT_TIMEOUT_MS;
|
|
73
77
|
}
|
|
78
|
+
function timeoutNetworkError(url, effectiveTimeoutMs) {
|
|
79
|
+
return new NetworkError(`Request to ${url} timed out after ${effectiveTimeoutMs}ms (override with TRAWL_TIMEOUT env var, ms)`, true);
|
|
80
|
+
}
|
|
81
|
+
function connectionNetworkError(url, err, mayHaveCompleted = false) {
|
|
82
|
+
const e = err;
|
|
83
|
+
const cause = e?.cause;
|
|
84
|
+
const code = e?.code ?? cause?.code;
|
|
85
|
+
const detail = code ? ` (${code})` : cause?.message ? ` (${cause.message})` : '';
|
|
86
|
+
return new NetworkError(`Network error reaching ${url}${detail}: ${e?.message ?? String(err)}`, mayHaveCompleted);
|
|
87
|
+
}
|
|
74
88
|
async function safeFetch(url, options, effectiveTimeoutMs) {
|
|
75
89
|
try {
|
|
76
90
|
return await fetch(url, options);
|
|
@@ -78,11 +92,9 @@ async function safeFetch(url, options, effectiveTimeoutMs) {
|
|
|
78
92
|
catch (err) {
|
|
79
93
|
const e = err;
|
|
80
94
|
if (e?.name === 'TimeoutError' || e?.name === 'AbortError') {
|
|
81
|
-
throw
|
|
95
|
+
throw timeoutNetworkError(url, effectiveTimeoutMs);
|
|
82
96
|
}
|
|
83
|
-
|
|
84
|
-
const causeDetail = cause?.code ? ` (${cause.code})` : cause?.message ? ` (${cause.message})` : '';
|
|
85
|
-
throw new NetworkError(`Network error reaching ${url}${causeDetail}: ${e?.message ?? String(err)}`);
|
|
97
|
+
throw connectionNetworkError(url, err);
|
|
86
98
|
}
|
|
87
99
|
}
|
|
88
100
|
function extractErrorMessage(raw, statusText) {
|
|
@@ -331,6 +343,233 @@ function probeJson(path, opts) {
|
|
|
331
343
|
req.end();
|
|
332
344
|
});
|
|
333
345
|
}
|
|
346
|
+
function toHeaders(raw) {
|
|
347
|
+
const headers = new Headers();
|
|
348
|
+
for (const [key, value] of Object.entries(raw)) {
|
|
349
|
+
try {
|
|
350
|
+
if (Array.isArray(value))
|
|
351
|
+
for (const one of value)
|
|
352
|
+
headers.append(key, one);
|
|
353
|
+
else if (value !== undefined)
|
|
354
|
+
headers.append(key, value);
|
|
355
|
+
}
|
|
356
|
+
catch {
|
|
357
|
+
}
|
|
358
|
+
}
|
|
359
|
+
return headers;
|
|
360
|
+
}
|
|
361
|
+
const NULL_BODY_STATUS = new Set([204, 205, 304]);
|
|
362
|
+
function toResponse({ status, statusText, headers, text }) {
|
|
363
|
+
const body = NULL_BODY_STATUS.has(status) ? null : text;
|
|
364
|
+
const safeStatus = status >= 200 && status <= 599 ? status : 502;
|
|
365
|
+
try {
|
|
366
|
+
return new Response(body, { status: safeStatus, statusText, headers });
|
|
367
|
+
}
|
|
368
|
+
catch {
|
|
369
|
+
return new Response(body, { status: safeStatus, headers });
|
|
370
|
+
}
|
|
371
|
+
}
|
|
372
|
+
const MAX_REDIRECTS = 3;
|
|
373
|
+
const REDIRECT_KEEPS_METHOD = new Set([307, 308]);
|
|
374
|
+
const REDIRECT_DROPS_BODY = new Set([301, 302, 303]);
|
|
375
|
+
const MAX_DECODED_BYTES = 64 * 1024 * 1024;
|
|
376
|
+
const DECODE_LIMIT = { maxOutputLength: MAX_DECODED_BYTES };
|
|
377
|
+
function isOutputOverflow(err) {
|
|
378
|
+
return err?.code === 'ERR_BUFFER_TOO_LARGE';
|
|
379
|
+
}
|
|
380
|
+
function inflateEitherForm(body) {
|
|
381
|
+
try {
|
|
382
|
+
return inflateSync(body, DECODE_LIMIT);
|
|
383
|
+
}
|
|
384
|
+
catch (err) {
|
|
385
|
+
if (isOutputOverflow(err))
|
|
386
|
+
throw err;
|
|
387
|
+
return inflateRawSync(body, DECODE_LIMIT);
|
|
388
|
+
}
|
|
389
|
+
}
|
|
390
|
+
const CONTENT_DECODERS = Object.freeze({
|
|
391
|
+
gzip: (body) => gunzipSync(body, DECODE_LIMIT),
|
|
392
|
+
'x-gzip': (body) => gunzipSync(body, DECODE_LIMIT),
|
|
393
|
+
deflate: inflateEitherForm,
|
|
394
|
+
br: (body) => brotliDecompressSync(body, DECODE_LIMIT),
|
|
395
|
+
});
|
|
396
|
+
function decodeBody(body, contentEncoding) {
|
|
397
|
+
const encoding = (contentEncoding ?? '').trim().toLowerCase();
|
|
398
|
+
if (encoding === '' || encoding === 'identity')
|
|
399
|
+
return body;
|
|
400
|
+
const decoder = CONTENT_DECODERS[encoding];
|
|
401
|
+
if (!decoder)
|
|
402
|
+
return { reason: `unsupported Content-Encoding "${encoding}"` };
|
|
403
|
+
try {
|
|
404
|
+
return decoder(body);
|
|
405
|
+
}
|
|
406
|
+
catch (err) {
|
|
407
|
+
if (isOutputOverflow(err)) {
|
|
408
|
+
return {
|
|
409
|
+
reason: `Content-Encoding "${encoding}" expanded past the ${MAX_DECODED_BYTES / (1024 * 1024)} MB decode ceiling`,
|
|
410
|
+
};
|
|
411
|
+
}
|
|
412
|
+
return { reason: `Content-Encoding "${encoding}" failed to decode (${err.message})` };
|
|
413
|
+
}
|
|
414
|
+
}
|
|
415
|
+
function postRaw(url, payload, headers, timeoutMs) {
|
|
416
|
+
return new Promise((resolve, reject) => {
|
|
417
|
+
let target;
|
|
418
|
+
try {
|
|
419
|
+
target = new URL(url);
|
|
420
|
+
}
|
|
421
|
+
catch {
|
|
422
|
+
reject(connectionNetworkError(url, Object.assign(new Error('invalid API base URL'), { code: 'ERR_INVALID_URL' })));
|
|
423
|
+
return;
|
|
424
|
+
}
|
|
425
|
+
if (target.protocol !== 'http:' && target.protocol !== 'https:') {
|
|
426
|
+
reject(connectionNetworkError(url, Object.assign(new Error(`unsupported protocol "${target.protocol}"`), { code: 'ERR_INVALID_PROTOCOL' })));
|
|
427
|
+
return;
|
|
428
|
+
}
|
|
429
|
+
let settled = false;
|
|
430
|
+
let timedOut = false;
|
|
431
|
+
let mayHaveCompleted = false;
|
|
432
|
+
let timer;
|
|
433
|
+
let active;
|
|
434
|
+
const settleWith = (err) => {
|
|
435
|
+
if (settled)
|
|
436
|
+
return;
|
|
437
|
+
settled = true;
|
|
438
|
+
clearTimeout(timer);
|
|
439
|
+
active?.destroy();
|
|
440
|
+
reject(err);
|
|
441
|
+
};
|
|
442
|
+
const fail = (err) => {
|
|
443
|
+
settleWith(timedOut ? timeoutNetworkError(url, timeoutMs) : connectionNetworkError(url, err, mayHaveCompleted));
|
|
444
|
+
};
|
|
445
|
+
const done = (value) => {
|
|
446
|
+
if (settled)
|
|
447
|
+
return;
|
|
448
|
+
settled = true;
|
|
449
|
+
clearTimeout(timer);
|
|
450
|
+
resolve(value);
|
|
451
|
+
};
|
|
452
|
+
const onResponse = (res, from, hopsLeft, failThisHop) => {
|
|
453
|
+
const status = res.statusCode ?? 0;
|
|
454
|
+
const statusText = res.statusMessage ?? '';
|
|
455
|
+
res.on('error', failThisHop);
|
|
456
|
+
if (REDIRECT_KEEPS_METHOD.has(status)) {
|
|
457
|
+
res.resume();
|
|
458
|
+
const location = res.headers.location;
|
|
459
|
+
const next = location ? safeUrl(location, from) : null;
|
|
460
|
+
const refusal = !next
|
|
461
|
+
? 'it carries no usable Location header'
|
|
462
|
+
: next.protocol !== 'http:' && next.protocol !== 'https:'
|
|
463
|
+
? `its target uses an unsupported protocol ("${next.protocol}")`
|
|
464
|
+
: next.hostname !== from.hostname
|
|
465
|
+
? `its target is a different host (${next.hostname}) — this request carries your credential`
|
|
466
|
+
: !sameEffectivePort(from, next)
|
|
467
|
+
? `its target is a different port (${effectivePort(next)}) — this request carries your credential`
|
|
468
|
+
: from.protocol === 'https:' && next.protocol === 'http:'
|
|
469
|
+
? 'it downgrades https to http — this request carries your credential'
|
|
470
|
+
: hopsLeft <= 0
|
|
471
|
+
? `it exceeds ${MAX_REDIRECTS} redirects`
|
|
472
|
+
: null;
|
|
473
|
+
if (refusal || !next) {
|
|
474
|
+
settleWith(new ApiError(status, `${status} ${statusText}: refused to follow this redirect because ${refusal}. Point TRAWL_API_URL at the final URL.`));
|
|
475
|
+
return;
|
|
476
|
+
}
|
|
477
|
+
send(next, hopsLeft - 1);
|
|
478
|
+
return;
|
|
479
|
+
}
|
|
480
|
+
if (REDIRECT_DROPS_BODY.has(status)) {
|
|
481
|
+
res.resume();
|
|
482
|
+
settleWith(new ApiError(status, `${status} ${statusText}: this endpoint redirected the POST to ${res.headers.location ?? 'an unspecified location'}, and that status rewrites it to a bodyless GET — only 307/308 are followed. Point TRAWL_API_URL at the final URL.`));
|
|
483
|
+
return;
|
|
484
|
+
}
|
|
485
|
+
mayHaveCompleted = status >= 200 && status < 300;
|
|
486
|
+
const chunks = [];
|
|
487
|
+
res.on('data', (chunk) => chunks.push(chunk));
|
|
488
|
+
res.on('end', () => {
|
|
489
|
+
const decoded = decodeBody(Buffer.concat(chunks), res.headers['content-encoding']);
|
|
490
|
+
if (!Buffer.isBuffer(decoded)) {
|
|
491
|
+
settleWith(new NetworkError(`Unreadable response from ${url} (HTTP ${status}): ${decoded.reason} — the server ignored "Accept-Encoding: identity"`, mayHaveCompleted));
|
|
492
|
+
return;
|
|
493
|
+
}
|
|
494
|
+
done({ status, statusText, headers: toHeaders(res.headers), text: decoded.toString('utf8') });
|
|
495
|
+
});
|
|
496
|
+
res.on('close', () => {
|
|
497
|
+
if (!res.complete)
|
|
498
|
+
failThisHop(new Error('response aborted before it finished'));
|
|
499
|
+
});
|
|
500
|
+
};
|
|
501
|
+
const send = (to, hopsLeft) => {
|
|
502
|
+
mayHaveCompleted = false;
|
|
503
|
+
const transport = to.protocol === 'http:' ? httpRequest : httpsRequest;
|
|
504
|
+
let req;
|
|
505
|
+
const failThisHop = (err) => {
|
|
506
|
+
if (active !== req)
|
|
507
|
+
return;
|
|
508
|
+
fail(err);
|
|
509
|
+
};
|
|
510
|
+
try {
|
|
511
|
+
req = transport(to, { method: 'POST', headers: { 'Accept-Encoding': 'identity', ...headers } }, (res) => onResponse(res, to, hopsLeft, failThisHop));
|
|
512
|
+
}
|
|
513
|
+
catch (err) {
|
|
514
|
+
fail(err);
|
|
515
|
+
return;
|
|
516
|
+
}
|
|
517
|
+
active = req;
|
|
518
|
+
req.on('error', failThisHop);
|
|
519
|
+
if (payload !== undefined)
|
|
520
|
+
req.write(payload);
|
|
521
|
+
req.end();
|
|
522
|
+
};
|
|
523
|
+
timer = setTimeout(() => {
|
|
524
|
+
timedOut = true;
|
|
525
|
+
active?.destroy(new Error(`request to ${url} timed out after ${timeoutMs}ms`));
|
|
526
|
+
}, timeoutMs);
|
|
527
|
+
send(target, MAX_REDIRECTS);
|
|
528
|
+
});
|
|
529
|
+
}
|
|
530
|
+
export function effectivePort(u) {
|
|
531
|
+
return u.port !== '' ? u.port : u.protocol === 'https:' ? '443' : '80';
|
|
532
|
+
}
|
|
533
|
+
export function sameEffectivePort(from, next) {
|
|
534
|
+
if (from.port === '' && next.port === '')
|
|
535
|
+
return true;
|
|
536
|
+
return effectivePort(from) === effectivePort(next);
|
|
537
|
+
}
|
|
538
|
+
function safeUrl(raw, base) {
|
|
539
|
+
try {
|
|
540
|
+
return new URL(raw, base);
|
|
541
|
+
}
|
|
542
|
+
catch {
|
|
543
|
+
return null;
|
|
544
|
+
}
|
|
545
|
+
}
|
|
546
|
+
async function postLongRunning(path, body, reqOpts = {}) {
|
|
547
|
+
const token = getToken();
|
|
548
|
+
if (!token)
|
|
549
|
+
throw notLoggedInError();
|
|
550
|
+
const url = `${getApiUrl()}${path}`;
|
|
551
|
+
const timeoutMs = getTimeoutMs(reqOpts.timeoutMs);
|
|
552
|
+
const payload = body ? JSON.stringify(body) : undefined;
|
|
553
|
+
const raw = await postRaw(url, payload, {
|
|
554
|
+
'Content-Type': 'application/json',
|
|
555
|
+
'User-Agent': USER_AGENT,
|
|
556
|
+
...(payload !== undefined ? { 'Content-Length': String(Buffer.byteLength(payload)) } : {}),
|
|
557
|
+
...authHeaders(token),
|
|
558
|
+
}, timeoutMs);
|
|
559
|
+
await throwIfError(toResponse(raw), false, getAuthMode(token));
|
|
560
|
+
try {
|
|
561
|
+
if (!raw.text)
|
|
562
|
+
return {};
|
|
563
|
+
const parsed = parseServerJson(raw.text);
|
|
564
|
+
if (parsed !== null && typeof parsed === 'object' && 'data' in parsed) {
|
|
565
|
+
return parsed.data;
|
|
566
|
+
}
|
|
567
|
+
return parsed;
|
|
568
|
+
}
|
|
569
|
+
catch {
|
|
570
|
+
throw new NetworkError('Invalid JSON in server response', true);
|
|
571
|
+
}
|
|
572
|
+
}
|
|
334
573
|
async function getText(path, reqOpts = {}) {
|
|
335
574
|
const token = getToken();
|
|
336
575
|
if (!token)
|
|
@@ -356,6 +595,7 @@ export const api = {
|
|
|
356
595
|
method: 'POST',
|
|
357
596
|
body: body ? JSON.stringify(body) : undefined,
|
|
358
597
|
}, opts),
|
|
598
|
+
postLongRunning: (path, body, opts) => postLongRunning(path, body, opts),
|
|
359
599
|
put: (path, body, opts) => request(path, {
|
|
360
600
|
method: 'PUT',
|
|
361
601
|
body: body ? JSON.stringify(body) : undefined,
|
package/dist/lib/config.js
CHANGED
|
@@ -1,8 +1,10 @@
|
|
|
1
|
+
import { chmodSync, statSync } from 'node:fs';
|
|
1
2
|
import Conf from 'conf';
|
|
2
3
|
const configDir = process.env['TRAWL_CONFIG_DIR']?.trim();
|
|
3
4
|
const config = new Conf({
|
|
4
5
|
projectName: 'trawl-cli',
|
|
5
6
|
...(configDir ? { cwd: configDir } : {}),
|
|
7
|
+
configFileMode: 0o600,
|
|
6
8
|
defaults: {
|
|
7
9
|
apiUrl: 'https://api.trawl.me',
|
|
8
10
|
token: '',
|
|
@@ -12,6 +14,16 @@ const config = new Conf({
|
|
|
12
14
|
skillsNudgeShownAt: 0,
|
|
13
15
|
},
|
|
14
16
|
});
|
|
17
|
+
if (process.platform !== 'win32') {
|
|
18
|
+
try {
|
|
19
|
+
const stat = statSync(config.path, { throwIfNoEntry: false });
|
|
20
|
+
if (stat && (stat.mode & 0o077) !== 0) {
|
|
21
|
+
chmodSync(config.path, 0o600);
|
|
22
|
+
}
|
|
23
|
+
}
|
|
24
|
+
catch {
|
|
25
|
+
}
|
|
26
|
+
}
|
|
15
27
|
export function getApiUrl() {
|
|
16
28
|
const override = process.env['TRAWL_API_URL']?.trim();
|
|
17
29
|
return override ? override : config.get('apiUrl');
|
package/docs/agent-quickstart.md
CHANGED
|
@@ -101,23 +101,30 @@ first run, and auto-fix on failure (default on; `--no-autofix` disables it).
|
|
|
101
101
|
`success` is an honest outcome of that first run, not "did the HTTP call
|
|
102
102
|
succeed" — a failed first run is still a 200 response (the scrap was still
|
|
103
103
|
created), and the CLI exits `1` in that case even though `--json` always
|
|
104
|
-
prints the raw payload verbatim. The call legitimately runs
|
|
105
|
-
server-side (AI generation + a
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
104
|
+
prints the raw payload verbatim. The call legitimately runs many minutes
|
|
105
|
+
server-side (AI generation + a proxy-tier probe ladder + an iteration loop +
|
|
106
|
+
a real run) — the CLI arms a 1500s ceiling for this endpoint alone, above the
|
|
107
|
+
300s `run`/`data --fresh`/`trigger --wait` use and well above the generic 30s
|
|
108
|
+
default. That one call also uses a transport of its own: Node's global
|
|
109
|
+
`fetch` cannot wait longer than ~300s no matter what ceiling is asked for.
|
|
110
|
+
|
|
111
|
+
> **Not idempotent, and not a one-off run.** A client-side transport failure
|
|
112
|
+
> (exit `5`, a `NetworkError`) — a timeout, or a connection dropped after the
|
|
113
|
+
> API already answered — does not mean the wizard failed server-side: scrap
|
|
111
114
|
> creation + the first run keep going after the CLI gives up waiting, so the
|
|
112
|
-
> scrap may already exist.
|
|
115
|
+
> scrap may already exist. ⚠ The stderr warning about this is HUMAN-mode
|
|
116
|
+
> only: under `--json` you get exit `5` and nothing else — no warning line,
|
|
117
|
+
> no extra envelope field — so treat ANY exit `5` from `create` as "maybe
|
|
118
|
+
> created". Check `trawl list --json` for a matching URL/title
|
|
113
119
|
> **before** retrying — a blind retry creates a duplicate scrap and burns
|
|
114
120
|
> quota a second time for the same goal. Separately, the created scrap is
|
|
115
121
|
> scheduled to re-run every day at 07:00 UTC by default
|
|
116
122
|
> (`cron: "0 7 * * *"`, unrelated to `--no-autofix`) — each recurring run
|
|
117
123
|
> consumes execute quota. Disable or change it once you've reviewed the
|
|
118
124
|
> scrap: `trawl scraps update <id> --no-cron` (or `--cron <expr>`). Finally,
|
|
119
|
-
> if `TRAWL_TIMEOUT` is set globally for a tighter budget than
|
|
120
|
-
> clamps `create`'s ceiling too (env always wins) —
|
|
125
|
+
> if `TRAWL_TIMEOUT` is set globally for a tighter budget than 1500s, it
|
|
126
|
+
> clamps `create`'s ceiling too (env always wins) — so a value picked for a
|
|
127
|
+
> quick call, e.g. 600s, aborts `create` mid-wizard. Unset it or raise it
|
|
121
128
|
> before calling `create`.
|
|
122
129
|
|
|
123
130
|
> **No breaking change:** every verb above is also still reachable under its
|