@trawlme/cli 3.8.2 → 3.9.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -36,7 +36,7 @@ Four methods:
36
36
  3. **Token flag** — `trawl login --token <jwt>` (CI/CD, direct JWT)
37
37
  4. **API key** — `TRAWL_API_KEY=trawl_xxx trawl list` (scoped, revocable one-by-one — the recommended credential for an agent driving this CLI; see [docs/agent-quickstart.md](docs/agent-quickstart.md))
38
38
 
39
- A `TRAWL_API_KEY` is a `trawl_*`-prefixed credential (create/revoke one in the Trawl dashboard) sent as `Authorization: Bearer` instead of the session `Cookie: TOKEN=` a JWT uses — `trawl` picks the right one automatically based on the credential's own shape, never a flag. `TRAWL_API_KEY` wins over `TRAWL_TOKEN` when both happen to be set. It is scoped server-side — in practice to a set of scraps, since the dashboard's key-create form sets only the scrap allow-list; narrowing which *actions* a key may perform needs an explicit `scopes` array at creation over the API, and a key without one has every action granted. It works on most of the Core tier — `create`/`list`/`get`/`data`/`history`/`run-info`/`run`/`trigger`/`ping`, including the `--watch` **polling flag** on `run`/`trigger` — plus, outside Core, `scraps account status`/`scraps doctor`/`scraps autofix` (all three only ever read routes trawl_node opened to keys). It does **not** work on `whoami`, `scraps update`/`delete`, `scraps account set`/`delete`/`clear-session`/`session set`, `scraps banner`, `scraps snapshot` (the scrap lookup it starts from is dual-auth, but the `html-snapshot` route it downloads from isn't), or the standalone SSE **command** `scraps watch` (do not conflate the two: `--watch` is a flag on `run`/`trigger` and works under a key; `scraps watch` is a separate command and is JWT-only) — those stay JWT-only and fail with a `kind:"auth"` envelope (exit `3`) pointing at `trawl login` under a key. This list mirrors trawl_node's route wiring as of this writing, not a frozen guarantee — for anything not named here, trust the real `--json` envelope over this paragraph. Never send both a key and a JWT on the same request — the CLI only ever attaches one.
39
+ A `TRAWL_API_KEY` is a `trawl_*`-prefixed credential (create/revoke one in the Trawl dashboard) sent as `Authorization: Bearer` instead of the session `Cookie: TOKEN=` a JWT uses — `trawl` picks the right one automatically based on the credential's own shape, never a flag. `TRAWL_API_KEY` wins over `TRAWL_TOKEN` when both happen to be set. It is scoped server-side — in practice to a set of scraps, since the dashboard's key-create form sets only the scrap allow-list; narrowing which *actions* a key may perform needs an explicit `scopes` array at creation over the API, and a key without one has every action granted. It works on most of the Core tier — `create`/`list`/`get`/`data`/`history`/`run-info`/`run`/`trigger`/`ping`, including the `--watch` **polling flag** on `run`/`trigger` — plus, outside Core, `scraps account status`/`scraps doctor`/`scraps autofix` (all three only ever read routes trawl_node opened to keys). It does **not** work on `whoami`, `scraps update`/`delete`, `scraps account set`/`delete`/`clear-session`/`session set`, `scraps banner`, `scraps snapshot` (the scrap lookup it starts from is dual-auth, but the `html-snapshot` route it downloads from isn't), or the standalone SSE **command** `scraps watch` (do not conflate the two: `--watch` is a flag on `run`/`trigger` and works under a key; `scraps watch` is a separate command and is JWT-only) — those stay JWT-only and fail with a `kind:"auth"` envelope (exit `3`) pointing at `trawl login` under a key. This list mirrors trawl_node's route wiring as of this writing, not a frozen guarantee — for anything not named here, trust the real `--json` envelope over this paragraph. Never send both a key and a JWT on the same request — the CLI only ever attaches one. `list --unhealthy` (and the plain health badge on `list`/`get`) rides the same `GET /api/scraps`/`GET /api/scraps/:id` routes as the rest of the group — no separate auth path — so it works under a key exactly like plain `list`/`get`, and a scrap-scoped key still only ever sees its own allow-listed scraps through the filter.
40
40
 
41
41
  Custom API URL: `trawl login --url https://self-hosted.example.com`
42
42
 
@@ -51,7 +51,7 @@ All commands accept a global `--debug` flag to show full error stack traces on f
51
51
  ```
52
52
  trawl create <url> --prompt <goal> [--no-autofix] [--json] # url also accepted as --url <url> (#116)
53
53
  Create a persistent, self-healing scrap from a URL + a goal (AI-generated)
54
- trawl list|ls [--json] [--status <success|failure|never|running|regression>] [--limit <n>] [--page <n>]
54
+ trawl list|ls [--json] [--status <success|failure|never|running|regression>] [--unhealthy] [--limit <n>] [--page <n>]
55
55
  trawl get <id> [--json] Get scrap details
56
56
  trawl run <id> [--watch] [--json] Run a scrap (blocks until finished)
57
57
  trawl trigger <id> [--watch] [--wait] [--json] Launch a scrap as a background worker (returns immediately)
@@ -76,6 +76,7 @@ trawl spec [--json] Print a versioned, machine-readab
76
76
  - `history` lists past runs (newest first); `run-info <hid>` shows details of a single run from that history.
77
77
  - `data` returns the last persisted run payload (no execute quota); `--fresh` runs the scrap live instead (consumes execute quota); `--errors` shows the last run's error detail (`--json` on a never-run scrap returns `{"status":"no_runs"}`, exit 0, matching `scraps doctor --json`). `[]` on stdout means a genuine zero-item successful run — a scrap that has never run, whose last run failed, or whose payload aged out of retention returns a `--json` error envelope (exit 4/1/4 respectively) instead. Two more honest states: a run still **in flight** (`status: null` server-side) returns a `kind:"in_progress"` error envelope (exit 1, "retry shortly" — never suggests `--fresh`, which would just 429 against the run already holding the lock); a run whose item count **regressed** vs baseline (`statusDetail: "regression"`) still returns the real, non-empty items on stdout (exit 0) plus a stderr warning pointing at `scraps doctor <id>` — the data itself is genuine even though the run is flagged.
78
78
  - `get` (and anything reading through it, like `data`'s default path) embeds only the newest 100 history rows on the returned scrap object — `run-info` and `scraps doctor` fetch a single run directly and are unaffected by that cap. `list`/`get` show a distinct amber `▼` "regression" badge, never the red `✗` a genuine failure gets (matches `scraps doctor`'s own badge).
79
+ - `list`/`get` also show a **health** badge next to the status badge. By default it reads the scrap's own `lastCronOutcome` field (always present, no extra cost): `⚠ cron paused` when the last scheduled tick was skipped for being unhealthy, `—` otherwise — a breadcrumb of the last tick, not a live read (never set on a no-cron scrap, can lag by one tick). `list --unhealthy` asks the API to filter to scraps whose *current* consecutive-failure streak is 3+ (trawl_node's own threshold) and annotates each with the exact `consecutiveFailedRuns`/`unhealthySince` — shown instead of the badge whenever present, and never re-derived client-side from history. `--json` carries whichever fields the request produced: `lastCronOutcome` always, `consecutiveFailedRuns`/`unhealthySince` only when `--unhealthy` was passed. **Old-server safety:** a server predating trawl_node#1952 silently ignores the unknown `--unhealthy` query param and returns everything, unfiltered — the CLI detects that (no returned item carries `consecutiveFailedRuns`) and prints a stderr warning instead of presenting the full list as "your unhealthy scraps".
79
80
  - **Long-running calls (`create`, `run`, `data --fresh`, `trigger --wait`):** these hit server-side paths that can legitimately take 30–250s+ (AI generation + scrap creation + a first run for `create`; proxy tier escalation + AI-fix retries for the other three) — the CLI arms a 300s timeout for exactly these four call sites instead of the generic 30s default. `TRAWL_TIMEOUT` (see below) still overrides ALL requests, including these — set it if you need a tighter or looser ceiling than 300s, but note a global override that tight also clamps `create`.
80
81
  - **`--watch` is poll-based, not a live stream:** the activities SSE endpoint has no backlog and, for the default async `trigger` (no `--wait`), runs in a separate cron-consumer pod whose events never reach the API pod holding the SSE connection — a naive "await the run, then open SSE" shows nothing. `run --watch` and `trigger --watch` instead poll `GET /api/scraps/:id` (terminal status) and the activities REST list until the run finishes, printing each new activity line as it appears. The watched run's outcome drives the exit code too, in BOTH human and `--json` mode: a genuinely failed terminal run, a poll timeout, or a persistently unreachable API all exit non-zero — a clean successful run is the only exit `0`. A run that never reaches a terminal status within 300s prints an honest timeout notice pointing at `scraps doctor <id>` (human mode) — see the `--json` shape below. `scraps watch <id>` (the standalone command, no trigger) is unchanged — it still opens the live SSE stream directly.
81
82
  - `spec --json` prints the CLI's own command tree — `{specVersion, cliVersion, commands[], exitCodes, errorKinds, kindExitCodes}` — DERIVED at runtime by walking the live commander tree (never a hand-maintained file, which would silently drift from reality). Each entry in `commands[]` carries its full path (e.g. `"scraps account session set"`), description, `hidden` (the legacy `scraps <verb>` aliases above), `leaf` (false for a pure namespace/group node like `scraps`/`skills`/`telemetry` — invoking one directly is a guaranteed-failing tool, not a real command), `aliases`, `arguments`, and `options`; `exitCodes`/`errorKinds`/`kindExitCodes` are read from the exact same source `classifyError` uses (see [Exit codes](#exit-codes)) — never a second copy. `kindExitCodes` is the inverse of `exitCodes`: `kind -> exitCode`, since exit code `1` alone is a shared bucket (`api`/`refused`/`unknown`/`in_progress`/`run_failed`/`upgrade_failed`) that `exitCodes`' flat label can't disambiguate. Without `--json`, `spec` prints one short human line (version + visible command count) pointing at `--json`.
@@ -96,6 +97,8 @@ trawl scraps delete <id> [--force] [--json] Alias: rm
96
97
 
97
98
  - `-u/--url` is **the site the scrap targets**, not necessarily the exact URL fetched at run time — a scrap whose real target is computed inside its `-r/--request` script (a templated search query, an interpolated item id, e.g. `https://www.ebay.com/sch/i.html?_nkw=${query}`) should still supply a base/site URL. Every domain-level policy (proxy tier ceiling, provider routing, per-domain memoize) keys off this field, so a blank value silently disables all of it. `create` requires it — validated client-side (must parse as an `http`/`https` URL) so a missing or malformed value fails fast with a readable usage error (exit 2) instead of a raw server 400/422; `update` keeps it optional (omitting it is still a no-op, matching the API) but a value that **is** passed is validated the same way.
98
99
  - `--tier` forces a proxy tier; `--force-tier` raises the proxy-tier ceiling past the auto-cap (history-gated: may be refused or cost more). `create --json`/`update --json` print the full scrap object (including the `_tierOverride` outcome) on stdout; a refused tier override exits 1 with a standard `--json` error envelope (`kind:"refused"` — distinct from `"unknown"`, so a script can branch on "the server said no"). When a tier was requested but the server's response carries no `_tierOverride` at all (an older server that can't confirm what actually got applied), a stderr warning is printed either way, and under `--json` the emitted object also carries `"_tierUnconfirmed": true` — the machine-readable counterpart to that warning, since a `--json` caller has no reliable reason to read stderr.
100
+
101
+ `list --unhealthy` has the same shape of guard: against a server that predates the filter, the command warns on stderr and, under `--json`, wraps the payload as `{ "scraps": [...], "_healthFilterUnconfirmed": true }` instead of emitting a bare array — so a script can tell an unfiltered list from a genuinely-unhealthy one. A confirming server returns the bare array unchanged.
99
102
  - `scraps doctor` diagnoses the last run (error, failed selector, block status, page state, autofix outcome); `--autofix` includes the full autofix diff/dry-run/knowledge. A run that is still in flight (`status: null`, server-side) shows a `running` badge, never `failed`; a run whose item count regressed vs baseline (`statusDetail: "regression"`) shows its own amber `regression` badge, never `failed` either.
100
103
  - `scraps autofix` shows the last auto-fix attempt on its own (decision, diff, dry-run, knowledge). `--json` on a scrap that has **never run** returns `{"status":"no_runs"}` (exit 0) — distinct from `null`, which means a run exists but had no auto-fix attempt.
101
104
  - `scraps snapshot --error` fetches the error-path snapshot instead of the normal one; `-o <file>` writes to a file instead of stdout. On a scrap that has **never run**: `--json` returns `{"status":"no_runs"}` (exit 0), matching `doctor`/`autofix`; `-o <file>` (without `--json`) exits 4 (`not_found`) instead of silently exiting 0 with nothing written — a script checking the exit code alone must be able to tell "no file was produced" from success. `--json` takes priority when both are passed.
@@ -25,19 +25,31 @@ export interface Run {
25
25
  selectors?: Record<string, number>;
26
26
  } | null;
27
27
  blocked?: boolean;
28
+ block?: {
29
+ kind?: string | null;
30
+ } | null;
31
+ /** @deprecated trawl_node#1950 renamed this to `block.kind`. Kept as a read
32
+ * fallback (see detectWallVendor) for the window where this CLI is
33
+ * published ahead of the trawl_node prod tag — a prod backend served from
34
+ * the pre-#1950 tag still returns this flat field, not `block.kind`. Drop
35
+ * once prod is confirmed on a tag containing #1950. */
28
36
  blockType?: string | null;
29
37
  proxyTier?: string | null;
30
- regressionDetected?: boolean;
31
38
  baselineLength?: number | null;
32
39
  fixVersionId?: string | null;
33
40
  createdAt?: string;
34
41
  time?: number | null;
35
42
  triggeredBy?: string | null;
43
+ /** trawl_node#1975 — freeform failure classification (`'auth'` = login-wall
44
+ * empty run, cookies are the fix). Not a TS union — trawl_node's own set
45
+ * is additive/open, so this CLI must stay read-safe against a future
46
+ * value it doesn't know about yet, same posture as `block.kind` above. */
47
+ failureKind?: string | null;
36
48
  }
37
49
  /**
38
- * Resolve a known anti-bot vendor (or auth) name from a worker `blockType`
39
- * string. Returns null when the run succeeded, when there is no `blockType`
40
- * signal at all, or when `blockType` names something other than a known
50
+ * Resolve a known anti-bot vendor name from a worker `block.kind`
51
+ * string. Returns null when the run succeeded, when there is no `block.kind`
52
+ * signal at all, or when `block.kind` names something other than a known
41
53
  * vendor (e.g. `proxy-domain-gate`, `rate_limited_per_host`) — those stay on
42
54
  * the genuine-error path since we can't honestly attribute them to a specific
43
55
  * "no reliable bypass" wall.
@@ -45,12 +57,18 @@ export interface Run {
45
57
  * `blocked` is deliberately NOT consulted here (#177): it is a mid-run,
46
58
  * attempt-level signal the worker only stamps on one envelope shape, so it
47
59
  * reads `false` on the great majority of genuinely walled runs (the early-block
48
- * throw path sets `blockType` but never `blocked` — 63 of 64 walled runs
60
+ * throw path sets `block.kind` but never `blocked` — 63 of 64 walled runs
49
61
  * measured on prod). `status === true` wins instead: a run that ultimately
50
- * returned data was not walled, even if an earlier tier's `blockType` stamp
62
+ * returned data was not walled, even if an earlier tier's `block.kind` stamp
51
63
  * survived on the row.
64
+ *
65
+ * trawl_node#1950 renamed the flat `blockType` field to nested `block.kind`.
66
+ * Read `block?.kind` first, falling back to the deprecated flat `blockType` —
67
+ * this CLI can be published (and talk to prod) before the trawl_node prod tag
68
+ * containing #1950 is cut (see the `Run.blockType` doc comment), so a prod
69
+ * response can still be the pre-#1950 flat shape for a while.
52
70
  */
53
- export declare function detectWallVendor(run: Pick<Run, 'status' | 'statusDetail' | 'blockType'>): string | null;
71
+ export declare function detectWallVendor(run: Pick<Run, 'status' | 'statusDetail' | 'block' | 'blockType'>): string | null;
54
72
  /**
55
73
  * Autofix activity metadata — from the persisted ai_fix_end activity.
56
74
  * aiUsage (cost) is stripped server-side; all diagnostics are kept.
@@ -2,23 +2,25 @@ import { api } from '../lib/api.js';
2
2
  import chalk from 'chalk';
3
3
  import { formatDate } from '../lib/format.js';
4
4
  /**
5
- * Known anti-bot vendors (+ auth) that the worker's `blockType` field may name.
6
- * Ordered by first-match; `blockType` is a freeform worker string, not an enum
7
- * (see the `Run.blockType` doc comment), so this is a best-effort substring
8
- * match against the real field — never an invented/mocked value.
5
+ * Known anti-bot vendors that the worker's `block.kind` field may name
6
+ * (ex-flat `blockType`, trawl_node#1950). Ordered by first-match;
7
+ * `block.kind` is a freeform worker string, not an enum (see the `Run.block`
8
+ * doc comment), so this is a best-effort substring match against the real
9
+ * field — never an invented/mocked value.
9
10
  *
10
11
  * `datadome`/`perimeterx`/`akamai`/`cloudflare` are literal substrings the
11
- * current worker emits (detect.js). `kasada` and `auth` are forward-compatible:
12
- * the issue (#62 / wall-registry WS-7) names them as genuine walls, but the
13
- * current worker vocabulary does not yet emit them — these branches light up
14
- * automatically if/when a future worker classifier does, without a CLI change.
12
+ * current worker emits (detect.js). `kasada` is forward-compatible: the issue
13
+ * (#62 / wall-registry WS-7) names it as a genuine wall, but the current
14
+ * worker vocabulary does not yet emit it — this branch lights up
15
+ * automatically if/when a future worker classifier does, without a CLI
16
+ * change.
15
17
  *
16
- * The `auth` pattern is delimiter-anchored (`^`/`$`/`-_:`) rather than a bare
17
- * substring so it matches worker-style tokens (`auth`, `auth-wall`, `auth_wall`,
18
- * `login:auth`) without false-matching `oauth`/`authorization`/`author`. It uses
19
- * `[-_:]` (not `\b`) because `_` is a JS `\w` char, so `\bauth\b` would miss the
20
- * underscore-delimited `auth_wall` form the worker's `rate_limited_per_host`-style
21
- * naming favors.
18
+ * trawl_cli#182 — an `auth` entry used to live here, rendering a login wall
19
+ * as "walled: auth — no reliable bypass". Removed: unlike a real anti-bot
20
+ * vendor, a login wall's own reliable bypass is the user's session cookies,
21
+ * so that verdict was wrong, not just early. `failureKind==='auth'`
22
+ * (trawl_node#1975) is the honest signal for a login wall — see the
23
+ * dedicated branch in `formatDoctor` below.
22
24
  */
23
25
  const WALL_VENDOR_PATTERNS = [
24
26
  [/datadome/i, 'DataDome'],
@@ -26,12 +28,11 @@ const WALL_VENDOR_PATTERNS = [
26
28
  [/perimeterx/i, 'PerimeterX'],
27
29
  [/akamai/i, 'Akamai'],
28
30
  [/cloudflare/i, 'Cloudflare'],
29
- [/(?:^|[-_:])auth(?:$|[-_:])/i, 'auth'],
30
31
  ];
31
32
  /**
32
- * Resolve a known anti-bot vendor (or auth) name from a worker `blockType`
33
- * string. Returns null when the run succeeded, when there is no `blockType`
34
- * signal at all, or when `blockType` names something other than a known
33
+ * Resolve a known anti-bot vendor name from a worker `block.kind`
34
+ * string. Returns null when the run succeeded, when there is no `block.kind`
35
+ * signal at all, or when `block.kind` names something other than a known
35
36
  * vendor (e.g. `proxy-domain-gate`, `rate_limited_per_host`) — those stay on
36
37
  * the genuine-error path since we can't honestly attribute them to a specific
37
38
  * "no reliable bypass" wall.
@@ -39,24 +40,31 @@ const WALL_VENDOR_PATTERNS = [
39
40
  * `blocked` is deliberately NOT consulted here (#177): it is a mid-run,
40
41
  * attempt-level signal the worker only stamps on one envelope shape, so it
41
42
  * reads `false` on the great majority of genuinely walled runs (the early-block
42
- * throw path sets `blockType` but never `blocked` — 63 of 64 walled runs
43
+ * throw path sets `block.kind` but never `blocked` — 63 of 64 walled runs
43
44
  * measured on prod). `status === true` wins instead: a run that ultimately
44
- * returned data was not walled, even if an earlier tier's `blockType` stamp
45
+ * returned data was not walled, even if an earlier tier's `block.kind` stamp
45
46
  * survived on the row.
47
+ *
48
+ * trawl_node#1950 renamed the flat `blockType` field to nested `block.kind`.
49
+ * Read `block?.kind` first, falling back to the deprecated flat `blockType` —
50
+ * this CLI can be published (and talk to prod) before the trawl_node prod tag
51
+ * containing #1950 is cut (see the `Run.blockType` doc comment), so a prod
52
+ * response can still be the pre-#1950 flat shape for a while.
46
53
  */
47
54
  export function detectWallVendor(run) {
48
55
  // A run that returned data was not walled. `status === true` is not the whole
49
56
  // predicate: trawl_node flips a degraded-but-non-empty run to
50
57
  // `status:false, statusDetail:'regression'` (scraps.service.js, #1112) via a
51
- // patch that never clears `blockType` — so keying on `status` alone would
58
+ // patch that never clears `block.kind` — so keying on `status` alone would
52
59
  // print "no reliable bypass" next to "Regression: length N vs baseline M",
53
60
  // which is the same contradiction this guard exists to remove.
54
61
  if (run.status === true || run.statusDetail === 'regression')
55
62
  return null;
56
- if (!run.blockType)
63
+ const kind = run.block?.kind ?? run.blockType ?? null;
64
+ if (!kind)
57
65
  return null; // no signal at all
58
66
  for (const [pattern, label] of WALL_VENDOR_PATTERNS) {
59
- if (pattern.test(run.blockType))
67
+ if (pattern.test(kind))
60
68
  return label;
61
69
  }
62
70
  return null;
@@ -70,8 +78,8 @@ const TIER_LABELS = {
70
78
  };
71
79
  const RUN_ALLOWLIST = [
72
80
  '_id', 'status', 'statusDetail', 'length', 'errorMessage', 'errorSnapshot',
73
- 'emptyContext', 'blocked', 'blockType', 'proxyTier', 'regressionDetected', 'baselineLength',
74
- 'fixVersionId', 'createdAt', 'time', 'triggeredBy',
81
+ 'emptyContext', 'blocked', 'block', 'blockType', 'proxyTier', 'baselineLength',
82
+ 'fixVersionId', 'createdAt', 'time', 'triggeredBy', 'failureKind',
75
83
  ];
76
84
  const FIX_ALLOWLIST = [
77
85
  'outcome', 'classification', 'reason', 'fixDiff', 'dryRunResults',
@@ -124,14 +132,81 @@ export function formatDoctor(scrapTitle, run, fix = null, scrapId) {
124
132
  lines.push(`${chalk.bold(scrapTitle)} ${badge}${run.statusDetail ? ` (${run.statusDetail})` : ''}`);
125
133
  lines.push(chalk.dim(` Run ID: ${run._id}`));
126
134
  // Error message — an honest accept-wall string for known-walled scraps
127
- // (DataDome/Kasada/PerimeterX/Akamai/auth terminal verdict from the worker),
135
+ // (DataDome/Kasada/PerimeterX/Akamai terminal verdict from the worker),
128
136
  // otherwise the real error (genuine transient failure).
137
+ //
138
+ // trawl_cli#182 — the wall-vendor verdict (this CLI's own
139
+ // WALL_VENDOR_PATTERNS matched against block.kind/blockType) and the
140
+ // login-wall hint (node's failureKind==='auth', trawl_node#1975) are two
141
+ // INDEPENDENT server-derived signals with no shared contract: node's block
142
+ // classifier (modules/historys/helpers/failureKind.js BLOCK_TYPE_PATTERN)
143
+ // and this CLI's vendor table live in different repos and can legitimately
144
+ // disagree. Concretely, `kasada` matches this CLI's table but is absent
145
+ // from node's block pattern, and the deprecated flat `blockType` fallback
146
+ // is a field node's classifier no longer reads at all — so a run can come
147
+ // back `failureKind:'auth'` (the server says login wall) while also
148
+ // matching a vendor here (this CLI says Kasada/DataDome/etc), both true
149
+ // signals about the same run, neither one wrong.
150
+ //
151
+ // trawl_cli#169 precedent: never assert client-side what only the server
152
+ // knows. Silently picking a winner between two disagreeing signals is
153
+ // exactly that — printing "no reliable bypass" alone asserts the run is
154
+ // NOT an auth wall (may stop someone from trying a fixable cookie
155
+ // problem); printing the cookie hint alone asserts the run is NOT a
156
+ // vendor wall (may send someone chasing cookies against a wall no cookie
157
+ // fixes). This CLI cannot tell which is true without either importing
158
+ // node's BLOCK_TYPE_PATTERN (a second source of truth for the same list —
159
+ // how this defect got here) or re-detecting the login URL itself (the
160
+ // client-side re-derivation this fix deliberately rejects). So when both
161
+ // fire, render ONE hedged line naming both readings instead of picking:
162
+ // never emit the bare "no reliable bypass" verdict in that case — it is
163
+ // exactly the false precision this branch exists to avoid.
164
+ //
165
+ // `conflictingSignals` is deliberately NOT gated on `scrapId` the way the
166
+ // plain hint below is: the false "no reliable bypass" verdict is wrong
167
+ // regardless of whether we also have an id to build an actionable command
168
+ // for, so a scrapId-less conflicting run must still avoid it. Only the
169
+ // action line degrades (to a generic Settings pointer) when there's no id.
170
+ //
171
+ // Single-signal cases are untouched: a plain vendor match (no
172
+ // failureKind:'auth') still gets "no reliable bypass"; a plain
173
+ // failureKind:'auth' (no vendor match) still gets the plain hint block
174
+ // below. `errMsg` stays suppressed whenever any hint form (conflicting or
175
+ // plain) is about to render — node's own `errorMessage` for an auth run
176
+ // (historys.service.js ~line 774) IS the hint copy verbatim, literal
177
+ // `<id>` placeholder and all, so printing it here would just be a second
178
+ // rendering of the same information. If neither hint form renders (no
179
+ // vendor conflict and no scrapId for the plain hint), fall back to node's
180
+ // raw message rather than dropping the error entirely.
129
181
  const wallVendor = detectWallVendor(run);
130
182
  const errMsg = run.errorMessage ?? run.errorSnapshot?.errorMessage;
131
- if (wallVendor) {
183
+ // trawl_cli#182 — same `status`/`statusDetail` guard as `detectWallVendor`
184
+ // above, and for the same reason (see its doc comment): `failureKind` is a
185
+ // terminal classification trawl_node stamps once
186
+ // (modules/historys/services/historys.service.js ~line 837), but
187
+ // `patchForRegression` (modules/scraps/services/scraps.service.js ~line
188
+ // 1972) can flip `status`/`statusDetail` to success/regression LATER,
189
+ // without ever clearing `failureKind`. Left unguarded, a run that
190
+ // ultimately succeeded or degraded could still carry a stale
191
+ // `failureKind:'auth'` and render the login-wall hint (or the
192
+ // conflicting-signals hedge) right next to a green "success" badge or a
193
+ // "Regression: length N vs baseline M" line — the exact contradiction
194
+ // `detectWallVendor`'s guard already exists to prevent for `block.kind`.
195
+ // One rule, not two coincidences: both checks guard the same two fields
196
+ // against the same after-the-fact patch.
197
+ const authWall = run.failureKind === 'auth' && run.status !== true && run.statusDetail !== 'regression';
198
+ const authHintWillRender = authWall && Boolean(scrapId);
199
+ const conflictingSignals = Boolean(wallVendor) && authWall;
200
+ if (conflictingSignals) {
201
+ lines.push(chalk.dim(' Error: ') + chalk.yellow(`conflicting signals — the server reported ${wallVendor} on this run, but also classified it as a login wall; can't tell which is true here`));
202
+ lines.push(chalk.dim(scrapId
203
+ ? ` → cookies might help, but this may still be a genuine ${wallVendor} block: trawl scraps account session set ${scrapId} -c cookies.json`
204
+ : ` → cookies might help, but this may still be a genuine ${wallVendor} block (see app Settings → Account)`));
205
+ }
206
+ else if (wallVendor) {
132
207
  lines.push(chalk.dim(' Error: ') + chalk.red(`walled: ${wallVendor} — no reliable bypass`));
133
208
  }
134
- else if (errMsg) {
209
+ else if (errMsg && !authHintWillRender) {
135
210
  lines.push(chalk.dim(' Error: ') + chalk.red(errMsg));
136
211
  }
137
212
  // Failed selector
@@ -149,8 +224,32 @@ export function formatDoctor(scrapTitle, run, fix = null, scrapId) {
149
224
  const page = run.emptyContext.page;
150
225
  lines.push(chalk.dim(' Empty context: ') + `url=${page.url ?? '?'} anchors=${page.totalAnchors ?? '?'}`);
151
226
  }
152
- // Regression detail
153
- if (run.regressionDetected) {
227
+ // Login-wall hint (trawl_node#1975 `failureKind==='auth'`) — trust node's
228
+ // verdict, never re-detect the login URL client-side. Hypothesis-framed
229
+ // ("likely fix"), not a promise: measured cases (Reddit/X/Instagram) had
230
+ // valid cookies and stayed walled. Gated on `scrapId` alone (via
231
+ // `authHintWillRender` above, shared with the generic error line so the
232
+ // two can't drift apart), NOT the `scrapId ?? run._id` fallback used
233
+ // elsewhere here — `run._id` is a history id, and the session-set command
234
+ // on one silently targets the wrong document.
235
+ //
236
+ // trawl_cli#182 — excludes `conflictingSignals`: when a wall-vendor match
237
+ // also fired, the combined hedged line above already covers the cookie
238
+ // suggestion. Rendering this plain, unhedged block too would stack a
239
+ // second verdict under the first (and directly contradict it, since this
240
+ // block asserts the run IS a login wall with no caveat) — exactly the
241
+ // two-verdicts-for-one-run defect this fix removes. Exactly one verdict
242
+ // block renders per run.
243
+ if (authHintWillRender && !conflictingSignals) {
244
+ lines.push('');
245
+ lines.push(chalk.yellow(' Login wall (hypothesis): ') + 'this looks like a login redirect — your own session cookies are the likely fix.');
246
+ lines.push(chalk.dim(` → trawl scraps account session set ${scrapId} -c cookies.json (or app Settings → Account)`));
247
+ }
248
+ // Regression detail. trawl_node#1950 dropped the redundant `regressionDetected`
249
+ // boolean — it was `true` on every row iff `statusDetail === 'regression'`,
250
+ // so read that directly (works against a pre- or post-#1950 backend alike,
251
+ // no CLI/backend sequencing gap: `statusDetail` isn't part of this rename).
252
+ if (run.statusDetail === 'regression') {
154
253
  lines.push(chalk.dim(' Regression: ') + `length ${run.length ?? '?'} vs baseline ${run.baselineLength ?? '?'}`);
155
254
  }
156
255
  // Autofix summary (when fix exists)
@@ -136,6 +136,67 @@ function statusIcon(status) {
136
136
  return chalk.hex('#FFA500')('▼');
137
137
  return chalk.dim('—');
138
138
  }
139
+ /**
140
+ * #179 — health badge, next to the existing status icon. Two distinct
141
+ * sources, deliberately never blended into one derived verdict:
142
+ * 1. `consecutiveFailedRuns`/`unhealthySince` — present ONLY when `--unhealthy`
143
+ * asked the API to attach them (see attachListCommand). This is the
144
+ * live, exact server verdict — shown whenever available, since it is
145
+ * strictly more informative than the breadcrumb below.
146
+ * 2. `lastCronOutcome === 'skipped_unhealthy'` — always present (a normal
147
+ * scrap field, no extra query cost), but a STALE breadcrumb of the last
148
+ * scheduled tick, not a live read: never set on a no-cron scrap, lags
149
+ * by up to one tick, and flips back to 'ok'/'error' once the recovery
150
+ * probe fires. Labeled as what it is ("cron paused"), never upgraded to
151
+ * "unhealthy" — that word is reserved for the live verdict above.
152
+ * Never re-derives either signal from `history[]` — both come straight off
153
+ * the server (#1952 in trawl_node).
154
+ */
155
+ function healthBadge(scrap) {
156
+ if (typeof scrap.consecutiveFailedRuns === 'number' && scrap.consecutiveFailedRuns > 0) {
157
+ const n = scrap.consecutiveFailedRuns;
158
+ return chalk.yellow(`⚠ failing ${n} ${n === 1 ? 'run' : 'runs'} (since ${formatDate(scrap.unhealthySince)})`);
159
+ }
160
+ if (scrap.lastCronOutcome === 'skipped_unhealthy') {
161
+ return chalk.yellow('⚠ cron paused');
162
+ }
163
+ return chalk.dim('—');
164
+ }
165
+ /**
166
+ * #179 — old-server fallback for `list --unhealthy`, same shape as
167
+ * `warnIfUnconfirmedTier` above (#86 finding 4b): an older trawl_node that
168
+ * predates #1952 ignores the unknown `minConsecutiveFailures` query param
169
+ * entirely and returns the FULL unfiltered list — silently presenting every
170
+ * scrap as "unhealthy" is exactly the kind of unconfirmed-outcome lie that
171
+ * precedent exists to prevent. A confirming server attaches
172
+ * `consecutiveFailedRuns` to EVERY item once the param is set (the
173
+ * controller writes `s.consecutiveFailedRuns || 0` unconditionally in that
174
+ * branch — see trawl_node scraps.controller.js), so its total absence across
175
+ * a non-empty response is the honest tell. Stderr only (safe under --json —
176
+ * stdout purity is untouched), never blocks the (unreliable) results below.
177
+ */
178
+ function healthFilterUnconfirmed(data, unhealthyWasRequested) {
179
+ if (!unhealthyWasRequested || data.length === 0)
180
+ return false;
181
+ return !data.some((s) => s.consecutiveFailedRuns !== undefined);
182
+ }
183
+ function warnIfUnconfirmedHealthFilter(data, unhealthyWasRequested) {
184
+ if (!healthFilterUnconfirmed(data, unhealthyWasRequested))
185
+ return;
186
+ console.error(chalk.yellow('⚠ Server did not confirm the --unhealthy filter (older server, predates trawl_node#1952) — results below are UNFILTERED, not just the unhealthy scraps.'));
187
+ }
188
+ /**
189
+ * #179 — the machine-readable counterpart to the stderr warning above, in the
190
+ * exact shape `withTierUnconfirmed` already uses for the tier case. A --json
191
+ * caller has no reliable reason to read stderr, so without this a script sees
192
+ * a full unfiltered list and cannot tell it is not the unhealthy set.
193
+ * Emitted ONLY when the filter went unconfirmed — never fabricated otherwise.
194
+ */
195
+ function withHealthFilterUnconfirmed(rows, data, unhealthyWasRequested) {
196
+ if (!healthFilterUnconfirmed(data, unhealthyWasRequested))
197
+ return rows;
198
+ return { scraps: rows, _healthFilterUnconfirmed: true };
199
+ }
139
200
  export const scraps = new Command('scraps').description('Manage scraps');
140
201
  // shared SSE streaming helper. `asJson` (#107) emits one raw JSON object per
141
202
  // line (NDJSON) on stdout instead of the human-formatted timestamped text —
@@ -426,6 +487,15 @@ export async function pollRunProgress(id, before, opts = {}) {
426
487
  // keeps `--status`'s allowed values and its error message in one place,
427
488
  // mirrored from lastStatus()'s own return type so the two can never drift.
428
489
  const VALID_LAST_STATUSES = ['success', 'failure', 'never', 'running', 'regression'];
490
+ // #179 — mirrors `UNHEALTHY_STREAK_LENGTH` in trawl_node
491
+ // modules/scraps/helpers/scrapCronHealth.js. `--unhealthy` is a boolean
492
+ // convenience flag for that ONE product-defined threshold (also what the
493
+ // server's own cron-skip gate and email use), not a general numeric filter —
494
+ // the API's `minConsecutiveFailures` param accepts any threshold, but this
495
+ // CLI only ever asks for the one value the rest of the product means by
496
+ // "unhealthy". If trawl_node's threshold ever changes, update this constant
497
+ // to match.
498
+ const UNHEALTHY_STREAK_LENGTH = 3;
429
499
  // list — promoted to a top-level verb (#108, see the AttachOptions comment above)
430
500
  export function attachListCommand(parent, attachOpts = {}) {
431
501
  return parent
@@ -443,6 +513,14 @@ export function attachListCommand(parent, attachOpts = {}) {
443
513
  // (~line 660) and report the actual bad input in the error message.
444
514
  .option('--limit <n>', 'Show only the first N results')
445
515
  .option('--page <n>', 'Fetch a specific page only (50 per page, no auto-pagination)')
516
+ // #179 — backed by the server's `minConsecutiveFailures` filter
517
+ // (trawl_node#1952): restricts the result set server-side to scraps
518
+ // whose CURRENT consecutive-failure streak is >= UNHEALTHY_STREAK_LENGTH,
519
+ // and annotates each returned scrap with the exact `consecutiveFailedRuns`/
520
+ // `unhealthySince` the aggregation computed. Deliberately never derived
521
+ // client-side from `history[]` — that would drift from what the web
522
+ // dashboard shows (comes-io/trawl_vue#1299 reads the identical filter).
523
+ .option('--unhealthy', `Only show scraps failing their last ${UNHEALTHY_STREAK_LENGTH}+ runs (server-computed)`)
446
524
  .action(async (opts, cmd) => {
447
525
  // #149 item 3 — validate --status against its enum the same way
448
526
  // --tier already validates (usageError + return, checked first, before
@@ -473,19 +551,23 @@ export function attachListCommand(parent, attachOpts = {}) {
473
551
  return;
474
552
  }
475
553
  }
554
+ // #179 — same param on every request this action can issue (both
555
+ // branches below), so a filtered result stays filtered across the
556
+ // fetch-all pagination loop, not just its first page.
557
+ const healthFilter = opts.unhealthy ? `&minConsecutiveFailures=${UNHEALTHY_STREAK_LENGTH}` : '';
476
558
  let data;
477
559
  try {
478
560
  data = await spin(async () => {
479
561
  if (page !== undefined) {
480
562
  // Single-page mode: explicit page requested, no loop
481
- return api.get(`/api/scraps?perPage=50&page=${page}`);
563
+ return api.get(`/api/scraps?perPage=50&page=${page}${healthFilter}`);
482
564
  }
483
565
  // Fetch-all mode: paginate until a page returns < 200 items
484
566
  const perPage = 200;
485
567
  let result = [];
486
568
  let pageNum = 1;
487
569
  while (true) {
488
- const batch = await api.get(`/api/scraps?perPage=${perPage}&page=${pageNum}`);
570
+ const batch = await api.get(`/api/scraps?perPage=${perPage}&page=${pageNum}${healthFilter}`);
489
571
  result = result.concat(batch);
490
572
  if (batch.length < perPage)
491
573
  break;
@@ -519,21 +601,28 @@ export function attachListCommand(parent, attachOpts = {}) {
519
601
  process.exitCode = exitCode;
520
602
  return;
521
603
  }
604
+ // #179 — check the RAW response, before --status narrows it further,
605
+ // so a genuinely empty (all-healthy) result is never mistaken for an
606
+ // unconfirmed filter.
607
+ warnIfUnconfirmedHealthFilter(data, opts.unhealthy);
522
608
  if (opts.status)
523
609
  data = data.filter((s) => lastStatus(s) === opts.status);
524
610
  const totalMatched = data.length;
525
611
  const rows = limit !== undefined ? data.slice(0, limit) : data;
526
612
  if (opts.json)
527
- return json(rows);
613
+ return json(withHealthFilterUnconfirmed(rows, data, opts.unhealthy));
528
614
  const tableRows = rows.map((s) => ({
529
615
  id: s._id,
530
616
  title: s.title || '(untitled)',
531
617
  cron: s.cron || '—',
532
618
  status: statusIcon(lastStatus(s)),
619
+ // #179 — next to the existing status badge, per the issue's own
620
+ // wording (see healthBadge() above for the two source signals).
621
+ health: healthBadge(s),
533
622
  'last run': lastRun(s),
534
623
  updated: formatDate(s.updatedAt),
535
624
  }));
536
- table(tableRows, ['id', 'title', 'cron', 'status', 'last run', 'updated']);
625
+ table(tableRows, ['id', 'title', 'cron', 'status', 'health', 'last run', 'updated']);
537
626
  // Print footer when --limit truncates
538
627
  if (limit !== undefined && rows.length < totalMatched) {
539
628
  console.log(chalk.dim(`Showing ${rows.length} of ${totalMatched} — omit --limit to see all`));
@@ -556,6 +645,10 @@ export function attachGetCommand(parent, attachOpts = {}) {
556
645
  console.log(chalk.dim(` ID: `) + data._id);
557
646
  console.log(chalk.dim(` Cron: `) + (data.cron || '—'));
558
647
  console.log(chalk.dim(` Status: `) + statusIcon(lastStatus(data)));
648
+ // #179 — same badge `list` renders (see healthBadge() above); `get`
649
+ // never sends `minConsecutiveFailures`, so this always reads the
650
+ // free `lastCronOutcome` breadcrumb, never the live per-run count.
651
+ console.log(chalk.dim(` Health: `) + healthBadge(data));
559
652
  console.log(chalk.dim(` Last run: `) + lastRun(data));
560
653
  console.log(chalk.dim(` Updated: `) + formatDate(data.updatedAt));
561
654
  });
@@ -1141,7 +1234,9 @@ export function attachHistoryCommand(parent, attachOpts = {}) {
1141
1234
  time: h.time ?? null,
1142
1235
  tier: h.proxyTier ?? null,
1143
1236
  failureKind: h.failureKind ?? null,
1144
- blockType: h.blockType ?? null,
1237
+ // trawl_node#1950 renamed blockType to block.kind; fall back to the
1238
+ // deprecated flat field for a prod backend not yet on that tag.
1239
+ blockType: h.block?.kind ?? h.blockType ?? null,
1145
1240
  createdAt: h.createdAt ?? null,
1146
1241
  }));
1147
1242
  if (opts.json) {
@@ -1167,7 +1262,9 @@ export function attachRunInfoCommand(parent, attachOpts = {}) {
1167
1262
  time: h.time ?? null,
1168
1263
  tier: h.proxyTier ?? null,
1169
1264
  failureKind: h.failureKind ?? null,
1170
- blockType: h.blockType ?? null,
1265
+ // trawl_node#1950 renamed blockType to block.kind; fall back to the
1266
+ // deprecated flat field for a prod backend not yet on that tag.
1267
+ blockType: h.block?.kind ?? h.blockType ?? null,
1171
1268
  errorMessage: h.errorSnapshot?.errorMessage ?? null,
1172
1269
  selector: h.errorSnapshot?.selector ?? null,
1173
1270
  emptyContext: h.emptyContext ?? null,
@@ -81,7 +81,7 @@ first-class on every one, and none of them ever blocks on a prompt (see
81
81
  ```
82
82
  trawl create <url> --prompt <goal> [--no-autofix] [--json] # url also accepted as --url <url> (#116)
83
83
  Create a persistent, self-healing scrap from a URL + a goal (AI-generated)
84
- trawl list|ls [--json] [--status <s>] [--limit <n>] [--page <n>] List all scraps
84
+ trawl list|ls [--json] [--status <s>] [--unhealthy] [--limit <n>] [--page <n>] List all scraps
85
85
  trawl get <id> [--json] Get scrap details
86
86
  trawl run <id> [--watch] [--json] Run a scrap (blocks until finished)
87
87
  trawl trigger <id> [--watch] [--wait] [--json] Launch a scrap as a background worker
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@trawlme/cli",
3
- "version": "3.8.2",
3
+ "version": "3.9.1",
4
4
  "description": "Trawl CLI — manage scraps from the terminal",
5
5
  "type": "module",
6
6
  "bin": {