@trawlme/cli 3.10.0 → 3.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -79,7 +79,7 @@ trawl spec [--json] Print a versioned, machine-readab
79
79
  - `list`/`get` also show a **health** badge next to the status badge. By default it reads the scrap's own `lastCronOutcome` field (always present, no extra cost): `⚠ cron paused` when the last scheduled tick was skipped for being unhealthy, `—` otherwise — a breadcrumb of the last tick, not a live read (never set on a no-cron scrap, can lag by one tick). `list --unhealthy` asks the API to filter to scraps whose *current* consecutive-failure streak is 3+ (trawl_node's own threshold) and annotates each with the exact `consecutiveFailedRuns`/`unhealthySince` — shown instead of the badge whenever present, and never re-derived client-side from history. `--json` carries whichever fields the request produced: `lastCronOutcome` always, `consecutiveFailedRuns`/`unhealthySince` only when `--unhealthy` was passed. **Old-server safety:** a server predating trawl_node#1952 silently ignores the unknown `--unhealthy` query param and returns everything, unfiltered — the CLI detects that (no returned item carries `consecutiveFailedRuns`) and prints a stderr warning instead of presenting the full list as "your unhealthy scraps".
80
80
  - **Long-running calls (`create`, `run`, `data --fresh`, `trigger --wait`):** these hit server-side paths that can legitimately take 30–250s+ (AI generation + scrap creation + a first run for `create`; proxy tier escalation + AI-fix retries for the other three) — the CLI arms a 300s timeout for exactly these four call sites instead of the generic 30s default. `TRAWL_TIMEOUT` (see below) still overrides ALL requests, including these — set it if you need a tighter or looser ceiling than 300s, but note a global override that tight also clamps `create`.
81
81
  - **`--watch` is poll-based, not a live stream:** the activities SSE endpoint has no backlog and, for the default async `trigger` (no `--wait`), runs in a separate cron-consumer pod whose events never reach the API pod holding the SSE connection — a naive "await the run, then open SSE" shows nothing. `run --watch` and `trigger --watch` instead poll `GET /api/scraps/:id` (terminal status) and the activities REST list until the run finishes, printing each new activity line as it appears. The watched run's outcome drives the exit code too, in BOTH human and `--json` mode: a genuinely failed terminal run, a poll timeout, or a persistently unreachable API all exit non-zero — a clean successful run is the only exit `0`. A run that never reaches a terminal status within 300s prints an honest timeout notice pointing at `scraps doctor <id>` (human mode) — see the `--json` shape below. `scraps watch <id>` (the standalone command, no trigger) is unchanged — it still opens the live SSE stream directly.
82
- - `spec --json` prints the CLI's own command tree — `{specVersion, cliVersion, commands[], exitCodes, errorKinds, kindExitCodes}` — DERIVED at runtime by walking the live commander tree (never a hand-maintained file, which would silently drift from reality). Each entry in `commands[]` carries its full path (e.g. `"scraps account session set"`), description, `hidden` (the legacy `scraps <verb>` aliases above), `leaf` (false for a pure namespace/group node like `scraps`/`skills`/`telemetry` — invoking one directly is a guaranteed-failing tool, not a real command), `aliases`, `arguments`, and `options`; `exitCodes`/`errorKinds`/`kindExitCodes` are read from the exact same source `classifyError` uses (see [Exit codes](#exit-codes)) — never a second copy. `kindExitCodes` is the inverse of `exitCodes`: `kind -> exitCode`, since exit code `1` alone is a shared bucket (`api`/`refused`/`unknown`/`in_progress`/`run_failed`/`upgrade_failed`) that `exitCodes`' flat label can't disambiguate. Without `--json`, `spec` prints one short human line (version + visible command count) pointing at `--json`.
82
+ - `spec --json` prints the CLI's own command tree — `{specVersion, cliVersion, commands[], exitCodes, errorKinds, kindExitCodes, docsUrl?, llmsUrl?}` — DERIVED at runtime by walking the live commander tree (never a hand-maintained file, which would silently drift from reality). Each entry in `commands[]` carries its full path (e.g. `"scraps account session set"`), description, `hidden` (the legacy `scraps <verb>` aliases above), `leaf` (false for a pure namespace/group node like `scraps`/`skills`/`telemetry` — invoking one directly is a guaranteed-failing tool, not a real command), `aliases`, `arguments`, `options`, and an optional per-command `docs` deep link (e.g. every `scraps account *` command points at the account-sessions guide); `exitCodes`/`errorKinds`/`kindExitCodes` are read from the exact same source `classifyError` uses (see [Exit codes](#exit-codes)) — never a second copy. `kindExitCodes` is the inverse of `exitCodes`: `kind -> exitCode`, since exit code `1` alone is a shared bucket (`api`/`refused`/`unknown`/`in_progress`/`run_failed`/`upgrade_failed`) that `exitCodes`' flat label can't disambiguate. `docsUrl` is resolved server-first (`externalDocs.url` on the configured API base's own OpenAPI document, when declared), else derived from a KNOWN first-party `trawl.me` API host, else **omitted entirely** — never a guessed URL pointed at the wrong docs host for a self-hosted install. `llmsUrl` and every per-command `docs` link are narrower: they only ever come from that same known-host derivation, NEVER from a server-declared `docsUrl` — this CLI's own guide slugs have no reason to exist on a third party's own docs root, so a self-hosted server that declares `externalDocs.url` gets a correct top-level `docsUrl` but no `llmsUrl` and no per-command `docs` links. A `failureKind:"auth"` run's JSON payload (`doctor`/`data --errors`/`run-info`) carries the same `docs` field, by the same known-host-only rule. Without `--json`, `spec` prints one short human line (version + visible command count) pointing at `--json`.
83
83
  - `run --json`/`trigger --json` bypass the spinner and print the raw launch/trigger payload on stdout; combined with `--watch`, every intermediate progress line stays suppressed (stdout stays pure JSON) and, once the watch reaches its outcome, exactly ONE final NDJSON line is emitted: `{"runId","status"}` (the honest terminal status — `success`/`error`/`empty`/`regression`/…), `{"runId","status":"timeout"}` on a poll timeout, or `{"runId","status":"poll_error","error"}` if the API stays unreachable for several consecutive polls — `process.exitCode` is non-zero for all three except a genuine success. `scraps watch --json` emits one raw JSON object per activity line (NDJSON) instead of the formatted `[time] message` text — there's no single final payload to wait for on a live stream.
84
84
 
85
85
  ### Scrap management
@@ -113,9 +113,12 @@ trawl scraps account delete <id> [--force] [--json]
113
113
  trawl scraps account clear-session <id> [--json]
114
114
  trawl scraps account status <id> [--json]
115
115
  trawl scraps account session set <id> -c <file> [--json]
116
+ trawl scraps account session capture <id> [--chrome <path>] [--json]
116
117
  ```
117
118
 
118
- `account session set` uploads a Puppeteer cookie JSON array to bootstrap a logged-in session without storing credentials (BYO-cookies). A missing `-u/--username`/`-p/--password` on `account set` follows the same [non-interactive rule](#non-interactive-rule) as `login`.
119
+ `account session set` uploads a cookie JSON array — or a `{ cookies, origins }` storageState file (see `session capture` below) — to bootstrap a logged-in session without storing credentials (BYO-cookies). A missing `-u/--username`/`-p/--password` on `account set` follows the same [non-interactive rule](#non-interactive-rule) as `login`.
120
+
121
+ `account session capture <id>` is the one-command version of the same idea: it opens a real, visible Chrome window at the scrap's target URL, waits for **you** to log in there — 2FA included — then reads the resulting session over the Chrome DevTools Protocol (cookies **and** per-origin localStorage) and uploads it. You stay authenticated as yourself; Trawl never sees your credentials, and responsibility for lawful use of the captured session stays with you. The capture is scoped to the scrap's own target host by RFC 6265 domain-matching, the same rule a real browser uses to decide which cookies to send: a cookie explicitly scoped to a domain (e.g. `Domain=example.com`) is in scope for that host and its subdomains, while a host-only cookie (no `Domain` attribute) is in scope only for the exact host it was set on — never a sibling subdomain, even one sharing a parent domain, and never another tenant on the same multi-tenant hosting domain (e.g. a different `*.github.io` site). localStorage follows the same anchor: an origin is in scope only if it's the target host itself or a subdomain of it. The one deliberate gap: a host-only cookie set on a sibling host (e.g. `auth.example.com` while the scrap targets `app.example.com`) is excluded — a real browser would never send it to the target host either, so nothing replay-relevant is lost unless the scrap script itself later navigates to that sibling. Completion is explicit: press Enter in the terminal once you're signed in — the primary path, and the only one guaranteed cross-platform. Closing the Chrome window also completes the capture (cookies only, not localStorage) when Chrome itself stays running after its last window closes, which this CLI verified on macOS; on a platform where closing the last window quits Chrome entirely, the command reports that fact instead of a session, and pressing Enter is the reliable path there. This command needs a local interactive terminal with a real display; it does not work headless, in CI, or over a plain SSH session, and auto-detects a system Chrome/Chromium (override with `--chrome <path>` or `TRAWL_CHROME_PATH`). Never prints the captured session itself, in any mode — `--json` echoes only the server's response plus capture counts.
119
122
 
120
123
  ### Claude Code skills
121
124
 
@@ -12,6 +12,36 @@ import { formatDoctor, formatAutofix, fetchRunAndFix, pickRun, pickFix, isAuthWa
12
12
  import { renderPinch, pinchEnabled } from '../lib/pinch.js';
13
13
  import { maybeShowReferralTip } from '../lib/tips.js';
14
14
  import { maybeSuggestSkillsForAuthWall } from '../lib/skillsNudge.js';
15
+ import { captureSession, checkInteractiveEnvironment, LOCALSTORAGE_READ_TIMEOUT_MS, } from '../lib/session-capture.js';
16
+ import { assertSecureTransport } from '../lib/secure-transport.js';
17
+ import { getApiUrl } from '../lib/config.js';
18
+ import { resolveDocsUrls, resolveFailureKindDocsUrl } from '../lib/docs.js';
19
+ /**
20
+ * #185 — `{ docs }` (spreadable, empty object when nothing to add) for a
21
+ * run's JSON payload (`doctor`/`data --errors`/`run-info`), keyed on
22
+ * trawl_node's `failureKind` (NOT errors.ts's unrelated `ErrorEnvelope.kind`
23
+ * — see docs.ts's `FAILURE_KIND_DOC_PATHS` doc comment for why those two
24
+ * `'auth'`s must never be conflated). Reuses `isAuthWall` — the SAME
25
+ * staleness guard `formatDoctor`'s human hint and the #184 skills nudge
26
+ * already share (a `failureKind:'auth'` stamp survives a later
27
+ * success/regression patch, see `isAuthWall`'s own doc comment) — so a run
28
+ * that ultimately succeeded or degraded never carries a stale `docs` link
29
+ * either, exactly the invariant `detectWallVendor`/`isAuthWall` already
30
+ * enforce for the human-facing surfaces. One call site, one rule, reused by
31
+ * all three JSON surfaces below instead of three copies that could drift.
32
+ */
33
+ function runDocsField(run) {
34
+ if (!isAuthWall(run))
35
+ return {};
36
+ // Whole resolved object (never just `.docsUrl`) so resolveFailureKindDocsUrl
37
+ // can gate on rung provenance — see docs.ts's DocsUrls.docsUrlIsDerived.
38
+ // This call never supplies `externalDocsUrl`, so `docsUrl` here can only
39
+ // ever be rung-2-derived or absent — never a rung-1 server-declared root —
40
+ // but the gate stays explicit rather than relying on that as an invariant.
41
+ const resolved = resolveDocsUrls({ apiBaseUrl: getApiUrl() });
42
+ const docs = resolveFailureKindDocsUrl(run.failureKind, resolved);
43
+ return docs ? { docs } : {};
44
+ }
15
45
  /**
16
46
  * Print a usage/validation error consistently: human text to stderr, or a
17
47
  * machine envelope on stdout under --json (never both — reportError is the
@@ -1065,8 +1095,12 @@ export function attachDataCommand(parent, attachOpts = {}) {
1065
1095
  // --json is honored for BOTH outcomes (success or failure) — an agent
1066
1096
  // parsing `data --errors --json` must always get the flat run object,
1067
1097
  // never prose gated behind a status check. (#71 finding 13)
1098
+ // #185 — `docs` rides alongside as a sibling field on the same flat
1099
+ // object (never nested under `run`, since there's no wrapper shape
1100
+ // here), present only for a live mapped failureKind (see
1101
+ // runDocsField's doc comment).
1068
1102
  if (opts.json)
1069
- return json(pickRun(result.run));
1103
+ return json({ ...pickRun(result.run), ...runDocsField(result.run) });
1070
1104
  if (result.run.status === true) {
1071
1105
  console.log(chalk.green('✓ Last run succeeded. No errors to show.'));
1072
1106
  return;
@@ -1270,6 +1304,11 @@ export function attachRunInfoCommand(parent, attachOpts = {}) {
1270
1304
  selector: h.errorSnapshot?.selector ?? null,
1271
1305
  emptyContext: h.emptyContext ?? null,
1272
1306
  createdAt: h.createdAt ?? null,
1307
+ // #185 — present only for a live mapped failureKind (see
1308
+ // runDocsField's doc comment); omitted, never `docs: undefined`, for
1309
+ // every other run so this object's own JSON.stringify never grows an
1310
+ // unexpected key on existing consumers.
1311
+ ...runDocsField(h),
1273
1312
  };
1274
1313
  if (opts.json) {
1275
1314
  json(info);
@@ -1558,32 +1597,69 @@ account
1558
1597
  const accountSession = account
1559
1598
  .command('session')
1560
1599
  .description('Manage scrap account session cookies (flavour B BYO-cookies)');
1600
+ /** Structural check only (bare origin format, {name,value} shape) — the
1601
+ * server's sessionShape.js does the authoritative validation. No domain
1602
+ * scoping here: unlike `capture`'s CDP path, a hand-authored/exported file
1603
+ * has no "target URL" to scope against, and the human supplying it already
1604
+ * chose what to include. */
1605
+ function isValidOriginsShape(value) {
1606
+ return (Array.isArray(value) &&
1607
+ value.every((o) => o &&
1608
+ typeof o === 'object' &&
1609
+ typeof o.origin === 'string' &&
1610
+ Array.isArray(o.localStorage) &&
1611
+ o.localStorage.every((kv) => kv && typeof kv === 'object' && typeof kv.name === 'string' && typeof kv.value === 'string')));
1612
+ }
1561
1613
  // account session set
1562
1614
  accountSession
1563
1615
  .command('set <id>')
1564
- .description('Upload browser session cookies for a scrap (Puppeteer cookie JSON array)')
1565
- .requiredOption('-c, --cookies <file>', 'Path to a Puppeteer cookie JSON array file')
1616
+ .description('Upload a browser session for a scrap — a cookie JSON array, or a { cookies, origins } storageState file (see `session capture`)')
1617
+ .requiredOption('-c, --cookies <file>', 'Path to a cookie JSON array, or a storageState file ({ cookies, origins? })')
1566
1618
  .option('--json', 'Output as JSON')
1567
1619
  .action(async (id, opts) => {
1568
1620
  validateObjectId(id);
1621
+ // A session is a bearer-equivalent secret — refuse to PUT it over
1622
+ // plain HTTP (trawl_cli#183 review finding 5). Checked before anything
1623
+ // else in this action, including the file read below.
1624
+ const secureTransportError = assertSecureTransport(getApiUrl());
1625
+ if (secureTransportError) {
1626
+ usageError(secureTransportError, { json: opts.json });
1627
+ return;
1628
+ }
1569
1629
  const { existsSync, readFileSync } = await import('fs');
1570
1630
  if (!existsSync(opts.cookies)) {
1571
1631
  usageError(`File not found: ${opts.cookies}`, { json: opts.json });
1572
1632
  return;
1573
1633
  }
1574
1634
  let cookies;
1635
+ let origins;
1575
1636
  try {
1576
1637
  const raw = readFileSync(opts.cookies, 'utf-8');
1577
- cookies = JSON.parse(raw);
1638
+ const parsed = JSON.parse(raw);
1639
+ if (Array.isArray(parsed)) {
1640
+ cookies = parsed;
1641
+ }
1642
+ else if (parsed && typeof parsed === 'object' && Array.isArray(parsed.cookies)) {
1643
+ // storageState-shaped file — trawl_cli#183's `session capture` output shape.
1644
+ const shaped = parsed;
1645
+ cookies = shaped.cookies;
1646
+ if (shaped.origins !== undefined) {
1647
+ if (!isValidOriginsShape(shaped.origins)) {
1648
+ usageError('origins must be an array of { origin: string, localStorage: [{name, value}] }', { json: opts.json });
1649
+ return;
1650
+ }
1651
+ origins = shaped.origins;
1652
+ }
1653
+ }
1654
+ else {
1655
+ usageError('Cookies file must contain a JSON array, or a { cookies, origins? } storageState object', { json: opts.json });
1656
+ return;
1657
+ }
1578
1658
  }
1579
1659
  catch (e) {
1580
1660
  usageError(`Failed to parse cookies file: ${e.message}`, { json: opts.json });
1581
1661
  return;
1582
1662
  }
1583
- if (!Array.isArray(cookies)) {
1584
- usageError('Cookies file must contain a JSON array', { json: opts.json });
1585
- return;
1586
- }
1587
1663
  if (cookies.length === 0) {
1588
1664
  usageError('Cookies array must not be empty', { json: opts.json });
1589
1665
  return;
@@ -1592,7 +1668,10 @@ accountSession
1592
1668
  usageError('Each cookie must have a name (string) and value (string)', { json: opts.json });
1593
1669
  return;
1594
1670
  }
1595
- const call = () => api.put(`/api/scraps/${id}/account/session`, { cookies });
1671
+ const body = { cookies };
1672
+ if (origins)
1673
+ body.origins = origins;
1674
+ const call = () => api.put(`/api/scraps/${id}/account/session`, body);
1596
1675
  if (opts.json) {
1597
1676
  const data = await call();
1598
1677
  json(data);
@@ -1605,6 +1684,164 @@ accountSession
1605
1684
  const acc = data.account;
1606
1685
  console.log(chalk.dim(' Session: ') + (acc.hasSession ? chalk.green('✓ active') : chalk.dim('none')));
1607
1686
  });
1687
+ /**
1688
+ * @desc Map a failed `captureSession()` result onto this CLI's existing
1689
+ * exit-code taxonomy (#71/#88/#107). `no_cookies_in_scope` is the one
1690
+ * outcome where everything ran correctly but nothing was found to upload —
1691
+ * a business-level refusal (exit 1, kind:"refused"), not a usage mistake.
1692
+ *
1693
+ * `capture_failed` (trawl_cli#183 R4) is a SECOND exception, for the
1694
+ * opposite reason: launch + handshake already succeeded — Chrome was
1695
+ * reachable over CDP, a page was already open — and the failure is a
1696
+ * genuine bug in this CLI (or an unanticipated non-page-scoped failure),
1697
+ * never a "cannot run in this environment" problem. Routing it through
1698
+ * `usageError`'s exit 2 / "cannot complete in the current environment"
1699
+ * framing would misclassify it exactly the way a bare `launch_failed` used
1700
+ * to before the R4 fix. Reported instead through `reportError` with a bare
1701
+ * `Error` — `classifyError` has no dedicated `kind` for it, so it falls
1702
+ * into the generic `kind:"unknown"`/exit 1 bucket, the SAME code an
1703
+ * unhandled bug reaching index.ts's own top-level catch would already get.
1704
+ * Never `RefusalError` (this is not a business-logic refusal, everything
1705
+ * up to this point may have gone fine) and never `UsageError` (nothing
1706
+ * about the invocation or the environment was wrong).
1707
+ *
1708
+ * Every remaining reason (non_interactive, no_chrome, launch_failed, the CDP
1709
+ * pipe closing before capture, a corrupt CDP frame tearing it down) means
1710
+ * THIS invocation cannot complete in the current environment — exit 2, the
1711
+ * same bucket `confirmDestructive`'s own non-interactive refusal uses.
1712
+ */
1713
+ function reportCaptureFailure(result, opts) {
1714
+ if (result.reason === 'no_cookies_in_scope') {
1715
+ process.exitCode = reportError(new RefusalError(result.message), { json: opts.json });
1716
+ return;
1717
+ }
1718
+ if (result.reason === 'capture_failed') {
1719
+ process.exitCode = reportError(new Error(result.message), { json: opts.json });
1720
+ return;
1721
+ }
1722
+ usageError(result.message, { json: opts.json });
1723
+ }
1724
+ // account session capture — trawl_cli#183
1725
+ accountSession
1726
+ .command('capture <id>')
1727
+ .description('Open a headed Chrome to the scrap\'s target URL, wait for you to log in (2FA included), and capture the session over CDP — cookies AND localStorage. Requires a local interactive terminal with a display; does not work headless, in CI, or over a plain SSH session. You stay authenticated as yourself — Trawl never sees your credentials. Responsibility for lawful use of the captured session stays with you.')
1728
+ .option('--chrome <path>', 'Path to a Chrome/Chromium executable (auto-detected if omitted)')
1729
+ .option('--json', 'Output as JSON (counts only — a session is a bearer secret and is never printed, in any mode)')
1730
+ .action(async (id, opts) => {
1731
+ validateObjectId(id);
1732
+ // A captured session is a bearer-equivalent secret — refuse to PUT it
1733
+ // over plain HTTP (trawl_cli#183 review finding 5). Checked before
1734
+ // anything else, including the scrap lookup below (which itself
1735
+ // already sends the auth token over whatever transport is configured).
1736
+ const secureTransportError = assertSecureTransport(getApiUrl());
1737
+ if (secureTransportError) {
1738
+ usageError(secureTransportError, { json: opts.json });
1739
+ return;
1740
+ }
1741
+ const scrap = await api.get(`/api/scraps/${id}`);
1742
+ if (!scrap.url) {
1743
+ usageError(`Scrap ${id} has no target URL configured — nothing to open a browser to. A URL can be set with: trawl scraps update ${id} -u <url>`, { json: opts.json });
1744
+ return;
1745
+ }
1746
+ // Validated here, before captureSession ever creates a temp profile:
1747
+ // `existsSync` alone doesn't mean `spawn()` can run the file — a
1748
+ // non-executable path fails asynchronously deep inside chrome-launch.ts
1749
+ // instead. That failure is still caught there (an 'error' listener is
1750
+ // the safety net for anything not caught by this pre-flight — a race,
1751
+ // a permission change after this check, PUPPETEER_EXECUTABLE_PATH/
1752
+ // CHROME_PATH), but a pre-flight check gives a faster, more specific
1753
+ // message for the common case: an explicit path the user pointed at.
1754
+ if (opts.chrome) {
1755
+ const { existsSync, accessSync, constants } = await import('fs');
1756
+ if (!existsSync(opts.chrome)) {
1757
+ usageError(`Chrome not found at: ${opts.chrome}`, { json: opts.json });
1758
+ return;
1759
+ }
1760
+ try {
1761
+ accessSync(opts.chrome, constants.X_OK);
1762
+ }
1763
+ catch {
1764
+ usageError(`Chrome exists at ${opts.chrome} but is not executable.`, { json: opts.json });
1765
+ return;
1766
+ }
1767
+ }
1768
+ else if (process.env.TRAWL_CHROME_PATH) {
1769
+ // Same check, same reasoning, for the env-var form of an explicit
1770
+ // path — findChrome() gives TRAWL_CHROME_PATH top precedence and
1771
+ // would otherwise hand this same non-executable path straight to
1772
+ // captureSession.
1773
+ const trawlChromePath = process.env.TRAWL_CHROME_PATH;
1774
+ const { existsSync, accessSync, constants } = await import('fs');
1775
+ if (existsSync(trawlChromePath)) {
1776
+ try {
1777
+ accessSync(trawlChromePath, constants.X_OK);
1778
+ }
1779
+ catch {
1780
+ usageError(`Chrome exists at ${trawlChromePath} (TRAWL_CHROME_PATH) but is not executable.`, { json: opts.json });
1781
+ return;
1782
+ }
1783
+ }
1784
+ }
1785
+ // Check the same guard captureSession runs internally BEFORE printing
1786
+ // anything about a Chrome window that may never open — otherwise a
1787
+ // non-interactive run saw "a window will open" immediately followed by
1788
+ // "this cannot work headless" (trawl_cli#183 gap).
1789
+ const nonInteractiveReason = checkInteractiveEnvironment();
1790
+ if (nonInteractiveReason) {
1791
+ reportCaptureFailure({ ok: false, reason: 'non_interactive', message: nonInteractiveReason }, opts);
1792
+ return;
1793
+ }
1794
+ // Progress/prompts are stderr-only, in every mode — stdout under
1795
+ // --json must stay a single parseable document (#88/#107's rule,
1796
+ // restated for this command by trawl_cli#183's own hard rules).
1797
+ console.error(chalk.dim(`Chrome will open at ${scrap.url} for you to log in there (2FA included); capture completes when you press Enter back in this terminal.`));
1798
+ const result = await captureSession(scrap.url, opts.chrome ? { findChrome: () => opts.chrome ?? null } : {});
1799
+ if (!result.ok) {
1800
+ reportCaptureFailure(result, opts);
1801
+ return;
1802
+ }
1803
+ const call = () => api.put(`/api/scraps/${id}/account/session`, {
1804
+ cookies: result.storageState.cookies,
1805
+ origins: result.storageState.origins,
1806
+ });
1807
+ if (opts.json) {
1808
+ const data = await call();
1809
+ // Never the session itself — only what the server echoes back, plus
1810
+ // counts (#183 hard rule: no cookie/localStorage VALUE, ever, under
1811
+ // --json or otherwise).
1812
+ json({ account: data.account, targetDomain: result.targetDomain, capture: result.counts });
1813
+ return;
1814
+ }
1815
+ const data = await spin(call, {
1816
+ text: `Uploading captured session for scrap ${chalk.bold(id)}…`,
1817
+ successText: `Session captured for scrap ${chalk.bold(id)}: ${result.counts.cookiesCaptured} cookie(s), ${result.counts.originsCaptured} origin(s) in scope for ${result.targetDomain}`,
1818
+ });
1819
+ const acc = data.account;
1820
+ console.log(chalk.dim(' Session: ') + (acc.hasSession ? chalk.green('✓ active') : chalk.dim('none')));
1821
+ if (result.counts.cookiesDroppedOutOfScope > 0 || result.counts.originsDroppedOutOfScope > 0) {
1822
+ console.log(chalk.dim(` Scoped to ${result.targetDomain}: dropped ${result.counts.cookiesDroppedOutOfScope} cookie(s) and ${result.counts.originsDroppedOutOfScope} origin(s) outside that domain.`));
1823
+ }
1824
+ if (result.counts.closedEarly) {
1825
+ console.log(chalk.dim(' The browser window was closed before Enter — localStorage was not captured (cookies only).'));
1826
+ }
1827
+ if (result.counts.originsUnreadable > 0) {
1828
+ // Impossible to overlook (chalk.yellow + ⚠, not the routine chalk.dim
1829
+ // scope-drop note above) — this is a DEGRADED capture: cookies still
1830
+ // uploaded, but at least one in-scope origin's localStorage did not.
1831
+ // Two distinct causes share this one count: the page threw reading
1832
+ // `window.localStorage` (e.g. a SecurityError on partitioned
1833
+ // storage), or the read never got a response within
1834
+ // LOCALSTORAGE_READ_TIMEOUT_MS (the page's renderer was blocked — a
1835
+ // native dialog, a synchronous script, a paused debugger;
1836
+ // trawl_cli#183 review finding 3's own pipe-level timeout fix
1837
+ // reopened finding 4's exact bug class on this one path, since a
1838
+ // timeout does not close the pipe the way a corrupt frame does). The
1839
+ // two are not split apart here: both mean the same actionable fact
1840
+ // ("re-run once nothing is blocking that tab") and neither one names
1841
+ // the page URL/title/exception text — a page controls that content.
1842
+ console.log(chalk.yellow(` ⚠ ${result.counts.originsUnreadable} origin(s) could not be captured — threw reading localStorage, or gave no response within ${LOCALSTORAGE_READ_TIMEOUT_MS / 1000}s (a blocked tab). Not the same as being empty; cookies were still captured.`));
1843
+ }
1844
+ });
1608
1845
  // account status
1609
1846
  account
1610
1847
  .command('status <id>')
@@ -1662,8 +1899,11 @@ scraps
1662
1899
  console.log(chalk.dim('No runs yet.'));
1663
1900
  return;
1664
1901
  }
1902
+ // #185 — `docs` sits alongside `run`/`fix` (present only for a live
1903
+ // mapped failureKind — see runDocsField's doc comment), never nested
1904
+ // inside the allowlisted `run` object.
1665
1905
  if (opts.json)
1666
- return json({ run: pickRun(result.run), fix: pickFix(result.fix) });
1906
+ return json({ run: pickRun(result.run), fix: pickFix(result.fix), ...runDocsField(result.run) });
1667
1907
  console.log(formatDoctor(result.scrap.title, result.run, result.fix, id));
1668
1908
  if (opts.autofix && result.fix) {
1669
1909
  console.log('\n' + formatAutofix(result.fix));
@@ -1,4 +1,5 @@
1
1
  import { Command } from 'commander';
2
+ import { type DocsUrls } from '../lib/docs.js';
2
3
  /**
3
4
  * `trawl spec --json` (#170) — a versioned, machine-readable description of
4
5
  * the command tree, so an agent can learn the CLI's surface without parsing
@@ -54,6 +55,18 @@ export interface CliSpecCommand {
54
55
  aliases: string[];
55
56
  arguments: CliSpecArgument[];
56
57
  options: CliSpecOption[];
58
+ /**
59
+ * #185 — a deep link into a specific guide when one is mapped for this
60
+ * command (docs.ts's `COMMAND_DOC_PATHS`, e.g. every `scraps account *`
61
+ * command -> the account-sessions guide). Omitted (never a guessed/generic
62
+ * link) whenever no guide is mapped, the top-level `docsUrl` itself
63
+ * couldn't be resolved, OR `docsUrl` resolved from rung 1 (a server-
64
+ * declared `externalDocs.url`) rather than rung 2 (this CLI's own known-host
65
+ * derivation) — appending our guide slug onto a third party's own docs root
66
+ * is exactly the confidently-wrong-404 defect this omission closes. See
67
+ * `resolveCommandDocsUrl` / `DocsUrls.docsUrlIsDerived`.
68
+ */
69
+ docs?: string;
57
70
  }
58
71
  export interface CliSpec {
59
72
  specVersion: 1;
@@ -81,12 +94,30 @@ export interface CliSpec {
81
94
  * (errors.ts's KIND_EXIT_CODES) — never a second copy.
82
95
  */
83
96
  kindExitCodes: Record<string, number>;
97
+ /**
98
+ * #185 — where the human-readable guides live, resolved via docs.ts's
99
+ * ladder (server `externalDocs.url` > derive from the API base's known
100
+ * `trawl.me` host > omit). Absent (never a guessed URL) when neither rung
101
+ * resolves — e.g. a self-hosted install with no `externalDocs` declared.
102
+ */
103
+ docsUrl?: string;
104
+ /** #185 — the llms.txt entry point alongside `docsUrl`, when resolvable.
105
+ * See docs.ts's module doc comment for why this never rides along with a
106
+ * server-declared `docsUrl` it wasn't itself derived from. */
107
+ llmsUrl?: string;
84
108
  }
85
109
  /**
86
110
  * Build the full spec from a live, already-constructed program (e.g.
87
111
  * `createProgram()`'s return value). Includes every node in the tree —
88
112
  * hidden legacy aliases (`scraps list`, …) included, flagged via `hidden`,
89
113
  * so a consumer that wants only the advertised surface can filter on it.
114
+ *
115
+ * #185 — `docs` is an already-RESOLVED `DocsUrls` (docs.ts's ladder run to
116
+ * completion), never computed in here: `buildSpec` stays pure/synchronous
117
+ * (no network, no config read) exactly as before, so every existing caller
118
+ * (the README-vs-spec contract test, the synthetic-tree unit tests) keeps
119
+ * working unchanged by simply omitting the parameter. Only `spec.ts`'s own
120
+ * action resolves rung 1 (a bounded fetch) before calling this.
90
121
  */
91
- export declare function buildSpec(program: Command): CliSpec;
122
+ export declare function buildSpec(program: Command, docs?: DocsUrls): CliSpec;
92
123
  export declare const spec: Command;
@@ -1,6 +1,9 @@
1
1
  import { Command, Help } from 'commander';
2
2
  import { json } from '../lib/format.js';
3
3
  import { EXIT_CODE_LABELS, ENVELOPE_KINDS, KIND_EXIT_CODES } from '../lib/errors.js';
4
+ import { getApiUrl } from '../lib/config.js';
5
+ import { api } from '../lib/api.js';
6
+ import { resolveDocsUrls, resolveCommandDocsUrl } from '../lib/docs.js';
4
7
  function buildOption(option) {
5
8
  const entry = {
6
9
  long: option.long ?? '',
@@ -15,8 +18,8 @@ function buildOption(option) {
15
18
  entry.default = option.defaultValue;
16
19
  return entry;
17
20
  }
18
- function buildCommandEntry(cmd, name, hidden) {
19
- return {
21
+ function buildCommandEntry(cmd, name, hidden, docs = {}) {
22
+ const entry = {
20
23
  name,
21
24
  description: cmd.description(),
22
25
  hidden,
@@ -29,20 +32,29 @@ function buildCommandEntry(cmd, name, hidden) {
29
32
  })),
30
33
  options: cmd.options.map(buildOption),
31
34
  };
35
+ const docsLink = resolveCommandDocsUrl(name, docs);
36
+ if (docsLink)
37
+ entry.docs = docsLink;
38
+ return entry;
32
39
  }
33
40
  /**
34
41
  * Walk `cmd.commands` recursively, collecting one entry per node under its
35
42
  * FULL path name. `hidden` is read via commander's own
36
43
  * `Help#visibleCommands()` — the same idiom this codebase already uses
37
44
  * (scraps.test.ts's #108 describe block) — rather than reaching for the
38
- * private, untyped `_hidden` field directly.
45
+ * private, untyped `_hidden` field directly. The resolved `docs` object,
46
+ * when present, threads through so every entry can carry its own per-command
47
+ * `docs` deep link (#185) off the SAME resolved root — never a second lookup
48
+ * per node, and never flattened to a bare string (see docs.ts's
49
+ * `DocsUrls.docsUrlIsDerived` — the rung provenance must survive the walk
50
+ * unchanged so `resolveCommandDocsUrl` can still gate on it at each leaf).
39
51
  */
40
- function walk(cmd, prefix, out) {
52
+ function walk(cmd, prefix, out, docs = {}) {
41
53
  const visible = new Set(new Help().visibleCommands(cmd));
42
54
  for (const sub of cmd.commands) {
43
55
  const name = prefix ? `${prefix} ${sub.name()}` : sub.name();
44
- out.push(buildCommandEntry(sub, name, !visible.has(sub)));
45
- walk(sub, name, out);
56
+ out.push(buildCommandEntry(sub, name, !visible.has(sub), docs));
57
+ walk(sub, name, out, docs);
46
58
  }
47
59
  }
48
60
  /**
@@ -50,11 +62,18 @@ function walk(cmd, prefix, out) {
50
62
  * `createProgram()`'s return value). Includes every node in the tree —
51
63
  * hidden legacy aliases (`scraps list`, …) included, flagged via `hidden`,
52
64
  * so a consumer that wants only the advertised surface can filter on it.
65
+ *
66
+ * #185 — `docs` is an already-RESOLVED `DocsUrls` (docs.ts's ladder run to
67
+ * completion), never computed in here: `buildSpec` stays pure/synchronous
68
+ * (no network, no config read) exactly as before, so every existing caller
69
+ * (the README-vs-spec contract test, the synthetic-tree unit tests) keeps
70
+ * working unchanged by simply omitting the parameter. Only `spec.ts`'s own
71
+ * action resolves rung 1 (a bounded fetch) before calling this.
53
72
  */
54
- export function buildSpec(program) {
73
+ export function buildSpec(program, docs = {}) {
55
74
  const commands = [];
56
- walk(program, '', commands);
57
- return {
75
+ walk(program, '', commands, docs);
76
+ const spec = {
58
77
  specVersion: 1,
59
78
  cliVersion: program.version() ?? 'unknown',
60
79
  commands,
@@ -62,18 +81,71 @@ export function buildSpec(program) {
62
81
  errorKinds: [...ENVELOPE_KINDS],
63
82
  kindExitCodes: { ...KIND_EXIT_CODES },
64
83
  };
84
+ if (docs.docsUrl)
85
+ spec.docsUrl = docs.docsUrl;
86
+ if (docs.llmsUrl)
87
+ spec.llmsUrl = docs.llmsUrl;
88
+ return spec;
89
+ }
90
+ /**
91
+ * #185 — bounds the SOCKET, not just the promise, so a slow/dead server can
92
+ * never perceptibly slow `spec --json` down. Mirrors lib/tips.ts's
93
+ * FLAG_CHECK_TIMEOUT_MS for the same reason: this is the one surface allowed
94
+ * a network call at all, and only because it's a deliberate, once-per-
95
+ * invocation agent probe.
96
+ *
97
+ * This is a REAL wall-clock ceiling, not merely a promise-level one: an
98
+ * earlier version of this fetch used `api.publicGet` (Node's global
99
+ * `fetch()` + `AbortSignal.timeout()`), which bounds the promise but not the
100
+ * OS-level connection — against a routable-but-silently-dropping host (a
101
+ * dropped SYN, no RST; a realistic firewalled/air-gapped shape, the very
102
+ * audience this feature cites) the promise rejected at ~2005ms while the
103
+ * PROCESS didn't exit until ~10.5s later, waiting out Node's own internal
104
+ * connect-timeout. `api.probeJson` (see its own doc comment in lib/api.ts)
105
+ * fixes this by managing the raw socket directly — an independent timer
106
+ * that `req.destroy()`s the connection, never relying on the idle-based
107
+ * `http.request` `timeout` option. Re-measured after the fix, same
108
+ * black-holed host, via the exact repro command (`time env
109
+ * TRAWL_API_URL=https://192.0.2.1 node dist/index.js spec --json`): total
110
+ * process time ~2.43s/~2.47s across two runs (the 2000ms bound plus Node
111
+ * startup and spec-tree construction), not ~10.68s.
112
+ */
113
+ const EXTERNAL_DOCS_FETCH_TIMEOUT_MS = 2_000;
114
+ /**
115
+ * Rung 1 — read `externalDocs.url` off the OpenAPI document the configured
116
+ * API base itself serves at `/api/spec.json`. Unset (null) server-side
117
+ * today, so this resolves to `null` in practice; forward-looking for when
118
+ * it isn't. ANY failure — network error, timeout, a malformed/unexpected
119
+ * response shape — is indistinguishable from "not declared" here: resolves
120
+ * to `null`, never throws (api.probeJson's own contract — see its doc
121
+ * comment). `spec --json` must stay a pure, reliable, side-effect-free probe
122
+ * even against an unreachable or ancient server.
123
+ */
124
+ async function fetchExternalDocsUrl() {
125
+ const data = await api.probeJson('/api/spec.json', {
126
+ timeoutMs: EXTERNAL_DOCS_FETCH_TIMEOUT_MS,
127
+ });
128
+ const url = data?.externalDocs?.url;
129
+ return typeof url === 'string' && url ? url : null;
65
130
  }
66
131
  export const spec = new Command('spec')
67
132
  .description('Print a versioned, machine-readable description of the command tree')
68
133
  .option('--json', 'Output as JSON')
69
- .action((opts) => {
134
+ .action(async (opts) => {
70
135
  // `spec` is registered as a direct child of the program root
71
136
  // (index.ts's createProgram), so `.parent` IS that root by the time this
72
137
  // action ever runs — commander sets it in `addCommand()`. Falling back
73
138
  // to `spec` itself only matters for an isolated unit invocation (e.g. a
74
139
  // test driving `spec.parseAsync()` directly, unattached to a program).
75
140
  const root = spec.parent ?? spec;
76
- const cliSpec = buildSpec(root);
141
+ const apiBaseUrl = getApiUrl();
142
+ // #185 — rung 1 (the server fetch) only runs under --json: the
143
+ // plain-text branch below never reads docsUrl at all, so paying a
144
+ // network round-trip for it would be pure waste. Rungs 2/3 (sync,
145
+ // no network) still apply either way via the fallback below.
146
+ const externalDocsUrl = opts.json ? await fetchExternalDocsUrl() : null;
147
+ const docs = resolveDocsUrls({ apiBaseUrl, externalDocsUrl });
148
+ const cliSpec = buildSpec(root, docs);
77
149
  if (opts.json) {
78
150
  json(cliSpec);
79
151
  return;
package/dist/index.d.ts CHANGED
@@ -32,6 +32,15 @@ export declare function resolveCommandName(actionCommand: Command | undefined):
32
32
  */
33
33
  export declare function collectCommandNames(root: Command): string[];
34
34
  export declare function createProgram(): Command;
35
+ /**
36
+ * #185 — see the `addHelpText('afterAll', …)` call above for why this is
37
+ * unconditional (no TTY/--json gate) and why a thrown error must never
38
+ * escape: commander calls this synchronously while already writing help
39
+ * output, so a throw here would crash a `--help` invocation, the one code
40
+ * path this whole feature is not allowed to touch (see docs.ts's own
41
+ * `docsFooterLine` for the "never a guessed URL" contract this composes).
42
+ */
43
+ export declare function docsFooterText(): string;
35
44
  /**
36
45
  * True when this module is the process entrypoint (not merely imported by a test).
37
46
  *
@@ -96,6 +105,14 @@ export declare function isBareInvocation(argv: string[]): boolean;
96
105
  * overstates its own invariant is worse than no comment, because the next
97
106
  * reader trusts it.
98
107
  *
108
+ * #185 — `spec --json` (not plain `spec`) now DOES make its own bounded
109
+ * (2s), swallowed-on-failure GET for `externalDocs.url` (see spec.ts's
110
+ * `fetchExternalDocsUrl`) — this is not a regression of the "no mutation"
111
+ * invariant above (a GET mutates nothing, and any failure — including no
112
+ * egress at all — degrades to the sync host-derivation fallback, never a
113
+ * thrown error), just a second, narrower kind of side effect this guard was
114
+ * never meant to suppress in the first place.
115
+ *
99
116
  * Before this, `trawl spec --json` silently re-synced
100
117
  * `~/.claude/skills` (and `./.claude/skills`) and printed `trawl: re-synced
101
118
  * skill …` lines to stderr BEFORE the JSON payload — an agent's very first