@haystackeditor/cli 0.17.1 → 0.19.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -80,41 +80,87 @@ haystack triage <ref> --json # poll findings later / after --no-wait
80
80
  `dismiss`, `mark_reviewed`, `undismiss`, `request_review`, `trigger_review`,
81
81
  and `schema`.
82
82
 
83
- **`haystack verify`** verifies the current change using the authenticated
84
- Haystack fleet. Inside a git checkout, push your commit and run:
83
+ **`haystack verify`** shows the product blast radius of your current change:
84
+ where it shows up in the running app, and what broke. Inside a git checkout
85
+ (nothing needs to be committed or pushed), run:
85
86
 
86
87
  ```bash
87
88
  haystack verify
89
+ haystack verify --no-wait
88
90
  haystack verify --json
89
91
  ```
90
92
 
91
- The repository defaults to `owner/repo` from the `origin` remote URL. The head
92
- is the current `HEAD` commit. The base is `git merge-base HEAD
93
- origin/<default branch>`, using `origin/HEAD` to discover the default branch
94
- and falling back to `origin/main`. If the current commit is not reachable from
95
- any remote-tracking branch, the command asks you to push first.
96
- Explicit `--repo`, `--base`, and `--head` values override their defaults; supplying
97
- all three also works outside a checkout:
98
-
99
- ```bash
100
- haystack verify --repo owner/repo --base main --head feature --json
101
- ```
102
-
103
- Intent is optional: `--intent-file <path>` accepts JSON with exactly `problem`,
104
- `goal`, and `intended_outcomes`. The CLI uses your saved account and waits up to
105
- 35 minutes by default. Pass `--no-wait` to return after the fleet run is queued.
106
- Repeating the same repository, refs, and intent reuses the same run; choose a new
107
- `--idempotency-key` when a fresh execution is intentional. A wait that expires
108
- prints `timed_out: true` and `next_command`, then exits 2. Resume with:
93
+ It captures the checkout, committed and uncommitted changes together, exactly
94
+ as the stop hook does: the repository from `origin`, the base at `git
95
+ merge-base HEAD origin/<default branch>`, and the working tree as a
96
+ deterministic snapshot commit, so the same tree always names the same capture.
97
+ It then finds your crawl of that exact capture: the same code and the same
98
+ change title (HEAD's subject, which tells the crawl's judge what the change is;
99
+ committing or rewording keeps the code but changes the title). When there is
100
+ none, or the one it finds was cancelled, stopped before finishing or is being
101
+ cancelled, it submits the capture through the same request `haystack verify
102
+ precompute` sends (the service starts the crawl, or runs the stopped one again
103
+ once it has shut down), says so in one line, and follows the crawl the service
104
+ acknowledged; it never shows a crawl of another revision or another title.
105
+
106
+ A crawl builds the app with and without your change, starts from the changed
107
+ code, reaches it in the running app, explores outward with both builds side by
108
+ side, and double-checks and judges every difference. `haystack verify` waits
109
+ for the crawl, printing each step as it happens and each bug the moment the
110
+ crawl finds it (each read is held by the service until the crawl changes, so
111
+ nothing waits on a polling interval). It finishes when the crawl answers, as
112
+ soon as its time is up, without waiting for its machines to shut down; there
113
+ is no time limit, and Ctrl-C stops waiting, never the crawl. It then prints:
114
+
115
+ - a headline: how many bugs it found, none, or why the crawl could not finish;
116
+ - **Where your change shows up in the app**: every changed spot as
117
+ `file:line`, how far the crawl got with it (never reached, loaded but never
118
+ ran, shown on screen, ran, ran when its control was pressed, a difference
119
+ seen, a difference seen and confirmed, judged a bug), and the steps and page
120
+ that reached it;
121
+ - **What the crawl found**: bugs first, each with a one-sentence reason and the
122
+ steps to see it;
123
+ - **What the crawl never ran**: changed files no path ran, and how many of the
124
+ functions your change affects ran.
125
+
126
+ `--no-wait` prints the crawl's current state and returns. `--account <login>`
127
+ picks a saved account, and `--repo owner/repo` names the repository when
128
+ `origin` does not. `--json` prints one document, `{ "schema_version", "crawl" }`,
129
+ where `crawl` is the crawl as the service returns it, with its `answer` (what it
130
+ found so far, or its answer once its time is up) while no sealed `manifest`
131
+ exists (or `null` when no crawl of the capture could be started);
132
+ `haystack schema verify` prints its schema. Exit codes follow
133
+ `haystack case-batch status`: 0 when the crawl answered or finished (the bugs
134
+ it found are in the output) or is still running under `--no-wait`; 2 when it ended
135
+ without finishing (stopped early, or cancelled because a newer stop in the
136
+ repository replaced it) or the machines it used could not be proven shut down;
137
+ 1 when the command failed or no crawl could be started.
138
+
139
+ **The stop hook.** `haystack hooks install-session --cli claude` installs a
140
+ Claude Code Stop hook that runs `haystack verify precompute --hook` whenever an
141
+ agent turn stops. It captures the checkout within five seconds and hands the
142
+ request to a detached sender; the service then starts the change's analysis
143
+ and, for a repository set up for crawling, its crawl in the background, so
144
+ `haystack verify` usually finds the crawl already running or finished. A newer stop in the same repository replaces an
145
+ older crawl that has not finished. `haystack verify precompute` (without
146
+ `--hook`) submits the same capture in the foreground and prints one line for
147
+ the crawl it started.
148
+
149
+ The hosted fleet run of exact pushed commits stays available:
109
150
 
110
151
  ```bash
152
+ haystack verify hosted start owner/repo --base <sha> --head <sha> --json
111
153
  haystack verify hosted status cv_<48-lowercase-hex-characters> --wait --json
112
154
  haystack verify history owner/repo --limit 20
113
155
  ```
114
156
 
115
- Use the returned account-bound `next_command` verbatim on machines with multiple
116
- saved accounts. `haystack verify hosted start owner/repo --base <sha> --head
117
- <sha>` starts the same fleet execution through an explicit subcommand.
157
+ Intent is optional there: `--intent-file <path>` accepts JSON with exactly
158
+ `problem`, `goal`, and `intended_outcomes`. It waits up to 35 minutes by
159
+ default; pass `--no-wait` to return after the fleet run is queued. Repeating the
160
+ same repository, commits, and intent reuses the same run; choose a new
161
+ `--idempotency-key` when a fresh execution is intentional. A wait that expires
162
+ prints `timed_out: true` and `next_command`, then exits 2. Use the returned
163
+ account-bound `next_command` verbatim on machines with multiple saved accounts.
118
164
 
119
165
  Retained fleet cases remain available through their exact case IDs:
120
166
 
@@ -462,7 +508,8 @@ haystack hooks install --force # Overwrite existing hooks
462
508
  # Status
463
509
  haystack hooks status # Check installation status
464
510
 
465
- # Session hooks (triage on CLI start)
511
+ # Session hooks (triage on CLI start; Claude Code's Stop hook starts a crawl,
512
+ # see `haystack verify`)
466
513
  haystack hooks install-session # Auto-detect CLIs
467
514
  haystack hooks install-session --cli claude # Claude Code only
468
515
  haystack hooks install-session --cli all # All detected CLIs
@@ -520,6 +567,145 @@ behind an `UNKNOWN` that is not an arrival or wall outcome, and behind an
520
567
  is `null` when the outcome has none. A case in `--cases` may carry `repeatIndex` (1-16), which
521
568
  must be part of its `contentDigest`, to run as a repeat next to its original.
522
569
 
570
+ ### `haystack db profile`
571
+
572
+ Describe your production Postgres database so Haystack can build a full-size
573
+ stand-in for testing, without any personal data leaving your environment. You
574
+ run the profiler yourself, inside your own network, against your own database;
575
+ it writes a JSON file you read before you send it.
576
+
577
+ ```bash
578
+ # The connection string is read from the variable you name, never from a flag.
579
+ export DATABASE_URL='postgres://profiler_readonly@db.internal:5432/app'
580
+ haystack db profile --url-env DATABASE_URL --owner-table public.users --out profile.json
581
+
582
+ # Read exactly what would be sent, in plain English.
583
+ haystack db profile show profile.json
584
+
585
+ # Send it for one repository.
586
+ haystack db profile upload profile.json --repo acme/app
587
+ ```
588
+
589
+ **Which database it is used for.** Uploaded without `--dependency`, the profile
590
+ belongs to the repository, and Haystack builds the stand-in from it for the
591
+ repository's one Postgres database, whatever id onboarding gave it. When the
592
+ repository has several (or none), `haystack verify hosted start` says the
593
+ stand-in was not used, and why, naming the ids; upload again with `--dependency <id>` to pick
594
+ one. A profile uploaded with `--dependency <id>` is used for that database; when
595
+ both apply, the newer upload is used. `haystack verify hosted start` prints one
596
+ line saying which profile the stand-in was built from, or why it was not used, for every
597
+ repository that has uploaded a profile.
598
+
599
+ **Connection strings.** Any `postgres://` or `postgresql://` URL that
600
+ node-postgres accepts works, including a Unix socket
601
+ (`postgresql://profiler_readonly@/app?host=/var/run/postgresql`).
602
+
603
+ **What it does to your database.** It only reads. Every query is a single
604
+ `SELECT` in its own `READ ONLY` transaction (the session default is read only
605
+ too, and the transaction's read-only state is checked before each query), with
606
+ a statement timeout and a 5-second lock timeout so it never queues behind a
607
+ migration. Each transaction is rolled back. All of them read one `REPEATABLE
608
+ READ` snapshot, so every number describes the same moment even while your
609
+ application writes; one extra connection holds that snapshot open for the run
610
+ (on a primary, rows deleted during the run are cleaned up by vacuum only after
611
+ it ends). A role with `SELECT` on the
612
+ application tables is enough, and is what we recommend; tables with row-level
613
+ security active for that role, or that the role cannot read, stop the run with
614
+ their names. It works on Postgres 12 and later (tested on 16).
615
+
616
+ **Run it on a read replica.** It works on a streaming-replication hot standby
617
+ (tested on Postgres 16): nothing it runs writes, takes a write lock or creates a
618
+ temporary table, and each transaction names its isolation level (`REPEATABLE
619
+ READ`), so a replica whose default is `serializable` still answers. A
620
+ replica does not count dead rows, so a profile taken there marks each table's
621
+ dead rows as unavailable and says why; everything else is the same as on the
622
+ primary. A replica may cancel a long query when replaying the primary's changes
623
+ cannot wait for it (`canceling statement due to conflict with recovery`), and
624
+ the run's snapshot is held for its whole length; the run then stops and names
625
+ the fix: turn on `hot_standby_feedback` on the replica, or raise its
626
+ `max_standby_streaming_delay` above `--statement-timeout`.
627
+
628
+ **What leaves your database.** The database computes every count and pattern;
629
+ no row ever reaches the profiler. A column's actual values are written only when
630
+ all three hold:
631
+
632
+ 1. the column is a category: at most `--category-max-distinct` distinct values,
633
+ counted exactly over the whole table even when the table is otherwise
634
+ sampled (never estimated, so a column with many rare values cannot pass);
635
+ 2. the column is not personal. Personal means names, emails, phones, addresses,
636
+ free text (notes, comments, descriptions, messages), secrets, tokens and
637
+ passwords, IP and network addresses, dates of birth and government IDs. It is
638
+ decided from the column's name, from its type (`inet`, `cidr`, `macaddr`), and
639
+ from value shapes the database measures (the share of values shaped like an
640
+ email, phone number, IP address, secret, SSN, card number or "First Last"
641
+ name, and the number of words per value). When in doubt a column is personal,
642
+ and the profile records why;
643
+ 3. each value is shared by at least `--category-min-owners` owners. Owners are the
644
+ distinct rows of `--owner-table` (your users or organizations) reached by
645
+ following foreign keys from the column's table, and only owners that exist
646
+ there count (a key left dangling under a `NOT VALID` foreign key owns
647
+ nothing). A table with no foreign-key path to the owner table keeps no
648
+ values, shapes, JSON keys, percentiles or shares of true. Without
649
+ `--owner-table`, each row counts as an owner and the profile says so; a value
650
+ one heavy user repeats in thousands of rows then passes, so name an owner
651
+ table whenever you have one.
652
+
653
+ Values your schema declares (enum labels, `CHECK (col IN (...))` lists) need no
654
+ data and are kept with their share of rows. Everything else leaves only as
655
+ numbers: row counts (and whether each was counted exactly or is the planner's
656
+ estimate), how each table was read (in full, or the sample percentage and
657
+ seed), table/index/TOAST sizes, dead-row share (or why it is unavailable), null
658
+ share, distinct count, average width, the share of `true` in each boolean
659
+ column, and the 5th/25th/50th/75th/95th percentiles of lengths and numbers
660
+ (never the minimum or maximum; percentiles and the share of `true` are withheld
661
+ for a column with fewer than `--category-min-owners` owners, and number
662
+ percentiles are withheld for personal columns, because a percentile of birth
663
+ dates is a birth date). Character shapes (`Aa Aa`, `+9 (9) 9-9`) and JSON key
664
+ paths leave only when shared by at least `--category-min-owners` owners, and a
665
+ JSON key path never leaves when any key on it looks like an email, phone number,
666
+ IP address, SSN, card number or secret (`15551234567.verified` stays home). Odd
667
+ rows (wrong type, oversize, bad encoding, rare nulls, rare shapes, missing JSON
668
+ keys) are counted and described without their values, even for a single row.
669
+ Index definitions are sent with every literal replaced by `?`, since a partial
670
+ index can name a value (`WHERE email <> 'ceo@example.com'` is sent as `WHERE
671
+ (email <> ?::text)`); the column names, operators and types stay. Settings are
672
+ an allowlist of planner and storage settings; connection, authentication, SSL
673
+ and file settings are never read.
674
+
675
+ `show` prints every string the file carries. `upload` refuses a file with any
676
+ field `show` does not print, a file whose percentiles include anything but
677
+ p5 to p95, and a file with a literal left in an index definition. Haystack's
678
+ server checks the file again and also refuses one made with a looser rule than
679
+ at most 50 distinct values and at least 50 owners, a kept value its row share
680
+ shows fewer than `--category-min-owners` rows hold, and any kept value shaped
681
+ like an email, phone number, IP address, SSN, card number, token or password
682
+ hash, whatever its column's type, or that does not fit its column's declared
683
+ type (an integer column keeps only integers, a boolean column `true` or
684
+ `false`, a date or time column dates or times; a column of any other type than
685
+ text, numbers, booleans, dates and times keeps no data values, and enum labels
686
+ leave as schema values). The profiler never writes such a value itself:
687
+ however many owners share it (a date like `2031-07-19` has a phone number's
688
+ shape, and so does any number of seven or more digits), it is withheld like a
689
+ value too few owners share. An accepted file is stored once, under its
690
+ SHA-256, in Haystack's results storage.
691
+
692
+ Knobs and what turning them does:
693
+
694
+ | Flag | Default | Turning it up | Turning it down |
695
+ |------|---------|---------------|-----------------|
696
+ | `--owner-table <schema.table>` | none (rows are owners) | n/a | n/a; naming one makes the owner threshold count people or organizations instead of rows, which is stricter and what we recommend. |
697
+ | `--schema <name>` (repeatable) | every non-system schema | n/a | Profiles fewer schemas; foreign-key paths to the owner table only run through profiled tables. |
698
+ | `--category-max-distinct <n>` | 50 (1 to 1000) | More columns count as categories and can keep values (each still needs enough owners); `upload` accepts at most 50. | Fewer columns keep values; more leave only as numbers and shapes. |
699
+ | `--category-min-owners <n>` | 50 (5 to 1,000,000) | Stricter: fewer values, shapes, JSON keys and percentiles pass. | More rare values pass; each one is shared by fewer people. Not below 5; `upload` requires at least 50. |
700
+ | `--scan-max-mb <mb>` | 256 | Bigger tables are read in full: exact counts and more rare-but-common-enough values, at the cost of more reading on your database. | Tables above this are read through `TABLESAMPLE SYSTEM` of about this size (same pages for every query). Owner counts in a sample are lower than the truth, so fewer values pass (never more). Row counts of sampled tables come from the planner's statistics, so a sampled table must have been analyzed; the profile records the sample percentage and seed and marks the row count as an estimate. Distinct counts of a sampled table are estimated from the sample (a single-column primary or unique key counts one per row, and a validated single-column foreign key counts the parent rows it references), except a column that could be a category: its distinct values are counted exactly, which reads that table in full once for all such columns (the run says which tables). |
701
+ | `--concurrency <n>` | 2 (1 to 8) | Faster, more simultaneous load on your database. | Gentler; 1 runs one query at a time. Either way one more connection holds the run's snapshot. |
702
+ | `--statement-timeout <seconds>` | 120 | Allows longer queries on big tables. | Caps each query sooner; a query that hits it stops the run and names the table. |
703
+
704
+ Every failure is one line naming what failed. Nothing is written unless the
705
+ whole profile succeeds. On a primary, dead-row shares come from
706
+ `pg_stat_user_tables`; a table with rows but no counts there (statistics reset)
707
+ stops the run and asks for `ANALYZE` on that table.
708
+
523
709
  ### `haystack policy`
524
710
 
525
711
  Manage review policies (`.haystack/review-policy.md`):
@@ -162,8 +162,8 @@ export function integer(value, what, minimum, maximum) {
162
162
  }
163
163
  return value;
164
164
  }
165
- /** Bounded JSON, as the combination-search hook already bounds its inputs:
166
- * finite numbers, no prototype-polluting keys, depth 64, 8 MiB serialized. */
165
+ /** Bounded JSON: finite numbers, no prototype-polluting keys, depth 64, 8 MiB
166
+ * serialized. */
167
167
  export function boundedCaseBatchJson(value, what, maximumBytes = CASE_BATCH_MAX_INPUT_BYTES) {
168
168
  const visit = (item, depth) => {
169
169
  if (depth > 64)
@@ -19,19 +19,8 @@ import { resolveAuthContext } from '../utils/auth.js';
19
19
  import { classifyHttpError, HaystackApiError, haystackApiUrl } from '../utils/haystack-api.js';
20
20
  import { boundedCaseBatchJson, buildCaseBatchRequest, CASE_BATCH_MAX_CASES, CASE_BATCH_MAX_CONCURRENT_CASES, CASE_BATCH_MAX_INPUT_BYTES, CASE_BATCH_MAX_CASE_WALL_MS, CASE_BATCH_MAX_TOTAL_BUDGET_MS, CASE_BATCH_MIN_CASE_WALL_MS, CASE_BATCH_MIN_TOTAL_BUDGET_MS, CaseBatchRequestValidationError, CaseBatchResponseError, isJsonObject, isTerminalCaseBatchStatus, parseCaseBatchCommit, parseCaseBatchIdempotencyKey, parseCaseBatchLimits, parseCaseBatchRepository, parseCaseBatchSnapshot, parseCaseBatchSource, parseCaseBatchWorld, parseProductCases, } from './case-batch-contract.js';
21
21
  /** The CLI gateway mount. The same worker handlers are `/v1/case-batches` and
22
- * `/api/cloud-verifier/case-batches`; CLI callers use the agent gateway, as
23
- * they already do for `/api/agent/cloud-verifier/searches`.
24
- *
25
- * BLOCKED ON THE SERVER LANE. No handler serves this path in this repository
26
- * yet, and the auth worker does not proxy it: `isSearchProxyPath` in
27
- * `infra/auth-worker/index.js` covers `/api/agent/cloud-verifier/searches`
28
- * only, and `agent/cloudflare/src/combination-search-route.ts` has no
29
- * case-batch sibling. Until Lane A of `docs/case-batch-coordinator.md` lands
30
- * `POST/GET /v1/case-batches`, `GET/DELETE /v1/case-batches/:runId`,
31
- * `GET /v1/case-batches/:runId/bundle[/path]` and the matching auth-worker
32
- * proxy entry, every command in this file reaches a 404 in production. That
33
- * lane must merge first; this file is written against the frozen contract
34
- * ahead of it, which is what the implementation plan assigns to Lane D.
22
+ * `/api/cloud-verifier/case-batches` (`agent/cloudflare/src/case-batch-route.ts`);
23
+ * CLI callers use the agent gateway, which `infra/auth-worker/index.js` proxies.
35
24
  *
36
25
  * Three things the server side must hold for this file to work as written:
37
26
  * the 8 MiB request bound, `repository`-scoped reads with `limit` and
@@ -40,8 +29,8 @@ import { boundedCaseBatchJson, buildCaseBatchRequest, CASE_BATCH_MAX_CASES, CASE
40
29
  * is the one artifact with no parent digest to check it against). */
41
30
  const GATEWAY = '/api/agent/cloud-verifier/case-batches';
42
31
  const RUN_ID = /^cv_[0-9a-f]{48}$/;
43
- /** Page size for the bounded case pagination; the search read route caps
44
- * `limit` at 500 and the batch read mirrors it. */
32
+ /** Page size for the bounded case pagination; the batch read route caps
33
+ * `limit` at 500. */
45
34
  const CASE_PAGE_LIMIT = 500;
46
35
  /** A hard stop on pagination independent of what the server reports, so a
47
36
  * cursor that never terminates cannot spin the CLI forever. */
@@ -101,9 +90,12 @@ function requireRunId(runId) {
101
90
  }
102
91
  return runId;
103
92
  }
93
+ /** How long any gateway request may take unless its caller says otherwise. */
94
+ export const GATEWAY_TIMEOUT_MS = 120_000;
104
95
  /** Every outbound request. The header set is built here and nowhere else, so
105
- * the credential surface of this command is one line: the login bearer. */
106
- async function gatewayFetch(path, token, init = { method: 'GET' }) {
96
+ * the credential surface of this command is one line: the login bearer.
97
+ * `haystack verify` reads crawls through this same client. */
98
+ export async function gatewayFetch(path, token, init = { method: 'GET' }) {
107
99
  const headers = new Headers();
108
100
  headers.set('Authorization', `Bearer ${token}`);
109
101
  headers.set('Accept', init.accept ?? 'application/json');
@@ -114,7 +106,7 @@ async function gatewayFetch(path, token, init = { method: 'GET' }) {
114
106
  method: init.method,
115
107
  headers,
116
108
  ...(init.body === undefined ? {} : { body: init.body }),
117
- signal: AbortSignal.timeout(120_000),
109
+ signal: AbortSignal.timeout(init.timeoutMs ?? GATEWAY_TIMEOUT_MS),
118
110
  });
119
111
  }
120
112
  async function gatewayJson(path, token, init = { method: 'GET' }) {
@@ -0,0 +1,9 @@
1
+ export const CRAWL_TITLE_MAX_CHARS = 200;
2
+ /** Amendment 8: the time a crawl may be asked to take, whole seconds from 1 to 30 minutes. */
3
+ export const CRAWL_BUDGET_MIN_MS = 60_000;
4
+ export const CRAWL_BUDGET_MAX_MS = 30 * 60_000;
5
+ export const CRAWL_MAX_FINDINGS = 50;
6
+ export const CRAWL_MAX_FINDING_STEPS = 64;
7
+ /** Amendment 9: what a crawl has found so far (`exploring`) or its answer at the budget (`answered`), before the sealed manifest;
8
+ * findings carry no images. A read may pass `waitAfter=<updatedAt>` to be held until the crawl changes (at most CRAWL_WAIT_MAX_MS). */
9
+ export const CRAWL_WAIT_MAX_MS = 25_000;