@orkestrel/scaffold 0.0.66 → 0.0.68

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (74) hide show
  1. package/dist/bin/main.js +67 -44
  2. package/dist/bin/main.js.map +1 -1
  3. package/dist/host/agents/templates/brief.md +9 -0
  4. package/dist/host/claude/agents/orkestrel.md +8 -8
  5. package/dist/host/claude/rules/names.md +15 -0
  6. package/dist/host/claude/rules/tests.md +33 -4
  7. package/dist/host/claude/rules/workspace.md +14 -2
  8. package/dist/host/dotfiles/prettierignore +3 -0
  9. package/dist/host/guides/README.md +65 -0
  10. package/dist/host/guides/abort.md +169 -0
  11. package/dist/host/guides/agent.md +1509 -0
  12. package/dist/host/guides/brief.md +1266 -0
  13. package/dist/host/guides/browser.md +2200 -0
  14. package/dist/host/guides/budget.md +196 -0
  15. package/dist/host/guides/codec.md +519 -0
  16. package/dist/host/guides/console.md +785 -0
  17. package/dist/host/guides/contract.md +1193 -0
  18. package/dist/host/guides/csv.md +541 -0
  19. package/dist/host/guides/database.md +2518 -0
  20. package/dist/host/guides/emitter.md +233 -0
  21. package/dist/host/guides/form.md +1791 -0
  22. package/dist/host/guides/html.md +717 -0
  23. package/dist/host/guides/indexeddb.md +505 -0
  24. package/dist/host/guides/interpret.md +1029 -0
  25. package/dist/host/guides/lsp.md +515 -0
  26. package/dist/host/guides/markdown.md +964 -0
  27. package/dist/host/guides/mcp.md +5554 -0
  28. package/dist/host/guides/middleware.md +927 -0
  29. package/dist/host/guides/msg.md +440 -0
  30. package/dist/host/guides/ndjson.md +120 -0
  31. package/dist/host/guides/ollama.md +380 -0
  32. package/dist/host/guides/pool.md +280 -0
  33. package/dist/host/guides/probe.md +1210 -0
  34. package/dist/host/guides/process.md +1620 -0
  35. package/dist/host/guides/program.md +1110 -0
  36. package/dist/host/guides/qualifier.md +854 -0
  37. package/dist/host/guides/queue.md +370 -0
  38. package/dist/host/guides/rater.md +330 -0
  39. package/dist/host/guides/reason.md +1122 -0
  40. package/dist/host/guides/relation.md +373 -0
  41. package/dist/host/guides/router.md +753 -0
  42. package/dist/host/guides/scaffold.md +192 -31
  43. package/dist/host/guides/sea.md +383 -0
  44. package/dist/host/guides/server.md +752 -0
  45. package/dist/host/guides/sqlite.md +330 -0
  46. package/dist/host/guides/sse.md +187 -0
  47. package/dist/host/guides/supervisor.md +4890 -0
  48. package/dist/host/guides/table.md +1556 -0
  49. package/dist/host/guides/template.md +280 -0
  50. package/dist/host/guides/terminal.md +1145 -0
  51. package/dist/host/guides/test.md +2969 -0
  52. package/dist/host/guides/timeout.md +252 -0
  53. package/dist/host/guides/tool.md +311 -0
  54. package/dist/host/guides/toolbox.md +1038 -0
  55. package/dist/host/guides/websocket.md +282 -0
  56. package/dist/host/guides/worker.md +615 -0
  57. package/dist/host/guides/workflow.md +1507 -0
  58. package/dist/host/guides/workspace.md +595 -0
  59. package/dist/host/manifest.json +1218 -10
  60. package/dist/host/tests/policy.test.ts +279 -2
  61. package/dist/host/tests/setupPolicy.ts +437 -6
  62. package/dist/src/core/index.cjs +44 -22
  63. package/dist/src/core/index.cjs.map +1 -1
  64. package/dist/src/core/index.d.cts +33 -9
  65. package/dist/src/core/index.d.ts +33 -9
  66. package/dist/src/core/index.js +43 -23
  67. package/dist/src/core/index.js.map +1 -1
  68. package/dist/src/server/index.cjs +1750 -1567
  69. package/dist/src/server/index.cjs.map +1 -1
  70. package/dist/src/server/index.d.cts +106 -24
  71. package/dist/src/server/index.d.ts +106 -24
  72. package/dist/src/server/index.js +1751 -1570
  73. package/dist/src/server/index.js.map +1 -1
  74. package/package.json +9 -9
@@ -0,0 +1,541 @@
1
+ # CSV
2
+
3
+ > A types-first RFC 4180 CSV parser and renderer — a hand-written,
4
+ > single-pass tokenizer that turns CSV text into a typed `CSVTable`, and a
5
+ > stateful `CSV` workspace that wraps that table with query, rewrite,
6
+ > streaming, and export operations.
7
+
8
+ Parse once, then treat every read as a projection of the parsed table.
9
+ `parseCSV` runs a tokenizer phase — `readRecords`, a character scanner
10
+ honoring quoting, escaping, and CRLF, LF, and CR line endings — then a
11
+ table-building phase of header mapping, ragged-row handling, and optional
12
+ whole-column type inference, and returns a `CSVParseResult` pairing the table
13
+ with any `CSVError`s collected along the way. The renderer, `renderCSV`, is a
14
+ separate downstream projection from a table (or plain row list) back to text;
15
+ it never assumes its input came from `parseCSV`. Every row is built with a
16
+ null prototype, so a hostile header name (`__proto__`) can never reach
17
+ `Object.prototype`, and `renderCSV` sanitizes every field against CSV formula
18
+ injection by default. Parsing never throws on malformed data — a `CSVError`
19
+ (a machine-readable `code` plus `line` / `column` / `offset`) is collected
20
+ into the result's `errors` instead, unless `strict` is set, in which case the
21
+ first collected error throws immediately. An invalid option
22
+ (`INVALID_OPTION`) always throws — that is a programmer error, not a parse
23
+ malformation. Source: [`src/core`](../src/core). Surfaced through the
24
+ `@src/core` barrel.
25
+
26
+ ## Surface
27
+
28
+ Parse a document with `createCSV`, infer its column types, and read the rows
29
+ as typed records:
30
+
31
+ ```ts
32
+ import { createCSV } from '@orkestrel/csv'
33
+
34
+ const csv = createCSV('name,age\nAda,36\nGrace,85', { infer: true })
35
+ csv.rows // [{ name: 'Ada', age: 36 }, { name: 'Grace', age: 85 }]
36
+ ```
37
+
38
+ ### Types
39
+
40
+ The full parse/render/export shape, from [`types.ts`](../src/core/types.ts). An
41
+ interface's call-signature members are documented under
42
+ [`## Methods`](#methods).
43
+
44
+ A `Shape` cell holds an interface's data members as bare names in braces, `?`
45
+ marking an optional member and `plus` introducing its call-signature members,
46
+ and a type alias's own type literal with a union's arms escaped as `\|`.
47
+
48
+ | Type | Kind | Shape | Summary |
49
+ | ----------------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
50
+ | `Row` | type | `Record<string, unknown>` | Represents a CSV row — a plain record of column values keyed by column name. |
51
+ | `CSVTable` | interface | `{ columns, rows }` | Represents a parsed CSV table — the typed rows plus the column order they were parsed (or declared) in. |
52
+ | `RawField` | interface | `{ value, quoted }` | Represents one raw parsed field — the value exactly as it appeared in a record, before type inference or column mapping, plus whether it was quoted in the source. |
53
+ | `Position` | interface | `{ offset, line, column }` | Represents a cursor position in a parsed source text — relative to the input after byte-order-mark removal. |
54
+ | `RawRecord` | interface | `{ fields, start }` | Represents one raw parsed record — its ordered `RawField`s plus where the record begins in the source, before header mapping. |
55
+ | `FieldScan` | interface | `{ field, next, errors }` | Represents one scanned field — a single `RawField` the tokenizer produced, the `Position` immediately after it, and any malformations found while scanning it. |
56
+ | `RecordScan` | interface | `{ record, next, errors }` | Represents one scanned record — a single `RawRecord` the tokenizer produced, the `Position` immediately after it, and any malformations found while scanning it. |
57
+ | `HeaderResult` | interface | `{ columns, body, errors }` | Represents the result of resolving a header record — the disambiguated column names, the remaining body records, and any header-related errors. |
58
+ | `RowResult` | interface | `{ row?, error? }` | Represents the result of building one `RawRecord` into a typed `Row` — the row, the error that excluded it, or both when `ParseOptions.ragged` is `'collect'`. |
59
+ | `RecordsResult` | interface | `{ records, errors }` | Represents the result of the record-splitting phase — every `RawRecord` the tokenizer produced plus any `CSVError`s collected along the way. |
60
+ | `CSVParseResult` | interface | `{ table, errors }` | Represents the result of a full parse — the assembled `CSVTable` plus any `CSVError`s collected along the way. |
61
+ | `EscapeStyle` | type | `'double' \| 'backslash'` | Names how an embedded quote character is escaped inside a quoted field. |
62
+ | `QuoteStyle` | type | `'minimal' \| 'always' \| 'nonnumeric'` | Names the renderer's quoting policy — which fields get wrapped in quotes. |
63
+ | `RaggedPolicy` | type | `'collect' \| 'pad' \| 'error'` | Names how the parser treats a record whose field count does not match the header. |
64
+ | `ColumnType` | type | `'text' \| 'integer' \| 'real' \| 'boolean' \| 'json' \| 'blob'` | Names a portable storage type for a column — the same literal set `@orkestrel/database` declares as `ColumnStorage` (never imported), so a CSV column map and a database table schema stay drop-in interchangeable. |
65
+ | `Columns` | type | `Readonly<Record<string, ContractShape>>` | Represents a CSV's declared columns — a map of column name to its value `ContractShape`. |
66
+ | `ParseOptions` | interface | `{ delimiter?, quote?, escape?, header?, comment?, blanks?, trim?, ragged?, infer?, limit?, strict? }` | Represents the options for parsing CSV text into a `CSVTable`. |
67
+ | `ResolvedParseOptions` | type | `Required<Omit<ParseOptions, 'comment'>> & Pick<ParseOptions, 'comment'>` | Represents the fully-resolved parse configuration every tokenizer and table-building helper takes — `ParseOptions` with every member defaulted except `comment`, which has no default and stays optional. |
68
+ | `RenderOptions` | interface | `{ delimiter?, quote?, escape?, newline?, header?, columns?, quotes?, blank?, sanitize?, bom? }` | Represents the options for rendering a `CSVTable` (or row list) back to CSV text. |
69
+ | `ResolvedRenderOptions` | type | `Required<Omit<RenderOptions, 'columns'>> & Pick<RenderOptions, 'columns'>` | Represents the fully-resolved render configuration every quoting and rendering helper takes — `RenderOptions` with every member defaulted except `columns`, which has no default and stays optional. |
70
+ | `ExportOptions` | interface | `{ key?, columns? }` | Represents the options for `CSVInterface.export`. |
71
+ | `TableExport` | interface | `{ key, columns, schema }` | Represents a CSV's portable definition, produced by `CSVInterface.export` — the unit of schema exchange across environments. |
72
+ | `CSVErrorCode` | type | `'UNTERMINATED_QUOTE' \| 'BAD_QUOTE' \| 'RAGGED_ROW' \| 'DUPLICATE_HEADER' \| 'EMPTY_HEADER' \| 'LIMIT_EXCEEDED' \| 'INVALID_OPTION'` | Names a machine-readable `CSVError` code. |
73
+ | `CSVInterface` | interface | `{ table, rows, errors } plus find, filter, map, reduce, stream, toJSON, export` | Represents a parsed, queryable CSV document — the typed `CSVTable` plus the query, rewrite, and export operations over it. |
74
+
75
+ ### Errors
76
+
77
+ From [`errors.ts`](../src/core/errors.ts). An invalid option or programmer
78
+ error always throws a `CSVError`; a parse-time malformation is collected
79
+ into a result's `errors` unless `strict` is set.
80
+
81
+ | Error | Kind | Signature | Summary |
82
+ | ------------ | -------- | --------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |
83
+ | `CSVError` | class | `extends Error` | Represents an error surfaced by the CSV layer — either thrown for a programmer error / `strict`-mode parse failure, or collected into a result's `errors` list. |
84
+ | `isCSVError` | function | `(value: unknown) => value is CSVError` | Narrows an unknown caught value to a `CSVError`. |
85
+
86
+ ### Constants
87
+
88
+ Centralized, frozen data the parser/renderer draw their defaults and
89
+ canonical patterns from, from [`constants.ts`](../src/core/constants.ts).
90
+
91
+ A `Shape` cell holds the constant's declared type.
92
+
93
+ | Constant | Kind | Shape | Summary |
94
+ | -------------------------- | ----- | ------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
95
+ | `BOM` | const | `string` | Names the UTF-8 byte-order-mark character, `'\uFEFF'`, prepended when `RenderOptions.bom` is `true`. |
96
+ | `DEFAULT_PARSE_OPTIONS` | const | `Required<Omit<ParseOptions, 'comment'>>` | Holds the resolved default `ParseOptions` (everything but `comment`, which has no default) — what `parseCSV` uses for any option left unspecified: `{ delimiter: ',', quote: '"', escape: 'double', header: true, blanks: true, trim: false, ragged: 'collect', infer: false, limit: 0, strict: false }`. |
97
+ | `DEFAULT_RENDER_OPTIONS` | const | `Required<Omit<RenderOptions, 'columns'>>` | Holds the resolved default `RenderOptions` (everything but `columns`, which has no default) — what `renderCSV` uses for any option left unspecified: `{ delimiter: ',', quote: '"', escape: 'double', newline: '\r\n', header: true, quotes: 'minimal', blank: '', sanitize: true, bom: false }`. |
98
+ | `SANITIZE_PREFIXES` | const | `ReadonlySet<string>` | Lists the leading characters the OWASP CSV-injection guard treats as formula-triggering — a field starting with any of these is prefixed with a protective `'` when `RenderOptions.sanitize` is `true`. |
99
+ | `POSITIONAL_COLUMN_PREFIX` | const | `string` | Names the prefix used for positional columns (`column1`, `column2`, …) when `ParseOptions.header` is `false`, or a header field is empty — 1-based, `'column'`. |
100
+ | `SANITIZE_ESCAPE` | const | `string` | Names the protective prefix `sanitizeField` prepends to a field starting with a formula-triggering character (the OWASP CSV-injection guidance), `"'"`. |
101
+ | `SUFFIX_SEPARATOR` | const | `string` | Names the separator between a disambiguated column name and its collision counter (`name` → `name_2`, `name_3`, …) — see `uniqueName`, `'_'`. |
102
+ | `INTEGER_PATTERN` | const | `RegExp` | Matches a canonical integer only — an optional leading `-`, no leading zeros (except the bare digit `0`), digits only. No `+` sign, no whitespace. |
103
+ | `REAL_PATTERN` | const | `RegExp` | Matches a canonical decimal only — an optional leading `-`, an integer part with no leading zeros (except the bare digit `0`), an optional `.` followed by at least one digit. No scientific notation, no `NaN` / `Infinity`, no decimal comma, no trailing dot. |
104
+ | `NUMERIC_PATTERN` | const | `RegExp` | Matches what the renderer treats as a plain number for the `'nonnumeric'` `QuoteStyle` and the sanitize `+` / `-` exemption — like `REAL_PATTERN` but also allowing a leading `+`. |
105
+ | `BOOLEAN_TRUE` | const | `string` | Names the canonical serialized form of the boolean `true` — the string `'true'`. |
106
+ | `BOOLEAN_FALSE` | const | `string` | Names the canonical serialized form of the boolean `false` — the string `'false'`. |
107
+ | `MAX_ERRORS` | const | `number` | Sets the maximum number of `CSVError`s collected into a parse result, `100` — once reached, error collection stops (earlier records already parsed are kept, later malformations are silently no longer recorded). |
108
+
109
+ ### Helpers
110
+
111
+ Pure, total, zero-dependency leaves from
112
+ [`helpers.ts`](../src/core/helpers.ts) — the option resolvers, the
113
+ hand-written tokenizer and table builders `parsers.ts` composes, and the
114
+ rendering projections callers reach for directly. Every function is
115
+ unit-testable in isolation.
116
+
117
+ | Helper | Kind | Signature | Summary |
118
+ | ----------------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
119
+ | `assertValidSeparators` | function | `(delimiter: string, quote: string) => void` | Validates a delimiter / quote pair shared by both `resolveParseOptions` and `resolveRenderOptions` — each must be exactly one character, they must differ, and neither may be CR, LF, or the BOM character. |
120
+ | `resolveParseOptions` | function | `(options?: ParseOptions) => ResolvedParseOptions` | Merges `options` over `DEFAULT_PARSE_OPTIONS` into a fully-resolved parse configuration. |
121
+ | `resolveRenderOptions` | function | `(options?: RenderOptions) => ResolvedRenderOptions` | Merges `options` over `DEFAULT_RENDER_OPTIONS` into a fully-resolved render configuration. |
122
+ | `uniqueName` | function | `(name: string, taken: ReadonlySet<string>) => string` | Disambiguates a single column name against the names already taken — the collision leaf `uniqueColumns` composes over an entire header. |
123
+ | `uniqueColumns` | function | `(names: readonly string[]) => readonly string[]` | Disambiguates a header's column names deterministically — an empty (or whitespace-only) name becomes positional, and a name that repeats an earlier kept name is suffixed `_2`, `_3`, … until unique. |
124
+ | `sanitizeField` | function | `(field: string) => string` | Guards a field against CSV/spreadsheet formula injection (the OWASP CSV-injection guidance) — a field starting with a formula-triggering character is prefixed with a protective `SANITIZE_ESCAPE`. |
125
+ | `serializeCell` | function | `(value: unknown, blank: string) => string` | Serializes one cell value to its rendered text — the renderer's stringify leaf, applied before sanitize/quote. |
126
+ | `deriveColumns` | function | `(rows: readonly Row[]) => readonly string[]` | Derives a column order from a plain row list — the first-seen union of every row's keys, in encounter order. |
127
+ | `needsQuote` | function | `(field: string, options: ResolvedRenderOptions) => boolean` | Checks `field` against the correctness floor every `QuoteStyle` policy respects — a field containing the delimiter, the quote character, CR, or LF must ALWAYS be quoted regardless of policy. |
128
+ | `wrapQuoted` | function | `(field: string, options: ResolvedRenderOptions) => string` | Wraps `field` in quotes, escaping per `options.escape` — the shared quote-and-escape step every quoting policy applies once it decides `field` needs quoting; it IS the `'always'` `QuoteStyle` as well (every field quoted unconditionally). |
129
+ | `quoteMinimal` | function | `(field: string, options: ResolvedRenderOptions) => string` | Implements the `'minimal'` `QuoteStyle` — quotes a field only when `needsQuote` requires it. |
130
+ | `quoteNonnumeric` | function | `(field: string, options: ResolvedRenderOptions) => string` | Implements the `'nonnumeric'` `QuoteStyle` — quotes every field whose value is not a plain number (or that `needsQuote` requires regardless). |
131
+ | `renderRecord` | function | `(row: Row, columns: readonly string[], options: ResolvedRenderOptions, quote: (field: string, options: ResolvedRenderOptions) => string) => string` | Renders one row to one delimited line — serializes every column's cell, optionally sanitizes it, then applies the given quoting policy. |
132
+ | `quoteStyleToPolicy` | function | `(quotes: ResolvedRenderOptions['quotes']) => (field: string, options: ResolvedRenderOptions) => string` | Selects the quoting-policy function for a resolved `options.quotes`. |
133
+ | `isRowList` | function | `(source: CSVTable \| readonly Row[]) => source is readonly Row[]` | Narrows a `CSVTable \| readonly Row[]` union to its row-list member. |
134
+ | `renderCSV` | function | `(input: CSVTable \| readonly Row[], options?: RenderOptions) => string` | Renders a `CSVTable` (or a plain row list) to CSV text. |
135
+ | `advancePosition` | function | `(position: Position, count?: number) => Position` | Advances a `Position` by `count` NON-line-break characters. |
136
+ | `isBreakChar` | function | `(char: string) => boolean` | Checks whether `char` starts a record separator (CR or LF). |
137
+ | `scanBreak` | function | `(source: string, position: Position) => Position \| undefined` | Consumes exactly one line break (CRLF, bare LF, or bare CR) at `position` — a CRLF pair counts as ONE break. |
138
+ | `scanComment` | function | `(source: string, position: Position, options: ResolvedParseOptions) => Position \| undefined` | Consumes a comment line at `position`, when `options.comment` names one starting there — through the end of that line INCLUDING its break (or end-of-input). |
139
+ | `scanUnquoted` | function | `(source: string, position: Position, options: ResolvedParseOptions) => FieldScan` | Scans one unquoted field starting at `position` — content runs until the delimiter, a line break, or end-of-input. |
140
+ | `scanQuoted` | function | `(source: string, position: Position, options: ResolvedParseOptions) => FieldScan` | Scans one quoted field starting at `position` — `position` must be AT the opening quote character. |
141
+ | `scanField` | function | `(source: string, position: Position, options: ResolvedParseOptions) => FieldScan` | Scans one field at `position` — dispatches to `scanQuoted` when the character there is `options.quote`, else `scanUnquoted`. |
142
+ | `scanRecord` | function | `(source: string, position: Position, options: ResolvedParseOptions) => RecordScan` | Scans one full record at `position` — fields separated by `options.delimiter`, ending at a break (consumed through `scanBreak`) or end-of-input. |
143
+ | `readRecords` | function | `(input: string, options?: ParseOptions) => RecordsResult` | Splits `input` into raw, un-mapped `RawRecord`s — the tokenizer phase beneath `parseCSV`. |
144
+ | `deriveHeader` | function | `(records: readonly RawRecord[], options: ResolvedParseOptions) => HeaderResult` | Resolves a table's header from its raw records — disambiguates the first record's names when `options.header` is `true`, or generates positional names sized to the widest record otherwise. |
145
+ | `buildRow` | function | `(record: RawRecord, columns: readonly string[], options: ResolvedParseOptions) => RowResult` | Builds one `RawRecord` into one null-prototype `Row`, padding or truncating to `columns.length` per `options.ragged`. |
146
+
147
+ ### Inferers
148
+
149
+ Whole-column type inference from
150
+ [`inferers.ts`](../src/core/inferers.ts) — reads the raw cell text a column
151
+ holds, rules which `ColumnType` it carries, and applies that ruling cell by
152
+ cell.
153
+
154
+ | Inferer | Kind | Signature | Summary |
155
+ | ----------------- | -------- | ---------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
156
+ | `inferColumnType` | function | `(values: readonly string[]) => ColumnType` | Infers a whole column's `ColumnType` conservatively from its raw string values — never `'json'` or `'blob'` (those require an explicit `Columns` declaration). Empty-string cells are ignored entirely (they neither confirm nor demote a type); a column with no non-empty cells is `'text'`. |
157
+ | `coerceInferred` | function | `(value: string, type: ColumnType) => unknown` | Coerces one string cell to `type`'s typed representation — the exhaustive per-cell dispatch `inferRows` applies once a column's type is known. |
158
+ | `inferRows` | function | `(rows: readonly Row[], columns: readonly string[]) => readonly Row[]` | Applies whole-column type inference to a built row set — per column, infers its `ColumnType` from its string cells, then coerces every cell of that type through `coerceInferred`. |
159
+
160
+ ### Parsers
161
+
162
+ The coercers, from [`parsers.ts`](../src/core/parsers.ts) — the `parseCSV`
163
+ entry point, which composes the `helpers.ts` tokenizer, the `helpers.ts` table
164
+ builders, and the `inferers.ts` column inference into a `CSVParseResult`, plus
165
+ the flat cell coercers that inference dispatches to.
166
+
167
+ | Parser | Kind | Signature | Summary |
168
+ | -------------- | -------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
169
+ | `parseCSV` | function | `(input: string, options?: ParseOptions) => CSVParseResult` | Parses `input` into a typed `CSVParseResult` — header mapping, ragged-row handling, and optional type inference on top of `readRecords`, `deriveHeader`, `buildRow`, and `inferRows`. |
170
+ | `parseInteger` | function | `(value: string) => number \| undefined` | Parses a raw cell string into a canonical integer — `undefined` for anything else (leading zeros, decimals, out-of-safe-range magnitude, non-numeric text). |
171
+ | `parseReal` | function | `(value: string) => number \| undefined` | Parses a raw cell string into a canonical decimal (or integer) — `undefined` for anything else. |
172
+ | `parseBoolean` | function | `(value: string) => boolean \| undefined` | Parses a raw cell string into a strict boolean — `undefined` for anything other than the exact canonical forms. |
173
+
174
+ ### Shapers
175
+
176
+ Declarative `ContractShape` values (from `@orkestrel/contract`), from
177
+ [`shapers.ts`](../src/core/shapers.ts) — one shape compiles into a guard,
178
+ coercing parser, JSON Schema, and seeded generator. `deriveShapes` builds a
179
+ whole `Columns` map of them from a table's own cell values.
180
+
181
+ | Shaper | Kind | Signature | Summary |
182
+ | ----------------- | -------- | ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
183
+ | `columnTypeShape` | function | `(type: ColumnType) => ContractShape` | Returns the `ContractShape` a `ColumnType`'s values must satisfy. |
184
+ | `csvTableShape` | const | `ContractShape` | Represents the `ContractShape` of a `CSVTable`'s JSON-serializable projection — an ordered `columns` list of strings plus `rows`, each an open record of JSON values. |
185
+ | `deriveShapes` | function | `(table: CSVTable) => Columns` | Derives one `ContractShape` per table column from that column's cell values across all rows (excluding `undefined`/empty-string cells) — the schema-inference leaf behind `CSVInterface.export` when no explicit `Columns` is given. |
186
+
187
+ ### Validators
188
+
189
+ Guards from [`validators.ts`](../src/core/validators.ts) — total, never
190
+ throw, return `false` for any off-shape input.
191
+
192
+ In a guard table a `Shape` cell holds the type the guard narrows to.
193
+
194
+ | Guard | Kind | Shape | Summary |
195
+ | -------------- | ----- | ------------ | --------------------------------------------------------------------------------------------------------------- |
196
+ | `isCSVTable` | const | `CSVTable` | Determines whether an arbitrary value is a valid `CSVTable` — an array of column names plus an array of `Row`s. |
197
+ | `isColumnType` | const | `ColumnType` | Determines whether a value is a valid `ColumnType` literal. |
198
+
199
+ ### Classes
200
+
201
+ The implementing class of `CSVInterface`, from [`CSV.ts`](../src/core/CSV.ts) —
202
+ documented in full under its own heading following this table.
203
+
204
+ | Name | Kind | Summary |
205
+ | ----- | ----- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
206
+ | `CSV` | class | Wraps a typed `CSVTable` with the query (`find` / `filter` / `reduce`), rewrite (`map`), streaming, and export operations `CSVInterface` declares. |
207
+
208
+ ### `CSV`
209
+
210
+ A `CSV` is constructed from a CSV `string` (which runs `parseCSV`) or from an
211
+ already-parsed `CSVTable` (adopted AS-IS, not re-validated — `errors` is empty
212
+ in that case). It exposes its parsed state through the `readonly table`,
213
+ `readonly rows`, and `readonly errors` members, and it is immutable: `map`
214
+ never mutates the stored table, it returns a new `CSV`. See
215
+ [`## Methods`](#methods) for its public call-signature surface.
216
+
217
+ ### Factories
218
+
219
+ From [`factories.ts`](../src/core/factories.ts).
220
+
221
+ | Factory | Kind | Signature | Summary |
222
+ | --------------------- | -------- | --------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
223
+ | `createCSV` | function | `(input: string \| CSVTable, options?: ParseOptions) => CSVInterface` | Creates a working `CSVInterface` from a CSV string or an already-parsed `CSVTable`. |
224
+ | `createTableContract` | function | `(columns: Columns) => ContractInterface<Row>` | Compiles a `Columns` map into a `ContractInterface` for a `Row` — a guard, coercing parser, JSON Schema, and seeded generator from one shape declaration. |
225
+
226
+ ## Methods
227
+
228
+ The public methods of `CSVInterface`, keyed by its backticked name.
229
+
230
+ #### `CSVInterface`
231
+
232
+ | Method | Returns | Summary |
233
+ | -------- | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
234
+ | `find` | `Row \| undefined` | Finds the first row matching `predicate`, called with each row and its index in table order. |
235
+ | `filter` | `readonly Row[]` | Collects every row matching `predicate`, called with each row and its index in table order. |
236
+ | `map` | `CSVInterface` | Rewrites every row (copy-on-write) and returns a new `CSVInterface`, never mutating this one. |
237
+ | `reduce` | `T` | Folds the rows, in table order, into an accumulator through `callback`. |
238
+ | `stream` | `ReadableStream<Row>` | Returns a web-standard `ReadableStream` over the table's rows (source order) — a lazy, pull-based, backpressure-respecting source that enqueues one row per `pull`. A fresh, independently-replayable stream every call; never mutates the table. |
239
+ | `toJSON` | `CSVTable` | Returns the stored `CSVTable` — the JSON-serializable projection. |
240
+ | `export` | `TableExport` | Produces a portable `TableExport` for moving this CSV's schema elsewhere. |
241
+
242
+ ## RFC 4180 and dialects
243
+
244
+ `parseCSV` / `readRecords` honor the RFC 4180 grammar — quoted fields, an
245
+ embedded delimiter/quote/newline inside a quoted field, and doubled quotes as
246
+ the escape convention — while accepting the common real-world dialect
247
+ variants: `\r\n`, bare `\n`, and bare `\r` line endings are all recognized (a
248
+ CRLF pair counts as one line break), and a mix of them within the same
249
+ document is handled record-by-record. A single leading UTF-8 byte-order-mark
250
+ is always stripped before scanning, regardless of `options`. `delimiter`,
251
+ `quote`, and `escape` (`'double'` doubles an embedded quote, `'backslash'`
252
+ prefixes it) are all caller-configurable knobs, validated by
253
+ `assertValidSeparators` (each exactly one character, distinct from each
254
+ other, and never CR/LF/BOM). Tab-separated output is a `renderCSV` dialect,
255
+ not a separate parser: call `renderCSV` with `delimiter: '\t'`.
256
+
257
+ ## Total parsing and the error model
258
+
259
+ `parseCSV` never throws on malformed DATA — every malformation (an
260
+ unterminated quote, a bad quote placement, a ragged row, a duplicate or
261
+ empty header, the record limit) is collected as a `CSVError` into the
262
+ result's `errors` list, capped at `MAX_ERRORS` (further malformations past
263
+ the cap are silently no longer recorded; scanning still continues). Each
264
+ `CSVError` carries a machine-readable `code` (a `CSVErrorCode`) plus, for a
265
+ parse-time malformation, the 1-based `line`/`column` and 0-based `offset`
266
+ into the (post-BOM) source, with `column` and `offset` measured in UTF-16
267
+ code units. Setting `strict: true` flips this to throw-on-
268
+ first-error: the first collected error throws immediately instead of being
269
+ returned. An invalid OPTION (`INVALID_OPTION` — a malformed delimiter/quote
270
+ pair, an empty `comment`, a negative `limit`, a bad `newline`) is always a
271
+ thrown programmer error, never collected, regardless of `strict`. A ragged
272
+ row — a record whose field count does not match the header — is handled per
273
+ `RaggedPolicy`: `'collect'` pads/truncates the row AND records
274
+ `RAGGED_ROW`; `'pad'` does the same silently (no error recorded); `'error'`
275
+ excludes the row entirely (still recording `RAGGED_ROW`). A duplicate or
276
+ empty header name is deterministically renamed through `uniqueColumns` (a repeat
277
+ gets a `_2`, `_3`, … suffix; a blank name becomes positional) so the table
278
+ always has a full, unique column list even when the header itself was
279
+ malformed.
280
+
281
+ ## Security
282
+
283
+ Every parsed row is built with `Object.create(null)` — a null-prototype
284
+ object — so a hostile header name (`__proto__`, `constructor`, `prototype`)
285
+ becomes a plain OWN property on the row that can never reach
286
+ `Object.prototype`; there is no prototype-pollution path through a CSV
287
+ header, however adversarial. On the render side, `renderCSV` guards against
288
+ CSV/spreadsheet formula injection (the OWASP CSV-injection guidance): a field
289
+ whose first character is one of `SANITIZE_PREFIXES` (`=`, `+`, `-`, `@`, tab,
290
+ CR, LF) is prefixed with a protective `'` when `RenderOptions.sanitize` is
291
+ `true` (the default) — EXCEPT a `+`/`-`-led field that is a plain signed
292
+ number (`NUMERIC_PATTERN`), which is left untouched so legitimate numeric
293
+ data round-trips unmodified. This is a known, intentionally-scoped mitigation
294
+ — it defends against the ASCII formula-trigger characters the OWASP guidance
295
+ names, not against homoglyph or zero-width-character bypasses (a
296
+ lookalike `=` or a zero-width-joined `=` would not match
297
+ `SANITIZE_PREFIXES`); that class of evasion is out of scope for this layer.
298
+
299
+ ## Conservative inference
300
+
301
+ Type inference (`ParseOptions.infer`) is OFF by default — every field parses
302
+ as a `string` unless a caller opts in. When enabled, `inferColumnType`
303
+ decides a type for a WHOLE column at once (never per-cell), so a column with
304
+ even one non-conforming value stays `'text'` entirely. Several common traps
305
+ stay text deliberately: a leading-zero numeral (`'007'`) fails
306
+ `INTEGER_PATTERN` (which permits no leading zeros beyond a bare `0`) and so
307
+ never infers as a number — a phone number or zip code is preserved verbatim;
308
+ scientific notation (`'1e5'`) and `NaN`/`Infinity` are not matched by either
309
+ numeric pattern and stay text; a value outside `Number.isSafeInteger` range
310
+ stays text even though its digits match `INTEGER_PATTERN`; a date string and
311
+ a decimal-comma number (`'3,14'`) both fail both numeric patterns and stay
312
+ text. The two numeric outcomes split on whether any cell carries a fractional
313
+ part: all-integer cells infer `'integer'`, any decimal cell present promotes
314
+ the WHOLE column to `'real'`. `inferColumnType` never infers `'json'` or
315
+ `'blob'` — those require an explicit `Columns` declaration naming the shape.
316
+
317
+ ## Database interop without dependency
318
+
319
+ A `Row` is `Record<string, unknown>` — a plain record any database `Table`'s
320
+ `set`/`add` primitives can accept directly, with no adapter layer and no
321
+ runtime dependency on `@orkestrel/database` (this package never imports it).
322
+ `CSVInterface.toJSON` returns the stored `CSVTable` — the JSON-serializable
323
+ seam a CSV round-trips through when crossing a process boundary or a
324
+ `JSON.stringify` call. `CSVInterface.export` produces a `TableExport` —
325
+ `{ key, columns, schema }`. `@orkestrel/database` declares
326
+ `TableDefinition { primary, columns, schema }`: this package's `Columns` map
327
+ is structurally identical to that package's `ColumnMap`, and `schema` is the
328
+ same JSON Schema `@orkestrel/contract` compiles from it. The two name the key
329
+ column differently — `key` here, `primary` there. The interop is structural:
330
+ no import crosses the package boundary in either direction.
331
+
332
+ ## Streaming boundary
333
+
334
+ `CSVInterface.stream` returns a web-standard `ReadableStream<Row>` — a fresh,
335
+ pull-based stream every call, enqueuing one already-parsed row per `pull` so
336
+ a slow consumer's backpressure is respected. This is a POST-PARSE row
337
+ stream, not chunked ingestion: the entire CSV text is parsed up front (by
338
+ `parseCSV`, synchronously, into a complete `CSVTable`) before `stream()` ever
339
+ enqueues a row. The package parses a whole string. It has no incremental
340
+ parser that consumes a text stream and emits rows as they arrive, so a
341
+ caller with a very large file reads it fully into memory first.
342
+
343
+ ## Patterns
344
+
345
+ Every feature below has a compact, runnable example.
346
+
347
+ ### Parse and query
348
+
349
+ Parses a small CSV string with inference, then reads and filters its rows:
350
+
351
+ ```ts
352
+ import { createCSV } from '@orkestrel/csv'
353
+
354
+ const csv = createCSV('name,age\nAda,36\nGrace,85', { infer: true })
355
+ csv.table // { columns: ['name', 'age'], rows: [{ name: 'Ada', age: 36 }, { name: 'Grace', age: 85 }] }
356
+
357
+ const ada = csv.find((row) => row.name === 'Ada') // Row | undefined
358
+ const adults = csv.filter((row) => Number(row.age) >= 40) // readonly Row[]
359
+ ```
360
+
361
+ ### Rewrite with `map`, then render back
362
+
363
+ Rewrites every row through `map`, then renders the new table back to CSV text:
364
+
365
+ ```ts
366
+ import { createCSV } from '@orkestrel/csv'
367
+ import { renderCSV } from '@orkestrel/csv'
368
+
369
+ const csv = createCSV('name,age\nAda,36', { infer: true })
370
+ const older = csv.map((row) => ({ ...row, age: Number(row.age) + 1 }))
371
+
372
+ renderCSV(older.toJSON()) // 'name,age\r\nAda,37'
373
+ ```
374
+
375
+ Each `map` call returns a NEW `CSVInterface` — the original `csv` is never
376
+ mutated.
377
+
378
+ ### Reduce into an accumulator
379
+
380
+ Folds every row into a running total through `reduce`:
381
+
382
+ ```ts
383
+ import { createCSV } from '@orkestrel/csv'
384
+
385
+ const csv = createCSV('amount\n10\n20\n30', { infer: true })
386
+
387
+ const total = csv.reduce<number>((sum, row) => sum + Number(row.amount), 0) // 60
388
+ ```
389
+
390
+ ### Streaming rows
391
+
392
+ Drains the table through `stream` as a web-standard `ReadableStream`:
393
+
394
+ ```ts
395
+ import { createCSV } from '@orkestrel/csv'
396
+
397
+ const csv = createCSV('a\n1\n2\n3')
398
+
399
+ const reader = csv.stream().getReader()
400
+ const values: string[] = []
401
+ for (let result = await reader.read(); !result.done; result = await reader.read()) {
402
+ values.push(String(result.value.a))
403
+ }
404
+ // values: ['1', '2', '3']
405
+ ```
406
+
407
+ ### Handling errors without `strict`
408
+
409
+ Collects a ragged-row malformation into `errors` instead of throwing:
410
+
411
+ ```ts
412
+ import { createCSV, isCSVError } from '@orkestrel/csv'
413
+
414
+ const csv = createCSV('a,b\n1,2,3') // ragged row — collected, not thrown
415
+ csv.errors.length > 0 // true
416
+ for (const error of csv.errors) {
417
+ if (isCSVError(error)) console.warn(error.code, error.line)
418
+ }
419
+ ```
420
+
421
+ ### `strict` mode throws the first error
422
+
423
+ Throws the first collected error immediately when `strict` is set:
424
+
425
+ ```ts
426
+ import { createCSV, isCSVError } from '@orkestrel/csv'
427
+
428
+ try {
429
+ createCSV('a,b\n1,2,3', { strict: true })
430
+ } catch (error) {
431
+ if (isCSVError(error)) error.code // 'RAGGED_ROW'
432
+ }
433
+ ```
434
+
435
+ ### Exporting a portable schema
436
+
437
+ Exports the parsed table's inferred schema as a portable `TableExport`:
438
+
439
+ ```ts
440
+ import { createCSV } from '@orkestrel/csv'
441
+
442
+ const csv = createCSV('id,name\n1,Ada\n2,Grace', { infer: true })
443
+ const table = csv.export() // { key: 'id', columns: {...}, schema: {...} }
444
+ table.schema // a JSON Schema describing every column
445
+ ```
446
+
447
+ ### Contract-backed row validation
448
+
449
+ Compiles a `Columns` map into a `ContractInterface` and validates a row against it:
450
+
451
+ ```ts
452
+ import { createTableContract, columnTypeShape } from '@orkestrel/csv'
453
+
454
+ const contract = createTableContract({
455
+ id: columnTypeShape('integer'),
456
+ name: columnTypeShape('text'),
457
+ })
458
+ contract.is({ id: 1, name: 'Ada' }) // true
459
+ contract.is({ id: 'x', name: 'Ada' }) // false
460
+ ```
461
+
462
+ ### Guarding an adopted table
463
+
464
+ Guards an unknown value before adopting it as a `CSVTable`:
465
+
466
+ ```ts
467
+ import { createCSV, isCSVTable } from '@orkestrel/csv'
468
+
469
+ function adopt(candidate: unknown) {
470
+ if (!isCSVTable(candidate)) return undefined // total guard - never throws
471
+ return createCSV(candidate) // adopted AS-IS, not re-parsed
472
+ }
473
+ ```
474
+
475
+ ### Tokenizer leaves directly
476
+
477
+ Calls the tokenizer and inference leaves directly, without going through `parseCSV`:
478
+
479
+ ```ts
480
+ import {
481
+ coerceInferred,
482
+ isBreakChar,
483
+ isRowList,
484
+ resolveParseOptions,
485
+ scanField,
486
+ } from '@orkestrel/csv'
487
+
488
+ isBreakChar('\n') // true
489
+ isBreakChar('a') // false
490
+
491
+ const scan = scanField('ab,c', { offset: 0, line: 1, column: 1 }, resolveParseOptions())
492
+ scan.field // { value: 'ab', quoted: false }
493
+
494
+ coerceInferred('42', 'integer') // 42
495
+
496
+ isRowList([{ a: 1 }]) // true
497
+ isRowList({ columns: ['a'], rows: [{ a: 1 }] }) // false
498
+ ```
499
+
500
+ ## Tests
501
+
502
+ - [`tests/guides.test.ts`](../tests/guides.test.ts) — the `## Surface` ↔
503
+ `src/core` bijection (value and type exports), the `CSVInterface` ↔ `CSV`
504
+ method bijection, and the equality gate: every `Summary` cell against its
505
+ declaration's description paragraph, the titled `Parse and query` fence
506
+ against the `@example` block of that title (pinned so the titled pair cannot
507
+ be retired silently), and the README pitch against this guide's tagline. It
508
+ also runs the flagship fences and asserts the values their comments claim.
509
+ - [`tests/src/core/CSV.test.ts`](../tests/src/core/CSV.test.ts) —
510
+ construction from a string vs. an adopted table, `find`/`filter`/`reduce`,
511
+ and `map` copy-on-write behavior.
512
+ - [`tests/src/core/factories.test.ts`](../tests/src/core/factories.test.ts) —
513
+ `createCSV` and `createTableContract` return working, correctly-typed
514
+ results.
515
+ - [`tests/src/core/helpers.test.ts`](../tests/src/core/helpers.test.ts) —
516
+ separator validation and option resolution (incl. `INVALID_OPTION`
517
+ throws), column disambiguation, sanitization, cell serialization, the
518
+ quoting policies, `isRowList` narrowing, `renderCSV` including its
519
+ tab-delimiter dialect, and the tokenizer and table-builder leaves
520
+ (`advancePosition` through `buildRow`).
521
+ - [`tests/src/core/inferers.test.ts`](../tests/src/core/inferers.test.ts) —
522
+ `inferColumnType` against the classic inference traps, `coerceInferred`
523
+ dispatch per `ColumnType`, and `inferRows` copy-on-write.
524
+ - [`tests/src/core/parsers.test.ts`](../tests/src/core/parsers.test.ts) —
525
+ the flat cell coercers and `parseCSV`, incl. ragged-row policies, header
526
+ handling, and `strict`-mode throwing.
527
+ - [`tests/src/core/shapers.test.ts`](../tests/src/core/shapers.test.ts) —
528
+ `columnTypeShape` per `ColumnType`, `csvTableShape` structural validation,
529
+ and `deriveShapes` column derivation.
530
+ - [`tests/src/core/validators.test.ts`](../tests/src/core/validators.test.ts) —
531
+ `isCSVTable` and `isColumnType` soundness on well-formed and off-shape
532
+ input, incl. its leniency-lock cases against `csvTableShape`.
533
+
534
+ ## See also
535
+
536
+ - [`guide.md`](guide.md) — the mirrored guide for `@orkestrel/guide`, the
537
+ devDependency powering this repo's guides-parity test suite.
538
+ - [`contract.md`](contract.md) — the mirrored guide for `@orkestrel/contract`,
539
+ this package's runtime dependency for shapes, guards, and compiled
540
+ contracts.
541
+ - [`README.md`](README.md) — the guides index.