@wafertools/testdata-parser 0.12.0 → 0.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,7 +4,7 @@
4
4
 
5
5
  Rust/WASM parsers for semiconductor test data formats: **STDF**, **ATDF**, **CSV**, **JSON**, and **Parquet**. Compiled to a single WASM module via `wasm-bindgen`; the same Rust source also builds natively (used by [tsmap](https://github.com/wafertools/tsmap)'s Tauri backend).
6
6
 
7
- All formats parse to one shared shape (`ParsedStdf` / `ScanResult`) — there is no format-specific output type on the JS side.
7
+ All formats parse to one shared shape (`ParsedStdf` / `ScanResult`) — there is no format-specific output type on the JS side. A full parse returns it as a compact columnar buffer, which `decodeParsed` (rows) or `decodeColumns` (columns) from `@wafertools/testdata-parser/columnar.js` turns into that shape.
8
8
 
9
9
  ## Install
10
10
 
@@ -18,10 +18,11 @@ The module must be initialized once before calling any parse function — it loa
18
18
 
19
19
  ```js
20
20
  import init, { parse_stdf } from '@wafertools/testdata-parser';
21
+ import { decodeParsed } from '@wafertools/testdata-parser/columnar.js';
21
22
 
22
23
  await init(); // fetches testdata_parser_bg.wasm relative to the module URL
23
24
  const bytes = new Uint8Array(await file.arrayBuffer());
24
- const parsed = parse_stdf(bytes); // ParsedStdf, or throws a ParserError
25
+ const parsed = decodeParsed(parse_stdf(bytes)); // ParsedStdf; parse_stdf throws a ParserError
25
26
  ```
26
27
 
27
28
  **Import `init` as the default export, not by name.** There is also a named `init`
@@ -34,7 +35,7 @@ code, display the message — messages are prose and may be reworded.
34
35
 
35
36
  ```js
36
37
  try {
37
- const parsed = parse_stdf(bytes);
38
+ const parsed = decodeParsed(parse_stdf(bytes));
38
39
  } catch (err) {
39
40
  if (err.code === 'not-stdf') { // the file is not the format its name claims
40
41
  // ...offer to try another parser
@@ -69,23 +70,44 @@ and this README to each other.
69
70
 
70
71
  ## API
71
72
 
72
- Every parse function takes raw file bytes (`Uint8Array`) and returns a plain JS object (via `serde-wasm-bindgen`), or throws a `ParserError` (see above). Gzip-compressed input (`.gz`) is transparently decompressed for every format.
73
+ Every parse function takes raw file bytes (`Uint8Array`), or throws a `ParserError` (see above). The full parses (`parse_*`) return a columnar buffer (`Uint8Array`) to decode as below; the scans and header reads return a plain JS object (via `serde-wasm-bindgen`). Gzip-compressed input (`.gz`) is transparently decompressed for every format.
73
74
 
74
75
  | Function | Signature | Returns |
75
76
  | --- | --- | --- |
76
- | `parse_stdf` | `(bytes: Uint8Array) => ParsedStdf` | Full parse of an STDF file |
77
- | `parse_atdf` | `(bytes: Uint8Array) => ParsedStdf` | Full parse of an ATDF file |
78
- | `parse_csv` | `(bytes: Uint8Array, mapping: CsvMapping) => ParsedStdf` | Full parse of a CSV, using an explicit column mapping |
79
- | `parse_json` | `(bytes: Uint8Array, mapping: CsvMapping) => ParsedStdf` | Full parse of a JSON array-of-records file, using the same mapping shape as CSV |
77
+ | `parse_stdf` | `(bytes: Uint8Array) => Uint8Array` | Full parse of an STDF file, as a columnar buffer |
78
+ | `parse_atdf` | `(bytes: Uint8Array) => Uint8Array` | Full parse of an ATDF file, as a columnar buffer |
79
+ | `parse_csv` | `(bytes: Uint8Array, mapping: CsvMapping) => Uint8Array` | Full parse of a CSV, using an explicit column mapping |
80
+ | `parse_json` | `(bytes: Uint8Array, mapping: CsvMapping) => Uint8Array` | Full parse of a JSON array-of-records file, using the same mapping shape as CSV |
80
81
  | `parquet_headers` | `(bytes: Uint8Array) => ParquetHeadersResult` | Schema + a sample of rows, for a column-mapping UI (see below — unlike CSV/JSON, this one *is* a WASM export) |
81
- | `parse_parquet` | `(bytes: Uint8Array, mapping: CsvMapping) => ParsedStdf` | Full parse of a Parquet file, using the same mapping shape as CSV/JSON |
82
+ | `parse_parquet` | `(bytes: Uint8Array, mapping: CsvMapping) => Uint8Array` | Full parse of a Parquet file, using the same mapping shape as CSV/JSON |
82
83
  | `stdf_test_names` | `(bytes: Uint8Array) => ScanResult` | Fast first-pass scan: test definitions + die count, no die accumulation |
83
84
  | `atdf_test_names` | `(bytes: Uint8Array) => ScanResult` | Same first-pass scan for ATDF |
84
85
  | `stdf_file_meta` | `(bytes: Uint8Array) => FileMeta` | Lot metadata, wafer count, first/last timestamps and site count from an MIR/SDR/WIR/WRR-only scan — no PTR/FTR/PIR/PRR walk, so it stays cheap across a batch of files |
85
86
  | `atdf_file_meta` | `(bytes: Uint8Array) => FileMeta` | Same metadata-only scan for ATDF |
86
87
  | `parquet_distinct_count` | `(bytes: Uint8Array, columns: string[]) => number` | How many distinct combinations of those columns the file holds — a wafer count from `['lot','wafer']` without a full parse, read as a column projection. A column missing from the schema is an error, not a count of zero. Parquet only: CSV/JSON have no equivalent shortcut |
87
- | `parse_stdf_filtered` | `(bytes: Uint8Array, selected: number[]) => ParsedStdf` | Full parse, skipping per-site accumulation for test numbers not in `selected` |
88
- | `parse_atdf_filtered` | `(bytes: Uint8Array, selected: number[]) => ParsedStdf` | Same filtered parse for ATDF |
88
+ | `parse_stdf_filtered` | `(bytes: Uint8Array, selected: number[]) => Uint8Array` | Full parse, skipping per-site accumulation for test numbers not in `selected` |
89
+ | `parse_atdf_filtered` | `(bytes: Uint8Array, selected: number[]) => Uint8Array` | Same filtered parse for ATDF |
90
+
91
+ ### Decoding a parse: `decodeParsed` and `decodeColumns`
92
+
93
+ A full parse returns one `Uint8Array`: typed columns, one per field and one per test, followed by
94
+ a JSON header (everything in `ParsedStdf` except the dies) and its length. It is small, and it transfers from a
95
+ Worker without a copy (`postMessage(buf, [buf.buffer])`). Decode it with one of two functions
96
+ from `@wafertools/testdata-parser/columnar.js`, which ships in the package:
97
+
98
+ - **`decodeParsed(buf)`** gives the `ParsedStdf` described below, with one `DieResult` object
99
+ per die.
100
+ - **`decodeColumns(buf)`** gives the same `ParsedStdf`, except that each wafer's `results`
101
+ stays columns: positions, bins and site as typed arrays (−32768 = no position, 65535 = no
102
+ bin or site), and test values and verdicts as `{ indices, values }` per test number (only the
103
+ records that have one). That is the shape `@wafertools/wafermap`'s `buildWaferMap` accepts as
104
+ `results` (`DieColumns`), so a large lot is built without an object per die: pass
105
+ `wafer.results` as it is.
106
+ Within a wafer, tests with values (or verdicts) on the same records share one `indices`
107
+ array, so treat the arrays as read-only.
108
+
109
+ Test values are sent as 32-bit floats when every value in the column is exactly representable
110
+ as one (STDF readings are), and as 64-bit otherwise, so no value changes.
89
111
 
90
112
  ### Two-pass parsing (STDF/ATDF)
91
113
 
@@ -116,8 +138,10 @@ interface CsvMapping {
116
138
  testnameCol?: string; // for "tall" CSVs: column holding the test name per row
117
139
  testnumberCol?: string; // for "tall" CSVs: column holding the test's real number per row
118
140
  testvalueCol?: string; // for "tall" CSVs: column holding the test value per row
119
- loLimitCol?: string;
141
+ loLimitCol?: string; // for "tall" CSVs: test limits (STDF LO_LIMIT/HI_LIMIT) per row
120
142
  hiLimitCol?: string;
143
+ loSpecCol?: string; // for "tall" CSVs: spec limits (STDF LO_SPEC/HI_SPEC) per row — a separate pair
144
+ hiSpecCol?: string;
121
145
  unitsCol?: string;
122
146
  passBins: number[]; // hbin/sbin values treated as a pass for pass/fail summary
123
147
  }
@@ -220,11 +244,13 @@ type ParserWarningCode =
220
244
  | 'unpositioned-dies' // dies with no X/Y — real data, not placeable
221
245
  | 'bin-invalid' // a bin outside STDF's 0–32767, or a missing hard bin
222
246
  | 'coordinate-invalid' // an X/Y outside STDF's -32767..32767
247
+ | 'site-invalid' // a site number outside STDF's 0–255 (flat formats)
223
248
  | 'result-unusable' // results the tester flagged unusable (value left out, verdict kept)
224
249
  | 'records-not-read' // records this parser does not read yet (MPR)
225
250
  | 'wafer-end-missing' // a wafer had no WRR; closed at the next wafer or end of file
226
251
  | 'file-truncated' // the file ends part-way through a record
227
252
  | 'record-malformed' // a PRR too short to hold its required fields
253
+ | 'test-number-invalid' // an ATDF test number that is not a u32 (record left out)
228
254
  | 'values-not-numeric' // a mapped column held values that would not coerce
229
255
  | 'retests-assumed' // repeated positions read as retests
230
256
  | 'wafer-split-by-column' // one wafer per value of a mapped column
@@ -262,7 +288,7 @@ an empty array, which asserts that nothing passes.
262
288
  **`warnings` carries a stable `code`, prose, and a severity** — branch on the code, display
263
289
  the message, and never match on the prose. `severity: 'error'` means a number or a plot
264
290
  built from this result can mislead, because data was dropped or a value was substituted
265
- (`unpositioned-dies`, `bin-invalid`, `coordinate-invalid`, `record-malformed`, `records-not-read`, `values-not-numeric`); `'warning'` means the
291
+ (`unpositioned-dies`, `bin-invalid`, `coordinate-invalid`, `site-invalid`, `record-malformed`, `records-not-read`, `test-number-invalid`, `values-not-numeric`); `'warning'` means the
266
292
  parse applied a documented rule or interpretation — one you may want to change, or, like
267
293
  `result-unusable`, the spec's own rule for leaving out values the tester flagged — and the
268
294
  result means what the file says. Nothing here is fatal — the parse succeeded. Surface them: a silently discarded
@@ -375,6 +401,15 @@ pub struct CsvHeadersResult {
375
401
  - **Byte readers are panic-free.** STDF/ATDF field readers are bounds-checked and return `Option`/`Result` rather than panicking on truncated input — a panic inside WASM aborts the whole module with no recovery, so this is a hard requirement, not a style preference.
376
402
  - **Big-endian and little-endian STDF** are both supported (detected from the FAR record's `CPU_TYPE`).
377
403
  - **Gzip is transparent** — every entry point decompresses `.gz` input automatically by sniffing the magic bytes; callers don't need to branch on compression.
404
+ - **MPR records are not read yet.** A multiple-result parametric record (STDF `MPR`, ATDF `MPR:`)
405
+ carries several results for one test; both parsers skip them — in the full parse and the
406
+ test-name scan alike — and count them in a `records-not-read` warning, so the missing tests are
407
+ never silent. PTR and FTR records are read in full.
408
+ - **Every format applies STDF V4's value ranges.** A bin, coordinate or site number STDF cannot
409
+ store, or a test value that is not finite, is missing in CSV, JSON and Parquet exactly as in
410
+ STDF and ATDF, and is reported under the same codes (`bin-invalid`, `coordinate-invalid`,
411
+ `site-invalid`, `result-unusable`). The rule lives in one place (`SpecCheck` in `types.rs`),
412
+ which the flat formats reach through `flat_wafers::into_parsed`.
378
413
  - **CSV/JSON/Parquet test numbers fall back to a deterministic hash only when the file itself carries no real one.** Neither format has a *mandatory* STDF-style test number the way STDF/ATDF do, but a real one is used whenever the source data has it — see "Test identity — real number vs. synthesized one" above for the wide/tall rules. Hashing is the fallback, not the default: it fires per test only when no real number was mapped or the mapped column's value didn't parse (`test_identity::stable_test_number`, FNV-1a with a fixed seed and a reserved floor, collision-probed so two tests in one file can never collide, and never colliding with a genuine numeric-header/`testnumberCol` value either). Deliberately not sequential/encounter-order: a hash means the number for a given test doesn't change if the file is reordered or a column is added — a *hashed* number is otherwise meaningless and callers should never rely on its value, only on it being stable and unique within one parse. `order` (see `TestDef` above) carries the file's own display order instead.
379
414
  - **Parquet reads through a row-oriented API, not Arrow.** `parquet::record::Row`/`Field` rather than the `arrow` feature — a closer fit for this crate's row-based `DieResult` model, and a smaller WASM bundle (no Arrow array machinery pulled in). A typed Parquet cell is coerced to `f64` for numeric roles and to a plain string otherwise; a value that fails to coerce (e.g. a numeric role mapped to a genuinely string-typed column) is skipped and surfaced as one summarised entry in `warnings`, not a panic or a silent zero.
380
415
  - **Parquet's `zstd` codec is native-only.** `snappy`, `gzip`, `lz4`, and `brotli` build for `wasm32-unknown-unknown` with no extra toolchain; `zstd`'s C library needs a real C cross-compiler targeting wasm32, which a plain `wasm-pack build` doesn't assume is available. A `zstd`-compressed Parquet file parses natively but fails clearly on the WASM build.
package/columnar.d.ts ADDED
@@ -0,0 +1,42 @@
1
+ import type { ParsedStdf } from './testdata_parser.js';
2
+
3
+ /**
4
+ * Turn the buffer a parse function returns (`parse_stdf`, `parse_csv`, …) into a
5
+ * `ParsedStdf`. Accepts the `Uint8Array` the function returned, or the
6
+ * `ArrayBuffer` it was transferred or delivered as.
7
+ */
8
+ export function decodeParsed(input: Uint8Array | ArrayBuffer): ParsedStdf;
9
+
10
+ /** A wafer's records as columns: the shape `@wafertools/wafermap` accepts as `results` (`DieColumns`). */
11
+ export interface ParsedColumns {
12
+ /** Number of records. */
13
+ count: number;
14
+ /** Per record; −32768 = no position. */
15
+ x?: Int32Array;
16
+ y?: Int32Array;
17
+ /** Per record; 65535 = none. */
18
+ hbin?: Uint32Array;
19
+ sbin?: Uint32Array;
20
+ siteNum?: Uint32Array;
21
+ partId?: Array<number | string | undefined>;
22
+ supersedes?: Array<'partId' | 'position' | undefined>;
23
+ /** Per test number: the records that have a value, and their values. */
24
+ testValues?: Record<number, { indices: Int32Array; values: Float32Array | Float64Array }>;
25
+ /** Per test number: the records that have a recorded verdict; 1 = pass, 0 = fail. */
26
+ testPass?: Record<number, { indices: Int32Array; values: Int8Array }>;
27
+ }
28
+
29
+ /** A `ParsedStdf` whose wafers carry their records as columns. */
30
+ export type ParsedStdfColumns = Omit<ParsedStdf, 'wafers'> & {
31
+ wafers: Array<Omit<ParsedStdf['wafers'][number], 'results'> & { results: ParsedColumns }>;
32
+ };
33
+
34
+ /**
35
+ * Like `decodeParsed`, but each wafer's `results` stays columns, with no object
36
+ * per die: positions, bins and site as typed arrays with STDF's missing values,
37
+ * test values and verdicts sparse. Pass a wafer's `results` straight to
38
+ * `buildWaferMap`. Nothing refers to the input buffer afterwards. Within a
39
+ * wafer, columns with the same records present share one `indices` array:
40
+ * treat them as read-only.
41
+ */
42
+ export function decodeColumns(input: Uint8Array | ArrayBuffer): ParsedStdfColumns;
package/columnar.js ADDED
@@ -0,0 +1,256 @@
1
+ // Decodes the columnar buffer the parse functions return into a `ParsedStdf`.
2
+ // The encoder, and the full layout, is `src/columnar.rs`; the two ship together
3
+ // in this package so they cannot drift apart.
4
+
5
+ /** The format this decoder reads — `FORMAT` in `src/columnar.rs`. */
6
+ const FORMAT = 2;
7
+
8
+ const MISSING_I32 = -2147483648;
9
+ const MISSING_U32 = 4294967295;
10
+ const SUPERSEDES = [undefined, 'partId', 'position'];
11
+
12
+ /**
13
+ * The buffer's header, and each wafer's columns as typed-array views over the
14
+ * buffer (no copy). Shared by `decodeParsed` and `decodeColumns`.
15
+ */
16
+ function readColumns(input) {
17
+ let bytes = input instanceof ArrayBuffer ? new Uint8Array(input) : input;
18
+ // Typed-array views need aligned offsets; a buffer that is itself a view at an
19
+ // odd offset is copied once so every column below can be viewed in place.
20
+ if (bytes.byteOffset % 8 !== 0) bytes = bytes.slice();
21
+ const { buffer, byteOffset } = bytes;
22
+
23
+ // Columns first, then the JSON header, then its length: see `src/columnar.rs`.
24
+ const headerLen = new DataView(buffer, byteOffset + bytes.length - 4, 4).getUint32(0, true);
25
+ const header = JSON.parse(new TextDecoder().decode(bytes.subarray(bytes.length - 4 - headerLen, bytes.length - 4)));
26
+ if (header.columnarFormat !== FORMAT) {
27
+ throw new Error(`testdata-parser: columnar format ${header.columnarFormat} is not the format ${FORMAT} this decoder reads — the parser and its decoder come from different versions`);
28
+ }
29
+ delete header.columnarFormat;
30
+ const body = byteOffset;
31
+ checkLayout(header.wafers, bytes.length - 4 - headerLen);
32
+
33
+ const wafers = header.wafers.map((wafer) => {
34
+ const n = wafer.dieCount;
35
+ const view = (c) => {
36
+ const at = body + c.offset;
37
+ switch (c.kind) {
38
+ case 'i32': return new Int32Array(buffer, at, n);
39
+ case 'u32': return new Uint32Array(buffer, at, n);
40
+ case 'u8': return new Uint8Array(buffer, at, n);
41
+ case 'i8': return new Int8Array(buffer, at, n);
42
+ case 'f32': return new Float32Array(buffer, at, n);
43
+ case 'f64': return new Float64Array(buffer, at, n);
44
+ default: throw new Error(`testdata-parser: unknown column kind ${c.kind}`);
45
+ }
46
+ };
47
+ const col = {};
48
+ const values = [];
49
+ const verdicts = [];
50
+ for (const c of wafer.columns) {
51
+ if (c.field === 'testValues') values.push([c.test, view(c)]);
52
+ else if (c.field === 'testPass') verdicts.push([c.test, view(c)]);
53
+ else col[c.field] = view(c);
54
+ }
55
+ const partIds = wafer.partIds;
56
+ delete wafer.dieCount;
57
+ delete wafer.columns;
58
+ delete wafer.partIds;
59
+ return { wafer, n, col, values, verdicts, partIds };
60
+ });
61
+ return { header, wafers };
62
+ }
63
+
64
+ const WIDTH = { i32: 4, u32: 4, f32: 4, f64: 8, u8: 1, i8: 1 };
65
+
66
+ /**
67
+ * Every column must lie inside the body (before the header) and no two may
68
+ * share bytes. A typed array only refuses a column that runs off the buffer; one
69
+ * at a wrong offset inside it would decode as plausible, wrong data.
70
+ */
71
+ function checkLayout(wafers, bodyLen) {
72
+ if (!(bodyLen >= 0)) throw new Error('testdata-parser: columnar header is longer than the buffer');
73
+ const spans = [];
74
+ for (const wafer of wafers) {
75
+ for (const c of wafer.columns) {
76
+ const width = WIDTH[c.kind];
77
+ if (width === undefined) throw new Error(`testdata-parser: unknown column kind ${c.kind}`);
78
+ const end = c.offset + wafer.dieCount * width;
79
+ if (!Number.isInteger(c.offset) || c.offset < 0 || end > bodyLen) {
80
+ throw new Error(`testdata-parser: column ${c.field}${c.test === undefined ? '' : ` ${c.test}`} lies outside the column data`);
81
+ }
82
+ spans.push([c.offset, end]);
83
+ }
84
+ }
85
+ spans.sort((a, b) => a[0] - b[0]);
86
+ for (let i = 1; i < spans.length; i++) {
87
+ if (spans[i][0] < spans[i - 1][1]) throw new Error('testdata-parser: two columns share bytes — the buffer is corrupt');
88
+ }
89
+ }
90
+
91
+ /**
92
+ * @param {Uint8Array | ArrayBuffer} input
93
+ * @returns {import('./testdata_parser.js').ParsedStdf}
94
+ */
95
+ export function decodeParsed(input) {
96
+ const { header, wafers } = readColumns(input);
97
+ for (const { wafer, n, col, values, verdicts, partIds } of wafers) {
98
+ const { x, y, dieIndex, hbin, sbin, siteNum, supersedes } = col;
99
+
100
+ // Fields assigned in one fixed order, so every die shares a shape.
101
+ const results = new Array(n);
102
+ for (let i = 0; i < n; i++) {
103
+ const d = {};
104
+ if (x && x[i] !== MISSING_I32) d.x = x[i];
105
+ if (y && y[i] !== MISSING_I32) d.y = y[i];
106
+ if (dieIndex && dieIndex[i] !== MISSING_U32) d.dieIndex = dieIndex[i];
107
+ if (hbin && hbin[i] !== MISSING_U32) d.hbin = hbin[i];
108
+ if (sbin && sbin[i] !== MISSING_U32) d.sbin = sbin[i];
109
+ if (siteNum && siteNum[i] !== MISSING_U32) d.siteNum = siteNum[i];
110
+ if (partIds && partIds[i] != null) d.partId = partIds[i];
111
+ if (supersedes && supersedes[i] !== 0) d.supersedes = SUPERSEDES[supersedes[i]];
112
+ results[i] = d;
113
+ }
114
+ // Highest test number first. V8 picks an object's element storage from the
115
+ // first integer-like key it receives: a small one (1001 is below its
116
+ // 1024-slot gap limit) gets a holey array sized to the largest key, ~6 KB
117
+ // for 50 readings at 1001–1050; a large one gets a dictionary, ~1.5 KB.
118
+ // Enumeration order is unaffected — integer-like keys always enumerate
119
+ // ascending — so this changes memory only.
120
+ const byKeyDesc = (a, b) => Number(b[0]) - Number(a[0]);
121
+ values.sort(byKeyDesc);
122
+ verdicts.sort(byKeyDesc);
123
+ attachValues(results, values);
124
+ attachVerdicts(results, verdicts);
125
+ wafer.results = results;
126
+ }
127
+ return header;
128
+ }
129
+
130
+ /**
131
+ * Like `decodeParsed`, but each wafer's `results` stays columns: the shape
132
+ * `@wafertools/wafermap`'s `buildWaferMap` accepts as `results` (`DieColumns`),
133
+ * with no object per die. Positions, bins and site are per-record typed
134
+ * arrays with STDF's missing values (−32768 for coordinates, 65535 for bins and
135
+ * site); test values and verdicts are sparse (`indices` of the records that
136
+ * have one, and the `values`), values at the precision they were sent.
137
+ *
138
+ * Nothing refers to the input buffer afterwards, so it can be freed.
139
+ *
140
+ * @param {Uint8Array | ArrayBuffer} input
141
+ */
142
+ export function decodeColumns(input) {
143
+ const { header, wafers } = readColumns(input);
144
+ for (const { wafer, n, col, values, verdicts, partIds } of wafers) {
145
+ const results = { count: n };
146
+ if (col.x) results.x = stdfInts(col.x, MISSING_I32, COORD_MISSING, Int32Array);
147
+ if (col.y) results.y = stdfInts(col.y, MISSING_I32, COORD_MISSING, Int32Array);
148
+ if (col.hbin) results.hbin = stdfInts(col.hbin, MISSING_U32, U16_MISSING, Uint32Array);
149
+ if (col.sbin) results.sbin = stdfInts(col.sbin, MISSING_U32, U16_MISSING, Uint32Array);
150
+ if (col.siteNum) results.siteNum = stdfInts(col.siteNum, MISSING_U32, U16_MISSING, Uint32Array);
151
+ if (partIds) results.partId = partIds.map(v => (v == null ? undefined : v));
152
+ if (col.supersedes) results.supersedes = Array.from(col.supersedes, c => SUPERSEDES[c]);
153
+ const shared = sharedIndices(n);
154
+ if (values.length) {
155
+ results.testValues = {};
156
+ for (const [key, dense] of values) results.testValues[key] = sparse(dense, v => v === v, shared); // NaN is missing
157
+ }
158
+ if (verdicts.length) {
159
+ results.testPass = {};
160
+ for (const [key, dense] of verdicts) results.testPass[key] = sparse(dense, v => v !== -1, shared);
161
+ }
162
+ wafer.results = results;
163
+ }
164
+ return header;
165
+ }
166
+
167
+ const COORD_MISSING = -32768;
168
+ const U16_MISSING = 65535;
169
+
170
+ /** A copy of an integer column with the wire's missing value replaced by STDF's. */
171
+ function stdfInts(column, wireMissing, stdfMissing, Type) {
172
+ const out = new Type(column.length);
173
+ for (let i = 0; i < column.length; i++) out[i] = column[i] === wireMissing ? stdfMissing : column[i];
174
+ return out;
175
+ }
176
+
177
+ /**
178
+ * A dense column as `{ indices, values }` of the entries `present` accepts,
179
+ * values in the column's own type. Columns with the same entries present share
180
+ * one `indices` array (from `shared`): in most lots every die has every test,
181
+ * so a wafer's columns all hold the same indices, and one copy of them about halves
182
+ * the memory the columns take. Nothing writes to an `indices` array.
183
+ */
184
+ function sparse(dense, present, shared) {
185
+ const scratch = shared.scratch;
186
+ let k = 0;
187
+ let hash = 0;
188
+ for (let i = 0; i < dense.length; i++) {
189
+ if (present(dense[i])) { scratch[k++] = i; hash = (Math.imul(hash, 31) + i) | 0; }
190
+ }
191
+ const values = new dense.constructor(k);
192
+ for (let j = 0; j < k; j++) values[j] = dense[scratch[j]];
193
+ return { indices: shared.get(k, hash), values };
194
+ }
195
+
196
+ /**
197
+ * The index arrays of one wafer's columns: `get` returns an earlier array
198
+ * holding the same indices as `scratch[0..count)`, or a copy of them.
199
+ */
200
+ function sharedIndices(n) {
201
+ const scratch = new Int32Array(n);
202
+ const seen = new Map(); // `${count}:${hash}` → arrays with that count and hash
203
+ return {
204
+ scratch,
205
+ get(count, hash) {
206
+ const key = `${count}:${hash}`;
207
+ const candidates = seen.get(key);
208
+ if (candidates) {
209
+ for (const c of candidates) {
210
+ let same = true;
211
+ for (let j = 0; j < count; j++) if (c[j] !== scratch[j]) { same = false; break; }
212
+ if (same) return c;
213
+ }
214
+ }
215
+ const indices = scratch.slice(0, count);
216
+ if (candidates) candidates.push(indices); else seen.set(key, [indices]);
217
+ return indices;
218
+ },
219
+ };
220
+ }
221
+
222
+ // Each die's map is built in one go, not a column at a time reaching back into
223
+ // every die once per test, and in functions of their own: as one long loop
224
+ // inside `decodeParsed`, WebKit's engine (the desktop app on Linux and macOS)
225
+ // left the work in its slowest tier and it took ~7 s on a 266k-die lot.
226
+
227
+ /** A test key as the property key to write: its number when it is one, so no string is parsed per die. */
228
+ const propKey = (key) => (String(Number(key)) === key ? Number(key) : key);
229
+
230
+ function attachValues(results, columns) {
231
+ const keys = columns.map(([key]) => propKey(key));
232
+ const cols = columns.map(([, v]) => v);
233
+ const t = cols.length;
234
+ for (let i = 0; i < results.length; i++) {
235
+ let map;
236
+ for (let j = 0; j < t; j++) {
237
+ const value = cols[j][i];
238
+ if (value === value) (map ??= {})[keys[j]] = value; // NaN is missing
239
+ }
240
+ if (map) results[i].testValues = map;
241
+ }
242
+ }
243
+
244
+ function attachVerdicts(results, columns) {
245
+ const keys = columns.map(([key]) => propKey(key));
246
+ const cols = columns.map(([, v]) => v);
247
+ const t = cols.length;
248
+ for (let i = 0; i < results.length; i++) {
249
+ let map;
250
+ for (let j = 0; j < t; j++) {
251
+ const verdict = cols[j][i];
252
+ if (verdict !== -1) (map ??= {})[keys[j]] = verdict === 1;
253
+ }
254
+ if (map) results[i].testPass = map;
255
+ }
256
+ }
package/llms.txt CHANGED
@@ -4,11 +4,12 @@ Rust/WASM parsers for semiconductor test data: STDF, ATDF, CSV, JSON, Parquet.
4
4
  One Rust source compiles two ways — to WebAssembly for the browser, and natively for a
5
5
  Tauri/Rust backend — and both produce identical results from identical input. All five
6
6
  formats parse to one shared shape (`ParsedStdf`), so a consuming app writes its rendering
7
- and analysis code once.
7
+ and analysis code once. A full parse returns a columnar buffer (`Uint8Array`); decode it with
8
+ `decodeParsed` or `decodeColumns` from `@wafertools/testdata-parser/columnar.js`.
8
9
 
9
10
  Output feeds `@wafertools/wafermap` (https://wafertools.github.io/wafermap/) directly:
10
- parse to `ParsedStdf`, hand the dies to `buildWaferMap()`, render. See "Handing the result
11
- to wafermap" below — the fields that make yield correct are easy to drop on the floor.
11
+ parse, `decodeColumns`, hand each wafer's `results` to `buildWaferMap()`, render. See "Handing
12
+ the result to wafermap" below — the fields that make yield correct are easy to drop on the floor.
12
13
 
13
14
  BEFORE INTEGRATING: tsmap (https://wafertools.github.io/tsmap/) is a finished, free,
14
15
  MIT-licensed desktop and browser application built on this parser and wafermap — it opens
@@ -25,7 +26,8 @@ resolve.
25
26
  - [tsmap](https://wafertools.github.io/tsmap/) - the application built on this package.
26
27
 
27
28
  ## The API in one line each
28
- - `parse_stdf(bytes)` / `parse_atdf(bytes)` - full parse. No mapping needed; the format carries its own test identity.
29
+ - `parse_stdf(bytes)` / `parse_atdf(bytes)` - full parse, returned as a columnar buffer. No mapping needed; the format carries its own test identity.
30
+ - `decodeParsed(buf)` / `decodeColumns(buf)` (from `@wafertools/testdata-parser/columnar.js`) - a full parse's buffer as `ParsedStdf` with one object per die, or with each wafer's `results` as columns in wafermap's `DieColumns` shape. Every `parse_*` full parse needs one of them.
29
31
  - `parse_csv(bytes, mapping)` / `parse_json(bytes, mapping)` / `parse_parquet(bytes, mapping)` - full parse with an explicit `CsvMapping`. There is no header auto-detection.
30
32
  - `stdf_test_names(bytes)` / `atdf_test_names(bytes)` - fast scan: test definitions + die count, no per-die accumulation.
31
33
  - `parse_stdf_filtered(bytes, selected)` / `parse_atdf_filtered(bytes, selected)` - full parse that accumulates values only for the selected test numbers.
@@ -46,11 +48,13 @@ resolve.
46
48
  - **Parse off the main thread in a browser.** The WASM build is slower than native and a large STDF will freeze the page otherwise. tsmap's `parserWorker.ts` is a worked example, including the `new URL(..., import.meta.url)` resolution that bundlers get finicky about.
47
49
 
48
50
  ## Handing the result to wafermap
51
+ With `decodeColumns`, each wafer's `results` is already wafermap's `results` (`DieColumns`):
52
+ pass it as it is, which builds even a large lot without an object per die. Tests present on the same records share one `indices` array, so never write to one. The rest of
49
53
  `ParsedStdf` is close to wafermap's input but is not the same object — convert deliberately,
50
54
  and carry these across or the map is plausibly wrong:
51
55
  - `passHbins` -> wafermap's `passBins`. Omit it and wafermap defaults to `[1]`, which decides both the yield number and the wording of its label whether or not bin 1 is this program's pass bin. Absent when no HBR record carried a usable Pass flag — then leave wafermap's default alone rather than passing an empty array, which would mean "nothing passes".
52
56
  - `hbinDefs` / `sbinDefs` -> wafermap's `hbinDefs` / `sbinDefs`, giving bins their real names.
53
- - `testPass` -> wafermap's `die.testPass`: recorded per-test verdicts, `true` = pass. Functional (FTR) tests live here only — they have no measured value. Read them through wafermap's `getTestPassStatus()`; a missing verdict is no-data, never a fail.
57
+ - `testPass` (in `results`, rows or columns) -> wafermap's `die.testPass`: recorded per-test verdicts, pass = `true` (`1` in columns). Functional (FTR) tests live here only — they have no measured value. Read them through wafermap's `getTestPassStatus()`; a missing verdict is no-data, never a fail.
54
58
  - `testDefs` is a `Record<string, TestDef>` keyed by test number as a string, and `TestDef` here has no `testNumber` field. wafermap wants an array of `TestDef` each carrying a numeric `testNumber`, so the key becomes the field during conversion.
55
59
  - `hbin` / `sbin` are already numbers — never `?? 0` them. A missing bin is not bin 0.
56
60
  - A parse `warning` with `severity: 'error'` is worth showing next to the map, not just logging: `unpositioned-dies` means the map holds fewer dies than the file, and `soft-bin-mirrored` means a soft-bin map shows numbers the file never stated.
package/package.json CHANGED
@@ -2,7 +2,7 @@
2
2
  "name": "@wafertools/testdata-parser",
3
3
  "type": "module",
4
4
  "description": "Rust/WASM parsers for semiconductor test data formats (STDF, ATDF, CSV, JSON, Parquet)",
5
- "version": "0.12.0",
5
+ "version": "0.14.0",
6
6
  "license": "MIT",
7
7
  "repository": {
8
8
  "type": "git",
@@ -12,7 +12,9 @@
12
12
  "testdata_parser_bg.wasm",
13
13
  "testdata_parser.js",
14
14
  "testdata_parser.d.ts",
15
- "llms.txt"
15
+ "llms.txt",
16
+ "columnar.js",
17
+ "columnar.d.ts"
16
18
  ],
17
19
  "main": "testdata_parser.js",
18
20
  "types": "testdata_parser.d.ts",
@@ -87,11 +87,13 @@ export type ParserWarningCode =
87
87
  | "unpositioned-dies"
88
88
  | "bin-invalid"
89
89
  | "coordinate-invalid"
90
+ | "site-invalid"
90
91
  | "result-unusable"
91
92
  | "records-not-read"
92
93
  | "wafer-end-missing"
93
94
  | "file-truncated"
94
95
  | "record-malformed"
96
+ | "test-number-invalid"
95
97
  | "values-not-numeric"
96
98
  | "retests-assumed"
97
99
  | "wafer-split-by-column"
@@ -193,8 +195,13 @@ export interface CsvMapping {
193
195
  testnumberCol?: string | null;
194
196
  /** Tall format: the column holding each row's measured value. */
195
197
  testvalueCol?: string | null;
198
+ /** Tall format: the columns holding each test's test limits (STDF LO_LIMIT/HI_LIMIT). */
196
199
  loLimitCol?: string | null;
197
200
  hiLimitCol?: string | null;
201
+ /** Tall format: the columns holding each test's spec limits (STDF LO_SPEC/HI_SPEC) —
202
+ * a separate pair from the test limits, never mixed with them. */
203
+ loSpecCol?: string | null;
204
+ hiSpecCol?: string | null;
198
205
  unitsCol?: string | null;
199
206
  /** Bins counted as a pass in this file's own pass/fail summary. */
200
207
  passBins: number[];
@@ -237,19 +244,19 @@ export function parquet_distinct_count(bytes: Uint8Array, columns: string[]): nu
237
244
 
238
245
  export function parquet_headers(bytes: Uint8Array): ParquetHeadersResult;
239
246
 
240
- export function parse_atdf(bytes: Uint8Array): ParsedStdf;
247
+ export function parse_atdf(bytes: Uint8Array): Uint8Array;
241
248
 
242
- export function parse_atdf_filtered(bytes: Uint8Array, selected: number[]): ParsedStdf;
249
+ export function parse_atdf_filtered(bytes: Uint8Array, selected: number[]): Uint8Array;
243
250
 
244
- export function parse_csv(bytes: Uint8Array, mapping: CsvMapping): ParsedStdf;
251
+ export function parse_csv(bytes: Uint8Array, mapping: CsvMapping): Uint8Array;
245
252
 
246
- export function parse_json(bytes: Uint8Array, mapping: CsvMapping): ParsedStdf;
253
+ export function parse_json(bytes: Uint8Array, mapping: CsvMapping): Uint8Array;
247
254
 
248
- export function parse_parquet(bytes: Uint8Array, mapping: CsvMapping): ParsedStdf;
255
+ export function parse_parquet(bytes: Uint8Array, mapping: CsvMapping): Uint8Array;
249
256
 
250
- export function parse_stdf(bytes: Uint8Array): ParsedStdf;
257
+ export function parse_stdf(bytes: Uint8Array): Uint8Array;
251
258
 
252
- export function parse_stdf_filtered(bytes: Uint8Array, selected: number[]): ParsedStdf;
259
+ export function parse_stdf_filtered(bytes: Uint8Array, selected: number[]): Uint8Array;
253
260
 
254
261
  export function stdf_file_meta(bytes: Uint8Array): FileMeta;
255
262
 
@@ -264,13 +271,13 @@ export interface InitOutput {
264
271
  readonly init: () => void;
265
272
  readonly parquet_distinct_count: (a: number, b: number, c: any) => [number, number, number];
266
273
  readonly parquet_headers: (a: number, b: number) => [number, number, number];
267
- readonly parse_atdf: (a: number, b: number) => [number, number, number];
268
- readonly parse_atdf_filtered: (a: number, b: number, c: any) => [number, number, number];
269
- readonly parse_csv: (a: number, b: number, c: any) => [number, number, number];
270
- readonly parse_json: (a: number, b: number, c: any) => [number, number, number];
271
- readonly parse_parquet: (a: number, b: number, c: any) => [number, number, number];
272
- readonly parse_stdf: (a: number, b: number) => [number, number, number];
273
- readonly parse_stdf_filtered: (a: number, b: number, c: any) => [number, number, number];
274
+ readonly parse_atdf: (a: number, b: number) => [number, number, number, number];
275
+ readonly parse_atdf_filtered: (a: number, b: number, c: any) => [number, number, number, number];
276
+ readonly parse_csv: (a: number, b: number, c: any) => [number, number, number, number];
277
+ readonly parse_json: (a: number, b: number, c: any) => [number, number, number, number];
278
+ readonly parse_parquet: (a: number, b: number, c: any) => [number, number, number, number];
279
+ readonly parse_stdf: (a: number, b: number) => [number, number, number, number];
280
+ readonly parse_stdf_filtered: (a: number, b: number, c: any) => [number, number, number, number];
274
281
  readonly stdf_file_meta: (a: number, b: number) => [number, number, number];
275
282
  readonly stdf_test_names: (a: number, b: number) => [number, number, number];
276
283
  readonly __wbindgen_malloc: (a: number, b: number) => number;
@@ -64,105 +64,119 @@ export function parquet_headers(bytes) {
64
64
 
65
65
  /**
66
66
  * @param {Uint8Array} bytes
67
- * @returns {ParsedStdf}
67
+ * @returns {Uint8Array}
68
68
  */
69
69
  export function parse_atdf(bytes) {
70
70
  const ptr0 = passArray8ToWasm0(bytes, wasm.__wbindgen_malloc);
71
71
  const len0 = WASM_VECTOR_LEN;
72
72
  const ret = wasm.parse_atdf(ptr0, len0);
73
- if (ret[2]) {
74
- throw takeFromExternrefTable0(ret[1]);
73
+ if (ret[3]) {
74
+ throw takeFromExternrefTable0(ret[2]);
75
75
  }
76
- return takeFromExternrefTable0(ret[0]);
76
+ var v2 = getArrayU8FromWasm0(ret[0], ret[1]).slice();
77
+ wasm.__wbindgen_free(ret[0], ret[1] * 1, 1);
78
+ return v2;
77
79
  }
78
80
 
79
81
  /**
80
82
  * @param {Uint8Array} bytes
81
83
  * @param {number[]} selected
82
- * @returns {ParsedStdf}
84
+ * @returns {Uint8Array}
83
85
  */
84
86
  export function parse_atdf_filtered(bytes, selected) {
85
87
  const ptr0 = passArray8ToWasm0(bytes, wasm.__wbindgen_malloc);
86
88
  const len0 = WASM_VECTOR_LEN;
87
89
  const ret = wasm.parse_atdf_filtered(ptr0, len0, selected);
88
- if (ret[2]) {
89
- throw takeFromExternrefTable0(ret[1]);
90
+ if (ret[3]) {
91
+ throw takeFromExternrefTable0(ret[2]);
90
92
  }
91
- return takeFromExternrefTable0(ret[0]);
93
+ var v2 = getArrayU8FromWasm0(ret[0], ret[1]).slice();
94
+ wasm.__wbindgen_free(ret[0], ret[1] * 1, 1);
95
+ return v2;
92
96
  }
93
97
 
94
98
  /**
95
99
  * @param {Uint8Array} bytes
96
100
  * @param {CsvMapping} mapping
97
- * @returns {ParsedStdf}
101
+ * @returns {Uint8Array}
98
102
  */
99
103
  export function parse_csv(bytes, mapping) {
100
104
  const ptr0 = passArray8ToWasm0(bytes, wasm.__wbindgen_malloc);
101
105
  const len0 = WASM_VECTOR_LEN;
102
106
  const ret = wasm.parse_csv(ptr0, len0, mapping);
103
- if (ret[2]) {
104
- throw takeFromExternrefTable0(ret[1]);
107
+ if (ret[3]) {
108
+ throw takeFromExternrefTable0(ret[2]);
105
109
  }
106
- return takeFromExternrefTable0(ret[0]);
110
+ var v2 = getArrayU8FromWasm0(ret[0], ret[1]).slice();
111
+ wasm.__wbindgen_free(ret[0], ret[1] * 1, 1);
112
+ return v2;
107
113
  }
108
114
 
109
115
  /**
110
116
  * @param {Uint8Array} bytes
111
117
  * @param {CsvMapping} mapping
112
- * @returns {ParsedStdf}
118
+ * @returns {Uint8Array}
113
119
  */
114
120
  export function parse_json(bytes, mapping) {
115
121
  const ptr0 = passArray8ToWasm0(bytes, wasm.__wbindgen_malloc);
116
122
  const len0 = WASM_VECTOR_LEN;
117
123
  const ret = wasm.parse_json(ptr0, len0, mapping);
118
- if (ret[2]) {
119
- throw takeFromExternrefTable0(ret[1]);
124
+ if (ret[3]) {
125
+ throw takeFromExternrefTable0(ret[2]);
120
126
  }
121
- return takeFromExternrefTable0(ret[0]);
127
+ var v2 = getArrayU8FromWasm0(ret[0], ret[1]).slice();
128
+ wasm.__wbindgen_free(ret[0], ret[1] * 1, 1);
129
+ return v2;
122
130
  }
123
131
 
124
132
  /**
125
133
  * @param {Uint8Array} bytes
126
134
  * @param {CsvMapping} mapping
127
- * @returns {ParsedStdf}
135
+ * @returns {Uint8Array}
128
136
  */
129
137
  export function parse_parquet(bytes, mapping) {
130
138
  const ptr0 = passArray8ToWasm0(bytes, wasm.__wbindgen_malloc);
131
139
  const len0 = WASM_VECTOR_LEN;
132
140
  const ret = wasm.parse_parquet(ptr0, len0, mapping);
133
- if (ret[2]) {
134
- throw takeFromExternrefTable0(ret[1]);
141
+ if (ret[3]) {
142
+ throw takeFromExternrefTable0(ret[2]);
135
143
  }
136
- return takeFromExternrefTable0(ret[0]);
144
+ var v2 = getArrayU8FromWasm0(ret[0], ret[1]).slice();
145
+ wasm.__wbindgen_free(ret[0], ret[1] * 1, 1);
146
+ return v2;
137
147
  }
138
148
 
139
149
  /**
140
150
  * @param {Uint8Array} bytes
141
- * @returns {ParsedStdf}
151
+ * @returns {Uint8Array}
142
152
  */
143
153
  export function parse_stdf(bytes) {
144
154
  const ptr0 = passArray8ToWasm0(bytes, wasm.__wbindgen_malloc);
145
155
  const len0 = WASM_VECTOR_LEN;
146
156
  const ret = wasm.parse_stdf(ptr0, len0);
147
- if (ret[2]) {
148
- throw takeFromExternrefTable0(ret[1]);
157
+ if (ret[3]) {
158
+ throw takeFromExternrefTable0(ret[2]);
149
159
  }
150
- return takeFromExternrefTable0(ret[0]);
160
+ var v2 = getArrayU8FromWasm0(ret[0], ret[1]).slice();
161
+ wasm.__wbindgen_free(ret[0], ret[1] * 1, 1);
162
+ return v2;
151
163
  }
152
164
 
153
165
  /**
154
166
  * @param {Uint8Array} bytes
155
167
  * @param {number[]} selected
156
- * @returns {ParsedStdf}
168
+ * @returns {Uint8Array}
157
169
  */
158
170
  export function parse_stdf_filtered(bytes, selected) {
159
171
  const ptr0 = passArray8ToWasm0(bytes, wasm.__wbindgen_malloc);
160
172
  const len0 = WASM_VECTOR_LEN;
161
173
  const ret = wasm.parse_stdf_filtered(ptr0, len0, selected);
162
- if (ret[2]) {
163
- throw takeFromExternrefTable0(ret[1]);
174
+ if (ret[3]) {
175
+ throw takeFromExternrefTable0(ret[2]);
164
176
  }
165
- return takeFromExternrefTable0(ret[0]);
177
+ var v2 = getArrayU8FromWasm0(ret[0], ret[1]).slice();
178
+ wasm.__wbindgen_free(ret[0], ret[1] * 1, 1);
179
+ return v2;
166
180
  }
167
181
 
168
182
  /**
Binary file