@wafertools/testdata-parser 0.6.0 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  <img src="https://raw.githubusercontent.com/wafertools/tsmap/main/packages/parsers/testdata-parser-readme-header-256.png" width="64" height="64" alt="testdata-parser icon">
4
4
 
5
- Rust/WASM parsers for semiconductor test data formats: **STDF**, **ATDF**, **CSV**, and **JSON**. Compiled to a single WASM module via `wasm-bindgen`; the same Rust source also builds natively (used by [tsmap](https://github.com/wafertools/tsmap)'s Tauri backend).
5
+ Rust/WASM parsers for semiconductor test data formats: **STDF**, **ATDF**, **CSV**, **JSON**, and **Parquet**. Compiled to a single WASM module via `wasm-bindgen`; the same Rust source also builds natively (used by [tsmap](https://github.com/wafertools/tsmap)'s Tauri backend).
6
6
 
7
7
  All formats parse to one shared shape (`ParsedStdf` / `ScanResult`) — there is no format-specific output type on the JS side.
8
8
 
@@ -38,6 +38,8 @@ Every parse function takes raw file bytes (`Uint8Array`) and returns a plain JS
38
38
  | `parse_atdf` | `(bytes: Uint8Array) => ParsedStdf` | Full parse of an ATDF file |
39
39
  | `parse_csv` | `(bytes: Uint8Array, mapping: CsvMapping) => ParsedStdf` | Full parse of a CSV, using an explicit column mapping |
40
40
  | `parse_json` | `(bytes: Uint8Array, mapping: CsvMapping) => ParsedStdf` | Full parse of a JSON array-of-records file, using the same mapping shape as CSV |
41
+ | `parquet_headers` | `(bytes: Uint8Array) => ParquetHeadersResult` | Schema + a sample of rows, for a column-mapping UI (see below — unlike CSV/JSON, this one *is* a WASM export) |
42
+ | `parse_parquet` | `(bytes: Uint8Array, mapping: CsvMapping) => ParsedStdf` | Full parse of a Parquet file, using the same mapping shape as CSV/JSON |
41
43
  | `stdf_test_names` | `(bytes: Uint8Array) => ScanResult` | Fast first-pass scan: test definitions + die count, no die accumulation |
42
44
  | `atdf_test_names` | `(bytes: Uint8Array) => ScanResult` | Same first-pass scan for ATDF |
43
45
  | `parse_stdf_filtered` | `(bytes: Uint8Array, selected: number[]) => ParsedStdf` | Full parse, skipping per-site accumulation for test numbers not in `selected` |
@@ -55,7 +57,7 @@ STDF and ATDF files can be large and contain far more tests than a caller wants
55
57
 
56
58
  ### CsvMapping
57
59
 
58
- `parse_csv` and `parse_json` require an explicit mapping — there's no header auto-detection. Column mapping fields (all are source column names, matched against the file's header row):
60
+ `parse_csv`, `parse_json`, and `parse_parquet` all require an explicit mapping — there's no header auto-detection. Column mapping fields (all are source column names, matched against the file's header row):
59
61
 
60
62
  ```ts
61
63
  interface CsvMapping {
@@ -70,6 +72,7 @@ interface CsvMapping {
70
72
  meta: string[]; // extra columns to surface as generic per-row metadata
71
73
  splitBy: string[]; // columns to additionally facet wafers by (beyond `wafer`)
72
74
  testnameCol?: string; // for "tall" CSVs: column holding the test name per row
75
+ testnumberCol?: string; // for "tall" CSVs: column holding the test's real number per row
73
76
  testvalueCol?: string; // for "tall" CSVs: column holding the test value per row
74
77
  loLimitCol?: string;
75
78
  hiLimitCol?: string;
@@ -84,7 +87,12 @@ interface CsvTestCol {
84
87
  }
85
88
  ```
86
89
 
87
- Two ways to describe test columns are supported: a **fixed set** of `tests` (one column per test, "wide" format), or a **tall** layout (`testnameCol`/`testvalueCol` — one row per die×test, with the test identity read from a column rather than the header).
90
+ Two ways to describe test columns are supported: a **fixed set** of `tests` (one column per test, "wide" format), or a **tall** layout (`testnameCol`/`testnumberCol`/`testvalueCol` — one row per die×test, with the test identity read from a column rather than the header).
91
+
92
+ **Test identity — real number vs. synthesized one.** Neither format has a mandatory real STDF-style test number, so one gets synthesized by default (see "Design notes" below) — but a caller that *does* have real numbers in the source data shouldn't lose them:
93
+
94
+ - **Tall layout**: `testnameCol` and `testnumberCol` are independent — set either alone, or both together. Number alone is a legitimate, fully-supported case (some exports carry only a numeric test ID, no descriptive name) — the test's display name then falls back to the number itself, stringified. Name alone keeps the pre-existing hash-based behavior. Both together: the real number from `testnumberCol` is used as the key (not hashed), paired with the given name — this is the common "I have both and want them both honoured" case. Whichever columns are set, at least one of `testnameCol`/`testnumberCol` plus `testvalueCol` is required to trigger tall-layout parsing at all.
95
+ - **Wide layout** (`tests: CsvTestCol[]`): `testNumber` is assigned by the caller building the mapping (tsmap's `mappingUI.ts` does this before calling in), not by this crate — but the same principle applies there: if a column's own header is itself a bare number (a real-world convention — columns literally named `1001`, `1002`), that number should be used directly rather than hashed. `test_identity`'s reserved band (below) exists specifically so a real number like this can never collide with a hashed one.
88
96
 
89
97
  ### Return shape — `ParsedStdf`
90
98
 
@@ -153,9 +161,20 @@ interface ScanResult {
153
161
  }
154
162
  ```
155
163
 
156
- ### Column headers are not a WASM export
164
+ ### Column headers: CSV/JSON vs Parquet
165
+
166
+ There is no `csv_headers`/`json_headers` in the WASM API — a browser caller that needs to show the user a column-mapping UI before parsing a CSV/JSON file can just read the header row itself in plain JS (this is what tsmap's web build does; the desktop build calls the native functions below instead). The byte-based Rust functions exist (`csv_headers_from_bytes`, and `json_headers_sync`'s logic), they are simply not wired through `wasm-bindgen` for these two formats.
167
+
168
+ **Parquet is different: `parquet_headers` *is* a WASM export.** Its binary, footer-based schema has no equivalent plain-JS shortcut — a caller genuinely needs the parser to read it. `ParquetHeadersResult` extends the same headers/sample/rowCount shape with `columnTypes` (coarse `"number" | "bool" | "string"` per column, inferred from the first sampled row), since Parquet's columns are natively typed unlike CSV/JSON's all-text cells:
157
169
 
158
- There is no `csv_headers`/`json_headers` in the WASM API — a browser caller that needs to show the user a column-mapping UI before parsing has to read the header row itself in JS (this is what tsmap's web build does; the desktop build calls the native functions below). The byte-based Rust functions exist (`csv_headers_from_bytes`, and `json_headers_sync`'s logic), they are simply not wired through `wasm-bindgen` yet.
170
+ ```ts
171
+ interface ParquetHeadersResult {
172
+ headers: string[];
173
+ sample: Record<string, string>[];
174
+ rowCount: number;
175
+ columnTypes: Record<string, 'number' | 'bool' | 'string'>;
176
+ }
177
+ ```
159
178
 
160
179
  ## Native (non-WASM) usage
161
180
 
@@ -171,6 +190,8 @@ The crate also builds as a native Rust library (used directly by tsmap's Tauri c
171
190
  | `parse_csv_inner(path: String, mapping: CsvMapping) -> Result<ParsedStdf, String>` | `parse_csv` |
172
191
  | `json_headers_sync(path: String) -> Result<JsonHeadersResult, String>` | `parse_json` |
173
192
  | `parse_json_sync(path: String, mapping: CsvMapping) -> Result<ParsedStdf, String>` | `parse_json` |
193
+ | `parquet_headers_inner(path: String) -> Result<ParquetHeadersResult, String>` | `parse_parquet` |
194
+ | `parse_parquet_inner(path: String, mapping: CsvMapping) -> Result<ParsedStdf, String>` | `parse_parquet` |
174
195
  | `read_bytes(path: &str) -> Result<Vec<u8>, String>` | `read_file` |
175
196
  | `read_text(path: &str) -> Result<String, String>` | `read_file` |
176
197
 
@@ -187,6 +208,8 @@ The crate also builds as a native Rust library (used directly by tsmap's Tauri c
187
208
  | `csv_headers_from_bytes(&[u8]) -> Result<CsvHeadersResult, String>` | `parse_csv` |
188
209
  | `parse_csv_from_bytes(&[u8], mapping: CsvMapping) -> Result<ParsedStdf, String>` | `parse_csv` |
189
210
  | `parse_json_from_bytes(&[u8], mapping: CsvMapping) -> Result<ParsedStdf, String>` | `parse_json` |
211
+ | `parquet_headers_from_bytes(&[u8]) -> Result<ParquetHeadersResult, String>` | `parse_parquet` |
212
+ | `parse_parquet_from_bytes(&[u8], mapping: CsvMapping) -> Result<ParsedStdf, String>` | `parse_parquet` |
190
213
  | `decompress_if_gzip(Vec<u8>) -> Result<Vec<u8>, String>` | `read_file` |
191
214
 
192
215
  `CsvHeadersResult` and `JsonHeadersResult` are the same shape — the header row plus enough of the file to preview a mapping:
@@ -204,7 +227,9 @@ pub struct CsvHeadersResult {
204
227
  - **Byte readers are panic-free.** STDF/ATDF field readers are bounds-checked and return `Option`/`Result` rather than panicking on truncated input — a panic inside WASM aborts the whole module with no recovery, so this is a hard requirement, not a style preference.
205
228
  - **Big-endian and little-endian STDF** are both supported (detected from the FAR record's `CPU_TYPE`).
206
229
  - **Gzip is transparent** — every entry point decompresses `.gz` input automatically by sniffing the magic bytes; callers don't need to branch on compression.
207
- - **CSV/JSON test numbers are a deterministic hash, not a real STDF test number.** STDF/ATDF have a real test number in the file; CSV/JSON don't, so one is synthesized — from the source column for wide format, from the test name for long format (`test_identity::stable_test_number`, FNV-1a with a fixed seed and a reserved floor, collision-probed so two tests in one file can never collide). Deliberately not sequential/encounter-order: a hash means the number for a given test doesn't change if the file is reordered or a column is added — the number is otherwise meaningless and callers should never rely on its value, only on it being stable and unique within one parse. `order` (see `TestDef` above) carries the file's own display order instead.
230
+ - **CSV/JSON/Parquet test numbers are a deterministic hash, not a real STDF test number.** STDF/ATDF have a real test number in the file; the other three don't, so one is synthesized — from the source column for wide format, from the test name for long format (`test_identity::stable_test_number`, FNV-1a with a fixed seed and a reserved floor, collision-probed so two tests in one file can never collide). Deliberately not sequential/encounter-order: a hash means the number for a given test doesn't change if the file is reordered or a column is added — the number is otherwise meaningless and callers should never rely on its value, only on it being stable and unique within one parse. `order` (see `TestDef` above) carries the file's own display order instead.
231
+ - **Parquet reads through a row-oriented API, not Arrow.** `parquet::record::Row`/`Field` rather than the `arrow` feature — a closer fit for this crate's row-based `DieResult` model, and a smaller WASM bundle (no Arrow array machinery pulled in). A typed Parquet cell is coerced to `f64` for numeric roles and to a plain string otherwise; a value that fails to coerce (e.g. a numeric role mapped to a genuinely string-typed column) is skipped and surfaced as one summarised entry in `warnings`, not a panic or a silent zero.
232
+ - **Parquet's `zstd` codec is native-only.** `snappy`, `gzip`, `lz4`, and `brotli` build for `wasm32-unknown-unknown` with no extra toolchain; `zstd`'s C library needs a real C cross-compiler targeting wasm32, which a plain `wasm-pack build` doesn't assume is available. A `zstd`-compressed Parquet file parses natively but fails clearly on the WASM build.
208
233
 
209
234
  ## Versioning
210
235
 
package/package.json CHANGED
@@ -1,8 +1,8 @@
1
1
  {
2
2
  "name": "@wafertools/testdata-parser",
3
3
  "type": "module",
4
- "description": "Rust/WASM parsers for semiconductor test data formats (STDF, ATDF, CSV, JSON)",
5
- "version": "0.6.0",
4
+ "description": "Rust/WASM parsers for semiconductor test data formats (STDF, ATDF, CSV, JSON, Parquet)",
5
+ "version": "0.7.0",
6
6
  "license": "MIT",
7
7
  "repository": {
8
8
  "type": "git",
@@ -5,6 +5,8 @@ export function atdf_test_names(bytes: Uint8Array): any;
5
5
 
6
6
  export function init(): void;
7
7
 
8
+ export function parquet_headers(bytes: Uint8Array): any;
9
+
8
10
  export function parse_atdf(bytes: Uint8Array): any;
9
11
 
10
12
  export function parse_atdf_filtered(bytes: Uint8Array, selected: any): any;
@@ -13,6 +15,8 @@ export function parse_csv(bytes: Uint8Array, mapping: any): any;
13
15
 
14
16
  export function parse_json(bytes: Uint8Array, mapping: any): any;
15
17
 
18
+ export function parse_parquet(bytes: Uint8Array, mapping: any): any;
19
+
16
20
  export function parse_stdf(bytes: Uint8Array): any;
17
21
 
18
22
  export function parse_stdf_filtered(bytes: Uint8Array, selected: any): any;
@@ -25,10 +29,12 @@ export interface InitOutput {
25
29
  readonly memory: WebAssembly.Memory;
26
30
  readonly atdf_test_names: (a: number, b: number) => [number, number, number];
27
31
  readonly init: () => void;
32
+ readonly parquet_headers: (a: number, b: number) => [number, number, number];
28
33
  readonly parse_atdf: (a: number, b: number) => [number, number, number];
29
34
  readonly parse_atdf_filtered: (a: number, b: number, c: any) => [number, number, number];
30
35
  readonly parse_csv: (a: number, b: number, c: any) => [number, number, number];
31
36
  readonly parse_json: (a: number, b: number, c: any) => [number, number, number];
37
+ readonly parse_parquet: (a: number, b: number, c: any) => [number, number, number];
32
38
  readonly parse_stdf: (a: number, b: number) => [number, number, number];
33
39
  readonly parse_stdf_filtered: (a: number, b: number, c: any) => [number, number, number];
34
40
  readonly stdf_test_names: (a: number, b: number) => [number, number, number];
@@ -18,6 +18,20 @@ export function init() {
18
18
  wasm.init();
19
19
  }
20
20
 
21
+ /**
22
+ * @param {Uint8Array} bytes
23
+ * @returns {any}
24
+ */
25
+ export function parquet_headers(bytes) {
26
+ const ptr0 = passArray8ToWasm0(bytes, wasm.__wbindgen_malloc);
27
+ const len0 = WASM_VECTOR_LEN;
28
+ const ret = wasm.parquet_headers(ptr0, len0);
29
+ if (ret[2]) {
30
+ throw takeFromExternrefTable0(ret[1]);
31
+ }
32
+ return takeFromExternrefTable0(ret[0]);
33
+ }
34
+
21
35
  /**
22
36
  * @param {Uint8Array} bytes
23
37
  * @returns {any}
@@ -77,6 +91,21 @@ export function parse_json(bytes, mapping) {
77
91
  return takeFromExternrefTable0(ret[0]);
78
92
  }
79
93
 
94
+ /**
95
+ * @param {Uint8Array} bytes
96
+ * @param {any} mapping
97
+ * @returns {any}
98
+ */
99
+ export function parse_parquet(bytes, mapping) {
100
+ const ptr0 = passArray8ToWasm0(bytes, wasm.__wbindgen_malloc);
101
+ const len0 = WASM_VECTOR_LEN;
102
+ const ret = wasm.parse_parquet(ptr0, len0, mapping);
103
+ if (ret[2]) {
104
+ throw takeFromExternrefTable0(ret[1]);
105
+ }
106
+ return takeFromExternrefTable0(ret[0]);
107
+ }
108
+
80
109
  /**
81
110
  * @param {Uint8Array} bytes
82
111
  * @returns {any}
@@ -316,6 +345,11 @@ function __wbg_get_imports() {
316
345
  const ret = getStringFromWasm0(arg0, arg1);
317
346
  return ret;
318
347
  },
348
+ __wbindgen_cast_0000000000000003: function(arg0) {
349
+ // Cast intrinsic for `U64 -> Externref`.
350
+ const ret = BigInt.asUintN(64, arg0);
351
+ return ret;
352
+ },
319
353
  __wbindgen_init_externref_table: function() {
320
354
  const table = wasm.__wbindgen_externrefs;
321
355
  const offset = table.grow(4);
Binary file