@wafertools/testdata-parser 0.5.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,5 +1,7 @@
1
1
  # @wafertools/testdata-parser
2
2
 
3
+ <img src="https://raw.githubusercontent.com/wafertools/tsmap/main/packages/parsers/testdata-parser-readme-header-256.png" width="64" height="64" alt="testdata-parser icon">
4
+
3
5
  Rust/WASM parsers for semiconductor test data formats: **STDF**, **ATDF**, **CSV**, and **JSON**. Compiled to a single WASM module via `wasm-bindgen`; the same Rust source also builds natively (used by [tsmap](https://github.com/wafertools/tsmap)'s Tauri backend).
4
6
 
5
7
  All formats parse to one shared shape (`ParsedStdf` / `ScanResult`) — there is no format-specific output type on the JS side.
@@ -120,6 +122,7 @@ interface TestDef {
120
122
  loLimit?: number;
121
123
  hiLimit?: number;
122
124
  units?: string;
125
+ order?: number; // display order, independent of the key — see below
123
126
  }
124
127
 
125
128
  interface LotMeta {
@@ -150,25 +153,58 @@ interface ScanResult {
150
153
  }
151
154
  ```
152
155
 
156
+ ### Column headers are not a WASM export
157
+
158
+ There is no `csv_headers`/`json_headers` in the WASM API — a browser caller that needs to show the user a column-mapping UI before parsing has to read the header row itself in JS (this is what tsmap's web build does; the desktop build calls the native functions below). The byte-based Rust functions exist (`csv_headers_from_bytes`, and `json_headers_sync`'s logic), they are simply not wired through `wasm-bindgen` yet.
159
+
153
160
  ## Native (non-WASM) usage
154
161
 
155
- The crate also builds as a native Rust library (used directly by tsmap's Tauri commands, bypassing WASM entirely). Native-only entry points read from a file path instead of a byte buffer and are synchronous:
162
+ The crate also builds as a native Rust library (used directly by tsmap's Tauri commands, bypassing WASM entirely). Enable the `native` feature (the default); the `wasm` feature gates the `wasm-bindgen` exports above. See `Cargo.toml` for the full feature list, including `bench` (enables `parse_stdf_from_bytes_timed`, a timed parse variant used by the perf benchmarks).
163
+
164
+ **Path-based** — `native` feature only, synchronous, read from a file path rather than a byte buffer:
156
165
 
157
166
  | Function | Module |
158
167
  | --- | --- |
159
168
  | `parse_stdf_sync(path: String) -> Result<ParsedStdf, String>` | `parse_stdf` |
160
169
  | `parse_atdf_sync(path: String) -> Result<ParsedStdf, String>` | `parse_atdf` |
161
- | `csv_headers_inner(path) -> Result<CsvHeadersResult, String>` | `parse_csv` |
162
- | `parse_csv_inner(path, mapping) -> Result<ParsedStdf, String>` | `parse_csv` |
163
- | `json_headers_sync(path) -> Result<Vec<String>, String>` | `parse_json` |
170
+ | `csv_headers_inner(path: String) -> Result<CsvHeadersResult, String>` | `parse_csv` |
171
+ | `parse_csv_inner(path: String, mapping: CsvMapping) -> Result<ParsedStdf, String>` | `parse_csv` |
172
+ | `json_headers_sync(path: String) -> Result<JsonHeadersResult, String>` | `parse_json` |
173
+ | `parse_json_sync(path: String, mapping: CsvMapping) -> Result<ParsedStdf, String>` | `parse_json` |
174
+ | `read_bytes(path: &str) -> Result<Vec<u8>, String>` | `read_file` |
175
+ | `read_text(path: &str) -> Result<String, String>` | `read_file` |
176
+
177
+ **Byte-based** — available on every target, and what the WASM exports wrap. Use these from Rust when you already hold the bytes:
164
178
 
165
- Enable the `native` feature (default) for these; the `wasm` feature gates the `wasm-bindgen` exports above. See `Cargo.toml` for the full feature list, including `bench` (enables a timed parse variant used by the perf benchmarks).
179
+ | Function | Module |
180
+ | --- | --- |
181
+ | `parse_stdf_from_bytes(&[u8]) -> Result<ParsedStdf, String>` | `parse_stdf` |
182
+ | `parse_atdf_from_bytes(&[u8]) -> Result<ParsedStdf, String>` | `parse_atdf` |
183
+ | `parse_stdf_test_names(&[u8]) -> Result<ScanResult, String>` | `parse_stdf` |
184
+ | `parse_atdf_test_names(&[u8]) -> Result<ScanResult, String>` | `parse_atdf` |
185
+ | `parse_stdf_from_bytes_filtered(&[u8], &HashSet<u32>) -> Result<ParsedStdf, String>` | `parse_stdf` |
186
+ | `parse_atdf_from_bytes_filtered(&[u8], &HashSet<u32>) -> Result<ParsedStdf, String>` | `parse_atdf` |
187
+ | `csv_headers_from_bytes(&[u8]) -> Result<CsvHeadersResult, String>` | `parse_csv` |
188
+ | `parse_csv_from_bytes(&[u8], mapping: CsvMapping) -> Result<ParsedStdf, String>` | `parse_csv` |
189
+ | `parse_json_from_bytes(&[u8], mapping: CsvMapping) -> Result<ParsedStdf, String>` | `parse_json` |
190
+ | `decompress_if_gzip(Vec<u8>) -> Result<Vec<u8>, String>` | `read_file` |
191
+
192
+ `CsvHeadersResult` and `JsonHeadersResult` are the same shape — the header row plus enough of the file to preview a mapping:
193
+
194
+ ```rust
195
+ pub struct CsvHeadersResult {
196
+ pub headers: Vec<String>,
197
+ pub sample: Vec<HashMap<String, String>>, // first few rows, for a preview UI
198
+ pub row_count: usize,
199
+ }
200
+ ```
166
201
 
167
202
  ## Design notes
168
203
 
169
204
  - **Byte readers are panic-free.** STDF/ATDF field readers are bounds-checked and return `Option`/`Result` rather than panicking on truncated input — a panic inside WASM aborts the whole module with no recovery, so this is a hard requirement, not a style preference.
170
205
  - **Big-endian and little-endian STDF** are both supported (detected from the FAR record's `CPU_TYPE`).
171
206
  - **Gzip is transparent** — every entry point decompresses `.gz` input automatically by sniffing the magic bytes; callers don't need to branch on compression.
207
+ - **CSV/JSON test numbers are a deterministic hash, not a real STDF test number.** STDF/ATDF have a real test number in the file; CSV/JSON don't, so one is synthesized — from the source column for wide format, from the test name for long format (`test_identity::stable_test_number`, FNV-1a with a fixed seed and a reserved floor, collision-probed so two tests in one file can never collide). Deliberately not sequential/encounter-order: a hash means the number for a given test doesn't change if the file is reordered or a column is added — the number is otherwise meaningless and callers should never rely on its value, only on it being stable and unique within one parse. `order` (see `TestDef` above) carries the file's own display order instead.
172
208
 
173
209
  ## Versioning
174
210
 
package/package.json CHANGED
@@ -2,7 +2,7 @@
2
2
  "name": "@wafertools/testdata-parser",
3
3
  "type": "module",
4
4
  "description": "Rust/WASM parsers for semiconductor test data formats (STDF, ATDF, CSV, JSON)",
5
- "version": "0.5.0",
5
+ "version": "0.6.0",
6
6
  "license": "MIT",
7
7
  "repository": {
8
8
  "type": "git",
Binary file