splitgrid 0.1.0__tar.gz → 0.1.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (25) hide show
  1. {splitgrid-0.1.0 → splitgrid-0.1.2}/Makefile +14 -1
  2. {splitgrid-0.1.0/src/splitgrid.egg-info → splitgrid-0.1.2}/PKG-INFO +58 -21
  3. splitgrid-0.1.2/README.md +144 -0
  4. {splitgrid-0.1.0 → splitgrid-0.1.2}/pyproject.toml +1 -1
  5. {splitgrid-0.1.0 → splitgrid-0.1.2}/setup.py +3 -2
  6. {splitgrid-0.1.0 → splitgrid-0.1.2}/src/splitgrid/__init__.py +1 -1
  7. {splitgrid-0.1.0 → splitgrid-0.1.2}/src/splitgrid/codec.py +14 -1
  8. {splitgrid-0.1.0 → splitgrid-0.1.2}/src/splitgrid/pack.c +2 -0
  9. {splitgrid-0.1.0 → splitgrid-0.1.2/src/splitgrid.egg-info}/PKG-INFO +58 -21
  10. {splitgrid-0.1.0 → splitgrid-0.1.2}/tests/test_payload_codec.py +39 -0
  11. splitgrid-0.1.0/README.md +0 -107
  12. {splitgrid-0.1.0 → splitgrid-0.1.2}/LICENSE +0 -0
  13. {splitgrid-0.1.0 → splitgrid-0.1.2}/MANIFEST.in +0 -0
  14. {splitgrid-0.1.0 → splitgrid-0.1.2}/setup.cfg +0 -0
  15. {splitgrid-0.1.0 → splitgrid-0.1.2}/src/splitgrid/deal_shim.py +0 -0
  16. {splitgrid-0.1.0 → splitgrid-0.1.2}/src/splitgrid/pack.pyx +0 -0
  17. {splitgrid-0.1.0 → splitgrid-0.1.2}/src/splitgrid.egg-info/SOURCES.txt +0 -0
  18. {splitgrid-0.1.0 → splitgrid-0.1.2}/src/splitgrid.egg-info/dependency_links.txt +0 -0
  19. {splitgrid-0.1.0 → splitgrid-0.1.2}/src/splitgrid.egg-info/requires.txt +0 -0
  20. {splitgrid-0.1.0 → splitgrid-0.1.2}/src/splitgrid.egg-info/top_level.txt +0 -0
  21. {splitgrid-0.1.0 → splitgrid-0.1.2}/tests/test_cython_parity.py +0 -0
  22. {splitgrid-0.1.0 → splitgrid-0.1.2}/tests/test_payload_codec_policy_verification.py +0 -0
  23. {splitgrid-0.1.0 → splitgrid-0.1.2}/tests/test_serialization_ab.py +0 -0
  24. {splitgrid-0.1.0 → splitgrid-0.1.2}/tests/test_serialization_ab_support.py +0 -0
  25. {splitgrid-0.1.0 → splitgrid-0.1.2}/tests/test_serialization_verification.py +0 -0
@@ -1,6 +1,6 @@
1
1
  # SplitGrid
2
2
 
3
- .PHONY: test test-all test-slow verify install-test
3
+ .PHONY: test test-all test-slow verify install-test bench bench-child profile-pack
4
4
 
5
5
  # Prefer a local venv so `make test` / `make verify` work without activating it.
6
6
  PYTEST := $(if $(wildcard .venv/bin/pytest),.venv/bin/pytest,python3 -m pytest)
@@ -22,3 +22,16 @@ test-slow:
22
22
  # deal contracts + hypothesis round-trips (excludes CrossHair subprocess).
23
23
  verify:
24
24
  $(PYTEST) tests/test_serialization_verification.py tests/test_payload_codec_policy_verification.py -m "not slow" -q
25
+
26
+ PYTHON := $(if $(wildcard .venv/bin/python),.venv/bin/python,python3)
27
+
28
+ # Asymmetric host/child serialization bench (needs numpy).
29
+ bench:
30
+ $(PYTHON) scripts/bench_serialization.py --direction both
31
+
32
+ bench-child:
33
+ $(PYTHON) scripts/bench_serialization.py --child-only
34
+
35
+ profile-pack:
36
+ $(PYTHON) scripts/profile_pack.py
37
+ $(PYTHON) scripts/profile_nones.py
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: splitgrid
3
- Version: 0.1.0
3
+ Version: 0.1.2
4
4
  Summary: Asymmetric split-grid serialization: host stdlib flatten + optional Cython, child NumPy unpack
5
5
  Author-email: Keith Cu <keithcu@gmail.com>
6
6
  License: GPL-3.0-or-later
@@ -35,23 +35,33 @@ Dynamic: license-file
35
35
 
36
36
  # SplitGrid
37
37
 
38
+ [![PyPI](https://img.shields.io/pypi/v/splitgrid.svg)](https://pypi.org/project/splitgrid/)
39
+
38
40
  Asymmetric **split-grid** serialization for rectangular numeric and mixed-type grids.
39
41
 
40
- This package was **pulled out of [WriterAgent](https://github.com/KeithCu/writeragent)** (LibreOffice Writer/Calc/Draw AI extension). It is the same host flatten + child unpack codec that lived in `plugin.scripting.payload_codec`, plus the optional Cython flatten accelerator from `native/writeragent_vec`. WriterAgent can later depend on `splitgrid` without changing the wire dict.
42
+ This package was **pulled out of [WriterAgent](https://github.com/KeithCu/writeragent)** (LibreOffice Writer/Calc/Draw AI extension). It is the host flatten + child unpack codec that lived in `plugin.scripting.payload_codec`, plus the optional Cython flatten accelerator from `native/writeragent_vec`.
43
+
44
+ ## Why this exists
45
+
46
+ A spreadsheet host often cannot import NumPy: LibreOffice’s embedded Python is a different interpreter from the user’s venv, and loading the user’s C extensions into the host is an ABI footgun. Heavy compute therefore runs in a **child** process that *does* have NumPy. The range still has to cross that boundary.
47
+
48
+ On the wire, a Calc-style range is a nested list: `list[list[float | int | str | None]]`. Standard pickle walks **one heap object per cell**. On a 20,000 × 5 grid that is ~12 ms for dump + load. Pack the same numbers into a contiguous `float64` buffer and Pickle Protocol 5 moves the bytes in ~0.015 ms; the child can `np.frombuffer` in ~0.002 ms.
49
+
50
+ Almost all of the remaining time is **host flatten** — turning nested Python objects into that buffer without NumPy (~8.3 ms pure Python on that shape; ~3 ms with the Cython helper). SplitGrid is that flatten/unpack. Length-prefixed Pickle 5 framing stays in the application; this package is the codec only.
41
51
 
42
- Repo: [github.com/KeithCu/writeragent](https://github.com/KeithCu/writeragent)
52
+ Column-wise blobs and JSON + Base64 were tried first. Transposing columns on the host builds extra object graphs; Base64 and per-column reconstructs lose the `frombuffer` win. One row-major float64 buffer plus a sparse string map is faster and simpler.
43
53
 
44
- ## What split-grid is
54
+ ## How it works
45
55
 
46
- The compute path is **asymmetric by design**:
56
+ The path is **asymmetric**:
47
57
 
48
- - **Host pack** (LibreOffice’s embedded Python, or any NumPy-free interpreter) flattens a 1D or 2D grid into a contiguous `float64` buffer plus a sparse **integer-keyed** `strings` map. Empty cells become `NaN`. Zip codes like `"02138"` stay strings — they are never parsed as floats.
49
- - **Child unpack** (venv with NumPy) materializes a **pure-numeric** grid (`strings == {}`) with `np.frombuffer` (ndarray). Mixed grids use a vectorized object-masking path and return nested lists, restoring `None` for NaN holes.
50
- - **Host unpack** preserves `float('nan')` from the buffer (does **not** coerce holes to `None`). That is the locked egress policy: NaN becomes a Calc error, not a silent blank.
58
+ - **Host pack** (stdlib only) flattens a 1D or 2D rectangular grid into a contiguous `float64` buffer plus a sparse **integer-keyed** `strings` map. Empty cells become `NaN`. Zip codes like `"02138"` stay strings — they are never parsed as floats.
59
+ - **Child unpack** (NumPy) materializes a **pure-numeric** grid (`strings == {}`) with `np.frombuffer` (ndarray). Mixed grids use a vectorized object-masking path and return nested lists, restoring `None` for NaN holes.
60
+ - **Host unpack** preserves `float('nan')` from the buffer (does **not** coerce holes to `None`). That is the locked egress policy: NaN becomes a spreadsheet error, not a silent blank.
51
61
 
52
62
  Grids with fewer than **100 cells** (`BINARY_MIN_CELLS`) stay nested Python lists. Force `"always"` / `"never"` overrides that threshold (used by A/B tests).
53
63
 
54
- Wire envelope (Pickle5-friendly dict, same keys as WriterAgent):
64
+ Wire envelope (Pickle5-friendly dict):
55
65
 
56
66
  ```python
57
67
  {
@@ -64,9 +74,30 @@ Wire envelope (Pickle5-friendly dict, same keys as WriterAgent):
64
74
  }
65
75
  ```
66
76
 
77
+ | Cell value | `buffer` (float64) | `strings` |
78
+ |------------|--------------------|-----------|
79
+ | `None` (empty cell) | `NaN` | — |
80
+ | `int` / `float` | numeric value | — |
81
+ | `bool` | `0.0` / `1.0` | — |
82
+ | `str` (including `"02138"`) | `NaN` | text by flat index |
83
+
67
84
  There is **no datetime lane** on the float64 buffer. Python `datetime` objects stringify into `strings`. Do not add a `'date'` column kind.
68
85
 
69
- Jagged 2D grids raise `ValueError` (Calc ranges are rectangular).
86
+ Jagged 2D grids raise `ValueError` (spreadsheet ranges are rectangular).
87
+
88
+ ## Numbers
89
+
90
+ Median timings from an asymmetric bench (host = stdlib pack; child = deserialize + materialize).
91
+
92
+ Ingress, **20,000 × 5** (100k cells):
93
+
94
+ | Format | Pack | Dump | Load | Materialize | Total |
95
+ |--------|------|------|------|-------------|-------|
96
+ | JSON nested lists | 6.2 ms | 22.8 ms | 14.6 ms | 2.1 ms | 45.7 ms |
97
+ | Pickle 5 nested lists | 6.1 ms | 1.4 ms | 2.9 ms | 2.0 ms | 12.4 ms |
98
+ | **Pickle 5 + split-grid** | 8.3 ms | **0.013 ms** | **0.015 ms** | **0.002 ms** | **8.3 ms** |
99
+
100
+ Child materialize of a **100 × 100** numeric grid: nested-list pickle then `np.array` ~0.6 ms; split-grid `frombuffer` ~**0.016 ms**. Below 100 cells the envelope is not worth it — that is why `BINARY_MIN_CELLS` exists. Host pack still dominates large ingress; Cython only speeds that loop.
70
101
 
71
102
  ## Python vs Cython
72
103
 
@@ -78,7 +109,7 @@ Host flatten is an optimized **pure-Python** loop:
78
109
  - lazy column-state upgrades
79
110
  - rectangular validation before the hot loop
80
111
 
81
- The optional **Cython** module `splitgrid.pack` exposes `fast_flatten_grid_1d` / `fast_flatten_grid_2d`. It is loaded dynamically and **canary-tested** at import. If the extension is missing or fails the canary, the codec uses pure Python. Importing `splitgrid` never requires a compiler.
112
+ The optional **Cython** module `splitgrid.pack` exposes `fast_flatten_grid_1d` / `fast_flatten_grid_2d`. It is loaded dynamically and **canary-tested** at import. If the extension is missing or fails the canary, the codec uses pure Python. Importing `splitgrid` never requires a compiler. Extension builds use release flags (`-O3 -DNDEBUG -g0` on Unix, `/O2 /DNDEBUG` on Windows).
82
113
 
83
114
  ## Install
84
115
 
@@ -108,7 +139,21 @@ make verify # deal contracts + hypothesis round-trips, no CrossHair
108
139
 
109
140
  Default pytest **does not require** the compiled extension. When `splitgrid.pack` is built, `tests/test_cython_parity.py` compares Cython and Python flatten components on the same grids.
110
141
 
111
- ## Public API (WriterAgent-compatible)
142
+ ## Bench
143
+
144
+ Requires NumPy (and the optional Cython extension if you want the accelerated pack row). From a source checkout:
145
+
146
+ ```bash
147
+ python scripts/bench_serialization.py --direction both
148
+ python scripts/bench_serialization.py --child-only
149
+ python scripts/bench_unpacking_opt.py
150
+ python scripts/profile_pack.py
151
+ python scripts/profile_nones.py
152
+ python scripts/run_serialization_ab.py --all
153
+ make bench
154
+ ```
155
+
156
+ ## Public API
112
157
 
113
158
  ```python
114
159
  from splitgrid import (
@@ -129,14 +174,6 @@ from splitgrid import (
129
174
 
130
175
  `host_pack_data(..., force="auto"|"always"|"never")` chooses split-grid vs nested list. `host_pack_multi_data` is a thin `multi_data` wrapper over the same per-grid packing.
131
176
 
132
- ## How WriterAgent will consume this
133
-
134
- In [WriterAgent](https://github.com/KeithCu/writeragent), replace `from plugin.scripting.payload_codec import host_pack_data, ...` with `from splitgrid import host_pack_data, ...`. The envelope tag stays `"split_grid"`, dtype `"float64"`, integer-keyed `strings`, and `column_kinds` `int`/`float`/`bool`. Pickle Protocol 5 framing stays in WriterAgent’s `ipc.py` — this package is the codec only.
135
-
136
- ## Releasing
137
-
138
- Tags matching `v*` (for example `v0.1.0`) run `.github/workflows/publish.yml`: cibuildwheel + sdist, then trusted publishing to PyPI (GitHub environment `pypi`, no API token). The workflow file must already be on `main` before you push the tag.
139
-
140
177
  ## License
141
178
 
142
- GPL-3.0-or-later (same as WriterAgent).
179
+ GPL-3.0-or-later.
@@ -0,0 +1,144 @@
1
+ # SplitGrid
2
+
3
+ [![PyPI](https://img.shields.io/pypi/v/splitgrid.svg)](https://pypi.org/project/splitgrid/)
4
+
5
+ Asymmetric **split-grid** serialization for rectangular numeric and mixed-type grids.
6
+
7
+ This package was **pulled out of [WriterAgent](https://github.com/KeithCu/writeragent)** (LibreOffice Writer/Calc/Draw AI extension). It is the host flatten + child unpack codec that lived in `plugin.scripting.payload_codec`, plus the optional Cython flatten accelerator from `native/writeragent_vec`.
8
+
9
+ ## Why this exists
10
+
11
+ A spreadsheet host often cannot import NumPy: LibreOffice’s embedded Python is a different interpreter from the user’s venv, and loading the user’s C extensions into the host is an ABI footgun. Heavy compute therefore runs in a **child** process that *does* have NumPy. The range still has to cross that boundary.
12
+
13
+ On the wire, a Calc-style range is a nested list: `list[list[float | int | str | None]]`. Standard pickle walks **one heap object per cell**. On a 20,000 × 5 grid that is ~12 ms for dump + load. Pack the same numbers into a contiguous `float64` buffer and Pickle Protocol 5 moves the bytes in ~0.015 ms; the child can `np.frombuffer` in ~0.002 ms.
14
+
15
+ Almost all of the remaining time is **host flatten** — turning nested Python objects into that buffer without NumPy (~8.3 ms pure Python on that shape; ~3 ms with the Cython helper). SplitGrid is that flatten/unpack. Length-prefixed Pickle 5 framing stays in the application; this package is the codec only.
16
+
17
+ Column-wise blobs and JSON + Base64 were tried first. Transposing columns on the host builds extra object graphs; Base64 and per-column reconstructs lose the `frombuffer` win. One row-major float64 buffer plus a sparse string map is faster and simpler.
18
+
19
+ ## How it works
20
+
21
+ The path is **asymmetric**:
22
+
23
+ - **Host pack** (stdlib only) flattens a 1D or 2D rectangular grid into a contiguous `float64` buffer plus a sparse **integer-keyed** `strings` map. Empty cells become `NaN`. Zip codes like `"02138"` stay strings — they are never parsed as floats.
24
+ - **Child unpack** (NumPy) materializes a **pure-numeric** grid (`strings == {}`) with `np.frombuffer` (ndarray). Mixed grids use a vectorized object-masking path and return nested lists, restoring `None` for NaN holes.
25
+ - **Host unpack** preserves `float('nan')` from the buffer (does **not** coerce holes to `None`). That is the locked egress policy: NaN becomes a spreadsheet error, not a silent blank.
26
+
27
+ Grids with fewer than **100 cells** (`BINARY_MIN_CELLS`) stay nested Python lists. Force `"always"` / `"never"` overrides that threshold (used by A/B tests).
28
+
29
+ Wire envelope (Pickle5-friendly dict):
30
+
31
+ ```python
32
+ {
33
+ "__wa_payload__": "split_grid",
34
+ "dtype": "float64",
35
+ "column_kinds": ["int", "float"], # per column: int / float / bool
36
+ "shape": [rows, cols], # or [n] for 1D
37
+ "buffer": b"...", # row-major float64 bytes
38
+ "strings": {7: "banana"}, # integer keys, not str(idx)
39
+ }
40
+ ```
41
+
42
+ | Cell value | `buffer` (float64) | `strings` |
43
+ |------------|--------------------|-----------|
44
+ | `None` (empty cell) | `NaN` | — |
45
+ | `int` / `float` | numeric value | — |
46
+ | `bool` | `0.0` / `1.0` | — |
47
+ | `str` (including `"02138"`) | `NaN` | text by flat index |
48
+
49
+ There is **no datetime lane** on the float64 buffer. Python `datetime` objects stringify into `strings`. Do not add a `'date'` column kind.
50
+
51
+ Jagged 2D grids raise `ValueError` (spreadsheet ranges are rectangular).
52
+
53
+ ## Numbers
54
+
55
+ Median timings from an asymmetric bench (host = stdlib pack; child = deserialize + materialize).
56
+
57
+ Ingress, **20,000 × 5** (100k cells):
58
+
59
+ | Format | Pack | Dump | Load | Materialize | Total |
60
+ |--------|------|------|------|-------------|-------|
61
+ | JSON nested lists | 6.2 ms | 22.8 ms | 14.6 ms | 2.1 ms | 45.7 ms |
62
+ | Pickle 5 nested lists | 6.1 ms | 1.4 ms | 2.9 ms | 2.0 ms | 12.4 ms |
63
+ | **Pickle 5 + split-grid** | 8.3 ms | **0.013 ms** | **0.015 ms** | **0.002 ms** | **8.3 ms** |
64
+
65
+ Child materialize of a **100 × 100** numeric grid: nested-list pickle then `np.array` ~0.6 ms; split-grid `frombuffer` ~**0.016 ms**. Below 100 cells the envelope is not worth it — that is why `BINARY_MIN_CELLS` exists. Host pack still dominates large ingress; Cython only speeds that loop.
66
+
67
+ ## Python vs Cython
68
+
69
+ Host flatten is an optimized **pure-Python** loop:
70
+
71
+ - identity type checks (`type(val) is float`, `val is None`)
72
+ - bound-method capture (`buf_append = buf.append`)
73
+ - `None` as NaN on the fast path
74
+ - lazy column-state upgrades
75
+ - rectangular validation before the hot loop
76
+
77
+ The optional **Cython** module `splitgrid.pack` exposes `fast_flatten_grid_1d` / `fast_flatten_grid_2d`. It is loaded dynamically and **canary-tested** at import. If the extension is missing or fails the canary, the codec uses pure Python. Importing `splitgrid` never requires a compiler. Extension builds use release flags (`-O3 -DNDEBUG -g0` on Unix, `/O2 /DNDEBUG` on Windows).
78
+
79
+ ## Install
80
+
81
+ ```bash
82
+ pip install splitgrid
83
+ ```
84
+
85
+ PyPI wheels include the compiled Cython flatten accelerator. From a source checkout (Cython is built if a compiler is available):
86
+
87
+ ```bash
88
+ pip install .
89
+ pip install -e ".[test]" # editable + pytest, hypothesis, deal, numpy
90
+ pip install -e ".[numpy]" # child unpack/pack
91
+ pip install -e ".[verify]" # crosshair-tool
92
+ ```
93
+
94
+ Host pack stays NumPy-free. Child `child_unpack_*` / `child_pack_*` import NumPy locally.
95
+
96
+ ## Test
97
+
98
+ ```bash
99
+ pytest # default: -m "not slow"
100
+ pytest -m "not slow" # same
101
+ pytest -m slow # CrossHair subprocess checks (optional)
102
+ make verify # deal contracts + hypothesis round-trips, no CrossHair
103
+ ```
104
+
105
+ Default pytest **does not require** the compiled extension. When `splitgrid.pack` is built, `tests/test_cython_parity.py` compares Cython and Python flatten components on the same grids.
106
+
107
+ ## Bench
108
+
109
+ Requires NumPy (and the optional Cython extension if you want the accelerated pack row). From a source checkout:
110
+
111
+ ```bash
112
+ python scripts/bench_serialization.py --direction both
113
+ python scripts/bench_serialization.py --child-only
114
+ python scripts/bench_unpacking_opt.py
115
+ python scripts/profile_pack.py
116
+ python scripts/profile_nones.py
117
+ python scripts/run_serialization_ab.py --all
118
+ make bench
119
+ ```
120
+
121
+ ## Public API
122
+
123
+ ```python
124
+ from splitgrid import (
125
+ BINARY_MIN_CELLS,
126
+ host_pack_split_grid,
127
+ host_unpack_split_grid,
128
+ child_unpack_split_grid,
129
+ child_pack_split_grid,
130
+ host_pack_data,
131
+ host_unpack_data,
132
+ child_unpack_data,
133
+ child_pack_result,
134
+ is_split_grid,
135
+ load_cython_accelerator,
136
+ get_cython_status_info,
137
+ )
138
+ ```
139
+
140
+ `host_pack_data(..., force="auto"|"always"|"never")` chooses split-grid vs nested list. `host_pack_multi_data` is a thin `multi_data` wrapper over the same per-grid packing.
141
+
142
+ ## License
143
+
144
+ GPL-3.0-or-later.
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
4
4
 
5
5
  [project]
6
6
  name = "splitgrid"
7
- version = "0.1.0"
7
+ version = "0.1.2"
8
8
  description = "Asymmetric split-grid serialization: host stdlib flatten + optional Cython, child NumPy unpack"
9
9
  readme = "README.md"
10
10
  requires-python = ">=3.11"
@@ -10,9 +10,10 @@ system = platform.system()
10
10
  machine = platform.machine().lower()
11
11
 
12
12
  if system == "Windows":
13
- extra_compile_args.append("/O2")
13
+ extra_compile_args.extend(["/O2", "/DNDEBUG"])
14
14
  else:
15
- extra_compile_args.append("-O3")
15
+ # -g0 overrides manylinux/CPython default -g so wheels are not full of debug symbols.
16
+ extra_compile_args.extend(["-O3", "-DNDEBUG", "-g0"])
16
17
 
17
18
  # Only apply SPLITGRID_ARCH / WRITERAGENT_ARCH logic on Linux x86_64.
18
19
  # Generic x86-64 (not v3): flatten is memory-bound; SIMD floors buy ~1%.
@@ -55,7 +55,7 @@ from splitgrid.codec import (
55
55
  wire_cell_count,
56
56
  )
57
57
 
58
- __version__ = "0.1.0"
58
+ __version__ = "0.1.2"
59
59
 
60
60
 
61
61
  def __getattr__(name: str):
@@ -813,6 +813,10 @@ def _flatten_update_column_state(column_states: list[int], c: int, val: Any) ->
813
813
  column_states[c] = 2
814
814
  elif tname.startswith("float"):
815
815
  column_states[c] = 3
816
+ else:
817
+ # Decimal/Fraction/etc. already survived float(val). Default to float,
818
+ # matching Cython _update_column_state — not int (state 0).
819
+ column_states[c] = 3
816
820
 
817
821
 
818
822
  def _flatten_append_cell_slow(
@@ -1421,7 +1425,16 @@ def _child_unpack_single_data(wire: Any) -> Any:
1421
1425
  else:
1422
1426
  grid = list(unpacked)
1423
1427
  if is_numeric_grid(grid):
1424
- arr = np.array(grid, dtype=np.float64)
1428
+ # is_numeric_coercible treats whitespace/"" as Calc blanks, but
1429
+ # np.float64 cannot convert those strings (ValueError). Keep the list.
1430
+ try:
1431
+ arr = np.array(grid, dtype=np.float64)
1432
+ except ValueError:
1433
+ log.debug(
1434
+ "splitgrid child_unpack json_list as-is (non-floatable blanks) %s",
1435
+ describe_wire_value(unpacked),
1436
+ )
1437
+ return grid
1425
1438
  log.debug(
1426
1439
  "splitgrid child_unpack json_list -> ndarray shape=%s",
1427
1440
  arr.shape,
@@ -6,6 +6,8 @@
6
6
  "depends": [],
7
7
  "extra_compile_args": [
8
8
  "-O3",
9
+ "-DNDEBUG",
10
+ "-g0",
9
11
  "-march=x86-64"
10
12
  ],
11
13
  "name": "splitgrid.pack",
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: splitgrid
3
- Version: 0.1.0
3
+ Version: 0.1.2
4
4
  Summary: Asymmetric split-grid serialization: host stdlib flatten + optional Cython, child NumPy unpack
5
5
  Author-email: Keith Cu <keithcu@gmail.com>
6
6
  License: GPL-3.0-or-later
@@ -35,23 +35,33 @@ Dynamic: license-file
35
35
 
36
36
  # SplitGrid
37
37
 
38
+ [![PyPI](https://img.shields.io/pypi/v/splitgrid.svg)](https://pypi.org/project/splitgrid/)
39
+
38
40
  Asymmetric **split-grid** serialization for rectangular numeric and mixed-type grids.
39
41
 
40
- This package was **pulled out of [WriterAgent](https://github.com/KeithCu/writeragent)** (LibreOffice Writer/Calc/Draw AI extension). It is the same host flatten + child unpack codec that lived in `plugin.scripting.payload_codec`, plus the optional Cython flatten accelerator from `native/writeragent_vec`. WriterAgent can later depend on `splitgrid` without changing the wire dict.
42
+ This package was **pulled out of [WriterAgent](https://github.com/KeithCu/writeragent)** (LibreOffice Writer/Calc/Draw AI extension). It is the host flatten + child unpack codec that lived in `plugin.scripting.payload_codec`, plus the optional Cython flatten accelerator from `native/writeragent_vec`.
43
+
44
+ ## Why this exists
45
+
46
+ A spreadsheet host often cannot import NumPy: LibreOffice’s embedded Python is a different interpreter from the user’s venv, and loading the user’s C extensions into the host is an ABI footgun. Heavy compute therefore runs in a **child** process that *does* have NumPy. The range still has to cross that boundary.
47
+
48
+ On the wire, a Calc-style range is a nested list: `list[list[float | int | str | None]]`. Standard pickle walks **one heap object per cell**. On a 20,000 × 5 grid that is ~12 ms for dump + load. Pack the same numbers into a contiguous `float64` buffer and Pickle Protocol 5 moves the bytes in ~0.015 ms; the child can `np.frombuffer` in ~0.002 ms.
49
+
50
+ Almost all of the remaining time is **host flatten** — turning nested Python objects into that buffer without NumPy (~8.3 ms pure Python on that shape; ~3 ms with the Cython helper). SplitGrid is that flatten/unpack. Length-prefixed Pickle 5 framing stays in the application; this package is the codec only.
41
51
 
42
- Repo: [github.com/KeithCu/writeragent](https://github.com/KeithCu/writeragent)
52
+ Column-wise blobs and JSON + Base64 were tried first. Transposing columns on the host builds extra object graphs; Base64 and per-column reconstructs lose the `frombuffer` win. One row-major float64 buffer plus a sparse string map is faster and simpler.
43
53
 
44
- ## What split-grid is
54
+ ## How it works
45
55
 
46
- The compute path is **asymmetric by design**:
56
+ The path is **asymmetric**:
47
57
 
48
- - **Host pack** (LibreOffice’s embedded Python, or any NumPy-free interpreter) flattens a 1D or 2D grid into a contiguous `float64` buffer plus a sparse **integer-keyed** `strings` map. Empty cells become `NaN`. Zip codes like `"02138"` stay strings — they are never parsed as floats.
49
- - **Child unpack** (venv with NumPy) materializes a **pure-numeric** grid (`strings == {}`) with `np.frombuffer` (ndarray). Mixed grids use a vectorized object-masking path and return nested lists, restoring `None` for NaN holes.
50
- - **Host unpack** preserves `float('nan')` from the buffer (does **not** coerce holes to `None`). That is the locked egress policy: NaN becomes a Calc error, not a silent blank.
58
+ - **Host pack** (stdlib only) flattens a 1D or 2D rectangular grid into a contiguous `float64` buffer plus a sparse **integer-keyed** `strings` map. Empty cells become `NaN`. Zip codes like `"02138"` stay strings — they are never parsed as floats.
59
+ - **Child unpack** (NumPy) materializes a **pure-numeric** grid (`strings == {}`) with `np.frombuffer` (ndarray). Mixed grids use a vectorized object-masking path and return nested lists, restoring `None` for NaN holes.
60
+ - **Host unpack** preserves `float('nan')` from the buffer (does **not** coerce holes to `None`). That is the locked egress policy: NaN becomes a spreadsheet error, not a silent blank.
51
61
 
52
62
  Grids with fewer than **100 cells** (`BINARY_MIN_CELLS`) stay nested Python lists. Force `"always"` / `"never"` overrides that threshold (used by A/B tests).
53
63
 
54
- Wire envelope (Pickle5-friendly dict, same keys as WriterAgent):
64
+ Wire envelope (Pickle5-friendly dict):
55
65
 
56
66
  ```python
57
67
  {
@@ -64,9 +74,30 @@ Wire envelope (Pickle5-friendly dict, same keys as WriterAgent):
64
74
  }
65
75
  ```
66
76
 
77
+ | Cell value | `buffer` (float64) | `strings` |
78
+ |------------|--------------------|-----------|
79
+ | `None` (empty cell) | `NaN` | — |
80
+ | `int` / `float` | numeric value | — |
81
+ | `bool` | `0.0` / `1.0` | — |
82
+ | `str` (including `"02138"`) | `NaN` | text by flat index |
83
+
67
84
  There is **no datetime lane** on the float64 buffer. Python `datetime` objects stringify into `strings`. Do not add a `'date'` column kind.
68
85
 
69
- Jagged 2D grids raise `ValueError` (Calc ranges are rectangular).
86
+ Jagged 2D grids raise `ValueError` (spreadsheet ranges are rectangular).
87
+
88
+ ## Numbers
89
+
90
+ Median timings from an asymmetric bench (host = stdlib pack; child = deserialize + materialize).
91
+
92
+ Ingress, **20,000 × 5** (100k cells):
93
+
94
+ | Format | Pack | Dump | Load | Materialize | Total |
95
+ |--------|------|------|------|-------------|-------|
96
+ | JSON nested lists | 6.2 ms | 22.8 ms | 14.6 ms | 2.1 ms | 45.7 ms |
97
+ | Pickle 5 nested lists | 6.1 ms | 1.4 ms | 2.9 ms | 2.0 ms | 12.4 ms |
98
+ | **Pickle 5 + split-grid** | 8.3 ms | **0.013 ms** | **0.015 ms** | **0.002 ms** | **8.3 ms** |
99
+
100
+ Child materialize of a **100 × 100** numeric grid: nested-list pickle then `np.array` ~0.6 ms; split-grid `frombuffer` ~**0.016 ms**. Below 100 cells the envelope is not worth it — that is why `BINARY_MIN_CELLS` exists. Host pack still dominates large ingress; Cython only speeds that loop.
70
101
 
71
102
  ## Python vs Cython
72
103
 
@@ -78,7 +109,7 @@ Host flatten is an optimized **pure-Python** loop:
78
109
  - lazy column-state upgrades
79
110
  - rectangular validation before the hot loop
80
111
 
81
- The optional **Cython** module `splitgrid.pack` exposes `fast_flatten_grid_1d` / `fast_flatten_grid_2d`. It is loaded dynamically and **canary-tested** at import. If the extension is missing or fails the canary, the codec uses pure Python. Importing `splitgrid` never requires a compiler.
112
+ The optional **Cython** module `splitgrid.pack` exposes `fast_flatten_grid_1d` / `fast_flatten_grid_2d`. It is loaded dynamically and **canary-tested** at import. If the extension is missing or fails the canary, the codec uses pure Python. Importing `splitgrid` never requires a compiler. Extension builds use release flags (`-O3 -DNDEBUG -g0` on Unix, `/O2 /DNDEBUG` on Windows).
82
113
 
83
114
  ## Install
84
115
 
@@ -108,7 +139,21 @@ make verify # deal contracts + hypothesis round-trips, no CrossHair
108
139
 
109
140
  Default pytest **does not require** the compiled extension. When `splitgrid.pack` is built, `tests/test_cython_parity.py` compares Cython and Python flatten components on the same grids.
110
141
 
111
- ## Public API (WriterAgent-compatible)
142
+ ## Bench
143
+
144
+ Requires NumPy (and the optional Cython extension if you want the accelerated pack row). From a source checkout:
145
+
146
+ ```bash
147
+ python scripts/bench_serialization.py --direction both
148
+ python scripts/bench_serialization.py --child-only
149
+ python scripts/bench_unpacking_opt.py
150
+ python scripts/profile_pack.py
151
+ python scripts/profile_nones.py
152
+ python scripts/run_serialization_ab.py --all
153
+ make bench
154
+ ```
155
+
156
+ ## Public API
112
157
 
113
158
  ```python
114
159
  from splitgrid import (
@@ -129,14 +174,6 @@ from splitgrid import (
129
174
 
130
175
  `host_pack_data(..., force="auto"|"always"|"never")` chooses split-grid vs nested list. `host_pack_multi_data` is a thin `multi_data` wrapper over the same per-grid packing.
131
176
 
132
- ## How WriterAgent will consume this
133
-
134
- In [WriterAgent](https://github.com/KeithCu/writeragent), replace `from plugin.scripting.payload_codec import host_pack_data, ...` with `from splitgrid import host_pack_data, ...`. The envelope tag stays `"split_grid"`, dtype `"float64"`, integer-keyed `strings`, and `column_kinds` `int`/`float`/`bool`. Pickle Protocol 5 framing stays in WriterAgent’s `ipc.py` — this package is the codec only.
135
-
136
- ## Releasing
137
-
138
- Tags matching `v*` (for example `v0.1.0`) run `.github/workflows/publish.yml`: cibuildwheel + sdist, then trusted publishing to PyPI (GitHub environment `pypi`, no API token). The workflow file must already be on `main` before you push the tag.
139
-
140
177
  ## License
141
178
 
142
- GPL-3.0-or-later (same as WriterAgent).
179
+ GPL-3.0-or-later.
@@ -15,6 +15,7 @@ from __future__ import annotations
15
15
 
16
16
  import ast
17
17
  import math
18
+ from decimal import Decimal
18
19
  from pathlib import Path
19
20
  from unittest.mock import patch
20
21
 
@@ -41,6 +42,7 @@ from splitgrid.codec import (
41
42
  should_use_binary_envelope,
42
43
  wire_cell_count,
43
44
  )
45
+ from tests.serialization_ab_support import cython_accelerator_context
44
46
  from tests.payload_codec_test_support import (
45
47
  MIXED_LABEL_GRID,
46
48
  MIXED_WITH_ZIP,
@@ -512,6 +514,43 @@ def test_mixed_grid_preserves_non_numeric_string() -> None:
512
514
  assert out[0][1] == "hello"
513
515
 
514
516
 
517
+ def test_whitespace_only_cell_does_not_crash_child_unpack() -> None:
518
+ """Pasted ' ' is numeric-coercible for Calc but np.float64 cannot convert it."""
519
+ pytest.importorskip("numpy")
520
+ grid = [[1.0, " ", 3.0, 4.0], [5.0, 6.0, 7.0, 8.0], [9.0, 10.0, 11.0, 12.0]]
521
+ out = child_unpack_data(host_pack_data(grid, force="always"))
522
+ assert isinstance(out, list)
523
+ assert out[0][1] == " "
524
+
525
+
526
+ def test_empty_string_mixed_split_grid_does_not_crash_child_unpack() -> None:
527
+ """Bare '' on mixed split_grid must not raise (Calc usually maps '' to None first)."""
528
+ pytest.importorskip("numpy")
529
+ grid = [[1.0, "", 3.0, 4.0], [5.0, 6.0, 7.0, 8.0], [9.0, 10.0, 11.0, 12.0]]
530
+ out = child_unpack_data(host_pack_data(grid, force="always"))
531
+ assert isinstance(out, list)
532
+ assert out[0][1] == ""
533
+
534
+
535
+ def test_mixed_grid_real_nan_becomes_none_on_child() -> None:
536
+ """Documented: mixed-grid ingress has no blank-vs-NaN wire bit."""
537
+ pytest.importorskip("numpy")
538
+ grid = [[1.0, float("nan")], ["label", 4.0]]
539
+ out = child_unpack_data(host_pack_data(grid, force="always"))
540
+ assert out[0][1] is None
541
+ assert out[1][0] == "label"
542
+
543
+
544
+ def test_decimal_split_grid_stays_float_not_truncated_int() -> None:
545
+ """stdlib flatten must label Decimal columns float (Cython already did)."""
546
+ pytest.importorskip("numpy")
547
+ grid = [[Decimal("1.5"), Decimal("2.25")], [Decimal("3.0"), Decimal("4.75")]]
548
+ with cython_accelerator_context(enabled=False):
549
+ out = child_unpack_data(host_pack_data(grid, force="always"))
550
+ assert out[0][0] == pytest.approx(1.5)
551
+ assert out[0][1] == pytest.approx(2.25)
552
+
553
+
515
554
  def test_bool_cells_round_trip_in_numeric_grid() -> None:
516
555
  """Calc booleans in an all-numeric grid become 0.0/1.0 in child ndarray (float64 lane)."""
517
556
  np = pytest.importorskip("numpy")
splitgrid-0.1.0/README.md DELETED
@@ -1,107 +0,0 @@
1
- # SplitGrid
2
-
3
- Asymmetric **split-grid** serialization for rectangular numeric and mixed-type grids.
4
-
5
- This package was **pulled out of [WriterAgent](https://github.com/KeithCu/writeragent)** (LibreOffice Writer/Calc/Draw AI extension). It is the same host flatten + child unpack codec that lived in `plugin.scripting.payload_codec`, plus the optional Cython flatten accelerator from `native/writeragent_vec`. WriterAgent can later depend on `splitgrid` without changing the wire dict.
6
-
7
- Repo: [github.com/KeithCu/writeragent](https://github.com/KeithCu/writeragent)
8
-
9
- ## What split-grid is
10
-
11
- The compute path is **asymmetric by design**:
12
-
13
- - **Host pack** (LibreOffice’s embedded Python, or any NumPy-free interpreter) flattens a 1D or 2D grid into a contiguous `float64` buffer plus a sparse **integer-keyed** `strings` map. Empty cells become `NaN`. Zip codes like `"02138"` stay strings — they are never parsed as floats.
14
- - **Child unpack** (venv with NumPy) materializes a **pure-numeric** grid (`strings == {}`) with `np.frombuffer` (ndarray). Mixed grids use a vectorized object-masking path and return nested lists, restoring `None` for NaN holes.
15
- - **Host unpack** preserves `float('nan')` from the buffer (does **not** coerce holes to `None`). That is the locked egress policy: NaN becomes a Calc error, not a silent blank.
16
-
17
- Grids with fewer than **100 cells** (`BINARY_MIN_CELLS`) stay nested Python lists. Force `"always"` / `"never"` overrides that threshold (used by A/B tests).
18
-
19
- Wire envelope (Pickle5-friendly dict, same keys as WriterAgent):
20
-
21
- ```python
22
- {
23
- "__wa_payload__": "split_grid",
24
- "dtype": "float64",
25
- "column_kinds": ["int", "float"], # per column: int / float / bool
26
- "shape": [rows, cols], # or [n] for 1D
27
- "buffer": b"...", # row-major float64 bytes
28
- "strings": {7: "banana"}, # integer keys, not str(idx)
29
- }
30
- ```
31
-
32
- There is **no datetime lane** on the float64 buffer. Python `datetime` objects stringify into `strings`. Do not add a `'date'` column kind.
33
-
34
- Jagged 2D grids raise `ValueError` (Calc ranges are rectangular).
35
-
36
- ## Python vs Cython
37
-
38
- Host flatten is an optimized **pure-Python** loop:
39
-
40
- - identity type checks (`type(val) is float`, `val is None`)
41
- - bound-method capture (`buf_append = buf.append`)
42
- - `None` as NaN on the fast path
43
- - lazy column-state upgrades
44
- - rectangular validation before the hot loop
45
-
46
- The optional **Cython** module `splitgrid.pack` exposes `fast_flatten_grid_1d` / `fast_flatten_grid_2d`. It is loaded dynamically and **canary-tested** at import. If the extension is missing or fails the canary, the codec uses pure Python. Importing `splitgrid` never requires a compiler.
47
-
48
- ## Install
49
-
50
- ```bash
51
- pip install splitgrid
52
- ```
53
-
54
- PyPI wheels include the compiled Cython flatten accelerator. From a source checkout (Cython is built if a compiler is available):
55
-
56
- ```bash
57
- pip install .
58
- pip install -e ".[test]" # editable + pytest, hypothesis, deal, numpy
59
- pip install -e ".[numpy]" # child unpack/pack
60
- pip install -e ".[verify]" # crosshair-tool
61
- ```
62
-
63
- Host pack stays NumPy-free. Child `child_unpack_*` / `child_pack_*` import NumPy locally.
64
-
65
- ## Test
66
-
67
- ```bash
68
- pytest # default: -m "not slow"
69
- pytest -m "not slow" # same
70
- pytest -m slow # CrossHair subprocess checks (optional)
71
- make verify # deal contracts + hypothesis round-trips, no CrossHair
72
- ```
73
-
74
- Default pytest **does not require** the compiled extension. When `splitgrid.pack` is built, `tests/test_cython_parity.py` compares Cython and Python flatten components on the same grids.
75
-
76
- ## Public API (WriterAgent-compatible)
77
-
78
- ```python
79
- from splitgrid import (
80
- BINARY_MIN_CELLS,
81
- host_pack_split_grid,
82
- host_unpack_split_grid,
83
- child_unpack_split_grid,
84
- child_pack_split_grid,
85
- host_pack_data,
86
- host_unpack_data,
87
- child_unpack_data,
88
- child_pack_result,
89
- is_split_grid,
90
- load_cython_accelerator,
91
- get_cython_status_info,
92
- )
93
- ```
94
-
95
- `host_pack_data(..., force="auto"|"always"|"never")` chooses split-grid vs nested list. `host_pack_multi_data` is a thin `multi_data` wrapper over the same per-grid packing.
96
-
97
- ## How WriterAgent will consume this
98
-
99
- In [WriterAgent](https://github.com/KeithCu/writeragent), replace `from plugin.scripting.payload_codec import host_pack_data, ...` with `from splitgrid import host_pack_data, ...`. The envelope tag stays `"split_grid"`, dtype `"float64"`, integer-keyed `strings`, and `column_kinds` `int`/`float`/`bool`. Pickle Protocol 5 framing stays in WriterAgent’s `ipc.py` — this package is the codec only.
100
-
101
- ## Releasing
102
-
103
- Tags matching `v*` (for example `v0.1.0`) run `.github/workflows/publish.yml`: cibuildwheel + sdist, then trusted publishing to PyPI (GitHub environment `pypi`, no API token). The workflow file must already be on `main` before you push the tag.
104
-
105
- ## License
106
-
107
- GPL-3.0-or-later (same as WriterAgent).
File without changes
File without changes
File without changes