splitgrid 0.1.0__tar.gz → 0.1.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (25) hide show
  1. {splitgrid-0.1.0/src/splitgrid.egg-info → splitgrid-0.1.1}/PKG-INFO +42 -21
  2. splitgrid-0.1.1/README.md +128 -0
  3. {splitgrid-0.1.0 → splitgrid-0.1.1}/pyproject.toml +1 -1
  4. {splitgrid-0.1.0 → splitgrid-0.1.1}/setup.py +3 -2
  5. {splitgrid-0.1.0 → splitgrid-0.1.1}/src/splitgrid/__init__.py +1 -1
  6. {splitgrid-0.1.0 → splitgrid-0.1.1}/src/splitgrid/pack.c +2 -0
  7. {splitgrid-0.1.0 → splitgrid-0.1.1/src/splitgrid.egg-info}/PKG-INFO +42 -21
  8. splitgrid-0.1.0/README.md +0 -107
  9. {splitgrid-0.1.0 → splitgrid-0.1.1}/LICENSE +0 -0
  10. {splitgrid-0.1.0 → splitgrid-0.1.1}/MANIFEST.in +0 -0
  11. {splitgrid-0.1.0 → splitgrid-0.1.1}/Makefile +0 -0
  12. {splitgrid-0.1.0 → splitgrid-0.1.1}/setup.cfg +0 -0
  13. {splitgrid-0.1.0 → splitgrid-0.1.1}/src/splitgrid/codec.py +0 -0
  14. {splitgrid-0.1.0 → splitgrid-0.1.1}/src/splitgrid/deal_shim.py +0 -0
  15. {splitgrid-0.1.0 → splitgrid-0.1.1}/src/splitgrid/pack.pyx +0 -0
  16. {splitgrid-0.1.0 → splitgrid-0.1.1}/src/splitgrid.egg-info/SOURCES.txt +0 -0
  17. {splitgrid-0.1.0 → splitgrid-0.1.1}/src/splitgrid.egg-info/dependency_links.txt +0 -0
  18. {splitgrid-0.1.0 → splitgrid-0.1.1}/src/splitgrid.egg-info/requires.txt +0 -0
  19. {splitgrid-0.1.0 → splitgrid-0.1.1}/src/splitgrid.egg-info/top_level.txt +0 -0
  20. {splitgrid-0.1.0 → splitgrid-0.1.1}/tests/test_cython_parity.py +0 -0
  21. {splitgrid-0.1.0 → splitgrid-0.1.1}/tests/test_payload_codec.py +0 -0
  22. {splitgrid-0.1.0 → splitgrid-0.1.1}/tests/test_payload_codec_policy_verification.py +0 -0
  23. {splitgrid-0.1.0 → splitgrid-0.1.1}/tests/test_serialization_ab.py +0 -0
  24. {splitgrid-0.1.0 → splitgrid-0.1.1}/tests/test_serialization_ab_support.py +0 -0
  25. {splitgrid-0.1.0 → splitgrid-0.1.1}/tests/test_serialization_verification.py +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: splitgrid
3
- Version: 0.1.0
3
+ Version: 0.1.1
4
4
  Summary: Asymmetric split-grid serialization: host stdlib flatten + optional Cython, child NumPy unpack
5
5
  Author-email: Keith Cu <keithcu@gmail.com>
6
6
  License: GPL-3.0-or-later
@@ -37,21 +37,29 @@ Dynamic: license-file
37
37
 
38
38
  Asymmetric **split-grid** serialization for rectangular numeric and mixed-type grids.
39
39
 
40
- This package was **pulled out of [WriterAgent](https://github.com/KeithCu/writeragent)** (LibreOffice Writer/Calc/Draw AI extension). It is the same host flatten + child unpack codec that lived in `plugin.scripting.payload_codec`, plus the optional Cython flatten accelerator from `native/writeragent_vec`. WriterAgent can later depend on `splitgrid` without changing the wire dict.
40
+ This package was **pulled out of [WriterAgent](https://github.com/KeithCu/writeragent)** (LibreOffice Writer/Calc/Draw AI extension). It is the host flatten + child unpack codec that lived in `plugin.scripting.payload_codec`, plus the optional Cython flatten accelerator from `native/writeragent_vec`.
41
41
 
42
- Repo: [github.com/KeithCu/writeragent](https://github.com/KeithCu/writeragent)
42
+ ## Why this exists
43
43
 
44
- ## What split-grid is
44
+ A spreadsheet host often cannot import NumPy: LibreOffice’s embedded Python is a different interpreter from the user’s venv, and loading the user’s C extensions into the host is an ABI footgun. Heavy compute therefore runs in a **child** process that *does* have NumPy. The range still has to cross that boundary.
45
45
 
46
- The compute path is **asymmetric by design**:
46
+ On the wire, a Calc-style range is a nested list: `list[list[float | int | str | None]]`. Standard pickle walks **one heap object per cell**. On a 20,000 × 5 grid that is ~12 ms for dump + load. Pack the same numbers into a contiguous `float64` buffer and Pickle Protocol 5 moves the bytes in ~0.015 ms; the child can `np.frombuffer` in ~0.002 ms.
47
47
 
48
- - **Host pack** (LibreOffice’s embedded Python, or any NumPy-free interpreter) flattens a 1D or 2D grid into a contiguous `float64` buffer plus a sparse **integer-keyed** `strings` map. Empty cells become `NaN`. Zip codes like `"02138"` stay strings — they are never parsed as floats.
49
- - **Child unpack** (venv with NumPy) materializes a **pure-numeric** grid (`strings == {}`) with `np.frombuffer` (ndarray). Mixed grids use a vectorized object-masking path and return nested lists, restoring `None` for NaN holes.
50
- - **Host unpack** preserves `float('nan')` from the buffer (does **not** coerce holes to `None`). That is the locked egress policy: NaN becomes a Calc error, not a silent blank.
48
+ Almost all of the remaining time is **host flatten** — turning nested Python objects into that buffer without NumPy (~8.3 ms pure Python on that shape; ~3 ms with the Cython helper). SplitGrid is that flatten/unpack. Length-prefixed Pickle 5 framing stays in the application; this package is the codec only.
49
+
50
+ Column-wise blobs and JSON + Base64 were tried first. Transposing columns on the host builds extra object graphs; Base64 and per-column reconstructs lose the `frombuffer` win. One row-major float64 buffer plus a sparse string map is faster and simpler.
51
+
52
+ ## How it works
53
+
54
+ The path is **asymmetric**:
55
+
56
+ - **Host pack** (stdlib only) flattens a 1D or 2D rectangular grid into a contiguous `float64` buffer plus a sparse **integer-keyed** `strings` map. Empty cells become `NaN`. Zip codes like `"02138"` stay strings — they are never parsed as floats.
57
+ - **Child unpack** (NumPy) materializes a **pure-numeric** grid (`strings == {}`) with `np.frombuffer` (ndarray). Mixed grids use a vectorized object-masking path and return nested lists, restoring `None` for NaN holes.
58
+ - **Host unpack** preserves `float('nan')` from the buffer (does **not** coerce holes to `None`). That is the locked egress policy: NaN becomes a spreadsheet error, not a silent blank.
51
59
 
52
60
  Grids with fewer than **100 cells** (`BINARY_MIN_CELLS`) stay nested Python lists. Force `"always"` / `"never"` overrides that threshold (used by A/B tests).
53
61
 
54
- Wire envelope (Pickle5-friendly dict, same keys as WriterAgent):
62
+ Wire envelope (Pickle5-friendly dict):
55
63
 
56
64
  ```python
57
65
  {
@@ -64,9 +72,30 @@ Wire envelope (Pickle5-friendly dict, same keys as WriterAgent):
64
72
  }
65
73
  ```
66
74
 
75
+ | Cell value | `buffer` (float64) | `strings` |
76
+ |------------|--------------------|-----------|
77
+ | `None` (empty cell) | `NaN` | — |
78
+ | `int` / `float` | numeric value | — |
79
+ | `bool` | `0.0` / `1.0` | — |
80
+ | `str` (including `"02138"`) | `NaN` | text by flat index |
81
+
67
82
  There is **no datetime lane** on the float64 buffer. Python `datetime` objects stringify into `strings`. Do not add a `'date'` column kind.
68
83
 
69
- Jagged 2D grids raise `ValueError` (Calc ranges are rectangular).
84
+ Jagged 2D grids raise `ValueError` (spreadsheet ranges are rectangular).
85
+
86
+ ## Numbers
87
+
88
+ Median timings from an asymmetric bench (host = stdlib pack; child = deserialize + materialize).
89
+
90
+ Ingress, **20,000 × 5** (100k cells):
91
+
92
+ | Format | Pack | Dump | Load | Materialize | Total |
93
+ |--------|------|------|------|-------------|-------|
94
+ | JSON nested lists | 6.2 ms | 22.8 ms | 14.6 ms | 2.1 ms | 45.7 ms |
95
+ | Pickle 5 nested lists | 6.1 ms | 1.4 ms | 2.9 ms | 2.0 ms | 12.4 ms |
96
+ | **Pickle 5 + split-grid** | 8.3 ms | **0.013 ms** | **0.015 ms** | **0.002 ms** | **8.3 ms** |
97
+
98
+ Child materialize of a **100 × 100** numeric grid: nested-list pickle then `np.array` ~0.6 ms; split-grid `frombuffer` ~**0.016 ms**. Below 100 cells the envelope is not worth it — that is why `BINARY_MIN_CELLS` exists. Host pack still dominates large ingress; Cython only speeds that loop.
70
99
 
71
100
  ## Python vs Cython
72
101
 
@@ -78,7 +107,7 @@ Host flatten is an optimized **pure-Python** loop:
78
107
  - lazy column-state upgrades
79
108
  - rectangular validation before the hot loop
80
109
 
81
- The optional **Cython** module `splitgrid.pack` exposes `fast_flatten_grid_1d` / `fast_flatten_grid_2d`. It is loaded dynamically and **canary-tested** at import. If the extension is missing or fails the canary, the codec uses pure Python. Importing `splitgrid` never requires a compiler.
110
+ The optional **Cython** module `splitgrid.pack` exposes `fast_flatten_grid_1d` / `fast_flatten_grid_2d`. It is loaded dynamically and **canary-tested** at import. If the extension is missing or fails the canary, the codec uses pure Python. Importing `splitgrid` never requires a compiler. Extension builds use release flags (`-O3 -DNDEBUG -g0` on Unix, `/O2 /DNDEBUG` on Windows).
82
111
 
83
112
  ## Install
84
113
 
@@ -108,7 +137,7 @@ make verify # deal contracts + hypothesis round-trips, no CrossHair
108
137
 
109
138
  Default pytest **does not require** the compiled extension. When `splitgrid.pack` is built, `tests/test_cython_parity.py` compares Cython and Python flatten components on the same grids.
110
139
 
111
- ## Public API (WriterAgent-compatible)
140
+ ## Public API
112
141
 
113
142
  ```python
114
143
  from splitgrid import (
@@ -129,14 +158,6 @@ from splitgrid import (
129
158
 
130
159
  `host_pack_data(..., force="auto"|"always"|"never")` chooses split-grid vs nested list. `host_pack_multi_data` is a thin `multi_data` wrapper over the same per-grid packing.
131
160
 
132
- ## How WriterAgent will consume this
133
-
134
- In [WriterAgent](https://github.com/KeithCu/writeragent), replace `from plugin.scripting.payload_codec import host_pack_data, ...` with `from splitgrid import host_pack_data, ...`. The envelope tag stays `"split_grid"`, dtype `"float64"`, integer-keyed `strings`, and `column_kinds` `int`/`float`/`bool`. Pickle Protocol 5 framing stays in WriterAgent’s `ipc.py` — this package is the codec only.
135
-
136
- ## Releasing
137
-
138
- Tags matching `v*` (for example `v0.1.0`) run `.github/workflows/publish.yml`: cibuildwheel + sdist, then trusted publishing to PyPI (GitHub environment `pypi`, no API token). The workflow file must already be on `main` before you push the tag.
139
-
140
161
  ## License
141
162
 
142
- GPL-3.0-or-later (same as WriterAgent).
163
+ GPL-3.0-or-later.
@@ -0,0 +1,128 @@
1
+ # SplitGrid
2
+
3
+ Asymmetric **split-grid** serialization for rectangular numeric and mixed-type grids.
4
+
5
+ This package was **pulled out of [WriterAgent](https://github.com/KeithCu/writeragent)** (LibreOffice Writer/Calc/Draw AI extension). It is the host flatten + child unpack codec that lived in `plugin.scripting.payload_codec`, plus the optional Cython flatten accelerator from `native/writeragent_vec`.
6
+
7
+ ## Why this exists
8
+
9
+ A spreadsheet host often cannot import NumPy: LibreOffice’s embedded Python is a different interpreter from the user’s venv, and loading the user’s C extensions into the host is an ABI footgun. Heavy compute therefore runs in a **child** process that *does* have NumPy. The range still has to cross that boundary.
10
+
11
+ On the wire, a Calc-style range is a nested list: `list[list[float | int | str | None]]`. Standard pickle walks **one heap object per cell**. On a 20,000 × 5 grid that is ~12 ms for dump + load. Pack the same numbers into a contiguous `float64` buffer and Pickle Protocol 5 moves the bytes in ~0.015 ms; the child can `np.frombuffer` in ~0.002 ms.
12
+
13
+ Almost all of the remaining time is **host flatten** — turning nested Python objects into that buffer without NumPy (~8.3 ms pure Python on that shape; ~3 ms with the Cython helper). SplitGrid is that flatten/unpack. Length-prefixed Pickle 5 framing stays in the application; this package is the codec only.
14
+
15
+ Column-wise blobs and JSON + Base64 were tried first. Transposing columns on the host builds extra object graphs; Base64 and per-column reconstructs lose the `frombuffer` win. One row-major float64 buffer plus a sparse string map is faster and simpler.
16
+
17
+ ## How it works
18
+
19
+ The path is **asymmetric**:
20
+
21
+ - **Host pack** (stdlib only) flattens a 1D or 2D rectangular grid into a contiguous `float64` buffer plus a sparse **integer-keyed** `strings` map. Empty cells become `NaN`. Zip codes like `"02138"` stay strings — they are never parsed as floats.
22
+ - **Child unpack** (NumPy) materializes a **pure-numeric** grid (`strings == {}`) with `np.frombuffer` (ndarray). Mixed grids use a vectorized object-masking path and return nested lists, restoring `None` for NaN holes.
23
+ - **Host unpack** preserves `float('nan')` from the buffer (does **not** coerce holes to `None`). That is the locked egress policy: NaN becomes a spreadsheet error, not a silent blank.
24
+
25
+ Grids with fewer than **100 cells** (`BINARY_MIN_CELLS`) stay nested Python lists. Force `"always"` / `"never"` overrides that threshold (used by A/B tests).
26
+
27
+ Wire envelope (Pickle5-friendly dict):
28
+
29
+ ```python
30
+ {
31
+ "__wa_payload__": "split_grid",
32
+ "dtype": "float64",
33
+ "column_kinds": ["int", "float"], # per column: int / float / bool
34
+ "shape": [rows, cols], # or [n] for 1D
35
+ "buffer": b"...", # row-major float64 bytes
36
+ "strings": {7: "banana"}, # integer keys, not str(idx)
37
+ }
38
+ ```
39
+
40
+ | Cell value | `buffer` (float64) | `strings` |
41
+ |------------|--------------------|-----------|
42
+ | `None` (empty cell) | `NaN` | — |
43
+ | `int` / `float` | numeric value | — |
44
+ | `bool` | `0.0` / `1.0` | — |
45
+ | `str` (including `"02138"`) | `NaN` | text by flat index |
46
+
47
+ There is **no datetime lane** on the float64 buffer. Python `datetime` objects stringify into `strings`. Do not add a `'date'` column kind.
48
+
49
+ Jagged 2D grids raise `ValueError` (spreadsheet ranges are rectangular).
50
+
51
+ ## Numbers
52
+
53
+ Median timings from an asymmetric bench (host = stdlib pack; child = deserialize + materialize).
54
+
55
+ Ingress, **20,000 × 5** (100k cells):
56
+
57
+ | Format | Pack | Dump | Load | Materialize | Total |
58
+ |--------|------|------|------|-------------|-------|
59
+ | JSON nested lists | 6.2 ms | 22.8 ms | 14.6 ms | 2.1 ms | 45.7 ms |
60
+ | Pickle 5 nested lists | 6.1 ms | 1.4 ms | 2.9 ms | 2.0 ms | 12.4 ms |
61
+ | **Pickle 5 + split-grid** | 8.3 ms | **0.013 ms** | **0.015 ms** | **0.002 ms** | **8.3 ms** |
62
+
63
+ Child materialize of a **100 × 100** numeric grid: nested-list pickle then `np.array` ~0.6 ms; split-grid `frombuffer` ~**0.016 ms**. Below 100 cells the envelope is not worth it — that is why `BINARY_MIN_CELLS` exists. Host pack still dominates large ingress; Cython only speeds that loop.
64
+
65
+ ## Python vs Cython
66
+
67
+ Host flatten is an optimized **pure-Python** loop:
68
+
69
+ - identity type checks (`type(val) is float`, `val is None`)
70
+ - bound-method capture (`buf_append = buf.append`)
71
+ - `None` as NaN on the fast path
72
+ - lazy column-state upgrades
73
+ - rectangular validation before the hot loop
74
+
75
+ The optional **Cython** module `splitgrid.pack` exposes `fast_flatten_grid_1d` / `fast_flatten_grid_2d`. It is loaded dynamically and **canary-tested** at import. If the extension is missing or fails the canary, the codec uses pure Python. Importing `splitgrid` never requires a compiler. Extension builds use release flags (`-O3 -DNDEBUG -g0` on Unix, `/O2 /DNDEBUG` on Windows).
76
+
77
+ ## Install
78
+
79
+ ```bash
80
+ pip install splitgrid
81
+ ```
82
+
83
+ PyPI wheels include the compiled Cython flatten accelerator. From a source checkout (Cython is built if a compiler is available):
84
+
85
+ ```bash
86
+ pip install .
87
+ pip install -e ".[test]" # editable + pytest, hypothesis, deal, numpy
88
+ pip install -e ".[numpy]" # child unpack/pack
89
+ pip install -e ".[verify]" # crosshair-tool
90
+ ```
91
+
92
+ Host pack stays NumPy-free. Child `child_unpack_*` / `child_pack_*` import NumPy locally.
93
+
94
+ ## Test
95
+
96
+ ```bash
97
+ pytest # default: -m "not slow"
98
+ pytest -m "not slow" # same
99
+ pytest -m slow # CrossHair subprocess checks (optional)
100
+ make verify # deal contracts + hypothesis round-trips, no CrossHair
101
+ ```
102
+
103
+ Default pytest **does not require** the compiled extension. When `splitgrid.pack` is built, `tests/test_cython_parity.py` compares Cython and Python flatten components on the same grids.
104
+
105
+ ## Public API
106
+
107
+ ```python
108
+ from splitgrid import (
109
+ BINARY_MIN_CELLS,
110
+ host_pack_split_grid,
111
+ host_unpack_split_grid,
112
+ child_unpack_split_grid,
113
+ child_pack_split_grid,
114
+ host_pack_data,
115
+ host_unpack_data,
116
+ child_unpack_data,
117
+ child_pack_result,
118
+ is_split_grid,
119
+ load_cython_accelerator,
120
+ get_cython_status_info,
121
+ )
122
+ ```
123
+
124
+ `host_pack_data(..., force="auto"|"always"|"never")` chooses split-grid vs nested list. `host_pack_multi_data` is a thin `multi_data` wrapper over the same per-grid packing.
125
+
126
+ ## License
127
+
128
+ GPL-3.0-or-later.
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
4
4
 
5
5
  [project]
6
6
  name = "splitgrid"
7
- version = "0.1.0"
7
+ version = "0.1.1"
8
8
  description = "Asymmetric split-grid serialization: host stdlib flatten + optional Cython, child NumPy unpack"
9
9
  readme = "README.md"
10
10
  requires-python = ">=3.11"
@@ -10,9 +10,10 @@ system = platform.system()
10
10
  machine = platform.machine().lower()
11
11
 
12
12
  if system == "Windows":
13
- extra_compile_args.append("/O2")
13
+ extra_compile_args.extend(["/O2", "/DNDEBUG"])
14
14
  else:
15
- extra_compile_args.append("-O3")
15
+ # -g0 overrides manylinux/CPython default -g so wheels are not full of debug symbols.
16
+ extra_compile_args.extend(["-O3", "-DNDEBUG", "-g0"])
16
17
 
17
18
  # Only apply SPLITGRID_ARCH / WRITERAGENT_ARCH logic on Linux x86_64.
18
19
  # Generic x86-64 (not v3): flatten is memory-bound; SIMD floors buy ~1%.
@@ -55,7 +55,7 @@ from splitgrid.codec import (
55
55
  wire_cell_count,
56
56
  )
57
57
 
58
- __version__ = "0.1.0"
58
+ __version__ = "0.1.1"
59
59
 
60
60
 
61
61
  def __getattr__(name: str):
@@ -6,6 +6,8 @@
6
6
  "depends": [],
7
7
  "extra_compile_args": [
8
8
  "-O3",
9
+ "-DNDEBUG",
10
+ "-g0",
9
11
  "-march=x86-64"
10
12
  ],
11
13
  "name": "splitgrid.pack",
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: splitgrid
3
- Version: 0.1.0
3
+ Version: 0.1.1
4
4
  Summary: Asymmetric split-grid serialization: host stdlib flatten + optional Cython, child NumPy unpack
5
5
  Author-email: Keith Cu <keithcu@gmail.com>
6
6
  License: GPL-3.0-or-later
@@ -37,21 +37,29 @@ Dynamic: license-file
37
37
 
38
38
  Asymmetric **split-grid** serialization for rectangular numeric and mixed-type grids.
39
39
 
40
- This package was **pulled out of [WriterAgent](https://github.com/KeithCu/writeragent)** (LibreOffice Writer/Calc/Draw AI extension). It is the same host flatten + child unpack codec that lived in `plugin.scripting.payload_codec`, plus the optional Cython flatten accelerator from `native/writeragent_vec`. WriterAgent can later depend on `splitgrid` without changing the wire dict.
40
+ This package was **pulled out of [WriterAgent](https://github.com/KeithCu/writeragent)** (LibreOffice Writer/Calc/Draw AI extension). It is the host flatten + child unpack codec that lived in `plugin.scripting.payload_codec`, plus the optional Cython flatten accelerator from `native/writeragent_vec`.
41
41
 
42
- Repo: [github.com/KeithCu/writeragent](https://github.com/KeithCu/writeragent)
42
+ ## Why this exists
43
43
 
44
- ## What split-grid is
44
+ A spreadsheet host often cannot import NumPy: LibreOffice’s embedded Python is a different interpreter from the user’s venv, and loading the user’s C extensions into the host is an ABI footgun. Heavy compute therefore runs in a **child** process that *does* have NumPy. The range still has to cross that boundary.
45
45
 
46
- The compute path is **asymmetric by design**:
46
+ On the wire, a Calc-style range is a nested list: `list[list[float | int | str | None]]`. Standard pickle walks **one heap object per cell**. On a 20,000 × 5 grid that is ~12 ms for dump + load. Pack the same numbers into a contiguous `float64` buffer and Pickle Protocol 5 moves the bytes in ~0.015 ms; the child can `np.frombuffer` in ~0.002 ms.
47
47
 
48
- - **Host pack** (LibreOffice’s embedded Python, or any NumPy-free interpreter) flattens a 1D or 2D grid into a contiguous `float64` buffer plus a sparse **integer-keyed** `strings` map. Empty cells become `NaN`. Zip codes like `"02138"` stay strings — they are never parsed as floats.
49
- - **Child unpack** (venv with NumPy) materializes a **pure-numeric** grid (`strings == {}`) with `np.frombuffer` (ndarray). Mixed grids use a vectorized object-masking path and return nested lists, restoring `None` for NaN holes.
50
- - **Host unpack** preserves `float('nan')` from the buffer (does **not** coerce holes to `None`). That is the locked egress policy: NaN becomes a Calc error, not a silent blank.
48
+ Almost all of the remaining time is **host flatten** — turning nested Python objects into that buffer without NumPy (~8.3 ms pure Python on that shape; ~3 ms with the Cython helper). SplitGrid is that flatten/unpack. Length-prefixed Pickle 5 framing stays in the application; this package is the codec only.
49
+
50
+ Column-wise blobs and JSON + Base64 were tried first. Transposing columns on the host builds extra object graphs; Base64 and per-column reconstructs lose the `frombuffer` win. One row-major float64 buffer plus a sparse string map is faster and simpler.
51
+
52
+ ## How it works
53
+
54
+ The path is **asymmetric**:
55
+
56
+ - **Host pack** (stdlib only) flattens a 1D or 2D rectangular grid into a contiguous `float64` buffer plus a sparse **integer-keyed** `strings` map. Empty cells become `NaN`. Zip codes like `"02138"` stay strings — they are never parsed as floats.
57
+ - **Child unpack** (NumPy) materializes a **pure-numeric** grid (`strings == {}`) with `np.frombuffer` (ndarray). Mixed grids use a vectorized object-masking path and return nested lists, restoring `None` for NaN holes.
58
+ - **Host unpack** preserves `float('nan')` from the buffer (does **not** coerce holes to `None`). That is the locked egress policy: NaN becomes a spreadsheet error, not a silent blank.
51
59
 
52
60
  Grids with fewer than **100 cells** (`BINARY_MIN_CELLS`) stay nested Python lists. Force `"always"` / `"never"` overrides that threshold (used by A/B tests).
53
61
 
54
- Wire envelope (Pickle5-friendly dict, same keys as WriterAgent):
62
+ Wire envelope (Pickle5-friendly dict):
55
63
 
56
64
  ```python
57
65
  {
@@ -64,9 +72,30 @@ Wire envelope (Pickle5-friendly dict, same keys as WriterAgent):
64
72
  }
65
73
  ```
66
74
 
75
+ | Cell value | `buffer` (float64) | `strings` |
76
+ |------------|--------------------|-----------|
77
+ | `None` (empty cell) | `NaN` | — |
78
+ | `int` / `float` | numeric value | — |
79
+ | `bool` | `0.0` / `1.0` | — |
80
+ | `str` (including `"02138"`) | `NaN` | text by flat index |
81
+
67
82
  There is **no datetime lane** on the float64 buffer. Python `datetime` objects stringify into `strings`. Do not add a `'date'` column kind.
68
83
 
69
- Jagged 2D grids raise `ValueError` (Calc ranges are rectangular).
84
+ Jagged 2D grids raise `ValueError` (spreadsheet ranges are rectangular).
85
+
86
+ ## Numbers
87
+
88
+ Median timings from an asymmetric bench (host = stdlib pack; child = deserialize + materialize).
89
+
90
+ Ingress, **20,000 × 5** (100k cells):
91
+
92
+ | Format | Pack | Dump | Load | Materialize | Total |
93
+ |--------|------|------|------|-------------|-------|
94
+ | JSON nested lists | 6.2 ms | 22.8 ms | 14.6 ms | 2.1 ms | 45.7 ms |
95
+ | Pickle 5 nested lists | 6.1 ms | 1.4 ms | 2.9 ms | 2.0 ms | 12.4 ms |
96
+ | **Pickle 5 + split-grid** | 8.3 ms | **0.013 ms** | **0.015 ms** | **0.002 ms** | **8.3 ms** |
97
+
98
+ Child materialize of a **100 × 100** numeric grid: nested-list pickle then `np.array` ~0.6 ms; split-grid `frombuffer` ~**0.016 ms**. Below 100 cells the envelope is not worth it — that is why `BINARY_MIN_CELLS` exists. Host pack still dominates large ingress; Cython only speeds that loop.
70
99
 
71
100
  ## Python vs Cython
72
101
 
@@ -78,7 +107,7 @@ Host flatten is an optimized **pure-Python** loop:
78
107
  - lazy column-state upgrades
79
108
  - rectangular validation before the hot loop
80
109
 
81
- The optional **Cython** module `splitgrid.pack` exposes `fast_flatten_grid_1d` / `fast_flatten_grid_2d`. It is loaded dynamically and **canary-tested** at import. If the extension is missing or fails the canary, the codec uses pure Python. Importing `splitgrid` never requires a compiler.
110
+ The optional **Cython** module `splitgrid.pack` exposes `fast_flatten_grid_1d` / `fast_flatten_grid_2d`. It is loaded dynamically and **canary-tested** at import. If the extension is missing or fails the canary, the codec uses pure Python. Importing `splitgrid` never requires a compiler. Extension builds use release flags (`-O3 -DNDEBUG -g0` on Unix, `/O2 /DNDEBUG` on Windows).
82
111
 
83
112
  ## Install
84
113
 
@@ -108,7 +137,7 @@ make verify # deal contracts + hypothesis round-trips, no CrossHair
108
137
 
109
138
  Default pytest **does not require** the compiled extension. When `splitgrid.pack` is built, `tests/test_cython_parity.py` compares Cython and Python flatten components on the same grids.
110
139
 
111
- ## Public API (WriterAgent-compatible)
140
+ ## Public API
112
141
 
113
142
  ```python
114
143
  from splitgrid import (
@@ -129,14 +158,6 @@ from splitgrid import (
129
158
 
130
159
  `host_pack_data(..., force="auto"|"always"|"never")` chooses split-grid vs nested list. `host_pack_multi_data` is a thin `multi_data` wrapper over the same per-grid packing.
131
160
 
132
- ## How WriterAgent will consume this
133
-
134
- In [WriterAgent](https://github.com/KeithCu/writeragent), replace `from plugin.scripting.payload_codec import host_pack_data, ...` with `from splitgrid import host_pack_data, ...`. The envelope tag stays `"split_grid"`, dtype `"float64"`, integer-keyed `strings`, and `column_kinds` `int`/`float`/`bool`. Pickle Protocol 5 framing stays in WriterAgent’s `ipc.py` — this package is the codec only.
135
-
136
- ## Releasing
137
-
138
- Tags matching `v*` (for example `v0.1.0`) run `.github/workflows/publish.yml`: cibuildwheel + sdist, then trusted publishing to PyPI (GitHub environment `pypi`, no API token). The workflow file must already be on `main` before you push the tag.
139
-
140
161
  ## License
141
162
 
142
- GPL-3.0-or-later (same as WriterAgent).
163
+ GPL-3.0-or-later.
splitgrid-0.1.0/README.md DELETED
@@ -1,107 +0,0 @@
1
- # SplitGrid
2
-
3
- Asymmetric **split-grid** serialization for rectangular numeric and mixed-type grids.
4
-
5
- This package was **pulled out of [WriterAgent](https://github.com/KeithCu/writeragent)** (LibreOffice Writer/Calc/Draw AI extension). It is the same host flatten + child unpack codec that lived in `plugin.scripting.payload_codec`, plus the optional Cython flatten accelerator from `native/writeragent_vec`. WriterAgent can later depend on `splitgrid` without changing the wire dict.
6
-
7
- Repo: [github.com/KeithCu/writeragent](https://github.com/KeithCu/writeragent)
8
-
9
- ## What split-grid is
10
-
11
- The compute path is **asymmetric by design**:
12
-
13
- - **Host pack** (LibreOffice’s embedded Python, or any NumPy-free interpreter) flattens a 1D or 2D grid into a contiguous `float64` buffer plus a sparse **integer-keyed** `strings` map. Empty cells become `NaN`. Zip codes like `"02138"` stay strings — they are never parsed as floats.
14
- - **Child unpack** (venv with NumPy) materializes a **pure-numeric** grid (`strings == {}`) with `np.frombuffer` (ndarray). Mixed grids use a vectorized object-masking path and return nested lists, restoring `None` for NaN holes.
15
- - **Host unpack** preserves `float('nan')` from the buffer (does **not** coerce holes to `None`). That is the locked egress policy: NaN becomes a Calc error, not a silent blank.
16
-
17
- Grids with fewer than **100 cells** (`BINARY_MIN_CELLS`) stay nested Python lists. Force `"always"` / `"never"` overrides that threshold (used by A/B tests).
18
-
19
- Wire envelope (Pickle5-friendly dict, same keys as WriterAgent):
20
-
21
- ```python
22
- {
23
- "__wa_payload__": "split_grid",
24
- "dtype": "float64",
25
- "column_kinds": ["int", "float"], # per column: int / float / bool
26
- "shape": [rows, cols], # or [n] for 1D
27
- "buffer": b"...", # row-major float64 bytes
28
- "strings": {7: "banana"}, # integer keys, not str(idx)
29
- }
30
- ```
31
-
32
- There is **no datetime lane** on the float64 buffer. Python `datetime` objects stringify into `strings`. Do not add a `'date'` column kind.
33
-
34
- Jagged 2D grids raise `ValueError` (Calc ranges are rectangular).
35
-
36
- ## Python vs Cython
37
-
38
- Host flatten is an optimized **pure-Python** loop:
39
-
40
- - identity type checks (`type(val) is float`, `val is None`)
41
- - bound-method capture (`buf_append = buf.append`)
42
- - `None` as NaN on the fast path
43
- - lazy column-state upgrades
44
- - rectangular validation before the hot loop
45
-
46
- The optional **Cython** module `splitgrid.pack` exposes `fast_flatten_grid_1d` / `fast_flatten_grid_2d`. It is loaded dynamically and **canary-tested** at import. If the extension is missing or fails the canary, the codec uses pure Python. Importing `splitgrid` never requires a compiler.
47
-
48
- ## Install
49
-
50
- ```bash
51
- pip install splitgrid
52
- ```
53
-
54
- PyPI wheels include the compiled Cython flatten accelerator. From a source checkout (Cython is built if a compiler is available):
55
-
56
- ```bash
57
- pip install .
58
- pip install -e ".[test]" # editable + pytest, hypothesis, deal, numpy
59
- pip install -e ".[numpy]" # child unpack/pack
60
- pip install -e ".[verify]" # crosshair-tool
61
- ```
62
-
63
- Host pack stays NumPy-free. Child `child_unpack_*` / `child_pack_*` import NumPy locally.
64
-
65
- ## Test
66
-
67
- ```bash
68
- pytest # default: -m "not slow"
69
- pytest -m "not slow" # same
70
- pytest -m slow # CrossHair subprocess checks (optional)
71
- make verify # deal contracts + hypothesis round-trips, no CrossHair
72
- ```
73
-
74
- Default pytest **does not require** the compiled extension. When `splitgrid.pack` is built, `tests/test_cython_parity.py` compares Cython and Python flatten components on the same grids.
75
-
76
- ## Public API (WriterAgent-compatible)
77
-
78
- ```python
79
- from splitgrid import (
80
- BINARY_MIN_CELLS,
81
- host_pack_split_grid,
82
- host_unpack_split_grid,
83
- child_unpack_split_grid,
84
- child_pack_split_grid,
85
- host_pack_data,
86
- host_unpack_data,
87
- child_unpack_data,
88
- child_pack_result,
89
- is_split_grid,
90
- load_cython_accelerator,
91
- get_cython_status_info,
92
- )
93
- ```
94
-
95
- `host_pack_data(..., force="auto"|"always"|"never")` chooses split-grid vs nested list. `host_pack_multi_data` is a thin `multi_data` wrapper over the same per-grid packing.
96
-
97
- ## How WriterAgent will consume this
98
-
99
- In [WriterAgent](https://github.com/KeithCu/writeragent), replace `from plugin.scripting.payload_codec import host_pack_data, ...` with `from splitgrid import host_pack_data, ...`. The envelope tag stays `"split_grid"`, dtype `"float64"`, integer-keyed `strings`, and `column_kinds` `int`/`float`/`bool`. Pickle Protocol 5 framing stays in WriterAgent’s `ipc.py` — this package is the codec only.
100
-
101
- ## Releasing
102
-
103
- Tags matching `v*` (for example `v0.1.0`) run `.github/workflows/publish.yml`: cibuildwheel + sdist, then trusted publishing to PyPI (GitHub environment `pypi`, no API token). The workflow file must already be on `main` before you push the tag.
104
-
105
- ## License
106
-
107
- GPL-3.0-or-later (same as WriterAgent).
File without changes
File without changes
File without changes
File without changes