bpp-format 0.3.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Furkan
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,327 @@
1
+ Metadata-Version: 2.4
2
+ Name: bpp-format
3
+ Version: 0.3.0
4
+ Summary: A token-efficient text format for feeding structured data and plans to LLMs, with a lossless JSON/YAML/CSV converter.
5
+ Author: Furkan
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/E7lektronXF/bpp
8
+ Project-URL: Documentation, https://github.com/E7lektronXF/bpp/blob/main/SPEC.md
9
+ Project-URL: Benchmark, https://github.com/E7lektronXF/bpp/blob/main/BENCHMARK.md
10
+ Project-URL: Issues, https://github.com/E7lektronXF/bpp/issues
11
+ Keywords: llm,tokens,json,toon,prompt,serialization,csv,yaml,markdown
12
+ Classifier: Development Status :: 4 - Beta
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: Programming Language :: Python :: 3
15
+ Classifier: Programming Language :: Python :: 3 :: Only
16
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
17
+ Classifier: Topic :: Text Processing
18
+ Requires-Python: >=3.9
19
+ Description-Content-Type: text/markdown
20
+ License-File: LICENSE
21
+ Requires-Dist: PyYAML>=5.4
22
+ Provides-Extra: stats
23
+ Requires-Dist: tiktoken>=0.7; extra == "stats"
24
+ Provides-Extra: api
25
+ Requires-Dist: anthropic>=0.40; extra == "api"
26
+ Provides-Extra: dev
27
+ Requires-Dist: pytest>=7; extra == "dev"
28
+ Requires-Dist: hypothesis>=6; extra == "dev"
29
+ Requires-Dist: tiktoken>=0.7; extra == "dev"
30
+ Dynamic: license-file
31
+
32
+ # bpp — give LLMs your data in ~63% fewer tokens
33
+
34
+ **bpp** converts JSON, YAML, CSV and Markdown plans into `.bpp`, a compact text format that
35
+ language models read with far fewer tokens. You can convert it back without losing anything.
36
+
37
+ * **~63% fewer tokens than pretty JSON, ~36% fewer than [TOON](https://github.com/toon-format/toon)**
38
+ across the benchmark set. The format beats minified JSON, YAML, CSV and even raw Markdown on
39
+ every example.
40
+ * **Lossless round trip:** `decode(encode(x)) == x` for JSON and YAML, with types preserved.
41
+ CSV cells come back byte for byte. Markdown plans keep their structure.
42
+ * **Measured, not invented.** Every syntax choice was benchmarked on two tokenizers. There are no
43
+ made-up symbols or binary tricks, only patterns models already know from YAML, CSV, JSON and
44
+ TypeScript.
45
+ * **One command, one dependency.** `pip install`, then `bpp data.json`.
46
+
47
+ **▶ [Try it in your browser](https://e7lektronxf.github.io/bpp/):** paste your own data and compare
48
+ token counts, with no install needed.
49
+
50
+ 🇹🇷 Türkçe: [README.tr.md](https://github.com/E7lektronXF/bpp/blob/main/README.tr.md)
51
+
52
+ ```
53
+ JSON, pretty-printed (146 tokens) bpp (53 tokens)
54
+
55
+ { bpp3
56
+ "store": "Downtown Branch", store Downtown Branch
57
+ "orders": [ orders[3]{id product qty price customer}
58
+ {"id": 1, "customer": "Alice Johnson", 1 Headphones 2 79.9 Alice Johnson
59
+ "product": "Headphones", "qty": 2, 2 Keyboard 1 129.0 Bob Smith
60
+ "price": 79.9}, 3 Headphones 1 79.9 Carol White
61
+ ...
62
+ ```
63
+
64
+ ## Results
65
+
66
+ Token counts per format for six example files (o200k tokenizer; fewer is better):
67
+
68
+ | data | JSON | minified JSON | YAML | TOON | Markdown | **bpp** | bpp vs best alternative |
69
+ |---|---:|---:|---:|---:|---:|---:|---|
70
+ | 60×10 table (CSV) | 5538 | 3562 | 4388 | 2277 | – | **2042** | −6% vs CSV |
71
+ | nested config (YAML) | 601 | 365 | 443 | 399 | – | **323** | −12% vs minified JSON |
72
+ | project plan with deps (JSON) | 1609 | 990 | 1212 | 1228 | – | **636** | −36% vs minified JSON |
73
+ | Markdown checklist plan | 1501 | 1025 | 1119 | 1093 | 837 | **792** | −5% vs Markdown |
74
+ | API response with nested objects | 3524 | 2231 | 2636 | 2276 | – | **1151** | −48% vs minified JSON |
75
+ | repetitive logs | 5629 | 4109 | 4587 | 3348 | – | **1857** | −42% vs CSV |
76
+ | **total** | 18402 | 12282 | 14385 | 10621 | | **6801** | **−63% vs JSON, −36% vs TOON** |
77
+
78
+ Anthropic's published Claude tokenizer (`claude2`) shows the same picture: −62% vs JSON, −36% vs
79
+ TOON. The full tables, the method and the cases where bpp wins by less are in
80
+ [BENCHMARK.md](https://github.com/E7lektronXF/bpp/blob/main/BENCHMARK.md).
81
+
82
+ ## Quick start
83
+
84
+ ### 1. Install
85
+
86
+ You need Python 3.9 or newer (check with `python --version`). Then run:
87
+
88
+ ```bash
89
+ pip install https://github.com/E7lektronXF/bpp/archive/HEAD.zip
90
+ ```
91
+
92
+ This installs straight from GitHub. You don't need git. If `pip` isn't found, use `python3 -m pip`
93
+ on macOS/Linux or `py -m pip` on Windows.
94
+
95
+ > ⚠️ **Don't run `pip install bpp`.** The package called `bpp` on PyPI is an unrelated project.
96
+
97
+ Optional: `pip install tiktoken` to get exact token counts instead of estimates.
98
+
99
+ ```bash
100
+ bpp --version # bpp 0.3.0
101
+ ```
102
+
103
+ ### 2. Convert a file
104
+
105
+ ```console
106
+ $ bpp quickstart.json
107
+ quickstart.json -> quickstart.bpp (o200k: 122 -> 53 tokens, -57%)
108
+ ```
109
+
110
+ This works with `.json`, `.yaml`, `.csv` and `.md` files. Want a sample to try?
111
+ `curl -O https://raw.githubusercontent.com/E7lektronXF/bpp/HEAD/examples/quickstart.json`
112
+
113
+ ### 3. Give it to your LLM
114
+
115
+ Open `quickstart.bpp`, then paste it into ChatGPT, Claude or your prompt template. If the model
116
+ hasn't seen the format before, add `--primer`. It prepends a one-line explanation (~85 tokens):
117
+
118
+ ```bash
119
+ bpp quickstart.json --primer -o - # print to the terminal instead of writing a file
120
+ ```
121
+
122
+ ```
123
+ bpp3
124
+ # bpp3: JSON as 'key value' lines, 1-space indent nests. k[N]{a b}: N rows of values in column order, last column = rest of line; x? = optional ('-' if absent), x?= columns appear as x=v. "..." = JSON string, *n = &n.
125
+ store Downtown Branch
126
+ ...
127
+ ```
128
+
129
+ ### 4. Convert back
130
+
131
+ ```console
132
+ $ bpp quickstart.bpp -o roundtrip.json
133
+ ```
134
+
135
+ `roundtrip.json` holds exactly the original data. Use `-o file.yaml`, `-o file.csv` or `--to md`
136
+ for other formats. bpp never silently overwrites an existing file; pass `--force` to allow it.
137
+
138
+ ### Compare formats yourself
139
+
140
+ ```console
141
+ $ bpp stats quickstart.json
142
+ format chars o200k claude2 vs JSON (o200k) vs JSON (claude2)
143
+ JSON (indent 2) 434 146 136 +0.0% +0.0%
144
+ JSON (minified) 271 82 81 -43.8% -40.4%
145
+ YAML 261 102 82 -30.1% -39.7%
146
+ bpp 163 53 48 -63.7% -64.7%
147
+ bpp + primer 402 137 137 -6.2% +0.7%
148
+ ```
149
+
150
+ On a file this small the primer eats most of the savings. Use it on larger inputs.
151
+
152
+ ### From Python
153
+
154
+ ```python
155
+ import bpp
156
+
157
+ text = bpp.dumps(data) # Python object -> bpp text (send this to the LLM)
158
+ data = bpp.loads(text) # bpp text -> Python object
159
+
160
+ data = bpp.load("config.yaml") # reads .json .yaml .csv .md .bpp
161
+ bpp.dump(data, "config.bpp") # format chosen by extension
162
+ ```
163
+
164
+ ## The format in one minute
165
+
166
+ ```
167
+ bpp3
168
+ server
169
+ host 0.0.0.0
170
+ port 8080
171
+ tls
172
+ enabled true
173
+ replicas[2]{host port weight}
174
+ db-1.internal 5432 2
175
+ db-2.internal 5432 1
176
+ regions [TR,DE,NL]
177
+ ```
178
+
179
+ * **`key value` lines.** A bare `key` opens a nested object, and nesting is one space of
180
+ indentation. There are no braces, no colons and no quotes unless a value needs them.
181
+ * **Tables:** `name[N]{a b c}` is followed by N rows of values in column order. Keys are written
182
+ once instead of on every object. The last column takes the rest of the line, so free text
183
+ needs no quotes.
184
+ * **Nested data:** `customer.name` is key `name` inside object `customer`, and `>items{...}`
185
+ means the rows indented under a row are its `items`, with their own columns.
186
+ * **Optional columns:** `x?` sits in place, with `-` when absent. `x?=` appears only when present,
187
+ as `x=value`.
188
+ * **Trees:** `>steps` means rows indented under a row are its children. This is how plans are
189
+ written.
190
+ * **Dictionary:** a long value repeated many times is written once as `&0 value` and referenced
191
+ as `*0`, in YAML anchor style. It is used only when it measurably saves tokens.
192
+ * **Strings** are quoted JSON-style only when they would be ambiguous. Unicode stays raw UTF-8.
193
+
194
+ A plan with dependencies (`examples/plan.json`: 1609 tokens as JSON, 636 as bpp):
195
+
196
+ ```
197
+ steps[6]{id:str status priority deps:str owner?= note?= title}>steps
198
+ 1 done P1 [] owner=Ayşe Gereksinim analizi
199
+ 1.1 done P1 [] Paydaş görüşmeleri
200
+ 1.2 done P0 [] Regülasyon incelemesi (BDDK, PCI-DSS)
201
+ 1.3 done P1 [1.1] Kabul kriterlerinin yazılması
202
+ 2 done P1 [1] owner=Mehmet Mimari tasarım
203
+ ```
204
+
205
+ An API response with nested objects and sub-lists (`examples/orders.json`: 3524 tokens as JSON,
206
+ 2231 minified, 1151 as bpp). `customer.city` is a key inside the `customer` object, and each order's
207
+ `items` are the indented rows under it, with their own columns:
208
+
209
+ ```
210
+ orders[20]{order_id customer.city status shipping_method total note customer.name}>items{sku qty unit_price name}
211
+ ORD-2026-00001 Ankara paid standard 2897.68 *4 Ayşe Arslan
212
+ SKU-254 3 335.6 *3
213
+ SKU-940 1 1890.88 Laptop Standı
214
+ ```
215
+
216
+ A Markdown checklist (`examples/project_plan.md`: 837 tokens as Markdown, 792 as bpp). Headings
217
+ and nested lists become a tree; `[ ]` `[x]` `[/]` `[-]` become `todo` `done` `doing` `cancelled`:
218
+
219
+ ```
220
+ steps[5]{status? note?= title}>steps
221
+ - Keşif ve envanter
222
+ done note="Toplam 312 Airflow DAG'i, 48 Spark işi ve 17 Hive veritabanı tespit edildi." Mevcut iş akışlarının envanteri
223
+ done Veri sahipleriyle görüşmeler
224
+ done Pazarlama analitiği
225
+ ```
226
+
227
+ The full grammar and the measurement behind each rule are in [SPEC.md](https://github.com/E7lektronXF/bpp/blob/main/SPEC.md).
228
+
229
+ ## Why it is smaller
230
+
231
+ LLMs read tokens, and tokenizers are trained on English, code and JSON. Invented symbols or binary
232
+ encodings usually cost more tokens and hurt understanding. bpp saves tokens by:
233
+
234
+ 1. **Removing repetition.** Object keys are written once per table, not once per row.
235
+ 2. **Dropping punctuation the tokenizer charges for.** A space merges into the next token; `:`,
236
+ `,` and `"` usually don't. Switching table rows from commas to spaces alone saved 12–22%.
237
+ 3. **Referencing long repeated values** through a small dictionary, but only when the estimated
238
+ gain is positive.
239
+ 4. **Keeping structure explicit:** row counts (`[N]`), column names and indentation give the model
240
+ a frame to read against.
241
+
242
+ ## Guarantees
243
+
244
+ | input | round trip |
245
+ |---|---|
246
+ | JSON / YAML | `loads(dumps(x)) == x`, types included (`1` ≠ `1.0`, `"42"` ≠ `42`, `null`). With `--keep-order`, the JSON text comes back identical, key order included. YAML comments and anchors are not data and are not kept. Dates stay strings. |
247
+ | CSV | Cells and column order come back byte for byte. Numbers are typed only when writing them back gives the same text (`007` and `1.50` stay strings). |
248
+ | Markdown | Structure is kept, formatting is not. Output is normalized Markdown that parses back to the same tree. |
249
+
250
+ Backed by 217 tests: edge cases (empty containers, 80-level nesting, delimiters inside strings,
251
+ multi-line text, Unicode, number-like strings) and hypothesis property tests on random JSON, CSV,
252
+ YAML and Markdown trees.
253
+
254
+ ## Honest caveats
255
+
256
+ * **Understanding is not measured yet.** Token savings are measured. Whether models answer
257
+ questions about bpp as accurately as about JSON is not yet known. A ready-to-run comprehension
258
+ benchmark (10 auto-graded questions per dataset, every format) is included:
259
+ `ANTHROPIC_API_KEY=... python bench/run_qa.py`. The biggest risk is the dictionary: the model
260
+ has to resolve `*3` to its definition.
261
+ * **Token counts are proxies.** They come from tiktoken `o200k_base` and Anthropic's older public
262
+ Claude tokenizer. Current Claude models use a different tokenizer. Set `ANTHROPIC_API_KEY` and
263
+ `bpp stats` adds real `count_tokens` numbers.
264
+ * **Small gains in some cases:** plain flat tables are only ~6% smaller than CSV on o200k (17% on
265
+ claude2). The primer (~85 tokens) cancels the savings on documents of a few hundred tokens.
266
+
267
+ ## Command reference
268
+
269
+ `bpp FILE` covers most uses. It encodes, or decodes if FILE ends in `.bpp`. The full subcommands:
270
+
271
+ ```bash
272
+ bpp encode data.json -o data.bpp [--primer none|short|long] [--no-refs] [--keep-order]
273
+ bpp encode - --from yaml < config.yaml # read from stdin
274
+ bpp decode data.bpp -o data.yaml # output format from the extension
275
+ bpp decode data.bpp --to json --indent -1 # minified JSON
276
+ bpp decode plan.bpp --to md
277
+ bpp stats data.json [--markdown]
278
+ ```
279
+
280
+ | option | effect |
281
+ |---|---|
282
+ | `--primer` | Prepend a format explanation (`short` ~85, `long` ~135 tokens). |
283
+ | `--no-refs` | Disable the `&n`/`*n` dictionary. |
284
+ | `--keep-order` | Never reorder keys. By default a table may move a free-text column such as `title` to the end of each row, which changes JSON key order but not the data. Always on for CSV input. |
285
+ | `-o -` | Write to stdout. |
286
+ | `-f`, `--force` | Allow overwriting an existing file in one-step mode. |
287
+
288
+ ## Troubleshooting
289
+
290
+ | problem | fix |
291
+ |---|---|
292
+ | `bpp: command not found` | Use `python -m bpp ...`, or open a new terminal. |
293
+ | `pip: command not found` | `python3 -m pip install ...` (macOS/Linux) or `py -m pip install ...` (Windows). |
294
+ | `externally-managed-environment` | `pipx install https://github.com/E7lektronXF/bpp/archive/HEAD.zip`, or install inside a virtualenv (`python3 -m venv .venv && . .venv/bin/activate`). |
295
+ | `cannot infer format` | Use one of these extensions: `.json .yaml .yml .csv .md .bpp`. |
296
+ | `... exists; use -o ...` | bpp refused to overwrite your original file. Pick another name with `-o`. |
297
+
298
+ Upgrade with `pip install --upgrade https://github.com/E7lektronXF/bpp/archive/HEAD.zip`. Uninstall
299
+ with `pip uninstall bpp`.
300
+
301
+ ## Development
302
+
303
+ ```bash
304
+ git clone https://github.com/E7lektronXF/bpp.git && cd bpp
305
+ pip install -e ".[dev]" # + pytest, hypothesis, tiktoken
306
+ pytest -q # 217 tests
307
+ python bench/run_tokens.py # regenerate the token benchmark
308
+ python bench/experiments.py # the design experiments behind SPEC.md
309
+ python bench/run_qa.py # comprehension benchmark (needs ANTHROPIC_API_KEY)
310
+ ```
311
+
312
+ The TOON column needs Node.js and `cd bench/toon && npm install`.
313
+
314
+ ```
315
+ SPEC.md format specification and the measurement behind every rule
316
+ js/ JavaScript port (byte-identical output, tested against Python)
317
+ site/ browser playground source; `python site/build.py` writes docs/index.html
318
+ BENCHMARK.md token results, comprehension test, losses and proposed revisions
319
+ src/bpp/ encoder, decoder, converters, CLI
320
+ examples/ sample inputs next to their .bpp output
321
+ bench/ experiments, token benchmark, comprehension benchmark
322
+ tests/ pytest + hypothesis
323
+ ```
324
+
325
+ ## License
326
+
327
+ [MIT](https://github.com/E7lektronXF/bpp/blob/main/LICENSE)
@@ -0,0 +1,296 @@
1
+ # bpp — give LLMs your data in ~63% fewer tokens
2
+
3
+ **bpp** converts JSON, YAML, CSV and Markdown plans into `.bpp`, a compact text format that
4
+ language models read with far fewer tokens. You can convert it back without losing anything.
5
+
6
+ * **~63% fewer tokens than pretty JSON, ~36% fewer than [TOON](https://github.com/toon-format/toon)**
7
+ across the benchmark set. The format beats minified JSON, YAML, CSV and even raw Markdown on
8
+ every example.
9
+ * **Lossless round trip:** `decode(encode(x)) == x` for JSON and YAML, with types preserved.
10
+ CSV cells come back byte for byte. Markdown plans keep their structure.
11
+ * **Measured, not invented.** Every syntax choice was benchmarked on two tokenizers. There are no
12
+ made-up symbols or binary tricks, only patterns models already know from YAML, CSV, JSON and
13
+ TypeScript.
14
+ * **One command, one dependency.** `pip install`, then `bpp data.json`.
15
+
16
+ **▶ [Try it in your browser](https://e7lektronxf.github.io/bpp/):** paste your own data and compare
17
+ token counts, with no install needed.
18
+
19
+ 🇹🇷 Türkçe: [README.tr.md](https://github.com/E7lektronXF/bpp/blob/main/README.tr.md)
20
+
21
+ ```
22
+ JSON, pretty-printed (146 tokens) bpp (53 tokens)
23
+
24
+ { bpp3
25
+ "store": "Downtown Branch", store Downtown Branch
26
+ "orders": [ orders[3]{id product qty price customer}
27
+ {"id": 1, "customer": "Alice Johnson", 1 Headphones 2 79.9 Alice Johnson
28
+ "product": "Headphones", "qty": 2, 2 Keyboard 1 129.0 Bob Smith
29
+ "price": 79.9}, 3 Headphones 1 79.9 Carol White
30
+ ...
31
+ ```
32
+
33
+ ## Results
34
+
35
+ Token counts per format for six example files (o200k tokenizer; fewer is better):
36
+
37
+ | data | JSON | minified JSON | YAML | TOON | Markdown | **bpp** | bpp vs best alternative |
38
+ |---|---:|---:|---:|---:|---:|---:|---|
39
+ | 60×10 table (CSV) | 5538 | 3562 | 4388 | 2277 | – | **2042** | −6% vs CSV |
40
+ | nested config (YAML) | 601 | 365 | 443 | 399 | – | **323** | −12% vs minified JSON |
41
+ | project plan with deps (JSON) | 1609 | 990 | 1212 | 1228 | – | **636** | −36% vs minified JSON |
42
+ | Markdown checklist plan | 1501 | 1025 | 1119 | 1093 | 837 | **792** | −5% vs Markdown |
43
+ | API response with nested objects | 3524 | 2231 | 2636 | 2276 | – | **1151** | −48% vs minified JSON |
44
+ | repetitive logs | 5629 | 4109 | 4587 | 3348 | – | **1857** | −42% vs CSV |
45
+ | **total** | 18402 | 12282 | 14385 | 10621 | | **6801** | **−63% vs JSON, −36% vs TOON** |
46
+
47
+ Anthropic's published Claude tokenizer (`claude2`) shows the same picture: −62% vs JSON, −36% vs
48
+ TOON. The full tables, the method and the cases where bpp wins by less are in
49
+ [BENCHMARK.md](https://github.com/E7lektronXF/bpp/blob/main/BENCHMARK.md).
50
+
51
+ ## Quick start
52
+
53
+ ### 1. Install
54
+
55
+ You need Python 3.9 or newer (check with `python --version`). Then run:
56
+
57
+ ```bash
58
+ pip install https://github.com/E7lektronXF/bpp/archive/HEAD.zip
59
+ ```
60
+
61
+ This installs straight from GitHub. You don't need git. If `pip` isn't found, use `python3 -m pip`
62
+ on macOS/Linux or `py -m pip` on Windows.
63
+
64
+ > ⚠️ **Don't run `pip install bpp`.** The package called `bpp` on PyPI is an unrelated project.
65
+
66
+ Optional: `pip install tiktoken` to get exact token counts instead of estimates.
67
+
68
+ ```bash
69
+ bpp --version # bpp 0.3.0
70
+ ```
71
+
72
+ ### 2. Convert a file
73
+
74
+ ```console
75
+ $ bpp quickstart.json
76
+ quickstart.json -> quickstart.bpp (o200k: 122 -> 53 tokens, -57%)
77
+ ```
78
+
79
+ This works with `.json`, `.yaml`, `.csv` and `.md` files. Want a sample to try?
80
+ `curl -O https://raw.githubusercontent.com/E7lektronXF/bpp/HEAD/examples/quickstart.json`
81
+
82
+ ### 3. Give it to your LLM
83
+
84
+ Open `quickstart.bpp`, then paste it into ChatGPT, Claude or your prompt template. If the model
85
+ hasn't seen the format before, add `--primer`. It prepends a one-line explanation (~85 tokens):
86
+
87
+ ```bash
88
+ bpp quickstart.json --primer -o - # print to the terminal instead of writing a file
89
+ ```
90
+
91
+ ```
92
+ bpp3
93
+ # bpp3: JSON as 'key value' lines, 1-space indent nests. k[N]{a b}: N rows of values in column order, last column = rest of line; x? = optional ('-' if absent), x?= columns appear as x=v. "..." = JSON string, *n = &n.
94
+ store Downtown Branch
95
+ ...
96
+ ```
97
+
98
+ ### 4. Convert back
99
+
100
+ ```console
101
+ $ bpp quickstart.bpp -o roundtrip.json
102
+ ```
103
+
104
+ `roundtrip.json` holds exactly the original data. Use `-o file.yaml`, `-o file.csv` or `--to md`
105
+ for other formats. bpp never silently overwrites an existing file; pass `--force` to allow it.
106
+
107
+ ### Compare formats yourself
108
+
109
+ ```console
110
+ $ bpp stats quickstart.json
111
+ format chars o200k claude2 vs JSON (o200k) vs JSON (claude2)
112
+ JSON (indent 2) 434 146 136 +0.0% +0.0%
113
+ JSON (minified) 271 82 81 -43.8% -40.4%
114
+ YAML 261 102 82 -30.1% -39.7%
115
+ bpp 163 53 48 -63.7% -64.7%
116
+ bpp + primer 402 137 137 -6.2% +0.7%
117
+ ```
118
+
119
+ On a file this small the primer eats most of the savings. Use it on larger inputs.
120
+
121
+ ### From Python
122
+
123
+ ```python
124
+ import bpp
125
+
126
+ text = bpp.dumps(data) # Python object -> bpp text (send this to the LLM)
127
+ data = bpp.loads(text) # bpp text -> Python object
128
+
129
+ data = bpp.load("config.yaml") # reads .json .yaml .csv .md .bpp
130
+ bpp.dump(data, "config.bpp") # format chosen by extension
131
+ ```
132
+
133
+ ## The format in one minute
134
+
135
+ ```
136
+ bpp3
137
+ server
138
+ host 0.0.0.0
139
+ port 8080
140
+ tls
141
+ enabled true
142
+ replicas[2]{host port weight}
143
+ db-1.internal 5432 2
144
+ db-2.internal 5432 1
145
+ regions [TR,DE,NL]
146
+ ```
147
+
148
+ * **`key value` lines.** A bare `key` opens a nested object, and nesting is one space of
149
+ indentation. There are no braces, no colons and no quotes unless a value needs them.
150
+ * **Tables:** `name[N]{a b c}` is followed by N rows of values in column order. Keys are written
151
+ once instead of on every object. The last column takes the rest of the line, so free text
152
+ needs no quotes.
153
+ * **Nested data:** `customer.name` is key `name` inside object `customer`, and `>items{...}`
154
+ means the rows indented under a row are its `items`, with their own columns.
155
+ * **Optional columns:** `x?` sits in place, with `-` when absent. `x?=` appears only when present,
156
+ as `x=value`.
157
+ * **Trees:** `>steps` means rows indented under a row are its children. This is how plans are
158
+ written.
159
+ * **Dictionary:** a long value repeated many times is written once as `&0 value` and referenced
160
+ as `*0`, in YAML anchor style. It is used only when it measurably saves tokens.
161
+ * **Strings** are quoted JSON-style only when they would be ambiguous. Unicode stays raw UTF-8.
162
+
163
+ A plan with dependencies (`examples/plan.json`: 1609 tokens as JSON, 636 as bpp):
164
+
165
+ ```
166
+ steps[6]{id:str status priority deps:str owner?= note?= title}>steps
167
+ 1 done P1 [] owner=Ayşe Gereksinim analizi
168
+ 1.1 done P1 [] Paydaş görüşmeleri
169
+ 1.2 done P0 [] Regülasyon incelemesi (BDDK, PCI-DSS)
170
+ 1.3 done P1 [1.1] Kabul kriterlerinin yazılması
171
+ 2 done P1 [1] owner=Mehmet Mimari tasarım
172
+ ```
173
+
174
+ An API response with nested objects and sub-lists (`examples/orders.json`: 3524 tokens as JSON,
175
+ 2231 minified, 1151 as bpp). `customer.city` is a key inside the `customer` object, and each order's
176
+ `items` are the indented rows under it, with their own columns:
177
+
178
+ ```
179
+ orders[20]{order_id customer.city status shipping_method total note customer.name}>items{sku qty unit_price name}
180
+ ORD-2026-00001 Ankara paid standard 2897.68 *4 Ayşe Arslan
181
+ SKU-254 3 335.6 *3
182
+ SKU-940 1 1890.88 Laptop Standı
183
+ ```
184
+
185
+ A Markdown checklist (`examples/project_plan.md`: 837 tokens as Markdown, 792 as bpp). Headings
186
+ and nested lists become a tree; `[ ]` `[x]` `[/]` `[-]` become `todo` `done` `doing` `cancelled`:
187
+
188
+ ```
189
+ steps[5]{status? note?= title}>steps
190
+ - Keşif ve envanter
191
+ done note="Toplam 312 Airflow DAG'i, 48 Spark işi ve 17 Hive veritabanı tespit edildi." Mevcut iş akışlarının envanteri
192
+ done Veri sahipleriyle görüşmeler
193
+ done Pazarlama analitiği
194
+ ```
195
+
196
+ The full grammar and the measurement behind each rule are in [SPEC.md](https://github.com/E7lektronXF/bpp/blob/main/SPEC.md).
197
+
198
+ ## Why it is smaller
199
+
200
+ LLMs read tokens, and tokenizers are trained on English, code and JSON. Invented symbols or binary
201
+ encodings usually cost more tokens and hurt understanding. bpp saves tokens by:
202
+
203
+ 1. **Removing repetition.** Object keys are written once per table, not once per row.
204
+ 2. **Dropping punctuation the tokenizer charges for.** A space merges into the next token; `:`,
205
+ `,` and `"` usually don't. Switching table rows from commas to spaces alone saved 12–22%.
206
+ 3. **Referencing long repeated values** through a small dictionary, but only when the estimated
207
+ gain is positive.
208
+ 4. **Keeping structure explicit:** row counts (`[N]`), column names and indentation give the model
209
+ a frame to read against.
210
+
211
+ ## Guarantees
212
+
213
+ | input | round trip |
214
+ |---|---|
215
+ | JSON / YAML | `loads(dumps(x)) == x`, types included (`1` ≠ `1.0`, `"42"` ≠ `42`, `null`). With `--keep-order`, the JSON text comes back identical, key order included. YAML comments and anchors are not data and are not kept. Dates stay strings. |
216
+ | CSV | Cells and column order come back byte for byte. Numbers are typed only when writing them back gives the same text (`007` and `1.50` stay strings). |
217
+ | Markdown | Structure is kept, formatting is not. Output is normalized Markdown that parses back to the same tree. |
218
+
219
+ Backed by 217 tests: edge cases (empty containers, 80-level nesting, delimiters inside strings,
220
+ multi-line text, Unicode, number-like strings) and hypothesis property tests on random JSON, CSV,
221
+ YAML and Markdown trees.
222
+
223
+ ## Honest caveats
224
+
225
+ * **Understanding is not measured yet.** Token savings are measured. Whether models answer
226
+ questions about bpp as accurately as about JSON is not yet known. A ready-to-run comprehension
227
+ benchmark (10 auto-graded questions per dataset, every format) is included:
228
+ `ANTHROPIC_API_KEY=... python bench/run_qa.py`. The biggest risk is the dictionary: the model
229
+ has to resolve `*3` to its definition.
230
+ * **Token counts are proxies.** They come from tiktoken `o200k_base` and Anthropic's older public
231
+ Claude tokenizer. Current Claude models use a different tokenizer. Set `ANTHROPIC_API_KEY` and
232
+ `bpp stats` adds real `count_tokens` numbers.
233
+ * **Small gains in some cases:** plain flat tables are only ~6% smaller than CSV on o200k (17% on
234
+ claude2). The primer (~85 tokens) cancels the savings on documents of a few hundred tokens.
235
+
236
+ ## Command reference
237
+
238
+ `bpp FILE` covers most uses. It encodes, or decodes if FILE ends in `.bpp`. The full subcommands:
239
+
240
+ ```bash
241
+ bpp encode data.json -o data.bpp [--primer none|short|long] [--no-refs] [--keep-order]
242
+ bpp encode - --from yaml < config.yaml # read from stdin
243
+ bpp decode data.bpp -o data.yaml # output format from the extension
244
+ bpp decode data.bpp --to json --indent -1 # minified JSON
245
+ bpp decode plan.bpp --to md
246
+ bpp stats data.json [--markdown]
247
+ ```
248
+
249
+ | option | effect |
250
+ |---|---|
251
+ | `--primer` | Prepend a format explanation (`short` ~85, `long` ~135 tokens). |
252
+ | `--no-refs` | Disable the `&n`/`*n` dictionary. |
253
+ | `--keep-order` | Never reorder keys. By default a table may move a free-text column such as `title` to the end of each row, which changes JSON key order but not the data. Always on for CSV input. |
254
+ | `-o -` | Write to stdout. |
255
+ | `-f`, `--force` | Allow overwriting an existing file in one-step mode. |
256
+
257
+ ## Troubleshooting
258
+
259
+ | problem | fix |
260
+ |---|---|
261
+ | `bpp: command not found` | Use `python -m bpp ...`, or open a new terminal. |
262
+ | `pip: command not found` | `python3 -m pip install ...` (macOS/Linux) or `py -m pip install ...` (Windows). |
263
+ | `externally-managed-environment` | `pipx install https://github.com/E7lektronXF/bpp/archive/HEAD.zip`, or install inside a virtualenv (`python3 -m venv .venv && . .venv/bin/activate`). |
264
+ | `cannot infer format` | Use one of these extensions: `.json .yaml .yml .csv .md .bpp`. |
265
+ | `... exists; use -o ...` | bpp refused to overwrite your original file. Pick another name with `-o`. |
266
+
267
+ Upgrade with `pip install --upgrade https://github.com/E7lektronXF/bpp/archive/HEAD.zip`. Uninstall
268
+ with `pip uninstall bpp`.
269
+
270
+ ## Development
271
+
272
+ ```bash
273
+ git clone https://github.com/E7lektronXF/bpp.git && cd bpp
274
+ pip install -e ".[dev]" # + pytest, hypothesis, tiktoken
275
+ pytest -q # 217 tests
276
+ python bench/run_tokens.py # regenerate the token benchmark
277
+ python bench/experiments.py # the design experiments behind SPEC.md
278
+ python bench/run_qa.py # comprehension benchmark (needs ANTHROPIC_API_KEY)
279
+ ```
280
+
281
+ The TOON column needs Node.js and `cd bench/toon && npm install`.
282
+
283
+ ```
284
+ SPEC.md format specification and the measurement behind every rule
285
+ js/ JavaScript port (byte-identical output, tested against Python)
286
+ site/ browser playground source; `python site/build.py` writes docs/index.html
287
+ BENCHMARK.md token results, comprehension test, losses and proposed revisions
288
+ src/bpp/ encoder, decoder, converters, CLI
289
+ examples/ sample inputs next to their .bpp output
290
+ bench/ experiments, token benchmark, comprehension benchmark
291
+ tests/ pytest + hypothesis
292
+ ```
293
+
294
+ ## License
295
+
296
+ [MIT](https://github.com/E7lektronXF/bpp/blob/main/LICENSE)