arcaeon-distill 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- arcaeon_distill-0.1.0/.gitignore +8 -0
- arcaeon_distill-0.1.0/LICENSE +21 -0
- arcaeon_distill-0.1.0/PKG-INFO +247 -0
- arcaeon_distill-0.1.0/README.md +228 -0
- arcaeon_distill-0.1.0/arcaeon_distill/__init__.py +710 -0
- arcaeon_distill-0.1.0/arcaeon_distill/mcp_server.py +165 -0
- arcaeon_distill-0.1.0/arcaeon_distill/selftest.py +100 -0
- arcaeon_distill-0.1.0/pyproject.toml +31 -0
- arcaeon_distill-0.1.0/test_distill.py +254 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Arcaeon
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,247 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: arcaeon-distill
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Deterministic, cache-stable tool-output distiller for AI agents. Fits more signal in the context budget; keeps a drop receipt for what got cut.
|
|
5
|
+
Project-URL: Homepage, https://arcaeon.io
|
|
6
|
+
Project-URL: Source, https://github.com/Arcaeon-io/arcaeon-distill
|
|
7
|
+
Author: Arcaeon
|
|
8
|
+
License: MIT
|
|
9
|
+
License-File: LICENSE
|
|
10
|
+
Keywords: agents,ai,compaction,context,context-window,deterministic,llm,mcp,prompt-caching,token,tool-output
|
|
11
|
+
Classifier: Intended Audience :: Developers
|
|
12
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Topic :: Software Development :: Libraries
|
|
15
|
+
Requires-Python: >=3.9
|
|
16
|
+
Provides-Extra: ledger
|
|
17
|
+
Requires-Dist: arcaeon-ledger; extra == 'ledger'
|
|
18
|
+
Description-Content-Type: text/markdown
|
|
19
|
+
|
|
20
|
+
# arcaeon-distill
|
|
21
|
+
|
|
22
|
+
<!-- mcp-name: io.arcaeon/distill -->
|
|
23
|
+
|
|
24
|
+
**A big tool output doesn't need to be a big tool output. `arcaeon-distill`
|
|
25
|
+
compacts it under a budget — deterministically, so it doesn't fight your
|
|
26
|
+
provider's prompt cache — and keeps a receipt for exactly what it cut.**
|
|
27
|
+
|
|
28
|
+
```
|
|
29
|
+
pip install arcaeon-distill # then: from arcaeon_distill import distill
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
```python
|
|
33
|
+
from arcaeon_distill import distill
|
|
34
|
+
|
|
35
|
+
huge = call_some_api() # 40k tokens of JSON, most of it noise
|
|
36
|
+
result = distill(huge, budget=500)
|
|
37
|
+
|
|
38
|
+
result.content # the compacted structure — keys kept, values capped
|
|
39
|
+
result.receipt # DropReceipt: prove what got cut, re-fetch if it mattered
|
|
40
|
+
print(result) # "distilled via json: 41302 -> 483 est. tokens (~99% cut, ...)"
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
## Read this before the features: token reduction is not cost reduction
|
|
44
|
+
|
|
45
|
+
A July 2026 paper, **"Token Reduction Is Not Cost Reduction"** (arXiv
|
|
46
|
+
[2607.12161](https://arxiv.org/abs/2607.12161)), ran three token-reduction
|
|
47
|
+
approaches against an unmodified Claude Code baseline and found the
|
|
48
|
+
aggressive setup cut delivered tool-output tokens by **38.4%** while
|
|
49
|
+
**increasing billed cost by 6.8%.** Quoting the abstract directly:
|
|
50
|
+
|
|
51
|
+
> "The largest compression setup reduced delivered tool-output tokens by
|
|
52
|
+
> 38.4% but increased billed cost by 6.8%, while lighter compression produced
|
|
53
|
+
> only small and statistically uncertain savings. Across tasks, token
|
|
54
|
+
> reduction was weakly correlated with cost reduction (Pearson r = 0.15).
|
|
55
|
+
> Cost decomposition shows that prompt-cache creation and reads dominate the
|
|
56
|
+
> measured input-side cost... compression can alter agent trajectories
|
|
57
|
+
> through additional retrieval, diagnosis, testing, and turns, offsetting
|
|
58
|
+
> local token savings. On a SWE-bench Go subset, aggressive compression also
|
|
59
|
+
> reduced successful patch application."
|
|
60
|
+
|
|
61
|
+
Why: providers bill prompt-cache **writes and reads**, not raw token count.
|
|
62
|
+
A compressor whose output shifts from call to call — even for the *same*
|
|
63
|
+
underlying tool result — busts the cached prefix and forces a full
|
|
64
|
+
cache-write every time. Aggressive, unstable pruning also changes agent
|
|
65
|
+
trajectories: extra retrieval and diagnosis turns that eat the local token
|
|
66
|
+
savings, and in that paper's benchmark, sometimes broke the task outright.
|
|
67
|
+
|
|
68
|
+
**So this library makes no cost-savings claim.** It is positioned on
|
|
69
|
+
**context-budget headroom and task reliability** — fit more real signal in
|
|
70
|
+
the window, fewer truncation-driven failures — not on dollars. If you came
|
|
71
|
+
here for a "$ saved" number, that number is not one this library will give
|
|
72
|
+
you, because the paper above shows it usually isn't reliably true. What it
|
|
73
|
+
gives you instead: a smaller, **cache-stable** context footprint, and a
|
|
74
|
+
receipt that tells you when the compaction cut something that mattered.
|
|
75
|
+
|
|
76
|
+
## Deterministic, on purpose
|
|
77
|
+
|
|
78
|
+
The engineering constraint that flows straight from the finding above: **the
|
|
79
|
+
same input at the same budget produces byte-identical output, every run,
|
|
80
|
+
every machine.** No LLM on the hot path (nothing to be nondeterministic
|
|
81
|
+
about — a generative summarizer can't guarantee a byte-identical rewrite
|
|
82
|
+
and will paraphrase IDs into garbage anyway). No wall-clock, no randomness,
|
|
83
|
+
no unstable tie-breaking anywhere in the ranking or truncation logic. Call
|
|
84
|
+
`distill()` twice on the same tool output and a provider's prompt cache sees
|
|
85
|
+
the same bytes twice — not a new prefix to write and bill for.
|
|
86
|
+
|
|
87
|
+
## Three strategies, picked from the input's shape
|
|
88
|
+
|
|
89
|
+
```python
|
|
90
|
+
distill(tool_output, budget=2000, schema_hint=None, query=None, receipt=True)
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
- **json** — dict/list input (or a str that parses as JSON). Every key is
|
|
94
|
+
kept; long string values are truncated with a `"...+412 more chars"`
|
|
95
|
+
count; long list values keep a head/tail slice with an
|
|
96
|
+
`"...+412 more items"` marker where the middle used to be.
|
|
97
|
+
- **tabular** — list-of-dicts, list-of-lists, or CSV/TSV/markdown-table
|
|
98
|
+
text. Keeps the header, a head slice and a tail slice of rows, and a
|
|
99
|
+
dropped-row count between them.
|
|
100
|
+
- **text** — free text. **Deterministic extractive** sentence selection —
|
|
101
|
+
score by position (lead + conclusion weighted over the middle) and, if you
|
|
102
|
+
pass `query`, keyword overlap with it. **Not an LLM call.** Kept sentences
|
|
103
|
+
are reassembled in their original order with a `[...]` gap marker, so the
|
|
104
|
+
surviving text still reads as prose, not a shuffled highlight reel.
|
|
105
|
+
|
|
106
|
+
`schema_hint="json" | "tabular" | "text"` forces a strategy instead of
|
|
107
|
+
auto-detecting. `budget` is an approximate **token** budget (see
|
|
108
|
+
`estimate_tokens`, below — a heuristic, not a real tokenizer count).
|
|
109
|
+
|
|
110
|
+
## The honesty hook: the drop receipt
|
|
111
|
+
|
|
112
|
+
Deterministic extraction is not semantic understanding. Position+keyword
|
|
113
|
+
sentence ranking, head/tail slicing, and length-based truncation are
|
|
114
|
+
mechanical rules — they can, and sometimes will, cut the one line that
|
|
115
|
+
actually mattered. That's exactly why `distill()` doesn't just cut quietly:
|
|
116
|
+
|
|
117
|
+
```python
|
|
118
|
+
result = distill(incident_log, budget=200)
|
|
119
|
+
result.receipt.full # {"digest": "sha256:...", "bytes": 41302}
|
|
120
|
+
result.receipt.distilled # {"digest": "sha256:...", "bytes": 812}
|
|
121
|
+
result.receipt.drops
|
|
122
|
+
# [{"kind": "string_truncated", "path": "body",
|
|
123
|
+
# "digest": "sha256:raw-bytes:v1:...", "dropped_bytes": 40100,
|
|
124
|
+
# "dropped_count": 40100}, ...]
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
Digests are of the **dropped content only** — self-describing
|
|
128
|
+
(`sha256:<recipe>:<version>:<hex>`), compatible with
|
|
129
|
+
[`arcaeon-ledger`](https://pypi.org/project/arcaeon-ledger/)'s format so a
|
|
130
|
+
receipt travels cleanly into a chain, but `arcaeon_distill` never *requires*
|
|
131
|
+
`arcaeon-ledger` to be installed. The kept content is never re-hashed into
|
|
132
|
+
the receipt and the cut content is never carried verbatim — a receipt can be
|
|
133
|
+
logged, shipped, or handed to a third party without leaking what was cut,
|
|
134
|
+
only proving that something was and how much.
|
|
135
|
+
|
|
136
|
+
```python
|
|
137
|
+
from arcaeon_distill import verify_receipt
|
|
138
|
+
|
|
139
|
+
verify_receipt(result.receipt) # self-consistency: schema, digests
|
|
140
|
+
# well-formed, truncated agrees with drops
|
|
141
|
+
```
|
|
142
|
+
|
|
143
|
+
If an agent reads a distilled result and something looks off — a field it
|
|
144
|
+
expected is missing, a count doesn't add up — the receipt's `full.digest`
|
|
145
|
+
lets it prove that *this* receipt describes the tool output it's holding,
|
|
146
|
+
and the drop manifest tells it exactly what to re-fetch. **"Distill, but
|
|
147
|
+
keep the receipt."** That's the differentiator: distillers that lose data
|
|
148
|
+
silently, versus one that's tamper-evidently honest about the loss.
|
|
149
|
+
|
|
150
|
+
Chain a receipt onto a tamper-evident ledger (optional — this is the only
|
|
151
|
+
place `arcaeon-ledger` is ever touched):
|
|
152
|
+
|
|
153
|
+
```python
|
|
154
|
+
pip install arcaeon-distill[ledger] # or: pip install arcaeon-ledger
|
|
155
|
+
|
|
156
|
+
result.receipt.seal("receipts.jsonl", distiller="my-agent-v3")
|
|
157
|
+
# -> chained row, same tamper-evidence as any other arcaeon-ledger entry
|
|
158
|
+
```
|
|
159
|
+
|
|
160
|
+
## `estimate_tokens()` — cheap, and it says so
|
|
161
|
+
|
|
162
|
+
```python
|
|
163
|
+
from arcaeon_distill import estimate_tokens
|
|
164
|
+
estimate_tokens("some text") # ~len(text) // 4
|
|
165
|
+
```
|
|
166
|
+
|
|
167
|
+
A heuristic, not a tokenizer call: no dependency, no model-specific
|
|
168
|
+
vocabulary. It will be wrong, sometimes by a lot, on code, non-English text,
|
|
169
|
+
and highly repetitive strings. Use it to size a budget cheaply — never to
|
|
170
|
+
predict a bill.
|
|
171
|
+
|
|
172
|
+
## Drop it into any MCP agent
|
|
173
|
+
|
|
174
|
+
```json
|
|
175
|
+
{
|
|
176
|
+
"mcpServers": {
|
|
177
|
+
"distill": {
|
|
178
|
+
"command": "python",
|
|
179
|
+
"args": ["-m", "arcaeon_distill.mcp_server"]
|
|
180
|
+
}
|
|
181
|
+
}
|
|
182
|
+
}
|
|
183
|
+
```
|
|
184
|
+
|
|
185
|
+
One tool, `distill_tool_output(tool_output, budget, schema_hint, query,
|
|
186
|
+
receipt)`, returning content + strategy + token estimates + the drop
|
|
187
|
+
receipt. Zero dependencies — MCP is JSON-RPC over stdio and this speaks it
|
|
188
|
+
directly, no SDK. Import-guarded: `arcaeon_distill` itself never imports
|
|
189
|
+
`mcp_server`, so `distill()` works with zero MCP awareness and the server is
|
|
190
|
+
only touched if you run it.
|
|
191
|
+
|
|
192
|
+
## What this does NOT do — read before you assume
|
|
193
|
+
|
|
194
|
+
Being precise about the boundary is the product, not a disclaimer.
|
|
195
|
+
|
|
196
|
+
**1. It does not guarantee cost reduction.** Covered at the top, and worth
|
|
197
|
+
repeating because it's the whole reason this library is shaped the way it
|
|
198
|
+
is: raw token count and billed cost are only weakly correlated under prompt
|
|
199
|
+
caching (arXiv 2607.12161 measured Pearson r = 0.15 across tasks). This
|
|
200
|
+
library's claim is context-budget headroom and reliability, never a dollar
|
|
201
|
+
figure — and it's built deterministic specifically so it doesn't accidentally
|
|
202
|
+
make the caching-cost problem worse.
|
|
203
|
+
|
|
204
|
+
**2. Deterministic extraction is not semantic understanding.** Nothing here
|
|
205
|
+
reads for meaning. Position+keyword sentence ranking and length-based value
|
|
206
|
+
truncation are mechanical rules that can drop the one fact that mattered —
|
|
207
|
+
which is exactly why the drop receipt exists. A pass is not a promise
|
|
208
|
+
nothing important was lost; it's a promise you can check.
|
|
209
|
+
|
|
210
|
+
**3. Budget is best-effort, not a hard cap on pathological input.**
|
|
211
|
+
`distill()` shrinks its internal caps across a bounded number of passes and
|
|
212
|
+
stops. Deeply nested structures, one enormous atomic value with no natural
|
|
213
|
+
cut point below the floor, or degenerate inputs can land over budget. It
|
|
214
|
+
will never loop forever chasing an unreachable target and it will never take
|
|
215
|
+
a different number of shrink passes on the same input twice — determinism
|
|
216
|
+
holds even when the budget isn't hit — but it does not promise the number
|
|
217
|
+
never overshoots.
|
|
218
|
+
|
|
219
|
+
## Complements, doesn't replace
|
|
220
|
+
|
|
221
|
+
- [`arcaeon-dedup`](https://pypi.org/project/arcaeon-dedup/) strips
|
|
222
|
+
near-duplicate text across *multiple* tool outputs (SimHash, zero-dep).
|
|
223
|
+
`arcaeon-distill` shrinks *one* tool output under a budget. Run dedup
|
|
224
|
+
first across a batch, then distill what's left, and you've addressed both
|
|
225
|
+
the "same thing twice" and the "one thing too big" failure modes.
|
|
226
|
+
- [`arcaeon-compact`](https://pypi.org/project/arcaeon-compact/) receipts
|
|
227
|
+
*conversation/memory compaction* (an LLM or heuristic summarizer's claim
|
|
228
|
+
about what it kept). `arcaeon-distill` receipts *single tool-call*
|
|
229
|
+
extraction. Different layer, same honesty pattern, same digest format.
|
|
230
|
+
- [`arcaeon-ledger`](https://pypi.org/project/arcaeon-ledger/) is the
|
|
231
|
+
tamper-evident chain either receipt can seal onto.
|
|
232
|
+
|
|
233
|
+
## Status
|
|
234
|
+
|
|
235
|
+
Pure stdlib (`json`, `hashlib`, `re`, `dataclasses`) — no required
|
|
236
|
+
dependencies. `arcaeon-ledger` is optional, only imported by
|
|
237
|
+
`DropReceipt.seal()`. Tested for determinism (repeated runs collapse to one
|
|
238
|
+
byte-identical output), budget adherence per strategy, and receipt honesty
|
|
239
|
+
(a planted claim mismatch — `truncated=True` with an empty drop list, or the
|
|
240
|
+
reverse — is caught by `verify_receipt`). Ships a runnable self-test with
|
|
241
|
+
frozen golden digest vectors:
|
|
242
|
+
|
|
243
|
+
```
|
|
244
|
+
python -m arcaeon_distill.selftest
|
|
245
|
+
```
|
|
246
|
+
|
|
247
|
+
MIT. Built by Arcaeon — the evidence layer for AI.
|
|
@@ -0,0 +1,228 @@
|
|
|
1
|
+
# arcaeon-distill
|
|
2
|
+
|
|
3
|
+
<!-- mcp-name: io.arcaeon/distill -->
|
|
4
|
+
|
|
5
|
+
**A big tool output doesn't need to be a big tool output. `arcaeon-distill`
|
|
6
|
+
compacts it under a budget — deterministically, so it doesn't fight your
|
|
7
|
+
provider's prompt cache — and keeps a receipt for exactly what it cut.**
|
|
8
|
+
|
|
9
|
+
```
|
|
10
|
+
pip install arcaeon-distill # then: from arcaeon_distill import distill
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
```python
|
|
14
|
+
from arcaeon_distill import distill
|
|
15
|
+
|
|
16
|
+
huge = call_some_api() # 40k tokens of JSON, most of it noise
|
|
17
|
+
result = distill(huge, budget=500)
|
|
18
|
+
|
|
19
|
+
result.content # the compacted structure — keys kept, values capped
|
|
20
|
+
result.receipt # DropReceipt: prove what got cut, re-fetch if it mattered
|
|
21
|
+
print(result) # "distilled via json: 41302 -> 483 est. tokens (~99% cut, ...)"
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
## Read this before the features: token reduction is not cost reduction
|
|
25
|
+
|
|
26
|
+
A July 2026 paper, **"Token Reduction Is Not Cost Reduction"** (arXiv
|
|
27
|
+
[2607.12161](https://arxiv.org/abs/2607.12161)), ran three token-reduction
|
|
28
|
+
approaches against an unmodified Claude Code baseline and found the
|
|
29
|
+
aggressive setup cut delivered tool-output tokens by **38.4%** while
|
|
30
|
+
**increasing billed cost by 6.8%.** Quoting the abstract directly:
|
|
31
|
+
|
|
32
|
+
> "The largest compression setup reduced delivered tool-output tokens by
|
|
33
|
+
> 38.4% but increased billed cost by 6.8%, while lighter compression produced
|
|
34
|
+
> only small and statistically uncertain savings. Across tasks, token
|
|
35
|
+
> reduction was weakly correlated with cost reduction (Pearson r = 0.15).
|
|
36
|
+
> Cost decomposition shows that prompt-cache creation and reads dominate the
|
|
37
|
+
> measured input-side cost... compression can alter agent trajectories
|
|
38
|
+
> through additional retrieval, diagnosis, testing, and turns, offsetting
|
|
39
|
+
> local token savings. On a SWE-bench Go subset, aggressive compression also
|
|
40
|
+
> reduced successful patch application."
|
|
41
|
+
|
|
42
|
+
Why: providers bill prompt-cache **writes and reads**, not raw token count.
|
|
43
|
+
A compressor whose output shifts from call to call — even for the *same*
|
|
44
|
+
underlying tool result — busts the cached prefix and forces a full
|
|
45
|
+
cache-write every time. Aggressive, unstable pruning also changes agent
|
|
46
|
+
trajectories: extra retrieval and diagnosis turns that eat the local token
|
|
47
|
+
savings, and in that paper's benchmark, sometimes broke the task outright.
|
|
48
|
+
|
|
49
|
+
**So this library makes no cost-savings claim.** It is positioned on
|
|
50
|
+
**context-budget headroom and task reliability** — fit more real signal in
|
|
51
|
+
the window, fewer truncation-driven failures — not on dollars. If you came
|
|
52
|
+
here for a "$ saved" number, that number is not one this library will give
|
|
53
|
+
you, because the paper above shows it usually isn't reliably true. What it
|
|
54
|
+
gives you instead: a smaller, **cache-stable** context footprint, and a
|
|
55
|
+
receipt that tells you when the compaction cut something that mattered.
|
|
56
|
+
|
|
57
|
+
## Deterministic, on purpose
|
|
58
|
+
|
|
59
|
+
The engineering constraint that flows straight from the finding above: **the
|
|
60
|
+
same input at the same budget produces byte-identical output, every run,
|
|
61
|
+
every machine.** No LLM on the hot path (nothing to be nondeterministic
|
|
62
|
+
about — a generative summarizer can't guarantee a byte-identical rewrite
|
|
63
|
+
and will paraphrase IDs into garbage anyway). No wall-clock, no randomness,
|
|
64
|
+
no unstable tie-breaking anywhere in the ranking or truncation logic. Call
|
|
65
|
+
`distill()` twice on the same tool output and a provider's prompt cache sees
|
|
66
|
+
the same bytes twice — not a new prefix to write and bill for.
|
|
67
|
+
|
|
68
|
+
## Three strategies, picked from the input's shape
|
|
69
|
+
|
|
70
|
+
```python
|
|
71
|
+
distill(tool_output, budget=2000, schema_hint=None, query=None, receipt=True)
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
- **json** — dict/list input (or a str that parses as JSON). Every key is
|
|
75
|
+
kept; long string values are truncated with a `"...+412 more chars"`
|
|
76
|
+
count; long list values keep a head/tail slice with an
|
|
77
|
+
`"...+412 more items"` marker where the middle used to be.
|
|
78
|
+
- **tabular** — list-of-dicts, list-of-lists, or CSV/TSV/markdown-table
|
|
79
|
+
text. Keeps the header, a head slice and a tail slice of rows, and a
|
|
80
|
+
dropped-row count between them.
|
|
81
|
+
- **text** — free text. **Deterministic extractive** sentence selection —
|
|
82
|
+
score by position (lead + conclusion weighted over the middle) and, if you
|
|
83
|
+
pass `query`, keyword overlap with it. **Not an LLM call.** Kept sentences
|
|
84
|
+
are reassembled in their original order with a `[...]` gap marker, so the
|
|
85
|
+
surviving text still reads as prose, not a shuffled highlight reel.
|
|
86
|
+
|
|
87
|
+
`schema_hint="json" | "tabular" | "text"` forces a strategy instead of
|
|
88
|
+
auto-detecting. `budget` is an approximate **token** budget (see
|
|
89
|
+
`estimate_tokens`, below — a heuristic, not a real tokenizer count).
|
|
90
|
+
|
|
91
|
+
## The honesty hook: the drop receipt
|
|
92
|
+
|
|
93
|
+
Deterministic extraction is not semantic understanding. Position+keyword
|
|
94
|
+
sentence ranking, head/tail slicing, and length-based truncation are
|
|
95
|
+
mechanical rules — they can, and sometimes will, cut the one line that
|
|
96
|
+
actually mattered. That's exactly why `distill()` doesn't just cut quietly:
|
|
97
|
+
|
|
98
|
+
```python
|
|
99
|
+
result = distill(incident_log, budget=200)
|
|
100
|
+
result.receipt.full # {"digest": "sha256:...", "bytes": 41302}
|
|
101
|
+
result.receipt.distilled # {"digest": "sha256:...", "bytes": 812}
|
|
102
|
+
result.receipt.drops
|
|
103
|
+
# [{"kind": "string_truncated", "path": "body",
|
|
104
|
+
# "digest": "sha256:raw-bytes:v1:...", "dropped_bytes": 40100,
|
|
105
|
+
# "dropped_count": 40100}, ...]
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
Digests are of the **dropped content only** — self-describing
|
|
109
|
+
(`sha256:<recipe>:<version>:<hex>`), compatible with
|
|
110
|
+
[`arcaeon-ledger`](https://pypi.org/project/arcaeon-ledger/)'s format so a
|
|
111
|
+
receipt travels cleanly into a chain, but `arcaeon_distill` never *requires*
|
|
112
|
+
`arcaeon-ledger` to be installed. The kept content is never re-hashed into
|
|
113
|
+
the receipt and the cut content is never carried verbatim — a receipt can be
|
|
114
|
+
logged, shipped, or handed to a third party without leaking what was cut,
|
|
115
|
+
only proving that something was and how much.
|
|
116
|
+
|
|
117
|
+
```python
|
|
118
|
+
from arcaeon_distill import verify_receipt
|
|
119
|
+
|
|
120
|
+
verify_receipt(result.receipt) # self-consistency: schema, digests
|
|
121
|
+
# well-formed, truncated agrees with drops
|
|
122
|
+
```
|
|
123
|
+
|
|
124
|
+
If an agent reads a distilled result and something looks off — a field it
|
|
125
|
+
expected is missing, a count doesn't add up — the receipt's `full.digest`
|
|
126
|
+
lets it prove that *this* receipt describes the tool output it's holding,
|
|
127
|
+
and the drop manifest tells it exactly what to re-fetch. **"Distill, but
|
|
128
|
+
keep the receipt."** That's the differentiator: distillers that lose data
|
|
129
|
+
silently, versus one that's tamper-evidently honest about the loss.
|
|
130
|
+
|
|
131
|
+
Chain a receipt onto a tamper-evident ledger (optional — this is the only
|
|
132
|
+
place `arcaeon-ledger` is ever touched):
|
|
133
|
+
|
|
134
|
+
```python
|
|
135
|
+
pip install arcaeon-distill[ledger] # or: pip install arcaeon-ledger
|
|
136
|
+
|
|
137
|
+
result.receipt.seal("receipts.jsonl", distiller="my-agent-v3")
|
|
138
|
+
# -> chained row, same tamper-evidence as any other arcaeon-ledger entry
|
|
139
|
+
```
|
|
140
|
+
|
|
141
|
+
## `estimate_tokens()` — cheap, and it says so
|
|
142
|
+
|
|
143
|
+
```python
|
|
144
|
+
from arcaeon_distill import estimate_tokens
|
|
145
|
+
estimate_tokens("some text") # ~len(text) // 4
|
|
146
|
+
```
|
|
147
|
+
|
|
148
|
+
A heuristic, not a tokenizer call: no dependency, no model-specific
|
|
149
|
+
vocabulary. It will be wrong, sometimes by a lot, on code, non-English text,
|
|
150
|
+
and highly repetitive strings. Use it to size a budget cheaply — never to
|
|
151
|
+
predict a bill.
|
|
152
|
+
|
|
153
|
+
## Drop it into any MCP agent
|
|
154
|
+
|
|
155
|
+
```json
|
|
156
|
+
{
|
|
157
|
+
"mcpServers": {
|
|
158
|
+
"distill": {
|
|
159
|
+
"command": "python",
|
|
160
|
+
"args": ["-m", "arcaeon_distill.mcp_server"]
|
|
161
|
+
}
|
|
162
|
+
}
|
|
163
|
+
}
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
One tool, `distill_tool_output(tool_output, budget, schema_hint, query,
|
|
167
|
+
receipt)`, returning content + strategy + token estimates + the drop
|
|
168
|
+
receipt. Zero dependencies — MCP is JSON-RPC over stdio and this speaks it
|
|
169
|
+
directly, no SDK. Import-guarded: `arcaeon_distill` itself never imports
|
|
170
|
+
`mcp_server`, so `distill()` works with zero MCP awareness and the server is
|
|
171
|
+
only touched if you run it.
|
|
172
|
+
|
|
173
|
+
## What this does NOT do — read before you assume
|
|
174
|
+
|
|
175
|
+
Being precise about the boundary is the product, not a disclaimer.
|
|
176
|
+
|
|
177
|
+
**1. It does not guarantee cost reduction.** Covered at the top, and worth
|
|
178
|
+
repeating because it's the whole reason this library is shaped the way it
|
|
179
|
+
is: raw token count and billed cost are only weakly correlated under prompt
|
|
180
|
+
caching (arXiv 2607.12161 measured Pearson r = 0.15 across tasks). This
|
|
181
|
+
library's claim is context-budget headroom and reliability, never a dollar
|
|
182
|
+
figure — and it's built deterministic specifically so it doesn't accidentally
|
|
183
|
+
make the caching-cost problem worse.
|
|
184
|
+
|
|
185
|
+
**2. Deterministic extraction is not semantic understanding.** Nothing here
|
|
186
|
+
reads for meaning. Position+keyword sentence ranking and length-based value
|
|
187
|
+
truncation are mechanical rules that can drop the one fact that mattered —
|
|
188
|
+
which is exactly why the drop receipt exists. A pass is not a promise
|
|
189
|
+
nothing important was lost; it's a promise you can check.
|
|
190
|
+
|
|
191
|
+
**3. Budget is best-effort, not a hard cap on pathological input.**
|
|
192
|
+
`distill()` shrinks its internal caps across a bounded number of passes and
|
|
193
|
+
stops. Deeply nested structures, one enormous atomic value with no natural
|
|
194
|
+
cut point below the floor, or degenerate inputs can land over budget. It
|
|
195
|
+
will never loop forever chasing an unreachable target and it will never take
|
|
196
|
+
a different number of shrink passes on the same input twice — determinism
|
|
197
|
+
holds even when the budget isn't hit — but it does not promise the number
|
|
198
|
+
never overshoots.
|
|
199
|
+
|
|
200
|
+
## Complements, doesn't replace
|
|
201
|
+
|
|
202
|
+
- [`arcaeon-dedup`](https://pypi.org/project/arcaeon-dedup/) strips
|
|
203
|
+
near-duplicate text across *multiple* tool outputs (SimHash, zero-dep).
|
|
204
|
+
`arcaeon-distill` shrinks *one* tool output under a budget. Run dedup
|
|
205
|
+
first across a batch, then distill what's left, and you've addressed both
|
|
206
|
+
the "same thing twice" and the "one thing too big" failure modes.
|
|
207
|
+
- [`arcaeon-compact`](https://pypi.org/project/arcaeon-compact/) receipts
|
|
208
|
+
*conversation/memory compaction* (an LLM or heuristic summarizer's claim
|
|
209
|
+
about what it kept). `arcaeon-distill` receipts *single tool-call*
|
|
210
|
+
extraction. Different layer, same honesty pattern, same digest format.
|
|
211
|
+
- [`arcaeon-ledger`](https://pypi.org/project/arcaeon-ledger/) is the
|
|
212
|
+
tamper-evident chain either receipt can seal onto.
|
|
213
|
+
|
|
214
|
+
## Status
|
|
215
|
+
|
|
216
|
+
Pure stdlib (`json`, `hashlib`, `re`, `dataclasses`) — no required
|
|
217
|
+
dependencies. `arcaeon-ledger` is optional, only imported by
|
|
218
|
+
`DropReceipt.seal()`. Tested for determinism (repeated runs collapse to one
|
|
219
|
+
byte-identical output), budget adherence per strategy, and receipt honesty
|
|
220
|
+
(a planted claim mismatch — `truncated=True` with an empty drop list, or the
|
|
221
|
+
reverse — is caught by `verify_receipt`). Ships a runnable self-test with
|
|
222
|
+
frozen golden digest vectors:
|
|
223
|
+
|
|
224
|
+
```
|
|
225
|
+
python -m arcaeon_distill.selftest
|
|
226
|
+
```
|
|
227
|
+
|
|
228
|
+
MIT. Built by Arcaeon — the evidence layer for AI.
|