tokenmizer 0.3.1__tar.gz → 0.3.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (119) hide show
  1. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/.claude-plugin/plugin.json +1 -1
  2. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/.github/workflows/ci.yml +26 -16
  3. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/.github/workflows/release.yml +14 -0
  4. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/CHANGELOG.md +161 -2
  5. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/PKG-INFO +137 -13
  6. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/README.md +136 -12
  7. tokenmizer-0.3.2/docs/DEMO_SCRIPT.md +142 -0
  8. tokenmizer-0.3.2/docs/assets/demo.gif +0 -0
  9. tokenmizer-0.3.2/docs/superpowers/plans/2026-07-10-audit-fixes.md +90 -0
  10. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/pyproject.toml +1 -1
  11. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/scripts/gen_demo_gif.py +10 -26
  12. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/server.json +2 -2
  13. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/integration/test_api_endpoint.py +70 -2
  14. tokenmizer-0.3.2/tests/unit/test_decision_tracker.py +174 -0
  15. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/unit/test_graph.py +59 -2
  16. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/unit/test_hybrid_extractor.py +80 -0
  17. tokenmizer-0.3.2/tests/unit/test_mcp_server.py +198 -0
  18. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/unit/test_security.py +23 -0
  19. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/unit/test_validator.py +46 -0
  20. tokenmizer-0.3.2/tests/unit/test_version_consistency.py +68 -0
  21. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/__init__.py +1 -1
  22. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/api/app.py +25 -8
  23. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/cli.py +13 -0
  24. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/compression/engine.py +12 -1
  25. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/filters/file_intelligence.py +15 -1
  26. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/graph_memory/decision_tracker.py +108 -24
  27. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/graph_memory/graph.py +85 -4
  28. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/graph_memory/hybrid_extractor.py +47 -4
  29. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/graph_memory/validator.py +28 -0
  30. tokenmizer-0.3.2/tokenmizer/graph_memory/visualization.py +559 -0
  31. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/mcp/server.py +154 -60
  32. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/security/redaction.py +21 -0
  33. tokenmizer-0.3.1/docs/assets/demo.gif +0 -0
  34. tokenmizer-0.3.1/tokenmizer/graph_memory/visualization.py +0 -313
  35. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/.claude-plugin/marketplace.json +0 -0
  36. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/.claude-plugin/skills/analyze/SKILL.md +0 -0
  37. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/.claude-plugin/skills/checkpoint/SKILL.md +0 -0
  38. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/.claude-plugin/skills/resume/SKILL.md +0 -0
  39. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/.claude-plugin/skills/stats/SKILL.md +0 -0
  40. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/.github/ISSUE_TEMPLATE/bug_report.md +0 -0
  41. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/.github/ISSUE_TEMPLATE/extraction_miss.md +0 -0
  42. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/.github/PULL_REQUEST_TEMPLATE.md +0 -0
  43. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/.gitignore +0 -0
  44. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/.mcp.json +0 -0
  45. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/CONTRIBUTING.md +0 -0
  46. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/Dockerfile +0 -0
  47. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/LICENSE +0 -0
  48. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/SECURITY.md +0 -0
  49. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/TESTING.md +0 -0
  50. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/USAGE.md +0 -0
  51. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/benchmarks/__init__.py +0 -0
  52. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/benchmarks/checkpoint_accuracy/__init__.py +0 -0
  53. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/benchmarks/checkpoint_accuracy/runner.py +0 -0
  54. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/benchmarks/checkpoint_accuracy/runner_v2.py +0 -0
  55. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/benchmarks/checkpoint_accuracy/runner_v3.py +0 -0
  56. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/benchmarks/graph_retrieval/__init__.py +0 -0
  57. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/benchmarks/graph_retrieval/runner.py +0 -0
  58. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/benchmarks/latency/__init__.py +0 -0
  59. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/benchmarks/latency/runner.py +0 -0
  60. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/benchmarks/resume_quality/__init__.py +0 -0
  61. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/docker-compose.yml +0 -0
  62. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/docs/assets/architecture.svg +0 -0
  63. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/docs/assets/logo.svg +0 -0
  64. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/examples/basic_usage.py +0 -0
  65. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/scripts/install.sh +0 -0
  66. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/scripts/mcp_e2e_check.py +0 -0
  67. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/scripts/run_stdlib_tests.py +0 -0
  68. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/scripts/setup.sh +0 -0
  69. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/scripts/static_audit.py +0 -0
  70. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/__init__.py +0 -0
  71. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/chaos/__init__.py +0 -0
  72. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/chaos/test_recovery.py +0 -0
  73. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/conftest.py +0 -0
  74. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/integration/__init__.py +0 -0
  75. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/integration/test_checkpoint.py +0 -0
  76. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/memory_accuracy/__init__.py +0 -0
  77. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/memory_accuracy/test_retention.py +0 -0
  78. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/unit/__init__.py +0 -0
  79. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/unit/test_cache.py +0 -0
  80. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/unit/test_compression_correctness.py +0 -0
  81. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/unit/test_decision_cache_async.py +0 -0
  82. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/unit/test_file_intelligence.py +0 -0
  83. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/unit/test_graph_persistence.py +0 -0
  84. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/unit/test_rate_limiter.py +0 -0
  85. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tests/unit/test_tokenizer.py +0 -0
  86. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/agents/__init__.py +0 -0
  87. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/analytics/__init__.py +0 -0
  88. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/analytics/engine.py +0 -0
  89. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/api/__init__.py +0 -0
  90. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/api/rate_limiter.py +0 -0
  91. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/checkpoints/__init__.py +0 -0
  92. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/checkpoints/manager.py +0 -0
  93. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/compression/__init__.py +0 -0
  94. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/compression/output_trimmer.py +0 -0
  95. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/compression/window.py +0 -0
  96. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/config/__init__.py +0 -0
  97. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/config/settings.py +0 -0
  98. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/core/__init__.py +0 -0
  99. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/core/dto.py +0 -0
  100. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/core/errors.py +0 -0
  101. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/core/tokenizer.py +0 -0
  102. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/dashboard/__init__.py +0 -0
  103. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/dashboard/page.py +0 -0
  104. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/filters/__init__.py +0 -0
  105. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/graph_memory/__init__.py +0 -0
  106. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/graph_memory/helpers.py +0 -0
  107. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/graph_memory/types.py +0 -0
  108. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/mcp/__init__.py +0 -0
  109. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/providers/__init__.py +0 -0
  110. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/providers/providers.py +0 -0
  111. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/security/__init__.py +0 -0
  112. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/security/auth.py +0 -0
  113. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/security/middleware.py +0 -0
  114. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/semantic_cache/__init__.py +0 -0
  115. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/semantic_cache/cache.py +0 -0
  116. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/state/__init__.py +0 -0
  117. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/state/backend.py +0 -0
  118. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer/storage/__init__.py +0 -0
  119. {tokenmizer-0.3.1 → tokenmizer-0.3.2}/tokenmizer.yaml +0 -0
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "tokenmizer",
3
- "version": "0.2.3",
3
+ "version": "0.3.2",
4
4
  "description": "Never lose your AI context again. Graph-backed memory, session checkpointing, and file intelligence for Claude Code and any LLM.",
5
5
  "author": {
6
6
  "name": "Shweta Mishra",
@@ -8,10 +8,22 @@ on:
8
8
 
9
9
  jobs:
10
10
  test:
11
- runs-on: ubuntu-latest
11
+ # Windows is in the matrix because the 2026-07-10 audit found THREE
12
+ # Windows-only bugs Linux CI could never catch: a cp1252
13
+ # UnicodeEncodeError crashing `tokenmizer --help`, a leaked SQLite
14
+ # handle blocking corrupt-DB recovery (WinError 32), and console
15
+ # encoding issues in scripts. One Windows leg (latest Python only —
16
+ # Windows runners are slow) keeps that whole bug class from ever
17
+ # shipping again.
18
+ runs-on: ${{ matrix.os }}
12
19
  strategy:
20
+ fail-fast: false
13
21
  matrix:
22
+ os: [ubuntu-latest]
14
23
  python-version: ["3.10", "3.11", "3.12"]
24
+ include:
25
+ - os: windows-latest
26
+ python-version: "3.12"
15
27
 
16
28
  steps:
17
29
  - uses: actions/checkout@v4
@@ -20,14 +32,7 @@ jobs:
20
32
  uses: actions/setup-python@v5
21
33
  with:
22
34
  python-version: ${{ matrix.python-version }}
23
-
24
- - name: Cache pip
25
- uses: actions/cache@v4
26
- with:
27
- path: ~/.cache/pip
28
- key: ${{ runner.os }}-pip-${{ hashFiles('pyproject.toml') }}
29
- restore-keys: |
30
- ${{ runner.os }}-pip-
35
+ cache: pip
31
36
 
32
37
  - name: Install dependencies
33
38
  run: |
@@ -40,16 +45,21 @@ jobs:
40
45
  # Coverage threshold lives in pyproject.toml [tool.coverage.report]
41
46
  # fail_under — single source of truth, do NOT override it here.
42
47
  - name: Run tests with coverage
43
- run: |
44
- pytest tests/ \
45
- --cov=tokenmizer \
46
- --cov-report=term-missing \
47
- --cov-report=xml \
48
- -v
48
+ run: pytest tests/ --cov=tokenmizer --cov-report=term-missing --cov-report=xml -v
49
+
50
+ # Regression guard for the cp1252 crash: --help must work on a
51
+ # non-UTF-8 Windows console, not just under pytest's captured IO.
52
+ - name: CLI smoke test (Windows encoding regression)
53
+ if: runner.os == 'Windows'
54
+ run: python -m tokenmizer.cli --help
55
+
56
+ - name: MCP stdio e2e
57
+ if: matrix.python-version == '3.12'
58
+ run: python scripts/mcp_e2e_check.py
49
59
 
50
60
  - name: Upload coverage
51
61
  uses: codecov/codecov-action@v4
52
- if: matrix.python-version == '3.12'
62
+ if: matrix.python-version == '3.12' && runner.os == 'Linux'
53
63
  with:
54
64
  file: ./coverage.xml
55
65
  continue-on-error: true
@@ -24,6 +24,20 @@ jobs:
24
24
  with:
25
25
  python-version: "3.12"
26
26
 
27
+ # The release tag (v0.3.2) must match the package version — otherwise
28
+ # a stale tag publishes the wrong code under the wrong number. Version
29
+ # pins across server.json / plugin.json / app.py are enforced by
30
+ # tests/unit/test_version_consistency.py in the test run below.
31
+ - name: Verify tag matches package version
32
+ run: |
33
+ pip install -e .
34
+ PKG_VERSION=$(python -c "import tokenmizer; print(tokenmizer.__version__)")
35
+ TAG_VERSION="${GITHUB_REF_NAME#v}"
36
+ if [ "$PKG_VERSION" != "$TAG_VERSION" ]; then
37
+ echo "::error::Tag $GITHUB_REF_NAME does not match package version $PKG_VERSION"
38
+ exit 1
39
+ fi
40
+
27
41
  - name: Run tests first — never publish a broken build
28
42
  run: |
29
43
  pip install -e ".[dev]"
@@ -1,5 +1,160 @@
1
1
  # Changelog
2
2
 
3
+ ## [0.3.2] — 2026-07-10 — full-repo audit: graph memory, MCP server, visualization
4
+
5
+ ### Critical — the LLM extraction path never worked
6
+ - **`api/app.py`:** the background extraction called
7
+ `HybridExtractor(provider_fn=_pfn)` — a kwarg `__init__` never accepted —
8
+ so it raised `TypeError` on EVERY call, and even without that,
9
+ `ext.extract(_msgs)` omitted `provider_fn`, which defaults to `None` and
10
+ skips the LLM pass. Net effect: **`use_llm_extraction: true` has never
11
+ once produced an LLM extraction** — every call raised, was caught by the
12
+ broad except, and logged as if it were a transient provider failure.
13
+ Fixed (`HybridExtractor()` + `extract(_msgs, provider_fn=_pfn)`) with a
14
+ regression test that replicates the exact call pattern.
15
+
16
+ ### Graph memory — topic classifier
17
+ - **"Go" the verb collided with Go the language:** "Go with tRPC for the
18
+ API layer" classified as `language` (first-single-word-hit-wins never
19
+ reached "trpc"), so a later "use gRPC instead" never superseded it.
20
+ Bare "go" removed from language keywords; unambiguous forms kept
21
+ ("golang", "in go", "use go", ...).
22
+ - **Vocabulary drift:** hybrid_extractor.py knew supabase/clerk/etc., the
23
+ classifier didn't — those decisions classified as None and supersession
24
+ silently never fired. New buckets: backend_platform, auth_provider,
25
+ payments, observability, state_management, package_manager, styling.
26
+ - **Multi-topic decisions collapsed to one topic:** "Use FastAPI with
27
+ SQLAlchemy and PostgreSQL" returned only `web_framework`; a later
28
+ Postgres→SQLite switch was never detected as contradicting it. New
29
+ `classify_topics()` returns ALL matched topics; contradiction detection
30
+ is now set-intersection. Bigrams match first and consume their words
31
+ ("session store" no longer leaks a spurious `auth_mechanism`).
32
+ `classify_topic()` kept as a backward-compatible wrapper.
33
+ - **`ARCHIVED` was unreachable:** documented, fully wired into decay/
34
+ prune/query logic — and nothing ever set it. The README advertised a
35
+ 4-state model whose 4th state could not occur. SUPERSEDED decisions now
36
+ age into ARCHIVED after 7 days (from supersession time, not creation).
37
+ - `/api/decision/invalidate` substring-matched across ALL decision nodes
38
+ regardless of status — a short label could flip already-SUPERSEDED
39
+ history nodes to INVALIDATED, destroying their supersession record.
40
+ Now only ACTIVE decisions are eligible, and the response lists exactly
41
+ which nodes were affected.
42
+
43
+ ### MCP server — hardened like a client integration
44
+ - **`isError` was string-sniffing** (`result_text.startswith("❌")`): a
45
+ missing required argument produced `"Tool error: 'session_id'"` with
46
+ `isError: false` — MCP clients saw a *successful* result. Handlers now
47
+ return `(text, is_error)` structurally; missing/invalid args produce
48
+ typed validation errors with `isError: true`.
49
+ - **Three server-killer inputs fixed:** an exception inside `initialize`
50
+ crashed the whole stdio loop with no JSON-RPC error (client hung, then
51
+ watched the subprocess die); a valid-JSON-but-not-an-object line
52
+ (`[1,2]`) raised AttributeError and killed the loop; malformed JSON was
53
+ silently dropped (`continue` — the request hung forever, no log, no
54
+ error). Per-request try/except now returns `-32603`, non-objects get
55
+ `-32600`, parse failures get `-32700` + a warning log.
56
+ - **`logger` was dead code** — defined, never called once. Every caught
57
+ exception was invisible to the operator. `run_stdio_server()` now
58
+ configures stderr logging (stdout stays protocol-only), controlled by
59
+ `TOKENMIZER_MCP_LOG_LEVEL`.
60
+ - Input validation: `session_id`/`file_path` presence+type, `level` enum,
61
+ `token_budget` positive-int (bool explicitly rejected — it's an int
62
+ subclass), non-dict `arguments`.
63
+ - 18 new unit tests (`tests/unit/test_mcp_server.py`) covering all of the
64
+ above; `scripts/mcp_e2e_check.py` still ALL PASS.
65
+ - `server.json` version synced (was 0.2.6, three releases stale).
66
+
67
+ ### Silent failures made visible (continuing the v0.2.x hardening)
68
+ - `hybrid_extractor.llm_extract`: debug→warning (a provider outage
69
+ silently degraded every turn to heuristic-only).
70
+ - `compression/engine.filter_json`: swallowed ALL exceptions with zero
71
+ logging (returned unfiltered content). Real failures now log at
72
+ warning; the benign not-JSON case stays quiet.
73
+ - `graph._load_transitions`: DB corruption looked identical to the benign
74
+ first-run case (both debug). Corruption now warns; first-run stays debug.
75
+ - `file_intelligence.detect_file_type`: sniff failure silently reclassified
76
+ files as "text" (worse extraction, no signal) — now warns. PDFs where
77
+ NO page yields text now warn once (scanned/image-only detection).
78
+ - `CheckpointManager`: new `persistence_broken` flag — previously the
79
+ constructor swallowed a triple init failure and reported healthy while
80
+ every subsequent save was doomed. `get_latest()` now raises
81
+ `StorageError` on DB read failure instead of returning None ("no
82
+ checkpoint found, 404" vs "your checkpoints exist but are unreadable"
83
+ are different problems; callers can finally tell them apart).
84
+ - **`_db_connect` leaked the connection when the WAL PRAGMA failed on a
85
+ corrupt DB file** (both graph.py and manager.py) — on Windows the open
86
+ handle blocked the documented delete-and-recreate recovery path
87
+ (WinError 32). Found because the chaos test failed for the RIGHT reason
88
+ once get_latest stopped swallowing errors.
89
+
90
+ ### Security — redaction gaps
91
+ - URL-embedded credentials (`postgres://admin:pass@host/db`) matched NO
92
+ pattern — no `password=` literal, no recognized prefix. Now redacted
93
+ (credential part only; host survives for readability).
94
+ - Added OpenRouter / Hugging Face / xAI / Together key patterns.
95
+ - Checkpoint `next_action` (a raw 200-char message slice persisted to
96
+ SQLite) is now independently redacted — defense-in-depth so the
97
+ single-point-of-application assumption in chat_completions() is not the
98
+ only thing between a pasted key and the checkpoint DB.
99
+
100
+ ### Visualization — redesigned (was a generic node soup)
101
+ - The old shareable HTML exported `transitions` (the supersession
102
+ history — the one thing this product tracks that a generic graph view
103
+ doesn't) and then never rendered them. It also loaded D3 from a CDN, so
104
+ the "self-contained" artifact broke offline.
105
+ - New artifact (zero external deps, hand-rolled force layout):
106
+ supersession arcs (dashed red, old→new, arrowheads), a clickable
107
+ "Decision history" timeline panel (struck-through old label → new label
108
+ with trigger/reason/timestamp; click = spotlight both nodes + center),
109
+ glow rings on active decisions vs dashed rings + strikethrough on
110
+ superseded/archived, red rings on invalidated, per-type filter chips,
111
+ "Active only" toggle, text search, wheel-zoom/pan, one-click PNG export.
112
+ - Verified in a real browser (not just unit-tested): filters, search,
113
+ timeline spotlight, and layout all exercised via DOM inspection.
114
+
115
+ ### CLI
116
+ - `tokenmizer --help` crashed on Windows (cp1252) with UnicodeEncodeError
117
+ from the 🧠 emoji — found by dry-running the README quick start. Same
118
+ UTF-8 reconfigure fix the benchmark runners got in v0.2.4; the CLI had
119
+ been missed.
120
+
121
+ ### Round 2 — extraction-quality residuals (same audit, follow-up pass)
122
+ - **Near-duplicate decision nodes merged instead of self-superseding:**
123
+ one message ("Decided: use React for the frontend.") could emit two
124
+ decision variants ("Use React" + "use React for the frontend.") via
125
+ different regex passes; they became two nodes and one superseded the
126
+ other — a bogus "Changed:" line in every resume. `_is_same_decision`
127
+ now recognizes containment (smaller label's words ⊆ larger's, min 2
128
+ words so "Use PostgreSQL" is NOT collapsed into "Switch from PostgreSQL
129
+ to SQLite"), and `add_node` fuzzy-merges same-decision variants: keeps
130
+ the existing node, upgrades to the longer label, backfills the summary,
131
+ never resurrects SUPERSEDED status. Verified end-to-end: the demo
132
+ scenario now produces exactly one transition (React→Next.js) and a
133
+ clean resume block.
134
+ - **Validator now honors extractor corroboration confidence:** it used to
135
+ recompute confidence purely from label length/wording, so a doubly-
136
+ corroborated short decision (0.95) could be rejected while a verbose
137
+ weakly-sourced one passed. `validate()` gains `extractor_confidence`;
138
+ final = max(heuristic, (heuristic+extractor)/2) — monotone (evidence
139
+ only raises), not an override (heuristic-only 0.65 still fails a 0.65
140
+ threshold), and hard rejects remain absolute (0.95 cannot resurrect
141
+ junk). Wired from `add_node` via the existing confidence≠0.7 sentinel.
142
+ - **`HybridExtractor.min_confidence` is no longer dead code:** extract()
143
+ now filters merged output by merge()'s confidence tiers — default 0.55
144
+ keeps every tier (behavior unchanged); 0.7 drops heuristic-only items;
145
+ 0.9 keeps only corroborated ones. Decisions filter per-item, simple
146
+ lists per-category.
147
+ - Test fixture fix: `test_prune_preserves_decisions` used ten decisions
148
+ differing only by a trailing digit — 83% word-overlap, which
149
+ `_is_same_decision` always considered the same decision; the merge fix
150
+ made that judgment consequential, collapsing the fixture. Replaced with
151
+ ten genuinely distinct decisions across different topic buckets.
152
+
153
+ ### Tests
154
+ - 220 → 275 (55 new: MCP server 18, classifier+dedup 18, archival 3,
155
+ invalidate-scope 2, redaction 5, LLM-pass regression 1, validator
156
+ blending 4, min_confidence filter 4).
157
+
3
158
  ## [0.3.1] — 2026-07-03 — shareable graph visualization
4
159
 
5
160
  - NEW `GET /api/graph/{session_id}/html` — self-contained dark interactive
@@ -111,7 +266,9 @@
111
266
  removed; benchmark table updated to freshly measured v0.2.4 numbers with
112
267
  date and sample-size caveat.
113
268
 
114
- ## [Unreleased] — tokenizer, cache, and version consistency fixes
269
+ ## Shipped in [0.2.4] — tokenizer, cache, and version consistency fixes
270
+
271
+ (Header fixed 2026-07-10: this section was left titled "[Unreleased]" after it shipped.)
115
272
 
116
273
  ### Correctness — core value proposition
117
274
  - **`core/tokenizer.py`:** Claude/Anthropic models were counted with tiktoken
@@ -175,7 +332,9 @@
175
332
  Quick Start code example (previously only mentioned deep in the CLI
176
333
  section) — Cursor/Continue.dev users were hitting an unexplained HTTP 501.
177
334
 
178
- ## [Unreleased] — security/correctness audit pass
335
+ ## Shipped in [0.2.4] — security/correctness audit pass
336
+
337
+ (Header fixed 2026-07-10: this section was left titled "[Unreleased]" after it shipped.)
179
338
 
180
339
  Full senior-level audit covering security, silent failures, dead code,
181
340
  and benchmark honesty. See `TESTING.md` for what's verified vs. not, and
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: tokenmizer
3
- Version: 0.3.1
3
+ Version: 0.3.2
4
4
  Summary: Reduce AI context loss by 2x. Graph-backed checkpoint and resume for any LLM session.
5
5
  Project-URL: Homepage, https://github.com/Shweta-Mishra-ai/tokenmizer
6
6
  Project-URL: Repository, https://github.com/Shweta-Mishra-ai/tokenmizer
@@ -92,6 +92,7 @@ Description-Content-Type: text/markdown
92
92
  <a href="https://registry.modelcontextprotocol.io/v0/servers?search=tokenmizer"><img src="https://img.shields.io/badge/MCP%20Registry-published-5ee7c8?style=flat-square" alt="MCP Registry"/></a>
93
93
  <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-4ade80?style=flat-square"/></a>
94
94
  <a href="https://github.com/Shweta-Mishra-ai/tokenmizer/stargazers"><img src="https://img.shields.io/github/stars/Shweta-Mishra-ai/tokenmizer?style=flat-square&color=f9d84a" alt="Stars"/></a>
95
+ <a href="https://glama.ai/mcp/servers/Shweta-Mishra-ai/tokenmizer"><img src="https://glama.ai/mcp/servers/Shweta-Mishra-ai/tokenmizer/badges/score.svg" alt="Glama Score"/></a>
95
96
  </p>
96
97
 
97
98
  <p>
@@ -150,35 +151,122 @@ Your App → TokenMizer (:8000) → Claude / GPT / Gemini / any LLM
150
151
  | 🟢 `ACTIVE` | Current — in effect | ✅ Always |
151
152
  | 🟡 `SUPERSEDED` | Replaced by newer decision | ⚠️ 7 days |
152
153
  | 🔴 `INVALIDATED` | Explicitly wrong/cancelled | ⚠️ Always (warning) |
153
- | ⬜ `ARCHIVED` | Old but valid, not relevant | ❌ Never |
154
+ | ⬜ `ARCHIVED` | Superseded >7 days ago — aged out | ❌ Never |
154
155
 
155
156
  History is **never deleted**. "Why did we switch from React to Next.js?" — always answerable.
156
157
 
158
+ > Honesty note (2026-07-10 audit): before v0.3.2, `ARCHIVED` was advertised
159
+ > here but **unreachable** — the decay/prune/query logic all handled it, yet
160
+ > no code path ever set it. Superseded decisions now age into `ARCHIVED`
161
+ > automatically after 7 days (`GraphMemory.ARCHIVE_SUPERSEDED_AFTER_DAYS`),
162
+ > which is what makes the "⚠️ 7 days" row above actually true.
163
+
157
164
  ---
158
165
 
159
166
  ## Quick Start
160
167
 
168
+ <details>
169
+ <summary><b>🟢 Complete step-by-step setup (start here if you're new — 5 minutes, no code reading needed)</b></summary>
170
+
171
+ <br/>
172
+
173
+ **Step 0 — Check Python** (need 3.10 or newer)
174
+
175
+ Open a terminal (Windows: press Win, type "PowerShell", Enter · Mac: Cmd+Space, type "Terminal"):
176
+
177
+ ```
178
+ python --version
179
+ ```
180
+
181
+ You should see `Python 3.10` or higher. If not: install from [python.org/downloads](https://python.org/downloads) (Windows: tick **"Add Python to PATH"** during install).
182
+
183
+ **Step 1 — Install TokenMizer**
184
+
185
+ ```
186
+ pip install "tokenmizer[anthropic,cache]"
187
+ ```
188
+
189
+ ✅ You should see: `Successfully installed tokenmizer-...`
190
+
191
+ **Step 2 — Add your API key** (get one at [console.anthropic.com](https://console.anthropic.com) → API Keys)
192
+
193
+ Windows PowerShell:
194
+ ```powershell
195
+ setx TOKENMIZER_ANTHROPIC_API_KEY "sk-ant-YOUR-KEY"
196
+ ```
197
+ then **close and reopen** the terminal.
198
+
199
+ Mac/Linux:
200
+ ```bash
201
+ export TOKENMIZER_ANTHROPIC_API_KEY=sk-ant-YOUR-KEY
202
+ ```
203
+
204
+ *(No key? Use free local Ollama instead — see "No API key?" below.)*
205
+
206
+ **Step 3 — Start TokenMizer**
207
+
208
+ ```
209
+ tokenmizer serve
210
+ ```
211
+
212
+ ✅ You should see: `Proxy: http://localhost:8000/v1/chat/completions`
213
+ Leave this terminal open — TokenMizer runs here.
214
+
215
+ **Step 4 — Verify it's alive**
216
+
217
+ Open [http://localhost:8000](http://localhost:8000) in your browser → the TokenMizer dashboard appears. That's it — the proxy works.
218
+
219
+ **Step 5 — Connect your tool** (pick yours)
220
+
221
+ - **Cursor:** Settings → Models → OpenAI API → Base URL: `http://localhost:8000/v1`
222
+ - **Claude Desktop / Claude Code:** see [Claude Code Integration](#claude-code-integration) below (copy one JSON block, restart the app)
223
+ - **Your own Python code:** see "Use — change one line" below
224
+
225
+ **Something failed?** `pip` not found → reinstall Python with "Add to PATH". Port 8000 busy → `tokenmizer serve --port 8001`. Anything else → [open an issue](https://github.com/Shweta-Mishra-ai/tokenmizer/issues) with the error text — median response < 1 day.
226
+
227
+ </details>
228
+
161
229
  ### 1. Install
162
230
 
231
+ Works on **Windows, macOS, and Linux** (Python 3.10+). Same command everywhere:
232
+
163
233
  ```bash
164
234
  # Recommended
165
235
  pip install "tokenmizer[anthropic,cache]"
166
236
 
167
237
  # All providers
168
238
  pip install "tokenmizer[anthropic,openai,gemini,cohere,cache]"
239
+ ```
240
+
241
+ <details>
242
+ <summary><b>No API key? Use Ollama (free, local)</b></summary>
243
+
244
+ ```bash
245
+ # macOS: brew install ollama
246
+ # Windows: winget install Ollama.Ollama (or download from ollama.com)
247
+ # Linux: curl -fsSL https://ollama.com/install.sh | sh
169
248
 
170
- # No key? Use Ollama (free, local)
171
- brew install ollama && ollama pull llama3
249
+ ollama pull llama3
172
250
  pip install tokenmizer
251
+ # then set provider: ollama in tokenmizer.yaml
173
252
  ```
253
+ </details>
174
254
 
175
255
  ### 2. Set your API key
176
256
 
257
+ **macOS / Linux (bash, zsh):**
177
258
  ```bash
178
259
  export TOKENMIZER_ANTHROPIC_API_KEY=sk-ant-...
179
- # or: TOKENMIZER_OPENAI_API_KEY, TOKENMIZER_GEMINI_API_KEY, etc.
180
260
  ```
181
261
 
262
+ **Windows (PowerShell):**
263
+ ```powershell
264
+ $env:TOKENMIZER_ANTHROPIC_API_KEY = "sk-ant-..." # current session
265
+ setx TOKENMIZER_ANTHROPIC_API_KEY "sk-ant-..." # persistent (new terminals)
266
+ ```
267
+
268
+ Other providers: `TOKENMIZER_OPENAI_API_KEY`, `TOKENMIZER_GEMINI_API_KEY`, etc. — full table in [Supported Providers](#supported-providers).
269
+
182
270
  ### 3. Start
183
271
 
184
272
  ```bash
@@ -233,11 +321,11 @@ Then use skills directly:
233
321
  /tokenmizer:stats → token savings report
234
322
  ```
235
323
 
236
- ### Option B — MCP server
324
+ ### Option B — MCP server (Claude Desktop, Claude Code, Cursor, VS Code, Zed)
237
325
 
238
326
  mcp-name: io.github.Shweta-Mishra-ai/tokenmizer
239
327
 
240
- Add to `~/.claude/settings.json`:
328
+ Add this `mcpServers` block to your client's MCP config file:
241
329
 
242
330
  ```json
243
331
  {
@@ -250,7 +338,31 @@ Add to `~/.claude/settings.json`:
250
338
  }
251
339
  ```
252
340
 
253
- (`tokenmizer-mcp` is installed with the package; `python3 -m tokenmizer.mcp.server` also works.)
341
+ Where the config file lives:
342
+
343
+ | Client | Config file |
344
+ |---|---|
345
+ | **Claude Desktop** (Windows) | `%APPDATA%\Claude\claude_desktop_config.json` |
346
+ | **Claude Desktop** (macOS) | `~/Library/Application Support/Claude/claude_desktop_config.json` |
347
+ | **Claude Code** | `.mcp.json` in your project, or `~/.claude/settings.json` |
348
+ | **Cursor** | Settings → MCP → Add server (same JSON) |
349
+ | **VS Code / Zed** | their MCP settings — same `command` + `env` |
350
+ | **OpenAI Codex CLI** | `~/.codex/config.toml` — TOML format, see below |
351
+
352
+ <details>
353
+ <summary>Codex CLI config (TOML, not JSON)</summary>
354
+
355
+ ```toml
356
+ [mcp_servers.tokenmizer]
357
+ command = "tokenmizer-mcp"
358
+ env = { TOKENMIZER_URL = "http://localhost:8000" }
359
+ ```
360
+ </details>
361
+
362
+ Then restart the client. Keep `tokenmizer serve` running for the
363
+ checkpoint/resume/stats tools (file analysis works without it).
364
+ If `tokenmizer-mcp` isn't on your PATH, use `"command": "python"`,
365
+ `"args": ["-m", "tokenmizer.mcp.server"]` instead.
254
366
 
255
367
  ---
256
368
 
@@ -348,7 +460,13 @@ default_model: claude-sonnet-4-6
348
460
  graph_checkpoint:
349
461
  enabled: true
350
462
  trigger_at_percent: 0.85
351
- use_llm_extraction: false # true = 80%+ recall, needs key (~$0.001/turn)
463
+ use_llm_extraction: false # true = hybrid LLM+heuristic extraction
464
+ # (needs a provider key, ~$0.001/turn).
465
+ # NOTE: before v0.3.2 this flag silently did
466
+ # nothing — the call site passed provider_fn
467
+ # to the wrong function and raised TypeError
468
+ # on every turn, falling back to heuristics.
469
+ # Fixed + regression-tested; see CHANGELOG.
352
470
 
353
471
  compression:
354
472
  enabled: true
@@ -388,7 +506,7 @@ TOKENMIZER_API_KEY=strong-key docker-compose up
388
506
  | `/api/checkpoint` | POST | Manual checkpoint |
389
507
  | `/api/decision/invalidate` | POST | Mark decision as invalid |
390
508
  | `/api/graph/{id}` | GET | Session graph stats |
391
- | `/api/graph/{id}/html` | GET | **Interactive graph page** — open, drag, zoom, share |
509
+ | `/api/graph/{id}/html` | GET | **Interactive graph page** — decision-history timeline, supersession arcs, type/status filters, search, zoom/pan, PNG export. Zero external dependencies (works offline) |
392
510
  | `/api/stats` | GET | Token savings analytics |
393
511
  | `/health` | GET | Health check |
394
512
  | `/docs` | GET | Swagger UI |
@@ -401,7 +519,13 @@ TOKENMIZER_API_KEY=strong-key docker-compose up
401
519
  - Secret/PII redaction applied once at ingestion, before graph storage,
402
520
  checkpoint storage, AND every LLM call (main chat *and* the background
403
521
  extraction model — these are separate, the redaction gap between them
404
- was a real bug, now fixed)
522
+ was a real bug, now fixed). Patterns cover Anthropic/OpenAI/Google/
523
+ GitHub/AWS/Slack/Stripe/JWT/OpenRouter/HF/xAI keys, URL-embedded
524
+ credentials (`postgres://user:pass@host` — a 2026-07 audit gap, fixed),
525
+ and generic `key=`/`password=` assignments. Best-effort by nature —
526
+ an unrecognized format with no keyword context can still slip through.
527
+ The checkpoint layer independently re-redacts what it persists
528
+ (defense-in-depth, added in the same audit).
405
529
  - Session-isolated cache (sensitive data never shared across sessions)
406
530
  - Basic prompt-injection keyword filter — catches copy-pasted jailbreak
407
531
  templates only; **not** a security boundary against a motivated
@@ -535,7 +659,7 @@ Contributions welcome — this project merges fast (median PR review < 1 day).
535
659
  git clone https://github.com/Shweta-Mishra-ai/tokenmizer
536
660
  cd tokenmizer
537
661
  pip install -e ".[dev]"
538
- pytest tests/ -v && ruff check tokenmizer/ # 218 tests, must stay green
662
+ pytest tests/ -v && ruff check tokenmizer/ # 262 tests, must stay green
539
663
  python scripts/mcp_e2e_check.py # full-pipeline e2e check
540
664
  ```
541
665
 
@@ -568,5 +692,5 @@ MIT © [Shweta Mishra](https://github.com/Shweta-Mishra-ai)
568
692
  <div align="center">
569
693
  <sub>Built for developers who spend too much time re-explaining their projects to AI.</sub>
570
694
  <br/><br/>
571
- <a href="https://github.com/Shweta-Mishra-ai/tokenmizer"><img src="https://img.shields.io/github/stars/Shweta-Mishra-ai/tokenmizer?style=social" alt="GitHub stars"/></a>
695
+ <a href="https://github.com/Shweta-Mishra-ai/tokenmizer/stargazers"><img src="https://img.shields.io/github/stars/Shweta-Mishra-ai/tokenmizer?style=flat-square&color=f9d84a&label=%E2%AD%90%20Star%20on%20GitHub" alt="GitHub stars"/></a>
572
696
  </div>