turnsafe 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- turnsafe-0.1.0/.gitignore +12 -0
- turnsafe-0.1.0/AGENT.md +99 -0
- turnsafe-0.1.0/CHANGELOG.md +39 -0
- turnsafe-0.1.0/CONTRIBUTING.md +54 -0
- turnsafe-0.1.0/LICENSE +21 -0
- turnsafe-0.1.0/PKG-INFO +361 -0
- turnsafe-0.1.0/README.md +328 -0
- turnsafe-0.1.0/SECURITY.md +21 -0
- turnsafe-0.1.0/docs/research.md +164 -0
- turnsafe-0.1.0/docs/rules.md +138 -0
- turnsafe-0.1.0/examples/anthropic_agent.py +50 -0
- turnsafe-0.1.0/examples/offline_demo.py +82 -0
- turnsafe-0.1.0/examples/openai_agent.py +57 -0
- turnsafe-0.1.0/pyproject.toml +74 -0
- turnsafe-0.1.0/src/turnsafe/__init__.py +66 -0
- turnsafe-0.1.0/src/turnsafe/_match.py +126 -0
- turnsafe-0.1.0/src/turnsafe/_providers.py +351 -0
- turnsafe-0.1.0/src/turnsafe/_thinking.py +99 -0
- turnsafe-0.1.0/src/turnsafe/_types.py +119 -0
- turnsafe-0.1.0/src/turnsafe/fit.py +169 -0
- turnsafe-0.1.0/src/turnsafe/py.typed +0 -0
- turnsafe-0.1.0/src/turnsafe/repair.py +381 -0
- turnsafe-0.1.0/src/turnsafe/tokens.py +51 -0
- turnsafe-0.1.0/src/turnsafe/validate.py +196 -0
- turnsafe-0.1.0/tests/__init__.py +0 -0
- turnsafe-0.1.0/tests/builders.py +57 -0
- turnsafe-0.1.0/tests/test_fit.py +174 -0
- turnsafe-0.1.0/tests/test_live.py +93 -0
- turnsafe-0.1.0/tests/test_properties.py +240 -0
- turnsafe-0.1.0/tests/test_regressions.py +132 -0
- turnsafe-0.1.0/tests/test_repair.py +219 -0
- turnsafe-0.1.0/tests/test_validate.py +223 -0
turnsafe-0.1.0/AGENT.md
ADDED
|
@@ -0,0 +1,99 @@
|
|
|
1
|
+
# AGENT.md
|
|
2
|
+
|
|
3
|
+
A guide for AI coding agents and human contributors working on `turnsafe`.
|
|
4
|
+
|
|
5
|
+
## What this project is
|
|
6
|
+
|
|
7
|
+
`turnsafe` is a zero-dependency Python library that validates, repairs and trims LLM chat
|
|
8
|
+
histories. Its purpose is to make sure a tool call is never separated from its result. It works on
|
|
9
|
+
plain provider message dicts: OpenAI Chat Completions and Anthropic Messages in v0.1.
|
|
10
|
+
|
|
11
|
+
Why it exists, and what was rejected in its favour, is in [docs/research.md](docs/research.md).
|
|
12
|
+
Every provider rule is listed with its source in [docs/rules.md](docs/rules.md).
|
|
13
|
+
|
|
14
|
+
## Commands
|
|
15
|
+
|
|
16
|
+
```bash
|
|
17
|
+
uv venv && uv pip install -e '.[dev]' # set up
|
|
18
|
+
pytest -q # all offline tests, including property tests (~2 s)
|
|
19
|
+
mypy # strict, src only
|
|
20
|
+
ruff check src tests && ruff format src tests
|
|
21
|
+
TURNSAFE_LIVE=1 pytest tests/test_live.py # real API contract tests; needs API keys
|
|
22
|
+
python -m build && twine check dist/* # build and check the release artefacts
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
A change is done only when `pytest`, `mypy` and `ruff check` all pass.
|
|
26
|
+
|
|
27
|
+
## Layout
|
|
28
|
+
|
|
29
|
+
| File | Responsibility |
|
|
30
|
+
|---|---|
|
|
31
|
+
| `src/turnsafe/_providers.py` | **All** format-specific knowledge. It has one `Adapter` per provider: where calls and results live, the allowed "zone" for results, what starts a turn, and how to build a placeholder. It also contains `normalize()` (deep copy, SDK objects to dicts) and `detect()`. |
|
|
32
|
+
| `src/turnsafe/_match.py` | Pairs every call with its result. `validate` and `repair` both use it, so they cannot disagree. |
|
|
33
|
+
| `src/turnsafe/validate.py` | Turns a `Match` plus provider-only rules into `Violation`s. |
|
|
34
|
+
| `src/turnsafe/repair.py` | Rebuilds the history from a `Match`. |
|
|
35
|
+
| `src/turnsafe/fit.py` | Splits the history into turns and drops whole turns, oldest first. |
|
|
36
|
+
| `src/turnsafe/_thinking.py` | Opt-in removal of Anthropic thinking blocks after the first change (prefix binding). |
|
|
37
|
+
| `src/turnsafe/tokens.py` | Offline token estimate. |
|
|
38
|
+
| `src/turnsafe/_types.py` | Public dataclasses, violation codes and exceptions. |
|
|
39
|
+
| `tests/builders.py` | Short constructors (`ou`, `oa`, `ot`, `au`, `aa`, `use`, `res`, `thinking`). Use them in tests. |
|
|
40
|
+
| `tests/test_properties.py` | Hypothesis tests of the invariants below. |
|
|
41
|
+
|
|
42
|
+
## Invariants: never break these
|
|
43
|
+
|
|
44
|
+
The property tests check all of them. If a change needs one relaxed, stop and ask a
|
|
45
|
+
maintainer first.
|
|
46
|
+
|
|
47
|
+
1. **Pure.** No public function mutates its input. `normalize()` deep-copies first.
|
|
48
|
+
2. **Repaired output validates.** `validate(repair(x))` has no errors except
|
|
49
|
+
`system_in_messages` on the first message, which only the caller can move.
|
|
50
|
+
3. **Fixed point.** If `validate(x)` has no violations, then `repair(x) == x` and there are no fixes.
|
|
51
|
+
4. **Idempotent.** `repair(repair(x)) == repair(x)`.
|
|
52
|
+
5. **Assistant messages are sacred.** With the default `thinking="keep"`, `repair()` never
|
|
53
|
+
edits an assistant message, except to remove a repeated tool call id or under
|
|
54
|
+
`on_missing="strip"`. With `thinking="strip_after_change"`, the only extra edit is removing
|
|
55
|
+
thinking blocks that come after the first changed block, because on models with prefix
|
|
56
|
+
binding those blocks are invalid. Read the thinking section of `docs/rules.md` before
|
|
57
|
+
touching this.
|
|
58
|
+
6. **Fitted output validates, and is a suffix.** `fit()` output is valid, and apart from pinned
|
|
59
|
+
system messages and an optional pinned first turn, it is a contiguous tail of the input
|
|
60
|
+
(with `thinking="strip_after_change"`, thinking blocks are removed once anything was dropped).
|
|
61
|
+
It never splits a turn.
|
|
62
|
+
7. **Unknown content passes through.** Blocks and keys that turnsafe doesn't understand are
|
|
63
|
+
copied unchanged, never dropped.
|
|
64
|
+
8. **No runtime dependencies.** Only the standard library. SDKs may be used in tests only.
|
|
65
|
+
9. **Linear time.** Matching, validation, repair and fit are linear in the history size.
|
|
66
|
+
`tests/test_regressions.py` has timing guards for the pathological cases (many orphans,
|
|
67
|
+
reused ids, thousands of parallel calls). Don't add per-turn scans over the whole history.
|
|
68
|
+
|
|
69
|
+
## Rules for changing provider behaviour
|
|
70
|
+
|
|
71
|
+
- **Every rule needs a source.** Add a row to `docs/rules.md` with the official doc URL and, where
|
|
72
|
+
possible, the exact API error string. Don't encode a rule you can't source. If a rule is only
|
|
73
|
+
sometimes enforced, make it a `warning`.
|
|
74
|
+
- **Put the rule in the right place.** Format knowledge goes in an `Adapter` method. A
|
|
75
|
+
provider-only check that isn't about pairing goes in `validate.py` (see
|
|
76
|
+
`_anthropic_rules`). Don't add `if provider == ...` branches to `_match.py` or `fit.py`.
|
|
77
|
+
- **Add a live case.** When you add or change a rule, add a case to `tests/test_live.py` that
|
|
78
|
+
the real API rejects before repair and accepts after.
|
|
79
|
+
- **Adding a provider** (Responses, Gemini) means writing a new `Adapter` subclass, adding it
|
|
80
|
+
to `ADAPTERS`, adding detection markers, adding a strategy in `test_properties.py`, and adding
|
|
81
|
+
a section in `docs/rules.md`.
|
|
82
|
+
|
|
83
|
+
## Style
|
|
84
|
+
|
|
85
|
+
- Match the existing code: small functions, type hints everywhere (`mypy --strict`), docstrings
|
|
86
|
+
on public functions, and comments only where the reason isn't obvious.
|
|
87
|
+
- Public API changes must update `__all__`, `README.md` and `CHANGELOG.md`.
|
|
88
|
+
- Violation codes and `Fix.action` strings are public API. Don't rename them outside a
|
|
89
|
+
major release.
|
|
90
|
+
- Error and violation messages say what is wrong and where. `InvalidHistoryError` points the
|
|
91
|
+
user to `repair()`.
|
|
92
|
+
|
|
93
|
+
## Releasing
|
|
94
|
+
|
|
95
|
+
1. Bump `version` in `pyproject.toml` and `__version__` in `src/turnsafe/__init__.py`.
|
|
96
|
+
2. Move the "Unreleased" entries in `CHANGELOG.md` under the new version.
|
|
97
|
+
3. Tag `vX.Y.Z` and push the tag. `.github/workflows/release.yml` builds the package and
|
|
98
|
+
publishes it to PyPI through trusted publishing. Never publish from a laptop with a
|
|
99
|
+
long-lived token.
|
|
@@ -0,0 +1,39 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to this project are documented here. The format follows
|
|
4
|
+
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and the project uses
|
|
5
|
+
[Semantic Versioning](https://semver.org/).
|
|
6
|
+
|
|
7
|
+
## [0.1.0] - 2026-09-28
|
|
8
|
+
|
|
9
|
+
### Added
|
|
10
|
+
|
|
11
|
+
- `validate()`: checks an OpenAI Chat Completions or Anthropic Messages history and returns a
|
|
12
|
+
`Report` of typed `Violation`s. The report says what the API would reject, where, and for
|
|
13
|
+
which tool call.
|
|
14
|
+
- `repair()` and `repair_detailed()`: rebuild a valid history. They move misplaced results,
|
|
15
|
+
drop orphaned and duplicate results, and add placeholder results for unanswered calls (or
|
|
16
|
+
remove those calls when `on_missing="strip"`). They also fix the Anthropic block-order and
|
|
17
|
+
empty-content rules.
|
|
18
|
+
- `fit()` and `fit_detailed()`: trim a history to a token budget by dropping whole turns, so a
|
|
19
|
+
tool call is never split from its results. Options: `keep_system`, `keep_first_turn` and a
|
|
20
|
+
custom `counter`.
|
|
21
|
+
- `thinking=` on `repair()` and `fit()`. The default `"keep"` never touches Anthropic thinking
|
|
22
|
+
blocks. `"strip_after_change"` removes them from the first changed block onward, for Claude
|
|
23
|
+
models with thinking-block prefix binding (Fable 5.1, Opus 5.5).
|
|
24
|
+
- Anthropic rules verified against the docs on 2026-09-28:
|
|
25
|
+
- Consecutive user messages, and consecutive assistant messages, are each treated as one turn.
|
|
26
|
+
- `tool_use` ids must be unique across the history.
|
|
27
|
+
- Blank `tool_result` content and empty assistant content are accepted.
|
|
28
|
+
- A mid-conversation `system` message is a warning; it's an error only as the first message.
|
|
29
|
+
- A result that ended up after an unrelated assistant message is moved back to its call instead
|
|
30
|
+
of being replaced by a placeholder. The search stops at the next call that reuses the same id.
|
|
31
|
+
- Tool call and result ids must be non-empty strings. Anything else raises a `ValueError` that
|
|
32
|
+
names the position, instead of silently escaping the rules.
|
|
33
|
+
- Validation, repair and trimming are linear in the size of the history, including for
|
|
34
|
+
histories with many orphaned results, reused ids or thousands of parallel calls.
|
|
35
|
+
- `approx_tokens()`: an offline token estimate that counts images as images.
|
|
36
|
+
- `detect_provider()` and `provider="auto"`.
|
|
37
|
+
- Accepts SDK message objects that have `model_dump()`.
|
|
38
|
+
- Property-based tests for the invariants listed in `AGENT.md`, and opt-in live contract tests
|
|
39
|
+
against both APIs.
|
|
@@ -0,0 +1,54 @@
|
|
|
1
|
+
# Contributing to turnsafe
|
|
2
|
+
|
|
3
|
+
Thanks for helping. This guide covers setting up, what a good change looks like, and how to
|
|
4
|
+
report a provider rule turnsafe gets wrong.
|
|
5
|
+
|
|
6
|
+
## Set up
|
|
7
|
+
|
|
8
|
+
```bash
|
|
9
|
+
git clone https://github.com/Yasir-Khan-7/turnsafe && cd turnsafe
|
|
10
|
+
uv venv && uv pip install -e '.[dev]' # or: python -m venv .venv && pip install -e '.[dev]'
|
|
11
|
+
pytest -q # fully offline, about 3 seconds
|
|
12
|
+
```
|
|
13
|
+
|
|
14
|
+
## Before you open a pull request
|
|
15
|
+
|
|
16
|
+
```bash
|
|
17
|
+
pytest -q
|
|
18
|
+
mypy
|
|
19
|
+
ruff check src tests && ruff format src tests
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
CI runs all of these on Python 3.10 to 3.14.
|
|
23
|
+
|
|
24
|
+
## What a good change looks like
|
|
25
|
+
|
|
26
|
+
- **Tests.** Every behaviour change comes with a test. If it touches `repair()` or `fit()`, also
|
|
27
|
+
check that the property tests in `tests/test_properties.py` still hold, or extend their
|
|
28
|
+
generators.
|
|
29
|
+
- **Cited rules.** Every provider rule needs a source. Add it to `docs/rules.md` with the
|
|
30
|
+
official docs link, and where possible the exact API error text. If you can't source a rule,
|
|
31
|
+
make it a warning, not an error.
|
|
32
|
+
- **Format code lives in one place.** Format-specific code goes in an `Adapter` in
|
|
33
|
+
`src/turnsafe/_providers.py`, not in `if provider == ...` branches elsewhere.
|
|
34
|
+
- **Stable names.** Public names, violation codes and `Fix.action` strings are stable API. Note
|
|
35
|
+
any change to them in `CHANGELOG.md`.
|
|
36
|
+
|
|
37
|
+
The invariants the code must keep are listed in [AGENT.md](AGENT.md#invariants-never-break-these).
|
|
38
|
+
|
|
39
|
+
## Reporting a wrong rule
|
|
40
|
+
|
|
41
|
+
The most useful issue you can open includes:
|
|
42
|
+
1. The exact API error text.
|
|
43
|
+
2. The model and platform: first-party API, Bedrock, Vertex or a gateway.
|
|
44
|
+
3. The smallest `messages` list that reproduces it. Redact the content, since only the
|
|
45
|
+
structure matters.
|
|
46
|
+
|
|
47
|
+
## Live tests
|
|
48
|
+
|
|
49
|
+
`tests/test_live.py` sends real requests to check that repaired histories are accepted. It
|
|
50
|
+
costs a few tokens and is skipped by default:
|
|
51
|
+
|
|
52
|
+
```bash
|
|
53
|
+
TURNSAFE_LIVE=1 OPENAI_API_KEY=... ANTHROPIC_API_KEY=... pytest tests/test_live.py
|
|
54
|
+
```
|
turnsafe-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Yasir Khan
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
turnsafe-0.1.0/PKG-INFO
ADDED
|
@@ -0,0 +1,361 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: turnsafe
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Validate, repair and trim LLM chat history without orphaning tool calls. Fixes 'tool_use ids were found without tool_result blocks' and 'messages with role tool must be a response to a preceeding message with tool_calls'.
|
|
5
|
+
Project-URL: Homepage, https://github.com/Yasir-Khan-7/turnsafe
|
|
6
|
+
Project-URL: Documentation, https://github.com/Yasir-Khan-7/turnsafe#readme
|
|
7
|
+
Project-URL: Issues, https://github.com/Yasir-Khan-7/turnsafe/issues
|
|
8
|
+
Project-URL: Changelog, https://github.com/Yasir-Khan-7/turnsafe/blob/main/CHANGELOG.md
|
|
9
|
+
Author: Yasir Khan
|
|
10
|
+
License-Expression: MIT
|
|
11
|
+
License-File: LICENSE
|
|
12
|
+
Keywords: agents,anthropic,chat-history,claude,context-window,function-calling,llm,openai,tool-calling,tool_result,tool_use,trim-messages
|
|
13
|
+
Classifier: Development Status :: 4 - Beta
|
|
14
|
+
Classifier: Intended Audience :: Developers
|
|
15
|
+
Classifier: Operating System :: OS Independent
|
|
16
|
+
Classifier: Programming Language :: Python :: 3
|
|
17
|
+
Classifier: Programming Language :: Python :: 3 :: Only
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
20
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
21
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
22
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
23
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
24
|
+
Classifier: Topic :: Software Development :: Libraries :: Python Modules
|
|
25
|
+
Classifier: Typing :: Typed
|
|
26
|
+
Requires-Python: >=3.10
|
|
27
|
+
Provides-Extra: dev
|
|
28
|
+
Requires-Dist: hypothesis>=6.100; extra == 'dev'
|
|
29
|
+
Requires-Dist: mypy>=1.10; extra == 'dev'
|
|
30
|
+
Requires-Dist: pytest>=8; extra == 'dev'
|
|
31
|
+
Requires-Dist: ruff>=0.6; extra == 'dev'
|
|
32
|
+
Description-Content-Type: text/markdown
|
|
33
|
+
|
|
34
|
+
<div align="center">
|
|
35
|
+
|
|
36
|
+
# turnsafe
|
|
37
|
+
|
|
38
|
+
**Never send a broken chat history to an LLM again.**
|
|
39
|
+
|
|
40
|
+
Validate, repair and trim OpenAI and Anthropic message lists, without ever separating a tool
|
|
41
|
+
call from its result.
|
|
42
|
+
|
|
43
|
+
[](https://github.com/Yasir-Khan-7/turnsafe/actions/workflows/ci.yml)
|
|
44
|
+
[](https://pypi.org/project/turnsafe/)
|
|
45
|
+
[](https://pypi.org/project/turnsafe/)
|
|
46
|
+
[](https://mypy.readthedocs.io/)
|
|
47
|
+
[](https://github.com/Yasir-Khan-7/turnsafe/blob/main/pyproject.toml)
|
|
48
|
+
[](https://github.com/Yasir-Khan-7/turnsafe/blob/main/LICENSE)
|
|
49
|
+
|
|
50
|
+
</div>
|
|
51
|
+
|
|
52
|
+
---
|
|
53
|
+
|
|
54
|
+
## Did one of these errors bring you here?
|
|
55
|
+
|
|
56
|
+
```text
|
|
57
|
+
anthropic.BadRequestError: messages.4: `tool_use` ids were found without `tool_result` blocks
|
|
58
|
+
immediately after: toolu_01A... Each `tool_use` block must have a corresponding `tool_result`
|
|
59
|
+
block in the next message.
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
```text
|
|
63
|
+
openai.BadRequestError: Invalid parameter: messages with role 'tool' must be a response to a
|
|
64
|
+
preceeding message with 'tool_calls'.
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
```text
|
|
68
|
+
openai.BadRequestError: An assistant message with 'tool_calls' must be followed by tool
|
|
69
|
+
messages responding to each 'tool_call_id'. The following tool_call_ids did not have
|
|
70
|
+
response messages: call_abc123
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
Each one means **a tool call got separated from its result**. Common causes:
|
|
74
|
+
|
|
75
|
+
- You trimmed the history to fit the context window (`messages[-20:]`) and the cut landed in
|
|
76
|
+
the middle of a tool call.
|
|
77
|
+
- You saved the conversation and reloaded it, and something got lost or reordered.
|
|
78
|
+
- The agent crashed or was cancelled while a tool was running.
|
|
79
|
+
- The user sent a message while tools were still running.
|
|
80
|
+
|
|
81
|
+
After that, **every later request in the conversation fails** until someone fixes the history.
|
|
82
|
+
|
|
83
|
+
`turnsafe` fixes the history before you send it. It works on the plain `messages` list you
|
|
84
|
+
already pass to the OpenAI and Anthropic SDKs. It has no dependencies and needs no framework or
|
|
85
|
+
network access.
|
|
86
|
+
|
|
87
|
+
## Contents
|
|
88
|
+
|
|
89
|
+
- [Install](#install)
|
|
90
|
+
- [Quickstart](#quickstart)
|
|
91
|
+
- [Try it without an API key](#try-it-without-an-api-key)
|
|
92
|
+
- [How it works](#how-it-works)
|
|
93
|
+
- [API](#api)
|
|
94
|
+
- [Anthropic thinking blocks](#anthropic-thinking-blocks)
|
|
95
|
+
- [How it compares](#how-it-compares)
|
|
96
|
+
- [FAQ](#faq)
|
|
97
|
+
- [Status and roadmap](#status-and-roadmap)
|
|
98
|
+
- [Contributing](#contributing)
|
|
99
|
+
|
|
100
|
+
## Install
|
|
101
|
+
|
|
102
|
+
```bash
|
|
103
|
+
pip install turnsafe
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
It needs Python 3.10 or newer and has zero runtime dependencies.
|
|
107
|
+
|
|
108
|
+
## Quickstart
|
|
109
|
+
|
|
110
|
+
Add one line before each API call in your agent loop:
|
|
111
|
+
|
|
112
|
+
```python
|
|
113
|
+
import anthropic
|
|
114
|
+
import turnsafe
|
|
115
|
+
|
|
116
|
+
client = anthropic.Anthropic()
|
|
117
|
+
|
|
118
|
+
while True:
|
|
119
|
+
request = turnsafe.fit(turnsafe.repair(messages), max_tokens=150_000)
|
|
120
|
+
response = client.messages.create(
|
|
121
|
+
model="claude-sonnet-5", max_tokens=1024, tools=tools, messages=request
|
|
122
|
+
)
|
|
123
|
+
...
|
|
124
|
+
```
|
|
125
|
+
|
|
126
|
+
- `repair()` heals a broken history.
|
|
127
|
+
- `fit()` trims whole turns so the history fits your budget.
|
|
128
|
+
|
|
129
|
+
They work the same way with `openai.OpenAI().chat.completions.create(...)`. The provider is
|
|
130
|
+
detected from the message format, or you can pass `provider="openai"` or `"anthropic"`.
|
|
131
|
+
|
|
132
|
+
For complete runnable agents that save and reload their history, see
|
|
133
|
+
[`examples/anthropic_agent.py`](https://github.com/Yasir-Khan-7/turnsafe/blob/main/examples/anthropic_agent.py)
|
|
134
|
+
and [`examples/openai_agent.py`](https://github.com/Yasir-Khan-7/turnsafe/blob/main/examples/openai_agent.py).
|
|
135
|
+
|
|
136
|
+
## Try it without an API key
|
|
137
|
+
|
|
138
|
+
```bash
|
|
139
|
+
git clone https://github.com/Yasir-Khan-7/turnsafe && cd turnsafe
|
|
140
|
+
pip install -e . && python examples/offline_demo.py
|
|
141
|
+
```
|
|
142
|
+
|
|
143
|
+
```text
|
|
144
|
+
=== Anthropic: agent crashed between two parallel tool calls ===
|
|
145
|
+
anthropic history has 1 error(s), 0 warning(s):
|
|
146
|
+
[error] messages[1] unanswered_tool_call: tool call 'toolu_2' has no tool result in the next user message
|
|
147
|
+
fix: messages[1] added_placeholder: no result for 'toolu_2'
|
|
148
|
+
after repair: anthropic history is valid.
|
|
149
|
+
|
|
150
|
+
=== OpenAI: a user message landed between a call and its result ===
|
|
151
|
+
openai history has 1 error(s), 0 warning(s):
|
|
152
|
+
[error] messages[3] misplaced_tool_result: result for tool call 'call_7' must be in the tool messages directly after the call
|
|
153
|
+
fix: messages[1] moved_result: moved result for 'call_7' from messages[3]
|
|
154
|
+
after repair: openai history is valid.
|
|
155
|
+
|
|
156
|
+
=== fit() to 1,000 tokens ===
|
|
157
|
+
kept 24 of 80 messages (~977 tokens), dropped 14 whole turns; valid: True
|
|
158
|
+
```
|
|
159
|
+
|
|
160
|
+
## How it works
|
|
161
|
+
|
|
162
|
+
```text
|
|
163
|
+
user ─┐
|
|
164
|
+
│ turn 1 fit() drops whole turns, oldest first.
|
|
165
|
+
asst │ (tool call) ─┐ A tool call and its results are always in the
|
|
166
|
+
tool │ (result) ◄──┘ same turn, so they leave together or stay together.
|
|
167
|
+
asst ─┘
|
|
168
|
+
user ─┐
|
|
169
|
+
│ turn 2 repair() pairs every call with its result and fixes
|
|
170
|
+
asst │ (tool call) ─┐ anything the API would reject: missing results,
|
|
171
|
+
user │ "hurry!" │ orphaned results, misplaced results, duplicates.
|
|
172
|
+
tool │ (result) ◄──┘ ← misplaced: repair() moves it back next to its call
|
|
173
|
+
asst ─┘
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
Every rule turnsafe checks comes from the provider's documentation, or from the exact API error
|
|
177
|
+
text when the docs are silent. All sources are listed in
|
|
178
|
+
[`docs/rules.md`](https://github.com/Yasir-Khan-7/turnsafe/blob/main/docs/rules.md).
|
|
179
|
+
|
|
180
|
+
The core promises are checked by property-based tests that run thousands of randomly damaged
|
|
181
|
+
histories on every change:
|
|
182
|
+
|
|
183
|
+
- `repair()` output always passes `validate()`.
|
|
184
|
+
- A valid history comes back unchanged.
|
|
185
|
+
- `repair()` never mutates its input.
|
|
186
|
+
- `fit()` never splits a tool call from its result.
|
|
187
|
+
|
|
188
|
+
## API
|
|
189
|
+
|
|
190
|
+
### `validate(messages, provider="auto") -> Report`
|
|
191
|
+
|
|
192
|
+
`validate()` tells you what the API would reject, without changing anything.
|
|
193
|
+
|
|
194
|
+
```python
|
|
195
|
+
report = turnsafe.validate(messages)
|
|
196
|
+
if not report: # falsy when there are errors
|
|
197
|
+
print(report.explain())
|
|
198
|
+
report.raise_for_errors() # or fail fast, for example in tests or CI
|
|
199
|
+
```
|
|
200
|
+
|
|
201
|
+
Each `Violation` has:
|
|
202
|
+
- a stable `code`, such as `"orphan_tool_result"`
|
|
203
|
+
- the message `index`
|
|
204
|
+
- the `tool_call_id`
|
|
205
|
+
- a `severity`, either `"error"` or `"warning"`
|
|
206
|
+
|
|
207
|
+
### `repair(messages, provider="auto", *, on_missing="placeholder", placeholder=..., thinking="keep")`
|
|
208
|
+
|
|
209
|
+
`repair()` returns a fixed copy of the history. Use `repair_detailed()` to also get a `Fix` for
|
|
210
|
+
every change, which is useful for logging.
|
|
211
|
+
|
|
212
|
+
```python
|
|
213
|
+
messages, fixes = turnsafe.repair_detailed(messages)
|
|
214
|
+
for fix in fixes:
|
|
215
|
+
log.warning("turnsafe: %s", fix) # messages[1] added_placeholder: no result for 'toolu_2'
|
|
216
|
+
```
|
|
217
|
+
|
|
218
|
+
| Problem | What `repair()` does |
|
|
219
|
+
|---|---|
|
|
220
|
+
| A tool result whose call is gone, for example trimmed away | Drops it and keeps any other content in the same message. |
|
|
221
|
+
| A tool call with no result, for example after an interrupted run | Adds an error result saying the tool didn't run. With `on_missing="strip"` it removes the call instead. |
|
|
222
|
+
| A result in the wrong place, for example a user message got in between | Moves it back right after its call. |
|
|
223
|
+
| A second result for the same call | Keeps the first. |
|
|
224
|
+
| Anthropic: text before the `tool_result` blocks | Moves the results first. |
|
|
225
|
+
| Anthropic: the same `tool_use` id used twice in the history | Renames the later call and its result. |
|
|
226
|
+
| Anthropic: an empty user message, or a blank text block | Removes it. |
|
|
227
|
+
|
|
228
|
+
`repair()` also guarantees:
|
|
229
|
+
- It never mutates your list.
|
|
230
|
+
- A valid history comes back unchanged.
|
|
231
|
+
- The output passes `validate()`. The one exception is an Anthropic `system` message placed
|
|
232
|
+
first, which you have to move to the `system` parameter yourself.
|
|
233
|
+
- By default it never rewrites an assistant message, so Anthropic thinking blocks and their
|
|
234
|
+
signatures survive.
|
|
235
|
+
|
|
236
|
+
### `fit(messages, max_tokens, provider="auto", *, counter=approx_tokens, keep_system=True, keep_first_turn=False, thinking="keep")`
|
|
237
|
+
|
|
238
|
+
`fit()` drops the oldest whole turns until the history fits in `max_tokens`.
|
|
239
|
+
|
|
240
|
+
```python
|
|
241
|
+
result = turnsafe.fit_detailed(messages, max_tokens=100_000, keep_first_turn=True)
|
|
242
|
+
print(result.tokens, result.dropped_turns)
|
|
243
|
+
```
|
|
244
|
+
|
|
245
|
+
- `keep_system` keeps leading system and developer messages.
|
|
246
|
+
- `keep_first_turn` keeps the first turn, which is usually the original task.
|
|
247
|
+
- If even the latest turn doesn't fit, `fit()` raises `ContextBudgetError` instead of returning
|
|
248
|
+
a broken history.
|
|
249
|
+
- If the input breaks the rules, it raises `InvalidHistoryError`, so call `repair()` first.
|
|
250
|
+
- `max_tokens` covers the messages only. Leave room for tools, the system prompt and the reply.
|
|
251
|
+
- The default counter is an offline estimate: about 4 characters per token, with images counted
|
|
252
|
+
by size rather than by their base64 length. Pass `counter=` for exact numbers, for example
|
|
253
|
+
with `tiktoken` or Anthropic's token-counting endpoint.
|
|
254
|
+
|
|
255
|
+
### Rules checked
|
|
256
|
+
|
|
257
|
+
| Code | Providers | Meaning |
|
|
258
|
+
|---|---|---|
|
|
259
|
+
| `unanswered_tool_call` | both | A tool call has no result |
|
|
260
|
+
| `orphan_tool_result` | both | A result answers no call in the preceding assistant turn |
|
|
261
|
+
| `misplaced_tool_result` | both | The result exists but isn't directly after its call |
|
|
262
|
+
| `duplicate_tool_result` | both | A call has more than one result |
|
|
263
|
+
| `duplicate_tool_call_id` | both | A tool call id is repeated. For Anthropic, this means anywhere in the history. |
|
|
264
|
+
| `tool_result_not_first` | Anthropic | Text comes before the `tool_result` blocks in the user turn |
|
|
265
|
+
| `system_in_messages` | Anthropic | An error when first. Elsewhere a warning, because only Opus 4.8+, Opus 5, Fable and Mythos accept it. |
|
|
266
|
+
| `empty_content` | Anthropic | An empty user message, or a blank text block outside a `tool_result` |
|
|
267
|
+
| `first_message_not_user` | Anthropic | Warning: Bedrock and some deployments require a user message first |
|
|
268
|
+
|
|
269
|
+
The Anthropic API merges consecutive messages with the same role, and turnsafe does the same.
|
|
270
|
+
Results split over several user messages in a row, or calls split over several assistant
|
|
271
|
+
messages, are valid.
|
|
272
|
+
|
|
273
|
+
## Anthropic thinking blocks
|
|
274
|
+
|
|
275
|
+
By default, turnsafe never touches `thinking` or `redacted_thinking` blocks.
|
|
276
|
+
|
|
277
|
+
**On Claude Fable 5.1 and Opus 5.5**, a replayed thinking block may be valid only while every
|
|
278
|
+
message before it is unchanged. This is enforced by default for accounts created on or after
|
|
279
|
+
2026-08-31. Trimming the front of the history then makes the API reject later thinking blocks
|
|
280
|
+
with ``Invalid `signature` in `thinking` block``. There are two fixes:
|
|
281
|
+
|
|
282
|
+
- Set `thinking.block_binding.prefix_mismatch_behavior` to `"drop_block"` on the request. That's
|
|
283
|
+
the field the API error names, and the API then drops the stale blocks itself.
|
|
284
|
+
- Or pass `thinking="strip_after_change"` to `fit()` and `repair()`. turnsafe then removes
|
|
285
|
+
thinking blocks after the first change, which is the documented recovery.
|
|
286
|
+
|
|
287
|
+
Don't strip on other models. They expect the latest turn's thinking to be replayed. See the
|
|
288
|
+
[details](https://github.com/Yasir-Khan-7/turnsafe/blob/main/docs/rules.md#anthropic-thinking-blocks).
|
|
289
|
+
|
|
290
|
+
## How it compares
|
|
291
|
+
|
|
292
|
+
| | turnsafe | litellm `modify_params` | LangChain `trim_messages` | pydantic-ai history repair |
|
|
293
|
+
|---|---|---|---|---|
|
|
294
|
+
| Works on plain SDK dicts, with no framework | ✅ | ❌ inside `litellm.completion()` | ❌ LangChain messages | ❌ pydantic-ai messages |
|
|
295
|
+
| OpenAI **and** Anthropic formats | ✅ | Anthropic path only | ✅ | ✅ |
|
|
296
|
+
| Reports problems without changing anything | ✅ `validate()` | ❌ | ❌ | ❌ |
|
|
297
|
+
| Repairs broken pairs | ✅ | ✅ | ❌ | ✅ |
|
|
298
|
+
| Trims to a token budget | ✅ | ❌ | ✅ | ❌ |
|
|
299
|
+
| Trimming can't orphan a tool call | ✅ property-tested | – | ❌ [#33245](https://github.com/langchain-ai/langchain/issues/33245) | – |
|
|
300
|
+
| Tells you what it changed | ✅ | ❌ | ❌ | ❌ |
|
|
301
|
+
|
|
302
|
+
If your framework already handles this for you, keep using it. `turnsafe` is for agent loops
|
|
303
|
+
written directly on the provider SDKs, for histories you store and reload yourself, and for
|
|
304
|
+
anyone who wants to check a history before sending it. For how the problem was chosen and what
|
|
305
|
+
we learned from existing tools, see
|
|
306
|
+
[`docs/research.md`](https://github.com/Yasir-Khan-7/turnsafe/blob/main/docs/research.md).
|
|
307
|
+
|
|
308
|
+
## FAQ
|
|
309
|
+
|
|
310
|
+
**Does it call any API or send my data anywhere?**
|
|
311
|
+
No. It's pure Python working on your list in memory.
|
|
312
|
+
|
|
313
|
+
**Can I pass SDK objects instead of dicts?**
|
|
314
|
+
Yes. Messages and content blocks that have `model_dump()` are accepted, for example
|
|
315
|
+
`response.choices[0].message` or `response.content`. The output is always plain dicts.
|
|
316
|
+
|
|
317
|
+
**What if a tool call id isn't a string?**
|
|
318
|
+
turnsafe raises a `ValueError` naming the position, for example
|
|
319
|
+
`messages[3].tool_calls[0].id must be a non-empty string`. A call without a usable id can't be
|
|
320
|
+
matched to its result, so it isn't silently ignored.
|
|
321
|
+
|
|
322
|
+
**What does the placeholder result say?**
|
|
323
|
+
It says the tool call wasn't executed or its result was lost, and that the model shouldn't
|
|
324
|
+
assume it succeeded. Change it with `placeholder="..."`.
|
|
325
|
+
|
|
326
|
+
**My history ends with an assistant tool call because I'm about to run the tools. Is that
|
|
327
|
+
an error?**
|
|
328
|
+
`validate()` reports it, because the API would reject it if you sent it. Don't `repair()`
|
|
329
|
+
a history while you're about to add the real results yourself.
|
|
330
|
+
|
|
331
|
+
**Should I use this with server-side compaction?**
|
|
332
|
+
Yes, they work together. Anthropic context editing and OpenAI Responses compaction make the
|
|
333
|
+
context smaller, and turnsafe makes sure what you send is valid.
|
|
334
|
+
|
|
335
|
+
## Status and roadmap
|
|
336
|
+
|
|
337
|
+
- **v0.1:** OpenAI Chat Completions and Anthropic Messages.
|
|
338
|
+
- **Planned:**
|
|
339
|
+
- OpenAI Responses API items
|
|
340
|
+
- Gemini `contents`
|
|
341
|
+
- an optional LangChain adapter
|
|
342
|
+
- Legacy OpenAI `function_call` messages aren't checked.
|
|
343
|
+
|
|
344
|
+
The version follows [Semantic Versioning](https://semver.org/). Violation codes and fix actions
|
|
345
|
+
are public API. Changes are listed in the
|
|
346
|
+
[CHANGELOG](https://github.com/Yasir-Khan-7/turnsafe/blob/main/CHANGELOG.md).
|
|
347
|
+
|
|
348
|
+
## Contributing
|
|
349
|
+
|
|
350
|
+
Issues and pull requests are welcome. See
|
|
351
|
+
[CONTRIBUTING.md](https://github.com/Yasir-Khan-7/turnsafe/blob/main/CONTRIBUTING.md) to get
|
|
352
|
+
set up. If you work with an AI coding agent, point it at
|
|
353
|
+
[AGENT.md](https://github.com/Yasir-Khan-7/turnsafe/blob/main/AGENT.md). To report a security
|
|
354
|
+
issue, see [SECURITY.md](https://github.com/Yasir-Khan-7/turnsafe/blob/main/SECURITY.md).
|
|
355
|
+
|
|
356
|
+
Found a provider rule turnsafe gets wrong? Open an issue with the exact API error text. That's
|
|
357
|
+
the most valuable report you can send.
|
|
358
|
+
|
|
359
|
+
## License
|
|
360
|
+
|
|
361
|
+
[MIT](https://github.com/Yasir-Khan-7/turnsafe/blob/main/LICENSE) © Yasir Khan
|