corollary 0.1.0a1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (71) hide show
  1. corollary-0.1.0a1/.gitignore +40 -0
  2. corollary-0.1.0a1/CHANGELOG.md +54 -0
  3. corollary-0.1.0a1/LICENSE +21 -0
  4. corollary-0.1.0a1/PKG-INFO +367 -0
  5. corollary-0.1.0a1/README.md +325 -0
  6. corollary-0.1.0a1/docs/api.md +271 -0
  7. corollary-0.1.0a1/docs/architecture.md +169 -0
  8. corollary-0.1.0a1/docs/assets/corollary-demo.gif +0 -0
  9. corollary-0.1.0a1/docs/concepts.md +167 -0
  10. corollary-0.1.0a1/docs/faq.md +66 -0
  11. corollary-0.1.0a1/docs/getting-started.md +184 -0
  12. corollary-0.1.0a1/docs/guides/agents.md +211 -0
  13. corollary-0.1.0a1/docs/guides/belief-base.md +248 -0
  14. corollary-0.1.0a1/docs/guides/confidence.md +344 -0
  15. corollary-0.1.0a1/docs/guides/conflicts.md +111 -0
  16. corollary-0.1.0a1/docs/guides/contract.md +151 -0
  17. corollary-0.1.0a1/docs/guides/documents.md +57 -0
  18. corollary-0.1.0a1/docs/guides/models.md +117 -0
  19. corollary-0.1.0a1/docs/guides/persistence.md +70 -0
  20. corollary-0.1.0a1/docs/guides/time.md +102 -0
  21. corollary-0.1.0a1/docs/guides/verification.md +123 -0
  22. corollary-0.1.0a1/docs/index.md +68 -0
  23. corollary-0.1.0a1/docs/related-work.md +83 -0
  24. corollary-0.1.0a1/examples/agent_claude.py +43 -0
  25. corollary-0.1.0a1/examples/agent_offline.py +124 -0
  26. corollary-0.1.0a1/examples/self_repairing_report.py +186 -0
  27. corollary-0.1.0a1/pyproject.toml +106 -0
  28. corollary-0.1.0a1/src/corollary/__init__.py +114 -0
  29. corollary-0.1.0a1/src/corollary/_version.py +1 -0
  30. corollary-0.1.0a1/src/corollary/agent.py +634 -0
  31. corollary-0.1.0a1/src/corollary/belief.py +280 -0
  32. corollary-0.1.0a1/src/corollary/changes.py +110 -0
  33. corollary-0.1.0a1/src/corollary/conflict.py +99 -0
  34. corollary-0.1.0a1/src/corollary/contract.py +267 -0
  35. corollary-0.1.0a1/src/corollary/errors.py +67 -0
  36. corollary-0.1.0a1/src/corollary/formula.py +120 -0
  37. corollary-0.1.0a1/src/corollary/justification.py +127 -0
  38. corollary-0.1.0a1/src/corollary/kernel.py +1705 -0
  39. corollary-0.1.0a1/src/corollary/ledger.py +221 -0
  40. corollary-0.1.0a1/src/corollary/models/__init__.py +7 -0
  41. corollary-0.1.0a1/src/corollary/models/anthropic.py +108 -0
  42. corollary-0.1.0a1/src/corollary/models/base.py +111 -0
  43. corollary-0.1.0a1/src/corollary/models/openai.py +71 -0
  44. corollary-0.1.0a1/src/corollary/projector.py +156 -0
  45. corollary-0.1.0a1/src/corollary/proof.py +298 -0
  46. corollary-0.1.0a1/src/corollary/py.typed +0 -0
  47. corollary-0.1.0a1/src/corollary/resolvers.py +116 -0
  48. corollary-0.1.0a1/src/corollary/rules.py +66 -0
  49. corollary-0.1.0a1/src/corollary/textmatch.py +69 -0
  50. corollary-0.1.0a1/src/corollary/tools.py +179 -0
  51. corollary-0.1.0a1/src/corollary/trust.py +91 -0
  52. corollary-0.1.0a1/src/corollary/verify.py +383 -0
  53. corollary-0.1.0a1/tests/__init__.py +0 -0
  54. corollary-0.1.0a1/tests/conftest.py +54 -0
  55. corollary-0.1.0a1/tests/test_agent.py +519 -0
  56. corollary-0.1.0a1/tests/test_belief.py +69 -0
  57. corollary-0.1.0a1/tests/test_confidence.py +273 -0
  58. corollary-0.1.0a1/tests/test_contract.py +87 -0
  59. corollary-0.1.0a1/tests/test_examples.py +29 -0
  60. corollary-0.1.0a1/tests/test_formula.py +57 -0
  61. corollary-0.1.0a1/tests/test_kernel.py +424 -0
  62. corollary-0.1.0a1/tests/test_ledger.py +114 -0
  63. corollary-0.1.0a1/tests/test_models.py +125 -0
  64. corollary-0.1.0a1/tests/test_persistence.py +75 -0
  65. corollary-0.1.0a1/tests/test_projector.py +65 -0
  66. corollary-0.1.0a1/tests/test_proof.py +85 -0
  67. corollary-0.1.0a1/tests/test_resolvers.py +79 -0
  68. corollary-0.1.0a1/tests/test_textmatch.py +47 -0
  69. corollary-0.1.0a1/tests/test_time.py +52 -0
  70. corollary-0.1.0a1/tests/test_tools.py +86 -0
  71. corollary-0.1.0a1/tests/test_verify.py +156 -0
@@ -0,0 +1,40 @@
1
+ # Byte-compiled and cache files
2
+ __pycache__/
3
+ *.py[cod]
4
+ .pytest_cache/
5
+ .mypy_cache/
6
+ .ruff_cache/
7
+
8
+ # Build and packaging
9
+ build/
10
+ dist/
11
+ *.egg-info/
12
+ .eggs/
13
+
14
+ # Virtual environments
15
+ .venv/
16
+ venv/
17
+ env/
18
+
19
+ # Test and coverage output
20
+ .coverage
21
+ .coverage.*
22
+ coverage.xml
23
+ htmlcov/
24
+
25
+ # Documentation build
26
+ site/
27
+
28
+ # Editors and OS
29
+ .idea/
30
+ .vscode/
31
+ *.swp
32
+ .DS_Store
33
+ Thumbs.db
34
+
35
+ # Local belief base snapshots and secrets
36
+ *.beliefs.json
37
+ .env
38
+
39
+ # Local folders
40
+ local/
@@ -0,0 +1,54 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented here. The format follows
4
+ [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and the project uses
5
+ [Semantic Versioning](https://semver.org/). Until 1.0, minor versions may contain breaking changes.
6
+
7
+ ## [Unreleased]
8
+
9
+ ## [0.1.0a1] - 2026-10-02
10
+
11
+ First public pre-release.
12
+
13
+ ### Added
14
+
15
+ - **Kernel.** `BeliefBase`: a justification-based truth maintenance system over versioned beliefs.
16
+ Incremental labeling restricted to the affected region, well-founded support (cycles alone never keep a
17
+ belief `IN`), non-monotonic `unless` justifications with odd-loop detection and rollback, retraction and
18
+ restoration, and a change log with reasons for every transition.
19
+ - **Re-derivation.** `propagate()` re-runs rules automatically and model-derived beliefs through a
20
+ re-deriver, with early cutoff when a recomputed value is unchanged.
21
+ - **Rules.** The `@rule` decorator for deterministic derivations, replayable by the verifier.
22
+ - **Conflicts.** Value conflicts and user-defined `Constraint`s are first-class `Conflict` objects
23
+ carrying both support chains. `resolve()`, plus the resolvers `PreferHigherConfidence`, `PreferNewest`,
24
+ `PreferSource` and `AskHuman`.
25
+ - **Time.** Validity windows (`ttl`, `valid_until`) on evidence, `refresh()`, `stale()`, and in-place
26
+ renewal that avoids needless cascades.
27
+ - **Documents.** `add_document()` and `cite()` with span-checked quotes and value checks.
28
+ - **Proofs.** `Proof` snapshots with tree rendering, Mermaid and Graphviz export, JSON round-trip and diffs.
29
+ - **Verification.** `Verifier` with structure, grounding, arithmetic (formula and rule replay), citation,
30
+ temporal and numeric-provenance checks, plus a `Check` protocol for custom checks.
31
+ - **Agent runtime.** `Agent` with the claim contract, runtime-executed tools that emit premises, a
32
+ projector that builds every context from `IN` beliefs only, conservative and declared dependency
33
+ policies, `repair()`, `reverify()` and ablation-based `narrow()`.
34
+ - **Models.** `AnthropicModel` (Claude, via the official SDK), `OpenAIModel` (any Chat Completions
35
+ endpoint), `CallableModel` and `ScriptedModel`.
36
+ - **Learned reliability.** `TrustLedger` records when sources turn out right or wrong and re-estimates
37
+ their reliability from the trust policy's prior, optionally forgetting old outcomes
38
+ (`memory_half_life`). Ledgers can be shared across belief bases and are saved with a base that owns one.
39
+ - **Outcomes.** `retract(..., fault="source")`, `resolve(..., learn=True)`, `AskHuman` decisions,
40
+ independent confirmations, `record_outcome()`, and the agent's verified formulas and citations
41
+ (`learn_from_checks`) all feed the ledger.
42
+ - **Corroboration.** Agreeing premises from independent origins combine by noisy-OR. Sources declare a
43
+ shared origin with `origin=` (on `assert_`, `Source` and `@tool`); `TrustPolicy(corroboration=False)`
44
+ turns it off.
45
+ - **Evidence decay.** `half_life=` on `assert_` and `@tool` fades confidence without changing status.
46
+ `faded()` lists faded premises, and `Agent.reverify()` now refreshes them too.
47
+ - **Self-consistency.** `Agent(self_consistency=k)` samples unverified claims `k` times and uses the
48
+ agreement rate as the step's certainty.
49
+ - A confidence guide in the documentation.
50
+ - **Persistence.** JSON snapshots with `save()` / `load()`, including the trust ledger.
51
+ - Examples: a 20-conclusion self-repairing report, an offline agent, and a Claude agent.
52
+
53
+ [Unreleased]: https://github.com/gabe-santana/corollary/compare/v0.1.0a1...HEAD
54
+ [0.1.0a1]: https://github.com/gabe-santana/corollary/releases/tag/v0.1.0a1
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Gabriel Santana and the Corollary contributors
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,367 @@
1
+ Metadata-Version: 2.5
2
+ Name: corollary
3
+ Version: 0.1.0a1
4
+ Summary: An agent runtime where the unit of state is a belief, not a message.
5
+ Project-URL: Homepage, https://github.com/gabe-santana/corollary
6
+ Project-URL: Documentation, https://gabe-santana.github.io/corollary/
7
+ Project-URL: Repository, https://github.com/gabe-santana/corollary
8
+ Project-URL: Issues, https://github.com/gabe-santana/corollary/issues
9
+ Project-URL: Changelog, https://github.com/gabe-santana/corollary/blob/main/CHANGELOG.md
10
+ Author: Gabriel Santana
11
+ License-Expression: MIT
12
+ License-File: LICENSE
13
+ Keywords: agents,belief-revision,knowledge-representation,llm,provenance,truth-maintenance,verification
14
+ Classifier: Development Status :: 2 - Pre-Alpha
15
+ Classifier: Intended Audience :: Developers
16
+ Classifier: Intended Audience :: Science/Research
17
+ Classifier: Operating System :: OS Independent
18
+ Classifier: Programming Language :: Python :: 3
19
+ Classifier: Programming Language :: Python :: 3 :: Only
20
+ Classifier: Programming Language :: Python :: 3.10
21
+ Classifier: Programming Language :: Python :: 3.11
22
+ Classifier: Programming Language :: Python :: 3.12
23
+ Classifier: Programming Language :: Python :: 3.13
24
+ Classifier: Programming Language :: Python :: 3.14
25
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
26
+ Classifier: Topic :: Software Development :: Libraries :: Python Modules
27
+ Classifier: Typing :: Typed
28
+ Requires-Python: >=3.10
29
+ Provides-Extra: anthropic
30
+ Requires-Dist: anthropic>=1.0; extra == 'anthropic'
31
+ Provides-Extra: dev
32
+ Requires-Dist: mypy>=1.10; extra == 'dev'
33
+ Requires-Dist: pytest-cov>=5.0; extra == 'dev'
34
+ Requires-Dist: pytest>=8.0; extra == 'dev'
35
+ Requires-Dist: ruff>=0.6; extra == 'dev'
36
+ Provides-Extra: docs
37
+ Requires-Dist: mkdocs-material>=9.5; extra == 'docs'
38
+ Requires-Dist: mkdocs<2,>=1.6; extra == 'docs'
39
+ Provides-Extra: openai
40
+ Requires-Dist: openai>=1.0; extra == 'openai'
41
+ Description-Content-Type: text/markdown
42
+
43
+ <div align="center">
44
+
45
+ # Corollary
46
+
47
+ **An agent runtime where the unit of state is a belief, not a message.**
48
+
49
+ Every conclusion your agent reaches carries its proof.
50
+ Correct one fact, and everything that followed from it updates itself.
51
+
52
+ [![CI](https://github.com/gabe-santana/corollary/actions/workflows/ci.yml/badge.svg)](https://github.com/gabe-santana/corollary/actions/workflows/ci.yml)
53
+ [![Status: pre-alpha](https://img.shields.io/badge/status-pre--alpha-orange)](#roadmap)
54
+ [![License: MIT](https://img.shields.io/badge/license-MIT-blue)](LICENSE)
55
+ [![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue)](#installation)
56
+ [![Typed](https://img.shields.io/badge/typing-mypy%20strict-blue)](pyproject.toml)
57
+ [![PRs welcome](https://img.shields.io/badge/PRs-welcome-brightgreen)](CONTRIBUTING.md)
58
+
59
+ [Documentation](docs/index.md) · [Getting started](docs/getting-started.md) · [Examples](examples/) · [Contributing](CONTRIBUTING.md)
60
+
61
+ <br>
62
+
63
+ <img src="docs/assets/corollary-demo.gif" alt="Animated diagram: beliefs linked by what they follow from. Q2 revenue is corrected, every conclusion that depended on it goes OUT, and only those are re-derived, while an independent risk belief stays untouched." width="860">
64
+
65
+ </div>
66
+
67
+ ---
68
+
69
+ > ⚠️ **Corollary is pre-alpha.** The v0.1 kernel and agent runtime are implemented and tested, and most of
70
+ > v0.2 and v0.3 are in. Interfaces may still change before 1.0. If the idea resonates, star the repo, open
71
+ > an issue, or help shape the belief schema.
72
+
73
+ ## Why "Corollary"
74
+
75
+ In mathematics, a corollary is a result that follows directly from something already proven. It stands only as long as the theorem behind it stands.
76
+
77
+ That is the contract Corollary enforces for AI agents: every conclusion must follow from evidence, and when the evidence falls, the conclusion falls with it.
78
+
79
+ ## The problem
80
+
81
+ Every agent framework today stores state the same way: as a **message log**. A chat transcript, growing step by step.
82
+
83
+ That single design choice is why long-running agents rot:
84
+
85
+ - A hallucination at step 3 is just text. At step 40 the agent reads it with the same trust as a verified tool result.
86
+ - Nobody can answer *why* the agent believes something without re-reading the entire transcript.
87
+ - When one input turns out to be wrong, the only fix is to re-run everything, or hope a human catches every downstream mistake.
88
+ - Two sources disagree, and the model silently picks one. You never learn there was a conflict.
89
+
90
+ Better models reduce these failures. They cannot eliminate them, because the problem isn't the model. It's the data structure.
91
+
92
+ ## The idea
93
+
94
+ Corollary replaces the message log with a **belief base**.
95
+
96
+ A belief is a claim plus everything needed to trust it:
97
+
98
+ ```text
99
+ Belief
100
+ ├── claim "Q3 revenue grew 4.65% over Q2"
101
+ ├── confidence 0.92
102
+ ├── source tool:sec_filings | model:any-llm | human:alice
103
+ ├── valid_until 2026-12-31
104
+ └── follows_from [ belief:revenue_Q2, belief:revenue_Q3, rule:growth ]
105
+ ```
106
+
107
+ Underneath sits a **Truth Maintenance System** (Doyle, 1979): every belief records what supports it, and when that support is withdrawn, dependent beliefs are retracted automatically. A message log still exists, but only as a *view* derived from the belief base, never as the source of truth.
108
+
109
+ Truth maintenance never took off for general reasoning because humans had to write every justification by hand. LLMs can now propose them at scale. That was the missing piece.
110
+
111
+ ## What this unlocks
112
+
113
+ **Retraction cascades.** A tool returns a corrected figure. Every conclusion derived from the old value is retracted and re-derived, surgically. You fix one fact, not a transcript.
114
+
115
+ **Contradictions as first-class objects.** When evidence conflicts, Corollary raises a `Conflict` carrying both support chains instead of letting the model guess. Resolve it by policy, by source rank, or by asking a human.
116
+
117
+ **Proof-carrying answers.** Every answer ships with its proof: the graph of beliefs it follows from. A deterministic verifier checks it. Cited spans exist, arithmetic re-executes, dates are consistent, and every leaf traces back to a tool, a document, or a person. Trust moves from the model to the proof.
118
+
119
+ **Beliefs that expire.** A stock price is valid for a minute; a company's headquarters for a year. Stale beliefs trigger re-verification instead of silent reuse.
120
+
121
+ **Trust that is earned.** Every source, the model included, is trusted as much as its track record justifies. When a source is shown to be wrong, everything it reported counts for less. Independent sources that agree reinforce each other. See [Confidence](docs/guides/confidence.md).
122
+
123
+ **Auditable memory.** New sessions inherit verified beliefs with their proofs attached, not raw transcripts.
124
+
125
+ **Multi-agent by argument** *(planned)*. Agents exchange beliefs with their support. The receiver can accept, reject, or demand proof.
126
+
127
+ ## A taste of the API
128
+
129
+ ```python
130
+ from corollary import Agent, BeliefBase, tool
131
+
132
+ kb = BeliefBase()
133
+
134
+
135
+ @tool(trust="high")
136
+ def get_revenue(quarter: str) -> float:
137
+ """Quarterly revenue in USD from SEC filings."""
138
+ ...
139
+
140
+
141
+ agent = Agent(model="anthropic:claude-opus-5-5", beliefs=kb, tools=[get_revenue])
142
+
143
+ report = agent.run("Compare Q2 and Q3 revenue and assess the growth trend.")
144
+
145
+ print(report.answer)
146
+ print(report.proof) # the graph of beliefs the answer follows from
147
+ report.verify() # deterministic checks over the proof
148
+ ```
149
+
150
+ Now an upstream figure is corrected:
151
+
152
+ ```python
153
+ kb.retract("revenue:Q2", reason="restated in 10-K/A")
154
+ kb.assert_("revenue:Q2", 4.1e9, source="tool:get_revenue")
155
+
156
+ for change in agent.repair():
157
+ print(change)
158
+ # OUT revenue:Q2 (retracted: restated in 10-K/A)
159
+ # OUT growth:Q3_vs_Q2 (lost support: revenue:Q2)
160
+ # OUT trend:Q3 (lost support: revenue:Q2, growth:Q3_vs_Q2)
161
+ # OUT answer (lost support: revenue:Q2, growth:Q3_vs_Q2, trend:Q3)
162
+ # IN revenue:Q2 (asserted by tool:get_revenue)
163
+ # IN growth:Q3_vs_Q2 (re-derived: 9.7561)
164
+ # IN trend:Q3 (re-derived: 'strong')
165
+ # IN answer (re-derived: 'Q3 revenue grew 9.76% over Q2: strong growth.')
166
+
167
+ print(report.answer) # the repaired answer
168
+ ```
169
+
170
+ No re-run. No transcript archaeology. A diff of what changed. During repair, the model sees only the *current* inputs of each belief it re-derives. The retracted figure never reaches it again.
171
+
172
+ The kernel works without a model too. In [`examples/self_repairing_report.py`](examples/self_repairing_report.py), a 20-conclusion financial report repairs itself after a restatement: 12 conclusions change, 7 are never touched, the growth trend is recomputed, found unchanged, and stops the cascade, all in 13 rule evaluations and zero model calls.
173
+
174
+ ## How it works
175
+
176
+ Corollary is not a system prompt asking a model to "track its reasoning." A prompt is a request; a model can ignore it. Every guarantee in Corollary is **enforced by code outside the model**.
177
+
178
+ The key inversion: **the model doesn't own the state, the runtime does.** The model is a stateless *proposer*. It never sees a transcript and never writes state directly. It receives a context built by the runtime and can only respond with structured claims, which the runtime validates before accepting.
179
+
180
+ ```mermaid
181
+ flowchart LR
182
+ U[Task] --> P[LLM proposer]
183
+ P -->|claims + dependencies| C[Contract validator]
184
+ T[Tools] -->|executed by runtime| PR[Premises]
185
+ PR --> K
186
+ C --> K[(Kernel: belief graph + TMS)]
187
+ K -->|IN beliefs only| V[Projector]
188
+ V --> P
189
+ K --> X[Conflict detector]
190
+ X --> R[Resolver: policy / human]
191
+ R --> K
192
+ K --> A[Answer + proof]
193
+ A --> VER[Verifier]
194
+ ```
195
+
196
+ **1. The kernel (pure code, no LLM).** A dependency graph plus a labeling algorithm. Each node is a belief; each edge is a justification. The kernel computes which beliefs are `IN` or `OUT` from the current graph. It is deterministic, fast, and testable like any data structure. Retraction walks only the affected region, so cost scales with the change, not the size of the base.
197
+
198
+ **2. The contract (structured output, enforced).** The model must answer in a schema: the claim, its value, and the IDs of the beliefs it follows from. The runtime rejects anything that doesn't parse or that depends on a belief that doesn't exist or is currently `OUT`. Tool calls go *through the runtime*, which executes them and records the results as premises itself. The model cannot fabricate a tool result.
199
+
200
+ **3. The projector (what makes retraction real).** Every model call receives a context assembled from `IN` beliefs only. When a belief is retracted it doesn't just get a flag; it disappears from everything the model can see. There is no transcript for a wrong fact to leak from. Forgetting is structural, not requested.
201
+
202
+ **4. The verifier.** Answers are exported with their proof and checked deterministically: arithmetic re-executes, citations are span-matched, temporal claims are checked for consistency.
203
+
204
+ ### What is LLM and what is code
205
+
206
+ | The LLM does | Code does |
207
+ |---|---|
208
+ | Proposes claims | Records premises from tools, documents, humans |
209
+ | Chooses which tools to call | Executes tools and stores their results |
210
+ | Re-derives conclusions after a retraction | Validates every claim against the contract |
211
+ | | Verifies arithmetic, citations, dates |
212
+ | | Computes `IN` / `OUT` for every belief |
213
+ | | Builds every context window |
214
+ | | Cascades retractions |
215
+
216
+ Corollary doesn't make the model more honest. It makes the model's honesty irrelevant to whether bad facts propagate.
217
+
218
+ ### The hard part, stated honestly
219
+
220
+ A model may declare that a conclusion depends on A when it also used B. Corollary handles this in layers:
221
+
222
+ - **Conservative default.** A claim is assumed to depend on *everything in the context it was generated from* (`dependencies="conservative"`). The projector guarantees the model could not have used anything else.
223
+ - **Narrowing by ablation.** `agent.narrow(key)` regenerates a claim with one belief removed from context at a time. If the answer doesn't change, that dependency is pruned.
224
+ - **Deterministic checks** on numeric and cited claims: formulas are re-executed, quotes are span-matched, and numbers that appear from nowhere are flagged.
225
+
226
+ The resulting guarantee is **over-retraction, never under-retraction**. You may occasionally re-derive something unnecessarily, but a retracted fact can never silently survive in a conclusion.
227
+
228
+ What code can't catch is listed too: see [limits, stated honestly](docs/architecture.md#limits-stated-honestly).
229
+
230
+ ## Core concepts
231
+
232
+ | Concept | What it is |
233
+ |---|---|
234
+ | `Belief` | A claim with confidence, source, and revision. Its status (`IN` / `OUT`) is computed by the kernel. |
235
+ | `Justification` | A link from a set of antecedent beliefs (and optionally a rule or formula) to a conclusion. |
236
+ | `Premise` | A belief grounded directly in a tool, document span, or human. No antecedents. |
237
+ | `Rule` | A deterministic derivation, re-run automatically when its inputs change. |
238
+ | `Conflict` | Two `IN` beliefs that cannot both hold, with both support chains attached. |
239
+ | `Constraint` | An invariant over beliefs; a violation raises a `Conflict`. |
240
+ | `Proof` | The justification subgraph behind an answer. Exportable, verifiable, diffable. |
241
+ | `Projector` | Builds every model context from `IN` beliefs. |
242
+ | `TrustPolicy` | How much each source type is believed, and what is hidden from the model. |
243
+ | `TrustLedger` | Learns how reliable each source really is from its track record. |
244
+
245
+ Read more in [Core concepts](docs/concepts.md).
246
+
247
+ ## How Corollary differs
248
+
249
+ | | Message-log agents | Memory layers | Prompted confidence scoring | **Corollary** |
250
+ |---|---|---|---|---|
251
+ | Unit of state | Message | Fact or summary | Annotated message | Belief with justifications |
252
+ | Knows *why* it believes X | ❌ | Source only | Self-reported | ✅ full derivation |
253
+ | Fix one input, update dependents | Re-run | Manual | ❌ | ✅ automatic cascade |
254
+ | Retracted facts removed from context | ❌ | ❌ | ❌ | ✅ enforced by projector |
255
+ | Surfaces contradictions | ❌ | May overwrite | ❌ | ✅ first-class `Conflict` |
256
+ | Verifiable output | ❌ | ❌ | ❌ | ✅ proof + deterministic verifier |
257
+ | Where guarantees live | — | Storage | The prompt | The runtime |
258
+
259
+ Corollary is not a memory plugin and not a prompting strategy. The belief base *is* the agent; memory, context, tool results and inter-agent messages are all projections of it. See [Related work](docs/related-work.md) for the prior art it builds on.
260
+
261
+ ## Installation
262
+
263
+ Corollary requires Python 3.10+. The core has **no runtime dependencies**; model SDKs are optional extras.
264
+
265
+ ```bash
266
+ pip install corollary # the kernel and the agent runtime
267
+ pip install "corollary[anthropic]" # + the Claude adapter (official Anthropic SDK)
268
+ pip install "corollary[openai]" # + the OpenAI-compatible adapter
269
+ ```
270
+
271
+ To develop:
272
+
273
+ ```bash
274
+ git clone https://github.com/gabe-santana/corollary.git
275
+ cd corollary
276
+ pip install -e ".[dev]"
277
+ pytest
278
+ ```
279
+
280
+ Then follow [Getting started](docs/getting-started.md), or run an example. Neither of the first two needs an API key:
281
+
282
+ ```bash
283
+ python examples/self_repairing_report.py # 20 conclusions repair themselves, zero model calls
284
+ python examples/agent_offline.py # a full agent run with a scripted model
285
+ python examples/agent_claude.py # the same with Claude (needs ANTHROPIC_API_KEY)
286
+ ```
287
+
288
+ ## Documentation
289
+
290
+ | | |
291
+ |---|---|
292
+ | [Getting started](docs/getting-started.md) | Install, a self-repairing belief base, an offline agent, Claude |
293
+ | [Core concepts](docs/concepts.md) | Beliefs, revisions, justifications, labels, confidence, conflicts, proofs |
294
+ | [Guides](docs/guides/) | [Belief base](docs/guides/belief-base.md) · [Agents](docs/guides/agents.md) · [Contract](docs/guides/contract.md) · [Verification](docs/guides/verification.md) · [Conflicts](docs/guides/conflicts.md) · [Documents](docs/guides/documents.md) · [Time](docs/guides/time.md) · [Models](docs/guides/models.md) · [Persistence](docs/guides/persistence.md) |
295
+ | [Architecture and guarantees](docs/architecture.md) | How the kernel works, what is guaranteed, and the limits |
296
+ | [API reference](docs/api.md) | Every public class and method |
297
+ | [FAQ](docs/faq.md) | Quick answers |
298
+
299
+ ## Roadmap
300
+
301
+ **v0.1: the kernel** ✅
302
+ - [x] `Belief`, `Justification`, `BeliefBase` with JTMS labeling and retraction
303
+ - [x] Claim contract (structured output) with model-agnostic adapters
304
+ - [x] Runtime-executed tools that emit premises
305
+ - [x] Projector: context built from `IN` beliefs only
306
+ - [x] Flagship demo: a 20-conclusion research report that repairs itself when one input is corrected
307
+
308
+ **v0.2: proof and verification**
309
+ - [x] Deterministic verifiers: arithmetic re-execution, citation spans, temporal consistency
310
+ - [x] Proof export (JSON, Mermaid, Graphviz)
311
+ - [ ] Interactive proof graph viewer
312
+ - [x] Ablation-based dependency narrowing
313
+ - [x] Conflict detection and pluggable resolvers
314
+
315
+ **v0.3: time and memory**
316
+ - [x] Validity windows and automatic re-verification of stale beliefs
317
+ - [x] Persistent belief snapshots (JSON)
318
+ - [ ] Persistent belief stores (SQLite, Postgres)
319
+ - [ ] Cross-session inheritance policies for verified beliefs
320
+
321
+ **v0.4: many agents**
322
+ - [ ] Belief exchange protocol: accept / reject / demand proof
323
+ - [ ] Assumption-based (ATMS) mode for exploring alternative hypotheses in parallel
324
+ - [ ] MCP server so any agent can use a Corollary belief base
325
+
326
+ ## Open problems we want help with
327
+
328
+ These are the hard parts. Each one is a research contribution waiting to happen.
329
+
330
+ - **Confabulated dependencies.** How cheaply can we detect a declared dependency the model didn't actually use?
331
+ - **Granularity.** Too fine and the graph explodes; too coarse and cascades become blunt. What is the right atomic claim?
332
+ - **Equivalence.** When are two differently worded claims the same belief?
333
+ - **Developer experience.** It has to feel as simple as appending to a message list, or nobody will switch.
334
+ - **Benchmarks.** There is no standard evaluation for how well an agent recovers from a corrected input. We want to build one.
335
+
336
+ ## Contributing
337
+
338
+ Corollary is in the design phase, which is the best time to shape it. Good first contributions:
339
+
340
+ - Comment on the belief schema in [Discussions](https://github.com/gabe-santana/corollary/discussions)
341
+ - Propose a demo scenario from your domain (finance, law, science, operations)
342
+ - Implement a [verifier](docs/guides/verification.md#writing-a-check)
343
+ - Write a [model adapter](docs/guides/models.md#writing-an-adapter) for your favorite provider
344
+
345
+ See [CONTRIBUTING.md](CONTRIBUTING.md) to get started. Everyone taking part is expected to follow the [Code of Conduct](CODE_OF_CONDUCT.md). To report a vulnerability, see [SECURITY.md](SECURITY.md).
346
+
347
+ ## Background
348
+
349
+ Corollary stands on decades of work in knowledge representation:
350
+
351
+ - Jon Doyle, *A Truth Maintenance System* (1979)
352
+ - Johan de Kleer, *An Assumption-based TMS* (1986)
353
+ - Alchourrón, Gärdenfors & Makinson, the AGM theory of belief revision (1985)
354
+
355
+ What's new is the pairing: classical reason maintenance as the kernel, LLMs as the engine that proposes beliefs and justifications at scale.
356
+
357
+ ## License
358
+
359
+ MIT. See [LICENSE](LICENSE). You can use, modify and distribute Corollary freely, including in commercial and closed-source projects.
360
+
361
+ ---
362
+
363
+ <div align="center">
364
+
365
+ **Agents shouldn't just conclude. Their conclusions should follow.**
366
+
367
+ </div>