tiny-datalog 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- tiny_datalog-0.1.0/LICENSE +21 -0
- tiny_datalog-0.1.0/PKG-INFO +360 -0
- tiny_datalog-0.1.0/README.md +348 -0
- tiny_datalog-0.1.0/pyproject.toml +29 -0
- tiny_datalog-0.1.0/setup.cfg +4 -0
- tiny_datalog-0.1.0/tiny_datalog/__init__.py +42 -0
- tiny_datalog-0.1.0/tiny_datalog/containment.py +228 -0
- tiny_datalog-0.1.0/tiny_datalog/datalog.py +1455 -0
- tiny_datalog-0.1.0/tiny_datalog/incremental.py +424 -0
- tiny_datalog-0.1.0/tiny_datalog/magic.py +175 -0
- tiny_datalog-0.1.0/tiny_datalog/prolog.py +335 -0
- tiny_datalog-0.1.0/tiny_datalog/semantics.py +182 -0
- tiny_datalog-0.1.0/tiny_datalog/semiring.py +383 -0
- tiny_datalog-0.1.0/tiny_datalog/subsumption.py +475 -0
- tiny_datalog-0.1.0/tiny_datalog/tabling.py +220 -0
- tiny_datalog-0.1.0/tiny_datalog.egg-info/PKG-INFO +360 -0
- tiny_datalog-0.1.0/tiny_datalog.egg-info/SOURCES.txt +18 -0
- tiny_datalog-0.1.0/tiny_datalog.egg-info/dependency_links.txt +1 -0
- tiny_datalog-0.1.0/tiny_datalog.egg-info/entry_points.txt +8 -0
- tiny_datalog-0.1.0/tiny_datalog.egg-info/top_level.txt +1 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Andrew Goodchild
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,360 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: tiny-datalog
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: A Datalog engine small enough to read in an afternoon
|
|
5
|
+
Author: Andrew Goodchild
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/andrewgoodchild/tiny-datalog
|
|
8
|
+
Requires-Python: >=3.9
|
|
9
|
+
Description-Content-Type: text/markdown
|
|
10
|
+
License-File: LICENSE
|
|
11
|
+
Dynamic: license-file
|
|
12
|
+
|
|
13
|
+
# tiny-datalog
|
|
14
|
+
|
|
15
|
+
[](https://github.com/andrewgoodchild/tiny-datalog/actions/workflows/ci.yml)
|
|
16
|
+
|
|
17
|
+
**A logic engine small enough to read in an afternoon, and a course
|
|
18
|
+
that builds it up from nothing.**
|
|
19
|
+
|
|
20
|
+
## What is Datalog?
|
|
21
|
+
|
|
22
|
+
A query language from the early 1980s, with three properties worth
|
|
23
|
+
memorising:
|
|
24
|
+
|
|
25
|
+
- **Declarative.** You state what follows from what, never how to
|
|
26
|
+
compute it. No loops, no ordering, no stopping condition.
|
|
27
|
+
- **Recursive.** Rules may refer to themselves, so "at any depth"
|
|
28
|
+
questions are the native shape rather than something bolted on.
|
|
29
|
+
- **Terminating.** Every program finishes. Always. That is a theorem,
|
|
30
|
+
not a convention.
|
|
31
|
+
|
|
32
|
+
Datalog is useful when a query is recursive and SQL is fighting you,
|
|
33
|
+
or when a rule set has grown past the point where anyone can review
|
|
34
|
+
it. Here is what that buys you.
|
|
35
|
+
|
|
36
|
+
You deploy 12 services, sitting on 160 packages joined by 292
|
|
37
|
+
dependency edges. A CVE (Common Vulnerabilities and Exposures entry, a
|
|
38
|
+
published security flaw) lands on one package. Which services are
|
|
39
|
+
exposed?
|
|
40
|
+
|
|
41
|
+
You write down **facts**, things simply true:
|
|
42
|
+
|
|
43
|
+
```prolog
|
|
44
|
+
depends(pkg4, pkg13). % pkg4 pulls in pkg13
|
|
45
|
+
service(pkg4). % pkg4 is something we deploy
|
|
46
|
+
vulnerable(pkg21, cve_2026_0001). % pkg21 has the CVE
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
...and **rules** for deriving new facts. `:-` means "if", the comma
|
|
50
|
+
means "and", capitals are variables:
|
|
51
|
+
|
|
52
|
+
```prolog
|
|
53
|
+
% programs/supply-chain.dl
|
|
54
|
+
uses(X, Y) :- depends(X, Y). % you use what you depend on
|
|
55
|
+
uses(X, Z) :- depends(X, Y), uses(Y, Z). % ...and whatever that uses
|
|
56
|
+
|
|
57
|
+
exposed(S, C) :- service(S), uses(S, L), vulnerable(L, C).
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
The middle rule reads *X uses Z if X depends on Y and Y uses Z*. It
|
|
61
|
+
refers to itself, and that is the whole trick: applied over and over it
|
|
62
|
+
walks the graph to any depth, without you saying how deep or when to
|
|
63
|
+
stop.
|
|
64
|
+
|
|
65
|
+
```
|
|
66
|
+
$ python3 datalog.py -q 'exposed(S, C)' programs/supply-chain.dl
|
|
67
|
+
?- exposed(S, C)
|
|
68
|
+
exposed(pkg0, cve_2026_0001).
|
|
69
|
+
exposed(pkg4, cve_2026_0001).
|
|
70
|
+
exposed(pkg5, cve_2026_0001).
|
|
71
|
+
exposed(pkg8, cve_2026_0001).
|
|
72
|
+
(4 answers)
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
Four of twelve, and nothing you can see in the file says which four.
|
|
76
|
+
The 292 edges you stated imply 8,457 `uses` facts, and the answer lives
|
|
77
|
+
in those. Ask why, and you get the derivation the answer came out of,
|
|
78
|
+
which is also the remediation path:
|
|
79
|
+
|
|
80
|
+
```
|
|
81
|
+
$ python3 datalog.py --explain 'exposed(pkg4, cve_2026_0001)' programs/supply-chain.dl
|
|
82
|
+
?- explain exposed(pkg4, cve_2026_0001)
|
|
83
|
+
exposed(pkg4, cve_2026_0001) [via exposed(S, C) :- service(S), uses(S, L), vulnerable(L, C).]
|
|
84
|
+
service(pkg4) (base fact)
|
|
85
|
+
uses(pkg4, pkg21) [via uses(X, Z) :- depends(X, Y), uses(Y, Z).]
|
|
86
|
+
depends(pkg4, pkg13) (base fact)
|
|
87
|
+
uses(pkg13, pkg21) [via uses(X, Y) :- depends(X, Y).]
|
|
88
|
+
depends(pkg13, pkg21) (base fact)
|
|
89
|
+
vulnerable(pkg21, cve_2026_0001) (base fact)
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
When tomorrow's CVE arrives, the engine repairs those 8,457 facts
|
|
93
|
+
instead of recomputing them:
|
|
94
|
+
|
|
95
|
+
```
|
|
96
|
+
$ python3 incremental.py programs/supply-chain.dl -u 'vulnerable(pkg100, cve_2026_0002).'
|
|
97
|
+
materialised: 8766 facts
|
|
98
|
+
vulnerable(pkg100, cve_2026_0002).
|
|
99
|
+
-> {'inserted': 1, 'derived': 12} in 0.034s
|
|
100
|
+
(a from-scratch rebuild of this program: 0.822s)
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
Nothing to install:
|
|
104
|
+
|
|
105
|
+
```sh
|
|
106
|
+
git clone https://github.com/andrewgoodchild/tiny-datalog && cd tiny-datalog
|
|
107
|
+
python3 tests.py # 234 tests, ~12s
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
## Why the language choice decides what you can ask later
|
|
111
|
+
|
|
112
|
+
Frontier language models are getting good at reasoning and logic. They
|
|
113
|
+
have been trained on every logic trope in the canon, they can usually
|
|
114
|
+
spot a contradiction in prose, and when they meet more data than fits
|
|
115
|
+
in a context window they do the sensible thing: they write a program to
|
|
116
|
+
solve it.
|
|
117
|
+
|
|
118
|
+
So the question was never whether to run code. It is what the code
|
|
119
|
+
should be, and that choice decides which questions you can still ask
|
|
120
|
+
afterwards.
|
|
121
|
+
|
|
122
|
+
Rules that encode policy tend to outlive the query that prompted them.
|
|
123
|
+
They get reviewed, audited, inherited by someone who did not write
|
|
124
|
+
them, and changed under pressure. Sooner or later somebody asks why a
|
|
125
|
+
particular decision came out the way it did, and somebody else asks
|
|
126
|
+
whether the rules are even coherent before trusting any answer at all.
|
|
127
|
+
|
|
128
|
+
(The sign-off case has its own demonstration:
|
|
129
|
+
[lesson 17](lessons/17-writing-rules.md) writes a lending policy badly
|
|
130
|
+
twice, and `--explain` names which of two rules wrongly let a
|
|
131
|
+
suspended staff member borrow — three lines, no debugger.)
|
|
132
|
+
|
|
133
|
+
Both are questions about the rules, not about a run, and most languages
|
|
134
|
+
cannot answer them. Datalog can, because it gave things up:
|
|
135
|
+
|
|
136
|
+
| Question about the rules themselves | Datalog | A general-purpose program |
|
|
137
|
+
|---|---|---|
|
|
138
|
+
| Does it terminate? | yes, by construction | undecidable |
|
|
139
|
+
| Are the rules circular? | decidable, and it names the cycle | n/a, recursion is ordinary |
|
|
140
|
+
| Is there exactly one consistent answer? | decidable | not even well-formed |
|
|
141
|
+
| Is one rule redundant given another? | decidable for the non-recursive fragment | undecidable |
|
|
142
|
+
| Are two rule sets equivalent? | decidable for the non-recursive fragment | undecidable |
|
|
143
|
+
|
|
144
|
+
(Containment and equivalence become undecidable once recursion is
|
|
145
|
+
involved — Shmueli, 1993, which is why `containment.py` handles
|
|
146
|
+
conjunctive queries and refuses the rest.
|
|
147
|
+
[Lesson 16](lessons/16-containment.md) covers the boundary.)
|
|
148
|
+
|
|
149
|
+
Being declarative, recursive and terminating is not a feature list. It
|
|
150
|
+
is the trade that makes rules analysable, and three things follow from
|
|
151
|
+
it:
|
|
152
|
+
|
|
153
|
+
- **The artifact is reviewable by someone who is not a programmer.**
|
|
154
|
+
`eligible(P) :- member(P, H), qualifying_household(H), not employed(P).`
|
|
155
|
+
is nearly the policy sentence. Loops and mutable state are not.
|
|
156
|
+
- **Termination is a safety property**, not a nicety, when the rules
|
|
157
|
+
are going to be executed by something you do not supervise.
|
|
158
|
+
- **Provenance and incremental maintenance come from the semantics**
|
|
159
|
+
rather than from extra code that would itself need verifying.
|
|
160
|
+
|
|
161
|
+
Every feature in this repository is a check of that kind: stratification
|
|
162
|
+
catches circularity, `--models` catches contradiction and ambiguity,
|
|
163
|
+
`containment.py` catches redundancy, and `--explain` produces the
|
|
164
|
+
derivation somebody has to sign off on.
|
|
165
|
+
|
|
166
|
+
Two limits. An engine checks coherence, not intent: a rule set can pass
|
|
167
|
+
every check above and still be the wrong policy. And the trade only
|
|
168
|
+
pays when the rules genuinely outlive the query. For a one-off
|
|
169
|
+
question, or anything arithmetic-heavy, forty lines of ordinary code is
|
|
170
|
+
the better tool, and that covers most problems.
|
|
171
|
+
|
|
172
|
+
## What else you can ask it
|
|
173
|
+
|
|
174
|
+
The worked example above is one row of a table. The full version — 20
|
|
175
|
+
questions, each with the command that answers it and the lesson that
|
|
176
|
+
builds the machinery — is in
|
|
177
|
+
[getting started](lessons/getting-started.md), beside the reading
|
|
178
|
+
paths.
|
|
179
|
+
|
|
180
|
+
## Claims you can check
|
|
181
|
+
|
|
182
|
+
Beyond the benchmarks: **every shell command quoted in every lesson is
|
|
183
|
+
executed in CI and its quoted output diffed against reality** — the
|
|
184
|
+
course cannot silently rot, which is a rarer property than anything
|
|
185
|
+
else on this page. Every performance claim below is reproducible from
|
|
186
|
+
the shipped generator, including the one that goes the wrong way. Timings are on an
|
|
187
|
+
Apple M1 Pro, CPython 3.10, single core:
|
|
188
|
+
|
|
189
|
+
```sh
|
|
190
|
+
python3 benchmarks/generate.py chain 50 > chain50.dl
|
|
191
|
+
python3 benchmarks/generate.py chain 100 > chain100.dl
|
|
192
|
+
python3 benchmarks/generate.py chain 150 > chain150.dl
|
|
193
|
+
```
|
|
194
|
+
|
|
195
|
+
| Claim | Measured |
|
|
196
|
+
|---|---|
|
|
197
|
+
| Semi-naive beats naive, and the gap grows | chain-50: 0.7s → 0.07s (**10×**); chain-100: 10.7s → 0.22s (**48×**) |
|
|
198
|
+
| Magic sets makes a *selective* query goal-directed | `path(n140, X)`: **66 facts vs 11,175**, 0.05s vs 0.62s |
|
|
199
|
+
| Magic sets is not a free lunch | `path(n1, X)`: **11,325 facts vs 11,175**, 2.19s vs 0.62s |
|
|
200
|
+
|
|
201
|
+
The last two are the same rewriting on the same program. Magic sets
|
|
202
|
+
pays in proportion to how much the query's bindings prune; when demand
|
|
203
|
+
is the whole relation the guards are pure overhead.
|
|
204
|
+
[Lesson 7](lessons/07-magic-sets.md) works through why.
|
|
205
|
+
|
|
206
|
+
Correctness is checked by a seeded differential fuzzer that generates
|
|
207
|
+
stratified programs and demands semi-naive, naive, magic-sets and
|
|
208
|
+
tabled evaluation all agree, and that incremental maintenance matches
|
|
209
|
+
recomputation under random updates. 400 programs per run;
|
|
210
|
+
`TINY_DATALOG_FUZZ=3000 python3 tests.py DifferentialFuzzTests` soaks.
|
|
211
|
+
|
|
212
|
+
## Learning Datalog
|
|
213
|
+
|
|
214
|
+
`lessons/` is a complete course, no prior exposure assumed, every
|
|
215
|
+
example a runnable file, following the field's own history from 1977 to
|
|
216
|
+
the current research threads. The field's own recent lecture notes
|
|
217
|
+
observe that the literature advises people building Datalog engines
|
|
218
|
+
better than people trying to *use* one; this course does both halves
|
|
219
|
+
on purpose — sixteen lessons where the engine is the explanation, then
|
|
220
|
+
a lesson on authoring rules that survive review. And it is built to be
|
|
221
|
+
inherited: `git clone`, no dependencies, no hosted anything, and every
|
|
222
|
+
quoted transcript re-verified by CI — the exercises cannot rot. (For where each
|
|
223
|
+
technique ships — CodeQL, RDFox, Feldera, SNOMED and the rest —
|
|
224
|
+
[lesson 0](lessons/00-what-is-datalog.md) ends with the deployments.)
|
|
225
|
+
[lessons/getting-started.md](lessons/getting-started.md) has the titles
|
|
226
|
+
and the reading order, and
|
|
227
|
+
[lessons/glossary.md](lessons/glossary.md) defines every technical term
|
|
228
|
+
the course uses, and
|
|
229
|
+
[lessons/references.md](lessons/references.md) collects every work the
|
|
230
|
+
lessons cite.
|
|
231
|
+
|
|
232
|
+
Three of them teach things that are hard to find taught well anywhere
|
|
233
|
+
else, and they are the reason the course exists rather than just the
|
|
234
|
+
engine:
|
|
235
|
+
|
|
236
|
+
- **[Lesson 8](lessons/08-semirings.md)** proves that why-provenance
|
|
237
|
+
cannot be specialised into derivation counts, with a program that
|
|
238
|
+
prints the disproof: two facts with identical provenance and different
|
|
239
|
+
counts. That settles "materialise provenance once, specialise later,"
|
|
240
|
+
which is a real design-review question with a real answer.
|
|
241
|
+
- **[Lesson 16](lessons/16-containment.md)** shows that the containment
|
|
242
|
+
test you need for query minimisation is the search already sitting in
|
|
243
|
+
`datalog.py`: `_match` maps a rule body into a database,
|
|
244
|
+
`find_homomorphism` maps a rule body into another rule body. Same
|
|
245
|
+
backtracking, one level up.
|
|
246
|
+
- **[Lesson 4](lessons/04-closed-and-open-worlds.md)** contrasts the
|
|
247
|
+
two reasoners in this repository, which disagree about what absence
|
|
248
|
+
means, and leaves you with a habit: when you see `not`, ask whose
|
|
249
|
+
authority says this is absent.
|
|
250
|
+
|
|
251
|
+
Every lesson ends with exercises, and every exercise has a worked answer
|
|
252
|
+
in `exercises/` — runnable where the answer is a program, and executed
|
|
253
|
+
by the test suite so the answers cannot rot. `cases/` lets anyone add a
|
|
254
|
+
regression test without writing Python.
|
|
255
|
+
|
|
256
|
+
## What this is not, and what is missing on purpose
|
|
257
|
+
|
|
258
|
+
Not an engine to build a product on. Joins are nested-loop,
|
|
259
|
+
stable-model search is exhaustive, evaluation is batch. For real
|
|
260
|
+
workloads see Soufflé, clingo, RDFox or Feldera.
|
|
261
|
+
|
|
262
|
+
Deliberate omissions, because saying why teaches more than lacking them
|
|
263
|
+
quietly:
|
|
264
|
+
|
|
265
|
+
- **Arithmetic and comparisons.** A built-in isn't a relation you can
|
|
266
|
+
enumerate, so it must be *evaluated* the moment its operands bind —
|
|
267
|
+
which entangles correctness with join order and forces terms to
|
|
268
|
+
become trees. [Lesson 14](lessons/14-arithmetic.md) is the whole
|
|
269
|
+
story, including what to do instead.
|
|
270
|
+
- **Indexes and join planning.** Every join is a nested loop so the
|
|
271
|
+
algorithms stay one-screen readable. It is also why the magic-sets
|
|
272
|
+
timing above goes the way it does.
|
|
273
|
+
- **⊤ and role hierarchies** in the classifier — what ships is EL⊥
|
|
274
|
+
(disjointness and unsatisfiability detection included). SNOMED needs
|
|
275
|
+
ELH (EL plus role hierarchies) with right identities, which is what
|
|
276
|
+
the ELK and Snorocket reasoners implement and this does not.
|
|
277
|
+
- **A REPL (interactive prompt) and packaging.** `git clone` and run.
|
|
278
|
+
|
|
279
|
+
Aggregation used to be on this list;
|
|
280
|
+
[lesson 13](lessons/13-aggregation.md) is what promoting an omission
|
|
281
|
+
into a feature looks like.
|
|
282
|
+
|
|
283
|
+
## Using it in your own project
|
|
284
|
+
|
|
285
|
+
The engine is on PyPI, with no dependencies beyond the standard library:
|
|
286
|
+
|
|
287
|
+
```sh
|
|
288
|
+
pip install tiny-datalog
|
|
289
|
+
```
|
|
290
|
+
|
|
291
|
+
```python
|
|
292
|
+
from tiny_datalog import run_program, explain
|
|
293
|
+
|
|
294
|
+
engine = run_program(open("supply-chain.dl").read())
|
|
295
|
+
for service, cve in sorted(engine.rels["exposed"]):
|
|
296
|
+
print("\n".join(explain(engine, "exposed", (service, cve))))
|
|
297
|
+
```
|
|
298
|
+
|
|
299
|
+
The command-line interface installs too, as `tiny-datalog` (and
|
|
300
|
+
`tiny-datalog-semiring`, `-tabling`, `-incremental`, `-subsumption`,
|
|
301
|
+
`-containment`, `-prolog` for the satellites):
|
|
302
|
+
|
|
303
|
+
```sh
|
|
304
|
+
tiny-datalog -q 'exposed(S, C)' supply-chain.dl
|
|
305
|
+
```
|
|
306
|
+
|
|
307
|
+
The course material — lessons, programs, exercises — is not part of the
|
|
308
|
+
installed package; clone the repository for that.
|
|
309
|
+
|
|
310
|
+
## Layout
|
|
311
|
+
|
|
312
|
+
```
|
|
313
|
+
tiny_datalog/ the engine and its satellites — the code you read:
|
|
314
|
+
datalog.py the core: AST, parser, safety checks, stratification,
|
|
315
|
+
the semi-naive evaluator, and the CLI
|
|
316
|
+
magic.py the magic-sets rewriting (a program-to-program pass)
|
|
317
|
+
semantics.py grounding, stable models, the well-founded model
|
|
318
|
+
semiring.py semiring-valued evaluation (costs, counts, provenance,
|
|
319
|
+
probabilities)
|
|
320
|
+
incremental.py insertions + DRed deletions over a live materialisation
|
|
321
|
+
prolog.py top-down SLD resolution with function symbols
|
|
322
|
+
tabling.py tabled top-down evaluation (iterative QSQR)
|
|
323
|
+
subsumption.py KL-ONE-style EL classifier, compiled to Datalog
|
|
324
|
+
containment.py query containment and minimisation by homomorphism
|
|
325
|
+
*.py three-line launchers, so `python3 datalog.py ...` works
|
|
326
|
+
straight from a checkout with nothing installed
|
|
327
|
+
programs/ teaching programs, numbered by the lesson that uses
|
|
328
|
+
them (00-* are the README's examples)
|
|
329
|
+
lessons/ getting started, glossary, and lessons 0–18
|
|
330
|
+
exercises/ worked answers, verified by the test suite
|
|
331
|
+
cases/ golden test cases — add one without writing Python
|
|
332
|
+
benchmarks/ scaled input generators (chain/tree/clique/grid)
|
|
333
|
+
tests.py 234 tests: every shipped program and exercise answer is
|
|
334
|
+
executed, a conformance suite runs every query through
|
|
335
|
+
every applicable strategy, and a seeded fuzzer checks
|
|
336
|
+
the same property on random programs
|
|
337
|
+
```
|
|
338
|
+
|
|
339
|
+
The code is part of the course: comments explain the algorithms as
|
|
340
|
+
they happen, and the lessons that introduce machinery end with an
|
|
341
|
+
*Under the hood* section reading the piece of the implementation they
|
|
342
|
+
used.
|
|
343
|
+
|
|
344
|
+
### How big is it, honestly
|
|
345
|
+
|
|
346
|
+
The evaluator is about 850 lines (`tiny_datalog/datalog.py`, up to the
|
|
347
|
+
command-line interface), the CLI, `--explain` and why-not another 600, and the eight satellite modules about
|
|
348
|
+
2,400. Call it 3.9k lines of toolkit and 2.6k of tests, roughly a
|
|
349
|
+
quarter of it commentary.
|
|
350
|
+
|
|
351
|
+
"Tiny" is a claim about the evaluator, and about each satellite module
|
|
352
|
+
singly: none of the eight exceeds 500 lines, which a test asserts. It is not a claim about
|
|
353
|
+
the repository, which is nine modules because it teaches nine things.
|
|
354
|
+
|
|
355
|
+
There is no dead code to golf away (checked); shrinking further means
|
|
356
|
+
deleting either a technique or an explanation.
|
|
357
|
+
|
|
358
|
+
## License
|
|
359
|
+
|
|
360
|
+
MIT.
|