tiny-datalog 0.1.2__tar.gz → 0.3.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (21) hide show
  1. {tiny_datalog-0.1.2/tiny_datalog.egg-info → tiny_datalog-0.3.0}/PKG-INFO +43 -18
  2. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/README.md +41 -16
  3. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/pyproject.toml +3 -2
  4. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/tiny_datalog/__init__.py +3 -2
  5. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/tiny_datalog/datalog.py +61 -3
  6. tiny_datalog-0.3.0/tiny_datalog/defeasible.py +462 -0
  7. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/tiny_datalog/tabling.py +75 -31
  8. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0/tiny_datalog.egg-info}/PKG-INFO +43 -18
  9. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/tiny_datalog.egg-info/SOURCES.txt +1 -0
  10. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/tiny_datalog.egg-info/entry_points.txt +1 -0
  11. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/LICENSE +0 -0
  12. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/setup.cfg +0 -0
  13. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/tiny_datalog/containment.py +0 -0
  14. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/tiny_datalog/incremental.py +0 -0
  15. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/tiny_datalog/magic.py +0 -0
  16. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/tiny_datalog/prolog.py +0 -0
  17. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/tiny_datalog/semantics.py +0 -0
  18. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/tiny_datalog/semiring.py +0 -0
  19. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/tiny_datalog/subsumption.py +0 -0
  20. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/tiny_datalog.egg-info/dependency_links.txt +0 -0
  21. {tiny_datalog-0.1.2 → tiny_datalog-0.3.0}/tiny_datalog.egg-info/top_level.txt +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: tiny-datalog
3
- Version: 0.1.2
3
+ Version: 0.3.0
4
4
  Summary: A Datalog engine small enough to read in an afternoon
5
5
  Author: Andrew Goodchild
6
6
  License-Expression: MIT
@@ -9,7 +9,7 @@ Project-URL: Source, https://github.com/andrewgoodchild/tiny-datalog
9
9
  Project-URL: Issues, https://github.com/andrewgoodchild/tiny-datalog/issues
10
10
  Project-URL: Changelog, https://github.com/andrewgoodchild/tiny-datalog/releases
11
11
  Project-URL: Lessons, https://github.com/andrewgoodchild/tiny-datalog/tree/main/lessons
12
- Keywords: datalog,logic programming,deductive database,semi-naive evaluation,magic sets,provenance,stable models,teaching
12
+ Keywords: datalog,logic programming,deductive database,semi-naive evaluation,magic sets,provenance,stable models,defeasible logic,teaching
13
13
  Classifier: Development Status :: 3 - Alpha
14
14
  Classifier: Intended Audience :: Developers
15
15
  Classifier: Intended Audience :: Education
@@ -126,7 +126,7 @@ Nothing to install:
126
126
 
127
127
  ```sh
128
128
  git clone https://github.com/andrewgoodchild/tiny-datalog && cd tiny-datalog
129
- python3 tests.py # 235 tests, ~12s
129
+ python3 tests.py # 254 tests, ~12s
130
130
  ```
131
131
 
132
132
  ## Why the language choice decides what you can ask later
@@ -235,8 +235,10 @@ Answers are also checked against engines this repository did not write.
235
235
  [`conformance/`](https://github.com/andrewgoodchild/tiny-datalog/tree/main/conformance)
236
236
  runs the [datalog-conformance](https://pypi.org/project/datalog-conformance/)
237
237
  corpus, harvested from Soufflé, Nemo and Crepe, through all four
238
- strategies: 89 cases pass. The 25 it skips need arithmetic or 10⁵-fact
239
- joins, and both are omissions this course makes on purpose.
238
+ strategies, and its defeasible theories, from SPINdle and the papers,
239
+ through `defeasible.py`: 195 cases pass. Every skip is named with its
240
+ reason: arithmetic and 10⁵-fact joins, which this course omits on
241
+ purpose, or cases written for a different logic.
240
242
 
241
243
  ## Learning Datalog
242
244
 
@@ -246,7 +248,8 @@ the current research threads. The field's own recent lecture notes
246
248
  observe that the literature advises people building Datalog engines
247
249
  better than people trying to *use* one; this course does both halves
248
250
  on purpose — sixteen lessons where the engine is the explanation, then
249
- a lesson on authoring rules that survive review. And it is built to be
251
+ a lesson on authoring rules that survive review, and one on a logic
252
+ built for rules with exceptions. And it is built to be
250
253
  inherited: `git clone`, no dependencies, no hosted anything, and every
251
254
  quoted transcript re-verified by CI — the exercises cannot rot. (For where each
252
255
  technique ships — CodeQL, RDFox, Feldera, SNOMED and the rest —
@@ -286,7 +289,9 @@ regression test without writing Python.
286
289
 
287
290
  Not an engine to build a product on. Joins are nested-loop,
288
291
  stable-model search is exhaustive, evaluation is batch. For real
289
- workloads see Soufflé, clingo, RDFox or Feldera.
292
+ workloads see Soufflé, clingo, RDFox or Feldera; for Datalog inside a
293
+ Python application, with arithmetic and queries over SQL databases,
294
+ see [pyDatalog](https://pypi.org/project/pyDatalog/).
290
295
 
291
296
  Deliberate omissions, because saying why teaches more than lacking them
292
297
  quietly:
@@ -303,7 +308,7 @@ quietly:
303
308
  (disjointness and unsatisfiability detection included). SNOMED needs
304
309
  ELH (EL plus role hierarchies) with right identities, which is what
305
310
  the ELK and Snorocket reasoners implement and this does not.
306
- - **A REPL (interactive prompt) and packaging.** `git clone` and run.
311
+ - **A REPL (interactive prompt).** Run a file, or call it from Python.
307
312
 
308
313
  Aggregation used to be on this list;
309
314
  [lesson 13](https://github.com/andrewgoodchild/tiny-datalog/blob/main/lessons/13-aggregation.md) is what promoting an omission
@@ -321,13 +326,32 @@ pip install tiny-datalog
321
326
  from tiny_datalog import run_program, explain
322
327
 
323
328
  engine = run_program(open("supply-chain.dl").read())
324
- for service, cve in sorted(engine.rels["exposed"]):
325
- print("\n".join(explain(engine, "exposed", (service, cve))))
329
+ for answer in engine.query("exposed(S, C)"):
330
+ print(answer) # {'S': 'pkg0', 'C': 'cve_2026_0001'}
331
+ print("\n".join(explain(engine, "exposed", ("pkg0", "cve_2026_0001"))))
332
+ ```
333
+
334
+ Facts can come straight from Python data (a CSV file, a database
335
+ query) instead of program text. Each row is checked as if it had been
336
+ parsed:
337
+
338
+ ```python
339
+ rules = """
340
+ uses(X, Y) :- depends(X, Y).
341
+ uses(X, Z) :- depends(X, Y), uses(Y, Z).
342
+ exposed(S, C) :- service(S), uses(S, L), vulnerable(L, C).
343
+ """
344
+ engine = run_program(rules, facts={
345
+ "depends": [("app", "lib"), ("lib", "core")],
346
+ "service": ["app"], # one-column rows may be bare
347
+ "vulnerable": [("core", "cve_2026_0001")],
348
+ })
349
+ engine.query("exposed(S, C)") # [{'S': 'app', 'C': 'cve_2026_0001'}]
326
350
  ```
327
351
 
328
352
  The command-line interface installs too, as `tiny-datalog` (and
329
353
  `tiny-datalog-semiring`, `-tabling`, `-incremental`, `-subsumption`,
330
- `-containment`, `-prolog` for the satellites):
354
+ `-containment`, `-defeasible`, `-prolog` for the satellites):
331
355
 
332
356
  ```sh
333
357
  tiny-datalog -q 'exposed(S, C)' supply-chain.dl
@@ -351,17 +375,18 @@ tiny_datalog/ the engine and its satellites — the code you read:
351
375
  tabling.py tabled top-down evaluation (iterative QSQR)
352
376
  subsumption.py KL-ONE-style EL classifier, compiled to Datalog
353
377
  containment.py query containment and minimisation by homomorphism
378
+ defeasible.py defeasible logic: exceptions, priorities, defeaters
354
379
  *.py three-line launchers, so `python3 datalog.py ...` works
355
380
  straight from a checkout with nothing installed
356
381
  programs/ teaching programs, numbered by the lesson that uses
357
382
  them (00-* are the README's examples)
358
- lessons/ getting started, glossary, and lessons 0–18
383
+ lessons/ getting started, glossary, and lessons 0–19
359
384
  exercises/ worked answers, verified by the test suite
360
385
  cases/ golden test cases — add one without writing Python
361
386
  conformance/ the external datalog-conformance corpus (Soufflé, Nemo,
362
- Crepe), run against all four evaluation strategies
387
+ Crepe, SPINdle), run against every evaluation strategy
363
388
  benchmarks/ scaled input generators (chain/tree/clique/grid)
364
- tests.py 235 tests: every shipped program and exercise answer is
389
+ tests.py 254 tests: every shipped program and exercise answer is
365
390
  executed, a conformance suite runs every query through
366
391
  every applicable strategy, and a seeded fuzzer checks
367
392
  the same property on random programs
@@ -375,13 +400,13 @@ used.
375
400
  ### How big is it, honestly
376
401
 
377
402
  The evaluator is about 850 lines (`tiny_datalog/datalog.py`, up to the
378
- command-line interface), the CLI, `--explain` and why-not another 600, and the eight satellite modules about
379
- 2,400. Call it 3.9k lines of toolkit and 2.6k of tests, roughly a
403
+ command-line interface), the CLI, `--explain` and why-not another 600, and the nine satellite modules about
404
+ 2,900. Call it 4.3k lines of toolkit and 2.7k of tests, roughly a
380
405
  quarter of it commentary.
381
406
 
382
407
  "Tiny" is a claim about the evaluator, and about each satellite module
383
- singly: none of the eight exceeds 500 lines, which a test asserts. It is not a claim about
384
- the repository, which is nine modules because it teaches nine things.
408
+ singly: none of the nine exceeds 500 lines, which a test asserts. It is not a claim about
409
+ the repository, which is ten modules because it teaches ten things.
385
410
 
386
411
  There is no dead code to golf away (checked); shrinking further means
387
412
  deleting either a technique or an explanation.
@@ -93,7 +93,7 @@ Nothing to install:
93
93
 
94
94
  ```sh
95
95
  git clone https://github.com/andrewgoodchild/tiny-datalog && cd tiny-datalog
96
- python3 tests.py # 235 tests, ~12s
96
+ python3 tests.py # 254 tests, ~12s
97
97
  ```
98
98
 
99
99
  ## Why the language choice decides what you can ask later
@@ -202,8 +202,10 @@ Answers are also checked against engines this repository did not write.
202
202
  [`conformance/`](https://github.com/andrewgoodchild/tiny-datalog/tree/main/conformance)
203
203
  runs the [datalog-conformance](https://pypi.org/project/datalog-conformance/)
204
204
  corpus, harvested from Soufflé, Nemo and Crepe, through all four
205
- strategies: 89 cases pass. The 25 it skips need arithmetic or 10⁵-fact
206
- joins, and both are omissions this course makes on purpose.
205
+ strategies, and its defeasible theories, from SPINdle and the papers,
206
+ through `defeasible.py`: 195 cases pass. Every skip is named with its
207
+ reason: arithmetic and 10⁵-fact joins, which this course omits on
208
+ purpose, or cases written for a different logic.
207
209
 
208
210
  ## Learning Datalog
209
211
 
@@ -213,7 +215,8 @@ the current research threads. The field's own recent lecture notes
213
215
  observe that the literature advises people building Datalog engines
214
216
  better than people trying to *use* one; this course does both halves
215
217
  on purpose — sixteen lessons where the engine is the explanation, then
216
- a lesson on authoring rules that survive review. And it is built to be
218
+ a lesson on authoring rules that survive review, and one on a logic
219
+ built for rules with exceptions. And it is built to be
217
220
  inherited: `git clone`, no dependencies, no hosted anything, and every
218
221
  quoted transcript re-verified by CI — the exercises cannot rot. (For where each
219
222
  technique ships — CodeQL, RDFox, Feldera, SNOMED and the rest —
@@ -253,7 +256,9 @@ regression test without writing Python.
253
256
 
254
257
  Not an engine to build a product on. Joins are nested-loop,
255
258
  stable-model search is exhaustive, evaluation is batch. For real
256
- workloads see Soufflé, clingo, RDFox or Feldera.
259
+ workloads see Soufflé, clingo, RDFox or Feldera; for Datalog inside a
260
+ Python application, with arithmetic and queries over SQL databases,
261
+ see [pyDatalog](https://pypi.org/project/pyDatalog/).
257
262
 
258
263
  Deliberate omissions, because saying why teaches more than lacking them
259
264
  quietly:
@@ -270,7 +275,7 @@ quietly:
270
275
  (disjointness and unsatisfiability detection included). SNOMED needs
271
276
  ELH (EL plus role hierarchies) with right identities, which is what
272
277
  the ELK and Snorocket reasoners implement and this does not.
273
- - **A REPL (interactive prompt) and packaging.** `git clone` and run.
278
+ - **A REPL (interactive prompt).** Run a file, or call it from Python.
274
279
 
275
280
  Aggregation used to be on this list;
276
281
  [lesson 13](https://github.com/andrewgoodchild/tiny-datalog/blob/main/lessons/13-aggregation.md) is what promoting an omission
@@ -288,13 +293,32 @@ pip install tiny-datalog
288
293
  from tiny_datalog import run_program, explain
289
294
 
290
295
  engine = run_program(open("supply-chain.dl").read())
291
- for service, cve in sorted(engine.rels["exposed"]):
292
- print("\n".join(explain(engine, "exposed", (service, cve))))
296
+ for answer in engine.query("exposed(S, C)"):
297
+ print(answer) # {'S': 'pkg0', 'C': 'cve_2026_0001'}
298
+ print("\n".join(explain(engine, "exposed", ("pkg0", "cve_2026_0001"))))
299
+ ```
300
+
301
+ Facts can come straight from Python data (a CSV file, a database
302
+ query) instead of program text. Each row is checked as if it had been
303
+ parsed:
304
+
305
+ ```python
306
+ rules = """
307
+ uses(X, Y) :- depends(X, Y).
308
+ uses(X, Z) :- depends(X, Y), uses(Y, Z).
309
+ exposed(S, C) :- service(S), uses(S, L), vulnerable(L, C).
310
+ """
311
+ engine = run_program(rules, facts={
312
+ "depends": [("app", "lib"), ("lib", "core")],
313
+ "service": ["app"], # one-column rows may be bare
314
+ "vulnerable": [("core", "cve_2026_0001")],
315
+ })
316
+ engine.query("exposed(S, C)") # [{'S': 'app', 'C': 'cve_2026_0001'}]
293
317
  ```
294
318
 
295
319
  The command-line interface installs too, as `tiny-datalog` (and
296
320
  `tiny-datalog-semiring`, `-tabling`, `-incremental`, `-subsumption`,
297
- `-containment`, `-prolog` for the satellites):
321
+ `-containment`, `-defeasible`, `-prolog` for the satellites):
298
322
 
299
323
  ```sh
300
324
  tiny-datalog -q 'exposed(S, C)' supply-chain.dl
@@ -318,17 +342,18 @@ tiny_datalog/ the engine and its satellites — the code you read:
318
342
  tabling.py tabled top-down evaluation (iterative QSQR)
319
343
  subsumption.py KL-ONE-style EL classifier, compiled to Datalog
320
344
  containment.py query containment and minimisation by homomorphism
345
+ defeasible.py defeasible logic: exceptions, priorities, defeaters
321
346
  *.py three-line launchers, so `python3 datalog.py ...` works
322
347
  straight from a checkout with nothing installed
323
348
  programs/ teaching programs, numbered by the lesson that uses
324
349
  them (00-* are the README's examples)
325
- lessons/ getting started, glossary, and lessons 0–18
350
+ lessons/ getting started, glossary, and lessons 0–19
326
351
  exercises/ worked answers, verified by the test suite
327
352
  cases/ golden test cases — add one without writing Python
328
353
  conformance/ the external datalog-conformance corpus (Soufflé, Nemo,
329
- Crepe), run against all four evaluation strategies
354
+ Crepe, SPINdle), run against every evaluation strategy
330
355
  benchmarks/ scaled input generators (chain/tree/clique/grid)
331
- tests.py 235 tests: every shipped program and exercise answer is
356
+ tests.py 254 tests: every shipped program and exercise answer is
332
357
  executed, a conformance suite runs every query through
333
358
  every applicable strategy, and a seeded fuzzer checks
334
359
  the same property on random programs
@@ -342,13 +367,13 @@ used.
342
367
  ### How big is it, honestly
343
368
 
344
369
  The evaluator is about 850 lines (`tiny_datalog/datalog.py`, up to the
345
- command-line interface), the CLI, `--explain` and why-not another 600, and the eight satellite modules about
346
- 2,400. Call it 3.9k lines of toolkit and 2.6k of tests, roughly a
370
+ command-line interface), the CLI, `--explain` and why-not another 600, and the nine satellite modules about
371
+ 2,900. Call it 4.3k lines of toolkit and 2.7k of tests, roughly a
347
372
  quarter of it commentary.
348
373
 
349
374
  "Tiny" is a claim about the evaluator, and about each satellite module
350
- singly: none of the eight exceeds 500 lines, which a test asserts. It is not a claim about
351
- the repository, which is nine modules because it teaches nine things.
375
+ singly: none of the nine exceeds 500 lines, which a test asserts. It is not a claim about
376
+ the repository, which is ten modules because it teaches ten things.
352
377
 
353
378
  There is no dead code to golf away (checked); shrinking further means
354
379
  deleting either a technique or an explanation.
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
4
4
 
5
5
  [project]
6
6
  name = "tiny-datalog"
7
- version = "0.1.2"
7
+ version = "0.3.0"
8
8
  description = "A Datalog engine small enough to read in an afternoon"
9
9
  readme = "README.md"
10
10
  license = "MIT"
@@ -14,7 +14,7 @@ authors = [{name = "Andrew Goodchild"}]
14
14
  dependencies = []
15
15
  keywords = ["datalog", "logic programming", "deductive database",
16
16
  "semi-naive evaluation", "magic sets", "provenance",
17
- "stable models", "teaching"]
17
+ "stable models", "defeasible logic", "teaching"]
18
18
  # No "License ::" classifier: `license` above is an SPDX expression, and
19
19
  # setuptools refuses a build that declares the licence both ways.
20
20
  classifiers = [
@@ -51,6 +51,7 @@ tiny-datalog-tabling = "tiny_datalog.tabling:main"
51
51
  tiny-datalog-incremental = "tiny_datalog.incremental:main"
52
52
  tiny-datalog-subsumption = "tiny_datalog.subsumption:main"
53
53
  tiny-datalog-containment = "tiny_datalog.containment:main"
54
+ tiny-datalog-defeasible = "tiny_datalog.defeasible:main"
54
55
 
55
56
  [tool.setuptools.packages.find]
56
57
  include = ["tiny_datalog*"]
@@ -10,7 +10,8 @@ Everything else in the package is a satellite built on top of it:
10
10
  models), `semiring` (provenance-weighted evaluation), `incremental`
11
11
  (maintenance under updates), `tabling` (top-down with memoing),
12
12
  `subsumption` (KL-ONE style classification), `containment` (query
13
- containment), and `prolog` (a Prolog reader for the same syntax).
13
+ containment), `defeasible` (rules with exceptions and priorities), and
14
+ `prolog` (a Prolog reader for the same syntax).
14
15
  Import those by name: `from tiny_datalog import semiring`.
15
16
  """
16
17
 
@@ -29,7 +30,7 @@ from tiny_datalog.datalog import (
29
30
  format_atom, format_fact,
30
31
  )
31
32
 
32
- __version__ = "0.1.2"
33
+ __version__ = "0.3.0"
33
34
 
34
35
  __all__ = [
35
36
  "Var", "Const", "Struct", "Atom", "Literal", "Rule",
@@ -821,6 +821,27 @@ class Engine:
821
821
  groups[key].append(s[var.name])
822
822
  return groups
823
823
 
824
+ def query(self, goal):
825
+ """The answers to a query, one dict per answer, in sorted order:
826
+
827
+ engine.query("exposed(S, C)")
828
+ -> [{"S": "pkg0", "C": "cve_2026_0001"}, ...]
829
+
830
+ `goal` is query text or an Atom. Anonymous variables are left
831
+ out of the answers; a query with no variables answers [{}] if it
832
+ holds and [] if it does not."""
833
+ atom = parse_goal(goal) if isinstance(goal, str) else goal
834
+ check_query_atom(atom, self.program.arity)
835
+ names = []
836
+ for a in atom.args:
837
+ if isinstance(a, Var) and not a.anonymous and a.name not in names:
838
+ names.append(a.name)
839
+ rows = {tuple(s[n] for n in names)
840
+ for s in (_match(atom.args, tup, {})
841
+ for tup in self.rels.get(atom.pred, ()))
842
+ if s is not None}
843
+ return [dict(zip(names, row)) for row in sorted(rows, key=_sort_key)]
844
+
824
845
  @staticmethod
825
846
  def _instantiate(atom, subst):
826
847
  return tuple(a.value if isinstance(a, Const) else subst[a.name]
@@ -851,13 +872,50 @@ def _fold(func, values, rule):
851
872
  return agg
852
873
 
853
874
 
854
- def run_program(text):
855
- """Parse, stratify, and evaluate a program; return the Engine."""
856
- engine = Engine(Program(parse(text)))
875
+ def run_program(text, facts=None):
876
+ """Parse, stratify, and evaluate a program; return the Engine.
877
+
878
+ `facts` adds base facts from Python data, so rows from a CSV file or
879
+ a database query can feed the rules without first being written out
880
+ as Datalog text:
881
+
882
+ run_program(rules, facts={"depends": [("pkg4", "pkg13")],
883
+ "service": ["pkg4"]})
884
+
885
+ Each row is a tuple of str, int or float (a bare value is a one-
886
+ column row), and every fact is checked exactly as if it had been
887
+ parsed: arity, groundness, and agreement with the rules."""
888
+ engine = Engine(Program(parse(text) + python_facts(facts or {})))
857
889
  engine.run()
858
890
  return engine
859
891
 
860
892
 
893
+ _PREDICATE = re.compile(r"[a-z][A-Za-z0-9_]*\Z")
894
+
895
+
896
+ def python_facts(facts):
897
+ """{predicate: rows} as fact clauses (see run_program)."""
898
+ clauses = []
899
+ for pred, rows in facts.items():
900
+ if not isinstance(pred, str) or not _PREDICATE.match(pred) \
901
+ or pred == "not":
902
+ raise DatalogError("%r cannot name a predicate: a predicate is "
903
+ "a lowercase identifier" % (pred,))
904
+ for row in rows:
905
+ if not isinstance(row, (tuple, list)):
906
+ row = (row,)
907
+ for v in row:
908
+ # bool is an int in Python, and would print as a constant
909
+ # named True; nan and inf print as constants too
910
+ if (isinstance(v, bool) or not isinstance(v, (str, int, float))
911
+ or (isinstance(v, float) and not math.isfinite(v))):
912
+ raise DatalogError(
913
+ "fact %s%r: values must be str, int or finite float, "
914
+ "not %r" % (pred, tuple(row), v))
915
+ clauses.append(Rule(Atom(pred, tuple(Const(v) for v in row)), ()))
916
+ return clauses
917
+
918
+
861
919
  # ---------------------------------------------------------------------------
862
920
  # CLI
863
921
  # ---------------------------------------------------------------------------
@@ -0,0 +1,462 @@
1
+ #!/usr/bin/env python3
2
+ """
3
+ defeasible.py — defeasible logic: rules with exceptions and priorities.
4
+ (Lesson 18, which ends with a tour of this module.)
5
+
6
+ Lesson 3 met defaults through `not`: birds fly unless abnormal. That
7
+ works for one exception, written by hand into the rule it overrides.
8
+ Real rule sets — statutes, policies, eligibility criteria — are made of
9
+ generalisations, exceptions to them, exceptions to the exceptions, and
10
+ an order saying which wins. Defeasible logic (Nute, 1994) makes all of
11
+ that first-class:
12
+
13
+ r1: penguin(X) -> bird(X). % strict: no exceptions, ever
14
+ r2: bird(X) => flies(X). % defeasible: usually
15
+ r3: penguin(X) => ~flies(X). % a rule for the opposite
16
+ r4: injured(X) ~> ~flies(X). % a defeater: can only block
17
+ r3 > r2. % superiority: r3 beats r2
18
+
19
+ `~` is *strong* negation — `~flies(opus)` is a claim, not an absence —
20
+ and a conflict is exactly a pair of complementary literals. There is
21
+ no `not` here at all: exceptions are rules that attack, priorities say
22
+ who wins, and an unresolved conflict yields neither side instead of
23
+ whichever rule happened to have the `not` written into it.
24
+
25
+ Conclusions carry four tags (Antoniou, Billington, Governatori & Maher
26
+ 2001):
27
+
28
+ +Δ q q is definitely provable: facts and strict rules alone
29
+ −Δ q q is demonstrably not definitely provable
30
+ +∂ q q is defeasibly provable: some applicable rule supports q,
31
+ ~q is not definite, and every rule for ~q is either
32
+ inapplicable or beaten by an applicable rule for q
33
+ −∂ q q is demonstrably not defeasibly provable
34
+
35
+ "Demonstrably": a −tag is a finite proof of failure, not the mere
36
+ absence of a +tag. On a loop like `a => b. b => a.` neither +∂ b nor
37
+ −∂ b is ever established, and the module reports b as *undecided* —
38
+ where Lesson 5's well-founded semantics would settle it false, as does
39
+ the well-founded variant of this logic (Maher & Governatori 1999).
40
+ Standard defeasible logic is characterised instead by Kunen's
41
+ three-valued semantics of a logic program that encodes the theory —
42
+ which is why it cannot see that such a loop has no foundation.
43
+
44
+ The proof conditions are implemented one-to-one in `_conclude`, over
45
+ the theory's grounding, as a fixpoint: tags are only ever added, so
46
+ iteration stops. Grounding reuses the core engine. Every rule, read
47
+ as if strict, is a positive Datalog rule, and its least model bounds
48
+ everything a +tag could be about; see `Theory.ground` for the one
49
+ consequence of grounding that way. `--propagating` switches from
50
+ ambiguity blocking to ambiguity propagation (`programs/ambiguity.dfl`
51
+ shows the difference).
52
+ """
53
+
54
+ import argparse
55
+ import re
56
+ import sys
57
+ from collections import defaultdict
58
+ from itertools import combinations
59
+
60
+ from tiny_datalog.datalog import (
61
+ Atom, DatalogError, Engine, Literal, ParseError, Program, Rule,
62
+ SafetyError, Var, format_atom, match_answers, parse_goal, read_program,
63
+ _sort_key)
64
+
65
+ STRICT, DEFEASIBLE, DEFEATER = "->", "=>", "~>"
66
+
67
+ _TOKEN = re.compile(r"""
68
+ (?P<comment>%[^\n]*) | (?P<space>\s+)
69
+ | (?P<string>"[^"\n]*"|'[^'\n]*')
70
+ | (?P<number>-?[0-9]+(?:\.[0-9]+)?(?:[eE][-+]?[0-9]+)?)
71
+ | (?P<arrow>->|=>|~>) | (?P<punct>[:,.()>~])
72
+ | (?P<name>[A-Za-z_][A-Za-z0-9_]*)
73
+ """, re.VERBOSE)
74
+
75
+
76
+ def complement(lit):
77
+ """`p` <-> `~p`. A literal is (pred, args); strong negation lives
78
+ in the predicate name, so the core engine stores ~p as a relation
79
+ of its own and never needs to know it is special."""
80
+ pred, args = lit
81
+ return (pred[1:] if pred.startswith("~") else "~" + pred), args
82
+
83
+
84
+ def show(lit):
85
+ return format_atom(*lit)
86
+
87
+
88
+ class Theory:
89
+ """A parsed defeasible theory: facts, labelled rules, superiority."""
90
+
91
+ def __init__(self):
92
+ self.facts = [] # core Atoms; a ~ fact has pred "~p"
93
+ self.rules = [] # (label, kind, head Atom, body Atoms)
94
+ self.superior = set() # (stronger label, weaker label)
95
+
96
+ # -- reading ------------------------------------------------------------
97
+
98
+ @classmethod
99
+ def parse(cls, text):
100
+ theory = cls()
101
+ tokens = _tokens(text)
102
+ statement = []
103
+ for tok in tokens:
104
+ statement.append(tok)
105
+ if tok[1] == "." and _depth(statement) == 0:
106
+ theory._statement(statement[:-1], tok[2])
107
+ statement = []
108
+ if statement:
109
+ raise ParseError("line %d: expected '.', got end of input"
110
+ % statement[-1][2])
111
+ theory._check()
112
+ return theory
113
+
114
+ def _statement(self, toks, line):
115
+ text = [t[1] for t in toks]
116
+ if len(toks) == 3 and text[1] == ">": # r3 > r2.
117
+ self.superior.add((text[0], text[2]))
118
+ return
119
+ label = None
120
+ if len(toks) > 1 and toks[0][0] == "name" and text[1] == ":":
121
+ label, toks = text[0], toks[2:]
122
+ arrows = [i for i, t in enumerate(toks) if t[0] == "arrow"]
123
+ if not arrows:
124
+ if label is not None:
125
+ raise ParseError("line %d: a labelled statement must be a "
126
+ "rule (->, => or ~>)" % line)
127
+ self.facts.append(_atom(toks, line))
128
+ return
129
+ if len(arrows) > 1:
130
+ raise ParseError("line %d: one arrow per rule" % line)
131
+ i = arrows[0]
132
+ body = [_atom(part, line) for part in _split(toks[:i])]
133
+ head = _atom(toks[i + 1:], line)
134
+ self.rules.append((label or "_r%d" % (len(self.rules) + 1),
135
+ toks[i][1], head, tuple(_fresh_anonymous(body))))
136
+
137
+ def _check(self):
138
+ labels = [r[0] for r in self.rules]
139
+ dup = {l for l in labels if labels.count(l) > 1}
140
+ if dup:
141
+ raise ParseError("rule label used twice: %s" % ", ".join(sorted(dup)))
142
+ for pair in self.superior:
143
+ for l in pair:
144
+ if l not in labels:
145
+ raise ParseError("superiority names unknown rule %s" % l)
146
+ # the superiority relation must be acyclic, or "r beats s"
147
+ # could end up meaning r beats itself
148
+ beats = defaultdict(set)
149
+ for a, b in self.superior:
150
+ beats[a].add(b)
151
+ for start in list(beats):
152
+ stack, seen = list(beats[start]), set()
153
+ while stack:
154
+ x = stack.pop()
155
+ if x == start:
156
+ raise SafetyError("superiority is cyclic through %s"
157
+ % start)
158
+ if x not in seen:
159
+ seen.add(x)
160
+ stack.extend(beats[x])
161
+
162
+ # -- grounding ----------------------------------------------------------
163
+
164
+ def ground(self):
165
+ """Ground rules (label, kind, head, body) over literals, and the
166
+ set of facts.
167
+
168
+ Every rule is first read as a strict positive Datalog rule; the
169
+ core engine's least model of that program — the *envelope* —
170
+ holds every literal any rule could ever establish. A rule
171
+ instance is kept when its variables can all be bound by body
172
+ literals inside the envelope; the body literals it does not
173
+ match there simply fail. So `p, q => r` with q underivable is
174
+ kept, and r is refuted (−∂) as the proof theory says, while an
175
+ instance no derivable literal could bind at all is never made.
176
+
177
+ That is the one place this departs from the proof theory, which
178
+ ranges over every constant. A loop nothing starts, written with
179
+ variables — `a(X) => b(X). b(X) => a(X).` — is not reported at
180
+ all, where the proof theory would leave each instance undecided.
181
+ Written without variables it needs no binding, is kept, and
182
+ comes out undecided (`programs/loops.dfl`)."""
183
+ rules = [Rule(head, tuple(Literal(b, False) for b in body))
184
+ for _l, _k, head, body in self.rules]
185
+ program = Program([Rule(a, ()) for a in self.facts] + rules)
186
+ arities = defaultdict(set)
187
+ for pred, n in program.arity.items():
188
+ arities[pred.lstrip("~")].add(n)
189
+ for pred, ns in sorted(arities.items()):
190
+ if len(ns) > 1:
191
+ raise SafetyError("predicate %s used with arities %s (its "
192
+ "negation counts too)"
193
+ % (pred, " and ".join(map(str, sorted(ns)))))
194
+ engine = Engine(program)
195
+ engine.run()
196
+ facts = {(a.pred, tuple(x.value for x in a.args)) for a in self.facts}
197
+ ground, seen = [], set()
198
+ for label, kind, head, body in self.rules:
199
+ needed = _variables(body)
200
+ # match every subset of the body that binds all variables
201
+ for n in range(len(body) + 1):
202
+ for part in combinations(body, n):
203
+ if _variables(part) != needed:
204
+ continue
205
+ sub = Rule(head, tuple(Literal(b, False) for b in part))
206
+ for subst in engine._rule_substitutions(sub):
207
+ g = (label, kind,
208
+ (head.pred, engine._instantiate(head, subst)),
209
+ tuple((b.pred, engine._instantiate(b, subst))
210
+ for b in body))
211
+ if g not in seen:
212
+ seen.add(g)
213
+ ground.append(g)
214
+ return facts, ground
215
+
216
+ # -- the proof theory ---------------------------------------------------
217
+
218
+ def conclusions(self, policy="blocking"):
219
+ """{tag: set of literals} for the tags +Δ, −Δ, +∂, −∂, and
220
+ 'undecided' for literals that got neither +∂ nor −∂."""
221
+ if policy not in ("blocking", "propagating"):
222
+ raise DatalogError("unknown policy %r: use 'blocking' or "
223
+ "'propagating'" % (policy,))
224
+ facts, ground = self.ground()
225
+ return _conclude(facts, ground, self.superior, policy)
226
+
227
+
228
+ def _tokens(text):
229
+ """(kind, text, line) triples; comments and whitespace dropped."""
230
+ tokens, pos = [], 0
231
+ while pos < len(text):
232
+ m = _TOKEN.match(text, pos)
233
+ if m is None:
234
+ raise ParseError("line %d: unexpected character %r"
235
+ % (text.count("\n", 0, pos) + 1, text[pos]))
236
+ if m.lastgroup not in ("space", "comment"):
237
+ tokens.append((m.lastgroup, m.group(),
238
+ text.count("\n", 0, pos) + 1))
239
+ pos = m.end()
240
+ return tokens
241
+
242
+
243
+ def _variables(atoms):
244
+ return {a.name for atom in atoms for a in atom.args if isinstance(a, Var)}
245
+
246
+
247
+ def _depth(toks):
248
+ return sum({"(": 1, ")": -1}.get(t[1], 0) for t in toks)
249
+
250
+
251
+ def _split(toks):
252
+ """Comma-separated parts at bracket depth 0."""
253
+ parts, cur, depth = [], [], 0
254
+ for t in toks:
255
+ depth += {"(": 1, ")": -1}.get(t[1], 0)
256
+ if t[1] == "," and depth == 0:
257
+ parts.append(cur)
258
+ cur = []
259
+ else:
260
+ cur.append(t)
261
+ return parts + [cur] if cur else parts
262
+
263
+
264
+ def _atom(toks, line):
265
+ """`[~] atom` -> core Atom, pred prefixed '~' when negated. The atom
266
+ itself goes through the core parser, so constants, strings,
267
+ variables and its error messages are the ones the course knows."""
268
+ if not toks:
269
+ raise ParseError("line %d: expected a literal" % line)
270
+ neg = toks[0][1] == "~"
271
+ toks = toks[1:] if neg else toks
272
+ if not toks or toks[0][0] != "name" or toks[0][1][0].isupper():
273
+ raise ParseError("line %d: expected a literal, got %r"
274
+ % (line, " ".join(t[1] for t in toks) or "nothing"))
275
+ atom = parse_goal(" ".join(t[1] for t in toks))
276
+ return Atom(("~" if neg else "") + atom.pred, atom.args)
277
+
278
+
279
+ def _fresh_anonymous(atoms):
280
+ """Each literal was parsed on its own, so each numbered its `_`s from
281
+ 1; renumber so two `_` in one rule stay two different variables."""
282
+ n = 0
283
+ out = []
284
+ for a in atoms:
285
+ args = []
286
+ for x in a.args:
287
+ if isinstance(x, Var) and x.anonymous:
288
+ n += 1
289
+ x = Var("_#%d" % n)
290
+ args.append(x)
291
+ out.append(Atom(a.pred, tuple(args)))
292
+ return out
293
+
294
+
295
+ def _conclude(facts, ground, superior, policy):
296
+ """The proof conditions, as a fixpoint (Antoniou, Billington,
297
+ Governatori & Maher 2001). For a literal q, R[q] is every rule for
298
+ q, Rs[q] the strict ones, Rsd[q] strict or defeasible — defeaters
299
+ may attack, never support — and A(r) is rule r's body:
300
+
301
+ +Δq q is a fact, or some r in Rs[q] has A(r) all +Δ.
302
+ −Δq q is not a fact, and every r in Rs[q] has some a in A(r) −Δ.
303
+ +∂q +Δq; or (1) some r in Rsd[q] has A(r) all +∂, (2) ~q is −Δ,
304
+ and (3) every attacker s in R[~q] is either dead — some a in
305
+ A(s) is −∂ — or beaten: some t in Rsd[q] with A(t) all +∂
306
+ and t > s. The t may differ per attacker: team defeat.
307
+ −∂q −Δq, and: every r in Rsd[q] has some a in A(r) −∂; or ~q is
308
+ +Δ; or some attacker s in R[~q] is live — A(s) all +∂ — and
309
+ every t in Rsd[q] has a body literal −∂ or is not above s.
310
+
311
+ That is ambiguity *blocking*: an attacker only counts once its body
312
+ is proved. Under ambiguity *propagation* (Maher 2012) the same
313
+ conditions hold with attackers judged by mere *support*, σ:
314
+
315
+ +σq +Δq, or some r in Rsd[q] has A(r) all +σ, and no attacker
316
+ s in R[~q] with no −∂ body literal is above r.
317
+ −σq −Δq, and every r in Rsd[q] has a body literal −σ, or is
318
+ below some attacker s in R[~q] with A(s) all +∂.
319
+
320
+ and in +∂ an attacker is dead only if a body literal is −σ, while in
321
+ −∂ it is live as soon as A(s) is all +σ. Supported-but-unproved
322
+ attackers then still block, so doubt spreads downstream."""
323
+ by_head = defaultdict(list)
324
+ for rule in ground:
325
+ by_head[rule[2]].append(rule)
326
+ # report on every literal the theory mentions; decide its complement
327
+ # too, since every +∂ / −∂ condition looks at ~q
328
+ mentioned = set(facts) | set(by_head) | {a for r in ground for a in r[3]}
329
+ universe = mentioned | {complement(q) for q in mentioned}
330
+
331
+ def strict(q):
332
+ return [r for r in by_head[q] if r[1] == STRICT]
333
+
334
+ def supporting(q):
335
+ return [r for r in by_head[q] if r[1] != DEFEATER]
336
+
337
+ def body_in(r, tagged):
338
+ return all(a in tagged for a in r[3])
339
+
340
+ def body_hits(r, tagged):
341
+ return any(a in tagged for a in r[3])
342
+
343
+ def above(t, s):
344
+ return (t[0], s[0]) in superior
345
+
346
+ plus_d, minus_d = set(), set() # +Δ, −Δ: the strict part first
347
+ changed = True
348
+ while changed:
349
+ changed = False
350
+ for q in universe:
351
+ if q not in plus_d and (q in facts or any(
352
+ body_in(r, plus_d) for r in strict(q))):
353
+ plus_d.add(q)
354
+ changed = True
355
+ if q not in minus_d and q not in facts and all(
356
+ body_hits(r, minus_d) for r in strict(q)):
357
+ minus_d.add(q)
358
+ changed = True
359
+
360
+ plus, minus = set(), set() # +∂, −∂
361
+ if policy == "propagating":
362
+ s_plus, s_minus = set(), set() # +σ, −σ: who is merely supported
363
+ else:
364
+ s_plus, s_minus = plus, minus # blocking: support = proof
365
+ changed = True
366
+ while changed:
367
+ changed = False
368
+ for q in universe:
369
+ nq = complement(q)
370
+ attackers = by_head[nq]
371
+ if q not in plus and (q in plus_d or (
372
+ any(body_in(r, plus) for r in supporting(q))
373
+ and nq in minus_d
374
+ and all(body_hits(s, s_minus) or any(
375
+ body_in(t, plus) and above(t, s)
376
+ for t in supporting(q)) for s in attackers))):
377
+ plus.add(q)
378
+ changed = True
379
+ if q not in minus and q in minus_d and (
380
+ all(body_hits(r, minus) for r in supporting(q))
381
+ or nq in plus_d
382
+ or any(body_in(s, s_plus) and all(
383
+ body_hits(t, minus) or not above(t, s)
384
+ for t in supporting(q)) for s in attackers)):
385
+ minus.add(q)
386
+ changed = True
387
+ if policy != "propagating":
388
+ continue
389
+ if q not in s_plus and (q in plus_d or any(
390
+ body_in(r, s_plus) and all(
391
+ body_hits(s, minus) or not above(s, r)
392
+ for s in attackers)
393
+ for r in supporting(q))):
394
+ s_plus.add(q)
395
+ changed = True
396
+ if q not in s_minus and q in minus_d and all(
397
+ body_hits(r, s_minus) or any(
398
+ body_in(s, plus) and above(s, r) for s in attackers)
399
+ for r in supporting(q)):
400
+ s_minus.add(q)
401
+ changed = True
402
+ return {"+Δ": plus_d & mentioned, "−Δ": minus_d & mentioned,
403
+ "+∂": plus & mentioned, "−∂": minus & mentioned,
404
+ "undecided": mentioned - plus - minus}
405
+
406
+
407
+ def load(text):
408
+ return Theory.parse(text)
409
+
410
+
411
+ TAG_NAMES = [("+Δ", "definitely"), ("+∂", "defeasibly"),
412
+ ("−∂", "not defeasibly"), ("undecided", "undecided")]
413
+
414
+
415
+ def main(argv=None):
416
+ ap = argparse.ArgumentParser(
417
+ description="Defeasible logic: strict rules (->), defeasible rules "
418
+ "(=>), defeaters (~>), superiority (r1 > r2).")
419
+ ap.add_argument("file", help="defeasible theory")
420
+ ap.add_argument("-q", "--query", action="append", default=[],
421
+ metavar="LITERAL",
422
+ help="show the tags of matching literals (repeatable), "
423
+ "e.g. -q 'flies(X)' or -q '~flies(X)'")
424
+ ap.add_argument("--propagating", action="store_true",
425
+ help="ambiguity propagating instead of blocking")
426
+ args = ap.parse_args(argv)
427
+ try:
428
+ theory = load(read_program(args.file))
429
+ result = theory.conclusions(
430
+ "propagating" if args.propagating else "blocking")
431
+ queries = [_atom(_tokens(q.rstrip(". ")), 1) for q in args.query]
432
+ except DatalogError as exc:
433
+ print("error: %s" % exc, file=sys.stderr)
434
+ return 1
435
+
436
+ def ordered(lits):
437
+ return sorted(lits, key=lambda l: (l[0].lstrip("~"), l[0],
438
+ _sort_key(l[1])))
439
+
440
+ if queries:
441
+ for q in queries:
442
+ print("?- %s" % (q.pred if not q.args else "%s(%s)" % (
443
+ q.pred, ", ".join(str(a) for a in q.args))))
444
+ hits = [l for l in ordered(set().union(*result.values()))
445
+ if l[0] == q.pred and match_answers(q, [l[1]])]
446
+ for lit in hits:
447
+ tags = [t for t, _n in TAG_NAMES if lit in result[t]]
448
+ print(" %-28s %s" % (show(lit), " ".join(tags)))
449
+ if not hits:
450
+ print(" (no rule or fact mentions it)")
451
+ return 0
452
+ for tag, name in TAG_NAMES:
453
+ lits = ordered(result[tag])
454
+ print("%s %s (%d)" % (tag, name, len(lits)) if tag != "undecided"
455
+ else "%s (%d)" % (name, len(lits)))
456
+ for lit in lits:
457
+ print(" " + show(lit))
458
+ return 0
459
+
460
+
461
+ if __name__ == "__main__":
462
+ sys.exit(main())
@@ -30,9 +30,15 @@ the magic predicates' contents. They are the same sets — magic sets is
30
30
  tabling performed at compile time, tabling is magic sets performed at
31
31
  run time.
32
32
 
33
- Positive programs only (tabling under negation is SLG resolution, which
34
- computes the well-founded semantics — XSB's whole claim to fame — and
35
- is beyond this teaching module).
33
+ Negation is supported when the program is stratified, which is the
34
+ simplification pyDatalog also makes. A negated subgoal `not q(a)`
35
+ belongs to a lower stratum, so q's tables can be *completed* first —
36
+ grown to fixpoint on their own, since nothing in them depends on the
37
+ rule asking — and then looked up. Negation through recursion needs
38
+ full SLG resolution (Chen & Warren 1996), which suspends calls instead
39
+ of re-running them, detects when a group of tables is complete, and
40
+ *delays* negative literals it cannot yet decide; it computes the
41
+ well-founded semantics, and Lesson 15 says why it is not here.
36
42
 
37
43
  CLI
38
44
  ---
@@ -47,12 +53,14 @@ import sys
47
53
  from collections import defaultdict
48
54
 
49
55
  from tiny_datalog.datalog import (
50
- check_query_atom, Const, DatalogError, format_fact, parse, parse_goal,
51
- read_program, validate, _aggregate_of, _match, _sort_key)
56
+ check_query_atom, Const, DatalogError, StratificationError, format_fact,
57
+ parse, parse_goal, read_program, stratify, validate, _aggregate_of,
58
+ _match, _sort_key)
52
59
 
53
60
 
54
61
  class TabledEngine:
55
- """Iterative QSQR: tables keyed by call pattern, filled to fixpoint.
62
+ """Iterative QSQR: tables keyed by call pattern, filled to fixpoint;
63
+ a negated subgoal's tables are completed first, one stratum down.
56
64
 
57
65
  After query(), `tables` maps (pred, pattern) — pattern has a constant
58
66
  per bound argument and None per free one — to the set of full answer
@@ -65,18 +73,27 @@ class TabledEngine:
65
73
  raise DatalogError("retraction is incremental.py's job: %s" % c)
66
74
  if c.body and _aggregate_of(c.head):
67
75
  raise DatalogError(
68
- "tabled aggregation needs completion detection (real "
69
- "SLG); this module is positive-rules-only: %s" % c)
70
- for lit in c.body:
71
- if lit.negated:
72
- raise DatalogError(
73
- "tabling under negation is SLG resolution (the "
74
- "well-founded semantics, XSB) — beyond this "
75
- "module: %s" % c)
76
+ "tabled aggregation is not implemented here — use "
77
+ "datalog.py: %s" % c)
78
+ try:
79
+ self.strata = stratify(clauses)
80
+ except StratificationError as exc:
81
+ raise StratificationError(
82
+ "%s. Tabling under unstratified negation is full SLG "
83
+ "resolution, which computes the well-founded semantics "
84
+ "(Lesson 15); datalog.py --models computes it bottom-up."
85
+ % str(exc).rstrip("."), exc.cycle) from exc
76
86
  self.by_pred = defaultdict(list) # (pred, arity) -> clauses
77
87
  for c in clauses:
78
- self.by_pred[(c.head.pred, len(c.head.args))].append(c)
88
+ # positive literals first: they bind; negations then test
89
+ # ground atoms (safety guarantees they are ground by then)
90
+ body = tuple(sorted(c.body, key=lambda lit: lit.negated))
91
+ self.by_pred[(c.head.pred, len(c.head.args))].append(
92
+ (c.head, body))
79
93
  self.tables = {}
94
+ self.complete = set() # tables known to hold every answer
95
+ self.open = {} # ...and the rest, in creation order (a
96
+ # dict: a set's order would vary per run)
80
97
  self.rounds = 0
81
98
 
82
99
  # -- call patterns ------------------------------------------------------
@@ -98,8 +115,8 @@ class TabledEngine:
98
115
  def _table(self, pred, pattern):
99
116
  key = (pred, pattern)
100
117
  if key not in self.tables:
101
- self.tables[key] = set() # discovered a new subgoal
102
- self._grew = True # ...which the fixpoint must revisit
118
+ self.tables[key] = set() # discovered a new subgoal, which
119
+ self.open[key] = None # the fixpoint will now revisit
103
120
  return self.tables[key]
104
121
 
105
122
  # -- one round of top-down solving --------------------------------------
@@ -112,8 +129,17 @@ class TabledEngine:
112
129
  yield subst
113
130
  return
114
131
  lit, rest = body[0], body[1:]
115
- table = self._table(lit.atom.pred, self._pattern(lit.atom, subst))
116
- for ans in table:
132
+ key = (lit.atom.pred, self._pattern(lit.atom, subst))
133
+ table = self._table(*key)
134
+ if lit.negated:
135
+ # `not q(a)` may only read a finished table. q sits in a
136
+ # lower stratum, so finishing it cannot need this rule.
137
+ if key not in self.complete:
138
+ self._fixpoint(below=self.strata.get(lit.atom.pred, 0))
139
+ if key[1] not in table: # the pattern is the ground atom
140
+ yield from self._prove(rest, subst)
141
+ return
142
+ for ans in list(table): # a negation below may add to it
117
143
  s = _match(lit.atom.args, ans, subst)
118
144
  if s is not None:
119
145
  yield from self._prove(rest, s)
@@ -122,11 +148,11 @@ class TabledEngine:
122
148
  """Re-derive a subgoal's answers from its clauses, one step of
123
149
  head unification plus a tabled body proof."""
124
150
  pred, pattern = key
125
- for clause in self.by_pred.get((pred, len(pattern)), ()):
151
+ for head, body in self.by_pred.get((pred, len(pattern)), ()):
126
152
  # unify the head with the call pattern (bound args only)
127
153
  seed = {}
128
154
  ok = True
129
- for a, v in zip(clause.head.args, pattern):
155
+ for a, v in zip(head.args, pattern):
130
156
  if v is None:
131
157
  continue
132
158
  if isinstance(a, Const):
@@ -140,9 +166,9 @@ class TabledEngine:
140
166
  seed[a.name] = v
141
167
  if not ok:
142
168
  continue
143
- for s in self._prove(list(clause.body), seed):
169
+ for s in self._prove(body, seed):
144
170
  yield tuple(a.value if isinstance(a, Const) else s[a.name]
145
- for a in clause.head.args)
171
+ for a in head.args)
146
172
 
147
173
  # -- the fixpoint --------------------------------------------------------
148
174
 
@@ -152,21 +178,39 @@ class TabledEngine:
152
178
  check_query_atom(atom, self.arity)
153
179
  root = (atom.pred, self._pattern(atom, {}))
154
180
  self.tables = {root: set()}
181
+ self.complete = set()
182
+ self.open = {root: None}
155
183
  self.rounds = 0
184
+ self._fixpoint()
185
+ return {t for t in self.tables[root]
186
+ if _match(atom.args, t, {}) is not None}
187
+
188
+ def _fixpoint(self, below=None):
189
+ """Re-solve tables until none grows and no new subgoal appears.
190
+ With `below`, only the tables of strata up to that one — the
191
+ nested fixpoint a negation runs to complete what it reads; they
192
+ are then marked complete. Without, every table: the query's
193
+ own fixpoint, whose rounds are the ones counted."""
156
194
  changed = True
157
195
  while changed:
196
+ known = len(self.tables)
158
197
  changed = False
159
- self._grew = False
160
- self.rounds += 1
161
- for key in list(self.tables):
198
+ if below is None:
199
+ self.rounds += 1
200
+ for key in list(self.open): # complete tables cannot grow
201
+ if below is not None and self.strata.get(key[0], 0) > below:
202
+ continue
162
203
  table = self.tables[key]
163
204
  for ans in list(self._answers_for(key)):
164
205
  if ans not in table:
165
206
  table.add(ans)
166
207
  changed = True
167
- changed = changed or self._grew
168
- return {t for t in self.tables[root]
169
- if _match(atom.args, t, {}) is not None}
208
+ changed = changed or len(self.tables) != known
209
+ done = {k for k in self.open
210
+ if below is None or self.strata.get(k[0], 0) <= below}
211
+ self.complete |= done
212
+ for k in done:
213
+ del self.open[k]
170
214
 
171
215
 
172
216
  # ---------------------------------------------------------------------------
@@ -175,9 +219,9 @@ class TabledEngine:
175
219
 
176
220
  def main(argv=None):
177
221
  ap = argparse.ArgumentParser(
178
- description="Tabled top-down (QSQR) evaluation of positive "
222
+ description="Tabled top-down (QSQR) evaluation of stratified "
179
223
  "Datalog — handles left recursion SLD cannot.")
180
- ap.add_argument("file", help="Datalog program (.dl), positive rules only")
224
+ ap.add_argument("file", help="Datalog program (.dl), stratified")
181
225
  ap.add_argument("-q", "--query", action="append", default=[],
182
226
  metavar="ATOM", help="goal to solve (repeatable)")
183
227
  ap.add_argument("-t", "--tables", action="store_true",
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: tiny-datalog
3
- Version: 0.1.2
3
+ Version: 0.3.0
4
4
  Summary: A Datalog engine small enough to read in an afternoon
5
5
  Author: Andrew Goodchild
6
6
  License-Expression: MIT
@@ -9,7 +9,7 @@ Project-URL: Source, https://github.com/andrewgoodchild/tiny-datalog
9
9
  Project-URL: Issues, https://github.com/andrewgoodchild/tiny-datalog/issues
10
10
  Project-URL: Changelog, https://github.com/andrewgoodchild/tiny-datalog/releases
11
11
  Project-URL: Lessons, https://github.com/andrewgoodchild/tiny-datalog/tree/main/lessons
12
- Keywords: datalog,logic programming,deductive database,semi-naive evaluation,magic sets,provenance,stable models,teaching
12
+ Keywords: datalog,logic programming,deductive database,semi-naive evaluation,magic sets,provenance,stable models,defeasible logic,teaching
13
13
  Classifier: Development Status :: 3 - Alpha
14
14
  Classifier: Intended Audience :: Developers
15
15
  Classifier: Intended Audience :: Education
@@ -126,7 +126,7 @@ Nothing to install:
126
126
 
127
127
  ```sh
128
128
  git clone https://github.com/andrewgoodchild/tiny-datalog && cd tiny-datalog
129
- python3 tests.py # 235 tests, ~12s
129
+ python3 tests.py # 254 tests, ~12s
130
130
  ```
131
131
 
132
132
  ## Why the language choice decides what you can ask later
@@ -235,8 +235,10 @@ Answers are also checked against engines this repository did not write.
235
235
  [`conformance/`](https://github.com/andrewgoodchild/tiny-datalog/tree/main/conformance)
236
236
  runs the [datalog-conformance](https://pypi.org/project/datalog-conformance/)
237
237
  corpus, harvested from Soufflé, Nemo and Crepe, through all four
238
- strategies: 89 cases pass. The 25 it skips need arithmetic or 10⁵-fact
239
- joins, and both are omissions this course makes on purpose.
238
+ strategies, and its defeasible theories, from SPINdle and the papers,
239
+ through `defeasible.py`: 195 cases pass. Every skip is named with its
240
+ reason: arithmetic and 10⁵-fact joins, which this course omits on
241
+ purpose, or cases written for a different logic.
240
242
 
241
243
  ## Learning Datalog
242
244
 
@@ -246,7 +248,8 @@ the current research threads. The field's own recent lecture notes
246
248
  observe that the literature advises people building Datalog engines
247
249
  better than people trying to *use* one; this course does both halves
248
250
  on purpose — sixteen lessons where the engine is the explanation, then
249
- a lesson on authoring rules that survive review. And it is built to be
251
+ a lesson on authoring rules that survive review, and one on a logic
252
+ built for rules with exceptions. And it is built to be
250
253
  inherited: `git clone`, no dependencies, no hosted anything, and every
251
254
  quoted transcript re-verified by CI — the exercises cannot rot. (For where each
252
255
  technique ships — CodeQL, RDFox, Feldera, SNOMED and the rest —
@@ -286,7 +289,9 @@ regression test without writing Python.
286
289
 
287
290
  Not an engine to build a product on. Joins are nested-loop,
288
291
  stable-model search is exhaustive, evaluation is batch. For real
289
- workloads see Soufflé, clingo, RDFox or Feldera.
292
+ workloads see Soufflé, clingo, RDFox or Feldera; for Datalog inside a
293
+ Python application, with arithmetic and queries over SQL databases,
294
+ see [pyDatalog](https://pypi.org/project/pyDatalog/).
290
295
 
291
296
  Deliberate omissions, because saying why teaches more than lacking them
292
297
  quietly:
@@ -303,7 +308,7 @@ quietly:
303
308
  (disjointness and unsatisfiability detection included). SNOMED needs
304
309
  ELH (EL plus role hierarchies) with right identities, which is what
305
310
  the ELK and Snorocket reasoners implement and this does not.
306
- - **A REPL (interactive prompt) and packaging.** `git clone` and run.
311
+ - **A REPL (interactive prompt).** Run a file, or call it from Python.
307
312
 
308
313
  Aggregation used to be on this list;
309
314
  [lesson 13](https://github.com/andrewgoodchild/tiny-datalog/blob/main/lessons/13-aggregation.md) is what promoting an omission
@@ -321,13 +326,32 @@ pip install tiny-datalog
321
326
  from tiny_datalog import run_program, explain
322
327
 
323
328
  engine = run_program(open("supply-chain.dl").read())
324
- for service, cve in sorted(engine.rels["exposed"]):
325
- print("\n".join(explain(engine, "exposed", (service, cve))))
329
+ for answer in engine.query("exposed(S, C)"):
330
+ print(answer) # {'S': 'pkg0', 'C': 'cve_2026_0001'}
331
+ print("\n".join(explain(engine, "exposed", ("pkg0", "cve_2026_0001"))))
332
+ ```
333
+
334
+ Facts can come straight from Python data (a CSV file, a database
335
+ query) instead of program text. Each row is checked as if it had been
336
+ parsed:
337
+
338
+ ```python
339
+ rules = """
340
+ uses(X, Y) :- depends(X, Y).
341
+ uses(X, Z) :- depends(X, Y), uses(Y, Z).
342
+ exposed(S, C) :- service(S), uses(S, L), vulnerable(L, C).
343
+ """
344
+ engine = run_program(rules, facts={
345
+ "depends": [("app", "lib"), ("lib", "core")],
346
+ "service": ["app"], # one-column rows may be bare
347
+ "vulnerable": [("core", "cve_2026_0001")],
348
+ })
349
+ engine.query("exposed(S, C)") # [{'S': 'app', 'C': 'cve_2026_0001'}]
326
350
  ```
327
351
 
328
352
  The command-line interface installs too, as `tiny-datalog` (and
329
353
  `tiny-datalog-semiring`, `-tabling`, `-incremental`, `-subsumption`,
330
- `-containment`, `-prolog` for the satellites):
354
+ `-containment`, `-defeasible`, `-prolog` for the satellites):
331
355
 
332
356
  ```sh
333
357
  tiny-datalog -q 'exposed(S, C)' supply-chain.dl
@@ -351,17 +375,18 @@ tiny_datalog/ the engine and its satellites — the code you read:
351
375
  tabling.py tabled top-down evaluation (iterative QSQR)
352
376
  subsumption.py KL-ONE-style EL classifier, compiled to Datalog
353
377
  containment.py query containment and minimisation by homomorphism
378
+ defeasible.py defeasible logic: exceptions, priorities, defeaters
354
379
  *.py three-line launchers, so `python3 datalog.py ...` works
355
380
  straight from a checkout with nothing installed
356
381
  programs/ teaching programs, numbered by the lesson that uses
357
382
  them (00-* are the README's examples)
358
- lessons/ getting started, glossary, and lessons 0–18
383
+ lessons/ getting started, glossary, and lessons 0–19
359
384
  exercises/ worked answers, verified by the test suite
360
385
  cases/ golden test cases — add one without writing Python
361
386
  conformance/ the external datalog-conformance corpus (Soufflé, Nemo,
362
- Crepe), run against all four evaluation strategies
387
+ Crepe, SPINdle), run against every evaluation strategy
363
388
  benchmarks/ scaled input generators (chain/tree/clique/grid)
364
- tests.py 235 tests: every shipped program and exercise answer is
389
+ tests.py 254 tests: every shipped program and exercise answer is
365
390
  executed, a conformance suite runs every query through
366
391
  every applicable strategy, and a seeded fuzzer checks
367
392
  the same property on random programs
@@ -375,13 +400,13 @@ used.
375
400
  ### How big is it, honestly
376
401
 
377
402
  The evaluator is about 850 lines (`tiny_datalog/datalog.py`, up to the
378
- command-line interface), the CLI, `--explain` and why-not another 600, and the eight satellite modules about
379
- 2,400. Call it 3.9k lines of toolkit and 2.6k of tests, roughly a
403
+ command-line interface), the CLI, `--explain` and why-not another 600, and the nine satellite modules about
404
+ 2,900. Call it 4.3k lines of toolkit and 2.7k of tests, roughly a
380
405
  quarter of it commentary.
381
406
 
382
407
  "Tiny" is a claim about the evaluator, and about each satellite module
383
- singly: none of the eight exceeds 500 lines, which a test asserts. It is not a claim about
384
- the repository, which is nine modules because it teaches nine things.
408
+ singly: none of the nine exceeds 500 lines, which a test asserts. It is not a claim about
409
+ the repository, which is ten modules because it teaches ten things.
385
410
 
386
411
  There is no dead code to golf away (checked); shrinking further means
387
412
  deleting either a technique or an explanation.
@@ -4,6 +4,7 @@ pyproject.toml
4
4
  tiny_datalog/__init__.py
5
5
  tiny_datalog/containment.py
6
6
  tiny_datalog/datalog.py
7
+ tiny_datalog/defeasible.py
7
8
  tiny_datalog/incremental.py
8
9
  tiny_datalog/magic.py
9
10
  tiny_datalog/prolog.py
@@ -1,6 +1,7 @@
1
1
  [console_scripts]
2
2
  tiny-datalog = tiny_datalog.datalog:main
3
3
  tiny-datalog-containment = tiny_datalog.containment:main
4
+ tiny-datalog-defeasible = tiny_datalog.defeasible:main
4
5
  tiny-datalog-incremental = tiny_datalog.incremental:main
5
6
  tiny-datalog-prolog = tiny_datalog.prolog:main
6
7
  tiny-datalog-semiring = tiny_datalog.semiring:main
File without changes
File without changes