eyeprolog 1.5.10 → 1.5.11
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +2 -0
- package/package.json +5 -2
- package/src/ARCHITECTURE.md +7 -0
- package/test/bench/benchmarks.json +184 -0
- package/test/bench/eyelet-forward.pl +58 -0
- package/test/benchmark-core.mjs +94 -0
- package/test/benchmark-worker.mjs +28 -0
- package/test/benchmark.mjs +228 -0
- package/test/run-benchmark-tests.mjs +73 -0
- package/the-art-of-eyeprolog.md +30 -1
package/README.md
CHANGED
|
@@ -85,6 +85,8 @@ The checked [Symbiotic Knowledge Graphs example](examples/symbiotic-knowledge-gr
|
|
|
85
85
|
|
|
86
86
|
The same RDF → Prolog → RDF boundary is exercised by five additional checked scenarios: [cross-organization data sharing](https://eyereasoner.github.io/eyeprolog/examples/deck/cross-organization-data-sharing), [explainable EV-depot configuration](https://eyereasoner.github.io/eyeprolog/examples/deck/explainable-ev-depot-configuration), [operational incident response](https://eyereasoner.github.io/eyeprolog/examples/deck/operational-incident-response), [software supply-chain vulnerability response](https://eyereasoner.github.io/eyeprolog/examples/deck/sbom-vulnerability-response), and a [scientific evidence graph](https://eyereasoner.github.io/eyeprolog/examples/deck/scientific-evidence-graph). Together they cover policy decisions, reversible configuration reasoning, dependency-graph diagnosis, transitive SBOM exposure, and evidence aggregation with explicit disagreement.
|
|
87
87
|
|
|
88
|
+
## Benchmarks
|
|
89
|
+
EyeProlog has 20 checksum-protected wall-clock benchmarks spanning recursion/indexing, constraints, tabling/WFS, DCGs, Eyelet, search, term I/O, attributes, and rewriting. Short workloads are adaptively batched before timing so millisecond-scale noise is not mistaken for a regression. Run `npm run benchmark`; create a machine-local comparison point with `npm run benchmark:baseline`; use `npm run test:benchmark` for harness checks. Details are in [*The Art of EyeProlog*](the-art-of-eyeprolog.md).
|
|
88
90
|
## Development
|
|
89
91
|
|
|
90
92
|
```sh
|
package/package.json
CHANGED
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
"publishConfig": {
|
|
4
4
|
"access": "public"
|
|
5
5
|
},
|
|
6
|
-
"version": "1.5.
|
|
6
|
+
"version": "1.5.11",
|
|
7
7
|
"description": "EyeProlog turns facts and rules into answers and proofs.",
|
|
8
8
|
"type": "module",
|
|
9
9
|
"main": "./index.js",
|
|
@@ -66,6 +66,9 @@
|
|
|
66
66
|
"report:wg17-syntax": "node tools/report-wg17-syntax-coverage.mjs",
|
|
67
67
|
"preversion": "npm test && node test/run-conformance-report.mjs conformance-report.md",
|
|
68
68
|
"postversion": "git push origin HEAD --follow-tags",
|
|
69
|
-
"test:architecture": "node test/run-architecture.mjs"
|
|
69
|
+
"test:architecture": "node test/run-architecture.mjs",
|
|
70
|
+
"benchmark": "node test/benchmark.mjs",
|
|
71
|
+
"benchmark:baseline": "node test/benchmark.mjs --save .benchmarks/baseline.json",
|
|
72
|
+
"test:benchmark": "node test/run-benchmark-tests.mjs"
|
|
70
73
|
}
|
|
71
74
|
}
|
package/src/ARCHITECTURE.md
CHANGED
|
@@ -88,3 +88,10 @@ retains the same data structures and selection functions.
|
|
|
88
88
|
Large solver fast paths deliberately remain co-located in `solver.js` until a
|
|
89
89
|
split can demonstrate benchmark parity. A cleaner file layout is not worth a
|
|
90
90
|
runtime regression.
|
|
91
|
+
|
|
92
|
+
Performance claims use the wall-clock workloads in `test/bench/`, not predicate,
|
|
93
|
+
inference, or host-call counts as a substitute for elapsed time. Each benchmark
|
|
94
|
+
runs in its own Node process, warms the engine before measured parse+execute
|
|
95
|
+
runs, reports the median, and verifies a committed output SHA-256 before its
|
|
96
|
+
timing is accepted. Machine-specific timing baselines live under `.benchmarks/`
|
|
97
|
+
and are intentionally not versioned.
|
|
@@ -0,0 +1,184 @@
|
|
|
1
|
+
[
|
|
2
|
+
{
|
|
3
|
+
"name": "deep-taxonomy",
|
|
4
|
+
"group": "recursion/indexing",
|
|
5
|
+
"file": "examples/deep-taxonomy-100000.pl",
|
|
6
|
+
"goals": [
|
|
7
|
+
"result(report, success)"
|
|
8
|
+
],
|
|
9
|
+
"expectedSha256": "a0aa24fcbc87c394a4da789ac861b8e18cabad5e0d73b25ddc02df5dd4376c1b"
|
|
10
|
+
},
|
|
11
|
+
{
|
|
12
|
+
"name": "clpz-sudoku",
|
|
13
|
+
"group": "clpz",
|
|
14
|
+
"file": "examples/clpz-sudoku-9x9.pl",
|
|
15
|
+
"goals": [
|
|
16
|
+
"sudoku9_solution(X0)"
|
|
17
|
+
],
|
|
18
|
+
"expectedSha256": "0eed3bcf38498cb514143adb2e475d8e48b0326bcf329fbedb6a3b0e87646731"
|
|
19
|
+
},
|
|
20
|
+
{
|
|
21
|
+
"name": "clpb-weighted-planning",
|
|
22
|
+
"group": "clpb",
|
|
23
|
+
"file": "examples/clpb-weighted-planning.pl",
|
|
24
|
+
"goals": [
|
|
25
|
+
"best_release_plan(X0, X1)"
|
|
26
|
+
],
|
|
27
|
+
"expectedSha256": "52f2cc7d22873accef7968f89aea3584fb3304b2dceb555428e79477b41f0b5d"
|
|
28
|
+
},
|
|
29
|
+
{
|
|
30
|
+
"name": "clpb-cardinality",
|
|
31
|
+
"group": "clpb",
|
|
32
|
+
"file": "examples/clpb-cardinality.pl",
|
|
33
|
+
"goals": [
|
|
34
|
+
"review_quorum_count(X0)"
|
|
35
|
+
],
|
|
36
|
+
"expectedSha256": "a643a582baefc95da37b0ccf9609a2241aeb7cfe559e2636bb439c05c50e2177"
|
|
37
|
+
},
|
|
38
|
+
{
|
|
39
|
+
"name": "tabling-cycle",
|
|
40
|
+
"group": "tabling",
|
|
41
|
+
"file": "examples/cyclic-path.pl",
|
|
42
|
+
"goals": [
|
|
43
|
+
"path(X0, X1)"
|
|
44
|
+
],
|
|
45
|
+
"expectedSha256": "5580a2782c6896ffdf36bd6cc242931e96e1722b154976c1a20e7cad1206566f"
|
|
46
|
+
},
|
|
47
|
+
{
|
|
48
|
+
"name": "tabling-chart-parser",
|
|
49
|
+
"group": "tabling",
|
|
50
|
+
"file": "examples/chart-parser.pl",
|
|
51
|
+
"goals": [
|
|
52
|
+
"chart_parser_answer(X0, X1)"
|
|
53
|
+
],
|
|
54
|
+
"expectedSha256": "70707d3026aabc65982b82b9617a15d3de6cbb5fa6cdb2ee955bedd4f29f0e5a"
|
|
55
|
+
},
|
|
56
|
+
{
|
|
57
|
+
"name": "wfs-policy",
|
|
58
|
+
"group": "wfs",
|
|
59
|
+
"file": "examples/odrl-policy-reasoning.pl",
|
|
60
|
+
"goals": [
|
|
61
|
+
"wfsQuestion(X0, X1)"
|
|
62
|
+
],
|
|
63
|
+
"expectedSha256": "b4a153234f3daf1e2ba6c26843349a1bc1919ebd9f9e98559835eb5ef915716b"
|
|
64
|
+
},
|
|
65
|
+
{
|
|
66
|
+
"name": "dcg-command",
|
|
67
|
+
"group": "dcg",
|
|
68
|
+
"file": "examples/dcg-command-parser.pl",
|
|
69
|
+
"goals": [
|
|
70
|
+
"dcg_example(X0, X1)"
|
|
71
|
+
],
|
|
72
|
+
"expectedSha256": "6f640fb0d327eb9139d3568cf7a86ad55a52afdbb0208ca35a70419be2bf1b9c"
|
|
73
|
+
},
|
|
74
|
+
{
|
|
75
|
+
"name": "dcg-expression",
|
|
76
|
+
"group": "dcg",
|
|
77
|
+
"file": "examples/dcg-expression-language.pl",
|
|
78
|
+
"goals": [
|
|
79
|
+
"dcg_expression_example(X0, X1)"
|
|
80
|
+
],
|
|
81
|
+
"expectedSha256": "562985450257f1db8c67682c6184102622b4e6946f2cb0abd3ac15bde493f5d5"
|
|
82
|
+
},
|
|
83
|
+
{
|
|
84
|
+
"name": "eyelet-forward",
|
|
85
|
+
"group": "eyelet",
|
|
86
|
+
"file": "test/bench/eyelet-forward.pl",
|
|
87
|
+
"goals": [],
|
|
88
|
+
"expectedSha256": "127277f2ba5c02d41143b22ac0ae069596a4a7dd4dc71079f3dff1319ca019cc"
|
|
89
|
+
},
|
|
90
|
+
{
|
|
91
|
+
"name": "takeuchi",
|
|
92
|
+
"group": "recursive-control",
|
|
93
|
+
"file": "examples/takeuchi.pl",
|
|
94
|
+
"goals": [
|
|
95
|
+
"tak([34, 13, 8], X0)"
|
|
96
|
+
],
|
|
97
|
+
"expectedSha256": "a7c33872b86aa7877b9c38ec41c447fc96a328d40f512428993d662f28a097ad"
|
|
98
|
+
},
|
|
99
|
+
{
|
|
100
|
+
"name": "clpz-register-allocation",
|
|
101
|
+
"group": "clpz",
|
|
102
|
+
"file": "examples/register-allocation.pl",
|
|
103
|
+
"goals": [
|
|
104
|
+
"registerAnswer(X0, X1)"
|
|
105
|
+
],
|
|
106
|
+
"expectedSha256": "1ed0b4a90deff2908463774d0025210d58b89366efb5d82b94ff5b51daca046b"
|
|
107
|
+
},
|
|
108
|
+
{
|
|
109
|
+
"name": "clpb-feature-model",
|
|
110
|
+
"group": "clpb",
|
|
111
|
+
"file": "examples/clpb-feature-model.pl",
|
|
112
|
+
"goals": [
|
|
113
|
+
"feature_plan(X0)"
|
|
114
|
+
],
|
|
115
|
+
"expectedSha256": "47abfd8a87efb595ffb11c6c8cbda3cb5855e95124bc574e3cae387c12318f33"
|
|
116
|
+
},
|
|
117
|
+
{
|
|
118
|
+
"name": "attributed-variables",
|
|
119
|
+
"group": "attributes",
|
|
120
|
+
"file": "examples/attributed-variables.pl",
|
|
121
|
+
"goals": [
|
|
122
|
+
"accepts_required_value(X0)",
|
|
123
|
+
"alias_preserves_attribute(X0)"
|
|
124
|
+
],
|
|
125
|
+
"expectedSha256": "3d60cd3464f2aae44ad01859e838fc2b8754e35916bafe3bba66ca745f77c236"
|
|
126
|
+
},
|
|
127
|
+
{
|
|
128
|
+
"name": "dif-constraints",
|
|
129
|
+
"group": "constraints",
|
|
130
|
+
"file": "examples/dif-constraints.pl",
|
|
131
|
+
"goals": [
|
|
132
|
+
"allowed_pair(X0, X1)",
|
|
133
|
+
"specialization(X0, X1)"
|
|
134
|
+
],
|
|
135
|
+
"expectedSha256": "b6268a6653d02f7312f73406612b451d030a3766931e093003c3642cc170009d"
|
|
136
|
+
},
|
|
137
|
+
{
|
|
138
|
+
"name": "dijkstra-findall-sort",
|
|
139
|
+
"group": "search/collections",
|
|
140
|
+
"file": "examples/dijkstra-findall-sort.pl",
|
|
141
|
+
"goals": [
|
|
142
|
+
"shortestPath(X0, X1)",
|
|
143
|
+
"cost(X0, X1)"
|
|
144
|
+
],
|
|
145
|
+
"expectedSha256": "1318aaaca44aa19f295f988d87b09165cd0fc279e94ae14c10e2d1f146d544f3"
|
|
146
|
+
},
|
|
147
|
+
{
|
|
148
|
+
"name": "iso-term-io",
|
|
149
|
+
"group": "term-io",
|
|
150
|
+
"file": "examples/iso-term-io.pl",
|
|
151
|
+
"goals": [
|
|
152
|
+
"report(X0, X1)"
|
|
153
|
+
],
|
|
154
|
+
"expectedSha256": "0719932d697e7587285082b558f4c39c04ec1f2222131073e73e6b6fe37e0320"
|
|
155
|
+
},
|
|
156
|
+
{
|
|
157
|
+
"name": "symbolic-derivative",
|
|
158
|
+
"group": "term-rewriting",
|
|
159
|
+
"file": "examples/symbolic-derivative.pl",
|
|
160
|
+
"goals": [
|
|
161
|
+
"derivative_result(X0, X1)"
|
|
162
|
+
],
|
|
163
|
+
"expectedSha256": "49f9f33e89e8c8c63ce2eb64ea2639e416869d6ef5db185cab70452d1f77f342"
|
|
164
|
+
},
|
|
165
|
+
{
|
|
166
|
+
"name": "sat-dpll",
|
|
167
|
+
"group": "search/control",
|
|
168
|
+
"file": "examples/sat-solver-dpll.pl",
|
|
169
|
+
"goals": [
|
|
170
|
+
"satModel(X0)",
|
|
171
|
+
"satValue(X0, X1)"
|
|
172
|
+
],
|
|
173
|
+
"expectedSha256": "a8c5cdddac5efcb7d5b1031200d3bbc5b78c1668928a45e064377432f05b3c9f"
|
|
174
|
+
},
|
|
175
|
+
{
|
|
176
|
+
"name": "type-inference",
|
|
177
|
+
"group": "unification",
|
|
178
|
+
"file": "examples/type-inference.pl",
|
|
179
|
+
"goals": [
|
|
180
|
+
"type_answer(X0, X1)"
|
|
181
|
+
],
|
|
182
|
+
"expectedSha256": "acad9ddb995d9ddef091e17cc91659d36ed7f9014d6714fd40100d85a53ae203"
|
|
183
|
+
}
|
|
184
|
+
]
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
% Synthetic Eyelet fixed-point benchmark.
|
|
2
|
+
% The chain is intentionally moderate: large enough to exercise repeated
|
|
3
|
+
% forward-rule rounds without making the default benchmark suite unwieldy.
|
|
4
|
+
|
|
5
|
+
reach(0).
|
|
6
|
+
edge(0, 1).
|
|
7
|
+
edge(1, 2).
|
|
8
|
+
edge(2, 3).
|
|
9
|
+
edge(3, 4).
|
|
10
|
+
edge(4, 5).
|
|
11
|
+
edge(5, 6).
|
|
12
|
+
edge(6, 7).
|
|
13
|
+
edge(7, 8).
|
|
14
|
+
edge(8, 9).
|
|
15
|
+
edge(9, 10).
|
|
16
|
+
edge(10, 11).
|
|
17
|
+
edge(11, 12).
|
|
18
|
+
edge(12, 13).
|
|
19
|
+
edge(13, 14).
|
|
20
|
+
edge(14, 15).
|
|
21
|
+
edge(15, 16).
|
|
22
|
+
edge(16, 17).
|
|
23
|
+
edge(17, 18).
|
|
24
|
+
edge(18, 19).
|
|
25
|
+
edge(19, 20).
|
|
26
|
+
edge(20, 21).
|
|
27
|
+
edge(21, 22).
|
|
28
|
+
edge(22, 23).
|
|
29
|
+
edge(23, 24).
|
|
30
|
+
edge(24, 25).
|
|
31
|
+
edge(25, 26).
|
|
32
|
+
edge(26, 27).
|
|
33
|
+
edge(27, 28).
|
|
34
|
+
edge(28, 29).
|
|
35
|
+
edge(29, 30).
|
|
36
|
+
edge(30, 31).
|
|
37
|
+
edge(31, 32).
|
|
38
|
+
edge(32, 33).
|
|
39
|
+
edge(33, 34).
|
|
40
|
+
edge(34, 35).
|
|
41
|
+
edge(35, 36).
|
|
42
|
+
edge(36, 37).
|
|
43
|
+
edge(37, 38).
|
|
44
|
+
edge(38, 39).
|
|
45
|
+
edge(39, 40).
|
|
46
|
+
edge(40, 41).
|
|
47
|
+
edge(41, 42).
|
|
48
|
+
edge(42, 43).
|
|
49
|
+
edge(43, 44).
|
|
50
|
+
edge(44, 45).
|
|
51
|
+
edge(45, 46).
|
|
52
|
+
edge(46, 47).
|
|
53
|
+
edge(47, 48).
|
|
54
|
+
edge(48, 49).
|
|
55
|
+
edge(49, 50).
|
|
56
|
+
|
|
57
|
+
reach(Y) :+ reach(X), edge(X, Y).
|
|
58
|
+
true :+ reach(50).
|
|
@@ -0,0 +1,94 @@
|
|
|
1
|
+
import crypto from 'node:crypto';
|
|
2
|
+
import { performance } from 'node:perf_hooks';
|
|
3
|
+
import { run } from '../index.js';
|
|
4
|
+
|
|
5
|
+
function digest(text) {
|
|
6
|
+
return crypto.createHash('sha256').update(text).digest('hex');
|
|
7
|
+
}
|
|
8
|
+
|
|
9
|
+
function nextPowerOfTwo(value) {
|
|
10
|
+
let power = 1;
|
|
11
|
+
while (power < value && power < 1024) power *= 2;
|
|
12
|
+
return power;
|
|
13
|
+
}
|
|
14
|
+
|
|
15
|
+
export function measureBenchmark(source, goals, { runs = 5, warmup = 1, targetMs = 400 } = {}) {
|
|
16
|
+
let referenceStdout = null;
|
|
17
|
+
let referenceDigest = null;
|
|
18
|
+
let answerLines = null;
|
|
19
|
+
let outputBytes = null;
|
|
20
|
+
|
|
21
|
+
function execute() {
|
|
22
|
+
return run(source, { goals });
|
|
23
|
+
}
|
|
24
|
+
|
|
25
|
+
function validateResult(result) {
|
|
26
|
+
if (result.haltCode != null) throw new Error(`benchmark halted with code ${result.haltCode}`);
|
|
27
|
+
if (referenceStdout == null) {
|
|
28
|
+
referenceStdout = result.stdout;
|
|
29
|
+
referenceDigest = digest(result.stdout);
|
|
30
|
+
answerLines = result.stdout.split('\n').filter((line) => line.length > 0).length;
|
|
31
|
+
outputBytes = Buffer.byteLength(result.stdout);
|
|
32
|
+
} else if (result.stdout !== referenceStdout) {
|
|
33
|
+
throw new Error(`non-deterministic output digest: ${referenceDigest} != ${digest(result.stdout)}`);
|
|
34
|
+
}
|
|
35
|
+
}
|
|
36
|
+
|
|
37
|
+
function runBatch(batchSize) {
|
|
38
|
+
const results = new Array(batchSize);
|
|
39
|
+
const started = performance.now();
|
|
40
|
+
for (let i = 0; i < batchSize; i++) results[i] = execute();
|
|
41
|
+
const elapsed = performance.now() - started;
|
|
42
|
+
for (const result of results) validateResult(result);
|
|
43
|
+
return elapsed;
|
|
44
|
+
}
|
|
45
|
+
|
|
46
|
+
let batchSize = 1;
|
|
47
|
+
let calibrationMs = 0;
|
|
48
|
+
if (targetMs > 0) {
|
|
49
|
+
runBatch(1); // prime parser/JIT/module state before sizing the measured batch
|
|
50
|
+
calibrationMs = runBatch(1);
|
|
51
|
+
if (calibrationMs < targetMs) {
|
|
52
|
+
batchSize = nextPowerOfTwo(Math.ceil(targetMs / Math.max(calibrationMs, 0.01)));
|
|
53
|
+
calibrationMs = runBatch(batchSize);
|
|
54
|
+
while (calibrationMs < targetMs * 0.8 && batchSize < 1024) {
|
|
55
|
+
batchSize = Math.min(batchSize * 2, 1024);
|
|
56
|
+
calibrationMs = runBatch(batchSize);
|
|
57
|
+
}
|
|
58
|
+
while (calibrationMs > targetMs * 2 && batchSize > 1) {
|
|
59
|
+
batchSize = Math.max(Math.floor(batchSize / 2), 1);
|
|
60
|
+
calibrationMs = runBatch(batchSize);
|
|
61
|
+
}
|
|
62
|
+
}
|
|
63
|
+
}
|
|
64
|
+
|
|
65
|
+
for (let i = 0; i < warmup; i++) runBatch(batchSize);
|
|
66
|
+
|
|
67
|
+
// Warm-up can materially change short-workload cost. Re-check the chosen batch
|
|
68
|
+
// in the warmed state and grow it until samples are long enough to be useful.
|
|
69
|
+
if (targetMs > 0) {
|
|
70
|
+
calibrationMs = runBatch(batchSize);
|
|
71
|
+
while (calibrationMs < targetMs * 0.8 && batchSize < 1024) {
|
|
72
|
+
batchSize = Math.min(batchSize * 2, 1024);
|
|
73
|
+
calibrationMs = runBatch(batchSize);
|
|
74
|
+
}
|
|
75
|
+
}
|
|
76
|
+
|
|
77
|
+
const samplesMs = [];
|
|
78
|
+
const batchSamplesMs = [];
|
|
79
|
+
for (let i = 0; i < runs; i++) {
|
|
80
|
+
const elapsed = runBatch(batchSize);
|
|
81
|
+
batchSamplesMs.push(elapsed);
|
|
82
|
+
samplesMs.push(elapsed / batchSize);
|
|
83
|
+
}
|
|
84
|
+
|
|
85
|
+
return {
|
|
86
|
+
digest: referenceDigest,
|
|
87
|
+
answerLines,
|
|
88
|
+
outputBytes,
|
|
89
|
+
batchSize,
|
|
90
|
+
calibrationMs,
|
|
91
|
+
batchSamplesMs,
|
|
92
|
+
samplesMs,
|
|
93
|
+
};
|
|
94
|
+
}
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
import fs from 'node:fs/promises';
|
|
2
|
+
import path from 'node:path';
|
|
3
|
+
import process from 'node:process';
|
|
4
|
+
import { measureBenchmark } from './benchmark-core.mjs';
|
|
5
|
+
|
|
6
|
+
function parseInteger(value, name, minimum) {
|
|
7
|
+
const number = Number(value);
|
|
8
|
+
if (!Number.isInteger(number) || number < minimum) throw new Error(`${name} must be an integer >= ${minimum}`);
|
|
9
|
+
return number;
|
|
10
|
+
}
|
|
11
|
+
|
|
12
|
+
function parseNumber(value, name, minimum) {
|
|
13
|
+
const number = Number(value);
|
|
14
|
+
if (!Number.isFinite(number) || number < minimum) throw new Error(`${name} must be a number >= ${minimum}`);
|
|
15
|
+
return number;
|
|
16
|
+
}
|
|
17
|
+
|
|
18
|
+
const [fileArg, goalsJson = '[]', runsArg = '5', warmupArg = '1', targetMsArg = '400'] = process.argv.slice(2);
|
|
19
|
+
if (!fileArg) throw new Error('benchmark worker requires a file');
|
|
20
|
+
const goals = JSON.parse(goalsJson);
|
|
21
|
+
if (!Array.isArray(goals) || !goals.every((goal) => typeof goal === 'string')) throw new Error('benchmark worker goals must be a JSON array of strings');
|
|
22
|
+
const source = await fs.readFile(path.resolve(fileArg), 'utf8');
|
|
23
|
+
const result = measureBenchmark(source, goals, {
|
|
24
|
+
runs: parseInteger(runsArg, 'runs', 1),
|
|
25
|
+
warmup: parseInteger(warmupArg, 'warmup', 0),
|
|
26
|
+
targetMs: parseNumber(targetMsArg, 'targetMs', 0),
|
|
27
|
+
});
|
|
28
|
+
process.stdout.write(`${JSON.stringify(result)}\n`);
|
|
@@ -0,0 +1,228 @@
|
|
|
1
|
+
import fs from 'node:fs/promises';
|
|
2
|
+
import path from 'node:path';
|
|
3
|
+
import process from 'node:process';
|
|
4
|
+
import { spawnSync } from 'node:child_process';
|
|
5
|
+
import { fileURLToPath } from 'node:url';
|
|
6
|
+
|
|
7
|
+
const root = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '..');
|
|
8
|
+
const manifestPath = path.join(root, 'test', 'bench', 'benchmarks.json');
|
|
9
|
+
const defaultBaselinePath = path.join(root, '.benchmarks', 'baseline.json');
|
|
10
|
+
|
|
11
|
+
function usage() {
|
|
12
|
+
process.stdout.write(`EyeProlog benchmark runner\n\nUsage:\n node test/benchmark.mjs [options]\n\nOptions:\n --runs N Measured batches per benchmark (default: 5)\n --warmup N Warm-up batches per benchmark (default: 1)\n --target-ms N Minimum target for short measured batches (default: 400)\n --filter TEXT Run benchmarks whose name or group contains TEXT\n --baseline FILE Compare against a saved timing baseline\n --save FILE Save current medians as a timing baseline\n --json Print machine-readable JSON\n --list List benchmark names and exit\n -h, --help Show this help\n\nAfter one untimed priming execution, short workloads are automatically repeated\nin independent run() executions until a measured batch lasts about --target-ms. The report divides batch time by the\niteration count, so all results remain milliseconds per workload execution.\nOutput digests are always checked against the committed semantic checksums in\ntest/bench/benchmarks.json.\n`);
|
|
13
|
+
}
|
|
14
|
+
|
|
15
|
+
function parseInteger(value, name, minimum) {
|
|
16
|
+
const number = Number(value);
|
|
17
|
+
if (!Number.isInteger(number) || number < minimum) {
|
|
18
|
+
throw new Error(`${name} must be an integer >= ${minimum}`);
|
|
19
|
+
}
|
|
20
|
+
return number;
|
|
21
|
+
}
|
|
22
|
+
|
|
23
|
+
function parseNumber(value, name, minimum) {
|
|
24
|
+
const number = Number(value);
|
|
25
|
+
if (!Number.isFinite(number) || number < minimum) {
|
|
26
|
+
throw new Error(`${name} must be a number >= ${minimum}`);
|
|
27
|
+
}
|
|
28
|
+
return number;
|
|
29
|
+
}
|
|
30
|
+
|
|
31
|
+
const options = {
|
|
32
|
+
runs: 5,
|
|
33
|
+
warmup: 1,
|
|
34
|
+
targetMs: 400,
|
|
35
|
+
filter: null,
|
|
36
|
+
baseline: null,
|
|
37
|
+
save: null,
|
|
38
|
+
json: false,
|
|
39
|
+
list: false,
|
|
40
|
+
};
|
|
41
|
+
|
|
42
|
+
for (let i = 2; i < process.argv.length; i++) {
|
|
43
|
+
const arg = process.argv[i];
|
|
44
|
+
if (arg === '--runs') options.runs = parseInteger(process.argv[++i], '--runs', 1);
|
|
45
|
+
else if (arg === '--warmup') options.warmup = parseInteger(process.argv[++i], '--warmup', 0);
|
|
46
|
+
else if (arg === '--target-ms') options.targetMs = parseNumber(process.argv[++i], '--target-ms', 0);
|
|
47
|
+
else if (arg === '--filter') options.filter = process.argv[++i] ?? '';
|
|
48
|
+
else if (arg === '--baseline') options.baseline = path.resolve(process.argv[++i] ?? '');
|
|
49
|
+
else if (arg === '--save') options.save = path.resolve(process.argv[++i] ?? '');
|
|
50
|
+
else if (arg === '--json') options.json = true;
|
|
51
|
+
else if (arg === '--list') options.list = true;
|
|
52
|
+
else if (arg === '--help' || arg === '-h') { usage(); process.exit(0); }
|
|
53
|
+
else throw new Error(`unknown option: ${arg}`);
|
|
54
|
+
}
|
|
55
|
+
|
|
56
|
+
const manifest = JSON.parse(await fs.readFile(manifestPath, 'utf8'));
|
|
57
|
+
if (!Array.isArray(manifest) || manifest.length === 0) throw new Error('benchmark manifest is empty');
|
|
58
|
+
|
|
59
|
+
const names = new Set();
|
|
60
|
+
for (const item of manifest) {
|
|
61
|
+
if (!item || typeof item.name !== 'string' || typeof item.group !== 'string' || typeof item.file !== 'string') {
|
|
62
|
+
throw new Error('invalid benchmark manifest entry');
|
|
63
|
+
}
|
|
64
|
+
if (names.has(item.name)) throw new Error(`duplicate benchmark name: ${item.name}`);
|
|
65
|
+
names.add(item.name);
|
|
66
|
+
if (!Array.isArray(item.goals) || !item.goals.every((goal) => typeof goal === 'string')) {
|
|
67
|
+
throw new Error(`invalid goals for benchmark ${item.name}`);
|
|
68
|
+
}
|
|
69
|
+
if (!/^[0-9a-f]{64}$/.test(item.expectedSha256)) {
|
|
70
|
+
throw new Error(`invalid expectedSha256 for benchmark ${item.name}`);
|
|
71
|
+
}
|
|
72
|
+
}
|
|
73
|
+
|
|
74
|
+
const selected = manifest.filter((item) => {
|
|
75
|
+
if (options.filter == null) return true;
|
|
76
|
+
const needle = options.filter.toLowerCase();
|
|
77
|
+
return item.name.toLowerCase().includes(needle) || item.group.toLowerCase().includes(needle);
|
|
78
|
+
}).sort((a, b) => a.name.localeCompare(b.name));
|
|
79
|
+
if (selected.length === 0) throw new Error(`no benchmarks match filter: ${options.filter}`);
|
|
80
|
+
|
|
81
|
+
if (options.list) {
|
|
82
|
+
for (const item of selected) process.stdout.write(`${item.name}\t${item.group}\t${item.file}\n`);
|
|
83
|
+
process.exit(0);
|
|
84
|
+
}
|
|
85
|
+
|
|
86
|
+
let baselinePath = options.baseline;
|
|
87
|
+
if (baselinePath == null) {
|
|
88
|
+
try {
|
|
89
|
+
await fs.access(defaultBaselinePath);
|
|
90
|
+
baselinePath = defaultBaselinePath;
|
|
91
|
+
} catch (_) {
|
|
92
|
+
baselinePath = null;
|
|
93
|
+
}
|
|
94
|
+
}
|
|
95
|
+
|
|
96
|
+
let baseline = null;
|
|
97
|
+
let baselineWarning = null;
|
|
98
|
+
if (baselinePath != null) {
|
|
99
|
+
const loaded = JSON.parse(await fs.readFile(baselinePath, 'utf8'));
|
|
100
|
+
if (loaded.format === 2) baseline = loaded;
|
|
101
|
+
else {
|
|
102
|
+
baselineWarning = `Ignoring legacy timing baseline format ${loaded.format ?? 'unknown'}; regenerate it with npm run benchmark:baseline.`;
|
|
103
|
+
baselinePath = null;
|
|
104
|
+
}
|
|
105
|
+
}
|
|
106
|
+
|
|
107
|
+
function median(values) {
|
|
108
|
+
const sorted = [...values].sort((a, b) => a - b);
|
|
109
|
+
const middle = Math.floor(sorted.length / 2);
|
|
110
|
+
return sorted.length % 2 === 0 ? (sorted[middle - 1] + sorted[middle]) / 2 : sorted[middle];
|
|
111
|
+
}
|
|
112
|
+
|
|
113
|
+
function formatMs(value) {
|
|
114
|
+
if (value == null) return '—';
|
|
115
|
+
return `${value.toFixed(1)} ms`;
|
|
116
|
+
}
|
|
117
|
+
|
|
118
|
+
function changePercent(medianMs, baselineMs) {
|
|
119
|
+
if (baselineMs == null || baselineMs === 0) return null;
|
|
120
|
+
return ((medianMs - baselineMs) / baselineMs) * 100;
|
|
121
|
+
}
|
|
122
|
+
|
|
123
|
+
function changeText(item) {
|
|
124
|
+
if (item.changePercent == null) return '—';
|
|
125
|
+
const change = item.changePercent;
|
|
126
|
+
const value = `${change >= 0 ? '+' : ''}${change.toFixed(1)}%`;
|
|
127
|
+
if (Math.abs(change) < 5) return `≈ ${value}`;
|
|
128
|
+
return `${change < 0 ? '↓' : '↑'} ${value}`;
|
|
129
|
+
}
|
|
130
|
+
|
|
131
|
+
function runWorker(item) {
|
|
132
|
+
const child = spawnSync(process.execPath, [
|
|
133
|
+
path.join(root, 'test', 'benchmark-worker.mjs'),
|
|
134
|
+
path.join(root, item.file),
|
|
135
|
+
JSON.stringify(item.goals),
|
|
136
|
+
String(options.runs),
|
|
137
|
+
String(options.warmup),
|
|
138
|
+
String(options.targetMs),
|
|
139
|
+
], { cwd: root, encoding: 'utf8', maxBuffer: 10 * 1024 * 1024, timeout: 120_000 });
|
|
140
|
+
if (child.error) throw new Error(`${item.name} failed: ${child.error.message}`);
|
|
141
|
+
if (child.status !== 0) {
|
|
142
|
+
throw new Error(`${item.name} failed${child.stderr ? `: ${child.stderr.trim()}` : ''}`);
|
|
143
|
+
}
|
|
144
|
+
try { return JSON.parse(child.stdout); }
|
|
145
|
+
catch (error) { throw new Error(`${item.name} returned invalid worker output: ${error.message}`); }
|
|
146
|
+
}
|
|
147
|
+
|
|
148
|
+
const results = [];
|
|
149
|
+
for (const item of selected) {
|
|
150
|
+
if (!options.json) process.stderr.write(`benchmark ${item.name}...\n`);
|
|
151
|
+
const worker = runWorker(item);
|
|
152
|
+
if (worker.digest !== item.expectedSha256) {
|
|
153
|
+
throw new Error(`${item.name} semantic checksum changed: expected ${item.expectedSha256}, got ${worker.digest}`);
|
|
154
|
+
}
|
|
155
|
+
const medianMs = median(worker.samplesMs);
|
|
156
|
+
const minMs = Math.min(...worker.samplesMs);
|
|
157
|
+
const maxMs = Math.max(...worker.samplesMs);
|
|
158
|
+
const baselineItem = baseline?.benchmarks?.[item.name] ?? null;
|
|
159
|
+
results.push({
|
|
160
|
+
name: item.name,
|
|
161
|
+
group: item.group,
|
|
162
|
+
file: item.file,
|
|
163
|
+
goals: item.goals,
|
|
164
|
+
medianMs,
|
|
165
|
+
minMs,
|
|
166
|
+
maxMs,
|
|
167
|
+
batchSize: worker.batchSize,
|
|
168
|
+
calibrationMs: worker.calibrationMs,
|
|
169
|
+
batchSamplesMs: worker.batchSamplesMs,
|
|
170
|
+
samplesMs: worker.samplesMs,
|
|
171
|
+
answerLines: worker.answerLines,
|
|
172
|
+
outputBytes: worker.outputBytes,
|
|
173
|
+
sha256: worker.digest,
|
|
174
|
+
baselineMs: baselineItem?.medianMs ?? null,
|
|
175
|
+
changePercent: changePercent(medianMs, baselineItem?.medianMs ?? null),
|
|
176
|
+
});
|
|
177
|
+
}
|
|
178
|
+
|
|
179
|
+
if (options.save != null) {
|
|
180
|
+
await fs.mkdir(path.dirname(options.save), { recursive: true });
|
|
181
|
+
const saved = {
|
|
182
|
+
format: 2,
|
|
183
|
+
generatedAt: new Date().toISOString(),
|
|
184
|
+
node: process.version,
|
|
185
|
+
platform: `${process.platform}-${process.arch}`,
|
|
186
|
+
runs: options.runs,
|
|
187
|
+
warmup: options.warmup,
|
|
188
|
+
targetMs: options.targetMs,
|
|
189
|
+
benchmarks: Object.fromEntries(results.map((item) => [item.name, {
|
|
190
|
+
medianMs: item.medianMs,
|
|
191
|
+
batchSize: item.batchSize,
|
|
192
|
+
sha256: item.sha256,
|
|
193
|
+
}])),
|
|
194
|
+
};
|
|
195
|
+
await fs.writeFile(options.save, `${JSON.stringify(saved, null, 2)}\n`);
|
|
196
|
+
}
|
|
197
|
+
|
|
198
|
+
if (options.json) {
|
|
199
|
+
process.stdout.write(`${JSON.stringify({
|
|
200
|
+
runs: options.runs,
|
|
201
|
+
warmup: options.warmup,
|
|
202
|
+
targetMs: options.targetMs,
|
|
203
|
+
baseline: baselinePath,
|
|
204
|
+
baselineWarning,
|
|
205
|
+
results,
|
|
206
|
+
}, null, 2)}\n`);
|
|
207
|
+
} else {
|
|
208
|
+
const headers = ['Benchmark', 'Median/op', 'Range/op', 'Batch', 'Baseline', 'Change', 'Answers'];
|
|
209
|
+
const rows = results.map((item) => [
|
|
210
|
+
item.name,
|
|
211
|
+
formatMs(item.medianMs),
|
|
212
|
+
`${formatMs(item.minMs)}–${formatMs(item.maxMs)}`,
|
|
213
|
+
String(item.batchSize),
|
|
214
|
+
formatMs(item.baselineMs),
|
|
215
|
+
changeText(item),
|
|
216
|
+
String(item.answerLines),
|
|
217
|
+
]);
|
|
218
|
+
const widths = headers.map((header, index) => Math.max(header.length, ...rows.map((row) => row[index].length)));
|
|
219
|
+
const printRow = (row) => process.stdout.write(`${row.map((cell, index) => cell.padEnd(widths[index])).join(' ')}\n`);
|
|
220
|
+
printRow(headers);
|
|
221
|
+
printRow(widths.map((width) => '-'.repeat(width)));
|
|
222
|
+
for (const row of rows) printRow(row);
|
|
223
|
+
process.stdout.write(`\n${results.length} benchmarks; ${options.runs} measured batch${options.runs === 1 ? '' : 'es'} each after ${options.warmup} warm-up batch${options.warmup === 1 ? '' : 'es'}, calibrated after one priming execution toward ${options.targetMs} ms per batch.\n`);
|
|
224
|
+
if (baselineWarning) process.stdout.write(`${baselineWarning}\n`);
|
|
225
|
+
if (baselinePath == null && !baselineWarning) process.stdout.write('No timing baseline found; run npm run benchmark:baseline to create .benchmarks/baseline.json.\n');
|
|
226
|
+
if (options.save != null) process.stdout.write(`Saved timing baseline: ${path.relative(root, options.save)}\n`);
|
|
227
|
+
process.stdout.write('Change compares the current median/op directly with the saved baseline median/op; the measured range is shown separately.\n');
|
|
228
|
+
}
|
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
import fs from 'node:fs/promises';
|
|
2
|
+
import path from 'node:path';
|
|
3
|
+
import process from 'node:process';
|
|
4
|
+
import { spawn } from 'node:child_process';
|
|
5
|
+
import { fileURLToPath } from 'node:url';
|
|
6
|
+
|
|
7
|
+
const root = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '..');
|
|
8
|
+
const manifest = JSON.parse(await fs.readFile(path.join(root, 'test', 'bench', 'benchmarks.json'), 'utf8'));
|
|
9
|
+
|
|
10
|
+
let passed = 0;
|
|
11
|
+
function ok(condition, message) {
|
|
12
|
+
if (!condition) throw new Error(message);
|
|
13
|
+
passed++;
|
|
14
|
+
}
|
|
15
|
+
|
|
16
|
+
function spawnJson(args) {
|
|
17
|
+
return new Promise((resolve, reject) => {
|
|
18
|
+
const child = spawn(process.execPath, args, { cwd: root, stdio: ['ignore', 'pipe', 'pipe'] });
|
|
19
|
+
let stdout = '';
|
|
20
|
+
let stderr = '';
|
|
21
|
+
child.stdout.setEncoding('utf8');
|
|
22
|
+
child.stderr.setEncoding('utf8');
|
|
23
|
+
child.stdout.on('data', (chunk) => { stdout += chunk; });
|
|
24
|
+
child.stderr.on('data', (chunk) => { stderr += chunk; });
|
|
25
|
+
child.on('error', reject);
|
|
26
|
+
child.on('close', (code) => {
|
|
27
|
+
if (code !== 0) return reject(new Error(stderr || `child exited ${code}`));
|
|
28
|
+
try { resolve(JSON.parse(stdout)); }
|
|
29
|
+
catch (error) { reject(new Error(`invalid child JSON: ${error.message}`)); }
|
|
30
|
+
});
|
|
31
|
+
});
|
|
32
|
+
}
|
|
33
|
+
|
|
34
|
+
function runWorker(item) {
|
|
35
|
+
return spawnJson([
|
|
36
|
+
path.join(root, 'test', 'benchmark-worker.mjs'),
|
|
37
|
+
path.join(root, item.file),
|
|
38
|
+
JSON.stringify(item.goals),
|
|
39
|
+
'1',
|
|
40
|
+
'0',
|
|
41
|
+
'0',
|
|
42
|
+
]);
|
|
43
|
+
}
|
|
44
|
+
|
|
45
|
+
ok(Array.isArray(manifest) && manifest.length === 20, 'benchmark manifest should contain exactly 20 representative workloads');
|
|
46
|
+
ok(new Set(manifest.map((item) => item.name)).size === manifest.length, 'benchmark names should be unique');
|
|
47
|
+
|
|
48
|
+
for (const item of manifest) {
|
|
49
|
+
ok(/^[0-9a-f]{64}$/.test(item.expectedSha256), `${item.name} should have a committed semantic checksum`);
|
|
50
|
+
await fs.access(path.join(root, item.file));
|
|
51
|
+
passed++;
|
|
52
|
+
|
|
53
|
+
const worker = await runWorker(item);
|
|
54
|
+
ok(worker.digest === item.expectedSha256, `${item.name} should preserve its committed semantic checksum`);
|
|
55
|
+
ok(worker.batchSize === 1, `${item.name} digest test should use one execution per batch`);
|
|
56
|
+
ok(Array.isArray(worker.samplesMs) && worker.samplesMs.length === 1, `${item.name} should return one requested sample`);
|
|
57
|
+
ok(worker.samplesMs[0] >= 0, `${item.name} should return a non-negative wall time`);
|
|
58
|
+
}
|
|
59
|
+
|
|
60
|
+
const adaptive = await spawnJson([
|
|
61
|
+
path.join(root, 'test', 'benchmark.mjs'),
|
|
62
|
+
'--filter', 'dcg-command',
|
|
63
|
+
'--runs', '1',
|
|
64
|
+
'--warmup', '0',
|
|
65
|
+
'--target-ms', '50',
|
|
66
|
+
'--json',
|
|
67
|
+
]);
|
|
68
|
+
ok(adaptive.results.length === 1, 'adaptive benchmark smoke test should select one workload');
|
|
69
|
+
ok(adaptive.results[0].batchSize > 1, 'adaptive benchmark smoke test should batch a short workload');
|
|
70
|
+
ok(adaptive.results[0].sha256 === manifest.find((item) => item.name === 'dcg-command').expectedSha256,
|
|
71
|
+
'adaptive batching should preserve the semantic checksum');
|
|
72
|
+
|
|
73
|
+
process.stdout.write(`Benchmark harness tests: ${passed}/${passed} passed.\n`);
|
package/the-art-of-eyeprolog.md
CHANGED
|
@@ -3147,7 +3147,12 @@ Before replacing one definition with another, record:
|
|
|
3147
3147
|
|
|
3148
3148
|
Then compare both versions. `--stats` can show fewer calls or unifications, but
|
|
3149
3149
|
performance evidence comes after semantic evidence. A faster program that
|
|
3150
|
-
silently drops a mode is a different program.
|
|
3150
|
+
silently drops a mode is a different program. Predicate and inference counts are
|
|
3151
|
+
also not reliable substitutes for elapsed time: one expensive host call can cost
|
|
3152
|
+
more than thousands of cheap Prolog calls. Repository-level performance work
|
|
3153
|
+
therefore uses `npm run benchmark`, which measures median wall-clock parse+execute
|
|
3154
|
+
time over representative workloads and rejects any run whose answer checksum no
|
|
3155
|
+
longer matches the committed result.
|
|
3151
3156
|
|
|
3152
3157
|
**Exercises.**
|
|
3153
3158
|
|
|
@@ -7858,6 +7863,30 @@ with:
|
|
|
7858
7863
|
npm test
|
|
7859
7864
|
```
|
|
7860
7865
|
|
|
7866
|
+
Performance is checked separately so ordinary correctness tests stay deterministic
|
|
7867
|
+
and fast:
|
|
7868
|
+
|
|
7869
|
+
```sh
|
|
7870
|
+
npm run benchmark
|
|
7871
|
+
npm run benchmark:baseline
|
|
7872
|
+
```
|
|
7873
|
+
|
|
7874
|
+
The benchmark suite contains 20 representative workloads and stores their
|
|
7875
|
+
semantic output digests in the repository, while wall-clock baselines remain
|
|
7876
|
+
machine-local under `.benchmarks/` because absolute timings are machine-specific.
|
|
7877
|
+
Each benchmark runs in its own fresh Node worker. Inside that worker, one untimed
|
|
7878
|
+
execution primes parser, module, and JIT state; short workloads are then repeated
|
|
7879
|
+
with independent `run()` calls until a batch is roughly 400 ms long. After one
|
|
7880
|
+
warm-up batch, five measured batches are reported as milliseconds per workload
|
|
7881
|
+
execution. Naturally long workloads keep a batch size of one.
|
|
7882
|
+
|
|
7883
|
+
The report shows the median, per-operation range, chosen batch size, saved
|
|
7884
|
+
baseline median, and the percentage change between the current median and baseline
|
|
7885
|
+
median. The range remains visible as context, but it does not suppress or reinterpret
|
|
7886
|
+
the median-to-median comparison. Older unbatched baseline files are ignored with a
|
|
7887
|
+
request to regenerate them. This keeps the rule simple: unchanged answers first,
|
|
7888
|
+
sufficiently long wall-clock samples second, and direct median-versus-median change.
|
|
7889
|
+
|
|
7861
7890
|
When adding an example:
|
|
7862
7891
|
|
|
7863
7892
|
1. choose a filename that names the mathematical or domain idea;
|