rigorrun 0.1.0 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +51 -3
- package/dist/rigorrun.mjs +1930 -265
- package/package.json +2 -2
- package/ui/assets/DemoPage-DvY0tdpJ.js +295 -0
- package/ui/assets/Evidence-B65d5xlS.js +1 -0
- package/ui/assets/index-CziJAnNn.css +2 -0
- package/ui/assets/index-DEjJuxpK.js +30 -0
- package/ui/index.html +25 -2
- package/ui/robots.txt +8 -0
- package/ui/assets/DemoPage-B1s79hcz.js +0 -295
- package/ui/assets/Proof-Doh3bIfF.js +0 -1
- package/ui/assets/index-CR8r1ZFu.css +0 -2
- package/ui/assets/index-DAiDcuMt.js +0 -14
package/README.md
CHANGED
|
@@ -12,10 +12,31 @@ npx rigorrun
|
|
|
12
12
|
|
|
13
13
|
Open the URL it prints. That is the whole install.
|
|
14
14
|
|
|
15
|
-
**Early Access · v0.
|
|
15
|
+
**Early Access · v0.2.** It does what this page says and it is young. The
|
|
16
16
|
limits are written down and marked one by one, rather than left for you to
|
|
17
17
|
find: [the v1 gap audit](https://github.com/Konuktor/rigorrun/blob/master/docs/V1_GAP_AUDIT.md).
|
|
18
18
|
|
|
19
|
+
## Or start with a server you already use
|
|
20
|
+
|
|
21
|
+
An MCP server can annotate a tool `readOnlyHint: true`. Nothing checks that.
|
|
22
|
+
|
|
23
|
+
```bash
|
|
24
|
+
npx rigorrun verify npm:@modelcontextprotocol/server-memory@2026.8.31
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
No project, no browser, no agent, nothing to configure first. It pins the server
|
|
28
|
+
to the exact bytes the registry published, runs it in a container with no
|
|
29
|
+
network and no access to your machine, calls each tool with arguments derived
|
|
30
|
+
from its own schema, and reads the filesystem before and after to see what
|
|
31
|
+
actually changed — then compares that against what the server declared.
|
|
32
|
+
|
|
33
|
+
Exit `0` verified · `1` a declaration was contradicted · `2` it could not run ·
|
|
34
|
+
`3` it ran and established too little to be worth much.
|
|
35
|
+
|
|
36
|
+
**Needs Docker.** Only `npm:` and `dir:` references, and only stdio servers.
|
|
37
|
+
Against four published servers nobody here wrote, it exercised **19 of 37
|
|
38
|
+
tools**; the rest are named with reasons in every record it writes.
|
|
39
|
+
|
|
19
40
|
## What it does
|
|
20
41
|
|
|
21
42
|
You have an agent that calls tools. You need to know whether it can do a real
|
|
@@ -32,6 +53,30 @@ it and reads your system to find out what actually happened.
|
|
|
32
53
|
RigorRun: you do the job once → RigorRun writes the tests
|
|
33
54
|
```
|
|
34
55
|
|
|
56
|
+
## How strongly was it verified?
|
|
57
|
+
|
|
58
|
+
RigorRun tells you how strongly each result was verified, on every result:
|
|
59
|
+
|
|
60
|
+
- **AUTHORITATIVE** — checked against direct, trusted state.
|
|
61
|
+
- **PARTIAL** — verified through the reads your system exposes. A normal
|
|
62
|
+
connected MCP server, whose state RigorRun reads back through the tools you
|
|
63
|
+
nominated, is **PARTIAL** — the common, honest case, not a defect.
|
|
64
|
+
- **OBSERVATIONAL** — actions were observed but the final state could not be
|
|
65
|
+
independently proven (e.g. a browser with nothing readable attached).
|
|
66
|
+
|
|
67
|
+
`AUTHORITATIVE` is claimed only where RigorRun genuinely has authoritative state
|
|
68
|
+
access. Against your own system the honest label is usually `PARTIAL`, and it is
|
|
69
|
+
shown on the same line as the verdict.
|
|
70
|
+
|
|
71
|
+
## What it can build depends on your system
|
|
72
|
+
|
|
73
|
+
RigorRun generates every case it can safely and reproducibly verify, and tells
|
|
74
|
+
you what it could not test. A system it can **seed and reset** yields the
|
|
75
|
+
richest suite (boundary and adversarial cases, repeated destructive checks). A
|
|
76
|
+
system without a reset still works, but produces fewer cases, disables repeated
|
|
77
|
+
mutating cases, reports isolation `NONE`, and verifies `PARTIAL`. Best results:
|
|
78
|
+
a staging or scratch environment with read-back and a reset.
|
|
79
|
+
|
|
35
80
|
## Why it runs locally
|
|
36
81
|
|
|
37
82
|
Your MCP server, your internal API and your staging box are usually not
|
|
@@ -48,12 +93,15 @@ infrastructure, because there is no path by which they could.
|
|
|
48
93
|
instance, with a way to reset it.
|
|
49
94
|
- **An agent.** If it speaks MCP it works unchanged; RigorRun hands it a URL —
|
|
50
95
|
whether your agent listens on an address or is a command RigorRun runs. If it
|
|
51
|
-
does not speak MCP, about ten lines of
|
|
96
|
+
does not speak MCP, about ten lines of plain HTTP — there is no package to
|
|
97
|
+
install; the protocol is documented at
|
|
98
|
+
[docs/HTTP_AGENT.md](https://github.com/Konuktor/rigorrun/blob/master/docs/HTTP_AGENT.md).
|
|
52
99
|
|
|
53
100
|
## Commands
|
|
54
101
|
|
|
55
102
|
```bash
|
|
56
103
|
npx rigorrun # start the runner and open the interface
|
|
104
|
+
npx rigorrun verify <server-ref> # what does this server's tools actually do?
|
|
57
105
|
npx rigorrun doctor # check this machine and every project
|
|
58
106
|
npx rigorrun projects # what is on this machine
|
|
59
107
|
npx rigorrun run --project <id> # run the suite
|
|
@@ -71,7 +119,7 @@ Setting a project up needs the interface; running and gating it does not. It has
|
|
|
71
119
|
been used successfully by the people who wrote it and is now looking for people
|
|
72
120
|
who did not.
|
|
73
121
|
|
|
74
|
-
Every capability is marked WORKING, PARTIAL or MISSING in
|
|
122
|
+
Every capability is marked WORKING, PARTIAL, DEMO-ONLY, BROKEN or MISSING in
|
|
75
123
|
[the v1 gap audit](https://github.com/Konuktor/rigorrun/blob/master/docs/V1_GAP_AUDIT.md),
|
|
76
124
|
with how each one was checked. Read it before you rely on this for anything
|
|
77
125
|
that matters.
|