rigorrun 0.1.0 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -12,10 +12,31 @@ npx rigorrun
12
12
 
13
13
  Open the URL it prints. That is the whole install.
14
14
 
15
- **Early Access · v0.1.** It does what this page says and it is young. The
15
+ **Early Access · v0.2.** It does what this page says and it is young. The
16
16
  limits are written down and marked one by one, rather than left for you to
17
17
  find: [the v1 gap audit](https://github.com/Konuktor/rigorrun/blob/master/docs/V1_GAP_AUDIT.md).
18
18
 
19
+ ## Or start with a server you already use
20
+
21
+ An MCP server can annotate a tool `readOnlyHint: true`. Nothing checks that.
22
+
23
+ ```bash
24
+ npx rigorrun verify npm:@modelcontextprotocol/server-memory@2026.8.31
25
+ ```
26
+
27
+ No project, no browser, no agent, nothing to configure first. It pins the server
28
+ to the exact bytes the registry published, runs it in a container with no
29
+ network and no access to your machine, calls each tool with arguments derived
30
+ from its own schema, and reads the filesystem before and after to see what
31
+ actually changed — then compares that against what the server declared.
32
+
33
+ Exit `0` verified · `1` a declaration was contradicted · `2` it could not run ·
34
+ `3` it ran and established too little to be worth much.
35
+
36
+ **Needs Docker.** Only `npm:` and `dir:` references, and only stdio servers.
37
+ Against four published servers nobody here wrote, it exercised **19 of 37
38
+ tools**; the rest are named with reasons in every record it writes.
39
+
19
40
  ## What it does
20
41
 
21
42
  You have an agent that calls tools. You need to know whether it can do a real
@@ -32,6 +53,30 @@ it and reads your system to find out what actually happened.
32
53
  RigorRun: you do the job once → RigorRun writes the tests
33
54
  ```
34
55
 
56
+ ## How strongly was it verified?
57
+
58
+ RigorRun tells you how strongly each result was verified, on every result:
59
+
60
+ - **AUTHORITATIVE** — checked against direct, trusted state.
61
+ - **PARTIAL** — verified through the reads your system exposes. A normal
62
+ connected MCP server, whose state RigorRun reads back through the tools you
63
+ nominated, is **PARTIAL** — the common, honest case, not a defect.
64
+ - **OBSERVATIONAL** — actions were observed but the final state could not be
65
+ independently proven (e.g. a browser with nothing readable attached).
66
+
67
+ `AUTHORITATIVE` is claimed only where RigorRun genuinely has authoritative state
68
+ access. Against your own system the honest label is usually `PARTIAL`, and it is
69
+ shown on the same line as the verdict.
70
+
71
+ ## What it can build depends on your system
72
+
73
+ RigorRun generates every case it can safely and reproducibly verify, and tells
74
+ you what it could not test. A system it can **seed and reset** yields the
75
+ richest suite (boundary and adversarial cases, repeated destructive checks). A
76
+ system without a reset still works, but produces fewer cases, disables repeated
77
+ mutating cases, reports isolation `NONE`, and verifies `PARTIAL`. Best results:
78
+ a staging or scratch environment with read-back and a reset.
79
+
35
80
  ## Why it runs locally
36
81
 
37
82
  Your MCP server, your internal API and your staging box are usually not
@@ -48,12 +93,15 @@ infrastructure, because there is no path by which they could.
48
93
  instance, with a way to reset it.
49
94
  - **An agent.** If it speaks MCP it works unchanged; RigorRun hands it a URL —
50
95
  whether your agent listens on an address or is a command RigorRun runs. If it
51
- does not speak MCP, about ten lines of the agent SDK.
96
+ does not speak MCP, about ten lines of plain HTTP — there is no package to
97
+ install; the protocol is documented at
98
+ [docs/HTTP_AGENT.md](https://github.com/Konuktor/rigorrun/blob/master/docs/HTTP_AGENT.md).
52
99
 
53
100
  ## Commands
54
101
 
55
102
  ```bash
56
103
  npx rigorrun # start the runner and open the interface
104
+ npx rigorrun verify <server-ref> # what does this server's tools actually do?
57
105
  npx rigorrun doctor # check this machine and every project
58
106
  npx rigorrun projects # what is on this machine
59
107
  npx rigorrun run --project <id> # run the suite
@@ -71,7 +119,7 @@ Setting a project up needs the interface; running and gating it does not. It has
71
119
  been used successfully by the people who wrote it and is now looking for people
72
120
  who did not.
73
121
 
74
- Every capability is marked WORKING, PARTIAL or MISSING in
122
+ Every capability is marked WORKING, PARTIAL, DEMO-ONLY, BROKEN or MISSING in
75
123
  [the v1 gap audit](https://github.com/Konuktor/rigorrun/blob/master/docs/V1_GAP_AUDIT.md),
76
124
  with how each one was checked. Read it before you rely on this for anything
77
125
  that matters.