rigorrun 0.2.0 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +1 -1
- package/README.md +83 -96
- package/dist/rigorrun.mjs +3438 -875
- package/package.json +1 -1
- package/ui/assets/DemoPage-BgHJFGif.js +303 -0
- package/ui/assets/index-B0Grxuu3.js +14 -0
- package/ui/assets/index-BnnN5Oot.css +2 -0
- package/ui/favicon.svg +8 -0
- package/ui/fonts/geist-mono-latin.woff2 +0 -0
- package/ui/fonts/instrument-sans-latin-italic.woff2 +0 -0
- package/ui/fonts/instrument-sans-latin.woff2 +0 -0
- package/ui/index.html +20 -28
- package/ui/assets/DemoPage-DvY0tdpJ.js +0 -295
- package/ui/assets/Evidence-B65d5xlS.js +0 -1
- package/ui/assets/index-CziJAnNn.css +0 -2
- package/ui/assets/index-DEjJuxpK.js +0 -30
- package/ui/robots.txt +0 -8
package/LICENSE
CHANGED
package/README.md
CHANGED
|
@@ -1,135 +1,122 @@
|
|
|
1
|
+
<div align="center">
|
|
2
|
+
|
|
1
3
|
# RigorRun
|
|
2
4
|
|
|
3
|
-
**
|
|
5
|
+
**Your agent said it worked. RigorRun checks what it actually did.**
|
|
6
|
+
|
|
7
|
+
Show RigorRun a job once. It turns that into a repeatable acceptance suite and decides whether your
|
|
8
|
+
agent is safe to ship by reading the system it changed — never by trusting what it says about itself.
|
|
9
|
+
|
|
10
|
+
**Early Access · v0.2** — parts of it are honestly unfinished, and they are listed rather than hidden.
|
|
11
|
+
|
|
12
|
+
[rigorrun.xyz](https://rigorrun.xyz) · [Documentation](https://docs.rigorrun.xyz) · [Evidence](https://rigorrun.xyz/evidence) · [What is and is not built](https://rigorrun.xyz/what-is-built)
|
|
4
13
|
|
|
5
|
-
|
|
6
|
-
RigorRun proves whether the agent can do that job safely — by reading the
|
|
7
|
-
system it changed, never by trusting what it says about itself.
|
|
14
|
+
</div>
|
|
8
15
|
|
|
9
16
|
```bash
|
|
10
17
|
npx rigorrun
|
|
11
18
|
```
|
|
12
19
|
|
|
13
|
-
Open the URL it prints.
|
|
20
|
+
Open the URL it prints. Everything runs on your machine: there is no account, and no hosted
|
|
21
|
+
component to send your systems to.
|
|
14
22
|
|
|
15
|
-
|
|
16
|
-
limits are written down and marked one by one, rather than left for you to
|
|
17
|
-
find: [the v1 gap audit](https://github.com/Konuktor/rigorrun/blob/master/docs/V1_GAP_AUDIT.md).
|
|
23
|
+
---
|
|
18
24
|
|
|
19
|
-
##
|
|
25
|
+
## Why
|
|
20
26
|
|
|
21
|
-
An
|
|
27
|
+
An agent that reports success and an agent that achieved it are indistinguishable from the
|
|
28
|
+
transcript. They are trivially distinguishable from the database.
|
|
29
|
+
|
|
30
|
+
Evaluation harnesses score what the model wrote. RigorRun compares system state before and after,
|
|
31
|
+
through read operations you nominate, and asks the questions that have a fact behind them: does the
|
|
32
|
+
refund exist, is the amount right, is the ticket attached, was approval required, and did anything
|
|
33
|
+
change that should not have.
|
|
22
34
|
|
|
23
|
-
```
|
|
24
|
-
|
|
35
|
+
```
|
|
36
|
+
most tools: you write the tests → the tool runs them
|
|
37
|
+
RigorRun: you do the job once → RigorRun writes the tests
|
|
25
38
|
```
|
|
26
39
|
|
|
27
|
-
|
|
28
|
-
to the exact bytes the registry published, runs it in a container with no
|
|
29
|
-
network and no access to your machine, calls each tool with arguments derived
|
|
30
|
-
from its own schema, and reads the filesystem before and after to see what
|
|
31
|
-
actually changed — then compares that against what the server declared.
|
|
40
|
+
## How
|
|
32
41
|
|
|
33
|
-
|
|
34
|
-
|
|
42
|
+
1. **Connect your system** — an MCP server you already run, an OpenAPI document, or a web
|
|
43
|
+
application through a browser.
|
|
44
|
+
2. **Do the job once.** RigorRun reads your system before and after and derives what the rules must
|
|
45
|
+
be.
|
|
46
|
+
3. **Rule on what it worked out.** It shows the evidence behind each proposed rule. A rule you
|
|
47
|
+
reject cannot fail your agent.
|
|
48
|
+
4. **Connect your agent** — an HTTP endpoint, a local command, or your own loop pulling work.
|
|
49
|
+
5. **Run it.** A verdict, with how strongly each answer could be verified.
|
|
50
|
+
6. **Gate the next change** in CI.
|
|
35
51
|
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
tools**; the rest are named with reasons in every record it writes.
|
|
52
|
+
Setting a project up happens in the local interface. Running, gating and comparing are also
|
|
53
|
+
available from the command line, which is the half CI needs.
|
|
39
54
|
|
|
40
|
-
|
|
55
|
+
```bash
|
|
56
|
+
rigorrun gate --project <id> # exit 1 stops the build
|
|
57
|
+
```
|
|
41
58
|
|
|
42
|
-
|
|
43
|
-
job in your real system without doing something unsafe — and you need to know
|
|
44
|
-
again next week, after somebody changes a prompt.
|
|
59
|
+
## Verify a published MCP server
|
|
45
60
|
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
the answers into an executable acceptance suite. Then it runs your agent against
|
|
49
|
-
it and reads your system to find out what actually happened.
|
|
61
|
+
An MCP server can annotate a tool `readOnlyHint: true`. Nothing checks that, and agents use it to
|
|
62
|
+
decide whether a tool may be called without asking.
|
|
50
63
|
|
|
64
|
+
```bash
|
|
65
|
+
rigorrun verify npm:@modelcontextprotocol/server-memory@2026.8.31
|
|
51
66
|
```
|
|
52
|
-
most tools: you write the tests → the tool runs them
|
|
53
|
-
RigorRun: you do the job once → RigorRun writes the tests
|
|
54
|
-
```
|
|
55
|
-
|
|
56
|
-
## How strongly was it verified?
|
|
57
67
|
|
|
58
|
-
|
|
68
|
+
Pins the server to the exact bytes the registry published, runs it in a container with no network
|
|
69
|
+
and no access to your machine, calls each tool with arguments derived from its own schema, and reads
|
|
70
|
+
the filesystem before and after. Needs a container runtime; nothing else does.
|
|
59
71
|
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
connected MCP server, whose state RigorRun reads back through the tools you
|
|
63
|
-
nominated, is **PARTIAL** — the common, honest case, not a defect.
|
|
64
|
-
- **OBSERVATIONAL** — actions were observed but the final state could not be
|
|
65
|
-
independently proven (e.g. a browser with nothing readable attached).
|
|
72
|
+
Run against four published servers, it exercised **19 of 37** tools. The other 18 are named, each
|
|
73
|
+
with the reason it was not reached — [see the records](https://rigorrun.xyz/evidence).
|
|
66
74
|
|
|
67
|
-
|
|
68
|
-
access. Against your own system the honest label is usually `PARTIAL`, and it is
|
|
69
|
-
shown on the same line as the verdict.
|
|
75
|
+
## What you should know before relying on it
|
|
70
76
|
|
|
71
|
-
|
|
77
|
+
- **A verdict against a real system is `PARTIAL`, by design.** RigorRun reads back what your
|
|
78
|
+
nominated reads return and no more, and says so on every result.
|
|
79
|
+
- **A browser cannot verify itself.** A page saying "done" is a claim by the system that would have
|
|
80
|
+
to be wrong for it not to be done, so a browser connection is `OBSERVATIONAL`. Attach a readable
|
|
81
|
+
API or MCP connection for the same system to strengthen it.
|
|
82
|
+
- **Isolation is `DECLARED`, not `RESET`, when you nominate a reset tool** — RigorRun has not run it
|
|
83
|
+
twice and compared.
|
|
84
|
+
- **Three of the six verification sources are not emitted yet.** Model-judged and human-review
|
|
85
|
+
evaluators exist in the schema and are unreachable in practice.
|
|
86
|
+
- **No external team has used this yet.** There are no customers.
|
|
72
87
|
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
system without a reset still works, but produces fewer cases, disables repeated
|
|
77
|
-
mutating cases, reports isolation `NONE`, and verifies `PARTIAL`. Best results:
|
|
78
|
-
a staging or scratch environment with read-back and a reset.
|
|
88
|
+
The full list is at [rigorrun.xyz/what-is-built](https://rigorrun.xyz/what-is-built), and the
|
|
89
|
+
version carrying how each line was checked is
|
|
90
|
+
[docs/V1_GAP_AUDIT.md](https://github.com/Konuktor/rigorrun/blob/master/docs/V1_GAP_AUDIT.md).
|
|
79
91
|
|
|
80
|
-
##
|
|
92
|
+
## Privacy
|
|
81
93
|
|
|
82
|
-
Your
|
|
83
|
-
|
|
84
|
-
`http://127.0.0.1`. So RigorRun's interface is served by this process, on your
|
|
85
|
-
machine. Your credentials, recordings and systems never touch anybody's
|
|
86
|
-
infrastructure, because there is no path by which they could.
|
|
94
|
+
Your systems, your credentials and your recordings stay on this machine. There is no account, and
|
|
95
|
+
nothing is uploaded unless you ask for it.
|
|
87
96
|
|
|
88
|
-
|
|
97
|
+
Three things can make a network request, all because you asked: `rigorrun verify` downloads the named
|
|
98
|
+
package from the npm registry; connecting your own system sends requests to your own system; and an
|
|
99
|
+
LLM-backed agent, which exists only if you set a provider key, talks to that provider. None of them
|
|
100
|
+
reach RigorRun. [How it is built](https://rigorrun.xyz/security).
|
|
89
101
|
|
|
90
|
-
|
|
91
|
-
- **A way in to the system you want to test**: an MCP server, or an OpenAPI
|
|
92
|
-
document and the address it is served from. Ideally staging or a scratch
|
|
93
|
-
instance, with a way to reset it.
|
|
94
|
-
- **An agent.** If it speaks MCP it works unchanged; RigorRun hands it a URL —
|
|
95
|
-
whether your agent listens on an address or is a command RigorRun runs. If it
|
|
96
|
-
does not speak MCP, about ten lines of plain HTTP — there is no package to
|
|
97
|
-
install; the protocol is documented at
|
|
98
|
-
[docs/HTTP_AGENT.md](https://github.com/Konuktor/rigorrun/blob/master/docs/HTTP_AGENT.md).
|
|
102
|
+
## Verifying what you installed
|
|
99
103
|
|
|
100
|
-
|
|
104
|
+
Published by a GitHub Actions workflow that holds no npm token, using npm trusted publishing, with a
|
|
105
|
+
SLSA build provenance attestation.
|
|
101
106
|
|
|
102
107
|
```bash
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
npx rigorrun doctor # check this machine and every project
|
|
106
|
-
npx rigorrun projects # what is on this machine
|
|
107
|
-
npx rigorrun run --project <id> # run the suite
|
|
108
|
-
npx rigorrun gate --project <id> # run it, and exit non-zero if it misses the bar
|
|
109
|
-
npx rigorrun compare-runs --project <id> <runId>
|
|
110
|
-
npx rigorrun feedback export # a sanitised bundle for a bug report
|
|
108
|
+
npm audit signatures
|
|
109
|
+
npm view rigorrun dist.integrity
|
|
111
110
|
```
|
|
112
111
|
|
|
113
|
-
##
|
|
112
|
+
## Requirements
|
|
114
113
|
|
|
115
|
-
|
|
116
|
-
applications through a browser — though a browser cannot verify itself, so a
|
|
117
|
-
verdict from one is OBSERVATIONAL unless something readable is attached.
|
|
118
|
-
Setting a project up needs the interface; running and gating it does not. It has
|
|
119
|
-
been used successfully by the people who wrote it and is now looking for people
|
|
120
|
-
who did not.
|
|
114
|
+
Node 20.11 or newer. Docker only for `rigorrun verify`.
|
|
121
115
|
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
that matters.
|
|
126
|
-
|
|
127
|
-
If it goes wrong, `npx rigorrun feedback export` produces a bundle that contains
|
|
128
|
-
no credentials, no tool arguments and no results — only what is needed to work
|
|
129
|
-
out where it broke.
|
|
130
|
-
|
|
131
|
-
## Documentation
|
|
116
|
+
```bash
|
|
117
|
+
rigorrun doctor # what this machine has, and what it does not
|
|
118
|
+
```
|
|
132
119
|
|
|
133
|
-
|
|
120
|
+
---
|
|
134
121
|
|
|
135
|
-
MIT
|
|
122
|
+
MIT · Built by Erbol Tahirov · [Report a problem](https://github.com/Konuktor/rigorrun/issues)
|