rigorrun 0.2.0 → 0.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE CHANGED
@@ -1,6 +1,6 @@
1
1
  MIT License
2
2
 
3
- Copyright (c) 2026 RigorRun
3
+ Copyright (c) 2026 Erbol Tahirov
4
4
 
5
5
  Permission is hereby granted, free of charge, to any person obtaining a copy
6
6
  of this software and associated documentation files (the "Software"), to deal
package/README.md CHANGED
@@ -1,135 +1,128 @@
1
+ <div align="center">
2
+
1
3
  # RigorRun
2
4
 
3
- **Acceptance testing for tool-using AI agents.**
5
+ **Your agent said it worked. RigorRun checks what it actually did.**
6
+
7
+ Show RigorRun a job once. It turns that into a repeatable acceptance suite and decides whether your
8
+ agent is safe to ship by reading the system it changed — never by trusting what it says about itself.
9
+
10
+ **Early Access · v0.3** — parts of it are honestly unfinished, and they are listed rather than hidden.
4
11
 
5
- Connect your system. Show RigorRun how one job is done. Connect your agent.
6
- RigorRun proves whether the agent can do that job safely — by reading the
7
- system it changed, never by trusting what it says about itself.
12
+ [rigorrun.xyz](https://rigorrun.xyz) · [Documentation](https://docs.rigorrun.xyz) · [Evidence](https://rigorrun.xyz/evidence) · [What is and is not built](https://rigorrun.xyz/what-is-built)
13
+
14
+ </div>
8
15
 
9
16
  ```bash
10
- npx rigorrun
17
+ npx rigorrun demo # a real recorded run, replayed offline in a second
18
+ npx rigorrun # your own system and agent, in the local interface
11
19
  ```
12
20
 
13
- Open the URL it prints. That is the whole install.
21
+ `demo` replays a real model working a bundled support desk: what it said beside what the system held
22
+ afterwards. The same run, case by case, is at [rigorrun.xyz/replay](https://rigorrun.xyz/replay).
14
23
 
15
- **Early Access · v0.2.** It does what this page says and it is young. The
16
- limits are written down and marked one by one, rather than left for you to
17
- find: [the v1 gap audit](https://github.com/Konuktor/rigorrun/blob/master/docs/V1_GAP_AUDIT.md).
24
+ `npx rigorrun` opens a local interface. Everything runs on your machine: there is no account, and no
25
+ hosted component to send your systems to.
18
26
 
19
- ## Or start with a server you already use
27
+ ---
20
28
 
21
- An MCP server can annotate a tool `readOnlyHint: true`. Nothing checks that.
29
+ ## Why
30
+
31
+ An agent that reports success and an agent that achieved it are indistinguishable from the
32
+ transcript. They are trivially distinguishable from the database.
33
+
34
+ Evaluation harnesses score what the model wrote. RigorRun compares system state before and after,
35
+ through read operations you nominate, and asks the questions that have a fact behind them: does the
36
+ refund exist, is the amount right, is the ticket attached, was approval required, and did anything
37
+ change that should not have.
22
38
 
23
- ```bash
24
- npx rigorrun verify npm:@modelcontextprotocol/server-memory@2026.8.31
39
+ ```
40
+ most tools: you write the tests → the tool runs them
41
+ RigorRun: you do the job once → RigorRun writes the tests
25
42
  ```
26
43
 
27
- No project, no browser, no agent, nothing to configure first. It pins the server
28
- to the exact bytes the registry published, runs it in a container with no
29
- network and no access to your machine, calls each tool with arguments derived
30
- from its own schema, and reads the filesystem before and after to see what
31
- actually changed — then compares that against what the server declared.
44
+ ## How
32
45
 
33
- Exit `0` verified · `1` a declaration was contradicted · `2` it could not run ·
34
- `3` it ran and established too little to be worth much.
46
+ 1. **Connect your system** — an MCP server you already run, an OpenAPI document, or a web
47
+ application through a browser.
48
+ 2. **Do the job once.** RigorRun reads your system before and after and derives what the rules must
49
+ be.
50
+ 3. **Rule on what it worked out.** It shows the evidence behind each proposed rule. A rule you
51
+ reject cannot fail your agent.
52
+ 4. **Connect your agent** — unchanged, wherever it runs: RigorRun sends each case's work to a URL it
53
+ already serves and reads the result from your system (a black box). Or an HTTP endpoint, a local
54
+ command, or your own loop pulling work.
55
+ 5. **Run it.** A verdict, with how strongly each answer could be verified.
56
+ 6. **Gate the next change** in CI.
35
57
 
36
- **Needs Docker.** Only `npm:` and `dir:` references, and only stdio servers.
37
- Against four published servers nobody here wrote, it exercised **19 of 37
38
- tools**; the rest are named with reasons in every record it writes.
58
+ Setting a project up happens in the local interface. Running, gating and comparing are also
59
+ available from the command line, which is the half CI needs.
39
60
 
40
- ## What it does
61
+ ```bash
62
+ rigorrun gate --project <id> # exit 1 stops the build
63
+ ```
41
64
 
42
- You have an agent that calls tools. You need to know whether it can do a real
43
- job in your real system without doing something unsafe — and you need to know
44
- again next week, after somebody changes a prompt.
65
+ ## Verify a published MCP server
45
66
 
46
- RigorRun watches a person do that job once, reads the system before and after,
47
- works out what the rules must be, asks about what it can only guess, and turns
48
- the answers into an executable acceptance suite. Then it runs your agent against
49
- it and reads your system to find out what actually happened.
67
+ An MCP server can annotate a tool `readOnlyHint: true`. Nothing checks that, and agents use it to
68
+ decide whether a tool may be called without asking.
50
69
 
70
+ ```bash
71
+ rigorrun verify npm:@modelcontextprotocol/server-memory@2026.8.31
51
72
  ```
52
- most tools: you write the tests → the tool runs them
53
- RigorRun: you do the job once → RigorRun writes the tests
54
- ```
55
-
56
- ## How strongly was it verified?
57
73
 
58
- RigorRun tells you how strongly each result was verified, on every result:
74
+ Pins the server to the exact bytes the registry published, runs it in a container with no network
75
+ and no access to your machine, calls each tool with arguments derived from its own schema, and reads
76
+ the filesystem before and after. Needs a container runtime; nothing else does.
59
77
 
60
- - **AUTHORITATIVE** — checked against direct, trusted state.
61
- - **PARTIAL** — verified through the reads your system exposes. A normal
62
- connected MCP server, whose state RigorRun reads back through the tools you
63
- nominated, is **PARTIAL** — the common, honest case, not a defect.
64
- - **OBSERVATIONAL** — actions were observed but the final state could not be
65
- independently proven (e.g. a browser with nothing readable attached).
78
+ Run against four published servers, it exercised **19 of 37** tools. The other 18 are named, each
79
+ with the reason it was not reached — [see the records](https://rigorrun.xyz/evidence).
66
80
 
67
- `AUTHORITATIVE` is claimed only where RigorRun genuinely has authoritative state
68
- access. Against your own system the honest label is usually `PARTIAL`, and it is
69
- shown on the same line as the verdict.
81
+ ## What you should know before relying on it
70
82
 
71
- ## What it can build depends on your system
83
+ - **A verdict against a real system is `PARTIAL`, by design.** RigorRun reads back what your
84
+ nominated reads return and no more, and says so on every result.
85
+ - **A browser cannot verify itself.** A page saying "done" is a claim by the system that would have
86
+ to be wrong for it not to be done, so a browser connection is `OBSERVATIONAL`. Attach a readable
87
+ API or MCP connection for the same system to strengthen it.
88
+ - **Isolation is `DECLARED`, not `RESET`, when you nominate a reset tool** — RigorRun has not run it
89
+ twice and compared.
90
+ - **Three of the six verification sources are not emitted yet.** Model-judged and human-review
91
+ evaluators exist in the schema and are unreachable in practice.
92
+ - **No external team has used this yet.** There are no customers.
72
93
 
73
- RigorRun generates every case it can safely and reproducibly verify, and tells
74
- you what it could not test. A system it can **seed and reset** yields the
75
- richest suite (boundary and adversarial cases, repeated destructive checks). A
76
- system without a reset still works, but produces fewer cases, disables repeated
77
- mutating cases, reports isolation `NONE`, and verifies `PARTIAL`. Best results:
78
- a staging or scratch environment with read-back and a reset.
94
+ The full list is at [rigorrun.xyz/what-is-built](https://rigorrun.xyz/what-is-built), and the
95
+ version carrying how each line was checked is
96
+ [docs/V1_GAP_AUDIT.md](https://github.com/Konuktor/rigorrun/blob/master/docs/V1_GAP_AUDIT.md).
79
97
 
80
- ## Why it runs locally
98
+ ## Privacy
81
99
 
82
- Your MCP server, your internal API and your staging box are usually not
83
- reachable from the public internet, and a page served over `https` cannot fetch
84
- `http://127.0.0.1`. So RigorRun's interface is served by this process, on your
85
- machine. Your credentials, recordings and systems never touch anybody's
86
- infrastructure, because there is no path by which they could.
100
+ Your systems, your credentials and your recordings stay on this machine. There is no account, and
101
+ nothing is uploaded unless you ask for it.
87
102
 
88
- ## What you need
103
+ Three things can make a network request, all because you asked: `rigorrun verify` downloads the named
104
+ package from the npm registry; connecting your own system sends requests to your own system; and an
105
+ LLM-backed agent, which exists only if you set a provider key, talks to that provider. None of them
106
+ reach RigorRun. [How it is built](https://rigorrun.xyz/security).
89
107
 
90
- - **Node 20.11 or newer.**
91
- - **A way in to the system you want to test**: an MCP server, or an OpenAPI
92
- document and the address it is served from. Ideally staging or a scratch
93
- instance, with a way to reset it.
94
- - **An agent.** If it speaks MCP it works unchanged; RigorRun hands it a URL —
95
- whether your agent listens on an address or is a command RigorRun runs. If it
96
- does not speak MCP, about ten lines of plain HTTP — there is no package to
97
- install; the protocol is documented at
98
- [docs/HTTP_AGENT.md](https://github.com/Konuktor/rigorrun/blob/master/docs/HTTP_AGENT.md).
108
+ ## Verifying what you installed
99
109
 
100
- ## Commands
110
+ Published by a GitHub Actions workflow that holds no npm token, using npm trusted publishing, with a
111
+ SLSA build provenance attestation.
101
112
 
102
113
  ```bash
103
- npx rigorrun # start the runner and open the interface
104
- npx rigorrun verify <server-ref> # what does this server's tools actually do?
105
- npx rigorrun doctor # check this machine and every project
106
- npx rigorrun projects # what is on this machine
107
- npx rigorrun run --project <id> # run the suite
108
- npx rigorrun gate --project <id> # run it, and exit non-zero if it misses the bar
109
- npx rigorrun compare-runs --project <id> <runId>
110
- npx rigorrun feedback export # a sanitised bundle for a bug report
114
+ npm audit signatures
115
+ npm view rigorrun dist.integrity
111
116
  ```
112
117
 
113
- ## This is early access
118
+ ## Requirements
114
119
 
115
- It connects to MCP servers, to HTTP APIs with an OpenAPI document, and to web
116
- applications through a browser — though a browser cannot verify itself, so a
117
- verdict from one is OBSERVATIONAL unless something readable is attached.
118
- Setting a project up needs the interface; running and gating it does not. It has
119
- been used successfully by the people who wrote it and is now looking for people
120
- who did not.
120
+ Node 20.11 or newer. Docker only for `rigorrun verify`.
121
121
 
122
- Every capability is marked WORKING, PARTIAL, DEMO-ONLY, BROKEN or MISSING in
123
- [the v1 gap audit](https://github.com/Konuktor/rigorrun/blob/master/docs/V1_GAP_AUDIT.md),
124
- with how each one was checked. Read it before you rely on this for anything
125
- that matters.
126
-
127
- If it goes wrong, `npx rigorrun feedback export` produces a bundle that contains
128
- no credentials, no tool arguments and no results — only what is needed to work
129
- out where it broke.
130
-
131
- ## Documentation
122
+ ```bash
123
+ rigorrun doctor # what this machine has, and what it does not
124
+ ```
132
125
 
133
- <https://github.com/Konuktor/rigorrun/tree/master/docs>
126
+ ---
134
127
 
135
- MIT licensed.
128
+ MIT · Built by Erbol Tahirov · [Report a problem](https://github.com/Konuktor/rigorrun/issues)