rigorrun 0.1.1 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE CHANGED
@@ -1,6 +1,6 @@
1
1
  MIT License
2
2
 
3
- Copyright (c) 2026 RigorRun
3
+ Copyright (c) 2026 Erbol Tahirov
4
4
 
5
5
  Permission is hereby granted, free of charge, to any person obtaining a copy
6
6
  of this software and associated documentation files (the "Software"), to deal
package/README.md CHANGED
@@ -1,113 +1,122 @@
1
+ <div align="center">
2
+
1
3
  # RigorRun
2
4
 
3
- **Acceptance testing for tool-using AI agents.**
5
+ **Your agent said it worked. RigorRun checks what it actually did.**
6
+
7
+ Show RigorRun a job once. It turns that into a repeatable acceptance suite and decides whether your
8
+ agent is safe to ship by reading the system it changed — never by trusting what it says about itself.
9
+
10
+ **Early Access · v0.2** — parts of it are honestly unfinished, and they are listed rather than hidden.
4
11
 
5
- Connect your system. Show RigorRun how one job is done. Connect your agent.
6
- RigorRun proves whether the agent can do that job safely — by reading the
7
- system it changed, never by trusting what it says about itself.
12
+ [rigorrun.xyz](https://rigorrun.xyz) · [Documentation](https://docs.rigorrun.xyz) · [Evidence](https://rigorrun.xyz/evidence) · [What is and is not built](https://rigorrun.xyz/what-is-built)
13
+
14
+ </div>
8
15
 
9
16
  ```bash
10
17
  npx rigorrun
11
18
  ```
12
19
 
13
- Open the URL it prints. That is the whole install.
20
+ Open the URL it prints. Everything runs on your machine: there is no account, and no hosted
21
+ component to send your systems to.
14
22
 
15
- **Early Access · v0.1.** It does what this page says and it is young. The
16
- limits are written down and marked one by one, rather than left for you to
17
- find: [the v1 gap audit](https://github.com/Konuktor/rigorrun/blob/master/docs/V1_GAP_AUDIT.md).
23
+ ---
18
24
 
19
- ## What it does
25
+ ## Why
20
26
 
21
- You have an agent that calls tools. You need to know whether it can do a real
22
- job in your real system without doing something unsafe — and you need to know
23
- again next week, after somebody changes a prompt.
27
+ An agent that reports success and an agent that achieved it are indistinguishable from the
28
+ transcript. They are trivially distinguishable from the database.
24
29
 
25
- RigorRun watches a person do that job once, reads the system before and after,
26
- works out what the rules must be, asks about what it can only guess, and turns
27
- the answers into an executable acceptance suite. Then it runs your agent against
28
- it and reads your system to find out what actually happened.
30
+ Evaluation harnesses score what the model wrote. RigorRun compares system state before and after,
31
+ through read operations you nominate, and asks the questions that have a fact behind them: does the
32
+ refund exist, is the amount right, is the ticket attached, was approval required, and did anything
33
+ change that should not have.
29
34
 
30
35
  ```
31
- most tools: you write the tests → the tool runs them
32
- RigorRun: you do the job once → RigorRun writes the tests
36
+ most tools: you write the tests → the tool runs them
37
+ RigorRun: you do the job once → RigorRun writes the tests
33
38
  ```
34
39
 
35
- ## How strongly was it verified?
40
+ ## How
36
41
 
37
- RigorRun tells you how strongly each result was verified, on every result:
42
+ 1. **Connect your system** — an MCP server you already run, an OpenAPI document, or a web
43
+ application through a browser.
44
+ 2. **Do the job once.** RigorRun reads your system before and after and derives what the rules must
45
+ be.
46
+ 3. **Rule on what it worked out.** It shows the evidence behind each proposed rule. A rule you
47
+ reject cannot fail your agent.
48
+ 4. **Connect your agent** — an HTTP endpoint, a local command, or your own loop pulling work.
49
+ 5. **Run it.** A verdict, with how strongly each answer could be verified.
50
+ 6. **Gate the next change** in CI.
38
51
 
39
- - **AUTHORITATIVE** — checked against direct, trusted state.
40
- - **PARTIAL** — verified through the reads your system exposes. A normal
41
- connected MCP server, whose state RigorRun reads back through the tools you
42
- nominated, is **PARTIAL** — the common, honest case, not a defect.
43
- - **OBSERVATIONAL** — actions were observed but the final state could not be
44
- independently proven (e.g. a browser with nothing readable attached).
52
+ Setting a project up happens in the local interface. Running, gating and comparing are also
53
+ available from the command line, which is the half CI needs.
45
54
 
46
- `AUTHORITATIVE` is claimed only where RigorRun genuinely has authoritative state
47
- access. Against your own system the honest label is usually `PARTIAL`, and it is
48
- shown on the same line as the verdict.
55
+ ```bash
56
+ rigorrun gate --project <id> # exit 1 stops the build
57
+ ```
49
58
 
50
- ## What it can build depends on your system
59
+ ## Verify a published MCP server
51
60
 
52
- RigorRun generates every case it can safely and reproducibly verify, and tells
53
- you what it could not test. A system it can **seed and reset** yields the
54
- richest suite (boundary and adversarial cases, repeated destructive checks). A
55
- system without a reset still works, but produces fewer cases, disables repeated
56
- mutating cases, reports isolation `NONE`, and verifies `PARTIAL`. Best results:
57
- a staging or scratch environment with read-back and a reset.
61
+ An MCP server can annotate a tool `readOnlyHint: true`. Nothing checks that, and agents use it to
62
+ decide whether a tool may be called without asking.
58
63
 
59
- ## Why it runs locally
64
+ ```bash
65
+ rigorrun verify npm:@modelcontextprotocol/server-memory@2026.8.31
66
+ ```
60
67
 
61
- Your MCP server, your internal API and your staging box are usually not
62
- reachable from the public internet, and a page served over `https` cannot fetch
63
- `http://127.0.0.1`. So RigorRun's interface is served by this process, on your
64
- machine. Your credentials, recordings and systems never touch anybody's
65
- infrastructure, because there is no path by which they could.
68
+ Pins the server to the exact bytes the registry published, runs it in a container with no network
69
+ and no access to your machine, calls each tool with arguments derived from its own schema, and reads
70
+ the filesystem before and after. Needs a container runtime; nothing else does.
66
71
 
67
- ## What you need
72
+ Run against four published servers, it exercised **19 of 37** tools. The other 18 are named, each
73
+ with the reason it was not reached — [see the records](https://rigorrun.xyz/evidence).
68
74
 
69
- - **Node 20.11 or newer.**
70
- - **A way in to the system you want to test**: an MCP server, or an OpenAPI
71
- document and the address it is served from. Ideally staging or a scratch
72
- instance, with a way to reset it.
73
- - **An agent.** If it speaks MCP it works unchanged; RigorRun hands it a URL —
74
- whether your agent listens on an address or is a command RigorRun runs. If it
75
- does not speak MCP, about ten lines of plain HTTP — there is no package to
76
- install; the protocol is documented at
77
- [docs/HTTP_AGENT.md](https://github.com/Konuktor/rigorrun/blob/master/docs/HTTP_AGENT.md).
75
+ ## What you should know before relying on it
78
76
 
79
- ## Commands
77
+ - **A verdict against a real system is `PARTIAL`, by design.** RigorRun reads back what your
78
+ nominated reads return and no more, and says so on every result.
79
+ - **A browser cannot verify itself.** A page saying "done" is a claim by the system that would have
80
+ to be wrong for it not to be done, so a browser connection is `OBSERVATIONAL`. Attach a readable
81
+ API or MCP connection for the same system to strengthen it.
82
+ - **Isolation is `DECLARED`, not `RESET`, when you nominate a reset tool** — RigorRun has not run it
83
+ twice and compared.
84
+ - **Three of the six verification sources are not emitted yet.** Model-judged and human-review
85
+ evaluators exist in the schema and are unreachable in practice.
86
+ - **No external team has used this yet.** There are no customers.
80
87
 
81
- ```bash
82
- npx rigorrun # start the runner and open the interface
83
- npx rigorrun doctor # check this machine and every project
84
- npx rigorrun projects # what is on this machine
85
- npx rigorrun run --project <id> # run the suite
86
- npx rigorrun gate --project <id> # run it, and exit non-zero if it misses the bar
87
- npx rigorrun compare-runs --project <id> <runId>
88
- npx rigorrun feedback export # a sanitised bundle for a bug report
89
- ```
88
+ The full list is at [rigorrun.xyz/what-is-built](https://rigorrun.xyz/what-is-built), and the
89
+ version carrying how each line was checked is
90
+ [docs/V1_GAP_AUDIT.md](https://github.com/Konuktor/rigorrun/blob/master/docs/V1_GAP_AUDIT.md).
91
+
92
+ ## Privacy
93
+
94
+ Your systems, your credentials and your recordings stay on this machine. There is no account, and
95
+ nothing is uploaded unless you ask for it.
90
96
 
91
- ## This is early access
97
+ Three things can make a network request, all because you asked: `rigorrun verify` downloads the named
98
+ package from the npm registry; connecting your own system sends requests to your own system; and an
99
+ LLM-backed agent, which exists only if you set a provider key, talks to that provider. None of them
100
+ reach RigorRun. [How it is built](https://rigorrun.xyz/security).
92
101
 
93
- It connects to MCP servers, to HTTP APIs with an OpenAPI document, and to web
94
- applications through a browser — though a browser cannot verify itself, so a
95
- verdict from one is OBSERVATIONAL unless something readable is attached.
96
- Setting a project up needs the interface; running and gating it does not. It has
97
- been used successfully by the people who wrote it and is now looking for people
98
- who did not.
102
+ ## Verifying what you installed
99
103
 
100
- Every capability is marked WORKING, PARTIAL, DEMO-ONLY, BROKEN or MISSING in
101
- [the v1 gap audit](https://github.com/Konuktor/rigorrun/blob/master/docs/V1_GAP_AUDIT.md),
102
- with how each one was checked. Read it before you rely on this for anything
103
- that matters.
104
+ Published by a GitHub Actions workflow that holds no npm token, using npm trusted publishing, with a
105
+ SLSA build provenance attestation.
104
106
 
105
- If it goes wrong, `npx rigorrun feedback export` produces a bundle that contains
106
- no credentials, no tool arguments and no results — only what is needed to work
107
- out where it broke.
107
+ ```bash
108
+ npm audit signatures
109
+ npm view rigorrun dist.integrity
110
+ ```
108
111
 
109
- ## Documentation
112
+ ## Requirements
113
+
114
+ Node 20.11 or newer. Docker only for `rigorrun verify`.
115
+
116
+ ```bash
117
+ rigorrun doctor # what this machine has, and what it does not
118
+ ```
110
119
 
111
- <https://github.com/Konuktor/rigorrun/tree/master/docs>
120
+ ---
112
121
 
113
- MIT licensed.
122
+ MIT · Built by Erbol Tahirov · [Report a problem](https://github.com/Konuktor/rigorrun/issues)