@icntswm/skillcheck 0.1.0 → 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +9 -18
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -12,22 +12,12 @@ some of your requests load the wrong skill. Nothing warns you: the agent still
|
|
|
12
12
|
answers, just with the wrong instructions. skillcheck catches this before your
|
|
13
13
|
users do.
|
|
14
14
|
|
|
15
|
-
|
|
16
|
-
$ skillcheck run examples/demo/skillcheck.yaml --skill test-guard,find-bug
|
|
17
|
-
5 cases × 1 repeat × 1 agent = 5 runs
|
|
18
|
-
ok #bug-500 POST /orders returns 500 with 'nil pointer deref… → find-bug
|
|
19
|
-
FAIL #flaky-ci-only this test is red in CI roughly once a day but I … → find-bug · not loaded test-guard
|
|
20
|
-
FAIL #flaky-retry TestCartMerge fails sometimes and goes green w… → find-bug · not loaded test-guard; forbidden find-bug
|
|
21
|
-
ok #bug-consistent TestCheckoutTotal fails on every run since this … → find-bug
|
|
22
|
-
ok #perf-latency the /search endpoint went from 80ms to 900ms aft… → perf-profile
|
|
23
|
-
|
|
24
|
-
confusion:
|
|
25
|
-
expected test-guard → got find-bug (2)
|
|
26
|
-
2 failed of 5 · runs 5 · cost $0.25
|
|
27
|
-
```
|
|
15
|
+

|
|
28
16
|
|
|
29
|
-
|
|
30
|
-
|
|
17
|
+
A real run on the [demo](examples/demo): one skill's description got wider,
|
|
18
|
+
and it started taking requests that belong to its neighbour. The free `lint`
|
|
19
|
+
flags a suspicious description, and one batch call finds the two misrouted
|
|
20
|
+
requests.
|
|
31
21
|
|
|
32
22
|
## What you get
|
|
33
23
|
|
|
@@ -46,8 +36,9 @@ wider, and it started taking requests that belong to its neighbour.
|
|
|
46
36
|
from a broken one instead of letting you guess.
|
|
47
37
|
- 🏷️ **Catches renames and typos.** Case names are checked against the skills
|
|
48
38
|
the agent really has, so a renamed skill can't pass silently.
|
|
49
|
-
- ⚙️ **CI-ready.**
|
|
50
|
-
|
|
39
|
+
- ⚙️ **CI-ready.** A GitHub Action (`uses: icntswm/skillcheck@v1`), JUnit
|
|
40
|
+
and JSON reports, clear exit codes, a spending cap (`--budget`), and
|
|
41
|
+
`--config-dir` to test only the skills in your repository.
|
|
51
42
|
- 📄 **Plain YAML, one dependency.** Cases are readable by anyone on the team
|
|
52
43
|
and live next to the skills they test.
|
|
53
44
|
|
|
@@ -86,7 +77,7 @@ Start at the cheapest level and go up only when it finds nothing.
|
|
|
86
77
|
| `skillcheck run --batch` | one per 25 cases | which skill the model *says* it would load |
|
|
87
78
|
| `skillcheck run` | one per case | which skill the model *actually* loads |
|
|
88
79
|
|
|
89
|
-
On the demo suite one batch call covered all 12 cases for $0.05 and agreed
|
|
80
|
+
On the demo suite one batch call covered all 12 cases for $0.05–0.10 and agreed
|
|
90
81
|
with the normal run in 34 checks out of 34. More in [docs/cost.md](docs/cost.md).
|
|
91
82
|
|
|
92
83
|
## Does it really catch regressions?
|