agentskills-tools 0.4.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- agentskills_tools-0.4.0/PKG-INFO +394 -0
- agentskills_tools-0.4.0/README.md +367 -0
- agentskills_tools-0.4.0/agentskills_tools/__init__.py +20 -0
- agentskills_tools-0.4.0/agentskills_tools/__main__.py +10 -0
- agentskills_tools-0.4.0/agentskills_tools/cli.py +444 -0
- agentskills_tools-0.4.0/agentskills_tools/cost.py +391 -0
- agentskills_tools-0.4.0/agentskills_tools/discovery.py +93 -0
- agentskills_tools-0.4.0/agentskills_tools/evals.py +467 -0
- agentskills_tools-0.4.0/agentskills_tools/evalspec.py +376 -0
- agentskills_tools-0.4.0/agentskills_tools/findings.py +63 -0
- agentskills_tools-0.4.0/agentskills_tools/inspection.py +82 -0
- agentskills_tools-0.4.0/agentskills_tools/lint.py +141 -0
- agentskills_tools-0.4.0/agentskills_tools/py.typed +0 -0
- agentskills_tools-0.4.0/agentskills_tools/render.py +153 -0
- agentskills_tools-0.4.0/agentskills_tools/scaffold.py +155 -0
- agentskills_tools-0.4.0/agentskills_tools/serve.py +51 -0
- agentskills_tools-0.4.0/agentskills_tools/validate.py +153 -0
- agentskills_tools-0.4.0/pyproject.toml +37 -0
|
@@ -0,0 +1,394 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: agentskills-tools
|
|
3
|
+
Version: 0.4.0
|
|
4
|
+
Summary: Command line tools for authoring and validating Agent Skills (https://agentskills.io)
|
|
5
|
+
License: MIT
|
|
6
|
+
Author: Pratik Panda
|
|
7
|
+
Requires-Python: >=3.12,<4.0
|
|
8
|
+
Classifier: Development Status :: 3 - Alpha
|
|
9
|
+
Classifier: Environment :: Console
|
|
10
|
+
Classifier: Intended Audience :: Developers
|
|
11
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
12
|
+
Classifier: Programming Language :: Python :: 3
|
|
13
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
16
|
+
Classifier: Topic :: Software Development :: Libraries
|
|
17
|
+
Provides-Extra: serve
|
|
18
|
+
Requires-Dist: agentskills-adapters (>=0.3.0,<1.0)
|
|
19
|
+
Requires-Dist: agentskills-core (>=0.4.0,<1.0)
|
|
20
|
+
Requires-Dist: agentskills-fs (>=0.3.0,<1.0)
|
|
21
|
+
Requires-Dist: agentskills-mcp-server (>=0.3.0,<1.0) ; extra == "serve"
|
|
22
|
+
Requires-Dist: pyyaml (>=6.0,<7.0)
|
|
23
|
+
Project-URL: Homepage, https://agentskills.io
|
|
24
|
+
Project-URL: Repository, https://github.com/pratikxpanda/agentskills-sdk
|
|
25
|
+
Description-Content-Type: text/markdown
|
|
26
|
+
|
|
27
|
+
# agentskills-tools
|
|
28
|
+
|
|
29
|
+
Command line tools for authoring and validating [Agent Skills](https://agentskills.io).
|
|
30
|
+
|
|
31
|
+
Part of the [Agent Skills SDK](https://github.com/pratikxpanda/agentskills-sdk).
|
|
32
|
+
|
|
33
|
+
## Install
|
|
34
|
+
|
|
35
|
+
```bash
|
|
36
|
+
pip install agentskills-tools
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
The `serve` command needs the MCP server, which is an optional extra so that
|
|
40
|
+
validating skills in CI does not pull in `mcp` and `pydantic`:
|
|
41
|
+
|
|
42
|
+
```bash
|
|
43
|
+
pip install "agentskills-tools[serve]"
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
## Commands
|
|
47
|
+
|
|
48
|
+
Every command takes either one skill folder or a folder of skill folders — the
|
|
49
|
+
one containing `SKILL.md`, or the one containing directories that do.
|
|
50
|
+
|
|
51
|
+
| Command | What it does |
|
|
52
|
+
| --- | --- |
|
|
53
|
+
| `agentskills init <name>` | Scaffold a skill that already validates. |
|
|
54
|
+
| `agentskills validate <path>` | Check skills against the specification. Exits `1` on any error. |
|
|
55
|
+
| `agentskills lint <path>` | Report what is legal but still costly. |
|
|
56
|
+
| `agentskills inspect <path>` | Show what an agent would actually receive, and what it costs. |
|
|
57
|
+
| `agentskills eval <path>` | Measure what difference a skill makes. |
|
|
58
|
+
| `agentskills serve <path>` | Run an MCP server over a folder of skills. |
|
|
59
|
+
|
|
60
|
+
### `init`
|
|
61
|
+
|
|
62
|
+
```bash
|
|
63
|
+
agentskills init incident-response --path ./skills
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
Creates `skills/incident-response/` with a valid `SKILL.md` and empty
|
|
67
|
+
`references/`, `scripts/`, and `assets/` directories. The template is validated
|
|
68
|
+
before anything is written, so an unusable name is refused rather than
|
|
69
|
+
scaffolded.
|
|
70
|
+
|
|
71
|
+
### `validate`
|
|
72
|
+
|
|
73
|
+
```bash
|
|
74
|
+
agentskills validate ./skills
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
```text
|
|
78
|
+
skills/incident-response
|
|
79
|
+
ok
|
|
80
|
+
skills/broken-skill
|
|
81
|
+
error frontmatter-invalid-yaml (line 3): frontmatter is not valid YAML: mapping values are not allowed here
|
|
82
|
+
|
|
83
|
+
2 skills checked, 1 error, 0 warnings
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
Frontmatter is parsed by the CLI before the skill reaches the SDK's validator.
|
|
87
|
+
The SDK's parser is deliberately forgiving — malformed YAML yields an empty
|
|
88
|
+
mapping — which downstream reads as "no name, no description" and tells you
|
|
89
|
+
nothing about the colon you missed.
|
|
90
|
+
|
|
91
|
+
### `lint`
|
|
92
|
+
|
|
93
|
+
```bash
|
|
94
|
+
agentskills lint ./skills --strict --max-body-tokens 4000
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
| Code | Warning |
|
|
98
|
+
| --- | --- |
|
|
99
|
+
| `missing-version` | No `version`, so consumers cannot pin the skill or detect drift. |
|
|
100
|
+
| `description-too-long-for-catalog` | Catalog entries sit in context every turn. |
|
|
101
|
+
| `body-over-token-budget` | Body is large enough that detail belongs in `references/`. |
|
|
102
|
+
| `unreferenced-resource` | A file the body never mentions is a file no agent will load. |
|
|
103
|
+
|
|
104
|
+
Warnings do not fail the command unless `--strict` is passed.
|
|
105
|
+
|
|
106
|
+
### `inspect`
|
|
107
|
+
|
|
108
|
+
```bash
|
|
109
|
+
agentskills inspect ./skills/incident-response
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
Prints the metadata, the resource list, the catalog entry the agent sees on
|
|
113
|
+
every turn, and the body it loads on demand — each with an estimated token
|
|
114
|
+
cost, so you can see the price before shipping.
|
|
115
|
+
|
|
116
|
+
#### Token cost
|
|
117
|
+
|
|
118
|
+
```bash
|
|
119
|
+
agentskills inspect ./skills --cost
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
```text
|
|
123
|
+
skills/incident-response (incident-response)
|
|
124
|
+
counted with tiktoken/cl100k_base
|
|
125
|
+
catalog entry 66 every turn
|
|
126
|
+
body 439 on load
|
|
127
|
+
Incident Response 14
|
|
128
|
+
When to Declare an Incident 50
|
|
129
|
+
Roles 70
|
|
130
|
+
General Triage Steps 121
|
|
131
|
+
references/escalation-policy.md 448 on demand
|
|
132
|
+
assets/escalation-flowchart.mermaid 235 on demand
|
|
133
|
+
per turn 66, per load 505, all resources 2,197
|
|
134
|
+
|
|
135
|
+
1 skill, 66 tokens charged every turn
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
The right-hand column is the point. A catalog entry is injected on **every
|
|
139
|
+
turn** whether or not the skill is ever used; a body is charged **once per
|
|
140
|
+
load**; a reference is charged **only if the agent goes and reads it**. Authors
|
|
141
|
+
reliably get this backwards, trimming a body while ignoring a description that
|
|
142
|
+
costs a hundred tokens a turn forever.
|
|
143
|
+
|
|
144
|
+
Sections do not nest — a heading owns its own text up to the next heading of
|
|
145
|
+
any level — so the parts sum to the body exactly. Depth shows in the indent
|
|
146
|
+
instead. A `#` inside a fenced code block is a shell comment, not a heading.
|
|
147
|
+
|
|
148
|
+
A resource that is not UTF-8 text reports its size in bytes and no token count,
|
|
149
|
+
because an image has a size but not a token cost.
|
|
150
|
+
|
|
151
|
+
| Flag | Effect |
|
|
152
|
+
| --- | --- |
|
|
153
|
+
| `--budget N` | Exit `1` when catalog entry plus body exceeds `N` tokens. |
|
|
154
|
+
| `--turn-budget N` | Exit `1` when the catalog entry alone exceeds `N` tokens. |
|
|
155
|
+
| `--tokenizer` | `auto` (default), `tiktoken`, or `heuristic`. |
|
|
156
|
+
|
|
157
|
+
Two budgets rather than one, for the same reason: a single threshold is
|
|
158
|
+
dominated by the body, so the per-turn cost stays invisible to exactly the gate
|
|
159
|
+
meant to catch it.
|
|
160
|
+
|
|
161
|
+
Counting is exact when [`tiktoken`](https://pypi.org/project/tiktoken/) is
|
|
162
|
+
installed and a four-characters-per-token estimate otherwise. It is not a
|
|
163
|
+
dependency here: it ships a compiled wheel and fetches its vocabulary over the
|
|
164
|
+
network on first use, which is a poor trade for a tool whose main job is
|
|
165
|
+
reading YAML in CI. Install it yourself if you want exact numbers.
|
|
166
|
+
|
|
167
|
+
Whichever counter ran is named in every report, and `--tokenizer tiktoken`
|
|
168
|
+
refuses to fall back — a budget gate that quietly changes its arithmetic
|
|
169
|
+
depending on what happens to be installed is worse than no gate. Pin it in CI
|
|
170
|
+
and leave `auto` for the terminal.
|
|
171
|
+
|
|
172
|
+
`lint --max-body-tokens` keeps the estimate regardless, so its verdict never
|
|
173
|
+
depends on the machine it ran on.
|
|
174
|
+
|
|
175
|
+
### `eval`
|
|
176
|
+
|
|
177
|
+
A skill is a prompt, and nobody measures whether a given prompt makes an agent
|
|
178
|
+
better. Authors ship on intuition, reviewers approve on prose quality, and
|
|
179
|
+
editing a body can degrade task success with no signal anywhere.
|
|
180
|
+
|
|
181
|
+
Write cases beside the skill, in `evals/` inside the skill folder:
|
|
182
|
+
|
|
183
|
+
```yaml
|
|
184
|
+
# skills/incident-response/evals/triage.yaml
|
|
185
|
+
skill: incident-response # optional; checked against the folder
|
|
186
|
+
judge_model: gpt-4o # required if any case uses `judge`
|
|
187
|
+
cases:
|
|
188
|
+
- name: declares-and-triages
|
|
189
|
+
prompt: Checkout is returning 500s for a third of users.
|
|
190
|
+
repeat: 3 # models are not deterministic
|
|
191
|
+
threshold: 0.67 # fraction of repeats that must pass
|
|
192
|
+
expect:
|
|
193
|
+
- contains: "Incident Commander"
|
|
194
|
+
- not_contains: "I don't have access"
|
|
195
|
+
- regex: "(?i)severity"
|
|
196
|
+
- judge: "Tells the responder to assess severity before attempting a fix"
|
|
197
|
+
```
|
|
198
|
+
|
|
199
|
+
`repeat` defaults to `1` and `threshold` to `1.0`. Every expectation must hold
|
|
200
|
+
for a repeat to pass.
|
|
201
|
+
|
|
202
|
+
Eval files are checked by `agentskills validate`, with no model and no API key,
|
|
203
|
+
so a broken case fails in CI beside the skill rather than the first time
|
|
204
|
+
somebody pays to run it.
|
|
205
|
+
|
|
206
|
+
```bash
|
|
207
|
+
agentskills eval ./skills --model mypkg.evals:openai_client
|
|
208
|
+
```
|
|
209
|
+
|
|
210
|
+
```text
|
|
211
|
+
incident-response (triage.yaml)
|
|
212
|
+
pass declares-and-triages: with 100%, without 33%, delta +67%
|
|
213
|
+
FAIL postmortem-window: with 0%, without 0%, delta +0%
|
|
214
|
+
unmet contains: 48 hours
|
|
215
|
+
suite delta +33% on gpt-4o
|
|
216
|
+
|
|
217
|
+
2 cases run, 1 failed, mean delta +33%
|
|
218
|
+
```
|
|
219
|
+
|
|
220
|
+
Every case runs twice: once with the skill's body in the system prompt, once
|
|
221
|
+
without. Absolute pass rates mostly measure the underlying model, so the number
|
|
222
|
+
that means anything is the difference. A skill whose cases pass equally well
|
|
223
|
+
without it is not earning its tokens.
|
|
224
|
+
|
|
225
|
+
#### Bringing your own model
|
|
226
|
+
|
|
227
|
+
`--model` takes `module:factory` — a dotted path to a zero-argument callable
|
|
228
|
+
returning a client. Nothing in this project depends on a provider SDK, and a
|
|
229
|
+
ten-line adapter is a smaller ask than an opinion about which vendor you should
|
|
230
|
+
install:
|
|
231
|
+
|
|
232
|
+
```python
|
|
233
|
+
# mypkg/evals.py
|
|
234
|
+
from openai import AsyncOpenAI
|
|
235
|
+
from agentskills_tools.evals import ModelResponse
|
|
236
|
+
|
|
237
|
+
|
|
238
|
+
class OpenAIModel:
|
|
239
|
+
model_id = "gpt-4o"
|
|
240
|
+
|
|
241
|
+
def __init__(self) -> None:
|
|
242
|
+
self._client = AsyncOpenAI()
|
|
243
|
+
|
|
244
|
+
async def complete(self, *, system: str, prompt: str) -> ModelResponse:
|
|
245
|
+
reply = await self._client.chat.completions.create(
|
|
246
|
+
model=self.model_id,
|
|
247
|
+
temperature=0,
|
|
248
|
+
messages=[
|
|
249
|
+
{"role": "system", "content": system},
|
|
250
|
+
{"role": "user", "content": prompt},
|
|
251
|
+
],
|
|
252
|
+
)
|
|
253
|
+
return ModelResponse(reply.choices[0].message.content or "")
|
|
254
|
+
|
|
255
|
+
|
|
256
|
+
openai_client = OpenAIModel
|
|
257
|
+
```
|
|
258
|
+
|
|
259
|
+
`model_id` is part of every report and of the cache key, because a pass rate
|
|
260
|
+
without the model that produced it is not a measurement. Set temperature to
|
|
261
|
+
zero if your provider allows it; this side has no opinion it could enforce.
|
|
262
|
+
|
|
263
|
+
`--judge` names a second client for `judge` expectations and defaults to the
|
|
264
|
+
model under test — the cheapest judge and the least independent one. When
|
|
265
|
+
`repeat` is above `1`, the report flags cases whose repeats disagreed, because
|
|
266
|
+
a case that passes three times in five has measured sampling noise rather than
|
|
267
|
+
a skill.
|
|
268
|
+
|
|
269
|
+
#### Cost
|
|
270
|
+
|
|
271
|
+
These calls hit real APIs and cost real money. `eval` is never part of
|
|
272
|
+
`pytest`: it runs only when you invoke it, with credentials you supply.
|
|
273
|
+
Completions are cached under `.agentskills/eval-cache` by model, system prompt,
|
|
274
|
+
user prompt, and repeat index — so editing a skill re-buys its runs, while
|
|
275
|
+
tightening an expectation re-grades the answers already bought. `--no-cache`
|
|
276
|
+
turns that off; `--cache-dir` moves it.
|
|
277
|
+
|
|
278
|
+
### `serve`
|
|
279
|
+
|
|
280
|
+
```bash
|
|
281
|
+
agentskills serve ./skills --transport stdio
|
|
282
|
+
```
|
|
283
|
+
|
|
284
|
+
Runs the MCP server over a folder of skills without hand-writing a config
|
|
285
|
+
file. For anything beyond a single filesystem root — HTTP providers,
|
|
286
|
+
per-skill options, environment placeholders — use
|
|
287
|
+
[agentskills-mcp-server](https://github.com/pratikxpanda/agentskills-sdk/tree/main/packages/integrations/agentskills-mcp-server)
|
|
288
|
+
with a `server.json`.
|
|
289
|
+
|
|
290
|
+
## Exit codes
|
|
291
|
+
|
|
292
|
+
| Code | Meaning |
|
|
293
|
+
| --- | --- |
|
|
294
|
+
| `0` | Ran, found nothing wrong. |
|
|
295
|
+
| `1` | Ran, found errors — or warnings under `--strict`, or a cost over budget. |
|
|
296
|
+
| `2` | Could not run: bad path, missing extra, unwritable directory. |
|
|
297
|
+
|
|
298
|
+
The distinction matters in CI: `1` means a skill is broken, `2` means the
|
|
299
|
+
invocation is.
|
|
300
|
+
|
|
301
|
+
## JSON output
|
|
302
|
+
|
|
303
|
+
`validate`, `lint`, `inspect`, and `eval` accept `--format json`. The schema is
|
|
304
|
+
a published contract; `schemaVersion` is bumped only for a breaking change, and
|
|
305
|
+
new fields are added rather than existing ones repurposed.
|
|
306
|
+
|
|
307
|
+
```json
|
|
308
|
+
{
|
|
309
|
+
"schemaVersion": 1,
|
|
310
|
+
"command": "validate",
|
|
311
|
+
"ok": false,
|
|
312
|
+
"summary": { "skills": 2, "errors": 1, "warnings": 0 },
|
|
313
|
+
"skills": [
|
|
314
|
+
{
|
|
315
|
+
"id": "broken-skill",
|
|
316
|
+
"path": "skills/broken-skill",
|
|
317
|
+
"ok": false,
|
|
318
|
+
"findings": [
|
|
319
|
+
{
|
|
320
|
+
"severity": "error",
|
|
321
|
+
"code": "frontmatter-invalid-yaml",
|
|
322
|
+
"message": "frontmatter is not valid YAML: mapping values are not allowed here",
|
|
323
|
+
"line": 3,
|
|
324
|
+
"file": "skills/broken-skill/SKILL.md"
|
|
325
|
+
}
|
|
326
|
+
]
|
|
327
|
+
}
|
|
328
|
+
]
|
|
329
|
+
}
|
|
330
|
+
```
|
|
331
|
+
|
|
332
|
+
`ok` mirrors the exit code, so a consumer never has to re-derive the
|
|
333
|
+
strictness rules. `line` is `null` unless the problem can be attributed to one
|
|
334
|
+
line. `file` is the skill's `SKILL.md` unless the finding is about another file
|
|
335
|
+
in the folder, such as an eval case file.
|
|
336
|
+
|
|
337
|
+
`inspect --cost --format json` reports each skill's `perTurn`, `perLoad` and
|
|
338
|
+
`onDemand` totals, the `sections` and `resources` they were summed from, the
|
|
339
|
+
`overBudget` messages, and the `counter` that produced the numbers — including
|
|
340
|
+
whether it was `exact`. A consumer that charts these over time needs to know
|
|
341
|
+
when the unit changed underneath it.
|
|
342
|
+
|
|
343
|
+
## Continuous integration
|
|
344
|
+
|
|
345
|
+
The published action wraps `validate` and `lint` and annotates every finding
|
|
346
|
+
on the pull request diff:
|
|
347
|
+
|
|
348
|
+
```yaml
|
|
349
|
+
- uses: pratikxpanda/agentskills-sdk/actions/validate@v1
|
|
350
|
+
with:
|
|
351
|
+
path: ./skills
|
|
352
|
+
fail-on-lint: false
|
|
353
|
+
```
|
|
354
|
+
|
|
355
|
+
To run it yourself, `validate` and `lint` also accept `--format github`, which
|
|
356
|
+
emits [workflow commands](https://docs.github.com/actions/reference/workflow-commands-for-github-actions)
|
|
357
|
+
instead of a report:
|
|
358
|
+
|
|
359
|
+
```bash
|
|
360
|
+
agentskills validate ./skills --format github
|
|
361
|
+
```
|
|
362
|
+
|
|
363
|
+
```text
|
|
364
|
+
::error file=skills/deploy/SKILL.md,line=3,title=frontmatter-invalid-yaml::frontmatter is not valid YAML
|
|
365
|
+
```
|
|
366
|
+
|
|
367
|
+
Anywhere else, the exit code is enough:
|
|
368
|
+
|
|
369
|
+
```yaml
|
|
370
|
+
- run: pip install agentskills-tools
|
|
371
|
+
- run: agentskills validate ./skills
|
|
372
|
+
```
|
|
373
|
+
|
|
374
|
+
## Logging
|
|
375
|
+
|
|
376
|
+
Pass `-v` to send the SDK's debug logs to stderr, leaving stdout parseable:
|
|
377
|
+
|
|
378
|
+
```bash
|
|
379
|
+
agentskills validate ./skills --format json -v > report.json
|
|
380
|
+
```
|
|
381
|
+
|
|
382
|
+
## Security
|
|
383
|
+
|
|
384
|
+
Agent Skills are **equivalent to executable code** — skill content is injected
|
|
385
|
+
into an LLM agent's context verbatim. Validating a skill does not make it safe
|
|
386
|
+
to run. **Only load skills from sources you trust.**
|
|
387
|
+
|
|
388
|
+
See
|
|
389
|
+
[SECURITY.md](https://github.com/pratikxpanda/agentskills-sdk/blob/main/SECURITY.md).
|
|
390
|
+
|
|
391
|
+
## License
|
|
392
|
+
|
|
393
|
+
MIT
|
|
394
|
+
|