agentskills-tools 0.4.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,394 @@
1
+ Metadata-Version: 2.4
2
+ Name: agentskills-tools
3
+ Version: 0.4.0
4
+ Summary: Command line tools for authoring and validating Agent Skills (https://agentskills.io)
5
+ License: MIT
6
+ Author: Pratik Panda
7
+ Requires-Python: >=3.12,<4.0
8
+ Classifier: Development Status :: 3 - Alpha
9
+ Classifier: Environment :: Console
10
+ Classifier: Intended Audience :: Developers
11
+ Classifier: License :: OSI Approved :: MIT License
12
+ Classifier: Programming Language :: Python :: 3
13
+ Classifier: Programming Language :: Python :: 3.12
14
+ Classifier: Programming Language :: Python :: 3.13
15
+ Classifier: Programming Language :: Python :: 3.14
16
+ Classifier: Topic :: Software Development :: Libraries
17
+ Provides-Extra: serve
18
+ Requires-Dist: agentskills-adapters (>=0.3.0,<1.0)
19
+ Requires-Dist: agentskills-core (>=0.4.0,<1.0)
20
+ Requires-Dist: agentskills-fs (>=0.3.0,<1.0)
21
+ Requires-Dist: agentskills-mcp-server (>=0.3.0,<1.0) ; extra == "serve"
22
+ Requires-Dist: pyyaml (>=6.0,<7.0)
23
+ Project-URL: Homepage, https://agentskills.io
24
+ Project-URL: Repository, https://github.com/pratikxpanda/agentskills-sdk
25
+ Description-Content-Type: text/markdown
26
+
27
+ # agentskills-tools
28
+
29
+ Command line tools for authoring and validating [Agent Skills](https://agentskills.io).
30
+
31
+ Part of the [Agent Skills SDK](https://github.com/pratikxpanda/agentskills-sdk).
32
+
33
+ ## Install
34
+
35
+ ```bash
36
+ pip install agentskills-tools
37
+ ```
38
+
39
+ The `serve` command needs the MCP server, which is an optional extra so that
40
+ validating skills in CI does not pull in `mcp` and `pydantic`:
41
+
42
+ ```bash
43
+ pip install "agentskills-tools[serve]"
44
+ ```
45
+
46
+ ## Commands
47
+
48
+ Every command takes either one skill folder or a folder of skill folders — the
49
+ one containing `SKILL.md`, or the one containing directories that do.
50
+
51
+ | Command | What it does |
52
+ | --- | --- |
53
+ | `agentskills init <name>` | Scaffold a skill that already validates. |
54
+ | `agentskills validate <path>` | Check skills against the specification. Exits `1` on any error. |
55
+ | `agentskills lint <path>` | Report what is legal but still costly. |
56
+ | `agentskills inspect <path>` | Show what an agent would actually receive, and what it costs. |
57
+ | `agentskills eval <path>` | Measure what difference a skill makes. |
58
+ | `agentskills serve <path>` | Run an MCP server over a folder of skills. |
59
+
60
+ ### `init`
61
+
62
+ ```bash
63
+ agentskills init incident-response --path ./skills
64
+ ```
65
+
66
+ Creates `skills/incident-response/` with a valid `SKILL.md` and empty
67
+ `references/`, `scripts/`, and `assets/` directories. The template is validated
68
+ before anything is written, so an unusable name is refused rather than
69
+ scaffolded.
70
+
71
+ ### `validate`
72
+
73
+ ```bash
74
+ agentskills validate ./skills
75
+ ```
76
+
77
+ ```text
78
+ skills/incident-response
79
+ ok
80
+ skills/broken-skill
81
+ error frontmatter-invalid-yaml (line 3): frontmatter is not valid YAML: mapping values are not allowed here
82
+
83
+ 2 skills checked, 1 error, 0 warnings
84
+ ```
85
+
86
+ Frontmatter is parsed by the CLI before the skill reaches the SDK's validator.
87
+ The SDK's parser is deliberately forgiving — malformed YAML yields an empty
88
+ mapping — which downstream reads as "no name, no description" and tells you
89
+ nothing about the colon you missed.
90
+
91
+ ### `lint`
92
+
93
+ ```bash
94
+ agentskills lint ./skills --strict --max-body-tokens 4000
95
+ ```
96
+
97
+ | Code | Warning |
98
+ | --- | --- |
99
+ | `missing-version` | No `version`, so consumers cannot pin the skill or detect drift. |
100
+ | `description-too-long-for-catalog` | Catalog entries sit in context every turn. |
101
+ | `body-over-token-budget` | Body is large enough that detail belongs in `references/`. |
102
+ | `unreferenced-resource` | A file the body never mentions is a file no agent will load. |
103
+
104
+ Warnings do not fail the command unless `--strict` is passed.
105
+
106
+ ### `inspect`
107
+
108
+ ```bash
109
+ agentskills inspect ./skills/incident-response
110
+ ```
111
+
112
+ Prints the metadata, the resource list, the catalog entry the agent sees on
113
+ every turn, and the body it loads on demand — each with an estimated token
114
+ cost, so you can see the price before shipping.
115
+
116
+ #### Token cost
117
+
118
+ ```bash
119
+ agentskills inspect ./skills --cost
120
+ ```
121
+
122
+ ```text
123
+ skills/incident-response (incident-response)
124
+ counted with tiktoken/cl100k_base
125
+ catalog entry 66 every turn
126
+ body 439 on load
127
+ Incident Response 14
128
+ When to Declare an Incident 50
129
+ Roles 70
130
+ General Triage Steps 121
131
+ references/escalation-policy.md 448 on demand
132
+ assets/escalation-flowchart.mermaid 235 on demand
133
+ per turn 66, per load 505, all resources 2,197
134
+
135
+ 1 skill, 66 tokens charged every turn
136
+ ```
137
+
138
+ The right-hand column is the point. A catalog entry is injected on **every
139
+ turn** whether or not the skill is ever used; a body is charged **once per
140
+ load**; a reference is charged **only if the agent goes and reads it**. Authors
141
+ reliably get this backwards, trimming a body while ignoring a description that
142
+ costs a hundred tokens a turn forever.
143
+
144
+ Sections do not nest — a heading owns its own text up to the next heading of
145
+ any level — so the parts sum to the body exactly. Depth shows in the indent
146
+ instead. A `#` inside a fenced code block is a shell comment, not a heading.
147
+
148
+ A resource that is not UTF-8 text reports its size in bytes and no token count,
149
+ because an image has a size but not a token cost.
150
+
151
+ | Flag | Effect |
152
+ | --- | --- |
153
+ | `--budget N` | Exit `1` when catalog entry plus body exceeds `N` tokens. |
154
+ | `--turn-budget N` | Exit `1` when the catalog entry alone exceeds `N` tokens. |
155
+ | `--tokenizer` | `auto` (default), `tiktoken`, or `heuristic`. |
156
+
157
+ Two budgets rather than one, for the same reason: a single threshold is
158
+ dominated by the body, so the per-turn cost stays invisible to exactly the gate
159
+ meant to catch it.
160
+
161
+ Counting is exact when [`tiktoken`](https://pypi.org/project/tiktoken/) is
162
+ installed and a four-characters-per-token estimate otherwise. It is not a
163
+ dependency here: it ships a compiled wheel and fetches its vocabulary over the
164
+ network on first use, which is a poor trade for a tool whose main job is
165
+ reading YAML in CI. Install it yourself if you want exact numbers.
166
+
167
+ Whichever counter ran is named in every report, and `--tokenizer tiktoken`
168
+ refuses to fall back — a budget gate that quietly changes its arithmetic
169
+ depending on what happens to be installed is worse than no gate. Pin it in CI
170
+ and leave `auto` for the terminal.
171
+
172
+ `lint --max-body-tokens` keeps the estimate regardless, so its verdict never
173
+ depends on the machine it ran on.
174
+
175
+ ### `eval`
176
+
177
+ A skill is a prompt, and nobody measures whether a given prompt makes an agent
178
+ better. Authors ship on intuition, reviewers approve on prose quality, and
179
+ editing a body can degrade task success with no signal anywhere.
180
+
181
+ Write cases beside the skill, in `evals/` inside the skill folder:
182
+
183
+ ```yaml
184
+ # skills/incident-response/evals/triage.yaml
185
+ skill: incident-response # optional; checked against the folder
186
+ judge_model: gpt-4o # required if any case uses `judge`
187
+ cases:
188
+ - name: declares-and-triages
189
+ prompt: Checkout is returning 500s for a third of users.
190
+ repeat: 3 # models are not deterministic
191
+ threshold: 0.67 # fraction of repeats that must pass
192
+ expect:
193
+ - contains: "Incident Commander"
194
+ - not_contains: "I don't have access"
195
+ - regex: "(?i)severity"
196
+ - judge: "Tells the responder to assess severity before attempting a fix"
197
+ ```
198
+
199
+ `repeat` defaults to `1` and `threshold` to `1.0`. Every expectation must hold
200
+ for a repeat to pass.
201
+
202
+ Eval files are checked by `agentskills validate`, with no model and no API key,
203
+ so a broken case fails in CI beside the skill rather than the first time
204
+ somebody pays to run it.
205
+
206
+ ```bash
207
+ agentskills eval ./skills --model mypkg.evals:openai_client
208
+ ```
209
+
210
+ ```text
211
+ incident-response (triage.yaml)
212
+ pass declares-and-triages: with 100%, without 33%, delta +67%
213
+ FAIL postmortem-window: with 0%, without 0%, delta +0%
214
+ unmet contains: 48 hours
215
+ suite delta +33% on gpt-4o
216
+
217
+ 2 cases run, 1 failed, mean delta +33%
218
+ ```
219
+
220
+ Every case runs twice: once with the skill's body in the system prompt, once
221
+ without. Absolute pass rates mostly measure the underlying model, so the number
222
+ that means anything is the difference. A skill whose cases pass equally well
223
+ without it is not earning its tokens.
224
+
225
+ #### Bringing your own model
226
+
227
+ `--model` takes `module:factory` — a dotted path to a zero-argument callable
228
+ returning a client. Nothing in this project depends on a provider SDK, and a
229
+ ten-line adapter is a smaller ask than an opinion about which vendor you should
230
+ install:
231
+
232
+ ```python
233
+ # mypkg/evals.py
234
+ from openai import AsyncOpenAI
235
+ from agentskills_tools.evals import ModelResponse
236
+
237
+
238
+ class OpenAIModel:
239
+ model_id = "gpt-4o"
240
+
241
+ def __init__(self) -> None:
242
+ self._client = AsyncOpenAI()
243
+
244
+ async def complete(self, *, system: str, prompt: str) -> ModelResponse:
245
+ reply = await self._client.chat.completions.create(
246
+ model=self.model_id,
247
+ temperature=0,
248
+ messages=[
249
+ {"role": "system", "content": system},
250
+ {"role": "user", "content": prompt},
251
+ ],
252
+ )
253
+ return ModelResponse(reply.choices[0].message.content or "")
254
+
255
+
256
+ openai_client = OpenAIModel
257
+ ```
258
+
259
+ `model_id` is part of every report and of the cache key, because a pass rate
260
+ without the model that produced it is not a measurement. Set temperature to
261
+ zero if your provider allows it; this side has no opinion it could enforce.
262
+
263
+ `--judge` names a second client for `judge` expectations and defaults to the
264
+ model under test — the cheapest judge and the least independent one. When
265
+ `repeat` is above `1`, the report flags cases whose repeats disagreed, because
266
+ a case that passes three times in five has measured sampling noise rather than
267
+ a skill.
268
+
269
+ #### Cost
270
+
271
+ These calls hit real APIs and cost real money. `eval` is never part of
272
+ `pytest`: it runs only when you invoke it, with credentials you supply.
273
+ Completions are cached under `.agentskills/eval-cache` by model, system prompt,
274
+ user prompt, and repeat index — so editing a skill re-buys its runs, while
275
+ tightening an expectation re-grades the answers already bought. `--no-cache`
276
+ turns that off; `--cache-dir` moves it.
277
+
278
+ ### `serve`
279
+
280
+ ```bash
281
+ agentskills serve ./skills --transport stdio
282
+ ```
283
+
284
+ Runs the MCP server over a folder of skills without hand-writing a config
285
+ file. For anything beyond a single filesystem root — HTTP providers,
286
+ per-skill options, environment placeholders — use
287
+ [agentskills-mcp-server](https://github.com/pratikxpanda/agentskills-sdk/tree/main/packages/integrations/agentskills-mcp-server)
288
+ with a `server.json`.
289
+
290
+ ## Exit codes
291
+
292
+ | Code | Meaning |
293
+ | --- | --- |
294
+ | `0` | Ran, found nothing wrong. |
295
+ | `1` | Ran, found errors — or warnings under `--strict`, or a cost over budget. |
296
+ | `2` | Could not run: bad path, missing extra, unwritable directory. |
297
+
298
+ The distinction matters in CI: `1` means a skill is broken, `2` means the
299
+ invocation is.
300
+
301
+ ## JSON output
302
+
303
+ `validate`, `lint`, `inspect`, and `eval` accept `--format json`. The schema is
304
+ a published contract; `schemaVersion` is bumped only for a breaking change, and
305
+ new fields are added rather than existing ones repurposed.
306
+
307
+ ```json
308
+ {
309
+ "schemaVersion": 1,
310
+ "command": "validate",
311
+ "ok": false,
312
+ "summary": { "skills": 2, "errors": 1, "warnings": 0 },
313
+ "skills": [
314
+ {
315
+ "id": "broken-skill",
316
+ "path": "skills/broken-skill",
317
+ "ok": false,
318
+ "findings": [
319
+ {
320
+ "severity": "error",
321
+ "code": "frontmatter-invalid-yaml",
322
+ "message": "frontmatter is not valid YAML: mapping values are not allowed here",
323
+ "line": 3,
324
+ "file": "skills/broken-skill/SKILL.md"
325
+ }
326
+ ]
327
+ }
328
+ ]
329
+ }
330
+ ```
331
+
332
+ `ok` mirrors the exit code, so a consumer never has to re-derive the
333
+ strictness rules. `line` is `null` unless the problem can be attributed to one
334
+ line. `file` is the skill's `SKILL.md` unless the finding is about another file
335
+ in the folder, such as an eval case file.
336
+
337
+ `inspect --cost --format json` reports each skill's `perTurn`, `perLoad` and
338
+ `onDemand` totals, the `sections` and `resources` they were summed from, the
339
+ `overBudget` messages, and the `counter` that produced the numbers — including
340
+ whether it was `exact`. A consumer that charts these over time needs to know
341
+ when the unit changed underneath it.
342
+
343
+ ## Continuous integration
344
+
345
+ The published action wraps `validate` and `lint` and annotates every finding
346
+ on the pull request diff:
347
+
348
+ ```yaml
349
+ - uses: pratikxpanda/agentskills-sdk/actions/validate@v1
350
+ with:
351
+ path: ./skills
352
+ fail-on-lint: false
353
+ ```
354
+
355
+ To run it yourself, `validate` and `lint` also accept `--format github`, which
356
+ emits [workflow commands](https://docs.github.com/actions/reference/workflow-commands-for-github-actions)
357
+ instead of a report:
358
+
359
+ ```bash
360
+ agentskills validate ./skills --format github
361
+ ```
362
+
363
+ ```text
364
+ ::error file=skills/deploy/SKILL.md,line=3,title=frontmatter-invalid-yaml::frontmatter is not valid YAML
365
+ ```
366
+
367
+ Anywhere else, the exit code is enough:
368
+
369
+ ```yaml
370
+ - run: pip install agentskills-tools
371
+ - run: agentskills validate ./skills
372
+ ```
373
+
374
+ ## Logging
375
+
376
+ Pass `-v` to send the SDK's debug logs to stderr, leaving stdout parseable:
377
+
378
+ ```bash
379
+ agentskills validate ./skills --format json -v > report.json
380
+ ```
381
+
382
+ ## Security
383
+
384
+ Agent Skills are **equivalent to executable code** — skill content is injected
385
+ into an LLM agent's context verbatim. Validating a skill does not make it safe
386
+ to run. **Only load skills from sources you trust.**
387
+
388
+ See
389
+ [SECURITY.md](https://github.com/pratikxpanda/agentskills-sdk/blob/main/SECURITY.md).
390
+
391
+ ## License
392
+
393
+ MIT
394
+