@mrciphersmith/keryx 0.2.164 → 0.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +4 -1
- package/dist/cli.js +82540 -50300
- package/dist/core.js +28967 -18937
- package/package.json +2 -2
- package/src/gdgraph/affected-report.ts +141 -0
- package/src/gdgraph/build.ts +170 -23
- package/src/gdgraph/service.ts +6 -0
- package/src/gdgraph/staleness.ts +253 -45
- package/src/gdskills/bundled/agents/codebase-navigator.md +55 -0
- package/src/gdskills/bundled/agents/design-advisor.md +64 -0
- package/src/gdskills/bundled/agents/docs-maintainer.md +56 -0
- package/src/gdskills/bundled/agents/end-to-end-tester.md +56 -0
- package/src/gdskills/bundled/agents/error-path-auditor.md +57 -0
- package/src/gdskills/bundled/agents/go-build-fixer.md +52 -0
- package/src/gdskills/bundled/agents/go-code-auditor.md +49 -0
- package/src/gdskills/bundled/agents/performance-auditor.md +63 -0
- package/src/gdskills/bundled/agents/python-build-fixer.md +52 -0
- package/src/gdskills/bundled/agents/python-code-auditor.md +49 -0
- package/src/gdskills/bundled/agents/refactoring-steward.md +61 -0
- package/src/gdskills/bundled/agents/security-auditor.md +62 -0
- package/src/gdskills/bundled/agents/test-first-driver.md +61 -0
- package/src/gdskills/bundled/agents/work-planner.md +62 -0
- package/src/gdskills/bundled/install-manifest.json +797 -0
- package/src/gdskills/bundled/rules/core/model-selection.mdc +51 -0
- package/src/gdskills/bundled/rules/core/skill-lifecycle.mdc +29 -1
- package/src/gdskills/bundled/rules/core/skills-storage-workflow.mdc +2 -2
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/review/review-jev-rules/SKILL.md +267 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.detail.md +26 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +75 -247
- package/src/gdskills/bundled/skills/review/review-orchestrator/output-contract.schema.json +19 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/reviewer-finding.schema.json +10 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/reviewer-input.schema.json +5 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/templates/pr-comment-backend.md +50 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/templates/pr-comment-frontend.md +52 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/templates/review-report.md +143 -0
- package/src/gdskills/bundled/stacks/angular/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/angular/governance/eval.json +1751 -0
- package/src/gdskills/bundled/stacks/angular/governance/scout.json +32 -0
- package/src/gdskills/bundled/stacks/angular/pack.json +55 -0
- package/src/gdskills/bundled/stacks/angular/rules/coding-style.mdc +82 -0
- package/src/gdskills/bundled/stacks/angular/rules/patterns.mdc +84 -0
- package/src/gdskills/bundled/stacks/angular/rules/security.mdc +70 -0
- package/src/gdskills/bundled/stacks/angular/rules/testing.mdc +73 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-build-fix/SKILL.md +127 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-build-fix/evals.json +72 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-code-review/SKILL.md +98 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-code-review/evals.json +73 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-implementation/SKILL.md +112 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-implementation/evals.json +74 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-testing/SKILL.md +102 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-testing/evals.json +71 -0
- package/src/gdskills/bundled/stacks/go/agent-refs.json +3 -0
- package/src/gdskills/bundled/stacks/go/governance/eval.json +1745 -0
- package/src/gdskills/bundled/stacks/go/governance/scout.json +31 -0
- package/src/gdskills/bundled/stacks/go/pack.json +41 -0
- package/src/gdskills/bundled/stacks/go/rules/coding-style.mdc +85 -0
- package/src/gdskills/bundled/stacks/go/rules/patterns.mdc +65 -0
- package/src/gdskills/bundled/stacks/go/rules/security.mdc +73 -0
- package/src/gdskills/bundled/stacks/go/rules/testing.mdc +68 -0
- package/src/gdskills/bundled/stacks/go/skills/go-build-fix/SKILL.md +138 -0
- package/src/gdskills/bundled/stacks/go/skills/go-build-fix/evals.json +75 -0
- package/src/gdskills/bundled/stacks/go/skills/go-code-review/SKILL.md +121 -0
- package/src/gdskills/bundled/stacks/go/skills/go-code-review/evals.json +72 -0
- package/src/gdskills/bundled/stacks/go/skills/go-implementation/SKILL.md +122 -0
- package/src/gdskills/bundled/stacks/go/skills/go-implementation/evals.json +76 -0
- package/src/gdskills/bundled/stacks/go/skills/go-testing/SKILL.md +126 -0
- package/src/gdskills/bundled/stacks/go/skills/go-testing/evals.json +73 -0
- package/src/gdskills/bundled/stacks/mobx/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/mobx/governance/eval.json +904 -0
- package/src/gdskills/bundled/stacks/mobx/governance/scout.json +18 -0
- package/src/gdskills/bundled/stacks/mobx/pack.json +28 -0
- package/src/gdskills/bundled/stacks/mobx/rules/coding-style.mdc +91 -0
- package/src/gdskills/bundled/stacks/mobx/rules/patterns.mdc +122 -0
- package/src/gdskills/bundled/stacks/mobx/rules/security.mdc +56 -0
- package/src/gdskills/bundled/stacks/mobx/rules/testing.mdc +63 -0
- package/src/gdskills/bundled/stacks/mobx/skills/mobx-observable-testing/SKILL.md +124 -0
- package/src/gdskills/bundled/stacks/mobx/skills/mobx-observable-testing/evals.json +73 -0
- package/src/gdskills/bundled/stacks/mobx/skills/mobx-store-implementation/SKILL.md +149 -0
- package/src/gdskills/bundled/stacks/mobx/skills/mobx-store-implementation/evals.json +74 -0
- package/src/gdskills/bundled/stacks/nestjs/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/nestjs/governance/eval.json +1308 -0
- package/src/gdskills/bundled/stacks/nestjs/governance/scout.json +34 -0
- package/src/gdskills/bundled/stacks/nestjs/pack.json +53 -0
- package/src/gdskills/bundled/stacks/nestjs/rules/coding-style.mdc +70 -0
- package/src/gdskills/bundled/stacks/nestjs/rules/patterns.mdc +83 -0
- package/src/gdskills/bundled/stacks/nestjs/rules/security.mdc +73 -0
- package/src/gdskills/bundled/stacks/nestjs/rules/testing.mdc +69 -0
- package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-build-fix/SKILL.md +157 -0
- package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-build-fix/evals.json +70 -0
- package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-implementation/SKILL.md +129 -0
- package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-implementation/evals.json +71 -0
- package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-testing/SKILL.md +143 -0
- package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-testing/evals.json +69 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/governance/eval.json +2413 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/governance/scout.json +42 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/pack.json +42 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/coding-style.mdc +69 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/patterns.mdc +88 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/security.mdc +72 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/testing.mdc +64 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-build-fix/SKILL.md +147 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-build-fix/evals.json +75 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-code-review/SKILL.md +118 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-code-review/evals.json +76 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-implementation/SKILL.md +135 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-implementation/evals.json +78 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-testing/SKILL.md +116 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-testing/evals.json +75 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-upgrade-migration/SKILL.md +134 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-upgrade-migration/evals.json +76 -0
- package/src/gdskills/bundled/stacks/python/agent-refs.json +3 -0
- package/src/gdskills/bundled/stacks/python/governance/eval.json +1758 -0
- package/src/gdskills/bundled/stacks/python/governance/scout.json +34 -0
- package/src/gdskills/bundled/stacks/python/pack.json +41 -0
- package/src/gdskills/bundled/stacks/python/rules/coding-style.mdc +63 -0
- package/src/gdskills/bundled/stacks/python/rules/patterns.mdc +88 -0
- package/src/gdskills/bundled/stacks/python/rules/security.mdc +84 -0
- package/src/gdskills/bundled/stacks/python/rules/testing.mdc +77 -0
- package/src/gdskills/bundled/stacks/python/skills/python-build-fix/SKILL.md +144 -0
- package/src/gdskills/bundled/stacks/python/skills/python-build-fix/evals.json +74 -0
- package/src/gdskills/bundled/stacks/python/skills/python-code-review/SKILL.md +155 -0
- package/src/gdskills/bundled/stacks/python/skills/python-code-review/evals.json +72 -0
- package/src/gdskills/bundled/stacks/python/skills/python-implementation/SKILL.md +143 -0
- package/src/gdskills/bundled/stacks/python/skills/python-implementation/evals.json +78 -0
- package/src/gdskills/bundled/stacks/python/skills/python-testing/SKILL.md +132 -0
- package/src/gdskills/bundled/stacks/python/skills/python-testing/evals.json +73 -0
- package/src/gdskills/bundled/stacks/react/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/react/governance/eval.json +2188 -0
- package/src/gdskills/bundled/stacks/react/governance/scout.json +40 -0
- package/src/gdskills/bundled/stacks/react/pack.json +42 -0
- package/src/gdskills/bundled/stacks/react/rules/coding-style.mdc +58 -0
- package/src/gdskills/bundled/stacks/react/rules/patterns.mdc +79 -0
- package/src/gdskills/bundled/stacks/react/rules/security.mdc +70 -0
- package/src/gdskills/bundled/stacks/react/rules/testing.mdc +60 -0
- package/src/gdskills/bundled/stacks/react/skills/react-build-fix/SKILL.md +139 -0
- package/src/gdskills/bundled/stacks/react/skills/react-build-fix/evals.json +72 -0
- package/src/gdskills/bundled/stacks/react/skills/react-code-review/SKILL.md +148 -0
- package/src/gdskills/bundled/stacks/react/skills/react-code-review/evals.json +74 -0
- package/src/gdskills/bundled/stacks/react/skills/react-implementation/SKILL.md +140 -0
- package/src/gdskills/bundled/stacks/react/skills/react-implementation/evals.json +74 -0
- package/src/gdskills/bundled/stacks/react/skills/react-testing/SKILL.md +142 -0
- package/src/gdskills/bundled/stacks/react/skills/react-testing/evals.json +83 -0
- package/src/gdskills/bundled/stacks/react/skills/react-upgrade-migration/SKILL.md +155 -0
- package/src/gdskills/bundled/stacks/react/skills/react-upgrade-migration/evals.json +74 -0
- package/src/gdskills/bundled/stacks/ts-js-node/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/ts-js-node/governance/eval.json +2155 -0
- package/src/gdskills/bundled/stacks/ts-js-node/governance/scout.json +40 -0
- package/src/gdskills/bundled/stacks/ts-js-node/pack.json +41 -0
- package/src/gdskills/bundled/stacks/ts-js-node/rules/coding-style.mdc +73 -0
- package/src/gdskills/bundled/stacks/ts-js-node/rules/patterns.mdc +61 -0
- package/src/gdskills/bundled/stacks/ts-js-node/rules/security.mdc +71 -0
- package/src/gdskills/bundled/stacks/ts-js-node/rules/testing.mdc +63 -0
- package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-build-fix/SKILL.md +137 -0
- package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-build-fix/evals.json +73 -0
- package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-code-review/SKILL.md +124 -0
- package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-code-review/evals.json +74 -0
- package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-esm-migration/SKILL.md +152 -0
- package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-esm-migration/evals.json +71 -0
- package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-implementation/SKILL.md +127 -0
- package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-implementation/evals.json +72 -0
- package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-testing/SKILL.md +134 -0
- package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-testing/evals.json +70 -0
- package/src/gdskills/bundled/stacks/vue/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/vue/governance/eval.json +2215 -0
- package/src/gdskills/bundled/stacks/vue/governance/scout.json +42 -0
- package/src/gdskills/bundled/stacks/vue/pack.json +42 -0
- package/src/gdskills/bundled/stacks/vue/rules/coding-style.mdc +73 -0
- package/src/gdskills/bundled/stacks/vue/rules/patterns.mdc +84 -0
- package/src/gdskills/bundled/stacks/vue/rules/security.mdc +60 -0
- package/src/gdskills/bundled/stacks/vue/rules/testing.mdc +69 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-build-fix/SKILL.md +137 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-build-fix/evals.json +72 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-code-review/SKILL.md +120 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-code-review/evals.json +71 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-implementation/SKILL.md +122 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-implementation/evals.json +72 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-testing/SKILL.md +115 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-testing/evals.json +72 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue2-to-vue3-migration/SKILL.md +135 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue2-to-vue3-migration/evals.json +71 -0
|
@@ -0,0 +1,155 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: python-code-review
|
|
3
|
+
description: "Use when reviewing Python changes for correctness and safety risks -- checks mutable default arguments, broad except clauses, resource leaks missing a with block, blocking calls inside async functions, typing holes (Any, missing Optional), N+1/ORM query misuse, and security sinks. Read-only: reports findings, does not edit code."
|
|
4
|
+
triggers:
|
|
5
|
+
- "review this python diff"
|
|
6
|
+
- "check this python pr for bugs"
|
|
7
|
+
- "review python code changes"
|
|
8
|
+
- "audit this python module"
|
|
9
|
+
- "check for python security issues"
|
|
10
|
+
- "review this async python code"
|
|
11
|
+
metadata:
|
|
12
|
+
origin: authored
|
|
13
|
+
category: review
|
|
14
|
+
version: "1.0.0"
|
|
15
|
+
compatible_harnesses: "claude,codex,cursor,zed,opencode"
|
|
16
|
+
license: "MIT"
|
|
17
|
+
---
|
|
18
|
+
|
|
19
|
+
# Python code review
|
|
20
|
+
|
|
21
|
+
Review a set of Python changes for correctness, resource-safety, typing, and
|
|
22
|
+
security risk. Read-only: this skill reports findings, it never edits code.
|
|
23
|
+
Scoped to Python-specific defects; for generic architecture/style review use
|
|
24
|
+
the catalog's `review-*` skills, for a language-agnostic security sweep use
|
|
25
|
+
`review-security-code`, for fixing what this skill finds use
|
|
26
|
+
`python-build-fix` (checker failures) or hand the report to the author.
|
|
27
|
+
|
|
28
|
+
## Workflow
|
|
29
|
+
|
|
30
|
+
### Step 1: Scope the review
|
|
31
|
+
|
|
32
|
+
1. Identify the changed Python files (`git diff` against the review base,
|
|
33
|
+
or the files the requester names).
|
|
34
|
+
2. Read `pyproject.toml` for the project's configured `ruff`/`mypy`/
|
|
35
|
+
`pyright` rules — a finding this skill would raise that the project's
|
|
36
|
+
own linter already enforces and passes on is lower priority than one the
|
|
37
|
+
linter cannot catch (logic, resource, or security issues).
|
|
38
|
+
3. Read enough of the surrounding module (call sites, class definition) to
|
|
39
|
+
judge whether a flagged pattern is actually a bug in context, not just a
|
|
40
|
+
pattern match.
|
|
41
|
+
|
|
42
|
+
### Step 2: Check each changed file against these categories
|
|
43
|
+
|
|
44
|
+
**Mutable defaults and shared state**
|
|
45
|
+
- `def f(items=[])`/`def f(config={})` — a mutable default is shared and
|
|
46
|
+
mutated across every call that omits the argument.
|
|
47
|
+
- A module-level or class-level mutable used as implicit shared state
|
|
48
|
+
across requests/calls without synchronization.
|
|
49
|
+
|
|
50
|
+
**Exception handling**
|
|
51
|
+
- Bare `except:` or `except Exception:` that swallows an error the caller
|
|
52
|
+
needed to see, especially one that also catches
|
|
53
|
+
`asyncio.CancelledError`/`KeyboardInterrupt`/`SystemExit`.
|
|
54
|
+
- A re-raise inside an `except` block written as bare `raise NewError(...)`
|
|
55
|
+
with no `from exc`/`from None`. Python keeps the original exception either
|
|
56
|
+
way through implicit chaining (`__context__`, printed as "During handling
|
|
57
|
+
of the above exception..."), so this is not a lost traceback — it's a
|
|
58
|
+
missing statement of intent: `from exc` says the new exception was caused
|
|
59
|
+
by the original, `from None` says the original is deliberately suppressed
|
|
60
|
+
and shouldn't be shown.
|
|
61
|
+
- An `except` block that logs and continues where the correct behavior was
|
|
62
|
+
to propagate.
|
|
63
|
+
|
|
64
|
+
**Resource management**
|
|
65
|
+
- A file, socket, DB connection, lock, or temp resource opened without
|
|
66
|
+
`with`/`async with`, or closed only on the happy path (no `finally`/
|
|
67
|
+
context manager covering the exception path).
|
|
68
|
+
- A context manager whose `__exit__`/`finally` doesn't actually run cleanup
|
|
69
|
+
when the body raises.
|
|
70
|
+
|
|
71
|
+
**Async correctness**
|
|
72
|
+
- A blocking call (`requests.get`, `time.sleep`, synchronous file I/O, a
|
|
73
|
+
CPU-bound loop) directly inside an `async def` instead of the async
|
|
74
|
+
client, `asyncio.sleep`, or `asyncio.to_thread`.
|
|
75
|
+
- A coroutine created but never awaited (`asyncio.create_task` result
|
|
76
|
+
discarded with no reference kept, or a bare `coro()` call with no
|
|
77
|
+
`await`).
|
|
78
|
+
- `asyncio.gather`/manual task tracking where a `TaskGroup` would give
|
|
79
|
+
correct sibling-cancellation semantics, if the project targets 3.11+.
|
|
80
|
+
|
|
81
|
+
**Typing**
|
|
82
|
+
- A new/touched public function with no type hints, or a hint that is
|
|
83
|
+
`Any` where a `Protocol`/union/`TypeVar` would express the real
|
|
84
|
+
constraint.
|
|
85
|
+
- A parameter or return that can be `None` at some call site but is typed
|
|
86
|
+
without `Optional`/`| None`.
|
|
87
|
+
- A type: ignore/noqa added to silence a real typing/lint issue rather than
|
|
88
|
+
fixing it.
|
|
89
|
+
|
|
90
|
+
**Data access**
|
|
91
|
+
- A loop that issues one query per iteration (N+1) where a single
|
|
92
|
+
batched/joined query or `select_related`/`prefetch_related` (Django) or
|
|
93
|
+
equivalent eager-load would do.
|
|
94
|
+
- An ORM query built by interpolating a value into a raw SQL string instead
|
|
95
|
+
of using parameter binding or the ORM's query builder.
|
|
96
|
+
|
|
97
|
+
**Security** (full list: `rules/security.mdc`)
|
|
98
|
+
- `subprocess` with `shell=True` or a string command.
|
|
99
|
+
- `eval`/`exec` on anything that could carry untrusted input.
|
|
100
|
+
- `pickle.load`/`yaml.load` (not `safe_load`) on data not fully controlled
|
|
101
|
+
by the project.
|
|
102
|
+
- SQL built by string interpolation instead of parameters.
|
|
103
|
+
- A network call with no `timeout`, or `verify=False`.
|
|
104
|
+
- A hard-coded secret, or a token generated with `random` instead of
|
|
105
|
+
`secrets`.
|
|
106
|
+
|
|
107
|
+
### Step 3: Report
|
|
108
|
+
|
|
109
|
+
For each finding: file:line, category, what's wrong, and the safe
|
|
110
|
+
alternative (cite the exact API/pattern, e.g. "use `secrets.token_urlsafe`
|
|
111
|
+
instead of `random.random()`"). Group by severity — a security sink or a
|
|
112
|
+
resource leak outranks a missing type hint.
|
|
113
|
+
|
|
114
|
+
```
|
|
115
|
+
python-code-review: 3 findings
|
|
116
|
+
[security] auth.py:42 — `subprocess.run(cmd, shell=True)`; pass an
|
|
117
|
+
argument list instead
|
|
118
|
+
[resource] client.py:18 — `open()` with no `with`; the handle leaks if
|
|
119
|
+
`json.load` raises
|
|
120
|
+
[typing] models.py:9 — `def find(id) -> User:` has no param type and can
|
|
121
|
+
return `None`; add `id: int` and `-> User | None`
|
|
122
|
+
```
|
|
123
|
+
|
|
124
|
+
## Rules
|
|
125
|
+
|
|
126
|
+
- NEVER edit the files under review — report findings only.
|
|
127
|
+
- Cite the specific line and the specific safe alternative; a vague "this
|
|
128
|
+
could be an issue" finding is not actionable.
|
|
129
|
+
- Do not duplicate a finding the project's own configured `ruff`/`mypy`
|
|
130
|
+
rules already enforce and would catch on their own — focus on what static
|
|
131
|
+
tooling misses (resource lifetime across exception paths, N+1 patterns,
|
|
132
|
+
logic bugs, security sinks needing call-site context).
|
|
133
|
+
- Distinguish a real bug from a stylistic preference; a stylistic point
|
|
134
|
+
belongs in `rules/coding-style.mdc`/`rules/patterns.mdc`, not a review
|
|
135
|
+
finding blocking the change.
|
|
136
|
+
|
|
137
|
+
## Red Flags
|
|
138
|
+
|
|
139
|
+
| Rationalization | Why it is wrong |
|
|
140
|
+
|---|---|
|
|
141
|
+
| "The `except Exception` here is fine, it just logs" | Logging and continuing after swallowing an exception still hides the failure from the caller and from any monitoring keyed off the exception propagating |
|
|
142
|
+
| "It's just a review, I'll fix the mutable default myself since it's a one-liner" | This skill is read-only; even a trivial fix belongs to the author or `python-build-fix`, not a silent edit during review |
|
|
143
|
+
| "The N+1 loop only ever runs over 3 items in tests" | Test data size does not bound production data size; flag it regardless of the loop's current call sites |
|
|
144
|
+
|
|
145
|
+
## Verification
|
|
146
|
+
|
|
147
|
+
Do not report the review done until all of the following hold:
|
|
148
|
+
|
|
149
|
+
- Every changed Python file in scope was checked against all seven
|
|
150
|
+
categories in Step 2.
|
|
151
|
+
- No finding duplicates something the project's own configured linter/type
|
|
152
|
+
checker already flags and enforces.
|
|
153
|
+
- Every finding names a file:line, the specific problem, and a specific
|
|
154
|
+
fix — no vague findings.
|
|
155
|
+
- No file under review was modified.
|
|
@@ -0,0 +1,72 @@
|
|
|
1
|
+
{
|
|
2
|
+
"triggers": {
|
|
3
|
+
"positive": [
|
|
4
|
+
"Review this Python diff for mutable default argument bugs",
|
|
5
|
+
"Check this pull request for broad except clauses swallowing errors",
|
|
6
|
+
"Audit this Python module for resource leaks and unclosed files",
|
|
7
|
+
"Review this async Python code for blocking calls inside async def",
|
|
8
|
+
"Check this Python code review for typing holes like Any or missing Optional",
|
|
9
|
+
"Review this Django ORM code for N+1 query issues"
|
|
10
|
+
],
|
|
11
|
+
"negative": [
|
|
12
|
+
"Implement a new Python feature that fetches user data",
|
|
13
|
+
"Write pytest tests for this module",
|
|
14
|
+
"Fix the mypy type errors so the build passes",
|
|
15
|
+
"Review this TypeScript React component for MobX store misuse",
|
|
16
|
+
"Run review-security-code on our whole repository for OWASP issues",
|
|
17
|
+
"Fix this ModuleNotFoundError in our Python package"
|
|
18
|
+
]
|
|
19
|
+
},
|
|
20
|
+
"scenarios": [
|
|
21
|
+
{
|
|
22
|
+
"id": "mutable-default-finding",
|
|
23
|
+
"prompt": "Review this Python function for bugs:\n\ndef add_item(item, items=[]):\n items.append(item)\n return items",
|
|
24
|
+
"strictness": "high",
|
|
25
|
+
"expected_behavior": [
|
|
26
|
+
{
|
|
27
|
+
"grader": "judge",
|
|
28
|
+
"rubric": "A correct review flags `items=[]` as a mutable default argument that is created once and shared/mutated across every call that omits `items`, and gives the standard fix, without editing the code itself.",
|
|
29
|
+
"pass_criteria": [
|
|
30
|
+
"identifies that `items=[]` is a mutable default argument evaluated once at function definition, so it is shared and accumulates state across calls that don't pass their own list",
|
|
31
|
+
"gives the standard fix and shows the corrected code: default the parameter to `None` and create a new list inside the function body when it is `None`, e.g. `def add_item(item, items=None): if items is None: items = []` -- not just a description of the fix in the abstract",
|
|
32
|
+
"presents this as a specific finding on this function, not a generic style remark, and does not modify the code -- this skill only reports findings"
|
|
33
|
+
],
|
|
34
|
+
"fail_criteria": [
|
|
35
|
+
"declares the function has no bug or is fine as written, missing the mutable-default issue entirely"
|
|
36
|
+
]
|
|
37
|
+
}
|
|
38
|
+
],
|
|
39
|
+
"calibration": {
|
|
40
|
+
"known_right": "Finding: `add_item`'s `items=[]` default is a mutable default argument. Python evaluates default argument values once, at function-definition time, so every call to `add_item(item)` that omits `items` shares and mutates the exact same list object -- items appended in one call are still there on the next call, which is almost certainly not the intended behavior and will surface as a confusing bug wherever this is called more than once without an explicit `items=`. Fix: change the signature to `def add_item(item, items=None):` and create a fresh list inside the body when needed: `if items is None: items = []`. That guarantees each call that doesn't pass its own list gets a brand-new empty one instead of the shared default. This is a correctness bug, not a style nit, so I'd flag it as the primary finding here rather than suggesting an edit myself, since this review is read-only.",
|
|
41
|
+
"known_wrong": "This function looks fine to me -- `add_item` just appends an item to a list and returns it, which is exactly what it's supposed to do. The `items=[]` default is a convenient shortcut so callers don't have to pass an empty list every time; Python handles default arguments the same way most other languages do, evaluating them fresh per call. No changes needed here, the logic is straightforward and there's nothing that would cause unexpected behavior in normal use.",
|
|
42
|
+
"vague": "There's a subtle bug with the default list argument here -- mutable defaults in Python are generally risky and worth avoiding.",
|
|
43
|
+
"subtle_wrong": "Worth a note: `items=[]` is evaluated once at definition time, so calling `add_item` repeatedly without passing `items` would reuse the same list. I'd suggest switching to `items=None` with a fallback inside the body, though in practice this is more of a minor style nit than a real bug unless the function actually gets called that way in this codebase."
|
|
44
|
+
}
|
|
45
|
+
},
|
|
46
|
+
{
|
|
47
|
+
"id": "broad-except-finding",
|
|
48
|
+
"prompt": "Review this Python code:\n\ntry:\n process(data)\nexcept Exception:\n pass",
|
|
49
|
+
"strictness": "high",
|
|
50
|
+
"expected_behavior": [
|
|
51
|
+
{
|
|
52
|
+
"grader": "judge",
|
|
53
|
+
"rubric": "A correct review flags `except Exception: pass` as swallowing whatever error `process` raises rather than treating a broad catch that silently discards the error as acceptable.",
|
|
54
|
+
"pass_criteria": [
|
|
55
|
+
"identifies that `except Exception: pass` swallows the error instead of letting it propagate, being logged, or being handled specifically",
|
|
56
|
+
"explains the consequence: the caller and any monitoring keyed off the exception never see the failure, so real bugs go unnoticed",
|
|
57
|
+
"states this as a concrete finding, naming a specific alternative action -- log the exception with context (e.g. `logging.exception(...)`), re-raise, or catch a specific exception type -- not a vague 'handle it better'"
|
|
58
|
+
],
|
|
59
|
+
"fail_criteria": [
|
|
60
|
+
"dismisses the bare `except Exception: pass` as fine or acceptable because it simply logs or prevents a crash, without flagging that the error is being hidden"
|
|
61
|
+
]
|
|
62
|
+
}
|
|
63
|
+
],
|
|
64
|
+
"calibration": {
|
|
65
|
+
"known_right": "Finding: the bare `except Exception: pass` around `process(data)` swallows any error the call raises with no logging, no re-raise, and no fallback -- the failure just vanishes. If `process` fails for a genuinely unexpected reason (bad data, a downstream service error, a bug), nothing here surfaces it: the caller sees a silent no-op success, and any monitoring or alerting that depends on the exception propagating never fires. This also catches things a broad handler shouldn't, since `Exception` covers essentially every non-system-exiting error. At minimum this should log the exception with context before continuing, e.g. `logging.exception(\"process failed\")`, and ideally it should catch only the specific exception types `process` can legitimately raise and let anything else propagate, so real bugs aren't hidden behind a passing test run.",
|
|
66
|
+
"known_wrong": "This is fine -- catching `Exception` and passing just means a failure in `process(data)` won't crash the rest of the program, which is generally what you want in production code. It's a defensive pattern: if `process` has some edge case that throws, the surrounding code keeps running instead of taking down the whole request. I wouldn't flag this as a problem; if anything it shows the author was being careful about robustness. No changes needed.",
|
|
67
|
+
"vague": "This broad except could hide problems, so it would be safer to handle errors a bit more carefully here.",
|
|
68
|
+
"subtle_wrong": "Catching `Exception` is fine as a safety net so a failure in `process` doesn't take down the caller. I'd just add a one-line `logging.exception(\"process failed\")` before the `pass` so at least it's not completely silent -- keeps the same broad catch but gives you some visibility if it ever fires."
|
|
69
|
+
}
|
|
70
|
+
}
|
|
71
|
+
]
|
|
72
|
+
}
|
|
@@ -0,0 +1,143 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: python-implementation
|
|
3
|
+
description: "Use when implementing or extending a feature in a modern Python (3.12/3.13) codebase or service -- covers project tooling discovery (pyproject.toml, uv/poetry/pip, ruff, mypy/pyright), typing (generics, Protocol, TypedDict, dataclasses), context managers, exception chaining, asyncio TaskGroup, request/event logging with the standard logging module, and src/-layout packaging."
|
|
4
|
+
triggers:
|
|
5
|
+
- "implement this in python"
|
|
6
|
+
- "add a python feature"
|
|
7
|
+
- "write a python function"
|
|
8
|
+
- "add a dataclass"
|
|
9
|
+
- "implement asyncio taskgroup"
|
|
10
|
+
- "add type hints to this module"
|
|
11
|
+
- "package this as a python module"
|
|
12
|
+
- "add a protocol class"
|
|
13
|
+
- "log requests in this python service"
|
|
14
|
+
metadata:
|
|
15
|
+
origin: authored
|
|
16
|
+
category: implement
|
|
17
|
+
version: "1.0.0"
|
|
18
|
+
compatible_harnesses: "claude,codex,cursor,zed,opencode"
|
|
19
|
+
license: "MIT"
|
|
20
|
+
---
|
|
21
|
+
|
|
22
|
+
# Python implementation
|
|
23
|
+
|
|
24
|
+
Implement or extend a feature in a modern Python (3.12/3.13) codebase:
|
|
25
|
+
discovering the project's own tooling and conventions before writing code,
|
|
26
|
+
then applying current-practice typing, resource management, error handling,
|
|
27
|
+
concurrency, logging, and packaging patterns. Scoped to writing production
|
|
28
|
+
code — for writing/fixing tests use `python-testing`, for reviewing a diff
|
|
29
|
+
without editing it use `python-code-review`, for fixing a broken build/lint/
|
|
30
|
+
type-check without adding a feature use `python-build-fix`.
|
|
31
|
+
|
|
32
|
+
## Workflow
|
|
33
|
+
|
|
34
|
+
### Step 1: Discover the project's own tooling and conventions
|
|
35
|
+
|
|
36
|
+
1. Read `pyproject.toml` for: build backend, dependency manager (`uv`,
|
|
37
|
+
`poetry`, plain `pip`+`requirements.txt`), configured tools
|
|
38
|
+
(`[tool.ruff]`, `[tool.mypy]`/`[tool.pyright]`, `[tool.pytest.ini_options]`),
|
|
39
|
+
and the declared minimum Python version (`requires-python`).
|
|
40
|
+
2. Confirm the package layout: `src/<package>/` (modern) vs. a flat
|
|
41
|
+
`<package>/` at repo root — place new modules consistently with what is
|
|
42
|
+
already there, don't introduce the other layout.
|
|
43
|
+
3. Read 1-2 neighboring modules for: docstring style, typing style (PEP 604
|
|
44
|
+
`X | None` vs. `Optional[X]`), import grouping, and whether the project
|
|
45
|
+
already uses `from __future__ import annotations`.
|
|
46
|
+
4. Check `requires-python`/CI config for the actual supported Python
|
|
47
|
+
versions before using a 3.12+-only feature (type-parameter syntax
|
|
48
|
+
`class Box[T]:`, `except*`) in a project that must also run on 3.11 or
|
|
49
|
+
earlier.
|
|
50
|
+
|
|
51
|
+
### Step 2: Design the change
|
|
52
|
+
|
|
53
|
+
- Prefer extending an existing module/class over adding a new one for a
|
|
54
|
+
small change; add a new module when the change introduces a genuinely new
|
|
55
|
+
concern.
|
|
56
|
+
- Pick the data-modeling shape for the job: `@dataclass` for a fixed set of
|
|
57
|
+
related fields your own code constructs, `TypedDict` for a dict shape
|
|
58
|
+
that must stay a `dict` (e.g. JSON), `Protocol` for "anything with this
|
|
59
|
+
method" instead of a concrete base class, `Generic`/type-parameter syntax
|
|
60
|
+
for a container whose element type varies by call site. See
|
|
61
|
+
`rules/patterns.mdc` for the full set.
|
|
62
|
+
- Decide error handling up front: which specific exception types the new
|
|
63
|
+
code raises, and which (if any) it must catch — never plan around a bare
|
|
64
|
+
`except:`/`except Exception:`.
|
|
65
|
+
|
|
66
|
+
### Step 3: Implement
|
|
67
|
+
|
|
68
|
+
1. Type-hint every new public function signature, including `| None`/
|
|
69
|
+
`Optional` where a parameter or return can be absent (match the
|
|
70
|
+
project's PEP 604 vs. `Optional` convention from Step 1).
|
|
71
|
+
2. Acquire any closable/lockable resource with `with`/`async with`; write a
|
|
72
|
+
custom context manager with `@contextlib.contextmanager` unless the type
|
|
73
|
+
also needs other methods.
|
|
74
|
+
3. Raise a specific exception type, chaining with `raise NewError(...) from
|
|
75
|
+
exc` when re-raising inside an `except` block.
|
|
76
|
+
4. For concurrent work, group related awaitables with `asyncio.TaskGroup`
|
|
77
|
+
and let `asyncio.CancelledError` propagate through cleanup rather than
|
|
78
|
+
swallowing it; never call a blocking function inside `async def` (use
|
|
79
|
+
the async client, `asyncio.sleep`, or `asyncio.to_thread`).
|
|
80
|
+
5. Log through the standard `logging` module (or the project's structured
|
|
81
|
+
logger) with lazy `%s` interpolation, not `print` or an f-string passed
|
|
82
|
+
to the logging call.
|
|
83
|
+
6. Follow `rules/coding-style.mdc` for naming/formatting/imports and
|
|
84
|
+
`rules/patterns.mdc` for the rest of the idiom; check `rules/security.mdc`
|
|
85
|
+
before touching subprocess calls, deserialization, SQL, file paths, or
|
|
86
|
+
secrets.
|
|
87
|
+
|
|
88
|
+
### Step 4: Verify
|
|
89
|
+
|
|
90
|
+
```bash
|
|
91
|
+
ruff check .
|
|
92
|
+
ruff format --check .
|
|
93
|
+
mypy . # or: pyright
|
|
94
|
+
pytest -x -q
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
Prefix each command with the project's own run prefix when it uses one
|
|
98
|
+
(`uv run ruff check .`, `poetry run pytest -x -q`) — discovered in Step 1
|
|
99
|
+
from `pyproject.toml`/lockfile presence, not assumed.
|
|
100
|
+
|
|
101
|
+
### Step 5: Report
|
|
102
|
+
|
|
103
|
+
```
|
|
104
|
+
Implemented: src/mypkg/feature.py
|
|
105
|
+
- added `Widget` dataclass + `build_widget()`
|
|
106
|
+
- ruff/mypy/pytest: all green
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
## Rules
|
|
110
|
+
|
|
111
|
+
- ALWAYS discover and match the project's own tooling (Step 1) before
|
|
112
|
+
assuming `uv`/`poetry`/`pip`, `mypy`/`pyright`, or a specific Python
|
|
113
|
+
version.
|
|
114
|
+
- NEVER add `# type: ignore` or `# noqa` to route around a real typing or
|
|
115
|
+
lint problem — fix the underlying code, or narrow the suppression to the
|
|
116
|
+
one line with a comment explaining why it is correct as written.
|
|
117
|
+
- NEVER use a Python version feature the project's `requires-python` does
|
|
118
|
+
not support.
|
|
119
|
+
- Match the project's existing docstring style rather than introducing a
|
|
120
|
+
second one in the same file.
|
|
121
|
+
|
|
122
|
+
## Red Flags
|
|
123
|
+
|
|
124
|
+
| Rationalization | Why it is wrong |
|
|
125
|
+
|---|---|
|
|
126
|
+
| "I'll just use `Any` here, typing this properly is fiddly" | Defeats the type checker for every downstream caller; use `Protocol`/`Generic`/a union instead, or `TypeVar` if the shape is genuinely generic |
|
|
127
|
+
| "This helper doesn't need a `with` block, I'll just call `.close()` at the end" | Skips cleanup the moment an exception is raised before that line runs; use `with`/`try`/`finally` |
|
|
128
|
+
| "I'll catch `Exception` broadly so nothing crashes" | Hides bugs in unrelated code paths as silently-ignored failures; catch the specific exception type you can actually handle |
|
|
129
|
+
| "3.12 syntax is cleaner, I'll use it even though `requires-python` says 3.10" | Breaks on every environment still running the declared minimum version |
|
|
130
|
+
|
|
131
|
+
## Verification
|
|
132
|
+
|
|
133
|
+
Do not report the work done until all of the following hold:
|
|
134
|
+
|
|
135
|
+
- New/touched public functions are type-hinted per Step 3.1, matching the
|
|
136
|
+
project's `Optional`/`| None` convention from Step 1.
|
|
137
|
+
- `ruff check .`, `ruff format --check .`, and `mypy .`/`pyright` (or the
|
|
138
|
+
project's own configured equivalents) exit 0 on the touched files.
|
|
139
|
+
- `pytest -x -q` (or the project's configured runner) passes; if a feature
|
|
140
|
+
needs new tests, hand off to `python-testing` rather than writing test
|
|
141
|
+
files as part of this skill's own change set when out of scope.
|
|
142
|
+
- No bare `except:`/`except Exception:`, no unclosed resource, no blocking
|
|
143
|
+
call inside `async def`, introduced by this change.
|
|
@@ -0,0 +1,78 @@
|
|
|
1
|
+
{
|
|
2
|
+
"triggers": {
|
|
3
|
+
"positive": [
|
|
4
|
+
"Implement a Widget dataclass and a build_widget() function in src/mypkg/widgets.py",
|
|
5
|
+
"Add a new async function that fetches user records with asyncio.TaskGroup",
|
|
6
|
+
"Write a Python module that reads a config file and returns a typed Protocol object",
|
|
7
|
+
"I need to add type hints to this untyped Python function",
|
|
8
|
+
"Implement a context manager for our database connection pool in Python",
|
|
9
|
+
"Add a new feature to this Python service that logs each request with the logging module",
|
|
10
|
+
"Package this Python code with a src/ layout and pyproject.toml entry point"
|
|
11
|
+
],
|
|
12
|
+
"negative": [
|
|
13
|
+
"Write pytest tests for the widgets module I just added",
|
|
14
|
+
"Review this Python pull request for bugs before I merge it",
|
|
15
|
+
"Fix this ModuleNotFoundError when importing mypkg.util",
|
|
16
|
+
"Implement this new feature in TypeScript for our Node service",
|
|
17
|
+
"Review the frontend React component for MobX store violations",
|
|
18
|
+
"Run a generic security audit across the whole repository"
|
|
19
|
+
]
|
|
20
|
+
},
|
|
21
|
+
"scenarios": [
|
|
22
|
+
{
|
|
23
|
+
"id": "async-taskgroup-feature",
|
|
24
|
+
"prompt": "Implement a Python async function `fetch_all(urls)` that fetches several URLs concurrently and returns their results. Use current best practice for grouping the concurrent calls.",
|
|
25
|
+
"strictness": "high",
|
|
26
|
+
"expected_behavior": [
|
|
27
|
+
{
|
|
28
|
+
"grader": "regex",
|
|
29
|
+
"value": "[Tt]ask[Gg]roup"
|
|
30
|
+
},
|
|
31
|
+
{
|
|
32
|
+
"grader": "judge",
|
|
33
|
+
"rubric": "A correct answer implements `fetch_all` using an async-compatible HTTP client and groups the concurrent fetches with `asyncio.TaskGroup` (structured concurrency), never a blocking synchronous call inside the async function.",
|
|
34
|
+
"pass_criteria": [
|
|
35
|
+
"groups the concurrent fetches with `asyncio.TaskGroup` in actual code, not merely stated as an intention -- a bare list of coroutines, `asyncio.gather`, or unmanaged `create_task` calls does not count",
|
|
36
|
+
"performs the HTTP fetch with an async-capable client (e.g. `httpx.AsyncClient`, `aiohttp`) shown in code, rather than a blocking call inside the coroutine",
|
|
37
|
+
"returns the aggregated results for all the URLs once the group completes"
|
|
38
|
+
],
|
|
39
|
+
"fail_criteria": [
|
|
40
|
+
"calls a blocking function such as `requests.get` directly inside the async function instead of an async client. Mentioning it only to warn against it is not a failure."
|
|
41
|
+
]
|
|
42
|
+
}
|
|
43
|
+
],
|
|
44
|
+
"calibration": {
|
|
45
|
+
"known_right": "```python\nimport asyncio\nimport httpx\n\nasync def fetch_all(urls: list[str]) -> list[httpx.Response]:\n results: list[httpx.Response] = [None] * len(urls) # type: ignore[list-item]\n async with httpx.AsyncClient() as client:\n async with asyncio.TaskGroup() as tg:\n async def fetch(i: int, url: str) -> None:\n results[i] = await client.get(url)\n for i, url in enumerate(urls):\n tg.create_task(fetch(i, url))\n return results\n```\nThis groups every fetch under one `asyncio.TaskGroup`, which gives correct sibling-cancellation: if one request fails, the group cancels the rest and raises an `ExceptionGroup` instead of leaving orphaned tasks running. Fetches go through `httpx.AsyncClient`, an async HTTP client -- never `requests.get`, which is a blocking call and would stall the event loop and defeat the concurrency entirely. Results are collected by index so the return order matches the input `urls` order.",
|
|
46
|
+
"known_wrong": "```python\nimport asyncio\nimport requests\n\nasync def fetch_all(urls):\n results = []\n for url in urls:\n results.append(requests.get(url))\n return results\n```\nThis loops through each URL and fetches it with `requests.get`, which is simple and reliable since `requests` is the standard, well-tested HTTP library everyone already knows. Wrapping the function in `async def` lets callers `await fetch_all(urls)` from other async code without needing a separate sync wrapper, and there's no need to pull in `asyncio.TaskGroup` or an async HTTP client just for a handful of URLs -- `requests.get` is a blocking function, but the sequential requests will still complete and return everything the caller needs.",
|
|
47
|
+
"vague": "Use asyncio to fire off the requests concurrently and gather the results with proper structured concurrency instead of doing them one at a time.",
|
|
48
|
+
"subtle_wrong": "```python\nimport asyncio\nimport httpx\n\nasync def fetch_all(urls):\n async with httpx.AsyncClient() as client:\n tasks = [asyncio.create_task(client.get(u)) for u in urls]\n return await asyncio.gather(*tasks)\n```\nThis fires all the requests concurrently with `asyncio.create_task` and `asyncio.gather`, which is async and non-blocking throughout -- `gather` does the same job here with less new syntax, and this codebase hasn't adopted the newer structured-concurrency APIs yet."
|
|
49
|
+
}
|
|
50
|
+
},
|
|
51
|
+
{
|
|
52
|
+
"id": "resource-with-block",
|
|
53
|
+
"prompt": "Write a Python function that opens a file, reads its JSON contents, and returns the parsed data.",
|
|
54
|
+
"strictness": "high",
|
|
55
|
+
"expected_behavior": [
|
|
56
|
+
{
|
|
57
|
+
"grader": "judge",
|
|
58
|
+
"rubric": "A correct answer guarantees the file handle is closed on every path, including when parsing raises, parses the contents with the standard `json` module, and returns the parsed data.",
|
|
59
|
+
"pass_criteria": [
|
|
60
|
+
"the file handle is closed on every path, including when parsing raises -- a `with` block around `open(...)`, or an API that closes internally such as `Path.read_text()`/`Path.read_bytes()`, both satisfy this; a bare `open()` followed by a manual `.close()` that only runs on the happy path does not",
|
|
61
|
+
"parses the file's contents with `json.load`/`json.loads` and returns the resulting data",
|
|
62
|
+
"shows the actual code implementing this, not just a description of the approach"
|
|
63
|
+
],
|
|
64
|
+
"fail_criteria": [
|
|
65
|
+
"opens the file without a `with` block and instead calls `.close()` manually at the end, so the handle leaks if an exception occurs before that line runs. Mentioning it only to warn against it is not a failure."
|
|
66
|
+
]
|
|
67
|
+
}
|
|
68
|
+
],
|
|
69
|
+
"calibration": {
|
|
70
|
+
"known_right": "```python\nimport json\nfrom pathlib import Path\n\ndef read_json(path: str | Path) -> object:\n with open(path, \"r\", encoding=\"utf-8\") as f:\n return json.load(f)\n```\nThe `with open(...)` context manager acquires the file handle and guarantees `f.close()` runs in every case -- including when `json.load(f)` raises because the file contains invalid JSON. I deliberately didn't write this as `f = open(path); data = json.load(f); f.close()`, since calling `.close()` at the end only runs on the happy path and leaks the handle whenever the line before it raises. `encoding=\"utf-8\"` is passed explicitly rather than relying on a platform default.",
|
|
71
|
+
"known_wrong": "```python\nimport json\n\ndef read_json(path):\n f = open(path)\n data = json.load(f)\n f.close()\n return data\n```\nThis opens the file, hands it to `json.load` to parse, and closes it right after before returning the parsed data -- straightforward and easy to follow line by line. Since `.close()` is called right after `json.load`, the handle doesn't stay open for long, so this should be fine for a normal config or data file without needing the extra `with` syntax.",
|
|
72
|
+
"vague": "Make sure the file handle gets closed properly even if the JSON turns out to be malformed, then return whatever was parsed.",
|
|
73
|
+
"subtle_wrong": "```python\nimport json\n\ndef read_json(path):\n return json.loads(open(path).read())\n```\nThis opens the file and reads it in one expression -- Python's reference counting closes the handle as soon as the temporary file object goes out of scope, so there's no separate `.close()` call to remember, and it's a common one-liner for a quick read."
|
|
74
|
+
},
|
|
75
|
+
"anti_patterns": [".close()"]
|
|
76
|
+
}
|
|
77
|
+
]
|
|
78
|
+
}
|
|
@@ -0,0 +1,132 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: python-testing
|
|
3
|
+
description: "Use when you write, extend, or fix a Python project's pytest suite -- add pytest.mark.parametrize cases, design conftest.py fixtures at the right scope, use pytest.mark.asyncio for coroutines, and patch external calls with mocker.patch/monkeypatch to close coverage gaps in a failing or incomplete test file. Not for auditing test conventions without changing files (see review-testing-practices)."
|
|
4
|
+
triggers:
|
|
5
|
+
- "write pytest tests"
|
|
6
|
+
- "add python test coverage"
|
|
7
|
+
- "fix failing pytest test"
|
|
8
|
+
- "pytest fixture"
|
|
9
|
+
- "parametrize test"
|
|
10
|
+
metadata:
|
|
11
|
+
origin: authored
|
|
12
|
+
category: test
|
|
13
|
+
version: "1.0.0"
|
|
14
|
+
compatible_harnesses: "claude,codex,cursor,zed,opencode"
|
|
15
|
+
license: "MIT"
|
|
16
|
+
---
|
|
17
|
+
|
|
18
|
+
# Python testing (pytest)
|
|
19
|
+
|
|
20
|
+
Write, extend, or fix a Python project's `pytest` test suite. Scoped to
|
|
21
|
+
pytest specifically — its fixture system, `parametrize`, and `monkeypatch`/
|
|
22
|
+
`unittest.mock` conventions differ enough from a generic test-generation
|
|
23
|
+
workflow (`quality/test-gen`) that a pytest-specific one earns its own
|
|
24
|
+
skill; see `governance/scout.json` for why a fork of an existing skill was
|
|
25
|
+
not enough.
|
|
26
|
+
|
|
27
|
+
## Workflow
|
|
28
|
+
|
|
29
|
+
### Step 1: Discover the project's pytest conventions
|
|
30
|
+
|
|
31
|
+
1. Read `pyproject.toml`/`setup.cfg`/`pytest.ini` for `[tool.pytest.ini_options]`
|
|
32
|
+
or `[pytest]` — test paths, markers, `addopts` (e.g. `--strict-markers`,
|
|
33
|
+
coverage flags already configured).
|
|
34
|
+
2. Find the test layout: `tests/` mirroring `src/`, or co-located
|
|
35
|
+
`test_*.py`/`*_test.py` next to the module under test. Match whichever
|
|
36
|
+
the project already uses.
|
|
37
|
+
3. Read 1-2 neighboring test files for: fixture style (local `conftest.py`
|
|
38
|
+
vs. inline), naming (`test_<behavior>`), assertion style (plain
|
|
39
|
+
`assert` vs. a matcher library), and how mocks are constructed.
|
|
40
|
+
|
|
41
|
+
### Step 2: Plan fixtures before test cases
|
|
42
|
+
|
|
43
|
+
- A fixture needed by just one test module belongs in that module — define
|
|
44
|
+
it directly in the test file. Promote it to a `conftest.py` only once a
|
|
45
|
+
second test file needs the same fixture, and then place it in the
|
|
46
|
+
narrowest `conftest.py` that covers every file needing it (the shared
|
|
47
|
+
directory, not the suite root).
|
|
48
|
+
- Prefer a fixture's natural scope (`function` is the pytest default) over
|
|
49
|
+
widening to `module`/`session` for convenience; a wider-scoped fixture
|
|
50
|
+
that mutates state leaks between tests that assumed isolation.
|
|
51
|
+
- Use `pytest.fixture(params=[...])` or a `@pytest.mark.parametrize` on the
|
|
52
|
+
test itself for input variations — parametrize the test when only the
|
|
53
|
+
inputs vary, use a fixture when setup/teardown logic itself varies.
|
|
54
|
+
|
|
55
|
+
### Step 3: Plan test cases
|
|
56
|
+
|
|
57
|
+
**Functions:** happy path, edge cases (empty/`None`/zero/negative),
|
|
58
|
+
exception cases (`pytest.raises(SpecificError)`), boundary values.
|
|
59
|
+
|
|
60
|
+
**Fixtures/context managers:** setup ran, teardown ran even when the body
|
|
61
|
+
raises, correct value yielded.
|
|
62
|
+
|
|
63
|
+
**Async code:** `pytest.mark.asyncio` (or the project's configured async
|
|
64
|
+
plugin) — do not write a sync test that silently never awaits the coroutine
|
|
65
|
+
under test.
|
|
66
|
+
|
|
67
|
+
**Mocks:** patch at the point of use (`mocker.patch("mypkg.mod.dep")`, not
|
|
68
|
+
the definition site), and only external dependencies — an internal
|
|
69
|
+
collaborator mocked away stops the test verifying real integration.
|
|
70
|
+
|
|
71
|
+
### Step 4: Write
|
|
72
|
+
|
|
73
|
+
1. Create/extend the test file at the project's own convention path.
|
|
74
|
+
2. Import fixtures via `conftest.py` discovery, not manual re-import.
|
|
75
|
+
3. One assertion concept per test; a multi-assertion test states in its
|
|
76
|
+
name what single behavior it verifies.
|
|
77
|
+
4. Use `pytest.mark.parametrize` with explicit `ids=` when parameter tuples
|
|
78
|
+
are not self-describing in pytest's default output.
|
|
79
|
+
|
|
80
|
+
### Step 5: Run and fix
|
|
81
|
+
|
|
82
|
+
```bash
|
|
83
|
+
keryx test run --changed --strict
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
`src/testing/service.ts` detects the project's own test runner from its
|
|
87
|
+
lockfile/scripts — do not hard-code `pytest` invocation flags here beyond
|
|
88
|
+
what Step 1 already discovered. On a project with no keryx testing config,
|
|
89
|
+
run the project's own configured `pytest` invocation instead.
|
|
90
|
+
|
|
91
|
+
Fix failing tests (max 3 iterations) — fix the test, not the source under
|
|
92
|
+
test.
|
|
93
|
+
|
|
94
|
+
### Step 6: Report
|
|
95
|
+
|
|
96
|
+
```
|
|
97
|
+
Generated: tests/test_helper.py
|
|
98
|
+
- 9 test cases (3 parametrized), all passing
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
## Rules
|
|
102
|
+
|
|
103
|
+
- ALWAYS match the project's existing fixture/parametrize/mock conventions
|
|
104
|
+
found in Step 1-2, not a different project's pytest style.
|
|
105
|
+
- NEVER modify source code — only test files and `conftest.py`.
|
|
106
|
+
- Mock external dependencies (network, filesystem, other services), not
|
|
107
|
+
internal modules under the same package.
|
|
108
|
+
- If no `pytest` is configured, suggest adding it; do not add it
|
|
109
|
+
unasked.
|
|
110
|
+
|
|
111
|
+
## Red Flags
|
|
112
|
+
|
|
113
|
+
| Rationalization | Why it is wrong |
|
|
114
|
+
|---|---|
|
|
115
|
+
| "This fixture is only used once but I'll put it in the root `conftest.py` anyway" | Widens its visible scope and discoverability for no reason; a fixture needed by just one test file belongs directly in that file, not in any `conftest.py`, until a second file needs it |
|
|
116
|
+
| "The assertion keeps failing; I'll patch the source under test to make it pass" | This skill writes test files only. A source change buried in a test-authoring run is an unreviewed fix that also hides the real bug |
|
|
117
|
+
| "I'll mock the internal helper so the test is simpler" | Mock external dependencies, not internal ones — a test whose internal collaborators are all mocked only checks that the mocks agree with each other |
|
|
118
|
+
| "Still failing after three iterations; I'll loosen the assertion" | A test that asserts nothing covers nothing while reporting coverage. After 3 iterations, stop and report the failing case instead |
|
|
119
|
+
|
|
120
|
+
## Verification
|
|
121
|
+
|
|
122
|
+
Do not report the work done until all of the following hold:
|
|
123
|
+
|
|
124
|
+
- The test file sits at the project's own convention path, matching the
|
|
125
|
+
fixture/parametrize/mock style read in Step 1-2.
|
|
126
|
+
- `keryx test run --changed --strict` — or, with no keryx testing config,
|
|
127
|
+
the project's own discovered `pytest` command — exits 0 with every
|
|
128
|
+
generated test passing.
|
|
129
|
+
- `git status` shows only test files (and `conftest.py`, if touched)
|
|
130
|
+
added or modified; no source file under test changed.
|
|
131
|
+
- Every exported function/class/endpoint identified in Step 1 that lacked
|
|
132
|
+
coverage now has at least one test, or the report says why it does not.
|