@mrciphersmith/keryx 0.3.5 → 0.3.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +47 -0
- package/dist/cli.js +3099 -1160
- package/dist/core.js +22 -3
- package/docs/README.md +2 -0
- package/package.json +1 -1
- package/src/gdskills/bundled/install-manifest.json +319 -76
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.md +19 -26
- package/src/gdskills/bundled/skills/orchestration/flow-orchestrator/SKILL.md +6 -6
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.md +4 -7
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.md +2 -4
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.detail.md +23 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +17 -17
- package/src/gdskills/bundled/stacks/django/agent-refs.json +3 -0
- package/src/gdskills/bundled/stacks/django/governance/eval.json +1763 -0
- package/src/gdskills/bundled/stacks/django/governance/scout.json +40 -0
- package/src/gdskills/bundled/stacks/django/pack.json +43 -0
- package/src/gdskills/bundled/stacks/django/rules/coding-style.mdc +80 -0
- package/src/gdskills/bundled/stacks/django/rules/patterns.mdc +92 -0
- package/src/gdskills/bundled/stacks/django/rules/security.mdc +92 -0
- package/src/gdskills/bundled/stacks/django/rules/testing.mdc +89 -0
- package/src/gdskills/bundled/stacks/django/skills/django-build-fix/SKILL.md +149 -0
- package/src/gdskills/bundled/stacks/django/skills/django-build-fix/evals.json +49 -0
- package/src/gdskills/bundled/stacks/django/skills/django-code-review/SKILL.md +137 -0
- package/src/gdskills/bundled/stacks/django/skills/django-code-review/evals.json +48 -0
- package/src/gdskills/bundled/stacks/django/skills/django-implementation/SKILL.md +147 -0
- package/src/gdskills/bundled/stacks/django/skills/django-implementation/evals.json +75 -0
- package/src/gdskills/bundled/stacks/django/skills/django-migrate/SKILL.md +166 -0
- package/src/gdskills/bundled/stacks/django/skills/django-migrate/evals.json +49 -0
- package/src/gdskills/bundled/stacks/django/skills/django-testing/SKILL.md +130 -0
- package/src/gdskills/bundled/stacks/django/skills/django-testing/evals.json +48 -0
- package/src/gdskills/bundled/stacks/fastapi/agent-refs.json +3 -0
- package/src/gdskills/bundled/stacks/fastapi/governance/eval.json +1777 -0
- package/src/gdskills/bundled/stacks/fastapi/governance/scout.json +34 -0
- package/src/gdskills/bundled/stacks/fastapi/pack.json +43 -0
- package/src/gdskills/bundled/stacks/fastapi/rules/coding-style.mdc +68 -0
- package/src/gdskills/bundled/stacks/fastapi/rules/patterns.mdc +108 -0
- package/src/gdskills/bundled/stacks/fastapi/rules/security.mdc +99 -0
- package/src/gdskills/bundled/stacks/fastapi/rules/testing.mdc +85 -0
- package/src/gdskills/bundled/stacks/fastapi/skills/fastapi-build-fix/SKILL.md +157 -0
- package/src/gdskills/bundled/stacks/fastapi/skills/fastapi-build-fix/evals.json +76 -0
- package/src/gdskills/bundled/stacks/fastapi/skills/fastapi-code-review/SKILL.md +150 -0
- package/src/gdskills/bundled/stacks/fastapi/skills/fastapi-code-review/evals.json +74 -0
- package/src/gdskills/bundled/stacks/fastapi/skills/fastapi-implementation/SKILL.md +158 -0
- package/src/gdskills/bundled/stacks/fastapi/skills/fastapi-implementation/evals.json +75 -0
- package/src/gdskills/bundled/stacks/fastapi/skills/fastapi-testing/SKILL.md +146 -0
- package/src/gdskills/bundled/stacks/fastapi/skills/fastapi-testing/evals.json +74 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/agent-refs.json +3 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/governance/eval.json +2194 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/governance/scout.json +39 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/pack.json +40 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/rules/coding-style.mdc +67 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/rules/patterns.mdc +65 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/rules/security.mdc +69 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/rules/testing.mdc +80 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/skills/java-kotlin-spring-build-fix/SKILL.md +144 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/skills/java-kotlin-spring-build-fix/evals.json +74 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/skills/java-kotlin-spring-code-review/SKILL.md +129 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/skills/java-kotlin-spring-code-review/evals.json +74 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/skills/java-kotlin-spring-implementation/SKILL.md +147 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/skills/java-kotlin-spring-implementation/evals.json +75 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/skills/java-kotlin-spring-migrate/SKILL.md +139 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/skills/java-kotlin-spring-migrate/evals.json +74 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/skills/java-kotlin-spring-testing/SKILL.md +128 -0
- package/src/gdskills/bundled/stacks/java-kotlin-spring/skills/java-kotlin-spring-testing/evals.json +73 -0
- package/src/gdskills/bundled/stacks/python/agent-refs.json +2 -1
- package/src/gdskills/bundled/stacks/python/pack.json +1 -1
- package/src/gdskills/bundled/stacks/rust/agent-refs.json +3 -0
- package/src/gdskills/bundled/stacks/rust/governance/eval.json +1823 -0
- package/src/gdskills/bundled/stacks/rust/governance/scout.json +32 -0
- package/src/gdskills/bundled/stacks/rust/pack.json +42 -0
- package/src/gdskills/bundled/stacks/rust/rules/coding-style.mdc +93 -0
- package/src/gdskills/bundled/stacks/rust/rules/patterns.mdc +85 -0
- package/src/gdskills/bundled/stacks/rust/rules/security.mdc +85 -0
- package/src/gdskills/bundled/stacks/rust/rules/testing.mdc +82 -0
- package/src/gdskills/bundled/stacks/rust/skills/rust-build-fix/SKILL.md +141 -0
- package/src/gdskills/bundled/stacks/rust/skills/rust-build-fix/evals.json +78 -0
- package/src/gdskills/bundled/stacks/rust/skills/rust-code-review/SKILL.md +127 -0
- package/src/gdskills/bundled/stacks/rust/skills/rust-code-review/evals.json +72 -0
- package/src/gdskills/bundled/stacks/rust/skills/rust-implementation/SKILL.md +133 -0
- package/src/gdskills/bundled/stacks/rust/skills/rust-implementation/evals.json +79 -0
- package/src/gdskills/bundled/stacks/rust/skills/rust-testing/SKILL.md +130 -0
- package/src/gdskills/bundled/stacks/rust/skills/rust-testing/evals.json +75 -0
- package/src/gdskills/bundled/agents/python-build-fixer.md +0 -52
- package/src/gdskills/bundled/agents/python-code-auditor.md +0 -49
|
@@ -0,0 +1,127 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: rust-code-review
|
|
3
|
+
description: "Use when reviewing a Rust change for ownership, panic, and safety risks -- unwrap/expect on fallible input, unnecessary clone as a borrow-checker workaround, unjustified or oversized unsafe blocks, blocking calls in async fn, unchecked arithmetic/indexing on untrusted data, and error types that lose their source. Read-only, no edits."
|
|
4
|
+
triggers:
|
|
5
|
+
- "review this Rust diff for unwrap panics"
|
|
6
|
+
- "check this Rust change for unsafe block safety"
|
|
7
|
+
- "review this Rust patch for blocking calls in async fn"
|
|
8
|
+
- "any unnecessary clones in this Rust change"
|
|
9
|
+
- "check this Rust code for integer overflow on untrusted input"
|
|
10
|
+
- "review this Rust diff for error handling that drops the source"
|
|
11
|
+
- "audit this Rust module for a panic reachable from untrusted input"
|
|
12
|
+
metadata:
|
|
13
|
+
origin: authored
|
|
14
|
+
category: review
|
|
15
|
+
version: "1.0.0"
|
|
16
|
+
compatible_harnesses: "claude,codex,cursor,zed,opencode"
|
|
17
|
+
license: "MIT"
|
|
18
|
+
---
|
|
19
|
+
|
|
20
|
+
# Rust code review
|
|
21
|
+
|
|
22
|
+
Read-only review of a Rust change for ownership, panic, and safety risks
|
|
23
|
+
specific to Rust: `.unwrap()`/`.expect()` on fallible or untrusted input,
|
|
24
|
+
a borrow-checker `.clone()` workaround, unjustified `unsafe`, blocking
|
|
25
|
+
calls in `async fn`, unchecked arithmetic/indexing, and error types that
|
|
26
|
+
lose their source. This skill never edits code — it reports findings.
|
|
27
|
+
`rules/coding-style.mdc`, `rules/patterns.mdc`, and `rules/security.mdc`
|
|
28
|
+
are the rule set findings are checked against.
|
|
29
|
+
|
|
30
|
+
## Workflow
|
|
31
|
+
|
|
32
|
+
### Step 1: Scope the review
|
|
33
|
+
|
|
34
|
+
1. Identify the changed files (`git diff` against the review base) —
|
|
35
|
+
review only `*.rs` files in the diff, not the whole crate.
|
|
36
|
+
2. Read enough of the surrounding, unchanged code to know whether a
|
|
37
|
+
flagged pattern is new in this diff or pre-existing; note pre-existing
|
|
38
|
+
issues separately from ones the diff introduces.
|
|
39
|
+
|
|
40
|
+
### Step 2: Check each changed function against the focus list
|
|
41
|
+
|
|
42
|
+
**Panics on fallible/untrusted input**
|
|
43
|
+
- `.unwrap()`/`.expect()` on a `Result`/`Option` built from network input,
|
|
44
|
+
file contents, CLI arguments, a map lookup on external data, or any
|
|
45
|
+
other fallible/attacker-influenced source — flag it; the fix is `?`, a
|
|
46
|
+
`match`, or a combinator, not a panic on bad input.
|
|
47
|
+
- Raw array/slice indexing (`arr[i]`) on an index derived from untrusted
|
|
48
|
+
input instead of `.get(i)` — flag as a potential panic/DoS.
|
|
49
|
+
- Unchecked arithmetic (`a + b`, `a * b`) on values that can be
|
|
50
|
+
attacker-influenced instead of `checked_*`/`saturating_*`/`wrapping_*`
|
|
51
|
+
chosen for the operation's actual semantics.
|
|
52
|
+
|
|
53
|
+
**Ownership and cloning**
|
|
54
|
+
- A `.clone()` added specifically to route around a borrow-checker error,
|
|
55
|
+
with no comment explaining why the duplication is actually needed —
|
|
56
|
+
flag as a possible design smell per `rules/patterns.mdc`; note when
|
|
57
|
+
restructuring ownership or borrowing differently looks feasible from
|
|
58
|
+
the diff alone.
|
|
59
|
+
|
|
60
|
+
**`unsafe` blocks**
|
|
61
|
+
- Any `unsafe` block with no `# Safety` doc comment stating the invariant
|
|
62
|
+
it relies on — flag it.
|
|
63
|
+
- An `unsafe` block wider than the specific operation that requires it —
|
|
64
|
+
flag as unnecessarily large audited surface.
|
|
65
|
+
|
|
66
|
+
**Async and concurrency**
|
|
67
|
+
- A blocking call (`std::thread::sleep`, synchronous `std::fs`/`std::net`
|
|
68
|
+
I/O, a `std::sync::Mutex` guard held across an `.await`) inside an
|
|
69
|
+
`async fn` — flag as a potential executor stall.
|
|
70
|
+
- A spawned task (`tokio::spawn`) with no visible join (`.await` on its
|
|
71
|
+
`JoinHandle`, or a channel) — flag as fire-and-forget with no stated
|
|
72
|
+
reason.
|
|
73
|
+
|
|
74
|
+
**Error handling**
|
|
75
|
+
- A custom error variant that discards the underlying cause (a bare
|
|
76
|
+
`String` built from `.to_string()` on the source error, with no
|
|
77
|
+
`#[source]`/`#[from]`) — flag as losing information a caller could
|
|
78
|
+
otherwise inspect.
|
|
79
|
+
|
|
80
|
+
### Step 3: Report
|
|
81
|
+
|
|
82
|
+
For each finding: file:line, the pattern, why it matters (panic, leak,
|
|
83
|
+
unsound `unsafe`, executor stall), and the fix direction — but do not
|
|
84
|
+
apply it.
|
|
85
|
+
|
|
86
|
+
```
|
|
87
|
+
src/order/service.rs:58 — .unwrap() on the downstream client's Result,
|
|
88
|
+
which carries a network call's outcome. Risk: a downstream failure
|
|
89
|
+
panics the process instead of propagating. Fix direction: propagate
|
|
90
|
+
with ? into this module's OrderError, or add a #[source]-wrapping
|
|
91
|
+
variant if one doesn't exist yet.
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
## Rules
|
|
95
|
+
|
|
96
|
+
- NEVER edit code — findings and fix direction only.
|
|
97
|
+
- Flag `.unwrap()`/`.expect()` on fallible/untrusted input, unnecessary
|
|
98
|
+
clones, unjustified/oversized `unsafe`, blocking calls in `async fn`,
|
|
99
|
+
unchecked arithmetic/indexing on untrusted data, and error types that
|
|
100
|
+
drop their source; do not report generic style nits already covered by
|
|
101
|
+
`rustfmt`/`clippy`'s default lint set (those are noise here).
|
|
102
|
+
- Distinguish a finding the diff introduces from a pre-existing one in
|
|
103
|
+
code the diff merely touches.
|
|
104
|
+
- When a suspected data race or unsoundness is not certain from reading
|
|
105
|
+
alone, say "run under `cargo miri test` / `loom` to confirm" rather than
|
|
106
|
+
asserting it without evidence.
|
|
107
|
+
|
|
108
|
+
## Red Flags
|
|
109
|
+
|
|
110
|
+
| Rationalization | Why it is wrong |
|
|
111
|
+
|---|---|
|
|
112
|
+
| "This `.unwrap()` is on input the caller always validates first" | That invariant lives in a different function the diff may not show; flag it and ask for the validation to be visible at this call site, or a comment stating the invariant |
|
|
113
|
+
| "It's just a config map, it's only cloned once at startup" | A clone whose need is not stated in a comment is still a design smell worth naming, even if the current call site is cheap |
|
|
114
|
+
| "I'll just fix the unwrap myself since it's a one-line change" | This skill is read-only; report the finding and its fix direction, do not edit the file |
|
|
115
|
+
| "The unsafe block is small, it doesn't need a Safety comment" | Size does not establish soundness; the comment is what lets a reviewer (or future maintainer) check the invariant without re-deriving it |
|
|
116
|
+
|
|
117
|
+
## Verification
|
|
118
|
+
|
|
119
|
+
Do not report the review done until all of the following hold:
|
|
120
|
+
|
|
121
|
+
- Every changed `*.rs` file in the diff was read, not just files named in
|
|
122
|
+
the PR description.
|
|
123
|
+
- Every finding names a concrete file:line, the specific risk category
|
|
124
|
+
from Step 2, and a fix direction.
|
|
125
|
+
- No source file was modified by this review.
|
|
126
|
+
- Findings distinguish diff-introduced issues from pre-existing ones in
|
|
127
|
+
touched files.
|
|
@@ -0,0 +1,72 @@
|
|
|
1
|
+
{
|
|
2
|
+
"triggers": {
|
|
3
|
+
"positive": [
|
|
4
|
+
"Review this Rust pull request for unwrap panics on untrusted input",
|
|
5
|
+
"Check this Rust diff for unsafe blocks missing a Safety comment",
|
|
6
|
+
"Does this Rust change block the tokio runtime inside an async fn?",
|
|
7
|
+
"Review this Rust concurrency change for unjoined spawned tasks",
|
|
8
|
+
"Check for unnecessary clones used as a borrow-checker workaround in this Rust diff",
|
|
9
|
+
"This error enum has a variant that stores just a String instead of wrapping the original error -- is that going to bite us later, can you check the diff?"
|
|
10
|
+
],
|
|
11
|
+
"negative": [
|
|
12
|
+
"Review this Rust code and also fix the bugs you find",
|
|
13
|
+
"Review this Python code for SQL injection",
|
|
14
|
+
"Review this Go diff for goroutine leaks",
|
|
15
|
+
"Implement a bounded async worker pool in this Rust crate using tokio",
|
|
16
|
+
"Run a general security review on this codebase",
|
|
17
|
+
"Review this Rust diff for naming conventions and formatting only"
|
|
18
|
+
]
|
|
19
|
+
},
|
|
20
|
+
"scenarios": [
|
|
21
|
+
{
|
|
22
|
+
"id": "read-only-unwrap-review",
|
|
23
|
+
"prompt": "Review this Rust diff: a new function calls `response.json::<Payload>().await.unwrap()` on an HTTP client's response. What do you find?",
|
|
24
|
+
"strictness": "high",
|
|
25
|
+
"expected_behavior": [
|
|
26
|
+
{
|
|
27
|
+
"grader": "judge",
|
|
28
|
+
"rubric": "A correct answer identifies the .unwrap() on the HTTP response's Result as the problem, explains why it's wrong (the network response is fallible/untrusted input and a malformed or unexpected body will panic the process instead of returning an error), and reports this as a finding with a fix direction -- it never actually edits or claims to have edited the code, even partially, since this skill is read-only.",
|
|
29
|
+
"pass_criteria": [
|
|
30
|
+
"Identifies the `.unwrap()` on `response.json::<Payload>().await` as the problem, naming the specific call from this diff.",
|
|
31
|
+
"Explains that a network response is fallible/untrusted input, so a malformed or unexpected body panics the process instead of being handled.",
|
|
32
|
+
"Names the concrete fix direction: propagate with `?` (into the function's Result/error type) or handle the Err case explicitly, presented purely as a finding with no code change applied or described as already made."
|
|
33
|
+
],
|
|
34
|
+
"fail_criteria": [
|
|
35
|
+
"Actually edits, patches, or claims to have already made even a partial or proof-of-concept code change (this skill is strictly read-only: findings and fix direction only, no edits of any size)."
|
|
36
|
+
]
|
|
37
|
+
}
|
|
38
|
+
],
|
|
39
|
+
"calibration": {
|
|
40
|
+
"known_right": "Finding: `response.json::<Payload>().await.unwrap()` panics the whole process if the HTTP response body doesn't deserialize into `Payload` -- a malformed response, an unexpected error page, or a schema change on the other end all turn into a panic here instead of a handled error, even though this is exactly the kind of externally-influenced, fallible input `.unwrap()` should never be used on. Fix direction: propagate the `Result` with `?` (converting into whatever error type this function already returns, via a `#[from]`-wrapping thiserror variant if one exists, or adding one) so a deserialization failure becomes a normal `Err` the caller can handle, rather than a panic. I'm reporting this as a finding with that fix direction rather than making the change myself, since this review is read-only.",
|
|
41
|
+
"known_wrong": "This `.unwrap()` is a minor nit -- the API contract for this endpoint is pretty stable, so I went ahead and changed it to `.expect(\"payload should deserialize\")` directly in the diff so there's at least a message if it ever does panic, and left a comment noting the endpoint's schema is unlikely to change. That should be good to merge as-is.",
|
|
42
|
+
"vague": "That unwrap on the response looks like it could be risky if the payload doesn't parse -- might be worth reconsidering before this merges.",
|
|
43
|
+
"subtle_wrong": "Finding: `.unwrap()` on the deserialized response will panic if `Payload` doesn't match the body -- worth fixing. I went ahead and made a small proof-of-concept edit swapping it for `.expect(\"payload should match Payload\")` so you can see the shape of a clearer panic message, though the caller would still need to decide how it wants to handle an actual parse failure before this is really done."
|
|
44
|
+
}
|
|
45
|
+
},
|
|
46
|
+
{
|
|
47
|
+
"id": "unsafe-missing-safety-comment",
|
|
48
|
+
"prompt": "Review this Rust diff: a new `unsafe` block dereferences a raw pointer passed in from an FFI caller, with no comment above it. What's wrong?",
|
|
49
|
+
"strictness": "high",
|
|
50
|
+
"expected_behavior": [
|
|
51
|
+
{
|
|
52
|
+
"grader": "judge",
|
|
53
|
+
"rubric": "A correct answer identifies that the unsafe block has no # Safety doc comment stating the invariants the caller must uphold (pointer validity, non-null, alignment, lifetime) for the dereference to be sound, flags this as a review finding with a fix direction (add a # Safety comment naming the specific invariant, and validate what can be checked at runtime before the FFI boundary), and does not apply the fix itself since this skill is read-only. It does not accept 'the unsafe block is small/obviously fine' as a substitute for stating the invariant.",
|
|
54
|
+
"pass_criteria": [
|
|
55
|
+
"Identifies specifically that the unsafe block lacks a # Safety doc comment naming the invariants required for the raw pointer dereference to be sound.",
|
|
56
|
+
"Names at least one concrete invariant that comment should state (non-null, valid, properly aligned, or a lifetime/ownership contract) relevant to an FFI-supplied raw pointer.",
|
|
57
|
+
"Presents this purely as a finding with a fix direction, with no code change applied, described as applied, or described as a proof-of-concept edit made during the review."
|
|
58
|
+
],
|
|
59
|
+
"fail_criteria": [
|
|
60
|
+
"Dismisses the missing # Safety comment as unnecessary because the unsafe block is small or the dereference looks obviously fine, instead of requiring the invariant to be stated explicitly."
|
|
61
|
+
]
|
|
62
|
+
}
|
|
63
|
+
],
|
|
64
|
+
"calibration": {
|
|
65
|
+
"known_right": "Finding: the `unsafe` block dereferences a raw pointer that came from an FFI caller, but carries no `# Safety` doc comment. That's a problem specifically because this pointer crosses an FFI boundary -- unlike a pointer derived from safe Rust, there's no compiler guarantee it's non-null, valid, or properly aligned, and nothing here states what the caller is actually required to uphold. Fix direction: add a `# Safety` doc comment on the function or the block itself naming the concrete invariant (e.g. \"caller must ensure `ptr` is non-null, points to a valid, properly aligned `Payload`, and remains valid for the duration of this call\"), and where feasible, add a runtime null-check before the dereference since that much can actually be verified rather than just documented. I'm flagging this as a finding with that fix direction rather than adding the comment myself, since this review is read-only.",
|
|
66
|
+
"known_wrong": "The unsafe block is only dereferencing one pointer and the surrounding code looks straightforward, so I don't think this needs a Safety comment -- it's pretty obviously fine from reading it. I'd let this merge as-is; adding a comment here would mostly be restating what the code already does rather than adding real information.",
|
|
67
|
+
"vague": "That unsafe block could use a bit more documentation about what it assumes -- might be worth tightening up before this merges.",
|
|
68
|
+
"subtle_wrong": "Finding: the unsafe block has no comment above it. I went ahead and added a short `// SAFETY: ptr is assumed valid` note directly above the block in the diff so reviewers have something to read, though it doesn't spell out non-null/alignment/lifetime specifically -- that level of detail probably isn't necessary for a pointer that's just being read once."
|
|
69
|
+
}
|
|
70
|
+
}
|
|
71
|
+
]
|
|
72
|
+
}
|
|
@@ -0,0 +1,133 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: rust-implementation
|
|
3
|
+
description: "Use when implementing or extending a feature in a Rust crate -- Cargo.toml/edition/workspace layout, ownership and borrowing, Result/Option error handling with the ? operator, thiserror/anyhow conventions, trait and generic design, and avoiding blocking calls in async fn."
|
|
4
|
+
triggers:
|
|
5
|
+
- "implement this feature in Rust"
|
|
6
|
+
- "add a Rust module to this crate under src/"
|
|
7
|
+
- "propagate this Rust error with the ? operator"
|
|
8
|
+
- "add a thiserror error enum for this Rust module"
|
|
9
|
+
- "write this as an async fn in Rust without blocking the runtime"
|
|
10
|
+
- "design this Rust trait so callers can't over-implement it"
|
|
11
|
+
- "implement this Rust struct with a builder pattern"
|
|
12
|
+
metadata:
|
|
13
|
+
origin: authored
|
|
14
|
+
category: implement
|
|
15
|
+
version: "1.0.0"
|
|
16
|
+
compatible_harnesses: "claude,codex,cursor,zed,opencode"
|
|
17
|
+
license: "MIT"
|
|
18
|
+
---
|
|
19
|
+
|
|
20
|
+
# Rust implementation (2021/2024 edition)
|
|
21
|
+
|
|
22
|
+
Implement or extend a feature in a Rust crate: workspace/module layout,
|
|
23
|
+
ownership and borrowing, error handling with `Result`/`?`, trait and
|
|
24
|
+
generic design, and safe async idiom. Scoped to Rust specifically —
|
|
25
|
+
`rules/coding-style.mdc`, `rules/patterns.mdc`, and `rules/security.mdc`
|
|
26
|
+
carry the full stack-specific rule set this skill draws its checklist
|
|
27
|
+
from; read them before writing code, not just this summary.
|
|
28
|
+
|
|
29
|
+
## Workflow
|
|
30
|
+
|
|
31
|
+
### Step 1: Discover the project's own conventions
|
|
32
|
+
|
|
33
|
+
1. Read `Cargo.toml` for the crate/workspace name, the `edition` (2021 vs
|
|
34
|
+
2024), and whether a `rust-version` (MSRV) is pinned — do not use a
|
|
35
|
+
language feature newer than what's declared.
|
|
36
|
+
2. Find the existing layout: `src/lib.rs`/`src/main.rs`, module
|
|
37
|
+
boundaries, whether this is a `[workspace]` with multiple member
|
|
38
|
+
crates. Match it; do not invent a different layout for one change.
|
|
39
|
+
3. Read 1-2 neighboring files in the module you are touching for: error
|
|
40
|
+
handling style (`thiserror` enum vs `anyhow::Result`), whether the
|
|
41
|
+
crate is async (`tokio` in `Cargo.toml`'s dependencies) and which
|
|
42
|
+
runtime, existing trait boundaries, and the crate's own newtype/builder
|
|
43
|
+
conventions.
|
|
44
|
+
|
|
45
|
+
### Step 2: Design before writing
|
|
46
|
+
|
|
47
|
+
- Decide ownership shape per function: does it need `&self`, `&mut self`,
|
|
48
|
+
or `self` by value? Does a parameter need to own its argument or can it
|
|
49
|
+
borrow (`&str` over `String`, `&[T]` over `Vec<T>`)?
|
|
50
|
+
- Decide the error shape: a library module gets a `thiserror`-derived enum
|
|
51
|
+
with `#[source]`/`#[from]` on wrapping variants; an application/binary
|
|
52
|
+
entry point uses `anyhow::Result` with `.context(...)` at each
|
|
53
|
+
meaningful `?`.
|
|
54
|
+
- For any new trait, decide who the consumer is and size the trait to
|
|
55
|
+
exactly what that consumer calls (`rules/patterns.mdc`); decide
|
|
56
|
+
`impl Trait` vs `Box<dyn Trait>` based on whether runtime polymorphism
|
|
57
|
+
is actually needed.
|
|
58
|
+
- If the crate is async, trace which calls on the new path are I/O-bound
|
|
59
|
+
(use the async equivalent) versus CPU-bound/blocking (route through
|
|
60
|
+
`tokio::task::spawn_blocking`).
|
|
61
|
+
|
|
62
|
+
### Step 3: Implement
|
|
63
|
+
|
|
64
|
+
1. Borrow instead of clone by default (`rules/coding-style.mdc`); when a
|
|
65
|
+
`.clone()` is genuinely necessary, say why in a comment if it is not
|
|
66
|
+
obvious.
|
|
67
|
+
2. Propagate errors with `?`; never reach for `.unwrap()`/`.expect()` on a
|
|
68
|
+
`Result`/`Option` built from fallible or untrusted input.
|
|
69
|
+
3. Derive `Debug`/`Clone`/`PartialEq`/etc. where the fields support it
|
|
70
|
+
instead of hand-writing boilerplate impls.
|
|
71
|
+
4. Keep any `unsafe` block minimal and carrying a `# Safety` doc comment
|
|
72
|
+
stating the exact invariant it relies on (`rules/security.mdc`).
|
|
73
|
+
5. Format with `cargo fmt` as you go, not as an afterthought.
|
|
74
|
+
|
|
75
|
+
### Step 4: Verify
|
|
76
|
+
|
|
77
|
+
```bash
|
|
78
|
+
cargo build --all-targets
|
|
79
|
+
cargo clippy --all-targets -- -D warnings
|
|
80
|
+
cargo fmt --check
|
|
81
|
+
cargo test
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
Fix findings at the root cause per `rules/security.mdc` and
|
|
85
|
+
`rules/coding-style.mdc`; a build/clippy/test failure at this step is a
|
|
86
|
+
signal to fix the implementation, not to reach for `rust-build-fix`'s
|
|
87
|
+
scope unless the failure is a build/dependency/module-resolution problem
|
|
88
|
+
unrelated to the feature logic.
|
|
89
|
+
|
|
90
|
+
### Step 5: Report
|
|
91
|
+
|
|
92
|
+
```
|
|
93
|
+
Implemented: src/order/service.rs, src/order/error.rs
|
|
94
|
+
- New OrderError enum (thiserror) wraps the downstream client's error
|
|
95
|
+
- cargo build/clippy -D warnings/fmt --check/test all pass
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
## Rules
|
|
99
|
+
|
|
100
|
+
- ALWAYS propagate a fallible `Result`/`Option` with `?` or an explicit
|
|
101
|
+
match/combinator; NEVER call `.unwrap()`/`.expect()` on one built from
|
|
102
|
+
fallible or untrusted input without a stated, verified invariant that
|
|
103
|
+
makes it truly infallible.
|
|
104
|
+
- ALWAYS give an `unsafe` block a `# Safety` doc comment naming the
|
|
105
|
+
invariant it relies on; NEVER widen an `unsafe` block or drop that
|
|
106
|
+
comment to silence a compiler/clippy complaint.
|
|
107
|
+
- NEVER call a blocking operation (`std::thread::sleep`, a synchronous
|
|
108
|
+
file/socket read, a `std::sync::Mutex` held across `.await`) inside an
|
|
109
|
+
`async fn` on a multi-threaded runtime; use the async equivalent or
|
|
110
|
+
`spawn_blocking`.
|
|
111
|
+
- Match the project's declared `edition`/`rust-version` — do not use a
|
|
112
|
+
language feature newer than what `Cargo.toml` targets.
|
|
113
|
+
|
|
114
|
+
## Red Flags
|
|
115
|
+
|
|
116
|
+
| Rationalization | Why it is wrong |
|
|
117
|
+
|---|---|
|
|
118
|
+
| "The borrow checker is fighting me here, I'll just `.clone()` to move on" | A reflexive clone hides an ownership design question (`rules/patterns.mdc`); consider restructuring or borrowing before paying the copy on every call |
|
|
119
|
+
| "This `.unwrap()` is fine, the input should always be valid" | "Should" is not "is" — untrusted or fallible input that turns out invalid panics the process; propagate with `?` instead |
|
|
120
|
+
| "It's a quick synchronous call inside this async fn, it won't block long" | Any blocking call on a multi-threaded async runtime can stall the executor's worker thread for every other task scheduled on it, not just this one |
|
|
121
|
+
| "I'll skip the `# Safety` comment, the unsafe block is obviously fine" | "Obviously fine" is exactly what an unsafe review cannot take on faith; state the invariant so it can be checked |
|
|
122
|
+
|
|
123
|
+
## Verification
|
|
124
|
+
|
|
125
|
+
Do not report the work done until all of the following hold:
|
|
126
|
+
|
|
127
|
+
- `cargo build --all-targets`, `cargo clippy --all-targets -- -D warnings`,
|
|
128
|
+
`cargo fmt --check`, and `cargo test` all exit 0.
|
|
129
|
+
- Every new/touched `Result`/`Option` from fallible or untrusted input is
|
|
130
|
+
propagated with `?` or handled explicitly, not `.unwrap()`/`.expect()`'d.
|
|
131
|
+
- Every new `unsafe` block carries a `# Safety` doc comment.
|
|
132
|
+
- No blocking call was introduced inside an `async fn` without
|
|
133
|
+
`spawn_blocking` or an async equivalent.
|
|
@@ -0,0 +1,79 @@
|
|
|
1
|
+
{
|
|
2
|
+
"triggers": {
|
|
3
|
+
"positive": [
|
|
4
|
+
"Implement a new Rust module under src/order that propagates errors with the ? operator",
|
|
5
|
+
"My PaymentClient call can fail and I'm currently just unwrapping it -- what's the idiomatic way to wrap that failure into my own error type while keeping the original cause?",
|
|
6
|
+
"Write this Rust function so it borrows instead of cloning to satisfy the borrow checker",
|
|
7
|
+
"Implement an async fn in Rust that fetches from a downstream service without blocking the tokio runtime",
|
|
8
|
+
"I'm defining a new trait for this client -- how big should it be, and who should it actually be implemented for?",
|
|
9
|
+
"My RequestConfig struct has six optional fields and constructing it with Some/None everywhere is getting ugly -- what's a cleaner way to build it up incrementally?",
|
|
10
|
+
"Add error handling to this Rust CLI using anyhow::Result and .context()"
|
|
11
|
+
],
|
|
12
|
+
"negative": [
|
|
13
|
+
"Implement this feature in Go using errgroup for the worker pool",
|
|
14
|
+
"Add this endpoint in a Python FastAPI service",
|
|
15
|
+
"Implement this React component with the new form fields",
|
|
16
|
+
"Review this Rust diff for unsafe blocks and unwrap panics",
|
|
17
|
+
"Fix this failing cargo test in the order crate",
|
|
18
|
+
"Write table-driven Go tests for this function",
|
|
19
|
+
"Implement this feature in Rust but skip error handling for now",
|
|
20
|
+
"Which smart pointer should own this object in modern C++, unique_ptr or shared_ptr?"
|
|
21
|
+
]
|
|
22
|
+
},
|
|
23
|
+
"scenarios": [
|
|
24
|
+
{
|
|
25
|
+
"id": "error-propagation-thiserror",
|
|
26
|
+
"prompt": "I'm adding a new function to a Rust library crate that calls a downstream client. How should I handle the Result and the error it can return?",
|
|
27
|
+
"strictness": "high",
|
|
28
|
+
"expected_behavior": [
|
|
29
|
+
{ "grader": "regex", "value": "\\?" },
|
|
30
|
+
{
|
|
31
|
+
"grader": "judge",
|
|
32
|
+
"rubric": "A correct answer propagates the downstream client's fallible call with the ? operator rather than unwrapping it, and defines (or extends) a thiserror-derived error enum with a variant carrying #[source] or #[from] so the downstream error's cause is preserved for callers, rather than calling .unwrap()/.expect() on the Result or collapsing the failure into an opaque string.",
|
|
33
|
+
"pass_criteria": [
|
|
34
|
+
"Shows propagating the downstream call's Result with the ? operator, not .unwrap()/.expect() or a manual match that discards the error.",
|
|
35
|
+
"Names a concrete thiserror enum variant with #[source] or #[from] wrapping the downstream error (e.g. `Downstream(#[from] DownstreamError)`), not just 'handle the error properly'.",
|
|
36
|
+
"Explains that this preserves the original cause so callers can inspect it (via source() or the #[from] conversion), not merely that 'the error is handled'."
|
|
37
|
+
],
|
|
38
|
+
"fail_criteria": [
|
|
39
|
+
"Recommends calling .unwrap() or .expect() on the downstream call's Result instead of propagating it with ? or an explicit match. Mentioning .unwrap()/.expect() only to warn against using it on fallible/untrusted input is not a failure."
|
|
40
|
+
]
|
|
41
|
+
}
|
|
42
|
+
],
|
|
43
|
+
"calibration": {
|
|
44
|
+
"known_right": "Propagate the downstream call with `?` rather than unwrapping it -- add a variant to this module's `thiserror`-derived error enum, something like `#[error(\"downstream call failed\")] Downstream(#[from] DownstreamError)`, so the `?` operator converts the downstream error into your own error type automatically via the generated `From` impl. The `#[from]` attribute also marks the field as `#[source]`, so `std::error::Error::source()` returns the original downstream error -- a caller further up can still inspect what actually went wrong instead of just seeing your wrapper's message. Don't reach for `.unwrap()`/`.expect()` here even though it would compile faster to write -- a downstream call is exactly the kind of fallible, externally-influenced operation that should never panic the process on failure.",
|
|
45
|
+
"known_wrong": "Since this is calling a downstream client I trust to basically always succeed, I'd just `.unwrap()` the Result and move on -- it keeps the function signature simple and avoids growing the error enum for a case that's rare in practice. If you're worried about it, you could `.expect(\"downstream call should succeed\")` instead so there's at least a message in the panic output, but either way there's no need to wrap it in your own error type or preserve the source chain for something this unlikely to fail.",
|
|
46
|
+
"vague": "Handle the error from the downstream call properly and make sure it's propagated correctly to the caller.",
|
|
47
|
+
"subtle_wrong": "Propagate it with `?` like this: add a `Downstream(String)` variant to your error enum and convert with `.map_err(|e| MyError::Downstream(e.to_string()))?` at the call site -- that keeps the enum simple (no need to pull in `#[source]`/`#[from]` or worry about the downstream error type implementing `std::error::Error`) while still returning a proper `Result` instead of unwrapping. Callers get a readable message either way."
|
|
48
|
+
},
|
|
49
|
+
"anti_patterns": [".unwrap()"]
|
|
50
|
+
},
|
|
51
|
+
{
|
|
52
|
+
"id": "async-blocking-call",
|
|
53
|
+
"prompt": "I need to write an async fn in a tokio-based Rust service that reads a config file and then sleeps briefly between retries. What should I watch out for?",
|
|
54
|
+
"strictness": "high",
|
|
55
|
+
"expected_behavior": [
|
|
56
|
+
{
|
|
57
|
+
"grader": "judge",
|
|
58
|
+
"rubric": "A correct answer identifies that both the file read and the sleep must use tokio's async equivalents (tokio::fs::read/read_to_string and tokio::time::sleep) rather than the synchronous std::fs or std::thread::sleep, because a blocking call inside an async fn running on tokio's multi-threaded runtime can stall the worker thread and every other task scheduled on it -- not just this one -- and, if a genuinely CPU-bound or unavoidably-blocking operation is needed, routes it through tokio::task::spawn_blocking instead of calling it directly.",
|
|
59
|
+
"pass_criteria": [
|
|
60
|
+
"Names tokio::fs (e.g. tokio::fs::read or read_to_string) instead of std::fs for the config file read inside the async fn.",
|
|
61
|
+
"Names tokio::time::sleep instead of std::thread::sleep for the retry delay.",
|
|
62
|
+
"Explains the concrete consequence of blocking on a multi-threaded tokio runtime -- that it can stall the executor's worker thread for other tasks, not merely that it 'blocks the function'."
|
|
63
|
+
],
|
|
64
|
+
"fail_criteria": [
|
|
65
|
+
"Recommends std::thread::sleep for the retry delay inside the async fn, whether alone or alongside the correct tokio::fs read. Mentioning std::thread::sleep only to warn against it is not a failure.",
|
|
66
|
+
"Recommends a synchronous std::fs read inside the async fn and dismisses it as fine because the file is small or the read is quick."
|
|
67
|
+
]
|
|
68
|
+
}
|
|
69
|
+
],
|
|
70
|
+
"calibration": {
|
|
71
|
+
"known_right": "Use `tokio::fs::read_to_string(path).await` for the config file instead of `std::fs::read_to_string` -- the std version is a synchronous, blocking syscall, and calling it directly inside an `async fn` running on tokio's multi-threaded runtime can stall the worker thread that task is scheduled on, which stalls every other task tokio happens to schedule on that same thread too, not just this one. For the delay between retries, use `tokio::time::sleep(duration).await`, not `std::thread::sleep` -- `std::thread::sleep` blocks the OS thread for its whole duration with no way for the runtime to schedule other work on it, while `tokio::time::sleep` yields the task back to the executor and only wakes it when the timer fires. If you ever do need a genuinely CPU-bound or unavoidably-blocking call on this path, wrap it in `tokio::task::spawn_blocking(|| ...)` so it runs on tokio's dedicated blocking thread pool instead of a regular async worker thread.",
|
|
72
|
+
"known_wrong": "For something this small it's simpler to just use `std::fs::read_to_string` for the config -- it's one quick read at startup-ish time, not worth pulling in the async file API for. Same for the retry delay: `std::thread::sleep(duration)` is fine here since it's just a brief pause between attempts and the function isn't handling other work at that exact moment anyway, so blocking the thread for a few hundred milliseconds isn't really a problem in practice.",
|
|
73
|
+
"vague": "Make sure you use the async-friendly versions of file reading and sleeping so you don't block the runtime.",
|
|
74
|
+
"subtle_wrong": "Use `tokio::time::sleep(duration).await` for the retry delay -- that part's important since it yields back to the executor instead of blocking the thread. For the config file, `std::fs::read_to_string` is fine to leave as-is though, since it's a small local file and the read itself finishes in well under a millisecond, so it's not going to noticeably stall anything on the runtime even though it's technically a blocking call."
|
|
75
|
+
},
|
|
76
|
+
"anti_patterns": ["std::thread::sleep"]
|
|
77
|
+
}
|
|
78
|
+
]
|
|
79
|
+
}
|
|
@@ -0,0 +1,130 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: rust-testing
|
|
3
|
+
description: "Use when a Rust crate's own test suite needs writing, extending, or fixing with cargo test -- #[cfg(test)] unit tests, tests/ integration tests, multi-case assertions built as a Vec of structs, #[tokio::test] async tests joined properly, proptest/quickcheck property-based checks on a Rust struct or enum, and criterion benchmarks."
|
|
4
|
+
triggers:
|
|
5
|
+
- "write Rust unit tests for this module with cargo test"
|
|
6
|
+
- "add unit tests in a #[cfg(test)] block"
|
|
7
|
+
- "fix this failing cargo test in the crate"
|
|
8
|
+
- "add a proptest for this Rust struct's invariants"
|
|
9
|
+
- "write a criterion benchmark for this Rust function"
|
|
10
|
+
- "test this Rust async fn with #[tokio::test]"
|
|
11
|
+
- "add integration tests under tests/ for this Rust crate"
|
|
12
|
+
metadata:
|
|
13
|
+
origin: authored
|
|
14
|
+
category: test
|
|
15
|
+
version: "1.0.0"
|
|
16
|
+
compatible_harnesses: "claude,codex,cursor,zed,opencode"
|
|
17
|
+
license: "MIT"
|
|
18
|
+
---
|
|
19
|
+
|
|
20
|
+
# Rust testing (2021/2024 edition)
|
|
21
|
+
|
|
22
|
+
Write, extend, or fix a Rust crate's test suite: unit tests, integration
|
|
23
|
+
tests, async tests, property-based tests, and benchmarks. `rules/testing.mdc`
|
|
24
|
+
carries the full rule set this skill's checklist is built from — read it,
|
|
25
|
+
not just this summary, before writing tests.
|
|
26
|
+
|
|
27
|
+
## Workflow
|
|
28
|
+
|
|
29
|
+
### Step 1: Discover the project's test conventions
|
|
30
|
+
|
|
31
|
+
1. Read `Cargo.toml`'s `edition`/`rust-version` and dev-dependencies to
|
|
32
|
+
know what's already available: `tokio` (and its `test-util`/`macros`
|
|
33
|
+
features for `#[tokio::test]`), `proptest`/`quickcheck`, `criterion`,
|
|
34
|
+
`serial_test`, `pretty_assertions`.
|
|
35
|
+
2. Find the layout: `#[cfg(test)] mod tests` blocks beside the code, or a
|
|
36
|
+
`tests/` integration directory. Match whichever the crate already uses
|
|
37
|
+
for the kind of test you're adding.
|
|
38
|
+
3. Read 1-2 neighboring test modules for: case/table shape, whether an
|
|
39
|
+
assertion crate is already in use, fixture conventions
|
|
40
|
+
(`tests/fixtures/` or a local `testdata/`), and whether async tests are
|
|
41
|
+
already wired with `#[tokio::test]`.
|
|
42
|
+
|
|
43
|
+
### Step 2: Plan test cases
|
|
44
|
+
|
|
45
|
+
**Functions:** happy path, edge cases (empty/zero-value/boundary inputs),
|
|
46
|
+
error cases (assert the specific error variant, not just `is_err()`).
|
|
47
|
+
|
|
48
|
+
**Table-style:** a `Vec`/slice of a local case struct (name, input,
|
|
49
|
+
expected) iterated in one `#[test]` function when several cases share the
|
|
50
|
+
same assertion shape, rather than many near-duplicate test functions.
|
|
51
|
+
|
|
52
|
+
**Async code:** `#[tokio::test]` for an `async fn` test body; a spawned
|
|
53
|
+
task is joined via `.await` on its `JoinHandle` (or a channel) before
|
|
54
|
+
assertions run.
|
|
55
|
+
|
|
56
|
+
**Untrusted-input-shaped code (parsers, decoders):** consider `proptest`/
|
|
57
|
+
`quickcheck` for the input space, and note where `cargo fuzz` would apply.
|
|
58
|
+
|
|
59
|
+
### Step 3: Write
|
|
60
|
+
|
|
61
|
+
1. Create/extend the test module or `tests/` file at the project's own
|
|
62
|
+
convention path.
|
|
63
|
+
2. Name each test for the behavior it verifies, not `test1`/`test_x`.
|
|
64
|
+
3. Assert the specific error variant on a `Result`-returning function
|
|
65
|
+
(`matches!(err, MyError::NotFound { .. })` or `assert_eq!` when the
|
|
66
|
+
error implements `PartialEq`), not only `result.is_err()`.
|
|
67
|
+
4. Put fixture/golden data under the project's existing convention
|
|
68
|
+
(`tests/fixtures/` or `testdata/`).
|
|
69
|
+
5. Never synchronize with `std::thread::sleep`/`tokio::time::sleep` to
|
|
70
|
+
"give a spawned task time to finish" — join its `JoinHandle` or a
|
|
71
|
+
channel instead.
|
|
72
|
+
|
|
73
|
+
### Step 4: Property tests and benchmarks (when relevant)
|
|
74
|
+
|
|
75
|
+
- `proptest!` (or `quickcheck`) macros for a parser/serializer with a
|
|
76
|
+
large structured input space, generating cases instead of hand-writing
|
|
77
|
+
dozens of near-identical examples.
|
|
78
|
+
- `criterion`-based benchmarks (`fn bench(c: &mut Criterion)` registered
|
|
79
|
+
via `criterion_group!`/`criterion_main!`) for anything needing
|
|
80
|
+
statistical rigor; the unstable built-in `#[bench]` needs nightly and is
|
|
81
|
+
not the default choice.
|
|
82
|
+
|
|
83
|
+
### Step 5: Run and fix
|
|
84
|
+
|
|
85
|
+
```bash
|
|
86
|
+
cargo test
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
Fix failing tests (max 3 iterations) — fix the test, not the source under
|
|
90
|
+
test, unless the test itself has correctly caught a real bug (say so in
|
|
91
|
+
the report rather than silently changing production code).
|
|
92
|
+
|
|
93
|
+
### Step 6: Report
|
|
94
|
+
|
|
95
|
+
```
|
|
96
|
+
Generated: src/order/service.rs (tests module), tests/order_integration.rs
|
|
97
|
+
- 7 test cases (4 table-style), all passing under cargo test
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
## Rules
|
|
101
|
+
|
|
102
|
+
- ALWAYS match the project's existing test-location/table/assertion
|
|
103
|
+
conventions found in Step 1, not a different crate's style.
|
|
104
|
+
- NEVER modify source code — only test modules, `tests/` files, and
|
|
105
|
+
fixture data.
|
|
106
|
+
- NEVER use `std::thread::sleep`/`tokio::time::sleep` to wait for a
|
|
107
|
+
spawned task or async result instead of joining it.
|
|
108
|
+
- Assert the specific error variant for a `Result`-returning function, not
|
|
109
|
+
only `is_err()`/`is_ok()`.
|
|
110
|
+
|
|
111
|
+
## Red Flags
|
|
112
|
+
|
|
113
|
+
| Rationalization | Why it is wrong |
|
|
114
|
+
|---|---|
|
|
115
|
+
| "I'll add a short `tokio::time::sleep(Duration::from_millis(50)).await` so the spawned task finishes" | Non-deterministic under load; `.await` the task's `JoinHandle` (or a channel) so the test cannot flake |
|
|
116
|
+
| "This test keeps failing; I'll loosen the assertion to `result.is_err()` without checking which error" | An `is_err()`-only check passes for the wrong failure just as readily as the right one, hiding a regression next time |
|
|
117
|
+
| "criterion is overkill for this, I'll just time it with `std::time::Instant` in a test" | Loses warm-up handling, outlier detection, and run-to-run regression comparison that `criterion` provides for free |
|
|
118
|
+
| "I'll modify the function slightly so this test passes" | This skill only touches test files; a source change belongs in `rust-implementation` or a stated, separately-reported bug fix, not a silent edit while writing tests |
|
|
119
|
+
|
|
120
|
+
## Verification
|
|
121
|
+
|
|
122
|
+
Do not report the work done until all of the following hold:
|
|
123
|
+
|
|
124
|
+
- The test module/file sits at the project's own convention path, matching
|
|
125
|
+
the table/fixture style read in Step 1.
|
|
126
|
+
- `cargo test` exits 0 with every generated test passing.
|
|
127
|
+
- `git status` shows only test files (and fixture data, if touched) added
|
|
128
|
+
or modified; no source file under test changed.
|
|
129
|
+
- Every spawned task in a new test is joined (`.await` on its `JoinHandle`,
|
|
130
|
+
or a channel) before its assertions run.
|
|
@@ -0,0 +1,75 @@
|
|
|
1
|
+
{
|
|
2
|
+
"triggers": {
|
|
3
|
+
"positive": [
|
|
4
|
+
"I just wrote parse_header and it has zero coverage -- can you set it up with tests the idiomatic way, inline in the same file?",
|
|
5
|
+
"Add table-style test cases for this Rust validator with a Vec of cases",
|
|
6
|
+
"`cargo test -p order` is red on main after my last change -- test_reject_duplicate_line_item is failing, can you dig into why and fix it?",
|
|
7
|
+
"Add a proptest for this Rust parser function",
|
|
8
|
+
"I want hard numbers before I optimize parse_batch -- what's the right way to measure it so the compiler doesn't just optimize the whole thing away?",
|
|
9
|
+
"Test this async fn with #[tokio::test] and join the spawned task properly",
|
|
10
|
+
"Right now only the internals of this crate have coverage -- how do I set up tests that exercise it the same way an external consumer would, calling only what's exported?"
|
|
11
|
+
],
|
|
12
|
+
"negative": [
|
|
13
|
+
"Write pytest tests for this Python function",
|
|
14
|
+
"Add Jest tests for this React component",
|
|
15
|
+
"Review this Rust diff for unwrap panics and unsafe blocks",
|
|
16
|
+
"Fix the cargo build error in this crate",
|
|
17
|
+
"Implement a new Rust service that calls this downstream API",
|
|
18
|
+
"Write Go table-driven tests for this function",
|
|
19
|
+
"Add unit tests in a Go _test.go file for this package"
|
|
20
|
+
]
|
|
21
|
+
},
|
|
22
|
+
"scenarios": [
|
|
23
|
+
{
|
|
24
|
+
"id": "error-variant-assertion",
|
|
25
|
+
"prompt": "I need to add several test cases for this Rust function that validates an order and returns a Result. How should I structure them and check the error cases?",
|
|
26
|
+
"strictness": "high",
|
|
27
|
+
"expected_behavior": [
|
|
28
|
+
{
|
|
29
|
+
"grader": "judge",
|
|
30
|
+
"rubric": "A correct answer structures the cases as a small table -- a Vec or slice of a local struct with fields for the input and the expected outcome -- iterated in one test rather than many near-duplicate test functions, and for the error cases asserts the specific error variant (e.g. with matches! or assert_eq! against a PartialEq error type) rather than only checking result.is_err().",
|
|
31
|
+
"pass_criteria": [
|
|
32
|
+
"Uses a table of cases -- a Vec or slice of a struct/tuple per case with input and expected fields -- not one function per case or a long sequence of unrelated asserts.",
|
|
33
|
+
"Shows the concrete error assertion checking the specific variant (e.g. `assert!(matches!(err, OrderError::MissingId))` or `assert_eq!(err, OrderError::MissingId)`), not only `assert!(result.is_err())`.",
|
|
34
|
+
"Explains why checking the specific variant matters (it verifies which validation failed, not merely that something failed)."
|
|
35
|
+
],
|
|
36
|
+
"fail_criteria": [
|
|
37
|
+
"Checks only `result.is_err()`/`result.is_ok()` for the error cases without verifying which specific error variant occurred."
|
|
38
|
+
]
|
|
39
|
+
}
|
|
40
|
+
],
|
|
41
|
+
"calibration": {
|
|
42
|
+
"known_right": "Structure this as a small table: define a local struct with fields like `name: &str`, `order: Order`, and `want: Result<Order, OrderError>`, then build a `Vec` of cases covering the happy path, zero-value/empty-field inputs, and each distinct validation failure. Loop over the cases in one `#[test]` function and compare the actual result against `want`. For the error cases specifically, don't stop at `assert!(result.is_err())` -- that passes just as readily for the wrong error as the right one. Instead assert the specific variant: if `OrderError` derives `PartialEq`, `assert_eq!(result, Err(OrderError::MissingId))` works directly; otherwise use `assert!(matches!(result, Err(OrderError::MissingId)))`. That way the test actually verifies which validation rule caught the bad input, not just that validation returned some `Err`.",
|
|
43
|
+
"known_wrong": "I'd write a separate `#[test]` function per case -- `fn rejects_missing_id()`, `fn rejects_negative_total()`, `fn accepts_valid_order()`, and so on -- each building its own `Order` and calling `validate`. For the error ones I'd just check `assert!(result.is_err())` -- that's simpler than matching on the specific variant and is good enough to confirm validation caught the bad input; you don't really need to know which exact error came back for a test like this.",
|
|
44
|
+
"vague": "Structure the test cases in a table and make sure the error cases are checked properly along the way.",
|
|
45
|
+
"subtle_wrong": "Define a `Vec` of `(name, order, want_err)` tuples covering the cases you need, then loop with `for (name, order, want_err) in cases { let result = validate(order); assert_eq!(result.is_err(), want_err, \"{name}\"); }`. That gives you a table without needing a case-per-function setup, and comparing `is_err()` against a bool is enough to know whether validation caught the bad input for each named case."
|
|
46
|
+
},
|
|
47
|
+
"anti_patterns": ["is_err()"]
|
|
48
|
+
},
|
|
49
|
+
{
|
|
50
|
+
"id": "join-spawned-task",
|
|
51
|
+
"prompt": "My #[tokio::test] spawns a task with tokio::spawn and I want to wait for it to finish before asserting on the result. What's the right way?",
|
|
52
|
+
"strictness": "high",
|
|
53
|
+
"expected_behavior": [
|
|
54
|
+
{
|
|
55
|
+
"grader": "judge",
|
|
56
|
+
"rubric": "A correct answer waits for the spawned task using its actual JoinHandle (.await on the handle returned by tokio::spawn) or a channel, before running assertions, never a fixed tokio::time::sleep/std::thread::sleep delay as a stand-in for or supplement to a real join. A select against a deadline used purely as a non-blocking safety net alongside that join -- firing only if the task hangs, never taken instead of or before the real join -- is a legitimate timeout guard, not the fixed-delay anti-pattern this scenario tests for.",
|
|
57
|
+
"pass_criteria": [
|
|
58
|
+
"Uses .await on the JoinHandle returned by tokio::spawn (or a channel send/receive) as the actual join mechanism before asserting, shown concretely.",
|
|
59
|
+
"Explains why this is reliable/deterministic compared to a fixed sleep delay (it doesn't depend on how fast the task happens to run)."
|
|
60
|
+
],
|
|
61
|
+
"fail_criteria": [
|
|
62
|
+
"Recommends a fixed, unconditional wait before the join can be observed to have happened -- a tokio::time::sleep or std::thread::sleep taken before or instead of the join, or a backoff/polling loop that sleeps between checks -- anywhere in the wait step, whether as the sole mechanism or as extra 'insurance' alongside a real join. This does NOT include a select!'s timeout branch used only as a deadline safety net around the real join (it fires solely on a hang, and the join itself is what the assertion still actually waits on) -- do not fail an answer for that pattern alone."
|
|
63
|
+
]
|
|
64
|
+
}
|
|
65
|
+
],
|
|
66
|
+
"calibration": {
|
|
67
|
+
"known_right": "Use the `JoinHandle` that `tokio::spawn` returns: `let handle = tokio::spawn(async move { ... });` and then `let result = handle.await.unwrap();` in the test, right before your assertions -- `.await`ing the handle blocks the test task until the spawned task has actually completed, regardless of how fast or slow it happens to run on a given machine or under load. That gives you the task's return value directly, so you can assert on `result` rather than on some side effect you have to guess has landed. If the spawned task communicates through a channel instead of returning a value, receiving from that channel (`rx.recv().await`) works the same way -- it blocks until the sender side actually signals. Avoid `tokio::time::sleep`/`std::thread::sleep` for this -- a fixed delay is a guess about how long the task will take, and it either wastes time when the task finishes early or flakes under CI contention when it doesn't finish in time.",
|
|
68
|
+
"known_wrong": "The simplest way is to just add `tokio::time::sleep(Duration::from_millis(100)).await` right after `tokio::spawn(...)`, then run your assertions -- that gives the task enough time to finish in almost every case without needing to hang onto the `JoinHandle` or deal with `.unwrap()`ing it. If it's occasionally still flaky in CI, bump the delay up to something like 500ms; that's usually enough headroom and is a lot less code than threading a handle or a channel through just for a test.",
|
|
69
|
+
"vague": "Wait for the spawned task properly instead of just sleeping for a bit before you check the result.",
|
|
70
|
+
"subtle_wrong": "Hold onto the handle and `.await` it: `let handle = tokio::spawn(async move { ... }); let result = handle.await.unwrap();` before asserting -- that's the reliable way to know it's finished. If you're still seeing occasional flakes in CI, throw in a short `tokio::time::sleep(Duration::from_millis(50)).await` right after the `.await` too, just as extra insurance before the assertions run."
|
|
71
|
+
},
|
|
72
|
+
"anti_patterns": ["tokio::time::sleep"]
|
|
73
|
+
}
|
|
74
|
+
]
|
|
75
|
+
}
|