pakhale 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +72 -0
- package/assets/instructions/AGENTS.md +29 -0
- package/assets/statusline/claude-code.sh +100 -0
- package/dist/cli.js +2037 -0
- package/package.json +50 -0
- package/skills/architecture/SKILL.md +14 -0
- package/skills/consistency-check/DIMENSIONS.md +91 -0
- package/skills/consistency-check/SKILL.md +74 -0
- package/skills/consistency-check/scripts/consistency-workflow.js +158 -0
- package/skills/deslop/SKILL.md +153 -0
- package/skills/deslop/references/python.md +65 -0
- package/skills/deslop/references/react.md +71 -0
- package/skills/deslop/references/typescript.md +62 -0
- package/skills/sitedrop/SKILL.md +40 -0
- package/skills/stackup/SKILL.md +96 -0
- package/skills/stackup/references/hooks-ci.md +132 -0
- package/skills/stackup/references/python.md +81 -0
- package/skills/stackup/references/react.md +79 -0
- package/skills/stackup/references/typescript.md +114 -0
package/package.json
ADDED
|
@@ -0,0 +1,50 @@
|
|
|
1
|
+
{
|
|
2
|
+
"name": "pakhale",
|
|
3
|
+
"version": "0.1.0",
|
|
4
|
+
"description": "Pratik's personal CLI — coding-agent workflow setup and more",
|
|
5
|
+
"type": "module",
|
|
6
|
+
"license": "MIT",
|
|
7
|
+
"author": "Pratik Pakhale",
|
|
8
|
+
"homepage": "https://github.com/pratikpakhale/pakhale#readme",
|
|
9
|
+
"repository": {
|
|
10
|
+
"type": "git",
|
|
11
|
+
"url": "git+https://github.com/pratikpakhale/pakhale.git"
|
|
12
|
+
},
|
|
13
|
+
"bugs": {
|
|
14
|
+
"url": "https://github.com/pratikpakhale/pakhale/issues"
|
|
15
|
+
},
|
|
16
|
+
"keywords": [
|
|
17
|
+
"cli",
|
|
18
|
+
"claude-code",
|
|
19
|
+
"opencode",
|
|
20
|
+
"coding-agent",
|
|
21
|
+
"agent-skills",
|
|
22
|
+
"mcp",
|
|
23
|
+
"dotfiles",
|
|
24
|
+
"setup"
|
|
25
|
+
],
|
|
26
|
+
"bin": {
|
|
27
|
+
"pakhale": "./dist/cli.js"
|
|
28
|
+
},
|
|
29
|
+
"files": [
|
|
30
|
+
"dist",
|
|
31
|
+
"skills",
|
|
32
|
+
"assets"
|
|
33
|
+
],
|
|
34
|
+
"engines": {
|
|
35
|
+
"node": ">=20"
|
|
36
|
+
},
|
|
37
|
+
"scripts": {
|
|
38
|
+
"build": "tsdown",
|
|
39
|
+
"test": "bun test --timeout 20000",
|
|
40
|
+
"typecheck": "tsc --noEmit",
|
|
41
|
+
"prepublishOnly": "bun run typecheck && bun run test && bun run build"
|
|
42
|
+
},
|
|
43
|
+
"devDependencies": {
|
|
44
|
+
"@clack/prompts": "^1.7.0",
|
|
45
|
+
"@types/bun": "latest",
|
|
46
|
+
"picocolors": "^1.1.1",
|
|
47
|
+
"tsdown": "^0.15.0",
|
|
48
|
+
"typescript": "^5"
|
|
49
|
+
}
|
|
50
|
+
}
|
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: architecture
|
|
3
|
+
description: Visualize the codebase architecture provided a target.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Architecture document
|
|
7
|
+
|
|
8
|
+
Read the `improve-codebase-architecture` skill and its `HTML-REPORT.md`, using that as reference we just want to visualize the existing architecture provided a target.
|
|
9
|
+
|
|
10
|
+
Depending on target, read lot of code around it and then proceed. If the target is a code diff, then present the visualization as before / after in architecture using same HTML REPORT patterns.
|
|
11
|
+
|
|
12
|
+
Few important things to note :
|
|
13
|
+
|
|
14
|
+
- Try to render callstack diffs as much as possible
|
|
@@ -0,0 +1,91 @@
|
|
|
1
|
+
# Consistency dimensions
|
|
2
|
+
|
|
3
|
+
One sub-agent per dimension. Each agent: (1) learn the established convention from the
|
|
4
|
+
**existing/unchanged** code and **measure prevalence with counts**, (2) inspect **only** the
|
|
5
|
+
changed files, (3) return findings with `file:line`, the deviation, the established form + its
|
|
6
|
+
count, severity, and confidence. Skip accepted dual-conventions. Pick the dimensions that fit the
|
|
7
|
+
stack — drop the irrelevant ones, add domain-specific ones.
|
|
8
|
+
|
|
9
|
+
For each dimension below: **Check** = what to look for · **Measure** = how to prove prevalence ·
|
|
10
|
+
**Common false positives** = things that look wrong but are accepted.
|
|
11
|
+
|
|
12
|
+
---
|
|
13
|
+
|
|
14
|
+
## 1. Code idioms
|
|
15
|
+
- **Check:** error handling shape (throw vs structured result), early-return vs nesting, async
|
|
16
|
+
patterns, `for (;;)` vs `while (true)`, optional chaining/nullish style, import grouping &
|
|
17
|
+
ordering, default vs named exports, `type` vs `interface`, **reuse of existing helpers instead
|
|
18
|
+
of re-implementing them** (a private `formatBytes`/`clsx`/date helper that already exists centrally).
|
|
19
|
+
- **Measure:** `grep -rn` the duplicated helper's name and the central one; count call sites.
|
|
20
|
+
- **False positives:** two idioms both common across the repo; a local helper that genuinely
|
|
21
|
+
differs in behavior from the central one.
|
|
22
|
+
|
|
23
|
+
## 2. UI / visual
|
|
24
|
+
- **Check:** reuse of UI primitives (Button, Dialog, Spinner, SearchBar, Tooltip) vs hand-rolled
|
|
25
|
+
markup; spacing scale; loader component vs raw `animate-spin` icon; empty-state pattern; scrollbar
|
|
26
|
+
utility vs inline scrollbar-hiding CSS; responsive breakpoints.
|
|
27
|
+
- **Measure:** count primitive usages vs raw equivalents (`grep -rl`). Count the utility class vs the
|
|
28
|
+
inline form it replaces.
|
|
29
|
+
- **False positives:** Tailwind `size-N` vs `h-N w-N` (often *both* dominant — measure!). Two spinner
|
|
30
|
+
components both in wide use — only flag mixing *within one new feature*.
|
|
31
|
+
|
|
32
|
+
## 3. Design tokens
|
|
33
|
+
- **Check:** semantic color tokens (`text-primary`, `bg-muted`, `border-border`) vs hardcoded
|
|
34
|
+
hex/rgb or raw palette (`text-gray-500`); radius/shadow tokens; dark-mode parity; z-index scale.
|
|
35
|
+
- **Measure:** `grep -rnE '#[0-9a-fA-F]{3,6}|rgb\('` in changed files; compare with token usage counts.
|
|
36
|
+
- **False positives:** one-off brand colors that are hardcoded everywhere by design; canvas/chart
|
|
37
|
+
code that legitimately needs literal colors.
|
|
38
|
+
|
|
39
|
+
## 4. i18n / copy
|
|
40
|
+
- **Check:** every user-facing string via the i18n framework (no hardcoded literals in JSX/aria);
|
|
41
|
+
**all locale files updated** for new keys; key naming & namespacing matches siblings; placeholders
|
|
42
|
+
and `aria-label`s localized; reuse of an existing translated key instead of a near-duplicate.
|
|
43
|
+
- **Measure:** diff the key set across all locale files (`python3` to load each, compare key paths);
|
|
44
|
+
grep for hardcoded `aria-label="..."` and string literals in changed `.tsx`.
|
|
45
|
+
- **False positives:** example URLs/IDs that are identical across locales by convention; dev-only or
|
|
46
|
+
console strings; values intentionally untranslated.
|
|
47
|
+
|
|
48
|
+
## 5. State & data-flow
|
|
49
|
+
- **Check:** store conventions (slice shape, selector use, action naming); query-key factory &
|
|
50
|
+
query-hook conventions; **no `useState`+`useEffect` to mirror server/store state** (read the store
|
|
51
|
+
directly, update on event); fetch/error-handling parity with sibling services.
|
|
52
|
+
- **Measure:** compare the new slice/query against an existing one side-by-side; grep for
|
|
53
|
+
`useEffect(` that only calls a setter mirroring a prop/store value.
|
|
54
|
+
- **False positives:** genuinely local UI state; an effect with real external synchronization.
|
|
55
|
+
|
|
56
|
+
## 6. Architecture / layering
|
|
57
|
+
- **Check:** module placed in the right layer (pure parser lib must not import app/store types;
|
|
58
|
+
API/wire types live with API types, not in a util lib); dependency direction (no UI imported by
|
|
59
|
+
libs); server vs client boundary (`'use client'`, no server-only imports in client code); types
|
|
60
|
+
living next to their consumers.
|
|
61
|
+
- **Measure:** inspect import statements of new/moved modules; trace what depends on what.
|
|
62
|
+
- **False positives:** an intentional shared type in a `shared` package; a re-export barrel.
|
|
63
|
+
|
|
64
|
+
## 7. API / types / wire
|
|
65
|
+
- **Check:** request/response shapes match the server contract; field casing at the boundary
|
|
66
|
+
(snake_case wire vs camelCase domain) handled consistently with sibling endpoints; structured
|
|
67
|
+
error results vs thrown exceptions match the service's existing pattern; no breaking change hidden
|
|
68
|
+
behind a back-compat shim.
|
|
69
|
+
- **Measure:** compare the new service method against neighbors in the same file; check the server
|
|
70
|
+
schema if reachable.
|
|
71
|
+
- **False positives:** snake/camel mix that *must* match the wire; deliberate structured-result design.
|
|
72
|
+
|
|
73
|
+
## 8. Naming & file structure
|
|
74
|
+
- **Check:** file naming (kebab vs camel), component/hook/util naming conventions, directory
|
|
75
|
+
placement, export naming, test file co-location.
|
|
76
|
+
- **Measure:** list sibling files in the same directory; compare casing/prefix conventions.
|
|
77
|
+
- **False positives:** a domain term that legitimately breaks the casing pattern.
|
|
78
|
+
|
|
79
|
+
## 9. Accessibility & semantics
|
|
80
|
+
- **Check:** `aria-label`/role parity with sibling interactive elements; keyboard handlers present
|
|
81
|
+
where siblings have them; focus management in modals/menus; alt text; semantic elements vs `div`s.
|
|
82
|
+
- **Measure:** compare the new interactive component against the closest existing one.
|
|
83
|
+
- **False positives:** decorative elements that correctly have no a11y semantics.
|
|
84
|
+
|
|
85
|
+
---
|
|
86
|
+
|
|
87
|
+
## Reporting contract for each agent
|
|
88
|
+
|
|
89
|
+
Return JSON: `{ findings: [{ title, file, line, deviation, established, prevalence, severity, confidence }] }`
|
|
90
|
+
where `severity ∈ {correctness, behavior, i18n, a11y, naming, cosmetic}` and `prevalence` states the
|
|
91
|
+
counts you measured (e.g. `"established 349×, deviation 2× (the diff)"`). No prevalence number → not a finding.
|
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: consistency-check
|
|
3
|
+
description: Audit a set of changes (a branch, PR, or working tree) for consistency with the codebase's established conventions across many dimensions — code idioms, UI/visual, design tokens, i18n/copy, state & data-flow, architecture/layering, API/types, naming, and accessibility — by fanning out parallel sub-agents, one per dimension. Each agent measures how prevalent a convention actually is before flagging, so accepted dual-conventions are not reported as defects. Use when the user asks to check consistency, audit a diff/PR against existing patterns, "find every inconsistency", or sanity-check a feature branch before merge.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Consistency check
|
|
7
|
+
|
|
8
|
+
Fan out sub-agents to compare a diff against the codebase's *own* established
|
|
9
|
+
conventions, then verify and classify each finding before reporting. The point is
|
|
10
|
+
a high-signal report, not a pile of cosmetic nits.
|
|
11
|
+
|
|
12
|
+
## The four rules (read before anything else)
|
|
13
|
+
|
|
14
|
+
1. **Scope to the diff.** Only code the change *introduced or touched* is in scope.
|
|
15
|
+
Never flag pre-existing untouched code. Identify unrelated/concurrent working-tree
|
|
16
|
+
changes (other people's in-flight work) and exclude them explicitly.
|
|
17
|
+
2. **Measure before flagging.** A convention is only "the" convention if it dominates.
|
|
18
|
+
Count *both* forms (`grep -c`) in the existing codebase. If both are widely used
|
|
19
|
+
(e.g. Tailwind `size-4` 219× vs `h-4 w-4` 349×), it is an **accepted dual-convention** —
|
|
20
|
+
NOT a finding. Most false positives die here.
|
|
21
|
+
3. **Verify against source.** Read the actual lines before reporting. A finding needs
|
|
22
|
+
`file:line`, the deviation, the established alternative, and its prevalence count.
|
|
23
|
+
4. **Classify and justify.** Every candidate ends as Genuine / False-positive / Out-of-scope
|
|
24
|
+
with a one-line reason. Report skips too — silent omission reads as "all clear".
|
|
25
|
+
|
|
26
|
+
## Workflow
|
|
27
|
+
|
|
28
|
+
1. **Scope the diff** (inline, deterministic):
|
|
29
|
+
- `base=$(git merge-base <base-branch> HEAD)`
|
|
30
|
+
- `git diff --stat $base...HEAD` and `git status --short` → changed + untracked files.
|
|
31
|
+
- Separate *this change* from unrelated concurrent edits. Carry the in-scope file list forward.
|
|
32
|
+
2. **Fan out one agent per dimension.** Use [DIMENSIONS.md](DIMENSIONS.md) as the catalog.
|
|
33
|
+
- If the user has opted into orchestration (this skill counts as opt-in for a review
|
|
34
|
+
workflow), run [scripts/consistency-workflow.js](scripts/consistency-workflow.js) via the
|
|
35
|
+
Workflow tool, passing `args: { base, changedFiles, repo }` — add `only: [<keys>]` to run a
|
|
36
|
+
subset of dimensions. It pipelines review → adversarial verify → synthesize.
|
|
37
|
+
- Otherwise fan out with the Agent tool (`Explore` for find, default for verify), one per
|
|
38
|
+
dimension, in a single message so they run in parallel.
|
|
39
|
+
- Each find-agent: learn the convention + prevalence from the **existing** code, then inspect
|
|
40
|
+
**only** the changed files, returning structured findings (file:line, deviation, established
|
|
41
|
+
form + count, severity, confidence).
|
|
42
|
+
3. **Adversarially verify each candidate.** A second agent reads the source and confirms it is
|
|
43
|
+
(a) introduced by the diff, (b) a genuine deviation and not an accepted dual-convention
|
|
44
|
+
(re-measure), (c) in scope. Default to *refuted* when uncertain.
|
|
45
|
+
4. **Synthesize.** Dedup by file+line across dimensions. Rank by severity:
|
|
46
|
+
correctness > behavior parity > i18n/a11y gap > naming > cosmetic.
|
|
47
|
+
5. **Report** (default). Apply fixes only if the user asks. When fixing, never churn a cosmetic
|
|
48
|
+
dual-convention, and prefer fixing at the source over downstream remapping.
|
|
49
|
+
|
|
50
|
+
## Severity ladder
|
|
51
|
+
|
|
52
|
+
- **correctness** — wrong behavior/output, broken types, wrong copy a user sees.
|
|
53
|
+
- **behavior parity** — same operation done differently than its siblings (e.g. one code path
|
|
54
|
+
derives a value with a fallback that another omits).
|
|
55
|
+
- **i18n / a11y gap** — hardcoded user-facing/screen-reader string in an otherwise localized app;
|
|
56
|
+
missing locale key; untranslated `aria-label`.
|
|
57
|
+
- **naming / structure** — module placed in the wrong layer, type living far from its use,
|
|
58
|
+
helper reinvented instead of imported.
|
|
59
|
+
- **cosmetic** — interchangeable with an equally-common existing form. Usually *not* worth a fix;
|
|
60
|
+
only flag when it breaks consistency *within the same new feature*.
|
|
61
|
+
|
|
62
|
+
## Output shape
|
|
63
|
+
|
|
64
|
+
```
|
|
65
|
+
## Consistency report — <branch> vs <base> (N files in scope)
|
|
66
|
+
### Genuine (ranked)
|
|
67
|
+
- [severity] file:line — <deviation>; codebase uses <established> (Mx vs Nx). Fix: <one line>
|
|
68
|
+
### False positives (measured, not defects)
|
|
69
|
+
- file:line — <form> is an accepted dual-convention (Mx vs Nx). No change.
|
|
70
|
+
### Out of scope
|
|
71
|
+
- file — unrelated concurrent change / pre-existing. Untouched.
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
See [DIMENSIONS.md](DIMENSIONS.md) for the per-dimension checklists and grep recipes.
|
|
@@ -0,0 +1,158 @@
|
|
|
1
|
+
export const meta = {
|
|
2
|
+
name: 'consistency-check',
|
|
3
|
+
description:
|
|
4
|
+
'Fan out one agent per consistency dimension to audit a diff against the codebase\'s own conventions, adversarially verify each candidate, then synthesize a ranked report.',
|
|
5
|
+
phases: [
|
|
6
|
+
{ title: 'Review', detail: 'one agent per consistency dimension' },
|
|
7
|
+
{ title: 'Verify', detail: 'adversarially confirm each candidate finding' },
|
|
8
|
+
{ title: 'Synthesize', detail: 'dedup, classify, rank by severity' }
|
|
9
|
+
]
|
|
10
|
+
}
|
|
11
|
+
|
|
12
|
+
// args: { base: string, changedFiles: string[], repo?: string }
|
|
13
|
+
// Gather these inline (git merge-base + git diff --stat) BEFORE invoking, so the
|
|
14
|
+
// script only orchestrates. Drop dimensions that don't fit the stack via args.only.
|
|
15
|
+
const { base = 'main', changedFiles = [], repo = '.', only = null } = args || {}
|
|
16
|
+
const fileList = changedFiles.length
|
|
17
|
+
? changedFiles.map((f) => `- ${f}`).join('\n')
|
|
18
|
+
: '(none provided — derive from `git diff --name-only ' + base + '...HEAD`)'
|
|
19
|
+
|
|
20
|
+
const ALL_DIMENSIONS = [
|
|
21
|
+
{
|
|
22
|
+
key: 'code-idioms',
|
|
23
|
+
prompt:
|
|
24
|
+
'Error-handling shape, early-return style, async patterns, import grouping, default vs named exports, type vs interface, and especially HELPERS REINVENTED instead of imported from a central module (e.g. a private formatBytes when one already exists).'
|
|
25
|
+
},
|
|
26
|
+
{
|
|
27
|
+
key: 'ui-visual',
|
|
28
|
+
prompt:
|
|
29
|
+
'Reuse of UI primitives (Button/Dialog/Spinner/SearchBar) vs hand-rolled markup; loader component vs raw animate-spin icon; scrollbar utility vs inline scrollbar-hiding CSS; spacing scale. NOTE: size-N vs h-N w-N are often BOTH dominant — count both, do not flag a dual-convention.'
|
|
30
|
+
},
|
|
31
|
+
{
|
|
32
|
+
key: 'design-tokens',
|
|
33
|
+
prompt:
|
|
34
|
+
'Semantic color tokens (text-primary, bg-muted) vs hardcoded hex/rgb or raw palette; radius/shadow tokens; dark-mode parity. Exclude canvas/chart code that needs literal colors.'
|
|
35
|
+
},
|
|
36
|
+
{
|
|
37
|
+
key: 'i18n-copy',
|
|
38
|
+
prompt:
|
|
39
|
+
'Every user-facing string localized (no hardcoded JSX/aria literals); ALL locale files updated for new keys; key naming/namespacing matches siblings; placeholders and aria-labels localized; an existing translated key reused instead of a near-duplicate.'
|
|
40
|
+
},
|
|
41
|
+
{
|
|
42
|
+
key: 'state-dataflow',
|
|
43
|
+
prompt:
|
|
44
|
+
'Store slice/selector/action conventions; query-key factory & hook conventions; NO useState+useEffect mirroring server/store state (read store directly, update on event); fetch/error parity with sibling services.'
|
|
45
|
+
},
|
|
46
|
+
{
|
|
47
|
+
key: 'architecture',
|
|
48
|
+
prompt:
|
|
49
|
+
'Module in the right layer (pure lib must not import app/store types; wire/API types live with API types, not in a util lib); dependency direction; server vs client boundary; types living next to consumers.'
|
|
50
|
+
},
|
|
51
|
+
{
|
|
52
|
+
key: 'api-types',
|
|
53
|
+
prompt:
|
|
54
|
+
'Request/response shapes match the server contract; boundary casing (snake wire vs camel domain) handled like siblings; structured-result vs thrown-error matches the service pattern; no hidden back-compat shim.'
|
|
55
|
+
},
|
|
56
|
+
{
|
|
57
|
+
key: 'naming-structure',
|
|
58
|
+
prompt:
|
|
59
|
+
'File naming (kebab vs camel), component/hook/util naming, directory placement, export naming, test co-location — compared against sibling files in the same directory.'
|
|
60
|
+
},
|
|
61
|
+
{
|
|
62
|
+
key: 'a11y-semantics',
|
|
63
|
+
prompt:
|
|
64
|
+
'aria-label/role parity with sibling interactive elements; keyboard handlers where siblings have them; modal focus management; semantic elements vs divs.'
|
|
65
|
+
}
|
|
66
|
+
]
|
|
67
|
+
|
|
68
|
+
const DIMENSIONS = only
|
|
69
|
+
? ALL_DIMENSIONS.filter((d) => only.includes(d.key))
|
|
70
|
+
: ALL_DIMENSIONS
|
|
71
|
+
|
|
72
|
+
const FINDINGS_SCHEMA = {
|
|
73
|
+
type: 'object',
|
|
74
|
+
additionalProperties: false,
|
|
75
|
+
properties: {
|
|
76
|
+
findings: {
|
|
77
|
+
type: 'array',
|
|
78
|
+
items: {
|
|
79
|
+
type: 'object',
|
|
80
|
+
additionalProperties: false,
|
|
81
|
+
properties: {
|
|
82
|
+
title: { type: 'string' },
|
|
83
|
+
file: { type: 'string' },
|
|
84
|
+
line: { type: 'number' },
|
|
85
|
+
deviation: { type: 'string' },
|
|
86
|
+
established: { type: 'string' },
|
|
87
|
+
prevalence: {
|
|
88
|
+
type: 'string',
|
|
89
|
+
description: 'measured counts, e.g. "established 349x, deviation 2x"'
|
|
90
|
+
},
|
|
91
|
+
severity: {
|
|
92
|
+
type: 'string',
|
|
93
|
+
enum: ['correctness', 'behavior', 'i18n', 'a11y', 'naming', 'cosmetic']
|
|
94
|
+
},
|
|
95
|
+
confidence: { type: 'number' }
|
|
96
|
+
},
|
|
97
|
+
required: ['title', 'file', 'deviation', 'established', 'prevalence', 'severity']
|
|
98
|
+
}
|
|
99
|
+
}
|
|
100
|
+
},
|
|
101
|
+
required: ['findings']
|
|
102
|
+
}
|
|
103
|
+
|
|
104
|
+
const VERDICT_SCHEMA = {
|
|
105
|
+
type: 'object',
|
|
106
|
+
additionalProperties: false,
|
|
107
|
+
properties: {
|
|
108
|
+
real: { type: 'boolean' },
|
|
109
|
+
reason: { type: 'string' },
|
|
110
|
+
classification: {
|
|
111
|
+
type: 'string',
|
|
112
|
+
enum: ['genuine', 'false-positive', 'out-of-scope']
|
|
113
|
+
}
|
|
114
|
+
},
|
|
115
|
+
required: ['real', 'reason', 'classification']
|
|
116
|
+
}
|
|
117
|
+
|
|
118
|
+
const RULES = `Repo: ${repo}. Diff base: ${base}.
|
|
119
|
+
IN-SCOPE CHANGED FILES (only these):
|
|
120
|
+
${fileList}
|
|
121
|
+
|
|
122
|
+
Hard rules:
|
|
123
|
+
- Learn the established convention from the EXISTING (unchanged) codebase and MEASURE prevalence with counts (grep -c both forms). If both forms are widely used, it is an accepted dual-convention — DO NOT flag it.
|
|
124
|
+
- Flag only deviations the diff INTRODUCED. Never flag pre-existing untouched code or unrelated concurrent edits.
|
|
125
|
+
- Every finding needs file:line, the deviation, the established alternative, and a prevalence count. No count → not a finding.`
|
|
126
|
+
|
|
127
|
+
log(`Auditing ${changedFiles.length} files across ${DIMENSIONS.length} dimensions vs ${base}`)
|
|
128
|
+
|
|
129
|
+
const reviewed = await pipeline(
|
|
130
|
+
DIMENSIONS,
|
|
131
|
+
(d) =>
|
|
132
|
+
agent(`${RULES}\n\nDimension: ${d.key}\nCheck: ${d.prompt}`, {
|
|
133
|
+
label: `review:${d.key}`,
|
|
134
|
+
phase: 'Review',
|
|
135
|
+
schema: FINDINGS_SCHEMA,
|
|
136
|
+
agentType: 'Explore'
|
|
137
|
+
}),
|
|
138
|
+
(review, d) =>
|
|
139
|
+
parallel(
|
|
140
|
+
(review?.findings || []).map((f) => () =>
|
|
141
|
+
agent(
|
|
142
|
+
`Adversarially verify this consistency finding. READ the actual source lines first. Confirm ALL of: (1) introduced by the diff vs ${base}; (2) a genuine deviation and NOT an accepted dual-convention — re-measure prevalence yourself; (3) in scope (a changed file, not pre-existing/concurrent work). Default to real=false when uncertain.\n\nFinding:\n${JSON.stringify(f, null, 2)}`,
|
|
143
|
+
{ label: `verify:${d.key}`, phase: 'Verify', schema: VERDICT_SCHEMA }
|
|
144
|
+
).then((v) => ({ ...f, dimension: d.key, verdict: v }))
|
|
145
|
+
)
|
|
146
|
+
)
|
|
147
|
+
)
|
|
148
|
+
|
|
149
|
+
const all = reviewed.flat().filter(Boolean)
|
|
150
|
+
const confirmed = all.filter((f) => f.verdict?.real && f.verdict?.classification === 'genuine')
|
|
151
|
+
const dismissed = all.filter((f) => !(f.verdict?.real && f.verdict?.classification === 'genuine'))
|
|
152
|
+
|
|
153
|
+
const report = await agent(
|
|
154
|
+
`Write the final consistency report. Dedup by file+line. Rank GENUINE findings by severity (correctness > behavior > i18n/a11y > naming > cosmetic), each as: "[severity] file:line — deviation; codebase uses <established> (<prevalence>). Fix: <one line>". Then list dismissed candidates under "False positives / out of scope" with their one-line reason. Recommend fixes; DO NOT apply them.\n\nGENUINE:\n${JSON.stringify(confirmed, null, 2)}\n\nDISMISSED:\n${JSON.stringify(dismissed.map((d) => ({ file: d.file, line: d.line, title: d.title, verdict: d.verdict })), null, 2)}`,
|
|
155
|
+
{ label: 'synthesize', phase: 'Synthesize' }
|
|
156
|
+
)
|
|
157
|
+
|
|
158
|
+
return { base, scope: changedFiles.length, confirmed, dismissed, report }
|
|
@@ -0,0 +1,153 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: deslop
|
|
3
|
+
description: "Clean up generated code changes — strip slop (dead code, one-use abstractions, redundant guards, comment noise, unnecessary effects), fix what the change broke, and report what can't be safely fixed. No unrelated changes. Use when the user says 'deslop', 'mdeslop', 'clean this up', 'review my changes', or after generating/modifying a chunk of code."
|
|
4
|
+
metadata:
|
|
5
|
+
author: pratikpakhale
|
|
6
|
+
version: "4.0.0"
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
# Deslop
|
|
10
|
+
|
|
11
|
+
Strip the slop out of the change just made, fix what it broke, report what you shouldn't
|
|
12
|
+
touch yourself. Fixing is the default — you are not writing a review.
|
|
13
|
+
|
|
14
|
+
## Rules
|
|
15
|
+
|
|
16
|
+
- **Scope is the diff.** Only code the change introduced or touched. Pre-existing code is
|
|
17
|
+
off-limits unless this change actively broke it.
|
|
18
|
+
- **No unrelated changes.** No opportunistic refactors, renames, dep bumps, or reformatting.
|
|
19
|
+
- **Fix quietly, flag loudly.** Apply cleanups without narrating each one. Surface only what
|
|
20
|
+
needed a judgment call you couldn't make.
|
|
21
|
+
- **Trace before deleting.** One `rg` for every symbol you're about to remove.
|
|
22
|
+
|
|
23
|
+
## Load what applies
|
|
24
|
+
|
|
25
|
+
Read the reference for every stack present in the diff, **before** starting the hunt:
|
|
26
|
+
|
|
27
|
+
| In the diff | Read |
|
|
28
|
+
| --- | --- |
|
|
29
|
+
| `.tsx`, React, Next.js, RSC | [references/react.md](references/react.md) |
|
|
30
|
+
| `.ts`, Node, Bun, backend TS | [references/typescript.md](references/typescript.md) |
|
|
31
|
+
| `.py` | [references/python.md](references/python.md) |
|
|
32
|
+
|
|
33
|
+
Nothing matches? Run the hunt below on its own — it's language-agnostic.
|
|
34
|
+
|
|
35
|
+
## The hunt
|
|
36
|
+
|
|
37
|
+
Four passes, in order. The map decides what counts as slop; the trace is where the real
|
|
38
|
+
bugs are.
|
|
39
|
+
|
|
40
|
+
### 1. Map — what was this change *for*?
|
|
41
|
+
|
|
42
|
+
Read the diff (`git diff`, plus untracked files) and write the change's intent as one
|
|
43
|
+
sentence. Keep it in front of you: **anything in the diff that doesn't serve that sentence
|
|
44
|
+
is a removal candidate.** Separate out unrelated in-flight edits now and leave them alone.
|
|
45
|
+
|
|
46
|
+
Cheap and worth it first: does the change actually do what it claimed? A feature that's
|
|
47
|
+
wired but never called, or a branch that can't be reached from any entry point, is the
|
|
48
|
+
biggest possible piece of slop.
|
|
49
|
+
|
|
50
|
+
### 2. Trace — what did it break?
|
|
51
|
+
|
|
52
|
+
The only pass where you read *outside* the diff. For every function, component, type,
|
|
53
|
+
constant, or table column the diff modified, `rg -w <name>` and read each call site.
|
|
54
|
+
|
|
55
|
+
Look for callers that assumed the old shape:
|
|
56
|
+
|
|
57
|
+
- return type widened/narrowed, or a field renamed
|
|
58
|
+
- a default value or optional param changed
|
|
59
|
+
- errors that used to throw now return `null` (or the reverse)
|
|
60
|
+
- shared state written by the new code and read somewhere else
|
|
61
|
+
- a component whose props changed but whose other usages weren't updated
|
|
62
|
+
|
|
63
|
+
Fix these. This is where regressions actually live.
|
|
64
|
+
|
|
65
|
+
### 3. Strip — read each hunk against these tells
|
|
66
|
+
|
|
67
|
+
Hunt for *tells*, not categories. Each one below is something you can literally see in the
|
|
68
|
+
diff, paired with the question that settles it.
|
|
69
|
+
|
|
70
|
+
**One-use abstraction.** Tell: a helper, hook, wrapper, interface, factory, or options
|
|
71
|
+
object with exactly one caller — or a wrapper that forwards its arguments unchanged.
|
|
72
|
+
Ask: if I inline this, what gets worse? Nothing → inline it.
|
|
73
|
+
|
|
74
|
+
**Dead weight.** Tell: a symbol the diff introduced. `rg -w` it — if the definition is the
|
|
75
|
+
only hit, it's dead. Same for exports nobody imports, params never read, branches whose
|
|
76
|
+
condition can't be true, and imports left behind by an edit.
|
|
77
|
+
|
|
78
|
+
**Impossible guard.** Tell: `if (!x)` where `x` is non-nullable, a `try/catch` that only
|
|
79
|
+
rethrows, a re-check of something the caller already checked, a fallback for a state the
|
|
80
|
+
type system forbids. Ask: write the concrete input that makes this branch fire. Can't? Delete it.
|
|
81
|
+
|
|
82
|
+
**Half-done work.** Tell: `TODO`, `FIXME`, stub returns, hardcoded values that should be
|
|
83
|
+
arguments, `catch {}` that swallows, error paths that log and continue as if nothing happened.
|
|
84
|
+
Finish it or flag it — never leave it silent.
|
|
85
|
+
|
|
86
|
+
**Comment noise.** Tell: a comment that restates the line under it, a JSDoc block repeating
|
|
87
|
+
the signature, section banners, `// Added for X` changelog notes, commented-out code.
|
|
88
|
+
Delete all of it. Keep only comments explaining a non-obvious *why*.
|
|
89
|
+
|
|
90
|
+
**Verbosity.** Tell: a 12-line block that's one expression, an intermediate variable used
|
|
91
|
+
once on the next line, an if/else assigning the same variable, repeated near-identical
|
|
92
|
+
blocks that differ by one value.
|
|
93
|
+
Ask: does the shorter form lose any clarity? No → shorten it.
|
|
94
|
+
|
|
95
|
+
**User-facing gap.** Tell: a new string literal rendered to a user in a project whose sibling
|
|
96
|
+
files call `t()` / `useTranslations`; a `catch` that renders the raw error object; a mutating
|
|
97
|
+
action with no pending or success state.
|
|
98
|
+
Ask: does someone on a slow connection, in another locale, hitting the failure path, see
|
|
99
|
+
something sensible?
|
|
100
|
+
|
|
101
|
+
**Doesn't look like the neighbors.** The highest-yield tell, and the one most often skipped —
|
|
102
|
+
don't judge the new code on its own, open the nearest sibling first and diff the shape. For
|
|
103
|
+
UI that means the sibling *screen*, not just the sibling file.
|
|
104
|
+
|
|
105
|
+
- Code: naming (`getX` vs `fetchX`, `handleY` vs `onY`), error handling, file placement,
|
|
106
|
+
import order, reuse of the existing helper/type/constant instead of a fresh one.
|
|
107
|
+
- UI: a button/input/modal/table hand-rolled when the design system already exports one; a
|
|
108
|
+
one-off `className` where the component takes a `variant`; raw hex/px/radius values where
|
|
109
|
+
tokens exist; spacing off the project's scale; a different icon set; loading, empty, and
|
|
110
|
+
error states shaped differently from the sibling screen's; copy in another voice
|
|
111
|
+
(`Sign in` vs `Log In`, sentence case vs Title Case).
|
|
112
|
+
|
|
113
|
+
Ask: reading only the new lines, would someone who knows this repo think it came from here?
|
|
114
|
+
For UI, put the new screen next to its sibling — would a user notice two different people
|
|
115
|
+
built them?
|
|
116
|
+
|
|
117
|
+
### 4. Flag — what you must not fix yourself
|
|
118
|
+
|
|
119
|
+
Report these; don't act on them unless the user says to.
|
|
120
|
+
|
|
121
|
+
- **Security.** XSS from unsanitized input, hardcoded secrets or tokens, `eval` /
|
|
122
|
+
`innerHTML` / `dangerouslySetInnerHTML`, a new endpoint or action with no auth guard,
|
|
123
|
+
SQL/command injection, path traversal. Say what's exploitable and how — the fix is often
|
|
124
|
+
a design decision.
|
|
125
|
+
- **Architecture.** Is this production-ready or a hack that gets rewritten in a month? Is a
|
|
126
|
+
breaking change wired end-to-end or left half-migrated? Describe the better shape; don't build it.
|
|
127
|
+
- **Ambiguity.** Anything where the correct answer depends on intent you don't have.
|
|
128
|
+
|
|
129
|
+
## Large diffs
|
|
130
|
+
|
|
131
|
+
Above ~10 files or ~500 changed lines, fan out rather than reading serially.
|
|
132
|
+
|
|
133
|
+
1. Partition changed files into **disjoint** groups — by module or feature, keeping callers
|
|
134
|
+
with their callees. Disjoint is load-bearing: two agents editing one file clobber each other.
|
|
135
|
+
2. Spawn one Agent per group in a single message. Give each: its file list, the intent
|
|
136
|
+
sentence from pass 1, the reference file(s) for its stack, and the boundary — *fix only
|
|
137
|
+
inside your files, report anything crossing the boundary back to me*.
|
|
138
|
+
3. Do pass 2 and the consistency tell yourself, across the whole diff, once they return.
|
|
139
|
+
Cross-group breakage and duplicated logic are invisible from inside a partition.
|
|
140
|
+
|
|
141
|
+
Use a Workflow only if the user explicitly opted into orchestration ("use a workflow",
|
|
142
|
+
ultracode) — then pipeline each group through fix → verify and synthesize. Otherwise plain
|
|
143
|
+
Agent fan-out.
|
|
144
|
+
|
|
145
|
+
## Output
|
|
146
|
+
|
|
147
|
+
Short.
|
|
148
|
+
|
|
149
|
+
1. **Stripped** — what you removed and fixed, grouped, one line each.
|
|
150
|
+
2. **Flagged** — pass 4 findings, severity-ordered, each with `file:line`.
|
|
151
|
+
3. **Open questions** — only if a decision is genuinely yours to ask about.
|
|
152
|
+
|
|
153
|
+
No preamble, no restating the diff back.
|
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
# Python slop
|
|
2
|
+
|
|
3
|
+
Stack-specific tells. Run these alongside the main hunt, not instead of it.
|
|
4
|
+
|
|
5
|
+
## Exception handling
|
|
6
|
+
|
|
7
|
+
`rg -n "except" <changed files>` and check each one the diff added.
|
|
8
|
+
|
|
9
|
+
- `except Exception:` / bare `except:` — what specific error is expected here? Narrow it.
|
|
10
|
+
- `except ... : pass` — silently eats real failures, including `KeyboardInterrupt` on a bare
|
|
11
|
+
clause.
|
|
12
|
+
- `try/except` that only logs and re-raises — delete, let it propagate.
|
|
13
|
+
- A `try` block wrapping ten lines when one of them raises — shrink it to that line.
|
|
14
|
+
- `except` returning a success-shaped value (`return []`, `return None`) so callers can't tell
|
|
15
|
+
failure from empty.
|
|
16
|
+
|
|
17
|
+
## Defaults & mutation
|
|
18
|
+
|
|
19
|
+
- Mutable default arguments (`def f(x=[])`, `={}`) — classic shared-state bug.
|
|
20
|
+
- A function mutating a list/dict passed in by the caller without saying so in the name.
|
|
21
|
+
- Module-level mutable state written at import time.
|
|
22
|
+
- `dataclass` field with a mutable default missing `field(default_factory=...)`.
|
|
23
|
+
|
|
24
|
+
## Redundant checks
|
|
25
|
+
|
|
26
|
+
- `if x is not None:` on a value the type hint declares non-optional and the caller guarantees.
|
|
27
|
+
- `if len(xs) > 0:` — use `if xs:`.
|
|
28
|
+
- `if x == True:` / `if x != None:`.
|
|
29
|
+
- `hasattr` / `in dict` guards for keys the constructor always sets.
|
|
30
|
+
- Re-validating an argument the caller already validated.
|
|
31
|
+
|
|
32
|
+
## Verbosity
|
|
33
|
+
|
|
34
|
+
- A `for` loop appending to a list that's a comprehension.
|
|
35
|
+
- A comprehension nested three deep, or one spanning multiple lines with a condition —
|
|
36
|
+
that's a loop, write the loop.
|
|
37
|
+
- `dict(a=1, b=2)` where a literal reads better; `list()`/`dict()` where `[]`/`{}` do.
|
|
38
|
+
- Manual index tracking instead of `enumerate`; parallel indexing instead of `zip`.
|
|
39
|
+
- String building with `+` in a loop instead of `join`.
|
|
40
|
+
- An `if/else` returning `True`/`False` — return the expression.
|
|
41
|
+
|
|
42
|
+
## Types & structure
|
|
43
|
+
|
|
44
|
+
- Type hints on some functions in the file and not others — match the file's convention.
|
|
45
|
+
- `Any` / `dict` / bare `list` where a `TypedDict`, `dataclass`, or Pydantic model exists in
|
|
46
|
+
the codebase already.
|
|
47
|
+
- `Optional[X]` added everywhere to satisfy the checker rather than because `None` is real.
|
|
48
|
+
- A new module-level helper duplicating something in the project's utils package — `rg` the
|
|
49
|
+
behavior, not the name.
|
|
50
|
+
- Imports inside a function without a circular-import or lazy-load reason.
|
|
51
|
+
- A class with one method and no state — that's a function.
|
|
52
|
+
|
|
53
|
+
## I/O & resources
|
|
54
|
+
|
|
55
|
+
- `open()` without a context manager; a file/socket/cursor never closed on the error path.
|
|
56
|
+
- Requests in a loop where a batched call or `asyncio.gather` fits.
|
|
57
|
+
- ORM query in a loop (N+1) — use `select_related` / `prefetch_related` / a single filtered query.
|
|
58
|
+
- `requests` / `httpx` calls with no timeout.
|
|
59
|
+
- Blocking I/O inside an `async def`.
|
|
60
|
+
|
|
61
|
+
## Tests
|
|
62
|
+
|
|
63
|
+
- A test asserting nothing, or asserting only that the call didn't raise.
|
|
64
|
+
- Fixtures added but unused; `mock` patches whose target no longer exists.
|
|
65
|
+
- `time.sleep` used to sequence async behavior.
|