sherpa-mcp 0.1.0 → 0.1.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +310 -0
- package/package.json +3 -3
package/README.md
ADDED
|
@@ -0,0 +1,310 @@
|
|
|
1
|
+
English | [Español](./README.es.md)
|
|
2
|
+
|
|
3
|
+
# sherpa
|
|
4
|
+
|
|
5
|
+
**A Claude Code plugin that offloads the heavy lifting — reading, searching, and batch-editing code — to your local LLM.**
|
|
6
|
+
|
|
7
|
+
[](./LICENSE)
|
|
8
|
+
[](https://nodejs.org)
|
|
9
|
+
[](https://github.com/Tongas/sherpa-mcp/pulls)
|
|
10
|
+
|
|
11
|
+
Use your Claude subscription and your local AI at the same time, inside
|
|
12
|
+
Claude Code. Claude does the thinking; your local model does the heavy
|
|
13
|
+
reading. You save tokens.
|
|
14
|
+
|
|
15
|
+
## How it works
|
|
16
|
+
|
|
17
|
+
When you delegate a task, your local model reads, searches, or rewrites
|
|
18
|
+
files on your own hardware. Only a short summary comes back into Claude
|
|
19
|
+
Code's context — the file contents never enter it. The sherpa carries the
|
|
20
|
+
weight; you still decide the route.
|
|
21
|
+
|
|
22
|
+
sherpa ships as a Claude Code plugin: an MCP server with five tools, a
|
|
23
|
+
skill that teaches Claude when delegating is worth it, and a
|
|
24
|
+
`/sherpa-status` command for diagnosing your setup.
|
|
25
|
+
|
|
26
|
+
## The evidence
|
|
27
|
+
|
|
28
|
+
Exploring the same 3-layer Python codebase, same prompt, two ways:
|
|
29
|
+
|
|
30
|
+
| | Direct (Claude reads the files) | With sherpa |
|
|
31
|
+
|---|---|---|
|
|
32
|
+
| Tokens entering Claude's context | ~190,000 (the file content itself) | ~600 (just the summary) |
|
|
33
|
+
| Tokens processed locally | 0 | ~149,000 |
|
|
34
|
+
| Wall-clock time | ~39s | ~100s |
|
|
35
|
+
| Files covered | 41 | 24 |
|
|
36
|
+
|
|
37
|
+
Two things matter more than the ratio:
|
|
38
|
+
|
|
39
|
+
**The delegated summary held up.** Claude had already mapped the same
|
|
40
|
+
codebase manually in an earlier turn. When the sherpa summary came back,
|
|
41
|
+
it matched: same check count, same internal patterns, same data-model
|
|
42
|
+
chain. The local model didn't produce a vaguer answer — it produced the
|
|
43
|
+
same answer without spending orchestrator context.
|
|
44
|
+
|
|
45
|
+
**Delegating is slower, not faster.** Roughly 2–4x in wall-clock time.
|
|
46
|
+
The value here is context saved, not speed. See
|
|
47
|
+
[When NOT to use it](#when-not-to-use-it).
|
|
48
|
+
|
|
49
|
+
**This is one measurement, on one repo, with one model — not a
|
|
50
|
+
benchmark.** Your ratio depends on file sizes, the instruction, and the
|
|
51
|
+
model you run.
|
|
52
|
+
|
|
53
|
+
## Quickstart
|
|
54
|
+
|
|
55
|
+
Prerequisites: Node.js ≥ 18, [ripgrep](https://github.com/BurntSushi/ripgrep#installation)
|
|
56
|
+
(`rg`) on `PATH`, and a local backend running (Ollama, llama.cpp server,
|
|
57
|
+
or LM Studio).
|
|
58
|
+
|
|
59
|
+
### Main path: install the plugin
|
|
60
|
+
|
|
61
|
+
This gets you the MCP server, the skill, and `/sherpa-status` together,
|
|
62
|
+
in one shot:
|
|
63
|
+
|
|
64
|
+
```
|
|
65
|
+
/plugin marketplace add Tongas/sherpa-mcp
|
|
66
|
+
/plugin install sherpa@sherpa-mcp
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
**Configure the backend.** A plugin-provided MCP server inherits Claude
|
|
70
|
+
Code's own process environment — there's no documented way to attach a
|
|
71
|
+
per-plugin `env` block after a marketplace install. That makes shell
|
|
72
|
+
exports fragile: if you launch Claude Code from a GUI launcher instead
|
|
73
|
+
of a terminal, it doesn't inherit anything from `~/.bashrc`/`~/.zshrc`.
|
|
74
|
+
|
|
75
|
+
**Recommended: a config file**, which sherpa reads regardless of how
|
|
76
|
+
Claude Code was launched. Create `~/.claude/sherpa/config.json` (applies
|
|
77
|
+
everywhere) or `./sherpa.config.json` in a specific project (same keys,
|
|
78
|
+
camelCase):
|
|
79
|
+
|
|
80
|
+
```json
|
|
81
|
+
{
|
|
82
|
+
"backend": "openai-compatible",
|
|
83
|
+
"baseUrl": "http://localhost:8080",
|
|
84
|
+
"model": "qwen2.5-coder-14b",
|
|
85
|
+
"contextWindowOverride": 32768,
|
|
86
|
+
"maxOutputTokensOverride": 8192
|
|
87
|
+
}
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
(For Ollama, drop `backend`/`contextWindowOverride`/`maxOutputTokensOverride`
|
|
91
|
+
— just `baseUrl` and `model` are enough; see
|
|
92
|
+
[Configuration](#configuration) below for the full key list.)
|
|
93
|
+
|
|
94
|
+
**Quick alternative:** if you're always launching Claude Code from a
|
|
95
|
+
shell, exporting env vars works too:
|
|
96
|
+
|
|
97
|
+
```bash
|
|
98
|
+
export SHERPA_BASE_URL="http://localhost:11434" # Ollama default
|
|
99
|
+
export SHERPA_MODEL="qwen2.5-coder:14b" # whatever you have pulled
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
For an `openai-compatible` backend (llama.cpp server, LM Studio), also
|
|
103
|
+
export `SHERPA_BACKEND=openai-compatible` plus `SHERPA_CONTEXT_WINDOW`
|
|
104
|
+
and `SHERPA_MAX_OUTPUT_TOKENS` set to your server's real values. There's
|
|
105
|
+
no standard endpoint to discover the context window, so without one of
|
|
106
|
+
these two config methods sherpa falls back to a conservative 4096/2048,
|
|
107
|
+
which makes `delegate_transform` skip files over roughly 200 lines.
|
|
108
|
+
|
|
109
|
+
**Verify:** open a new session and run `/sherpa-status`. It shows the
|
|
110
|
+
active backend, the loaded model, and where each config value actually
|
|
111
|
+
came from, so a typo doesn't go unnoticed.
|
|
112
|
+
|
|
113
|
+
### Alternative: MCP server only, via npx
|
|
114
|
+
|
|
115
|
+
Use this if you just want the tools — for example, wiring sherpa into
|
|
116
|
+
something other than Claude Code. **You won't get the skill or
|
|
117
|
+
`/sherpa-status`**, which means no automatic guidance on when delegating
|
|
118
|
+
is worth it and no built-in way to check what's configured; you'll need
|
|
119
|
+
to invoke the tools explicitly and know your own setup.
|
|
120
|
+
|
|
121
|
+
Add this to `~/.claude.json` (or a project-level `.mcp.json`) — here the
|
|
122
|
+
`env` block is explicit and does work, since you're registering the MCP
|
|
123
|
+
server directly rather than through a plugin:
|
|
124
|
+
|
|
125
|
+
```json
|
|
126
|
+
{
|
|
127
|
+
"mcpServers": {
|
|
128
|
+
"sherpa": {
|
|
129
|
+
"command": "npx",
|
|
130
|
+
"args": ["-y", "sherpa-mcp"],
|
|
131
|
+
"env": {
|
|
132
|
+
"SHERPA_BASE_URL": "http://localhost:11434",
|
|
133
|
+
"SHERPA_MODEL": "qwen2.5-coder:14b"
|
|
134
|
+
}
|
|
135
|
+
}
|
|
136
|
+
}
|
|
137
|
+
}
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
For an `openai-compatible` backend, add `SHERPA_BACKEND`,
|
|
141
|
+
`SHERPA_CONTEXT_WINDOW`, and `SHERPA_MAX_OUTPUT_TOKENS` to that same
|
|
142
|
+
`env` block, same as above.
|
|
143
|
+
|
|
144
|
+
Tested with llama.cpp server. Also supports Ollama and LM Studio through
|
|
145
|
+
the same OpenAI-compatible interface.
|
|
146
|
+
|
|
147
|
+
#### Installing from a clone (development)
|
|
148
|
+
|
|
149
|
+
If you're working on `sherpa` itself, build locally instead of using
|
|
150
|
+
`npx`:
|
|
151
|
+
|
|
152
|
+
```bash
|
|
153
|
+
cd mcp-server
|
|
154
|
+
npm install
|
|
155
|
+
npm run build
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
Then point `command`/`args` at the built entrypoint instead of `npx`:
|
|
159
|
+
|
|
160
|
+
```json
|
|
161
|
+
{
|
|
162
|
+
"command": "node",
|
|
163
|
+
"args": ["/path/to/sherpa-mcp/mcp-server/dist/index.js"]
|
|
164
|
+
}
|
|
165
|
+
```
|
|
166
|
+
|
|
167
|
+
## Example prompts
|
|
168
|
+
|
|
169
|
+
These are prompts we actually ran, not hypothetical ones. Name sherpa in
|
|
170
|
+
the prompt — see [Automatic vs. explicit invocation](#automatic-vs-explicit-invocation)
|
|
171
|
+
for why.
|
|
172
|
+
|
|
173
|
+
**1. Codebase reconnaissance — the main case:**
|
|
174
|
+
|
|
175
|
+
> Use sherpa to explore this project and tell me how the audit logic is
|
|
176
|
+
> structured
|
|
177
|
+
|
|
178
|
+
Measured: 24 files, ~149k tokens processed locally, ~100s, and only the
|
|
179
|
+
summary entered Claude's context.
|
|
180
|
+
|
|
181
|
+
**2. Search with synthesis:**
|
|
182
|
+
|
|
183
|
+
> Use sherpa to find every usage of the old logging API and summarize
|
|
184
|
+
> what needs to change
|
|
185
|
+
|
|
186
|
+
ripgrep does the searching; the local model only synthesizes the matches.
|
|
187
|
+
|
|
188
|
+
**3. Batch transformation:**
|
|
189
|
+
|
|
190
|
+
> Use sherpa to rename the config key `oldName` to `newName` across the
|
|
191
|
+
> project
|
|
192
|
+
|
|
193
|
+
`dry_run` is the default: you get a reviewable proposal with per-file
|
|
194
|
+
diffs, nothing is written. Then `apply_transform` writes exactly what you
|
|
195
|
+
reviewed — no regeneration, and it refuses any file that changed on disk
|
|
196
|
+
in the meantime.
|
|
197
|
+
|
|
198
|
+
## Automatic vs. explicit invocation
|
|
199
|
+
|
|
200
|
+
The plugin ships a skill that teaches Claude when delegating is worth it,
|
|
201
|
+
and it does fire on its own sometimes. But in practice, automatic
|
|
202
|
+
activation is not reliable for exploration: Claude often reaches for its
|
|
203
|
+
native tools (`Read`, `Grep`, the Explore agent) even with the skill
|
|
204
|
+
loaded and the tools visible.
|
|
205
|
+
|
|
206
|
+
**For a guarantee, name sherpa in the prompt.** Explicit invocation works
|
|
207
|
+
consistently.
|
|
208
|
+
|
|
209
|
+
Measured on Claude Code v2.x. This may change in future versions.
|
|
210
|
+
|
|
211
|
+
## When NOT to use it
|
|
212
|
+
|
|
213
|
+
If Claude already has the project map in context — files already read
|
|
214
|
+
this session, or a small codebase — direct tools win every time.
|
|
215
|
+
Delegating costs a round trip; reading two known files does not.
|
|
216
|
+
|
|
217
|
+
Delegating is slower, not faster. The value is context saved, not speed.
|
|
218
|
+
|
|
219
|
+
And v1 doesn't write new code: `delegate_transform` performs mechanical
|
|
220
|
+
transformations on files that already exist (renaming, repetitive
|
|
221
|
+
boilerplate), not new features or logic. See `skills/sherpa/SKILL.md` for
|
|
222
|
+
the full when-to/when-not-to table.
|
|
223
|
+
|
|
224
|
+
## The tools
|
|
225
|
+
|
|
226
|
+
| Tool | What it does |
|
|
227
|
+
|---|---|
|
|
228
|
+
| `health_check` | Checks whether the local backend is available and which model is loaded. |
|
|
229
|
+
| `delegate_exploration` | Reads many files/directories and returns a synthesis, without that content passing through Claude's context. |
|
|
230
|
+
| `delegate_search` | Runs ripgrep over the given paths and synthesizes the matches according to an instruction. |
|
|
231
|
+
| `delegate_transform` | Proposes a per-file batch transformation (never writes directly — generates a reviewable proposal). |
|
|
232
|
+
| `apply_transform` | Writes to disk exactly what a prior `delegate_transform` proposed, with a staleness check. |
|
|
233
|
+
|
|
234
|
+
## Configuration
|
|
235
|
+
|
|
236
|
+
Works with zero files, env vars only (see Quickstart for where they go).
|
|
237
|
+
Precedence: MCP server env > `sherpa.config.json` (project) >
|
|
238
|
+
`~/.claude/sherpa/config.json` (user) > defaults.
|
|
239
|
+
|
|
240
|
+
| Variable | Default | Purpose |
|
|
241
|
+
|---|---|---|
|
|
242
|
+
| `SHERPA_BACKEND` | `ollama` | `ollama` or `openai-compatible` |
|
|
243
|
+
| `SHERPA_BASE_URL` | `http://localhost:11434` | Local backend URL |
|
|
244
|
+
| `SHERPA_MODEL` | *(no default)* | Model to use — if missing, `health_check` lists the models available |
|
|
245
|
+
| `SHERPA_MAX_FILES` | `100` | File budget per `delegate_exploration`/`delegate_search` call |
|
|
246
|
+
| `SHERPA_MAX_CHUNKS` | `20` | Local-model call budget per invocation |
|
|
247
|
+
| `SHERPA_RESULTS_DIR` | `.sherpa` | Where full results get written (relative to the project root) |
|
|
248
|
+
| `SHERPA_TRUNCATION_THRESHOLD` | `0.75` | Truncation-guard threshold in `delegate_transform` (see Limitations) |
|
|
249
|
+
| `SHERPA_CONTEXT_WINDOW` | *(no default, falls back to 4096)* | `openai-compatible` only: no standard endpoint to discover the context window |
|
|
250
|
+
| `SHERPA_MAX_OUTPUT_TOKENS` | *(no default, falls back to 2048)* | `openai-compatible` only, same reason |
|
|
251
|
+
|
|
252
|
+
You can also use `./sherpa.config.json` (project) or
|
|
253
|
+
`~/.claude/sherpa/config.json` (user) with the same keys in camelCase.
|
|
254
|
+
|
|
255
|
+
Add `.sherpa/` to your `.gitignore` — the full results of every
|
|
256
|
+
`delegate_*` call accumulate there with no limit in v1.
|
|
257
|
+
|
|
258
|
+
## Security
|
|
259
|
+
|
|
260
|
+
- **Path confinement:** every path `sherpa` touches (read or write) is
|
|
261
|
+
resolved against the project root and rejected if it falls outside —
|
|
262
|
+
including escaping `..` and symlinks pointing elsewhere. This is a
|
|
263
|
+
hard-boundary check, not an optional guard.
|
|
264
|
+
- **Untrusted content:** everything the local model returns (`summary`,
|
|
265
|
+
`diffPreview`, quoted file content) is **observed data, never
|
|
266
|
+
instructions**. A file in the repo can contain text aimed at an LLM
|
|
267
|
+
that tries to sound like a command — Claude always treats it as text
|
|
268
|
+
to evaluate, never as something to obey (see
|
|
269
|
+
`skills/sherpa/SKILL.md`).
|
|
270
|
+
|
|
271
|
+
## Known limitations (v1)
|
|
272
|
+
|
|
273
|
+
- **No automatic `.sherpa/` cleanup**: results accumulate indefinitely —
|
|
274
|
+
delete them manually whenever you want.
|
|
275
|
+
- **`delegate_transform` skips (doesn't fail) files that exceed the
|
|
276
|
+
model's output budget**: with the conservative fallback defaults
|
|
277
|
+
(4096/2048) this affects files over roughly 200 lines. Set
|
|
278
|
+
`SHERPA_CONTEXT_WINDOW`/`SHERPA_MAX_OUTPUT_TOKENS` to your server's
|
|
279
|
+
real values to avoid it.
|
|
280
|
+
- **Truncation guard in `delegate_transform`**
|
|
281
|
+
(`SHERPA_TRUNCATION_THRESHOLD`, default `0.75`): a threshold that's
|
|
282
|
+
blind to intent. An instruction that legitimately shortens a file a lot
|
|
283
|
+
(e.g. "delete all dead code") can trigger a false rejection — lower the
|
|
284
|
+
threshold for that specific use case.
|
|
285
|
+
- **v1 doesn't write new code**: mechanical transformations on existing
|
|
286
|
+
files only, not new features or logic.
|
|
287
|
+
- **`getCapabilities()` on `openai-compatible` backends**: no standard
|
|
288
|
+
endpoint to discover the context window on llama.cpp server / LM
|
|
289
|
+
Studio. If you change the loaded model without updating the env vars,
|
|
290
|
+
chunking will use stale values.
|
|
291
|
+
- **No automatic fallback or retry** if the local backend doesn't
|
|
292
|
+
respond: this is intentional — Claude does the task itself and moves
|
|
293
|
+
on, without interrupting your session.
|
|
294
|
+
- **TOCTOU in `apply_transform`'s staleness guard**: there's an
|
|
295
|
+
unavoidable window between reading the hash and writing the file
|
|
296
|
+
(Node's sync `fs` APIs don't offer an atomic "check and write"
|
|
297
|
+
operation for this case).
|
|
298
|
+
- **`file-enumeration.ts` doesn't follow symlinks**: a source tree that
|
|
299
|
+
is (or contains) a symlink returns zero files instead of an error —
|
|
300
|
+
there's no explicit warning that a symlink is being enumerated.
|
|
301
|
+
|
|
302
|
+
## License
|
|
303
|
+
|
|
304
|
+
MIT — see [LICENSE](./LICENSE).
|
|
305
|
+
|
|
306
|
+
---
|
|
307
|
+
|
|
308
|
+
Built by [Gastón Parravicini](https://github.com/Tongas).
|
|
309
|
+
|
|
310
|
+
Built with AI assistance.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "sherpa-mcp",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.2",
|
|
4
4
|
"description": "MCP server that delegates high-volume file exploration, search, and transformation to a local LLM backend",
|
|
5
5
|
"author": "Gastón Parravicini",
|
|
6
6
|
"license": "MIT",
|
|
@@ -15,7 +15,7 @@
|
|
|
15
15
|
},
|
|
16
16
|
"type": "module",
|
|
17
17
|
"bin": {
|
|
18
|
-
"sherpa-mcp": "
|
|
18
|
+
"sherpa-mcp": "dist/index.js"
|
|
19
19
|
},
|
|
20
20
|
"main": "./dist/index.js",
|
|
21
21
|
"files": [
|
|
@@ -24,7 +24,7 @@
|
|
|
24
24
|
"scripts": {
|
|
25
25
|
"build": "tsc -p tsconfig.json",
|
|
26
26
|
"prepare": "npm run build",
|
|
27
|
-
"prepublishOnly": "npm run build",
|
|
27
|
+
"prepublishOnly": "cp ../README.md ./README.md && npm run build",
|
|
28
28
|
"test": "vitest run",
|
|
29
29
|
"start": "node dist/index.js"
|
|
30
30
|
},
|