sherpa-mcp 0.1.0 → 0.1.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (2) hide show
  1. package/README.md +310 -0
  2. package/package.json +3 -3
package/README.md ADDED
@@ -0,0 +1,310 @@
1
+ English | [Español](./README.es.md)
2
+
3
+ # sherpa
4
+
5
+ **A Claude Code plugin that offloads the heavy lifting — reading, searching, and batch-editing code — to your local LLM.**
6
+
7
+ [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](./LICENSE)
8
+ [![Node](https://img.shields.io/badge/node-%E2%89%A518-brightgreen.svg)](https://nodejs.org)
9
+ [![PRs welcome](https://img.shields.io/badge/PRs-welcome-brightgreen.svg)](https://github.com/Tongas/sherpa-mcp/pulls)
10
+
11
+ Use your Claude subscription and your local AI at the same time, inside
12
+ Claude Code. Claude does the thinking; your local model does the heavy
13
+ reading. You save tokens.
14
+
15
+ ## How it works
16
+
17
+ When you delegate a task, your local model reads, searches, or rewrites
18
+ files on your own hardware. Only a short summary comes back into Claude
19
+ Code's context — the file contents never enter it. The sherpa carries the
20
+ weight; you still decide the route.
21
+
22
+ sherpa ships as a Claude Code plugin: an MCP server with five tools, a
23
+ skill that teaches Claude when delegating is worth it, and a
24
+ `/sherpa-status` command for diagnosing your setup.
25
+
26
+ ## The evidence
27
+
28
+ Exploring the same 3-layer Python codebase, same prompt, two ways:
29
+
30
+ | | Direct (Claude reads the files) | With sherpa |
31
+ |---|---|---|
32
+ | Tokens entering Claude's context | ~190,000 (the file content itself) | ~600 (just the summary) |
33
+ | Tokens processed locally | 0 | ~149,000 |
34
+ | Wall-clock time | ~39s | ~100s |
35
+ | Files covered | 41 | 24 |
36
+
37
+ Two things matter more than the ratio:
38
+
39
+ **The delegated summary held up.** Claude had already mapped the same
40
+ codebase manually in an earlier turn. When the sherpa summary came back,
41
+ it matched: same check count, same internal patterns, same data-model
42
+ chain. The local model didn't produce a vaguer answer — it produced the
43
+ same answer without spending orchestrator context.
44
+
45
+ **Delegating is slower, not faster.** Roughly 2–4x in wall-clock time.
46
+ The value here is context saved, not speed. See
47
+ [When NOT to use it](#when-not-to-use-it).
48
+
49
+ **This is one measurement, on one repo, with one model — not a
50
+ benchmark.** Your ratio depends on file sizes, the instruction, and the
51
+ model you run.
52
+
53
+ ## Quickstart
54
+
55
+ Prerequisites: Node.js ≥ 18, [ripgrep](https://github.com/BurntSushi/ripgrep#installation)
56
+ (`rg`) on `PATH`, and a local backend running (Ollama, llama.cpp server,
57
+ or LM Studio).
58
+
59
+ ### Main path: install the plugin
60
+
61
+ This gets you the MCP server, the skill, and `/sherpa-status` together,
62
+ in one shot:
63
+
64
+ ```
65
+ /plugin marketplace add Tongas/sherpa-mcp
66
+ /plugin install sherpa@sherpa-mcp
67
+ ```
68
+
69
+ **Configure the backend.** A plugin-provided MCP server inherits Claude
70
+ Code's own process environment — there's no documented way to attach a
71
+ per-plugin `env` block after a marketplace install. That makes shell
72
+ exports fragile: if you launch Claude Code from a GUI launcher instead
73
+ of a terminal, it doesn't inherit anything from `~/.bashrc`/`~/.zshrc`.
74
+
75
+ **Recommended: a config file**, which sherpa reads regardless of how
76
+ Claude Code was launched. Create `~/.claude/sherpa/config.json` (applies
77
+ everywhere) or `./sherpa.config.json` in a specific project (same keys,
78
+ camelCase):
79
+
80
+ ```json
81
+ {
82
+ "backend": "openai-compatible",
83
+ "baseUrl": "http://localhost:8080",
84
+ "model": "qwen2.5-coder-14b",
85
+ "contextWindowOverride": 32768,
86
+ "maxOutputTokensOverride": 8192
87
+ }
88
+ ```
89
+
90
+ (For Ollama, drop `backend`/`contextWindowOverride`/`maxOutputTokensOverride`
91
+ — just `baseUrl` and `model` are enough; see
92
+ [Configuration](#configuration) below for the full key list.)
93
+
94
+ **Quick alternative:** if you're always launching Claude Code from a
95
+ shell, exporting env vars works too:
96
+
97
+ ```bash
98
+ export SHERPA_BASE_URL="http://localhost:11434" # Ollama default
99
+ export SHERPA_MODEL="qwen2.5-coder:14b" # whatever you have pulled
100
+ ```
101
+
102
+ For an `openai-compatible` backend (llama.cpp server, LM Studio), also
103
+ export `SHERPA_BACKEND=openai-compatible` plus `SHERPA_CONTEXT_WINDOW`
104
+ and `SHERPA_MAX_OUTPUT_TOKENS` set to your server's real values. There's
105
+ no standard endpoint to discover the context window, so without one of
106
+ these two config methods sherpa falls back to a conservative 4096/2048,
107
+ which makes `delegate_transform` skip files over roughly 200 lines.
108
+
109
+ **Verify:** open a new session and run `/sherpa-status`. It shows the
110
+ active backend, the loaded model, and where each config value actually
111
+ came from, so a typo doesn't go unnoticed.
112
+
113
+ ### Alternative: MCP server only, via npx
114
+
115
+ Use this if you just want the tools — for example, wiring sherpa into
116
+ something other than Claude Code. **You won't get the skill or
117
+ `/sherpa-status`**, which means no automatic guidance on when delegating
118
+ is worth it and no built-in way to check what's configured; you'll need
119
+ to invoke the tools explicitly and know your own setup.
120
+
121
+ Add this to `~/.claude.json` (or a project-level `.mcp.json`) — here the
122
+ `env` block is explicit and does work, since you're registering the MCP
123
+ server directly rather than through a plugin:
124
+
125
+ ```json
126
+ {
127
+ "mcpServers": {
128
+ "sherpa": {
129
+ "command": "npx",
130
+ "args": ["-y", "sherpa-mcp"],
131
+ "env": {
132
+ "SHERPA_BASE_URL": "http://localhost:11434",
133
+ "SHERPA_MODEL": "qwen2.5-coder:14b"
134
+ }
135
+ }
136
+ }
137
+ }
138
+ ```
139
+
140
+ For an `openai-compatible` backend, add `SHERPA_BACKEND`,
141
+ `SHERPA_CONTEXT_WINDOW`, and `SHERPA_MAX_OUTPUT_TOKENS` to that same
142
+ `env` block, same as above.
143
+
144
+ Tested with llama.cpp server. Also supports Ollama and LM Studio through
145
+ the same OpenAI-compatible interface.
146
+
147
+ #### Installing from a clone (development)
148
+
149
+ If you're working on `sherpa` itself, build locally instead of using
150
+ `npx`:
151
+
152
+ ```bash
153
+ cd mcp-server
154
+ npm install
155
+ npm run build
156
+ ```
157
+
158
+ Then point `command`/`args` at the built entrypoint instead of `npx`:
159
+
160
+ ```json
161
+ {
162
+ "command": "node",
163
+ "args": ["/path/to/sherpa-mcp/mcp-server/dist/index.js"]
164
+ }
165
+ ```
166
+
167
+ ## Example prompts
168
+
169
+ These are prompts we actually ran, not hypothetical ones. Name sherpa in
170
+ the prompt — see [Automatic vs. explicit invocation](#automatic-vs-explicit-invocation)
171
+ for why.
172
+
173
+ **1. Codebase reconnaissance — the main case:**
174
+
175
+ > Use sherpa to explore this project and tell me how the audit logic is
176
+ > structured
177
+
178
+ Measured: 24 files, ~149k tokens processed locally, ~100s, and only the
179
+ summary entered Claude's context.
180
+
181
+ **2. Search with synthesis:**
182
+
183
+ > Use sherpa to find every usage of the old logging API and summarize
184
+ > what needs to change
185
+
186
+ ripgrep does the searching; the local model only synthesizes the matches.
187
+
188
+ **3. Batch transformation:**
189
+
190
+ > Use sherpa to rename the config key `oldName` to `newName` across the
191
+ > project
192
+
193
+ `dry_run` is the default: you get a reviewable proposal with per-file
194
+ diffs, nothing is written. Then `apply_transform` writes exactly what you
195
+ reviewed — no regeneration, and it refuses any file that changed on disk
196
+ in the meantime.
197
+
198
+ ## Automatic vs. explicit invocation
199
+
200
+ The plugin ships a skill that teaches Claude when delegating is worth it,
201
+ and it does fire on its own sometimes. But in practice, automatic
202
+ activation is not reliable for exploration: Claude often reaches for its
203
+ native tools (`Read`, `Grep`, the Explore agent) even with the skill
204
+ loaded and the tools visible.
205
+
206
+ **For a guarantee, name sherpa in the prompt.** Explicit invocation works
207
+ consistently.
208
+
209
+ Measured on Claude Code v2.x. This may change in future versions.
210
+
211
+ ## When NOT to use it
212
+
213
+ If Claude already has the project map in context — files already read
214
+ this session, or a small codebase — direct tools win every time.
215
+ Delegating costs a round trip; reading two known files does not.
216
+
217
+ Delegating is slower, not faster. The value is context saved, not speed.
218
+
219
+ And v1 doesn't write new code: `delegate_transform` performs mechanical
220
+ transformations on files that already exist (renaming, repetitive
221
+ boilerplate), not new features or logic. See `skills/sherpa/SKILL.md` for
222
+ the full when-to/when-not-to table.
223
+
224
+ ## The tools
225
+
226
+ | Tool | What it does |
227
+ |---|---|
228
+ | `health_check` | Checks whether the local backend is available and which model is loaded. |
229
+ | `delegate_exploration` | Reads many files/directories and returns a synthesis, without that content passing through Claude's context. |
230
+ | `delegate_search` | Runs ripgrep over the given paths and synthesizes the matches according to an instruction. |
231
+ | `delegate_transform` | Proposes a per-file batch transformation (never writes directly — generates a reviewable proposal). |
232
+ | `apply_transform` | Writes to disk exactly what a prior `delegate_transform` proposed, with a staleness check. |
233
+
234
+ ## Configuration
235
+
236
+ Works with zero files, env vars only (see Quickstart for where they go).
237
+ Precedence: MCP server env > `sherpa.config.json` (project) >
238
+ `~/.claude/sherpa/config.json` (user) > defaults.
239
+
240
+ | Variable | Default | Purpose |
241
+ |---|---|---|
242
+ | `SHERPA_BACKEND` | `ollama` | `ollama` or `openai-compatible` |
243
+ | `SHERPA_BASE_URL` | `http://localhost:11434` | Local backend URL |
244
+ | `SHERPA_MODEL` | *(no default)* | Model to use — if missing, `health_check` lists the models available |
245
+ | `SHERPA_MAX_FILES` | `100` | File budget per `delegate_exploration`/`delegate_search` call |
246
+ | `SHERPA_MAX_CHUNKS` | `20` | Local-model call budget per invocation |
247
+ | `SHERPA_RESULTS_DIR` | `.sherpa` | Where full results get written (relative to the project root) |
248
+ | `SHERPA_TRUNCATION_THRESHOLD` | `0.75` | Truncation-guard threshold in `delegate_transform` (see Limitations) |
249
+ | `SHERPA_CONTEXT_WINDOW` | *(no default, falls back to 4096)* | `openai-compatible` only: no standard endpoint to discover the context window |
250
+ | `SHERPA_MAX_OUTPUT_TOKENS` | *(no default, falls back to 2048)* | `openai-compatible` only, same reason |
251
+
252
+ You can also use `./sherpa.config.json` (project) or
253
+ `~/.claude/sherpa/config.json` (user) with the same keys in camelCase.
254
+
255
+ Add `.sherpa/` to your `.gitignore` — the full results of every
256
+ `delegate_*` call accumulate there with no limit in v1.
257
+
258
+ ## Security
259
+
260
+ - **Path confinement:** every path `sherpa` touches (read or write) is
261
+ resolved against the project root and rejected if it falls outside —
262
+ including escaping `..` and symlinks pointing elsewhere. This is a
263
+ hard-boundary check, not an optional guard.
264
+ - **Untrusted content:** everything the local model returns (`summary`,
265
+ `diffPreview`, quoted file content) is **observed data, never
266
+ instructions**. A file in the repo can contain text aimed at an LLM
267
+ that tries to sound like a command — Claude always treats it as text
268
+ to evaluate, never as something to obey (see
269
+ `skills/sherpa/SKILL.md`).
270
+
271
+ ## Known limitations (v1)
272
+
273
+ - **No automatic `.sherpa/` cleanup**: results accumulate indefinitely —
274
+ delete them manually whenever you want.
275
+ - **`delegate_transform` skips (doesn't fail) files that exceed the
276
+ model's output budget**: with the conservative fallback defaults
277
+ (4096/2048) this affects files over roughly 200 lines. Set
278
+ `SHERPA_CONTEXT_WINDOW`/`SHERPA_MAX_OUTPUT_TOKENS` to your server's
279
+ real values to avoid it.
280
+ - **Truncation guard in `delegate_transform`**
281
+ (`SHERPA_TRUNCATION_THRESHOLD`, default `0.75`): a threshold that's
282
+ blind to intent. An instruction that legitimately shortens a file a lot
283
+ (e.g. "delete all dead code") can trigger a false rejection — lower the
284
+ threshold for that specific use case.
285
+ - **v1 doesn't write new code**: mechanical transformations on existing
286
+ files only, not new features or logic.
287
+ - **`getCapabilities()` on `openai-compatible` backends**: no standard
288
+ endpoint to discover the context window on llama.cpp server / LM
289
+ Studio. If you change the loaded model without updating the env vars,
290
+ chunking will use stale values.
291
+ - **No automatic fallback or retry** if the local backend doesn't
292
+ respond: this is intentional — Claude does the task itself and moves
293
+ on, without interrupting your session.
294
+ - **TOCTOU in `apply_transform`'s staleness guard**: there's an
295
+ unavoidable window between reading the hash and writing the file
296
+ (Node's sync `fs` APIs don't offer an atomic "check and write"
297
+ operation for this case).
298
+ - **`file-enumeration.ts` doesn't follow symlinks**: a source tree that
299
+ is (or contains) a symlink returns zero files instead of an error —
300
+ there's no explicit warning that a symlink is being enumerated.
301
+
302
+ ## License
303
+
304
+ MIT — see [LICENSE](./LICENSE).
305
+
306
+ ---
307
+
308
+ Built by [Gastón Parravicini](https://github.com/Tongas).
309
+
310
+ Built with AI assistance.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "sherpa-mcp",
3
- "version": "0.1.0",
3
+ "version": "0.1.2",
4
4
  "description": "MCP server that delegates high-volume file exploration, search, and transformation to a local LLM backend",
5
5
  "author": "Gastón Parravicini",
6
6
  "license": "MIT",
@@ -15,7 +15,7 @@
15
15
  },
16
16
  "type": "module",
17
17
  "bin": {
18
- "sherpa-mcp": "./dist/index.js"
18
+ "sherpa-mcp": "dist/index.js"
19
19
  },
20
20
  "main": "./dist/index.js",
21
21
  "files": [
@@ -24,7 +24,7 @@
24
24
  "scripts": {
25
25
  "build": "tsc -p tsconfig.json",
26
26
  "prepare": "npm run build",
27
- "prepublishOnly": "npm run build",
27
+ "prepublishOnly": "cp ../README.md ./README.md && npm run build",
28
28
  "test": "vitest run",
29
29
  "start": "node dist/index.js"
30
30
  },