hackshop-mcp 0.0.2 → 0.0.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (38) hide show
  1. package/README.md +63 -9
  2. package/SAMPLING.md +12 -6
  3. package/catalog.json +1877 -54
  4. package/dist/catalog/schema.d.ts +510 -0
  5. package/dist/catalog/schema.js +47 -0
  6. package/dist/catalog/schema.js.map +1 -1
  7. package/dist/platforms/index.d.ts +26 -0
  8. package/dist/platforms/index.js +51 -0
  9. package/dist/platforms/index.js.map +1 -0
  10. package/dist/platforms/load.d.ts +4 -0
  11. package/dist/platforms/load.js +78 -0
  12. package/dist/platforms/load.js.map +1 -0
  13. package/dist/platforms/schema.d.ts +383 -0
  14. package/dist/platforms/schema.js +53 -0
  15. package/dist/platforms/schema.js.map +1 -0
  16. package/dist/sampling.d.ts +4 -0
  17. package/dist/sampling.js +17 -19
  18. package/dist/sampling.js.map +1 -1
  19. package/dist/server.js +135 -23
  20. package/dist/server.js.map +1 -1
  21. package/dist/telemetry.d.ts +29 -0
  22. package/dist/telemetry.js +100 -0
  23. package/dist/telemetry.js.map +1 -0
  24. package/dist/tools/assess_hackability.d.ts +4 -0
  25. package/dist/tools/assess_hackability.js +44 -4
  26. package/dist/tools/assess_hackability.js.map +1 -1
  27. package/dist/tools/plan_gadget.d.ts +66 -0
  28. package/dist/tools/plan_gadget.js +406 -0
  29. package/dist/tools/plan_gadget.js.map +1 -0
  30. package/dist/tools/propose_hardware.d.ts +2 -0
  31. package/dist/tools/propose_hardware.js +38 -2
  32. package/dist/tools/propose_hardware.js.map +1 -1
  33. package/dist/tools/simulate_assembly.d.ts +315 -0
  34. package/dist/tools/simulate_assembly.js +145 -0
  35. package/dist/tools/simulate_assembly.js.map +1 -0
  36. package/package.json +4 -1
  37. package/platforms.json +448 -0
  38. package/tags.md +1 -0
package/README.md CHANGED
@@ -8,11 +8,13 @@ You describe a project. The agent surfaces 3-5 hackable hardware options you wou
8
8
 
9
9
  A tinkerer has an idea. The idea would be cooler with the right piece of hardware attached: an old screen, an abandoned smart speaker, a bricked frame, a hackable handheld. The tinkerer doesn't know what hardware exists, what's hackable, or what would creatively *fit* the idea. So the idea stays purely software, or gets paired with a Raspberry Pi.
10
10
 
11
- This is a hardware-knowledge layer on top of LLMs. Two tools, fifty hand-vetted devices, one closed-set tag vocabulary, and a brick-risk safety rule that won't let the agent fabricate a score for hardware classes where bricks are unrecoverable.
11
+ This is a hardware-knowledge layer on top of LLMs. Four tools, 80 researched devices, one closed-set tag vocabulary, and a brick-risk safety rule that won't let the agent fabricate a score for hardware classes where bricks are unrecoverable. `simulate_assembly` drops a proposed robot into a MuJoCo physics world and tells you, honestly, whether it would actually move.
12
+
13
+ Hackshop now knows about Meta's Muse Gadgets SDK: ESP32 boards and Linux machines that can become a physical body for Muse, Meta's personal AI agent. Muse recommendations include the SDK tier, setup path, supported features, printable stand/enclosure links when available, and the required terms caveat: personal, non-commercial use only, at most 50 devices per token, no selling or public marketplace listing, and revocable access.
12
14
 
13
15
  ## Status
14
16
 
15
- V0.0.2 — published on npm. Install with `npx hackshop-mcp` or add to your MCP client config.
17
+ v0.0.4 — published on npm. Install with `npx hackshop-mcp` or add it to your MCP client config. Four tools (`propose_hardware`, `assess_hackability`, `plan_gadget`, `simulate_assembly`), 80 devices. The simulation layer is live at [hackshop.dev](https://hackshop.dev).
16
18
 
17
19
  ## Install in 30 seconds
18
20
 
@@ -53,11 +55,54 @@ Returns 3-5 hardware proposals, each with:
53
55
 
54
56
  Lookup by id, exact name, or substring. Returns the same shape as a single proposal. Use when you have a device in mind and want to verify hackability before searching for one to buy.
55
57
 
58
+ ### `plan_gadget(idea, platform?, budget_usd?, owned_device_ids?, needs?, limit?)`
59
+
60
+ Deterministically plans a physical gadget for an AI agent, with Meta Muse Gadgets as the first supported platform. It infers needs such as voice, screen, camera, air sensors, e-paper, round display, home-network tunnel, or Linux control; ranks supported boards; and returns:
61
+
62
+ - `inferred_needs` and ranked `picks`
63
+ - each pick's Muse platform, support level, tier, score, concrete `why`, gaps, price label, firmware/build links, setup steps, and caveats
64
+ - `fabrication.printables` with STL/STEP/SVG/fab.json URLs when a stand or enclosure exists
65
+ - Muse SDK `terms` for every platform represented in the picks
66
+ - concrete `next_steps`
67
+
68
+ This tool does not call an LLM and does not use the network. It never suggests selling Muse devices; the Muse SDK token terms are personal and non-commercial.
69
+
70
+ ### `simulate_assembly(assembly)`
71
+
72
+ Takes an **Assembly IR** — `{ idea, components[{ref,device_id,name,role}], edges[], goal{kind,spec,success_metric}, world{template,goal_xy?} }` (build it from the site's assembly output or by hand) — drops it into a MuJoCo physics world, and runs a **bounded, synchronous** rollout (`duration_s` ≤ 10, default 8) on the sim-worker. It returns:
73
+
74
+ - `success` — did the rollout pass the typed position, collision, and upright acceptance criteria
75
+ - `summary` / `post_mortem` — natural-language verdict plus honest failure theatre (stuck / tipped / collisions / heading-oscillation)
76
+ - `artifacts` — hosted URLs for the rendered `video`, `scene` (MJCF), `control` (control.py), and `telemetry.json`
77
+ - `metric_value`, `telemetry`, `authored_by`, `world_desc`
78
+
79
+ Today it simulates the diff-drive **`navigate`** slice; other goal kinds return an honest `unsupported` rather than faking a pass. Set `SIM_WORKER_URL` to point at a running sim-worker (defaults to `http://127.0.0.1:8000`). The bounded rollout here is intentionally small so it fits in a single tool call; the **rich, longer, agent-driven runs happen via the web app** at [hackshop.dev](https://hackshop.dev), backed by the worker at [hackshop-sim.fly.dev](https://hackshop-sim.fly.dev).
80
+
81
+ ## Simulation (v2)
82
+
83
+ The site turns a proposal into a watchable robot: **proposal → select one complete build → deterministic feasibility check → MuJoCo rollout → interactive 3D replay**, with a shareable summary page you can link to. Honest by design — a robot that gets stuck on a ramp gets a post-mortem, not a green checkmark.
84
+
85
+ The current navigation slice deliberately has a narrow fidelity contract:
86
+
87
+ - alternative chassis are separate candidates, never merged into one BOM;
88
+ - Create 3 uses an explicit Pi + RPLIDAR build, while TurtleBot 4 Lite preserves its factory-integrated Pi/camera/lidar stack;
89
+ - versioned manifests provide real outer dimensions and mass to product-specific primitive proxies (not pretend CAD);
90
+ - the controller only runs when the assembly declares the 2D-lidar observations it consumes;
91
+ - typed position/collision/upright criteria drive the verdict; and
92
+ - the browser replay supports orbit, zoom, playback/scrubbing, world geometry, path/goal overlays, and collision/failure markers.
93
+
94
+ - **Live:** [https://hackshop.dev](https://hackshop.dev)
95
+ - **Worker:** [https://hackshop-sim.fly.dev](https://hackshop-sim.fly.dev)
96
+ - **Design doc:** [`docs/v2-simulation-plan.md`](docs/v2-simulation-plan.md)
97
+
98
+ The `simulate_assembly` MCP tool above is the bounded, single-call entry point into this same physics worker.
99
+
56
100
  ## Architecture
57
101
 
58
102
  - TypeScript + `@modelcontextprotocol/sdk`
59
- - LLM reasoning delegated to the host via `sampling/createMessage` (no Anthropic SDK bundled, no BYO key)
60
- - Catalog stored as `catalog.json` in the repo (JSON, version-controllable, ~50 devices in V1)
103
+ - LLM reasoning delegated to the host via `sampling/createMessage` first; falls back to a direct Anthropic API call (`@anthropic-ai/sdk`) when `ANTHROPIC_API_KEY` is set and the host lacks sampling
104
+ - `simulate_assembly` calls out to a separate Python MuJoCo **sim-worker** over HTTP (`SIM_WORKER_URL`); the worker isn't bundled in the npm package
105
+ - Catalog stored as `catalog.json` in the repo (JSON, version-controllable, 80 devices and growing)
61
106
  - Tag vocabulary in `tags.md`, validated at boot — server refuses to start on tag drift
62
107
  - eBay integration is **not** in this server. Compose with [`ebay-mcp`](https://github.com/YosefHayim/ebay-mcp) at the host level.
63
108
 
@@ -72,9 +117,9 @@ npm test # safety + schema + lookup tests
72
117
  npm run build # tsc -> dist/
73
118
  ```
74
119
 
75
- ## Day-0 Smoke Test (REQUIRED before scaffolding more)
120
+ ## Troubleshooting: verify sampling support
76
121
 
77
- This server depends on `sampling/createMessage`. Some MCP hosts don't support it. Verify yours does first.
122
+ `propose_hardware` reasons via `sampling/createMessage`. Some MCP hosts don't support it (and without an `ANTHROPIC_API_KEY` fallback you'll get degraded, raw-catalog responses). If proposals come back without reasoning, verify your host supports sampling with this quick smoke check.
78
123
 
79
124
  ```bash
80
125
  npm install
@@ -93,11 +138,11 @@ Add this to your Claude Desktop config (`~/Library/Application Support/Claude/cl
93
138
  }
94
139
  ```
95
140
 
96
- Restart Claude Desktop. Ask Claude to call the `smoke_check` tool. If it returns "Smoke OK," your architecture works. If it fails, stop here — `hackshop-mcp` won't work in this host.
141
+ Restart Claude Desktop. Ask Claude to call the `smoke_check` tool. If it returns "Smoke OK," sampling works in your host. If it fails, set an `ANTHROPIC_API_KEY` (see install above) or expect degraded responses from `propose_hardware`.
97
142
 
98
- ## Install (target host)
143
+ ## Install (local build → host)
99
144
 
100
- After dev is done and smoke passes, add to your MCP client config:
145
+ To run a locally built copy instead of `npx`, add to your MCP client config:
101
146
 
102
147
  ```json
103
148
  {
@@ -120,6 +165,15 @@ The founder had an Electric Objects EO1 picture frame. The company shut down; th
120
165
 
121
166
  Bricking unrecoverable hardware is the single failure mode that ends this product. The catalog tracks brick-risk provenance: `founder-verified | community-reported | llm-inferred`. For categories where bricks are unrecoverable (`handheld`, `sbc`), the server **refuses to surface LLM-inferred brick-risk scores**. It returns "brick-risk unknown — research before flashing" instead. This is a tested release gate. See `src/safety.ts` and `test/safety.test.ts`.
122
167
 
168
+ ## Telemetry
169
+
170
+ Starting with v0.0.4 the MCP server sends an anonymous ping when it starts and after each tool call. It exists so the maintainer can tell whether anyone is actually using the server.
171
+
172
+ - **Sent:** the event name, tool name, success/degraded flag, duration, `hackshop-mcp` version, MCP client name/version (e.g. `claude-code`), OS platform, Node major version, and a random install id stored in `~/.config/hackshop-mcp/telemetry.json`.
173
+ - **Never sent:** tool arguments, your idea text, device names, results, API keys, or file paths.
174
+ - **Destination:** `https://www.hackshop.dev/api/telemetry/mcp`, which forwards to the project's PostHog.
175
+ - **Opt out:** set `HACKSHOP_TELEMETRY=0` (or `DO_NOT_TRACK=1`) in the server's `env`. Telemetry is also off automatically in CI and under test runners. The implementation is `src/telemetry.ts`.
176
+
123
177
  ## Contributing
124
178
 
125
179
  See `CONTRIBUTING.md`. New devices come in via PR; tag changes require a `tags.md` edit; `community-reported` is the default provenance for community contributions.
package/SAMPLING.md CHANGED
@@ -74,11 +74,17 @@ JSON only, no markdown. Code-fence allowed (the parser strips it).
74
74
 
75
75
  The retry is in `sampleJson` itself (`src/sampling.ts`).
76
76
 
77
- ### What the server NEVER does
77
+ ### What the server does (v0.0.2 contract)
78
78
 
79
- - Bundle an Anthropic / OpenAI / other-provider SDK.
80
- - Read a user-side `ANTHROPIC_API_KEY` env var.
81
- - Make HTTP calls to provider endpoints directly.
82
- - Crash on sampling failure.
79
+ - **Always tries host sampling first** via `sampling/createMessage`. No key needed; host pays.
80
+ - **Reads `ANTHROPIC_API_KEY` only as a fallback** when host sampling fails twice. Optional: if the env var is unset, the server returns a degraded response instead of calling Anthropic directly.
81
+ - **Bundles `@anthropic-ai/sdk`** as a runtime dependency for the fallback path. (V0.0.1 promised "no SDK" — V0.0.2 chose pragmatism after live testing showed Claude Desktop's sampling support is spotty.)
82
+ - **Never crashes on sampling failure** — degraded response is the floor.
83
+ - **Never makes provider calls when `ANTHROPIC_API_KEY` is unset** — keyless installs stay keyless.
83
84
 
84
- If a future feature seems to need a direct provider call, treat that as a redesign signal — not a quick patch.
85
+ ### What the server still NEVER does
86
+
87
+ - Bundle an OpenAI / other non-Anthropic SDK.
88
+ - Hard-require `ANTHROPIC_API_KEY` (it's optional, host sampling is the preferred path).
89
+ - Persist user prompts or proposals.
90
+ - Make outbound calls beyond the Anthropic API + the user's MCP host.