mapsmith 0.2.0__tar.gz → 0.2.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (61) hide show
  1. {mapsmith-0.2.0 → mapsmith-0.2.1}/.gitignore +6 -0
  2. mapsmith-0.2.1/CHANGELOG.md +141 -0
  3. {mapsmith-0.2.0 → mapsmith-0.2.1}/CLAUDE.md +4 -0
  4. {mapsmith-0.2.0 → mapsmith-0.2.1}/MANIFESTO.md +13 -11
  5. mapsmith-0.2.1/PKG-INFO +410 -0
  6. mapsmith-0.2.1/README.md +374 -0
  7. mapsmith-0.2.1/SECURITY.md +83 -0
  8. {mapsmith-0.2.0 → mapsmith-0.2.1}/docker-compose.yml +7 -3
  9. {mapsmith-0.2.0 → mapsmith-0.2.1}/docs/benchmarks.md +10 -2
  10. {mapsmith-0.2.0 → mapsmith-0.2.1}/examples/03_validated_plans.ipynb +24 -24
  11. {mapsmith-0.2.0 → mapsmith-0.2.1}/pyproject.toml +4 -1
  12. {mapsmith-0.2.0 → mapsmith-0.2.1}/server.json +3 -3
  13. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/__init__.py +1 -1
  14. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/engines/dispatch.py +2 -1
  15. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/engines/duckdb_engine.py +70 -22
  16. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/engines/raster.py +2 -2
  17. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/engines/vector.py +14 -14
  18. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/plans/executor.py +1 -0
  19. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/preview.py +291 -272
  20. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/ui.py +385 -367
  21. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/verify.py +63 -6
  22. {mapsmith-0.2.0 → mapsmith-0.2.1}/tests/test_duckdb_sandbox.py +124 -3
  23. mapsmith-0.2.1/tests/test_showcase.py +303 -0
  24. mapsmith-0.2.1/tests/test_verification_status.py +135 -0
  25. {mapsmith-0.2.0 → mapsmith-0.2.1}/tests/test_verify_repair.py +30 -0
  26. {mapsmith-0.2.0 → mapsmith-0.2.1}/tests/test_whitebox_encoding.py +7 -2
  27. mapsmith-0.2.0/CHANGELOG.md +0 -77
  28. mapsmith-0.2.0/PKG-INFO +0 -288
  29. mapsmith-0.2.0/README.md +0 -252
  30. mapsmith-0.2.0/SECURITY.md +0 -53
  31. {mapsmith-0.2.0 → mapsmith-0.2.1}/CONTRIBUTING.md +0 -0
  32. {mapsmith-0.2.0 → mapsmith-0.2.1}/Dockerfile +0 -0
  33. {mapsmith-0.2.0 → mapsmith-0.2.1}/LICENSE +0 -0
  34. {mapsmith-0.2.0 → mapsmith-0.2.1}/TRADEMARKS.md +0 -0
  35. {mapsmith-0.2.0 → mapsmith-0.2.1}/examples/01_verified_geoprocessing.ipynb +0 -0
  36. {mapsmith-0.2.0 → mapsmith-0.2.1}/examples/02_terrain_hydrology.ipynb +0 -0
  37. {mapsmith-0.2.0 → mapsmith-0.2.1}/examples/README.md +0 -0
  38. {mapsmith-0.2.0 → mapsmith-0.2.1}/examples/fixtures/mount_st_helens_dem.tif +0 -0
  39. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/catalog.py +0 -0
  40. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/engines/__init__.py +0 -0
  41. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/engines/sedona_engine.py +0 -0
  42. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/engines/whitebox_engine.py +0 -0
  43. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/jobs.py +0 -0
  44. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/plans/__init__.py +0 -0
  45. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/plans/models.py +0 -0
  46. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/plans/registry.py +0 -0
  47. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/plans/validator.py +0 -0
  48. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/provenance.py +0 -0
  49. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/server.py +0 -0
  50. {mapsmith-0.2.0 → mapsmith-0.2.1}/src/mapsmith/workspace.py +0 -0
  51. {mapsmith-0.2.0 → mapsmith-0.2.1}/tests/test_catalog.py +0 -0
  52. {mapsmith-0.2.0 → mapsmith-0.2.1}/tests/test_catalog_retrieval.py +0 -0
  53. {mapsmith-0.2.0 → mapsmith-0.2.1}/tests/test_engines.py +0 -0
  54. {mapsmith-0.2.0 → mapsmith-0.2.1}/tests/test_plans.py +0 -0
  55. {mapsmith-0.2.0 → mapsmith-0.2.1}/tests/test_preview.py +0 -0
  56. {mapsmith-0.2.0 → mapsmith-0.2.1}/tests/test_raster.py +0 -0
  57. {mapsmith-0.2.0 → mapsmith-0.2.1}/tests/test_smoke.py +0 -0
  58. {mapsmith-0.2.0 → mapsmith-0.2.1}/tests/test_ui.py +0 -0
  59. {mapsmith-0.2.0 → mapsmith-0.2.1}/tests/test_verify.py +0 -0
  60. {mapsmith-0.2.0 → mapsmith-0.2.1}/tests/test_whitebox.py +0 -0
  61. {mapsmith-0.2.0 → mapsmith-0.2.1}/tests/test_workspace.py +0 -0
@@ -22,6 +22,12 @@ venv/
22
22
  *.provenance.json
23
23
  data/
24
24
 
25
+ # Scratch databases: probing the DuckDB sandbox (ATTACH, COPY) drops these in
26
+ # the repo root, and one of them reached a public commit before we noticed.
27
+ *.db
28
+ *.duckdb
29
+ *.wal
30
+
25
31
  # OS / editors
26
32
  .DS_Store
27
33
  Thumbs.db
@@ -0,0 +1,141 @@
1
+ # Changelog
2
+
3
+ All notable changes to MapSmith are documented here, in the format of
4
+ [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). This project follows
5
+ [semantic versioning](https://semver.org/).
6
+
7
+ ## [0.2.1] — 2026-08-20
8
+
9
+ Three fixes you would rather not find yourself. All came from reviewing 0.2.0
10
+ *after* it shipped, and all were reproduced through the real MCP tools instead
11
+ of read off a diff.
12
+
13
+ ### Fixed
14
+
15
+ - **An empty spatial join no longer crashes before writing its manifest.**
16
+ DuckDB writes no GeoParquet `geo` metadata for a zero-row result, so reading
17
+ the output back raised — on the default engine path, for exactly the case the
18
+ verification checks exist to explain, while the tool description promised a
19
+ warning. A zero-row result is now written as a valid empty GeoParquet with the
20
+ analysis CRS and the joined schema, and the join goes through the same audited
21
+ writer as everything else, so the manifest exists even when a check fails.
22
+ - **A GeoParquet declaring `crs: null` is no longer read as CRS84.** MapSmith
23
+ invented a coordinate system and recorded it in the manifest as fact, with
24
+ `verified: true` — the worst class of bug a provenance tool can have. The
25
+ GeoParquet spec distinguishes an *absent* `crs` field, which does mean CRS84,
26
+ from an explicit null, which means unknown; so does MapSmith now, and an
27
+ unknown CRS is refused by the preconditions like any other missing CRS.
28
+ - **The DuckDB sandbox locks its configuration in every mode.** The lock used to
29
+ apply only under `MAPSMITH_WORKSPACE`, so a multi-statement call could switch
30
+ extension autoloading back on, and an explicit `LOAD httpfs` was never blocked
31
+ at all. Locking is now unconditional and DuckDB's HTTP and S3 filesystems are
32
+ disabled, while local reads keep working.
33
+ - **The security documentation no longer claims that unconfined mode blocks
34
+ network egress.** It does not, and it never did: GDAL carries its own HTTP
35
+ client, so `ST_Read('/vsicurl/https://…')` in raw SQL reads whatever the host
36
+ can reach — internal services and metadata endpoints included — and the URL it
37
+ names can carry data out. Remote reads are deliberately available while the
38
+ server is unconfined, because cloud-native data is a feature; the claim that
39
+ the network was closed anyway was the bug. README and SECURITY.md now state
40
+ the price of that choice, and two tests pin both halves of it: without a
41
+ workspace the read succeeds, and with one it is refused *before any request
42
+ leaves*, asserted by counting requests at a loopback server instead of
43
+ matching an error message. If you do not trust your `run_sql` input, set a
44
+ workspace.
45
+ - **The multi-layer guard fails closed.** A container whose layer list could not
46
+ be read looked like a single-layer file, and mechanical geometry repair would
47
+ then have destroyed the other layers while recording success.
48
+ - **`execute_plan` reports `repairs` per step.** Geometry MapSmith rewrote was
49
+ visible in a single-operation result and invisible at plan level.
50
+ - **Manifests record `EPSG:32632`, not 2.5 KB of PROJJSON**, when a GeoParquet
51
+ input carries its CRS as an embedded projection object.
52
+ - **The example `docker-compose.yml` binds MinIO to loopback.** It published
53
+ ports 9000/9001 on every interface with the documented development
54
+ credentials.
55
+
56
+ ### Changed
57
+
58
+ - **The provenance badge has three states instead of two.** "All critical checks
59
+ passed" is vacuously true when no critical check ran — a `run_sql` manifest,
60
+ for one — so an output whose only check had *failed* rendered with the same
61
+ green tick as a verified buffer. `provenance_summary` now reports `verified`,
62
+ `failed` or `unchecked`, with the reason, and the map panel renders all three.
63
+ The `verified` boolean stays in the payload, computed correctly, so a client
64
+ reading it gets a fix rather than a breaking change.
65
+ - **The README says when *not* to use MapSmith**, and the
66
+ [benchmark results](docs/benchmarks.md) are linked from it — they were public
67
+ for a day with nothing pointing at them.
68
+
69
+ ## [0.2.0] — 2026-08-20
70
+
71
+ The first release you can point an agent at and trust the answer: results are
72
+ verified on the way in and on the way out, plans are checked before anything
73
+ runs, and the server can be confined to a single directory.
74
+
75
+ ### Added
76
+
77
+ - **Interactive map inside the chat.** `preview_map` renders your layers on a
78
+ pan/zoom map panel in any client implementing the
79
+ [MCP Apps](https://modelcontextprotocol.io/extensions/apps/overview)
80
+ extension (field-tested on Claude Desktop), with an OpenStreetMap backdrop
81
+ and a provenance card per layer showing operation, engine and one of three
82
+ states: `verified ✓`, `verification failed`, or `not verifiable` when no
83
+ critical check ran. Fully self-contained; on clients without MCP Apps the
84
+ same call returns structured data.
85
+ - **Typed plans.** `validate_plan` statically checks a multi-step analysis —
86
+ operations exist and are installed, arguments complete and well-typed,
87
+ `$step` references resolve backwards, input files exist, outputs don't
88
+ collide, and the CRS of every intermediate is simulated from the real
89
+ inputs — and returns machine-actionable error codes. `execute_plan` then
90
+ runs the validated plan with per-step provenance plus a plan-level manifest
91
+ fingerprinting the exact plan that produced the result.
92
+ - **Terrain and hydrology** on the Whitebox Workflows engine (`[whitebox]`
93
+ extra): `hillshade`, `flow_accumulation` (D8, with depression filling) and
94
+ `watershed` (many pour points at once).
95
+ - **Zonal statistics** with exact fractional pixel coverage via exactextract
96
+ (`[raster]` extra).
97
+ - **A searchable operation catalog.** `list_operations` ranks capabilities by
98
+ relevance (BM25) so an agent can discover what exists — including what is
99
+ planned but not yet available — instead of guessing from a wall of tools.
100
+ - **Workspace confinement.** Set `MAPSMITH_WORKSPACE` and every path argument
101
+ must resolve inside it, `run_sql`'s DuckDB connection is sandboxed to that
102
+ directory with extension loading refused and memory/temp-disk capped, and
103
+ UNC hosts and NTFS alternate data streams are refused in every mode.
104
+ - **Verification on the way in.** Operations check their inputs for the
105
+ failures that produce plausible junk: a missing CRS is refused outright, and
106
+ empty inputs or extents that cannot possibly intersect come back as named
107
+ warnings with hints — in the tool result, not only in the manifest.
108
+ - **Bounded deterministic repair.** Mechanically broken output geometry is
109
+ repaired (`make_valid`, at most two rounds, written atomically) and every
110
+ attempt is recorded in the manifest: a repaired output never looks like one
111
+ that was right the first time.
112
+ - **A notebook gallery** (`examples/`) and a
113
+ [benchmarks page](docs/benchmarks.md) with the harness that produced it.
114
+
115
+ ### Fixed
116
+
117
+ - **Wrong terrain results from ordinary compressed rasters.** Whitebox
118
+ Workflows 2.x does not undo the TIFF predictor when reading, so any DEM
119
+ saved with `PREDICTOR=2` or `3` — the standard encoding for elevation data —
120
+ produced hillshades and flow accumulations that looked like terrain and were
121
+ not. MapSmith now detects the predictor and converts the input first,
122
+ recording it in the manifest.
123
+ ([upstream report](https://github.com/jblindsay/whitebox_next_gen/issues/32))
124
+ - **GeoParquet outputs from the vector engines** were written through a GDAL
125
+ path that produced unreadable files.
126
+ - `run_sql` materialisations and the DuckDB/SedonaDB join fast paths now run
127
+ the same deterministic verification as every other writer, and record their
128
+ CRS decisions.
129
+
130
+ ### Changed
131
+
132
+ - `spatial_join` with `engine="auto"` falls back to GeoPandas when the inputs'
133
+ CRS differ or are unknown, instead of joining mismatched coordinates.
134
+ - The planning-failure figure quoted in the docs is stated as the upper bound
135
+ it is ("up to ~47%"), since the underlying study counts errors multi-label.
136
+
137
+ ## [0.1.0] — 2026-08-18
138
+
139
+ First public release: the engine dispatcher (SedonaDB / DuckDB / GeoPandas),
140
+ `run_sql`, the job ledger, stateless Streamable HTTP transport, and provenance
141
+ manifests with deterministic verification on every writer.
@@ -19,6 +19,10 @@ MapSmith gives AI agents professional-grade geoprocessing via MCP, with **verifi
19
19
  - Version pins are deliberate: `mcp>=1.26,<2` (1.26 is the floor for resource `meta`, needed by the MCP Apps map panel; v2 renamed FastMCP→MCPServer; migration planned with MCP Tasks), `ruff>=0.16,<0.17` (new ruff minors add default rules and break CI).
20
20
  - Tests use closed-form expected values (e.g., a known 5×5 raster block → mean=22, sum=550) plus rejection-path tests; `pytest.importorskip` for extras.
21
21
  - Run `python -m ruff check .` before committing; CI runs lint + tests on Python 3.10/3.12 + Docker build.
22
+ - **The floor is Python 3.10**, and local runs happen on a newer interpreter, so stdlib added
23
+ after 3.10 breaks only in CI: no `tomllib`, `datetime.UTC`, `StrEnum`, `contextlib.chdir`,
24
+ `ExceptionGroup` (3.11), no `itertools.batched`, `typing.override` (3.12). Ruff's
25
+ `target-version` catches too-new *syntax*; too-new *modules* are on you.
22
26
  - Docker (or `uvx` where wheels work) is the only supported install path — keep it true in docs.
23
27
  - Verify external-library APIs against primary documentation before coding against them; this repo has already been bitten by from-memory APIs three times.
24
28
 
@@ -12,7 +12,7 @@ name under. These are the commitments that follow from taking that seriously.
12
12
  ## 1. Engines compute. Models orchestrate.
13
13
 
14
14
  Every geometry and every number in a MapSmith result comes from executing a
15
- deterministic engine — GDAL, GeoPandas, DuckDB Spatial, WhiteboxTools,
15
+ deterministic engine — GDAL, GeoPandas, DuckDB Spatial, Whitebox Workflows,
16
16
  SedonaDB. None
17
17
  of it is generated by a model, ever. The model's job is to decide *what* to
18
18
  run; it is never asked to produce a coordinate. This is the line that
@@ -38,8 +38,9 @@ that passed.
38
38
 
39
39
  ## 4. Coordinate systems are stated, never assumed.
40
40
 
41
- Silent CRS mismatch is the largest single cause of confidently wrong
42
- geospatial analysis. MapSmith refuses inputs without a CRS, never runs metric
41
+ A silent CRS mismatch is the classic way a geospatial analysis comes out
42
+ confidently wrong: the arithmetic succeeds, the map looks plausible, and the
43
+ numbers mean nothing. MapSmith refuses inputs without a CRS, never runs metric
43
44
  operations on degrees without a declared reprojection, and writes the reason
44
45
  for every transformation into the manifest.
45
46
 
@@ -71,14 +72,15 @@ engine actually ran goes in the manifest.
71
72
  ## 7. Your data stays on your machine.
72
73
 
73
74
  No telemetry, no phoning home, no dataset uploads: MapSmith runs beside your
74
- data — on a laptop, in a container, on your own infrastructure. Two outbound
75
- requests exist and we would rather name them than have you find them: the
76
- in-chat map panel fetches OpenStreetMap background tiles (which reveals the
77
- map view you are looking at, and falls back to a plain background when the
78
- host blocks it), and the SQL engine downloads its spatial extension once per
79
- environment. Neither carries your datasets, and both are avoidable — skip the
80
- panel, pre-install the extension. The
81
- provenance manifests it produces belong to you, and they happen to be exactly
75
+ data — on a laptop, in a container, on your own infrastructure. MapSmith itself
76
+ makes two outbound requests, and we would rather name them than have you find
77
+ them: the in-chat map panel fetches OpenStreetMap background tiles (which
78
+ reveals the map view you are looking at, and falls back to a plain background
79
+ when the host blocks it), and the SQL engine downloads its spatial extension
80
+ once per environment. Neither carries your datasets, and both are avoidable —
81
+ skip the panel, pre-install the extension. What an *agent* can reach through
82
+ the SQL engine is a separate question, answered precisely in `SECURITY.md`.
83
+ The provenance manifests MapSmith produces belong to you, and they happen to be exactly
82
84
  the kind of record an audit asks for: what ran, on what inputs, with what
83
85
  parameters, verified how. The EU AI Act's traceability obligations for
84
86
  high-risk systems are not a burden if your tooling produces that trail as a
@@ -0,0 +1,410 @@
1
+ Metadata-Version: 2.5
2
+ Name: mapsmith
3
+ Version: 0.2.1
4
+ Summary: Professional-grade geoprocessing for AI agents via MCP, with verifiable provenance
5
+ Project-URL: Homepage, https://github.com/mapsmith-ai/MapSmith
6
+ Project-URL: Repository, https://github.com/mapsmith-ai/MapSmith
7
+ Author-email: MapSmith <mapsmith@proton.me>
8
+ License-Expression: AGPL-3.0-or-later
9
+ License-File: LICENSE
10
+ Keywords: ai-agents,geoprocessing,geospatial,gis,mcp,provenance
11
+ Classifier: Development Status :: 3 - Alpha
12
+ Classifier: Intended Audience :: Science/Research
13
+ Classifier: Programming Language :: Python :: 3
14
+ Classifier: Topic :: Scientific/Engineering :: GIS
15
+ Requires-Python: >=3.10
16
+ Requires-Dist: duckdb>=1.4
17
+ Requires-Dist: geopandas>=1.1.2
18
+ Requires-Dist: mcp<2,>=1.26
19
+ Requires-Dist: pyarrow>=17
20
+ Requires-Dist: pyogrio>=0.9
21
+ Requires-Dist: pyproj>=3.6
22
+ Requires-Dist: shapely>=2.0
23
+ Provides-Extra: postgres
24
+ Requires-Dist: psycopg[binary]>=3.2; extra == 'postgres'
25
+ Provides-Extra: raster
26
+ Requires-Dist: exactextract>=0.2; extra == 'raster'
27
+ Requires-Dist: rasterio>=1.3; extra == 'raster'
28
+ Provides-Extra: sedona
29
+ Requires-Dist: apache-sedona[db]>=1.9; extra == 'sedona'
30
+ Provides-Extra: test
31
+ Requires-Dist: pytest>=8.0; extra == 'test'
32
+ Requires-Dist: ruff<0.17,>=0.16; extra == 'test'
33
+ Provides-Extra: whitebox
34
+ Requires-Dist: whitebox-workflows<3,>=2.0.6; extra == 'whitebox'
35
+ Description-Content-Type: text/markdown
36
+
37
+ # MapSmith
38
+
39
+ [![CI](https://github.com/mapsmith-ai/MapSmith/actions/workflows/ci.yml/badge.svg)](https://github.com/mapsmith-ai/MapSmith/actions/workflows/ci.yml)
40
+ [![PyPI](https://img.shields.io/pypi/v/mapsmith)](https://pypi.org/project/mapsmith/)
41
+ [![Container](https://img.shields.io/badge/ghcr.io-mapsmith--ai%2Fmapsmith-2496ED?logo=docker&logoColor=white)](https://github.com/mapsmith-ai/MapSmith/pkgs/container/mapsmith)
42
+ [![MCP](https://img.shields.io/badge/Model_Context_Protocol-server-654FF0)](https://modelcontextprotocol.io)
43
+ [![License: AGPL-3.0](https://img.shields.io/badge/license-AGPL--3.0-blue)](LICENSE)
44
+
45
+ **Professional-grade geoprocessing for AI agents — with provenance you can verify.**
46
+
47
+ MapSmith is an open-source [MCP](https://modelcontextprotocol.io) server that gives an AI
48
+ agent real GIS analysis — buffers, overlays, reprojections, zonal statistics, terrain and
49
+ hydrology — executed by GeoPandas, DuckDB Spatial, exactextract and Whitebox Workflows,
50
+ never written by the model. Every dataset it produces lands on disk next to a lineage
51
+ manifest: inputs with checksums, the exact parameters, the CRS decisions and *why*, engine
52
+ versions, and the deterministic checks that ran on the result.
53
+
54
+ > Ask for the result. The agent picks the tools. You can check the work afterwards.
55
+
56
+ Evidence before promises: an [A/B on GABench](docs/benchmarks.md) whose headline is a null
57
+ result — with the analysis that took our own positive number apart — [notebooks](examples/)
58
+ on a real USGS DEM of Mount St. Helens, and an
59
+ [in-chat map panel](#see-results-inside-the-chat) that shows the verification status of
60
+ every layer it draws.
61
+
62
+ ## Quickstart
63
+
64
+ Add MapSmith to any MCP client over stdio (Claude Desktop, Claude Code, Cursor, VS Code):
65
+
66
+ ```json
67
+ {
68
+ "mcpServers": {
69
+ "mapsmith": {
70
+ "command": "uvx",
71
+ "args": ["mapsmith"]
72
+ }
73
+ }
74
+ }
75
+ ```
76
+
77
+ Docker is the supported path, and confines the server to the directory you mount:
78
+
79
+ ```json
80
+ {
81
+ "mcpServers": {
82
+ "mapsmith": {
83
+ "command": "docker",
84
+ "args": ["run", "-i", "--rm",
85
+ "-v", "/absolute/path/to/your/data:/data",
86
+ "-e", "MAPSMITH_WORKSPACE=/data",
87
+ "ghcr.io/mapsmith-ai/mapsmith"]
88
+ }
89
+ }
90
+ }
91
+ ```
92
+
93
+ One-click installs:
94
+
95
+ [![Install in Cursor](https://cursor.com/deeplink/mcp-install-dark.svg)](https://cursor.com/install-mcp?name=mapsmith&config=eyJjb21tYW5kIjoidXZ4IiwiYXJncyI6WyJtYXBzbWl0aCJdfQ%3D%3D)
96
+ [![Install in VS Code](https://img.shields.io/badge/VS_Code-Install_MapSmith-0098FF?logo=visualstudiocode&logoColor=white)](https://insiders.vscode.dev/redirect/mcp/install?name=mapsmith&config=%7B%22command%22%3A%22uvx%22%2C%22args%22%3A%5B%22mapsmith%22%5D%7D)
97
+
98
+ or from a terminal: `code --add-mcp '{"name":"mapsmith","command":"uvx","args":["mapsmith"]}'`
99
+
100
+ To check it runs before wiring a client, `uvx mapsmith` starts the server on stdio
101
+ (Ctrl-C to quit) — it speaks MCP, not a CLI, so a silent prompt means it is working.
102
+
103
+ Then ask your agent things like:
104
+
105
+ > "Take parcels.gpkg, keep only the parcels within 300 m of the river in rivers.gpkg, and
106
+ > give me the result with the analysis lineage."
107
+
108
+ The Docker image includes the `[raster]` and `[whitebox]` extras. With `uvx`, pick your
109
+ own: `uvx --from "mapsmith[raster,whitebox]" mapsmith`. **Docker — or `uvx` on a machine
110
+ with working wheels — is the only supported installation path**: geospatial native
111
+ dependencies across three OSes are a support black hole, and issues about broken local
112
+ environments will be redirected here.
113
+
114
+ ## What you get back
115
+
116
+ Every dataset comes with the file below, written next to it as
117
+ `<output>.provenance.json` — enough to re-run the analysis without the model that asked
118
+ for it:
119
+
120
+ ```json
121
+ {
122
+ "mapsmith_version": "0.2.1",
123
+ "operation": "buffer_layer",
124
+ "parameters": {"distance_meters": 300.0},
125
+ "inputs": [{"path": "rivers.gpkg", "sha256": "9f2c…", "crs": "EPSG:4326"}],
126
+ "crs_decisions": {"analysis_crs": "EPSG:32632", "reason": "estimated UTM zone for metric buffering"},
127
+ "engine": {"name": "geopandas", "version": "1.0.1"},
128
+ "started_at": "2026-08-18T10:15:03Z",
129
+ "finished_at": "2026-08-18T10:15:04Z"
130
+ }
131
+ ```
132
+
133
+ The full manifest also carries the verification checks that ran, and any geometry MapSmith
134
+ had to repair. `get_provenance` returns it for any output.
135
+
136
+ ## Why MapSmith
137
+
138
+ - **Real geoprocessing, not map CRUD.** Built on the proven open geospatial stack: GDAL,
139
+ GeoPandas, Shapely, DuckDB Spatial, Whitebox Workflows and exactextract ship today
140
+ (more to come: PDAL, QGIS Processing via sidecar).
141
+ - **Provenance by design.** Every layer MapSmith produces ships with a machine-readable
142
+ lineage manifest — source datasets with checksums, tools executed, exact parameters, CRS
143
+ decisions, software versions, timestamps. Everything needed to re-run the analysis
144
+ without the LLM is in there. No AI slop.
145
+ - **The engines compute, the model orchestrates.** Geometry and numbers only ever come
146
+ from deterministic tool executions — never from model output.
147
+ - **Semantic tools, not a tool dump.** 16 goal-level tools plus a searchable operation
148
+ catalog (progressive discovery), because agent accuracy collapses when you expose
149
+ hundreds of raw tools.
150
+ - **Model-agnostic infrastructure.** Claude, GPT, Qwen, Kimi, GLM — anything that speaks
151
+ MCP, cloud or local. The leverage is better contracts (typed plans, actionable error
152
+ codes, a searchable catalog), not weights we would have to maintain. See
153
+ [the manifesto](MANIFESTO.md).
154
+
155
+ ## Tools
156
+
157
+ | Tool | What it does |
158
+ |---|---|
159
+ | `describe_dataset` | CRS, geometry types, schema, extent, feature count of any vector dataset |
160
+ | `buffer_layer` | Metric buffer with automatic UTM estimation for geographic CRS |
161
+ | `clip_layer` | Clip a layer with a mask layer |
162
+ | `reproject_layer` | Reproject to any CRS (EPSG code or WKT) |
163
+ | `spatial_join` | Join by spatial predicate, auto-routed to the fastest engine (SedonaDB > DuckDB > GeoPandas) |
164
+ | `run_sql` | Spatial SQL (DuckDB dialect) over GeoParquet and GDAL formats |
165
+ | `zonal_statistics` | Raster statistics per vector zone with exact fractional pixel coverage (`[raster]` extra) |
166
+ | `hillshade` | Shaded relief from a DEM, in-memory Whitebox engine (`[whitebox]` extra) |
167
+ | `flow_accumulation` | D8 flow accumulation with automatic depression filling (`[whitebox]` extra) |
168
+ | `watershed` | Watershed delineation from a DEM and pour points (`[whitebox]` extra) |
169
+ | `preview_map` | Interactive in-chat map (MCP Apps) of any datasets, with a provenance card and verification status per layer |
170
+ | `validate_plan` | Statically validate a multi-step plan before running anything: operations, arguments, references, input files, simulated CRS flow |
171
+ | `execute_plan` | Validate then run a plan step by step, with per-step provenance and a plan-level manifest |
172
+ | `get_provenance` | Return the full lineage manifest of any MapSmith output |
173
+ | `list_operations` | BM25-ranked catalog search; `detail=true` returns parameters and worked examples |
174
+ | `server_info` | Version, license, available engines |
175
+
176
+ ## Verification, in and out
177
+
178
+ Every tool that writes a dataset also writes `<output>.provenance.json` beside it and
179
+ verifies its own work — CRS agreement, geometry validity, raster dimensions, count and
180
+ extent invariants — recording the results in the manifest *before* raising anything, so
181
+ the audit trail survives the error.
182
+
183
+ Verification runs on the way in as well. Before an operation touches your data, MapSmith
184
+ checks the failures that produce *plausible* junk: an input with no CRS is refused
185
+ outright, because metric maths on unknown units is how a confidently wrong answer gets
186
+ made; an empty input, or two layers whose extents cannot possibly overlap, comes back as
187
+ a named warning with a hint — in the tool result, not only in the manifest, so the agent
188
+ sees it instead of assuming success. (The join fast paths, DuckDB and SedonaDB, only ever
189
+ receive inputs that already share a known CRS; they verify their output and diagnose an
190
+ empty join.)
191
+
192
+ An output whose geometry is *mechanically* broken — typically invalidity inherited from an
193
+ invalid input — is repaired deterministically: `make_valid`, at most two rounds, written
194
+ to a temporary file and swapped in only once it is complete, and skipped rather than
195
+ risked where a rewrite could drop data (a multi-layer GeoPackage is refused, not
196
+ rewritten). Every attempt lands in the manifest *and* in the tool result, because a
197
+ repaired output must never look like one that was right the first time. Failures that need
198
+ judgement are never "fixed": an empty result, or geometries eroded away by a wrong
199
+ distance, come back as warnings with hints for the agent to act on.
200
+
201
+ ## See results inside the chat
202
+
203
+ ![MapSmith's interactive map panel rendered inside Claude Desktop: OSM basemap, buffer and zone layers, and per-layer provenance cards with verification status](docs/images/map-panel.png)
204
+
205
+ `preview_map` renders your layers on an interactive map panel *inside* the chat — pan,
206
+ zoom, toggle layers, and read each layer's provenance card (operation, engine, and one of
207
+ three honest states: `verified ✓`, `verification failed`, or `not verifiable` when no
208
+ critical check ran) right next to the geometry it explains. Field-tested on Claude
209
+ Desktop; it renders in any client that implements the official
210
+ [MCP Apps](https://modelcontextprotocol.io/extensions/apps/overview) extension, and on
211
+ clients without it the same call returns the preview as structured data.
212
+
213
+ The panel is self-contained — no CDN, no bundled libraries, no telemetry — with one
214
+ outbound request named here rather than buried: the OpenStreetMap background tiles, which
215
+ reveal the map view you are looking at (never your data) and which the panel drops to a
216
+ plain backdrop when the host blocks them. The preview is deliberately lossy (simplified
217
+ geometry, capped feature counts): the dataset of record stays on disk with its manifest.
218
+
219
+ ## Plans: reject wrong analyses before they run
220
+
221
+ In [GISAgentBench](https://arxiv.org/abs/2608.01645) — 349 practitioner-sourced tasks over
222
+ 128 GIS APIs — the best frontier agent completes 32.7% of tasks under strict scoring, and
223
+ planning defects dominate the failures: missing operations in 28.3% of failed runs and
224
+ wrong operation order in 18.4% (multi-label, so up to ~47% involve a planning mistake),
225
+ against 7.8% for parameter errors. MapSmith attacks this where it is cheapest: the agent
226
+ submits a **typed plan**, and static validation rejects unknown operations (with
227
+ suggestions), missing arguments, forward references, absent input files and CRS-unsuitable
228
+ steps **before anything executes** — with machine-actionable error codes the agent can
229
+ repair.
230
+
231
+ ```json
232
+ {
233
+ "goal": "buildings within 300 m of rivers",
234
+ "steps": [
235
+ {"id": "buf", "operation": "buffer_layer",
236
+ "arguments": {"input_path": "rivers.gpkg", "distance_meters": 300,
237
+ "output_path": "rivers_300m.parquet"}},
238
+ {"id": "cut", "operation": "clip_layer",
239
+ "arguments": {"input_path": "buildings.parquet", "mask_path": "$buf",
240
+ "output_path": "at_risk.parquet"}}
241
+ ]
242
+ }
243
+ ```
244
+
245
+ `"$buf"` consumes the output of step `buf`; references may only point backwards, so plans
246
+ are acyclic by construction. `validate_plan` also simulates the CRS of every intermediate
247
+ dataset from the real input files. `execute_plan` then runs the chain with per-step
248
+ provenance plus a plan-level manifest (`<output>.plan.json`) fingerprinting the exact plan
249
+ that produced the result.
250
+
251
+ ## Confinement
252
+
253
+ UNC hosts and NTFS alternate data streams are refused in every path *argument* of every
254
+ tool call, before anything touches the filesystem (on Windows even an existence check on a
255
+ UNC path talks to an attacker-chosen host). Remote and virtual forms — GDAL `/vsi*`,
256
+ `https://` COGs — stay available while the server is unconfined, because cloud-native data
257
+ is a feature, and are refused once a workspace is set. Validated plans are stricter by
258
+ design and reject every non-local form.
259
+
260
+ Set `MAPSMITH_WORKSPACE=/data` to confine the server to one directory:
261
+
262
+ - every path argument of every tool must resolve inside the workspace (checked at the MCP
263
+ boundary, and again by plan validation with stable error codes);
264
+ - the `run_sql` DuckDB connection is sandboxed in the engine itself, because SQL text is
265
+ out of reach of a textual path check: filesystem whitelisted to the workspace
266
+ (`allowed_directories` + external access off, which also covers GDAL-backed `ST_Read`),
267
+ extension install and load refused, memory and temp disk capped
268
+ (`MAPSMITH_DUCKDB_MEMORY`, default 4GB; `MAPSMITH_DUCKDB_TEMP_LIMIT`, default 8GB),
269
+ configuration locked. SQL can name any path it likes; the engine refuses to open it.
270
+
271
+ Without a workspace, *file* access is deliberately unconfined — fine for a local stdio
272
+ server on your own files — and plan validation flags `run_sql` steps with a
273
+ `SQL_NOT_SANDBOXED` warning. Code execution is still closed: extension autoloading and
274
+ community extensions are off (`shellfs` turns a filename into a shell command), unsigned
275
+ extensions are refused, DuckDB's HTTP and S3 filesystems are disabled, and the
276
+ configuration is locked, so untrusted SQL cannot turn file access into code execution.
277
+
278
+ **The network is not closed in that mode, by the same decision that keeps cloud-native
279
+ data working**: GDAL carries its own HTTP client, so `ST_Read('/vsicurl/https://…')` in raw
280
+ SQL reads whatever the host can reach — internal services and link-local metadata
281
+ endpoints included — and the URL is chosen by whoever wrote the SQL, so it can also carry a
282
+ string out. Setting `MAPSMITH_WORKSPACE` closes all of it, GDAL included, and the test
283
+ suite asserts both halves (`tests/test_duckdb_sandbox.py`). If your `run_sql` input is not
284
+ trusted and your host sits somewhere interesting, set a workspace. The full threat model —
285
+ and what is explicitly *not* covered — is in [SECURITY.md](SECURITY.md).
286
+
287
+ Fine print, because it changes how you deploy this: the path jail assumes a single trusted
288
+ writer of the workspace filesystem (paths are resolved at check time, so a symlink swap by
289
+ another local process is out of scope); the DuckDB spatial extension is fetched once per
290
+ environment, so on air-gapped machines pre-install it (`python -c "import duckdb;
291
+ duckdb.connect().install_extension('spatial')"`) before locking the network down; and the
292
+ HTTP transport has no authentication in this release, so keep it on loopback or a trusted
293
+ network. For real isolation, run the container and mount only the data you want it to see.
294
+
295
+ ## We measured whether this actually helps
296
+
297
+ Claims about agent performance are cheap, so
298
+ [**docs/benchmarks.md**](docs/benchmarks.md) reports an A/B on
299
+ [GABench](https://github.com/GeoX-Lab/GABench) — 57 executable GIS tasks over a
300
+ 133-tool server, scored by its deterministic evaluator — where the *only*
301
+ variable is whether the agent's typed plan is validated before the solver runs.
302
+
303
+ The honest headline is a **null result**, on a frontier model and on a small
304
+ one, and the interesting part is why:
305
+
306
+ | | Arm A (no gate) | Arm B (gate) |
307
+ |---|---|---|
308
+ | Sonnet 5 — TAO / PEA | 0.824 / 0.430 | 0.781 / 0.425 |
309
+ | Haiku 4.5 — TAO / PEA | 0.660 / 0.320 | 0.714 / 0.366 |
310
+
311
+ Haiku looks like a clean win until you notice the gate only fired on 4 of 57
312
+ plans, and that the 53 tasks it never touched moved by just as much: the
313
+ aggregate delta is run-to-run variance, and measuring that noise floor
314
+ (2–5 points per metric on a single repetition) is the reusable result. What
315
+ survives is narrower — on the plans it did repair, tool selection improved by
316
+ +0.19 TAO — and it points at where the failures actually are: PEA around 0.4
317
+ in every arm, i.e. wrong parameters and missing outputs at *execution* time,
318
+ which is why MapSmith enforces its plans at the execution boundary and verifies
319
+ inputs and outputs at runtime rather than advising an agent that improvises.
320
+
321
+ The harness is in [`benchmarks/gabench-ab/`](benchmarks/gabench-ab/), including
322
+ the `split_analysis.py` that took our own win apart.
323
+
324
+ ## Notebook gallery
325
+
326
+ Three executable walkthroughs in [`examples/`](examples/): verified buffer+clip with
327
+ provenance manifests, terrain and hydrology on a real 520×520 USGS DEM of **Mount St.
328
+ Helens**, and a deliberately wrong plan rejected before execution and then repaired. The
329
+ terrain notebook also shows what happens when reality bites: that DEM is stored with the
330
+ standard TIFF predictor, which Whitebox Workflows 2.x does not undo when reading
331
+ ([upstream report](https://github.com/jblindsay/whitebox_next_gen/issues/32)), so MapSmith
332
+ detects it, converts the input first, and discloses the workaround in the manifest.
333
+
334
+ ## Architecture
335
+
336
+ ```
337
+ AI agent (Claude / ChatGPT / Copilot / your app)
338
+ │ MCP (stdio local · Streamable HTTP remote)
339
+ ▼
340
+ ┌─────────────────────────────────────────────┐
341
+ │ MapSmith server │
342
+ │ · semantic tools + operation catalog │
343
+ │ · parameter validation, CRS discipline │
344
+ │ · provenance recorder (lineage manifests) │
345
+ ├─────────────────────────────────────────────┤
346
+ │ Engines │
347
+ │ · vector: GeoPandas/Shapely (built-in) │
348
+ │ · SQL/analytics: DuckDB Spatial (built-in) │
349
+ │ · heavy joins: SedonaDB ([sedona] extra) │
350
+ │ · zonal stats: exactextract ([raster]) │
351
+ │ · terrain/hydro: Whitebox NG ([whitebox]) │
352
+ │ · qgis_process / GRASS sidecar (roadmap, │
353
+ │ GPL-isolated via subprocess) │
354
+ └─────────────────────────────────────────────┘
355
+ ```
356
+
357
+ ## When not to use MapSmith
358
+
359
+ - **You need an authenticated remote server today.** The Streamable HTTP transport has no
360
+ authentication in this release: anyone who can reach the endpoint can run every tool
361
+ against everything the process can see. Loopback or a trusted network only
362
+ ([SECURITY.md](SECURITY.md)).
363
+ - **You want a sandbox for arbitrary agent code.** MapSmith confines paths and the SQL
364
+ engine; there is no code-execution tool yet, and a path jail is not a container.
365
+ - **You need cartography.** No styling, no layouts, no print composer. Outputs are
366
+ datasets, plus a lossy read-only preview panel — not maps you publish.
367
+ - **Your data lives in a database.** MapSmith reads and writes files (GeoParquet,
368
+ GeoPackage, anything GDAL opens). There is no PostGIS engine and no database catalog —
369
+ the `[postgres]` extra is for the optional job ledger, not for data.
370
+ - **You want the full breadth of a desktop GIS.** 16 tools plus a catalog that tells the
371
+ agent what does *not* exist yet. The ~900 QGIS Processing algorithms are on the roadmap,
372
+ not in the box.
373
+ - **You expect plan validation to make a weak model strong.** Our own A/B says advisory
374
+ validation upstream of an improvising solver does approximately nothing at aggregate
375
+ level; MapSmith's answer is enforcement at the execution boundary, and that hypothesis
376
+ is not measured yet.
377
+ - **You want us to debug your local geospatial toolchain.** Docker, or `uvx` where the
378
+ wheels work, are the only supported paths; a hand-built native GDAL stack is not, on
379
+ purpose.
380
+
381
+ ## Roadmap
382
+
383
+ - [x] Zonal statistics (exactextract, exact fractional coverage)
384
+ - [x] Whitebox Next Gen adapter: hillshade, flow accumulation, watershed (in-memory, open tier)
385
+ - [x] Typed analysis plans: static validation against the operation registry + simulated CRS flow before execution
386
+ - [x] Runtime verification: input preconditions, warnings with hints in the tool result, bounded deterministic repair recorded in the manifest
387
+ - [x] MCP Apps in-chat map panel with provenance cards (self-contained, works under the default host sandbox)
388
+ - [ ] Agent-loop repair: hand verification failures back to the agent for a bounded number of retries
389
+ - [ ] More terrain & hydrology: slope/aspect, stream network extraction
390
+ - [ ] QGIS Processing sidecar (subprocess-isolated): ~900 algorithms
391
+ - [ ] Sandboxed code-execution tool for the long tail
392
+ - [ ] Map panel: MapLibre vector rendering and shareable viewer URLs (raster OSM tiles already ship)
393
+ - [ ] Authenticated remote mode (OAuth on the existing Streamable HTTP transport) and long-job progress via MCP Tasks
394
+
395
+ ## License and project
396
+
397
+ - MapSmith server and engines: **AGPL-3.0-or-later** (see [LICENSE](LICENSE))
398
+ - Client SDK and tool-schema definitions (future `sdk/`): **Apache-2.0**
399
+
400
+ You can self-host MapSmith freely, forever. If you modify it and offer it as a service, the
401
+ AGPL asks you to share your changes — or [talk to us](mailto:mapsmith@proton.me) about a
402
+ commercial license.
403
+
404
+ Release notes are in [CHANGELOG.md](CHANGELOG.md), how to contribute in
405
+ [CONTRIBUTING.md](CONTRIBUTING.md), how to report a vulnerability in
406
+ [SECURITY.md](SECURITY.md). "MapSmith" is a trademark of the MapSmith project — see
407
+ [TRADEMARKS.md](TRADEMARKS.md). Updates: [@mapsmith_ai](https://x.com/mapsmith_ai) ·
408
+ [Bluesky](https://bsky.app/profile/mapsmith.bsky.social).
409
+
410
+ <!-- mcp-name: io.github.mapsmith-ai/mapsmith -->