beans-picker 0.3.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- beans_picker-0.3.0/.env.example +6 -0
- beans_picker-0.3.0/.gitignore +37 -0
- beans_picker-0.3.0/CHANGELOG.md +51 -0
- beans_picker-0.3.0/CONTRIBUTING.md +59 -0
- beans_picker-0.3.0/LICENSE +21 -0
- beans_picker-0.3.0/PKG-INFO +227 -0
- beans_picker-0.3.0/README.md +198 -0
- beans_picker-0.3.0/SECURITY.md +49 -0
- beans_picker-0.3.0/bench/__init__.py +0 -0
- beans_picker-0.3.0/bench/fixture-app/build.sh +19 -0
- beans_picker-0.3.0/bench/fixture-app/main.m +183 -0
- beans_picker-0.3.0/bench/judge.py +145 -0
- beans_picker-0.3.0/bench/results/runs.jsonl +48 -0
- beans_picker-0.3.0/bench/results/summary.md +27 -0
- beans_picker-0.3.0/bench/run.py +584 -0
- beans_picker-0.3.0/bench/summarize.py +108 -0
- beans_picker-0.3.0/bench/tasks.json +61 -0
- beans_picker-0.3.0/pyproject.toml +118 -0
- beans_picker-0.3.0/scripts/check_wheel.py +36 -0
- beans_picker-0.3.0/scripts/smoke_installed.py +66 -0
- beans_picker-0.3.0/skills/beans-picker/SKILL.md +117 -0
- beans_picker-0.3.0/src/beans_picker/__init__.py +1 -0
- beans_picker-0.3.0/src/beans_picker/__main__.py +6 -0
- beans_picker-0.3.0/src/beans_picker/_aio.py +32 -0
- beans_picker-0.3.0/src/beans_picker/_json.py +101 -0
- beans_picker-0.3.0/src/beans_picker/_numbers.py +116 -0
- beans_picker-0.3.0/src/beans_picker/_proc.py +92 -0
- beans_picker-0.3.0/src/beans_picker/_stdin.py +66 -0
- beans_picker-0.3.0/src/beans_picker/_text.py +97 -0
- beans_picker-0.3.0/src/beans_picker/act/__init__.py +0 -0
- beans_picker-0.3.0/src/beans_picker/act/execute.py +487 -0
- beans_picker-0.3.0/src/beans_picker/act/pixel.py +204 -0
- beans_picker-0.3.0/src/beans_picker/candidates/__init__.py +0 -0
- beans_picker-0.3.0/src/beans_picker/candidates/build.py +340 -0
- beans_picker-0.3.0/src/beans_picker/candidates/describe.py +91 -0
- beans_picker-0.3.0/src/beans_picker/candidates/keypad.py +117 -0
- beans_picker-0.3.0/src/beans_picker/candidates/menu.py +61 -0
- beans_picker-0.3.0/src/beans_picker/candidates/prune.py +77 -0
- beans_picker-0.3.0/src/beans_picker/candidates/safety.py +161 -0
- beans_picker-0.3.0/src/beans_picker/candidates/types.py +68 -0
- beans_picker-0.3.0/src/beans_picker/cli.py +83 -0
- beans_picker-0.3.0/src/beans_picker/config.py +170 -0
- beans_picker-0.3.0/src/beans_picker/driver/__init__.py +0 -0
- beans_picker-0.3.0/src/beans_picker/driver/app.py +143 -0
- beans_picker-0.3.0/src/beans_picker/driver/lock.py +22 -0
- beans_picker-0.3.0/src/beans_picker/driver/markdown.py +108 -0
- beans_picker-0.3.0/src/beans_picker/driver/mcp.py +374 -0
- beans_picker-0.3.0/src/beans_picker/driver/sentinel.py +140 -0
- beans_picker-0.3.0/src/beans_picker/driver/types.py +366 -0
- beans_picker-0.3.0/src/beans_picker/errors.py +96 -0
- beans_picker-0.3.0/src/beans_picker/jev/__init__.py +0 -0
- beans_picker-0.3.0/src/beans_picker/jev/client.py +213 -0
- beans_picker-0.3.0/src/beans_picker/jev/questions.py +55 -0
- beans_picker-0.3.0/src/beans_picker/jev/rank.py +143 -0
- beans_picker-0.3.0/src/beans_picker/jev/state.py +225 -0
- beans_picker-0.3.0/src/beans_picker/log.py +48 -0
- beans_picker-0.3.0/src/beans_picker/menus/__init__.py +0 -0
- beans_picker-0.3.0/src/beans_picker/menus/keyequiv.py +194 -0
- beans_picker-0.3.0/src/beans_picker/menus/menukeys.py +288 -0
- beans_picker-0.3.0/src/beans_picker/native/__init__.py +0 -0
- beans_picker-0.3.0/src/beans_picker/native/axtext.m +81 -0
- beans_picker-0.3.0/src/beans_picker/native/menukeys.m +94 -0
- beans_picker-0.3.0/src/beans_picker/observe/__init__.py +0 -0
- beans_picker-0.3.0/src/beans_picker/observe/exacttext.py +210 -0
- beans_picker-0.3.0/src/beans_picker/observe/helpers.py +179 -0
- beans_picker-0.3.0/src/beans_picker/observe/identity.py +32 -0
- beans_picker-0.3.0/src/beans_picker/observe/menudiff.py +105 -0
- beans_picker-0.3.0/src/beans_picker/observe/normalize.py +58 -0
- beans_picker-0.3.0/src/beans_picker/observe/png.py +164 -0
- beans_picker-0.3.0/src/beans_picker/observe/signature.py +29 -0
- beans_picker-0.3.0/src/beans_picker/observe/snapshot.py +394 -0
- beans_picker-0.3.0/src/beans_picker/observe/types.py +107 -0
- beans_picker-0.3.0/src/beans_picker/observe/visual.py +179 -0
- beans_picker-0.3.0/src/beans_picker/paths.py +31 -0
- beans_picker-0.3.0/src/beans_picker/py.typed +0 -0
- beans_picker-0.3.0/src/beans_picker/server.py +192 -0
- beans_picker-0.3.0/src/beans_picker/tools/__init__.py +0 -0
- beans_picker-0.3.0/src/beans_picker/tools/act.py +371 -0
- beans_picker-0.3.0/src/beans_picker/tools/args.py +339 -0
- beans_picker-0.3.0/src/beans_picker/tools/extract.py +230 -0
- beans_picker-0.3.0/src/beans_picker/tools/observe.py +65 -0
- beans_picker-0.3.0/src/beans_picker/tools/present.py +69 -0
- beans_picker-0.3.0/src/beans_picker/tools/session.py +250 -0
- beans_picker-0.3.0/src/beans_picker/verify/__init__.py +0 -0
- beans_picker-0.3.0/src/beans_picker/verify/effect.py +168 -0
- beans_picker-0.3.0/tests/__init__.py +0 -0
- beans_picker-0.3.0/tests/act_support.py +209 -0
- beans_picker-0.3.0/tests/conftest.py +33 -0
- beans_picker-0.3.0/tests/fake_driver.py +124 -0
- beans_picker-0.3.0/tests/fakes.py +104 -0
- beans_picker-0.3.0/tests/fixtures/calculator.window-state.json +2446 -0
- beans_picker-0.3.0/tests/fixtures/initialize.json +13 -0
- beans_picker-0.3.0/tests/fixtures/textedit.window-state.json +2563 -0
- beans_picker-0.3.0/tests/fixtures/tools_list.json +180 -0
- beans_picker-0.3.0/tests/fixtures/validation_errors.jsonl +47 -0
- beans_picker-0.3.0/tests/helpers.py +88 -0
- beans_picker-0.3.0/tests/raw_mcp.py +76 -0
- beans_picker-0.3.0/tests/test_act.py +338 -0
- beans_picker-0.3.0/tests/test_act_tool_paths.py +302 -0
- beans_picker-0.3.0/tests/test_aio.py +68 -0
- beans_picker-0.3.0/tests/test_app.py +294 -0
- beans_picker-0.3.0/tests/test_args.py +133 -0
- beans_picker-0.3.0/tests/test_bench.py +722 -0
- beans_picker-0.3.0/tests/test_candidates.py +1047 -0
- beans_picker-0.3.0/tests/test_cli.py +80 -0
- beans_picker-0.3.0/tests/test_client.py +250 -0
- beans_picker-0.3.0/tests/test_config.py +208 -0
- beans_picker-0.3.0/tests/test_desktop.py +207 -0
- beans_picker-0.3.0/tests/test_driver_mcp.py +465 -0
- beans_picker-0.3.0/tests/test_driver_types.py +182 -0
- beans_picker-0.3.0/tests/test_e2e.py +172 -0
- beans_picker-0.3.0/tests/test_effect.py +90 -0
- beans_picker-0.3.0/tests/test_errors.py +79 -0
- beans_picker-0.3.0/tests/test_exacttext.py +301 -0
- beans_picker-0.3.0/tests/test_execute.py +573 -0
- beans_picker-0.3.0/tests/test_extract.py +186 -0
- beans_picker-0.3.0/tests/test_golden.py +81 -0
- beans_picker-0.3.0/tests/test_helpers.py +212 -0
- beans_picker-0.3.0/tests/test_identity.py +30 -0
- beans_picker-0.3.0/tests/test_json.py +55 -0
- beans_picker-0.3.0/tests/test_judge.py +220 -0
- beans_picker-0.3.0/tests/test_lock.py +46 -0
- beans_picker-0.3.0/tests/test_log.py +73 -0
- beans_picker-0.3.0/tests/test_markdown.py +164 -0
- beans_picker-0.3.0/tests/test_menudiff.py +149 -0
- beans_picker-0.3.0/tests/test_menukeys.py +383 -0
- beans_picker-0.3.0/tests/test_native_build.py +24 -0
- beans_picker-0.3.0/tests/test_native_sources.py +59 -0
- beans_picker-0.3.0/tests/test_normalize.py +94 -0
- beans_picker-0.3.0/tests/test_numbers.py +144 -0
- beans_picker-0.3.0/tests/test_output_shapes.py +105 -0
- beans_picker-0.3.0/tests/test_paths.py +18 -0
- beans_picker-0.3.0/tests/test_pixel.py +139 -0
- beans_picker-0.3.0/tests/test_png.py +134 -0
- beans_picker-0.3.0/tests/test_proc.py +98 -0
- beans_picker-0.3.0/tests/test_rank.py +179 -0
- beans_picker-0.3.0/tests/test_sentinel.py +270 -0
- beans_picker-0.3.0/tests/test_server.py +204 -0
- beans_picker-0.3.0/tests/test_server_cancel.py +72 -0
- beans_picker-0.3.0/tests/test_session.py +321 -0
- beans_picker-0.3.0/tests/test_signature.py +99 -0
- beans_picker-0.3.0/tests/test_snapshot.py +660 -0
- beans_picker-0.3.0/tests/test_source_rules.py +95 -0
- beans_picker-0.3.0/tests/test_state.py +210 -0
- beans_picker-0.3.0/tests/test_stdin.py +32 -0
- beans_picker-0.3.0/tests/test_stdio.py +111 -0
- beans_picker-0.3.0/tests/test_summarize.py +109 -0
- beans_picker-0.3.0/tests/test_text.py +116 -0
- beans_picker-0.3.0/tests/test_verify.py +208 -0
- beans_picker-0.3.0/tests/test_visual.py +322 -0
- beans_picker-0.3.0/tests/tool_fakes.py +144 -0
- beans_picker-0.3.0/uv.lock +1149 -0
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
# Environments and secrets
|
|
2
|
+
.venv/
|
|
3
|
+
.env*
|
|
4
|
+
!.env.example
|
|
5
|
+
|
|
6
|
+
# Build output
|
|
7
|
+
dist/
|
|
8
|
+
dist-from-sdist/
|
|
9
|
+
build/
|
|
10
|
+
*.egg-info/
|
|
11
|
+
|
|
12
|
+
# Caches
|
|
13
|
+
__pycache__/
|
|
14
|
+
*.py[cod]
|
|
15
|
+
.pytest_cache/
|
|
16
|
+
.mypy_cache/
|
|
17
|
+
.ruff_cache/
|
|
18
|
+
.coverage
|
|
19
|
+
.coverage.*
|
|
20
|
+
coverage.xml
|
|
21
|
+
htmlcov/
|
|
22
|
+
|
|
23
|
+
# Bench scratch space and built fixture app
|
|
24
|
+
bench/.work/
|
|
25
|
+
bench/fixture-app/*.app
|
|
26
|
+
|
|
27
|
+
# macOS
|
|
28
|
+
.DS_Store
|
|
29
|
+
|
|
30
|
+
# Local agent instructions and settings
|
|
31
|
+
CLAUDE.md
|
|
32
|
+
CLAUDE.local.md
|
|
33
|
+
AGENTS.md
|
|
34
|
+
GEMINI.md
|
|
35
|
+
.claude/
|
|
36
|
+
.cursor/
|
|
37
|
+
.cursorrules
|
|
@@ -0,0 +1,51 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to this project are documented here. The format follows
|
|
4
|
+
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and versions follow
|
|
5
|
+
[Semantic Versioning](https://semver.org/).
|
|
6
|
+
|
|
7
|
+
## [Unreleased]
|
|
8
|
+
|
|
9
|
+
To be released as 0.3.0, the first release on PyPI. Versions 0.1 and 0.2 were never published.
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- The `beans-picker` command and `python -m beans_picker`: an MCP server over stdio with the tools `observe`,
|
|
14
|
+
`act` and `extract`. It runs on macOS; on other systems it exits with an error.
|
|
15
|
+
- `beans-picker --version` and `beans-picker --help`, which work on every system.
|
|
16
|
+
- `BEANS_PICKER_LOG_LEVEL` sets the level of the server's log on stderr (default `WARNING`).
|
|
17
|
+
- The Jev key and other settings are read from the environment, then from `.env.local` / `.env`
|
|
18
|
+
in the checkout when the package runs from a beans-picker checkout, then from `~/.config/beans-picker/`
|
|
19
|
+
(`$XDG_CONFIG_HOME/beans-picker`), so an installed server finds its key without a checkout.
|
|
20
|
+
- Time limits on every request: `CUA_DRIVER_TIMEOUT` (default 120 s) for cua-driver's startup and
|
|
21
|
+
each call, `JEV_CONNECT_TIMEOUT` (10 s) and `JEV_READ_TIMEOUT` (120 s) for Jev. A request that
|
|
22
|
+
runs out of time fails with `driver_timeout` or `jev_unavailable` instead of blocking later calls.
|
|
23
|
+
At shutdown, the final `end_session` to cua-driver waits at most 5 seconds.
|
|
24
|
+
|
|
25
|
+
### Changed
|
|
26
|
+
|
|
27
|
+
- Failures that are not a tool's own refusal have their own codes: `driver_unavailable`,
|
|
28
|
+
`driver_timeout`, `driver_error` and `jev_bad_response`, besides `internal`. They are logged with
|
|
29
|
+
their traceback at `WARNING`.
|
|
30
|
+
- A call to an unknown tool is a JSON-RPC error (-32602), not a tool result.
|
|
31
|
+
- `pid` and `windowId` must be at least 1.
|
|
32
|
+
- Unknown command-line arguments print the usage and exit with status 2 instead of starting the
|
|
33
|
+
server.
|
|
34
|
+
- An empty `JEV_API_KEY` no longer hides `TYPESAFE_API_KEY`.
|
|
35
|
+
- The exact-text helper also reads switches (`AXSwitch`). Its source changed, so it is rebuilt:
|
|
36
|
+
run `beans-picker grant-ax` again after upgrading.
|
|
37
|
+
|
|
38
|
+
### Fixed
|
|
39
|
+
|
|
40
|
+
- A call canceled by the client while it was still waiting for its turn no longer runs.
|
|
41
|
+
- A step in `then` that raises no longer discards the steps before it: it is reported as a `failed`
|
|
42
|
+
step, with the completed steps and the `skipped` count.
|
|
43
|
+
- A Jev answer with a probability or confidence outside [0, 1], or a ranking without
|
|
44
|
+
probabilities, is rejected as `jev_bad_response`.
|
|
45
|
+
- `~` in `CUA_DRIVER_BIN` is expanded.
|
|
46
|
+
- A screenshot that cua-driver saved to a path of its own is deleted on every failure path too.
|
|
47
|
+
- A server made by `create_server()` without a session closes the session it made when it stops.
|
|
48
|
+
- A text field named only by its placeholder keeps its candidate ids once it is typed in, so text
|
|
49
|
+
entered into it is checked in that field instead of ending as `mismatch` ("the target field is gone").
|
|
50
|
+
|
|
51
|
+
[Unreleased]: https://github.com/mimo-3/beans-picker/commits/main
|
|
@@ -0,0 +1,59 @@
|
|
|
1
|
+
# Contributing
|
|
2
|
+
|
|
3
|
+
## Setup
|
|
4
|
+
|
|
5
|
+
You need [uv](https://docs.astral.sh/uv/). The test suite runs on macOS and Linux; using the
|
|
6
|
+
server needs macOS, cua-driver and a Jev key (see the README).
|
|
7
|
+
|
|
8
|
+
```sh
|
|
9
|
+
uv sync
|
|
10
|
+
```
|
|
11
|
+
|
|
12
|
+
## Checks
|
|
13
|
+
|
|
14
|
+
Run all of these before opening a pull request; CI runs the same ones.
|
|
15
|
+
|
|
16
|
+
```sh
|
|
17
|
+
uv run ruff check
|
|
18
|
+
uv run ruff format --check
|
|
19
|
+
uv run mypy
|
|
20
|
+
uv run pytest
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
mypy runs in strict mode over `src`, `tests` and `bench`, and coverage must stay at 90% or more.
|
|
24
|
+
|
|
25
|
+
## Ground rules
|
|
26
|
+
|
|
27
|
+
- **Behavior is the contract.** Tool names, input schemas, statuses, codes, messages, JSON key
|
|
28
|
+
order and number formatting are what clients rely on. A change to any of them needs a test
|
|
29
|
+
that pins the new behavior and a CHANGELOG entry.
|
|
30
|
+
- **stdout belongs to MCP.** The server writes nothing but protocol frames to stdout. Log with the
|
|
31
|
+
`beans_picker` logger; it goes to stderr.
|
|
32
|
+
- **No OS access in tests.** Subprocesses, cua-driver and the cache directory sit behind small
|
|
33
|
+
seams (`runner=`, `Paths`, fakes in `tests/fakes.py`), so the suite runs anywhere. Tests that
|
|
34
|
+
need macOS are marked `macos` and skipped elsewhere; the end-to-end test also needs
|
|
35
|
+
`BEANS_PICKER_E2E=1` and a real session with cua-driver.
|
|
36
|
+
- **Exact text.** Lengths are counted in UTF-16 units, the way the accessibility API counts them,
|
|
37
|
+
and whitespace means the fixed set in `beans_picker._text`. Use its regex classes instead of `\s`,
|
|
38
|
+
`\w`, `\d` and `.`.
|
|
39
|
+
- Code, comments and docs are in English.
|
|
40
|
+
|
|
41
|
+
## Native helpers
|
|
42
|
+
|
|
43
|
+
`src/beans_picker/native/*.m` are compiled on the user's machine, on first use. Keep them small,
|
|
44
|
+
read-only and free of dependencies beyond the macOS SDK. A change to either file changes its
|
|
45
|
+
cache key, so the next run rebuilds it.
|
|
46
|
+
|
|
47
|
+
## Benchmark
|
|
48
|
+
|
|
49
|
+
`bench/` compares Claude Code with cua-driver alone against Claude Code with beans-picker on eight
|
|
50
|
+
tasks. It needs a Mac with cua-driver, the `claude` CLI and an unlocked screen.
|
|
51
|
+
|
|
52
|
+
```sh
|
|
53
|
+
sh bench/fixture-app/build.sh # the fixture app, into ~/Library/Caches/beans-picker/fixture
|
|
54
|
+
uv run python -m bench.run --reps 3 # appends to bench/results/runs.jsonl
|
|
55
|
+
uv run python -m bench.summarize > bench/results/summary.md
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
Replace `bench/results/runs.jsonl` rather than appending to it when you publish new numbers, and
|
|
59
|
+
update the README tables from the new `summary.md`.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 mimo-3
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,227 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: beans-picker
|
|
3
|
+
Version: 0.3.0
|
|
4
|
+
Summary: An MCP server that handpicks the right control for agents driving macOS apps, acts without stealing focus, and checks the effect.
|
|
5
|
+
Project-URL: Repository, https://github.com/mimo-3/beans-picker
|
|
6
|
+
Project-URL: Issues, https://github.com/mimo-3/beans-picker/issues
|
|
7
|
+
Project-URL: Changelog, https://github.com/mimo-3/beans-picker/blob/main/CHANGELOG.md
|
|
8
|
+
Author: mimo-3
|
|
9
|
+
License-Expression: MIT
|
|
10
|
+
License-File: LICENSE
|
|
11
|
+
Keywords: accessibility,agents,automation,macos,mcp
|
|
12
|
+
Classifier: Development Status :: 4 - Beta
|
|
13
|
+
Classifier: Environment :: MacOS X
|
|
14
|
+
Classifier: Intended Audience :: Developers
|
|
15
|
+
Classifier: Operating System :: MacOS
|
|
16
|
+
Classifier: Programming Language :: Python :: 3 :: Only
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
20
|
+
Classifier: Topic :: Desktop Environment
|
|
21
|
+
Classifier: Typing :: Typed
|
|
22
|
+
Requires-Python: >=3.12
|
|
23
|
+
Requires-Dist: anyio<5,>=4.9
|
|
24
|
+
Requires-Dist: httpx2<3,>=2.13
|
|
25
|
+
Requires-Dist: mcp<2.3,>=2.2
|
|
26
|
+
Requires-Dist: pydantic<3,>=2.13
|
|
27
|
+
Requires-Dist: typesafe-sdk<0.8,>=0.7.1
|
|
28
|
+
Description-Content-Type: text/markdown
|
|
29
|
+
|
|
30
|
+
# beans-picker
|
|
31
|
+
|
|
32
|
+
Handpicks the right control. ~2× faster, ~1/6 the cost.
|
|
33
|
+
|
|
34
|
+
Not affiliated with Cua / trycua or TypeSafe; cua-driver and Jev are separate projects this server talks to.
|
|
35
|
+
|
|
36
|
+
An MCP server that lets Claude Code or Codex pick the right control in a macOS app, act on it without bringing the app to the front, and check that it worked. It does for native apps what Stagehand does for the browser:
|
|
37
|
+
|
|
38
|
+
- **The caller thinks.** Claude Code or Codex breaks the task down, picks the next step and decides what to do when something goes wrong.
|
|
39
|
+
- **Jev picks the element.** [TypeSafe Jev](https://typesafe.ai) chooses among candidate actions built from the accessibility tree. It never writes text, so it cannot invent a target.
|
|
40
|
+
- **cua-driver acts.** [cua-driver](https://github.com/trycua/cua) clicks, types and presses keys without bringing the app to the front.
|
|
41
|
+
|
|
42
|
+
The server itself never calls a generative model (no `claude -p`, no LLM API).
|
|
43
|
+
|
|
44
|
+
## Tools
|
|
45
|
+
|
|
46
|
+
| tool | input | what it does |
|
|
47
|
+
|---|---|---|
|
|
48
|
+
| `observe` | `app` \| `pid`, `windowId?`, `instruction?`, `limit?` | Lists the window's candidate actions (`id`, `kind`, what it does). With `instruction`, Jev ranks them and each comes with its probability `p`. |
|
|
49
|
+
| `act` | `app` \| `pid`, `windowId?`, `instruction`, `text?`, `candidateId?`, `allowDestructive?`, `modifiers?`, `then?` | Performs one action and checks its effect, then each step in `then` (same fields, up to 12) the same way, stopping at the first that is not `done` or `unverified`; each step's result is in `steps`. Jev picks the action for `instruction`, unless `candidateId` is given. `text` is entered exactly as given. `modifiers` (`shift`, `cmd`, `option`, `ctrl`) are held during a click or toggle, as a pixel click on a control visible on the window. |
|
|
50
|
+
| `extract` | `app` \| `pid`, `windowId?`, `instruction` | Returns the text or value of the element Jev picks, exactly as read. A table or list with no value of its own comes back as `rows` (each row's texts), any other container as `text`. Its `status` is `done`, `ambiguous` (the shortlist comes back with each element's value) or `not_found`. |
|
|
51
|
+
|
|
52
|
+
`act` returns a `status`:
|
|
53
|
+
|
|
54
|
+
| status | meaning |
|
|
55
|
+
|---|---|
|
|
56
|
+
| `done` | The effect asked for is observed. For text, the field reads **exactly** the expected text (see "Exact checks"). |
|
|
57
|
+
| `unverified` | The field changed, but its exact text could not be read (the helper has no Accessibility permission). |
|
|
58
|
+
| `no_effect` | The window's signature did not change over `BEANS_PICKER_EFFECT_RETAKES` fresh snapshots (default 5). Decided by count, not by waiting. |
|
|
59
|
+
| `mismatch` | Something changed, but not what was asked (for example the text landed in another field, or a trailing space was lost). |
|
|
60
|
+
| `ambiguous` | Jev had no clear leader. The top candidates are returned; call `act` again with the right `candidateId`. |
|
|
61
|
+
| `needs_confirmation` | The action may not be undoable (delete, close, send, quit …). Call again with `candidateId` and `allowDestructive: true`. |
|
|
62
|
+
| `not_found` | No candidate fits, or the `candidateId` is not on the window any more. |
|
|
63
|
+
| `failed` | cua-driver refused, the command needs the foreground, or the app came to the front (`foreground_violation`). |
|
|
64
|
+
|
|
65
|
+
A call that cannot run at all returns `isError` with `{"status": "failed", "code", "message"}`. Codes include `bad_target`, `window_not_found`, `screen_locked`, `jev_unavailable` (no key, Jev unreachable or too slow), `jev_bad_response`, `driver_unavailable` (cua-driver could not be started), `driver_timeout`, `driver_error` (cua-driver refused a call the tool needed) and `internal`. Anything but a tool's own refusal is also logged with its traceback at `WARNING`.
|
|
66
|
+
|
|
67
|
+
Each result also carries the action that ran (`action`, with the cua-driver route) and a summary of what changed on the window (`change`).
|
|
68
|
+
|
|
69
|
+
Candidate ids are derived from a control's stable identity (role, identifier, label and ancestors), not from the snapshot's element index, so an id from `observe` still works after the window changes, as long as the control is there.
|
|
70
|
+
|
|
71
|
+
## Install
|
|
72
|
+
|
|
73
|
+
Requirements: macOS, Python 3.12+, [cua-driver](https://github.com/trycua/cua) at `~/.local/bin/cua-driver` with its Accessibility and Screen Recording permissions, a Jev API key, and Xcode Command Line Tools (`clang`) for the two small native helpers, which are compiled on first use into `~/Library/Caches/beans-picker`.
|
|
74
|
+
|
|
75
|
+
```sh
|
|
76
|
+
uv tool install beans-picker
|
|
77
|
+
mkdir -p ~/.config/beans-picker && echo 'JEV_API_KEY=...' > ~/.config/beans-picker/.env.local
|
|
78
|
+
claude mcp add beans-picker -- "$(command -v beans-picker)" # absolute path: GUI-launched clients often lack ~/.local/bin on PATH
|
|
79
|
+
beans-picker grant-ax # exact text (optional, see below)
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
Without installing, let `uvx` fetch it on each start: `claude mcp add beans-picker -- "$(command -v uvx)" beans-picker`, and `uvx beans-picker grant-ax`.
|
|
83
|
+
|
|
84
|
+
For Codex, add the same absolute path as an MCP server in `~/.codex/config.toml` (`command -v beans-picker` prints it):
|
|
85
|
+
|
|
86
|
+
```toml
|
|
87
|
+
[mcp_servers.beans-picker]
|
|
88
|
+
command = "/Users/you/.local/bin/beans-picker"
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
From a checkout:
|
|
92
|
+
|
|
93
|
+
```sh
|
|
94
|
+
uv sync
|
|
95
|
+
echo 'JEV_API_KEY=...' > .env.local # read from the checkout, whatever the caller's cwd
|
|
96
|
+
claude mcp add beans-picker -- "$(command -v uv)" run --directory "$PWD" beans-picker
|
|
97
|
+
uv run beans-picker grant-ax
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
`beans-picker grant-ax` builds the read-only helper and asks macOS to list it; then turn on **beans-picker axtext** under System Settings > Privacy & Security > Accessibility. The helper is built from its source, so an upgrade that changes that source builds a new helper that macOS has not been told about: run `beans-picker grant-ax` again after upgrading if text results turn `unverified`. `beans-picker --version` prints the version and `beans-picker --help` the usage.
|
|
101
|
+
|
|
102
|
+
The agent skill in `skills/beans-picker` tells a model how to use these tools well: when cua-driver fits better, what to do with each `act` status, and recipes for combo boxes, number fields, long tables and checking the result. Link it where your agent looks for skills, e.g. `ln -s "$PWD/skills/beans-picker" ~/.claude/skills/beans-picker` (Codex: `~/.codex/skills`). Leave it out when benchmarking the bare tools.
|
|
103
|
+
|
|
104
|
+
## Configuration
|
|
105
|
+
|
|
106
|
+
| variable | meaning |
|
|
107
|
+
|---|---|
|
|
108
|
+
| `JEV_API_KEY` | Required for `observe` with an instruction, `act` without `candidateId`, and `extract`. `TYPESAFE_API_KEY` is used when it is not set or empty. |
|
|
109
|
+
| `BEANS_PICKER_MODEL` | The Jev model (default `jev-latest`). |
|
|
110
|
+
| `CUA_DRIVER_BIN` | Path to cua-driver (default `~/.local/bin/cua-driver`; a leading `~` is expanded). |
|
|
111
|
+
| `CUA_DRIVER_TIMEOUT` | Seconds to wait for cua-driver to start, and for its answer to each call (default 120). |
|
|
112
|
+
| `JEV_CONNECT_TIMEOUT` | Seconds to wait for a connection to Jev (default 10). |
|
|
113
|
+
| `JEV_READ_TIMEOUT` | Seconds to wait for Jev's answer to one request (default 120). |
|
|
114
|
+
| `BEANS_PICKER_EFFECT_RETAKES` | Fresh snapshots taken to see an effect (default 5). |
|
|
115
|
+
| `BEANS_PICKER_LOG_LEVEL` | Level of the server's log on stderr (default `WARNING`). |
|
|
116
|
+
| `TYPESAFE_BASE_URL` | The Jev endpoint (default `https://api.typesafe.ai`). |
|
|
117
|
+
|
|
118
|
+
Variables are read from the environment first, then from `.env.local` and `.env` in the checkout (only when the package runs from a beans-picker checkout), then from `.env.local` and `.env` in `~/.config/beans-picker` (`$XDG_CONFIG_HOME/beans-picker` when that is set). A value that is already set is never overridden. See `.env.example`.
|
|
119
|
+
|
|
120
|
+
## Rules the server keeps
|
|
121
|
+
|
|
122
|
+
- **Background only.** `bring_to_front`, `invoke_menu`, `move_cursor` and `delivery_mode: "foreground"` are refused before they reach cua-driver. Menu commands run as their keyboard shortcut sent to the app's pid; a command without a known shortcut returns `failed` / `foreground_required`. During `act`, the frontmost app is sampled continuously; if the target app comes to the front, the call stops with `foreground_violation`.
|
|
123
|
+
- **Time limits.** Every request to cua-driver and to Jev has a generous limit (see Configuration), so one that stalls ends its call with `driver_timeout` or `jev_unavailable` instead of blocking the queue. An action whose answer timed out may still have happened: observe before repeating it. At shutdown, the final `end_session` gets 5 seconds.
|
|
124
|
+
- **Screen lock.** While `CGSSessionScreenIsLocked` is set, every tool refuses with `screen_locked` and does nothing.
|
|
125
|
+
- **One call at a time.** Tool calls are queued, so two actions never interleave on the desktop.
|
|
126
|
+
- **Destructive actions** are marked, not hidden: `act` returns `needs_confirmation` unless `allowDestructive` is set, whoever picked the action. A checkbox is never destructive, since it can be switched back.
|
|
127
|
+
|
|
128
|
+
## How `act` works
|
|
129
|
+
|
|
130
|
+
```
|
|
131
|
+
snapshot ─► candidates ─► Jev picks (or candidateId) ─► gate ─► cua-driver ─► fresh snapshots ─► status
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
1. **Snapshot** (`src/beans_picker/observe`). `get_window_state` returns the AX tree. The structured elements are joined with the tree markdown, which holds identifiers, help text and static text. The menu bar is split off, and a signature of what matters (title, field values, visible text, other windows) is computed.
|
|
135
|
+
2. **Candidates** (`src/beans_picker/candidates`). Clicks, toggles, text entry (`set_value`, `type_into` at the caret, `append` at the end), pop-up choices, an on-screen keypad sequence for `text` ("12×7="), menu commands with their background shortcut, one step up or down on a slider or number field (an arrow key sent to it), the options of a web list box (a combo box's suggestions), a page down or up on a scroll view, table, list or web page, a named row's, link's or image's context menu (AXShowMenu, which a web page receives as a right-click), and Return / Escape / Tab / Space / the arrow keys / Shift+F10. With `text`, only the actions that enter text are in the running.
|
|
136
|
+
3. **Jev** (`src/beans_picker/jev`). One `system_one` request asks two choice questions over the same candidates: one with a `none` option, one forced. More than 60 candidates are sharded (best lexical match first), and the leaders of each shard go to a runoff. The gate acts on a leader at p ≥ 0.8, or at p ≥ 0.5 when the forced question agrees and the leader has twice the runner-up's probability. Otherwise the result is `ambiguous` (or `not_found` when `none` dominates).
|
|
137
|
+
4. **cua-driver** (`src/beans_picker/act`). AX press, pixel clicks for keypads, a pop-up pressed open and then its item pressed (a menu left open over the page, as Chrome's is, is closed with Escape), `set_value` / `type_text` (a web page's number field ignores an AXValue write, so it is retyped: End, ⇧Home, then `type_text`; once an app's page has ignored a write, its text fields are typed into straight away), ⌘↓ then `type_text` for `append`, a wheel event at the area for a scroll, and the menu shortcut as a pid-targeted hotkey. Stale element tokens are rebound by stable key.
|
|
138
|
+
5. **Effect** (`src/beans_picker/verify/effect.py`). Fresh snapshots are taken, up to the retake count, until the effect shows.
|
|
139
|
+
|
|
140
|
+
## Exact checks
|
|
141
|
+
|
|
142
|
+
A false success is the worst failure a desktop agent can have: text judged by `trim` or by "contains", a search field mistaken for the document, a trailing space silently lost. `act` therefore checks text by **exact equality, in the targeted field only**:
|
|
143
|
+
|
|
144
|
+
- `set_value`: the field reads exactly `text`.
|
|
145
|
+
- `append`: the field reads exactly its previous text followed by `text`.
|
|
146
|
+
- `type_into`: the field reads its previous text with `text` inserted whole at one position.
|
|
147
|
+
|
|
148
|
+
cua-driver 0.8 does not give the exact text: it trims whitespace at both ends of a value, shows an empty field's placeholder as its value, and leaves out a checkbox's state. The read-only helper `src/beans_picker/native/axtext.m` reads AXValue as it is. It needs the Accessibility permission, which is why it runs as its own tiny background app (`beans-picker axtext`) that you can allow on its own, without granting anything to your terminal. Without it, a text action that visibly changed the field returns `unverified`, never `done`. A checkbox whose state is not in the tree is judged by its own pixels, captured in the background before and after the click.
|
|
149
|
+
|
|
150
|
+
## Privacy
|
|
151
|
+
|
|
152
|
+
Window text leaves the machine only when a step needs Jev. [SECURITY.md](https://github.com/mimo-3/beans-picker/blob/main/SECURITY.md) lists exactly what is sent, what runs locally, and how to report a vulnerability.
|
|
153
|
+
|
|
154
|
+
## Development
|
|
155
|
+
|
|
156
|
+
```sh
|
|
157
|
+
uv sync
|
|
158
|
+
uv run ruff check
|
|
159
|
+
uv run ruff format --check
|
|
160
|
+
uv run mypy
|
|
161
|
+
uv run pytest
|
|
162
|
+
```
|
|
163
|
+
|
|
164
|
+
The benchmark is in `bench/` (see below). `bench/fixture-app` is a small AppKit window used only by the benchmark; `sh bench/fixture-app/build.sh` builds it into `~/Library/Caches/beans-picker/fixture`. `python -m bench.run`, `python -m bench.judge` and `python -m bench.summarize` run from the repository root. See [CONTRIBUTING.md](https://github.com/mimo-3/beans-picker/blob/main/CONTRIBUTING.md).
|
|
165
|
+
|
|
166
|
+
## Benchmark
|
|
167
|
+
|
|
168
|
+
Run on 2026-09-24 with an unpublished pre-release of beans-picker (0.2.0) and `--reps 3`: 8 tasks × 2 conditions × 3 repetitions = 48 runs, Claude Code headless (`claude -p`) with `--model sonnet`, one run at a time. Raw records: [`bench/results/runs.jsonl`](https://github.com/mimo-3/beans-picker/blob/main/bench/results/runs.jsonl); the tables below are `python -m bench.summarize` of them.
|
|
169
|
+
|
|
170
|
+
**Conditions.** Both get the same prompt (task, target pid and window id, "work in the background", and a final `RESULT: success|failure` line), no built-in tools (`--tools ""`) and only one MCP server:
|
|
171
|
+
|
|
172
|
+
- **(a) cua-driver only**: `cua-driver mcp`, with `bring_to_front`, `move_cursor`, `kill_app`, `get_desktop_state` and a few other tools denied.
|
|
173
|
+
- **(b) beans-picker**: this server only.
|
|
174
|
+
|
|
175
|
+
**Tasks** ([`bench/tasks.json`](https://github.com/mimo-3/beans-picker/blob/main/bench/tasks.json)). Five run on `bench/fixture-app`, a small AppKit window made for the bench (so no user document is touched): a Name with leading and trailing spaces, appending to a note body next to a search field, a pop-up and Save, clearing a search field next to a destructive "Delete note" button, and fixing an email plus a checkbox. Three run on Calculator: `(48 + 16) / 8`, 15% of 80, and `7 − 19` then change sign. TextEdit and Notes were not used.
|
|
176
|
+
|
|
177
|
+
**Judging.** Success is decided only by [`bench/judge.py`](https://github.com/mimo-3/beans-picker/blob/main/bench/judge.py), a separate script that compares the final state with the task's `expected` JSON by exact equality: every key of the fixture's state file (the controls' values, written by the app itself), or Calculator's display read over accessibility (bidi marks removed, nothing else). The agent's `RESULT:` line is used only to count false success claims. The front app was sampled every 200 ms during each run to count focus steals.
|
|
178
|
+
|
|
179
|
+
#### Overall
|
|
180
|
+
|
|
181
|
+
| | success | false success / success claims | tool calls (median) | time s (median) | Claude tokens (median, incl. cache) | output tokens (median) | USD (median) | Jev calls (total) | focus steals |
|
|
182
|
+
|---|---|---|---|---|---|---|---|---|---|
|
|
183
|
+
| (a) cua-driver only | 23/24 | 0/22 | 10.0 | 49.7 | 646.0k | 1876 | 0.388 | 0 | 0 |
|
|
184
|
+
| (b) beans-picker | 24/24 | 0/24 | 5.5 | 28.1 | 117.9k | 962 | 0.060 | 265 | 0 |
|
|
185
|
+
|
|
186
|
+
#### Per task
|
|
187
|
+
|
|
188
|
+
| task / condition | success | false success / success claims | tool calls (median) | time s (median) | Claude tokens (median, incl. cache) | output tokens (median) | USD (median) | Jev calls (total) | focus steals |
|
|
189
|
+
|---|---|---|---|---|---|---|---|---|---|
|
|
190
|
+
| fx-name-spaces a | 2/3 | 0/1 | 13.0 | 135.6 | 1150.9k | 10704 | 0.673 | 0 | 0 |
|
|
191
|
+
| fx-name-spaces b | 3/3 | 0/3 | 3.0 | 18.5 | 69.8k | 1213 | 0.056 | 9 | 0 |
|
|
192
|
+
| fx-body-not-search a | 3/3 | 0/3 | 10.0 | 57.6 | 693.4k | 3514 | 0.437 | 0 | 0 |
|
|
193
|
+
| fx-body-not-search b | 3/3 | 0/3 | 3.0 | 12.9 | 71.3k | 690 | 0.054 | 6 | 0 |
|
|
194
|
+
| fx-size-save a | 3/3 | 0/3 | 8.0 | 27.3 | 580.9k | 1160 | 0.389 | 0 | 0 |
|
|
195
|
+
| fx-size-save b | 3/3 | 0/3 | 7.0 | 28.6 | 152.1k | 1144 | 0.084 | 10 | 0 |
|
|
196
|
+
| fx-clear-search-keep-note a | 3/3 | 0/3 | 3.0 | 12.8 | 202.0k | 489 | 0.187 | 0 | 0 |
|
|
197
|
+
| fx-clear-search-keep-note b | 3/3 | 0/3 | 3.0 | 14.3 | 70.9k | 709 | 0.053 | 5 | 0 |
|
|
198
|
+
| fx-email-fix-newsletter a | 3/3 | 0/3 | 6.0 | 37.3 | 400.6k | 1179 | 0.292 | 0 | 0 |
|
|
199
|
+
| fx-email-fix-newsletter b | 3/3 | 0/3 | 4.0 | 22.4 | 90.4k | 871 | 0.061 | 6 | 0 |
|
|
200
|
+
| calc-chain a | 3/3 | 0/3 | 14.0 | 72.0 | 874.7k | 1952 | 0.353 | 0 | 0 |
|
|
201
|
+
| calc-chain b | 3/3 | 0/3 | 9.0 | 68.1 | 185.5k | 1492 | 0.091 | 108 | 0 |
|
|
202
|
+
| calc-percent a | 3/3 | 0/3 | 13.0 | 69.9 | 1129.4k | 2467 | 0.557 | 0 | 0 |
|
|
203
|
+
| calc-percent b | 3/3 | 0/3 | 9.0 | 49.4 | 185.3k | 1690 | 0.069 | 67 | 0 |
|
|
204
|
+
| calc-negate a | 3/3 | 0/3 | 10.0 | 44.4 | 633.1k | 1424 | 0.332 | 0 | 0 |
|
|
205
|
+
| calc-negate b | 3/3 | 0/3 | 6.0 | 36.1 | 124.0k | 756 | 0.042 | 54 | 0 |
|
|
206
|
+
|
|
207
|
+
Tokens are what Claude Code reported (input + output + cache reads + cache writes); the USD columns are Claude only. Jev is billed separately at $0.042 per 1M input tokens (output free; [TypeSafe models](https://docs.typesafe.ai/models), as of 2026-09-27): 608,908 Jev input tokens ≈ $0.026. Totals including that Jev cost: (a) USD 10.50 and 1,541 s for 24 runs; (b) USD 1.62 and 820 s for 24 runs (median $0.060 per run), from 265 Jev calls.
|
|
208
|
+
|
|
209
|
+
**What this shows**
|
|
210
|
+
|
|
211
|
+
- Both conditions nearly always succeeded. (b) passed 24/24, (a) 23/24: one Name-with-spaces run failed, and the agent said so. In another (a) run the agent reported failure although the judge found the task done.
|
|
212
|
+
- Neither condition claimed success on a failed run in the counted runs. So this benchmark does **not** show that beans-picker reduces false success claims. In a pilot run before the benchmark (not counted), (a) typed `48+16÷8`, got 50 and claimed success.
|
|
213
|
+
- (b) used fewer tool calls (median 5.5 vs 10), less time (28 s vs 50 s) and about 1/5 of the Claude tokens, mostly because cua-driver's `get_window_state` returns the whole tree (and a screenshot) to the model at every step, while beans-picker returns a short candidate list and the change summary.
|
|
214
|
+
|
|
215
|
+
**Where beans-picker did not win**
|
|
216
|
+
|
|
217
|
+
- **Time on simple tasks.** `fx-size-save` (28.6 s vs 27.3 s) and `fx-clear-search-keep-note` (14.3 s vs 12.8 s) were slower with beans-picker: a Jev call plus fresh snapshots cost more than one direct click. One `calc-chain` run took 126 s and 18 tool calls with beans-picker, the slowest run of that task in either condition.
|
|
218
|
+
- **Output tokens** were higher with beans-picker on `fx-clear-search-keep-note` (709 vs 489) and about equal on `fx-size-save`.
|
|
219
|
+
- **Jev cost is small.** 265 Jev calls add ≈ $0.026 to (b)'s total ($1.62 with Jev vs $1.59 Claude-only); they are not in the USD columns above.
|
|
220
|
+
- **Pop-ups.** `choose_option` failed in all three `fx-size-save` runs: cua-driver cannot `set_value` a closed `NSPopUpButton` ("has no AX children"). The caller recovered by clicking the pop-up and then the menu item, which is why that task took 6–7 calls in both conditions. Since then `choose_option` presses the pop-up open and then the item titled exactly `text`, in the background (checked on the fixture and on a `<select>` in Chrome).
|
|
221
|
+
- **Exact text was not confirmed.** The benchmark ran without the Accessibility permission for the `beans-picker axtext` helper. Of the 242 `act` results in (b), 206 were `done`, 18 `unverified` (text entered; the judge later found it exactly right), 6 `failed` (the pop-up above), 6 `ambiguous`, 4 `not_found` and 2 `no_effect`.
|
|
222
|
+
|
|
223
|
+
**Caveats.** 3 repetitions per task and one model (Sonnet 5) are a small sample, and one person's Mac. The tasks and the fixture app were written by the same author as beans-picker. An earlier attempt at this run was stopped after 3 runs because the fixture only recorded typed input, not values set over accessibility (so it failed a correct `set_value`). The fixture was fixed to write the controls' actual values, and the 48 runs above were all made after that. The baseline is Claude Code with cua-driver only; no other agent was compared.
|
|
224
|
+
|
|
225
|
+
## License
|
|
226
|
+
|
|
227
|
+
MIT
|
|
@@ -0,0 +1,198 @@
|
|
|
1
|
+
# beans-picker
|
|
2
|
+
|
|
3
|
+
Handpicks the right control. ~2× faster, ~1/6 the cost.
|
|
4
|
+
|
|
5
|
+
Not affiliated with Cua / trycua or TypeSafe; cua-driver and Jev are separate projects this server talks to.
|
|
6
|
+
|
|
7
|
+
An MCP server that lets Claude Code or Codex pick the right control in a macOS app, act on it without bringing the app to the front, and check that it worked. It does for native apps what Stagehand does for the browser:
|
|
8
|
+
|
|
9
|
+
- **The caller thinks.** Claude Code or Codex breaks the task down, picks the next step and decides what to do when something goes wrong.
|
|
10
|
+
- **Jev picks the element.** [TypeSafe Jev](https://typesafe.ai) chooses among candidate actions built from the accessibility tree. It never writes text, so it cannot invent a target.
|
|
11
|
+
- **cua-driver acts.** [cua-driver](https://github.com/trycua/cua) clicks, types and presses keys without bringing the app to the front.
|
|
12
|
+
|
|
13
|
+
The server itself never calls a generative model (no `claude -p`, no LLM API).
|
|
14
|
+
|
|
15
|
+
## Tools
|
|
16
|
+
|
|
17
|
+
| tool | input | what it does |
|
|
18
|
+
|---|---|---|
|
|
19
|
+
| `observe` | `app` \| `pid`, `windowId?`, `instruction?`, `limit?` | Lists the window's candidate actions (`id`, `kind`, what it does). With `instruction`, Jev ranks them and each comes with its probability `p`. |
|
|
20
|
+
| `act` | `app` \| `pid`, `windowId?`, `instruction`, `text?`, `candidateId?`, `allowDestructive?`, `modifiers?`, `then?` | Performs one action and checks its effect, then each step in `then` (same fields, up to 12) the same way, stopping at the first that is not `done` or `unverified`; each step's result is in `steps`. Jev picks the action for `instruction`, unless `candidateId` is given. `text` is entered exactly as given. `modifiers` (`shift`, `cmd`, `option`, `ctrl`) are held during a click or toggle, as a pixel click on a control visible on the window. |
|
|
21
|
+
| `extract` | `app` \| `pid`, `windowId?`, `instruction` | Returns the text or value of the element Jev picks, exactly as read. A table or list with no value of its own comes back as `rows` (each row's texts), any other container as `text`. Its `status` is `done`, `ambiguous` (the shortlist comes back with each element's value) or `not_found`. |
|
|
22
|
+
|
|
23
|
+
`act` returns a `status`:
|
|
24
|
+
|
|
25
|
+
| status | meaning |
|
|
26
|
+
|---|---|
|
|
27
|
+
| `done` | The effect asked for is observed. For text, the field reads **exactly** the expected text (see "Exact checks"). |
|
|
28
|
+
| `unverified` | The field changed, but its exact text could not be read (the helper has no Accessibility permission). |
|
|
29
|
+
| `no_effect` | The window's signature did not change over `BEANS_PICKER_EFFECT_RETAKES` fresh snapshots (default 5). Decided by count, not by waiting. |
|
|
30
|
+
| `mismatch` | Something changed, but not what was asked (for example the text landed in another field, or a trailing space was lost). |
|
|
31
|
+
| `ambiguous` | Jev had no clear leader. The top candidates are returned; call `act` again with the right `candidateId`. |
|
|
32
|
+
| `needs_confirmation` | The action may not be undoable (delete, close, send, quit …). Call again with `candidateId` and `allowDestructive: true`. |
|
|
33
|
+
| `not_found` | No candidate fits, or the `candidateId` is not on the window any more. |
|
|
34
|
+
| `failed` | cua-driver refused, the command needs the foreground, or the app came to the front (`foreground_violation`). |
|
|
35
|
+
|
|
36
|
+
A call that cannot run at all returns `isError` with `{"status": "failed", "code", "message"}`. Codes include `bad_target`, `window_not_found`, `screen_locked`, `jev_unavailable` (no key, Jev unreachable or too slow), `jev_bad_response`, `driver_unavailable` (cua-driver could not be started), `driver_timeout`, `driver_error` (cua-driver refused a call the tool needed) and `internal`. Anything but a tool's own refusal is also logged with its traceback at `WARNING`.
|
|
37
|
+
|
|
38
|
+
Each result also carries the action that ran (`action`, with the cua-driver route) and a summary of what changed on the window (`change`).
|
|
39
|
+
|
|
40
|
+
Candidate ids are derived from a control's stable identity (role, identifier, label and ancestors), not from the snapshot's element index, so an id from `observe` still works after the window changes, as long as the control is there.
|
|
41
|
+
|
|
42
|
+
## Install
|
|
43
|
+
|
|
44
|
+
Requirements: macOS, Python 3.12+, [cua-driver](https://github.com/trycua/cua) at `~/.local/bin/cua-driver` with its Accessibility and Screen Recording permissions, a Jev API key, and Xcode Command Line Tools (`clang`) for the two small native helpers, which are compiled on first use into `~/Library/Caches/beans-picker`.
|
|
45
|
+
|
|
46
|
+
```sh
|
|
47
|
+
uv tool install beans-picker
|
|
48
|
+
mkdir -p ~/.config/beans-picker && echo 'JEV_API_KEY=...' > ~/.config/beans-picker/.env.local
|
|
49
|
+
claude mcp add beans-picker -- "$(command -v beans-picker)" # absolute path: GUI-launched clients often lack ~/.local/bin on PATH
|
|
50
|
+
beans-picker grant-ax # exact text (optional, see below)
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
Without installing, let `uvx` fetch it on each start: `claude mcp add beans-picker -- "$(command -v uvx)" beans-picker`, and `uvx beans-picker grant-ax`.
|
|
54
|
+
|
|
55
|
+
For Codex, add the same absolute path as an MCP server in `~/.codex/config.toml` (`command -v beans-picker` prints it):
|
|
56
|
+
|
|
57
|
+
```toml
|
|
58
|
+
[mcp_servers.beans-picker]
|
|
59
|
+
command = "/Users/you/.local/bin/beans-picker"
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
From a checkout:
|
|
63
|
+
|
|
64
|
+
```sh
|
|
65
|
+
uv sync
|
|
66
|
+
echo 'JEV_API_KEY=...' > .env.local # read from the checkout, whatever the caller's cwd
|
|
67
|
+
claude mcp add beans-picker -- "$(command -v uv)" run --directory "$PWD" beans-picker
|
|
68
|
+
uv run beans-picker grant-ax
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
`beans-picker grant-ax` builds the read-only helper and asks macOS to list it; then turn on **beans-picker axtext** under System Settings > Privacy & Security > Accessibility. The helper is built from its source, so an upgrade that changes that source builds a new helper that macOS has not been told about: run `beans-picker grant-ax` again after upgrading if text results turn `unverified`. `beans-picker --version` prints the version and `beans-picker --help` the usage.
|
|
72
|
+
|
|
73
|
+
The agent skill in `skills/beans-picker` tells a model how to use these tools well: when cua-driver fits better, what to do with each `act` status, and recipes for combo boxes, number fields, long tables and checking the result. Link it where your agent looks for skills, e.g. `ln -s "$PWD/skills/beans-picker" ~/.claude/skills/beans-picker` (Codex: `~/.codex/skills`). Leave it out when benchmarking the bare tools.
|
|
74
|
+
|
|
75
|
+
## Configuration
|
|
76
|
+
|
|
77
|
+
| variable | meaning |
|
|
78
|
+
|---|---|
|
|
79
|
+
| `JEV_API_KEY` | Required for `observe` with an instruction, `act` without `candidateId`, and `extract`. `TYPESAFE_API_KEY` is used when it is not set or empty. |
|
|
80
|
+
| `BEANS_PICKER_MODEL` | The Jev model (default `jev-latest`). |
|
|
81
|
+
| `CUA_DRIVER_BIN` | Path to cua-driver (default `~/.local/bin/cua-driver`; a leading `~` is expanded). |
|
|
82
|
+
| `CUA_DRIVER_TIMEOUT` | Seconds to wait for cua-driver to start, and for its answer to each call (default 120). |
|
|
83
|
+
| `JEV_CONNECT_TIMEOUT` | Seconds to wait for a connection to Jev (default 10). |
|
|
84
|
+
| `JEV_READ_TIMEOUT` | Seconds to wait for Jev's answer to one request (default 120). |
|
|
85
|
+
| `BEANS_PICKER_EFFECT_RETAKES` | Fresh snapshots taken to see an effect (default 5). |
|
|
86
|
+
| `BEANS_PICKER_LOG_LEVEL` | Level of the server's log on stderr (default `WARNING`). |
|
|
87
|
+
| `TYPESAFE_BASE_URL` | The Jev endpoint (default `https://api.typesafe.ai`). |
|
|
88
|
+
|
|
89
|
+
Variables are read from the environment first, then from `.env.local` and `.env` in the checkout (only when the package runs from a beans-picker checkout), then from `.env.local` and `.env` in `~/.config/beans-picker` (`$XDG_CONFIG_HOME/beans-picker` when that is set). A value that is already set is never overridden. See `.env.example`.
|
|
90
|
+
|
|
91
|
+
## Rules the server keeps
|
|
92
|
+
|
|
93
|
+
- **Background only.** `bring_to_front`, `invoke_menu`, `move_cursor` and `delivery_mode: "foreground"` are refused before they reach cua-driver. Menu commands run as their keyboard shortcut sent to the app's pid; a command without a known shortcut returns `failed` / `foreground_required`. During `act`, the frontmost app is sampled continuously; if the target app comes to the front, the call stops with `foreground_violation`.
|
|
94
|
+
- **Time limits.** Every request to cua-driver and to Jev has a generous limit (see Configuration), so one that stalls ends its call with `driver_timeout` or `jev_unavailable` instead of blocking the queue. An action whose answer timed out may still have happened: observe before repeating it. At shutdown, the final `end_session` gets 5 seconds.
|
|
95
|
+
- **Screen lock.** While `CGSSessionScreenIsLocked` is set, every tool refuses with `screen_locked` and does nothing.
|
|
96
|
+
- **One call at a time.** Tool calls are queued, so two actions never interleave on the desktop.
|
|
97
|
+
- **Destructive actions** are marked, not hidden: `act` returns `needs_confirmation` unless `allowDestructive` is set, whoever picked the action. A checkbox is never destructive, since it can be switched back.
|
|
98
|
+
|
|
99
|
+
## How `act` works
|
|
100
|
+
|
|
101
|
+
```
|
|
102
|
+
snapshot ─► candidates ─► Jev picks (or candidateId) ─► gate ─► cua-driver ─► fresh snapshots ─► status
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
1. **Snapshot** (`src/beans_picker/observe`). `get_window_state` returns the AX tree. The structured elements are joined with the tree markdown, which holds identifiers, help text and static text. The menu bar is split off, and a signature of what matters (title, field values, visible text, other windows) is computed.
|
|
106
|
+
2. **Candidates** (`src/beans_picker/candidates`). Clicks, toggles, text entry (`set_value`, `type_into` at the caret, `append` at the end), pop-up choices, an on-screen keypad sequence for `text` ("12×7="), menu commands with their background shortcut, one step up or down on a slider or number field (an arrow key sent to it), the options of a web list box (a combo box's suggestions), a page down or up on a scroll view, table, list or web page, a named row's, link's or image's context menu (AXShowMenu, which a web page receives as a right-click), and Return / Escape / Tab / Space / the arrow keys / Shift+F10. With `text`, only the actions that enter text are in the running.
|
|
107
|
+
3. **Jev** (`src/beans_picker/jev`). One `system_one` request asks two choice questions over the same candidates: one with a `none` option, one forced. More than 60 candidates are sharded (best lexical match first), and the leaders of each shard go to a runoff. The gate acts on a leader at p ≥ 0.8, or at p ≥ 0.5 when the forced question agrees and the leader has twice the runner-up's probability. Otherwise the result is `ambiguous` (or `not_found` when `none` dominates).
|
|
108
|
+
4. **cua-driver** (`src/beans_picker/act`). AX press, pixel clicks for keypads, a pop-up pressed open and then its item pressed (a menu left open over the page, as Chrome's is, is closed with Escape), `set_value` / `type_text` (a web page's number field ignores an AXValue write, so it is retyped: End, ⇧Home, then `type_text`; once an app's page has ignored a write, its text fields are typed into straight away), ⌘↓ then `type_text` for `append`, a wheel event at the area for a scroll, and the menu shortcut as a pid-targeted hotkey. Stale element tokens are rebound by stable key.
|
|
109
|
+
5. **Effect** (`src/beans_picker/verify/effect.py`). Fresh snapshots are taken, up to the retake count, until the effect shows.
|
|
110
|
+
|
|
111
|
+
## Exact checks
|
|
112
|
+
|
|
113
|
+
A false success is the worst failure a desktop agent can have: text judged by `trim` or by "contains", a search field mistaken for the document, a trailing space silently lost. `act` therefore checks text by **exact equality, in the targeted field only**:
|
|
114
|
+
|
|
115
|
+
- `set_value`: the field reads exactly `text`.
|
|
116
|
+
- `append`: the field reads exactly its previous text followed by `text`.
|
|
117
|
+
- `type_into`: the field reads its previous text with `text` inserted whole at one position.
|
|
118
|
+
|
|
119
|
+
cua-driver 0.8 does not give the exact text: it trims whitespace at both ends of a value, shows an empty field's placeholder as its value, and leaves out a checkbox's state. The read-only helper `src/beans_picker/native/axtext.m` reads AXValue as it is. It needs the Accessibility permission, which is why it runs as its own tiny background app (`beans-picker axtext`) that you can allow on its own, without granting anything to your terminal. Without it, a text action that visibly changed the field returns `unverified`, never `done`. A checkbox whose state is not in the tree is judged by its own pixels, captured in the background before and after the click.
|
|
120
|
+
|
|
121
|
+
## Privacy
|
|
122
|
+
|
|
123
|
+
Window text leaves the machine only when a step needs Jev. [SECURITY.md](https://github.com/mimo-3/beans-picker/blob/main/SECURITY.md) lists exactly what is sent, what runs locally, and how to report a vulnerability.
|
|
124
|
+
|
|
125
|
+
## Development
|
|
126
|
+
|
|
127
|
+
```sh
|
|
128
|
+
uv sync
|
|
129
|
+
uv run ruff check
|
|
130
|
+
uv run ruff format --check
|
|
131
|
+
uv run mypy
|
|
132
|
+
uv run pytest
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
The benchmark is in `bench/` (see below). `bench/fixture-app` is a small AppKit window used only by the benchmark; `sh bench/fixture-app/build.sh` builds it into `~/Library/Caches/beans-picker/fixture`. `python -m bench.run`, `python -m bench.judge` and `python -m bench.summarize` run from the repository root. See [CONTRIBUTING.md](https://github.com/mimo-3/beans-picker/blob/main/CONTRIBUTING.md).
|
|
136
|
+
|
|
137
|
+
## Benchmark
|
|
138
|
+
|
|
139
|
+
Run on 2026-09-24 with an unpublished pre-release of beans-picker (0.2.0) and `--reps 3`: 8 tasks × 2 conditions × 3 repetitions = 48 runs, Claude Code headless (`claude -p`) with `--model sonnet`, one run at a time. Raw records: [`bench/results/runs.jsonl`](https://github.com/mimo-3/beans-picker/blob/main/bench/results/runs.jsonl); the tables below are `python -m bench.summarize` of them.
|
|
140
|
+
|
|
141
|
+
**Conditions.** Both get the same prompt (task, target pid and window id, "work in the background", and a final `RESULT: success|failure` line), no built-in tools (`--tools ""`) and only one MCP server:
|
|
142
|
+
|
|
143
|
+
- **(a) cua-driver only**: `cua-driver mcp`, with `bring_to_front`, `move_cursor`, `kill_app`, `get_desktop_state` and a few other tools denied.
|
|
144
|
+
- **(b) beans-picker**: this server only.
|
|
145
|
+
|
|
146
|
+
**Tasks** ([`bench/tasks.json`](https://github.com/mimo-3/beans-picker/blob/main/bench/tasks.json)). Five run on `bench/fixture-app`, a small AppKit window made for the bench (so no user document is touched): a Name with leading and trailing spaces, appending to a note body next to a search field, a pop-up and Save, clearing a search field next to a destructive "Delete note" button, and fixing an email plus a checkbox. Three run on Calculator: `(48 + 16) / 8`, 15% of 80, and `7 − 19` then change sign. TextEdit and Notes were not used.
|
|
147
|
+
|
|
148
|
+
**Judging.** Success is decided only by [`bench/judge.py`](https://github.com/mimo-3/beans-picker/blob/main/bench/judge.py), a separate script that compares the final state with the task's `expected` JSON by exact equality: every key of the fixture's state file (the controls' values, written by the app itself), or Calculator's display read over accessibility (bidi marks removed, nothing else). The agent's `RESULT:` line is used only to count false success claims. The front app was sampled every 200 ms during each run to count focus steals.
|
|
149
|
+
|
|
150
|
+
#### Overall
|
|
151
|
+
|
|
152
|
+
| | success | false success / success claims | tool calls (median) | time s (median) | Claude tokens (median, incl. cache) | output tokens (median) | USD (median) | Jev calls (total) | focus steals |
|
|
153
|
+
|---|---|---|---|---|---|---|---|---|---|
|
|
154
|
+
| (a) cua-driver only | 23/24 | 0/22 | 10.0 | 49.7 | 646.0k | 1876 | 0.388 | 0 | 0 |
|
|
155
|
+
| (b) beans-picker | 24/24 | 0/24 | 5.5 | 28.1 | 117.9k | 962 | 0.060 | 265 | 0 |
|
|
156
|
+
|
|
157
|
+
#### Per task
|
|
158
|
+
|
|
159
|
+
| task / condition | success | false success / success claims | tool calls (median) | time s (median) | Claude tokens (median, incl. cache) | output tokens (median) | USD (median) | Jev calls (total) | focus steals |
|
|
160
|
+
|---|---|---|---|---|---|---|---|---|---|
|
|
161
|
+
| fx-name-spaces a | 2/3 | 0/1 | 13.0 | 135.6 | 1150.9k | 10704 | 0.673 | 0 | 0 |
|
|
162
|
+
| fx-name-spaces b | 3/3 | 0/3 | 3.0 | 18.5 | 69.8k | 1213 | 0.056 | 9 | 0 |
|
|
163
|
+
| fx-body-not-search a | 3/3 | 0/3 | 10.0 | 57.6 | 693.4k | 3514 | 0.437 | 0 | 0 |
|
|
164
|
+
| fx-body-not-search b | 3/3 | 0/3 | 3.0 | 12.9 | 71.3k | 690 | 0.054 | 6 | 0 |
|
|
165
|
+
| fx-size-save a | 3/3 | 0/3 | 8.0 | 27.3 | 580.9k | 1160 | 0.389 | 0 | 0 |
|
|
166
|
+
| fx-size-save b | 3/3 | 0/3 | 7.0 | 28.6 | 152.1k | 1144 | 0.084 | 10 | 0 |
|
|
167
|
+
| fx-clear-search-keep-note a | 3/3 | 0/3 | 3.0 | 12.8 | 202.0k | 489 | 0.187 | 0 | 0 |
|
|
168
|
+
| fx-clear-search-keep-note b | 3/3 | 0/3 | 3.0 | 14.3 | 70.9k | 709 | 0.053 | 5 | 0 |
|
|
169
|
+
| fx-email-fix-newsletter a | 3/3 | 0/3 | 6.0 | 37.3 | 400.6k | 1179 | 0.292 | 0 | 0 |
|
|
170
|
+
| fx-email-fix-newsletter b | 3/3 | 0/3 | 4.0 | 22.4 | 90.4k | 871 | 0.061 | 6 | 0 |
|
|
171
|
+
| calc-chain a | 3/3 | 0/3 | 14.0 | 72.0 | 874.7k | 1952 | 0.353 | 0 | 0 |
|
|
172
|
+
| calc-chain b | 3/3 | 0/3 | 9.0 | 68.1 | 185.5k | 1492 | 0.091 | 108 | 0 |
|
|
173
|
+
| calc-percent a | 3/3 | 0/3 | 13.0 | 69.9 | 1129.4k | 2467 | 0.557 | 0 | 0 |
|
|
174
|
+
| calc-percent b | 3/3 | 0/3 | 9.0 | 49.4 | 185.3k | 1690 | 0.069 | 67 | 0 |
|
|
175
|
+
| calc-negate a | 3/3 | 0/3 | 10.0 | 44.4 | 633.1k | 1424 | 0.332 | 0 | 0 |
|
|
176
|
+
| calc-negate b | 3/3 | 0/3 | 6.0 | 36.1 | 124.0k | 756 | 0.042 | 54 | 0 |
|
|
177
|
+
|
|
178
|
+
Tokens are what Claude Code reported (input + output + cache reads + cache writes); the USD columns are Claude only. Jev is billed separately at $0.042 per 1M input tokens (output free; [TypeSafe models](https://docs.typesafe.ai/models), as of 2026-09-27): 608,908 Jev input tokens ≈ $0.026. Totals including that Jev cost: (a) USD 10.50 and 1,541 s for 24 runs; (b) USD 1.62 and 820 s for 24 runs (median $0.060 per run), from 265 Jev calls.
|
|
179
|
+
|
|
180
|
+
**What this shows**
|
|
181
|
+
|
|
182
|
+
- Both conditions nearly always succeeded. (b) passed 24/24, (a) 23/24: one Name-with-spaces run failed, and the agent said so. In another (a) run the agent reported failure although the judge found the task done.
|
|
183
|
+
- Neither condition claimed success on a failed run in the counted runs. So this benchmark does **not** show that beans-picker reduces false success claims. In a pilot run before the benchmark (not counted), (a) typed `48+16÷8`, got 50 and claimed success.
|
|
184
|
+
- (b) used fewer tool calls (median 5.5 vs 10), less time (28 s vs 50 s) and about 1/5 of the Claude tokens, mostly because cua-driver's `get_window_state` returns the whole tree (and a screenshot) to the model at every step, while beans-picker returns a short candidate list and the change summary.
|
|
185
|
+
|
|
186
|
+
**Where beans-picker did not win**
|
|
187
|
+
|
|
188
|
+
- **Time on simple tasks.** `fx-size-save` (28.6 s vs 27.3 s) and `fx-clear-search-keep-note` (14.3 s vs 12.8 s) were slower with beans-picker: a Jev call plus fresh snapshots cost more than one direct click. One `calc-chain` run took 126 s and 18 tool calls with beans-picker, the slowest run of that task in either condition.
|
|
189
|
+
- **Output tokens** were higher with beans-picker on `fx-clear-search-keep-note` (709 vs 489) and about equal on `fx-size-save`.
|
|
190
|
+
- **Jev cost is small.** 265 Jev calls add ≈ $0.026 to (b)'s total ($1.62 with Jev vs $1.59 Claude-only); they are not in the USD columns above.
|
|
191
|
+
- **Pop-ups.** `choose_option` failed in all three `fx-size-save` runs: cua-driver cannot `set_value` a closed `NSPopUpButton` ("has no AX children"). The caller recovered by clicking the pop-up and then the menu item, which is why that task took 6–7 calls in both conditions. Since then `choose_option` presses the pop-up open and then the item titled exactly `text`, in the background (checked on the fixture and on a `<select>` in Chrome).
|
|
192
|
+
- **Exact text was not confirmed.** The benchmark ran without the Accessibility permission for the `beans-picker axtext` helper. Of the 242 `act` results in (b), 206 were `done`, 18 `unverified` (text entered; the judge later found it exactly right), 6 `failed` (the pop-up above), 6 `ambiguous`, 4 `not_found` and 2 `no_effect`.
|
|
193
|
+
|
|
194
|
+
**Caveats.** 3 repetitions per task and one model (Sonnet 5) are a small sample, and one person's Mac. The tasks and the fixture app were written by the same author as beans-picker. An earlier attempt at this run was stopped after 3 runs because the fixture only recorded typed input, not values set over accessibility (so it failed a correct `set_value`). The fixture was fixed to write the controls' actual values, and the 48 runs above were all made after that. The baseline is Claude Code with cua-driver only; no other agent was compared.
|
|
195
|
+
|
|
196
|
+
## License
|
|
197
|
+
|
|
198
|
+
MIT
|