beans-picker 0.3.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (152) hide show
  1. beans_picker-0.3.0/.env.example +6 -0
  2. beans_picker-0.3.0/.gitignore +37 -0
  3. beans_picker-0.3.0/CHANGELOG.md +51 -0
  4. beans_picker-0.3.0/CONTRIBUTING.md +59 -0
  5. beans_picker-0.3.0/LICENSE +21 -0
  6. beans_picker-0.3.0/PKG-INFO +227 -0
  7. beans_picker-0.3.0/README.md +198 -0
  8. beans_picker-0.3.0/SECURITY.md +49 -0
  9. beans_picker-0.3.0/bench/__init__.py +0 -0
  10. beans_picker-0.3.0/bench/fixture-app/build.sh +19 -0
  11. beans_picker-0.3.0/bench/fixture-app/main.m +183 -0
  12. beans_picker-0.3.0/bench/judge.py +145 -0
  13. beans_picker-0.3.0/bench/results/runs.jsonl +48 -0
  14. beans_picker-0.3.0/bench/results/summary.md +27 -0
  15. beans_picker-0.3.0/bench/run.py +584 -0
  16. beans_picker-0.3.0/bench/summarize.py +108 -0
  17. beans_picker-0.3.0/bench/tasks.json +61 -0
  18. beans_picker-0.3.0/pyproject.toml +118 -0
  19. beans_picker-0.3.0/scripts/check_wheel.py +36 -0
  20. beans_picker-0.3.0/scripts/smoke_installed.py +66 -0
  21. beans_picker-0.3.0/skills/beans-picker/SKILL.md +117 -0
  22. beans_picker-0.3.0/src/beans_picker/__init__.py +1 -0
  23. beans_picker-0.3.0/src/beans_picker/__main__.py +6 -0
  24. beans_picker-0.3.0/src/beans_picker/_aio.py +32 -0
  25. beans_picker-0.3.0/src/beans_picker/_json.py +101 -0
  26. beans_picker-0.3.0/src/beans_picker/_numbers.py +116 -0
  27. beans_picker-0.3.0/src/beans_picker/_proc.py +92 -0
  28. beans_picker-0.3.0/src/beans_picker/_stdin.py +66 -0
  29. beans_picker-0.3.0/src/beans_picker/_text.py +97 -0
  30. beans_picker-0.3.0/src/beans_picker/act/__init__.py +0 -0
  31. beans_picker-0.3.0/src/beans_picker/act/execute.py +487 -0
  32. beans_picker-0.3.0/src/beans_picker/act/pixel.py +204 -0
  33. beans_picker-0.3.0/src/beans_picker/candidates/__init__.py +0 -0
  34. beans_picker-0.3.0/src/beans_picker/candidates/build.py +340 -0
  35. beans_picker-0.3.0/src/beans_picker/candidates/describe.py +91 -0
  36. beans_picker-0.3.0/src/beans_picker/candidates/keypad.py +117 -0
  37. beans_picker-0.3.0/src/beans_picker/candidates/menu.py +61 -0
  38. beans_picker-0.3.0/src/beans_picker/candidates/prune.py +77 -0
  39. beans_picker-0.3.0/src/beans_picker/candidates/safety.py +161 -0
  40. beans_picker-0.3.0/src/beans_picker/candidates/types.py +68 -0
  41. beans_picker-0.3.0/src/beans_picker/cli.py +83 -0
  42. beans_picker-0.3.0/src/beans_picker/config.py +170 -0
  43. beans_picker-0.3.0/src/beans_picker/driver/__init__.py +0 -0
  44. beans_picker-0.3.0/src/beans_picker/driver/app.py +143 -0
  45. beans_picker-0.3.0/src/beans_picker/driver/lock.py +22 -0
  46. beans_picker-0.3.0/src/beans_picker/driver/markdown.py +108 -0
  47. beans_picker-0.3.0/src/beans_picker/driver/mcp.py +374 -0
  48. beans_picker-0.3.0/src/beans_picker/driver/sentinel.py +140 -0
  49. beans_picker-0.3.0/src/beans_picker/driver/types.py +366 -0
  50. beans_picker-0.3.0/src/beans_picker/errors.py +96 -0
  51. beans_picker-0.3.0/src/beans_picker/jev/__init__.py +0 -0
  52. beans_picker-0.3.0/src/beans_picker/jev/client.py +213 -0
  53. beans_picker-0.3.0/src/beans_picker/jev/questions.py +55 -0
  54. beans_picker-0.3.0/src/beans_picker/jev/rank.py +143 -0
  55. beans_picker-0.3.0/src/beans_picker/jev/state.py +225 -0
  56. beans_picker-0.3.0/src/beans_picker/log.py +48 -0
  57. beans_picker-0.3.0/src/beans_picker/menus/__init__.py +0 -0
  58. beans_picker-0.3.0/src/beans_picker/menus/keyequiv.py +194 -0
  59. beans_picker-0.3.0/src/beans_picker/menus/menukeys.py +288 -0
  60. beans_picker-0.3.0/src/beans_picker/native/__init__.py +0 -0
  61. beans_picker-0.3.0/src/beans_picker/native/axtext.m +81 -0
  62. beans_picker-0.3.0/src/beans_picker/native/menukeys.m +94 -0
  63. beans_picker-0.3.0/src/beans_picker/observe/__init__.py +0 -0
  64. beans_picker-0.3.0/src/beans_picker/observe/exacttext.py +210 -0
  65. beans_picker-0.3.0/src/beans_picker/observe/helpers.py +179 -0
  66. beans_picker-0.3.0/src/beans_picker/observe/identity.py +32 -0
  67. beans_picker-0.3.0/src/beans_picker/observe/menudiff.py +105 -0
  68. beans_picker-0.3.0/src/beans_picker/observe/normalize.py +58 -0
  69. beans_picker-0.3.0/src/beans_picker/observe/png.py +164 -0
  70. beans_picker-0.3.0/src/beans_picker/observe/signature.py +29 -0
  71. beans_picker-0.3.0/src/beans_picker/observe/snapshot.py +394 -0
  72. beans_picker-0.3.0/src/beans_picker/observe/types.py +107 -0
  73. beans_picker-0.3.0/src/beans_picker/observe/visual.py +179 -0
  74. beans_picker-0.3.0/src/beans_picker/paths.py +31 -0
  75. beans_picker-0.3.0/src/beans_picker/py.typed +0 -0
  76. beans_picker-0.3.0/src/beans_picker/server.py +192 -0
  77. beans_picker-0.3.0/src/beans_picker/tools/__init__.py +0 -0
  78. beans_picker-0.3.0/src/beans_picker/tools/act.py +371 -0
  79. beans_picker-0.3.0/src/beans_picker/tools/args.py +339 -0
  80. beans_picker-0.3.0/src/beans_picker/tools/extract.py +230 -0
  81. beans_picker-0.3.0/src/beans_picker/tools/observe.py +65 -0
  82. beans_picker-0.3.0/src/beans_picker/tools/present.py +69 -0
  83. beans_picker-0.3.0/src/beans_picker/tools/session.py +250 -0
  84. beans_picker-0.3.0/src/beans_picker/verify/__init__.py +0 -0
  85. beans_picker-0.3.0/src/beans_picker/verify/effect.py +168 -0
  86. beans_picker-0.3.0/tests/__init__.py +0 -0
  87. beans_picker-0.3.0/tests/act_support.py +209 -0
  88. beans_picker-0.3.0/tests/conftest.py +33 -0
  89. beans_picker-0.3.0/tests/fake_driver.py +124 -0
  90. beans_picker-0.3.0/tests/fakes.py +104 -0
  91. beans_picker-0.3.0/tests/fixtures/calculator.window-state.json +2446 -0
  92. beans_picker-0.3.0/tests/fixtures/initialize.json +13 -0
  93. beans_picker-0.3.0/tests/fixtures/textedit.window-state.json +2563 -0
  94. beans_picker-0.3.0/tests/fixtures/tools_list.json +180 -0
  95. beans_picker-0.3.0/tests/fixtures/validation_errors.jsonl +47 -0
  96. beans_picker-0.3.0/tests/helpers.py +88 -0
  97. beans_picker-0.3.0/tests/raw_mcp.py +76 -0
  98. beans_picker-0.3.0/tests/test_act.py +338 -0
  99. beans_picker-0.3.0/tests/test_act_tool_paths.py +302 -0
  100. beans_picker-0.3.0/tests/test_aio.py +68 -0
  101. beans_picker-0.3.0/tests/test_app.py +294 -0
  102. beans_picker-0.3.0/tests/test_args.py +133 -0
  103. beans_picker-0.3.0/tests/test_bench.py +722 -0
  104. beans_picker-0.3.0/tests/test_candidates.py +1047 -0
  105. beans_picker-0.3.0/tests/test_cli.py +80 -0
  106. beans_picker-0.3.0/tests/test_client.py +250 -0
  107. beans_picker-0.3.0/tests/test_config.py +208 -0
  108. beans_picker-0.3.0/tests/test_desktop.py +207 -0
  109. beans_picker-0.3.0/tests/test_driver_mcp.py +465 -0
  110. beans_picker-0.3.0/tests/test_driver_types.py +182 -0
  111. beans_picker-0.3.0/tests/test_e2e.py +172 -0
  112. beans_picker-0.3.0/tests/test_effect.py +90 -0
  113. beans_picker-0.3.0/tests/test_errors.py +79 -0
  114. beans_picker-0.3.0/tests/test_exacttext.py +301 -0
  115. beans_picker-0.3.0/tests/test_execute.py +573 -0
  116. beans_picker-0.3.0/tests/test_extract.py +186 -0
  117. beans_picker-0.3.0/tests/test_golden.py +81 -0
  118. beans_picker-0.3.0/tests/test_helpers.py +212 -0
  119. beans_picker-0.3.0/tests/test_identity.py +30 -0
  120. beans_picker-0.3.0/tests/test_json.py +55 -0
  121. beans_picker-0.3.0/tests/test_judge.py +220 -0
  122. beans_picker-0.3.0/tests/test_lock.py +46 -0
  123. beans_picker-0.3.0/tests/test_log.py +73 -0
  124. beans_picker-0.3.0/tests/test_markdown.py +164 -0
  125. beans_picker-0.3.0/tests/test_menudiff.py +149 -0
  126. beans_picker-0.3.0/tests/test_menukeys.py +383 -0
  127. beans_picker-0.3.0/tests/test_native_build.py +24 -0
  128. beans_picker-0.3.0/tests/test_native_sources.py +59 -0
  129. beans_picker-0.3.0/tests/test_normalize.py +94 -0
  130. beans_picker-0.3.0/tests/test_numbers.py +144 -0
  131. beans_picker-0.3.0/tests/test_output_shapes.py +105 -0
  132. beans_picker-0.3.0/tests/test_paths.py +18 -0
  133. beans_picker-0.3.0/tests/test_pixel.py +139 -0
  134. beans_picker-0.3.0/tests/test_png.py +134 -0
  135. beans_picker-0.3.0/tests/test_proc.py +98 -0
  136. beans_picker-0.3.0/tests/test_rank.py +179 -0
  137. beans_picker-0.3.0/tests/test_sentinel.py +270 -0
  138. beans_picker-0.3.0/tests/test_server.py +204 -0
  139. beans_picker-0.3.0/tests/test_server_cancel.py +72 -0
  140. beans_picker-0.3.0/tests/test_session.py +321 -0
  141. beans_picker-0.3.0/tests/test_signature.py +99 -0
  142. beans_picker-0.3.0/tests/test_snapshot.py +660 -0
  143. beans_picker-0.3.0/tests/test_source_rules.py +95 -0
  144. beans_picker-0.3.0/tests/test_state.py +210 -0
  145. beans_picker-0.3.0/tests/test_stdin.py +32 -0
  146. beans_picker-0.3.0/tests/test_stdio.py +111 -0
  147. beans_picker-0.3.0/tests/test_summarize.py +109 -0
  148. beans_picker-0.3.0/tests/test_text.py +116 -0
  149. beans_picker-0.3.0/tests/test_verify.py +208 -0
  150. beans_picker-0.3.0/tests/test_visual.py +322 -0
  151. beans_picker-0.3.0/tests/tool_fakes.py +144 -0
  152. beans_picker-0.3.0/uv.lock +1149 -0
@@ -0,0 +1,6 @@
1
+ JEV_API_KEY=
2
+ # BEANS_PICKER_MODEL=jev-1.13.0
3
+ # CUA_DRIVER_BIN=~/.local/bin/cua-driver
4
+ # CUA_DRIVER_TIMEOUT=120
5
+ # JEV_CONNECT_TIMEOUT=10
6
+ # JEV_READ_TIMEOUT=120
@@ -0,0 +1,37 @@
1
+ # Environments and secrets
2
+ .venv/
3
+ .env*
4
+ !.env.example
5
+
6
+ # Build output
7
+ dist/
8
+ dist-from-sdist/
9
+ build/
10
+ *.egg-info/
11
+
12
+ # Caches
13
+ __pycache__/
14
+ *.py[cod]
15
+ .pytest_cache/
16
+ .mypy_cache/
17
+ .ruff_cache/
18
+ .coverage
19
+ .coverage.*
20
+ coverage.xml
21
+ htmlcov/
22
+
23
+ # Bench scratch space and built fixture app
24
+ bench/.work/
25
+ bench/fixture-app/*.app
26
+
27
+ # macOS
28
+ .DS_Store
29
+
30
+ # Local agent instructions and settings
31
+ CLAUDE.md
32
+ CLAUDE.local.md
33
+ AGENTS.md
34
+ GEMINI.md
35
+ .claude/
36
+ .cursor/
37
+ .cursorrules
@@ -0,0 +1,51 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented here. The format follows
4
+ [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and versions follow
5
+ [Semantic Versioning](https://semver.org/).
6
+
7
+ ## [Unreleased]
8
+
9
+ To be released as 0.3.0, the first release on PyPI. Versions 0.1 and 0.2 were never published.
10
+
11
+ ### Added
12
+
13
+ - The `beans-picker` command and `python -m beans_picker`: an MCP server over stdio with the tools `observe`,
14
+ `act` and `extract`. It runs on macOS; on other systems it exits with an error.
15
+ - `beans-picker --version` and `beans-picker --help`, which work on every system.
16
+ - `BEANS_PICKER_LOG_LEVEL` sets the level of the server's log on stderr (default `WARNING`).
17
+ - The Jev key and other settings are read from the environment, then from `.env.local` / `.env`
18
+ in the checkout when the package runs from a beans-picker checkout, then from `~/.config/beans-picker/`
19
+ (`$XDG_CONFIG_HOME/beans-picker`), so an installed server finds its key without a checkout.
20
+ - Time limits on every request: `CUA_DRIVER_TIMEOUT` (default 120 s) for cua-driver's startup and
21
+ each call, `JEV_CONNECT_TIMEOUT` (10 s) and `JEV_READ_TIMEOUT` (120 s) for Jev. A request that
22
+ runs out of time fails with `driver_timeout` or `jev_unavailable` instead of blocking later calls.
23
+ At shutdown, the final `end_session` to cua-driver waits at most 5 seconds.
24
+
25
+ ### Changed
26
+
27
+ - Failures that are not a tool's own refusal have their own codes: `driver_unavailable`,
28
+ `driver_timeout`, `driver_error` and `jev_bad_response`, besides `internal`. They are logged with
29
+ their traceback at `WARNING`.
30
+ - A call to an unknown tool is a JSON-RPC error (-32602), not a tool result.
31
+ - `pid` and `windowId` must be at least 1.
32
+ - Unknown command-line arguments print the usage and exit with status 2 instead of starting the
33
+ server.
34
+ - An empty `JEV_API_KEY` no longer hides `TYPESAFE_API_KEY`.
35
+ - The exact-text helper also reads switches (`AXSwitch`). Its source changed, so it is rebuilt:
36
+ run `beans-picker grant-ax` again after upgrading.
37
+
38
+ ### Fixed
39
+
40
+ - A call canceled by the client while it was still waiting for its turn no longer runs.
41
+ - A step in `then` that raises no longer discards the steps before it: it is reported as a `failed`
42
+ step, with the completed steps and the `skipped` count.
43
+ - A Jev answer with a probability or confidence outside [0, 1], or a ranking without
44
+ probabilities, is rejected as `jev_bad_response`.
45
+ - `~` in `CUA_DRIVER_BIN` is expanded.
46
+ - A screenshot that cua-driver saved to a path of its own is deleted on every failure path too.
47
+ - A server made by `create_server()` without a session closes the session it made when it stops.
48
+ - A text field named only by its placeholder keeps its candidate ids once it is typed in, so text
49
+ entered into it is checked in that field instead of ending as `mismatch` ("the target field is gone").
50
+
51
+ [Unreleased]: https://github.com/mimo-3/beans-picker/commits/main
@@ -0,0 +1,59 @@
1
+ # Contributing
2
+
3
+ ## Setup
4
+
5
+ You need [uv](https://docs.astral.sh/uv/). The test suite runs on macOS and Linux; using the
6
+ server needs macOS, cua-driver and a Jev key (see the README).
7
+
8
+ ```sh
9
+ uv sync
10
+ ```
11
+
12
+ ## Checks
13
+
14
+ Run all of these before opening a pull request; CI runs the same ones.
15
+
16
+ ```sh
17
+ uv run ruff check
18
+ uv run ruff format --check
19
+ uv run mypy
20
+ uv run pytest
21
+ ```
22
+
23
+ mypy runs in strict mode over `src`, `tests` and `bench`, and coverage must stay at 90% or more.
24
+
25
+ ## Ground rules
26
+
27
+ - **Behavior is the contract.** Tool names, input schemas, statuses, codes, messages, JSON key
28
+ order and number formatting are what clients rely on. A change to any of them needs a test
29
+ that pins the new behavior and a CHANGELOG entry.
30
+ - **stdout belongs to MCP.** The server writes nothing but protocol frames to stdout. Log with the
31
+ `beans_picker` logger; it goes to stderr.
32
+ - **No OS access in tests.** Subprocesses, cua-driver and the cache directory sit behind small
33
+ seams (`runner=`, `Paths`, fakes in `tests/fakes.py`), so the suite runs anywhere. Tests that
34
+ need macOS are marked `macos` and skipped elsewhere; the end-to-end test also needs
35
+ `BEANS_PICKER_E2E=1` and a real session with cua-driver.
36
+ - **Exact text.** Lengths are counted in UTF-16 units, the way the accessibility API counts them,
37
+ and whitespace means the fixed set in `beans_picker._text`. Use its regex classes instead of `\s`,
38
+ `\w`, `\d` and `.`.
39
+ - Code, comments and docs are in English.
40
+
41
+ ## Native helpers
42
+
43
+ `src/beans_picker/native/*.m` are compiled on the user's machine, on first use. Keep them small,
44
+ read-only and free of dependencies beyond the macOS SDK. A change to either file changes its
45
+ cache key, so the next run rebuilds it.
46
+
47
+ ## Benchmark
48
+
49
+ `bench/` compares Claude Code with cua-driver alone against Claude Code with beans-picker on eight
50
+ tasks. It needs a Mac with cua-driver, the `claude` CLI and an unlocked screen.
51
+
52
+ ```sh
53
+ sh bench/fixture-app/build.sh # the fixture app, into ~/Library/Caches/beans-picker/fixture
54
+ uv run python -m bench.run --reps 3 # appends to bench/results/runs.jsonl
55
+ uv run python -m bench.summarize > bench/results/summary.md
56
+ ```
57
+
58
+ Replace `bench/results/runs.jsonl` rather than appending to it when you publish new numbers, and
59
+ update the README tables from the new `summary.md`.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 mimo-3
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,227 @@
1
+ Metadata-Version: 2.5
2
+ Name: beans-picker
3
+ Version: 0.3.0
4
+ Summary: An MCP server that handpicks the right control for agents driving macOS apps, acts without stealing focus, and checks the effect.
5
+ Project-URL: Repository, https://github.com/mimo-3/beans-picker
6
+ Project-URL: Issues, https://github.com/mimo-3/beans-picker/issues
7
+ Project-URL: Changelog, https://github.com/mimo-3/beans-picker/blob/main/CHANGELOG.md
8
+ Author: mimo-3
9
+ License-Expression: MIT
10
+ License-File: LICENSE
11
+ Keywords: accessibility,agents,automation,macos,mcp
12
+ Classifier: Development Status :: 4 - Beta
13
+ Classifier: Environment :: MacOS X
14
+ Classifier: Intended Audience :: Developers
15
+ Classifier: Operating System :: MacOS
16
+ Classifier: Programming Language :: Python :: 3 :: Only
17
+ Classifier: Programming Language :: Python :: 3.12
18
+ Classifier: Programming Language :: Python :: 3.13
19
+ Classifier: Programming Language :: Python :: 3.14
20
+ Classifier: Topic :: Desktop Environment
21
+ Classifier: Typing :: Typed
22
+ Requires-Python: >=3.12
23
+ Requires-Dist: anyio<5,>=4.9
24
+ Requires-Dist: httpx2<3,>=2.13
25
+ Requires-Dist: mcp<2.3,>=2.2
26
+ Requires-Dist: pydantic<3,>=2.13
27
+ Requires-Dist: typesafe-sdk<0.8,>=0.7.1
28
+ Description-Content-Type: text/markdown
29
+
30
+ # beans-picker
31
+
32
+ Handpicks the right control. ~2× faster, ~1/6 the cost.
33
+
34
+ Not affiliated with Cua / trycua or TypeSafe; cua-driver and Jev are separate projects this server talks to.
35
+
36
+ An MCP server that lets Claude Code or Codex pick the right control in a macOS app, act on it without bringing the app to the front, and check that it worked. It does for native apps what Stagehand does for the browser:
37
+
38
+ - **The caller thinks.** Claude Code or Codex breaks the task down, picks the next step and decides what to do when something goes wrong.
39
+ - **Jev picks the element.** [TypeSafe Jev](https://typesafe.ai) chooses among candidate actions built from the accessibility tree. It never writes text, so it cannot invent a target.
40
+ - **cua-driver acts.** [cua-driver](https://github.com/trycua/cua) clicks, types and presses keys without bringing the app to the front.
41
+
42
+ The server itself never calls a generative model (no `claude -p`, no LLM API).
43
+
44
+ ## Tools
45
+
46
+ | tool | input | what it does |
47
+ |---|---|---|
48
+ | `observe` | `app` \| `pid`, `windowId?`, `instruction?`, `limit?` | Lists the window's candidate actions (`id`, `kind`, what it does). With `instruction`, Jev ranks them and each comes with its probability `p`. |
49
+ | `act` | `app` \| `pid`, `windowId?`, `instruction`, `text?`, `candidateId?`, `allowDestructive?`, `modifiers?`, `then?` | Performs one action and checks its effect, then each step in `then` (same fields, up to 12) the same way, stopping at the first that is not `done` or `unverified`; each step's result is in `steps`. Jev picks the action for `instruction`, unless `candidateId` is given. `text` is entered exactly as given. `modifiers` (`shift`, `cmd`, `option`, `ctrl`) are held during a click or toggle, as a pixel click on a control visible on the window. |
50
+ | `extract` | `app` \| `pid`, `windowId?`, `instruction` | Returns the text or value of the element Jev picks, exactly as read. A table or list with no value of its own comes back as `rows` (each row's texts), any other container as `text`. Its `status` is `done`, `ambiguous` (the shortlist comes back with each element's value) or `not_found`. |
51
+
52
+ `act` returns a `status`:
53
+
54
+ | status | meaning |
55
+ |---|---|
56
+ | `done` | The effect asked for is observed. For text, the field reads **exactly** the expected text (see "Exact checks"). |
57
+ | `unverified` | The field changed, but its exact text could not be read (the helper has no Accessibility permission). |
58
+ | `no_effect` | The window's signature did not change over `BEANS_PICKER_EFFECT_RETAKES` fresh snapshots (default 5). Decided by count, not by waiting. |
59
+ | `mismatch` | Something changed, but not what was asked (for example the text landed in another field, or a trailing space was lost). |
60
+ | `ambiguous` | Jev had no clear leader. The top candidates are returned; call `act` again with the right `candidateId`. |
61
+ | `needs_confirmation` | The action may not be undoable (delete, close, send, quit …). Call again with `candidateId` and `allowDestructive: true`. |
62
+ | `not_found` | No candidate fits, or the `candidateId` is not on the window any more. |
63
+ | `failed` | cua-driver refused, the command needs the foreground, or the app came to the front (`foreground_violation`). |
64
+
65
+ A call that cannot run at all returns `isError` with `{"status": "failed", "code", "message"}`. Codes include `bad_target`, `window_not_found`, `screen_locked`, `jev_unavailable` (no key, Jev unreachable or too slow), `jev_bad_response`, `driver_unavailable` (cua-driver could not be started), `driver_timeout`, `driver_error` (cua-driver refused a call the tool needed) and `internal`. Anything but a tool's own refusal is also logged with its traceback at `WARNING`.
66
+
67
+ Each result also carries the action that ran (`action`, with the cua-driver route) and a summary of what changed on the window (`change`).
68
+
69
+ Candidate ids are derived from a control's stable identity (role, identifier, label and ancestors), not from the snapshot's element index, so an id from `observe` still works after the window changes, as long as the control is there.
70
+
71
+ ## Install
72
+
73
+ Requirements: macOS, Python 3.12+, [cua-driver](https://github.com/trycua/cua) at `~/.local/bin/cua-driver` with its Accessibility and Screen Recording permissions, a Jev API key, and Xcode Command Line Tools (`clang`) for the two small native helpers, which are compiled on first use into `~/Library/Caches/beans-picker`.
74
+
75
+ ```sh
76
+ uv tool install beans-picker
77
+ mkdir -p ~/.config/beans-picker && echo 'JEV_API_KEY=...' > ~/.config/beans-picker/.env.local
78
+ claude mcp add beans-picker -- "$(command -v beans-picker)" # absolute path: GUI-launched clients often lack ~/.local/bin on PATH
79
+ beans-picker grant-ax # exact text (optional, see below)
80
+ ```
81
+
82
+ Without installing, let `uvx` fetch it on each start: `claude mcp add beans-picker -- "$(command -v uvx)" beans-picker`, and `uvx beans-picker grant-ax`.
83
+
84
+ For Codex, add the same absolute path as an MCP server in `~/.codex/config.toml` (`command -v beans-picker` prints it):
85
+
86
+ ```toml
87
+ [mcp_servers.beans-picker]
88
+ command = "/Users/you/.local/bin/beans-picker"
89
+ ```
90
+
91
+ From a checkout:
92
+
93
+ ```sh
94
+ uv sync
95
+ echo 'JEV_API_KEY=...' > .env.local # read from the checkout, whatever the caller's cwd
96
+ claude mcp add beans-picker -- "$(command -v uv)" run --directory "$PWD" beans-picker
97
+ uv run beans-picker grant-ax
98
+ ```
99
+
100
+ `beans-picker grant-ax` builds the read-only helper and asks macOS to list it; then turn on **beans-picker axtext** under System Settings > Privacy & Security > Accessibility. The helper is built from its source, so an upgrade that changes that source builds a new helper that macOS has not been told about: run `beans-picker grant-ax` again after upgrading if text results turn `unverified`. `beans-picker --version` prints the version and `beans-picker --help` the usage.
101
+
102
+ The agent skill in `skills/beans-picker` tells a model how to use these tools well: when cua-driver fits better, what to do with each `act` status, and recipes for combo boxes, number fields, long tables and checking the result. Link it where your agent looks for skills, e.g. `ln -s "$PWD/skills/beans-picker" ~/.claude/skills/beans-picker` (Codex: `~/.codex/skills`). Leave it out when benchmarking the bare tools.
103
+
104
+ ## Configuration
105
+
106
+ | variable | meaning |
107
+ |---|---|
108
+ | `JEV_API_KEY` | Required for `observe` with an instruction, `act` without `candidateId`, and `extract`. `TYPESAFE_API_KEY` is used when it is not set or empty. |
109
+ | `BEANS_PICKER_MODEL` | The Jev model (default `jev-latest`). |
110
+ | `CUA_DRIVER_BIN` | Path to cua-driver (default `~/.local/bin/cua-driver`; a leading `~` is expanded). |
111
+ | `CUA_DRIVER_TIMEOUT` | Seconds to wait for cua-driver to start, and for its answer to each call (default 120). |
112
+ | `JEV_CONNECT_TIMEOUT` | Seconds to wait for a connection to Jev (default 10). |
113
+ | `JEV_READ_TIMEOUT` | Seconds to wait for Jev's answer to one request (default 120). |
114
+ | `BEANS_PICKER_EFFECT_RETAKES` | Fresh snapshots taken to see an effect (default 5). |
115
+ | `BEANS_PICKER_LOG_LEVEL` | Level of the server's log on stderr (default `WARNING`). |
116
+ | `TYPESAFE_BASE_URL` | The Jev endpoint (default `https://api.typesafe.ai`). |
117
+
118
+ Variables are read from the environment first, then from `.env.local` and `.env` in the checkout (only when the package runs from a beans-picker checkout), then from `.env.local` and `.env` in `~/.config/beans-picker` (`$XDG_CONFIG_HOME/beans-picker` when that is set). A value that is already set is never overridden. See `.env.example`.
119
+
120
+ ## Rules the server keeps
121
+
122
+ - **Background only.** `bring_to_front`, `invoke_menu`, `move_cursor` and `delivery_mode: "foreground"` are refused before they reach cua-driver. Menu commands run as their keyboard shortcut sent to the app's pid; a command without a known shortcut returns `failed` / `foreground_required`. During `act`, the frontmost app is sampled continuously; if the target app comes to the front, the call stops with `foreground_violation`.
123
+ - **Time limits.** Every request to cua-driver and to Jev has a generous limit (see Configuration), so one that stalls ends its call with `driver_timeout` or `jev_unavailable` instead of blocking the queue. An action whose answer timed out may still have happened: observe before repeating it. At shutdown, the final `end_session` gets 5 seconds.
124
+ - **Screen lock.** While `CGSSessionScreenIsLocked` is set, every tool refuses with `screen_locked` and does nothing.
125
+ - **One call at a time.** Tool calls are queued, so two actions never interleave on the desktop.
126
+ - **Destructive actions** are marked, not hidden: `act` returns `needs_confirmation` unless `allowDestructive` is set, whoever picked the action. A checkbox is never destructive, since it can be switched back.
127
+
128
+ ## How `act` works
129
+
130
+ ```
131
+ snapshot ─► candidates ─► Jev picks (or candidateId) ─► gate ─► cua-driver ─► fresh snapshots ─► status
132
+ ```
133
+
134
+ 1. **Snapshot** (`src/beans_picker/observe`). `get_window_state` returns the AX tree. The structured elements are joined with the tree markdown, which holds identifiers, help text and static text. The menu bar is split off, and a signature of what matters (title, field values, visible text, other windows) is computed.
135
+ 2. **Candidates** (`src/beans_picker/candidates`). Clicks, toggles, text entry (`set_value`, `type_into` at the caret, `append` at the end), pop-up choices, an on-screen keypad sequence for `text` ("12×7="), menu commands with their background shortcut, one step up or down on a slider or number field (an arrow key sent to it), the options of a web list box (a combo box's suggestions), a page down or up on a scroll view, table, list or web page, a named row's, link's or image's context menu (AXShowMenu, which a web page receives as a right-click), and Return / Escape / Tab / Space / the arrow keys / Shift+F10. With `text`, only the actions that enter text are in the running.
136
+ 3. **Jev** (`src/beans_picker/jev`). One `system_one` request asks two choice questions over the same candidates: one with a `none` option, one forced. More than 60 candidates are sharded (best lexical match first), and the leaders of each shard go to a runoff. The gate acts on a leader at p ≥ 0.8, or at p ≥ 0.5 when the forced question agrees and the leader has twice the runner-up's probability. Otherwise the result is `ambiguous` (or `not_found` when `none` dominates).
137
+ 4. **cua-driver** (`src/beans_picker/act`). AX press, pixel clicks for keypads, a pop-up pressed open and then its item pressed (a menu left open over the page, as Chrome's is, is closed with Escape), `set_value` / `type_text` (a web page's number field ignores an AXValue write, so it is retyped: End, ⇧Home, then `type_text`; once an app's page has ignored a write, its text fields are typed into straight away), ⌘↓ then `type_text` for `append`, a wheel event at the area for a scroll, and the menu shortcut as a pid-targeted hotkey. Stale element tokens are rebound by stable key.
138
+ 5. **Effect** (`src/beans_picker/verify/effect.py`). Fresh snapshots are taken, up to the retake count, until the effect shows.
139
+
140
+ ## Exact checks
141
+
142
+ A false success is the worst failure a desktop agent can have: text judged by `trim` or by "contains", a search field mistaken for the document, a trailing space silently lost. `act` therefore checks text by **exact equality, in the targeted field only**:
143
+
144
+ - `set_value`: the field reads exactly `text`.
145
+ - `append`: the field reads exactly its previous text followed by `text`.
146
+ - `type_into`: the field reads its previous text with `text` inserted whole at one position.
147
+
148
+ cua-driver 0.8 does not give the exact text: it trims whitespace at both ends of a value, shows an empty field's placeholder as its value, and leaves out a checkbox's state. The read-only helper `src/beans_picker/native/axtext.m` reads AXValue as it is. It needs the Accessibility permission, which is why it runs as its own tiny background app (`beans-picker axtext`) that you can allow on its own, without granting anything to your terminal. Without it, a text action that visibly changed the field returns `unverified`, never `done`. A checkbox whose state is not in the tree is judged by its own pixels, captured in the background before and after the click.
149
+
150
+ ## Privacy
151
+
152
+ Window text leaves the machine only when a step needs Jev. [SECURITY.md](https://github.com/mimo-3/beans-picker/blob/main/SECURITY.md) lists exactly what is sent, what runs locally, and how to report a vulnerability.
153
+
154
+ ## Development
155
+
156
+ ```sh
157
+ uv sync
158
+ uv run ruff check
159
+ uv run ruff format --check
160
+ uv run mypy
161
+ uv run pytest
162
+ ```
163
+
164
+ The benchmark is in `bench/` (see below). `bench/fixture-app` is a small AppKit window used only by the benchmark; `sh bench/fixture-app/build.sh` builds it into `~/Library/Caches/beans-picker/fixture`. `python -m bench.run`, `python -m bench.judge` and `python -m bench.summarize` run from the repository root. See [CONTRIBUTING.md](https://github.com/mimo-3/beans-picker/blob/main/CONTRIBUTING.md).
165
+
166
+ ## Benchmark
167
+
168
+ Run on 2026-09-24 with an unpublished pre-release of beans-picker (0.2.0) and `--reps 3`: 8 tasks × 2 conditions × 3 repetitions = 48 runs, Claude Code headless (`claude -p`) with `--model sonnet`, one run at a time. Raw records: [`bench/results/runs.jsonl`](https://github.com/mimo-3/beans-picker/blob/main/bench/results/runs.jsonl); the tables below are `python -m bench.summarize` of them.
169
+
170
+ **Conditions.** Both get the same prompt (task, target pid and window id, "work in the background", and a final `RESULT: success|failure` line), no built-in tools (`--tools ""`) and only one MCP server:
171
+
172
+ - **(a) cua-driver only**: `cua-driver mcp`, with `bring_to_front`, `move_cursor`, `kill_app`, `get_desktop_state` and a few other tools denied.
173
+ - **(b) beans-picker**: this server only.
174
+
175
+ **Tasks** ([`bench/tasks.json`](https://github.com/mimo-3/beans-picker/blob/main/bench/tasks.json)). Five run on `bench/fixture-app`, a small AppKit window made for the bench (so no user document is touched): a Name with leading and trailing spaces, appending to a note body next to a search field, a pop-up and Save, clearing a search field next to a destructive "Delete note" button, and fixing an email plus a checkbox. Three run on Calculator: `(48 + 16) / 8`, 15% of 80, and `7 − 19` then change sign. TextEdit and Notes were not used.
176
+
177
+ **Judging.** Success is decided only by [`bench/judge.py`](https://github.com/mimo-3/beans-picker/blob/main/bench/judge.py), a separate script that compares the final state with the task's `expected` JSON by exact equality: every key of the fixture's state file (the controls' values, written by the app itself), or Calculator's display read over accessibility (bidi marks removed, nothing else). The agent's `RESULT:` line is used only to count false success claims. The front app was sampled every 200 ms during each run to count focus steals.
178
+
179
+ #### Overall
180
+
181
+ | | success | false success / success claims | tool calls (median) | time s (median) | Claude tokens (median, incl. cache) | output tokens (median) | USD (median) | Jev calls (total) | focus steals |
182
+ |---|---|---|---|---|---|---|---|---|---|
183
+ | (a) cua-driver only | 23/24 | 0/22 | 10.0 | 49.7 | 646.0k | 1876 | 0.388 | 0 | 0 |
184
+ | (b) beans-picker | 24/24 | 0/24 | 5.5 | 28.1 | 117.9k | 962 | 0.060 | 265 | 0 |
185
+
186
+ #### Per task
187
+
188
+ | task / condition | success | false success / success claims | tool calls (median) | time s (median) | Claude tokens (median, incl. cache) | output tokens (median) | USD (median) | Jev calls (total) | focus steals |
189
+ |---|---|---|---|---|---|---|---|---|---|
190
+ | fx-name-spaces a | 2/3 | 0/1 | 13.0 | 135.6 | 1150.9k | 10704 | 0.673 | 0 | 0 |
191
+ | fx-name-spaces b | 3/3 | 0/3 | 3.0 | 18.5 | 69.8k | 1213 | 0.056 | 9 | 0 |
192
+ | fx-body-not-search a | 3/3 | 0/3 | 10.0 | 57.6 | 693.4k | 3514 | 0.437 | 0 | 0 |
193
+ | fx-body-not-search b | 3/3 | 0/3 | 3.0 | 12.9 | 71.3k | 690 | 0.054 | 6 | 0 |
194
+ | fx-size-save a | 3/3 | 0/3 | 8.0 | 27.3 | 580.9k | 1160 | 0.389 | 0 | 0 |
195
+ | fx-size-save b | 3/3 | 0/3 | 7.0 | 28.6 | 152.1k | 1144 | 0.084 | 10 | 0 |
196
+ | fx-clear-search-keep-note a | 3/3 | 0/3 | 3.0 | 12.8 | 202.0k | 489 | 0.187 | 0 | 0 |
197
+ | fx-clear-search-keep-note b | 3/3 | 0/3 | 3.0 | 14.3 | 70.9k | 709 | 0.053 | 5 | 0 |
198
+ | fx-email-fix-newsletter a | 3/3 | 0/3 | 6.0 | 37.3 | 400.6k | 1179 | 0.292 | 0 | 0 |
199
+ | fx-email-fix-newsletter b | 3/3 | 0/3 | 4.0 | 22.4 | 90.4k | 871 | 0.061 | 6 | 0 |
200
+ | calc-chain a | 3/3 | 0/3 | 14.0 | 72.0 | 874.7k | 1952 | 0.353 | 0 | 0 |
201
+ | calc-chain b | 3/3 | 0/3 | 9.0 | 68.1 | 185.5k | 1492 | 0.091 | 108 | 0 |
202
+ | calc-percent a | 3/3 | 0/3 | 13.0 | 69.9 | 1129.4k | 2467 | 0.557 | 0 | 0 |
203
+ | calc-percent b | 3/3 | 0/3 | 9.0 | 49.4 | 185.3k | 1690 | 0.069 | 67 | 0 |
204
+ | calc-negate a | 3/3 | 0/3 | 10.0 | 44.4 | 633.1k | 1424 | 0.332 | 0 | 0 |
205
+ | calc-negate b | 3/3 | 0/3 | 6.0 | 36.1 | 124.0k | 756 | 0.042 | 54 | 0 |
206
+
207
+ Tokens are what Claude Code reported (input + output + cache reads + cache writes); the USD columns are Claude only. Jev is billed separately at $0.042 per 1M input tokens (output free; [TypeSafe models](https://docs.typesafe.ai/models), as of 2026-09-27): 608,908 Jev input tokens ≈ $0.026. Totals including that Jev cost: (a) USD 10.50 and 1,541 s for 24 runs; (b) USD 1.62 and 820 s for 24 runs (median $0.060 per run), from 265 Jev calls.
208
+
209
+ **What this shows**
210
+
211
+ - Both conditions nearly always succeeded. (b) passed 24/24, (a) 23/24: one Name-with-spaces run failed, and the agent said so. In another (a) run the agent reported failure although the judge found the task done.
212
+ - Neither condition claimed success on a failed run in the counted runs. So this benchmark does **not** show that beans-picker reduces false success claims. In a pilot run before the benchmark (not counted), (a) typed `48+16÷8`, got 50 and claimed success.
213
+ - (b) used fewer tool calls (median 5.5 vs 10), less time (28 s vs 50 s) and about 1/5 of the Claude tokens, mostly because cua-driver's `get_window_state` returns the whole tree (and a screenshot) to the model at every step, while beans-picker returns a short candidate list and the change summary.
214
+
215
+ **Where beans-picker did not win**
216
+
217
+ - **Time on simple tasks.** `fx-size-save` (28.6 s vs 27.3 s) and `fx-clear-search-keep-note` (14.3 s vs 12.8 s) were slower with beans-picker: a Jev call plus fresh snapshots cost more than one direct click. One `calc-chain` run took 126 s and 18 tool calls with beans-picker, the slowest run of that task in either condition.
218
+ - **Output tokens** were higher with beans-picker on `fx-clear-search-keep-note` (709 vs 489) and about equal on `fx-size-save`.
219
+ - **Jev cost is small.** 265 Jev calls add ≈ $0.026 to (b)'s total ($1.62 with Jev vs $1.59 Claude-only); they are not in the USD columns above.
220
+ - **Pop-ups.** `choose_option` failed in all three `fx-size-save` runs: cua-driver cannot `set_value` a closed `NSPopUpButton` ("has no AX children"). The caller recovered by clicking the pop-up and then the menu item, which is why that task took 6–7 calls in both conditions. Since then `choose_option` presses the pop-up open and then the item titled exactly `text`, in the background (checked on the fixture and on a `<select>` in Chrome).
221
+ - **Exact text was not confirmed.** The benchmark ran without the Accessibility permission for the `beans-picker axtext` helper. Of the 242 `act` results in (b), 206 were `done`, 18 `unverified` (text entered; the judge later found it exactly right), 6 `failed` (the pop-up above), 6 `ambiguous`, 4 `not_found` and 2 `no_effect`.
222
+
223
+ **Caveats.** 3 repetitions per task and one model (Sonnet 5) are a small sample, and one person's Mac. The tasks and the fixture app were written by the same author as beans-picker. An earlier attempt at this run was stopped after 3 runs because the fixture only recorded typed input, not values set over accessibility (so it failed a correct `set_value`). The fixture was fixed to write the controls' actual values, and the 48 runs above were all made after that. The baseline is Claude Code with cua-driver only; no other agent was compared.
224
+
225
+ ## License
226
+
227
+ MIT
@@ -0,0 +1,198 @@
1
+ # beans-picker
2
+
3
+ Handpicks the right control. ~2× faster, ~1/6 the cost.
4
+
5
+ Not affiliated with Cua / trycua or TypeSafe; cua-driver and Jev are separate projects this server talks to.
6
+
7
+ An MCP server that lets Claude Code or Codex pick the right control in a macOS app, act on it without bringing the app to the front, and check that it worked. It does for native apps what Stagehand does for the browser:
8
+
9
+ - **The caller thinks.** Claude Code or Codex breaks the task down, picks the next step and decides what to do when something goes wrong.
10
+ - **Jev picks the element.** [TypeSafe Jev](https://typesafe.ai) chooses among candidate actions built from the accessibility tree. It never writes text, so it cannot invent a target.
11
+ - **cua-driver acts.** [cua-driver](https://github.com/trycua/cua) clicks, types and presses keys without bringing the app to the front.
12
+
13
+ The server itself never calls a generative model (no `claude -p`, no LLM API).
14
+
15
+ ## Tools
16
+
17
+ | tool | input | what it does |
18
+ |---|---|---|
19
+ | `observe` | `app` \| `pid`, `windowId?`, `instruction?`, `limit?` | Lists the window's candidate actions (`id`, `kind`, what it does). With `instruction`, Jev ranks them and each comes with its probability `p`. |
20
+ | `act` | `app` \| `pid`, `windowId?`, `instruction`, `text?`, `candidateId?`, `allowDestructive?`, `modifiers?`, `then?` | Performs one action and checks its effect, then each step in `then` (same fields, up to 12) the same way, stopping at the first that is not `done` or `unverified`; each step's result is in `steps`. Jev picks the action for `instruction`, unless `candidateId` is given. `text` is entered exactly as given. `modifiers` (`shift`, `cmd`, `option`, `ctrl`) are held during a click or toggle, as a pixel click on a control visible on the window. |
21
+ | `extract` | `app` \| `pid`, `windowId?`, `instruction` | Returns the text or value of the element Jev picks, exactly as read. A table or list with no value of its own comes back as `rows` (each row's texts), any other container as `text`. Its `status` is `done`, `ambiguous` (the shortlist comes back with each element's value) or `not_found`. |
22
+
23
+ `act` returns a `status`:
24
+
25
+ | status | meaning |
26
+ |---|---|
27
+ | `done` | The effect asked for is observed. For text, the field reads **exactly** the expected text (see "Exact checks"). |
28
+ | `unverified` | The field changed, but its exact text could not be read (the helper has no Accessibility permission). |
29
+ | `no_effect` | The window's signature did not change over `BEANS_PICKER_EFFECT_RETAKES` fresh snapshots (default 5). Decided by count, not by waiting. |
30
+ | `mismatch` | Something changed, but not what was asked (for example the text landed in another field, or a trailing space was lost). |
31
+ | `ambiguous` | Jev had no clear leader. The top candidates are returned; call `act` again with the right `candidateId`. |
32
+ | `needs_confirmation` | The action may not be undoable (delete, close, send, quit …). Call again with `candidateId` and `allowDestructive: true`. |
33
+ | `not_found` | No candidate fits, or the `candidateId` is not on the window any more. |
34
+ | `failed` | cua-driver refused, the command needs the foreground, or the app came to the front (`foreground_violation`). |
35
+
36
+ A call that cannot run at all returns `isError` with `{"status": "failed", "code", "message"}`. Codes include `bad_target`, `window_not_found`, `screen_locked`, `jev_unavailable` (no key, Jev unreachable or too slow), `jev_bad_response`, `driver_unavailable` (cua-driver could not be started), `driver_timeout`, `driver_error` (cua-driver refused a call the tool needed) and `internal`. Anything but a tool's own refusal is also logged with its traceback at `WARNING`.
37
+
38
+ Each result also carries the action that ran (`action`, with the cua-driver route) and a summary of what changed on the window (`change`).
39
+
40
+ Candidate ids are derived from a control's stable identity (role, identifier, label and ancestors), not from the snapshot's element index, so an id from `observe` still works after the window changes, as long as the control is there.
41
+
42
+ ## Install
43
+
44
+ Requirements: macOS, Python 3.12+, [cua-driver](https://github.com/trycua/cua) at `~/.local/bin/cua-driver` with its Accessibility and Screen Recording permissions, a Jev API key, and Xcode Command Line Tools (`clang`) for the two small native helpers, which are compiled on first use into `~/Library/Caches/beans-picker`.
45
+
46
+ ```sh
47
+ uv tool install beans-picker
48
+ mkdir -p ~/.config/beans-picker && echo 'JEV_API_KEY=...' > ~/.config/beans-picker/.env.local
49
+ claude mcp add beans-picker -- "$(command -v beans-picker)" # absolute path: GUI-launched clients often lack ~/.local/bin on PATH
50
+ beans-picker grant-ax # exact text (optional, see below)
51
+ ```
52
+
53
+ Without installing, let `uvx` fetch it on each start: `claude mcp add beans-picker -- "$(command -v uvx)" beans-picker`, and `uvx beans-picker grant-ax`.
54
+
55
+ For Codex, add the same absolute path as an MCP server in `~/.codex/config.toml` (`command -v beans-picker` prints it):
56
+
57
+ ```toml
58
+ [mcp_servers.beans-picker]
59
+ command = "/Users/you/.local/bin/beans-picker"
60
+ ```
61
+
62
+ From a checkout:
63
+
64
+ ```sh
65
+ uv sync
66
+ echo 'JEV_API_KEY=...' > .env.local # read from the checkout, whatever the caller's cwd
67
+ claude mcp add beans-picker -- "$(command -v uv)" run --directory "$PWD" beans-picker
68
+ uv run beans-picker grant-ax
69
+ ```
70
+
71
+ `beans-picker grant-ax` builds the read-only helper and asks macOS to list it; then turn on **beans-picker axtext** under System Settings > Privacy & Security > Accessibility. The helper is built from its source, so an upgrade that changes that source builds a new helper that macOS has not been told about: run `beans-picker grant-ax` again after upgrading if text results turn `unverified`. `beans-picker --version` prints the version and `beans-picker --help` the usage.
72
+
73
+ The agent skill in `skills/beans-picker` tells a model how to use these tools well: when cua-driver fits better, what to do with each `act` status, and recipes for combo boxes, number fields, long tables and checking the result. Link it where your agent looks for skills, e.g. `ln -s "$PWD/skills/beans-picker" ~/.claude/skills/beans-picker` (Codex: `~/.codex/skills`). Leave it out when benchmarking the bare tools.
74
+
75
+ ## Configuration
76
+
77
+ | variable | meaning |
78
+ |---|---|
79
+ | `JEV_API_KEY` | Required for `observe` with an instruction, `act` without `candidateId`, and `extract`. `TYPESAFE_API_KEY` is used when it is not set or empty. |
80
+ | `BEANS_PICKER_MODEL` | The Jev model (default `jev-latest`). |
81
+ | `CUA_DRIVER_BIN` | Path to cua-driver (default `~/.local/bin/cua-driver`; a leading `~` is expanded). |
82
+ | `CUA_DRIVER_TIMEOUT` | Seconds to wait for cua-driver to start, and for its answer to each call (default 120). |
83
+ | `JEV_CONNECT_TIMEOUT` | Seconds to wait for a connection to Jev (default 10). |
84
+ | `JEV_READ_TIMEOUT` | Seconds to wait for Jev's answer to one request (default 120). |
85
+ | `BEANS_PICKER_EFFECT_RETAKES` | Fresh snapshots taken to see an effect (default 5). |
86
+ | `BEANS_PICKER_LOG_LEVEL` | Level of the server's log on stderr (default `WARNING`). |
87
+ | `TYPESAFE_BASE_URL` | The Jev endpoint (default `https://api.typesafe.ai`). |
88
+
89
+ Variables are read from the environment first, then from `.env.local` and `.env` in the checkout (only when the package runs from a beans-picker checkout), then from `.env.local` and `.env` in `~/.config/beans-picker` (`$XDG_CONFIG_HOME/beans-picker` when that is set). A value that is already set is never overridden. See `.env.example`.
90
+
91
+ ## Rules the server keeps
92
+
93
+ - **Background only.** `bring_to_front`, `invoke_menu`, `move_cursor` and `delivery_mode: "foreground"` are refused before they reach cua-driver. Menu commands run as their keyboard shortcut sent to the app's pid; a command without a known shortcut returns `failed` / `foreground_required`. During `act`, the frontmost app is sampled continuously; if the target app comes to the front, the call stops with `foreground_violation`.
94
+ - **Time limits.** Every request to cua-driver and to Jev has a generous limit (see Configuration), so one that stalls ends its call with `driver_timeout` or `jev_unavailable` instead of blocking the queue. An action whose answer timed out may still have happened: observe before repeating it. At shutdown, the final `end_session` gets 5 seconds.
95
+ - **Screen lock.** While `CGSSessionScreenIsLocked` is set, every tool refuses with `screen_locked` and does nothing.
96
+ - **One call at a time.** Tool calls are queued, so two actions never interleave on the desktop.
97
+ - **Destructive actions** are marked, not hidden: `act` returns `needs_confirmation` unless `allowDestructive` is set, whoever picked the action. A checkbox is never destructive, since it can be switched back.
98
+
99
+ ## How `act` works
100
+
101
+ ```
102
+ snapshot ─► candidates ─► Jev picks (or candidateId) ─► gate ─► cua-driver ─► fresh snapshots ─► status
103
+ ```
104
+
105
+ 1. **Snapshot** (`src/beans_picker/observe`). `get_window_state` returns the AX tree. The structured elements are joined with the tree markdown, which holds identifiers, help text and static text. The menu bar is split off, and a signature of what matters (title, field values, visible text, other windows) is computed.
106
+ 2. **Candidates** (`src/beans_picker/candidates`). Clicks, toggles, text entry (`set_value`, `type_into` at the caret, `append` at the end), pop-up choices, an on-screen keypad sequence for `text` ("12×7="), menu commands with their background shortcut, one step up or down on a slider or number field (an arrow key sent to it), the options of a web list box (a combo box's suggestions), a page down or up on a scroll view, table, list or web page, a named row's, link's or image's context menu (AXShowMenu, which a web page receives as a right-click), and Return / Escape / Tab / Space / the arrow keys / Shift+F10. With `text`, only the actions that enter text are in the running.
107
+ 3. **Jev** (`src/beans_picker/jev`). One `system_one` request asks two choice questions over the same candidates: one with a `none` option, one forced. More than 60 candidates are sharded (best lexical match first), and the leaders of each shard go to a runoff. The gate acts on a leader at p ≥ 0.8, or at p ≥ 0.5 when the forced question agrees and the leader has twice the runner-up's probability. Otherwise the result is `ambiguous` (or `not_found` when `none` dominates).
108
+ 4. **cua-driver** (`src/beans_picker/act`). AX press, pixel clicks for keypads, a pop-up pressed open and then its item pressed (a menu left open over the page, as Chrome's is, is closed with Escape), `set_value` / `type_text` (a web page's number field ignores an AXValue write, so it is retyped: End, ⇧Home, then `type_text`; once an app's page has ignored a write, its text fields are typed into straight away), ⌘↓ then `type_text` for `append`, a wheel event at the area for a scroll, and the menu shortcut as a pid-targeted hotkey. Stale element tokens are rebound by stable key.
109
+ 5. **Effect** (`src/beans_picker/verify/effect.py`). Fresh snapshots are taken, up to the retake count, until the effect shows.
110
+
111
+ ## Exact checks
112
+
113
+ A false success is the worst failure a desktop agent can have: text judged by `trim` or by "contains", a search field mistaken for the document, a trailing space silently lost. `act` therefore checks text by **exact equality, in the targeted field only**:
114
+
115
+ - `set_value`: the field reads exactly `text`.
116
+ - `append`: the field reads exactly its previous text followed by `text`.
117
+ - `type_into`: the field reads its previous text with `text` inserted whole at one position.
118
+
119
+ cua-driver 0.8 does not give the exact text: it trims whitespace at both ends of a value, shows an empty field's placeholder as its value, and leaves out a checkbox's state. The read-only helper `src/beans_picker/native/axtext.m` reads AXValue as it is. It needs the Accessibility permission, which is why it runs as its own tiny background app (`beans-picker axtext`) that you can allow on its own, without granting anything to your terminal. Without it, a text action that visibly changed the field returns `unverified`, never `done`. A checkbox whose state is not in the tree is judged by its own pixels, captured in the background before and after the click.
120
+
121
+ ## Privacy
122
+
123
+ Window text leaves the machine only when a step needs Jev. [SECURITY.md](https://github.com/mimo-3/beans-picker/blob/main/SECURITY.md) lists exactly what is sent, what runs locally, and how to report a vulnerability.
124
+
125
+ ## Development
126
+
127
+ ```sh
128
+ uv sync
129
+ uv run ruff check
130
+ uv run ruff format --check
131
+ uv run mypy
132
+ uv run pytest
133
+ ```
134
+
135
+ The benchmark is in `bench/` (see below). `bench/fixture-app` is a small AppKit window used only by the benchmark; `sh bench/fixture-app/build.sh` builds it into `~/Library/Caches/beans-picker/fixture`. `python -m bench.run`, `python -m bench.judge` and `python -m bench.summarize` run from the repository root. See [CONTRIBUTING.md](https://github.com/mimo-3/beans-picker/blob/main/CONTRIBUTING.md).
136
+
137
+ ## Benchmark
138
+
139
+ Run on 2026-09-24 with an unpublished pre-release of beans-picker (0.2.0) and `--reps 3`: 8 tasks × 2 conditions × 3 repetitions = 48 runs, Claude Code headless (`claude -p`) with `--model sonnet`, one run at a time. Raw records: [`bench/results/runs.jsonl`](https://github.com/mimo-3/beans-picker/blob/main/bench/results/runs.jsonl); the tables below are `python -m bench.summarize` of them.
140
+
141
+ **Conditions.** Both get the same prompt (task, target pid and window id, "work in the background", and a final `RESULT: success|failure` line), no built-in tools (`--tools ""`) and only one MCP server:
142
+
143
+ - **(a) cua-driver only**: `cua-driver mcp`, with `bring_to_front`, `move_cursor`, `kill_app`, `get_desktop_state` and a few other tools denied.
144
+ - **(b) beans-picker**: this server only.
145
+
146
+ **Tasks** ([`bench/tasks.json`](https://github.com/mimo-3/beans-picker/blob/main/bench/tasks.json)). Five run on `bench/fixture-app`, a small AppKit window made for the bench (so no user document is touched): a Name with leading and trailing spaces, appending to a note body next to a search field, a pop-up and Save, clearing a search field next to a destructive "Delete note" button, and fixing an email plus a checkbox. Three run on Calculator: `(48 + 16) / 8`, 15% of 80, and `7 − 19` then change sign. TextEdit and Notes were not used.
147
+
148
+ **Judging.** Success is decided only by [`bench/judge.py`](https://github.com/mimo-3/beans-picker/blob/main/bench/judge.py), a separate script that compares the final state with the task's `expected` JSON by exact equality: every key of the fixture's state file (the controls' values, written by the app itself), or Calculator's display read over accessibility (bidi marks removed, nothing else). The agent's `RESULT:` line is used only to count false success claims. The front app was sampled every 200 ms during each run to count focus steals.
149
+
150
+ #### Overall
151
+
152
+ | | success | false success / success claims | tool calls (median) | time s (median) | Claude tokens (median, incl. cache) | output tokens (median) | USD (median) | Jev calls (total) | focus steals |
153
+ |---|---|---|---|---|---|---|---|---|---|
154
+ | (a) cua-driver only | 23/24 | 0/22 | 10.0 | 49.7 | 646.0k | 1876 | 0.388 | 0 | 0 |
155
+ | (b) beans-picker | 24/24 | 0/24 | 5.5 | 28.1 | 117.9k | 962 | 0.060 | 265 | 0 |
156
+
157
+ #### Per task
158
+
159
+ | task / condition | success | false success / success claims | tool calls (median) | time s (median) | Claude tokens (median, incl. cache) | output tokens (median) | USD (median) | Jev calls (total) | focus steals |
160
+ |---|---|---|---|---|---|---|---|---|---|
161
+ | fx-name-spaces a | 2/3 | 0/1 | 13.0 | 135.6 | 1150.9k | 10704 | 0.673 | 0 | 0 |
162
+ | fx-name-spaces b | 3/3 | 0/3 | 3.0 | 18.5 | 69.8k | 1213 | 0.056 | 9 | 0 |
163
+ | fx-body-not-search a | 3/3 | 0/3 | 10.0 | 57.6 | 693.4k | 3514 | 0.437 | 0 | 0 |
164
+ | fx-body-not-search b | 3/3 | 0/3 | 3.0 | 12.9 | 71.3k | 690 | 0.054 | 6 | 0 |
165
+ | fx-size-save a | 3/3 | 0/3 | 8.0 | 27.3 | 580.9k | 1160 | 0.389 | 0 | 0 |
166
+ | fx-size-save b | 3/3 | 0/3 | 7.0 | 28.6 | 152.1k | 1144 | 0.084 | 10 | 0 |
167
+ | fx-clear-search-keep-note a | 3/3 | 0/3 | 3.0 | 12.8 | 202.0k | 489 | 0.187 | 0 | 0 |
168
+ | fx-clear-search-keep-note b | 3/3 | 0/3 | 3.0 | 14.3 | 70.9k | 709 | 0.053 | 5 | 0 |
169
+ | fx-email-fix-newsletter a | 3/3 | 0/3 | 6.0 | 37.3 | 400.6k | 1179 | 0.292 | 0 | 0 |
170
+ | fx-email-fix-newsletter b | 3/3 | 0/3 | 4.0 | 22.4 | 90.4k | 871 | 0.061 | 6 | 0 |
171
+ | calc-chain a | 3/3 | 0/3 | 14.0 | 72.0 | 874.7k | 1952 | 0.353 | 0 | 0 |
172
+ | calc-chain b | 3/3 | 0/3 | 9.0 | 68.1 | 185.5k | 1492 | 0.091 | 108 | 0 |
173
+ | calc-percent a | 3/3 | 0/3 | 13.0 | 69.9 | 1129.4k | 2467 | 0.557 | 0 | 0 |
174
+ | calc-percent b | 3/3 | 0/3 | 9.0 | 49.4 | 185.3k | 1690 | 0.069 | 67 | 0 |
175
+ | calc-negate a | 3/3 | 0/3 | 10.0 | 44.4 | 633.1k | 1424 | 0.332 | 0 | 0 |
176
+ | calc-negate b | 3/3 | 0/3 | 6.0 | 36.1 | 124.0k | 756 | 0.042 | 54 | 0 |
177
+
178
+ Tokens are what Claude Code reported (input + output + cache reads + cache writes); the USD columns are Claude only. Jev is billed separately at $0.042 per 1M input tokens (output free; [TypeSafe models](https://docs.typesafe.ai/models), as of 2026-09-27): 608,908 Jev input tokens ≈ $0.026. Totals including that Jev cost: (a) USD 10.50 and 1,541 s for 24 runs; (b) USD 1.62 and 820 s for 24 runs (median $0.060 per run), from 265 Jev calls.
179
+
180
+ **What this shows**
181
+
182
+ - Both conditions nearly always succeeded. (b) passed 24/24, (a) 23/24: one Name-with-spaces run failed, and the agent said so. In another (a) run the agent reported failure although the judge found the task done.
183
+ - Neither condition claimed success on a failed run in the counted runs. So this benchmark does **not** show that beans-picker reduces false success claims. In a pilot run before the benchmark (not counted), (a) typed `48+16÷8`, got 50 and claimed success.
184
+ - (b) used fewer tool calls (median 5.5 vs 10), less time (28 s vs 50 s) and about 1/5 of the Claude tokens, mostly because cua-driver's `get_window_state` returns the whole tree (and a screenshot) to the model at every step, while beans-picker returns a short candidate list and the change summary.
185
+
186
+ **Where beans-picker did not win**
187
+
188
+ - **Time on simple tasks.** `fx-size-save` (28.6 s vs 27.3 s) and `fx-clear-search-keep-note` (14.3 s vs 12.8 s) were slower with beans-picker: a Jev call plus fresh snapshots cost more than one direct click. One `calc-chain` run took 126 s and 18 tool calls with beans-picker, the slowest run of that task in either condition.
189
+ - **Output tokens** were higher with beans-picker on `fx-clear-search-keep-note` (709 vs 489) and about equal on `fx-size-save`.
190
+ - **Jev cost is small.** 265 Jev calls add ≈ $0.026 to (b)'s total ($1.62 with Jev vs $1.59 Claude-only); they are not in the USD columns above.
191
+ - **Pop-ups.** `choose_option` failed in all three `fx-size-save` runs: cua-driver cannot `set_value` a closed `NSPopUpButton` ("has no AX children"). The caller recovered by clicking the pop-up and then the menu item, which is why that task took 6–7 calls in both conditions. Since then `choose_option` presses the pop-up open and then the item titled exactly `text`, in the background (checked on the fixture and on a `<select>` in Chrome).
192
+ - **Exact text was not confirmed.** The benchmark ran without the Accessibility permission for the `beans-picker axtext` helper. Of the 242 `act` results in (b), 206 were `done`, 18 `unverified` (text entered; the judge later found it exactly right), 6 `failed` (the pop-up above), 6 `ambiguous`, 4 `not_found` and 2 `no_effect`.
193
+
194
+ **Caveats.** 3 repetitions per task and one model (Sonnet 5) are a small sample, and one person's Mac. The tasks and the fixture app were written by the same author as beans-picker. An earlier attempt at this run was stopped after 3 runs because the fixture only recorded typed input, not values set over accessibility (so it failed a correct `set_value`). The fixture was fixed to write the controls' actual values, and the 48 runs above were all made after that. The baseline is Claude Code with cua-driver only; no other agent was compared.
195
+
196
+ ## License
197
+
198
+ MIT