toolroll 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +1113 -0
- package/THIRD_PARTY_NOTICES.md +449 -0
- package/dist/accent-colors.d.ts +48 -0
- package/dist/accent-colors.js +122 -0
- package/dist/action-ledger.d.ts +26 -0
- package/dist/action-ledger.js +75 -0
- package/dist/agent-fence.d.ts +42 -0
- package/dist/agent-fence.js +183 -0
- package/dist/agentconfig.d.ts +216 -0
- package/dist/agentconfig.js +454 -0
- package/dist/api-tokens.d.ts +16 -0
- package/dist/api-tokens.js +27 -0
- package/dist/approval-policy.d.ts +69 -0
- package/dist/approval-policy.js +137 -0
- package/dist/approval-rules-ui.d.ts +25 -0
- package/dist/approval-rules-ui.js +39 -0
- package/dist/assignment-adapters.d.ts +180 -0
- package/dist/assignment-adapters.js +239 -0
- package/dist/assignment-brief.d.ts +71 -0
- package/dist/assignment-brief.js +129 -0
- package/dist/assignment-delivery.d.ts +56 -0
- package/dist/assignment-delivery.js +160 -0
- package/dist/assignment-presentation.d.ts +26 -0
- package/dist/assignment-presentation.js +80 -0
- package/dist/assignment-status.d.ts +62 -0
- package/dist/assignment-status.js +152 -0
- package/dist/assignment-ui.d.ts +71 -0
- package/dist/assignment-ui.js +103 -0
- package/dist/assignment.d.ts +222 -0
- package/dist/assignment.js +399 -0
- package/dist/attest.d.ts +56 -0
- package/dist/attest.js +153 -0
- package/dist/backend.d.ts +97 -0
- package/dist/backend.js +166 -0
- package/dist/backup-ui.d.ts +22 -0
- package/dist/backup-ui.js +59 -0
- package/dist/backup.d.ts +84 -0
- package/dist/backup.js +421 -0
- package/dist/beads.d.ts +34 -0
- package/dist/beads.js +135 -0
- package/dist/bin.d.ts +2 -0
- package/dist/bin.js +22 -0
- package/dist/board.d.ts +150 -0
- package/dist/board.js +210 -0
- package/dist/boot-identity.d.ts +63 -0
- package/dist/boot-identity.js +99 -0
- package/dist/browser/THIRD_PARTY_NOTICES.txt +2295 -0
- package/dist/browser/workspace.css +4 -0
- package/dist/browser/workspace.js +225 -0
- package/dist/browser-crew.d.ts +11 -0
- package/dist/browser-crew.js +50 -0
- package/dist/browser-shell.d.ts +8 -0
- package/dist/browser-shell.js +60 -0
- package/dist/browser-workspace.d.ts +780 -0
- package/dist/browser-workspace.js +47 -0
- package/dist/budget-alerts.d.ts +19 -0
- package/dist/budget-alerts.js +46 -0
- package/dist/builder.d.ts +370 -0
- package/dist/builder.js +3262 -0
- package/dist/capscan.d.ts +35 -0
- package/dist/capscan.js +113 -0
- package/dist/chat-acceptance.d.ts +71 -0
- package/dist/chat-acceptance.js +120 -0
- package/dist/chat-actions.d.ts +265 -0
- package/dist/chat-actions.js +1273 -0
- package/dist/chat-channel.d.ts +109 -0
- package/dist/chat-channel.js +652 -0
- package/dist/chat-continuity.d.ts +32 -0
- package/dist/chat-continuity.js +317 -0
- package/dist/chat-controls.d.ts +113 -0
- package/dist/chat-controls.js +53 -0
- package/dist/chat-delivery-state.d.ts +159 -0
- package/dist/chat-delivery-state.js +309 -0
- package/dist/chat-delivery.d.ts +38 -0
- package/dist/chat-delivery.js +693 -0
- package/dist/chat-display.d.ts +1 -0
- package/dist/chat-display.js +18 -0
- package/dist/chat-evidence.d.ts +123 -0
- package/dist/chat-evidence.js +207 -0
- package/dist/chat-flow.d.ts +44 -0
- package/dist/chat-flow.js +215 -0
- package/dist/chat-inbox.d.ts +57 -0
- package/dist/chat-inbox.js +123 -0
- package/dist/chat-polish.d.ts +16 -0
- package/dist/chat-polish.js +135 -0
- package/dist/chat-review.d.ts +46 -0
- package/dist/chat-review.js +98 -0
- package/dist/chat-rooms.d.ts +123 -0
- package/dist/chat-rooms.js +216 -0
- package/dist/chat-task-actions.d.ts +44 -0
- package/dist/chat-task-actions.js +79 -0
- package/dist/check-progress.d.ts +51 -0
- package/dist/check-progress.js +290 -0
- package/dist/child-database.d.ts +12 -0
- package/dist/child-database.js +30 -0
- package/dist/claim.d.ts +688 -0
- package/dist/claim.js +1740 -0
- package/dist/cli.d.ts +137 -0
- package/dist/cli.js +1461 -0
- package/dist/codex-limits.d.ts +15 -0
- package/dist/codex-limits.js +145 -0
- package/dist/coding-context.d.ts +55 -0
- package/dist/coding-context.js +92 -0
- package/dist/coding-handoff.d.ts +57 -0
- package/dist/coding-handoff.js +277 -0
- package/dist/coding-provider.d.ts +72 -0
- package/dist/coding-provider.js +424 -0
- package/dist/coding-shipping-ui.d.ts +4 -0
- package/dist/coding-shipping-ui.js +12 -0
- package/dist/coding-types.d.ts +62 -0
- package/dist/coding-types.js +1 -0
- package/dist/coding-ui.d.ts +17 -0
- package/dist/coding-ui.js +364 -0
- package/dist/coding-update.d.ts +15 -0
- package/dist/coding-update.js +149 -0
- package/dist/coding-workspace.d.ts +124 -0
- package/dist/coding-workspace.js +768 -0
- package/dist/container-state.d.ts +5 -0
- package/dist/container-state.js +62 -0
- package/dist/containment.d.ts +197 -0
- package/dist/containment.js +559 -0
- package/dist/contest.d.ts +318 -0
- package/dist/contest.js +753 -0
- package/dist/control-setup.d.ts +35 -0
- package/dist/control-setup.js +40 -0
- package/dist/control-ui.d.ts +29 -0
- package/dist/control-ui.js +53 -0
- package/dist/controller-service.d.ts +1 -0
- package/dist/controller-service.js +19 -0
- package/dist/controller-supervisor.d.ts +30 -0
- package/dist/controller-supervisor.js +104 -0
- package/dist/converse.d.ts +304 -0
- package/dist/converse.js +849 -0
- package/dist/coordinator-proposals.d.ts +29 -0
- package/dist/coordinator-proposals.js +137 -0
- package/dist/coordinator.d.ts +186 -0
- package/dist/coordinator.js +447 -0
- package/dist/credentials-ui.d.ts +25 -0
- package/dist/credentials-ui.js +48 -0
- package/dist/daemon.d.ts +209 -0
- package/dist/daemon.js +604 -0
- package/dist/decision.d.ts +120 -0
- package/dist/decision.js +388 -0
- package/dist/demo.d.ts +55 -0
- package/dist/demo.js +992 -0
- package/dist/desktop-access.d.ts +28 -0
- package/dist/desktop-access.js +88 -0
- package/dist/desktop-bundle.d.ts +16 -0
- package/dist/desktop-bundle.js +74 -0
- package/dist/desktop-host.d.ts +54 -0
- package/dist/desktop-host.js +508 -0
- package/dist/desktop-update-gate.d.ts +12 -0
- package/dist/desktop-update-gate.js +91 -0
- package/dist/desktop-update-recovery.d.ts +14 -0
- package/dist/desktop-update-recovery.js +243 -0
- package/dist/desktop-update.d.ts +110 -0
- package/dist/desktop-update.js +740 -0
- package/dist/discord-api.d.ts +20 -0
- package/dist/discord-api.js +138 -0
- package/dist/discord-chat.d.ts +16 -0
- package/dist/discord-chat.js +381 -0
- package/dist/discord-settings.d.ts +7 -0
- package/dist/discord-settings.js +75 -0
- package/dist/discord.d.ts +10 -0
- package/dist/discord.js +169 -0
- package/dist/discover.d.ts +75 -0
- package/dist/discover.js +150 -0
- package/dist/dispatch.d.ts +100 -0
- package/dist/dispatch.js +311 -0
- package/dist/dispose.d.ts +161 -0
- package/dist/dispose.js +803 -0
- package/dist/email-settings.d.ts +34 -0
- package/dist/email-settings.js +50 -0
- package/dist/envelope.d.ts +47 -0
- package/dist/envelope.js +105 -0
- package/dist/evidence-pack.d.ts +185 -0
- package/dist/evidence-pack.js +258 -0
- package/dist/evidence.d.ts +345 -0
- package/dist/evidence.js +782 -0
- package/dist/exec.d.ts +285 -0
- package/dist/exec.js +1841 -0
- package/dist/exhaustion.d.ts +85 -0
- package/dist/exhaustion.js +141 -0
- package/dist/export-ui.d.ts +4 -0
- package/dist/export-ui.js +17 -0
- package/dist/export.d.ts +48 -0
- package/dist/export.js +305 -0
- package/dist/flow-actions.d.ts +77 -0
- package/dist/flow-actions.js +203 -0
- package/dist/flow-code.d.ts +50 -0
- package/dist/flow-code.js +159 -0
- package/dist/flow-draft.d.ts +38 -0
- package/dist/flow-draft.js +80 -0
- package/dist/flow-engine.d.ts +66 -0
- package/dist/flow-engine.js +416 -0
- package/dist/flow-insights.d.ts +95 -0
- package/dist/flow-insights.js +149 -0
- package/dist/flow-live.d.ts +24 -0
- package/dist/flow-live.js +150 -0
- package/dist/flow-people.d.ts +24 -0
- package/dist/flow-people.js +91 -0
- package/dist/flow-replies.d.ts +48 -0
- package/dist/flow-replies.js +113 -0
- package/dist/flow-scripts.d.ts +38 -0
- package/dist/flow-scripts.js +73 -0
- package/dist/flow-secrets.d.ts +13 -0
- package/dist/flow-secrets.js +45 -0
- package/dist/flow-sort.d.ts +82 -0
- package/dist/flow-sort.js +153 -0
- package/dist/flow-steps.d.ts +40 -0
- package/dist/flow-steps.js +408 -0
- package/dist/flow-triggers.d.ts +244 -0
- package/dist/flow-triggers.js +959 -0
- package/dist/flows-ui.d.ts +20 -0
- package/dist/flows-ui.js +175 -0
- package/dist/flows.d.ts +229 -0
- package/dist/flows.js +968 -0
- package/dist/fonts.d.ts +19 -0
- package/dist/fonts.js +19 -0
- package/dist/gaps.d.ts +35 -0
- package/dist/gaps.js +101 -0
- package/dist/gate-failure.d.ts +40 -0
- package/dist/gate-failure.js +66 -0
- package/dist/git.d.ts +54 -0
- package/dist/git.js +94 -0
- package/dist/google-mail.d.ts +58 -0
- package/dist/google-mail.js +151 -0
- package/dist/grant.d.ts +136 -0
- package/dist/grant.js +238 -0
- package/dist/graph.d.ts +164 -0
- package/dist/graph.js +383 -0
- package/dist/guides.d.ts +22 -0
- package/dist/guides.js +363 -0
- package/dist/held.d.ts +150 -0
- package/dist/held.js +799 -0
- package/dist/invoke.d.ts +112 -0
- package/dist/invoke.js +780 -0
- package/dist/issues.d.ts +41 -0
- package/dist/issues.js +131 -0
- package/dist/job-object-helper.ps1 +252 -0
- package/dist/jsonl-discriminants.d.ts +11 -0
- package/dist/jsonl-discriminants.js +93 -0
- package/dist/keys.d.ts +104 -0
- package/dist/keys.js +229 -0
- package/dist/kits-ui.d.ts +17 -0
- package/dist/kits-ui.js +46 -0
- package/dist/kits.d.ts +86 -0
- package/dist/kits.js +215 -0
- package/dist/knowledge-cli.d.ts +101 -0
- package/dist/knowledge-cli.js +92 -0
- package/dist/knowledge-ui.d.ts +24 -0
- package/dist/knowledge-ui.js +40 -0
- package/dist/lead-context.d.ts +8 -0
- package/dist/lead-context.js +56 -0
- package/dist/lead-follow.d.ts +22 -0
- package/dist/lead-follow.js +167 -0
- package/dist/lead-status.d.ts +75 -0
- package/dist/lead-status.js +285 -0
- package/dist/ledger-chain.d.ts +46 -0
- package/dist/ledger-chain.js +187 -0
- package/dist/ledger-csv.d.ts +4 -0
- package/dist/ledger-csv.js +11 -0
- package/dist/ledger-view.d.ts +24 -0
- package/dist/ledger-view.js +58 -0
- package/dist/limits-ui.d.ts +26 -0
- package/dist/limits-ui.js +67 -0
- package/dist/link.d.ts +72 -0
- package/dist/link.js +217 -0
- package/dist/live.d.ts +96 -0
- package/dist/live.js +366 -0
- package/dist/liveness.d.ts +30 -0
- package/dist/liveness.js +42 -0
- package/dist/log.d.ts +7 -0
- package/dist/log.js +25 -0
- package/dist/mailbox.d.ts +57 -0
- package/dist/mailbox.js +150 -0
- package/dist/maintenance.d.ts +11 -0
- package/dist/maintenance.js +35 -0
- package/dist/mate-cli.d.ts +52 -0
- package/dist/mate-cli.js +345 -0
- package/dist/mate-contract.d.ts +10 -0
- package/dist/mate-contract.js +30 -0
- package/dist/mate-doors.d.ts +69 -0
- package/dist/mate-doors.js +548 -0
- package/dist/mate-progress.d.ts +29 -0
- package/dist/mate-progress.js +121 -0
- package/dist/mate-tools.d.ts +94 -0
- package/dist/mate-tools.js +1878 -0
- package/dist/mate.d.ts +92 -0
- package/dist/mate.js +435 -0
- package/dist/mcp-connect.d.ts +103 -0
- package/dist/mcp-connect.js +252 -0
- package/dist/mcp.d.ts +27 -0
- package/dist/mcp.js +651 -0
- package/dist/memory-cli.d.ts +466 -0
- package/dist/memory-cli.js +111 -0
- package/dist/memory-pass.d.ts +160 -0
- package/dist/memory-pass.js +409 -0
- package/dist/metrics.d.ts +3 -0
- package/dist/metrics.js +56 -0
- package/dist/mobile-viewport.d.ts +3 -0
- package/dist/mobile-viewport.js +37 -0
- package/dist/model-catalog.d.ts +118 -0
- package/dist/model-catalog.js +376 -0
- package/dist/models-cli.d.ts +102 -0
- package/dist/models-cli.js +66 -0
- package/dist/models-ui.d.ts +34 -0
- package/dist/models-ui.js +68 -0
- package/dist/modes.d.ts +80 -0
- package/dist/modes.js +182 -0
- package/dist/monitoring-settings.d.ts +44 -0
- package/dist/monitoring-settings.js +173 -0
- package/dist/monitoring-ui.d.ts +20 -0
- package/dist/monitoring-ui.js +51 -0
- package/dist/monitoring.d.ts +34 -0
- package/dist/monitoring.js +221 -0
- package/dist/names.d.ts +67 -0
- package/dist/names.js +102 -0
- package/dist/observations.d.ts +40 -0
- package/dist/observations.js +211 -0
- package/dist/oidc.d.ts +71 -0
- package/dist/oidc.js +182 -0
- package/dist/onboard.d.ts +108 -0
- package/dist/onboard.js +325 -0
- package/dist/openrouter-models.d.ts +30 -0
- package/dist/openrouter-models.js +120 -0
- package/dist/operate.d.ts +118 -0
- package/dist/operate.js +11321 -0
- package/dist/peek-cli.d.ts +97 -0
- package/dist/peek-cli.js +215 -0
- package/dist/peek.d.ts +134 -0
- package/dist/peek.js +578 -0
- package/dist/phase-routing.d.ts +307 -0
- package/dist/phase-routing.js +658 -0
- package/dist/plan-auto.d.ts +14 -0
- package/dist/plan-auto.js +92 -0
- package/dist/plan.d.ts +186 -0
- package/dist/plan.js +401 -0
- package/dist/planner-source.d.ts +251 -0
- package/dist/planner-source.js +460 -0
- package/dist/planner.d.ts +75 -0
- package/dist/planner.js +992 -0
- package/dist/policy-ui.d.ts +20 -0
- package/dist/policy-ui.js +41 -0
- package/dist/policy.d.ts +111 -0
- package/dist/policy.js +231 -0
- package/dist/prepared-evidence.d.ts +16 -0
- package/dist/prepared-evidence.js +96 -0
- package/dist/pricing.d.ts +45 -0
- package/dist/pricing.js +77 -0
- package/dist/principal.d.ts +42 -0
- package/dist/principal.js +82 -0
- package/dist/probe.d.ts +46 -0
- package/dist/probe.js +79 -0
- package/dist/process-custody.d.ts +17 -0
- package/dist/process-custody.js +91 -0
- package/dist/process-liveness.d.ts +10 -0
- package/dist/process-liveness.js +81 -0
- package/dist/process-recovery-anchor.d.ts +75 -0
- package/dist/process-recovery-anchor.js +188 -0
- package/dist/process-recovery-coalition.d.ts +25 -0
- package/dist/process-recovery-coalition.js +105 -0
- package/dist/process-recovery-eligibility.d.ts +178 -0
- package/dist/process-recovery-eligibility.js +170 -0
- package/dist/process-recovery-native.d.ts +172 -0
- package/dist/process-recovery-native.js +592 -0
- package/dist/process-recovery-provenance.d.ts +163 -0
- package/dist/process-recovery-provenance.js +292 -0
- package/dist/process-recovery-services.d.ts +56 -0
- package/dist/process-recovery-services.js +191 -0
- package/dist/process-recovery-settlement.d.ts +32 -0
- package/dist/process-recovery-settlement.js +132 -0
- package/dist/process-recovery.d.ts +41 -0
- package/dist/process-recovery.js +91 -0
- package/dist/process-tree.d.ts +41 -0
- package/dist/process-tree.js +264 -0
- package/dist/project-access.d.ts +14 -0
- package/dist/project-access.js +34 -0
- package/dist/project-cli.d.ts +18 -0
- package/dist/project-cli.js +104 -0
- package/dist/project-delete-ui.d.ts +22 -0
- package/dist/project-delete-ui.js +48 -0
- package/dist/project-delete.d.ts +55 -0
- package/dist/project-delete.js +413 -0
- package/dist/project-knowledge.d.ts +111 -0
- package/dist/project-knowledge.js +241 -0
- package/dist/project-learning.d.ts +100 -0
- package/dist/project-learning.js +438 -0
- package/dist/project-memory.d.ts +81 -0
- package/dist/project-memory.js +264 -0
- package/dist/project-skills.d.ts +152 -0
- package/dist/project-skills.js +660 -0
- package/dist/project-tools.d.ts +219 -0
- package/dist/project-tools.js +796 -0
- package/dist/project.d.ts +61 -0
- package/dist/project.js +125 -0
- package/dist/prompt.d.ts +22 -0
- package/dist/prompt.js +72 -0
- package/dist/proof.d.ts +513 -0
- package/dist/proof.js +1140 -0
- package/dist/proposal.d.ts +82 -0
- package/dist/proposal.js +210 -0
- package/dist/provider-connection.d.ts +22 -0
- package/dist/provider-connection.js +98 -0
- package/dist/provider-limits.d.ts +35 -0
- package/dist/provider-limits.js +95 -0
- package/dist/provider.d.ts +279 -0
- package/dist/provider.js +935 -0
- package/dist/publish.d.ts +107 -0
- package/dist/publish.js +650 -0
- package/dist/pulls.d.ts +118 -0
- package/dist/pulls.js +240 -0
- package/dist/push.d.ts +95 -0
- package/dist/push.js +353 -0
- package/dist/quality.d.ts +8 -0
- package/dist/quality.js +6 -0
- package/dist/recipe-ui.d.ts +18 -0
- package/dist/recipe-ui.js +152 -0
- package/dist/recipes.d.ts +81 -0
- package/dist/recipes.js +332 -0
- package/dist/remote.d.ts +67 -0
- package/dist/remote.js +122 -0
- package/dist/render.d.ts +64 -0
- package/dist/render.js +587 -0
- package/dist/report-summary.d.ts +13 -0
- package/dist/report-summary.js +15 -0
- package/dist/repos.d.ts +65 -0
- package/dist/repos.js +277 -0
- package/dist/repository-context-ui.d.ts +2 -0
- package/dist/repository-context-ui.js +10 -0
- package/dist/repository-context.d.ts +67 -0
- package/dist/repository-context.js +376 -0
- package/dist/restart-certification.d.ts +134 -0
- package/dist/restart-certification.js +226 -0
- package/dist/result-actions.d.ts +28 -0
- package/dist/result-actions.js +270 -0
- package/dist/result-completion.d.ts +18 -0
- package/dist/result-completion.js +66 -0
- package/dist/result-review.d.ts +213 -0
- package/dist/result-review.js +430 -0
- package/dist/retention-ui.d.ts +12 -0
- package/dist/retention-ui.js +39 -0
- package/dist/retention.d.ts +72 -0
- package/dist/retention.js +288 -0
- package/dist/review-context.d.ts +243 -0
- package/dist/review-context.js +1043 -0
- package/dist/review-evidence.d.ts +33 -0
- package/dist/review-evidence.js +56 -0
- package/dist/reviewer.d.ts +97 -0
- package/dist/reviewer.js +202 -0
- package/dist/routine.d.ts +228 -0
- package/dist/routine.js +764 -0
- package/dist/runner.d.ts +266 -0
- package/dist/runner.js +399 -0
- package/dist/scan.d.ts +32 -0
- package/dist/scan.js +92 -0
- package/dist/scope.d.ts +695 -0
- package/dist/scope.js +1176 -0
- package/dist/scout-report.d.ts +42 -0
- package/dist/scout-report.js +108 -0
- package/dist/scout.d.ts +69 -0
- package/dist/scout.js +351 -0
- package/dist/serve.d.ts +313 -0
- package/dist/serve.js +21776 -0
- package/dist/session-brief.d.ts +8 -0
- package/dist/session-brief.js +24 -0
- package/dist/session-cli.d.ts +13 -0
- package/dist/session-cli.js +284 -0
- package/dist/session-contract.d.ts +183 -0
- package/dist/session-contract.js +122 -0
- package/dist/session-http.d.ts +19 -0
- package/dist/session-http.js +154 -0
- package/dist/session-server.d.ts +12 -0
- package/dist/session-server.js +37 -0
- package/dist/session-service.d.ts +22 -0
- package/dist/session-service.js +157 -0
- package/dist/setup-guide.d.ts +39 -0
- package/dist/setup-guide.js +93 -0
- package/dist/sign-in-guard.d.ts +45 -0
- package/dist/sign-in-guard.js +74 -0
- package/dist/skills-ui.d.ts +11 -0
- package/dist/skills-ui.js +58 -0
- package/dist/skills.d.ts +97 -0
- package/dist/skills.js +276 -0
- package/dist/slack-api.d.ts +55 -0
- package/dist/slack-api.js +229 -0
- package/dist/slack-chat.d.ts +28 -0
- package/dist/slack-chat.js +413 -0
- package/dist/slack-settings.d.ts +7 -0
- package/dist/slack-settings.js +75 -0
- package/dist/slack-state.d.ts +13 -0
- package/dist/slack-state.js +9 -0
- package/dist/slack.d.ts +13 -0
- package/dist/slack.js +167 -0
- package/dist/spend-ui.d.ts +28 -0
- package/dist/spend-ui.js +95 -0
- package/dist/spend.d.ts +142 -0
- package/dist/spend.js +313 -0
- package/dist/sqlite-runtime.d.ts +3 -0
- package/dist/sqlite-runtime.js +6 -0
- package/dist/sso-settings.d.ts +33 -0
- package/dist/sso-settings.js +103 -0
- package/dist/sso-ui.d.ts +22 -0
- package/dist/sso-ui.js +39 -0
- package/dist/storage.d.ts +15 -0
- package/dist/storage.js +84 -0
- package/dist/store.d.ts +7754 -0
- package/dist/store.js +23279 -0
- package/dist/structured-output.d.ts +58 -0
- package/dist/structured-output.js +103 -0
- package/dist/style-asset.d.ts +7 -0
- package/dist/style-asset.js +41 -0
- package/dist/subscription-chat.d.ts +46 -0
- package/dist/subscription-chat.js +275 -0
- package/dist/summary.d.ts +33 -0
- package/dist/summary.js +71 -0
- package/dist/supervisor.mjs +297 -0
- package/dist/surface.d.ts +63 -0
- package/dist/surface.js +352 -0
- package/dist/sync.d.ts +94 -0
- package/dist/sync.js +279 -0
- package/dist/task-composer.d.ts +13 -0
- package/dist/task-composer.js +52 -0
- package/dist/task-control.d.ts +165 -0
- package/dist/task-control.js +214 -0
- package/dist/task-outcome-cli.d.ts +10 -0
- package/dist/task-outcome-cli.js +92 -0
- package/dist/task-text.d.ts +30 -0
- package/dist/task-text.js +41 -0
- package/dist/team-cli.d.ts +15 -0
- package/dist/team-cli.js +618 -0
- package/dist/team-contract.d.ts +92 -0
- package/dist/team-contract.js +1 -0
- package/dist/team-http.d.ts +17 -0
- package/dist/team-http.js +144 -0
- package/dist/team-leads.d.ts +81 -0
- package/dist/team-leads.js +565 -0
- package/dist/team-runtime.d.ts +27 -0
- package/dist/team-runtime.js +264 -0
- package/dist/team-ui.d.ts +3 -0
- package/dist/team-ui.js +9 -0
- package/dist/team-updates.d.ts +8 -0
- package/dist/team-updates.js +57 -0
- package/dist/teammate-admin.d.ts +53 -0
- package/dist/teammate-admin.js +111 -0
- package/dist/teammate-desk.d.ts +60 -0
- package/dist/teammate-desk.js +154 -0
- package/dist/teammate-memory.d.ts +64 -0
- package/dist/teammate-memory.js +196 -0
- package/dist/teammate-question.d.ts +59 -0
- package/dist/teammate-question.js +175 -0
- package/dist/teammate-tools.d.ts +108 -0
- package/dist/teammate-tools.js +271 -0
- package/dist/teammate-week.d.ts +68 -0
- package/dist/teammate-week.js +132 -0
- package/dist/teammate-work.d.ts +76 -0
- package/dist/teammate-work.js +280 -0
- package/dist/teammates-ui.d.ts +20 -0
- package/dist/teammates-ui.js +166 -0
- package/dist/teammates.d.ts +180 -0
- package/dist/teammates.js +253 -0
- package/dist/teams-api.d.ts +33 -0
- package/dist/teams-api.js +204 -0
- package/dist/teams-chat.d.ts +26 -0
- package/dist/teams-chat.js +215 -0
- package/dist/teams-settings.d.ts +12 -0
- package/dist/teams-settings.js +57 -0
- package/dist/teams.d.ts +21 -0
- package/dist/teams.js +124 -0
- package/dist/telegram-flow.d.ts +49 -0
- package/dist/telegram-flow.js +110 -0
- package/dist/telegram-mate.d.ts +137 -0
- package/dist/telegram-mate.js +707 -0
- package/dist/telegram-progress.d.ts +14 -0
- package/dist/telegram-progress.js +129 -0
- package/dist/telegram-settings.d.ts +9 -0
- package/dist/telegram-settings.js +36 -0
- package/dist/telegram-status.d.ts +75 -0
- package/dist/telegram-status.js +265 -0
- package/dist/telegram-team.d.ts +85 -0
- package/dist/telegram-team.js +359 -0
- package/dist/telegram.d.ts +230 -0
- package/dist/telegram.js +1931 -0
- package/dist/templates.d.ts +65 -0
- package/dist/templates.js +118 -0
- package/dist/tool-launcher.d.ts +1 -0
- package/dist/tool-launcher.js +101 -0
- package/dist/tools-ui.d.ts +30 -0
- package/dist/tools-ui.js +40 -0
- package/dist/transitions-recipes.d.ts +1 -0
- package/dist/transitions-recipes.js +203 -0
- package/dist/tree-proof.d.ts +34 -0
- package/dist/tree-proof.js +58 -0
- package/dist/verification-evidence.d.ts +44 -0
- package/dist/verification-evidence.js +238 -0
- package/dist/version.d.ts +2 -0
- package/dist/version.js +10 -0
- package/dist/webhooks.d.ts +91 -0
- package/dist/webhooks.js +283 -0
- package/dist/work-index.d.ts +68 -0
- package/dist/work-index.js +384 -0
- package/dist/work-summary.d.ts +57 -0
- package/dist/work-summary.js +116 -0
- package/dist/workspace-motion.d.ts +4 -0
- package/dist/workspace-motion.js +133 -0
- package/dist/workspace-revision.d.ts +19 -0
- package/dist/workspace-revision.js +92 -0
- package/dist/workspace-ui.d.ts +204 -0
- package/dist/workspace-ui.js +411 -0
- package/dist/worktree-notices.d.ts +13 -0
- package/dist/worktree-notices.js +17 -0
- package/dist/worktree.d.ts +249 -0
- package/dist/worktree.js +720 -0
- package/package.json +122 -0
- package/scripts/canary-assertions.mjs +42 -0
- package/scripts/crash-canary.mjs +253 -0
- package/scripts/fixtures/crash-process.mjs +74 -0
- package/scripts/fixtures/pilot-scenarios.mjs +85 -0
- package/scripts/fixtures/restart-service.mjs +27 -0
- package/scripts/launchd-certification.mjs +143 -0
- package/scripts/pilot.mjs +121 -0
- package/scripts/proof-preflight.mjs +140 -0
- package/scripts/provider-canary.mjs +379 -0
- package/scripts/recovery-canary.mjs +135 -0
- package/scripts/restart-certification.mjs +98 -0
package/README.md
ADDED
|
@@ -0,0 +1,1113 @@
|
|
|
1
|
+
<div align="center">
|
|
2
|
+
|
|
3
|
+
<picture>
|
|
4
|
+
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/ap9000/standing-orders/main/docs/media/wordmark-dark.svg">
|
|
5
|
+
<img src="https://raw.githubusercontent.com/ap9000/standing-orders/main/docs/media/wordmark-light.svg" alt="Toolroll — a control plane for unattended coding agents" width="480">
|
|
6
|
+
</picture>
|
|
7
|
+
|
|
8
|
+
**Queue twelve tasks, walk away, come back to pull requests —
|
|
9
|
+
interrupted only for decisions that genuinely need a human.**
|
|
10
|
+
|
|
11
|
+
[](https://github.com/ap9000/standing-orders/actions/workflows/ci.yml)
|
|
12
|
+
[](https://www.npmjs.com/package/toolroll)
|
|
13
|
+

|
|
14
|
+
[](LICENSE)
|
|
15
|
+
|
|
16
|
+
[Guides](docs/guide/README.md) · [Design](docs/DESIGN.md) · [Never Stuck contract](docs/NEVER_STUCK.md) · [Priorities](docs/PRIORITIES.md) · [Ledger](docs/PROGRESS.md) · [Contributing](CONTRIBUTING.md) · [Issues](https://github.com/ap9000/standing-orders/issues) · [npm](https://www.npmjs.com/package/toolroll)
|
|
17
|
+
|
|
18
|
+
<img src="https://raw.githubusercontent.com/ap9000/standing-orders/main/docs/media/ui/unified-chat.png" alt="Toolroll unified chat showing a live portfolio overview across projects, active builds, decisions, and proposed next actions." width="920">
|
|
19
|
+
|
|
20
|
+
<sub>One conversation across every project, backed by durable tasks—not a chat-only copy of the work.</sub>
|
|
21
|
+
|
|
22
|
+
</div>
|
|
23
|
+
|
|
24
|
+
## Quick start
|
|
25
|
+
|
|
26
|
+
```sh
|
|
27
|
+
curl -fsSL https://raw.githubusercontent.com/ap9000/standing-orders/main/install.sh | sh
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
macOS or Linux, with Node.js 22.13+ and git. It installs the command and opens
|
|
31
|
+
the console; `npx toolroll demo` shows a seeded sandbox first. Then read
|
|
32
|
+
[Getting started](docs/guide/getting-started.md), [Flows](docs/guide/flows.md)
|
|
33
|
+
and [Security](docs/guide/security.md).
|
|
34
|
+
|
|
35
|
+
## One command center, the whole loop
|
|
36
|
+
|
|
37
|
+
Tell Toolroll what outcome you want. Its planner reads the repository,
|
|
38
|
+
drafts the scope and proof rubric, and asks only when an answer would materially
|
|
39
|
+
change the work. You approve the exact contract once; long-running agents can
|
|
40
|
+
build it while the control plane handles queues, dependencies,
|
|
41
|
+
crashes, and decisions. The result comes back with the diff, checks, screenshots,
|
|
42
|
+
and a criterion-by-criterion verdict.
|
|
43
|
+
|
|
44
|
+
### Hand off in one prompt; approve exactly what will run
|
|
45
|
+
|
|
46
|
+
Project and quality stay close to the prompt; expert fields appear only when
|
|
47
|
+
you open them. Before execution, the approval card restates the goal,
|
|
48
|
+
boundaries, evidence requirements, model, and permissions.
|
|
49
|
+
|
|
50
|
+
<div align="center">
|
|
51
|
+
<img src="https://raw.githubusercontent.com/ap9000/standing-orders/main/docs/media/ui/task-handoff-mobile.png" alt="Mobile chat-first task composer with one outcome prompt and progressive details." width="260">
|
|
52
|
+
|
|
53
|
+
<img src="https://raw.githubusercontent.com/ap9000/standing-orders/main/docs/media/ui/scope-approval.png" alt="Desktop scope approval card showing the goal, boundaries, evidence requirement, model, permissions, and approval action." width="640">
|
|
54
|
+
</div>
|
|
55
|
+
|
|
56
|
+
### One lead, durable crew work
|
|
57
|
+
|
|
58
|
+
Start in **Chat**, inspect work in **Tasks**, and manage guidance in **Projects**.
|
|
59
|
+
Browser and CLI chat share the lead conversation. A local database catch-up shows
|
|
60
|
+
current decisions and results before you authorize any model usage.
|
|
61
|
+
|
|
62
|
+
Crew work ends at **Ready**. Open the result, mark it **Complete**, or request a
|
|
63
|
+
specific revision of the same task. Actual checks remain visible; there is no
|
|
64
|
+
separate model reviewer or automatic evidence resubmission loop. Deployment is
|
|
65
|
+
reported separately and still requires the exact passing native machine check.
|
|
66
|
+
|
|
67
|
+
Enable **Automatic crew updates** in a conversation, or use `toolroll chat
|
|
68
|
+
--follow`. Meaningful updates reach the same lead within its saved permissions
|
|
69
|
+
and limits. Idle scanning uses no model. Restart recovery reuses saved responses;
|
|
70
|
+
a failed response does not rerun a task. `--no-follow` pauses automatic responses.
|
|
71
|
+
This wakes Toolroll's own lead, not an unrelated external agent session.
|
|
72
|
+
|
|
73
|
+
`toolroll brief` reads the local database. `knowledge search`, `knowledge
|
|
74
|
+
impact`, and `knowledge refresh` add bounded source retrieval; see
|
|
75
|
+
[repository context](docs/REPOSITORY_CONTEXT.md). Curated knowledge stays in the
|
|
76
|
+
database, and crew context stays attached to its original run.
|
|
77
|
+
|
|
78
|
+
Select a saved checkout once, then use the same task records from the terminal:
|
|
79
|
+
|
|
80
|
+
```sh
|
|
81
|
+
toolroll project use /path/to/project
|
|
82
|
+
toolroll project show
|
|
83
|
+
toolroll brief
|
|
84
|
+
toolroll assignment show <task>
|
|
85
|
+
toolroll task complete <task>
|
|
86
|
+
toolroll task revise <task> --feedback "Keep the filter selected after reload."
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
Completion and revisions use the local sign-in already saved by `up`. Completion
|
|
90
|
+
reports the exact run, commit and check outcome it handled. Agents use a scoped
|
|
91
|
+
credential (`project use /path/to/project --token-file /path/to/credential`) and
|
|
92
|
+
complete only with the exact `--digest` from `assignment show`; JSON completion
|
|
93
|
+
also requires that digest. The profile stores a file reference, never a secret,
|
|
94
|
+
and grants no additional access. It supplies defaults for database briefs,
|
|
95
|
+
knowledge commands, assignment commands, and these two task actions.
|
|
96
|
+
|
|
97
|
+
For an exact revision replay, reuse `--run`, `--source` and `--key` from its JSON
|
|
98
|
+
result. A revision keeps the original task history and approval boundaries; a
|
|
99
|
+
stale result cannot create another current version. Actual failed checks remain
|
|
100
|
+
failed when a result is marked complete.
|
|
101
|
+
|
|
102
|
+
### What is in the current build
|
|
103
|
+
|
|
104
|
+
- **Guided workflow recipes.** Choose a starter, customize its outcome and
|
|
105
|
+
success checks, and preview the steps before creating one-time or scheduled
|
|
106
|
+
work. Save project recipes for teammates, reuse a task's scope, or share a
|
|
107
|
+
portable JSON definition. Creation retries return the same work; existing
|
|
108
|
+
approval and recovery rules apply. [Get started](docs/WORKFLOW_RECIPES.md).
|
|
109
|
+
- **A creator for your usual work.** Write instructions once and add questions
|
|
110
|
+
for the parts that change, such as a module or feature. Save the recipe,
|
|
111
|
+
then use a short answer form to kick off another run. Saved recipes are
|
|
112
|
+
searchable, show recently used work first, and can be shared with teammates.
|
|
113
|
+
- **Flows: your process on a canvas.** Draw zones (triage, a person's
|
|
114
|
+
go-ahead, build, review, tell the team) and drop cards into them. Each zone
|
|
115
|
+
runs its step: build and research zones file ordinary tasks, so every
|
|
116
|
+
approval, check and agent fence still applies; decision zones wait for a
|
|
117
|
+
named person; a send back becomes a revision of the same work. The engine
|
|
118
|
+
is model-free and runs in the worker's pass. Or describe the process in
|
|
119
|
+
chat: the lead drafts the flow, and adds, moves or decides cards, as
|
|
120
|
+
cards you confirm from the console or your phone.
|
|
121
|
+
- **Triggers start cards on their own.** A button with a few questions, a
|
|
122
|
+
schedule, GitHub (new issues, a label being added, new pull requests,
|
|
123
|
+
failed checks), Linear (a team, a state, a label), or another flow's cards
|
|
124
|
+
reaching a zone. GitHub is checked through your `gh` login and Linear
|
|
125
|
+
with an API key kept on this computer; neither spends model tokens. Each
|
|
126
|
+
issue or run makes at most one card; text from outside the repository's
|
|
127
|
+
team is left out unless you say anyone. With a public address, GitHub,
|
|
128
|
+
Linear or any service can post to a trigger's secret webhook instead
|
|
129
|
+
(see [Webhooks](#webhooks-through-a-reverse-proxy)). A button can also be
|
|
130
|
+
shared as a secret form link, so someone without an account can report
|
|
131
|
+
something straight into a flow.
|
|
132
|
+
- **Steps with no AI, and where flows break.** A project keeps a library of
|
|
133
|
+
scripts — `run-tests`, `lint`, `smoke-staging` — written on a flow's
|
|
134
|
+
Scripts panel or drafted by the lead in chat, and any flow runs one from a
|
|
135
|
+
"Run a script" zone in a fresh copy of the card's work: exit 0 passes,
|
|
136
|
+
anything else takes the failure path with the log kept. A Build whose
|
|
137
|
+
project checks fail takes its failure path too, and an "Update the issue"
|
|
138
|
+
zone comments on and closes the GitHub or Linear issue a card came from.
|
|
139
|
+
Each flow's Insights show, per zone, how many cards passed, failed or were
|
|
140
|
+
sent back and how long they waited, how each script does, and every run's
|
|
141
|
+
log; the lead reads the same numbers to tell you where things break.
|
|
142
|
+
- **Sorting with Jev.** A "Sort" zone asks [Jev](https://openrouter.ai/typesafe),
|
|
143
|
+
TypeSafe's decision model, one question about each card through your own
|
|
144
|
+
OpenRouter key: which of the zone's answers fits, how sure it is, and a few
|
|
145
|
+
scores or yes/no notes (how urgent, asking for a refund). It answers in
|
|
146
|
+
well under a second for a fraction of a cent, and can only ever pick one
|
|
147
|
+
of your answers. Cards it's sure about go where their answer leads; the
|
|
148
|
+
rest wait for a person, and every card shows what Jev decided. Insights
|
|
149
|
+
count how often people moved a sorted card elsewhere, by how sure Jev
|
|
150
|
+
was, so you can see when to trust it more. Five templates start from it:
|
|
151
|
+
issue triage, a spam filter for public forms, lead routing, effort routing
|
|
152
|
+
(small changes straight to a build, big ones through a plan) and exception
|
|
153
|
+
routing (orders, invoices, deliveries).
|
|
154
|
+
- **Drafts, approved from your phone.** A "Draft" zone has Claude write a
|
|
155
|
+
reply, summary or note from the card in a few seconds, with no repository
|
|
156
|
+
and no tools, through the lead chat's sign-in. Nothing is sent by itself: the
|
|
157
|
+
next "Person decides" zone puts the draft in front of the flow's owner (or
|
|
158
|
+
whoever it names) in their chat app. On Telegram the message carries
|
|
159
|
+
Approve, Edit and Send back: Edit takes your own version as a reply and
|
|
160
|
+
brings it back to approve, and Send back takes a note and has Claude try
|
|
161
|
+
again. In the console the same draft is an editable box on the card. Each
|
|
162
|
+
flow has an owner (whoever made it, until handed on) whom these decisions
|
|
163
|
+
go to.
|
|
164
|
+
- **Steps that reach outside.** A "Web request" zone calls an API with the
|
|
165
|
+
card's details (its host is written out, so a card can't redirect it;
|
|
166
|
+
secrets you save on the step go only into its headers). A "Send email" zone
|
|
167
|
+
mails through your own mail server (Settings → Email: Gmail, Outlook,
|
|
168
|
+
Fastmail, Resend, SES or any SMTP), and {{card.email}} is the address a card
|
|
169
|
+
mentions. A "Use a tool" zone calls one of the project's MCP servers from
|
|
170
|
+
the Tools page, like posting to Slack or adding a page to Notion. The Email
|
|
171
|
+
replies template puts it together: Claude drafts, the flow's owner approves
|
|
172
|
+
from their phone, and the reply is emailed.
|
|
173
|
+
- **Unified portfolio chat.** Read every project, prioritize queues, answer
|
|
174
|
+
decisions, repair failed or cancelled dependencies, and confirm rich action
|
|
175
|
+
cards from one conversation. The chat proposes; durable workflow state
|
|
176
|
+
remains the source of truth.
|
|
177
|
+
- **Chat-first task handoff.** The default form is one outcome prompt. The
|
|
178
|
+
repository-aware planner drafts the goal, boundaries, acceptance criteria,
|
|
179
|
+
likely files, and a concise execution plan with milestones, dependencies,
|
|
180
|
+
risks, and proof. Review or refine that plan before starting; approval locks
|
|
181
|
+
the exact revision the builder receives. Expert controls remain under
|
|
182
|
+
**Edit details**.
|
|
183
|
+
- **A plan that adapts without quietly widening what you signed.** While a
|
|
184
|
+
build runs, it checkpoints which milestone is pending, in progress, done,
|
|
185
|
+
or blocked — reported by the agent, never counted as completion proof. When
|
|
186
|
+
the repository shows a stated dependency, risk, or approach was wrong, the
|
|
187
|
+
builder can file one evidence-linked replacement plan and pause at a safe
|
|
188
|
+
point. A plan-only refinement appends an immutable revision and resumes on
|
|
189
|
+
its own; anything that would touch the goal, boundaries, touches,
|
|
190
|
+
acceptance, permissions, quality, budget, or publication authority stays
|
|
191
|
+
paused for your accept or reject. The task page and focused chat always
|
|
192
|
+
show the same live progress, the same revision, and the same pending
|
|
193
|
+
decision.
|
|
194
|
+
- **Structured handoffs repair their shape, not their meaning.** A malformed
|
|
195
|
+
planner reply gets conservative syntax normalization, then at
|
|
196
|
+
most two correction turns in the same session with the exact validation
|
|
197
|
+
errors. Corrections cannot invent scope or criteria; every reply is sealed
|
|
198
|
+
for audit, including replies rejected by provider or session checks, and
|
|
199
|
+
workspace integrity is re-proved
|
|
200
|
+
after each turn.
|
|
201
|
+
- **Long-running, recoverable execution.** There is no arbitrary task
|
|
202
|
+
countdown. Installed workers survive terminal closure and reboot, recover
|
|
203
|
+
expired claims, and continue until a terminal result, a real decision, or a
|
|
204
|
+
signed no-progress/runaway breaker.
|
|
205
|
+
- **Subscription-native agents.** Use the Codex and Claude logins already on
|
|
206
|
+
the machine. Dollar caps are optional; subscription usage is labeled as an
|
|
207
|
+
API-price equivalent, never presented as an API charge.
|
|
208
|
+
- **Per-task autonomy with a global default.** Choose **Auto** or **Full
|
|
209
|
+
access** for the installation, then override it on any task. The exact
|
|
210
|
+
provider permission mode is sealed into the approved scope.
|
|
211
|
+
- **Explicit quality choices.** Default uses configured everyday agents;
|
|
212
|
+
Strict / release requests stronger configured agents. The repository check
|
|
213
|
+
remains authoritative for checks. Quality never starts a reviewer loop.
|
|
214
|
+
- **Revisions retain context.** Feedback, source result and selected project
|
|
215
|
+
context remain attached to the task. Each requested revision receives its
|
|
216
|
+
own scope and approval; earlier results remain available.
|
|
217
|
+
- **Honest completion.** Ready results include actual checks and saved work.
|
|
218
|
+
The lead or user marks Complete after inspection. Failed checks stay failed;
|
|
219
|
+
absent optional historical assessments do not block completion.
|
|
220
|
+
|
|
221
|
+
<div align="center">
|
|
222
|
+
<img src="https://raw.githubusercontent.com/ap9000/standing-orders/main/docs/media/ui/verified-result.png" alt="A verified Toolroll result card with checks, follow-up notes, and evidence-backed completion details." width="720">
|
|
223
|
+
<br>
|
|
224
|
+
<sub>Inspect the saved result and actual checks, then mark Complete or request changes.</sub>
|
|
225
|
+
</div>
|
|
226
|
+
|
|
227
|
+
## Install
|
|
228
|
+
|
|
229
|
+
Choose your projects folder the first time:
|
|
230
|
+
|
|
231
|
+
```sh
|
|
232
|
+
npx toolroll up --project-root ~/Projects # or: bunx toolroll up
|
|
233
|
+
```
|
|
234
|
+
|
|
235
|
+
That is the install and the setup. It needs Node 22.13 or newer on the
|
|
236
|
+
machine (Bun's runtime has no `node:sqlite`; `bunx` hands the shebang to
|
|
237
|
+
Node, so it works too). `npm install -g toolroll` gives you the
|
|
238
|
+
bare `toolroll` command for later (`standing-orders`, the older name, still
|
|
239
|
+
works).
|
|
240
|
+
|
|
241
|
+
`up` prints your login once (and saves it beside the database as
|
|
242
|
+
`up-login.txt`), opens the app in your browser, and connects this machine as
|
|
243
|
+
the builder. Add an existing folder or a GitHub repository from **Projects**;
|
|
244
|
+
the builder and unified chat pick it up while the app keeps running. The
|
|
245
|
+
projects folder and every added repository are remembered. Later,
|
|
246
|
+
`toolroll up` can be run from any directory and reconnects all of them.
|
|
247
|
+
|
|
248
|
+
To reach it from your phone over a tailnet:
|
|
249
|
+
`toolroll up --host 0.0.0.0 --allow-host <your-machine>.ts.net:4180`.
|
|
250
|
+
|
|
251
|
+
If the inbox says **Builder disconnected**, reopen Toolroll on the
|
|
252
|
+
machine where the projects live; queued work resumes automatically. You do not
|
|
253
|
+
run `up` separately in each project.
|
|
254
|
+
|
|
255
|
+
Advanced deployments can start the console alone with
|
|
256
|
+
`toolroll serve --repo .`: with no account yet it
|
|
257
|
+
prints a six-digit setup code, and the login page offers **create the
|
|
258
|
+
first account** — enter the code, pick a username and password, and you are
|
|
259
|
+
in. A second person joins by invite link from the people page, never by
|
|
260
|
+
another setup code.
|
|
261
|
+
|
|
262
|
+
```sh
|
|
263
|
+
npx toolroll demo # a seeded sandbox — see it working in 90 seconds, zero spend
|
|
264
|
+
npx toolroll # what's in flight across your repos — read-only, zero config
|
|
265
|
+
```
|
|
266
|
+
|
|
267
|
+
## Getting started
|
|
268
|
+
|
|
269
|
+
There is one normal road: keep one `toolroll up` running on the machine.
|
|
270
|
+
It is the app and the builder for every saved project. The separate console,
|
|
271
|
+
worker, and OS service commands documented later are advanced deployment tools
|
|
272
|
+
for people splitting those parts across machines.
|
|
273
|
+
|
|
274
|
+
### In the console
|
|
275
|
+
|
|
276
|
+
1. **Sign in** with the login `up` printed. You land on the **inbox**:
|
|
277
|
+
everything that waits on you, and nothing else.
|
|
278
|
+
2. **Describe the outcome** with **+ new task**. The normal path is one
|
|
279
|
+
ChatGPT-style prompt: the planner inspects the open repository, drafts the
|
|
280
|
+
goal, boundaries, acceptance rubric, and implementation approach, then asks
|
|
281
|
+
only when a missing answer would materially change the work. Project and
|
|
282
|
+
quality stay in the compact footer; **Edit details** reveals the full
|
|
283
|
+
contract, research-only mode, permissions, dependencies, and expert fields.
|
|
284
|
+
3. **Review and approve the proposed scope.** A scope needs at least one
|
|
285
|
+
criterion — a plain outcome statement and the evidence kind (check,
|
|
286
|
+
screenshot, changed-path, or manual review) that will answer it — before it
|
|
287
|
+
can be signed. The task page leads with a concise approval card; the full
|
|
288
|
+
structured, editable execution plan and full contract remain one click
|
|
289
|
+
away. It restates the exact scope AND rubric you are signing, the provider
|
|
290
|
+
and model it will run on, and your password. Nothing spends a token until
|
|
291
|
+
this yes.
|
|
292
|
+
4. **Watch it build.** The **board** moves the card to *building*; the
|
|
293
|
+
card's own page shows the stage, the live transcript, and — once the
|
|
294
|
+
agent checkpoints one — which milestone is pending, in progress, done,
|
|
295
|
+
or blocked; **peek** (`/peek`, or *peek at the live ones →* on the builds
|
|
296
|
+
page) shows every live agent at once. If the repository disproves a
|
|
297
|
+
named dependency or risk, the task page shows the replacement plan with
|
|
298
|
+
its evidence; a plan-only fix resumes on its own, while anything that
|
|
299
|
+
would touch what you signed waits for your accept or reject, right
|
|
300
|
+
there next to the plan.
|
|
301
|
+
5. **Answer when asked.** An agent that hits a judgement call parks a typed
|
|
302
|
+
decision — question, options, consequences, which are reversible. It
|
|
303
|
+
arrives in the inbox, on `/next`, and on your phone if Telegram is
|
|
304
|
+
paired; one tap answers it and the build resumes.
|
|
305
|
+
6. **Collect the result.** A build becomes Ready with its saved diff, checks,
|
|
306
|
+
available screenshots and limitations. Open the result to mark Complete or
|
|
307
|
+
request a specific revision. A scout returns a report. Publication requires
|
|
308
|
+
its own approved grant; task completion alone never publishes or deploys.
|
|
309
|
+
|
|
310
|
+
Verification can recover one common environment failure without hiding it.
|
|
311
|
+
Run `toolroll verify set ... --self-heal` without `--yes` first. The
|
|
312
|
+
preview shows the exact approved setup and its digest; confirm only that
|
|
313
|
+
preview by rerunning with `--setup-digest <shown> --yes`. If the project
|
|
314
|
+
check cannot start because a required project executable is missing,
|
|
315
|
+
Toolroll may run that setup once and retry the exact check once. It
|
|
316
|
+
does not recover ordinary test failures, timeouts, or an executable that
|
|
317
|
+
exists but cannot run. Every step stays in the check log. Recovery stops if
|
|
318
|
+
the setup or project check changes, setup fails, files change, the checkout
|
|
319
|
+
moves, unchanged files cannot be confirmed, the worker loses custody, or
|
|
320
|
+
the executable is still missing.
|
|
321
|
+
7. **Inspect without the transcript.** Open the task's result for its summary,
|
|
322
|
+
changes and checks. Annotate exact lines and request a revision when needed.
|
|
323
|
+
Saved history and diagnostics remain available without becoming another
|
|
324
|
+
mandatory review step. Older results retain their own links.
|
|
325
|
+
|
|
326
|
+
Specialized views remain in **Tools** and **Settings**: the activity ledger, the review cockpit,
|
|
327
|
+
routines (tasks that file themselves on a schedule), the fleet,
|
|
328
|
+
people (invite a second approver), the operating mode (a signed, expiring
|
|
329
|
+
envelope that pre-approves your own filings), and **chat** — the mate, one
|
|
330
|
+
conversation across every project, which only ever proposes.
|
|
331
|
+
|
|
332
|
+
### In the terminal
|
|
333
|
+
|
|
334
|
+
```sh
|
|
335
|
+
toolroll task add "Give outbound webhooks a bounded retry policy" --id retries --repo .
|
|
336
|
+
toolroll task scope retries --goal "Exponential backoff, dead-letter after 24h, no payload changes" \
|
|
337
|
+
--acceptance "A failing webhook retries with exponential backoff and dead-letters after 24h.|check"
|
|
338
|
+
toolroll task show retries --json # the scope's digest is what you sign
|
|
339
|
+
toolroll task approve retries --as you --digest <digest> --yes # asks for your password
|
|
340
|
+
|
|
341
|
+
toolroll peek # one pane per live agent; q leaves
|
|
342
|
+
toolroll decide <id> --choose <option> # answer a parked decision
|
|
343
|
+
toolroll task show retries # attempts, outcome, where the branch is
|
|
344
|
+
|
|
345
|
+
toolroll task add "Why does the login test flake?" --id flaky --report # a scout
|
|
346
|
+
toolroll chat --say "what is waiting on me across every project?" # the mate
|
|
347
|
+
```
|
|
348
|
+
|
|
349
|
+
Every command takes `--json` and answers with one envelope; every mutation
|
|
350
|
+
takes `--key` so a retry never files twice. `toolroll --help` and
|
|
351
|
+
`toolroll skills get console` are the live references — the second
|
|
352
|
+
is what your coding agent reads when you ask it how something works.
|
|
353
|
+
|
|
354
|
+
An agent that hits a judgement call **parks a typed decision instead of
|
|
355
|
+
guessing** — answer it from the terminal, the console, or a Telegram tap,
|
|
356
|
+
and the freed build resumes in seconds. Built work leaves only as a pushed
|
|
357
|
+
branch and a pull request, under a publication grant whose exact terms you
|
|
358
|
+
approved.
|
|
359
|
+
|
|
360
|
+
## Unattended is not auto-accept
|
|
361
|
+
|
|
362
|
+
Every tool in this category has a mode where the agent stops asking —
|
|
363
|
+
usually named something like *auto-accept*, or worse. Here the boundaries
|
|
364
|
+
do not loosen when you leave the room:
|
|
365
|
+
|
|
366
|
+
| An agent here can never | Enforced by |
|
|
367
|
+
|---|---|
|
|
368
|
+
| touch a default branch | builds land on `toolroll/<task>` in a leased worktree; push + PR happen only under a publication grant naming the exact repo, branch prefix, and base |
|
|
369
|
+
| approve its own work | approval nonces are minted only on screens that restate the digest-bound terms, and require your approver token typed again — **no LLM sits in any approval path** — and the [agent fence](#what-an-agent-can-reach) keeps your remembered login, runner tokens and the database out of the agent's reach |
|
|
370
|
+
| act on an irreversible option | `reversible` is a schema field; irreversible choices never auto-apply, and answering one from a phone takes a second minted confirmation tap |
|
|
371
|
+
| see Toolroll's secrets | the [agent fence](#what-an-agent-can-reach) blocks your login, runner and bot tokens, stored provider keys and project tool secrets at the operating system; other providers' keys are stripped from each agent's environment; secrets live in 0600 files, never in the database, URLs, or logs |
|
|
372
|
+
| spend while idle | **an LLM never polls** — the daemon does every no-judgement chore at zero token cost and wakes an agent only on a real event |
|
|
373
|
+
| spend without being counted | every provider spawn is stamped *before* it spends, so cost is measured, never asserted |
|
|
374
|
+
| guess at a judgement call | it parks a typed decision — recap, options with reversibility, recommendation, evidence — and the other eleven tasks keep going |
|
|
375
|
+
|
|
376
|
+
The whole claim is executable: one test,
|
|
377
|
+
[`src/unattended.test.ts`](src/unattended.test.ts), queues twelve tasks,
|
|
378
|
+
walks away, and comes back to pull requests.
|
|
379
|
+
|
|
380
|
+
The name comes from a captain's night orders — the written standing instructions left for the officer of the watch: *proceed on this course without me, and wake me under exactly these conditions.* That is the product, and it is not about the hour: it is for **long-running work that outlasts your attention** — an afternoon of errands, a weekend, or yes, a night.
|
|
381
|
+
|
|
382
|
+
Toolroll is a control plane for coding agents, optimized for the stretch where **nobody is watching**. It owns the scheduler, the attention surface — the typed queue of things waiting on a human — and an append-only event log.
|
|
383
|
+
|
|
384
|
+
It owns a deliberately small local task store, adapts richer trackers when they are already there, and owns no worktree pool, no review gate, and no agents. Those are adapters over [`beads`](https://github.com/gastownhall/beads), [`treehouse`](https://github.com/kunchenguid/treehouse), [`no-mistakes`](https://github.com/kunchenguid/no-mistakes), `claude`, and `codex`.
|
|
385
|
+
|
|
386
|
+
## Two claims
|
|
387
|
+
|
|
388
|
+
**Sixty seconds to first value.** No init, no daemon start, no wizard, no OAuth app.
|
|
389
|
+
|
|
390
|
+
```sh
|
|
391
|
+
npx toolroll ~/code # or: git clone … && npm install && npm run dev -- ~/code
|
|
392
|
+
```
|
|
393
|
+
|
|
394
|
+
It walks the filesystem for `.git` and reads every repo through the `git` credentials already on your machine, then shows what is in flight:
|
|
395
|
+
|
|
396
|
+
```
|
|
397
|
+
10 branches in flight across 24 repositories
|
|
398
|
+
|
|
399
|
+
vamarketplacenew main
|
|
400
|
+
feat/wise-payouts upstream gone 4d ago
|
|
401
|
+
feature/public-api-v1 ahead 17 1mo ago
|
|
402
|
+
api-pricing-impl behind 84 2mo ago
|
|
403
|
+
|
|
404
|
+
oddcircle redesign/instrumentation-cash-flag
|
|
405
|
+
main ahead 3, behind 56 23d ago
|
|
406
|
+
```
|
|
407
|
+
|
|
408
|
+
No agent has run. Nothing has been configured, written, or installed. Every other tool in this space starts from an empty database it expects you to fill. `--json` emits the same thing as `{ scannedAt, roots, repos }`, because half the intended audience is an agent.
|
|
409
|
+
|
|
410
|
+
Reads are priced before they are made. Listing refs is O(refs) and finishes in milliseconds; `git status` is O(working tree) and was measured at over two minutes on a real repo, so it is off by default behind `--dirty`. Computing ahead/behind walks history — 22s cold on a 304MB repo — so it is bounded at 5s and degrades to a branch list that says what it withheld. Every call goes through `--no-optional-locks`, so a scan never takes the index lock from an editor you have open.
|
|
411
|
+
|
|
412
|
+
`toolroll pulls` answers the narrower question of what is waiting on a person, and `toolroll graph` says which work graph is already here:
|
|
413
|
+
|
|
414
|
+
```
|
|
415
|
+
Work graph — detected in your repos
|
|
416
|
+
|
|
417
|
+
▸ beads 2 repos · 47 ready · native deps · runtime ok (1.4.0)
|
|
418
|
+
GitHub Issues 112 open · native deps · 2.67.0 too old, needs 2.94.0
|
|
419
|
+
|
|
420
|
+
Suggested: beads — the only work graph in your repos, and its runtime answers.
|
|
421
|
+
Nothing is enrolled, and detection grants nothing.
|
|
422
|
+
```
|
|
423
|
+
|
|
424
|
+
Backends are chosen by looking rather than asking, but **detection is not authorization** — finding a populated tracker says it exists, not that anyone wants an agent scheduling or closing what is in it. Data and runtime are detected separately, so a tracker whose binary is missing is reported as real work this machine cannot dispatch, which is a visible gap at 9am instead of a dead loop at 3am. Two populated trackers means neither is chosen: task count is not authority, and the biggest one may be the abandoned one. Where a fact is not established — Backlog.md's dependency edges, for instance — it is marked unverified and **fails closed**, because a private dependency graph other tools cannot see is shadow data.
|
|
425
|
+
|
|
426
|
+
**Nothing is ever installed for you.** `bd init` stages files, edits agent integrations, and can create a commit, so Toolroll prints the command and its side effects and lets you run it.
|
|
427
|
+
|
|
428
|
+
## Queueing work, and taking it
|
|
429
|
+
|
|
430
|
+
The built-in store is the fallback backend, and the commands over it are written for an agent first — because the agent is what runs them ten thousand times while you are away.
|
|
431
|
+
|
|
432
|
+
```sh
|
|
433
|
+
toolroll task add "migrate the payouts schema" --id schema
|
|
434
|
+
toolroll task add "wire the payouts API" --id api
|
|
435
|
+
toolroll task block api --on schema # api waits for schema
|
|
436
|
+
|
|
437
|
+
toolroll ready --json # what could be dispatched now
|
|
438
|
+
toolroll claim schema --runner builder-1 --key dispatch-schema
|
|
439
|
+
toolroll heartbeat <lease> # still working
|
|
440
|
+
toolroll release <lease> # done holding it
|
|
441
|
+
```
|
|
442
|
+
|
|
443
|
+
Four properties make that loop safe to run unattended.
|
|
444
|
+
|
|
445
|
+
**Every outcome is data.** `--json` returns the same envelope from every command, failures included: `{ ok, command, reason, message }`. `reason` is a stable token — `held`, `fenced`, `unknown-task` — because prose gets reworded and anything branching on it breaks silently. The binary teaches its own surface: `contract --commands` dumps the declared command guide, and `skills get <name>` serves version-matched operating guides straight from the exact build an agent is driving — never a stale snapshot.
|
|
446
|
+
|
|
447
|
+
**Exit codes separate "no" from "broken".** `0` got it · `1` something broke · `2` bad usage · `3` ran fine, the answer is no. Losing a claim race and finding the ready set empty are correct answers, not errors, and a loop that stops on them is as wrong as one that ignores real breakage.
|
|
448
|
+
|
|
449
|
+
**Every mutation takes `--key`.** An agent whose command succeeded but whose output was lost *will* retry. With a key that retry returns the first answer instead of queueing a second task or taking a second lease. Mutations that changed nothing are never recorded, so a refusal never becomes a permanent no.
|
|
450
|
+
|
|
451
|
+
**`fenced` means stop.** A runner whose machine slept, whose lease expired, and whose task was reclaimed will be told exactly that at its next heartbeat — long before it finishes work nobody will accept. Dispatch is a compare-and-swap on `(task, lease_generation)`, enforced by the database rather than by anything the caller remembers to check.
|
|
452
|
+
|
|
453
|
+
## The unattended pass
|
|
454
|
+
|
|
455
|
+
`toolroll tick` is the loop above with nobody typing it, once per invocation:
|
|
456
|
+
|
|
457
|
+
```sh
|
|
458
|
+
toolroll tick --runner builder-1 --token <t> --repo ~/code/thing --max 1
|
|
459
|
+
```
|
|
460
|
+
|
|
461
|
+
One pass: take the ready set, skip what nobody approved, claim what is left — re-proving readiness inside the same transaction as the claim, because the world moves between a list and a take — build each task in a leased worktree on `toolroll/<task-id>`, and commit. **Tick itself never pushes** and cannot touch the default branch; pushing and opening the pull request happen only under a publication grant whose exact repository, branch prefix, and base you approved — and merging stays yours, on GitHub.
|
|
462
|
+
|
|
463
|
+
It is deliberately a pass and not a daemon: point cron at it and the fences make repetition safe — a second pass finds the first's work done and converges to `empty` (exit 3) instead of building anything twice. A broken build marks its task `failed` and the pass exits 1 even if other tasks succeeded, because exit 0 has to mean "nothing needs you". Refusals that are really a person's pending decision — a scope nobody approved, or one that changed after approval — leave the task queued and untouched.
|
|
464
|
+
|
|
465
|
+
## Advanced: a separate background builder
|
|
466
|
+
|
|
467
|
+
Normal local use does not require this section: `toolroll up` is the
|
|
468
|
+
product command. For a remote or split deployment, the builder loop can manage
|
|
469
|
+
itself as an OS service — launchd on macOS, systemd on
|
|
470
|
+
Linux, Task Scheduler on Windows, chosen automatically — so "set it
|
|
471
|
+
running" is one command, and reboots and crashes are the supervisor's
|
|
472
|
+
problem:
|
|
473
|
+
|
|
474
|
+
```sh
|
|
475
|
+
toolroll daemon install --runner builder-1 --token <runner-token> --repo ~/code/thing
|
|
476
|
+
toolroll daemon status # running, as which pid, logs where
|
|
477
|
+
toolroll daemon logs # the file to tail
|
|
478
|
+
toolroll daemon uninstall # take it back off
|
|
479
|
+
```
|
|
480
|
+
|
|
481
|
+
The service restarts after a crash and after an unexpected clean exit alike
|
|
482
|
+
(launchd `KeepAlive`, systemd `Restart=always`, a restart loop under Task
|
|
483
|
+
Scheduler); `daemon install` is idempotent on a healthy running service and
|
|
484
|
+
really reloads a changed definition; `daemon uninstall` unloads and disables
|
|
485
|
+
it. `--containment observed|preferred|required` chooses how provider, setup
|
|
486
|
+
and check processes are bounded: a delegated cgroup v2 on Linux or a Job
|
|
487
|
+
Object on Windows when native, the observational tree scan otherwise —
|
|
488
|
+
`required` refuses to spawn rather than silently downgrading (macOS has no
|
|
489
|
+
native equivalent; the status says so and names the Linux route). The
|
|
490
|
+
contract and boundaries are in
|
|
491
|
+
[docs/PROCESS_CONTAINMENT.md](docs/PROCESS_CONTAINMENT.md).
|
|
492
|
+
|
|
493
|
+
Under the hood it runs `toolroll watch`: a work-conserving loop that
|
|
494
|
+
composes the same passes cron would call — but wakes on events (a decision
|
|
495
|
+
answered from your phone dispatches the next build in seconds), recovers
|
|
496
|
+
its own predecessor's mid-flight work after a crash, and spends zero tokens
|
|
497
|
+
while idle. The runner token lives in a 0600 file beside the database; the
|
|
498
|
+
service unit never carries it. Cron remains first-class if you prefer it —
|
|
499
|
+
`reconcile && tick ; bridge telegram` on a schedule does the same jobs at
|
|
500
|
+
cron's cadence, and a stray cron tick alongside a watch is safe (ordinary
|
|
501
|
+
claims settle the race), it just is not needed.
|
|
502
|
+
|
|
503
|
+
An honesty note for Windows: every pull request now type-checks, builds, and
|
|
504
|
+
runs the native Task Scheduler/link tests plus the core dispatch contract on
|
|
505
|
+
Windows with Node 22 and 24. Approved setup and verification commands use
|
|
506
|
+
Windows' native command shell. The scheduled-task definition follows the Task
|
|
507
|
+
Scheduler XML schema and every `schtasks` interaction is covered by scripted
|
|
508
|
+
tests. A physical Windows install has not yet been certified; the exact
|
|
509
|
+
real-provider and post-reboot checklist is in the
|
|
510
|
+
[Never Stuck release certification](https://github.com/ap9000/standing-orders/blob/main/docs/CERTIFICATION.md).
|
|
511
|
+
|
|
512
|
+
### Unattended permissions
|
|
513
|
+
|
|
514
|
+
The console's **Settings → unattended permissions** control chooses the
|
|
515
|
+
starting policy for new tasks. **Auto** lets routine repository commands and
|
|
516
|
+
edits proceed while the provider may stop on a risky permission request.
|
|
517
|
+
**Full access** runs Claude with `--dangerously-skip-permissions`, Codex (and
|
|
518
|
+
its OpenRouter transport) in its own sandbox widened to write anywhere with
|
|
519
|
+
network on, and Gemini with `--approval-mode yolo`, so permission prompts
|
|
520
|
+
cannot pause work while you are away. A Full access agent can change files
|
|
521
|
+
anywhere on your computer outside the [agent fence](#what-an-agent-can-reach).
|
|
522
|
+
Use Full access only for repositories and setup commands you trust.
|
|
523
|
+
|
|
524
|
+
### What an agent can reach
|
|
525
|
+
|
|
526
|
+
Agents run as your own user, so Toolroll fences its own secrets off
|
|
527
|
+
from them at the operating system, in every permission mode. The fence covers
|
|
528
|
+
the state folder beside the database (your remembered login `up-login.txt`,
|
|
529
|
+
runner and coordinator tokens, chat bot tokens, the database and its backups,
|
|
530
|
+
other runs' evidence) — everything there except the build's own worktree —
|
|
531
|
+
and `~/.toolroll` (stored provider keys and project tool secrets).
|
|
532
|
+
|
|
533
|
+
| | Auto | Full access |
|
|
534
|
+
|---|---|---|
|
|
535
|
+
| **Codex / OpenRouter** (any OS) | Codex's workspace sandbox (workspace writes, no network) with the fence | Codex's sandbox widened to write anywhere with network, with the fence |
|
|
536
|
+
| **Claude** | macOS: the agent runs inside a sandbox that denies the fence. Linux/Windows: Claude's own file tools refuse the fence, but its shell is not fenced yet | same |
|
|
537
|
+
| **Gemini** | macOS: fenced like Claude. Linux/Windows: not fenced yet | same |
|
|
538
|
+
|
|
539
|
+
Each attempt records how it was fenced. Two limits to know: in **API-key
|
|
540
|
+
mode** an agent's own provider key is in its environment, so its shell can
|
|
541
|
+
read it (subscription mode, the default, puts no key there); and a project's
|
|
542
|
+
own tool secrets reach that project's tools. Reviews and the lead chat run
|
|
543
|
+
with no tools and are confined to their own files. See [SECURITY.md](SECURITY.md).
|
|
544
|
+
|
|
545
|
+
Every new-task and task-scope form has the same two-choice control. A task's
|
|
546
|
+
choice is durable through planning rewrites and is sealed into the approved
|
|
547
|
+
execution profile; changing the installation default never broadens an
|
|
548
|
+
existing scope or approval.
|
|
549
|
+
|
|
550
|
+
### Quality modes
|
|
551
|
+
|
|
552
|
+
The **quality mode** setting selects routine or stronger configured agents.
|
|
553
|
+
**Strict / release** requests the stronger tier; it does not start an isolated
|
|
554
|
+
reviewer or automatic repair. The choice is signed into each scope, survives
|
|
555
|
+
later global changes, and remains visible in approval and run details.
|
|
556
|
+
Repository checks still run through the approved verifier. Publication and
|
|
557
|
+
deployment require their own authority and checks.
|
|
558
|
+
|
|
559
|
+
### Explainable phase routing
|
|
560
|
+
|
|
561
|
+
Which agent plans, builds, and repairs a task is decided once, from
|
|
562
|
+
signed facts, and written down with its reasons. The route reads the task's
|
|
563
|
+
declared **risk** (routine, elevated, high), its quality mode, what the
|
|
564
|
+
acceptance rubric demands (screenshots, manual review), how far a live
|
|
565
|
+
operating mode may carry the result unattended, and the agents you configured for each phase. Strength is never
|
|
566
|
+
inferred from a model's name: the ordinary phase row is the routine tier, and
|
|
567
|
+
`config set <phase> --tier strong --provider <p> --model <m>` names the agent
|
|
568
|
+
high-risk, strict, screenshot-proof, and automerge routes reach for. With no
|
|
569
|
+
strong row, a demanding task keeps the default and says so.
|
|
570
|
+
|
|
571
|
+
```
|
|
572
|
+
toolroll task scope <id> --goal … --acceptance … --risk high
|
|
573
|
+
toolroll task route <id> # every leg, its reason, its readiness
|
|
574
|
+
toolroll task route <id> --phase build --provider codex --model gpt-5-codex --as you --token <t>
|
|
575
|
+
toolroll task route <id> --clear-phase build --as you --token <t>
|
|
576
|
+
toolroll providers --report --runner <name> --token <t> # this machine's readiness
|
|
577
|
+
```
|
|
578
|
+
|
|
579
|
+
Every leg is **exact**: approvals bind a provider *and* a model id for the
|
|
580
|
+
planner, builder, and repair alike, so each phase names its model
|
|
581
|
+
once (`config set plan --provider claude --model <m>`, and the same for
|
|
582
|
+
`build`; repair inherits the build's model unless a same-provider
|
|
583
|
+
repair row names another). A phase without an exact model, a repair row on
|
|
584
|
+
another provider, or an unknown provider files the scope **unresolved** with
|
|
585
|
+
the words to fix it — nothing is guessed or substituted. Every override names
|
|
586
|
+
an exact model too, and a plan override becomes the planner pin.
|
|
587
|
+
|
|
588
|
+
Every override is recorded under the approver's name, in one transaction that
|
|
589
|
+
checks the scope you were reading is still the one on file (`--digest`, or the
|
|
590
|
+
form's own field). Approval seals the route — routine-shaped routes included —
|
|
591
|
+
together with any configured fallback chain, exactly as it seals the execution
|
|
592
|
+
profile; a later risk change or override re-files the scope and the old
|
|
593
|
+
approval reads stale, while a global configuration change can never rewrite a
|
|
594
|
+
sealed route. A row filed before routing existed is recognised by a durable
|
|
595
|
+
marker and stays governed by its sealed profile; a routed row whose route data
|
|
596
|
+
is missing or unreadable is refused everywhere until re-filed and approved
|
|
597
|
+
again. Repairs always stay on the build provider. Runners report provider
|
|
598
|
+
readiness without spending — at startup and with `providers --report`, never
|
|
599
|
+
on a timer — and the task page, the chat, and `task show` say **ready**,
|
|
600
|
+
**unavailable** (with the runner's own words), or **unknown** beside the
|
|
601
|
+
agents, outside what the approval signs; an unavailable provider halts before
|
|
602
|
+
any claim and is never substituted, except that an unavailable primary under
|
|
603
|
+
an approved fallback chain moves to the exact approved next entry (and only
|
|
604
|
+
under a live mode that allows paid fallback). Fallback admission re-check readiness and the exact leg inside their transactions.
|
|
605
|
+
Every run is stamped with its route and the actual provider and model at
|
|
606
|
+
admission — set once; a run that would spend as anything else refuses.
|
|
607
|
+
|
|
608
|
+
Admission proves that stamp before any run row exists: its shape (a known
|
|
609
|
+
phase, provenance word, and provider; a digest; an exact, argv-safe model id),
|
|
610
|
+
its phase against the run's role, its provider and model against what the run
|
|
611
|
+
would spend as, and its provenance against the authority the task actually
|
|
612
|
+
holds — the sealed route's digest and that phase's exact leg, the one approved
|
|
613
|
+
fallback-chain entry the run is bound to for a `fallback` run (two entries
|
|
614
|
+
that share a provider and model but differ in auth mode or repair model are
|
|
615
|
+
different authorities, and an ambiguous pair proves nothing), or a proven
|
|
616
|
+
pre-routing row for `legacy`, whose stamp must name the very sealed profile
|
|
617
|
+
(and, for a build or repair, its exact pair) — the bare word `legacy` belongs
|
|
618
|
+
only to a task with no scope. A task filed under routing never opens an
|
|
619
|
+
unstamped run, and the store dictates nothing: every planner, builder, scout,
|
|
620
|
+
reviewer, repair turn, and fallback **presents** the exact authority it holds
|
|
621
|
+
(`routeAuthorityFor` puts it in the caller's hands, in words), and a missing,
|
|
622
|
+
forged, stale, or inexact stamp opens no row — the refusal names what would
|
|
623
|
+
have had to be presented. Chain custody is proved and written in the same
|
|
624
|
+
insert: base custody opens the task's fallback cycle with the row, a
|
|
625
|
+
parked-resume takes the parked tail's custody through the proven transfer,
|
|
626
|
+
a repair turn inherits exactly its same-task parent's binding, and a fallback
|
|
627
|
+
entry is admitted only when everything it is told — the task, the live cycle,
|
|
628
|
+
the approved chain, the index, the entry digest, provider, model, auth mode,
|
|
629
|
+
repair model, and `fallback` provenance — re-proves against durable state; a
|
|
630
|
+
binding that cannot be proved rolls the insert back, and a mismatch creates no
|
|
631
|
+
run and consumes no edge. A reviewer after a fallback spawns under the review
|
|
632
|
+
leg and takes no custody. A run admitted under a route the scope has since
|
|
633
|
+
re-sealed away refuses to spend. Malformed authority — corrupt route, profile,
|
|
634
|
+
chain, or fallback JSON, a model id that is not one, a turn bound or clock
|
|
635
|
+
that is not a positive whole number, a run row whose chain index or auth mode
|
|
636
|
+
does not read — fails closed in words and never shrinks: a fallback row that
|
|
637
|
+
cannot be read files the scope unresolved rather than sealing a single profile
|
|
638
|
+
nobody configured, and an unreadable binding is no binding, never the base
|
|
639
|
+
entry or the subscription credential.
|
|
640
|
+
|
|
641
|
+
Codex resumes carry their sandbox as a configuration override (`-c
|
|
642
|
+
sandbox_mode=…`): `codex exec resume` has no `--sandbox` flag, so a structured
|
|
643
|
+
correction or repair-by-resume handed one exited before it initialized. A
|
|
644
|
+
resume the harness refuses by its own protocol, or never comes up for, ends
|
|
645
|
+
the planning attempt with its typed reason — the recorded session is not
|
|
646
|
+
resumed twice — and the next attempt, after the planning backoff, is a fresh
|
|
647
|
+
planner root in a fresh session.
|
|
648
|
+
|
|
649
|
+
A **routine** freezes its agents too. Filing a standing order resolves its
|
|
650
|
+
four-role route from the configuration of that moment and binds it into the
|
|
651
|
+
digest you sign; approval seals the snapshot, and every firing re-hashes that
|
|
652
|
+
snapshot against the approval, holds the build and repair legs to the sealed
|
|
653
|
+
profile's exact provider and model, copies it onto the instance verbatim, and
|
|
654
|
+
rolls the whole firing back unless the instance seals under it — a later
|
|
655
|
+
`config set` cannot re-route a firing. One integrity projection answers
|
|
656
|
+
every question about a standing order's agents — what the page says, whether
|
|
657
|
+
a password may be minted, whether a yes may land, whether a firing may proceed
|
|
658
|
+
— read before any write: an approval is *live* only when its frozen snapshot
|
|
659
|
+
reads back, hashes with the stored terms to the digest the approver signed,
|
|
660
|
+
states no leg problem, and agrees with the sealed profile's build and repair
|
|
661
|
+
pairs. A routine approved before agents were frozen (an upgrade from before
|
|
662
|
+
v48), one whose snapshot cannot be read, or one whose snapshot no longer
|
|
663
|
+
verifies fires nothing and pages once; its page and `routine show` say so and
|
|
664
|
+
offer the one road: `toolroll routine refresh <name>` (or the page's
|
|
665
|
+
**Refresh agents** button) re-resolves the agents from today's configuration,
|
|
666
|
+
withdraws an approval that is not live even when the working agents are
|
|
667
|
+
unchanged, and approves nothing — you read the exact agents it now names and
|
|
668
|
+
approve it again with your password, and only that yes fires the new
|
|
669
|
+
snapshot. An authentic v47 database, or one whose v47→v48 upgrade was
|
|
670
|
+
interrupted, upgrades without re-running older data passes, changing ids,
|
|
671
|
+
backfilling a route, or approving anything.
|
|
672
|
+
|
|
673
|
+
Every approval surface — the task page, the focused chat, and `/next` — shows
|
|
674
|
+
the same concise line of exact agents above the password, with the declared
|
|
675
|
+
risk explained in plain words and the runtime limits one tap away; changing
|
|
676
|
+
any agent invalidates the approval every surface signed under. Where no yes
|
|
677
|
+
could bind — an unreadable route, a route that cannot run, a standing order
|
|
678
|
+
without frozen agents, or a pre-routing row whose old approval no longer
|
|
679
|
+
stands — those surfaces mint no nonce and show no password or approve button,
|
|
680
|
+
only the reason and the act that opens it (re-file the scope, change the
|
|
681
|
+
agents, refresh the routine) — the inbox row reads *needs attention* rather
|
|
682
|
+
than *review & approve*, and the routine's recovery is one labelled, described
|
|
683
|
+
button with nothing to type; an old approval already on a pre-routing row is
|
|
684
|
+
grandfathered, but no new yes lands on it. The demo sandbox shows no seeded
|
|
685
|
+
conversation: chat evidence is a real subscription-backed plane. The task page's controls offer,
|
|
686
|
+
per role, only the agents you configured *for that role* (gemini never
|
|
687
|
+
reviews; repairs stay on the build provider), a current agent the
|
|
688
|
+
configuration no longer names is shown for what runs today and never offered
|
|
689
|
+
again, and each risk level says what it does — truthfully under strict
|
|
690
|
+
quality and screenshot proof too. Chat reads the same route (`get_agents`)
|
|
691
|
+
and proposes one confirmation-gated change (`propose_agents`) — a risk, one
|
|
692
|
+
role switched to a listed agent, or a hand-picked role cleared — which lands,
|
|
693
|
+
when you confirm the card, through the same authenticated route edit the page
|
|
694
|
+
uses; the chosen agent is re-proved against the role's configured choices
|
|
695
|
+
*inside* that transaction, so a card drafted against yesterday's
|
|
696
|
+
configuration changes nothing.
|
|
697
|
+
|
|
698
|
+
## The phone, both directions
|
|
699
|
+
|
|
700
|
+
The Telegram bridge closes the loop without a terminal: a parked decision
|
|
701
|
+
arrives as a message with one button per option, and a tap answers it
|
|
702
|
+
through the same authenticated path as the CLI and the web view — the hold
|
|
703
|
+
lifts, and the next pass resumes the task with the answer in the agent's
|
|
704
|
+
brief. No LLM is anywhere in this path.
|
|
705
|
+
|
|
706
|
+
Setup, once:
|
|
707
|
+
|
|
708
|
+
1. In Telegram, message **@BotFather**: `/newbot`, pick a name and a
|
|
709
|
+
username. Copy the token it hands you.
|
|
710
|
+
2. `toolroll bridge telegram token <that-token>` — stored in a 0600 file
|
|
711
|
+
beside the database (or set `TOOLROLL_TELEGRAM_TOKEN`, which wins;
|
|
712
|
+
or paste it into `serve`'s settings card from your phone).
|
|
713
|
+
3. `toolroll approver add you --password <yours>` if you have no
|
|
714
|
+
sign-in yet — that name and password are the login for the console and
|
|
715
|
+
every approving act. (Omit `--password` and a high-entropy one is
|
|
716
|
+
minted and printed once instead — better for API/bearer use.)
|
|
717
|
+
4. `toolroll bridge telegram pair --as you --token <approver-token>` —
|
|
718
|
+
prints a one-time code, good for ten minutes.
|
|
719
|
+
5. From your phone, open your bot's chat, press Start, send
|
|
720
|
+
`/pair <that-code>`, then run `toolroll bridge telegram` once to
|
|
721
|
+
complete it. The bot replies with who the chat now answers as.
|
|
722
|
+
|
|
723
|
+
Then cron the pass next to `tick`:
|
|
724
|
+
|
|
725
|
+
```sh
|
|
726
|
+
toolroll bridge telegram # sends pending, applies taps, exits
|
|
727
|
+
```
|
|
728
|
+
|
|
729
|
+
Once paired, an ordinary message in that chat talks to the same assistant
|
|
730
|
+
as the console's chat page and `toolroll chat`: the same saved
|
|
731
|
+
thread, the same proposal cards (Confirm or Dismiss under each one), the
|
|
732
|
+
same confirm doors, with the phone recorded as the source. `/status`,
|
|
733
|
+
`/task <id>` and `/help` stay cheap, model-free reads. Reply to a result
|
|
734
|
+
message to ask for changes to that exact result; approvals that take a
|
|
735
|
+
password, cancelling and publishing still happen on the computer. Chat from
|
|
736
|
+
the phone uses your configured membership provider only; a direct-API
|
|
737
|
+
configuration is never spent from Telegram.
|
|
738
|
+
|
|
739
|
+
`bridge telegram status` shows the token source, the binding, and what is
|
|
740
|
+
waiting. For answers in seconds instead of at the next cron firing,
|
|
741
|
+
`toolroll bridge telegram --follow` stays on the wire — one long-poll
|
|
742
|
+
actor holding the same poll lease, so a cron pass overlapping it simply
|
|
743
|
+
loses the race. `toolroll watch` embeds the same follower automatically
|
|
744
|
+
when a bot token is configured: a tap on your phone answers the decision,
|
|
745
|
+
the answer wakes the loop, and the freed task resumes — phone to build,
|
|
746
|
+
no timer in between.
|
|
747
|
+
|
|
748
|
+
**Away mode.** `toolroll bridge telegram digest --every 2h` (or
|
|
749
|
+
`--off`, or the console's settings card) holds routine facts — merges,
|
|
750
|
+
reports, retries, plans ready — and sends them as one digest on that
|
|
751
|
+
cadence. A decision, and anything that needs a person now (a stalled
|
|
752
|
+
task, a malformed payload, a gap that blocks work), still pages the
|
|
753
|
+
moment it lands. `bridge telegram status` says how many facts are held.
|
|
754
|
+
|
|
755
|
+
**Check work from your phone.** In the paired private chat, send `/status`
|
|
756
|
+
for recent work across enrolled projects, `/task <id>` for one task's exact
|
|
757
|
+
state, recorded evidence, delivery status and next step, or `/help` for the
|
|
758
|
+
available commands. These read existing workflow records without a model call
|
|
759
|
+
or task mutation. A saved branch, a pending review, and a merged PR stay
|
|
760
|
+
distinct. `/status` covers the newest 60 tasks, explicitly says when older
|
|
761
|
+
work is omitted, and shows up to two rows per group; `/task` can look up older
|
|
762
|
+
tasks directly. The computer and bridge must be awake and connected. These
|
|
763
|
+
commands are not yet free-form task creation or revision chat; use the console
|
|
764
|
+
for those. See [phone status and its limits](docs/PHONE_STATUS.md).
|
|
765
|
+
|
|
766
|
+
A chat is not a person: pairing binds one private chat and one immutable
|
|
767
|
+
Telegram user id to one approver credential. Buttons carry opaque one-time
|
|
768
|
+
tokens whose meaning lives in the local database — a stolen bot token can
|
|
769
|
+
read what was sent and repaint keyboards, but it cannot mint a token,
|
|
770
|
+
answer as you, or arm an irreversible choice, which takes a second minted
|
|
771
|
+
confirmation tap. Rotating your approver credential strands the chat and
|
|
772
|
+
every outstanding button, and the bot token itself is stripped from every
|
|
773
|
+
agent's environment.
|
|
774
|
+
|
|
775
|
+
## Peeking at the agents
|
|
776
|
+
|
|
777
|
+
```sh
|
|
778
|
+
toolroll peek # one pane per live run: stage, clock, what the agent is saying
|
|
779
|
+
toolroll peek 42 # follow one run until it finishes
|
|
780
|
+
toolroll peek --tmux # a real tmux session, one window per run
|
|
781
|
+
```
|
|
782
|
+
|
|
783
|
+
The panes tail each run's live transcript, the same file the console's
|
|
784
|
+
run page follows: the text the agent said and the kind of tool it reached
|
|
785
|
+
for (editing files, running a command, searching the code), never file
|
|
786
|
+
contents or command lines, with credential-shaped lines redacted at
|
|
787
|
+
write. Digits focus one pane, `a` shows them all, `q` leaves. Outside a
|
|
788
|
+
terminal, or with `--json`, it prints one snapshot and exits.
|
|
789
|
+
|
|
790
|
+
## The mate, and the gateway
|
|
791
|
+
|
|
792
|
+
`/chat` in the console (or `toolroll chat` in the terminal) is one
|
|
793
|
+
conversation across every project you serve. The mate reads the fleet
|
|
794
|
+
and **only proposes**: file a task (or a scout), move one to the front,
|
|
795
|
+
reserve it for a worker, hold it, rewrite a scope, retry/replace/unlink a
|
|
796
|
+
terminal dependency, guide a task's next attempt, cancel, or suggest an answer to a parked decision. Every
|
|
797
|
+
proposal is a card you confirm, with
|
|
798
|
+
every consequence and the builder's recommendation shown beside the
|
|
799
|
+
mate's pick; a scope the mate wrote never seals under an operating mode.
|
|
800
|
+
Absolute paths, internal digests and account names are redacted; source context includes relative file citations. Direct API use spends
|
|
801
|
+
against a ceiling you set per conversation. The console keeps a live pulse for every
|
|
802
|
+
admitted project beside that shared thread, with one-click fleet questions
|
|
803
|
+
and a direct road to each project's board.
|
|
804
|
+
|
|
805
|
+
Every task has an **Overview / Ask** switch. **Ask** opens a focused companion
|
|
806
|
+
to that task without creating another conversation: Toolroll attaches
|
|
807
|
+
the current task to each new message, keeps the live status beside the thread,
|
|
808
|
+
and offers plain-language starters for status, scope revision, steering, and
|
|
809
|
+
result inspection. Proposed guidance is inert until you confirm its card, then it
|
|
810
|
+
reaches the next attempt without interrupting work already running.
|
|
811
|
+
|
|
812
|
+
New work starts the same way: describe the outcome once in ordinary language.
|
|
813
|
+
The mate infers a concise title, narrow scope, safe non-goals, and proof
|
|
814
|
+
criteria; it asks only when the project or outcome is genuinely ambiguous, or
|
|
815
|
+
when an irreversible or compatibility tradeoff changes what should be built.
|
|
816
|
+
Questions are grouped, carry a recommended default, and stop for reversible
|
|
817
|
+
choices when you say **use your judgment**. The task card says whether Standing
|
|
818
|
+
Orders will inspect the repository and draft a plan before asking you to
|
|
819
|
+
approve anything.
|
|
820
|
+
|
|
821
|
+
When a task finishes, both views lead with the same Ready status and result.
|
|
822
|
+
Open the work and its actual checks, then mark Complete or request changes.
|
|
823
|
+
Historical diagnostics remain available in details; optional missing assessments
|
|
824
|
+
do not create another review stage. Chat cannot rewrite the stored result.
|
|
825
|
+
|
|
826
|
+
The evidence page opens the sealed patch in a clean **View** mode. Switch to
|
|
827
|
+
**Annotate** only when you need a change: select the exact line, leave plain-
|
|
828
|
+
language feedback, and collect as many notes as needed. Creating a revision is
|
|
829
|
+
a separate, optional act that seals the exact annotation batch into one scoped
|
|
830
|
+
task for approval; ordinary result review never requires it.
|
|
831
|
+
|
|
832
|
+
The chat setup screen defaults to **Codex membership · default model**.
|
|
833
|
+
Run `codex login` once on the machine serving Toolroll, choose that
|
|
834
|
+
provider, and there is no Toolroll dollar maximum. The conversation
|
|
835
|
+
stays live until you end it; the daily turn limit and your plan's own upstream
|
|
836
|
+
limits still apply. Anthropic
|
|
837
|
+
membership works the same way after signing in with the `claude` CLI. Direct
|
|
838
|
+
`anthropic-api` and `openrouter-api` modes remain available; only those modes
|
|
839
|
+
ask for weekly and per-conversation dollar ceilings.
|
|
840
|
+
|
|
841
|
+
```sh
|
|
842
|
+
codex login
|
|
843
|
+
toolroll config set chat --provider codex-subscription --as you --token <password>
|
|
844
|
+
toolroll serve --repo /path/to/project-a --repo /path/to/project-b
|
|
845
|
+
# Open /chat, type your password once to start the conversation, then talk.
|
|
846
|
+
```
|
|
847
|
+
|
|
848
|
+
Coding agents you run elsewhere reach the same plane through the MCP
|
|
849
|
+
gateway: `toolroll mcp` serves a coordinator credential you mint,
|
|
850
|
+
bound to named repositories, that can read the fleet, file quarantined
|
|
851
|
+
proposals, and propose the same guarded acts — `toolroll proposals`
|
|
852
|
+
and the task page are where you confirm them. Both roads keep the one
|
|
853
|
+
rule: the plane never acts on a model's word.
|
|
854
|
+
|
|
855
|
+
## The console
|
|
856
|
+
|
|
857
|
+
`toolroll serve --repo <path>` is no longer just the decision view — it
|
|
858
|
+
is the whole built-in queue, operable from a phone: an inbox of everything
|
|
859
|
+
waiting on you, a live activity report (run counts, measured spend,
|
|
860
|
+
decisions, incidents, stranded work, gaps), every task with its scope, holds, runs, decisions and
|
|
861
|
+
incidents on one screen, run pages with the economics and the agent's
|
|
862
|
+
concluding words, and read-only capabilities. Adding a task, holding,
|
|
863
|
+
requeuing, cancelling, and editing a scope all happen from the page — each
|
|
864
|
+
re-proved server-side, so a stale tab never erases what the world did in
|
|
865
|
+
the meantime, and a task a runner is building right now refuses to be
|
|
866
|
+
cancelled out from under it.
|
|
867
|
+
|
|
868
|
+
Approving a scope is deliberately heavier than a click: the form restates
|
|
869
|
+
the goal, the exclusions, and the touched paths — exactly the fields the
|
|
870
|
+
approval digest binds — and requires your approver token typed again. A
|
|
871
|
+
logged-in session alone can read everything and approve nothing.
|
|
872
|
+
|
|
873
|
+
Plain HTTP, so keep it on localhost or a tailnet and put TLS in front for
|
|
874
|
+
anything else.
|
|
875
|
+
|
|
876
|
+
### Webhooks through a reverse proxy
|
|
877
|
+
|
|
878
|
+
Flow triggers can take webhooks from GitHub, Linear or any service at a
|
|
879
|
+
secret address under `/hooks/`. Keep the console itself private (on your
|
|
880
|
+
tailnet or localhost) and expose only that path. With Caddy:
|
|
881
|
+
|
|
882
|
+
```caddy
|
|
883
|
+
hooks.example.com {
|
|
884
|
+
handle /hooks/* {
|
|
885
|
+
reverse_proxy 127.0.0.1:4180 {
|
|
886
|
+
header_up Host {upstream_hostport}
|
|
887
|
+
}
|
|
888
|
+
}
|
|
889
|
+
respond 404
|
|
890
|
+
}
|
|
891
|
+
```
|
|
892
|
+
|
|
893
|
+
Then save `https://hooks.example.com` as the public webhook address on a
|
|
894
|
+
flow's Triggers panel. GitHub deliveries are proved with the secret the
|
|
895
|
+
panel shows once; Linear deliveries with Linear's own signing secret, pasted
|
|
896
|
+
on the panel behind your password. Addresses are never stored readable: a
|
|
897
|
+
lost one is replaced with **New address**, which retires the old.
|
|
898
|
+
|
|
899
|
+
## Steering a fleet, not just a task
|
|
900
|
+
|
|
901
|
+
Everything below ships in 0.4.0:
|
|
902
|
+
|
|
903
|
+
- **The queue screen** — every worker's up-next list as columns, like a
|
|
904
|
+
music queue: drag to reorder, drag into a worker's column to reserve a
|
|
905
|
+
task for it (each column wears an editable theme note), top is taken
|
|
906
|
+
first. A worker drains its own column, then the shared queue. The
|
|
907
|
+
reservation is enforced in the claim primitive itself — the wrong
|
|
908
|
+
worker's claim gets a typed `reserved` refusal, however it asks.
|
|
909
|
+
- **Chains and "this one first"** — `task block/unblock` wires
|
|
910
|
+
dependencies (cycles refused), "starts after" on the filing form,
|
|
911
|
+
`task next` moves work to the front of its own queue. Scheduling,
|
|
912
|
+
never authority: approvals are untouched by any of it.
|
|
913
|
+
- **Tournaments** — race 2–4 agents on one task under native dollar
|
|
914
|
+
caps, compare their verified diffs side by side, and pick one through
|
|
915
|
+
a password ceremony; the losers' branches and evidence are kept.
|
|
916
|
+
- **The live peek** — a running build's page shows what is changing in
|
|
917
|
+
its checkout right now (names and counts, never contents), through a
|
|
918
|
+
native reader that executes nothing — no git command ever runs
|
|
919
|
+
against an agent-controlled worktree.
|
|
920
|
+
- **External dispatch** — enroll a GitHub repository with an explicit
|
|
921
|
+
dispatch grant and its labeled issues become ordinary local tasks:
|
|
922
|
+
scoped and approved HERE (issue bodies are never imported — a tracker
|
|
923
|
+
anyone can write to is a prompt-injection surface), built unattended,
|
|
924
|
+
answered back with a PR-link comment under exactly the write classes
|
|
925
|
+
you granted. An issue closed mid-build can never publish: the
|
|
926
|
+
completion transaction disowns it, keeps the branch as evidence, and
|
|
927
|
+
says so. Done stays done — remote closure never regresses a completed
|
|
928
|
+
dependency. Revisions ride review comments: mark up the finished
|
|
929
|
+
diff (or let granted reviewers do it from the PR) and seal the batch
|
|
930
|
+
into one new approval-bound task.
|
|
931
|
+
- **Operating modes** — a per-repository, password-signed, expiring
|
|
932
|
+
envelope that pre-authorizes the SIGNER'S OWN future acts: your
|
|
933
|
+
filings approve themselves the moment you file them, watched sessions
|
|
934
|
+
start without retyping your password, merges fire themselves when CI
|
|
935
|
+
is seen green on the exact authorized commit (only through a merge
|
|
936
|
+
grant, never around one), and daily run/dollar rails bound the spend.
|
|
937
|
+
Every term renders in words at the signing ceremony — including the
|
|
938
|
+
sentence "your signed-in browser session becomes a spend credential
|
|
939
|
+
for this repository" — and ending it all is one click, for any
|
|
940
|
+
approver, at any moment. No mode signed means nothing changes: every
|
|
941
|
+
act keeps its own ceremony, the default forever.
|
|
942
|
+
- **The reviewer** — an agent pass over a finished build's sealed diff,
|
|
943
|
+
and nothing else: no worktree, no repository access — the patch is
|
|
944
|
+
re-verified against its recorded hash, comments are proven
|
|
945
|
+
patch-local, and they land beside your own for YOU to prune and seal
|
|
946
|
+
into a revision task. A revision keeps the source's contract — its
|
|
947
|
+
goal, exclusions, touches, exact rubric, declared risk, quality,
|
|
948
|
+
permission posture, and budget ceiling — however the installation's
|
|
949
|
+
defaults have changed since, re-resolves its agents for a fresh
|
|
950
|
+
approval, and inherits no approval, session, publication, or merge
|
|
951
|
+
grant (the policy is [docs/REVISION_TERMS.md](docs/REVISION_TERMS.md)).
|
|
952
|
+
One successful review per build, ever; a mode
|
|
953
|
+
can run one on every finished build automatically. A review that
|
|
954
|
+
fails or is interrupted may be retried EXPLICITLY — `task review
|
|
955
|
+
<run>` again, or the task/result page's **Retry review** — at most
|
|
956
|
+
twice (three root attempts in all). Each retry is a fresh request and
|
|
957
|
+
a fresh reviewer admitted under the build's current sealed route with
|
|
958
|
+
every sealed input re-verified; the failed attempts stay on record;
|
|
959
|
+
a review that succeeded, or one still queued or running, is never
|
|
960
|
+
retried; nothing retries by itself, and the source build is never
|
|
961
|
+
rerun. A retry is an operator's act only: the mode's and the Strict /
|
|
962
|
+
release scope's automatic review asks are one shot per build, and a
|
|
963
|
+
replayed one is refused at request and again at admission, before any
|
|
964
|
+
money.
|
|
965
|
+
- **People** — invite someone with a single-use link that pins their
|
|
966
|
+
powers at mint (watch everything, or approve and act), see who is
|
|
967
|
+
doing what, and remove access with one ceremony that actually severs:
|
|
968
|
+
sessions, invites, and every mode they signed end together, while
|
|
969
|
+
history stays attributed forever.
|
|
970
|
+
- **Watched sessions, plural** — run several attended sessions per
|
|
971
|
+
worker under one signed ceiling, each with its own model and
|
|
972
|
+
permission posture chosen at mint, the whole execution profile under
|
|
973
|
+
the signature.
|
|
974
|
+
- **Scout tasks** — `task add … --report` (or the "scout" checkbox, or
|
|
975
|
+
the mate's `propose_task` with `report: true`) files a task whose
|
|
976
|
+
deliverable is a report, never a branch: once its scope is approved,
|
|
977
|
+
a read-only session investigates the goal as a question and hands
|
|
978
|
+
back a title, a summary, a document, and up to five follow-ups, each
|
|
979
|
+
of which files as a task with one tap. The workspace is proven
|
|
980
|
+
untouched before a byte of the report is read; a scout that changed
|
|
981
|
+
anything gets nothing ingested.
|
|
982
|
+
|
|
983
|
+
## Writing to a tracker you already have
|
|
984
|
+
|
|
985
|
+
Detection tells you what is there; a grant is what lets anything be written to it.
|
|
986
|
+
|
|
987
|
+
```sh
|
|
988
|
+
toolroll enroll . --backend github-issues --paths owner/name # shows the terms
|
|
989
|
+
toolroll enroll . --backend github-issues --paths owner/name --yes
|
|
990
|
+
toolroll grants # what has been granted, and to what
|
|
991
|
+
toolroll revoke . # take it back
|
|
992
|
+
|
|
993
|
+
toolroll ready --backend github-issues # reads need no grant
|
|
994
|
+
toolroll task add "..." --backend beads # writes do
|
|
995
|
+
```
|
|
996
|
+
|
|
997
|
+
The grant is not a boolean. It records which paths or repositories may be touched, which mutation classes are allowed, which tasks are covered, which credential scope applies, and whether the writes will turn up in `git status` — that last one asked of `git check-ignore` rather than assumed. Two defaults carry weight: only tasks Toolroll created or was given, because enrolling a repo with four hundred open issues is not volunteering all four hundred; and `close` is withheld, because closing what somebody else filed is not the same act as transitioning your own task.
|
|
998
|
+
|
|
999
|
+
Every backend goes through the same contract, and the authorization wraps it rather than living inside each adapter — an adapter written later inherits the check instead of having to remember it.
|
|
1000
|
+
|
|
1001
|
+
**Edges are never emulated.** beads has native dependencies and they are used. This GitHub adapter has not confirmed the dependency endpoint against a live repository, so `addEdge` refuses rather than storing a graph only Toolroll can see — one that would read as ready to every human on the repo. That is the design's rule, and the refusal says so.
|
|
1002
|
+
|
|
1003
|
+
The beads adapter is built to beads' own documentation and exercised against a stubbed runner; it has never run against a real installation, because `bd` was not present on the machine it was written on. Commands whose flags could not be established — a general status update, in particular — refuse rather than guess.
|
|
1004
|
+
|
|
1005
|
+
The materialised snapshot M0 promised shipped as **external dispatch**: a `sync` pass mirrors a tracker's nominated issues into the queue as ordinary local tasks (titles only, validated; bodies never), so the scheduler's hot path never touches the network — see below.
|
|
1006
|
+
|
|
1007
|
+
**It survives the night, cheaply.** Work dispatches itself from a dependency graph, fails safely, and parks a *typed* decision — recap, options with reversibility, a recommendation, evidence — instead of guessing. Parking never stalls the loop; the blocked task steps aside and eleven others keep going.
|
|
1008
|
+
|
|
1009
|
+
And it costs nothing while idle. **An LLM never polls.** The daemon handles everything that needs no judgement — ticks, capability probes, lease reaping, CI polling, notifications — at zero token cost, and wakes an agent only on a real event. The target is a testable invariant: an eight-hour run with twelve tasks shows near-zero token spend across idle windows.
|
|
1010
|
+
|
|
1011
|
+
## What breaks overnight, and the answer
|
|
1012
|
+
|
|
1013
|
+
| Failure | Mechanism |
|
|
1014
|
+
|---|---|
|
|
1015
|
+
| Expired key found at 3am after 40k wasted tokens | capabilities probed *before* dispatch; gaps ranked by tasks unblocked |
|
|
1016
|
+
| A runner dies holding a worktree | `Claim` with an immutable lease id and a fencing generation; late completions rejected |
|
|
1017
|
+
| A build fails, then fails the same way again | typed strikes with doubling backoff; three strikes stalls the task for a person, who exits it with `task requeue` |
|
|
1018
|
+
| You wake to five transcripts | one briefing: what ran, what is blocked, what needs deciding |
|
|
1019
|
+
|
|
1020
|
+
## Status
|
|
1021
|
+
|
|
1022
|
+
**0.4.0 — schema v29, suite 1,369.** The M4 loop plus tournaments, the
|
|
1023
|
+
live peek and live transcript, chains and queue ordering, per-worker
|
|
1024
|
+
queue columns, external dispatch, merge grants with observed-green CI
|
|
1025
|
+
and a fourteen-transition merge machine (per-merge human authorization
|
|
1026
|
+
by default; a signed automerge mode may substitute for exactly that
|
|
1027
|
+
yes, re-proved at the moment of firing), the phone PWA with
|
|
1028
|
+
zero-dependency push, the attended core (governed live sessions:
|
|
1029
|
+
signed terms, mid-session conversation, crash custody, continuation,
|
|
1030
|
+
N parallel sessions per worker), the attested runtime (four providers
|
|
1031
|
+
— Claude, Codex, OpenRouter, Gemini — the last admitted by versioned
|
|
1032
|
+
conformance, never a registry row), labeled cross-runtime comparisons
|
|
1033
|
+
with honest per-lane money, operating modes with their daily rails,
|
|
1034
|
+
the artifact-only reviewer, and multi-user instances with invite
|
|
1035
|
+
links, roles, and a People screen. Project access now supports viewer/operator
|
|
1036
|
+
invitations restricted to selected repositories, with an action ledger for
|
|
1037
|
+
people and unattended work. See [Project access and action ledger v1](docs/PROJECT_ACCESS_LEDGER.md)
|
|
1038
|
+
|
|
1039
|
+
[Automatic approvals](docs/AUTO_APPROVAL.md) explains the signed project policy for scope filings, unchanged plans, reviews, repairs, and merges.
|
|
1040
|
+
|
|
1041
|
+
for its permissions, history coverage, and migration behavior.
|
|
1042
|
+
Earlier arcs shipped behind their own adversarial review rounds;
|
|
1043
|
+
docs/PROGRESS.md records those findings.
|
|
1044
|
+
|
|
1045
|
+
**M4 built.** The whole loop runs: `toolroll watch` (or `daemon
|
|
1046
|
+
install` — no crontab) dispatches approved work, spends nothing while idle,
|
|
1047
|
+
survives crashes by recovering exactly its own predecessor's claims, and
|
|
1048
|
+
stops taking work on the first signal. Failures are typed — strikes,
|
|
1049
|
+
doubling backoff, three-strike stalls a person exits with `task requeue` —
|
|
1050
|
+
and every provider spawn is stamped before it spends, so cost is measured,
|
|
1051
|
+
never asserted. CI on published PRs is watched as episodes that never call
|
|
1052
|
+
silence green. The whole unattended stretch is one test,
|
|
1053
|
+
[`src/unattended.test.ts`](src/unattended.test.ts): queue twelve, walk
|
|
1054
|
+
away, come back to PRs. Architecture: [`docs/DESIGN.md`](docs/DESIGN.md);
|
|
1055
|
+
the item-by-item ledger: [`docs/PROGRESS.md`](docs/PROGRESS.md).
|
|
1056
|
+
|
|
1057
|
+
## Milestones
|
|
1058
|
+
|
|
1059
|
+
| | | |
|
|
1060
|
+
|---|---|---|
|
|
1061
|
+
| M0 | discovery, graph adapters, leases, CLI | `npx toolroll` shows what is in flight — **useful before it is autonomous** |
|
|
1062
|
+
| M1 | runners, worktrees, first builder | one task goes queued → branch → commit unattended |
|
|
1063
|
+
| M2 | capability probes, secrets, briefing | fill one gap, three tasks start |
|
|
1064
|
+
| M3 | decisions, evidence, web view | a park renders as one screen, answerable on a phone — **and it does, executably** |
|
|
1065
|
+
| M4 | the loop | **queue twelve, sleep, wake to PRs — with near-zero idle spend** |
|
|
1066
|
+
|
|
1067
|
+
Deferred until M4 earns them: the spatial board, multiplayer, in-browser terminals, Postgres, RBAC.
|
|
1068
|
+
|
|
1069
|
+
M4 is the product. M0 is what makes anyone install it long enough to reach M4.
|
|
1070
|
+
|
|
1071
|
+
## Not competing with
|
|
1072
|
+
|
|
1073
|
+
[**agor**](https://github.com/preset-io/agor) owns the execution-plane category and does it well — browser UI, six runtimes, multiplayer, a spatial board. It optimizes for a team steering agents *live*; we optimize for nobody being awake. It is BSL 1.1; this is MIT.
|
|
1074
|
+
|
|
1075
|
+
[**firstmate**](https://github.com/kunchenguid/firstmate) proves the orchestrator role works as conventions plus tmux, with no UI and no schema. Its event-driven bash watcher is where the zero-token supervision rule came from. Our bet is that the same role is better with a typed decision record and a browser you can answer from.
|
|
1076
|
+
|
|
1077
|
+
## Contributing
|
|
1078
|
+
|
|
1079
|
+
The most useful surface is a **provider adapter** — `src/provider.ts` is
|
|
1080
|
+
the only module that names an agent binary, and
|
|
1081
|
+
[CONTRIBUTING.md](CONTRIBUTING.md) walks the contract. Bug reports want
|
|
1082
|
+
`--json` output; the issue forms say what else. Every behavior lands with
|
|
1083
|
+
a test — the suite is the specification.
|
|
1084
|
+
|
|
1085
|
+
`npm run e2e:flows` checks flows end to end with nothing stubbed: a throwaway
|
|
1086
|
+
instance (the real CLI, console and worker) on a scratch repository, driven
|
|
1087
|
+
through a real browser, with real Claude turns for the lead and one real
|
|
1088
|
+
build. It needs `claude` and `gh` logged in and Playwright's Chromium, takes
|
|
1089
|
+
about seven minutes, and writes `report.md`, screenshots and both logs to
|
|
1090
|
+
`output/e2e/`.
|
|
1091
|
+
|
|
1092
|
+
## Credits
|
|
1093
|
+
|
|
1094
|
+
The workflow this formalizes comes from [Jason Ku's agentic engineering session](https://youtu.be/Ukju3maxbEQ) and his [`agents-md-snippets`](https://github.com/jasonku09/agents-md-snippets), plus Kun Chen's `treehouse`, `no-mistakes`, `gnhf`, `tasks-axi`, and `axi`. The design was reviewed adversarially by Codex; the appendix in `docs/DESIGN.md` lists every claim that review falsified, because the corrections are more useful than a clean spec would have been.
|
|
1095
|
+
|
|
1096
|
+
## License
|
|
1097
|
+
|
|
1098
|
+
[MIT](LICENSE).
|
|
1099
|
+
|
|
1100
|
+
### Desktop app and project setup
|
|
1101
|
+
|
|
1102
|
+
A local macOS shell, guided project setup, provider/model discovery, and weekly
|
|
1103
|
+
schedules with timezones use the same controller and approval flow as the CLI.
|
|
1104
|
+
The local native build also includes **File → Install app update** and
|
|
1105
|
+
**Update status**: same-schema updates drain current work, verify a private
|
|
1106
|
+
backup, atomically swap the app, and restore the previous app on failed health
|
|
1107
|
+
checks without replacing newer task data. An independent temporary macOS job
|
|
1108
|
+
automatically recovers an interrupted updater without reopening the window;
|
|
1109
|
+
bounded retries and a persistent Stop request prevent endless recovery or
|
|
1110
|
+
restarting work you stopped. Signed-release permission persistence
|
|
1111
|
+
and physical reboot acceptance remain release gates, not claims from unit tests.
|
|
1112
|
+
See [build and usage instructions](docs/control-app.md) and
|
|
1113
|
+
[the integration assessment](docs/CONTROL_APP_INTEGRATION.md).
|