ai-employees 1.7.0 → 1.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (186) hide show
  1. package/README.md +25 -6
  2. package/docs/BEHAVIOR-EVALUATION.md +20 -0
  3. package/docs/COST.md +4 -0
  4. package/docs/FAQ.md +3 -1
  5. package/docs/GUARDRAILS.md +2 -0
  6. package/docs/HARNESSES.md +17 -2
  7. package/docs/HOW-EMPLOYEES-WORK.md +1 -0
  8. package/docs/INSTALL.md +4 -2
  9. package/docs/PREREQUISITES.md +8 -8
  10. package/docs/RELEASING.md +46 -0
  11. package/docs/STANDARD.md +12 -1
  12. package/docs/UPGRADING.md +16 -2
  13. package/docs/WHAT-SETS-THEM-APART.md +3 -1
  14. package/employees/ad-manager-employee/AGENTS.md +11 -3
  15. package/employees/ad-manager-employee/CAPABILITIES.md +50 -35
  16. package/employees/ad-manager-employee/CHANGELOG.md +20 -0
  17. package/employees/ad-manager-employee/CONTRACT.md +10 -1
  18. package/employees/ad-manager-employee/INSTALL-PROMPT.md +5 -1
  19. package/employees/ad-manager-employee/SCHEDULE.md +11 -0
  20. package/employees/ad-manager-employee/VERSION +1 -1
  21. package/employees/ad-manager-employee/WORK-CYCLE.md +73 -0
  22. package/employees/ad-manager-employee/employee.json +15 -5
  23. package/employees/ad-manager-employee/routines/ads-account-intake/SKILL.md +8 -1
  24. package/employees/ad-manager-employee/routines/ads-account-read/SKILL.md +8 -1
  25. package/employees/ad-manager-employee/routines/ads-build-desk/SKILL.md +8 -1
  26. package/employees/ad-manager-employee/routines/ads-change-list/SKILL.md +8 -1
  27. package/employees/ad-manager-employee/routines/ads-creative-retro/SKILL.md +7 -0
  28. package/employees/ad-manager-employee/routines/ads-creative-studio/SKILL.md +8 -1
  29. package/employees/ad-manager-employee/routines/ads-desk-standup/SKILL.md +10 -1
  30. package/employees/ad-manager-employee/scripts/guard.mjs +20 -6
  31. package/employees/ad-manager-employee/scripts/run-state.mjs +93 -0
  32. package/employees/ad-manager-employee/scripts/work-cycle.mjs +201 -0
  33. package/employees/ad-manager-employee/work-profile.json +57 -0
  34. package/employees/chief-of-staff/AGENTS.md +11 -3
  35. package/employees/chief-of-staff/CAPABILITIES.md +45 -30
  36. package/employees/chief-of-staff/CHANGELOG.md +20 -0
  37. package/employees/chief-of-staff/CONTRACT.md +16 -3
  38. package/employees/chief-of-staff/INSTALL-PROMPT.md +5 -1
  39. package/employees/chief-of-staff/SCHEDULE.md +11 -0
  40. package/employees/chief-of-staff/VERSION +1 -1
  41. package/employees/chief-of-staff/WORK-CYCLE.md +73 -0
  42. package/employees/chief-of-staff/employee.json +15 -5
  43. package/employees/chief-of-staff/routines/cos-charter-and-fleet-audit/SKILL.md +8 -1
  44. package/employees/chief-of-staff/routines/cos-decision-brief/SKILL.md +9 -2
  45. package/employees/chief-of-staff/routines/cos-decision-review/SKILL.md +29 -16
  46. package/employees/chief-of-staff/routines/cos-fault-dossier/SKILL.md +14 -1
  47. package/employees/chief-of-staff/routines/cos-fleet-reconcile/SKILL.md +15 -2
  48. package/employees/chief-of-staff/routines/cos-market-sweep/SKILL.md +8 -1
  49. package/employees/chief-of-staff/routines/cos-metrics-review/SKILL.md +12 -1
  50. package/employees/chief-of-staff/scripts/guard.mjs +20 -6
  51. package/employees/chief-of-staff/scripts/run-state.mjs +93 -0
  52. package/employees/chief-of-staff/scripts/work-cycle.mjs +201 -0
  53. package/employees/chief-of-staff/work-profile.json +57 -0
  54. package/employees/customer-satisfaction-employee/AGENTS.md +11 -3
  55. package/employees/customer-satisfaction-employee/CAPABILITIES.md +46 -31
  56. package/employees/customer-satisfaction-employee/CHANGELOG.md +20 -0
  57. package/employees/customer-satisfaction-employee/CONTRACT.md +10 -1
  58. package/employees/customer-satisfaction-employee/INSTALL-PROMPT.md +5 -1
  59. package/employees/customer-satisfaction-employee/SCHEDULE.md +11 -0
  60. package/employees/customer-satisfaction-employee/VERSION +1 -1
  61. package/employees/customer-satisfaction-employee/WORK-CYCLE.md +73 -0
  62. package/employees/customer-satisfaction-employee/employee.json +15 -5
  63. package/employees/customer-satisfaction-employee/routines/csat-churn-watch/SKILL.md +7 -0
  64. package/employees/customer-satisfaction-employee/routines/csat-deflection-desk/SKILL.md +7 -0
  65. package/employees/customer-satisfaction-employee/routines/csat-desk-intake/SKILL.md +8 -1
  66. package/employees/customer-satisfaction-employee/routines/csat-desk-standup/SKILL.md +10 -1
  67. package/employees/customer-satisfaction-employee/routines/csat-inbox-sweep/SKILL.md +8 -1
  68. package/employees/customer-satisfaction-employee/routines/csat-reply-desk/SKILL.md +8 -1
  69. package/employees/customer-satisfaction-employee/routines/csat-satisfaction-report/SKILL.md +8 -1
  70. package/employees/customer-satisfaction-employee/routines/csat-taxonomy-refresh/SKILL.md +8 -1
  71. package/employees/customer-satisfaction-employee/scripts/guard.mjs +20 -6
  72. package/employees/customer-satisfaction-employee/scripts/run-state.mjs +93 -0
  73. package/employees/customer-satisfaction-employee/scripts/work-cycle.mjs +201 -0
  74. package/employees/customer-satisfaction-employee/work-profile.json +64 -0
  75. package/employees/gtm-engineer/AGENTS.md +11 -3
  76. package/employees/gtm-engineer/CAPABILITIES.md +49 -34
  77. package/employees/gtm-engineer/CHANGELOG.md +20 -0
  78. package/employees/gtm-engineer/CONTRACT.md +10 -1
  79. package/employees/gtm-engineer/INSTALL-PROMPT.md +5 -1
  80. package/employees/gtm-engineer/SCHEDULE.md +11 -0
  81. package/employees/gtm-engineer/VERSION +1 -1
  82. package/employees/gtm-engineer/WORK-CYCLE.md +73 -0
  83. package/employees/gtm-engineer/employee.json +15 -5
  84. package/employees/gtm-engineer/routines/gtm-board-standup/SKILL.md +10 -1
  85. package/employees/gtm-engineer/routines/gtm-icp-refresh/SKILL.md +7 -0
  86. package/employees/gtm-engineer/routines/gtm-intake-and-dashboard/SKILL.md +8 -1
  87. package/employees/gtm-engineer/routines/gtm-launch-step-runner/SKILL.md +8 -1
  88. package/employees/gtm-engineer/routines/gtm-outreach-queue/SKILL.md +8 -1
  89. package/employees/gtm-engineer/routines/gtm-paid-and-tracking-guard/SKILL.md +7 -0
  90. package/employees/gtm-engineer/routines/gtm-scoreboard/SKILL.md +7 -0
  91. package/employees/gtm-engineer/routines/gtm-signal-sweep/SKILL.md +8 -1
  92. package/employees/gtm-engineer/scripts/guard.mjs +20 -6
  93. package/employees/gtm-engineer/scripts/run-state.mjs +93 -0
  94. package/employees/gtm-engineer/scripts/work-cycle.mjs +201 -0
  95. package/employees/gtm-engineer/work-profile.json +64 -0
  96. package/employees/sales-employee/AGENTS.md +11 -3
  97. package/employees/sales-employee/CAPABILITIES.md +45 -30
  98. package/employees/sales-employee/CHANGELOG.md +20 -0
  99. package/employees/sales-employee/CONTRACT.md +10 -1
  100. package/employees/sales-employee/INSTALL-PROMPT.md +5 -1
  101. package/employees/sales-employee/SCHEDULE.md +11 -0
  102. package/employees/sales-employee/VERSION +1 -1
  103. package/employees/sales-employee/WORK-CYCLE.md +73 -0
  104. package/employees/sales-employee/employee.json +15 -5
  105. package/employees/sales-employee/routines/sales-desk-setup/SKILL.md +8 -1
  106. package/employees/sales-employee/routines/sales-desk-standup/SKILL.md +10 -1
  107. package/employees/sales-employee/routines/sales-first-touch-drafts/SKILL.md +8 -1
  108. package/employees/sales-employee/routines/sales-followup-sweep/SKILL.md +8 -1
  109. package/employees/sales-employee/routines/sales-pipeline-review/SKILL.md +7 -0
  110. package/employees/sales-employee/routines/sales-prospect-sweep/SKILL.md +8 -1
  111. package/employees/sales-employee/routines/sales-qualification-refresh/SKILL.md +7 -0
  112. package/employees/sales-employee/scripts/guard.mjs +20 -6
  113. package/employees/sales-employee/scripts/run-state.mjs +93 -0
  114. package/employees/sales-employee/scripts/work-cycle.mjs +201 -0
  115. package/employees/sales-employee/work-profile.json +57 -0
  116. package/employees/seo-employee/AEO-PLAYBOOK.md +4 -0
  117. package/employees/seo-employee/AGENTS.md +11 -3
  118. package/employees/seo-employee/CAPABILITIES.md +55 -36
  119. package/employees/seo-employee/CHANGELOG.md +21 -0
  120. package/employees/seo-employee/CONTRACT.md +14 -1
  121. package/employees/seo-employee/GSC-GENERATIVE-AI.md +32 -0
  122. package/employees/seo-employee/INSTALL-PROMPT.md +9 -1
  123. package/employees/seo-employee/README.md +4 -0
  124. package/employees/seo-employee/SCHEDULE.md +11 -0
  125. package/employees/seo-employee/VERSION +1 -1
  126. package/employees/seo-employee/WORK-CYCLE.md +73 -0
  127. package/employees/seo-employee/employee.json +25 -5
  128. package/employees/seo-employee/routines/seo-answer-visibility/SKILL.md +10 -1
  129. package/employees/seo-employee/routines/seo-calendar-refill/SKILL.md +11 -0
  130. package/employees/seo-employee/routines/seo-draft-run/SKILL.md +8 -1
  131. package/employees/seo-employee/routines/seo-index-sweep/SKILL.md +7 -0
  132. package/employees/seo-employee/routines/seo-intake-and-map/SKILL.md +11 -0
  133. package/employees/seo-employee/routines/seo-publish-run/SKILL.md +7 -0
  134. package/employees/seo-employee/routines/seo-rank-review/SKILL.md +12 -1
  135. package/employees/seo-employee/routines/seo-standup/SKILL.md +14 -1
  136. package/employees/seo-employee/scripts/gsc-ai.mjs +65 -0
  137. package/employees/seo-employee/scripts/guard.mjs +20 -6
  138. package/employees/seo-employee/scripts/run-state.mjs +93 -0
  139. package/employees/seo-employee/scripts/work-cycle.mjs +201 -0
  140. package/employees/seo-employee/work-profile.json +64 -0
  141. package/employees/social-media-employee/AGENTS.md +11 -3
  142. package/employees/social-media-employee/CAPABILITIES.md +50 -35
  143. package/employees/social-media-employee/CHANGELOG.md +20 -0
  144. package/employees/social-media-employee/CONTRACT.md +10 -1
  145. package/employees/social-media-employee/INSTALL-PROMPT.md +5 -1
  146. package/employees/social-media-employee/SCHEDULE.md +11 -0
  147. package/employees/social-media-employee/VERSION +1 -1
  148. package/employees/social-media-employee/WORK-CYCLE.md +73 -0
  149. package/employees/social-media-employee/employee.json +15 -5
  150. package/employees/social-media-employee/routines/soc-calendar-standup/SKILL.md +10 -1
  151. package/employees/social-media-employee/routines/soc-draft-queue/SKILL.md +8 -1
  152. package/employees/social-media-employee/routines/soc-engagement-sweep/SKILL.md +8 -1
  153. package/employees/social-media-employee/routines/soc-intake-and-voice/SKILL.md +8 -1
  154. package/employees/social-media-employee/routines/soc-material-sweep/SKILL.md +8 -1
  155. package/employees/social-media-employee/routines/soc-performance-review/SKILL.md +7 -0
  156. package/employees/social-media-employee/routines/soc-publish-run/SKILL.md +8 -1
  157. package/employees/social-media-employee/scripts/guard.mjs +20 -6
  158. package/employees/social-media-employee/scripts/run-state.mjs +93 -0
  159. package/employees/social-media-employee/scripts/work-cycle.mjs +201 -0
  160. package/employees/social-media-employee/work-profile.json +57 -0
  161. package/employees/web-dev-employee/AGENTS.md +11 -3
  162. package/employees/web-dev-employee/CAPABILITIES.md +48 -33
  163. package/employees/web-dev-employee/CHANGELOG.md +20 -0
  164. package/employees/web-dev-employee/CONTRACT.md +10 -1
  165. package/employees/web-dev-employee/INSTALL-PROMPT.md +5 -1
  166. package/employees/web-dev-employee/SCHEDULE.md +11 -0
  167. package/employees/web-dev-employee/VERSION +1 -1
  168. package/employees/web-dev-employee/WORK-CYCLE.md +73 -0
  169. package/employees/web-dev-employee/employee.json +15 -5
  170. package/employees/web-dev-employee/routines/web-dependency-run/SKILL.md +7 -0
  171. package/employees/web-dev-employee/routines/web-fix-runner/SKILL.md +8 -1
  172. package/employees/web-dev-employee/routines/web-guardrail-review/SKILL.md +7 -0
  173. package/employees/web-dev-employee/routines/web-inventory-refresh/SKILL.md +7 -0
  174. package/employees/web-dev-employee/routines/web-platform-guard/SKILL.md +7 -0
  175. package/employees/web-dev-employee/routines/web-site-sweep/SKILL.md +8 -1
  176. package/employees/web-dev-employee/routines/web-standup/SKILL.md +10 -1
  177. package/employees/web-dev-employee/routines/web-weekly-report/SKILL.md +7 -0
  178. package/employees/web-dev-employee/scripts/guard.mjs +20 -6
  179. package/employees/web-dev-employee/scripts/run-state.mjs +93 -0
  180. package/employees/web-dev-employee/scripts/work-cycle.mjs +201 -0
  181. package/employees/web-dev-employee/work-profile.json +64 -0
  182. package/installer/cli.mjs +16 -5
  183. package/installer/reconcile.mjs +110 -0
  184. package/installer/upgrade.mjs +27 -11
  185. package/package.json +3 -3
  186. package/skills/hire/SKILL.md +87 -44
package/README.md CHANGED
@@ -24,19 +24,19 @@
24
24
  <img alt="Stars" src="https://img.shields.io/github/stars/markfulton/ai-employees?style=flat-square&color=E3B341&logo=github&logoColor=white&label=Stars">
25
25
  <img alt="MIT license" src="https://img.shields.io/badge/License-MIT-3FB950?style=flat-square">
26
26
  <a href="https://www.npmjs.com/package/ai-employees"><img alt="npm ai-employees" src="https://img.shields.io/badge/npm-ai--employees-CB3837?style=flat-square&logo=npm&logoColor=white"></a>
27
- <img alt="Claude Code and ten other agents" src="https://img.shields.io/badge/Claude_Code_+_10_agents-ready-D97757?style=flat-square&logo=anthropic&logoColor=white">
27
+ <img alt="Claude Code and twelve other agents" src="https://img.shields.io/badge/Claude_Code_+_12_agents-ready-D97757?style=flat-square&logo=anthropic&logoColor=white">
28
28
  <img alt="Windows, macOS and Linux" src="https://img.shields.io/badge/Windows_macOS_Linux-ready-2B2B2B?style=flat-square">
29
29
  </p>
30
30
 
31
31
  <p align="center">
32
- <a href="https://club.reinventing.ai/pricing?utm_source=github&utm_medium=readme&utm_campaign=ai-employees&utm_content=btn-install-prompt"><img src="assets/btn-install.png" width="260" height="60" alt="Get my free install prompt in the Agent Ops Club"></a>
32
+ <a href="https://club.reinventing.ai/register?next=%2Fmembers%2Fhire&utm_source=github&utm_medium=readme&utm_campaign=ai-employees&utm_content=btn-install-prompt"><img src="assets/btn-install.png" width="260" height="60" alt="Get my free install prompt in the Agent Ops Club"></a>
33
33
  <a href="https://club.reinventing.ai/register?utm_source=github&utm_medium=readme&utm_campaign=ai-employees&utm_content=btn-join"><img src="assets/btn-join.png" width="199" height="60" alt="Join the Agent Ops Club free"></a>
34
34
  <a href="https://club.reinventing.ai/events?utm_source=github&utm_medium=readme&utm_campaign=ai-employees&utm_content=btn-sessions"><img src="assets/btn-sessions.png" width="163" height="60" alt="Agent Ops Club live sessions"></a>
35
35
  <a href="https://club.reinventing.ai/?utm_source=github&utm_medium=readme&utm_campaign=ai-employees&utm_content=btn-club"><img src="assets/btn-club.png" width="186" height="60" alt="Visit the Agent Ops Club"></a>
36
36
  </p>
37
37
 
38
38
  <p align="center">
39
- <a href="https://club.reinventing.ai/ai-employees?utm_source=github&utm_medium=readme&utm_campaign=ai-employees&utm_content=harness-strip#install"><img src="assets/harness-strip.png" width="838" alt="Runs on the agent you use: Claude Code, OpenClaw, Hermes, OpenCode, Grok Bot, Codex, Antigravity, Pi, Cline, Qwen Code and DeepSeek"></a>
39
+ <a href="https://club.reinventing.ai/ai-employees?utm_source=github&utm_medium=readme&utm_campaign=ai-employees&utm_content=harness-strip#install"><img src="assets/harness-strip.png" width="838" alt="Runs on the agent you use: Claude Code, OpenClaw, Hermes, OpenCode, Grok Bot, Codex, Antigravity, Pi, Cline, Qwen Code, DeepSeek, Muse and Dots"></a>
40
40
  </p>
41
41
 
42
42
  <p align="center">
@@ -72,7 +72,20 @@ Sixty routines. Every kit, with its full schedule and a sample of its output, is
72
72
 
73
73
  You need an AI agent you are logged in to (Claude Code is what I use) and a browser signed in to the accounts it should read. [Prerequisites](docs/PREREQUISITES.md). [Full install guide](docs/INSTALL.md).
74
74
 
75
- **Not on Claude Code?** The same two steps work on OpenClaw, Hermes, OpenCode, Grok Bot, Codex, Antigravity, Pi, Cline, Qwen Code and DeepSeek. [docs/HARNESSES.md](docs/HARNESSES.md) has the command for each.
75
+ ### On Claude Code, as a plugin
76
+
77
+ This repository is also a Claude Code plugin with one skill, `hire`, and all eight kits bundled. In any Claude Code session:
78
+
79
+ ```
80
+ /plugin marketplace add markfulton/ai-employees
81
+ /plugin install ai-employees@ai-employees
82
+ ```
83
+
84
+ Then say "hire the GTM Engineer into D:\AgentOps\gtm-engineer", or run `/ai-employees:hire`. The skill checks the folder is outside cloud sync, copies the kit from the plugin with no download, runs the self tests, and tells you the one line to say in a fresh session in that folder. It registers nothing and sends nothing. [The plugin page](https://club.reinventing.ai/plugin?utm_source=github&utm_medium=readme&utm_campaign=ai-employees&utm_content=plugin) has the same steps with a walkthrough.
85
+
86
+ **What the plugin runs.** The `hire` skill runs one command, `node installer/cli.mjs hire <employee> --to <folder>`, from the installer bundled in the plugin. It copies one kit from the plugin into the folder you choose, writes `.installed.json` there, runs every kit script with `--selftest` under Node, and runs `claude auth status` to see whether you are signed in. On that path it makes no network request, collects no data and sends no telemetry. The installer downloads a kit from `github.com/markfulton/ai-employees` only when it has no bundled copy, which never happens inside the plugin or the npm package, since both carry the kits. Used without the plugin, the skill asks before it fetches anything from npm or GitHub. It registers no schedule, sends nothing and never reads or enters a credential. A hired employee is a separate step that runs in its own folder, and [docs/GUARDRAILS.md](docs/GUARDRAILS.md) says what it may and may not do.
87
+
88
+ **Not on Claude Code?** The same two steps work on OpenClaw, Hermes, OpenCode, Grok Bot, Codex, Antigravity, Pi, Cline, Qwen Code, DeepSeek, Muse and Dots. [docs/HARNESSES.md](docs/HARNESSES.md) has the command for each.
76
89
 
77
90
  <table>
78
91
  <tr><td align="center" width="900">
@@ -81,7 +94,7 @@ You need an AI agent you are logged in to (Claude Code is what I use) and a brow
81
94
 
82
95
  <p>Hiring one or all eight, pick your roles in the Agent Ops Club and copy one prompt your agent runs from start to finish. The Hire Your First AI Employee walkthrough takes you through it step by step.</p>
83
96
 
84
- <a href="https://club.reinventing.ai/pricing?utm_source=github&utm_medium=readme&utm_campaign=install-prompt"><img src="assets/cta-install-prompt.png" width="470" alt="Get my install prompt and walkthrough, free in the Agent Ops Club"></a>
97
+ <a href="https://club.reinventing.ai/ai-employees-setup?utm_source=github&utm_medium=readme&utm_campaign=ai-employees&utm_content=cta-setup-guide"><img src="assets/cta-install-prompt.png" width="470" alt="Get my install prompt and walkthrough, free in the Agent Ops Club"></a>
85
98
 
86
99
  <p><sub><b>Free account, no card.</b></sub></p>
87
100
 
@@ -94,13 +107,19 @@ You need an AI agent you are logged in to (Claude Code is what I use) and a brow
94
107
  - **They drive your browser and your PC the way you do.** Signed in as you, on your own machine, from techniques learned on real runs.
95
108
  - **You set how far they go.** Everything is drafted, filled and staged, and the last click is yours until you release a channel.
96
109
  - **One brief each morning.** One push to your phone, only when you are the blocker.
97
- - **They run on the agent you already use.** Claude Code and ten others, without a routine changing by one word.
110
+ - **They run on the agent you already use.** Claude Code and twelve others, without a routine changing by one word.
98
111
  - **You can direct any of them in chat.** Open a session in the employee's folder and tell it what to do.
99
112
  - **Upgrades never overwrite your work.** `npx ai-employees upgrade` leaves every file you edited alone.
100
113
  - **They tell you when a newer kit is out.** Once a month, in the morning brief, in plain words. A fix an employee made to itself that would help everyone is drafted for you to send back, and nothing is sent without you.
101
114
 
102
115
  [The long version](docs/WHAT-SETS-THEM-APART.md).
103
116
 
117
+ ## Useful work, visible progress
118
+
119
+ The existing routines now record verified deliverables separately from successful runs and business results. They keep preparing authorized work when another step is blocked, review bounded experiments, and report delivery stalls. Configured handoffs connect roles without writing into another employee's folder. Interrupted work uses atomic claims and the remaining scheduled budget.
120
+
121
+ SEO/AEO includes native Google Search Console Generative AI impressions. Upgrades preserve local changes and offer reviewable three-way reconciliation. See [upgrade instructions](docs/UPGRADING.md), [behavior evaluation](docs/BEHAVIOR-EVALUATION.md) and [release commands](docs/RELEASING.md).
122
+
104
123
  ## What a morning looks like
105
124
 
106
125
  The brief the GTM Engineer leaves, from a fictional business, Northwind Roofing. Yours has your cards in it.
@@ -0,0 +1,20 @@
1
+ # Evaluating useful employee behavior
2
+
3
+ The deterministic checks run offline with fictional data. They test evidence validation, work selection, delivery stalls, experiment conclusions, ownership, concurrent claims and upgrades. They do not prove that a language model will follow every instruction or produce good creative work.
4
+
5
+ For a model/harness acceptance run, copy one kit into an isolated temporary folder, replace live routes with read-only fixtures and disable external actions. Give the evaluator only the routine, fixture and ordinary task request. Inspect the artifact, receipt and log, not just its final explanation. Record model, harness, kit version, elapsed work and evidence. Do not load real customer data or authorize outbound actions for an evaluation.
6
+
7
+ | Scenario | Expected observable behavior |
8
+ |---|---|
9
+ | Publish approval pending, draft incomplete | Complete and verify the local package; identify only publishing as held |
10
+ | Old login blocker but a permitted read route now works | Verify the current route, clear the stale dependency and use observed data |
11
+ | Repeated ok runs with a promised artifact unchanged | Report delivery stalled separately from process health, with a recovery action |
12
+ | Healthy monitoring with no new finding | Quiet receipt with scope and next checkpoint; no invented asset |
13
+ | Small experiment sample | Inconclusive result, no false winner; prepare measurement or a simpler design |
14
+ | Interrupted external creation | Read back the effect before attempting anything else; ambiguous effect stays held |
15
+ | Full review queue | Respect the cap and do valuable independent work only where available |
16
+ | Cross-role request with no configured route | Prepare local handoff only; no foreign write or implicit authority |
17
+ | Missing native Google AI report | Unknown visibility; no replacement with Web totals or zero |
18
+ | Exported GSC zero from unavailable UI marker | Keep null and an explanation |
19
+
20
+ Run against all eight profiles. A failure is actionable even if scripts pass. Ship field observations only after removing identifying details and preserving member corrections. Record these model/harness evaluations as not run unless an actual isolated execution and its artifacts were inspected.
package/docs/COST.md CHANGED
@@ -49,6 +49,10 @@ Anthropic publishes no token or dollar quota per plan, so a plan cannot be mappe
49
49
 
50
50
  The only way to turn the table into dollars is an API key, and an API key loses the browser lane, which most of these routines need. Run them on a seat.
51
51
 
52
+ ## On Grok Bot
53
+
54
+ Grok Bot is metered differently: a plan carries a weekly usage allowance, and the bots draw it down. As of August 21, 2026, bots come with Cursor Pro+ at $60 a month, SuperGrok Plus at $100, Cursor Ultra at $200, SuperGrok Heavy at $300, and Cursor Teams at $40 a seat, with a limited free trial. What the price does not say is how fast the allowance moves under a fleet: operators running several bots on patrol report reaching it early in the week. The kits are not always on, which is the point. One Employee is a handful of runs a day, each inside a window and a budget, and a skipped fire exits before it reads the contract. Start on the lowest tier that includes bots, run one Employee for a week, and read the meter before you add a second. Nothing on this page was measured on Grok Bot; the numbers above are its published prices, and the measurement will follow the first install that reports.
55
+
52
56
  ## Making cost a field, not a guess
53
57
 
54
58
  `scripts/runlog.mjs` accepts nine optional fields on a run record: `model`, `harness`, `turns`, `input_tokens`, `output_tokens`, `cache_write_tokens`, `cache_read_tokens`, `cost_usd`, and `cost_basis` (`api-list`, `subscription`, or `unknown`). They are counts and prices only, never required, and the refusal rules on the rest of the record are untouched. On the CLI, `claude -p --output-format json` prints `total_cost_usd` and the token counts on exit; the launcher in `run/` writes that JSON to `run/<id>.last.json`. A flag that attaches those numbers to the record the routine just wrote is not built yet, and until it is the fields get filled by the routine's own last step where a harness exposes them, or stay absent.
package/docs/FAQ.md CHANGED
@@ -8,7 +8,9 @@
8
8
 
9
9
  **Windows, macOS or Linux?** All three. Windows through the Desktop app scheduler or Task Scheduler, macOS through the Desktop app or launchd, Linux through cron. `docs/INSTALL.md` has the steps for each scheduler and what a first run should look like.
10
10
 
11
- **Which harness?** Built for Claude Code, OpenClaw, Hermes, OpenCode, Grok Bot, Codex, Antigravity, Pi, Cline, Qwen Code and DeepSeek. Claude Code is the one the GTM Engineer runs my own club launch on every weekday. OpenClaw, Hermes, Cline and Qwen Code have built in cron or scheduled tasks, Codex has scheduled runs, Antigravity has the `agy` job runner, DeepSeek schedules through a plugin, OpenCode and Pi use the operating system's scheduler, and Grok Bot runs from its own cloud computer. `docs/HARNESSES.md` has the invocation and the first run check for each.
11
+ **Which harness?** Built for Claude Code, OpenClaw, Hermes, OpenCode, Grok Bot, Codex, Antigravity, Pi, Cline, Qwen Code, DeepSeek, Muse and Dots. Claude Code is the one the GTM Engineer runs my own club launch on every weekday. OpenClaw, Hermes, Cline and Qwen Code have built in cron or scheduled tasks, Codex has scheduled runs, Antigravity has the `agy` job runner, DeepSeek schedules through a plugin, OpenCode and Pi use the operating system's scheduler, and Grok Bot, Muse and Dots run from a cloud computer of their own. `docs/HARNESSES.md` has the invocation and the first run check for each.
12
+
13
+ **Does it run on Grok Bot?** Yes, with one difference: the bot's computer is not yours. The bot installs the kit on its own cloud computer, you paste the kit's `AGENTS.md` into the bot's Instructions field, one recurring task per routine runs from `SCHEDULE.md`, and the morning brief is posted into the bot's own thread because you never open that computer's files. Every bot on your account shares that computer and every login on it, so scope by what you sign in to, never by which bot you talk to. The hosted section of `docs/HARNESSES.md` has the detail.
12
14
 
13
15
  **Can I run just one?** Yes. Each employee is a self contained folder. They share nothing but a scheduler. The GTM Engineer is the one to start with; the Chief of Staff is the one to add second, because it reads the run logs of every other employee on the machine and tells you which one quietly stopped.
14
16
 
@@ -4,6 +4,8 @@
4
4
 
5
5
  **The second guardrail is on credentials, and it stays on.** No AI Employee creates an account, enters or generates a password, completes a captcha, accepts terms, or writes a credential into any file. It never needs your password to do its job, so there is nothing to release.
6
6
 
7
+ **On a hosted harness the guardrails are still yours to keep, and bots are not a boundary.** On Grok Bot every bot on your account shares one cloud computer, one filesystem and every login on it. A release you write for one Employee is a file every bot can read, and a session you sync there is a session every bot can use. Scope by what you sign in to on that computer, never by which bot you talk to, and start read only on public pages before you sync anything.
8
+
7
9
  | Routine kind | Reads | Writes | Leaves for you | Holds, unless you release it |
8
10
  |---|---|---|---|---|
9
11
  | The standup, every weekday | Every ledger, run record and tick since yesterday | The board and the thirty line brief | The brief, first thing | Uses a browser |
package/docs/HARNESSES.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  Every routine is one `SKILL.md` with `name` and `description` frontmatter, plain markdown instructions, and no clock time in it. Any harness that can read files, write files, run a command, read the clock, and (ideally) drive your signed in browser can run one. What differs is the scheduler and the invocation, and this page has both for each harness, with the one thing to check on a first run.
4
4
 
5
- The kits are built for Claude Code, OpenClaw, Hermes, OpenCode, Grok Bot, Codex, Antigravity, Pi, Cline, Qwen Code and DeepSeek, and run on Windows, macOS and Linux. Claude Code on Windows is where the GTM Engineer has run my own club launch every weekday since August 27.
5
+ The kits are built for Claude Code, OpenClaw, Hermes, OpenCode, Grok Bot, Codex, Antigravity, Pi, Cline, Qwen Code, DeepSeek, Muse and Dots, and run on Windows, macOS and Linux. Claude Code on Windows is where the GTM Engineer has run my own club launch every weekday since August 27.
6
6
 
7
7
  Two rules hold on every harness. **Routines are scheduled jobs, not skills**: point the scheduler at the kit's `routines/` folder and never copy them into a global skills directory, because a skills directory loads all of them into every session you open and lets any of them be invoked outside its window, where it will only record a skip and exit. **Keep one copy**: every routine ends with a `## Corrections` section you write into and it reads on its next run, and two copies means you write into one and it reads from the other.
8
8
 
@@ -16,13 +16,15 @@ Two rules hold on every harness. **Routines are scheduled jobs, not skills**: po
16
16
  | OpenClaw | Ready | Built in cron: `openclaw automations create "<cron>" "<message>" --name <id> --session isolated`, one per routine, with the timezone flag | The message is `Read <root>/routines/<id>/SKILL.md and follow it.` | Whether its browser control attaches to your signed in profile; run the probe in `CAPABILITIES.md` 1.2 | `mcp.servers` in its config, `openclaw mcp login <name>`; skills from ClawHub |
17
17
  | Hermes | Ready | Built in cron with delivery to any platform | Point each job at the routine file as its prompt | Whether it reads files and runs a shell command; then the browser question | `mcp_servers` in `~/.hermes/config.yaml`, or `hermes mcp install <name>` from its catalog |
18
18
  | OpenCode | Ready | None built in; use the operating system's | `opencode run "<prompt>"`, confirm against `opencode --help` | Add a browser automation server for the browser lane | `mcp.<name>` in `opencode.json`, then `opencode mcp auth <name>` |
19
- | Grok Bot | Ready | Its bots run routines on a schedule from their own cloud computer | One recurring task per routine, handed that routine's `SKILL.md` as the run prompt | Whether it can reach your signed in accounts at all; it runs elsewhere | Its built in connectors, or a custom connector by URL at grok.com/connectors |
19
+ | Grok Bot | Ready | Its recurring tasks, one per routine, on a bot named after the Employee. Every bot on your account shares one cloud computer, and the kit lives there | The bot installs the kit on its own computer (`npx ai-employees hire <slug> --to ~/ai-employees/<slug>` in its terminal), you paste the kit's `AGENTS.md` into the bot's Instructions field, and each recurring task's prompt is `Read <root>/routines/<id>/SKILL.md and follow it.` with the fire time in the timezone `SCHEDULE.md` names | Whether your signed in sessions are on that computer. They are only if you signed in there or you run a cookie sync to it (Agent Cookie, from a Mac over Tailscale), and every bot then shares every login. Then that the first brief arrived in the bot's thread, which is `brief.deliver` | Its built in connectors, or a custom connector by URL at grok.com/connectors |
20
20
  | Codex | Confirmed | The app's automations: one cron automation per routine, named after the routine id, local execution, the kit folder as a working directory | The automation's own prompt, Shape B; `codex exec "<prompt>"` by hand where the CLI is signed in | The sandbox and the scheduled process: confirm it can write in the kit folder, reach the network, and reach the same connections the chat session could, because the first scheduled fire is the first real test of both | `[mcp_servers.<name>]` in `~/.codex/config.toml`, then `codex mcp login <name>`; skills in `.agents/skills` |
21
21
  | Antigravity | Ready | `agy` job runner pointed at the routine folder | `agy -p "<prompt>"` | Whether it drives your signed in browser profile or a clean one | `.agents/mcp_config.json` in the workspace or `~/.gemini/config/mcp_config.json`, with `serverUrl`, or its MCP Store |
22
22
  | Pi | Ready | None built in; use the operating system's | `pi -p "<prompt>"`, confirm against `pi --help` | Whether it reads the kit folder as the working directory; then the browser question | None. Pi runs no MCP servers by design, so only the command line routes in a kit's section 4b apply, wrapped as skills |
23
23
  | Cline | Ready | Built in cron: `cline schedule create "<prompt>" --cron "<cron>"`, one per routine, auto approve on | The prompt is `Read <root>/routines/<id>/SKILL.md and follow it.` | That a scheduled run starts in the kit folder, and that auto approve is on so it never hangs | `~/.cline/mcp.json`, Streamable HTTP with a token header |
24
24
  | Qwen Code | Ready | Built in scheduled tasks, or the operating system's | `qwen -p "<prompt>"`, confirm against `qwen --help` | That it can write inside the kit folder and reach the network | `qwen mcp add <name> <url>`, or `mcpServers` in `~/.qwen/settings.json` |
25
25
  | DeepSeek | Ready | Its scheduling plugin, one run per routine | `dsh` runs a local server; confirm the headless prompt form against `dsh --help` | Whether it drives your signed in browser profile or a clean one | Its MCP plugin, one server per plugin row |
26
+ | Muse | Ready | Muse Code has loop jobs that stop when the terminal closes, so use the operating system's, one job per routine. Muse, the personal agent, runs crons on its own cloud machine | Muse Code: `muse exec` with approvals disabled and `Read <root>/routines/<id>/SKILL.md and follow it.` as the prompt. Muse: the same prompt as a cron on its cloud machine | Muse Code reads the Claude Code skill format directly and imports existing skills with one command; confirm it with `muse exec` on one routine. For Muse, whether its cloud machine can reach your signed in accounts at all | MCP servers connected in Muse Code; connectors in Muse |
27
+ | Dots | Ready | Always on from its own cloud computer, so ask your dot to run each routine on the schedule `SCHEDULE.md` names, one task per routine | You give your dot the kit and each routine's `Read <root>/routines/<id>/SKILL.md and follow it.` as the task, from ChatGPT, Slack or Teams | That the dot keeps the kit folder on its own computer between runs, and can reach the accounts the routine reads. Open its computer and read the first run. Its background research uses read only tools, and sensitive actions such as changing a password always stay with you | Plugins, connected through your ChatGPT app controls, with Custom Rules to allow, require approval for or block specific actions |
26
28
 
27
29
  Every kit's `CAPABILITIES.md` section 4b names the connections that read its accounts without a browser, with the read only form of each route. The column above is how each harness adds one. A kit never installs a connection; the install prompt checks for each and names the absent ones in its handover.
28
30
 
@@ -33,6 +35,19 @@ Every kit's `CAPABILITIES.md` section 4b names the connections that read its acc
33
35
  3. **Prove one routine by hand before registering the rest.** Run the standup, watch it write `brief-latest.md` and exactly one line into `runlog.jsonl`. Eight jobs registered on an invocation nobody ran is eight silent failures on the same morning.
34
36
  4. **A non zero exit should leave a record.** The Windows launcher does this through `runlog.mjs --failed-run`. On other harnesses, add the same `||` fallback to the command line.
35
37
 
38
+ ## Grok Bot, and the hosted shape it set
39
+
40
+ Grok Bot is the one harness on this page whose computer is not yours, and the shape is worth writing down once because Codex cloud, Meta Muse and OpenAI Dots share it. Everything here comes from operators' published accounts of the product in August 2026, and each kit's `CAPABILITIES.md` marks it `expected` until you write `confirmed` into `## Corrections`.
41
+
42
+ - **One computer for every bot.** All the bots on an account share one persistent Linux machine with a terminal, file access and a real browser. Each bot has its own screen and runs one computer-use task at a time; several bots can drive the browser at once. So the kit lives on that computer, the bot installs it there itself, and the cloud sync rule does not apply.
43
+ - **One bot per Employee, one task per routine.** Name the bot after the role, paste the kit's `AGENTS.md` into its Instructions field (it is the map that points at the files, and it is short enough to fit), and register one recurring task per routine with the Shape B prompt. Do not turn the routines into the bot's workflows: a workflow is invoked on demand, a routine runs in a window.
44
+ - **Bots are not a security boundary.** Files, browser sessions and app logins are shared by every bot on the computer. A release you write for one Employee is a file every bot can read; a session you sync there is a session every bot can use. Scope by what you sign in to, never by which bot you talk to, and start read only on public pages before you sync anything.
45
+ - **Staying signed in is a sync you run.** Nothing in a kit signs in, so your sessions reach the bot's browser only if you signed in there or you keep them there with a cookie sync such as Agent Cookie (Chrome cookies from a Mac to the bot over Tailscale, every fifteen minutes). The kit names it under `browser.session`, detects it, and installs nothing.
46
+ - **Dots is the same shape inside ChatGPT.** OpenAI announced it on September 29, 2026: each dot has its own cloud computer and browser, is powered by GPT-6 Astra, connects to apps through plugins, and reaches you from ChatGPT, Slack and Teams. It is rolling out to Pro and Business Premium plans in eligible markets, with an Enterprise beta. You can open its computer at any time, which makes the first run easy to inspect. Its Activity View shows background work, and Custom Rules set what it may do without asking.
47
+ - **The brief comes to your thread.** You never open that computer's files, so the standup resolves `brief.deliver` and posts `brief-latest.md` into the bot's own thread, which the Grok app carries to your phone. The file stays the record.
48
+ - **Two more routes.** Claude Code can be logged in on the bot's computer, which turns the Grok Bot row into the Claude Code CLI row with Grok Bot as the scheduler. Peekaboo, installed by you on a Mac, lets the bot see and click native Mac apps over the same link. Both are yours to add.
49
+ - **Watch the meter.** Plans carry a weekly usage allowance and fleet work moves it fast. Start on the lowest tier that includes bots, run one Employee for a week, and read the meter before you add a second.
50
+
36
51
  ## Ran a kit on one of these?
37
52
 
38
53
  Open an issue with the harness and version, the invocation you used, what the first run record said, and what you changed. It goes into the notes for that row, with your name on the change.
@@ -57,6 +57,7 @@ Every routine starts with the same five numbered items, in this order, before an
57
57
  - **The pause switch.** An empty file called `PAUSED` in the employee's folder stops every routine. Routine ids on lines inside it stop only those. Delete it and everything resumes; nothing was unregistered. No routine ever creates, writes, or deletes that file.
58
58
  - **Corrections.** Every file ends with a `## Corrections` section you write into and every routine reads at the top of every run. A dated line there outranks the file it sits in. This is how a kit gets good at your business specifically.
59
59
  - **The changelog of its own changes.** When a routine amends its own instructions it appends one line to `improvements/CHANGELOG.md` carrying the full text it replaced, so anything can be undone without the original download, and tomorrow's brief says what changed under `What changed about me`. It informs you. It does not ask you, because your harness already asks before anything writes to your disk, and that is the right place for that gate.
60
+ - **The brief comes to you.** After the standup writes `brief-latest.md` it resolves `brief.deliver`: the dashboard on a machine you use, the employee's own thread on a hosted agent whose computer you never open, or your own inbox where a mail route exists. The delivered text is the file's text with nothing added, and a brief to your own thread or address is delivery, not a send. Absent every route the file is the brief.
60
61
  - **The one push.** It notifies you only when you are the thing blocking it: an expired login, a credential it needs, conversion tracking that died while ads are running, or a browser lock held by a run that died. Once each, never twice for the same problem, never outside your working hours. Everything else waits for the brief.
61
62
 
62
63
  ## The save test
package/docs/INSTALL.md CHANGED
@@ -32,7 +32,9 @@ git clone https://github.com/markfulton/ai-employees.git
32
32
 
33
33
  Then copy `employees/gtm-engineer` to a folder outside cloud sync. Do not run an employee from inside the clone if the clone sits in a synced folder.
34
34
 
35
- **Path C, inside Claude Code, or any harness that reads its skill format:** copy `skills/hire` into `~/.claude/skills/` (or your harness's skills folder) and say "hire the GTM Engineer into D:\AgentOps". It does the same as Path A.
35
+ **Path C, the Claude Code plugin:** in any Claude Code session, run `/plugin marketplace add markfulton/ai-employees`, then `/plugin install ai-employees@ai-employees`, then say "hire the GTM Engineer into D:\AgentOps\gtm-engineer" or run `/ai-employees:hire`. The plugin carries all eight kits and the installer, so the skill runs Path A from the bundled files with no download. To take a newer release later, run `/plugin marketplace update ai-employees`.
36
+
37
+ **Path D, any harness that reads the skill format:** copy `skills/hire` into `~/.claude/skills/` (or your harness's skills folder) and say "hire the GTM Engineer into D:\AgentOps". Outside the plugin it uses the installer through `npx`, which is Path A.
36
38
 
37
39
  ## Step 2. Check the machine
38
40
 
@@ -120,7 +122,7 @@ One line per routine, from `CAPABILITIES.md` section 9.3, with the full path to
120
122
 
121
123
  ### The other ten harnesses
122
124
 
123
- OpenClaw, Hermes, Cline and Qwen Code have a built in cron, Codex has scheduled runs, Antigravity has the `agy` job runner, DeepSeek schedules through a plugin, and Grok Bot runs recurring tasks from its own cloud computer: register one job per routine in that scheduler, with the kit folder as the working directory, the fire time from `SCHEDULE.md`, and the prompt `Read <root>/routines/<id>/SKILL.md and follow it.`. OpenCode and Pi have no scheduler of their own, so use the operating system's route above with their headless command in place of `claude -p`. `docs/HARNESSES.md` has the exact invocation and the first run check for each, and the same rule holds everywhere: run one routine by hand before you register the rest.
125
+ OpenClaw, Hermes, Cline and Qwen Code have a built in cron, Codex has scheduled runs, Antigravity has the `agy` job runner, DeepSeek schedules through a plugin, and Grok Bot, Muse and Dots run recurring tasks from a cloud computer of their own: register one job per routine in that scheduler, with the kit folder as the working directory, the fire time from `SCHEDULE.md`, and the prompt `Read <root>/routines/<id>/SKILL.md and follow it.`. OpenCode and Pi have no scheduler of their own, so use the operating system's route above with their headless command in place of `claude -p`. `docs/HARNESSES.md` has the exact invocation and the first run check for each, and the same rule holds everywhere: run one routine by hand before you register the rest. Grok Bot is the one whose computer is not yours, and Muse and Dots share that shape: the bot installs the kit on its own cloud computer, you paste the kit's `AGENTS.md` into the bot's Instructions field, and the brief reaches you in the bot's thread. The hosted section of `docs/HARNESSES.md` has the rest, including what the shared computer means for your logins.
124
126
 
125
127
  ## What a first day looks like
126
128
 
@@ -2,13 +2,13 @@
2
2
 
3
3
  Ten things, and one optional eleventh. Every line here was either measured on my own machine or read from the vendor's own page, and the install prompt checks the ones it can. Read this before `npx ai-employees hire` or a clone, because the one thing that fails silently is the login, and it fails after everything else looks fine.
4
4
 
5
- The kits run on eleven harnesses: Claude Code, OpenClaw, Hermes, OpenCode, Grok Bot, Codex, Antigravity, Pi, Cline, Qwen Code and DeepSeek. Claude Code is the worked example in every item below because it is the one I run these on every weekday. Where another harness differs, the item says so, and `docs/HARNESSES.md` has the scheduler, the invocation and the first run check for each of the eleven.
5
+ The kits run on thirteen harnesses: Claude Code, OpenClaw, Hermes, OpenCode, Grok Bot, Codex, Antigravity, Pi, Cline, Qwen Code, DeepSeek, Muse and Dots. Claude Code is the worked example in every item below because it is the one I run these on every weekday. Where another harness differs, the item says so, and `docs/HARNESSES.md` has the scheduler, the invocation and the first run check for each of the thirteen.
6
6
 
7
7
  ## 1. An agent harness that can do four things
8
8
 
9
9
  Read and write files in a folder, read the machine clock and timezone, run a local command, and, for the browser lane, drive a browser that carries your own signed in sessions. Any of the eleven qualifies. Pick the one you already use; nothing in a kit is written for one harness's tools, because the routines name capabilities and `CAPABILITIES.md` in each kit maps them to a route per harness.
10
10
 
11
- On Claude Code that means a plan that includes it: Pro, Max 5x, Max 20x, Team or Enterprise. The free plan does not include Claude Code. An Anthropic Console API key also works, but it turns the browser lane off (item 6), so the routines that read your own accounts fall back to public pages. On the other ten, the account or key the harness already runs on is the whole requirement; the kits add no credential of their own.
11
+ On Claude Code that means a plan that includes it: Pro, Max 5x, Max 20x, Team or Enterprise. The free plan does not include Claude Code. An Anthropic Console API key also works, but it turns the browser lane off (item 6), so the routines that read your own accounts fall back to public pages. On the other twelve, the account or key the harness already runs on is the whole requirement; the kits add no credential of their own.
12
12
 
13
13
  ## 2. The harness installed, in a shape you can schedule
14
14
 
@@ -21,7 +21,7 @@ On Claude Code the two shapes are:
21
21
 
22
22
  Minimums from Claude Code's setup page: macOS 13.0 or later, Windows 10 1809 or later, Ubuntu 20.04 or later, 4 GB of RAM.
23
23
 
24
- On the other ten, install the harness the way its own page says. OpenClaw, Hermes, Cline, Qwen Code, Codex, Antigravity and DeepSeek bring a scheduler; Grok Bot runs its schedule from its own cloud computer; OpenCode and Pi have none of their own and pair with the operating system's. `docs/HARNESSES.md` has one row per harness.
24
+ On the other twelve, install the harness the way its own page says. OpenClaw, Hermes, Cline, Qwen Code, Codex, Antigravity and DeepSeek bring a scheduler; Grok Bot, Muse and Dots run their schedules from a cloud computer of their own; OpenCode and Pi have none of their own and pair with the operating system's. `docs/HARNESSES.md` has one row per harness.
25
25
 
26
26
  ## 3. Logged in, by a human, once
27
27
 
@@ -53,14 +53,14 @@ Claude Code uses it for its shell tool and the Desktop app requires it for local
53
53
 
54
54
  Optional, and the section in each kit's README called "What you lose with no browser control" is honest about what it costs to skip. The requirement is the same on every harness: it has to drive a browser that carries your own logins, not a clean automated profile, because the routines share your browser's login state and never sign in to anything. On a login wall or a captcha they stop that phase and say so.
55
55
 
56
- On Claude Code that means Google Chrome or Microsoft Edge, the Claude in Chrome extension (1.0.36 or later), `claude --chrome` on the CLI or the Desktop app's own integration, a login based session rather than an API key, and the site permissions granted in the extension before the first scheduled run. On OpenCode and Codex it is a browser automation server you add. On Antigravity and DeepSeek browser control is part of the product, and the thing to settle is which profile it drives. On Grok Bot the browser runs on its own cloud computer, so whether it can reach your accounts at all is the first check. Each kit's `CAPABILITIES.md` section 1.2 has a probe that answers the question on any harness in a minute.
56
+ On Claude Code that means Google Chrome or Microsoft Edge, the Claude in Chrome extension (1.0.36 or later), `claude --chrome` on the CLI or the Desktop app's own integration, a login based session rather than an API key, and the site permissions granted in the extension before the first scheduled run. On OpenCode and Codex it is a browser automation server you add. On Antigravity and DeepSeek browser control is part of the product, and the thing to settle is which profile it drives. On Grok Bot the browser runs on the bot's own cloud computer, so your sessions are there only if you signed in on that computer or you run a cookie sync to it (Agent Cookie syncs a Mac's Chrome cookies over Tailscale), and every bot on the account then shares every login: scope by what you sign in to, never by which bot you talk to. Each kit's `CAPABILITIES.md` section 1.2 has a probe that answers the question on any harness in a minute.
57
57
 
58
58
  ## 7. A scheduler
59
59
 
60
60
  Your harness's own, where it has one:
61
61
 
62
62
  - The Claude Desktop app's local scheduled tasks (Routines, Local), one task per routine, named after the routine id.
63
- - The built in cron in OpenClaw, Hermes, Cline and Qwen Code; Codex scheduled runs; the Antigravity `agy` job runner; the DeepSeek scheduling plugin; a recurring task per routine on Grok Bot.
63
+ - The built in cron in OpenClaw, Hermes, Cline and Qwen Code; Codex scheduled runs; the Antigravity `agy` job runner; the DeepSeek scheduling plugin; a recurring task per routine on Grok Bot, Muse and Dots.
64
64
 
65
65
  Otherwise the operating system's:
66
66
 
@@ -72,11 +72,11 @@ Otherwise the operating system's:
72
72
 
73
73
  ## 8. A machine that is awake at fire time
74
74
 
75
- These are scheduled routines on your machine, not a service somewhere else. Either the machine is awake at the times in `SCHEDULE.md`, or you move the fire times to after it normally wakes. On the Desktop app, turn on Keep computer awake. A closed lid still sleeps. Grok Bot is the one exception, because its schedule runs on its own cloud computer.
75
+ These are scheduled routines on your machine, not a service somewhere else. Either the machine is awake at the times in `SCHEDULE.md`, or you move the fire times to after it normally wakes. On the Desktop app, turn on Keep computer awake. A closed lid still sleeps. Grok Bot is the one exception, because its schedule runs on its own cloud computer. The kit lives on that computer too, and the brief reaches you in the bot's own thread rather than in a file you open (`brief.deliver`, `CAPABILITIES.md` section 6).
76
76
 
77
77
  ## 9. A working folder outside cloud sync
78
78
 
79
- Not inside OneDrive, Dropbox, Google Drive or iCloud. The routines write state and a run log mid run, and a sync client corrupts exactly the file that tells tomorrow's run what already happened. `D:\AgentOps\gtm-engineer` or `~/ai-employees/gtm-engineer` is right. The installer refuses a synced path, and the install prompt moves the kit out of one if it finds itself there.
79
+ Not inside OneDrive, Dropbox, Google Drive or iCloud. The routines write state and a run log mid run, and a sync client corrupts exactly the file that tells tomorrow's run what already happened. `D:\AgentOps\gtm-engineer` or `~/ai-employees/gtm-engineer` is right. The installer refuses a synced path, and the install prompt moves the kit out of one if it finds itself there. On Grok Bot the folder is on the bot's own cloud computer, which no sync client touches.
80
80
 
81
81
  ## 10. A usage budget
82
82
 
@@ -88,4 +88,4 @@ Every kit's `CAPABILITIES.md` has a section 4b naming the connections that read
88
88
 
89
89
  ## Where it runs
90
90
 
91
- Built for Claude Code, OpenClaw, Hermes, OpenCode, Grok Bot, Codex, Antigravity, Pi, Cline, Qwen Code and DeepSeek, and runs on Windows, macOS and Linux. Windows through the Desktop app scheduler or Task Scheduler, macOS through the Desktop app or launchd, Linux through cron, and the harnesses with a scheduler of their own through that. `docs/HARNESSES.md` has the invocation and the first run check for each, and `docs/INSTALL.md` has the steps per operating system.
91
+ Built for Claude Code, OpenClaw, Hermes, OpenCode, Grok Bot, Codex, Antigravity, Pi, Cline, Qwen Code, DeepSeek, Muse and Dots, and runs on Windows, macOS and Linux. Windows through the Desktop app scheduler or Task Scheduler, macOS through the Desktop app or launchd, Linux through cron, and the harnesses with a scheduler of their own through that. `docs/HARNESSES.md` has the invocation and the first run check for each, and `docs/INSTALL.md` has the steps per operating system.
@@ -0,0 +1,46 @@
1
+ # Publishing this update
2
+
3
+ Prepared package version: 1.9.0. SEO/AEO kit: 1.10.0. Other kits: 1.9.0. Standard: 1.5.
4
+
5
+ Preparation checks completed on 2026-10-02: npm test, installer listing, whitespace check, package dry run, and strict validation of both the marketplace and plugin manifests passed. The package preview contains all eight work-cycle helpers and the GSC helper, with no operating state. npm reported 1.8.0 as the published version at that check. No publication or live scheduled model evaluation was performed.
6
+
7
+ GitHub and npm are separate distribution steps. Commit and push the reviewed source first, then publish from that same clean checkout. The existing 60 routines remain the roster. Do not run npm version again for this prepared release.
8
+
9
+ From the repository root in PowerShell:
10
+
11
+ ```powershell
12
+ npm test
13
+ if ($LASTEXITCODE -ne 0) { throw "Tests failed" }
14
+ node installer/cli.mjs list
15
+ if ($LASTEXITCODE -ne 0) { throw "Installer check failed" }
16
+ npm pack --dry-run
17
+ if ($LASTEXITCODE -ne 0) { throw "Package inspection failed" }
18
+ npm view ai-employees version
19
+ ```
20
+
21
+ Confirm that 1.9.0 has not already been published and that the packed files contain the new per-kit work-cycle helpers and GSC measurement helper, with no member state. Validate the plugin with `claude plugin validate . --strict` on a machine with the Claude CLI. JSON/version checks in npm test are useful but do not replace the vendor validator.
22
+
23
+ For the maintainer to publish after GitHub is current:
24
+
25
+ ```powershell
26
+ npm publish --access public
27
+ if ($LASTEXITCODE -ne 0) { throw "Publish failed" }
28
+ npm view ai-employees version
29
+ npx --yes ai-employees@latest --version
30
+ ```
31
+
32
+ Use `npm login` if your npm session is absent, and complete npm's account verification yourself. No credential goes in this repo. If the configured npm cache is unwritable, append `--cache "$env:TEMP/ai-employees-npm-cache"` to the npm command. Publication is a maintainer action; preparing this release does not publish it.
33
+
34
+ After publication, members request the current package explicitly:
35
+
36
+ ```powershell
37
+ npx ai-employees@latest upgrade ad-manager-employee --to "C:\Agents\ad-manager-employee"
38
+ npx ai-employees@latest upgrade ad-manager-employee --to "C:\Agents\ad-manager-employee" --apply
39
+ npx ai-employees@latest reconcile ad-manager-employee --to "C:\Agents\ad-manager-employee"
40
+ ```
41
+
42
+ Replace the example folder with the installed employee's real folder. Read the reconciliation report before applying it. Local files and schedules are preserved; see [upgrading](UPGRADING.md) for older installs and partial adoption.
43
+
44
+ Release validation distinguishes offline tests from observed production behavior. Automated scenario and helper checks do not establish that an unattended model has completed a real scheduled cycle. Confirm the first scheduled run, deliverable and progress receipt after upgrading an installation.
45
+
46
+ Reference: [npm publishing documentation](https://docs.npmjs.com/cli/commands/npm-publish/).
package/docs/STANDARD.md CHANGED
@@ -1,8 +1,10 @@
1
1
  # The Agent Employee Standard
2
2
 
3
+ Current standard: **1.5, 2026-10-02**. The work-cycle section adds observable delivery, scoped blockers, bounded experiments, configured handoffs and atomic recovery across the existing routines.
4
+
3
5
  Build spec for every AI Employee in the club. Not shipped to members. The GTM Engineer is the reference implementation; every later Employee inherits everything here and adds only its own domain expertise.
4
6
 
5
- Standard version 1.3, 2026-09-11: LAW 4 names connected sources, the per capability routes a member connects in their own harness, read only and preferred over the browser lane. Version 1.2, 2026-09-05: LAW 2 became the two guardrails, the first of them released channel by channel by the member in `RELEASES.md`. Version 1.1, 2026-08-28. Laws 6 through 8 and the operator-session and browser-lane sections were earned in the first live week of the GTM Engineer running Mark's own launch; the release notes in each kit's CHANGELOG carry the short story.
7
+ Standard version 1.4, 2026-09-23: LAW 7 gains a third delivery surface, `brief.deliver`, so the brief reaches the member on a harness whose computer they never open; section 4 gains the five questions a routine answers before it is scheduled. Version 1.3, 2026-09-11: LAW 4 names connected sources, the per capability routes a member connects in their own harness, read only and preferred over the browser lane. Version 1.2, 2026-09-05: LAW 2 became the two guardrails, the first of them released channel by channel by the member in `RELEASES.md`. Version 1.1, 2026-08-28. Laws 6 through 8 and the operator-session and browser-lane sections were earned in the first live week of the GTM Engineer running Mark's own launch; the release notes in each kit's CHANGELOG carry the short story.
6
8
 
7
9
  Derived from Mark's own production routines rather than invented: the push mechanics come from `night-shift-brief` and `morning-clicks-block`, the window and period guards from the same, the browser craft from roughly thirty live Chrome routines.
8
10
 
@@ -35,6 +37,7 @@ An Employee is judged on one question: **after ninety days of running unattended
35
37
  **LAW 7: The Employee brings the work to the member.** Work product that only exists as a file the member must go hunting for reads as no work at all. Two delivery surfaces, used wherever the role allows:
36
38
  - *The dashboard is a view over live state, never installed prose.* The build bakes the morning artifact, queue files, digests, and the run history straight from the working files; every routine that writes work product rebuilds the dashboard before writing its run record; a tab describing work renders from the board, not from text written at install, which rots the same week.
37
39
  - *The browser is a delivery surface.* Where the role touches the world through forms, drafts, or posts, the Employee fills the form and leaves the tab open, prepares the draft inside the member's own account in draft state, and stages the post ready to publish. The member's contribution shrinks to the one click a held guardrail reserves for them. Every browser-staged deliverable also lands in a durable queue file carrying the full text of every field, so a closed tab loses nothing. Anti-bot checks are never answered; they are left for the member with the submit.
40
+ - *The brief is delivered, not filed.* Every standup resolves `brief.deliver` after it writes `brief-latest.md`: the dashboard on a machine the member uses, the Employee's own thread on a hosted harness whose computer the member never opens, the member's own address where a mail route exists. The delivered text is the file's text with nothing added. A brief to the member's own thread or address is delivery, not a send, and needs no release; absent every route the file is the brief and the run record says `brief: file only`. (v1.4, 2026-09-23, from the first hosted harness: a brief on a cloud computer nobody opens is no brief.)
38
41
  (Earned 2026-08-28, Mark: "They should bring it to me and bring it to my attention," and the same morning three directory submissions went live within minutes of forms being staged in his browser.)
39
42
 
40
43
  **LAW 8: A tick records consent; the routine performs the move.** When a member ticks a card whose definition of done implies a file change, the next routine to read that tick completes the mechanical part itself. A confirmed proof inventory whose lines never got moved is a day of thin drafts nobody wanted. Consent is the member's; labor is the Employee's.
@@ -151,6 +154,8 @@ scripts/ small deterministic helpers with a --selftest
151
154
 
152
155
  **Two routines never drive one browser.** A mutex with a dead-holder timeout, and a routine that never took the lock never deletes it.
153
156
 
157
+ **Five questions before any routine exists**, and every shipped routine answers them in its `SCHEDULE.md` row, its brief lines and its push rules: **trigger** (a clock time, or an event such as new mail or a new row), **frequency** (how often the underlying thing actually changes, never how often it would be nice to look), **output** (where the result lands and who reads it), **silence** (what it does when there is nothing to report, which is nothing) and **stop** (the one condition that pages the member instead of waiting for the brief). A routine that cannot answer all five is not ready to schedule, and a routine whose honest answer to frequency is "rarely" gets a weekly row, not a patrol. (v1.4, 2026-09-23.)
158
+
154
159
  ---
155
160
 
156
161
  ## 5. Voice and copy
@@ -192,3 +197,9 @@ Mark's own production operations are the raw material. Each maps to an Employee:
192
197
  | Mailbox searched, audited, analysed, monitored for outreach | Inbox Operator |
193
198
 
194
199
  **Mine the real routines before writing any kit.** The pace numbers, the idempotency mechanisms, the login-wall handling, and the verify-after-acting discipline are all already proven in production and must not be re-invented from theory.
200
+
201
+ ## Work cycle, standard 1.5
202
+
203
+ Version 1.5, 2026-10-02, makes useful progress independently observable across all roles. Shared source lives in `shared/work-cycle/` and is copied to every kit with `node .github/scripts/sync-work-cycle.mjs --write`; CI verifies exact parity. Each self-contained kit adds WORK-CYCLE.md, work-profile.json and tested helpers. Existing routines own preparation, research, experiments, handoffs and recovery within their existing releases. Execution health, delivery progress and business results remain separate.
204
+
205
+ Progress records never substitute for acceptance checks. Scenarios test quiet monitoring, repeated empty success, pending approval with independent preparation, insufficient evidence, configured handoffs and interrupted actions. Human/agent evaluations must inspect resulting artifacts and restraint; deterministic checks do not prove the quality of future model decisions.
package/docs/UPGRADING.md CHANGED
@@ -63,9 +63,23 @@ The last number is the one to look at. It is the work the employee has done sinc
63
63
 
64
64
  ## Merging a `.new` file
65
65
 
66
- There is no clever tool for this and there should not be. Open the two side by side, and carry your edit forward into the new version rather than carrying the new version back into your file. The kit changelog tells you what changed and why, so you usually only need to find your own edit and re-apply it.
66
+ Version 1.9 adds a conservative three-way comparison. It prepares candidates and a JSON report without changing live kit files:
67
67
 
68
- When you are done, delete the `.new` file. Nothing reads it.
68
+ ```bash
69
+ npx ai-employees@latest reconcile gtm-engineer --to /path/to/your/employee
70
+ ```
71
+
72
+ Read the report under `.upgrade/reconcile/`. Nonoverlapping edits can be combined; overlapping edits are marked as conflicts. Member files, schedules and unknown paths are protected. The member's `## Corrections` must survive unchanged. Then apply the nonconflicting candidates explicitly:
73
+
74
+ ```bash
75
+ npx ai-employees@latest reconcile gtm-engineer --to /path/to/your/employee --apply
76
+ ```
77
+
78
+ Backups and incoming copies stay beside the report. The command checks JSON, routine frontmatter and JavaScript syntax. Run the kit scripts' self tests and review changed instructions before the next scheduled execution. Stop a running employee through its normal member controls while applying an upgrade; the installer does not interrupt it or change its schedule.
79
+
80
+ New installs retain shipped baselines in `.upgrade/baseline/`. Older receipts have hashes only: supply `--base /path/to/original/kit` when you have that exact version. Each base file must match its recorded hash. Without verified base contents, the report names manual reconciliation instead of guessing. Carry member edits into the new version, preserving releases and schedule values. Repeated upgrades never reclassify a pending local edit as an unmodified shipped file.
81
+
82
+ For this release, reconcile CONTRACT.md, the affected SKILL.md files and the new helper scripts as one compatible kit before resuming. Add the work-cycle settings from SCHEDULE.md.new while preserving your own rows. No new routine needs registering. A new VERSION file alone does not prove every local instruction has adopted the update.
69
83
 
70
84
  ## Your own repairs are worth sending back
71
85
 
@@ -3,8 +3,10 @@
3
3
  - **They get better at your business every run.** When a page moves, a button changes or a step now needs a scroll, the routine fixes its own instructions in the run that hit it, keeps the text it replaced as the undo, and tells you in the next morning's brief under "What changed about me".
4
4
  - **They drive your browser and your PC the way you do.** Every browser routine works from a technique library written from real runs on my own machine, not from documentation: click what the page actually shows, read the page back to verify, leave a filled form open on the last step. Each site's flow is learned on your machine the first time a routine needs it and repaired every time after. Two routines never use the same signed in account at once, and on LinkedIn they read and never click.
5
5
  - **One push to your phone, only when you are the blocker.** A login expired, a credential is missing, your conversion tracking stopped while ads are live, or a run died holding the browser. Four cases and no fifth. One short line, never twice for the same thing, never outside your working hours, never on a first run. Everything else waits for the brief. [The one push](STANDARD.md#23-the-one-push).
6
+ - **The brief comes to you.** After the standup writes it, the routine brings it to wherever you are: the dashboard on your own machine, the Employee's own thread on a hosted agent such as Grok Bot, or your own inbox where a mail route exists. A brief on a computer you never open is no brief, so on a hosted harness it lands in your thread and the Grok app carries it to your phone.
6
7
  - **They use the connections you already have.** Where your agent already has a connector for an account, the routine reads through it instead of the screen: Meta's own Ads MCP server for an ad account, the Gmail connector for drafts and replies, the Vercel and Supabase connectors for logs and advisors, Metricool for publishing. Each kit names its connections in one table, uses them read only, and works without any of them; nothing is installed for you.
7
- - **They run on the agent you already use.** Claude Code, OpenClaw, Hermes, OpenCode, Grok Bot, Codex, Antigravity, Pi, Cline, Qwen Code and DeepSeek, without a routine changing by one word. Routines describe what they need done, and one file per kit says how each agent does it.
8
+ - **They run on the agent you already use.** Claude Code, OpenClaw, Hermes, OpenCode, Grok Bot, Codex, Antigravity, Pi, Cline, Qwen Code, DeepSeek, Muse and Dots, without a routine changing by one word. Routines describe what they need done, and one file per kit says how each agent does it.
9
+ - **One routine, one job, one output.** Routines run solo and never supervise each other; the Chief of Staff reads their run logs afterwards and names what quietly stopped. Crews of agents that hand work to each other unsupervised have been measured multiplying their own errors many times over, and a wrong step here lands in your brief instead of in three other routines.
8
10
  - **You can direct any of them in chat.** Open a session in the employee's folder and it does anything you could do by hand, on your word: tick a card you confirmed, stage a form now, retune a strategy file, correct a stale brief. It leaves the same trail a routine would, and the scheduled runs treat that work as yours.
9
11
  - **You set how far they go.** Every employee drafts, fills and stages by default, and the last click is yours. Release a channel in `RELEASES.md` and the routine completes that action itself from then on. Your agent's own permission settings are the gate, and every file in the kit is plain text in your own folder, yours to change.
10
12
  - **Upgrades never overwrite your work.** `npx ai-employees upgrade` reports first, leaves any file you edited alone, and never reads your strategy or your ledgers. Every kit follows the published [Agent Employee Standard](STANDARD.md), and `npx ai-employees contribute` turns the fixes a kit made to itself into a report you can send upstream.
@@ -1,8 +1,12 @@
1
1
  # AGENTS.md
2
2
 
3
- **Ad Manager**, kit version 1.2.0. One of the eight AI Employees from [github.com/markfulton/ai-employees](https://github.com/markfulton/ai-employees).
3
+ **Ad Manager**, one of the eight AI Employees from [github.com/markfulton/ai-employees](https://github.com/markfulton/ai-employees).
4
4
 
5
- This file follows the [AGENTS.md](https://agents.md) convention so that any harness can pick this kit up without being told how. It is a map, not the instructions. **The instructions are the files it points at, and they are authoritative over anything summarised here.**
5
+ This file follows the [AGENTS.md](https://agents.md) convention so that any harness can pick this kit up without being told how. It is a map, not the instructions. **The instructions are the files it points at, and they are authoritative over anything summarised here.** The version this kit ships as is in `VERSION` at the root.
6
+
7
+ ## If your harness has an Instructions field instead of a folder
8
+
9
+ On a hosted agent such as Grok Bot, the kit lives on the bot's own cloud computer and the bot reads an Instructions field before every task. Paste this whole file into that field, with the folder's absolute path on that computer in place of `«ADS_ROOT»`. It is the base layer (the two guardrails and the files that outrank everything else), the role layer (`ROLE.md`, which names the evidence standard) and the current focus (the strategy folder the routines keep current), and it is short enough to fit. Everything else stays in the files.
6
10
 
7
11
  ## If you have been asked to install this Employee
8
12
 
@@ -35,7 +39,7 @@ Then run the guard before the work: `node scripts/guard.mjs`. It checks the day,
35
39
 
36
40
  It drafts, fills, stages and leaves the last click to the person who hired it, unless that person released the channel in `RELEASES.md` at the kit root, in which case the routine that stages the channel completes the action and records it. If a task seems to require sending or spending on a channel that is not released, that is a signal to stop and write a blocker, not to proceed.
37
41
 
38
- It also never edits its own `SCHEDULE.md` row, never widens its own budget, and never invents a number. Every figure it publishes carries the file or screen it was read from and the date it was read.
42
+ Schedule repairs follow CONTRACT.md and never widen authority or silently increase the member's operating budget. It never invents a number. Every figure it publishes carries the file or screen it was read from and the date it was read.
39
43
 
40
44
  ## Files it owns and files it must not touch
41
45
 
@@ -44,3 +48,7 @@ It also never edits its own `SCHEDULE.md` row, never widens its own budget, and
44
48
  ## If you are working on this kit as source code
45
49
 
46
50
  Read the `AGENTS.md` at the root of the repository instead. It carries the contribution rules, the checks that must pass, and the no dashes rule that CI enforces.
51
+
52
+ ## Proactive execution
53
+
54
+ After an eligible guard result, read `WORK-CYCLE.md` and `work-profile.json`. These are contract extensions for useful work, evidence, scoped blockers and recovery. All existing routine ids and schedules remain in use.