contextos-agents 1.6.1 → 2.0.0-beta.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (323) hide show
  1. package/.agents/AGENTS.md +6 -1
  2. package/.agents/adapters/aider/export.js +117 -97
  3. package/.agents/adapters/claude/export.js +68 -26
  4. package/.agents/adapters/copilot/export.js +90 -51
  5. package/.agents/adapters/cursor/export.js +83 -68
  6. package/.agents/adapters/drift-detector.js +196 -0
  7. package/.agents/adapters/gemini/export.js +76 -45
  8. package/.agents/adapters/pure-compiler.js +443 -0
  9. package/.agents/adapters/zed/export.js +109 -62
  10. package/.agents/compiled/registry.v2.json +504 -0
  11. package/.agents/compiled/registry.v2.sha256 +1 -0
  12. package/.agents/compiler/manifest-compiler.js +963 -0
  13. package/.agents/compiler/vendor/yaml.LICENSE.txt +13 -0
  14. package/.agents/compiler/vendor/yaml.SBOM.json +6 -0
  15. package/.agents/compiler/vendor/yaml.js +139 -0
  16. package/.agents/core/profiles/init.yaml +25 -0
  17. package/.agents/core/skills/context-manager/references/context-rules.md +59 -0
  18. package/.agents/core/skills/context-manager/skill.yaml +10 -5
  19. package/.agents/core/skills/context-os/SKILL.md +3 -6
  20. package/.agents/core/skills/context-os/skill.yaml +14 -8
  21. package/.agents/core/skills/engineering-workflow/SKILL.md +1 -1
  22. package/.agents/core/skills/engineering-workflow/skill.yaml +7 -7
  23. package/.agents/core/skills/gemini-precision/SKILL.md +4 -0
  24. package/.agents/core/skills/gemini-precision/skill.yaml +5 -6
  25. package/.agents/core/skills/gstack-roles/SKILL.md +3 -1
  26. package/.agents/core/skills/gstack-roles/skill.yaml +9 -6
  27. package/.agents/core/skills/ponytail-mindset/skill.yaml +7 -7
  28. package/.agents/core/skills/security/skill.yaml +21 -2
  29. package/.agents/ctx.js +587 -111
  30. package/.agents/customization-dx.js +282 -0
  31. package/.agents/doctor.js +877 -33
  32. package/.agents/filesystem/index.js +71 -0
  33. package/.agents/filesystem/journaled-transaction.js +451 -0
  34. package/.agents/filesystem/lockfile-v2.js +275 -0
  35. package/.agents/filesystem/platform-hardening.js +222 -0
  36. package/.agents/filesystem/project-lock.js +218 -0
  37. package/.agents/filesystem/safe-path.js +256 -0
  38. package/.agents/generated/claude/skills/context-os/SKILL.md +1 -1
  39. package/.agents/generated/claude/skills/engineering-workflow/SKILL.md +1 -1
  40. package/.agents/generated/claude/skills/gemini-precision/SKILL.md +4 -0
  41. package/.agents/generated/claude/skills/gstack-roles/SKILL.md +3 -1
  42. package/.agents/generated/gemini/skills/context-os/SKILL.md +2 -2
  43. package/.agents/generated/gemini/skills/engineering-workflow/SKILL.md +1 -1
  44. package/.agents/generated/gemini/skills/gemini-precision/SKILL.md +4 -0
  45. package/.agents/generated/gemini/skills/gstack-roles/SKILL.md +3 -1
  46. package/.agents/plugins/contextos/hooks.json +25 -0
  47. package/.agents/plugins/contextos/plugin.json +19 -0
  48. package/.agents/plugins.js +432 -73
  49. package/.agents/profiles.js +507 -46
  50. package/.agents/resolver.js +50 -414
  51. package/.agents/schemas/attestation.review.v1.json +111 -0
  52. package/.agents/schemas/attestation.verification.v1.json +85 -0
  53. package/.agents/schemas/lockfile.v2.schema.json +134 -0
  54. package/.agents/schemas/profile.v2.schema.json +114 -0
  55. package/.agents/schemas/runtime.thread.v1.json +192 -0
  56. package/.agents/schemas/skill.manifest.v2.json +177 -0
  57. package/.agents/schemas/verification.spec.v1.json +39 -0
  58. package/.agents/schemas/workspace.graph.schema.json +106 -0
  59. package/.agents/stats.js +22 -9
  60. package/.agents/transaction-core/event-store.js +288 -0
  61. package/.agents/transaction-core/idempotency.js +129 -0
  62. package/.agents/transaction-core/ipc-lock.js +311 -0
  63. package/.agents/transaction-core/plugin-supply-chain-bundle.js +436 -0
  64. package/.agents/validate.js +143 -14
  65. package/.agents/watch.js +354 -102
  66. package/.agents/workspace/workspace-graph.js +778 -0
  67. package/README.md +59 -387
  68. package/benchmarks/v2/analysis/statistics.js +140 -0
  69. package/benchmarks/v2/analysis/stats.js +69 -0
  70. package/benchmarks/v2/arms/arm-definitions.js +79 -0
  71. package/benchmarks/v2/dataset.schema.json +34 -0
  72. package/benchmarks/v2/evaluators/index.js +25 -0
  73. package/benchmarks/v2/evaluators/verified-success.js +116 -0
  74. package/benchmarks/v2/harness/runner.js +88 -0
  75. package/benchmarks/v2/pilot-tasks.json +392 -0
  76. package/bin/commands/recover.js +88 -0
  77. package/bin/commands/uninstall.js +207 -0
  78. package/bin/commands/update.js +325 -0
  79. package/bin/commands.js +342 -0
  80. package/bin/index.js +326 -149
  81. package/bin/lib/detector.js +106 -0
  82. package/bin/lib/lockfile.js +253 -0
  83. package/bin/lib/safe-writer.js +290 -0
  84. package/package.json +85 -73
  85. package/registry.json +15 -7
  86. package/registry.schema.json +3 -1
  87. package/registry.v2.schema.json +86 -0
  88. package/.agents/core/profiles/backend.yaml +0 -47
  89. package/.agents/core/profiles/enterprise.yaml +0 -46
  90. package/.agents/core/profiles/frontend.yaml +0 -46
  91. package/.agents/core/profiles/hackathon.yaml +0 -45
  92. package/.agents/core/profiles/mvp.yaml +0 -44
  93. package/.agents/core/profiles/startup.yaml +0 -48
  94. package/.agents/core/skills/adapters/EXAMPLES.md +0 -19
  95. package/.agents/core/skills/adapters/SKILL.md +0 -105
  96. package/.agents/core/skills/adapters/TROUBLESHOOTING.md +0 -7
  97. package/.agents/core/skills/adapters/VALIDATION.json +0 -12
  98. package/.agents/core/skills/adapters/skill.yaml +0 -10
  99. package/.agents/core/skills/architecture-diagrams/SKILL.md +0 -108
  100. package/.agents/core/skills/architecture-diagrams/VALIDATION.json +0 -12
  101. package/.agents/core/skills/architecture-diagrams/skill.yaml +0 -8
  102. package/.agents/core/skills/brutalist-design/SKILL.md +0 -150
  103. package/.agents/core/skills/brutalist-design/VALIDATION.json +0 -12
  104. package/.agents/core/skills/brutalist-design/skill.yaml +0 -8
  105. package/.agents/core/skills/database/EXAMPLES.md +0 -74
  106. package/.agents/core/skills/database/SKILL.md +0 -101
  107. package/.agents/core/skills/database/TROUBLESHOOTING.md +0 -18
  108. package/.agents/core/skills/database/VALIDATION.json +0 -11
  109. package/.agents/core/skills/database/skill.yaml +0 -25
  110. package/.agents/core/skills/ddd/EXAMPLES.md +0 -42
  111. package/.agents/core/skills/ddd/SKILL.md +0 -247
  112. package/.agents/core/skills/ddd/TROUBLESHOOTING.md +0 -19
  113. package/.agents/core/skills/ddd/VALIDATION.json +0 -12
  114. package/.agents/core/skills/ddd/ddd.md +0 -178
  115. package/.agents/core/skills/ddd/skill.yaml +0 -10
  116. package/.agents/core/skills/decisions/EXAMPLES.md +0 -35
  117. package/.agents/core/skills/decisions/SKILL.md +0 -90
  118. package/.agents/core/skills/decisions/TROUBLESHOOTING.md +0 -13
  119. package/.agents/core/skills/decisions/VALIDATION.json +0 -12
  120. package/.agents/core/skills/decisions/skill.yaml +0 -10
  121. package/.agents/core/skills/docker/EXAMPLES.md +0 -56
  122. package/.agents/core/skills/docker/SKILL.md +0 -63
  123. package/.agents/core/skills/docker/TROUBLESHOOTING.md +0 -18
  124. package/.agents/core/skills/docker/VALIDATION.json +0 -11
  125. package/.agents/core/skills/docker/skill.yaml +0 -23
  126. package/.agents/core/skills/fastapi/EXAMPLES.md +0 -36
  127. package/.agents/core/skills/fastapi/SKILL.md +0 -148
  128. package/.agents/core/skills/fastapi/TROUBLESHOOTING.md +0 -19
  129. package/.agents/core/skills/fastapi/VALIDATION.json +0 -12
  130. package/.agents/core/skills/fastapi/fastapi.md +0 -112
  131. package/.agents/core/skills/fastapi/skill.yaml +0 -10
  132. package/.agents/core/skills/generators/EXAMPLES.md +0 -19
  133. package/.agents/core/skills/generators/SKILL.md +0 -112
  134. package/.agents/core/skills/generators/TROUBLESHOOTING.md +0 -7
  135. package/.agents/core/skills/generators/VALIDATION.json +0 -12
  136. package/.agents/core/skills/generators/skill.yaml +0 -10
  137. package/.agents/core/skills/generators/templates/API.md +0 -77
  138. package/.agents/core/skills/generators/templates/ARCHITECTURE.md +0 -70
  139. package/.agents/core/skills/generators/templates/DATABASE.md +0 -42
  140. package/.agents/core/skills/generators/templates/DECISION.md +0 -46
  141. package/.agents/core/skills/generators/templates/PRD.md +0 -67
  142. package/.agents/core/skills/generators/templates/PROJECT_GRAPH.md +0 -56
  143. package/.agents/core/skills/generators/templates/ROADMAP.md +0 -51
  144. package/.agents/core/skills/generators/templates/TASKS.md +0 -43
  145. package/.agents/core/skills/generators/templates/UI.md +0 -73
  146. package/.agents/core/skills/graphify/EXAMPLES.md +0 -73
  147. package/.agents/core/skills/graphify/SKILL.md +0 -130
  148. package/.agents/core/skills/graphify/VALIDATION.json +0 -12
  149. package/.agents/core/skills/graphify/skill.yaml +0 -13
  150. package/.agents/core/skills/impeccable-design/EXAMPLES.md +0 -26
  151. package/.agents/core/skills/impeccable-design/SKILL.md +0 -201
  152. package/.agents/core/skills/impeccable-design/TROUBLESHOOTING.md +0 -19
  153. package/.agents/core/skills/impeccable-design/VALIDATION.json +0 -12
  154. package/.agents/core/skills/impeccable-design/skill.yaml +0 -14
  155. package/.agents/core/skills/interview-me/SKILL.md +0 -97
  156. package/.agents/core/skills/interview-me/VALIDATION.json +0 -12
  157. package/.agents/core/skills/interview-me/skill.yaml +0 -8
  158. package/.agents/core/skills/microservices/EXAMPLES.md +0 -38
  159. package/.agents/core/skills/microservices/SKILL.md +0 -164
  160. package/.agents/core/skills/microservices/TROUBLESHOOTING.md +0 -19
  161. package/.agents/core/skills/microservices/VALIDATION.json +0 -12
  162. package/.agents/core/skills/microservices/microservices.md +0 -119
  163. package/.agents/core/skills/microservices/skill.yaml +0 -10
  164. package/.agents/core/skills/minimalist-design/SKILL.md +0 -113
  165. package/.agents/core/skills/minimalist-design/VALIDATION.json +0 -12
  166. package/.agents/core/skills/minimalist-design/skill.yaml +0 -8
  167. package/.agents/core/skills/nestjs/EXAMPLES.md +0 -40
  168. package/.agents/core/skills/nestjs/SKILL.md +0 -139
  169. package/.agents/core/skills/nestjs/TROUBLESHOOTING.md +0 -19
  170. package/.agents/core/skills/nestjs/VALIDATION.json +0 -12
  171. package/.agents/core/skills/nestjs/nestjs.md +0 -103
  172. package/.agents/core/skills/nestjs/skill.yaml +0 -10
  173. package/.agents/core/skills/nextjs/EXAMPLES.md +0 -40
  174. package/.agents/core/skills/nextjs/SKILL.md +0 -163
  175. package/.agents/core/skills/nextjs/TROUBLESHOOTING.md +0 -19
  176. package/.agents/core/skills/nextjs/VALIDATION.json +0 -12
  177. package/.agents/core/skills/nextjs/nextjs.md +0 -67
  178. package/.agents/core/skills/nextjs/skill.yaml +0 -10
  179. package/.agents/core/skills/node/EXAMPLES.md +0 -80
  180. package/.agents/core/skills/node/SKILL.md +0 -128
  181. package/.agents/core/skills/node/TROUBLESHOOTING.md +0 -19
  182. package/.agents/core/skills/node/VALIDATION.json +0 -12
  183. package/.agents/core/skills/node/node.md +0 -87
  184. package/.agents/core/skills/node/skill.yaml +0 -10
  185. package/.agents/core/skills/performance/EXAMPLES.md +0 -30
  186. package/.agents/core/skills/performance/SKILL.md +0 -75
  187. package/.agents/core/skills/performance/TROUBLESHOOTING.md +0 -19
  188. package/.agents/core/skills/performance/VALIDATION.json +0 -12
  189. package/.agents/core/skills/performance/performance.md +0 -52
  190. package/.agents/core/skills/performance/skill.yaml +0 -10
  191. package/.agents/core/skills/react/EXAMPLES.md +0 -79
  192. package/.agents/core/skills/react/SKILL.md +0 -132
  193. package/.agents/core/skills/react/TROUBLESHOOTING.md +0 -19
  194. package/.agents/core/skills/react/VALIDATION.json +0 -12
  195. package/.agents/core/skills/react/react.md +0 -93
  196. package/.agents/core/skills/react/skill.yaml +0 -10
  197. package/.agents/core/skills/react-best-practices/SKILL.md +0 -155
  198. package/.agents/core/skills/react-best-practices/VALIDATION.json +0 -12
  199. package/.agents/core/skills/react-best-practices/skill.yaml +0 -10
  200. package/.agents/core/skills/redesign-audit/SKILL.md +0 -117
  201. package/.agents/core/skills/redesign-audit/VALIDATION.json +0 -12
  202. package/.agents/core/skills/redesign-audit/skill.yaml +0 -8
  203. package/.agents/core/skills/soft-design/SKILL.md +0 -108
  204. package/.agents/core/skills/soft-design/VALIDATION.json +0 -12
  205. package/.agents/core/skills/soft-design/skill.yaml +0 -8
  206. package/.agents/core/skills/state-management/EXAMPLES.md +0 -56
  207. package/.agents/core/skills/state-management/SKILL.md +0 -48
  208. package/.agents/core/skills/state-management/TROUBLESHOOTING.md +0 -18
  209. package/.agents/core/skills/state-management/VALIDATION.json +0 -11
  210. package/.agents/core/skills/state-management/skill.yaml +0 -22
  211. package/.agents/core/skills/subagent-orchestrator/SKILL.md +0 -100
  212. package/.agents/core/skills/subagent-orchestrator/VALIDATION.json +0 -12
  213. package/.agents/core/skills/subagent-orchestrator/skill.yaml +0 -8
  214. package/.agents/core/skills/system-design/EXAMPLES.md +0 -75
  215. package/.agents/core/skills/system-design/SKILL.md +0 -419
  216. package/.agents/core/skills/system-design/TROUBLESHOOTING.md +0 -19
  217. package/.agents/core/skills/system-design/VALIDATION.json +0 -12
  218. package/.agents/core/skills/system-design/skill.yaml +0 -13
  219. package/.agents/core/skills/system-design/system-design.md +0 -112
  220. package/.agents/core/skills/testing/EXAMPLES.md +0 -71
  221. package/.agents/core/skills/testing/SKILL.md +0 -70
  222. package/.agents/core/skills/testing/TROUBLESHOOTING.md +0 -18
  223. package/.agents/core/skills/testing/VALIDATION.json +0 -11
  224. package/.agents/core/skills/testing/skill.yaml +0 -26
  225. package/.agents/core/skills/typescript/EXAMPLES.md +0 -64
  226. package/.agents/core/skills/typescript/SKILL.md +0 -112
  227. package/.agents/core/skills/typescript/TROUBLESHOOTING.md +0 -19
  228. package/.agents/core/skills/typescript/VALIDATION.json +0 -12
  229. package/.agents/core/skills/typescript/skill.yaml +0 -10
  230. package/.agents/core/skills/typescript/typescript.md +0 -71
  231. package/.agents/core/skills/ui-design/EXAMPLES.md +0 -21
  232. package/.agents/core/skills/ui-design/SKILL.md +0 -124
  233. package/.agents/core/skills/ui-design/TROUBLESHOOTING.md +0 -19
  234. package/.agents/core/skills/ui-design/VALIDATION.json +0 -12
  235. package/.agents/core/skills/ui-design/skill.yaml +0 -10
  236. package/.agents/core/skills/ui-design/ui.md +0 -88
  237. package/.agents/core/skills/ui-ux-pro/EXAMPLES.md +0 -62
  238. package/.agents/core/skills/ui-ux-pro/SKILL.md +0 -375
  239. package/.agents/core/skills/ui-ux-pro/TROUBLESHOOTING.md +0 -19
  240. package/.agents/core/skills/ui-ux-pro/VALIDATION.json +0 -12
  241. package/.agents/core/skills/ui-ux-pro/skill.yaml +0 -13
  242. package/.agents/core/skills/ux-design/EXAMPLES.md +0 -36
  243. package/.agents/core/skills/ux-design/SKILL.md +0 -116
  244. package/.agents/core/skills/ux-design/TROUBLESHOOTING.md +0 -19
  245. package/.agents/core/skills/ux-design/VALIDATION.json +0 -12
  246. package/.agents/core/skills/ux-design/skill.yaml +0 -10
  247. package/.agents/core/skills/ux-design/ux.md +0 -80
  248. package/.agents/core/skills/vercel-optimize/SKILL.md +0 -83
  249. package/.agents/core/skills/vercel-optimize/VALIDATION.json +0 -12
  250. package/.agents/core/skills/vercel-optimize/skill.yaml +0 -10
  251. package/.agents/core/skills/web-accessibility/EXAMPLES.md +0 -39
  252. package/.agents/core/skills/web-accessibility/SKILL.md +0 -170
  253. package/.agents/core/skills/web-accessibility/TROUBLESHOOTING.md +0 -19
  254. package/.agents/core/skills/web-accessibility/VALIDATION.json +0 -12
  255. package/.agents/core/skills/web-accessibility/accessibility.md +0 -63
  256. package/.agents/core/skills/web-accessibility/skill.yaml +0 -10
  257. package/.agents/generated/claude/skills/adapters/SKILL.md +0 -126
  258. package/.agents/generated/claude/skills/architecture-diagrams/SKILL.md +0 -101
  259. package/.agents/generated/claude/skills/brutalist-design/SKILL.md +0 -145
  260. package/.agents/generated/claude/skills/database/SKILL.md +0 -191
  261. package/.agents/generated/claude/skills/ddd/SKILL.md +0 -305
  262. package/.agents/generated/claude/skills/decisions/SKILL.md +0 -134
  263. package/.agents/generated/claude/skills/docker/SKILL.md +0 -135
  264. package/.agents/generated/claude/skills/fastapi/SKILL.md +0 -200
  265. package/.agents/generated/claude/skills/generators/SKILL.md +0 -133
  266. package/.agents/generated/claude/skills/graphify/SKILL.md +0 -198
  267. package/.agents/generated/claude/skills/impeccable-design/SKILL.md +0 -241
  268. package/.agents/generated/claude/skills/interview-me/SKILL.md +0 -90
  269. package/.agents/generated/claude/skills/microservices/SKILL.md +0 -218
  270. package/.agents/generated/claude/skills/minimalist-design/SKILL.md +0 -108
  271. package/.agents/generated/claude/skills/nestjs/SKILL.md +0 -195
  272. package/.agents/generated/claude/skills/nextjs/SKILL.md +0 -219
  273. package/.agents/generated/claude/skills/node/SKILL.md +0 -224
  274. package/.agents/generated/claude/skills/performance/SKILL.md +0 -121
  275. package/.agents/generated/claude/skills/react/SKILL.md +0 -227
  276. package/.agents/generated/claude/skills/react-best-practices/SKILL.md +0 -146
  277. package/.agents/generated/claude/skills/redesign-audit/SKILL.md +0 -112
  278. package/.agents/generated/claude/skills/soft-design/SKILL.md +0 -103
  279. package/.agents/generated/claude/skills/state-management/SKILL.md +0 -120
  280. package/.agents/generated/claude/skills/subagent-orchestrator/SKILL.md +0 -93
  281. package/.agents/generated/claude/skills/system-design/SKILL.md +0 -507
  282. package/.agents/generated/claude/skills/testing/SKILL.md +0 -157
  283. package/.agents/generated/claude/skills/typescript/SKILL.md +0 -192
  284. package/.agents/generated/claude/skills/ui-design/SKILL.md +0 -161
  285. package/.agents/generated/claude/skills/ui-ux-pro/SKILL.md +0 -451
  286. package/.agents/generated/claude/skills/ux-design/SKILL.md +0 -168
  287. package/.agents/generated/claude/skills/vercel-optimize/SKILL.md +0 -76
  288. package/.agents/generated/claude/skills/web-accessibility/SKILL.md +0 -225
  289. package/.agents/generated/gemini/skills/adapters/SKILL.md +0 -135
  290. package/.agents/generated/gemini/skills/architecture-diagrams/SKILL.md +0 -107
  291. package/.agents/generated/gemini/skills/brutalist-design/SKILL.md +0 -151
  292. package/.agents/generated/gemini/skills/database/SKILL.md +0 -200
  293. package/.agents/generated/gemini/skills/ddd/SKILL.md +0 -314
  294. package/.agents/generated/gemini/skills/decisions/SKILL.md +0 -143
  295. package/.agents/generated/gemini/skills/docker/SKILL.md +0 -144
  296. package/.agents/generated/gemini/skills/fastapi/SKILL.md +0 -209
  297. package/.agents/generated/gemini/skills/generators/SKILL.md +0 -142
  298. package/.agents/generated/gemini/skills/graphify/SKILL.md +0 -205
  299. package/.agents/generated/gemini/skills/impeccable-design/SKILL.md +0 -250
  300. package/.agents/generated/gemini/skills/interview-me/SKILL.md +0 -96
  301. package/.agents/generated/gemini/skills/microservices/SKILL.md +0 -227
  302. package/.agents/generated/gemini/skills/minimalist-design/SKILL.md +0 -114
  303. package/.agents/generated/gemini/skills/nestjs/SKILL.md +0 -204
  304. package/.agents/generated/gemini/skills/nextjs/SKILL.md +0 -298
  305. package/.agents/generated/gemini/skills/node/SKILL.md +0 -323
  306. package/.agents/generated/gemini/skills/performance/SKILL.md +0 -185
  307. package/.agents/generated/gemini/skills/react/SKILL.md +0 -332
  308. package/.agents/generated/gemini/skills/react-best-practices/SKILL.md +0 -152
  309. package/.agents/generated/gemini/skills/redesign-audit/SKILL.md +0 -118
  310. package/.agents/generated/gemini/skills/soft-design/SKILL.md +0 -109
  311. package/.agents/generated/gemini/skills/state-management/SKILL.md +0 -129
  312. package/.agents/generated/gemini/skills/subagent-orchestrator/SKILL.md +0 -99
  313. package/.agents/generated/gemini/skills/system-design/SKILL.md +0 -631
  314. package/.agents/generated/gemini/skills/testing/SKILL.md +0 -166
  315. package/.agents/generated/gemini/skills/typescript/SKILL.md +0 -275
  316. package/.agents/generated/gemini/skills/ui-design/SKILL.md +0 -170
  317. package/.agents/generated/gemini/skills/ui-ux-pro/SKILL.md +0 -460
  318. package/.agents/generated/gemini/skills/ux-design/SKILL.md +0 -177
  319. package/.agents/generated/gemini/skills/vercel-optimize/SKILL.md +0 -82
  320. package/.agents/generated/gemini/skills/web-accessibility/SKILL.md +0 -300
  321. package/.agents/mcp/runtime.py +0 -470
  322. package/.agents/mcp/server.mjs +0 -189271
  323. package/benchmarks/gemini-issues.js +0 -533
package/README.md CHANGED
@@ -3,209 +3,94 @@
3
3
  [![npm version](https://img.shields.io/npm/v/contextos-agents.svg)](https://www.npmjs.com/package/contextos-agents)
4
4
  [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
5
5
  [![Node.js](https://img.shields.io/badge/node-%3E%3D18.0.0-brightgreen.svg)](https://nodejs.org/)
6
- [![Tests](https://img.shields.io/badge/tests-passing-brightgreen.svg)](#testing)
6
+ [![CI](https://github.com/kok-o/contextos-agents/actions/workflows/validate-skills.yml/badge.svg)](https://github.com/kok-o/contextos-agents/actions/workflows/validate-skills.yml)
7
7
 
8
- This is an open-source set of skills and behavioral rules for AI assistants. The package automatically installs an `.agents` folder into your project, teaching your AI assistant software development best practices (UI Design, Architecture, Security, and more).
8
+ **One source of truth for every coding agent.**
9
+
10
+ ContextOS is a deterministic context and policy compiler for AI coding agents. It transforms your team's version-controlled engineering rules into a minimal, verifiable context payload for Cursor, Claude Code, GitHub Copilot, Aider, and Zed—and detects configuration drift in CI.
9
11
 
10
12
  ## Installation
11
13
 
12
14
  You do not need to clone anything manually. Just open your terminal in the root of your project and run:
13
15
 
14
16
  ```bash
15
- npx contextos
16
- # or: npx contextos-agents
17
+ npx contextos-agents init
17
18
  ```
18
19
 
19
- The script will automatically detect your project tech stack, create the `.agents` folder, configure skills, and compile them for your AI agent.
20
+ The script will automatically detect your project tech stack, create the `.agents` folder, configure a neutral bootstrap profile, and compile it for your AI agent.
20
21
 
21
22
  ### Options
22
23
 
23
24
  ```bash
24
- npx contextos --help # Show all options
25
- npx contextos --version # Show version
26
- npx contextos --minimal # Install only 5 core skills (lightweight footprint)
27
- npx contextos --profile mvp # Install with specific profile (mvp, startup, enterprise, frontend, backend)
28
- npx contextos --auto # Auto-detect tech stack and apply recommended profile
29
- npx contextos --with-mcp # Install with MCP execution server enabled (.agents/mcp/)
30
- npx contextos setup-mcp # Add MCP server to an existing .agents/ project
31
- npx contextos --dry-run # Preview what will be installed
32
- npx contextos --force # Overwrite an existing .agents/ folder
33
- npx contextos --skip-compile # Skip auto-compilation step
25
+ npx contextos-agents --help # Show all options
26
+ npx contextos-agents --version # Show version
27
+ npx contextos-agents --minimal # Install only the core bootstrap skills
28
+ npx contextos-agents --profile init # Install with specific profile
29
+ npx contextos-agents --auto # Auto-detect tech stack and apply recommended profile
30
+ npx contextos-agents --dry-run # Preview what will be installed
31
+ npx contextos-agents --force # Overwrite an existing .agents/ folder
32
+ npx contextos-agents --skip-compile # Skip auto-compilation step
34
33
  ```
35
34
 
36
35
  ## Why ContextOS?
37
36
 
38
- Most AI coding assistants suffer from two extremes: they either operate in a vacuum with zero knowledge of your architectural standards, or they are choked with massive monolithic system prompts that cause context overflow, hallucinated dependencies, and lazy code stubs (`// TODO`).
39
-
40
- **ContextOS transforms chaotic AI code generation into a disciplined, senior-level software engineering team.**
41
-
42
- ### The Problem vs. The Solution
43
-
44
- | Without ContextOS (Everyday AI Frustrations) | With ContextOS (Engineering Discipline) |
45
- |---|---|
46
- | **Prompt Bloat & Token Waste:** Pasting giant system prompts burns tokens, slows responses, and degrades reasoning. | **Dynamic Context Resolution:** Dynamically resolves only 2–3 required skills per task (`ctx.js resolve`), saving up to 70–80% in prompt tokens. |
47
- | **Lazy Code & Slop:** Output full of `// TODO: implement later`, missing imports, and broken refactorings. | **Zero-Placeholder Invariant:** Strict guardrails enforce 100% complete, drop-in ready code with verified syntax and error boundaries. |
48
- | **Tool Zoo Fragmentation:** Inconsistent rules across Cursor (`.cursorrules`), Zed (`.zed/`), Aider, and Claude Code. | **Single Source of Truth:** Author skills once in markdown; ContextOS exports optimized configurations for all major AI editors (`ctx.js export all`). |
49
- | **Destructive File Rewrites:** Agents overwrite hundreds of lines without reading existing code first. | **Surgical Blast Radius & Sandboxing:** Changes are confined to planned lines or executed safely in isolated Git worktrees via ContextOS MCP. |
50
- | **"Black Box" Hallucinations:** You only see the start and end, with no insight into the agent's decisions. | **Transparent Pair Programming:** The agent outlines technical decisions, adheres to strict phases (DEFINE → PLAN → BUILD → VERIFY), and proves work with test runs. |
37
+ Most AI coding assistants suffer from two extremes: they either operate in a vacuum with zero knowledge of your architectural standards, or they are choked with massive monolithic system prompts that cause context overflow and lazy code stubs (`// TODO`).
51
38
 
52
- ### Key Developer Advantages
39
+ **ContextOS is not another coding agent.** It governs the context and policies used by the agents your team already has.
53
40
 
54
- - **Zero-Config Onboarding:** Run `npx contextos-agents` in your repository. It auto-detects your stack (React, Node, Python, etc.) and sets up the ideal profile in seconds.
55
- - **Tailored Project Profiles:** Use `mvp` for lean, rapid prototyping without bloated microservices boilerplate, or `enterprise` for strict TDD, DDD, and security auditing.
56
- - **Autonomous Multi-Agent Worktrees:** Run parallel tasks safely with the bundled MCP server—subagents work in isolated Git worktrees without corrupting your active workspace.
57
- - **Verifiable Benchmarks:** Backed by reproducible side-by-side benchmarks demonstrating measurable code quality improvements and reduced token usage.
41
+ ### The Three Pillars
58
42
 
59
- ## Project Profiles & Stack Auto-Detection
43
+ 1. **Portable:** Define your engineering rules once. ContextOS exports optimized configurations for all major AI editors (Cursor, Claude Code, Copilot, Aider, Zed).
44
+ 2. **Minimal:** The dynamic resolver selects only the relevant skills and rules needed for a specific task, eliminating prompt bloat and token waste.
45
+ 3. **Verifiable:** Lockfiles, provenance, drift detection, and CI gates make generated agent configuration reproducible and auditable.
60
46
 
61
- ContextOS allows you to tailor your AI rules to the project lifecycle and architecture:
47
+ ## How it works
62
48
 
63
- | Profile | Focus | Excluded / Filtered Skills | Ideal For |
64
- |---|---|---|---|
65
- | `mvp` | Maximum speed & minimalism | `microservices`, `ddd`, `cqrs`, `kubernetes` | Hackathons, prototypes, fast validation |
66
- | `startup` | Balanced agile stack | `microservices`, `kubernetes` | SaaS startups, modular monoliths |
67
- | `enterprise` | Maximum rigor & compliance | _(none)_ — full TDD, DDD, Security Audit | Large scale teams, strict audit requirements |
68
- | `frontend` | Dedicated UI/UX & React | `fastapi`, `nestjs`, `microservices`, `ddd` | Next.js, React, Design systems, SPAs |
69
- | `backend` | Server-side & APIs | `ui-ux-pro`, `impeccable-design`, `ui-design` | API servers, microservices, databases |
70
-
71
- ### Profile Commands
49
+ 1. Define version-controlled engineering policies once.
50
+ 2. Resolve only the policies relevant to the current task.
51
+ 3. Compile native configuration for each coding agent.
52
+ 4. Detect configuration drift in CI.
72
53
 
73
54
  ```bash
74
- # Auto-detect tech stack in the current project
75
- contextos detect
76
- # or: node .agents/ctx.js detect
55
+ contextos resolve "review authentication changes" \
56
+ --files src/auth/session.ts \
57
+ --explain
58
+ ```
77
59
 
78
- # List available profiles and current active profile
79
- contextos profile list
60
+ Selected:
61
+ security explicit task match
62
+ engineering-workflow required dependency
80
63
 
81
- # Apply a profile
82
- contextos profile apply mvp
64
+ Excluded:
65
+ context-manager context budget
83
66
 
84
- # Recompile all agent exports for the active profile
85
- contextos export all
86
- ```
87
-
88
- ## What's Inside?
89
-
90
- ### Master Orchestrator
91
-
92
- - **AGENTS.md** — The core ruleset. Automatically routes skills by task type and technology detected in your codebase.
93
-
94
- ### Skills (39 total)
95
-
96
- | Category | Skill | What It Does |
97
- | ---------- | ------- | ------------- |
98
- | Core | `engineering-workflow` | Enforces DEFINE→PLAN→BUILD→VERIFY→REVIEW→SHIP pipeline and slash commands |
99
- | Core | `gstack-roles` | 23 specialist roles (PM, Architect, QA Lead, etc.) — AI declares its role before each task |
100
- | Core | `ponytail-mindset` | 7-rung decision ladder before writing any code. Eliminates premature abstraction |
101
- | Core | `interview-me` | Progressive single-question requirements elicitation before drafting specs |
102
- | Core | `subagent-orchestrator` | Multi-agent task decomposition, context boundary isolation, and merge synthesis |
103
- | Core | `gemini-precision` | High-precision engineering guardrails, zero-assumption verification, and zero-placeholder output |
104
- | Frontend | `ui-ux-pro` | Planning guide for UI: color systems, typography, Tailwind v4 `@theme`, Framer Motion |
105
- | Frontend | `impeccable-design` | 50 deterministic QA rules for design review (typography, color, layout, animation) |
106
- | Frontend | `react` | Modern React 19, concurrency, state colocation, `useOptimistic`, and render optimization |
107
- | Frontend | `react-best-practices` | Vercel engineering standards, eliminating async waterfalls, bundle trace optimization |
108
- | Frontend | `nextjs` | Next.js 15+ App Router, RSC, `after()`, `React.cache()`, Server Actions, and PPR |
109
- | Frontend | `typescript` | Type-safe code, generics, config, and invariant type assertions |
110
- | Frontend | `state-management` | Zustand, TanStack Query, client/server state separation |
111
- | Frontend | `ui-design` | Component library design, design tokens, and shadcn/ui patterns |
112
- | Frontend | `ux-design` | User flow design, interaction patterns, and user journey optimization |
113
- | Frontend | `web-accessibility` | ARIA dialogs, focus traps, WCAG 2.1 compliance, and `:focus-visible` standards |
114
- | Frontend | `brutalist-design` | Raw mechanical interfaces, Swiss print typography, and high-contrast styling |
115
- | Frontend | `minimalist-design` | Clean, content-first editorial interfaces with generous negative space |
116
- | Frontend | `soft-design` | Warm, low-contrast premium surfaces with subtle atmospheric depth |
117
- | Frontend | `redesign-audit` | Systematic UI codebase auditing and refactoring without breaking existing features |
118
- | Backend | `system-design` | DDIA patterns (Outbox, CDC, Idempotency), serverless pooling, and CAP trade-offs |
119
- | Backend | `database` | Zero-downtime migrations (expand/contract), PostgreSQL indexing, and serverless pooling |
120
- | Backend | `node` | Node.js asynchronous event loop and server runtime best practices |
121
- | Backend | `fastapi` | FastAPI and Pydantic v2 high-performance Python backends |
122
- | Backend | `nestjs` | Enterprise modular backend architecture and dependency injection |
123
- | Backend | `microservices` | Service boundaries, Saga orchestration/choreography, and Dead Letter Queues |
124
- | Backend | `ddd` | Domain-Driven Design, Aggregate invariants, Domain Events, and Clean Architecture |
125
- | Cross | `security` | Zero-trust auth, OWASP API Top 10, SSRF IP blocking, and Prompt Injection defense |
126
- | Cross | `performance` | Core Web Vitals 2026 (INP < 200ms, LCP < 2.5s), waterfall elimination |
127
- | Cross | `vercel-optimize` | Edge caching, stale-while-revalidate, and Vercel platform optimizations |
128
- | Cross | `testing` | Vitest, React Testing Library behavior testing, and Playwright E2E suites |
129
- | Cross | `docker` | Multi-stage Dockerfiles, non-root security, and container standards |
130
- | Cross | `decisions` | Architectural Decision Records (ADR) format and evaluation |
131
- | Cross | `architecture-diagrams` | Animated, interactive SVG/HTML architecture, sequence, and data-flow diagrams |
132
- | Cross | `adapters` | Multi-agent system export and configuration generation |
133
- | Cross | `generators` | Automated PRD, Architecture, and Task generation |
134
- | Cross | `context-manager` | Smart context token selection and optimization |
135
- | Cross | `context-os` | ContextOS compiler meta-skill |
136
- | Cross | `graphify` | Codebase knowledge graph, Tree-sitter AST dependency mapping, and blast-radius analysis |
137
-
138
- ## Slash Command Workflows
139
-
140
- ContextOS maps development phases directly to slash commands in your AI chat:
141
-
142
- | Command | Role Activated | What It Does |
143
- |:---|:---|:---|
144
- | `/spec` | Product Manager | Turn vague ideas into structured requirements and acceptance criteria |
145
- | `/plan` | Architect | Decompose the spec into atomic, testable tasks (< 2 hours each) |
146
- | `/build` | Senior Developer | Implement code task-by-task with TDD and minimal blast radius |
147
- | `/test` | QA Lead | Run unit, integration, and E2E behavioral tests covering edge cases |
148
- | `/simplify` | Staff Engineer | Run the Ponytail 7-rung ladder to strip over-engineering and dead abstractions |
149
- | `/review` | Staff Engineer + Designer | 5-axis quality gate (correctness, architecture, security, performance, design) |
150
- | `/ship` | Release Engineer | Verify clean CI, lint checks, docs, and rollback plan before merging |
67
+ Risk: high
68
+ Estimated context: 2,840 tokens
151
69
 
152
70
  ## Dynamic Skill Resolution & Unified CLI (`contextos` / `ctx.js`)
153
71
 
154
- ContextOS provides a unified CLI (`contextos` or `npx contextos`) and local engine (`.agents/ctx.js`) to resolve minimal skills on the fly, run health diagnostics, and compile exports for AI assistants.
155
-
156
- ### Dynamic Skill Resolution (`resolve` & `index`)
72
+ ContextOS provides a unified CLI (`contextos` or `npx contextos-agents`) and local engine (`.agents/ctx.js`) to resolve minimal skills on the fly, run health diagnostics, and compile exports for AI assistants.
157
73
 
158
- To prevent context bloat, ContextOS dynamically resolves the exact 2–4 skills needed for any prompt or file:
74
+ ### Dynamic Skill Resolution (`resolve`)
159
75
 
160
76
  ```bash
161
- # Resolve skills for a task description (English):
77
+ # Resolve skills for a task description:
162
78
  contextos resolve "Build an accessible modal component with React and Tailwind"
163
79
 
164
80
  # Output:
165
81
  # [DOMAIN: Frontend] [PHASE: Build] [ROLE: Senior Developer]
166
- # Skills loaded: ponytail-mindset, engineering-workflow, react, ui-ux-pro, web-accessibility
167
-
168
- # Multilingual support (Russian):
169
- contextos resolve "создай модальное окно авторизации и напиши юнит-тесты"
170
-
171
- # Output:
172
- # [DOMAIN: Frontend] [PHASE: Build] [ROLE: Senior Developer]
173
- # Skills loaded: ponytail-mindset, engineering-workflow, react, ui-ux-pro, security, testing
82
+ # Skills loaded: ponytail-mindset, engineering-workflow, gemini-precision
174
83
 
175
- # Resolve skills based on active files (hybrid AST & config analysis):
176
- contextos resolve --files "app/api/auth/route.ts"
177
-
178
- # Generate/update progressive lightweight skills index:
179
- contextos index
180
-
181
- # Clean up lingering .swarm-worktrees directories and orphaned swarm/* git branches:
182
- contextos clean-worktrees
84
+ # Resolve with full evidence scoring explanation:
85
+ contextos resolve "security review" --files apps/web/app/login/page.tsx --explain
183
86
  ```
184
87
 
185
88
  ### Diagnostic Health Check (`contextos doctor`)
186
89
 
187
- Run a comprehensive pre-flight verification across your repository to ensure valid skills, profile alignment, symlinks, git worktree status, and compiler synchronization:
90
+ Run a comprehensive pre-flight verification across your repository to ensure valid skills, profile alignment, and compiler synchronization:
188
91
 
189
92
  ```bash
190
93
  contextos doctor
191
- # or: npx contextos doctor
192
- ```
193
-
194
- ### Context Savings Analytics (`contextos stats`)
195
-
196
- Measure your real token savings. Compares monolithic prompt injection against ContextOS dynamic skill resolution across frontend, backend, security, and full-stack tasks:
197
-
198
- ```bash
199
- contextos stats
200
- ```
201
-
202
- ### Continuous Auto-Sync Daemon (`contextos watch`)
203
-
204
- Watch your source skills in `.agents/core/skills/` and automatically recompile adapter outputs (`.cursorrules`, `.zed/rules.md`, `.github/copilot-instructions.md`, etc.) upon saving:
205
-
206
- ```bash
207
- contextos watch
208
- # or: npm run watch
209
94
  ```
210
95
 
211
96
  ### Supported Agents & Compilation
@@ -220,265 +105,52 @@ contextos watch
220
105
  | **Zed IDE** | `contextos export zed` | `.zed/rules.md` + `.zed/prompts/*.md` |
221
106
 
222
107
  ```bash
223
- contextos export all # Compile for all agents (or: node .agents/ctx.js export all)
224
- contextos export gemini # Compile for Gemini / Antigravity
225
- contextos export claude # Compile for Claude Code
226
- contextos export cursor # Compile for Cursor (.cursor/rules/*.mdc)
227
- contextos export copilot # Compile for GitHub Copilot
228
- contextos export aider # Compile for Aider
229
- contextos export zed # Compile for Zed IDE
108
+ contextos export all # Compile for all agents
230
109
  ```
231
110
 
232
- ### Pre-Compiled Artifacts & Git Architecture
233
-
234
- ContextOS commits generated adapter configurations (`.cursorrules`, `.cursor/rules/*.mdc`, `.github/copilot-instructions.md`, `.aider.conf.yml`, `CONVENTIONS.md`, `.zed/rules.md`) directly into Git:
235
-
236
- - **Zero-Build Onboarding:** AI assistants (Cursor, Claude Code, GitHub Copilot, Zed, Aider) activate instantly upon repository clone without requiring `npm install` or separate build steps.
237
- - **Git-Native Context:** Assistant engines index project rules using native file matchers and git tree walking without depending on background daemon processes.
238
- - **Automated Sync & Drift Prevention:** CI strictly validates that generated exports match source skills (`node .agents/ctx.js validate`). Any uncommitted adapter drift fails CI checks via `git diff --exit-code`.
239
- - **Contributor Workflow:** Source rules are authored exclusively in `.agents/core/skills/<name>/SKILL.md`. Running `node .agents/ctx.js export all` regenerates all assistant configurations deterministically.
240
-
241
111
  ### CI Quality Gate Action (`contextos-gate`)
242
112
 
243
- You can guard your repository against skill drift, secret leaks, and rule regressions using the official reusable GitHub Composite Action:
113
+ Guard your repository against skill drift, secret leaks, and rule regressions using the official GitHub Composite Action:
244
114
 
245
115
  ```yaml
246
116
  # .github/workflows/pr-gate.yml
247
117
  name: ContextOS Quality Gate
248
-
249
- on:
250
- pull_request:
251
- branches: [main]
252
- push:
253
- branches: [main]
254
-
118
+ on: [pull_request, push]
255
119
  jobs:
256
120
  gate:
257
121
  runs-on: ubuntu-latest
258
122
  steps:
259
123
  - uses: actions/checkout@v4
260
124
  - uses: kok-o/contextos-agents/.github/actions/contextos-gate@main
261
- with:
262
- node-version: '20'
263
125
  ```
264
126
 
265
- The action validates skill frontmatter integrity, checks for adapter configuration drift, scans for accidental secrets or API keys, and runs your test suite.
266
-
267
- ### Plugin Skills & Validation
268
-
269
- You can expand your `.agents` folder with community plugins or validate your own custom skills using the top-level commands:
270
-
271
- ```bash
272
- # Launch the interactive skill installer to browse and install community skills
273
- contextos install-skill
274
- # or: npx contextos install-skill
275
-
276
- # Or install a specific skill from a GitHub repository automatically
277
- contextos install-skill --from-repo kok-o/awesome-skill
278
-
279
- # Validate your local skills (checks frontmatter, dependencies, and sync)
280
- contextos audit
281
- ```
282
-
283
- ## ContextOS MCP Server & Autonomous Multi-Agent Swarm
284
-
285
- ContextOS includes a standalone **Model Context Protocol (MCP)** execution server located in `contextos-mcp/` and bundled as `.agents/mcp/server.mjs`. It allows orchestrator agents (like Antigravity, Claude Code, or Cursor) to safely delegate coding tasks to parallel subagents running in isolated Git worktrees.
286
-
287
- > [!TIP]
288
- > **Lightweight by default:** Standard installation (`npx contextos-agents`) installs only lightweight skills, adapters, and behavioral rules (~400 KB) without copying the bundled MCP runtime. To enable MCP worktrees and subagents, pass `--with-mcp` during installation, or run `npx contextos-agents setup-mcp` at any time.
289
-
290
- ### Installing & Enabling MCP
291
-
292
- To add the MCP execution server to an existing `.agents/` project:
293
-
294
- ```bash
295
- npx contextos-agents setup-mcp
296
- ```
297
-
298
- Or install a new project with MCP enabled from the start:
299
-
300
- ```bash
301
- npx contextos-agents --with-mcp
302
- ```
303
-
304
- ### MCP Server Configuration
305
-
306
- Add ContextOS to your IDE's MCP settings (e.g. in `.agents/mcp_config.json`):
307
-
308
- ```json
309
- {
310
- "mcpServers": {
311
- "contextos": {
312
- "command": "node",
313
- "args": [
314
- "./.agents/mcp/server.mjs",
315
- "--dir",
316
- "."
317
- ]
318
- }
319
- }
320
- }
321
- ```
127
+ ## ContextOS MCP Bridge (Beta)
322
128
 
323
- ### Exposed MCP Tools
129
+ ContextOS provides a read-only **Model Context Protocol (MCP)** server to allow compatible agents (like Claude Desktop) to dynamically read project rules, resolve context, and check project status.
324
130
 
325
- | Tool | Purpose | Key Parameters |
326
- |---|---|---|
327
- | `contextos_delegate` | Spawns multiple AI agents in parallel in isolated git worktrees with automatic ContextOS skill injection | `task`, `agents` (model, provider, backend), `wait` (sync/async), `verify_command` (in-worktree test) |
328
- | `contextos_status` | Inspects thread progress, statuses, and diff summaries from memory and persistent disk journal | `dir`, `task_id`, `thread_id` |
329
- | `contextos_diff` | Captures unified git diff and changes for a specific thread | `thread_id`, `dir` |
330
- | `contextos_compare` | Compares multi-agent solutions side-by-side with token cost and execution duration metrics | `thread_ids`, `dir` |
331
- | `contextos_merge` | Merges completed thread branches back into the main working tree with conflict detection | `thread_id`, `dir`, `delete_branch` |
332
- | `contextos_cleanup` | Destroys worktrees, frees sessions, and purges orphaned branches and leftover directories | `dir`, `purge_orphans` |
131
+ > [!WARNING]
132
+ > **Separation of Concerns:** The MCP Bridge and experimental runtime orchestration tools are distributed separately in the `@contextos/mcp` package (Beta). The `--with-mcp` flag in the Core CLI is deprecated.
333
133
 
334
- ### Enterprise Architecture Guarantees
134
+ Install the optional MCP Bridge:
335
135
 
336
- - **Git Worktree Sandboxing:** Each subagent operates in a private git worktree (`.swarm-worktrees/`). The developer's active workspace cannot be corrupted by experimental changes or failing tests.
337
- - **Disk-Backed Session Persistence:** Active and completed threads are recorded in `.swarm-worktrees/session-state.json`. If the MCP process is restarted, tasks and diffs can be recovered without losing work.
338
- - **Automated In-Worktree Verification (`verify_command`):** Runs test commands (`npm test`, `cargo test`, `pytest`) inside the isolated worktree before marking tasks as successful.
339
- - **Context Token Compression:** The server extracts essential rules, constraints, and checklists (`extractEssentialSkillContent`), eliminating verbose samples and reducing prompt overhead.
340
-
341
- ## Testing
342
-
343
- Tests use the **Node.js built-in test runner** for the core framework and **Vitest** for the MCP engine — zero external test bloat.
344
-
345
- ### 1. Root Test Suite (131 tests)
346
-
347
- ```bash
348
- npm test
349
- ```
350
-
351
- ```text
352
- # tests 131
353
- # suites 27
354
- # pass 131
355
- # fail 0
356
- ```
357
-
358
- ### 2. MCP Server Test Suite (440 tests)
359
-
360
- ```bash
361
- cd contextos-mcp && npm test
362
- ```
363
-
364
- ```text
365
- Test Files 25 passed | 1 skipped (26)
366
- Tests 440 passed | 13 skipped (453)
367
- ```
368
-
369
- **Test coverage:**
370
-
371
- - `tests/install.test.js` — installer CLI flags (--help, --minimal, --dry-run, --force)
372
- - `tests/export.test.js` — ctx.js export for gemini, claude, cursor (.mdc rules), copilot, aider, zed
373
- - `tests/skills.test.js` — validates all skill source files and frontmatter
374
- - `tests/profile.test.js` — profile resolution, stack auto-detection, and skill filtering
375
- - `tests/validate.test.js` — validator rules, dependency graph, and sync checks
376
- - `tests/plugins.test.js` — plugin lockfile, registry fetching, and security checks
377
- - `tests/resolver.test.js` — dynamic skill resolution, AST import graph analysis, progressive index, and bilingual prompt matching
378
- - `tests/benchmark.test.js` — benchmark scoring engine, static AST checks, runtime sandbox, and reporters
379
- - `contextos-mcp/tests/unit/session-persistence.test.ts` — session disk persistence, thread state tracking, and orphan purge
380
- - `contextos-mcp/tests/unit/contextos-tools.test.ts` — all 6 MCP tool handlers and validation
381
-
382
- ## Benchmark: With Skills vs. Without Skills
383
-
384
- The repository includes a paired, reproducible code-quality benchmark suite supporting OpenAI (GPT-4o, GPT-5, o1, o3-mini), Google Gemini, Anthropic Claude, and custom gateways (AgentRouter, OpenRouter).
385
-
386
- The benchmark evaluates real-world code quality, security vulnerabilities, timing attacks, ARIA accessibility contracts, DDD business invariants, and error isolation between baseline LLMs and ContextOS-assisted agents.
387
-
388
- ### Live Benchmark Execution
389
-
390
- ```bash
391
- # 1. Run live benchmark with OpenAI (GPT-4o, GPT-5, o3-mini):
392
- set OPENAI_API_KEY=sk-... # PowerShell: $env:OPENAI_API_KEY = "sk-..."
393
- npm run benchmark:live -- --provider openai --model gpt-4o
394
-
395
- # 2. Run live benchmark with Google Gemini:
396
- set GEMINI_API_KEY=... # PowerShell: $env:GEMINI_API_KEY = "..."
397
- npm run benchmark:live -- --provider gemini --model gemini-2.5-flash
398
-
399
- # 3. Run live benchmark with Anthropic Claude:
400
- set ANTHROPIC_API_KEY=... # PowerShell: $env:ANTHROPIC_API_KEY = "..."
401
- npm run benchmark:live -- --provider anthropic --model claude-3-7-sonnet-20250219
402
-
403
- # 4. Run with custom OpenAI-compatible router (OpenRouter, AgentRouter, Local vLLM):
404
- node benchmarks/run-live-benchmark.js --base-url "https://agentrouter.org/v1" --api-key "sk-..." --model "gpt-5.6-sol" --open
405
- ```
406
-
407
- ### Execution-Backed Runtime Benchmark (Real Sandbox Test Assertions)
408
-
409
- In addition to static checks, ContextOS features an **execution-backed runtime benchmark suite**. It compiles model-generated code in an isolated Node.js V8 sandbox (`node:vm`) and runs rigorous behavioral unit assertions (`node:assert`):
410
-
411
- ```bash
412
- # 1. Run runtime benchmark with OpenRouter (Google Gemini 3.8 Flash):
413
- node benchmarks/run-runtime-benchmark.js --base-url "https://openrouter.ai/api/v1" --api-key "sk-or-v1-..." --model "google/gemini-3.8-flash" --open
414
-
415
- # 2. Run runtime benchmark with AgentRouter (GPT-5.6-sol):
416
- node benchmarks/run-runtime-benchmark.js --base-url "https://agentrouter.org/v1" --api-key "sk-..." --model "gpt-5.6-sol" --open
417
-
418
- # 3. Run specific scenario (auth-security, ddd-order-invariants, or resilient-api-client):
419
- npm run benchmark:runtime -- --base-url "https://agentrouter.org/v1" --api-key "sk-..." --model "gpt-5.6-sol" --task auth-security
420
- ```
421
-
422
- ### Evaluation Methodology
423
-
424
- Submissions are evaluated using a strict, multi-stage verification pipeline:
425
-
426
- 1. **Sandboxed V8 Runtime Execution (Primary Ground Truth):** Compiles TypeScript into CommonJS via native AST type stripping (`node:module.stripTypeScriptTypes`) and executes in an isolated sandbox with timeout and assertion checks (`node:assert`).
427
- 2. **Behavioral Invariant Testing:** Stress-tests timing attacks (`crypto.timingSafeEqual`), brute-force IP/Account rate-limiting, error stack redaction, immutable Value Objects, domain event dispatch, and circuit breaker state transitions.
428
- 3. **Deterministic Static Analysis:** AST verification checking for zero ORM/HTTP transport leakage in domain layers and contract compliance.
429
-
430
- ### Production Scenarios Evaluated
431
-
432
- The runtime sandbox evaluates model outputs against real-world engineering invariants:
433
-
434
- | Scenario | Category | Skills Activated | Key Technical Invariant Proved |
435
- |---|---|---|---|
436
- | **Secure Auth & Rate Limiting** | Security & Backend | `security`, `node`, `ponytail-mindset` | Constant-time password verification (`timingSafeEqual`), dual-key rate-limiting, strict email/credential sanitization, zero stack-trace leak in 500s. |
437
- | **DDD Order Aggregate Root** | Architecture & DDD | `ddd`, `system-design`, `decisions` | Immutable `Money` Value Object, state-machine invariants (PENDING → PAID → SHIPPED), explicit Domain Event classes with queue draining. |
438
- | **Resilient API Client** | Reliability & Async | `typescript`, `system-design`, `performance` | 3-state Circuit Breaker (CLOSED → OPEN → HALF-OPEN), `AbortController` timeouts, typed error taxonomy without credential leakage. |
439
-
440
- ### Running Benchmarks Locally & in CI
441
-
442
- You can run the benchmark suite locally with your own API keys:
443
-
444
- ```bash
445
- # Run runtime sandbox benchmark with Google Gemini:
446
- $env:GEMINI_API_KEY = "your-key"
447
- npm run benchmark:runtime -- --provider gemini --model gemini-2.5-flash
448
-
449
- # Run with OpenAI:
450
- $env:OPENAI_API_KEY = "sk-..."
451
- npm run benchmark:runtime -- --provider openai --model gpt-4o
452
-
453
- # Run complete multi-model runtime matrix (Gemini, OpenRouter, AgentRouter):
454
- GEMINI_API_KEY=... OPENROUTER_API_KEY=... AGENTROUTER_API_KEY=... node benchmarks/run-multi-runtime.cjs
455
- ```
456
-
457
- When executed, reports are generated in `benchmarks/results/` (`.html`, `.md`, `.json`) and tracked so results are visible and shareable.
458
-
459
- > [!NOTE]
460
- > **Why Runtime Benchmarks Are Dispatch-Only in CI:** Standard CI checks (`validate-skills.yml`) run hermetically without external API calls to avoid flaky network dependencies and API token expenditures on every pull request. Live runtime evaluation is triggered on demand via GitHub Actions **Workflow Dispatch** ([`benchmark-runtime.yml`](.github/workflows/benchmark-runtime.yml)) using secure repository secrets.
136
+ `ash
137
+ npm install --save-dev @contextos/mcp
138
+ npx contextos-mcp --dir .
139
+ `
461
140
 
141
+ The MCP Bridge is read-only by default. Experimental execution tools require
142
+ the explicit --enable-runtime flag and should only be used in trusted repositories.
462
143
 
463
144
  ## Security — Third-Party Skills
464
145
 
465
146
  ContextOS skills are **executable context** — they become part of the system prompt that controls your AI agent's behavior. A malicious skill could instruct the AI agent to exfiltrate environment variables, modify files, or ignore your project's security policies.
466
147
 
467
148
  > [!CAUTION]
468
- > **Install skills only from repositories you trust as you would trust executable code.** Skills installed via `ctx.js skill add` from npm or GitHub are not sandboxed. ContextOS includes a built-in prompt injection scanner, but it cannot guarantee safety of arbitrary third-party content.
149
+ > **Install skills only from repositories you trust as you would trust executable code.** Skills installed via `ctx.js skill add` from npm or GitHub are not sandboxed.
469
150
 
470
151
  ## Contributing
471
152
 
472
- We are open to pull requests! See [CONTRIBUTING.md](./CONTRIBUTING.md) for a step-by-step guide on how to add a new skill.
473
-
474
- Quick start:
475
-
476
- 1. Fork the repository
477
- 2. Create your feature branch (`git checkout -b feature/AmazingSkill`)
478
- 3. Add your skill in `.agents/core/skills/<name>/SKILL.md`
479
- 4. Run `npm test` — all tests must pass
480
- 5. Commit your changes (`git commit -m 'feat: add AmazingSkill'`)
481
- 6. Push and open a Pull Request
153
+ We are open to pull requests! See [CONTRIBUTING.md](./CONTRIBUTING.md) for a step-by-step guide.
482
154
 
483
155
  ## License
484
156