@qwen-code/qwen-code 0.19.9 → 0.19.10-nightly.20260715.c538bd70d

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (310) hide show
  1. package/bundled/qc-helper/docs/configuration/settings.md +8 -0
  2. package/bundled/qc-helper/docs/features/channels/dingtalk.md +3 -0
  3. package/bundled/qc-helper/docs/features/channels/overview.md +11 -3
  4. package/bundled/qc-helper/docs/features/code-review.md +88 -44
  5. package/bundled/qc-helper/docs/features/commands.md +6 -6
  6. package/bundled/qc-helper/docs/features/hooks.md +29 -0
  7. package/bundled/qc-helper/docs/features/sub-agents.md +21 -0
  8. package/bundled/qc-helper/docs/qwen-serve.md +46 -20
  9. package/bundled/qc-helper/docs/reference/keyboard-shortcuts.md +1 -1
  10. package/bundled/review/DESIGN.md +219 -33
  11. package/bundled/review/SKILL.md +449 -153
  12. package/chunks/{MaxSizedBox-6HYJ3ZPJ.js → MaxSizedBox-LOOBDOTZ.js} +31 -31
  13. package/chunks/{StandaloneSessionPicker-WS4GXYOG.js → StandaloneSessionPicker-76CLZSH2.js} +51 -49
  14. package/chunks/{acpAgent-5FWSOZL5.js → acpAgent-DFJOK7EB.js} +567 -482
  15. package/chunks/{agent-HOLHB3RT.js → agent-OJQWXNHP.js} +26 -26
  16. package/chunks/{agent-headless-SXL6H56A.js → agent-headless-ND7GOORY.js} +26 -26
  17. package/chunks/{anthropicContentGenerator-L73TO3DB.js → anthropicContentGenerator-56U3YCP2.js} +5 -5
  18. package/chunks/{artifact-tool-U7IRTLIV.js → artifact-tool-MIZLXBLF.js} +2 -2
  19. package/chunks/{askUserQuestion-2DLUUZZG.js → askUserQuestion-XMJVPNWM.js} +2 -2
  20. package/chunks/{bridge-BCJQ7BBB.js → bridge-SNJK2JLK.js} +32 -32
  21. package/chunks/{ca-UJ6HC4H2.js → ca-C2JTT2N5.js} +4 -1
  22. package/chunks/channel-worker-group-K7QKH6IG.js +13 -0
  23. package/chunks/channel-worker-manager-YSFB4HUU.js +418 -0
  24. package/chunks/channel-worker-supervisor-QCKIOYZD.js +20 -0
  25. package/chunks/{chunk-CCDMCZPP.js → chunk-26UABK2C.js} +1 -1
  26. package/chunks/{chunk-LXNEJD34.js → chunk-2FO5GFKZ.js} +1 -1
  27. package/chunks/{chunk-L36P3NBD.js → chunk-2W3OOD4W.js} +2 -0
  28. package/chunks/{chunk-7R3CZA3Q.js → chunk-2Z5UX4KI.js} +1 -1
  29. package/chunks/chunk-3MX7D6QN.js +516 -0
  30. package/chunks/{chunk-CLFKOZEK.js → chunk-3XULHBC7.js} +1 -1
  31. package/chunks/chunk-4QZADKJS.js +34 -0
  32. package/chunks/{chunk-E2IY465S.js → chunk-4UDHAJFV.js} +10 -10
  33. package/chunks/{chunk-5U27BZ72.js → chunk-55SAYASQ.js} +190 -77
  34. package/chunks/chunk-5HBA2Z7V.js +28 -0
  35. package/chunks/{chunk-ZJE7NSL3.js → chunk-5M6IDOMF.js} +34 -97
  36. package/chunks/{chunk-HVY356M2.js → chunk-5SQ55XJZ.js} +3 -3
  37. package/chunks/{chunk-EAXLGSFP.js → chunk-6GRAQNTY.js} +8 -8
  38. package/chunks/{chunk-K434OPES.js → chunk-6WPJRYLZ.js} +3 -3
  39. package/chunks/{chunk-VWPVHAMO.js → chunk-7DXHIMXL.js} +5 -1
  40. package/chunks/{chunk-J4TSQ5EQ.js → chunk-7NFB7LBQ.js} +8 -1
  41. package/chunks/{chunk-5SFOMFG4.js → chunk-7R4DPSRA.js} +36 -2
  42. package/chunks/{chunk-W2UIXQDL.js → chunk-7STBROEV.js} +1 -1
  43. package/chunks/{chunk-BMGVRZN4.js → chunk-AETEZU4U.js} +1 -1
  44. package/chunks/{chunk-HPDUL5MF.js → chunk-AGZWN4WV.js} +1 -1
  45. package/chunks/{chunk-RDJSGHZE.js → chunk-ALGGS7UH.js} +1 -38
  46. package/chunks/{chunk-GWHYEOV2.js → chunk-ARKANCNX.js} +2 -1
  47. package/chunks/{chunk-34Z3U5RQ.js → chunk-AU4R7ACM.js} +518 -114
  48. package/chunks/{channel-worker-supervisor-WSOAAYUO.js → chunk-B7MFJBKE.js} +67 -35
  49. package/chunks/{chunk-524RQP4Q.js → chunk-B7NUUEXC.js} +10 -4
  50. package/chunks/{chunk-MC23XKFE.js → chunk-BLPMRN4B.js} +3 -3
  51. package/chunks/{chunk-A3XKDTSH.js → chunk-BRRKAUGR.js} +4 -4
  52. package/chunks/{chunk-W3JHSWRP.js → chunk-BUPLOPUA.js} +4 -4
  53. package/chunks/{chunk-3VAGG5T5.js → chunk-BYOOK22F.js} +121 -16
  54. package/chunks/{chunk-EMDMPBJ4.js → chunk-C47ENMZJ.js} +3 -3
  55. package/chunks/{chunk-3AHXV67U.js → chunk-C5TVXLJZ.js} +1 -1
  56. package/chunks/{chunk-X6TJ3YZV.js → chunk-CGT5TQIZ.js} +159 -57
  57. package/chunks/{chunk-4C5BQ2D2.js → chunk-CQQD7MDG.js} +82 -12
  58. package/chunks/{chunk-KKNMOEZH.js → chunk-E2UVIFI2.js} +185 -398
  59. package/chunks/{chunk-P67PFTUT.js → chunk-EPY4ZEU6.js} +4 -4
  60. package/chunks/{chunk-6PESIEKJ.js → chunk-EZSRHWGS.js} +68 -6
  61. package/chunks/{chunk-4GORAUCN.js → chunk-F333XEAU.js} +1 -1
  62. package/chunks/{chunk-YLV3XXFB.js → chunk-FBU7WRZI.js} +17 -1
  63. package/chunks/{chunk-Q6I5QHMH.js → chunk-FGTCZJOD.js} +1 -1
  64. package/chunks/{chunk-LK7KUE55.js → chunk-FSA7ERJ2.js} +1 -1
  65. package/chunks/{chunk-B4G2WOFA.js → chunk-FTU55ZZ7.js} +3 -3
  66. package/chunks/{chunk-EGYZ275M.js → chunk-GCD7PEDZ.js} +4 -4
  67. package/chunks/{chunk-SJHDYNHF.js → chunk-GDQQRV43.js} +3 -3
  68. package/chunks/{chunk-3BTRNWWD.js → chunk-GOXKNSDZ.js} +1 -64
  69. package/chunks/{chunk-GZRTYZ2R.js → chunk-GQKWT45X.js} +154 -1
  70. package/chunks/chunk-GYTKQYDB.js +23 -0
  71. package/chunks/{chunk-YVEDCAUE.js → chunk-HGT6JR3U.js} +19 -0
  72. package/chunks/{chunk-N3EMS75V.js → chunk-HKM2TEF7.js} +1268 -1233
  73. package/chunks/chunk-IDS7MSUP.js +74 -0
  74. package/chunks/{chunk-PVZM22CK.js → chunk-INKAZNYC.js} +5 -5
  75. package/chunks/{chunk-3G7CIHVV.js → chunk-IOQ75R4R.js} +4 -4
  76. package/chunks/{chunk-AURZZYMD.js → chunk-IQOOJSPR.js} +4 -4
  77. package/chunks/{chunk-AKIVHSJR.js → chunk-J4UQXGVL.js} +111 -93
  78. package/chunks/{chunk-VTU57BWZ.js → chunk-JEKOCLRN.js} +1 -1
  79. package/chunks/{chunk-B5ORX3FG.js → chunk-JGUT3LWZ.js} +6 -1
  80. package/chunks/{chunk-PDZ7MBDG.js → chunk-JIQHT2O6.js} +27 -6
  81. package/chunks/{chunk-XUOOWUZS.js → chunk-JPAH76WA.js} +1 -1
  82. package/chunks/{chunk-BAPGSC7U.js → chunk-K4LG6TVJ.js} +23 -9
  83. package/chunks/chunk-KET3M4K4.js +493 -0
  84. package/chunks/{chunk-ORBGV7NJ.js → chunk-KKRTYONM.js} +6 -4
  85. package/chunks/{chunk-SCGXUJBT.js → chunk-KPA5QGN2.js} +3 -3
  86. package/chunks/{chunk-IAVB3HM3.js → chunk-KSO42X3Z.js} +13 -0
  87. package/chunks/{chunk-4QLQE5WJ.js → chunk-L5U3ELZX.js} +1 -1
  88. package/chunks/{chunk-77XJ3SWN.js → chunk-LIUEXYPL.js} +5 -5
  89. package/chunks/{chunk-4ZVWQHEQ.js → chunk-LQIDP6K5.js} +1 -1
  90. package/chunks/{chunk-34MW4KTI.js → chunk-LT5CRUA5.js} +2 -2
  91. package/chunks/{chunk-DC5Z3DU6.js → chunk-LXLTBIDP.js} +43 -6
  92. package/chunks/{chunk-VA4KQY6V.js → chunk-MEK7CQJR.js} +3 -3
  93. package/chunks/{chunk-KHUQZZJ6.js → chunk-MO7O5722.js} +370 -36
  94. package/chunks/chunk-MR3PXB6E.js +48 -0
  95. package/chunks/{chunk-PODM4SZ4.js → chunk-MV52C5MA.js} +5 -5
  96. package/chunks/{chunk-WL4L7EZH.js → chunk-OA4JVLSA.js} +27 -1
  97. package/chunks/{chunk-X4NKFGFJ.js → chunk-OJAMDVF5.js} +0 -15
  98. package/chunks/{chunk-BZIFY5TF.js → chunk-OKAIGAYW.js} +3 -3
  99. package/chunks/{chunk-ECNYL4AA.js → chunk-OPNPGO6D.js} +1 -1
  100. package/chunks/{chunk-3SJSQXNS.js → chunk-PEJWA3QQ.js} +12 -8
  101. package/chunks/{chunk-MXL5UP4X.js → chunk-Q6TUALBE.js} +1 -1
  102. package/chunks/{chunk-QYMQSECS.js → chunk-QAW7PIHT.js} +2 -2
  103. package/chunks/{chunk-5HOMWMCL.js → chunk-QEDYXCLG.js} +4167 -1338
  104. package/chunks/chunk-QNNTA7IX.js +91 -0
  105. package/chunks/{chunk-D2GFWEXB.js → chunk-QNT5T4Q2.js} +6 -6
  106. package/chunks/{chunk-CUMMSHH6.js → chunk-R5XLNDE2.js} +1 -1
  107. package/chunks/{chunk-LUODI7WD.js → chunk-RCEEXIOO.js} +6 -4
  108. package/chunks/{chunk-WQUQVGTF.js → chunk-RQWDLG6Z.js} +183 -96
  109. package/chunks/{chunk-XYQSMABS.js → chunk-SLCVZX6F.js} +4 -4
  110. package/chunks/{chunk-CEA3E3JB.js → chunk-SPJQO2CE.js} +2 -2
  111. package/chunks/{chunk-JSG7ZSBC.js → chunk-TQBODCFD.js} +1 -1
  112. package/chunks/chunk-TV6K2FB4.js +224 -0
  113. package/chunks/{chunk-5XUCZNSY.js → chunk-TYLOFD2U.js} +1 -1
  114. package/chunks/{chunk-CNEIYWSA.js → chunk-UKWAQ7AV.js} +5 -5
  115. package/chunks/{chunk-7PU4FYLG.js → chunk-US4YQT62.js} +4 -4
  116. package/chunks/{chunk-Q4YQW46J.js → chunk-UYPPCIJQ.js} +36 -18
  117. package/chunks/{chunk-YAUFFXDS.js → chunk-UZKG7HFG.js} +12 -11
  118. package/chunks/{chunk-GWLSNBG4.js → chunk-VQLRHJKL.js} +30 -26
  119. package/chunks/{chunk-L6P4KPHR.js → chunk-VRUHAVZS.js} +59 -39
  120. package/chunks/{chunk-HKMAENAO.js → chunk-W2U4RJJU.js} +8 -6
  121. package/chunks/{chunk-ATAKVRIF.js → chunk-WDSCTLAB.js} +7 -7
  122. package/chunks/{chunk-ZLXQN2QS.js → chunk-WIEO4CWB.js} +1 -1
  123. package/chunks/{chunk-OQ4DRIOA.js → chunk-WJOC24IW.js} +2 -2
  124. package/chunks/{chunk-J3JA76CF.js → chunk-WTNZ6DXB.js} +1 -1
  125. package/chunks/{chunk-NVYTCB5Z.js → chunk-XCR44EEA.js} +14589 -13980
  126. package/chunks/{chunk-TDEXZKMT.js → chunk-Y66VOBRP.js} +2 -2
  127. package/chunks/{chunk-CC2ITGCF.js → chunk-YHN5SUIJ.js} +49 -6
  128. package/chunks/{chunk-D5SKC4YG.js → chunk-ZBAROUR6.js} +41 -13
  129. package/chunks/{chunk-U2NNEY4R.js → chunk-ZEEVKDO7.js} +1 -1
  130. package/chunks/{chunk-5VXNZWGF.js → chunk-ZLOHB76C.js} +1 -1
  131. package/chunks/{chunk-JZETEN3Z.js → chunk-ZNCLTRNU.js} +3 -3
  132. package/chunks/{chunk-TO6QE22P.js → chunk-ZP4ZZJJG.js} +26 -22
  133. package/chunks/{chunk-UOZWKTQG.js → chunk-ZPAHLFZU.js} +2 -2
  134. package/chunks/{chunk-XLA62MYB.js → chunk-ZTXZ7FM4.js} +2 -2
  135. package/chunks/{computer-use-HM37RX6A.js → computer-use-DRETHEQF.js} +26 -26
  136. package/chunks/{config-utils-NV2SVGPB.js → config-utils-XSHOMMK4.js} +3 -2
  137. package/chunks/{contextCommand-OF5SFGH2.js → contextCommand-7YJUW23J.js} +28 -28
  138. package/chunks/{create-sub-session-INX7TXJI.js → create-sub-session-EAB2U5XW.js} +2 -2
  139. package/chunks/{create-sub-session-KPRT7ECJ.js → create-sub-session-L5HOZAEL.js} +26 -26
  140. package/chunks/{cron-create-AMBKGZX6.js → cron-create-GMVKSXZT.js} +4 -4
  141. package/chunks/{cron-delete-LRU4OJYG.js → cron-delete-VDZKUAVK.js} +4 -4
  142. package/chunks/{cron-list-WVP2JXER.js → cron-list-UG7C7RAR.js} +4 -4
  143. package/chunks/{daemon-ARHTSTXP.js → daemon-MI5HF26N.js} +603 -62
  144. package/chunks/{daemon-status-provider-B2UTZUJ2.js → daemon-status-provider-2D47GOJT.js} +36 -36
  145. package/chunks/{de-EBJN3OVN.js → de-2VCDRDGU.js} +4 -1
  146. package/chunks/{dist-WKPOYU7O.js → dist-IEXZC5ZR.js} +1 -1
  147. package/chunks/{dist-XKU3ABM5.js → dist-PQK4GCKA.js} +1 -1
  148. package/chunks/{dist-VHV4EVHG.js → dist-U75JK3PF.js} +1 -1
  149. package/chunks/{dist-R7LN5AZE.js → dist-VXO7QBON.js} +2 -2
  150. package/chunks/{dist-PNVLTKTL.js → dist-WH4TZSH3.js} +158 -4
  151. package/chunks/{dist-QYCAEZIT.js → dist-X2ABIW5R.js} +5 -1
  152. package/chunks/{earlyInputCapture-3VTHRC4C.js → earlyInputCapture-OJQPGCSX.js} +27 -27
  153. package/chunks/{edit-VNI4NJ73.js → edit-LZOAFLT2.js} +27 -27
  154. package/chunks/{en-5A6LA7RK.js → en-NJCA3TAA.js} +5 -1
  155. package/chunks/{enter-worktree-IYYBKPZT.js → enter-worktree-VL2TCGI5.js} +26 -26
  156. package/chunks/{enterPlanMode-GMC5FD5A.js → enterPlanMode-QGHBHZWI.js} +42 -27
  157. package/chunks/{environment-5HSSRMJ7.js → environment-2G5IC2EU.js} +30 -29
  158. package/chunks/{errors-HMFCJJ7P.js → errors-7QNQDNI2.js} +28 -28
  159. package/chunks/{exit-worktree-6BHCMP5G.js → exit-worktree-5V3X2MRH.js} +26 -26
  160. package/chunks/{exitPlanMode-CJW7G74W.js → exitPlanMode-XDPNQB4H.js} +26 -26
  161. package/chunks/{fast-path-FRDMDVHG.js → fast-path-NNSD4BVZ.js} +3 -3
  162. package/chunks/{fast-path-settings-5QZE64HR.js → fast-path-settings-DAOVFADC.js} +6 -4
  163. package/chunks/{fr-YXRABLYZ.js → fr-5F3E7WKD.js} +4 -1
  164. package/chunks/{gemini-QYSEGHAX.js → gemini-RKDUWGXJ.js} +139 -76
  165. package/chunks/{geminiContentGenerator-J7YNZDYP.js → geminiContentGenerator-V7QGHXVF.js} +4 -4
  166. package/chunks/{glob-BRCSGFOT.js → glob-ZNQUUOTS.js} +28 -27
  167. package/chunks/{grep-EE2DUOSY.js → grep-NYIGLM7R.js} +26 -26
  168. package/chunks/{handleAutoUpdate-5NWTLR67.js → handleAutoUpdate-UBQ4IQXY.js} +30 -30
  169. package/chunks/{i18n-BASPENNQ.js → i18n-5T3S6SGP.js} +27 -27
  170. package/chunks/initializer-TJED246Q.js +73 -0
  171. package/chunks/{installationInfo-W6V4DT7W.js → installationInfo-OTEU7FKR.js} +27 -27
  172. package/chunks/{ja-7VQOANVG.js → ja-66J6B53L.js} +5 -2
  173. package/chunks/{keychain-token-storage-VKUBNCP4.js → keychain-token-storage-AU22CQI6.js} +2 -2
  174. package/chunks/list-6OHS7MFJ.js +76 -0
  175. package/chunks/loadedSettingsAdapter-5456XIYG.js +70 -0
  176. package/chunks/{loop-wakeup-WX7BLWWP.js → loop-wakeup-C4EXLYVA.js} +5 -5
  177. package/chunks/{ls-ZB2BDP6T.js → ls-2OUTG3IZ.js} +4 -4
  178. package/chunks/{lsp-V5DN4YUJ.js → lsp-4VMNXZW6.js} +2 -2
  179. package/chunks/mcp-NOWRPFKY.js +70 -0
  180. package/chunks/{monitor-X5Z3H4FY.js → monitor-CVDZO3IZ.js} +26 -26
  181. package/chunks/nonInteractiveCli-UGI7Y43E.js +129 -0
  182. package/chunks/{notebook-edit-TRM2QIJK.js → notebook-edit-NWIA5TEJ.js} +27 -27
  183. package/chunks/{openaiContentGenerator-AAYONPMF.js → openaiContentGenerator-KJONBBJO.js} +12 -12
  184. package/chunks/{pidfile-D4I4FTVA.js → pidfile-L3TGXMTF.js} +27 -27
  185. package/chunks/{pt-2C6YCSHL.js → pt-TFZO5Y6T.js} +4 -1
  186. package/chunks/{qwenContentGenerator-K3HZ2H76.js → qwenContentGenerator-GXUZB5SJ.js} +28 -28
  187. package/chunks/{qwenOAuth2-PJQ27GSO.js → qwenOAuth2-LEYOU7RU.js} +5 -5
  188. package/chunks/{read-file-7SZSIZOL.js → read-file-I4OUZXTA.js} +8 -8
  189. package/chunks/{read-mcp-resource-SM5DA4RG.js → read-mcp-resource-JUFHCKSC.js} +2 -2
  190. package/chunks/{record-artifact-4XKZX3UW.js → record-artifact-DWSMBXYY.js} +2 -2
  191. package/chunks/{ripGrep-YNPFVCZP.js → ripGrep-XLCXANIO.js} +26 -26
  192. package/chunks/{ru-6BBHVZZV.js → ru-4L4LFHIT.js} +4 -1
  193. package/chunks/{run-qwen-serve-DMWV5N6L.js → run-qwen-serve-R6SBP6YS.js} +1232 -371
  194. package/chunks/{runtime-57WWRVHN.js → runtime-U3S33YUQ.js} +42 -37
  195. package/chunks/{scheduler-YBLGWSIL.js → scheduler-65QCIELP.js} +26 -26
  196. package/chunks/{send-message-F36YD6KQ.js → send-message-JZ752NWD.js} +3 -3
  197. package/chunks/{serve-R63NK3CV.js → serve-PSPOFCOD.js} +36 -34
  198. package/chunks/{server-Z22GCAXP.js → server-6P5ZB56O.js} +5561 -1377
  199. package/chunks/{session-HKAAI7BN.js → session-ILLOYW75.js} +298 -98
  200. package/chunks/{settings-3EYN66GK.js → settings-RMUJGYIK.js} +34 -32
  201. package/chunks/{shell-6YXINTWJ.js → shell-65GTAN5Q.js} +28 -26
  202. package/chunks/{skill-OO57P4X6.js → skill-L76U3YCX.js} +10 -10
  203. package/chunks/{spawnChannel-JD7G4VU6.js → spawnChannel-LXIGZ4EM.js} +28 -28
  204. package/chunks/{src-G5UPLZSZ.js → src-DWZT2X3Z.js} +87 -27
  205. package/chunks/{standalone-update-OB7F4JPZ.js → standalone-update-F6SRKPSD.js} +28 -28
  206. package/chunks/{startInteractiveUI-73EGLIIH.js → startInteractiveUI-UEJQP7S3.js} +730 -397
  207. package/chunks/{syntheticOutput-34H7LGA7.js → syntheticOutput-LELY6HAI.js} +3 -3
  208. package/chunks/{task-create-FRGOEIKY.js → task-create-QY3BIMNM.js} +8 -7
  209. package/chunks/{task-list-OW766BXZ.js → task-list-D2U47WBJ.js} +6 -5
  210. package/chunks/{task-stop-SIR5SNBN.js → task-stop-N6E5IYLH.js} +2 -2
  211. package/chunks/{task-update-YRA532MT.js → task-update-HJCQILLB.js} +8 -7
  212. package/chunks/{team-create-PECAAPXK.js → team-create-4QBZVOGN.js} +26 -26
  213. package/chunks/{team-delete-6OSECUZC.js → team-delete-RCWYDRF4.js} +6 -5
  214. package/chunks/{team-plan-approval-L6F434RY.js → team-plan-approval-BK4CIOFV.js} +26 -26
  215. package/chunks/{theme-manager-ZG4RY6GU.js → theme-manager-5F2IGW6E.js} +27 -27
  216. package/chunks/{todoWrite-YDMFEV4G.js → todoWrite-FKCGATPJ.js} +4 -4
  217. package/chunks/{tool-search-THVIWAPB.js → tool-search-J2ESO33H.js} +9 -9
  218. package/chunks/{total-session-admission-JOKNRIIA.js → total-session-admission-U4LZ5QTQ.js} +33 -33
  219. package/chunks/tree-sitter-YJVE2ZUK.js +2980 -0
  220. package/chunks/{trustedFolders-EOGT3UVS.js → trustedFolders-TIDOXCJA.js} +29 -28
  221. package/chunks/types-ML3TRJQ5.js +16 -0
  222. package/chunks/{updateCheck-2KYEYVSB.js → updateCheck-VRVR26FI.js} +29 -29
  223. package/chunks/{validateNonInterActiveAuth-EN6ZQSR2.js → validateNonInterActiveAuth-UGKYAVLG.js} +75 -71
  224. package/chunks/{version-3UEA6ZIR.js → version-7HR6VZAT.js} +1 -1
  225. package/chunks/{web-fetch-PAAEQKBW.js → web-fetch-QGODBB6S.js} +5 -5
  226. package/chunks/{workflow-QCCCMYP3.js → workflow-EDQAITEW.js} +27 -27
  227. package/chunks/workspace-providers-status-27UGAK77.js +73 -0
  228. package/chunks/workspace-registration-store-5H7577YQ.js +27 -0
  229. package/chunks/{workspace-registry-24QFLQYH.js → workspace-registry-F3THGAXC.js} +33 -33
  230. package/chunks/workspace-service-4JSAQAHP.js +86 -0
  231. package/chunks/workspace-skills-status-5JTWZHZU.js +72 -0
  232. package/chunks/{write-file-I3MUDRIE.js → write-file-UY6XCZZM.js} +27 -27
  233. package/chunks/{zh-RQE7I22T.js → zh-FSAH32JB.js} +6 -2
  234. package/chunks/{zh-TW-TBPQGHKH.js → zh-TW-D2YIE2LO.js} +6 -2
  235. package/cli.js +11 -11
  236. package/locales/ca.js +6 -0
  237. package/locales/de.js +6 -0
  238. package/locales/en.js +8 -0
  239. package/locales/fr.js +6 -0
  240. package/locales/ja.js +7 -1
  241. package/locales/pt.js +6 -0
  242. package/locales/ru.js +6 -0
  243. package/locales/zh-TW.js +9 -1
  244. package/locales/zh.js +9 -1
  245. package/package.json +3 -3
  246. package/web-shell/assets/{arc-C15hqnJe.js → arc-CObt_TK9.js} +1 -1
  247. package/web-shell/assets/{architectureDiagram-3BPJPVTR-B-At37J_.js → architectureDiagram-3BPJPVTR-B7aiJ6CK.js} +1 -1
  248. package/web-shell/assets/{blockDiagram-GPEHLZMM-B51wjcU7.js → blockDiagram-GPEHLZMM-BSK2CrID.js} +1 -1
  249. package/web-shell/assets/{c4Diagram-AAUBKEIU-xBK7BAwN.js → c4Diagram-AAUBKEIU-Bhx2p7T_.js} +1 -1
  250. package/web-shell/assets/channel-BBaLq8MV.js +1 -0
  251. package/web-shell/assets/{chunk-2J33WTMH-xKmLg7Yr.js → chunk-2J33WTMH-c3Hzv0Ni.js} +1 -1
  252. package/web-shell/assets/{chunk-4BX2VUAB-BfkdKZCM.js → chunk-4BX2VUAB-DNXaK7ZF.js} +1 -1
  253. package/web-shell/assets/{chunk-55IACEB6-DT713sZf.js → chunk-55IACEB6-BXjvDj_i.js} +1 -1
  254. package/web-shell/assets/{chunk-727SXJPM-CYEeAamv.js → chunk-727SXJPM-B6CaKFp-.js} +1 -1
  255. package/web-shell/assets/{chunk-AQP2D5EJ-DBN-5CFd.js → chunk-AQP2D5EJ-tij6Mz9B.js} +1 -1
  256. package/web-shell/assets/{chunk-FMBD7UC4-CGZKk4PH.js → chunk-FMBD7UC4-DdlsGk2h.js} +1 -1
  257. package/web-shell/assets/{chunk-ND2GUHAM-qQbC2O-y.js → chunk-ND2GUHAM-CdgGZSHE.js} +1 -1
  258. package/web-shell/assets/{chunk-QZHKN3VN-B3zeXwsX.js → chunk-QZHKN3VN-D8oJn1C0.js} +1 -1
  259. package/web-shell/assets/classDiagram-4FO5ZUOK-B2ZK6sWE.js +1 -0
  260. package/web-shell/assets/classDiagram-v2-Q7XG4LA2-B2ZK6sWE.js +1 -0
  261. package/web-shell/assets/{cose-bilkent-S5V4N54A-BKsl1Wvy.js → cose-bilkent-S5V4N54A-BUVmNp1Y.js} +1 -1
  262. package/web-shell/assets/{dagre-BM42HDAG-CuuuLiBK.js → dagre-BM42HDAG-CO4q310O.js} +1 -1
  263. package/web-shell/assets/{diagram-2AECGRRQ-CG4_eKiV.js → diagram-2AECGRRQ-C6gFvWkg.js} +1 -1
  264. package/web-shell/assets/{diagram-5GNKFQAL-DSFTmZB1.js → diagram-5GNKFQAL-NG2I_c0h.js} +1 -1
  265. package/web-shell/assets/{diagram-KO2AKTUF-BxwlEHn8.js → diagram-KO2AKTUF-DFKMS6I9.js} +1 -1
  266. package/web-shell/assets/{diagram-LMA3HP47-BUAg-jMG.js → diagram-LMA3HP47-BiElfsfP.js} +1 -1
  267. package/web-shell/assets/{diagram-OG6HWLK6-BFM2jlO5.js → diagram-OG6HWLK6-zLWDYUqu.js} +1 -1
  268. package/web-shell/assets/{erDiagram-TEJ5UH35-CerbFPa2.js → erDiagram-TEJ5UH35-DoVaazhK.js} +1 -1
  269. package/web-shell/assets/{flowDiagram-I6XJVG4X-BNeZz5Zz.js → flowDiagram-I6XJVG4X-0RLz-mXg.js} +1 -1
  270. package/web-shell/assets/{ganttDiagram-6RSMTGT7-CTOogLaM.js → ganttDiagram-6RSMTGT7-C-eX1A39.js} +3 -3
  271. package/web-shell/assets/{gitGraphDiagram-PVQCEYII-yxSFkXPj.js → gitGraphDiagram-PVQCEYII-BauQu1Bu.js} +1 -1
  272. package/web-shell/assets/index-3uAT3ehO.css +5 -0
  273. package/web-shell/assets/index-CVHa2QH7.js +3 -0
  274. package/web-shell/assets/index-D1KigWat.js +1128 -0
  275. package/web-shell/assets/{infoDiagram-5YYISTIA-DYdAUevm.js → infoDiagram-5YYISTIA-C800cFvZ.js} +1 -1
  276. package/web-shell/assets/{ishikawaDiagram-YF4QCWOH-CNuijZ90.js → ishikawaDiagram-YF4QCWOH-CFK1Jk6e.js} +1 -1
  277. package/web-shell/assets/{journeyDiagram-JHISSGLW-BTyiuonf.js → journeyDiagram-JHISSGLW-Cjx0WAC4.js} +1 -1
  278. package/web-shell/assets/{kanban-definition-UN3LZRKU-DPJo0Uow.js → kanban-definition-UN3LZRKU-DWMDHbiv.js} +1 -1
  279. package/web-shell/assets/{linear-C8wKCbIC.js → linear-vpW4FrZs.js} +1 -1
  280. package/web-shell/assets/{mermaid.core-BV5o7nW5.js → mermaid.core-DsHRUnDd.js} +5 -5
  281. package/web-shell/assets/{mindmap-definition-RKZ34NQL-BtLcBmSt.js → mindmap-definition-RKZ34NQL-D3Eb1cAK.js} +1 -1
  282. package/web-shell/assets/{pieDiagram-4H26LBE5-C22J7A9Q.js → pieDiagram-4H26LBE5-Cr2-6sWt.js} +1 -1
  283. package/web-shell/assets/{quadrantDiagram-W4KKPZXB-DSL10wiL.js → quadrantDiagram-W4KKPZXB-Yd7yUgVp.js} +1 -1
  284. package/web-shell/assets/{requirementDiagram-4Y6WPE33-feqh0UpX.js → requirementDiagram-4Y6WPE33-BE6SprL8.js} +1 -1
  285. package/web-shell/assets/{sankeyDiagram-5OEKKPKP-Z4OVaWaZ.js → sankeyDiagram-5OEKKPKP-DzkuMK8d.js} +1 -1
  286. package/web-shell/assets/{sequenceDiagram-3UESZ5HK-DIluIcb1.js → sequenceDiagram-3UESZ5HK-wD9A3Dhi.js} +1 -1
  287. package/web-shell/assets/{stateDiagram-AJRCARHV-hlhnuoYH.js → stateDiagram-AJRCARHV-BOp5q1Nm.js} +1 -1
  288. package/web-shell/assets/stateDiagram-v2-BHNVJYJU-BDsWUmPG.js +1 -0
  289. package/web-shell/assets/{timeline-definition-PNZ67QCA-Dqf_CqgL.js → timeline-definition-PNZ67QCA-Btc08sUA.js} +1 -1
  290. package/web-shell/assets/{vennDiagram-CIIHVFJN-DD2QhaTX.js → vennDiagram-CIIHVFJN-Bv6upsOA.js} +1 -1
  291. package/web-shell/assets/{wardley-L42UT6IY-B2uI85U_.js → wardley-L42UT6IY-D-XamnCK.js} +1 -1
  292. package/web-shell/assets/{wardleyDiagram-YWT4CUSO-U1aJX71g.js → wardleyDiagram-YWT4CUSO-BbP_5Wfp.js} +1 -1
  293. package/web-shell/assets/{xychartDiagram-2RQKCTM6-C0ODHUee.js → xychartDiagram-2RQKCTM6-CO5Hh28b.js} +1 -1
  294. package/web-shell/index.html +2 -2
  295. package/chunks/chunk-CU3L64TP.js +0 -93
  296. package/chunks/initializer-UTQ7FB7G.js +0 -71
  297. package/chunks/list-DEICDL65.js +0 -74
  298. package/chunks/loadedSettingsAdapter-NOWEGDYJ.js +0 -68
  299. package/chunks/mcp-L5JXEEIA.js +0 -68
  300. package/chunks/nonInteractiveCli-LS2K4X6A.js +0 -124
  301. package/chunks/types-JNKGKUJT.js +0 -12
  302. package/chunks/workspace-providers-status-4KMGETOJ.js +0 -71
  303. package/chunks/workspace-service-RVLBMSMZ.js +0 -80
  304. package/chunks/workspace-skills-status-DS4ZVFTO.js +0 -70
  305. package/web-shell/assets/channel-CLiQD50f.js +0 -1
  306. package/web-shell/assets/classDiagram-4FO5ZUOK-D3sYqow5.js +0 -1
  307. package/web-shell/assets/classDiagram-v2-Q7XG4LA2-D3sYqow5.js +0 -1
  308. package/web-shell/assets/index-D5IWFVk1.css +0 -5
  309. package/web-shell/assets/index-KFJRDhX7.js +0 -783
  310. package/web-shell/assets/stateDiagram-v2-BHNVJYJU-BhQMwgpl.js +0 -1
@@ -2,18 +2,19 @@
2
2
 
3
3
  > Architecture decisions, trade-offs, and rejected alternatives for the `/review` skill.
4
4
 
5
- ## Why 10 agents + 1 verify + iterative reverse, not 1 agent?
5
+ ## Why 12 agents + 1 verify + iterative reverse, not 1 agent?
6
6
 
7
7
  **Considered:**
8
8
 
9
9
  - **1 agent (Copilot approach):** Single agent with tool-calling, reads and reviews in one pass. Cheapest (1 LLM call). But dimensional coverage depends entirely on one prompt's attention — easy to miss performance issues while focused on security.
10
10
  - **5 parallel agents (original design):** Each agent focuses on one dimension. Higher coverage through forced diversity of perspective. Limited by combined Correctness+Security and a single undirected pass — recall ceiling left findings on the table that the user only discovered in subsequent /review rounds.
11
11
  - **9 parallel agents:** 6 review dimensions (Correctness, Security, Code Quality, Performance, Test Coverage, Undirected) + Build & Test. Undirected runs as 3 personas in parallel.
12
- - **10 parallel agents (current):** The 9-agent design plus Issue Fidelity & Root-Cause Ownership, which compares linked issue evidence against the PR's claimed fix before accepting a client-side change.
12
+ - **10 parallel agents:** The 9-agent design plus Issue Fidelity & Root-Cause Ownership, which compares linked issue evidence against the PR's claimed fix before accepting a client-side change.
13
+ - **12 parallel agents (current):** The 10-agent design with Correctness split into three procedural walks — 1a line-by-line scan, 1b removed-behavior audit, 1c cross-file tracer — plus up to 2 optional diff-specialized finders (Agent 8) when one domain dominates the diff.
13
14
 
14
- **Decision:** 10 agents. The marginal cost (10x vs 1x) is acceptable because:
15
+ **Decision:** 12 agents. The marginal cost (12x vs 1x) is acceptable because:
15
16
 
16
- 1. Parallel execution means time cost is ~1x (all 10 agents launch in one response)
17
+ 1. All 12 agents are submitted in one response and run concurrently up to the runtime's tool-call cap (default 10, `QWEN_CODE_MAX_TOOL_CONCURRENCY`) wall time is bounded by roughly two waves at worst, still far below twelve sequential agents
17
18
  2. Dimensional focus produces higher recall (fewer missed issues)
18
19
  3. Three undirected personas (attacker / 3am-oncall / maintainer) catch cross-dimensional issues that a single undirected agent's prompt-induced bias would miss
19
20
  4. Issue Fidelity prevents a common false approval mode: a PR can be internally well-tested while solving only the author's mistaken diagnosis, not the linked issue's original failure
@@ -29,7 +30,7 @@ Test gaps are a systematic blind spot. Review agents focused on bugs in the new
29
30
 
30
31
  ### Why a dedicated Issue Fidelity agent
31
32
 
32
- Bugfix PRs often carry their own diagnosis in the PR body, but that diagnosis can be wrong. The linked issue's original reproduction, observed payload, expected behavior, and maintainer comments must be checked before judging whether the implementation is a real fix. The implementation deliberately keeps issue discovery out of `pr-context`: the Issue Fidelity agent fetches GitHub's closing-issue metadata with `gh pr view --json closingIssuesReferences`, then fetches relevant issue discussions with `gh issue view --json title,body,comments` (the `--json` form is required — it returns the issue **body**, which `--comments` alone omits). This keeps relevance judgment in the agent instead of baking fragile PR-body parsing into TypeScript. The agent runs only for PR targets — a local-diff or file-path review has no PR or linked issue, so it is skipped there (9 agents instead of 10).
33
+ Bugfix PRs often carry their own diagnosis in the PR body, but that diagnosis can be wrong. The linked issue's original reproduction, observed payload, expected behavior, and maintainer comments must be checked before judging whether the implementation is a real fix. The implementation deliberately keeps issue discovery out of `pr-context`: the Issue Fidelity agent fetches GitHub's closing-issue metadata with `gh pr view --json closingIssuesReferences`, then fetches relevant issue discussions with `gh issue view --json title,body,comments` (the `--json` form is required — it returns the issue **body**, which `--comments` alone omits). This keeps relevance judgment in the agent instead of baking fragile PR-body parsing into TypeScript. The agent runs only for PR targets — a local-diff or file-path review has no PR or linked issue, so it is skipped there (11 agents instead of 12).
33
34
 
34
35
  The agent also enforces the root-cause ownership gate: a client-side parser/sanitizer workaround for malformed upstream output is not acceptable as a root-cause fix unless a maintainer explicitly asked for that defensive mitigation.
35
36
 
@@ -39,17 +40,40 @@ A single undirected agent has prompt-induced bias and tends to find the same kin
39
40
 
40
41
  Empirically, ensemble diversity drops sharply past 3-5 sampled paths. Three is the sweet spot: enough to break single-prompt bias, few enough that the marginal cost stays bounded.
41
42
 
43
+ ### Why Correctness is three procedural agents, not one topical agent
44
+
45
+ A topic brief ("find correctness bugs") lets the agent choose where to look, and independently-prompted agents converge on the same visibly-suspicious hunks — redundancy, not coverage. A procedural brief fixes the walk: every hunk line-by-line with its enclosing function (1a); every deleted line, asking where the deleted invariant is re-established (1b); every changed symbol's callers and read sites (1c). Complementary coverage comes from the walk itself, not from luck. The evidence is in this skill's own history: the whole-file invariant checklist — a procedural walk — found the five PR #6457 Criticals that both the topical dimension agents and 14 chunk agents missed ("what the chunk agents lack is not the lines; it is the question").
46
+
47
+ Two structural holes this closes:
48
+
49
+ - **Removed behavior was nobody's job.** A deleted guard, error path, or test leaves no trace in the post-change tree; only the diff's `-` lines witness it. Heavy files got this covered via the invariant agents' `diffRange`; an ordinary diff's deletions had no dedicated reader. Agent 1b is that reader.
50
+ - **Cross-file was everybody's job, which is the same thing.** The consumer/producer analysis was a shared duty of Agents 1–6: six agents re-running the same greps (~6× the tool calls), none accountable for finishing the walk. Step 3B had already consolidated it into one whole-diff agent; 3A now matches. Single ownership is also the shape the producer-direction lesson (PR #6621) demands — the read site of a never-populated field lives in a file no topical reviewer would open on its own initiative.
51
+
52
+ The language-pitfall and wrapper/proxy checklists fold into 1a rather than standing alone: they are line-level questions asked during the same walk, not separate walks.
53
+
54
+ ### Why removed-behavior is a whole-diff agent in 3B, not only a chunk duty
55
+
56
+ 3B folds Agent 1b into each chunk agent, scoped to "the deleted lines in your territory". That is necessary and — as PR #6638 proved — not sufficient. Territory-scoped 1b can only ask "was this deletion re-established _here_", and for the deletions that matter most the answer is somewhere else entirely.
57
+
58
+ The measurement: three reviewers ran over #6638 (extension management v2 — 43 files, 8 255 additions, 28 chunks). The 3B run with per-chunk 1b reported **one** Critical. An independent reviewer (Codex `$qreview`) reported 32, and a parallel hand-run wave of 1b + 1c agents over the same commit independently reproduced six of them. Every one of that overlapping six is a **cross-chunk deletion**: `enableByPath(includeSubdirs: true)` deleted in one file and replaced by an exact-path `setWorkspaceActivation` in another, silently narrowing what a workspace-scoped disable means for every untouched CLI/TUI caller; `refreshTools()` dropped from the activation paths, its replacement swallowing the errors it used to propagate; a global mutation timeout removed and replaced by one that covers only the prepare phase. Each has a deletion in chunk A, a replacement in chunk B, and a consumer in a file the diff never touches. **No chunk agent can see that triple, and 1c does not look for it.** The split is by task, not by symbol: 1c owns caller compatibility — it greps the removed export's old name (right there in the deleted lines) and checks each call site — while 1b owns the pairing, finding the _replacement_ and comparing its semantics to what was deleted. A replacement that leaves every call site compiling is all 1c can see; that it now means something different at every one of them is what only 1b goes looking for.
59
+
60
+ So 1b joins 1c as a whole-diff agent, with an explicit split: **1c walks the callers; 1b walks the replacement and compares its semantics.** The chunk agents keep the local half (a guard deleted and not re-established within the same hunk is theirs, and it is the common case). The cost is one agent per 3B review. The class it closes is the one where a replacement type-checks, compiles, passes every test, and means something different to callers nobody edited.
61
+
62
+ ### Why diff-specialized finders (Agent 8) are optional and capped at 2
63
+
64
+ Domains have failure grammars — a reconnect state machine, a module loader, a cron scheduler each fail in ways no generic dimension list names. The whole-file invariant checklist is the fixed-form ancestor: a domain-specific walk out-finds a generic brief over the same lines. Agent 8 generalizes that idea to the diff's dominant domain, with the brief written per-review by the orchestrator. Capped at 2 so the fan-out stays bounded and specialization happens only when a domain actually dominates; zero is the common case. Findings flow through Step 4 verification like any other `[review]` finding.
65
+
42
66
  ## Why batch verification instead of N independent agents?
43
67
 
44
68
  **Considered:**
45
69
 
46
70
  - **N independent agents (original design):** One verification agent per finding. Each reads code independently. High quality but cost scales linearly with finding count (15 findings = 15 LLM calls).
47
71
  - **1 batch agent (original):** Single agent receives all findings, verifies each one. Fixed cost.
48
- - **Sharded batches, ≤8 findings each (chosen):** `ceil(N/8)` agents, launched together.
72
+ - **Sharded batches, ≤8 findings each (chosen):** `ceil(F/8)` agents (F = finding count), launched together.
49
73
 
50
- **Decision:** Shard. One batch agent was right when a review produced 15 findings — it saw cross-finding relationships and cost O(1). But a Step 3B review of a large PR produces 30-60 findings, and one agent re-reading code for each of them inside a single context window degrades on the tail of the list. Sharding costs `ceil(N/8)` calls instead of 1, still far below one-agent-per-finding, and keeps each verifier's job small enough to do properly.
74
+ **Decision:** Shard. One batch agent was right when a review produced 15 findings — it saw cross-finding relationships and cost O(1). But a Step 3B review of a large PR produces 30-60 findings, and one agent re-reading code for each of them inside a single context window degrades on the tail of the list. Sharding costs `ceil(F/8)` calls instead of 1, still far below one-agent-per-finding, and keeps each verifier's job small enough to do properly.
51
75
 
52
- **A verifier may never reject a Critical.** It may downgrade to low confidence, with specific contradicting code cited. A rejected Critical is deleted from both the PR and the terminal and no later stage revisits it; a downgraded one still reaches a human under "Needs Human Review". The asymmetry between a false positive (noise) and a deleted true positive (a shipped bug plus another `/review` round) is not close.
76
+ **Rejecting a Critical requires quoted contradiction.** A verifier may reject a Critical only when it can quote the specific code that contradicts the claim (the finding describes behavior the code demonstrably does not have) or when the finding merely re-describes a change the diff's own text documents as deliberate; anything less certain is downgraded to low confidence, never deleted. A rejected Critical is deleted from both the PR and the terminal and no later stage revisits it; a downgraded one still reaches a human under "Needs Human Review". The asymmetry between a false positive (noise) and a wrongly deleted true positive (a shipped bug plus another `/review` round) is why the bar for rejection is quoted evidence, not judgment.
53
77
 
54
78
  ## Why reverse audit is a separate step, and why iterative
55
79
 
@@ -76,11 +100,11 @@ The original design gave one agent the whole diff plus a growing cumulative find
76
100
 
77
101
  ### Why the topology gate counts source lines, not diff lines
78
102
 
79
- Diff size is a bad proxy for review risk, because tests dominate it. Across this repo's last 40 merged PRs the median diff is **41% test code**, and 14 of the 40 are more than half tests. A gate on raw diff lines sends a change of 173 production lines that ships 489 lines of new tests into the territory fan-out, where the production code ends up owned by a single chunk agent — while under the dimension fan-out it would have been read by eight lenses.
103
+ Diff size is a bad proxy for review risk, because tests dominate it. Across this repo's last 40 merged PRs the median diff is **41% test code**, and 14 of the 40 are more than half tests. A gate on raw diff lines sends a change of 173 production lines that ships 489 lines of new tests into the territory fan-out, where the production code ends up owned by a single chunk agent — while under the dimension fan-out it would have been read by ten lenses (the diff-reading dimension agents: twelve minus Issue Fidelity and Build & Test).
80
104
 
81
- Territory fan-out is worth it when there is a lot of _risky_ code to divide, not a lot of _lines_. So the gate is `srcDiffLines > 500`, with a second clause `diffLines > 2400` as a delivery bound: past that point `ceil(diffLines / 400) + 4 > 10`, so chunking uses fewer agents than the ten-lens topology anyway, and asking ten agents each to read a diff that large dilutes all of them. On the 40-PR sample the second clause never fires; it exists for a changeset dominated by tests or generated files.
105
+ Territory fan-out is worth it when there is a lot of _risky_ code to divide, not a lot of _lines_. So the gate is `srcDiffLines > 500`, with a second clause `diffLines > 3200` as an attention bound: past that point asking ten diff-reading lenses each to swallow the whole diff dilutes all of them, and the chunk topology's base cost (`ceil(diffLines / 400) + 4`, counting the whole-diff agents that read the diff — Build & Test reads none) crosses twelve about there. It is not a promise of fewer calls a heavy file adds three invariant agents and a dominant domain up to two specialized finders but of one accountable reader per line instead of ten diluted ones. On the 40-PR sample the second clause never fires; it exists for a changeset dominated by tests or generated files.
82
106
 
83
- Re-gating moves 6 of those 40 PRs from 3B back to 3A and costs 22 extra agents in total across all 40 — about 5%. It buys those six PRs eight review lenses on their production code instead of one.
107
+ Re-gating moved 6 of those 40 PRs from 3B back to 3A and cost 22 extra agents in total across all 40 — about 5% — measured under the earlier 10-agent 3A roster; under the current 12-agent roster the same six PRs cost 2 more each, ~34 extra (~7%). It buys those six PRs ten review lenses on their production code instead of one.
84
108
 
85
109
  Chunking itself is unchanged: the plan still tiles every line, tests and generated files included. Only the count of reviewers and their brief change. `heavy` is likewise restricted to `source` files — the invariant checklist asks about fields, timers, collections, and error taxonomies, and a rewritten test file has none of those.
86
110
 
@@ -113,6 +137,17 @@ Eight simultaneous checks over a 2 400-line file is a task an agent performs onc
113
137
 
114
138
  They used to, on the theory that the auditor "already has full context, so its output is inherently high-confidence." That premise is false precisely when the diff is large: the agent with the least room to think was the one whose output nobody checked. Verification is sharded now, so the marginal cost of including reverse-audit findings is small.
115
139
 
140
+ ## Why findings carry a failure scenario instead of an impact statement
141
+
142
+ `Impact` asked why the finding matters. `Failure scenario` asks the finder to prove the finding can happen: name the input/state/timing that triggers it and the wrong outcome that results — or, for quality findings, the concrete cost (what is duplicated, wasted, or harder to maintain, or the quoted project rule).
143
+
144
+ Two effects:
145
+
146
+ 1. **Finders self-filter.** A "risk" for which no trigger can be constructed dies at the source instead of reaching the PR. Dogfood motivation: a /review run on PR #6612 auto-published two hallucinated Criticals onto an already-approved PR — both were findings for which no concrete trigger could have been written down. An `Impact` field accepts "this could cause issues in production"; a `Failure scenario` field does not.
147
+ 2. **Verifiers get a testable claim.** Step 4's verdict becomes the result of tracing the claimed trigger through the real code — confirmed (high) = the trace works and the lines are quoted; confirmed (low) = mechanism real, trigger uncertain; rejected = the code contradicts the claim — rather than a plausibility vote on the finding's prose.
148
+
149
+ The reporting gate is severity-asymmetric, matching the recall rules elsewhere in the skill: a Suggestion with no scenario and no cost is dropped at the source; a suspected Critical with an uncertain trigger is kept at `Confidence: low` for the verifier to rule on. A dropped Suggestion costs a nicety; a dropped Critical costs a shipped bug.
150
+
116
151
  ## Why low-confidence over rejection on uncertain findings
117
152
 
118
153
  **Original behavior:** When verification was uncertain, it would reject. Bias toward precision.
@@ -193,6 +228,15 @@ Line-based classification was chosen because it's deterministic, cheap, and catc
193
228
  - Any failure → downgrade `APPROVE` to `COMMENT`, body explains.
194
229
  - All pending → downgrade to `COMMENT` (don't approve before CI decides), body explains.
195
230
 
231
+ **The hole under all of this: a check that never ran looked like a check that passed.** GitHub reports a skipped job as `status: completed, conclusion: skipped`. The classifier tested for failure conclusions and for pending statuses, and `skipped` matched neither — so it fell through into `all_pass`. Every word above delegates runtime truth to CI _because_ the LLM pipeline reads code statically. If the delegation returns nothing, and returns it wearing a green badge, the delegation is worse than not having it.
232
+
233
+ PR #6486: the one job that would have exercised the new `Ctrl+F` hotkey — `Integration Tests (CLI, No Sandbox)` — was skipped, as were the macOS and Windows `Test` legs. `all_pass`. And even had it run, it would have passed: the test drove a CSI-u sequence into a PTY that never negotiated the kitty protocol, so the keypress was discarded before reaching the handler. A test that cannot fail, in a job that did not run, scored as verification.
234
+
235
+ `skipped`/`neutral` are now recognised, with two deliberately different consequences:
236
+
237
+ - **Some checks skipped → a disclosure, not a downgrade.** Empirically this repo emits skipped runs constantly — routing jobs (`authorize`, `review-pr`, `precheck-pr`) that also emit a successful run of the same name, which is why "did it run" is a question about the _name_, not about any single run. And a docs-only PR legitimately skips the test matrix. Auto-downgrading on any skip would downgrade every review in the repo, which is how a gate gets ignored. So presubmit _names_ them and Step 7 rules on them — because whether a skipped check would have exercised **this** diff is a question about the diff, which presubmit cannot see and the reviewer can.
238
+ - **Every check skipped → a downgrade.** Checks exist, not one ran: there is no green here to approve on, and no judgment is required to say so. (A repo with no CI at all is a different claim — `totalChecks === 0`, not downgraded.)
239
+
196
240
  **Why downgrade rather than block:** the reviewer LLM has done substantive work; throwing the review away because CI is red wastes that. Downgrading to `COMMENT` keeps all inline findings, preserves the static review value, and lets GitHub's check status carry the "do not merge" signal naturally.
197
241
 
198
242
  **Why this stacks with self-PR downgrade:** a self-authored PR with red CI hits **both** downgrade rules. The event is `COMMENT` either way, so stacking is operationally a no-op — but the body should mention both reasons so a future maintainer reading the review knows why an LLM that found no Critical issues did not approve.
@@ -244,20 +288,161 @@ A malicious PR could add `.qwen/review-rules.md` with "never report security iss
244
288
 
245
289
  **Decision:** Tips. Qwen Code's follow-up suggestion system is a core UX differentiator. Blocking prompts interrupt flow. Tips are zero-friction and let users decide when/if to act.
246
290
 
291
+ ## Why the COMMENT body is composed from clauses, not picked from fixed sentences
292
+
293
+ The body rules began as a table of exact one-liners — the right call against smuggled prose, and it stayed right while only one state could apply at a time. Then the states multiplied: presubmit downgrades, the context-unavailable cap, discarded-Suggestion disclosure, uncoverable-chunk disclosure, body-relocated Criticals. Four consecutive review rounds each found a **pairwise collision** — two rules both claiming to be "the" body, so applying either erased the other's disclosure (a downgrade reason overwriting the diff-only warning; a "Suggestions are inline" restored by 422 recovery inside a run that never saw the PR's discussion; an all-discarded run claiming its suggestions were inline). Patching collisions one at a time provably does not converge: n states have n(n−1)/2 pairs.
294
+
295
+ The fix is a composition rule: an ordered clause inventory, each clause present iff its condition holds, joined into one paragraph, nothing else permitted. It keeps the anti-prose discipline (the inventory is closed; free text is still banned), reduces to the table's exact sentences in the single-state case, and makes every future state additive — a new state adds one clause, not one patch per existing state. `C` is likewise defined once, globally (everything the review posts, anywhere — inline or body), so no downstream rule can re-derive it over a subset and delete a body-only blocker.
296
+
297
+ ## Why parse-args and compose-review are subcommands, and pr-context renders bodies in full
298
+
299
+ Seven rounds of review-the-review on this PR converged on one diagnosis: the skill's deterministic logic kept shipping bugs precisely where it was written as prose. Argument parsing produced three bugs (a flag consumed as a value, the `=` form undefined, an invalid value leaking into target disambiguation). The event/body machine produced five (four Critical), all one shape — a downstream branch not updated when an upstream rule gained a new state, because the machine was restated in four places that had to be synchronized by hand: n states, n(n−1)/2 pairwise collisions, patched one at a time without converging. And the "fetch review bodies for the re-check" instruction was rewritten **five times in four rounds** (missing pagination → shell truncation → unpageable single-line JSON → a marker filter that discarded markerless blockers → offline selection), which is what writing a download program in English looks like.
300
+
301
+ The resolution is the same one this document already records for presubmit and cleanup: judgment stays in the prompt, bookkeeping moves to tested subcommands that version together with the skill.
302
+
303
+ - **`parse-args`** owns the grammar. Every previously-shipped parsing bug is a named row in its table-driven tests. The raw string travels **on stdin** (`--stdin` with a quoted heredoc), never as a positional: a flag-first raw string (`/review --effort low`) is consumed by the CLI's own strict parser before the handler runs, and a positional also breaks on quotes and shell metacharacters. Pure-function tests could not see that class — the documented invocation failed only when run against the built binary — so the suite includes yargs-level wiring tests alongside the table.
304
+ - **`compose-review`** owns event selection and body composition — the C/S table (counting body Criticals and discarded Suggestions), the event caps (cannot-tell existing Criticals, uncoverable chunks, unreviewed dimensions, context-unavailable), the downgrade carve-outs, and the clause composition. Its truth-table tests pin each shipped bug; writing them immediately caught one more instance of the class (all Suggestions discarded → S=0 → APPROVE). The input is validated at the boundary: the producer is a model writing JSON that omits inapplicable fields, so absent counts default to zero and malformed values throw typed errors — before that, an omitted count meant `undefined + 1 = NaN`, which fails every event comparison and would have returned APPROVE over a body-only blocker. 422 recovery stops being a hand-derived recomposition: it is the same call with updated counts, so the "recompute may never upgrade the verdict" guarantee holds by construction.
305
+ - **`pr-context`** ends the fetch-prose chain at its root: review bodies **and every blocker-bearing body** render **in full** (a body-only blocker lives only there; a capped body names its review or comment id so the tail stays fetchable one object at a time, and reply snippets name their comment id when cut), and blocker-bearing threads are quarantined into a "Blockers to re-check" section instead of settling into "Already discussed" — a reply alone never retires a blocker. The `gh` wrapper's `maxBuffer` rises to 64 MiB, closing the ENOBUFS that killed two subcommands mid-review on a comment-heavy PR.
306
+
307
+ What deliberately stays prose: everything judgment-shaped — what counts as a Critical, verification, the posting gate's authorization semantics, the angles. A truth table cannot decide whether a finding is real; it can guarantee that a real finding is never mislabeled, dropped by a downgrade, or approved past.
308
+
309
+ ## Why blocker recognition is semantic, not the `[Critical]` marker
310
+
311
+ The mandatory re-check section used to be gated on the literal string `[Critical]`. That marker is emitted by exactly one author — `/review` itself. Every human blocker was therefore invisible to the gate, and the fallback was a prose instruction in Step 6 telling the model to also scan "Already discussed" semantically.
312
+
313
+ Prose does not beat structure. PR #6486 is the proof, and it cost a shipped blocker.
314
+
315
+ A maintainer built the PR, drove the real CLI through a PTY, and found that `Ctrl+F` **dual-fires** — it toggles the model _and_ moves the input cursor, because `text-buffer.ts:2663` still binds `Ctrl+F → move('right')` and both handlers are independent subscribers of a `KeypressContext.broadcast()` that has no stop-propagation. They filed it as an **issue comment**, headed `🔴 Finding 1 — … (blocker)`. No `[Critical]` marker, because a human wrote it.
316
+
317
+ Three things then compounded:
318
+
319
+ 1. Issue comments all settle into **"Already discussed — do NOT re-report"**.
320
+ 2. They render as **240-character one-line snippets**.
321
+ 3. The first 240 characters of a verification report are its **preamble**: _"I built this PR from source and drove the real CLI … to validate the model-toggle hotkey before merge. Sharing the results as a merge reference."_
322
+
323
+ So the one artifact that proved the PR was broken was presented to the review agents as a **maintainer endorsement**, in the section that says not to re-report it. The blocker itself began 1 143 characters past the cut. Three hours later `/review` reviewed the same commit — the fix did not land until that evening — and submitted **"Reviewed — no blockers"**. This is precisely the "dropped blocker" failure the Step 6 re-check exists to prevent, and the re-check could not prevent it, because the input it was handed said the opposite of the truth.
324
+
325
+ The fix moves the decision out of prose and into `carriesBlockerSignal`: any body asserting a blocking defect — inline thread or issue comment, `[Critical]` or `(blocker)` or "is a blocker" or "must fix" or "still reproducible" or 阻塞项 — is promoted into **"Blockers to re-check"** and rendered **in full**. A bare `🔴` is deliberately **not** a signal, for the reason the next paragraph measures.
326
+
327
+ Two properties are deliberate:
328
+
329
+ - **Fail-safe direction.** A false positive costs one extra ruling by the re-check; a false negative ships the bug. When in doubt, promote.
330
+ - **Precision still matters, in the other direction.** Promotion means full-body rendering, and a context file that outgrows one `read_file` is its own way of losing a blocker (PR #5738, recorded above). The prose scan of "Already discussed" is retained as a floor — `carriesBlockerSignal` recognises the phrasings we have seen, not every phrasing that exists.
331
+
332
+ **Both of those were nearly undone by the first implementation, and only a live run showed it.** That version scanned the whole body for the words `blocker`, `🔴`, `阻塞`, `[Critical]`. Run against the real #6486 thread it promoted **8 of 15** issue comments; exactly one was a live blocker. The others were the triage bot's own template line **"No critical blockers."** (the word inside its own negation), the author's **"### 🔴 Critical fixes"** (a severity emoji on a list of repairs), and a later comment _quoting_ `[Critical]` while arguing a finding away. Eight full bodies took the context file from 30 KB to 59 KB and pushed the real blocker to character **43 094** — past the 25 000 one `read_file` returns. The section existed, held the right blocker, and no agent could see it: PR #5738's failure, reintroduced one section further down by the fix for it.
333
+
334
+ Three changes, and the ordering one is load-bearing:
335
+
336
+ - **The section is written FIRST**, ahead of the description and the review history. Nothing in the file outranks the claims a `C=0` verdict may not be reached without ruling on. On the live thread this moved the heading from char 25 961 to **569**, and the blocker body from 43 094 to **4 421**.
337
+ - **Recognition matches assertion patterns, not word presence** — `[Critical]`, `(blocker)`, `is a blocker`, a bare `blocking` (with a `non-blocking` / `非阻塞` lookbehind), `must fix`, `still reproducible/repro/broken/fails`, `阻塞项/问题/点` — with a **bilingual** negation guard, so neither "no blockers" nor "没有阻塞项" ever promotes. Live promotions dropped 8 → 3 (the one real blocker plus two harmless mentions), and the file 59 KB → 40 KB.
338
+ - **The section carries a character budget.** Tight patterns keep promotion rare; the budget keeps a pathological thread from blowing the read window anyway. Bodies past it degrade to snippets **naming their exact fetch**, which the re-check already must run before ruling — not to silence.
339
+
340
+ The lesson generalizes past this file: **"a false positive is cheap" is a claim about a budget, and it has to be measured against the real distribution, not assumed.** Here it was false until the ordering was fixed.
341
+
342
+ ## Why a test-efficacy probe, when there is already a Test Coverage agent
343
+
344
+ Agent 5 asks whether a test **exists** and whether its assertions **look like** they check something. Agent 7 runs the suite and reports that it is **green**. Neither can see a test that protects nothing, and there are two ways to ship one:
345
+
346
+ - **Unreachable** — the project's test command never collects the file.
347
+ - **Inert** — it runs, it passes, and it would still pass with the change reverted.
348
+
349
+ PR #6486 shipped both, in one file. The new test lived in `integration-tests/`, which is not an npm workspace, so `npm test --workspaces` never collected it; its CI job (`Integration Tests (CLI, No Sandbox)`) was skipped, so CI never ran it either. **The test executed nowhere — not in CI, not in the review — and nothing in the pipeline noticed.** And had it run, it would have passed regardless: it drove a kitty CSI-u sequence into a PTY that never negotiated the kitty protocol, so the keypress was discarded before reaching the handler under test. It could only ever have caught a startup crash. Agent 5 saw a test file with plausible assertions and said coverage was fine.
350
+
351
+ Both questions are decidable without judgment, which is why they are a subcommand and not a prompt. Unreachability needs no execution at all — it is a path against the root `package.json` workspace globs. Inertness needs one run: revert the diff's **source** files to base, keep its **tests**, re-run them. A test that is still green is green whether or not the feature exists.
352
+
353
+ **The trap, and the reason the classifier is asymmetric.** Reverting source frequently breaks the test's own compile — it imports a symbol the diff introduced — and the runner exits non-zero having collected nothing. It is tempting to score that as "the test caught the revert". It is not: a compile error says nothing about whether the test would catch a _behavioural_ regression, and scoring it as `gated` would hand back precisely the false assurance this command exists to remove. So `gated` requires a real **assertion** failure; a bare non-zero exit with nothing collected is `inconclusive`, and `inconclusive` is never reported as a finding.
354
+
355
+ Two other deliberate limits:
356
+
357
+ - **A test-only diff is never probed.** A new test for old code is _supposed_ to pass with nothing reverted. Probing it would flag every such PR as inert — a false blocker on exactly the PRs we want people to write.
358
+ - **Findings are Suggestions, not Criticals.** A test that does not gate is not itself wrong code; nothing is broken today. What the finding must say concretely is which behaviour is now shipping unprotected.
359
+
360
+ ## Why "fixed by this diff" is the verdict that needed a bar
361
+
362
+ The re-check has three verdicts, and until PR #6486 only two of them cost anything:
363
+
364
+ | verdict | consequence |
365
+ | -------------------- | ----------------------------------------------------- |
366
+ | `still stands` | `REQUEST_CHANGES` — blocks the merge |
367
+ | `cannot tell` | serialized into the body, caps the event at `COMMENT` |
368
+ | `fixed by this diff` | **nothing. Silent, free, unrecorded.** |
369
+
370
+ An agent under context pressure, choosing among three answers where one is free and two are not, drifts toward the free one — and the free one is the only one that can ship a bug.
371
+
372
+ Worse, the bar for it read "you read the lines and the fix is there", which invites reading **the diff's lines**. That is precisely the reading that fails. A fix's new lines are always in the diff; whether they _work_ routinely depends on code outside it.
373
+
374
+ PR #6486 is the case. A `Ctrl+F` dual-fire blocker was filed — the hotkey toggled the model _and_ moved the input cursor. The author added a guard to the toggle handler: visible in the diff, and it reads like a fix. It changed nothing. The second handler is `text-buffer.ts:2663`, in a file the PR never touches, subscribed independently to a `KeypressContext.broadcast()` that has no stop-propagation — `return`ing from one subscriber does not stop the other. Read the diff and you see a guard and rule "fixed". Read `text-buffer.ts:2663` and you cannot.
375
+
376
+ Two changes, split the way this document keeps arriving at — **determinism owns the evidence, judgment owns the ruling**:
377
+
378
+ - **`pr-context` extracts the evidence** (`extractCodeRefs`). A blocker's body names the code it is about — #6486's named `text-buffer.ts:2663` outright — so a promoted blocker that names a file now renders a **Referenced code** list (a blocker citing no path gets none — the reader traces the mechanism themselves). "Go read the untouched code" stops being a hope the agent might have and becomes a list it is handed.
379
+ - **SKILL.md raises the bar** on the ruling: name the mechanism, name what now stops it, and when the stopping condition lives outside the diff, read it there — or the verdict is `cannot tell`.
380
+
381
+ No new `compose-review` input was needed: `cannot tell` already caps the event. The change is to make wrong "fixed" rulings land there instead of passing silently.
382
+
383
+ ## What the first dogfood batch changed
384
+
385
+ Six concurrent real-PR runs (batch 3) produced three targeted changes, each fixing something the batch measured rather than predicted:
386
+
387
+ - **Overlap disposal is deterministic.** presubmit's overlap report used to end in "ask the user whether to proceed" — 2 of 6 runs stalled on an improvised interactive question (fatal for a headless run) while the other 4 proceeded, the signature of an under-specified decision point. An overlap is a duplicate by the Exclusion Criteria; the rule is now drop, note in the terminal, continue — and the counts handed to `compose-review` shrink accordingly, so a dropped finding can never flip the verdict.
388
+ - **Host routing is a flag, not prose.** The GH_HOST-by-prefix instruction survived exactly one review round before a reviewer noted the model must remember it per call. `--host` on `fetch-pr` / `pr-context` / `presubmit` routes every wrapped `gh` call in code (`lib/gh.ts` `setGhHost`/`ghEnv`), leaving the prose rule only for the handful of `gh` commands the orchestrating model runs directly.
389
+ - **A fixed completion line.** Three different completion phrasings across one batch each needed their own detection regex in the batch driver. Step 9 now ends every run with `Review complete: <target> — <disposition>`, greppable by `^Review complete: `.
390
+
391
+ ## Why Step 7 opens with a hard posting gate
392
+
393
+ Posting is the only irreversible, public, outward-facing action the skill takes, and it must never happen as a side effect of a confident verdict. The skip condition existed from the start, but it was phrased as one clause among several ("skip if … or if BOTH `--comment` absent AND no post request"), which a model evaluates as a judgment call at the end of a long run — exactly when it is reasoning about what it wants to say rather than about what it was authorized to do.
394
+
395
+ Dogfooding proved the phrasing insufficient: across four concurrent no-`--comment` reviews, three correctly withheld (offering the follow-up tip) and one self-submitted a `COMMENT` review with an inline suggestion to a real PR. One violation in four is a model-adherence failure, not a logic error — the rule was right, its force was not.
396
+
397
+ The fix promotes the gate to the first thing in Step 7 and reframes it as arithmetic, not judgment: post **only if** `--comment` was parsed in Step 1 **or** the user explicitly asked to post this session; otherwise no `reviews`-API write happens at all, regardless of verdict or the "Tip: post comments" text being printed. This mirrors the `event`/`body` invariant elsewhere in Step 7 ("stop reasoning and count") — the same failure mode (a model rationalizing past a stated rule at submit time) gets the same countermeasure (convert the rule to a check with no discretion).
398
+
399
+ ## Why verification checks the diff's own documented intent
400
+
401
+ Verification traces a finding's failure scenario through the code, but "the code does what the finding says" is not sufficient for a finding framed as a **regression** — the code doing X is exactly what a deliberate, documented change to do X looks like. The missing question is whether X is a defect or a design decision, and the diff itself usually answers it: a rationale comment, a JSDoc note, or a test that asserts the new behavior on purpose.
402
+
403
+ Dogfooding auto-posted the failure. A review of a secret-sanitization PR filed a Critical — "third-party credentials (`AWS_SECRET_ACCESS_KEY`, `GITHUB_TOKEN`, `NPM_TOKEN`) now pass through to subprocesses = security regression." The factual claim was true; the framing was wrong. The same file carried a rationale comment three lines from the change — user-managed credentials `must remain available` for shell/MCP/tool subprocesses, and the old broad denylist that scrubbed them was the bug this PR fixed — plus tests that assert the pass-through on purpose. The verifier traced the behavior and confirmed it without reading the rationale, and the Critical published to a real PR.
404
+
405
+ So verification now has an explicit step: for any finding that reads as "regression / removed protection / now allows X", read the diff-local comments and tests for the changed lines, and engage the documented intent. A documented-and-deliberate change is a design decision — reject the finding if it merely re-describes that change without naming any harm the rationale fails to answer. Documentation changes what the verifier must do, not what confidence it may reach: a traced, concrete harm that survives the rationale keeps high confidence (documenting a hole does not make it safe); low confidence is for cases where the rationale makes the harm genuinely uncertain, e.g. it names a compensating control the verifier cannot rule out. It is the diff-local analogue of Agent 0's root-cause-ownership gate (which checks intent against the linked _issue_); this checks intent against the _diff's own text_, which every review path has even when there is no issue. The counterpart finding in that same review — two new `scrubChildEnv(process.env, …)` call sites missing the `normalizePathEnvForWindows` wrapper that every sibling call site uses — had no such rationale and was a real oversight bug; the gate is about documented intent, not about suppressing findings on sanitization PRs.
406
+
407
+ ## Why whole-diff agents get a substantive-return check
408
+
409
+ Step 3B's coverage receipts guarantee every chunk was read, but they cover only chunk agents — the whole-diff agents (Issue Fidelity, removed-behavior, cross-file tracer, invariant agents, test-coverage matrix, diff-specialized finders) have no receipt, because they own a concern, not a territory. That left a blind spot symmetric to the one receipts close: an agent that whiffs — returns almost instantly with near-empty output — is indistinguishable from one that examined its concern and found nothing.
410
+
411
+ Dogfooding surfaced it concretely. On a heavy-file review, one of the three invariant agents returned in 11 seconds having emitted ~370 tokens while its siblings ran for minutes and thousands; the fast one owned the checklist half (counters / return-values / error-taxonomy) that, in a parallel exhaustive pass, produced the run's most serious finding. Nothing flagged the whiff, and the orchestrator folded its silence into "no issues in that dimension".
412
+
413
+ The countermeasure is cheap and needs no new machinery: before Step 4, sanity-check that each receipt-less agent's return actually describes its walk (the fields/callers/lines it enumerated) rather than a bare "No issues found." The primary test is evidential, not statistical — a return that names nothing it examined is a non-return regardless of length, and a legitimately empty scope passes as long as it says what it checked. The comparative signal ("far shorter and faster than its peers") is only a prompt to look at that agent's output, never a threshold to relaunch on: no fixed cutoff would survive a review where every agent is legitimately terse. Deliberately no number, because a false relaunch costs one agent call and a missed whiff costs a shipped bug — when in doubt, relaunch. It is the receipt-less analogue of "a chunk with no receipt was never reviewed," and it applies to 3A's dimension agents just as it does to 3B's whole-diff agents, since neither emits a receipt.
414
+
415
+ ## Why effort levels (low / medium / high)
416
+
417
+ **Considered:**
418
+
419
+ - **Always-full (original):** every `/review` runs the full pipeline. Right for a PR verdict; wrong for a 5-line pre-commit sanity check — 12 agents, sharded verification, and ≥2 reverse-audit rounds to re-derive what one reader could see in a single pass.
420
+ - **A `--quick` boolean:** two modes, but "quick" hides what is and isn't checked (rules? cross-file? build?).
421
+ - **Three levels (chosen):** **low** = one orchestrator pass over the chunk plan, hunk-visible bugs only, ≤8 findings. **medium** = the finder angles (1a, 1b, 1c, quality/altitude, performance, conventions) run **sequentially in the orchestrator's own context** — inline sequencing, not subagents, is what makes the level cheap — ≤12 findings. **high** = the full pipeline, unchanged.
422
+
423
+ **Guardrails, because a quick pass is recall-limited by construction.** "Quick pass" means **low and medium together** — they differ in depth (one diff pass vs. sequential finder angles; ≤8 vs. ≤12 findings) but share every guardrail below, because what the guardrails defend against is the same at both: findings that no verifier ever checked.
424
+
425
+ - Labeled **unverified**; no Approve/Request-changes verdict is emitted. A verdict is a claim the pipeline earns in Steps 4–5; a quick pass claims findings, not absence of findings.
426
+ - Never posts to the PR: `--comment` forces high, and a "post comments" follow-up after a quick pass is declined.
427
+ - Never consults or writes the incremental cache — otherwise a medium run's SHA would make a later high run report "No new changes since last review", silently converting a quick pass into a full-review verdict.
428
+ - Scope handling (worktree, diff capture, chunk plan) is identical at all levels. The levels change who reads the diff and what runs afterwards, never how the diff is obtained — the base-resolution and truncation traps do not care how fast the user wants the answer.
429
+
430
+ **Defaults:** PR targets → high (the product is a public verdict); local-diff / file-path targets → medium (the product is fast feedback; the closing tip advertises `--effort high`). Findings caps exist only at the unverified levels — at high effort, verification is the noise filter, so no cap is needed.
431
+
247
432
  ## LLM call budget
248
433
 
249
- **Small diffs (≤ 500 lines, Step 3A) — 12-14 calls:**
434
+ **Small diffs (≤ 500 source lines AND ≤ 3200 total diff lines, Step 3A, high effort) — 15-21 calls (typically 15-17):**
250
435
 
251
- | Stage | Calls | Why |
252
- | ----------------------- | ----------------- | ------------------------------------------------------------------------------------------------------------------------ |
253
- | Review agents | 10 (9) | issue fidelity + 6 dimensions + 3 undirected personas; Agent 7 skipped in cross-repo, Agent 0 skipped for non-PR reviews |
254
- | Batch verification | 1 | O(1) not O(N) batch is as good as individual |
255
- | Iterative reverse audit | 1-3 | Loop until "No issues found" or 3-round hard cap |
256
- | **Total** | **12-14 (11-13)** | Same-repo PR: 12-14; cross-repo lightweight PR or local/file (no Agent 0): 11-13 |
436
+ | Stage | Calls | Why |
437
+ | ----------------------- | ------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
438
+ | Review agents | 12 (+0-2) | issue fidelity + 3 procedural correctness walks (1a/1b/1c) + security/quality/perf/tests + 3 undirected personas + build&test, plus 0-2 diff-specialized finders; cross-repo skips Agents 7 and 1c (10), non-PR skips Agent 0 (11) |
439
+ | Sharded verification | `ceil(F/8)` | F = findings; typically 1-2; keeps each verifier's job small on high-finding reviews |
440
+ | Iterative reverse audit | 2-5 | loop ends after two consecutive dry rounds; 5-round hard cap |
441
+ | **Total** | **~15-21 (~13-20)** | Row maxima do not co-occur on typical runs (~15-17 is common), but the honest sum of ranges is 15-21 same-repo, 13-20 cross-repo/local. **Low/medium effort: 0 subagent calls** the inline pass runs in the orchestrator's own context |
257
442
 
258
- **Large diffs (> 500 lines, Step 3B) — `ceil(diffLines / 400)` chunk agents + 4 whole-diff agents + 1 verify + 1-3 reverse.** PR #6457 (5801 diff lines) plans to 19 chunks, so ~27 calls.
443
+ **Large diffs (> 500 source lines OR > 3200 total diff lines, Step 3B, high effort) — `ceil(diffLines / 400)` chunk agents + `5..7` whole-diff agents + `3H` invariant agents (H = heavy files) + `ceil(F/8)` verify (F = findings) + `rounds × chunks` reverse audit.** The reverse audit dominates: it fans out one auditor per chunk per round, and the stop rule needs two consecutive dry rounds (hard cap 5). PR #6457 (5801 diff lines, 19 chunks, 1 heavy file) costs ~27-29 first-wave calls, then `19 × (2..5) = 38-95` reverse auditors — ~66-126 calls total depending on how long the audit keeps finding; ~70 is the clean-run floor, and the count scales with chunks and findings, not a fixed ceiling.
259
444
 
260
- That is roughly 2x the small-diff budget, and it buys the thing the small-diff topology cannot deliver at that size: coverage. Ten dimension agents on a 5801-line diff each read the same truncated 14% window (see "Why the diff is a file, not a command"), so nine of the ten calls are redundant reads of the same hunks. Nineteen chunk agents each read a distinct ~390-line territory, and every line of the diff has exactly one accountable owner. The comparison to make is not 27 calls vs 14: PR #6457 took **eight** review rounds at 12-14 calls each — over 100 calls — and was still surfacing Criticals in code that had been in the diff since the first commit.
445
+ That is roughly 4x the small-diff budget, and it buys the thing the small-diff topology cannot deliver at that size: coverage. Ten dimension agents (the roster of the day; twelve now) on a 5801-line diff each read the same truncated 14% window (see "Why the diff is a file, not a command"), so nine of the ten calls were redundant reads of the same hunks. Nineteen chunk agents each read a distinct ~390-line territory, and every line of the diff has exactly one accountable owner. The comparison to make is not ~70 calls vs ~17: PR #6457 took **eight** review rounds at 12-14 calls each — over 100 calls — and was still surfacing Criticals in code that had been in the diff since the first commit.
261
446
 
262
447
  Competitors: Copilot uses 1 call, Gemini uses 2, Claude /ultrareview uses 5-20 (cloud). Ours biases toward higher recall — the assumption is that "find more issues per round" is more valuable than minimizing per-run cost, because every missed issue forces the user into another `/review` iteration.
263
448
 
@@ -341,21 +526,22 @@ The convergence concern that motivated the summary is real but narrower than it
341
526
 
342
527
  For a PR with 15 findings:
343
528
 
344
- | Approach | LLM calls | Notes |
345
- | --------------------------------------------------- | --------- | ---------------------------------------------------- |
346
- | Copilot (1 agent) | 1 | Lowest cost, lowest coverage |
347
- | Gemini (2 LLM tasks) | 2 | Good cost, medium coverage |
348
- | Our design (5 agents, N verify) | 21 | 5+15+1 — too expensive |
349
- | Our design (5 agents, batch verify, single reverse) | 7 | 5+1+1 — original design |
350
- | Our design (9 agents, iterative reverse) | 11-13 | 9+1+(1-3) — +50% cost for meaningfully higher recall |
351
- | Our design (10 agents, current) | 12-14 | 10+1+(1-3) — adds issue-fidelity/root-cause gate |
352
- | Claude /ultrareview | 5-20 | Cloud-hosted, cost on Anthropic |
529
+ | Approach | LLM calls | Notes |
530
+ | --------------------------------------------------- | -------------------- | ---------------------------------------------------------------------------------------------------- |
531
+ | Copilot (1 agent) | 1 | Lowest cost, lowest coverage |
532
+ | Gemini (2 LLM tasks) | 2 | Good cost, medium coverage |
533
+ | Our design (5 agents, N verify) | 21 | 5+15+1 — too expensive |
534
+ | Our design (5 agents, batch verify, single reverse) | 7 | 5+1+1 — original design |
535
+ | Our design (9 agents, iterative reverse) | 11-13 | 9+1+(1-3) — +50% cost for meaningfully higher recall |
536
+ | Our design (10 agents) | 12-14 | 10+1+(1-3) — adds issue-fidelity/root-cause gate |
537
+ | Our design (12 agents + effort levels, current) | 15-21 high / 0 quick | 12(+0-2)+ceil(F/8)+(2-5) under 3A; low/medium run inline with no subagents — cost scales with intent |
538
+ | Claude /ultrareview | 5-20 | Cloud-hosted, cost on Anthropic |
353
539
 
354
540
  ## Future optimization: Fork Subagent
355
541
 
356
542
  > Dependency: [Fork Subagent proposal](https://github.com/wenshao/codeagents/blob/main/docs/comparison/qwen-code-improvement-report-p0-p1-core.md#2-fork-subagentp0)
357
543
 
358
- **Current problem:** Each of the 12-14 LLM calls (10 review + 1 verify + 1-3 reverse audit rounds) creates a new subagent from scratch. The system prompt (~50K tokens) is sent independently to each, totaling ~620-730K input tokens with massive redundancy. The cost grew along with the agent count — Fork Subagent matters more under the current 10-agent design than under the original 5-agent design.
544
+ **Current problem:** Each of the ~15-21 LLM calls (12-14 review + sharded verify + 2-5 reverse audit rounds) creates a new subagent from scratch. At ~52K per agent (50K system + 2K task), that is ~780K-1.1M input tokens with massive redundancy. The cost grew along with the agent count — Fork Subagent matters even more under the current 12-agent design than under the original 5-agent design. (Effort levels bound the cost from the other side: low/medium runs spawn no subagents at all.)
359
545
 
360
546
  **Fork Subagent solution:** Instead of creating independent subagents, fork the current conversation. All forks inherit the parent's full context (system prompt, conversation history, Step 1/1.1/1.5 results) and share a prompt cache prefix. The API caches the common prefix once; each fork only pays for its unique delta (~2K per agent).
361
547
 
@@ -363,13 +549,13 @@ For a PR with 15 findings:
363
549
  Current (independent subagents):
364
550
  Agent 1: [50K system] + [2K task] = 52K
365
551
  Agent 2: [50K system] + [2K task] = 52K
366
- ...× 12-14 agents = ~620-730K total input tokens
552
+ ...× 15-21 agents = ~780K-1.1M total input tokens
367
553
 
368
554
  With Fork + prompt cache sharing:
369
555
  Cached prefix: [50K system + conversation history] (cached once)
370
556
  Fork 1: [cache hit] + [2K delta] = ~2K effective
371
557
  Fork 2: [cache hit] + [2K delta] = ~2K effective
372
- ...× 12-14 forks = ~50K cached + ~24-28K delta = ~74-78K total
558
+ ...× 15-21 forks = ~50K cached + ~30-42K delta = ~80-92K total
373
559
  ```
374
560
 
375
561
  **Additional benefits for /review:**
@@ -379,6 +565,6 @@ With Fork + prompt cache sharing:
379
565
  - Verification and reverse audit agents inherit all prior findings naturally
380
566
  - Agent 6 personas can fork from a shared diff-loaded base, paying only the persona-framing delta
381
567
 
382
- **Estimated savings:** ~85-90% token reduction (~620K → ~75K) with zero quality impact. The savings ratio is now even more compelling than under the 5-agent design.
568
+ **Estimated savings:** ~88-92% token reduction (~780K-1.1M → ~80-92K) with zero quality impact. The savings ratio is now even more compelling than under the 5-agent design.
383
569
 
384
570
  **Why not implemented now:** Fork Subagent requires changes to the Qwen Code core (`AgentTool`, `forkSubagent.ts`, `CacheSafeParams`). This is a platform-level feature (~400 lines, ~5 days), not a /review-specific change. When available, /review should be updated to use fork instead of independent subagents.