@qwen-code/qwen-code 0.21.2 → 0.21.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (351) hide show
  1. package/bundled/qc-helper/docs/configuration/settings.md +41 -38
  2. package/bundled/qc-helper/docs/features/channels/github.md +7 -0
  3. package/bundled/qc-helper/docs/features/channels/gitlab.md +6 -0
  4. package/bundled/qc-helper/docs/features/code-review.md +78 -23
  5. package/bundled/qc-helper/docs/features/commands.md +2 -0
  6. package/bundled/qc-helper/docs/features/hooks.md +115 -18
  7. package/bundled/qc-helper/docs/features/memory.md +27 -0
  8. package/bundled/qc-helper/docs/features/skills.md +46 -1
  9. package/bundled/qc-helper/docs/features/sub-agents.md +42 -9
  10. package/bundled/qc-helper/docs/features/tool-use-summaries.md +7 -7
  11. package/bundled/qc-helper/docs/qwen-serve.md +21 -1
  12. package/bundled/qc-helper/docs/reference/keyboard-shortcuts.md +13 -13
  13. package/bundled/qc-helper/docs/support/troubleshooting.md +6 -1
  14. package/bundled/review/DESIGN.md +167 -32
  15. package/bundled/review/SKILL.md +156 -32
  16. package/chunks/{MaxSizedBox-TQBQ247P.js → MaxSizedBox-OJR636GP.js} +40 -38
  17. package/chunks/{StandaloneSessionPicker-2HPTKCLX.js → StandaloneSessionPicker-AMMBF5UY.js} +62 -59
  18. package/chunks/{acp-startup-profiler-2C4A5ZTJ.js → acp-startup-profiler-3RICG2YW.js} +2 -2
  19. package/chunks/{acpAgent-PS6EDGH3.js → acpAgent-M57LVIQA.js} +662 -560
  20. package/chunks/{agent-VDOUB35F.js → agent-PE37ARLM.js} +36 -34
  21. package/chunks/{agent-headless-EO53EFS2.js → agent-headless-45ASME4R.js} +36 -34
  22. package/chunks/{anthropicContentGenerator-OFHSZJBF.js → anthropicContentGenerator-WCBYB6JG.js} +337 -27
  23. package/chunks/{artifact-tool-RDN2W7ND.js → artifact-tool-TBOYEKTY.js} +2 -2
  24. package/chunks/{askUserQuestion-OZERWH5R.js → askUserQuestion-23Y5WZFJ.js} +2 -2
  25. package/chunks/{bridge-W7AZSRZQ.js → bridge-K3ZRPO4D.js} +42 -40
  26. package/chunks/{ca-6W3SK7OQ.js → ca-4OJN75WT.js} +46 -5
  27. package/chunks/{channel-management-service-5S7YCWX5.js → channel-management-service-IXJ2MZTS.js} +7 -7
  28. package/chunks/{channel-settings-store-JFVRSNER.js → channel-settings-store-PEAR3IME.js} +43 -41
  29. package/chunks/{channel-worker-group-OVWGYMZS.js → channel-worker-group-NDG4WOOZ.js} +8 -8
  30. package/chunks/{channel-worker-manager-H55GKOQA.js → channel-worker-manager-ZCW6UW34.js} +8 -8
  31. package/chunks/{channel-worker-supervisor-CABBXYHK.js → channel-worker-supervisor-ATGFTZRN.js} +4 -4
  32. package/chunks/{chunk-W3UKBMSI.js → chunk-22YFMJBW.js} +84 -8
  33. package/chunks/{chunk-6MPDY6ZG.js → chunk-24ER7UV4.js} +1 -1
  34. package/chunks/{chunk-QDK4F2N5.js → chunk-2DNKEBXY.js} +1 -1
  35. package/chunks/{chunk-D6MDRSZ6.js → chunk-2LZCLQIS.js} +22 -11
  36. package/chunks/{chunk-JYBU6YUE.js → chunk-2M5N3DHY.js} +1 -1
  37. package/chunks/{chunk-KR32NIAN.js → chunk-2MI7ZTRT.js} +10 -10
  38. package/chunks/{chunk-LHKCY7DY.js → chunk-2UISUFHQ.js} +3 -3
  39. package/chunks/{chunk-JXVC2PL3.js → chunk-33FNCQSY.js} +2 -2
  40. package/chunks/{chunk-NBD76HUW.js → chunk-3LGZEYOF.js} +2 -2
  41. package/chunks/{chunk-CO2U743O.js → chunk-3MWNROHB.js} +4 -4
  42. package/chunks/{chunk-E2BHWIDT.js → chunk-3RXUYQJI.js} +6 -2
  43. package/chunks/{chunk-E5Z2AVNV.js → chunk-3SH53ANB.js} +0 -11
  44. package/chunks/{chunk-LYI2PZDK.js → chunk-3XF4WT3B.js} +2 -2
  45. package/chunks/{chunk-N4ACHWNH.js → chunk-432XCUHF.js} +1 -1
  46. package/chunks/{chunk-FNDXMN7I.js → chunk-4AGGSHCH.js} +774 -205
  47. package/chunks/{chunk-DUNI4EBW.js → chunk-4K3RV3OA.js} +80 -50
  48. package/chunks/{chunk-L2E5GFAG.js → chunk-4LKVQFHM.js} +2 -2
  49. package/chunks/{chunk-YDLJGZHL.js → chunk-4NFY2S7N.js} +0 -33
  50. package/chunks/{chunk-JLZ3RBQG.js → chunk-5JPVOYWA.js} +144 -11
  51. package/chunks/{chunk-6TAASXWL.js → chunk-5VYFXHY2.js} +3 -3
  52. package/chunks/{chunk-EU5TBFJI.js → chunk-6NE5NTHJ.js} +2 -2
  53. package/chunks/{chunk-XQK4XXTN.js → chunk-6O5DVH27.js} +1 -1
  54. package/chunks/{chunk-2RFKHEXL.js → chunk-6PLDPT2C.js} +1 -1
  55. package/chunks/{chunk-3APUUXRF.js → chunk-A5YZ3FA2.js} +2 -1
  56. package/chunks/{chunk-SHN4KHDK.js → chunk-A6SX5TAH.js} +3 -3
  57. package/chunks/{chunk-SZ3KHTDR.js → chunk-AKA24BGO.js} +1 -1
  58. package/chunks/{chunk-Q3T5KA2V.js → chunk-BFHXDAR7.js} +1 -1
  59. package/chunks/{chunk-L55JWS76.js → chunk-BGXZBI5B.js} +1 -1
  60. package/chunks/{chunk-P5VC5USY.js → chunk-BKGA75MU.js} +1 -1
  61. package/chunks/{chunk-2KKCCLPT.js → chunk-BKXTHMOZ.js} +1 -1
  62. package/chunks/{chunk-V4XNUDPE.js → chunk-BSKUTHLO.js} +15 -7
  63. package/chunks/{chunk-MM7TH7ZM.js → chunk-CFDR3GNL.js} +18 -3
  64. package/chunks/{chunk-E2FN6SM7.js → chunk-CIUUYB23.js} +5 -5
  65. package/chunks/{chunk-A7Y5H4TX.js → chunk-D2GPKXCQ.js} +7 -7
  66. package/chunks/{chunk-JCJ7LKNA.js → chunk-DIEFUBV6.js} +1644 -1570
  67. package/chunks/{chunk-GL4JOZXU.js → chunk-DSGJLHM3.js} +4 -4
  68. package/chunks/{chunk-675H6PLP.js → chunk-DYA25Y7U.js} +1 -1
  69. package/chunks/{chunk-3I6DUGQD.js → chunk-E4WHKETF.js} +32 -4
  70. package/chunks/{chunk-CU47KXB5.js → chunk-EGYJRIDF.js} +6 -6
  71. package/chunks/{chunk-6CQWRWUC.js → chunk-EIMKVOSF.js} +3 -3
  72. package/chunks/{chunk-QTERJVPZ.js → chunk-EKTCKZV5.js} +1 -0
  73. package/chunks/{chunk-NHFEMWAZ.js → chunk-FF75UUB2.js} +1 -1
  74. package/chunks/{chunk-L32VV3ZJ.js → chunk-FG63DZQA.js} +12 -12
  75. package/chunks/{chunk-FCYYAB4S.js → chunk-FHPKHXHT.js} +3 -3
  76. package/chunks/{chunk-GJ2GBPBD.js → chunk-FXVHVCNA.js} +5 -5
  77. package/chunks/{chunk-56Z7QEAQ.js → chunk-GMSBXEH4.js} +16 -16
  78. package/chunks/chunk-GOFAQQZA.js +132 -0
  79. package/chunks/{chunk-YUFIEIWC.js → chunk-GTM6IHBB.js} +2 -2
  80. package/chunks/{chunk-HIVKI6O6.js → chunk-H35Q5CEG.js} +9 -7
  81. package/chunks/{chunk-FJ5WUMUQ.js → chunk-H36SETQS.js} +1 -1
  82. package/chunks/{chunk-AUOULH6Q.js → chunk-H6FIBBXK.js} +1610 -2196
  83. package/chunks/{chunk-VEMNQMHA.js → chunk-H6ZYP74A.js} +5 -13
  84. package/chunks/{chunk-TTJZAIFP.js → chunk-HEBW74ZW.js} +3 -3
  85. package/chunks/{chunk-6FMJLI5M.js → chunk-HFVW54NL.js} +41 -6
  86. package/chunks/{chunk-44F4LBWZ.js → chunk-HIXVWF7I.js} +56 -18
  87. package/chunks/{chunk-EQNUDTL6.js → chunk-HUIQ4FYC.js} +2 -2
  88. package/chunks/{chunk-5NNHICSA.js → chunk-I7JDGNG6.js} +3 -3
  89. package/chunks/{chunk-PVNQRENX.js → chunk-IF6K42YJ.js} +1 -1
  90. package/chunks/{chunk-PJOHNFIY.js → chunk-IWGTGH4G.js} +5 -5
  91. package/chunks/{chunk-B4AJNNOL.js → chunk-J26HHFI5.js} +1 -1
  92. package/chunks/{chunk-CNTR4GR6.js → chunk-JELIWHXD.js} +1 -1
  93. package/chunks/{chunk-V25K5ID7.js → chunk-K2J5JZ6B.js} +1 -1
  94. package/chunks/{chunk-IGTHU3T4.js → chunk-KKN2IE7Q.js} +1 -1
  95. package/chunks/{chunk-TZTAV5RH.js → chunk-KVUBYJDM.js} +2 -2
  96. package/chunks/{chunk-3F2WDAV6.js → chunk-KXWERR6P.js} +14 -14
  97. package/chunks/{chunk-JJPKFJBT.js → chunk-L7FTDKCD.js} +3 -3
  98. package/chunks/chunk-LBOVL47Y.js +29 -0
  99. package/chunks/{chunk-FS674JW4.js → chunk-LBWNSLJK.js} +3 -3
  100. package/chunks/{chunk-I5BT7ADR.js → chunk-LGT3YMLN.js} +4 -4
  101. package/chunks/{chunk-UV6KF7T3.js → chunk-LZV2BZPL.js} +3 -3
  102. package/chunks/{chunk-P755N4CS.js → chunk-M5LIY7Y6.js} +6 -6
  103. package/chunks/{chunk-YGE3IXJP.js → chunk-M744OFVE.js} +1 -1
  104. package/chunks/{chunk-NTDH4IML.js → chunk-MCXMWLFT.js} +7 -7
  105. package/chunks/chunk-MISUEWB4.js +8749 -0
  106. package/chunks/{chunk-TEGJSJDF.js → chunk-MOHAWIAW.js} +2 -2
  107. package/chunks/{chunk-KXDONCPU.js → chunk-MTR4JWQJ.js} +4 -4
  108. package/chunks/{chunk-RSVPYAGK.js → chunk-MUEJ4R3E.js} +5 -5
  109. package/chunks/{chunk-PCK4B6LR.js → chunk-N42C2FU3.js} +1 -1
  110. package/chunks/{chunk-WWCMSZII.js → chunk-NW57GHNL.js} +71 -7
  111. package/chunks/{chunk-PQNHRFJC.js → chunk-NYX53SV4.js} +0 -2
  112. package/chunks/{chunk-QARJNL2V.js → chunk-O2ZYHG33.js} +5 -4
  113. package/chunks/{chunk-V7RNNPGC.js → chunk-P5Y23G2L.js} +18 -1
  114. package/chunks/{chunk-GURGAJTX.js → chunk-PDDYAXL2.js} +4 -1
  115. package/chunks/chunk-PWCQRLM5.js +1170 -0
  116. package/chunks/{chunk-VMX5RZQH.js → chunk-Q3G5L5KF.js} +18 -9
  117. package/chunks/{chunk-DQ2O5QOJ.js → chunk-QDWLUXKP.js} +1 -1
  118. package/chunks/chunk-QL4TN4HS.js +59 -0
  119. package/chunks/{chunk-E6CJ2SMM.js → chunk-QMBEW7EW.js} +2 -2
  120. package/chunks/{chunk-U4VTUMMK.js → chunk-QONSRAEV.js} +1 -1
  121. package/chunks/{chunk-DEEABGH2.js → chunk-QPSYIV2O.js} +3320 -536
  122. package/chunks/{chunk-Q5WZTTWE.js → chunk-QWDVZH24.js} +6 -8
  123. package/chunks/{chunk-2LEOZTBT.js → chunk-QWW6I2UO.js} +111 -26
  124. package/chunks/{chunk-LOY7RWQG.js → chunk-QYPT3QUD.js} +3 -3
  125. package/chunks/{chunk-6PQPKTXP.js → chunk-R7XNRHYD.js} +1 -1
  126. package/chunks/{chunk-7SJN2IWT.js → chunk-RCXCGZG6.js} +5 -5
  127. package/chunks/{chunk-HX6RLYUO.js → chunk-RL6J3BPE.js} +4 -4
  128. package/chunks/{chunk-Z3S5EMOW.js → chunk-RPYYTL7E.js} +60 -12
  129. package/chunks/{chunk-MB32MS4U.js → chunk-RR224BUP.js} +215 -174
  130. package/chunks/{chunk-H6I4PGRV.js → chunk-SKSF3LGQ.js} +4 -4
  131. package/chunks/{chunk-77V3VIKL.js → chunk-SQ3YD5MI.js} +9 -9
  132. package/chunks/{chunk-CBBP4BMC.js → chunk-SRY26X4H.js} +72 -206
  133. package/chunks/{chunk-RDJOB6S3.js → chunk-TXABHVZD.js} +1 -1
  134. package/chunks/{chunk-R2XWGR75.js → chunk-TYERNYH4.js} +2 -2
  135. package/chunks/chunk-TZTBZ25J.js +222 -0
  136. package/chunks/{chunk-45IOTAUS.js → chunk-U25CMJYY.js} +1 -1
  137. package/chunks/{chunk-7ULQIS27.js → chunk-UEJESCS7.js} +1 -1
  138. package/chunks/{chunk-HC45LTEN.js → chunk-UH7Z7Y45.js} +1 -1
  139. package/chunks/{chunk-CPZAH673.js → chunk-UKZEXS6K.js} +2 -0
  140. package/chunks/{chunk-54YCNSK5.js → chunk-UQDFEU47.js} +4113 -10152
  141. package/chunks/{chunk-XP4L6KS3.js → chunk-URN76TKR.js} +2 -2
  142. package/chunks/{chunk-XUGZYLDH.js → chunk-UWSXFI6E.js} +1 -1
  143. package/chunks/{chunk-AOHY3QJ3.js → chunk-VA5NLVPT.js} +3 -3
  144. package/chunks/{chunk-PZEXDV3H.js → chunk-VUNRNRN6.js} +71 -43
  145. package/chunks/{chunk-FANRZEAJ.js → chunk-VWOVHOTT.js} +1 -1
  146. package/chunks/{chunk-QTHQ6Z5X.js → chunk-WK4TJUEA.js} +1 -1
  147. package/chunks/{chunk-ZPZ7ZJEE.js → chunk-WUC6ZFXO.js} +1 -1
  148. package/chunks/{chunk-7KAXXBER.js → chunk-WY2ENRDN.js} +3 -3
  149. package/chunks/{chunk-TIRRPCFS.js → chunk-X6QNZOTC.js} +1 -1
  150. package/chunks/{chunk-KWHKHSIS.js → chunk-X6WUYMOV.js} +8 -8
  151. package/chunks/{chunk-IIUGIWWY.js → chunk-XC2VAFVZ.js} +1 -1
  152. package/chunks/{chunk-D4HTUMGX.js → chunk-XC656O42.js} +3 -3
  153. package/chunks/{chunk-RG64CDT2.js → chunk-XMJO6C44.js} +131 -48
  154. package/chunks/{chunk-ZMOJ7ISP.js → chunk-XUMAD5IT.js} +1 -1
  155. package/chunks/{chunk-THVU3RNH.js → chunk-YAVY33G4.js} +1 -1
  156. package/chunks/{chunk-YANIWTRZ.js → chunk-YPUHMQXV.js} +1 -1
  157. package/chunks/{chunk-QCXGC7QB.js → chunk-YU6TBRTA.js} +3 -3
  158. package/chunks/{chunk-OFZKFPYG.js → chunk-YWMDTMAU.js} +3 -3
  159. package/chunks/{chunk-DSFHVTWD.js → chunk-YXEVDA66.js} +2 -2
  160. package/chunks/{chunk-JE2CEOBO.js → chunk-ZGSUFIL5.js} +4 -4
  161. package/chunks/{chunk-BBIQAX4V.js → chunk-ZJES4C3I.js} +30 -18
  162. package/chunks/{computer-use-5QWCFNKD.js → computer-use-UJMK6JQU.js} +36 -34
  163. package/chunks/{config-utils-WADIDWGG.js → config-utils-7COTJMWK.js} +4 -4
  164. package/chunks/{contextCommand-4Z7T7TUM.js → contextCommand-4YABGF4S.js} +40 -38
  165. package/chunks/{core-runtime-6N4EFC2Q.js → core-runtime-7M4VM3DA.js} +36 -34
  166. package/chunks/{create-sub-session-AEREHZBQ.js → create-sub-session-HOFDW4N7.js} +39 -37
  167. package/chunks/{create-sub-session-62YZWVAY.js → create-sub-session-J32WC64U.js} +2 -2
  168. package/chunks/{cron-create-6E7GIQTG.js → cron-create-HMFSHIAR.js} +4 -4
  169. package/chunks/{cron-delete-JD26XY3W.js → cron-delete-ZJFHOX57.js} +4 -4
  170. package/chunks/{cron-list-CUWFB44G.js → cron-list-ZEX2NBMN.js} +4 -4
  171. package/chunks/{daemon-PCNCI5YU.js → daemon-CCIX4BMG.js} +47 -6
  172. package/chunks/{daemon-status-provider-DDF7C5M7.js → daemon-status-provider-W2U5Q3J5.js} +47 -45
  173. package/chunks/{daemon-trust-policy-Q4KSOJTR.js → daemon-trust-policy-IFICWPXP.js} +42 -40
  174. package/chunks/{daemon-trust-policy-monitor-V6FTE3IY.js → daemon-trust-policy-monitor-GX7EO6FW.js} +42 -40
  175. package/chunks/{de-SY6O76BF.js → de-A6TI4LBB.js} +46 -5
  176. package/chunks/{deferred-core-runtime-GHJ5GQHK.js → deferred-core-runtime-U36OOOKE.js} +36 -34
  177. package/chunks/display-image-VS3TXYQH.js +184 -0
  178. package/chunks/{dist-ZHXOBXDN.js → dist-3BMEUGTG.js} +1 -1
  179. package/chunks/{dist-26BIMQT6.js → dist-4KAX7CCL.js} +289 -27
  180. package/chunks/{dist-OAOC5OV4.js → dist-7JGAAQSC.js} +2 -2
  181. package/chunks/{dist-SL2LUMML.js → dist-A7H2KKUC.js} +1 -1
  182. package/chunks/{dist-2PERFI23.js → dist-DPPHJAVL.js} +3 -3
  183. package/chunks/{dist-PVEAFLHY.js → dist-HVBMEYWK.js} +1 -1
  184. package/chunks/{dist-ULIG4M5H.js → dist-NEBRI7WO.js} +67 -15
  185. package/chunks/{dist-WGMNT3JQ.js → dist-SFQ34F4M.js} +24 -6
  186. package/chunks/{earlyInputCapture-TQXLLISD.js → earlyInputCapture-5ZWDH3SB.js} +37 -35
  187. package/chunks/{edit-FJEUV6KZ.js → edit-J7SW6UP6.js} +36 -34
  188. package/chunks/{en-NBK3JKCE.js → en-L4UQDLWW.js} +57 -6
  189. package/chunks/{enter-worktree-7GE3TZ4P.js → enter-worktree-NRLOHRB6.js} +36 -34
  190. package/chunks/{enterPlanMode-KR5TMPCJ.js → enterPlanMode-GDSTEFSM.js} +36 -34
  191. package/chunks/{environment-SWL3CMB3.js → environment-KHYEKG5W.js} +39 -37
  192. package/chunks/{errors-XQE7BGZV.js → errors-VNLAJV3O.js} +39 -37
  193. package/chunks/{exit-worktree-CHKTYNCV.js → exit-worktree-ARVID7DF.js} +36 -34
  194. package/chunks/{exitPlanMode-QYAG6LWX.js → exitPlanMode-4ZFY7KUH.js} +36 -34
  195. package/chunks/{fast-path-Z2K7W6DC.js → fast-path-KA65VREJ.js} +9 -5
  196. package/chunks/{fr-4KG5V7UT.js → fr-QF7RI5ZE.js} +46 -5
  197. package/chunks/{gemini-YU4XKEDG.js → gemini-5PMCNPF7.js} +149 -90
  198. package/chunks/{geminiContentGenerator-ZZX6NBO4.js → geminiContentGenerator-NFSYJ4HA.js} +6 -6
  199. package/chunks/{glob-VYF5CXSX.js → glob-F3EAUDAX.js} +36 -34
  200. package/chunks/{goal-tools-HGO6VXPB.js → goal-tools-SPQ5M6D3.js} +3 -3
  201. package/chunks/{grep-5DP5ZUTW.js → grep-UJMPVZER.js} +36 -34
  202. package/chunks/{handleAutoUpdate-UVBJ56DL.js → handleAutoUpdate-MG4FRZPN.js} +42 -40
  203. package/chunks/{i18n-UQNYHEUD.js → i18n-EVHV4ZF3.js} +38 -36
  204. package/chunks/{image-gen-WHPIZNWF.js → image-gen-UC7XVHI7.js} +9 -9
  205. package/chunks/{initializer-GMIICF4W.js → initializer-O5B6HUBI.js} +43 -41
  206. package/chunks/{installationInfo-4MGG4WVI.js → installationInfo-2LIKCNVX.js} +37 -35
  207. package/chunks/{ja-F3GSQXMF.js → ja-URQESVGW.js} +46 -5
  208. package/chunks/{keychain-token-storage-IKO4G53C.js → keychain-token-storage-7BT4TQ2A.js} +2 -2
  209. package/chunks/{kittyProtocolDetector-MBRTDRBK.js → kittyProtocolDetector-BFLAKYXH.js} +1 -1
  210. package/chunks/{list-CPN2W334.js → list-BPKIE3Z4.js} +45 -43
  211. package/chunks/{list-agents-PXJ7ZUFB.js → list-agents-TWEGRJWO.js} +2 -2
  212. package/chunks/{loadedSettingsAdapter-ZQAV5NDO.js → loadedSettingsAdapter-NGLFF4V4.js} +42 -40
  213. package/chunks/{loggingContentGenerator-5E7W4IEP.js → loggingContentGenerator-7TFESRG5.js} +52 -37
  214. package/chunks/{loop-wakeup-Y26NQZ5I.js → loop-wakeup-6LLEOQYV.js} +5 -5
  215. package/chunks/{ls-ONQTTMQR.js → ls-H5WX22VK.js} +4 -4
  216. package/chunks/{lsp-YGHHRGMD.js → lsp-G6FXYEIS.js} +2 -2
  217. package/chunks/{managed-npm-update-MKKAXYXA.js → managed-npm-update-CLZGGV7F.js} +39 -37
  218. package/chunks/{mcp-V5NTJF6X.js → mcp-X5OZOOXX.js} +42 -40
  219. package/chunks/{monitor-CO3ASY6E.js → monitor-TGSVSAZZ.js} +36 -34
  220. package/chunks/nonInteractiveCli-GUMJ5JLG.js +150 -0
  221. package/chunks/{notebook-edit-Z7AEE6FM.js → notebook-edit-EZISE7IC.js} +36 -34
  222. package/chunks/{openaiContentGenerator-CA3MQOP3.js → openaiContentGenerator-ZAKS5A6P.js} +18 -18
  223. package/chunks/{pidfile-NL7MKDRU.js → pidfile-PCMBX4LR.js} +37 -35
  224. package/chunks/{processUtils-PB5K2DYH.js → processUtils-2CSVE5BW.js} +2 -2
  225. package/chunks/{pt-EQ33KWEN.js → pt-EEEQXRUI.js} +46 -5
  226. package/chunks/{qwenContentGenerator-JYFNDCOS.js → qwenContentGenerator-S3H6PIC4.js} +42 -40
  227. package/chunks/{qwenOAuth2-W2ALKQDD.js → qwenOAuth2-MNU425VE.js} +8 -8
  228. package/chunks/{read-file-IEOR5H5W.js → read-file-HXJXHBI6.js} +10 -10
  229. package/chunks/{read-mcp-resource-U2IJN5HM.js → read-mcp-resource-DRZSI3HE.js} +2 -2
  230. package/chunks/{record-artifact-MGPJ7EG4.js → record-artifact-IPIRWMFY.js} +3 -3
  231. package/chunks/{resumeHistoryUtils-BNBTP5AS.js → resumeHistoryUtils-4QAMC2S6.js} +44 -42
  232. package/chunks/{ripGrep-UMQODKEB.js → ripGrep-MCETPF2E.js} +36 -34
  233. package/chunks/{ru-7OZAUKGG.js → ru-WGZDYQAB.js} +46 -5
  234. package/chunks/{run-qwen-serve-RN2R6QQW.js → run-qwen-serve-3NMODTXV.js} +219 -55
  235. package/chunks/{runtime-QFJ2MJF5.js → runtime-AQOV7DVP.js} +49 -45
  236. package/chunks/{scheduler-AMYUFOZI.js → scheduler-DN7SZ6ZE.js} +36 -34
  237. package/chunks/{sdk-exporters-http-4BQANSJ6.js → sdk-exporters-http-IKNWQJNV.js} +2 -2
  238. package/chunks/{sdk-impl-Z4VXWVFX.js → sdk-impl-7CUUMQDB.js} +2 -2
  239. package/chunks/{send-message-PMP4ATVM.js → send-message-HGCVVDTZ.js} +5 -4
  240. package/chunks/{serve-EIIYJVX3.js → serve-WZB45HVT.js} +42 -40
  241. package/chunks/{server-EP2LLAPD.js → server-GOTZFEEZ.js} +616 -387
  242. package/chunks/{session-GMPZBLOW.js → session-23LOFLWN.js} +149 -87
  243. package/chunks/{settings-LN7RNKBS.js → settings-LI2ORVXC.js} +41 -39
  244. package/chunks/{shell-EM2OWU5X.js → shell-GLC6OP3C.js} +36 -34
  245. package/chunks/{skill-T5VPSK73.js → skill-C7W6M4GO.js} +27 -14
  246. package/chunks/{skill-settings-GMR3GMNY.js → skill-settings-W25NCK5I.js} +41 -39
  247. package/chunks/{spawnChannel-YPCU62DI.js → spawnChannel-OHJLEBLD.js} +38 -36
  248. package/chunks/{standalone-update-JAKIM2NF.js → standalone-update-E4V5V7WT.js} +39 -37
  249. package/chunks/{startInteractiveUI-YETQWGYF.js → startInteractiveUI-FKILLW7Q.js} +1815 -1034
  250. package/chunks/{syntheticOutput-D7E6ZFKR.js → syntheticOutput-EYL7DWO7.js} +3 -3
  251. package/chunks/{task-create-NFCYD43L.js → task-create-CSCEG4SM.js} +10 -9
  252. package/chunks/{task-list-TAVMGEVD.js → task-list-6CVQPORX.js} +6 -6
  253. package/chunks/{task-stop-ZDF2UJRM.js → task-stop-Y5ZTFVRX.js} +2 -2
  254. package/chunks/{task-update-6L7LHGEX.js → task-update-C52LNFU5.js} +10 -9
  255. package/chunks/{team-create-LFJ6HGNY.js → team-create-HBIJ7NHT.js} +36 -34
  256. package/chunks/{team-delete-DPGGRQBU.js → team-delete-CE7DC53R.js} +6 -6
  257. package/chunks/{team-plan-approval-C2KLSDGO.js → team-plan-approval-IBJOKS4J.js} +36 -34
  258. package/chunks/terminal-image-renderer-BOQJPVLB.js +89 -0
  259. package/chunks/{theme-manager-JQ3ZGBP5.js → theme-manager-NSS7VLUD.js} +37 -35
  260. package/chunks/{todoWrite-2UO5G6G6.js → todoWrite-YKAA57W2.js} +107 -36
  261. package/chunks/{tool-search-FAZBFRJZ.js → tool-search-DEAS54VG.js} +13 -12
  262. package/chunks/{total-session-admission-5CALIVYO.js → total-session-admission-4LWTX3WD.js} +43 -41
  263. package/chunks/{trustedFolders-KOPSP44S.js → trustedFolders-NF3J2TA7.js} +38 -36
  264. package/chunks/{update-relaunch-6QL2TICS.js → update-relaunch-WUHRH5V3.js} +5 -5
  265. package/chunks/{updateCheck-4Q6RIKFW.js → updateCheck-7EG4BLBH.js} +41 -39
  266. package/chunks/{useAutoAcceptIndicator-Q5BMKZV3.js → useAutoAcceptIndicator-ZRV56GVP.js} +46 -44
  267. package/chunks/{validateNonInterActiveAuth-GKOC5F6Z.js → validateNonInterActiveAuth-WC2FPE3I.js} +86 -84
  268. package/chunks/{version-VSATY2LD.js → version-6SS7KCUP.js} +1 -1
  269. package/chunks/{web-fetch-JCLKL6NL.js → web-fetch-Y5QQMPKA.js} +13 -13
  270. package/chunks/{web-search-DSQTFUBX.js → web-search-QKKEU723.js} +8 -8
  271. package/chunks/{workflow-M6QAVKVJ.js → workflow-PWIU6Q2M.js} +418 -304
  272. package/chunks/{workspace-providers-status-YKLTLMJA.js → workspace-providers-status-B2DM6FNK.js} +45 -43
  273. package/chunks/{workspace-registration-store-5DMGKU64.js → workspace-registration-store-7Y33TTHJ.js} +2 -2
  274. package/chunks/{workspace-registry-PTYUSYXX.js → workspace-registry-TSNJKLPC.js} +43 -41
  275. package/chunks/{workspace-service-25ML3T47.js → workspace-service-4I2QHDX4.js} +50 -48
  276. package/chunks/{workspace-skills-status-YKZWOYRX.js → workspace-skills-status-HVXNWVGR.js} +44 -42
  277. package/chunks/{workspace-trust-reconciler-5FI2S3OS.js → workspace-trust-reconciler-RDBBGMWP.js} +49 -47
  278. package/chunks/{write-file-AY7VLP5X.js → write-file-RQ4DG5RY.js} +36 -34
  279. package/chunks/{zh-H3YDK5QW.js → zh-72XPUG5P.js} +54 -6
  280. package/chunks/{zh-TW-JYCP7XUA.js → zh-TW-LTB4SRH3.js} +54 -6
  281. package/chunks/{zoom-image-S4LJZMD5.js → zoom-image-O5YCB7WN.js} +10 -10
  282. package/cli.js +32 -48
  283. package/locales/ca.js +60 -4
  284. package/locales/de.js +62 -4
  285. package/locales/en.js +76 -6
  286. package/locales/fr.js +64 -4
  287. package/locales/ja.js +61 -4
  288. package/locales/pt.js +60 -4
  289. package/locales/ru.js +63 -4
  290. package/locales/zh-TW.js +67 -6
  291. package/locales/zh.js +67 -6
  292. package/package.json +3 -3
  293. package/web-shell/assets/{arc-xjArKHU2.js → arc-Bd7d781i.js} +1 -1
  294. package/web-shell/assets/{architectureDiagram-3BPJPVTR-BgrQDNYG.js → architectureDiagram-3BPJPVTR-CpnOJqNE.js} +1 -1
  295. package/web-shell/assets/{blockDiagram-GPEHLZMM-CXxWC-1_.js → blockDiagram-GPEHLZMM-NxVTiWnk.js} +1 -1
  296. package/web-shell/assets/{c4Diagram-AAUBKEIU-Dx1tLm3v.js → c4Diagram-AAUBKEIU-BWT75l82.js} +1 -1
  297. package/web-shell/assets/channel-A434sWK5.js +1 -0
  298. package/web-shell/assets/{chunk-2J33WTMH-CBxjogVR.js → chunk-2J33WTMH-Dzz03hk6.js} +1 -1
  299. package/web-shell/assets/{chunk-4BX2VUAB-izb-rY5W.js → chunk-4BX2VUAB-DRF-egfm.js} +1 -1
  300. package/web-shell/assets/{chunk-55IACEB6-NzLyvSb8.js → chunk-55IACEB6-Co15vq5x.js} +1 -1
  301. package/web-shell/assets/{chunk-727SXJPM-DVyKb0dl.js → chunk-727SXJPM-DHSUk1J7.js} +1 -1
  302. package/web-shell/assets/{chunk-AQP2D5EJ-BtqL7NzB.js → chunk-AQP2D5EJ-Br8u_E2R.js} +1 -1
  303. package/web-shell/assets/{chunk-FMBD7UC4-iNpdKIO0.js → chunk-FMBD7UC4-DXbPTnqF.js} +1 -1
  304. package/web-shell/assets/{chunk-ND2GUHAM-DI7cbFmX.js → chunk-ND2GUHAM-BleQdi4d.js} +1 -1
  305. package/web-shell/assets/{chunk-QZHKN3VN-3Y_-xFYs.js → chunk-QZHKN3VN-CqgBf-95.js} +1 -1
  306. package/web-shell/assets/classDiagram-4FO5ZUOK-CwPuhldB.js +1 -0
  307. package/web-shell/assets/classDiagram-v2-Q7XG4LA2-CwPuhldB.js +1 -0
  308. package/web-shell/assets/{cose-bilkent-S5V4N54A--raRatZn.js → cose-bilkent-S5V4N54A-WS7b13Yg.js} +1 -1
  309. package/web-shell/assets/{dagre-BM42HDAG-B8SU70mO.js → dagre-BM42HDAG-v_55ki8o.js} +1 -1
  310. package/web-shell/assets/{diagram-2AECGRRQ-ewgoakXo.js → diagram-2AECGRRQ-DaYAjbdK.js} +1 -1
  311. package/web-shell/assets/{diagram-5GNKFQAL-Bm_JLTnS.js → diagram-5GNKFQAL-D_Hy-L64.js} +1 -1
  312. package/web-shell/assets/{diagram-KO2AKTUF-DEIUjUy9.js → diagram-KO2AKTUF-dB06cZyD.js} +1 -1
  313. package/web-shell/assets/{diagram-LMA3HP47-C2Wmrpj-.js → diagram-LMA3HP47-DzSluD4F.js} +1 -1
  314. package/web-shell/assets/{diagram-OG6HWLK6-CtYe6bFx.js → diagram-OG6HWLK6-Cvb6CeXQ.js} +1 -1
  315. package/web-shell/assets/{erDiagram-TEJ5UH35-DfpHcWLV.js → erDiagram-TEJ5UH35-7eB2Hmov.js} +1 -1
  316. package/web-shell/assets/{flowDiagram-I6XJVG4X-CAiUhio5.js → flowDiagram-I6XJVG4X-B08iGCw_.js} +1 -1
  317. package/web-shell/assets/{ganttDiagram-6RSMTGT7-CYc7ezr4.js → ganttDiagram-6RSMTGT7-BgsrCHq-.js} +1 -1
  318. package/web-shell/assets/{gitGraphDiagram-PVQCEYII-70lbkcRS.js → gitGraphDiagram-PVQCEYII-DQLurQ2i.js} +1 -1
  319. package/web-shell/assets/{index-Ww86nT9f.js → index-B_1Z0Mgr.js} +1 -1
  320. package/web-shell/assets/index-Bi1dP2mU.css +5 -0
  321. package/web-shell/assets/index-DJ-z1Oba.js +1756 -0
  322. package/web-shell/assets/{infoDiagram-5YYISTIA-C7e8DWPa.js → infoDiagram-5YYISTIA-DXZGO8qB.js} +1 -1
  323. package/web-shell/assets/{ishikawaDiagram-YF4QCWOH-BYLr2op-.js → ishikawaDiagram-YF4QCWOH-g7gbuH2w.js} +1 -1
  324. package/web-shell/assets/{journeyDiagram-JHISSGLW-hEibAHKX.js → journeyDiagram-JHISSGLW-euefICGn.js} +1 -1
  325. package/web-shell/assets/{kanban-definition-UN3LZRKU-F2rt8Uhp.js → kanban-definition-UN3LZRKU-Dws8-mjG.js} +1 -1
  326. package/web-shell/assets/{linear-DPhyNezd.js → linear-C63F0tyh.js} +1 -1
  327. package/web-shell/assets/{mermaid.core-ZUtnug0n.js → mermaid.core-fLrhp_4-.js} +5 -5
  328. package/web-shell/assets/{mindmap-definition-RKZ34NQL-DDWcMoi8.js → mindmap-definition-RKZ34NQL-qRkydfT-.js} +1 -1
  329. package/web-shell/assets/{pieDiagram-4H26LBE5-CTRKzw1_.js → pieDiagram-4H26LBE5-BZ0wQaKN.js} +1 -1
  330. package/web-shell/assets/{quadrantDiagram-W4KKPZXB-CAEwYG_z.js → quadrantDiagram-W4KKPZXB-Fjq6GbSZ.js} +1 -1
  331. package/web-shell/assets/{requirementDiagram-4Y6WPE33-CRjnH9D9.js → requirementDiagram-4Y6WPE33-CqWvjLHU.js} +1 -1
  332. package/web-shell/assets/{sankeyDiagram-5OEKKPKP-LPQT_sOd.js → sankeyDiagram-5OEKKPKP-VcXX6PVZ.js} +1 -1
  333. package/web-shell/assets/{sequenceDiagram-3UESZ5HK-Ht2Iqvb0.js → sequenceDiagram-3UESZ5HK-BJSRtxFr.js} +1 -1
  334. package/web-shell/assets/{stateDiagram-AJRCARHV-CIUdqOXu.js → stateDiagram-AJRCARHV-Dt9zUwmZ.js} +1 -1
  335. package/web-shell/assets/stateDiagram-v2-BHNVJYJU-CCxQt56w.js +1 -0
  336. package/web-shell/assets/{timeline-definition-PNZ67QCA-BcilneZF.js → timeline-definition-PNZ67QCA-DQwlNRws.js} +1 -1
  337. package/web-shell/assets/{vennDiagram-CIIHVFJN-BljwiiH1.js → vennDiagram-CIIHVFJN-DwxjUEtQ.js} +1 -1
  338. package/web-shell/assets/{wardley-L42UT6IY-DeKjkYjq.js → wardley-L42UT6IY-D3QdMSpJ.js} +1 -1
  339. package/web-shell/assets/{wardleyDiagram-YWT4CUSO-L3U1H7KG.js → wardleyDiagram-YWT4CUSO-BddN8oK8.js} +1 -1
  340. package/web-shell/assets/{xychartDiagram-2RQKCTM6-IJnWsgMT.js → xychartDiagram-2RQKCTM6-BUNuBrRJ.js} +1 -1
  341. package/web-shell/index.html +2 -2
  342. package/chunks/chunk-4KP2F2TE.js +0 -82
  343. package/chunks/chunk-HGT6JR3U.js +0 -177
  344. package/chunks/nonInteractiveCli-ZJYEOO7U.js +0 -148
  345. package/web-shell/assets/channel-B1cyk94G.js +0 -1
  346. package/web-shell/assets/classDiagram-4FO5ZUOK-BpBYZoRE.js +0 -1
  347. package/web-shell/assets/classDiagram-v2-Q7XG4LA2-BpBYZoRE.js +0 -1
  348. package/web-shell/assets/index-C1k3PDuU.css +0 -5
  349. package/web-shell/assets/index-DZLIXILO.js +0 -1750
  350. package/web-shell/assets/insert-DmRbh9Rv.svg +0 -1
  351. package/web-shell/assets/stateDiagram-v2-BHNVJYJU-BmqqQZ6w.js +0 -1
@@ -2,7 +2,7 @@
2
2
 
3
3
  > Architecture decisions, trade-offs, and rejected alternatives for the `/review` skill.
4
4
 
5
- ## Why 12 agents + 1 verify + iterative reverse, not 1 agent?
5
+ ## Why 14 agents + 1 verify + iterative reverse, not 1 agent?
6
6
 
7
7
  **Considered:**
8
8
 
@@ -10,11 +10,12 @@
10
10
  - **5 parallel agents (original design):** Each agent focuses on one dimension. Higher coverage through forced diversity of perspective. Limited by combined Correctness+Security and a single undirected pass — recall ceiling left findings on the table that the user only discovered in subsequent /review rounds.
11
11
  - **9 parallel agents:** 6 review dimensions (Correctness, Security, Code Quality, Performance, Test Coverage, Undirected) + Build & Test. Undirected runs as 3 personas in parallel.
12
12
  - **10 parallel agents:** The 9-agent design plus Issue Fidelity & Root-Cause Ownership, which compares linked issue evidence against the PR's claimed fix before accepting a client-side change.
13
- - **12 parallel agents (current):** The 10-agent design with Correctness split into three procedural walks — 1a line-by-line scan, 1b removed-behavior audit, 1c cross-file tracer — plus up to 2 optional diff-specialized finders (Agent 8) when one domain dominates the diff.
13
+ - **12 parallel agents:** The 10-agent design with Correctness split into three procedural walks — 1a line-by-line scan, 1b removed-behavior audit, 1c cross-file tracer — plus up to 2 optional diff-specialized finders (Agent 8) when one domain dominates the diff.
14
+ - **14 parallel agents (current):** The 12-agent design with Code Quality split into three checklist slices on the same evidence that split Correctness and the invariant checklist — 3a reuse & duplication, 3b altitude & abstraction fit, 3c consistency & clarity. One agent holding a six-item quality checklist finishes one item (measured on PR #6457: one agent with an eight-item checklist found 1 of 5 defects; the same model split three ways found all 5).
14
15
 
15
- **Decision:** 12 agents. The marginal cost (12x vs 1x) is acceptable because:
16
+ **Decision:** 14 agents. The marginal cost (14x vs 1x) is acceptable because:
16
17
 
17
- 1. All 12 agents are submitted in one response and run concurrently up to the runtime's tool-call cap (default 10, `QWEN_CODE_MAX_TOOL_CONCURRENCY`) — wall time is bounded by roughly two waves at worst, still far below twelve sequential agents
18
+ 1. All 14 agents are submitted in one response and run concurrently up to the runtime's tool-call cap (default 10, `QWEN_CODE_MAX_TOOL_CONCURRENCY`) — wall time is bounded by roughly two waves at worst, still far below fourteen sequential agents
18
19
  2. Dimensional focus produces higher recall (fewer missed issues)
19
20
  3. Three undirected personas (attacker / 3am-oncall / maintainer) catch cross-dimensional issues that a single undirected agent's prompt-induced bias would miss
20
21
  4. Issue Fidelity prevents a common false approval mode: a PR can be internally well-tested while solving only the author's mistaken diagnosis, not the linked issue's original failure
@@ -30,7 +31,7 @@ Test gaps are a systematic blind spot. Review agents focused on bugs in the new
30
31
 
31
32
  ### Why a dedicated Issue Fidelity agent
32
33
 
33
- Bugfix PRs often carry their own diagnosis in the PR body, but that diagnosis can be wrong. The linked issue's original reproduction, observed payload, expected behavior, and maintainer comments must be checked before judging whether the implementation is a real fix. The implementation deliberately keeps issue discovery out of `pr-context`: the Issue Fidelity agent fetches GitHub's closing-issue metadata with `gh pr view --json closingIssuesReferences`, then fetches relevant issue discussions with `gh issue view --json title,body,comments` (the `--json` form is required — it returns the issue **body**, which `--comments` alone omits). This keeps relevance judgment in the agent instead of baking fragile PR-body parsing into TypeScript. The agent runs only for PR targets — a local-diff or file-path review has no PR or linked issue, so it is skipped there (11 agents instead of 12).
34
+ Bugfix PRs often carry their own diagnosis in the PR body, but that diagnosis can be wrong. The linked issue's original reproduction, observed payload, expected behavior, and maintainer comments must be checked before judging whether the implementation is a real fix. The implementation deliberately keeps issue discovery out of `pr-context`: the Issue Fidelity agent fetches GitHub's closing-issue metadata with `gh pr view --json closingIssuesReferences`, then fetches relevant issue discussions with `gh issue view --json title,body,comments` (the `--json` form is required — it returns the issue **body**, which `--comments` alone omits). This keeps relevance judgment in the agent instead of baking fragile PR-body parsing into TypeScript. The agent runs only for PR targets — a local-diff or file-path review has no PR or linked issue, so it is skipped there (13 agents instead of 14).
34
35
 
35
36
  The agent also enforces the root-cause ownership gate: a client-side parser/sanitizer workaround for malformed upstream output is not acceptable as a root-cause fix unless a maintainer explicitly asked for that defensive mitigation.
36
37
 
@@ -100,11 +101,11 @@ The original design gave one agent the whole diff plus a growing cumulative find
100
101
 
101
102
  ### Why the topology gate counts source lines, not diff lines
102
103
 
103
- Diff size is a bad proxy for review risk, because tests dominate it. Across this repo's last 40 merged PRs the median diff is **41% test code**, and 14 of the 40 are more than half tests. A gate on raw diff lines sends a change of 173 production lines that ships 489 lines of new tests into the territory fan-out, where the production code ends up owned by a single chunk agent — while under the dimension fan-out it would have been read by ten lenses (the diff-reading dimension agents: twelve minus Issue Fidelity and Build & Test).
104
+ Diff size is a bad proxy for review risk, because tests dominate it. Across this repo's last 40 merged PRs the median diff is **41% test code**, and 14 of the 40 are more than half tests. A gate on raw diff lines sends a change of 173 production lines that ships 489 lines of new tests into the territory fan-out, where the production code ends up owned by a single chunk agent — while under the dimension fan-out it would have been read by twelve lenses (the diff-reading dimension agents: fourteen minus Issue Fidelity and Build & Test).
104
105
 
105
- Territory fan-out is worth it when there is a lot of _risky_ code to divide, not a lot of _lines_. So the gate is `srcDiffLines > 500`, with a second clause `diffLines > 3200` as an attention bound: past that point asking ten diff-reading lenses each to swallow the whole diff dilutes all of them, and the chunk topology's base cost (`ceil(diffLines / 400) + 4`, counting the whole-diff agents that read the diff — Build & Test reads none) crosses twelve about there. It is not a promise of fewer calls — a heavy file adds three invariant agents and a dominant domain up to two specialized finders — but of one accountable reader per line instead of ten diluted ones. On the 40-PR sample the second clause never fires; it exists for a changeset dominated by tests or generated files.
106
+ Territory fan-out is worth it when there is a lot of _risky_ code to divide, not a lot of _lines_. So the gate is `srcDiffLines > 500`, with a second clause `diffLines > 3200` as an attention bound: past that point asking the thirteen diff-reading lenses each to swallow the whole diff dilutes all of them, and the chunk topology's base cost (`ceil(diffLines / 400) + 4`, counting the whole-diff agents that read the diff — Build & Test reads none) crosses that count nearer 3 600. The gate stays at 3 200 rather than moving with the roster — fanning out slightly before the crossover errs toward one accountable reader per line, and a gate that drifts every time a dimension is split or merged is a gate nobody can reason about. It is not a promise of fewer calls — a heavy file adds three invariant agents and a dominant domain up to two specialized finders — but of one accountable reader per line instead of thirteen diluted ones. On the 40-PR sample the second clause never fires; it exists for a changeset dominated by tests or generated files.
106
107
 
107
- Re-gating moved 6 of those 40 PRs from 3B back to 3A and cost 22 extra agents in total across all 40 — about 5% — measured under the earlier 10-agent 3A roster; under the current 12-agent roster the same six PRs cost 2 more each, ~34 extra (~7%). It buys those six PRs ten review lenses on their production code instead of one.
108
+ Re-gating moved 6 of those 40 PRs from 3B back to 3A and cost 22 extra agents in total across all 40 — about 5% — measured under the earlier 10-agent 3A roster; under the current 14-agent roster the same six PRs cost 4 more each, ~46 extra (~10%). It buys those six PRs twelve review lenses on their production code instead of one.
108
109
 
109
110
  Chunking itself is unchanged: the plan still tiles every line, tests and generated files included. Only the count of reviewers and their brief change. `heavy` is likewise restricted to `source` files — the invariant checklist asks about fields, timers, collections, and error taxonomies, and a rewritten test file has none of those.
110
111
 
@@ -376,6 +377,140 @@ Two other deliberate limits:
376
377
  - **A test-only diff is never probed.** A new test for old code is _supposed_ to pass with nothing reverted. Probing it would flag every such PR as inert — a false blocker on exactly the PRs we want people to write.
377
378
  - **Findings are Suggestions, not Criticals.** A test that does not gate is not itself wrong code; nothing is broken today. What the finding must say concretely is which behaviour is now shipping unprotected.
378
379
 
380
+ ## Why a review that only ever had one tree needed the other one
381
+
382
+ Every step in this pipeline looks at a single tree. The agents read the PR's code. The verifier traces a failure scenario through the PR's code. Even the probe capability — the one thing here that _runs_ rather than reads — runs against the PR's code alone. The merge base has been known since the first `fetch-pr` (`mergeBaseSha`, resolved and recorded), and it was used for exactly one thing: choosing the diff range. Nothing ever built it.
383
+
384
+ That is fine for most findings, because most findings are claims about the code in front of you: this branch is unreachable, this variable is undefined here, this lock is never released. But it leaves a class the review can only ever guess at, and it is a large one, because it is the class the diff itself is _about_:
385
+
386
+ - "This changes the output format."
387
+ - "This only adds a field; existing consumers are unaffected."
388
+ - "This silently drops the error message."
389
+ - "Before this, a cancelled call and a failed call were indistinguishable."
390
+
391
+ Every one is a statement about the **difference between two programs**, and the review has had one of them. So the difference gets recovered by reading the diff — and that is precisely the reading that fails, for the reason this document keeps rediscovering in other contexts: a diff's new lines are always present and always look correct, and whether they change what anyone observes routinely turns on code the diff never touches. It is the same shape as the `fixed by this diff` trap ("the diff adds a fix" is not "the defect can no longer fire") and the same shape as the documented-intent trap. Both were closed by making the verifier go read something outside the diff. This one cannot be closed that way, because what is outside the diff here is not a _file_ — it is a _build_.
392
+
393
+ **With a built base tree the question stops being an argument and becomes an observation.** Feed the same input to both, compare the two outputs. That is a different kind of evidence from anything else in this pipeline: not a better-traced claim, but a measurement, and the only kind that settles a disagreement about what a program used to do.
394
+
395
+ Three deliberate limits:
396
+
397
+ - **The command builds; it does not run.** Standing up the tree is the expensive, failure-prone half — a detached worktree at the right SHA, a stale sibling from a crashed run, the minimal build set, the widening loop, deadlines a real build can meet — and all of it is decidable, so it is code. _What_ to run is not: it depends entirely on the claim under test, and a fixed scenario would fit almost none of them. The report hands back a path and stops.
398
+ - **It is per-finding, not per-review.** A cold checkout means an install and a build — the honest price, and why the command's idempotent fast path reuses an already-built tree instead of letting concurrent verifiers each pay it (or worse, sweep it out from under each other mid-A/B). Paid on a review with a comparative claim it is cheap for what it settles; paid on every review it is a tax most of them get nothing for. So it lives in the verifier's brief as an option, next to the probe, on the same terms.
399
+ - **Unavailable is never a finding.** No merge base, a merge base that may be stale (`baseFetchFailed` — an A/B against the wrong base attributes the base branch's own commits to this PR, the two-dot-diff error in another shape), or a base tree that will not compile: each is a fact about the harness. The base failing to build says nothing whatsoever about the PR, and a review that filed it as one would be reporting on its own infrastructure.
400
+
401
+ ## Why rendering claims get a real renderer — and only with a user-designated scratch repo
402
+
403
+ A sanitizer PR's guarantees are claims about GitHub: what its comment pipeline decodes, what its allowlist strips, when its notification path fires. A live verification measured the gap between the model and the authority exactly where it hurts: an `@` → `@` defusal that every local reading called sound, because GitHub decodes character references _before_ the mention filter runs — a fact no local markdown library reproduces and the real renderer demonstrated in one posted comment (the mention registered, the subscription fired). Judging a sanitizer against a local model of GitHub is the same parser-divergence failure the sanitizer itself is being reviewed for.
404
+
405
+ So the verifier gets the authority itself — under three constraints that keep it from eroding the write ban:
406
+
407
+ - **User-designated, or nothing.** The capability exists only when `QWEN_REVIEW_SCRATCH_REPO` names a repo the user chose for disposable posts. There is no default, no fallback, no "any repo I can write to": the review does not pick its own outward write destination, ever.
408
+ - **Payload-minimal.** What gets posted is the markdown shape under test, never the report, the diff, or anything naming the PR or its authors — a scratch post that leaked review content would be a disclosure, not a measurement.
409
+ - **Honest without it.** Absent the setting, a rendering claim caps at low confidence / `cannot tell`. The alternative — "confirmed" off a local approximation — is precisely the false assurance the live case measured.
410
+
411
+ Step 7's write ban names the carve-out explicitly rather than relying on "the scratch repo is not the PR": the ban's strength is that a compressor cannot narrow it, so an exception it does not name is an exception a compressed run cannot trust.
412
+
413
+ ## Why extract-step exists, and why it stubs nothing
414
+
415
+ The strongest workflow verification in this repo's review history ran the real composer step from both arms with a stubbed `gh` and byte-compared outputs against a real posted comment. Everything judgment-shaped in that harness — what to stub, what input to feed, what to diff — stays with the verifier. What moves into code is the half that is mechanical and quietly error-prone by hand: finding the right job, the right step among same-named siblings, and carrying the settings whose values change the script's behaviour. The command emits the script **verbatim** plus metadata: the `${{ … }}` sites (listed, never evaluated — any value this command inserted would be an invention), the env (as comments, never half-substituted exports — an unbound variable should fail loudly), and a heuristic list of invoked commands as the stubbing starting point. Combined with `base-tree`, the two arms of the by-hand harness are now two invocations.
416
+
417
+ Three details of that mechanical half are worth naming, because each was a way the command could have been quietly wrong about the script it claims to reproduce verbatim:
418
+
419
+ - **`env:`, `shell:` and `working-directory:` are three-level settings** — workflow, job, step, nearest wins — and only the step level appears in the step's own text. Reading step level alone would reproduce by machine the transcription error the command exists to remove; measured on this repo's own workflows, **195 of 434 `run:` steps inherit env from an outer level**, so it would have been wrong about 45% of them. The metadata carries the merged result plus the level each key came from, ordered nearest-first so a step's own vars are not buried under an inherited block.
420
+ - **A `shell:` declared as `bash` is not the runner's default `bash`.** The default is `bash -e {0}`; declaring `shell: bash` (at any level) makes it `bash --noprofile --norc -eo pipefail {0}`. A pipeline whose middle stage fails aborts under one and not the other, so the emitted header carries `set -eo pipefail` or `set -e` accordingly — a distinction that decides whether the extraction measures the same script the runner ran.
421
+ - **The header must be inert, line by line.** A `env:` value can be a YAML block scalar; commenting only the entry's first line left its continuation lines in command position, and under the header's own `set -e` the step died in its preamble before its body ran. Four steps in this repo produced exactly that. The test oracle asserts the property directly — the file is the header plus the body verbatim, and every line before the body is a comment or a named directive — rather than filtering the output for lines that look executable, a filter that could not tell the header's `set -e` from one the body legitimately contains.
422
+
423
+ ## Why the round-3 lenses are prose, not detectors
424
+
425
+ Seven lenses joined the briefs from one verification round, and each is a judgment with a crisp trigger rather than a decidable predicate — which is what separates a brief lens from a subcommand here:
426
+
427
+ - **A borrowed idiom, missing what made it work at home.** The `@` rewrite was lifted from a workflow where the _code ancestor_ did the protecting; the entity was belt-and-braces, and only the braces were copied. The check — read the source context of a lifted defensive construct, name what it provided — requires understanding which surrounding condition was load-bearing.
428
+ - **A second parser is a divergence hunt.** A sanitizer's model of markdown against GitHub's parse: every input the two read differently is a bypass. Finding the divergent input is the work; a tool can only confirm one once named.
429
+ - **A `fixed` ruling on a divergence-class defect is a ruling about the family.** Round 6 of the same verification closed the fence-shaped entrance into a raw-HTML block and left the code-span entrance beside it open — same divergence, adjacent syntax. The re-check bar in Step 6 now says it outright: enumerate the sibling entrances before ruling `fixed`, report a still-open sibling as a new finding, and keep the two rulings separate (the original's `fixed` stands when its own input is closed). Judgment again: knowing which inputs are "the same mechanism" is the understanding a detector cannot supply.
430
+ - **A threshold fix's coverage is a number, and the number wants measuring.** A ratio-guard fix verified live was covering exactly half of its linked issue's reported shapes — provable only by holding the issue's own preamble fixed and binary-searching the payload size where recovery flips (~473 chars). The recipe (fix the variables, scan the guarded one, put the boundary next to the issue's report) is in the verifier's brief; choosing which variable to scan is the judgment half.
431
+ - **The sharpest parser-differential corner is the format's own delimiters as payload.** A no-escaping extractor fed a value containing its close tag truncates silently — measured live as a file written truncated with no warning. Named explicitly in the lens so the first probe is the strongest one.
432
+ - **A deliberate-design defence covers the states it argues, not the gate it shares.** An input-hold correct and argued for the _active_ state silently froze three idle states nothing had argued for — the sibling-entrance rule applied to a state machine. The documented-intent step now says it: enumerate the states a shared gate serves, and treat every unargued one on its own merits.
433
+ - **Mechanism-pinning tests, and oracles that mirror the implementation.** A test asserting `@` appears pins the mechanism while the guarantee fails; a fold-balance test whose helper re-implements the sanitizer's own scanner shares its blind spot by construction. Both shapes need the reviewer to ask what the _effect_ assertion would be and where an independent oracle would come from.
434
+
435
+ ## Why test failures are attributed by measurement, and why the delta is over file sets
436
+
437
+ Agent 7's brief has always carried a path rule: a failure in a file the diff changed is a Critical, one in a file it did not touch is pre-existing. It was the best rule available when the review had one tree, and it misclassifies in both directions — an environment-sensitive test failing in a touched file gets filed as a Critical the PR did not cause, and a PR that breaks a test in an untouched file (the exact shape the base-tree section above is about) gets waved through. The first live run of this pipeline hit the benign half: three env-sensitive core failures the model had to _reason_ into "pre-existing, not in diff", correctly but on judgment.
438
+
439
+ With `base-tree` standing, attribution is decidable: `test-delta` reruns the same failed command in the built merge base and diffs the outcomes. The two design points that matter:
440
+
441
+ - **File sets, not counts.** Measured on a live re-verification: the same branch's flaky suite failed _different test names_ on two consecutive runs, so counts (and names) are noise. The failing-file set is the stable unit, and an empty net-new set is the strongest "all pre-existing" statement obtainable.
442
+ - **Failures only, and base attributes nothing it did not finish.** A green PR-side suite has nothing to attribute, and base's suite was green before the PR existed — so the base run costs exactly one rerun per PR-side failure. A base rerun that times out attributes _nothing_: promoting PR-side failures to net-new off an unfinished run would manufacture the command's strongest evidence out of an infrastructure timeout (this shipped briefly in review of the command itself, caught because the test that "covered" it asserted only the note text). The same holds for every other way the base side can end up unmeasured — a rerun that failed without naming a file, a command the budget could not fit, a base tree that would not build — and the report names each with its own reason, because "we could not measure" and "we measured nothing" are different facts to the author.
443
+
444
+ ## Why the cache carries a findings ledger
445
+
446
+ A human reviewer's round-2 comment opens with "M1 is fixed"; the pipeline's round-2 opened with a fresh list, because the incremental cache stored a _count_ and a _verdict_ — enough to scope the diff to `lastCommitSha..HEAD`, nothing with which to say what became of round 1's findings. The author was left to diff two reports by hand, which inverts who is doing the review.
447
+
448
+ So the cache now carries the findings themselves, with round-scoped ids (`R1-2`), and an incremental re-review owes each entry a ruling under the same bar the open-Criticals re-check already enforces — _fixed_ requires tracing that the mechanism can no longer fire, not observing that the diff contains a fix. Two boundaries keep the ledger honest: only **confirmed high-confidence** findings enter (next round re-asserts each entry by id, so the ledger holds claims the review stands behind), and a finding ruled fixed _leaves_ (the cache is what the next round must check, not history — the report already told the story). The fail-closed rules are unchanged: a run that must not advance the cache does not advance the ledger either.
449
+
450
+ ## Why the ledger's authoritative copy rides the posted review, not the cache
451
+
452
+ The round ledger shipped as a local cache file and its first multi-round live use exposed the flaw immediately: four model-comparison rounds reviewed the same two PRs from the same machine at medium effort, and every round opened from scratch — medium never reads the cache, and had the rounds run from CI or another clone there would have been no cache to read. Meanwhile the one artifact every environment can see — the posted review — carried nothing machine-readable, even though the human it imitates opens round 2 with "M1 is fixed" precisely because the previous report is right there on the PR.
453
+
454
+ So the authoritative ledger now travels in the posted body as an HTML-comment marker: invisible on the PR page, durable as the comment, recovered by the next round's `pr-context` wherever it runs, with the local cache demoted to fallback for rounds that never posted. Three boundaries keep it honest:
455
+
456
+ - **Own-account, latest round only.** The ledger claims "these are the findings the previous /review stood behind", and only this account's reviews can make that claim — another user's marker is data about _their_ tooling. Each posted round embeds a fresh full copy, so the newest marker is the whole state.
457
+ - **Data, not authority.** Every recovered entry is owed a Step 6 ruling against the code — the ledger routes work, it never rules. A tampered or stale marker therefore costs a few wasted rulings, not a wrong verdict, which is why parsing is fail-quiet and the round number feeds `compose-review` from a CLI-written side file rather than a model's memory.
458
+ - **Medium reads, high writes.** Recovering the ledger is free (the reviews were already fetched), so the default-effort re-review finally opens like a round-2 comment; the cache write and the posting that carries the marker keep their existing effort gates untouched.
459
+
460
+ Two consequences of those boundaries are worth naming rather than discovering. Ids are **carried, not renumbered**: a still-standing finding is re-reported under the id it already has, that id is written into the comment right after the severity marker, and `buildLedger` reads it back — because a ledger that renumbered by position would key the next round's work list to ids the report riding beside it never used, and `R1-2 names the same claim in every round` is the entire payoff. And own-account recovery means a PR reviewed from **two** accounts — a maintainer locally, a bot in CI — keeps two independent ledgers, each with its own round counter and its own `R2-1`; that is the honest reading of "only this account's reviews can claim what this account stood behind", but it does mean the ids are scoped to the account that wrote them, not to the PR.
461
+
462
+ ## Why three more mutation operators, and why each is shaped the way it is
463
+
464
+ Statement deletion with a safety-verb filter was the first operator because it has the cleanest survivor semantics. But a live maintainer re-verification produced a survivor list the deletion operator cannot express — and every entry mapped to one of three shapes, each with equally crisp semantics:
465
+
466
+ - **`?? fallback` dropped.** The surviving case was the one line preventing a previously-fixed regression from returning through a different path — a coalesce to `getModel()` that nothing tested. A coalesce survivor means the miss path is unexercised, and the miss path is frequently the entire safety property.
467
+ - **Guard condition → `true`.** The surviving case was the round-2 fix _itself_ — a skip-condition shipped in response to review, tested by nothing. Restricted to comparison-bearing `if`s on purpose: forcing `if (ready)` to `true` survives trivially everywhere and means nothing; forcing `a !== b` to `true` surviving means no test pins when the guard must _not_ fire, which is precisely the untested half of any guard.
468
+ - **`+ UPPER_CONST` dropped.** The surviving case was a reserve term in a budget estimate. A term-drop survivor means the constant never decides any test's outcome — the boundary is unpinned.
469
+
470
+ Mechanically they are **replacements**, not deletions, which bought one bug worth recording: the selector's per-line code view is trimmed and literal-blanked, and an edit index computed there and applied to the raw line spliced `iftrue 0)` into a guard. The fix is the conservative equivalence the selector now enforces — a line whose raw text and code view disagree (it carries a string or comment) yields no candidate at all, because a mangled mutant is worse than a skipped one: its compile error reads as `inconclusive` and quietly spends a cap slot. Deletion mutants keep cap priority (they have the track record); the operators queue behind them and every skip is counted.
471
+
472
+ ## Why the quality brief checks documentation parity, not documentation
473
+
474
+ "This flag needs docs" is a reviewer's preference; "three of this flag's four siblings have a docs entry and it does not" is the codebase's own convention, broken. The lens is deliberately the second shape: no documented sibling, no finding, and the finding must name the precedent file — so the Suggestion arrives as the house standard rather than taste. The trigger that earns it a place at all is the compounding case from a live review: a surface whose behaviour can _silently change_ (an automatic model swap with a warning) shipping undocumented leaves the user staring at a message with nowhere to look it up.
475
+
476
+ ## Why the Test Plan is checked — and why a count mismatch is never a contradiction
477
+
478
+ Every other input this pipeline reads is something the review has to derive: the diff, the linked issue, the existing threads, the build's exit code. A Test Plan is different. It is a list of falsifiable assertions the author **already wrote down** and handed over, and until `test-plan` existed the review read none of them.
479
+
480
+ Not for want of the text — `pr-context` renders the PR body in full. But its consumer is Agent 0, and Agent 0's question is root-cause fidelity: is this the right fix for the linked issue? "The author says 471 tests pass; do they?" is a different question, nobody owned it, and the answer is frequently no in a way that costs the next reader real time — a path from a commit that got amended away, an `npm run test:unit` that was renamed, a count copied from the first push.
481
+
482
+ The split follows this document's recurring line — determinism owns the evidence, judgment owns the ruling — but the interesting part is where it says determinism owns **nothing**. Two claim kinds are decidable here with no model and no false positives:
483
+
484
+ - **A path that is not there.** Checkable against the reviewed tree. Absent from the diff _and_ absent from the worktree means the sentence describes some other commit. (Present-but-untouched is not a defect: "ran the existing suite at X" is a normal thing to write, and the ruling says so.)
485
+ - **An npm script that does not exist.** Checkable against the workspace manifests. If no package defines it, a reviewer who follows the Test Plan cannot run it.
486
+
487
+ **A test count is the third kind, it is the one that motivated the command, and it is deliberately not ruled a contradiction.** The temptation is obvious — the count is right there, `build-test` observed a count, compare them. It is wrong, because a count is only falsifiable against the suite the author meant, and a Test Plan almost never names one. `build-test` runs the subset of workspaces the diff touched; the author ran whatever they ran. `471 ≠ 472` is then a fact about two different measurements, and filing it as a defect is filing arithmetic the command cannot do. So the verdict is `differs`: both numbers, side by side, framed as claimed-versus-observed, and the reader decides. That is what the observation was worth in the first place — a note to the author, never a blocker. The real 471-vs-472 case that prompted this was the mildest item in a four-item review, and the fix was "bump the number".
488
+
489
+ **Nothing here blocks and nothing caps**, which makes `testPlanGate` the first gate in this file that is pure disclosure. Both halves are deliberate. A Test Plan defect is not a code defect — the diff is unaffected, and the verdict is about the code; spending the review's one irreversible public action on a documentation nit is exactly the "cry wolf" cost the design philosophy exists to avoid. And capping on a **missing** report would cap essentially every PR, because most produce no notes at all. That is the deferred-checker precedent from `script-lint`, for the identical reason: a limitation the author cannot fix must not make their PR un-Approvable forever. A stale report is dropped in silence rather than failed closed, since there is no cap to fall back to and a note about a previous commit's Test Plan is worse than no note.
490
+
491
+ ## Why the probe is also per-hunk, when there are already mutants
492
+
493
+ The efficacy command asks "does anything gate this change?" three ways, and the third exists because the first two leave a gap that is easy to miss:
494
+
495
+ | probe | neutralises | answers |
496
+ | ------ | -------------------------------------------------------- | ------------------------ |
497
+ | revert | **all** the diff's source, at once | is ANY of this gated? |
498
+ | mutant | **one statement**, from a high-precision safety-verb set | is THIS statement gated? |
499
+ | hunk | **one hunk** | is THIS change gated? |
500
+
501
+ The revert probe is all-or-nothing, and the live dogfood that motivated the mutants showed exactly what that costs: a file with six well-tested behaviours and one untested safety statement reverts red on the six, reports `gated`, and the seventh — the PR's headline invariant — is invisible. The mutants close that, but only for statements the safety-verb set recognises: calls that discard, detach or reset state, and reassignment to an empty collection. That set is deliberately narrow, because a wide one produces mutants nobody should act on.
502
+
503
+ So a diff made of **condition changes, return-value changes, format changes, off-by-one fixes** — which is most diffs — generates **zero mutants**, and its only signal is the all-or-nothing revert. A hunk is the natural unit for the missing question: it is the granularity the author wrote and the granularity a reviewer reads, and reverting one at a time is the only way to attribute a still-green suite to a **particular** change rather than to the diff at large.
504
+
505
+ Four things this gets right by construction, three of them borrowed from the mutants:
506
+
507
+ - **The patch, not a checkout.** `git checkout base -- <file>` reverts the whole file and the verdict belongs to no particular change — the revert probe's limitation, one level down. Reverse-applying the hunk's own patch keeps the attribution, and `git` does the line-offset arithmetic so a later hunk lands in the right place without this code tracking offsets.
508
+ - **The third outcome is still asymmetric.** A patch that will not apply, or a tree that will not compile without the hunk, is `inconclusive` and **never** `killed`. A compile error says nothing about whether a test would have caught a behavioural regression, and scoring it as "a test caught it" is the precise false assurance this whole command exists to remove.
509
+ - **Restore by content, never by re-applying forward.** A forward re-apply can fail on its own and would leave the tree neutralised for every probe after it, turning one bad restore into a run of false survivors.
510
+ - **Hunks a mutant already covers are skipped**, and the probes run **last**, out of the mutants' leftover budget. The ordering is the priority statement: the safety-verb mutant is the higher-precision experiment, so it is bought first; a hunk probe is what the remainder buys. Both skip counts — cap and budget — are reported, because a hunk probe that never ran must never be readable as a hunk that came back clean.
511
+
512
+ The gating mistake worth recording, because it inverted the feature while every test stayed green: the hunk loop first lived **inside** the mutant branch, so it ran only when the diff already had a safety-verb candidate. The one class of diff per-hunk probing exists for — no mutants at all — got nothing. Selection now happens beside the mutants' and the phase runs whenever **either** kind has candidates.
513
+
379
514
  ## Why "fixed by this diff" is the verdict that needed a bar
380
515
 
381
516
  The re-check has three verdicts, and until PR #6486 only two of them cost anything:
@@ -435,11 +570,11 @@ The countermeasure is cheap and needs no new machinery: before Step 4, sanity-ch
435
570
 
436
571
  **Considered:**
437
572
 
438
- - **Always-full (original):** every `/review` runs the full pipeline. Right for a PR verdict; wrong for a 5-line pre-commit sanity check — 12 agents, sharded verification, and ≥2 reverse-audit rounds to re-derive what one reader could see in a single pass.
573
+ - **Always-full (original):** every `/review` runs the full pipeline. Right for a PR verdict; wrong for a 5-line pre-commit sanity check — 14 agents, sharded verification, and ≥2 reverse-audit rounds to re-derive what one reader could see in a single pass.
439
574
  - **A `--quick` boolean:** two modes, but "quick" hides what is and isn't checked (rules? cross-file? build?).
440
- - **Three levels (chosen):** **low** = one orchestrator pass over the chunk plan, hunk-visible bugs only, ≤8 findings. **medium** = the finder angles (1a, 1b, 1c, quality/altitude, performance, conventions) run **sequentially in the orchestrator's own context** inline sequencing, not subagents, is what makes the level cheap≤12 findings. **high** = the full pipeline, unchanged.
575
+ - **Three levels (chosen):** **low** = 3-6 directed angles (per `plan.budget.inlineAngles`) plus a gap sweep, all in the orchestrator's own context over the chunk plan hunk-visible bugs only, ≤10 unverified findings. **medium** = the high pipeline minus its most expensive passes: the parallel finder fan-out over a reduced dimension set (no adversarial personas, no Agent 8), build & test, and a single verification pass verified findings, Approve capped at Comment, no reverse audit. **high** = the full pipeline, unchanged.
441
576
 
442
- **Guardrails, because a quick pass is recall-limited by construction.** "Quick pass" means **low and medium together** they differ in depth (one diff pass vs. sequential finder angles; ≤8 vs. ≤12 findings) but share every guardrail below, because what the guardrails defend against is the same at both: findings that no verifier ever checked.
577
+ **Guardrails, because an unverified pass is recall-limited by construction.** These guardrails defend against findings that no verifier ever checked, which since medium became a verified fan-out means **low alone**; medium shares only the cache and posting rules (its Approve cap is Step 6's own rule, not one of these).
443
578
 
444
579
  - Labeled **unverified**; no Approve/Request-changes verdict is emitted. A verdict is a claim the pipeline earns in Steps 4–5; a quick pass claims findings, not absence of findings.
445
580
  - Never posts to the PR: `--comment` forces high, and a "post comments" follow-up after a quick pass is declined.
@@ -450,18 +585,18 @@ The countermeasure is cheap and needs no new machinery: before Step 4, sanity-ch
450
585
 
451
586
  ## LLM call budget
452
587
 
453
- **Small diffs (≤ 500 source lines AND ≤ 3200 total diff lines, Step 3A, high effort) — 15-21 calls (typically 15-17):**
588
+ **Small diffs (≤ 500 source lines AND ≤ 3200 total diff lines, Step 3A, high effort) — 17-23 calls (typically 17-19):**
454
589
 
455
- | Stage | Calls | Why |
456
- | ----------------------- | ------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
457
- | Review agents | 12 (+0-2) | issue fidelity + 3 procedural correctness walks (1a/1b/1c) + security/quality/perf/tests + 3 undirected personas + build&test, plus 0-2 diff-specialized finders; cross-repo skips Agents 7 and 1c (10), non-PR skips Agent 0 (11) |
458
- | Sharded verification | `ceil(F/8)` | F = findings; typically 1-2; keeps each verifier's job small on high-finding reviews |
459
- | Iterative reverse audit | 2-5 | loop ends after two consecutive dry rounds; 5-round hard cap |
460
- | **Total** | **~15-21 (~13-20)** | Row maxima do not co-occur on typical runs (~15-17 is common), but the honest sum of ranges is 15-21 same-repo, 13-20 cross-repo/local. **Low/medium effort: 0 subagent calls** — the inline pass runs in the orchestrator's own context |
590
+ | Stage | Calls | Why |
591
+ | ----------------------- | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
592
+ | Review agents | 14 (+0-2) | issue fidelity + 3 procedural correctness walks (1a/1b/1c) + security + 3 quality slices (3a/3b/3c) + perf/tests + 3 undirected personas + build&test, plus 0-2 diff-specialized finders; cross-repo skips Agents 7 and 1c (12), non-PR skips Agent 0 (13) |
593
+ | Sharded verification | `ceil(F/8)` | F = findings; typically 1-2; keeps each verifier's job small on high-finding reviews |
594
+ | Iterative reverse audit | 2-5 | loop ends after two consecutive dry rounds; 5-round hard cap |
595
+ | **Total** | **~17-23 (~15-22)** | Row maxima do not co-occur on typical runs (~17-19 is common), but the honest sum of ranges is 17-23 same-repo, 15-22 cross-repo/local. **Low effort: 0 subagent calls** — the angle rotation runs in the orchestrator's own context; medium launches its reduced fan-out |
461
596
 
462
597
  **Large diffs (> 500 source lines OR > 3200 total diff lines, Step 3B, high effort) — `ceil(diffLines / 400)` chunk agents + `5..7` whole-diff agents + `3H` invariant agents (H = heavy files) + `ceil(F/8)` verify (F = findings) + `rounds × chunks` reverse audit.** The reverse audit dominates: it fans out one auditor per chunk per round, and the stop rule needs two consecutive dry rounds (hard cap 5). PR #6457 (5801 diff lines, 19 chunks, 1 heavy file) costs ~27-29 first-wave calls, then `19 × (2..5) = 38-95` reverse auditors — ~66-126 calls total depending on how long the audit keeps finding; ~70 is the clean-run floor, and the count scales with chunks and findings, not a fixed ceiling.
463
598
 
464
- That is roughly 4x the small-diff budget, and it buys the thing the small-diff topology cannot deliver at that size: coverage. Ten dimension agents (the roster of the day; twelve now) on a 5801-line diff each read the same truncated 14% window (see "Why the diff is a file, not a command"), so nine of the ten calls were redundant reads of the same hunks. Nineteen chunk agents each read a distinct ~390-line territory, and every line of the diff has exactly one accountable owner. The comparison to make is not ~70 calls vs ~17: PR #6457 took **eight** review rounds at 12-14 calls each — over 100 calls — and was still surfacing Criticals in code that had been in the diff since the first commit.
599
+ That is roughly 4x the small-diff budget, and it buys the thing the small-diff topology cannot deliver at that size: coverage. Ten dimension agents (the roster of the day; fourteen now) on a 5801-line diff each read the same truncated 14% window (see "Why the diff is a file, not a command"), so nine of the ten calls were redundant reads of the same hunks. Nineteen chunk agents each read a distinct ~390-line territory, and every line of the diff has exactly one accountable owner. The comparison to make is not ~70 calls vs ~17: PR #6457 took **eight** review rounds at 12-14 calls each — over 100 calls — and was still surfacing Criticals in code that had been in the diff since the first commit.
465
600
 
466
601
  Competitors: Copilot uses 1 call, Gemini uses 2, Claude /ultrareview uses 5-20 (cloud). Ours biases toward higher recall — the assumption is that "find more issues per round" is more valuable than minimizing per-run cost, because every missed issue forces the user into another `/review` iteration.
467
602
 
@@ -545,22 +680,22 @@ The convergence concern that motivated the summary is real but narrower than it
545
680
 
546
681
  For a PR with 15 findings:
547
682
 
548
- | Approach | LLM calls | Notes |
549
- | --------------------------------------------------- | -------------------- | ---------------------------------------------------------------------------------------------------- |
550
- | Copilot (1 agent) | 1 | Lowest cost, lowest coverage |
551
- | Gemini (2 LLM tasks) | 2 | Good cost, medium coverage |
552
- | Our design (5 agents, N verify) | 21 | 5+15+1 — too expensive |
553
- | Our design (5 agents, batch verify, single reverse) | 7 | 5+1+1 — original design |
554
- | Our design (9 agents, iterative reverse) | 11-13 | 9+1+(1-3) — +50% cost for meaningfully higher recall |
555
- | Our design (10 agents) | 12-14 | 10+1+(1-3) — adds issue-fidelity/root-cause gate |
556
- | Our design (12 agents + effort levels, current) | 15-21 high / 0 quick | 12(+0-2)+ceil(F/8)+(2-5) under 3A; low/medium run inline with no subagents — cost scales with intent |
557
- | Claude /ultrareview | 5-20 | Cloud-hosted, cost on Anthropic |
683
+ | Approach | LLM calls | Notes |
684
+ | --------------------------------------------------- | ------------------ | ----------------------------------------------------------------------------------------------------------------------- |
685
+ | Copilot (1 agent) | 1 | Lowest cost, lowest coverage |
686
+ | Gemini (2 LLM tasks) | 2 | Good cost, medium coverage |
687
+ | Our design (5 agents, N verify) | 21 | 5+15+1 — too expensive |
688
+ | Our design (5 agents, batch verify, single reverse) | 7 | 5+1+1 — original design |
689
+ | Our design (9 agents, iterative reverse) | 11-13 | 9+1+(1-3) — +50% cost for meaningfully higher recall |
690
+ | Our design (10 agents) | 12-14 | 10+1+(1-3) — adds issue-fidelity/root-cause gate |
691
+ | Our design (14 agents + effort levels, current) | 17-23 high / 0 low | 14(+0-2)+ceil(F/8)+(2-5) under 3A; low runs inline with no subagents, 3-6 angles by diff size — cost scales with intent |
692
+ | Claude /ultrareview | 5-20 | Cloud-hosted, cost on Anthropic |
558
693
 
559
694
  ## Future optimization: Fork Subagent
560
695
 
561
696
  > Dependency: [Fork Subagent proposal](https://github.com/wenshao/codeagents/blob/main/docs/comparison/qwen-code-improvement-report-p0-p1-core.md#2-fork-subagentp0)
562
697
 
563
- **Current problem:** Each of the ~15-21 LLM calls (12-14 review + sharded verify + 2-5 reverse audit rounds) creates a new subagent from scratch. At ~52K per agent (50K system + 2K task), that is ~780K-1.1M input tokens with massive redundancy. The cost grew along with the agent count — Fork Subagent matters even more under the current 12-agent design than under the original 5-agent design. (Effort levels bound the cost from the other side: low/medium runs spawn no subagents at all.)
698
+ **Current problem:** Each of the ~17-23 LLM calls (14-16 review + sharded verify + 2-5 reverse audit rounds) creates a new subagent from scratch. At ~52K per agent (50K system + 2K task), that is ~880K-1.2M input tokens with massive redundancy. The cost grew along with the agent count — Fork Subagent matters even more under the current 14-agent design than under the original 5-agent design. (Effort levels bound the cost from the other side: low runs spawn no subagents at all, and medium spawns the reduced fan-out.)
564
699
 
565
700
  **Fork Subagent solution:** Instead of creating independent subagents, fork the current conversation. All forks inherit the parent's full context (system prompt, conversation history, Step 1/1.1/1.5 results) and share a prompt cache prefix. The API caches the common prefix once; each fork only pays for its unique delta (~2K per agent).
566
701
 
@@ -568,13 +703,13 @@ For a PR with 15 findings:
568
703
  Current (independent subagents):
569
704
  Agent 1: [50K system] + [2K task] = 52K
570
705
  Agent 2: [50K system] + [2K task] = 52K
571
- ...× 15-21 agents = ~780K-1.1M total input tokens
706
+ ...× 17-23 agents = ~880K-1.2M total input tokens
572
707
 
573
708
  With Fork + prompt cache sharing:
574
709
  Cached prefix: [50K system + conversation history] (cached once)
575
710
  Fork 1: [cache hit] + [2K delta] = ~2K effective
576
711
  Fork 2: [cache hit] + [2K delta] = ~2K effective
577
- ...× 15-21 forks = ~50K cached + ~30-42K delta = ~80-92K total
712
+ ...× 17-23 forks = ~50K cached + ~34-46K delta = ~84-96K total
578
713
  ```
579
714
 
580
715
  **Additional benefits for /review:**