@qwen-code/qwen-code 0.22.2 → 0.22.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (324) hide show
  1. package/bundled/computer-use/SKILL.md +3 -3
  2. package/bundled/qc-helper/docs/configuration/settings.md +37 -36
  3. package/bundled/qc-helper/docs/features/channels/dws.md +2 -0
  4. package/bundled/qc-helper/docs/features/channels/overview.md +34 -3
  5. package/bundled/qc-helper/docs/features/code-review.md +3 -3
  6. package/bundled/qc-helper/docs/features/computer-use.md +2 -2
  7. package/bundled/qc-helper/docs/features/tool-use-summaries.md +1 -1
  8. package/bundled/qc-helper/docs/qwen-serve.md +1 -1
  9. package/bundled/review/SKILL.md +81 -31
  10. package/bundled/review/references/persistence.md +7 -3
  11. package/bundled/review/references/posting.md +10 -4
  12. package/chunks/{MaxSizedBox-X5FZSNAJ.js → MaxSizedBox-X42SXBJH.js} +33 -35
  13. package/chunks/{StandaloneSessionPicker-AVLI6VJP.js → StandaloneSessionPicker-2FJAOLGG.js} +51 -53
  14. package/chunks/{acpAgent-MXEAN54C.js → acpAgent-2ZMIJAFV.js} +268 -186
  15. package/chunks/{agent-PJIPDWHB.js → agent-X5WX3L2I.js} +24 -26
  16. package/chunks/{agent-headless-B4JVD7UK.js → agent-headless-IZM25LJA.js} +24 -26
  17. package/chunks/{anthropicContentGenerator-HGMBACVE.js → anthropicContentGenerator-YVTB6QX7.js} +43 -47
  18. package/chunks/{artifact-tool-5YT4QF7Y.js → artifact-tool-XSMPHZBT.js} +9 -1
  19. package/chunks/{bridge-FKZTVIS3.js → bridge-ME6QXZNA.js} +37 -39
  20. package/chunks/{channel-management-service-X4AZ6IXX.js → channel-management-service-HQ6KDS7C.js} +6 -6
  21. package/chunks/{channel-settings-store-4UFCLBBG.js → channel-settings-store-PEOI7XCG.js} +47 -44
  22. package/chunks/{channel-worker-group-27HMGD5S.js → channel-worker-group-6NA6BBZV.js} +5 -5
  23. package/chunks/{channel-worker-manager-QEGX646Z.js → channel-worker-manager-NEZ2J6RK.js} +5 -5
  24. package/chunks/{channel-worker-supervisor-UEW2WWCC.js → channel-worker-supervisor-FKU7DIHJ.js} +4 -4
  25. package/chunks/{chunk-P5GGM4KS.js → chunk-2HYKTO7V.js} +3 -3
  26. package/chunks/{chunk-XDFHY4R3.js → chunk-2LNNVAUV.js} +2 -2
  27. package/chunks/{chunk-ZFKOCECU.js → chunk-2OBKDPZZ.js} +2 -2
  28. package/chunks/{chunk-3DXFGHAJ.js → chunk-2OUD5T67.js} +1 -1
  29. package/chunks/{chunk-EOGELB3H.js → chunk-2SXSLSQE.js} +1 -1
  30. package/chunks/{chunk-4JWNNLPT.js → chunk-2YOHKHYR.js} +2 -2
  31. package/chunks/{chunk-JU4EEXC7.js → chunk-36GSQ3MM.js} +2 -2
  32. package/chunks/chunk-36LGBEM6.js +130 -0
  33. package/chunks/{chunk-H2WDGQ6D.js → chunk-3JXM2CVW.js} +1 -1
  34. package/chunks/{chunk-G34IR3D6.js → chunk-3SZC4VP4.js} +26 -109
  35. package/chunks/{chunk-HVEYF6VT.js → chunk-3WK3QDNV.js} +1 -1
  36. package/chunks/{chunk-5XBFCMWD.js → chunk-3XKLXHGK.js} +17 -9
  37. package/chunks/{chunk-SFPGAQUL.js → chunk-42IDLQWS.js} +24 -18
  38. package/chunks/{chunk-43GGFFLY.js → chunk-4K7KNVWH.js} +1 -1
  39. package/chunks/{chunk-F6FV3C5K.js → chunk-4NDLQAY2.js} +24 -272
  40. package/chunks/{chunk-VVW4ZNFY.js → chunk-4QHPPXK2.js} +15 -7
  41. package/chunks/{chunk-IJKTMLBE.js → chunk-4TT3IVGA.js} +1 -1
  42. package/chunks/{chunk-C6K43ZDE.js → chunk-4VPCTGH6.js} +5 -5
  43. package/chunks/chunk-4VQBQLIZ.js +293 -0
  44. package/chunks/{chunk-OIVXBW3W.js → chunk-62GQFYID.js} +1 -1
  45. package/chunks/{chunk-BRVWYMKV.js → chunk-6MBXY6WM.js} +1 -1
  46. package/chunks/{chunk-AGR5UROZ.js → chunk-6X7EU5TI.js} +17 -29
  47. package/chunks/{chunk-PFCQ62V4.js → chunk-6ZPDLGHN.js} +6 -6
  48. package/chunks/{chunk-RIEFGNTP.js → chunk-734C6JGI.js} +207 -43
  49. package/chunks/{chunk-FBSSIAAQ.js → chunk-73RM4POK.js} +7 -7
  50. package/chunks/{chunk-KGJGEEVR.js → chunk-7FA2II6K.js} +10 -1
  51. package/chunks/chunk-7J6OTNGO.js +424 -0
  52. package/chunks/{chunk-O2UN57NX.js → chunk-7LOELM2I.js} +19 -2
  53. package/chunks/{chunk-DIMEPKCW.js → chunk-7S226YTK.js} +1 -1
  54. package/chunks/{chunk-CPGJJJLU.js → chunk-7T52W4SW.js} +1 -1
  55. package/chunks/{chunk-5O72BNDD.js → chunk-7VMIMAVB.js} +7 -7
  56. package/chunks/{chunk-XRSOI3DL.js → chunk-A3LWD5MK.js} +1 -1
  57. package/chunks/{chunk-F7TNAPZ4.js → chunk-AIAOIFY4.js} +5 -5
  58. package/chunks/{chunk-4CDSK5AZ.js → chunk-AYFA7MZV.js} +2 -2
  59. package/chunks/{chunk-3JGZSIDA.js → chunk-B466ZSHZ.js} +1 -1
  60. package/chunks/{chunk-L6BZRIUL.js → chunk-B5PHIJGI.js} +182 -4
  61. package/chunks/{chunk-3I6UTTDX.js → chunk-BKWNKLZB.js} +19 -19
  62. package/chunks/{chunk-3FDFA4MZ.js → chunk-BOWFGCEU.js} +1 -1
  63. package/chunks/{chunk-3AFMQUTI.js → chunk-C2X7KY45.js} +1 -1
  64. package/chunks/{chunk-FY76G3G5.js → chunk-C3PEFRKD.js} +4 -4
  65. package/chunks/{chunk-FDY5RMGH.js → chunk-CFJY4DGX.js} +1 -1
  66. package/chunks/{chunk-CCBZDVUA.js → chunk-CIZCL2MW.js} +33 -20
  67. package/chunks/{chunk-LDP5OD6N.js → chunk-CLWO3B77.js} +2 -2
  68. package/chunks/chunk-DJBQCO2E.js +842 -0
  69. package/chunks/{chunk-4J657AJR.js → chunk-DJECZIJX.js} +40 -40
  70. package/chunks/{chunk-F5SZP77N.js → chunk-DJEZ55D4.js} +71 -65
  71. package/chunks/{chunk-BAJKAAOY.js → chunk-DJPASAUV.js} +3840 -1055
  72. package/chunks/{chunk-RUGNCYNO.js → chunk-DMLVSW7L.js} +532 -22
  73. package/chunks/{chunk-4FTKQNWJ.js → chunk-DYXFD5RO.js} +181 -2
  74. package/chunks/{chunk-PDMJ3KGS.js → chunk-EB6QRGHM.js} +1 -1
  75. package/chunks/{chunk-DZWPESIF.js → chunk-EFWRMA2I.js} +3 -3
  76. package/chunks/{chunk-4MQEPP3Y.js → chunk-EQSBPBZ3.js} +5 -119
  77. package/chunks/{chunk-Z7TGYN5F.js → chunk-FROAFEFB.js} +2 -2
  78. package/chunks/{chunk-MZS7GEC7.js → chunk-FTF2YIZ5.js} +1 -1
  79. package/chunks/{chunk-Z7VOMGDC.js → chunk-FTJMWDND.js} +0 -24
  80. package/chunks/{chunk-43UYNAWX.js → chunk-G6FECKTJ.js} +2 -2
  81. package/chunks/{chunk-SJ3ZUWFK.js → chunk-GBJBJ4HX.js} +4 -4
  82. package/chunks/{chunk-IOPCFOTF.js → chunk-GPGU4S2N.js} +2 -2
  83. package/chunks/{chunk-6LFJEACX.js → chunk-GVIVASYK.js} +7 -7
  84. package/chunks/{chunk-Z4X5MTRT.js → chunk-HCOBAG2L.js} +2 -4
  85. package/chunks/{chunk-A7YP63XN.js → chunk-HP5LR5HL.js} +5 -5
  86. package/chunks/{chunk-7JPW6IYH.js → chunk-HVC3PNVC.js} +1 -1
  87. package/chunks/{chunk-TRZTNIQ6.js → chunk-I6LYRJPB.js} +2 -2
  88. package/chunks/{chunk-E55H6BZ7.js → chunk-ICURCKIU.js} +1335 -119
  89. package/chunks/{chunk-4HV6HH33.js → chunk-ISKH4QX3.js} +9 -9
  90. package/chunks/{chunk-WKK5BQNP.js → chunk-JHYH6PYQ.js} +43 -420
  91. package/chunks/{chunk-6IUNAPLR.js → chunk-K5KXUXWC.js} +1 -1
  92. package/chunks/{chunk-JHS74YAB.js → chunk-KID24ZFN.js} +1 -1
  93. package/chunks/{chunk-EFUM7RVY.js → chunk-KJO2LBUK.js} +942 -16
  94. package/chunks/{chunk-B7CDU2SL.js → chunk-KYWTPLBJ.js} +17 -20
  95. package/chunks/{chunk-JSHHYW6T.js → chunk-LHBUKJG2.js} +3 -3
  96. package/chunks/{chunk-P2SU6ZTI.js → chunk-LUEHOUHM.js} +1 -1
  97. package/chunks/{chunk-66KDY3LV.js → chunk-M3QRM3DE.js} +1 -1
  98. package/chunks/{chunk-V4QXXQJ2.js → chunk-MF7HCQWZ.js} +2 -2
  99. package/chunks/{chunk-W5MJJ2V7.js → chunk-MK2GIB46.js} +1 -1
  100. package/chunks/{chunk-Y6LSMB6J.js → chunk-MO3FC2WT.js} +1 -1
  101. package/chunks/{chunk-L6YT5H34.js → chunk-MPNBVAEU.js} +4 -4
  102. package/chunks/{chunk-CSMI4FP5.js → chunk-MSBO7X2W.js} +1 -1
  103. package/chunks/{chunk-5PA6UEYA.js → chunk-MZCJWS5Y.js} +1 -1
  104. package/chunks/{chunk-BGXZBI5B.js → chunk-N4X7G4J2.js} +2 -1
  105. package/chunks/{chunk-6LFTG244.js → chunk-NMOD2FXQ.js} +15 -5
  106. package/chunks/{chunk-KUX5D6DK.js → chunk-NV55WWD6.js} +11330 -10971
  107. package/chunks/{chunk-Q26AYA62.js → chunk-OU5APVWO.js} +2 -2
  108. package/chunks/{chunk-2K6ORIHJ.js → chunk-P2G476QN.js} +11 -11
  109. package/chunks/{chunk-OIP7C4FE.js → chunk-PADPMMYM.js} +4 -4
  110. package/chunks/{chunk-SAH4BD2J.js → chunk-PQEISIKS.js} +67 -0
  111. package/chunks/{chunk-TGWAQ5RB.js → chunk-PZ66FRIC.js} +2071 -2124
  112. package/chunks/{chunk-7UYV53BC.js → chunk-Q2MSPKIJ.js} +1 -1
  113. package/chunks/{chunk-L2M46FP5.js → chunk-QYMAKPY5.js} +5 -5
  114. package/chunks/{chunk-JVPFTDBK.js → chunk-RFIDDFAL.js} +1 -1
  115. package/chunks/chunk-SEONHUS3.js +275 -0
  116. package/chunks/{chunk-ZGXYO7DY.js → chunk-SS4MRMKE.js} +122 -13
  117. package/chunks/{chunk-T32W2JHI.js → chunk-SS5SDCD3.js} +2 -2
  118. package/chunks/{chunk-3SGB2UHX.js → chunk-ST5MDGH3.js} +23 -11
  119. package/chunks/{chunk-MIIVI25Q.js → chunk-ST5V3ZOC.js} +2 -2
  120. package/chunks/{chunk-4K7NQDCG.js → chunk-T4EFG2P5.js} +18 -9
  121. package/chunks/{chunk-TWSKM647.js → chunk-T7KMERIX.js} +1 -1
  122. package/chunks/{chunk-77RG5SFQ.js → chunk-TEBDFQBZ.js} +1 -1
  123. package/chunks/{chunk-6CGL7E7G.js → chunk-TVC6HQ2U.js} +2 -2
  124. package/chunks/{chunk-BBWV7ONL.js → chunk-TZBSGJ6C.js} +6 -6
  125. package/chunks/{chunk-T3WXYEZ4.js → chunk-U54AYGEC.js} +38 -5
  126. package/chunks/{chunk-JCQD3INC.js → chunk-URE3QBBB.js} +1 -1
  127. package/chunks/{chunk-Y3QL45LS.js → chunk-V4264FEG.js} +8 -8
  128. package/chunks/{chunk-KEKZFOL6.js → chunk-V7EESTJV.js} +1 -1
  129. package/chunks/{chunk-R4PMVSR5.js → chunk-VBCPNIQD.js} +18 -20
  130. package/chunks/{chunk-RJC5EIKY.js → chunk-VGIFXISB.js} +3 -3
  131. package/chunks/{chunk-Z7FLBQPR.js → chunk-W3JH33JE.js} +7 -7
  132. package/chunks/{chunk-K2NQPH5D.js → chunk-WGY3AVGM.js} +3 -3
  133. package/chunks/{chunk-PKAYJDB3.js → chunk-WNPDGUJO.js} +21 -2
  134. package/chunks/{chunk-IZIVM7LZ.js → chunk-WXEA74YB.js} +1 -1
  135. package/chunks/{chunk-YRLW2MSX.js → chunk-XBVNNDPK.js} +11 -4
  136. package/chunks/{chunk-IK5BW3CS.js → chunk-XHF7WQ6B.js} +2 -2
  137. package/chunks/{chunk-2NYCN6F2.js → chunk-XSG6G3G6.js} +43 -26
  138. package/chunks/{chunk-HVQ3B7FF.js → chunk-Y5S7LBOZ.js} +43 -6
  139. package/chunks/{chunk-LXIBMVRN.js → chunk-YLSQLSH7.js} +13 -7
  140. package/chunks/{chunk-JSC3H3TD.js → chunk-ZEYFMJQA.js} +14 -6
  141. package/chunks/chunk-ZWJDJ6RU.js +2597 -0
  142. package/chunks/{chunk-QP4C6WCE.js → chunk-ZYMLPEZQ.js} +2 -2
  143. package/chunks/{chunk-BNMF3K5X.js → chunk-ZZMSC3RG.js} +52 -36
  144. package/chunks/{chunk-XDLEZ5HX.js → chunk-ZZUKP7HV.js} +1 -1
  145. package/chunks/{config-utils-BBOCOGZW.js → config-utils-R6M3TA34.js} +39 -39
  146. package/chunks/{contextCommand-EET67MQ2.js → contextCommand-ZCAP23J6.js} +36 -38
  147. package/chunks/{core-runtime-677KFX6Q.js → core-runtime-AXB6SK5D.js} +41 -43
  148. package/chunks/{create-sub-session-JKZAYZX7.js → create-sub-session-4QNE64MR.js} +34 -36
  149. package/chunks/{daemon-MJQUO7RA.js → daemon-3GNHB3M2.js} +495 -7
  150. package/chunks/{daemon-git-worktree-guard-6MVAAZZZ.js → daemon-git-worktree-guard-IVAJJWYK.js} +33 -35
  151. package/chunks/{daemon-status-provider-RFKZOK3S.js → daemon-status-provider-K3SJ5C2U.js} +42 -44
  152. package/chunks/{daemon-trust-policy-DRZPFB4Q.js → daemon-trust-policy-DOR3YG62.js} +41 -43
  153. package/chunks/{daemon-trust-policy-monitor-B5U6TQX5.js → daemon-trust-policy-monitor-KGGROSIV.js} +41 -43
  154. package/chunks/{deferred-core-runtime-BBVIOP24.js → deferred-core-runtime-WOYLUQQ7.js} +33 -35
  155. package/chunks/{dist-DJBNNPLN.js → dist-4WXLCBDR.js} +1 -1
  156. package/chunks/{dist-TJUDJTCB.js → dist-5VXFRTWU.js} +280 -48
  157. package/chunks/{dist-QODVOVWI.js → dist-6BPYWHVD.js} +31 -30
  158. package/chunks/{dist-ZBVOXSKF.js → dist-GFN5DCBX.js} +74 -65
  159. package/chunks/{dist-SATKXKUR.js → dist-HZNT6GQS.js} +1 -1
  160. package/chunks/{dist-ILDBUWDB.js → dist-LVZCXVPF.js} +1 -1
  161. package/chunks/{dist-7CKN54NL.js → dist-N6W6E753.js} +71 -64
  162. package/chunks/{dist-4Q7NTUNB.js → dist-NSV4NBPK.js} +31 -14
  163. package/chunks/{dist-UQUFZDIY.js → dist-SWQMYP2Q.js} +18 -16
  164. package/chunks/{earlyInputCapture-2JOMMNQ4.js → earlyInputCapture-OHBEVO4K.js} +33 -35
  165. package/chunks/{edit-BSV6YRVW.js → edit-LEQLAGPZ.js} +24 -26
  166. package/chunks/{enter-worktree-MQXFFPKM.js → enter-worktree-35QCAAFP.js} +3 -3
  167. package/chunks/{enterPlanMode-PKNF35FE.js → enterPlanMode-ONQN3WHE.js} +27 -29
  168. package/chunks/{environment-MOZMPUHR.js → environment-UANTR7A4.js} +37 -39
  169. package/chunks/{errors-DMPQVYJL.js → errors-WAK6KFBC.js} +36 -38
  170. package/chunks/{exit-worktree-L344UY7K.js → exit-worktree-XQD6D543.js} +3 -3
  171. package/chunks/{exitPlanMode-UE4FGCY4.js → exitPlanMode-UEYBG6P5.js} +24 -26
  172. package/chunks/{fast-path-773PV7PU.js → fast-path-WUL7UZ6H.js} +4 -4
  173. package/chunks/{fast-path-settings-XMFBYFUH.js → fast-path-settings-XH426LUU.js} +3 -3
  174. package/chunks/{glob-CZHOWUEA.js → glob-Y54QUWDY.js} +24 -26
  175. package/chunks/{goal-tools-YZQJXTGH.js → goal-tools-PTD3MBA7.js} +2 -1
  176. package/chunks/{handleAutoUpdate-SGXKTQM6.js → handleAutoUpdate-FP3NWNSG.js} +37 -39
  177. package/chunks/{i18n-B42HKHKH.js → i18n-INIU5X7A.js} +35 -37
  178. package/chunks/{image-gen-HAYOEQ4O.js → image-gen-J6FDGUZP.js} +4 -6
  179. package/chunks/{initializer-FB66WYR3.js → initializer-TYODV6GE.js} +41 -43
  180. package/chunks/{installationInfo-Y7AULNMX.js → installationInfo-AAPKYDSZ.js} +33 -35
  181. package/chunks/{list-2HWGSVRA.js → list-TUODDN6J.js} +44 -46
  182. package/chunks/{gemini-IP4FJAP5.js → llm-NSAVK5V3.js} +84 -82
  183. package/chunks/{geminiContentGenerator-PUNEBWJW.js → llm-content-generator-AY2QKXHE.js} +25 -25
  184. package/chunks/{loadedSettingsAdapter-7Y7IEI6T.js → loadedSettingsAdapter-KJCLZ3C4.js} +41 -43
  185. package/chunks/{loggingContentGenerator-2GGZNGLG.js → loggingContentGenerator-DSFQHXIF.js} +305 -27
  186. package/chunks/{managed-npm-update-5TEV2E3J.js → managed-npm-update-TSWENTK5.js} +33 -35
  187. package/chunks/{mcp-QGG5NMJJ.js → mcp-MWSAPFOL.js} +41 -43
  188. package/chunks/{monitor-FI7PY36Z.js → monitor-RW4GWJHB.js} +24 -26
  189. package/chunks/{nonInteractiveCli-4LO2TL72.js → nonInteractiveCli-JXKL343Q.js} +77 -78
  190. package/chunks/{notebook-edit-TPGMGBST.js → notebook-edit-CQNYS5OR.js} +24 -26
  191. package/chunks/{open-with-auth-4UI3KHUR.js → open-with-auth-YN3QZHVM.js} +1 -1
  192. package/chunks/{openaiContentGenerator-SUCYANUZ.js → openaiContentGenerator-J6QIF7EX.js} +13 -15
  193. package/chunks/{pidfile-JMF2CLMY.js → pidfile-QOIQZFT2.js} +33 -35
  194. package/chunks/process-registry-OAEG6WGC.js +662 -0
  195. package/chunks/{processUtils-IQZBYWSB.js → processUtils-GPHNANKZ.js} +2 -2
  196. package/chunks/{prompt-terminal-ledger-DCZO6S7C.js → prompt-terminal-ledger-FOMFTYOM.js} +34 -36
  197. package/chunks/{qwenContentGenerator-J4TXFGJT.js → qwenContentGenerator-WWDXPYQI.js} +28 -30
  198. package/chunks/{qwenOAuth2-RF33DNZW.js → qwenOAuth2-OF5WRBQT.js} +4 -6
  199. package/chunks/{read-file-QE2B5FTX.js → read-file-FQUQHC6N.js} +5 -5
  200. package/chunks/{record-artifact-RZGZBENU.js → record-artifact-VW23GS4F.js} +2 -2
  201. package/chunks/{report-findings-Q3SYKKW5.js → report-findings-AO5VYFLR.js} +7 -3
  202. package/chunks/{resumeHistoryUtils-CZ4JRDVW.js → resumeHistoryUtils-NSGF3UCB.js} +37 -39
  203. package/chunks/{ripGrep-R6B6HWPX.js → ripGrep-4CAT6SKV.js} +7 -7
  204. package/chunks/{run-qwen-serve-TPHZMSXJ.js → run-qwen-serve-5X6IDZER.js} +271 -86
  205. package/chunks/{runtime-TTIGSBM4.js → runtime-HK27GMPF.js} +46 -48
  206. package/chunks/{scheduled-tasks-QB4ZVUQS.js → scheduled-tasks-MN2ZZ5P6.js} +39 -41
  207. package/chunks/{scheduler-YRUG22M7.js → scheduler-WD4MHMO5.js} +35 -37
  208. package/chunks/{sdk-exporters-http-KCJ4ZFG7.js → sdk-exporters-http-TTU2CHBU.js} +1 -1
  209. package/chunks/{sdk-impl-HLSEDAEP.js → sdk-impl-UUMVYMAL.js} +10 -4
  210. package/chunks/{send-message-Y5I6NYAM.js → send-message-AJ4OCUBV.js} +18 -2
  211. package/chunks/{serve-MXXCDAQN.js → serve-2XCHPPGG.js} +41 -43
  212. package/chunks/{server-DFUSSSHI.js → server-GO2N5DH7.js} +4148 -3348
  213. package/chunks/{session-3TEHNENZ.js → session-IFNHTUBT.js} +81 -82
  214. package/chunks/session-pr-refresh-I3UBWMVP.js +364 -0
  215. package/chunks/{settings-VOZYKLTT.js → settings-L6N54CU6.js} +40 -42
  216. package/chunks/{shell-6RTKZVMI.js → shell-4Y5LLW6N.js} +24 -26
  217. package/chunks/{skill-FEOKFX26.js → skill-YOWTRAKQ.js} +58 -14
  218. package/chunks/{skill-settings-VA4XVYVL.js → skill-settings-ZL4JY5UE.js} +41 -43
  219. package/chunks/{spawnChannel-F6WVP723.js → spawnChannel-ZST6XJVJ.js} +35 -37
  220. package/chunks/{standalone-update-RGILUQVK.js → standalone-update-THVU3OKY.js} +36 -38
  221. package/chunks/{startInteractiveUI-WHJPFOGH.js → startInteractiveUI-XDFSHBZD.js} +401 -334
  222. package/chunks/{stdioHelpers-7A2EE64Z.js → stdioHelpers-AY6XCDY2.js} +3 -1
  223. package/chunks/{task-create-TMSV7RD5.js → task-create-QN4LYRSE.js} +3 -3
  224. package/chunks/{task-list-W55RBFRX.js → task-list-JPLGK45U.js} +2 -2
  225. package/chunks/{task-update-N536I4GQ.js → task-update-IWEOJFV2.js} +3 -3
  226. package/chunks/{team-create-QYSWCYCL.js → team-create-OOPXCDBA.js} +24 -26
  227. package/chunks/{team-delete-FJRFX33M.js → team-delete-RZVVSGCY.js} +2 -2
  228. package/chunks/{team-plan-approval-P7DAD644.js → team-plan-approval-DSX5EWAT.js} +24 -26
  229. package/chunks/{terminal-image-renderer-SMQE46UJ.js → terminal-image-renderer-KS25G33E.js} +33 -35
  230. package/chunks/{theme-manager-RFO5WWFF.js → theme-manager-YMZNG45F.js} +33 -35
  231. package/chunks/{tool-search-F2EMIAPV.js → tool-search-J5Q6U7DN.js} +9 -9
  232. package/chunks/{total-session-admission-HTVTP4BS.js → total-session-admission-PVN64JVV.js} +37 -39
  233. package/chunks/{trustedFolders-7OCNMUQE.js → trustedFolders-FATSC4YI.js} +35 -37
  234. package/chunks/{types-WYCFCEHR.js → types-HKLPXZB3.js} +1 -3
  235. package/chunks/{update-relaunch-4GF3OXMB.js → update-relaunch-SIKHLHVR.js} +6 -6
  236. package/chunks/{updateCheck-7QBDXZRT.js → updateCheck-D3YG2WCD.js} +36 -38
  237. package/chunks/{useAutoAcceptIndicator-UI4SNL2O.js → useAutoAcceptIndicator-ILINOSAS.js} +43 -45
  238. package/chunks/{validateNonInterActiveAuth-BPPA563H.js → validateNonInterActiveAuth-JUDQP6V4.js} +74 -75
  239. package/chunks/{version-ANP4TS6O.js → version-C63GDDWN.js} +1 -1
  240. package/chunks/{web-fetch-QJADSAH3.js → web-fetch-FOQZMNSC.js} +8 -10
  241. package/chunks/{web-search-7TTQKEKM.js → web-search-PUJTBFUC.js} +5 -7
  242. package/chunks/{web-shell-static-A6QMUNNV.js → web-shell-static-HNB4FDOE.js} +2 -2
  243. package/chunks/{workflow-5XN6XSOT.js → workflow-Y5IPINVH.js} +47 -58
  244. package/chunks/{workspace-providers-status-PDFZWL5Q.js → workspace-providers-status-IJ4VHFMB.js} +45 -47
  245. package/chunks/{workspace-registration-store-XKVW2YVH.js → workspace-registration-store-SMJWI2AI.js} +1 -1
  246. package/chunks/{workspace-registry-23R7XRB5.js → workspace-registry-BTGZIMPK.js} +37 -39
  247. package/chunks/{workspace-service-MUYO2QIC.js → workspace-service-SIN67ITC.js} +46 -51
  248. package/chunks/{workspace-skills-status-IKECTOOQ.js → workspace-skills-status-F3RKCUIX.js} +44 -46
  249. package/chunks/{workspace-trust-reconciler-KLGOB53D.js → workspace-trust-reconciler-PLIDLJZX.js} +45 -47
  250. package/chunks/{write-file-GGRBB5TZ.js → write-file-PAIGB33C.js} +26 -28
  251. package/chunks/{zoom-image-6ASG6J4C.js → zoom-image-LYQD7BQA.js} +5 -5
  252. package/cli.js +14 -14
  253. package/package.json +3 -3
  254. package/web-shell/assets/{abnfDiagram-VCTEODGH-uS-vtfs3.js → abnfDiagram-VCTEODGH-TRCMSA_6.js} +1 -1
  255. package/web-shell/assets/{arc-BSTnk5pu.js → arc-D9Uls5b3.js} +1 -1
  256. package/web-shell/assets/{architectureDiagram-5GKGNRK7-DCMyiqzy.js → architectureDiagram-5GKGNRK7-Bwb8jPXg.js} +1 -1
  257. package/web-shell/assets/{blockDiagram-NRAW4CY4-C7_gt0r1.js → blockDiagram-NRAW4CY4-DHj71TSX.js} +1 -1
  258. package/web-shell/assets/{c4Diagram-UCG6FXSJ-BeQg2L8B.js → c4Diagram-UCG6FXSJ-uKBBdbyr.js} +1 -1
  259. package/web-shell/assets/channel-CXjVLaBa.js +1 -0
  260. package/web-shell/assets/{chunk-2Q5K7J3B-DnV9bKBe.js → chunk-2Q5K7J3B-BZd0Phrl.js} +1 -1
  261. package/web-shell/assets/{chunk-5VM5RSS4-rT1GAcT7.js → chunk-5VM5RSS4-7KNeNBqN.js} +1 -1
  262. package/web-shell/assets/{chunk-F27PBJKO-BcgSxqGA.js → chunk-F27PBJKO-BeboS7Fi.js} +1 -1
  263. package/web-shell/assets/{chunk-G27WJ6UU-Bw1ZhyZq.js → chunk-G27WJ6UU-RV3JONoj.js} +1 -1
  264. package/web-shell/assets/{chunk-JWPE2WC7-D4iDw-96.js → chunk-JWPE2WC7-Yw9FCFsP.js} +1 -1
  265. package/web-shell/assets/{chunk-LCL6LL3I-D3MB3GTZ.js → chunk-LCL6LL3I-GIMnydDe.js} +1 -1
  266. package/web-shell/assets/{chunk-POPQ4Y6H-Cep2ndrS.js → chunk-POPQ4Y6H-Cfrux7bL.js} +1 -1
  267. package/web-shell/assets/{chunk-SVP7TREG-BuahMHJc.js → chunk-SVP7TREG-BlvIN0CO.js} +1 -1
  268. package/web-shell/assets/{chunk-XXDRQBXY-BdujIvJY.js → chunk-XXDRQBXY-DMs7Rr4t.js} +1 -1
  269. package/web-shell/assets/classDiagram-DTDB5LWJ-DANEijO3.js +1 -0
  270. package/web-shell/assets/classDiagram-v2-JRS7N3AN-DANEijO3.js +1 -0
  271. package/web-shell/assets/{cose-bilkent-JH36ORCC-CGLLUbF8.js → cose-bilkent-JH36ORCC-DNGqTNwI.js} +1 -1
  272. package/web-shell/assets/{cynefin-OW5HDTMX-D6kDsRSh.js → cynefin-OW5HDTMX-CojMpaoH.js} +1 -1
  273. package/web-shell/assets/{cynefinDiagram-5FMLGOSQ-CCy78qZx.js → cynefinDiagram-5FMLGOSQ-DF07YXGo.js} +1 -1
  274. package/web-shell/assets/{dagre-3AP2YEHR-D5gJuJsf.js → dagre-3AP2YEHR-17-4WtCH.js} +1 -1
  275. package/web-shell/assets/{diagram-S7CK7UJ4-ZwEZ5W__.js → diagram-S7CK7UJ4-G3Wuyz9m.js} +1 -1
  276. package/web-shell/assets/{diagram-UQ7AKVKN-B036APjv.js → diagram-UQ7AKVKN-nyWIKZ8P.js} +1 -1
  277. package/web-shell/assets/{diagram-VSXAHHWV-BfDJ57Fq.js → diagram-VSXAHHWV-tKHhmYSC.js} +1 -1
  278. package/web-shell/assets/{diagram-VX7I27RA-Wha0p5rT.js → diagram-VX7I27RA-CpDf1hYk.js} +1 -1
  279. package/web-shell/assets/{diagram-Z3DM3KII-L20bz2K1.js → diagram-Z3DM3KII-Cah2IFXG.js} +1 -1
  280. package/web-shell/assets/{ebnfDiagram-PWID7BFC-D1ZM0Cp7.js → ebnfDiagram-PWID7BFC-DNaCkgxy.js} +1 -1
  281. package/web-shell/assets/{erDiagram-SSCWMZ5O-CMtA19v8.js → erDiagram-SSCWMZ5O-CAdRgcf1.js} +1 -1
  282. package/web-shell/assets/{flowDiagram-A5DVABFB-D96m326S.js → flowDiagram-A5DVABFB-CMUJAShL.js} +1 -1
  283. package/web-shell/assets/{ganttDiagram-EL5Y4UJY-BuHiYG2o.js → ganttDiagram-EL5Y4UJY-o-_MUNey.js} +1 -1
  284. package/web-shell/assets/{gitGraphDiagram-WWUBYQGX-Ck3NmkSw.js → gitGraphDiagram-WWUBYQGX-_aG-2Z7O.js} +1 -1
  285. package/web-shell/assets/{index-CBGz4Qjl.js → index-BK_NKxiJ.js} +1 -1
  286. package/web-shell/assets/index-Ci-z4zWY.js +1913 -0
  287. package/web-shell/assets/index-LDlXZRHA.css +36 -0
  288. package/web-shell/assets/{infoDiagram-RXCK75RN-B6DAOyXe.js → infoDiagram-RXCK75RN-qJPeZbGP.js} +1 -1
  289. package/web-shell/assets/{ishikawaDiagram-5VMMS53U-Bpg-joFo.js → ishikawaDiagram-5VMMS53U-7Om4vNo1.js} +1 -1
  290. package/web-shell/assets/{journeyDiagram-EYS64GPL-BclrHzGa.js → journeyDiagram-EYS64GPL-BNtXqmBW.js} +1 -1
  291. package/web-shell/assets/{kanban-definition-3QL26DDD-5gJzNTYG.js → kanban-definition-3QL26DDD-COgFec8P.js} +1 -1
  292. package/web-shell/assets/{layout-CEz_1L8z.js → layout-B9n81PC9.js} +1 -1
  293. package/web-shell/assets/{linear-CBjTq9Hc.js → linear-Bb6nXZQq.js} +1 -1
  294. package/web-shell/assets/{mermaid.core-Dzhbxx_I.js → mermaid.core-Cgrm3JVT.js} +6 -6
  295. package/web-shell/assets/{mindmap-definition-FBJOCRG2-DQdVqlBF.js → mindmap-definition-FBJOCRG2-TKwLTjnG.js} +1 -1
  296. package/web-shell/assets/{pegDiagram-XKGWAZYB-D1JZlYXL.js → pegDiagram-XKGWAZYB-BJPuvlpi.js} +1 -1
  297. package/web-shell/assets/{pieDiagram-E7YTZNPT-BmkpYL0w.js → pieDiagram-E7YTZNPT-D01VkNsW.js} +1 -1
  298. package/web-shell/assets/{quadrantDiagram-AXDQQJYC-DRvD6co5.js → quadrantDiagram-AXDQQJYC-C4hHsiv-.js} +1 -1
  299. package/web-shell/assets/{railroadDiagram-O6MQD6OU-CEdjDuoo.js → railroadDiagram-O6MQD6OU-CuEO4xK2.js} +1 -1
  300. package/web-shell/assets/{requirementDiagram-EFPCY7ZU-BCHG-ie8.js → requirementDiagram-EFPCY7ZU-D5J_vIo5.js} +1 -1
  301. package/web-shell/assets/{sankeyDiagram-P5KCCOFB-C9cZMnye.js → sankeyDiagram-P5KCCOFB-DQKY223V.js} +1 -1
  302. package/web-shell/assets/{sequenceDiagram-WJ2MYXX4-DYHRV4zL.js → sequenceDiagram-WJ2MYXX4-D27YV_cm.js} +1 -1
  303. package/web-shell/assets/{sizeCapture-X5ZJPWSS-DlK21aye.js → sizeCapture-X5ZJPWSS-TJg8UcwV.js} +1 -1
  304. package/web-shell/assets/{stateDiagram-HBIQ2CUA-BrmZssvk.js → stateDiagram-HBIQ2CUA-CmnPzUIx.js} +1 -1
  305. package/web-shell/assets/stateDiagram-v2-4QOOHH4V-DNKPyrDc.js +1 -0
  306. package/web-shell/assets/{swimlanes-XN3QIQJK-D9xfcayR.js → swimlanes-XN3QIQJK-ByeP3_Nu.js} +1 -1
  307. package/web-shell/assets/swimlanesDiagram-VK2B7HYN-ZL1HPC6w.js +8 -0
  308. package/web-shell/assets/{timeline-definition-24CTP7MA-D4gdDL_d.js → timeline-definition-24CTP7MA-BsFA5juR.js} +1 -1
  309. package/web-shell/assets/{vennDiagram-4TSXK5OY-DdhIExjI.js → vennDiagram-4TSXK5OY-DREP2ECo.js} +1 -1
  310. package/web-shell/assets/{wardleyDiagram-VM6X3IG4-C9RS_Zof.js → wardleyDiagram-VM6X3IG4-iyd9lKi2.js} +1 -1
  311. package/web-shell/assets/{xychartDiagram-S5SC5T6Z-orRgmxWk.js → xychartDiagram-S5SC5T6Z-CVr1_j9q.js} +1 -1
  312. package/web-shell/index.html +3 -3
  313. package/chunks/chunk-7DJCPZE3.js +0 -945
  314. package/chunks/chunk-J2OSJFP3.js +0 -202
  315. package/chunks/chunk-K2OJUPOE.js +0 -78
  316. package/chunks/chunk-P3QQPMQA.js +0 -19
  317. package/chunks/process-registry-PMJOA5CO.js +0 -172
  318. package/web-shell/assets/channel-CENAhErJ.js +0 -1
  319. package/web-shell/assets/classDiagram-DTDB5LWJ-CM6FArGH.js +0 -1
  320. package/web-shell/assets/classDiagram-v2-JRS7N3AN-CM6FArGH.js +0 -1
  321. package/web-shell/assets/index-8hXIpvo4.js +0 -1865
  322. package/web-shell/assets/index-DclEv91i.css +0 -5
  323. package/web-shell/assets/stateDiagram-v2-4QOOHH4V-B2Q7XDaE.js +0 -1
  324. package/web-shell/assets/swimlanesDiagram-VK2B7HYN-DpZPOSbX.js +0 -8
@@ -65,13 +65,13 @@ You cannot fix this yourself: the skill you are reading comes from that same bun
65
65
  It prints a JSON verdict; use it **verbatim**:
66
66
 
67
67
  - `target` — `{type: "pr-number", number}` | `{type: "pr-url", url, host, owner, repo, number}` | `{type: "file", path}` | `{type: "local"}`. A `pr-url` arrives validated and canonicalized (scheme/host lowercased, query and fragment dropped, the number required to end its path segment — `/pull/42oops` is not PR 42) with host/owner/repo/number extracted; do not re-classify tokens by hand. A token that merely looks like a URL is refused with a warning and reported in `extraTokens`, never guessed into a target.
68
- - `effort` + `effortSource` — the resolved level after defaults (**high** for PR targets, **medium** for local/file) and the `--comment` override (an **effective** `--comment` forces `high`; an ignored one on a non-PR target changes nothing). Two `settings.json` keys feed the defaults: `review.effort` replaces the built-in default when `--effort` is absent (`effortSource: "configured"`), and `review.comment: true` makes every PR review behave as if `--comment` was passed — the forcings above still apply. Both resolve from operator scopes only (system/user); a repository's `.qwen/settings.json` cannot set them. Do not re-derive it.
68
+ - `effort` + `effortSource` — the resolved level after remembered/configured defaults (**high** for PR targets, **medium** for local/file) and the `--comment` override (an **effective** `--comment` forces `high`; an ignored one on a non-PR target changes nothing). `last_used` means the project reused the last level the user explicitly typed, and it outranks `review.effort`. Two `settings.json` keys feed the configured defaults: `review.effort` replaces the built-in target default when neither an explicit nor remembered level applies (`effortSource: "configured"`), and `review.comment: true` makes every PR review behave as if `--comment` was passed — the forcings above still apply. Both resolve from operator scopes only (system/user); a repository's `.qwen/settings.json` cannot set them. Do not re-derive it.
69
69
  - `comment.requested` / `comment.effective` — `effective` is what gates Step 7 (true also when only the `review.comment` setting is on); `requested && !effective` means the user asked on a non-PR target, and the warning for that is already in `warnings`.
70
70
  - `fix.requested` / `fix.effective` — `--fix` is `--comment` reflected, and gated on the opposite target. `--comment` writes to a **pull request**, so it needs one; `--fix` writes to a **working tree**, so it needs one that outlives the review. A PR review's tree is the ephemeral worktree `fetch-pr` creates and Step 9 deletes, so `--fix` on a PR target is ignored with a warning — edits there are discarded minutes later, and reporting findings as "fixed" into a directory that no longer exists is worse than not fixing them. `effective` is what gates Step 6B. An effective `--fix` also floors the effort at **medium**: it edits the user's files, and low runs no verification, so applying an unverified finding is the same mistake as posting one, aimed at their working tree instead of a pull request. It does not force **high** — medium's findings are verified, and the reverse audit high adds hunts for findings that are _missing_, which is not what deciding whether to apply one turns on.
71
- - `severityFloor` + `severityFloorSource` — the posting floor for a PR review: `critical` posts only Criticals (otherwise-postable high-confidence Suggestions are recorded and deferred — Step 6's convergence posture; low-confidence and Nice-to-have findings stay terminal-only as ever), `suggestion` posts Criticals and Suggestions at every round, and `auto` — the default — is the **round-adaptive rule you resolve in Step 6**, where the round is known: `suggestion` through round 5, `critical` from round 6 — **or `critical` from any round once the recovered ledger's `flatRounds` streak has reached its bar** (Step 6's signal-driven trigger: the first-time-finding rate has not fallen for that many consecutive rounds, so the loop is re-deriving the same set and the floor stems it early). The parser cannot resolve `auto` itself (the round comes from the previous posted round's ledger, not fetched yet), so carry the verdict's value forward and resolve it there. Explicit flag beats the `review.severityFloor` setting beats `auto`; a non-PR target has no rounds, so the flag warns and is ignored there. The floor governs what the review **posts**, never what it finds, verifies, or reports in the terminal.
71
+ - `severityFloor` + `severityFloorSource` — the posting floor for a PR review: `critical` posts only Criticals (otherwise-postable high-confidence Suggestions are recorded and deferred — Step 6's convergence posture — and so is the one Critical shape the floor defers by its axes, fails-closed on new surface; low-confidence and Nice-to-have findings stay terminal-only as ever), `suggestion` posts Criticals and Suggestions at every round, and `auto` — the default — is the **round-adaptive rule you resolve in Step 6**, where the round is known: `suggestion` through round 5, `critical` from round 6 — **or `critical` from any round once the recovered ledger's `flatRounds` streak has reached its bar** (Step 6's signal-driven trigger: the first-time-finding rate has not fallen for that many consecutive rounds, so the loop is re-deriving the same set and the floor stems it early). The parser cannot resolve `auto` itself (the round comes from the previous posted round's ledger, not fetched yet), so carry the verdict's value forward and resolve it there. Explicit flag beats the `review.severityFloor` setting beats `auto`; a non-PR target has no rounds, so the flag warns and is ignored there. The floor governs what the review **posts**, never what it finds, verifies, or reports in the terminal.
72
72
  - `topology` + `topologySource` — the shape of the run. `auto` (the default) runs the standing effort-driven pipeline described below. `minimal` runs the single-pass A/B comparison arm (Step 3M) instead — and when it is set, it OVERRIDES the effort dispatch entirely. In this step you run `parse-args` and the **diff capture only** (`fetch-pr` for a same-repo PR, the lightweight `fetch-diff` for a cross-repo PR, or the local capture for a local/file target — exactly as below), then jump straight to **Step 3M**. You SKIP the rest of Step 1's setup — the rules load, `pr-context`, `comment-status`, and the incremental-cache check — and you skip the fan-out, verification, reverse audit, and posting. `minimal` is terminal-only; the parser has already forced `comment.effective`, `fix.effective`, and `resume.effective` to false, and its warnings for that are in `warnings`. There is no configured topology — it is only ever an explicit flag.
73
- - `resume.requested` / `resume.effective` — `--resume` continues an interrupted run of the same PR instead of starting over. `effective` is what gates the resume branch below, and it is a TARGET-SHAPE gate rather than a promise: a cross-repo `pr-url` with no matching remote is `effective: true` but routes to lightweight mode, which never calls `fetch-pr` — item 3 below owns telling the user the flag is inert there. `requested && !effective` means a local or file target, or `--topology minimal` (a fresh single pass neither continues nor consumes an interrupted run), already warned in `warnings`. It never changes the effort: a continuation is pinned to the interrupted run's recorded level, and an explicitly different `--effort` makes `fetch-pr` refuse the resume and run fresh at the requested one.
74
- - `warnings` — surface every entry to the user, word for word.
73
+ - `resume.requested` / `resume.effective` — `--resume` continues an interrupted run of the same PR instead of starting over. `effective` is what gates the resume branch below, and it is a TARGET-SHAPE gate rather than a promise: a cross-repo `pr-url` with no matching remote is `effective: true` but routes to lightweight mode, which never calls `fetch-pr` — item 3 below owns telling the user the flag is inert there. `requested && !effective` means a local or file target, or `--topology minimal` (a fresh single pass neither continues nor consumes an interrupted run), already warned in `warnings`. The resolved effort source controls continuity: a target default is omitted so the interrupted run stays pinned to its recorded level; an explicit, remembered, configured, or comment-forced level is passed through, and a mismatch makes `fetch-pr` refuse the resume and run fresh at that required level.
74
+ - `warnings` — surface every entry to the user, word for word. When a warning says the last explicitly typed effort was reused, relay it as the opening line before starting the review.
75
75
  - `extraTokens` / `unknownFlags` — leftover input the parser refused to guess about; mention them to the user rather than silently dropping them.
76
76
 
77
77
  **Reference files, gated by this verdict.** This skill's conditional territory lives in `references/` beside it, and the verdict above already decides which of them this run needs — read each applicable one with `read_file` from this skill's base directory before the step that owns it:
@@ -112,7 +112,13 @@ For an **Aone Code** target — a `…/codereview/<id>` URL, a `pr-url` whose ve
112
112
  Based on the parsed `target.type`:
113
113
 
114
114
  - **`local`**: Review local uncommitted changes — staged, unstaged, **and untracked**. Capture them with `qwen review capture-local` (below); do not run `git diff` yourself. A `git diff` of any form reports changes to files git already **tracks**, and a file the user created but has not `git add`ed is in neither the index nor HEAD — so it appears in no `git diff` output at all. Reviews have skipped brand-new files this way — not judged low-risk, simply unseen (measured; DESIGN.md — The unseen untracked file).
115
- - If the capture's plan is empty (`chunks: []` — nothing staged, nothing unstaged, nothing untracked), inform the user there are no changes to review and stop heredo not proceed to the review agents
115
+ - **At medium effort, the cache is a LEDGER, never an anchor**: read the `findings` of the cache the capture names in its plan (`cachePath`)**read that field, do not compute the name**: `target` is derived inside the command and `safeTarget` is not hand-reproducible (past 64 characters it suffixes a digest, and symlink canonicalisation diverges from any hand recipe), so a predicted name misses exactly the spellings the canonicalisation exists for and the round then rules on zero entries over a Critical that still stands. At this effort the capture runs without `--cache`, so run it first and read the field off the plan Step 6 owes each entry a ruling at medium too, and a medium round that cannot see the previous high round's open Critical presents zero blockers over a blocker that still stands. Do NOT pass `--cache` to the capture and do NOT write the cache: incremental scoping and the cache write stay high-only, for the PR cache's exact reasons.
116
+ - **Incremental local rounds** (high effort only — the same gate, and the same reasons, as the PR cache): append `--cache .qwen/review-cache` to the `capture-local` command — **the DIRECTORY, not a file name you compute**. For a plain local round the file is `local.json` and either form works; for a FILE review the name is namespaced by the source path (`file-<target>-<digest>.json`), and `target` is derived inside the command from `--file`, so it does not exist yet when this step runs. Predicting it is the same hand-derivation the capture block forbids, wrong by construction — the name carries a digest only the command computes — and wrong in exactly the spelling classes canonicalisation exists for: `ln -s src srclink` then a review of `srclink/foo.ts` predicts from `srclink/foo.ts` while the command canonicalises to `src/foo.ts`, so the prediction misses and the round silently loses BOTH incremental scoping and the findings ledger, with no refusal line printed. Given the directory, the command resolves the file from the target it derived, and a directory holding no cache for this target reads as no anchor. **Do not pass a model**: the command rules the same-model gate over the identity the runtime published, not over a token you carry. A hand-carried one was wrong every time it was written, because `{{model}}` interpolates the BARE model id while the identity the CLI records is provider-qualified — two provider configurations exposing one model name compared equal and passed each other's gate, which is the whole contract. The command enforces the gates itself — same identity, same HEAD, content actually unchanged — and on any refusal falls back to the full capture with the reason on stderr; **repeat that line to the user**, whichever way it went. When it does scope incrementally, the plan carries an `incremental` block (changed files + one-import-hop interaction files, the rest left out) and the chunk briefs direct each agent accordingly; the rest of the flow reads the same plan shape it always did. **Also read the cache's `findings` ledger**: those are the previous local round's findings with their ids, and Step 6 owes each of them a ruling this round, exactly as on the PR path.
117
+ - If the plan carries `nothingToReview: { reason: "unchanged-since-last-round" }` — the field, not the stderr sentence; the capture writes it and `qwen review run` reads it, so a decided stop no longer reaches the parent as "Review did not complete" — **first check the cache's `findings` for open entries.** The state is byte-identical to the round that recorded them, so every open finding still stands VERBATIM — render the still-open list with ids and titles (no re-ruling is needed; nothing they describe can have changed), keeping severities distinct: open Criticals remain the round's blockers, open Suggestions are re-listed as open suggestions and block nothing. Then stop. Only when the cached ledger has no open findings does the stop read as clean: inform the user nothing changed since the previous round's clean review — name that round's verdict — and stop here. This is NOT the clean-tree case below: the tree is dirty, but it is byte-identical to the state the previous round already reviewed.
118
+ - If the plan carries `nothingToReview: { reason: "scope-emptied" }`, the round is decided the same way, for a different reason: the incremental slice kept zero sections — each anchored path has since been REMOVED (a file deleted, or the change discarded) or sits BYTE-IDENTICAL to what the previous round reviewed, and the stop gate does not distinguish the two. So split the cache's still-open findings by their CITED PATHS against the plan's `incremental.scope.supersededPaths` — the capture publishes exactly the paths whose recorded change is gone, and file PRESENCE cannot answer this (a discarded change leaves the file present with the cited bytes gone): a finding whose cited file IS IN `supersededPaths` is SUPERSEDED — the bytes it cited no longer exist and there is nothing left for it to block; **Never render these findings as still-standing blockers** and do not re-rule them — a verdict that rendered them as standing would repeat that contradiction every round, until HEAD or the model changes. A finding whose cited file is NOT in the list sits byte-identical to what the previous round reviewed — render it as still-standing exactly as the `unchanged-since-last-round` bullet above does (open Criticals remain the round's blockers; open Suggestions are re-listed as open suggestions and block nothing). Then stop. (Without this bullet the shape had no branch at all: `chunks: []` with an `incremental` block, so neither stop fired, `agent-prompt --roster` threw on the first diff-reading role, and the parent reported "Review did not complete" over a decided round.)
119
+ - If the plan has `chunks: []` and a NON-EMPTY `skippedFiles` and NO `nothingToReview`, that is not a stop and must never be reported as one: the capture read nothing AND could not read what it skipped. Report every skipped entry under "Not reviewed" with its reason, tell the user the working tree was not reviewed, and end the round WITHOUT a clean verdict — the absent field is the capture refusing to call this decided, and the round owes the user that distinction.
120
+ - If the plan has `chunks: []` and an EMPTY `skippedFiles` and NO `nothingToReview` on a plain local round, the capture withheld the stop field — a stop is a DECIDED outcome, and none of the shapes that land here is decided. In one, the tree MOVED while the capture was hashing it (`WARNING: 0 chunks, but the working tree changed while the capture was being hashed`); in another, a cached path DROPPED OUT of the capture while still on disk and diverges from HEAD (`WARNING: 0 chunks, but a cached path dropped out of this capture while still on disk and diverges from HEAD`) — an edit git cannot see (`git update-index --assume-unchanged` is the live case), which the anchor refusal above already named; in the third, tracked paths carry an `--assume-unchanged`/`--skip-worktree` bit (or the bits could not be enumerated), and `git diff` is blind to any edit on them (`WARNING: 0 chunks, but … carry an --assume-unchanged/--skip-worktree bit`, or the same sentence on stderr from an incremental round whose stop it withheld); in the fourth, the round ran with `--no-untracked`, so the untracked half was never enumerated — the clean-tree stop's third clause, checked by nobody, which is exactly the shape the oversized-skip recovery re-run lands in (`the tracked tree is clean, but untracked files were not enumerated (--no-untracked)`), and the two INCREMENTAL stops carry the same exclusion and withhold under the same flag (`The incremental scope kept nothing to review, but untracked files were not enumerated (--no-untracked)`): their comparisons cover tracked content only, and the gate admits no narrower round than the cache, so the cached round ran narrow too and a brand-new file is invisible to both. Never report nothing-to-review on any of these shapes and never take the clean-tree branch: for the `--no-untracked` shape do NOT re-run — report the untracked scope under "Not reviewed" and end the round without a clean verdict; for the others re-run `capture-local` once, and if the warning repeats tell the user — for the moved tree, that their tree is being modified while the review captures it; for the dropped-out path, that a file diverges from HEAD invisibly to git (an `--assume-unchanged`/`--skip-worktree` bit, or an ignore rule) and needs their inspection; for the visibility bits, which paths carry them and that clearing them (`git update-index --no-assume-unchanged` / `--no-skip-worktree`) restores reviewability — and end the round without a verdict. (A FILE review reaching this shape takes the no-diff branch below instead: a whole-file review reads the current state either way.)
121
+ - If the plan carries `nothingToReview: { reason: "clean-tree" }` (`chunks: []` — nothing staged, nothing unstaged, nothing untracked), inform the user there are no changes to review and stop here — do not proceed to the review agents. Read the FIELD, not the chunk count: a capture that SKIPPED files also has no chunks, and that round could not read what it skipped, so the capture withholds the field there and the round owes a "Not reviewed" section instead of a stop. `qwen review run` reads the same field, so this stop no longer reaches the parent as "Review did not complete". **First, the same ledger carve-out the no-changes stop above carries**: read the cache at the plan's `cachePath` and, if it holds open findings, render them with ids and severities before stopping — open Criticals as the round's still-standing blockers, open Suggestions as still-open suggestions. A clean tree is not a resolution: the common shape is a user who COMMITS the change without fixing the blocker, leaving a permanently clean tree, and without this the finding is never surfaced on any later round
116
122
 
117
123
  - **`pr-number`, or `pr-url` with a matching remote** (cross-repo `pr-url`s are handled by the lightweight mode above):
118
124
 
@@ -129,10 +135,9 @@ Based on the parsed `target.type`:
129
135
  # compose-review's own coverage recomputation — reads it from there, so they
130
136
  # cannot disagree about which agents a medium review owed. Omit it only if
131
137
  # the parser resolved the default high. On a FRESH run passing it always
132
- # is harmless; on a RESUME it is not — the ruling cannot tell a passed-
133
- # through default from a user's explicit choice, so follow the resume
134
- # bullet below: pass --effort only when the user chose a level in THIS
135
- # invocation.
138
+ # is harmless; on a RESUME it is not — pass it for explicit, last_used,
139
+ # configured, or forced-by-comment and omit it only for default, as
140
+ # detailed in the resume bullet below.
136
141
  # High-effort re-review with a cached anchor: append --since <lastCommitSha>
137
142
  # (the incremental check below) — the CLI validates the anchor and scopes
138
143
  # the diff and plan; never run git against an anchor yourself.
@@ -158,20 +163,20 @@ Based on the parsed `target.type`:
158
163
 
159
164
  Guessing the owner/repo here is not a recoverable mistake — a guessed repo has already stopped a review before it read a line of code (measured; DESIGN.md — The guessed fork repo). If `meta` fails, or the matcher exits 6 (no remote matches) or 7 (several do), say so and stop rather than picking one.
160
165
 
161
- Read `.qwen/tmp/qwen-review-pr-<n>-fetch.json` for: `worktreePath`, `baseRefName`, `headRefName`, `fetchedSha` (use as the **HEAD commit SHA** for Step 7), `isCrossRepository`, `diffStat` (files / additions / deletions), `emptyDiff` (**stop here**: the branch tree is byte-identical to its merge base — the work already landed or was superseded; tell the user and recommend close-as-superseded instead of fanning out agents over zero hunks — but first run `"${QWEN_CODE_CLI:-qwen}" review cleanup pr-<n>` to release the lease and remove the worktree just created, same as the same-SHA stop below: this stop is clean, yet without the cleanup the lease survives process exit and every later review of this PR refuses until it is deleted by hand), `collapsedFromUpstream` (disclose in the summary: overlapping merged PRs have collapsed this one to a residual — the review scope is the recomputed diff, and body claims about the rest are description-of-history, which Agent 0 should read accordingly), `prDescriptionHasHan` (the PR description contains Chinese — every posted inline comment must then be bilingual; see Step 7), and — when `--since` was passed — `incremental` (the anchor ruling the incremental-review check below acts on: `effective`/`upToDate`/`reason`) If the command fails (auth, network, PR not found), inform the user and stop. One failure needs a specific relay: a **lease conflict** says another session is already reviewing this PR. Same-PR reviews share one worktree path, so `fetch-pr` refuses rather than destroy the other session's worktree mid-run (#9205). Tell the user the PR is under review by another session and stop — do NOT delete the lease file to force the fetch: that file is the only protection the other session's state has, and removing it re-opens exactly the destruction this refusal prevents.
166
+ Read `.qwen/tmp/qwen-review-pr-<n>-fetch.json` for: `worktreePath`, `baseRefName`, `headRefName`, `fetchedSha` (use as the **HEAD commit SHA** for Step 7), `isCrossRepository`, `diffStat` (files / additions / deletions), `emptyDiff` (**stop here**: the branch tree is byte-identical to its merge base — the work already landed or was superseded; tell the user and recommend close-as-superseded instead of fanning out agents over zero hunks — but first write the stop sidecar exactly as the up-to-date stop below does (reason `empty-diff`, runId from `QWEN_REVIEW_RUN_ID`, skipped without the variable) and run `"${QWEN_CODE_CLI:-qwen}" review cleanup pr-<n>` to release the lease and remove the worktree just created, same as the same-SHA stop below: this stop is clean, yet without the cleanup the lease survives process exit and every later review of this PR refuses until it is deleted by hand), `collapsedFromUpstream` (disclose in the summary: overlapping merged PRs have collapsed this one to a residual — the review scope is the recomputed diff, and body claims about the rest are description-of-history, which Agent 0 should read accordingly), `prDescriptionHasHan` (the PR description contains Chinese — every posted inline comment must then be bilingual; see Step 7), and — when `--since` was passed — `incremental` (the anchor ruling the incremental-review check below acts on: `effective`/`upToDate`/`reason`) If the command fails (auth, network, PR not found), inform the user and stop. One failure needs a specific relay: a **lease conflict** says another session is already reviewing this PR. Same-PR reviews share one worktree path, so `fetch-pr` refuses rather than destroy the other session's worktree mid-run (#9205). Tell the user the PR is under review by another session and stop — do NOT delete the lease file to force the fetch: that file is the only protection the other session's state has, and removing it re-opens exactly the destruction this refusal prevents.
162
167
 
163
168
  Worktree isolation: all subsequent steps (agents, build/test) operate inside `worktreePath`, not the user's working tree. Cache and reports (Step 8) are written to the **main project directory**, not the worktree.
164
169
 
165
170
  - **Incremental review check** (high effort only — neither low nor medium consults or updates the cache): read `.qwen/review-cache/pr-<n>.json` **before** `fetch-pr` (it is a local file; nothing about it needs the fetch) and, when it holds a `lastCommitSha`, pass BOTH fields to the fetch verbatim: `--since <lastCommitSha> --since-model <lastModelId>` (omit `--since-model` when the cache has no `lastModelId`; do not substitute anything for it). **Copy them; do not compare them to anything.** The same-model gate is ruled inside `fetch-pr`, over the identity the runtime published — "clean up to `lastCommitSha`" is the recorded identity's verdict, and the command validates an anchor against the HISTORY, never against who certified it, so an anchor from another identity is ancestrally perfect and would scope this round past code it never reviewed. A hand-applied version of that gate was wrong every time it was written, because `{{model}}` interpolates the BARE model id while every identity the CLI records is provider-qualified: two provider configurations exposing one model name compared equal and passed each other's gate. When the gate refuses, the report says `cross-model-anchor` and the round reviews the full diff. Read the cache's `findings` ledger either way (Step 6 owes each entry a ruling; the work list carries across models, only the anchor does not). **You never run `git` against an anchor yourself** — no `git diff <sha>..HEAD`, no `cat-file`, no `merge-base --is-ancestor`: the command validates the anchor against the fetched history and computes the scoped diff and chunk plan in one pass, because a hand-run check is one a run can skip, and the hand-computed delta was exactly the shape this skill forbids everywhere else (the diff is a file the CLI writes, never a command you run). The report's `incremental` field is the decision; act on it with `lastModelId` from the cache and the current model ID (`{{model}}`):
166
171
  - `effective: true` (no `upToDate`) → the report's diff and plan ARE the incremental scope (`since..head`); continue with them exactly as with a full plan. The file set is **widened by one import hop**: a still-clean source file that imports a changed one re-enters the scope with its own full-range hunks, because the round before cleared it against the callee's OLD shape. `incremental.scope` names each file's class — `deltaFiles` (touched since the anchor), `interaction[]` (widened back in, each with the edges that did it), `contextFileCount` (weighed and passed over) — and a chunk brief built for an interaction file points its agent at that seam instead of a from-scratch re-review. **Also read the cache's `findings` ledger** (older caches have none — then there is nothing to track): these are the previous round's findings with their ids, and Step 6 owes each of them a ruling this round. (Reachable only under a matching identity: the gate inside the command is what keeps a cross-model anchor from scoping anything.)
167
- - `upToDate: true` **and** `comment.effective` is false (no `--comment` flag, and `review.comment` not enabled in settings) → inform the user "No new changes since last review" (this branch consumes no plan, so it holds even when `diffPath` is null), run `"${QWEN_CODE_CLI:-qwen}" review cleanup pr-<n>` to remove the worktree just created, and stop. **This branch does not apply on a resumed run** (`resumed: true` from the resume branch below): a continuation's `incremental` field is the interrupted attempt's history, not this run's decision, and taking the stop/cleanup here would destroy the very state `--resume` reused.
172
+ - `upToDate: true` **and** `comment.effective` is false (no `--comment` flag, and `review.comment` not enabled in settings) → inform the user "No new changes since last review" (this branch consumes no plan, so it holds even when `diffPath` is null). **Before the cleanup, write the stop sidecar** so `qwen review run` reads the round as DECIDED instead of exiting 1 "Review did not complete" over it: when the environment carries `QWEN_REVIEW_RUN_ID`, write `.qwen/tmp/qwen-review-pr-<n>-stop.json` containing exactly `{"reason": "<up-to-date|empty-diff>", "runId": "<the QWEN_REVIEW_RUN_ID value>"}` — the same reason+runId contract `capture-local` writes for local stops, runId copied verbatim (the parent's reader is nonce-fenced and discards any other stamp); without that variable no parent is reading and the file is not written. `cleanup` deliberately KEEPS this run's sidecar (it spares a `stop.json` whose `runId` matches the environment) so the parent can still read the decision after the child exits — do not remove it by hand; the next run's cleanup collects it. Then run `"${QWEN_CODE_CLI:-qwen}" review cleanup pr-<n>` to remove the worktree just created, and stop. **This branch does not apply on a resumed run** (`resumed: true` from the resume branch below): a continuation's `incremental` field is the interrupted attempt's history, not this run's decision, and taking the stop/cleanup here would destroy the very state `--resume` reused.
168
173
  - `upToDate: true` **but** `comment.effective` is true (the `--comment` flag or the `review.comment` setting) → run the full review anyway — the report already holds the full-range diff and plan for exactly this flow, unless `diffPath` is null, which is the ordinary degraded state (partial coverage, disclosed) rather than a scoping fact. Inform the user: "No new code changes. Running review to post inline comments."
169
174
  - `reason: cross-model-anchor` → the cached anchor was certified by another identity, so it was not used. Continue on the full-range plan (or, when `diffPath` is null, on the degraded state its siblings name). The command already said which identity certified it and which is running; repeat that to the user rather than restating it from the cache.
170
175
  - `effective: false` → the anchor was refused and the report says why. **Every reason names a CAUSE** — `not-an-ancestor` (a rebase or force-push); `unknown-commit`; `behind-merge-base` (the base moved past the anchor, e.g. a partial merge landed, and scoping to it would review base history the PR does not contain); `nothing-to-narrow` (the narrowing found nothing it could publish — all deterministic and all safe, because the round keeps the full range: an ordinary "undo per feedback" revert that puts lines back the way the base had them, so the PR's own diff no longer displays the undone FILE at all (a file the PR still displays does not refuse — the join fails closed and publishes its section whole instead); a capture on either side whose bytes do not survive a UTF-8 round trip; a delta the parser cannot read; and a fail-closed refusal where the two captures key the same change differently — a path or a rename git resolves differently across the two ranges — so narrowing would drop a change the PR's diff displays); `base-untrusted` (the base could not be fetched, so the clamp that keeps an anchor from scoping wider than the PR's diff could not be ruled); `capture-failed` (a capture threw, or the base fetch or merge-base resolution failed); `partition-failed` (the diff would not tile). **Whether a PLAN exists is a separate field: `diffPath`.** Non-null → the diff and plan are the full range; continue as a full review. Null → no diff exists at all: that is the `diffPath: null` degraded state (partial coverage, disclosed), whatever the reason says. Do not read one field for both facts — a reason that meant "planless" as well as "why" is what put deterministic refusals into the retry class below. The previous round's ledger is still owed its rulings in every refusal.
171
176
 
172
177
  - **When the cache has no anchor, the PR itself carries one** (high effort only, same as the cache). The file being absent is the NORMAL state everywhere except the machine that ran the last review — CI, another clone, a colleague's checkout — and it used to mean the incremental range silently degraded to the full diff every time, which is precisely the cost incremental review exists to avoid. The anchor now rides the posted review: the machine ledger's marker carries `sha`, the head the last clean round reviewed, and `pr-context` writes it into the side file `qwen-review-pr-<n>-prev-ledger.json` with the rest of the ledger. So when the cache had no anchor to pass — including the case where it HELD one that the cache-path gate withheld, because `lastModelId` was another model's: the marker may carry an anchor THIS model certified, and a round that stops at the cache would never look — **or the anchor it passed was refused** (`incremental.effective: false` — a rebase or force-push retires a cached anchor exactly when another environment may have posted a newer round whose marker still holds a valid one): proceed with the setup batch as usual, and when the side file lands with a `sha` — **different from the one already refused, OR the same sha when the refusal was infrastructure** (`base-untrusted`, `capture-failed`: the anchor was never ruled invalid, and the component that failed — a base fetch, a merge-base resolution, a capture — is re-run by the re-run. One shape of `capture-failed` retries ONCE, not forever: a base-less refusal (a null `mergeBaseSha`) means the base fetch failed (`baseFetchFailed: true`) and no local base ref remained, or `git merge-base` itself failed on a non-answer exit. The failed component IS re-run by the re-run, but the exit status cannot split the members — git exits 128 identically for a transient fetch fault and for a deterministic refusal (the base branch deleted on the remote — the refspec fetch fails every time), and the merge-base probe folds its surface failures the same way — so a second refusal of the same shape on the same sha is the deterministic member. Retry that one, once. Every other reason is deterministic for the same sha and must NOT be retried: a validity refusal re-refuses; a planless `partition-failed` always carries a `mergeBaseSha` — with no base nothing is captured and an empty diff cannot fail to tile — so both ranges were in hand and both refused to tile, which the re-run reproduces exactly, do not retry it; `nothing-to-narrow` re-narrows identically: the same two captures select the same hunks, and a capture that failed a UTF-8 round trip fails it again — and its base-less shape (a null `mergeBaseSha` with `baseFetchFailed: false`) is NOT retryable: the fetch succeeded and `git merge-base` found no common ancestor at all (a cross-fork PR with unrelated history), which a re-run reproduces exactly) —, **re-run the `fetch-pr` command from above with `--since <sha>` — REPLACING any `--since` it already carries, never appending a second one** (a repeated flag is one flag with two values; the CLI takes the last, but a command that reads as two anchors is a command nobody can check) — the PR ref is already fetched so the re-run is cheap, and it rebuilds the worktree, diff and chunk plan scoped to the delta, with the validation the old flow asked you to hand-run (`cat-file`, `merge-base --is-ancestor`) inside the command where it cannot be skipped. Then act on the new report's `incremental` field exactly as the cache path above does (**the same-model gate on this path is RULED FOR YOU, not left to you to apply**: the marker carries `model` beside its `sha` — the identity that certified the range — and `pr-context`'s ledger section states the verdict outright, either "the same-model contract HOLDS" or "**Do NOT pass the anchor above as `--since`**". Obey that sentence and do not compare the two identities yourself: the marker's `model` is a PROVIDER-QUALIFIED identity (`<model>@<digest>`) while `{{model}}` above is the bare model id, so they are not the same kind of string — comparing them by hand either never matches, which throws away this whole recovery path, or matches loosely, which accepts another provider's same-named model and scopes past code it never reviewed. A ledger section that states no verdict — because the side file survived from an earlier round the recovery could not re-vouch — is a mismatch: review the full range. The ledger's round is used only for precedence, and an `upToDate` anchor from the side file stops only when `comment.effective` is false **and the side file carries no `anchorFromRound`** — a grafted anchor that resolves to the head means the round it was carried for closed at a head its source had already certified, so `sha..HEAD` re-covers nothing, and the stop would abandon that round's owed work list without a ruling, with every later round at the same head repeating the same stop: proceed instead as when `comment.effective` is true (the re-run report already holds the full-range diff and plan) and rule every ledger entry). The decision lands AFTER the setup batch but BEFORE any agent launches, which is where the money is (a same-SHA stop still runs `cleanup`; it just fires three cheap commands later than the cache's fast path would have). An anchor that fails validation falls back to the full diff with the reason in the report, exactly as a rebased cache sha does. Two edges, both decided for you: if the side file's `round` is **higher** than the cache's, prefer the side file's sha — the cache is stale by a round some other environment posted; and a side file with no `sha` field means no anchor is recoverable. When the last posted round was fail-closed (`compose-review` withholds the anchor then — Step 8 names the conditions) and its work list survived whole, `pr-context` grafts the anchor forward from the most recent EARLIER own marker that carries one — the withhold is about the fail-closed round's own range, while the earlier round's "clean up to `sha`" stays true, and scoping `sha..HEAD` re-covers the gap (the ledger section says "anchoring at", never "reviewed at", when the anchor was carried forward this way, and names the round it was carried from). So a missing `sha` means a shape the graft refuses or cannot reach — the winning work list was truncated by the marker's size caps (a partial work list must not certify a range — the dropped entries would fall outside the grafted scope and retire silently), the only anchored own marker is the winner's own round (one round cannot both certify and withhold), the winner ran at the same head the candidate sha certifies (grafting it would hand Step 1 a same-sha stop that abandons the work list the winner still owes), every own round on the PR closed without an anchor, the only markers are other accounts' (the sha never crosses accounts), or the markers predate the field — and the review is full-range. (The side file may also carry `commitId` — the previous review's own `commit_id`. That is Step 6's **age reference** for the convergence posture, present even on fail-closed rounds; it is never an anchor, and scoping the diff to it would skip exactly the range a fail-closed round could not certify.)
173
178
 
174
- - **Resuming an interrupted run (`--resume`)**: when `parse-args` reported `resume.effective: true`, append `--resume` to the `fetch-pr` command above, and decide `--effort` off `effortSource`, not off whether the word `--effort` was typed. Pass the resolved level whenever `effortSource` is `explicit` **or `forced-by-comment`** (the `--comment` flag or the `review.comment` setting forces high — parse-args announces "running at high effort"); omit it ONLY when `effortSource` is `default`. `fetch-pr` cannot tell a passed-through default from a chosen level: the interrupted run may have recorded a different one, and handing it the resolved default refuses the resume (`effort-mismatch`) whose fresh fall-through discards the very state `--resume` exists to save — blaming an effort nobody asked for. Omitted, the continuation pins to the recorded level. A level this invocation actually requires — a user's explicit `--effort`, or the high that `--comment` forces — that differs from the recorded one is NOT a passed-through default: pass it, so a recorded lower level refuses (`effort-mismatch`) and runs fresh at the level this invocation needs. That is right — different effort is different work, and posting authority raising the required depth is different work too, never a silent pin. Omitting a `forced-by-comment` high is the trap: `fetch-pr` has no `--comment` input and reads `requestedEffort` only from `--effort`, so the null would pin the continuation at the recorded sub-high level while `--comment` stays effective — the "effective comment at medium effort" state the medium-tier rules call impossible, posting nothing (medium skips posting) or posting from a pipeline missing the high-only passes the forcing exists to guarantee. `fetch-pr` rules on the interrupted attempt's on-disk state itself (worktree still at `fetchedSha` and clean, diff bytes unchanged, PR head unmoved, resume cap unspent — every probe is a fact it gathers, none is yours to assert) and prints one JSON line on stdout. Branch on it:
179
+ - **Resuming an interrupted run (`--resume`)**: when `parse-args` reported `resume.effective: true`, append `--resume` to the `fetch-pr` command above, and decide `--effort` off `effortSource`, not off whether the word `--effort` was typed. Pass the resolved level whenever `effortSource` is `explicit`, `last_used`, `configured`, or `forced-by-comment` (the `--comment` flag or the `review.comment` setting forces high — parse-args announces "running at high effort"); omit it ONLY when `effortSource` is `default`. `fetch-pr` cannot tell a passed-through default from a chosen level: the interrupted run may have recorded a different one, and handing it the resolved default refuses the resume (`effort-mismatch`) whose fresh fall-through discards the very state `--resume` exists to save — blaming an effort nobody asked for. Omitted, the continuation pins to the recorded level. A level this invocation actually requires — a user's explicit `--effort`, the project's remembered level, a configured `review.effort`, or the high that `--comment` forces — that differs from the recorded one is NOT a passed-through default: pass it, so a mismatch refuses the resume (`effort-mismatch`) and runs fresh at the level this invocation needs. That is right — different effort is different work, and posting authority raising the required depth is different work too, never a silent pin. Omitting a `forced-by-comment` high is the trap: `fetch-pr` has no `--comment` input and reads `requestedEffort` only from `--effort`, so the null would pin the continuation at the recorded sub-high level while `--comment` stays effective — the "effective comment at medium effort" state the medium-tier rules call impossible, posting nothing (medium skips posting) or posting from a pipeline missing the high-only passes the forcing exists to guarantee. `fetch-pr` rules on the interrupted attempt's on-disk state itself (worktree still at `fetchedSha` and clean, diff bytes unchanged, PR head unmoved, resume cap unspent — every probe is a fact it gathers, none is yours to assert) and prints one JSON line on stdout. Branch on it:
175
180
  - **`{"resumed": true, ...}`** — this run continues the interrupted one. The report at the `--out` path is the PREVIOUS attempt's, deliberately left untouched (its mtime is the run epoch every downstream fence keys on); read it for the worktree, plan and diff, which are all reused. The report's `incremental` field is now HISTORY, not a decision to re-take: a resumed run proceeds on the reused plan and does NOT re-enter the incremental check above — in particular it never takes the `upToDate: true` stop/cleanup branch, which runs `cleanup pr-<n>` and would destroy the exact worktree and lease `--resume` just saved (the interrupted attempt was a `--comment` full review of an up-to-date PR; resuming it without `--comment` effective in THIS invocation would otherwise route it straight into "No new changes since last review" and abandon it). Then rebuild your working state from disk before launching anything:
176
181
 
177
182
  ```bash
@@ -218,7 +223,7 @@ Based on the parsed `target.type`:
218
223
  - **Attach repository context** at medium or high effort, before `agent-prompt --roster` (and therefore before launching agents): run `qwen review repo-context` with absolute `--plan`, `--worktree`, and `--out` paths. See the repository-context step in the Diff capture section below; for same-repo PRs the manifest is read from the trusted merge base recorded by `fetch-pr`.
219
224
 
220
225
  - **`file`** (e.g., `src/foo.ts`):
221
- - Run `"${QWEN_CODE_CLI:-qwen}" review capture-local --file <file> --target <filename> --out .qwen/tmp/qwen-review-<filename>-plan.json` to get its changes (`--out` is required — see the capture block below for the full form). An **untracked** target file is captured whole (every line reads as added), which is the right frame for a file that does not exist upstream yet. The path is taken relative to **your** working directory and must be inside the repo.
226
+ - Run `"${QWEN_CODE_CLI:-qwen}" review capture-local --file <file> --out .qwen/tmp/file-review-<first 24 chars of the basename>-<HHMMSS>-plan.json` to get its changes (`--out` is required, and the 24-char truncation is not optional a POSIX basename may run to 255 bytes, the decoration adds 29, and the full spelling dies with ENAMETOOLONG before the capture runs; the capture block below carries the same form and the reason). **A file review carries the same ledger and incremental rules as `local` above — read those four bullets and apply them here**: append `--cache .qwen/review-cache` at high effort (the DIRECTORY; the command resolves this target's file from the target it derives, and that name is namespaced by source path so it is not yours to spell), read the cache's `findings` at medium and high alike, and branch on `nothingToReview` exactly as they say. Without this the file-path ledger was write-only: Step 8 wrote it and nothing ever read it back, so round 2 of a high-effort file review presented zero blockers over a Critical round 1 had recorded as open. **Do not pass `--target` for a file review and do not compute one**: the command derives it from `--file`, using the same repo-relative canonicalisation and flattening `qwen review run` uses to name the artifacts it waits for. Applying that recipe by hand is what made the two disagree — the hand version normalises characters but does not canonicalise, so `ln -s src srclink` then a review of `srclink/foo.ts` had the parent waiting on one name while every child artifact carried another, and a review that had already run reported no verdict. An **untracked** target file is captured whole (every line reads as added), which is the right frame for a file that does not exist upstream yet. The path is taken relative to **your** working directory and must be inside the repo.
222
227
  - If the plan is empty (the file is tracked and unmodified), read the file and review its current state — see the no-diff branch below
223
228
 
224
229
  ### Diff capture and the review topology
@@ -247,8 +252,49 @@ For **local-diff and file-path reviews**, capture and plan in one command:
247
252
  ```bash
248
253
  "${QWEN_CODE_CLI:-qwen}" review capture-local --effort <effort> --out .qwen/tmp/qwen-review-local-plan.json
249
254
  # for a file-path review:
250
- "${QWEN_CODE_CLI:-qwen}" review capture-local --file <file> --target <filename> --effort <effort> \
251
- --out .qwen/tmp/qwen-review-<filename>-plan.json
255
+ "${QWEN_CODE_CLI:-qwen}" review capture-local --file <file> --effort <effort> \
256
+ --out .qwen/tmp/file-review-<first 24 chars of the basename>-<HHMMSS>-plan.json
257
+ # The plan's own `--out` is the ONE name you may choose: you write it and you
258
+ # read it back, so it cannot diverge from anything. Make it UNIQUE to this
259
+ # run and keep it BOUNDED — at most the first 24 characters of the basename
260
+ # plus a time suffix. Bounded, not merely "short": the decoration around it
261
+ # is 29 characters, a basename is itself allowed up to 255, and the plan
262
+ # write dies with ENAMETOOLONG before the capture runs — every round, for
263
+ # that target. The family deliberately does NOT start with `qwen-review-`:
264
+ # Step 9's `cleanup` sweeps `.qwen/tmp/qwen-review-<target>-*`, and any
265
+ # `qwen-review-…` family is inside SOME target's sweep — a file literally
266
+ # named `file` (or `file-<X>`) cleaned up while another file review ran
267
+ # swept that review's live plan mid-round and killed it on its next plan
268
+ # read. `file-review-…` is outside every sweep prefix, which is what makes
269
+ # the "cleanup must never glob its family" contract in Step 9 true. Truncating cannot collide within a run (the time suffix
270
+ # separates), and across runs it does not matter: you write this name and
271
+ # you read it back. Never the full PATH flattened into one name, for the
272
+ # same ceiling one level worse. One fixed name is not safe here: a file
273
+ # review takes no lease (leases are PR-only) and the plan is re-read all
274
+ # round long (`repo-context --plan`, `agent-prompt --roster`,
275
+ # `check-coverage`, `compose-review`, and Step 8's
276
+ # `cachePath`/`cacheCandidatePath`), so two concurrent file reviews
277
+ # overwrite each other's central artifact mid-run — the second round then
278
+ # reviews the first's file and merges its findings into the wrong ledger.
279
+ # It does not have to match anything the CLI derives; it only has to differ
280
+ # from another run's.
281
+ #
282
+ # Every OTHER artifact of this round — the roster, coverage,
283
+ # compose-review's `--out`, Step 8's cache name, Step 9's
284
+ # `cleanup <target>` — must carry the token the CLI derived, and the report
285
+ # hands it to you as **`target`**. READ IT; do not recompute it. `qwen review
286
+ # run` pins the artifact name it waits for from the same canonicalisation, and
287
+ # a stem flattened by hand agrees with it only where the two happen to: put a
288
+ # symlink below the repo root (`ln -s src srclink`, then review
289
+ # `srclink/foo.ts`) and every artifact you name misses the poll, so a review
290
+ # that has already run — and with --comment, already posted — reports that no
291
+ # verdict was produced.
292
+ #
293
+ # Never the basename either: the target keys the tmp stems AND the review
294
+ # cache, and `src/index.ts` and `test/index.ts` sharing the target
295
+ # `index.ts` would overwrite each other's cache, the second review erasing
296
+ # the first file's still-open findings. The CLI's token never collides that
297
+ # way; a hand-picked one can.
252
298
  # <effort> is the resolved level (local defaults to medium). It is recorded in
253
299
  # the plan so the roster, check-coverage and compose-review all read one value.
254
300
  ```
@@ -621,7 +667,7 @@ Then skip Steps 4 and 5 entirely and go to Step 6 with these adjustments:
621
667
 
622
668
  ### Deduplication
623
669
 
624
- Before verification, merge findings that refer to the same issue (same file, same line range, same root cause) even if reported by different agents. Keep the most detailed description and note which agents flagged it. When severities differ across merged items, use the **highest severity** — never let deduplication downgrade severity. **If a merged finding includes any deterministic source** (`[build]`, `[test]`), treat the entire merged finding as pre-confirmed — retain all source tags for reporting, preserve deterministic severity as authoritative, and skip verification.
670
+ Before verification, merge findings that refer to the same issue (same file, same line range, same root cause) even if reported by different agents. Keep the most detailed description and note which agents flagged it. When severities differ across merged items, use the **highest severity** — never let deduplication downgrade severity. **Deduplication merges the fix side too: keep every `fixWitness` and every sourced `fixConstraint` the merged findings carry.** Combine consistent constraints into one; when two conflict, adjudicate explicitly — re-read the named sources and keep the constraint the code actually bears — instead of silently discarding one with the less-detailed report. The most-detailed-description pick is about the claim's wording and cannot see a fix-side sentence only another agent's copy recorded, and canonicalization receives only the deduplicated record: a witness or a constraint dropped here reads as absent at posting, leaving the unwitnessed guard or the unconstrained fix these fields exist to prevent. **If a merged finding includes any deterministic source** (`[build]`, `[test]`), treat the entire merged finding as pre-confirmed — retain all source tags for reporting, preserve deterministic severity as authoritative, and skip verification.
625
671
 
626
672
  ### Batch verification
627
673
 
@@ -656,7 +702,7 @@ The A/B's version axis is git, and it is not the only one. A claim that the code
656
702
 
657
703
  The brief also carries **`extract-step`**, which is the A/B's counterpart for a claim about a **workflow**. A `run:` script is a shell program that happens to live inside YAML, and reviewing one in place fails in a way reading normal code does not: the body is indented inside a block scalar, the `env:` that decides its behaviour is spread over three levels — workflow, job, step, nearest wins, and two of them sit nowhere near the step — and every `${{ … }}` is a hole the reader silently fills in. `qwen review extract-step` lifts the script out **verbatim** as an executable and reports what the runner would have supplied around it: the merged three-level `env:` with each key's level named, every `${{ … }}` site listed unevaluated (the stub list — the command refuses to invent values), the resolved `shell` and `working-directory`, and a heuristic list of invoked commands. What to stub and what to feed it stays with the verifier, which is the judgment half; with `base-tree`, the two arms of a workflow A/B become two invocations. A `uses:` step has no `run:` and is refused rather than simulated.
658
704
 
659
- **The witness rule.** The capabilities above exist so a verdict can be something a run produced instead of something a reading concluded, and for a **Critical** that difference is the verdict: a confirmed Critical carries a **witness** — the observed output that settled it, quoted and trimmed to the deciding lines — or one line saying why none could run (`witness: not run — <why>`: the claim needs infrastructure the harness lacks, a timing window no probe can pin, state only production holds). The forms a witness takes are exactly the capabilities' outputs: the probe's flip (both sides), the A/B's two quoted outputs, an extract-step run, the failing build/test text a `[build]`/`[test]` finding already carries, the render read-back, the **version axis**'s two-version pair (above), and — all below — the **impact sweep**, its **table sweep** specialization, and an **isolation by elimination** pair. A confirmed Critical carrying neither the witness nor the one-line reason is not confirmed at the bar this pipeline posts at: sort it **low confidence** — terminal-only, "Needs Human Review" — whatever the verifier's prose says. The demotion is deliberately mechanical, the same shape as the `— [unverified]` tag — and like that tag it has a machine half, not just this rule: `qwen review findings` (Step 6) demotes any high-confidence `[review]`-source Critical that arrives without the `witness` field and names each demotion on stderr, so a sort you miss here is caught at canonicalization rather than posted. Deterministic sources are exempt there by construction — a `[build]`/`[test]`/`[probe]` finding IS a run's output. This is the double-execute lesson made the default instead of the option (measured; DESIGN.md — The double-execute the probe caught), and it is what maintainer dogfooding measured at scale from the other side: in the review rounds that held up, every posted hard finding quoted executed output, and the one claim written from a reading alone was retracted publicly a round later when its first measurement came back zero (measured; DESIGN.md — The read-only claim retracted in round 2 (PR #8225)).
705
+ **The witness rule.** The capabilities above exist so a verdict can be something a run produced instead of something a reading concluded, and for anything this review can **post** that difference is the verdict: a confirmed Critical — and, on the same terms, a confirmed Suggestion — carries a **witness** — the observed output that settled it, quoted and trimmed to the deciding lines — or one line saying why none could run (`witness: not run — <the capability that came closest, and why it could not>`: the claim needs infrastructure the harness lacks, a timing window no probe can pin, state only production holds — the named capability is the escape hatch's toll, and a reason-less line counts as no witness at all). Both postable severities on purpose: an unexecuted claim rides onto the author's screen through the Suggestion door exactly as it would through the Critical one, and only `Nice to have` — terminal-only by construction — is exempt. The forms a witness takes are exactly the capabilities' outputs: the probe's flip (both sides), the A/B's two quoted outputs — including the **paired live-stack captures** `ab-drive` hands back when the claim needs a running product on both arms — an extract-step run, the failing build/test text a `[build]`/`[test]` finding already carries, the render read-back, the **version axis**'s two-version pair (above), the **hunk-necessity pair** (the same probe run intact and with one hunk reverted via `revert-hunk` the load-bearing question, measured), and — all below — the **impact sweep**, its **table sweep** specialization, and an **isolation by elimination** pair. A confirmed Critical or Suggestion carrying neither the witness nor the one-line reason is not confirmed at the bar this pipeline posts at: sort it **low confidence** — terminal-only, "Needs Human Review" — whatever the verifier's prose says. The demotion is deliberately mechanical, the same shape as the `— [unverified]` tag — and like that tag it has a machine half, not just this rule: `qwen review findings` (Step 6) demotes any high-confidence `[review]`-source Critical or Suggestion that arrives without the `witness` field (a reason-less `not run` line included) and names each demotion on stderr, so a sort you miss here is caught at canonicalization rather than posted. Deterministic sources are exempt there by construction — a `[build]`/`[test]`/`[probe]` finding IS a run's output. This is the double-execute lesson made the default instead of the option (measured; DESIGN.md — The double-execute the probe caught), and it is what maintainer dogfooding measured at scale from the other side: in the review rounds that held up, every posted hard finding quoted executed output, and the one claim written from a reading alone was retracted publicly a round later when its first measurement came back zero (measured; DESIGN.md — The read-only claim retracted in round 2 (PR #8225)).
660
706
 
661
707
  **The impact sweep** is the witness form for a defect that is mechanically enumerable — a pattern misused, a predicate that misclassifies, a parser that mishandles a shape. Instead of confirming the one reported instance, run the check over the repo's **real population** (every workflow step body, every call site, every input the predicate will actually see) and quote the count. "195 of 434 real `run:` bodies reach this path" is at once the confirmation, the severity evidence, and a number the author can re-run rather than argue with — and "0 of 434" is the retraction that keeps a false Critical off the PR. Two guards keep a sweep evidence rather than theatre: its oracle must be an **external authority** — the real parser, the real tool, `bash -n` — never a reimplementation of the logic under test, because a mirror of the implementation shares its blind spots and mirrored sweeps have manufactured false findings twice (measured; DESIGN.md — The mirrored oracle's false positives (PR #8225)); and a nonzero count is spot-checked by reading one hit before it is quoted.
662
708
 
@@ -664,7 +710,9 @@ The brief also carries **`extract-step`**, which is the A/B's counterpart for a
664
710
 
665
711
  **Isolation by elimination** is the witness form for a claim about an **aggregate** — a summed gauge, a maximum across children, a count over a fleet. The instinct is to add a per-component dump and read that, and the verdict is then a reading of code the review itself wrote. The cheaper move runs the other way: **shrink the contributing population instead of instrumenting the reader**. Take the aggregate with every contributor live, remove exactly one — kill the process, unregister the workspace, drop the feed — and take it again; both numbers come out of unmodified code. Read the pair for the combining rule rather than as a subtraction: doubling with the population is a sum, holding flat is not one, and reducing the population to a single contributor makes the reading that contributor's own value outright. The **difference** is a contributor's value only under a sum — under a maximum, removing a non-holder moves nothing and removing the holder exposes the next-largest. It settles the questions an aggregate cannot answer about itself, which is a larger class than it looks: whether a total is a sum or a maximum (a two-child daemon whose summed RSS moved 193.6 → 377.5 MB while its reported heap peak moved 103.5 → 103.7 MB has answered it), and whether a field is per-component or fleet-wide. It does not settle every question of that family: whether a contributor reporting nothing is skipped or folded in as a zero is invisible under a sum and a maximum alike, and shows only in a figure a zero would move — a count, a denominator, an average. Identify the contributor you remove by something the product did not choose for you — a process's own working directory, its port, its registered id — because removing the one you assumed is how this quietly answers a different question than the one asked.
666
712
 
667
- **After verification:** remove all rejected findings. Separate confirmed findings into two groups: high-confidence and low-confidence, applying the witness rule as you sort — a Critical whose confirmation carries neither witness nor the one-line reason lands in the low-confidence group. The witness rides the finding from here on — into the findings artifact (`witness`, Step 6), the terminal report, and, on a posting run, the inline comment body (Step 7) — because the evidence that settled the verdict is the one part of a finding the author can act on without re-deriving the bug. Low-confidence findings appear **only in terminal output** (under "Needs Human Review") and are **never posted as PR inline comments** — this preserves the "Silence is better than noise" principle for PR interactions.
713
+ **After verification:** remove all rejected findings. Separate confirmed findings into two groups: high-confidence and low-confidence, applying the witness rule as you sort — a Critical or Suggestion whose confirmation carries neither witness nor the one-line reason lands in the low-confidence group. The witness rides the finding from here on — into the findings artifact (`witness`, Step 6), the terminal report, and, on a posting run, the inline comment body (Step 7) — because the evidence that settled the verdict is the one part of a finding the author can act on without re-deriving the bug. **So do the two decision axes the verifier read off that witness** — `direction` (`certifies-falsely` | `fails-closed`) and `baseline` (`regression` | `new-surface`): they ride the artifact (Step 6), the Critical's claim line as bracket tags (Step 7), and the ledger marker, because Step 6's convergence posture routes a Critical by them. Copy each axis exactly as the verifier stated it; an axis the verifier omitted stays absent — never fill one in from the finding's prose, because an unclassified Critical posts at any floor while a guess on EITHER axis of the pair — a `fails-closed` beside a settled `new-surface`, or a `new-surface` beside a settled `fails-closed` — takes a blocker off the pull request, since the deferral needs both. Low-confidence findings appear **only in terminal output** (under "Needs Human Review") and are **never posted as PR inline comments** — this preserves the "Silence is better than noise" principle for PR interactions.
714
+
715
+ **A verifier's report may end with an `### Incidental findings` section** — what its runs tripped over on the way. The brief bounds the channel (zero extra budget, never self-confirmed, verdicts first); your half is to treat the entries as finder candidates, never as verdicts: dedup them against the cumulative list under Step 4's own rules, and merge the survivors in carrying the `— [unverified]` tag. At high effort they ride the next verification round exactly as Step 5's new findings do (`--round <k>` shards — a fresh verifier by construction, so no run ever confirms its own discovery); at medium effort, which has no later round, they surface terminal-only under "Needs Human Review" as low-confidence entries and are never posted. This channel exists because the maintainer verifications this step borrows its run capabilities from kept surfacing their sharpest notes as side effects of driving the stack — an error message no user could ever see, bookkeeping growing without bound, a test-plan step the product cannot exhibit — none reachable by reading the diff (measured; DESIGN.md — The three notes only a live stack surfaced (PR #9131)).
668
716
 
669
717
  ### Pattern aggregation
670
718
 
@@ -685,6 +733,7 @@ For each pattern group:
685
733
  - **Witness:** <the representative instance's witness — often the one sweep or probe that confirmed the whole pattern — or the group's shared `not run — <reason>` line; the witness rule reads an aggregate exactly as it reads a standalone finding>
686
734
  - **Suggested fix:** <general fix approach>
687
735
  - **Fix witness:** <the group's shared acceptance criterion — the test that must go red if the general fix is removed, or N/A>
736
+ - **Fix constraint:** <the existing fact the general fix must not violate, with its source — omit the line when none was observed>
688
737
  - **Severity:** <highest severity among the group>
689
738
 
690
739
  **Aggregation must not drop the anchors.** Each merged finding arrived with its own `Anchor`, and Step 7 posts one comment per location — so it needs one anchor per location, not one for the group. An aggregated entry sent to `resolve-anchors` with no `anchor` is a hard failure: the subcommand validates every entry and **throws on the whole batch**, so a single anchorless aggregate takes down the resolution of every other finding in the review. Carry the anchors through into the aggregate's `locations[]` — one entry per location, each with its own `anchor` — and Step 6's `findings --to-anchors` performs the expansion mechanically: one resolver request per location, ids suffixed `<id>-1`, `<id>-2`, … (resolutions are joined back to findings by id, so these must be unique — a suffix that collides with another finding's id is refused at projection, and the subcommand rejects duplicates besides).
@@ -712,18 +761,18 @@ One anomaly the builder flags but does not refuse (#9242): a per-chunk build on
712
761
  **The convergence pair — 3A (whole-diff form).** Rounds 1 and 2 launch **in one response** — together with Step 4's verifier shards (Step 4 names this) — each built by its own `agent-prompt` call: `--round 1` and `--round 2`, the **same** `--findings` file. This is not a loosened criterion; it is the serial shape's own arithmetic made concurrent: a dry round leaves the cumulative list unchanged, so round 2's launch input was already substantively identical to round 1's — the same entries, at most with verification tags the merge had cleared in between — an independent rerun that the serial shape bought with a full round of wall clock, and that one budget-gated run could no longer afford at all, shipping a capped verdict for want of a second dry audit it had time to run in parallel but not in series (measured; DESIGN.md — The serial convergence pair). What the two-consecutive-dry criterion demands is unchanged: two independent, substantively-dry audits of the whole diff. The one delta the pair does introduce is the same one-round suppression window the pipelined loop already accepts (the merge bullet in the termination rules): the round-2 member audits with entries a verifier may be rejecting mid-flight still on its do-not-re-report list.
713
762
 
714
763
  - **Both members dry** (substantive receipts, per the termination rules): the audit has converged. Wait for the riding verifiers' verdicts, apply the final merge, and proceed to Step 6.
715
- - **Either member reports findings**: the pair is one reporting round. Its members could not see each other, so first dedup the pair against itself (same defect, same location, same root cause keeps one, at the highest severity), merge into the cumulative list, and continue serially: the pair's verifiers ride with round 3's auditor — verify builds over the **deduped union**, sharded per Step 4's `verifyShard` exactly as any reporting round's findings are, **every shard passed as `--round 2`** (the pair's later label; never one build per member — the dedup already merged cross-member findings, and a per-member split would put one entry in front of two verifiers) — and convergence now needs two consecutive dry rounds from round 3 on. A dry member of a reporting pair is **not** carried forward as half of that evidence — its dry predates the other member's findings entering the list. One exception, and it is the retroactively-dry rule below, not a third rule: if a later merge retires the pair in full — every finding from both members rejected — the pair counts as the dry predecessor, and round 3's dry return ends the loop.
764
+ - **Either member reports findings**: the pair is one reporting round. Its members could not see each other, so first dedup the pair against itself (same defect, same location, same root cause keeps one, at the highest severity; a `fixWitness`/sourced `fixConstraint` on either copy survives the merge — Step 4's rule), merge into the cumulative list, and continue serially: the pair's verifiers ride with round 3's auditor — verify builds over the **deduped union**, sharded per Step 4's `verifyShard` exactly as any reporting round's findings are, **every shard passed as `--round 2`** (the pair's later label; never one build per member — the dedup already merged cross-member findings, and a per-member split would put one entry in front of two verifiers) — and convergence now needs two consecutive dry rounds from round 3 on. A dry member of a reporting pair is **not** carried forward as half of that evidence — its dry predates the other member's findings entering the list. One exception, and it is the retroactively-dry rule below, not a third rule: if a later merge retires the pair in full — every finding from both members rejected — the pair counts as the dry predecessor, and round 3's dry return ends the loop.
716
765
  - The substantive-return check applies per member, relaunch-once included. A twice-whiffed member makes the pair not dry — silence is not convergence evidence — and its scope joins the outstanding-whiffed-scopes list exactly as for any round.
717
766
  - If the deadline gate refuses one of the pair's builds (exit 4) and admits the other, launch the admitted member alone and treat the refusal as the budget stop it is (the termination rules below). If it refuses BOTH builds, nothing launches: the remaining budget cannot cover even one round plus the reserve, the first refusal's stop marker is the stop, and the two refusals each name their own round's stop entry — proceed to Step 6 and relay the MARKER's entry only (it holds the first refusal, and it is the one `compose-review` renders). The single-refusal split is defensive only: while the runtime's tool-concurrency pool holds both whole-diff members at once, the gate prices the paired round 2 at one round's wall, so it admits no dearer than the round 1 just admitted and that split cannot currently fire — the rule exists so a future pricing change degrades to the serial shape instead of to a guess.
718
767
 
719
768
  **The convergence pair — 3B (per-chunk form).** On 3B the pair applies per chunk. Launch `--all-chunks --round 1` **and** `--all-chunks --round 2` **in the same response** — both fan out to every chunk (rounds 1 and 2 always do, and the retirement schedule only reads history from round 3, so round 2's build needs nothing round 1 has produced yet), so each chunk's two establishing audits run concurrently instead of a round-wall apart. This is the same arithmetic as 3A read per territory: a chunk dry in round 1 leaves its slice of the cumulative list unchanged, so that chunk's round-2 auditor re-runs substantively the same audit — one round's wall the serial shape paid on every chunked review (measured; DESIGN.md — The serial 3B convergence rounds). The convergence contract is unchanged and reads per chunk through the retirement ledger: a chunk dry in both members holds its two-consecutive-dry certificate, and a pair dry on **every** chunk converges at the round-3 `--all-chunks` build (`CONVERGED`, exit 5) exactly as an all-dry pair does on 3A. Same one-round suppression window, per chunk (a round-2 auditor audits with entries a verifier may be clearing mid-flight). The launch coupling holds too: both members ride with the Step 4 verifier shards (Step 4 names this).
720
769
 
721
- - **Any auditor in either member reports findings**: the pair is one reporting round, exactly as on 3A — wait for BOTH fan-outs to return in full before the dedup (every chunk has an auditor in each member, and members cannot see each other across rounds either), dedup the pair against itself across rounds **and** chunks (same defect, same location, same root cause keeps one, at the highest severity), and merge the union into the cumulative list once. The pair's verifiers ride round 3's `--all-chunks` build: one batch over the **deduped union**, sharded per Step 4's `verifyShard`, **every shard passed as `--round 2`** (the pair's later label — never one build per member). Round 2's auditors are already in flight when round 1's returns land, so the pipelined k/k+1 rule below does not launch them again; this bullet is the pair's only transition. Convergence then reads per chunk through the retirement ledger as above: a chunk that reported in either member holds no certificate and stays under every-round audit, and the pair counts as one reporting round for the retroactively-dry rule — retired only when every finding from **both** members is rejected.
770
+ - **Any auditor in either member reports findings**: the pair is one reporting round, exactly as on 3A — wait for BOTH fan-outs to return in full before the dedup (every chunk has an auditor in each member, and members cannot see each other across rounds either), dedup the pair against itself across rounds **and** chunks (same defect, same location, same root cause keeps one, at the highest severity; a `fixWitness`/sourced `fixConstraint` on any copy survives the merge — Step 4's rule), and merge the union into the cumulative list once. The pair's verifiers ride round 3's `--all-chunks` build: one batch over the **deduped union**, sharded per Step 4's `verifyShard`, **every shard passed as `--round 2`** (the pair's later label — never one build per member). Round 2's auditors are already in flight when round 1's returns land, so the pipelined k/k+1 rule below does not launch them again; this bullet is the pair's only transition. Convergence then reads per chunk through the retirement ledger as above: a chunk that reported in either member holds no certificate and stays under every-round audit, and the pair counts as one reporting round for the retroactively-dry rule — retired only when every finding from **both** members is rejected.
722
771
  - If the deadline gate refuses one member's `--all-chunks` build (exit 4) and admits the other's, launch the admitted member alone and take the stop. The gate prices the round-2 build as the pair's wall — both fan-outs in waves of the runtime's tool-concurrency pool — so this split fires exactly when the pair plus the reserve does not fit but one round still does, and the admitted round alone keeps the serial shape. If it refuses BOTH builds, nothing launches: the remaining budget cannot cover even one round plus the reserve, the first refusal's stop marker is the stop, and the two refusals each name their own round's stop entry — proceed to Step 6 and relay the MARKER's entry only (it holds the first refusal, and it is the one `compose-review` renders).
723
772
 
724
773
  **Do not write the reverse auditor's prompt. Ask for it — and hand it the findings so far so it prints the whole block:**
725
774
 
726
- Write **the cumulative list of every finding reported so far** (Steps 3-4 plus all prior rounds — verified or still under verification; entries a verifier rejected are removed) to a file, so the auditor hunts what is not already on it. **Every entry not yet through Step 4 carries a trailing `— [unverified]` tag** — added at the merge that admits it, removed by the merge after its verdict lands. An early round on a clean review may have nothing confirmed yet — pass the file anyway (empty is fine; the command tells the auditor so). Then:
775
+ Write **the cumulative list of every finding reported so far** (Steps 3-4 plus all prior rounds — finder and verifier-incidental entries alike, verified or still under verification; entries a verifier rejected are removed) to a file, so the auditor hunts what is not already on it. **Every entry not yet through Step 4 carries a trailing `— [unverified]` tag** — added at the merge that admits it, removed by the merge after its verdict lands. An early round on a clean review may have nothing confirmed yet — pass the file anyway (empty is fine; the command tells the auditor so). Then:
727
776
 
728
777
  ```bash
729
778
  # Step 3A (small diff): one auditor per round, the whole diff. The convergence
@@ -798,11 +847,12 @@ For each **individual** finding, include:
798
847
  2. **Source tag** — `[build]`, `[test]`, or `[review]`
799
848
  3. **What's wrong** — Clear description of the issue
800
849
  4. **Failure scenario** — the concrete trigger and wrong outcome (for quality findings, the concrete cost or the quoted rule)
801
- 5. **Witness** — for a Critical: the observed output that settled the verdict, trimmed to the deciding lines — the probe's two sides, the A/B's quote pair, the sweep count over the real population, the failing test text — or the verifier's `not run — <reason>` line (Step 4's witness rule). A Suggestion carries one when a run produced it; it is not owed one.
850
+ 5. **Witness** — for a Critical or Suggestion: the observed output that settled the verdict, trimmed to the deciding lines — the probe's two sides, the A/B's quote pair, the sweep count over the real population, the failing test text — or the verifier's `not run — <reason>` line naming the capability that came closest (Step 4's witness rule; both postable severities are held to it). A Nice to have carries one when a run produced it; it is not owed one.
802
851
  6. **Suggested fix** — Concrete code suggestion when possible
803
852
  7. **Fix witness** — the test that must go RED if that fix is removed (file + the behaviour it pins), or `N/A` when the fix adds no guard, branch or behaviour a test can pin. This is the ACCEPTANCE CRITERION for whoever fixes it, not the reviewer's evidence — `Witness` above is the evidence, and the two never substitute for each other.
853
+ 8. **Fix constraint** — an existing fact the fix must not violate, with its source (the quoted constant or `file:line`): a configured limit a new bound must stay within, a second site that reads the field a shape change touches, a uniqueness a newly shared resource's key must keep. `Fix witness` pins the fix's claim; this pins its premises — the class that passes a witnessed test and is still wrong. Omit it when none was observed — never `N/A` — and never without a source: a caution with no quoted fact ("be careful about concurrency") is not a constraint, and a wrong one is misdirection the fixer will follow.
804
854
 
805
- For **pattern-aggregated** findings, use the aggregated format from Step 4 (Pattern, Occurrences, Example, Failure scenario, Witness, Suggested fix, Fix witness, Severity) with the source tag added.
855
+ For **pattern-aggregated** findings, use the aggregated format from Step 4 (Pattern, Occurrences, Example, Failure scenario, Witness, Suggested fix, Fix witness, Fix constraint, Severity) with the source tag added.
806
856
 
807
857
  Group high-confidence findings first. Then add a separate section:
808
858
 
@@ -820,7 +870,7 @@ If there are none of these, omit this section.
820
870
 
821
871
  ### Previous round's findings (incremental re-review only)
822
872
 
823
- The ledger has two sources, in priority order: **the PR itself** — `pr-context` recovers the machine ledger embedded in this account's last posted review and renders it as the "Previous /review round (machine ledger)" section (also written beside the context file as `qwen-review-pr-<n>-prev-ledger.json`) — and, as fallback for rounds that never posted, the local cache. The PR copy is authoritative because it survives what the cache cannot: CI, another machine, a fresh clone. **This ruling section runs at medium effort too** — recovering the ledger costs nothing (pr-context already fetched the reviews), and a re-review that ignores what it told the author last round is the amnesia this exists to end; medium still writes no cache and posts nothing, exactly as before. When either source loaded a ledger, this review is **round N+1 of the same PR**, and the single most useful thing it can tell the reader is what happened to round N's findings — a re-reviewer who only lists new findings leaves the author to diff two reports by hand. Rule on **every** ledger entry against the code at the reviewed commit, exactly the way the open-Criticals re-check below rules (trace the mechanism; the diff containing a fix is not the same claim as the defect no longer firing):
873
+ The ledger has two sources, in priority order: **the PR itself** — `pr-context` recovers the machine ledger embedded in this account's last posted review and renders it as the "Previous /review round (machine ledger)" section (also written beside the context file as `qwen-review-pr-<n>-prev-ledger.json`) — and, as fallback for rounds that never posted, the local cache. (A local or file-path review has no PR to post to, so its only source IS its cache — read at the plan's `cachePath`, never a name you compute (the capture derives `target` inside itself, and a file review's cache name carries a source-path digest besides); Step 1's incremental check already read it; the rulings below apply to its entries unchanged.) The PR copy is authoritative because it survives what the cache cannot: CI, another machine, a fresh clone. **This ruling section runs at medium effort too** — recovering the ledger costs nothing (pr-context already fetched the reviews), and a re-review that ignores what it told the author last round is the amnesia this exists to end; medium still writes no cache and posts nothing, exactly as before. When either source loaded a ledger, this review is **round N+1 of the same PR**, and the single most useful thing it can tell the reader is what happened to round N's findings — a re-reviewer who only lists new findings leaves the author to diff two reports by hand. Rule on **every** ledger entry against the code at the reviewed commit, exactly the way the open-Criticals re-check below rules (trace the mechanism; the diff containing a fix is not the same claim as the defect no longer firing):
824
874
 
825
875
  - **fixed** — the mechanism can no longer fire. Say so, by id, in one line: `R1-2 fixed by <what>`. Do not re-report it as a finding. The sibling-entrance rule from the re-check below applies here unchanged: for a divergence-class entry, `fixed` is a ruling about the family's entrances, checked one by one — for a **bounded** family a still-open sibling becomes a fresh `R<round>-<n>` entry (for an unbounded surface, apply the bounded/unbounded rule below instead of filing the sibling), never a reason to withhold the original's `fixed`.
826
876
  - **still stands** — re-report it **under its original id**, updating the location if the code moved. It keeps its severity; a still-standing Critical blocks exactly as a new one would. Write that id into the re-report itself, immediately after the severity marker — `**[Critical]** R1-2: <the claim>` — and into the body entry if it cannot be anchored (`R1-2 <the claim>`). That prefix is not decoration: `compose-review` reads it back out of the comment when it builds the marker, and it is the only way an id survives into the machine ledger the next round recovers. Omit it and the same claim comes back renumbered, which is exactly what carrying the id forward exists to prevent.
@@ -850,7 +900,7 @@ Render the rulings as a short table at the top of the Findings section — id, o
850
900
 
851
901
  **Resolve the floor first.** The Step 1 verdict's `severityFloor` is `critical`, `suggestion`, or `auto`. Explicit values are the operator's call: `critical` applies the Critical-only posture from round 1; `suggestion` turns the posture **off** — every round posts Suggestions, and the code-age rule below does not run. `auto` — the default — resolves here, where the round is known: **this review is round `prev ledger round + 1`**, and the round that decides the posture is the SIDE FILE's — the same read `compose-review` stamps into the marker and the deferral clause; the local cache's round scopes the diff but never decides the posture, or the body and the marker would disagree about which round ran (no recovered ledger → round 1 → no posture). Through round 5 the floor is `suggestion`; **from round 6 it is `critical`** — **and it is `critical` from ANY round once the side file's `flatRounds` is at its bar of 2**. That streak is the signal-driven early trigger: `compose-review` measures each round's first-time-finding rate against the previous round's, stamps the consecutive not-falling count into the marker as `flatRounds`, and engages the floor ahead of schedule when the count reaches 2 — acting on the convergence paragraph's own "drop to `--severity-floor critical`" advice instead of only printing it. You cannot evaluate that trend yourself (it is a deterministic join over the ledger, which is exactly why the module owns it), so your routing follows the **marker**: `flatRounds >= 2` in the side file means the floor is `critical` for this round and every later round of this PR — route Suggestions to the deferral channel accordingly. On the round the streak first reaches the bar you will usually have drafted under the open posture; the enforcement backstop below moves those Suggestions mechanically and the posted body discloses the move with the streak that armed it — that is the trigger working, not a lost finding. Once engaged the trigger **latches**: the streak is pinned in the marker rather than re-measured (the floor itself quiets the posted-set trend it reads), so it does not release on a quiet round — an explicit `--severity-floor suggestion` remains the only way back to full posting. In the **context-unavailable** state the round is unknowable — the ledger this rule counts from could not be recovered by a run that could not read the PR — so treat `auto` as round 1: no posture, full posting, and say so in the terminal report (the deterministic marker still stamps its own count from the side file; a posting bar in doubt fails open, bookkeeping does not). Carry the **verdict's `severityFloor` into the compose state UNRESOLVED** — explicit values as they are, and `auto` as the literal string `auto`, never as the level it resolved to this round: the module licenses `auto` by the round it derives itself, and a round-resolved `suggestion` is indistinguishable from the operator's explicit posture-off override — passing it would turn every legal rounds-2–5 age-rule deferral into an unlicensed one. The resolution in this paragraph decides what YOU post; the state field carries the policy. **The module also enforces the floor itself**: a Suggestion still drafted inline past a resolved `critical` floor is moved into the deferral list mechanically by `compose-review`/`submit` (the composed result's `floorEnforced` names the moved indices, the posted body discloses the move, and `submit` drops those comments from the write). Your Step 6 routing stays the primary path — the enforcement is the backstop that keeps the posted set lawful when the routing drifts, so a submit report showing fewer inline comments than you drafted under a critical floor is the floor working, not a lost finding. Three consequences of it being mechanical: the backstop classifies by the drafted severity MARKER alone — it cannot re-derive confidence or a Nice-to-have, so keeping low-confidence and Nice-to-have findings OUT of the drafted comments (as this step already mandates) is what keeps them out of the published deferral list too; **leave moved comments IN the comments file and the submit payload** — the CLI removes them from the write itself, and hand-removing them "to match" makes both boundaries recompute over the reduced set and erases the deferral record the move exists to keep; and the floor it enforces is the RESOLVED one (an explicit `critical`, `auto` from round 6, or `auto` with the `flatRounds` streak at its bar), recovered where possible from the CLI's own record of the invocation rather than the state field alone.
852
902
 
853
- **At floor `critical`, a non-Critical finding that would otherwise post is recorded, not requested.** The deferrable set is exactly the set the floor takes away: **high-confidence Suggestions** — the findings a `suggestion`-floor round would have drafted inline. Low-confidence findings and Nice-to-haves were never posted at any floor and **stay terminal-only exactly as before**: routing them through the deferral list would _publish_ to the PR what the review contract keeps out of it, and inflate the list the posture exists to keep small. A deferred finding has been through Step 4 like any posted one — the deferral list publishes its one-line claims in the body, so `compose-review`'s verifier-delivery floor counts deferred findings exactly as posted ones; an unverified claim does not become publishable by being deferred. (Deterministic findings are the exception on both sides at once: a `[build]`/`[test]`/`[probe]` finding is pre-confirmed, Step 4 launches no verifier for it, and the floor excludes it — by its `source` field.) Each deferred finding stays in the findings artifact and the terminal report under its own grouping — "Deferred (convergence posture)" — and enters the compose state's `deferredSuggestions` as a **TYPED entry, one object per finding, copied from the artifact's own fields**: `{"file": "src/a.ts", "line": 42, "source": "test", "severity": "Suggestion", "title": "mutation survivor on the retry guard"}` (`line` optional; a pattern aggregate adds `"locations": N` for its further locations). This is a data field, not a sentence: `compose-review` derives deterministic from `source`, relocates a `severity: "Critical"` entry into the body Criticals (a Critical is never deferred), refuses a `"Nice to have"` (terminal-only) or any malformed entry, and RENDERS the human line `file:line — [source] title` itself — never write that line into the state, and never re-type the fields: read them out of the findings artifact you just wrote. It is **not** drafted into the `comments` array, **not** counted toward `S`, and casts no vote on the event: `compose-review` renders the list as a disclosed, non-capping paragraph — up to 20 entries, each capped at 240 characters, with an overflow count pointing at the run report — so the deferral is on the PR record without opening a thread that regenerates a round, and anything past the rendered cap survives in full in the findings artifact and the terminal report (say so there when the cap trims the list). A previous-round **non-Critical** ledger entry that still stands is ruled in the status table as `still stands — deferred (convergence posture)` and is likewise not re-posted; it leaves the machine ledger (`buildLedger` ingests only posted findings), and the deferral list plus the original round's thread remain its record. **A Critical is never deferredany round, any floor**: new Criticals post, still-standing ledger Criticals re-post under their original ids, and every Critical ruling above runs unchanged. An APPROVE composed over a non-empty deferral list opens "No blocking issues" instead of "No issues found" — `compose-review` owns that wording.
903
+ **At floor `critical`, a non-Critical finding that would otherwise post is recorded, not requested.** The deferrable set is exactly the set the floor takes away: **high-confidence Suggestions** — the findings a `suggestion`-floor round would have drafted inline — plus, at floor `critical` only, the fails-closed/new-surface Criticals described below. Low-confidence findings and Nice-to-haves were never posted at any floor and **stay terminal-only exactly as before**: routing them through the deferral list would _publish_ to the PR what the review contract keeps out of it, and inflate the list the posture exists to keep small. A deferred finding has been through Step 4 like any posted one — the deferral list publishes its one-line claims in the body, so `compose-review`'s verifier-delivery floor counts deferred findings exactly as posted ones; an unverified claim does not become publishable by being deferred. (Deterministic findings are the exception on the verifier's side, and for Suggestions on the floor's side too: a `[build]`/`[test]`/`[probe]` finding is pre-confirmed, Step 4 launches no verifier for it, and the floor's source exclusion leaves a deterministic Suggestion inline — by its `source` field; a deterministic Critical the axes classify defers like any other axes-Critical, its source riding the entry.) Each deferred finding stays in the findings artifact and the terminal report under its own grouping — "Deferred (convergence posture)" — and enters the compose state's `deferredSuggestions` as a **TYPED entry, one object per finding, copied from the artifact's own fields**: `{"file": "src/a.ts", "line": 42, "source": "test", "severity": "Suggestion", "title": "mutation survivor on the retry guard"}` (`line` optional; a pattern aggregate adds `"locations": N` for its further locations). This is a data field, not a sentence: `compose-review` derives deterministic from `source`, relocates a `severity: "Critical"` entry into the body Criticals unless it is the fails-closed, new-surface shape at floor `critical` (below), refuses a `"Nice to have"` (terminal-only) or any malformed entry, and RENDERS the human line `file:line — [source] title` itself — never write that line into the state, and never re-type the fields: read them out of the findings artifact you just wrote. It is **not** drafted into the `comments` array, **not** counted toward `S`, and casts no vote on the event: `compose-review` renders the list as a disclosed, non-capping paragraph — up to 20 entries, each capped at 240 characters, with an overflow count pointing at the run report — so the deferral is on the PR record without opening a thread that regenerates a round, and anything past the rendered cap survives in full in the findings artifact and the terminal report (say so there when the cap trims the list). A previous-round **non-Critical** ledger entry that still stands is ruled in the status table as `still stands — deferred (convergence posture)` and is likewise not re-posted; it leaves the machine ledger (`buildLedger` ingests only posted findings), and the deferral list plus the original round's thread remain its record. **A Critical is deferred by its axes, never by its severity and only at floor `critical`.** The severity bit alone carried three decisions in one — which way the defect fails, what it is measured against, how often it triggers — and past the convergence rounds everything that mattered still landed on the floor, so the floor filtered nothing and the loop oscillated instead of settling (measured; DESIGN.md — The floor that could not floor (#9659)). Two of those axes now travel with the finding (Step 4's verifier states them off its witness; the artifact carries them as `direction` and `baseline`), and the floor reads them: a Critical whose artifact entry carries `direction: fails-closed` AND `baseline: new-surface` — the change narrows what works, in a surface the merge base never had, so merging it certifies nothing false and regresses nothing — is recorded, not requested, exactly like a Suggestion: a typed `deferredSuggestions` entry with `severity: "Critical"` and both axes copied from the artifact, under its own `D<round>-<n>` artifact id, its `title` opening with the original `R<round>-<n>` id when it carries a still-standing entry forward (the closure mint reads the id there — an id-less re-post silences that round's lineage). Every other Critical posts: `certifies-falsely` at either baseline (the code lies — that is the core promise broken, whatever surface it lives in), `regression` in either direction (the merge base did it right, and a merge gate grades against the merge base), a Critical with either axis missing or self-contradicting (the floor cannot classify it, and a blocker in doubt posts), and every Critical at any floor below `critical` — the rounds-2–5 code-age rule never touches a Critical. `compose-review` holds the same rule in code: a `Critical` entry that is not both `fails-closed` and `new-surface`, or that arrives when the floor is not in effect, is relocated into the body Criticals and posts; and the enforcement backstop moves a drafted `**[Critical]**` comment whose claim line carries both the `[fails-closed]` and `[new-surface]` tags (Step 7 puts them there from the artifact) exactly as it moves a Suggestion, naming the move by severity in the disclosure. The deferred Critical's record is the same as a deferred Suggestion's — the posted deferral line (which names it `Critical` and shows its tags), the findings artifact entry, the terminal report — and it is follow-up work the author files as an issue, not work this round requests; no issue is filed by the review. Everything else about Criticals is unchanged: new Criticals that post still post, still-standing ledger Criticals re-post under their original ids, and every Critical ruling above runs unchanged — with one addition to the routing: the side file's work-list table shows a carried Critical's recorded axes beside its severity (`Critical (fails-closed, new-surface)`), so a still-standing entry of that shape at a `critical` floor goes to the deferral channel rather than being re-posted. An APPROVE composed over a non-empty deferral list opens "No blocking issues" instead of "No issues found" — `compose-review` owns that wording.
854
904
 
855
905
  **Rounds 2–5 carry a narrower gate: the code-age rule.** With an `auto` floor resolved to `suggestion` — never under an explicit `--severity-floor suggestion`, which turns the posture off, this rule included — a **new otherwise-postable finding — the same deferrable set as above, high-confidence Suggestions only, never low-confidence or Nice-to-have entries** — anchored on code **unchanged since the previous round's reviewed head** is deferred the same way — the previous round read that code and did not flag it, so filing a nit on it now is re-derivation churn, not signal. (Carried-forward entries keep their original ids and are not "new"; this gates first appearances only.) The age reference is the side file's `commitId` — the previous review's own `commit_id`, set by GitHub when the round posted. It is an **age reference, never an incremental anchor**: the ledger's `sha` stays the only range certification, withheld on fail-closed rounds on purpose, while `commit_id` exists on every posted round — a posting bar needs a reference point, not a certification, which is exactly why a full-range re-review (still the shape whenever no own anchor is usable — no own marker on the PR carries one, the graft's certifier mismatches this round's identity, or the markers predate the field) can still apply this rule. Validate it inside the worktree — `git cat-file -e <commitId>^{commit}` and `git merge-base --is-ancestor <commitId> HEAD` — and decide age with `git --literal-pathspecs diff <commitId>..HEAD --unified=0 -- '<file>'`: a finding whose anchor line falls inside a changed hunk is new-code and posts. **Two diff-output doubt states fail OPEN like every other arm, never toward suppression**: run the command from the worktree ROOT, and before reading its silence, prove the pathspec matches — `git cat-file -e HEAD:'<file>'` (tree-relative, cwd-independent); a non-matching pathspec means the diff's emptiness is about the PATH, not the code — skip the age rule for that finding, it posts. And a NON-empty diff with zero `@@` hunks (a `.gitattributes` `binary`/`-diff` mark, which the PR controls) is a file-level CHANGE — the finding posts; only a matching pathspec with a genuinely empty diff reads as unchanged. **A pattern aggregate is aged per location**: it posts (as the usual aggregated comment) if ANY of its `locations[]` falls inside a changed hunk — the changed entrance is new-code and must not ride out a round inside a deferral line — and defers only when EVERY location is unchanged and covered; its deferral line names the root anchor with the location count (`a.ts:10 (+2 locations)`). **Both operands are hostile-input-hardened, and neither hardening is optional.** The path is PR-controlled: unquoted, a filename like `x;touch PWNED` ends the argument and executes the tail as a command, so the path rides in single quotes (a `'` inside the name becomes `'\''`); and without `--literal-pathspecs` (a global option — it must precede `diff`) a name carrying glob metacharacters is a wildcard pathspec, so `foo[1].ts` matches the _sibling_ `foo1.ts` and the finding is aged against the wrong file's hunks. **The rule also needs the previous round to have actually read the code it vouches for.** Its premise is "the previous round saw this code and did not flag it" — so before deferring, check the previous round's own review body: **the review whose id the side file's `reviewId` names** (pr-context renders review bodies whole up to an 8,000-character cap, with a fetch note at the cut; with several summaries on the PR, the id decides which body's disclosures bind — checking a different body can vouch for code the true previous round never read). A body whose render carries the truncation note is consulted only after running that note's fetch, redirected to a file exactly as the blocker re-check prescribes — a "Not reviewed" disclosure past the cap is invisible, and ruling on the visible prefix would defer a finding on code nobody read. A body that cannot be read whole: skip the age rule. One absence is benign and decided, not skipped: a previous round that converged clean posts the canonical LGTM body, which pr-context filters from the render — that body has no disclosures BY DEFINITION (a capped or partial round never composes it), so a `reviewId` whose body is absent because it matched the canonical LGTM filter is disclosure-free, and the age rule proceeds. A finding whose file falls in scope that round disclosed as not reviewed — a named unread chunk or dimension covering it, or the scope-wide "could not certify that any of this diff was reviewed" opener — gets no age suppression; the premise is false there, and a first-time Suggestion in code nobody read must post like any round-1 finding. When the `commitId` field is absent (older rounds, or a run whose recovery came up empty — pr-context strips a stale file's `commitId` then), the recorded `commitId` fails the validation above (rebase), there is no worktree (lightweight mode), or Step 1 set the **context-unavailable** state (this run's pr-context failed, so the side file may be a previous run's leftovers), **skip the age rule, not the review** — full posting, exactly as before. The Exclusion Criteria's newly-reachable exception extends across rounds unchanged: a finding on unchanged code that this round's changes make **newly reachable or newly wrong** is new-code by that fact, and posts.
856
906
 
@@ -938,9 +988,9 @@ Write every confirmed finding — high and low confidence alike — as a JSON ar
938
988
 
939
989
  **One finding, one name.** A high-effort PR review also writes the incremental cache's cross-round `findings` ledger (Step 8), whose ids are `R<round>-<n>` — use those same ids here: a finding that will enter the ledger gets its `R<round>-<n>` as the artifact `id`, and a carried-forward finding keeps the id it already has. Two id schemes for one finding is how "R1-2" in next round's report and "f7" in this round's outcome ledger turn out to be the same defect that nobody can join. A finding the convergence posture deferred is still a confirmed finding and enters this artifact with all its fields — the deferral is a posting decision recorded in the compose state, never a severity change and never a reason to leave the artifact — but under its own id sequence, `D<round>-<n>`, **never consuming an `R<round>-<n>`**: the `R` counter must predict `buildLedger`, which numbers POSTED findings only, and a deferred finding holding `R6-2` would hand next round a ledger whose `R6-2` names a different defect than this round's artifact — the exact join "one finding, one name" exists to keep.
940
990
 
941
- Each entry carries `id` (unique — outcomes and resolved anchors both join on it), `severity`, `confidence`, `source`, `summary`, `failureScenario`, and either `file`/`line`/`anchor` or, for a pattern aggregate, a `locations[]` array with **one entry per location** (`suggestedFix`, `fixWitness`, `category`, `shortSummary` and `witness` are optional; `shortSummary` is derived from `summary` when absent; `witness` is the Step 4 witness — the executed evidence, or its `not run — <reason>` line — carried as data so the report and the comment bodies quote one recorded string instead of transcribing it twice more; `fixWitness` is the acceptance criterion the finding format asks for — the test that must go red if the suggested fix is removed, or `N/A` — carried for the same reason and read back by Step 7's comment body). The command validates the shape, refuses a duplicate id, refuses a finding with no failure scenario, sorts by severity → confidence → file → line → id, and writes counts nobody then recomputes by hand. Read the artifact for the numbers you quote in the Summary. This is a **canonicalization**, not a gate: it does not decide the verdict — `compose-review` does that, from the same findings — and it does not run at low effort, where the pass is unverified and emits no verdict.
991
+ Each entry carries `id` (unique — outcomes and resolved anchors both join on it), `severity`, `confidence`, `source`, `summary`, `failureScenario`, and either `file`/`line`/`anchor` or, for a pattern aggregate, a `locations[]` array with **one entry per location** (`suggestedFix`, `fixWitness`, `fixConstraint`, `category`, `shortSummary`, `witness`, `direction` and `baseline` are optional; `shortSummary` is derived from `summary` when absent; `direction` (`certifies-falsely` | `fails-closed`) and `baseline` (`regression` | `new-surface`) are the two decision axes Step 4's verifier stated for a confirmed Critical, copied exactly — an axis the verifier omitted stays absent, and a misspelled one is refused; `witness` is the Step 4 witness — the executed evidence, or its `not run — <reason>` line — carried as data so the report and the comment bodies quote one recorded string instead of transcribing it twice more; `fixWitness` is the acceptance criterion the finding format asks for — the test that must go red if the suggested fix is removed, or `N/A` — carried for the same reason and read back by Step 7's comment body; `fixConstraint` is the existing fact the fix must not violate, with its source — present only when the finder observed one, with no `N/A` form (the command drops the literal), and read back by the same comment body). The command validates the shape, refuses a duplicate id, refuses a finding with no failure scenario, sorts by severity → confidence → file → line → id, and writes counts nobody then recomputes by hand. Read the artifact for the numbers you quote in the Summary. This is a **canonicalization**, not a gate: it does not decide the verdict — `compose-review` does that, from the same findings — and it does not run at low effort, where the pass is unverified and emits no verdict.
942
992
 
943
- **Then speak the same list to the client, in-band — one `report_findings` tool call.** The artifact is the canonical record, but it is a file on disk registered after the fact (Step 8); every client rendering this session live — the TUI, the Web Shell transcript, an ACP host — otherwise sees only the prose restatement, which is the transcription surface the artifact exists to close. Immediately after the artifact is written, call the `report_findings` tool once (load it via `tool_search` if it is not in your tool list) — each call replaces the whole list, and Step 6B re-issues it with outcomes after a fix run — with `level` set to this review's effort and one entry per finding **copied from the artifact you just wrote** — `id`, `severity`, `confidence`, `source`, `file`/`line` (a pattern aggregate passes its first location; the artifact keeps the rest), `summary`, `shortSummary`, `failureScenario`, `category` — never re-typed from the terminal prose: the artifact is the oracle, and a re-derived severity here is the same drift the marker rule below closes. A finding the convergence posture deferred is still a finding — report it under its `D<round>-<n>` id like any other. **The tool's contract is harder-bounded than the artifact's, and a violation refuses the whole call**: at most 50 findings, with per-field length caps the schema states. When the artifact outgrows those bounds, do not let the call die on them — pass the first 50 findings in artifact order (the artifact is already sorted most-severe-first) and say in the terminal summary how many the cap cut, and shorten an over-cap `summary`/`failureScenario` — or `outcomeNote` on the Step 6B re-report — to fit rather than dropping the entry (the artifact keeps the full-length text, so nothing is lost by a delivery-only shortening). This is the one sanctioned departure from copy-verbatim, and it is a departure of length only, never of severity, confidence, or meaning — a bounded list delivered beats a complete list refused. This call is UI delivery, not bookkeeping: it persists nothing and decides nothing, and a failure (or an environment where the tool is not registered and `tool_search` cannot find it) is disclosed and moved past — never a reason to touch the artifact, the compose state, or the verdict, exactly the rule `record_artifact` follows in Step 8.
993
+ **Then speak the same list to the client, in-band — one `report_findings` tool call.** The artifact is the canonical record, but it is a file on disk registered after the fact (Step 8); every client rendering this session live — the TUI, the Web Shell transcript, an ACP host — otherwise sees only the prose restatement, which is the transcription surface the artifact exists to close. Immediately after the artifact is written, call the `report_findings` tool once (load it via `tool_search` if it is not in your tool list) — each call replaces the whole list, and Step 6B re-issues it with outcomes after a fix run — with `level` set to this review's effort and one entry per finding **copied from the artifact you just wrote** — `id`, `severity`, `confidence`, `source`, `file`/`line` (a pattern aggregate passes its first location; the artifact keeps the rest), `summary`, `shortSummary`, `failureScenario`, `category`, `direction`, `baseline` — never re-typed from the terminal prose: the artifact is the oracle, and a re-derived severity here is the same drift the marker rule below closes. A finding the convergence posture deferred is still a finding — report it under its `D<round>-<n>` id like any other. **The tool's contract is harder-bounded than the artifact's, and a violation refuses the whole call**: at most 50 findings, with per-field length caps the schema states. When the artifact outgrows those bounds, do not let the call die on them — pass the first 50 findings in artifact order (the artifact is already sorted most-severe-first) and say in the terminal summary how many the cap cut, and shorten an over-cap `summary`/`failureScenario` — or `outcomeNote` on the Step 6B re-report — to fit rather than dropping the entry (the artifact keeps the full-length text, so nothing is lost by a delivery-only shortening). This is the one sanctioned departure from copy-verbatim, and it is a departure of length only, never of severity, confidence, or meaning — a bounded list delivered beats a complete list refused. This call is UI delivery, not bookkeeping: it persists nothing and decides nothing, and a failure (or an environment where the tool is not registered and `tool_search` cannot find it) is disclosed and moved past — never a reason to touch the artifact, the compose state, or the verdict, exactly the rule `record_artifact` follows in Step 8.
944
994
 
945
995
  **The severities in this artifact are the canonical ones — draft the inline markers and the compose state FROM it, not from the list you typed by hand.** Ordering alone does not close the loop: `compose-review` reads `comments.json` and `compose.json`, both hand-written, so a hold that lowered a severity here still ships as `**[Critical]**` in the payload if the marker was copied from the draft instead of the artifact. Read `severity` out of `findings.json` for every marker and for the body Criticals.
946
996
 
@@ -965,11 +1015,11 @@ Each entry carries `id` (unique — outcomes and resolved anchors both join on i
965
1015
  It prints a `Verdict:` line to stderr. **That line is the verdict — print it, and nothing else.** It writes nothing, posts nothing, and needs no authorisation, so run it on every verified review — **high and medium** — whether or not you are going to post. The state file is the same one Step 7 uses (every field is listed just below): your findings and the states you established — the body Criticals, the discarded suggestions, the `cannot tell` blockers, the unreviewed dimensions, the `planPath`, the `findingsPath` (high effort — the cumulative reverse-audit findings file, for the `— [unverified]` check), the presubmit flags, the model id. It does **not** take the coverage or the inline counts, and it **refuses** a state JSON carrying `criticalsInline`/`suggestionsInline`. It derives coverage from the harness's transcripts, and it **counts** the inline findings from `--comments`: write the drafted inline comments to that file first — the same `[{path, line, body, …}]` array the Step 7 payload will carry, each body opening with its `**[Critical]**`/`**[Suggestion]**` marker; a review with nothing anchored inline passes a file containing `[]`. A report-only run has read Approve over a blocker its own report listed (measured; DESIGN.md — The Approve over a relocated Critical); counted from the draft, that finding cannot fall out of the computation. **If the comment set changes after composing** — an anchor fails to resolve, a finding relocates to the body, a comment is dropped — update the comments file (and the state), and run `compose-review` again: the verdict must be computed from the set you actually post, and Step 7's `submit` recounts from the payload to hold you to it.
966
1016
 
967
1017
  - **Not `criticalsInline` / `suggestionsInline`.** `submit` counts those off the `**[Critical]**` / `**[Suggestion]**` prefixes of the comments you attached — a number beside a list is a number that can disagree with the list, and one did. A `state` that supplies either is refused.
968
- - `bodyCriticals` — descriptions of unmappable or 422-relocated Criticals (their only copy lives in the body; they count toward `C` like anchored ones); a `Critical` entry placed in `deferredSuggestions` is relocated here, never deferred.
1018
+ - `bodyCriticals` — descriptions of unmappable or 422-relocated Criticals (their only copy lives in the body; they count toward `C` like anchored ones); a `Critical` entry placed in `deferredSuggestions` is relocated here unless the floor is `critical` and the entry is `fails-closed` on `new-surface` (Step 6's posture section — the one Critical shape the floor defers); an entry whose finding carries a `fixWitness` or a `fixConstraint` appends the corresponding sentence, copied from the artifact — the only published copy of the finding must not post without the fix's witness or the premise it rests on; the deferral channel's disclosed line is the one exception: it carries neither, and the entry's full record, witness and constraint included, survives in the findings artifact.
969
1019
  - `suggestionsDiscarded` — how MANY Suggestions lost their anchors to offline validation or the 422 recovery: a count (non-negative integer). The list of discarded items itself is also accepted and counted by its length (`[]` is zero). They still count toward `S`: dropping every anchor must never upgrade the verdict.
970
1020
  - `suggestionsDroppedAsDuplicates` — one entry per **confirmed** Suggestion you did not re-post because it is already reported on the PR (a prior round, a concurrent reviewer, an overlap drop), each naming the finding and where it already lives — never the finding's own text: Suggestion text must never appear in the review `body`, because `.github/workflows/qwen-autofix.yml` does not filter review bodies, so a Suggestion copied into the body would be handed to the autofix bot (full rule in `references/posting.md`); the carve-out for this account is exactly that name + location, e.g. `R1-2 loose review-config pins — already reported (comment 3788857379)`. Use this INSTEAD of bumping `suggestionsDiscarded` for duplicate drops: the two render different sentences, and the discarded one asserts an anchor failure that never happened. They still count toward `S`.
971
1021
  - `cannotTellCriticals` — one line per existing PR Critical whose Step 6 re-check landed on `cannot tell` (location + what could not be determined).
972
- - `deferredSuggestions` — the findings the convergence posture deferred, as **typed entries** `{file, line?, source, severity, title, locations?}` copied from the findings artifact (Step 6's posture section — **high-confidence Suggestions that would otherwise post**, never low-confidence or Nice-to-have entries, which stay terminal-only; a `Critical` entry is relocated into the body Criticals, a malformed or free-text entry is refused). Deferred findings are **not** drafted into `comments` and are **not** counted toward `S` — the body renders them as a disclosed, non-capping list (up to 20 entries × 240 chars, overflow counted; the full set lives in the findings artifact), so the deferral is on the PR record without regenerating a review round. Non-deterministic entries **do** count toward the verifier-delivery floor — a deferred claim still publishes — while `source: build|test|probe` entries are excluded by that field exactly as body Criticals are by their tag: they are pre-confirmed, no verifier ever exists for them, and demanding one would cap the verdict with a gap no repair can close. A deferral never withholds the ledger anchor.
1022
+ - `deferredSuggestions` — the findings the convergence posture deferred, as **typed entries** `{file, line?, source, severity, direction?, baseline?, title, locations?}` copied from the findings artifact (Step 6's posture section — **high-confidence Suggestions that would otherwise post**, never low-confidence or Nice-to-have entries, which stay terminal-only; a `Critical` entry is relocated into the body Criticals unless it carries `direction: "fails-closed"` and `baseline: "new-surface"` under a floor resolved to `critical` — then it defers, Step 6's posture section; a malformed or free-text entry, or a misspelled axis, is refused). Deferred findings are **not** drafted into `comments` and are **not** counted toward `S` — the body renders them as a disclosed, non-capping list (up to 20 entries × 240 chars, overflow counted; the full set lives in the findings artifact), so the deferral is on the PR record without regenerating a review round. Non-deterministic entries **do** count toward the verifier-delivery floor — a deferred claim still publishes — while `source: build|test|probe` entries are excluded by that field exactly as body Criticals are by their tag: they are pre-confirmed, no verifier ever exists for them, and demanding one would cap the verdict with a gap no repair can close. A deferral never withholds the ledger anchor.
973
1023
  - `convergence` — this round's census from Step 6's fix-induced rule, as `{"fresh": N, "induced": M}`: how many defects this round newly identified (not the marker's `fresh`, which counts comments posted for the first time), and how many of those the fix-induced rule attributed to a previous entry's fix (the ATTRIBUTED count, not the count of findings on newly pushed lines). Two integers, `induced <= fresh`; a malformed pair, a float, a negative, or a numerator larger than its denominator is read as no census at all. **Omit the field when the round could not measure it** — absence, a malformed pair, and a census too small to be a trend (fewer than 4 `fresh`, zeros included) all carry the churn streak forward untouched; only a measured below-bar census with at least 4 `fresh` resets it — zeros written for an unmeasured round state a measurement the round never made, so omit them too. `compose-review` owns everything downstream: the bar (half or more of `fresh`, and at least 4 `fresh`), the streak it stamps into the marker as `churnRounds`, and the body Critical it files itself on the second round counted against the bar.
974
1024
  - `severityFloor` — the Step 1 verdict's floor, carried UNRESOLVED (`critical`, `suggestion`, or the literal `auto` — never `auto`'s per-round resolution, which would masquerade as the operator's explicit override). This is the deferral channel's licence check: a non-empty `deferredSuggestions` under an explicit `suggestion` floor (posture off) or on round 1 under `auto` (no posture, no age reference) is an unlicensed deferral — `compose-review` renders the list but CAPS the verdict and says so, the same fail-closed treatment as unreviewed scope: the findings stay visible, nothing certifies past them, and the round is never lost to a refusal.
975
1025
  - `planPath` — the plan report from Step 1. **Coverage is not an input.** `submit` recomputes it from the harness's transcripts, because a `coverage` object you typed is a document you write — and the last time this skill trusted one, it was fabricated.
@@ -1059,7 +1109,7 @@ Run the bundled cleanup subcommand:
1059
1109
  "${QWEN_CODE_CLI:-qwen}" review cleanup <target>
1060
1110
  ```
1061
1111
 
1062
- `<target>` is the same suffix used throughout (`pr-<n>`, `local`, or filename). The command removes the worktree at `.qwen/tmp/review-pr-<n>` (PR targets only), deletes the local branch ref `qwen-review/pr-<n>`, and clears any `.qwen/tmp/qwen-review-<target>-*` side files (review JSON, PR context, presubmit / findings reports). It is idempotent — missing files are silent OK. It is also lease-guarded: when another session still holds this PR's worktree lease, cleanup skips the target wholesale and prints a `note:` line saying so (#9205) — relay that note verbatim and leave the lease file alone; the holder's own cleanup releases it. For PR targets it first **audits the review window**: any issue comment the reviewing account posted — or edited — since `fetch-pr` opened the window (the boundary reaches back across drift restarts and a clock-skew allowance), and any **review** the account submitted that `submit`'s receipt does not vouch for, is flagged with `warning:` lines, because submit's one sanctioned write is receipt-recorded and never touches issue comments (Step 7's write ban) — so such a comment is most likely an external same-account write — something the user did by hand from another terminal, or **another workflow posting under the same account** (in CI the review shares the bot identity with precheck/triage; their marker-stamped comments are filtered out automatically, but this reading stays real for anything unmarked) — and is a write that bypassed the gate only if its content is this review's own output. On an **Aone target** the audit runs through the `a1` CLI and the ruling keys on comment ids instead of review ids, because there the sanctioned submit POSTS COMMENTS (the inline findings and the summary — Aone has no review object): any MR comment the authenticated account posted — or edited — inside the window whose id the submit receipt does not vouch for is flagged the same way (a marker-stamped comment is filtered as on GitHub; a submitted comment whose id was never read back is unvouchable and may draw a flag — over-flagging is the fail-safe direction). Because the default listing hides RESOLVED comments, the audit unions it with a `--resolved` query — a bypass posted-then-resolved inside the window is still flagged — but a resolved comment is judged by its CREATION only (a resolution bumps `updatedAt` exactly like an edit, so it is not edit evidence). Five disclosed residuals: an edit of a submit-posted (receipt-vouched) comment is outside the tripwire's sight; an edit of an UNVOUCHED pre-window comment is invisible once its discussion is resolved (a resolved comment is judged by its creation only — a resolution bump is not edit evidence); resolved replies have no a1 listing at all; the comment listing is unpaged (one `comment list` per query — if a1 caps a page, comments past the cap stay invisible); and `a1 repo mr approve` / `a1 repo mr edit` writes are banned in Step 7 but outside this tripwire's coverage (the recorded a1 surface exposes no listing an audit could query for them). **Relay those `warning:` lines verbatim in your terminal summary** — the user can dismiss their own comment; a bypass they were never told about, they cannot. The audit is best-effort: when it cannot run (offline, unauthenticated, no report) it says so once on stderr — `note: bypass audit skipped (…)` — so a skipped audit is never mistaken for a clean one. Also remove `.qwen/tmp/qwen-review-parse-args.json` and the session args directory `.qwen/tmp/s-<session>/` (the path from the `<skill-args>` note) — both are written before the target suffix is known, so the pattern above misses them. (Leave the args file in place if you had to fall back to writing it yourself and the run failed: it is the only record of what the review was actually asked to do.)
1112
+ `<target>` is the same suffix used throughout (`pr-<n>`, `local`, or filename). **A FILE review whose derived token collides with a RESERVED one — `local`, `pr`, or `pr-<n>` — must NOT run this command at all** (a repo-root directory or file literally named `local` — or `pr`, whose sweep prefix engulfs EVERY PR family and whose lease guard never runs on the bare token; the CLI refuses that one itself — derives exactly such a token): the sweep is a prefix match over a shared namespace, so `cleanup local` from a file review deletes a concurrent whole-tree round's live plan and its `-prompts` records mid-round, and neither is lease-guarded. Skip the command, remove only the artifacts you wrote (the plan `--out` and its `-prompts` directory, per the paragraph below), and say so in the terminal — leaking this target's other side files is the affordable side of that trade. (The general prefix-collision class, and the namespace fix that ends it, is tracked in issue #10057.) The command removes the worktree at `.qwen/tmp/review-pr-<n>` (PR targets only), deletes the local branch ref `qwen-review/pr-<n>`, and clears any `.qwen/tmp/qwen-review-<target>-*` side files (review JSON, PR context, presubmit / findings reports). It is idempotent — missing files are silent OK. It is also lease-guarded: when another session still holds this PR's worktree lease, cleanup skips the target wholesale and prints a `note:` line saying so (#9205) — relay that note verbatim and leave the lease file alone; the holder's own cleanup releases it. For PR targets it first **audits the review window**: any issue comment the reviewing account posted — or edited — since `fetch-pr` opened the window (the boundary reaches back across drift restarts and a clock-skew allowance), and any **review** the account submitted that `submit`'s receipt does not vouch for, is flagged with `warning:` lines, because submit's one sanctioned write is receipt-recorded and never touches issue comments (Step 7's write ban) — so such a comment is most likely an external same-account write — something the user did by hand from another terminal, or **another workflow posting under the same account** (in CI the review shares the bot identity with precheck/triage; their marker-stamped comments are filtered out automatically, but this reading stays real for anything unmarked) — and is a write that bypassed the gate only if its content is this review's own output. On an **Aone target** the audit runs through the `a1` CLI and the ruling keys on comment ids instead of review ids, because there the sanctioned submit POSTS COMMENTS (the inline findings and the summary — Aone has no review object): any MR comment the authenticated account posted — or edited — inside the window whose id the submit receipt does not vouch for is flagged the same way (a marker-stamped comment is filtered as on GitHub; a submitted comment whose id was never read back is unvouchable and may draw a flag — over-flagging is the fail-safe direction). Because the default listing hides RESOLVED comments, the audit unions it with a `--resolved` query — a bypass posted-then-resolved inside the window is still flagged — but a resolved comment is judged by its CREATION only (a resolution bumps `updatedAt` exactly like an edit, so it is not edit evidence). Five disclosed residuals: an edit of a submit-posted (receipt-vouched) comment is outside the tripwire's sight; an edit of an UNVOUCHED pre-window comment is invisible once its discussion is resolved (a resolved comment is judged by its creation only — a resolution bump is not edit evidence); resolved replies have no a1 listing at all; the comment listing is unpaged (one `comment list` per query — if a1 caps a page, comments past the cap stay invisible); and `a1 repo mr approve` / `a1 repo mr edit` writes are banned in Step 7 but outside this tripwire's coverage (the recorded a1 surface exposes no listing an audit could query for them). **Relay those `warning:` lines verbatim in your terminal summary** — the user can dismiss their own comment; a bypass they were never told about, they cannot. The audit is best-effort: when it cannot run (offline, unauthenticated, no report) it says so once on stderr — `note: bypass audit skipped (…)` — so a skipped audit is never mistaken for a clean one. Also remove `.qwen/tmp/qwen-review-parse-args.json` and the session args directory `.qwen/tmp/s-<session>/` (the path from the `<skill-args>` note) — both are written before the target suffix is known, so the pattern above misses them. (Leave the args file in place if you had to fall back to writing it yourself and the run failed: it is the only record of what the review was actually asked to do.) A FILE review's plan falls outside the pattern for the opposite reason: its `--out` is the one name you chose yourself — unique to your run, precisely because no lease guards a file review — so `cleanup` cannot know it and must never glob its family. The family therefore deliberately does NOT start with `qwen-review-` — the prefix every cleanup sweep matches — because a sweep could not tell a live concurrent plan from its own run's residue, and a target literally named `file` or `file-<X>` sweeping `qwen-review-file-*` deleted concurrent file reviews' live plans mid-round. Remove the plan `--out` you wrote, and the `-prompts` directory beside it — **unless the reverse-audit loop stopped without converging**: a `budget-stop.json` marker inside the `-prompts` directory is that stop's record, and the directory is then the only certification history there is — keep the plan AND the directory, and tell the user to remove them once diagnosed, exactly as cleanup's `Kept` line does for the swept families (#9206; this instruction is the file family's only remover, so the retention duty rides with it). On a converged run: `agent-prompt` records every launch prompt under the plan's own name with `.json` replaced by `-prompts`, so the record rides the one family cleanup must never glob, and nothing else removes it.
1063
1113
 
1064
1114
  This step runs **after** Step 7 and Step 8 to ensure all review outputs are saved before cleanup.
1065
1115
 
@@ -1106,7 +1156,7 @@ These criteria apply to both Step 3 (review agents) and Step 4 (verification age
1106
1156
  - Focus on the diff, not pre-existing issues in unchanged code.
1107
1157
  - Keep the review concise. Don't repeat the same point for every occurrence — use pattern aggregation.
1108
1158
  - When suggesting a fix, show the actual code change.
1109
- - A Critical you post carries its witness — the observed output that proved it — or says in one line why none could run (Step 4's witness rule).
1159
+ - A Critical or Suggestion you post carries its witness — the observed output that proved it — or says in one line why none could run, naming the capability that came closest (Step 4's witness rule).
1110
1160
  - A comment whose fix adds a guard or a branch asks for the test that pins it — one sentence naming the test that must go red without the fix (Step 7's fix-witness rule). Roughly a third of a re-review's findings are introduced by the fix round before it; the acceptance criterion is what closes them a round earlier.
1111
1161
  - Flag any exposed secrets, credentials, API keys, or tokens in the diff as **Critical**.
1112
1162
  - Silence is better than noise. If you have nothing important to say, say nothing.