@qwen-code/qwen-code 0.19.8 → 0.19.9-preview.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bundled/qc-helper/SKILL.md +1 -0
- package/bundled/qc-helper/docs/configuration/auth.md +40 -55
- package/bundled/qc-helper/docs/configuration/model-providers.md +224 -218
- package/bundled/qc-helper/docs/configuration/settings.md +68 -51
- package/bundled/qc-helper/docs/features/_meta.ts +1 -0
- package/bundled/qc-helper/docs/features/channels/dingtalk.md +12 -2
- package/bundled/qc-helper/docs/features/channels/overview.md +82 -7
- package/bundled/qc-helper/docs/features/channels/wecom.md +6 -0
- package/bundled/qc-helper/docs/features/code-review.md +109 -41
- package/bundled/qc-helper/docs/features/commands.md +45 -44
- package/bundled/qc-helper/docs/features/computer-use.md +77 -0
- package/bundled/qc-helper/docs/features/hooks.md +29 -0
- package/bundled/qc-helper/docs/features/sub-agents.md +21 -0
- package/bundled/qc-helper/docs/features/tool-use-summaries.md +19 -23
- package/bundled/qc-helper/docs/integration-vscode.md +2 -2
- package/bundled/qc-helper/docs/qwen-serve.md +66 -23
- package/bundled/qc-helper/docs/reference/keyboard-shortcuts.md +2 -2
- package/bundled/review/DESIGN.md +303 -46
- package/bundled/review/SKILL.md +670 -165
- package/chunks/MaxSizedBox-BEVSAWZ6.js +70 -0
- package/chunks/{StandaloneSessionPicker-EYRGH6PJ.js → StandaloneSessionPicker-2DA2MDDE.js} +81 -79
- package/chunks/{acpAgent-GRFTTRHW.js → acpAgent-G3WJU3PS.js} +1318 -465
- package/chunks/agent-DD3WSHXE.js +68 -0
- package/chunks/agent-headless-MHI7U2CU.js +62 -0
- package/chunks/{anthropicContentGenerator-CPF22MA7.js → anthropicContentGenerator-56U3YCP2.js} +187 -31
- package/chunks/{artifact-tool-D233YJFH.js → artifact-tool-MIZLXBLF.js} +5 -5
- package/chunks/{askUserQuestion-MKL5ZTMC.js → askUserQuestion-XMJVPNWM.js} +6 -6
- package/chunks/bridge-HKEWPNB5.js +78 -0
- package/chunks/{ca-MKD32FQL.js → ca-C2JTT2N5.js} +68 -6
- package/chunks/channel-worker-group-K7QKH6IG.js +13 -0
- package/chunks/channel-worker-manager-YSFB4HUU.js +418 -0
- package/chunks/channel-worker-supervisor-QCKIOYZD.js +20 -0
- package/chunks/chunk-2FMQQXMM.js +38 -0
- package/chunks/{chunk-CSOTA7KX.js → chunk-2FO5GFKZ.js} +2 -2
- package/chunks/{chunk-QQDPRDVW.js → chunk-2HRYPZT5.js} +1 -1
- package/chunks/{chunk-UXEW557D.js → chunk-2J4ZWME4.js} +3 -3
- package/chunks/{chunk-P3XEEI3X.js → chunk-2TQXXQMI.js} +36 -7
- package/chunks/{chunk-GSHQYYIV.js → chunk-2UW5RFRP.js} +6 -26
- package/chunks/{chunk-6IAWPIQ6.js → chunk-2W3OOD4W.js} +5 -3
- package/chunks/{chunk-2PSWA5ID.js → chunk-2X56YUSZ.js} +1 -1
- package/chunks/{chunk-NQKLOAVJ.js → chunk-2Z5UX4KI.js} +2 -2
- package/chunks/chunk-3MX7D6QN.js +516 -0
- package/chunks/{chunk-EEN2TKST.js → chunk-43RXPU5Q.js} +2 -2
- package/chunks/{chunk-55ZMG67I.js → chunk-4LD5APFL.js} +4 -4
- package/chunks/{chunk-WF4FAXDL.js → chunk-4QZ3XHWE.js} +62 -62
- package/chunks/chunk-4QZADKJS.js +34 -0
- package/chunks/{chunk-UIDXQNMV.js → chunk-4WGODMVR.js} +1 -1
- package/chunks/{chunk-E6US47KI.js → chunk-4WKAA4KS.js} +1 -1
- package/chunks/{chunk-TI3BWD4D.js → chunk-54UZQ57N.js} +39 -19
- package/chunks/{chunk-3GONHQOA.js → chunk-57OAHC2Y.js} +1 -1
- package/chunks/{chunk-JZSY4WP3.js → chunk-5DTTYNB7.js} +1 -1
- package/chunks/chunk-5HBA2Z7V.js +28 -0
- package/chunks/{chunk-SKAZWEV5.js → chunk-5M6IDOMF.js} +35 -98
- package/chunks/{chunk-A4BMJM77.js → chunk-5O2XNYP6.js} +1 -0
- package/chunks/{chunk-FQNC62B2.js → chunk-5VHASII5.js} +11 -11
- package/chunks/{chunk-GO6LNQXT.js → chunk-6AM5D4ME.js} +1 -1
- package/chunks/{chunk-BJBWRCSK.js → chunk-6H67XLET.js} +1 -1
- package/chunks/{chunk-O7FIRWJV.js → chunk-6JPDCEDB.js} +3 -3
- package/chunks/{chunk-3UP777ZS.js → chunk-6RLROE6G.js} +1 -1
- package/chunks/{chunk-GN3F6PIU.js → chunk-6WPJRYLZ.js} +5 -5
- package/chunks/{chunk-ZERZSAZL.js → chunk-75DOP5OR.js} +2 -2
- package/chunks/chunk-7DKX4A5Z.js +91 -0
- package/chunks/{chunk-M7S4GK5Z.js → chunk-7IDCQ35Q.js} +147 -30
- package/chunks/{chunk-22MNNNXE.js → chunk-7NFB7LBQ.js} +11 -3
- package/chunks/{chunk-ZL5HN75Q.js → chunk-7R4DPSRA.js} +37 -3
- package/chunks/{chunk-VZ7HY3YU.js → chunk-7STBROEV.js} +3 -3
- package/chunks/{chunk-SAIILPQ3.js → chunk-7X7TEQNX.js} +5 -3
- package/chunks/{chunk-RMLZAUWH.js → chunk-A3U2TQUL.js} +10 -12
- package/chunks/{chunk-IRJWHZWU.js → chunk-A4UCTEB5.js} +2 -2
- package/chunks/{chunk-BOBRZZKD.js → chunk-A7W4GQRW.js} +5 -5
- package/chunks/{chunk-O33TBKDC.js → chunk-AAUZBE5Y.js} +155 -19
- package/chunks/{chunk-SQ2UATK7.js → chunk-ADLYZJYZ.js} +11 -11
- package/chunks/{chunk-QYUE6W3T.js → chunk-ALGGS7UH.js} +2 -39
- package/chunks/{chunk-WMPVYQ4P.js → chunk-ARKANCNX.js} +3 -2
- package/chunks/{chunk-IBY3Q3HG.js → chunk-AU4R7ACM.js} +967 -181
- package/chunks/{channel-worker-supervisor-NZQUOEFJ.js → chunk-B7MFJBKE.js} +190 -34
- package/chunks/{chunk-HW4N45EE.js → chunk-B7NUUEXC.js} +38 -11
- package/chunks/{chunk-3B26UVCX.js → chunk-BLPMRN4B.js} +39 -10
- package/chunks/{chunk-ZN5T4BHI.js → chunk-BPZALHVR.js} +2 -2
- package/chunks/{chunk-BVYU3ZTI.js → chunk-BXT6PQ73.js} +1020 -942
- package/chunks/{chunk-OFEVLU4C.js → chunk-CARU2RR2.js} +1 -1
- package/chunks/{chunk-HDOH3QYE.js → chunk-CEGKYZDB.js} +2 -2
- package/chunks/{chunk-GLE2YUPU.js → chunk-COH43MRJ.js} +1 -1
- package/chunks/{chunk-E5A7LHNN.js → chunk-CPBF7KYF.js} +1 -1
- package/chunks/{chunk-UWCTAVOD.js → chunk-CR3C7WXL.js} +1 -1
- package/chunks/{chunk-IWKSG2AR.js → chunk-CXAOG665.js} +1 -1
- package/chunks/{chunk-64WXLC72.js → chunk-DJ2GSRLV.js} +1 -1
- package/chunks/{chunk-QN5NZ3UQ.js → chunk-DMTGGOSA.js} +2 -2
- package/chunks/{chunk-27BHARDE.js → chunk-DMVA55W5.js} +2124 -30
- package/chunks/{chunk-JPALE3PA.js → chunk-DQCCLVQK.js} +6 -6
- package/chunks/{chunk-X2XNQQJI.js → chunk-E2UVIFI2.js} +186 -399
- package/chunks/{chunk-KHDZHZMH.js → chunk-E5Z2AVNV.js} +1 -1
- package/chunks/{chunk-E7EMCSUG.js → chunk-EAHMOXLL.js} +64 -42
- package/chunks/{chunk-VGGOPZUL.js → chunk-EJ3KGYWL.js} +36 -8
- package/chunks/{chunk-DIWVZ3VM.js → chunk-ERTHKDKN.js} +5 -5
- package/chunks/{chunk-UX6OTBVD.js → chunk-EX5OULP3.js} +2 -2
- package/chunks/{chunk-5MBO76SL.js → chunk-F333XEAU.js} +3 -3
- package/chunks/{chunk-OACJLMLD.js → chunk-FBU7WRZI.js} +18 -2
- package/chunks/{chunk-5IFG2VC4.js → chunk-FCMNKVXT.js} +1 -1
- package/chunks/chunk-FEOBU3FA.js +253 -0
- package/chunks/{chunk-LAXNEOBG.js → chunk-FI2WSOWX.js} +19 -3
- package/chunks/{chunk-E7KNELIZ.js → chunk-FSA7ERJ2.js} +2 -2
- package/chunks/{chunk-44NROQYV.js → chunk-GDQQRV43.js} +4 -4
- package/chunks/{chunk-ZE6E424U.js → chunk-GN5HJJBO.js} +2 -2
- package/chunks/{chunk-O2YT7VB4.js → chunk-GOXKNSDZ.js} +2 -65
- package/chunks/{chunk-V35MD7KN.js → chunk-GRC5HGPI.js} +1 -1
- package/chunks/{chunk-ON3JA3RU.js → chunk-GTENQOEJ.js} +7 -9
- package/chunks/{chunk-S6RLAIUR.js → chunk-GYTKQYDB.js} +1 -1
- package/chunks/{chunk-A2X3KLC3.js → chunk-GZXCZHCT.js} +3 -3
- package/chunks/{chunk-K5PGHDBN.js → chunk-H6XPXXMH.js} +1 -1
- package/chunks/{chunk-PMFULMLX.js → chunk-HCJTURZD.js} +2262 -384
- package/chunks/chunk-HGG2RRXD.js +445 -0
- package/chunks/{chunk-II73RK2S.js → chunk-HGT6JR3U.js} +20 -1
- package/chunks/{chunk-DICGXFAI.js → chunk-HJCVNZLQ.js} +11 -9
- package/chunks/{chunk-3GNOQZDC.js → chunk-HJUDVNCO.js} +17663 -14877
- package/chunks/{chunk-AYNCLGTW.js → chunk-HRUGYOJN.js} +8 -51
- package/chunks/{chunk-KYMBIKIW.js → chunk-HTHUX2T5.js} +1 -1
- package/chunks/{chunk-LQ7TMOCE.js → chunk-HTO4JFDZ.js} +1 -1
- package/chunks/{chunk-BUZ6HGL4.js → chunk-HWTL6MPR.js} +357 -59
- package/chunks/chunk-IDS7MSUP.js +74 -0
- package/chunks/{chunk-4CGOWHKM.js → chunk-J4MOFEOK.js} +2 -2
- package/chunks/chunk-JCD4FM7K.js +296 -0
- package/chunks/{chunk-22IFUCVR.js → chunk-JESCGQM3.js} +1 -1
- package/chunks/{chunk-JZHIQJOE.js → chunk-JGO6RFI2.js} +6 -6
- package/chunks/{chunk-X3YGAX7V.js → chunk-JGUT3LWZ.js} +9 -2
- package/chunks/{chunk-JRQ7EVGE.js → chunk-JW7OOJC6.js} +2 -2
- package/chunks/{chunk-OQXDAMKY.js → chunk-KF5MMGG4.js} +8 -8
- package/chunks/{chunk-42XGFRJS.js → chunk-KKRTYONM.js} +8 -6
- package/chunks/{chunk-GGQQA4JW.js → chunk-KSO42X3Z.js} +15 -2
- package/chunks/{chunk-MRO43B25.js → chunk-KW7NOTN6.js} +1 -1
- package/chunks/{chunk-54YZOP7P.js → chunk-L573Z2QZ.js} +11894 -5190
- package/chunks/{chunk-TLWA7TPR.js → chunk-L5U3ELZX.js} +2 -2
- package/chunks/{chunk-YHKAK72D.js → chunk-LHEOLOXU.js} +251 -63
- package/chunks/{chunk-H6EMK6QK.js → chunk-LQIDP6K5.js} +11913 -5207
- package/chunks/{chunk-I6WGFJJ3.js → chunk-LSEJG6AM.js} +73 -19
- package/chunks/{chunk-MZ7BABX3.js → chunk-LXLTBIDP.js} +44 -7
- package/chunks/{chunk-HLKG6CZN.js → chunk-M5GB774H.js} +1 -1
- package/chunks/{chunk-JCH35JB3.js → chunk-MKK4XGLF.js} +2 -2
- package/chunks/{chunk-YKZIAK7C.js → chunk-MO7O5722.js} +424 -40
- package/chunks/chunk-MR3PXB6E.js +48 -0
- package/chunks/{chunk-A5JOJ77H.js → chunk-MRYWYBQ4.js} +38 -15
- package/chunks/{chunk-332PWN27.js → chunk-NHDKSOEY.js} +1 -1
- package/chunks/{chunk-TSPB77TG.js → chunk-NRAFASYX.js} +1016 -2967
- package/chunks/{chunk-BMFMJINR.js → chunk-OA4JVLSA.js} +28 -2
- package/chunks/{chunk-P7M4GA57.js → chunk-OCPBI7J5.js} +1 -1
- package/chunks/chunk-OJAMDVF5.js +66 -0
- package/chunks/{chunk-IJRL2S3D.js → chunk-OKAIGAYW.js} +6 -6
- package/chunks/{chunk-FBET5XZ2.js → chunk-PEJWA3QQ.js} +29 -10
- package/chunks/{chunk-L4AMBVQG.js → chunk-PMKW35AT.js} +54 -22
- package/chunks/{chunk-AGHU5VRK.js → chunk-PZL23GTM.js} +1 -1
- package/chunks/{chunk-PC4ZXHU2.js → chunk-Q6TUALBE.js} +3 -3
- package/chunks/{chunk-FXGLL2HL.js → chunk-Q7YDKOGY.js} +1 -1
- package/chunks/{chunk-XOHNA2S6.js → chunk-QAW7PIHT.js} +3 -3
- package/chunks/{chunk-XYR5RUA5.js → chunk-QE2WJ7FS.js} +10 -3
- package/chunks/{chunk-LGFAYIKX.js → chunk-QHTIBUWB.js} +1 -1
- package/chunks/{chunk-P3MYAZBY.js → chunk-QMGX2KO2.js} +1 -1
- package/chunks/{chunk-Z3I3QJK3.js → chunk-QRUIS3DN.js} +1 -1
- package/chunks/{chunk-ITOGNELQ.js → chunk-QSWDGAIC.js} +2 -2
- package/chunks/{chunk-PUVCVELY.js → chunk-QUJOI7SR.js} +70 -8
- package/chunks/{chunk-GHOTR7HL.js → chunk-QUPXZXLV.js} +1 -1
- package/chunks/{chunk-6VK6FIMQ.js → chunk-R5D5JOAM.js} +415 -379
- package/chunks/{chunk-CD6USWHZ.js → chunk-R5XLNDE2.js} +2 -2
- package/chunks/{chunk-W7EEOR7D.js → chunk-RCEEXIOO.js} +8 -6
- package/chunks/{chunk-TLRYABYP.js → chunk-RKUWKYED.js} +1 -1
- package/chunks/{chunk-RNBYOUGV.js → chunk-RXFQM6FQ.js} +1 -1
- package/chunks/{chunk-QXPI7FEC.js → chunk-S3TWXGKA.js} +4 -4
- package/chunks/{chunk-2EM7ECZD.js → chunk-SMSFWUFY.js} +309 -80
- package/chunks/{chunk-FV7425LN.js → chunk-SPJQO2CE.js} +18 -5
- package/chunks/{chunk-MLZQVCF3.js → chunk-SZRN75WQ.js} +1 -1
- package/chunks/{chunk-7RYW5LQV.js → chunk-T2NCM2ET.js} +2 -2
- package/chunks/{chunk-Z2Z3GUXZ.js → chunk-TB6UDU4T.js} +1 -1
- package/chunks/{chunk-EXPMGZZV.js → chunk-TEGEBB2I.js} +1 -1
- package/chunks/{chunk-MX2YXRER.js → chunk-TJTKSVIV.js} +1 -1
- package/chunks/{chunk-CCGOFQV3.js → chunk-TQBODCFD.js} +8 -7
- package/chunks/chunk-TV6K2FB4.js +224 -0
- package/chunks/{chunk-H6BD2ELD.js → chunk-TYAMGABM.js} +2 -2
- package/chunks/{chunk-QYNNN7BK.js → chunk-TYLOFD2U.js} +2 -2
- package/chunks/{chunk-VRDMOSQS.js → chunk-UDXSXLER.js} +2 -2
- package/chunks/{chunk-DHOLNQLC.js → chunk-UFDX27PL.js} +60 -40
- package/chunks/{chunk-WHJQ3JUS.js → chunk-UGN6CTNZ.js} +10 -10
- package/chunks/{chunk-CHB5SLZZ.js → chunk-UNLRKDMW.js} +4 -4
- package/chunks/{chunk-CJCKDMWO.js → chunk-URK542T4.js} +1 -1
- package/chunks/{chunk-AHNXYU4O.js → chunk-V6WHFGST.js} +1 -1
- package/chunks/{chunk-ZTQ26VBE.js → chunk-VDDGDZMA.js} +1 -1
- package/chunks/{chunk-BOBOCR5H.js → chunk-VEETSXYX.js} +4 -4
- package/chunks/{chunk-OMX7CUOE.js → chunk-VGC4I5JJ.js} +1 -1
- package/chunks/{chunk-43W4YSO6.js → chunk-VJNBRRZG.js} +2 -2
- package/chunks/{chunk-WQTHLB5T.js → chunk-VKS7EIOA.js} +7 -7
- package/chunks/{chunk-J7RCB6N5.js → chunk-VVDNGM6X.js} +2 -2
- package/chunks/{chunk-ER3BKOLB.js → chunk-W7FJ3N32.js} +1 -1
- package/chunks/{chunk-3HX5LZ6R.js → chunk-WBL3FJEU.js} +2 -2
- package/chunks/{chunk-7LTB54MK.js → chunk-WIEO4CWB.js} +2 -2
- package/chunks/chunk-WMC6HH7Y.js +493 -0
- package/chunks/{chunk-FEJ2FZ3U.js → chunk-WTNZ6DXB.js} +2 -2
- package/chunks/{chunk-L46XKEGM.js → chunk-X6ODDESK.js} +5 -8
- package/chunks/chunk-XDPYM5DE.js +71 -0
- package/chunks/{chunk-3RW4AZJV.js → chunk-XFSFZCM5.js} +1 -1
- package/chunks/{chunk-2MPVVENX.js → chunk-XU457J4V.js} +2 -2
- package/chunks/{chunk-SIUQ3YYX.js → chunk-XVNQMZ2I.js} +1 -1
- package/chunks/{chunk-SYCJMSIJ.js → chunk-XWRJCPHC.js} +1 -1
- package/chunks/{chunk-SUAKHHK7.js → chunk-Y66VOBRP.js} +3 -3
- package/chunks/chunk-YAF2MQZ6.js +19 -0
- package/chunks/chunk-YHN5SUIJ.js +219 -0
- package/chunks/{chunk-BR4QREVK.js → chunk-YQ3U5MUC.js} +1 -1
- package/chunks/{chunk-Y6Z2O3WR.js → chunk-YUZI3WAC.js} +1 -1
- package/chunks/{chunk-YDOXPLYL.js → chunk-ZA74ZTQU.js} +4564 -1326
- package/chunks/{chunk-DCZWC4PF.js → chunk-ZGE3222J.js} +5 -5
- package/chunks/chunk-ZP4ZZJJG.js +448 -0
- package/chunks/{chunk-AHNJGLFV.js → chunk-ZPAHLFZU.js} +3 -3
- package/chunks/cli-entry-path-4VJ3Y6T2.js +10 -0
- package/chunks/{computer-use-WMY3UOBX.js → computer-use-SHA6C3ES.js} +50 -50
- package/chunks/config-utils-XSHOMMK4.js +24 -0
- package/chunks/contextCommand-HOAF5LUC.js +68 -0
- package/chunks/create-sub-session-EAB2U5XW.js +207 -0
- package/chunks/create-sub-session-WJOWBFEP.js +375 -0
- package/chunks/{cron-create-JSXEIPVS.js → cron-create-GMVKSXZT.js} +9 -9
- package/chunks/{cron-delete-OIFVEPGU.js → cron-delete-VDZKUAVK.js} +7 -7
- package/chunks/{cron-list-D4ZX4IOE.js → cron-list-UG7C7RAR.js} +14 -10
- package/chunks/{daemon-WPAENEBW.js → daemon-MI5HF26N.js} +1074 -61
- package/chunks/daemon-status-provider-KFL3LRTP.js +76 -0
- package/chunks/{de-ZFPPQVSF.js → de-2VCDRDGU.js} +68 -6
- package/chunks/{devtools-FM6GJPYG.js → devtools-4QFYJT6U.js} +3 -3
- package/chunks/{dist-O2IRTRPP.js → dist-IEXZC5ZR.js} +10 -10
- package/chunks/{dist-SCFLLYMK.js → dist-PQK4GCKA.js} +12 -11
- package/chunks/{dist-63IS3ZMI.js → dist-PZV5RKM6.js} +4 -4
- package/chunks/{dist-E6UANPIM.js → dist-U75JK3PF.js} +1011 -197
- package/chunks/{dist-H2AEJCZ5.js → dist-VXO7QBON.js} +5 -5
- package/chunks/{dist-Z34QPAQW.js → dist-WH4TZSH3.js} +529 -79
- package/chunks/{dist-BEXOZXLQ.js → dist-X2ABIW5R.js} +36 -15
- package/chunks/earlyInputCapture-PKAQJMIP.js +69 -0
- package/chunks/{edit-OTYWAXGP.js → edit-6TXLYHJ2.js} +51 -51
- package/chunks/{en-N4EPZAQI.js → en-NJCA3TAA.js} +78 -6
- package/chunks/{enter-worktree-HNYBDVCX.js → enter-worktree-ZTILGPJX.js} +50 -50
- package/chunks/{enterPlanMode-CJCNGT63.js → enterPlanMode-62HRAMCH.js} +66 -51
- package/chunks/environment-OYA3ZVEL.js +89 -0
- package/chunks/errors-BFXOUUUR.js +75 -0
- package/chunks/{exit-worktree-DUUKOE3X.js → exit-worktree-HHJ5M6MV.js} +50 -50
- package/chunks/{exitPlanMode-AMMSXOAJ.js → exitPlanMode-M3SWHCL7.js} +51 -51
- package/chunks/{fast-path-OV6NIP3K.js → fast-path-3VMPLM74.js} +8 -8
- package/chunks/{fast-path-settings-7RIFI3PK.js → fast-path-settings-DAOVFADC.js} +8 -6
- package/chunks/{fileFromPath-IBEHA3CO.js → fileFromPath-UDALK7FM.js} +3 -3
- package/chunks/{fr-5WWMY53B.js → fr-5F3E7WKD.js} +68 -6
- package/chunks/{gemini-NQ3WWBKB.js → gemini-ZCGVYOJL.js} +181 -110
- package/chunks/{geminiContentGenerator-CDUNEJQE.js → geminiContentGenerator-XBYBAIDQ.js} +9 -9
- package/chunks/{getMachineId-bsd-4CASPIU4.js → getMachineId-bsd-EXV7SWPA.js} +3 -3
- package/chunks/{getMachineId-darwin-HPQPEMZR.js → getMachineId-darwin-IUOMFXU3.js} +3 -3
- package/chunks/{getMachineId-linux-AUARKYHL.js → getMachineId-linux-K7XJYFHL.js} +2 -2
- package/chunks/{getMachineId-unsupported-S32ZDA2T.js → getMachineId-unsupported-EA5FDRJ6.js} +2 -2
- package/chunks/{getMachineId-win-4EFLHYIJ.js → getMachineId-win-KBA6RAI5.js} +3 -3
- package/chunks/{glob-ROV3YS4R.js → glob-UOEYQNHN.js} +97 -75
- package/chunks/{grep-NRITKDCD.js → grep-OQM3CH7W.js} +50 -50
- package/chunks/handleAutoUpdate-ORBYVJHZ.js +69 -0
- package/chunks/i18n-NBP574CO.js +84 -0
- package/chunks/initializer-CWD7TCCD.js +73 -0
- package/chunks/installationInfo-CHPDUPWC.js +65 -0
- package/chunks/{ja-CUHZOLEJ.js → ja-66J6B53L.js} +69 -7
- package/chunks/{keychain-token-storage-UJV5XP4Z.js → keychain-token-storage-AU22CQI6.js} +4 -4
- package/chunks/{kittyProtocolDetector-OJRCJLIU.js → kittyProtocolDetector-HCNNKJTI.js} +2 -2
- package/chunks/list-ZTTO362W.js +76 -0
- package/chunks/loadedSettingsAdapter-V4ZTRKUO.js +70 -0
- package/chunks/{loop-wakeup-OHWEMFP5.js → loop-wakeup-C4EXLYVA.js} +10 -10
- package/chunks/{lowlight-FYAAUU5J.js → lowlight-J2OZNPCA.js} +1 -1
- package/chunks/{ls-7MFU2QYW.js → ls-2OUTG3IZ.js} +6 -6
- package/chunks/{lsp-XHDTPTGV.js → lsp-4VMNXZW6.js} +4 -4
- package/chunks/mcp-UH5PSGIK.js +70 -0
- package/chunks/{monitor-FL2ZO33W.js → monitor-F4NYEU5J.js} +50 -50
- package/chunks/{multipart-parser-AJ4WASWR.js → multipart-parser-AWZKXSQN.js} +3 -3
- package/chunks/nonInteractiveCli-7IHPXVNH.js +129 -0
- package/chunks/{notebook-edit-6PULTHL5.js → notebook-edit-HPDXNGQH.js} +51 -51
- package/chunks/openaiContentGenerator-OBDNNOE2.js +56 -0
- package/chunks/pidfile-AUMGYIOK.js +73 -0
- package/chunks/{pt-OLYN42DS.js → pt-TFZO5Y6T.js} +68 -6
- package/chunks/{qwenContentGenerator-NX2E7PB7.js → qwenContentGenerator-LH2U2YRG.js} +52 -52
- package/chunks/{qwenOAuth2-ZV7YLMND.js → qwenOAuth2-LEYOU7RU.js} +9 -9
- package/chunks/read-file-HUUFBLEO.js +30 -0
- package/chunks/{read-mcp-resource-CNQHBSJM.js → read-mcp-resource-JUFHCKSC.js} +5 -5
- package/chunks/{read-package-up-ER5OJUGP.js → read-package-up-TCM6I7S2.js} +2 -2
- package/chunks/{record-artifact-SQ2OKRJW.js → record-artifact-DWSMBXYY.js} +4 -4
- package/chunks/ripGrep-C7X33LU6.js +60 -0
- package/chunks/{ru-LDM3GWVQ.js → ru-4L4LFHIT.js} +68 -6
- package/chunks/{run-qwen-serve-JK2DL7PI.js → run-qwen-serve-DGW5AFDJ.js} +1465 -283
- package/chunks/runtime-NYUSLBEZ.js +99 -0
- package/chunks/{scheduler-O4IGVZKK.js → scheduler-5VKDXULT.js} +51 -51
- package/chunks/{send-message-76GT3JMC.js → send-message-JZ752NWD.js} +7 -7
- package/chunks/serve-UTAHSFXH.js +76 -0
- package/chunks/{server-SGNPKX2O.js → server-VUHSYSUV.js} +9088 -1817
- package/chunks/{session-5BNM25EZ.js → session-QSGSSNFB.js} +340 -135
- package/chunks/settings-X24RYEED.js +120 -0
- package/chunks/shell-4MSI2R4O.js +70 -0
- package/chunks/{skill-I22QK5XS.js → skill-JR3GTNVZ.js} +24 -24
- package/chunks/spawnChannel-RFAAO44J.js +70 -0
- package/chunks/{src-7XL4G4DC.js → src-VLI3TKEY.js} +4 -4
- package/chunks/{src-BBQSM46M.js → src-YDBNMIIM.js} +183 -53
- package/chunks/standalone-update-4CY4HKGC.js +79 -0
- package/chunks/{startInteractiveUI-JII36UGN.js → startInteractiveUI-FOD7MMVB.js} +2368 -1923
- package/chunks/stdioHelpers-UMY72QRP.js +14 -0
- package/chunks/{syntheticOutput-F7LSDCAR.js → syntheticOutput-LELY6HAI.js} +5 -5
- package/chunks/task-create-QY3BIMNM.js +23 -0
- package/chunks/{task-list-TGGJX7EA.js → task-list-D2U47WBJ.js} +11 -10
- package/chunks/{task-stop-W7NGTMYQ.js → task-stop-N6E5IYLH.js} +4 -4
- package/chunks/{task-update-XYWMHJGJ.js → task-update-HJCQILLB.js} +13 -12
- package/chunks/{team-create-A7NUU5AQ.js → team-create-ZLQSU2E3.js} +50 -50
- package/chunks/{team-delete-UHQQ2QQX.js → team-delete-RCWYDRF4.js} +11 -10
- package/chunks/{team-plan-approval-TOQRJ2KI.js → team-plan-approval-YUAYBYLS.js} +50 -50
- package/chunks/theme-manager-SKNEBMOG.js +64 -0
- package/chunks/{todoWrite-C4ZAMSR3.js → todoWrite-FKCGATPJ.js} +6 -6
- package/chunks/{tool-search-SZ4EFS2T.js → tool-search-V5J3TGWY.js} +24 -24
- package/chunks/total-session-admission-2WIH72I3.js +72 -0
- package/chunks/tree-sitter-YJVE2ZUK.js +2980 -0
- package/chunks/trustedFolders-JF265CRI.js +80 -0
- package/chunks/types-ML3TRJQ5.js +16 -0
- package/chunks/updateCheck-AMXNPBP6.js +69 -0
- package/chunks/updateEventEmitter-5MKC7BLQ.js +10 -0
- package/chunks/validateNonInterActiveAuth-6YHQO7PD.js +194 -0
- package/chunks/{version-NO2XSNNP.js → version-UEJJK7PE.js} +3 -3
- package/chunks/{web-fetch-UYGMOHW4.js → web-fetch-QGODBB6S.js} +8 -8
- package/chunks/{workflow-WO2H6NRI.js → workflow-5YFA5V4M.js} +51 -51
- package/chunks/workspace-providers-status-LMBLBNR7.js +73 -0
- package/chunks/workspace-registration-store-VV2S4MSV.js +27 -0
- package/chunks/workspace-registry-BVMPYRLN.js +76 -0
- package/chunks/workspace-service-PUYIZ3J4.js +86 -0
- package/chunks/workspace-skills-status-CJ7J2KPP.js +72 -0
- package/chunks/{write-file-PX2Z7IBE.js → write-file-OV77M6X7.js} +51 -51
- package/chunks/{yargs-B64EYTCF.js → yargs-6H2AUULL.js} +2 -2
- package/chunks/{zh-J4F6QWQV.js → zh-FSAH32JB.js} +79 -7
- package/chunks/{zh-TW-XVDWOLSW.js → zh-TW-D2YIE2LO.js} +79 -7
- package/cli.js +18 -17
- package/locales/ca.js +91 -7
- package/locales/de.js +90 -7
- package/locales/en.js +109 -7
- package/locales/fr.js +91 -7
- package/locales/ja.js +93 -7
- package/locales/pt.js +90 -7
- package/locales/ru.js +90 -7
- package/locales/zh-TW.js +106 -7
- package/locales/zh.js +106 -7
- package/package.json +3 -6
- package/web-shell/assets/{arc-BWyoJpY2.js → arc-CObt_TK9.js} +1 -1
- package/web-shell/assets/{architectureDiagram-3BPJPVTR-BT0mGrz2.js → architectureDiagram-3BPJPVTR-B7aiJ6CK.js} +1 -1
- package/web-shell/assets/{blockDiagram-GPEHLZMM-BLgUWvyf.js → blockDiagram-GPEHLZMM-BSK2CrID.js} +1 -1
- package/web-shell/assets/{c4Diagram-AAUBKEIU-DYx9M7oU.js → c4Diagram-AAUBKEIU-Bhx2p7T_.js} +1 -1
- package/web-shell/assets/channel-BBaLq8MV.js +1 -0
- package/web-shell/assets/{chunk-2J33WTMH-CfMHWbuZ.js → chunk-2J33WTMH-c3Hzv0Ni.js} +1 -1
- package/web-shell/assets/{chunk-4BX2VUAB-B1URoqEV.js → chunk-4BX2VUAB-DNXaK7ZF.js} +1 -1
- package/web-shell/assets/{chunk-55IACEB6-DAIaM4QF.js → chunk-55IACEB6-BXjvDj_i.js} +1 -1
- package/web-shell/assets/{chunk-727SXJPM-B2X34knw.js → chunk-727SXJPM-B6CaKFp-.js} +1 -1
- package/web-shell/assets/{chunk-AQP2D5EJ-Cif-zLTQ.js → chunk-AQP2D5EJ-tij6Mz9B.js} +1 -1
- package/web-shell/assets/{chunk-FMBD7UC4-H18WCXdY.js → chunk-FMBD7UC4-DdlsGk2h.js} +1 -1
- package/web-shell/assets/{chunk-ND2GUHAM-YFtetHi0.js → chunk-ND2GUHAM-CdgGZSHE.js} +1 -1
- package/web-shell/assets/{chunk-QZHKN3VN-By2hiH3r.js → chunk-QZHKN3VN-D8oJn1C0.js} +1 -1
- package/web-shell/assets/classDiagram-4FO5ZUOK-B2ZK6sWE.js +1 -0
- package/web-shell/assets/classDiagram-v2-Q7XG4LA2-B2ZK6sWE.js +1 -0
- package/web-shell/assets/{cose-bilkent-S5V4N54A-VrcnnLZu.js → cose-bilkent-S5V4N54A-BUVmNp1Y.js} +1 -1
- package/web-shell/assets/{dagre-BM42HDAG-CGo-IaVV.js → dagre-BM42HDAG-CO4q310O.js} +1 -1
- package/web-shell/assets/{diagram-2AECGRRQ-Dl4QjzIB.js → diagram-2AECGRRQ-C6gFvWkg.js} +1 -1
- package/web-shell/assets/{diagram-5GNKFQAL-qMLmsO0u.js → diagram-5GNKFQAL-NG2I_c0h.js} +1 -1
- package/web-shell/assets/{diagram-KO2AKTUF-CHlmuxPI.js → diagram-KO2AKTUF-DFKMS6I9.js} +1 -1
- package/web-shell/assets/{diagram-LMA3HP47-BUwU3GwT.js → diagram-LMA3HP47-BiElfsfP.js} +1 -1
- package/web-shell/assets/{diagram-OG6HWLK6-HEhwPDUM.js → diagram-OG6HWLK6-zLWDYUqu.js} +1 -1
- package/web-shell/assets/{erDiagram-TEJ5UH35-1n0llBmM.js → erDiagram-TEJ5UH35-DoVaazhK.js} +1 -1
- package/web-shell/assets/{flowDiagram-I6XJVG4X-BN74uMnt.js → flowDiagram-I6XJVG4X-0RLz-mXg.js} +1 -1
- package/web-shell/assets/{ganttDiagram-6RSMTGT7-BoOq78i3.js → ganttDiagram-6RSMTGT7-C-eX1A39.js} +3 -3
- package/web-shell/assets/{gitGraphDiagram-PVQCEYII-CN8b00c8.js → gitGraphDiagram-PVQCEYII-BauQu1Bu.js} +1 -1
- package/web-shell/assets/index-3uAT3ehO.css +5 -0
- package/web-shell/assets/index-CVHa2QH7.js +3 -0
- package/web-shell/assets/index-D1KigWat.js +1128 -0
- package/web-shell/assets/{infoDiagram-5YYISTIA-CCurnIna.js → infoDiagram-5YYISTIA-C800cFvZ.js} +1 -1
- package/web-shell/assets/{ishikawaDiagram-YF4QCWOH-AwuRqdUP.js → ishikawaDiagram-YF4QCWOH-CFK1Jk6e.js} +1 -1
- package/web-shell/assets/{journeyDiagram-JHISSGLW-B7IICk-E.js → journeyDiagram-JHISSGLW-Cjx0WAC4.js} +1 -1
- package/web-shell/assets/{kanban-definition-UN3LZRKU-BYF0OKhx.js → kanban-definition-UN3LZRKU-DWMDHbiv.js} +1 -1
- package/web-shell/assets/{linear-GWWguGfD.js → linear-vpW4FrZs.js} +1 -1
- package/web-shell/assets/{mermaid.core-JqtkxvCl.js → mermaid.core-DsHRUnDd.js} +5 -5
- package/web-shell/assets/{mindmap-definition-RKZ34NQL-B0MmHsZ6.js → mindmap-definition-RKZ34NQL-D3Eb1cAK.js} +1 -1
- package/web-shell/assets/{pieDiagram-4H26LBE5-DhImUZLQ.js → pieDiagram-4H26LBE5-Cr2-6sWt.js} +1 -1
- package/web-shell/assets/{quadrantDiagram-W4KKPZXB-DGeSC3DT.js → quadrantDiagram-W4KKPZXB-Yd7yUgVp.js} +1 -1
- package/web-shell/assets/{requirementDiagram-4Y6WPE33-DkspY6Re.js → requirementDiagram-4Y6WPE33-BE6SprL8.js} +1 -1
- package/web-shell/assets/{sankeyDiagram-5OEKKPKP-CaOofS2x.js → sankeyDiagram-5OEKKPKP-DzkuMK8d.js} +1 -1
- package/web-shell/assets/{sequenceDiagram-3UESZ5HK-DFQ_kFnN.js → sequenceDiagram-3UESZ5HK-wD9A3Dhi.js} +1 -1
- package/web-shell/assets/{stateDiagram-AJRCARHV-D3snw0Tz.js → stateDiagram-AJRCARHV-BOp5q1Nm.js} +1 -1
- package/web-shell/assets/stateDiagram-v2-BHNVJYJU-BDsWUmPG.js +1 -0
- package/web-shell/assets/{timeline-definition-PNZ67QCA-l4zH8nLM.js → timeline-definition-PNZ67QCA-Btc08sUA.js} +1 -1
- package/web-shell/assets/{vennDiagram-CIIHVFJN-BnRlA98M.js → vennDiagram-CIIHVFJN-Bv6upsOA.js} +1 -1
- package/web-shell/assets/{wardley-L42UT6IY-BS0xUdmd.js → wardley-L42UT6IY-D-XamnCK.js} +1 -1
- package/web-shell/assets/{wardleyDiagram-YWT4CUSO-DKGBUWyQ.js → wardleyDiagram-YWT4CUSO-BbP_5Wfp.js} +1 -1
- package/web-shell/assets/{xychartDiagram-2RQKCTM6-DVb2HuOB.js → xychartDiagram-2RQKCTM6-CO5Hh28b.js} +1 -1
- package/web-shell/index.html +16 -4
- package/chunks/MaxSizedBox-5DT2SD7B.js +0 -71
- package/chunks/agent-H2DGKS77.js +0 -68
- package/chunks/agent-headless-KFMFIT6J.js +0 -62
- package/chunks/bridge-NA3ZQQKJ.js +0 -78
- package/chunks/chunk-5P5XGNYH.js +0 -93
- package/chunks/chunk-6SHO7WFF.js +0 -79
- package/chunks/chunk-IFEWJXFS.js +0 -37
- package/chunks/chunk-XNXVCNAT.js +0 -88
- package/chunks/contextCommand-LXFZRP4L.js +0 -68
- package/chunks/daemon-status-provider-GJRXY5EN.js +0 -76
- package/chunks/earlyInputCapture-KNOYZ4MJ.js +0 -69
- package/chunks/environment-ATFTYMLZ.js +0 -88
- package/chunks/errors-A54HAADA.js +0 -75
- package/chunks/handleAutoUpdate-CEZJAARW.js +0 -65
- package/chunks/initializer-BRFJKRRX.js +0 -71
- package/chunks/list-KVHF7O4X.js +0 -73
- package/chunks/loadedSettingsAdapter-LL5PSFKX.js +0 -68
- package/chunks/mcp-3GCXRTOQ.js +0 -68
- package/chunks/nonInteractiveCli-PTQI3DV6.js +0 -119
- package/chunks/openaiContentGenerator-YE42PKCY.js +0 -56
- package/chunks/pidfile-NZYSCFRK.js +0 -73
- package/chunks/read-file-33XGP5HF.js +0 -30
- package/chunks/ripGrep-QWZLGTAN.js +0 -60
- package/chunks/serve-JTMIGANH.js +0 -74
- package/chunks/settings-3MI4XXP6.js +0 -118
- package/chunks/shell-YJKBE67T.js +0 -68
- package/chunks/spawnChannel-UNV6W4DU.js +0 -70
- package/chunks/task-create-WCHTB435.js +0 -22
- package/chunks/theme-manager-WZ4RRMMB.js +0 -64
- package/chunks/total-session-admission-ULXJWXFV.js +0 -72
- package/chunks/trustedFolders-LAZZPZ3D.js +0 -79
- package/chunks/types-QX5C3CHJ.js +0 -12
- package/chunks/updateCheck-KONLM73X.js +0 -65
- package/chunks/validateNonInterActiveAuth-MPKSRCXD.js +0 -184
- package/chunks/workspace-providers-status-3W7E7DFW.js +0 -71
- package/chunks/workspace-registry-SXJZNN3O.js +0 -74
- package/chunks/workspace-service-CZRQZ4Q4.js +0 -80
- package/chunks/workspace-skills-status-K7WFREPD.js +0 -70
- package/node_modules/@qwen-code/audio-capture/dist/index.d.ts +0 -35
- package/node_modules/@qwen-code/audio-capture/dist/index.js +0 -48
- package/node_modules/@qwen-code/audio-capture/dist/index.js.map +0 -1
- package/node_modules/@qwen-code/audio-capture/dist/platform.d.ts +0 -7
- package/node_modules/@qwen-code/audio-capture/dist/platform.js +0 -18
- package/node_modules/@qwen-code/audio-capture/dist/platform.js.map +0 -1
- package/node_modules/@qwen-code/audio-capture/node_modules/node-gyp-build/LICENSE +0 -21
- package/node_modules/@qwen-code/audio-capture/node_modules/node-gyp-build/README.md +0 -58
- package/node_modules/@qwen-code/audio-capture/node_modules/node-gyp-build/SECURITY.md +0 -5
- package/node_modules/@qwen-code/audio-capture/node_modules/node-gyp-build/bin.js +0 -84
- package/node_modules/@qwen-code/audio-capture/node_modules/node-gyp-build/build-test.js +0 -19
- package/node_modules/@qwen-code/audio-capture/node_modules/node-gyp-build/index.js +0 -6
- package/node_modules/@qwen-code/audio-capture/node_modules/node-gyp-build/node-gyp-build.js +0 -207
- package/node_modules/@qwen-code/audio-capture/node_modules/node-gyp-build/optional.js +0 -7
- package/node_modules/@qwen-code/audio-capture/node_modules/node-gyp-build/package.json +0 -43
- package/node_modules/@qwen-code/audio-capture/package.json +0 -30
- package/node_modules/@qwen-code/audio-capture/prebuilds/darwin-arm64/@qwen-code+audio-capture.node +0 -0
- package/node_modules/@qwen-code/audio-capture/prebuilds/darwin-x64/@qwen-code+audio-capture.node +0 -0
- package/node_modules/@qwen-code/audio-capture/prebuilds/linux-arm64/@qwen-code+audio-capture.node +0 -0
- package/node_modules/@qwen-code/audio-capture/prebuilds/linux-x64/@qwen-code+audio-capture.node +0 -0
- package/node_modules/@qwen-code/audio-capture/prebuilds/win32-x64/@qwen-code+audio-capture.node +0 -0
- package/web-shell/assets/channel-wmPUGHtg.js +0 -1
- package/web-shell/assets/classDiagram-4FO5ZUOK-BA2ljEjW.js +0 -1
- package/web-shell/assets/classDiagram-v2-Q7XG4LA2-BA2ljEjW.js +0 -1
- package/web-shell/assets/index-CNwoO1Cn.js +0 -783
- package/web-shell/assets/index-CWQ0z6kZ.css +0 -5
- package/web-shell/assets/stateDiagram-v2-BHNVJYJU-BbSm_9d0.js +0 -1
package/bundled/review/DESIGN.md
CHANGED
|
@@ -2,18 +2,19 @@
|
|
|
2
2
|
|
|
3
3
|
> Architecture decisions, trade-offs, and rejected alternatives for the `/review` skill.
|
|
4
4
|
|
|
5
|
-
## Why
|
|
5
|
+
## Why 12 agents + 1 verify + iterative reverse, not 1 agent?
|
|
6
6
|
|
|
7
7
|
**Considered:**
|
|
8
8
|
|
|
9
9
|
- **1 agent (Copilot approach):** Single agent with tool-calling, reads and reviews in one pass. Cheapest (1 LLM call). But dimensional coverage depends entirely on one prompt's attention — easy to miss performance issues while focused on security.
|
|
10
10
|
- **5 parallel agents (original design):** Each agent focuses on one dimension. Higher coverage through forced diversity of perspective. Limited by combined Correctness+Security and a single undirected pass — recall ceiling left findings on the table that the user only discovered in subsequent /review rounds.
|
|
11
11
|
- **9 parallel agents:** 6 review dimensions (Correctness, Security, Code Quality, Performance, Test Coverage, Undirected) + Build & Test. Undirected runs as 3 personas in parallel.
|
|
12
|
-
- **10 parallel agents
|
|
12
|
+
- **10 parallel agents:** The 9-agent design plus Issue Fidelity & Root-Cause Ownership, which compares linked issue evidence against the PR's claimed fix before accepting a client-side change.
|
|
13
|
+
- **12 parallel agents (current):** The 10-agent design with Correctness split into three procedural walks — 1a line-by-line scan, 1b removed-behavior audit, 1c cross-file tracer — plus up to 2 optional diff-specialized finders (Agent 8) when one domain dominates the diff.
|
|
13
14
|
|
|
14
|
-
**Decision:**
|
|
15
|
+
**Decision:** 12 agents. The marginal cost (12x vs 1x) is acceptable because:
|
|
15
16
|
|
|
16
|
-
1.
|
|
17
|
+
1. All 12 agents are submitted in one response and run concurrently up to the runtime's tool-call cap (default 10, `QWEN_CODE_MAX_TOOL_CONCURRENCY`) — wall time is bounded by roughly two waves at worst, still far below twelve sequential agents
|
|
17
18
|
2. Dimensional focus produces higher recall (fewer missed issues)
|
|
18
19
|
3. Three undirected personas (attacker / 3am-oncall / maintainer) catch cross-dimensional issues that a single undirected agent's prompt-induced bias would miss
|
|
19
20
|
4. Issue Fidelity prevents a common false approval mode: a PR can be internally well-tested while solving only the author's mistaken diagnosis, not the linked issue's original failure
|
|
@@ -29,7 +30,7 @@ Test gaps are a systematic blind spot. Review agents focused on bugs in the new
|
|
|
29
30
|
|
|
30
31
|
### Why a dedicated Issue Fidelity agent
|
|
31
32
|
|
|
32
|
-
Bugfix PRs often carry their own diagnosis in the PR body, but that diagnosis can be wrong. The linked issue's original reproduction, observed payload, expected behavior, and maintainer comments must be checked before judging whether the implementation is a real fix. The implementation deliberately keeps issue discovery out of `pr-context`: the Issue Fidelity agent fetches GitHub's closing-issue metadata with `gh pr view --json closingIssuesReferences`, then fetches relevant issue discussions with `gh issue view --json title,body,comments` (the `--json` form is required — it returns the issue **body**, which `--comments` alone omits). This keeps relevance judgment in the agent instead of baking fragile PR-body parsing into TypeScript. The agent runs only for PR targets — a local-diff or file-path review has no PR or linked issue, so it is skipped there (
|
|
33
|
+
Bugfix PRs often carry their own diagnosis in the PR body, but that diagnosis can be wrong. The linked issue's original reproduction, observed payload, expected behavior, and maintainer comments must be checked before judging whether the implementation is a real fix. The implementation deliberately keeps issue discovery out of `pr-context`: the Issue Fidelity agent fetches GitHub's closing-issue metadata with `gh pr view --json closingIssuesReferences`, then fetches relevant issue discussions with `gh issue view --json title,body,comments` (the `--json` form is required — it returns the issue **body**, which `--comments` alone omits). This keeps relevance judgment in the agent instead of baking fragile PR-body parsing into TypeScript. The agent runs only for PR targets — a local-diff or file-path review has no PR or linked issue, so it is skipped there (11 agents instead of 12).
|
|
33
34
|
|
|
34
35
|
The agent also enforces the root-cause ownership gate: a client-side parser/sanitizer workaround for malformed upstream output is not acceptable as a root-cause fix unless a maintainer explicitly asked for that defensive mitigation.
|
|
35
36
|
|
|
@@ -39,14 +40,40 @@ A single undirected agent has prompt-induced bias and tends to find the same kin
|
|
|
39
40
|
|
|
40
41
|
Empirically, ensemble diversity drops sharply past 3-5 sampled paths. Three is the sweet spot: enough to break single-prompt bias, few enough that the marginal cost stays bounded.
|
|
41
42
|
|
|
43
|
+
### Why Correctness is three procedural agents, not one topical agent
|
|
44
|
+
|
|
45
|
+
A topic brief ("find correctness bugs") lets the agent choose where to look, and independently-prompted agents converge on the same visibly-suspicious hunks — redundancy, not coverage. A procedural brief fixes the walk: every hunk line-by-line with its enclosing function (1a); every deleted line, asking where the deleted invariant is re-established (1b); every changed symbol's callers and read sites (1c). Complementary coverage comes from the walk itself, not from luck. The evidence is in this skill's own history: the whole-file invariant checklist — a procedural walk — found the five PR #6457 Criticals that both the topical dimension agents and 14 chunk agents missed ("what the chunk agents lack is not the lines; it is the question").
|
|
46
|
+
|
|
47
|
+
Two structural holes this closes:
|
|
48
|
+
|
|
49
|
+
- **Removed behavior was nobody's job.** A deleted guard, error path, or test leaves no trace in the post-change tree; only the diff's `-` lines witness it. Heavy files got this covered via the invariant agents' `diffRange`; an ordinary diff's deletions had no dedicated reader. Agent 1b is that reader.
|
|
50
|
+
- **Cross-file was everybody's job, which is the same thing.** The consumer/producer analysis was a shared duty of Agents 1–6: six agents re-running the same greps (~6× the tool calls), none accountable for finishing the walk. Step 3B had already consolidated it into one whole-diff agent; 3A now matches. Single ownership is also the shape the producer-direction lesson (PR #6621) demands — the read site of a never-populated field lives in a file no topical reviewer would open on its own initiative.
|
|
51
|
+
|
|
52
|
+
The language-pitfall and wrapper/proxy checklists fold into 1a rather than standing alone: they are line-level questions asked during the same walk, not separate walks.
|
|
53
|
+
|
|
54
|
+
### Why removed-behavior is a whole-diff agent in 3B, not only a chunk duty
|
|
55
|
+
|
|
56
|
+
3B folds Agent 1b into each chunk agent, scoped to "the deleted lines in your territory". That is necessary and — as PR #6638 proved — not sufficient. Territory-scoped 1b can only ask "was this deletion re-established _here_", and for the deletions that matter most the answer is somewhere else entirely.
|
|
57
|
+
|
|
58
|
+
The measurement: three reviewers ran over #6638 (extension management v2 — 43 files, 8 255 additions, 28 chunks). The 3B run with per-chunk 1b reported **one** Critical. An independent reviewer (Codex `$qreview`) reported 32, and a parallel hand-run wave of 1b + 1c agents over the same commit independently reproduced six of them. Every one of that overlapping six is a **cross-chunk deletion**: `enableByPath(includeSubdirs: true)` deleted in one file and replaced by an exact-path `setWorkspaceActivation` in another, silently narrowing what a workspace-scoped disable means for every untouched CLI/TUI caller; `refreshTools()` dropped from the activation paths, its replacement swallowing the errors it used to propagate; a global mutation timeout removed and replaced by one that covers only the prepare phase. Each has a deletion in chunk A, a replacement in chunk B, and a consumer in a file the diff never touches. **No chunk agent can see that triple, and 1c does not look for it.** The split is by task, not by symbol: 1c owns caller compatibility — it greps the removed export's old name (right there in the deleted lines) and checks each call site — while 1b owns the pairing, finding the _replacement_ and comparing its semantics to what was deleted. A replacement that leaves every call site compiling is all 1c can see; that it now means something different at every one of them is what only 1b goes looking for.
|
|
59
|
+
|
|
60
|
+
So 1b joins 1c as a whole-diff agent, with an explicit split: **1c walks the callers; 1b walks the replacement and compares its semantics.** The chunk agents keep the local half (a guard deleted and not re-established within the same hunk is theirs, and it is the common case). The cost is one agent per 3B review. The class it closes is the one where a replacement type-checks, compiles, passes every test, and means something different to callers nobody edited.
|
|
61
|
+
|
|
62
|
+
### Why diff-specialized finders (Agent 8) are optional and capped at 2
|
|
63
|
+
|
|
64
|
+
Domains have failure grammars — a reconnect state machine, a module loader, a cron scheduler each fail in ways no generic dimension list names. The whole-file invariant checklist is the fixed-form ancestor: a domain-specific walk out-finds a generic brief over the same lines. Agent 8 generalizes that idea to the diff's dominant domain, with the brief written per-review by the orchestrator. Capped at 2 so the fan-out stays bounded and specialization happens only when a domain actually dominates; zero is the common case. Findings flow through Step 4 verification like any other `[review]` finding.
|
|
65
|
+
|
|
42
66
|
## Why batch verification instead of N independent agents?
|
|
43
67
|
|
|
44
68
|
**Considered:**
|
|
45
69
|
|
|
46
70
|
- **N independent agents (original design):** One verification agent per finding. Each reads code independently. High quality but cost scales linearly with finding count (15 findings = 15 LLM calls).
|
|
47
|
-
- **1 batch agent (
|
|
71
|
+
- **1 batch agent (original):** Single agent receives all findings, verifies each one. Fixed cost.
|
|
72
|
+
- **Sharded batches, ≤8 findings each (chosen):** `ceil(F/8)` agents (F = finding count), launched together.
|
|
48
73
|
|
|
49
|
-
**Decision:**
|
|
74
|
+
**Decision:** Shard. One batch agent was right when a review produced 15 findings — it saw cross-finding relationships and cost O(1). But a Step 3B review of a large PR produces 30-60 findings, and one agent re-reading code for each of them inside a single context window degrades on the tail of the list. Sharding costs `ceil(F/8)` calls instead of 1, still far below one-agent-per-finding, and keeps each verifier's job small enough to do properly.
|
|
75
|
+
|
|
76
|
+
**Rejecting a Critical requires quoted contradiction.** A verifier may reject a Critical only when it can quote the specific code that contradicts the claim (the finding describes behavior the code demonstrably does not have) or when the finding merely re-describes a change the diff's own text documents as deliberate; anything less certain is downgraded to low confidence, never deleted. A rejected Critical is deleted from both the PR and the terminal and no later stage revisits it; a downgraded one still reaches a human under "Needs Human Review". The asymmetry between a false positive (noise) and a wrongly deleted true positive (a shipped bug plus another `/review` round) is why the bar for rejection is quoted evidence, not judgment.
|
|
50
77
|
|
|
51
78
|
## Why reverse audit is a separate step, and why iterative
|
|
52
79
|
|
|
@@ -59,13 +86,67 @@ Verification is targeted (check specific claims at specific locations). Reverse
|
|
|
59
86
|
|
|
60
87
|
### Why iterative (multi-round)
|
|
61
88
|
|
|
62
|
-
A single reverse audit pass leaves whatever the reverse audit agent itself missed. Each new round receives the cumulative finding list from prior rounds, so it focuses on what's left undiscovered.
|
|
89
|
+
A single reverse audit pass leaves whatever the reverse audit agent itself missed. Each new round receives the cumulative finding list from prior rounds, so it focuses on what's left undiscovered.
|
|
90
|
+
|
|
91
|
+
### Why the stop rule is two consecutive dry rounds, not one
|
|
92
|
+
|
|
93
|
+
One dry round was the original rule, and PR #6457 shows why it is unsound. The per-round Critical yield across its eight review rounds was `2, 2, 7, 0, 0, 5, 3, 1`. The review returned "no blockers" **twice**, and the next round surfaced five Criticals — three of them in code that had been in the diff since the first commit. A yield of zero is evidence about one round's agents, not about the code.
|
|
94
|
+
|
|
95
|
+
Requiring two consecutive dry rounds makes a single lazy or context-starved agent unable to end the loop. The hard cap moves from 3 rounds to 5, and when the cap is what stopped the loop the output must say so rather than implying convergence.
|
|
96
|
+
|
|
97
|
+
### Why the reverse audit fans out per chunk
|
|
98
|
+
|
|
99
|
+
The original design gave one agent the whole diff plus a growing cumulative finding list. On a 5 800-line diff that is the most context-starved agent in the pipeline — exactly on the PRs where reverse audit matters most. Under Step 3B each round runs one auditor per chunk, each with the full cumulative finding list but only its own territory to re-read.
|
|
100
|
+
|
|
101
|
+
### Why the topology gate counts source lines, not diff lines
|
|
102
|
+
|
|
103
|
+
Diff size is a bad proxy for review risk, because tests dominate it. Across this repo's last 40 merged PRs the median diff is **41% test code**, and 14 of the 40 are more than half tests. A gate on raw diff lines sends a change of 173 production lines that ships 489 lines of new tests into the territory fan-out, where the production code ends up owned by a single chunk agent — while under the dimension fan-out it would have been read by ten lenses (the diff-reading dimension agents: twelve minus Issue Fidelity and Build & Test).
|
|
104
|
+
|
|
105
|
+
Territory fan-out is worth it when there is a lot of _risky_ code to divide, not a lot of _lines_. So the gate is `srcDiffLines > 500`, with a second clause `diffLines > 3200` as an attention bound: past that point asking ten diff-reading lenses each to swallow the whole diff dilutes all of them, and the chunk topology's base cost (`ceil(diffLines / 400) + 4`, counting the whole-diff agents that read the diff — Build & Test reads none) crosses twelve about there. It is not a promise of fewer calls — a heavy file adds three invariant agents and a dominant domain up to two specialized finders — but of one accountable reader per line instead of ten diluted ones. On the 40-PR sample the second clause never fires; it exists for a changeset dominated by tests or generated files.
|
|
106
|
+
|
|
107
|
+
Re-gating moved 6 of those 40 PRs from 3B back to 3A and cost 22 extra agents in total across all 40 — about 5% — measured under the earlier 10-agent 3A roster; under the current 12-agent roster the same six PRs cost 2 more each, ~34 extra (~7%). It buys those six PRs ten review lenses on their production code instead of one.
|
|
108
|
+
|
|
109
|
+
Chunking itself is unchanged: the plan still tiles every line, tests and generated files included. Only the count of reviewers and their brief change. `heavy` is likewise restricted to `source` files — the invariant checklist asks about fields, timers, collections, and error taxonomies, and a rewritten test file has none of those.
|
|
110
|
+
|
|
111
|
+
### Why `plan-diff` exists
|
|
112
|
+
|
|
113
|
+
Step 3B's chunk agents are defined as "one per entry in `chunks[]`", and only `fetch-pr` produced a chunk plan. A local-diff review, or a cross-repo review in lightweight mode, therefore routed into a topology it had no chunk list for: no receipts, no tiling guarantee, and the orchestrator left to improvise line ranges. Two of the four review paths were promised a mechanism the skill could not deliver.
|
|
114
|
+
|
|
115
|
+
`qwen review plan-diff <diff-file>` reads a captured diff and emits the same `chunks[]`, `files[]` and topology counts. Redirecting `git diff` or `gh pr diff` to a file already bypasses the 30 000-char shell cap, so all four paths now share one code path. It cannot decide `heavy` — that needs a tree to read the post-change file from — so a bare diff gets chunk agents but no invariant agents.
|
|
116
|
+
|
|
117
|
+
### Why the topology gate ignores prose
|
|
118
|
+
|
|
119
|
+
`docs/**` and root-level markdown classify as `docs` and stay out of `srcDiffLines`. A translation PR carries no runtime risk, and gating on raw size would fan chunk agents across it. Markdown _inside a source tree_ stays `source`: this repo's bundled skill prompts are `packages/core/src/skills/**/SKILL.md`, and they are executable behaviour. Coverage is unaffected either way — every line is still chunked and receipted.
|
|
120
|
+
|
|
121
|
+
### Why the invariant checklist is split across three agents
|
|
63
122
|
|
|
64
|
-
|
|
123
|
+
Measured on PR #6457's `QQChannel.ts` (1551 → 2643 lines, 65% rewritten), at its first commit, against the nine defects maintainers later confirmed in that commit:
|
|
65
124
|
|
|
66
|
-
|
|
125
|
+
| Reviewer | Invariant-class defects found |
|
|
126
|
+
| ----------------------------------------- | ----------------------------- |
|
|
127
|
+
| One agent, all eight checks | 1 of 5 |
|
|
128
|
+
| Three agents, 2-3 checks each, same model | 5 of 5 |
|
|
129
|
+
| 14 chunk agents (Step 3B), same diff | 0 of 5 |
|
|
130
|
+
| 8 dimension agents on the truncated diff | 2 of 5 |
|
|
67
131
|
|
|
68
|
-
|
|
132
|
+
The chunk agents _saw_ every one of those five defects — the code was inside their territory — and reported none of them. Visibility is necessary and not sufficient. What the chunk agents lack is not the lines; it is the question. "Review this diff for bugs" and "list every retry counter, then check the increment at every call site" are not the same instruction, and only the second one finds an unreachable ceiling.
|
|
133
|
+
|
|
134
|
+
Eight simultaneous checks over a 2 400-line file is a task an agent performs once, shallowly. Three agents with two or three checks each perform it three times, deeply. The cost is two extra calls per heavy file.
|
|
135
|
+
|
|
136
|
+
### Why reverse audit findings no longer skip verification
|
|
137
|
+
|
|
138
|
+
They used to, on the theory that the auditor "already has full context, so its output is inherently high-confidence." That premise is false precisely when the diff is large: the agent with the least room to think was the one whose output nobody checked. Verification is sharded now, so the marginal cost of including reverse-audit findings is small.
|
|
139
|
+
|
|
140
|
+
## Why findings carry a failure scenario instead of an impact statement
|
|
141
|
+
|
|
142
|
+
`Impact` asked why the finding matters. `Failure scenario` asks the finder to prove the finding can happen: name the input/state/timing that triggers it and the wrong outcome that results — or, for quality findings, the concrete cost (what is duplicated, wasted, or harder to maintain, or the quoted project rule).
|
|
143
|
+
|
|
144
|
+
Two effects:
|
|
145
|
+
|
|
146
|
+
1. **Finders self-filter.** A "risk" for which no trigger can be constructed dies at the source instead of reaching the PR. Dogfood motivation: a /review run on PR #6612 auto-published two hallucinated Criticals onto an already-approved PR — both were findings for which no concrete trigger could have been written down. An `Impact` field accepts "this could cause issues in production"; a `Failure scenario` field does not.
|
|
147
|
+
2. **Verifiers get a testable claim.** Step 4's verdict becomes the result of tracing the claimed trigger through the real code — confirmed (high) = the trace works and the lines are quoted; confirmed (low) = mechanism real, trigger uncertain; rejected = the code contradicts the claim — rather than a plausibility vote on the finding's prose.
|
|
148
|
+
|
|
149
|
+
The reporting gate is severity-asymmetric, matching the recall rules elsewhere in the skill: a Suggestion with no scenario and no cost is dropped at the source; a suspected Critical with an uncertain trigger is kept at `Confidence: low` for the verifier to rule on. A dropped Suggestion costs a nicety; a dropped Critical costs a shipped bug.
|
|
69
150
|
|
|
70
151
|
## Why low-confidence over rejection on uncertain findings
|
|
71
152
|
|
|
@@ -147,6 +228,15 @@ Line-based classification was chosen because it's deterministic, cheap, and catc
|
|
|
147
228
|
- Any failure → downgrade `APPROVE` to `COMMENT`, body explains.
|
|
148
229
|
- All pending → downgrade to `COMMENT` (don't approve before CI decides), body explains.
|
|
149
230
|
|
|
231
|
+
**The hole under all of this: a check that never ran looked like a check that passed.** GitHub reports a skipped job as `status: completed, conclusion: skipped`. The classifier tested for failure conclusions and for pending statuses, and `skipped` matched neither — so it fell through into `all_pass`. Every word above delegates runtime truth to CI _because_ the LLM pipeline reads code statically. If the delegation returns nothing, and returns it wearing a green badge, the delegation is worse than not having it.
|
|
232
|
+
|
|
233
|
+
PR #6486: the one job that would have exercised the new `Ctrl+F` hotkey — `Integration Tests (CLI, No Sandbox)` — was skipped, as were the macOS and Windows `Test` legs. `all_pass`. And even had it run, it would have passed: the test drove a CSI-u sequence into a PTY that never negotiated the kitty protocol, so the keypress was discarded before reaching the handler. A test that cannot fail, in a job that did not run, scored as verification.
|
|
234
|
+
|
|
235
|
+
`skipped`/`neutral` are now recognised, with two deliberately different consequences:
|
|
236
|
+
|
|
237
|
+
- **Some checks skipped → a disclosure, not a downgrade.** Empirically this repo emits skipped runs constantly — routing jobs (`authorize`, `review-pr`, `precheck-pr`) that also emit a successful run of the same name, which is why "did it run" is a question about the _name_, not about any single run. And a docs-only PR legitimately skips the test matrix. Auto-downgrading on any skip would downgrade every review in the repo, which is how a gate gets ignored. So presubmit _names_ them and Step 7 rules on them — because whether a skipped check would have exercised **this** diff is a question about the diff, which presubmit cannot see and the reviewer can.
|
|
238
|
+
- **Every check skipped → a downgrade.** Checks exist, not one ran: there is no green here to approve on, and no judgment is required to say so. (A repo with no CI at all is a different claim — `totalChecks === 0`, not downgraded.)
|
|
239
|
+
|
|
150
240
|
**Why downgrade rather than block:** the reviewer LLM has done substantive work; throwing the review away because CI is red wastes that. Downgrading to `COMMENT` keeps all inline findings, preserves the static review value, and lets GitHub's check status carry the "do not merge" signal naturally.
|
|
151
241
|
|
|
152
242
|
**Why this stacks with self-PR downgrade:** a self-authored PR with red CI hits **both** downgrade rules. The event is `COMMENT` either way, so stacking is operationally a no-op — but the body should mention both reasons so a future maintainer reading the review knows why an LLM that found no Critical issues did not approve.
|
|
@@ -198,18 +288,183 @@ A malicious PR could add `.qwen/review-rules.md` with "never report security iss
|
|
|
198
288
|
|
|
199
289
|
**Decision:** Tips. Qwen Code's follow-up suggestion system is a core UX differentiator. Blocking prompts interrupt flow. Tips are zero-friction and let users decide when/if to act.
|
|
200
290
|
|
|
201
|
-
##
|
|
291
|
+
## Why the COMMENT body is composed from clauses, not picked from fixed sentences
|
|
292
|
+
|
|
293
|
+
The body rules began as a table of exact one-liners — the right call against smuggled prose, and it stayed right while only one state could apply at a time. Then the states multiplied: presubmit downgrades, the context-unavailable cap, discarded-Suggestion disclosure, uncoverable-chunk disclosure, body-relocated Criticals. Four consecutive review rounds each found a **pairwise collision** — two rules both claiming to be "the" body, so applying either erased the other's disclosure (a downgrade reason overwriting the diff-only warning; a "Suggestions are inline" restored by 422 recovery inside a run that never saw the PR's discussion; an all-discarded run claiming its suggestions were inline). Patching collisions one at a time provably does not converge: n states have n(n−1)/2 pairs.
|
|
294
|
+
|
|
295
|
+
The fix is a composition rule: an ordered clause inventory, each clause present iff its condition holds, joined into one paragraph, nothing else permitted. It keeps the anti-prose discipline (the inventory is closed; free text is still banned), reduces to the table's exact sentences in the single-state case, and makes every future state additive — a new state adds one clause, not one patch per existing state. `C` is likewise defined once, globally (everything the review posts, anywhere — inline or body), so no downstream rule can re-derive it over a subset and delete a body-only blocker.
|
|
296
|
+
|
|
297
|
+
## Why parse-args and compose-review are subcommands, and pr-context renders bodies in full
|
|
298
|
+
|
|
299
|
+
Seven rounds of review-the-review on this PR converged on one diagnosis: the skill's deterministic logic kept shipping bugs precisely where it was written as prose. Argument parsing produced three bugs (a flag consumed as a value, the `=` form undefined, an invalid value leaking into target disambiguation). The event/body machine produced five (four Critical), all one shape — a downstream branch not updated when an upstream rule gained a new state, because the machine was restated in four places that had to be synchronized by hand: n states, n(n−1)/2 pairwise collisions, patched one at a time without converging. And the "fetch review bodies for the re-check" instruction was rewritten **five times in four rounds** (missing pagination → shell truncation → unpageable single-line JSON → a marker filter that discarded markerless blockers → offline selection), which is what writing a download program in English looks like.
|
|
300
|
+
|
|
301
|
+
The resolution is the same one this document already records for presubmit and cleanup: judgment stays in the prompt, bookkeeping moves to tested subcommands that version together with the skill.
|
|
302
|
+
|
|
303
|
+
- **`parse-args`** owns the grammar. Every previously-shipped parsing bug is a named row in its table-driven tests. The raw string travels **on stdin** (`--stdin` with a quoted heredoc), never as a positional: a flag-first raw string (`/review --effort low`) is consumed by the CLI's own strict parser before the handler runs, and a positional also breaks on quotes and shell metacharacters. Pure-function tests could not see that class — the documented invocation failed only when run against the built binary — so the suite includes yargs-level wiring tests alongside the table.
|
|
304
|
+
- **`compose-review`** owns event selection and body composition — the C/S table (counting body Criticals and discarded Suggestions), the event caps (cannot-tell existing Criticals, uncoverable chunks, unreviewed dimensions, context-unavailable), the downgrade carve-outs, and the clause composition. Its truth-table tests pin each shipped bug; writing them immediately caught one more instance of the class (all Suggestions discarded → S=0 → APPROVE). The input is validated at the boundary: the producer is a model writing JSON that omits inapplicable fields, so absent counts default to zero and malformed values throw typed errors — before that, an omitted count meant `undefined + 1 = NaN`, which fails every event comparison and would have returned APPROVE over a body-only blocker. 422 recovery stops being a hand-derived recomposition: it is the same call with updated counts, so the "recompute may never upgrade the verdict" guarantee holds by construction.
|
|
305
|
+
- **`pr-context`** ends the fetch-prose chain at its root: review bodies **and every blocker-bearing body** render **in full** (a body-only blocker lives only there; a capped body names its review or comment id so the tail stays fetchable one object at a time, and reply snippets name their comment id when cut), and blocker-bearing threads are quarantined into a "Blockers to re-check" section instead of settling into "Already discussed" — a reply alone never retires a blocker. The `gh` wrapper's `maxBuffer` rises to 64 MiB, closing the ENOBUFS that killed two subcommands mid-review on a comment-heavy PR.
|
|
306
|
+
|
|
307
|
+
What deliberately stays prose: everything judgment-shaped — what counts as a Critical, verification, the posting gate's authorization semantics, the angles. A truth table cannot decide whether a finding is real; it can guarantee that a real finding is never mislabeled, dropped by a downgrade, or approved past.
|
|
308
|
+
|
|
309
|
+
## Why blocker recognition is semantic, not the `[Critical]` marker
|
|
310
|
+
|
|
311
|
+
The mandatory re-check section used to be gated on the literal string `[Critical]`. That marker is emitted by exactly one author — `/review` itself. Every human blocker was therefore invisible to the gate, and the fallback was a prose instruction in Step 6 telling the model to also scan "Already discussed" semantically.
|
|
312
|
+
|
|
313
|
+
Prose does not beat structure. PR #6486 is the proof, and it cost a shipped blocker.
|
|
314
|
+
|
|
315
|
+
A maintainer built the PR, drove the real CLI through a PTY, and found that `Ctrl+F` **dual-fires** — it toggles the model _and_ moves the input cursor, because `text-buffer.ts:2663` still binds `Ctrl+F → move('right')` and both handlers are independent subscribers of a `KeypressContext.broadcast()` that has no stop-propagation. They filed it as an **issue comment**, headed `🔴 Finding 1 — … (blocker)`. No `[Critical]` marker, because a human wrote it.
|
|
316
|
+
|
|
317
|
+
Three things then compounded:
|
|
318
|
+
|
|
319
|
+
1. Issue comments all settle into **"Already discussed — do NOT re-report"**.
|
|
320
|
+
2. They render as **240-character one-line snippets**.
|
|
321
|
+
3. The first 240 characters of a verification report are its **preamble**: _"I built this PR from source and drove the real CLI … to validate the model-toggle hotkey before merge. Sharing the results as a merge reference."_
|
|
322
|
+
|
|
323
|
+
So the one artifact that proved the PR was broken was presented to the review agents as a **maintainer endorsement**, in the section that says not to re-report it. The blocker itself began 1 143 characters past the cut. Three hours later `/review` reviewed the same commit — the fix did not land until that evening — and submitted **"Reviewed — no blockers"**. This is precisely the "dropped blocker" failure the Step 6 re-check exists to prevent, and the re-check could not prevent it, because the input it was handed said the opposite of the truth.
|
|
324
|
+
|
|
325
|
+
The fix moves the decision out of prose and into `carriesBlockerSignal`: any body asserting a blocking defect — inline thread or issue comment, `[Critical]` or `(blocker)` or "is a blocker" or "must fix" or "still reproducible" or 阻塞项 — is promoted into **"Blockers to re-check"** and rendered **in full**. A bare `🔴` is deliberately **not** a signal, for the reason the next paragraph measures.
|
|
326
|
+
|
|
327
|
+
Two properties are deliberate:
|
|
328
|
+
|
|
329
|
+
- **Fail-safe direction.** A false positive costs one extra ruling by the re-check; a false negative ships the bug. When in doubt, promote.
|
|
330
|
+
- **Precision still matters, in the other direction.** Promotion means full-body rendering, and a context file that outgrows one `read_file` is its own way of losing a blocker (PR #5738, recorded above). The prose scan of "Already discussed" is retained as a floor — `carriesBlockerSignal` recognises the phrasings we have seen, not every phrasing that exists.
|
|
331
|
+
|
|
332
|
+
**Both of those were nearly undone by the first implementation, and only a live run showed it.** That version scanned the whole body for the words `blocker`, `🔴`, `阻塞`, `[Critical]`. Run against the real #6486 thread it promoted **8 of 15** issue comments; exactly one was a live blocker. The others were the triage bot's own template line **"No critical blockers."** (the word inside its own negation), the author's **"### 🔴 Critical fixes"** (a severity emoji on a list of repairs), and a later comment _quoting_ `[Critical]` while arguing a finding away. Eight full bodies took the context file from 30 KB to 59 KB and pushed the real blocker to character **43 094** — past the 25 000 one `read_file` returns. The section existed, held the right blocker, and no agent could see it: PR #5738's failure, reintroduced one section further down by the fix for it.
|
|
333
|
+
|
|
334
|
+
Three changes, and the ordering one is load-bearing:
|
|
335
|
+
|
|
336
|
+
- **The section is written FIRST**, ahead of the description and the review history. Nothing in the file outranks the claims a `C=0` verdict may not be reached without ruling on. On the live thread this moved the heading from char 25 961 to **569**, and the blocker body from 43 094 to **4 421**.
|
|
337
|
+
- **Recognition matches assertion patterns, not word presence** — `[Critical]`, `(blocker)`, `is a blocker`, a bare `blocking` (with a `non-blocking` / `非阻塞` lookbehind), `must fix`, `still reproducible/repro/broken/fails`, `阻塞项/问题/点` — with a **bilingual** negation guard, so neither "no blockers" nor "没有阻塞项" ever promotes. Live promotions dropped 8 → 3 (the one real blocker plus two harmless mentions), and the file 59 KB → 40 KB.
|
|
338
|
+
- **The section carries a character budget.** Tight patterns keep promotion rare; the budget keeps a pathological thread from blowing the read window anyway. Bodies past it degrade to snippets **naming their exact fetch**, which the re-check already must run before ruling — not to silence.
|
|
339
|
+
|
|
340
|
+
The lesson generalizes past this file: **"a false positive is cheap" is a claim about a budget, and it has to be measured against the real distribution, not assumed.** Here it was false until the ordering was fixed.
|
|
341
|
+
|
|
342
|
+
## Why a test-efficacy probe, when there is already a Test Coverage agent
|
|
343
|
+
|
|
344
|
+
Agent 5 asks whether a test **exists** and whether its assertions **look like** they check something. Agent 7 runs the suite and reports that it is **green**. Neither can see a test that protects nothing, and there are two ways to ship one:
|
|
345
|
+
|
|
346
|
+
- **Unreachable** — the project's test command never collects the file.
|
|
347
|
+
- **Inert** — it runs, it passes, and it would still pass with the change reverted.
|
|
348
|
+
|
|
349
|
+
PR #6486 shipped both, in one file. The new test lived in `integration-tests/`, which is not an npm workspace, so `npm test --workspaces` never collected it; its CI job (`Integration Tests (CLI, No Sandbox)`) was skipped, so CI never ran it either. **The test executed nowhere — not in CI, not in the review — and nothing in the pipeline noticed.** And had it run, it would have passed regardless: it drove a kitty CSI-u sequence into a PTY that never negotiated the kitty protocol, so the keypress was discarded before reaching the handler under test. It could only ever have caught a startup crash. Agent 5 saw a test file with plausible assertions and said coverage was fine.
|
|
350
|
+
|
|
351
|
+
Both questions are decidable without judgment, which is why they are a subcommand and not a prompt. Unreachability needs no execution at all — it is a path against the root `package.json` workspace globs. Inertness needs one run: revert the diff's **source** files to base, keep its **tests**, re-run them. A test that is still green is green whether or not the feature exists.
|
|
352
|
+
|
|
353
|
+
**The trap, and the reason the classifier is asymmetric.** Reverting source frequently breaks the test's own compile — it imports a symbol the diff introduced — and the runner exits non-zero having collected nothing. It is tempting to score that as "the test caught the revert". It is not: a compile error says nothing about whether the test would catch a _behavioural_ regression, and scoring it as `gated` would hand back precisely the false assurance this command exists to remove. So `gated` requires a real **assertion** failure; a bare non-zero exit with nothing collected is `inconclusive`, and `inconclusive` is never reported as a finding.
|
|
354
|
+
|
|
355
|
+
Two other deliberate limits:
|
|
356
|
+
|
|
357
|
+
- **A test-only diff is never probed.** A new test for old code is _supposed_ to pass with nothing reverted. Probing it would flag every such PR as inert — a false blocker on exactly the PRs we want people to write.
|
|
358
|
+
- **Findings are Suggestions, not Criticals.** A test that does not gate is not itself wrong code; nothing is broken today. What the finding must say concretely is which behaviour is now shipping unprotected.
|
|
359
|
+
|
|
360
|
+
## Why "fixed by this diff" is the verdict that needed a bar
|
|
361
|
+
|
|
362
|
+
The re-check has three verdicts, and until PR #6486 only two of them cost anything:
|
|
363
|
+
|
|
364
|
+
| verdict | consequence |
|
|
365
|
+
| -------------------- | ----------------------------------------------------- |
|
|
366
|
+
| `still stands` | `REQUEST_CHANGES` — blocks the merge |
|
|
367
|
+
| `cannot tell` | serialized into the body, caps the event at `COMMENT` |
|
|
368
|
+
| `fixed by this diff` | **nothing. Silent, free, unrecorded.** |
|
|
369
|
+
|
|
370
|
+
An agent under context pressure, choosing among three answers where one is free and two are not, drifts toward the free one — and the free one is the only one that can ship a bug.
|
|
371
|
+
|
|
372
|
+
Worse, the bar for it read "you read the lines and the fix is there", which invites reading **the diff's lines**. That is precisely the reading that fails. A fix's new lines are always in the diff; whether they _work_ routinely depends on code outside it.
|
|
373
|
+
|
|
374
|
+
PR #6486 is the case. A `Ctrl+F` dual-fire blocker was filed — the hotkey toggled the model _and_ moved the input cursor. The author added a guard to the toggle handler: visible in the diff, and it reads like a fix. It changed nothing. The second handler is `text-buffer.ts:2663`, in a file the PR never touches, subscribed independently to a `KeypressContext.broadcast()` that has no stop-propagation — `return`ing from one subscriber does not stop the other. Read the diff and you see a guard and rule "fixed". Read `text-buffer.ts:2663` and you cannot.
|
|
375
|
+
|
|
376
|
+
Two changes, split the way this document keeps arriving at — **determinism owns the evidence, judgment owns the ruling**:
|
|
377
|
+
|
|
378
|
+
- **`pr-context` extracts the evidence** (`extractCodeRefs`). A blocker's body names the code it is about — #6486's named `text-buffer.ts:2663` outright — so a promoted blocker that names a file now renders a **Referenced code** list (a blocker citing no path gets none — the reader traces the mechanism themselves). "Go read the untouched code" stops being a hope the agent might have and becomes a list it is handed.
|
|
379
|
+
- **SKILL.md raises the bar** on the ruling: name the mechanism, name what now stops it, and when the stopping condition lives outside the diff, read it there — or the verdict is `cannot tell`.
|
|
380
|
+
|
|
381
|
+
No new `compose-review` input was needed: `cannot tell` already caps the event. The change is to make wrong "fixed" rulings land there instead of passing silently.
|
|
382
|
+
|
|
383
|
+
## What the first dogfood batch changed
|
|
384
|
+
|
|
385
|
+
Six concurrent real-PR runs (batch 3) produced three targeted changes, each fixing something the batch measured rather than predicted:
|
|
386
|
+
|
|
387
|
+
- **Overlap disposal is deterministic.** presubmit's overlap report used to end in "ask the user whether to proceed" — 2 of 6 runs stalled on an improvised interactive question (fatal for a headless run) while the other 4 proceeded, the signature of an under-specified decision point. An overlap is a duplicate by the Exclusion Criteria; the rule is now drop, note in the terminal, continue — and the counts handed to `compose-review` shrink accordingly, so a dropped finding can never flip the verdict.
|
|
388
|
+
- **Host routing is a flag, not prose.** The GH_HOST-by-prefix instruction survived exactly one review round before a reviewer noted the model must remember it per call. `--host` on `fetch-pr` / `pr-context` / `presubmit` routes every wrapped `gh` call in code (`lib/gh.ts` `setGhHost`/`ghEnv`), leaving the prose rule only for the handful of `gh` commands the orchestrating model runs directly.
|
|
389
|
+
- **A fixed completion line.** Three different completion phrasings across one batch each needed their own detection regex in the batch driver. Step 9 now ends every run with `Review complete: <target> — <disposition>`, greppable by `^Review complete: `.
|
|
390
|
+
|
|
391
|
+
## Why Step 7 opens with a hard posting gate
|
|
392
|
+
|
|
393
|
+
Posting is the only irreversible, public, outward-facing action the skill takes, and it must never happen as a side effect of a confident verdict. The skip condition existed from the start, but it was phrased as one clause among several ("skip if … or if BOTH `--comment` absent AND no post request"), which a model evaluates as a judgment call at the end of a long run — exactly when it is reasoning about what it wants to say rather than about what it was authorized to do.
|
|
394
|
+
|
|
395
|
+
Dogfooding proved the phrasing insufficient: across four concurrent no-`--comment` reviews, three correctly withheld (offering the follow-up tip) and one self-submitted a `COMMENT` review with an inline suggestion to a real PR. One violation in four is a model-adherence failure, not a logic error — the rule was right, its force was not.
|
|
396
|
+
|
|
397
|
+
The fix promotes the gate to the first thing in Step 7 and reframes it as arithmetic, not judgment: post **only if** `--comment` was parsed in Step 1 **or** the user explicitly asked to post this session; otherwise no `reviews`-API write happens at all, regardless of verdict or the "Tip: post comments" text being printed. This mirrors the `event`/`body` invariant elsewhere in Step 7 ("stop reasoning and count") — the same failure mode (a model rationalizing past a stated rule at submit time) gets the same countermeasure (convert the rule to a check with no discretion).
|
|
398
|
+
|
|
399
|
+
## Why verification checks the diff's own documented intent
|
|
400
|
+
|
|
401
|
+
Verification traces a finding's failure scenario through the code, but "the code does what the finding says" is not sufficient for a finding framed as a **regression** — the code doing X is exactly what a deliberate, documented change to do X looks like. The missing question is whether X is a defect or a design decision, and the diff itself usually answers it: a rationale comment, a JSDoc note, or a test that asserts the new behavior on purpose.
|
|
402
|
+
|
|
403
|
+
Dogfooding auto-posted the failure. A review of a secret-sanitization PR filed a Critical — "third-party credentials (`AWS_SECRET_ACCESS_KEY`, `GITHUB_TOKEN`, `NPM_TOKEN`) now pass through to subprocesses = security regression." The factual claim was true; the framing was wrong. The same file carried a rationale comment three lines from the change — user-managed credentials `must remain available` for shell/MCP/tool subprocesses, and the old broad denylist that scrubbed them was the bug this PR fixed — plus tests that assert the pass-through on purpose. The verifier traced the behavior and confirmed it without reading the rationale, and the Critical published to a real PR.
|
|
404
|
+
|
|
405
|
+
So verification now has an explicit step: for any finding that reads as "regression / removed protection / now allows X", read the diff-local comments and tests for the changed lines, and engage the documented intent. A documented-and-deliberate change is a design decision — reject the finding if it merely re-describes that change without naming any harm the rationale fails to answer. Documentation changes what the verifier must do, not what confidence it may reach: a traced, concrete harm that survives the rationale keeps high confidence (documenting a hole does not make it safe); low confidence is for cases where the rationale makes the harm genuinely uncertain, e.g. it names a compensating control the verifier cannot rule out. It is the diff-local analogue of Agent 0's root-cause-ownership gate (which checks intent against the linked _issue_); this checks intent against the _diff's own text_, which every review path has even when there is no issue. The counterpart finding in that same review — two new `scrubChildEnv(process.env, …)` call sites missing the `normalizePathEnvForWindows` wrapper that every sibling call site uses — had no such rationale and was a real oversight bug; the gate is about documented intent, not about suppressing findings on sanitization PRs.
|
|
406
|
+
|
|
407
|
+
## Why whole-diff agents get a substantive-return check
|
|
408
|
+
|
|
409
|
+
Step 3B's coverage receipts guarantee every chunk was read, but they cover only chunk agents — the whole-diff agents (Issue Fidelity, removed-behavior, cross-file tracer, invariant agents, test-coverage matrix, diff-specialized finders) have no receipt, because they own a concern, not a territory. That left a blind spot symmetric to the one receipts close: an agent that whiffs — returns almost instantly with near-empty output — is indistinguishable from one that examined its concern and found nothing.
|
|
410
|
+
|
|
411
|
+
Dogfooding surfaced it concretely. On a heavy-file review, one of the three invariant agents returned in 11 seconds having emitted ~370 tokens while its siblings ran for minutes and thousands; the fast one owned the checklist half (counters / return-values / error-taxonomy) that, in a parallel exhaustive pass, produced the run's most serious finding. Nothing flagged the whiff, and the orchestrator folded its silence into "no issues in that dimension".
|
|
412
|
+
|
|
413
|
+
The countermeasure is cheap and needs no new machinery: before Step 4, sanity-check that each receipt-less agent's return actually describes its walk (the fields/callers/lines it enumerated) rather than a bare "No issues found." The primary test is evidential, not statistical — a return that names nothing it examined is a non-return regardless of length, and a legitimately empty scope passes as long as it says what it checked. The comparative signal ("far shorter and faster than its peers") is only a prompt to look at that agent's output, never a threshold to relaunch on: no fixed cutoff would survive a review where every agent is legitimately terse. Deliberately no number, because a false relaunch costs one agent call and a missed whiff costs a shipped bug — when in doubt, relaunch. It is the receipt-less analogue of "a chunk with no receipt was never reviewed," and it applies to 3A's dimension agents just as it does to 3B's whole-diff agents, since neither emits a receipt.
|
|
414
|
+
|
|
415
|
+
## Why effort levels (low / medium / high)
|
|
416
|
+
|
|
417
|
+
**Considered:**
|
|
418
|
+
|
|
419
|
+
- **Always-full (original):** every `/review` runs the full pipeline. Right for a PR verdict; wrong for a 5-line pre-commit sanity check — 12 agents, sharded verification, and ≥2 reverse-audit rounds to re-derive what one reader could see in a single pass.
|
|
420
|
+
- **A `--quick` boolean:** two modes, but "quick" hides what is and isn't checked (rules? cross-file? build?).
|
|
421
|
+
- **Three levels (chosen):** **low** = one orchestrator pass over the chunk plan, hunk-visible bugs only, ≤8 findings. **medium** = the finder angles (1a, 1b, 1c, quality/altitude, performance, conventions) run **sequentially in the orchestrator's own context** — inline sequencing, not subagents, is what makes the level cheap — ≤12 findings. **high** = the full pipeline, unchanged.
|
|
422
|
+
|
|
423
|
+
**Guardrails, because a quick pass is recall-limited by construction.** "Quick pass" means **low and medium together** — they differ in depth (one diff pass vs. sequential finder angles; ≤8 vs. ≤12 findings) but share every guardrail below, because what the guardrails defend against is the same at both: findings that no verifier ever checked.
|
|
424
|
+
|
|
425
|
+
- Labeled **unverified**; no Approve/Request-changes verdict is emitted. A verdict is a claim the pipeline earns in Steps 4–5; a quick pass claims findings, not absence of findings.
|
|
426
|
+
- Never posts to the PR: `--comment` forces high, and a "post comments" follow-up after a quick pass is declined.
|
|
427
|
+
- Never consults or writes the incremental cache — otherwise a medium run's SHA would make a later high run report "No new changes since last review", silently converting a quick pass into a full-review verdict.
|
|
428
|
+
- Scope handling (worktree, diff capture, chunk plan) is identical at all levels. The levels change who reads the diff and what runs afterwards, never how the diff is obtained — the base-resolution and truncation traps do not care how fast the user wants the answer.
|
|
429
|
+
|
|
430
|
+
**Defaults:** PR targets → high (the product is a public verdict); local-diff / file-path targets → medium (the product is fast feedback; the closing tip advertises `--effort high`). Findings caps exist only at the unverified levels — at high effort, verification is the noise filter, so no cap is needed.
|
|
431
|
+
|
|
432
|
+
## LLM call budget
|
|
433
|
+
|
|
434
|
+
**Small diffs (≤ 500 source lines AND ≤ 3200 total diff lines, Step 3A, high effort) — 15-21 calls (typically 15-17):**
|
|
435
|
+
|
|
436
|
+
| Stage | Calls | Why |
|
|
437
|
+
| ----------------------- | ------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
438
|
+
| Review agents | 12 (+0-2) | issue fidelity + 3 procedural correctness walks (1a/1b/1c) + security/quality/perf/tests + 3 undirected personas + build&test, plus 0-2 diff-specialized finders; cross-repo skips Agents 7 and 1c (10), non-PR skips Agent 0 (11) |
|
|
439
|
+
| Sharded verification | `ceil(F/8)` | F = findings; typically 1-2; keeps each verifier's job small on high-finding reviews |
|
|
440
|
+
| Iterative reverse audit | 2-5 | loop ends after two consecutive dry rounds; 5-round hard cap |
|
|
441
|
+
| **Total** | **~15-21 (~13-20)** | Row maxima do not co-occur on typical runs (~15-17 is common), but the honest sum of ranges is 15-21 same-repo, 13-20 cross-repo/local. **Low/medium effort: 0 subagent calls** — the inline pass runs in the orchestrator's own context |
|
|
442
|
+
|
|
443
|
+
**Large diffs (> 500 source lines OR > 3200 total diff lines, Step 3B, high effort) — `ceil(diffLines / 400)` chunk agents + `5..7` whole-diff agents + `3H` invariant agents (H = heavy files) + `ceil(F/8)` verify (F = findings) + `rounds × chunks` reverse audit.** The reverse audit dominates: it fans out one auditor per chunk per round, and the stop rule needs two consecutive dry rounds (hard cap 5). PR #6457 (5801 diff lines, 19 chunks, 1 heavy file) costs ~27-29 first-wave calls, then `19 × (2..5) = 38-95` reverse auditors — ~66-126 calls total depending on how long the audit keeps finding; ~70 is the clean-run floor, and the count scales with chunks and findings, not a fixed ceiling.
|
|
444
|
+
|
|
445
|
+
That is roughly 4x the small-diff budget, and it buys the thing the small-diff topology cannot deliver at that size: coverage. Ten dimension agents (the roster of the day; twelve now) on a 5801-line diff each read the same truncated 14% window (see "Why the diff is a file, not a command"), so nine of the ten calls were redundant reads of the same hunks. Nineteen chunk agents each read a distinct ~390-line territory, and every line of the diff has exactly one accountable owner. The comparison to make is not ~70 calls vs ~17: PR #6457 took **eight** review rounds at 12-14 calls each — over 100 calls — and was still surfacing Criticals in code that had been in the diff since the first commit.
|
|
446
|
+
|
|
447
|
+
Competitors: Copilot uses 1 call, Gemini uses 2, Claude /ultrareview uses 5-20 (cloud). Ours biases toward higher recall — the assumption is that "find more issues per round" is more valuable than minimizing per-run cost, because every missed issue forces the user into another `/review` iteration.
|
|
448
|
+
|
|
449
|
+
## Why the diff is a file, not a command
|
|
450
|
+
|
|
451
|
+
Agents used to be handed `git diff main...HEAD` and told to run it. Shell tool output passes through `truncateToolOutput` with `ShellTool.maxOutputChars = 30_000` and `keep: 'both'`, which allocates `threshold / 5` characters to the head and the remainder to the tail.
|
|
452
|
+
|
|
453
|
+
On PR #6457's 211 000-character diff that yields a 6 000-char head (`QQChannel.ts` lines 41-250) and a 24 000-char tail (`stream.test.ts` and `types.ts`, which sort last by path and together changed 9 lines). 85.8% of the diff — including 19 of the 20 Criticals eventually reported on that PR — was replaced by a `[CONTENT TRUNCATED]` marker. Every agent saw the same window, so the ten-way dimension fan-out multiplied redundancy rather than coverage, and each round of `/review` sampled a different subset of the bugs depending on which files an agent happened to `read_file` on its own initiative.
|
|
454
|
+
|
|
455
|
+
`fetch-pr` now writes the diff to `.qwen/tmp/qwen-review-pr-<n>-diff.txt` and emits a chunk plan. `read_file` overrides `maxOutputChars` to `Infinity`, so it escapes the scheduler's head/tail mangling — but `processSingleFileContent` still caps one read at `truncateToolOutputThreshold` (25 000 chars), sets `isTruncated`, and expects the caller to page. Writing the diff to a file is therefore necessary but **not sufficient**: a single `read_file` over PR #6457's diff returns lines 1-611 and stops.
|
|
456
|
+
|
|
457
|
+
The chunk plan is what closes the gap. Chunks are bounded by **both** a line budget (attention) and a character budget (`MAX_CHUNK_CHARS`, 20 000 — under the 25 000 read cap, so a chunk never comes back short), and they tile the diff exactly (`chunksCoverDiff` asserts no gap, no overlap). Exact tiling is what makes the Step 3B coverage receipts checkable: a chunk with no receipt is a territory nobody reviewed.
|
|
202
458
|
|
|
203
|
-
|
|
204
|
-
| ----------------------- | ----------------- | ------------------------------------------------------------------------------------------------------------------------ |
|
|
205
|
-
| Review agents | 10 (9) | issue fidelity + 6 dimensions + 3 undirected personas; Agent 7 skipped in cross-repo, Agent 0 skipped for non-PR reviews |
|
|
206
|
-
| Batch verification | 1 | O(1) not O(N) — batch is as good as individual |
|
|
207
|
-
| Iterative reverse audit | 1-3 | Loop until "No issues found" or 3-round hard cap |
|
|
208
|
-
| **Total** | **12-14 (11-13)** | Same-repo PR: 12-14; cross-repo lightweight PR or local/file (no Agent 0): 11-13 |
|
|
459
|
+
Measured on PR #6457's real 211 000-char diff, driving the production `truncateAndSaveToFile` and `processSingleFileContent`:
|
|
209
460
|
|
|
210
|
-
|
|
461
|
+
| What the agent is given | Chars delivered | Diff covered | Of the 20 Criticals eventually found, in view |
|
|
462
|
+
| ------------------------------------------ | --------------- | ------------ | --------------------------------------------- |
|
|
463
|
+
| `git diff` via shell (the old way) | 30 468 | 14.4% | 1 |
|
|
464
|
+
| Diff in a file, read whole (no chunk plan) | 25 015 | 10.5% | — |
|
|
465
|
+
| Diff in a file + 19-chunk plan | 210 900 | **100%** | **20** |
|
|
211
466
|
|
|
212
|
-
|
|
467
|
+
Chunk boundaries fall on hunk boundaries wherever they can, because a boundary inside a hunk risks cutting a function in half. A hunk larger than the target is the exception: it is split, but only at a column-0 source line preceded by a blank line — a top-level declaration. A brand-new file arrives as one enormous hunk (`events.test.ts` was a single 1535-line hunk), so treating hunks as strictly atomic would hand one agent a 50 000-char territory and defeat the whole point. When no such boundary exists the hunk stays whole and the chunk is flagged `oversized`.
|
|
213
468
|
|
|
214
469
|
## Why cross-repo uses lightweight mode
|
|
215
470
|
|
|
@@ -228,28 +483,29 @@ Key implementation detail: Step 7 must use the owner/repo extracted from the URL
|
|
|
228
483
|
|
|
229
484
|
**Decision:** Auto-discovery. Every project already defines its tool chain in CI config. Reading those files leverages existing knowledge without asking users to duplicate it. The LLM is capable of parsing YAML workflow files and extracting the relevant commands. Falls back gracefully: if no CI config exists, the build/test discovery is simply skipped and LLM agents still review the diff.
|
|
230
485
|
|
|
231
|
-
## Why Suggestion-level findings
|
|
486
|
+
## Why Suggestion-level findings are posted as inline comments, like Critical
|
|
232
487
|
|
|
233
488
|
**Considered:**
|
|
234
489
|
|
|
235
|
-
- **
|
|
236
|
-
- **Critical inline, Suggestion in
|
|
237
|
-
- **
|
|
490
|
+
- **Critical inline, Suggestion in the review `body`:** splits by severity, but the review body is a frozen artifact of one review submission — every new /review run appends a new review with its own body, so Suggestion lists accumulate across runs and never converge.
|
|
491
|
+
- **Critical inline, Suggestion in one updatable issue comment:** Suggestion findings go to a single PR issue comment located by author + embedded marker and PATCHed in place on every run, so the list refreshes rather than grows. Shipped for a while; reverted for the reasons below.
|
|
492
|
+
- **Both severities inline, distinguished by a `**[Critical]**`/`**[Suggestion]**` body prefix (chosen):** every high-confidence finding is pinned to its code line and carries a one-click ` ```suggestion ` block. Severity is communicated in the comment text, not by the channel it arrives on.
|
|
238
493
|
|
|
239
|
-
**Decision:**
|
|
494
|
+
**Decision:** Both inline. The updatable-summary design optimized for a convergence problem, but it paid for that with two costs that turned out to dominate:
|
|
240
495
|
|
|
241
|
-
|
|
496
|
+
1. **A summary comment can never collapse.** GitHub marks an inline review thread **Outdated** and folds it away as soon as the author edits the line it is anchored to. So an addressed inline finding removes itself from the page. An issue comment has no such lifecycle — it sits in the PR conversation permanently, one extra comment whether or not its rows still apply. PATCHing it to "all suggestions addressed" replaces the content but not the comment. The very mechanism intended to prevent clutter _was_ the clutter.
|
|
497
|
+
2. **A Markdown table cannot carry a one-click fix.** GitHub renders a ` ```suggestion ` fence as an applicable change only inside a review comment on a diff line; in an issue comment it degrades to a plain code block. Suggestion-level findings — mechanical, localized cleanups — are precisely the class that benefits most from one-click apply, so the split withheld the feature from the findings that most needed it. The table's cramped "Suggested fix" column also degraded badly as the suggestion count grew.
|
|
242
498
|
|
|
243
|
-
|
|
499
|
+
The convergence concern that motivated the summary is real but narrower than it looked: GitHub's Outdated-collapse handles every suggestion the author actually acts on, which is the common case. What remains is a suggestion the author declines and leaves untouched — its line does not change, so the thread stays open and a later run can post a near-duplicate. That residue is bounded by the presubmit Overlap check (`blockOnExistingComments`), which blocks submission when a new finding lands on the same `(path, line)` as a live Qwen comment on the same commit.
|
|
244
500
|
|
|
245
501
|
**Trade-off:**
|
|
246
502
|
|
|
247
|
-
- ✅
|
|
248
|
-
- ✅
|
|
249
|
-
- ✅
|
|
250
|
-
- ❌
|
|
251
|
-
-
|
|
252
|
-
- ❌ Pattern-aggregated Suggestion findings (the multi-occurrence `Pattern:` form)
|
|
503
|
+
- ✅ Suggestion findings regain one-click ` ```suggestion ` apply and sit next to the code in "Files changed."
|
|
504
|
+
- ✅ Addressed findings self-collapse via GitHub's Outdated mechanism; no permanent extra comment on the PR page.
|
|
505
|
+
- ✅ One posting path for both severities — the `comments` array — instead of a review submission plus a second issue-comment API call.
|
|
506
|
+
- ❌ Suggestions now share the atomic `POST /pulls/{n}/reviews` call with Criticals. That call is all-or-nothing: one entry anchored to a line outside the diff 422s the whole review, so a mis-anchored Suggestion can suppress a Critical blocker. Previously Suggestions travelled on a separate, line-agnostic issue-comment call where a bad anchor was impossible. Step 7 mitigates with a 422 fallback rather than pre-validating every anchor up front: GitHub's 422 does not identify the offending entry, so the fallback has the model recheck each anchor against the diff, relocate failing Criticals into `body` (failing Suggestions are discarded — Suggestion text must stay off the `body` channel, which `qwen-autofix.yml` does not filter), and resubmit — degrading to an all-prose review of the blockers rather than posting nothing.
|
|
507
|
+
- ❌ A declined suggestion on an unchanged line can be re-posted by a later run on a new commit: the presubmit Overlap check only compares against comments whose `commit_id` matches the commit under review, so prior comments are bucketed `stale` after any push. Closing this fully needs a resolve/minimize step (GraphQL `resolveReviewThread` / `minimizeComment`) that folds our own superseded threads before submitting a new review.
|
|
508
|
+
- ❌ Pattern-aggregated Suggestion findings (the multi-occurrence `Pattern:` form) must pick a representative line to anchor to; the full structured aggregation remains visible in the terminal output.
|
|
253
509
|
|
|
254
510
|
## Rejected alternatives
|
|
255
511
|
|
|
@@ -270,21 +526,22 @@ Critical stays inline because blockers must be pinned to the exact code line and
|
|
|
270
526
|
|
|
271
527
|
For a PR with 15 findings:
|
|
272
528
|
|
|
273
|
-
| Approach | LLM calls
|
|
274
|
-
| --------------------------------------------------- |
|
|
275
|
-
| Copilot (1 agent) | 1
|
|
276
|
-
| Gemini (2 LLM tasks) | 2
|
|
277
|
-
| Our design (5 agents, N verify) | 21
|
|
278
|
-
| Our design (5 agents, batch verify, single reverse) | 7
|
|
279
|
-
| Our design (9 agents, iterative reverse) | 11-13
|
|
280
|
-
| Our design (10 agents
|
|
281
|
-
|
|
|
529
|
+
| Approach | LLM calls | Notes |
|
|
530
|
+
| --------------------------------------------------- | -------------------- | ---------------------------------------------------------------------------------------------------- |
|
|
531
|
+
| Copilot (1 agent) | 1 | Lowest cost, lowest coverage |
|
|
532
|
+
| Gemini (2 LLM tasks) | 2 | Good cost, medium coverage |
|
|
533
|
+
| Our design (5 agents, N verify) | 21 | 5+15+1 — too expensive |
|
|
534
|
+
| Our design (5 agents, batch verify, single reverse) | 7 | 5+1+1 — original design |
|
|
535
|
+
| Our design (9 agents, iterative reverse) | 11-13 | 9+1+(1-3) — +50% cost for meaningfully higher recall |
|
|
536
|
+
| Our design (10 agents) | 12-14 | 10+1+(1-3) — adds issue-fidelity/root-cause gate |
|
|
537
|
+
| Our design (12 agents + effort levels, current) | 15-21 high / 0 quick | 12(+0-2)+ceil(F/8)+(2-5) under 3A; low/medium run inline with no subagents — cost scales with intent |
|
|
538
|
+
| Claude /ultrareview | 5-20 | Cloud-hosted, cost on Anthropic |
|
|
282
539
|
|
|
283
540
|
## Future optimization: Fork Subagent
|
|
284
541
|
|
|
285
542
|
> Dependency: [Fork Subagent proposal](https://github.com/wenshao/codeagents/blob/main/docs/comparison/qwen-code-improvement-report-p0-p1-core.md#2-fork-subagentp0)
|
|
286
543
|
|
|
287
|
-
**Current problem:** Each of the
|
|
544
|
+
**Current problem:** Each of the ~15-21 LLM calls (12-14 review + sharded verify + 2-5 reverse audit rounds) creates a new subagent from scratch. At ~52K per agent (50K system + 2K task), that is ~780K-1.1M input tokens with massive redundancy. The cost grew along with the agent count — Fork Subagent matters even more under the current 12-agent design than under the original 5-agent design. (Effort levels bound the cost from the other side: low/medium runs spawn no subagents at all.)
|
|
288
545
|
|
|
289
546
|
**Fork Subagent solution:** Instead of creating independent subagents, fork the current conversation. All forks inherit the parent's full context (system prompt, conversation history, Step 1/1.1/1.5 results) and share a prompt cache prefix. The API caches the common prefix once; each fork only pays for its unique delta (~2K per agent).
|
|
290
547
|
|
|
@@ -292,13 +549,13 @@ For a PR with 15 findings:
|
|
|
292
549
|
Current (independent subagents):
|
|
293
550
|
Agent 1: [50K system] + [2K task] = 52K
|
|
294
551
|
Agent 2: [50K system] + [2K task] = 52K
|
|
295
|
-
...×
|
|
552
|
+
...× 15-21 agents = ~780K-1.1M total input tokens
|
|
296
553
|
|
|
297
554
|
With Fork + prompt cache sharing:
|
|
298
555
|
Cached prefix: [50K system + conversation history] (cached once)
|
|
299
556
|
Fork 1: [cache hit] + [2K delta] = ~2K effective
|
|
300
557
|
Fork 2: [cache hit] + [2K delta] = ~2K effective
|
|
301
|
-
...×
|
|
558
|
+
...× 15-21 forks = ~50K cached + ~30-42K delta = ~80-92K total
|
|
302
559
|
```
|
|
303
560
|
|
|
304
561
|
**Additional benefits for /review:**
|
|
@@ -308,6 +565,6 @@ With Fork + prompt cache sharing:
|
|
|
308
565
|
- Verification and reverse audit agents inherit all prior findings naturally
|
|
309
566
|
- Agent 6 personas can fork from a shared diff-loaded base, paying only the persona-framing delta
|
|
310
567
|
|
|
311
|
-
**Estimated savings:** ~
|
|
568
|
+
**Estimated savings:** ~88-92% token reduction (~780K-1.1M → ~80-92K) with zero quality impact. The savings ratio is now even more compelling than under the 5-agent design.
|
|
312
569
|
|
|
313
570
|
**Why not implemented now:** Fork Subagent requires changes to the Qwen Code core (`AgentTool`, `forkSubagent.ts`, `CacheSafeParams`). This is a platform-level feature (~400 lines, ~5 days), not a /review-specific change. When available, /review should be updated to use fork instead of independent subagents.
|