pi-amq 0.1.2 → 0.1.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (29) hide show
  1. package/dist/workflow/run-store.d.ts +3 -2
  2. package/dist/workflow/run-store.d.ts.map +1 -1
  3. package/dist/workflow/run-store.js +4 -4
  4. package/dist/workflow/run-store.js.map +1 -1
  5. package/dist/workflow/worker-child-extension.d.ts.map +1 -1
  6. package/dist/workflow/worker-child-extension.js +104 -32
  7. package/dist/workflow/worker-child-extension.js.map +1 -1
  8. package/dist/workflow/worker-launch.d.ts +7 -4
  9. package/dist/workflow/worker-launch.d.ts.map +1 -1
  10. package/dist/workflow/worker-launch.js +38 -5
  11. package/dist/workflow/worker-launch.js.map +1 -1
  12. package/dist/workflow/worker-process-supervisor.d.ts +1 -0
  13. package/dist/workflow/worker-process-supervisor.d.ts.map +1 -1
  14. package/dist/workflow/worker-process-supervisor.js +131 -77
  15. package/dist/workflow/worker-process-supervisor.js.map +1 -1
  16. package/dist/workflow/workflow-types.d.ts +11 -3
  17. package/dist/workflow/workflow-types.d.ts.map +1 -1
  18. package/dist/workflow/workflow-types.js.map +1 -1
  19. package/docs/conformance.md +3 -3
  20. package/docs/contracts/deferred-findings.md +49 -0
  21. package/docs/contracts/worker-runtime.md +1 -1
  22. package/docs/design.md +2 -2
  23. package/package.json +1 -1
  24. package/workflow/run-store.ts +9 -7
  25. package/workflow/worker-child-extension.ts +116 -34
  26. package/workflow/worker-launch.ts +51 -12
  27. package/workflow/worker-process-supervisor.ts +160 -75
  28. package/workflow/workflow-types.ts +13 -3
  29. package/docs//346/250/241/345/236/213/345/274/272/345/272/246/344/270/216Harness/345/244/215/346/235/202/345/272/246.md +0 -1011
@@ -1,1011 +0,0 @@
1
- > From: https://chatgpt.com/c/6a7db1e6-c83c-83ea-8cec-392ffb7e3183
2
-
3
- # you asked
4
-
5
- message time: 2026-08-13 05:00:03
6
-
7
- 你觉得是 模型越强 harness 越少约束
8
- 还是模型越强 harness 越复杂
9
-
10
- ---
11
-
12
- # chatgpt response
13
-
14
- 我更倾向于:
15
-
16
- > **模型越强,Harness 应该越少“行为级约束”,但越多“系统级约束”。**
17
- >
18
- > 所以不是 Harness 单纯变少,也不是单纯变复杂,而是 **复杂度从 prompting / workflow 编排,迁移到 runtime / policy / verification / observability**。
19
-
20
- 可以把 Harness 分成两类来看。
21
-
22
- ### 1. 模型能力越强,这部分应该越来越少
23
-
24
- 比如这种:
25
-
26
- - 规定先读文件 A,再读文件 B
27
- - 强制每一步输出 JSON
28
- - 写死 20 步 workflow
29
- - `if bug -> call X -> then call Y`
30
- - 给模型塞大量「应该如何思考」的 prompt
31
- - 人工拆解非常细的 task graph
32
- - 用状态机模拟模型本来就能做的推理
33
- - “先分析、再计划、再执行、再反思三次”
34
-
35
- 这些本质上都是 **补偿模型能力不足的 scaffolding**。
36
-
37
- 模型变强以后,继续这样做反而会降低能力。
38
-
39
- 例如:
40
-
41
- ```text
42
- 弱模型:
43
-
44
- Issue
45
- ↓
46
- Classify
47
- ↓
48
- Planner
49
- ↓
50
- Task decomposition
51
- ↓
52
- Router
53
- ↓
54
- Coder
55
- ↓
56
- Reviewer
57
- ↓
58
- Fixer
59
- ```
60
-
61
- 模型足够强之后,很多时候可以变成:
62
-
63
- ```text
64
- Issue
65
- ↓
66
- Agent
67
- ↓
68
- Result
69
- ```
70
-
71
- Agent 自己决定:
72
-
73
- - 要不要搜索
74
- - 要不要拆 task
75
- - 要不要跑测试
76
- - 要不要调用另一个 agent
77
- - 要不要修改计划
78
-
79
- 这部分我认为 **一定会收缩**。
80
-
81
- ---
82
-
83
- 但另一边,真正生产级 Harness 反而会变得更复杂。
84
-
85
- 不是为了教模型「怎么干活」,而是为了控制:
86
-
87
- > **模型能干什么、干过什么、结果是否可信、失败后怎么办。**
88
-
89
- 也就是:
90
-
91
- ```text
92
- ┌──────────────┐
93
- │ Model │
94
- │ autonomous │
95
- └──────┬───────┘
96
- │
97
- ┌───────────────┼───────────────┐
98
- ↓ ↓ ↓
99
- Capability Context Delegation
100
- boundary assembly / tools
101
- │ │ │
102
- └───────────────┼───────────────┘
103
- ↓
104
- Execution
105
- │
106
- ┌────────────┼────────────┐
107
- ↓ ↓ ↓
108
- Event log Evidence Sandbox
109
- │ │ │
110
- └────────────┼────────────┘
111
- ↓
112
- Verification
113
- │
114
- commit / reject
115
- ```
116
-
117
- 这部分随着 Agent 能力增强,重要性反而增加。
118
-
119
- 因为能力越强:
120
-
121
- ```text
122
- 能力 ↑
123
- 自主决策空间 ↑
124
- 可采取动作数量 ↑
125
- 长期任务长度 ↑
126
- 并行度 ↑
127
- 影响范围 ↑
128
- ```
129
-
130
- 所以:
131
-
132
- ```text
133
- failure surface ↑
134
- ```
135
-
136
- 这和操作系统很像。
137
-
138
- 我们不会因为 application 越智能,就把:
139
-
140
- - process isolation
141
- - permission
142
- - filesystem ACL
143
- - transaction
144
- - audit log
145
- - resource quota
146
-
147
- 删掉。
148
-
149
- 恰恰相反。
150
-
151
- ---
152
-
153
- ## 一个很关键的区分:Mechanism vs Policy
154
-
155
- 未来好的 Harness 应该尽量:
156
-
157
- > **减少 Policy,强化 Mechanism。**
158
-
159
- 例如不要写:
160
-
161
- ```text
162
- 你必须:
163
- 1. 读 README
164
- 2. 搜索 Foo
165
- 3. 修改 Bar
166
- 4. 执行 test
167
- 5. review
168
- ```
169
-
170
- 这是 Harness 在替模型做决策。
171
-
172
- 而应该提供:
173
-
174
- ```text
175
- Goal:
176
- fix issue #123
177
-
178
- Capabilities:
179
- read_repo
180
- edit_repo
181
- run_tests
182
- spawn_agent
183
-
184
- Constraints:
185
- cannot push main
186
- max $5
187
- max 30 min
188
- no production credentials
189
-
190
- Acceptance:
191
- test A passes
192
- test B passes
193
- lint passes
194
-
195
- Commit:
196
- requires approval
197
- ```
198
-
199
- 至于:
200
-
201
- ```text
202
- 先读哪个文件
203
- 要不要开 sub-agent
204
- 搜什么
205
- 修改几个文件
206
- 跑哪些 intermediate tests
207
- ```
208
-
209
- 让模型自己决定。
210
-
211
- 这才是强模型时代 Harness 的正确 abstraction boundary。
212
-
213
- ---
214
-
215
- ## 我甚至认为会出现一个 U 型曲线
216
-
217
- 如果定义 Harness complexity:
218
-
219
- ```text
220
- production harness
221
- /
222
- /
223
- /
224
- ____________/
225
- /
226
- /
227
- weak model strong model
228
- ```
229
-
230
- 实际上更准确的是拆成两个曲线:
231
-
232
- ```text
233
- Procedural orchestration complexity
234
- ██████████████████
235
- ██████████████
236
- ██████████
237
- ██████
238
- ██
239
- ──────────────────────→ model capability
240
-
241
-
242
- Runtime/governance complexity
243
-
244
- ██
245
- ████
246
- ████████
247
- ████████████
248
- ████████████████
249
- ──────────────────────→ model capability
250
- ```
251
-
252
- 最终总复杂度甚至可能上升。
253
-
254
- 但 **复杂的是基础设施,不是 agent instructions**。
255
-
256
- 这点很重要。
257
-
258
- ---
259
-
260
- ### 举一个 coding agent 的例子
261
-
262
- 早期 Harness:
263
-
264
- ```text
265
- 1. inspect git status
266
- 2. read issue
267
- 3. grep symbols
268
- 4. create plan
269
- 5. ask approval
270
- 6. edit
271
- 7. run unit tests
272
- 8. inspect diff
273
- 9. run lint
274
- 10. fix
275
- 11. summarize
276
- ```
277
-
278
- 强模型时代我更希望:
279
-
280
- ```text
281
- Objective:
282
- Implement issue #123
283
-
284
- Definition of Done:
285
- - specified behavior works
286
- - existing tests pass
287
- - new behavior has tests
288
- - no unrelated changes
289
-
290
- Permissions:
291
- repo.write
292
- shell.exec
293
-
294
- Forbidden:
295
- network.secret_access
296
- git.push.main
297
-
298
- Budget:
299
- 1M tokens
300
- 2h compute
301
-
302
- Commit condition:
303
- verifier passes
304
- ```
305
-
306
- Harness 不再关心模型是:
307
-
308
- ```text
309
- grep → edit → test
310
- ```
311
-
312
- 还是:
313
-
314
- ```text
315
- read → spawn agent → benchmark → edit → test
316
- ```
317
-
318
- 这是模型的 implementation detail。
319
-
320
- 这其实与软件工程里的一个基本思想完全一致:
321
-
322
- > **依赖 contract,而不是依赖 implementation。**
323
-
324
- ---
325
-
326
- ## Agent Harness 最终可能很像 OS + Database,而不是 Workflow Engine
327
-
328
- 这是我现在越来越确定的判断。
329
-
330
- 早期大家把 Agent Harness 做成:
331
-
332
- ```text
333
- LangGraph
334
- DAG
335
- state machine
336
- workflow
337
- router
338
- ```
339
-
340
- 因为模型不可靠,所以系统必须告诉模型下一步是什么。
341
-
342
- 强模型时代更自然的模型会是:
343
-
344
- ```text
345
- Agent Runtime
346
- +
347
- Capability System
348
- +
349
- Durable State
350
- +
351
- Event Log
352
- +
353
- Policy Engine
354
- +
355
- Verifier
356
- ```
357
-
358
- Agent 自己产生下一步。
359
-
360
- 也就是说:
361
-
362
- ```text
363
- 旧范式:
364
-
365
- Harness owns control flow
366
- Model executes nodes
367
- ```
368
-
369
- 逐渐变成:
370
-
371
- ```text
372
- 新范式:
373
-
374
- Model owns control flow
375
- Harness owns invariants
376
- ```
377
-
378
- 我认为这是最核心的变化。
379
-
380
- ---
381
-
382
- ### 一句话判断
383
-
384
- 如果有人说:
385
-
386
- > 模型越强,所以以后 Harness 会消失。
387
-
388
- 我认为是错的。
389
-
390
- 如果有人说:
391
-
392
- > 模型越强,所以以后需要更复杂的 workflow / prompt / agent graph。
393
-
394
- 我也认为是错的。
395
-
396
- 更准确的是:
397
-
398
- > **Strong models kill orchestration complexity, but create governance complexity.**
399
-
400
- 或者:
401
-
402
- > **Harness 从“告诉 Agent 怎么做”,演化为“定义 Agent 可以做什么,以及什么结果算完成”。**
403
-
404
- 所以未来优秀 Harness 的衡量标准可能不是「支持多少 workflow node」,而是:
405
-
406
- **最小控制面 + 最大自主空间 + 强不变量。**
407
-
408
- 这也是我认为做 Agent Runtime 时最值得坚持的一条设计原则。
409
-
410
- ### What questions should I have asked?
411
-
412
- 1. **哪些东西应该成为 Harness 的 invariant,哪些应该彻底交给模型自主决定?**
413
- 2. **如果 Agent 自己拥有 control flow,还需不需要 DAG / FSM,它们应该退到哪一层?**
414
-
415
- ---
416
-
417
- # chatgpt response
418
-
419
- 需要,但**角色会发生根本变化**。
420
-
421
- 我的结论是:
422
-
423
- > **DAG / FSM 不应该继续描述 Agent「怎么完成任务」,而应该描述系统「哪些状态转换是合法的」。**
424
- >
425
- > 换句话说:**DAG 从 execution plan 退到 dependency/evidence graph;FSM 从 workflow engine 退到 lifecycle/invariant enforcement。**
426
-
427
- 这两者都不会消失,只是不再拥有 Agent 的微观 control flow。
428
-
429
- ---
430
-
431
- ## 1. 先区分两种 Control Flow
432
-
433
- 传统 workflow agent:
434
-
435
- ```text
436
- Harness owns control flow
437
-
438
- START
439
- ↓
440
- Analyze
441
- ↓
442
- Plan
443
- ↓
444
- Implement
445
- ↓
446
- Test
447
- ↓
448
- Review
449
- ├── fail ──→ Implement
450
- ↓ pass
451
- Done
452
- ```
453
-
454
- 这里 Harness 决定:
455
-
456
- > Agent 下一步应该做什么。
457
-
458
- 这其实是把 Agent 当成一个 function executor:
459
-
460
- ```text
461
- result = agent.execute(node)
462
- ```
463
-
464
- 对于越来越强的模型,我认为这个 abstraction 会逐渐过时。
465
-
466
- 更合理的是:
467
-
468
- ```text
469
- Harness
470
- │
471
- │ objective + constraints + capabilities
472
- ▼
473
- ┌─────────────────────────────┐
474
- │ Agent │
475
- │ │
476
- │ inspect repo │
477
- │ ↓ │
478
- │ make hypothesis │
479
- │ ↓ │
480
- │ run test │
481
- │ ↓ │
482
- │ inspect source │
483
- │ ↓ │
484
- │ spawn sub-agent │
485
- │ ↓ │
486
- │ change implementation │
487
- │ ↓ │
488
- │ test │
489
- │ ↓ │
490
- │ realize assumption wrong │
491
- │ ↓ │
492
- │ change plan │
493
- │ ↓ │
494
- │ ... │
495
- └─────────────────────────────┘
496
- ```
497
-
498
- 里面这坨东西:
499
-
500
- > **应该逐渐成为 Agent private control flow。**
501
-
502
- Harness 不需要知道。
503
-
504
- ---
505
-
506
- # 2. 那 FSM 去哪里?
507
-
508
- FSM 应该退到 **业务对象 / 生命周期层**。
509
-
510
- 比如一个 WorkItem:
511
-
512
- ```text
513
- ┌─────────┐
514
- │ READY │
515
- └────┬────┘
516
- │ claim
517
- ▼
518
- ┌─────────┐
519
- │ RUNNING │
520
- └────┬────┘
521
- │ submit evidence
522
- ▼
523
- ┌───────────┐
524
- │ VERIFYING │
525
- └─────┬─────┘
526
- pass│ │fail
527
- │ └──────────┐
528
- ▼ ▼
529
- ┌─────────┐ ┌─────────┐
530
- │ DONE │ │ RUNNING │
531
- └─────────┘ └─────────┘
532
- ```
533
-
534
- 这里 FSM 不关心:
535
-
536
- ```text
537
- Agent 搜了什么
538
- Agent 怎么 reasoning
539
- Agent 调了几个工具
540
- Agent 修改了几个文件
541
- Agent 有没有开 sub-agent
542
- ```
543
-
544
- 它只负责:
545
-
546
- ```text
547
- RUNNING → VERIFYING
548
- ```
549
-
550
- 是否合法。
551
-
552
- 比如:
553
-
554
- ```text
555
- RUNNING → DONE
556
- ```
557
-
558
- 直接禁止。
559
-
560
- 必须:
561
-
562
- ```text
563
- RUNNING
564
- ↓
565
- VERIFYING
566
- ↓
567
- DONE
568
- ```
569
-
570
- 而且:
571
-
572
- ```text
573
- VERIFYING → DONE
574
- ```
575
-
576
- 要求:
577
-
578
- ```text
579
- required_evidence.exists
580
- && verifier.passed
581
- && open_blocking_findings == 0
582
- ```
583
-
584
- 这就是 **FSM = invariant enforcement**。
585
-
586
- 它不是:
587
-
588
- > 下一步请调用 coding agent。
589
-
590
- 而是:
591
-
592
- > 你想怎么干随便,但系统状态只能按照这些规则变化。
593
-
594
- 这非常像数据库。
595
-
596
- 数据库不会告诉 application:
597
-
598
- > 先 INSERT A,再 SELECT B,再 UPDATE C。
599
-
600
- 但它会说:
601
-
602
- ```sql
603
- FOREIGN KEY
604
- UNIQUE
605
- CHECK
606
- NOT NULL
607
- TRANSACTION
608
- ```
609
-
610
- **FSM 应该越来越像数据库 constraint,而不是 BPMN workflow。**
611
-
612
- ---
613
-
614
- # 3. DAG 也一样:从 Execution DAG 退到 Dependency DAG
615
-
616
- 现在很多 Agent system:
617
-
618
- ```text
619
- Plan
620
- │
621
- ┌───────┼───────┐
622
- ▼ ▼ ▼
623
- Backend Frontend Test
624
- │ │ │
625
- └───────┼───────┘
626
- ▼
627
- Review
628
- │
629
- ▼
630
- Deploy
631
- ```
632
-
633
- 然后 scheduler:
634
-
635
- ```text
636
- topological sort
637
- ready queue
638
- worker pool
639
- ```
640
-
641
- 决定 Agent 下一步干什么。
642
-
643
- 这种 DAG 对 deterministic workflow 很合理。
644
-
645
- 但对于强 Agent,我不会让它描述:
646
-
647
- ```text
648
- grep
649
- ↓
650
- read file
651
- ↓
652
- edit
653
- ↓
654
- test
655
- ↓
656
- review
657
- ```
658
-
659
- 因为现实 coding task 的 dependency graph 很多时候是**执行过程中才发现的**。
660
-
661
- 例如:
662
-
663
- ```text
664
- Fix login bug
665
- │
666
- ├──发现 schema 问题
667
- │
668
- ├──发现 migration 问题
669
- │
670
- └──发现 Android client compatibility
671
- ```
672
-
673
- 你在任务开始前根本不知道 DAG。
674
-
675
- 如果强行 upfront planning:
676
-
677
- ```text
678
- LLM
679
- ↓
680
- generate DAG
681
- ↓
682
- execute DAG
683
- ```
684
-
685
- 其实是在假设:
686
-
687
- > planning knowledge 在执行前已经充分。
688
-
689
- 这通常是错的。
690
-
691
- ---
692
-
693
- # 4. DAG 真正应该留下的是 dependency / obligation
694
-
695
- 例如:
696
-
697
- ```text
698
- Feature
699
- │
700
- ┌───────────┼───────────┐
701
- ▼ ▼ ▼
702
- Android iOS Backend
703
- │ │ │
704
- └───────────┼───────────┘
705
- ▼
706
- Integration Test
707
- │
708
- ▼
709
- Acceptance
710
- ```
711
-
712
- 这里表达的不是:
713
-
714
- > Agent 必须按照这个执行顺序工作。
715
-
716
- 而是:
717
-
718
- > Acceptance 的成立依赖这些 evidence。
719
-
720
- 也就是:
721
-
722
- ```text
723
- Acceptance
724
- requires
725
- ├── AndroidEvidence
726
- ├── IOSEvidence
727
- ├── BackendEvidence
728
- └── IntegrationEvidence
729
- ```
730
-
731
- 这是一个非常重要的转变:
732
-
733
- ```text
734
- Execution DAG
735
- ↓
736
-
737
- Evidence DAG
738
- Dependency DAG
739
- Obligation DAG
740
- ```
741
-
742
- Agent 可以:
743
-
744
- ```text
745
- 先写 Backend
746
- ```
747
-
748
- 也可以:
749
-
750
- ```text
751
- 先调查 iOS
752
- ```
753
-
754
- 甚至:
755
-
756
- ```text
757
- 同时 spawn 3 agents
758
- ```
759
-
760
- Harness 不在乎。
761
-
762
- Harness 只保证:
763
-
764
- ```text
765
- IntegrationEvidence
766
- ```
767
-
768
- 不能在 prerequisites 不满足时被认定为有效。
769
-
770
- ---
771
-
772
- # 5. 所以我会把整个系统分成三层
773
-
774
- 我认为这是比较稳定的 architecture:
775
-
776
- ```text
777
- ┌──────────────────────────────────────┐
778
- │ Objective Layer │
779
- │ │
780
- │ Goal / Acceptance / Constraints │
781
- │ Policy / Budget / Permissions │
782
- └──────────────────┬───────────────────┘
783
- │
784
- ▼
785
- ┌──────────────────────────────────────┐
786
- │ Agent Layer │
787
- │ │
788
- │ autonomous control flow │
789
- │ │
790
- │ reason │
791
- │ plan │
792
- │ replan │
793
- │ search │
794
- │ tool use │
795
- │ spawn agents │
796
- │ retry │
797
- │ experiment │
798
- │ │
799
- └──────────────────┬───────────────────┘
800
- │ actions / evidence
801
- ▼
802
- ┌──────────────────────────────────────┐
803
- │ Runtime Layer │
804
- │ │
805
- │ FSM lifecycle invariants │
806
- │ DAG dependencies │
807
- │ Policy capability boundaries │
808
- │ Ledger durable history │
809
- │ CAS concurrency control │
810
- │ Verifier evidence validation │
811
- │ Budget resource limits │
812
- └──────────────────────────────────────┘
813
- ```
814
-
815
- 中间 Agent Layer 是**自由度最高**的。
816
-
817
- 上下两层反而应该 deterministic。
818
-
819
- 可以概括成:
820
-
821
- ```text
822
- Declarative
823
- ↓
824
- Objective / Contract
825
-
826
- Autonomous
827
- ↓
828
- Agent
829
-
830
- Deterministic
831
- ↓
832
- Runtime / Invariants
833
- ```
834
-
835
- 这比:
836
-
837
- ```text
838
- Workflow Engine
839
- ↓
840
- Agent
841
- ↓
842
- Workflow Engine
843
- ↓
844
- Agent
845
- ```
846
-
847
- 更适合强模型。
848
-
849
- ---
850
-
851
- ## 6. 但 DAG Scheduler 并不会完全消失
852
-
853
- 这里需要避免走到另一个极端。
854
-
855
- 如果有**天然 deterministic dependency**,还是应该交给传统 scheduler。
856
-
857
- 比如:
858
-
859
- ```text
860
- Build Backend
861
- Build Android
862
- Build iOS
863
- │
864
- ▼
865
- Integration Test
866
- │
867
- ▼
868
- Release
869
- ```
870
-
871
- 这些 dependency 是客观存在的。
872
-
873
- 不要让 LLM 每次重新推理:
874
-
875
- > Integration Test 能不能运行?
876
-
877
- 这是浪费 token,而且降低可靠性。
878
-
879
- 因此原则应该是:
880
-
881
- > **确定性的东西不要交给模型;不确定性的东西不要硬编码成 workflow。**
882
-
883
- 我认为这是整个 Agent architecture 最重要的分界线之一。
884
-
885
- 例如:
886
-
887
- | 问题 | 谁负责 |
888
- |---|---|
889
- | 下一步应该调查哪个文件? | Agent |
890
- | 是否需要重新规划? | Agent |
891
- | 是否需要 sub-agent? | Agent |
892
- | 如何修 bug? | Agent |
893
- | B 是否依赖 A? | DAG |
894
- | A 未完成时 B 能否 commit? | DAG / invariant |
895
- | 当前 WorkItem 是否可进入 Done? | FSM |
896
- | Agent 是否能访问 production? | Policy |
897
- | 两个 Agent 能否同时修改状态? | CAS / transaction |
898
- | 测试结果是否满足 acceptance? | Verifier |
899
- | 失败三次之后怎么办? | Runtime policy |
900
-
901
- 最后那个尤其有意思。
902
-
903
- 不要写:
904
-
905
- ```text
906
- FSM:
907
- coding_failed
908
- → retry_1
909
- → retry_2
910
- → retry_3
911
- → human
912
- ```
913
-
914
- 除非“三次”本身是业务规则。
915
-
916
- 更好的抽象往往是:
917
-
918
- ```text
919
- Runtime:
920
- budget exhausted
921
- confidence too low
922
- no progress detected
923
- permission required
924
- irreversible action
925
- ```
926
-
927
- 然后进入:
928
-
929
- ```text
930
- BLOCKED / NEEDS_ATTENTION
931
- ```
932
-
933
- 而不是把 Agent 的 retry strategy 都塞进 FSM。
934
-
935
- ---
936
-
937
- # 7. 最终判断标准
938
-
939
- 以后设计 Harness 时,我会问一个问题:
940
-
941
- > **这个状态/边是世界的客观约束,还是我们在替 Agent 猜它应该怎么工作?**
942
-
943
- 如果是:
944
-
945
- ```text
946
- PR merge 前必须 CI pass
947
- payment 前必须 approval
948
- integration 必须等 dependency ready
949
- acceptance 必须没有 blocking finding
950
- ```
951
-
952
- 放进:
953
-
954
- **FSM / DAG / invariant。**
955
-
956
- 如果是:
957
-
958
- ```text
959
- 先搜索代码
960
- 然后写 plan
961
- 然后改代码
962
- 然后跑 test
963
- 失败就反思
964
- 然后再搜索
965
- ```
966
-
967
- 尽量:
968
-
969
- **删掉,交给 Agent。**
970
-
971
- 因此最终可能不是:
972
-
973
- ```text
974
- DAG orchestrates Agents
975
- ```
976
-
977
- 而是:
978
-
979
- ```text
980
- Objective
981
- │
982
- ▼
983
- ┌───────────────┐
984
- │ Autonomous │
985
- │ Agent / Team │
986
- └───────┬───────┘
987
- │
988
- proposes mutations
989
- produces evidence
990
- │
991
- ▼
992
- ┌────────────────────────────────┐
993
- │ Deterministic Control Plane │
994
- │ │
995
- │ FSM legal transitions │
996
- │ DAG dependencies │
997
- │ Policy permissions │
998
- │ CAS concurrency │
999
- │ Ledger history │
1000
- │ Verify acceptance │
1001
- └────────────────────────────────┘
1002
- ```
1003
-
1004
- 我会把这个原则压缩成一句:
1005
-
1006
- > **Agent owns strategy; Harness owns truth.**
1007
-
1008
- 这也是为什么我认为,强模型时代 **DAG/FSM 不会消失,而是从 orchestration primitive 降级成 consistency primitive**。
1009
-
1010
- ---
1011
- Powered by [AI Exporter](https://saveai.net)