skill-systems-thinking 0.1.0 → 0.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +37 -25
- package/LICENSE +21 -21
- package/README.md +482 -396
- package/SKILL.md +125 -125
- package/assets/runtime-prompt.txt +21 -21
- package/evals/checks.json +88 -88
- package/init.ps1 +68 -0
- package/package.json +11 -4
- package/references/bounded-control.md +82 -82
- package/references/closed-loop-workflow.md +89 -89
- package/references/discrete-systems.md +79 -79
- package/references/disturbance.md +98 -98
- package/references/modeling.md +87 -87
- package/references/multivariable.md +112 -112
- package/references/original-text.md +58 -58
- package/references/stability.md +97 -97
- package/references/state-and-control.md +84 -84
- package/templates/change-proposal.md +51 -51
- package/templates/debugging-checklist.md +44 -44
|
@@ -1,58 +1,58 @@
|
|
|
1
|
-
# Reference: Original Text — Engineering Cybernetics
|
|
2
|
-
|
|
3
|
-
Direct excerpts from Qian Xuesen (钱学森) & Song Jian (宋健), *Engineering Cybernetics*
|
|
4
|
-
(工程控制论). Used here under fair use for scholarship and skill development. The original
|
|
5
|
-
book and its mathematical foundations are the authoritative source; this skill is a faithful
|
|
6
|
-
interpretation, not a replacement.
|
|
7
|
-
|
|
8
|
-
---
|
|
9
|
-
|
|
10
|
-
## §1.1 稳定性的基本要求(节选)
|
|
11
|
-
|
|
12
|
-
> 对于控制工程来说,稳定性是第一个要求。一个系统即使能在某些工作点给出正确
|
|
13
|
-
> 的输出,但在扰动下发散、振荡或泄漏,这个系统就是**不稳定**的——无论其设计
|
|
14
|
-
> 多么巧妙,都是没有实用价值的。
|
|
15
|
-
|
|
16
|
-
This is the anchor for the law "Stability before everything." An LLM must verify
|
|
17
|
-
perturbation behavior before proposing any feature or optimization.
|
|
18
|
-
|
|
19
|
-
---
|
|
20
|
-
|
|
21
|
-
## §1.5 控制量与受控量(节选)
|
|
22
|
-
|
|
23
|
-
> 受控量 x 是必须保持正确的量;控制量 u 是我们有自由施加作用的量。线性
|
|
24
|
-
> 调节理论假定变量无界——但物理系统都有界。忽略边界会使"过渡过程时间与
|
|
25
|
-
> 步长无关"这一结论在实践中不成立。
|
|
26
|
-
|
|
27
|
-
This is the anchor for the principle "Name x before u; respect bounds." Every actuator
|
|
28
|
-
saturates; every loop must have an explicit bound.
|
|
29
|
-
|
|
30
|
-
---
|
|
31
|
-
|
|
32
|
-
## §3.7 反馈的基本作用(节选)
|
|
33
|
-
|
|
34
|
-
> 反馈同时提高了准确度和响应速度。开环系统没有从误差中恢复的路径;闭环
|
|
35
|
-
> 系统通过"作用→观测→校正"实现收敛。
|
|
36
|
-
|
|
37
|
-
This is the anchor for "Closed loop beats open loop." After every non-trivial change,
|
|
38
|
-
observe the actual state before declaring success.
|
|
39
|
-
|
|
40
|
-
---
|
|
41
|
-
|
|
42
|
-
## §6.3 多变量控制的解耦(节选)
|
|
43
|
-
|
|
44
|
-
> 当两个变量之间不应相互干扰时,应当使交叉项为零——隔离、单一职责、纯函数、
|
|
45
|
-
> 无共享可变状态。一旦对角化,每个子系统就可以作为独立的单变量系统设计,
|
|
46
|
-
> 从而大幅简化问题。
|
|
47
|
-
|
|
48
|
-
This is the anchor for "Decouple what hurts, coordinate what helps."
|
|
49
|
-
|
|
50
|
-
---
|
|
51
|
-
|
|
52
|
-
## §11 时滞系统(节选)
|
|
53
|
-
|
|
54
|
-
> 时滞是稳定性的敌人。在回路延迟很大时,提高增益往往不是解决滞后响应的
|
|
55
|
-
> 方法——它会 pushing the system toward instability。应当降低回路增益,或添加
|
|
56
|
-
> 阻尼,或缩短回路。
|
|
57
|
-
|
|
58
|
-
This is the anchor for "Never crank gain on a delayed loop; add damping instead."
|
|
1
|
+
# Reference: Original Text — Engineering Cybernetics
|
|
2
|
+
|
|
3
|
+
Direct excerpts from Qian Xuesen (钱学森) & Song Jian (宋健), *Engineering Cybernetics*
|
|
4
|
+
(工程控制论). Used here under fair use for scholarship and skill development. The original
|
|
5
|
+
book and its mathematical foundations are the authoritative source; this skill is a faithful
|
|
6
|
+
interpretation, not a replacement.
|
|
7
|
+
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
## §1.1 稳定性的基本要求(节选)
|
|
11
|
+
|
|
12
|
+
> 对于控制工程来说,稳定性是第一个要求。一个系统即使能在某些工作点给出正确
|
|
13
|
+
> 的输出,但在扰动下发散、振荡或泄漏,这个系统就是**不稳定**的——无论其设计
|
|
14
|
+
> 多么巧妙,都是没有实用价值的。
|
|
15
|
+
|
|
16
|
+
This is the anchor for the law "Stability before everything." An LLM must verify
|
|
17
|
+
perturbation behavior before proposing any feature or optimization.
|
|
18
|
+
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
## §1.5 控制量与受控量(节选)
|
|
22
|
+
|
|
23
|
+
> 受控量 x 是必须保持正确的量;控制量 u 是我们有自由施加作用的量。线性
|
|
24
|
+
> 调节理论假定变量无界——但物理系统都有界。忽略边界会使"过渡过程时间与
|
|
25
|
+
> 步长无关"这一结论在实践中不成立。
|
|
26
|
+
|
|
27
|
+
This is the anchor for the principle "Name x before u; respect bounds." Every actuator
|
|
28
|
+
saturates; every loop must have an explicit bound.
|
|
29
|
+
|
|
30
|
+
---
|
|
31
|
+
|
|
32
|
+
## §3.7 反馈的基本作用(节选)
|
|
33
|
+
|
|
34
|
+
> 反馈同时提高了准确度和响应速度。开环系统没有从误差中恢复的路径;闭环
|
|
35
|
+
> 系统通过"作用→观测→校正"实现收敛。
|
|
36
|
+
|
|
37
|
+
This is the anchor for "Closed loop beats open loop." After every non-trivial change,
|
|
38
|
+
observe the actual state before declaring success.
|
|
39
|
+
|
|
40
|
+
---
|
|
41
|
+
|
|
42
|
+
## §6.3 多变量控制的解耦(节选)
|
|
43
|
+
|
|
44
|
+
> 当两个变量之间不应相互干扰时,应当使交叉项为零——隔离、单一职责、纯函数、
|
|
45
|
+
> 无共享可变状态。一旦对角化,每个子系统就可以作为独立的单变量系统设计,
|
|
46
|
+
> 从而大幅简化问题。
|
|
47
|
+
|
|
48
|
+
This is the anchor for "Decouple what hurts, coordinate what helps."
|
|
49
|
+
|
|
50
|
+
---
|
|
51
|
+
|
|
52
|
+
## §11 时滞系统(节选)
|
|
53
|
+
|
|
54
|
+
> 时滞是稳定性的敌人。在回路延迟很大时,提高增益往往不是解决滞后响应的
|
|
55
|
+
> 方法——它会 pushing the system toward instability。应当降低回路增益,或添加
|
|
56
|
+
> 阻尼,或缩短回路。
|
|
57
|
+
|
|
58
|
+
This is the anchor for "Never crank gain on a delayed loop; add damping instead."
|
package/references/stability.md
CHANGED
|
@@ -1,97 +1,97 @@
|
|
|
1
|
-
# Reference: Stability
|
|
2
|
-
|
|
3
|
-
Book anchors: §1.1 (constant-coefficient linear systems), §4.1 (Lyapunov stability definition
|
|
4
|
-
and direct method), §1.3 (nonlinear systems have multiple equilibria).
|
|
5
|
-
|
|
6
|
-
## Stability is the first requirement
|
|
7
|
-
|
|
8
|
-
(§1.1) For a control system, stability is demanded *before* anything else. A program that
|
|
9
|
-
gives the right answer *sometimes* but diverges, hangs, leaks, or corrupts state under
|
|
10
|
-
perturbation is **unstable** — and therefore useless, however clever.
|
|
11
|
-
|
|
12
|
-
### Lyapunov definition (§4.1), in plain terms
|
|
13
|
-
|
|
14
|
-
A motion is *stable* if: for any tolerated deviation ε > 0, there exists a δ > 0 such that,
|
|
15
|
-
whenever the initial state starts within δ of the target, the trajectory stays within ε of the
|
|
16
|
-
target for all t ≥ t₀. *Asymptotically stable* if it also converges back to the target.
|
|
17
|
-
|
|
18
|
-
- Translate: **small perturbations must produce bounded deviation, not runaway.**
|
|
19
|
-
- "For all t" matters. A system that looks fine for 5 minutes then explodes is not stable.
|
|
20
|
-
|
|
21
|
-
## LLM behavior pattern: Stability before everything
|
|
22
|
-
|
|
23
|
-
> "A system that diverges under perturbation is wrong, however clever the optimization."
|
|
24
|
-
|
|
25
|
-
The LLM using this skill **must** verify stability before proposing any feature or
|
|
26
|
-
optimization. This is not a suggestion; it is the second cognitive law.
|
|
27
|
-
|
|
28
|
-
- Before generating performance optimizations, the LLM must reason: "Is the base system
|
|
29
|
-
stable under perturbation? If not, I reject the optimization and fix stability first."
|
|
30
|
-
- Perturbations to consider: empty input, huge input, null, concurrent calls, network flake,
|
|
31
|
-
dependency down, clock skew.
|
|
32
|
-
|
|
33
|
-
### Test for stability, not just correctness
|
|
34
|
-
|
|
35
|
-
- Probe with **perturbations**, not only golden paths: empty input, huge input, null,
|
|
36
|
-
concurrent calls, network flake, one dependency down, clock skew.
|
|
37
|
-
- If deviation **grows without bound** under a small kick → unstable. Fix before features.
|
|
38
|
-
- Watch the **signatures of instability**:
|
|
39
|
-
- unbounded growth → memory/CPU/disk creep, state explosion;
|
|
40
|
-
- oscillation → retry storms, ping-pong state between two components;
|
|
41
|
-
- limit cycle → two routines mutually undoing each other every tick.
|
|
42
|
-
|
|
43
|
-
## Local vs global stability (§1.3, §4.1)
|
|
44
|
-
|
|
45
|
-
- Linear systems are all-or-nothing: either every motion is stable or none is. So "is the
|
|
46
|
-
system stable?" is a meaningful question *only* for linear systems.
|
|
47
|
-
- Nonlinear systems (and almost everything real, §1.4) can be stable in one region and
|
|
48
|
-
unstable in another — each equilibrium has its own stability. "Works on my machine / small
|
|
49
|
-
data" is **local stability only**. Probe the edges of the operating envelope.
|
|
50
|
-
|
|
51
|
-
## LLM behavior pattern: Reject unstable trajectories
|
|
52
|
-
|
|
53
|
-
> "If you observe a deviation that grows over iterations, you have lost stability. Stop. Do not paper over it with a 'slightly better' patch. Return to step 1 (Identify) and re-analyze the plant."
|
|
54
|
-
|
|
55
|
-
An LLM must not declare success when a trajectory is unstable. Growing deviation is a
|
|
56
|
-
hard-stop signal, not a minor concern.
|
|
57
|
-
|
|
58
|
-
## Worked example: retry storm → damped
|
|
59
|
-
|
|
60
|
-
**Before (unstable):**
|
|
61
|
-
```python
|
|
62
|
-
def fetch(url):
|
|
63
|
-
for attempt in range(999): # unbounded loop
|
|
64
|
-
try:
|
|
65
|
-
return requests.get(url, timeout=3)
|
|
66
|
-
except requests.Timeout:
|
|
67
|
-
continue # high gain, zero damping
|
|
68
|
-
```
|
|
69
|
-
No backoff, no circuit breaker, no jitter. Under load this oscillates: concurrent
|
|
70
|
-
retries amplify the load spike → each times out → more retries. Trajectory diverges.
|
|
71
|
-
|
|
72
|
-
**After (stable):**
|
|
73
|
-
```python
|
|
74
|
-
def fetch(url, max_retries=3):
|
|
75
|
-
for attempt in range(max_retries):
|
|
76
|
-
try:
|
|
77
|
-
return requests.get(url, timeout=3)
|
|
78
|
-
except requests.Timeout:
|
|
79
|
-
delay = (2 ** attempt) + random.uniform(0, 0.1) # exponential + jitter
|
|
80
|
-
time.sleep(delay) # damping
|
|
81
|
-
raise CircuitBreakerOpen() # open loop before saturation
|
|
82
|
-
```
|
|
83
|
-
|
|
84
|
-
- Bound on actuator (`max_retries`).
|
|
85
|
-
- Backoff = damping; jitter = break synchronization.
|
|
86
|
-
- Circuit breaker = open loop when service is unhealthy.
|
|
87
|
-
|
|
88
|
-
---
|
|
89
|
-
|
|
90
|
-
## Stabilizing a system
|
|
91
|
-
|
|
92
|
-
- Linear feedback (§5): `u = −Kx` pulls state back toward target. But simple feedback is not
|
|
93
|
-
always enough (§5, §9.6): near actuator saturation or with bad phase margin, you need
|
|
94
|
-
damping / lead correction, not just gain.
|
|
95
|
-
- Adding gain to "fix" a lagging or oscillating system often *destabilizes* it (§11): reduce
|
|
96
|
-
loop gain or add damping instead.
|
|
97
|
-
- Keep stability margins explicit: how much worse can the input/get worse before it diverges?
|
|
1
|
+
# Reference: Stability
|
|
2
|
+
|
|
3
|
+
Book anchors: §1.1 (constant-coefficient linear systems), §4.1 (Lyapunov stability definition
|
|
4
|
+
and direct method), §1.3 (nonlinear systems have multiple equilibria).
|
|
5
|
+
|
|
6
|
+
## Stability is the first requirement
|
|
7
|
+
|
|
8
|
+
(§1.1) For a control system, stability is demanded *before* anything else. A program that
|
|
9
|
+
gives the right answer *sometimes* but diverges, hangs, leaks, or corrupts state under
|
|
10
|
+
perturbation is **unstable** — and therefore useless, however clever.
|
|
11
|
+
|
|
12
|
+
### Lyapunov definition (§4.1), in plain terms
|
|
13
|
+
|
|
14
|
+
A motion is *stable* if: for any tolerated deviation ε > 0, there exists a δ > 0 such that,
|
|
15
|
+
whenever the initial state starts within δ of the target, the trajectory stays within ε of the
|
|
16
|
+
target for all t ≥ t₀. *Asymptotically stable* if it also converges back to the target.
|
|
17
|
+
|
|
18
|
+
- Translate: **small perturbations must produce bounded deviation, not runaway.**
|
|
19
|
+
- "For all t" matters. A system that looks fine for 5 minutes then explodes is not stable.
|
|
20
|
+
|
|
21
|
+
## LLM behavior pattern: Stability before everything
|
|
22
|
+
|
|
23
|
+
> "A system that diverges under perturbation is wrong, however clever the optimization."
|
|
24
|
+
|
|
25
|
+
The LLM using this skill **must** verify stability before proposing any feature or
|
|
26
|
+
optimization. This is not a suggestion; it is the second cognitive law.
|
|
27
|
+
|
|
28
|
+
- Before generating performance optimizations, the LLM must reason: "Is the base system
|
|
29
|
+
stable under perturbation? If not, I reject the optimization and fix stability first."
|
|
30
|
+
- Perturbations to consider: empty input, huge input, null, concurrent calls, network flake,
|
|
31
|
+
dependency down, clock skew.
|
|
32
|
+
|
|
33
|
+
### Test for stability, not just correctness
|
|
34
|
+
|
|
35
|
+
- Probe with **perturbations**, not only golden paths: empty input, huge input, null,
|
|
36
|
+
concurrent calls, network flake, one dependency down, clock skew.
|
|
37
|
+
- If deviation **grows without bound** under a small kick → unstable. Fix before features.
|
|
38
|
+
- Watch the **signatures of instability**:
|
|
39
|
+
- unbounded growth → memory/CPU/disk creep, state explosion;
|
|
40
|
+
- oscillation → retry storms, ping-pong state between two components;
|
|
41
|
+
- limit cycle → two routines mutually undoing each other every tick.
|
|
42
|
+
|
|
43
|
+
## Local vs global stability (§1.3, §4.1)
|
|
44
|
+
|
|
45
|
+
- Linear systems are all-or-nothing: either every motion is stable or none is. So "is the
|
|
46
|
+
system stable?" is a meaningful question *only* for linear systems.
|
|
47
|
+
- Nonlinear systems (and almost everything real, §1.4) can be stable in one region and
|
|
48
|
+
unstable in another — each equilibrium has its own stability. "Works on my machine / small
|
|
49
|
+
data" is **local stability only**. Probe the edges of the operating envelope.
|
|
50
|
+
|
|
51
|
+
## LLM behavior pattern: Reject unstable trajectories
|
|
52
|
+
|
|
53
|
+
> "If you observe a deviation that grows over iterations, you have lost stability. Stop. Do not paper over it with a 'slightly better' patch. Return to step 1 (Identify) and re-analyze the plant."
|
|
54
|
+
|
|
55
|
+
An LLM must not declare success when a trajectory is unstable. Growing deviation is a
|
|
56
|
+
hard-stop signal, not a minor concern.
|
|
57
|
+
|
|
58
|
+
## Worked example: retry storm → damped
|
|
59
|
+
|
|
60
|
+
**Before (unstable):**
|
|
61
|
+
```python
|
|
62
|
+
def fetch(url):
|
|
63
|
+
for attempt in range(999): # unbounded loop
|
|
64
|
+
try:
|
|
65
|
+
return requests.get(url, timeout=3)
|
|
66
|
+
except requests.Timeout:
|
|
67
|
+
continue # high gain, zero damping
|
|
68
|
+
```
|
|
69
|
+
No backoff, no circuit breaker, no jitter. Under load this oscillates: concurrent
|
|
70
|
+
retries amplify the load spike → each times out → more retries. Trajectory diverges.
|
|
71
|
+
|
|
72
|
+
**After (stable):**
|
|
73
|
+
```python
|
|
74
|
+
def fetch(url, max_retries=3):
|
|
75
|
+
for attempt in range(max_retries):
|
|
76
|
+
try:
|
|
77
|
+
return requests.get(url, timeout=3)
|
|
78
|
+
except requests.Timeout:
|
|
79
|
+
delay = (2 ** attempt) + random.uniform(0, 0.1) # exponential + jitter
|
|
80
|
+
time.sleep(delay) # damping
|
|
81
|
+
raise CircuitBreakerOpen() # open loop before saturation
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
- Bound on actuator (`max_retries`).
|
|
85
|
+
- Backoff = damping; jitter = break synchronization.
|
|
86
|
+
- Circuit breaker = open loop when service is unhealthy.
|
|
87
|
+
|
|
88
|
+
---
|
|
89
|
+
|
|
90
|
+
## Stabilizing a system
|
|
91
|
+
|
|
92
|
+
- Linear feedback (§5): `u = −Kx` pulls state back toward target. But simple feedback is not
|
|
93
|
+
always enough (§5, §9.6): near actuator saturation or with bad phase margin, you need
|
|
94
|
+
damping / lead correction, not just gain.
|
|
95
|
+
- Adding gain to "fix" a lagging or oscillating system often *destabilizes* it (§11): reduce
|
|
96
|
+
loop gain or add damping instead.
|
|
97
|
+
- Keep stability margins explicit: how much worse can the input/get worse before it diverges?
|
|
@@ -1,84 +1,84 @@
|
|
|
1
|
-
# Reference: State, Control, Observability & Bounds
|
|
2
|
-
|
|
3
|
-
Book anchor: §1.5 (controlled vs control quantities; constraints on both).
|
|
4
|
-
|
|
5
|
-
## Two kinds of variables
|
|
6
|
-
|
|
7
|
-
In every system you touch, separate:
|
|
8
|
-
|
|
9
|
-
- **Controlled variables x** — the things that must be correct: output values, internal
|
|
10
|
-
state, invariants, latencies, error rates. You cannot set them directly; you can only
|
|
11
|
-
*influence* them through the equations of motion.
|
|
12
|
-
- **Control variables u** — the levers you actually manipulate: function arguments, config
|
|
13
|
-
values, the code you write, the inputs you feed. Only on u do you have freedom to act.
|
|
14
|
-
|
|
15
|
-
## LLM behavior pattern: Name x before u
|
|
16
|
-
|
|
17
|
-
> "Name the controlled variable x before proposing any change. Unnamed targets are unregulated."
|
|
18
|
-
|
|
19
|
-
An LLM using this skill **must** name the controlled variable and the control variable before
|
|
20
|
-
any proposal. This is not optional framing—it is the first cognitive law. If you cannot name
|
|
21
|
-
x, you cannot verify success; you are operating open-loop.
|
|
22
|
-
|
|
23
|
-
- "Make it work" is not a control problem; "reduce error rate below 1%" is.
|
|
24
|
-
- Before generating a patch, the LLM must internally resolve: what is x, what is u, and how
|
|
25
|
-
will I measure x after u moves?
|
|
26
|
-
|
|
27
|
-
## LLM behavior pattern: Demand observability before acting
|
|
28
|
-
|
|
29
|
-
> "A failure mode you cannot observe will drift until it bites. Add the sensor first."
|
|
30
|
-
|
|
31
|
-
If the LLM identifies a plausible failure mode that is invisible to existing tests/logs, it
|
|
32
|
-
must **propose instrumentation before proposing the fix**. Acting without observability is
|
|
33
|
-
guessing; guessing is not control theory.
|
|
34
|
-
|
|
35
|
-
- If observability is impossible in the current environment, state that explicitly as a risk.
|
|
36
|
-
- Prefer mechanisms that *measure* outcomes over mechanisms that *assert* they happened.
|
|
37
|
-
|
|
38
|
-
## Bounds are physical, not optional (§1.5)
|
|
39
|
-
|
|
40
|
-
Linear regulation theory silently assumes variables are unbounded. Reality is not:
|
|
41
|
-
|
|
42
|
-
- Control quantities are bounded: thrust limits, voltage/power limits, **API rate limits**,
|
|
43
|
-
**context-window size**, **stack depth**.
|
|
44
|
-
- Controlled quantities are bounded: integer ranges, buffer sizes, **memory**, **deadlines**.
|
|
45
|
-
|
|
46
|
-
Near a bound, linear intuition fails (saturation, overflow, OOM, token cutoff). Design for the
|
|
47
|
-
bound, not the asymptote. The book notes that ignoring limits makes "transition time
|
|
48
|
-
independent of step size" — which is simply false in practice. Same trap: assuming a function
|
|
49
|
-
behaves the same at n=10 and n=10^7.
|
|
50
|
-
|
|
51
|
-
## Observability vs controllability (§5.5)
|
|
52
|
-
|
|
53
|
-
- A state you can measure → you can close a loop on it (observable).
|
|
54
|
-
- A quantity you can change → you can steer it (controllable).
|
|
55
|
-
- Some states are neither easily measured nor directly changed — they must be *estimated*
|
|
56
|
-
from observables. When debugging "impossible" faults, suspect a state that is observable
|
|
57
|
-
only indirectly. Build the estimator (instrumentation) before concluding the bug is
|
|
58
|
-
untraceable.
|
|
59
|
-
|
|
60
|
-
## Worked example: unregulated target → open-loop patch
|
|
61
|
-
|
|
62
|
-
**Before (no x, no u, no observability):**
|
|
63
|
-
```python
|
|
64
|
-
# "Make it better" — what is x? what is u? how will we know?
|
|
65
|
-
def process(data):
|
|
66
|
-
return data # placeholder; "fix" the symptom later
|
|
67
|
-
```
|
|
68
|
-
Without naming x and u, every subsequent change is a guess.
|
|
69
|
-
|
|
70
|
-
**After (closed-loop from the start):**
|
|
71
|
-
```python
|
|
72
|
-
def process(data, max_size=1000):
|
|
73
|
-
# x = queue depth / error rate; u = max_size
|
|
74
|
-
# observability: raise if exceeded, count in tests
|
|
75
|
-
if len(data) > max_size:
|
|
76
|
-
raise ValueError(f"x exceeded bound: {len(data)} > {max_size}")
|
|
77
|
-
return data
|
|
78
|
-
```
|
|
79
|
-
|
|
80
|
-
- x = `len(data)` (controlled variable, must stay ≤ `max_size`)
|
|
81
|
-
- u = `max_size` (the lever we adjust)
|
|
82
|
-
- observability = exception raised, testable
|
|
83
|
-
- bound = `max_size` (actuator saturation respected)
|
|
84
|
-
|
|
1
|
+
# Reference: State, Control, Observability & Bounds
|
|
2
|
+
|
|
3
|
+
Book anchor: §1.5 (controlled vs control quantities; constraints on both).
|
|
4
|
+
|
|
5
|
+
## Two kinds of variables
|
|
6
|
+
|
|
7
|
+
In every system you touch, separate:
|
|
8
|
+
|
|
9
|
+
- **Controlled variables x** — the things that must be correct: output values, internal
|
|
10
|
+
state, invariants, latencies, error rates. You cannot set them directly; you can only
|
|
11
|
+
*influence* them through the equations of motion.
|
|
12
|
+
- **Control variables u** — the levers you actually manipulate: function arguments, config
|
|
13
|
+
values, the code you write, the inputs you feed. Only on u do you have freedom to act.
|
|
14
|
+
|
|
15
|
+
## LLM behavior pattern: Name x before u
|
|
16
|
+
|
|
17
|
+
> "Name the controlled variable x before proposing any change. Unnamed targets are unregulated."
|
|
18
|
+
|
|
19
|
+
An LLM using this skill **must** name the controlled variable and the control variable before
|
|
20
|
+
any proposal. This is not optional framing—it is the first cognitive law. If you cannot name
|
|
21
|
+
x, you cannot verify success; you are operating open-loop.
|
|
22
|
+
|
|
23
|
+
- "Make it work" is not a control problem; "reduce error rate below 1%" is.
|
|
24
|
+
- Before generating a patch, the LLM must internally resolve: what is x, what is u, and how
|
|
25
|
+
will I measure x after u moves?
|
|
26
|
+
|
|
27
|
+
## LLM behavior pattern: Demand observability before acting
|
|
28
|
+
|
|
29
|
+
> "A failure mode you cannot observe will drift until it bites. Add the sensor first."
|
|
30
|
+
|
|
31
|
+
If the LLM identifies a plausible failure mode that is invisible to existing tests/logs, it
|
|
32
|
+
must **propose instrumentation before proposing the fix**. Acting without observability is
|
|
33
|
+
guessing; guessing is not control theory.
|
|
34
|
+
|
|
35
|
+
- If observability is impossible in the current environment, state that explicitly as a risk.
|
|
36
|
+
- Prefer mechanisms that *measure* outcomes over mechanisms that *assert* they happened.
|
|
37
|
+
|
|
38
|
+
## Bounds are physical, not optional (§1.5)
|
|
39
|
+
|
|
40
|
+
Linear regulation theory silently assumes variables are unbounded. Reality is not:
|
|
41
|
+
|
|
42
|
+
- Control quantities are bounded: thrust limits, voltage/power limits, **API rate limits**,
|
|
43
|
+
**context-window size**, **stack depth**.
|
|
44
|
+
- Controlled quantities are bounded: integer ranges, buffer sizes, **memory**, **deadlines**.
|
|
45
|
+
|
|
46
|
+
Near a bound, linear intuition fails (saturation, overflow, OOM, token cutoff). Design for the
|
|
47
|
+
bound, not the asymptote. The book notes that ignoring limits makes "transition time
|
|
48
|
+
independent of step size" — which is simply false in practice. Same trap: assuming a function
|
|
49
|
+
behaves the same at n=10 and n=10^7.
|
|
50
|
+
|
|
51
|
+
## Observability vs controllability (§5.5)
|
|
52
|
+
|
|
53
|
+
- A state you can measure → you can close a loop on it (observable).
|
|
54
|
+
- A quantity you can change → you can steer it (controllable).
|
|
55
|
+
- Some states are neither easily measured nor directly changed — they must be *estimated*
|
|
56
|
+
from observables. When debugging "impossible" faults, suspect a state that is observable
|
|
57
|
+
only indirectly. Build the estimator (instrumentation) before concluding the bug is
|
|
58
|
+
untraceable.
|
|
59
|
+
|
|
60
|
+
## Worked example: unregulated target → open-loop patch
|
|
61
|
+
|
|
62
|
+
**Before (no x, no u, no observability):**
|
|
63
|
+
```python
|
|
64
|
+
# "Make it better" — what is x? what is u? how will we know?
|
|
65
|
+
def process(data):
|
|
66
|
+
return data # placeholder; "fix" the symptom later
|
|
67
|
+
```
|
|
68
|
+
Without naming x and u, every subsequent change is a guess.
|
|
69
|
+
|
|
70
|
+
**After (closed-loop from the start):**
|
|
71
|
+
```python
|
|
72
|
+
def process(data, max_size=1000):
|
|
73
|
+
# x = queue depth / error rate; u = max_size
|
|
74
|
+
# observability: raise if exceeded, count in tests
|
|
75
|
+
if len(data) > max_size:
|
|
76
|
+
raise ValueError(f"x exceeded bound: {len(data)} > {max_size}")
|
|
77
|
+
return data
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
- x = `len(data)` (controlled variable, must stay ≤ `max_size`)
|
|
81
|
+
- u = `max_size` (the lever we adjust)
|
|
82
|
+
- observability = exception raised, testable
|
|
83
|
+
- bound = `max_size` (actuator saturation respected)
|
|
84
|
+
|
|
@@ -1,51 +1,51 @@
|
|
|
1
|
-
# Change Proposal (control-systems thinking scaffold)
|
|
2
|
-
|
|
3
|
-
Fill this in before a non-trivial change. It forces the control-theoretic mindset: name the
|
|
4
|
-
plant, the setpoint, the actuator, and how you will know it converged.
|
|
5
|
-
|
|
6
|
-
## Target
|
|
7
|
-
- **Controlled variable x** (what must stay correct):
|
|
8
|
-
_________________________________
|
|
9
|
-
- **Setpoint x*** (target / invariant / acceptable range):
|
|
10
|
-
_________________________________
|
|
11
|
-
|
|
12
|
-
## Plant & model
|
|
13
|
-
- **Current behavior** (measured, not assumed):
|
|
14
|
-
_________________________________
|
|
15
|
-
- **Mental model of the plant** (and where it may be wrong):
|
|
16
|
-
_________________________________
|
|
17
|
-
- **Model validity range** (inputs/sizes/conditions where the model holds):
|
|
18
|
-
_________________________________
|
|
19
|
-
|
|
20
|
-
## Control action
|
|
21
|
-
- **Control variable u** (the lever I will change):
|
|
22
|
-
_________________________________
|
|
23
|
-
- **Bounds on u** (max retries / depth / rate / size):
|
|
24
|
-
_________________________________
|
|
25
|
-
- **Sensors** (how I will observe x after the change — tests/logs/metrics):
|
|
26
|
-
_________________________________
|
|
27
|
-
|
|
28
|
-
## Disturbances
|
|
29
|
-
- **Known external kicks** (load, network, other services):
|
|
30
|
-
_________________________________
|
|
31
|
-
- **Feedforward (if measurable)** / **feedback only (if not)**:
|
|
32
|
-
_________________________________
|
|
33
|
-
|
|
34
|
-
## Convergence
|
|
35
|
-
- **Pass condition** (bounded deviation ||ε|| under which perturbation?):
|
|
36
|
-
_________________________________
|
|
37
|
-
- **Test / observation that proves it**:
|
|
38
|
-
_________________________________
|
|
39
|
-
|
|
40
|
-
## Coupling
|
|
41
|
-
- **Subsystems that must stay separate** (decouple harmful coupling):
|
|
42
|
-
_________________________________
|
|
43
|
-
- **Subsystems that must stay coordinated** (keep beneficial coupling, regulate relations):
|
|
44
|
-
_________________________________
|
|
45
|
-
|
|
46
|
-
---
|
|
47
|
-
Anti-patterns to avoid:
|
|
48
|
-
- Unbounded retry/recursion/allocation → saturation → instability
|
|
49
|
-
- Feature creep before the base is stable → divergence
|
|
50
|
-
- Optimization before correctness → wrong trajectory
|
|
51
|
-
- Guessing the cause without measurement → open-loop patch
|
|
1
|
+
# Change Proposal (control-systems thinking scaffold)
|
|
2
|
+
|
|
3
|
+
Fill this in before a non-trivial change. It forces the control-theoretic mindset: name the
|
|
4
|
+
plant, the setpoint, the actuator, and how you will know it converged.
|
|
5
|
+
|
|
6
|
+
## Target
|
|
7
|
+
- **Controlled variable x** (what must stay correct):
|
|
8
|
+
_________________________________
|
|
9
|
+
- **Setpoint x*** (target / invariant / acceptable range):
|
|
10
|
+
_________________________________
|
|
11
|
+
|
|
12
|
+
## Plant & model
|
|
13
|
+
- **Current behavior** (measured, not assumed):
|
|
14
|
+
_________________________________
|
|
15
|
+
- **Mental model of the plant** (and where it may be wrong):
|
|
16
|
+
_________________________________
|
|
17
|
+
- **Model validity range** (inputs/sizes/conditions where the model holds):
|
|
18
|
+
_________________________________
|
|
19
|
+
|
|
20
|
+
## Control action
|
|
21
|
+
- **Control variable u** (the lever I will change):
|
|
22
|
+
_________________________________
|
|
23
|
+
- **Bounds on u** (max retries / depth / rate / size):
|
|
24
|
+
_________________________________
|
|
25
|
+
- **Sensors** (how I will observe x after the change — tests/logs/metrics):
|
|
26
|
+
_________________________________
|
|
27
|
+
|
|
28
|
+
## Disturbances
|
|
29
|
+
- **Known external kicks** (load, network, other services):
|
|
30
|
+
_________________________________
|
|
31
|
+
- **Feedforward (if measurable)** / **feedback only (if not)**:
|
|
32
|
+
_________________________________
|
|
33
|
+
|
|
34
|
+
## Convergence
|
|
35
|
+
- **Pass condition** (bounded deviation ||ε|| under which perturbation?):
|
|
36
|
+
_________________________________
|
|
37
|
+
- **Test / observation that proves it**:
|
|
38
|
+
_________________________________
|
|
39
|
+
|
|
40
|
+
## Coupling
|
|
41
|
+
- **Subsystems that must stay separate** (decouple harmful coupling):
|
|
42
|
+
_________________________________
|
|
43
|
+
- **Subsystems that must stay coordinated** (keep beneficial coupling, regulate relations):
|
|
44
|
+
_________________________________
|
|
45
|
+
|
|
46
|
+
---
|
|
47
|
+
Anti-patterns to avoid:
|
|
48
|
+
- Unbounded retry/recursion/allocation → saturation → instability
|
|
49
|
+
- Feature creep before the base is stable → divergence
|
|
50
|
+
- Optimization before correctness → wrong trajectory
|
|
51
|
+
- Guessing the cause without measurement → open-loop patch
|