math-skill 1.0.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (87) hide show
  1. package/README.en-US.md +313 -0
  2. package/README.md +313 -278
  3. package/agents/math-critic.en.md +235 -0
  4. package/agents/math-critic.md +237 -203
  5. package/commands/abstraction.md +11 -34
  6. package/commands/algorithmic-thinking.md +11 -34
  7. package/commands/ask.md +18 -21
  8. package/commands/axiomatization.md +11 -34
  9. package/commands/causal-inference.md +11 -34
  10. package/commands/discrete-combinatorial.md +11 -34
  11. package/commands/game-theory.md +11 -34
  12. package/commands/induction-analogy.md +11 -34
  13. package/commands/information-theory.md +11 -34
  14. package/commands/logic-deduction.md +11 -34
  15. package/commands/modeling.md +11 -37
  16. package/commands/optimization.md +11 -33
  17. package/commands/probability-statistics.md +11 -36
  18. package/commands/symmetry-invariance.md +11 -34
  19. package/commands/topological-thinking.md +11 -33
  20. package/commands/transformation.md +11 -33
  21. package/knowledge-base/overview.en.md +228 -0
  22. package/knowledge-base/overview.md +230 -230
  23. package/package.json +73 -59
  24. package/references/agentic-workflow.en.md +53 -0
  25. package/references/agentic-workflow.md +55 -0
  26. package/references/books/abstract-algebra.md +124 -0
  27. package/references/books/algebraic-geometry-rising-sea.md +171 -0
  28. package/references/books/differential-geometry.md +140 -0
  29. package/references/books/matrix-analysis.md +146 -0
  30. package/references/books/micro-lie-theory.md +116 -0
  31. package/references/books/optimization-ml.md +164 -0
  32. package/references/books/smooth-manifolds.md +105 -0
  33. package/references/gpu-friendly-math.en.md +65 -0
  34. package/references/gpu-friendly-math.md +67 -0
  35. package/references/inspiration.en.md +113 -0
  36. package/{docs → references}/inspiration.md +2 -0
  37. package/skills/abstraction/SKILL.en.md +117 -0
  38. package/skills/abstraction/SKILL.md +121 -264
  39. package/skills/abstraction/original-texts.en.md +163 -0
  40. package/skills/algorithmic-thinking/SKILL.en.md +132 -0
  41. package/skills/algorithmic-thinking/SKILL.md +138 -371
  42. package/skills/algorithmic-thinking/original-texts.en.md +253 -0
  43. package/skills/axiomatization/SKILL.en.md +144 -0
  44. package/skills/axiomatization/SKILL.md +151 -213
  45. package/skills/axiomatization/original-texts.en.md +154 -0
  46. package/skills/causal-inference/SKILL.en.md +147 -0
  47. package/skills/causal-inference/SKILL.md +151 -374
  48. package/skills/causal-inference/original-texts.en.md +136 -0
  49. package/skills/discrete-combinatorial/SKILL.en.md +124 -0
  50. package/skills/discrete-combinatorial/SKILL.md +131 -286
  51. package/skills/discrete-combinatorial/original-texts.en.md +184 -0
  52. package/skills/game-theory/SKILL.en.md +117 -0
  53. package/skills/game-theory/SKILL.md +123 -318
  54. package/skills/game-theory/original-texts.en.md +131 -0
  55. package/skills/induction-analogy/SKILL.en.md +145 -0
  56. package/skills/induction-analogy/SKILL.md +152 -310
  57. package/skills/induction-analogy/original-texts.en.md +140 -0
  58. package/skills/information-theory/SKILL.en.md +134 -0
  59. package/skills/information-theory/SKILL.md +140 -242
  60. package/skills/information-theory/original-texts.en.md +127 -0
  61. package/skills/logic-deduction/SKILL.en.md +130 -0
  62. package/skills/logic-deduction/SKILL.md +135 -280
  63. package/skills/logic-deduction/original-texts.en.md +160 -0
  64. package/skills/math-research-activator/SKILL.en.md +132 -0
  65. package/skills/math-research-activator/SKILL.md +136 -0
  66. package/skills/math-research-activator/original-texts.en.md +105 -0
  67. package/skills/{meta-selector → math-research-activator}/original-texts.md +104 -104
  68. package/skills/modeling/SKILL.en.md +135 -0
  69. package/skills/modeling/SKILL.md +139 -318
  70. package/skills/modeling/original-texts.en.md +162 -0
  71. package/skills/optimization/SKILL.en.md +129 -0
  72. package/skills/optimization/SKILL.md +135 -292
  73. package/skills/optimization/original-texts.en.md +167 -0
  74. package/skills/probability-statistics/SKILL.en.md +146 -0
  75. package/skills/probability-statistics/SKILL.md +151 -312
  76. package/skills/probability-statistics/original-texts.en.md +191 -0
  77. package/skills/symmetry-invariance/SKILL.en.md +135 -0
  78. package/skills/symmetry-invariance/SKILL.md +139 -358
  79. package/skills/symmetry-invariance/original-texts.en.md +206 -0
  80. package/skills/topological-thinking/SKILL.en.md +124 -0
  81. package/skills/topological-thinking/SKILL.md +128 -273
  82. package/skills/topological-thinking/original-texts.en.md +134 -0
  83. package/skills/transformation/SKILL.en.md +120 -0
  84. package/skills/transformation/SKILL.md +124 -264
  85. package/skills/transformation/original-texts.en.md +204 -0
  86. package/docs/CLAUDE.md +0 -187
  87. package/skills/meta-selector/SKILL.md +0 -188
@@ -1,292 +1,135 @@
1
- ---
2
- name: optimization
3
- description: |
4
- 触发:当问题涉及资源分配、取舍、最大化/最小化目标,或在限制条件下做决策,或需要判断问题的凸性、使用拉格朗日方法、分析对偶结构时调用。
5
- 模式:科研模式适用于数学推导、论文审查、实验设计优化;生活模式适用于日常决策、时间管理、人际互动、生活规划。
6
- Trigger when a problem involves resource allocation, trade-offs, maximizing/minimizing objectives, decision-making under constraints, or when convexity analysis, Lagrangian methods, or duality structures are relevant.
7
- Mode: Research mode for mathematical derivation, paper review, experimental design optimization; Life mode for daily decisions, time management, interpersonal interactions, life planning.
8
- ---
9
-
10
- # ⚖️ 优化思想
11
-
12
- > "在最一般的约束条件下,寻找目标函数的极值——凸性决定难度,KKT 给出必要条件,对偶揭示结构。"
13
- > "Under the most general constraints, find extrema of the objective — convexity determines difficulty, KKT gives necessity, duality reveals structure."
14
- >
15
- > —— 最优化理论与运筹学
16
- > —— Optimization Theory & Operations Research
17
-
18
- ## 核心原则 / Core Principle
19
-
20
- **任何决策问题都可以表述为优化问题:在约束条件下最大化(或最小化)某个目标。优化的本质不是追求'最好',而是在约束下追求'可行中的最好'。**
21
-
22
- **Any decision problem can be formulated as an optimization problem: maximizing (or minimizing) an objective subject to constraints. The essence of optimization is not pursuing 'the best' in the abstract, but pursuing 'the best among the feasible' under constraints.**
23
-
24
- > 人生就是一场最优化——目标是你追求的价值,起点是你的出身,约束是时间、健康、规则。局部最优是当前环境中看似最好的选择,全局最优或许在另一个领域。探索与利用的权衡是人生的根本张力。重要的不是找到传说中的全局最优,而是在每一步迭代中保持方向大致正确、接受噪声与约束、允许目标随阅历增长而优雅地演变。过程即意义,优化本身就是生活。
25
-
26
- > Life itself is an optimization — the objective is what you value, the starting point is where you begin, constraints are time, health, and rules. Local optima are the best-looking choices in your current environment; the global optimum might be in a completely different field. The explore-exploit tradeoff is life's fundamental tension. The point is not finding a legendary global optimum, but keeping the direction roughly right at each iteration, accepting noise and constraints, and letting the objective evolve gracefully as experience grows. The process is the meaning; optimization itself is life.
27
-
28
- 优化的三个核心要素:
29
- - **目标(Objective)**:你要最大化或最小化什么
30
- - **约束(Constraints)**:不可违反的限制
31
- - **可行选项(Feasible Options)**:满足所有约束的选择范围
32
-
33
- > **数学形式化 / Mathematical Formalization**(科研模式参考)
34
- >
35
- > 一般优化问题:$\min_{x \in \mathbb{R}^n} f(x) \quad \text{s.t.} \quad g_i(x) \leq 0, \; i=1,\dots,m; \quad h_j(x) = 0, \; j=1,\dots,p$
36
- >
37
- > 其中 $f(x)$ 为目标函数,$g_i(x)$ 为不等式约束,$h_j(x)$ 为等式约束。
38
- >
39
- > 拉格朗日函数:$L(x, \lambda, \mu) = f(x) + \sum_{i=1}^{m} \lambda_i g_i(x) + \sum_{j=1}^{p} \mu_j h_j(x)$
40
- >
41
- > KKT 条件(核心必要条件):若 $x^*$ 为最优解且约束满足某种正则性条件(如 Slater 条件),则存在 $\lambda^* \geq 0$, $\mu^*$ 使得:
42
- > 1. **驻点条件**:$\nabla_x L(x^*, \lambda^*, \mu^*) = 0$
43
- > 2. **原始可行性**:$g_i(x^*) \leq 0$, $h_j(x^*) = 0$
44
- > 3. **对偶可行性**:$\lambda_i^* \geq 0$
45
- > 4. **互补松弛性**:$\lambda_i^* g_i(x^*) = 0$
46
- >
47
- > 凸性定义与关键性质:若 $f$ 和所有 $g_i$ 为凸函数,所有 $h_j$ 为线性函数,则问题为凸优化。凸优化中,**KKT 条件变为充分条件**,且局部最优 = 全局最优。
48
-
49
- ## 不适用场景 / When NOT to Use
50
-
51
- - **没有明确的评价标准**(不知道什么是"好")——先定义目标再谈优化 `[通用]`
52
- - **纯粹的执行性任务**(如"把这段代码格式化")——没有优化空间 `[通用]`
53
- - **用户已经确定了方案**——优化已由用户完成,无需重复 `[通用]`
54
- - **问题本质是定性判断而非定量极值**——应先建模再优化 `[科研]`
55
-
56
- ## 何时使用 / When to Use
57
-
58
- ### 科研触发条件 / Research Triggers
59
-
60
- - 需要判断问题是否为凸优化以决定求解难度
61
- - 科研项目的时间规划或实验设计优化
62
- - 需要在约束条件下做出理性决策,涉及可量化目标函数
63
- - 不确定当前策略是否最优,想要系统分析
64
-
65
- ### 生活触发条件 / Life Triggers
66
-
67
- - 需要在有限资源(时间、精力、金钱)之间做分配
68
- - 面临多目标取舍(如质量 vs 速度、短期 vs 长期)
69
- - 想搞清楚"在当前限制下,我最好的选择是什么"
70
- - 需要判断该优先争取什么资源(哪些限制一旦放宽收益最大)
71
- - 感觉自己陷入"眼前看似最好但可能不是真正最好"的困境
72
-
73
- ## 方法流程 / Method
74
-
75
- ### 第一步:定义目标 / Define the Objective
76
-
77
- 明确你要最大化或最小化什么。这是最重要的一步——目标不清晰,优化没有方向。
78
-
79
- #### 科研模式 / Research Mode
80
-
81
- **关键问题**:
82
- - 单一目标还是多目标?
83
- - 目标是可量化的吗?如果不能,能否找到代理变量?
84
- - 目标是时间相关的吗?(静态优化 vs 动态优化)
85
- - **$f$ 是否为凸函数?** 若凸,则局部最优即全局最优;若非凸,需警惕局部极值。
86
-
87
- #### 生活模式 / Life Mode
88
-
89
- 明确你要优化什么——这不是数学公式,而是人生方向。
90
-
91
- - 你真正想要的是什么?(一个明确的目标,还是多个相互矛盾的追求?)
92
- - 这个目标可以衡量吗?如果不能,用什么代理指标来判断"够好"?
93
- - 目标会不会随时间变化?(短期目标 vs 长期方向)
94
- - 核心追问:你追求的是当前环境下"看起来最好的",还是可能存在一个完全不同的领域更契合你?
95
-
96
- #### 共通要点 / Common Key Points
97
-
98
- 目标定义是优化的方向——方向错了,走得越远越偏。必须先回答"我要什么"再回答"怎么要"。
99
-
100
- ### 第二步:列出约束 / List the Constraints
101
-
102
- #### 科研模式 / Research Mode
103
-
104
- 区分**硬约束**与**软约束**,同时对约束进行数学分类:
105
-
106
- - **硬约束 [硬]**:物理限制、硬性 deadline、预算上限
107
- - **软约束 [软]**:个人偏好、质量下限、舒适区
108
- - **不等式约束 $g_i(x) \leq 0$**:定义可行域边界,是优化中最常见的约束类型
109
- - **等式约束 $h_j(x) = 0$**:减少可行域维度,通常对应物理守恒或精确要求
110
- - **线性约束 vs 非线性约束**:线性约束保持可行域为凸多面体;非线性约束可能使可行域非凸
111
-
112
- #### 生活模式 / Life Mode
113
-
114
- 搞清楚哪些限制是不可逾越的,哪些是可以协商的。
115
-
116
- - **不可逾越的限制**:法律法规、物理极限、硬性截止日期——你不能突破这些
117
- - **可以协商的限制**:个人偏好、舒适区、习惯——这些看似限制但其实是软约束,你可以选择突破
118
- - 核心追问:你以为的限制中,有多少是真正的硬约束?有多少只是你给自己设的软约束?
119
-
120
- #### 共通要点 / Common Key Points
121
-
122
- 约束定义了你的选择范围。区分硬约束和软约束至关重要——混淆两者要么让你错过可行选项,要么让你追逐不可行方案。
123
-
124
- ### 第三步:类型分类 / Classify the Problem Type
125
-
126
- #### 科研模式 / Research Mode
127
-
128
- 根据目标函数和约束的结构,判定问题类别:
129
-
130
- | 类型 | 目标 | 约束 | 核心性质 | 典型方法 |
131
- |------|------|------|----------|----------|
132
- | LP(线性规划) | 线性 | 线性不等式 | 全局最优在顶点 | 单纯形法 |
133
- | QP(二次规划) | 二次 | 线性 | 正定 QP 为凸 | 内点法 |
134
- | 凸优化 | 凸 | 凸不等式 + 线性等式 | 局部=全局 | 梯度下降、内点法 |
135
- | 非凸优化 | 非凸 | 任意 | 多局部极值 | 全局搜索、模拟退火 |
136
- | 组合优化 | 离散域 | 任意 | NP-hard 常见 | 分支定界、启发式 |
137
- | 随机优化 | 含随机项 | 可能含随机 | 期望最优 vs 随机可行 | SAA、鲁棒优化 |
138
-
139
- #### 生活模式 / Life Mode
140
-
141
- 判断你的问题属于哪种类型——这决定了你的求解策略:
142
-
143
- - **有明确最优答案**的问题:目标清晰、限制明确,存在一个"最好的"选择——直接找到它
144
- - **看似最好但未必真正最好的问题**:你可能在局部最优中——当前环境下最好的选择,换一个环境可能更好
145
- - **没有完美方案的问题**:多个目标互相矛盾,不可能同时满足——必须取舍
146
- - **信息不足的问题**:不确定选项的后果——找到"足够好"的方案即可,不必追求最优
147
-
148
- #### 共通要点 / Common Key Points
149
-
150
- 问题类型决定求解策略。不同类型需要不同方法——对"没有完美方案"的问题追求最优是浪费时间。
151
-
152
- ### 第四步:寻找最优解 / Find the Optimal Solution
153
-
154
- #### 科研模式 / Research Mode
155
-
156
- 根据类型分类选择求解策略:
157
- - **LP/QP/凸优化**:利用凸性,梯度类方法或内点法可保证收敛到全局最优
158
- - **非凸优化**:需警惕局部最优——尝试多起点、全局搜索策略,或松弛为凸近似
159
- - **组合优化**:精确求解往往 NP-hard——分支定界求小规模精确解,启发式求大规模近似解
160
- - **随机优化**:样本平均近似(SAA)将随机问题转化为确定性近似问题
161
- - **信息不足**:使用满意解(satisficing)——找到"足够好"的解即可
162
-
163
- #### 生活模式 / Life Mode
164
-
165
- 根据问题类型采取不同策略:
166
-
167
- - **有明确最优答案**:系统比较各选项,直接选最优
168
- - **可能陷入局部最优**:多探索不同领域,避免过早锁定——"多试几个方向再决定"比"在当前领域精雕细琢"有时更有效
169
- - **没有完美方案**:明确各目标的权重,在取舍中找平衡——接受"不完美但足够好"
170
- - **信息不足**:先收集关键信息,或采用"满意解"策略——找到"足够好"就行动,不必等"最优"
171
-
172
- #### 共通要点 / Common Key Points
173
-
174
- 求解策略必须匹配问题类型。对没有完美方案的问题追求最优是徒劳——满意解往往比最优解更务实。
175
-
176
- ### 第五步:灵敏度分析 / Sensitivity Analysis
177
-
178
- #### 科研模式 / Research Mode
179
-
180
- 拉格朗日乘子的经济学含义:$\lambda_i^*$ 为第 $i$ 个不等式约束的**影子价格**——约束放松一个单位,目标函数改善约 $\lambda_i^*$ 个单位。互补松弛性表明:若 $\lambda_i^* = 0$,则该约束对最优解无影响(非活跃约束);若 $\lambda_i^* > 0$,则约束活跃且最优解恰好在其边界上。
181
-
182
- **关键问题**:如果约束条件或目标函数发生微小变化,最优解会如何变化?哪些约束是活跃的?影子价格是多少?
183
-
184
- #### 生活模式 / Life Mode
185
-
186
- 搞清楚哪个限制最重要——如果某个限制放宽一点,你的选择空间会大幅扩展。
187
-
188
- - **关键限制的优先级**:哪个限制一旦放宽,你获得的选择改善最大?——这告诉你应该优先争取什么资源
189
- - **限制的活跃性**:有些限制根本不影响你的最佳选择(你离它的边界很远),有些限制恰好卡住你(你就在它的边界上)——只关注活跃的限制
190
- - 核心追问:如果我只能放宽一个限制,应该选哪个?这往往指向你该优先争取的资源
191
-
192
- #### 共通要点 / Common Key Points
193
-
194
- 灵敏度分析告诉你"什么最重要"——不是所有限制都同等重要,活跃限制对结果的影响远大于非活跃限制。
195
-
196
- ### 第六步:多目标与帕累托 / Multi-Objective and Pareto
197
-
198
- #### 科研模式 / Research Mode
199
-
200
- 当存在多个目标 $f_1, f_2, \dots, f_k$ 时,通常不存在单一最优解。帕累托最优解集定义:不存在另一个可行解使所有目标同时改善。
201
-
202
- 常用方法:
203
- - **加权求和法**:$\min \sum w_i f_i(x)$,权重 $w_i > 0$——不同权重对应帕累托前沿上不同点
204
- - **$\epsilon$-约束法**:将一个目标作为主目标,其余转为不等式约束 $\min f_1(x)$ s.t. $f_i(x) \leq \epsilon_i$——遍历 $\epsilon_i$ 可覆盖帕累托前沿
205
-
206
- #### 生活模式 / Life Mode
207
-
208
- 当追求多个目标且它们互相矛盾时,没有完美方案——必须取舍。
209
-
210
- - **取舍的本质**:有些目标之间必须牺牲一个来成就另一个,不存在同时满足所有目标的完美方案
211
- - **如何取舍**:明确各目标的权重——哪个对你更重要?或者先把最重要的目标做到位,其余只要"不太差"就行
212
- - 核心追问:你的各目标之间,哪些可以同时满足,哪些必须取舍?取舍时你优先保哪个?
213
-
214
- #### 共通要点 / Common Key Points
215
-
216
- 多目标问题没有唯一最优解——取舍是不可避免的。关键不是"找到完美方案",而是"在取舍中做出有意识的选择"。
217
-
218
- ### 第七步:监控约束变化 / Monitor Constraint Changes
219
-
220
- #### 科研模式 / Research Mode
221
-
222
- 最优解不是永恒的——当约束条件变化时,需要重新优化。活跃约束的变化对最优解影响最大(影子价格高),非活跃约束的微小变化通常不影响最优解。
223
-
224
- #### 生活模式 / Life Mode
225
-
226
- 你的最佳选择不是一成不变的——当限制条件变了,你该重新审视选择。
227
-
228
- - **什么变化最值得关注**:那些恰好卡住你的限制如果变化了,你的最优选择会随之改变;离你很远的限制即使变化,对你的选择也没什么影响
229
- - **何时需要重新决策**:当你面临的限制发生了明显变化(新的机会出现、旧的约束消失),就该重新审视选择
230
- - 核心追问:最近我的限制条件有没有变化?如果有,我的选择是否还合适?
231
-
232
- #### 共通要点 / Common Key Points
233
-
234
- 最优解依赖约束——约束变了,最优解就变了。定期审视限制条件是否发生变化,是持续优化的关键。
235
-
236
- ## 常见错误 / Common Errors
237
-
238
- | 错误 / Error | 批评 / Critique | 正确做法 / Correct Approach | 模式 |
239
- |-------------|-------------------------------|---------------------------|------|
240
- | 没有明确目标就优化 | 目标不清晰,方向不确定 | 先精确定义目标 | `[通用]` |
241
- | 忽略隐式约束 | 未发现的约束使"最优解"不可行 | 穷尽检查所有可能的约束 | `[通用]` |
242
- | 陷入局部最优 | 非凸问题中贪心策略不保证全局最优 | 验证凸性;非凸时用多起点或全局方法 | `[科研]` |
243
- | 把最优当唯一 | 最优解可能不唯一 | 检查是否存在多个等价最优解 | `[科研]` |
244
- | 多目标用单目标方法 | 不同目标需 trade-off | 使用帕累托分析 | `[科研]` |
245
- | 忘记重新优化 | 约束变了但未更新 | 定期检查约束是否变化 | `[通用]` |
246
- | 未验证凸性 | 非凸问题误用凸优化方法 | 先判断凸性,再选方法 | `[科研]` |
247
- | 忽略对偶理论 | 对偶问题可能更易求解 | 构造对偶问题,利用强对偶性 | `[科研]` |
248
- | 混淆可行与最优 | 可行解不一定最优;最优解必须可行 | 先验证可行性,再验证最优性 | `[科研]` |
249
- | 忽略计算复杂度 | 组合优化问题可能 NP-hard | 对大规模问题使用启发式或近似算法 | `[科研]` |
250
- | 把软约束当硬约束 | 以为不可突破的限制其实可以协商 | 严格区分不可逾越的限制和可以协商的限制 | `[生活]` |
251
- | 追求完美方案 | 多目标问题不存在完美方案,取舍不可避免 | 有意识地做取舍,接受"足够好" | `[生活]` |
252
- | 不关注哪个限制最关键 | 所有限制看似同等重要,但活跃限制的影响远大于非活跃限制 | 找出哪个限制一旦放宽收益最大,优先争取 | `[生活]` |
253
- | 在当前领域死磕 | 可能陷入局部最优——换个领域可能更好 | 多探索不同方向,不要过早锁定 | `[生活]` |
254
-
255
- ## 操作规程 / Operating Procedure
256
-
257
- **模式选择**:根据问题性质自动选择——
258
- - **科研模式**触发条件:涉及数学推导、定理证明、算法设计、论文审查、实验设计优化、需要判断凸性或使用对偶理论
259
- - **生活模式**触发条件:涉及日常决策、人际互动、时间管理、生活规划、资源分配、取舍判断
260
-
261
- 当本 skill 被触发时,根据选择的模式执行以下步骤:
262
-
263
- ### 科研模式输出格式 / Research Mode Output Format
264
-
265
- 1. **目标函数**:用一句话明确"我们要最大化/最小化什么" `[目标]: [描述]`,并判断凸性 `[凸性]: [凸/非凸/未知]`
266
- 2. **约束清单**:标注每个约束为 `[硬约束]` 或 `[软约束]`,并分类 `[不等式/等式]` `[线性/非线性]`
267
- 3. **类型分类**:判定问题类别 `[类型]: [LP/QP/凸/非凸/组合/随机]`
268
- 4. **可行域分析**:在约束下,哪些选项是可行的?活跃约束是哪些?
269
- 5. **最优解/满意解**:标注使用的策略 `[策略]: [梯度法/内点法/全局搜索/satisficing/帕累托]`
270
- 6. **灵敏度分析**:关键约束的影子价格是多少?若关键约束变化 X%,结论如何变化?
271
- 7. **行动建议**:明确写出"接下来我将……"
272
-
273
- **科研模式输出必须包含以上 7 项,不得只输出分析性文字而不给出结论。**
274
-
275
- ### 生活模式输出格式 / Life Mode Output Format
276
-
277
- 1. **核心追求**:一句话说清"我最想要什么" `[追求]: [描述]`——单一目标还是多个矛盾目标?
278
- 2. **现实限制**:标注每个限制 `[不可逾越]` 或 `[可以协商]`——哪些是真正的硬约束,哪些只是你给自己设的软约束?
279
- 3. **可选范围**:在这些限制下,我实际有哪些选择?最受限的选择维度是什么?
280
- 4. **最优选择/满意选择**:标注策略 `[策略]: [系统比较/多方向探索/取舍平衡/满意即可]`
281
- 5. **关键限制的优先级**:哪个限制一旦放宽收益最大?——这告诉我该优先争取什么资源
282
- 6. **行动建议**:明确写出"接下来我将……"
283
-
284
- **生活模式输出必须包含以上 6 项,不得只输出分析性文字而不给出结论。**
285
-
286
- ## 与其他 skill 的关系 / Relations to Other Skills
287
-
288
- - **建模思想**:优化前需要先建模——定义目标函数和约束本身就是一个建模过程
289
- - **概率与统计**:在不确定性下做优化需要随机优化或鲁棒优化方法
290
- - **变换思想**:有时变换到对偶问题更容易求解;对偶性是优化中最深刻的变换
291
- - **博弈思想**:当多个决策者同时优化时,问题变为博弈论中的均衡问题——纳什均衡即多人优化中的稳定点
292
- - **算法思想**:优化问题的求解依赖于算法设计——凸优化可用梯度法,组合优化需分支定界或启发式算法
1
+ ---
2
+ name: optimization
3
+ description: |
4
+ 触发:问题涉及资源分配、取舍、最大化/最小化目标、约束下决策;或需判断凸性、使用拉格朗日/KKT、分析对偶结构;或为算法/算子/训练设计选择优化方法时调用。
5
+ English: Trigger when a problem involves resource allocation, trade-offs, maximizing/minimizing objectives, decisions under constraints; or needs convexity analysis, Lagrangian/KKT methods, duality structure; or choosing optimization methods for algorithm/operator/training design.
6
+ ---
7
+
8
+ > **语言路由**:若用户消息为英文,请读取并遵循同目录下的 `SKILL.en.md`,按其操作规程以英文输出;中文消息则继续使用本文件。
9
+
10
+ # ⚖️ 优化思想 / Optimization
11
+
12
+ > "在最一般的约束条件下,寻找目标函数的极值——凸性决定难度,KKT 给出必要条件,对偶揭示结构。"
13
+ > "Under the most general constraints, find extrema of the objective — convexity determines difficulty, KKT gives necessity, duality reveals structure."
14
+ >
15
+ > —— 最优化理论与运筹学 / Optimization Theory & Operations Research
16
+
17
+ ## 核心原则 / Core Principle
18
+
19
+ **任何决策问题都可以表述为优化问题:在约束条件下最大化(或最小化)某个目标。优化的本质不是追求'最好',而是在约束下追求'可行中的最好'。**
20
+
21
+ **Any decision problem can be formulated as optimization: maximizing (or minimizing) an objective subject to constraints. The essence is not 'the best' in the abstract, but 'the best among the feasible'.**
22
+
23
+ 优化的三个核心要素:**目标(Objective)**、**约束(Constraints)**、**可行域(Feasible set)**。
24
+
25
+ > **数学形式化 / Mathematical Formalization**
26
+ >
27
+ > 一般优化问题:$\min_{x \in \mathbb{R}^n} f(x) \quad \text{s.t.} \quad g_i(x) \leq 0,\; i=1,\dots,m; \quad h_j(x) = 0,\; j=1,\dots,p$
28
+ >
29
+ > 拉格朗日函数:$L(x, \lambda, \mu) = f(x) + \sum_i \lambda_i g_i(x) + \sum_j \mu_j h_j(x)$
30
+ >
31
+ > KKT 条件(核心必要条件,Slater 等正则性下):① 驻点 $\nabla_x L = 0$;② 原始可行 $g_i \le 0, h_j = 0$;③ 对偶可行 $\lambda_i \ge 0$;④ 互补松弛 $\lambda_i g_i = 0$。
32
+ >
33
+ > 凸性:若 $f$ 与各 $g_i$ 凸、各 $h_j$ 线性,则问题凸;此时 **KKT 充分**,局部最优 = 全局最优。
34
+
35
+ ## GPU 友好性 / GPU-Friendliness(横切检查)
36
+
37
+ 当优化用于**算法/算子/训练设计**时,求解方法本身必须过 `../../references/gpu-friendly-math.md` 八维门:
38
+
39
+ - **一阶法(SGD/Adam)**:GEMM 友好、可并行、低精度可行;注意优化器状态精度与分布式通信开销。
40
+ - **二阶/牛顿法**:Hessian 求逆 $O(n^3)$、显存爆炸——典型"美但不可算"→ 改造为 **K-FAC / 低秩 / 对角近似**(见 `../../references/books/optimization-ml.md`、`matrix-analysis.md`)。
41
+ - **约束投影**:投影是否有闭式、可张量化?迭代投影警惕串行依赖。
42
+ - **分布式**:计算/通信能否 overlap;是否需要梯度压缩。
43
+
44
+ 八维最低判定(正式术语):**张量化**看目标/约束/梯度能否批量;**GEMM 可映射**看主计算是否为矩阵乘、HVP、K-FAC 小矩阵;**复杂度**明确一阶/二阶/组合求解阶;**显存与 KV-Cache**检查优化器状态、Hessian、激活保存;**低精度稳定**检查条件数、阻尼、loss scaling;**并行与通信**检查梯度同步和通信 overlap;**稀疏结构**看预条件/约束是否块结构化;**算子融合**看更新、裁剪、正则能否融合。
45
+
46
+ > 配合 `../../references/books/optimization-ml.md`(Chong/Lu/Żak)与 `../../references/books/matrix-analysis.md`。
47
+
48
+ ## 不适用场景 / When NOT to Use
49
+
50
+ - **没有明确评价标准**(不知道什么是"好")——先定义目标再优化。
51
+ - **纯执行性任务**(如格式化代码)——没有优化空间。
52
+ - **用户已确定方案**——优化已完成。
53
+ - **本质是定性判断而非定量极值**——应先建模再优化。
54
+
55
+ ## 何时使用 / When to Use
56
+
57
+ - 需要判断问题是否为凸优化以决定求解难度。
58
+ - 为算法/算子/训练设计选择优化方法,并评估其 GPU 可行性。
59
+ - 在约束条件下做可量化目标的理性决策。
60
+ - 实验设计、资源分配、超参/结构搜索的系统优化。
61
+ - 不确定当前策略是否最优,想系统分析(凸性、对偶、灵敏度)。
62
+
63
+ ## 方法流程 / Method
64
+
65
+ ### 第一步:定义目标 / Define the Objective
66
+ 明确最大化/最小化什么。关键:单目标还是多目标?是否可量化(否则找代理变量)?静态还是动态?**$f$ 是否凸**(凸则局部=全局,非凸需警惕局部极值)?方向错了,走得越远越偏。
67
+
68
+ ### 第二步:列出约束 / List the Constraints
69
+ 区分**硬约束**(物理/预算/deadline)与**软约束**(偏好/质量下限);数学上分类:不等式 $g_i(x)\le 0$(定义可行域边界)、等式 $h_j(x)=0$(降维)、线性(可行域为凸多面体)vs 非线性(可能非凸)。
70
+
71
+ ### 第三步:类型分类 / Classify the Problem Type
72
+
73
+ | 类型 | 目标 | 约束 | 核心性质 | 典型方法 |
74
+ |------|------|------|----------|----------|
75
+ | LP 线性规划 | 线性 | 线性不等式 | 最优在顶点 | 单纯形法 |
76
+ | QP 二次规划 | 二次 | 线性 | 正定 QP 为凸 | 内点法 |
77
+ | 凸优化 | 凸 | 凸不等式+线性等式 | 局部=全局 | 梯度下降、内点法 |
78
+ | 非凸优化 | 非凸 | 任意 | 多局部极值 | 全局搜索、模拟退火 |
79
+ | 组合优化 | 离散域 | 任意 | NP-hard 常见 | 分支定界、启发式 |
80
+ | 随机优化 | 含随机项 | 可能含随机 | 期望最优 vs 随机可行 | SAA、鲁棒优化 |
81
+
82
+ ### 第四步:寻找最优解 / Find the Optimal Solution
83
+ - **LP/QP/凸**:利用凸性,梯度类或内点法保证收敛到全局最优。
84
+ - **非凸**:多起点、全局搜索,或松弛为凸近似。
85
+ - **组合**:精确解常 NP-hard——小规模分支定界,大规模启发式/近似。
86
+ - **随机**:样本平均近似(SAA)转化为确定性近似。
87
+ - **信息不足**:满意解(satisficing)即可。
88
+
89
+ ### 第五步:灵敏度分析 / Sensitivity Analysis
90
+ 拉格朗日乘子 $\lambda_i^*$ 为第 $i$ 个约束的**影子价格**——约束放松一单位,目标改善约 $\lambda_i^*$。互补松弛:$\lambda_i^*=0$ 为非活跃约束(对最优解无影响);$\lambda_i^*>0$ 为活跃约束(最优解恰在其边界)。关注:约束/目标微小变化时最优解如何变、哪些约束活跃。
91
+
92
+ ### 第六步:多目标与帕累托 / Multi-Objective & Pareto
93
+ 多目标 $f_1,\dots,f_k$ 通常无单一最优解。帕累托最优:不存在使所有目标同时改善的可行解。方法:**加权求和** $\min\sum w_i f_i$(不同权重对应前沿不同点);**$\epsilon$-约束法** $\min f_1$ s.t. $f_i\le\epsilon_i$(遍历 $\epsilon_i$ 覆盖前沿)。
94
+
95
+ ### 第七步:监控约束变化 / Monitor Constraint Changes
96
+ 最优解依赖约束——约束变了就要重新优化。活跃约束的变化影响最大(影子价格高),非活跃约束的微小变化通常不影响最优解。
97
+
98
+ ## 常见错误 / Common Errors
99
+
100
+ | 错误 / Error | 批评 / Critique | 正确做法 / Correct Approach |
101
+ |-------------|----------------|---------------------------|
102
+ | 没有明确目标就优化 | 方向不确定 | 先精确定义目标 |
103
+ | 忽略隐式约束 | "最优解"实则不可行 | 穷尽检查所有约束 |
104
+ | 陷入局部最优 | 非凸贪心不保证全局 | 验证凸性;非凸用多起点/全局法 |
105
+ | 把最优当唯一 | 最优解可能不唯一 | 检查是否存在多个等价最优解 |
106
+ | 多目标用单目标方法 | 不同目标需 trade-off | 使用帕累托分析 |
107
+ | 未验证凸性 | 非凸误用凸方法 | 先判断凸性再选方法 |
108
+ | 忽略对偶理论 | 对偶问题可能更易解 | 构造对偶,利用强对偶性 |
109
+ | 混淆可行与最优 | 可行不一定最优 | 先验证可行性再验证最优性 |
110
+ | 忽略计算/GPU 复杂度 | 二阶法/组合优化可能不可算 | 评估复杂度,过 GPU 八维门,必要时近似 |
111
+ | 忘记重新优化 | 约束变了未更新 | 定期检查约束变化 |
112
+
113
+ ## 操作规程 / Operating Procedure
114
+
115
+ 当本 skill 被触发时,输出必须包含:
116
+
117
+ 1. **目标函数**:`[目标]: [描述]` + `[凸性]: [凸/非凸/未知]`
118
+ 2. **约束清单**:每条标 `[硬/软]` 与 `[不等式/等式]` `[线性/非线性]`
119
+ 3. **类型分类**:`[类型]: [LP/QP/凸/非凸/组合/随机]`
120
+ 4. **可行域分析**:哪些选项可行?活跃约束是哪些?
121
+ 5. **最优解/满意解**:`[策略]: [梯度法/内点法/全局搜索/satisficing/帕累托]`
122
+ 6. **灵敏度分析**:关键约束影子价格?变化 X% 结论如何变?
123
+ 7. **GPU 可行性**(若用于算法/算子/训练):求解方法过八维门,标注友好/可改造/不友好 + 改造建议。
124
+ 8. **行动建议**:明确写出"接下来我将……"
125
+
126
+ **输出不得只给分析而无结论。**
127
+
128
+ ## 与其他 skill 的关系 / Relations to Other Skills
129
+
130
+ - **建模思想**:优化前需先建模——定义目标与约束本身就是建模。
131
+ - **概率与统计**:不确定性下的优化需随机/鲁棒优化。
132
+ - **变换思想**:变换到对偶问题常更易求解;对偶是优化中最深刻的变换。
133
+ - **博弈思想**:多决策者同时优化即博弈,纳什均衡为多人优化的稳定点。
134
+ - **算法思想**:求解依赖算法设计——凸优化用梯度法,组合优化需分支定界/启发式。
135
+ - **现代数学激活**:`../../references/books/optimization-ml.md`(GPU 友好优化器、二阶法可行性)、`matrix-analysis.md`(条件数、低秩、预条件)。
@@ -0,0 +1,167 @@
1
+ # Mathematical Sources and Classic Texts
2
+
3
+ ## Euler-Lagrange Variational Equation & Brachistochrone (1696)
4
+
5
+ > In 1696, Johann Bernoulli posed the brachistochrone problem: under gravity alone, along which curve does a particle descend from A to B in the shortest time? The answer is the cycloid.
6
+
7
+ **Core equation**: For the functional J[y] = ∫_{a}^{b} F(x, y, y') dx, the extremizing function satisfies
8
+
9
+ > ∂F/∂y - d/dx(∂F/∂y') = 0 (Euler-Lagrange equation)
10
+
11
+ For the brachistochrone, F = √(1 + y'²) / √(2g(y₀ - y)). Substituting into the Euler-Lagrange equation yields the parametric equations of the cycloid: x = a(θ - sin θ), y = a(1 - cos θ).
12
+
13
+ **Key idea**: The calculus of variations reduces "optimization over an entire class of functions" to solving a differential equation — this is the origin of infinite-dimensional optimization. Euler (1744) systematized the calculus of variations; Lagrange (1788) introduced the δ notation and the method of multipliers, both in the same lineage. The calculus of variations later evolved into optimal control theory (Pontryagin's Maximum Principle, 1961), with wide applications in spacecraft trajectory design, robotic path planning, and many other fields.
14
+
15
+ ## Fermat & Weierstrass Conditions for Unconstrained Optimization
16
+
17
+ > Fermat (1636): A necessary condition for a differentiable function f to attain a local extremum at x* is ∇f(x*) = 0.
18
+ > Weierstrass (19th c.): If f is twice differentiable and ∇f(x*) = 0, then the Hessian H(x*) is positive definite ⇔ x* is a strict local minimum; H(x*) is negative definite ⇔ x* is a strict local maximum.
19
+
20
+ For convex functions, ∇f(x*) = 0 is not only necessary but also sufficient — and x* is the global minimum.
21
+
22
+ **Key idea**: First-order conditions locate stationary points; second-order conditions determine the nature of the extremum — the most fundamental theorem of unconstrained optimization. The condition ∇f = 0 is the convergence target of virtually all iterative optimization algorithms (gradient descent, Newton's method).
23
+
24
+ ## Lagrange Multipliers (1788)
25
+
26
+ > Finding the extremum of f(x) subject to g(x) = 0 is equivalent to finding the unconstrained extremum of L(x,λ) = f(x) - λg(x).
27
+
28
+ **Key idea**: Transform a constrained optimization problem into an unconstrained one. At the optimal solution, the gradient of the objective function is parallel to the gradient of the constraint function: ∇f(x*) = λ∇g(x*).
29
+
30
+ ## Pareto Optimality (Vilfredo Pareto, ~1906)
31
+
32
+ > In multi-objective optimization, a solution x is Pareto optimal if there does not exist y such that f_i(y) ≤ f_i(x) for all i with at least one f_j(y) < f_j(x) (strict improvement).
33
+
34
+ In *Manuale di economia politica* (1906), Pareto introduced this concept into economics: the efficiency frontier of resource allocation. The Pareto frontier {f(x) : x is Pareto optimal} characterizes all trade-off solutions that cannot be simultaneously improved. Mathematically, for objectives f₁,...,fₘ, the Pareto frontier is {y ∈ ℝᵐ : y = (f₁(x),...,fₘ(x)), x Pareto optimal} — it is a (generally non-convex) surface in objective space.
35
+
36
+ > Weighted scalarization: Convert multi-objective min f₁,...,fₘ into single-objective min Σw_i f_i (w_i > 0); each weight combination corresponds to a point on the Pareto frontier. This is a standard method for computing the Pareto frontier, but it cannot enumerate all Pareto-optimal solutions on non-convex frontiers.
37
+
38
+ **Key idea**: Multi-objective optimization has no single "optimal solution" but rather a set of Pareto-optimal solutions among which trade-offs must be made. This is ubiquitous in economics (efficiency vs. equity), engineering (performance vs. cost), and machine learning (accuracy vs. complexity).
39
+
40
+ ## Game Theory & Optimization (von Neumann, 1928; Nash, 1950)
41
+
42
+ > Nash equilibrium: A strategy profile (s₁*, ..., sₙ*) such that for each player i, s_i* is the best response given the other players' strategies — i.e., no player has an incentive to unilaterally deviate.
43
+
44
+ Von Neumann (1928) proved the minimax theorem for zero-sum games: max_{x} min_{y} f(x,y) = min_{y} max_{x} f(x,y), which is essentially a precursor of linear programming duality. Nash (1950) generalized the equilibrium concept to non-zero-sum games; a Nash equilibrium is an " mutually optimal" optimization — each player's optimization problem takes the other players' strategies as constraints.
45
+
46
+ > Existence of Nash equilibrium: Nash (1950) used the Brouwer fixed-point theorem to prove that every finite game has at least one mixed-strategy Nash equilibrium. This reveals a deep connection between game theory and topology.
47
+
48
+ **Key idea**: Game theory extends optimization from "single-agent decision-making" to "multi-agent interactive decision-making" — optimality is no longer with respect to nature, but the best response to other rational agents.
49
+
50
+ ## Karush-Kuhn-Tucker Conditions (1939/1951)
51
+
52
+ Necessary conditions for inequality-constrained optimization:
53
+
54
+ > For min f(x) s.t. g_i(x) ≤ 0, h_j(x) = 0, the optimal solution x* satisfies:
55
+ > - Stationarity: ∇f(x*) + Σμ_i∇g_i(x*) + Σλ_j∇h_j(x*) = 0
56
+ > - Primal feasibility: g_i(x*) ≤ 0, h_j(x*) = 0
57
+ > - Complementary slackness: μ_i · g_i(x*) = 0, μ_i ≥ 0
58
+
59
+ Karush (1939, master's thesis) first derived these conditions; Kuhn & Tucker (1951) independently rediscovered and widely disseminated them.
60
+
61
+ **Key idea**: At the optimal solution, either a constraint is inactive (μ_i = 0) or it is tight (g_i(x*) = 0). Complementary slackness is the cornerstone of duality theory — it ties the KKT conditions to a zero duality gap. When a constraint qualification holds (e.g., LICQ: the gradients of active constraints are linearly independent), the KKT conditions are necessary for optimality.
62
+
63
+ ## Linear Programming Duality (Farkas, 1902; Dantzig-von Neumann, 1947)
64
+
65
+ > Farkas' Lemma (1902): Ax = b, x ≥ 0 has a solution ⟺ Aᵀy ≥ 0, bᵀy < 0 has no solution — this is the geometric foundation of LP duality.
66
+
67
+ > Strong duality theorem: If the primal LP min cᵀx s.t. Ax = b, x ≥ 0 has a feasible solution, then the optimal values of the primal and the dual max bᵀy s.t. Aᵀy ≤ c are equal. The components of the dual variable y are called shadow prices — they measure the "marginal value" of each constraint.
68
+
69
+ **Key idea**: Every LP has a "mirror" problem — the primal views the problem in terms of cost, the dual in terms of value. Zero duality gap is a profound fact: the minimum cost of resources equals the maximum revenue valued at shadow prices.
70
+
71
+ ## Simplex Method (Dantzig, 1947)
72
+
73
+ The fundamental algorithm for linear programming. It searches for the optimal solution by moving along vertices of the feasible region.
74
+
75
+ > Klee-Minty cube (1973): There exist LP instances for which the simplex method visits all 2ⁿ vertices — exponential worst-case complexity.
76
+
77
+ > Interior-Point Methods: Karmarkar (1984) proposed a polynomial-time LP algorithm that converges along the central path through the interior of the feasible region, avoiding vertex enumeration. Modern interior-point methods (e.g., Mehrotra predictor-corrector) perform excellently in both practice and theory.
78
+
79
+ **Key idea**: The optimal solution of a linear program always lies at a vertex of the feasible region. The simplex method exploits this structure by moving along edges; interior-point methods approach from the inside — the two are complementary.
80
+
81
+ ## Duality Theory
82
+
83
+ > Weak duality: For any optimization problem, the dual optimal value d* ≤ primal optimal value p* (duality gap = p* - d* ≥ 0).
84
+
85
+ > Strong duality (Slater's condition): If a convex optimization problem has a strictly feasible point (Slater point: g_i(x) < 0 strictly), then d* = p*, and the duality gap is zero.
86
+
87
+ > Duality provides lower bounds (the primal optimal value ≥ the dual value), which is a critical tool in branch-and-bound, cutting-plane, and related algorithms.
88
+
89
+ **Key idea**: Duality connects "minimizing cost" with "maximizing value" — two perspectives on the same problem. Slater's condition guarantees "no information loss" in convex optimization.
90
+
91
+ ## Gradient Descent (Cauchy, 1847; Robbins-Monro, 1951)
92
+
93
+ > Cauchy (1847) proposed the method of steepest descent: x_{k+1} = x_k - α_k ∇f(x_k), iterating along the negative gradient direction with step size α_k. For L-smooth convex functions (Lipschitz-continuous gradient), with step size α = 1/L, f(x_k) - f(x*) ≤ L‖x₀ - x*‖² / (2k) — a convergence rate of O(1/k).
94
+
95
+ > Newton's Method: x_{k+1} = x_k - H(x_k)⁻¹ ∇f(x_k), using second-order information, achieves quadratic convergence O(‖x_k - x*‖²) near the optimum.
96
+
97
+ > Stochastic Gradient Descent (SGD): Robbins & Monro (1951) proposed stochastic approximation, replacing the true gradient with a noisy estimate g_k ≈ ∇f(x_k). The conditions Σα_k = ∞, Σα_k² < ∞ guarantee convergence. SGD converges in expectation at the cost of variance — minibatching is a trade-off between variance and efficiency.
98
+
99
+ **Key idea**: Gradient descent is the most naive form of optimization — "go where it is steepest." SGD extends this idea to large-data settings: one need not examine all the data, only a random sample suffices to approximate the direction. The training of deep neural networks relies almost entirely on SGD and its variants (Adam, AdaGrad, etc.).
100
+
101
+ ## Integer Programming & NP-Hardness (Cook, 1971; Karp, 1972)
102
+
103
+ > Integer Linear Programming (ILP): min cᵀx s.t. Ax ≤ b, x ∈ ℤⁿ — variables take integer values. Even 0-1 ILP (x ∈ {0,1}ⁿ) is NP-hard.
104
+
105
+ > Cook (1971) proved that SAT is NP-complete (the Cook-Levin theorem); Karp (1972) listed 21 NP-complete problems, including integer programming, the knapsack problem, and the traveling salesman problem. ILP is NP-hard — no polynomial-time exact algorithm exists (under the assumption P ≠ NP).
106
+
107
+ > Convex relaxation: Relaxing x ∈ ℤⁿ to x ∈ ℝⁿ yields an LP whose optimal value provides a lower bound for the ILP — this is precisely the core mechanism of branch-and-bound.
108
+
109
+ **Key idea**: Discrete optimization is fundamentally harder than continuous optimization. The seemingly minor constraint of "integrality" pushes a problem from polynomial-time solvability to NP-hardness. Branch-and-bound and cutting-plane methods are the principal exact methods; heuristics are more commonly used in practice.
110
+
111
+ ## Convex Optimization (Boyd & Vandenberghe, 2004)
112
+
113
+ > Convex optimization: min f(x) s.t. g_i(x) ≤ 0 (convex), h_j(x) = 0 (affine), x ∈ C (convex set). When f is convex, a local minimum is a global minimum.
114
+
115
+ > Key properties: (1) Local minimum = global minimum; (2) The feasible region is a convex set; (3) Strong duality holds under Slater's condition; (4) Interior-point methods can solve it in polynomial time.
116
+
117
+ Important subclasses encompassed by convex optimization:
118
+ - LP (Linear Programming): f linear, constraints linear
119
+ - QP (Quadratic Programming): f quadratic, constraints linear
120
+ - SOCP (Second-Order Cone Programming): constraints include second-order cones ‖Ax + b‖ ≤ cᵀx + d
121
+ - SDP (Semidefinite Programming): constraints include positive semidefinite matrix conditions X ≥ 0
122
+ - GP (Geometric Programming): convex after logarithmic transformation
123
+
124
+ Boyd & Vandenberghe's *Convex Optimization* (2004) unified all of the above problem classes and became the standard textbook and engineering reference for modern optimization. Convex optimization marks the boundary of "efficiently solvable optimization" — convexity guarantees algorithmic convergence and global optimality of solutions.
125
+
126
+ **Key idea**: Convex optimization is the "sweet spot" of optimization theory — rich enough to encompass a vast number of practical problems, yet structured enough to guarantee tractability. Non-convex problems are often approximated via convex relaxation.
127
+
128
+ ## Bellman's Principle of Optimality (1957)
129
+
130
+ > "An optimal policy has the property that whatever the initial state and initial decision are, the remaining decisions must constitute an optimal policy with regard to the state resulting from the first decision."
131
+
132
+ > Recursive equation: V(s) = max_a { R(s,a) + γ · V(s') }, where s' is the result of the state transition.
133
+
134
+ **Key idea**: The core of dynamic programming — substructures of an optimal solution are themselves optimal. The Bellman equation decomposes multi-stage decision-making into single-stage subproblems and is the theoretical foundation of reinforcement learning.
135
+
136
+ ## The Philosophical Significance of Optimization
137
+
138
+ The core of optimization thinking is **making the best choice under constraints** — this is not only a mathematical question but also a question about life. Most decisions in life can be formulated as optimization problems:
139
+
140
+ - **Objective function**: What do you want to maximize? (Happiness? Achievement? Freedom?)
141
+ - **Constraints**: What limitations do you face? (Time? Money? Ability?)
142
+ - **Feasible region**: What choices are available to you under the constraints?
143
+ - **Optimal solution**: Among all choices, which one maximizes the objective?
144
+ - **Dual perspective**: Viewing the problem in terms of "cost" or "value" yields the same answer — this is the philosophy of strong duality.
145
+ - **Convex vs. non-convex**: Convex optimization has a unique global optimum; non-convex optimization has multiple local optima — the predicament of life lies precisely in the fact that the world is non-convex.
146
+
147
+ ## Timeline of Optimization
148
+
149
+ | Year | Event |
150
+ |------|-------|
151
+ | 1636 | Fermat proposes the necessary condition ∇f = 0 for unconstrained extrema |
152
+ | 1696 | Johann Bernoulli poses the brachistochrone problem; the seeds of the calculus of variations |
153
+ | 1744 | Euler systematizes the calculus of variations |
154
+ | 1788 | Lagrange publishes the method of multipliers in *Mécanique analytique* |
155
+ | 1902 | Farkas' Lemma — the geometric foundation of LP duality |
156
+ | ~1906 | Pareto introduces the efficiency frontier concept into economics |
157
+ | 1928 | Von Neumann proves the minimax theorem (zero-sum games) |
158
+ | 1939 | Karush derives the KKT conditions (master's thesis) |
159
+ | 1947 | Dantzig invents the simplex method; von Neumann establishes LP duality |
160
+ | 1950 | Nash proves the existence of equilibria in non-zero-sum games |
161
+ | 1951 | Kuhn & Tucker publish the KKT conditions; Robbins-Monro propose stochastic approximation |
162
+ | 1957 | Bellman publishes dynamic programming and the principle of optimality |
163
+ | 1971 | Cook proves NP-completeness (the Cook-Levin theorem) |
164
+ | 1972 | Karp lists 21 NP-complete problems |
165
+ | 1973 | Klee-Minty prove exponential worst-case complexity of the simplex method |
166
+ | 1984 | Karmarkar proposes the polynomial-time interior-point method |
167
+ | 2004 | Boyd & Vandenberghe publish *Convex Optimization* |