math-skill 1.0.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (87) hide show
  1. package/README.en-US.md +313 -0
  2. package/README.md +313 -278
  3. package/agents/math-critic.en.md +235 -0
  4. package/agents/math-critic.md +237 -203
  5. package/commands/abstraction.md +11 -34
  6. package/commands/algorithmic-thinking.md +11 -34
  7. package/commands/ask.md +18 -21
  8. package/commands/axiomatization.md +11 -34
  9. package/commands/causal-inference.md +11 -34
  10. package/commands/discrete-combinatorial.md +11 -34
  11. package/commands/game-theory.md +11 -34
  12. package/commands/induction-analogy.md +11 -34
  13. package/commands/information-theory.md +11 -34
  14. package/commands/logic-deduction.md +11 -34
  15. package/commands/modeling.md +11 -37
  16. package/commands/optimization.md +11 -33
  17. package/commands/probability-statistics.md +11 -36
  18. package/commands/symmetry-invariance.md +11 -34
  19. package/commands/topological-thinking.md +11 -33
  20. package/commands/transformation.md +11 -33
  21. package/knowledge-base/overview.en.md +228 -0
  22. package/knowledge-base/overview.md +230 -230
  23. package/package.json +73 -59
  24. package/references/agentic-workflow.en.md +53 -0
  25. package/references/agentic-workflow.md +55 -0
  26. package/references/books/abstract-algebra.md +124 -0
  27. package/references/books/algebraic-geometry-rising-sea.md +171 -0
  28. package/references/books/differential-geometry.md +140 -0
  29. package/references/books/matrix-analysis.md +146 -0
  30. package/references/books/micro-lie-theory.md +116 -0
  31. package/references/books/optimization-ml.md +164 -0
  32. package/references/books/smooth-manifolds.md +105 -0
  33. package/references/gpu-friendly-math.en.md +65 -0
  34. package/references/gpu-friendly-math.md +67 -0
  35. package/references/inspiration.en.md +113 -0
  36. package/{docs → references}/inspiration.md +2 -0
  37. package/skills/abstraction/SKILL.en.md +117 -0
  38. package/skills/abstraction/SKILL.md +121 -264
  39. package/skills/abstraction/original-texts.en.md +163 -0
  40. package/skills/algorithmic-thinking/SKILL.en.md +132 -0
  41. package/skills/algorithmic-thinking/SKILL.md +138 -371
  42. package/skills/algorithmic-thinking/original-texts.en.md +253 -0
  43. package/skills/axiomatization/SKILL.en.md +144 -0
  44. package/skills/axiomatization/SKILL.md +151 -213
  45. package/skills/axiomatization/original-texts.en.md +154 -0
  46. package/skills/causal-inference/SKILL.en.md +147 -0
  47. package/skills/causal-inference/SKILL.md +151 -374
  48. package/skills/causal-inference/original-texts.en.md +136 -0
  49. package/skills/discrete-combinatorial/SKILL.en.md +124 -0
  50. package/skills/discrete-combinatorial/SKILL.md +131 -286
  51. package/skills/discrete-combinatorial/original-texts.en.md +184 -0
  52. package/skills/game-theory/SKILL.en.md +117 -0
  53. package/skills/game-theory/SKILL.md +123 -318
  54. package/skills/game-theory/original-texts.en.md +131 -0
  55. package/skills/induction-analogy/SKILL.en.md +145 -0
  56. package/skills/induction-analogy/SKILL.md +152 -310
  57. package/skills/induction-analogy/original-texts.en.md +140 -0
  58. package/skills/information-theory/SKILL.en.md +134 -0
  59. package/skills/information-theory/SKILL.md +140 -242
  60. package/skills/information-theory/original-texts.en.md +127 -0
  61. package/skills/logic-deduction/SKILL.en.md +130 -0
  62. package/skills/logic-deduction/SKILL.md +135 -280
  63. package/skills/logic-deduction/original-texts.en.md +160 -0
  64. package/skills/math-research-activator/SKILL.en.md +132 -0
  65. package/skills/math-research-activator/SKILL.md +136 -0
  66. package/skills/math-research-activator/original-texts.en.md +105 -0
  67. package/skills/{meta-selector → math-research-activator}/original-texts.md +104 -104
  68. package/skills/modeling/SKILL.en.md +135 -0
  69. package/skills/modeling/SKILL.md +139 -318
  70. package/skills/modeling/original-texts.en.md +162 -0
  71. package/skills/optimization/SKILL.en.md +129 -0
  72. package/skills/optimization/SKILL.md +135 -292
  73. package/skills/optimization/original-texts.en.md +167 -0
  74. package/skills/probability-statistics/SKILL.en.md +146 -0
  75. package/skills/probability-statistics/SKILL.md +151 -312
  76. package/skills/probability-statistics/original-texts.en.md +191 -0
  77. package/skills/symmetry-invariance/SKILL.en.md +135 -0
  78. package/skills/symmetry-invariance/SKILL.md +139 -358
  79. package/skills/symmetry-invariance/original-texts.en.md +206 -0
  80. package/skills/topological-thinking/SKILL.en.md +124 -0
  81. package/skills/topological-thinking/SKILL.md +128 -273
  82. package/skills/topological-thinking/original-texts.en.md +134 -0
  83. package/skills/transformation/SKILL.en.md +120 -0
  84. package/skills/transformation/SKILL.md +124 -264
  85. package/skills/transformation/original-texts.en.md +204 -0
  86. package/docs/CLAUDE.md +0 -187
  87. package/skills/meta-selector/SKILL.md +0 -188
@@ -0,0 +1,162 @@
1
+ # Mathematical Sources and Classic Texts
2
+
3
+ ## Newton's *Principia Mathematica* (1687)
4
+
5
+ > "The same effects of nature must always be assigned to the same causes."
6
+
7
+ **The most important achievement in the history of mathematical modeling**: Newton unified celestial and terrestrial motion within a single mathematical framework — the law of universal gravitation F = GMm/r² and the three laws of motion. This was humanity's first systematic use of mathematical models to precisely describe the physical world, marking the birth of scientific modeling. The *Principia* established the methodological paradigm of "mathematical model → physical prediction → experimental verification," which remains the foundation of all modeling work to this day.
8
+
9
+ ---
10
+
11
+ ## Fourier's Heat Equation (1822)
12
+
13
+ > ∂u/∂t = κ ∇²u
14
+
15
+ In *Théorie analytique de la chaleur*, Fourier proposed a mathematical model for heat diffusion and, in doing so, invented Fourier analysis — the method of decomposing arbitrary functions into trigonometric series. This modeling achievement has a twofold significance: it provided a precise mathematical description of heat conduction, and it gave rise to the entire field of harmonic analysis. The Fourier transform remains a core tool in signal processing, quantum mechanics, image compression, and many other domains.
16
+
17
+ **Modeling insight**: The new mathematics invented to solve a modeling problem often has a more profound and lasting impact than the original model itself.
18
+
19
+ ---
20
+
21
+ ## Maxwell's Equations (1865)
22
+
23
+ > ∇·E = ρ/ε₀, ∇×E = -∂B/∂t
24
+ > ∇·B = 0, ∇×B = μ₀J + μ₀ε₀∂E/∂t
25
+
26
+ Maxwell unified electricity and magnetism into four partial differential equations, and from the mathematical model alone predicted the existence of electromagnetic waves — a prediction later confirmed experimentally by Hertz. This is one of the most brilliant examples in the history of modeling: the discovery of an entirely new physical phenomenon through pure mathematical deduction. The equations also imply the speed of light c = 1/√(μ₀ε₀), subsuming optics into electromagnetic theory — arguably the pinnacle of unification in modeling.
27
+
28
+ ---
29
+
30
+ ## Lotka-Volterra Predator-Prey Model (1925-1926)
31
+
32
+ The classical population dynamics model:
33
+
34
+ > dx/dt = αx - βxy (prey growth - predation)
35
+ > dy/dt = δxy - γy (predator growth - natural death)
36
+
37
+ **Modeling insight**: Through two simple differential equations, this model successfully explains the periodic oscillations observed in predator and prey populations in nature. It is the epitome of a "good model" — simple yet useful.
38
+
39
+ ---
40
+
41
+ ## Kermack-McKendrick SIR Epidemic Model (1927)
42
+
43
+ > dS/dt = -βSI (decrease in susceptibles)
44
+ > dI/dt = βSI - γI (change in infected)
45
+ > dR/dt = γI (increase in recovered)
46
+
47
+ **Modeling insight**: The basic reproduction number R₀ = β/γ determines whether an epidemic will take off. Kermack and McKendrick first proposed this model in their 1927 paper "A Contribution to the Mathematical Theory of Epidemics," laying the theoretical foundation for mathematical modeling of infectious diseases. This simple model was widely deployed during the COVID-19 pandemic, demonstrating the enduring value of classical models.
48
+
49
+ ---
50
+
51
+ ## Buckingham Pi Theorem (1914)
52
+
53
+ > If a physical relationship involves n variables with k independent dimensions, the relationship can be reduced to one among (n-k) dimensionless Π groups.
54
+
55
+ **Modeling insight**: The Buckingham Pi theorem formalizes dimensional analysis, an extremely important simplification tool in modeling. It tells us that any physical model can be cast in dimensionless form, thereby reducing the number of parameters and revealing essential structure. For example, the Reynolds number Re = ρvL/μ in fluid mechanics is a single Π quantity that unifies countless seemingly different flow regimes.
56
+
57
+ ---
58
+
59
+ ## Turing's Reaction-Diffusion Model (1952)
60
+
61
+ > ∂u/∂t = D_u ∇²u + f(u,v)
62
+ > ∂v/∂t = D_v ∇²v + g(u,v)
63
+
64
+ In his paper "The Chemical Basis of Morphogenesis," Turing proved that two chemical substances diffusing at different rates and interacting can generate stable spatial patterns — spots, stripes, spirals — from an initially uniform state. This is a mathematical model of morphogenesis, explaining a wide range of phenomena from leopard spots to seashell patterns.
65
+
66
+ **Modeling insight**: Mathematical models can explain "how order emerges from disorder" — no pre-existing pattern is required; patterns arise spontaneously from the dynamics of the equations alone. Turing thereby founded the theory of pattern formation.
67
+
68
+ ---
69
+
70
+ ## Lorenz System (1963)
71
+
72
+ > dx/dt = σ(y - x)
73
+ > dy/dt = x(ρ - y) - xz
74
+ > dz/dt = xy - βz
75
+
76
+ While numerically simulating an atmospheric convection model, Lorenz discovered that tiny differences in initial conditions lead to completely different long-term behavior — the "butterfly effect." Three seemingly simple equations revealed the essential feature of chaos: the long-term unpredictability of deterministic systems.
77
+
78
+ **Modeling insight**: Chaos theory profoundly changed the philosophy of modeling — even if a model is perfectly correct, long-term prediction may be fundamentally impossible. Modelers must distinguish between "predictable timescales" and "unpredictable chaotic regimes."
79
+
80
+ ---
81
+
82
+ ## Kalman Filtering (1960)
83
+
84
+ > x̂_{k|k} = x̂_{k|k-1} + K_k(z_k - H x̂_{k|k-1})
85
+
86
+ The Kalman filter is a recursive estimation method based on state-space models: it uses the system dynamics model to predict the state and then corrects the prediction with observational data. It unifies "model prediction" and "data updating" within a single mathematical framework and is a cornerstone of modern control theory and signal processing.
87
+
88
+ **Modeling insight**: Good modeling is not only about "building a model" but also about "making optimal estimates based on the model." The Kalman filter demonstrates how models and data work together — the model provides the prior, the data provide the correction, and neither is dispensable.
89
+
90
+ ---
91
+
92
+ ## Pólya's *How to Solve It* (1945)
93
+
94
+ > "The first step in mathematical modeling is understanding the problem, the second is devising a plan, the third is carrying out the plan, and the fourth is looking back."
95
+
96
+ Pólya's problem-solving framework is a precursor to modeling thinking: transforming an unfamiliar problem into a known mathematical problem.
97
+
98
+ ---
99
+
100
+ ## Akaike Information Criterion (1974)
101
+
102
+ > AIC = -2 ln(L_max) + 2k
103
+
104
+ In 1974, Akaike proposed AIC, formalizing the model selection problem for the first time: striking an optimal balance between goodness of fit (-2 ln L) and model complexity (2k), where k is the number of parameters and L_max is the maximum likelihood.
105
+
106
+ **Modeling insight**: AIC established a mathematical standard for the "principle of parsimony" — neither the most complex nor the simplest model is best; rather, the optimal trade-off lies in minimizing information loss while keeping parameters parsimonious. This is the quantitative version of Box's dictum that "some are useful."
107
+
108
+ ---
109
+
110
+ ## Black-Scholes Model (1973)
111
+
112
+ > C = S N(d₁) - K e^{-rT} N(d₂)
113
+ > d₁ = [ln(S/K) + (r + σ²/2)T] / (σ√T)
114
+ > d₂ = d₁ - σ√T
115
+
116
+ Black, Scholes, and Merton developed a mathematical model for option pricing. Based on geometric Brownian motion dS = μS dt + σS dW and the no-arbitrage principle, they derived the partial differential equation ∂C/∂t + ½σ²S²∂²C/∂S² + rS∂C/∂S - rC = 0. This model is the most celebrated result in financial mathematics; Scholes and Merton were awarded the 1997 Nobel Prize in Economics for this work.
117
+
118
+ **Modeling insight**: The Black-Scholes model is the best illustration of Box's dictum — its assumptions (constant volatility, continuous trading, frictionless markets) are all "wrong" in reality, yet it provides the core framework for pricing and risk management, and is therefore "useful."
119
+
120
+ ---
121
+
122
+ ## George Box's Famous Quote
123
+
124
+ > "All models are wrong, but some are useful."
125
+
126
+ **Modeling philosophy**:
127
+ - A model is not reality — it is necessarily a simplification of reality
128
+ - The value of a model lies not in its "truthfulness" but in its predictive and explanatory power
129
+ - Criteria for a good model: parsimonious, testable, predictive
130
+
131
+ ---
132
+
133
+ ## General Principles of Modeling
134
+
135
+ 1. **Start simple**: Begin with the simplest model, then incrementally add complexity (Lorenz revealed chaos with three equations; Lotka-Volterra explained oscillations with two)
136
+ 2. **State assumptions explicitly**: Record and test every assumption (Black-Scholes assumptions, though wrong, are explicit — and so the model remains usable)
137
+ 3. **Validate and falsify**: Test models with independent data, not just by fitting (AIC quantifies the risk of overfitting)
138
+ 4. **Know the scope**: Every model has a domain of validity; beyond it, the model fails (Newtonian mechanics fails at high velocities and requires Einstein's correction)
139
+ 5. **Iterate**: Modeling is a cyclical process, not a one-shot endeavor (Pólya's "looking back" step)
140
+ 6. **Unify dimensions**: Use the Buckingham Pi theorem to cast models in dimensionless form, reducing parameters and revealing structure
141
+ 7. **Synergize models and data**: The Kalman filter demonstrates how model prediction and data correction work in a feedback loop
142
+
143
+ ---
144
+
145
+ ## Timeline of Mathematical Modeling
146
+
147
+ | Year | Achievement | Field |
148
+ |------|-------------|-------|
149
+ | 1687 | Newton *Principia* | Mechanics |
150
+ | 1822 | Fourier heat equation | Heat diffusion |
151
+ | 1865 | Maxwell's equations | Electromagnetism |
152
+ | 1914 | Buckingham Pi theorem | Dimensional analysis |
153
+ | 1925-26 | Lotka-Volterra | Population dynamics |
154
+ | 1927 | Kermack-McKendrick SIR | Epidemiology |
155
+ | 1945 | Pólya *How to Solve It* | Methodology |
156
+ | 1952 | Turing reaction-diffusion | Morphogenesis |
157
+ | 1960 | Kalman filter | Estimation & control |
158
+ | 1963 | Lorenz system | Chaos theory |
159
+ | 1973 | Black-Scholes | Finance |
160
+ | 1974 | Akaike AIC | Model selection |
161
+
162
+ This timeline reveals a central pattern: great modeling achievements often transcend disciplinary boundaries. Newton unified the heavens and the earth, Maxwell unified electricity and magnetism, Turing unified chemistry and biology — the power of mathematical models lies precisely in their cross-domain universality.
@@ -0,0 +1,129 @@
1
+ ---
2
+ name: optimization
3
+ description: |
4
+ Trigger when a problem involves resource allocation, trade-offs, maximizing/minimizing objectives, decisions under constraints; or needs convexity analysis, Lagrangian/KKT methods, duality structure; or choosing optimization methods for algorithm/operator/training design.
5
+ ---
6
+
7
+ # Optimization
8
+
9
+ > "Under the most general constraints, find extrema of the objective -- convexity determines difficulty, KKT gives necessity, duality reveals structure."
10
+ >
11
+ > -- Optimization Theory & Operations Research
12
+
13
+ ## Core Principle
14
+
15
+ **Any decision problem can be formulated as an optimization problem: maximizing (or minimizing) an objective subject to constraints. The essence of optimization is not the pursuit of "the best" in the abstract, but "the best among the feasible."**
16
+
17
+ The three core elements of optimization: **Objective**, **Constraints**, **Feasible set**.
18
+
19
+ > **Mathematical Formalization**
20
+ >
21
+ > General optimization problem: $\min_{x \in \mathbb{R}^n} f(x) \quad \text{s.t.} \quad g_i(x) \leq 0,\; i=1,\dots,m; \quad h_j(x) = 0,\; j=1,\dots,p$
22
+ >
23
+ > Lagrangian: $L(x, \lambda, \mu) = f(x) + \sum_i \lambda_i g_i(x) + \sum_j \mu_j h_j(x)$
24
+ >
25
+ > KKT conditions (essential necessary conditions, under Slater-type constraint qualifications): (1) Stationarity $\nabla_x L = 0$; (2) Primal feasibility $g_i \le 0, h_j = 0$; (3) Dual feasibility $\lambda_i \ge 0$; (4) Complementary slackness $\lambda_i g_i = 0$.
26
+ >
27
+ > Convexity: If $f$ and each $g_i$ are convex and each $h_j$ is linear, the problem is convex; in this case **KKT is sufficient**, and local optimality implies global optimality.
28
+
29
+ ## GPU-Friendliness (Cross-Cutting Check)
30
+
31
+ When optimization is used for **algorithm/operator/training design**, the solution method itself must pass the eight-dimensional gate in `../../references/gpu-friendly-math.md`:
32
+
33
+ - **First-order methods (SGD/Adam)**: GEMM-friendly, parallelizable, viable in low precision; pay attention to optimizer state precision and distributed communication overhead.
34
+ - **Second-order / Newton methods**: Hessian inversion $O(n^3)$, memory explosion -- the classic "beautiful but incomputable" case -- adapt to **K-FAC / low-rank / diagonal approximations** (see `../../references/books/optimization-ml.md`, `matrix-analysis.md`).
35
+ - **Constraint projection**: Does the projection admit a closed form and is it tensorizable? Iterative projection requires caution regarding serial dependencies.
36
+ - **Distribution**: Can computation and communication overlap? Is gradient compression needed?
37
+
38
+ Eight-dimensional minimum assessment (formal terms): **Tensorization** -- whether objectives / constraints / gradients can be batched; **GEMM-mappability** -- whether the primary computation is matrix multiplication, HVP, or small K-FAC matrices; **Complexity** -- order of first-order / second-order / combinatorial solvers; **Memory & KV-Cache** -- optimizer states, Hessian storage, activation retention; **Low-precision stability** -- condition numbers, damping, loss scaling; **Parallelism & communication** -- gradient synchronization and communication overlap; **Sparse structure** -- whether preconditioners / constraints are block-structured; **Operator fusion** -- whether updates, clipping, and regularization can be fused.
39
+
40
+ > Cross-reference `../../references/books/optimization-ml.md` (Chong/Lu/Żak) and `../../references/books/matrix-analysis.md`.
41
+
42
+ ## When NOT to Use
43
+
44
+ - **No clear evaluation criterion** (one does not know what "good" means) -- define the objective before optimizing.
45
+ - **Purely execution-oriented tasks** (e.g., formatting code) -- there is no optimization space.
46
+ - **The user has already decided on a plan** -- optimization is already complete.
47
+ - **The problem is essentially qualitative judgment rather than quantitative extremum** -- model first, then optimize.
48
+
49
+ ## When to Use
50
+
51
+ - When one needs to determine whether a problem is convex in order to assess solution difficulty.
52
+ - Choosing optimization methods for algorithm/operator/training design and evaluating their GPU viability.
53
+ - Making rational decisions with quantifiable objectives under constraints.
54
+ - Systematic optimization of experimental design, resource allocation, and hyperparameter / architecture search.
55
+ - When one is uncertain whether the current strategy is optimal and wishes to perform a systematic analysis (convexity, duality, sensitivity).
56
+
57
+ ## Method
58
+
59
+ ### Step 1: Define the Objective
60
+ Clarify what is to be maximized / minimized. Key questions: Single-objective or multi-objective? Is it quantifiable (otherwise, find proxy variables)? Static or dynamic? **Is $f$ convex** (convex implies local = global; non-convex requires vigilance against local extrema)? A wrong objective leads further astray the further one proceeds.
61
+
62
+ ### Step 2: List the Constraints
63
+ Distinguish **hard constraints** (physical / budget / deadline) from **soft constraints** (preferences / quality floors); mathematically classify: inequality $g_i(x)\le 0$ (defines the boundary of the feasible set), equality $h_j(x)=0$ (reduces dimensionality), linear (feasible set is a convex polyhedron) vs. nonlinear (may be non-convex).
64
+
65
+ ### Step 3: Classify the Problem Type
66
+
67
+ | Type | Objective | Constraints | Key Property | Typical Method |
68
+ |------|-----------|-------------|--------------|----------------|
69
+ | LP (Linear Programming) | Linear | Linear inequalities | Optimum at a vertex | Simplex method |
70
+ | QP (Quadratic Programming) | Quadratic | Linear | Positive-definite QP is convex | Interior-point method |
71
+ | Convex optimization | Convex | Convex inequalities + linear equalities | Local = global | Gradient descent, interior-point methods |
72
+ | Non-convex optimization | Non-convex | Arbitrary | Multiple local extrema | Global search, simulated annealing |
73
+ | Combinatorial optimization | Discrete domain | Arbitrary | Frequently NP-hard | Branch-and-bound, heuristics |
74
+ | Stochastic optimization | Contains random terms | May include stochastic constraints | Expected optimum vs. stochastic feasibility | SAA, robust optimization |
75
+
76
+ ### Step 4: Find the Optimal Solution
77
+ - **LP/QP/Convex**: Exploit convexity; gradient-based or interior-point methods guarantee convergence to the global optimum.
78
+ - **Non-convex**: Multi-start strategies, global search, or relaxation to convex approximations.
79
+ - **Combinatorial**: Exact solutions are often NP-hard -- branch-and-bound for small scale, heuristics / approximations for large scale.
80
+ - **Stochastic**: Sample Average Approximation (SAA) converts to a deterministic approximation.
81
+ - **Insufficient information**: A satisficing solution suffices.
82
+
83
+ ### Step 5: Sensitivity Analysis
84
+ The Lagrange multiplier $\lambda_i^*$ is the **shadow price** of the $i$-th constraint -- relaxing the constraint by one unit improves the objective by approximately $\lambda_i^*$. Complementary slackness: $\lambda_i^*=0$ indicates an inactive constraint (no effect on the optimal solution); $\lambda_i^*>0$ indicates an active constraint (the optimal solution lies precisely at its boundary). Focus on: how the optimal solution changes under small perturbations of constraints / objective, and which constraints are active.
85
+
86
+ ### Step 6: Multi-Objective & Pareto
87
+ Multi-objective problems $f_1,\dots,f_k$ generally have no single optimal solution. Pareto optimality: no feasible solution exists that improves all objectives simultaneously. Methods: **Weighted sum** $\min\sum w_i f_i$ (different weights trace different points on the Pareto front); **$\epsilon$-constraint method** $\min f_1$ s.t. $f_i\le\epsilon_i$ (sweep $\epsilon_i$ to cover the front).
88
+
89
+ ### Step 7: Monitor Constraint Changes
90
+ The optimal solution depends on the constraints -- when constraints change, re-optimization is required. Changes in active constraints have the greatest impact (high shadow prices); small changes in inactive constraints typically do not affect the optimal solution.
91
+
92
+ ## Common Errors
93
+
94
+ | Error | Critique | Correct Approach |
95
+ |-------|----------|------------------|
96
+ | Optimizing without a clear objective | Direction is undefined | Precisely define the objective first |
97
+ | Ignoring implicit constraints | The "optimal solution" is actually infeasible | Exhaustively verify all constraints |
98
+ | Getting trapped in local optima | Greedy methods on non-convex problems do not guarantee global optimality | Verify convexity; use multi-start / global methods for non-convex problems |
99
+ | Treating the optimum as unique | The optimal solution may not be unique | Check for the existence of multiple equivalent optima |
100
+ | Using single-objective methods for multi-objective problems | Different objectives require trade-offs | Employ Pareto analysis |
101
+ | Failing to verify convexity | Applying convex methods to non-convex problems | Determine convexity before selecting a method |
102
+ | Ignoring duality theory | The dual problem may be easier to solve | Construct the dual and exploit strong duality |
103
+ | Confusing feasibility with optimality | Feasible does not imply optimal | Verify feasibility first, then verify optimality |
104
+ | Ignoring computational / GPU complexity | Second-order methods / combinatorial optimization may be incomputable | Assess complexity, pass the GPU eight-dimensional gate, approximate when necessary |
105
+ | Forgetting to re-optimize | Failing to update when constraints change | Periodically check for constraint changes |
106
+
107
+ ## Operating Procedure
108
+
109
+ When this skill is triggered, the output must include:
110
+
111
+ 1. **Objective function**: `[Objective]: [Description]` + `[Convexity]: [Convex / Non-convex / Unknown]`
112
+ 2. **Constraint list**: Each labeled `[Hard / Soft]` and `[Inequality / Equality]` `[Linear / Nonlinear]`
113
+ 3. **Problem type classification**: `[Type]: [LP / QP / Convex / Non-convex / Combinatorial / Stochastic]`
114
+ 4. **Feasible set analysis**: Which options are feasible? Which constraints are active?
115
+ 5. **Optimal / satisficing solution**: `[Strategy]: [Gradient method / Interior-point / Global search / Satisficing / Pareto]`
116
+ 6. **Sensitivity analysis**: Shadow prices of key constraints? How do conclusions change under X% variation?
117
+ 7. **GPU viability** (if used for algorithm/operator/training): Does the solution method pass the eight-dimensional gate? Label as friendly / retrofittable / unfriendly, with adaptation recommendations.
118
+ 8. **Action recommendations**: Explicitly state "Next, I will..."
119
+
120
+ **Output must not consist of analysis alone without conclusions.**
121
+
122
+ ## Relations to Other Skills
123
+
124
+ - **Modeling**: Optimization requires prior modeling -- defining objectives and constraints is itself an act of modeling.
125
+ - **Probability and statistics**: Optimization under uncertainty requires stochastic / robust optimization.
126
+ - **Transformation**: Transforming to the dual problem often makes optimization easier; duality is the most profound transformation in optimization.
127
+ - **Game-theoretic thinking**: Multiple decision-makers simultaneously optimizing constitutes a game; the Nash equilibrium is the stable point of multi-player optimization.
128
+ - **Algorithmic thinking**: Solution methods depend on algorithm design -- convex optimization uses gradient methods, combinatorial optimization requires branch-and-bound / heuristics.
129
+ - **Modern mathematics activation**: `../../references/books/optimization-ml.md` (GPU-friendly optimizers, feasibility of second-order methods), `matrix-analysis.md` (condition numbers, low-rank, preconditioning).