pwn 0.5.669 → 0.5.673

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (69) hide show
  1. checksums.yaml +4 -4
  2. data/.rubocop.yml +1 -1
  3. data/Gemfile +1 -1
  4. data/README.md +13 -9
  5. data/documentation/AI-Integration.md +1 -1
  6. data/documentation/Agent-Tool-Registry.md +17 -5
  7. data/documentation/Configuration.md +19 -8
  8. data/documentation/Diagrams.md +2 -2
  9. data/documentation/Home.md +2 -2
  10. data/documentation/How-PWN-Works.md +12 -10
  11. data/documentation/Installation.md +3 -1
  12. data/documentation/Mistakes.md +3 -0
  13. data/documentation/Persistence.md +4 -1
  14. data/documentation/Reinforcement-Learning.md +90 -73
  15. data/documentation/Skills-Memory-Learning.md +28 -4
  16. data/documentation/What-is-PWN.md +8 -7
  17. data/documentation/Why-PWN.md +3 -2
  18. data/documentation/diagrams/agent-tool-registry.svg +188 -160
  19. data/documentation/diagrams/dot/agent-tool-registry.dot +7 -4
  20. data/documentation/diagrams/dot/memory-skills-detailed.dot +12 -5
  21. data/documentation/diagrams/dot/overall-pwn-architecture.dot +4 -3
  22. data/documentation/diagrams/dot/persistence-filesystem.dot +2 -1
  23. data/documentation/diagrams/dot/pwn-ai-feedback-learning-loop.dot +11 -4
  24. data/documentation/diagrams/dot/reinforcement-learning.dot +8 -3
  25. data/documentation/diagrams/dot/task-summarizer.dot +24 -12
  26. data/documentation/diagrams/memory-skills-detailed.svg +252 -210
  27. data/documentation/diagrams/overall-pwn-architecture.svg +21 -12
  28. data/documentation/diagrams/persistence-filesystem.svg +127 -113
  29. data/documentation/diagrams/pwn-ai-feedback-learning-loop.svg +449 -397
  30. data/documentation/diagrams/reinforcement-learning.svg +276 -239
  31. data/documentation/diagrams/task-summarizer.svg +178 -126
  32. data/documentation/pwn-ai-Agent.md +66 -32
  33. data/lib/pwn/ai/agent/curriculum.rb +13 -4
  34. data/lib/pwn/ai/agent/dispatch.rb +3 -0
  35. data/lib/pwn/ai/agent/learning.rb +100 -14
  36. data/lib/pwn/ai/agent/loop.rb +692 -25
  37. data/lib/pwn/ai/agent/metrics.rb +52 -4
  38. data/lib/pwn/ai/agent/mistakes.rb +158 -1
  39. data/lib/pwn/ai/agent/policy.rb +935 -0
  40. data/lib/pwn/ai/agent/prompt_builder.rb +82 -6
  41. data/lib/pwn/ai/agent/reflect.rb +11 -3
  42. data/lib/pwn/ai/agent/registry.rb +18 -3
  43. data/lib/pwn/ai/agent/reward.rb +360 -48
  44. data/lib/pwn/ai/agent/task_summarizer.rb +415 -33
  45. data/lib/pwn/ai/agent/tool_guard.rb +157 -0
  46. data/lib/pwn/ai/agent/tools/policy.rb +76 -0
  47. data/lib/pwn/ai/agent/tools/ruby_eval.rb +27 -1
  48. data/lib/pwn/ai/agent/tools/shell.rb +28 -32
  49. data/lib/pwn/ai/agent.rb +2 -0
  50. data/lib/pwn/config.rb +14 -2
  51. data/lib/pwn/memory.rb +188 -0
  52. data/lib/pwn/sessions.rb +9 -4
  53. data/lib/pwn/version.rb +1 -1
  54. data/spec/integration/reinforced_feedback_loop_spec.rb +35 -9
  55. data/spec/lib/pwn/ai/agent/learning_spec.rb +119 -3
  56. data/spec/lib/pwn/ai/agent/loop_spec.rb +215 -3
  57. data/spec/lib/pwn/ai/agent/metrics_spec.rb +10 -0
  58. data/spec/lib/pwn/ai/agent/mistakes_spec.rb +33 -0
  59. data/spec/lib/pwn/ai/agent/policy_spec.rb +165 -0
  60. data/spec/lib/pwn/ai/agent/prompt_builder_spec.rb +51 -0
  61. data/spec/lib/pwn/ai/agent/reward_spec.rb +75 -0
  62. data/spec/lib/pwn/ai/agent/signal_hygiene_spec.rb +118 -0
  63. data/spec/lib/pwn/ai/agent/task_summarizer_spec.rb +111 -6
  64. data/spec/lib/pwn/ai/agent/tool_guard_spec.rb +61 -0
  65. data/spec/lib/pwn/ai/agent/tools/policy_spec.rb +18 -0
  66. data/spec/lib/pwn/memory_spec.rb +62 -0
  67. data/spec/support/sandbox.rb +2 -0
  68. data/third_party/pwn_rdoc.jsonl +108 -4
  69. metadata +10 -3
@@ -1,51 +1,77 @@
1
1
  # Reinforcement Learning in pwn-ai
2
2
 
3
- pwn-ai runs an in-context learning loop that can also export training data for
4
- a local model. On hosts **with a trainer and GPU**, the path
3
+ pwn-ai learns while you work. Most of that learning stays **in context**: the
4
+ agent writes what happened to disk and puts the useful bits back into the next
5
+ prompt. On a host **with a trainer and a GPU**, the same data can also train a
6
+ local adapter.
5
7
 
6
8
  `Curriculum.practice` → `Reward.export_dpo` → `Curriculum.train_and_gate`
7
9
 
8
- can promote a new LoRA. Without a trainer the same path is **export-ready**
9
- (datasets plus a manual CLI). Live improvement still happens in-context on
10
- every turn.
10
+ That path can promote a new LoRA when the candidate beats the current one.
11
+ Without a trainer it still **exports** the datasets and a manual CLI. Live
12
+ improvement does not wait on weights.
11
13
 
12
14
  ![Reinforcement-learning loop](diagrams/reinforcement-learning.svg)
13
15
 
14
16
  ```
15
17
  +------------------------------------------------+
16
18
  request -----> | Loop.run |
17
- | plan_first -> Curriculum.red_team_plan (S4) |
18
- | Dispatch -> Reward.semantic_ok (R4) |
19
- | -> Mistakes.record(cause:) |
20
- | guard -> Curriculum.counterfactual (S2) |--> preference ledger (W1)
21
- | final -> Curriculum.critic (S3) |
22
- | -> Reward.judge (outcome) (R1) |--> verify_as_reward (E3)
23
- | -> Reward.prm (process) (R2) |--> Sessions[step_reward] (C4)
24
- | -> Curriculum.hindsight (C3) |
25
- | -> Curriculum.calibrate (W3) |--> Metrics.calibration
26
- | -> Reward.sentinel (R3) |--> Mistakes(reward_signal)
19
+ | plan_first -> Curriculum.red_team_plan (S4) |
20
+ | Dispatch -> Reward.semantic_ok (R4) |
21
+ | -> Mistakes.record(cause:) (E1) |
22
+ | guard -> Curriculum.counterfactual (S2) |--> preference ledger (W1)
23
+ | final -> Curriculum.critic (S3) |
24
+ | -> Reward.judge (outcome) (R1) |--> verify_as_reward (E3)
25
+ | -> Reward.prm (process) (R2) |--> Sessions[step_reward] (C4)
26
+ | -> Curriculum.hindsight (C3) |
27
+ | -> Curriculum.calibrate (W3) |--> Metrics.calibration
28
+ | -> Reward.sentinel (R3) |--> Mistakes(reward_signal)
27
29
  +------------------------------------------------+
28
30
  |
29
- Learning.consolidate (M1 semantic merge, M3 importance eviction)
31
+ Learning.consolidate (M1 merge, M3 importance-evict)
30
32
  MemoryIndex.recall_semantic (M2 similarity x recency x importance)
31
- Registry.rank (C1 keyword fit + advantage + UCB)
33
+ Registry.rank (C1 keyword + UCB + Q-advantage)
34
+ Policy (R5 live Q / REINFORCE on a judge-scored MDP)
32
35
  Learning.exemplars_for (C2 prioritized replay, C4 minimal trace)
33
36
  |
34
- nightly cron --> Curriculum.practice (S1) --> Mistakes.resolve --> preference ledger (W1)
37
+ nightly cron --> Curriculum.practice (S1) --> Mistakes.resolve --> preference (W1)
35
38
  weekly cron --> Curriculum.train_and_gate (W2) --> optional LoRA --> A/B gate --> promote
36
39
  |
37
- Extrospection.correlate (E2 world vs self join)
38
- Metrics.changepoints (E1) --> Mistakes(cause: :env_drift)
40
+ Extrospection.correlate (E2 world vs self)
41
+ Metrics.changepoints (E1 CUSUM) --> Mistakes(cause: :env_drift)
39
42
  ```
40
43
 
44
+ ## Live Policy on every turn (R5 · `PWN::AI::Agent::Policy`)
45
+
46
+ This is the live numeric controller. It does not replace planning.
47
+
48
+ | Piece | What it is |
49
+ |---|---|
50
+ | State | request kind, task family, plan quality, answer completeness, usable-result, last action, fail bin, and engine |
51
+ | Action | tool name, or `final` |
52
+ | Step reward | `Reward.semantic_ok` hygiene (`+0.05` / `-0.20`) |
53
+ | Terminal reward | `Reward.judge` score (skipped when the cheap proxy is untrusted and there is no judge) |
54
+ | Updates | Q-learning (`alpha=0.15`, `gamma=0.85`) and REINFORCE (`alpha=0.05`). Stored trajectories replay twice on warmup so a short table is not empty advice. |
55
+ | Budget | Eight finished episodes (live or warmup-credited) unlock greedy suggestions. Until then the prompt omits them. |
56
+ | Steer | Q-advantage in `Registry.rank` once the episode budget is met; keyword fit and CORE_TOOLS still come first |
57
+ | Files | `~/.pwn/policy.json`, `~/.pwn/policy_traj.jsonl` |
58
+ | Tools | `policy_stats` · `policy_evaluate` · `policy_recommend` (inspect only) |
59
+ | Off switch | `ai.agent.policy: false` |
60
+
61
+ Loop calls `begin_episode` before the first rank, `observe_step` after each
62
+ tool, and `finish` from `Learning.auto_introspect` (or Loop if introspect is
63
+ skipped). `Learning.gc_stores` can trim old trajectories without dropping
64
+ high-return / high-score episodes.
65
+
41
66
  ## Reward signal (`PWN::AI::Agent::Reward`)
42
67
 
43
68
  | ID | Method | What it does |
44
69
  |----|--------|--------------|
45
- | **R1** | `.judge` | Outcome score on `(request, final)` → `{score:0..1, verdict:, rationale:, key_step:}`. Replaces brittle success regexes. |
70
+ | **R1** | `.judge` | Cheap LLM outcome score on `(request, final)` → `{score:0..1, verdict:, rationale:, key_step:, source:}`. Calls the active engine `.chat` with a short timeout (default 12s). `Reflect.on` is used only when `module_reflection` is on. Fallback scores completeness, plan cover, claims, and tool-trace echo. Token overlap is only a small on-topic gate. |
46
71
  | **R2** | `.prm` | Process reward - per-tool-step `+1/0/-1` written into `Sessions[:step_reward]`. |
47
72
  | **R3** | `.sentinel` | Compares proxy success rate vs judge mean vs user-correction rate. A large gap fingerprints `reward_signal` so the agent distrusts a lying proxy. |
48
- | **R4** | `.semantic_ok` | Treats informational non-zero exits (e.g. `grep`/`rg` with no match) as benign. Metrics count them as OK; Mistakes only see true dispatch failures. |
73
+ | **R4** | `.semantic_ok` | Treats informational non-zero exits (for example `grep` / `rg` with no match) as benign. Metrics count them as OK; Mistakes only see true dispatch failures. |
74
+ | **R5** | `Policy` (live MDP) | Tabular Q-learning + REINFORCE on real Loop turns. Q-advantage is an advisory `Registry.rank` term and never replaces TaskSummarizer or plan_first. |
49
75
  | - | `.warm_sentinel` | Backfills the sentinel window from scored Learning outcomes so local hosts can engage proxy distrust without waiting for live remote introspect. |
50
76
  | **W1** | `.record_preference` / `.export_dpo` | Preference ledger (`~/.pwn/preferences.jsonl`) from user corrections, resolve, counterfactual, critic, and practice. Caps per source; keeps trajectory-shaped pairs (winning traces / revised answers), not fix commentary. |
51
77
  | - | `.scrub_preferences` / `.preference_balance` / `.generator_mix` | Ledger hygiene and source-mix health so one channel cannot flood preference export. |
@@ -54,9 +80,9 @@ every turn.
54
80
 
55
81
  | ID | Where | What |
56
82
  |----|-------|------|
57
- | **C1** | `Registry.rank` + `Metrics.{ucb,thompson,advantage,prm_advantage}` | Live tool choice blends keyword fit, historical advantage, exploration bonus, and process-reward signal. |
83
+ | **C1** | `Registry.rank` + `Metrics.{ucb,thompson,advantage,prm_advantage}` + `Policy.advantage` | Live tool choice blends keyword fit, historical advantage, exploration bonus, process-reward signal, and (once the episode budget is met) Q(s,a)-V(s). Neighbor states fill cold (s,a) pairs. |
58
84
  | **C2** | `Learning.exemplars_for` | Prior successful traces ranked by judge score, recency, and keyword fit. Low-score "proxy success" rows are dropped when judge distrust is high. |
59
- | **C3** | `Curriculum.hindsight` | On a failed goal, relabel what the trajectory *did* achieve (`success: 'soft'`). Soft rows stay out of hard SFT and are down-weighted in exemplars. |
85
+ | **C3** | `Curriculum.hindsight` | On a failed goal, relabel what the trajectory *did* achieve (`success: 'soft'`). Soft rows stay out of hard supervised export and are down-weighted in exemplars. |
60
86
  | **C4** | `Learning.compress_exemplar` / skill build | Keep steps with positive `step_reward` so few-shot traces stay short. |
61
87
 
62
88
  ## Memory that stays high-signal
@@ -77,9 +103,9 @@ every turn.
77
103
  | **S3** | `.critic` | Constitutional critic of the final answer (can use tools). Under budget pressure, runs text-only so it cannot burn the remaining iterations. |
78
104
  | **S4** | `.red_team_plan` | Adversarial review of the plan-first outline using Metrics / Mistakes / drift. |
79
105
  | - | `.offline_judge` | Score recent sessions with outcome + process judges, warm the sentinel, optional ledger scrub. Meant for nightly cron on local hosts that only introspect failures live. |
80
- | **C3** | `.hindsight` | HER relabel described above. |
106
+ | **C3** | `.hindsight` | Relabel described above. |
81
107
  | **W3** | `.calibrate` | Plan `p(success)=` vs actual outcome → per-engine Brier / overconfidence. Overconfidence can force plan_first + critic and tighten `max_iters`. |
82
- | **W2** | `.train_and_gate` | Export SFT + DPO → optional LoRA train → promote only if the candidate wins on resolved margin, mean judge, smoke set, and a healthy preference diet. Without a trainer: `weight_loop: :export_ready`. |
108
+ | **W2** | `.train_and_gate` | Export supervised + preference data → optional LoRA train → promote only if the candidate wins on resolved margin, mean judge, smoke set, and a healthy preference diet. Without a trainer: `weight_loop: :export_ready`. |
83
109
  | - | `.practice_kpi` | Week-over-week trend of repeating mistakes (outer curriculum health). |
84
110
 
85
111
  ## Budget pressure (iteration ceiling)
@@ -88,7 +114,7 @@ When unresolved `agent_loop` / `assistant_answer` budget-exhaustion fingerprints
88
114
  dominate, the loop marks the budget path hot and tightens the live turn:
89
115
 
90
116
  - lower effective `max_iters` (stricter on local/ollama engines than remote)
91
- - **Last-iter force-final**: tools=nil on the final iteration so a text answer is required
117
+ - last-iter force-final: tools=nil on the final iteration so a text answer is required
92
118
  - skip counterfactual / red-team forks that would spend more tool rounds
93
119
  - still flush TaskSummarizer state and Learning on the exhaust path
94
120
  - end-of-turn critic runs text-only under the same pressure
@@ -96,27 +122,50 @@ dominate, the loop marks the budget path hot and tightens the live turn:
96
122
  Practice prioritizes those scars with short-horizon "finish the task" prompts.
97
123
  Raising `ai.agent.max_iters` or resolving the scar returns normal runway.
98
124
 
125
+ ## Design-priority STATUS (post P14-P25)
126
+
127
+ This table is the flag authority. Stop minting new P-numbers for chores already
128
+ covered. Track these outcomes instead of hunting comments in the source.
129
+
130
+ | Pri | ID | Control | Module(s) | Success criterion |
131
+ |-----|----|---------|-----------|-------------------|
132
+ | **P0** | W1 generator diversity | `Reward::TARGET_SOURCE_MIX` + `generator_mix` + mix-urgent force on critic/counterfactual | `reward.rb`, `curriculum.rb` | `generator_mix.healthy` OR `recommendation` not stuck on `suppress:mistakes_resolve`; trajectory_fraction ≥ 0.5 |
133
+ | **P0** | Introspect budget | `Learning::INTROSPECT_SOFT_MS` / `HARD_MS`; stage skip under soft/hard / `budget_exhaustion_hot?` | `learning.rb` | `auto_introspect` returns `stages_skipped` when over soft; post-answer path cannot re-thrash tool critic |
134
+ | **P1** | Local judge calibration | Heuristic score shrinkage + `confidence`; `Metrics.effective_rate` scales distrust by `judge_confidence` | `reward.rb`, `metrics.rb` | distrust × heuristic no longer fully replaces proxy; local no-trace highs capped |
135
+ | **ops** | Cheap LLM ORM | Direct engine `.chat` for `Reward.judge` / `.prm`; optional `reward_model` / `reward_llm_timeout`; evidence-prior fallback last | `reward.rb` | remote turns grade with ORM even when `module_reflection` is off |
136
+ | **ops** | Outcome-signal haircut | `judge_sample_weight` (ORM 1.0, heuristic 0.25) in sentinel / `Learning.stats` / `Metrics.judge_rate` | `reward.rb`, `learning.rb`, `metrics.rb` | proxy distrust blends toward ORM scores, not bag-of-words overlap |
137
+ | **P1** | Practice outer KPI | `Curriculum.practice_kpi` / `repeating_trend` → `~/.pwn/curriculum_kpi.jsonl` | `curriculum.rb` | week-over-week `delta_repeating` ≤ 0 on budget fingerprints after practice nights |
138
+ | **P2** | PRM sample efficiency | `PRM_MIN_N=5`, shrinkage to `PRM_FULL_N=20`, fleet coverage gate in `Registry.rank` | `metrics.rb`, `registry.rb` | `prm_advantage=0` until n≥5; rank delta=0 until ≥3 tools ready |
139
+ | **P2** | STATUS over flag archaeology | This table | docs | New work cites Pri/ID here, not fresh P26+ comments for the same theme |
140
+ | **ops** | Nightly diet close | `offline_judge` → `scrub_preferences` + `generator_mix` + `practice_kpi` | `curriculum.rb` | Cron path returns `scrub`/`generator_mix`/`practice_kpi`; raw resolve prose does not survive the night |
141
+ | **ops** | Shape backfill | `Reward.infer_shape` + scrub rewrite | `reward.rb` | Legacy shapeless rows get `winning_trace`/`revised_answer` when content warrants; traj_f measurable |
142
+ | **ops** | Mix in prompt | `Metrics.to_context` emits `W1 MIX:` when unhealthy | `metrics.rb` | Unhealthy diet visible every turn without a tool call |
143
+ | **P0** | Budget exhaust deepen | Last-iter force-final (tools=nil); skip CF when `budget_exhaustion_hot?`; tighter caps 24 local / 75 remote; exhaust path `append_session`+`auto_introspect` | `loop.rb` | Exhaust returns a judged final, not a bare string; CF cannot re-enter under hot; last iter cannot tool-call |
144
+
99
145
  ## Intro and extro join
100
146
 
101
- | Where | What |
102
- |-------|------|
103
- | `Metrics.changepoints` + `Loop.attribute_cause` (**E1**) | Env-drift-attributed failures get `cause: :env_drift` and do not inflate `[REPEATING]`. |
104
- | `Extrospection.correlate` (**E2**) | Lead-lag style joins ("tool X started failing after toolchain Y changed"). |
105
- | `Reward.verify_as_reward` (**E3**) | Browser-backed claim checks can floor/cap the outcome score. |
147
+ | ID | Where | What |
148
+ |----|-------|------|
149
+ | **E1** | `Metrics.changepoints` + `Loop.attribute_cause` | Env-drift-attributed failures get `cause: :env_drift` and do not inflate `[REPEATING]`. |
150
+ | **E2** | `Extrospection.correlate` | Lead-lag style joins ("tool X started failing after toolchain Y changed"). |
151
+ | **E3** | `Reward.verify_as_reward` | Browser-backed claim checks can floor/cap the outcome score. |
106
152
 
107
153
  ## Config (`PWN::Env[:ai][:agent]`)
108
154
 
109
155
  ```yaml
110
156
  :ai:
111
- :module_reflection: false # gates Reflect lesson writing (not ORM alone)
157
+ :module_reflection: false # gates Reflect lesson writing (not the judge alone)
112
158
  :agent:
113
159
  :critic: null # S3 - nil = ON for remote engines, OFF for ollama
114
- :red_team_plan: null # S4 - same auto policy
115
- :counterfactual: null # S2 - same auto policy
160
+ :red_team_plan: null # S4 - same auto rule
161
+ :counterfactual: null # S2 - same auto rule
116
162
  :hindsight: true # C3 - hindsight relabel on failed turns (default true)
163
+ :policy: true # R5 - live tabular Q / REINFORCE (advisory rank only)
117
164
  :verify_as_reward: null # E3 - nil = auto sample on claim-shaped answers
118
165
  :reward_llm: null # nil = outcome/process judges use LLM teacher on remote
119
- :local_introspect: :failure_only # ollama cost policy; remote always introspects
166
+ :reward_model: null # optional cheaper model id for ORM/PRM (nil = engine default)
167
+ :reward_llm_timeout: 12 # cheap ORM chat timeout seconds (clamped 2..30)
168
+ :local_introspect: :failure_only # ollama cost rule; remote always introspects
120
169
  :introspect_every_n: 3
121
170
  :max_iters: 25 # hard cap; budget pressure may lower effective value
122
171
  ```
@@ -141,17 +190,18 @@ PWN::Cron.install_defaults
141
190
  `reward_preferences` · `reward_scrub_preferences` · `reward_preference_balance` ·
142
191
  `reward_export_dpo` · `reward_generator_mix` · `curriculum_practice` ·
143
192
  `curriculum_train` · `curriculum_hindsight` · `curriculum_offline_judge` ·
144
- `curriculum_preference_balance` · `curriculum_practice_kpi` ·
145
- `learning_purge_noise`
193
+ `curriculum_preference_balance` · `curriculum_practice_kpi` · `policy_stats` ·
194
+ `policy_evaluate` · `policy_recommend` · `learning_purge_noise`
146
195
 
147
- ## Design claims
196
+ ## What this stack actually does
148
197
 
149
198
  1. Process reward on real security tool traces, not only math demos (**R2**).
150
- 2. Automatic blame attribution: self vs environment drift (**E1** + **E2**).
199
+ 2. Automatic blame: self vs environment drift (**E1** + **E2**).
151
200
  3. Reward-hacking self-detection when proxy success diverges from the judge (**R3**).
152
201
  4. Mistake-driven curriculum with regression-gated LoRA promotion when a trainer exists (**S1** + **W2**).
153
202
  5. Preference pairs from normal agent work (corrections, resolve, critic, practice) with no separate human labelling queue (**W1**).
154
203
  6. Export and promote only when the preference diet is diverse and trajectory-shaped.
204
+ 7. Live tabular Q-learning + REINFORCE on real Loop turns, used only as advice next to planning (**R5**).
155
205
 
156
206
  ## Preference signal quality
157
207
 
@@ -166,39 +216,6 @@ keep it honest by:
166
216
  5. Requiring smoke checks and mean judge improvement before LoRA promote.
167
217
  6. Treating budget exhaustion as a first-class practice target so the agent learns to finish.
168
218
 
169
- ## Design-priority STATUS
170
-
171
- Living checklist for the reinforced feedback loop. Cite the Pri/ID here instead
172
- of inventing new milestone labels for the same theme.
173
-
174
- | Pri | ID | Control | Module(s) | Success criterion |
175
- |-----|----|---------|-----------|-------------------|
176
- | **P0** | W1 generator diversity | `Reward::TARGET_SOURCE_MIX` + `generator_mix` + mix-urgent force on critic/counterfactual | `reward.rb`, `curriculum.rb` | `generator_mix.healthy` OR `recommendation` not stuck on `suppress:mistakes_resolve`; trajectory_fraction ≥ 0.5 |
177
- | **P0** | Introspect budget | `Learning::INTROSPECT_SOFT_MS` / `HARD_MS`; stage skip under soft/hard / budget-pressure mode | `learning.rb` | `auto_introspect` returns `stages_skipped` when over soft; post-answer path cannot re-thrash tool critic |
178
- | **P0** | Local judge calibration | Heuristic score shrinkage + `confidence`; `Metrics.effective_rate` scales distrust by `judge_confidence` | `reward.rb`, `metrics.rb` | distrust×heuristic no longer fully replaces proxy; local no-trace highs capped |
179
- | **P0** | Practice outer KPI | `Curriculum.practice_kpi` / `repeating_trend` → `~/.pwn/curriculum_kpi.jsonl` | `curriculum.rb` | week-over-week `delta_repeating` ≤ 0 on budget fingerprints after practice nights |
180
- | **P0** | PRM sample efficiency | `PRM_MIN_N=5`, shrinkage to `PRM_FULL_N=20`, fleet coverage gate in `Registry.rank` | `metrics.rb`, `registry.rb` | `prm_advantage=0` until n≥5; rank delta=0 until ≥3 tools ready |
181
- | **P0** | STATUS over flag archaeology | This table | docs | New work cites Pri/ID here, not fresh comments for the same theme |
182
- | **ops** | Nightly diet close | `offline_judge` → `scrub_preferences` + `generator_mix` + `practice_kpi` | `curriculum.rb` | Cron path returns `scrub`/`generator_mix`/`practice_kpi`; raw resolve prose does not survive the night |
183
- | **ops** | Shape backfill | `Reward.infer_shape` + scrub rewrite | `reward.rb` | Legacy shapeless rows get `winning_trace`/`revised_answer` when content warrants; traj_f measurable |
184
- | **ops** | Mix in prompt | `Metrics.to_context` emits `W1 MIX:` when unhealthy | `metrics.rb` | Unhealthy diet visible every turn without a tool call |
185
- | **P0** | Budget exhaust deepen | Last-iter force-final (tools=nil); skip CF under budget-pressure mode; hot caps 24 ollama / 75 remote; exhaust path `append_session`+`auto_introspect` | `loop.rb` | Exhaust returns a judged final, not a bare string; CF cannot re-enter under hot; last iter cannot tool-call |
186
-
187
- ### Config additions
188
-
189
- ```yaml
190
- :ai:
191
- :agent:
192
- # Introspect budget (ms wall-clock inside auto_introspect)
193
- # INTROSPECT_SOFT_MS / HARD_MS are constants; override only via code/reload today.
194
- :critic: null # also force-on when generator_mix.urgent includes critic
195
- :counterfactual: null # also force-on when generator_mix.urgent includes counterfactual
196
- ```
197
-
198
- ### Tools (design-priority)
199
-
200
- `reward_generator_mix` · `curriculum_practice_kpi` (plus existing reward/curriculum set)
201
-
202
219
  ---
203
220
 
204
221
  **See also:** [pwn-ai Agent](pwn-ai-Agent.md) · [Mistakes](Mistakes.md) ·
@@ -1,4 +1,4 @@
1
- # Memory · Skills · Learning · Mistakes · Metrics - Introspection
1
+ # Memory · Skills · Learning · Mistakes · Metrics · Policy - Introspection
2
2
 
3
3
  The **inward-facing** half of the pwn-ai feedback loop: how the agent measures
4
4
  its own performance, turns wins into permanent capability, and - critically -
@@ -6,12 +6,13 @@ its own performance, turns wins into permanent capability, and - critically -
6
6
 
7
7
  ![Memory / Skills detail](diagrams/memory-skills-detailed.svg)
8
8
 
9
- ## The five stores
9
+ ## The six stores
10
10
 
11
11
  | Store | File | Write tool | Read tool | Injected as |
12
12
  |---|---|---|---|---|
13
13
  | **Memory** | `memory.json` (+ `memory.idx`) | `memory_remember` | `memory_recall` · `PWN::MemoryIndex.recall_semantic` | `MEMORY` block - durable facts / prefs / lessons / env. **Relevance-ranked** for the current request via a local embedding index when `ai.ollama.embed_model` is available; falls back to newest-first otherwise. |
14
14
  | **Skills** | `skills/<name>/SKILL.md` | `skill_create` · `skill_migrate_legacy` · `learning_distill_skill` | `skill_list` · `skill_view` | `SKILLS` list - reusable procedures + `references:` (CWE/CVE/ATT&CK/NIST/URL) |
15
+ | **Policy** | `policy.json` + `policy_traj.jsonl` | Loop `begin_episode` / `observe_step` / `finish` | `policy_stats` · `policy_evaluate` · `policy_recommend` | `POLICY` block - live tabular Q / REINFORCE. Advisory rank only. Disable with `ai.agent.policy: false`. |
15
16
  | **Learning** | `learning.jsonl` | `learning_note_outcome` · `learning_reflect` | `learning_outcomes` · `learning_stats` · `Learning.exemplars_for` | `LEARNING` block - recent outcomes + success_rate. Prior *successful* traces are also spliced in as **few-shot exemplars** for local models. |
16
17
  | **Mistakes** | `mistakes.json` | `mistakes_record` · `mistakes_resolve` · *auto on failure* | `mistakes_list` | `KNOWN MISTAKES` + `KNOWN FIXES` blocks - do-NOT-repeat + do-THIS-instead |
17
18
  | **Metrics** | `metrics.json` | *automatic* (every Dispatch) | `metrics_summary` | `TOOL EFFECTIVENESS` block - steer tool choice. **Segmented per engine** (`engine=...`) so a local model's telemetry never blends with a frontier model's. |
@@ -33,6 +34,7 @@ wiping durable facts, preferences, and lessons. Clearing is still available via
33
34
  abstract lessons for a small model.
34
35
  3. (local model) Loop.plan_first forces a numbered tool plan BEFORE dispatch.
35
36
  4. Dispatch runs a tool → Metrics.record(tool, ok?, ms, engine:)
37
+ ↳ Policy.observe_step → hygiene reward (semantic_ok) into the live MDP episode
36
38
  ↳ tool FAILED? → Mistakes.record(tool, error) (count++, cross-session)
37
39
  ↳ same sig ≥3×? → guard_repeated_failure + inline correction_hint
38
40
  ↳ (local) ≥ ESCALATE_AFTER_FAILS → Swarm.ask(escalation_persona) → 3-line frontier hint
@@ -40,6 +42,7 @@ wiping durable facts, preferences, and lessons. Clearing is still available via
40
42
  ↳ extro_verify → :refuted → Mistakes.record(tool:'assumption', ...) # proactive
41
43
  ↳ extro_verify → :confirmed → observe(:intel, ttl:30d)
42
44
  6. Final answer produced → Learning.auto_introspect(session_id)
45
+ ↳ Reward.judge (cheap LLM ORM) → Policy.finish (terminal reward; Q + REINFORCE update)
43
46
  ↳ (local) fact_check_local_final → auto extro_verify every CVE/version claim in the answer
44
47
  ↳ if auto_extrospect enabled → Extrospection.auto_extrospect # AUTO_SECTIONS only
45
48
  7. Reflect.on(engine: reflect_engine)→ Memory.remember(lesson_xxxx, ...) # teacher-student: a
@@ -51,7 +54,7 @@ wiping durable facts, preferences, and lessons. Clearing is still available via
51
54
  model via Curriculum.train_and_gate (resolved margin + mean judge + smoke) - the ONLY step that
52
55
  changes weights, not just the scaffold. Without a trainer this stays export-ready.
53
56
  11. Next launch: PromptBuilder injects the budgeted blocks → the model already knows:
54
- MEMORY · SKILLS · LEARNING · KNOWN MISTAKES/FIXES · TOOL EFFECTIVENESS · EXTROSPECTION
57
+ MEMORY · SKILLS · LEARNING · KNOWN MISTAKES/FIXES · TOOL EFFECTIVENESS · POLICY · EXTROSPECTION · RECENT TURNS
55
58
  ```
56
59
 
57
60
  `extro_correlate` is the **join** - it tells the agent whether a failure was
@@ -190,6 +193,7 @@ when deciding which side of the loop to exercise.
190
193
  | "Reflect on this session / extract lessons" | `learning_reflect` / `sessions_current` |
191
194
  | "Don't do that again / that was wrong / resolve..." | `mistakes_record` / `mistakes_resolve` / `mistakes_list` |
192
195
  | "Which tools are unhealthy / avg duration" | `metrics_summary` |
196
+ | "What did the live policy learn / suggest a tool" | `policy_stats` / `policy_evaluate` / `policy_recommend` |
193
197
  | "What did we run in session X / active session" | `sessions_view` / `sessions_current` |
194
198
  | "Disable reflection while we fuzz" | `learning_auto_introspect_toggle` |
195
199
 
@@ -203,8 +207,28 @@ and which procedures to promote permanently.
203
207
  [Sessions](Sessions.md) · [Persistence](Persistence.md)
204
208
 
205
209
 
210
+ ## Live Policy (advisory Q / REINFORCE)
211
+
212
+ When `ai.agent.policy` is on (the default), every Loop turn is one learning
213
+ episode:
214
+
215
+ 1. `begin_episode` opens before the first tool rank.
216
+ 2. `observe_step` writes a hygiene reward after each tool (`semantic_ok`).
217
+ 3. `finish` (from `Learning.auto_introspect`, or Loop if introspect is skipped)
218
+ adds the `Reward.judge` terminal score and updates Q and REINFORCE.
219
+
220
+ State is a short key: request kind, a hash of the active English task, last
221
+ action, fail count bin, and engine. Files:
222
+
223
+ - `~/.pwn/policy.json` - Q table, REINFORCE logits, visit counts, recent returns
224
+ - `~/.pwn/policy_traj.jsonl` - episode log
225
+
226
+ Tools `policy_stats`, `policy_evaluate`, and `policy_recommend` are read-only. Reset is Ruby-only (`PWN::AI::Agent::Policy.reset`) so a tool call cannot wipe the table. `Registry.rank` may add a small Q-advantage after a pair has been visited at least twice. Planning still owns the task list.
227
+
228
+ Turn it off with `ai.agent.policy: false` in `~/.pwn/pwn.yaml`.
229
+
206
230
  ## Task briefs vs learning
207
231
 
208
- `TaskSummarizer` is **UX**, not a persistence layer. Plan/`about_to` lines are ephemeral TUI briefs (deduped by `last_brief_fp`). Durable learning still flows through Learning · Mistakes · Reward · sessions only. See [pwn-ai Agent § Task summaries](pwn-ai-Agent.md#task-summaries-long-autonomous-turns).
232
+ `TaskSummarizer` is **UX**, not a persistence layer. Plan/`about_to` lines are ephemeral TUI briefs (deduped by `last_brief_fp`). Durable learning still flows through Learning · Mistakes · Reward · Policy · sessions. See [pwn-ai Agent § Task summaries](pwn-ai-Agent.md#task-summaries-long-autonomous-turns).
209
233
 
210
234
  [← Home](Home.md)
@@ -20,7 +20,7 @@ with a **tool-calling AI agent** on top that can run the same methods.
20
20
  | `PWN::FFI::*` | **8** | Native DSP/RF backends: Volk · Liquid · FFTW · RTLSdr · HackRF · AdalmPluto · SoapySDR · Stdio |
21
21
  | `PWN::AI::*` | **6** engines | OpenAI, Anthropic, Grok (OAuth device-flow), Gemini, Ollama, Open WebUI |
22
22
  | `bin/pwn_*` | **53** | Headless CLI drivers for CI/CD |
23
- | Agent toolsets | **12** · **82 tools** | terminal · pwn · memory · skills · sessions · learning · metrics · extrospection · cron · swarm · reward · curriculum |
23
+ | Agent toolsets | **13** · **85 tools** | terminal · pwn · memory · skills · sessions · learning · metrics · policy · extrospection · cron · swarm · reward · curriculum |
24
24
 
25
25
  ## Three ways to use it
26
26
 
@@ -36,12 +36,13 @@ with a **tool-calling AI agent** on top that can run the same methods.
36
36
  - **Everything is Ruby, everything is a method.** No YAML DSLs, no black-box
37
37
  plugins. If you can call it in the REPL, the AI agent can call it, a driver
38
38
  can call it, and a cron job can call it.
39
- - **Closed feedback loop.** Metrics + Learning + **Reward** (ORM/PRM judge,
40
- sentinel) + **Curriculum** (mistake-driven self-play, HER, regression-gated
41
- LoRA) on the introspection side; Snapshot + Drift + Intel + Verify on the
42
- extrospection side; joined by `extro_correlate`, which tells the agent
43
- whether a failure was *its* fault or *the world* changed - and writes the
44
- lesson back into the next prompt.
39
+ - **Closed feedback loop.** Metrics + Learning + **Reward** (outcome and process
40
+ judges, sentinel) + **Curriculum** (mistake-driven self-play, hindsight relabel,
41
+ export-ready LoRA gate) + **Policy** (live tabular Q-learning and REINFORCE on
42
+ each turn, advisory rank only) on the introspection side; Snapshot + Drift +
43
+ Intel + Verify on the extrospection side; joined by `extro_correlate`, which
44
+ tells the agent whether a failure was *its* fault or *the world* changed - and
45
+ writes the lesson back into the next prompt.
45
46
  - **Native multi-agent.** `PWN::AI::Agent::Swarm` runs personas (each a full
46
47
  tool-calling agent, optionally on a *different* LLM engine) that debate,
47
48
  broadcast, and share an append-only bus - no IRC daemon, no external service.
@@ -39,8 +39,9 @@ critical when the caller is an autonomous agent.
39
39
 
40
40
  A pentest framework that does not learn repeats the same dead-end scans
41
41
  forever. PWN records **per-tool success rate**, **per-task outcome**, **host
42
- drift**, and **external CVE intel**, then correlates them so tomorrow's run
43
- can start where today's left off. See
42
+ drift**, **external CVE intel**, and a **live Policy** table (Q / REINFORCE
43
+ advice next to planning), then correlates them so tomorrow's run can start
44
+ where today's left off. See
44
45
  [Skills, Memory & Learning](Skills-Memory-Learning.md),
45
46
  [Reinforcement Learning](Reinforcement-Learning.md), and
46
47
  [Extrospection](Extrospection.md).