pwn 0.5.669 → 0.5.673
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.rubocop.yml +1 -1
- data/Gemfile +1 -1
- data/README.md +13 -9
- data/documentation/AI-Integration.md +1 -1
- data/documentation/Agent-Tool-Registry.md +17 -5
- data/documentation/Configuration.md +19 -8
- data/documentation/Diagrams.md +2 -2
- data/documentation/Home.md +2 -2
- data/documentation/How-PWN-Works.md +12 -10
- data/documentation/Installation.md +3 -1
- data/documentation/Mistakes.md +3 -0
- data/documentation/Persistence.md +4 -1
- data/documentation/Reinforcement-Learning.md +90 -73
- data/documentation/Skills-Memory-Learning.md +28 -4
- data/documentation/What-is-PWN.md +8 -7
- data/documentation/Why-PWN.md +3 -2
- data/documentation/diagrams/agent-tool-registry.svg +188 -160
- data/documentation/diagrams/dot/agent-tool-registry.dot +7 -4
- data/documentation/diagrams/dot/memory-skills-detailed.dot +12 -5
- data/documentation/diagrams/dot/overall-pwn-architecture.dot +4 -3
- data/documentation/diagrams/dot/persistence-filesystem.dot +2 -1
- data/documentation/diagrams/dot/pwn-ai-feedback-learning-loop.dot +11 -4
- data/documentation/diagrams/dot/reinforcement-learning.dot +8 -3
- data/documentation/diagrams/dot/task-summarizer.dot +24 -12
- data/documentation/diagrams/memory-skills-detailed.svg +252 -210
- data/documentation/diagrams/overall-pwn-architecture.svg +21 -12
- data/documentation/diagrams/persistence-filesystem.svg +127 -113
- data/documentation/diagrams/pwn-ai-feedback-learning-loop.svg +449 -397
- data/documentation/diagrams/reinforcement-learning.svg +276 -239
- data/documentation/diagrams/task-summarizer.svg +178 -126
- data/documentation/pwn-ai-Agent.md +66 -32
- data/lib/pwn/ai/agent/curriculum.rb +13 -4
- data/lib/pwn/ai/agent/dispatch.rb +3 -0
- data/lib/pwn/ai/agent/learning.rb +100 -14
- data/lib/pwn/ai/agent/loop.rb +692 -25
- data/lib/pwn/ai/agent/metrics.rb +52 -4
- data/lib/pwn/ai/agent/mistakes.rb +158 -1
- data/lib/pwn/ai/agent/policy.rb +935 -0
- data/lib/pwn/ai/agent/prompt_builder.rb +82 -6
- data/lib/pwn/ai/agent/reflect.rb +11 -3
- data/lib/pwn/ai/agent/registry.rb +18 -3
- data/lib/pwn/ai/agent/reward.rb +360 -48
- data/lib/pwn/ai/agent/task_summarizer.rb +415 -33
- data/lib/pwn/ai/agent/tool_guard.rb +157 -0
- data/lib/pwn/ai/agent/tools/policy.rb +76 -0
- data/lib/pwn/ai/agent/tools/ruby_eval.rb +27 -1
- data/lib/pwn/ai/agent/tools/shell.rb +28 -32
- data/lib/pwn/ai/agent.rb +2 -0
- data/lib/pwn/config.rb +14 -2
- data/lib/pwn/memory.rb +188 -0
- data/lib/pwn/sessions.rb +9 -4
- data/lib/pwn/version.rb +1 -1
- data/spec/integration/reinforced_feedback_loop_spec.rb +35 -9
- data/spec/lib/pwn/ai/agent/learning_spec.rb +119 -3
- data/spec/lib/pwn/ai/agent/loop_spec.rb +215 -3
- data/spec/lib/pwn/ai/agent/metrics_spec.rb +10 -0
- data/spec/lib/pwn/ai/agent/mistakes_spec.rb +33 -0
- data/spec/lib/pwn/ai/agent/policy_spec.rb +165 -0
- data/spec/lib/pwn/ai/agent/prompt_builder_spec.rb +51 -0
- data/spec/lib/pwn/ai/agent/reward_spec.rb +75 -0
- data/spec/lib/pwn/ai/agent/signal_hygiene_spec.rb +118 -0
- data/spec/lib/pwn/ai/agent/task_summarizer_spec.rb +111 -6
- data/spec/lib/pwn/ai/agent/tool_guard_spec.rb +61 -0
- data/spec/lib/pwn/ai/agent/tools/policy_spec.rb +18 -0
- data/spec/lib/pwn/memory_spec.rb +62 -0
- data/spec/support/sandbox.rb +2 -0
- data/third_party/pwn_rdoc.jsonl +108 -4
- metadata +10 -3
|
@@ -1,51 +1,77 @@
|
|
|
1
1
|
# Reinforcement Learning in pwn-ai
|
|
2
2
|
|
|
3
|
-
pwn-ai
|
|
4
|
-
|
|
3
|
+
pwn-ai learns while you work. Most of that learning stays **in context**: the
|
|
4
|
+
agent writes what happened to disk and puts the useful bits back into the next
|
|
5
|
+
prompt. On a host **with a trainer and a GPU**, the same data can also train a
|
|
6
|
+
local adapter.
|
|
5
7
|
|
|
6
8
|
`Curriculum.practice` → `Reward.export_dpo` → `Curriculum.train_and_gate`
|
|
7
9
|
|
|
8
|
-
can promote a new LoRA
|
|
9
|
-
|
|
10
|
-
|
|
10
|
+
That path can promote a new LoRA when the candidate beats the current one.
|
|
11
|
+
Without a trainer it still **exports** the datasets and a manual CLI. Live
|
|
12
|
+
improvement does not wait on weights.
|
|
11
13
|
|
|
12
14
|

|
|
13
15
|
|
|
14
16
|
```
|
|
15
17
|
+------------------------------------------------+
|
|
16
18
|
request -----> | Loop.run |
|
|
17
|
-
| plan_first -> Curriculum.red_team_plan
|
|
18
|
-
| Dispatch -> Reward.semantic_ok
|
|
19
|
-
| -> Mistakes.record(cause:)
|
|
20
|
-
| guard -> Curriculum.counterfactual
|
|
21
|
-
| final -> Curriculum.critic
|
|
22
|
-
| -> Reward.judge (outcome)
|
|
23
|
-
| -> Reward.prm (process)
|
|
24
|
-
| -> Curriculum.hindsight
|
|
25
|
-
| -> Curriculum.calibrate
|
|
26
|
-
| -> Reward.sentinel
|
|
19
|
+
| plan_first -> Curriculum.red_team_plan (S4) |
|
|
20
|
+
| Dispatch -> Reward.semantic_ok (R4) |
|
|
21
|
+
| -> Mistakes.record(cause:) (E1) |
|
|
22
|
+
| guard -> Curriculum.counterfactual (S2) |--> preference ledger (W1)
|
|
23
|
+
| final -> Curriculum.critic (S3) |
|
|
24
|
+
| -> Reward.judge (outcome) (R1) |--> verify_as_reward (E3)
|
|
25
|
+
| -> Reward.prm (process) (R2) |--> Sessions[step_reward] (C4)
|
|
26
|
+
| -> Curriculum.hindsight (C3) |
|
|
27
|
+
| -> Curriculum.calibrate (W3) |--> Metrics.calibration
|
|
28
|
+
| -> Reward.sentinel (R3) |--> Mistakes(reward_signal)
|
|
27
29
|
+------------------------------------------------+
|
|
28
30
|
|
|
|
29
|
-
Learning.consolidate (M1
|
|
31
|
+
Learning.consolidate (M1 merge, M3 importance-evict)
|
|
30
32
|
MemoryIndex.recall_semantic (M2 similarity x recency x importance)
|
|
31
|
-
Registry.rank (C1 keyword
|
|
33
|
+
Registry.rank (C1 keyword + UCB + Q-advantage)
|
|
34
|
+
Policy (R5 live Q / REINFORCE on a judge-scored MDP)
|
|
32
35
|
Learning.exemplars_for (C2 prioritized replay, C4 minimal trace)
|
|
33
36
|
|
|
|
34
|
-
nightly cron --> Curriculum.practice (S1) --> Mistakes.resolve --> preference
|
|
37
|
+
nightly cron --> Curriculum.practice (S1) --> Mistakes.resolve --> preference (W1)
|
|
35
38
|
weekly cron --> Curriculum.train_and_gate (W2) --> optional LoRA --> A/B gate --> promote
|
|
36
39
|
|
|
|
37
|
-
Extrospection.correlate (E2 world vs self
|
|
38
|
-
Metrics.changepoints (E1) --> Mistakes(cause: :env_drift)
|
|
40
|
+
Extrospection.correlate (E2 world vs self)
|
|
41
|
+
Metrics.changepoints (E1 CUSUM) --> Mistakes(cause: :env_drift)
|
|
39
42
|
```
|
|
40
43
|
|
|
44
|
+
## Live Policy on every turn (R5 · `PWN::AI::Agent::Policy`)
|
|
45
|
+
|
|
46
|
+
This is the live numeric controller. It does not replace planning.
|
|
47
|
+
|
|
48
|
+
| Piece | What it is |
|
|
49
|
+
|---|---|
|
|
50
|
+
| State | request kind, task family, plan quality, answer completeness, usable-result, last action, fail bin, and engine |
|
|
51
|
+
| Action | tool name, or `final` |
|
|
52
|
+
| Step reward | `Reward.semantic_ok` hygiene (`+0.05` / `-0.20`) |
|
|
53
|
+
| Terminal reward | `Reward.judge` score (skipped when the cheap proxy is untrusted and there is no judge) |
|
|
54
|
+
| Updates | Q-learning (`alpha=0.15`, `gamma=0.85`) and REINFORCE (`alpha=0.05`). Stored trajectories replay twice on warmup so a short table is not empty advice. |
|
|
55
|
+
| Budget | Eight finished episodes (live or warmup-credited) unlock greedy suggestions. Until then the prompt omits them. |
|
|
56
|
+
| Steer | Q-advantage in `Registry.rank` once the episode budget is met; keyword fit and CORE_TOOLS still come first |
|
|
57
|
+
| Files | `~/.pwn/policy.json`, `~/.pwn/policy_traj.jsonl` |
|
|
58
|
+
| Tools | `policy_stats` · `policy_evaluate` · `policy_recommend` (inspect only) |
|
|
59
|
+
| Off switch | `ai.agent.policy: false` |
|
|
60
|
+
|
|
61
|
+
Loop calls `begin_episode` before the first rank, `observe_step` after each
|
|
62
|
+
tool, and `finish` from `Learning.auto_introspect` (or Loop if introspect is
|
|
63
|
+
skipped). `Learning.gc_stores` can trim old trajectories without dropping
|
|
64
|
+
high-return / high-score episodes.
|
|
65
|
+
|
|
41
66
|
## Reward signal (`PWN::AI::Agent::Reward`)
|
|
42
67
|
|
|
43
68
|
| ID | Method | What it does |
|
|
44
69
|
|----|--------|--------------|
|
|
45
|
-
| **R1** | `.judge` |
|
|
70
|
+
| **R1** | `.judge` | Cheap LLM outcome score on `(request, final)` → `{score:0..1, verdict:, rationale:, key_step:, source:}`. Calls the active engine `.chat` with a short timeout (default 12s). `Reflect.on` is used only when `module_reflection` is on. Fallback scores completeness, plan cover, claims, and tool-trace echo. Token overlap is only a small on-topic gate. |
|
|
46
71
|
| **R2** | `.prm` | Process reward - per-tool-step `+1/0/-1` written into `Sessions[:step_reward]`. |
|
|
47
72
|
| **R3** | `.sentinel` | Compares proxy success rate vs judge mean vs user-correction rate. A large gap fingerprints `reward_signal` so the agent distrusts a lying proxy. |
|
|
48
|
-
| **R4** | `.semantic_ok` | Treats informational non-zero exits (
|
|
73
|
+
| **R4** | `.semantic_ok` | Treats informational non-zero exits (for example `grep` / `rg` with no match) as benign. Metrics count them as OK; Mistakes only see true dispatch failures. |
|
|
74
|
+
| **R5** | `Policy` (live MDP) | Tabular Q-learning + REINFORCE on real Loop turns. Q-advantage is an advisory `Registry.rank` term and never replaces TaskSummarizer or plan_first. |
|
|
49
75
|
| - | `.warm_sentinel` | Backfills the sentinel window from scored Learning outcomes so local hosts can engage proxy distrust without waiting for live remote introspect. |
|
|
50
76
|
| **W1** | `.record_preference` / `.export_dpo` | Preference ledger (`~/.pwn/preferences.jsonl`) from user corrections, resolve, counterfactual, critic, and practice. Caps per source; keeps trajectory-shaped pairs (winning traces / revised answers), not fix commentary. |
|
|
51
77
|
| - | `.scrub_preferences` / `.preference_balance` / `.generator_mix` | Ledger hygiene and source-mix health so one channel cannot flood preference export. |
|
|
@@ -54,9 +80,9 @@ every turn.
|
|
|
54
80
|
|
|
55
81
|
| ID | Where | What |
|
|
56
82
|
|----|-------|------|
|
|
57
|
-
| **C1** | `Registry.rank` + `Metrics.{ucb,thompson,advantage,prm_advantage}` | Live tool choice blends keyword fit, historical advantage, exploration bonus,
|
|
83
|
+
| **C1** | `Registry.rank` + `Metrics.{ucb,thompson,advantage,prm_advantage}` + `Policy.advantage` | Live tool choice blends keyword fit, historical advantage, exploration bonus, process-reward signal, and (once the episode budget is met) Q(s,a)-V(s). Neighbor states fill cold (s,a) pairs. |
|
|
58
84
|
| **C2** | `Learning.exemplars_for` | Prior successful traces ranked by judge score, recency, and keyword fit. Low-score "proxy success" rows are dropped when judge distrust is high. |
|
|
59
|
-
| **C3** | `Curriculum.hindsight` | On a failed goal, relabel what the trajectory *did* achieve (`success: 'soft'`). Soft rows stay out of hard
|
|
85
|
+
| **C3** | `Curriculum.hindsight` | On a failed goal, relabel what the trajectory *did* achieve (`success: 'soft'`). Soft rows stay out of hard supervised export and are down-weighted in exemplars. |
|
|
60
86
|
| **C4** | `Learning.compress_exemplar` / skill build | Keep steps with positive `step_reward` so few-shot traces stay short. |
|
|
61
87
|
|
|
62
88
|
## Memory that stays high-signal
|
|
@@ -77,9 +103,9 @@ every turn.
|
|
|
77
103
|
| **S3** | `.critic` | Constitutional critic of the final answer (can use tools). Under budget pressure, runs text-only so it cannot burn the remaining iterations. |
|
|
78
104
|
| **S4** | `.red_team_plan` | Adversarial review of the plan-first outline using Metrics / Mistakes / drift. |
|
|
79
105
|
| - | `.offline_judge` | Score recent sessions with outcome + process judges, warm the sentinel, optional ledger scrub. Meant for nightly cron on local hosts that only introspect failures live. |
|
|
80
|
-
| **C3** | `.hindsight` |
|
|
106
|
+
| **C3** | `.hindsight` | Relabel described above. |
|
|
81
107
|
| **W3** | `.calibrate` | Plan `p(success)=` vs actual outcome → per-engine Brier / overconfidence. Overconfidence can force plan_first + critic and tighten `max_iters`. |
|
|
82
|
-
| **W2** | `.train_and_gate` | Export
|
|
108
|
+
| **W2** | `.train_and_gate` | Export supervised + preference data → optional LoRA train → promote only if the candidate wins on resolved margin, mean judge, smoke set, and a healthy preference diet. Without a trainer: `weight_loop: :export_ready`. |
|
|
83
109
|
| - | `.practice_kpi` | Week-over-week trend of repeating mistakes (outer curriculum health). |
|
|
84
110
|
|
|
85
111
|
## Budget pressure (iteration ceiling)
|
|
@@ -88,7 +114,7 @@ When unresolved `agent_loop` / `assistant_answer` budget-exhaustion fingerprints
|
|
|
88
114
|
dominate, the loop marks the budget path hot and tightens the live turn:
|
|
89
115
|
|
|
90
116
|
- lower effective `max_iters` (stricter on local/ollama engines than remote)
|
|
91
|
-
-
|
|
117
|
+
- last-iter force-final: tools=nil on the final iteration so a text answer is required
|
|
92
118
|
- skip counterfactual / red-team forks that would spend more tool rounds
|
|
93
119
|
- still flush TaskSummarizer state and Learning on the exhaust path
|
|
94
120
|
- end-of-turn critic runs text-only under the same pressure
|
|
@@ -96,27 +122,50 @@ dominate, the loop marks the budget path hot and tightens the live turn:
|
|
|
96
122
|
Practice prioritizes those scars with short-horizon "finish the task" prompts.
|
|
97
123
|
Raising `ai.agent.max_iters` or resolving the scar returns normal runway.
|
|
98
124
|
|
|
125
|
+
## Design-priority STATUS (post P14-P25)
|
|
126
|
+
|
|
127
|
+
This table is the flag authority. Stop minting new P-numbers for chores already
|
|
128
|
+
covered. Track these outcomes instead of hunting comments in the source.
|
|
129
|
+
|
|
130
|
+
| Pri | ID | Control | Module(s) | Success criterion |
|
|
131
|
+
|-----|----|---------|-----------|-------------------|
|
|
132
|
+
| **P0** | W1 generator diversity | `Reward::TARGET_SOURCE_MIX` + `generator_mix` + mix-urgent force on critic/counterfactual | `reward.rb`, `curriculum.rb` | `generator_mix.healthy` OR `recommendation` not stuck on `suppress:mistakes_resolve`; trajectory_fraction ≥ 0.5 |
|
|
133
|
+
| **P0** | Introspect budget | `Learning::INTROSPECT_SOFT_MS` / `HARD_MS`; stage skip under soft/hard / `budget_exhaustion_hot?` | `learning.rb` | `auto_introspect` returns `stages_skipped` when over soft; post-answer path cannot re-thrash tool critic |
|
|
134
|
+
| **P1** | Local judge calibration | Heuristic score shrinkage + `confidence`; `Metrics.effective_rate` scales distrust by `judge_confidence` | `reward.rb`, `metrics.rb` | distrust × heuristic no longer fully replaces proxy; local no-trace highs capped |
|
|
135
|
+
| **ops** | Cheap LLM ORM | Direct engine `.chat` for `Reward.judge` / `.prm`; optional `reward_model` / `reward_llm_timeout`; evidence-prior fallback last | `reward.rb` | remote turns grade with ORM even when `module_reflection` is off |
|
|
136
|
+
| **ops** | Outcome-signal haircut | `judge_sample_weight` (ORM 1.0, heuristic 0.25) in sentinel / `Learning.stats` / `Metrics.judge_rate` | `reward.rb`, `learning.rb`, `metrics.rb` | proxy distrust blends toward ORM scores, not bag-of-words overlap |
|
|
137
|
+
| **P1** | Practice outer KPI | `Curriculum.practice_kpi` / `repeating_trend` → `~/.pwn/curriculum_kpi.jsonl` | `curriculum.rb` | week-over-week `delta_repeating` ≤ 0 on budget fingerprints after practice nights |
|
|
138
|
+
| **P2** | PRM sample efficiency | `PRM_MIN_N=5`, shrinkage to `PRM_FULL_N=20`, fleet coverage gate in `Registry.rank` | `metrics.rb`, `registry.rb` | `prm_advantage=0` until n≥5; rank delta=0 until ≥3 tools ready |
|
|
139
|
+
| **P2** | STATUS over flag archaeology | This table | docs | New work cites Pri/ID here, not fresh P26+ comments for the same theme |
|
|
140
|
+
| **ops** | Nightly diet close | `offline_judge` → `scrub_preferences` + `generator_mix` + `practice_kpi` | `curriculum.rb` | Cron path returns `scrub`/`generator_mix`/`practice_kpi`; raw resolve prose does not survive the night |
|
|
141
|
+
| **ops** | Shape backfill | `Reward.infer_shape` + scrub rewrite | `reward.rb` | Legacy shapeless rows get `winning_trace`/`revised_answer` when content warrants; traj_f measurable |
|
|
142
|
+
| **ops** | Mix in prompt | `Metrics.to_context` emits `W1 MIX:` when unhealthy | `metrics.rb` | Unhealthy diet visible every turn without a tool call |
|
|
143
|
+
| **P0** | Budget exhaust deepen | Last-iter force-final (tools=nil); skip CF when `budget_exhaustion_hot?`; tighter caps 24 local / 75 remote; exhaust path `append_session`+`auto_introspect` | `loop.rb` | Exhaust returns a judged final, not a bare string; CF cannot re-enter under hot; last iter cannot tool-call |
|
|
144
|
+
|
|
99
145
|
## Intro and extro join
|
|
100
146
|
|
|
101
|
-
| Where | What |
|
|
102
|
-
|
|
103
|
-
| `Metrics.changepoints` + `Loop.attribute_cause`
|
|
104
|
-
| `Extrospection.correlate`
|
|
105
|
-
| `Reward.verify_as_reward`
|
|
147
|
+
| ID | Where | What |
|
|
148
|
+
|----|-------|------|
|
|
149
|
+
| **E1** | `Metrics.changepoints` + `Loop.attribute_cause` | Env-drift-attributed failures get `cause: :env_drift` and do not inflate `[REPEATING]`. |
|
|
150
|
+
| **E2** | `Extrospection.correlate` | Lead-lag style joins ("tool X started failing after toolchain Y changed"). |
|
|
151
|
+
| **E3** | `Reward.verify_as_reward` | Browser-backed claim checks can floor/cap the outcome score. |
|
|
106
152
|
|
|
107
153
|
## Config (`PWN::Env[:ai][:agent]`)
|
|
108
154
|
|
|
109
155
|
```yaml
|
|
110
156
|
:ai:
|
|
111
|
-
:module_reflection: false # gates Reflect lesson writing (not
|
|
157
|
+
:module_reflection: false # gates Reflect lesson writing (not the judge alone)
|
|
112
158
|
:agent:
|
|
113
159
|
:critic: null # S3 - nil = ON for remote engines, OFF for ollama
|
|
114
|
-
:red_team_plan: null # S4 - same auto
|
|
115
|
-
:counterfactual: null # S2 - same auto
|
|
160
|
+
:red_team_plan: null # S4 - same auto rule
|
|
161
|
+
:counterfactual: null # S2 - same auto rule
|
|
116
162
|
:hindsight: true # C3 - hindsight relabel on failed turns (default true)
|
|
163
|
+
:policy: true # R5 - live tabular Q / REINFORCE (advisory rank only)
|
|
117
164
|
:verify_as_reward: null # E3 - nil = auto sample on claim-shaped answers
|
|
118
165
|
:reward_llm: null # nil = outcome/process judges use LLM teacher on remote
|
|
119
|
-
:
|
|
166
|
+
:reward_model: null # optional cheaper model id for ORM/PRM (nil = engine default)
|
|
167
|
+
:reward_llm_timeout: 12 # cheap ORM chat timeout seconds (clamped 2..30)
|
|
168
|
+
:local_introspect: :failure_only # ollama cost rule; remote always introspects
|
|
120
169
|
:introspect_every_n: 3
|
|
121
170
|
:max_iters: 25 # hard cap; budget pressure may lower effective value
|
|
122
171
|
```
|
|
@@ -141,17 +190,18 @@ PWN::Cron.install_defaults
|
|
|
141
190
|
`reward_preferences` · `reward_scrub_preferences` · `reward_preference_balance` ·
|
|
142
191
|
`reward_export_dpo` · `reward_generator_mix` · `curriculum_practice` ·
|
|
143
192
|
`curriculum_train` · `curriculum_hindsight` · `curriculum_offline_judge` ·
|
|
144
|
-
`curriculum_preference_balance` · `curriculum_practice_kpi` ·
|
|
145
|
-
`learning_purge_noise`
|
|
193
|
+
`curriculum_preference_balance` · `curriculum_practice_kpi` · `policy_stats` ·
|
|
194
|
+
`policy_evaluate` · `policy_recommend` · `learning_purge_noise`
|
|
146
195
|
|
|
147
|
-
##
|
|
196
|
+
## What this stack actually does
|
|
148
197
|
|
|
149
198
|
1. Process reward on real security tool traces, not only math demos (**R2**).
|
|
150
|
-
2. Automatic blame
|
|
199
|
+
2. Automatic blame: self vs environment drift (**E1** + **E2**).
|
|
151
200
|
3. Reward-hacking self-detection when proxy success diverges from the judge (**R3**).
|
|
152
201
|
4. Mistake-driven curriculum with regression-gated LoRA promotion when a trainer exists (**S1** + **W2**).
|
|
153
202
|
5. Preference pairs from normal agent work (corrections, resolve, critic, practice) with no separate human labelling queue (**W1**).
|
|
154
203
|
6. Export and promote only when the preference diet is diverse and trajectory-shaped.
|
|
204
|
+
7. Live tabular Q-learning + REINFORCE on real Loop turns, used only as advice next to planning (**R5**).
|
|
155
205
|
|
|
156
206
|
## Preference signal quality
|
|
157
207
|
|
|
@@ -166,39 +216,6 @@ keep it honest by:
|
|
|
166
216
|
5. Requiring smoke checks and mean judge improvement before LoRA promote.
|
|
167
217
|
6. Treating budget exhaustion as a first-class practice target so the agent learns to finish.
|
|
168
218
|
|
|
169
|
-
## Design-priority STATUS
|
|
170
|
-
|
|
171
|
-
Living checklist for the reinforced feedback loop. Cite the Pri/ID here instead
|
|
172
|
-
of inventing new milestone labels for the same theme.
|
|
173
|
-
|
|
174
|
-
| Pri | ID | Control | Module(s) | Success criterion |
|
|
175
|
-
|-----|----|---------|-----------|-------------------|
|
|
176
|
-
| **P0** | W1 generator diversity | `Reward::TARGET_SOURCE_MIX` + `generator_mix` + mix-urgent force on critic/counterfactual | `reward.rb`, `curriculum.rb` | `generator_mix.healthy` OR `recommendation` not stuck on `suppress:mistakes_resolve`; trajectory_fraction ≥ 0.5 |
|
|
177
|
-
| **P0** | Introspect budget | `Learning::INTROSPECT_SOFT_MS` / `HARD_MS`; stage skip under soft/hard / budget-pressure mode | `learning.rb` | `auto_introspect` returns `stages_skipped` when over soft; post-answer path cannot re-thrash tool critic |
|
|
178
|
-
| **P0** | Local judge calibration | Heuristic score shrinkage + `confidence`; `Metrics.effective_rate` scales distrust by `judge_confidence` | `reward.rb`, `metrics.rb` | distrust×heuristic no longer fully replaces proxy; local no-trace highs capped |
|
|
179
|
-
| **P0** | Practice outer KPI | `Curriculum.practice_kpi` / `repeating_trend` → `~/.pwn/curriculum_kpi.jsonl` | `curriculum.rb` | week-over-week `delta_repeating` ≤ 0 on budget fingerprints after practice nights |
|
|
180
|
-
| **P0** | PRM sample efficiency | `PRM_MIN_N=5`, shrinkage to `PRM_FULL_N=20`, fleet coverage gate in `Registry.rank` | `metrics.rb`, `registry.rb` | `prm_advantage=0` until n≥5; rank delta=0 until ≥3 tools ready |
|
|
181
|
-
| **P0** | STATUS over flag archaeology | This table | docs | New work cites Pri/ID here, not fresh comments for the same theme |
|
|
182
|
-
| **ops** | Nightly diet close | `offline_judge` → `scrub_preferences` + `generator_mix` + `practice_kpi` | `curriculum.rb` | Cron path returns `scrub`/`generator_mix`/`practice_kpi`; raw resolve prose does not survive the night |
|
|
183
|
-
| **ops** | Shape backfill | `Reward.infer_shape` + scrub rewrite | `reward.rb` | Legacy shapeless rows get `winning_trace`/`revised_answer` when content warrants; traj_f measurable |
|
|
184
|
-
| **ops** | Mix in prompt | `Metrics.to_context` emits `W1 MIX:` when unhealthy | `metrics.rb` | Unhealthy diet visible every turn without a tool call |
|
|
185
|
-
| **P0** | Budget exhaust deepen | Last-iter force-final (tools=nil); skip CF under budget-pressure mode; hot caps 24 ollama / 75 remote; exhaust path `append_session`+`auto_introspect` | `loop.rb` | Exhaust returns a judged final, not a bare string; CF cannot re-enter under hot; last iter cannot tool-call |
|
|
186
|
-
|
|
187
|
-
### Config additions
|
|
188
|
-
|
|
189
|
-
```yaml
|
|
190
|
-
:ai:
|
|
191
|
-
:agent:
|
|
192
|
-
# Introspect budget (ms wall-clock inside auto_introspect)
|
|
193
|
-
# INTROSPECT_SOFT_MS / HARD_MS are constants; override only via code/reload today.
|
|
194
|
-
:critic: null # also force-on when generator_mix.urgent includes critic
|
|
195
|
-
:counterfactual: null # also force-on when generator_mix.urgent includes counterfactual
|
|
196
|
-
```
|
|
197
|
-
|
|
198
|
-
### Tools (design-priority)
|
|
199
|
-
|
|
200
|
-
`reward_generator_mix` · `curriculum_practice_kpi` (plus existing reward/curriculum set)
|
|
201
|
-
|
|
202
219
|
---
|
|
203
220
|
|
|
204
221
|
**See also:** [pwn-ai Agent](pwn-ai-Agent.md) · [Mistakes](Mistakes.md) ·
|
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
# Memory · Skills · Learning · Mistakes · Metrics - Introspection
|
|
1
|
+
# Memory · Skills · Learning · Mistakes · Metrics · Policy - Introspection
|
|
2
2
|
|
|
3
3
|
The **inward-facing** half of the pwn-ai feedback loop: how the agent measures
|
|
4
4
|
its own performance, turns wins into permanent capability, and - critically -
|
|
@@ -6,12 +6,13 @@ its own performance, turns wins into permanent capability, and - critically -
|
|
|
6
6
|
|
|
7
7
|

|
|
8
8
|
|
|
9
|
-
## The
|
|
9
|
+
## The six stores
|
|
10
10
|
|
|
11
11
|
| Store | File | Write tool | Read tool | Injected as |
|
|
12
12
|
|---|---|---|---|---|
|
|
13
13
|
| **Memory** | `memory.json` (+ `memory.idx`) | `memory_remember` | `memory_recall` · `PWN::MemoryIndex.recall_semantic` | `MEMORY` block - durable facts / prefs / lessons / env. **Relevance-ranked** for the current request via a local embedding index when `ai.ollama.embed_model` is available; falls back to newest-first otherwise. |
|
|
14
14
|
| **Skills** | `skills/<name>/SKILL.md` | `skill_create` · `skill_migrate_legacy` · `learning_distill_skill` | `skill_list` · `skill_view` | `SKILLS` list - reusable procedures + `references:` (CWE/CVE/ATT&CK/NIST/URL) |
|
|
15
|
+
| **Policy** | `policy.json` + `policy_traj.jsonl` | Loop `begin_episode` / `observe_step` / `finish` | `policy_stats` · `policy_evaluate` · `policy_recommend` | `POLICY` block - live tabular Q / REINFORCE. Advisory rank only. Disable with `ai.agent.policy: false`. |
|
|
15
16
|
| **Learning** | `learning.jsonl` | `learning_note_outcome` · `learning_reflect` | `learning_outcomes` · `learning_stats` · `Learning.exemplars_for` | `LEARNING` block - recent outcomes + success_rate. Prior *successful* traces are also spliced in as **few-shot exemplars** for local models. |
|
|
16
17
|
| **Mistakes** | `mistakes.json` | `mistakes_record` · `mistakes_resolve` · *auto on failure* | `mistakes_list` | `KNOWN MISTAKES` + `KNOWN FIXES` blocks - do-NOT-repeat + do-THIS-instead |
|
|
17
18
|
| **Metrics** | `metrics.json` | *automatic* (every Dispatch) | `metrics_summary` | `TOOL EFFECTIVENESS` block - steer tool choice. **Segmented per engine** (`engine=...`) so a local model's telemetry never blends with a frontier model's. |
|
|
@@ -33,6 +34,7 @@ wiping durable facts, preferences, and lessons. Clearing is still available via
|
|
|
33
34
|
abstract lessons for a small model.
|
|
34
35
|
3. (local model) Loop.plan_first forces a numbered tool plan BEFORE dispatch.
|
|
35
36
|
4. Dispatch runs a tool → Metrics.record(tool, ok?, ms, engine:)
|
|
37
|
+
↳ Policy.observe_step → hygiene reward (semantic_ok) into the live MDP episode
|
|
36
38
|
↳ tool FAILED? → Mistakes.record(tool, error) (count++, cross-session)
|
|
37
39
|
↳ same sig ≥3×? → guard_repeated_failure + inline correction_hint
|
|
38
40
|
↳ (local) ≥ ESCALATE_AFTER_FAILS → Swarm.ask(escalation_persona) → 3-line frontier hint
|
|
@@ -40,6 +42,7 @@ wiping durable facts, preferences, and lessons. Clearing is still available via
|
|
|
40
42
|
↳ extro_verify → :refuted → Mistakes.record(tool:'assumption', ...) # proactive
|
|
41
43
|
↳ extro_verify → :confirmed → observe(:intel, ttl:30d)
|
|
42
44
|
6. Final answer produced → Learning.auto_introspect(session_id)
|
|
45
|
+
↳ Reward.judge (cheap LLM ORM) → Policy.finish (terminal reward; Q + REINFORCE update)
|
|
43
46
|
↳ (local) fact_check_local_final → auto extro_verify every CVE/version claim in the answer
|
|
44
47
|
↳ if auto_extrospect enabled → Extrospection.auto_extrospect # AUTO_SECTIONS only
|
|
45
48
|
7. Reflect.on(engine: reflect_engine)→ Memory.remember(lesson_xxxx, ...) # teacher-student: a
|
|
@@ -51,7 +54,7 @@ wiping durable facts, preferences, and lessons. Clearing is still available via
|
|
|
51
54
|
model via Curriculum.train_and_gate (resolved margin + mean judge + smoke) - the ONLY step that
|
|
52
55
|
changes weights, not just the scaffold. Without a trainer this stays export-ready.
|
|
53
56
|
11. Next launch: PromptBuilder injects the budgeted blocks → the model already knows:
|
|
54
|
-
MEMORY · SKILLS · LEARNING · KNOWN MISTAKES/FIXES · TOOL EFFECTIVENESS · EXTROSPECTION
|
|
57
|
+
MEMORY · SKILLS · LEARNING · KNOWN MISTAKES/FIXES · TOOL EFFECTIVENESS · POLICY · EXTROSPECTION · RECENT TURNS
|
|
55
58
|
```
|
|
56
59
|
|
|
57
60
|
`extro_correlate` is the **join** - it tells the agent whether a failure was
|
|
@@ -190,6 +193,7 @@ when deciding which side of the loop to exercise.
|
|
|
190
193
|
| "Reflect on this session / extract lessons" | `learning_reflect` / `sessions_current` |
|
|
191
194
|
| "Don't do that again / that was wrong / resolve..." | `mistakes_record` / `mistakes_resolve` / `mistakes_list` |
|
|
192
195
|
| "Which tools are unhealthy / avg duration" | `metrics_summary` |
|
|
196
|
+
| "What did the live policy learn / suggest a tool" | `policy_stats` / `policy_evaluate` / `policy_recommend` |
|
|
193
197
|
| "What did we run in session X / active session" | `sessions_view` / `sessions_current` |
|
|
194
198
|
| "Disable reflection while we fuzz" | `learning_auto_introspect_toggle` |
|
|
195
199
|
|
|
@@ -203,8 +207,28 @@ and which procedures to promote permanently.
|
|
|
203
207
|
[Sessions](Sessions.md) · [Persistence](Persistence.md)
|
|
204
208
|
|
|
205
209
|
|
|
210
|
+
## Live Policy (advisory Q / REINFORCE)
|
|
211
|
+
|
|
212
|
+
When `ai.agent.policy` is on (the default), every Loop turn is one learning
|
|
213
|
+
episode:
|
|
214
|
+
|
|
215
|
+
1. `begin_episode` opens before the first tool rank.
|
|
216
|
+
2. `observe_step` writes a hygiene reward after each tool (`semantic_ok`).
|
|
217
|
+
3. `finish` (from `Learning.auto_introspect`, or Loop if introspect is skipped)
|
|
218
|
+
adds the `Reward.judge` terminal score and updates Q and REINFORCE.
|
|
219
|
+
|
|
220
|
+
State is a short key: request kind, a hash of the active English task, last
|
|
221
|
+
action, fail count bin, and engine. Files:
|
|
222
|
+
|
|
223
|
+
- `~/.pwn/policy.json` - Q table, REINFORCE logits, visit counts, recent returns
|
|
224
|
+
- `~/.pwn/policy_traj.jsonl` - episode log
|
|
225
|
+
|
|
226
|
+
Tools `policy_stats`, `policy_evaluate`, and `policy_recommend` are read-only. Reset is Ruby-only (`PWN::AI::Agent::Policy.reset`) so a tool call cannot wipe the table. `Registry.rank` may add a small Q-advantage after a pair has been visited at least twice. Planning still owns the task list.
|
|
227
|
+
|
|
228
|
+
Turn it off with `ai.agent.policy: false` in `~/.pwn/pwn.yaml`.
|
|
229
|
+
|
|
206
230
|
## Task briefs vs learning
|
|
207
231
|
|
|
208
|
-
`TaskSummarizer` is **UX**, not a persistence layer. Plan/`about_to` lines are ephemeral TUI briefs (deduped by `last_brief_fp`). Durable learning still flows through Learning · Mistakes · Reward · sessions
|
|
232
|
+
`TaskSummarizer` is **UX**, not a persistence layer. Plan/`about_to` lines are ephemeral TUI briefs (deduped by `last_brief_fp`). Durable learning still flows through Learning · Mistakes · Reward · Policy · sessions. See [pwn-ai Agent § Task summaries](pwn-ai-Agent.md#task-summaries-long-autonomous-turns).
|
|
209
233
|
|
|
210
234
|
[← Home](Home.md)
|
|
@@ -20,7 +20,7 @@ with a **tool-calling AI agent** on top that can run the same methods.
|
|
|
20
20
|
| `PWN::FFI::*` | **8** | Native DSP/RF backends: Volk · Liquid · FFTW · RTLSdr · HackRF · AdalmPluto · SoapySDR · Stdio |
|
|
21
21
|
| `PWN::AI::*` | **6** engines | OpenAI, Anthropic, Grok (OAuth device-flow), Gemini, Ollama, Open WebUI |
|
|
22
22
|
| `bin/pwn_*` | **53** | Headless CLI drivers for CI/CD |
|
|
23
|
-
| Agent toolsets | **
|
|
23
|
+
| Agent toolsets | **13** · **85 tools** | terminal · pwn · memory · skills · sessions · learning · metrics · policy · extrospection · cron · swarm · reward · curriculum |
|
|
24
24
|
|
|
25
25
|
## Three ways to use it
|
|
26
26
|
|
|
@@ -36,12 +36,13 @@ with a **tool-calling AI agent** on top that can run the same methods.
|
|
|
36
36
|
- **Everything is Ruby, everything is a method.** No YAML DSLs, no black-box
|
|
37
37
|
plugins. If you can call it in the REPL, the AI agent can call it, a driver
|
|
38
38
|
can call it, and a cron job can call it.
|
|
39
|
-
- **Closed feedback loop.** Metrics + Learning + **Reward** (
|
|
40
|
-
sentinel) + **Curriculum** (mistake-driven self-play,
|
|
41
|
-
LoRA)
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
39
|
+
- **Closed feedback loop.** Metrics + Learning + **Reward** (outcome and process
|
|
40
|
+
judges, sentinel) + **Curriculum** (mistake-driven self-play, hindsight relabel,
|
|
41
|
+
export-ready LoRA gate) + **Policy** (live tabular Q-learning and REINFORCE on
|
|
42
|
+
each turn, advisory rank only) on the introspection side; Snapshot + Drift +
|
|
43
|
+
Intel + Verify on the extrospection side; joined by `extro_correlate`, which
|
|
44
|
+
tells the agent whether a failure was *its* fault or *the world* changed - and
|
|
45
|
+
writes the lesson back into the next prompt.
|
|
45
46
|
- **Native multi-agent.** `PWN::AI::Agent::Swarm` runs personas (each a full
|
|
46
47
|
tool-calling agent, optionally on a *different* LLM engine) that debate,
|
|
47
48
|
broadcast, and share an append-only bus - no IRC daemon, no external service.
|
data/documentation/Why-PWN.md
CHANGED
|
@@ -39,8 +39,9 @@ critical when the caller is an autonomous agent.
|
|
|
39
39
|
|
|
40
40
|
A pentest framework that does not learn repeats the same dead-end scans
|
|
41
41
|
forever. PWN records **per-tool success rate**, **per-task outcome**, **host
|
|
42
|
-
drift**,
|
|
43
|
-
|
|
42
|
+
drift**, **external CVE intel**, and a **live Policy** table (Q / REINFORCE
|
|
43
|
+
advice next to planning), then correlates them so tomorrow's run can start
|
|
44
|
+
where today's left off. See
|
|
44
45
|
[Skills, Memory & Learning](Skills-Memory-Learning.md),
|
|
45
46
|
[Reinforcement Learning](Reinforcement-Learning.md), and
|
|
46
47
|
[Extrospection](Extrospection.md).
|