pwn 0.5.643 → 0.5.650

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -81,7 +81,7 @@ $ pwn --ai "run bin/pwn_sast against ./src and push findings to DefectDojo"
81
81
 
82
82
  ## What the agent can call
83
83
 
84
- 10 toolsets · **71 tools** - full table at
84
+ 12 toolsets · **78 tools** - full table at
85
85
  [Agent Tool Registry](Agent-Tool-Registry.md).
86
86
 
87
87
  The two that matter most:
@@ -118,10 +118,11 @@ full `Loop.run` under a persona overlay) that share a JSONL bus. See
118
118
  local.
119
119
  - `PWN::AI::Agent::Learning.export_finetune` + `Reward.export_dpo` turn every
120
120
  successful session and every preference pair into supervised / DPO
121
- datasets under `~/.pwn/finetune/` - `Curriculum.train_and_gate` then
122
- LoRA-tunes the local model on them and only promotes the new adapter if it
123
- beats the old one on the current `Mistakes.top` set. See
124
- [Reinforcement Learning](Reinforcement-Learning.md).
121
+ datasets under `~/.pwn/finetune/` — `Curriculum.train_and_gate` then
122
+ LoRA-tunes the local model and promotes only under **gate v2** (resolved
123
+ margin + mean judge + frozen smoke set). Preference pairs use trajectory
124
+ geometry (P9/P14/P15): revised answers / winning traces, not `CORRECTION:` prose; `scrub_preferences` + export filter; practice lands `shape: :winning_trace`.
125
+ See [Reinforcement Learning](Reinforcement-Learning.md).
125
126
 
126
127
  ## RL feature flags (`PWN::Env[:ai][:agent]`)
127
128