pwn 0.5.683 → 0.5.685

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (82) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +17 -10
  3. data/documentation/Agent-Tool-Registry.md +3 -5
  4. data/documentation/CLI-Drivers.md +27 -28
  5. data/documentation/Configuration.md +5 -7
  6. data/documentation/Contributing.md +2 -2
  7. data/documentation/Cron.md +11 -7
  8. data/documentation/Diagrams.md +2 -2
  9. data/documentation/General-PWN-Usage.md +1 -1
  10. data/documentation/Home.md +8 -7
  11. data/documentation/How-PWN-Works.md +10 -10
  12. data/documentation/Installation.md +19 -3
  13. data/documentation/Persistence.md +3 -2
  14. data/documentation/Plugins.md +3 -3
  15. data/documentation/Reinforcement-Learning.md +102 -100
  16. data/documentation/SAST.md +1 -1
  17. data/documentation/Session-Workflow.md +132 -0
  18. data/documentation/Sessions.md +2 -1
  19. data/documentation/Skills-Memory-Learning.md +20 -0
  20. data/documentation/Troubleshooting.md +2 -2
  21. data/documentation/What-is-PWN.md +3 -3
  22. data/documentation/Why-PWN.md +1 -1
  23. data/documentation/diagrams/agent-tool-registry.svg +152 -152
  24. data/documentation/diagrams/cron-scheduling.svg +141 -121
  25. data/documentation/diagrams/dot/agent-tool-registry.dot +4 -4
  26. data/documentation/diagrams/dot/cron-scheduling.dot +10 -7
  27. data/documentation/diagrams/dot/driver-framework.dot +1 -1
  28. data/documentation/diagrams/dot/memory-skills-detailed.dot +1 -1
  29. data/documentation/diagrams/dot/overall-pwn-architecture.dot +5 -5
  30. data/documentation/diagrams/dot/persistence-filesystem.dot +2 -2
  31. data/documentation/diagrams/dot/plugin-ecosystem.dot +2 -2
  32. data/documentation/diagrams/dot/sessions-cron-automation.dot +1 -1
  33. data/documentation/diagrams/dot/task-summarizer.dot +18 -16
  34. data/documentation/diagrams/driver-framework.svg +1 -1
  35. data/documentation/diagrams/memory-skills-detailed.svg +85 -84
  36. data/documentation/diagrams/overall-pwn-architecture.svg +93 -93
  37. data/documentation/diagrams/persistence-filesystem.svg +5 -5
  38. data/documentation/diagrams/plugin-ecosystem.svg +101 -100
  39. data/documentation/diagrams/sessions-cron-automation.svg +7 -6
  40. data/documentation/diagrams/task-summarizer.svg +166 -156
  41. data/documentation/pwn-ai-Agent.md +13 -11
  42. data/etc/default_skills/bug-bounty-hunting/SKILL.md +70 -0
  43. data/etc/default_skills/deep-exploitation/SKILL.md +74 -0
  44. data/etc/default_skills/hardware-and-firmware-testing/SKILL.md +83 -0
  45. data/etc/default_skills/osint/SKILL.md +147 -0
  46. data/etc/default_skills/penetration-testing/SKILL.md +81 -0
  47. data/etc/default_skills/red-teaming/SKILL.md +70 -0
  48. data/etc/default_skills/reverse-engineering-binaries/SKILL.md +70 -0
  49. data/etc/default_skills/sast-code-scans/SKILL.md +66 -0
  50. data/etc/default_skills/social-engineering/SKILL.md +67 -0
  51. data/etc/default_skills/vulnerability-research-fundamentals/SKILL.md +74 -0
  52. data/etc/default_skills/web-application-penetration-testing/SKILL.md +68 -0
  53. data/lib/pwn/ai/agent/dispatch.rb +52 -1
  54. data/lib/pwn/ai/agent/loop.rb +170 -190
  55. data/lib/pwn/ai/agent/mistakes.rb +6 -6
  56. data/lib/pwn/ai/agent/open_goal.rb +83 -0
  57. data/lib/pwn/ai/agent/prompt_builder.rb +3 -5
  58. data/lib/pwn/ai/agent/tool_guard.rb +1 -46
  59. data/lib/pwn/ai/agent/tools/ruby_eval.rb +0 -10
  60. data/lib/pwn/ai/agent/tools/shell.rb +0 -6
  61. data/lib/pwn/ai/agent.rb +1 -0
  62. data/lib/pwn/config.rb +63 -2
  63. data/lib/pwn/cron.rb +18 -5
  64. data/lib/pwn/migrate.rb +2 -0
  65. data/lib/pwn/plugins/repl.rb +9 -4
  66. data/lib/pwn/plugins/tty_spinner.rb +64 -17
  67. data/lib/pwn/version.rb +1 -1
  68. data/spec/integration/config_spec.rb +37 -0
  69. data/spec/integration/persistence_roundtrip_spec.rb +2 -1
  70. data/spec/integration/reinforced_feedback_loop_spec.rb +6 -11
  71. data/spec/lib/pwn/ai/agent/dispatch_spec.rb +38 -82
  72. data/spec/lib/pwn/ai/agent/loop_spec.rb +230 -17
  73. data/spec/lib/pwn/ai/agent/mistakes_spec.rb +7 -0
  74. data/spec/lib/pwn/ai/agent/open_goal_spec.rb +37 -0
  75. data/spec/lib/pwn/ai/agent/signal_hygiene_spec.rb +5 -8
  76. data/spec/lib/pwn/ai/agent/tools/shell_spec.rb +10 -18
  77. data/spec/lib/pwn/cron_spec.rb +13 -3
  78. data/spec/lib/pwn/plugins/repl_spec.rb +17 -0
  79. data/spec/lib/pwn/plugins/tty_spinner_spec.rb +14 -0
  80. data/third_party/pwn_rdoc.jsonl +25 -5
  81. metadata +15 -2
  82. data/documentation/PWN_Contributors_and_Users.png +0 -0
@@ -5,10 +5,10 @@ agent writes what happened to disk and puts the useful bits back into the next
5
5
  prompt. On a host **with a trainer and a GPU**, the same data can also train a
6
6
  local adapter.
7
7
 
8
- `Curriculum.practice` → `Reward.export_dpo` → `Curriculum.train_and_gate`
8
+ `Curriculum.practice` -> `Reward.export_dpo` -> `Curriculum.train_and_gate`
9
9
 
10
- That path can promote a new LoRA when the candidate beats the current one.
11
- Without a trainer it still **exports** the datasets and a manual CLI. Live
10
+ That path can promote a new local adapter when the candidate beats the current
11
+ one. Without a trainer it still **exports** the datasets and a manual CLI. Live
12
12
  improvement does not wait on weights.
13
13
 
14
14
  ![Reinforcement-learning loop](diagrams/reinforcement-learning.svg)
@@ -16,32 +16,32 @@ improvement does not wait on weights.
16
16
  ```
17
17
  +------------------------------------------------+
18
18
  request -----> | Loop.run |
19
- | plan_first -> Curriculum.red_team_plan (S4) |
20
- | Dispatch -> Reward.semantic_ok (R4) |
21
- | -> Mistakes.record(cause:) (E1) |
22
- | guard -> Curriculum.counterfactual (S2) |--> preference ledger (W1)
23
- | final -> Curriculum.critic (S3) |
24
- | -> Reward.judge (outcome) (R1) |--> verify_as_reward (E3)
25
- | -> Reward.prm (process) (R2) |--> Sessions[step_reward] (C4)
26
- | -> Curriculum.hindsight (C3) |
27
- | -> Curriculum.calibrate (W3) |--> Metrics.calibration
28
- | -> Reward.sentinel (R3) |--> Mistakes(reward_signal)
19
+ | plan_first -> Curriculum.red_team_plan |
20
+ | Dispatch -> Reward.semantic_ok |
21
+ | -> Mistakes.record(cause:) |
22
+ | guard -> Curriculum.counterfactual |--> preference ledger
23
+ | final -> Curriculum.critic |
24
+ | -> Reward.judge (outcome) |--> verify_as_reward
25
+ | -> Reward.prm (process) |--> Sessions[step_reward]
26
+ | -> Curriculum.hindsight |
27
+ | -> Curriculum.calibrate |--> Metrics.calibration
28
+ | -> Reward.sentinel |--> Mistakes(reward_signal)
29
29
  +------------------------------------------------+
30
30
  |
31
- Learning.consolidate (M1 merge, M3 importance-evict)
32
- MemoryIndex.recall_semantic (M2 similarity x recency x importance)
33
- Registry.rank (C1 keyword + UCB + Q-advantage + tool_preference)
34
- Policy (R5 live Q / REINFORCE on a judge-scored MDP)
35
- Learning.exemplars_for (C2 prioritized replay, C4 minimal trace)
31
+ Learning.consolidate (merge + importance-evict)
32
+ MemoryIndex.recall_semantic (similarity x recency x importance)
33
+ Registry.rank (keyword + UCB + Q-advantage + tool_preference)
34
+ Policy (live Q / REINFORCE on a judge-scored MDP)
35
+ Learning.exemplars_for (prioritized replay, short traces)
36
36
  |
37
- nightly cron --> Curriculum.practice (S1) --> Mistakes.resolve --> preference (W1)
38
- weekly cron --> Curriculum.train_and_gate (W2) --> optional LoRA --> A/B gate --> promote
37
+ nightly cron --> Curriculum.practice --> Mistakes.resolve --> preference
38
+ weekly cron --> Curriculum.train_and_gate --> optional adapter --> A/B gate --> promote
39
39
  |
40
- Extrospection.correlate (E2 world vs self)
41
- Metrics.changepoints (E1 CUSUM) --> Mistakes(cause: :env_drift)
40
+ Extrospection.correlate (world vs self)
41
+ Metrics.changepoints (CUSUM) --> Mistakes(cause: :env_drift)
42
42
  ```
43
43
 
44
- ## Live Policy on every turn (R5 · `PWN::AI::Agent::Policy`)
44
+ ## Live Policy on every turn (`PWN::AI::Agent::Policy`)
45
45
 
46
46
  This is the live numeric controller. It does not replace planning.
47
47
 
@@ -65,89 +65,91 @@ high-return / high-score episodes.
65
65
 
66
66
  ## Reward signal (`PWN::AI::Agent::Reward`)
67
67
 
68
- | ID | Method | What it does |
69
- |----|--------|--------------|
70
- | **R1** | `.judge` | Cheap LLM outcome score on `(request, final)` → `{score:0..1, verdict:, rationale:, key_step:, source:}`. Calls the active engine `.chat` with a short timeout (default 12s). `Reflect.on` is used only when `module_reflection` is on. Fallback scores completeness, plan cover, claims, and tool-trace echo. Token overlap is only a small on-topic gate. |
71
- | **R2** | `.prm` | Process reward - per-tool-step `+1/0/-1` written into `Sessions[:step_reward]`. |
72
- | **R3** | `.sentinel` | Compares proxy success rate vs judge mean vs user-correction rate. A large gap fingerprints `reward_signal` so the agent distrusts a lying proxy. |
73
- | **R4** | `.semantic_ok` | Treats informational non-zero exits (for example `grep` / `rg` with no match) as benign. Metrics count them as OK; Mistakes only see true dispatch failures. |
74
- | **R5** | `Policy` (live MDP) | Tabular Q-learning + REINFORCE on real Loop turns. Q-advantage is an advisory `Registry.rank` term and never replaces TaskSummarizer or plan_first. |
75
- | - | `.warm_sentinel` | Backfills the sentinel window from scored Learning outcomes so local hosts can engage proxy distrust without waiting for live remote introspect. |
76
- | **W1** | `.record_preference` / `.export_dpo` | Preference ledger (`~/.pwn/preferences.jsonl`) from user corrections, resolve, counterfactual, critic, and practice. Caps per source; keeps trajectory-shaped pairs (winning traces / revised answers), not fix commentary. |
77
- | - | `.scrub_preferences` / `.preference_balance` / `.generator_mix` | Ledger hygiene and source-mix health so one channel cannot flood preference export. |
68
+ | Method | What it does |
69
+ |--------|--------------|
70
+ | **R1** `.judge` | Cheap LLM outcome score on `(request, final)` -> `{score:0..1, verdict:, rationale:, key_step:, source:}`. Calls the active engine `.chat` with a short timeout (default 12s). `Reflect.on` is used only when `module_reflection` is on. Fallback scores completeness, plan cover, claims, and tool-trace echo. Token overlap is only a small on-topic gate. |
71
+ | **R2** `.prm` | Process reward - per-tool-step `+1/0/-1` written into `Sessions[:step_reward]`. |
72
+ | **R3** `.sentinel` | Compares proxy success rate vs judge mean vs user-correction rate. A large gap fingerprints `reward_signal` so the agent distrusts a lying proxy. |
73
+ | **R4** `.semantic_ok` | Treats informational non-zero exits (for example `grep` / `rg` with no match) as benign. Metrics count them as OK; Mistakes only see true dispatch failures. |
74
+ | **R5** `Policy` (live MDP) | Tabular Q-learning + REINFORCE on real Loop turns. Q-advantage is an advisory `Registry.rank` term and never replaces TaskSummarizer or plan_first. |
75
+ | `.warm_sentinel` | Backfills the sentinel window from scored Learning outcomes so local hosts can engage proxy distrust without waiting for live remote introspect. |
76
+ | `.record_preference` / `.export_dpo` | Preference ledger (`~/.pwn/preferences.jsonl`) from user corrections, resolve, counterfactual, critic, and practice. Caps per source; keeps trajectory-shaped pairs (winning traces / revised answers), not fix commentary. |
77
+ | `.scrub_preferences` / `.preference_balance` / `.generator_mix` | Ledger hygiene and source-mix health so one channel cannot flood preference export. |
78
78
 
79
79
  ## Credit assignment and replay
80
80
 
81
- | ID | Where | What |
82
- |----|-------|------|
83
- | **C1** | `Registry.rank` + `Metrics.{ucb,thompson,advantage,prm_advantage}` + `Policy.advantage` | Live tool choice blends keyword fit, historical advantage, exploration bonus, process-reward signal, and (once the episode budget is met) Q(s,a)-V(s). Neighbor states fill cold (s,a) pairs. |
84
- | **C2** | `Learning.exemplars_for` | Prior successful traces ranked by judge score, recency, and keyword fit. Low-score "proxy success" rows are dropped when judge distrust is high. |
85
- | **C3** | `Curriculum.hindsight` | On a failed goal, relabel what the trajectory *did* achieve (`success: 'soft'`). Soft rows stay out of hard supervised export and are down-weighted in exemplars. |
86
- | **C4** | `Learning.compress_exemplar` / skill build | Keep steps with positive `step_reward` so few-shot traces stay short. |
81
+ | Where | What |
82
+ |-------|------|
83
+ | **C1** `Registry.rank` + `Metrics.{ucb,thompson,advantage,prm_advantage}` + `Policy.advantage` | Live tool choice blends keyword fit, historical advantage, exploration bonus, process-reward signal, and (once the episode budget is met) Q(s,a)-V(s). Neighbor states fill cold (s,a) pairs. |
84
+ | **C2** `Learning.exemplars_for` | Prior successful traces ranked by judge score, recency, and keyword fit. Low-score "proxy success" rows are dropped when judge distrust is high. |
85
+ | **C3** `Curriculum.hindsight` | On a failed goal, relabel what the trajectory *did* achieve (`success: 'soft'`). Soft rows stay out of hard supervised export and are down-weighted in exemplars. |
86
+ | **C4** `Learning.compress_exemplar` / skill build | Keep steps with positive `step_reward` so few-shot traces stay short. |
87
87
 
88
88
  ## Memory that stays high-signal
89
89
 
90
- | ID | Where | What |
91
- |----|-------|------|
92
- | **M1** | `Learning.consolidate` | Semantic merge of near-duplicate lessons; importance-weighted eviction. |
93
- | **M2** | `MemoryIndex.recall_semantic` | Rank MEMORY by similarity × recency × importance when embeddings are available. |
94
- | **M3** | `Memory.remember` | Supports source, confidence, importance, and TTL so garbage self-evicts. |
95
- | **M4** | `Learning.note_outcome` | Task outcomes go to `learning.jsonl` only. Memory `:lesson` is reserved for reflect, resolve, and human notes. `purge_noise` cleans older noisy lesson shapes. |
90
+ | Where | What |
91
+ |-------|------|
92
+ | **M1** `Learning.consolidate` | Semantic merge of near-duplicate lessons; importance-weighted eviction. |
93
+ | **M2** `MemoryIndex.recall_semantic` | Rank MEMORY by similarity x recency x importance when embeddings are available. |
94
+ | **M3** `Memory.remember` | Supports source, confidence, importance, and TTL so garbage self-evicts. |
95
+ | **M4** `Learning.note_outcome` | Task outcomes go to `learning.jsonl` only. Memory `:lesson` is reserved for reflect, resolve, and human notes. `purge_noise` cleans older noisy lesson shapes. |
96
96
 
97
97
  ## Curriculum and self-play (`PWN::AI::Agent::Curriculum`)
98
98
 
99
- | ID | Method | What |
100
- |----|--------|------|
101
- | **S1** | `.practice` | Mine top unresolved Mistakes → natural reproducers → self-play under the judge → auto-resolve on strong holdouts. Budget-exhaustion scars get "finish under N steps" prompts. |
102
- | **S2** | `.counterfactual` | On repeated failure, fork an alt-persona branch, judge both, emit a preference pair. Skipped when the iteration budget is already under pressure. |
103
- | **S3** | `.critic` | Constitutional critic of the final answer (can use tools). Under budget pressure, runs text-only so it cannot burn the remaining iterations. |
104
- | **S4** | `.red_team_plan` | Adversarial review of the plan-first outline using Metrics / Mistakes / drift. |
105
- | - | `.offline_judge` | Score recent sessions with outcome + process judges, warm the sentinel, optional ledger scrub. Meant for nightly cron on local hosts that only introspect failures live. |
106
- | **C3** | `.hindsight` | Relabel described above. |
107
- | **W3** | `.calibrate` | Plan `p(success)=` vs actual outcome → per-engine Brier / overconfidence. Overconfidence can force plan_first + critic and tighten `max_iters`. |
108
- | **W2** | `.train_and_gate` | Export supervised + preference data → optional LoRA train → promote only if the candidate wins on resolved margin, mean judge, smoke set, and a healthy preference diet. Without a trainer: `weight_loop: :export_ready`. |
109
- | - | `.practice_kpi` | Week-over-week trend of repeating mistakes (outer curriculum health). |
99
+ | Method | What |
100
+ |--------|------|
101
+ | **S1** `.practice` | Mine top unresolved Mistakes -> natural reproducers -> self-play under the judge -> auto-resolve on strong holdouts. Budget-exhaustion scars get "finish under N steps" prompts. |
102
+ | **S2** `.counterfactual` | On repeated failure, fork an alt-persona branch, judge both, emit a preference pair. Skipped when the iteration budget is already under pressure. |
103
+ | **S3** `.critic` | Constitutional critic of the final answer (can use tools). Under budget pressure, runs text-only so it cannot burn the remaining iterations. |
104
+ | **S4** `.red_team_plan` | Adversarial review of the plan-first outline using Metrics / Mistakes / drift. |
105
+ | `.offline_judge` | Score recent sessions with outcome + process judges, warm the sentinel, optional ledger scrub. Meant for nightly cron on local hosts that only introspect failures live. |
106
+ | `.hindsight` | Relabel described above. |
107
+ | **W3** `.calibrate` | Plan `p(success)=` vs actual outcome -> per-engine Brier / overconfidence. Overconfidence can force plan_first + critic. It does **not** shrink this request's `max_iters`. |
108
+ | **W2** `.train_and_gate` | Export supervised + preference data -> optional local adapter train -> promote only if the candidate wins on resolved margin, mean judge, smoke set, and a healthy preference diet. Without a trainer: `weight_loop: :export_ready`. |
109
+ | **W1** `.record_preference` | Preference ledger from corrections, resolve, counterfactual, critic, and practice. |
110
+ | `.practice_kpi` | Week-over-week trend of repeating mistakes (outer curriculum health). |
110
111
 
111
112
  ## Budget pressure (iteration ceiling)
112
113
 
113
114
  When unresolved `agent_loop` / `assistant_answer` budget-exhaustion fingerprints
114
- dominate, the loop marks the budget path hot and tightens the live turn:
115
+ dominate, the loop marks the budget path hot and tightens *side work* only:
115
116
 
116
- - lower effective `max_iters` (stricter on local/ollama engines than remote)
117
- - last-iter force-final: tools=nil on the final iteration so a text answer is required
118
117
  - skip counterfactual / red-team forks that would spend more tool rounds
118
+ - last-iter strips tools only when the original request is already satisfied
119
119
  - still flush TaskSummarizer state and Learning on the exhaust path
120
120
  - end-of-turn critic runs text-only under the same pressure
121
+ - **do not shrink `max_iters`** - yesterday's scars must not abort this request
121
122
 
122
123
  Practice prioritizes those scars with short-horizon "finish the task" prompts.
123
- Raising `ai.agent.max_iters` or resolving the scar returns normal runway.
124
+ Raising `ai.agent.max_iters` still sets the hard ceiling.
124
125
 
125
126
  ## Design-priority STATUS
126
127
 
127
- This table is the live control list. Track the outcomes, not source comments.
128
-
129
- | Pri | ID | Control | Module(s) | Success criterion |
130
- |-----|----|---------|-----------|-------------------|
131
- | **P0** | W1 generator diversity | `Reward::TARGET_SOURCE_MIX` + `generator_mix` + mix-urgent force on critic/counterfactual | `reward.rb`, `curriculum.rb` | `generator_mix.healthy` OR `recommendation` not stuck on `suppress:mistakes_resolve`; trajectory_fraction ≥ 0.5 |
132
- | **P0** | Introspect budget | `Learning::INTROSPECT_SOFT_MS` / `HARD_MS`; stage skip under soft/hard / `budget_exhaustion_hot?` | `learning.rb` | `auto_introspect` returns `stages_skipped` when over soft; post-answer path cannot re-thrash tool critic |
133
- | **P1** | Local judge calibration | Heuristic score shrinkage + `confidence`; `Metrics.effective_rate` scales distrust by `judge_confidence` | `reward.rb`, `metrics.rb` | distrust × heuristic no longer fully replaces proxy; local no-trace highs capped |
134
- | **ops** | Cheap LLM ORM | Direct engine `.chat` for `Reward.judge` / `.prm`; optional `reward_model` / `reward_llm_timeout`; evidence-prior fallback last | `reward.rb` | remote turns grade with ORM even when `module_reflection` is off |
135
- | **ops** | Outcome-signal haircut | `judge_sample_weight` (ORM 1.0, heuristic 0.25) in sentinel / `Learning.stats` / `Metrics.judge_rate` | `reward.rb`, `learning.rb`, `metrics.rb` | proxy distrust blends toward ORM scores, not bag-of-words overlap |
136
- | **P1** | Practice outer KPI | `Curriculum.practice_kpi` / `repeating_trend` → `~/.pwn/curriculum_kpi.jsonl` | `curriculum.rb` | week-over-week `delta_repeating` ≤ 0 on budget fingerprints after practice nights |
137
- | **P2** | PRM sample efficiency | `PRM_MIN_N=5`, shrinkage to `PRM_FULL_N=20`, fleet coverage gate in `Registry.rank` | `metrics.rb`, `registry.rb` | `prm_advantage=0` until n≥5; rank delta=0 until ≥3 tools ready |
138
- | **P2** | STATUS over flag archaeology | This table | docs | New work cites Pri/ID here, not fresh P26+ comments for the same theme |
139
- | **ops** | Nightly diet close | `offline_judge` → `scrub_preferences` + `generator_mix` + `practice_kpi` | `curriculum.rb` | Cron path returns `scrub`/`generator_mix`/`practice_kpi`; raw resolve prose does not survive the night |
140
- | **ops** | Shape backfill | `Reward.infer_shape` + scrub rewrite | `reward.rb` | Legacy shapeless rows get `winning_trace`/`revised_answer` when content warrants; traj_f measurable |
141
- | **ops** | Mix in prompt | `Metrics.to_context` emits `W1 MIX:` when unhealthy | `metrics.rb` | Unhealthy diet visible every turn without a tool call |
142
- | **P0** | Budget exhaust deepen | Last-iter force-final (tools=nil); skip CF when `budget_exhaustion_hot?`; tighter caps 24 local / 75 remote; exhaust path `append_session`+`auto_introspect` | `loop.rb` | Exhaust returns a judged final, not a bare string; CF cannot re-enter under hot; last iter cannot tool-call |
128
+ This is the live control list. Track the outcomes, not source comments.
129
+
130
+ `PRM_MIN_N` / `PRM_FULL_N` gate process-reward rank until enough samples exist.
131
+
132
+ | Priority | Control | Success look |
133
+ |----------|---------|--------------|
134
+ | must | Preference-source diversity (`Reward::TARGET_SOURCE_MIX` + `generator_mix`) | Mix is healthy, or the recommendation is not stuck suppressing `mistakes_resolve`; trajectory-shaped pairs stay at least half the ledger |
135
+ | must | Introspect budget (`Learning::INTROSPECT_SOFT_MS` / `HARD_MS`) | `auto_introspect` skips extra stages when over the soft cap so the post-answer path cannot re-thrash a tool critic |
136
+ | high | Local judge calibration | Heuristic scores shrink by confidence; a local no-trace high does not fully replace the proxy |
137
+ | ops | Cheap LLM outcome judge | Remote turns grade with the engine chat even when `module_reflection` is off; token overlap is last-resort |
138
+ | ops | Outcome-signal haircut (`judge_sample_weight`) | Proxy distrust blends toward engine scores, not bag-of-words overlap |
139
+ | high | Practice outer KPI (`Curriculum.practice_kpi`) | Week-over-week repeating-mistake trend on budget fingerprints is not rising after practice nights |
140
+ | next | Process-reward sample efficiency | `prm_advantage` stays 0 until enough samples exist; rank does not swing on a tiny fleet |
141
+ | ops | Nightly diet close | `offline_judge` then scrub + mix + KPI so raw resolve prose does not survive the night |
142
+ | ops | Shape backfill (`Reward.infer_shape`) | Legacy shapeless rows become `winning_trace` / `revised_answer` when the content warrants |
143
+ | ops | Mix in prompt | `Metrics.to_context` emits `MIX:` when preference sources are unhealthy |
144
+ | must | Budget exhaust | Last-iter force-final only when the original request is already satisfied; skip extra forks when the budget path is hot; **do not shrink max_iters**; still flush the session + introspect |
143
145
 
144
146
  ## Intro and extro join
145
147
 
146
- | ID | Where | What |
147
- |----|-------|------|
148
- | **E1** | `Metrics.changepoints` + `Loop.attribute_cause` | Env-drift-attributed failures get `cause: :env_drift` and do not inflate `[REPEATING]`. |
149
- | **E2** | `Extrospection.correlate` | Lead-lag style joins ("tool X started failing after toolchain Y changed"). |
150
- | **E3** | `Reward.verify_as_reward` | Browser-backed claim checks can floor/cap the outcome score. |
148
+ | Where | What |
149
+ |-------|------|
150
+ | **E1** `Metrics.changepoints` + `Loop.attribute_cause` | Env-drift-attributed failures get `cause: :env_drift` and do not inflate `[REPEATING]`. |
151
+ | **E2** `Extrospection.correlate` | Lead-lag style joins ("tool X started failing after toolchain Y changed"). |
152
+ | **E3** `Reward.verify_as_reward` | Browser-backed claim checks can floor/cap the outcome score. |
151
153
 
152
154
  ## Config (`PWN::Env[:ai][:agent]`)
153
155
 
@@ -155,18 +157,18 @@ This table is the live control list. Track the outcomes, not source comments.
155
157
  :ai:
156
158
  :module_reflection: false # gates Reflect lesson writing (not the judge alone)
157
159
  :agent:
158
- :critic: null # S3 - nil = ON for remote engines, OFF for ollama
159
- :red_team_plan: null # S4 - same auto rule
160
- :counterfactual: null # S2 - same auto rule
161
- :hindsight: true # C3 - hindsight relabel on failed turns (default true)
162
- :policy: true # R5 - live tabular Q / REINFORCE (advisory rank only)
163
- :verify_as_reward: null # E3 - nil = auto sample on claim-shaped answers
160
+ :critic: null # nil = ON for remote engines, OFF for ollama
161
+ :red_team_plan: null # same auto rule
162
+ :counterfactual: null # same auto rule
163
+ :hindsight: true # hindsight relabel on failed turns (default true)
164
+ :policy: true # live tabular Q / REINFORCE (advisory rank only)
165
+ :verify_as_reward: null # nil = auto sample on claim-shaped answers
164
166
  :reward_llm: null # nil = outcome/process judges use LLM teacher on remote
165
167
  :reward_model: null # optional cheaper model id for ORM/PRM (nil = engine default)
166
168
  :reward_llm_timeout: 12 # cheap ORM chat timeout seconds (clamped 2..30)
167
169
  :local_introspect: :failure_only # ollama cost rule; remote always introspects
168
170
  :introspect_every_n: 3
169
- :max_iters: 75 # hard cap; budget pressure may lower effective value
171
+ :max_iters: 777 # hard cap for this request; scars do not lower it
170
172
  :defer_introspect: true # post-answer Learning after the user-visible reply
171
173
  :prompt_cache: true # engine-native prefix cache (not ollama / openwebui)
172
174
  :tool_preference: [memory_recall, session_recall, skills_recall, pwn_eval, shell, mistakes_record, mistakes_resolve, learning_note_outcome, memory_remember]
@@ -177,10 +179,10 @@ This table is the live control list. Track the outcomes, not source comments.
177
179
  ```ruby
178
180
  # Seeded idempotently by PWN::Cron.install_defaults (pwn setup --migrate):
179
181
  PWN::Cron.install_defaults
180
- # → curriculum_practice_nightly 0 3 * * * Curriculum.practice(limit: 3)
181
- # → curriculum_offline_judge 30 3 * * * Curriculum.offline_judge(since_hours: 24, limit: 40)
182
- # → curriculum_train_weekly 0 4 * * 0 Curriculum.train_and_gate(dry_run: true) # false only with trainer+GPU
183
- # → learning_consolidate_nightly 0 5 * * * Learning.consolidate
182
+ # -> curriculum_practice_nightly 0 3 * * * Curriculum.practice(limit: 3)
183
+ # -> curriculum_offline_judge 30 3 * * * Curriculum.offline_judge(since_hours: 24, limit: 40)
184
+ # -> curriculum_train_weekly 0 4 * * 0 Curriculum.train_and_gate(dry_run: true) # false only with trainer+GPU
185
+ # -> learning_consolidate_nightly 0 5 * * * Learning.consolidate
184
186
  ```
185
187
 
186
188
  `install_defaults` treats the legacy name `offline_judge_nightly` as an alias of
@@ -199,13 +201,13 @@ without a per-job crontab line.
199
201
 
200
202
  ## What this stack actually does
201
203
 
202
- 1. Process reward on real security tool traces, not only math demos (**R2**).
203
- 2. Automatic blame: self vs environment drift (**E1** + **E2**).
204
- 3. Reward-hacking self-detection when proxy success diverges from the judge (**R3**).
205
- 4. Mistake-driven curriculum with regression-gated LoRA promotion when a trainer exists (**S1** + **W2**).
206
- 5. Preference pairs from normal agent work (corrections, resolve, critic, practice) with no separate human labelling queue (**W1**).
204
+ 1. Process reward on real security tool traces, not only math demos.
205
+ 2. Automatic blame: self vs environment drift.
206
+ 3. Reward-hacking self-detection when proxy success diverges from the judge.
207
+ 4. Mistake-driven curriculum with regression-gated adapter promotion when a trainer exists.
208
+ 5. Preference pairs from normal agent work (corrections, resolve, critic, practice) with no separate human labelling queue.
207
209
  6. Export and promote only when the preference diet is diverse and trajectory-shaped.
208
- 7. Live tabular Q-learning + REINFORCE on real Loop turns, used only as advice next to planning (**R5**).
210
+ 7. Live tabular Q-learning + REINFORCE on real Loop turns, used only as advice next to planning.
209
211
 
210
212
  ## Preference signal quality
211
213
 
@@ -215,9 +217,9 @@ keep it honest by:
215
217
 
216
218
  1. Capping how much any one preference source can dominate the ledger.
217
219
  2. Preferring winning tool traces and revised full answers over fix commentary.
218
- 3. Scrubbing historical prose-only pairs before DPO export.
220
+ 3. Scrubbing historical prose-only pairs before preference export.
219
221
  4. Warming the sentinel from offline scores so local hosts are not stuck cold.
220
- 5. Requiring smoke checks and mean judge improvement before LoRA promote.
222
+ 5. Requiring smoke checks and mean judge improvement before adapter promote.
221
223
  6. Treating budget exhaustion as a first-class practice target so the agent learns to finish.
222
224
 
223
225
  ---
@@ -1,6 +1,6 @@
1
1
  # `PWN::SAST` - Static Application Security Testing
2
2
 
3
- 48 language-aware rule modules + a `TestCaseEngine` + a `Factory` loader.
3
+ 48 modules in `PWN::SAST`: 45 language-aware scan rules plus `Factory`, `TestCaseEngine`, and `Version`.
4
4
  Source: `lib/pwn/sast/*.rb`. CLI: `bin/pwn_sast`.
5
5
 
6
6
  ![SAST pipeline](diagrams/code-scanning-sast.svg)
@@ -0,0 +1,132 @@
1
+ # Session Workflow
2
+
3
+ How a `pwn-ai` line is routed, when a goal is considered unfinished, and
4
+ which phrases flip those switches. Exact matching lives in
5
+ `PWN::AI::Agent::Loop` and `PWN::AI::Agent::OpenGoal`. This page is the
6
+ operator cheat sheet.
7
+
8
+ **See also:** [Sessions](Sessions.md) · [pwn-ai Agent](pwn-ai-Agent.md) ·
9
+ [Persistence](Persistence.md)
10
+
11
+ ---
12
+
13
+ ## Life of a line
14
+
15
+ 1. Entering `pwn-ai` always creates a **new** `~/.pwn/sessions/<timestamp>_<hex>.jsonl`.
16
+ 2. Cheap intents (greeting, how-to, recall) never open a host-work goal.
17
+ 3. Everything else is an autonomous goal. `Loop.run` keeps CORE_TOOLS until
18
+ the original request is done or truly blocked.
19
+ 4. An unfinished host-work request is written to `~/.pwn/open_goal.json`.
20
+ 5. An accepted final answer deletes that file. Budget exhaust leaves it.
21
+
22
+ ```
23
+ you type a line
24
+ |
25
+ +-- greeting / how-to / recall -----> short answer, no tools
26
+ |
27
+ +-- continue / resume / keep going --> reload open_goal.json
28
+ |
29
+ +-- last / previous session ---------> prior JSONL, not this line
30
+ |
31
+ +-- everything else -----------------> one Loop.run
32
+ write? then read it back
33
+ listing is not done
34
+ ```
35
+
36
+ ---
37
+
38
+ ## Unfinished goals
39
+
40
+ File: `~/.pwn/open_goal.json` (`PWN::AI::Agent::OpenGoal`).
41
+
42
+ | You type (whole line) | Effect |
43
+ |---|---|
44
+ | `continue` | Reload the saved request and keep working |
45
+ | `resume` | same |
46
+ | `keep going` | same |
47
+ | `pick up` | same |
48
+ | `carry on` | same |
49
+ | `continue please` | same |
50
+ | `resume the goal` / `resume the task` / `resume work` | same |
51
+
52
+ The line must be **only** that phrase (optional `please` / `the goal|task|work`
53
+ and a trailing `.` / `!`). A new sentence is a new goal and **replaces** the
54
+ saved request.
55
+
56
+ `continue scanning the lab` is a new ask, not a resume.
57
+
58
+ Delete `~/.pwn/open_goal.json` to drop a stuck checkpoint.
59
+
60
+ ---
61
+
62
+ ## This session vs last session
63
+
64
+ Entering `pwn-ai` is a new transcript. "Last session" is the newest other
65
+ `~/.pwn/sessions/*.jsonl`.
66
+
67
+ | Phrase | Where it reads |
68
+ |---|---|
69
+ | `what did I just say?` | **This** session |
70
+ | `what did I just ask?` / `what was my last request?` | this session |
71
+ | `how did you respond?` / `what did you just say?` | this session |
72
+ | `what did I just say in the last session?` | **Previous** JSONL |
73
+ | `in the last session` / `previous session` / `prior session` | previous JSONL |
74
+ | `session_recall` (tool) | older transcripts, skips the current id |
75
+
76
+ If there is no previous file: `I do not have a previous session transcript yet.`
77
+
78
+ ---
79
+
80
+ ## Cheap short-circuits (no tools)
81
+
82
+ Whole-line greetings only, so `hi, scan this host` stays a goal:
83
+
84
+ | Trigger | Examples |
85
+ |---|---|
86
+ | Greeting | `hi` `hello` `hey` `howdy` `yo` `good morning` |
87
+ | How-to | `how to` `how do I` `how can I` `syntax for` `usage of` `man page` |
88
+ | Recall | table above |
89
+
90
+ ---
91
+
92
+ ## Host work stays open until the artefact is checked
93
+
94
+ | Need | Done only when |
95
+ |---|---|
96
+ | Write / update / regenerate / docs / fix | A **write** effect, then a later **read** (`cat`, `ruby -c`, eval read) |
97
+ | Browser / navigate / `TransparentBrowser` | A **browse** effect (goto / dump_links / close) |
98
+ | Hostname / uname / cwd / whoami | Any live **read** |
99
+ | Other long goals (bounty, scrape, recon) | **write**, **browse**, or **eval** (`ls` alone is not enough) |
100
+
101
+ These do **not** count as finishing a write: `memory_remember`,
102
+ `learning_note_outcome`, `mistakes_*`, rspec/rubocop green, a README listing,
103
+ or "I will do that next time."
104
+
105
+ After you mutate a file, read it back before a final.
106
+
107
+ ---
108
+
109
+ ## Phrases that are **not** a final answer
110
+
111
+ Loop keeps calling tools if the model emits any of:
112
+
113
+ - `shall I` / `should I` / `want me to` / `proceed?` / `continue?`
114
+ - `# Remaining block` / heading-only outlines
115
+ - `were not applied` / `not written to disk` / `next time`
116
+ - narrated next tool (`Wait, let's try...`) with no `tool_calls`
117
+
118
+ Type `continue` yourself only to resume a **saved** open goal after the REPL
119
+ died. Do not use it as a mid-turn "ok, go on" - the loop should already be
120
+ going.
121
+
122
+ ---
123
+
124
+ ## Files
125
+
126
+ | Path | Role |
127
+ |---|---|
128
+ | `~/.pwn/sessions/<id>.jsonl` | This activation's transcript |
129
+ | `~/.pwn/open_goal.json` | Unfinished host-work request |
130
+ | `~/.pwn/memory.json` | Durable facts (not "last line") |
131
+
132
+ [← Home](Home.md)
@@ -38,7 +38,8 @@ appended as JSON-per-line to `~/.pwn/sessions/<id>.jsonl`.
38
38
  {"role":"tool","ts":"...","name":"shell","content":"...","step_reward":1}
39
39
  ```
40
40
 
41
- **See also:** [Skills, Memory & Learning](Skills-Memory-Learning.md) ·
41
+ **See also:** [Session Workflow](Session-Workflow.md) ·
42
+ [Skills, Memory & Learning](Skills-Memory-Learning.md) ·
42
43
  [Swarm](Swarm.md) · [Persistence](Persistence.md)
43
44
 
44
45
  [← Home](Home.md)
@@ -103,6 +103,26 @@ them in-place to the spec-conformant `<name>/SKILL.md` layout with
103
103
  back-filled `name`/`description` front-matter. `skill_create` always writes
104
104
  the new format.
105
105
 
106
+ A fresh install (first `pwn` launch or `pwn setup --migrate`) copies these
107
+ bundled skills into `~/.pwn/skills/` when the name is missing:
108
+
109
+ | Skill | For |
110
+ |---|---|
111
+ | `vulnerability-research-fundamentals` | First-pass research methodology with PWN plugins |
112
+ | `deep-exploitation` | Crash / primitive to reliable PoC |
113
+ | `bug-bounty-hunting` | Program scope, Burp, authz replay, report |
114
+ | `sast-code-scans` | `PWN::SAST::Factory` / `pwn_sast` + reports |
115
+ | `reverse-engineering-binaries` | checksec, disasm, `PWN::Plugins::Assembly` |
116
+ | `penetration-testing` | PTES / NIST / OSSTMM / ISSAF engagement loop |
117
+ | `web-application-penetration-testing` | OWASP WSTG / ASVS / API Top 10 |
118
+ | `red-teaming` | MITRE ATT&CK, TIBER-EU, Kill Chain |
119
+ | `hardware-and-firmware-testing` | OWASP FSTM / ISTG, serial, firmware |
120
+ | `social-engineering` | OSSTMM human channel, ATT&CK phishing |
121
+ | `osint` | `extro_osint` kinds, feeds, keys, pivots |
122
+
123
+ Source: `etc/default_skills/` in the gem. Edits in `~/.pwn/skills` are never
124
+ overwritten.
125
+
106
126
  ## Housekeeping
107
127
 
108
128
  | Tool | When |
@@ -138,7 +138,7 @@ list, then install equivalents by hand.
138
138
 
139
139
  **Cause:** older dual-emit paths or `about_to` restating the full goal every batch.
140
140
 
141
- **Current behavior (v0.5.660+):**
141
+ **Current behavior (v0.5.684+):**
142
142
 
143
143
  1. `Loop.task_summary_about_to!` is the only about_to entry.
144
144
  2. `TaskSummarizer.about_to` builds distinct lines with `tool_counts_phrase` + `intent_phrase`.
@@ -149,6 +149,6 @@ If you still see doubles, confirm you are on a build with `lib/pwn/ai/agent/task
149
149
 
150
150
  ## Agent stops early / "iteration budget exhausted"
151
151
 
152
- When unresolved budget-exhaustion mistakes dominate, the Loop tightens effective `max_iters` (stricter on local/ollama than remote), and the last iterations are text-only so a final answer is forced without starving multi-step autonomy. Check `mistakes_list`, resolve fixed signatures, or raise `ai.agent.max_iters` after the scar cools. Exhaust paths still run Learning + task_summary flush.
152
+ When unresolved budget-exhaustion mistakes dominate, the Loop does **not** shrink `max_iters` (default 777). It only trims side work: skip extra counterfactual forks, and strip tools on the last iteration only if the original request is already satisfied. Check `mistakes_list`, resolve fixed signatures, or raise `ai.agent.max_iters` if a long task still hits the cap. Exhaust paths still run Learning and the task-summary flush.
153
153
 
154
154
  [← Home](Home.md)
@@ -12,15 +12,15 @@ with a **tool-calling AI agent** on top that can run the same methods.
12
12
 
13
13
  | Namespace | Count | What it is |
14
14
  |---|---|---|
15
- | `PWN::Plugins::*` | **66** | Wrappers for external and native tooling (Burp, Nmap, Metasploit, Shodan, browsers, serial, ...) |
15
+ | `PWN::Plugins::*` | **67** | Wrappers for external and native tooling (Burp, Nmap, Metasploit, Shodan, browsers, serial, ...) |
16
16
  | `PWN::SAST::*` | **48** | Static-analysis rules across C/Java/Go/Python/Ruby/Scala/PHP/TS |
17
17
  | `PWN::AWS::*` | **90** | One module per AWS service for cloud enumeration |
18
18
  | `PWN::WWW::*` | **22** | Site-specific browser automations (HackerOne, BugCrowd, GitHub, Google, LinkedIn, ...) |
19
19
  | `PWN::SDR::*` | **6** (+ **20** protocol decoders + Base/DSP) | GQRX, FlipperZero, RFIDler, SonMicro, band tables, `Decoder::{ADSB,POCSAG,RDS,LoRa,...}` |
20
20
  | `PWN::FFI::*` | **8** | Native DSP/RF backends: Volk · Liquid · FFTW · RTLSdr · HackRF · AdalmPluto · SoapySDR · Stdio |
21
21
  | `PWN::AI::*` | **6** engines | OpenAI, Anthropic, Grok (OAuth device-flow), Gemini, Ollama, Open WebUI |
22
- | `bin/pwn_*` | **53** | Headless CLI drivers for CI/CD |
23
- | Agent toolsets | **13** · **85 tools** | terminal · pwn · memory · skills · sessions · learning · metrics · policy · extrospection · cron · swarm · reward · curriculum |
22
+ | `bin/pwn_*` + `pwn` | **54** | Headless CLI executables for CI/CD |
23
+ | Agent toolsets | **13** · **87 tools** | terminal · pwn · memory · skills · sessions · learning · metrics · policy · extrospection · cron · swarm · reward · curriculum |
24
24
 
25
25
  ## Three ways to use it
26
26
 
@@ -33,7 +33,7 @@ critical when the caller is an autonomous agent.
33
33
  | "Trust our scanner" | `cat lib/pwn/sast/sql.rb` - read the regex yourself |
34
34
  | Per-seat license for the glue | MIT-licensed glue; bring your own Burp Pro / Nessus key |
35
35
  | Agent output is a PDF | Agent output is a `PWN::Reports` object *and* a distilled skill *and* a memory entry |
36
- | One vendor's model | Five interchangeable engines; swarm can pit them against each other |
36
+ | One vendor's model | Six interchangeable engines; swarm can pit them against each other |
37
37
 
38
38
  ## Why the feedback loop matters
39
39