eklavya 1.30.3 → 1.30.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/assets/tutor/references/grading.md +6 -0
- package/dist/assets/tutor/references/writing-mcq.md +9 -0
- package/dist/hooks/checkpoint-quiz.js +1 -1
- package/dist/plugin/.claude-plugin/plugin.json +1 -1
- package/dist/plugin/agents/tutor.md +2 -1
- package/dist/plugin/skills/tutor/references/grading.md +6 -0
- package/dist/plugin/skills/tutor/references/writing-mcq.md +9 -0
- package/package.json +1 -1
|
@@ -41,6 +41,12 @@ Within the cap, still grade honestly:
|
|
|
41
41
|
| 1 | picked a distractor built on a misconception |
|
|
42
42
|
| 0 | "Other" with *I don't know* (`outcome: dont_know` — **teach it**), or a decline (`outcome: declined`) |
|
|
43
43
|
|
|
44
|
+
**A defensible pick is a correct answer.** If the option they chose also
|
|
45
|
+
answers the stem — or they argue convincingly that it does — the question was
|
|
46
|
+
flawed, not the learner. Grade it 4 as if it were the key, say plainly that the
|
|
47
|
+
question had two right answers, and never grade them down for your own
|
|
48
|
+
ambiguity.
|
|
49
|
+
|
|
44
50
|
**If they want to explain, let them, and say so.** Someone who picks "Other" and
|
|
45
51
|
types a real answer has just given you better evidence than the multiple choice
|
|
46
52
|
could. Grade that as the free answer it is — **omit `format`**, and the cap does
|
|
@@ -43,6 +43,14 @@ learner can dismiss without thinking is a free point, and it teaches nothing.
|
|
|
43
43
|
When three that good will not come, the stem is too vague to have a near-miss:
|
|
44
44
|
rewrite the stem and the distractors follow.
|
|
45
45
|
|
|
46
|
+
**Exactly one option answers the stem.** Plausible is not the same as
|
|
47
|
+
defensible: a distractor must be believable *and wrong as an answer to this
|
|
48
|
+
question*. Row two is where this breaks — a true statement answers most "why"
|
|
49
|
+
stems well enough, and a learner who picks it is marked wrong for being right.
|
|
50
|
+
Before asking, read each distractor as if it were the key: if a senior reviewer
|
|
51
|
+
could argue for it, narrow the stem ("why *here*", "what does *this line*
|
|
52
|
+
prevent") until it cannot, or replace it with a different row.
|
|
53
|
+
|
|
46
54
|
**4. One clause of `description` per option.** This is where a near-miss earns
|
|
47
55
|
its place — the sentence that makes the wrong answer tempting.
|
|
48
56
|
|
|
@@ -90,6 +98,7 @@ telling you which part to rebuild.
|
|
|
90
98
|
grammar. A visibly longer or more careful option reads as the correct one, and
|
|
91
99
|
gets picked without engaging — the same leak as always answering first.
|
|
92
100
|
- Each of the three wrong options came from a different row of the table above.
|
|
101
|
+
- Only the correct option answers the stem; no expert could argue for another.
|
|
93
102
|
- The correct option sits at `answer_position`.
|
|
94
103
|
- The tool renders the labels, so the stem does not number them.
|
|
95
104
|
- The stem is the question and nothing else — the dials are in the developer's
|
|
@@ -149,7 +149,7 @@ Concept: ${row.concept}
|
|
|
149
149
|
|
|
150
150
|
Do exactly this, then get straight back to the task:
|
|
151
151
|
1. get_session_quiz_plan with max: 1 and ignore_cooldown: true (the pacing is already decided -- this hook is the cooldown).
|
|
152
|
-
2. Ask that ONE question with AskUserQuestion: four options, one correct, three plausible, and put the correct one in the slot answer_position names.
|
|
152
|
+
2. Ask that ONE question with AskUserQuestion: four options, exactly one correct, three plausible but wrong for this stem, and put the correct one in the slot answer_position names.
|
|
153
153
|
${attributionRule()}
|
|
154
154
|
3. Grade it with record_attempt: format "mcq", the labels in "options", the stem alone in "question".
|
|
155
155
|
4. Tell them the verdict: right, or wrong and what the right answer is, with one line of why. Never skip this -- an answer with no verdict teaches nothing.
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "eklavya",
|
|
3
3
|
"displayName": "Eklavya",
|
|
4
|
-
"version": "1.30.
|
|
4
|
+
"version": "1.30.4",
|
|
5
5
|
"description": "Learn while your agent works. Turns coding-agent generation time into adaptive, Socratic learning grounded in the code being written.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Ajay Kumar"
|
|
@@ -56,7 +56,8 @@ replying "A" without reading.
|
|
|
56
56
|
a blank, not a decline: grade 0 with `outcome: "dont_know"`, then teach.
|
|
57
57
|
|
|
58
58
|
Same rules as the tutor skill's `skills/tutor/references/writing-mcq.md`: four
|
|
59
|
-
options,
|
|
59
|
+
options, exactly one that answers the stem, three plausible distractors that do
|
|
60
|
+
not, the correct one at `answer_position`, the
|
|
60
61
|
stem alone in `record_attempt`'s `question`, the labels in `options`,
|
|
61
62
|
`format: "mcq"`, and the grade capped at 4. The only thing that changes is who
|
|
62
63
|
draws the box.
|
|
@@ -41,6 +41,12 @@ Within the cap, still grade honestly:
|
|
|
41
41
|
| 1 | picked a distractor built on a misconception |
|
|
42
42
|
| 0 | "Other" with *I don't know* (`outcome: dont_know` — **teach it**), or a decline (`outcome: declined`) |
|
|
43
43
|
|
|
44
|
+
**A defensible pick is a correct answer.** If the option they chose also
|
|
45
|
+
answers the stem — or they argue convincingly that it does — the question was
|
|
46
|
+
flawed, not the learner. Grade it 4 as if it were the key, say plainly that the
|
|
47
|
+
question had two right answers, and never grade them down for your own
|
|
48
|
+
ambiguity.
|
|
49
|
+
|
|
44
50
|
**If they want to explain, let them, and say so.** Someone who picks "Other" and
|
|
45
51
|
types a real answer has just given you better evidence than the multiple choice
|
|
46
52
|
could. Grade that as the free answer it is — **omit `format`**, and the cap does
|
|
@@ -43,6 +43,14 @@ learner can dismiss without thinking is a free point, and it teaches nothing.
|
|
|
43
43
|
When three that good will not come, the stem is too vague to have a near-miss:
|
|
44
44
|
rewrite the stem and the distractors follow.
|
|
45
45
|
|
|
46
|
+
**Exactly one option answers the stem.** Plausible is not the same as
|
|
47
|
+
defensible: a distractor must be believable *and wrong as an answer to this
|
|
48
|
+
question*. Row two is where this breaks — a true statement answers most "why"
|
|
49
|
+
stems well enough, and a learner who picks it is marked wrong for being right.
|
|
50
|
+
Before asking, read each distractor as if it were the key: if a senior reviewer
|
|
51
|
+
could argue for it, narrow the stem ("why *here*", "what does *this line*
|
|
52
|
+
prevent") until it cannot, or replace it with a different row.
|
|
53
|
+
|
|
46
54
|
**4. One clause of `description` per option.** This is where a near-miss earns
|
|
47
55
|
its place — the sentence that makes the wrong answer tempting.
|
|
48
56
|
|
|
@@ -90,6 +98,7 @@ telling you which part to rebuild.
|
|
|
90
98
|
grammar. A visibly longer or more careful option reads as the correct one, and
|
|
91
99
|
gets picked without engaging — the same leak as always answering first.
|
|
92
100
|
- Each of the three wrong options came from a different row of the table above.
|
|
101
|
+
- Only the correct option answers the stem; no expert could argue for another.
|
|
93
102
|
- The correct option sits at `answer_position`.
|
|
94
103
|
- The tool renders the labels, so the stem does not number them.
|
|
95
104
|
- The stem is the question and nothing else — the dials are in the developer's
|
package/package.json
CHANGED