@nodaro/prompts 1.4.0 → 1.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,226 @@
1
+ /**
2
+ * REFERENCE RULES — the two short blocks that decide whether a multi-reference
3
+ * image obeys its references, and whether it looks like a film frame or a
4
+ * posed photograph.
5
+ *
6
+ * Both were measured on gpt-image-2 at 2K, against one deliberately hard brief:
7
+ * four references (two people, a street-fashion shot, and a composite holding a
8
+ * wine glass, a smartphone, a hotel suite AND two other women's faces), with
9
+ * wardrobe swapped BETWEEN the two people. 36 draws, 2026-08-06. Scored on five
10
+ * criteria: both faces correct and unblended, each person's wardrobe, the
11
+ * props, the location, and one coherent composition.
12
+ *
13
+ * ─── REFERENCE_RULES ───────────────────────────────────────────────────────
14
+ *
15
+ * Two wordings already existed in this codebase and NEITHER was best. Each was
16
+ * written without knowledge of the other:
17
+ *
18
+ * the `reference-lock` snippet — default-deny + likeness + compose.
19
+ * **0 of 4** on the hardest criterion (move
20
+ * a garment from one reference to another).
21
+ * Flawless on the other four, which is what
22
+ * made it feel like it worked.
23
+ * gvp's `referenceRules` — default-deny + face rules + likeness +
24
+ * "expression, gaze and pose follow the
25
+ * scene". **1 of 4**.
26
+ * the two MERGED — **4 of 4** (Fisher exact vs the snippet,
27
+ * p ≈ 0.03).
28
+ *
29
+ * The arms isolate: swapping the performance clause for the compose clause took
30
+ * it 1/4 → 4/4, and adding the face rules took 0/4 → 4/4. Each half does real
31
+ * work and neither is sufficient alone. The performance clause is dropped here
32
+ * because a composite has no scene to perform — it belongs to gvp's SCENE lane,
33
+ * where a subject must act the beat.
34
+ *
35
+ * ─── SCENE_FRAME_RULE ──────────────────────────────────────────────────────
36
+ *
37
+ * The separate failure Tal reported: everyone faces the lens and the result
38
+ * reads as an artfully posed photograph rather than a frame out of a film.
39
+ * Four fixes were measured against the block above (which is 0/4 on gaze by
40
+ * itself):
41
+ *
42
+ * "This image is a scene start frame of a video." 0/4 gaze, **0/3 identity**
43
+ * "Film still from a feature film." gaze ✗ (sampled)
44
+ * "A candid moment, unposed, nobody aware of…" identity ✗ (sampled)
45
+ * rewriting the verbs (drinking, not holding) 4/4 gaze, 3/4 identity
46
+ * **"Nobody looks at the camera."** **4/4 gaze, 4/4 identity,
47
+ * 4/4 wardrobe**
48
+ *
49
+ * Only the short negative is free. Everything longer — a medium declaration, a
50
+ * genre label, a mood sentence — competes with the reference bindings and the
51
+ * references lose: the "scene start frame" arm put a woman from INSIDE a
52
+ * reference into the lead role in every draw. This is the same effect gvp's
53
+ * block already recorded on a different brief ("more instruction bought LESS
54
+ * compliance"), now measured twice.
55
+ *
56
+ * KEPT SEPARATE FROM THE RULES, deliberately. A portrait, a piece to camera, or
57
+ * a product shot WANTS the eyeline; suppressing it is a creative choice, not a
58
+ * correctness rule. They are two controls, defaulting independently.
59
+ */
60
+
61
+ /**
62
+ * Default-deny + likeness + compose. Goes ahead of the scene it governs.
63
+ *
64
+ * OPT-IN, NOT INJECTED. An earlier cut of this work defaulted it on for every
65
+ * image node; Tal's call was that the snippets are enough, and he is right for
66
+ * a wording still being learned — a platform-wide default would apply the
67
+ * current best guess to every job at once, and the evidence for WHICH block is
68
+ * best is still split (see REFERENCE_RULES_MULTI_PERSON).
69
+ *
70
+ * ONE STRING, TWO CONSUMERS — the `reference-lock` factory snippet and
71
+ * gvp/recast's own grounding. They drifted apart once already (the snippet and
72
+ * gvp each carried a different version, and neither knew the other existed),
73
+ * which is the whole reason this is a constant rather than two literals.
74
+ */
75
+ export const REFERENCE_RULES =
76
+ "Do not use anything from reference images unless specified explicitly. " +
77
+ "All elements taken from reference images must preserve likeness. " +
78
+ "Compose them naturally into a single image."
79
+
80
+ /**
81
+ * The same rules PLUS the two face clauses — for briefs that move elements
82
+ * BETWEEN people.
83
+ *
84
+ * NOT THE DEFAULT, and the reason is a genuine conflict in the evidence that
85
+ * should not be quietly resolved by whoever edits this next.
86
+ *
87
+ * The controlled comparison: on a four-reference brief with two faces and a
88
+ * garment crossing from one person to the other, this block moved that garment
89
+ * 4 times in 4 draws and {@link REFERENCE_RULES} moved it 0 in 4 (Fisher exact,
90
+ * p ≈ 0.03). On the other four criteria — faces correct, own wardrobe, props,
91
+ * location, composition — the two were identical, 12/12 both ways.
92
+ *
93
+ * Tal's counter-evidence: across his own volume of real jobs, the shorter block
94
+ * works better IN GENERAL. Both hold. The face clauses earn their place exactly
95
+ * when two faces are in play and elements cross between them; on a single
96
+ * subject, a product or a landscape, "do not alter face structure" is dead
97
+ * weight, and this codebase has already measured twice that more instruction
98
+ * buys less compliance.
99
+ *
100
+ * So the DEFAULT follows the population (the short block) and this is the tool
101
+ * you reach for on the composition case. Settling "in general" properly needs
102
+ * the same method applied across brief TYPES, not more draws of one brief.
103
+ */
104
+ export const REFERENCE_RULES_MULTI_PERSON =
105
+ "Do not take anything from the reference images unless specified explicitly. " +
106
+ "Do not alter face structure. Do not blend faces. " +
107
+ "Preserve the likeness of every element taken. " +
108
+ "Compose them naturally into a single image."
109
+
110
+ /**
111
+ * THE FRAMING PREFIX — a sentence fragment that swallows the scene after it.
112
+ *
113
+ * "Medium wide film still of" + "The person from reference image A wears…"
114
+ * reads as one phrase, and that is the whole trick. Lead with the SHOT SIZE
115
+ * (see {@link filmStillPrefix}).
116
+ *
117
+ * NO "CINEMATIC", and that word was in here for about ten minutes. "Film still"
118
+ * describes the KIND of picture — a frame lifted out of moving footage, which
119
+ * is what a UGC clip, a product video and a documentary all are too. "Cinematic"
120
+ * describes a REGISTER, and imposing one is wrong for most briefs. It is also
121
+ * the exact category of word every measured arm punished: a genre claim. The same idea as a standalone SENTENCE
122
+ * ("Film still from a feature film." / "This image is a scene start frame of a
123
+ * video.") measured badly-to-catastrophically: the sentence competes with the
124
+ * reference bindings and the references lose — the "scene start frame" arm put
125
+ * a face from INSIDE a reference into the lead role in 3 of 3 draws. The prefix
126
+ * costs nothing because it never makes a separate claim.
127
+ *
128
+ * What it buys is not just the eyeline. Tal's distinction, which is sharper
129
+ * than the binary this was first scored on: a subject may be turned toward the
130
+ * lens and still be IN the scene rather than presenting to the viewer. The
131
+ * prefix also moves staging, depth and the quality of light — things an
132
+ * eyeline rule cannot reach.
133
+ */
134
+ export const FILM_STILL_PREFIX = "Film still of"
135
+
136
+ /**
137
+ * The prefix with a SHOT SIZE in front — "Extreme wide cinematic film still of
138
+ * …", "Medium close-up cinematic film still of …".
139
+ *
140
+ * Leading with the framing is standard practice and Tal reports it better
141
+ * again. HONEST STATUS: the POSITION is measured (a prefix costs nothing where
142
+ * a standalone claim cost the lead's identity 3 of 3); the shot-size word is
143
+ * his experience, not a controlled arm. The shot size is safe in a way a genre
144
+ * label is not — it says how the picture is FRAMED, which the scene needs
145
+ * anyway, rather than what kind of production it belongs to.
146
+ */
147
+ export function filmStillPrefix(shotSize?: string): string {
148
+ const shot = shotSize?.trim()
149
+ return shot ? `${shot} film still of` : FILM_STILL_PREFIX
150
+ }
151
+
152
+ /**
153
+ * The eyeline suppressor. Five words, measured free — it costs nothing on
154
+ * identity, wardrobe or composition, which none of the longer phrasings
155
+ * managed.
156
+ *
157
+ * DO NOT EXPAND THIS. Every attempt to say more about what the picture IS cost
158
+ * reference fidelity; the sentence works BECAUSE it constrains exactly one
159
+ * thing and claims nothing about medium, genre or mood.
160
+ */
161
+ export const SCENE_FRAME_RULE = "Nobody looks at the camera."
162
+
163
+ /**
164
+ * A LOOK TAIL — film stock, lens, light, palette — appended AFTER everything.
165
+ *
166
+ * An example to edit, not a universal: a different film wants a different
167
+ * stock. What generalises is the POSITION.
168
+ *
169
+ * THIS one is allowed to say "cinematic" — imposing a register is its entire
170
+ * job, and a user opts in by name. {@link filmStillPrefix} may not, because it
171
+ * is a DEFAULT: "film still" describes the kind of picture (a frame out of
172
+ * moving footage — true of a UGC clip and a documentary too), while
173
+ * "cinematic" describes a register most briefs did not ask for.
174
+ *
175
+ * ─── THE ONE RULE ALL OF THIS TURNED OUT TO BE ─────────────────────────────
176
+ *
177
+ * Position decides whether an instruction helps or fights the references:
178
+ *
179
+ * 1. RULES FIRST — default-deny, ahead of the bindings it governs.
180
+ * 2. FRAMING AS A PREFIX that swallows the subject ("Film still of …").
181
+ * 3. LOOK LAST — stock, lens, lighting, palette.
182
+ * 4. A separate CLAIM in the middle competes with the reference bindings,
183
+ * and the references lose.
184
+ *
185
+ * That is why "This image is a scene start frame of a video." (a standalone
186
+ * claim, mid-prompt) cost the lead's identity in 3 of 3 draws while "Film still
187
+ * of" (a prefix) costs nothing, and why this tail is free at the end. It
188
+ * matches the published guidance independently — Subject → Action →
189
+ * Surroundings → Camera/Lighting → Atmosphere — and gvp's engine had already
190
+ * found rule 3 the hard way: "THE MEDIUM GOES LAST… a rendering directive at
191
+ * the end has nothing after it to argue with."
192
+ *
193
+ * RECAST DOES NOT NEED THIS SNIPPET. `lookDirective` already appends a tail
194
+ * built from the look the ANALYSER observed in the source film — stock, grade,
195
+ * lens, lighting — which beats a hand-written one because it is that film's
196
+ * own look. This exists for the platform's image nodes, which have no analyser.
197
+ */
198
+ export const CINEMATIC_LOOK_TAIL =
199
+ "Shot on Super 16mm Kodak 7298 with Canon K35 lenses, soft naturalistic window light mixed with " +
200
+ "dim tungsten practicals and muted fluorescent spill, earthy muted palette with faded greens, " +
201
+ "warm skin tones and gentle shadow fall-off"
202
+
203
+ /**
204
+ * Compose the blocks a caller wants, in a stable order, ready to prepend.
205
+ * Returns `""` when everything is off, so a caller can prepend blind.
206
+ *
207
+ * No route calls this today — the platform ships these as SNIPPETS a user
208
+ * inserts. It exists for gvp/recast (which builds its grounding block in code)
209
+ * and for whatever calls this next, so the ordering rule lives in one place:
210
+ * rules first, eyeline rule after them, scene after both.
211
+ */
212
+ export function referenceRulesBlock(opts?: {
213
+ /** Default-deny + likeness + compose. Absent = ON. */
214
+ referenceRules?: boolean
215
+ /** Add the two face clauses — for briefs moving elements between people. */
216
+ multiPerson?: boolean
217
+ /** "Nobody looks at the camera." Absent = OFF (a portrait wants the eyeline). */
218
+ sceneFrame?: boolean
219
+ }): string {
220
+ const parts: string[] = []
221
+ if (opts?.referenceRules !== false) {
222
+ parts.push(opts?.multiPerson === true ? REFERENCE_RULES_MULTI_PERSON : REFERENCE_RULES)
223
+ }
224
+ if (opts?.sceneFrame === true) parts.push(SCENE_FRAME_RULE)
225
+ return parts.join(" ")
226
+ }