@softspark/ai-toolkit 4.1.0 → 4.2.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,204 @@
1
+ # Google AI Pipeline
2
+
3
+ Reference for `seo-validate` — understanding how Google processes content for AI Search. Informs Category 6 (GEO) and Category 10 (Topical Authority) findings.
4
+
5
+ Source: Piotr Smargol, Indygo Agency — SEO Copywriting 2026.
6
+
7
+ ---
8
+
9
+ ## 1. The Four-Stage Pipeline
10
+
11
+ Every query in Google AI Search (AI Overviews, AI Mode, and classical SERPs) passes through the same pipeline. The presentation layer differs; the engine is the same.
12
+
13
+ ```
14
+ User Query
15
+
16
+
17
+ ┌─────────────────────────────────────────────────┐
18
+ │ Stage 1: PREPARE │
19
+ │ NLU — synonym mapping — intent classification │
20
+ │ Query Fan Out (synthetic sub-queries) │
21
+ └─────────────────────────────────────────────────┘
22
+
23
+
24
+ ┌─────────────────────────────────────────────────┐
25
+ │ Stage 2: RETRIEVE │
26
+ │ Content divided into chunks (≤500 tokens) │
27
+ │ Layout parsed — embeddings computed (Gecko) │
28
+ └─────────────────────────────────────────────────┘
29
+
30
+
31
+ ┌─────────────────────────────────────────────────┐
32
+ │ Stage 3: SIGNAL │
33
+ │ 7 ranking signals applied per chunk candidate │
34
+ └─────────────────────────────────────────────────┘
35
+
36
+
37
+ ┌─────────────────────────────────────────────────┐
38
+ │ Stage 4: SERVE │
39
+ │ Gemini 2.5 Flash generates the final answer │
40
+ │ Inline source citations added │
41
+ └─────────────────────────────────────────────────┘
42
+ ```
43
+
44
+ **Key implication for copywriters:** AI does not read your article — it extracts a chunk and pastes it as an answer. Each section must be a self-contained answer unit. If the answer is scattered across the article, AI cannot assemble it.
45
+
46
+ ---
47
+
48
+ ## 2. The Seven Ranking Signals
49
+
50
+ Google applies seven signals to decide which chunk from which source appears in the answer.
51
+
52
+ ### Signal 1 — Base Ranking
53
+ Classical relevance algorithm (PageRank descendants). SEO fundamentals still matter. High-authority domains get a floor advantage, but content structure can compensate.
54
+
55
+ **Copywriter implication:** Standard on-page SEO (title, H1, internal linking) remains the foundation.
56
+
57
+ ---
58
+
59
+ ### Signal 2 — Gecko Score (Semantic Embedding Similarity)
60
+ Gecko is Google's embedding model. It measures how semantically close your chunk is to the user's query — not keyword overlap, but meaning overlap. A chunk about "mattress firmness" will score well for "which mattress is best for back pain" even without that exact phrase.
61
+
62
+ **Copywriter implication:** Cover the topic from multiple angles and personas. Topical depth beats keyword density. A single cluster of related articles scores higher than one over-optimised page.
63
+
64
+ ---
65
+
66
+ ### Signal 3 — Jetstream (Cross-Attention)
67
+ Jetstream is the model's ability to understand context, contrasts, negations, and comparisons within a chunk. It specifically rewards:
68
+
69
+ - `X vs Y` comparisons
70
+ - Negations: "not recommended for…", "avoid if…", "unlike X…"
71
+ - Conditional statements: "best for…", "ideal for…", "if X, choose Y"
72
+ - Contrast phrases: "in contrast to X, Y does…", "unlike foam, latex…"
73
+
74
+ **Copywriter implication:** These constructions are not just reader-friendly — they are a direct Jetstream signal. Content with explicit comparisons and decision frameworks is over-represented in AI citations.
75
+
76
+ ---
77
+
78
+ ### Signal 4 — BM25 (Keyword Matching)
79
+ The classic TF-IDF/BM25 algorithm is still active. Exact keyword matches in headings and early paragraphs contribute. Not dominant, but present.
80
+
81
+ **Copywriter implication:** Include the primary keyword in H2 headings and in the first sentence of each section. Do not keyword-stuff — BM25 has diminishing returns and Gecko penalises unnatural density.
82
+
83
+ ---
84
+
85
+ ### Signal 5 — PCTR (Predicted Click-Through Rate)
86
+ Google estimates how likely users are to click your result based on historical CTR patterns for that title format, topic, and position. High PCTR boosts a chunk's probability of being included in the AI answer.
87
+
88
+ **Copywriter implication:** Title and meta description quality directly affects AI citation probability — not just traditional click traffic. Use benefit-led titles. Avoid clickbait (Google's AI pipeline detects and suppresses it). Brand at the end of the title.
89
+
90
+ ---
91
+
92
+ ### Signal 6 — Freshness
93
+ Content age is weighted by topic type:
94
+
95
+ | Topic type | Freshness weight |
96
+ |---|---|
97
+ | Breaking news | Dominant — stale content is excluded |
98
+ | Technology / AI / tools | High — 13-week threshold observed |
99
+ | Evergreen how-to | Moderate — update signals matter |
100
+ | Historical / definitions | Low — accuracy over recency |
101
+
102
+ **The 13-week rule:** Research across top AI-cited content shows 50% of cited results were published or updated within the last 13 weeks. Evergreen pages that are not refreshed drop out of AI citation pools gradually.
103
+
104
+ **Copywriter implication:** Add visible `dateModified` to all indexable pages. Refresh key articles before the 13-week threshold. Use "Updated: [date]" visible to users, not just in JSON-LD.
105
+
106
+ ---
107
+
108
+ ### Signal 7 — Boost/Bury Rules
109
+ Manual editorial adjustments applied by Google. Certain sources are systematically boosted (e.g., Reddit, YouTube, Wikipedia, government and academic domains). Certain patterns are buried (AI-generated content without human editing, thin content, pages without authors).
110
+
111
+ **Copywriter implication:** Named authors with bylines and bios protect content from Bury Rules. Generic AI-generated text without editorial review is actively suppressed.
112
+
113
+ ---
114
+
115
+ ## 3. Query Fan Out
116
+
117
+ When a user submits a query, the AI does not search only for that query. It generates dozens of synthetic sub-queries internally and retrieves candidate chunks for each.
118
+
119
+ ```
120
+ User query: "which laptop for remote work"
121
+
122
+ ┌──────────┴──────────┐
123
+ ▼ ▼
124
+ "best laptops remote work 2026" "lightweight laptop long battery"
125
+ │ │
126
+ ▼ ▼
127
+ "MacBook vs ThinkPad remote" "quiet fan laptop home office"
128
+ │ │
129
+ ▼ ▼
130
+ "budget laptop for Zoom" "laptop under 1000 remote work"
131
+
132
+
133
+ "best keyboard laptop coding"
134
+ … (50+ more sub-queries)
135
+ ```
136
+
137
+ ### The 95% problem
138
+
139
+ **95% of these sub-queries have zero Monthly Search Volume** in Semrush, Ahrefs, Senuto, or any keyword tool. Classical keyword research shows only 5% of the actual retrieval surface.
140
+
141
+ **Copywriter implication:**
142
+ - Stop optimising for a single phrase per page.
143
+ - Cover the topic from multiple persona angles: beginner vs expert, budget vs premium, use case A vs use case B.
144
+ - Answer implicit sub-questions within the same article using H2/H3 sections.
145
+ - Cluster architecture (pillar + cluster articles) is the structural response to Query Fan Out.
146
+
147
+ ---
148
+
149
+ ## 4. Probabilistic Ranking
150
+
151
+ In classical SEO, "position 1" was deterministic — the same for every user. In AI Search, ranking is probabilistic: two users sending the identical query can receive different answers.
152
+
153
+ AI personalises based on:
154
+ - Search history and session context
155
+ - Location and language
156
+ - Inferred demographic profile
157
+ - Device type
158
+ - Stochastic variation in model generation
159
+
160
+ **Implication:** There is no single position to win. There is a **probability of being cited**. Maximise that probability by:
161
+ 1. Covering the topic from multiple perspectives (different personas, contexts, formats).
162
+ 2. Being present in multiple chunk positions within a long article (each H2 is a citation candidate).
163
+ 3. Publishing cluster articles that address sub-queries individually.
164
+ 4. Maintaining freshness so your chunk stays in the retrieval pool.
165
+
166
+ ---
167
+
168
+ ## 5. Format Routing
169
+
170
+ The AI routes queries to content formats. If your format does not match the query type, you lose — even with superior content.
171
+
172
+ | Query pattern | Optimal content format | Schema signal |
173
+ |---|---|---|
174
+ | `how to X` | Numbered step-by-step tutorial | `HowTo` |
175
+ | `X vs Y` | Side-by-side comparison table | `Article` + table markup |
176
+ | `show me X` | Visual: image gallery or video | Image alt + video schema |
177
+ | `latest / best X in [year]` | Freshly updated article with date | `Article` + `dateModified` |
178
+ | `reviews of X` / `what do people think of X` | UGC aggregation or forum discussion | `Review`, `AggregateRating` |
179
+ | `what is X` | Definition paragraph + entity context | `Article` with entity-first opening |
180
+ | `X for [persona]` | Persona-targeted guide | `Article` + `FAQPage` |
181
+ | `is X worth it` | Opinionated recommendation with evidence | `Article` with author + data |
182
+
183
+ **Copywriter implication:** Before writing, determine the dominant query pattern for the page's topic. Structure the page to match that format. Mismatched format = invisible to AI routing.
184
+
185
+ ---
186
+
187
+ ## 6. The Two-Audience Reality
188
+
189
+ | Audience | Wins with | Metric |
190
+ |---|---|---|
191
+ | Fast AI (simple information queries) | Structured chunks, direct answers, semantic triples | AI citation probability |
192
+ | Human readers (purchase decisions, complex topics) | Narrative, trust signals, social proof, UGC | Engagement, conversion |
193
+
194
+ **Strategy:** Pages must serve both simultaneously. Chunk-optimised structure aids AI extraction without harming human readability — they are the same discipline (clear headings, direct answers, no filler).
195
+
196
+ ---
197
+
198
+ ## References
199
+
200
+ - Google Search Central — How Google Search works: https://developers.google.com/search/docs/fundamentals/how-search-works
201
+ - Google Gecko embedding model: https://research.google/pubs/gecko-versatile-text-embeddings-distilled-from-large-language-models/
202
+ - Blyskall study — 40M AI Overviews results, backlink vs content signals: referenced in SEO Copywriting 2026 (Smargol, Indygo Agency)
203
+ - Senuto topical authority study — 212K phrases, 7,200 semantic groups
204
+ - Google AI Mode announcement (March 2025): https://blog.google/products/search/google-ai-mode-search/
@@ -0,0 +1,372 @@
1
+ # Content Citability Patterns
2
+
3
+ Reference for `seo-validate` Category 6 (GEO). Covers the structural and linguistic patterns that maximise AI citation probability. Companion to [geo-guidelines.md](geo-guidelines.md) and [ai-pipeline.md](ai-pipeline.md).
4
+
5
+ Source: Piotr Smargol, Indygo Agency — SEO Copywriting 2026.
6
+
7
+ ---
8
+
9
+ ## 1. Chunk Architecture
10
+
11
+ ### The 500-token boundary
12
+
13
+ Google's retrieval stage splits content into chunks of **maximum 500 tokens (~375 words)**. AI does not summarise your article — it extracts one chunk and uses it as the answer. If the answer to a question is distributed across multiple sections or buried after a long preamble, AI cannot assemble it.
14
+
15
+ **Rules:**
16
+ - Each H2 section should stay within ~375 words.
17
+ - If a section exceeds this, split with an H3 sub-heading — each H3 becomes its own chunk candidate.
18
+ - Never put the answer at the end of a section after a long setup.
19
+
20
+ ### Anatomy of a citable chunk
21
+
22
+ Every section should follow this structure:
23
+
24
+ ```
25
+ H2: [Question or keyword-rich title]
26
+
27
+ ├── Direct answer — 2–3 sentences. No preamble. The fact first.
28
+
29
+ ├── Elaboration — data, context, nuance. 3–5 sentences.
30
+
31
+ ├── Visual element — list, table, or code block.
32
+
33
+ └── TL;DR (optional, for long sections) — 1-sentence summary.
34
+ ```
35
+
36
+ ### Before / after examples
37
+
38
+ **Before (unchunkable):**
39
+ ```
40
+ H2: Choosing the Right Mattress
41
+
42
+ When you're looking for a new mattress, there are many factors to consider.
43
+ The market offers a wide variety of options, and it can be overwhelming to
44
+ navigate. Let's explore the key aspects you should think about before making
45
+ a purchase decision, because getting this right is important for your sleep
46
+ quality and long-term health. After all, we spend a third of our lives in bed.
47
+
48
+ Eventually, firmness is one of the most important factors. For people under
49
+ 70 kg, a softer H1–H2 rating works well.
50
+ ```
51
+ *Problem: answer buried after 80-word preamble; AI extracts the preamble as the "answer".*
52
+
53
+ **After (chunk-optimised):**
54
+ ```
55
+ H2: How to Choose Mattress Firmness by Body Weight
56
+
57
+ Match firmness to your weight: under 70 kg → H1 or H2; 70–90 kg → H3;
58
+ over 90 kg → H4. Higher body weight needs firmer support to maintain
59
+ spinal alignment.
60
+
61
+ | Weight | Recommended firmness |
62
+ |--------|----------------------|
63
+ | < 70 kg | H1–H2 (soft–medium) |
64
+ | 70–90 kg | H3 (medium-firm) |
65
+ | > 90 kg | H4 (firm) |
66
+
67
+ **TL;DR:** Firmness = body weight ÷ 10, rounded up to the nearest H-rating.
68
+ ```
69
+ *Result: AI extracts the first two sentences as the answer; table is bonus precision.*
70
+
71
+ ---
72
+
73
+ ## 2. Semantic Triples
74
+
75
+ ### What they are
76
+
77
+ A semantic triple is the simplest factual statement AI can parse without inference or guessing:
78
+
79
+ ```
80
+ Subject → Predicate → Object
81
+ ```
82
+
83
+ LLMs are trained on structured knowledge graphs built from triples. Content written as triples is extracted with high confidence and low hallucination risk.
84
+
85
+ ### Triple structure
86
+
87
+ | Component | Role | Example |
88
+ |---|---|---|
89
+ | Subject | What/who is described | "Premium Mattress" |
90
+ | Predicate | Property or relationship | "has firmness rating" |
91
+ | Object | Value or target | "H3" |
92
+
93
+ Assembled: "The Premium Mattress has a firmness rating of H3."
94
+
95
+ ### Stacking triples
96
+
97
+ One product/topic should generate multiple triples covering all queryable properties:
98
+
99
+ ```
100
+ Product X → has firmness → H3
101
+ Product X → measures → 160 × 200 cm
102
+ Product X → has height → 22 cm
103
+ Product X → is ideal for → side sleepers weighing 70–90 kg
104
+ Product X → is not recommended for → stomach sleepers
105
+ Product X → pairs with → slatted base with max 5 cm gap
106
+ ```
107
+
108
+ Each triple answers a different AI sub-query. Six triples = six citation opportunities.
109
+
110
+ ### Marketing prose vs semantic triples
111
+
112
+ | Style | Example | AI extractable? |
113
+ |---|---|---|
114
+ | Marketing prose | "Our exceptional collection will enchant you with its elegance and superior comfort." | No — zero extractable facts |
115
+ | Semantic triples | "Firmness: H3. Dimensions: 160×200 cm. Height: 22 cm. Ideal for side sleepers, 70–90 kg. Not recommended for stomach sleepers." | Yes — 5 distinct facts |
116
+
117
+ **Rule:** Every product description, category page, and guide should contain a minimum of three semantic triples in the opening paragraph or a specification table.
118
+
119
+ ### Grep patterns
120
+
121
+ ```
122
+ # Detect marketing filler (high prose density, low factual density)
123
+ Pattern: \b(exceptional|extraordinary|unique|enchant|fascinate|remarkable|unparalleled|outstanding)\b
124
+ Files: *.html, *.md, *.tsx, *.jsx, *.vue, *.svelte, *.astro
125
+ Flag: marketing superlative without associated factual triple
126
+ ```
127
+
128
+ ---
129
+
130
+ ## 3. Opinionated Content vs Hedging Language
131
+
132
+ ### Why AI skips hedged claims
133
+
134
+ AI answer engines extract concrete statements to synthesise answers. A hedged claim ("this might be a good choice for many people") cannot be used as a direct answer to "which mattress should I buy?" AI silently passes over it.
135
+
136
+ **Google's Jetstream signal explicitly rewards** opinionated, declarative content over neutral, hedged content.
137
+
138
+ ### Hedging patterns to eliminate
139
+
140
+ | Hedging phrase | Why it fails | Replacement |
141
+ |---|---|---|
142
+ | "may be a good choice" | Cannot be cited as a recommendation | "We recommend X for Y" |
143
+ | "might work well for" | Conditional — AI skips | "Works best for" |
144
+ | "worth considering" | No stance | "Our top pick for Z is X" |
145
+ | "for many people" | Undefined persona | "for side sleepers weighing 70–90 kg" |
146
+ | "could be ideal" | Speculative | "Is ideal for" |
147
+ | "one option is" | Non-committal | "Choose X if you need Y" |
148
+ | "it depends" (without resolution) | No extractable answer | "It depends on your weight: under 70 kg → H2, over 90 kg → H4" |
149
+
150
+ ### Opinionated writing rules
151
+
152
+ 1. **Take a position.** "We recommend X for Y" over "X is worth considering."
153
+ 2. **Name the persona.** "Ideal for side sleepers weighing 70–90 kg" over "suitable for most users."
154
+ 3. **Justify the claim.** AI cites "best for back pain because H3 stabilises the lumbar spine" — not "best for back pain" alone.
155
+ 4. **Use first-person plural for brand voice.** "We tested X and found…" establishes E-E-A-T Experience signal.
156
+
157
+ ### Grep patterns for hedging detection
158
+
159
+ ```
160
+ # High-priority hedging (recommendation context)
161
+ Pattern: \b(may be|might be|could be|worth considering|for many|for most people|it depends)\b
162
+ Files: *.html, *.md, *.tsx, *.jsx, *.vue, *.astro
163
+ Context: within 50 words of product name, category heading, or recommendation heading
164
+ Severity: INFO
165
+ ```
166
+
167
+ ---
168
+
169
+ ## 4. Decision Frameworks
170
+
171
+ ### Why they dominate AI citations
172
+
173
+ "If X → choose Y" is the most frequently cited construction in AI Search. It directly answers the user's decision-making intent, requires no interpretation, and maps cleanly to Query Fan Out sub-queries ("which X for Y?").
174
+
175
+ ### Formats
176
+
177
+ **Conditional (if/then):**
178
+ ```
179
+ If you sleep on your side → choose H2–H3 firmness.
180
+ If you share a bed with a partner of significantly different weight → choose a split mattress.
181
+ If you have chronic lower back pain → choose H3 with lumbar zone reinforcement.
182
+ ```
183
+
184
+ **Use-case mapping:**
185
+ ```
186
+ For home office video calls → prioritise battery life and weight under 1.5 kg.
187
+ For graphic design work → prioritise display colour gamut (>95% DCI-P3).
188
+ For travel → prioritise weight and keyboard quality for extended typing.
189
+ ```
190
+
191
+ **Elimination framework (negative definition):**
192
+ ```
193
+ NOT for: stomach sleepers, children under 12, anyone over 100 kg.
194
+ ```
195
+
196
+ ### Placement
197
+
198
+ Decision frameworks belong:
199
+ - In H2 "How to Choose" or "Which X for You" sections.
200
+ - In product/category opening paragraphs (chunk position 1).
201
+ - In FAQ answers (each answer = one framework).
202
+ - In comparison tables (column "Best for:").
203
+
204
+ ### Grep patterns
205
+
206
+ ```
207
+ # Detect absence of decision framework in guide/category pages
208
+ Detect: (?i)<h[2-3][^>]*>(how to choose|which .* for|guide to|best .* for)
209
+ Then check: (?i)\bif\b.{1,50}\b(choose|select|pick|go with|opt for)\b OR (?i)\bfor .{3,30}\b(choose|recommend|ideal|best)\b
210
+ Flag if: guide heading found, no decision framework pattern within 500 words
211
+ Severity: INFO
212
+ ```
213
+
214
+ ---
215
+
216
+ ## 5. Contrast Patterns (Jetstream Signal)
217
+
218
+ Contrasts and comparisons are a direct Jetstream (cross-attention) signal. Content with explicit "X vs Y" structures is over-represented in AI answers for comparative queries.
219
+
220
+ ### Contrast constructions
221
+
222
+ | Construction | Example |
223
+ |---|---|
224
+ | `X vs Y` heading | "Foam vs Latex Mattress: Which is Better?" |
225
+ | "Unlike X, Y…" | "Unlike foam, latex responds instantly to movement." |
226
+ | "In contrast to X…" | "In contrast to bonnell springs, pocket springs isolate motion." |
227
+ | "X is better than Y for Z" | "Latex is better than foam for hot sleepers because it has open-cell structure." |
228
+ | "Compared to X, Y offers…" | "Compared to H2 firmness, H3 provides 40% more lumbar support." |
229
+ | Comparison table | Explicit columns for two or more options |
230
+
231
+ ### When to use contrasts
232
+
233
+ - Every product category page should contain at least one "Category A vs Category B" section or table.
234
+ - Blog guides should compare at least two options before making a recommendation.
235
+ - Product descriptions should contrast with the alternative ("unlike [competitor type], this product…").
236
+
237
+ ### Grep patterns
238
+
239
+ ```
240
+ # Detect comparative content missing contrast constructions
241
+ Detect: (?i)\b(vs|versus|compare|comparison|difference between|X or Y)\b in headings/titles
242
+ Then check: (?i)(unlike|in contrast|compared to|better than|vs|versus) in body text
243
+ Then check: <table present in the content
244
+ Flag if: comparative heading found, no contrast language AND no comparison table
245
+ Severity: INFO
246
+ ```
247
+
248
+ ---
249
+
250
+ ## 6. Negative Definitions
251
+
252
+ Defining who/what a product is NOT for is a high-value signal for two reasons:
253
+
254
+ 1. **AI exclusion queries:** "which mattress is not for stomach sleepers" — only content with explicit negative definitions answers this.
255
+ 2. **Targeting precision:** Narrow persona definition increases citation probability for the right audience query.
256
+
257
+ ### Patterns
258
+
259
+ ```
260
+ Not recommended for: stomach sleepers, children under 12, bodyweight over 120 kg.
261
+ Not suitable for: use on adjustable bases without rigid slat support.
262
+ Avoid if: you prefer a cloud-soft feel — this mattress is designed for firm support.
263
+ This product is not designed for: outdoor use, temperatures below 5°C.
264
+ ```
265
+
266
+ ### Placement
267
+
268
+ - Product descriptions: one "Not recommended for:" block near the persona definition.
269
+ - Category pages: "Who this category is NOT for" subsection.
270
+ - FAQ: "Is X right for me?" answer should include negative persona.
271
+
272
+ ### Grep patterns
273
+
274
+ ```
275
+ # Check product/category pages for negative definition
276
+ Detect: (?i)(not recommended for|not suitable for|avoid if|not designed for|not ideal for)
277
+ Flag if: product or category page (detected by Product schema or URL pattern) has no negative definition
278
+ Severity: INFO
279
+ ```
280
+
281
+ ---
282
+
283
+ ## 7. Justified Superlatives
284
+
285
+ Superlatives without justification are ignored by AI ("best mattress" is marketing noise). Superlatives with evidence are cited.
286
+
287
+ ### Formula
288
+
289
+ ```
290
+ [Superlative claim] + [specific reason] + [evidence or data]
291
+ ```
292
+
293
+ ### Examples
294
+
295
+ | Bad (ignored) | Good (citable) |
296
+ |---|---|
297
+ | "The best mattress for back pain" | "The best mattress for back pain because H3 firmness maintains neutral lumbar alignment — validated by our customer data showing 78% pain reduction after 30 days." |
298
+ | "Most durable laptop battery" | "Longest real-world battery life in the under-1.5 kg category — 14.2 hours in our standardised office test." |
299
+ | "Premium quality" | "Premium build: aluminium chassis, ISO military-grade drop rating (MIL-STD-810H), 3-year warranty." |
300
+
301
+ ### Rule
302
+
303
+ Every superlative must be followed within the same sentence by "because", "with", "achieving", or a colon introducing a specific measurement.
304
+
305
+ ---
306
+
307
+ ## 8. The 13-Week Freshness Threshold
308
+
309
+ Research shows 50% of top AI-cited content was published or updated within the last 13 weeks (Blyskall, 40M AI Overviews results). After 13 weeks without a visible update, citation probability drops.
310
+
311
+ ### Freshness strategy
312
+
313
+ | Action | Frequency | Priority pages |
314
+ |---|---|---|
315
+ | Full content review and update | Every 12 weeks | Top-traffic pillar pages |
316
+ | Date + minor fact refresh | Every 13 weeks | High-competition cluster articles |
317
+ | Add new FAQ or data point | On new data availability | Product pages, comparison guides |
318
+ | Update `dateModified` in JSON-LD AND visible text | On every meaningful change | All indexable pages |
319
+
320
+ ### Visible freshness signals
321
+
322
+ Both are required:
323
+ 1. JSON-LD `dateModified` field.
324
+ 2. Visible "Updated: [date]" or "Last reviewed: [date]" text near the article byline.
325
+
326
+ JSON-LD alone is insufficient — AI engines parse visible text; users trust visible dates.
327
+
328
+ ---
329
+
330
+ ## 9. Author E-E-A-T (Experience Signal)
331
+
332
+ Google's most important E-E-A-T shift in 2025 was elevating **Experience** — content from someone who actually used the product, visited the place, or performed the task. This signal is impossible for pure AI generation to replicate.
333
+
334
+ ### Author requirements for citability
335
+
336
+ | Requirement | Implementation | AI signal |
337
+ |---|---|---|
338
+ | Real name (not "Admin" or "Team") | `<span class="author">Jane Doe</span>` + `Person` schema | Identity trust |
339
+ | Role/credentials in bio | ≥30 words of bio near author name | Expertise signal |
340
+ | `Person` schema with `sameAs` | LinkedIn, GitHub, ORCID, or personal site URL | Authoritativeness |
341
+ | First-person experience language | "I tested…", "In our lab…", "After 6 months of use…" | Experience signal |
342
+ | Date visible to user | "Published [date] · Updated [date]" | Freshness + trust |
343
+
344
+ ### Author schema
345
+
346
+ ```json
347
+ {
348
+ "@context": "https://schema.org",
349
+ "@type": "Article",
350
+ "author": {
351
+ "@type": "Person",
352
+ "name": "Jane Doe",
353
+ "jobTitle": "Senior Product Tester",
354
+ "sameAs": [
355
+ "https://linkedin.com/in/janedoe",
356
+ "https://github.com/janedoe"
357
+ ]
358
+ },
359
+ "datePublished": "2026-03-01",
360
+ "dateModified": "2026-04-20"
361
+ }
362
+ ```
363
+
364
+ ---
365
+
366
+ ## References
367
+
368
+ - Blyskall study — 40M AI Overviews, backlinks vs content structure signals
369
+ - Smargol, P. (2026). SEO Copywriting 2026. Indygo Agency. https://indygo.agency
370
+ - Aggarwal et al. (2023). GEO: Generative Engine Optimization. Princeton/Georgia Tech. https://arxiv.org/abs/2311.09735
371
+ - Google E-E-A-T: https://developers.google.com/search/docs/fundamentals/creating-helpful-content
372
+ - See also: [ai-pipeline.md](ai-pipeline.md), [geo-guidelines.md](geo-guidelines.md), [geo-aeo-patterns.md](geo-aeo-patterns.md)