citeget 0.2.2__tar.gz → 0.2.3__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (43) hide show
  1. citeget-0.2.3/.claude/skills/add-references/SKILL.md +389 -0
  2. citeget-0.2.3/.claude/skills/add-references/add-references.skill +0 -0
  3. citeget-0.2.3/.claude/skills/add-references/references/TEMPLATE.md +64 -0
  4. citeget-0.2.3/.claude/skills/add-references/references/ai-ml.md +191 -0
  5. citeget-0.2.3/.claude/skills/add-references/references/math-combinatorics.md +177 -0
  6. citeget-0.2.3/.claude/skills/add-references/references/parameters.md +200 -0
  7. citeget-0.2.3/.claude/skills/add-references/references/social-sciences.md +235 -0
  8. citeget-0.2.3/.claude/skills/add-references/references/tech-blog.md +161 -0
  9. {citeget-0.2.2 → citeget-0.2.3}/PKG-INFO +1 -1
  10. citeget-0.2.3/cookies.txt +4 -0
  11. {citeget-0.2.2 → citeget-0.2.3}/pyproject.toml +1 -1
  12. {citeget-0.2.2 → citeget-0.2.3}/.claude/skills/acquire-references/SKILL.md +0 -0
  13. {citeget-0.2.2 → citeget-0.2.3}/.claude/skills/check-submission-fit/SKILL.md +0 -0
  14. {citeget-0.2.2 → citeget-0.2.3}/.claude/skills/fetch-resources/SKILL.md +0 -0
  15. {citeget-0.2.2 → citeget-0.2.3}/.claude/skills/format-for-journal/SKILL.md +0 -0
  16. {citeget-0.2.2 → citeget-0.2.3}/.claude/skills/prepare-submission/SKILL.md +0 -0
  17. {citeget-0.2.2 → citeget-0.2.3}/.claude/skills/research-topic/SKILL.md +0 -0
  18. {citeget-0.2.2 → citeget-0.2.3}/.claude/skills/review-article/SKILL.md +0 -0
  19. {citeget-0.2.2 → citeget-0.2.3}/.gitattributes +0 -0
  20. {citeget-0.2.2 → citeget-0.2.3}/.github/workflows/ci.yml +0 -0
  21. {citeget-0.2.2 → citeget-0.2.3}/.gitignore +0 -0
  22. {citeget-0.2.2 → citeget-0.2.3}/CHANGELOG.md +0 -0
  23. {citeget-0.2.2 → citeget-0.2.3}/CLAUDE.md +0 -0
  24. {citeget-0.2.2 → citeget-0.2.3}/LICENSE +0 -0
  25. {citeget-0.2.2 → citeget-0.2.3}/README.md +0 -0
  26. {citeget-0.2.2 → citeget-0.2.3}/citeget/__init__.py +0 -0
  27. {citeget-0.2.2 → citeget-0.2.3}/citeget/__main__.py +0 -0
  28. {citeget-0.2.2 → citeget-0.2.3}/citeget/acquire_references.py +0 -0
  29. {citeget-0.2.2 → citeget-0.2.3}/citeget/article_pub/__init__.py +0 -0
  30. {citeget-0.2.2 → citeget-0.2.3}/citeget/article_pub/data/journal_profiles.json +0 -0
  31. {citeget-0.2.2 → citeget-0.2.3}/citeget/article_pub/scripts/check_article.py +0 -0
  32. {citeget-0.2.2 → citeget-0.2.3}/citeget/article_pub/scripts/extract_references.py +0 -0
  33. {citeget-0.2.2 → citeget-0.2.3}/citeget/article_pub/scripts/word_count.py +0 -0
  34. {citeget-0.2.2 → citeget-0.2.3}/citeget/cli.py +0 -0
  35. {citeget-0.2.2 → citeget-0.2.3}/citeget/core.py +0 -0
  36. {citeget-0.2.2 → citeget-0.2.3}/citeget/data/libgen_vg_report.md +0 -0
  37. {citeget-0.2.2 → citeget-0.2.3}/citeget/extract.py +0 -0
  38. {citeget-0.2.2 → citeget-0.2.3}/citeget/fetch.py +0 -0
  39. {citeget-0.2.2 → citeget-0.2.3}/citeget/resolve.py +0 -0
  40. {citeget-0.2.2 → citeget-0.2.3}/tests/__init__.py +0 -0
  41. {citeget-0.2.2 → citeget-0.2.3}/tests/test_extract.py +0 -0
  42. {citeget-0.2.2 → citeget-0.2.3}/tests/test_finalize.py +0 -0
  43. {citeget-0.2.2 → citeget-0.2.3}/tests/test_resolve.py +0 -0
@@ -0,0 +1,389 @@
1
+ ---
2
+ name: add-references
3
+ description: >
4
+ Find and add references to documents. Use this skill whenever a user provides text (usually markdown)
5
+ and wants references, citations, or sources added to it. This includes requests like "add references to
6
+ this document", "find citations for this", "source this article", "add links to back up the claims",
7
+ "reference this for publication", "find supporting evidence", or any variation where the user has
8
+ written content and wants it backed by credible sources. Also use when the user asks to generate a
9
+ references section, find sources for specific claims, or validate statements with citations. Trigger
10
+ even for indirect requests like "make this more credible", "this needs sources", or "prepare this for
11
+ submission". Works for blog posts, academic papers, grant proposals, reports, and any written content
12
+ that benefits from citations.
13
+ ---
14
+
15
+ # Reference Finder
16
+
17
+ Add high-quality references and citations to user-provided documents.
18
+
19
+ ## Overview
20
+
21
+ This skill takes user-provided text (typically markdown) and returns an equivalent version
22
+ with references added. How the referencing is done is governed by a set of parameters — the
23
+ skill's first job is to propose sensible defaults for these, get the user's buy-in, and
24
+ then proceed.
25
+
26
+ ## Parameters
27
+
28
+ Nine parameters control the referencing behavior. They are listed in detail in
29
+ `references/parameters.md`. Here is a summary:
30
+
31
+ | # | Parameter | What it controls | Example values |
32
+ |---|-----------|-----------------|----------------|
33
+ | 1 | **Context** | Domain, audience, venue | tech blog, NIH grant, journal article |
34
+ | 2 | **Citation scope** | What warrants a reference | empirical claims, tools, all non-obvious claims |
35
+ | 3 | **Citation style** | How refs appear in-text | Vancouver `[1]`, Author-Date `(Smith, 2023)`, inline hyperlinks, footnotes |
36
+ | 4 | **Bibliography format** | How the refs section looks | numbered, alphabetical, annotated, grouped, none |
37
+ | 5 | **Source criteria** | What counts as acceptable | peer-reviewed only, preprints OK, reputable blogs OK |
38
+ | 6 | **Research depth** | How hard to look | surface, standard, deep |
39
+ | 7 | **Output mode** | What to return | annotated document, refs list only, annotation report, diff |
40
+ | 8 | **Density** | How heavily to reference | light, moderate, comprehensive |
41
+ | 9 | **Gap handling** | What to do with unsourceable claims | flag, suggest softening, suggest removal, skip |
42
+
43
+ **Context (#1) drives defaults for all the others.** A user who says "tech blog" gets
44
+ inline hyperlinks, standard depth, moderate density, etc. — without having to think about
45
+ the rest. Any individual parameter can be overridden.
46
+
47
+ For the full menu of values each parameter can take, and what they mean, read
48
+ `references/parameters.md`.
49
+
50
+ Detailed profile-specific rules live in separate files in the `references/` directory.
51
+ Each file is a **referencing profile** — a set of citation conventions, source quality
52
+ criteria, density rules, and examples for a specific publication setting:
53
+
54
+ | Profile | File | When to use |
55
+ |---------|------|-------------|
56
+ | Technical blog | `references/tech-blog.md` | Software/AI/tech blog posts, Medium, dev.to, engineering blogs |
57
+ | Social sciences | `references/social-sciences.md` | Journal articles, grant proposals (NIH, NSF, ERC), health research |
58
+ | Mathematics | `references/math-combinatorics.md` | Combinatorics, graph theory, discrete math journals |
59
+ | AI/ML research | `references/ai-ml.md` | ML conferences (NeurIPS, ICML, ICLR), AI journals (JMLR, AIJ) |
60
+
61
+ To create a new profile, see "Maintaining profiles" at the end of this file.
62
+
63
+ ## Workflow
64
+
65
+ ### Step 1. Analyze and propose
66
+
67
+ Before any research, do a **cursory read** of the input text and any instructions the user
68
+ has provided. Also **probe for citeget** (see "citeget integration" below) — note
69
+ availability in your proposal if relevant.
70
+
71
+ From the cursory read, infer the best default for each parameter.
72
+
73
+ Then present a **configuration proposal** to the user. The proposal should be concise —
74
+ a quick table or bulleted list — with brief commentary on your reasoning.
75
+
76
+ **Highlight choices that are ambiguous or could go either way.** For example:
77
+ - The text reads like a blog post but has academic-level claims → note the tension
78
+ - The audience isn't clear → say what you're assuming and why
79
+ - The density is a judgment call → explain the tradeoff
80
+
81
+ The proposal should feel like a knowledgeable assistant saying "here's how I'd approach
82
+ this — anything you'd change?" Not a lengthy questionnaire.
83
+
84
+ **Example proposal** (adapt the format to what feels natural in context):
85
+
86
+ > Based on a quick read, here's how I'd approach the referencing:
87
+ >
88
+ > | Parameter | Proposed | Reasoning |
89
+ > |-----------|----------|-----------|
90
+ > | Context | Tech blog (AI/ML audience) | The tone and content suggest a Medium-style post for practitioners |
91
+ > | Citation scope | Empirical claims, named tools, benchmarks | Skipping well-known concepts (attention, fine-tuning) your audience knows |
92
+ > | Citation style | Inline hyperlinks | Standard for blog format |
93
+ > | Bibliography | None (links are self-contained) | Could add a "Further Reading" section if you want — let me know |
94
+ > | Source criteria | Official docs + papers; reputable blogs OK | No hard peer-review requirement given the venue |
95
+ > | Research depth | Standard | I'll verify sources actually support claims, but won't deep-dive into ref chains |
96
+ > | Output mode | Annotated document | Full text back with links woven in |
97
+ > | Density | Moderate (~8-12 refs for this length) | Enough to back key claims without cluttering |
98
+ > | Gap handling | Flag unsourceable claims | I'll call out anything I can't find a good source for |
99
+ >
100
+ > **One thing I'd flag:** You mention some performance numbers in section 3 without
101
+ > attribution — I'll try to find the original benchmarks, but if those are from your own
102
+ > experiments, let me know and I'll leave them unlinked.
103
+ >
104
+ > Want to adjust anything, or should I go ahead?
105
+
106
+ The user can then:
107
+ - **Approve** ("looks good", "go ahead") → proceed to Step 2
108
+ - **Adjust** ("make it more academic", "use Vancouver style", "go deeper on the ML
109
+ papers") → update the configuration and confirm
110
+ - **Discuss** ("what would change if this were for a journal?", "should I use footnotes
111
+ here?") → talk through the tradeoff, then confirm
112
+
113
+ Don't proceed to research until you have at least implicit confirmation.
114
+
115
+ **When to skip this step:** If the user has been very explicit about what they want (e.g.,
116
+ "add Vancouver-style references to this grant proposal for NIH, comprehensive density,
117
+ deep research"), you can compress the proposal to a single confirming sentence and proceed.
118
+ Use judgment — the goal is to save the user time, not to add a bureaucratic gate.
119
+
120
+ ### Step 2. Load profile rules
121
+
122
+ Read the appropriate profile from `references/` based on the confirmed configuration.
123
+ If the user's situation doesn't match an existing profile exactly, adapt the closest one
124
+ and note what you're doing.
125
+
126
+ ### Step 3. Scan and flag passages
127
+
128
+ Read the document carefully. Based on the confirmed **citation scope** and **density**,
129
+ identify passages that warrant references.
130
+
131
+ If the document already contains some citations, parse them first so you don't duplicate
132
+ work. If citeget is available, use `citeget.extract_references(text)` for this (it handles
133
+ multiple citation formats automatically). Otherwise, scan manually for existing `[N]`
134
+ brackets, inline hyperlinks, or author-date patterns.
135
+
136
+ Categorization of what to reference (adjusted by citation scope setting):
137
+
138
+ **Empirical claims** — statistics, data points, measurements, performance benchmarks.
139
+ Always reference these regardless of scope setting.
140
+
141
+ **Attributed ideas** — named studies, specific people's arguments, paraphrased work
142
+ from others. Always reference.
143
+
144
+ **Technical claims** — how a tool works, what a standard says, what an algorithm does.
145
+ Reference when scope is moderate or comprehensive.
146
+
147
+ **Conceptual origins** — the DRY principle, Agile methodology, Social Cognitive Theory.
148
+ Reference when scope is comprehensive, or when the audience may not know the origin.
149
+
150
+ **Tools and frameworks** — named software, libraries, platforms. Reference (link to
151
+ official source) when scope is moderate or comprehensive.
152
+
153
+ **Trend claims** — "adoption is growing", "the field is moving toward X". Reference
154
+ when any evidence exists; flag if no source found.
155
+
156
+ **Common knowledge** — skip unless the scope is explicitly set to comprehensive and the
157
+ audience is non-specialist.
158
+
159
+ ### Step 4. Research and find sources
160
+
161
+ This is the core work. The effort invested should match the confirmed **research depth**.
162
+
163
+ **Surface depth:**
164
+ - One or two searches per claim
165
+ - Take the first credible source that supports the claim
166
+ - Verify the URL loads and is relevant
167
+ - Appropriate for: blog posts about well-documented topics, internal reports
168
+
169
+ **Standard depth:**
170
+ - Multiple search queries per claim, varying terms to find the best source
171
+ - Prefer primary over secondary sources
172
+ - Fetch and verify that sources actually support the specific claim
173
+ - Check for more authoritative alternatives to the first hit
174
+ - Appropriate for: published blog posts, white papers, conference talks
175
+
176
+ **Deep depth:**
177
+ - Follow reference chains — read what sources cite, find the original
178
+ - Cross-reference across multiple databases (web search, arXiv, Google Scholar)
179
+ - Check for retractions, corrections, or superseding work
180
+ - Verify recency — is this still the best source or has it been updated?
181
+ - Appropriate for: journal articles, grant proposals, systematic reviews
182
+
183
+ Regardless of depth:
184
+
185
+ - **Never fabricate references.** No invented authors, titles, DOIs, or URLs.
186
+ - **Note when you can't find a source.** Handle according to the **gap handling** setting.
187
+ - **Prefer primary sources.** The original paper/docs, not someone's summary of it.
188
+ - **Supplement with citeget** (if available and context is academic): when web search
189
+ returns secondary coverage, try `citeget.search(title, topic="articles")` to find the
190
+ original paper. See "citeget integration" below.
191
+
192
+ ### Step 5. Assemble output
193
+
194
+ Return the result according to the confirmed **output mode**:
195
+
196
+ **Annotated document** — the full text with references woven in, formatted per the
197
+ confirmed citation style and bibliography format. Ready to use or nearly so.
198
+
199
+ **References list only** — a numbered/keyed list of sources, each tied to a passage or
200
+ claim in the text (by quoting a short phrase or giving a section/paragraph reference).
201
+ The source text is not modified.
202
+
203
+ **Annotation report** — a structured list: for each flagged passage, the passage text,
204
+ the suggested source(s), and any notes (e.g., "this claim may be overstated relative to
205
+ the source"). Useful when the author wants to insert citations themselves.
206
+
207
+ **Diff** — the annotated document, but with changes marked (e.g., using markdown
208
+ strikethrough/bold or a separate diff format). Useful for review before accepting.
209
+
210
+ ### Step 6. Summary
211
+
212
+ After the output, include a brief summary:
213
+ - How many references were added
214
+ - Any claims that couldn't be sourced (per gap handling policy)
215
+ - Passages where the author might want to add their own references (prior work, internal
216
+ data, personal experience)
217
+ - Any claims where the best available source is weaker than ideal
218
+ - If research depth was less than deep: note any areas that might benefit from deeper
219
+ digging
220
+ - If citeget is available and context is academic: offer to acquire PDFs for the
221
+ referenced papers (see "citeget integration" below)
222
+
223
+ ## Important guidelines
224
+
225
+ **Respect the author's voice.** You're adding references, not rewriting. Keep text
226
+ modifications to the minimum needed to incorporate citations.
227
+
228
+ **Accuracy over quantity.** Five well-chosen, verified references beat fifteen questionable
229
+ ones.
230
+
231
+ **Be transparent about uncertainty.** If a source is imperfect, say so. The author decides.
232
+
233
+ **Don't over-reference.** Even at comprehensive density, not every sentence needs a
234
+ citation. Judgment matters — the goal is a well-supported document, not a wall of brackets.
235
+
236
+ **Recency is contextual.** In AI, 2020 may be ancient. In social sciences, a 1977 seminal
237
+ paper may be exactly right.
238
+
239
+ ## citeget integration (optional)
240
+
241
+ This skill can optionally leverage [`citeget`](https://pypi.org/project/citeget/) — a
242
+ Python package for searching academic databases (Library Genesis), extracting references
243
+ from documents, and downloading papers. citeget is NOT required; the skill works fully
244
+ via web search alone. When available, it adds capabilities for academic/research contexts.
245
+
246
+ ### Probe for availability
247
+
248
+ At the start of any referencing session (before the parameter proposal), check whether
249
+ citeget is usable. In Claude Code, run:
250
+
251
+ ```python
252
+ try:
253
+ import citeget
254
+ _CITEGET = True
255
+ except ImportError:
256
+ _CITEGET = False
257
+ ```
258
+
259
+ In Claude.ai (no Python execution), assume citeget is NOT available unless the user
260
+ explicitly says otherwise.
261
+
262
+ If available, mention it in the parameter proposal — briefly, as one line:
263
+ > "citeget is installed — I can search academic databases and offer PDF acquisition."
264
+
265
+ ### Integration points
266
+
267
+ There are three specific places in the workflow where citeget adds value. Each is
268
+ independent — use whichever apply.
269
+
270
+ **1. Parse existing references (Step 3)**
271
+
272
+ If the input document already contains some citations, use citeget to parse them before
273
+ scanning for gaps:
274
+
275
+ ```python
276
+ from citeget import extract_references
277
+ result = extract_references(document_text)
278
+ existing_refs = result.references # list of Reference objects
279
+ ```
280
+
281
+ This avoids duplicating citations that are already present. The `extract_references`
282
+ function tries multiple extraction strategies automatically (section headers, `[N]`
283
+ patterns, bold-numbered entries) and reports confidence. It also supports a
284
+ `markdown_links` extractor for documents that use inline hyperlinks.
285
+
286
+ **2. Search academic databases (Step 4)**
287
+
288
+ When the context calls for scholarly references (social sciences, grants, research blogs)
289
+ and web search hasn't surfaced the right paper, use citeget's libgen search:
290
+
291
+ ```python
292
+ from citeget import search
293
+ results = search("cognitive behavioral therapy meta-analysis", topic="articles")
294
+ # Returns list of dicts: title, authors, year, doi, etc.
295
+ ```
296
+
297
+ This is a supplement to web search, not a replacement. Use it when:
298
+ - Web search returns news/blog coverage but you need the original paper
299
+ - You have a title or author and want to confirm it exists
300
+ - The user's context is academic and they'll want the actual PDF later
301
+
302
+ **3. Offer PDF acquisition (Step 6)**
303
+
304
+ After delivering the referenced document, if citeget is available AND the context is
305
+ academic (journal, grant, research), offer to acquire PDFs:
306
+
307
+ > "I added 23 references. Want me to acquire PDFs for them using citeget? I can
308
+ > search libgen, ArXiv, and other sources."
309
+
310
+ If the user accepts, you can either:
311
+ - Use the CLI: `citeget acquire <output_file> --preview`
312
+ - Use the Python API: `acquire_all_references(refs, download_dir=...)`
313
+
314
+ This is the user's choice — don't auto-download. Just offer.
315
+
316
+ ### When citeget is NOT available
317
+
318
+ Fall back entirely to web search. No functionality is lost for blog/web contexts. For
319
+ academic contexts, you can note in the summary that citeget could help acquire PDFs:
320
+ > "Tip: `pip install citeget` enables searching academic databases and downloading
321
+ > referenced papers as PDFs."
322
+
323
+ ### Checking for required dependencies
324
+
325
+ citeget itself may be installed but not fully functional. Its search and download features
326
+ require Playwright with Chromium. If you get errors about Playwright, note to the user:
327
+ > "citeget is installed but Playwright isn't set up. Run
328
+ > `python -m playwright install chromium` to enable academic paper search/download."
329
+
330
+ The `extract_references` function works without Playwright — it's pure Python regex.
331
+ So even a partial citeget install is useful for parsing existing references.
332
+
333
+ ## Usage in Claude Code
334
+
335
+ Works the same way. The user may provide input via a file path. Read the file, run the
336
+ workflow, write output to `{original_name}_referenced.md` (or as specified). The
337
+ propose-confirm step happens in the terminal conversation just as it would in Claude.ai.
338
+
339
+ In Claude Code specifically, citeget integration is most useful — you have Python
340
+ execution, filesystem access for downloads, and the user's local environment where
341
+ citeget is likely installed and configured.
342
+
343
+ ## Usage in Claude.ai
344
+
345
+ The skill works in Claude.ai with web search as the sole research tool. citeget is
346
+ generally NOT available here (no local Python environment). The propose-confirm workflow,
347
+ context-specific referencing rules, and all output modes work identically.
348
+
349
+ If the user has uploaded a Python environment or explicitly mentions citeget availability,
350
+ the code execution sandbox can be used — but this is an edge case.
351
+
352
+ ## Maintaining profiles
353
+
354
+ Each `.md` file in `references/` is a **referencing profile**. Profiles are the main
355
+ extension point for this skill — adding a new publication domain means adding a new file.
356
+
357
+ ### Creating a new profile
358
+
359
+ 1. Copy `references/TEMPLATE.md` to a new file (e.g., `references/legal-brief.md`)
360
+ 2. Fill in each section following the template's guidance
361
+ 3. Add a row to the profile table in this SKILL.md (the one in the Parameters section)
362
+
363
+ The template has six sections — fill in all of them:
364
+ - **Citation format**: How refs appear in-text and in the bibliography
365
+ - **Source quality criteria**: Tiered list of what counts as credible
366
+ - **Reference density**: How heavily to cite, with rough targets
367
+ - **Common patterns**: 3-4 examples showing typical citation usage
368
+ - **Recency guidelines**: How old is too old, by source type
369
+ - **Special considerations**: Domain-specific quirks
370
+
371
+ Research the target journals/venues before writing a profile. Check their author
372
+ guidelines for citation style requirements, and look at published papers to see actual
373
+ citation practice (which often diverges from stated guidelines).
374
+
375
+ ### Editing an existing profile
376
+
377
+ Just edit the file. Changes take effect immediately — no re-installation needed in
378
+ Claude Code (it reads from disk). For Claude.ai, re-zip and re-upload the skill.
379
+
380
+ ### Removing a profile
381
+
382
+ Delete the file from `references/` and remove its row from the profile table in this
383
+ SKILL.md. The skill will continue to work — it just won't offer that profile as an option.
384
+
385
+ ### Profile naming convention
386
+
387
+ Use lowercase kebab-case: `tech-blog.md`, `social-sciences.md`, `math-combinatorics.md`.
388
+ The filename should suggest both the domain and the venue type when they differ (e.g.,
389
+ `ai-ml.md` covers both conferences and journals because the conventions are similar).
@@ -0,0 +1,64 @@
1
+ # Profile: [Name] — [Short description]
2
+
3
+ Referencing conventions for [describe the domain, audience, and typical venues].
4
+
5
+ ## Citation format
6
+
7
+ [Describe the in-text citation format. Include examples in markdown code blocks.
8
+ Cover: how citations appear in running text, how multiple citations are handled,
9
+ and whether narrative vs. parenthetical forms are used.]
10
+
11
+ ### Bibliography/References section format
12
+
13
+ [Show the format for the references list at the end. Include 2-3 example entries
14
+ covering: journal article, book/chapter, web resource, conference paper.]
15
+
16
+ ## Source quality criteria
17
+
18
+ ### Tier 1 — Strongest/most preferred
19
+
20
+ - [List the gold-standard source types for this domain]
21
+
22
+ ### Tier 2 — Good supporting sources
23
+
24
+ - [List acceptable but non-primary source types]
25
+
26
+ ### Tier 3 — Use with justification
27
+
28
+ - [List sources that are OK in limited circumstances]
29
+
30
+ ### Avoid
31
+
32
+ - [List sources that should not be used in this context]
33
+
34
+ ## Reference density
35
+
36
+ [Describe expected citation density. Include rough targets per section or per
37
+ 1000 words. Note any sections that are naturally denser or sparser.]
38
+
39
+ ## Common patterns
40
+
41
+ [Provide 3-4 markdown examples showing how citations typically appear in this
42
+ domain. Cover the most frequent citation scenarios.]
43
+
44
+ ### [Pattern name 1]
45
+
46
+ ```markdown
47
+ [Example text with citations]
48
+ ```
49
+
50
+ ### [Pattern name 2]
51
+
52
+ ```markdown
53
+ [Example text with citations]
54
+ ```
55
+
56
+ ## Recency guidelines
57
+
58
+ [Describe how source age matters in this domain. Which types of sources need to
59
+ be recent? Which can be old if they're seminal/foundational?]
60
+
61
+ ## Special considerations
62
+
63
+ [Any domain-specific quirks: self-citation norms, preprint policies, software
64
+ citation requirements, data citation, conference vs. journal prestige, etc.]
@@ -0,0 +1,191 @@
1
+ # Profile: AI/ML Research — Conferences and Journals
2
+
3
+ Referencing conventions for machine learning and artificial intelligence research
4
+ published at top conferences (NeurIPS, ICML, ICLR, ACL, CVPR, AAAI, etc.) and journals
5
+ (JMLR, Artificial Intelligence, TMLR, IEEE TPAMI, etc.).
6
+
7
+ ## Citation format
8
+
9
+ AI/ML venues overwhelmingly use **author-date (natbib) style**, with flexibility on
10
+ the exact format as long as it's internally consistent. The two forms:
11
+
12
+ **Parenthetical** — when the citation is not part of the sentence grammar:
13
+
14
+ ```markdown
15
+ Transformers have become the dominant architecture for NLP (Vaswani et al., 2017)
16
+ and have since been applied to vision (Dosovitskiy et al., 2021), audio
17
+ (Radford et al., 2023), and multi-modal tasks (Alayrac et al., 2022).
18
+ ```
19
+
20
+ **Narrative** — when the author is the grammatical subject:
21
+
22
+ ```markdown
23
+ Vaswani et al. (2017) introduced the Transformer architecture, which replaced
24
+ recurrence with self-attention. Brown et al. (2020) later demonstrated that
25
+ scaling language models to 175B parameters yields strong few-shot performance.
26
+ ```
27
+
28
+ Multiple citations: `(Smith et al., 2020; Jones & Lee, 2021; Doe et al., 2022)` —
29
+ typically ordered chronologically or by relevance.
30
+
31
+ Most venues accept any consistent style, but in practice the community has converged on
32
+ natbib-style author-date. Numbered `[1]` style is acceptable but less common in ML.
33
+
34
+ ### References section format
35
+
36
+ Alphabetical by first author surname. Include all authors (no "et al." in the reference
37
+ list itself). Conference papers include the venue abbreviation and year:
38
+
39
+ ```markdown
40
+ ## References
41
+
42
+ Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X.,
43
+ Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S.,
44
+ Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers
45
+ for image recognition at scale. In *ICLR*, 2021.
46
+
47
+ Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N.,
48
+ Kaiser, Ł., and Polosukhin, I. Attention is all you need. In *Advances in
49
+ Neural Information Processing Systems (NeurIPS)*, 2017.
50
+ ```
51
+
52
+ For journal articles, include volume and page numbers. For arXiv-only papers, include
53
+ the arXiv ID. Always include URLs or DOIs for reproducibility.
54
+
55
+ ## Source quality criteria
56
+
57
+ ### Tier 1 — Core ML literature
58
+
59
+ - **Top-venue conference papers**: NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, ECCV,
60
+ AAAI, IJCAI — these are the primary publication venues in ML (unlike most fields
61
+ where journals dominate, ML is conference-driven)
62
+ - **Top journals**: JMLR, TMLR, Artificial Intelligence, IEEE TPAMI, TACL
63
+ - **Foundational papers**: Regardless of venue, papers that introduced key concepts
64
+ (backprop, attention, GANs, etc.)
65
+
66
+ ### Tier 2 — Strong supporting sources
67
+
68
+ - **Workshop papers** from top venues (NeurIPS workshops, ICML workshops) — acceptable
69
+ for very recent work, but check if a full version exists
70
+ - **arXiv preprints** — essential in ML. Many landmark results (GPT-4, Llama, etc.)
71
+ are only on arXiv. Always acceptable, but note if not peer-reviewed.
72
+ - **Technical reports** from major labs (Google Research, DeepMind, OpenAI, Meta AI,
73
+ Anthropic, Microsoft Research) — de facto primary sources for many systems
74
+ - **Survey papers** — valuable for positioning and context
75
+
76
+ ### Tier 3 — Use with care
77
+
78
+ - **Blog posts from major labs** (Google AI Blog, OpenAI blog, Anthropic research) —
79
+ acceptable for system announcements and capability claims, but prefer the
80
+ accompanying paper when one exists
81
+ - **Older conference papers** from less selective venues — check citation count and
82
+ community reception
83
+ - **Dissertations** — cite the resulting papers instead when possible
84
+
85
+ ### Avoid
86
+
87
+ - Blog posts from non-researchers (unless citing for a specific factual claim about
88
+ a product/service)
89
+ - Medium articles, tutorials, Stack Overflow (for research claims)
90
+ - Wikipedia
91
+ - Predatory or pay-to-publish venues (check acceptance rates, editorial board quality)
92
+ - Press coverage of research (cite the paper, not the TechCrunch article about it)
93
+
94
+ ## Reference density
95
+
96
+ ML papers are **moderately to heavily cited**. The field has a strong culture of
97
+ positioning work relative to prior art:
98
+
99
+ - **Introduction**: High density. Motivate the problem, cite the main approaches, and
100
+ clearly state what's new. 10-20 citations in a 1-1.5 page intro is typical.
101
+ - **Related work**: Very high density. This section exists specifically to cite and
102
+ discuss prior work. 15-40 citations depending on breadth.
103
+ - **Method**: Moderate. Cite the techniques, architectures, and training procedures
104
+ you build on.
105
+ - **Experiments**: Cite baselines, benchmarks, datasets, and evaluation metrics.
106
+ - **Conclusion/Discussion**: Sparse. Cite future directions if pointing to specific
107
+ open problems.
108
+
109
+ **Overall**: A typical 8-page ML conference paper (plus unlimited references) has
110
+ 30-60 references. Some survey-style or empirical papers have 80-100+.
111
+
112
+ ## Common patterns
113
+
114
+ ### Introducing a line of work
115
+
116
+ ```markdown
117
+ Large language models (LLMs) have shown remarkable capabilities across a wide
118
+ range of tasks (Brown et al., 2020; Chowdhery et al., 2023; Touvron et al.,
119
+ 2023). Recent work has focused on improving their reasoning abilities through
120
+ chain-of-thought prompting (Wei et al., 2022) and reinforcement learning from
121
+ human feedback (Ouyang et al., 2022; Bai et al., 2022).
122
+ ```
123
+
124
+ ### Citing a specific method or architecture
125
+
126
+ ```markdown
127
+ Our model builds on the Vision Transformer (ViT) architecture (Dosovitskiy
128
+ et al., 2021), using the standard patch embedding and multi-head self-attention
129
+ layers. We add a cross-attention module following Alayrac et al. (2022).
130
+ ```
131
+
132
+ ### Citing benchmarks and datasets
133
+
134
+ ```markdown
135
+ We evaluate on MMLU (Hendrycks et al., 2021), GSM8K (Cobbe et al., 2021),
136
+ and HumanEval (Chen et al., 2021). Results are reported using pass@k
137
+ following the protocol of Chen et al. (2021).
138
+ ```
139
+
140
+ ### Citing concurrent or very recent work
141
+
142
+ ```markdown
143
+ Concurrent with our work, Li et al. (2024) proposed a similar approach using
144
+ retrieval-augmented generation. Our method differs in that we...
145
+ ```
146
+
147
+ ## Recency guidelines
148
+
149
+ ML moves extremely fast. Recency norms:
150
+
151
+ - **Benchmarks and SOTA**: Must be current. Citing a 2020 SOTA when a 2024 result
152
+ exists is a significant omission that reviewers will catch.
153
+ - **Architectures and techniques**: Cite the original paper regardless of age (e.g.,
154
+ Hochreiter & Schmidhuber, 1997 for LSTM), but also cite recent extensions or the
155
+ version you actually use.
156
+ - **Datasets**: Cite the original dataset paper. Note if using a modified version.
157
+ - **Scaling results**: These become outdated rapidly. Always cite the most recent
158
+ scaling study if making claims about model capabilities.
159
+ - **arXiv preprints**: Check regularly — a paper you cited as a preprint may now be
160
+ published at a venue. Update the citation.
161
+
162
+ **General rule**: Reviewers expect to see citations from the last 1-2 years. A related
163
+ work section with nothing newer than 2022 (as of 2025-2026) signals the authors aren't
164
+ current with the literature.
165
+
166
+ ## Special considerations
167
+
168
+ **Conference-driven field**: Unlike most academic fields, ML's prestige hierarchy is
169
+ conference → journal, not the other way around. A NeurIPS paper is typically more
170
+ prestigious than a journal publication in all but the top journals (JMLR, TMLR).
171
+
172
+ **Double-blind submission**: When submitting to double-blind venues (NeurIPS, ICML,
173
+ ICLR), refer to your own work in third person: "Smith et al. (2023) showed..." not
174
+ "In our prior work (Smith et al., 2023), we showed..." This affects how you write
175
+ the citations, not which ones you include.
176
+
177
+ **arXiv culture**: Posting to arXiv before or simultaneously with conference submission
178
+ is standard. Citing arXiv preprints is fully accepted. Include the arXiv ID
179
+ (e.g., `arXiv:2301.12345`) for preprints that haven't appeared at a venue.
180
+
181
+ **Reproducibility citations**: Cite code repositories, datasets, and computational
182
+ resources. The ML community increasingly values reproducibility, and reviewers notice
183
+ when key implementation details lack citation.
184
+
185
+ **Software and framework citations**: When using specific frameworks (PyTorch, JAX,
186
+ Hugging Face Transformers), cite them. Standard practice:
187
+ `(Paszke et al., 2019)` for PyTorch, `(Wolf et al., 2020)` for Transformers.
188
+
189
+ **Benchmark gaming**: Don't cherry-pick which baselines to compare against. Cite and
190
+ compare with the current best results, even if your method doesn't beat them. Reviewers
191
+ view selective citation of baselines very negatively.