extract-webpage 1.2.20 β†’ 1.2.22

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (2) hide show
  1. package/README.md +87 -158
  2. package/package.json +4 -4
package/README.md CHANGED
@@ -2,211 +2,140 @@
2
2
  <img width="350px" src="https://i.imgur.com/8JvNmxU.jpeg" />
3
3
  </p>
4
4
  <p align="center">
5
- Being is Becoming<br />
6
- Whatever Research Can Be,<br />
7
- That is What It Must Become. <br />
8
- If AI is Humanity's Last Invention, <br />
9
- Then Vector Space is the Final Frontier.<br />
10
- </p>
11
- <p align="center">
12
- <br />
13
- <a href="https://npmjs.org/package/ai-research-agent">
14
- <img src="https://i.imgur.com/HZUZUP0.png"
15
- alt="NPM badge for ai-research-agent" />
16
- </a>
17
- </p>
18
- <p align="center">
19
- <a href="https://discord.gg/SJdBqBz3tV">
20
- <img src="https://img.shields.io/discord/1110227955554209923.svg?label=Chat&logo=Discord&colorB=7289da&style=flat"
21
- alt="Join Discord" />
5
+ <a href="https://npmjs.org/package/extract-webpage">
6
+ <img src="https://img.shields.io/npm/v/extract-webpage" alt="NPM version" />
22
7
  </a>
23
- <a href="https://github.com/vtempest/ai-research-agent/discussions">
24
- <img alt="GitHub Stars" src="https://img.shields.io/github/stars/vtempest/ai-research-agent" /></a>
25
- <a href="https://github.com/vtempest/ai-research-agent/discussions">
26
- <img alt="GitHub Discussions"
27
- src="https://img.shields.io/github/discussions/vtempest/ai-research-agent" />
8
+ <a href="https://npmjs.org/package/extract-webpage">
9
+ <img alt="NPM Downloads" src="https://img.shields.io/npm/dy/extract-webpage" />
28
10
  </a>
29
- <a href="https://npmjs.org/package/ai-research-agent"><img src="https://img.shields.io/npm/v/ai-research-agent"/></a>
30
- <a href="https://github.com/vtempest/ai-research-agent/pulse" alt="Activity">
31
- <img src="https://img.shields.io/github/commit-activity/m/vtempest/ai-research-agent" />
32
- </a>
33
- <img src="https://img.shields.io/github/last-commit/vtempest/ai-research-agent.svg" alt="GitHub last commit" />
34
- </p>
35
- <p align="center">
36
- <a href="https://npmjs.org/package/ai-research-agent">
37
- <img alt="NPM Downloads" src="https://img.shields.io/npm/dy/ai-research-agent" />
11
+ <a href="https://discord.gg/SJdBqBz3tV">
12
+ <img src="https://img.shields.io/discord/1110227955554209923.svg?label=Chat&logo=Discord&colorB=7289da&style=flat" alt="Join Discord" />
38
13
  </a>
39
- <a href="https://github.com/vtempest/ai-research-agent/actions/workflows/docs.yml">
40
- <img src="https://github.com/vtempest/ai-research-agent/actions/workflows/docs.yml/badge.svg" alt="Build Status" />
14
+ <a href="https://github.com/OpenSourceAGI/qwksearch-research-agent/discussions">
15
+ <img alt="GitHub Stars" src="https://img.shields.io/github/stars/OpenSourceAGI/qwksearch-research-agent" />
41
16
  </a>
42
17
  <a href="https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/proposing-changes-to-your-work-with-pull-requests/creating-a-pull-request">
43
- <img src="https://img.shields.io/badge/PRs-welcome-brightgreen.svg"
44
- alt="PRs Welcome" />
45
- </a>
46
- <a href="https://codespaces.new/vtempest/ai-research-agent">
47
- <img src="https://github.com/codespaces/badge.svg" width="150" height="20" />
18
+ <img src="https://img.shields.io/badge/PRs-welcome-brightgreen.svg" alt="PRs Welcome" />
48
19
  </a>
49
20
  </p>
50
- <h3 align="center"><a href="https://airesearch.js.org/"> πŸ“‘ Docs (airesearch.js.org)</a> <a href="https://qwksearch.com/"> πŸš€ Demo</a></h3>
21
+ <h3 align="center"><a href="https://airesearch.js.org/">πŸ“‘ Docs (airesearch.js.org)</a> <a href="https://qwksearch.com/">πŸš€ Demo</a></h3>
51
22
 
23
+ ## extract-webpage
52
24
 
53
- ## πŸ§ πŸ’» Reimagine the Internet as Self-Organizing Mind Map
25
+ Search, extract, cite, and outline the web for a topic with AI Research Agent. This package provides the core pipeline for turning URLs into structured content: fetching pages, extracting readable text, identifying citations, summarizing keyphrases, and tokenizing for search.
54
26
 
55
- <p align="center">
56
- <img src="https://i.imgur.com/Z9OJMwd.gif" />
57
- </p>
27
+ ```bash
28
+ npm install extract-webpage
29
+ ```
58
30
 
59
- πŸ“œ [Research Paper](https://drive.google.com/file/d/1EuV4fTKsBiBIyDDuc4ss47oFFUVjCsXX/view)
31
+ ---
60
32
 
61
- Critical times call for critical thinkers to create a crowdsourced argument reasoning dataset, for AI models to recommend research quotes, to evolve crowdsourced chain-of-thought reasoning, to unlock faster ways to read long articles, to monitor developments by topic modeling a knowledge base graph, and to provide a public service of answers to research.
33
+ ### πŸšœπŸ“œ Tractor the Text Extractor
62
34
 
63
- Language Models can distill the essence of collective thought into a vector space where every point has a weighted value representing its contribution to the overall decision-making process, leading to direct democratic AI economy where public votes reward influence. AI will show its reasoning based on what sentences and cites it used from the collective research, so that people can see it is aligned with our interests. Research Agents recommend articles for human researchers working alongside AI to develop a summarized topic outline as a public service. The agents monitor for any related articles via web searches for keywords associated with that Topic Model. Imagine uploading a research paper, then the app extracts full text of reference cites and creates topic model and keyword summaries, then monitors that literature base and stores highlights. People will make personal knowledge bases of what influences them to create AI assistants cloning their mind-uploaded perspective and interests in a self-organizing mind map. Similar apps are Anthropic Claude, Obsidian, SciSpace and Perplexity, showing that people need an emergence of this "self-organizing mind map" approach to manage the complexity of information flow.
64
-
65
- ![image](https://i.imgur.com/R2ARMyq.png)
35
+ <p align="center">
36
+ <img width="350px" src="https://i.imgur.com/o8NTXxY.png" />
37
+ </p>
66
38
 
67
- <img src="https://github.com/TutteInstitute/datamapplot/raw/main/examples/ArXiv_example.gif" width="800px"/>
39
+ [extract Docs](https://airesearch.js.org/functions/extractor/url-to-content/)
68
40
 
41
+ 1. **Main Content Detection**: Extract the main content from a URL by combining Mozilla Readability and Postlight Mercury algorithms, utilizing over 100 custom adapters for major sites for article, author, date HTML classes.
42
+ 2. **Basic HTML Standardization**: Transform complex HTML into a simplified reading-mode format of basic HTML, making it ideal for research note archival and focused reading, with headings, images, and links.
43
+ 3. **YouTube Transcript Processing**: When a YouTube video URL is detected, retrieve the complete video transcript including both manual captions and auto-generated subtitles, maintaining proper timestamp synchronization.
44
+ 4. **DOCX Extraction**: Extracts text and structure from Word documents.
45
+ 5. **PDF to HTML**: Extracts formatted text from PDF with parsing of linebreaks, page headers, footnotes, and section headings. Supports fonts, links, bold, italics, lists, headings, headers, footnotes, Table of Contents, Quotes, and Code Blocks. Uses [pdfjs-serverless](https://github.com/johannschopplich/pdfjs-serverless) to work in Cloudflare workers, serverless, Node.js, and front-end environments.
46
+ 6. **Cite**: Identify and extract citation metadata including author names, publication dates, sources, and titles using HTML meta tags and common class name patterns. Validates author names against a database of 90,000 first and last names to distinguish personal from organizational authors.
69
47
 
48
+ ---
70
49
 
71
- ### πŸ€–πŸ”Ž STREAM: Search with Top Result Extraction & Answer Model
50
+ ### πŸ•ΈοΈπŸ–₯️ Tardigrade the Web Crawler
72
51
 
73
52
  <p align="center">
74
- <img width="350px" src="https://i.imgur.com/s8gsYt1.png" />
53
+ <img src="https://i.imgur.com/iuzpcvD.png" width="350px" />
75
54
  </p>
76
55
 
77
- [searchSTREAM Docs](https://airesearch.js.org/functions/search/search-stream)
56
+ [scrapeURL Docs](https://airesearch.js.org/functions/extractor/url-to-content/scrape-url)
78
57
 
79
- 1. Search Web via metasearch of major engines or your custom data
80
- 2. Extract text of top results with Tractor the Text Extractor.
81
- 3. SEEKTOPIC: Extract Keyphrase Topics and Top Sentences that centralize those topics
82
- 4. Rerank documents's chunks based on relevance to query, using embeddings by convert text to concept vector, get cosine similarity of query to topic, returning the sentences central to key relevant parts of the article.
83
- 5. Research Agent prompt with key sentences from relevant sources to answer via Groq Llama, OpenAI, or other LLMs and suggest follow-ups
58
+ 1. **Fetch API first**: Scrape any domain's URL to get its HTML, JSON, or binary buffer. Includes timeout, redirects, default user agent, referer as Google, and bot detection checking.
59
+ 2. **Docker fallback**: If fetch does not return needed HTML, use a Docker container with proxy as backup.
60
+ 3. **Puppeteer rendering**: NodeJS server renders with Puppeteer DOM to get all HTML loaded by secondary in-page API requests after the initial page load, including user login and cookie storage.
61
+ 4. **Cloudflare bypass**: A webpage proxy that requests through Chromium (Puppeteer) to bypass Cloudflare anti-bot using cookie id JavaScript method.
62
+ 5. **Usage**: `http://localhost:3000/?url=https://example.org`
84
63
 
85
- ### πŸšœπŸ“œ Tractor the Text Extractor
64
+ ---
86
65
 
87
- <p align="center">
88
- <img width="350px" src="https://i.imgur.com/o8NTXxY.png" />
89
- </p>
66
+ ### πŸ” Web Search
90
67
 
91
- [extract Docs](https://airesearch.js.org/functions/extractor/url-to-content/)
68
+ Federated web search via multiple backends:
69
+
70
+ - **Tavily**: AI-optimized search API with result extraction.
71
+ - **SearXNG**: Open-source metasearch engine aggregating major search engines.
72
+ - **Meta Search Agent**: Unified interface over multiple search providers.
92
73
 
93
- 1. Main Content Detection: Extract the main content from a URL by combining Mozilla Readability and Postlight Mercury algorithms, utilizing over 100 custom adapters for major sites for article, author, date HTML classes.
94
- 3. YouTube Transcript Processing: When a YouTube video URL is detected, retrieve the complete video transcript including both manual captions and auto-generated subtitles, maintaining proper timestamp synchronization.
95
- 4. PDF to HTML: Extracts formatted text from PDF with parsing of linebreaks , page headers, footnotes, and section headings. Supports fonts, links, bold, italics, lists, headings, headers, footnotes, and Table of Contents, Quotes, and Code Blocks. Removes repeated headers, links footnote anchors to the footnote, and preserves number of the PDF page with invisible I element. This function uses [pdfjs-serverless](https://github.com/johannschopplich/pdfjs-serverless) to work in more environments than PDF.js-based tools: Cloudflare workers, serverless, node.js, and front-end only.
96
- 2. Basic HTML Standardization: Transform complex HTML into a simplified reading-mode format of basic HTML, making it ideal for research note archival and focused reading, with headings, images and links.
97
- 5. Cite: Identify and extract citation metadata including author names, publication dates, sources, and titles using HTML meta tags and common class name patterns. The system validates author names against a comprehensive database of 90,000 first and last names, distinguishing between personal and organizational authors to properly format citations.
74
+ [searchSTREAM Docs](https://airesearch.js.org/functions/search/search-stream)
75
+
76
+ ---
98
77
 
99
78
  ### πŸ”€πŸ“Š SEEKTOPIC: Summarization by Extracting Entities, Keyword Tokens, and Outline Phrases Important to Context
100
79
 
101
80
  <p align="center">
102
- <img width="350px" src="https://i.imgur.com/nMoDgz6.jpeg" />
81
+ <img width="350px" src="https://i.imgur.com/nMoDgz6.jpeg" />
103
82
  </p>
104
83
 
105
- [extractSEEKTOPIC Docs](https://airesearch.js.org/functions/topics/seektopic-keyphrases)
84
+ [extractSEEKTOPIC Docs](https://airesearch.js.org/functions/topics/seektopic-keyphrases)
106
85
  [SEEKTOPIC Sample Output](https://github.com/vtempest/ai-research-agent/blob/master/test/data/)
107
86
 
108
- SEEKTOPIC can be used to find unique, domain-specific keyphrases using noun Ngrams. The user can click on keyphrases or LLM can suggest questions based on them. The user can see highlighted just the most important sentences that centralize and tie in the core topics. It is possible to vectorize and compare the dot product similarity of query to keyphrases which are then mapped to parts of the document like section labels. This is more in line with how humans think of article organization into section headings and lead sentences which tie in concepts from others.
109
-
110
- SEEKTOPIC extracts unique, domain-specific key phrases from a document using noun n-grams and ranks sentences based on their centrality to the most frequently referenced key phrase concepts, enabling efficient extraction of domain-specific content and provides a flexible framework for summarization.
87
+ SEEKTOPIC extracts unique, domain-specific key phrases from a document using noun n-grams and ranks sentences based on their centrality to the most frequently referenced key phrase concepts.
111
88
 
112
- 1. Sentence Segmentation: Split the text into sentences, accounting for common abbreviations, numbers, URLs, and other exceptions.
113
- 2. Tokenization and Phrase Extraction: Employ a Wiki Phrases tokenizer to identify wiki topics, phrases, and nouns. This includes spell-checking and checks the root words using Porter Stemmer (e.g., "video gaming" is tokenized as "video game").
114
- 3. Noun N-gram Extraction: Generate noun edge-grams, allowing for stop words in the middle (e.g., "state of the art").
115
- 4. Key Phrase Consolidation: Merge smaller n-grams that are subsets of larger ones by comparing weights.
116
- 5. Domain Specificity Calculation: Determine named entities and phrase domain specificity using WikiIDF. This rewards unique key phrases specific to the document's field (e.g., "endocrinology" in medical texts or "thou shall" in religious texts).
117
- 6. Key Phrase Filtering: Select top key phrases based on a combination of frequency and word count.
118
- 7. Graph Construction: Create a double-ring weighted graph with key phrases in the central ring and sentences in the outer ring. Assign weights to links based on concept usage probability.
119
- 8. Sentence Weighting: Apply TextRank algorithm to weight sentences, identifying those that centralize and connect key phrase concepts most referenced by other sentences. This process, based on TextRank and PageRank, includes random surfing and jumping to avoid getting stuck in loops (like page headers).
120
- 9. Top Results Selection: Select top sentences and key phrases based on overall weight and graph centrality, using either a fixed number or percentage for larger documents.
121
- 10. Output Generation: Return top sentences (with associated key phrases) and top key phrases (with associated sentences).
122
- 11. Dynamic Reranking: If a user interacts with a key phrase or if there's a search query leading to the document, compare query similarity to key phrases, heavily weight the most similar key phrase, and reapply TextRank from step 8.
89
+ 1. **Sentence Segmentation**: Split text into sentences, accounting for common abbreviations, numbers, URLs, and other exceptions.
90
+ 2. **Tokenization and Phrase Extraction**: Employ a Wiki Phrases tokenizer to identify wiki topics, phrases, and nouns. Includes spell-checking and root word checking via Porter Stemmer.
91
+ 3. **Noun N-gram Extraction**: Generate noun edge-grams, allowing for stop words in the middle (e.g., "state of the art").
92
+ 4. **Key Phrase Consolidation**: Merge smaller n-grams that are subsets of larger ones by comparing weights.
93
+ 5. **Domain Specificity Calculation**: Determine named entities and phrase domain specificity using WikiIDF, rewarding unique key phrases specific to the document's field.
94
+ 6. **Graph Construction**: Create a double-ring weighted graph with key phrases in the central ring and sentences in the outer ring.
95
+ 7. **Sentence Weighting**: Apply TextRank algorithm to weight sentences, identifying those that centralize and connect key phrase concepts most referenced by other sentences.
96
+ 8. **Dynamic Reranking**: If a user interacts with a key phrase or a search query leads to the document, compare query similarity to key phrases, heavily weight the most similar key phrase, and reapply TextRank.
123
97
 
124
- ### πŸ•ΈοΈπŸ–₯️ Tardigrade the Web Crawler
125
- <p align="center">
126
- <img src="https://i.imgur.com/iuzpcvD.png" width="350px" />
127
- </p>
128
-
129
- [scrapeURL Docs](https://airesearch.js.org/functions/extractor/url-to-content/scrape-url)
98
+ ---
130
99
 
131
- 1. First Use Fetch API, check for bot detection. Scrape any domain's URL to get its HTML, JSON, or Binary Buffer.
132
- Scraping internet pages is a [free speech right](https://blog.apify.com/is-web-scraping-legal/) to access information.
133
- 2. Features: timeout, redirects, default UA, referer as google, and bot
134
- detection checking.
135
- 3. If fetch method does not get needed HTML, use Docker container wih proxy as backup.
136
- 4. [Setup Docker](https://github.com/vtempest/ai-research-agent/tree/master/packages/crawler)
137
- container with NodeJS server API renders with puppeteer DOM to get all HTML loaded by
138
- secondary in-page API requests after the initial page request, including user login and cookie storage.
139
- 5. Bypass Cloudflare bot check: A webpage proxy that request through Chromium (puppeteer) - can be used
140
- to bypass Cloudflare anti bot using cookie id javascript method.
141
- 6. Send your request to the server with the port 3000 and add your URL to the "url"
142
- query string like this: `http://localhost:3000/?url=https://example.org`. Pass in this proxy url in the fetch request.
143
-
144
- ### πŸŒπŸ“– WORLD: Wikipedia Outline Relational Lexicon & Dictionary
100
+ ### πŸ§©πŸ” Autocomplete & Query-to-Topic Phrase Tokenization
145
101
 
146
102
  <p align="center">
147
- <img width="350px" src="https://i.imgur.com/ffaU3s7.jpeg" />
103
+ <img width="350px" src="https://i.imgur.com/tMjFGe4.jpeg" />
148
104
  </p>
149
105
 
150
- [compileTopicModel Docs](https://airesearch.js.org/functions/datasets/compile-topic-model)
151
-
152
- Search and outline a research base using Wikipedia's 100k popular pages as the core topic phrases graph for LLM Research Agents. Most of the documents online (and by extension thinking in the collective conciousness) can revolve around core topic phrases linked as a graph. If all the available docs are nodes, the links in the graph can be extracted Wiki page entities and mappings of dictionary phrases to their wiki page. These can serve as topic labels, keywords, and suggestions for LLM followup questions. Documents can be linked in a graph with: 1. wiki page entity recognition 2. frequent keyphrases 3. html links 4. research paper references 5. keyphrases to query in global web search 6. site-specific recommendations. These can lay the foundation for LLM Research Agents to fully grok, summarize, and outline a research base.
153
-
154
- * 240K total words & phrases, first 117K first-word or single words to check every token against. 100K Wikipedia Page Titles and links - Wikipedia most popular pages titles. Also includes domain specificity score and what letters should be capital.
155
- * 84K words and 67K phrases in dictionary lexicon OpenEnglishWordNet, a better updated version of Wordnet - multiple definitions per term, 120k definitions, 45 concept categories
156
- * JSON Prefix Trie - arranged by sorting words and phrases for lookup by first word to tokenize by word, then find if it starts a phrase based on entries, for Phrase Extraction from a text. There is ["consensus"](https://johnresig.com/blog/javascript-trie-performance-analysis/) that Prefix Trie [O(1) lookups](https://github.com/daviddwlee84/LeetCode/blob/master/Notes/DataStructure/Trie_PrefixTree.md) (instead of thaving to loop through the index for each lookup) makes it the best data type for this task.
157
-
158
- ### πŸ“ˆπŸ“ WRITEFAT: Weigh Relevance by Inference of Topics, Entities, and Frequency Averages for Terms
159
- <p align="center">
160
- <img width="350px" src="https://i.imgur.com/e2uTpoh.png" />
161
- </p>
106
+ [suggestNextWordCompletions Docs](https://airesearch.js.org/functions/tokenize/suggest-complete-word)
162
107
 
163
- [weighRelevanceTermFrequency Docs](https://airesearch.js.org/functions/match/weigh-relevance-frequency)
108
+ Search-on-keystroke word and phrase completion, sorted by IDF frequency, for search autocomplete dropdowns. Tokenizes by phrase rather than individual word to improve search accuracy β€” "white house" or "state of the art" are extracted and searched as phrases rather than split into independent words.
164
109
 
165
- ![WRITEFAT Formula](https://i.imgur.com/gCjSQeV.png)
110
+ Additional tokenization utilities:
111
+ - **text-to-sentences**: Split text into sentence segments.
112
+ - **text-to-chunks**: Split text into LLM-ready chunks.
113
+ - **text-to-topic-tokens**: Extract topic tokens from text.
114
+ - **word-to-root-stem**: Porter Stemmer for root word normalization.
166
115
 
167
- Calculate term specificity for a single doc with BM25 formula by using Wikipedia term frequencies as the baseline Inverse Frequency across Documents. WikiBM25 solves the need to pass in all docs to compute against all documents in a database. The problem with BM25 and TF-IDF is that a large set of documents is needed to find the words that are repeated often across all. These overused words are often the same list of words, so using Wikipedia's term frequencies ensures a common sense baseline against a neutral corpus.
116
+ ---
168
117
 
169
- **Data Model**: All words in English Wikipedia are sorted by number of pages they are in for 325K words with frequencies of at least 32 wikipages, between 3 to 23 characters of Latin alphanumerics like az09, punctuation like .-, and diacritics like éï, but filtering out numbers and foreign language.
170
- Use this list to Replace or Combine with All Documents IDF - Many websites may have less than a hundred pages to search through and that is not enough to find which terms are domain-specific. They can score a single doc at a time to find the weight each word in query gets. Wikipedia IDf can be a baseline IDF to average with the All Docs IDF for uniqueness across the average public and the specific domain.
118
+ ## Usage
171
119
 
172
- **Example**: Given a query "Superbowl wins by year" we do not want to simply return docs filled with common words like year, but rather recognize Superbowl is more domain-specific. This requires precomputing IDF values across all docs, and for websites that may not have that many docs to start with may consider averaging their precomputed score with wikiIDF values to ensure most unique words get a score.
120
+ ```ts
121
+ import {
122
+ urlToContent,
123
+ urlToHtml,
124
+ htmlToContent,
125
+ htmlToCite,
126
+ seektopicKeyphrases,
127
+ searchWeb,
128
+ suggestCompleteWord,
129
+ } from "extract-webpage";
173
130
 
174
- **LLM RAG Use Case**: LLM RAG Chunk to Query Similarity - When we chunk a document into parts to find which to pass into a LLM prompt, they need to be weighed by relevance to the query. Semantic Embedding with a LLM not only takes resources to compute & store the vectors, it also [performs worse (video)](https://youtu.be/9QJXvNiJIG8?si=ey4GbqtV8tD5WV2P&t=725) than BM25 on its own. Hybrid BM25 & Embeddings RAG is best, but there may not be time to compute BM25 idf scores across all doc chunks. We need a fast way to distinguish more unique words to give them more weight rather than common short words that get repeated a lot in an edge case paragraph. WikiBM25 is the best in use cases like realtime web search where chunking the text cannot be done beforehand.
131
+ // Fetch and extract content from a URL
132
+ const content = await urlToContent("https://example.com/article");
175
133
 
176
- ### πŸ§©πŸ” Autocomplete & Query To Topic Phrase Tokenization
177
- <p align="center">
178
- <img width="350px" src="https://i.imgur.com/tMjFGe4.jpeg" />
179
- </p>
134
+ // Extract keyphrases and top sentences
135
+ const { keyphrases, sentences } = seektopicKeyphrases(content.text);
180
136
 
181
- [suggestNextWordCompletions Docs](https://airesearch.js.org/functions/tokenize/suggest-complete-word)
137
+ // Search the web
138
+ const results = await searchWeb("AI research agents");
139
+ ```
182
140
 
183
- Search-on-keystroke and load this JSON index for word and phrase completion, sorted by how common the terms are with IDF, for search autocomplete dropdown. Tokening by word can often have a meaning widely different than if it is part of a phrase, so it is better to extract phrases by first-word next-words pairings. Search results will be more accurate if we infer likely phrases and search for those words occuring together and not just split into words and find frequency. Examples are "white house" or "state of the art" which should be searched as a phrase but would return different context if split into words. As Led Zeppelin famously put it: β™« "'Cause you know sometimes words have two meanings."
184
-
185
- ## Further Research
186
-
187
- * [ThoughtSource Reasoning Datasets](https://github.com/OpenBioLink/ThoughtSource)
188
- * [Debate on Graph (arxiv)](https://arxiv.org/html/2409.03155v1)
189
- * [Awesome-LLMs-Datasets](https://github.com/lmmlzn/Awesome-LLMs-Datasets)
190
- * [AI Research Agent's NPM Dependecies](https://npmgraph.js.org/?q=ai-research-agent#hide=)
191
- * [Tensorflow.js Demos](https://www.tensorflow.org/js/demos)
192
- * [GPT Researcher](https://github.com/assafelovic/gpt-researcher)
193
- * [NLP Papers Latest Updates](https://index.quantumstat.com)
194
- * [Anthropic Persuation Overview](https://www.anthropic.com/research/measuring-model-persuasiveness)
195
- * [NLP Research Progress](https://github.com/sebastianruder/NLP-progress/)
196
- * [NLP Datasets](https://github.com/niderhoff/nlp-datasets?tab=readme-ov-file)
197
- * [Mastering Retrieval for LLMs - BM25, Fine-tuned Embeddings, and Re-Rankers](https://www.youtube.com/watch?v=9QJXvNiJIG8)
198
- * [Google Search Algorithm](https://searchengineland.com/google-search-document-leak-ranking-442617)
199
- * [Transformers Explained Visually (Part 3)](https://towardsdatascience.com/transformers-explained-visually-part-3-multi-head-attention-deep-dive-1c1ff1024853)
200
- * [Can LLMs Generate Novel Research Ideas?](https://arxiv.org/html/2409.04109v1)
201
- * [Graph Algorithms Playground](https://playground.memgraph.com)
202
- * [CommonCrawl C4 Download](https://huggingface.co/datasets/allenai/c4)
203
- * [Knowledge Graphs Prompts Papers](https://github.com/zjunlp/PromptKG)
204
- * [Paper - Iterative Research Idea Generation](https://arxiv.org/abs/2404.07738)
205
- * [How might LLMs store facts](https://www.youtube.com/watch?v=9-Jl0dxWQs8&t=70s)
206
- * [Awesome-LLM-Graph-Theory](https://github.com/XiaoxinHe/Awesome-Graph-LLM)
207
- * [Open Deep Search](https://arxiv.org/html/2503.20201v1)
208
- * [LangChain Hub](https://smith.langchain.com/hub) - A collection of reusable prompts, chains, and agents for building LLM applications
209
- * [LangChain Documentation](https://js.langchain.com/docs) - Comprehensive documentation for LangChain.js
210
-
211
- <img src="https://img.shields.io/badge/PRs-welcome-brightgreen.svg"
212
- alt="PRs Welcome" /> Please star this repo for updates! 🌟
141
+ <img src="https://img.shields.io/badge/PRs-welcome-brightgreen.svg" alt="PRs Welcome" /> Please star this repo for updates! 🌟
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "extract-webpage",
3
- "version": "1.2.20",
3
+ "version": "1.2.22",
4
4
  "module": "./dist/extract-webpage.es.js",
5
5
  "description": "Search, extract, cite, and outline the web for a topic with AI Research Agent.",
6
6
  "author": "vtempest <grokthiscontact@gmail.com>",
@@ -78,11 +78,11 @@
78
78
  "dependencies": {
79
79
  "@huggingface/transformers": "^3.8.1",
80
80
  "ai": "^5.0.0",
81
- "chat-agent-toolkit": "^1.2.18",
81
+ "chat-agent-toolkit": "^1.2.20",
82
82
  "chrono-node": "^2.9.0",
83
83
  "drizzle-orm": "^0.45.1",
84
- "extract-pdf": "^0.1.6",
85
- "extract-youtube": "^1.0.3",
84
+ "extract-pdf": "^0.1.8",
85
+ "extract-youtube": "^1.0.9",
86
86
  "highlight.js": "^11.11.1",
87
87
  "html-entities": "^2.6.0",
88
88
  "js-yaml": "^4.1.1",