domain-rank 0.1.21 → 0.1.23

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "domain-rank",
3
- "version": "0.1.21",
3
+ "version": "0.1.23",
4
4
  "description": "Domain Rank Top Million domains, source names, duplicates, and favicons",
5
5
  "repository": {
6
6
  "type": "git",
package/readme.md CHANGED
@@ -1,231 +1,231 @@
1
- ![logo](https://i.imgur.com/1MRYUqC.png)
2
-
3
- # domain-rank
4
-
5
- Look up top-ranked domains to get their name, human-readable source label, influence rank, and favicon.
6
-
7
- Rank data comes from the [Tranco List](https://tranco-list.eu/) (aggregates Cisco Umbrella, Majestic, Farsight, Chrome UX, and Cloudflare Radar) and [CommonCrawl](https://commoncrawl.org) backlink counts.
8
-
9
- **Use cases:** search/URL autocomplete, bookmark launchers, LLM web-app recommendations, domain reputation scoring.
10
-
11
- ## Installation
12
-
13
- ```bash
14
- npm install domain-rank
15
- # or
16
- pnpm add domain-rank
17
- # or
18
- yarn add domain-rank
19
- ```
20
-
21
- ## Usage
22
-
23
- ### Look up a domain
24
-
25
- ```js
26
- import { lookupDomain } from 'domain-rank';
27
-
28
- const result = lookupDomain('facebook.com');
29
- console.log(result);
30
- // {
31
- // domain: 'facebook.com',
32
- // name: 'Facebook',
33
- // rank: 1,
34
- // title: 'Facebook',
35
- // newsRank: undefined,
36
- // newsTitle: undefined,
37
- // langCode: undefined
38
- // }
39
- ```
40
-
41
- ### Get top domains
42
-
43
- ```js
44
- import { getTopDomains } from 'domain-rank';
45
-
46
- // Get top 10 domains
47
- const top10 = getTopDomains(10);
48
- console.log(top10);
49
- // [
50
- // { domain: 'facebook.com', name: 'Facebook', rank: 1, ... },
51
- // { domain: 'google.com', name: 'Google', rank: 2, ... },
52
- // { domain: 'instagram.com', name: 'Instagram', rank: 3, ... },
53
- // ...
54
- // ]
55
- ```
56
-
57
- ### Search domains
58
-
59
- ```js
60
- import { searchDomains } from 'domain-rank';
61
-
62
- const results = searchDomains('social', 5);
63
- // Returns up to 5 domains matching 'social' in name, domain, or title
64
- // sorted by rank
65
- ```
66
-
67
- ### Get favicon
68
-
69
- ```js
70
- import { getFaviconForDomain } from 'domain-rank';
71
-
72
- // Get as base64
73
- const base64Favicon = await getFaviconForDomain('google.com');
74
-
75
- // Get as URL
76
- const faviconURL = await getFaviconForDomain('google.com', false);
77
- // Returns: 'https://www.google.com/s2/favicons?domain=google.com'
78
- ```
79
-
80
- ### Format domain names
81
-
82
- ```js
83
- import { formatDomainAsTitle } from 'domain-rank';
84
-
85
- const formatted = formatDomainAsTitle('nytimes.com');
86
- // Returns: "NY Times" - human-readable name with proper capitalization
87
- ```
88
-
89
- ### Utility functions
90
-
91
- ```js
92
- import { convertURLToDomain, isURLValid, getTotalDomains } from 'domain-rank';
93
-
94
- // Extract domain from URL
95
- const domain = convertURLToDomain('https://en.wikipedia.org/wiki/Main_Page');
96
- // Returns: 'wikipedia'
97
-
98
- // Validate URL
99
- const isValid = isURLValid('https://example.com');
100
- // Returns: true
101
-
102
- // Get total number of domains in dataset
103
- const total = getTotalDomains();
104
- // Returns: ~100000
105
- ```
106
-
107
- ## Dataset Statistics
108
-
109
- - **Total domains**: 10,020 top-ranked domains
110
- - **Data file**: `data/domain-rank-merged.json` (297 KB)
111
- - **Format**: Pre-loaded in memory for instant lookups
112
- - **Bundle size**:
113
- - ESM: 828 KB (231 KB gzipped)
114
- - CJS: 828 KB (231 KB gzipped)
115
-
116
- ### Top 5 Domains
117
-
118
- | Rank | Domain | Name |
119
- |------|--------|------|
120
- | 1 | facebook.com | Facebook |
121
- | 2 | google.com | Google |
122
- | 3 | instagram.com | Instagram |
123
- | 4 | youtube.com | YouTube |
124
- | 5 | linkedin.com | LinkedIn |
125
-
126
- ## Data Format
127
-
128
- ### Source Data
129
-
130
- The raw data is stored in `data/domain-rank-merged.json` as a compact object format:
131
-
132
- ```json
133
- {
134
- "facebook.com": ["Facebook", 1],
135
- "google.com": ["Google", 2],
136
- "nytimes.com": [197, 38, "NY Times"]
137
- }
138
- ```
139
-
140
- Format: `[name, rank]` or `[newsRank, rank, title]` for news domains.
141
-
142
- ### Library API
143
-
144
- The library includes 10,020 top-ranked domains with the following information:
145
-
146
- - **domain**: The domain name (e.g., "facebook.com")
147
- - **name**: Human-readable name (e.g., "Facebook")
148
- - **rank**: Overall influence rank (lower is better, 1 is top)
149
- - **title**: Full title/description
150
- - **newsRank**: Rank for news/media domains (if applicable)
151
- - **newsTitle**: Media-specific title (if applicable)
152
- - **langCode**: Primary language code (if applicable)
153
-
154
- ### Data Sources
155
-
156
- The ranking combines multiple authoritative sources:
157
-
158
- 1. **[Tranco List](https://tranco-list.eu/)** - Aggregates:
159
- - Cisco Umbrella (DNS resolver data)
160
- - Majestic Million (backlink analysis)
161
- - Farsight (passive DNS data)
162
- - Chrome UX Report (real user metrics)
163
- - Cloudflare Radar (global network data)
164
-
165
- 2. **[CommonCrawl](https://commoncrawl.org)** - Web-wide backlink counts from petabytes of crawled data
166
-
167
- This multi-source approach provides a more stable and manipulation-resistant ranking than single-source lists.
168
-
169
- ## API Reference
170
-
171
- ### `lookupDomain(domain: string): DomainLookupResult | null`
172
- Look up information for a specific domain.
173
-
174
- ### `getTopDomains(n?: number): DomainLookupResult[]`
175
- Get top N domains by rank. Default: 100.
176
-
177
- ### `searchDomains(query: string, limit?: number): DomainLookupResult[]`
178
- Search domains by name/domain/title. Default limit: 10.
179
-
180
- ### `getAllDomains(): DomainLookupResult[]`
181
- Get all domains sorted by rank.
182
-
183
- ### `getTotalDomains(): number`
184
- Get total number of domains in the dataset.
185
-
186
- ### `getFaviconForDomain(urlOrDomain: string, formatBase64?: boolean): Promise<string>`
187
- Fetch favicon for a domain. Returns base64 by default, or URL if `formatBase64` is false.
188
-
189
- ### `formatDomainAsTitle(domain: string): string`
190
- Format domain into human-readable name with proper capitalization.
191
-
192
- ### `cleanSourceTitle(title: string): string | null`
193
- Clean and normalize a page title (removes common suffixes, HTML, etc.).
194
-
195
- ### `shouldRemoveDomain(domain: string): boolean`
196
- Check if a domain should be removed from rankings.
197
-
198
- ### `findMainDomain(domain: string): string | null`
199
- Find the main domain for a given alternate domain.
200
-
201
- ### `getTitleOverride(domain: string): string | null`
202
- Get a custom title override for a domain.
203
-
204
- ### `getSourceTitle(domain: string): Promise<string | null>`
205
- Fetch the page title from a domain's website.
206
-
207
- ### `convertURLToDomain(url: string): string`
208
- Extract domain from URL.
209
-
210
- ### `isURLValid(url: string): boolean`
211
- Validate URL format.
212
-
213
- ## Development
214
-
215
- ```bash
216
- # Install dependencies
217
- pnpm install
218
-
219
- # Type check
220
- pnpm run type-check
221
-
222
- # Build library
223
- pnpm run build
224
-
225
- # Run tests
226
- pnpm test
227
- ```
228
-
229
- ## License
230
-
231
- MIT
1
+ ![logo](https://i.imgur.com/1MRYUqC.png)
2
+
3
+ # domain-rank
4
+
5
+ Look up top-ranked domains to get their name, human-readable source label, influence rank, and favicon.
6
+
7
+ Rank data comes from the [Tranco List](https://tranco-list.eu/) (aggregates Cisco Umbrella, Majestic, Farsight, Chrome UX, and Cloudflare Radar) and [CommonCrawl](https://commoncrawl.org) backlink counts.
8
+
9
+ **Use cases:** search/URL autocomplete, bookmark launchers, LLM web-app recommendations, domain reputation scoring.
10
+
11
+ ## Installation
12
+
13
+ ```bash
14
+ npm install domain-rank
15
+ # or
16
+ pnpm add domain-rank
17
+ # or
18
+ yarn add domain-rank
19
+ ```
20
+
21
+ ## Usage
22
+
23
+ ### Look up a domain
24
+
25
+ ```js
26
+ import { lookupDomain } from 'domain-rank';
27
+
28
+ const result = lookupDomain('facebook.com');
29
+ console.log(result);
30
+ // {
31
+ // domain: 'facebook.com',
32
+ // name: 'Facebook',
33
+ // rank: 1,
34
+ // title: 'Facebook',
35
+ // newsRank: undefined,
36
+ // newsTitle: undefined,
37
+ // langCode: undefined
38
+ // }
39
+ ```
40
+
41
+ ### Get top domains
42
+
43
+ ```js
44
+ import { getTopDomains } from 'domain-rank';
45
+
46
+ // Get top 10 domains
47
+ const top10 = getTopDomains(10);
48
+ console.log(top10);
49
+ // [
50
+ // { domain: 'facebook.com', name: 'Facebook', rank: 1, ... },
51
+ // { domain: 'google.com', name: 'Google', rank: 2, ... },
52
+ // { domain: 'instagram.com', name: 'Instagram', rank: 3, ... },
53
+ // ...
54
+ // ]
55
+ ```
56
+
57
+ ### Search domains
58
+
59
+ ```js
60
+ import { searchDomains } from 'domain-rank';
61
+
62
+ const results = searchDomains('social', 5);
63
+ // Returns up to 5 domains matching 'social' in name, domain, or title
64
+ // sorted by rank
65
+ ```
66
+
67
+ ### Get favicon
68
+
69
+ ```js
70
+ import { getFaviconForDomain } from 'domain-rank';
71
+
72
+ // Get as base64
73
+ const base64Favicon = await getFaviconForDomain('google.com');
74
+
75
+ // Get as URL
76
+ const faviconURL = await getFaviconForDomain('google.com', false);
77
+ // Returns: 'https://www.google.com/s2/favicons?domain=google.com'
78
+ ```
79
+
80
+ ### Format domain names
81
+
82
+ ```js
83
+ import { formatDomainAsTitle } from 'domain-rank';
84
+
85
+ const formatted = formatDomainAsTitle('nytimes.com');
86
+ // Returns: "NY Times" - human-readable name with proper capitalization
87
+ ```
88
+
89
+ ### Utility functions
90
+
91
+ ```js
92
+ import { convertURLToDomain, isURLValid, getTotalDomains } from 'domain-rank';
93
+
94
+ // Extract domain from URL
95
+ const domain = convertURLToDomain('https://en.wikipedia.org/wiki/Main_Page');
96
+ // Returns: 'wikipedia'
97
+
98
+ // Validate URL
99
+ const isValid = isURLValid('https://example.com');
100
+ // Returns: true
101
+
102
+ // Get total number of domains in dataset
103
+ const total = getTotalDomains();
104
+ // Returns: ~100000
105
+ ```
106
+
107
+ ## Dataset Statistics
108
+
109
+ - **Total domains**: 10,020 top-ranked domains
110
+ - **Data file**: `data/domain-rank-merged.json` (297 KB)
111
+ - **Format**: Pre-loaded in memory for instant lookups
112
+ - **Bundle size**:
113
+ - ESM: 828 KB (231 KB gzipped)
114
+ - CJS: 828 KB (231 KB gzipped)
115
+
116
+ ### Top 5 Domains
117
+
118
+ | Rank | Domain | Name |
119
+ |------|--------|------|
120
+ | 1 | facebook.com | Facebook |
121
+ | 2 | google.com | Google |
122
+ | 3 | instagram.com | Instagram |
123
+ | 4 | youtube.com | YouTube |
124
+ | 5 | linkedin.com | LinkedIn |
125
+
126
+ ## Data Format
127
+
128
+ ### Source Data
129
+
130
+ The raw data is stored in `data/domain-rank-merged.json` as a compact object format:
131
+
132
+ ```json
133
+ {
134
+ "facebook.com": ["Facebook", 1],
135
+ "google.com": ["Google", 2],
136
+ "nytimes.com": [197, 38, "NY Times"]
137
+ }
138
+ ```
139
+
140
+ Format: `[name, rank]` or `[newsRank, rank, title]` for news domains.
141
+
142
+ ### Library API
143
+
144
+ The library includes 10,020 top-ranked domains with the following information:
145
+
146
+ - **domain**: The domain name (e.g., "facebook.com")
147
+ - **name**: Human-readable name (e.g., "Facebook")
148
+ - **rank**: Overall influence rank (lower is better, 1 is top)
149
+ - **title**: Full title/description
150
+ - **newsRank**: Rank for news/media domains (if applicable)
151
+ - **newsTitle**: Media-specific title (if applicable)
152
+ - **langCode**: Primary language code (if applicable)
153
+
154
+ ### Data Sources
155
+
156
+ The ranking combines multiple authoritative sources:
157
+
158
+ 1. **[Tranco List](https://tranco-list.eu/)** - Aggregates:
159
+ - Cisco Umbrella (DNS resolver data)
160
+ - Majestic Million (backlink analysis)
161
+ - Farsight (passive DNS data)
162
+ - Chrome UX Report (real user metrics)
163
+ - Cloudflare Radar (global network data)
164
+
165
+ 2. **[CommonCrawl](https://commoncrawl.org)** - Web-wide backlink counts from petabytes of crawled data
166
+
167
+ This multi-source approach provides a more stable and manipulation-resistant ranking than single-source lists.
168
+
169
+ ## API Reference
170
+
171
+ ### `lookupDomain(domain: string): DomainLookupResult | null`
172
+ Look up information for a specific domain.
173
+
174
+ ### `getTopDomains(n?: number): DomainLookupResult[]`
175
+ Get top N domains by rank. Default: 100.
176
+
177
+ ### `searchDomains(query: string, limit?: number): DomainLookupResult[]`
178
+ Search domains by name/domain/title. Default limit: 10.
179
+
180
+ ### `getAllDomains(): DomainLookupResult[]`
181
+ Get all domains sorted by rank.
182
+
183
+ ### `getTotalDomains(): number`
184
+ Get total number of domains in the dataset.
185
+
186
+ ### `getFaviconForDomain(urlOrDomain: string, formatBase64?: boolean): Promise<string>`
187
+ Fetch favicon for a domain. Returns base64 by default, or URL if `formatBase64` is false.
188
+
189
+ ### `formatDomainAsTitle(domain: string): string`
190
+ Format domain into human-readable name with proper capitalization.
191
+
192
+ ### `cleanSourceTitle(title: string): string | null`
193
+ Clean and normalize a page title (removes common suffixes, HTML, etc.).
194
+
195
+ ### `shouldRemoveDomain(domain: string): boolean`
196
+ Check if a domain should be removed from rankings.
197
+
198
+ ### `findMainDomain(domain: string): string | null`
199
+ Find the main domain for a given alternate domain.
200
+
201
+ ### `getTitleOverride(domain: string): string | null`
202
+ Get a custom title override for a domain.
203
+
204
+ ### `getSourceTitle(domain: string): Promise<string | null>`
205
+ Fetch the page title from a domain's website.
206
+
207
+ ### `convertURLToDomain(url: string): string`
208
+ Extract domain from URL.
209
+
210
+ ### `isURLValid(url: string): boolean`
211
+ Validate URL format.
212
+
213
+ ## Development
214
+
215
+ ```bash
216
+ # Install dependencies
217
+ pnpm install
218
+
219
+ # Type check
220
+ pnpm run type-check
221
+
222
+ # Build library
223
+ pnpm run build
224
+
225
+ # Run tests
226
+ pnpm test
227
+ ```
228
+
229
+ ## License
230
+
231
+ MIT