@devtune/ai-traffic 0.1.2 → 0.2.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,8 +1,6 @@
1
1
  # @devtune/ai-traffic
2
2
 
3
- `@devtune/ai-traffic` is the filtered edge-push sensor for DevTune AI Traffic. It matches incoming requests against DevTune's AI bot registry and forwards only matching machine hits to DevTune's ingest endpoint.
4
-
5
- Cloudflare-proxied sites should prefer the Cloudflare pull integration when available. For Vercel and self-hosted sites, this package keeps volume low by filtering at your edge instead of shipping full request logs.
3
+ Measure which AI crawlers visit your site with a dependency-free sensor that filters at the edge and forwards only matched machine traffic to DevTune.
6
4
 
7
5
  ## Install
8
6
 
@@ -10,118 +8,214 @@ Cloudflare-proxied sites should prefer the Cloudflare pull integration when avai
10
8
  pnpm add @devtune/ai-traffic
11
9
  ```
12
10
 
13
- Set your project-scoped server-side ingest key:
11
+ Create a project-scoped server-side ingest key and expose it only to your server runtime:
14
12
 
15
13
  ```bash
16
14
  DEVTUNE_AI_TRAFFIC_INGEST_KEY=dt_ingest_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
17
15
  ```
18
16
 
19
- ## Next.js Proxy / Middleware
17
+ ## Quickstarts
18
+
19
+ ### Next.js
20
20
 
21
- For Next.js 16, add `proxy.ts` and export `proxy()`. For older Next.js projects, use the same body from `middleware.ts` and export `middleware()` instead.
21
+ For Next.js 16, add `proxy.ts`. For older Next.js projects, use the same body in `middleware.ts` and export `middleware()` instead.
22
22
 
23
23
  ```ts
24
24
  // proxy.ts
25
- import { createDevTuneAiTraffic } from "@devtune/ai-traffic";
25
+ import { createDevTuneAiTrafficMiddleware } from "@devtune/ai-traffic";
26
26
  import { NextResponse, type NextFetchEvent, type NextRequest } from "next/server";
27
27
 
28
- const aiTraffic = createDevTuneAiTraffic({
28
+ const trackAiTraffic = createDevTuneAiTrafficMiddleware({
29
29
  ingestKey: process.env.DEVTUNE_AI_TRAFFIC_INGEST_KEY!,
30
- batchSize: 1,
31
- defaultStatus: null,
32
- flushIntervalMs: 0,
33
- waitUntilRegistryRefresh: false,
34
30
  });
35
31
 
36
32
  export function proxy(request: NextRequest, event: NextFetchEvent) {
37
- const response = NextResponse.next();
38
-
39
- aiTraffic.trackRequest(request, event);
33
+ trackAiTraffic(request, event);
40
34
 
41
- return response;
35
+ return NextResponse.next();
42
36
  }
43
37
  ```
44
38
 
45
- Use `defaultStatus: null` for pass-through proxy or middleware requests. Next.js proxy cannot observe the final status after route resolution, so this records the AI bot hit without guessing that every pass-through is a 200. Use `batchSize: 1` and `flushIntervalMs: 0` in middleware so Vercel does not hold middleware duration open for a delayed batch timer. Use `waitUntilRegistryRefresh: false` in middleware so hourly registry refreshes do not contribute to `waitUntil` duration. The package still schedules matched ingest work through `event.waitUntil()` when available, so the request path can return immediately after matching. Ordinary browser user agents are ignored before registry refresh work is scheduled.
46
-
47
- When your proxy or middleware returns a response directly, pass that exact status:
39
+ ### Express
48
40
 
49
41
  ```ts
50
- export function proxy(request: NextRequest, event: NextFetchEvent) {
51
- const response = new Response("Forbidden", { status: 403 });
42
+ import express from "express";
43
+ import { createDevTuneAiTrafficExpressMiddleware } from "@devtune/ai-traffic/express";
52
44
 
53
- aiTraffic.trackRequest(request, event, response.status);
45
+ const app = express();
54
46
 
55
- return response;
56
- }
47
+ app.use(
48
+ createDevTuneAiTrafficExpressMiddleware({
49
+ ingestKey: process.env.DEVTUNE_AI_TRAFFIC_INGEST_KEY!,
50
+ }),
51
+ );
57
52
  ```
58
53
 
59
- If you only need a minimal best-effort hook, the middleware helper is also available:
54
+ The middleware calls `next()` immediately and records the actual response status after Express emits `finish`.
55
+
56
+ ### Node HTTP
60
57
 
61
58
  ```ts
62
- import { createDevTuneAiTrafficMiddleware } from "@devtune/ai-traffic";
63
- import { NextResponse, type NextFetchEvent, type NextRequest } from "next/server";
59
+ import { createServer } from "node:http";
64
60
 
65
- const trackAiTraffic = createDevTuneAiTrafficMiddleware({
61
+ import { createDevTuneAiTrafficNodeHandler } from "@devtune/ai-traffic/node";
62
+
63
+ const trackAiTraffic = createDevTuneAiTrafficNodeHandler({
66
64
  ingestKey: process.env.DEVTUNE_AI_TRAFFIC_INGEST_KEY!,
67
65
  });
68
66
 
69
- export function proxy(request: NextRequest, event: NextFetchEvent) {
70
- trackAiTraffic(request, event);
67
+ createServer((request, response) => {
68
+ trackAiTraffic(request, response);
71
69
 
72
- return NextResponse.next();
73
- }
70
+ response.statusCode = 200;
71
+ response.end("ok");
72
+ }).listen(3000);
74
73
  ```
75
74
 
76
- The middleware helper applies the middleware-safe defaults shown above: unknown pass-through status, immediate single-hit sends, and registry refreshes that are not attached to `waitUntil`. The package does not import `next/server`; it uses the standard request shape and the `waitUntil` method that Next.js passes to proxy and middleware functions.
75
+ The hook records the final `response.statusCode` without blocking or changing the response.
77
76
 
78
- ## Status Capture
77
+ ## Privacy by Design
79
78
 
80
- Next.js proxy and middleware can report statuses they return directly, such as redirects, denials, and configured misses. They still cannot observe the final page response status after Next.js route resolution, so pass-through requests should use `defaultStatus: null` unless you intentionally want a best-effort default. When status is unknown, DevTune still counts the matched machine hit and treats status-specific reporting as unavailable for that event.
79
+ The sensor filters requests where your application runs. It uses a cheap user-agent hint before consulting the AI bot registry, and only matched machine hits are eligible to be forwarded. It uses no cookies, no fingerprinting, and no full request logs.
81
80
 
82
- Use the route wrapper where you own the handler response and need exact status capture:
81
+ For a matched crawler, DevTune receives only:
82
+
83
+ - `path`
84
+ - origin plus path, with query parameters removed
85
+ - user agent
86
+ - response status when the adapter can know it
87
+ - timestamp
88
+
89
+ Request bodies, cookies, IP addresses, unrelated headers, and query-derived identifiers are not sent.
90
+
91
+ ## Detect Without Sending
92
+
93
+ The standalone classifiers need no ingest key and never make a network request or push data to DevTune:
83
94
 
84
95
  ```ts
85
- // app/api/example/route.ts
86
- import { withDevTuneAiTrafficRoute } from "@devtune/ai-traffic";
87
-
88
- export const GET = withDevTuneAiTrafficRoute(
89
- async function GET() {
90
- return Response.json({ ok: true }, { status: 200 });
91
- },
92
- {
93
- ingestKey: process.env.DEVTUNE_AI_TRAFFIC_INGEST_KEY!,
94
- },
95
- );
96
+ import { detectAiCrawler, detectAiReferrer } from "@devtune/ai-traffic";
97
+
98
+ const crawler = detectAiCrawler(request);
99
+ const referrer = detectAiReferrer(request);
100
+ ```
101
+
102
+ Each function also accepts the relevant string directly: a user-agent for `detectAiCrawler()` or a referrer URL for `detectAiReferrer()`.
103
+
104
+ `detectAiCrawler()` returns the matched bot registry entry or `null`:
105
+
106
+ ```ts
107
+ type AiCrawlerDetection = {
108
+ uaPattern: string;
109
+ llmPlatform: string;
110
+ botClass: "training_crawler" | "index_bot" | "answer_fetcher" | "acting_agent";
111
+ label: string;
112
+ status?: "active" | "retired";
113
+ asnHints?: number[] | null;
114
+ notes?: string | null;
115
+ };
116
+ ```
117
+
118
+ `detectAiReferrer()` returns the matched referrer registry entry or `null`:
119
+
120
+ ```ts
121
+ type AiReferrerDetection = {
122
+ hostname: string;
123
+ llmPlatform: string;
124
+ label: string;
125
+ };
126
+ ```
127
+
128
+ Both use bundled registry snapshots by default, so they work offline and in CI. Pass `registryEntries` as the second argument when you need to classify against a pinned or private registry:
129
+
130
+ ```ts
131
+ const match = detectAiCrawler(userAgent, {
132
+ registryEntries: myRegistryEntries,
133
+ });
96
134
  ```
97
135
 
98
- For known misses in middleware, mark them explicitly:
136
+ ## Machine Traffic and Human Referrals
137
+
138
+ The sensor deliberately measures machine traffic only. DevTune gets human visits from AI products through its GA4 integration, where sessions, engagement, and conversions provide a richer picture than middleware referrer matching. The referrer detector is available for local classification, but the sensor adapters never forward human referral visits.
139
+
140
+ ## Production Notes
141
+
142
+ ### Batching and Runtime Lifetime
143
+
144
+ The default client configuration sends up to 10 matched events per unchanged ingest payload, with a 250 ms flush window for low-volume traffic. Set `batchSize: 1` and `flushIntervalMs: 0` to opt out and send each matched event immediately.
145
+
146
+ Express and Node servers are long-lived enough to benefit directly from the default. Next.js proxy and middleware pass the delayed send to `event.waitUntil()` when available, which prevents the response from waiting but may keep the middleware invocation alive for the short flush window. Use the explicit opt-out in short-lived runtimes that cannot reliably preserve delayed work, or when minimizing middleware duration matters more than request coalescing:
99
147
 
100
148
  ```ts
101
149
  const trackAiTraffic = createDevTuneAiTrafficMiddleware({
102
150
  ingestKey: process.env.DEVTUNE_AI_TRAFFIC_INGEST_KEY!,
103
- notFoundPathPatterns: [/^\/docs\/removed\//],
151
+ batchSize: 1,
152
+ flushIntervalMs: 0,
104
153
  });
105
154
  ```
106
155
 
107
- ## What Gets Sent
156
+ The middleware helper leaves registry refresh work outside `waitUntil` by default. Ordinary browser user agents are rejected before refresh or ingest work is scheduled.
108
157
 
109
- For matched AI bot requests only, the package sends:
158
+ ### Rate-Limit Retries
110
159
 
111
- - `path`
112
- - `url`
113
- - `ua`
114
- - `status` when known
115
- - `ts`
160
+ A `429` from the ingest endpoint is the one failure the server tells us is temporary, so the batch is redelivered rather than dropped. The client waits for the response's `Retry-After` — delay-seconds or an HTTP date — and falls back to exponential backoff from 250 ms when the header is absent or unusable. The batch being retried is held intact, so events that arrive during a backoff are sent separately rather than folded into it.
161
+
162
+ Retries are bounded so a short-lived runtime cannot be held open by a throttled endpoint. The budget belongs to one flush, not to each batch: `maxRetryAttempts` defaults to 3 and `maxRetryDelayMs` caps each wait at 5000 ms, including a `Retry-After` longer than the cap, so a flush adds at most 15 s no matter how deep the queue is. When the budget is spent the current batch is dropped with the usual sampled warning and the drain stops; whatever is behind it stays queued for the next flush, which gets a fresh budget. `maxRetryAttempts: 0` disables waiting entirely: a 429 drops the batch on its first rejection, as it did before this behaviour existed. The drain still stops there rather than dropping every batch behind it, so the rest stays queued. Tune both:
163
+
164
+ ```ts
165
+ const trackAiTraffic = createDevTuneAiTrafficMiddleware({
166
+ ingestKey: process.env.DEVTUNE_AI_TRAFFIC_INGEST_KEY!,
167
+ maxRetryAttempts: 2,
168
+ maxRetryDelayMs: 2_000,
169
+ });
170
+ ```
171
+
172
+ Because backing off lets the queue drain slower than traffic arrives, the queue itself is bounded. `maxQueuedEvents` defaults to 1000; past it the oldest events are shed with a sampled warning, so a sustained throttle cannot grow memory without limit.
173
+
174
+ Non-429 responses and network errors are unchanged: the batch is dropped after a single attempt, so a hard outage never keeps the queue alive.
175
+
176
+ ### Status Semantics
177
+
178
+ Express and Node adapters observe the completed response and report its actual status. Next.js proxy and middleware cannot observe the final route status after pass-through, so the convenience middleware omits status instead of guessing a 200.
179
+
180
+ When a Next.js proxy returns a response directly, use the client and pass the known status:
116
181
 
117
- It sends the batch to `https://devtune.ai/api/v1/llm-traffic/ingest` with `Authorization: Bearer <ingest key>`. It does not send cookies, request bodies, query-derived user identifiers, IP addresses, headers other than the user agent, or DevTune's client-side snippet key.
182
+ ```ts
183
+ import { createDevTuneAiTraffic } from "@devtune/ai-traffic";
184
+ import type { NextFetchEvent, NextRequest } from "next/server";
185
+
186
+ const aiTraffic = createDevTuneAiTraffic({
187
+ ingestKey: process.env.DEVTUNE_AI_TRAFFIC_INGEST_KEY!,
188
+ defaultStatus: null,
189
+ });
190
+
191
+ export function proxy(request: NextRequest, event: NextFetchEvent) {
192
+ const response = new Response("Forbidden", { status: 403 });
193
+
194
+ aiTraffic.trackRequest(request, event, response.status);
118
195
 
119
- ## Registry Cache
196
+ return response;
197
+ }
198
+ ```
199
+
200
+ Use `withDevTuneAiTrafficRoute()` where you own a Fetch-compatible route handler and want its exact response status captured automatically. `notFoundPathPatterns` and `statusResolver` remain available for applications that can provide additional status knowledge.
201
+
202
+ ### Registry Refresh and Fallback
203
+
204
+ Clients start with the bundled AI bot snapshot, refresh from `https://devtune.ai/api/v1/llm-traffic/registry`, cache active entries for about an hour, and use `ETag` revalidation. Failed refreshes are guarded and sampled; they never fail the application response, and the bundled snapshot remains usable.
120
205
 
121
- The package fetches `https://devtune.ai/api/v1/llm-traffic/registry`, caches the active registry for about an hour, and uses `ETag` revalidation. If the registry cannot be fetched, it falls back to a bundled snapshot so known AI bots still match.
206
+ User-agent matching is a conservative signal. Some agents spoof ordinary browsers or require network-level signals, so reported counts are a floor rather than exact bot truth. Cloudflare-proxied sites should prefer DevTune's Cloudflare pull integration when available because it can combine request, bot-score, and network signals without running middleware.
122
207
 
123
- User-agent matching is a v1 signal. Some agents spoof ordinary browsers or need ASN/TLS hints, so counts are floors rather than exact bot truth. Cloudflare pull should be used when available because Cloudflare can combine request, bot-score, and network-level signals.
208
+ ### Forwarded Origins
209
+
210
+ Express and Node adapters build event URLs from the request protocol and host, preferring the first value in standard `X-Forwarded-Proto` and `X-Forwarded-Host` chains. Only accept those headers from a trusted proxy. If they are unavailable or not trustworthy in your deployment, pass a fixed `origin`:
211
+
212
+ ```ts
213
+ const trackAiTraffic = createDevTuneAiTrafficNodeHandler({
214
+ ingestKey: process.env.DEVTUNE_AI_TRAFFIC_INGEST_KEY!,
215
+ origin: "https://www.example.com",
216
+ });
217
+ ```
124
218
 
125
- ## Vercel Pull Status
219
+ ### Failure Behavior
126
220
 
127
- Vercel Observability exposes bot and AI-crawler insights in the dashboard query builder, but the public Observability REST API currently covers Observability Plus project settings rather than event reads. Until Vercel ships a supported Observability read/query API, this middleware is the Vercel path.
221
+ Registry refresh and ingest sends are fire-and-forget and guarded. Network failures may drop telemetry, but they do not block, throw into, or alter the customer response. The ingest endpoint and batched wire format are unchanged.