Attribution guide / sources read 6 October 2026 UTC
Separate crawler traffic from agent visits.
A User-Agent token is a clue, not verified identity. Check the vendor's published IP ranges or signed requests before trusting it. A crawler visit does not establish that a buyer used an agent.
| Vendor | Search | Training | User fetch or agent browsing |
|---|---|---|---|
| OpenAI | OAI-SearchBot | GPTBot | ChatGPT-User identifies user-initiated fetches. Cloud browser requests use HTTP Message Signatures and Signature-Agent https://chatgpt.com. |
| Anthropic | Claude-SearchBot | ClaudeBot | Claude-User identifies user-triggered retrieval. These tokens do not identify all extension browsing. |
| Perplexity | PerplexityBot | No training token inferred here. | Perplexity-User identifies requested retrieval. These tokens do not prove a Comet browser visit. |
| Googlebot | Google-Extended is a robots control token, not a separate HTTP User-Agent. | Google-Agent identifies hosted agents performing user-requested browsing. It does not establish Gemini in Chrome activity. |
Use the current full strings and verification ranges in the primary sources below. Version numbers and browser strings change.
GA4: referrals you can observe
- Open Explore and create a free-form exploration.
- Add Session source / medium as a dimension and Sessions as a metric.
- Create a session segment where Session source matches
^(chatgpt\.com|perplexity\.ai|claude\.ai|gemini\.google\.com)$. - Compare that segment with all sessions for a stated date range. Treat the result as identified referrals, not all agent traffic.
Referrals from chatgpt.com, perplexity.ai, claude.ai and gemini.google.com arrive as referral or direct traffic depending on the client's setup. Browser agents may not execute analytics, and direct traffic has no reliable agent identity.
Cloudflare: edge requests
- Use the analytics view available on your plan for the relevant hostname and dates.
- Filter observed User-Agent tokens, then inspect the verified-bot signal or validated signature and bot category where available.
- Separate Search, Training and Agent requests. Compare response status and challenge or block actions.
- Review an access rule only when a recorded failed task supports it. Keep abuse controls and authentication in place.
Cloudflare's documented September 15, 2026 defaults for new domains block Training and Agent categories on pages with ads while allowing Search. This is not a claim that every site blocks agents.
Server logs: a first inspection
rg -i 'OAI-SearchBot|GPTBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|Google-Agent' access.log
This finds header claims only. Correlate the timestamp, requested page, status and verified identity. Redact personal data before sharing logs. This page installs no tracking script and collects no traffic data.
Primary sources
- Google Analytics: create and use segments
- OpenAI bots and Cloud browser request authentication
- Anthropic crawler purposes
- Perplexity crawlers and verification
- Google common crawlers and user-triggered fetchers
- Cloudflare bot categories and defaults and verified-bot variables
- Google AI features: no special AI file required