Amazonbot
Amazonbot: what it does and how to block it.
Amazon runs three separate crawlers. Here is what each one does with your content, how to spot it in your logs, and how to control it.
Short answer
What is this Amazon crawler?
Amazonbot is Amazon's web crawler. Amazon says it improves its products and services and may be used to train Amazon AI models. It honors robots.txt, so you can block Amazonbot there without affecting Amazon's two other crawlers.
In brief
Four things to know first.
Three crawlers, not one
Amazon documents three crawlers: its main bot, Amzn-SearchBot and Amzn-User. Each setting is independent of the others.
It may train models
Amazon says this crawler may be used to train Amazon AI models. The other two do not train generative AI.
It follows robots.txt
Amazon says automated crawling respects the Robots Exclusion Protocol. Changes can take about 24 hours.
Blocking has a trade-off
Allowing it may open Amazon Content Partners benefits. Blocking it removes that option.
What it does
Three Amazon crawlers, three jobs.
Amazon's developer page says each user agent setting is independent. Blocking one does not block the others, so decide for each what you want.
| Crawler | Amazon's own description | Model training |
|---|---|---|
| Main crawler | Improves Amazon products and services | May train Amazon AI models |
| Amzn-SearchBot | Improves search experiences, such as Alexa | No generative AI training |
| Amzn-User | Fetches live pages for user actions | No generative AI training |
Recognise it
Amazonbot user agent in your logs.
Amazon gives this user agent string for its main crawler, where W.X.Y.Z stands for the Chrome version. Search your server logs for the crawler token exactly as written in the string above.
Amazon also publishes the IP addresses it crawls from. Check a request's IP against that list before trusting the user agent alone.
User agent
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36
Verification
Compare the request IP with the crawler IP addresses Amazon publishes.
Control it
How to block Amazonbot in robots.txt.
Amazon says its crawlers honor the user-agent and allow/disallow directives. Rules apply per host, so a subdomain needs its own robots.txt.
Open your robots.txt
Edit the file at the root of each host you want to control.Name the crawler
Add a User-agent line for the main crawler, Amzn-SearchBot or Amzn-User. Each is set separately.Choose allow or disallow
Add a Disallow rule to block it, or an Allow rule to keep it.Wait for the change
Amazon says settings may take about 24 hours to take effect. It caches robots.txt for up to 30 days.
Not sure which AI crawlers to allow?
Crawler access is one part of Technical GEO. We check what assistants can reach and read on your site.
Limits
What robots.txt won't control.
Amazon says Amzn-User may not follow all robots.txt directives, because a user can initiate its actions. Crawl-delay is not supported by any of the three.
Page-level tags
Amazon honors noarchive (do not use for model training) and noindex or none (do not index).
Link-level rule
The crawlers respect rel=nofollow on links.
No crawl-delay
Amazon says these crawlers do not support the crawl-delay directive.
Should you block
Should you block it? The trade-off.
Amazon says allowing Amzn-SearchBot makes your content eligible to appear in search experiences such as Alexa. If Amzn-SearchBot is not named but other search bots are allowed, it follows their rules.
Amazon also says sites that allow this crawler may qualify for Amazon Content Partners, which offers a +1% affiliate commission boost, free hosting credits and AI traffic tools.
| Allow | Block | |
|---|---|---|
| Amazonbot | May feed AI training; may open Content Partners | Keeps content out of its reach |
| Amzn-SearchBot | Content eligible for Alexa-type search | Content not eligible there |
| Amzn-User | Live fetches can answer user queries | May still ignore some directives |
FAQ
Amazon's crawler: common questions.
What does Amazon use it for?
Amazon says this crawler is used to improve its products and services. Amazon adds that this helps it give customers more accurate information and that the crawler may be used to train Amazon AI models.
How do I block it?
To block it, add a robots.txt rule that names this crawler as the user agent and disallows the paths you want protected. Amazon says its crawlers honor these directives, and changes may take about 24 hours to apply.
What user agent does it send?
Amazon documents the full user agent string for this crawler and publishes the IP addresses it crawls from. The string carries Amazon's own crawler token with version 0.1, between Chrome and Safari version markers.
Does blocking it also block Amzn-SearchBot?
Blocking this crawler does not block Amzn-SearchBot or Amzn-User. Amazon states that each user agent setting is independent of the others, so you set a rule for each crawler you want to allow or block.
Does it respect crawl-delay?
Amazon's crawlers do not support the crawl-delay directive, according to Amazon's documentation. They do honor allow and disallow rules, rel=nofollow links, and the noarchive, noindex and none meta tags, so those are your controls instead.
Does Amzn-User always follow robots.txt?
Amzn-User may not follow all robots.txt directives. Amazon explains that a user can initiate its actions, such as fetching live information to answer an Alexa query. Amazon also says it does not crawl for generative AI training.
Make sure assistants can read and cite your site.
We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.