Amazonbot

Amazonbot: what it does and how to block it.

Amazon runs three separate crawlers. Here is what each one does with your content, how to spot it in your logs, and how to control it.

Short answer

What is this Amazon crawler?

Amazonbot is Amazon's web crawler. Amazon says it improves its products and services and may be used to train Amazon AI models. It honors robots.txt, so you can block Amazonbot there without affecting Amazon's two other crawlers.

In brief

Four things to know first.

  • Three crawlers, not one

    Amazon documents three crawlers: its main bot, Amzn-SearchBot and Amzn-User. Each setting is independent of the others.

  • It may train models

    Amazon says this crawler may be used to train Amazon AI models. The other two do not train generative AI.

  • It follows robots.txt

    Amazon says automated crawling respects the Robots Exclusion Protocol. Changes can take about 24 hours.

  • Blocking has a trade-off

    Allowing it may open Amazon Content Partners benefits. Blocking it removes that option.

What it does

Three Amazon crawlers, three jobs.

Amazon's developer page says each user agent setting is independent. Blocking one does not block the others, so decide for each what you want.

Amazon's three crawlers, as described on developer.amazon.com/amazon
CrawlerAmazon's own descriptionModel training
Main crawlerImproves Amazon products and servicesMay train Amazon AI models
Amzn-SearchBotImproves search experiences, such as AlexaNo generative AI training
Amzn-UserFetches live pages for user actionsNo generative AI training

Recognise it

Amazonbot user agent in your logs.

Amazon gives this user agent string for its main crawler, where W.X.Y.Z stands for the Chrome version. Search your server logs for the crawler token exactly as written in the string above.

Amazon also publishes the IP addresses it crawls from. Check a request's IP against that list before trusting the user agent alone.

  • User agent

    Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36

  • Verification

    Compare the request IP with the crawler IP addresses Amazon publishes.

Control it

How to block Amazonbot in robots.txt.

Amazon says its crawlers honor the user-agent and allow/disallow directives. Rules apply per host, so a subdomain needs its own robots.txt.

  1. Open your robots.txt

    Edit the file at the root of each host you want to control.
  2. Name the crawler

    Add a User-agent line for the main crawler, Amzn-SearchBot or Amzn-User. Each is set separately.
  3. Choose allow or disallow

    Add a Disallow rule to block it, or an Allow rule to keep it.
  4. Wait for the change

    Amazon says settings may take about 24 hours to take effect. It caches robots.txt for up to 30 days.

Not sure which AI crawlers to allow?

Crawler access is one part of Technical GEO. We check what assistants can reach and read on your site.

Limits

What robots.txt won't control.

Amazon says Amzn-User may not follow all robots.txt directives, because a user can initiate its actions. Crawl-delay is not supported by any of the three.

  • Page-level tags

    Amazon honors noarchive (do not use for model training) and noindex or none (do not index).

  • Link-level rule

    The crawlers respect rel=nofollow on links.

  • No crawl-delay

    Amazon says these crawlers do not support the crawl-delay directive.

Should you block

Should you block it? The trade-off.

Amazon says allowing Amzn-SearchBot makes your content eligible to appear in search experiences such as Alexa. If Amzn-SearchBot is not named but other search bots are allowed, it follows their rules.

Amazon also says sites that allow this crawler may qualify for Amazon Content Partners, which offers a +1% affiliate commission boost, free hosting credits and AI traffic tools.

AllowBlock
AmazonbotMay feed AI training; may open Content PartnersKeeps content out of its reach
Amzn-SearchBotContent eligible for Alexa-type searchContent not eligible there
Amzn-UserLive fetches can answer user queriesMay still ignore some directives

FAQ

Amazon's crawler: common questions.

What does Amazon use it for?

Amazon says this crawler is used to improve its products and services. Amazon adds that this helps it give customers more accurate information and that the crawler may be used to train Amazon AI models.

How do I block it?

To block it, add a robots.txt rule that names this crawler as the user agent and disallows the paths you want protected. Amazon says its crawlers honor these directives, and changes may take about 24 hours to apply.

What user agent does it send?

Amazon documents the full user agent string for this crawler and publishes the IP addresses it crawls from. The string carries Amazon's own crawler token with version 0.1, between Chrome and Safari version markers.

Does blocking it also block Amzn-SearchBot?

Blocking this crawler does not block Amzn-SearchBot or Amzn-User. Amazon states that each user agent setting is independent of the others, so you set a rule for each crawler you want to allow or block.

Does it respect crawl-delay?

Amazon's crawlers do not support the crawl-delay directive, according to Amazon's documentation. They do honor allow and disallow rules, rel=nofollow links, and the noarchive, noindex and none meta tags, so those are your controls instead.

Does Amzn-User always follow robots.txt?

Amzn-User may not follow all robots.txt directives. Amazon explains that a user can initiate its actions, such as fetching live information to answer an Alexa query. Amazon also says it does not crawl for generative AI training.

Make sure assistants can read and cite your site.

We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.