Meta crawler

Meta-ExternalAgent: what it does.

Meta runs several crawlers. Meta-ExternalAgent gathers content for AI. The fetcher opens links when a user asks. Here is how to spot them and control them.

Short answer

What does Meta's broad AI crawler do?

Meta-ExternalAgent is Meta's web crawler. Meta says it crawls the web for uses such as training foundation AI models or improving products by indexing content. Meta-ExternalFetcher is separate: it fetches individual links at a user's request.

In brief

Four things to know first.

  • Two different bots

    The first crawls the web for AI training or indexing. Meta-ExternalFetcher fetches single links a user asks for.

  • Easy to spot

    Each crawler names itself in the user agent string: meta-externalagent/1.1 for one, meta-externalfetcher/1.1 for the other.

  • robots.txt works, with a catch

    Meta says it respects robots.txt choices, but this fetcher may bypass them because a user asked for the fetch.

  • Blocking is a trade-off

    Blocking keeps content out of Meta's crawls. It also gives Meta less to work from when it answers.

Who runs it

Meta's crawlers and their jobs

Meta says it uses web crawlers, software that fetches content from websites, for several purposes. Its developer documentation lists the user agents that identify the most common ones.

Meta says it lets site owners state their preferences through industry-standard practices like robots.txt, rather than non-standard formats like NoAI tags.

Meta crawlers and what Meta says they do (Meta developer documentation)
CrawlerMeta's stated purpose
Broad crawlCrawls the web to train AI models or index content
User-requested fetchFetches single links at a user's request
Meta-WebIndexerImproves Meta AI search result quality
Meta-ExternalAdsImproves advertising and business products
FacebookExternalHitBuilds previews of shared links

Recognise it

Find Meta's crawlers in your logs

Meta publishes the user agent strings for each crawler. Search your server logs for the exact names below.

Meta advises allow-listing either the user agent strings or the IP addresses, and describes the IP addresses as the more secure option.

User agent strings published by Meta
CrawlerSpotting it in logs
Broad AI crawlerfacebookexternalhit/1.1, with or without the externalhit_uatext.php link
User-request fetcherFetcher name with version 1.1, with or without a documentation link

The fetcher

Meta-ExternalFetcher acts on request

Meta says this crawler fetches individual links at a user's request. It supports functions such as evaluating and improving agentic AI capabilities, including helping AI navigate websites to complete tasks for users.

Because a user asked for the fetch, Meta says this crawler may bypass robots.txt. A robots.txt rule alone may not stop it.

Meta's AI crawlerMeta's user-request fetcher
TriggerCrawls the web broadlyA user's request for a link
PurposeAI model training or indexingAgentic AI tasks for users
robots.txtSet a rule to block itMay bypass robots.txt rules

Not sure which AI crawlers to allow?

We check crawler access as part of Technical GEO, so the pages you want cited can be reached and read.

Control

Allow or block it in robots.txt

Meta says you block one of its crawlers by adding a disallow for it in robots.txt. Its documentation shows an example that opens with a User-agent line naming the crawler.

Allow up to 24 hours for changes to take effect, because crawlers may cache robots.txt for up to 24 hours.

  1. Find it in your logs

    Search your logs for the user agent strings of both Meta AI crawlers, the crawler and Meta's fetcher, to see which pages they request.
  2. Decide per crawler

    Choose separately for each of Meta's crawlers: the broad crawler, the user-request fetcher and Meta-WebIndexer. They do different jobs.
  3. Edit robots.txt

    Add a User-agent line for the crawler, then a Disallow line for the paths you want closed.
  4. Wait up to 24 hours

    Meta says crawlers may cache robots.txt for up to 24 hours before a change applies.
  5. Check your logs again

    Confirm the crawler stops requesting the blocked pages. For other bots, see our robots.txt guide.

The trade-off

Should you block it?

Blocking this crawler keeps your pages out of Meta's crawls for AI training or indexing. Meta does not publish figures on how this affects visibility, so the effect is unmeasured.

Meta asks site owners to allow Meta-WebIndexer so it can cite and link to their content in Meta AI responses. Blocking that crawler removes that option.

  • Block to opt out

    Choose this if you do not want your content used for Meta's AI training or indexing.

  • Allow to stay visible

    Allowing Meta-WebIndexer helps Meta cite and link to your content in Meta AI answers.

  • Do not rely on one rule

    This fetcher may bypass robots.txt, so a block may not stop user-requested fetches.

FAQ

Meta-ExternalAgent: common questions.

What is this Meta crawler?

Meta runs this web crawler. According to Meta's documentation, it crawls the web for use cases such as training foundation AI models or improving products by indexing content directly.

What is this fetcher and how is it different?

Meta's fetcher retrieves individual links at a user's request. Meta says it supports agentic AI, such as helping AI navigate websites to complete tasks for users. Because a user asked for each fetch, Meta says the fetcher may bypass robots.txt rules.

How do I block this crawler in robots.txt?

Add a disallow rule for this crawler to robots.txt, under a User-agent line that names it. Meta says changes can take up to 24 hours to take effect, because crawlers may cache the file.

How do I recognise this crawler in my logs?

In logs, the user agent shows the crawler's name and version 1.1, sometimes followed by a documentation path in parentheses. Meta also advises allow-listing by IP addresses, which it calls the more secure option.

Will robots.txt stop this fetcher?

A robots.txt rule may not stop this fetcher. Meta says it may bypass robots.txt because it performs fetches that a user requested. The other Meta crawler is the one Meta says to block with a disallow rule.

Does this crawler build link previews?

Meta runs a separate crawler, FacebookExternalHit, for shared links. According to Meta, FacebookExternalHit gathers, caches and displays a shared page's title, description and thumbnail image.

See which AI crawlers reach your pages.

We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.