Meta crawler
Meta-ExternalAgent: what it does.
Meta runs several crawlers. Meta-ExternalAgent gathers content for AI. The fetcher opens links when a user asks. Here is how to spot them and control them.
Short answer
What does Meta's broad AI crawler do?
Meta-ExternalAgent is Meta's web crawler. Meta says it crawls the web for uses such as training foundation AI models or improving products by indexing content. Meta-ExternalFetcher is separate: it fetches individual links at a user's request.
In brief
Four things to know first.
Two different bots
The first crawls the web for AI training or indexing. Meta-ExternalFetcher fetches single links a user asks for.
Easy to spot
Each crawler names itself in the user agent string: meta-externalagent/1.1 for one, meta-externalfetcher/1.1 for the other.
robots.txt works, with a catch
Meta says it respects robots.txt choices, but this fetcher may bypass them because a user asked for the fetch.
Blocking is a trade-off
Blocking keeps content out of Meta's crawls. It also gives Meta less to work from when it answers.
Who runs it
Meta's crawlers and their jobs
Meta says it uses web crawlers, software that fetches content from websites, for several purposes. Its developer documentation lists the user agents that identify the most common ones.
Meta says it lets site owners state their preferences through industry-standard practices like robots.txt, rather than non-standard formats like NoAI tags.
| Crawler | Meta's stated purpose |
|---|---|
| Broad crawl | Crawls the web to train AI models or index content |
| User-requested fetch | Fetches single links at a user's request |
| Meta-WebIndexer | Improves Meta AI search result quality |
| Meta-ExternalAds | Improves advertising and business products |
| FacebookExternalHit | Builds previews of shared links |
Recognise it
Find Meta's crawlers in your logs
Meta publishes the user agent strings for each crawler. Search your server logs for the exact names below.
Meta advises allow-listing either the user agent strings or the IP addresses, and describes the IP addresses as the more secure option.
| Crawler | Spotting it in logs |
|---|---|
| Broad AI crawler | facebookexternalhit/1.1, with or without the externalhit_uatext.php link |
| User-request fetcher | Fetcher name with version 1.1, with or without a documentation link |
The fetcher
Meta-ExternalFetcher acts on request
Meta says this crawler fetches individual links at a user's request. It supports functions such as evaluating and improving agentic AI capabilities, including helping AI navigate websites to complete tasks for users.
Because a user asked for the fetch, Meta says this crawler may bypass robots.txt. A robots.txt rule alone may not stop it.
| Meta's AI crawler | Meta's user-request fetcher | |
|---|---|---|
| Trigger | Crawls the web broadly | A user's request for a link |
| Purpose | AI model training or indexing | Agentic AI tasks for users |
| robots.txt | Set a rule to block it | May bypass robots.txt rules |
Not sure which AI crawlers to allow?
We check crawler access as part of Technical GEO, so the pages you want cited can be reached and read.
Control
Allow or block it in robots.txt
Meta says you block one of its crawlers by adding a disallow for it in robots.txt. Its documentation shows an example that opens with a User-agent line naming the crawler.
Allow up to 24 hours for changes to take effect, because crawlers may cache robots.txt for up to 24 hours.
Find it in your logs
Search your logs for the user agent strings of both Meta AI crawlers, the crawler and Meta's fetcher, to see which pages they request.Decide per crawler
Choose separately for each of Meta's crawlers: the broad crawler, the user-request fetcher and Meta-WebIndexer. They do different jobs.Edit robots.txt
Add a User-agent line for the crawler, then a Disallow line for the paths you want closed.Wait up to 24 hours
Meta says crawlers may cache robots.txt for up to 24 hours before a change applies.Check your logs again
Confirm the crawler stops requesting the blocked pages. For other bots, see our robots.txt guide.
The trade-off
Should you block it?
Blocking this crawler keeps your pages out of Meta's crawls for AI training or indexing. Meta does not publish figures on how this affects visibility, so the effect is unmeasured.
Meta asks site owners to allow Meta-WebIndexer so it can cite and link to their content in Meta AI responses. Blocking that crawler removes that option.
Block to opt out
Choose this if you do not want your content used for Meta's AI training or indexing.
Allow to stay visible
Allowing Meta-WebIndexer helps Meta cite and link to your content in Meta AI answers.
Do not rely on one rule
This fetcher may bypass robots.txt, so a block may not stop user-requested fetches.
FAQ
Meta-ExternalAgent: common questions.
What is this Meta crawler?
Meta runs this web crawler. According to Meta's documentation, it crawls the web for use cases such as training foundation AI models or improving products by indexing content directly.
What is this fetcher and how is it different?
Meta's fetcher retrieves individual links at a user's request. Meta says it supports agentic AI, such as helping AI navigate websites to complete tasks for users. Because a user asked for each fetch, Meta says the fetcher may bypass robots.txt rules.
How do I block this crawler in robots.txt?
Add a disallow rule for this crawler to robots.txt, under a User-agent line that names it. Meta says changes can take up to 24 hours to take effect, because crawlers may cache the file.
How do I recognise this crawler in my logs?
In logs, the user agent shows the crawler's name and version 1.1, sometimes followed by a documentation path in parentheses. Meta also advises allow-listing by IP addresses, which it calls the more secure option.
Will robots.txt stop this fetcher?
A robots.txt rule may not stop this fetcher. Meta says it may bypass robots.txt because it performs fetches that a user requested. The other Meta crawler is the one Meta says to block with a disallow rule.
Does this crawler build link previews?
Meta runs a separate crawler, FacebookExternalHit, for shared links. According to Meta, FacebookExternalHit gathers, caches and displays a shared page's title, description and thumbnail image.
See which AI crawlers reach your pages.
We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.