Meta crawler
facebookexternalhit: Meta's link preview crawler.
What the Facebook crawler does with your pages, how to recognise it in your logs, and how to allow or block it in robots.txt.
Short answer
What is facebookexternalhit?
facebookexternalhit is the Facebook crawler run by Meta. When someone shares your URL on Facebook, Instagram or Messenger, it fetches the page and caches its title, description and thumbnail image to build the link preview.
In brief
Four things to know first.
It builds link previews
Meta says it crawls content shared on its apps, then caches the title, description and thumbnail.
It is triggered by shares
Its visits follow someone sharing your link, whether pasted or shared through the Facebook social plugin.
It may skip robots.txt
Meta says it might bypass robots.txt during security or integrity checks, such as malware checks.
Other Meta bots differ
Meta documents separate crawlers for AI training, user-requested fetches and ads, each with its own name.
What it does
What the Facebook crawler does with your content.
Meta says FacebookExternalHit crawls the content of an app or website shared on its apps, such as Facebook, Instagram or Messenger. It gathers, caches and displays the title, description and thumbnail image.
This is a preview job. Meta lists it separately from its crawlers for AI training and AI search.
| Crawler | Stated purpose |
|---|---|
| Meta's link preview crawler | Crawls content shared on Meta apps for link previews |
| Meta-ExternalAgent | Training foundation AI models, indexing content directly |
| Meta-ExternalFetcher | Fetches individual links at a user's request |
| Meta-WebIndexer | Improves Meta AI search result quality |
| Meta-ExternalAds | Improves advertising and other business products |
Recognise it
Spot it in your logs.
Meta documents two user agent strings for this crawler. Match either one in your server logs.
Meta advises allow-listing by user agent string or by IP address, and calls the IP addresses the more secure option. Its documentation lists the current ranges.
Full string
externalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)
Short string
externalhit/1.1
Verify by IP
Check the address against the IP ranges Meta publishes before trusting a request.
Control it
Allow or block it in robots.txt.
Meta says to block a crawler by adding a disallow rule for it in robots.txt. Its example uses the meta-externalagent token. Allow up to 24 hours for changes, because crawlers may cache the file.
Find the visits
Search your logs for Meta's link preview user agent. Meta's documentation lists it in two forms: a full string that includes the externalhit_uatext.php link, and a shorter version 1.1 string without that link.Decide per crawler
Treat previews, AI training and user fetches as separate choices, each with its own user agent.Add a disallow rule
Add a User-agent line for the crawler, then a Disallow line for the paths to block.Wait up to 24 hours
Meta says crawlers may cache robots.txt for up to 24 hours, so changes are not instant.Re-check your logs
Confirm visits change as expected. Remember security checks may still bypass robots.txt.
Not sure which crawlers can read your site?
We check crawler access, rendering and schema as part of Technical GEO, so assistants can reach the pages you want cited.
Should you block it
What blocking it costs you.
Blocking this crawler risks broken previews: content that cannot be fetched cannot be displayed. Meta also says content must be crawlable within a few seconds, or Facebook cannot display it.
Meta's documentation does not tie this link preview crawler to AI answers. Meta-WebIndexer is the one it asks you to allow so Meta AI can cite and link to your content.
| Allow it | Block it | |
|---|---|---|
| Shared links | Preview shows title, description, image | Preview may fail to display |
| Slow pages | Must respond within a few seconds | Not relevant if blocked |
| Security checks | Run as Meta documents | Meta says they may bypass robots.txt |
| Missed content | Force a recrawl in Sharing Debugger | No preview refresh to request |
FAQ
facebookexternalhit: common questions.
Is Meta's link preview bot the same as the Facebook crawler?
Meta's link preview bot is the one most site owners see, but it is only one of several Meta crawlers. Meta-ExternalAgent, for example, is listed separately. Check the exact user agent string in your logs. Not every Meta visit is a link preview.
Does Meta's link preview crawler follow robots.txt?
Meta's link preview crawler generally follows robots.txt. Meta says it might bypass the file during security or integrity checks, such as checking for malware or malicious content. Meta also says robots.txt changes can take up to 24 hours to take effect.
How do I refresh a preview that shows old content?
Meta says that if content was not available when the crawler visited, you can force a recrawl by passing the URL through the Sharing Debugger tool or the Sharing API. This is how you refresh a stale or missing preview.
Why is my link preview not showing?
Meta says content must be crawlable within a few seconds, or Facebook cannot display it. Slow responses, blocked crawlers or a mishandled Range header can all prevent a preview. Check your logs, then recrawl with the Sharing Debugger.
Does blocking Meta's link preview crawler stop Meta AI training?
Blocking this link preview bot does not target AI training. Meta describes Meta-ExternalAgent as the crawler used for training foundation AI models. Each Meta crawler has its own user agent, so you set a separate robots.txt rule for each one.
How can I verify a request is really from Meta?
Meta advises allow-listing either the crawler's user agent strings or its IP addresses, and calls the IP addresses the more secure option. A user agent can be copied by anyone, so compare the request IP with Meta's published list.
See which crawlers reach your site.
We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.