Meta crawler

facebookexternalhit: Meta's link preview crawler.

What the Facebook crawler does with your pages, how to recognise it in your logs, and how to allow or block it in robots.txt.

Short answer

What is facebookexternalhit?

facebookexternalhit is the Facebook crawler run by Meta. When someone shares your URL on Facebook, Instagram or Messenger, it fetches the page and caches its title, description and thumbnail image to build the link preview.

In brief

Four things to know first.

  • It builds link previews

    Meta says it crawls content shared on its apps, then caches the title, description and thumbnail.

  • It is triggered by shares

    Its visits follow someone sharing your link, whether pasted or shared through the Facebook social plugin.

  • It may skip robots.txt

    Meta says it might bypass robots.txt during security or integrity checks, such as malware checks.

  • Other Meta bots differ

    Meta documents separate crawlers for AI training, user-requested fetches and ads, each with its own name.

What it does

What the Facebook crawler does with your content.

Meta says FacebookExternalHit crawls the content of an app or website shared on its apps, such as Facebook, Instagram or Messenger. It gathers, caches and displays the title, description and thumbnail image.

This is a preview job. Meta lists it separately from its crawlers for AI training and AI search.

Meta crawlers and their stated purpose, from Meta's crawler documentation
CrawlerStated purpose
Meta's link preview crawlerCrawls content shared on Meta apps for link previews
Meta-ExternalAgentTraining foundation AI models, indexing content directly
Meta-ExternalFetcherFetches individual links at a user's request
Meta-WebIndexerImproves Meta AI search result quality
Meta-ExternalAdsImproves advertising and other business products

Recognise it

Spot it in your logs.

Meta documents two user agent strings for this crawler. Match either one in your server logs.

Meta advises allow-listing by user agent string or by IP address, and calls the IP addresses the more secure option. Its documentation lists the current ranges.

  • Full string

    externalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)

  • Short string

    externalhit/1.1

  • Verify by IP

    Check the address against the IP ranges Meta publishes before trusting a request.

Control it

Allow or block it in robots.txt.

Meta says to block a crawler by adding a disallow rule for it in robots.txt. Its example uses the meta-externalagent token. Allow up to 24 hours for changes, because crawlers may cache the file.

  1. Find the visits

    Search your logs for Meta's link preview user agent. Meta's documentation lists it in two forms: a full string that includes the externalhit_uatext.php link, and a shorter version 1.1 string without that link.
  2. Decide per crawler

    Treat previews, AI training and user fetches as separate choices, each with its own user agent.
  3. Add a disallow rule

    Add a User-agent line for the crawler, then a Disallow line for the paths to block.
  4. Wait up to 24 hours

    Meta says crawlers may cache robots.txt for up to 24 hours, so changes are not instant.
  5. Re-check your logs

    Confirm visits change as expected. Remember security checks may still bypass robots.txt.

Not sure which crawlers can read your site?

We check crawler access, rendering and schema as part of Technical GEO, so assistants can reach the pages you want cited.

Should you block it

What blocking it costs you.

Blocking this crawler risks broken previews: content that cannot be fetched cannot be displayed. Meta also says content must be crawlable within a few seconds, or Facebook cannot display it.

Meta's documentation does not tie this link preview crawler to AI answers. Meta-WebIndexer is the one it asks you to allow so Meta AI can cite and link to your content.

Allow itBlock it
Shared linksPreview shows title, description, imagePreview may fail to display
Slow pagesMust respond within a few secondsNot relevant if blocked
Security checksRun as Meta documentsMeta says they may bypass robots.txt
Missed contentForce a recrawl in Sharing DebuggerNo preview refresh to request

FAQ

facebookexternalhit: common questions.

Is Meta's link preview bot the same as the Facebook crawler?

Meta's link preview bot is the one most site owners see, but it is only one of several Meta crawlers. Meta-ExternalAgent, for example, is listed separately. Check the exact user agent string in your logs. Not every Meta visit is a link preview.

Does Meta's link preview crawler follow robots.txt?

Meta's link preview crawler generally follows robots.txt. Meta says it might bypass the file during security or integrity checks, such as checking for malware or malicious content. Meta also says robots.txt changes can take up to 24 hours to take effect.

How do I refresh a preview that shows old content?

Meta says that if content was not available when the crawler visited, you can force a recrawl by passing the URL through the Sharing Debugger tool or the Sharing API. This is how you refresh a stale or missing preview.

Why is my link preview not showing?

Meta says content must be crawlable within a few seconds, or Facebook cannot display it. Slow responses, blocked crawlers or a mishandled Range header can all prevent a preview. Check your logs, then recrawl with the Sharing Debugger.

Does blocking Meta's link preview crawler stop Meta AI training?

Blocking this link preview bot does not target AI training. Meta describes Meta-ExternalAgent as the crawler used for training foundation AI models. Each Meta crawler has its own user agent, so you set a separate robots.txt rule for each one.

How can I verify a request is really from Meta?

Meta advises allow-listing either the crawler's user agent strings or its IP addresses, and calls the IP addresses the more secure option. A user agent can be copied by anyone, so compare the request IP with Meta's published list.

See which crawlers reach your site.

We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.