AI crawler

Bytespider: who runs it and how to block it.

This is ByteDance's crawler. ByteDance publishes no documentation for it, so this page relies on what Cloudflare reported.

Short answer

What is this crawler?

Bytespider is a crawler operated by ByteDance, the company that owns TikTok. Cloudflare reports it is used to gather training data for ByteDance's large language models. You can block Bytespider with a robots.txt rule or a firewall rule.

In brief

Four things to know first.

  • Who runs it

    ByteDance, the Chinese company that owns TikTok, according to Cloudflare.

  • What it does

    It reportedly gathers training data for ByteDance's LLMs, including those behind Doubao. It is not a search crawler.

  • No official docs

    ByteDance publishes no documentation for it, so every fact here is attributed to Cloudflare.

  • Blocking is a low-risk choice

    Google Search does not use this crawler. A rule for it is a separate decision from your Google Search visibility.

What it is

Who runs Bytespider, and why

ByteDance publishes no documentation for this crawler that we could cite. Everything below comes from Cloudflare's reporting, so treat the purpose as reported, not confirmed by the operator.

ByteDance's crawler at a glance (source: Cloudflare)
ItemWhat Cloudflare reports
OperatorByteDance, the Chinese company that owns TikTok
PurposeReportedly gathers training data for its LLMs
Related productIts ChatGPT rival, Doubao
TypeModel training crawler, not search or user-request

Scale

How heavily Bytespider crawls

Cloudflare says this crawler leads AI bots in request volume, how widely it crawls, and how often it is blocked. It ranked among the top four AI crawlers by requests, with Amazonbot, ClaudeBot and GPTBot.

  • Top four by requests

    ByteDance's crawler, Amazonbot, ClaudeBot and GPTBot, over the last year on Cloudflare's network.

  • Most blocked

    GPTBot ranks second to the ByteDance crawler in both crawling and being blocked.

  • Rarely named in robots.txt

    Cloudflare found top-10,000-domain files reference GPTBot and CCBot but rarely disallow this crawler.

Recognise it

Spotting the Bytespider user agent

Search your logs for Bytespider, the crawler name Cloudflare reports. We found no ByteDance user agent page or IP list. Names can be faked: Cloudflare says some operators lie about their user agent.

Cloudflare says operators have spoofed user agents to look like real browsers. Its machine learning model flags such traffic as bot activity even when the user agent is false.

  1. Search your logs

    Filter your server logs for this crawler's token in the user agent field.
  2. Treat the match as a claim

    A user agent can be faked, and no operator IP list is published to check it.
  3. Use bot scoring if you can

    Cloudflare says a WAF rule challenging bot scores below 30 blocked the evasive AI traffic it analyzed.

Not sure which bots to allow?

Our Technical GEO work sets up robots.txt for the AI crawlers that matter in your category and fixes crawler access, so the right bots can reach and read your most important pages.

Control

How to block Bytespider

Add a User-agent line with its token and Disallow: / to robots.txt. Then check whether your important pages can still be found and read by the AI crawlers you allow.

  1. Add the robots.txt rule

    In robots.txt, add a User-agent line with this crawler's token. On the next line, add: Disallow: /
  2. Add a firewall rule

    Block or challenge this user agent at your server, CDN or WAF, since robots.txt is only a request.
  3. Or use Cloudflare's toggle

    Cloudflare added a one-click 'AI Scrapers and Crawlers' setting under Security > Bots, free on every plan.
  4. Recheck your logs

    Watch for hits from it over the next weeks to confirm the block holds.

Trade-off

Should you block ByteDance's crawler?

Google Search does not use this crawler, so a rule for it is a separate decision from your Google Search visibility.

The sources here do not show whether Doubao answers cite your pages. Citation tracking can show which URLs AI responses cite, and whether those citations point to your own pages.

Allow this crawlerBlock this crawler
Training useContent may feed ByteDance's LLMsYour pages stay out of that crawl
Server loadAdds crawl requestsCuts those requests
Google, Bing, ChatGPT searchUnaffectedUnaffected

FAQ

ByteDance's crawler: common questions.

What is this crawler's user agent?

Search your logs for this crawler's user agent token. Our sources include no ByteDance documentation or IP list. You cannot check a match against official ranges, and user agents can be spoofed.

Does this crawler respect robots.txt?

ByteDance does not document this crawler's robots.txt behavior in the sources we have. Cloudflare says many AI bots follow robots.txt, but does not confirm this for it. If you need certainty, add a firewall rule beside your robots.txt Disallow line.

Will blocking this crawler hurt my Google rankings?

According to Google, Google Search does not use this crawler. Google also says many suggested GEO "hacks" are not effective or supported by how its Search works.

Why is this crawler so heavily blocked?

Site owners often block this crawler because of how much it crawls. Cloudflare says it leads AI bots in request volume, in how widely it crawls Internet properties, and in how often it is blocked.

How is this crawler different from GPTBot?

Both are training crawlers run by different companies. ByteDance runs one; OpenAI manages GPTBot, which collects training data for its LLMs. Cloudflare says GPTBot ranks second to the ByteDance crawler in both crawling and being blocked.

Can Cloudflare block this crawler for me?

Cloudflare offers a one-click 'AI Scrapers and Crawlers' toggle under Security > Bots that blocks this crawler. It is open to all customers, including the free tier, and Cloudflare says it updates the feature as it finds new bot fingerprints.

Decide which crawlers get your content.

We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.