เครื่องมือตรวจสอบ AI Crawler ฟรี
ตรวจสอบว่าเว็บไซต์ของคุณบล็อก AI crawler เช่น GPTBot, ClaudeBot และ Google-Extended หรือไม่ เราจะวิเคราะห์ robots.txt, meta tags และส่วนหัวของคุณทันที
AI Crawler คืออะไร?
AI crawler คือบอทเว็บที่ใช้โดยบริษัทอย่าง OpenAI, Anthropic และ Google เพื่อจัดทำดัชนีเนื้อหาสำหรับโมเดล AI ของพวกเขา เมื่อ crawler เหล่านี้ถูกบล็อก แบรนด์ของคุณจะมองไม่เห็นในผู้ช่วย AI — พวกเขาไม่สามารถอ้างอิงสิ่งที่อ่านไม่ได้
ทำไม robots.txt จึงสำคัญ
- ควบคุมว่าบอทใดสามารถเข้าถึงเนื้อหาของไซต์คุณได้
- การบล็อก AI crawler เป็นสาเหตุอันดับ 1 ของการมองไม่เห็นใน AI
- หลายไซต์บล็อก AI bot โดยไม่รู้ตัวด้วยกฎ wildcard
สิ่งที่เราตรวจสอบ
- กฎใน robots.txt สำหรับ user agent ของบอต AI 11 ตัว
- Meta robots tags (noindex, noai, noimageai)
- ส่วนหัว HTTP X-Robots-Tag
คำถามที่พบบ่อย
บอต AI คือบอตที่บริษัท AI ใช้ดึงหน้าเว็บ บางตัวใช้เทรนโมเดล บางตัวสร้างดัชนีที่อยู่เบื้องหลังคำตอบของผู้ช่วย และบางตัวเปิดลิงก์ทันทีที่ผู้ใช้ถาม ถ้า robots.txt ปิดกั้นตัวที่ทำหน้าที่ตอบ เนื้อหาของคุณจะไม่ปรากฏในคำตอบของ ChatGPT, Claude, Gemini หรือ Perplexity ไม่ว่าจะเขียนดีแค่ไหน นี่คือสาเหตุที่พบบ่อยที่สุดที่แบรนด์หายไปจากคำตอบของ AI
ไฟล์ robots.txt ของคุณควบคุมว่าบอทใดสามารถเข้าถึงเว็บไซต์ของคุณได้ หลายไซต์บล็อก AI crawler เช่น GPTBot, ClaudeBot และ Google-Extended โดยไม่รู้ตัว ไม่ว่าจะผ่านกฎ Disallow เฉพาะหรือการบล็อก wildcard แบบกว้าง ซึ่งหมายความว่าผู้ช่วย AI ไม่สามารถอ่านเนื้อหาของคุณได้และจะไม่อ้างอิงหรือแนะนำแบรนด์ของคุณ
ทั้งหมด 11 user agent — OpenAI: GPTBot (เทรนโมเดล), OAI-SearchBot (ดัชนีค้นหาของ ChatGPT), ChatGPT-User (ดึงหน้าเมื่อผู้ใช้สั่ง) Anthropic: ClaudeBot (เทรนโมเดล), Claude-SearchBot (ดัชนีค้นหา), Claude-User (ดึงหน้าเมื่อผู้ใช้สั่ง) และโทเคนเก่า anthropic-ai ที่ยังพบในไฟล์รุ่นเดิม นอกจากนี้ยังมี Google-Extended (การอ้างอิงและการเทรนของ Gemini), PerplexityBot, Bytespider (ByteDance) และ CCBot (Common Crawl) ทั้งหมดควบคุมแยกกัน คุณจึงอนุญาตเฉพาะตัวที่อ้างอิงคุณ และปฏิเสธตัวที่นำไปเทรนได้
นอกเหนือจาก robots.txt แล้ว หน้าเว็บสามารถจำกัดการเข้าถึงของ crawler ผ่าน <meta name="robots"> tags ใน HTML และส่วนหัว HTTP X-Robots-Tag คำสั่งเช่น noindex, nofollow, noai และ noimageai บอก crawler ไม่ให้จัดทำดัชนีหรือใช้เนื้อหาของคุณ เครื่องมือของเราตรวจสอบทั้งสามชั้นของการควบคุมการเข้าถึงของ crawler
ใช่ เครื่องมือตรวจสอบ AI crawler ของ GeoVector ฟรีทั้งหมดโดยไม่ต้องสมัครสมาชิก คุณสามารถสแกน URL ใดก็ได้และรับผลลัพธ์ทันทีที่แสดงสถานะ AI crawler ใน robots.txt, meta robots tags และส่วนหัว X-Robots-Tag สำหรับการตรวจสอบทั้งเว็บไซต์และการตรวจสอบอย่างต่อเนื่อง GeoVector มีแพ็กเกจพรีเมียมให้เลือก
บอต AI ที่เครื่องมือนี้ตรวจ
user agent 11 ตัว จัดกลุ่มตามผู้ให้บริการ ประเด็นสำคัญคือแต่ละตัวเป็นอิสระต่อกัน การอนุญาตเครื่องมือตอบคำถามพร้อมกับปฏิเสธบอตที่นำไปเทรนจึงเป็นการตั้งค่าที่รองรับอย่างเป็นทางการ ไม่ใช่ความขัดแย้ง
| Crawler | Operator | Used for | If you disallow it |
|---|---|---|---|
GPTBot | OpenAI | Crawling content that may be used to train OpenAI foundation models. | Your pages are excluded from that training data. |
OAI-SearchBot | OpenAI | Building the index behind ChatGPT's search answers. | You drop out of ChatGPT search answers, though you can still appear as a navigational link. |
ChatGPT-User | OpenAI | Fetching a page when a ChatGPT user or tool opens a link. | ChatGPT can't open your pages on request during a conversation. |
ClaudeBot | Anthropic | Collecting web content that may contribute to model training. | Future material is signalled as out of scope for training. |
Claude-SearchBot | Anthropic | Indexing pages to improve Claude's search results. | Your content stops being indexed for Claude search. |
Claude-User | Anthropic | Retrieving a page when a Claude user asks about it. | Claude can't pull your content into a user-directed answer. |
anthropic-ai | Anthropic | A legacy token that still appears in older robots.txt files; Anthropic no longer lists it among its active agents. | No effect on the three current agents. Harmless to leave in place. |
Google-Extended | Controlling whether your content is used for Gemini grounding and training. | You are removed from Gemini grounding. Google Search crawling, indexing, and ranking are unaffected. | |
PerplexityBot | Perplexity | Indexing pages so Perplexity can surface and cite them. | Perplexity stops citing your pages. |
Bytespider | ByteDance | Crawling for ByteDance's models and products. | Your content is excluded from ByteDance's collection. |
CCBot | Common Crawl | Building the open Common Crawl corpus that many model builders train on. | You are excluded from a dataset used well beyond any single vendor. |
รูปแบบ robots.txt สามแบบ
คัดลอกแบบที่ตรงกับนโยบายของคุณ robots.txt อยู่ที่รากของแต่ละโฮสต์ และซับโดเมนแต่ละตัวต้องมีไฟล์ของตัวเอง
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
Sitemap: https://example.com/sitemap.xml# Answer questions about us
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
# But do not train on us
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /เมื่อฝ่ายกฎหมายขอไม่ให้นำไปเทรน แต่ฝ่ายการตลาดอยากถูกอ้างอิง การแยกแบบนี้คือสิ่งที่ทีมส่วนใหญ่ต้องการจริง ๆ แต่ใช้ได้เฉพาะกับผู้ให้บริการที่แยกสองบทบาทนี้ออกจากกัน ซึ่ง OpenAI, Anthropic และ Google แยกไว้ ส่วน PerplexityBot กับ CCBot ไม่แยก จึงเหลือทางเลือกแค่อนุญาตหรือปฏิเสธทั้งหมด
User-agent: *
Disallow: /
User-agent: GPTBot
Allow: /ตาม RFC 9309 บอตจะทำตามกลุ่มกฎที่เจาะจงที่สุดซึ่งตรงกับชื่อของมันเท่านั้น และมองข้ามกลุ่มอื่นทั้งหมด ดังนั้น GPTBot ในตัวอย่างนี้อ่านเฉพาะกลุ่มของตัวเองและได้รับอนุญาต ขณะที่บอตทุกตัวที่ไม่มีกลุ่มของตัวเอง รวมถึง OAI-SearchBot และ Claude-SearchBot จะตกไปอยู่กลุ่มไวลด์การ์ดและถูกปิดกั้นทั้งหมด robots.txt ของเซิร์ฟเวอร์ทดสอบที่หลุดขึ้นโปรดักชันพังแบบนี้พอดี
การควบคุมระดับหน้า: meta robots และ X-Robots-Tag
robots.txt กำหนดว่าหน้านั้นดึงได้หรือไม่ ส่วนสองอย่างนี้กำหนดว่าเมื่อดึงไปแล้วทำอะไรได้บ้าง หน้าเว็บจึงอาจเข้าถึงได้เต็มที่แต่ยังไม่ถูกนำไปแสดงในผลลัพธ์
<meta name="robots" content="noindex, nofollow">X-Robots-Tag: noindexคุณจะเห็น noai และ noimageai ด้วย ทั้งสองมาจากแพลตฟอร์มโฮสต์รูปภาพ ไม่ได้อยู่ในข้อกำหนดใด และไม่มีบอตในรายการข้างต้นที่ประกาศว่ารองรับ เรารายงานให้เพราะมันบันทึกเจตนาของคุณ ไม่ใช่เพราะมันปิดกั้นอะไรได้จริง
แหล่งอ้างอิงและข้อกำหนด
ทุกอย่างในหน้านี้ตรวจสอบกับเอกสารต้นทางด้านล่าง หากผู้ให้บริการเปลี่ยนกฎ ที่นั่นคือที่แรกที่จะเห็นการเปลี่ยนแปลง
- RFC 9309: Robots Exclusion Protocol — IETF
- Overview of OpenAI crawlers — OpenAI
- Does Anthropic crawl data from the web? — Anthropic
- Google crawlers, fetchers, and user agents — Google Search Central
- PerplexityBot — Perplexity
- How Google interprets the robots.txt specification — Google Search Central
ตรวจทานล่าสุด เราตรวจหน้านี้เทียบกับข้อกำหนดต้นทางใหม่ทุกครั้งที่ผู้ให้บริการอัปเดต