Mô-đun 7d — The Agentic Internet: AI Crawler & Bot Control
Mục tiêu: Kiểm soát cách AI crawler và bot truy cập nội dung công khai của chính bạn — xem ai đang crawl, cho phép/chặn theo mục đích, áp dụng robots.txt, và (tùy chọn) kiếm tiền từ truy cập bằng Pay Per Crawl.
|
|
| 👤 Ai làm việc này |
Nhóm web / nội dung / bảo mật (chủ site) |
| ⏱️ Thời gian |
~30 phút |
| 🎯 Kết thúc bạn sẽ có |
Tầm nhìn về AI crawler trên các site của bạn và chính sách cho phép / chặn / tính phí có chủ đích theo từng crawler |
| ✋ Trước khi bắt đầu |
Một domain (zone) trên Cloudflare. Đây là Application Security (bảo vệ lưu lượng vào các site công khai của bạn) — khác với các mô-đun Zero Trust, nhưng thuộc cùng câu chuyện bảo mật AI. |
🧭 Một chiều khác. Mọi mô-đun khác trong hướng dẫn này bảo mật lưu lượng đi ra — nhân viên và agent của bạn tới internet, ứng dụng và mô hình AI. Mô-đun này là chiều ngược lại: AI crawler và agent đi vào các website công khai của bạn. Khi internet trở nên "agentic," luồng vào đó giờ là phần lớn lưu lượng — nên kiểm soát nó là một kỷ luật riêng.
Vì sao điều này quan trọng ngay bây giờ (báo cáo agentic-internet 2026 của Cloudflare)
- AI đang được áp dụng nhanh hơn smartphone khoảng 2× — 2,5 tỷ người dùng (30% nhân loại) trong 3,5 năm.
- Hơn 50% lưu lượng internet hiện không phải con người — bot và AI agent đã vượt ngưỡng lịch sử đó.
- 52% request crawler là để huấn luyện AI (tháng 6 năm 2026), tăng từ 22% mùa xuân 2025.
- Với mỗi giờ người ta dành để tìm kiếm trực tuyến, chỉ khoảng 15 phút là trên web mở — câu trả lời AI ngày càng thay thế lượt nhấp ("Google Zero").
- Sự trao đổi cũ — nội dung đổi lấy lưu lượng giới thiệu — đang tan rã, nên một nền kinh tế cấp phép nội dung đang hình thành (hơn 50 thỏa thuận nhà xuất bản–AI từ 2023).
Dịch ra cho bạn: mọi tổ chức xuất bản nội dung trực tuyến giờ cần quyết định, một cách có chủ đích, hệ thống AI nào được truy cập, vì mục đích gì, và theo điều khoản nào. AI Crawl Control của Cloudflare là nơi bạn làm việc đó.
Phần A — Xem ai đang crawl bạn (AI Audit)
Mọi site trên Cloudflare đều có quyền truy cập AI Audit trong AI Crawl Control.
- 👉 Trong Cloudflare dashboard, chọn account and domain của bạn, rồi vào AI Crawl Control.
- 👉 Mở tab Crawlers.
- 📺 Bạn sẽ thấy: bảng AI crawler đang truy cập site của bạn, với:
- Crawler và operator sở hữu nó (GPTBot/OpenAI, ClaudeBot/Anthropic, Bytespider/ByteDance…)
- Category (AI Crawler, AI Assistant, AI Search, Archiver…)
- Requests (được phép + không thành công, kèm xu hướng)
- Robots.txt violations
- 👉 Mở tab Metrics để vẽ biểu đồ hoạt động theo thời gian — nhóm by Crawler / Category / Operator / Host / Status code, và (trên gói trả phí) xem referrer analytics cho biết nhà vận hành AI nào gửi lưu lượng cho bạn.
✅ Điểm kiểm tra: bạn thấy AI crawler nào truy cập nội dung của bạn và mức độ đậm đặc thế nào.
💡 Chất lượng phát hiện: trên gói Free, crawler được nhận diện bằng user-agent (bắt các bot tự nhận diện, nổi tiếng). Gói trả phí bổ sung phát hiện Bot Management của Cloudflare cho crawler không tự nhận diện.
Phần B — Kiểm soát từng crawler (cho phép / chặn / tính phí)
Quyết định, theo từng crawler, nó được truy cập nội dung của bạn thế nào.
- 👉 Trong tab Crawlers, tìm một crawler và dùng cột Action:
- Allow — cho phép scrape (tốt với crawler gửi trích dẫn/giới thiệu, hoặc khi bạn có thỏa thuận).
- Block — dừng hoàn toàn. Bạn có thể configure the block response được trả về.
- Charge — yêu cầu thanh toán theo từng request (Pay Per Crawl — Phần D).
- 👉 Tùy chọn bật Enforce robots.txt để Cloudflare tự động thực thi quy tắc
robots.txt của bạn với crawler đó.
Đặt chính sách theo mục đích (khuyến nghị)
Thay vì từng crawler một, hãy đặt chính sách dựa trên hành vi:
- 👉 Vào Security → Settings → Configure AI bot policies.
- 👉 Chọn biện pháp cho từng behavior preset:
| Hành vi |
Nó là gì |
Lựa chọn phổ biến |
| Search |
Lập chỉ mục nội dung của bạn để trả lời câu hỏi sau này |
Allow (giữ bạn được tìm thấy) |
| Agent |
Hành động thời gian thực thay mặt một người (chat-fetch, browser-use) |
Allow hoặc Block trên trang có quảng cáo |
| Training |
Lấy nội dung để huấn luyện/tinh chỉnh mô hình (gồm Search+Training hỗn hợp) |
Block nếu bạn không muốn nuôi huấn luyện mô hình |
- 👉 Mỗi preset có Block (all pages), Block on pages with ads, hoặc Allow (do not block).
⚠️ Mặc định mới từ 15 tháng 9 năm 2026: với domain mới, crawler được phân loại Training hoặc Agent bị chặn trên các trang hiển thị quảng cáo, trong khi Search vẫn được phép; crawler đa mục đích (Search+Training) bị chặn bởi bất kỳ thiết lập chặn training nào. Công tắc một-nhấp cũ "Block AI bots" đang được thay bằng các chính sách hành vi này. Xem lại thiết lập của bạn để khớp với ý định.
✅ Điểm kiểm tra: mỗi mục đích crawler có hành động có chủ đích; crawler bị chặn nhận phản hồi bạn đã cấu hình.
💡 Cần kiểm soát tinh hơn? Dùng Cloudflare WAF custom rules (theo đường dẫn, theo crawler) hoặc Redirect Rules để dẫn crawler tới khu vực cụ thể — ví dụ cho phép crawl /blog nhưng chặn /pricing.
Phần C — Áp dụng robots.txt
robots.txt là chính sách bạn công bố; AI Crawl Control làm cho nó nhìn thấy được và thực thi được.
- 👉 Trong AI Crawl Control → Robots.txt, xem sức khỏe tệp, crawler violations và Agent Readiness.
- 👉 Bật managed robots.txt / enforcement để Cloudflare thực thi chỉ thị của bạn bằng WAF rule tự động (thay vì dựa vào crawler tự nguyện tuân thủ).
- 💡 Cân nhắc Content Signals Policy (contentsignals.org) để tuyên bố cách nội dung của bạn được dùng sau khi truy cập (ví dụ search thì có, huấn luyện AI thì không) ngay trong
robots.txt.
✅ Điểm kiểm tra: các vi phạm nhìn thấy được, và crawler không tuân thủ có thể bị chặn từ tab Crawlers hoặc qua WAF.
Phần D — Kiếm tiền từ truy cập (Pay Per Crawl)
Nếu nội dung của bạn có giá trị huấn luyện, bạn có thể tính phí AI crawler thay vì chỉ cho phép hoặc chặn — biến sự khan hiếm thành doanh thu.
🔒 Pay Per Crawl đang ở private/closed beta. Tham gia tại cloudflare.com/paypercrawl-signup hoặc hỏi giám đốc tài khoản Cloudflare của bạn.
Cách hoạt động (phía chủ site):
- 👉 Đặt price của bạn và chọn which crawlers cần tính phí.
- 👉 Kết nối Stripe để thanh toán và theo dõi delivery analytics.
- 📺 Khi crawler yêu cầu nội dung trả phí mà không đồng ý trả, Cloudflare trả
HTTP 402 Payment Required. Crawler chọn tham gia gửi ý định thanh toán qua HTTP header đã ký (crawler-max-price / crawler-exact-price, qua Web Bot Auth) và bị tính phí theo mỗi lần crawl thành công.
- 👉 Giữ trang khám phá miễn phí: một số đường dẫn luôn miễn phí (
/robots.txt, /sitemap.xml, /security.txt, /.well-known/security.txt, /crawlers.json), và bạn có thể miễn thêm qua Rules → Configuration Rules → Disable Pay Per Crawl trên một mẫu URI.
✅ Điểm kiểm tra: các crawler đã chọn bị tính phí (hoặc nhận 402) trong khi đường dẫn miễn phí/khám phá của bạn vẫn mở.
Phần E — Theo dõi xu hướng
- 👉 Cloudflare Radar (radar.cloudflare.com) công bố thông tin bot & AI-crawler trên internet — ngữ cảnh hữu ích về cách hành vi crawler đang đổi ở quy mô ngành.
- 👉 Quay lại AI Crawl Control → Metrics thường xuyên; hỗn hợp crawler đổi nhanh (tỷ lệ training-so-với-search đã đảo mạnh chỉ trong một năm).
Cách này khớp với bức tranh bảo mật AI
| Chiều |
Bạn đang bảo vệ gì |
Mô-đun |
| Đi ra — nhân viên & agent của bạn dùng AI |
Dữ liệu, danh tính, chi phí, quyền truy cập công cụ |
5d · 6 · 7 · 7b · 7c |
| Đi vào — AI truy cập nội dung & ứng dụng của bạn |
Nội dung, IP, kiếm tiền, lạm dụng ứng dụng |
7d (mô-đun này) + AI Security for Apps (WAF) |
Cùng nhau chúng bao phủ toàn bộ "agentic internet": bạn quản trị cách tổ chức tiêu thụ AI và cách hệ sinh thái AI tiêu thụ bạn.
✅ Hoàn thành Mô-đun 7d!
Bây giờ bạn có:
- ✅ Tầm nhìn về AI crawler nào truy cập các site của bạn (AI Audit)
- ✅ Hành động cho phép / chặn / tính phí có chủ đích theo từng crawler và theo mục đích (Search / Agent / Training)
- ✅
robots.txt được làm nhìn thấy được và thực thi được
- ✅ (Beta) Pay Per Crawl để kiếm tiền từ truy cập AI
- ✅ Nhận thức về thay đổi mặc định tháng 9 năm 2026 và nơi tinh chỉnh chúng
Khắc phục nhanh
| Vấn đề |
Cách khắc phục |
| Không có AI Crawl Control cho site của tôi |
Nó ở cấp zone — chọn account và domain trước; domain phải nằm trên Cloudflare |
| Crawler không được nhận diện |
Gói Free chỉ dùng user-agent; nâng cấp để có phát hiện Bot Management với bot không tự nhận diện |
| Đã chặn crawler nhưng nó vẫn lập chỉ mục |
Xác nhận Action = Block đã lưu, và không có WAF/rule ưu tiên cao hơn cho phép nó; một số bot bỏ qua robots.txt (thực thi qua Cloudflare, Phần C) |
| Muốn cho phép search nhưng chặn training |
Dùng các behavior preset Configure AI bot policies — Allow Search, Block Training (Phần B) |
| Thiếu tùy chọn Pay Per Crawl |
Đây là closed beta — đăng ký hoặc liên hệ đội ngũ tài khoản của bạn (Phần D) |
| Vô tình chặn bot tốt |
Kiểm tra Configure AI bot policies (mặc định tháng 9 năm 2026 có thể chặn Training/Agent trên trang có quảng cáo) và điều chỉnh |
Kết nối cả văn phòng và trung tâm dữ liệu với mạng của Cloudflare.
Module 7d — The Agentic Internet: AI Crawler & Bot Control
Goal: Control how AI crawlers and bots access your own public content — see who's crawling, allow/block them by purpose, enforce robots.txt, and (optionally) monetize access with Pay Per Crawl.
|
|
| 👤 Who does this |
Web / content / security team (site owners) |
| ⏱️ Time |
~30 minutes |
| 🎯 You'll finish with |
Visibility into AI crawlers on your sites and a deliberate allow / block / charge policy per crawler |
| ✋ Before you begin |
A domain (zone) on Cloudflare. This is Application Security (protecting inbound traffic to your public sites) — different from the Zero Trust modules, but part of the same AI security story. |
🧭 A different direction. Every other module in this guide secures traffic going out — your people and agents reaching the internet, apps, and AI models. This module is the opposite direction: AI crawlers and agents coming in to your public websites. As the internet becomes "agentic," that inbound flow is now the majority of traffic — so controlling it is its own discipline.
Why this matters now (Cloudflare's 2026 agentic-internet report)
- AI is being adopted ~2× faster than smartphones — 2.5 billion users (30% of humanity) in 3.5 years.
- More than 50% of internet traffic is now non-human — bots and AI agents crossed that historic threshold.
- 52% of crawler requests are for AI training (June 2026), up from 22% in spring 2025.
- For every hour people spend searching online, only ~15 minutes is on the open web — AI answers increasingly replace clicks ("Google Zero").
- The old exchange — content for referral traffic — is breaking down, so a content-licensing economy is emerging (50+ publisher–AI agreements since 2023).
Translation for you: any organization that publishes content online now needs to decide, deliberately, which AI systems may access it, for what purpose, and on what terms. Cloudflare's AI Crawl Control is where you do that.
Part A — See who's crawling you (AI Audit)
Every site on Cloudflare has access to AI Audit inside AI Crawl Control.
- 👉 In the Cloudflare dashboard, select your account and domain, then go to AI Crawl Control.
- 👉 Open the Crawlers tab.
- 📺 What you'll see: a table of AI crawlers hitting your site, with:
- Crawler and the operator that owns it (GPTBot/OpenAI, ClaudeBot/Anthropic, Bytespider/ByteDance…)
- Category (AI Crawler, AI Assistant, AI Search, Archiver…)
- Requests (allowed + unsuccessful, with a trend)
- Robots.txt violations
- 👉 Open the Metrics tab to chart activity over time — group by Crawler / Category / Operator / Host / Status code, and (on paid plans) see referrer analytics showing which AI operators send you traffic.
✅ Checkpoint: you can see which AI crawlers access your content and how heavily.
💡 Detection quality: on the Free plan, crawlers are identified by user-agent (catches well-known, self-identifying bots). Paid plans add Cloudflare's Bot Management detection for crawlers that don't self-identify.
Part B — Control each crawler (allow / block / charge)
Decide, per crawler, how it may access your content.
- 👉 In the Crawlers tab, find a crawler and use the Action column:
- Allow — let it scrape (good for crawlers that send citations/referrals, or where you have an agreement).
- Block — stop it entirely. You can configure the block response returned.
- Charge — require payment per request (Pay Per Crawl — Part D).
- 👉 Optionally toggle Enforce robots.txt so Cloudflare upholds your
robots.txt rules for that crawler automatically.
Set policy by purpose (recommended)
Rather than one crawler at a time, set behavior-based policy:
- 👉 Go to Security → Settings → Configure AI bot policies.
- 👉 Choose a mitigation for each behavior preset:
| Behavior |
What it is |
Common choice |
| Search |
Indexes your content to answer questions later |
Allow (keeps you discoverable) |
| Agent |
Real-time actions on a person's behalf (chat-fetch, browser-use) |
Allow or Block on ad pages |
| Training |
Takes content to train/fine-tune models (incl. mixed Search+Training) |
Block if you don't want to feed model training |
- 👉 Each preset offers Block (all pages), Block on pages with ads, or Allow (do not block).
⚠️ New defaults from 15 Sep 2026: for new domains, crawlers classified Training or Agent are blocked on pages that display ads, while Search stays allowed; mixed-purpose (Search+Training) crawlers are blocked by any training-block setting. The legacy one-click "Block AI bots" toggle is being replaced by these behavior policies. Review your setting so it matches your intent.
✅ Checkpoint: each crawler purpose has a deliberate action; a blocked crawler receives your configured response.
💡 Need finer control? Use Cloudflare WAF custom rules (per-path, per-crawler) or Redirect Rules to steer crawlers to specific areas — e.g. allow crawling of /blog but block /pricing.
Part C — Enforce robots.txt
robots.txt is your stated policy; AI Crawl Control makes it visible and enforceable.
- 👉 In AI Crawl Control → Robots.txt, review file health, crawler violations, and Agent Readiness.
- 👉 Turn on managed robots.txt / enforcement so Cloudflare can uphold your directives with an automatic WAF rule (rather than relying on crawlers to voluntarily obey).
- 💡 Consider the Content Signals Policy (contentsignals.org) to declare how your content may be used after access (e.g. search yes, AI training no) directly in
robots.txt.
✅ Checkpoint: violations are visible, and non-compliant crawlers can be blocked from the Crawlers tab or via WAF.
Part D — Monetize access (Pay Per Crawl)
If your content has training value, you can charge AI crawlers instead of just allowing or blocking — turning scarcity into revenue.
🔒 Pay Per Crawl is in private/closed beta. Join at cloudflare.com/paypercrawl-signup or ask your Cloudflare account executive.
How it works (site-owner side):
- 👉 Set your price and select which crawlers to charge.
- 👉 Connect Stripe for payments and watch delivery analytics.
- 📺 When a crawler requests paid content without agreeing to pay, Cloudflare returns
HTTP 402 Payment Required. Crawlers that opt in send payment intent via signed HTTP headers (crawler-max-price / crawler-exact-price, via Web Bot Auth) and are charged per successful crawl.
- 👉 Keep discovery pages free: some paths are always free (
/robots.txt, /sitemap.xml, /security.txt, /.well-known/security.txt, /crawlers.json), and you can exempt more via Rules → Configuration Rules → Disable Pay Per Crawl on a URI pattern.
✅ Checkpoint: chosen crawlers are charged (or receive 402) while your free/discovery paths stay open.
Part E — Monitor the trend
- 👉 Cloudflare Radar (radar.cloudflare.com) publishes bot & AI-crawler insights across the internet — useful context for how crawler behavior is shifting industry-wide.
- 👉 Revisit AI Crawl Control → Metrics regularly; crawler mix changes fast (the training-vs-search split shifted dramatically in a single year).
How this fits the AI security picture
| Direction |
What you're protecting |
Modules |
| Outbound — your people & agents using AI |
Data, identity, cost, tool access |
5d · 6 · 7 · 7b · 7c |
| Inbound — AI accessing your content & apps |
Content, IP, monetization, app abuse |
7d (this) + AI Security for Apps (WAF) |
Together these cover the full "agentic internet": you govern how your organization consumes AI and how the AI ecosystem consumes you.
✅ Module 7d complete!
You now have:
- ✅ Visibility into which AI crawlers access your sites (AI Audit)
- ✅ A deliberate allow / block / charge action per crawler and per purpose (Search / Agent / Training)
- ✅
robots.txt made visible and enforceable
- ✅ (Beta) Pay Per Crawl to monetize AI access
- ✅ Awareness of the Sept 2026 default changes and where to tune them
Quick troubleshooting
| Problem |
Fix |
| No AI Crawl Control for my site |
It's zone-level — select the account and domain first; the domain must be on Cloudflare |
| Crawlers not identified |
Free plan uses user-agent only; upgrade for Bot Management detection of non-self-identifying bots |
| Blocked a crawler but it's still indexing |
Confirm the Action = Block saved, and that no higher-priority WAF/rule allows it; some bots ignore robots.txt (enforce it via Cloudflare, Part C) |
| Want to allow search but block training |
Use Configure AI bot policies behavior presets — Allow Search, Block Training (Part B) |
| Pay Per Crawl options missing |
It's closed beta — sign up or contact your account team (Part D) |
| Accidentally blocking a good bot |
Check Configure AI bot policies (the Sept 2026 defaults may block Training/Agent on ad pages) and adjust |
Connect whole offices and data centers to Cloudflare's network.
Nguồn cộng đồng — không phải tài liệu chính thức của Cloudflare: https://zerotrust.cfsase.workers.dev
Community source — not an official Cloudflare publication: https://zerotrust.cfsase.workers.dev