Mô-đun 7c — AI Gateway: Protect & Implement
Mục tiêu: Đặt Cloudflare AI Gateway trước các mô hình AI mà ứng dụng và agent của bạn gọi — để mọi yêu cầu AI theo chương trình được xác thực, quan sát được, kiểm soát chi phí, kiểm duyệt nội dung không an toàn, và được quét dữ liệu nhạy cảm bằng DLP profiles Zero Trust của bạn.
|
|
| 👤 Ai làm việc này |
Nhóm nền tảng / kỹ thuật AI + bảo mật |
| ⏱️ Thời gian |
~40 phút |
| 🎯 Kết thúc bạn sẽ có |
Một AI Gateway đã xác thực với guardrails, DLP, ghi log và kiểm soát chi phí đứng trước các nhà cung cấp AI của bạn |
| ✋ Trước khi bắt đầu |
Một ứng dụng/agent gọi mô hình AI (OpenAI, Anthropic, Workers AI, Google…). Quét DLP dùng DLP profiles Zero Trust (Mô-đun 6) — Enterprise. |
🧭 AI Gateway nằm ở đâu. Mô-đun 7 quản trị người dùng AI trong trình duyệt. Mô-đun 7b quản trị AI agent kết nối tới công cụ (MCP). AI Gateway quản trị chiều còn lại: ứng dụng và agent của chính bạn gọi mô hình AI qua API. Đó là proxy nằm giữa mã của bạn và nhà cung cấp AI, nên bạn có kiểm soát và tầm nhìn mà không đổi mô hình hay nhà cung cấp — và quan trọng là, nó hoạt động ở lớp API, không cần client trên thiết bị hay TLS decryption.
Your app / agent / coding tool
│ (change base URL → gateway endpoint)
▼
┌─────────────────────────────────────────────┐
│ CLOUDFLARE AI GATEWAY │
│ Auth ▸ Guardrails ▸ DLP ▸ Rate/Spend limits│
│ ▸ Cache ▸ Logs ▸ Fallbacks │
└───────────────────┬─────────────────────────┘
▼
AI providers: OpenAI · Anthropic · Workers AI · Google · …
Phần A — Tạo & kết nối gateway
- 👉 Trong Cloudflare dashboard, vào AI → AI Gateway.
- 👉 Nhấp Create Gateway, đặt tên (ví dụ
production), rồi tạo.
- 📺 Bạn sẽ thấy OpenAI-compatible endpoint của gateway và hướng dẫn thiết lập.
- 💡 Lối tắt: bạn cũng có thể dùng gateway ID
default — AI Gateway tự tạo nó ở request đầu tiên.
- 👉 Hướng ứng dụng của bạn tới gateway. Trong mã, đổi base URL của nhà cung cấp AI thành endpoint gateway của bạn:
https://gateway.ai.cloudflare.com/v1/{account_id}/{gateway_id}/{provider}
Ví dụ, base URL của OpenAI client trở thành …/{account_id}/production/openai. Không đổi mã nào khác — gateway chuyển tiếp tới nhà cung cấp.
- 👉 Gửi một request thử từ ứng dụng của bạn.
✅ Điểm kiểm tra: Trong AI → AI Gateway → gateway của bạn, request xuất hiện kèm số liệu (số request, token, chi phí, độ trễ). Lưu lượng giờ đang chảy qua gateway.
💡 Không cần WARP client, không cần TLS decryption, không đổi mạng — AI Gateway là điểm kiểm soát lớp API. Đó là lý do nó phù hợp với dịch vụ backend, tác vụ hàng loạt và agent tự hành.
Phần B — Bảo vệ gateway (Authenticated Gateway)
Mặc định, bất kỳ ai biết URL endpoint đều có thể gửi request qua nó. Khóa lại để chỉ caller được ủy quyền của bạn mới dùng được.
- 👉 Trong gateway của bạn, mở Settings.
- 👉 Bật Authenticated Gateway.
- 👉 Tạo AI Gateway authentication token và lưu trong trình quản lý bí mật của bạn.
- 👉 Để ứng dụng gửi nó trên mọi request dưới dạng header:
cf-aig-authorization: Bearer <your-aig-token>
✅ Điểm kiểm tra: request không có token bị từ chối; request có token thành công. Dù URL endpoint bị lộ, nó cũng không thể bị lạm dụng.
⚠️ Lưu ý: token này xác thực việc dùng gateway; provider API keys của bạn là thứ riêng. Cân nhắc BYOK / Store Keys (Settings) để khóa nhà cung cấp nằm trong Cloudflare và không bao giờ đi kèm mã ứng dụng.
Phần C — Guardrails (chặn nội dung không an toàn)
Guardrails kiểm duyệt nội dung chảy qua gateway — bắt các prompt và phản hồi có hại hoặc không phù hợp trước khi chúng tới người dùng hoặc mô hình.
- 👉 Trong gateway → Settings → Guardrails → bật.
- 👉 Đặt evaluation scope: kiểm duyệt user prompts, model responses, hoặc both.
- 👉 Chọn các hazard categories cần theo dõi, và với mỗi loại quyết định Block (dừng lại) hoặc Flag (cho phép nhưng ghi log).
- 👉 Lưu.
✅ Điểm kiểm tra: một prompt/phản hồi thuộc hazard category bị chặn bị dừng (hoặc gắn cờ trong log), thấy được trong log của gateway.
💡 Guardrails so với DLP: Guardrails xử lý an toàn (nội dung độc hại, có hại, không an toàn); DLP (phần tiếp) xử lý dữ liệu nhạy cảm (PII, secret, mã nguồn). Dùng cả hai.
Phần D — DLP cho AI Gateway (quét dữ liệu nhạy cảm) (Enterprise)
Đây là tích hợp Zero Trust: AI Gateway có thể quét văn bản của request và response đối chiếu với DLP profiles Zero Trust của bạn — cùng các profile bạn đã xây trong Mô-đun 6 — với không cần lọc HTTP và không cần TLS decryption.
- 👉 Trong gateway → Features → DLP → Set up.
- 👉 Chọn các DLP profiles cần áp dụng — ví dụ Credentials and Secrets, PII, hoặc profile AI Prompt (PII, Source Code, Financial, chủ đề jailbreak/intent; thậm chí chủ đề prompt ngôn ngữ tự nhiên tùy chỉnh).
- 👉 Chọn quét gì bằng Check: Request (prompt gửi tới mô hình), Response (mô hình trả về), hoặc Both.
- 👉 Đặt hành động — flag (ghi log) hoặc block — khi khớp.
- 👉 Lưu.
✅ Điểm kiểm tra: một request chứa secret giả / PII thử nghiệm bị gắn cờ hoặc chặn, và phát hiện hiện trong log kèm profile đã khớp.
⚠️ Lưu ý streaming (quan trọng với độ trễ): nếu bạn bật quét Response, AI Gateway đệm toàn bộ phản hồi nhà cung cấp trước khi trả về, làm tăng thời gian tới token đầu tiên — dễ thấy với agent chat/coding dùng stream. Nếu bạn cần stream độ trễ thấp, đặt Check = Request only, hoặc định tuyến lưu lượng nhạy cảm về độ trễ qua gateway riêng.
💡 Lý tưởng cho coding agent: công cụ như Cursor/Claude Code gửi mã nguồn, cấu hình và secret tới nhà cung cấp mô hình. Vì AI Gateway nằm trên đường đi, DLP bắt secret/dữ liệu được quản lý đang rời đi (hoặc trở về) mà không sửa agent.
Phần E — Kiểm soát chi phí, tốc độ & quyền riêng tư
AI Gateway cũng cho bạn các guardrail vận hành:
| Kiểm soát |
Nó làm gì |
Ở đâu |
| Rate limiting |
Giới hạn số request trong một cửa sổ thời gian |
Settings → Rate limiting |
| Spend limits |
Giới hạn chi phí tính bằng đô la (theo mô hình, nhà cung cấp, hoặc metadata tùy chỉnh — ví dụ $200/ngày mỗi người dùng) |
Features → Spend limits |
| Caching |
Phục vụ request giống nhau từ cache — giảm độ trễ + chi phí nhà cung cấp |
Settings → Cache Responses |
| Fallbacks / retries |
Tự thử lại hoặc chuyển sang mô hình khác khi nhà cung cấp lỗi |
Configuration → Fallbacks |
| Log payload control |
Gửi cf-aig-collect-log-payload: false để chỉ ghi metadata, không ghi thân prompt/response |
header theo từng request |
💡 Mẹo quyền riêng tư: với dữ liệu được quản lý, dùng cf-aig-collect-log-payload: false để bạn giữ số liệu sử dụng (token, chi phí, mô hình, trạng thái) mà không lưu văn bản prompt/response nhạy cảm.
⚠️ Đừng bật caching hay rate limiting trên gateway dùng chung bởi một instance AI Search / RAG — nó can thiệp vào embedding/indexing. Dùng gateway riêng cho việc đó.
Phần F — Quan sát & xác minh
- 👉 Trong gateway, xem Analytics (request, token, chi phí, tỷ lệ cache hit, lỗi) và Logs (theo từng request: mô hình, nhà cung cấp, trạng thái, chi phí, thời lượng, user agent, và — trừ khi đã tắt — payload).
- 👉 Lọc log để xác nhận các hành động Auth, Guardrails và DLP đang kích hoạt như mong đợi.
- 👉 Kiểm tra đầu-cuối: gửi (a) request bình thường → thành công; (b) request thiếu auth token → bị từ chối; (c) prompt có nội dung không an toàn → guardrail chặn; (d) prompt có secret thử nghiệm → DLP gắn cờ/chặn.
✅ Điểm kiểm tra: cả bốn hành vi đều đúng và thấy được trong log.
✅ Hoàn thành Mô-đun 7c!
Bây giờ bạn có:
- ✅ Một AI Gateway đứng trước các nhà cung cấp AI (một endpoint, nhiều nhà cung cấp)
- ✅ Authenticated Gateway — chỉ caller được ủy quyền mới dùng được
- ✅ Guardrails chặn nội dung không an toàn trong prompt/response
- ✅ DLP quét prompt/response tìm dữ liệu nhạy cảm bằng Zero Trust profiles của bạn
- ✅ Kiểm soát chi phí, tốc độ, cache và quyền riêng tư
- ✅ Đầy đủ ghi log & analytics cho mọi cuộc gọi AI
Bức tranh bảo mật AI đầy đủ
| Lớp |
Quản trị |
Mô-đun |
| Dùng AI trên trình duyệt |
Người dán vào ChatGPT, v.v. |
7 — AI controls |
| AI agent → công cụ |
Máy chủ MCP & portal |
7b — Secure AI & MCP |
| Ứng dụng của bạn → mô hình AI |
AI theo chương trình/API |
7c (mô-đun này) |
| Phát hiện & phê duyệt ứng dụng AI |
Shadow AI |
5d — Shadow IT |
| Phát hiện dữ liệu nhạy cảm |
DLP profiles |
6 — DLP |
Khắc phục nhanh
| Vấn đề |
Cách khắc phục |
| Request bỏ qua gateway |
Xác nhận base URL của ứng dụng trỏ tới endpoint gateway (Phần A) |
| Bất kỳ ai cũng dùng được gateway |
Bật Authenticated Gateway + gửi header cf-aig-authorization (Phần B) |
| DLP → Set up không có |
DLP cho AI Gateway dùng Zero Trust DLP — xác nhận Enterprise + rằng các profile tồn tại (Mô-đun 6) |
| Streaming chậm sau khi bật DLP |
Quét Response đệm toàn bộ phản hồi — đặt Check = Request only hoặc dùng gateway riêng (Phần D) |
| Prompt nhạy cảm được lưu trong log |
Gửi cf-aig-collect-log-payload: false để chỉ ghi metadata (Phần E) |
| Độ chính xác AI Search giảm |
Đừng bật caching/rate limiting trên gateway mà instance RAG của bạn dùng (Phần E) |
Kết nối cả văn phòng và trung tâm dữ liệu với mạng của Cloudflare.
Module 7c — AI Gateway: Protect & Implement
Goal: Put Cloudflare AI Gateway in front of the AI models your applications and agents call — so every programmatic AI request is authenticated, observable, cost-controlled, moderated for unsafe content, and scanned for sensitive data with your Zero Trust DLP profiles.
|
|
| 👤 Who does this |
Platform / AI engineering + security team |
| ⏱️ Time |
~40 minutes |
| 🎯 You'll finish with |
An authenticated AI Gateway with guardrails, DLP, logging, and cost controls in front of your AI providers |
| ✋ Before you begin |
An app/agent that calls an AI model (OpenAI, Anthropic, Workers AI, Google…). DLP scanning uses Zero Trust DLP profiles (Module 6) — Enterprise. |
🧭 Where AI Gateway fits. Module 7 governs people using AI in a browser. Module 7b governs AI agents connecting to tools (MCP). AI Gateway governs the other direction: your own apps and agents calling AI models over the API. It's a proxy that sits between your code and the AI provider, so you get control and visibility without changing the model or the provider — and crucially, it works at the API layer with no device client or TLS decryption required.
Your app / agent / coding tool
│ (change base URL → gateway endpoint)
▼
┌─────────────────────────────────────────────┐
│ CLOUDFLARE AI GATEWAY │
│ Auth ▸ Guardrails ▸ DLP ▸ Rate/Spend limits│
│ ▸ Cache ▸ Logs ▸ Fallbacks │
└───────────────────┬─────────────────────────┘
▼
AI providers: OpenAI · Anthropic · Workers AI · Google · …
Part A — Create & connect a gateway
- 👉 In the Cloudflare dashboard, go to AI → AI Gateway.
- 👉 Click Create Gateway, give it a name (e.g.
production), and create it.
- 📺 You'll see the gateway's OpenAI-compatible endpoint and setup guidance.
- 💡 Shortcut: you can also use the gateway ID
default — AI Gateway auto-creates it on the first request.
- 👉 Point your app at the gateway. In your code, change the AI provider's base URL to your gateway endpoint:
https://gateway.ai.cloudflare.com/v1/{account_id}/{gateway_id}/{provider}
For example, an OpenAI client's base URL becomes …/{account_id}/production/openai. No other code changes — the gateway forwards to the provider.
- 👉 Send a test request from your app.
✅ Checkpoint: In AI → AI Gateway → your gateway, the request appears with metrics (request count, tokens, cost, latency). Traffic is now flowing through the gateway.
💡 No WARP client, no TLS decryption, no network changes — AI Gateway is an API-layer control point. That's what makes it the right tool for backend services, batch jobs, and autonomous agents.
Part B — Protect the gateway (Authenticated Gateway)
By default anyone who knows the endpoint URL could send requests through it. Lock it down so only your authorized callers can use it.
- 👉 In your gateway, open Settings.
- 👉 Enable Authenticated Gateway.
- 👉 Create an AI Gateway authentication token and store it in your secrets manager.
- 👉 Have your app send it on every request as a header:
cf-aig-authorization: Bearer <your-aig-token>
✅ Checkpoint: requests without the token are rejected; requests with it succeed. Even if the endpoint URL leaks, it can't be abused.
⚠️ Watch out: this token authenticates use of the gateway; your provider API keys are separate. Consider BYOK / Store Keys (Settings) so the provider key lives in Cloudflare and never ships in your app code.
Part C — Guardrails (block unsafe content)
Guardrails moderate the content flowing through the gateway — catching harmful or inappropriate prompts and responses before they reach users or models.
- 👉 In the gateway → Settings → Guardrails → enable.
- 👉 Set the evaluation scope: moderate user prompts, model responses, or both.
- 👉 Choose the hazard categories to monitor, and for each decide Block (stop it) or Flag (allow but log).
- 👉 Save.
✅ Checkpoint: a prompt/response in a blocked hazard category is stopped (or flagged in logs), visible in the gateway's logs.
💡 Guardrails vs DLP: Guardrails handle safety (toxic, harmful, unsafe content); DLP (next part) handles sensitive data (PII, secrets, source code). Use both.
Part D — DLP for AI Gateway (scan for sensitive data) (Enterprise)
This is the Zero Trust integration: AI Gateway can scan the text of requests and responses against your Zero Trust DLP profiles — the same profiles you built in Module 6 — with no HTTP filtering and no TLS decryption.
- 👉 In the gateway → Features → DLP → Set up.
- 👉 Select the DLP profiles to apply — e.g. Credentials and Secrets, PII, or an AI Prompt profile (PII, Source Code, Financial, jailbreak/intent topics; even custom natural-language prompt topics).
- 👉 Choose what to scan with Check: Request (prompts going to the model), Response (what the model returns), or Both.
- 👉 Set the action — flag (log) or block — on a match.
- 👉 Save.
✅ Checkpoint: a request containing a fake secret / test PII is flagged or blocked, and the detection shows in logs with the matched profile.
⚠️ Streaming caveat (important for latency): if you enable Response scanning, AI Gateway buffers the full provider response before returning it, which increases time-to-first-token — noticeable for chat/coding agents that stream. If you need low-latency streaming, set Check = Request only, or route latency-sensitive traffic through a separate gateway.
💡 Perfect for coding agents: tools like Cursor/Claude Code send source code, config, and secrets to model providers. Because AI Gateway sits in the path, DLP catches secrets/regulated data leaving (or returning) without modifying the agent.
Part E — Cost, rate & privacy controls
AI Gateway also gives you operational guardrails:
| Control |
What it does |
Where |
| Rate limiting |
Cap the number of requests over a window |
Settings → Rate limiting |
| Spend limits |
Cap dollar cost (by model, provider, or custom metadata — e.g. $200/day per user) |
Features → Spend limits |
| Caching |
Serve identical requests from cache — cuts latency + provider cost |
Settings → Cache Responses |
| Fallbacks / retries |
Auto-retry or fall back to another model on provider error |
Configuration → Fallbacks |
| Log payload control |
Send cf-aig-collect-log-payload: false to log metadata only, not prompt/response bodies |
per-request header |
💡 Privacy tip: for regulated data, use cf-aig-collect-log-payload: false so you keep usage metrics (tokens, cost, model, status) without persisting sensitive prompt/response text.
⚠️ Don't enable caching or rate limiting on a gateway shared by an AI Search / RAG instance — it interferes with embedding/indexing. Use a dedicated gateway for that.
Part F — Observe & verify
- 👉 In the gateway, review Analytics (requests, tokens, cost, cache hit rate, errors) and Logs (per-request: model, provider, status, cost, duration, user agent, and — unless disabled — payloads).
- 👉 Filter logs to confirm your Auth, Guardrails, and DLP actions are firing as expected.
- 👉 Test end-to-end: send (a) a normal request → succeeds; (b) a request missing the auth token → rejected; (c) a prompt with unsafe content → guardrail blocks; (d) a prompt with a test secret → DLP flags/blocks.
✅ Checkpoint: all four behave correctly and are visible in logs.
✅ Module 7c complete!
You now have:
- ✅ An AI Gateway in front of your AI providers (single endpoint, multi-provider)
- ✅ Authenticated Gateway — only authorized callers can use it
- ✅ Guardrails blocking unsafe content in prompts/responses
- ✅ DLP scanning prompts/responses for sensitive data using your Zero Trust profiles
- ✅ Cost, rate, cache, and privacy controls
- ✅ Full logging & analytics for every AI call
The complete AI security picture
| Layer |
Governs |
Module |
| Browser AI use |
People pasting into ChatGPT etc. |
7 — AI controls |
| AI agents → tools |
MCP servers & portals |
7b — Secure AI & MCP |
| Your apps → AI models |
Programmatic/API AI |
7c (this) |
| Discover & approve AI apps |
Shadow AI |
5d — Shadow IT |
| Sensitive-data detection |
DLP profiles |
6 — DLP |
Quick troubleshooting
| Problem |
Fix |
| Requests bypass the gateway |
Confirm the app's base URL points at the gateway endpoint (Part A) |
| Anyone can use the gateway |
Enable Authenticated Gateway + send the cf-aig-authorization header (Part B) |
| DLP → Set up unavailable |
DLP for AI Gateway uses Zero Trust DLP — confirm Enterprise + that profiles exist (Module 6) |
| Streaming feels slow after enabling DLP |
Response scanning buffers the full reply — set Check = Request only or use a separate gateway (Part D) |
| Sensitive prompts stored in logs |
Send cf-aig-collect-log-payload: false to log metadata only (Part E) |
| AI Search accuracy dropped |
Don't enable caching/rate limiting on the gateway your RAG instance uses (Part E) |
Connect whole offices and data centers to Cloudflare's network.
Nguồn cộng đồng — không phải tài liệu chính thức của Cloudflare: https://zerotrust.cfsase.workers.dev
Community source — not an official Cloudflare publication: https://zerotrust.cfsase.workers.dev