Lộ trình liên quan Related learning track Related learning track
Developer Platform Developer Platform Developer Platform
Build ứng dụng AI trên Cloudflare Build AI applications on Cloudflare Build AI applications on Cloudflare
Bạn muốn thêm inference, RAG hoặc gateway tới nhiều model mà không tự vận hành GPU cluster. You want inference, RAG, or multi-model gateways without operating your own GPU clusters. You want inference, RAG, or multi-model gateways without operating your own GPU clusters.
Kiến trúc gợi ý Suggested architecture Suggested architecture
User → Worker/Pages → AI Gateway / Workers AI → Vectorize + R2/KV User → Worker/Pages → AI Gateway / Workers AI → Vectorize + R2/KV User → Worker/Pages → AI Gateway / Workers AI → Vectorize + R2/KV
Sơ đồ tham chiếu (Cloudflare Docs) Reference diagrams (Cloudflare Docs) Reference diagrams (Cloudflare Docs)
Retrieval Augmented Generation (RAG) Retrieval Augmented Generation (RAG) Retrieval Augmented Generation (RAG)
RAG kết hợp retrieval (Vectorize/KV) với Workers AI để chatbot trả lời chính xác hơn — seeding knowledge và query path tách biệt. RAG combines retrieval with generative models for better text. It uses external knowledge to create factual, relevant responses, improving coherence and accuracy in NLP tasks like chatbots. RAG combines retrieval with generative models for better text. It uses external knowledge to create factual, relevant responses, improving coherence and accuracy in NLP tasks like chatbots.
Thuật ngữ: Concepts: Concepts: RAG · Vectorize · Workers AI · Knowledge seeding · Embeddings
Sơ đồ chính thức ↗ Official diagram ↗ Official diagram ↗ · AI Artificial Intelligence (AI) Artificial Intelligence (AI)
Quan sát và kiểm soát AI đa nhà cung cấp Multi-vendor AI observability and control Multi-vendor AI observability and control
Đưa rate limiting, cache và xử lý lỗi lên lớp proxy để áp cấu hình thống nhất cho nhiều dịch vụ và nhà cung cấp inference. By shifting features such as rate limiting, caching, and error handling to the proxy layer, organizations can apply unified configurations across services and inference service providers. By shifting features such as rate limiting, caching, and error handling to the proxy layer, organizations can apply unified configurations across services and inference service providers.
Sơ đồ chính thức ↗ Official diagram ↗ Official diagram ↗ · AI Artificial Intelligence (AI) Artificial Intelligence (AI)
Controls & stack Controls & stack Controls & stack
- Workers AI cho inference tại edge Workers AI for edge inference Workers AI for edge inference
- AI Gateway: routing, cache, observability tới LLM providers AI Gateway: routing, caching, observability to LLM providers AI Gateway: routing, caching, observability to LLM providers
- Vectorize cho RAG embeddings Vectorize for RAG embeddings Vectorize for RAG embeddings
- Durable Objects cho session/stateful chat Durable Objects for session/stateful chat Durable Objects for session/stateful chat
- R2/KV cho documents & config R2/KV for documents and configuration R2/KV for documents and configuration
Tình huống khác (cùng lộ trình) Other scenarios (same track) Other scenarios (same track)
- Build serverless app trên Cloudflare Build a serverless app on Cloudflare Build a serverless app on Cloudflare
- Deploy static site với Pages Deploy a static site with Pages Deploy a static site with Pages
- Build nền tảng SaaS multi-tenant Build a multi-tenant SaaS platform Build a multi-tenant SaaS platform
← Tất cả tình huống lộ trình này ← All scenarios in this track ← All scenarios in this track · Ba nhóm tình huống All three groups All three groups
Next step Next step Next step
Tiếp tục hành trình học của bạn. Continue your learning journey. Continue your learning journey.