Cloudflare Workers Cloudflare Workers Cloudflare Workers · Lab 6/7 Lab 6/7 Lab 6/7 · 10 phút · 10 min · 10 នាទី

06

AI Gateway AI Gateway AI Gateway

Đưa lời gọi Workers AI qua AI Gateway để cache, giám sát và kiểm soát chi phí. Route your Workers AI calls through AI Gateway for caching, monitoring, and cost control. A small configuration change with big operational benefits. បញ្ជូនការហៅ Workers AI តាម AI Gateway ដើម្បី cache តាមដាន និងគ្រប់គ្រងថ្លៃ។

Nội dung bước (lệnh, code) giữ nguyên tiếng Anh từ nguồn chính thức. Step body (commands, code) stays in English from the official source. ខ្លឹមសារជំហាន (ពាក្យបញ្ជា និង code) រក្សាភាសាអង់គ្លេសពីប្រភពផ្លូវការ។ labs.cloudflare.dev ↗

Cần trước Prerequisites តម្រូវការជាមុន

  • Đã xong bước 5 Completed Step 5 បានបញ្ចប់ជំហាន 5
  • bookmark-api đã gắn Workers AI bookmark-api with Workers AI bookmark-api បានភ្ជាប់ Workers AI

Bạn sẽ làm được Learning objectives គោលបំណងសិក្សា

  • Đưa inference AI qua gateway để quan sát và kiểm soát chi phí Route AI inference through a managed gateway for observability and cost control បញ្ជូន inference AI តាម gateway ដើម្បីមើល និងគ្រប់គ្រងថ្លៃ
  • Chọn TTL cache theo độ ổn định và độ mới của response Choose cache TTLs based on response stability and freshness requirements ជ្រើស TTL cache តាមស្ថិរភាព និងភាពថ្មីនៃ response
  • Đọc analytics Gateway để thấy chỗ tối ưu Interpret gateway analytics to identify optimization opportunities អាន analytics Gateway ដើម្បីឃើញចំណុចធ្វើឲ្យប្រសើរ

Bước 1: Tạo AI Gateway Step 1: Create the AI Gateway ជំហាន 1: បង្កើត AI Gateway

Chúng ta đang xây What we're building អ្វីដែលយើងកំពុងសង់
An AI Gateway configured in the Cloudflare Dashboard.
Vì sao quan trọng Why this matters ហេតុអ្វីសំខាន់
The gateway sits between your Worker and the AI models, providing observability and caching without changing your application logic.
  1. Go to the Cloudflare Dashboard
  2. Select AI > AI Gateway
  3. Click Create Gateway
  4. Name it bookmark-gateway and click Create

Bước 2: Sửa lời gọi AI Step 2: Update the AI Calls ជំហាន 2: កែការហៅ AI

Chúng ta đang xây What we're building អ្វីដែលយើងកំពុងសង់
All env.AI.run() calls routed through AI Gateway.
Vì sao quan trọng Why this matters ហេតុអ្វីសំខាន់
Adding the gateway option enables caching, monitoring, and rate limiting for every AI request.

The change is small. In src/index.ts, update the generateSummary function that calls env.AI.run():

The change is small. In src/index.ts, update the generateSummary function that calls env.AI.run():

The change is small. In src/index.ts, update the generateSummary function that calls env.AI.run():

Cập nhật generateSummary Update generateSummary ធ្វើបច្ចុប្បន្នភាព generateSummary

typescript
// CHANGED: added gateway option as third argument
async function generateSummary(title: string, url: string, env: Env): Promise<string> {
  try {
    const response = await env.AI.run('@cf/meta/llama-3.1-8b-instruct-fast', {
      messages: [
        {
          role: 'system',
          content: 'You are a helpful assistant that writes concise bookmark descriptions. Respond with exactly one sentence, no more than 20 words.'
        },
        {
          role: 'user',
          content: `Write a one-sentence description for this bookmark:\nTitle: ${title}\nURL: ${url}`
        }
      ]
    }, {
      gateway: {
        id: 'bookmark-gateway',
        skipCache: false,
        cacheTtl: 86400  // Cache summaries for 24 hours
      }
    });

    return response.response?.trim() || '';
  } catch (error) {
    console.error('AI summary failed:', error);
    return '';
  }
}

That is the entire code change. The rest of the file stays exactly the same.

That is the entire code change. The rest of the file stays exactly the same.

That is the entire code change. The rest of the file stays exactly the same.

Bước 3: Test cache Gateway Step 3: Test Gateway Caching ជំហាន 3: Test cache Gateway

Chúng ta đang xây What we're building អ្វីដែលយើងកំពុងសង់
Verified that duplicate AI requests are served from the gateway cache.
Vì sao quan trọng Why this matters ហេតុអ្វីសំខាន់
Cache hits mean faster responses and lower costs.
bash
npx wrangler dev --remote

Test cache Gateway với bookmark Test gateway caching with bookmarks Test cache Gateway ជាមួយ bookmark

Create a bookmark:

Create a bookmark:

Create a bookmark:

bash
curl -X POST http://localhost:8787/bookmarks \
  -H "Content-Type: application/json" \
  -d '{"url":"https://developers.cloudflare.com/workers/","title":"Workers Docs","tags":"docs"}'

Delete it and recreate with the same title and URL:

Delete it and recreate with the same title and URL:

Delete it and recreate with the same title and URL:

bash
curl -X DELETE http://localhost:8787/bookmarks/REPLACE_ID
bash
curl -X POST http://localhost:8787/bookmarks \
  -H "Content-Type: application/json" \
  -d '{"url":"https://developers.cloudflare.com/workers/","title":"Workers Docs","tags":"docs"}'

The second creation sends the same prompt to the AI model. Because the gateway caches by prompt, this request should return faster with a similar summary served from cache.

The second creation sends the same prompt to the AI model. Because the gateway caches by prompt, this request should return faster with a similar summary served from cache.

The second creation sends the same prompt to the AI model. Because the gateway caches by prompt, this request should return faster with a similar summary served from cache.

Bước 4: Xem analytics trên dashboard Step 4: View Analytics in the Dashboard ជំហាន 4: មើល analytics លើ dashboard

Chúng ta đang xây What we're building អ្វីដែលយើងកំពុងសង់
An understanding of the metrics available in the AI Gateway Dashboard.
Vì sao quan trọng Why this matters ហេតុអ្វីសំខាន់
Observability lets you optimize costs, debug issues, and understand usage patterns.

After deploying (npx wrangler deploy), go to AI > AI Gateway > bookmark-gateway in the Dashboard. You will see:

After deploying (npx wrangler deploy), go to AI > AI Gateway > bookmark-gateway in the Dashboard. You will see:

After deploying (npx wrangler deploy), go to AI > AI Gateway > bookmark-gateway in the Dashboard. You will see:

  • Request count - Total AI requests routed through the gateway
  • Cache hit rate - Percentage served from cache (your cost savings)
  • Latency - Average and p99 response times
  • Token usage - Input and output tokens consumed
  • Error rate - Failed AI requests

Cấu hình rate limit Configure rate limiting កំណត់ rate limit

In the gateway settings, you can set request limits per minute to prevent runaway costs from unexpected traffic.

In the gateway settings, you can set request limits per minute to prevent runaway costs from unexpected traffic.

In the gateway settings, you can set request limits per minute to prevent runaway costs from unexpected traffic.