SRE Lab · failover

Plan B

Pick a made-up incident and ask an AI model for the first three triage steps. Then break the primary model on purpose and watch a backup model, from a different provider, answer instead. Each attempt shows how it ended and how long it took.

1 · Incident (fictional)
2 · Primary model

Up to 10 runs per 10 minutes per visitor · shared daily cap

Result

Results appear here: each attempt, which model answered, and the answer.

How it works

Request path The page posts to a Pages Function, which calls AI Gateway with two steps. Step 0 goes to Workers AI. If it errors or times out, the gateway falls back to step 1, Anthropic. Everything except Anthropic runs on Cloudflare. runs on Cloudflare This pageincident + mode, no keys Pages Function/api/planb · limits · caps AI Gatewaysteps [0, 1] · no cache step 0 Workers AIllama-3.2-3b-instruct error / timeout step 1 Anthropicclaude-haiku-5-5

The page sends only an incident number and a mode. A Pages Function checks the per-visitor limit and the daily caps, then sends one request to Cloudflare AI Gateway with two steps. The gateway tries step 0 (Llama 3.2 3B on Workers AI). If that returns an error or runs past its timeout, the gateway retries the same question on step 1 (Claude Haiku 5.5 on Anthropic). There are no retries on step 0, so the fallback is visible.

The two-step request uses AI Gateway's Universal Endpoint format, which Cloudflare marks as deprecated (as of 10 October 2026); existing integrations keep working. Its recommended replacement for fallbacks, dynamic routing, configures the same kind of fallback in the gateway itself.

What this does and doesn't show

It shows recovery from a failed model request: a bad model id, a provider error or a slow response. The function and the gateway both run on Cloudflare, so if Cloudflare itself had a wide outage, this page, the function and the fallback would all be down together. Plan B is a backup model, not a backup platform.

Surviving a real platform outage takes more:

What changes in the agentic era

  • Model failover is now a reliability problem. A model API fails in its own ways. Claude's API returns 529 overloaded_error under high traffic across all users, a tier spend-cap 429 that has no retry-after and keeps failing, and errors that can arrive mid-stream after a 200. OpenAI likewise says quota 429s shouldn't be retried. Fallback logic has to tell these apart, because retrying a spend cap only burns the timeout.
  • A 200 isn't a good answer. A backup model can answer the same prompt shorter, in a different format or wrong. Agents make this sharper: one task makes many model calls and errors compound across steps, so a silent switch mid-task can change what the agent does. Record which model served each call, make failover a visible event someone reviews, and have the synthetic probes check the answer as well as the status code.
  • Failing over can mean another Region. Amazon Bedrock cross-Region inference routes requests within a geography or globally, and CloudTrail's additionalEventData.inferenceRegion shows where each one ran. Global routing can process requests in any supported commercial Region, so where a request may run is a data-residency decision for people, not just a reliability setting.

Sources

Guardrails and cost

Model names, prices, the free allowance and the Universal Endpoint's deprecation status are as of 10 October 2026. Check the current vendor docs: Claude models, Claude pricing, Workers AI pricing and free allocation, AI Gateway Universal Endpoint.