Plan B
Pick a made-up incident and ask an AI model for the first three triage steps. Then break the primary model on purpose and watch a backup model, from a different provider, answer instead. Each attempt shows how it ended and how long it took.
Result
Results appear here: each attempt, which model answered, and the answer.
How it works
The page sends only an incident number and a mode. A Pages Function checks the per-visitor limit and the daily caps, then sends one request to Cloudflare AI Gateway with two steps. The gateway tries step 0 (Llama 3.2 3B on Workers AI). If that returns an error or runs past its timeout, the gateway retries the same question on step 1 (Claude Haiku 5.5 on Anthropic). There are no retries on step 0, so the fallback is visible.
- Break it: down asks the gateway for a model that doesn't exist, so step 0 fails with a provider error.
- Break it: slow gives step 0 a 1 ms timeout, so it times out before the model can answer.
- Which model answered comes from the gateway's
cf-aig-stepresponse header when it's there, and otherwise from the shape of the answer (Anthropic and Workers AI reply in different formats). The result says which one it used. - The response cache is skipped, so every run is a real request. The API key for the backup stays on the server.
The two-step request uses AI Gateway's Universal Endpoint format, which Cloudflare marks as deprecated (as of 10 October 2026); existing integrations keep working. Its recommended replacement for fallbacks, dynamic routing, configures the same kind of fallback in the gateway itself.
What this does and doesn't show
It shows recovery from a failed model request: a bad model id, a provider error or a slow response. The function and the gateway both run on Cloudflare, so if Cloudflare itself had a wide outage, this page, the function and the fallback would all be down together. Plan B is a backup model, not a backup platform.
Surviving a real platform outage takes more:
- Failover outside the failing platform: a client or SDK that can call a second provider directly, or two independent edges (different CDN or cloud) behind DNS with health checks.
- Health-checked routing: synthetic probes that take a provider or region out of rotation before users notice, and put it back slowly.
- Degraded modes: a cached or rule-based answer, a queue for later, or a clear "AI help is unavailable" state, so the product still works without the model.
- Budgets on the backup: the fallback path gets real traffic only during incidents, so give it its own rate limits and spend caps, and test it regularly the way this page does.
What changes in the agentic era
- Model failover is now a reliability problem. A model API fails in its own ways. Claude's API returns 529
overloaded_errorunder high traffic across all users, a tier spend-cap 429 that has noretry-afterand keeps failing, and errors that can arrive mid-stream after a 200. OpenAI likewise says quota 429s shouldn't be retried. Fallback logic has to tell these apart, because retrying a spend cap only burns the timeout. - A 200 isn't a good answer. A backup model can answer the same prompt shorter, in a different format or wrong. Agents make this sharper: one task makes many model calls and errors compound across steps, so a silent switch mid-task can change what the agent does. Record which model served each call, make failover a visible event someone reviews, and have the synthetic probes check the answer as well as the status code.
- Failing over can mean another Region. Amazon Bedrock cross-Region inference routes requests within a geography or globally, and CloudTrail's
additionalEventData.inferenceRegionshows where each one ran. Global routing can process requests in any supported commercial Region, so where a request may run is a data-residency decision for people, not just a reliability setting.
Sources
Guardrails and cost
- Fixed, fictional incidents and a fixed prompt. Visitors select fixed server-side incidents; arbitrary visitor text is not inserted into the prompt.
- About 10 runs per 10 minutes per visitor (a salted, daily-rotating hash of the IP; raw IPs are never stored).
- Daily caps: about 300 runs and 100 backup calls. When a cap is hit, Plan B rests until midnight UTC.
- Answers are capped at 200 tokens. Workers AI runs inside the free daily allowance; a backup call to Claude Haiku 5.5 costs a small fraction of a cent.
Model names, prices, the free allowance and the Universal Endpoint's deprecation status are as of 10 October 2026. Check the current vendor docs: Claude models, Claude pricing, Workers AI pricing and free allocation, AI Gateway Universal Endpoint.