Why Fallbacks Matter
Provider Downtime
OpenAI outages took down ChatGPT, Cursor, and thousands of apps in 2024
Rate Limiting
Hit your quota during peak hours and block all users
Regional Issues
Provider performance varies by geography and time of day
Cost Optimization
Route to cheaper providers first, with expensive backups
How Fallbacks Work
The gateway automatically retries failed requests with alternative providers:- Tries OpenAI first
- If OpenAI fails → Instantly tries Azure
- If Azure fails → Tries all other available providers
- Returns first successful response
Automatic Failover Triggers
The gateway automatically fails over when it encounters:The gateway intelligently handles errors. Some 400 errors (like invalid model format) won’t trigger fallback since they’d fail on all providers.
Fallback Strategies
Default Automatic Routing
No configuration needed - the gateway automatically finds all providers:Explicit Provider Chain
Specify exact providers and order:Open-Ended Chain
Start with specific providers, then try all others:BYOK with Credit Fallback
Your provider keys automatically fall back to Helicone’s managed keys:Real-World Examples
Production Reliability
Maximize uptime with multi-provider redundancy:Cost Optimization
Use cheaper providers first, with premium backups:Regional Compliance
Prioritize EU providers for GDPR compliance:Mixed Model Fallbacks
Fall back to different models if preferred model fails:Avoid Problematic Providers
Exclude providers experiencing issues:Error Handling
Successful Fallback
When a fallback succeeds, the response includes metadata:- Which provider was tried first
- Why it failed
- Which provider ultimately succeeded
- Total latency including fallback time
All Providers Failed
If all providers in your chain fail, you get a detailed error:Fallback Best Practices
1. Start with Automatic Routing
Let the gateway handle everything:2. Use Open-Ended Chains
Always include a final catch-all:3. Monitor Fallback Rates
Check your Helicone dashboard regularly:- High fallback rates indicate provider issues
- Adjust your provider priorities based on reliability
- Consider excluding consistently failing providers
4. Test Your Fallback Chain
Verify your fallback logic works:5. Combine with Rate Limiting
Use Helicone’s custom rate limits to prevent provider quota exhaustion:Advanced: Model Provider Priority
The gateway uses intelligent provider selection:Priority Factors
- Your provider keys - Always tried first
- Provider cost - Cheaper providers prioritized
- Provider reliability - Historical uptime considered
- Load balancing - Equal providers rotated
Example Priority Order
Formodel: "gpt-4o-mini":
- Your OpenAI key (BYOK)
- Your Azure key (BYOK)
- Helicone’s OpenAI key (cheapest PTB)
- Helicone’s Azure key (PTB)
- Helicone’s Bedrock (PTB)
- Other providers
Fallback vs Rate Limiting
Understand when to use each:
Best practice: Use both together for maximum reliability and control.
Next Steps
Provider Routing
Learn how the gateway routes requests to providers
Custom Rate Limits
Prevent quota exhaustion with custom limits
Error Handling
Advanced error handling and retry strategies
Browse Models
See which providers support your models