Features
Automatic Fallback
Switches to lower-priority providers if the primary provider fails.
Latency-based Fallback
Optionally switches providers when a component stays above its latency budget for several consecutive turns.
Cooldown-based Retry
Implements a cooldown period before retrying a failed provider, preventing immediate repeated failures.
Auto-Recovery
Automatically switches back to a higher-priority provider once it becomes healthy again.
Permanent Disable
Permanently disables a provider after a configured number of failed recovery attempts.
Error-based Fallback
Here is how you can implement error-based fallback providers for STT, LLM, and TTS in your agent’sPipeline. When a provider fails or becomes unavailable, the system switches to the next configured provider.
Configuration Options
Set these on the wrapper for the slot they apply to. Each slot carries its own policy.float
default:"60"
The duration (in seconds) to wait before retrying a failed provider.
int
default:"3"
The maximum number of recovery attempts allowed before a provider is permanently disabled.
Latency-based Fallback
Beyond hard failures, the Fallback Adapter can switch providers when a healthy provider becomes too slow. This is useful for keeping conversations responsive when a provider degrades without erroring out.Latency-based fallback is off by default. Set
latency_threshold_ms on a component to enable it.- Each component measures a relevant latency metric: STT uses
stt_latency, LLM usesllm_ttft(time to first token), and TTS usesttfb(time to first byte). The budget on a wrapper is checked against its own slot’s metric. - A provider is only switched after it stays above the threshold for
consecutive_latency_hitsturns in a row, avoiding switches caused by a single slow turn. - Recovery and cooldown for a latency-disabled provider use the same
temporary_disable_secandpermanent_disable_after_attemptssettings as the error path.
latency_threshold_ms (and optionally consecutive_latency_hits) to the wrapper. Budget each slot on its own: a few hundred milliseconds is generous for STT and TTS and punishing for an LLM’s first token.
Configuration Options
You can configure the latency-based fallback behavior using the following parameters:int
This slot’s latency budget in milliseconds, checked against its own metric (STT
stt_latency, LLM llm_ttft, TTS ttfb). Off by default. Pass a value to enable latency-based fallback.int
default:"3"
The number of consecutive turns that must exceed
latency_threshold_ms before switching providers.References
- Python
- Node JS
Examples
Fallback Adapter
Checkout the full implementation on GitHub