The shared foundation
Configured stages, required facts, tool receipts, Manager, specialists, and Supervisor. Add native web widgets and optional recording. Carrier controls follow the configured phone connection.
Build voice agents with Ultravox, OpenAI, Gemini, Grok, or Smallest AI Hydra. Choose the speaking model, add Supervisor coaching, and keep one workflow for stages, tools, knowledge, and calls.
One Supafone API key for managed speech and supervision. Or bring your own keys for either.
Start with the familiar Supafone agent and the shared Agent Factory workflow. Keep Ultravox’s external TTS and compatible voice profiles.
new UltravoxS2S(client)Think “one API, many models” for live conversation. Speech-to-speech (S2S) takes caller audio in and returns spoken audio. Supafone gives you one place to choose the supported model and build the agent around it.
The S2S router connects your chosen speaking model. The Supervisor harness observes the call and returns quiet guidance. Agent Factory supplies the stages, tools, knowledge, and team. Use the same SupafoneS2S interface in Python or TypeScript.
create · apply · previewFive provider families, one saved workflow. You choose from the supported catalog. Saved model changes apply to future calls; configured native handoffs open a new provider session and keep call state. Access and voices depend on the provider.
Supervision and teamwork built around the call. Supervisor coaches the speaking model. Manager assigns specialists and proposes next steps. Explore the shared workflow ↓
Keep the stages, tools, and specialist team when you change the speaking model. Enable Supervisor coaching alongside it, or use standalone adapters with supported existing stacks.
Build a plan with 3–8 stages across every hosted speaking model. Set the facts and successful tool results each step needs. Supafone checks those requirements before moving on.
Collect the required details and save them as structured facts. A stage can require those fields before the call moves forward.
Required name + request → saved factsManager coordinates the work: choose an allowed specialist or propose the next stage. Specialists reason privately; Supafone checks each proposal against your saved plan. Supervisor remains the conversation coach.
Manager asks the configured booking specialist to review the caller’s request. The specialist reasons privately; your chosen S2S model keeps the voice.
Current stage: booking. Booking specialist is enabled for this stage.
A specialist’s advice is a proposal. It cannot run arbitrary tools or change the call plan.
Supervisor watches available call evidence and offers quiet guidance while the selected model speaks. Explore the same coaching loop across all five hosted families.
The default managed Ultravox agent can receive deferred Supervisor guidance during a call.
Supervisor seesThe caller chose a time. The booking tool is still running.
Guidance to the voice agent“Wait for the booking result before confirming the appointment.”
Supported across all five hosted speaking families. Enable Supervisor and configure its credentials separately from the speaking model. Native guidance is delivered on a tool request, without forcing an interruption.
View the support matrixSame agent. Same model interface. Your connection.
Start with a preview. Connect a managed number or your own carrier when you’re ready.
Preview your saved agent, then embed a native voice widget on your site. The server keeps provider credentials private and carries conversation context into the call.
Explore browser callsChoose the supported model and voice that fit the job. Every family connects to Agent Factory’s shared stages, tools, Manager, specialists, and Supervisor. Ultravox is the default; OpenAI, Gemini, Grok, and Hydra use their native browser and phone paths.
gpt-realtime-2.1 or gpt-live-1 with marin.gemini-3.1-flash-live-preview with Puck.grok-voice-latest with eve.hydra-v1.0 with sterling, or hydra-v1.1 with maya.Configured stages, required facts, tool receipts, Manager, specialists, and Supervisor. Add native web widgets and optional recording. Carrier controls follow the configured phone connection.
Allow exact models and voices, with at most three handoffs per call. Opt into one provider-connection recovery. Each handoff opens a new provider session while keeping the same facts and stage.
External TTS remains Ultravox-only, and Ultravox is outside native handoffs. Hydra is English-only with no native transcripts. Native voicemail detection is not supported.
Give your agent a job, ground it in your knowledge, and connect the tools that get work done. Add stages, a Manager, specialists, and Supervisor coaching, then test in the browser or connect a phone number.
Give the speaking agent a separate coach. Supervisor checks available transcripts, stage progress, and tool results, then proposes guidance within your rules. Choose managed supervision or your own supported reasoning provider independently of the speaking model.
Example guidance “Confirm the booking only after the scheduling tool succeeds.”
Follow the caller’s intent and the agent’s progress. Compare what the agent says with what its tools actually returned.
Send a short directive through the adapter’s supported control channel. Configure evidence thresholds, output limits, and operator guardrails.
Use available call evidence for scoring and QA, then improve the standing directive for future conversations.
Same coach, supported delivery paths. Enable Supervisor across all five hosted families with configured managed or BYOK credentials. Native models request guidance through check_guidance; Ultravox uses deferred delivery. Hydra supplies labeled model-reported context because it has no native transcript stream. Standalone SDK adapters have their own capability matrix.