Supafone Labs · Documentation
Documentation menu

Agent Factory#

Agent Factory creates a durable Supafone voice agent from a job description. Its S2S harness connects the selected speaking model to the agent's prompt, supported tools, stages, and browser or phone transport. You can change the model without rebuilding the surrounding native agent integration.

Shared S2S interface#

Use the same SupafoneS2S interface for Ultravox, OpenAI, Gemini, Grok, and Smallest AI Hydra. Provider subclasses supply the selection while Agent Factory keeps the agent and phone identity. Ultravox stays the default; the native catalog offers six model choices. See the common interface for the class contract and switching examples.

Ordinary model updates apply to new calls. Opt-in native session handoff is a separate runtime control. Capabilities remain provider-specific: Hydra has no native transcripts and cannot change its persona or voice mid-session. Check credentials and test the selected provider before a customer call.

Two voice-output choices in one Agent Factory#

Ultravox + custom TTS keeps the existing managed phone agent and lets a compatible Cartesia, ElevenLabs, Inworld, or Ultravox catalog voice supply its speech. Native S2S selects OpenAI, Gemini, Grok, or Hydra and uses that model's own voice list. Both are available through SupafoneS2S; external TTS is not a universal voice override for native models.

Compare both paths and create an Ultravox agent with custom TTS.

Choose an S2S option#

ProviderModelDefault voice
OpenAIgpt-realtime-2.1marin
OpenAIgpt-live-1marin
Googlegemini-3.1-flash-live-previewPuck
xAIgrok-voice-latesteve
Smallest AIhydra-v1.0sterling
Smallest AIhydra-v1.1maya

All six native catalog models support browser previews and the same five phone families: Supafone-managed, Twilio, Telnyx, Plivo, and SIP. Model availability and carrier readiness depend on the deployment's configuration.

Create a native agent#

In the Labs builder, create or open an agent and choose Ultravox, OpenAI, Gemini, Grok or Hydra in Setup. Native model and voice controls appear in place. Speaking model & voice shows the matching voice controls; compatible custom TTS remains on the Ultravox path.

Open Workflow to configure the plan, stage gates, Manager and specialists. Save the agent, then use Test browser voice. Unsaved editor changes must be saved before testing. The Supafone dashboard also supports saved agent model selection. The SDK follows the same Agent Factory contract:

ts
import { Supafone, HydraS2S } from "supafone-labs";

const supafone = new Supafone({ apiKey: process.env.SUPAFONE_TOKEN! });
const engine = new HydraS2S(supafone, { model: "hydra-v1.1", voice: "maya" });
const agent = await engine.create({
  agentKey: "northline-intake",
  name: "Northline intake",
  description: "Understand the request and book the right next step.",
  supervisor: true,
  telephony: { mode: "supafone_managed", provider: "supafone" },
});
const preview = await engine.testCall("northline-intake");

Start with one Supafone API key. The selected model uses a configured platform key by default; an encrypted account BYOK key overrides it. Check the model's runtime status and test a preview before dialing. Keep provider keys on the server. Managed phone and BYO carriers have separate readiness checks.

Keep the coach when you switch#

Set supervisor: true or the managed/BYOK Supervisor object on the agent. All five speaking families support coaching, independently from the voice choice. Native models receive guidance through check_guidance; Hydra passes model-reported context because it has no transcript stream. A supported coach must also be enabled and have its Supervisor credentials configured. See the hosted coaching contract.

Keep the workflow when you switch#

ts
import { OpenAIS2S } from "supafone-labs";

const next = new OpenAIS2S(supafone, { model: "gpt-realtime-2.1", voice: "marin" });
await next.apply("northline-intake");
const preview = await next.testCall("northline-intake");

Apply the selection before previewing: testCall uses the saved agent. The next session uses the new model. The agent identity, instructions, configured tools, team, custom stages and phone configuration remain together. Use a voice supported by the selected model. An active call keeps its frozen workflow even if the saved agent changes.

Manager, specialists and execution gates#

Enable manager: true for bounded reasoning about the next permitted stage or specialist. Configure agent_team with stage-scoped specialists; consultations return advice to the single speaking model. The Manager cannot invent a successful booking or bypass a stage gate. The server validates required saved fields, successful tool receipts, legal next stages and active permissions.

Read Shared runtime, Manager and teams for a complete booking-team example, budgets, recording, carrier controls and opt-in native model handoff.

The native S2S guide contains Python examples, credential status, browser audio, and carrier setup. The developer workflow covers creation and switching.

Keep compatible external voices with Ultravox#

Omitting realtime retains the managed Ultravox runtime. The planner, Manager, stage gates and tools are shared with native S2S. Ultravox additionally retains its compatible external TTS voices and opt-in language/voice profile router. Native providers use their own supported voices. Switching models does not make arbitrary external TTS or voice cloning available on every provider.

Managed compatibility workflow#

Start with the Supafone API key and hide provider keys until the user asks for advanced control.

ts
const supafone = new Supafone({
  apiKey: process.env.SUPAFONE_TOKEN!,
});

const agent = await supafone.labs.agents.createInboundWithNumber({
  agentKey: "northline-intake",
  name: "Northline intake",
  assistantName: "Maya",
  description: "Answer new inquiries, understand the request, and book the right next step.",
  websiteUrl: "https://northline.example",
  number: {
    search: { areaCode: "415" },
    numberStrategy: "default_pool"
  },
  labs: {
    enabled: true,
    mode: "supafone_managed",
    model: "gemma"
  }
});

console.log(agent.call_plan?.summary);
console.log(agent.call_plan?.call_stages); // the exact stages now running

Managed compatibility language and voice routing#

Live routing is an Agent Factory opt-in. It is off by default, so existing agents and manually built product agents keep their current language and voice behavior.

The smallest configuration enables English and Spanish with compatible voices selected from the account's live voice catalog:

ts
const agent = await supafone.labs.agents.createInbound({
  agentKey: "northline-bilingual",
  name: "Northline bilingual intake",
  languageVoiceRouting: true,
});

Use routingLanguages to configure two to four languages. The first language controls the greeting:

ts
const agent = await supafone.labs.agents.createInbound({
  agentKey: "northline-multilingual",
  name: "Northline multilingual intake",
  languageVoiceRouting: true,
  routingLanguages: ["es-MX", "en-US", "vi-VN"],
});

When the caller clearly requests or speaks another configured language, the same call continues with that language's selected voice. The current call stage, collected facts, campaign context, and available tools remain active. The server never routes from accent alone.

Voice selection is automatic by default. Advanced applications can provide a current catalog voice for each language:

ts
const agent = await supafone.labs.agents.createInbound({
  agentKey: "northline-curated",
  name: "Northline curated multilingual intake",
  languageVoiceRouting: true,
  languageProfiles: [
    { language: "en-US", voice: { provider: "cartesia", voiceId: "<catalog-voice-id>" } },
    { language: "es-MX", voice: { provider: "cartesia", voiceId: "<catalog-voice-id>" } },
  ],
});

Only the preference contract is published in the SDK. Detection policy, provider resolution, live call-state transitions, and telephony implementation remain server-side Supafone infrastructure.

The first configured language owns the opening. If it is not English, Agent Factory translates the supplied or generated greeting during provisioning and returns translation status with the resolved profiles. See Live Language and Voice Routing for REST, Python, TypeScript, MCP, PSTN, campaign, WebRTC, and troubleshooting details.

Python:

python
agent = supafone.labs.agents.create_inbound_with_number({
    "agentKey": "northline-intake",
    "name": "Northline intake",
    "assistantName": "Maya",
    "description": "Answer new inquiries, understand the request, and book the right next step.",
    "websiteUrl": "https://northline.example",
    "number": {
        "search": {"areaCode": "415"},
        "numberStrategy": "default_pool",
    },
    "labs": {
        "enabled": True,
        "mode": "supafone_managed",
        "model": "gemma",
    },
})

print(agent["call_plan"]["summary"])

Python uses the same opt-in:

python
agent = supafone.labs.agents.create_inbound({
    "agentKey": "northline-bilingual",
    "name": "Northline bilingual intake",
    "languageVoiceRouting": True,
    "routingLanguages": ["en-US", "es-MX"],
})

Outbound Agents#

Outbound is a first-class direction, not an inbound hack.

ts
const outbound = await supafone.labs.agents.createOutboundWithNumber({
  agentKey: "northline-speed-to-lead",
  name: "Northline speed to lead",
  assistantName: "Maya",
  goal: "Call new leads within five minutes and book a consult.",
  description: "Call warm, consented leads, understand fit and urgency, then book a consult without pressure.",
  number: { search: { areaCode: "415" } },
  labs: { enabled: true, mode: "supafone_managed", model: "gemma" },
});

Labs builder controls#

The Labs workspace edits the S2S selection and shared workflow together. Changing providers keeps the stages, team and Manager settings; save to apply the selection to future calls.

SectionControls
SetupAgent job, direction, speaking provider, native model and voice
Speaking model & voiceNative voice summary, or compatible Ultravox TTS catalog and advanced voice controls
WorkflowManaged or custom 3–8 stage plan, capture fields, stage rules, Manager, specialist team and timezone
ToolsConfigured business tools, with server-side stage and role permissions
RecordingOpt-in audio recording and optional post-call transcription
Test browser voiceSaved-agent microphone test, connection status, available transcripts and genuine Supervisor events
ExportTypeScript, Python, REST and JSON configuration

Workflow, Manager and specialists#

Choose Use the managed plan and a stage count, or Customize call stages to edit named stages. Custom rows expose goals and instructions, allowed business tools, required saved fields, required successful tool receipts, allowed next stages and an optional specialist ID. Define the captured field names in Fields this agent can save before using them as requirements.

For example, a booking stage can require a saved email and a successful book_appointment receipt before moving to confirmation. The server validates that evidence; neither the speaking model nor Manager can substitute a claim of success.

Default tool permissions use the configured tools. Choosing an explicit empty list permits no business tools. Default progression follows the next stage in order; choosing an explicit empty next-stage list prevents further transitions. Stage and specialist references must remain valid when renaming or removing rows.

Enable Manager allows separate reasoning about the next permitted stage or specialist. The builder defaults to 12 tasks per call, 2 parallel tasks and a 6-second timeout, with bounds of 1–30, 1–4 and 1–15 seconds respectively. Failed and timed-out work still uses the call budget.

Enable specialist team adds IDs, role instructions, stage scopes and tool restrictions. A blank stage scope allows all configured stages; choose a custom plan to use named scopes. Specialists cannot widen stage permissions. With Manager enabled, consultations run private reasoning tasks and return advice to the selected S2S speaker.

Manager coordinates permitted work; Supervisor coaches the speaker. The server's shared runtime controls stage progression and tool execution. These roles do not replace the speaking model or grant it additional permissions.

Use the current date and time supplies server time. An empty Timezone inherits the account setting; an explicit value must be an IANA timezone.

Browser test and supervision feed#

Save or create the agent before choosing Start browser call. The test uses the saved configuration and refuses to start a new test with unsaved changes. OpenAI, Gemini, Grok and Hydra use native PCM audio over a Supafone WebSocket; Ultravox uses its existing WebRTC client. Credentials remain on the server.

The transcript panel displays actual provider transcripts where supported. Hydra is English-only and has no native live transcript stream. Its Supervisor context is model-reported; optional transcription of a recording is produced after the call.

The Supafone Supervisor feed shows real server events and distinguishes observations from delivery acknowledgements. A display toggle only shows or hides the feed. An enabled Supervisor setting still requires configured reasoning credentials, and a connected voice session may have no guidance events yet.

Recording settings#

Record calls and Transcribe recordings are opt-in. On native S2S, post-call transcription requires recording and a configured Supafone Deepgram key. Live provider transcripts are separate from post-call artifacts.

json
{"recording": {"enabled": true, "transcribe": true, "max_duration_seconds": 900}}

The recording object accepts enabled, transcribe and an optional max_duration_seconds of 1–1800. The builder changes the two toggles and preserves an existing duration cap set through the API or SDK. This cap limits recorded audio rather than ending the call. These fields do not provide retention deletion, PII redaction or a consent announcement; retention_days is unsupported.

Credentials and capability boundaries#

Start with a Supafone API key. Configured platform provider keys are used by default, while an encrypted account BYOK key takes priority. Keep speaking provider, Supervisor reasoning, carrier and external TTS credentials separate. The Supervisor adapter catalog describes integrations for agents outside the hosted builder; it does not add extra hosted speaking models or carriers.

Native providers use their own voice lists. Ultravox keeps compatible external TTS and live language/voice routing. Configured carrier transfers, DTMF and media pause have carrier-specific requirements. Native voicemail detection is not supported. An ordinary model selection changes new calls; in-call native handoff needs a separate exact allowlist and opens a replacement provider session. See Shared runtime, Manager and teams.

Exported Code#

Export TypeScript:

ts
import { Supafone } from "supafone-labs";

const supafone = new Supafone({
  apiKey: process.env.SUPAFONE_TOKEN!,
});

await supafone.labs.agents.createInboundWithNumber({
  agentKey: "northline-intake",
  name: "Northline intake",
  assistantName: "Maya",
  websiteUrl: "https://northline.example",
  number: { search: { areaCode: "415" }, numberStrategy: "default_pool" },
  labs: { enabled: true, mode: "supafone_managed", model: "gemma" },
});

Export Python:

python
from supafone_labs import Supafone

supafone = Supafone(api_key=os.environ["SUPAFONE_TOKEN"])

supafone.labs.agents.create_inbound_with_number({
    "agentKey": "northline-intake",
    "name": "Northline intake",
    "assistantName": "Maya",
    "websiteUrl": "https://northline.example",
    "number": {"search": {"areaCode": "415"}, "numberStrategy": "default_pool"},
    "labs": {"enabled": True, "mode": "supafone_managed", "model": "gemma"},
})

Export JSON for replay/debugging:

json
{
  "agentKey": "northline-intake",
  "name": "Northline intake",
  "assistantName": "Maya",
  "description": "Answer new inquiries, understand the request, and book the right next step.",
  "websiteUrl": "https://northline.example",
  "number": { "search": { "areaCode": "415" }, "numberStrategy": "default_pool" },
  "labs": { "enabled": true, "mode": "supafone_managed", "model": "gemma" }
}

Convenience Defaults#

The Agent Factory should infer as much as possible:

Public API, Not an SDK-Only Feature#

The SDKs are typed conveniences over a normal authenticated REST contract:

http
POST https://api.supafone.ai/api/v1/labs/agent-plans
POST https://api.supafone.ai/api/v1/labs/agents
Authorization: Bearer $SUPAFONE_TOKEN

The first endpoint lets a product show the complete plan for review. The second creates the agent and installs that plan into the real stage runtime. The Python SDK, TypeScript SDK, and MCP server call these same endpoints, so a developer can choose the interface that fits their stack without losing capabilities.

Fixed Language and Voice Intent#

Agent Factory can resolve a current provider voice from plain-language intent:

bash
curl "https://api.supafone.ai/api/v1/labs/agents" \
  -H "Authorization: Bearer $SUPAFONE_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"name":"Spanish intake","preferredLanguage":"es-MX","voicePreference":{"description":"warm Latin American Spanish intake voice","configuredOnly":true}}'

The published SDK 0.7.0 also accepts preferredLanguage and voicePreference (with snake-case aliases) on managed agent creation. You can instead supply an explicit compatible voice selection as shown in the custom TTS guide. These preferences apply to the managed Ultravox voice path; native S2S providers use their own model voices.

preferredLanguage applies one validated language and compatible voice for the entire call. It does not add a language-switch tool or change voices mid-call. See Dynamic Voice Catalog and Selection.

Advanced BYOK Agent Factory#

Configure runtime, telephony, and TTS credentials in separate lanes. A provider key does not add a hosted runtime bridge: use only combinations supported by the selected runtime. The example below uses the Ultravox path with Cartesia TTS:

ts
await supafone.labs.agents.createOutbound({
  agentKey: "speed-to-lead-byok",
  name: "Speed to lead BYOK",
  goal: "Call new leads quickly, qualify fit, and book the next step.",
  labs: {
    enabled: true,
    mode: "byok",
    managedInfrastructure: false,
    llm: { provider: "openai", model: "gpt-4.1-mini" },
    stt: { provider: "deepgram", model: "nova-3" },
    tts: { provider: "cartesia", voiceId: "sonic-warm" }
  },
  byok: {
    agentProvider: {
      provider: "ultravox",
      apiKey: process.env.ULTRAVOX_API_KEY
    },
    telephony: {
      mode: "byok",
      provider: "telnyx",
      credentials: {
        apiKey: process.env.TELNYX_API_KEY,
        connectionId: process.env.TELNYX_CONNECTION_ID,
        fromNumber: "+14155550123"
      }
    },
    tts: {
      provider: "cartesia",
      apiKey: process.env.CARTESIA_API_KEY
    }
  }
});

Custom SIP trunks are first-class pass-through config:

ts
await supafone.labs.agents.createInbound({
  agentKey: "custom-sip-frontdesk",
  name: "Custom SIP front desk",
  telephony: {
    mode: "byok",
    provider: "sip",
    customSip: {
      sipTrunkUri: process.env.SIP_TRUNK_URI,
      username: process.env.SIP_USERNAME,
      password: process.env.SIP_PASSWORD
    }
  },
  ultravox: {
    customSip: {
      sipTrunkUri: process.env.SIP_TRUNK_URI
    }
  },
  labs: { enabled: true, mode: "supafone_managed" }
});

View raw Markdown