The question usually shows up as one line in a planning doc: “add a phone or voice interface.” Someone opens Vapi, LiveKit, and a couple of finished products in browser tabs, and the comparison starts. Which one has better telephony? Which one is cheaper per minute? Which one has the nicer SDK? Those are reasonable questions, but they compare products that do not do the same thing.
I have been on both sides of this. Tough Tongue AI is built on the same kind of real-time voice infrastructure, and most of our engineering time goes into the layer that sits on top of it. So this post is less a pitch than a map of where the work actually is, and a way to decide which side of the line your project belongs on.
Two kinds of products that both say “voice agent”
The first kind is infrastructure. LiveKit and Pipecat are open-source frameworks: you write the agent in code, pick a speech-to-text model, a language model, and a voice, and run it on your servers or their cloud. Vapi, Retell, and Bland are hosted platforms one step up. You configure providers, prompts, and call flows through an API or a dashboard, and they run the real-time plumbing for you. In both cases, you are building an agent. The platform handles the audio, and you handle everything the agent is supposed to accomplish.
The second kind is a finished agent. You start from the job, not the pipeline: call every new enquiry within minutes and book a counselling slot, answer the calls that come in after hours, run a product demo on a Zoom call with slides, screen a hundred candidates on the same rubric, or give a sales team a skeptical buyer to practice on. The conversation logic, tools, scoring, and deployment already exist. The person who knows the job describes it, and the agent does it. Tough Tongue AI is in this group.
This matters because the two groups fail in different ways. Choosing infrastructure for a job that a finished agent already does means months spent rebuilding things that already exist. Choosing a finished agent for a product that needs its own voice layer means working against a tool that was never meant to be your core. Neither is a quality problem. It is a category mistake, and it is easy to make when every homepage says “voice agent.”
What “build” means after the demo
The first day on any of these platforms is impressive. LiveKit says you can have a Python or Node.js voice agent running in under ten minutes, and that is roughly right. Vapi and Retell get you to a talking agent from a dashboard. You hear it answer a question, and it feels like the project is mostly done.
It is not, and the gap is easiest to see with an example. Here is a first version of a requirement that sounds finished:
The agent answers questions about our summer program.
Here is what the admissions team actually needs:
Call every new enquiry within five minutes of the form fill. Ask about start date, budget, and the student’s grade. If they fit, book a counselling slot on the right counsellor’s calendar. If nobody picks up, leave a short voicemail and try again tomorrow. Write a summary to the CRM, flag anyone who asked about scholarships, and let the counselling lead change the pitch without filing a ticket.
The first version is a prompt. The second is a product. Every clause in it is a piece of software that someone has to design, build, test, and keep working as models and providers change.
None of this is a knock on the infrastructure. LiveKit ships a test framework with LLM judges, Retell has simulation testing and call analytics, and Vapi has multi-agent squads and a flow builder. These are serious tools, and they get better every month. But they are tools for building an agent, and the agent is still yours to design, evaluate, and maintain.
The one question that decides it
The comparison I see most often is “which one has better telephony?” It is a fair question, but it rarely decides anything. A better one is: is voice the product, or is the conversation the job?
Take two teams. The first builds a field-service dispatch app. They want technicians to call in and update a job’s status, check tomorrow’s schedule, and reserve parts, all by voice. Every one of those actions has to go through the app’s own business rules: the status machine, the scheduling conflicts, the stock counts. The team has engineers, a TypeScript monorepo, and tests they trust. For them, voice is a new front end for their own product. They should build it, on LiveKit or Vapi or Pipecat, so the agent’s tools call their own service layer and live in their own repository.
The second team runs admissions for an edtech company. They have a steady flow of enquiries, not enough counsellors, and leads that go cold overnight. The person who knows what a good first call sounds like is the admissions lead, not an engineer. For them, the conversation is a job a person used to do. Building it on infrastructure would mean an engineering project to reproduce what a finished agent already does, and then a ticket queue every time the pitch changes.
A useful way to ask the same question: who will change this agent next Tuesday, and how? If the answer is “an engineer, through a pull request,” build. If it is “the person who owns the conversation, the same afternoon,” buy.
When building on Vapi, LiveKit, or Pipecat is the right call
Build when voice is part of what you sell. If the voice experience is part of your product’s value, you want to own it: the latency, the models, the edge cases, the roadmap. A finished agent will always be someone else’s product in the middle of yours.
Build when the agent’s actions have to go through your own business rules. If a call can change an order, move a job, or touch money, the logic belongs next to the rest of your code, with the same tests and the same review process. Infrastructure lets the agent’s tools call your service layer directly.
Build when you need to run and change the stack yourself. LiveKit’s server and agents framework are open source, and so is Pipecat. If your security review, your customers’ contracts, or data residency rules require the whole stack on your own servers, with your engineers able to modify any part of it, open-source infrastructure is the natural fit. Tough Tongue AI is hosted by default. Air-gapped deployment is available on Enterprise plans, but it is still a finished product, not a library you import into your own code.
Build when you need to control every component. Some teams need a specific speech model for a dialect, their own fine-tuned LLM, or custom turn-taking. Infrastructure is where that level of control lives.
Here is how the infrastructure options compare on what they publish. These are snapshots from each vendor’s pricing page in September 2026, and they change often.
| Option | What it is | Published pricing | How you build on it |
|---|---|---|---|
| LiveKit | Open-source agents framework (Apache 2.0) for Python and Node.js, plus LiveKit Cloud | Free plan includes 1,000 agent session minutes; paid plans charge $0.01 per minute after included minutes; models and telephony extra | Write the agent in code; run on LiveKit Cloud or your own servers |
| Pipecat | Open-source Python framework from Daily (BSD 2-Clause) | Free to use; you pay for hosting, models, and telephony | Write the agent in code and host it yourself or on a managed service |
| Vapi | Hosted voice agent API | About $0.05 per minute platform fee; speech, model, voice, and telephony billed separately | Configure through the API or dashboard; tools call your servers |
| Retell | Hosted voice agent platform with a visual flow builder | $0.07 to $0.31 per minute, depending on model and voice | Configure flows in the dashboard or API; functions call your servers |
| Bland | Hosted phone agent platform | $0.14 per minute on Start with no platform fee; $0.12 per minute plus $299 a month on Build; model, voice, and telephony included | Configure through the API or dashboard |
Per-minute prices are the easy part to compare, and they are not where most of the cost goes. On every row above, you still design the conversation, build the tools, decide how to evaluate calls, and build whatever interface your team uses to review them.
When buying a finished agent is the right call
Buy when the job is a known conversation. Answering the calls that come in, calling warm leads, qualifying them, demoing a product, following up, booking a meeting. Screening candidates. Running practice calls for a sales team. These conversations have a shape that many businesses share, even if every business’s pitch is different. A finished agent already has that shape built in, and you add your pitch, your rules, and your material.
Buy when the person who knows the job is not an engineer. This is the part that I think gets underestimated. The admissions lead knows which objection comes up in week three of the season. The sales manager knows the demo is running long. If every change goes through an engineer, the agent falls behind the business. In Tough Tongue AI, the person who owns the conversation writes and edits the agent directly: the brief, the knowledge base, the tools it can use, and the rubric it is scored against.
Buy when you need more than a voice on a phone line. A good sales call often turns into a demo. A finished agent can join a Google Meet or Zoom call, share slides, walk through a live product in a browser, and book the meeting at the end. On infrastructure, each of those is a separate integration to build.
Buy when you need to know how every call went. Transcripts are easy. Deciding whether a call was good is harder. Tough Tongue AI scores each session against a rubric you write, pulls out the fields you care about (start date, budget, objections raised), and sends them wherever your team works.
Buying does not mean giving up code
The most common misread of finished agents is that they are closed, no-code tools with nothing for an engineer to hold on to. That was true of a lot of early products. It is not how Tough Tongue AI works, and it matters for the build-or-buy decision, because many teams want the finished agent and still need to wire it into their own systems.
Here is what an engineer gets:
- Phone calls through your own carrier. Connect a SIP trunk from Twilio, Telnyx, Vonage, Plivo, Exotel, or a similar provider. Agents answer inbound calls on your numbers, and you place outbound calls, single or in batches, from the API. Answering machine detection can hang up or leave a voicemail you write.
- Meeting bots. Send an agent into a Google Meet or Zoom call from the API, with a name and a start time.
- Embedding. Put an agent inside your app as an iframe, in full, basic, or minimal form. It sends lifecycle events to your page, so your app knows when a session starts and ends.
- Your code, before and after the call. A custom function can run before a phone session starts, and another receives the finished session with its analysis, extracted fields, and recording link. The agent can also call MCP servers you connect during the conversation.
- Model choice. Pick the language model, speech-to-text, and voice providers for an agent, and bring your own API keys on the Pro plan and above.
- Versions. Save versions of an agent and pin a specific version to a call, so a change to the script does not surprise a campaign that is already running.
Placing a call is one request to https://api.toughtongueai.com/api/public:
curl -X POST https://api.toughtongueai.com/api/public/v2/sip/call \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"scenario_id": "SCENARIO_ID",
"phone_number": "+14155551234",
"user_name": "Priya",
"dynamic_vars": { "program": "Summer Coding Camp" }
}'
The dynamic_vars fill placeholders in the agent’s instructions, so one agent can call a thousand leads with each one’s own details. When the call ends, the transcript, the rubric scores, and the extracted fields come back through the API or a webhook, ready to write into your CRM. The iframe events guide and the Meet and Zoom bot guide go deeper on the other two surfaces.
Run it from Claude, ChatGPT, Codex, or Cursor
If you already work inside an AI coding assistant, you can do most of this without opening a dashboard. Tough Tongue AI runs a hosted MCP server with tools for creating agents, placing calls, scheduling meeting bots, and pulling session results. You add one URL and sign in through your browser.
claude mcp add --transport http ttai https://api.toughtongueai.com/api/public/mcpcodex mcp add ttai --url https://api.toughtongueai.com/api/public/mcp
codex mcp login ttaiIn ChatGPT, add the official ToughTongue AI app. For Cursor and other coding agents, npx plugins add tough-tongue/toughtongue-skills installs the MCP server along with skills for writing agents and analyzing sessions. Then ask in plain language:
Create an outbound agent that calls new enquiries for our summer coding camp,
qualifies them on start date, budget, and grade, and books a counselling slot.
Score each call on whether it booked a slot and handled the price question.
Then call these five leads tomorrow at 10 am.The next day, “Which leads booked, and what objections came up most?” pulls the sessions and their scores. For the full setup in each client, see the agents page.
Using both
The build-or-buy question is not always either-or. A lot of teams end up with both, and I think that is often the right answer.
Go back to the dispatch team. They should build the technician voice line on LiveKit or Vapi, because it runs on their own business rules. But the same company has other conversations that are not part of the product at all. The sales team demos the software to new field-service companies. The hiring team screens dispatchers. The support team takes the angry calls that the voice line escalates, and new support hires need practice before they take them. None of those need custom infrastructure. They need a finished agent that someone on each team can set up.
A short checklist before you pick
If you are still unsure, these five questions usually settle it:
- Is voice part of what you sell, or a way to get a job done? Part of what you sell points to building.
- Where do the agent’s actions have to run? If every action goes through your own business rules and tests, build. If the actions are booking, CRM updates, and handoffs, buy.
- Who will change the agent, and how often? Weekly changes by a non-engineer point to buying.
- Do your engineers need to run and modify the voice stack itself? If yes, build on open-source infrastructure.
- How will you know if a call went well? If you do not already have an answer, count the evaluation work as part of the build.
“Hi, this is Paris from admissions. Is now a good time to talk about the summer program?”
Start callMost teams can tell in one call whether they still need to build.
- Describe the call: who you are calling and what a good outcome looks like.
- Talk to it in the browser, then point it at your SIP trunk or a Meet link.
- Read the transcript, the rubric scores, and the fields it pulled out.
The layer you actually care about
Voice infrastructure has become very good, very quickly. LiveKit, Pipecat, Vapi, Retell, and Bland have made the audio part of a voice agent close to a solved problem, and that is a big part of why products like ours can exist at all. What has not become easy is the layer above: deciding what the agent should say, giving it the right tools, knowing whether each call went well, and letting the right person change it.
So I would not start by asking which platform has the best telephony. I would start by asking which layer you want to own. If it is the voice itself, build on infrastructure, and pick among the options above by how much control you need. If it is the conversation, try a finished agent first. You can hear one do the job in Tough Tongue AI in a few minutes, and if it is not the right fit, you will know exactly what you need to build. If you want to talk it through, book 15 minutes with me.