Speech to text API metaphor of a traveler hearing a songbird amid wind-rustled winter birches

The best speech to text API for a voice agent is the one that hears the important part of a request and gives your application a clear point at which to act. A fast caption is useful; a confidently wrong appointment time is expensive.

I compared the current documentation for seven options, focusing on streaming behavior, turn control, billing, and the integration work a small team inherits. This is a documentation-based comparison checked October 5, 2026, with a proposed pilot you can run on your own calls.

My recommendation is to shortlist by architecture first. Then test names, corrections, interruptions, and silence before you let a headline accuracy score choose the provider.

Speech to text APIs: the quick comparison

These products solve different parts of the listening problem. The table gives each a starting use case, rather than pretending there is one defensible ranking for every language and call type.

APIStart here whenLive integration questionPublic pricing reference
Deepgram FluxTurn-taking is the main engineering problemHow should end-of-turn events drive your agent?Flux English PAYG: current promotional $0.0065/min; multilingual $0.0078/min. Pricing
AssemblyAIYou want model choice and contextual transcriptionHow long will each billable socket remain open?Universal-3.6 Pro Realtime PAYG: $0.45/session hour, before add-ons. Pricing
ElevenLabs ScribeYou want an explicit partial-to-commit workflowWill your app commit manually or through VAD?Scribe v2 Realtime listed usage rate: $0.39/hour; verify plan allowances. Pricing
Mistral VoxtralYou need a hosted-to-self-hosted pathCan you operate the model and handle speaker attribution separately?Realtime hosted API: $0.006/input audio minute; hosting costs separate. Model and price
Google CloudYour application already lives in Google CloudDoes your language, region, and streaming feature combination fit?V2 Standard recognition: $0.016/min in the first monthly volume tier. Pricing
Microsoft AzureYou need Azure Speech deployment and customizationWhich region, recognizer, and paid features apply?Use the selected-region calculator; no verified numeric rate quoted here. Pricing
OpenAIYou want the current live transcription session interfaceWhere will your application decide the conversational turn?gpt-live-transcribe: estimated $0.017/min, rather than a fixed-rate guarantee. Pricing

Prices are USD references checked on the date above, not complete voice-agent budgets. Session time, processed audio, subscription allowances, and estimated model usage remain separate meters; telephony, reasoning, speech output, and storage are additional purchasing decisions.

If you want a ready-made phone agent rather than the recognition layer underneath one, start with our comparison of AI voice-agent platforms. The right purchase can save you an integration project.

Choose a speech to text API by the moment it becomes useful

I would separate three questions: what did the caller say, have they finished this thought, and may the application take the requested action? A transcript provider can help with the first two; your product still owns the third.

Consider: “Book Tuesday ... actually, Thursday afternoon.” An early partial can start a calendar lookup, but it should not book Tuesday while the caller is still correcting themselves.

Partial text is a working hypothesis. Final text is a stronger record of a segment, but even a final segment can precede a correction in the next turn. Confirmation rules belong around consequential actions.

That means the useful comparison is not just word error rate. Ask whether the system preserves the date, product name, negation, account identifier, and correction that determine the next step.

A transcript that misses a filler word might be harmless. A transcript that drops “don't” can reverse the instruction even while its average accuracy looks excellent.

Artificial Analysis's streaming benchmark usefully pairs word error rate with latency and distinguishes partial from final output after detected speech ends. Its results belong to its datasets, endpoint detector, and model versions; they do not establish your call-center winner.

I would use that benchmark to assemble candidates, then choose with a replayable sample from the actual workflow. Keep the agent prompt, response voice, network location, and decision rules fixed while you compare recognition.

The transcript-to-action contract I would write first

Before choosing a vendor, define the events your application needs: a new utterance, a provisional transcript, a committed segment, a correction, a turn boundary, and a connection failure. Those are product responsibilities, even when one SDK packages them neatly.

Give each utterance an identifier and preserve transcript revisions. Store the model and configuration with the event so a later investigation can distinguish a bad recognizer from a changed endpointing setting.

Allow provisional text to prepare reversible work, such as searching a knowledge base. Require the appropriate committed state and user confirmation before booking, sending, charging, or changing a record.

When an interruption arrives, cancel the response that is still playing and reconcile what the caller actually requested. A cancellation flag that reaches the language model but never reaches audio playback will still feel broken.

This is the same design principle behind keeping people in the approval loop for consequential actions. The voice interface changes the input speed; it does not remove the need for a trustworthy decision boundary.

Once that contract is explicit, a vendor switch becomes a translation between event formats. Without it, every provider's partial transcript can accidentally become a new instruction to your business system.

1. Deepgram: shortlist Flux for conversational turn control

I would shortlist Deepgram when the hardest part is deciding when a caller has finished, rather than simply generating a transcript. Flux gives you a specific conversational interface to evaluate against that problem.

Deepgram official speech-to-text product visual

The Flux documentation describes configurable end-of-turn events, eager end-of-turn behavior, and controls for forcing a turn boundary. Those are useful integration choices, not a guarantee that every pause in your calls will be interpreted correctly.

My first test would include a caller looking up an order number, a long hesitation before a correction, and a short acknowledgment while the agent speaks. Decide whether each event should launch a response, hold the turn open, or stop playback.

Keep Flux and Nova decisions separate. A feature documented for one recognition model should not be assumed to work with every other model simply because the API key is the same.

Current PAYG pricing lists Flux English at a temporary $0.0065/min, versus a regular $0.0077/min, and Flux multilingual at $0.0078/min. Check model-specific add-ons before treating the base rate as the invoice.

The tradeoff: more useful turn controls create more settings to tune. Save the settings alongside each test result, or you will be comparing configurations without realizing it.

2. AssemblyAI: compare the flagship and value tiers explicitly

AssemblyAI is worth a shortlist when language coverage, conversational context, and terminology matter more than finding the simplest base-rate column. I would decide which model family fits before comparing its cost.

AssemblyAI official company and speech recognition visual

The current model-selection documentation identifies Universal-3.6 Pro as the streaming flagship and default, with native code switching and partial transcripts. Pin the model parameter instead of relying on a default that can change.

For a client who regularly mixes a product name in English with a request in another language, I would compare that flagship with the appropriate value model on the same recordings. A generic multilingual label is not enough to settle the choice.

PAYG pricing lists the flagship at $0.45/hr and the value tier at $0.15/hr; flagship keyterms are included, while diarization adds $0.12/hr and prompting adds $0.05/hr. Verify the feature combination you actually enable.

The billing detail I would put in the implementation ticket is session duration: an open socket accrues time even while idle. Explicit termination is part of cost control.

The tradeoff: a conversational socket can be convenient to keep open, but abandoned connections and duplicate streams can undo an attractive hourly rate. Test cleanup as carefully as transcription.

3. ElevenLabs: make transcript commits an application decision

I would include ElevenLabs when the team wants a clear distinction between text that is appearing and text that has been committed. That is particularly useful for an interface showing live text while preparing a response.

ElevenLabs official Scribe speech-to-text product visual

Its Realtime commit guide documents partial and committed transcripts, with manual commits or voice activity detection. Manual control gives the app a boundary; VAD can commit when a silence threshold is reached.

I would test both approaches with the same hesitant request. The goal is to discover whether a pause ends a segment at a useful point, and whether the next segment correctly revises the request.

For an agent that needs to capture an email address, show the recognized address back to the caller before sending anything. Good transcript formatting does not establish that an identity or destination is correct.

The public API page lists Scribe v2 Realtime at $0.39/hr and identifies STT usage as audio-minute billing. PAYG and subscription choices have different allowances, so confirm the applicable plan and concurrency rather than subtracting an unrelated TTS credit bundle.

The tradeoff: do not transfer batch Scribe features to Realtime by assumption. Check speaker labels, vocabulary handling, and the exact commit events in the live interface you will use.

4. Mistral Voxtral: a real option when deployment control matters

Voxtral interests me when a team wants to start with a hosted recognizer but retain a path to operating the recognition model itself. That is a different buying criterion from the lowest introductory API price.

Mistral official Voxtral audio model visual

The current Voxtral Mini Transcribe Realtime model card identifies Apache 2.0 weights and a hosted price of $0.006 per input audio minute. Open weights remove a licensing barrier; they do not remove hardware, serving, monitoring, or engineering costs.

Read the Realtime interface before promising a meeting workflow: the current live mode is not compatible with the diarize parameter. Speaker attribution therefore needs a different design if it is essential.

I would begin with the hosted path and a fixed replay set. Only move to self-hosting after the team can state the reason, such as a deployment boundary or workload economics, and name who owns failures.

A private server does not automatically make the whole voice workflow private. Audio transport, log storage, the downstream language model, and any action destination still need an explicit design.

The tradeoff: deployment freedom is valuable only when someone can operate it. Keep a recovery route for overloaded or unavailable recognition rather than turning every infrastructure incident into a silent call.

5. Google Cloud: choose the exact Chirp 3 feature combination

Google Cloud belongs on the shortlist when the application already uses its project, identity, and infrastructure. I would favor a familiar operating environment when it meets the actual audio requirements.

Google Cloud official speech-to-text company visual

The Chirp 3 documentation covers streaming, synchronous, and batch recognition. Language, region, and feature availability are documented separately, so verify the exact streaming combination rather than borrowing a capability from the batch column.

This is especially important when a proposal promises speaker attribution, a particular language, and a required processing location together. Three individually supported features do not necessarily establish that the combination is available.

V2 Standard recognition pricing begins at $0.016/min for the first 500,000 minutes per month per account, with distinct higher-volume and commitment rates. Dynamic batch pricing belongs to a different workload and should not be used to cost a live agent.

Before the pilot, check request size, stream lifetime, renewal behavior, channels, and rounding for the selected API version. A clean recognizer result is only one part of keeping a longer conversation intact.

The tradeoff: cloud integration can simplify operations, while version and region combinations add configuration work. Keep a deployment checklist specific enough that another engineer can reproduce it.

6. Microsoft Azure: start with the deployment requirement

I would evaluate Microsoft first when the client's Azure environment and speech customization needs already constrain the decision. Familiar deployment and ownership can be more useful than saving a small amount on the base rate.

Current Azure Speech in Foundry Tools official product page

The Azure Speech in Foundry Tools overview distinguishes real-time and batch transcription and explains the service's recognition and customization paths. Choose the interface for the actual application instead of treating every speech product as the same endpoint.

For a specialized vocabulary, I would begin with a list of terms that currently cause business errors. Evaluate the benefit of customization on those terms, and check whether the improvement survives a caller with an unfamiliar accent.

The pricing page depends on region and configuration; the source view I checked exposed placeholders rather than a defensible selected-region price. I have left the numerical rate unquoted rather than presenting a generic estimate as your purchase price.

Request a quote or calculator export covering the recognizer, enhanced features, and any custom endpoint hosting. Keep the speech-input meter separate from Voice Live, speech output, and the language model.

The tradeoff: procurement familiarity is useful, but it can hide a complicated bill. Make the quote reproducible and ask which feature changes would move the workload into a different meter.

7. OpenAI: keep live transcription and turn policy separate

OpenAI is worth evaluating when you want its current live transcription interface alongside an existing OpenAI application. I would still keep recognition and the authority to act as separate components.

OpenAI official live transcription documentation visual

The current transcription guide uses gpt-live-transcribe in a transcription session over WebSocket or WebRTC. For that model, turn_detection is omitted or null; server_vad and semantic_vad are not supported.

That is a reason to read today's model-specific example instead of pasting an older Realtime configuration. A working speech-to-speech example does not automatically configure a standalone transcription session.

Use utterance identifiers to reconcile transcript revisions and final text. Then write down the event or policy that permits your agent to reply, especially when the caller pauses, corrects a number, or interrupts playback.

The pricing documentation lists an estimated $0.017/min for gpt-live-transcribe. Treat that as an estimate for this model, and validate account usage; older GPT-4o transcription estimates are not the price of the current live model.

The tradeoff: one provider can simplify credentials and operations, but a transcription session is still a session your application must supervise. Test reconnects, duplicate audio, and event reconciliation before allowing automatic writes.

Calculate the bill from the meter you will actually use

I would ask every shortlisted vendor to answer the same worksheet: what starts billing, what stops it, what rounds up, what multiplies it, and what costs extra? Put the answers beside the selected model and account plan.

For session billing, measure connection time. For audio billing, measure submitted or processed audio according to the provider's policy. For token-based or estimated model pricing, reconcile the usage report instead of asserting that every minute has a fixed cost.

Illustrative arithmetic: 20 billable session hours at AssemblyAI's listed $0.45/hr would be $9 before add-ons. If only 12 hours contained speech, that would not change its session-duration charge; both the hours and activity split are hypothetical, and the billing rule explains why.

Do not turn that example into a winner against an audio-metered provider. Different silence handling, channel count, retries, and billing floors can change which quantity reaches the invoice.

Include the language model, speech generation, carrier, hosting, retained recordings, and human escalation in the full budget. Our guide to AI-agent costs helps separate the model line from the cost of delivering a service.

A proposed 24-utterance pilot for your shortlist

Here is the pilot I would run before buying. It is an illustrative test design, not a claim that these providers achieved particular results.

Collect four consented examples in each of six categories: clean requests, background noise, accents, code switching, names and numbers, and pauses with corrections. Keep a separate expected transcript and expected business action for each.

Use the same audio and network location for every candidate. Freeze the model identifier, language setting, audio encoding, vocabulary prompt, endpointing configuration, and software version in the run record.

  1. First usable partial: record when enough correct information first arrives to prepare reversible work.
  2. Committed transcript: record when the application receives the text state it intends to rely on.
  3. Turn decision: record when the system decides the caller is ready for a response.
  4. First audible response: record when playback actually starts, after recognition, reasoning, and speech generation.

Pick one clock origin for each metric and keep it fixed. A vendor's first transcript measurement cannot be compared with your end-of-speech-to-audible-response measurement simply because both are expressed in milliseconds.

Score the important entities and intended action separately from overall word error rate. Include an explicit fail when an uncertain date, amount, address, or negation produces the wrong business action.

Review the slow and wrong cases individually, not only the average. A small pilot helps reject a poor fit; it cannot establish production reliability across every caller and condition.

For the next round, expand the cases that exposed failures and keep the earlier set as a regression check. Our guide to building an agent evaluation set shows how to turn recurring mistakes into tests you can repeat after an upgrade.

Keep the voice experiment attached to a useful workflow

A beautiful live transcript has little value if the agent cannot answer the question or route the request correctly. Start with one service, one knowledge source, and one action whose outcome you can inspect.

In Pickaxe, I would first establish the Agent's instructions and Knowledge Base, then connect the necessary external service through Actions. That gives the team a clear text workflow to evaluate before adding a separate realtime voice integration.

The recognition options above are not a promise that all seven ASR vendors can be selected directly in Pickaxe. Confirm the chosen voice frontend, its authentication, and its supported handoff into your application before quoting implementation work.

Our Actions and MCP integration guide covers the business-system connection. A persistent audio socket has a different lifecycle from a request to look up a customer record.

Make the workflow useful before adding a microphone.

Build an Agent in Pickaxe, define its knowledge, and test the action boundary.

Get started →

Questions I would settle before signing

Is the lowest word error rate automatically the best API?

No. Compare the transcript's important entities, the response timing, and the application's final action. A recognizer can be strong on a benchmark and still mishandle the words that matter in your particular service.

Do I need diarization for a phone agent?

Not always. A single caller's input channel may already distinguish the human from the agent, while a conference or shared microphone needs a more deliberate speaker design.

Should I self-host an open-weight recognizer?

I would do it for a concrete deployment or operating requirement, with someone responsible for capacity and incidents. Compare total operating cost and recovery behavior against the hosted service, not just the model license.

Can a good transcript remove confirmation steps?

It can make confirmation less painful, but it does not establish permission. Keep a clear confirmation step when a mistaken request could send information, spend money, or change an important record.

My recommendation: shortlist two, then test the action boundary

For a new conversational project, I would begin with one provider that offers the turn controls I need and one that fits the team's deployment environment. Add an open-weight option only when operating it serves a specific requirement.

Then run the same difficult requests through both, inspect the invoices, and review the failed business actions. Choose the smallest integration that satisfies the real workflow and leaves a comprehensible recovery path.

If the service itself is still uncertain, build the text version in Pickaxe first and give users a useful Agent through a Portal. Add voice when it improves that service, with a recognition layer and action policy you can explain.

Related Articles