Key Moments

Workshop: Async, Sync, or Realtime: Which model do I choose?

AssemblyAIAssemblyAI
Science & Technology6 min read31 min video
Oct 8, 2026|177 views|2
Save to Pod
TL;DR

Choosing the right speech-to-text API is crucial: async is cheaper for recorded audio, sync/realtime for live interaction, and realtime with Pro models offers advanced features for voice agents.

Key Insights

1

For recorded audio where response time isn't critical, asynchronous processing is the cheapest and offers higher accuracy, costing $0.21/hour with a 10-hour file limit.

2

Realtime streaming is best for live transcription and voice agents, with Universal 3.6 Pro offering advanced features like role-playing and context passing for more accurate interactions.

3

Synchronous and Dictation APIs are designed for shorter audio clips, with Dictation offering both raw and Large Language Model-revised transcriptions, making it suitable for medical fields.

4

AssemblyAI's Universal 3.6 Pro model ranks first among 30 ASR models for average word error rate, outperforming competitors in high-noise environments and code-switching scenarios.

5

Common mistakes include using async for live transcription, streaming pre-recorded files via realtime, and fragmenting long recordings for sync/dictation APIs, leading to higher costs and errors.

6

A $20 credit is offered to new users, covering a significant number of hours across all API types, encouraging experimentation with different models.

Understanding the core decision: Who is waiting for the text?

The fundamental choice between asynchronous, synchronous, and realtime speech-to-text processing hinges on two key factors: when the audio is available and who is waiting for the transcribed text. If no one is actively waiting for the transcription in real-time, asynchronous processing is the most suitable and cost-effective option. This is ideal for batch processing pre-recorded audio files where latency is not a primary concern. Conversely, if immediate text is required, the decision branches into realtime or synchronous models, with specific considerations for voice agents versus general transcription needs. The choice directly impacts latency, accuracy, and cost, ultimately shaping the user experience.

Asynchronous processing for cost-effective, high-accuracy transcription

Asynchronous processing is the go-to for recorded audio where there's no immediate need for the transcript. This method allows for the entire audio file to be processed, leading to higher accuracy and a lower cost of $0.21 per hour, with a maximum file size of 10 hours. It's perfect for use cases like call recordings, quality assurance, and training data preparation, where the audio is already captured and the transcript can be generated at a later time. Features like speaker diarization, which identifies who said what, can be added for an additional cost, approximately $0.02 per hour. The API is straightforward to use, allowing users to send a request and receive a webhook or poll for the result, with initial results often available within seconds.

Realtime and Sync APIs for immediate transcription needs

When text is needed instantly, realtime and synchronous (sync) APIs come into play. The realtime API uses a WebSocket connection to stream transcriptions as speech occurs, suitable for live applications like live captioning or robust voice agents. The sync API, on the other hand, provides a rapid request-response for short audio clips, ideal for scenarios where a user speaks and expects an immediate text output. Dictation is a variation of sync that also provides a Large Language Model-revised version of the transcript, proving particularly useful in fields like medicine where accuracy is paramount. A key distinction is that sync and dictation are designed for shorter segments, with a maximum of two minutes per request. Attempting to process longer files with these APIs necessitates manual segmentation, which should be done at phrase boundaries rather than arbitrary time intervals to maintain accuracy.

Advanced voice agents benefit from Universal 3.6 Pro

For sophisticated voice agents, AssemblyAI's Universal 3.6 Pro model is the recommended choice. It offers advanced capabilities such as role-playing, context passing, and conversation memory, which are crucial for creating natural and effective conversational AI. The Pro model allows fine-tuning of parameters related to role detection and context management, leading to significantly improved accuracy and reduced word error rates. Furthermore, it incorporates features like voice focus to isolate the primary speaker's voice from background noise, ensuring clarity even in noisy environments like parties or car rides. This advanced functionality is essential for applications requiring nuanced interaction and understanding, differentiating it from simpler transcription needs.

The Universal 3.6 Base model for simpler listening tasks

The Universal 3.6 Base model is a simpler, more cost-effective alternative designed for applications that primarily need pure audio transcription and speaker identification. It is ideal for use cases like meeting note-takers or live commentary systems where the complex features of the Pro model are not required. While it lacks the advanced conversational AI capabilities of the Pro version, it provides reliable transcription at a lower price point. It is available in preview and is suitable for production use, with a general release planned for late October.

Key performance metrics and competitive advantage

AssemblyAI emphasizes its strong performance against market competitors, particularly in asynchronous and realtime models. Independent tests by Cobalt show Universal 3.6 Pro ranking first among 30 ASR models for average word error rate. The company highlights its focus on reducing 'missing entity error rates'—instances where critical data like phone numbers or emails are not transcribed—which is vital for downstream processes. For instance, in the medical field, their ASR models demonstrate superior performance. The Universal 3.6 Pro also shows significant gains in handling high-noise environments, code-switching between languages, and accurately transcribing short English words like 'no', all contributing to a more robust and reliable transcription service.

Common pitfalls to avoid when choosing an API

Several common mistakes can lead to increased costs and reduced accuracy. One pitfall is using asynchronous processing when real-time interaction is needed, often by attempting to segment audio into very short clips, which can lead to 'violent cuts' mid-sentence. It's recommended to use a Voice Activity Detector (VAD) for segmentation and then opt for the synchronous API for near-instantaneous results. Another error is streaming pre-recorded files via the realtime API, which is billed by connection duration rather than data processed, making asynchronous processing a far more economical choice. Similarly, fragmenting long recordings for sync or dictation APIs, which have a two-minute limit, should be done at logical phrase boundaries. Finally, leaving realtime WebSocket connections open unnecessarily can lead to substantial charges, as billing is based on the duration the connection remains active, up to a 3-hour limit.

Getting started with AssemblyAI: Credits and resources

AssemblyAI offers a $20 credit to new users, which can cover a substantial number of hours across their various APIs: approximately 95 hours for async, 66 for sync, 44 for realtime, and 32 for dictation. This credit is designed to facilitate experimentation and help users identify the best API for their specific needs. Quick-start guides and documentation are readily available to assist users in making their first API calls within minutes. The company also plans to host future webinars dedicated to specific features like Text-to-Speech and Voice Agent APIs, providing in-depth dives into these advanced capabilities.

Choosing Your Speech-to-Text API Model

Practical takeaways from this episode

Do This

For recorded audio with no immediate need for transcription, use Async.
For real-time transcription needed during a call, use Realtime.
For voice agents requiring immediate responses and context, use Universal 3.6 Pro.
For simpler transcription needs like meeting notes, use Universal 3.6 Base.
When chunking audio for Sync API, cut at phrase boundaries, not arbitrary short intervals.
When implementing real-time needs with custom chunking, use a VAD like Silero and then the Sync API.
Ensure to close Websocket connections to avoid unnecessary charges.
Choose a model that fits your use case, not just based on price.

Avoid This

Do not use Async if someone is waiting for the text in real-time.
Do not arbitrarily cut audio into very short segments (e.g., 5-30 seconds) for Sync or Dictation.
Do not stream pre-recorded files via the Realtime API; use Async for longer files.
Do not assume all features are available across all APIs; check documentation.
Do not select a model solely based on price; prioritize the best fit for your use case.

API Model Comparison

Data extracted from this episode

FeatureAsyncSyncRealtimeDictation
Use CaseRecorded audio, no real-time needShort audio clips, quick transcriptionReal-time transcription, voice agentsShort audio clips with LLM refinement
Response Time10-32 seconds (file processing)Approx. 1 secondWords appear as spoken (0.3-0.4s post-utterance)Approx. 1 second
Max Audio Length10 hours2 minutes per request3 hours (per websocket connection)2 minutes per request
Cost per Hour$0.21Higher than AsyncHigher than SyncHigher than Sync
Key FeaturesSpeaker diarization (add-on)Fast, clean transcriptionReal-time streaming, role detection (Pro)Refined text via LLM

AssemblyAI Performance Benchmarks (Async)

Data extracted from this episode

MetricAssemblyAI AsyncCompetitor AvgRank (out of 11)
Price per Hour$0.21N/A1
Word Error Rate (WER)LowestN/A1
Missing Entities Rate (Medical)LowestN/A1

AssemblyAI Performance Benchmarks (Realtime)

Data extracted from this episode

MetricAssemblyAI Universal 3.60 ProCobalt Benchmark AvgRank (out of 30)
Average Word Error Rate (WER)LowestN/A1

Free Credit Breakdown ($20)

Data extracted from this episode

API TypeHours Covered
Async95
Sync66
Realtime44
Dictation32

Common Questions

Async is for pre-recorded audio with no real-time need, offering lower cost and higher accuracy. Sync is fast for short clips. Realtime streams audio as it's spoken, ideal for voice agents. Dictation provides a refined text output for short audio segments.

Topics

Mentioned in this video

More from AssemblyAI

View all 56 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free