Key Moments
Workshop: Async, Sync, or Realtime: Which model do I choose?
Key Moments
Choosing the right speech-to-text API is crucial: async is cheaper for recorded audio, sync/realtime for live interaction, and realtime with Pro models offers advanced features for voice agents.
Key Insights
For recorded audio where response time isn't critical, asynchronous processing is the cheapest and offers higher accuracy, costing $0.21/hour with a 10-hour file limit.
Realtime streaming is best for live transcription and voice agents, with Universal 3.6 Pro offering advanced features like role-playing and context passing for more accurate interactions.
Synchronous and Dictation APIs are designed for shorter audio clips, with Dictation offering both raw and Large Language Model-revised transcriptions, making it suitable for medical fields.
AssemblyAI's Universal 3.6 Pro model ranks first among 30 ASR models for average word error rate, outperforming competitors in high-noise environments and code-switching scenarios.
Common mistakes include using async for live transcription, streaming pre-recorded files via realtime, and fragmenting long recordings for sync/dictation APIs, leading to higher costs and errors.
A $20 credit is offered to new users, covering a significant number of hours across all API types, encouraging experimentation with different models.
Understanding the core decision: Who is waiting for the text?
The fundamental choice between asynchronous, synchronous, and realtime speech-to-text processing hinges on two key factors: when the audio is available and who is waiting for the transcribed text. If no one is actively waiting for the transcription in real-time, asynchronous processing is the most suitable and cost-effective option. This is ideal for batch processing pre-recorded audio files where latency is not a primary concern. Conversely, if immediate text is required, the decision branches into realtime or synchronous models, with specific considerations for voice agents versus general transcription needs. The choice directly impacts latency, accuracy, and cost, ultimately shaping the user experience.
Asynchronous processing for cost-effective, high-accuracy transcription
Asynchronous processing is the go-to for recorded audio where there's no immediate need for the transcript. This method allows for the entire audio file to be processed, leading to higher accuracy and a lower cost of $0.21 per hour, with a maximum file size of 10 hours. It's perfect for use cases like call recordings, quality assurance, and training data preparation, where the audio is already captured and the transcript can be generated at a later time. Features like speaker diarization, which identifies who said what, can be added for an additional cost, approximately $0.02 per hour. The API is straightforward to use, allowing users to send a request and receive a webhook or poll for the result, with initial results often available within seconds.
Realtime and Sync APIs for immediate transcription needs
When text is needed instantly, realtime and synchronous (sync) APIs come into play. The realtime API uses a WebSocket connection to stream transcriptions as speech occurs, suitable for live applications like live captioning or robust voice agents. The sync API, on the other hand, provides a rapid request-response for short audio clips, ideal for scenarios where a user speaks and expects an immediate text output. Dictation is a variation of sync that also provides a Large Language Model-revised version of the transcript, proving particularly useful in fields like medicine where accuracy is paramount. A key distinction is that sync and dictation are designed for shorter segments, with a maximum of two minutes per request. Attempting to process longer files with these APIs necessitates manual segmentation, which should be done at phrase boundaries rather than arbitrary time intervals to maintain accuracy.
Advanced voice agents benefit from Universal 3.6 Pro
For sophisticated voice agents, AssemblyAI's Universal 3.6 Pro model is the recommended choice. It offers advanced capabilities such as role-playing, context passing, and conversation memory, which are crucial for creating natural and effective conversational AI. The Pro model allows fine-tuning of parameters related to role detection and context management, leading to significantly improved accuracy and reduced word error rates. Furthermore, it incorporates features like voice focus to isolate the primary speaker's voice from background noise, ensuring clarity even in noisy environments like parties or car rides. This advanced functionality is essential for applications requiring nuanced interaction and understanding, differentiating it from simpler transcription needs.
The Universal 3.6 Base model for simpler listening tasks
The Universal 3.6 Base model is a simpler, more cost-effective alternative designed for applications that primarily need pure audio transcription and speaker identification. It is ideal for use cases like meeting note-takers or live commentary systems where the complex features of the Pro model are not required. While it lacks the advanced conversational AI capabilities of the Pro version, it provides reliable transcription at a lower price point. It is available in preview and is suitable for production use, with a general release planned for late October.
Key performance metrics and competitive advantage
AssemblyAI emphasizes its strong performance against market competitors, particularly in asynchronous and realtime models. Independent tests by Cobalt show Universal 3.6 Pro ranking first among 30 ASR models for average word error rate. The company highlights its focus on reducing 'missing entity error rates'—instances where critical data like phone numbers or emails are not transcribed—which is vital for downstream processes. For instance, in the medical field, their ASR models demonstrate superior performance. The Universal 3.6 Pro also shows significant gains in handling high-noise environments, code-switching between languages, and accurately transcribing short English words like 'no', all contributing to a more robust and reliable transcription service.
Common pitfalls to avoid when choosing an API
Several common mistakes can lead to increased costs and reduced accuracy. One pitfall is using asynchronous processing when real-time interaction is needed, often by attempting to segment audio into very short clips, which can lead to 'violent cuts' mid-sentence. It's recommended to use a Voice Activity Detector (VAD) for segmentation and then opt for the synchronous API for near-instantaneous results. Another error is streaming pre-recorded files via the realtime API, which is billed by connection duration rather than data processed, making asynchronous processing a far more economical choice. Similarly, fragmenting long recordings for sync or dictation APIs, which have a two-minute limit, should be done at logical phrase boundaries. Finally, leaving realtime WebSocket connections open unnecessarily can lead to substantial charges, as billing is based on the duration the connection remains active, up to a 3-hour limit.
Getting started with AssemblyAI: Credits and resources
AssemblyAI offers a $20 credit to new users, which can cover a substantial number of hours across their various APIs: approximately 95 hours for async, 66 for sync, 44 for realtime, and 32 for dictation. This credit is designed to facilitate experimentation and help users identify the best API for their specific needs. Quick-start guides and documentation are readily available to assist users in making their first API calls within minutes. The company also plans to host future webinars dedicated to specific features like Text-to-Speech and Voice Agent APIs, providing in-depth dives into these advanced capabilities.
Mentioned in This Episode
●Software & Apps
●Companies
Choosing Your Speech-to-Text API Model
Practical takeaways from this episode
Do This
Avoid This
API Model Comparison
Data extracted from this episode
| Feature | Async | Sync | Realtime | Dictation |
|---|---|---|---|---|
| Use Case | Recorded audio, no real-time need | Short audio clips, quick transcription | Real-time transcription, voice agents | Short audio clips with LLM refinement |
| Response Time | 10-32 seconds (file processing) | Approx. 1 second | Words appear as spoken (0.3-0.4s post-utterance) | Approx. 1 second |
| Max Audio Length | 10 hours | 2 minutes per request | 3 hours (per websocket connection) | 2 minutes per request |
| Cost per Hour | $0.21 | Higher than Async | Higher than Sync | Higher than Sync |
| Key Features | Speaker diarization (add-on) | Fast, clean transcription | Real-time streaming, role detection (Pro) | Refined text via LLM |
AssemblyAI Performance Benchmarks (Async)
Data extracted from this episode
| Metric | AssemblyAI Async | Competitor Avg | Rank (out of 11) |
|---|---|---|---|
| Price per Hour | $0.21 | N/A | 1 |
| Word Error Rate (WER) | Lowest | N/A | 1 |
| Missing Entities Rate (Medical) | Lowest | N/A | 1 |
AssemblyAI Performance Benchmarks (Realtime)
Data extracted from this episode
| Metric | AssemblyAI Universal 3.60 Pro | Cobalt Benchmark Avg | Rank (out of 30) |
|---|---|---|---|
| Average Word Error Rate (WER) | Lowest | N/A | 1 |
Free Credit Breakdown ($20)
Data extracted from this episode
| API Type | Hours Covered |
|---|---|
| Async | 95 |
| Sync | 66 |
| Realtime | 44 |
| Dictation | 32 |
Common Questions
Async is for pre-recorded audio with no real-time need, offering lower cost and higher accuracy. Sync is fast for short clips. Realtime streams audio as it's spoken, ideal for voice agents. Dictation provides a refined text output for short audio segments.
Topics
Mentioned in this video
More from AssemblyAI
View all 56 summaries
30 minEvent Recap: Build Smarter Voice Agents - New York Edition
37 minWorkshop: Building and optimizing dictation features
47 minBuild Smarter Voice Agents
60 minBuild a Voice Agent in an Hour with Claude Code | AssemblyAI Workshop
Ask anything from this episode.
Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.
Get Started Free