Back to Insights
Best Practice•2026-10-07•9 min read

Call Transcription: How to Transcribe Phone Calls

Call Transcription: How to Transcribe Phone Calls
TL
Team Laxis
Laxis Team @ Laxis

Run a webinar recording through any modern transcription tool and the result is close to flawless. Run last Tuesday's sales call through the same tool and half the surnames are wrong, the prices have drifted, and somewhere in the middle both people are apparently one person. Same software, same speakers, wildly different results — and the reason has almost nothing to do with the model.

Call transcription gets treated as a subset of ordinary transcription. It is closer to a separate discipline. Everything you are used to assuming about audio quality, speaker separation and where the file lives changes the moment the conversation travels down a telephone line instead of into a microphone. If you want the general mechanics of turning speech into text, we covered how transcription actually works separately. This is about the part phone calls break.

Why a phone call is a harder problem than a meeting

Start with what the network throws away. Traditional telephony is narrowband: the audio is sampled at 8 kHz, which means everything above roughly 3.4 kHz is discarded before it ever reaches a recorder. That band is not decorative. It is where the difference between an s and an f lives, where th separates from v, where a sibilant carries enough detail to tell fifteen from fifty. Wideband audio sampled at 16 kHz keeps it. A phone line does not.

Modern speech systems are trained overwhelmingly on wideband audio, and when you feed them 8 kHz input most of them upsample it first. Upsampling makes the file the right shape. It does not put back what was removed. There is no information in the reconstructed frequencies because there was none to begin with.

The consequences show up in the numbers, which is why published benchmarks mislead so consistently. Controlled studies that change nothing but the sample rate find that narrowband audio costs accuracy on its own. Real calls then stack codec compression, mobile networks, crosstalk and background noise on top, and error rates on production call centre audio are commonly reported at several times the leaderboard figures. Vendor benchmarks are almost always run on clean wideband recordings, so treat any headline accuracy figure as a ceiling for phone audio, not a forecast.

Compression adds its own layer. G.711, still the workhorse codec of the phone network, is narrowband by design. G.722, Opus and EVS carry wideband audio or better and sound obviously better, but they only help if every leg of the call supports them. One participant dialling in from a mobile drags the whole conversation down to the lowest common denominator, because the network has to transcode to something both ends understand.

A diagnosis you can run in thirty seconds. Open a recorded call in any audio editor and switch to its frequency view. If the energy stops dead in a flat line around 3.4 kHz, you are working with narrowband audio and no amount of switching transcription vendors will fix it. Changing how the call is carried — a VoIP leg instead of a mobile one, a headset instead of speakerphone — will.

One channel or two, and why it decides everything

The second thing that separates call audio from meeting audio is how many tracks you end up with, and it matters more than most buyers realise.

A call has two legs. If your recording captures them separately — caller on the left channel, agent on the right — then speaker separation is not a machine-learning problem at all. It is a fact about the file. Every word on channel one belongs to one person, every word on channel two belongs to the other, and the transcript comes out labelled correctly even when both people talk at once.

Mix those two legs into a single mono track and you have thrown that certainty away. Now the software has to tell the speakers apart by voice alone, using pitch, timbre and timing. On clean wideband audio that works reasonably well. On a narrowband phone line, where the very frequencies that distinguish one voice from another have been filtered out, it works considerably less well — and phone conversations happen to be full of the exact situation that breaks it, because without visual cues people interrupt each other far more than they do face to face.

Contact centre platforms have recorded dual-channel by default for years for precisely this reason. Plenty of business phone systems and most consumer recording apps still hand you a single mixed file. Before you conclude that a transcription tool is bad at your calls, check which one you are feeding it.

Three paths carry business calls, and they produce three different qualities of recording.

PSTN and mobile are the legacy path: narrowband, reliably so, with mobile adding its own compression on top. This is the floor. VoIP can be much better, since a softphone-to-softphone call over Opus or G.722 is genuinely wideband, but that only holds when both endpoints and every hop in between negotiate a wideband codec. Mixed calls fall back. Conferencing platforms such as Zoom, Google Meet and Microsoft Teams are wideband end to end when everyone joins from an app, which is the real reason meeting transcripts read so much better than call transcripts.

That last point has a practical edge to it. A Teams meeting with six people on laptops and one person on the dial-in number does not produce one bad speaker — it often produces a meeting where the dial-in participant is transcribed noticeably worse than everyone else, and where anything they say over someone else is likely to be lost. If a stakeholder always joins by phone, that is worth knowing before you rely on the record.

Five ways to get a call transcribed

There is no single answer here, only trade-offs, and the right one depends far more on where the transcript needs to end up than on which engine is fractionally more accurate this quarter.

RouteAccuracy on phone audioCost shapeWhere the transcript landsBest when
Native phone system or CRM callingDecent. Usually mono, rarely tuned to your vocabularyBundled into the seat, but gated. HubSpot only transcribes calls made on a Professional or Enterprise Sales or Service Hub seat; Starter includes 500 calling minutes a month per account, and free accounts get little or noneAutomatically on the contact or deal recordReps already dial from inside the CRM
Telephony or CPaaS APIHighest ceiling. You choose the engine and can request dual-channelMetered and stacked. Twilio, for example, bills transcription per minute, with recording and the call itself charged as separate linesWherever you write code to put itYou have engineers and unusual requirements
Meeting assistantStrong on app-based calls, weaker on dial-in legsPer user per month, often with a free allowanceIts own workspace, plus CRM sync where offeredMost calls happen on Zoom, Meet or Teams
Contact centre suiteBest in class on telephony. Dual-channel by default, tuned for narrowbandHeavy. Five9's voice tier lists around $159 per concurrent user per month with a 50-seat minimum; Genesys Cloud CX 4 lists at $240The platform, with live agent assist during the callYou run call volume, not meetings
Human or hybrid serviceHighest on genuinely bad audioHighest per call, and slowest to returnA delivered file you route yourselfThe call will be read by a court or a regulator

Prices are list prices as published in 2026; contact centre contracts in particular are routinely negotiated well below them at volume.

Most teams overthink the accuracy column and underthink the one about where the transcript lands. A contact centre suite is the technically correct answer for telephony and the wrong purchase for a twelve-person company that holds four discovery calls a week on Google Meet. For that shape of team the meeting-assistant route usually wins on cost and on the thing that actually matters, which is whether the output reaches the CRM. Tools such as the Laxis AI meeting assistant sit in that lane, capturing across Zoom, Google Meet and Microsoft Teams, accepting uploaded MP3, WAV and M4A files from calls recorded elsewhere, and, on the Business plan, pushing summaries and action items into HubSpot or Salesforce. The honest caveat is the same one that applies to every tool in that column: it is built around meetings and conversations rather than contact centre call volume, so if you are transcribing four thousand calls a day, buy the platform designed for it. There is a fuller roundup of note-taking tools if that is the column you are shopping in.

Live, or after everyone has hung up

Real-time and post-call transcription look like the same feature and behave nothing alike.

Streaming transcription returns text within a second or two of the words being spoken, which is what makes live agent assist, compliance alerts and on-screen translation possible. The price is accuracy. A streaming system has to commit to a guess before it has heard the end of the sentence, and while it revises earlier words as more audio arrives, it never gets the luxury of full context. Post-call processing does. It has the entire recording, can look forwards as well as backwards, and is generally both more accurate and cheaper per minute.

The useful move is to stop treating it as a single choice. Live output is for things that must happen during the call. Post-call output is for the record, the summary, the CRM entry and anything anyone will read later. Plenty of platforms do both, and there is no reason the same call cannot produce a rough live feed and a clean final transcript.

Two cheap upgrades that beat switching vendors. Load a custom vocabulary with your product names, competitor names, SKU formats and the surnames your reps mangle — proper nouns are where phone transcripts fail most visibly, and no model guesses them without help. Then get everyone onto headsets. Speakerphone in a hard-walled room adds echo, and echo confuses speaker separation faster than any accent does.

The part that decides whether any of it was worth it

Here is the uncomfortable truth about call transcription software: accuracy is the thing everybody evaluates and workflow is the thing that determines whether the purchase pays for itself.

A transcript sitting in a tool nobody opens has no value. A transcript attached to the right contact, on the right opportunity, searchable next to the rest of that account's history, changes how a deal gets run. The gap between those two outcomes is a plumbing question, and it is worth answering before you compare engines.

Three things make the difference. The transcript has to be attached to a record, not just stored — contact, company and open deal, resolved automatically from the phone number or the calendar invite. It has to be reduced to something readable, because nobody scrolls a forty-minute verbatim log looking for the objection; a summary and a list of commitments with names against them is what gets used. And it has to arrive without a human step, since any workflow that depends on a rep remembering to paste something will quietly stop working in week three.

That last requirement is where most setups fall down, and it is the reason native CRM calling and CRM-syncing assistants punch above their accuracy scores. Laxis, for instance, extracts action items with owners attached and, on its Business plan, syncs the result into HubSpot or Salesforce without anyone opening the transcript; processing happens in the cloud rather than on your device, which is worth checking against your own policy before you commit.

What to expect, and how to find out cheaply

Do not buy on a published accuracy figure. Take five real recordings — the bad line, the strong accent, the one with the client's procurement lead talking over your AE — and run them through every shortlisted tool. Count the errors that would actually cost you something: wrong numbers, wrong names, wrong attribution of who agreed to what. Typos in filler words are noise.

Set expectations accordingly. On a clean VoIP call with two people on headsets, good systems land in the same territory as meeting audio. On a mobile call with road noise and an accent the model has not heard much of, the transcript will be usable for recall and unreliable for quotation. Both can be true of the same vendor on the same day, which is exactly why the vendor cannot tell you a single number and mean it.

One thing that is not optional: get consent right before any of this runs, because the rules vary by jurisdiction and the penalties are not theoretical. Our guide to recording a phone call covers the practical mechanics and the consent rules, including the US states that require every party to agree.

Where this leaves you

The interesting shift in call transcription is not that the models got better, though they did. It is that the bottleneck moved. For most of the last decade the hard part was getting the words right. Now the hard part is a band of frequencies the telephone network has been discarding for decades and a set of decisions about how the recording is captured — one channel or two, app or dial-in, wideband or whatever the network could negotiate. Those choices are made by whoever configures the phone system, usually without anyone mentioning transcription.

Which makes this a rare situation where the cheapest available improvement is not a purchase. It is a conversation with whoever owns your telephony, asking two questions: can we record both legs separately, and can we keep more calls on a wideband path. Answer those and almost any competent engine will do.

Frequently asked questions

What is call transcription?

Call transcription is the conversion of a phone or VoIP conversation into searchable text, usually with each speaker labelled and each line timestamped. It differs from transcribing an ordinary audio file in one important way: the sound has travelled down a compressed telephone channel, and it may arrive as one mixed recording or as two separate legs.

How accurate is call transcription on phone audio?

Expect noticeably lower accuracy than the figures vendors publish, because those come from clean wideband recordings. Narrowband audio costs accuracy on its own, and real calls add compression, crosstalk and noise on top, so error rates on phone audio can run several times higher than the headline figure. Two-channel recording and custom vocabulary claw much of that back.

Can you transcribe phone calls automatically?

Yes. Most business phone systems, CRMs with native calling and contact-centre platforms transcribe automatically once recording is switched on, and telephony APIs will stream text back live if you build against them. The practical constraint is usually licensing rather than technology, since many plans gate transcription behind a paid seat tier or a monthly minute allowance.

What is the best way to transcribe sales calls?

Pick the route that lands the transcript on the right CRM record without anyone copying anything. In practice that means either a dialler with native CRM logging, or a meeting assistant that syncs to HubSpot or Salesforce for calls held on Zoom, Google Meet or Teams. Raw accuracy matters less than whether a rep ever opens the thing.

Does call transcription work in real time?

Yes, although live and post-call transcription are effectively different products. Streaming systems return text within a second or two and revise earlier words as more audio arrives, which caps their accuracy because they cannot see what comes next. Processing after the call has the whole recording available and is usually both more accurate and cheaper.

Where should call transcripts be stored?

Wherever the people who need them already work, which for revenue teams means the CRM, attached to the contact and the open opportunity rather than parked in a separate tool. Regulated teams layer retention windows and access controls on top. A transcript nobody can find while writing a follow-up has cost money rather than saved it.