# Call quality

The gap between a caller finishing and the agent answering, the two settings that shape the line, and the trade-off between them.

Source: https://docs.omazy.ai/how-to/voice/call-quality/

import Figure from '../../../../components/Figure.astro'
import { Aside } from '@astrojs/starlight/components'

A voice agent is judged on timing before it is judged on answers. A correct
reply that arrives after three seconds of dead air loses to a vaguer one that
arrives immediately.

## The gap, and what fills it

Between a caller finishing their sentence and the agent starting to speak,
several things have to happen: notice they stopped, turn the audio into words,
work out an answer, and turn that answer back into speech. None of it is
instant.

<Figure
  label="The pause after a caller speaks, with and without an acknowledgement"
  caption="The gap is the same length in both rows. The difference is whether the caller experiences it as the agent working or as the call having dropped."
>
<svg viewBox="0 0 700 200" xmlns="http://www.w3.org/2000/svg">
  <text x="0" y="14" class="d-eyebrow">THE GAP THE CALLER HEARS</text>

  <text x="0" y="36" class="d-sub">Without an acknowledgement</text>
  <rect x="0" y="44" width="200" height="34" rx="8" class="d-box" />
  <text x="14" y="66" class="d-sub">Caller speaks</text>
  <rect x="204" y="44" width="300" height="34" rx="8" class="d-box" stroke-dasharray="5 4" />
  <text x="286" y="66" class="d-sub">silence</text>
  <rect x="508" y="44" width="190" height="34" rx="8" class="d-box" />
  <text x="522" y="66" class="d-sub">Agent answers</text>

  <text x="0" y="122" class="d-sub">With one</text>
  <rect x="0" y="130" width="200" height="34" rx="8" class="d-box" />
  <text x="14" y="152" class="d-sub">Caller speaks</text>
  <rect x="204" y="130" width="300" height="34" rx="8" class="d-box-accent d-pulse" />
  <text x="248" y="152" class="d-sub">"Let me check that for you."</text>
  <rect x="508" y="130" width="190" height="34" rx="8" class="d-box" />
  <text x="522" y="152" class="d-sub">Agent answers</text>

  <text x="204" y="188" class="d-accent-text">same delay, entirely different call</text>
</svg>
</Figure>

This is why [acknowledgements](/how-to/voice/phrases/) are the highest-value phrases
you can write. They do not make anything faster. They change dead air into
something that sounds like a person thinking, which is the part the caller
actually reacts to.

## The two settings

Both live on the **Voice** tab. Both are trade-offs, and the defaults are
reasonable, so change them only against a specific complaint.

### Pacing lead

The agent's speech is sent slightly ahead of when the caller needs to hear it.
That small buffer absorbs network hiccups.

With no lead, at most a fraction of a second of audio is ever ahead of the
caller's ear, so any stall between the platform and the phone network is heard
as a gap in the middle of a word. Adding lead smooths those out.

The cost is that interrupting the agent takes marginally longer to take effect,
because there is more audio already in flight to discard.

**Raise it if** callers report choppy or stuttering speech. **Leave it alone
if** the line already sounds smooth.

### Echo suppression

If the caller is on a speakerphone, the agent's own voice comes back through
their microphone. Without protection, the agent hears itself, treats it as the
caller speaking, and answers its own sentence. It is unmistakable once you have
heard it: the agent starts talking to itself and the caller cannot get a word
in.

Echo suppression stops the agent listening while its own audio is still
reaching the caller, plus a short tail for the room.

<Aside type="caution" title="The trade-off, stated plainly">
While echo suppression is active, the caller cannot interrupt the agent
mid-sentence. You are choosing between a line that can be interrupted and a
line that never argues with itself. For speakerphone and car callers, the
second is the right choice. For a headset-only line, suppression can be turned
down and barge-in works throughout.
</Aside>

## What good sounds like

- **The greeting starts immediately.** Any pause before the first word is the
  worst pause on the call, because the caller does not yet know anyone is
  there.
- **Replies begin within a beat or two.** Some questions take longer, which is
  what acknowledgements are for.
- **Interrupting works**, unless you have deliberately traded it away.
- **No clipped word endings.** Speech cut off at the end usually means the
  agent decided the caller had stopped talking too eagerly.

## Things outside these settings

Some call quality problems are not settings problems, and recognising them
saves a lot of tuning.

**The caller's own connection.** A caller on a poor mobile signal will hear a
poor call regardless. Their audio reaching the platform intact matters as much
as anything on this page.

**Distance.** A call routed through a distant part of the network before it
reaches the agent carries that round trip on every single turn. If test calls
from one location are consistently slower than another, that is usually why,
and no amount of pacing adjustment fixes it.

**Speakerphones in hard rooms.** A tiled room with a speakerphone is the worst
case for both echo and recognition. It is worth testing one deliberately, since
a meaningful share of real calls arrive that way.

## Testing changes

Change one setting at a time and place a real call after each. These settings
interact, and two changes at once leaves you unable to attribute the result.

Test on the kind of connection your callers actually use. A change that
improves a call from a quiet office over good WiFi can make things worse for
someone on a phone in a car, which is where many of your calls come from.
