Clarity, also referred to as Clarity 1, is a real-time speech enhancement model from KugelAudio built for voice agents and call centers. It does two jobs at once on a live call: it removes background noise, and it extracts the primary speaker so that every other voice — café chatter, a colleague at the next desk, the crowd around the caller — is cut out. What remains is one clean voice. The purpose is straightforward: whoever is listening, whether a person or a voice bot, hears clear speech, and the agent responds to the person calling rather than to the room they are calling from. Clarity is aimed at teams building and running voice agents and at call centers that need the caller, not the environment, to come through.
The problem Clarity addresses is the environment callers actually call from. People call from cafés, shared offices and busy streets; images on the product site show callers on the street, on a construction site and on a train. The audio a voice agent receives is therefore a mix of the caller's voice, background noise and other people talking. Speech-to-text, turn detection and the language model behind a bot can only work with the audio they receive, so background talkers end up in transcripts and set off turn detection, and the bot ends up responding to the room instead of to the caller. Clarity is presented as the layer that fixes this before the rest of the pipeline ever sees the audio, so the models downstream work from one clean voice.
The first core capability is real-time background noise removal. Clarity strips out background noise from live audio rather than from a finished recording, which matters because conversations are live and the agent has to react while the caller is still speaking. The product materials place this against loud places specifically — clear calls in the café, on the street, on the site and on the train. Because the noise is removed before it reaches the agent, the agent is not asked to guess which parts of the incoming audio are speech and which are environment; the clean signal is what arrives.
The second core capability is primary speaker extraction. This goes a step beyond noise removal: it cuts out every other voice, from the café crowd to the person at the next desk, so only the caller comes through. Clarity explains how it decides which voice to keep — primary speaker extraction uses a short reference recording of the person you want to follow and cuts out every other voice. Without a reference, it keeps the loudest speaker, which gives teams a working default when no reference recording is available. Both behaviours are described as running in real time rather than as an offline post-processing step.
The third group of capabilities concerns how Clarity fits into a live pipeline. It processes audio in a continuous stream, as it arrives, and is built for live conversations rather than for waiting on a completed recording. Its chunk size is stated explicitly: Clarity processes audio in 240 ms chunks and needs about 50 ms for each one. If your turn detection already reads audio in 240 ms chunks, Clarity adds only those 50 ms. With smaller input frames, a sound can come out up to 290 ms after it went in, depending on where it falls in its chunk. Those numbers are published on the site so teams can judge whether the added latency fits their pipeline, and the headline figure quoted for voice agents is +50 ms of added latency in pipelines with 240 ms chunks.
How Clarity works overall is described as two jobs performed in real time on a stream: noise removal and primary speaker extraction, delivered as audio arrives. Its unique approach is doing both at once during a live call so that downstream components — speech-to-text, turn detection and the model behind the bot — all work from a single clean voice. KugelAudio also publishes comparative test results. On target-speaker quality (Libri2Mix, DNSMOS, higher is better), Clarity 1 at 240 ms is compared with StarTSE at 560 ms, scoring 3.56 on speech quality (SIG), 4.03 on background quality (BAK) and 3.27 overall (OVRL). On noise removal (DNS 2020 Challenge, DNSMOS, higher is better, with reverb), Clarity 1 at 240 ms scores 3.62 SIG, 4.11 BAK and 3.37 OVRL against unprocessed original audio, AI Acoustics at 15 ms, DeepFilterNet3 at 310 ms and GTCRN at 16 ms. These comparisons are presented on the site as evidence of quality against other approaches.
The benefits Clarity states for users follow directly from a single clean voice reaching the rest of the stack. Speech-to-text, turn detection and the language model behind a bot can only work with the audio they receive, so giving them the caller's voice alone means background talkers stop ending up in transcripts and stop setting off turn detection, and the bot responds to your caller instead of the room. For voice agents, the stated outcome is that your voice bot hears only the caller. For call centers, the stated outcome is that customers hear your agent and only your agent: on a busy call center floor, Clarity removes background noise, echo and the colleagues nearby, and enhances the agent's voice so every word comes through. In both cases the through-line is accuracy of understanding and clarity of output.
Concrete use cases described in the product content include voice agents on live calls, where Clarity removes background noise and chatter before the bot hears them so its turn detection and speech-to-text react only to the caller; call centers, where Clarity removes background noise, echo and nearby colleagues and enhances the agent's voice; and callers in demanding environments such as cafés, shared offices, busy streets, construction sites and trains. Teams can also hear it on their own audio: you can sign up and run your own recordings through Clarity in the dashboard, free until October 30th, and for a self-hosted deployment or a larger rollout, contact the team and bring examples from your callers' environments. The site also offers a sales conversation for teams that want to go further with the model.
Clarity is aimed at anyone building or operating voice agents and at call centers, as reflected in the two audiences named on the site: voice agents and call centers. It fits into pipelines that stream audio in chunks, with the stated latency budget depending on chunk size. Deployment options include running Clarity in KugelAudio's environment, or self-hosting and on-premise deployment available on Enterprise plans. Pricing is published as a free tier at €0 until October 30th, covering noise removal and primary speaker extraction in real time at no cost while you build, followed by pay as you go at €0.0015 per audio minute excluding VAT (€0.001785 including 19% VAT), where you pay only for the audio Clarity processes. Consumers in other EU countries pay their local VAT rate, shown before payment.
In summary, Clarity's value proposition is narrow and clearly stated: remove the noise, keep the voice. By removing background noise and extracting the primary speaker in real time on a live stream, it hands the rest of a voice pipeline one clean voice — the caller's — so speech-to-text, turn detection and the language model can do their jobs, and so callers reach an agent that hears them rather than the room around them.