Skip to content
tuxtone

foundations

RNNoise, explained: the small model under almost everything

Pixel drawing: a spectrum of about two dozen bands feeding a small neural network.

By the TuxTone crew · Updated 2026-09-20

Nearly every noise suppression path we cover — EasyEffects, the PipeWire filter-chain, OBS's filter, NoiseTorch, even an ffmpeg filter — is a different wrapper around the same core idea, and usually around the same code: RNNoise. Understanding it once means you understand the whole stack's behavior, its shared strengths and its shared failure modes. This page is that once.

What RNNoise is

RNNoise is a real-time noise suppression library from the Xiph.Org orbit — the same community that produced Opus and Vorbis — built around a deliberately tiny recurrent neural network. It was introduced by Jean-Marc Valin (of Opus fame) as a hybrid: classic signal processing does the heavy framing and analysis, and a small neural network makes the judgment calls that hand-tuned DSP historically got wrong. The code is open source under a BSD license, which is exactly why it could spread into every corner of the Linux audio world without anyone signing anything.

Maintenance reality, since that's our house obsession: the upstream repo (xiph/rnnoise) tagged v0.2 in April 2024 — bringing improved training and SSE/AVX runtime optimizations — and saw commits into 2025. It moves slowly, which for a finished, single-purpose library is normal, not alarming. The projects embedding it (EasyEffects, the LADSPA plugin, ffmpeg) are the actively-moving layer.

How it works, in plain terms

RNNoise processes mono audio at 48 kHz in 10-millisecond frames. For each frame it does three things:

  1. Summarize the spectrum. Instead of juggling hundreds of frequency bins, it collapses the frame into roughly two dozen bands spaced the way human hearing perceives pitch — fine resolution down low, coarse up high.
  2. Ask the network what's speech. A recurrent network — small enough to be described in tens of thousands of weights, not billions — looks at those bands over time and estimates, per band, how much is voice and how much is noise. The recurrence is the trick: because the network carries memory across frames, it can tell sustained fan drone from the sustained "ee" in your sentence.
  3. Turn the estimate into gains. Each band gets attenuated according to that estimate, with an extra pitch-filtering step to keep the harmonics of your voice from getting chewed. The result is reassembled into audio and shipped onward — all within milliseconds.

It also emits a voice-activity probability per frame, which wrappers expose as a "VAD threshold" control: below the threshold, the frame is treated as pure noise and muted entirely. That single knob explains most of the behavioral differences you'll notice between tools that are otherwise running identical code.

one frame, start to finish
  1. mono in, 48 kHzcut into 10 ms frames — 480 samples each
  2. ~two dozen bandsspaced the way hearing hears pitch
  3. small recurrent netvoice or noise, band by band, with memory across frames
  4. one gain per bandplus pitch filtering to spare your harmonics
  5. clean frame outwithin milliseconds

Side output: a voice-activity probability per frame — the “VAD threshold” knob every wrapper shows you.

The three steps above, plus the arithmetic: 48,000 samples a second × 10 ms = 480 samples per frame.

Why "tiny" is the point

RNNoise predates the everything-is-a-giant-model era, and its constraint aged beautifully: it was designed to run in real time on ordinary CPUs — the original write-up bragged about running fine on a Raspberry Pi. On a desktop, the cost disappears into the noise floor of your process list (the one place we like noise). That's why it can sit inside PipeWire's graph processing every frame of your mic all day without you ever thinking about it, and why five different projects embed it rather than inventing their own model.

The trade-off is honesty about ceilings. A 2017-vintage tiny model handles stationary noise superbly — fans, air conditioning, electrical hum, steady traffic. It struggles more with non-stationary noise: sudden clatter, a second voice, a keyboard hammered mid-sentence. Newer, heavier open models (like DeepFilterNet, which EasyEffects offers as its second engine) push that ceiling up at more CPU cost. Paid tools push it further still — though the flagship there doesn't ship a Linux app, so for most desktop Linux users the comparison is academic.

Where you'll actually meet RNNoise

WrapperWhat it wraps it asOur guide
EasyEffects"Noise Reduction" effect in the input pipelinesetup
noise-suppression-for-voice (werman)LADSPA/LV2 plugin, v1.21 (2026-05-29)filter-chain
OBS Studio"Noise Suppression" filter, RNNoise method (default)OBS guide
NoiseTorchVirtual denoised microphone (dormant project)status
ffmpegarnndn audio filter, loadable .rnnn modelsCLI guide
one core, 5 wrappers

RNNoiseone BSD-licensed core

  • EasyEffects"Noise Reduction" effect in the input pipeline
  • noise-suppression-for-voice (werman)LADSPA/LV2 plugin, v1.21 (2026-05-29)
  • OBS Studio"Noise Suppression" filter, RNNoise method (default)
  • NoiseTorchVirtual denoised microphone (dormant project)
  • ffmpegarnndn audio filter, loadable .rnnn models
The table above, drawn as wiring: stack two of these and RNNoise runs on its own output.

This shared ancestry has a practical consequence we repeat across the site: don't stack them. RNNoise feeding RNNoise doesn't double the cleaning; it doubles the artifacts, because pass two treats pass one's residue as signal to mangle. Pick one wrapper per audio path.

What it can't fix

  • Clipping and overload. If your input gain is too hot, the distortion is baked in before any model sees it. Fix levels first.
  • Reverb and room boom. Suppression removes noise, not reflections. A bare-walled room sounds like a bare-walled room, just a quieter one.
  • Another person talking. RNNoise is trained to keep speech. A second speaker is speech. It will faithfully keep your roommate's phone call.
  • Music you wanted to keep. The flip side of speech-only training: it eats music. Never leave voice suppression enabled while recording instruments.
  • Damaged recordings. Real-time frames can't use future context. For an existing file, an offline pass with proper tools does better — that's the CLI page.
Fact check for this page: upstream repo xiph/rnnoise, latest tagged release v0.2 (2024-04-15), repository not archived, commits into 2025; BSD-licensed. EasyEffects lists RNNoise as the basis of its Noise Reduction effect in its own README.verified against github.com/xiph/rnnoise and github.com/wwmm/easyeffects · 2026-09-20
Money note: RNNoise is BSD-licensed free software. There is nothing to buy, no affiliate program, and nobody paying for this page. If you want to support it, contribute upstream — github.com/xiph/rnnoise (no affiliate relationship; there's nothing to affiliate).