Fig. 1 — Speech sample, 1024-point FFT.
Clean up a recording — no account needed

Turn a raw recording into a finished episode tonight

Dead air, room hum, four takes to stitch together and a transcript for the show notes — Audio Magic handles all of it in your browser, running Whisper, DeepFilterNet and Demucs on your own machine. Your recording never leaves your computer, and nothing is metered per minute.

Free for 15 minutes of audio per tool per day, files up to 5 minutes. No card, no watermark.

§ 1

Every step of the cleanup, in one place

Cut the dead air and the long pauses. Strip the room hum, fan and air conditioning off the voices. Stitch your takes into one file, fade the ends, and get a transcript back for the show notes and the search engines. It is everything that stands between a recording and something you can publish.

§ 2

For interviews you cannot upload

Some recordings should never reach a cloud service — a source who needs protecting, a client interview, a patient session, an internal investigation. Your recording is never sent anywhere: it is read by the browser off your own disk, worked on there, and saved back. There is no processing agreement to sign and nothing left on a server to delete afterwards. The AI models themselves download to your browser the first time you use a tool, and are reused from then on.

§ 3

No upload, no queue, no waiting

A two-hour interview spends twenty minutes uploading to a browser tool before it starts doing anything, and then waits its turn behind everyone else. Here the work starts the second you drop the file in, because it happens on the machine you are already sitting at — on a train, on a plane, or with the Wi-Fi switched off.

§ 4

Free, and honest about what that means

15 minutes of audio per tool per day, and single files up to 5 minutes long, with no watermark on anything you export, no trial counting down and no account needed to start. If you outgrow either limit, a subscription removes both for $4.99 a month, with the first 7 days free. That is the entire business model, and it is the only thing an account gets you.

Index of tools — every one free to use today

Whisper for transcription, DeepFilterNet3 and RNNoise for noise, Demucs for stems, Kokoro and Chatterbox for speech — every one of them running on your own machine.

Start here

Transcribe Audio

Turn speech into text or timestamped subtitles, in over 90 languages, using an AI speech recognition model. It runs on your own device, so nothing is uploaded.

Needs WebGPU

Clone a Voice

Record about eight seconds of a voice, then type anything and hear it said in that voice. The model runs on your own device, so the recording is never uploaded. Needs WebGPU and a 1.5 GB one-off download, which is the trade for a voice model this capable running locally at all.

Free voices

Generate Speech

Type or paste text and hear it read aloud in a natural voice, then download it as WAV or MP3. Fifteen voices in American and British English — twenty-eight if you turn on the low-rated ones — each with a sample you can play first. The model runs on your own device, so nothing you write is uploaded — which also means generating takes about twice as long as the speech it produces.

New

Chain Tools

Build a chain of edits and run it over a batch in one pass — combine, split a stereo file into its channels, remove silence, strip background noise, bleep out words on a list, normalize, fade, picture the sound as a spectrogram, and write the result out as WAV or MP3. Arrange it as a guided list or as a node network you wire up yourself, play back what comes out of any block before you commit, save the chain to reuse next week, and run the whole thing on your own device like every other tool here.

Detect Sounds

Teach it a sound from a handful of short clips — a dog barking, a gunshot, a doorbell — then find every place that sound happens in a long recording. Exports the timestamps as a subtitle file. Training and detection both run on your own device.

Bleep Words

Give it a list of words and it finds every place they are said and covers each one — with the standard broadcast beep, with silence, or with a short sound of your own. Speech recognition runs on your own device to find the words, so the recording is never uploaded, and every run tells you exactly what it covered and when. There is no built-in list of profanity: it bleeps the words you list and no others.

Make a Spectrogram

Turn a recording into a picture of itself — frequency up the side, time across the bottom, loudness as colour. Select any region, pick a palette and an aspect ratio, and export it as a PNG with the axes on for analysis or off for a square you can hang on a wall. A stereo file is drawn as two plots stacked one above the other, sharing a time axis and a colour scale so the channels can be read against each other.

Remove Silence

Cut dead air and silent gaps out of a recording with exact, non-destructive DSP analysis. The waveform shows you exactly what goes before you commit.

Remove Background Noise

Strip hiss, hum, fans and room tone out of a voice recording using an AI speech-enhancement model that runs in your browser.

Separate Stems

Split a song into separate tracks with an AI model — take the vocals out for a karaoke backing track, or pull drums, bass and vocals apart. Runs on your own device.

Split Stereo

Pull a stereo recording apart into two mono files — one for the left channel, one for the right. Useful when a stereo file was never really stereo: two lav mics printed to the two sides, an interview with the guest on one channel and the room on the other. The samples are copied rather than processed, so both files come out exactly as they went in, and a file with more than two channels comes apart into one file per channel.

Convert Audio

Turn MP3 into WAV or WAV into MP3, and the same for M4A, OGG and FLAC. Pick a bitrate and see exactly how much smaller the file gets.

Merge Audio

Stitch a batch of recordings into one long file, or into evenly sized parts. Set how long each output runs before a new one starts, and how much silence sits between files.

Fade In & Out

Ease a clip in and out instead of cutting hard, using equal-power gain curves. Set one fade for every file, or give any file its own.

Normalize Volume

Bring a pile of recordings to the same loudness, measured the way broadcasters and streaming services measure it. Each file is checked against a true-peak ceiling on the way, so nothing clips getting there.

Label Speakers

Work out who spoke when across a recording with several people in it, and optionally transcribe it at the same time so every word is attached to whoever said it. You get a timeline and a transcript you can correct — rename the speakers, move any word that landed on the wrong one, listen to any turn — and export it as RTTM, a subtitle track, a spreadsheet or a speaker-labelled script. Runs entirely on your own device.