MSRX Tools

Change Speed

Play a recording faster or slower, with the voice left alone.

Everything runs inside your browser. Your files never leave your device.

Drop your file here

.mp3, .wav, .m4a, .aac, .ogg, .oga, .opus, .flac, .webm, .weba · they stay on this device

Ask about Change Speed

Questions about what this tool does, which option to pick, or what it can and cannot handle.

This is the one part of the site that uses a server. The question you type here is sent to an AI provider to be answered — your files and whatever you put in the tool above are not, and the assistant cannot see them. Answers are generated and can be wrong; the tool itself is not guessing.

About the Change Speed

Playing something faster used to mean making everybody sound like a cartoon. Speed and pitch were welded together, because the only way to make a tape finish sooner was to pull it past the head faster, and that raised every frequency on it.

They come apart here, and that is the whole point of the tool: you change one without changing the other. Leave "keep the original pitch" on and a lecture at 1.5× finishes in two thirds of the time with the lecturer still sounding like themselves. Turn it off and you get the tape behaviour, which is genuinely what you want occasionally — for an effect, or to hear detail in something slowed right down.

The technique behind the first one is called WSOLA, and it is worth knowing what it does because it explains the limits. The recording is cut into short overlapping grains, and the grains are laid back down at a different spacing: closer together to speed up, further apart to slow down. Done naively this warbles horribly, because consecutive grains land out of phase with each other. WSOLA fixes it by sliding each grain a few milliseconds to wherever it best matches the waveform that would naturally have followed the previous one. That similarity search is the entire trick, and it is why a stretched voice still sounds like a voice.

It is not magic. Somewhere past about half speed or double speed the grains start to be audible on complex material — a full mix will develop a slight flutter that a single voice will not. Speech survives much further than music does, which is convenient, because speech is what most people are speeding up.

Everything runs in the browser on samples it decoded itself. A recording of a meeting you are trying to get through faster is not a recording you should be uploading to a stranger's server to achieve it.

How to use it

  1. 1Drop in the recording.
  2. 2Set the speed — 1.5× and 2× are the usual choices for talk, 0.5× for transcribing something difficult.
  3. 3Leave the pitch lock on for anything with a voice in it; turn it off for the tape effect.
  4. 4Run it and download.

Questions

Will a sped-up voice sound like a chipmunk?
Not with the pitch lock on, which is the default. The speaker keeps their own pitch and simply talks faster. The chipmunk sound is what you get with it off, which is there because some people want exactly that.
How fast can I go before it sounds bad?
Speech holds up well to about 2×, and many people find 1.5× perfectly natural. Music is fussier — a full mix starts to flutter noticeably sooner, because there is more happening for the grain joins to disturb.
Does slowing down add detail?
No. It spreads the existing detail over more time, which genuinely helps when you are transcribing a mumbled sentence or learning a fast passage, but no information is created. What was not captured is still not there.
Why is the result not exactly the length I calculated?
The grains are laid down at whole-sample spacings, so the final length can land a fraction of a second either side of the arithmetic. The result panel shows the length it actually produced.