Voice Clone.
One voice, every line.
A voice clone tool I built to run entirely on my own Mac. Give it a few seconds of someone talking, type what you want them to say, and it hands back finished takes in that voice.
It exists for the things every edit needs: pickups, scratch VO, the line that has to change after the session wrapped. Every take renders at 48 kHz native as a numbered WAV with a manifest, straight to the timeline. No booth, no waiting on a recording, nothing uploaded anywhere.
How it works
Drop a sample
Five to fifteen seconds, one speaker, no music. A waveform preview shows exactly what the model is about to hear.
Say what it says
Type what the sample says, or hit Auto-transcribe. The model aligns to those words, which is what keeps the match tight.
Write the script
One line per render. Commas breathe, periods land, ellipses trail. Lines starting with # are skipped, and ⌘↵ renders.
Direct the performance, not just the words.
Choose how many takes you want per line (1, 3 or 5), then set the read: Tight sticks to the reference, Natural is the default, and Loose takes more chances. On top of that you steer the delivery with an expression preset.
Neutral uses continuation cloning, the closest match to the reference. The others trade a little of that for performance.
Four engines
Studio is the one you want for finals. Fast is for scratch. Everything else sits between them, and the language menu covers more than English.
best quality
quicker
Pacing, written into the script
Timing lives in the text, so a script reads like a script. Drop a code in a line where you want air.
[pause 1.2]a 1.2 second pause[beat]a short beat…trails off# noteskipped, for your own notes
Don’t like the read? Re-take the line.
Each line lands with its own stack of takes and a waveform for each one. Audition them, star the keepers, and if a line still isn’t right hit Redo line to re-roll just that one, without re-rendering the whole script.
Download line grabs one, Download all grabs everything, and Download starred grabs only your keepers, all in script order. Reveal in Finder jumps to the render folder.
Renders are kept in a history menu, so yesterday’s session, stars included, is one click away.
Under the hood
It runs on Apple Silicon with mlx-audio, VoxCPM2 and Qwen3-TTS for voices, mlx-whisper for transcription, and a small FastAPI app on 127.0.0.1 that only listens on this machine. The model stays warm between renders, so re-takes come back quickly. Voices, scripts and takes never leave the Mac.
Want to see it in action?
It’s not a public download, so the best way to see it is live. Reach out and I’ll walk you through it.