The same line, four ways — raw, cleaned up, synthetic and cloned.
The full script, word for word — 546 words, about 2 minutes to read. Current as at August 2026. AI tools change quickly; if something looks different when you try it, check the product's own help pages.
Hi, I'm Mel. The voice you hear in these videos is synthetic. That gives us one route to a voice track. We tested one short tour announcement as a raw recording, the recording after a text-based cleanup, a public synthetic voice, and a voice clone. You can hear the trade before I explain it.
The raw take is the starting point. The cleaned, synthetic and cloned versions use the finished wording: a welcome, the departure time and the check-in instruction. The comparison is close, though the raw take still contains its original fillers and restart. Listen for the pauses, the emphasis and whether the delivery feels connected to the sentence.
The Fern Hollow Lion Tour. Um, tours leave at twelve thirty every Saturday and Sunday. Please arrive at, um... Sorry, please check in fifteen minutes early.
Welcome to the Fern Hollow Lion Tour. Tours leave at twelve thirty every Saturday and Sunday. Please check in fifteen minutes early.
The raw file runs about seventeen seconds. The cleaned file runs about twelve. In Descript, the recording became editable text. Removing the two ums removed their audio. Cutting the restart and shortening the gaps tightened the delivery. The transcript heard Line; correcting it to Lion repaired the text while the spoken word stayed in place. Sound enhancement is a separate step, and this test does not claim it was applied. The performance still belongs to the recorded take.
Welcome to the Fern Hollow Lion Tour. Tours leave at twelve thirty every Saturday and Sunday. Please check in fifteen minutes early.
Welcome to the Fern Hollow Lion Tour. Tours leave at twelve thirty every Saturday and Sunday. Please check in fifteen minutes early.
The public synthetic voice is the quickest route. It reads new text without another recording, and its delivery comes from the selected voice and controls. A clone aims to keep the source voice while gaining the same revision advantage. Both generate a fresh performance from text. The cleaned recording keeps the pace, emphasis and intention from the take that was recorded. That is why cleanup deserves a test before you replace the speaker.
Use Descript when you are willing to record once and want to edit speech like a document. HeyGen combines public voices and personal voice clones with video production. ElevenLabs is a dedicated voice route with text-to-speech and cloning. A voice built into an editor is convenient for a short piece. A microphone and a quiet room remain a strong choice when the person carries the message. Choose by the job, then preview one sentence.
A voice is part of someone's identity. Clone your own, or use another person's only with explicit permission. HeyGen says it verifies ownership and requires written consent for a third-party voice. Licensing also changes by service and plan. ElevenLabs says its free plan does not include commercial use, and shared non-commercial output requires attribution. Check the terms that applied when the audio was generated. Disclose a synthetic or cloned voice when that fact could change trust, consent or interpretation.
Begin with the message. Use a synthetic voice for speed and repeatability, clean a real recording when the delivery carries meaning, and clone only with permission. Test the same sentence in two routes. Choose whichever preserves what the viewer needs.