Docs

User guide

Read one chapter at a time from the left menu. Steps match the live UI. Screenshots are from the English product interface.

Written for the current product UI · follow step by step · static pack

09 / 193 min

Audio nodes

Audio nodes cover voice and sound: text to speech, voice cloning and design, and the audio-processing tools that go with video.

Text to speech

From the Add node menu open Audio → Speech to create a text-to-speech node:

  • Type the text to read aloud, pick a voice and parameters, and generate a clip;
  • The bar has rate, pitch and volume prosody sliders and pause / emphasis markup helpers (some are model extensions; a model that doesn't support them reads the markup as plain text);
  • Text length is capped at about 50,000 characters, a soft on-screen counter.
Open Audio from the Add-node menu to create a speech node
Open Audio from the Add-node menu to create a speech node

Voice cloning and design

  • Voice cloning: upload a reference clip to register a custom voice you can then synthesize with;
  • Voice design: design a new voice from a description, without a reference clip;
  • Registering a voice charges a registration fee and runs asynchronously in the background; a voice that is Processing has to become Ready before you can use it.

Many audio capabilities actually live on the video node's toolbar (see Video nodes): split A/V, mix audio, vocal separation and the like all pull sound out or replace it, mostly for free.

TTS has prosody sliders and pause/emphasis markup · Cloning / designing a voice has a fee and becomes ready async · Split/mix and other audio tools are on the video node toolbar