Get in Touch
 Duration 14 hours

Course Outline

Introduction to Speech Synthesis and Voice Cloning

  • Overview of text-to-speech (TTS) and neural voice synthesis
  • Distinction between voice cloning and speech generation: use cases and limitations
  • Key models: Tacotron, WaveNet, FastSpeech, VITS

Utilising Commercial Platforms

  • Working with ElevenLabs and Resemble AI
  • Voice creation, cloning, and editing workflows
  • API access and text-to-speech implementation

Developing with Open-Source Tools

  • Installation and configuration of Coqui TTS
  • Training custom voices and managing datasets
  • Generating speech with precise control over pitch, speed, and emotion

Data Preparation and Voice Dataset Management

  • Collection and cleaning of voice samples
  • Segmentation, labelling, and transcript alignment
  • Ethical sourcing and voice consent management

Application Integration

  • Embedding TTS capabilities in websites and applications
  • Designing IVR systems and interactive bots
  • Generating synthetic dialogue for video and gaming content

Assessing Quality and Realism

  • MOS (Mean Opinion Score) and intelligibility testing
  • Managing expressiveness and prosody
  • Comparing latency, fidelity, and overall realism

Ethical, Legal, and Governance Considerations

  • Deepfake risks and responsible usage protocols
  • Consent, attribution, and copyright implications
  • Regulatory frameworks and organisational policies

Summary and Next Steps

Requirements

  • A solid understanding of machine learning fundamentals
  • Proficiency with audio file formats and editing software
  • Foundational skills in Python programming

Target Audience

  • AI developers and engineers focused on speech synthesis
  • Content creators and media technologists exploring voice generation
  • R&D teams developing personalised or dynamic audio systems

Number of participants


Price per participant

Provisional Upcoming Courses (Require 5+ participants)

Related Categories