Back to projects January 2023 - October 2024

RVC v2 Voice Model Training Pipeline

A training and release pipeline built on the RVC v2 framework, used to train and publicly release more than ten voice models.

  • PyTorch
  • Audio ML
  • Python
  • GPU Training
RVC v2 Voice Model Training Pipeline

Using RVC v2, an existing open source retrieval based voice conversion framework, I trained and publicly released more than ten voice models. The architecture and the retrieval method are not mine. What I built is the Python pipeline around them: sourcing and cleaning audio, preparing datasets, running GPU training, and running inference. I worked on this with a team of six.

Pipeline

Raw recordings are segmented, denoised, and passed through pitch and F0 extraction so the training data stays consistent. Training runs in PyTorch on the GPU, and each finished model is paired with a FAISS retrieval index, the RVC component that helps a converted voice hold the detail and character of the target speaker. Inference then converts a source recording into the target voice using the trained model and its index. Batch workflows let the team prepare large amounts of audio and train several models at once without repeating manual steps.

What I took from it

Handling audio at scale, keeping training runs reproducible, and tuning models for output quality were the core of the work.

Hear it in action

The same clip before and after conversion with a trained voice model.

Before Original recording
After Converted with the model