RVC v2 Voice Model Training Pipeline
A training and release pipeline built on the RVC v2 framework, used to train and publicly release more than ten voice models.
Using RVC v2, an existing open source retrieval based voice conversion framework, I trained and publicly released more than ten voice models. The architecture and the retrieval method are not mine. What I built is the Python pipeline around them: sourcing and cleaning audio, preparing datasets, running GPU training, and running inference. I worked on this with a team of six.
Pipeline
Raw recordings are segmented, denoised, and passed through pitch and F0 extraction so the training data stays consistent. Training runs in PyTorch on the GPU, and each finished model is paired with a FAISS retrieval index, the RVC component that helps a converted voice hold the detail and character of the target speaker. Inference then converts a source recording into the target voice using the trained model and its index. Batch workflows let the team prepare large amounts of audio and train several models at once without repeating manual steps.
What I took from it
Handling audio at scale, keeping training runs reproducible, and tuning models for output quality were the core of the work.
Hear it in action
The same clip before and after conversion with a trained voice model.