Исследование технологий для задачи клонирования голоса и методов их улучшения тема диссертации и автореферата по ВАК РФ 00.00.00, кандидат наук Садекова Таснима Равилевна
- Специальность ВАК РФ00.00.00
- Количество страниц 236
Оглавление диссертации кандидат наук Садекова Таснима Равилевна
1.5 Клонирование голоса
1.6 Дополнительный контроль синтеза и эффективность
1.7 Безопасность
1.8 Фокус данной работы
2 Диффузионная мультимодальная архитектура
2.1 Предварительные эксперименты
2.1.1 Базовые модели
2.1.2 Энкодер голоса
2.1.3 Модификации для клонирования голоса
2.1.4 Экспериментальные результаты
2.2 Основные понятия
2.2.1 Диффузионные модели
2.2.2 Копирование голоса
2.3 Описание предложенной мультимодальной системы
2.3.1 Мел-энкодер ф
2.3.2 Текстовый энкодер ф
2.3.3 Декодер
2.3.4 Вокодер
2.4 Сценарии использования
2.5 Похожие модели
2.6 Экспериментальные результаты
2.6.1 Модели для сравнения
2.6.2 Первый эксперимент
2.6.3 Второй эксперимент
3 Улучшение моделей клонирования голоса. Ускорение синтеза
3.1 Введение
3.1.1 Существующие подходы к ускорению диффузионных моделей
3.1.2 Преимущества и недостатки GAN и диффузионных моделей
3.1.3 Преимущества исследуемого метода
3.2 Описание модели
3.2.1 Базовая архитектура
3.2.2 Диффузионная генеративно-состязательная сеть
3.3 Модели для сравнения
3.4 Экспериментальные результаты
4 Улучшение моделей клонирования голоса. Контроль эмоций
4.1 Свойства векторов голоса
4.2 Описание модели
4.3 Метод
4.3.1 Алгоритм выбора компонент
4.3.2 Оптимизация гиперпараметров
4.4 Экспериментальные результаты
4.4.1 Проверка метода
4.4.2 MOS тест
Заключение
Список таблиц
Список изображений
Аппендикс А. Английский перевод диссертации
Введение
Тема диссертации
Современные модели синтеза речи позволяют генерировать записи хорошего качества при достаточном количестве обучающих данных. В зависимости от количества дикторов в данных есть возможность синтезировать речь одним или различными голосами. Для этого не обязательно обучать несколько моделей отдельно. Достаточно объединить обучающие данные и добавить небольшие архитектурные изменения. Одним из возможных вариантов является добавление обучаемой таблицы, в которой каждый вектор сопоставляется одному голосу. Этот вектор будет извлекаться по соответствующей метке диктора и передаваться модели. Однако такой способ позволяет выбирать голоса только из обучающего датасета. В некоторых же сценариях возникает необходимость добавления голоса, не присутствующего в обучающей выборке, при наличии небольшого количества аудиозаписей. Такая задача называется клонированием голоса (voice cloning). Существует два основных подхода - кодирование (speaker encoding; zero-shot cloning) и адаптация голоса (speaker adaptation; few-shot cloning) [1, 2].
Кодирование голоса подразумевает использование вектора, содержащего в себе основные характеристики голоса. Он извлекается из короткой записи и передается как дополнительный вход при синтезе. Такой вектор получают из специального блока - энкодера диктора, который можно обучать двумя способами: совместно с основной моделью либо отдельно на некоторой другой задаче, например, верификации голоса. Последний вариант показал более хорошие результаты ([1]), так как для отдельно обучаемой модели можно использовать больше данных без строгих требований к качеству и наличию текстов, необходимых для задачи синтеза речи. Это позволяет получить более разнообразные репрезентации голосов. Во время генерации речи достаточно взять короткую 5-секундную запись произвольного голоса для клонирования. Такой низкий порог по количеству данных и быстрое извлечение представления голоса являются большим преимуществом данного подхода. Однако в некоторых случаях одного вектора может быть недостаточно, чтобы передать важные особенности диктора.
Второй подход, адаптация голоса, подразумевает дообучение всей или части модели на записях нового диктора. Требование к количеству записей нужного голоса здесь значительно выше, а также необходимы ресурсы для обновления параметров. В то же время, такой вариант клонирования дает наиболее хорошие результаты по похожести.
Помимо базовых требований к качеству синтезируемой речи, важными критериями современных систем генерации являются их эффективность и управляемость.
С точки зрения эффективности, ключевым параметром выступает скорость синтеза, которую количественно оценивают с помощью коэффициента реального времени (Real-Time Factor, RTF). Данный показатель, рассчитываемый как отношение времени работы к длительности результирующего аудиосигнала, напрямую влияет на применимость модели. Оптимальным считается значение меньше 1, гарантирующее синтез не медленнее реального времени. Для интерактивных приложений, таких как голосовые помощники или диалоговые системы, чем ниже этот показатель, тем лучше.
Наряду с производительностью, важной характеристикой является возможность управления какими-либо атрибутами выходного сигнала. Поскольку задача генерации речи из текста относится к типу «один ко многим» (один текст может быть произнесен множеством различных способов), наличие дополнительных механизмов контроля повышает определенность процесса синтеза. В данной работе, например, рассматривается управление эмоциональной окраской речи, что позволяет получать более разнообразные записи и расширяет область применения системы.
Рекомендованный список диссертаций по специальности «Другие cпециальности», 00.00.00 шифр ВАК
Система разделения дикторов на основе вероятностного линейного дискриминантного анализа2014 год, кандидат наук Кудашев, Олег Юрьевич
Методы и средства выделения шепота в полилоге и конверсии шепотной речи в нормальную2025 год, кандидат наук Авдеева Анастасия Сергеевна
Модель, численная и программная реализация оценивания частоты основного тона речевого сигнала с помощью сингулярного спектрального анализа2015 год, кандидат наук Вольф Данияр Александрович
Сжатие речевых данных на основе субполосного анализа и синтеза речевых сигналов в области определения их косинус-преобразования2021 год, кандидат наук Трубицына Диана Игоревна
Генерация мимики и жестов по речи2022 год, кандидат наук Корзун Владислав Андреевич
Введение диссертации (часть автореферата) на тему «Исследование технологий для задачи клонирования голоса и методов их улучшения»
Актуальность работы
Задача клонирования голоса представляет собой одно из активно развивающихся направлений в области синтеза речи. Данная технология имеет широкий спектр потенциальных применений, включая следующие сценарии:
• упрощение создания сервисов генерации речи, использующих один голос. Так, например, для обучения модели синтеза для голосового ассистента требуется собрать многочасовой датасет разнообразной речи, записанный подготовленным диктором. Клонирование голоса позволяет значительно снизить эти
требования и использовать гораздо меньшее количество записей, адаптируя на них модель, предварительно обученную на многоголосых данных.
• перевод речи с языка на язык с сохранением голоса. Этот сценарий может использоваться для перевода образовательного и развлекательного видео контента, а также для дубляжа фильмов (при условии получения согласия от авторов или актеров). Технология позволяет расширить аудиторию, делая информацию более доступной, при этом сохраняя личность и уникальные особенности голоса исходного диктора и минимизируя потери при адаптации материала. Модели клонирования голоса при переводе видео уже внедрены в ряд платформ, таких как Яндекс Браузер и Уо^иЬе.
Кроме того такая задача может быть важна для устройств-переводчиков в путешествиях, позволяя клонировать голос пользователя и обеспечивая таким образом более естественное межъязыковое общение. В данном и подобных сценариях, когда возникает необходимость в адаптации под свой голос, полезным вариантом является возможность локального дообучения на пользовательском устройстве. Такой подход не требует передачи персональных аудиоданных на серверы, что гарантирует сохранение конфиденциальности.
• более быстрый способ создания аудио и видео контента. Так аудиокниги являются одним из способов получения полезной информации, однако количество книг, озвученных профессиональными чтецами, ограничено. Клонирование голоса открывает возможность для разработки сервисов, которые в реальном времени озвучивают книги с возможностью выбора из каталога предзапи-санных голосов для дополнительной персонализации. Подобные решения не только расширяют доступ к аудиокнигам, но и служат важным инструментом адаптации контента для людей с нарушениями зрения. Примером реализации данной идеи является сервис Воокша1е.
Схожей задачей является озвучка видео и анимационных фильмов. Модели синтеза и клонирования способны значительно ускорить этот процесс. Например, профессиональному диктору может быть достаточно записать выборку реплик для каждого персонажа, чтобы адаптировать модель под его голос, после чего весь остальной текст будет сгенерирован автоматически. При этом
выбранная модель должна обеспечивать не только разборчивость, но и естественную просодию, эмоциональную выразительность, а также способность передавать смысловые нюансы текста, что требует дальнейшего развития и улучшения технологий генерации.
• воссоздание голоса для людей, утративших способность говорить. На основе сохранившихся аудиозаписей модели клонирования позволяют восстановить индивидуальные голосовые характеристики человека.
• расширение голосового разнообразия для низкоресурсных языков. Для таких языков существует ограниченное количество записанных аудиоданных, что сужает выбор голосов для систем синтеза речи. Методы клонирования голоса позволяют решить эту проблему, обеспечивая возможность переноса голосовых характеристик из записей на других языках.
• аугментация данных для решения различных задач обработки речи. Сгенерированные аудио могут использоваться для расширения и обогащения наборов данных при обучении других речевых моделей.
Однако важно отметить, что технология клонирования голоса, помимо полезных применений, поднимает серьезные вопросы безопасности. Злоумышленники могут использовать синтезированный голос для мошенничества, обхода биометрической защиты или распространения дезинформации, что делает необходимость регулирования этой области особенно актуальной.
С этической точки зрения проблема затрагивает право человека на собственный голос, который признается уникальной биометрической характеристикой, обеспечивающей не только коммуникацию, но и личную идентификацию. В связи с этим возникает серьезная задача обеспечения конфиденциальности голосовых данных и защиты их от несанкционированного сбора, хранения и использования. Центральным этическим требованием становится получение информированного и осознанного согласия на любые манипуляции, включая запись, обработку и синтез. Однако для надежной защиты необходимы не только этические нормы, но и четкие, адаптированные правовые регуляторы, которые устанавливают границы допустимого использования технологии и предусматривают эффективные механиз-
мы ответственности за противоправные действия, такие как мошенничество или клевета.
С технической стороны параллельно с развитием моделей клонирования должны совершенствоваться и защитные механизмы, такие как системы верификации голоса и выявления синтезированных записей ([3, 4, 5, 6]). Они позволяют проводить дополнительную проверку в потенциально уязвимых ситуациях. Еще одним методом защиты могут служить специальные «водяные знаки»: их можно добавлять как в оригинальные записи, чтобы предотвратить несанкционированное клонирование ([7]), так и внедрять в сгенерированные аудио на этапе их создания, чтобы упростить последующее обнаружение ([8, 9]).
Таким образом, комплексное развитие технологий генерации и методов защиты позволит использовать все положительные возможности клонирования голоса, одновременно повышая безопасность в критически важных сферах.
В современных исследованиях, описывающих модели клонирования голоса, часто используют внутренние многоголосые датасеты большого объема для обучения как базовой модели синтеза, так и энкодера диктора. Кроме того, в них, как правило, не ставится жестких ограничений на объем данных, необходимых для адаптации. Снижение требований к количеству записей целевого голоса, не представленного в обучающей выборке, упростит сценарии использования и ускорит процесс дообучения. В данной работе для обучения основной модели используются исключительно открытые датасеты с разнообразным набором голосов. Для кодирования требуется только одно аудио нового диктора, а для сценария адаптации исследуется возможность переноса голоса с использованием одной минуты данных.
Большое количество современных нейронных архитектур синтеза речи состоит из двух моделей. Первая из них по входному тексту (символам/фонемам) предсказывает некоторые промежуточные признаки, чаще всего спектрограммы [10, 11, 12], вторая же, вокодер, восстанавливает аудио по этим промежуточным признакам [13, 14, 15]. Все исследуемые в данной диссертации модели будут следовать этой структуре.
Любая модель синтеза речи может быть модифицирована под задачу клонирования. Для варианта кодирования голоса достаточно предавать вектор диктора как дополнительный вход в один из блоков модели на этапе обучения. А для адап-
тации необходимо определить, какие части следует обновлять под новые данные. Так вокодер, обученный на многоголосом датасете, не требует обновления при клонировании голоса, что значительно упрощает задачу.
В первой части работы предложена диффузионная модель клонирования голоса. Дополнительно проводилось сравнение с моделями генерации речи, модифицированными для задачи клонирования. Были рассмотрены распространенные модели, основанные на рекуррентных сетях [10, 16], нормализующих потоках [17] и блоках «Трансформер» [11]. Для двух сценариев клонирования изучены оптимальные стратегии адаптации, а также определены минимальные требования к данным и вычислительным ресурсам. Проанализированы особенности каждого подхода с точки зрения устойчивости процесса дообучения и выбора параметров для обновления.
Вторая часть работы описывает методы улучшения некоторых характеристик исследованных моделей. Для диффузионных архитектур проблемным местом является скорость генерации из-за их итеративной природы. Это особенно заметно при использовании центрального процессора (CPU) вместо графического (GPU). Одним из предложенных методов ускорения для генерации изображений является объединение диффузионной модели с генеративно-состязательной сетью (GAN). Этот метод был перенесен на генерацию звука. Также важной характеристикой восприятия речи является ее экспрессивность, а генерация эмоциональных записей является актуальной задачей для современных моделей синтеза. В данной работе рассматривался универсальный метод модификации векторов диктора для добавления 4 эмоций.
В итоге целью данной работы является разработка собственной диффузионной архитектуры клонирования голоса, ее сравнение с другими моделями, а также подбор оптимального количества данных для оптимизации эффективности с сохранением качества. Дополнительное внимание уделяется улучшению таких характеристик, как скорости генерации и эмоциональности сгенерированной речи.
Основные результаты и выводы Вклад.
Ниже сформулированы основные результаты работы.
1. Предложена гибридная диффузионная архитектура, объединяющая в себе две задачи - клонирование и копирование голоса. Такая комбинация имеет некоторые преимущества, позволяя использовать для адаптации не размеченные данные. Более того, требования к количеству данных значительно снижаются, позволяя дообучаться на 15 секундах без значительных потерь в качестве. Данная модель сравнивалась с распространенными архитектурами синтеза речи, модифицированными для задачи клонирования голоса.
2. Исследован метод ускорения генерации диффузионной модели, основанный на объединении с генеративно-состязательными сетями, применительно к речевым моделям. Он сравнивался с рядом других методов. В результате экспериментов было установлено, что данный подход позволяет сократить количество необходимых шагов, однако для сохранения качества синтеза потребовалось увеличить размер модели. Это, в свою очередь, привело к лишь к незначительному повышению скорости генерации при выполнении на CPU.
3. Выявлено, что используемые векторные представления голоса содержат не только параметры, ответственные за индивидуальные характеристики тембра, но и кодируют информацию об эмоциональной окраске речи. На основании данного наблюдения предложен метод целенаправленной модификации векторов, обеспечивающий синтез речи с заданной одной из четырех эмоций (радость, грусть, злость, удивление)
Теоретическая и практическая значимость Результаты, выносимые на защиту
1. Диффузионная архитектура, выполняющая две задачи одновременно: клонирование и копирования голоса. Сравнительный анализ с другими моедлями.
2. Исследование по ускорению генерации речи диффузионной модели путем объединения с генеративно-состязательной сетью.
3. Метод модификации векторов диктора для синтеза эмоциональной речи при клонировании голоса.
Личный вклад в результаты, выносимые на защиту.
В первой публикации (раздел 2) автор диссертации реализовывал предложенную модель, адаптировал другие архитектуры генерации речи под задачу клонирования голоса для более широкого сравнения, обучал основные модели, а также проводил эксперименты по клонированию голоса. Во второй статье (раздел 3) основным вкладом является реализация одного из методов ускорения генерации диффузионных моделей, применимого к предложенной в данной диссертации. В третьей публикации (раздел 4) автор данной работы реализовывал метод поиска эмоциональных компонент для контроля эмоций и отвечал за проведение экспериментов и тестирование.
Публикации и апробация работы
Публикации повышенного уровня:
1. Sadekova T., Gogoryan V., Vovk I., Popov V., Kudinov M., Wei J. (2022) A Unified System for Voice Cloning and Voice Conversion through Diffusion Probabilistic Modeling Конференция Interspeech 2022, 3003-3007, doi: 10.21437/ Interspeech.2022-10879 Конференция ранга A по рейтингу CORE.
2. I.Vovk, T.Sadekova, V.Gogoryan, V.Popov, M.Kudinov, J.Wei "Fast Grad-TTS: Towards Efficient Diffusion-Based Speech Generation on CPU Interspeech, 2022 Конференция ранга A по рейтингу CORE
3. Shaheen, Z*., Sadekova, T.*, Matveeva, Y., Shirshova, A., & Kudinov, M. (2023). Exploiting Emotion Information in Speaker Embeddings for Expressive Text-to-Speech. INTERSPEECH 2023, 2038-2042. Конференция ранга A по рейтингу CORE.
1 Обзор литературы и области
Модели синтеза речи по тексту (text-to-speech synthesis, TTS) предназначены для генерации естественно звучащей речи на основе входного текста. Они являются фундаментом для создания интуитивных и многофункциональных интерфейсов. Повсеместное распространение голосовых ассистентов, интеллектуальных систем управления «умным домом», автомобильных навигационных систем, управляемых речью, наглядно демонстрирует эту тенденцию. Качество и характеристики голоса в синтезированных записях напрямую коррелирует с комфортом и доверием пользователя: натуральная интонация, эмоциональная окраска, отсутствие роботизированных артефактов и правильное смысловое ударение делают взаимодействие с технологией естественным процессом.
Возможность контроля голоса при генерации, доступная благодаря развитию технологий клонирования голоса позволяет расширить сферы применимости. Имея относительно небольшое количество записей интересующего диктора, можно адаптировать модель генерации речи под его голос, копируя особенности произношения и манеру речи. Это открывает двери для персонализированных решений в медиа-индустрии (озвучка контента, аудиокниг), кинопроизводства и создания видеоигр (быстрая генерация реплик для персонажей).
Параллельно с коммерческим и развлекательным применением, технология синтеза и клонирования речи может помочь и в ряде социальных задач. Так для людей с заболеваниями, лишающими возможности говорить, системы синтеза речи, способные использовать их собственный голос, становятся важнейшим средством коммуникации.
Стоит отметить и обратную сторону развития моделей клонирования, которая поднимает вопрос в сфере цифровой безопасности. Эта технология открывает возможности для злонамеренного применения в различных областях и, кроме того, подрывает надежность систем биометрической аутентификации, использующих голос в качестве уникального идентификатора. В связи с этим важным направлением в области обработки речевых сигналов становится интенсивное усовершенствование и разработка моделей, способных различать реальные записи человека и синтезированные генеративной моделью.
Область генерации речи претерпела сильные изменения с момента появления первых идей. Особенно явно прогресс ускорился с развитием глубоких нейронных сетей.
Первые модели представляли собой крайне упрощенные механистические прототипы голосового тракта человека, способные воспроизводить лишь ограниченный набор звуков и простых слов. Более поздние версии позволяли генерировать слова и предложения за счет соединения заранее записанных речевых сегментов (конка-тенативный синтез) или моделирования звуков с помощью специальных правил (формантный синтез). Следующее поколение, статистический параметрический синтез, приблизило технологию к современным решениям. Данный подход предполагал обучение нескольких подмодулей для предсказания речевых параметров, которые затем использовались для восстановления звуковой волны. В этой схеме часто применялись скрытые марковские модели (СММ). Далее с бурным прогрессом в области глубокого обучения за последние десятилетия были разработаны различные эффективные архитектуры нейронных сетей для синтеза речи. Эти модели значительно улучшили качество и естественность синтезированных записей.
Ранние исследования в данной области в основном фокусировались на системах с одним голосом (например, [18, 19, 20]). Однако вскоре возник интерес к моделям, способным синтезировать речь различными голосами ([21]). Обучение в таком режиме снижает требования к объему данных для отдельных дикторов, что более реалистично для открытых датасетов. Кроме того, подобные системы обеспечивают большую универсальность без увеличения вычислительных затрат. Клонирование голоса - это расширение возможностей моделей синтеза речи, позволяющее генерировать речь голосами, отсутствующими в обучающих данных, без полного переобучения.
Следует также отметить две важные темы, тесно связанные как с моделями синтеза речи, так и клонирования голоса:
• Первая касается их эффективности и скорости работы. Так для систем, где генерация речи используется как основной интерфейс взаимодействия, важно, чтобы ответ синтезировался насколько быстро, насколько это возможно. Данное требование побуждает к более аккуратному подбору архитектуры и размера модели. Дальнейшие ограничения по потреблению памяти и воз-
можности запуска на центральном процессоре могут возникнуть в сценарии работы моделей непосредственно на пользовательских устройствах при работе с конфиденциальными данными, что выглядит актуальным в условиях растущей распространенности умных домашних систем и использования мобильных телефонов.
• Вторая тема относится к возможности более гибкого контроля при генерации речи. Это может быть контроль над такими атрибутами как эмоциональная окраска, интонация, ритм, темп. Такая способность открывает значительный потенциал для разнообразных приложений, что является стимулом для развития различных методов и архитектур в последние десятилетия.
1.1 Способы представления аудио
Первым этапом при работе со звуковым сигналом является его преобразование из аналоговой формы в цифровую. Цифровая форма представляет из себя результат дискретизации и квантизации значений исходной звуковой волны, в результате которых она сохраняется в виде последовательности отсчетов амплитуды, отражающих изменения звукового давления во времени (Рис. 1а). Ключевыми параметрами являются частота дискретизации (количество отсчетов в секунду) и битовая глубина (количество бит для кодирования одного отсчета). При правильном выборе частоты, удовлетворяющей теореме Котельникова, такое представление содержит полную информацию и позволяет без потерь восстановить исходный аналоговый сигнал.
Волновая форма является наглядной и интуитивно понятной. Значение амплитуды отображает громкость звука, по самому графику можно определить участки, в которых произносятся некоторые характерные звуки (так можно понять, в каком временном интервале произносятся вокализованные и невокализованные звуки). Однако для описания одной секунды аудио с частотой дискретизации 16 кГц нужно хранить и работать с 16000 числами, что увеличивает требования к вычислительным ресурсам и памяти из-за большого количества отсчетов. Также при незначительном изменении произношения или добавления фонового шума звуковая волна может
сильно изменить свои вид.
(а) Временное представление - звуковая волна (Ь) Частотное представление - спектрограмма Рисунок 1: Способы представления речи.
Наиболее распространенным способом работы с аудио является частотное представление сигнала - спектрограмма (Рис. 1Ь). Звуковая волна может быть представлена в виде суммы более простых волн, например, синусоидальных, с различными частотами. Такое разложение можно выполнить путем применения преобразования Фурье. Рассмотрим процедуру получения частотных представлении поэтапно:
(1) Разделение на окна. Поскольку частотный состав звука меняется со временем, мы не можем применить преобразование Фурье ко всему сигналу сразу, так как с большой вероятностью мы получим ситуацию, в которой все частоты участвуют в формировании сигнала с некоторой мощностью. Поэтому аудио разбивается на небольшие перекрывающиеся сегменты с длиной окна 20-40 мс и с таким сдвигом, чтобы перекрытие составляло 50% ± 10%.
(2) Сглаживание. В результате процедуры разделения на фреймы, на границах полученных отрезков могут возникать частоты, которых нет в исходном сигнале. В связи с этим на данном этапе добавляется операция свертки со сглаживающим оконным преобразованием. Были предложены различные варианты окон, Ханна, Хэмминга, Блэкмана, однако их общей отличительной чертой является свойство плавно ослаблять сигнал к краям окна, что гарантирует близость амплитуды сигнала к нулю в начале и в конце фрейма. Пример
такого окна - окно Хэмминга (Рис. 2):
2 пп
w[n] = 0.54 - 0.46 cos (——-),
N — 1
где 0 < п < N — 1, N - длина окна.
0.8
а.
Е 0.4 <
Hamming
100 Samples
Рисунок 2: Окно Хэмминга
(3) Дискретное оконное преобразование Фурье позволяет прейти от временного представления аудио к частотному и получить разложение на частотные составляющие. К каждому фрейму применяется следующее преобразование:
N-1
xne N ,
n=0
(2)
где N - количество дискретных отсчетов во фрейме (длина окна), хп - значение отсчета, к - индекс по частоте, к = 0,1, ...^ — 1 (на практике значения к можно брать до N из-за симметричности полученных значений).
Из полученного вектора комплексных чисел можно вычислить амплитуду |Х(к)|2 и фазу каждой частотной составляющей сигнала во фрейме. Амплитуду можно изобразить в виде графика спектра, где по оси х - диапазон возможных частот, а по оси у - амплитуда, показывающая интенсивность частот в исходном аудио.
Для анализа изменения частот с течением времени используют спектрограммы. Это изображение, в котором частоты теперь отображаются по оси у, время по оси х, а цвет отображает значение амплитуды текущей частоты в спектре.
Для эффективного вычисления используют алгоритм Быстрого преобразования Фурье, которой позволяет сократить время вычисления с О(М2) до О(М1одМ).
(4) Перевод линейной шкалы частот в логрифмическую. Человеческий слух более чувствителен к низким частотам (речевой диапазон) и менее - к высоким. Поэтому чаще всего ось частот приводят к другому виду, чтобы отобразить эту нелинейность. Например, можно применять барк-шкалу, которая преобразует частоты по специальным правилам, либо наиболее часто используемую мел-шкалу. Она использует набор треугольных фильтров (от 40 до 128), которые собирают значения из различных диапазонов частот. Численно и визуально такое преобразование можно представить следующим образом (Рис. 3):
т = 2595 * ^10(1 + Х), (3)
где т - частота в мелах, f - частота в Гц.
Frequency
Рисунок 3: Окна для перевода в мел-шкалу
(5) Логарифмирование значений. Так как восприятие громкости человеком нелинейное, а скорее логарифмическое, то к значениями амплитуды (мощности) также применяется логарифмическое преобразование.
Полученное представление эффективно сжимает речевой сигнал, сохраняя важную информацию о частотах, является наглядным (так фонемы можно отличить по мел-спектрограмме на основе их формант) и более удобным для обучения нейро-сетевых моедлей. Все модели в данной работе используют логарифмированые мелили барк-спектрограммы.
Обратное преобразование Фурье
N-1
Хп = ^ Хкет, (4)
к=0
позволяет без потерь восстанавливать сигнал из спектра. Однако при вычислении спектрограммы фазовая информация исходного сигнала, как правило, отбрасывается, и учитывается только амплитуда. А так как в современных архитектура для синтеза речи чаще всего работают с мел-спектрограммами как представлениями аудио, то для восстановления информации о фазе обучают специальную модель-вокодер.
1.2 Общая структура моделей синтеза
Процесс генерации речи из текста обычно разделяется на несколько этапов (Рис.
4):
1. Во-первых, из текста извлекается необходимая лингвистическая информация, например, фонемы, положение ударных слогов.
2. Далее модель предсказания акустических признаков преобразует эту информацию в промежуточные акустические представления. Это могут быть мел-спектрограммы ([18, 22, 23, 21]), фонетические постериограммы ([24]) или иные латентные представления речи. Такая задача осложняется тем, что один и тот же текст может соответствовать разным способам и стилям произношения, поэтому модель должна учитывать эту неоднозначность и обеспечивать возможность моделирования разнообразных вариантов.
3. Завершающий этап - работа вокодера, который преобразует акустические признаки в итоговую звуковую волну([13, 25, 14, 26, 15]).
Иногда два последних компонента объединяются, и модель обучается за один проход, напрямую преобразуя входные лингвистические признаки в речь (еп^ 1ю-еп^ ([27, 11, 28]). Подобные системы упрощают процесс генерации, однако обеспечивают меньшую гибкость и встречаются реже.
Рисунок 4: Схема этапов генерации речи.
В последние годы большинство акустических моделей принимают некоторую дополнительную информацию на вход помимо текста. Это может быть вектор диктора, как в моделях клонирования голоса, метка языка или, например, целая референсная аудиозапись, из которой далее копируется стиль речи или эмоция. Такая дополнительная информация смягчает неоднозначность при генерации речи из текста, а также предоставляет дополнительный способ контроля.
Как уже упоминалось ранее, при использовании спектрограммы в качестве представления аудио акустическая модель восстанавливает магнитуду сигнала в частотном спектре, но не фазу. Одним из первых способов восстановления фазы является алгоритм Гриффина-Лима. Это метод, который по заданной спектрограмме итеративно восстанавливает информацию о фазе. На первом шаге применяется произвольная инициализация. Далее на каждом шаге по спектрограмме и текущим значениям фазы выполняется обратное преобразование Фурье для возврата, за которым следует прямое преобразование, возвращающее новое значение фазы, более близкое к реальному. Это простой алгоритм, не требующий никакого обучения, однако в синтезированных таким образом аудио остается много артефактов.
Похожие диссертационные работы по специальности «Другие cпециальности», 00.00.00 шифр ВАК
Математические модели импедансного типа в теории речеобразования и обработке речевых сигналов2016 год, кандидат наук Любимов, Николай Андреевич
Автоматическая оценка качества речевых сигналов для систем голосовой биометрии и антиспуфинга2022 год, кандидат наук Волкова Марина Викторовна
Устойчивость моделей глубокого обучения2026 год, кандидат наук Корж Дмитрий Сергеевич
Цифровая обработка изображений динамических сонограмм для нейтрализации спектральных искажений речевой информации2014 год, кандидат наук Алюшин, Виктор Михайлович
Синтез, анализ и практическая реализация алгоритмов распознавания и предобработки речевых сообщений2013 год, кандидат наук Выборнов, Сергей Владимирович
Список литературы диссертационного исследования кандидат наук Садекова Таснима Равилевна, 2026 год
Список таблиц
1 Шкала MOS для оценки качества звука и естественности речи. ... 48
2 Шкала MOS для оценки похожести голоса с эталонной записью (4-балльная)....................................................................49
3 MOS тест. Похожесть голоса ("П"), естественность ("Е") и качество звука ("К") моделей клонирования голоса................................53
4 MOS тест на 10 дикторах из LibriTTS. Похожесть голоса ("П"), качество звука ("К") ........................................................53
5 Предложенная модель (A) в сравнении с [115] (B) для клонирования/копирования голоса на нетранскрибированных данных..........74
6 Предложенная модель (A) в сравнении с [116] (B) для клонирования голоса на нетранскрибированных данных ..............................74
7 Копирование голоса на VCTK..............................................76
8 Клонирование голоса на VCTK............................................76
9 Клонирование голоса на внутреннем датасете ..........................77
10 MOS в зависимости от количества данных для адаптации..............78
11 Сравнение разных методов ускорения диффузий......................93
12 Субъективный тест на Amazon Mechanical на предпочтение базовой модели либо эмоциональной..............................................102
13 Субъективный тест на платформе Appen: Похожесть голоса (SMOS),
Качество звука (MOS) и точность распознавания эмоций. Значение с (*) получено при отдельном запуске на настоящих записях. EM (emotion modification) - модель с предложенной модификацией вектора диктора.................................... 105
1 MOS scale for assessing sound quality and speech naturalness...... 167
2 Mean Opinion Score (MOS) scale for evaluating voice similarity against reference recordings (4-point scale)..................... 169
3 Speaker similarity ('Sim"), naturalness ("Nat.") and sound quality ("Sound") of VC models........................... 173
4 Mean Opinion Score ("MOS") and Similarity of VC models. Test on 10 speakers from LibriTTS .......................... 173
5 Our model (A) compared with [115] (B) for voice cloning/voice conversion
on untranscribed data ......................................................189
6 Our model (A) compared with [116] (B) for voice cloning on untranscribed data ............................................................................191
7 Voice conversion comparison on VCTK......................................192
8 Voice cloning comparison on VCTK........................................192
9 Voice cloning comparison on the internal dataset..........................193
10 MOS depending on the amount of adaptation data........................193
11 Comparison of various models ..............................................207
12 Preference test on Amazon Mechanical Turk..............................215
13 Subjective tests on Appen: Speaker similarity (SMOS), Sound Quality (MOS) and Emotion recognition accuracy. Score with (*) was obtained
in a separate test run with emotional recordings............................218
Список иллюстраций
1 Способы представления речи....................... 16
2 Окно Хэмминга .............................. 17
3 Окна для перевода в мел-шкалу..................... 18
4 Схема этапов генерации речи....................... 20
5 Клонирование голоса. Обучение и генерация для кодирования (zero-shot) и адаптации (few-shot) голоса.................... 26
6 Схематичное изображение трех исследуемых базовых моделей и их модификация для задачи клонирования голоса............ 36
7 Вокодер LPCNet.............................. 42
8 Выравнивание текста и аудио ...................... 43
9 Энкодер голоса............................... 43
10 Процесс формирования данных при обучении энкодера диктора . . 45
11 Пример интерфейса страницы для асессоров на платформе AMT, использовавшейся для сбора оценок качества аудио...........51
12 Пример интерфейса страницы для асессоров на платформе AMT, использовавшейся для сбора оценок похожести голоса аудио и выбора эмоции (для секции 4) ........................... 52
13 Частые ошибки по моделям. Субъективная оценка по 25 фразам 10 оценщиками. Все модели были адаптированы на одном и том же дикторе из LibriTTS. -0 кодирование голоса (zero-shot); -N адаптация
в течении N итераций........................... 54
14 Общая схема диффузионных моделей (в дискретном виде)...... 55
15 Предложенная мультимодальная система. Среднее X* априорного распределения Хф = ф(Х0) получено или в режиме копирования голоса (вход - мел-спектрограмма Х0) или Х^ = ^(T) в режиме клонирования голоса (вход - текст T).................. 60
16 Алгоритм вычисления «усредненной по голосу» мел-спектрограммы X. 61
17 Мел-энкодер ................................ 63
18 Текстовый энкодер............................. 63
19 Диффузионный декодер sв. Принимает на вход: зашумленную до шага t спектрограмму Xt, сам момент времени t, «усредненную по голосу» мел-спектрограмму Xt, Y - запись целевого голоса...... 65
20 Вокодер HiFi-GAN. 4 ........................... 67
21 Мультимодальные архитектуры, участвовавшие в сравнении. Слева -модель NAUTILUS [115], справа - модель из работы [116]...... 69
22 Базовые модели генерации речи, адаптированные для задачи клонирования голоса и участвовавшие в сравнении.............. 72
23 Диффузионная генеративно-состязательная сеть для генерации речи. 86
24 Прогрессивная дистилляция в два этапа................ 88
25 Базовая схема LSGM для L=1 данные преобразуются в латентное пространство с помощью энкодера q(z0|x), а диффузионный процесс применяется в латентном пространстве (z0 ^ z1). Изображение из исходной статьи [127]. В данной работе z0 обозначается как z0, а zi
как 01.................................... 90
26 Архитектура модели............................ 96
27 LDA проекция векторов диктора, обеспечивающая наилучшее разделение между различными классами эмоций. a) Нейтральная vs. Злость; b) Нейтральная vs. Радость; с) Нейтральная vs. Удивление;
d) Нейтральная vs. Грусть........................ 98
28 Модель для сравнения Tacotron-GST................... 103
1 Audio representations............................ 140
2 Hamming Window..............................141
3 Filter bank on a Mel-Scale......................... 142
4 Text-to-speech pipeline............................ 144
5 Voice cloning. Training and inference for zero-shot and few-shot scenarios. 149
6 A schematic architectures of three studied models and their modification
for voice cloning. .............................. 157
7 Vocoder LPCNet.............................. 162
8 A text-audio alignment example...................... 163
9 A speaker encoeder ............................. 163
10 Batch construction during the speaker encoder training.......... 165
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
Example of the web page showed to assessors on AMT for naturalness evaluation.................................. 170
Example of the web page showed to assessors on AMT for speaker similarity evaluation and additional emotion selection for Section 4.4.2) . 171 Typical errors per model. Subjective evaluation by experts (25 phrases, 10 assessors). All models are adapted to the same speaker from LibriTTS. zsl stands for non-adapted zero-shot models; fsl-N stands for models
fine-tuned on the target speaker data for N iterations.......... 174
General scheme of diffusion process (discrete representation)....... 175
The proposed multimodal system. The mean X* of the prior in this DPM is either X^ = ^(X0) in voice conversion mode when the source spectrogram X0 is the input or X^ = ^(T) in voice cloning mode when the source text T is the input. The decoder conditioned on the trajectory Y of the reference mel-spectrogram Y under the forward diffusion is
trained and fine-tuned for the prior whose mean is X^.......... 178
"Average voice" mel-spectrogram X extraction algorithm........ 179
Mel encoder .................................. 181
Text encoder .................................. 181
Diffusion decoder sg. Inputs: a spectrogram Xt noised to step t, the time step t itself, a voice-averaged mel-spectrogram Xt, and Y - the target
voice recording................................ 182
Vocoder HiFi-GAN. 5............................ 184
The multimodal architectures included in the comparison. On the left -the NAUTILUS model [115], on the right - the model from the paper [116] 186 Baseline speech generation models adapted for the voice cloning task and
included in the comparison. ........................ 190
Denoising diffusion GAN for text-to-speech. ................ 201
Progressive distillation in two steps..................... 203
The base LSGM scheme for L = 1 involves transforming data into a latent space using an encoder q(z0|x), after which the diffusion process is applied within the latent space (z0 ^ z1). The image is taken from the original paper [127]. In this work, z0 is denoted as Zq, and Zi as zQ. . . 205
26 Model architecture.............................. 210
27 LDA projections of speaker embeddings giving the best separation between classes of emotions in ESD. a) Neutral vs. Angry; b) Neutral vs. Happy;
c) Neutral vs. Surprise; d) Neutral vs. Sad................ 212
28 Model for comparison Tacotron-GST.................... 216
Аппендикс А. Английский перевод диссертации
Introduction Topic of the thesis
Modern text-to-speech (TTS) synthesis models can generate high-quality audio when provided with sufficient training data. Depending on the number of speakers in the dataset, it is possible to synthesize speech with either a single voice or multiple voices. In the latter case, it is not necessary to train several models separately for each speaker. Instead, it is sufficient to combine the training data and make minor architectural changes. One possible option involves adding learnable lookup table in which each vector corresponds to a specific speaker. During inference, the model retrieves the appropriate vector based on a label from a set of seen voices. However, in some cases, there is a need to add a speaker absent in training data, given a small amount of audio recordings. This task is known as voice cloning (VC). There are two main approaches: speaker encoding (zero-shot cloning) and speaker adaptation (few-shot cloning) [1, 2].
Zero-shot voice cloning relies on speaker embeddings - vector representations that contain key vocal characteristics. These vectors are extracted from a short audio sample and used as an additional features during synthesis. For this purpose a special block, called the speaker encoder, is trained, either jointly with the synthesis model or separately on another task (e.g., speaker verification). Research has shown that independently trained speaker encoders yield superior results [1], as they can leverage larger datasets without the strict quality and transcription requirements typical of speech synthesis tasks. This results in more diverse voice representations. During speech generation, a short 5-second audio clip of any voice is sufficient for cloning. Such low data requirement and fast speaker embedding extraction process are significant advantages of speaker encoding. However, a single vector may not always fully capture individual speaker characteristics.
In contrast, the second approach, few-shot voice cloning, involves fine-tuning TTS model (or subset of its parameters) on recordings of a target speaker. The data
requirement for a new voice is significantly higher, and parameters update is resource-consuming. At the same time, this cloning method provides the best results in terms of similarity and copying specific pronunciation nuances.
In addition to basic speech quality requirements, key criteria for modern speech synthesis systems are their efficiency and controllability.
From an efficiency perspective, a critical parameter is synthesis speed, which is quantitatively measured by the Real-Time Factor (RTF). This metric, calculated as the ratio of processing time to the duration of the resulting audio, directly impacts the model's practical applicability. An RTF below 1 is considered optimal, ensuring synthesis is no slower than real time. For interactive applications such as voice assistants or dialogue systems, the lower this value, the better.
In addition to performance, a key feature is the ability to control specific attributes of the output signal. Since text-to-speech generation is a ¡¡one-to-many¿¿ mapping problem - where a single text can be spoken in many different ways - having additional control mechanisms increases the predictability and consistency of the synthesis process. In this study, for example, we examine the control of emotional tone, which not only enriches output diversity but also broadens the system's potential applications.
Relevance
The task of voice cloning represents one of the actively developing areas in speech synthesis. This technology offers a wide range of potential applications, including the following scenarios:
• simplifying the development of speech generation services based on a single voice. For instance, training a synthesis model for a voice assistant traditionally requires collecting hours of diverse speech recorded by a professional voice actor. Voice cloning can significantly reduce these requirements, enabling adaptation of a model pre-trained on multi-speaker data using only a small number of target speaker recordings.
• voice-preserving speech-to-speech translation. This application is suitable for translating educational and entertainment video content, as well as for film dubbing,
subject to obtaining consent from original creators or performers. The technology aids in expanding audience reach by enhancing content accessibility, while preserving the original speaker's identity and vocal characteristics and minimizing losses during adaptation. Voice cloning models for video translation have already been deployed on platforms such as Yandex Browser and YouTube.
Furthermore, this capability is also applicable to portable translation devices used in travel. It allows users to clone their own voice, facilitating more natural cross-lingual communication. In this and similar scenarios requiring personal voice adaptation, the option for on-device fine-tuning is particularly valuable. This approach eliminates the need to transmit personal audio data to external servers, thereby ensuring user privacy.
• a faster method for producing audio and video content. For example, while audiobooks are a popular medium for information consumption, the availability of titles narrated by professional voice actors remains limited. Voice cloning facilitates the creation of services capable of narrating books in real time, allowing users to select from a catalogue of pre-recorded voices for further personalization. Such solutions not only increase access to audiobooks but also function as a crucial tool for adapting content for individuals with visual impairments. A practical implementation of this concept is the service Bookmate.
A closely related use case is the dubbing of films and animated features. Here, voice synthesis and cloning models can substantially expedite production. For example, a professional actor might only need to record a limited sample of dialogue for a character. A model fine-tuned to their vocal profile could then generate the remaining lines automatically. For this application to be viable, the underlying model must ensure more than mere intelligibility; it must also replicate natural prosody, emotional expressiveness, and the capacity to convey subtle semantic nuances - an area that continues to demand significant technological advancement and refinement.
• restoring the voice for individuals who have lost the ability to speak. Using preserved previous audio recordings, voice cloning models can recreate a person's unique vocal characteristics.
• enhancing voice diversity for low-resource languages. These languages often suffer from a scarcity of recorded audio data, which severely limits the variety of voices available for speech synthesis systems. Voice cloning methods address this issue by enabling the transfer of vocal characteristics from recordings in other languages.
• data augmentation for diverse speech processing tasks. The generated audio can be used to expand and enrich datasets for training other speech-related models.
However, it is crucial to acknowledge that voice cloning technology also presents significant security risks alongside its beneficial applications. Malicious actors can potentially employ synthesized voices for fraudulent activities, such as circumventing biometric security systems or disseminating misinformation. Consequently, establishing robust regulatory measures for this field is particularly urgent.
From an ethical perspective, the technology raises fundamental questions concerning an individual's right to their own voice, the protection of personal privacy, and the imperative of obtaining informed consent for the use of vocal data. These considerations highlight the need for clear legal frameworks to govern the industry and prevent potential misuse.
From a technical perspective, the development of cloning models must be accompanied by advances in protective mechanisms, such as speaker verification systems and audio deepfake detection ([3, 4, 5, 6]). These systems allow for additional verification in potentially vulnerable situations. Another potential defense method involves specialized watermarks: they can be added to original recordings to prevent unauthorized cloning ([7]), or embedded into generated audio during creation to simplify subsequent detection
([8, 9]).
Therefore, the coordinated advancement of generative technologies and protective methods will allow for the full realization of the positive potential of voice cloning while simultaneously enhancing security in critical domains.
Recent research on voice cloning models frequently employs large-scale, proprietary multi-speaker datasets to train both the base speech synthesis model and the speaker encoder. Furthermore, such studies typically do not impose stringent limitations on the volume of data required for subsequent speaker adaptation. Reducing the data requirements for adapting to a target speaker not included in the original training set would simplify practical applications and accelerate the fine-tuning process. In this
work, the core model is trained exclusively on publicly available datasets encompassing a diverse range of voices. For encoding, only a single audio sample of a novel speaker is required. For speaker adaptation, the feasibility of voice transfer using just one minute of audio data is specifically investigated.
Many modern neural speech synthesis architectures consist of two parts. The first, feature predictor, generates some intermediate features, most often spectrograms [10, 11, 12], from the input text (symbols/phonemes). The second, vocoder, reconstructs the audio from these intermediate features [13, 14, 15]. All models studied in this dissertation follow this structure due to several advantages.
Any speech synthesis model can be modified for voice cloning. For the zero-shot, it is sufficient to pass the speaker embedding as an additional input to some blocks during training. For a few-shot, it is necessary to determine which parts of the models should be adapted on a new data. For example, a vocoder trained on a multi-speaker dataset does not require updating during speaker adaptation, that significantly simplifies the task. Furthermore, feature predictor's blocks related to the text processing may also be frozen.
The first part of the thesis is devoted to developing a hybrid diffusion probabilistic model for voice cloning. It was compared against a wide range of speech synthesis architectures adapted for the voice cloning task, including recurrent neural network (RNN)-based [10, 16], flow-based [17], transformer-based [11], and diffusion-based models [62]. Several issues critical to voice cloning were considered. The first involves minimizing the requirements for data, computational resources, and time for both inference and adaptation. The second concerns identifying the optimal subset of model parameters for updating in a few-shot learning scenario. Each model exhibits distinct characteristics regarding fine-tuning stability and the selection of parameters to adapt.
The second part of the work describes methods for improving certain characteristics of the studied models. For diffusion architectures, the speed of generation is a problematic area due to their iterative nature. This is especially noticeable when using a central processing unit (CPU) instead of a graphics processing unit (GPU). One proposed method for accelerating image generation is combining a diffusion model with a generative adversarial network (GAN). This method is adapted for speech synthesis by the author. Another important characteristic of speech perception is its expressiveness. However,
generating emotional audios is a challenging task. This work considered a universal method of modifying speaker embeddings to enable speech generation with four different emotions.
Therefore, the goal of this work is to develop a new hybrid diffusion-based voice cloning architecture, compare it with other architectures, and determine the optimal amount of data required to maximize efficiency while preserving quality. Additional focus is placed on improving key characteristics such as generation speed and the emotional expressiveness of the synthesized speech.
Key results and conclusions Contributions.
The main contributions of this work can be summarized as follows:
1. A new multimodal diffusion probabilistic model was introduced combining two tasks: voice cloning and conversion. Such hybrid structure offers several advantages, allowing to adapt only decoder part using untranscribed recordings. Moreover, the data requirements are significantly reduced, enabling fine-tuning with just 15 seconds of data without a significant loss in quality.
2. We adapted diffusion generative adversarial network architecture for speech synthesis to accelerate the diffusion model inference. This method was compared with several others. It was found that this approach reduces the number of necessary steps; however, it required increasing the model size to maintain a similar quality level, which only slightly improved the synthesis speed on a CPU.
3. We revealed the ability of utilized speaker embeddings to encode not only speaker related information, but also emotional one. We proposed the method to identify components responsible for different emotions and the best values for them. We tested this approach on zero-shot voice cloning and were able to synthesize speech with 4 emotions: sadness, happiness, anger and surprise.
Theoretical and Practical Significance
This dissertation presents a comprehensive study and comparative analysis of modern neural architectures for speech synthesis in the context of voice cloning. The research developed and proposes a novel hybrid diffusion architecture. To evaluate it objectively, a comparison was conducted with a range of contemporary models, examining their specific behaviors in cloning scenarios and assessing their adaptability using a limited data volume (1 minute). Particular focus was given to enhancing key characteristics for practical application, such as the generation speed of diffusion models and the ability to synthesize emotionally expressive speech.
Key aspects/ideas to be defended
1. A new hybrid diffusion model capable of voice cloning and voice conversion. Comparative analysis with other models.
2. Implementation of inference acceleration method of diffusion model which suppose a combination with GAN. Comparison with other methods was conducted.
3. A method of emotional voice cloning by modifying devoted speaker embedding components
Personal contribution. In the first publication (Sec. 4.4.2) the thesis' author implemented and trained the main models and conducted experiments. Training of other models for comparison was done by colleges. As for the second publication (Sec. 4.4.2)) the main contribution is implementation of one of the methods for increasing inference speed of diffusion models. And in the last paper (Sec. 4.4.2) the author was responsible for the proposed method of searching the best values for the components, responsible for emotions, as well as carried out all experiments.
Publications and approbation of the work
First-tier publications:
1. Sadekova T., Gogoryan V., Vovk I., Popov V., Kudinov M., Wei J. (2022) A Unified System for Voice Cloning and Voice Conversion through Diffusion Probabilistic Modeling Interspeech 2022, 3003-3007, doi: 10.21437/ Interspeech.2022-10879 (Core A)
2. I.Vovk, T.Sadekova, V.Gogoryan, V.Popov, M.Kudinov, J.Wei "Fast Grad-TTS: Towards Efficient Diffusion-Based Speech Generation on CPU", Interspeech, 2022 (Core A)
3. Shaheen, Z*., Sadekova, T.*, Matveeva, Y., Shirshova, A., & Kudinov, M. (2023). Exploiting Emotion Information in Speaker Embeddings for Expressive Text-to-Speech. INTERSPEECH 2023, 2038-2042. (Core A)
1 Literature and field review
Text-to-speech synthesis models are designed to generate natural-sounding speech from written text and are widely used in various applications, including personal assistants, smart devices, robotics, and entertainment. They form the foundation for creating intuitive user interfaces. Virtual assistants, smart home voice helpers, and navigation systems that possess a natural and expressive voice significantly enhance user comfort. The quality and characteristics of the voice in synthesized recordings directly correlate with user comfort and trust: natural intonation, emotional expression, the absence of robotic artifacts, and correct semantic emphasis make interacting with the technology a seamless experience.
The ability to control the voice during generation, made possible by advances in voice cloning technology, expands the scope of its applications. With a relatively small number of recordings from a target speaker, it is possible to adapt a speech generation model to replicate their voice, capturing their pronunciation characteristics and speech patterns. This opens the door to personalized solutions in the media industry (e.g., content dubbing, audiobooks), film production, and video game development (e.g., rapid generation of character dialogue).
Alongside commercial and entertainment uses, speech synthesis and cloning technology can also address a range of social needs. For individuals with medical conditions that impair their ability to speak, TTS systems capable of replicating their own voice become a crucial communication tool.
It is also essential to consider the converse implications of voice cloning technology, which raise significant concerns within the realm of digital security. This technology enables a range of malicious applications and, more critically, undermines the reliability of biometric authentication systems that rely on voice as a unique identifier. Consequently, a crucial focus in speech signal processing is the continued refinement and development of models capable of reliably distinguishing between authentic human recordings and those synthesized by generative AI.
The field of speech generation has evolved significantly since its inception. Progress has accelerated markedly with the advancement of deep neural networks.
The first models were very simplistic mechanistic prototypes of human vocal tract,
capable of producing only a limited set of sounds and basic words. The later ones allowed to generate words and sentences by joining pre-recorded speech segments (concatenative synthesis) or by modeling sounds with special rules (formant synthesis). The next generation, statistical parametric synthesis, brought the technology closer to modern solutions. This approach involved training multiple sub-modules to predict speech parameters, which were then used to reconstruct the audio waveform. Hidden Markov Models (HMMs) were commonly employed in this framework. Finally, with the huge progress deep learning has made in the last several decades, various effective neural network architectures have been developed for TTS. These models have significantly improved the quality and naturalness of synthesized speech.
Early research in text-to-speech synthesis primarily focused on single-speaker architectures (e.g., [18, 19, 20]). However, the interest in the multi-speaker generation soon emerged ([21]). Models trained in multi-speaker setting reduce the data requirements for individual speakers, what is more realistic for the open-source datasets. Furthermore, they offer greater versatility without increasing computational overhead. Voice cloning is the extension of TTS, which enables speech synthesis in voices not present in the training data without full model re-training.
Twoimportanttopicscloselyrelatedtobothspeechsynthesisandvoicecloningmodelsshouldalsobehighlighted:
There are two other topics worth noting here, which are closely connected with both speech synthesis and voice cloning models:
• The first concerns their efficiency and inference speed. For systems where speech generation serves as the primary interaction interface, it is crucial that responses are synthesized as quickly as possible. This requirement drives the need for careful selection of model architecture and size. Additional constraints related to memory consumption and the ability to run on central processing units (CPUs) may arise in scenarios where models operate directly on user devices - a relevant consideration given the growing prevalence of smart home systems and mobile phone usage, particularly when handling sensitive data.
• The second topic is fine-grained control over synthesized speech attributes, such as emotion, prosody, rhythm. This capability holds substantial potential for diverse applications, prompting the development of various methods and architectures
over the past decades.
1.1 Audio representation methods
The first step in working with a sound signal is to convert it from an analog form to a digital one. The digital form is the result of sampling the original sound wave, a process where it is stored as a sequence of amplitude values (samples) that reflect changes in sound pressure over time (Fig. 1a). The key parameters are the sampling rate (the number of samples per second) and the bit depth (the number of bits used to encode a single sample). If the sampling rate is chosen correctly, following the Nyquist-Shannon theorem, this representation contains complete information and allows for the lossless reconstruction of the original analog signal. The waveform is visual and intuitive. However, to describe just one second of audio with a sampling rate of 16 kHz, one needs to store and process 16000 numbers, which increases the demands on computational resources and memory due to the large number of samples. Furthermore, even a slight change in pronunciation or the addition of background noise can drastically alter the appearance of the sound wave.
(a) Time-domain representation - waveform
(b) Frequency-domain representation - spectrogram
Figure 1: Audio representations
The most common method for audio analysis is the frequency-domain representation of a signal - the spectrogram (Fig. 1b). A sound wave can be represented as a sum of simpler waves, such as sinusoids, with different frequencies. This decomposition is achieved by applying the Fourier transform. Let us consider the procedure for obtaining
frequency representations step by step:
(1) Framing. Since the frequency content of sound changes over time, we cannot apply the Fourier transform to the entire signal at once. Therefore, the audio is split into small, overlapping segments with a window length of 20-40 ms and an overlap of 50% ± 10%.
(2) Windowing. The framing procedure can introduce artificial frequencies at the boundaries of the resulting segments. To address this, a smoothing window function is applied via convolution. This function gradually smooths the signal at the edges, ensuring that the signal's amplitude is close to zero at the beginning and end of each frame. The Hamming window (Fig. 2) is an example of such a window:
w[n] = 0.54 - 0.46 cos ( where 0 < n < N — 1, N - window length.
2nn N- 1
(33)
Q.
£ 0.4 <
0.0
Hamming
Samples
Figure 2: Hamming Window
(3) Discrete Short-time Fourier Transform (STFT). For a single frame, this transform provides a complex matrix from which the magnitude and phase of each frequency component of the signal within the window can be calculated.
1 N-1
X(m,k) = — x(n + mH)e^w^1, (34)
n=0
where m - the frame number, H - the hop size, N - the window length, and k -the frequency bin, with k G [0, N] due to complex matrix symmetry.
The magnitude can be represented as a spectrum plot, where the x-axis shows the range of possible frequencies and the y-axis shows the amplitude, indicating the intensity of frequencies in the original audio. To analyze how frequencies change over time, spectrograms are used. This is an image where frequencies are displayed on the y-axis, time is on the x-axis, and color represents the magnitude value of a specific frequency in the spectrum y(m, k) = |X(m, k)|2.
(4) Filter banks. Human hearing is more sensitive to lower frequencies (the speech range) and less so to higher frequencies. Consequently, the frequency axis is often transformed to represent this nonlinearity. For instance, the Bark scale, which converts frequencies according to specific rules, can be applied, or the more commonly used Mel scale. This scale employs a set of triangular filters (typically from 40 to 128) that aggregate values from different frequency ranges. This transformation can be represented numerically and visually as follows (Fig. 3):
m = 2595 * log1o(1 + 700) where m - frequency in mel, f - frequency in Hz.
(35)
Figure 3: Filter bank on a Mel-Scale
(5) Logarithmic scaling of values. Since the human perception of loudness is also non-linear (and rather logarithmic), a logarithmic transform is applied to the magnitude (power spectrum) values as well.
The frequency-domain representation effectively compresses the speech signal while preserving essential frequency information. It is also visually interpretable, as phonemes
can be distinguished in mel-spectrograms based on their formants, and it is more suitable for training neural network models. All models in this work utilize logarithmic mel- or bark-spectrograms.
The inverse Fourier transform allows for the lossless reconstruction of the signal from its spectrum. However, when computing the spectrogram, the phase information of the original signal is typically discarded, with only the magnitude being retained. In modern architectures, a dedicated model is trained to recover the phase information.
1.2 General TTS pipeline
The whole process of speech generation from text is usually divided into several parts (Fig. 4). First of all, necessary linguistic information should be extracted from the text, for example phonemes or position of stressed syllables. Next, feature (acoustic) predictor model generates intermediate acoustic features from prepared linguistic information. These may be mel-spectrograms ([18, 22, 23, 21]), phonetic posteriograms (PPGs) ([24]) or some other latent speech representations. This task is complicated due to the fact that the same text may correspond to different pronunciations and speaking styles, so this model should cope with such ambiguity and perform good text-speech alignment. The final part, vocoder, is responsible for bridging the gap between the acoustic features and the actual audio waveform and generating high-quality, natural-sounding speech ([13, 25, 14, 26, 15]). Sometimes, two last components are joined and the model is trained end-to-end transforming input text directly to speech ([27, 11, 28]). Such systems simplify the generation pipeline, however, give less flexibility and are presented less often.
In the latest solutions, an acoustic model conditions on some additional information, such as speaker embedding in voice cloning models, the whole reference audio for speaker's style or emotion transfer, language identifier for multi-language systems. This auxiliary input eases an ambiguous of text-speech mapping problem, mentioned earlier, and provides more controllable way of speech generation.
When a spectrogram is used as an audio representation, the acoustic model reconstructs the signal's magnitude in the frequency spectrum but not its phase. One of
Figure 4: Text-to-speech pipeline.
the earliest methods for phase reconstruction is the Griffin-Lim algorithm. This is an iterative method that recovers phase information from a given spectrogram. The process begins with random phase initialization. In each subsequent step, an inverse Fourier transform is applied using the real magnitude spectrogram and the current phase estimation to reconstruct a time-domain signal, followed by a forward Fourier transform on this signal to derive updated, more accurate phase values. While this algorithm is straightforward and requires no training, the audio synthesized using it often contains significant artifacts.
In modern speech generation systems, a neural vocoder is used. This is a trainable neural network that is specifically designed to reconstruct the phase, enabling the generation of high-quality speech.
1.3 Early TTS models
The earliest systems capable of synthesizing whole sentences appeared in the middle of the past century and resembled the way of human sound formation. Formant synthesis ([29, 30, 31]) is a rule-based approach utilizing source-filter theory, which supposes that air flow from our lungs (source) is modified by a special way by our vocal tract (filter) to create a particular sound. The transformation process depends on the relative positions of the organs in the vocal tract and their specific characteristics, which vary from person to person. Both parts are mathematically formalized, and all necessary sounds are described with a list of rules.
Concatenative TTS ([32, 33]), which dominated the field in the late 1980s and early 2000s, uses a simple underlined idea: constructing speech by concatenating audio segments recorded in advance. They may correspond to phonemes or diphones (halves of two adjacent phonemes) and be stored in a large databases. In the 1990s with the development of computer technology and resources this method transformed to a more sophisticated and fine-grained one - unit selection ([34, 35, 36]). Now larger databases were collected with more detailed characteristics, such as pitch accent, lexical stress, part-of-speech information, linguistic contexts, and more complex algorithms were proposed for segments selection and concatenation (e.g. pitch synchronous overlap and add (PSOLA)). Despite the fact that this approach was used in some commercial solutions and was state-of-the art for many years it had several drawbacks. It requires large database of audio recording in order to cover all possible combinations of speech units for spoken word. Also selection algorithms are sometimes not very reliable that leads to bad quality and less smoothness in prosody, stress, etc. ([37])
The next generation of speech synthesis methods, statistical parametric, gained prominence in 2000s and tried to overcome these issues. Instead of constructing audio from parts we generate acoustic features that are necessary to produce speech (pitch, durations, mel-cepstral coefficients, etc.) and then pass them to a vocoder such as STRAIGHT ([38]) and WORLD ([39]) to form a final waveform. Hidden Markov model (HMM) was typically used as generative engine ([40, 41, 42, 43]) in early versions. It was trained on linguistic-acoustic feature pairs. As an advantages of such systems we may name better speech naturalness, flexibility in parameters control and lower computational and data cost. On the other hand they still suffered from robotic intonation and lack of expressiveness ([44]), what makes generated speech easily distinguishable from human-recorded.
1.4 Neural-based TTS models
One of the earliest applications of deep neural networks (DNNs) in speech synthesis was replacing HMMs in parametric systems ([45, 46]). Subsequently, DNNs enabled significant simplification of the acoustic model architecture by integrating features predictors into a single block. Increased modeling power of DNN made it possible to process characters or phonemes as linguistic representation directly and use high-dimensional
spectrograms as acoustic representations, while sequence-to-sequence paradigm allowed to learn text-speech alignment through attention mechanism. Additionally, neural networks became widely used as vocoders backbone. Neural-based TTS systems achieve superior voice quality in both intelligibility and naturalness, while requiring minimal manual preprocessing and feature engineering. In the following paragraphs, we examine the architectural evolution of these systems.
RNN-based
Recurrent neural networks (RNNs) demonstrate particular effectiveness in modeling sequential data and capturing long-range dependencies, making them well-suited for speech processing. The Tacotron family of models ([18, 10]) marked a significant breakthrough in the field by introducing an encoder-decoder framework with attention mechanism for mel-spectrograms generation. In this architecture, the encoder transforms linguistic inputs (characters or phonemes) into fixed-dimensional representations, while the decoder autoregressively predicts acoustic features.
These models were among the first to successfully employ attention mechanisms for learning text-to-speech alignment automatically. They substantially improved synthesized speech quality compared to previous approaches and established the foundation for numerous subsequent advancements. Later works were built upon this framework to enhance various aspects of synthesis, including robustness ([47, 16]) and expressiveness controllability ([48, 49, 50]).
CNN-based
Convolutional neural network-based (CNN) TTS models process entire data sequences simultaneously rather than frame-by-frame by applying filters for linguistic and acoustic representations. This makes them more computationally efficient during both training and inference. In addition, by stacking multiple convolutional layers with varying kernel sizes or dilation rates, CNNs can capture both long-range and short-range dependencies - a crucial capability for natural-sounding speech synthesis.
The first model from its family DeepVoice ([20]) appeared as an enhancement of
a parametric system, while maintaining the autoregressive nature of acoustic feature generation through its integration with the WaveNet vocoder ([13]). Subsequent versions [21, 12] introduced significant architectural modifications. Later developments included ClariNet [51], which implemented a fully convolutional architecture following the encoder-attention-decoder paradigm, while enabling direct end-to-end text-to-waveform generation. Another advancement, ParaNet ([52]), addressed efficiency concerns by introducing a non-autoregressive model that achieved faster mel-spectrogram generation while maintaining reasonable speech quality.
Transformer-based
The self-attention mechanism in Transformer architecture allows to capture long-context dependencies, crucial for both text and speech modeling in TTS systems in terms of prosody and rhythm. TranformerTTS ([19]) was the first model supporting Transformer architecture for encoder and decoder, while maintaining an autoregressive attention mechanism for frames generation. While drawing inspiration from Tacotron's architecture and achieving comparable audio quality, this model demonstrated significantly faster training. The next generation of models addressed key limitations of autoregressive approaches - inference speed and attention mechanism robustness. FastSpeech and FastSpeech2 ([22, 11]) proposed a length regulator block to explicitly model alignment between phonemes and frames. Furthermore, FastSpeech2 and FastPitch ([23]) further enhanced expressiveness through integrated pitch and energy predictors, enabling both improved naturalness and controllable speech characteristics. As a result, FastSpeech2 leverages extremely fast and robust speech generation with the quality on par with the previous autoregressive systems. This family of models also found broad application and became a base for many future works [53, 54].
Diffusion-based and other architectures
With the development of different ideas in deep learning they gradually were adapted for TTS models. Numerous studies have proposed generative adversarial networks (GANs)
and variational autoencoder (VAE)-based approaches for speech-related tasks [55, 56, 57, 58]. (GANs remain among the most effective solutions for vocoders [15]). Normalizing flow-based solutions found application in such models as Flowtron [59], Glow-TTS [17] and the popular end-to-end model VITS [28].
More recently, denoising diffusion probabilistic models (DDPMs) have demonstrated significant potential in modeling complex data distributions, particularly in image and graph generation. In the speech domain, DPMs initially emerged as vocoders achieving impressive results in fine-grained waveform reconstruction ([60, 61]). Subsequently, several feature generators, such as Grad-TTS [62] and Diff-TTS [63], were developed. They follow the standard TTS encoder-decoder structure, described earlier, but decoder is trained in the denoising diffusion paradigm. Remarkable generative capabilities of DPMs promises substantial advancements in the field, however, a weak point is slower inference speed, caused by their iterative denoising nature as well as larger models sizes. Though it may be not so problematic when running on GPU.
State-of-the-art speech generation models are now trained on large multi-speaker datasets consisting of tens of thousands hours, where zero-shot cloning of a voice from a short audio fragment is the standard operational mode. For instance, in language model-based approaches that autoregressively predict audio tokens, the capability for in-context learning is utilized ([valle, 64]). During generation, the model is conditioned on a reference audio recording, copying the target voice from it. In models based on the flow matching paradigm, the concept of ¡¡speech infilling¿¿ is employed, which involves masking a part of the audio and iteratively reconstructing this segment. During generation, the style and voice are also copied from a reference recording, which is treated as the unmasked segment ([65, 66, 67]).
1.5 Voice cloning
One of the first works referred to a question of multi-speaker speech generation by one model was DeepVoice2 ([21]). It employed a trainable low-dimensional embedding, stored in a lookup table and jointly optimized with the main model to encode speaker-specific characteristics. The majority of other parameters were shared across all speakers. However, a key limitation of this approach is its fixed-size speaker embedding table,
restricting synthesis to voices observed during training. A more challenging task is generating speech with an unseen voice using only a few reference samples, insufficient for training a model from scratch. This is the main objective of voice cloning.
Figure 5: Voice cloning. Training and inference for zero-shot and few-shot scenarios.
There are two main approaches: speaker encoding and speaker adaptation (zero- or few-shot voice cloning) [1](Fig. 5):
• The first one supposes using an embedding from a speaker encoder model, which may be trained jointly with the TTS model or separately on some other speech-related task. Compared to lookup table, it takes reference audio as input and thus is not limited by the speakers from the training dataset. Such idea extends multi-speaker TTS system's capability to new voices if it was trained in this paradigm on multi-speaker dataset.
This method is fast and do not require much additional time and computational cost during inference except for speaker embedding extraction. Furthermore, only one short audio sample is enough to perform zero-shot voice cloning. However, a drawback is the limited capacity of one vector to encode specific voice features in details.
• The second approach, speaker adaptation, involves fine-tuning all or a subset of the model's parameters to better match an unseen speaker's voice. While this yields more precise adaptation, it demands greater computational resources, more training data, and longer tuning time compared to the speaker encoding method.
Speaker-dependent modeling were initially explored in automatic speech recognition (ASR) to enhance performance by incorporating speaker-specific characteristics [68, 69, 70]. Insights from these studies later inspired solutions for voice cloning. Among the earliest models, VoiceLoop [71] modified a method of speaker embeddings table, jointly trained with TTS model, by proposing an idea of fine-tuning only this fixed-dimensional embedding during adaptation time. So this model may work only in few-shot voice cloning scenario.
Subsequent work [72] introduced a jointly trained speaker encoder neural network, enabling extraction of voice embeddings from arbitrary utterances during inference and thus zero-shot voice cloning. A further generalization of this concept emerged through independently trained speaker networks, which model the broader space of speaker characteristics ([1]). This approach decoupled speaker modeling from speech synthesis and relaxed data requirements: while the TTS model requires high-quality training data, the speaker encoder can be trained on untranscribed speech - including samples with reverberation and background noise - benefiting more from a large speaker pool.
These ideas were comprehensively compared in [1]. Building on this, authors of [2] proposed to utilize a speaker verification network to extract voice representations, and it turned out to be quite effective idea. Several other works studied the possibility of decreasing number of updated parameters during fine-tuning ([53]) and tried different adaption approaches, such as meta-learning ([73]).
It is worth noting that vocoders are typically universal models. Once trained on a multi-speaker dataset, they can synthesize speech for new voices without requiring any modifications. Consequently, in few-shot voice cloning, only the acoustic model in the TTS pipeline needs to be adapted to the target speaker.
1.6 Controllability and efficiency
Controllable TTS systems aim to regulate various aspects of synthesized speech, such as speed, energy, pitch, prosody, timbre, gender, emotion or high-level stylistic features. While voice cloning specifically addresses timbre control, other methods have been developed to manage these diverse parameters.
Early formant synthesis models employed predefined rules to explicitly control
duration, formant frequencies, and pitch to achieve desired acoustic characteristics during synthesis. Similarly, advanced concatenative TTS systems could manipulate prosody by leveraging extensive phoneme databases annotated with tags corresponding to specific speech attributes. These systems also enabled speaker selection by extracting appropriate units from the target speaker's voice data. HMM-based TTS approaches offered control over prosody, pitch, speaking rate, and timbre through the adjustment of statistical parameters.
Unlike traditional methods, neural-based TTS systems gave an opportunity to model more complex relationships between input text and speech output, enabling nuanced control over various speech characteristics. Architectures supporting fine-grained control of rhythm and intonation through duration, pitch and energy modeling emerged alongside the first best high-performance TTS models [22, 11, 23, 16]. Beyond these structured parameters, less formalized characteristics such as speaker style and speech manner have also been investigated, typically modeled using latent representations extracted from reference audio. For instance, [74] introduced a reference encoder block jointly trained with the Tacoton model and implicitly encoded speech stylistic information from the input prompt and then passed it to decoder as an additional input. Another extension of the Tacotron architecture, Global Style Tokens (GST) [48], employed combinations of style tokens to effectively extract information form the reference speech.
Further advancements include StyleSpeech [73], which utilized style-adaptive layer normalization as a style conditioning method, providing robust zero-shot voice cloning performance, and GenerSpeech [75], which addressed to a question of style transfer for out-of-domain custom voices and introduces a multi-level style adapter. Additionally, research has explored preserving speech style across different languages, while avoiding accent [49].
In more recent architectures, methods are proposed for encoding prosody into discrete representations and subsequently reconstructing them using a recurrent prediction block ([76]). Another interesting approach involves decomposing the speech audio signal into four constituent parts: linguistic, prosodic, acoustic information, and a speaker vector ([77]). Each component is reconstructed in a specific manner to produce the final audio output.
Another related group of works has focused on enabling speech synthesis with specific
emotional tones, for example happiness, sadness, anger. These systems extend beyond producing intelligible and natural sounding speech and focus on generating expressive output that aligns with desired emotional context. Studies such as [78, 79] tried to either predetermine embeddings corresponding to specific emotions or pre-trained emotional encoder to better follow target emotional tags or reference speech. [80] further advanced the field by detecting both emotion type and intensity for prosody modeling while disentangling timbre from the process - enabling the generation of emotional speech across different speaker voices.
1.7 Security
The advancement and refinement of speech synthesis and voice cloning models have sharply raised the issue of regulating their use for fraudulent purposes. Beyond the legitimate tasks these technologies are designed to solve, they can be exploited to bypass voice verification systems, misuse personal data without authorization, and spread false information. Such threats necessitate serious countermeasures.
To highlight the urgency of the problem and draw attention to it, specialized competitions are periodically held within the framework of scientific speech processing conferences. One of the most notable is ASVspoof (Automatic Speaker Verification Spoofing and Countermeasures). Held every two year since 2015, each iteration introduces increasingly complex scenarios and conditions. An important secondary outcome of this initiative has been the ASVspoof 2019 dataset ([81]), which remains one of the primary resources for training and testing detection models.
The ADD (Audio Deepfake Detection Challenge), held in 2023 as part of the ICASSP conference, should also be noted. It focused on modern neural network-based solutions and included two tasks: determining whether an entire audio clip or a part of it was generated. An important outcome of this competition was the publication of the associated dataset [82].
Speaker verification and audio deepfake detection models
Early research on security issues focused on enhancing speaker verification systems to protect against spoofing attacks. More broadly, the task can be formulated as classifying audio recordings as either synthetic or natural - audio deepfake detection. The basic
pipeline typically consists of two stages. The first involves extracting reliable features capable of accurately encoding the necessary information. The second stage entails training a classifier model.
The selection of acoustic features is a crucial step that significantly impacts the final performance of the entire system. Early studies employed various digital signal processing methods as primary features. For instance, different variants of frequency representations were widely used, such as spectrograms, mel-spectrograms, and cepstral coefficients ([83, 84, 85, 86, 87]).
Later, alternative approaches emerged that utilized learnable audio representations. Such features can be obtained through joint training with models, either task-specific or pre-trained on other objectives. These scenarios typically leverage large volumes of data, allowing the representations to be more robust to various recording conditions. Furthermore, they are capable of capturing complex acoustic patterns crucial for speaker verification that are difficult to detect using classical signal processing methods.
This idea led to a number of works where the two stages - feature extraction and classification - are integrated into a single architecture capable of extracting the necessary features directly from the audio waveform. Models of this type include SincNet ([88]), RawNet, RawNet2 ([89, 90]), and Statnet ([91]).
Another actively developing direction involves using representations derived from models pre-trained via self-supervised learning. Widely used models of this type include Wav2Vec2.0 ([92]), XLS-R ([93]), and WavLM ([94]). Examples of voice verification and synthetic audio detection research that leverage such representations can be found in [95, 4, 5].
Regarding classifier architecture, earlier solutions employed classical machine learning methods such as Gaussian Mixture Models, Logistic Regression, and Support Vector Machines ([96, 97]). With the advancement of neural networks, more modern approaches based on convolutional networks, Transformer blocks ([87, 86]), and graph neural networks ([98]) have been proposed for this task.
Alternative protection methods
An alternative approach to mitigating unauthorized voice cloning utilizes digital watermarks. These serve distinct protective functions. For instance, the AntiFake system ([7]) proposes safeguarding original audio content by embedding inaudible perturbations directly into user recordings. These modifications are designed to prevent voice cloning models from accurately replicating the speaker's voice without consent, instead causing the models to generate a voice that differs from the target speaker's. This method demonstrates robust protection, successfully defending against state-of-the-art synthesis and verification systems in over 95% of test cases and maintaining resilience against attempts to modify the protected audio.
Another approach involves deliberately embedding special markers during audio generation to facilitate the subsequent detection of synthesized recordings. This provides developers with assurance that their publicly available products will not be used for fraudulent purposes and that the generated audio can be easily distinguished by security systems.
An example of this concept is implemented in [8], which proposes the joint training of a generation model and a speaker verification model. Embedding watermarks in such a system must, on one hand, not degrade synthesis quality, and on the other, reliably distinguish synthesized recordings from natural ones. Experimental results indicate that the authors successfully achieved this balance.
A similar goal is pursued in [9], which introduces the Timbre Watermarking method. In this work, markers are added to frequency representations, and the architecture incorporates a specialized noise layer that simulates transformations typical of voice cloning models, thereby enhancing robustness. The paper investigates challenging detection scenarios, such as after applying a low-pass filter and audio compression. The proposed methodology enables the detection of voice-cloned recordings even when only 75% of the fine-tuning data contains the corresponding watermark.
It is important to acknowledge that, despite the efficacy of the aforementioned protective measures, the ongoing advancement of such technologies remains critically imperative. This necessity is driven by the continuous and parallel evolution of speech generation and voice cloning models.
Beyond purely technical countermeasures, the development and formal regulation of legal norms are equally essential. Legal and ethical protections against voice misuse are emerging at the intersection of biometric data regulation and principles of digital ethics. From a legal perspective, the voice is classified as a unique biometric identifier, granting it a distinct status under personal data protection statutes. The regulatory focus therefore extends beyond the mere collection of vocal data to encompass strict controls over its subsequent application. Within this framework, emerging legal norms mandate the labeling of synthesized content with special audio markers or disclosures. Furthermore, they establish clear legal liability for the creation and fraudulent use of voice models fine-tuned on personal data without the subject's explicit consent.
1.8 Focus of this work
While numerous studies have investigated key aspects of voice cloning model architectures and training methodologies, additional research remains necessary when considering real-world application challenges. In few-shot voice cloning scenario, adaptation often requires tens of minutes of audio data. An interesting research direction, therefore, involves establishing the minimal data requirements that can preserve both speech naturalness and speaker similarity.
Considering the potential deployment scenarios in interactive systems, special attention should be paid to the size and inference speed of the base architecture when selecting it, with further optimization where possible. The size and number of parameters requiring updates during adaptation to a new voice also impact the model's fine-tuning time.
Finally, controllability over speech characteristics, such as emotion and intonation, remains a valuable research direction for both TTS and voice cloning systems. Expanding model capabilities beyond timbre manipulation to incorporate control of speech attributes would enable broader applications, such as expressive audiobook narration or context-sensitive video dubbing.
2 Diffusion multimodal architecture
This chapter presents a novel diffusion-based architecture for voice cloning. It is a modification of the voice conversion model, which enables the replacement of speaker's characteristics in an audio recording while preserving its linguistic content. Consequently, the proposed system is hybrid and capable of performing both tasks following a single training stage.
The first successful voice cloning models demonstrating high-quality results were obtained by modifying architectures originally designed for speech generation. To assess the performance of these solutions and select the most promising ones as baselines for comparison, preliminary experiments were conducted, as described in the following section.
2.1 Preliminary experiments
In this dissertation, all models under consideration will follow a two-stage audio generation pipeline. An acoustic model will predict intermediate features—spectrograms—and a vocoder will subsequently reconstruct the audio waveform from them.
When adapting such systems for the voice cloning task, only the acoustic model requires architectural modifications. Since it primarily contributes to the synthesis system's cloning capability, various architectures were considered: the recurrent models Tacotron2 [99] and Non-attentive Tacotron [16], as well as a normalizing flows-based model GlowTTS [17]. To enable voice encoding and adaptation, these models were modified similarly to the approach in [2], and these modifications will be described later. The vocoder, on the other hand, only needs to be fine-tuned on a multi-speaker dataset to improve sound quality. For this section, the LPCNet vocoder [99] is used.
To optimize generation speed and adaptation time for a new voice, the base architectures were reduced in size where possible without significant quality loss, and fine-tuning was performed only on a subset of parameters, individually selected for each model. Furthermore, data requirements were lowered to just 1 minute.
Let us examine the architectures of the mentioned models in more detail, as two of them will be used for comparison with the proposed diffusion model.
2.1.1 Main models
Figure 6 shows their schematic structure with additional modifications for the voice cloning task.
Figure 6: A schematic architectures of three studied models and their modification for voice cloning.
Tacotron2
It is an encoder-decoder model with attention. It consists of 4 main parts:
• Encoder processes the input text sequence x = (xi, x2, ...,xTx) (phonemes) using convolutional and recurrent bi-LSTM layers. This combination enables to consider not only individual phonemes but also both short and long-range context, which
is crucial for correct word pronunciation. The output of this block is a sequence of states with the same length as the input: h = (hi, h2,..., hTx).
• Attention is commonly used in encoder-decoder architectures for predicting sequences of different lengths. Unlike classical architectures where all encoder-extracted information is compressed into a single context vector for the decoder, the attention mechanism dynamically determines which parts of the input sequence are most relevant at each generation step. Since the number of audio frames significantly exceeds the number of letters/phonemes in the text, and there's no predefined alignment between them, this concept enables alignment between these fundamentally different feature types.
The original Tacotron2 paper [10] employs Location Sensitive Attention (LSA), which incorporates attention weights from previous steps alongside the reweighted encoder hidden states. However, this configuration frequently leads to audio generation artifacts including repeated sounds, unclear pronunciation, prolonged pauses, and failure of stop token prediction. Therefore, following the approach in [99], this research replaces LSA with Stepwise Monotonic Attention (SMA, [100]), which enforces strict monotonic processing through the source text without returning. This modification substantially mitigates the aforementioned issues.
• Decoder consists of several recurrent layers and operates in an autoregressive manner. At each generation step, it takes as input its previous predictions processed through a small pre-net, along with a context vector c(h1, h2,..., hTx) from the attention mechanism. The output of the recurrent layer predicts a spectrogram frame yt and a stop token probability, indicating whether this prediction should be the final one. Generation terminates when the stoptokenprob exceeds a predefined threshold. The resulting sequence y = (y1,y2, ...,yTy) represents the initial spectrogram approximation.
• To further enhance the quality of acoustic features, a Posnet block composed of convolutional layers with residual connection is incorporated into the architecture. The convolutional layers process the generated spectrogram similarly to an image, enabling them to capture contextual information from surrounding frames. The final output of the entire model is the refined sequence y = (y1, y/2, ...,yTy).
The model is trained with mean squared error (MSE) between the predicted spectrogram and ground-truth one Ygt before and after the postnet Lfeat = — YgilH+llY' — Ygt ||2), where N - is the number of frames. The decoder works in teacher forcing regime during training utilizing ground-truth previous frame ygtt_ 1 instead of predicted one yt-i.
Parameters of Tacotron2 in this study differed from the original paper in order to make it more suitable for interactive application and are picked up to find a trade-off between quality and size. The final version had 18 million parameters. Encoder consisted of 3 convolutional layers with kernel size 3 and 1 bi-LSTM with 512 units. The model had decoder with 3-layer LSTM of 512 units and a decreased 4-layer postnet with kernel sizes [5, 5, 5, 5].
Non-attentive Tacotron (NAT)
Despite improvements, attention mechanisms often cause issues with pronunciation clarity and long pauses. Therefore, one research effort ([16]) proposed replacing it with a component that predicts the duration of each character d = (d1, d2,..., dTx). Then encoder outputs h = (h1, h2,..., hTx) are duplicated according to these predicted durations using a specific algorithm, and the autoregressive decoder receives features of length equal to the spectrogram duration, u = (u1,u2, ...,uTy). This method of aligning text and acoustic features has demonstrated more reliable generation results in other architectures ([11, 47, 17]) and was consequently implemented in this work.
The authors of the paper explored two methods for obtaining u: simple repetition of ht according to dt, and ¡¡Gaussian upsampling¿¿. In the latter approach, the duration prediction block additionally outputs a variance parameter a = (a1,a2,..., aTx) alongside the duration values d. Each ut is computed as a weighted sum of the encoder outputs h, where the weights are determined as follows:
di , V1^ N(t; ci,ai2) ,
C = ~o dj; wti = v^Tx Kr(.-2); ut wti, hi (36)
j= 1 1 N (t' Cj, a2) i=l
To compute the weight wit for each frame t, we use a Gaussian distribution centered at the midpoint of the input token i's interval with its corresponding predicted variance
ai. Unlike simple repetition, this upsampling method incorporates contextual information and produces more natural intonation, making it the preferred approach for our subsequent work.
To train the duration prediction module, we need the ground truth durations of phonemes in the training data (dgt). This is achieved using a pre-trained network that aligns audio with the corresponding text (Fig. 8). This work employs the Montreal Forced Aligner (MFA [101]) tool, which provides pre-trained models for various phoneme sets and languages. Our research uses an English language model with the ARPABET phoneme set, consisting of 68 phonemes plus two additional symbols for marking non-speech segments. The complete phoneme list is as follows:
AA0, AA1, AA2, AE0, AE1, AE2, AH0, AH1, AH2, AOO, AO1, AO2, AW0, AW1, AW2, AY0, AY1, AY2, B, CH, D, DH, EH0, EH1, EH2, ER0, ER1, ER2, EY0, EY1, EY2, F, G, HH, IH0, IH1, IH2, IY0, IY1, IY2, JH, K, L, M, N, NG, OW0, OW1, OW2, OY0, OY1, OY2, P, R, S, SH, T, TH, UH0, UH1, UH2, UW0, UW1, UW2, V, W, Y, Z, ZH, sp, sil
The numbers following vowel sounds indicate stress levels: 0 represents unstressed, 1 indicates primary stress, and 2 denotes secondary, weaker stress. The first step in obtaining actual phoneme durations involves phonemizing the input text, typically accomplished through either predefined rules or existing pronunciation dictionaries. Subsequently, the pre-trained model identifies the optimal alignment between the resulting phoneme sequence and the audio.
The loss function of NAT model comprises two components designed to enhance acoustic feature quality and train the duration prediction module. The first component, Lfeat, is similar to the Tacotron2 loss function, which modified using a combination of averaged L1 and L2 norms instead of only L2. The second component represents the mean squared error for predicted durations d: Ldur = N||d — dgt||2, with the duration prediction module being trained using teacher forcing.
We decreased the number of parameters in all modules of Non-attentive Tacotron compared to the original version. 2-layer Bi-LSTM with 512 units in the encoder was replaced with 1-layer Bi-LSTM with 256 units. The size of Bi-LSTM in the duration predictor was decreased to 256 from 512. Decoder 2-layer LSTM with 1024 units was changed to 3-layer LSTM with 512 units. For the postnet we used 4-layer neural network
with 1d-convolution, 256 channels and kernel sizes [5, 3, 3, 3]. Both Tacotron2 and Non-attentive Tacotron generated 3 feature outputs per step.
Glow-TTS
We considered Glow-TTS model as a potential alternative to Non-attentive Tacotron, because its training does not require golden durations. This is an important advantage for on-device training.
It is a flow-based architecture, which consists of text encoder, flow-based decoder and duration predictor. During training encoder fenc maps input text c into a sequence of parameters of a Gaussian distribution (^i,ai), while the decoder fdec maps the target spectrogram x into a sequence of latent variables Zj. The alignment A(i,j) between latent variables z2- and predicted parameters (^i, ai) is calculated via Monotonic Alignment Search (MAS), an algorithm that looks for the most probable monotonic alignment between linguistic and acoustic features. Its output is also used for training the duration predictor fdur. During inference the duration predictor controls the number of samples drawn from each distribution N(z; ai) and the decoder carries out inverse transform of the sampled latents z2-. We modify Glow-TTS to output normalized acoustic features for LPCNet and to work with the input speaker embedding generated by the speaker encoding network. Since Glow-TTS is designed to have a duration predictor as a separate module, we also make this duration predictor conditioned on speaker embedding. We should note though that we did not try to implement a compact version of Glow-TTS and did not change any parameters including temperature.
Vocoder
[14] was utilized as a vocoder for all feature generators described above (Fig. 7). In contrast to other popular neural vocoders [13, 25, 141], it combines classical digital signal processing methods based on Linear Predictive Coding (LPC) with modern deep neural networks, enabling high-quality synthesis with minimal computational requirements.
The LPCNet architecture is divided into two main components. The first part employs classical linear prediction to compute LPC coefficients and vocal tract oscillation parameters (pitch). These parameters are sufficient for synthesizing a rough speech signal that describes the spectral envelope and periodicity of the speech. This stage offloads the burden of predicting these parameters from the second component - a recurrent neural network. The RNN, in turn, utilizes these parameters along with its previous outputs to autoregressively predict the signal's excitation, which is then transformed into the final speech waveform.
Figure 7: Vocoder LPCNet
By delegating the computationally expensive task of spectral envelope reconstruction
Figure 8: A text-audio alignment example. Figure 9: A speaker encoeder
to the simple yet effective LPC filter, the neural network only needs to predict the relatively simple excitation signal. This enables real-time high-quality speech synthesis on standard CPUs (including mobile devices) without requiring powerful GPUs. Even a compact model (1.4 million parameters) can synthesize good quality audio at high speed. The input features used are mean- and variance-normalized bark-spectrograms. The model was trained on the multi-speaker LibriSpeech [102] dataset containing 460 hours of English speech (the ¡¡clean¿¿ subset). The audio sampling rate used for vocoder training was 16 kHz.
2.1.2 Speaker encoder
The speaker encoder plays a key role in voice cloning architectures, as it is responsible for extracting and compressing speaker characteristics into a compact vector representation. Its task is to identify and generalize individual speech attributes, such as timbre and pronunciation style, and represent them as a single embedding.
Two principal approaches for deriving this representation are identified in the literature. The first, proposed by [1], involves the joint training of a speaker encoder with an acoustic model, utilizing only a spectrogram reconstruction loss. In contrast, a second methodology leverages models pre-trained on specialized tasks that are more relevant for identifying voice characteristics. This work adopts the latter paradigm, aligning with the framework established in [2]. Thus, the speaker encoder used herein and in subsequent chapters is one that was pre-trained separately for the task of speaker
verification [6]6.
This task aims to verify the identity of a speaker from a voice sample, for purposes such as authorizing device access. A fundamental requirement for this task is model discriminability: in the learned feature space, embeddings derived from the same speaker must exhibit high similarity (minimal intra-class distance), whereas those from different speakers must be readily separable (maximal inter-class distance). To this end, training leverages specialized metric learning loss functions, such as triplet loss, to achieve this objective.
The chosen model adopts a novel training methodology for speaker verification based on the Generalized End-to-End (GE2E) loss function, which demonstrates superior performance over the earlier Tuple-Based End-to-End (TE2E) loss. The GE2E loss facilitates more efficient parameter updates by prioritizing challenging, hard-to-classify examples during each training step. Furthermore, it removes the necessity for a preliminary data selection phase.
During training, a similarity matrix is formed between each embedding vector and all speaker centroids in the batch, enabling the simultaneous consideration of both positive and negative interactions (Fig. 10). Two implementations of the loss function are proposed: one based on softmax and a contrastive loss, where the former is better suited for text-independent verification, and the latter for text-dependent verification. The GE2E loss pulls embedding vectors towards the centroid of their own cluster while pushing them away from the centroid of the nearest foreign cluster. A key improvement involves excluding the current embedding from the calculation of its own speaker's centroid, which stabilizes training and prevents degenerate solutions.
Joint training of speaker encoding models with TTS systems imposes stringent dataset requirements, as high-quality audio generation necessitates clean and accurately transcribed recordings. In contrast, training separately for the verification task offers distinct advantages: it benefits from acoustically diverse data and does not require transcriptions. This simplifies data collection and yields a model that encodes a broader range of speaker characteristics. Consequently, the encoder employed in this work was trained on a combination of the LibriSpeech [102], VoxCeleb [103], and VoxCeleb2 [104]
6 https://github.com/CorentinJ/Real-Time-Voice-Cloning
Data Utterance Batch of Features Embedding Vectors Similarity Matrix
q c-2 c3
Figure 10: Batch construction during the speaker encoder training.
datasets, comprising 3201 hours of audio from 8371 unique speakers.
The training dataset consists of speech audio recordings segmented into 1.6-second fragments, each annotated with speaker identification labels; textual transcripts are not used. The network input consists of 40-channel log mel-spectrograms. The network architecture comprises three consecutive LSTM layers with 768 cells each, followed by a linear layer that projects the features into a 256-dimensional space (Fig. 9). The final vector representation is created by L2-normalizing the output of the top layer at the final frame.
During inference, utterance of arbitrary length is divided into 800 ms windows with a 50% overlap. The network processes each window independently, and the final vector for the entire audio is formed by averaging and subsequently normalizing the resulting output vectors.
Despite the fact that the model was not specifically trained to extract speaker characteristics important for synthesis, the results show that the vector obtained during training on a speaker discrimination task is successfully used to convey identity to the synthesizing network.
2.1.3 Modifications for voice cloning
To apply the described speech synthesis models to a voice cloning task they were trained with several modifications:
• First, the training dataset must comprise a large number of speakers to enable the acoustic model to learn and represent vocal diversity.
• Second, the models accept additional input features, speaker embedding, for
controlled voice selection during generation. Such representations extracted by the speaker encoder described above and concatenated to the encoder outputs in all studied models (Fig. 6)
All models were trained on multi-speaker data from the LibriTTS [105] and VCTK [106] datasets with the following preprocessing: silent segments were removed, long recordings were splitted into shorter ones, noisy segments were filtered out, and speakers with a small number of recordings were excluded. The final training data contained 664 speakers from LibriTTS and 105 speakers from VCTK. Training process of every model took approximately 1 week on NVIDIA Tesla V100 32Gb.
The zero-shot models used speaker embeddings extracted from randomly chosen short target speaker recordings while the few-shot models were fine-tuned on few randomly chosen target speaker recordings with total duration of 1 minute. In the latter case, we averaged speaker embeddings extracted from all adaptation recordings. We also added 1 minute of recordings made by another speaker during adaptation of Tacotron2-based models as it improved attention stability. Golden durations for NAT were generated with Montreal Forced Aligner software. During adaptation of Tacotron2 only Decoder and Postnet modules were updated. For NAT models duration predictor was also fine-tuned. Adaptation time for Tacotron2-SMA and NAT models were 20 and 10 minutes correspondingly on 1 card NVIDIA Tesla V100 32Gb.
2.1.4 Experiments
We carried out two series of subjective evaluation tests. The goal of the first experiment was to identify the best architectures for both encoding and adaptation on in-field data with various acoustic conditions. For this experiment we used a private dataset consisting of fragments of radio shows and dialogues extracted from video clips of varying length and acoustic environment. The total number of speakers was 10 (4 male and 6 female). Test data for three of these speakers were parts of publicly available TTS datasets (for example, two speakers p280 and p315 from held-out VCTK dataset) while the remaining part contained small excerpts of radio programs or speech recordings extracted from video clips. The second one aimed at estimating the optimal number of fine-tuning steps and made use of 10 test speakers from LibriTTS.
MOS / Cate- Sound quality description Naturalness description
gory
5 - Excellent Sound is perfectly clear, completely Speech is completely natural, indis-
free from background noise, dis- tinguishable from a real person's
tortions or artifacts; studio-quality speech. Fully preserves smoothness
recording. and natural intonation.
4 - Good Sound quality is high, barely notice- Speech sounds generally natural and
able distortions or background noise pleasant. Very minor flaws in into-
may be present, but they don't dis- nation or pronunciation of individ-
tract or hinder perception. Sound is ual words may be present.
comfortable for extended listening.
3 - Fair Quality is acceptable, but back- Speech is intelligible, but clearly
ground noise or artifacts are notice- feels artificial. Noticeable
able. Speech comprehension doesn't monotony, unusual stresses,
require significant effort. minor pauses in unexpected places,
or slight "metallic" tone.
2 - Poor Low quality. Serious distortions, Speech is strongly "robotic". In-
strong noise. Causes discomfort to tonation is unnatural, rhythm is
the listener. choppy. Comprehension may re-
quire effort.
1 - Bad Quality is extremely low, speech Speech is completely mechanical,
is unintelligible or nearly unintel- monotonous and inexpressive. Very
ligible. Strong interference, con- difficult to understand.
stant interruptions, excessive hiss-
ing. Nearly impossible to listen to.
Table 1: MOS scale for assessing sound quality and speech naturalness.
2.1.5 Metrics
A number of objective metrics have been proposed for analyzing results in speech synthesis tasks; however, they do not always correlate with actual human perception. Therefore, the most commonly used metric is a subjective one - the Mean Opinion Score (MOS). During the evaluation process, a participant views a web page containing the generated results of a single sentence produced by various models under investigation. Their task is to evaluate different characteristics of the presented audio recordings. This study assessed the following properties: naturalness and sound quality on a 5-point scale, as well as voice similarity to a reference recording of the speaker's voice on a 4-point scale. The criteria for assigning scores are detailed in Tables 1 and 2. To increase rating precision, scores could be selected with 0.5 increments. As a recommendation, evalua-tors are advised to listen to the recordings in a quiet environment using headphones. Examples of the pages seen by the assessors are shown in Fig. 11 and Fig. 12.
For the second experiment in this section and all subsequent analyses, the perceptual attributes of sound quality and naturalness were merged into a single metric due to their high correlation named naturalness. Additionally, a 5-point scale was adopted for evaluating speaker similarity in later sections to align with the prevailing standard in the literature, thereby enabling direct comparison with other models.
In addition to synthesized recordings, some experiments included for comparison the original recording from the training dataset and its vocoder-resynthesized version (where a spectrogram is extracted from the original recording and then fed into the vocoder). This makes it possible to establish an upper bound for the values of the investigated characteristics and to assess the extent to which the metrics degrade during the audio generation stage by the vocoder.
The evaluation was carried out via the Amazon Mechanical Turk (AMT) platform. To ensure the reliability of the results, participation was restricted to pre-qualified Master workers. For additional quality control, besides the main recordings, a specially processed noisy recording with a distorted voice was included in the comparison set. This recording was designed to receive consistently low ratings. If a participant systematically assigned this recording a score above 2 points, all of their results were excluded from
MOS Similarity Voice Similarity Description
Category
4 Same: The synthesized voice is indistinguishable from the target
absolutely reference. Listeners express high confidence that both
sure samples originate from the same speaker, noting com-
plete alignment in timbral qualities, intonation patterns,
speaking style.
3 Same: The voice demonstrates general resemblance to the target,
moderately though perceptible differences are present. Listeners may
sure identify minor discrepancies in vocal characteristics while
still recognizing fundamental similarities.
2 Different: The voice bears limited resemblance to the target speaker.
moderately While some similarity is detectable, listeners perceive it
sure as originating from a different individual.
1 Different: The synthesized voice exhibits no perceptible similarity
absolutely to the target speaker. Absence of recognizable vocal
sure characteristics leads listeners to perceive it as originating
Обратите внимание, представленные выше научные тексты размещены для ознакомления и получены посредством распознавания оригинальных текстов диссертаций (OCR). В связи с чем, в них могут содержаться ошибки, связанные с несовершенством алгоритмов распознавания. В PDF файлах диссертаций и авторефератов, которые мы доставляем, подобных ошибок нет.