Fast Neural Speech Waveform Generative Models With Fully-Connected Layer-Based Upsampling
Although end-to-end (E2E) text-to-speech (TTS) models with HiFi-GAN-based neural vocoder (e.g. VITS and JETS) can achieve human-like speech quality with fast inference speed, these models still have room to further improve the inference speed with a CPU for practical implementations because HiFi-GAN...
Main Authors: | Haruki Yamashita, Takuma Okamoto, Ryoichi Takashima, Yamato Ohtani, Tetsuya Takiguchi, Tomoki Toda, Hisashi Kawai |
---|---|
Format: | Article |
Language: | English |
Published: |
IEEE
2024-01-01
|
Series: | IEEE Access |
Subjects: | |
Online Access: | https://ieeexplore.ieee.org/document/10438439/ |
Similar Items
-
MitoHiFi: a python pipeline for mitochondrial genome assembly from PacBio high fidelity reads
by: Marcela Uliano-Silva, et al.
Published: (2023-07-01) -
Unraveling metagenomics through long-read sequencing: a comprehensive review
by: Chankyung Kim, et al.
Published: (2024-01-01) -
Chromosome-Level Genome Assembly and Annotation of the Fiber Flax (Linum usitatissimum) Genome
by: Rula Sa, et al.
Published: (2021-09-01) -
Genome assembly of Melilotus officinalis provides a new reference genome for functional genomics
by: Aoran Meng, et al.
Published: (2024-04-01) -
Representing true plant genomes: haplotype-resolved hybrid pepper genome with trio-binning
by: Emily E. Delorean, et al.
Published: (2023-11-01)