Research

My research has focused on speech synthesis, speech signal processing, and statistical machine learning, with particular emphasis on statistical parametric speech synthesis, HMM-based speech synthesis, neural speech synthesis, singing voice synthesis, and voice conversion.

Overview

A central theme of my work has been to develop flexible, controllable, and theoretically grounded methods for generating and transforming speech and singing voice. My research has covered both statistical modeling and signal processing aspects of speech technology, ranging from vocoding and parameter generation to large-scale speech synthesis systems.

Major Research Areas

Statistical parametric speech synthesis

I have worked on statistical parametric speech synthesis, especially HMM-based speech synthesis, including acoustic modeling, parameter generation, speaker adaptation, and style control. This line of work contributed to the development of flexible and trainable speech synthesis systems.

Speech signal processing and vocoding

My research has also addressed speech analysis, synthesis filters, vocoding, and parameter representation. These studies provided signal-processing foundations for statistical and neural speech synthesis systems.

Neural speech synthesis

More recently, my interests have included neural text-to-speech, neural vocoders, sequence modeling, diffusion and flow-based generative models, and controllable neural speech generation.

Singing voice synthesis and voice conversion

I have also worked on singing voice synthesis and voice conversion, including methods for modeling speaker or singer characteristics, pitch, timing, and expressive variation.

Speech technology resources

I have been involved in the development and dissemination of speech technology resources and tools, including systems and software related to statistical speech synthesis and Japanese text-to-speech.

Current Interests

My current interests include controllable speech and singing voice generation, neural and language-model-based speech synthesis, integration of signal processing and deep learning, and speech technologies that remain interpretable, adaptable, and practically useful.