Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Generative Audio Synthesis in Sound, Speech, and Music

dc.contributor.authorNiu, Xinlei
dc.date.accessioned2026-08-13T05:38:18Z
dc.date.available2026-08-13T05:38:18Z
dc.date.issued2026
dc.description.abstractAudio is an important modality through which humans perceive and interpret the world. From speech and environmental sounds to music, it contains linguistic, emotional, and contextual information that can represent artistic expression. With the rapid progress of machine learning and generative AI, generative audio synthesis has emerged as a powerful approach to automatically producing realistic and diverse sounds, speech, and music content. However, current methods have challenges in requiring large computational resources, lacking flexibility across audio content, and offering limited controllability to users in practice. This thesis aims to tackle these limitations by developing audio synthesis techniques that are efficient, flexible, and controllable. This thesis explores how generative models can be designed to operate effectively under training resource constraints, adapt across different and diverse creative audio content, and allow for fine-grained and customized control to users. To achieve these goals, six complementary methods are proposed in this thesis. Among them, SoundLoCD in Chapter 3 and BDP in Chapter 5 focus on training efficiency, introducing a lightweight diffusion framework and a discrete latent optimal path modeling framework that preserve quality while reducing training computational requirements and strategies. HybridVC in Chapter 6 and SoundMorpher in Chapter 7 advance flexibility by enabling flexible voice conversion and seamless sound morphing across multi-modalities conditions. BVS in Chapter 4 and SteerMusic in Chapter 8 enhance synthesizing controllability, which allows fine-grained alignment between audio and visual context and empowers user-driven editing in music generation. In summary, these contributions form a unified perspective on AI-generated content of audio synthesis, which treats speech, sound, and music not as isolated problems but as interconnected domains of users immersive auditory experience. The resulting proposed methods move toward a future where audio generation is accessible, adaptive, and deeply integrated into creative, assistive, and interactive applications.
dc.identifier.urihttps://hdl.handle.net/1885/733814249
dc.language.isoen_AU
dc.titleGenerative Audio Synthesis in Sound, Speech, and Music
dc.typeThesis (PhD)
local.contributor.affiliationEngineering Computing Cybernetics, College of Systems and Society, The Australian National University
local.contributor.supervisor Martin, Charles
local.identifier.proquestYes
local.identifier.researcherIDPNG-3448-2026
local.mintdoimint
local.thesisANUonly.authore8774622-b999-4237-b6b5-092b1334f61f
local.thesisANUonly.key77ebc4e2-4ea0-2eba-81e0-bf39d3e72883
local.thesisANUonly.title000000026816_TC_1

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Niu_PhD_Thesis_2026.pdf
Size:
15.45 MB
Format:
Adobe Portable Document Format
Description:
Thesis Material