Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

A Multifaceted Analysis of Model Generalization in Challenging Environments

Loading...
Thumbnail Image

Date

Authors

Tu, Weijie

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Deep learning models have become central to modern artificial intelligence systems, achieving strong performance across a wide range of visual recognition and reasoning tasks. However, deploying these models reliably in real-world environments remains challenging. In practice, models often encounter distribution shifts, produce unreliable confidence estimates, and must operate in settings where labeled data for validation or model selection are unavailable. These challenges are particularly pronounced for large vision-language and multimodal models, whose complexity and open-ended applications demand a deeper understanding of robustness, uncertainty, and model behavior beyond conventional accuracy-based evaluation. This thesis investigates how trustworthiness can be analyzed and supported in deep models, with a particular focus on vision-language systems operating under realistic conditions. The thesis first studies the robustness of deep visual and vision-language models across diverse architectures, training distributions, and fine-tuning strategies in Chapter 3. Beyond overall accuracy, the analysis examines robustness with respect to specific visual factors, out-of-distribution conditions, and the interaction between vision and language encoders. It further incorporates safety-relevant aspects such as predictive uncertainty, out-of-distribution detection, and sensitivity to 3D perturbations, revealing systematic failure modes that arise under realistic testing scenarios. Building on these findings, Chapter 4 further delves into uncertainty estimation and calibration in vision-language models. The results show that strong zero-shot performance does not necessarily imply reliable confidence and that they are not inherently well calibrated. Nevertheless, simple post-hoc calibration methods can substantially improve uncertainty estimates, even under distribution shifts, across different label spaces, and with only limited calibration data. These results demonstrate that reliable uncertainty estimation of vision language models can be achieved in practice without extensive labeled supervision. The thesis then explores how model performance can be assessed when labeled evaluation data are unavailable. Chapters 5 and 6 study the problem of ranking models without labels. Chapter 5 shows that signals derived from softmax prediction probabilities can provide informative indicators of relative model performance on unlabeled data. Chapter 6 extends this idea to large multimodal models and further investigates the role of uncertainty signals in model ranking. Finally, Chapter 7 introduces a data-centric perspective through an unsupervised dataset representation that captures semantic structure without labels, supporting reasoning about dataset similarity, training set suitability, and test set difficulty. Together, this thesis provides a unified perspective on trustworthy machine learning for deep learning models, advancing methods for understanding model generalization in challenging environments.

Description

Keywords

Citation

Source

Book Title

Entity type

Access Statement

License Rights

DOI

Restricted until

Downloads

File
Description