A Multifaceted Analysis of Model Generalization in Challenging Environments
| dc.contributor.author | Tu, Weijie | |
| dc.date.accessioned | 2026-09-01T06:13:08Z | |
| dc.date.available | 2026-09-01T06:13:08Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | Deep learning models have become central to modern artificial intelligence systems, achieving strong performance across a wide range of visual recognition and reasoning tasks. However, deploying these models reliably in real-world environments remains challenging. In practice, models often encounter distribution shifts, produce unreliable confidence estimates, and must operate in settings where labeled data for validation or model selection are unavailable. These challenges are particularly pronounced for large vision-language and multimodal models, whose complexity and open-ended applications demand a deeper understanding of robustness, uncertainty, and model behavior beyond conventional accuracy-based evaluation. This thesis investigates how trustworthiness can be analyzed and supported in deep models, with a particular focus on vision-language systems operating under realistic conditions. The thesis first studies the robustness of deep visual and vision-language models across diverse architectures, training distributions, and fine-tuning strategies in Chapter 3. Beyond overall accuracy, the analysis examines robustness with respect to specific visual factors, out-of-distribution conditions, and the interaction between vision and language encoders. It further incorporates safety-relevant aspects such as predictive uncertainty, out-of-distribution detection, and sensitivity to 3D perturbations, revealing systematic failure modes that arise under realistic testing scenarios. Building on these findings, Chapter 4 further delves into uncertainty estimation and calibration in vision-language models. The results show that strong zero-shot performance does not necessarily imply reliable confidence and that they are not inherently well calibrated. Nevertheless, simple post-hoc calibration methods can substantially improve uncertainty estimates, even under distribution shifts, across different label spaces, and with only limited calibration data. These results demonstrate that reliable uncertainty estimation of vision language models can be achieved in practice without extensive labeled supervision. The thesis then explores how model performance can be assessed when labeled evaluation data are unavailable. Chapters 5 and 6 study the problem of ranking models without labels. Chapter 5 shows that signals derived from softmax prediction probabilities can provide informative indicators of relative model performance on unlabeled data. Chapter 6 extends this idea to large multimodal models and further investigates the role of uncertainty signals in model ranking. Finally, Chapter 7 introduces a data-centric perspective through an unsupervised dataset representation that captures semantic structure without labels, supporting reasoning about dataset similarity, training set suitability, and test set difficulty. Together, this thesis provides a unified perspective on trustworthy machine learning for deep learning models, advancing methods for understanding model generalization in challenging environments. | |
| dc.identifier.uri | https://hdl.handle.net/1885/733814669 | |
| dc.language.iso | en_AU | |
| dc.title | A Multifaceted Analysis of Model Generalization in Challenging Environments | |
| dc.type | Thesis (PhD) | |
| local.contributor.supervisor | Gedeon, Tamas | |
| local.identifier.proquest | Yes | |
| local.identifier.researcherID | NGS-4484-2025 | |
| local.mintdoi | mint | |
| local.thesisANUonly.author | 792f07df-77bf-45c0-ab67-ecf2042024bd | |
| local.thesisANUonly.key | eaa320a5-8476-c458-a9c9-237d66535b45 | |
| local.thesisANUonly.title | 000000027173_TC_1 |
Downloads
Original bundle
1 - 1 of 1
Loading...
- Name:
- Tu_PhD_thesis_2026.pdf
- Size:
- 5.77 MB
- Format:
- Adobe Portable Document Format
- Description:
- Thesis Material