Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Planning under Partial Observability and Object Discovery with Reusable Local Strategy Priors

Loading...
Thumbnail Image

Date

Authors

Tang, Zikang

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Partially observable online planning problems with dynamic object discovery and initially unstable identities are characterised by substantial uncertainty, strong sequential decision-making requirements, and a nontrivial exploration–execution trade-off. Network penetration provides a representative instance of this class of problems, since the agent initially lacks knowledge of the full network topology, node roles, stable identities, and exact connectivity, and must instead discover objects online, update an internal knowledge graph, and complete a target query under a limited budget. To address this setting, this thesis proposes a hierarchical online planning framework that combines knowledge-graph-based state representation, role beliefs, local subnet construction, and budgeted Monte Carlo Tree Search. To preserve local trade-offs among success probability, intermediate progress, and execution cost, the framework uses an ϵconstraint sweep to retain an approximate set of non-dominated local options rather than collapsing local decisions into a single weighted score. It further abstracts subnet-level search outcomes into action templates indexed by structural signatures in a generalized policy library, enabling cross-episode reuse of local strategies. At the global decision stage, rule-based actions, local subnet candidates, and library-retrieved candidates are aggregated within a unified selection process, while online feedback, slow forgetting, and successful trajectory replay are used to refine stored experience. Experimental results show that pure MCTS remains stronger in raw success robustness and wall-clock runtime. However, when the subnetbased framework succeeds, it often reaches terminal success in fewer interaction steps and preserves more remaining budget. Same-type policy-library transfer provides the strongest evidence for reusable local strategy priors, while cross-type transfer is useful but less reliable. Overall, the results suggest that combining local subnet search, approximate non-dominated option organisation, and abstract strategy reuse can improve interaction-level efficiency and same-type transfer in partially observable online planning problems with dynamic object discovery, while leaving computational efficiency and cross-type robustness as limitations for future work.

Description

the author deposited 31.08.2026 with approval from the supervisor.

Citation

Source

Book Title

Entity type

Access Statement

License Rights

DOI

Restricted until