Planning under Partial Observability and Object Discovery with Reusable Local Strategy Priors
| dc.contributor.author | Tang, Zikang | |
| dc.date.accessioned | 2026-08-31T04:54:59Z | |
| dc.date.available | 2026-08-31T04:54:59Z | |
| dc.date.issued | 2026 | |
| dc.description | the author deposited 31.08.2026 with approval from the supervisor. | |
| dc.description.abstract | Partially observable online planning problems with dynamic object discovery and initially unstable identities are characterised by substantial uncertainty, strong sequential decision-making requirements, and a nontrivial exploration–execution trade-off. Network penetration provides a representative instance of this class of problems, since the agent initially lacks knowledge of the full network topology, node roles, stable identities, and exact connectivity, and must instead discover objects online, update an internal knowledge graph, and complete a target query under a limited budget. To address this setting, this thesis proposes a hierarchical online planning framework that combines knowledge-graph-based state representation, role beliefs, local subnet construction, and budgeted Monte Carlo Tree Search. To preserve local trade-offs among success probability, intermediate progress, and execution cost, the framework uses an ϵconstraint sweep to retain an approximate set of non-dominated local options rather than collapsing local decisions into a single weighted score. It further abstracts subnet-level search outcomes into action templates indexed by structural signatures in a generalized policy library, enabling cross-episode reuse of local strategies. At the global decision stage, rule-based actions, local subnet candidates, and library-retrieved candidates are aggregated within a unified selection process, while online feedback, slow forgetting, and successful trajectory replay are used to refine stored experience. Experimental results show that pure MCTS remains stronger in raw success robustness and wall-clock runtime. However, when the subnetbased framework succeeds, it often reaches terminal success in fewer interaction steps and preserves more remaining budget. Same-type policy-library transfer provides the strongest evidence for reusable local strategy priors, while cross-type transfer is useful but less reliable. Overall, the results suggest that combining local subnet search, approximate non-dominated option organisation, and abstract strategy reuse can improve interaction-level efficiency and same-type transfer in partially observable online planning problems with dynamic object discovery, while leaving computational efficiency and cross-type robustness as limitations for future work. | |
| dc.identifier.uri | https://hdl.handle.net/1885/733814646 | |
| dc.language.iso | en | |
| dc.subject | Partially Observable Planning | |
| dc.subject | Dynamic Object Discovery | |
| dc.subject | Online Planning | |
| dc.subject | Monte Carlo Tree Search | |
| dc.subject | Knowledge Graph | |
| dc.subject | Local Strategy Abstraction | |
| dc.subject | Generalized Policy Library | |
| dc.subject | Strategy Transfer | |
| dc.title | Planning under Partial Observability and Object Discovery with Reusable Local Strategy Priors | |
| dc.type | Thesis (Masters) | |
| dcterms.valid | 2026 | |
| local.contributor.affiliation | School of Computing, ANU College of Systems & Society, The Australian National University | |
| local.contributor.supervisor | Haslum, Patrik | |
| local.identifier.proquest | Yes | |
| local.mintdoi | mint | |
| local.type.degree | Other |