Study-Aware Surrogate Modelling for Brain-Targeted Nanoparticle Optimization Across Heterogeneous Experimental Studies

CSEF · 2026 Computational Science (Senior Division)

Overview

Machine learning has, in recent years, gained popularity in biology, chemistry, and physics for the creation of surrogate predictors and optimization tools. Despite this promise, these algorithms currently see low adoption in real experimental pipelines due to large data requirements and unreliable generalization. In a field like nanomedicine, the average dataset can be as small as 400 points, suffering from extreme heterogeneity and sparsity. One defining feature of many such datasets is study identity, which indicates where the data came from. Current methods treat study identity as noise to be removed, as seen in ComBat and standard mixed-effects modeling. However, in sparse regimes, study identity is often a crucial and reliable signal. Treating it as such is a novel paradigm that maximizes the information present in small datasets. To utilize study identity effectively, I developed PRISM (Path-Retrieval In-Context Study-Conditional Model), a context retrieval architecture that constructs a CMI-weighted graph across training studies and uses Dijkstra shortest paths to identify the most transitively relevant studies for each test case. A fresh TabPFN is conditioned on this curated context at inference time, enabling study-conditional prediction without requiring study identity as an explicit feature. On brain nanoparticle delivery prediction, PRISM outperformed study-blind approaches and other leading methods under rigorous GroupKFold evaluation, demonstrating that path-sensitive graph retrieval is necessary for cross-study generalization. Using this surrogate, I constructed a multi-objective Bayesian optimization pipeline with Monte Carlo uncertainty propagation for context-aware generative design. Applied to nanomedicine, the system produces Pareto-optimal formulations for failed CNS therapeutics, with uncertainty-aware sampling providing theoretically-grounded improvement over trial-and-error approaches. This architecture reflects the data realities of experimental science, enabling reliable predictors and novel design tools across sparse, heterogeneous domains.

Competition history

  • CSEF 2026 Computational Science (Senior Division) · Entry S-07-33

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Browse more like this

Source: California Science & Engineering Fair public projects

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. An account also raises your daily allowance for “Has this been done?”, and lets you create a key for the MCP server with a much higher limit than anonymous use. Browsing stays public.

Continue with Google