CosmoFlow: a generative machine learning model that learns to compress large simulations of the early universe into small, interpretable representations that enable faster, easier scientific analyses
Project PI: Haewon Jeong, Ass't. Prof. ECE, Computer Engineering
Modern cosmology relies heavily on massive computer simulations to understand how the universe evolves. These simulations model how dark matter, an invisible substance that makes up most of the universe’s mass, assists in the formation of galaxies and other large-scale structures in the universe. High-quality simulations, while extremely informative, produce petabyte scale datasets that are extremely difficult to store, share, and analyze.
The CosmoFlow project tackles this bottleneck by using machine learning to compress these huge datasets into much smaller representations. Instead of storing every detail explicitly, the system learns a compact encoding that still preserves the important physical information. We do this by using one neural network to project the simulation to a smaller size, and then training another network to reconstruct the original simulation based on the output of the first network. The result is a compact, but scientifically meaningful representation of the original data. In particular, using our representations, scientists can infer information about the initial state of the universe, as well as generate samples of simulations run with different initial parameters, without having to run full simulations
Another key idea behind the project is making the model more understandable. Rather than acting like a black box, CosmoFlow organizes the compact representation so that different features correspond to structures at different scales in the universe—large cosmic filaments versus smaller clusters of dark matter, for example. This allows scientists to study how specific features encode information about the formation of our universe.
Haewon Jeong received her Ph.D. in Electrical and Computer Engineering from Carnegie Mellon University (’20), where her thesis laid key foundations for coded computing by applying information theory to design reliable large-scale computing systems. She then joined Harvard University as a postdoctoral fellow (’20–’22), shifting her focus to the reliability of machine learning systems—how to build models that people can trust and depend on. She was awarded the Harvard Data Science Initiative Postdoctoral Fellowship to research how machine learning systems may make biased decisions in educational settings. Now an assistant professor at UC Santa Barbara, Jeong’s research explores emerging ethical challenges of generative AI and applies AI methods to problems in physics. Her contributions have been recognized with the NSF CAREER Award, the JP Morgan Faculty Award, and a 2024 Hellman Fellowship.
RXIV: "CosmoFlow: Scale-Aware Representation Learning for Cosmology with Flow Matching"
Project Group: Jeong Lab
Project Contact: Sid Kannan <skannan@ucsb.edu