NASA and IBM Release Open-Source Lunar Foundation Model Trained on 17 Years of Moon Data
NASA and IBM have introduced an open-source multimodal foundation model for lunar science, trained on nearly two million data bundles across 17 years of orbiter observations to map craters and locate polar ice.

NASA and IBM Research, alongside several academic institutions, have unveiled the NASA-IBM Lunar Foundation Model. Designed as an open-source artificial intelligence system for planetary science, the model turns 17 years of raw lunar orbiter records into an adaptable tool for mapping craters, analyzing surface geology, and locating potential water ice reserves.
Observation data from space missions is abundant, but human-labeled examples are rare and labor-intensive to produce. Rather than relying on task-specific algorithms trained for single applications, this foundation model pretrains on massive volumes of unlabeled data. Researchers can subsequently fine-tune it for downstream lunar science objectives using minimal labeled samples.
Training on SomBench: 17 Years and Four Missions
The model was trained from scratch using SomBench, which the researchers describe as the largest co-registered multimodal lunar corpus assembled to date. The dataset features nearly 2 million tile bundles across 11 modalities and two spatial scales:
- Approximately 1 million high-resolution images captured by the Narrow Angle Camera (NAC) at roughly 1 meter per pixel.
- Nearly 964,000 multispectral images from the Wide Angle Camera (WAC) at 100 meters per pixel.
- More than 30 spatially aligned data layers from nine scientific instruments across four missions, including gravity data from GRAIL, hydrogen distribution from Lunar Prospector, and mineralogical measurements from JAXA's Kaguya/SELENE probe.
The majority of the training data stems from 17 years of observations by NASA’s Lunar Reconnaissance Orbiter (LRO)—a mission that has generated more raw data than all other NASA planetary missions combined. To prevent data leakage during training and evaluation, the research team partitioned the corpus geographically by map zones rather than randomly shuffling image tiles.
Explicit Geometry and Multi-Scale Vision
Architecturally, the lunar model is based on TerraMind, an Earth observation model IBM developed in 2025 in collaboration with the European Space Agency (ESA) and Forschungszentrum Jülich. Instead of fine-tuning the Earth model, however, the team trained the lunar network from scratch.
A key challenge on the Moon is that illumination angles and shadow lengths dramatically alter how the surface appears, often overpowering the true physical properties of the terrain. To address this, the model receives explicit imaging geometry—including illumination angles, sun position, and tile boundaries—as contextual input alongside image and elevation data.
The architecture breaks lunar imagery, elevation layers, and geometric data into tokens and learns their relationships by predicting masked sections. Using a technique called FlexiViT, the model learns from both coarse-scale and high-resolution imagery in a single training run, enabling it to adjust to different image patch sizes without retraining. It also processes each distinct data modality through dedicated paths, avoiding the common limitation of treating different sensor inputs as simple stacked channels.
Outperforming Baselines in Polar Ice Detection
The research team evaluated the model across several scientific benchmarks, including crater detection at 100-meter and 1-meter resolutions, the segmentation of Irregular Mare Patches (young volcanic features that challenge theories of lunar cooling), and the prediction of polar ice deposits.
Permanently shadowed regions at the lunar poles remain cold enough to trap water ice for billions of years, making them critical targets for prospective resources like drinking water, breathable oxygen, and rocket propellant. In polar ice deposit predictions, the lunar foundation model reduced prediction error by up to 22 percent compared to the leading baseline, SwinV2-B.
In coarse-scale (100-meter) crater detection, the model outperformed SwinV2-B by nearly 19 percent while using only half the labeled training data. On meter-scale crater detection and Irregular Mare Patch segmentation, the model performed comparably to top baselines, with a modest 3 percent margin reported on volcanic patches. The authors also found that parameter-efficient fine-tuning via LoRA (Low-Rank Adaptation) matched or exceeded full fine-tuning performance across most benchmarks.
Scientific Analysis, Not Geodetic Replacement
Despite its analytical capabilities, the model is not designed for exact spatial positioning. In generative evaluation tests, the model's predicted latitude and longitude coordinates deviated by dozens of degrees in certain cases, and reconstructed terrain occasionally exhibited shifted absolute height values despite maintaining correct morphological shapes.
IBM Research Europe director Juan Bernabé-Moreno noted that the model’s primary value lies in connecting diverse instrument readings to uncover patterns invisible in isolated datasets, emphasizing that it serves as an analytical foundation rather than a substitute for direct physical measurements.
The lunar model expands the ongoing "AI for Science" initiative established between NASA and IBM under a Space Act Agreement in early 2022. The collaboration previously yielded the Prithvi Earth observation model in August 2023 for flood and wildfire mapping.
The NASA-IBM Lunar Foundation Model, along with its pretraining datasets and benchmark suites, has been made publicly available on Hugging Face, with source code published on GitHub and integrated into the open-source TerraTorch toolkit.



