NASA tested the Gemma 3 model for image analysis directly aboard a satellite
NASA JPL and Loft Orbital tested the Gemma 3 model from Google directly aboard the YAM-9 satellite — the NAVI-Orbital system analyzes images on-board and sends a text summary instead of raw data, with 88% accuracy in a ground-based test.
NASA Jet Propulsion Laboratory, in collaboration with Loft Orbital, tested the NAVI-Orbital system, which ran the Gemma 3 model from Google aboard the YAM-9 satellite to analyze images captured by the satellite's own sensors. According to NASA, this is the first demonstration of a vision-language model processing image data directly in orbit. The system is managed by an agent framework built on LangGraph and uses a quantized 4-bit version of the open-weight Gemma 3 4B model, without additional training on the target data.
In a ground-based benchmark using 7 960 images, the system achieved 88% classification accuracy. During two live tests over Toulouse in France and the coast of Argentina, the model ran on an Nvidia Jetson Orin AGX computing module and required only 8 GB of memory, according to the authors. The satellite is powered by solar panels with an output of 150 to 500 watts, according to Loft Orbital. The model generated text descriptions of the images and answered predefined questions, such as whether an image showed commercial or residential development.
The project authors — Juan M. Delfa and Taran Cyriac John from NASA JPL and Andrew W. Herson from Loft Orbital — state that the approach allows structured commands for operators to be replaced with a simple text prompt uploaded to the satellite. Paul Lasserre from Loft Orbital describes this principle as “semantic compression”: instead of transmitting large volumes of image data to Earth, the satellite sends only a text summary, which the company says reduces transmission capacity requirements from megabytes to tens of kilobytes. Delfa cites fire detection as an example use case, where the current delay between capturing an image and evaluating it can be up to 90 minutes, according to NASA.
NAVI-Orbital is deliberately separate from the satellite's flight control software and has no access to its controls; it only analyzes images. Delfa mentions a longer-term vision of using a similar system as a language assistant for astronauts whose movement is restricted by a spacesuit, which he says requires further research. The source article provides no details on the timeline for further development.
Why it matters
If the approach proves effective more broadly, it will allow satellites to send only a brief text summary to Earth instead of large images, reducing transmission capacity requirements and, according to NASA, potentially shortening the delay in detecting phenomena such as fires from tens of minutes to practically zero. It also demonstrates that a relatively small, publicly available model (Gemma 3 4B with 4-bit quantization, 8 GB of memory) runs on low-power hardware such as Nvidia Jetson Orin AGX even under satellite conditions — targeting a new type of phenomenon then requires only a change to the text prompt, without retraining or revalidating the on-board software.
Check the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.