AI can label non-life as life after ~15 tweaks, Michigan State researchers warn
Avida-based tests show classifiers get “perfectly confident” about life they never see, raising risk for Mars, Europa, and exoplanet missions.

Christoph Adami and his student Ankit Gupta at Michigan State University tested AI classifiers on the Avida digital-life program and found AI can be repeatedly fooled. The consequence for decision-makers: AI may greenlight “life-like” chemistry when it is out of distribution, undermining mission claims without careful verification.
Artificial intelligence can be pushed into confidently calling non-life “life” after about 15 changes, according to Michigan State University researchers Christoph Adami and student Ankit Gupta. In their tests using the Avida digital-life simulation, the AI reached “perfectly confident” life classifications even though the researchers confirmed it was not life in those cases. The key failure mode, Adami says, is an AI Achilles heel when it is asked to classify things unlike its training examples, what he calls “out-of-distribution” samples.
This matters because many real-world searches for extraterrestrial biology will not be a simple “touch it and see” situation. Instead of a rover directly observing something microbial as we understand it, missions often rely on instrument data such as mass spectrometry, looking for molecules and processes that might correlate with life. If an AI on such a mission has been trained on Earth-like or otherwise limited biotic and abiotic samples, Adami argues it could still return a positive verdict on chemistry that has absolutely nothing to do with life, simply because it is out of distribution.
So what exactly did they do, and why does it convince? The researchers leaned on a long-running computational biologist idea from Adami: life can be defined by its ability to encode information and replicate it. To operationalize that, he devised Avida in 1993. Avida runs digital organisms written as code. Those programs can replicate by copying themselves, compete for resources in the form of CPU time, and evolve over time. While using digital life in evolutionary studies remains controversial, Avida gives Adami something crucial for this kind of AI experiment: a large, controllable dataset where the “life” versus “non-life” labels are knowable.
In their study, Adami and Gupta ran three months of computer analysis on a thousand parallel machines. They asked AI to determine which Avida programs had the properties of life and which did not. The complication is that “life” and “non-life” programs can have very similar coding patterns. That means the distinction is not always obvious, which is exactly the kind of ambiguity that can fool classifiers. The researchers started with programs representing a random sequence of molecules, then iteratively tweaked that sequence one change at a time. After each tweak, they checked whether the AI’s confidence that the input was life or non-life changed.
Here is the punchline. Adami said that within about 15 changes, the AI could be made perfectly confident of a life classification, even though none of the steps involved the inputs truly becoming life. And it was not a one-off trick. Adami said that no matter what starting sequence the researchers used, the AI was constantly fooled. This is the practical meaning of the “out-of-distribution” warning: the model performs well when asked about things it has essentially seen before, but it can behave very differently when it is handed something that lives in a different region of reality.
Adami draws a comparison to everyday machine learning examples. If you train AI to recognize pictures of apples and then ask it to pick out fruit from a dataset that includes other foods, it answers correctly virtually all the time. But if you replace the apple images with bananas and ask it to identify the fruit, the AI can struggle and misidentify things. The reason is straightforward: bananas represent out-of-distribution data relative to the training set. Adami argues that extraterrestrial microbes could be similarly unfamiliar. In other words, alien life could look, chemically or structurally, unlike the terrestrial microbes the AI was trained on. Without context, the AI has no reliable basis for separating life from not-life.
The researchers also frame the mission implication by walking through two detection pathways. If a Mars rover cuts open a rock and directly sees something that looks like microbial life as we know it, the answer could be more clear-cut and testable with traditional methods. But many other approaches will be indirect. The paper points to mass spectrometry as a common method for detecting molecules and processes linked to life, including attempts to detect life in the atmosphere of Venus, in the ocean of Europa, or on exoplanets via the Habitable Worlds Observatory, which NASA aims to launch in the 2040s to directly image exoplanets in the habitable zone of stars.
Adami’s concern is specifically about AI sitting in the workflow. If an AI on a mission examines mass spectrometry samples and has been trained on a set of biotic and abiotic examples on the ground, there is “a very great chance” it could return a positive verdict on samples that have nothing to do with life. The next step, Adami says, is to move out of the digital world and run the same kind of test with real-world data. Adami and Gupta are presenting their findings in August at the 2026 Conference on Artificial Life in Waterloo, Canada.
For executives and mission planners, the second-order issue is trust. AI is fast and good at pattern finding, and that is why it gets embedded into scientific workflows. But these results suggest that confidence scores can mislead when the inputs are outside training distribution. That means boards and leadership teams overseeing AI-enabled science systems should treat AI outputs as hypotheses, not verdicts, especially in high-stakes environments where the “unknown unknown” could be alien biology rather than sensor noise. The question is not whether AI can help search for life. It can. The question is whether the system is being designed with the right failure mode in mind, before a mission spends months or years optimizing for an answer it cannot verify.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Science
UVA researchers find a genome-analysis error, then release free ML tool to fix it
A UVA School of Medicine team pinpointed a widespread mistake in a popular genome method and built a machine-learning corrector for both bulk and single-cell data.

Nearly half of ransomware victims pay, as UK moves to block public payouts
A Sophos study finds ransom payments are common and demands are rising, while governments tighten the no-pay policy.

Two elephant-sized sauropods could rear like giants, then their bodies said no
Digital tests suggest their strong femurs supported upright rearing only while they were young and smaller.

