Scientists are increasingly dependent on powerful data sources and processing systems they cannot fully inspect, according to a new study published in BioScience. The research warns that the rise of proprietary and opaque systems — notably artificial intelligence and some remote-sensing products — is creating scientific "black boxes" that could erode reproducibility and public trust in research.
What the study found
An international team of researchers, led by Ivan Jarić at the University of Paris-Saclay, surveyed how modern tools reshape ecological and conservation science. While these technologies expand what is measurable — enabling continent-scale monitoring of biodiversity and threats or rapid analysis of vast datasets — the paper highlights a growing problem: many systems hide the processes that produce their outputs, limiting scrutiny.
"However, many of these tools represent true black boxes by keeping the processes behind those results largely hidden,"said Ivan Jarić, a researcher at the University of Paris-Saclay and lead author of the study.
The study identifies several categories of emerging black boxes used in ecology and conservation. Prominent among them are large language models and other AI systems that process and interpret satellite imagery, sensor data and complex ecological datasets. Researchers often lack access to the training data, algorithmic details or system-level testing that would be required to verify conclusions.
Why this matters
Reproducibility is a cornerstone of scientific method: independent researchers must be able to repeat analyses and obtain consistent results. When the tools used to generate findings are closed, proprietary or under commercial restrictions, that essential scrutiny becomes difficult or impossible. The paper argues this trend risks transforming scientific outputs into artefacts of particular commercial platforms rather than universally verifiable discoveries.
Those risks extend beyond AI. The study notes that many remote-sensing products and some wildlife tracking systems rely on proprietary processing pipelines or data sources that are not fully accessible to external researchers. Without transparent methods and data provenance, it becomes harder to interpret discrepancies, diagnose errors or fully assess uncertainty.
- Black-box categories include AI models (such as large language models), processed satellite products, proprietary remote-sensing pipelines and some wildlife-tracking platforms.
- Main concerns are lack of access to training data, closed algorithms, restricted testing and limited explanations of how outputs are produced.
- Consequences may include weaker reproducibility, reduced ability to assess uncertainty and diminished public and scientific confidence in findings.
| Type of tool | Examples / uses |
|---|---|
| AI and large language models | Analysing large datasets, interpreting imagery, modelling ecosystems |
| Satellite and remote sensing products | Continent-level monitoring, processed imagery with proprietary pipelines |
| Wildlife tracking / digital sensors | Telemetry and distributed sensor data where processing may be closed |
The paper stops short of rejecting the use of advanced tools. Instead it frames them as double-edged: they expand scientific reach but, if opaque, can make findings harder to verify. That tension is particularly acute as AI systems grow more capable and autonomous — the less transparent their internals, the harder it will be to interpret, replicate or challenge scientific claims based on their outputs.
Implications for policy and practice
The study’s findings have implications for funders, journals, research institutions and policy-makers. If scientific knowledge increasingly depends on privately controlled systems, the norms that require data sharing, method description and independent replication come under pressure. Ensuring the integrity of research may call for new expectations about disclosure of training data, algorithmic details and pipeline provenance — or stronger incentives for open alternatives.
For practitioners in ecology and conservation, the immediate tasks are practical: document dependencies on closed systems, report uncertainties linked to proprietary processing, and, where possible, prefer transparent tools or publish intermediate data that allow independent checks. For those who rely on research outcomes — policy-makers, conservation bodies and the public — the study underlines a need for cautious interpretation when findings rest on opaque technologies.
As scientific work becomes ever more computational and data-intensive, this study is a timely reminder that technological capability without transparency may weaken rather than strengthen the evidential base that society relies on.