TLDR: A new research paper introduces an agentic AI system capable of autonomously conducting scientific research, from hypothesis generation to manuscript preparation. The system successfully designed and executed three psychological studies on visual working memory, mental rotation, and imagery vividness, including online data collection with human participants. While demonstrating high efficiency and rigor, the research also highlights current limitations and raises important ethical questions about the future of AI in science.
Artificial intelligence is rapidly changing how scientific research is done, but most AI systems are designed for very specific tasks and still need a lot of human guidance. Imagine an AI that could handle the entire scientific process on its own, from coming up with ideas to writing the final paper. A new research paper, “Virtuous Machines: Towards Artificial General Science,” introduces just such a system, demonstrating its ability to independently conduct complex psychological studies.
An Autonomous Scientific Investigator
The researchers developed a domain-agnostic, agentic AI system capable of navigating the full scientific workflow. This means the AI can generate hypotheses, design experiments, collect data, analyze results, and even prepare complete manuscripts, all without continuous human intervention. This system represents a significant step towards Artificial General Science (AGS), where AI can drive scientific inquiry across various fields.
Real-World Psychological Studies
To test its capabilities, the AI system autonomously designed and executed three psychological studies focused on human cognition. These studies explored aspects of visual working memory, mental rotation, and imagery vividness. For one of the studies, the system even managed an online data collection with 288 participants. It also developed its own data analysis pipelines, working for over eight hours continuously to write and debug statistical code, and ultimately produced completed research papers for each study.
Key Findings and the Reliability Paradox
The studies yielded interesting results. For instance, one study found no significant correlation between visual working memory precision and mental rotation performance, challenging theories that suggest these abilities share common underlying resources. Another key finding highlighted the “reliability paradox” in individual differences research, where tasks that show strong effects at a group level might not be reliable for measuring differences between individuals. The AI’s analysis suggested that the poor reliability of some measurement parameters could explain the lack of correlation in its findings.
How the AI Works: A Multi-Agent Approach
The system operates using a hierarchical multi-agent architecture, much like a team of specialized scientists. A “master agent” coordinates the entire research project, delegating tasks to “orchestrator agents” (e.g., for methodology or data analysis) and “specialist agents” (for coding, troubleshooting, or review). This structure allows for complex tasks to be broken down and handled efficiently. The AI also incorporates human-inspired cognitive abilities such as abstraction (developing its own rules), metacognition (self-monitoring its thinking), decomposition (breaking down problems), autonomy (self-directed goal pursuit), and dynamic memory (selectively accessing relevant information, similar to a researcher consulting literature). This dynamic memory system, called d-RAG, allows the AI to search academic databases like Semantic Scholar, OpenAlex, and PubMed, and build its own knowledge base.
Also Read:
- Exploring the Evolution and Impact of AI Agents Across Industries
- Generative AI: Navigating the Path Between Progress and Peril
Efficiency and Ethical Considerations
The system completed full studies in approximately 17 hours on average, excluding data collection, at a marginal cost of about $114 USD per project (not including participant payments). This is a stark contrast to human-led research, which can take weeks or months. While highly efficient, the researchers acknowledge limitations, such as occasional imperfections in data visualizations and the “anchoring bias” where early conceptual errors can persist. The paper also raises important questions about the future of science, including the changing role of human researchers, the potential for democratizing high-quality research, and ethical considerations like scientific credit and the responsible use of autonomous systems. The ability of AI to generate and interpret real-world data marks a significant step towards AI actively participating in empirical investigations of the natural world. For more details, you can read the full research paper here.


