Benchmarking LLM-based agents for single-cell omics analysis

Бенчмаркинг агентов на основе больших языковых моделей для анализа одноклеточных омиксных данных
Yang Liu, Lu Zhou, Xiawei Du, Ruikun He, Xuguang Zhang, Rongbo Shen, Yixue Li
2026-02-25

LLM-based agentsbenchmarking evaluation systemmulti-agent frameworksretrieval-augmented generationsingle-cell omics analysis
BACKGROUND: The surge in single-cell omics data exposes limitations in traditional, manually defined analysis workflows. AI agents offer a paradigm shift, enabling adaptive planning, executable code generation, traceable decisions, and real-time knowledge fusion. However, the lack of a comprehensive benchmark critically hinders progress. RESULTS: We introduce a novel benchmarking evaluation system to rigorously assess agent capabilities in single-cell omics analysis. This system comprises: a unified platform compatible with diverse agent frameworks and LLMs; multidimensional metrics assessing cognitive program synthesis, collaboration, execution efficiency, bioinformatics knowledge integration, and task completion quality; and 50 diverse real-world single-cell omics analysis tasks spanning multi-omics, species, and sequencing technologies. Our evaluation reveals that Grok3-beta achieves state-of-the-art performance among tested agent frameworks. Multi-agent frameworks significantly enhance collaboration and execution efficiency over single-agent approaches through specialized role division. Attribution analyses of agent capabilities identify that high-quality code generation is crucial for task success, and self-reflection has the most significant overall impact, followed by retrieval-augmented generation (RAG) and planning. CONCLUSIONS: This work highlights persistent challenges in code generation, long-context handling, and context-aware knowledge retrieval, providing a critical empirical foundation and best practices for developing robust AI agents in computational biology.
1
Grok3-beta achieves the best performance among the evaluated agent frameworks.
2
High-quality code generation is crucial for task success, while self-reflection has the greatest overall impact, followed by retrieval-augmented generation and planning; persistent challenges include long-context handling and context-aware retrieval.
3
Introduces a unified benchmark for evaluating LLM-based agents in single-cell omics analysis across diverse agent frameworks and language models.
4
Multi-agent systems improve collaboration and execution efficiency compared with single-agent approaches through specialized role division.
5
The benchmark includes multidimensional metrics and 50 real-world tasks spanning multi-omics, species, and sequencing technologies.

LLM-based AI agents for single-cell omics analysis

Agent capabilities and performance in planning, code generation, collaboration, knowledge integration, execution efficiency, and task completion for single-cell omics analysis

Publication Details
Publication Date
2026-02-25
Journal
Publisher
ISSN
Cited by
0
Access Type
Author Information
Authors
Yang Liu
Lu Zhou
Xiawei Du
Ruikun He
Xuguang Zhang
Rongbo Shen
Yixue Li
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%