Business-facing domains
186 fine-grained scenarios covering professional services, industry workflows, and enterprise operations.
Benchmarking General AI Agents Across Diverse Scenarios
Peking University · Renmin University of China · Tsinghua University · Beijing Institute of Technology
Huawei Cloud Post-Training Team · Zhongguancun Academy
Contactscuuy05@gmail.com
Current frontier performance on the 644-task challenging set.
Overall is Pass@1 across all 644 tasks (micro-average). Similar final scores can reflect very different execution profiles.
| Rank | Model | Access | Effort | DAG | Solver | Program | DAG-S | Overall |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude-Sonnet-5Anthropic | Proprietary | max | 57.34 | 56.67 | 63.33 | 60.50 | 58.54 |
| 2 | GPT-5.6-solOpenAI | Proprietary | high | 55.37 | 65.00 | 50.00 | 59.00 | 57.14 |
| 3 | GLM-5.2GLM | Open | max | 54.80 | 26.67 | 60.00 | 69.00 | 56.83 |
| 4 | GPT-5.5OpenAI | Proprietary | high | 54.80 | 38.33 | 60.00 | 64.50 | 56.52 |
| 5 | DeepSeek-V4-ProDeepSeek | Open | max | 52.54 | 36.67 | 53.33 | 63.50 | 54.50 |
| 6 | Claude-Opus-4.7Anthropic | Proprietary | max | 53.39 | 43.33 | 63.33 | 57.50 | 54.19 |
| 7 | Kimi-K2.6-ThinkingKimi | Open | — | 49.72 | 45.00 | 63.33 | 57.50 | 52.33 |
| 8 | Qwen3.7-MaxQwen | Proprietary | — | 48.59 | 51.67 | 66.67 | 48.50 | 49.69 |
Choose up to six models. The default selection matches the five models shown in the paper.
Large language models are evolving from text generators into general agents that understand requests, invoke external tools, and complete complex tasks through interaction. Existing benchmarks, however, often focus on limited scenarios, tool ecosystems, or interaction formats.
OmniaBench evaluates agents across diverse application settings with explicit state spaces. Its taxonomy spans consumer, business, and employee settings, while executable environments require planning, tool use, state maintenance, and adaptation to feedback.

186 fine-grained scenarios covering professional services, industry workflows, and enterprise operations.
101 scenarios derived from real app ecosystems and everyday user needs.
67 scenarios grounded in industry-general knowledge work and workplace tasks.
The 644-task challenging set contains 354 DAG, 200 DAG-S, 60 Solver, and 30 Program tasks, all sharing executable environments and structured evaluation.
Multi-turn stateful interaction and tool-chain execution.
Single-turn tasks derived through query refinement.
Selection, scheduling, allocation, and optimization.
Procedural reasoning with branching, iteration, and debugging.

Ten capability dimensions reveal where agents succeed, where they fail, and why similar overall scores can conceal different practical strengths.


Claude-Sonnet-5 reaches 58.54 Overall Pass@1, showing substantial headroom even for the strongest evaluated systems.
Planning, decomposition, constraint maintenance, reflection, and adaptive correction account for most observed failures.
Longer trajectories often reflect redundant exploration or repeated replanning rather than more thorough execution.
Performance shifts substantially across business, consumer, employee, and fine-grained application domains.

The 644-task challenging set is available on Hugging Face, together with the open evaluation harness and interactive leaderboard.
Open the dataset ↗If OmniaBench supports your research, please cite our arXiv paper.
@misc{shen2026omniabenchbenchmarkinggeneralai,
title={OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios},
author={Chengyu Shen and Yujie Fu and Gangtao Xin and Yanheng Hou and Wenlong Fei and Guojie Zhu and Jiawei Li and Hongcheng Gao and Runming He and Zhen Hao Wong and Meiyi Qiang and Hao Liang and Zhao Cao and Hao Jiang and Chong Chen and Wentao Zhang},
year={2026},
eprint={2607.14989},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2607.14989},
}