Back to research & evidence

Evaluation record

Explore our benchmark results, upcoming evaluations, and research goals for software engineering and autonomous systems.

20-test subset result

SWE-bench Pro

20 of 20 tests resolved on initial attempts using the Taijitu AI Agent & Model.

  • Result: 100% resolution on initial attempts across a 20-test subset.
  • Planned evaluations: expansion to the 721-test suite and a comparison using SWE Agent paired with the Taijitu AI Model.

Planned evaluation

WebArena 2.0

Multi-step browser tasks, including navigation and interaction with changing web environments.

  • Evaluation focus: task completion through browser navigation, page interaction, and changes in environment state.

AOCA and model research

AOCA combines high-throughput time-series processing with autonomous operation. Its forthcoming AI Ops benchmark focuses on diagnostic quality and speed. Our model reliability and efficiency programme is in research and development.

Evaluation priorities for model research include task quality and the total cost of successful completion: compute, time to completion, failed attempts, retries, and operating effort.