ASTIR Introduces a Robotics Software Engineering Benchmark

ASTIR Introduces a Robotics Software Engineering Benchmark

  • 28. November 2025

ASTIR has developed the initial structure of a Robotics Software Engineering Benchmark for assessing AI-supported engineering methods and tools.

Generative AI systems are frequently evaluated using general coding or language benchmarks. These evaluations provide limited evidence about their suitability for robotics, where software interacts with physical systems, operational environments, sensors, actuators and safety constraints.

The ASTIR benchmark addresses this limitation by organising evaluation artefacts around real robotics software engineering activities. Its planned modules cover requirements, design, coding, testing and maintenance and are aligned with the project’s manufacturing, service robotics and drone use cases.

The benchmark follows principles of relevance, reproducibility, modularity, traceability and responsible data management. It also incorporates ethical, legal and social considerations so that technical performance is not treated separately from issues such as accountability, human oversight, cybersecurity and data protection.

During the next project phases, the benchmark will be used to support repeatable evaluation of ASTIR technologies. Where intellectual-property, security and data-protection conditions permit, benchmark resources will be prepared for reuse by the wider robotics and software engineering communities.