How can the benefits of generative AI for robotics software engineering be measured objectively? ASTIR has established an evaluation framework to answer this question.
The framework defines how the project will assess the performance, reliability, usability and practical added value of its technologies. Rather than evaluating individual AI outputs in isolation, ASTIR will examine complete engineering tasks in the project’s manufacturing, service robotics and drone use cases.
The approach adapts the Goal–Question–Metric method. Project objectives are translated into evaluation questions and measurable indicators covering areas such as development effort, software quality, reliability, security, test coverage, usability and developer confidence.
Depending on the engineering activity, evaluations may use before-and-after comparisons, parallel task execution, automated measurements, expert assessment and structured user feedback. This combination is intended to capture both quantitative improvements and factors that are harder to measure, such as whether generated explanations are understandable and useful to engineers.
The evaluation framework will be refined as the ASTIR components mature. It provides a common basis for demonstrating where AI augmentation offers genuine engineering value and where further technical or human validation remains necessary.

