started · updated
Artificial Analysis launches Optima for custom AI benchmarking
Artificial Analysis has launched Optima, a new platform designed to address limitations in public AI benchmarks by allowing users to create customized evaluations for specific use cases. While standard benchmarks use predefined tasks, Optima enables users to compare large language models (LLMs) based on their own data, workflows, or specific use-case descriptions.
Users can build benchmarks using several methods: uploading existing datasets from local files or Hugging Face, importing AI agent traces from platforms such as Arize, Braintrust, or Langfuse, or utilizing a developer skill that gathers information from coding environments. For those without existing data, the platform can generate suggested test inputs and evaluation criteria based on a description of the intended task.
The platform evaluates models not only on output quality but also on practical metrics including cost per task and processing time per task. Scoring is conducted through either rubric-based evaluations against objective criteria or a pairwise comparison method, where user preferences on sample response pairs are used to derive full rankings across a dataset. Optima is currently available for general use.