
Optima by Artificial Analysis
Share
Optima by Artificial Analysis
Create custom AI model comparisons with your own data. Evaluate performance, cost, and speed to choose the most efficient and cost-effective option.
General Information about Optima by Artificial Analysis
Optima by Artificial Analysis is an advanced platform designed for creating custom benchmarks that evaluate artificial intelligence models based on specific use cases. Unlike standardized tests that measure general capabilities, this tool facilitates the comparison of Large Language Models (LLMs) under specific performance, cost, and time-efficiency criteria. It is a solution geared toward businesses and developers who need to determine which technology best fits their actual workflows, such as contract review, data analysis, or technical content generation.
Optima’s operation is structured into three critical phases: construction, execution, and grading. In the initial phase, the user provides context through examples, files, or data traces. An integrated construction agent assists in drafting tasks and evaluation rubrics, ensuring the benchmark accurately reflects the workload intended for automation. The platform allows users to import data directly from their computer or use programming agents to define complex tasks.
The tool supports various evaluation styles to adapt to different technical needs:
- Objective Q&A: For answers with a known correct solution where grading is deterministic.
- Document Analysis: Evaluating model performance when processing and extracting information from uploaded files.
- Agentic Tasks: Tests where the model must use tools and complete deliverables autonomously.
- Interaction Simulation: Conversational scenarios to measure responses across different user profiles.
During execution, Optima subjects selected models (such as the Claude, GPT, or Gemini families) to the same tasks and conditions simultaneously. A distinguishing feature is the detailed logging of efficiency metrics, including exact token consumption, cost per task, and response time. This allows for the visualization of which model offers the best price-performance ratio for a specific project via scatter plots.
The grading phase uses AI-based judge panels or custom rubrics to score results. Users can choose between standard judges or frontier models for greater precision in subjective tasks. Additionally, the platform allows for the integration of proprietary agents via HTTP, making it easy for internally developed software to compete in the same benchmark against the most powerful commercial models.
The final result is a detailed, private leaderboard that eliminates uncertainty in AI adoption. Optima by Artificial Analysis provides empirical data on model behavior in production scenarios, optimizing technical and financial decision-making when choosing an inference provid
Features and Use Cases of Optima by Artificial Analysis
How Optima by Artificial Analysis Works
Frequently Asked Questions about Optima by Artificial Analysis
What exactly is Optima and what is it for?
Optima is a tool that allows you to create custom benchmarks to compare AI models based on their performance, cost, and speed for specific tasks.
How do I start building a benchmark in Optima?
To get started, simply describe your use case or import your own data and examples. The creation agent will then guide you through developing the tasks and evaluation rubrics.
What types of evaluations can I perform with Optima by Artificial Analysis?
You can choose from various task styles, such as Q&A sessions based on objective data, analysis of uploaded documents, agentic tasks that use tools, or real conversation simulations.
Can I compare my own AI agent against commercial models in Optima?
Yes, you have the option to connect your own agent via an HTTP request so it can compete under the same conditions and with the same judges as standard market models.
How are costs calculated when using Optima?
Test creation and execution are charged based on the actual token cost of the models used, while rubric-based grading has a fixed price per criterion or pairing.
What is the difference between standard and premium judges in evaluations?
Standard judges use high-capacity models for grading, while premium judges utilize more advanced frontier models that offer higher analytical precision.
What detailed information do I receive after running a benchmark with Optima?
Once finished, you will receive a comprehensive comparison that includes scores based on your custom rubric, the exact cost per task, and the execution time for each model analyzed.
Can I use my own documents for testing in Optima by Artificial Analysis?
Absolutely. You can upload specific files for the models to reference, allowing you to evaluate system accuracy using your actual corporate data.
Optima by Artificial Analysis Pricing
Optima (Pay-as-you-go)
This pricing model is based on token consumption and running evaluations to create custom benchmarks.
Benchmark creation and execution: Billed at the raw token cost of the models used, with no additional markups.
Rubric grading: $0.002 per criterion and model with standard judges, or $0.040 with premium judges.
Pairwise grading: $0.006 per pairing with standard judges, or $0.150 with premium judges.
Standard judges utilize high-capacity models, while premium judges use frontier models.
The system places a credit hold based on an estimate before each run, ultimately charging only for the tokens actually used.
Pro
$499 per month per seat (or $417 per month with annual billing).
API access for building tools.
Data export for local analysis.
Custom chart and table creation.
Access to industry reports and guides.
Email technical support.
Designed for individuals and small organizations.
Enterprise
Custom pricing (contact our official website).
Includes all Pro plan features.
Model, inference, and hardware benchmarking for proprietary use cases.
Higher API rate limits.
Access to compute market models with forecasting.
Workshops, education, and training.
AI strategy consulting.
Personalized support.
Geared toward organizations with 150+ employees.
Optima by Artificial Analysis Screenshots

