Install
Comparison and analysis of AI models and API hosting providers. Independent benchmarks across key performance metrics including quality, price, output speed & latency.
- 20articles · 30d
- 2+ day agolatest article
- Aug 20, 2026earliest in window
- 70%with images
- 245avg words
- science and technology 20
- EDJ 8
- CE 7
- SOD 7
- BI 5
- SCT 4
- JE 2
- LG 2
Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Announcing Artificial Analysis Intelligence Index v4.2
1+ week, 1+ day ago (557+ words) We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming + AA-Briefcase, our agentic knowledge work evaluation with a…...
Business Operations Manager
3+ week, 3+ day ago (157+ words) Our benchmarks and analysis are trusted by hundreds of thousands of users and are the go-to reference for leading AI labs including OpenAI, Google, Meta, NVIDIA and Anthropic, and major publications including the Wall Street Journal, Bloomberg, the Financial Times…...
GLM-5.3-Flash: API Provider Performance Benchmarking & Price Analysis
2+ week, 4+ day ago (648+ words) This analysis is intended to support you in choosing the best API provider of GLM-5.3-Flash for your use-case. Time to first answer token Blended price (per 1M tokens) GLM-5.3-Flash is available through 11 API providers, each offering different performance characteristics…...
Agnes 3.0 Flash: API Provider Performance Benchmarking & Price Analysis
2+ day, 13+ hour ago (611+ words) This analysis is intended to support you in choosing the best API provider of Agnes 3.0 Flash for your use-case. Time to first answer token Blended price (per 1M tokens) Agnes 3.0 Flash is available through Agnes AI. Update: Default performance benchmarking workload…...
Member of Technical Staff (Applied AI Research)
3+ week, 3+ day ago (499+ words) Our benchmarks and analysis are trusted by hundreds of thousands of users and are the go-to reference for leading AI labs including OpenAI, Google, Meta, NVIDIA and Anthropic, and major publications including the Wall Street Journal, Bloomberg, the Financial Times…...
Member of Technical Staff (Hardware)
3+ week, 3+ day ago (212+ words) Our benchmarks and analysis are trusted by hundreds of thousands of users and are the go-to reference for leading AI labs including OpenAI, Google, Meta, NVIDIA and Anthropic, and major publications including the Wall Street Journal, Bloomberg, the Financial Times…...
You are an administrative operations lead in a government de... | MicroEval
2+ week, 22+ hour ago (304+ words) You are an administrative operations lead in a government de... Artificial Analysis [+1] Addresses findings, risks, impacts or considerations identified in Key Findings in Implications for Government. [+1] Gives bullet points to each academic article’s Implications for Government. [+1] Gives bullet points of…...
Muse Spark 1.3: Meta reaches the frontier
1+ week, 3+ day ago (236+ words) Muse Spark 1.3 (xhigh) enters the Artificial Analysis Intelligence Index at 61, up 4 points from Muse Spark 1.2 (57, August) and 8 points from Muse Spark 1.1 (53, July). It enters tied with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high), and behind Claude Fable 5.1 (max,…...
Intelligence at pocket scale: Benchmarking small models and mobile phones
2+ week, 6+ day ago (559+ words) Models under a few billion parameters can now follow instructions, call tools, and answer questions on mobile phones. But available benchmark results often describe neither the quantized model artifact nor the device and runtime combination a user will run. Some…...
Artificial Analysis
3+ day, 16+ hour ago (571+ words) Independent evaluations of AI models across reasoning, knowledge, coding and agentic capabilities. Filter by the skills and knowledge domains. A composite benchmark aggregating ten challenging evaluations to provide a holistic measure of AI capabilities across mathematics, science, coding, and reasoning....