Skip to main content

Comparison Methodology

We document how prompts are chosen, which providers are in scope, and how editorial content separates verified facts from interpretive guidance. Product comparisons use identical user-supplied prompts—not undisclosed automated scoring.

Providers in scope

Live comparison integrates OpenAI, Anthropic, and Google AI via user API keys. Additional handlers may exist in code but require catalog and key configuration.

Evaluation criteria

Instruction following, accuracy on your prompts, reasoning clarity, coding usefulness, writing quality, latency, cost, privacy controls, and modality fit.

Limitations

Model versions change. We do not publish numeric leaderboard scores unless tied to a documented repeatable test. User results vary by prompt and configuration.

Corrections

Factual corrections are logged with updated dates on affected pages. See editorial and fact-checking policies.

Last verified: 2026-06-04

Compare models in the app

Use BYOK to test providers on your real prompts.