AI model leaderboard

Text, image and video models ranked on one page, on published third-party scores. The text table is every model the official plan serves; image and video are the boards themselves, reproduced.

OHarness runs no model evaluations. Every score is copied from a published third-party board, cited per table with the date it was read — 19 of the 44 models on the official plan carry one, and the rest show a dash rather than an estimate. The model list comes from the gateway that serves them, so nothing here advertises a model you cannot call. Last read 2026-09-16.

Esta página está escrita em inglês. Os termos que explica são os que se procuram em inglês, e um termo técnico traduzido é um termo diferente.

Claude Fable 5.1
Top text model · 53 Intelligence Index
GPT Image 2.5 Flare
Top image model · 1189 Arena Elo
Seedance 2.5
OHarness pick · 1482 Arena Elo

Ranked by modality

Scores only mean something within a modality — an intelligence index and an Elo rating measure different things by different methods — so each table ranks only against itself. Sort any column.

Text

Every model the official plan serves, scored by the index. Sort by either column.

Text models ranked by quality. Every model the official plan serves, scored by the index. Sort by either column.
#Model
1Claude Fable 5.1Anthropic5365 t/s
2GPT-6 AstraOpenAI5353 t/s
3Claude Opus 5Anthropic5149 t/s
4Claude Fable 5Anthropic5065 t/s
5GPT-5.6 SolOpenAI4765 t/s
6GLM 5.3Z.ai (GLM)4567 t/s
7Grok 4.6xAI4458 t/s
8Kimi K3Moonshot (Kimi)4435 t/s
9GPT-5.6 TerraOpenAI4299 t/s
10Gemini 3.8 FlashGoogle Gemini41332 t/s
11Qwen3.8 MaxQwen (Alibaba Cloud)4040 t/s
12Gemini 3.7 FlashGoogle Gemini39289 t/s
13Grok 4.5xAI3960 t/s
14Claude Sonnet 5Anthropic3880 t/s
15GPT-5.6 LunaOpenAI38115 t/s
16GLM 5.2Z.ai (GLM)3471 t/s
17DeepSeek V4 ProDeepSeek3186 t/s
18Gemini 3.1 Pro Preview1Google Gemini30108 t/s
19Claude Haiku 4.5Anthropic1885 t/s
20Claude Opus 4.8Anthropic——
21DeepSeek V4 FlashDeepSeek——
22DeepSeek V4.1 FlashDeepSeek——
23Gemini 2.5 FlashGoogle Gemini——
24Gemini 3.1 Flash-LiteGoogle Gemini——
25Gemini 3.5 FlashGoogle Gemini——
26Gemini 3.5 Flash-LiteGoogle Gemini——
27GLM 5.3 FlashZ.ai (GLM)——
28GPT-4oOpenAI——
29GPT-5.2OpenAI——
30GPT-5.4OpenAI——
31GPT-5.4 miniOpenAI——
32GPT-5.4 nanoOpenAI——
33GPT-5.5OpenAI——
34GPT-6 Astra ProOpenAI——
35Grok 4.7xAI——
36Kimi K2 0711Moonshot (Kimi)——
37Kimi K2 0905Moonshot (Kimi)——
38Kimi K2 ThinkingMoonshot (Kimi)——
39Kimi K2.5Moonshot (Kimi)——
40Kimi K2.6Moonshot (Kimi)——
41Kimi K2.7 CodeMoonshot (Kimi)——
42Qwen3.8 FlashQwen (Alibaba Cloud)——
43Sonar Deep ResearchPerplexity——
44Video UnderstandingOttoPort——
  1. 1Gemini 3.1 Pro Preview Scored as Gemini 3.1 Pro; the served entry is the preview build of the same model.

Image

Text-to-image models, ranked by blind head-to-head votes. OHarness does not serve these — the gateway answers completions and nothing else — so the board is reproduced here rather than sold.

Image models ranked by quality. Text-to-image models, ranked by blind head-to-head votes. OHarness does not serve these — the gateway answers completions and nothing else — so the board is reproduced here rather than sold.
#Model
1GPT Image 2.5 Flare1OpenAI1189
2GPT Image 2.52OpenAI1183
3GPT Image 23OpenAI1173
4Nano Banana 2Google1122
5GPT Image 1.54OpenAI1104
6Nano Banana ProGoogle1099
7Nano Banana LiteGoogle1088
8Qwen Image 3.0 ProAlibaba1087
9Seedream 5.0 ProByteDance1082
10Qwen Image 3.0Alibaba1073
11Seedream 5.0 LiteByteDance1000
12Wan 2.7 ProAlibaba979
13GPT Image 1 Mini5OpenAI919
14Midjourney v76Midjourney897
  1. 1GPT Image 2.5 Flare Ranked at the max quality setting.
  2. 2GPT Image 2.5 Board entry: GPT Image 2.5 Sunburst, max quality.
  3. 3GPT Image 2 Ranked at the high quality setting.
  4. 4GPT Image 1.5 Ranked at the high quality setting.
  5. 5GPT Image 1 Mini Ranked at the medium quality setting.
  6. 6Midjourney v7 Board entry: Midjourney v7 Alpha.

Video

Text-to-video models, ranked by blind head-to-head votes. Reproduced for the same reason as the image board: OHarness has no endpoint for them.

Video models ranked by quality. Text-to-video models, ranked by blind head-to-head votes. Reproduced for the same reason as the image board: OHarness has no endpoint for them.
#Model
1Seedance 2.51OHarness pickByteDance1482
2Gemini Omni 1.1 FlashGoogle1515
3Gemini Omni FlashGoogle1511
4FLUX 3 VideoBlack Forest Labs1494
5Wan 3.02Alibaba1494
6Seedance 2.03ByteDance1479
7MiniMax H3MiniMax1462
8Veo 3.14Google1364
9Veo 3.1 Fast5Google1362
10Wan 2.7Alibaba1341
11Wan 2.6Alibaba1328
12Wan 2.56Alibaba1246
13Hailuo 2.3MiniMax1205
14Hailuo 027MiniMax1181

Seedance 2.5 is placed first by OHarness, not by the board. OHarness places it first; the board does not. ByteDance's current flagship — native 30-second clips, 4K, up to 50 reference images — and the one model in this table whose measured Elo is visibly still catching up to it: arena.ai has it fifth at 1482, and Artificial Analysis has not scored it at all, because both boards are accumulating votes on a model that shipped after they last settled. The 1482 in its row is the board's number, unchanged. Only the position is ours.

  1. 1Seedance 2.5 Board entry: dreamina-seedance-2.5-720p. Not yet ranked by Artificial Analysis.
  2. 2Wan 3.0 Board entry: wan3.0. The Prime tier is not separately ranked.
  3. 3Seedance 2.0 Board entry: dreamina-seedance-2.0-720p.
  4. 4Veo 3.1 Board entry: veo-3.1-audio at 720p; the 1080p entry scores 1363.
  5. 5Veo 3.1 Fast Board entry: veo-3.1-fast-audio at 720p.
  6. 6Wan 2.5 Board entry: wan2.5-t2v-preview.
  7. 7Hailuo 02 Board entry: hailuo-02-standard; the Pro tier scores 1198.

How these rankings are built

Quality and speed: not ours

Text models carry the Artificial Analysis Intelligence Index; image and video carry Arena Elo from blind head-to-head votes. Elo pools are not interchangeable, so each column comes from exactly one board, named with the date it was read. The logos are their owners’ trademarks, here to identify the models beside them.

No prices, on purpose

A money column on a page about which model is best asks the reader to weigh two different questions in one row, and they get a worse answer to both. What each model costs — both rates, per million tokens, as the gateway meters them — is on the pricing page, which is a rate card and says so.

Picks: the one opinion here

Boards lag the market — a model can ship, lead its category, and wait months for an arena to gather enough votes to rank it. Where that happens the model is pinned to the top of the default view, badged, with the reason printed under the table. A pin never writes a score into a cell, and it disappears the moment you reverse the sort.

Image and video: reference only

The official plan serves chat completions and nothing else. Those two tables are reproduced from their boards so the field is visible, with no link into a playground that does not exist here. If that ever changes, they become tables of models you can actually call.

Questions

Who produced these scores?

Nobody here. Every quality figure is copied from a published third-party leaderboard, cited per table with the date it was read. OHarness sells access to some of these models, and a ranking from a seller is only worth reading if it says which numbers it measured and which it copied. It measured none of them.

Can I use the image and video models through OHarness?

No. The official plan serves chat completions and nothing else — every other path returns 404 — so those two tables reproduce their boards rather than advertising a catalogue that does not exist here. The text table is the one you can act on.

Why is there no price column?

Because it would make both questions harder to answer. Which model is best and what it costs are separate decisions, and a table that mixes them nudges the reader into ranking by money without saying so. The rate card is on the pricing page, where every model is listed at both rates, fetched from the gateway that meters the call.

Why is a model I use missing?

Either the plan does not serve it, or no board had published a score for it when this table was last read. A model with no score still appears in the text table with a dash: it is a model you can call, and an estimate in that cell would be worth less than the blank.

Why is a model ranked above one with a higher score?

Because it is pinned, and the badge on the row says so. It is the only judgement on this page: arenas lag the market, and a model can lead its category for months before a board has the votes to rank it there. A pin moves a row and nothing else — the score in the cell is still the board's, the reason is printed under the table, and reversing the sort drops the pin entirely.

Can I compare the numbers across tables?

No, and nothing on this page does. Text carries an intelligence index out of roughly a hundred; image and video carry Arena Elo, which has no meaningful zero and comes from a separate pool of votes per board. The scores rank within a table and mean nothing between them.