Model leaderboard

Ranked by what FreeTheAI actually served. Every number here is measured on real requests through our gateway, never taken from vendors or benchmarks.

Updated Sep 30, 2026, 1:23 PM UTC · 35.9K requests · 917.9M tokens in this window

95.4%Top success ratefta/mmx/minimax-m3
1166 msFastest median latencyfta/kai/poolside/laguna-xs-2.1:free
268.6MMost tokens servedfta/zai/glm-5.3

Ranking

Models ordered by tokens served in the window. Success rate, latency, and speed sit beside each one so you can weigh use against reliability.

RankModelTrafficReliabilityPrice per 1MSpeed
1
fta/zai/glm-5.3Zhipu AI · via Z.AI Coding · 1.0M context
268.6M29.3%11K requests
89.82% · 49sFree
55 t/s
2
fta/kimi/k3Moonshot · via Kimi Coding · 1.0M context
231M25.2%7K requests
90.29% · 26sFree
36 t/s
3
fta/zai/glm-5.2Zhipu AI · via Z.AI Coding · 1.0M context
122.4M13.3%5.7K requests
88.39% · 50sFree
56 t/s
4
fta/zai/glm-5.3-flashZhipu AI · via Z.AI Coding · 1.0M context
65.7M7.2%2K requests
91.66% · 39sFree
39 t/s
5
fta/bbl/gemini-3.5-flashGoogle · via FreeTheAI Mix · 1.0M context
59.2M6.4%1.4K requests
94.17% · 24sFree
375 t/s
6
fta/zai/glm-5.1Zhipu AI · via Z.AI Coding · 200.0K context
26.7M2.9%1.1K requests
92.19% · 45sFree
55 t/s
7
fta/mmx/minimax-m3MiniMax · via MiniMax · 1.0M context
24.6M2.7%1.3K requests
95.42% · 14sFree
60 t/s
8
fta/zai/glm-5Zhipu AI · via Z.AI Coding · 204.8K context
22.1M2.4%1.1K requests
90.44% · 31sFree
49 t/s
9
fta/zai/glm-4.7Zhipu AI · via Z.AI Coding · 204.8K context
18.6M2.0%851 requests
86.72% · 54sFree
37 t/s
10
fta/olm/kimi-k2.7-codeMoonshot · via Ollama Cloud · 262.1K context
16.8M1.8%879 requests
89.76% · 25sFree
67 t/s

Model makers

Share of tokens served in the window, by who made the model.

  • Zhipu AI58.7%
  • Moonshot27.0%
  • Google8.1%
  • MiniMax2.7%
  • xAI1.2%
  • Others2.3%

Methodology

Only our own traffic

Every request through the FreeTheAI gateway is counted once into time buckets: 5-minute, hourly, and daily. The 7-day and 30-day views use whole UTC days. No vendor numbers, benchmarks, or votes are used.

How rows are ranked

Models are ranked by tokens served (input, cache reads, and output). Only models listed in the catalog appear. The headline cards need at least 20 requests in the window.

What the numbers mean

Success rate counts requests the provider answered or failed; a caller's own bad request never counts against a model. Latency is the full request time; speed is output tokens over generation time after the first token.

Frequently asked questions

How is the leaderboard ranked?

By tokens FreeTheAI served for each model in the selected window. It is a measure of real use, not a quality score.

How fresh is it?

Requests are counted within about a minute, and the page refreshes its numbers every minute.

Why is a model missing?

A model appears once it is listed in the catalog and has served at least one request in the window.

Why does a free model rank so high?

Free models get a lot of traffic. Use the price filter to compare free and paid models separately; ranks stay the same so you can see where each sits overall.

Can I use this data?

Yes. Download it as JSON or CSV from the ranking section. Please credit FreeTheAI.