Model leaderboard

Ranked by what FreeTheAI actually served. Every number here is measured on real requests through our gateway, never taken from vendors or benchmarks.

Updated Sep 30, 2026, 2:19 PM UTC · 29.9K requests · 806.6M tokens in this window

1333 msFastest median latencyfta/kai/liquid/lfm-2.5-2.6b:free
235.7MMost tokens servedfta/zai/glm-5.3

Ranking

Models ordered by tokens served in the window. Success rate, latency, and speed sit beside each one so you can weigh use against reliability.

RankModelTrafficReliabilityPrice per 1MSpeed
1
fta/zai/glm-5.3Zhipu AI · via Z.AI Coding · 1.0M context
235.7M29.2%9K requests
90.66% · 47sFree
57 t/s
2
fta/kimi/k3Moonshot · via Kimi Coding · 1.0M context
199.1M24.7%5.5K requests
89.98% · 25sFree
37 t/s
3
fta/zai/glm-5.2Zhipu AI · via Z.AI Coding · 1.0M context
109.9M13.6%4.8K requests
89.28% · 50sFree
57 t/s
4
fta/zai/glm-5.3-flashZhipu AI · via Z.AI Coding · 1.0M context
58M7.2%1.8K requests
92.88% · 42sFree
39 t/s
5
fta/bbl/gemini-3.5-flashGoogle · via FreeTheAI Mix · 1.0M context
51.3M6.4%1.2K requests
93.36% · 24sFree
379 t/s
6
fta/zai/glm-5.1Zhipu AI · via Z.AI Coding · 200.0K context
22.6M2.8%868 requests
93.73% · 37sFree
56 t/s
7
fta/mmx/minimax-m3MiniMax · via MiniMax · 1.0M context
21.7M2.7%1.1K requests
94.55% · 14sFree
63 t/s
8
fta/zai/glm-5Zhipu AI · via Z.AI Coding · 204.8K context
19.9M2.5%873 requests
91.80% · 31sFree
50 t/s
9
fta/zai/glm-4.7Zhipu AI · via Z.AI Coding · 204.8K context
16.2M2.0%651 requests
88.51% · 56sFree
37 t/s
10
fta/olm/kimi-k2.7-codeMoonshot · via Ollama Cloud · 262.1K context
15.2M1.9%737 requests
88.38% · 25sFree
67 t/s

Model makers

Share of tokens served in the window, by who made the model.

  • Zhipu AI58.8%
  • Moonshot26.6%
  • Google7.9%
  • MiniMax2.7%
  • xAI1.4%
  • Others2.6%

Methodology

Only our own traffic

Every request through the FreeTheAI gateway is counted once into time buckets: 5-minute, hourly, and daily. The 7-day and 30-day views use whole UTC days. No vendor numbers, benchmarks, or votes are used.

How rows are ranked

Models are ranked by tokens served (input, cache reads, and output). Only models listed in the catalog appear. The headline cards need at least 20 requests in the window.

What the numbers mean

Success rate counts requests the provider answered or failed; a caller's own bad request never counts against a model. Latency is the full request time; speed is output tokens over generation time after the first token.

Frequently asked questions

How is the leaderboard ranked?

By tokens FreeTheAI served for each model in the selected window. It is a measure of real use, not a quality score.

How fresh is it?

Requests are counted within about a minute, and the page refreshes its numbers every minute.

Why is a model missing?

A model appears once it is listed in the catalog and has served at least one request in the window.

Why does a free model rank so high?

Free models get a lot of traffic. Use the price filter to compare free and paid models separately; ranks stay the same so you can see where each sits overall.

Can I use this data?

Yes. Download it as JSON or CSV from the ranking section. Please credit FreeTheAI.