Model leaderboard

Ranked by what FreeTheAI actually served. Every number here is measured on real requests through our gateway, never taken from vendors or benchmarks.

Updated Sep 30, 2026, 2:19 PM UTC · 37.5K requests · 964.7M tokens in this window

95.1%Top success ratefta/mmx/minimax-m3
1333 msFastest median latencyfta/kai/liquid/lfm-2.5-2.6b:free
285.4MMost tokens servedfta/zai/glm-5.3

Ranking

Models ordered by tokens served in the window. Success rate, latency, and speed sit beside each one so you can weigh use against reliability.

RankModelTrafficReliabilityPrice per 1MSpeed
1
fta/zai/glm-5.3Zhipu AI · via Z.AI Coding · 1.0M context
285.4M29.6%11.5K requests
90.18% · 49sFree
55 t/s
2
fta/kimi/k3Moonshot · via Kimi Coding · 1.0M context
236.8M24.5%7.3K requests
90.50% · 26sFree
36 t/s
3
fta/zai/glm-5.2Zhipu AI · via Z.AI Coding · 1.0M context
131.5M13.6%6K requests
88.77% · 50sFree
56 t/s
4
fta/zai/glm-5.3-flashZhipu AI · via Z.AI Coding · 1.0M context
69M7.2%2.1K requests
91.82% · 40sFree
39 t/s
5
fta/bbl/gemini-3.5-flashGoogle · via FreeTheAI Mix · 1.0M context
62.5M6.5%1.4K requests
93.78% · 24sFree
373 t/s
6
fta/zai/glm-5.1Zhipu AI · via Z.AI Coding · 200.0K context
28.2M2.9%1.1K requests
92.48% · 45sFree
55 t/s
7
fta/mmx/minimax-m3MiniMax · via MiniMax · 1.0M context
25.5M2.6%1.3K requests
95.14% · 14sFree
60 t/s
8
fta/zai/glm-5Zhipu AI · via Z.AI Coding · 204.8K context
23.1M2.4%1.1K requests
90.53% · 32sFree
50 t/s
9
fta/zai/glm-4.7Zhipu AI · via Z.AI Coding · 204.8K context
18.9M2.0%867 requests
86.86% · 55sFree
37 t/s
10
fta/olm/kimi-k2.7-codeMoonshot · via Ollama Cloud · 262.1K context
18M1.9%926 requests
90.18% · 26sFree
67 t/s

Model makers

Share of tokens served in the window, by who made the model.

  • Zhipu AI59.3%
  • Moonshot26.4%
  • Google8.1%
  • MiniMax2.6%
  • xAI1.3%
  • Others2.3%

Methodology

Only our own traffic

Every request through the FreeTheAI gateway is counted once into time buckets: 5-minute, hourly, and daily. The 7-day and 30-day views use whole UTC days. No vendor numbers, benchmarks, or votes are used.

How rows are ranked

Models are ranked by tokens served (input, cache reads, and output). Only models listed in the catalog appear. The headline cards need at least 20 requests in the window.

What the numbers mean

Success rate counts requests the provider answered or failed; a caller's own bad request never counts against a model. Latency is the full request time; speed is output tokens over generation time after the first token.

Frequently asked questions

How is the leaderboard ranked?

By tokens FreeTheAI served for each model in the selected window. It is a measure of real use, not a quality score.

How fresh is it?

Requests are counted within about a minute, and the page refreshes its numbers every minute.

Why is a model missing?

A model appears once it is listed in the catalog and has served at least one request in the window.

Why does a free model rank so high?

Free models get a lot of traffic. Use the price filter to compare free and paid models separately; ranks stay the same so you can see where each sits overall.

Can I use this data?

Yes. Download it as JSON or CSV from the ranking section. Please credit FreeTheAI.