Model leaderboard
Ranked by what FreeTheAI actually served. Every number here is measured on real requests through our gateway, never taken from vendors or benchmarks.
Updated Sep 30, 2026, 1:23 PM UTC · 35.9K requests · 917.9M tokens in this window
Ranking
Models ordered by tokens served in the window. Success rate, latency, and speed sit beside each one so you can weigh use against reliability.
| Rank | Model | Traffic | Reliability | Price per 1M | Speed |
|---|---|---|---|---|---|
| 1 | fta/zai/glm-5.3Zhipu AI · via Z.AI Coding · 1.0M context | 268.6M29.3%11K requests | 89.82% · 49s | Free | 55 t/s |
| 2 | fta/kimi/k3Moonshot · via Kimi Coding · 1.0M context | 231M25.2%7K requests | 90.29% · 26s | Free | 36 t/s |
| 3 | fta/zai/glm-5.2Zhipu AI · via Z.AI Coding · 1.0M context | 122.4M13.3%5.7K requests | 88.39% · 50s | Free | 56 t/s |
| 4 | fta/zai/glm-5.3-flashZhipu AI · via Z.AI Coding · 1.0M context | 65.7M7.2%2K requests | 91.66% · 39s | Free | 39 t/s |
| 5 | fta/bbl/gemini-3.5-flashGoogle · via FreeTheAI Mix · 1.0M context | 59.2M6.4%1.4K requests | 94.17% · 24s | Free | 375 t/s |
| 6 | fta/zai/glm-5.1Zhipu AI · via Z.AI Coding · 200.0K context | 26.7M2.9%1.1K requests | 92.19% · 45s | Free | 55 t/s |
| 7 | fta/mmx/minimax-m3MiniMax · via MiniMax · 1.0M context | 24.6M2.7%1.3K requests | 95.42% · 14s | Free | 60 t/s |
| 8 | fta/zai/glm-5Zhipu AI · via Z.AI Coding · 204.8K context | 22.1M2.4%1.1K requests | 90.44% · 31s | Free | 49 t/s |
| 9 | fta/zai/glm-4.7Zhipu AI · via Z.AI Coding · 204.8K context | 18.6M2.0%851 requests | 86.72% · 54s | Free | 37 t/s |
| 10 | fta/olm/kimi-k2.7-codeMoonshot · via Ollama Cloud · 262.1K context | 16.8M1.8%879 requests | 89.76% · 25s | Free | 67 t/s |
Model makers
Share of tokens served in the window, by who made the model.
- Zhipu AI58.7%
- Moonshot27.0%
- Google8.1%
- MiniMax2.7%
- xAI1.2%
- Others2.3%
Methodology
Only our own traffic
Every request through the FreeTheAI gateway is counted once into time buckets: 5-minute, hourly, and daily. The 7-day and 30-day views use whole UTC days. No vendor numbers, benchmarks, or votes are used.
How rows are ranked
Models are ranked by tokens served (input, cache reads, and output). Only models listed in the catalog appear. The headline cards need at least 20 requests in the window.
What the numbers mean
Success rate counts requests the provider answered or failed; a caller's own bad request never counts against a model. Latency is the full request time; speed is output tokens over generation time after the first token.
Frequently asked questions
How is the leaderboard ranked?
By tokens FreeTheAI served for each model in the selected window. It is a measure of real use, not a quality score.
How fresh is it?
Requests are counted within about a minute, and the page refreshes its numbers every minute.
Why is a model missing?
A model appears once it is listed in the catalog and has served at least one request in the window.
Why does a free model rank so high?
Free models get a lot of traffic. Use the price filter to compare free and paid models separately; ranks stay the same so you can see where each sits overall.
Can I use this data?
Yes. Download it as JSON or CSV from the ranking section. Please credit FreeTheAI.