The method
The easiest way to avoid stale lists is to ask OpenRouter directly. I used the public openrouter.ai/api/v1/models endpoint and filtered models where pricing.prompt == 0 and pricing.completion == 0.
On July 1, 2026, that filter returned 25 free models. Free does not mean unlimited, stable or production-ready. It simply means OpenRouter reports zero prompt and completion token pricing at the time of the snapshot.
Availability, rate limits and terms can change without warning. For any serious project, check the model in OpenRouter again right before using it.
Models to test first
The full list is useful, but I would not start by testing everything. I would pick a few candidates by use case.
| Use case | Model | Why |
|---|---|---|
| Coding agent | qwen/qwen3-coder:free | Large context, code-oriented, a strong candidate for agents and tools. |
| Long context | nvidia/nemotron-3-ultra-550b-a55b:free | 1M context window and 500B+ parameters, interesting for long documents, large repositories or exploratory RAG. |
| General baseline | meta-llama/llama-3.3-70b-instruct:free | A known model family, useful as a simple comparison point. |
| OpenAI open model | openai/gpt-oss-120b:free | Worth testing against other open-weight families. |
| Multimodal | google/gemma-4-26b-a4b-it:free | Text, image and video inputs to text output, handy for experimentation without a budget. |
| Automatic router | openrouter/free | Lets OpenRouter route to an available free model. |
For coding and AI agents, these free models can be ideal for testing an agentic loop at low cost, for example with a desktop tool like Hermes Agent. I would compare this with my tests of ZCode with GLM-5.2, MiMo Code and my 2026 LLM coding comparison. The model is no longer the only question: the surrounding harness, tools, project context, cost limits and iteration quality matter just as much.
The complete list of 25 free models
Here is the raw snapshot, sorted by context size. I keep the exact OpenRouter identifiers so you can copy them into your tests.
| Model | Context | Modality |
|---|---|---|
google/lyria-3-clip-preview | 1M | text+image → text+audio |
google/lyria-3-pro-preview | 1M | text+image → text+audio |
qwen/qwen3-coder:free | 1M | text → text |
nvidia/nemotron-3-super-120b-a12b:free | 1M | text → text |
nvidia/nemotron-3-ultra-550b-a55b:free | 1M | text → text |
google/gemma-4-26b-a4b-it:free | 262K | text+image+video → text |
google/gemma-4-31b-it:free | 262K | text+image+video → text |
poolside/laguna-m.1:free | 262K | text → text |
poolside/laguna-xs.2:free | 262K | text → text |
qwen/qwen3-next-80b-a3b-instruct:free | 262K | text → text |
cohere/north-mini-code:free | 256K | text → text |
nvidia/nemotron-3-nano-30b-a3b:free | 256K | text → text |
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 256K | text+image+audio+video → text |
openrouter/free | 200K | text+image → text |
meta-llama/llama-3.2-3b-instruct:free | 131K | text → text |
meta-llama/llama-3.3-70b-instruct:free | 131K | text → text |
nousresearch/hermes-3-llama-3.1-405b:free | 131K | text → text |
openai/gpt-oss-120b:free | 131K | text → text |
openai/gpt-oss-20b:free | 131K | text → text |
nvidia/nemotron-3.5-content-safety:free | 128K | text+image → text |
nvidia/nemotron-nano-12b-v2-vl:free | 128K | text+image+video → text |
nvidia/nemotron-nano-9b-v2:free | 128K | text → text |
cognitivecomputations/dolphin-mistral-24b-venice-edition:free | 32K | text → text |
liquid/lfm-2.5-1.2b-instruct:free | 32K | text → text |
liquid/lfm-2.5-1.2b-thinking:free | 32K | text → text |
The limits of free access
Free OpenRouter models are great for discovery, quick tests and small benchmarks. I would not treat them as a production foundation without a fallback.
- Rate limits and availability. A free model may be limited or unavailable depending on the provider behind OpenRouter.
- Variable quality. Some models are strong on one task and much weaker on long instructions or agentic workflows.
- Theoretical context. A 1M-token context window does not guarantee the model will use all of it well.
- Data and training. OpenRouter states that some models or providers may store inputs, or even use them to improve or train their own models. Before sending client code, sensitive data or a strategic prompt, check the provider policy, retention tags and Zero Data Retention options.
- Moving conditions. A model that is free today may change status tomorrow.
Use them as a playground and a selection filter. For stable work, keep a paid fallback, measure failures and monitor costs.
My starting point
If I had to pick one model to start with, I would probably use nvidia/nemotron-3-ultra-550b-a55b:free. It is powerful, has a very large context window, and becomes especially interesting once you put it inside a good harness: clean instructions, controlled project context, well-exposed tools, simple evaluations and a fast iteration loop.
That is what matters most now. Most recent models are already strong enough to succeed in a very wide range of situations. The difference is less and less about the model name alone, and more and more about the system around it: prompts, MCP, skills, memory, routing, tests, cost control and feedback quality.
OpenRouter is valuable as a comparison playground for free models. Not because it helps you find a magic model, but because it lets you identify good candidates, plug them into a complete harness, and measure which one gives the best balance of quality, speed, stability and cost. Once a model fits a serious use case, though, you should move to paid endpoints to get a real service guarantee, availability and throughput.



