๐Ÿง  Free AI Coding Model Landscape

August 2026 ยท Verified against live OpenRouter API, OpenCode config, free-coding-models sources.js, Kilo model registry, and koda agent config ยท Generated 2026-08-08
๐Ÿ“‹ Contents
  1. Executive Summary
  2. Current Free Model Availability โ€” Card Catalog
  3. Installed CLI Free Tier Map
  4. Frontier Models Accessible via Free Tiers
  5. Investigation: Meta Muse Spark 1.2 Free Hosting
  6. Model Comparison โ€” Workflow Scoring (with SWE-bench Scores)
  7. Agent Role Recommendations โ€” Free Agent Stack
  8. Local vs Hosted Comparison
  9. Risks & Caveats
  10. Final Recommendations
  11. Sources & Methodology

1. Executive Summary

The free AI coding model landscape in August 2026 is extraordinarily rich. Your installed CLIs โ€” OpenCode, Kilo, Grok, Koda, Gemini CLI, Mimo CLI, CodeBuddy, and the free-coding-models discovery tool โ€” collectively provide access to 221+ models across 10+ free-tier providers. This includes frontier-coding models (S+ tier with 80%+ SWE-bench Verified scores) routed through NVIDIA NIM's free tier, dedicated coding specialists from Poolside and Cohere, and local models via Ollama on Fedora.
The standout story of mid-2026: free tiers now reach frontier quality. NVIDIA NIM's free catalog includes GLM 5.2 (82.8% SWE-bench), Kimi K2.6 (80.2%), DeepSeek V4 Pro (80.6%), and Nemotron 3 Ultra (71.9%). Combined with Poolside's Laguna S 2.1 (coding-specialist) and Google Gemini's 1M-context Flash models, you can assemble a multi-agent coding stack that rivals paid setups โ€” without spending a dollar. The catch: rate limits and availability volatility.
Key findings:

2. Current Free Model Availability โ€” Card Catalog

All Models (28)
S+ Tier
Coding Specialists
Reasoning
Long Context (โ‰ฅ256K)
Local-Friendly

GLM 5.2

Zhipu AI ยท via NVIDIA NIM free tier
S+ Tier Coding Specialist Free (NIM)
๐Ÿ“ Context: 128,000 tokens
SWE-bench Verified: 82.8% โ€” highest among all free coding models
OpenCode โœ“ (NIM) NVIDIA NIM โœ“ OpenRouter โœ— Local โœ—
Strengths: #1 free model on SWE-bench. GLM family's latest coding-optimized release. Strong code generation and debugging.
Weaknesses: 128K context only. Hosted-only (not open-weight). NIM free tier rate limits.
๐ŸŽฏ Best for: Primary implementer when SWE-bench performance matters most. Complex bug fixes.

DeepSeek V4 Pro

DeepSeek ยท via NVIDIA NIM free tier / OpenAdapter
S+ Tier MoE Free (NIM)
๐Ÿ“ Context: 1,000,000 tokens (NIM) / 256K (OpenAdapter)
SWE-bench Verified: 80.6%
OpenCode โœ“ NVIDIA NIM โœ“ OpenRouter โœ— Local โœ—
Strengths: Frontier coding capability. 1M context via NIM. Strong agent workflows. Already configured in your OpenAdapter.
Weaknesses: OpenAdapter version has 256K context (not 1M). Hosted-only.
๐ŸŽฏ Best for: Heavy agent workflows. Complex multi-file tasks. When you need both S+ coding and long context.

DeepSeek V4 Flash

DeepSeek ยท via NVIDIA NIM free tier / OpenAdapter
S+ Tier MoE ยท 284B-13B Free (NIM)
๐Ÿ“ Context: 1,000,000 tokens (NIM) / 256K (OpenAdapter)
SWE-bench Verified: 79.0%
OpenCode โœ“ NVIDIA NIM โœ“ OpenRouter โœ— Local โœ—
Strengths: Near-Pro quality at Flash speed. 13B active params = fast inference. Frontier coding ability.
Weaknesses: Slightly below Pro on complex reasoning. Hosted-only.
๐ŸŽฏ Best for: Fast, high-quality code implementation. When you need S+ quality with lower latency.

Kimi K2.6

Moonshot AI ยท via NVIDIA NIM free tier
S+ Tier Free (NIM)
๐Ÿ“ Context: 262,144 tokens
SWE-bench Verified: 80.2%
OpenCode โœ“ (NIM) NVIDIA NIM โœ“ Kilo โœ“ Local โœ—
Strengths: Frontier coding. 262K context. Strong on long-form code generation. Available in Kilo.
Weaknesses: Less known ecosystem. NIM rate limits.
๐ŸŽฏ Best for: Long-form code generation. When you need an alternative S+ model for diversity.

Step 3.7 Flash

StepFun ยท via NVIDIA NIM free tier
S+ Tier Free (NIM)
๐Ÿ“ Context: 256,000 tokens
SWE-bench Verified: 74.4%
OpenCode โœ“ (NIM) NVIDIA NIM โœ“ OpenRouter โœ— Local โœ—
Strengths: S+ coding at Flash speed. 256K context. Efficient inference.
Weaknesses: Less known than DeepSeek/GLM. Limited ecosystem support.
๐ŸŽฏ Best for: Fast S+ tier coding. Alternative when other S+ models are rate-limited.

MiniMax M3

MiniMax ยท via NVIDIA NIM free tier / OpenAdapter
S+ Tier MoE Free (NIM)
๐Ÿ“ Context: 1,000,000 tokens (NIM) / 262K (OpenAdapter)
SWE-bench Verified: 78.4%
OpenCode โœ“ NVIDIA NIM โœ“ OpenRouter โœ— Local โœ—
Strengths: S+ coding with 1M context via NIM. Strong all-rounder. Already in OpenAdapter config.
Weaknesses: Less coding-specialized than GLM/DeepSeek. Chinese-market focus means English docs are sparse.
๐ŸŽฏ Best for: Long-context S+ coding. When you need both frontier quality and large context.

Mistral Medium 3.5 128B

Mistral ยท via NVIDIA NIM free tier
S+ Tier Dense ยท 128B Free (NIM)
๐Ÿ“ Context: 256,000 tokens
SWE-bench Verified: 77.6%
OpenCode โœ“ (NIM) NVIDIA NIM โœ“ OpenRouter โœ— Local โœ—
Strengths: Mistral's strongest free model. Dense = consistent quality. Strong European language support.
Weaknesses: 128B dense = slower inference than MoE peers. Not coding-specialized.
๐ŸŽฏ Best for: When you want Mistral ecosystem. European language coding. Consistent dense-model output.

Laguna S 2.1

Poolside ยท via OpenRouter
Coding Specialist MoE ยท 118B total Free
๐Ÿ“ Context: 262,144 tokens
SWE-bench Verified: not yet on SWE-bench leaderboard (new model)
OpenCode โœ“ OpenRouter โœ“ Kilo โš  Local โœ—
Strengths: Purpose-built coding agent. Excellent diff application and multi-file edits. Strong agentic coding workflow.
Weaknesses: Not open-weight. Weaker on general reasoning. Rate-limited on OpenRouter free tier.
๐ŸŽฏ Best for: Primary implementer. Diff/patch application. Multi-file refactoring.

Laguna XS 2.1

Poolside ยท via OpenRouter
Coding Specialist MoE ยท 33B-A3B Free
๐Ÿ“ Context: 262,144 tokens
OpenCode โœ“ OpenRouter โœ“ Kilo โš  Local โœ—
Strengths: Fast, lightweight coding specialist. Low latency. Good for quick edits.
Weaknesses: Less capable than S variant. Hosted only.
๐ŸŽฏ Best for: Fast code completions. Simple edits. Background worker tasks.

Laguna M 1

Poolside ยท via OpenRouter (OpenAdapter)
Coding Specialist MoE Free
๐Ÿ“ Context: 262,144 tokens
OpenCode โœ“ OpenRouter โœ“ Koda โœ“ Local โœ—
Strengths: Mid-size Laguna variant. Good balance of speed and capability.
Weaknesses: Older generation than 2.1 series. Less capable than Laguna S 2.1.
๐ŸŽฏ Best for: Mid-complexity coding. When Laguna S is rate-limited.

North Mini Code

Cohere ยท via OpenRouter
Coding Specialist MoE ยท Sparse Free
๐Ÿ“ Context: 256,000 tokens
OpenCode โœ“ OpenRouter โœ“ Kilo โš  Local โœ—
Strengths: Agentic-first design. Built for tool-calling. Strong codebase context understanding.
Weaknesses: Newer model, less battle-tested. Cohere's first coding model.
๐ŸŽฏ Best for: Agent orchestrator. Tool-use-heavy workflows. Codebase exploration.

Nemotron 3 Ultra

NVIDIA ยท via OpenRouter / NIM
S+ Tier Reasoning MoE ยท 550B-A55B Free
๐Ÿ“ Context: 1,000,000 tokens
SWE-bench Verified: 71.9%
OpenCode โœ“ (default) OpenRouter โœ“ NVIDIA NIM โœ“ Local โœ—
Strengths: 1M context. 71.9% SWE-bench. Reasoning+orchestration design. Best free model for full-repo analysis. Already your OpenCode default.
Weaknesses: Not coding-specialized (outperformed by GLM/DeepSeek on pure coding). Can be verbose.
๐ŸŽฏ Best for: Architecture review. Full-repo analysis. Planning/orchestration. Long-context diff review.

Nemotron 3 Super

NVIDIA ยท via OpenRouter / NIM
S Tier MoE ยท 120B-A12B Free
๐Ÿ“ Context: 1,000,000 tokens (OpenAdapter) / 262K (OpenRouter)
SWE-bench Verified: 60.5%
OpenCode โœ“ OpenRouter โœ“ NVIDIA NIM โœ“ Local โœ—
Strengths: Efficient MoE. 1M context via OpenAdapter config. Good balance of capability and speed.
Weaknesses: 60.5% SWE-bench โ€” below S+ models. Not coding-specialized.
๐ŸŽฏ Best for: General agent tasks. Fallback when Ultra is rate-limited. Fast 1M-context tasks.

Nemotron 3 Nano 30B A3B

NVIDIA ยท via OpenRouter / NIM
MoE ยท 30B-A3B Free
๐Ÿ“ Context: 1,000,000 tokens (via NIM)
SWE-bench Verified: 38.8%
OpenCode โœ“ OpenRouter โœ“ Local โœ—
Strengths: Most compute-efficient Nemotron. 1M context via NIM. Fast inference.
Weaknesses: 38.8% SWE-bench โ€” limited for serious coding.
๐ŸŽฏ Best for: Lightweight agent tasks. High-throughput parallel workers. Summarization.

Nemotron 3 Nano Omni (Reasoning)

NVIDIA ยท via OpenRouter / NIM
Reasoning MoE ยท 30B-A3B Free
๐Ÿ“ Context: 256,000 tokens
SWE-bench Verified: 52.0%
OpenCode โœ“ OpenRouter โœ“ Local โœ—
Strengths: Reasoning-tuned. Multimodal. 52% SWE-bench at 30B is impressive efficiency.
Weaknesses: Small active params limit complex reasoning depth.
๐ŸŽฏ Best for: Multimodal code review. UI implementation review. Perception-subagent.

GPT-OSS 120B

OpenAI ยท via Cerebras / OpenRouter / NIM
S Tier Dense ยท 120B Free
๐Ÿ“ Context: 131,072 tokens
SWE-bench Verified: 62.4%
OpenCode โœ“ Cerebras โœ“ (fast) OpenRouter โœ“ NVIDIA NIM โœ“ Local โœ—
Strengths: Largest open-weight dense model. 62.4% SWE-bench โ€” strong showing. Apache 2.0. Cerebras inference is blazing fast.
Weaknesses: 131K context only. Not coding-specialized.
๐ŸŽฏ Best for: Fast capable general tasks. When you want OpenAI-trained weights free.

GPT-OSS 20B

OpenAI ยท via OpenRouter / NIM
Dense ยท 21B Free
๐Ÿ“ Context: 131,072 tokens
SWE-bench Verified: 50.3%
OpenCode โœ“ OpenRouter โœ“ NVIDIA NIM โœ“ Local โš 
Strengths: Apache 2.0. 50.3% SWE-bench at 20B is efficient. Runnable locally on good hardware.
Weaknesses: 131K context. Outperformed by MoE peers at similar size.
๐ŸŽฏ Best for: Local deployment. License-permissive projects. Simple coding.

Gemma 4 31B

Google ยท via OpenRouter / NIM
Dense ยท 30.7B Free
๐Ÿ“ Context: 262,144 tokens
SWE-bench Verified: 52.0%
OpenCode โœ“ OpenRouter โœ“ NVIDIA NIM โœ“ Local โš 
Strengths: Dense = consistent output. 52% SWE-bench. Multimodal. Open-weight.
Weaknesses: Not coding-specialized. 31B dense is heavy for local.
๐ŸŽฏ Best for: General agent tasks. Multimodal review. Local fallback.

Hermes 3 Llama 3.1 405B

Nous Research / Meta ยท via OpenRouter
Dense ยท 405B Free
๐Ÿ“ Context: 131,072 tokens
OpenCode โœ“ OpenRouter โœ“ Local โœ—
Strengths: Largest dense free model. Strong general reasoning. Excellent for analysis and review.
Weaknesses: 131K context. Very slow (405B dense). Not coding-specialized.
๐ŸŽฏ Best for: Deep analysis. Architecture review. Second opinion on complex design.

Llama 3.3 70B

Meta ยท via Groq / OpenRouter / Cloudflare / NIM
Dense ยท 70B Free
๐Ÿ“ Context: 131,072 tokens
SWE-bench Verified: 22.0%
OpenCode โœ“ Groq โœ“ (fastest) Cloudflare โœ“ NVIDIA NIM โœ“ OpenRouter โœ“
Strengths: Most widely available free model. Groq = extremely fast. Reliable. Well-tested. Available everywhere.
Weaknesses: 22% SWE-bench โ€” far behind S+ models. 131K context. Older architecture.
๐ŸŽฏ Best for: Reliable general tasks. Fast Groq-powered agent work. Consistent, predictable output.

Qwen3 Next 80B A3B

Alibaba ยท via OpenRouter
MoE ยท 80B-A3B Free
๐Ÿ“ Context: 262,144 tokens
OpenCode โœ“ OpenRouter โœ“ Local โœ—
Strengths: Strong Qwen coding lineage. 262K context. Efficient MoE.
Weaknesses: Not explicitly coding-specialized. Qwen-Coder not free on OpenRouter.
๐ŸŽฏ Best for: General coding with Qwen-familiar outputs. Alternative to Laguna for non-specialist work.

DeepSeek R1 (NIM free)

DeepSeek ยท via NVIDIA NIM free tier
Reasoning Dense Free (NIM)
๐Ÿ“ Context: 131,072 tokens
OpenCode โœ“ NVIDIA NIM โœ“ OpenRouter โœ— Local โš  (7B/32B)
Strengths: Reasoning-first. Chain-of-thought by default. Strong on debugging and algorithms.
Weaknesses: 131K context. Reasoning overhead adds latency. Not for straightforward edits.
๐ŸŽฏ Best for: Complex debugging. Algorithm design. Step-by-step reasoning tasks.

Nex N2 Pro

Unknown ยท via OpenRouter (koda discovery)
Free MoE
๐Ÿ“ Context: unknown (listed in koda agent config)
Koda โœ“ OpenRouter โœ“ OpenCode โš 
Strengths: Newly available free model. Multimodal (text + image).
Weaknesses: Limited documentation. Unknown SWE-bench score. Unproven for coding.
๐ŸŽฏ Best for: Experimental use. Monitoring for capability. Backup option.

Owl Alpha

Unknown ยท via OpenRouter (koda discovery)
Free
๐Ÿ“ Context: unknown (listed in koda agent config)
Koda โœ“ OpenRouter โœ“ OpenCode โš 
Strengths: Newly available free model. Worth monitoring.
Weaknesses: Almost no public documentation. Unknown capabilities. Experimental.
๐ŸŽฏ Best for: Experimental testing only. Not recommended for production agent work yet.

Qwen2.5-Coder 32B (Local)

Alibaba ยท Local (Ollama)
Coding Specialist Dense ยท 32B Local
๐Ÿ“ Context: 32,768 tokens (Ollama default)
Local โœ“ (Ollama) OpenRouter โœ— OpenCode โœ—
Strengths: Best local coding model at 32B. Privacy-preserving. No rate limits. In your model-radar.
Weaknesses: 32K context โ€” severely limiting. Requires capable hardware. Not available hosted-free.
๐ŸŽฏ Best for: Privacy-sensitive code. Offline coding. Local agent worker.

DeepSeek R1 32B (Local)

DeepSeek ยท Local (Ollama)
Reasoning Dense ยท 32B Local
๐Ÿ“ Context: 32,768 tokens (Ollama default)
Local โœ“ (Ollama) OpenRouter โœ— OpenCode โœ—
Strengths: Reasoning on local hardware. Good for complex offline debugging. Privacy-preserving.
Weaknesses: 32K context. Reasoning overhead = slow on consumer hardware.
๐ŸŽฏ Best for: Offline debugging. Complex reasoning where privacy matters.

Qwen3 14B (Fedora)

Alibaba ยท Local (Ollama ยท Fedora)
Dense ยท 14B Local
๐Ÿ“ Context: 32,768 tokens (Ollama default)
OpenCode โœ“ (Fedora) Local โœ“ OpenRouter โœ—
Strengths: Runs on modest hardware. Already configured in Fedora provider. Good general coding.
Weaknesses: 32K context. 14B limited for complex tasks. Not coding-specialized.
๐ŸŽฏ Best for: Quick local tasks. Already deployed and ready. Privacy-sensitive work.

Gemini 3.6 Flash

Google ยท Gemini API free tier
Dense Free
๐Ÿ“ Context: 1,048,576 tokens
OpenCode โœ“ Gemini API โœ“ Gemini CLI โœ“ Local โœ—
Strengths: 1M context. Very fast. Latest Gemini generation. Strong multimodal. Generous free tier limits.
Weaknesses: Google free tier rate limits. Not coding-specialized. Privacy considerations.
๐ŸŽฏ Best for: Long-context codebase analysis. Fast general tasks. Multimodal review.

Gemini 3.1 Pro Preview

Google ยท Gemini API free tier
Dense Free
๐Ÿ“ Context: 1,048,576 tokens
OpenCode โœ“ Gemini API โœ“ Gemini CLI โœ“ Local โœ—
Strengths: Pro-tier capability at free tier. 1M context. Strong reasoning.
Weaknesses: Preview status = may change. Pro tier has stricter free limits (50 req/day).
๐ŸŽฏ Best for: When you need Google's best reasoning on the free tier. Architecture review.

3. Installed CLI Free Tier Map

You have 10 CLIs installed that provide access to free AI models. This section maps each CLI to the free models or providers it unlocks. Some CLIs (like OpenCode) are full agent environments; others (like free-coding-models) are discovery/routing tools.

๐Ÿ”ง OpenCode

~/.opencode/bin/opencode
8 free-tier providers configured:
  • OmniRoute (local router) โ€” auto/best-coding, auto/coding:free, etc.
  • OpenAdapter (OpenRouter proxy) โ€” 21 free/* models + 12 0G-* models
  • Groq (free tier) โ€” Llama 3.3 70B, Mixtral 8x7B
  • NVIDIA NIM (free) โ€” Llama 3.3 70B, DeepSeek R1
  • Cerebras (free) โ€” GPT-OSS 120B
  • Mistral (free tier) โ€” Mistral Small, Mistral Large
  • Google Gemini (free) โ€” 11 models (2.5 Flash through 3.6 Flash)
  • Cloudflare Workers AI (free) โ€” Llama 3.3 70B, Qwen2 72B, Hermes 2 Pro
  • Fedora Ollama (local) โ€” Qwen3 8B, Qwen3 14B
Total: 60+ models across 9 providers

๐Ÿ”ง Kilo Code

/opt/homebrew/bin/kilo
Model gateways configured:
  • Cloudflare AI Gateway โ€” routes to Anthropic (Claude 3 through Opus 5), OpenAI (GPT-3.5 through GPT-5.6), Moonshot (Kimi K3), Workers AI
  • OpenCode providers โ€” shares OpenCode's provider config (OpenRouter, NIM, etc.)
  • Workers AI โ€” Gemma SEA LION, BGE base, IndicTrans, etc.
Note: Most Cloudflare AI Gateway models are paid (routes to paid APIs). Workers AI models are free-tier.

๐Ÿง  Grok CLI

~/.local/bin/grok
Model access:
  • Grok's own model cache at ~/.grok/models_cache.json
  • Model discovery via its own mechanism
  • Currently has 1 cached model entry
Note: Primarily a xAI/Grok client โ€” free tier availability unclear without active session.

๐Ÿ”ง Koda

~/.config/koda/
OpenAdapter provider (26 free models):
  • NVIDIA: Nemotron Nano, Super, Ultra, Omni, Nano 9B/12B VL
  • Poolside: Laguna M 1, Laguna XS 2
  • Google: Gemma 4 26B, Gemma 4 31B
  • OpenAI: GPT-OSS 20B, GPT-OSS 120B
  • Meta: Llama 3.2 3B, Llama 3.3 70B
  • DeepSeek: V3, GLM-5, Qwen3.6, Qwen-VL
  • Cohere: North Mini Code (not in list, but OpenRouter has it)
  • New: free/nex-n2-pro, free/owl-alpha
  • Liquid: LFM 2.5 1.2B (instruct + thinking)
  • Dolphin: Mistral 24B Venice

๐Ÿค– Gemini CLI

@google/gemini-cli@0.53.1
Free-tier access:
  • Google Gemini models via API key
  • gemini gemma โ€” local Gemma model routing
  • MCP server support for tool extensions
  • Agent skills and hooks system
Note: Full-featured agent CLI with sandbox support. Free tier limits apply per Google's Gemini API terms.

๐Ÿค– Mimo CLI

@mimo-ai/cli@0.1.4
Model access:
  • Mimo AI's own models via their API
  • Free tier availability unclear โ€” requires account
Note: Newer tool (v0.1.4). Worth monitoring for free tier offers.

๐Ÿค– CodeBuddy

@tencent-ai/codebuddy-code@2.132.0
Model access:
  • Tencent's AI coding assistant
  • Free tier likely available (Tencent's model)
  • v2.132.0 suggests mature product
Note: Chinese-market focus. May have free-tier coding models worth investigating.

๐Ÿ” free-coding-models

free-coding-models@0.5.69 (global npm)
Comprehensive free model discovery:
  • Catalogs 221 models from NVIDIA NIM, Groq, Cerebras, SambaNova, OpenRouter, GitHub Models, Mistral, Codestral, Scaleway, Google AI, Z.AI, Qwen, Cloudflare, OVHcloud, OpenCode Zen, Kilo, llm7, Routeway, Novita, Ollama Cloud
  • Tiered from S+ (80%+ SWE-bench) to C (<20%)
  • Live ping/probe capability for uptime monitoring
  • Supports 30+ coding CLI targets (--opencode, --kilo, --aider, --cline, etc.)
  • Web dashboard + TUI + JSON output modes
  • FCM Router daemon for model routing
This is your most comprehensive discovery tool.

๐Ÿฆ™ Ollama

/usr/local/bin/ollama
Local model server:
  • Running on a local machine via Ollama
  • Currently pulled: Qwen3 8B, Qwen3 14B, nomic-embed-text
  • Available from model-radar: Qwen2.5-Coder 7B/14B/32B, DeepSeek R1 7B/32B, Llama 3.1 8B/70B, Mistral 7B, Mixtral 8x7B
Note: Privacy-preserving. No rate limits. Limited by hardware and 32K context default.

๐Ÿง  Claude CLI

~/.local/bin/claude
Access:
  • Anthropic Claude models (Opus 5, Sonnet 5, etc.)
  • Via OpenClaude (@gitlawb/openclaude@0.24.0) with OpenRouter model discovery
  • Model discovery cache includes OpenRouter entries
Note: Primarily for paid Claude access. Not a free-tier source itself, but OpenClaude can route through OpenRouter free models.

4. Frontier Models Accessible via Free Tiers

One of the most surprising findings of this audit: frontier S+ tier coding models are available for free through NVIDIA NIM's free tier. These are not stripped-down or lite versions โ€” they are the same models that achieve 70-83% on SWE-bench Verified.
ModelSWE-benchContextFree ViaAlso Available Paid On
GLM 5.2 82.8% 128K NVIDIA NIM free Z.AI direct, OpenRouter (paid)
DeepSeek V4 Pro 80.6% 1M NVIDIA NIM free DeepSeek direct, OpenRouter
Kimi K2.6 80.2% 262K NVIDIA NIM free Moonshot direct, Kilo (paid gateway)
DeepSeek V4 Flash 79.0% 1M NVIDIA NIM free DeepSeek direct, OpenRouter
MiniMax M3 78.4% 1M NVIDIA NIM free MiniMax direct
Mistral Medium 3.5 77.6% 256K NVIDIA NIM free Mistral API (paid),OpenRouter
Step 3.7 Flash 74.4% 256K NVIDIA NIM free StepFun direct
Nemotron 3 Ultra 71.9% 1M OpenRouter free, NIM free NVIDIA NIM (paid tier)
How this works: NVIDIA NIM offers a free tier with rate-limited access to their full model catalog โ€” including third-party models they host (DeepSeek, GLM, Kimi, MiniMax, Mistral, StepFun). This is a promotional play to drive NIM platform adoption. The free tier has request/min and request/day caps but no token-based billing.
Practical impact: You can run GLM 5.2 (82.8% SWE-bench โ€” higher than many paid coding assistants) for free through OpenCode's NVIDIA NIM provider. The rate limits mean you'll use it for the hardest tasks and fall back to other free models for routine work. This is the closest thing to "frontier coding for free" available in August 2026.

5. Investigation: Meta Muse Spark 1.2 โ€” Free Hosting?

Question: Is Meta Muse Spark 1.2 available for free anywhere?
Finding: NO. Meta Muse Spark 1.2 is not available for free on any platform we checked.
PlatformStatusPricing (per 1M tokens)
OpenRouterPaid only$1.25 input / $4.25 output
NVIDIA NIMNot listedNot in NIM catalog (free or paid)
Meta directNot hostedOpen-weight release only (self-host)
GroqNot listedNot in Groq catalog
CerebrasNot listedNot in Cerebras catalog
Cloudflare Workers AINot listedNot in Workers AI catalog
Google GeminiN/ACompeting model, not hosted
What Muse Spark 1.2 is: A reasoning model from Meta with 1M-token context, multimodal input (text, images, video, audio, PDF), designed for complex agentic tasks. It's available as open-weight (self-host) and as a paid API on OpenRouter.
Assessment: Muse Spark is the most notable absence from the free model ecosystem. As Meta's latest reasoning model with 1M context, it would be a strong candidate for the planner/reviewer role. However, its absence from NVIDIA NIM's free tier (unlike DeepSeek, GLM, and Kimi which are all there) suggests Meta hasn't struck a promotional distribution deal. Watch for it on Cloudflare Workers AI or Groq โ€” both have hosted Meta models for free before.
Bottom line: Muse Spark 1.2 is a paid model everywhere it's hosted. Self-hosting is possible (it's open-weight) but requires massive hardware given the model size. Not currently a viable free option.

6. Model Comparison โ€” Workflow Scoring with SWE-bench Scores

Scores are qualitative assessments based on model architecture, SWE-bench Verified scores (where available), documented capabilities, and community reports. They reflect suitability for coding agent workflows โ€” not general benchmark rankings.
Model SWE-bench A. Architecture
Review
B. Repo
Auditing
C. Refactoring D. UI Redesign
Impl
E. Terminal/
Tooling
F. Long-Context
Analysis
Overall
GLM 5.2 82.8% โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ—‹ 23/24
DeepSeek V4 Pro 80.6% โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ— 24/24
Kimi K2.6 80.2% โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ—‹ โ—โ—โ—โ— โ—โ—โ—โ—‹ 22/24
DeepSeek V4 Flash 79.0% โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ— 22/24
MiniMax M3 78.4% โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ— โ—โ—โ—โ—‹ โ—โ—โ—โ— โ—โ—โ—โ— 21/24
Mistral Medium 3.5 77.6% โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ 20/24
Step 3.7 Flash 74.4% โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ— โ—โ—โ—โ—‹ โ—โ—โ—โ— โ—โ—โ—โ—‹ 20/24
Nemotron 3 Ultra 71.9% โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ—‹ โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ— 19/24
Laguna S 2.1 N/A โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ— โ—โ—โ—โ—‹ 21/24
GPT-OSS 120B 62.4% โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—‹โ—‹ 16/24
Nemotron 3 Super 60.5% โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ— 17/24
Gemma 4 31B 52.0% โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ 18/24
Nemotron Nano Omni 52.0% โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ 18/24
GPT-OSS 20B 50.3% โ—โ—โ—‹โ—‹ โ—โ—โ—‹โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—‹โ—‹ 15/24
Llama 3.3 70B 22.0% โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—โ—‹ โ—โ—โ—‹โ—‹ 16/24
Scoring rubric: โ—โ—โ—โ— = Excellent fit for this workflow. โ—โ—โ—โ—‹ = Good, with minor gaps. โ—โ—โ—‹โ—‹ = Usable but limited. โ—โ—‹โ—‹โ—‹ = Not recommended. Scores are relative to free models โ€” not frontier paid models like Claude Opus 5 or GPT-5.6.

7. Agent Role Recommendations โ€” "Free Agent Stack"

Updated with the newly discovered S+ tier models. The free agent stack is now significantly more capable than when we only considered OpenRouter's free models.

๐Ÿง  Planner / Orchestrator

Architecture review ยท task decomposition ยท planning
DeepSeek V4 Pro (NIM)
Nemotron 3 Ultra (1M ctx), GLM 5.2 (82.8% SWE)
Why: 80.6% SWE-bench + 1M context via NIM. Better coding scores than Nemotron Ultra while matching its context window. The best free model for understanding an entire codebase and planning changes. Use Nemotron Ultra when you specifically need reasoning+orchestration design.

๐Ÿ”ง Implementer

Code generation ยท diff application ยท multi-file edits
GLM 5.2 (NIM)
Laguna S 2.1, DeepSeek V4 Flash (fast lane)
Why: 82.8% SWE-bench โ€” the highest-scoring free coding model. Purpose-built for code generation. For agentic coding workflows with tool calling, Laguna S 2.1 may still have an edge (built as a coding agent, not just a coding model). Use DeepSeek V4 Flash when you need speed.

๐Ÿ‘๏ธ Reviewer

Code review ยท diff analysis ยท quality gate
Nemotron 3 Ultra
DeepSeek V4 Pro (NIM), Gemini 3.1 Pro (multimodal)
Why: 1M context still wins for reviewing large PRs as a whole. The reasoning+orchestration design catches logical errors coding-specialized models might miss. DeepSeek V4 Pro is a strong alternative when you want higher SWE-bench accuracy in review.

๐Ÿ  Local Worker

Privacy-sensitive tasks ยท offline work
Qwen2.5-Coder 32B
DeepSeek R1 32B (reasoning), Qwen3 14B (lightweight ยท already deployed)
Why: Best local coding model at a runnable size. Privacy-preserving. No rate limits. Already in your model-radar. Runs on Fedora or macOS with sufficient RAM. The 32K context is the main limitation โ€” use hosted models for anything requiring broader context.

Complete Free Agent Stack (Updated)

RolePrimary ModelProviderFallbackWhen to use fallback
PlannerDeepSeek V4 ProNVIDIA NIMNemotron 3 UltraWhen NIM is rate-limited
Implementer (best)GLM 5.2NVIDIA NIMDeepSeek V4 FlashWhen speed > absolute quality
Implementer (agentic)Laguna S 2.1OpenRouterNorth Mini CodeWhen tool-use is primary need
Fast laneDeepSeek V4 FlashNVIDIA NIMLlama 3.3 70B (Groq)When NIM is rate-limited
ReviewerNemotron 3 UltraOpenRouterGemini 3.1 ProWhen multimodal review needed
Long-context analystNemotron 3 UltraOpenRouterGemini 3.6 FlashWhen 1M Gemini context preferred
Local workerQwen2.5-Coder 32BOllama (Fedora)DeepSeek R1 32BWhen reasoning > code gen
Generalist/routerGemini 3.6 FlashGoogle APIGPT-OSS 120B (Cerebras)When Google API is down
Architecture (keep paid)ClaudeAnthropicโ€”Don't replace this

8. Local vs Hosted Comparison

FactorHosted Free (OpenRouter/NIM/APIs)Local (Ollama on Fedora)
Best models availableGLM 5.2 (82.8% SWE), DeepSeek V4 Pro (80.6%), Kimi K2.6 (80.2%), Nemotron Ultra (71.9%)Qwen2.5-Coder 32B, DeepSeek R1 32B, Qwen3 14B
Max practical context1M tokens (Nemotron Ultra, DeepSeek V4 via NIM, Gemini Flash)~32K tokens (Ollama default; configurable but VRAM-limited)
SWE-bench ceiling82.8% (GLM 5.2)~40-50% (estimated for Qwen2.5-Coder 32B)
SpeedFast (API infra; Groq/Cerebras especially fast)Hardware-dependent; 32B ~5-15 tok/s on consumer GPU
PrivacyCode sent to third-party serversCode stays on your machine
Rate limitsYes โ€” NIM: ~30 req/min, ~500/day; OpenRouter free: ~20 req/minNone โ€” you own the hardware
Availability riskModels can go paid or be removed without noticeYou control the model; once downloaded, it's yours
Cost$0$0 + electricity + hardware depreciation
Updated verdict: The gap between hosted free and local has widened dramatically. NVIDIA NIM's free tier now hosts S+ models at 80%+ SWE-bench โ€” quality that would cost $3-15/M tokens on paid APIs. Local models at 32B simply cannot compete on capability. For your workflow, local models are now exclusively for privacy-sensitive tasks and offline fallback. Everything else should use hosted free models, layered behind OmniRoute for intelligent routing.

9. Risks & Caveats

โš ๏ธ Availability Changes (HIGH RISK)

NVIDIA NIM's free tier hosting GLM 5.2, DeepSeek V4 Pro, Kimi K2.6, and other S+ models is almost certainly a promotional strategy. NVIDIA can remove or restrict these at any time. OpenRouter free models have historically been converted to paid (several DeepSeek variants moved from free to paid in 2025-2026). Poolside's Laguna free tier is a customer acquisition play for their enterprise product.
The S+ free tier is likely temporary. NVIDIA isn't running a charity โ€” they're building NIM platform adoption. Once enterprise contracts are established, expect the free tier to narrow. The models most at risk: GLM 5.2 (highest SWE-bench free model) and DeepSeek V4 Pro (most capable all-rounder).

โš ๏ธ Rate Limits (MEDIUM-HIGH RISK)

NVIDIA NIM free: approximately 30 requests/minute, 500 requests/day (estimated โ€” not publicly documented). OpenRouter free: ~20 requests/minute, ~200 requests/day. Google Gemini free: 1,500 requests/day (Flash), 50 requests/day (Pro). Groq free: rate-limited by tokens/minute. Agent workflows making 30-50 sequential calls per complex task will hit daily limits quickly.
Strategy: Layer multiple free providers through OmniRoute. When NIM is rate-limited, fall back to OpenRouter. When OpenRouter is rate-limited, fall back to Gemini/Groq. Have at least 3 providers per role to avoid workflow interruption.

โš ๏ธ Privacy Considerations

All hosted free models process your code on third-party servers. NVIDIA NIM's terms of service for the free tier are less clearly documented than their enterprise terms. OpenRouter's privacy policy permits data use. Google's free tier allows data collection. Cloudflare Workers AI has a more enterprise-friendly data policy.
For proprietary or client code, use local models (Ollama on Fedora) or paid API tiers with data processing agreements. Free tiers are appropriate for open-source work, personal projects, and learning.

โš ๏ธ Provider Concentration Risk

NVIDIA NIM hosts 7 of the top 8 free models by SWE-bench score. If NIM changes its free tier policy, the free model landscape instantly degrades from S+ tier to A tier (Nemotron Ultra on OpenRouter at 71.9% becomes the best remaining option).
Mitigation: Don't build workflows that depend exclusively on NIM. Keep OpenRouter and Google Gemini free tiers as regularly exercised alternatives. This ensures a smooth transition if NIM's free tier changes.

โš ๏ธ Model Quality Variance

SWE-bench Verified scores are self-reported by model providers. Independent benchmarks (e.g., SWE-rebench) typically show lower scores. Free-tier models may use quantized versions (FP8/INT8) that reduce quality vs full-precision. NIM-hosted models may not be the exact same build as the provider's own API.
Test each model on your actual workflow before committing. A model with 82.8% SWE-bench might underperform a 71.9% model on your specific codebase and task patterns.

10. Final Recommendations

For Your Workflow

Your current setup โ€” Claude for architecture/review, OmniRoute for model routing, OpenCode for agent orchestration, Fedora for local AI โ€” is excellent. The free model ecosystem now provides S+ tier coding models that can handle implementation tasks at near-frontier quality.
Recommended configuration updates for OpenCode:
TaskCurrentRecommendedRationale
Default agent Nemotron 3 Ultra DeepSeek V4 Pro (NIM) 80.6% SWE-bench vs 71.9%. 1M context. Better coding. Keep Nemotron as fallback.
Code implementation Laguna S 2.1 GLM 5.2 (NIM) 82.8% SWE-bench. Best free coding model. Use Laguna for agentic tool-use tasks.
Fast lane Llama 3.3 70B (Groq) DeepSeek V4 Flash (NIM) 79% SWE-bench vs 22%. Comparable speed. Dramatically better coding.
Long-context analysis Nemotron 3 Ultra Keep Nemotron Ultra Still the best free reasoning+orchestration with 1M context.
Privacy/offline Qwen3 14B (Fedora) Qwen2.5-Coder 32B Pull this model to Fedora. 32B coding specialist. Worth the upgrade.
Architecture review Claude Keep Claude Don't replace this. Free models are not competitive with Claude for architectural thinking.

What to Do Now

  1. Add NVIDIA NIM S+ models to OmniRoute. Route coding tasks to GLM 5.2, planning to DeepSeek V4 Pro, and keep Nemotron Ultra for long-context reasoning.
  2. Pull Qwen2.5-Coder 32B to Fedora. It's already in your model-radar. 32B coding specialist will dramatically improve local agent work vs Qwen3 14B.
  3. Use free-coding-models for ongoing discovery. Run free-coding-models web to access the dashboard. Set up the FCM Router daemon if you want automatic model health monitoring.
  4. Monitor Meta Muse Spark availability. If it appears on Cloudflare Workers AI or Groq free tier, it becomes an instant candidate for the reviewer role.

What NOT to Do

What to Watch

11. Sources & Methodology

Methodology note: "Available" means the model appears in a live API query, is configured in one of your installed CLI configs, or is documented in free-coding-models sources.js as available via a free-tier provider. "Kilo โš " means the model uses an OpenAI-compatible API and should work with Kilo Code, but was not explicitly tested against Kilo's tool-calling format. SWE-bench scores are from the free-coding-models catalog (sources.js) and may differ from independently verified scores.