The Complete Guide to Free LLM APIs
and Low-Cost Inference in 2026
Token costs are climbing. Your Hermes Agent, OpenClaw, or Open Code project should not stall because of a billing alert. Here is every provider offering free API access right now, what the real limits are, and the cheapest unlimited-style subscriptions when you need to go further.
Quick Glossary: RPM Means Requests Per Minute, RPD Means Requests Per Day
Terms you will see throughout this guide
The number of API calls you can send in any 60-second window before the provider throttles or rejects the next one with a 429 error.
The total number of API calls allowed in a rolling or calendar day, regardless of how you space them out.
Tokens Per Minute. A separate ceiling some providers use instead of, or alongside, RPM, since one request can contain very different amounts of text.
The HTTP status code returned when you exceed a rate limit. It means slow down, not you are banned.
Quick Reference: All Free API Tiers
No credit card, no trial period limits, verified August 20, 2026
| Provider | Top Free Models | Daily / Monthly Limit | Rate Limit | Card Required | Specialty Modalities | How Long Free |
|---|---|---|---|---|---|---|
| Google AI Studio | Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, Gemini 3.1 Flash-Lite | Flash-Lite: up to 1,500 RPD; Flash: ~250-500 RPD; Pro models paid-only since Apr 1, 2026 | Flash-Lite ~15 RPM; Flash ~10-15 RPM, 250K TPM | No | Vision, Audio, Video, PDF input | Ongoing; Google has cut quotas before without notice |
| NVIDIA NIM | GLM-5.2, Nemotron 3 Ultra 550B, Nemotron 3.5 Lightning | No credit or token cap; gated only by rate limit | 40 RPM shared cap, varies with model load | No | Vision, embeddings, safety/guard models | Permanently free; NVIDIA removed its old credit ceiling |
| Groq | Llama 3.1 8B Instant, Llama 4 Scout, GPT-OSS 120B | 14,400 RPD (8B); 1,000 RPD (70B/Scout/GPT-OSS); 250 RPD (Compound) | 30-60 RPM; 6,000-30,000 TPM depending on model | No | Whisper Large v3 & Turbo STT (2,000 audio req/day) | Free Forever tier |
| Cerebras | gpt-oss-120b, Gemma 4 31B | $5 in credits, expires 30 days after granted | 5 RPM, 30K TPM, 1M TPD per model | Yes | Fastest raw inference speed (2,000+ tok/sec) | No longer a true free tier as of Aug 17, 2026; card required to unlock a 30-day credit trial |
| OpenRouter | GLM-5.2, NVIDIA Nemotron 3 Ultra 550B, Poolside Laguna S 2.1 | 50 RPD (no balance); 1,000 RPD (with $10+ lifetime credit) | 20 RPM | No | Vision and tool calling on select free models | Ongoing; 25+ free models rotate constantly |
| Mistral La Plateforme | Mistral Small 3.2, Devstral Small, Codestral | 1 req/sec (Experiment tier) | ~1 RPS | Phone verify | Vision (Pixtral), Code (Codestral) | Ongoing Experiment tier |
| GitHub Models | GPT-5-mini, Llama 4 Scout, DeepSeek-R1 | 10-150 RPD (model class dependent) | 10-15 RPM | No | Vision on GPT-4.1 / o4-mini class models | Ongoing for GitHub account holders |
| SambaNova Cloud | Llama 3.3 70B, DeepSeek-V3.1, gpt-oss-120b | Model-specific RPM caps | 10-30 RPM | No | Llama Vision (rate limited) | Ongoing free tier |
| Hugging Face Inference | Llama 3.2 11B Vision, Qwen2.5, FLUX.1 Dev | $0.10/month in credits (free users) | Cold-start dependent | No | Image gen, audio, STT | Ongoing (shared infra) |
| Cohere | Command A, Command R7B, Rerank 3.5 | 1,000 API calls/month | 20 RPM | No | Embed v4, Rerank, STT (5 RPM) | Ongoing, non-commercial only |
| DeepSeek | DeepSeek-V4-Flash, DeepSeek-R1 | Signup credit varies by promo | Low RPM without paid tier | After credit | 1M context, reasoning mode | One-time credit, then pay-as-you-go |
| Cloudflare Workers AI | Llama 4 Scout, Mistral 7B, Gemma models | 10,000 neurons/day | Standard Workers limits | No | Whisper STT, text-to-image, embeddings | Ongoing free plan |
| Fireworks AI | Apriel 1.6 15B Thinker, DeepCoder 14B, Sarvam M | $1 signup credit, plus 6 permanently free models | 10 RPM without card; 6,000 RPM with one | No | Image, speech, embeddings on paid tier | 6 models free forever; $1 credit is one-time |
| Novita AI | Llama 3.2 1B, Qwen2.5-7B, GLM-4-9B | ~$0.50 signup voucher (1yr); $10 per referral, up to $500 | ~60 RPM (third-party reported) | No | Image gen, embeddings (BGE-M3) | 5 models at $0/token ongoing; referral credits stack |
Provider Deep Dives
Top models ranked using live OpenRouter usage and Artificial Analysis benchmark data, pulled August 20, 2026
Do not use for confidential or client data. NVIDIA's free-tier terms permit using inputs and outputs to improve their own models.
Watch out: Google has cut free quotas before without warning; Pro-class models get no free access at all as of April 2026.
Note: Each model has its own RPD ceiling; the 8B Instant model gets by far the highest daily allowance.
Best for: Quick speed tests; move to the $10 tier fast if you need Cerebras for real work.
:free, or see the Free Router section below. The 50 vs. 1,000 RPD split applies to your total free-model usage across ALL free models combined, not per model.https://models.inference.ai.azure.comsn-).What Cursor, Codex, and Antigravity Actually Give You Free
The three coding environments founders ask about most, updated for August 2026
| Tool | Free Tier | What's Included | Paid Trigger |
|---|---|---|---|
| Cursor | Hobby plan, $0 | Limited Agent requests and limited Tab completions, no credit card. Cursor no longer publishes a fixed numeric quota for Hobby, so treat any specific number you see elsewhere as unreliable | Pro is $20/mo with unlimited Auto mode that never consumes credits; Pro+ $60/mo; Ultra $200/mo |
| OpenAI Codex | Bundled with ChatGPT Free and Go plans | Codex extension works in VS Code with a ChatGPT login; free accounts get reduced usage caps vs. Plus, Pro, Business, Enterprise | Higher limits and priority require ChatGPT Plus ($20/mo) or above |
| Google Antigravity | Free Individual tier | Google's own pricing page describes the free tier as having "Basic weekly rate limits" without publishing exact numbers; ships with Gemini 3.5 Flash access | Higher tiers advertise "more generous rate limits," also without published figures |
| Cline (VS Code extension) | Extension is free forever, open source | No tokens of its own. Bring your own key (BYOK) from any provider in this guide, or use its built-in usage-billing option where a small number of models are tagged FREE | Usage is billed by whichever provider key you connect, not by Cline itself |
The OpenRouter Free Models Router
One endpoint, zero cost, automatic model selection
OpenRouter ships a special model slug called openrouter/free that removes the need to pick a specific free model yourself.
How It Works
The Recommended Free Stack
How to cover most of your agent workload at $0/month
Agent Workflow Stack (August 2026)
Low-Cost Subscriptions for Unlimited Open Weights
What you actually pay, what it includes, and the top models in each plan, verified August 2026
These are flat-fee subscriptions, not pay-per-token pricing. Use these once you outgrow free tier rate limits and need guaranteed, repeatable throughput for a running agent.
| Provider / Plan | Price | What's Included | Top 3 Open Weight Models (August 2026) | Best For |
|---|---|---|---|---|
| Featherless Premium | $25/mo | Unlimited monthly requests, 4 concurrent units, access to any model in the catalogue including Kimi and GLM, private/anonymous usage, no logs | Kimi K2, GLM-5.3, DeepSeek V4 Flash | Founders who want one flat fee for any open model, any size |
| Featherless Token-Based Business | $25+/mo | Credit-based monthly billing that scales with usage, 8 concurrent units, 256K context, one agent sandbox included | DeepSeek V4 Pro, GLM-5.3, Nemotron 3 Ultra | Running persistent agents (Hermes Agent, OpenClaw) at scale |
| GLM Coding Lite | $18/mo (~$12.60/mo annual) | ~80 prompts per 5-hour cycle, ~400/week, 100 MCP calls/month, flagship GLM access via Claude Code, Cline, or OpenCode | GLM-5.3, GLM-5.2, GLM-4.7 | Day-to-day coding on small to mid repos |
| GLM Coding Pro | $72/mo (~$50.40/mo annual) | ~400 prompts per 5 hours, ~2,000/week, 1,000 MCP calls/month, priority queueing | GLM-5.3, GLM-5.2, GLM-4.7 | Heavier individual development workloads |
| GLM Coding Max | $160/mo (~$112/mo annual) | ~1,600 prompts per 5 hours, ~8,000/week, 4,000 MCP calls/month, peak-hour priority | GLM-5.3, GLM-5.2, GLM-4.7 | Teams or power users running agents continuously |
| Hugging Face PRO | $9/mo | 8x higher ZeroGPU quota, 25 min/day H200 compute, 1TB private storage, higher inference credits | Qwen2.5-72B, Llama 3.3 70B, FLUX.1-dev | Prototyping without cold-start delays |
Specialty Models: Image, Video, TTS/STT
Free access beyond text generation
| Modality | Provider | Free Model | Free Limit | Access |
|---|---|---|---|---|
| Image Generation | Hugging Face | FLUX.1 Dev | $0.10/month credit, shared infra | HF Inference API, read token |
| Image Generation | Cloudflare Workers AI | Stable Diffusion XL | 10K neurons/day | Workers AI API |
| Speech-to-Text | Groq | Whisper Large v3 / Turbo | 2,000 audio req/day | console.groq.com audio endpoint |
| Speech-to-Text | Cloudflare Workers AI | Whisper | Included in 10K neurons/day | Workers AI API |
| Speech-to-Text | Cohere | Audio Transcriptions | 5 RPM (trial key) | Cohere trial key |
| Text-to-Speech | ElevenLabs | Multilingual v2 / Flash / Turbo | 10,000 credits/month (~10 min audio), no commercial license | elevenlabs.io free plan |
| Vision / Multimodal | Google AI Studio | Gemini 3.5 Flash | ~15 RPM, up to 1,500 RPD | ai.google.dev |
| Vision / Multimodal | OpenRouter | Select :free vision models | 20 RPM / 50-1,000 RPD | openrouter.ai |
| Embeddings | Cohere | Embed v4 | Included in 1,000 calls/mo trial | Cohere trial key |
| Embeddings | Cloudflare Workers AI | BGE / Multilingual embeddings | 10K neurons/day | Workers AI API |
| Embeddings | Novita AI | BGE-M3 | $0/token, ongoing | api.novita.ai OpenAI-compatible endpoint |
| Reranking (RAG) | Cohere | Rerank 3.5 | Included in 1,000 calls/mo trial | Cohere trial key |
| Video Generation | Z.ai / GLM | CogVideoX-3 (paid, low cost) | Not free, roughly $0.20-0.40/video | docs.z.ai |
Benchmarks: Top Free Models Right Now
5 leading open weight models plus the top free Gemini models, live from OpenRouter's Artificial Analysis data, August 20, 2026
| Model | Type | Coding Index | Intelligence Index | Agentic Index | Context Window | Best Used For |
|---|---|---|---|---|---|---|
| GLM-5.2 (Z.ai) | Open Weight, Free | 68.8 | 52.6 | 45.7 | 256,000 tokens on OpenRouter (1M native) | Long-horizon coding and agent workflows; the top free open weight model by usage and score |
| NVIDIA Nemotron 3 Ultra 550B | Open Weight, Free | 49.3 | 38.3 | 27.5 | 1,000,000 tokens | Long-context orchestration tasks where a 1M window matters more than raw coding score |
| Poolside Laguna S 2.1 | Open Weight, Free | 70.2% Terminal-Bench 2.1 | Not yet indexed | Purpose-built coding agent | 262,144 tokens | Dedicated coding agent tasks; newest top-usage free model on OpenRouter |
| Google Gemma 4 31B | Open Weight, Free | 43.4 | 29.7 | 14.4 | 262,144 tokens | General reasoning with vision input at a smaller compute footprint |
| NVIDIA Nemotron 3 Super 120B | Open Weight, Free | 37.7 | 25.7 | 8.8 | 262,144 tokens | Complex multi-agent applications needing compute efficiency over raw benchmark scores |
| Gemini 3.5 Flash (Google) | Closed, Free Tier | 70.1 | 52.0 | 39.7 | 1,048,576 tokens | The best all-around free model for coding agents, multimodal input, and everyday reasoning |
| Gemini 3.5 Flash-Lite (Google) | Closed, Free Tier | 49.3 | 37.4 | 27.2 | 1,048,576 tokens | High-volume, latency-sensitive subagent tasks inside larger multi-agent workflows |
Fallback Benchmarks: Low-Cost Frontier Models
Where to go when free limits run out, live pricing and benchmarks from OpenRouter, August 20, 2026
| Model | Price (In/Out per 1M) | Coding Index | Intelligence Index | Agentic Index | Context Window | Best Used For |
|---|---|---|---|---|---|---|
| Gemini 3.5 Flash | $1.50 / $9.00 | 70.1 | 52.0 | 39.7 | 1,048,576 tokens | Best cost-to-capability ratio for coding agents and multimodal work |
| Gemini 3.6 Flash | $1.50 / $7.50 (standard); $0.75 / $3.75 batch | Not separately indexed on current catalog | Strong for its price tier | Google's current recommended default, succeeding 3.5 Flash | 1,048,576 tokens | Everyday business tasks, drafts, summarization at low frontier-adjacent cost |
| GPT-5.6 Luna | $0.20 / $1.20 | 71.4 | 52.3 | 46.9 | 1,050,000 tokens | Best value in this table; near-Terra coding quality at roughly a tenth of the price |
| Claude Sonnet 5 | $2.00 / $10.00 | 71.5 | 55.3 | 49.7 | 1,000,000 tokens | Highest coding accuracy per dollar among named frontier chat models; strong default for serious agent coding |
| Grok 4.5 | $2.00 / $6.00 (under 200K context; 2x above) | 72.4 | 55.8 | 48.9 | 500,000 tokens | Token-efficient coding agent with the highest coding index of this group at its price point |
| GPT-5.6 Terra | $2.00 / $12.00 | 76.7 | 56.6 | 50.2 | 1,050,000 tokens | Highest coding and agentic scores of any mid-priced model in this comparison |
5 Rules for Staying Inside Free Limits
Free LLM API Guide
Use more AI without the token bill.
Free monthly guide. Free weekly newsletter.You are subscribed
A few useful newsletters for you.
PromptHacker membership
The executive-tested AI skills library.
Practical skills for the decisions and actions executives need to move forward.Inside the skills library
- Executive strategyDecision briefs, scenarios, and resource allocation.
- MarketingSource-grounded briefs, campaigns, and performance diagnosis.
- SalesICP definition, discovery analysis, and business cases.
- ProductCustomer discovery, requirements, prioritization, and launches.
- OperationsProcess diagnosis, capacity planning, and incident action plans.
The full skills library, including the prompts and implementation notes behind every playbook.
The full skills library and archive, plus a private Slack community for executives applying AI across a team.
One useful extra
Keep this free resource too.
HubSpot's AI Search Sensor is ready if it helps your team.See where AI search is changing discovery for your business.
Sponsored resource.Your guide is ready
Confirm your secure access.
Open the secure link in your email. The Free LLM Guide will unlock immediately.We sent the sign-in link to the address you used to subscribe.