Files
penny-pincher-provider/README.md
T
tiennm99 1da36d245a docs(readme): add 13 free LLM providers and drop project claude settings
Adds LongCat, SenseNova, AMD Token Factory, Volcengine Ark, Baidu Qianfan,
iFlytek Spark, AIHubMix, OVHcloud AI Endpoints, LLM7.io, Token Harbor,
Aion Labs, Experiential Labs, and Empero to the free providers section,
each with limits, endpoint, source, and check date.

Removes .claude/settings.json; permissions now live in the ignored
settings.local.json.
2026-09-21 09:36:31 +07:00

638 lines
25 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Penny-Pincher Provider
Curated list of affordable and free LLM API providers — Claude Code guest passes,
coding-plan subscriptions, and free-tier APIs. Maintained for developers who want
capable models without premium pricing.
Contributions welcome — open a pull request or
[issue](https://github.com/tiennm99/penny-pincher-provider/issues/new) to add or
update an entry.
---
## Claude Code Guest Passes
One week of free Claude Pro (includes Claude Code). New users only; requires payment
info but can be cancelled before the trial ends.
**My passes:**
- ~~<https://claude.ai/referral/ZkoAngod1A>~~ — out of stock as of 2026-04-30
Have a spare pass? Open a PR adding your link, or open an issue.
## Claude AI Ecosystem
### [AgentKit](https://agentkit.best/)
AgentKit (formerly **ClaudeKit**, `claudekit.cc`) sells production-ready kits of skills, slash
commands, subagents, and workflows for coding agents — Claude Code, Codex, GitHub Copilot, and
others. **Engineer** ($99) covers frontend, backend, database, DevOps, code review, and debugging;
**Marketing** ($99) adds MCP integrations and subagents for lead research, SEO, and outreach;
the bundle is **$149** (108+ skills, 95+ commands, 45 subagents). A macOS/Windows desktop app for
the `ak` CLI is on a waitlist.
> Referral: **20% off** your first purchase via <https://agentkit.best/?ref=BWA910UK> (code: `BWA910UK`).
*Checked Aug 28, 2026.*
---
## Providers with Coding Plans
Monthly subscriptions with request-based (not token-based) quotas, compatible with
Claude Code, Cursor, Cline, and similar tools.
### [Z.ai](https://z.ai/subscribe)
Plans from **$18/month** (Lite/Pro/Max). OpenAI-compatible + Anthropic-compatible endpoint.
Models: GLM-5.1, GLM-5-Turbo, GLM-4.7, GLM-4.5-Air.
> Referral: <https://z.ai/subscribe?ic=PLKIAYEIPW>
*Checked Jun 5, 2026.*
### [MiniMax](https://platform.minimax.io)
Token Plan — priced per **API call**, not token. $20–$120/month.
- Plus ($20): 4-5 agents, all models on the API platform
- Max ($50): 6-7 agents, all models on the API platform
- Ultra ($120): 6-7 agents, all models on the API platform
Models: M3 (frontier multimodal coding, 1M context), M2.7 (language), M2.7-highspeed, speech-2.8-hd/turbo, Music-2.6, Hailuo 2.3 (video). Yearly plans ~17% off.
> Referral (10% off) until **Jul 1, 2026** — **For Referred Users:** 10% off subscription + become a dev ambassador. **For Referrers:** 10% back in API voucher per paid referral, usable across all MiniMax models, plus priority access to events and model previews. [View details](https://platform.minimax.io/subscribe/token-plan?code=CAQ5sxHAq6&source=link)
*Checked Jun 5, 2026.*
### [Kimi Code](https://www.kimi.com/code)
Moonshot's coding perk bundled with Kimi membership. Models: Kimi K-series. Rolling 5-hour quota window.
Tiers: Adagio (free), Andante, Presto. Pay-as-you-go also at `platform.moonshot.ai`.
> Referral (code: `C8CJ6F`) — sign up or subscribe via my link and we each get a guaranteed benefit, up to **1-Year Membership Credits**:
> - Sign up: <https://kimi-bot.com/activities/viral-referral/share?scenario=invite&from=share_poster&invitation_code=C8CJ6F>
> - Subscribe: <https://kimi-bot.com/activities/viral-referral/share?scenario=subscribe&from=share_poster&invitation_code=C8CJ6F>
*Checked Jun 5, 2026.*
### Alibaba Cloud Model Studio
> Referral: Up to **$1,700** in free trial credits via <https://www.alibabacloud.com/campaign/benefits?referral_code=A92LU5> (code: `A92LU5`).
#### [Token Plan (Team Edition)](https://www.alibabacloud.com/help/en/model-studio/token-plan-overview)
Credit-based, per-seat subscription (Singapore region only):
- Standard: **$30/seat/month** — 25,000 Credits/seat
- Advanced: **$100/seat/month** — 100,000 Credits/seat
- Premium: **$200/seat/month** — 250,000 Credits/seat
- Shared quota pack: **$700** — 625,000 Credits (1-month validity; unused Credits expire)
Models: qwen3.6-plus, glm-5, MiniMax-M2.5, deepseek-v3.2 (text); qwen-image-2.0/2.0-pro, wan2.7-image/image-pro (image).
#### [Coding Plan](https://www.alibabacloud.com/help/en/model-studio/coding-plan)
Pro plan: **$50/month** — 6,000 req/5-hour, 45,000 req/week, 90,000 req/month.
Models: qwen3.5-plus, kimi-k2.5, glm-5, MiniMax-M2.5.
Tools: Claude Code, Cursor, Cline, Codex, and more. Lite plan no longer accepting new subscribers.
**Note (as of Jun 5, 2026):** effectively unbuyable — perpetually out of stock. Claimed restock at 00:00 GMT+8, but repeated attempts still fail to purchase.
*Checked Jun 5, 2026.*
### [opencode — Go](https://opencode.ai/go)
Open-source `opencode` CLI subscription. $5 first month, **$10/month** thereafter.
Models: GLM-5/5.1, Kimi K2.5/K2.6, MiniMax M2.7/M3, Qwen3.5/3.6/3.7 Plus, MiMo-V2.5(-Pro), DeepSeek V4 Flash/Pro. Per-5-hour limits vary (200–10,200 req).
API key portable — works with Claude Code via LiteLLM proxy or `oc-go-cc`. Model format: `opencode-go/<model-id>`.
My referral:
> Invite friends to OpenCode Go. Earn $5 when a friend subscribes, and they'll get $5 too. Share your referral link; your friend joins and subscribes to Go; you both get a $5 usage credit to apply toward your Go usage limits.
>
> Referral link: <https://opencode.ai/go?ref=HE42WGS8BM>
*Checked Jun 5, 2026.*
### [Synthetic](https://synthetic.new/)
Privacy-first inference (no training on prompts/responses). **$30/month** subscription or pay-as-you-go.
Models: Kimi K2.6, MiniMax M2.5, GLM 5.1, GLM 4.7 Flash, vLLM-compatible open-source models.
OpenAI-compatible — works with Roo, Cline, Octofriend.
> Referral: **$10.00** in subscription credit via <https://synthetic.new/?referral=CNBFyw28zF0dZoj>
*Checked Jun 5, 2026.*
### [BigModel.cn — GLM Coding Plan](https://www.bigmodel.cn/glm-coding)
The Chinese (mainland) counterpart of Z.ai's GLM Coding Plan — same underlying Zhipu AI models, but billed in CNY through bigmodel.cn. Suited for users who can pay via Alipay / WeChat Pay or already have a 智谱 AI account.
Plans (monthly, after the 2026 price adjustment):
- **Lite**: ¥49/month
- **Pro**: ¥149/month
- **Max**: ¥469/month
All tiers support GLM-5.1, GLM-5-Turbo, GLM-4.7, and GLM-4.5-Air. Compatible with Claude Code, Cline, and 20+ coding tools via OpenAI-compatible API plus an Anthropic-compatible endpoint.
Referral program (challenge-based, resets every 30 invitees):
- Invited friend gets **5% off** their first GLM Coding Plan order.
- Referrer gets **10% cashback** once 3 friends subscribe, plus an **additional 10%** of the total paid amount for every 30 invitees.
- Rebate credit is usable for resource packs, API calls, and subscription renewals on the BigModel platform.
Source: <https://www.bigmodel.cn/glm-coding>, <https://docs.bigmodel.cn/cn/coding-plan/overview> [^bigmodel]
Homepage: <https://www.bigmodel.cn/glm-coding>
My referral:
>🚀 Join the GLM Coding Plan via my link — get 5% off your first order. Subscribe at https://www.bigmodel.cn/glm-coding?ic=VGRZKHKNKW (invitation code: `VGRZKHKNKW`).
[^bigmodel]: Checked on Jun 5, 2026
### [BytePlus ModelArk — Coding Plan](https://www.byteplus.com/en/activity/codingplan)
ByteDance. Lite: **$15/month**, Pro: **$35/month** (intro promo $5/$25 ended early 2026).
Models: ByteDance-Seed-2.0, DeepSeek-V3.2, GLM-5.1, Kimi-K2.5.
Tools: Claude Code, Cursor, Cline, Roo Code, OpenCode.
> Referral: <https://www.byteplus.com/activity/codingplan?ac=MMAUCIS9NT1S&rc=2739UWRE>
*Checked Jun 5, 2026.*
### [Xiaomi MiMo Open Platform](https://platform.xiaomimimo.com)
I'm on Xiaomi MiMo Open Platform — running Xiaomi's flagship MiMo V2.5 and the rest of the lineup. Sign up with my code and you'll instantly get $2 in API credits.
After signup, enter the code at the bottom-left of the console. Credits valid 40 days.
**Token Plan** (monthly, launched May 2026): Lite ¥39 (60M credits), Standard ¥99 (200M), Pro ¥329 (700M), Max ¥659 (1,600M). Models: MiMo-V2.5, MiMo-V2.5-Pro (2× credit cost). Annual plans discounted.
> Referral: Code `T8ESAY` · <https://platform.xiaomimimo.com?ref=T8ESAY>
*Checked Jun 5, 2026.*
### [GitHub Copilot Pro](https://github.com/features/copilot/plans)
Cheapest mainstream coding seat. **$10/month** ($100/year) — unmetered code completions and
chat on the base model, plus **1,500 premium requests/month** for frontier models
(Claude, GPT, Gemini). Copilot Pro+ is $39/month with 6,000 premium requests.
Free tier: 2,000 completions + 50 premium requests/month, no card.
Free for verified students, teachers, and maintainers of popular open-source repos.
Tools: VS Code, JetBrains, Neovim, Xcode, Visual Studio, `gh copilot` CLI, and the Copilot
coding agent on github.com.
*Checked Sep 11, 2026.*
---
## Free Providers
### [TokenRouter](https://www.tokenrouter.com/)
Unified AI gateway with OpenAI-compatible access to 300+ models.
- **Kimi K3 Free Version:** Use `moonshotai/kimi-k3-free` at $0 until **Aug 12, 2026**, running on TokenRouter's own B300 deployment.
- **API:** Use the OpenAI Chat Completions endpoint at `https://api.tokenrouter.com/v1` with a TokenRouter API key.
- **Limit:** Free compute capacity is limited, so stability and concurrency are not guaranteed.
Source: [TokenRouter models](https://www.tokenrouter.com/models) [^tokenrouter]
[^tokenrouter]: Check at Aug 5 2026
### [OrcaRouter](https://www.orcarouter.ai/)
OpenAI-compatible gateway with access to 200+ models, automatic routing, and failover.
> Referral: <https://www.orcarouter.ai/ref/ref_3976ba42abf37dc55c1d> (code: `ref_3976ba42abf37dc55c1d`).
- **Kimi K3:** Get $5 in credit as a new user before **Aug 6, 2026**; a payment card is required.
- **Tencent HY3:** Get $5 in credit before **Aug 21, 2026**.
- **Claude Opus 5:** Get a 60% deposit match, up to $500, when you join before **Aug 24, 2026**.
- **Free DeepSeek models:** Get 100 initial V4 Flash calls and 30 initial V4 Pro calls with no published deadline.
- **Grok 4.5:** The offer is sold out, but its waitlist is still open.
Offers can change quickly; check the [live offers page](https://www.orcarouter.ai/offers) before claiming. [^orcarouter]
[^orcarouter]: Check at Aug 5 2026
### [OpenRouter](https://openrouter.ai)
Free models (`:free` suffix): 20 RPM, 50 req/day (free accounts); 1,000 req/day after $10 top-up.
*Checked Jun 5, 2026.*
### [NVIDIA NIM](https://build.nvidia.com)
Free for NVIDIA Developer Program members, no credit card. ~40 RPM. 1,000 inference credits at signup (consumption-based, not a fixed monthly request cap).
Models: Kimi K2.5, GPT-OSS, DeepSeek-V3.2, Llama 3.x, Mistral, Phi, Nemotron. OpenAI-compatible.
*Checked Jun 5, 2026.*
### [OpenCode Zen](https://opencode.ai/docs/zen)
Hand-picked free models that change periodically. Optimized for coding agents.
| Model | Status |
|---|---|
| `Big Pickle` | Free |
| `DeepSeek V4 Flash Free` | Free |
| `MiMo-V2.5 Free` | Free |
| `Nemotron 3 Ultra Free` | Free |
OpenAI/Anthropic compatible endpoints at `https://opencode.ai/zen/v1/`.
*Checked Jun 5, 2026.*
### [Google AI Studio](https://aistudio.google.com/)
Google's developer platform for Gemini models. Generous free tier with pay-as-you-go available.
Free-tier limits vary by model (check the AI Studio dashboard for your project):
| Model (Free) | RPM | RPD |
|---|---|---|
| Gemini 2.5 Pro | 5 | 100 |
| Gemini 2.5 Flash | 10 | 250 |
| Gemini 2.5 Flash-Lite | 15 | 1,000 |
Paid tier: higher limits, usage-based billing.
Models: Gemini 2.5 Pro/Flash/Flash-Lite. OpenAI-compatible endpoint available.
**Warning:** In the Free tier, Google may use your prompts and responses to improve their products. Use the Paid tier or Vertex AI for privacy.
*Checked Jun 5, 2026.*
### [Kilo Code — Gateway](https://kilo.ai/gateway)
VS Code + JetBrains coding extension with built-in gateway.
Free: `:free` models, 200 req/hour per IP. First top-up: $20 bonus credits (60-day expiry).
BYOK supported — no Kilo subscription required for your own provider keys.
*Checked Jun 5, 2026.*
### [Cloudflare Workers AI](https://developers.cloudflare.com/workers-ai/)
10,000 Neurons/day (resets 00:00 UTC). ~150 LLM responses/day on Llama 3.3 70B.
Models: Llama 3.3 70B, Gemma 3, Qwen2.5-Coder, DeepSeek R1, BGE embeddings, Whisper.
*Checked Jun 5, 2026.*
### [DeepSeek Platform](https://platform.deepseek.com)
5M free tokens at signup (no code needed). OpenAI + Anthropic compatible.
PAYG: V4 Flash $0.14/M in, $0.28/M out; V4 Pro $0.435/M in, $0.87/M out (cached: $0.03/M).
*Checked Jun 5, 2026.*
### [Groq](https://console.groq.com)
LPU inference. Free tier, no credit card.
| Model | RPM | RPD | TPM | TPD |
|---|---|---|---|---|
| `llama-3.1-8b-instant` | 30 | 14,400 | 6K | 500K |
| `llama-3.3-70b-versatile` | 30 | 1,000 | 12K | 100K |
| `whisper-large-v3` | 20 | 2,000 | — | — |
*Checked Jun 5, 2026.*
### [GitHub Models](https://github.com/marketplace/models)
Free for all GitHub accounts. OpenAI-compatible at `https://models.github.ai/inference`.
Covers OpenAI, Anthropic, Llama, Mistral, DeepSeek, Grok, Phi. Rate limits vary per model
(e.g. GPT-4o: 10 RPM/50 RPD; DeepSeek-R1: 15 RPM/150 RPD). Requires PAT with `models:read`.
*Checked Jun 5, 2026.*
### [xAI Grok API](https://x.ai/api)
$25 signup credits + $150/month via Data Sharing Program (eligible countries).
Models: Grok 4, Grok 4.1 Fast (2M ctx), Grok Code Fast. OpenAI + Anthropic compatible.
**Warning:** Data Sharing opt-in is irreversible and lets xAI train on your prompts.
*Checked Jun 5, 2026.*
### [Mistral La Plateforme](https://console.mistral.ai)
Free Experiment plan — up to ~1B tokens/month, no credit card (phone verification only).
Models: Mistral Large 3, Medium 3, Small 3.1, Codestral, Pixtral, embeddings. OpenAI-compatible.
*Checked Jun 5, 2026.*
### [Google Cloud Vertex AI](https://cloud.google.com/vertex-ai)
$300 free credits for 90 days (new GCP customers). Not a recurring free tier.
Models: Gemini 3 Pro/Flash, Anthropic Claude on Vertex, DeepSeek, GLM, Qwen via MaaS.
Express Mode available without billing for limited evaluation quotas.
*Checked Jun 5, 2026.*
### [Hugging Face Inference Providers](https://huggingface.co/docs/inference-providers)
Routes across Together, Fireworks, Novita, Cerebras, Replicate, DeepInfra, Scaleway.
Free: 100K monthly credits. PRO ($9/month): 2M credits. OpenAI-compatible at `https://router.huggingface.co/v1`.
*Checked Jun 5, 2026.*
### [Cerebras Cloud](https://cloud.cerebras.ai)
Wafer-scale chip inference. Free tier, no credit card.
| Model | RPM | TPD |
|---|---|---|
| `gpt-oss-120b` | 30 | — |
| `llama3.1-8b-instant` | 30 | 500K |
| `qwen-3-235b-a22b-instruct-2507` | 30 | — |
| `zai-glm-4.7` | 10 | — |
1M tokens/day shared cap. 8K context limit on free tier.
*Checked Jun 5, 2026.*
### [BigModel.cn](https://www.bigmodel.cn/)
Zhipu AI (智谱 AI). 25M free tokens for new users to explore the API, playground, and AGI apps. GLM-4.7-Flash and GLM-4.5-Flash are permanently free.
> Referral: <https://www.bigmodel.cn/invite?icode=rIX6uZrLYfy8fQ6Urca4xf2gad6AKpjZefIo3dVEyA%3D>
*Checked Jun 5, 2026.*
### [Fireworks AI](https://fireworks.ai)
$1 starter credits. 50+ models. Function calling, MCP support. OpenAI-compatible.
*Checked Jun 5, 2026.*
### [Scaleway Generative APIs](https://www.scaleway.com/en/generative-apis/)
EU/GDPR, Paris. 1M free tokens for new customers (no time limit advertised).
Models: Qwen3 (235B/397B/coder), Llama 3.3 70B, Mistral Small 3.2, DeepSeek R1 distill, Pixtral.
*Checked Jun 5, 2026.*
### [SambaNova Cloud](https://cloud.sambanova.ai)
RDU (dataflow chip) inference. Developer tier, no credit card.
Free: **$5** in credits (expire after 30 days), 20 RPM, 200K tokens/day per model.
Models: DeepSeek, Llama 3.x/4, Qwen3, Whisper. OpenAI-compatible.
Source: <https://sambanova.ai/blog/sambanova-cloud-developer-tier-is-live>
*Checked Sep 11, 2026.*
### [Cohere](https://dashboard.cohere.com/api-keys)
Trial API key, no credit card. **1,000 calls/month**, 20 RPM.
Models: Command A+, Command R, Command R7B, Aya (11+ models), plus rerank and embeddings.
**Warning:** trial keys are **non-commercial use only** — production needs a production key (paid).
*Checked Sep 11, 2026.*
### [Vercel AI Gateway](https://vercel.com/docs/ai-gateway)
Single OpenAI-compatible endpoint routing to many providers, with failover and BYOK.
Free: **$5/month** in renewing credits (does not roll over). Free tier covers a subset of models and applies lower per-model rate limits than paid.
Source: <https://vercel.com/docs/ai-gateway/pricing>
*Checked Sep 11, 2026.*
### [Requesty](https://www.requesty.ai/)
LLM gateway with routing, caching, and spend controls. Works with Claude Code, Cline, Cursor, Roo.
Free: **200 req/day** on the free-model catalogue. No credit card, no trial clock — same platform as pay-as-you-go, just restricted to free models until you upgrade.
Source: <https://www.requesty.ai/free-models>, <https://www.requesty.ai/pricing>
*Checked Sep 11, 2026.*
### [SiliconFlow](https://cloud.siliconflow.cn/)
Chinese multi-model inference platform, 200+ LLM/image/audio/video models.
Free: a set of smaller open-source models is **permanently $0**; new international accounts get ~**$1** starter credit. 1,000 RPM / 50,000 TPM on the free models. Identity verification required.
Models (free): Qwen3-8B (128K ctx) and similar small open models. OpenAI-compatible.
*Checked Sep 11, 2026.*
### [ModelScope](https://modelscope.cn/)
Alibaba's model community (魔搭). API-Inference free for registered users.
Free: **2,000 req/day** total, **≤500 req/day per model**. Requires an Alibaba account with real-name verification (mainland China ID).
Models: Qwen, DeepSeek, GLM and 50+ other open models.
*Checked Sep 11, 2026.*
### [Ollama Cloud](https://ollama.com/cloud)
Hosted counterpart to the local `ollama` CLI — same commands and API, models run on Ollama's GPUs.
Free tier with per-session and weekly caps (exact numbers unpublished). No credit card.
Models: DeepSeek, Kimi, MiniMax, GPT-OSS, Qwen and 16 model families. OpenAI-compatible.
*Checked Sep 11, 2026.*
### [LongCat (Meituan)](https://longcat.chat/platform/docs/zh/)
Meituan's open platform for **LongCat-2.0** (1.6T-param MoE agentic model, **1M context**, 128K max output).
Both **OpenAI** (`https://api.longcat.chat/openai`) and **Anthropic** (`https://api.longcat.chat/anthropic`)
formats — works with Claude Code directly.
Free: signup grants a daily quota plus a one-time resource pack (exact figures only shown in the console;
third-party reports ~500K tokens/day, raisable on request). Cache hits don't consume quota. Returns 429
with exponential-backoff guidance when rate-limited.
Source: <https://longcat.chat/platform/docs/zh/>, <https://github.com/peter123023/awesome-free-llm-api>
*Checked Sep 21, 2026.*
### [SenseNova Token Plan](https://www.sensenova.cn/token-plan)
SenseTime (商汤). Public beta — **Free tier ¥0/month**, phone-number signup, no card, no ID verification.
OpenAI-compatible at `https://token.sensenova.cn/v1` plus an Anthropic-compatible endpoint.
Limits: **60,000 credits / rolling 5 h** and **600,000 credits / rolling week**, separately for the general
pool and the Flash-Lite pool. Up to 20 API keys.
Models at 0 credits: SenseNova 6.8 Flash-Lite, SenseNova U1 Fast, DeepSeek V4 Flash / V4 Pro, GLM-5.2, Kimi K3.
**Warning:** SenseTime says paid Lite/Pro tiers are "coming soon" with no end date for the free beta —
don't build production on it.
Source: <https://platform.sensenova.cn>, <https://github.com/peter123023/awesome-free-llm-api>
*Checked Sep 21, 2026.*
### [AMD Token Factory (Radeon Cloud)](https://developer.amd.com.cn/radeon/tokenfactory)
AMD's official inference platform on Radeon GPUs. **~$10-equivalent free quota per day**, resets daily
(does not roll over). OpenAI-compatible at `https://developer.amd.com.cn/radeon/api/v1`.
Free models: DeepSeek V4 Flash 0731, MiniCPM5-1B; GLM-5.3-Flash and Qwen3.8-Flash-Next as limited-time free.
High time-to-first-token (~20 s) reported.
Source: <https://github.com/peter123023/awesome-free-llm-api>
*Checked Sep 21, 2026.*
### [Volcengine Ark (ByteDance)](https://console.volcengine.com/ark)
ByteDance's model platform (火山引擎方舟). Permanent free tier: **2M tokens/day**, resets at midnight
(GMT+8). OpenAI-compatible at `https://ark.cn-beijing.volces.com/api/v3`.
Free models: Doubao-Lite, DeepSeek R2 / V3 (within the daily quota). Doubao-Pro and higher tiers are paid.
Requires a phone number with real-name verification (mainland China).
Source: <https://github.com/peter123023/awesome-free-llm-api>
*Checked Sep 21, 2026.*
### [Baidu Qianfan](https://cloud.baidu.com/product/wenxinworkshop)
Baidu's model platform (千帆). **ERNIE-Speed-128K and ERNIE-Lite are permanently free**, rate-limited.
OpenAI-compatible at `https://qianfan.baidubce.com/v2`. Real-name verification required.
Source: <https://github.com/peter123023/awesome-free-llm-api>
*Checked Sep 21, 2026.*
### [iFlytek Spark](https://xinghuo.xfyun.cn/sparkapi)
讯飞星火. **Spark Lite is permanently free with unlimited tokens**, capped at **2 QPS**.
OpenAI-compatible at `https://spark-api-open.xf-yun.com/v1` (APIKey/APISecret auth). Individual real-name
verification required.
Source: <https://github.com/peter123023/awesome-free-llm-api>
*Checked Sep 21, 2026.*
### [AIHubMix](https://aihubmix.com/models/free)
Gateway with **56 free models**, no credit card. Speaks Chat Completions, Messages, and Responses at
`https://aihubmix.com/v1`.
Free: 10 trial calls at signup (never expire). A **one-time $1 top-up** permanently unlocks
**100 req/day + 1M tokens/day** on the free catalogue (shared pool, resets daily).
Free models include glm-4.7-flash, hy3, minimax-m3, k2.6-code-preview, gpt-oss-20b, nemotron-3-ultra/super,
gemma-4-31b, mimo-v2.5(-pro), coding-glm-5.3, gpt-5.5, gemini-3.8-flash.
Source: <https://aihubmix.com/models/free>, <https://github.com/peter123023/awesome-free-llm-api>
*Checked Sep 21, 2026.*
### [OVHcloud AI Endpoints](https://www.ovhcloud.com/en/public-cloud/ai-endpoints/catalog/)
EU-hosted (France). **Anonymous free tier — no API key, no signup**: 2 RPM per IP per model.
OpenAI SDK-compatible at `https://oai.endpoints.kepler.ai.cloud.ovh.net/v1`.
Models (20+): Qwen3.5-397B-A17B, gpt-oss-120b/20b, Llama 3.3 70B, Qwen3.6-27B, Qwen3-Coder-30B,
Qwen2.5-VL-72B, Mistral Small 3.2, Mistral Nemo.
Source: <https://github.com/mnfst/awesome-free-llm-apis>
*Checked Sep 21, 2026.*
### [LLM7.io](https://token.llm7.io)
UK gateway. Anonymous access needs no key (10 RPM, 60 req/hour); a free token from `token.llm7.io` raises
the limits. OpenAI-compatible at `https://api.llm7.io/v1`.
Models: gpt-oss-20b, minimax-m2.7 (180K ctx), Mistral Nemo.
Source: <https://github.com/mnfst/awesome-free-llm-apis>
*Checked Sep 21, 2026.*
### [Token Harbor](https://tokenharbor.ai/pricing)
Small gateway. **Free tier $0/month** with a rolling 4-week quota (amount unpublished, unused quota carries
over). Agent Pass **$1.99/month** ($0.99 first month) tops up the same pool. OpenAI-compatible at
`https://tokenharbor.ai/v1`.
Free models: DeepSeek V4 Flash, DeepSeek V4.1 Flash, MiMo V2.5 ("promotional models added over time").
**Warning:** blocks requests from mainland China, Hong Kong, and Macau (`region_blocked`). Free requests may be logged.
Source: <https://tokenharbor.ai/pricing>, <https://github.com/peter123023/awesome-free-llm-api>
*Checked Sep 21, 2026.*
### [Aion Labs](https://www.aionlabs.ai/app/api-keys/)
Permanent free tier, no credit card. **15 RPM, 20K tokens/day**. OpenAI-compatible at `https://api.aionlabs.ai/v1`.
Models: aion-2.0, aion-3.0, aion-3.0-mini (128K ctx, reasoning), aion-rp-llama-3.1-8b. Tuned for
roleplay/storytelling rather than coding.
Source: <https://github.com/mnfst/awesome-free-llm-apis>
*Checked Sep 21, 2026.*
### [Experiential Labs](https://platform.experientiallabs.ai)
YC-backed "open-source OpenRouter". Free plan: **~500 credits/month** (1 credit = $0.01), hard stop when
exhausted, no auto-upgrade; new orgs get extra welcome credits. OpenAI-compatible at
`https://api.experientiallabs.ai/v1`.
$0-labelled models: Qwen3.8 27B, DeepSeek V4 Flash, GPT-5.6 Luna, GPT-6 Astra, Claude Fable 5.1
(all 1M ctx) — but the monthly credit cap still applies.
**Warning:** business model is free traffic in exchange for training traces; free-tier uptime is low
(72–100 % by model, ~4.5 s TTFT). Prototyping only.
Source: <https://github.com/peter123023/awesome-free-llm-api>
*Checked Sep 21, 2026.*
### [Empero](https://free.empero.org)
Community endpoint from German lab EmperoAI. **Completely free, no signup** — any string works as the API
key (convention: `free`). OpenAI-compatible at `https://free.empero.org/v1`.
Models rotate often: glm-5.3-flash, qwen3.8-flash, Qwen3.8-27B-FP8.
**Warning:** prompts and responses are logged (IP hashed) to train their open models — never send private
data. Frequent `upstream_down` / 503 under load.
Source: <https://github.com/peter123023/awesome-free-llm-api>
*Checked Sep 21, 2026.*
---
## License
Apache 2.0