tiennm99 eb85931584 feat: add 8 free providers, sort Free Providers by attractiveness
Add LongCat AI, Cerebras Cloud, Groq, Cloudflare Workers AI, Scaleway,
Kiro, Pollinations AI, and Google Cloud Vertex AI free credits.
Reorder Free Providers section by recurring quota size, model quality,
and signup friction.
2026-04-28 23:23:46 +07:00
2026-04-02 13:20:30 +07:00
2026-03-25 21:42:07 +07:00
2026-03-25 17:46:51 +07:00

Penny-Pincher Provider

A curated list of affordable (or almost free) LLM providers for people who are unwilling to pay premium prices for AI.

Note: Claude is arguably the best AI coding assistant out there — but it's expensive. If you want to try it before committing, you can use a guest pass for 1 week of free Claude Pro (includes Claude Code).

Guest passes are limited and available on a first-come, first-served basis. New users only — you'll need to enter payment info to activate, but you can cancel before the trial ends to avoid charges. Learn more.

My passes:

Friend's passes (click to reveal)

If you know of other providers, feel free to create a pull request or open an issue here. I will review and add them when possible. Thank you!

If you find this helpful, you can support me by donating or registering with my referral links. Thank you!

Providers with Coding Plans

These providers offer coding plan subscriptions. You can prepay a monthly fee and use their LLM APIs with usage limits.

Z.ai

Offers 3 plans:

  1. Lite: 3x usage of the Claude Pro plan
  2. Pro: 5x Lite plan usage
  3. Max: 4x Pro plan usage

Prices vary based on plan duration (monthly, quarterly, or yearly) and occasional promotional offers, so check the website for current pricing.

Source: https://z.ai/subscribe 1

Homepage: https://docs.z.ai/devpack/overview

My referral:

🚀 You've been invited to join the GLM Coding Plan! Enjoy full support for Claude Code, Cline, and 20+ top coding tools — starting at just $10/month. Subscribe now and grab the limited-time deal! Link: https://z.ai/subscribe?ic=PLKIAYEIPW

MiniMax

Offers a Token Plan — a unified subscription for multimodal AI (text, speech, video, image, music). Pricing is based on API calls, not tokens — very generous!

Plans (monthly):

  • Starter: $10/month — 1,500 M2.7 requests per 5-hour rolling window
  • Plus: $20/month — more requests + speech, image, music generation
  • Max: $50/month — even higher quotas across all models
  • Highspeed tiers ($40–$150/month) — dedicated M2.7-highspeed access, up to 30,000 requests per 5-hour window

Includes access to M2.7 language model, Speech 2.8, Image-01, Hailuo video, and Music-2.5. Yearly plans available with ~17% discount.

Source: https://platform.minimax.io/docs/token-plan/intro, https://platform.minimax.io/docs/guides/pricing-token-plan 2

Homepage: https://platform.minimax.io

My referral:

🎁 MiniMax Token Plan New Year Mega Offer! Invite friends and earn rewards for both! Exclusive 10% OFF for friends. Ready-to-use API vouchers for you! Token Plan Referral Program until May 1, 2026 — referred users get 10% off their subscription and join the dev ambassador community; referrers earn 10% back in API vouchers per paid referral, usable across all MiniMax models, plus priority access to events and model previews. 👉 Get your referral link: https://platform.minimax.io/subscribe/token-plan?code=CAQ5sxHAq6&source=link

Kimi Code

A coding-focused perk included with Kimi membership — drops into any dev workflow (terminal, IDE, or Kimi CLI) and is backed by Moonshot's Kimi K-series models, which are sharply priced per token.

Usage quotas are tracked on a rolling 5-hour window and scale with your Kimi membership tier. Since Kimi Code ships as part of Kimi membership rather than a standalone plan, check the membership page for the current tier list and pricing.

Source: https://www.kimi.com/code, https://www.kimi.com/membership/pricing 3

Homepage: https://www.kimi.com/code

Alibaba Cloud Model Studio — Coding Plan

Monthly subscription for AI coding tools — top Qwen/Kimi/GLM/MiniMax models at fixed, predictable pricing.

Pro plan: $50/month

  • 6,000 requests per 5-hour sliding window
  • 45,000 requests per week (resets Monday 00:00 UTC+8)
  • 90,000 requests per month (resets on subscription anniversary)

Models include qwen3.5-plus, qwen3-max, qwen3-coder, kimi-k2.5, glm-5, and MiniMax-M2.5.

Supported tools: Claude Code, Cursor, Cline (VS Code), OpenCode, Qwen Code, Kilo Code, Kilo CLI, OpenClaw, Codex, and more.

Note: the Lite plan stopped accepting new subscriptions on Mar 20, 2026.

Source: https://www.alibabacloud.com/help/en/model-studio/coding-plan 4

Homepage: https://www.alibabacloud.com/product/modelstudio

BytePlus ModelArk — Coding Plan

ByteDance's ModelArk coding subscription — flat monthly fee, works with mainstream coding tools, models swappable per task.

Standard plans:

  • Lite: $5/month
  • Pro: $25/month

Models include latest ByteDance-Seed-2.0-pro/lite, DeepSeek-V3.2, GLM-4.7, Kimi-K2.5, and GPT-OSS variants.

Supported tools: Claude Code, Cursor, Cline (VS Code), Kilo Code, Roo Code, OpenCode, TRAE, and more.

Note: new-user first-purchase promo pricing was suspended on Mar 17, 2026 — everyone now pays the list price.

Source: https://www.byteplus.com/en/activity/codingplan, https://docs.byteplus.com/en/docs/ModelArk/1925114 5

Homepage: https://console.byteplus.com/ark

opencode — Go

A subscription tier for the open-source opencode CLI that pools access to ~10 open-source coding models behind one flat price — aimed at developers who want generous request limits without premium-provider fees.

Pricing:

  • $5 first month, $10/month thereafter
  • Top up extra credit as needed; cancel anytime

Models include GLM-5.1, GLM-5, Kimi K2.6 (3× quotas through Apr 27), Kimi K2.5, MiMo-V2-Pro/Omni, Qwen3.5/3.6 Plus, MiniMax M2.5/M2.7.

Per-5-hour request limits vary by model tier (≈200 to 10,200).

Source: https://opencode.ai/go 6

Homepage: https://opencode.ai

Synthetic

Run open-source AI models for you in private, secure datacenters.

Privacy-first inference: Synthetic never trains on your data and doesn't store API prompts or completions.

Pricing:

  • Subscription: $30/month (app-based access)
  • Usage-based: pay-as-you-go, no charge for unused capacity

Models include Kimi K2.5, MiniMax M2.5, GLM 5.1, GLM 4.7 Flash, plus any vLLM-compatible open-source LLM.

OpenAI-compatible — works with Roo, Cline, Octofriend, and any other OpenAI-API-compatible client.

Source: https://synthetic.new/ 7

Homepage: https://synthetic.new

Free Providers

Sorted by attractiveness — biggest recurring free quota, model quality, and lowest friction first.

LongCat AI

Meituan's open-source LongCat models. API platform in public beta — no paid tier yet.

Free quota (resets daily 00:00 Beijing Time, no rollover):

  • LongCat-Flash-Lite: 50M tokens/day (no upgrade path — uniformly free)
  • LongCat-Flash-Chat: 500K tokens/day
  • LongCat-Flash-Thinking / Thinking-2601: 500K tokens/day each
  • LongCat-Flash-Omni-2603 (multimodal): 500K tokens/day
  • LongCat-2.0-Preview: 10M tokens / 2 hours (invite-only, 1M context)

Both OpenAI-compatible (https://api.longcat.chat/openai) and Anthropic-compatible (https://api.longcat.chat/anthropic) endpoints. 256K context on most models.

Source: https://longcat.chat/platform/docs/ 8

Homepage: https://longcat.chat/platform

Cerebras Cloud

World-fastest LLM inference (wafer-scale chip), OpenAI-compatible. Free tier — no credit card required.

Free tier limits:

  • 1M tokens/day shared cap across free models
  • 30 RPM most models (10 RPM for zai-glm-4.7)
  • 60K TPM per model
  • Free models: gpt-oss-120b, llama3.1-8b, qwen-3-235b-a22b-instruct-2507, zai-glm-4.7

Pay-as-you-go tier removes daily and per-minute caps for higher throughput.

Source: https://inference-docs.cerebras.ai/support/rate-limits 9

Homepage: https://www.cerebras.ai/inference

OpenRouter

Free usage limits: If you’re using a free model variant (with an ID ending in :free), you can make up to 20 requests per minute. The following per-day limits apply:

  • If you have purchased less than 10 credits, you’re limited to 50 :free model requests per day.
  • If you purchase at least 10 credits, your daily limit is increased to 1000 :free model requests per day.

Source: https://openrouter.ai/docs/api/reference/limits 10

Homepage: https://openrouter.ai

Groq

Fast LPU inference, OpenAI-compatible. Free tier with no credit card.

Free tier (per-model, organization-level):

  • llama-3.1-8b-instant: 30 RPM, 14.4K RPD, 6K TPM, 500K TPD
  • llama-3.3-70b-versatile: 30 RPM, 1K RPD, 12K TPM, 100K TPD
  • whisper-large-v3 (audio): 20 RPM, 2K RPD
  • Also free: gemma2-9b-it, allam-2-7b, gpt-oss variants

Upgrade to Developer plan for higher RPM/TPD, Batch and Flex processing.

Source: https://console.groq.com/docs/rate-limits 11

Homepage: https://groq.com

NVIDIA NIM

NVIDIA-hosted inference for 50+ open models — free for NVIDIA Developer Program members, no credit card required. OpenAI-compatible API at https://integrate.api.nvidia.com/v1 works out of the box with Cline, Roo, OpenCode, and any OpenAI-compatible client.

Free access:

  • Sign up for the free Developer Program → generate an nvapi- API key on build.nvidia.com
  • 1,000 inference credits on signup (some accounts report rate-limit-only model since early 2025)
  • Personal-account rate limits shown in dashboard top-right — typically ~40 RPM and 1,000 requests/month, resetting on the 1st
  • Models include Kimi K2.5, GPT-OSS, DeepSeek-V3.2, Llama 3.x, Mistral, Phi, and NVIDIA's own Nemotron family

Paid self-hosted NIM containers and pay-as-you-go API are available for higher throughput; the hosted free tier is fine for evaluation and light coding use.

Source: https://build.nvidia.com, https://developer.nvidia.com/nim 12

Homepage: https://build.nvidia.com

Cloudflare Workers AI

Serverless inference on Cloudflare's global edge network. Free tier on both Free and Paid Workers plans.

Free allowance:

  • 10,000 Neurons/day (resets 00:00 UTC); failures with 429 after exhaustion
  • 50+ models: LLMs (Llama 3.3 70B, Gemma 3, Qwen2.5-Coder, DeepSeek R1), embeddings (BGE), Whisper transcription, image generation
  • Example mileage: ~150 LLM responses/day on Llama 3.3 70B, or ~500 audio-seconds Whisper
  • Overage billed at $0.011 per 1,000 Neurons on paid plan

Requires Cloudflare account API token + Account ID.

Source: https://developers.cloudflare.com/workers-ai/platform/pricing/ 13

Homepage: https://developers.cloudflare.com/workers-ai/

Google Cloud Vertex AI (free trial credits)

Not a free-forever tier — but new GCP customers get $300 in free credits valid for 90 days, usable across Vertex AI for Gemini 3 Pro/Flash, Anthropic Claude on Vertex, and Vertex Partner models (DeepSeek, GLM, Qwen via MaaS).

Setup:

  • Sign up at https://cloud.google.com/free (credit card required for verification, not charged unless you upgrade)
  • Enable Vertex AI API in your GCP project
  • $300 expires after 90 days; account does not auto-convert to paid
  • Model access varies by region; Gemini 3 Pro/Flash available in most regions

Practical for short-term heavy evaluation; not a long-term free path.

Source: https://cloud.google.com/free, https://cloud.google.com/vertex-ai/generative-ai/pricing 14

Homepage: https://console.cloud.google.com/vertex-ai

Scaleway Generative APIs

EU/GDPR-compliant inference hosted in Paris, France. Privacy-first — provider does not log or train on inputs/outputs.

Free tier:

  • 1,000,000 free tokens for every new customer (no time limit advertised)
  • Models: Qwen3 (235B / 397B / coder-30B), Llama 3.3 70B, Mistral Small 3.2 24B, DeepSeek R1 distill, Pixtral, Gemma, embeddings
  • Higher rate limits unlocked after KYC + payment method on file
  • Beyond free tokens, paid pricing typically €0.20–€0.90 per 1M tokens

Source: https://www.scaleway.com/en/generative-apis/, https://www.scaleway.com/en/pricing/model-as-a-service/ 15

Homepage: https://console.scaleway.com

Kilo Code — Gateway

Open-source agentic coding extension for VS Code, JetBrains, and CLI. Its built-in Kilo Gateway routes LLM requests to any provider and ships with a genuine free path — no subscription required.

Free access:

  • Free models (IDs ending in :free) cost nothing — usage is tracked but not billed, rate-limited to 200 requests/hour per IP
  • kilo-auto/free auto-routes among available free models (e.g. GLM 4.7, MiniMax M2.1)
  • Signup grants $5 in free credits (one-time) usable on paid models
  • Bring-your-own-keys works for any provider — no Kilo subscription required

Paid Kilo Pass tiers are available for higher throughput on premium models (Starter $19, Pro $49, Expert $199/month), but the free path covers most casual coding use.

Source: https://kilo.ai/pricing, https://kilo.ai/docs/getting-started/using-kilo-for-free, https://kilo.ai/docs/gateway/usage-and-billing 16

Homepage: https://kilo.ai

Kiro (by AWS)

AWS's AI coding IDE/agent — backed by Claude Sonnet/Haiku/Opus and other frontier models. Free tier requires only an AWS Builder ID, social login, or AWS IAM Identity Center sign-in.

Free tier:

  • 50 credits/month (perpetual)
  • 500 bonus credits for new signups (usable within 30 days)
  • Default Auto agent picks among Sonnet 4.5 and other models; manual selection includes Sonnet 4 / 4.5, Haiku 4.5, Opus 4.5 / 4.6 / 4.7
  • Other models (GLM, MiniMax, Qwen, DeepSeek, Kimi) available regionally
  • Unused credits don't roll over; overage on paid tiers is $0.04/credit

Paid tiers: Pro $20, Pro+ $40, Power $200 — with 1,000–10,000 credits/month.

Source: https://kiro.dev/pricing/ 17

Homepage: https://kiro.dev

Pollinations AI

Open-source Gen-AI platform (Berlin) for text, image, audio, and video generation. OpenAI-compatible endpoints.

Free access (post-2026 key migration):

  • Publishable key (free, beta): 1 pollen/IP/hour — for client-side, demos, prototypes
  • Secret key: server-side only, no rate limit listed (still free during beta)
  • Sign up at https://enter.pollinations.ai; ~$1 ≈ 1 Pollen for paid pay-as-you-go
  • Models: DeepSeek V4 Flash/Pro, Flux, GPT Image, Seedream, Whisper, ElevenLabs voices, Veo (alpha)

Source: https://github.com/pollinations/pollinations 18

Homepage: https://pollinations.ai


  1. Checked on Mar 25, 2026 ↩︎

  2. Checked on Apr 25, 2026 ↩︎

  3. Checked on Apr 22, 2026 ↩︎

  4. Checked on Apr 22, 2026 ↩︎

  5. Checked on Apr 22, 2026 ↩︎

  6. Checked on Apr 22, 2026 ↩︎

  7. Checked on Apr 22, 2026 ↩︎

  8. Checked on Apr 28, 2026 ↩︎

  9. Checked on Apr 28, 2026 ↩︎

  10. Checked on Mar 25, 2026 ↩︎

  11. Checked on Apr 28, 2026 ↩︎

  12. Checked on Apr 25, 2026 ↩︎

  13. Checked on Apr 28, 2026 ↩︎

  14. Checked on Apr 28, 2026 ↩︎

  15. Checked on Apr 28, 2026 ↩︎

  16. Checked on Apr 22, 2026 ↩︎

  17. Checked on Apr 28, 2026 ↩︎

  18. Checked on Apr 28, 2026 ↩︎

S
Description
Curated list of affordable + free LLM API providers — coding-plan subs, free-tier APIs.
Readme Apache-2.0
459 KiB
0 Stars 1 Watchers 0 Forks
Languages
HTML 88.2%
Ruby 11.8%