diff --git a/CLAUDE.md b/CLAUDE.md index e78032c..ed20dc4 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -13,10 +13,11 @@ When adding a new provider to README.md, include: - Pricing/plan details (if applicable) - Free tier limits or usage restrictions - Source link for verification -- Footnote with the date information was verified (format: `[^providername]: Check at MMM DD YYYY`) +- A closing line with the date information was verified (format: `*Checked MMM DD, YYYY.*`) ## Structure README.md is organized into sections: -- **Providers support coding plans**: Subscription-based coding plans -- **Free providers**: Providers with free tiers or trials +- **Claude Code Guest Passes** and **Claude AI Ecosystem**: Claude-specific perks and referrals +- **Providers with Coding Plans**: Subscription-based coding plans +- **Free Providers**: Providers with free tiers or trials (card-gated trial credits are included and say so) diff --git a/README.md b/README.md index 90e83f7..1990e4f 100644 --- a/README.md +++ b/README.md @@ -12,8 +12,8 @@ update an entry. ## Claude Code Guest Passes -One week of free Claude Pro (includes Claude Code). New users only; requires payment -info but can be cancelled before the trial ends. +One week of free Claude Pro (includes Claude Code). Eligible subscribers get passes via the `/passes` +command in Claude Code. New users only; requires payment info but can be cancelled before the trial ends. **My passes:** @@ -21,176 +21,263 @@ info but can be cancelled before the trial ends. Have a spare pass? Open a PR adding your link, or open an issue. +*Checked Sep 29, 2026.* + ## Claude AI Ecosystem ### [AgentKit](https://agentkit.best/) AgentKit (formerly **ClaudeKit**, `claudekit.cc`) sells production-ready kits of skills, slash -commands, subagents, and workflows for coding agents — Claude Code, Codex, GitHub Copilot, and -others. **Engineer** ($99) covers frontend, backend, database, DevOps, code review, and debugging; -**Marketing** ($99) adds MCP integrations and subagents for lead research, SEO, and outreach; -the bundle is **$149** (108+ skills, 95+ commands, 45 subagents). A macOS/Windows desktop app for -the `ak` CLI is on a waitlist. +commands, subagents, and workflows for coding agents — native support for Claude Code, Codex, Antigravity, +Pi, and Oh My Pi; Cursor and DeepSeek Harness in beta; OpenCode, Grok, and GitHub Copilot in preview. +**Engineer** ($99) covers frontend, backend, database, DevOps, code review, and debugging; **Marketing** ($99) +adds research, SEO, competitor-intelligence, and copywriting agents; the bundle is **$149** (108+ skills, +95+ commands, 45 subagents). The **AgentKit App** desktop cockpit (macOS/Windows) is sold separately at +**$49/yr** (1 device) or **$99/yr** (3 devices). > Referral: **20% off** your first purchase via (code: `BWA910UK`). -*Checked Aug 28, 2026.* +*Checked Sep 29, 2026.* --- ## Providers with Coding Plans -Monthly subscriptions with request-based (not token-based) quotas, compatible with -Claude Code, Cursor, Cline, and similar tools. +Monthly subscriptions with a request, credit, or allowance quota instead of pure pay-per-token billing, +compatible with Claude Code, Cursor, Cline, and similar tools. ### [Z.ai](https://z.ai/subscribe) -Plans from **$18/month** (Lite/Pro/Max). OpenAI-compatible + Anthropic-compatible endpoint. -Models: GLM-5.1, GLM-5-Turbo, GLM-4.7, GLM-4.5-Air. +GLM Coding Plan, credit-based since Jul 30, 2026. **Lite $18 / Pro $80 / Max $168 per month** (quarterly −20%, +yearly −30%). Credits per 5 h / per week: Lite 2,000 / 10,000 · Pro 12,000 / 60,000 · Max 28,000 / 140,000; +off-peak usage (outside Mon–Fri 14:00–18:00 UTC+8) costs 50% fewer credits. + +Models: GLM-5.3, GLM-5.3-Flash (GLM-5.2/5.1 requests auto-route to GLM-5.3, GLM-4.7 to GLM-5.3-Flash). Includes +Vision, Web Search, Web Reader, and Zread MCP. +Endpoints: Anthropic `https://api.z.ai/api/anthropic`, OpenAI `https://api.z.ai/api/coding/paas/v4`. +Tools: Claude Code, Cursor, Cline, Roo Code, Kilo Code, OpenCode, OpenClaw, Crush, Goose (supported tools only). > Referral: -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* ### [MiniMax](https://platform.minimax.io) -Token Plan — priced per **API call**, not token. $20–$120/month. +Token Plan — usage-based deduction from one shared quota (5-hour rolling + weekly windows). $22–$132/month. -- Plus ($20): 4-5 agents, all models on the API platform -- Max ($50): 6-7 agents, all models on the API platform -- Ultra ($120): 6-7 agents, all models on the API platform +- Plus ($22): 3-4 agents +- Max ($55): 4-5 agents +- Ultra ($132): 6-7 agents -Models: M3 (frontier multimodal coding, 1M context), M2.7 (language), M2.7-highspeed, speech-2.8-hd/turbo, Music-2.6, Hailuo 2.3 (video). Yearly plans ~17% off. +Covers the full MiniMax lineup (M3 / M2.7 / image / speech); MiniMax H3 video, voice design, and rapid voice +cloning are excluded. Top-up Credits: 1,000 credits = $1, valid 365 days. +OpenAI- and Anthropic-compatible. Tools: Claude Code, Codex, Cursor, TRAE, Hermes Agent, OpenClaw, Pi. -> Referral (10% off) until **Jul 1, 2026** — **For Referred Users:** 10% off subscription + become a dev ambassador. **For Referrers:** 10% back in API voucher per paid referral, usable across all MiniMax models, plus priority access to events and model previews. [View details](https://platform.minimax.io/subscribe/token-plan?code=CAQ5sxHAq6&source=link) +> Referral (10% off) — **For Referred Users:** 10% off subscription + become a dev ambassador. **For Referrers:** 10% back in API voucher per paid referral, usable across all MiniMax models, plus priority access to events and model previews. [View details](https://platform.minimax.io/subscribe/token-plan?code=CAQ5sxHAq6&source=link) -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* ### [Kimi Code](https://www.kimi.com/code) -Moonshot's coding perk bundled with Kimi membership. Models: Kimi K-series. Rolling 5-hour quota window. +Moonshot's coding perk bundled with Kimi membership (desktop app, CLI, VS Code; Claude Code, OpenCode, Codex, +Hermes Agent via API key). Plans: Go ¥49 (no Kimi Code), **Plus ¥99**, **Pro ¥199**, **Max ¥699** per month; +annual billing saves up to ¥1,680. Rolling 5-hour window plus a monthly total (weekly cap removed for new members). -Tiers: Adagio (free), Andante, Presto. Pay-as-you-go also at `platform.moonshot.ai`. +Models: `k3` (K3, up to 1M context on Pro+), `k3-256k`, `kimi-for-coding` (K2.8 Preview), `kimi-for-coding-highspeed` +(K2.7 Code, Pro+). OpenAI- and Anthropic-compatible (`https://api.kimi.ai/coding/`). Pay-as-you-go also at +`platform.moonshot.ai`. > Referral (code: `C8CJ6F`) — sign up or subscribe via my link and we each get a guaranteed benefit, up to **1-Year Membership Credits**: > - Sign up: > - Subscribe: -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* ### Alibaba Cloud Model Studio > Referral: Up to **$1,700** in free trial credits via (code: `A92LU5`). -#### [Token Plan (Team Edition)](https://www.alibabacloud.com/help/en/model-studio/token-plan-overview) +#### [Token Plan](https://www.alibabacloud.com/help/en/model-studio/token-plan-overview) -Credit-based, per-seat subscription (Singapore region only): -- Standard: **$30/seat/month** — 25,000 Credits/seat -- Advanced: **$100/seat/month** — 100,000 Credits/seat -- Premium: **$200/seat/month** — 250,000 Credits/seat -- Shared quota pack: **$700** — 625,000 Credits (1-month validity; unused Credits expire) +Credit-based subscription (Singapore region only). Tools: Claude Code, Cursor, Qwen Code, Codex, Qoder, OpenClaw. -Models: qwen3.6-plus, glm-5, MiniMax-M2.5, deepseek-v3.2 (text); qwen-image-2.0/2.0-pro, wan2.7-image/image-pro (image). +**Personal Edition** (monthly quota, limited-time prices): +- Lite: **$6/month** (list $8) — 11,500 Credits +- Essential: **$10/month** (list $16) — 25,500 Credits +- Standard: **$18/month** (list $25) — 45,000 Credits +- Pro: **$68/month** (list $80) — 180,000 Credits +- Extra bundle: $15 — 20,000 Credits (up to 5, needs an active plan) + +**Team Edition** (per seat, no training on your data): +- Standard: **$20/seat/month** (list $30) — 25,000 Credits +- Pro: **$75/seat/month** (list $100) — 100,000 Credits +- Max: **$200/seat/month** — 250,000 Credits +- Shared quota pack: **$700** — 625,000 Credits + +Models (Personal): auto, qwen3.8-max, qwen3.8-flash, qwen3.7-max, qwen3.7-plus, qwen3.6-flash, deepseek-v4.1-flash, +deepseek-v4-pro, glm-5.3, glm-5.2, plus image (qwen-image-3.0-pro, wan2.7-image/-pro), audio, and HappyHorse video. +Team Edition adds kimi-k2.7-code/k2.6/k2.5, glm-5.1/5, MiniMax-M2.5, deepseek-v3.2. #### [Coding Plan](https://www.alibabacloud.com/help/en/model-studio/coding-plan) Pro plan: **$50/month** — 6,000 req/5-hour, 45,000 req/week, 90,000 req/month. -Models: qwen3.5-plus, kimi-k2.5, glm-5, MiniMax-M2.5. -Tools: Claude Code, Cursor, Cline, Codex, and more. Lite plan no longer accepting new subscribers. +Models: qwen3.7-plus, qwen3.6-plus, kimi-k2.5, glm-5, MiniMax-M2.5 (recommended); qwen3.5-plus, qwen3-max-2026-01-23, +qwen3-coder-next, qwen3-coder-plus, glm-4.7. +Endpoints: OpenAI `https://coding-intl.dashscope.aliyuncs.com/v1`, Anthropic `.../apps/anthropic`. +Tools: Claude Code, Cursor, Cline, Codex, OpenCode, Qwen Code, Qoder, Kilo CLI, OpenClaw, Hermes Agent, and more. +Lite closed to new subscribers (Mar 20, 2026) and to renewals/upgrades (Apr 13, 2026). -**Note (as of Jun 5, 2026):** effectively unbuyable — perpetually out of stock. Claimed restock at 00:00 GMT+8, but repeated attempts still fail to purchase. +**Note:** limited slots, restocked daily at 00:00 UTC+8 (first come, first served); Alibaba now recommends Token Plan +instead. As of Jun 5, 2026 it was effectively unbuyable (purchase not re-tested Sep 2026). -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* ### [opencode — Go](https://opencode.ai/go) -Open-source `opencode` CLI subscription. $5 first month, **$10/month** thereafter. +OpenCode subscription for curated open models. **Go $10/month**, **Go Plus $40/month** (higher limits). Limits are +monthly dollar amounts per model (5-hour = 20%, weekly = 50% of the monthly limit); e.g. Go allows ~220 GLM-5.3 or +~3,200 MiniMax M3 requests per 5 h. Falls back to Zen balance when enabled. -Models: GLM-5/5.1, Kimi K2.5/K2.6, MiniMax M2.7/M3, Qwen3.5/3.6/3.7 Plus, MiMo-V2.5(-Pro), DeepSeek V4 Flash/Pro. Per-5-hour limits vary (200–10,200 req). -API key portable — works with Claude Code via LiteLLM proxy or `oc-go-cc`. Model format: `opencode-go/`. +Models: GLM-5.3/5.3-Flash/5.2, Kimi K3/K2.7 Code/K2.6, MiniMax M3/M2.7, Qwen3.8 Max/Flash, Qwen3.7 Plus, +DeepSeek V4.1 Flash/V4 Pro/V4 Flash, MiMo-V2.6(-Pro/-Flash)/V2.5(-Pro), LongCat-2.0, Hy4 preview, Hy3, Grok 4.7/4.6, +GPT 6 Luna/5.6 Luna, plus limited-time free models. +Endpoints: `https://opencode.ai/zen/go/v1/{chat/completions,messages,responses}`. Claude Code works natively via the +Anthropic endpoint; also validated with Codex, Hermes, ZCode, Pi, jcode, Kilo Code CLI. Model format: `opencode-go/`. -My referral: +My referral (**program ended** — opencode's referral page now says links no longer earn credit for either side): -> Invite friends to OpenCode Go. Earn $5 when a friend subscribes, and they'll get $5 too. Share your referral link; your friend joins and subscribes to Go; you both get a $5 usage credit to apply toward your Go usage limits. +> ~~Invite friends to OpenCode Go. Earn $5 when a friend subscribes, and they'll get $5 too. Share your referral link; your friend joins and subscribes to Go; you both get a $5 usage credit to apply toward your Go usage limits.~~ > -> Referral link: +> ~~Referral link: ~~ -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* ### [Synthetic](https://synthetic.new/) -Privacy-first inference (no training on prompts/responses). **$30/month** subscription or pay-as-you-go. +Privacy-first inference (no training on prompts/responses). **$30/month per pack** ($1/day) — 500 requests/5 h, +1 concurrent request per model (buy more packs to raise both), or usage-based pay-per-token. -Models: Kimi K2.6, MiniMax M2.5, GLM 5.1, GLM 4.7 Flash, vLLM-compatible open-source models. -OpenAI-compatible — works with Roo, Cline, Octofriend. +Models: Kimi-K3, DeepSeek-V4.1-Flash (beta), GLM-5.3-Flash, GLM-4.7-Flash, Qwen3.8-27B, gpt-oss-120b, +NVIDIA Nemotron-3-Super-120B; nomic-embed-text-v1.5 embeddings included. +OpenAI-compatible (`https://api.synthetic.new/openai/v1`) and Anthropic-compatible +(`https://api.synthetic.new/anthropic/v1`) — guides for Claude Code, Crush, OpenCode, GitHub Copilot, OpenClaw, +Xcode, Roo, KiloCode, Octofriend. > Referral: **$10.00** in subscription credit via -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* ### [BigModel.cn — GLM Coding Plan](https://www.bigmodel.cn/glm-coding) The Chinese (mainland) counterpart of Z.ai's GLM Coding Plan — same underlying Zhipu AI models, but billed in CNY through bigmodel.cn. Suited for users who can pay via Alipay / WeChat Pay or already have a 智谱 AI account. -Plans (monthly, after the 2026 price adjustment): -- **Lite**: ¥49/month -- **Pro**: ¥149/month -- **Max**: ¥469/month +Credit-based plans since Jul 30, 2026 (monthly list price, per media reports — the official price page needs +JavaScript): **Lite ¥118**, **Pro ¥538**, **Max ¥1,078**. +Credits per 5 h / per week: Lite 2,000 / 10,000 · Pro 12,000 / 60,000 · Max 28,000 / 140,000; off-peak usage +(outside Mon–Fri 14:00–18:00 UTC+8) costs 50% fewer credits. Legacy V1/V2 subscribers keep ¥49 / ¥149 / ¥469. -All tiers support GLM-5.1, GLM-5-Turbo, GLM-4.7, and GLM-4.5-Air. Compatible with Claude Code, Cline, and 20+ coding tools via OpenAI-compatible API plus an Anthropic-compatible endpoint. +All tiers support GLM-5.3 and GLM-5.3-Flash (GLM-5.2/5.1 route to GLM-5.3; GLM-5-Turbo/4.7 route to GLM-5.3-Flash). +Anthropic endpoint `https://open.bigmodel.cn/api/anthropic`, OpenAI endpoint `https://open.bigmodel.cn/api/coding/paas/v4`. +Tools: Claude Code, Kilo Code, OpenClaw (lower priority), OpenCode, TRAE, CodeBuddy, and others on the supported list. -Referral program (challenge-based, resets every 30 invitees): +Referral program (challenge-based, resets every 30 invitees; terms as last seen, not re-verified Sep 2026): - Invited friend gets **5% off** their first GLM Coding Plan order. - Referrer gets **10% cashback** once 3 friends subscribe, plus an **additional 10%** of the total paid amount for every 30 invitees. - Rebate credit is usable for resource packs, API calls, and subscription renewals on the BigModel platform. -Source: , [^bigmodel] - -Homepage: +Source: , My referral: >🚀 Join the GLM Coding Plan via my link — get 5% off your first order. Subscribe at https://www.bigmodel.cn/glm-coding?ic=VGRZKHKNKW (invitation code: `VGRZKHKNKW`). -[^bigmodel]: Checked on Jun 5, 2026 +*Checked Sep 29, 2026.* ### [BytePlus ModelArk — Coding Plan](https://www.byteplus.com/en/activity/codingplan) -ByteDance. Lite: **$15/month**, Pro: **$35/month** (intro promo $5/$25 ended early 2026). +ByteDance. Lite: **$10/month** ($30/quarter), Pro: **$50/month** ($150/quarter). New-user first-purchase promo +($5/$25) suspended since Mar 17, 2026. +Limits: Lite ~1,900 req/5 h, ~12,000/week, ~24,000/month; Pro 5× Lite (~9,500 / ~60,000 / ~120,000). -Models: ByteDance-Seed-2.0, DeepSeek-V3.2, GLM-5.1, Kimi-K2.5. -Tools: Claude Code, Cursor, Cline, Roo Code, OpenCode. +Models: Auto, Dola-Seed-2.0-Pro/Lite/Code, ByteDance-Seed-Code, GLM-5.3-Flash, GLM-5.2, GLM-5.1, Kimi-K2.5, +DeepSeek-V4.1-Flash, DeepSeek-V4-Pro/Flash, GPT-OSS-120b. +Endpoints: OpenAI `https://ark.ap-southeast.bytepluses.com/api/coding/v3`, Anthropic `.../api/coding`. +Tools: Claude Code, Cursor, Cline, Codex, Roo Code, Kilo Code, OpenCode, OpenClaw, TraeCode, Hermes Agent. -> Referral: +> Referral: — campaign (10% off first +> order) is scheduled to end **Sep 30, 2026**. -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* + +### [Volcengine Ark — Coding Plan](https://www.volcengine.com/activity/codingplan) + +ByteDance's mainland-China sibling of the BytePlus plan (火山引擎方舟 Coding Plan), billed in CNY. Lite **¥40/month**, +Pro **¥200/month** (per Volcengine's own articles; the live price widget advertises "limited-time from ¥9.9" and +needs JavaScript). Lite ~1,200 req/5 h, 9,000/week, 18,000/month; Pro 5× Lite. + +Models: DeepSeek-V4.1-Flash, GLM-5.3 series, Doubao-Seed-Evolving, Kimi-K3, Kimi-K2.8-Preview. Tools: Claude Code, +Cursor, and others. Invite program: 5% voucher for the referrer, 5% off for the friend. + +Source: , + +*Checked Sep 29, 2026.* + +### [Atlas Cloud — Coding Plan](https://www.atlascloud.ai/coding-plan) + +Third-party aggregator plan with **full API access** (rare among coding plans). Starter **$10**, Lite **$20**, Plus **$50**, +Max **$100** per month — 16.5M / 33M / 82.5M / 165M points per week. + +Models: 18 LLMs incl. DeepSeek, GLM, Kimi, MiniMax. Tools: Claude Code, Codex, Cursor, OpenClaw. + +Source: + +*Checked Sep 29, 2026.* + +### [NanoGPT](https://nano-gpt.com/pricing) + +Pay-as-you-go gateway for every major model, plus an optional **Pro subscription: $12/month — 60 million included input +tokens per week** on subscription models (web + API), and 5% off eligible paid text models. + +Subscription models include GLM-5.3 / 5.3-Flash, Kimi K2.6 / K2.7 Code, MiniMax M3 / M2.7, DeepSeek V4 Flash / V4 Pro, +MiMo-V2.5(-Pro), Qwen3.8-27B, Nemotron 3 Ultra. OpenAI-compatible at `https://api.nano-gpt.com/api/v1` +(also `/messages` and `/responses`); use `https://api.nano-gpt.com/api/subscription/v1` to keep requests on the +subscription only. Pay-as-you-go deposits start at $1 (card). + +Source: , + +*Checked Sep 29, 2026.* ### [Xiaomi MiMo Open Platform](https://platform.xiaomimimo.com) -I'm on Xiaomi MiMo Open Platform — running Xiaomi's flagship MiMo V2.5 and the rest of the lineup. Sign up with my code and you'll instantly get $2 in API credits. +I'm on Xiaomi MiMo Open Platform — running Xiaomi's flagship MiMo V2.6 and the rest of the lineup. Sign up with my code and you'll instantly get $2 in API credits. After signup, enter the code at the bottom-left of the console. Credits valid 40 days. -**Token Plan** (monthly, launched May 2026): Lite ¥39 (60M credits), Standard ¥99 (200M), Pro ¥329 (700M), Max ¥659 (1,600M). Models: MiMo-V2.5, MiMo-V2.5-Pro (2× credit cost). Annual plans discounted. +**Token Plan** (monthly): Lite $6 / ¥39 (4.1B credits), Standard $16 / ¥99 (11B), Pro $50 / ¥329 (38B), +Max $100 / ¥659 (82B). Annual plans 12% off; 12% off first Individual purchase; 0.8× consumption 00:00–08:00 +Beijing time. Team Edition from $16/seat. Models: mimo-v2.6-pro, mimo-v2.6-flash, ASR/TTS (mimo-v2.5 and +mimo-v2.5-pro retire Oct 21, 2026). Works with OpenCode, OpenClaw, Claude Code. > Referral: Code `T8ESAY` · -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* ### [GitHub Copilot Pro](https://github.com/features/copilot/plans) -Cheapest mainstream coding seat. **$10/month** ($100/year) — unmetered code completions and -chat on the base model, plus **1,500 premium requests/month** for frontier models -(Claude, GPT, Gemini). Copilot Pro+ is $39/month with 6,000 premium requests. +Cheapest mainstream coding seat. **$10/month** — unlimited code completions and next-edit suggestions, plus +**1,500 GitHub AI Credits/month** (1,000 base + 500 flex; 1 credit = $0.01) for chat, agent mode, code review, +cloud agent, and CLI. Pro models include Claude Haiku 4.5, Sonnet 4.6/5/5.5, GPT-5.4/5.6 Luna/6 Luna, Gemini 3.x +Flash, Grok 4.x, Kimi K3 — Opus/Fable and GPT-6 Sol/Astra need Pro+ or Max. +Copilot Pro+ is $39/month (7,000 credits); Copilot Max is $100/month (20,000 credits). Extra usage billed at $0.01/credit. -Free tier: 2,000 completions + 50 premium requests/month, no card. -Free for verified students, teachers, and maintainers of popular open-source repos. +Free tier: 2,000 completions/month plus a small AI-credit allowance, auto model selection only, no card. +Free Copilot Student for verified students; free Pro for verified teachers and maintainers of popular open-source repos. -Tools: VS Code, JetBrains, Neovim, Xcode, Visual Studio, `gh copilot` CLI, and the Copilot -coding agent on github.com. +Tools: VS Code, JetBrains, Neovim, Xcode, Visual Studio, Copilot CLI, Copilot app, and the Copilot cloud agent on +github.com. -*Checked Sep 11, 2026.* +*Checked Sep 29, 2026.* --- @@ -198,437 +285,467 @@ coding agent on github.com. ### [TokenRouter](https://www.tokenrouter.com/) -Unified AI gateway with OpenAI-compatible access to 300+ models. +Unified AI gateway (145 models listed) exposing OpenAI-, Anthropic- and Gemini-format APIs behind one key. -- **Kimi K3 Free Version:** Use `moonshotai/kimi-k3-free` at $0 until **Aug 12, 2026**, running on TokenRouter's own B300 deployment. -- **API:** Use the OpenAI Chat Completions endpoint at `https://api.tokenrouter.com/v1` with a TokenRouter API key. -- **Limit:** Free compute capacity is limited, so stability and concurrency are not guaranteed. +- **Free model:** `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free` at $0/M input and output. +- **Launch promos:** TokenRouter runs short free windows on self-hosted launches (e.g. GLM 5.3 was free for all accounts, no card, until Sep 4, 2026). Watch the [blog](https://www.tokenrouter.com/blog). +- **API:** OpenAI Chat Completions at `https://api.tokenrouter.com/v1` with a TokenRouter API key. -Source: [TokenRouter models](https://www.tokenrouter.com/models) [^tokenrouter] +Source: [TokenRouter models](https://www.tokenrouter.com/models) -[^tokenrouter]: Check at Aug 5 2026 +*Checked Sep 29, 2026.* ### [OrcaRouter](https://www.orcarouter.ai/) -OpenAI-compatible gateway with access to 200+ models, automatic routing, and failover. +OpenAI-compatible gateway with access to 200+ models, automatic routing, and failover (also accepts Anthropic and Gemini formats). > Referral: (code: `ref_3976ba42abf37dc55c1d`). -- **Kimi K3:** Get $5 in credit as a new user before **Aug 6, 2026**; a payment card is required. -- **Tencent HY3:** Get $5 in credit before **Aug 21, 2026**. -- **Claude Opus 5:** Get a 60% deposit match, up to $500, when you join before **Aug 24, 2026**. -- **Free DeepSeek models:** Get 100 initial V4 Flash calls and 30 initial V4 Pro calls with no published deadline. -- **Grok 4.5:** The offer is sold out, but its waitlist is still open. +- **Always-free models ($0):** `deepseek/deepseek-v4-flash-free`, `z-ai/glm-5.3-flash-free`, `tencent/hy4-preview-free`, `tencent/hy3-free`, plus the `orcarouter/free` router. Requires a linked GitHub account with some history, or any paid purchase. +- **Hy4 Deposit Match:** 100% match on top-ups for `tencent/hy4-preview` (min $20 top-up, cap ≈$200 derived from API quota units), enroll by **Oct 28, 2026**. +- **DeepSeek V4.1 Flash Deposit Match:** 30% match (min $20 top-up, up to $100), enroll by **Oct 2, 2026**. +- **GPT-6 Astra Deposit Match:** 100% match for `openai/gpt-6-astra` (min $20, up to $100), card required, enroll by **Oct 5, 2026**. +- **API:** `https://api.orcarouter.ai/v1`. -Offers can change quickly; check the [live offers page](https://www.orcarouter.ai/offers) before claiming. [^orcarouter] +Offers can change quickly; check the [live offers page](https://www.orcarouter.ai/offers) before claiming. -[^orcarouter]: Check at Aug 5 2026 +*Checked Sep 29, 2026.* ### [OpenRouter](https://openrouter.ai) -Free models (`:free` suffix): 20 RPM, 50 req/day (free accounts); 1,000 req/day after $10 top-up. +Free models (`:free` suffix): 20 RPM, 50 req/day; 1,000 req/day once you have bought at least $10 in credits (all time). BYOK requests are not gated by the free-model cap. -*Checked Jun 5, 2026.* +Current free lineup includes NVIDIA Nemotron 3 Ultra 550B / Super 120B / 3.5 Lightning, Google Gemma 4 (26B-A4B, 31B), Qwen3.8-27B, Poolside Laguna S/XS 2.1, Thinking Machines Inkling / Inkling Small, Cohere North Mini Code, plus the `openrouter/free` auto-router. OpenAI-compatible at `https://openrouter.ai/api/v1`. + +*Checked Sep 29, 2026.* ### [NVIDIA NIM](https://build.nvidia.com) -Free for NVIDIA Developer Program members, no credit card. ~40 RPM. 1,000 inference credits at signup (consumption-based, not a fixed monthly request cap). +Free prototyping endpoints for NVIDIA Developer Program members, no credit card. Rate-limited per model (commonly ~40 RPM per community reports); NVIDIA does not publish a fixed quota and does not raise free-tier limits on request. -Models: Kimi K2.5, GPT-OSS, DeepSeek-V3.2, Llama 3.x, Mistral, Phi, Nemotron. OpenAI-compatible. +Models: Kimi K3, Kimi K2.6, DeepSeek V4.1 Flash, GLM 5.3 / 5.3 Flash, Nemotron 3 Ultra / Super / 3.5 Lightning, GPT-OSS-20B, Gemma 4 31B, Mistral Large. OpenAI-compatible at `https://integrate.api.nvidia.com/v1`. -*Checked Jun 5, 2026.* +**Warning:** trial terms say use is logged and may be used to improve NVIDIA products — do not send personal or confidential data. + +*Checked Sep 29, 2026.* ### [OpenCode Zen](https://opencode.ai/docs/zen) -Hand-picked free models that change periodically. Optimized for coding agents. +Hand-picked free models that change periodically (each is "free for a limited time"). Optimized for coding agents. -| Model | Status | -|---|---| -| `Big Pickle` | Free | -| `DeepSeek V4 Flash Free` | Free | -| `MiMo-V2.5 Free` | Free | -| `Nemotron 3 Ultra Free` | Free | +| Model | Model ID | Notes | +|---|---|---| +| Big Pickle | `big-pickle` | Stealth model; data may be used to improve it | +| Space Bunny Free | `space-bunny-free` | Stealth model; zero-retention provider | +| LongCat 2.5 Preview Free | `longcat-2.5-preview-free` | Zero-retention provider | +| MiMo-V2.6-Flash Free | `mimo-v2.6-flash-free` | Data may be used to improve the model | +| MiMo-V2.5 Free | `mimo-v2.5-free` | Data may be used to improve the model | +| Ling 3.0 Flash Fin Free | `ling-3.0-flash-fin-free` | Data may be used to improve the model | +| Nemotron 3 Ultra Free | `nemotron-3-ultra-free` | NVIDIA trial endpoint; logged | +| Nemotron 3.5 Lightning Free | `nemotron-3.5-lightning-free` | NVIDIA trial endpoint; logged | +| Muse Spark 1.3 Contributor Free | `muse-spark-1.3-contributor-free` | Prompts used to train Meta models | -OpenAI/Anthropic compatible endpoints at `https://opencode.ai/zen/v1/`. +OpenAI-compatible at `https://opencode.ai/zen/v1/chat/completions` (plus `/v1/responses`); Anthropic-format models use `https://opencode.ai/zen/v1/messages`. Model format in opencode: `opencode/`. -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* ### [Google AI Studio](https://aistudio.google.com/) -Google's developer platform for Gemini models. Generous free tier with pay-as-you-go available. +Google's developer platform for Gemini models. Free tier (no billing) with pay-as-you-go available. -Free-tier limits vary by model (check the AI Studio dashboard for your project): +Free-tier models: Gemini 3.8 / 3.7 / 3.6 / 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.1 Flash-Lite, Gemini 3 Flash Preview, plus Live/TTS variants and Gemma 4. **Gemini 3.1 Pro Preview is paid-only.** Gemini 2.5 models are now limited to projects that already used them. -| Model (Free) | RPM | RPD | -|---|---|---| -| Gemini 2.5 Pro | 5 | 100 | -| Gemini 2.5 Flash | 10 | 250 | -| Gemini 2.5 Flash-Lite | 15 | 1,000 | +Google no longer publishes a fixed free-tier table — limits are per project and shown in AI Studio. Third-party reports (Sep 2026): ~20 RPD on the 3.x Flash models, ~500 RPD on 3.5 / 3.1 Flash-Lite. RPD resets at midnight Pacific. -Paid tier: higher limits, usage-based billing. - -Models: Gemini 2.5 Pro/Flash/Flash-Lite. OpenAI-compatible endpoint available. +OpenAI-compatible endpoint: `https://generativelanguage.googleapis.com/v1beta/openai/`. **Warning:** In the Free tier, Google may use your prompts and responses to improve their products. Use the Paid tier or Vertex AI for privacy. -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* ### [Kilo Code — Gateway](https://kilo.ai/gateway) -VS Code + JetBrains coding extension with built-in gateway. +VS Code + JetBrains coding extension (and CLI) with a built-in OpenAI-compatible gateway at `https://api.kilo.ai/api/gateway`. Kilo was acquired by Anaconda. -Free: `:free` models, 200 req/hour per IP. First top-up: $20 bonus credits (60-day expiry). -BYOK supported — no Kilo subscription required for your own provider keys. +Free: `:free` models and the `kilo-auto/free` router, 200 req/hour per IP (anonymous or signed in). Current free models include Nemotron 3 Ultra / Super / 3.5 Lightning, Qwen3.8-27B, Laguna S/XS 2.1, Inkling Small, North Mini Code, and Step 3.7 Flash. +BYOK supported with no Kilo markup. Kilo Pass ($19/$49/$199 per month) adds up to 50% bonus credits. -*Checked Jun 5, 2026.* +**Warning:** Auto Free may route to providers that log prompts and use them for training (including NVIDIA trial endpoints). + +*Checked Sep 29, 2026.* ### [Cloudflare Workers AI](https://developers.cloudflare.com/workers-ai/) -10,000 Neurons/day (resets 00:00 UTC). ~150 LLM responses/day on Llama 3.3 70B. +10,000 Neurons/day free on both Workers Free and Paid (resets 00:00 UTC). Roughly ~49K output tokens/day on Llama 3.3 70B or ~147K on GPT-OSS-120B. -Models: Llama 3.3 70B, Gemma 3, Qwen2.5-Coder, DeepSeek R1, BGE embeddings, Whisper. +Models: GPT-OSS 120B/20B, Kimi K2.5, Llama 3.3 70B, Llama 4 Scout, Gemma 4 26B, Qwen3.8-27B, Nemotron 3 120B, GLM-4.7-Flash, Mistral Small 3.1, BGE embeddings, Whisper. Kimi K2.6/K2.7-Code, GLM 5.x and DeepSeek V4 require Workers Paid or AI Gateway credits. -*Checked Jun 5, 2026.* +OpenAI-compatible at `https://api.cloudflare.com/client/v4/accounts//ai/v1`. + +*Checked Sep 29, 2026.* ### [DeepSeek Platform](https://platform.deepseek.com) -5M free tokens at signup (no code needed). OpenAI + Anthropic compatible. +New accounts are widely reported to get 5M free tokens (granted balance, ~30 days, phone verification) — verify in your console. OpenAI (`https://api.deepseek.com`) + Anthropic (`https://api.deepseek.com/anthropic`) compatible. -PAYG: V4 Flash $0.14/M in, $0.28/M out; V4 Pro $0.435/M in, $0.87/M out (cached: $0.03/M). +PAYG (off-peak / peak, per 1M tokens): `deepseek-flash` (V4.1 Flash) $0.15/$0.30 in, $0.60/$1.20 out; `deepseek-v4-pro` $0.66/$1.32 in, $1.98/$3.96 out. Off-peak is half price — everything outside 01:00–04:00 and 06:00–10:00 UTC on weekdays. 1M context. -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* ### [Groq](https://console.groq.com) -LPU inference. Free tier, no credit card. +LPU inference. Free plan (no card reported; not restated in current docs). OpenAI-compatible at `https://api.groq.com/openai/v1`. | Model | RPM | RPD | TPM | TPD | |---|---|---|---|---| -| `llama-3.1-8b-instant` | 30 | 14,400 | 6K | 500K | -| `llama-3.3-70b-versatile` | 30 | 1,000 | 12K | 100K | -| `whisper-large-v3` | 20 | 2,000 | — | — | +| `openai/gpt-oss-120b` | 30 | 1K | 8K | 200K | +| `openai/gpt-oss-20b` | 30 | 1K | 8K | 200K | +| `qwen/qwen3.8-27b` | 30 | 1K | 8K | 200K | +| `whisper-large-v3` / `-turbo` | 20 | 2K | — | 7.2K audio-sec/hour | -*Checked Jun 5, 2026.* +Llama 3.1 8B and Llama 3.3 70B are now Enterprise-only (contact sales). -### [GitHub Models](https://github.com/marketplace/models) +*Checked Sep 29, 2026.* -Free for all GitHub accounts. OpenAI-compatible at `https://models.github.ai/inference`. +### [Mistral Studio](https://console.mistral.ai) -Covers OpenAI, Anthropic, Llama, Mistral, DeepSeek, Grok, Phi. Rate limits vary per model -(e.g. GPT-4o: 10 RPM/50 RPD; DeepSeek-R1: 15 RPM/150 RPD). Requires PAT with `models:read`. +Formerly "La Plateforme". Free plan (default for new accounts) includes **$10/month in API credits**, shared across Studio, the API, and Vibe Code; rate limits shown in the Admin Console. Enable pay-as-you-go to continue past the allowance. Pro ($14.99/month) includes $15/month in API credits (per the pricing page; verify). -*Checked Jun 5, 2026.* +Models (indicative): Mistral Medium 3.5, Mistral Large (2512), Mistral Small (2603), Devstral 2, Codestral, Ministral 3B/8B/14B, Voxtral, embeddings. OpenAI-compatible. -### [xAI Grok API](https://x.ai/api) +**Warning:** Model training on your data is opt-out, not opt-in. -$25 signup credits + $150/month via Data Sharing Program (eligible countries). - -Models: Grok 4, Grok 4.1 Fast (2M ctx), Grok Code Fast. OpenAI + Anthropic compatible. - -**Warning:** Data Sharing opt-in is irreversible and lets xAI train on your prompts. - -*Checked Jun 5, 2026.* - -### [Mistral La Plateforme](https://console.mistral.ai) - -Free Experiment plan — up to ~1B tokens/month, no credit card (phone verification only). - -Models: Mistral Large 3, Medium 3, Small 3.1, Codestral, Pixtral, embeddings. OpenAI-compatible. - -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* ### [Google Cloud Vertex AI](https://cloud.google.com/vertex-ai) -$300 free credits for 90 days (new GCP customers). Not a recurring free tier. +Now branded **Gemini Enterprise Agent Platform** (formerly Vertex AI). -Models: Gemini 3 Pro/Flash, Anthropic Claude on Vertex, DeepSeek, GLM, Qwen via MaaS. -Express Mode available without billing for limited evaluation quotas. +Free Trial: **$300** credit for 90 days (new GCP customers only; card verification). Not a recurring free tier. The $300 credit **cannot** pay for partner models offered as MaaS (e.g. Claude) or for Gemini API in AI Studio. -*Checked Jun 5, 2026.* +Express mode: `@gmail.com` accounts new to Google Cloud get a **90-day free tier with no billing info**, within express-mode quotas, on the APIs that support express mode (Gemini models). Separate from the $300 Free Trial. + +Models: Gemini 3.x (Pro / Flash / Flash-Lite); partner models (Claude and others) need a paid billing account. + +*Checked Sep 29, 2026.* ### [Hugging Face Inference Providers](https://huggingface.co/docs/inference-providers) -Routes across Together, Fireworks, Novita, Cerebras, Replicate, DeepInfra, Scaleway. +Router in front of 20 partners: Baseten, Cerebras, Cohere, DeepInfra, Fal AI, Featherless AI, Fireworks, Groq, HF Inference, Novita, Nscale, OVHcloud, Public AI, Replicate, Scaleway, Together, WaveSpeedAI, Z.ai and others. No markup over provider rates. -Free: 100K monthly credits. PRO ($9/month): 2M credits. OpenAI-compatible at `https://router.huggingface.co/v1`. +Free: **$0.10/month** in credits (subject to change). PRO (**$9/month**): **$2.00/month** in credits, usable across all HF compute. Pay-as-you-go beyond that requires buying credits. -*Checked Jun 5, 2026.* +OpenAI-compatible at `https://router.huggingface.co/v1` (~130 chat models). + +*Checked Sep 29, 2026.* ### [Cerebras Cloud](https://cloud.cerebras.ai) -Wafer-scale chip inference. Free tier, no credit card. +Wafer-scale chip inference. **Free Trial only**: **$5** in credits after adding a **verified payment method**, expiring 30 days after grant. No permanently free tier. OpenAI-compatible at `https://api.cerebras.ai/v1`. -| Model | RPM | TPD | -|---|---|---| -| `gpt-oss-120b` | 30 | — | -| `llama3.1-8b-instant` | 30 | 500K | -| `qwen-3-235b-a22b-instruct-2507` | 30 | — | -| `zai-glm-4.7` | 10 | — | +| Model | RPM | Uncached TPM | TPD | Context (free) | +|---|---|---|---|---| +| `gpt-oss-120b` | 5 | 30K | 1M | 65K | +| `qwen-3.8-27b` | 5 | 30K | 1M | 64K | -1M tokens/day shared cap. 8K context limit on free tier. +Total TPM (incl. cached) is 3x uncached (90K). Access stops when credits run out or expire until you buy credits. -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* ### [BigModel.cn](https://www.bigmodel.cn/) -Zhipu AI (智谱 AI). 25M free tokens for new users to explore the API, playground, and AGI apps. GLM-4.7-Flash and GLM-4.5-Flash are permanently free. +Zhipu AI (智谱 AI). New users get a free token package to explore the API, playground, and AGI apps (reported as 20M, 25M via invite). + +Permanently free models: **GLM-4.7-Flash** (200K ctx), GLM-4-Flash-250414 (128K), GLM-Z1-Flash (128K), GLM-4.6V-Flash, GLM-4.1V-Thinking-Flash, GLM-4V-Flash, plus CogView-3-Flash (image) and CogVideoX-Flash (video). Concurrency limits apply per model. GLM-5.3-Flash is paid on the pay-as-you-go API. + +OpenAI-compatible at `https://open.bigmodel.cn/api/paas/v4`; Anthropic-compatible at `https://open.bigmodel.cn/api/anthropic`. > Referral: -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* ### [Fireworks AI](https://fireworks.ai) -$1 starter credits. 50+ models. Function calling, MCP support. OpenAI-compatible. +**$1** in free starter credits for serverless inference. Without a payment method (or without credits) the account is capped at **10 RPM**; adding a payment method raises it up to 6,000 RPM. -*Checked Jun 5, 2026.* +Models: GLM 5.3 / 5.3 Flash, Kimi K3, DeepSeek V4 Flash, Qwen 3.8 27B and other open models. Function calling, MCP support. + +OpenAI-compatible at `https://api.fireworks.ai/inference/v1`; Anthropic-compatible at `https://api.fireworks.ai/inference` (works with Claude Code). + +*Checked Sep 29, 2026.* ### [Scaleway Generative APIs](https://www.scaleway.com/en/generative-apis/) -EU/GDPR, Paris. 1M free tokens for new customers (no time limit advertised). +EU/GDPR, Paris. Free tier for new customers: first **1,000,000 tokens** plus **60 minutes** of audio transcription (no time limit advertised). OpenAI-compatible at `https://api.scaleway.ai/v1`. -Models: Qwen3 (235B/397B/coder), Llama 3.3 70B, Mistral Small 3.2, DeepSeek R1 distill, Pixtral. +Models: GLM-5.2, DeepSeek V4 Flash, Qwen3.8-27B, Qwen3.5-397B, Qwen3.6-35B, Qwen3-235B, Qwen3-Coder-30B, Gemma 4 26B, Mistral Medium 3.5, Mistral Small 3.2, gpt-oss-120b, Llama 3.3 70B, Pixtral 12B, Whisper. -*Checked Jun 5, 2026.* +*Checked Sep 29, 2026.* ### [SambaNova Cloud](https://cloud.sambanova.ai) -RDU (dataflow chip) inference. Developer tier, no credit card. +RDU (dataflow chip) inference. **Free tier** applies when no payment method is linked: **20 RPM, 20 requests/day, 200K tokens/day** per model. Adding a card moves you to the Developer tier (pay-as-you-go, 60–240 RPM, 20M tokens/day across models). The $5 signup credit is no longer advertised, and the Plans page now says to add a payment method first — the two official pages disagree. -Free: **$5** in credits (expire after 30 days), 20 RPM, 200K tokens/day per model. +Free-tier models: DeepSeek-V3.1, DeepSeek-V3.2 (preview), Llama 3.3 70B, gpt-oss-120b, Gemma 4 31B (preview). -Models: DeepSeek, Llama 3.x/4, Qwen3, Whisper. OpenAI-compatible. +OpenAI-compatible and Anthropic-compatible at `https://api.sambanova.ai/v1`. -Source: +Source: -*Checked Sep 11, 2026.* +*Checked Sep 29, 2026.* ### [Cohere](https://dashboard.cohere.com/api-keys) -Trial API key, no credit card. **1,000 calls/month**, 20 RPM. +Trial API key, no credit card. **1,000 calls/month**, 20 RPM per chat model (Rerank 10 RPM, Embed 2,000 inputs/min). -Models: Command A+, Command R, Command R7B, Aya (11+ models), plus rerank and embeddings. +Models: Command A+, Command A Reasoning / Vision / Translate, Command A, Command R+, Command R, Command R7B, North Mini Code, Aya Expanse / Aya Vision, plus rerank and embeddings. -**Warning:** trial keys are **non-commercial use only** — production needs a production key (paid). +OpenAI-compatible at `https://api.cohere.ai/compatibility/v1`. -*Checked Sep 11, 2026.* +**Warning:** trial keys are **not permitted for production or commercial use** — production needs a production key (paid). New model variants (e.g. Command A Reasoning) stay at trial limits even on prod keys. + +*Checked Sep 29, 2026.* ### [Vercel AI Gateway](https://vercel.com/docs/ai-gateway) -Single OpenAI-compatible endpoint routing to many providers, with failover and BYOK. +Single endpoint routing to many providers, with failover and BYOK. OpenAI-compatible at `https://ai-gateway.vercel.sh/v1`; also Anthropic Messages, OpenResponses and Cohere-compatible APIs. -Free: **$5/month** in renewing credits (does not roll over). Free tier covers a subset of models and applies lower per-model rate limits than paid. +Free: a **monthly included credit** (reported as **$5 / 30 days**; does not roll over), starting with your first request. **Requires a valid payment method on the team.** Covers only free-tier-eligible models with lower per-model rate limits. Buying credits moves you to the paid tier permanently and the monthly free credit stops. Source: -*Checked Sep 11, 2026.* +*Checked Sep 29, 2026.* ### [Requesty](https://www.requesty.ai/) -LLM gateway with routing, caching, and spend controls. Works with Claude Code, Cline, Cursor, Roo. +LLM gateway with routing, caching, spend controls and EU data residency. Works with Claude Code, Cline, Cursor, Roo. -Free: **200 req/day** on the free-model catalogue. No credit card, no trial clock — same platform as pay-as-you-go, just restricted to free models until you upgrade. +Free: **200 req/day** on the free-model catalogue. No credit card, no trial clock — same platform as pay-as-you-go (600+ models, +5% fee), just restricted to free models until you upgrade. + +OpenAI-compatible at `https://router.requesty.ai/v1`; Claude Code via `ANTHROPIC_BASE_URL=https://router.requesty.ai`. Source: , -*Checked Sep 11, 2026.* +*Checked Sep 29, 2026.* ### [SiliconFlow](https://cloud.siliconflow.cn/) -Chinese multi-model inference platform, 200+ LLM/image/audio/video models. +Chinese multi-model inference platform, 200+ LLM/image/audio/video models. International site: . -Free: a set of smaller open-source models is **permanently $0**; new international accounts get ~**$1** starter credit. 1,000 RPM / 50,000 TPM on the free models. Identity verification required. +Free: a set of smaller open-source models is **permanently ¥0** with fixed rate limits (chat models from 1,000 RPM / 50K TPM); the international site gives **$1** starter credit. -Models (free): Qwen3-8B (128K ctx) and similar small open models. OpenAI-compatible. +Models (free, CN site): Qwen3-8B (128K), Qwen3.5-4B (256K), Qwen2.5-7B, GLM-4-9B-0414, GLM-Z1-9B-0414, DeepSeek-R1-0528-Qwen3-8B, Hunyuan-MT-7B, Xing4.0-29B, plus OCR, ASR and BGE embedding/rerank models. -*Checked Sep 11, 2026.* +OpenAI-compatible at `https://api.siliconflow.cn/v1`; Anthropic-compatible endpoint documented (Claude Code guide in docs). + +*Checked Sep 29, 2026.* ### [ModelScope](https://modelscope.cn/) Alibaba's model community (魔搭). API-Inference free for registered users. -Free: **2,000 req/day** total, **≤500 req/day per model**. Requires an Alibaba account with real-name verification (mainland China ID). +Free: **2,000 req/day** total, **≤500 req/day per model** (some large models lower; figures from secondary sources — the official limits page needs JavaScript). Requires binding an Alibaba Cloud account with real-name verification. Quotas may be adjusted at any time. -Models: Qwen, DeepSeek, GLM and 50+ other open models. +Models: ~35 API-Inference models, incl. DeepSeek V4 Pro / V4.1 Flash, Qwen3.5 / Qwen3.8, GLM-5.2, GLM-4.7-Flash, MiniMax-M3, Step-3.7-Flash, LongCat-Flash-Lite, Intern-S1. -*Checked Sep 11, 2026.* +OpenAI-compatible at `https://api-inference.modelscope.cn/v1`; Anthropic-compatible at `https://api-inference.modelscope.cn` (`/v1/messages`). + +*Checked Sep 29, 2026.* ### [Ollama Cloud](https://ollama.com/cloud) -Hosted counterpart to the local `ollama` CLI — same commands and API, models run on Ollama's GPUs. +Hosted counterpart to the local `ollama` CLI — same commands and API, models run on Ollama's GPUs. Prompts are not used for training. -Free tier with per-session and weekly caps (exact numbers unpublished). No credit card. +Free: a **starter amount of usage credits each month** (exact amount unpublished) on a smaller set of **starter models**, **1 concurrent request**, no credit card. Buying pay-as-you-go credits unlocks all cloud models. Session/weekly caps were removed on Aug 31, 2026. -Models: DeepSeek, Kimi, MiniMax, GPT-OSS, Qwen and 16 model families. OpenAI-compatible. +Models: DeepSeek V4, GLM-5.3, Kimi K3 / K2.7 Code, MiniMax M3, gpt-oss, Gemma 4, Nemotron 3, Mistral Large 3 (paid rates published per model). -*Checked Sep 11, 2026.* +OpenAI-compatible at `https://ollama.com/v1`; Anthropic-compatible at `https://ollama.com/v1/messages`. -### [LongCat (Meituan)](https://longcat.chat/platform/docs/zh/) - -Meituan's open platform for **LongCat-2.0** (1.6T-param MoE agentic model, **1M context**, 128K max output). -Both **OpenAI** (`https://api.longcat.chat/openai`) and **Anthropic** (`https://api.longcat.chat/anthropic`) -formats — works with Claude Code directly. - -Free: signup grants a daily quota plus a one-time resource pack (exact figures only shown in the console; -third-party reports ~500K tokens/day, raisable on request). Cache hits don't consume quota. Returns 429 -with exponential-backoff guidance when rate-limited. - -Source: , - -*Checked Sep 21, 2026.* +*Checked Sep 29, 2026.* ### [SenseNova Token Plan](https://www.sensenova.cn/token-plan) -SenseTime (商汤). Public beta — **Free tier ¥0/month**, phone-number signup, no card, no ID verification. +SenseTime (商汤). Public beta — **Free tier ¥0/month**, phone-number signup, no card. OpenAI-compatible at `https://token.sensenova.cn/v1` plus an Anthropic-compatible endpoint. -Limits: **60,000 credits / rolling 5 h** and **600,000 credits / rolling week**, separately for the general -pool and the Flash-Lite pool. Up to 20 API keys. +Limits: **60,000 credits / 5 h** (special models excepted). Up to 20 API keys. -Models at 0 credits: SenseNova 6.8 Flash-Lite, SenseNova U1 Fast, DeepSeek V4 Flash / V4 Pro, GLM-5.2, Kimi K3. +Models: SenseNova 6.8 Flash Lite (multimodal agent model) and SenseNova U1 Fast. Third-party models +(DeepSeek, GLM, Kimi) have been reported at 0 credits but are not listed on the official plan page. **Warning:** SenseTime says paid Lite/Pro tiers are "coming soon" with no end date for the free beta — don't build production on it. -Source: , +Source: -*Checked Sep 21, 2026.* +*Checked Sep 29, 2026.* ### [AMD Token Factory (Radeon Cloud)](https://developer.amd.com.cn/radeon/tokenfactory) -AMD's official inference platform on Radeon GPUs. **~$10-equivalent free quota per day**, resets daily -(does not roll over). OpenAI-compatible at `https://developer.amd.com.cn/radeon/api/v1`. +AMD's official inference platform on Radeon GPUs. Free shared endpoints after sign-in; usage is metered in +"points" against a daily allowance (reported as ~$10-equivalent/day, resets daily — not stated on the public page). +OpenAI-compatible at `https://developer.amd.com.cn/radeon/api/v1`. -Free models: DeepSeek V4 Flash 0731, MiniCPM5-1B; GLM-5.3-Flash and Qwen3.8-Flash-Next as limited-time free. -High time-to-first-token (~20 s) reported. +Free models: DeepSeek-V4.1-Flash, DeepSeek-V4-Flash (+ Vision-Exp), MiMo-V2.6-Flash, GLM-5.3-Flash, Qwen3.8-Flash-Next, +Qwen3.8-27B, MiniCPM5-2B (1M context on the DeepSeek/MiMo models). Marked "experimental" stability; high +time-to-first-token has been reported. -Source: +Source: -*Checked Sep 21, 2026.* - -### [Volcengine Ark (ByteDance)](https://console.volcengine.com/ark) - -ByteDance's model platform (火山引擎方舟). Permanent free tier: **2M tokens/day**, resets at midnight -(GMT+8). OpenAI-compatible at `https://ark.cn-beijing.volces.com/api/v3`. - -Free models: Doubao-Lite, DeepSeek R2 / V3 (within the daily quota). Doubao-Pro and higher tiers are paid. -Requires a phone number with real-name verification (mainland China). - -Source: - -*Checked Sep 21, 2026.* - -### [Baidu Qianfan](https://cloud.baidu.com/product/wenxinworkshop) - -Baidu's model platform (千帆). **ERNIE-Speed-128K and ERNIE-Lite are permanently free**, rate-limited. -OpenAI-compatible at `https://qianfan.baidubce.com/v2`. Real-name verification required. - -Source: - -*Checked Sep 21, 2026.* +*Checked Sep 29, 2026.* ### [iFlytek Spark](https://xinghuo.xfyun.cn/sparkapi) -讯飞星火. **Spark Lite is permanently free with unlimited tokens**, capped at **2 QPS**. -OpenAI-compatible at `https://spark-api-open.xf-yun.com/v1` (APIKey/APISecret auth). Individual real-name -verification required. +讯飞星火. **Spark Lite is free** (model id `lite`, 8K input / 4K output), throttled by per-second and concurrency limits +(third-party reports: 2 QPS, no token cap). OpenAI SDK-compatible at `https://spark-api-open.xf-yun.com/v1` +with `Authorization: Bearer ` from the console. Individual real-name verification required. -Source: +Source: -*Checked Sep 21, 2026.* +*Checked Sep 29, 2026.* ### [AIHubMix](https://aihubmix.com/models/free) -Gateway with **56 free models**, no credit card. Speaks Chat Completions, Messages, and Responses at -`https://aihubmix.com/v1`. +Gateway with **60 free models**, no credit card. Every free model speaks Chat Completions, Messages, and Responses at +`https://aihubmix.com/v1` (model IDs end in `-free`). -Free: 10 trial calls at signup (never expire). A **one-time $1 top-up** permanently unlocks -**100 req/day + 1M tokens/day** on the free catalogue (shared pool, resets daily). +Free: 10 trial calls at signup (never expire). A **one-time top-up of $1+** permanently unlocks +**100 req/day, 10 req/min, 1M tokens/day** on the free catalogue (resets daily). -Free models include glm-4.7-flash, hy3, minimax-m3, k2.6-code-preview, gpt-oss-20b, nemotron-3-ultra/super, -gemma-4-31b, mimo-v2.5(-pro), coding-glm-5.3, gpt-5.5, gemini-3.8-flash. +Free models include coding-glm-5.3(-flash), coding-kimi-k3, coding-minimax-m3, xiaomi-mimo-v2.6-pro/flash, mimo-v2.5-pro, +gpt-5.5, gemini-3.8-flash, qwen3.6-plus-preview, hy3, nemotron-3-ultra/super, gemma-4-31b-it, gpt-oss-20b, glm-4.7-flash. -Source: , +Source: -*Checked Sep 21, 2026.* +*Checked Sep 29, 2026.* ### [OVHcloud AI Endpoints](https://www.ovhcloud.com/en/public-cloud/ai-endpoints/catalog/) -EU-hosted (France). **Anonymous free tier — no API key, no signup**: 2 RPM per IP per model. -OpenAI SDK-compatible at `https://oai.endpoints.kepler.ai.cloud.ovh.net/v1`. +EU-hosted (France). **Anonymous free tier — no API key, no signup**: 2 RPM per IP per model +(with a key: 400 RPM per project per model, pay-per-token). OpenAI SDK-compatible at +`https://oai.endpoints.kepler.ai.cloud.ovh.net/v1`. -Models (20+): Qwen3.5-397B-A17B, gpt-oss-120b/20b, Llama 3.3 70B, Qwen3.6-27B, Qwen3-Coder-30B, -Qwen2.5-VL-72B, Mistral Small 3.2, Mistral Nemo. +Models: Qwen3.5-397B-A17B, Qwen3.8-27B, Qwen3.6-27B, Qwen3-Coder-30B-A3B, gpt-oss-120b/20b, Llama 3.3 70B, +Mistral Small 3.2, Mistral Nemo, Qwen2.5-VL-72B, plus Whisper, embeddings, and TTS. -Source: +Source: -*Checked Sep 21, 2026.* +*Checked Sep 29, 2026.* ### [LLM7.io](https://token.llm7.io) -UK gateway. Anonymous access needs no key (10 RPM, 60 req/hour); a free token from `token.llm7.io` raises -the limits. OpenAI-compatible at `https://api.llm7.io/v1`. +Gateway with keyless access. Anonymous: 10 req/min, 60 req/hour, **500K tokens/day**; a free token from +`token.llm7.io` raises this to 40 req/min, 100 req/hour, **1M tokens/day**. OpenAI-compatible at `https://api.llm7.io/v1`. -Models: gpt-oss-20b, minimax-m2.7 (180K ctx), Mistral Nemo. +Free (`turbo`-tier) models: DeepSeek-V4-Flash-0731 (400K ctx), codestral-latest, minimax-m2.7 (180K ctx), Mistral Nemo. +Everything else needs Pro ($12/month) or a topped-up balance. -Source: +Source: -*Checked Sep 21, 2026.* +*Checked Sep 29, 2026.* ### [Token Harbor](https://tokenharbor.ai/pricing) -Small gateway. **Free tier $0/month** with a rolling 4-week quota (amount unpublished, unused quota carries -over). Agent Pass **$1.99/month** ($0.99 first month) tops up the same pool. OpenAI-compatible at -`https://tokenharbor.ai/v1`. +Small gateway. **Free tier $0/month** with a rolling allowance (amount unpublished, unused allowance carries over). +Agent Pass **$1.99/month** ($0.99 first month) adds $10 of included usage. OpenAI-compatible at `https://tokenharbor.ai/v1`. -Free models: DeepSeek V4 Flash, DeepSeek V4.1 Flash, MiMo V2.5 ("promotional models added over time"). +Free models: DeepSeek V4.1 Flash, MiMo V2.6 Flash, TH-Rudder ("promotional models added over time"). -**Warning:** blocks requests from mainland China, Hong Kong, and Macau (`region_blocked`). Free requests may be logged. +**Warning:** free models are opt-in and their prompts/responses may be retained for diagnostics and model +improvement. Access is restricted in some regions (reported: mainland China, Hong Kong, Macau). -Source: , +Source: -*Checked Sep 21, 2026.* +*Checked Sep 29, 2026.* ### [Aion Labs](https://www.aionlabs.ai/app/api-keys/) Permanent free tier, no credit card. **15 RPM, 20K tokens/day**. OpenAI-compatible at `https://api.aionlabs.ai/v1`. -Models: aion-2.0, aion-3.0, aion-3.0-mini (128K ctx, reasoning), aion-rp-llama-3.1-8b. Tuned for -roleplay/storytelling rather than coding. +Models: aion-3.5 / aion-3.5-mini (256K ctx), aion-3.0 / aion-3.0-mini, aion-2.0 (128K ctx, reasoning), +aion-rp-llama-3.1-8b. Tuned for roleplay/storytelling rather than coding. -Source: +Source: -*Checked Sep 21, 2026.* +*Checked Sep 29, 2026.* ### [Experiential Labs](https://platform.experientiallabs.ai) -YC-backed "open-source OpenRouter". Free plan: **~500 credits/month** (1 credit = $0.01), hard stop when -exhausted, no auto-upgrade; new orgs get extra welcome credits. OpenAI-compatible at +"Open-source OpenRouter" gateway. Free plan: after a one-time **$1 card verification** (credited to your balance), the +balance **refills to 100 credits ($1) every 30 days**. OpenAI-compatible (Chat + Responses) and Anthropic-compatible at `https://api.experientiallabs.ai/v1`. -$0-labelled models: Qwen3.8 27B, DeepSeek V4 Flash, GPT-5.6 Luna, GPT-6 Astra, Claude Fable 5.1 -(all 1M ctx) — but the monthly credit cap still applies. +$0 models right now: GPT-6 Luna (promo, 100% off), MiMo-V2.6-Pro (free, zero-data-retention route), Jev. +Qwen3.8 27B and DeepSeek V4 Flash are 75% off. -**Warning:** business model is free traffic in exchange for training traces; free-tier uptime is low -(72–100 % by model, ~4.5 s TTFT). Prototyping only. +**Warning:** prompt/response capture is controlled by an org-wide switch — turn it off or require ZDR routes for private +code. Free-promo uptime varies by model (e.g. MiMo-V2.6-Pro ~76%). Prototyping only. -Source: +Source: , -*Checked Sep 21, 2026.* +*Checked Sep 29, 2026.* -### [Empero](https://free.empero.org) +### [Tencent TokenHub](https://cloud.tencent.com/document/product/1823/130053) -Community endpoint from German lab EmperoAI. **Completely free, no signup** — any string works as the API -key (convention: `free`). OpenAI-compatible at `https://free.empero.org/v1`. +Tencent Cloud's model platform (replaces the old Hunyuan platform, which shuts down Sep 30, 2026). -Models rotate often: glm-5.3-flash, qwen3.8-flash, Qwen3.8-27B-FP8. +- **Free trial:** **1M tokens per language/multimodal model**, claimed once per main account per model from the + Model Square "新用户福利" button; valid **1 year** from claim. Claim window ends **Dec 31, 2026**. +- **Token Plan (monthly):** Hy plan from **¥28** (Lite, 560 pts) to ¥468; universal plan ¥39–¥599. Models: DeepSeek-V4-Flash/Pro, + MiniMax-M2.7/M3, GLM-5.2/5.3/5.3-Flash, Kimi K2.7 Code, Kimi K3 (+ Hy3, Hy4 preview on the Hy plan). +- Plan endpoints: OpenAI `https://api.lkeap.cloud.tencent.com/plan/v3`, Anthropic `https://api.lkeap.cloud.tencent.com/plan/anthropic`. -**Warning:** prompts and responses are logged (IP hashed) to train their open models — never send private -data. Frequent `upstream_down` / 503 under load. +Which models qualify for the free trial is shown in the console. Real-name verification requirement: not confirmed. -Source: +Source: , -*Checked Sep 21, 2026.* +*Checked Sep 29, 2026.* + +### [Alibaba Cloud Model Studio — free quota](https://www.alibabacloud.com/help/en/model-studio/new-free-quota) + +New users get **1,000,000 free tokens per model** (typical), valid **90 days** from activation (or model release / +approval, whichever is later). Singapore region, international deployment scope only; real-time inference only +(no batch/fine-tune). Each model — and each dated snapshot — has its own quota; RAM users share the account's pool. +OpenAI-compatible at `https://dashscope-intl.aliyuncs.com/compatible-mode/v1`. + +Source: + +*Checked Sep 29, 2026.* + +### [Z.ai API — free Flash models](https://docs.z.ai/guides/overview/pricing) + +Zhipu's international platform (no mainland ID needed). **GLM-4.7-Flash**, **GLM-4.5-Flash** (text) and **GLM-4.6V-Flash** (vision) are priced +**Free** for input, cached input, and output. OpenAI-compatible at `https://api.z.ai/api/paas/v4`. + +Rate/concurrency limits for free models are shown only in the console (not published). + +Source: + +*Checked Sep 29, 2026.* + +### [Nebius Token Factory](https://tokenfactory.nebius.com/) + +Formerly Nebius AI Studio; EU-based open-model inference (60+ models). **$1 trial credit on first sign-up, valid 30 days**; joining the +free **Nebius Builder Program** adds a **$25 Token Factory credit** (open to everyone; meant for learning/testing). +OpenAI-compatible at `https://api.tokenfactory.nebius.com/v1`. + +**Warning:** setting up a billing account (bank card) is mandatory to finish onboarding. + +Source: , + +*Checked Sep 29, 2026.* + +### [Novita AI](https://novita.ai/pricing) + +Open-model inference platform. Two models are priced **$0 in / $0 out**: `inclusionai/ling-3.1-flash` and +`inclusionai/ling-3.0-flash-sante` (262K ctx). OpenAI-compatible at `https://api.novita.ai/openai`. + +Rate limits for the free models and any signup credit: not published on the pricing page. + +Source: + +*Checked Sep 29, 2026.* --- diff --git a/plans/reports/research-260929-1947-coding-plans-audit.md b/plans/reports/research-260929-1947-coding-plans-audit.md new file mode 100644 index 0000000..4868860 --- /dev/null +++ b/plans/reports/research-260929-1947-coding-plans-audit.md @@ -0,0 +1,537 @@ +# Coding-plan sections audit (Sep 29, 2026) + +Scope: README sections "Claude Code Guest Passes", "Claude AI Ecosystem" (AgentKit), and every entry under +"Providers with Coding Plans". README.md was not modified. Every claim below was checked against a live page on +2026-09-29 using curl or WebFetch; no browser was used. Where a price only loads through JavaScript, I name the +fallback source used, such as a docs markdown export, embedded page state, a JS bundle default, or a media report. + +## Summary + +| Entry | Status | Headline change | +|---|---|---| +| Claude Code Guest Passes | CURRENT | Program still live (`/passes`, Max-plan only). The user's own pass remains out of stock. | +| AgentKit | UPDATE | Desktop app now ships at $49/yr (1 device) or $99/yr (3 devices), no longer a waitlist. Runtime list changed. | +| Z.ai | UPDATE | Credit-based plans since Jul 30, 2026. Now $18 / $80 / $168. Models are now GLM-5.3 and GLM-5.3-Flash. | +| MiniMax | UPDATE | Now $22 / $55 / $132. Plus is 3-4 agents and Max is 4-5 agents. Usage-based deduction. The README's referral end date (Jul 1) has passed, but the program is still documented. | +| Kimi Code | UPDATE | Tiers renamed to Go / Plus / Pro / Max (¥49 / ¥99 / ¥199 / ¥699). The ¥49 Go tier has no Kimi Code. New K3 model. Weekly cap removed for new members. | +| Alibaba Token Plan | UPDATE | New Personal Edition (Lite $6 to Pro $68, promo prices). Team seats also discounted. Model list fully refreshed. | +| Alibaba Coding Plan | UPDATE | Still $50 Pro and still limited stock. Model list changed (qwen3.7-plus added). | +| opencode Go | UPDATE | $10 Go plus a new $40 Go Plus. **Referral program has ended.** The $5 first month is no longer shown. Dollar-based limits. | +| Synthetic | UPDATE | $30/mo per "pack", 500 req/5h. Model list changed. Now has an Anthropic endpoint. | +| BigModel.cn GLM Coding Plan | UPDATE | New prices ¥118 / ¥538 / ¥1,078 (from media reports). The ¥49 / ¥149 / ¥469 prices are now legacy-only. Models are GLM-5.3 and GLM-5.3-Flash. | +| BytePlus ModelArk | UPDATE | **README prices are wrong:** list price is $10 Lite and $50 Pro. **Referral campaign ends Sep 30, 2026.** | +| Xiaomi MiMo | UPDATE | USD prices are now listed ($6 / $16 / $50 / $100). Credit amounts changed (4.1B to 82B). Models are now MiMo-V2.6. | +| GitHub Copilot Pro | UPDATE | Premium requests replaced by GitHub AI Credits (Pro gets 1,500/mo). A new $100 Max tier. Pro+ now 7,000 credits. | + +--- + +## Claude Code Guest Passes + +**Status: CURRENT.** The program exists. The user's own link is still out of stock, and the README already says so. + +Proposed block (unchanged apart from the rules text and the date): + +```markdown +## Claude Code Guest Passes + +One week of free Claude Pro (includes Claude Code). Eligible subscribers (currently Max plans) get three passes via +the `/passes` command in Claude Code. New users only; requires payment info but can be cancelled before the trial +ends. + +**My passes:** + +- ~~~~ — out of stock as of 2026-04-30 + +Have a spare pass? Open a PR adding your link, or open an issue. + +*Checked Sep 29, 2026.* +``` + +What changed: +- Added how passes are issued. The official docs list `/passes` as "Share a free week of Claude Code with friends. Only visible if your account is eligible." +- The "Max plans only, three passes" detail comes from third-party guides, not from an Anthropic page. +- `claude.ai/referral/ZkoAngod1A` returns HTTP 403 to curl. That is a bot block, not proof the link is dead or alive, so the strikethrough stays. +- Added a checked-date line, which this section did not have before. + +Sources: +- https://code.claude.com/docs/en/commands.md (official, the `/passes` row) +- https://madrobot.blog/2026/09/27/claude-referral-link-guest-passes-how-to-refer/ (third party, dated Sep 27, 2026) +- https://growsurf.com/blog/anthropic-claude-referral-program/ (third party) + +## AgentKit + +**Status: UPDATE.** + +```markdown +### [AgentKit](https://agentkit.best/) + +AgentKit (formerly **ClaudeKit**, `claudekit.cc`) sells production-ready kits of skills, slash +commands, subagents, and workflows for coding agents — native support for Claude Code, Codex, Antigravity, Pi, and +Oh My Pi; Cursor and DeepSeek Harness in beta; OpenCode, Grok, and GitHub Copilot in preview. **Engineer** ($99) +covers frontend, backend, database, DevOps, code review, and debugging; **Marketing** ($99) adds research, SEO, +competitor-intelligence, and copywriting agents; the bundle is **$149** (108+ skills, 95+ commands, 45 subagents). +The **AgentKit App** desktop cockpit (macOS/Windows) is sold separately at **$49/yr** (1 device) or **$99/yr** +(3 devices). + +> Referral: **20% off** your first purchase via (code: `BWA910UK`). + +*Checked Sep 29, 2026.* +``` + +What changed: +- The desktop app is no longer on a waitlist. It is sold at $49/yr for one device or $99/yr for three devices. +- The runtime support list now has three levels: native, beta, and preview. GitHub Copilot is only at the "preview" level. +- Kit prices, skill counts, and command and subagent counts are unchanged. +- The referral program still exists. The leaderboard page shows referrer commission tiers from 20% to 40%. +- I could not verify the 20% buyer discount for `?ref=` links. The page shows a timer-based "30% special offer" banner, which may be separate from referrals. + +Sources: +- https://agentkit.best/ +- https://agentkit.best/leaderboard +- `https://claudekit.cc` now redirects to agentkit.best (checked with curl) + +## Z.ai + +**Status: UPDATE.** On Jul 30, 2026 the plans moved to a credit system. Legacy plans are no longer sold to new users. + +```markdown +### [Z.ai](https://z.ai/subscribe) + +GLM Coding Plan, credit-based since Jul 30, 2026. **Lite $18 / Pro $80 / Max $168 per month** (quarterly −20%, +yearly −30%). Credits per 5 h / per week: Lite 2,000 / 10,000 · Pro 12,000 / 60,000 · Max 28,000 / 140,000; +off-peak usage (outside Mon–Fri 14:00–18:00 UTC+8) costs 50% fewer credits. + +Models: GLM-5.3, GLM-5.3-Flash (GLM-5.2/5.1 requests auto-route to GLM-5.3, GLM-4.7 to GLM-5.3-Flash). Includes +Vision, Web Search, Web Reader, and Zread MCP. +Endpoints: Anthropic `https://api.z.ai/api/anthropic`, OpenAI `https://api.z.ai/api/coding/paas/v4`. +Tools: Claude Code, Cursor, Cline, Roo Code, Kilo Code, OpenCode, OpenClaw, Crush, Goose (supported tools only). + +> Referral: + +*Checked Sep 29, 2026.* +``` + +What changed: +- Plans are now credit-based. Prices rose from "$18+" to $18 / $80 / $168 per month. Quarterly costs are $43.20 / $192 / $403.20 and yearly $151.20 / $672 / $1,411.20. +- The model lineup changed from GLM-5.1 / 5-Turbo / 4.7 / 4.5-Air to GLM-5.3 and GLM-5.3-Flash. +- Plan quota may only be used in the supported tools. +- The invite program is still active: the invitee gets 10% off the first order, and the inviter gets 10% of that payment in credits once 3 friends have paid. +- A temporary promo runs to Oct 7: all-day off-peak rates, plus unlimited GLM-5.3-Flash overnight in ZCode and AutoClaw. It is left out of the block because it expires. +- Prices came from the defaults in the z.ai subscribe page's JS bundle (the `codePlansV3` table). Third-party guides report the same $18 / $80 / $168. + +Sources: +- https://docs.z.ai/devpack/overview.md +- https://docs.z.ai/devpack/notice/usage-revision.md +- https://docs.z.ai/devpack/quick-start.md +- https://docs.z.ai/devpack/credit-campaign-rules.md +- https://z.ai/subscribe (JS chunk defaults) +- https://www.aipricing.guru/z-ai-subscription-pricing/ (third party) + +## MiniMax + +**Status: UPDATE.** + +```markdown +### [MiniMax](https://platform.minimax.io) + +Token Plan — usage-based deduction from one shared quota (5-hour rolling + weekly windows). $22–$132/month. + +- Plus ($22): 3-4 agents +- Max ($55): 4-5 agents +- Ultra ($132): 6-7 agents + +Covers the full MiniMax lineup (M3 / M2.7 / image / speech); MiniMax H3 video, voice design, and rapid voice +cloning are excluded. Top-up Credits: 1,000 credits = $1, valid 365 days. +OpenAI- and Anthropic-compatible. Tools: Claude Code, Codex, Cursor, TRAE, Hermes Agent, OpenClaw, Pi. + +> Referral (10% off) until **Jul 1, 2026** — **For Referred Users:** 10% off subscription + become a dev ambassador. **For Referrers:** 10% back in API voucher per paid referral, usable across all MiniMax models, plus priority access to events and model previews. [View details](https://platform.minimax.io/subscribe/token-plan?code=CAQ5sxHAq6&source=link) + +*Checked Sep 29, 2026.* +``` + +What changed: +- Prices changed from $20 / $50 / $120 to $22 / $55 / $132. +- Agent counts changed. Plus went from 4-5 to 3-4 agents and Max from 6-7 to 4-5 agents. Ultra is unchanged at 6-7. +- Quota changed from per-call counting to deduction based on actual usage. +- The model list changed. The official docs now list M3 / M2.7 / image / speech. Hailuo 2.3 and Music-2.6 are not mentioned, and H3 video is excluded. +- I found no official mention of a yearly "~17% off" discount, so it was dropped. +- The referral program is still documented in the official FAQ, with no end date given. The invitee gets 10% off (all plans and upgrades) and "Builder" status. The referrer gets credits worth 10% of the payment, valid 90 days. +- The README's "until Jul 1, 2026" date has passed, but the blockquote is kept verbatim as instructed. The user should decide whether to edit their own wording. + +Sources: +- https://platform.minimax.io/docs/guides/pricing-token-plan.md +- https://platform.minimax.io/docs/token-plan/intro.md +- https://platform.minimax.io/docs/token-plan/faq.md +- https://platform.minimax.io/docs/llms.txt + +## Kimi Code + +**Status: UPDATE.** + +```markdown +### [Kimi Code](https://www.kimi.com/code) + +Moonshot's coding perk bundled with Kimi membership (desktop app, CLI, VS Code; Claude Code, OpenCode, Codex, +Hermes Agent via API key). New plans: Go ¥49 (no Kimi Code), **Plus ¥99**, **Pro ¥199**, **Max ¥699** per month; +annual billing saves up to ¥1,680. Rolling 5-hour window plus a monthly total (weekly cap removed for new members). + +Models: `k3` (K3, up to 1M context on Pro+), `k3-256k`, `kimi-for-coding` (K2.8 Preview), `kimi-for-coding-highspeed` +(K2.7 Code, Pro+). OpenAI- and Anthropic-compatible (`https://api.kimi.ai/coding/`). Pay-as-you-go also at +`platform.moonshot.ai`. + +> Referral (code: `C8CJ6F`) — sign up or subscribe via my link and we each get a guaranteed benefit, up to **1-Year Membership Credits**: +> - Sign up: +> - Subscribe: + +*Checked Sep 29, 2026.* +``` + +What changed: +- The tier names Adagio / Andante / Presto are stale. +- Kimi renamed its tiers to Go / Plus / Pro / Max around Sep 18 and kept prices the same. The help center still shows the legacy names at those prices: Andante ¥49, Moderato ¥99, Allegretto ¥199, Allegro ¥699. +- For new members, the ¥49 Go tier does not include Kimi Code. +- The model lineup is new: K3, K2.8 Preview, and K2.7 Code HighSpeed. +- The weekly quota was removed for new members. The 5-hour window and a monthly total remain. +- Extra Usage top-ups are available (¥25 minimum). +- The referral link still resolves. It redirects to `kimi-bot.com/activities/invite/share?...`, which is titled "Kimi 登月同行计划 - 邀请好友赢好礼". The reward details render only with JavaScript and could not be verified. +- USD prices for the international site were not verified. + +Sources: +- https://www.kimi.com/code/docs/en/kimi-code/membership.html +- https://www.kimi.com/code/docs/ (models and endpoints, Chinese) +- https://www.kimi.com/en/help/membership/membership-pricing +- https://www.kucoin.com/news/flash/kimi-launches-new-subscription-tiers-after-two-month-pause-kimi-code-not-sold-separately (media) +- https://codingplan.org/en/plans/kimi (third party) + +## Alibaba Cloud Model Studio + +### Token Plan + +**Status: UPDATE.** A Personal Edition was added and Team prices changed. Both docs pages were last updated Sep 28, 2026. + +```markdown +> Referral: Up to **$1,700** in free trial credits via (code: `A92LU5`). + +#### [Token Plan](https://www.alibabacloud.com/help/en/model-studio/token-plan-overview) + +Credit-based subscription (Singapore region only). Tools: Claude Code, Cursor, Qwen Code, Codex, Qoder, OpenClaw. + +**Personal Edition** (monthly quota, limited-time prices): +- Lite: **$6/month** (list $8) — 11,500 Credits +- Essential: **$10/month** (list $16) — 25,500 Credits +- Standard: **$18/month** (list $25) — 45,000 Credits +- Pro: **$68/month** (list $80) — 180,000 Credits +- Extra bundle: $15 — 20,000 Credits (up to 5, needs an active plan) + +**Team Edition** (per seat, no training on your data): +- Standard: **$20/seat/month** (list $30) — 25,000 Credits +- Pro: **$75/seat/month** (list $100) — 100,000 Credits +- Max: **$200/seat/month** — 250,000 Credits +- Shared quota pack: **$700** — 625,000 Credits + +Models (Personal): auto, qwen3.8-max, qwen3.8-flash, qwen3.7-max, qwen3.7-plus, qwen3.6-flash, deepseek-v4.1-flash, +deepseek-v4-pro, glm-5.3, glm-5.2, plus image (qwen-image-3.0-pro, wan2.7-image/-pro), audio, and HappyHorse video. +Team Edition adds kimi-k2.7-code/k2.6/k2.5, glm-5.1/5, MiniMax-M2.5, deepseek-v3.2. +``` + +What changed: +- The Personal Edition is new. +- Team Standard and Pro now have limited-time prices of $20 and $75 (list $30 and $100). "Premium" is now called "Max". +- The model list was replaced. qwen3.6-plus, glm-5, and MiniMax-M2.5 are now Team-only. +- On the account referral: the campaign page still loads with the code (HTTP 200) and shows an "Invite & Earn Plan". I could not find the "$1,700" figure on the page, only "$90 ECS credits" and "free tokens". The blockquote is kept verbatim, and the amount is marked unverified. + +### Coding Plan + +**Status: UPDATE.** The product is still listed, the Pro tier is still sold in limited quantities, and Lite is discontinued. + +```markdown +#### [Coding Plan](https://www.alibabacloud.com/help/en/model-studio/coding-plan) + +Pro plan: **$50/month** — 6,000 req/5-hour, 45,000 req/week, 90,000 req/month. + +Models: qwen3.7-plus, qwen3.6-plus, kimi-k2.5, glm-5, MiniMax-M2.5 (recommended); qwen3.5-plus, qwen3-max-2026-01-23, +qwen3-coder-next, qwen3-coder-plus, glm-4.7. +Endpoints: OpenAI `https://coding-intl.dashscope.aliyuncs.com/v1`, Anthropic `.../apps/anthropic`. +Tools: Claude Code, Cursor, Cline, Codex, OpenCode, Qwen Code, Qoder, Kilo CLI, OpenClaw, Hermes Agent, and more. +Lite closed to new subscribers (Mar 20, 2026) and to renewals/upgrades (Apr 13, 2026). + +**Note:** limited slots, restocked daily at 00:00 UTC+8 (first come, first served); Alibaba now recommends Token Plan +instead. As of Jun 5, 2026 it was effectively unbuyable. + +*Checked Sep 29, 2026.* +``` + +What changed: +- The model list changed. qwen3.7-plus and qwen3.6-plus are new and are the recommended models. +- Alibaba now officially confirms the stock limit ("limited-quantity offering and is no longer available once sold out"). +- I could not test whether a purchase goes through, because that needs a logged-in console. + +Sources: +- https://www.alibabacloud.com/help/en/model-studio/token-plan-overview +- https://www.alibabacloud.com/help/en/model-studio/token-plan-personal-overview +- https://www.alibabacloud.com/help/en/model-studio/token-plan-team-overview +- https://www.alibabacloud.com/help/en/model-studio/coding-plan +- https://www.alibabacloud.com/campaign/benefits?referral_code=A92LU5 + +## opencode — Go + +**Status: UPDATE.** The referral program has ended. + +```markdown +### [opencode — Go](https://opencode.ai/go) + +OpenCode subscription for curated open models. **Go $10/month**, **Go Plus $40/month** (higher limits). Limits are +monthly dollar amounts per model (5-hour = 20%, weekly = 50% of the monthly limit); e.g. Go allows ~220 GLM-5.3 or +~3,200 MiniMax M3 requests per 5 h. Falls back to Zen balance when enabled. + +Models: GLM-5.3/5.3-Flash/5.2, Kimi K3/K2.7 Code/K2.6, MiniMax M3/M2.7, Qwen3.8 Max/Flash, Qwen3.7 Plus, +DeepSeek V4.1 Flash/V4 Pro/V4 Flash, MiMo-V2.6(-Pro/-Flash)/V2.5(-Pro), LongCat-2.0, Hy4 preview, Hy3, Grok 4.7/4.6, +GPT 6 Luna/5.6 Luna, plus limited-time free models. +Endpoints: `https://opencode.ai/zen/go/v1/{chat/completions,messages,responses}`. Validated clients: Claude Code, +Codex, Hermes, ZCode, Pi, jcode, Kilo Code CLI. Model format: `opencode-go/`. + +My referral: + +> Invite friends to OpenCode Go. Earn $5 when a friend subscribes, and they'll get $5 too. Share your referral link; your friend joins and subscribes to Go; you both get a $5 usage credit to apply toward your Go usage limits. +> +> Referral link: + +*Checked Sep 29, 2026.* +``` + +What changed: +- **The referral program has ended.** `opencode.ai/go?ref=HE42WGS8BM` shows a warning: "The referral program has ended. Referral links no longer earn credit for you or the person who shared them." The blockquote is kept verbatim as instructed, but the user will probably want to strike it through or remove it. +- The "$5 first month" offer is no longer on the Go page, so it was dropped. +- Go Plus ($40/month) is new. +- Limits changed from per-request counts ("200–10,200 req") to dollar budgets for each model. +- The model list grew a lot. +- Claude Code now works natively, because Go has an Anthropic `/messages` endpoint and recognises Claude Code's session header. The LiteLLM or `oc-go-cc` workaround is no longer needed. + +Sources: +- https://opencode.ai/docs/go/ +- https://opencode.ai/go +- https://opencode.ai/go?ref=HE42WGS8BM (referral-ended notice in the HTML) + +## Synthetic + +**Status: UPDATE.** + +```markdown +### [Synthetic](https://synthetic.new/) + +Privacy-first inference (no training on prompts/responses). **$30/month per pack** ($1/day) — 500 requests/5 h, +1 concurrent request per model (buy more packs to raise both), or usage-based pay-per-token. + +Models: Kimi-K3, DeepSeek-V4.1-Flash (beta), GLM-5.3-Flash, GLM-4.7-Flash, Qwen3.8-27B, gpt-oss-120b, +NVIDIA Nemotron-3-Super-120B; nomic-embed-text-v1.5 embeddings included. +OpenAI-compatible (`https://api.synthetic.new/openai/v1`) and Anthropic-compatible +(`https://api.synthetic.new/anthropic/v1`) — guides for Claude Code, Crush, OpenCode, GitHub Copilot, OpenClaw, +Xcode, Roo, KiloCode, Octofriend. + +> Referral: **$10.00** in subscription credit via + +*Checked Sep 29, 2026.* +``` + +What changed: +- The $30/month price is now explained as one "pack" with 500 requests per 5 hours. +- The model list changed. Kimi K2.6, MiniMax M2.5, and GLM 5.1 are gone. Kimi-K3, GLM-5.3-Flash, DeepSeek-V4.1-Flash, and Qwen3.8-27B are new. +- An Anthropic-compatible endpoint now exists. +- The referral still works. The referred landing page shows "Subscribe today and get $10.00 off your first month!" The README's wording ("subscription credit") is close enough and was kept verbatim. + +Sources: +- https://synthetic.new/pricing +- https://dev.synthetic.new/docs/api/overview +- https://synthetic.new/?referral=CNBFyw28zF0dZoj + +## BigModel.cn — GLM Coding Plan + +**Status: UPDATE.** The plans changed on Jul 30, 2026, the same day as Z.ai's change. + +```markdown +### [BigModel.cn — GLM Coding Plan](https://www.bigmodel.cn/glm-coding) + +The Chinese (mainland) counterpart of Z.ai's GLM Coding Plan — same underlying Zhipu AI models, but billed in CNY through bigmodel.cn. Suited for users who can pay via Alipay / WeChat Pay or already have a 智谱 AI account. + +Credit-based plans since Jul 30, 2026 (monthly list price): **Lite ¥118**, **Pro ¥538**, **Max ¥1,078**. +Credits per 5 h / per week: Lite 2,000 / 10,000 · Pro 12,000 / 60,000 · Max 28,000 / 140,000; off-peak usage +(outside Mon–Fri 14:00–18:00 UTC+8) costs 50% fewer credits. Legacy V1/V2 subscribers keep ¥49 / ¥149 / ¥469. + +All tiers support GLM-5.3 and GLM-5.3-Flash (GLM-5.2/5.1 route to GLM-5.3; GLM-5-Turbo/4.7 route to GLM-5.3-Flash). +Anthropic endpoint `https://open.bigmodel.cn/api/anthropic`, OpenAI endpoint `https://open.bigmodel.cn/api/coding/paas/v4`. +Tools: Claude Code, Kilo Code, OpenClaw (lower priority), OpenCode, TRAE, CodeBuddy, and others on the supported list. + +Referral program (challenge-based, resets every 30 invitees): +- Invited friend gets **5% off** their first GLM Coding Plan order. +- Referrer gets **10% cashback** once 3 friends subscribe, plus an **additional 10%** of the total paid amount for every 30 invitees. +- Rebate credit is usable for resource packs, API calls, and subscription renewals on the BigModel platform. + +Source: , [^bigmodel] + +Homepage: + +My referral: + +>🚀 Join the GLM Coding Plan via my link — get 5% off your first order. Subscribe at https://www.bigmodel.cn/glm-coding?ic=VGRZKHKNKW (invitation code: `VGRZKHKNKW`). + +[^bigmodel]: Checked on Sep 29, 2026 +``` + +What changed: +- The ¥49 / ¥149 / ¥469 prices are now legacy-only. The official "老用户权益说明" notice lists them as the V2 prices that V1 and V2 holders keep. +- The new list prices, ¥118 / ¥538 / ¥1,078, come from Sina Finance and a Tencent Cloud developer article. IT之家 separately confirms "每月 118 元起". The official price page renders only with JavaScript, so I could not read it directly. +- The models changed to GLM-5.3 and GLM-5.3-Flash. +- The block keeps the entry's existing footnote style (`[^bigmodel]`) instead of `*Checked ...*` so its structure stays the same. +- I could not verify the referral terms on the current page. The README says 5% for the invitee, while Z.ai's version of the program gives 10%. + +Sources: +- https://docs.bigmodel.cn/cn/coding-plan/overview.md +- https://docs.bigmodel.cn/cn/coding-plan/notice/usage-revision.md +- https://docs.bigmodel.cn/cn/coding-plan/quick-start.md +- https://finance.sina.com.cn/stock/t/2026-07-31/doc-iniksxpi1210399.shtml (media) +- https://cloud.tencent.com/developer/article/2718987 (media) +- https://www.ithome.com/0/983/934.htm (media) + +## BytePlus ModelArk — Coding Plan + +**Status: UPDATE.** The prices in the README are wrong, and the referral campaign ends tomorrow. + +```markdown +### [BytePlus ModelArk — Coding Plan](https://www.byteplus.com/en/activity/codingplan) + +ByteDance. Lite: **$10/month** ($30/quarter), Pro: **$50/month** ($150/quarter). New-user first-purchase promo +($5/$25) suspended since Mar 17, 2026. +Limits: Lite ~1,900 req/5 h, ~12,000/week, ~24,000/month; Pro 5× Lite (~9,500 / ~60,000 / ~120,000). + +Models: Auto, Dola-Seed-2.0-Pro/Lite/Code, ByteDance-Seed-Code, GLM-5.3-Flash, GLM-5.2, GLM-5.1, Kimi-K2.5, +DeepSeek-V4.1-Flash, DeepSeek-V4-Pro/Flash, GPT-OSS-120b. +Endpoints: OpenAI `https://ark.ap-southeast.bytepluses.com/api/coding/v3`, Anthropic `.../api/coding`. +Tools: Claude Code, Cursor, Cline, Codex, Roo Code, Kilo Code, OpenCode, OpenClaw, TraeCode, Hermes Agent. + +> Referral: + +*Checked Sep 29, 2026.* +``` + +What changed: +- The prices were wrong. The README said $15 / $35, but the official list price is $10 / $50 per month. +- The README's promo note is also wrong. The first-purchase promo was $5 / $25 and was suspended on Mar 17, 2026, not "early 2026". +- Added the request quotas. +- The model list changed. ByteDance-Seed-2.0 is now listed as Dola-Seed-2.0, and GLM-5.3-Flash and DeepSeek-V4 are new. +- **The referral campaign runs from 2026-01-13 to 2026-09-30.** The invitee gets 10% off the first order, and the referrer gets a 10% voucher, valid 90 days. After tomorrow the user's link may stop giving a discount. Re-check on or after Oct 1. + +Sources: +- https://docs.byteplus.com/en/docs/ModelArk/1925114 (subscription overview, updated Sep 28, 2026, read from embedded `MDContent`) +- https://docs.byteplus.com/en/docs/modelark/1928265 (offer notice and prices) +- https://docs.byteplus.com/en/docs/modelark/2165246 (referral campaign period) +- https://www.byteplus.com/en/activity/codingplan + +## Xiaomi MiMo Open Platform + +**Status: UPDATE.** + +```markdown +### [Xiaomi MiMo Open Platform](https://platform.xiaomimimo.com) + +I'm on Xiaomi MiMo Open Platform — running Xiaomi's flagship MiMo V2.5 and the rest of the lineup. Sign up with my code and you'll instantly get $2 in API credits. + +After signup, enter the code at the bottom-left of the console. Credits valid 40 days. + +**Token Plan** (monthly): Lite $6 / ¥39 (4.1B credits), Standard $16 / ¥99 (11B), Pro $50 / ¥329 (38B), +Max $100 / ¥659 (82B). Annual plans 12% off; 12% off first Individual purchase; 0.8× consumption 00:00–08:00 +Beijing time. Team Edition from $16/seat. Models: mimo-v2.6-pro, mimo-v2.6-flash, ASR/TTS (mimo-v2.5 and +mimo-v2.5-pro retire Oct 21, 2026). Works with OpenCode, OpenClaw, Claude Code. + +> Referral: Code `T8ESAY` · + +*Checked Sep 29, 2026.* +``` + +What changed: +- USD prices are now published. +- Credit amounts changed from 60M / 200M / 700M / 1,600M to 4.1B / 11B / 38B / 82B. The CNY prices are unchanged. +- The models moved to MiMo-V2.6, and V2.5 retires on Oct 21, 2026. +- A Team Edition is new. +- The first paragraph is the user's own referral copy and was kept verbatim. It still says "MiMo V2.5", which the user may want to change to V2.6. +- The referral program still appears on the official site as an "Invite friends" page, but that page renders only with JavaScript. Search snippets of the official page say each side gets $2, the credits are valid 40 days, the invitee must have registered within 3 days, and there is 10% off the first plan within 30 days. I could not confirm this in the page body. +- Docs on `platform.xiaomimimo.com/docs/*` now redirect (302) to `mimo.mi.com`. + +Sources: +- https://mimo.mi.com/docs/en-US/price/token-plan (updated Sep 21, 2026) +- https://mimo.mi.com/docs/en-US/promotions/refer +- https://platform.xiaomimimo.com + +## GitHub Copilot Pro + +**Status: UPDATE.** Premium requests were replaced by usage-based GitHub AI Credits. + +```markdown +### [GitHub Copilot Pro](https://github.com/features/copilot/plans) + +Cheapest mainstream coding seat. **$10/month** — unlimited code completions and next-edit suggestions, plus +**1,500 GitHub AI Credits/month** (1,000 base + 500 flex; 1 credit = $0.01) for chat, agent mode, code review, +cloud agent, and CLI. Pro models include Claude Haiku 4.5, Sonnet 4.6/5/5.5, GPT-5.4/5.6 Luna/6 Luna, Gemini 3.x +Flash, Grok 4.x, Kimi K3 — Opus/Fable and GPT-6 Sol/Astra need Pro+ or Max. +Copilot Pro+ is $39/month (7,000 credits); Copilot Max is $100/month (20,000 credits). Extra usage billed at $0.01/credit. + +Free tier: 2,000 completions/month plus a small AI-credit allowance, auto model selection only, no card. +Free Copilot Student for verified students; free Pro for verified teachers and maintainers of popular open-source repos. + +Tools: VS Code, JetBrains, Neovim, Xcode, Visual Studio, Copilot CLI, Copilot app, and the Copilot cloud agent on +github.com. + +*Checked Sep 29, 2026.* +``` + +What changed: +- The allowance changed from "1,500 premium requests" to 1,500 AI credits, which are usage-based. +- Pro+ went from 6,000 premium requests to 7,000 AI credits. +- A new Max tier costs $100/month. +- The Free tier's "50 premium requests" became an unspecified AI-credit allowance. +- Claude Opus-class models are no longer on Pro. +- I could not verify the "$100/year" annual price on the current docs or plans page, so it was dropped. Add it back only if the user confirms it. + +Sources: +- https://docs.github.com/en/copilot/get-started/plans (via `docs.github.com/api/article/body`) +- https://docs.github.com/en/copilot/concepts/billing-and-usage/individuals/billing +- https://github.com/features/copilot/plans + +--- + +## Proposed additions + +These are ranked by how well they fit this list: a cheap subscription that works in Claude Code or Cursor, with a +live official page. + +1. **Volcengine Ark — Coding Plan (方舟 Coding Plan, mainland China)** + - The official activity page is live. It lists Lite and Pro plans with models DeepSeek-V4.1-Flash, the GLM-5.3 series, Doubao-Seed-Evolving, Kimi-K3, and Kimi-K2.8-Preview. It supports Claude Code and Cursor, and advertises "限时 9.9 元起". + - The invite terms are a 5% voucher for the referrer and 9.5折 (5% off) for the friend. + - Volcengine's own marketing articles give ¥40/month for Lite and ¥200/month for Pro. Lite allows about 1,200 requests per 5 hours, 9,000 per week, and 18,000 per month. Pro is 5× Lite. + - The live price widget renders only with JavaScript, so re-check prices before adding. It is the mainland sibling of BytePlus. + - Sources: https://www.volcengine.com/activity/codingplan, https://www.volcengine.com/article/37898 +2. **Atlas Cloud — Coding Plan** + - A third-party aggregator with four tiers: Starter $10, Lite $20, Plus $50, and Max $100 per month. They give 16.5M, 33M, 82.5M, and 165M points per week. + - It covers 18 LLMs, including DeepSeek, GLM, Kimi, and MiniMax, and works with Claude Code, Codex, Cursor, and OpenClaw. + - It advertises "Full API access", which is rare for coding plans. + - Source: https://www.atlascloud.ai/coding-plan +3. **Alibaba Token Plan Personal Edition** + - Not a separate entry. It is already folded into the Alibaba block above. At $6/month it is now the cheapest Alibaba route. + +Checked and not recommended: +- **Cerebras Code** (Pro $50, Max $200, GLM-4.7): both tiers show "sold out" on https://www.cerebras.ai/code. +- **DeepSeek**: has no official coding plan. Only a prepaid API exists (https://qcode.cc/en/deepseek-code-plan, third party). +- **Tencent Cloud Coding Plan**: exists according to third-party trackers, but `cloud.tencent.com/act/pro/codingplan` returns 404. I found no verified official URL. +- **Mistral, Together, NVIDIA**: I did not verify any coding-subscription product for them. + +## Unresolved questions + +1. **opencode Go referral**: the program has officially ended. Should the blockquote be struck through or removed? That is the user's decision. +2. **BytePlus referral**: it ends Sep 30, 2026. Re-check on Oct 1 whether BytePlus extends it. +3. **MiniMax referral**: the "until Jul 1, 2026" text in the user's blockquote is stale, but the program is still documented. Should the user reword it? +4. **BigModel.cn prices** (¥118 / ¥538 / ¥1,078) come from media reports because the official price page needs JavaScript. The current referral percentages (5% vs 10%) are also unverified. +5. **Alibaba "$1,700" trial credits**: this figure was not found on the campaign page. +6. **Kimi USD pricing** and the current value of the Kimi referral rewards (the page needs JavaScript) are unverified. +7. **Copilot Pro annual price** ($100/year) is no longer on the official pages. +8. **Endpoints**: I did not verify the Alibaba Token Plan endpoints or the MiniMax annual discount. +9. **AgentKit**: the 20% buyer discount for `?ref=` links is unverified. Only the referrer commission tiers are visible. +10. **Claude guest passes**: only third parties confirm that they are "Max-only, three passes". Anthropic's docs say only "if your account is eligible". diff --git a/plans/reports/research-260929-1947-free-providers-a-audit.md b/plans/reports/research-260929-1947-free-providers-a-audit.md new file mode 100644 index 0000000..1e7259c --- /dev/null +++ b/plans/reports/research-260929-1947-free-providers-a-audit.md @@ -0,0 +1,354 @@ +# Free Providers audit (batch A): 13 entries, checked 2026-09-29 + +## Verdict + +| Entry | Status | One-line reason | +|---|---|---| +| TokenRouter | UPDATE | Kimi K3 free ended (the model ID now returns 404). GLM 5.3's free week ended Sep 4. One $0 model is left. | +| OrcaRouter | UPDATE | All five dated offers are gone. Three deposit-match campaigns replaced them, and there is now a rotating set of $0 models. | +| OpenRouter | UPDATE | Limits unchanged (20 RPM, 50 or 1,000 RPD). Add the current `:free` lineup. | +| NVIDIA NIM | UPDATE | The model list changed completely. The "1,000 credits" claim cannot be verified. | +| OpenCode Zen | UPDATE | The free lineup changed. DeepSeek V4 Flash Free is gone from the docs. | +| Google AI Studio | UPDATE | The free lineup is now Gemini 3.x Flash and Flash-Lite. Gemini 2.5 is closed to new users. Google no longer publishes the limits. | +| Kilo Code Gateway | UPDATE | Still 200 req/hour per IP. The $20 first-top-up bonus was not found; Kilo Pass replaced it. Kilo was acquired by Anaconda. | +| Cloudflare Workers AI | UPDATE | Still 10k Neurons/day. The model list changed, some models now need a paid plan, and the "~150 responses" figure is wrong. | +| DeepSeek Platform | UPDATE | The model and pricing structure changed. The 5M signup tokens are not confirmed in official docs. | +| Groq | UPDATE | Both Llama models are now Enterprise-only. The free models are GPT-OSS, Qwen3.8-27B and Whisper. | +| GitHub Models | **STALE** | The service was fully retired on Jul 30, 2026. | +| xAI Grok API | UNVERIFIED | Official docs mention no free credits. Third-party sources disagree. Anthropic compatibility is deprecated. | +| Mistral La Plateforme | UPDATE | "Experiment ~1B tokens/month" is replaced by a Free plan with **$10/month in API credits**. | + +Formatting note: TokenRouter and OrcaRouter used footnotes (`[^tokenrouter]: Check at ...`). The blocks below switch them to the `*Checked ...*` line that the other entries use, as requested. The CLAUDE.md footnote convention would then no longer be followed for these two entries. Your call. + +--- + +## TokenRouter — UPDATE + +```markdown +### [TokenRouter](https://www.tokenrouter.com/) + +Unified AI gateway (145 models listed) exposing OpenAI-, Anthropic- and Gemini-format APIs behind one key. + +- **Free model:** `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free` at $0/M input and output. +- **Launch promos:** TokenRouter runs short free windows on self-hosted launches (e.g. GLM 5.3 was free for all accounts, no card, until Sep 4, 2026). Watch the [blog](https://www.tokenrouter.com/blog). +- **API:** OpenAI Chat Completions at `https://api.tokenrouter.com/v1` with a TokenRouter API key. + +Source: [TokenRouter models](https://www.tokenrouter.com/models) + +*Checked Sep 29, 2026.* +``` + +What changed: +- `moonshotai/kimi-k3-free` no longer exists; its model page returns 404. Kimi K3 is now paid at $1.80/M input and $9.00/M output. +- GLM 5.3 was free for one week, ending Sep 4, 2026. That promo has also expired. +- The models page shows 145 models, not "300+". +- The home page now advertises "OpenAI, Claude, and Gemini compatible APIs". + +Sources: https://www.tokenrouter.com/models, https://www.tokenrouter.com/models/nvidia/nemotron-3-nano-omni-30b-a3b-reasoning%3Afree/, https://www.tokenrouter.com/models/moonshotai/kimi-k3/, https://www.tokenrouter.com/blog/tokenrouter-announces-day-0-availability-of-self-hosted-glm-5-3, https://www.tokenrouter.com/ + +## OrcaRouter — UPDATE + +```markdown +### [OrcaRouter](https://www.orcarouter.ai/) + +OpenAI-compatible gateway with access to 200+ models, automatic routing, and failover (also accepts Anthropic and Gemini formats). + +> Referral: (code: `ref_3976ba42abf37dc55c1d`). + +- **Always-free models ($0):** `deepseek/deepseek-v4-flash-free`, `z-ai/glm-5.3-flash-free`, `tencent/hy4-preview-free`, `tencent/hy3-free`, plus the `orcarouter/free` router. Requires a linked GitHub account with some history, or any paid purchase. +- **Hy4 Deposit Match:** 100% match on top-ups for `tencent/hy4-preview` (min $20 top-up, up to $200), enroll by **Oct 28, 2026**. +- **DeepSeek V4.1 Flash Deposit Match:** 30% match (min $20 top-up, up to $100), enroll by **Oct 2, 2026**. +- **GPT-6 Astra Deposit Match:** 100% match for `openai/gpt-6-astra` (min $20, up to $100), card required, enroll by **Oct 5, 2026**. +- **API:** `https://api.orcarouter.ai/v1`. + +Offers can change quickly; check the [live offers page](https://www.orcarouter.ai/offers) before claiming. + +*Checked Sep 29, 2026.* +``` + +What changed: +- All five previous offers are gone from the live offers API. That covers the Kimi K3 $5, Tencent HY3 $5, Claude Opus 5 60% match, DeepSeek 100/30 calls, and the Grok 4.5 waitlist. +- Three deposit-match campaigns replaced them. They have start and end dates, and the match credit expires between Nov 4 and Nov 30, 2026. +- The dollar caps are my derivation. The API reports `reward_cap_quota` together with `quota_per_unit: 500000`: 100M ÷ 500k = $200, and 50M ÷ 500k = $100. +- The FAQ now says the $0 models need account history: "the workspace owner links a GitHub account with some history ... or the workspace makes a paid purchase of any amount". +- `/v1/models` lists 204 models. The `orcarouter/free` router supports the openai, openai-response, anthropic and gemini endpoint types. + +Sources: https://www.orcarouter.ai/api/offers (JSON), https://api.orcarouter.ai/v1/models, https://www.orcarouter.ai/offers, https://www.orcarouter.ai/llms.txt + +## OpenRouter — UPDATE + +```markdown +### [OpenRouter](https://openrouter.ai) + +Free models (`:free` suffix): 20 RPM, 50 req/day; 1,000 req/day once you have bought at least $10 in credits (all time). BYOK requests are not gated by the free-model cap. + +Current free lineup includes NVIDIA Nemotron 3 Ultra 550B / Super 120B / 3.5 Lightning, Google Gemma 4 (26B-A4B, 31B), Qwen3.8-27B, Poolside Laguna S/XS 2.1, Thinking Machines Inkling / Inkling Small, Cohere North Mini Code, plus the `openrouter/free` auto-router. OpenAI-compatible at `https://openrouter.ai/api/v1`. + +*Checked Sep 29, 2026.* +``` + +What changed: +- The limits are the same. The docs constants are `FREE_MODEL_RATE_LIMIT_RPM=20`, `FREE_MODEL_NO_CREDITS_RPD=50`, `FREE_MODEL_HAS_CREDITS_RPD=1e3` and `FREE_MODEL_CREDITS_THRESHOLD=10`. +- The docs now say the tier depends on "credits purchased (all time)". An account with a negative balance can get 402 errors, even on free models. +- The models API returns 20 zero-priced entries out of 460 models. The model list is new information for this entry. + +Sources: https://openrouter.ai/docs/api/reference/limits, https://openrouter.ai/api/v1/models + +## NVIDIA NIM — UPDATE + +```markdown +### [NVIDIA NIM](https://build.nvidia.com) + +Free prototyping endpoints for NVIDIA Developer Program members, no credit card. Rate-limited per model (commonly ~40 RPM); NVIDIA does not publish a fixed quota and does not raise free-tier limits on request. + +Models: Kimi K3, Kimi K2.6, DeepSeek V4.1 Flash, GLM 5.3 / 5.3 Flash, Nemotron 3 Ultra / Super / 3.5 Lightning, GPT-OSS-20B, Gemma 4 31B, Mistral Large. OpenAI-compatible at `https://integrate.api.nvidia.com/v1`. + +**Warning:** Trial terms: use is logged and may be used to improve NVIDIA products — do not send personal or confidential data. + +*Checked Sep 29, 2026.* +``` + +What changed: +- The model list is new. The public `/v1/models` endpoint returns 81 models. Kimi K2.5, GPT-OSS-120B and DeepSeek-V3.2 are no longer listed, and there are no Llama 3.x chat models apart from the 3.2 vision models. +- The "1,000 inference credits at signup" claim is unverified. Third-party trackers say the credit system was removed. Forum users were still asking for credit increases in Aug–Sep 2026, and the Trial ToS PDF still mentions credits. I removed the number instead of guessing. +- The ~40 RPM figure comes from community reports and forum thread titles. It is not official. +- On Sep 28, 2026, a forum user relayed a moderator saying there is "no official way ... to receive a rate limit increase on that same tier". +- The privacy warning is new. The logging language comes from the NVIDIA trial terms, as quoted on the OpenCode Zen and Kilo docs pages. + +Sources: https://integrate.api.nvidia.com/v1/models, https://docs.api.nvidia.com/nim/docs/api-catalog-quickstart-guide, https://forums.developer.nvidia.com/t/rate-limit-increase-request-40-rpm-200-rpm-build-nvidia-com-free-tier/384546, https://assets.ngc.nvidia.com/products/api-catalog/legal/NVIDIA%20API%20Trial%20Terms%20of%20Service.pdf, https://yangmao.ai/en/providers/nvidia-build/, https://kilo.ai/docs/gateway/models-and-providers + +## OpenCode Zen — UPDATE + +```markdown +### [OpenCode Zen](https://opencode.ai/docs/zen) + +Hand-picked free models that change periodically (each is "free for a limited time"). Optimized for coding agents. + +| Model | Model ID | Notes | +|---|---|---| +| Big Pickle | `big-pickle` | Stealth model; data may be used to improve it | +| Space Bunny Free | `space-bunny-free` | Stealth model; zero-retention provider | +| LongCat 2.5 Preview Free | `longcat-2.5-preview-free` | Zero-retention provider | +| MiMo-V2.6-Flash Free | `mimo-v2.6-flash-free` | Data may be used to improve the model | +| MiMo-V2.5 Free | `mimo-v2.5-free` | Data may be used to improve the model | +| Ling 3.0 Flash Fin Free | `ling-3.0-flash-fin-free` | Data may be used to improve the model | +| Nemotron 3 Ultra Free | `nemotron-3-ultra-free` | NVIDIA trial endpoint; logged | +| Nemotron 3.5 Lightning Free | `nemotron-3.5-lightning-free` | NVIDIA trial endpoint; logged | +| Muse Spark 1.3 Contributor Free | `muse-spark-1.3-contributor-free` | Prompts used to train Meta models | + +OpenAI-compatible at `https://opencode.ai/zen/v1/chat/completions` (plus `/v1/responses`); Anthropic-format models use `https://opencode.ai/zen/v1/messages`. Model format in opencode: `opencode/`. + +*Checked Sep 29, 2026.* +``` + +What changed: +- DeepSeek V4 Flash Free is no longer in the docs' free list or pricing table. The ID `deepseek-v4-flash-free` still appears in `/zen/v1/models`, so it may just be unlisted. +- Six free models were added: Space Bunny, LongCat 2.5 Preview, MiMo-V2.6-Flash, Ling 3.0 Flash Fin, Nemotron 3.5 Lightning and Muse Spark 1.3 Contributor. +- The privacy notes are new, taken from the docs' privacy section. +- `jev-1.13-free` is also free. I left it out because it uses a specialised `/zen/v1/systemone` classification API, not chat. + +Sources: https://opencode.ai/docs/zen, https://opencode.ai/zen/v1/models + +## Google AI Studio — UPDATE + +```markdown +### [Google AI Studio](https://aistudio.google.com/) + +Google's developer platform for Gemini models. Free tier (no billing) with pay-as-you-go available. + +Free-tier models: Gemini 3.8 / 3.7 / 3.6 / 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.1 Flash-Lite, Gemini 3 Flash Preview, plus Live/TTS variants and Gemma 4. **Gemini 3.1 Pro Preview is paid-only.** Gemini 2.5 models are now limited to projects that already used them. + +Google no longer publishes a fixed free-tier table — limits are per project and shown in AI Studio. Reported Sep 2026: ~20 RPD on the 3.x Flash models, ~500 RPD on 3.5 / 3.1 Flash-Lite. RPD resets at midnight Pacific. + +OpenAI-compatible endpoint: `https://generativelanguage.googleapis.com/v1beta/openai/`. + +**Warning:** In the Free tier, Google may use your prompts and responses to improve their products. Use the Paid tier or Vertex AI for privacy. + +*Checked Sep 29, 2026.* +``` + +What changed: +- The 2.5 Pro/Flash/Flash-Lite RPM/RPD table is gone. The official rate-limits page (updated 2026-09-02) no longer lists free-tier numbers and says to "View your active rate limits in AI Studio". +- The official pricing page marks these as "Free of charge" on the free tier: 3.8, 3.7, 3.6 and 3.5 Flash, 3.5 Flash-Lite, 3.1 Flash-Lite, 3 Flash Preview, and 2.5 Pro/Flash/Flash-Lite. 3.1 Pro Preview is "Not available" on the free tier. +- Changelog, Sep 18, 2026: "limiting access to the 2.5 models to users who have actively used them in the past." New users should treat 3.x as the free lineup. +- The ~20 and ~500 RPD figures come from third-party pages that say they were read from AI Studio. They are not official. +- The privacy row still says "Used to improve our products: Yes (Free) / No (Paid)". + +Sources: https://ai.google.dev/gemini-api/docs/pricing, https://ai.google.dev/gemini-api/docs/rate-limits, https://ai.google.dev/gemini-api/docs/changelog, https://ai.google.dev/gemini-api/docs/openai, https://www.scriptbyai.com/gemini-api-free-tier-limits/, https://github.com/robhunter/agentdeals/issues/2017 + +## Kilo Code — Gateway — UPDATE + +```markdown +### [Kilo Code — Gateway](https://kilo.ai/gateway) + +VS Code + JetBrains coding extension (and CLI) with a built-in OpenAI-compatible gateway at `https://api.kilo.ai/api/gateway`. Kilo was acquired by Anaconda. + +Free: `:free` models and the `kilo-auto/free` router, 200 req/hour per IP (anonymous or signed in). Current free models include Nemotron 3 Ultra / Super / 3.5 Lightning, Qwen3.8-27B, Laguna S/XS 2.1, Inkling Small, North Mini Code, and Step 3.7 Flash. +BYOK supported with no Kilo markup. Kilo Pass ($19/$49/$199 per month) adds up to 50% bonus credits. + +**Warning:** Auto Free may route to providers that log prompts and use them for training (including NVIDIA trial endpoints). + +*Checked Sep 29, 2026.* +``` + +What changed: +- The "first top-up: $20 bonus credits (60-day expiry)" offer does not appear on the pricing or gateway pages. The bonus mechanism is now Kilo Pass. +- The base URL and the `kilo-auto/free` router are new information. The gateway models API returns 19 zero-priced entries. +- There is a new data-handling warning in the docs. +- There is an Anaconda acquisition banner on every page. + +Sources: https://kilo.ai/docs/gateway, https://kilo.ai/docs/gateway/models-and-providers, https://kilo.ai/docs/gateway/usage-and-billing, https://kilo.ai/pricing, https://api.kilo.ai/api/gateway/models + +## Cloudflare Workers AI — UPDATE + +```markdown +### [Cloudflare Workers AI](https://developers.cloudflare.com/workers-ai/) + +10,000 Neurons/day free on both Workers Free and Paid (resets 00:00 UTC). Roughly ~49K output tokens/day on Llama 3.3 70B or ~147K on GPT-OSS-120B. + +Models: GPT-OSS 120B/20B, Kimi K2.5, Llama 3.3 70B, Llama 4 Scout, Gemma 4 26B, Qwen3.8-27B, Nemotron 3 120B, GLM-4.7-Flash, Mistral Small 3.1, BGE embeddings, Whisper. Kimi K2.6/K2.7-Code, GLM 5.x and DeepSeek V4 require Workers Paid or AI Gateway credits. + +OpenAI-compatible at `https://api.cloudflare.com/client/v4/accounts//ai/v1`. + +*Checked Sep 29, 2026.* +``` + +What changed: +- The allocation is unchanged. +- The "~150 LLM responses/day" figure had no basis in the docs. I replaced it with token figures computed from the published neuron rates: + - Llama 3.3 70B uses 204,805 neurons per million output tokens, so 10k neurons buys about 48.8K output tokens. + - GPT-OSS-120B uses 68,182 neurons per million output tokens, so 10k neurons buys about 147K. +- There is a new note that some models need a paid billing method. +- The model list was refreshed. Qwen2.5-Coder, Gemma 3 and DeepSeek-R1-distill are still listed but are older. +- The OpenAI-compatible base URL is new information. + +Sources: https://developers.cloudflare.com/workers-ai/platform/pricing/, https://developers.cloudflare.com/workers-ai/configuration/open-ai-compatibility/ + +## DeepSeek Platform — UPDATE + +```markdown +### [DeepSeek Platform](https://platform.deepseek.com) + +New accounts are widely reported to get 5M free tokens (granted balance, ~30 days, phone verification); verify in your console. OpenAI (`https://api.deepseek.com`) + Anthropic (`https://api.deepseek.com/anthropic`) compatible. + +PAYG (off-peak / peak, per 1M tokens): `deepseek-flash` (V4.1 Flash) $0.15/$0.30 in, $0.60/$1.20 out; `deepseek-v4-pro` $0.66/$1.32 in, $1.98/$3.96 out. Off-peak is half price — everything outside 01:00–04:00 and 06:00–10:00 UTC on weekdays. 1M context. + +*Checked Sep 29, 2026.* +``` + +What changed: +- Model names changed. `deepseek-flash` is now DeepSeek-V4.1-Flash. The legacy `deepseek-v4-flash` name is still accepted but served by V4.1. +- Pro is now `DeepSeek-V4-Pro-0813`. +- Pricing moved to peak and off-peak rates. The old flat prices ($0.14/$0.28, $0.435/$0.87) are wrong. +- The official pricing page mentions a "granted balance" used before topped-up balance. It does not state a signup amount. The 5M-token figure is third-party only, so I added "verify in your console". + +Sources: https://api-docs.deepseek.com/quick_start/pricing, https://api-docs.deepseek.com/, https://dev.to/tokenmixai/i-burned-through-deepseeks-5m-free-tokens-in-14-days-heres-the-exact-math-3n22, https://aicredits.dev/submissions/24-deepseek-5-million-free-tokens-for-new-users + +## Groq — UPDATE + +```markdown +### [Groq](https://console.groq.com) + +LPU inference. Free plan, no credit card. OpenAI-compatible at `https://api.groq.com/openai/v1`. + +| Model | RPM | RPD | TPM | TPD | +|---|---|---|---|---| +| `openai/gpt-oss-120b` | 30 | 1K | 8K | 200K | +| `openai/gpt-oss-20b` | 30 | 1K | 8K | 200K | +| `qwen/qwen3.8-27b` | 30 | 1K | 8K | 200K | +| `whisper-large-v3` / `-turbo` | 20 | 2K | — | 7.2K audio-sec/hour | + +Llama 3.1 8B and Llama 3.3 70B are now Enterprise-only (contact sales). + +*Checked Sep 29, 2026.* +``` + +What changed: +- Both Llama models are gone from the Free Plan table. The models page lists them as "Enterprise / Contact Sales". +- GPT-OSS 120B/20B and Qwen3.8-27B are the free chat models now. +- The table also lists Orpheus TTS and Prompt Guard models, which I omitted as niche. +- I did not independently confirm "no credit card" in the docs. It is carried over from the old entry. + +Sources: https://console.groq.com/docs/rate-limits (and `.md` variant), https://console.groq.com/docs/models + +## GitHub Models — STALE + +This entry should be removed. There is no replacement block. GitHub points users to Microsoft/Azure AI Foundry, which is not free, and to GitHub Copilot, which the README already covers under Coding Plans. The Copilot Free tier is already mentioned there. + +What changed: +- Jun 16, 2026: closed to new customers. +- Jul 30, 2026: fully retired. The playground, catalog, inference API and BYOK are all gone. + +Sources: https://docs.github.com/en/github-models/use-github-models/prototyping-with-ai-models, https://github.blog/changelog/2026-07-30-github-models-is-now-retired/, https://github.blog/changelog/2026-07-01-github-models-is-being-fully-retired-on-july-30-2026/, https://github.blog/changelog/2026-06-16-github-models-is-no-longer-available-to-new-customers/ + +## xAI Grok API — UNVERIFIED (candidate for removal) + +Use this block only if you want to keep the entry. I recommend removal unless you can confirm credits in console.x.ai. + +```markdown +### [xAI Grok API](https://docs.x.ai) + +No free tier is documented. Promotional signup credits and a data-sharing credit program have been reported by third parties but are not in xAI's docs — check console.x.ai before relying on them. + +Models: Grok 4.7 (500K ctx, $2/$6 per 1M), Grok 4.3 (1M ctx, $1.25/$2.50), Grok Build 0.1 (coding, $1/$2). OpenAI-compatible (Responses API recommended); Anthropic SDK compatibility is deprecated. + +**Warning:** If offered, Data Sharing opt-in is irreversible and lets xAI train on your prompts. + +*Checked Sep 29, 2026.* +``` + +What changed: +- The official pricing, billing and full docs (`llms-full.txt`) never mention free, signup or data-sharing credits. Billing is prepaid credits or invoice only. +- Third-party sources disagree: + - aitoolsrecap (corrected Sep 5, 2026) says $25 signup plus $150/month for data sharing is still active. + - yangmao.ai (Jun 2026) says the data-sharing program should be treated as ended in May 2025. +- All three models in the README were retired on May 15, 2026: Grok 4 (`grok-4-0709`), Grok 4.1 Fast and Grok Code Fast. Their IDs now redirect to `grok-4.3` and `grok-build-0.1`. +- The docs now say "The Anthropic SDK compatibility is fully deprecated." +- The docs brand the company as "SpaceXAI (xAI)". + +Sources: https://docs.x.ai/developers/pricing.md, https://docs.x.ai/docs/models, https://docs.x.ai/developers/migration/may-15-retirement.md, https://docs.x.ai/llms-full.txt, https://aitoolsrecap.com/Blog/how-to-get-free-grok-api-key-2026-step-by-step, https://yangmao.ai/en/questions/grok-api-free-credits/ + +## Mistral La Plateforme — UPDATE + +```markdown +### [Mistral Studio (API)](https://console.mistral.ai) + +Free plan (default for new accounts) includes **$10/month in API credits**, shared across Studio, the API, and Vibe Code; rate limits shown in the Admin Console. Enable pay-as-you-go to continue past the allowance. Pro ($14.99/month) includes $15/month in API credits. + +Models: Mistral Medium 3.5, Mistral Large (2512), Mistral Small (2603), Devstral 2, Codestral, Ministral 3B/8B/14B, Voxtral, embeddings. OpenAI-compatible. + +**Warning:** Model training on your data is opt-out, not opt-in. + +*Checked Sep 29, 2026.* +``` + +What changed: +- The "Experiment plan, ~1B tokens/month" wording is outdated. The docs now describe "Free mode ... included monthly usage", and the pricing page says Free gets "$10 /mo in API credits". +- Plan structure: Free, Pro/Education, Team, Enterprise. Pro shows $15/month in API credits and Education $30/month. My reading of the rendered page may have swapped these two, so check the live page. +- The pricing table shows "Model training: Opt-out". +- Phone verification and "no credit card" were not confirmed in the current docs. +- I took the model names from OpenRouter's `mistralai/*` catalog because Mistral's model page did not render. Treat them as indicative. + +Sources: https://mistral.ai/pricing, https://docs.mistral.ai/admin/billing-usage/usage-limits.md, https://docs.mistral.ai/admin/billing-usage/subscriptions.md, https://openrouter.ai/api/v1/models, https://pricepertoken.com/endpoints/mistral/free + +--- + +## New free offers from these vendors worth noting + +- **OpenCode Zen:** six new free models; see the table above. +- **Google:** Gemini 3.5 Flash-Lite and 3.1 Flash-Lite have the most generous free quota Google offers (reported at about 500 RPD). +- **TokenRouter:** runs a recurring pattern of one-week free windows on self-hosted model launches. None is live today. +- **OrcaRouter:** now has four always-free chat models (DeepSeek V4 Flash, GLM 5.3 Flash, Hy4 preview, Hy3), plus a free router. +- **Kilo and OpenRouter:** both carry the Nemotron 3 family and Qwen3.8-27B for free. + +## Unresolved questions + +1. **xAI:** keep the entry with the "no documented free tier" block, or remove it? Only a logged-in console.x.ai check can confirm the credits. +2. **NVIDIA NIM:** are signup credits still issued, or is it purely rate-limited now? This needs a logged-in check at build.nvidia.com. +3. **Google AI Studio:** the exact free RPM/RPD per model is only visible in a logged-in AI Studio project. Also, should the README still mention Gemini 2.5 for existing users? +4. **DeepSeek:** the 5M signup-token grant is not in official docs. Keep it with "verify", or drop it? +5. **Mistral:** does Free still require phone verification and no card? Is the Pro vs Education credit amount ($15 vs $30) right? Can the model names be confirmed from a Mistral-owned page? +6. **Groq:** "no credit card" is not restated in the current docs. +7. **OpenCode Zen:** is `deepseek-v4-flash-free` still callable at $0? It is in `/models` but not in the docs. +8. **Formatting:** should TokenRouter and OrcaRouter keep the CLAUDE.md footnote format, or use the `*Checked ...*` line as done here? +9. **Mistral heading:** should the heading be renamed from "La Plateforme" to "Mistral Studio"? The docs no longer use "La Plateforme". I proposed the rename, but you can keep the old title. diff --git a/plans/reports/research-260929-1947-free-providers-b-audit.md b/plans/reports/research-260929-1947-free-providers-b-audit.md new file mode 100644 index 0000000..7e801cb --- /dev/null +++ b/plans/reports/research-260929-1947-free-providers-b-audit.md @@ -0,0 +1,450 @@ +# Free Providers audit (batch B): 13 entries + +Date: Sep 29, 2026. Scope: 13 entries under "Free Providers" in README.md, from Google Cloud Vertex AI through Ollama Cloud. README.md was not modified. +Method: official docs and pricing pages fetched with curl or WebFetch, Mintlify `.md` doc endpoints, and public `/v1/models` JSON where an endpoint exposes one. Third-party sources are used only where the official page gives no number, and each such use is labeled. + +## Summary + +| Entry | Status | Headline | +|---|---|---| +| Google Cloud Vertex AI | UPDATE | Renamed "Gemini Enterprise Agent Platform". The $300 credit cannot pay for partner MaaS models such as Claude. | +| Hugging Face Inference Providers | UPDATE | The free tier is **$0.10/month**, not "100K". PRO is $9/mo with **$2** of credits, not "2M". | +| Cerebras Cloud | UPDATE (close to STALE) | Cerebras no longer offers a permanent free tier. It now gives a **$5 / 30-day trial that requires a payment method**, and serves only 2 models. | +| BigModel.cn | UPDATE | GLM-4.5-Flash is gone from the price list. The free list is now GLM-4.7-Flash, GLM-4-Flash-250414, GLM-Z1-Flash, plus vision and image models. | +| Fireworks AI | UPDATE | Still $1. Adds a 10 RPM cap without a payment method and an Anthropic-compatible endpoint. | +| Scaleway Generative APIs | UPDATE | Still 1M tokens, plus 60 audio-minutes. The model catalogue changed. | +| SambaNova Cloud | UPDATE (conflicting sources) | The $5 credit is no longer advertised. The no-card tier is now **20 RPM / 20 RPD / 200K TPD** on 5 models. | +| Cohere | CURRENT (minor) | Same 1,000 calls/mo, 20 RPM, and non-commercial rule. The model list and the OpenAI-compat URL are updated. | +| Vercel AI Gateway | UPDATE | Official docs now say you must **add a payment method** to use the free credits. The docs no longer state the $5 figure. | +| Requesty | CURRENT (minor) | Still 200 req/day with no card. Adds base URLs. | +| SiliconFlow | UPDATE | The free list is longer: Qwen3-8B, Qwen3.5-4B, GLM-4-9B, GLM-Z1-9B, R1-0528-Qwen3-8B, and others. There is also an Anthropic endpoint. | +| ModelScope | UPDATE (quota UNVERIFIED on the official page) | Only 35 models are exposed now, not "50+". There is an Anthropic `/v1/messages` endpoint. | +| Ollama Cloud | UPDATE | Pricing changed on Aug 31, 2026. Session and weekly caps are gone. Free now means a monthly starter credit on starter models, with 1 concurrent request. | + +No entry is outright STALE, because every vendor still offers some free allowance. **Cerebras** comes closest. It no longer has a no-card free tier and now looks like Fireworks-style trial credit. If the list's bar is "free without a card", remove it. If trials qualify, keep it. + +--- + +## Google Cloud Vertex AI + +**Status: UPDATE** + +```markdown +### [Google Cloud Vertex AI](https://cloud.google.com/vertex-ai) + +Now branded **Gemini Enterprise Agent Platform** (formerly Vertex AI). + +Free Trial: **$300** credit for 90 days (new GCP customers only; card verification). Not a recurring free tier. The $300 credit **cannot** pay for partner models offered as MaaS (e.g. Claude) or for Gemini API in AI Studio. + +Express mode: `@gmail.com` accounts new to Google Cloud get a **90-day free tier with no billing info**, within express-mode quotas, on the APIs that support express mode (Gemini models). Separate from the $300 Free Trial. + +Models: Gemini 3.x (Pro / Flash / Flash-Lite); partner models (Claude and others) need a paid billing account. + +*Checked Sep 29, 2026.* +``` + +What changed: +- Google rebranded the product. The product page title is "Gemini Enterprise Agent Platform (formerly Vertex AI)". +- The free-features doc states the restriction directly: "You can't access or use the $300 credit for a generative AI partner model that is offered as a managed API ... model as a service." The old README line implied you could spend the credit on Claude, DeepSeek, GLM, and Qwen through MaaS. That was wrong for Claude and for any other partner MaaS model. +- Express mode is described as a 90-day free tier for new users with a @gmail.com account, with no billing information needed. It is separate from the Free Trial. +- The Gemini line now runs up to 3.8 Flash, according to the docs navigation. + +Sources: +- https://cloud.google.com/vertex-ai +- https://cloud.google.com/free/docs/free-cloud-features +- https://cloud.google.com/vertex-ai/generative-ai/docs/start/express-mode/overview + +--- + +## Hugging Face Inference Providers + +**Status: UPDATE** + +```markdown +### [Hugging Face Inference Providers](https://huggingface.co/docs/inference-providers) + +Router in front of 20 partners: Baseten, Cerebras, Cohere, DeepInfra, Fal AI, Featherless AI, Fireworks, Groq, HF Inference, Novita, Nscale, OVHcloud, Public AI, Replicate, Scaleway, Together, WaveSpeedAI, Z.ai and others. No markup over provider rates. + +Free: **$0.10/month** in credits (subject to change). PRO (**$9/month**): **$2.00/month** in credits, usable across all HF compute. Pay-as-you-go beyond that requires buying credits. + +OpenAI-compatible at `https://router.huggingface.co/v1` (~130 chat models). + +*Checked Sep 29, 2026.* +``` + +What changed: +- The credit figures were wrong. The official table shows "Free Users $0.10, subject to change" and "PRO Users $2.00". The "100K / 2M credits" wording in the README does not match HF's units. +- The partner list grew. The old list named 7 providers. The router `/v1/models` endpoint currently returns 133 models from featherless-ai, deepinfra, novita, nscale, zai-org, together, cohere, fireworks-ai, baseten, scaleway, ovhcloud, publicai, groq, and cerebras. +- The PRO price is unchanged at $9/month. + +Sources: +- https://huggingface.co/docs/inference-providers/pricing +- https://huggingface.co/docs/inference-providers/index +- https://huggingface.co/pricing +- https://router.huggingface.co/v1/models + +--- + +## Cerebras Cloud + +**Status: UPDATE (effectively no longer a free tier; candidate for removal if the list requires no-card free access)** + +```markdown +### [Cerebras Cloud](https://cloud.cerebras.ai) + +Wafer-scale chip inference. **Free Trial only**: **$5** in credits after adding a **verified payment method**, expiring 30 days after grant. No permanently free tier. OpenAI-compatible at `https://api.cerebras.ai/v1`. + +| Model | RPM | Uncached TPM | TPD | Context (free) | +|---|---|---|---|---| +| `gpt-oss-120b` | 5 | 30K | 1M | 65K | +| `qwen-3.8-27b` | 5 | 30K | 1M | 64K | + +Total TPM (incl. cached) is 3x uncached (90K). Access stops when credits run out or expire until you buy credits. + +*Checked Sep 29, 2026.* +``` + +What changed: +- The free tier is gone. The official FAQ, asked "Is there a permanently free tier?", answers "No. The Free Trial is time- and credit-bounded: $5 in credits that expire 30 days after they're granted." It also says: "If you skip adding a payment method at sign-up, Playground and API access remain inactive." +- The model lineup shrank to 2 shared models, `gpt-oss-120b` and `qwen-3.8-27b`. `llama3.1-8b-instant`, `qwen-3-235b-a22b-instruct-2507`, and `zai-glm-4.7` are no longer in the Shared Inference catalogue. +- The rate limits changed to 5 RPM per model, 30K uncached TPM, 1M TPH, and 1M TPD. +- The free context limit is 64–65K, not 8K. + +Sources: +- https://inference-docs.cerebras.ai/support/rate-limits +- https://inference-docs.cerebras.ai/models/overview +- https://inference-docs.cerebras.ai/resources/openai.md + +--- + +## BigModel.cn + +**Status: UPDATE** + +```markdown +### [BigModel.cn](https://www.bigmodel.cn/) + +Zhipu AI (智谱 AI). New users get a free token package to explore the API, playground, and AGI apps (reported as 20M, 25M via invite). + +Permanently free models: **GLM-4.7-Flash** (200K ctx), GLM-4-Flash-250414 (128K), GLM-Z1-Flash (128K), GLM-4.6V-Flash, GLM-4.1V-Thinking-Flash, GLM-4V-Flash, plus CogView-3-Flash (image) and CogVideoX-Flash (video). Concurrency limits apply per model. + +OpenAI-compatible at `https://open.bigmodel.cn/api/paas/v4`; Anthropic-compatible at `https://open.bigmodel.cn/api/anthropic`. + +> Referral: + +*Checked Sep 29, 2026.* +``` + +What changed: +- **GLM-4.5-Flash is no longer on the official price list.** The only "4.5" rows are GLM-4.5-Air and GLM-4.5V, and both are paid. +- The list of models priced "免费" (free) is longer than the README showed. The additions are listed in the block above. +- The new GLM-5.3-Flash is **paid**, at ¥0.8 input / ¥2.8 output per 1M tokens. It is not free. In the pricing table, "限时免费" (free for a limited time) applies only to the cache-storage column. +- Coding Plan subscribers get a free night-time window for GLM-5.3-Flash from Sep 3 to Oct 7, 2026. That offer belongs to the Coding Plan entry, not this one. +- The new-user token amount is **UNVERIFIED** on an official page. Third-party posts cite 20M tokens, with an extra 25M for invitees. The README's "25M" figure is plausible only for accounts that sign up by invite. + +Sources: +- https://docs.bigmodel.cn/cn/guide/start/pricing.md +- https://docs.bigmodel.cn/cn/api/rate-limit.md +- https://docs.bigmodel.cn/cn/coding-plan/notice/event-glm-5.3-flash.md +- New-user gift (third-party): https://github.com/x2v-co/aiplans/issues/16 +- New-user gift (third-party): https://blog.csdn.net/2402_82616859/article/details/146219601 +- Endpoints probed: `open.bigmodel.cn/api/paas/v4/models` and `open.bigmodel.cn/api/anthropic/v1/messages` both return auth errors, which means both routes exist. + +--- + +## Fireworks AI + +**Status: UPDATE** + +```markdown +### [Fireworks AI](https://fireworks.ai) + +**$1** in free starter credits for serverless inference. Without a payment method (or without credits) the account is capped at **10 RPM**; adding a payment method raises it up to 6,000 RPM. + +Models: GLM 5.3 / 5.3 Flash, Kimi K3, DeepSeek V4 Flash, Qwen 3.8 27B and other open models. Function calling, MCP support. + +OpenAI-compatible at `https://api.fireworks.ai/inference/v1`; Anthropic-compatible at `https://api.fireworks.ai/inference` (works with Claude Code). + +*Checked Sep 29, 2026.* +``` + +What changed: +- The official pricing page still says "Get started with $1 in free credits." +- The account quota page now documents a 10 RPM limit for "No payment method or no credits". +- Fireworks now documents an Anthropic Messages endpoint, including a Claude Code setup. +- The model names above come from the current serverless and training price tables. + +Sources: +- https://fireworks.ai/pricing +- https://docs.fireworks.ai/guides/quotas_usage/account-quotas.md +- https://docs.fireworks.ai/tools-sdks/anthropic-compatibility.md + +--- + +## Scaleway Generative APIs + +**Status: UPDATE** + +```markdown +### [Scaleway Generative APIs](https://www.scaleway.com/en/generative-apis/) + +EU/GDPR, Paris. Free tier for new customers: first **1,000,000 tokens** plus **60 minutes** of audio transcription (no time limit advertised). OpenAI-compatible at `https://api.scaleway.ai/v1`. + +Models: GLM-5.2, DeepSeek V4 Flash, Qwen3.8-27B, Qwen3.5-397B, Qwen3.6-35B, Qwen3-235B, Qwen3-Coder-30B, Gemma 4 26B, Mistral Medium 3.5, Mistral Small 3.2, gpt-oss-120b, Llama 3.3 70B, Pixtral 12B, Whisper. + +*Checked Sep 29, 2026.* +``` + +What changed: +- The free tier now also includes 60 audio minutes. According to the docs FAQ, the free tier is applied to the most expensive usage first. +- DeepSeek R1 distill is no longer listed. +- New models are GLM-5.2, DeepSeek V4 Flash, Qwen3.8/3.6/3.5, Gemma 4, Mistral Medium 3.5, and gpt-oss-120b. +- Whether a card is needed to *activate* the free tier is not stated on these pages. Scaleway accounts generally require a payment method, so treat this as UNVERIFIED. + +Sources: +- https://www.scaleway.com/en/generative-apis/ +- https://www.scaleway.com/en/pricing/model-as-a-service/ +- https://www.scaleway.com/en/docs/generative-apis/faq/ + +--- + +## SambaNova Cloud + +**Status: UPDATE (official sources conflict)** + +```markdown +### [SambaNova Cloud](https://cloud.sambanova.ai) + +RDU (dataflow chip) inference. **Free tier** applies when no payment method is linked: **20 RPM, 20 requests/day, 200K tokens/day** per model. Adding a card moves you to the Developer tier (pay-as-you-go, 60–240 RPM, 20M tokens/day across models). + +Free-tier models: DeepSeek-V3.1, DeepSeek-V3.2 (preview), Llama 3.3 70B, gpt-oss-120b, Gemma 4 31B (preview). + +OpenAI-compatible and Anthropic-compatible at `https://api.sambanova.ai/v1`. + +*Checked Sep 29, 2026.* +``` + +What changed: +- The **$5 / 30-day credit is no longer advertised.** The Plans page's "Free" card now reads "Add a payment method and purchase credits to run your first requests." That contradicts the rate-limits doc, which still has a Free Tier tab "applied when there is no payment method linked." +- The free-tier limits are now 20 **RPD**, down from the old per-model daily figure, with 200K TPD. +- The model lineup changed. Whisper, Llama 4, and Qwen3 are gone. The public `/v1/models` endpoint returns DeepSeek-V3.1, DeepSeek-V3.2, Meta-Llama-3.3-70B-Instruct, MiniMax-M2.7, MiniMax-M3, gemma-4-31B-it, and gpt-oss-120b. The MiniMax models appear only in the Developer tier table. +- SambaNova now documents an Anthropic-compatible endpoint. +- The blog link in the README describes the 2024 launch credit. That page says credits expire in 3 months, not 30 days, and it is no longer current evidence. + +Sources: +- https://docs.sambanova.ai/docs/en/models/rate-limits +- https://cloud.sambanova.ai/plans +- https://cloud.sambanova.ai/pricing +- https://api.sambanova.ai/v1/models +- https://docs.sambanova.ai/docs/en/features/anthropic-compatibility.md + +--- + +## Cohere + +**Status: CURRENT (minor refresh)** + +```markdown +### [Cohere](https://dashboard.cohere.com/api-keys) + +Trial API key, no credit card. **1,000 calls/month**, 20 RPM per chat model (Rerank 10 RPM, Embed 2,000 inputs/min). + +Models: Command A+, Command A Reasoning / Vision / Translate, Command A, Command R+, Command R, Command R7B, North Mini Code, Aya Expanse / Aya Vision, plus rerank and embeddings. + +OpenAI-compatible at `https://api.cohere.ai/compatibility/v1`. + +**Warning:** trial keys are **not permitted for production or commercial use** — production needs a production key (paid). New model variants (e.g. Command A Reasoning) stay at trial limits even on prod keys. + +*Checked Sep 29, 2026.* +``` + +What changed: +- There was no material change to the free terms. The pricing FAQ still says "trial keys are rate limited and are not permitted to be used for production or commercial purposes." +- The model list gains North Mini Code and the Command A variants. +- The OpenAI Compatibility API URL is added. + +Sources: +- https://docs.cohere.com/docs/rate-limits.md +- https://cohere.com/pricing +- https://docs.cohere.com/v2/docs/how-does-cohere-pricing-work.md +- https://docs.cohere.com/docs/compatibility-api.md +- https://docs.cohere.com/docs/models.md + +--- + +## Vercel AI Gateway + +**Status: UPDATE** + +```markdown +### [Vercel AI Gateway](https://vercel.com/docs/ai-gateway) + +Single endpoint routing to many providers, with failover and BYOK. OpenAI-compatible at `https://ai-gateway.vercel.sh/v1`; also Anthropic Messages, OpenResponses and Cohere-compatible APIs. + +Free: a **monthly included credit** (reported as **$5 / 30 days**; does not roll over), starting with your first request. **Requires a valid payment method on the team.** Covers only free-tier-eligible models with lower per-model rate limits. Buying credits moves you to the paid tier permanently and the monthly free credit stops. + +Source: + +*Checked Sep 29, 2026.* +``` + +What changed: +- **A card is now required.** The Getting Started page says: "To use free AI Gateway Credits, add a valid payment method to your team." The FAQ lists error `403 customer_verification_required` with the explanation "must add a valid payment method before using free credits". A third-party post from Aug 16, 2026 said no card was needed, but the official docs now say otherwise. +- The official pricing, rate-limit, and FAQ pages no longer state a dollar amount. They say only "monthly included credit". The $5 figure now rests on third-party sources. +- The docs now say explicitly that buying credits ends the free credit. +- Vercel now documents Anthropic, OpenResponses, and Cohere-compatible APIs. + +Sources: +- https://vercel.com/docs/ai-gateway/pricing +- https://vercel.com/docs/ai-gateway/rate-limits +- https://vercel.com/docs/ai-gateway/faq +- https://vercel.com/docs/ai-gateway/getting-started +- $5 figure (third-party): https://agentjournal.dev/blog/vercel-ai-gateway-free/ +- $5 figure (third-party): https://freeaiapi.org/articles/vercel-api-key-guide + +--- + +## Requesty + +**Status: CURRENT (minor refresh)** + +```markdown +### [Requesty](https://www.requesty.ai/) + +LLM gateway with routing, caching, spend controls and EU data residency. Works with Claude Code, Cline, Cursor, Roo. + +Free: **200 req/day** on the free-model catalogue. No credit card, no trial clock — same platform as pay-as-you-go (600+ models, +5% fee), just restricted to free models until you upgrade. + +OpenAI-compatible at `https://router.requesty.ai/v1`; Claude Code via `ANTHROPIC_BASE_URL=https://router.requesty.ai`. + +Source: , + +*Checked Sep 29, 2026.* +``` + +What changed: +- The free terms are unchanged. +- The block adds the base URLs, the 600+ paid-model count, and the 5% pay-as-you-go fee. +- The public `/v1/models` endpoint lists 12 models at $0. Most are NVIDIA Nemotron 3 variants, plus Gemma 4 31B, Poolside Laguna, and Mistral Leanstral. I did not add them to the README block, because $0 in the model list may not equal free-tier eligibility. + +Sources: +- https://www.requesty.ai/free-models +- https://www.requesty.ai/pricing +- https://router.requesty.ai/v1/models + +--- + +## SiliconFlow + +**Status: UPDATE** + +```markdown +### [SiliconFlow](https://cloud.siliconflow.cn/) + +Chinese multi-model inference platform, 200+ LLM/image/audio/video models. International site: . + +Free: a set of smaller open-source models is **permanently ¥0** with fixed rate limits (chat models from 1,000 RPM / 50K TPM); the international site gives **$1** starter credit. + +Models (free, CN site): Qwen3-8B (128K), Qwen3.5-4B (256K), Qwen2.5-7B, GLM-4-9B-0414, GLM-Z1-9B-0414, DeepSeek-R1-0528-Qwen3-8B, Hunyuan-MT-7B, Xing4.0-29B, plus OCR, ASR and BGE embedding/rerank models. + +OpenAI-compatible at `https://api.siliconflow.cn/v1`; Anthropic-compatible at `https://api.siliconflow.cn/` (Claude Code guide in docs). + +*Checked Sep 29, 2026.* +``` + +What changed: +- I built the free-model list from the price-"0" entries in the CN pricing page's embedded model data. +- The docs say "The Rate Limits for free models are fixed." The chat rate range starts at 1,000 RPM / 50,000 TPM, consistent with the README. +- SiliconFlow now documents an Anthropic-compatible endpoint. +- The "identity verification required" claim is **UNVERIFIED**. I could not reach a docs page that states it. + +Sources: +- https://www.siliconflow.cn/pricing +- https://www.siliconflow.com/pricing +- https://docs.siliconflow.com/en/userguide/rate-limits/rate-limit-and-upgradation +- https://docs.siliconflow.cn/cn/usercases/use-siliconcloud-in-ClaudeCode + +--- + +## ModelScope + +**Status: UPDATE (quota numbers UNVERIFIED against the official page, which renders only with JavaScript)** + +```markdown +### [ModelScope](https://modelscope.cn/) + +Alibaba's model community (魔搭). API-Inference free for registered users. + +Free: **2,000 req/day** total, **≤500 req/day per model** (some large models lower). Requires binding an Alibaba Cloud account with real-name verification. Quotas may be adjusted at any time. + +Models: ~35 API-Inference models, incl. DeepSeek V4 Pro / V4.1 Flash, Qwen3.5 / Qwen3.8, GLM-5.2, GLM-4.7-Flash, MiniMax-M3, Step-3.7-Flash, LongCat-Flash-Lite, Intern-S1. + +OpenAI-compatible at `https://api-inference.modelscope.cn/v1`; Anthropic-compatible at `https://api-inference.modelscope.cn` (`/v1/messages`). + +*Checked Sep 29, 2026.* +``` + +What changed: +- The public `/v1/models` endpoint returns **35** models, not "50+". +- An Anthropic `/v1/messages` route exists. It answered an empty POST with a schema validation error, not a 404. +- The 2,000/day and 500/model quota comes only from secondary sources: Cherry Studio docs and search snippets. The official limits page did not render without JavaScript. + +Sources: +- https://api-inference.modelscope.cn/v1/models +- https://modelscope.cn/docs/model-service/API-Inference/limits (JavaScript only; content not extracted) +- Secondary: https://github.com/CherryHQ/cherry-studio-docs/blob/main/pre-basic/providers/modelscope.md + +--- + +## Ollama Cloud + +**Status: UPDATE** + +```markdown +### [Ollama Cloud](https://ollama.com/cloud) + +Hosted counterpart to the local `ollama` CLI — same commands and API, models run on Ollama's GPUs. Prompts are not used for training. + +Free: a **starter amount of usage credits each month** (exact amount unpublished) on a smaller set of **starter models**, **1 concurrent request**, no credit card. Buying pay-as-you-go credits unlocks all cloud models. Session/weekly caps were removed on Aug 31, 2026. + +Models: DeepSeek V4, GLM-5.3, Kimi K3 / K2.7 Code, MiniMax M3, gpt-oss, Gemma 4, Nemotron 3, Mistral Large 3 (paid rates published per model). + +OpenAI-compatible at `https://ollama.com/v1`; Anthropic-compatible at `https://ollama.com/v1/messages`. + +*Checked Sep 29, 2026.* +``` + +What changed: +- Ollama switched to credit-based pricing on Aug 31, 2026. The blog post says: "The free plan now includes a small amount of monthly usage for a set of starter models ... no 5-hour or weekly limits." +- The pricing FAQ confirms free concurrency is 1 and that free usage resets monthly from the signup date. +- Neither the starter amount nor the list of starter models is published. +- The "16 model families" count was not re-verified. The block lists the families shown in the price table instead. + +Sources: +- https://ollama.com/pricing +- https://ollama.com/blog/transparent-pricing +- https://docs.ollama.com/cloud.md +- https://docs.ollama.com/api/openai-compatibility.md +- https://docs.ollama.com/api/anthropic-compatibility.md +- https://ollama.com/v1/models + +--- + +## New free offers from these vendors + +- **Vertex express mode.** This is the one notable no-billing option, a 90-day free tier for Gemini. It is already mentioned in the entry and is now described more precisely. +- **BigModel's night-time GLM-5.3-Flash window.** This runs Sep 3 to Oct 7, 2026. It is a Coding Plan perk, not a free API tier, so if it gets a mention it belongs in the "BigModel.cn — GLM Coding Plan" section. +- None of the other vendors launched a new free tier. Several moved the other way: + - Cerebras, SambaNova, and Vercel now require a card to get their free allowance. + - Ollama replaced its caps with credits. + +## Unresolved questions + +1. **Cerebras inclusion policy.** Cerebras now offers a card-gated 30-day trial. Should it stay in "Free Providers", as Fireworks and Vertex do, or be removed? +2. **SambaNova.** The Plans page says you must buy credits, while the rate-limit doc still describes a no-card Free Tier. Which one is live cannot be settled without signing up for an account. +3. **Vercel $5 amount.** The official docs no longer state the dollar figure. It rests on third-party sources dated August 2026. +4. **ModelScope quotas** (2,000/day, 500/model, real-name requirement). The official limits page requires JavaScript and could not be read. +5. **BigModel new-user token gift** (20M, or 25M via invite). This was not confirmed on an official page. +6. **Ollama starter credit amount and starter model list.** Neither is published. +7. **Vertex express-mode quotas** (per-model RPM/RPD). These were not extracted from the docs. +8. **Scaleway and SiliconFlow sign-up requirements** (card for Scaleway, identity verification for SiliconFlow). These were not confirmed on any page I could reach. diff --git a/plans/reports/research-260929-1947-free-providers-c-and-new.md b/plans/reports/research-260929-1947-free-providers-c-and-new.md new file mode 100644 index 0000000..0f2a97c --- /dev/null +++ b/plans/reports/research-260929-1947-free-providers-c-and-new.md @@ -0,0 +1,483 @@ +# Research: Free Providers (batch C) refresh + new free/cheap LLM APIs + +Date: 2026-09-29 (Asia/Saigon). Scope: 13 existing README entries (LongCat … Empero) plus discovery of new providers. +Method: official pages, official docs, and anonymous `GET /v1/models` calls (no prompts sent) made this session. +Secondary sources (the peter123023 and mnfst lists, blogs) were used only to find candidates, never as the sole basis for a figure. +`cheahjs/free-llm-api-resources` now returns **404** (the repo was deleted or made private), so it was not usable. + +## Summary + +| Entry | Status | One-line reason | +|---|---|---| +| LongCat (Meituan) | **STALE** | Official docs no longer mention a free daily quota; paid billing (token packs + PAYG) launched 2026-06-30 | +| SenseNova Token Plan | UPDATE | Free ¥0 beta confirmed; official page lists only SenseNova 6.8 Flash Lite + U1 Fast | +| AMD Token Factory | UPDATE | Free catalogue changed (8 free + 1 limited-free); daily quota figure not published publicly | +| Volcengine Ark | **STALE** | "2M tokens/day permanent free" is really a data-sharing rebate on *paid* usage; its term ends 2026-09-30 | +| Baidu Qianfan | **STALE** | ERNIE-Speed-128K / ERNIE-Lite-8K retired 2026-01-27 (official retirement doc) | +| iFlytek Spark | UPDATE | Lite still free; auth is `Bearer `, 8K in / 4K out | +| AIHubMix | UPDATE | Now 60 free models; after the $1 top-up the limits are 100 req/day, 10 req/min, 1M tokens/day | +| OVHcloud AI Endpoints | UPDATE | Anonymous 2 RPM confirmed; model list refreshed | +| LLM7.io | UPDATE | Official limits and the free (`turbo`) model list changed | +| Token Harbor | UPDATE | Free models now DeepSeek V4.1 Flash, MiMo V2.6 Flash, TH-Rudder; Agent Pass includes $10 usage | +| Aion Labs | UPDATE | Limits unchanged; aion-3.5 / 3.5-mini added | +| Experiential Labs | UPDATE | Free credits are 100 (= $1) per 30 days after a $1 card verification, not ~500 | +| Empero | **STALE** | `/v1/models` returns 503 `maintenance`: "The endpoint is offline for now" | + +Proposed additions, ranked by usefulness for coding agents: **NanoGPT Pro**, **Tencent TokenHub**, **Alibaba Model Studio free quota**, **Z.ai free Flash models**, **Nebius Token Factory**, **Novita free models**. + +--- + +## Existing entries + +### LongCat (Meituan) — STALE + +Evidence: +- The official quick-start, FAQ, token-pack and pricing pages contain no free daily quota. +- The FAQ mentions only "活动赠送额度" (promotional gift credit), with no amount given. +- The changelog entry for 2026-06-30 reads "全新推出计费服务" (billing service launched): Token 资源包 (one-time token packs, valid 30 days, sold in limited daily flash sales at 10:00/16:00/21:00/23:00 CST) plus API 按量付费 (pay-as-you-go). +- Introductory PAYG price is ¥2 / ¥8 per 1M tokens (in / out), or $0.30 / $1.20. Cache hits cost ¥0.04 / $0.006. +- Models are now **LongCat-2.5-Preview** (added 2026-09-25, multimodal) and **LongCat-2.0**, both with 1M context and 128K max output. +- The old free Flash models were retired on 2026-05-29. +- The endpoint is alive: `https://api.longcat.chat/openai/v1/models` returns 401. + +Recommendation: remove the entry from Free Providers. If you want to keep it as a cheap PAYG option instead, use this block: + +```markdown +### [LongCat (Meituan)](https://longcat.chat/platform/docs/zh/) + +Meituan's API platform for **LongCat-2.5-Preview** (multimodal) and **LongCat-2.0** (1M context, 128K max output). +**OpenAI** (`https://api.longcat.chat/openai`) and **Anthropic** (`https://api.longcat.chat/anthropic`) formats — works with Claude Code. + +No standing free tier since paid billing launched (Jun 30, 2026). Pay-as-you-go launch price **$0.30 in / $1.20 out per 1M** +(cache hits $0.006/M); 30-day token packs sold in limited daily drops. Failed requests (401/403/429/500) are not billed. +Mainland-China users need real-name verification before paying. + +Source: , + +*Checked Sep 29, 2026.* +``` + +Sources: , , , + +### SenseNova Token Plan — UPDATE + +What changed: +- The official Token Plan page confirms "公测期完全免费开放,付费档位即将上线" (free during the public beta; paid tiers coming soon). +- It shows **Free ¥0/month**, **60,000 credits / 5 hours**, SenseNova 6.8 Flash Lite and SenseNova U1 Fast, and up to 20 API keys. +- The **600,000 credits/week** cap and the third-party models (DeepSeek V4, GLM-5.2, Kimi K3) do not appear on the official page. They come only from the peter123023 list. +- The peter123023 list removed SenseNova as a DeepSeek V4 Flash channel on 2026-09-22. +- The endpoint is alive: `https://token.sensenova.cn/v1/models` returns 401. + +```markdown +### [SenseNova Token Plan](https://www.sensenova.cn/token-plan) + +SenseTime (商汤). Public beta — **Free tier ¥0/month**, phone-number signup, no card. +OpenAI-compatible at `https://token.sensenova.cn/v1` plus an Anthropic-compatible endpoint. + +Limits: **60,000 credits / 5 h** (special models excepted). Up to 20 API keys. + +Models: SenseNova 6.8 Flash Lite (multimodal agent model) and SenseNova U1 Fast. Third-party models +(DeepSeek, GLM, Kimi) have been reported at 0 credits but are not listed on the official plan page. + +**Warning:** SenseTime says paid Lite/Pro tiers are "coming soon" with no end date for the free beta — +don't build production on it. + +Source: + +*Checked Sep 29, 2026.* +``` + +Sources: , + +### AMD Token Factory (Radeon Cloud) — UPDATE + +What changed: +- The live catalogue section `public_free` ("Public Free Model APIs") lists these models with badge **Free**: MiMo-V2.6-Flash, DeepSeek-V4.1-Flash, DeepSeek-V4-Flash, DeepSeek-V4-Flash-Vision-Exp, GLM-5.3-Flash, Qwen3.8-Flash-Next, Qwen3.8-27B and MiniCPM5-2B. +- MinerU2.5-Pro is listed as **Limited Free**. +- MiniCPM5-1B is gone from the list. +- Each model card says "Free to use. Points show relative usage—not a charge." +- The "~$10/day" quota is **not published** on any public page I could reach; it comes from the peter123023 list only. +- The base URL `https://developer.amd.com.cn/radeon/api/v1` is confirmed by the model cards and returns 401 without a key. + +```markdown +### [AMD Token Factory (Radeon Cloud)](https://developer.amd.com.cn/radeon/tokenfactory) + +AMD's official inference platform on Radeon GPUs. Free shared endpoints after sign-in; usage is metered in +"points" against a daily allowance (reported as ~$10-equivalent/day, resets daily — not stated on the public page). +OpenAI-compatible at `https://developer.amd.com.cn/radeon/api/v1`. + +Free models: DeepSeek-V4.1-Flash, DeepSeek-V4-Flash (+ Vision-Exp), MiMo-V2.6-Flash, GLM-5.3-Flash, Qwen3.8-Flash-Next, +Qwen3.8-27B, MiniCPM5-2B (1M context on the DeepSeek/MiMo models). Marked "experimental" stability; high +time-to-first-token has been reported. + +Source: + +*Checked Sep 29, 2026.* +``` + +Sources: (its catalogue JSON loads from `/radeon/api/tokenfactory/bootstrap`) + +### Volcengine Ark (ByteDance) — STALE + +Evidence from the official "协作奖励计划" (collaboration reward plan) doc, updated 2026-09-21: +- The plan is not a free tier. When you authorise data collection, Ark returns up to **5M tokens/model/day of your *paid* usage** the next day, as a free resource pack valid 30 days. +- You must have a real-name verified account with the no-charge 安心体验 (safe-trial) mode switched off. That means pay-as-you-go billing is on. +- The collected data "将授权提供给火山引擎用于…模型和算法优化" (is provided to Volcengine for model and algorithm optimisation), and Volcengine "可永久使用" (may use it permanently). +- The plan's term is "延长至2026年9月30日" (extended to 2026-09-30). Some models exited on 2026-09-01, and more exit on 2026-10-08. +- The entry point currently supports only Doubao-Seed-Evolving. +- What remains is the one-time new-user 免费推理额度 (free inference quota), counted per model. The amount appears only in the console; the doc's 500K figure is an illustrative example. +- "Doubao-Lite / DeepSeek R2 free within 2M/day" is not supported by any official page. +- The endpoint is alive: `https://ark.cn-beijing.volces.com/api/v3/models` returns 401. + +Recommendation: remove the entry. The only thing left is an unquantified one-time trial, which needs mainland real-name verification. + +Sources: , + +### Baidu Qianfan — STALE + +Evidence: Baidu's official model retirement doc lists **ERNIE-Speed-128K, ERNIE-Lite-8K and ERNIE-Tiny-8K as retired on 2026-01-27** ("no longer available for use"). It recommends DeepSeek-V3, ERNIE-Speed-Pro-128K or ERNIE-4.5-Turbo-32K instead, none of which it describes as free. I found no official page naming a current permanently free model. The endpoint `qianfan.baidubce.com/v2/models` returns 403 without auth. + +Recommendation: remove the entry. + +Sources: + +### iFlytek Spark — UPDATE + +What changed: +- The official HTTP doc still says Lite "支持**免费使用**" (free to use), with **8K max input and 4K max output** (model id `lite`). +- Auth is `Authorization: Bearer ` from the console, not APIKey/APISecret. +- The doc says "兼容openAI SDK" with base_url `https://spark-api-open.xf-yun.com/v1/`. +- Neither "unlimited tokens" nor "2 QPS" appears in the official doc. The doc only mentions per-second and concurrency throttle error codes (11202/11203). +- The endpoint is alive: `/v1/models` returns 401. + +```markdown +### [iFlytek Spark](https://xinghuo.xfyun.cn/sparkapi) + +讯飞星火. **Spark Lite is free** (model id `lite`, 8K input / 4K output), throttled by per-second and concurrency limits +(third-party reports: 2 QPS, no token cap). OpenAI SDK-compatible at `https://spark-api-open.xf-yun.com/v1` +with `Authorization: Bearer ` from the console. Individual real-name verification required. + +Source: + +*Checked Sep 29, 2026.* +``` + +### AIHubMix — UPDATE + +What changed: +- The official free page (updated 2026-09-28) lists **60 free models** from 16 authors, all at $0. +- Every model speaks Chat Completions, Messages and Responses. +- New accounts get 10 trial calls, no card needed, and the calls never expire. +- A **one-time top-up of $1 or more** switches every free model to daily quotas: **100 req/day, 10 req/min, 1M tokens/day**, reset daily. +- Model IDs carry a `-free` suffix. + +```markdown +### [AIHubMix](https://aihubmix.com/models/free) + +Gateway with **60 free models**, no credit card. Every free model speaks Chat Completions, Messages, and Responses at +`https://aihubmix.com/v1` (model IDs end in `-free`). + +Free: 10 trial calls at signup (never expire). A **one-time top-up of $1+** permanently unlocks +**100 req/day, 10 req/min, 1M tokens/day** on the free catalogue (resets daily). + +Free models include coding-glm-5.3(-flash), coding-kimi-k3, coding-minimax-m3, xiaomi-mimo-v2.6-pro/flash, mimo-v2.5-pro, +gpt-5.5, gemini-3.8-flash, qwen3.6-plus-preview, hy3, nemotron-3-ultra/super, gemma-4-31b-it, gpt-oss-20b, glm-4.7-flash. + +Source: + +*Checked Sep 29, 2026.* +``` + +### OVHcloud AI Endpoints — UPDATE + +What changed: +- The official getting-started doc says anonymous use is limited to "2 requests per minute, per IP and per model". With an API key the limit is "400 requests per minute, per PCI project and per model", billed. +- The live `GET https://oai.endpoints.kepler.ai.cloud.ovh.net/v1/models` (no key) returns 24 models. +- Qwen3.8-27B and Qwen3.5-9B are new. The chat models are Qwen3.5-397B-A17B, Qwen3.8-27B, Qwen3.6-27B, Qwen3.5-9B, Qwen3-Coder-30B-A3B, gpt-oss-120b/20b, Meta-Llama-3_3-70B, Mistral-Small-3.2-24B, Mistral-Nemo, Mistral-7B and Qwen2.5-VL-72B. +- The catalogue shows non-zero per-token prices; those apply to keyed use. +- A 2026-09-28 third-party report (xibodev/llmgw-core#40) found that anonymous chat calls to those models still work and return 429 when the quota runs out. + +```markdown +### [OVHcloud AI Endpoints](https://www.ovhcloud.com/en/public-cloud/ai-endpoints/catalog/) + +EU-hosted (France). **Anonymous free tier — no API key, no signup**: 2 RPM per IP per model +(with a key: 400 RPM per project per model, pay-per-token). OpenAI SDK-compatible at +`https://oai.endpoints.kepler.ai.cloud.ovh.net/v1`. + +Models: Qwen3.5-397B-A17B, Qwen3.8-27B, Qwen3.6-27B, Qwen3-Coder-30B-A3B, gpt-oss-120b/20b, Llama 3.3 70B, +Mistral Small 3.2, Mistral Nemo, Qwen2.5-VL-72B, plus Whisper, embeddings, and TTS. + +Source: + +*Checked Sep 29, 2026.* +``` + +### LLM7.io — UPDATE + +What changed: +- The official limits doc gives these limits: + - Anonymous: 1/s, 10/min, 60/hour, and **500K tokens per 24 h**. + - Free token: 2/s, 40/min, 100/hour, and **1M tokens per 24 h**. + - Pro: $12/month. +- Only `turbo`-tier models are open to anonymous and free-token users. +- The live `/v1/models` lists four turbo models: DeepSeek-V4-Flash-0731 (400K context, 71.5% availability in the last hour), codestral-latest, minimax-m2.7 (180K) and mistral-Nemo-Instruct-2407. +- gpt-oss-20b is no longer listed. +- I found nothing official supporting the "UK" location claim, so the proposed block drops it. + +```markdown +### [LLM7.io](https://token.llm7.io) + +Gateway with keyless access. Anonymous: 10 req/min, 60 req/hour, **500K tokens/day**; a free token from +`token.llm7.io` raises this to 40 req/min, 100 req/hour, **1M tokens/day**. OpenAI-compatible at `https://api.llm7.io/v1`. + +Free (`turbo`-tier) models: DeepSeek-V4-Flash-0731 (400K ctx), codestral-latest, minimax-m2.7 (180K ctx), Mistral Nemo. +Everything else needs Pro ($12/month) or a topped-up balance. + +Source: + +*Checked Sep 29, 2026.* +``` + +### Token Harbor — UPDATE + +What changed: +- The official pricing page describes the Free plan as $0/month with a monthly allowance counted in 4-week windows. The amount is still unpublished, and the allowance carries over. +- Free models are now **DeepSeek V4.1 Flash, MiMo V2.6 Flash and TH-Rudder**. +- Agent Pass costs $1.99/month ($0.99 for the first month). It includes **$10 of usage**, "boosted" to up to $20 on selected models, and adds GLM 5.3 Flash, GPT-6 Luna and Qwen3.8 Flash. +- Free models are opt-in, and while enabled Token Harbor "may retain those prompts and responses for … model or product improvement". +- The terms still restrict access from some jurisdictions. The specific CN/HK/Macau `region_blocked` wording was not re-verified. +- The endpoint is alive: `/v1/models` returns 401. + +```markdown +### [Token Harbor](https://tokenharbor.ai/pricing) + +Small gateway. **Free tier $0/month** with a rolling allowance (amount unpublished, unused allowance carries over). +Agent Pass **$1.99/month** ($0.99 first month) adds $10 of included usage. OpenAI-compatible at `https://tokenharbor.ai/v1`. + +Free models: DeepSeek V4.1 Flash, MiMo V2.6 Flash, TH-Rudder ("promotional models added over time"). + +**Warning:** free models are opt-in and their prompts/responses may be retained for diagnostics and model +improvement. Access is restricted in some regions (reported: mainland China, Hong Kong, Macau). + +Source: + +*Checked Sep 29, 2026.* +``` + +### Aion Labs — UPDATE (minor) + +What changed: +- The official rate-limits doc still gives the Free tier **15 RPM, 20,000 TPM and a 20,000-token daily limit**. +- Any top-up moves the account to Tier 1: 50 RPM, 1M TPM, and no daily limit. +- Public `/v1/models` now also lists **aion-3.5** and **aion-3.5-mini** (256K context, released 2026-09-23). +- The pricing page says "A daily credit allowance… No card required." + +```markdown +### [Aion Labs](https://www.aionlabs.ai/app/api-keys/) + +Permanent free tier, no credit card. **15 RPM, 20K tokens/day**. OpenAI-compatible at `https://api.aionlabs.ai/v1`. + +Models: aion-3.5 / aion-3.5-mini (256K ctx), aion-3.0 / aion-3.0-mini, aion-2.0 (128K ctx, reasoning), +aion-rp-llama-3.1-8b. Tuned for roleplay/storytelling rather than coding. + +Source: + +*Checked Sep 29, 2026.* +``` + +### Experiential Labs — UPDATE + +What changed: +- The official billing doc sets "One credit is $0.01". +- The Free plan's recurring benefit "starts with the card verification (a one-time $1 charge, credited to your balance)". After that, "the total balance replenishes up to **100 credits**" every 30 days. That is $1/month, not ~500 credits. +- The live model page shows three models at $0: **GPT-6 Luna** ("100% off"), **MiMo-V2.6-Pro** ("Free", ZDR) and **Jev**. Qwen3.8 27B and DeepSeek V4 Flash are "75% off", not free. +- The docs say captured prompts are governed by an org-wide switch, and ZDR routes are never captured. I did not find an official statement of the claim that free traffic is exchanged for training traces. +- Anthropic API and Responses are supported. The endpoint `/v1/models` returns 401. + +```markdown +### [Experiential Labs](https://platform.experientiallabs.ai) + +"Open-source OpenRouter" gateway. Free plan: after a one-time **$1 card verification** (credited to your balance), the +balance **refills to 100 credits ($1) every 30 days**. OpenAI-compatible (Chat + Responses) and Anthropic-compatible at +`https://api.experientiallabs.ai/v1`. + +$0 models right now: GPT-6 Luna (promo, 100% off), MiMo-V2.6-Pro (free, zero-data-retention route), Jev. +Qwen3.8 27B and DeepSeek V4 Flash are 75% off. + +**Warning:** prompt/response capture is controlled by an org-wide switch — turn it off or require ZDR routes for private +code. Free-promo uptime varies by model (e.g. MiMo-V2.6-Pro ~76%). Prototyping only. + +Source: , + +*Checked Sep 29, 2026.* +``` + +### Empero — STALE + +Evidence: `GET https://free.empero.org/v1/models` returns **HTTP 503** with `{"code":"maintenance","message":"The endpoint is offline for now and will be back soon…"}`. The landing page shows "00 — MAINTENANCE Qwen3.8-27B-FP8 … Quick restart in progress". The lab's site (empero.org) is alive, but the free API is down. + +Recommendation: remove the entry, or strike it through with an "offline since at least Sep 29, 2026" note, and re-check later. + +Sources: , + +--- + +## Proposed additions + +The ranking weighs coding-agent value: model strength, Anthropic or OpenAI compatibility, the size of the quota, and how much friction it takes to get a key. Each block states which README section it belongs in. + +### 1. NanoGPT — Pro subscription (section: Providers with Coding Plans) + +This is the strongest cheap coding option found. The model list is broad and includes current coding models, and the API speaks both OpenAI and Anthropic `/messages`. The live subscription catalogue (`/api/subscription/v1/models`) returns 292 models. + +```markdown +### [NanoGPT](https://nano-gpt.com/pricing) + +Pay-as-you-go gateway for every major model, plus an optional **Pro subscription: $12/month — 60 million included input +tokens per week** on subscription models (web + API), and 5% off eligible paid text models. + +Subscription models include GLM-5.3 / 5.3-Flash, Kimi K2.6 / K2.7 Code, MiniMax M3 / M2.7, DeepSeek V4 Flash / V4 Pro, +MiMo-V2.5(-Pro), Qwen3.8-27B, Nemotron 3 Ultra. OpenAI-compatible at `https://api.nano-gpt.com/api/v1` +(also `/messages` and `/responses`); use `https://api.nano-gpt.com/api/subscription/v1` to keep requests on the +subscription only. Pay-as-you-go deposits start at $1 (card). + +Source: , + +*Checked Sep 29, 2026.* +``` + +### 2. Tencent TokenHub (section: Free Providers; the Token Plan could also be listed under Coding Plans) + +The free offer is a one-time trial rather than a recurring free tier. It is still large: 1M tokens per model, valid for one year. The paid Hy Token Plan starts at ¥28/month and covers GLM-5.3, Kimi K3 and DeepSeek V4 Pro, over both OpenAI and Anthropic endpoints. TokenHub replaces the old Hunyuan platform, which shuts down 2026-09-30. + +```markdown +### [Tencent TokenHub](https://cloud.tencent.com/document/product/1823/130053) + +Tencent Cloud's model platform (replaces the old Hunyuan platform, which shuts down Sep 30, 2026). + +- **Free trial:** **1M tokens per language/multimodal model**, claimed once per main account per model from the + Model Square "新用户福利" button; valid **1 year** from claim. Claim window ends **Dec 31, 2026**. +- **Token Plan (monthly):** Hy plan from **¥28** (Lite, 560 pts) to ¥468; universal plan ¥39–¥599. Models: DeepSeek-V4-Flash/Pro, + MiniMax-M2.7/M3, GLM-5.2/5.3/5.3-Flash, Kimi K2.7 Code, Kimi K3 (+ Hy3, Hy4 preview on the Hy plan). +- Plan endpoints: OpenAI `https://api.lkeap.cloud.tencent.com/plan/v3`, Anthropic `https://api.lkeap.cloud.tencent.com/plan/anthropic`. + +Which models qualify for the free trial is shown in the console. Real-name verification requirement: not confirmed. + +Source: , + +*Checked Sep 29, 2026.* +``` + +### 3. Alibaba Cloud Model Studio — new-user free quota (section: Free Providers, or a sub-heading under the existing Alibaba entry) + +The README already covers Alibaba's paid plans but not the free quota. The quota is large and per-model, so it effectively multiplies across the Qwen coder and plus models. It expires after 90 days. + +```markdown +#### [Model Studio free quota](https://www.alibabacloud.com/help/en/model-studio/new-free-quota) + +New users get **1,000,000 free tokens per model** (typical), valid **90 days** from activation (or model release / +approval, whichever is later). Singapore region, international deployment scope only; real-time inference only +(no batch/fine-tune). Each model — and each dated snapshot — has its own quota; RAM users share the account's pool. +OpenAI-compatible at `https://dashscope-intl.aliyuncs.com/compatible-mode/v1`. + +Source: + +*Checked Sep 29, 2026.* +``` + +### 4. Z.ai — free Flash models (section: Free Providers) + +This is the international counterpart of the BigModel.cn free models: GLM Flash models at $0 with no mainland ID needed. The models are smaller than GLM-5.x, but they cost nothing to use. + +```markdown +### [Z.ai API — free Flash models](https://docs.z.ai/guides/overview/pricing) + +Zhipu's international platform. **GLM-4.7-Flash**, **GLM-4.5-Flash** (text) and **GLM-4.6V-Flash** (vision) are priced +**Free** for input, cached input, and output. OpenAI-compatible at `https://api.z.ai/api/paas/v4`. + +Rate/concurrency limits for free models are shown only in the console (not published). + +Source: + +*Checked Sep 29, 2026.* +``` + +### 5. Nebius Token Factory (section: Free Providers) + +This gives a small trial plus a notably larger $25 program credit, and the catalogue has 60+ open models. It needs a bank card. + +```markdown +### [Nebius Token Factory](https://tokenfactory.nebius.com/) + +Formerly Nebius AI Studio; EU-based open-model inference. **$1 trial credit on first sign-up, valid 30 days**; joining the +free **Nebius Builder Program** adds a **$25 Token Factory credit** (open to everyone; meant for learning/testing). +OpenAI-compatible at `https://api.tokenfactory.nebius.com/v1`. + +**Warning:** setting up a billing account (bank card) is mandatory to finish onboarding. + +Source: , + +*Checked Sep 29, 2026.* +``` + +### 6. Novita AI — free models (section: Free Providers) + +The free models are real, but they are the weakest for coding and their limits are unpublished. It is listed for completeness. + +```markdown +### [Novita AI](https://novita.ai/pricing) + +Open-model inference platform. Two models are priced **$0 in / $0 out**: `inclusionai/ling-3.1-flash` and +`inclusionai/ling-3.0-flash-sante` (262K ctx). OpenAI-compatible at `https://api.novita.ai/openai`. + +Rate limits for the free models and any signup credit: not published on the pricing page. + +Source: + +*Checked Sep 29, 2026.* +``` + +### Candidates rejected (verified this session) + +| Candidate | Reason | +|---|---| +| Together AI | Official docs: "does not currently offer free trials"; $5 minimum purchase | +| Kimi / Moonshot platform | No free credits; $1 minimum recharge, $5 voucher after $5 cumulative; Tier0 = 3 RPM, 1 concurrent request | +| Featherless | No free tier; plans from $25/month | +| Chutes | Plus $10 / Pro $20 give a "daily quota", but the page does not state the amount | +| Perplexity | API pricing page shows no free tier or subscriber credit | +| Upstage | Only 10 free Studio agent runs; no API credit stated | +| Inference.net | Pay-as-you-go plan has no signup credit; $50 credit only on the $250/month plan | +| Venice | Only a one-time $10 credit for Pro subscribers, or DIEM staking | +| Baseten | "New accounts come with credits"; no amount given | +| Hyperbolic | Docs are now GPU-rental focused; no LLM free credit found | +| Kluster | `api.kluster.ai` does not resolve (DNS failure) | +| iFlow | `apis.iflow.cn/v1/models` returns 404 | +| Pollinations | Live OpenAI-compatible API, but "Quest Pollen" free amounts are unpublished; legacy `pk_` keys are capped at 1 pollen/IP/hour | +| ZenMux | Lists `z-ai/glm-4.7-flash-free` and `z-ai/glm-4.6v-flash-free`; free-model limits undocumented (docs URL 404) | +| BazaarLink | 3 `:free` models at $0 (`auto:free`, qwen3.7-flash, deepseek-v4-flash-0731); allowance unpublished | +| Parasail, Targon, Crofai | Pricing pages gave no extractable free-tier text | +| Poe API | Docs return 403 to plain HTTP; could not verify | +| Codestral free endpoint | No official page confirming that `codestral.mistral.ai` is still free (only secondary blogs) | +| Cloudflare AI Gateway | A gateway only; no inference credit of its own | +| Onomeo | Daily check-in scheme described only by a secondary source; not verified | + +Not re-checked this session: DeepInfra, Lambda, Crusoe. + +--- + +## Flags on other README entries (outside this batch, not changed) + +- **Mistral La Plateforme:** the official pricing page now says the Free plan has "**$10 /mo in API credits**" with a training opt-out. The README's "~1B tokens/month" is out of date. +- **Vercel AI Gateway, NVIDIA NIM:** the peter123023 weekly audit of 2026-09-27 says a Vercel promo ended and that the NIM model lineup changed (DeepSeek V4 Flash, V4 Pro and MiniMax M3 delisted; GLM-5.3 and V4.1 Flash added). Not verified here. +- **Cerebras, GitHub Models:** a secondary search snippet claimed the Cerebras free tier returns 402 and GitHub Models returns 410. The OpenRouter comparison article (updated 2026-09-24) still lists both as active. This session, `models.github.ai/catalog/models` returned HTTP 200 with body `OK`, and `api.cerebras.ai/v1/models` returned 403 unauthenticated. **Inconclusive**, so re-verify both with real keys. + +## Unresolved questions + +1. **LongCat:** new accounts may still get a promotional gift quota ("活动赠送额度"). It is visible only after signup, so the STALE verdict rests on what the documentation leaves out. +2. **AMD Token Factory:** what the daily points allowance is, and whether it still equals ~$10. It is visible only after login. +3. **SenseNova:** whether the 600K/week cap and the 0-credit third-party models (DeepSeek V4, GLM-5.2, Kimi K3) are still current. Only the console or docs behind JavaScript show this. +4. **Z.ai free Flash models and Novita free models:** the rate and concurrency limits are console-only. +5. **Tencent TokenHub:** which language models qualify for the 1M-token trial, and whether real-name verification is needed to claim it. A developer-community article says the validity is 90 days, while the official doc says 1 year; the report uses the official doc. +6. **Empero:** whether the maintenance ends soon enough to keep the entry with a warning instead of removing it. +7. **Token Harbor:** whether the exact `region_blocked` list is still CN, HK and Macau. The terms mention regional restrictions only in general terms.