diff --git a/README.md b/README.md index 91df468..d0acfa4 100644 --- a/README.md +++ b/README.md @@ -23,59 +23,59 @@ This lists various services that provide free access or credits towards API-base -OpenRouter20 requests/minute
50 requests/day
1000 requests/day with $10 credit balance
Bytedance UI Tars 72B -DeepCoder 14B Preview -DeepHermes 3 Llama 3 8B Preview -DeepSeek R1 -DeepSeek R1 Distill Llama 70B -DeepSeek R1 Distill Qwen 14B -DeepSeek R1 Distill Qwen 32B -DeepSeek R1 Zero -DeepSeek V3 -DeepSeek V3 0324 -DeepSeek V3 Base -Dolphin 3.0 Mistral 24B -Dolphin 3.0 R1 Mistral 24B -Featherless Qwerky 72B -Gemini 2.5 Pro Experimental 03-25 -Gemma 2 9B Instruct -Gemma 3 12B Instruct -Gemma 3 1B Instruct -Gemma 3 27B Instruct -Gemma 3 4B Instruct -Kimi VL A3B Thinking -Llama 3.1 8B Instruct -Llama 3.1 Nemotron 70B Instruct -Llama 3.1 Nemotron Nano 8B v1 -Llama 3.1 Nemotron Ultra 253B v1 -Llama 3.2 11B Vision Instruct -Llama 3.2 1B Instruct -Llama 3.2 3B Instruct -Llama 3.3 70B Instruct -Llama 3.3 Nemotron Super 49B v1 -Llama 4 Maverick -Llama 4 Scout -Mistral 7B Instruct -Mistral Nemo -Mistral Small 24B Instruct 2501 -Mistral Small 3.1 24B Instruct -Molmo 7B D -Moonlight-16B-A3B-Instruct -OlympicCoder 32B -OlympicCoder 7B -QwQ 32B ArliAI RpR v1 -Qwen 2.5 72B Instruct -Qwen 2.5 7B Instruct -Qwen 2.5 VL 32B Instruct -Qwen 2.5 VL 3B Instruct -Qwen 2.5 VL 7B Instruct -Qwen QwQ 32B -Qwen QwQ 32B Preview -Qwen2.5 Coder 32B Instruct -Qwen2.5 VL 72B Instruct -Reka Flash 3 -Rogue Rose 103B v0.2 -Zephyr 7B Beta +OpenRouter20 requests/minute
50 requests/day
1000 requests/day with $10 credit balance
Bytedance UI Tars 72BShared Quota +DeepCoder 14B Preview +DeepHermes 3 Llama 3 8B Preview +DeepSeek R1 +DeepSeek R1 Distill Llama 70B +DeepSeek R1 Distill Qwen 14B +DeepSeek R1 Distill Qwen 32B +DeepSeek R1 Zero +DeepSeek V3 +DeepSeek V3 0324 +DeepSeek V3 Base +Dolphin 3.0 Mistral 24B +Dolphin 3.0 R1 Mistral 24B +Featherless Qwerky 72B +Gemini 2.5 Pro Experimental 03-25 +Gemma 2 9B Instruct +Gemma 3 12B Instruct +Gemma 3 1B Instruct +Gemma 3 27B Instruct +Gemma 3 4B Instruct +Kimi VL A3B Thinking +Llama 3.1 8B Instruct +Llama 3.1 Nemotron 70B Instruct +Llama 3.1 Nemotron Nano 8B v1 +Llama 3.1 Nemotron Ultra 253B v1 +Llama 3.2 11B Vision Instruct +Llama 3.2 1B Instruct +Llama 3.2 3B Instruct +Llama 3.3 70B Instruct +Llama 3.3 Nemotron Super 49B v1 +Llama 4 Maverick +Llama 4 Scout +Mistral 7B Instruct +Mistral Nemo +Mistral Small 24B Instruct 2501 +Mistral Small 3.1 24B Instruct +Molmo 7B D +Moonlight-16B-A3B-Instruct +OlympicCoder 32B +OlympicCoder 7B +QwQ 32B ArliAI RpR v1 +Qwen 2.5 72B Instruct +Qwen 2.5 7B Instruct +Qwen 2.5 VL 32B Instruct +Qwen 2.5 VL 3B Instruct +Qwen 2.5 VL 7B Instruct +Qwen QwQ 32B +Qwen QwQ 32B Preview +Qwen2.5 Coder 32B Instruct +Qwen2.5 VL 72B Instruct +Reka Flash 3 +Rogue Rose 103B v0.2 +Zephyr 7B Beta Google AI StudioData is used for training (when used outside of the UK/CH/EEA/EU).Gemini 2.5 Pro (Experimental)5,000,000 tokens/day
1,000,000 tokens/minute
25 requests/day
5 requests/minute Gemini 2.0 Flash1,000,000 tokens/minute
1,500 requests/day
15 requests/minute Gemini 2.0 Flash-Lite1,000,000 tokens/minute
1,500 requests/day
30 requests/minute @@ -157,46 +157,18 @@ This lists various services that provide free access or credits towards API-base Mixtral 8x7B Instruct v0.112 requests/minute Qwen 2.5 VL 72B Instruct12 requests/minute Qwen2.5 Coder 32B Instruct12 requests/minute - - Together - - Llama 3.2 11B Vision Instruct - - - - Llama 3.3 70B Instruct - - - - DeepSeek R1 Distil Llama 70B - - - Cohere - 20 requests/minute
1,000 requests/month
- Command-R - Shared Limit - - - Command-R+ - - - Command-R7B - - - Command-A - - - Aya Expanse 8B - - - Aya Expanse 32B - - - Aya Vision 8B - - - Aya Vision 32B - GitHub ModelsExtremely restrictive input/output token limits.
Rate limits dependent on Copilot subscription tier (Free/Pro/Business/Enterprise)AI21 Jamba 1.5 Large +TogetherUp to 60 requests/minuteLlama 3.2 11B Vision Instruct +Llama 3.3 70B Instruct +DeepSeek R1 Distil Llama 70B +Cohere20 requests/minute
1,000 requests/month
Command-AShared Limit +Command-R7B +Command-R+ +Command-R +Aya Expanse 8B +Aya Expanse 32B +Aya Vision 8B +Aya Vision 32B +GitHub ModelsExtremely restrictive input/output token limits.
Rate limits dependent on Copilot subscription tier (Free/Pro/Business/Enterprise)AI21 Jamba 1.5 Large AI21 Jamba 1.5 Mini Codestral 25.01 Cohere Command R @@ -317,45 +289,17 @@ This lists various services that provide free access or credits towards API-base TinyLlama 1.1B Chat v1.0 Una Cybertron 7B v2 (BF16) Zephyr 7B Beta (AWQ) - - Google Cloud Vertex AI - Very stringent payment verification for Google Cloud. - Llama 4 Maverick Instruct - Llama 4 API Service free during preview.
60 requests/minute - - - Llama 4 Scout Instruct - Llama 4 API Service free during preview.
60 requests/minute - - - Llama 3.1 70B Instruct - Llama 3.1 API Service free during preview.
60 requests/minute - - - Llama 3.1 8B Instruct - Llama 3.1 API Service free during preview.
60 requests/minute - - - Llama 3.2 90B Vision Instruct - Llama 3.2 API Service free during preview.
30 requests/minute - - - Llama 3.3 70B Instruct - Llama 3.3 API Service free during preview.
30 requests/minute - - - Gemini 2.5 Pro Experimental - Experimental Gemini model.
10 requests/minute - - - Gemini 2.0 Flash Experimental - - - Gemini 2.0 Flash Thinking Experimental - - - Gemini 2.0 Pro Experimental - +Google Cloud Vertex AIVery stringent payment verification for Google Cloud.Gemini 2.5 Pro (Experimental)10 requests/minute
Shared Quota +Gemini 2.0 Flash (Experimental) +Gemini 2.0 Flash Thinking (Experimental) +Gemini 2.0 Pro (Experimental) +Llama 4 Maverick Instruct60 requests/minute
Free during preview +Llama 4 Scout Instruct60 requests/minute
Free during preview +Llama 3.3 70B Instruct30 requests/minute
Free during preview +Llama 3.2 90B Vision Instruct30 requests/minute
Free during preview +Llama 3.1 70B Instruct60 requests/minute
Free during preview +Llama 3.1 8B Instruct60 requests/minute
Free during preview + ## Providers with trial credits diff --git a/src/pull_available_models.py b/src/pull_available_models.py index 5138c56..42f8eee 100644 --- a/src/pull_available_models.py +++ b/src/pull_available_models.py @@ -581,26 +581,83 @@ def main(): table += f'{get_human_limits(model)}
1000 requests/day with $10 credit balance
' table += f"{model['name']}" - table += "" + if idx == 0: + table += f'Shared Quota' table += "\n" gemini_text_models = [ - {"id": "gemini-2.5-pro-exp-03-25", "name": "Gemini 2.5 Pro (Experimental)", "limits": gemini_models["gemini-2.0-pro-exp"]}, - {"id": "gemini-2.0-flash", "name": "Gemini 2.0 Flash", "limits": gemini_models["gemini-2.0-flash"]}, - {"id": "gemini-2.0-flash-lite", "name": "Gemini 2.0 Flash-Lite", "limits": gemini_models["gemini-2.0-flash-lite"]}, - {"id": "gemini-2.0-flash-exp", "name": "Gemini 2.0 Flash (Experimental)", "limits": gemini_models["gemini-2.0-flash-exp"]}, - {"id": "gemini-1.5-flash", "name": "Gemini 1.5 Flash", "limits": gemini_models["gemini-1.5-flash"]}, - {"id": "gemini-1.5-flash-8b", "name": "Gemini 1.5 Flash-8B", "limits": gemini_models["gemini-1.5-flash-8b"]}, - {"id": "gemini-1.5-pro", "name": "Gemini 1.5 Pro", "limits": gemini_models["gemini-1.5-pro"]}, - {"id": "learnlm-1.5-pro-experimental", "name": "LearnLM 1.5 Pro (Experimental)", "limits": gemini_models["learnlm-1.5-pro-experimental"]}, - {"id": "gemma-3-27b-it", "name": "Gemma 3 27B Instruct", "limits": gemini_models["gemma-3-27b"]}, - {"id": "gemma-3-12b-it", "name": "Gemma 3 12B Instruct", "limits": gemini_models["gemma-3-12b"]}, - {"id": "gemma-3-4b-it", "name": "Gemma 3 4B Instruct", "limits": gemini_models["gemma-3-4b"]}, - {"id": "gemma-3-1b-it", "name": "Gemma 3 1B Instruct", "limits": gemini_models["gemma-3-1b"]}, + { + "id": "gemini-2.5-pro-exp-03-25", + "name": "Gemini 2.5 Pro (Experimental)", + "limits": gemini_models["gemini-2.0-pro-exp"], + }, + { + "id": "gemini-2.0-flash", + "name": "Gemini 2.0 Flash", + "limits": gemini_models["gemini-2.0-flash"], + }, + { + "id": "gemini-2.0-flash-lite", + "name": "Gemini 2.0 Flash-Lite", + "limits": gemini_models["gemini-2.0-flash-lite"], + }, + { + "id": "gemini-2.0-flash-exp", + "name": "Gemini 2.0 Flash (Experimental)", + "limits": gemini_models["gemini-2.0-flash-exp"], + }, + { + "id": "gemini-1.5-flash", + "name": "Gemini 1.5 Flash", + "limits": gemini_models["gemini-1.5-flash"], + }, + { + "id": "gemini-1.5-flash-8b", + "name": "Gemini 1.5 Flash-8B", + "limits": gemini_models["gemini-1.5-flash-8b"], + }, + { + "id": "gemini-1.5-pro", + "name": "Gemini 1.5 Pro", + "limits": gemini_models["gemini-1.5-pro"], + }, + { + "id": "learnlm-1.5-pro-experimental", + "name": "LearnLM 1.5 Pro (Experimental)", + "limits": gemini_models["learnlm-1.5-pro-experimental"], + }, + { + "id": "gemma-3-27b-it", + "name": "Gemma 3 27B Instruct", + "limits": gemini_models["gemma-3-27b"], + }, + { + "id": "gemma-3-12b-it", + "name": "Gemma 3 12B Instruct", + "limits": gemini_models["gemma-3-12b"], + }, + { + "id": "gemma-3-4b-it", + "name": "Gemma 3 4B Instruct", + "limits": gemini_models["gemma-3-4b"], + }, + { + "id": "gemma-3-1b-it", + "name": "Gemma 3 1B Instruct", + "limits": gemini_models["gemma-3-1b"], + }, ] gemini_embedding_models = [ - {"id": "text-embedding-004", "name": "text-embedding-004", "limits": gemini_models["project-embedding"]}, - {"id": "embedding-001", "name": "embedding-001", "limits": gemini_models["project-embedding"]}, + { + "id": "text-embedding-004", + "name": "text-embedding-004", + "limits": gemini_models["project-embedding"], + }, + { + "id": "embedding-001", + "name": "embedding-001", + "limits": gemini_models["project-embedding"], + }, ] for idx, model in enumerate(gemini_text_models): @@ -688,48 +745,59 @@ def main(): table += f"{get_human_limits(model)}" table += "\n" - table += """ - Together - - Llama 3.2 11B Vision Instruct - - - - Llama 3.3 70B Instruct - - - - DeepSeek R1 Distil Llama 70B - - """ + together_models = [ + { + "id": "meta-llama/Llama-Vision-Free", + "name": "Llama 3.2 11B Vision Instruct", + "urlId": "llama-3-2-11b-free", + }, + { + "id": "llmeta-llama/Llama-3.3-70B-Instruct-Turbo-Free", + "name": "Llama 3.3 70B Instruct", + "urlId": "llama-3-3-70b-free", + }, + { + "id": "deepseek-ai/DeepSeek-R1-Distill-Llama-70B-free", + "name": "DeepSeek R1 Distil Llama 70B", + "urlId": "deepseek-r1-distilled-llama-70b-free", + }, + ] - table += """ - Cohere - 20 requests/minute
1,000 requests/month
- Command-R - Shared Limit - - - Command-R+ - - - Command-R7B - - - Command-A - - - Aya Expanse 8B - - - Aya Expanse 32B - - - Aya Vision 8B - - - Aya Vision 32B - """ + for idx, model in enumerate(together_models): + table += "" + if idx == 0: + table += f'' + table += 'Together' + table += "" + table += ( + f'Up to 60 requests/minute' + ) + table += f"{model['name']}" + table += f"{get_human_limits(model)}" + table += "\n" + + cohere_models = [ + {"id": "command-a-03-2025", "name": "Command-A"}, + {"id": "command-r7b-12-2024", "name": "Command-R7B"}, + {"id": "command-r-plus", "name": "Command-R+"}, + {"id": "command-r", "name": "Command-R"}, + {"id": "c4ai-aya-expanse-8b", "name": "Aya Expanse 8B"}, + {"id": "c4ai-aya-expanse-32b", "name": "Aya Expanse 32B"}, + {"id": "c4ai-aya-vision-8b", "name": "Aya Vision 8B"}, + {"id": "c4ai-aya-vision-32b", "name": "Aya Vision 32B"}, + ] + + for idx, model in enumerate(cohere_models): + table += "" + if idx == 0: + table += f'' + table += 'Cohere' + table += "" + table += f'20 requests/minute
1,000 requests/month
' + table += f"{model['name']}" + if idx == 0: + table += 'Shared Limit' + table += "\n" for idx, model in enumerate(github_models): table += "" @@ -775,45 +843,86 @@ def main(): table += "" table += "\n" - table += """ - Google Cloud Vertex AI - Very stringent payment verification for Google Cloud. - Llama 4 Maverick Instruct - Llama 4 API Service free during preview.
60 requests/minute - - - Llama 4 Scout Instruct - Llama 4 API Service free during preview.
60 requests/minute - - - Llama 3.1 70B Instruct - Llama 3.1 API Service free during preview.
60 requests/minute - - - Llama 3.1 8B Instruct - Llama 3.1 API Service free during preview.
60 requests/minute - - - Llama 3.2 90B Vision Instruct - Llama 3.2 API Service free during preview.
30 requests/minute - - - Llama 3.3 70B Instruct - Llama 3.3 API Service free during preview.
30 requests/minute - - - Gemini 2.5 Pro Experimental - Experimental Gemini model.
10 requests/minute - - - Gemini 2.0 Flash Experimental - - - Gemini 2.0 Flash Thinking Experimental - - - Gemini 2.0 Pro Experimental - """ + vertex_llama_models = [ + { + "id": "llama-4-maverick-17b-128e-instruct-maas", + "name": "Llama 4 Maverick Instruct", + "urlId": "llama-4-maverick-17b-128e-instruct-maas", + "limits": {"requests/minute": 60}, + }, + { + "id": "llama-4-scout-17b-16e-instruct-maas", + "name": "Llama 4 Scout Instruct", + "urlId": "llama-4-maverick-17b-128e-instruct-maas", + "limits": {"requests/minute": 60}, + }, + { + "id": "llama-3.3-70b-instruct-maas", + "name": "Llama 3.3 70B Instruct", + "urlId": "llama-3-3-70b-instruct-maas", + "limits": {"requests/minute": 30}, + }, + { + "id": "llama-3.2-90b-vision-instruct-maas", + "name": "Llama 3.2 90B Vision Instruct", + "urlId": "llama-3-2-90b-vision-instruct-maas", + "limits": {"requests/minute": 30}, + }, + { + "id": "llama-3.1-70b-instruct-maas", + "name": "Llama 3.1 70B Instruct", + "urlId": "llama-3-1-405b-instruct-maas", + "limits": {"requests/minute": 60}, + }, + { + "id": "llama-3.1-8b-instruct-maas", + "name": "Llama 3.1 8B Instruct", + "urlId": "llama-3-1-405b-instruct-maas", + "limits": {"requests/minute": 60}, + }, + ] + vertex_gemini_models = [ + { + "id": "gemini-2.5-pro-exp-03-25", + "name": "Gemini 2.5 Pro (Experimental)", + "limits": {"requests/minute": 10}, + }, + { + "id": "gemini-2.0-flash-exp", + "name": "Gemini 2.0 Flash (Experimental)", + "limits": {"requests/minute": 10}, + }, + { + "id": "gemini-2.0-flash-thinking-exp-01-21", + "name": "Gemini 2.0 Flash Thinking (Experimental)", + "limits": {"requests/minute": 10}, + }, + { + "id": "gemini-exp-1206", + "name": "Gemini 2.0 Pro (Experimental)", + "limits": {"requests/minute": 10}, + }, + ] + + for idx, model in enumerate(vertex_gemini_models): + table += "" + if idx == 0: + table += ( + f'' + ) + table += 'Google Cloud Vertex AI' + table += "" + table += f'Very stringent payment verification for Google Cloud.' + table += f'{model['name']}' + if idx == 0: + table += f"{get_human_limits(model)}
Shared Quota" + table += "\n" + + for idx, model in enumerate(vertex_llama_models): + table += "" + table += f"{model['name']}" + table += f"{get_human_limits(model)}
Free during preview" + table += "\n" table += ""