AI Platforms & Models
Every major foundation model — frontier APIs, open weights, and small efficient models — scored, priced, and compared.
Claude Sonnet 4
Frontier reasoning with best-in-class coding.
Claude 3.5 Sonnet
Long-form writing, analysis, coding
Gemini 2.5 Pro
Million-token context with deep thinking.
DeepSeek R1
Open-weight reasoning that rivals the frontier.
Llama 3.1 70B
Self-hosted AI, privacy-critical apps
Gemini 1.5 Pro
Long-context document analysis and multimodal reasoning
ElevenLabs Multilingual
Lifelike multilingual text-to-speech and voice cloning
Kimi K2
Open trillion-parameter agentic intelligence.
Whisper v3
Accurate speech recognition and transcription across languages
DeepSeek V3
Budget-friendly frontier performance for coding and chat
Llama 4 Maverick
Meta’s open multimodal MoE for builders.
Mistral Large 2
Multilingual reasoning and code generation
Qwen 2.5 72B
Open-weight LLM for coding, math, and multilingual tasks
Qwen3 235B
Open MoE flagship with hybrid thinking.
Stable Diffusion XL
High-resolution image generation and artistic creation
Hugging Face Inference
Serverless deployment of 100,000+ open models
DeepSeek V2
Cost-efficient coding and mathematical reasoning
Gemini 2.0 Flash
Workhorse multimodal model at Flash speed.
Llama 3.3 70B
The self-host sweet spot: 70B, 405B-class quality.
Qwen 2 72B
Multilingual tasks and long-context understanding
Gemma 2 27B
High-quality open-weight reasoning and instruction following
Perplexity API
Real-time search-augmented answers with citations
Gemma 3 4B
Lightweight multimodal AI you can run on a laptop
Mistral Small 3
Low-latency chat assistants and lightweight agentic tasks
GPT-4o mini
Fast, cheap, and surprisingly capable.
Llama 4 Scout
10M context on a single H100.
Command R+
Enterprise retrieval-augmented generation and tool use
Claude 3.5 Haiku
Anthropic’s fastest model for scaled workloads.
Gemma 3
Open models from 1B to 27B, single-GPU friendly.
Groq API
Ultra-low latency inference for production applications
Phi-3 Medium
On-device AI and mobile applications
Phi-4
14B textbook-quality reasoning on a laptop.
Qwen2.5 72B
The fine-tuner’s favorite open base.
Together AI
Fine-tuning and serving open models at scale
Bedrock-native multimodal for AWS shops.
Fireworks AI
Fast, cost-efficient inference for open models
Cohere Command
Business text generation and embedding tasks
Mixtral 8x22B
The open MoE that started the wave.
Ollama
Local development and running models on personal hardware
Command R+
RAG-optimized enterprise model with citations.
Gemma 2 9B
Edge deployment and resource-constrained environments
Llama 3.2 1B
On-device and edge AI apps needing fast inference
LM Studio
GUI-based local model exploration and chat
Nemotron 70B
RLHF-tuned open model from NVIDIA.