7 world-class models, all running on Groq's LPU™ hardware for up to 10× faster inference than GPU-based APIs. Pick one — or compare two side-by-side.
7
Models Available
~450 tok/s
Max Speed
128K tokens
Max Context
Free
Cost
All models run on Groq's ultra-fast LPU™ inference engine — no GPU bottlenecks.
Meta
In-depth research, complex multi-step reasoning, and detailed technical answers.
Speed
~280 tok/s
Capability
128K tokens
Meta
Fast prototyping, quick Q&A, chat applications, and simple summarization.
Speed
~450 tok/s
Capability
128K tokens
DeepSeek
Mathematical proofs, logic puzzles, scientific reasoning, and chain-of-thought tasks.
Speed
~180 tok/s
Capability
128K tokens
Google DeepMind
Precise instruction following, structured JSON output, and concise focused responses.
Speed
~320 tok/s
Capability
8K tokens
Mistral AI
Non-English tasks, long-document analysis, diverse creative projects.
Speed
~240 tok/s
Capability
32K tokens
Alibaba Cloud
Problems requiring deep logical reasoning, advanced coding, and thorough analysis.
Speed
~200 tok/s
Capability
32K tokens
Meta
Analyzing images, reading screenshots, understanding charts, and visual question answering.
Speed
~150 tok/s
Capability
128K tokens
All the numbers in one place. Scroll right on mobile.
| Model | Provider | Context | Speed | Vision | Best For |
|---|---|---|---|---|---|
Llama 3.3 70B Most Capable | Meta | 128K tokens | ~280 tok/s | — | Complex reasoning, Long-form writing |
Llama 3.1 8B ⚡ Fastest | Meta | 128K tokens | ~450 tok/s | — | Quick answers, Chat |
DeepSeek R1 70B Best Reasoning | DeepSeek | 128K tokens | ~180 tok/s | — | Math, Step-by-step logic |
Gemma 2 9B Efficient | Google DeepMind | 8K tokens | ~320 tok/s | — | Instruction following, Structured output |
Mixtral 8×7B Multilingual | Mistral AI | 32K tokens | ~240 tok/s | — | Multilingual, Long context |
Qwen QwQ 32B Deep Thinker | Alibaba Cloud | 32K tokens | ~200 tok/s | — | Extended reasoning, Coding |
Llama 3.2 Vision Vision | Meta | 128K tokens | ~150 tok/s | Image understanding, Visual Q&A |
Quick recommendations based on your use case.
For everyday chat & Q&A
Fastest responses, great for conversational tasks and quick lookups.
For math & reasoning
Purpose-built for chain-of-thought reasoning, logic, and STEM problems.
For coding & deep analysis
Both excel at code generation, review, and complex multi-step analysis.
For image understanding
The only model with vision — reads screenshots, charts, and images.
Open the Studio and switch models instantly. Free, no login required.
Open Studio Free