AI Models

Choose the right model
for your task

7 world-class models, all running on Groq's LPU™ hardware for up to 10× faster inference than GPU-based APIs. Pick one — or compare two side-by-side.

7

Models Available

~450 tok/s

Max Speed

128K tokens

Max Context

Free

Cost

All Models

Explore all 7 models

All models run on Groq's ultra-fast LPU™ inference engine — no GPU bottlenecks.

Most Capable

Llama 3.3 70B

Meta

In-depth research, complex multi-step reasoning, and detailed technical answers.

Speed

~280 tok/s

Capability

128K tokens

Complex reasoningLong-form writingCode reviewAnalysis
Try in Studio
Fastest

Llama 3.1 8B ⚡

Meta

Fast prototyping, quick Q&A, chat applications, and simple summarization.

Speed

~450 tok/s

Capability

128K tokens

Quick answersChatSummariesSimple tasks
Try in Studio
Best Reasoning

DeepSeek R1 70B

DeepSeek

Mathematical proofs, logic puzzles, scientific reasoning, and chain-of-thought tasks.

Speed

~180 tok/s

Capability

128K tokens

MathStep-by-step logicSTEM problemsChain-of-thought
Try in Studio
Efficient

Gemma 2 9B

Google DeepMind

Precise instruction following, structured JSON output, and concise focused responses.

Speed

~320 tok/s

Capability

8K tokens

Instruction followingStructured outputConcise answers
Try in Studio
Multilingual

Mixtral 8×7B

Mistral AI

Non-English tasks, long-document analysis, diverse creative projects.

Speed

~240 tok/s

Capability

32K tokens

MultilingualLong contextCreative writingMoE architecture
Try in Studio
Deep Thinker

Qwen QwQ 32B

Alibaba Cloud

Problems requiring deep logical reasoning, advanced coding, and thorough analysis.

Speed

~200 tok/s

Capability

32K tokens

Extended reasoningCodingMathDeep analysis
Try in Studio
Vision

Llama 3.2 Vision

Meta

Analyzing images, reading screenshots, understanding charts, and visual question answering.

Speed

~150 tok/s

Capability

128K tokens

Image understandingVisual Q&AOCRCharts & diagrams Vision
Try in Studio
Comparison

Side-by-side spec sheet

All the numbers in one place. Scroll right on mobile.

ModelProviderContextSpeedVisionBest For

Llama 3.3 70B

Most Capable
Meta128K tokens
~280 tok/s
Complex reasoning, Long-form writing

Llama 3.1 8B ⚡

Fastest
Meta128K tokens
~450 tok/s
Quick answers, Chat

DeepSeek R1 70B

Best Reasoning
DeepSeek128K tokens
~180 tok/s
Math, Step-by-step logic

Gemma 2 9B

Efficient
Google DeepMind8K tokens
~320 tok/s
Instruction following, Structured output

Mixtral 8×7B

Multilingual
Mistral AI32K tokens
~240 tok/s
Multilingual, Long context

Qwen QwQ 32B

Deep Thinker
Alibaba Cloud32K tokens
~200 tok/s
Extended reasoning, Coding

Llama 3.2 Vision

Vision
Meta128K tokens
~150 tok/s
Image understanding, Visual Q&A
Guide

Not sure which to pick? We've got you.

Quick recommendations based on your use case.

For everyday chat & Q&A

Llama 3.1 8B ⚡

Fastest responses, great for conversational tasks and quick lookups.

For math & reasoning

DeepSeek R1 70B

Purpose-built for chain-of-thought reasoning, logic, and STEM problems.

For coding & deep analysis

Llama 3.3 70B or Qwen QwQ 32B

Both excel at code generation, review, and complex multi-step analysis.

For image understanding

Llama 3.2 Vision

The only model with vision — reads screenshots, charts, and images.

Ready to try them all?

Open the Studio and switch models instantly. Free, no login required.

Open Studio Free