Grok 4.5 is a Flagship coding/agent from xAI (United States).
One-line fit: best used for High-volume coding agents, cost-efficient SWE, office plugins.
Key specifications
| Provider | xAI |
|---|---|
| Model | Grok 4.5 |
| Type | Flagship coding/agent |
| Context window | 500K |
| Pricing (indicative) | $2 / $6 per 1M; cache ~$0.50 |
| Company category | Frontier Lab |
| Region | United States |
| Official site | https://x.ai |
Benchmarks and public standing
SWE Marathon 29% (#1 in xAI table); Terminal Bench 2.1 ~83.3%; SWE-Bench Pro ~64.7%; ~4.2x fewer tokens vs Opus 4.8 on some tasks
Treat scores as directional. SWE-bench, Terminal-Bench, LMSYS Arena, Artificial Analysis, and vendor cards use different harnesses and are not always comparable 1:1. Re-check the latest official model card before decisions.
What this model is good at
High-volume coding agents, cost-efficient SWE, office plugins
This aligns with xAI’s broader strengths:
- Intelligence per dollar
- Coding agents (Cursor flywheel)
- Realtime news/culture via X
- High TPS inference
Limitations and watch-outs
SWE-Pro trails Fable; EU lag at launch
- Rate limits, regional availability, and data retention policies vary by plan.
- Agent harness quality (Cursor, Claude Code, Codex, custom tools) can change outcomes more than raw model Elo.
- Open-weight availability (if any) is separate from hosted API quality and safety filters.
Ideal users
- Product engineers shipping features that match: High-volume coding agents, cost-efficient SWE, office plugins
- Teams standardizing on the xAI ecosystem
- Agent builders who need a Flagship coding/agent
When to pick something else
- Cheaper volume: compare lower tiers from the same lab or open Chinese/EU alternatives.
- Maximum hard SWE thoughtfulness: compare Claude Fable-class and other coding flagships.
- Giant multimodal corpora / Workspace: compare Gemini-class models.
- Self-host / open weights: Llama, Qwen, DeepSeek, GLM, Mistral open lines.
Parent company
xAI – Grok models emphasize real-time X data, coding agents, price/efficiency, and a distinct less-lecturing personality.
Related models from xAI
- Grok 4 / prior – Prior flagship
- Grok Imagine / media – Image/video gen
Data is curated for AIForumSphere model directory mid-2026 directional figures.