Grok 4.6 Released: Agentic Performance Matches GPT-5.6 Sol at Lower Cost

SpaceXAI released Grok 4.6 on August 12, 2026, positioning it as a frontier model for coding, long-running agent tasks, and knowledge work. The model builds directly on Grok 4.5, b

发布于 2026年8月13日generalGEO 评分: 09 次阅读
这张图片是Grok 4.6的相关宣传图,核心内容为展示Grok 4.6与GPT-5.6 Sol的对比,下方标注了对比涵盖的维度,包括基准测试、定价、API和智能体性能。背景为深黑色,左侧隐约呈现带有斜杠的环状标识,右侧隐约呈现螺旋状标识,整体契合文档中提到的AI开发者相关的简约科技风设计,用于直观呈现两款大模型的对比主题。

Grok 4.6 Released: Agentic Performance Matches GPT-5.6 Sol at Lower Cost

Introduction

SpaceXAI released Grok 4.6 on August 12, 2026, positioning it as a frontier model for coding, long-running agent tasks, and knowledge work.

The model builds directly on Grok 4.5, but the focus is narrower and more practical: staying effective across longer chains of work, using tools over many steps, and producing stronger first versions of interactive and visual applications.

According to SpaceXAI and Cursor, Grok 4.6 now matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, while posting especially strong results on real-world agentic evaluations such as GDPval-AA v2 and AA-Briefcase.

The original AIBase article also highlights Grok 4.6's price and serving speed. The price advantage is well supported by official documentation. The speed figure needs a little more nuance: AIBase reports 80 TPS, while Artificial Analysis currently measures the first-party Grok 4.6 API at roughly 68 output tokens per second. Runtime speed can change with infrastructure, prompt length, reasoning effort, and provider load.

Grok 4.6 Builds on Grok 4.5

Grok 4.6 is an incremental model generation rather than a completely new architecture announcement.

SpaceXAI says the new model received a longer supplemental training run than Grok 4.5. The training mix included:

  • Curated model-generated reasoning data
  • Advanced technical concepts
  • High-quality engineering data
  • A revised optimizer and training recipe
  • Supervised fine-tuning across multiple reasoning efforts and agent harnesses
  • Agentic reinforcement learning across coding, knowledge work, and specialized technical environments

The company says Grok 4.5 was also used to regenerate supervised fine-tuning trajectories across STEM, software engineering, and knowledge-work tasks, with model-based filters used to remove problematic traces.

The Grok 4.5 Foundation

Grok 4.5 had already been trained at very large scale.

SpaceXAI officially states that Grok 4.5 training ran across tens of thousands of NVIDIA GB300 GPUs. Its training data covered coding, science, engineering, and mathematics, while reinforcement learning focused heavily on multi-step software-engineering tasks and other long-running technical work.

Grok 4.6 extends that foundation rather than replacing it.

The most important change is the emphasis on sustained agent behavior: a model should not merely solve an isolated coding problem, but remain coherent while researching, editing files, using tools, testing its own work, and iterating through several rounds of feedback.

Training Focus: Long-Running Agent Tasks

SpaceXAI says Grok 4.6 was trained on a broad set of agentic reinforcement-learning tasks.

Those environments include:

  • Knowledge work
  • General coding
  • Kernel optimization
  • Web development
  • Computer-aided design
  • Other domain-specific technical tasks

This helps explain why several of the model's strongest results are on evaluations that involve extended work rather than short-answer reasoning.

Better First Passes on Visual and Interactive Projects

SpaceXAI and Cursor also emphasize the model's ability to turn a broad product idea into a usable first version.

In their internal testing, Grok 4.6 was able to:

  1. Research an unfamiliar domain.
  2. Structure an application.
  3. Implement core interactions.
  4. Establish a visual direction.
  5. Test parts of its own work.
  6. Continue refining the result over multiple rounds.

This does not mean every one-shot application will be production-ready. It means the company found the initial output quality stronger than Grok 4.5, especially when the task combined code, design, interaction, and iterative tool use.

Grok 4.6 Matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index

Artificial Analysis currently gives Grok 4.6 (high) a score of 61 on its Intelligence Index.

That is the same headline score as GPT-5.6 Sol (max) in the comparison published by SpaceXAI.

The index is a composite of nine evaluations rather than a single test, so it is best used as a broad comparative signal rather than proof that two models are identical in every workload.

图片为Grok 4.6高版本与Grok 4.5高版本、GPT - 5.6 Sol Max、Fable 5 Max在多个基准测试中的表现对比表。包含AA Intelligence Index、GDPVal - AA v2、CursorBench v3.2、DeepSWE v1.1、FrontierCode v1.1 Extended、APEX - Agents、Terminal - Bench v3.0、APEX - SWE、AA - Briefcase、Harvey LAB (Vals)等测试项目,各版本在各测试项目中的得分或占比以数字或百分比形式呈现。该表与上文提到的基准测试结果相呼应,直观展示了各版本在不同测试中的表现。

The benchmark table published with the launch includes the following results:

Benchmark Grok 4.6 High Grok 4.5 High GPT-5.6 Sol Max Fable 5 Max
AA Intelligence Index 61 56 61 62
GDPval-AA v2 1753 1526 1728 1741
CursorBench v3.2 69.9% 66.7% 67.2% 70.5%
DeepSWE v1.1 65.9% 54.0% 73.0% 70.0%
FrontierCode v1.1 Extended 61.3% 56.6% 60.6% 63.6%
APEX-Agents 57.5% 47.1% 56.7% 59.2%
Terminal-Bench v3.0 26.0% 15.7% 34.6% 34.1%
APEX-SWE 56.4% 53.6% 58.8%
AA-Briefcase 1577 1313 1502 1574
Harvey LAB (Vals) 15.8% 12.9% 2.5% 11.3%

Benchmark positions change as models, harnesses, and evaluation versions are updated. The table above reflects the launch-era comparison, not a permanent ranking.

GDPval-AA v2: Strong Real-World Knowledge Work

On GDPval-AA v2, Grok 4.6 scored 1753 Elo.

Artificial Analysis describes GDPval-AA v2 as an evaluation of real-world agentic knowledge work across many professional occupations.

At the time of the Grok 4.6 analysis, Artificial Analysis said the score was behind only the Claude Opus 5 family, while the confidence intervals overlapped with other leading models such as Claude Fable 5 and Qwen3.8 Max.

This is a more careful interpretation than simply saying Grok 4.6 is definitively second-best overall.

Elo-based evaluations carry uncertainty, and close scores can overlap statistically.

AA-Briefcase: Long-Horizon Agentic Work

Grok 4.6 also scored 1577 Elo on AA-Briefcase, Artificial Analysis' private benchmark for long-horizon agentic knowledge work.

That puts it at roughly Fable 5-tier performance in the launch snapshot, behind the Claude Opus 5 family.

The efficiency profile is notable as well.

Artificial Analysis reports that Grok 4.6 completed its AA-Briefcase tasks in roughly:

  • 53 turns on average
  • 0.5 billion input tokens on average

For comparison, Claude Opus 5 (max) reportedly used roughly:

  • 103 turns
  • 2.0 billion input tokens

That does not mean Grok 4.6 is universally more efficient than Claude Opus 5. It is a benchmark-specific observation. Still, for long-running agents, fewer turns and less accumulated context can materially reduce task cost.

Grok 4.6 Pricing Starts at $2 Input and $6 Output

The standard headline API pricing remains one of Grok 4.6's strongest competitive points.

For prompts below the long-context pricing threshold, SpaceXAI documents:

Token Type Price per 1M Tokens
Input $2.00
Cached input $0.50
Output $6.00

Artificial Analysis notes that this headline pricing is unchanged from Grok 4.5 even though the Intelligence Index score increased from 56 to 61.

By comparison, at the time of this article:

  • GPT-5.6 Sol standard API pricing is $5 input / $30 output per 1M tokens.
  • Claude Opus 5 is priced higher than Grok 4.6 on its headline token rates.

This makes Grok 4.6 especially interesting for high-volume agent loops where output tokens and repeated context can dominate total cost.

Long Prompts Cost More

The official SpaceXAI release notes add an important detail not mentioned in the short AIBase article.

For prompts above 200,000 tokens, Grok 4.6 uses a higher pricing tier:

Token Type Price per 1M Tokens Above 200k Prompt
Input $4.00
Cached input $1.00
Output $12.00

So "$2 input / $6 output" is the starting price, not the universal price for every possible request.

Fast Variant

Cursor says a fast variant is available at twice the standard price.

This is separate from the long-prompt pricing rule above. Depending on the platform and workload, users should check which speed tier and context tier apply before estimating production cost.

What About the Reported 80 TPS Speed?

The AIBase article says Grok 4.6 can run at 80 tokens per second.

That number should be treated carefully.

SpaceXAI's official Grok 4.6 announcement does not publish a fixed 80 TPS serving-speed guarantee. Artificial Analysis currently measures Grok 4.6 (high) on the first-party API at approximately 67.6 output tokens per second under its current benchmark workload.

By contrast, SpaceXAI's official Grok 4.5 launch page explicitly said Grok 4.5 was served at 80 TPS.

It is therefore possible that the short source article carried forward the earlier 80 TPS figure or referenced a different serving tier.

The practical takeaway is simpler:

Grok 4.6 is reasonably fast for a frontier reasoning model, but there is no single fixed TPS number that applies to every prompt, provider, reasoning effort, or speed tier.

500K Context on the SpaceXAI API

SpaceXAI's API documentation lists a 500,000-token context window for grok-4.6.

The model supports:

  • Text input
  • Image input
  • Text output
  • Low, medium, high, and xhigh reasoning effort
  • Function calling
  • Web search
  • X search
  • Code execution

The official model page also lists a knowledge cutoff of February 1, 2026.

Cursor Uses a Different Context Limit

Cursor's Grok 4.6 model documentation currently lists a 256k context window inside Cursor.

That is not a contradiction in the base model specification. It means the platform integration can expose a smaller context limit than the first-party SpaceXAI API.

When comparing model limits, always check the exact provider and integration being used.

How to Use Grok 4.6 Through the API

SpaceXAI's official model ID is:

grok-4.6

A minimal Responses API request looks like this:

curl https://api.x.ai/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -d '{
    "model": "grok-4.6",
    "input": "Find and fix the bug, then explain it: function median(a){a.sort();return a[a.length/2]}"
  }'

The same model is also available through the official xAI SDK and OpenAI-compatible API clients.

Python Example

import os
from xai_sdk import Client
from xai_sdk.chat import user

client = Client(api_key=os.getenv("XAI_API_KEY"))

chat = client.chat.create(model="grok-4.6")
chat.append(
    user(
        "Find and fix the bug, then explain it: "
        "function median(a){a.sort();return a[a.length/2]}"
    )
)

response = chat.sample()
print(response.content)

For long-running agent loops, SpaceXAI recommends using prompt caching and context compaction where appropriate to reduce repeated context cost.

Grok 4.6 Is Available in Cursor and Grok Build

At launch, Grok 4.6 became available in:

  • Cursor
  • Grok Build
  • SpaceXAI API
  • OpenRouter
  • Vercel
  • Cloudflare

SpaceXAI and Cursor also offered 2x included usage for the first week in Cursor and Grok Build.

That launch promotion is time-limited and should not be treated as permanent pricing.

Grok 4.6 in Cursor

Cursor exposes Grok 4.6 with four reasoning-effort levels:

  • xhigh
  • high
  • medium
  • low

high is the default reasoning effort.

On eligible Cursor plans, a fast speed tier is also available.

Inside Cursor, the model can use the platform's agent tools for tasks such as:

  • Searching files and folders
  • Reading files
  • Editing files
  • Running shell commands
  • Browsing the web
  • Browser-based application testing
  • Image generation
  • Fetching project rules

This is one reason the Grok 4.6 launch focuses so heavily on agentic benchmarks: its intended use is not limited to isolated chat completion.

Grok 4.6 Is Strongest When the Task Requires Sustained Work

The benchmark pattern suggests Grok 4.6's largest gains are not simply in static reasoning.

Its improvements are most visible when a task involves:

  1. Multiple steps
  2. Tool use
  3. Research
  4. Code or document manipulation
  5. Self-checking
  6. Long context
  7. Iterative refinement

That makes it especially relevant for:

  • Coding agents
  • Repository-scale software work
  • Research agents
  • Spreadsheet and document workflows
  • Technical analysis
  • Web application prototyping
  • Long-running knowledge-work tasks

For short, simple prompts, the advantage may be much less noticeable.

What the Benchmarks Do and Do Not Prove

The launch data supports several clear statements:

  • Grok 4.6 is materially stronger than Grok 4.5 on several agentic benchmarks.
  • It matches GPT-5.6 Sol on the current Artificial Analysis Intelligence Index launch comparison.
  • It performs particularly well on GDPval-AA v2 and AA-Briefcase.
  • Its standard headline API pricing is substantially lower than GPT-5.6 Sol's.

But the data does not prove that Grok 4.6 is universally better than GPT-5.6 Sol or Claude across every task.

For example, the same launch table shows GPT-5.6 Sol ahead on DeepSWE v1.1 and Terminal-Bench v3.0, while Claude Fable 5 leads several other coding and agent evaluations.

Model selection should therefore depend on the actual workload, not one aggregate score.

常见问题

What is Grok 4.6?

Grok 4.6 is SpaceXAI's frontier model for coding, agentic tasks, and knowledge work. It builds on Grok 4.5 with additional training focused on long-running agents, stronger first-pass application building, and more sustained multi-step work.

Is Grok 4.6 better than GPT-5.6 Sol?

It depends on the task. Grok 4.6 matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index in the launch comparison and leads it on some agentic evaluations, while GPT-5.6 Sol remains stronger on some coding benchmarks. The two models also differ significantly in pricing and context limits.

How much does the Grok 4.6 API cost?

Standard pricing starts at $2 per million input tokens, $0.50 per million cached input tokens, and $6 per million output tokens. For prompts above 200k tokens, SpaceXAI documents a higher $4 / $1 / $12 tier for input, cached input, and output respectively.

What is the Grok 4.6 context window?

The first-party SpaceXAI API supports a 500k-token context window. Cursor currently documents a 256k context window for its own Grok 4.6 integration, so the effective limit depends on the platform.

Does Grok 4.6 support images?

Yes. SpaceXAI lists Grok 4.6 as supporting text and image input with text output. The model also supports tools such as function calling, web search, X search, and code execution through the xAI API.

Can I use Grok 4.6 in Cursor?

Yes. Grok 4.6 is available in Cursor and supports xhigh, high, medium, and low reasoning effort. Cursor also gives it access to its agent toolset for file operations, shell commands, browsing, and other coding workflows.

Is Grok 4.6 open source?

No. Grok 4.6 is a proprietary model, and SpaceXAI has not published its model weights or parameter count.

How fast is Grok 4.6?

AIBase reported 80 TPS, but current Artificial Analysis measurements place the first-party Grok 4.6 API around 68 output tokens per second under its benchmark workload. Speed varies by reasoning effort, prompt size, infrastructure load, and platform tier.

相关工具

  • Grok 4.6: SpaceXAI's official launch page covering training, benchmarks, safety, and availability.
  • SpaceXAI API Console: The official console for creating API keys and accessing Grok models.
  • Grok Build: SpaceXAI's agentic coding and work environment powered by Grok models.
  • Cursor: AI coding environment that launched Grok 4.6 alongside SpaceXAI and exposes the model through its agent workflow.
  • Artificial Analysis: Independent model benchmarking for intelligence, cost, speed, context, and agentic performance.

Related Links

Summary

Grok 4.6 is a targeted upgrade to Grok 4.5 that focuses on long-running agents, coding, knowledge work, and better first-pass execution on interactive projects. Its launch results put it at the frontier of current model evaluations, including a score of 61 on the Artificial Analysis Intelligence Index—matching GPT-5.6 Sol in the launch comparison.

The strongest part of the release is the combination of agentic performance and price. Grok 4.6 starts at $2 per million input tokens and $6 per million output tokens, while also scoring strongly on GDPval-AA v2 and AA-Briefcase. Long prompts and fast tiers can cost more, so production estimates should use the exact provider configuration.

The reported 80 TPS figure deserves caution: current independent measurements are closer to 68 output tokens per second, and serving speed varies over time.

Grok 4.6's main advantage is not that it wins every benchmark, but that it now offers frontier-level agentic performance at a notably aggressive price.

Grok 4.6 发布:智能体性能媲美 GPT-5.6 Sol,成本更低