Grok 4.6 matches top models as SpaceX eyes worker data
4 takeaways from Grok 4.6’s 61-score launch, its pricing edge, and SpaceX’s plan to train on employee work.

How good is Grok 4.6, and what does SpaceX want to train it on?
Grok 4.6 matches leading models on benchmarks while SpaceX plans to use employee work as training data.
| Item | Intelligence Index | Input price | Output price |
|---|---|---|---|
| Grok 4.6 | 61 | $2 / 1M tokens | $6 / 1M tokens |
| OpenAI GPT-5.6 Sol | 61 | $5 / 1M tokens | $30 / 1M tokens |
| Anthropic Claude Fable 5 | 62 | $5 / 1M tokens | n/a |
| Claude Opus 5 | n/a | $5 / 1M tokens | $25 / 1M tokens |
1. Grok 4.6
Get the latest AI news in your inbox
Weekly picks of model releases, tools, and deep dives — no spam, unsubscribe anytime.
No spam. Unsubscribe at any time.
SpaceXAI’s newest model lands with a clear pitch: it improves without a bigger base model. The company says Grok 4.6 keeps the same 1.5-trillion-parameter V9 foundation as Grok 4.5, then gains ground through longer supplemental training, better supervised fine-tuning, and more reinforcement learning.

That matters because it shows how much performance can still come from post-training work. Instead of restarting pre-training from scratch, SpaceXAI used Grok 4.5 to regenerate and filter training trajectories, then pushed the model through agent-heavy environments such as kernel optimization, web development, and CAD.
- Foundation: 1.5T-parameter V9, unchanged from Grok 4.5
- Release date: Aug. 12, 2026
- Training focus: SFT, RL, and cleaner trajectory filtering
2. Benchmark result
Independent testing gave Grok 4.6 an Artificial Analysis Intelligence Index score of 61, up from 56 for Grok 4.5 and 38 for Grok 4.3. That puts it level with OpenAI’s GPT-5.6 Sol and within one point of Anthropic’s Claude Fable 5.
The agentic numbers are especially strong. On GDPval-AA v2, Grok 4.6 posted an Elo of 1753, placing it behind only Claude Opus 5 and in the same statistical tier as Claude Fable 5 and Qwen3.8 Max. For buyers who track model rankings closely, this is the part that turns a new launch into a real option.
- AI Index: 61
- Grok 4.5 comparison: 56
- Grok 4.3 comparison: 38
- Agent benchmark: GDPval-AA v2 Elo 1753
3. Token efficiency
Grok 4.6’s practical edge is not just score-based. Artificial Analysis said it completed knowledge-work tasks in about 53 turns and 0.5 billion input tokens on average, compared with Claude Opus 5’s 103 turns and 2.0 billion input tokens. Fewer turns can mean less orchestration overhead and lower real-world cost per task.

The pricing helps too. Grok 4.6 keeps the same rates as Grok 4.5 at $2 per million input tokens and $6 per million output tokens. That undercuts the listed pricing for both Anthropic’s Claude Opus 5 and OpenAI’s GPT-5.6 Sol, which makes Grok a stronger fit for teams watching API spend.
Grok 4.6: $2 in / $6 out per 1M tokens
Claude Opus 5: $5 in / $25 out
GPT-5.6 Sol: $5 in / $30 out
4. SpaceX employee data plan
The bigger story is not only the model launch. Elon Musk said SpaceX intends to train future Grok versions on “the sum total of all SpaceX information,” including what employees think and produce. He framed that as a way to pass on the judgment of SpaceX engineers to the model.
But the company has not said what data it will collect, how it will collect it, or whether workers can opt out. That ambiguity matters because the best training material for agentic AI is behavioral data: how people use tools, recover from errors, and complete multi-step work in real systems.
- Workforce size cited: about 14,000 to 15,000 employees
- No disclosed opt-out process
- No published list of data sources or safeguards
5. The Meta precedent
There is a recent warning sign here. Meta’s Model Capability Initiative tried to log employee computer activity for AI training, then ran into backlash and a data incident that exposed private conversations, performance data, and transcriptions company-wide. The program is still paused.
That history makes SpaceX’s plan easier to judge. If a company wants worker activity as training fuel, it needs tight scope, clear consent rules, and serious access controls. Without those, the risk is not theoretical, and the damage can spread fast once the data exists.
- Meta program launched: April 2026
- Incident severity: SEV 2
- Status: paused as of Aug. 12, 2026
How to decide
If you want the cheapest frontier-style API among the models named here, Grok 4.6 is the one to watch. If you care most about raw benchmark rank, Claude Fable 5 and Claude Opus 5 still sit at the top of the comparison set. If your main concern is workplace data policy, SpaceX’s plan is the part to scrutinize first.
For builders, the useful split is simple: Grok 4.6 looks attractive on cost and efficiency, while the employee-data story raises governance questions that could matter before any rollout reaches production.
// Related Articles
- [IND]
Grok 4.6 makes frontier AI cheaper for builders
- [IND]
Wall Street Should Put Real Assets On Blockchains, Not Just Pilot Them
- [IND]
WorkBuddy Skills: 5 picks that fit real office work
- [IND]
GLM-5.3’s coding gains came from post-training
- [IND]
Anthropic’s data center push now has big capital backing
- [IND]
DailyArxiv turns arXiv keywords into a daily paper feed