Grok 4.7: Benchmarks, Price, and Practical Use Cases

Grok 4.7 combines a 500K-token context window with competitive API pricing and strong vendor-reported results in long-running coding and knowledge-work tasks.

Grok 4.7 is xAI’s latest model for programming, knowledge work, and long-running tasks. According to xAI, it was trained to sustain reasoning for longer, verify its work more effectively, and handle large contexts.

The company reports meaningful gains across software engineering, terminal work, electrical engineering, and professional-task benchmarks. Artificial Analysis adds an independent perspective on the model’s price, throughput, and market position.

For Tess users, the practical takeaway is straightforward: Grok 4.7 is worth evaluating for workflows that need broad context, multi-step execution, and competitive API economics—particularly where long-running work matters more than instant output.

At a glance

  • Artificial Analysis Intelligence Index: 46 points.
  • Artificial Analysis overall rank: 16th of 202 models listed at the time of review.
  • Context window: 500,000 tokens.
  • Measured output speed: 39.3 output tokens per second.
  • Base API price: $2 per million input tokens and $6 per million output tokens.
  • Cache discount: 75%, according to the Artificial Analysis model page.
  • Fast variant: xAI states that a faster version delivers roughly twice the output speed at roughly twice the price.

Bottom line: Grok 4.7 combines a large context window and comparatively low API pricing with strong vendor-reported results in long-horizon coding and knowledge-work tasks.

What changed from Grok 4.6?

xAI says Grok 4.7 uses a larger base model than Grok 4.6 and received a longer reinforcement-learning phase. The training emphasis shifted toward harder tasks that may take hours to complete.

In practice, the company’s claim centers on three areas:

  1. Self-verification: stronger ability to review intermediate work, spot inconsistencies, and correct execution.
  2. Long-context management: improved handling of long conversations, files, instructions, and multi-stage tasks.
  3. Conversational and knowledge work: native training with the Grok Bot harness for interaction, context retrieval, and general professional tasks.

Grok 4.7 benchmark overview
Existing chart supplied for this article.

xAI-reported benchmark results

The results below were published by xAI in its Grok 4.7 announcement. They help explain how xAI positions the model, but they should be read as vendor-reported benchmarks rather than independent validation.

Area Benchmark Grok 4.7 Grok 4.6 Notes
Software engineering CursorBench 4.0 46.3% 40.4% Evaluates longer-form programming tasks.
Software engineering DeepSWE v1.1 71.0%* 65.2% xAI labels the Grok 4.7 result as high effort.
Electrical engineering EEBench 64.0% 53.0% Grok 4.7 leads the models presented in xAI’s comparison.
Long-horizon office work AA Briefcase v1.1 1,657 Elo 1,546 Elo Evaluates extended professional workflows.
Terminal work Terminal-Bench 4.0 38.0% 20.3% The largest relative gain versus Grok 4.6 in xAI’s table.
Legal work Harvey Legal Agent Benchmark 19.6% 15.8% Agent-style legal tasks.
Clinical reasoning HealthBench Professional 56.7% 48.5% Professional health-related reasoning tasks.

* Result labeled high effort by xAI.

Where the improvement appears most meaningful

The increase on Terminal-Bench 4.0, from 20.3% to 38.0%, stands out. This benchmark is designed around longer terminal tasks, where a model must navigate files, run commands, interpret output, and complete a sequence of steps.

Other notable gains include:

  • 40.4% to 46.3% on CursorBench 4.0;
  • 65.2% to 71.0% on DeepSWE v1.1;
  • 53.0% to 64.0% on EEBench; and
  • a 111 Elo-point gain on AA Briefcase v1.1.

Grok 4.7 coding benchmarks
Existing chart supplied for this article.

Knowledge work and professional deliverables

xAI also highlights performance on tasks such as document creation, presentations, and professional deliverables. In its announcement, the company cites evaluations including GDPval and AA Briefcase, which simulate work performed in fields such as law, nursing, and financial analysis.

In xAI’s published comparisons:

  • GDPval: Grok 4.7 reaches 1,695 Elo, compared with 1,605 Elo for Grok 4.6.
  • AA Briefcase: Grok 4.7 reaches 1,657 Elo, compared with 1,546 Elo for Grok 4.6.
  • EEBench: Grok 4.7 records 64.0%, up from 53.0%.

These results do not mean the model replaces specialist review. They do suggest it may be better suited to structured drafting, long-material analysis, presentation preparation, research assistance, and automations that involve multiple instructions.

Grok 4.7 knowledge-work benchmarks
Existing chart supplied for this article.

The independent Artificial Analysis view

Artificial Analysis lists Grok 4.7 (xhigh) at 46 points in its Artificial Analysis Intelligence Index, ranking it 16th among 202 models at the time of review.

The index aggregates ten evaluations, including AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1.

This score should not be compared directly with results from other indices or individual benchmarks. It is an aggregate view based on Artificial Analysis’s own methodology.

Speed and context

Artificial Analysis reports the following for Grok 4.7:

  • 39.3 output tokens per second;
  • 151st of 202 in its speed ranking;
  • a 500,000-token context window; and
  • an estimated capacity of about 750 A4 pages, as shown on the platform.

Speed is not the model’s relative strength in Artificial Analysis’s ranking. Its 500,000-token context window, however, is relevant to workflows involving extensive documentation, multiple files, support histories, technical specifications, and long-running processes.

Pricing and operating cost

The Artificial Analysis model page and xAI announcement report the same base prices:

Item Price
Input tokens $2 per million
Output tokens $6 per million
Cache reads 75% discount, according to Artificial Analysis
Fast variant Roughly 2× output speed and 2× the price, according to xAI

In Artificial Analysis’s cost ranking, Grok 4.7 is listed 74th of 202 models.

For teams running automations, total cost is shaped by more than the per-million-token rate: the context sent with each request, output size, cache utilization, retries, and task complexity all matter.

Practical reading: $2 / $6 per million tokens is competitive for a frontier model, while the measured 39.3 tokens per second may matter for products that require near-instant output.

Grok 4.7 price and performance
Existing chart supplied for this article.

Safety and cybersecurity: what xAI reports

xAI states that Grok 4.7 includes a new safeguard layer and has improved resistance to jailbreaks and unsafe completions.

Among the company-reported figures:

  • 62.4% on the LatchBio biosafety benchmark;
  • only 3.3% of risky dual-use prompts allowed on HackerBench v0.3; and
  • controlled red-team access for cybersecurity partners.

These figures are xAI claims and should be considered within the context of its methodology and governance. For sensitive deployments, model selection does not replace access policies, human review, audit records, and security controls.

Grok 4.7 safety benchmarks
Existing chart supplied for this article.

When Grok 4.7 is worth testing in Tess

Consider Grok 4.7 when your workflow includes:

  • long-running programming tasks, especially those involving terminals, repositories, and validation steps;
  • analysis of extensive documentation, using the 500,000-token context window;
  • document and presentation creation based on substantial source material;
  • knowledge-work automation requiring multiple instructions and review loops; or
  • API cost experiments where a competitive input/output price matters.

It may be less suitable as the first choice when:

  • the experience depends on very fast responses, given Artificial Analysis’s measured 39.3 output tokens per second;
  • the use case requires specific independent validation for safety, accuracy, or compliance; or
  • the task is simple enough that a smaller model can deliver sufficient output at lower operating cost.

Conclusion

Grok 4.7 arrives with a clear proposition: more capability for long-running tasks, more context, and stronger vendor-reported results in coding, terminal work, and professional workflows.

Artificial Analysis complements that picture: the model is listed at 46 points in its Intelligence Index, ranks 16th among 202 models, and combines a 500,000-token context window with $2 per million input tokens and $6 per million output tokens.

For Tess teams, the best next step is to compare Grok 4.7 with other available models using real work: a set of support tickets, a knowledge base, a repository, a complex spreadsheet, or a document-production workflow. Benchmarks help narrow the field; testing within the actual workflow determines the best choice.


Sources

Editorial note: benchmark results, pricing, and availability may change. Check the original sources before publication.

Keep reading

Explore more product news and best practices for teams building with Plataforma Tess pre_prod.

Build with TESS

Turn ideas from this article into working AI workflows.

Create agents, automations, and knowledge-powered workflows in one platform built for teams.