Grok 4.7 is SpaceXAI's most powerful model for coding and knowledge work, and the company describes it as its most capable model for these tasks. According to the announcement, Grok 4.7 works longer on difficult tasks, checks its own work more carefully, and comes with SpaceXAI's best-calibrated safeguards to date. It is served at the same price and speed as Grok 4.6, which the company says makes it highly competitive in its class, and it is pitched as twice as fast at half the price of comparable models. Grok 4.7 is available today in Cursor and Grok Build, through the Grok API, and across third-party coding harnesses, model routers, and cloud platforms.
The announcement frames Grok 4.7 around the demands of long, difficult work rather than short chat interactions. SpaceXAI states that the model uses a new, larger base model compared to Grok 4.6 and that it was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. This focus matters because long-running coding and knowledge tasks place unusual pressure on a model: it must stay coherent across extended context, avoid drifting from the original objective, and catch its own mistakes before handing work back. The company reports that Grok 4.7 is better at verifying its own work and managing longer context, which directly targets those failure modes. On CursorBench 4.0, a benchmark that stresses longer-running coding tasks, SpaceXAI says Grok 4.7 sits at the frontier in price-performance, with a chart comparing it to Fable 5.1, Opus 5, GPT-5.6 Sol, and Sonnet 5 on benchmark score against average cost per task.
Beyond raw size and training duration, SpaceXAI highlights a specific capability addition: Grok 4.7 was trained to natively understand the Grok Bot harness. The company states this makes the model better at conversational tasks and general knowledge work. In practical terms, a model that natively understands its harness can operate more naturally inside that environment, handling multi-step conversations and everyday knowledge tasks without the friction that comes from adapting an externally trained model to a new interface. That combination, a larger base model, a longer reinforcement learning run weighted toward multi-hour problems, improved self-verification, better long-context management, and native harness understanding, is the core of what distinguishes Grok 4.7 from its predecessor.
SpaceXAI publishes a detailed benchmark table comparing Grok 4.7 against Grok 4.6, GPT-5.6 Sol, and Fable 5.1. Grok 4.7 scores 46.3% on CursorBench 4.0, 71.0% on DeepSWE v1.1 with a high-effort marker, 64.0% on EEBench, 1,657 on AA Briefcase v1.1, 37.6% on Terminal-Bench 4.0, 19.6% on the Harvey Legal Agent Benchmark, and 56.7% on HealthBench Professional. The company also reports a GDPval Elo score of 1,695 for Grok 4.7 at xhigh effort, against 1,735 for Fable 5.1 at max, 1,605 for Grok 4.6 at high, and 1,542 for GPT-6 Astra at max. These benchmarks span software engineering, multi-hour terminal work, multi-hour office work, electrical engineering, legal work, and clinical reasoning, which reflects the breadth of tasks the model is positioned to handle.
On professional knowledge work, SpaceXAI states that Grok 4.7 is better at creating documents and presentations. In GDPval and AA Briefcase, the company explains, AI is asked to work on tasks done by professionals such as lawyers, nurses, and financial analysts. Grok 4.7 improves upon Grok 4.6 on both benchmarks and performs comparably to other frontier models. That means the model is not positioned only as a coding tool: it is also aimed at the document-heavy, multi-hour office work that these professions perform, where producing a usable deliverable matters more than producing a quick answer. The benchmark table supports this positioning, showing Grok 4.7's AA Briefcase score of 1,657 against 1,546 for Grok 4.6, 1,487 for GPT-5.6 Sol, and 1,678 for Fable 5.1.
Safety and cybersecurity are treated as a first-class part of the release. SpaceXAI says Grok 4.7 was built with an entirely new safeguard stack and that it is the strongest model the company has tested on refusals and jailbreak resistance. In dual-use domains such as cybersecurity and biological work, the company reports that it leads on both utility for benign tasks and safe refusal on dangerous ones, topping LatchBio's biosafety benchmark at 62.4%. The company also states that Grok 4.7 balances strong cyber defense capabilities with low refusal rates for legitimate use, showing the highest safety on HackerBench v0.3, its benchmark for risky and malicious cyber tasks, allowing only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. In addition, SpaceXAI has started giving select cybersecurity partners invite-only access to Grok 4.7's red-team capabilities for defense research.
The overall approach behind Grok 4.7 is a combination of scale, extended reinforcement learning, and deliberate alignment work rather than a single headline change. SpaceXAI describes a new, larger base model trained with a longer reinforcement learning run on a harder mix of tasks weighted toward problems that take many hours, which produces a model that verifies its own work and manages longer context more effectively. Native training on the Grok Bot harness adds conversational and general knowledge work strength, and an entirely new safeguard stack supplies refusal and jailbreak resistance alongside dual-use safety. The company then serves the result at the same price and speed as Grok 4.6, with an additional fast variant that runs at twice the output speed for twice the price. Each element reinforces the others: a stronger base model makes longer autonomous runs viable, and a stronger safeguard stack makes those runs safer to deploy.
Pricing and availability are explicit. Grok 4.7 is priced starting at $2 per million input tokens and $6 per million output tokens. SpaceXAI also serves a fast variant with twice the output speed at twice the price. The model is available today in Cursor and in Grok Build, and it is also available through the Grok API, third-party coding harnesses, and model routers and cloud platforms. For developers who want to try it before committing, the company offers a free try in Grok Build, and it publishes a terminal install command: curl -fsSL https://x.ai/cli/install.sh | bash. API keys can be created from the console, and documentation is available at docs.x.ai.
For users, the stated benefit is straightforward: frontier-class coding and knowledge work at a price-performance point that makes long-running tasks economically practical. Because Grok 4.7 is served at the same price and speed as Grok 4.6, teams that already use the previous model can adopt the new one without changing their cost structure, while gaining a larger base model, longer reinforcement learning, better self-verification, and stronger long-context handling. The improved document and presentation abilities extend that value beyond engineering into professional knowledge work, and the new safeguard stack gives organizations that operate in sensitive or dual-use domains a model that the company says rarely blocks legitimate security work while refusing dangerous requests.
Concrete use cases are visible in the benchmarks SpaceXAI selected. Software engineering and longer-running coding tasks are covered by CursorBench 4.0 and DeepSWE v1.1. Multi-hour terminal work is covered by Terminal-Bench 4.0. Multi-hour office work, producing documents and presentations, is covered by AA Briefcase v1.1 and GDPval. Legal work is measured by the Harvey Legal Agent Benchmark, clinical reasoning by HealthBench Professional, and electrical engineering by EEBench. Cybersecurity is addressed both through the HackerBench v0.3 safety results and through invite-only red-team access for selected cybersecurity partners conducting defense research. Together these scenarios describe an agent that can be pointed at a long task, left to work, and expected to check its own output.
The target audience follows from that positioning: software developers and engineering teams building in Cursor, Grok Build, or through the Grok API; organizations running third-party coding harnesses, model routers, and cloud platforms; and professionals whose multi-hour work involves documents, presentations, analysis, and research, such as those in legal, clinical, financial, and engineering roles. SpaceXAI also targets cybersecurity defenders, both through the model's balance of strong cyber defense capability with low refusal rates for legitimate use and through the invite-only red-team access it has begun granting to select partners.
Grok 4.7's takeaway is a single proposition: SpaceXAI's most capable model for coding and knowledge work, delivered at the same price and speed as Grok 4.6 and described as twice as fast at half the price of comparable models. It pairs a larger base model and longer reinforcement learning on multi-hour tasks with improved self-verification, better long-context management, native Grok Bot harness understanding, and an entirely new safeguard stack that leads the company's testing on refusals and jailbreak resistance. Available in Cursor, Grok Build, the Grok API, third-party harnesses, model routers, and cloud platforms from $2 per million input tokens and $6 per million output tokens, it is built for teams that need long, difficult work completed reliably and affordably.