Grok 4.7: Specs, Benchmarks, Capabilities, and Modelflare Pricing
A sourced guide to Grok 4.7 covering the published specification, reasoning levels, launch benchmarks, official token prices, and Modelflare grok-award and grok-stable rates checked on 21 September 2026.
xAI released Grok 4.7 on 21 September 2026. The API model ID is grok-4.7. It is xAI's current flagship for coding, agentic tool use, and knowledge work. Modelflare listed that same ID in the public pricing catalog on the same day.
This article keeps three layers apart. The specification is what xAI documents. The benchmark table is what xAI reported at launch. The Modelflare prices are what GET /api/pricing returned on 21 September 2026. A vendor score is not a Modelflare test, and a catalog row is not a promise that every key, route, or prompt will behave the same way.
xAI has not published a parameter count, a layer count, or a training-token total for Grok 4.7. "Larger than Grok 4.6" is the company's qualitative description of the new base model. The comparison below uses the figures that are actually public: context, token prices, reasoning levels, modalities, and the launch scores.
What was published
The launch post describes Grok 4.7 as a new, larger base model than Grok 4.6, trained with a longer reinforcement-learning run. The task mix is harder and weighted toward problems that take many hours. xAI says the model is better at checking its own work and at managing long context, and that it was trained to understand the Grok Bot harness used for conversational work.
Two sentences in that post should not be collapsed into one. The headline says the model is "twice as fast, at half the price of comparable models." The body says it is served at the same price and the same speed as Grok 4.6. The token table supports the second sentence against Grok 4.6: both list $2 per million input tokens and $6 per million output tokens. The headline is a comparison with other models in that table, whose list prices are higher. xAI does not publish a tokens-per-second number next to "twice as fast," so this article does not treat that phrase as a measured latency result.
grok-4.7 is on the xAI API, in Cursor, and is the default model in Grok Build. A separate Fast serving path exists only in Cursor and Grok Build. It is not on the public xAI API, and it was not in the Modelflare catalog checked for this article.
Specification comparison
Prices below are xAI's published list prices in USD per million tokens. "Short" means a prompt under 200,000 tokens. "Long" means a prompt that has reached 200,000 tokens. On the price table, crossing that line reprices every token in the request, not only the tokens past the line.
| Item | grok-4.7 | grok-4.6 | grok-4.5 |
|---|---|---|---|
| Context window | 500,000 tokens | 500,000 tokens | 500,000 tokens |
| Short input / cached input / output | $2.00 / $0.50 / $6.00 | $2.00 / $0.50 / $6.00 | $2.00 / $0.30 / $6.00 |
| Long input / cached input / output | $4.00 / $1.00 / $12.00 | $4.00 / $1.00 / $12.00 | $4.00 / $0.60 / $12.00 |
| Reasoning effort | low, medium, high (default), xhigh | low, medium, high (default), xhigh | low, medium, high (default) |
Grok 4.7 details checked on the model page the same day:
- Inputs are text and images. Output is text. The model does not generate images.
- The page states no fixed text-output limit. Client timeouts, account limits, and the remaining context window still apply.
- The knowledge cutoff is May 2026. Later events are invisible unless a search tool is enabled.
- Documented interfaces are the Responses API and Chat Completions.
- Documented tools include function calling, web search, X search, and code execution.
- Reasoning cannot be switched off. If you omit the effort, the default is
high. - On the Responses API,
grok-4.7returnsreasoning.encrypted_contenteven when the request does not ask for it. Later turns should send those reasoning items back unchanged. Chat Completions does not gain that behavior.
Grok 4.5 stops at high. Grok 4.6 and Grok 4.7 add xhigh. Cached-input list price also moved: Grok 4.5 is $0.30 on short prompts and $0.60 on long prompts; Grok 4.6 and Grok 4.7 are $0.50 and $1.00. Standard input and output list prices are unchanged across these three IDs.
Capability dimensions
Coding and long agent work
The launch post aims Grok 4.7 at work that runs for a long time: multi-step coding, checking a result before calling it done, and holding a long context together. These coding scores all come from xAI's launch table:
- CursorBench 4.0, which stresses longer coding tasks: 46.3% at xHigh. Grok 4.6 High is 40.4%, GPT-5.6 Sol Max is 41.7%, and Fable 5.1 Max is 51.8%.
- DeepSWE v1.1: 71.0%, marked by xAI as a high-effort score rather than xHigh. Grok 4.6 is 65.2%, Sol is 72.7%, and Fable 5.1 is 70.0%.
- Terminal-Bench 4.0: 38.0%. Grok 4.6 is 20.3%, Sol is 37.3%, and Fable 5.1 is 57.9%.
- EEBench, electrical engineering: 64.0%. Grok 4.6 is 53.0%, Sol is 39.4%, and Fable 5.1 is 56.4%.
Read the rows as a map of where the vendor says the model moved, not as the score your repository will get. On CursorBench, xAI calls the result frontier price-performance: the score is above Sol and Grok 4.6, below Fable 5.1, at a much lower list price than either Sol or Fable. DeepSWE is close to Fable and still below Sol, and it is not an xHigh number. Terminal work improved sharply versus Grok 4.6 and sits near Sol, while Fable 5.1 remains well ahead on that row. EEBench is the clearest lead in the table.
Professional knowledge work
xAI also reports office and professional tasks:
- AA Briefcase v1.1, multi-hour office work: 1,657. Grok 4.6 is 1,546, Sol is 1,487, and Fable 5.1 is 1,678.
- Harvey Legal Agent Benchmark: 19.6%. Grok 4.6 is 15.8%, Sol is 2.5%, and Fable 5.1 is 6.7%. Leading this table still means about one in five items. It is not evidence that the model can carry legal work unsupervised.
- HealthBench Professional, clinical reasoning: 56.7%. Grok 4.6 is 48.5%, Sol is 60.5%, and Fable 5.1 is 62.1%. On this row Grok 4.7 is ahead of its predecessor and behind both comparison models.
The post says Grok 4.7 is better at documents and presentations, and that it improves on Grok 4.6 on GDPval while landing in the same range as other frontier models. The prose does not print the GDPval point values, so this article does not supply any.
There is no single intelligence index in the Grok 4.7 launch text. Grok 4.6's earlier post published an Artificial Analysis Intelligence Index. This launch replaces that one number with the task table above. Use the row that matches the work. A model can lead electrical engineering and trail clinical reasoning in the same release.
Reasoning effort
reasoning_effort accepts low, medium, high, and xhigh. high is the default. xAI describes low as the lighter setting, medium as more thinking where latency is less sensitive, high as the setting for difficult multi-step problems, and xhigh as the deepest setting, with the highest latency, for problems where answer quality dominates response time.
More effort usually means more reasoning tokens, and output tokens are the expensive side of the list price. xhigh is not automatically the right production default. Compare task success, latency, and total tokens on your own prompts before you leave it on.
For a multi-turn Responses conversation, keep the encrypted reasoning items in the next request. Dropping them discards the context the model used to reach the previous answer. Chat Completions does not return that encrypted field.
Tools, images, and memory
Function calling is the tool path you implement. You declare the function, the model proposes a call, your application runs it, and you send the result back. Web search, X search, and code execution are hosted by xAI. The model can call them on its own, and each invocation is priced separately from tokens.
Images can be sent as input. The model-page summary does not list a separate image price. Image inputs consume context and appear in the request's token usage. Image generation stays on the Grok Imagine models, which are a different API.
Five hundred thousand tokens is the window, not a budget you should fill. xAI recommends a stable prompt_cache_key on the Responses API, and the x-grok-conv-id header on direct Chat Completions, so a conversation is more likely to land on the same server and hit cache. Without that affinity, a cache-cold server bills the full input price. Long loops are also expected to compact older turns. Compaction changes what the model can see. Check the summary against the original constraints, IDs, decisions, and open errors.
The May 2026 cutoff is a hard limit on memorized knowledge. Anything newer requires a search tool, and search is not included in the token price.
Safeguards
xAI says Grok 4.7 uses a new safeguard stack and is the strongest model it has tested on refusals and jailbreak resistance. Two figures are printed in the launch post:
- LatchBio's biosafety benchmark: 62.4%.
- HackerBench v0.3, which xAI describes as its benchmark of risky and malicious cyber tasks: 3.3% of risky dual-use prompts are allowed through. xAI also says legitimate security work is rarely blocked.
The post says selected cybersecurity partners have invite-only access to red-team capabilities for defense research. That access is not part of the public grok-4.7 API described here.
These figures are vendor evaluations. They are not an independent audit, and they are not a reason to skip human review on medical, legal, security, or other high-stakes answers.
Official prices
Token rates
The xAI price page, checked on 21 September 2026, lists USD per million tokens:
| Prompt size | Input | Cached input | Output |
|---|---|---|---|
| Under 200,000 prompt tokens | $2.00 | $0.50 | $6.00 |
| At or above 200,000 prompt tokens | $4.00 | $1.00 | $12.00 |
The long rate covers the whole request once the prompt reaches the threshold. The first 200,000 tokens do not keep the short rate.
Three official examples
Hosted tool fees are excluded.
| Request | Basis | Arithmetic | Total |
|---|---|---|---|
| 100,000 uncached input, 5,000 output | Short | 0.1 × $2 + 0.005 × $6 | $0.230 |
| 100,000 cached input, 5,000 output | Short | 0.1 × $0.50 + 0.005 × $6 | $0.080 |
| 220,000 uncached input, 10,000 output | Long | 0.22 × $4 + 0.01 × $12 | $1.000 |
The jump from the first row to the third is not "a bit more text." Crossing 200,000 tokens doubles every rate, and the extra tokens are billed at the doubled rate. A fully cached 100,000-token prompt in the second row costs about one third of the uncached short prompt, because cached input is one quarter of the input rate and the output charge does not change.
Hosted tools, Fast, and the US region
Hosted tool invocations on the same price page:
- Web search (
web_search): $5 per 1,000 calls. - X search (
x_search): $5 per 1,000 calls at the time of this check. Starting 21 September 2026 at 12:00 PT, xAI replaces that call price with $5 per 1,000 posts fetched and $10 per 1,000 user profiles fetched. Every returned post counts, including parents and quotes. Every returned profile counts. - Code execution (
code_executionandcode_interpreter): $5 per 1,000 calls. - File-attachment search: $10 per 1,000 calls.
- Collections search: $2.50 per 1,000 calls.
The model chooses how many hosted calls to make. A search-heavy answer can cost more in tool fees than in tokens.
The US regional endpoint https://us.api.x.ai/v1 keeps inference in the United States and multiplies input, cached input, and output by 1.1, including long-context rates. For grok-4.7 that is $2.20 / $0.55 / $6.60 under 200,000 prompt tokens, and $4.40 / $1.10 / $13.20 at or above that line.
Grok 4.7 Fast is the same model on faster infrastructure. xAI's prose says it is billed at twice the standard token rates. The price table prints $4.00 / $1.00 / $12.00 below 200,000 prompt tokens, and $6.00 / $1.50 / $18.00 at or above 200,000. Doubling the standard long-context row of $4 / $1 / $12 would be $8 / $2 / $24. The printed Fast long-context row does not match that doubling. Budget from the table, and recheck the price page before you rely on Fast. Fast is sold through Cursor and Grok Build. It is not on the public xAI API, it is not included in Grok Build's free tier, and the Modelflare catalog checked on 21 September 2026 had no Fast model ID. The Grok IDs present were grok-4.5, grok-4.6, and grok-4.7.
Launch benchmarks
The scores below are copied from xAI's launch post of 21 September 2026. Folkbench compares routes on availability, latency, and price, and keeps its notes on this model at folkbench.com/models/grok-4-7. Column headers keep the effort setting xAI printed. The other models' scores are the ones xAI placed in the same table. Modelflare did not re-run them.
| Evaluation | Grok 4.7 xHigh | Grok 4.6 High | GPT-5.6 Sol Max | Fable 5.1 Max |
|---|---|---|---|---|
| List input price, USD / 1M | 2 | 2 | 4 | 10 |
| List output price, USD / 1M | 6 | 6 | 20 | 50 |
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0% (high, not xHigh) | 65.2% | 72.7% | 70.0% |
| EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
The Grok 4.7 column is not one effort setting. DeepSWE is explicitly a high-effort score. The other Grok 4.7 cells sit under the xHigh header. Grok 4.6 is High. Sol and Fable are Max. A higher score can come from a deeper and more expensive reasoning setting. That is worth knowing, and it is not a matched comparison at a single effort level.
On this table Grok 4.7 leads EEBench and the Harvey benchmark. It is close to the leaders on CursorBench, DeepSWE, AA Briefcase, and on Terminal-Bench against Sol. It trails Fable 5.1 by a wide margin on Terminal-Bench, and it trails both Sol and Fable 5.1 on HealthBench Professional.
Input list price is half of Sol and one fifth of Fable. Output list price is 30% of Sol and 12% of Fable. "Half the price" fits the Sol input cell and does not fit every other cell. The table is the precise comparison.
Modelflare prices and how the groups differ
On 21 September 2026, Modelflare's public pricing response included grok-4.7. The catalog list rate matches xAI's short-context list price: input $2 per million tokens, a completion ratio of 3 so output is $6, and a cache ratio of 0.25 so cached input is $0.50. That row does not publish a second long-context tier. Do not multiply xAI's $4 / $1 / $12 long rate by a Modelflare group ratio and treat the product as the bill. The charge stored on the finished request is the figure that was applied.
The group is selected on the API key. Both groups call the same model ID.
Price table
| Group | Ratio | Input / 1M | Cached input / 1M | Output / 1M | Catalog description |
|---|---|---|---|---|---|
grok-award |
0.03 | $0.06 | $0.015 | $0.18 | Price-priority routing for budget-sensitive, retryable development, testing, and flexible workloads |
grok-stable |
0.11 | $0.22 | $0.055 | $0.66 | Everyday development, longer projects, and production applications for individual developers who want stability |
The arithmetic is the catalog list rate times the group ratio. Award is $2.00 × 0.03, $0.50 × 0.03, and $6.00 × 0.03. Stable is $2.00 × 0.11, $0.50 × 0.11, and $6.00 × 0.11. Against the short-context list price, those ratios are 3% and 11%. They are not 3% or 11% of the long-context list price, of the US regional price, or of hosted tool fees.
The Chinese catalog text for grok-award adds a sentence the English description does not: the very low price does not guarantee stability or model quality. Keep that next to the 0.03 ratio. grok-stable is the catalog's description for ordinary development and production. Neither sentence is a measured uptime agreement, a latency target, or a claim that the groups run different weights.
Three Modelflare examples
The token shapes match the official examples. They are priced on the catalog's single rate. Cached rows assume the usage record counts those input tokens as cache hits. Very small charges can also move when quota is rounded to the billing unit, so the request log remains the invoice.
| Request | grok-award | grok-stable |
|---|---|---|
| 100,000 uncached input, 5,000 output | $0.0069 | $0.0253 |
| 100,000 cached input, 5,000 output | $0.0024 | $0.0088 |
| 220,000 uncached input, 10,000 output, at the single catalog rate | $0.0150 | $0.0550 |
The 220,000-token row is intentionally not 0.03 × $1.00 or 0.11 × $1.00. One dollar is xAI's long-context list charge. If a later catalog revision adds a long-context tier, this row stops being the quote. The live page is the grok-4.7 price page. The index is Models and prices.
Calling the model through Modelflare, rather than opening a separate xAI account only for this ID, means:
- The model ID stays
grok-4.7. The client does not target a renamed alias. - The canonical base URL is the OpenAI-compatible
https://modelflare.dev/v1. An OpenAI SDK usually needs that base URL, a Modelflare key, and the model ID. - The key is placed in
grok-awardorgrok-stable. The same account can call other catalog models when the key's group allows them. - The price split is explicit. Award is the cheap route for work you can retry, with the stability and quality caveat above. Stable is the route the catalog describes for everyday development and production.
- The listed endpoint types for this ID are Chat Completions, Responses, and Responses compaction.
This price does not include a measured latency advantage over calling xAI directly, hosted search or code-execution fees, the Fast serving tier, or a service level beyond the group description.
How to call Grok 4.7 on Modelflare
Create an API key and select grok-award or grok-stable. The group lives on the key. A key in an unrelated group can fail on grok-4.7 even though the model is listed publicly. Send the model ID exactly. The canonical base URL is https://modelflare.dev/v1. The same API is also published at https://cf.modelflare.dev/v1 and https://origin.modelflare.dev/v1. The root hostname is the canonical site.
Chat Completions:
export MODELFLARE_API_KEY="YOUR_MODELFLARE_API_KEY"
curl -sS https://modelflare.dev/v1/chat/completions \
-H "Authorization: Bearer $MODELFLARE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-4.7",
"reasoning_effort": "high",
"messages": [
{
"role": "user",
"content": "Review this migration plan and list the three assumptions that most need a test."
}
]
}'
Responses:
curl -sS https://modelflare.dev/v1/responses \
-H "Authorization: Bearer $MODELFLARE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-4.7",
"reasoning": {"effort": "high"},
"input": "Review this migration plan and list the three assumptions that most need a test."
}'
On Chat Completions, reasoning_effort accepts low, medium, high, and xhigh. On Responses, set reasoning.effort to the same values. Image input uses the ordinary OpenAI content array on a user message. Your own functions use the ordinary tools field. POST /v1/responses/compact is listed for this model and is forwarded as compaction on grok-4.7, not as another model ID. Hosted xAI tools, including web search, are a provider behavior. Try them on the route you will actually use, and read the usage record, before you depend on them or on xAI's invocation fees.
For the difference between the two request shapes, see Responses API versus Chat Completions. For the cost method behind the tables, see how LLM token cost is calculated and AI API cost tracking.
Checks before you switch traffic
- Pin
grok-4.7. Do not accept a silent substitution of Grok 4.6 or Grok 4.5. - Confirm the key's group is
grok-awardorgrok-stable, and choose it from retry tolerance rather than from the ratio alone. - Replay one fixed, non-sensitive set at
low,medium,high, andxhigh. Record success, latency, input tokens, cached input, output tokens, and the charge. - Run Grok 4.6 on that same set before you replace it. The list price is similar. The score table is not uniform.
- Price one prompt just under 200,000 tokens and one just over it, then read the recorded charge. The xAI sheet and the Modelflare catalog row do not currently publish the same long-context rule.
- If the workflow needs images, function calls, streaming, or cancellation, test those paths. One text-only success does not cover them.
- Keep a rollback route. A completed HTTP response only shows that one call finished.
Sources
Checked on 21 September 2026.
- Introducing Grok 4.7
- Grok 4.7 developer guide
- xAI models and pricing
- xAI API pricing
- xAI release notes
- Reasoning effort
- Modelflare public pricing catalog,
GET https://modelflare.dev/api/pricing, modelgrok-4.7 - Earlier guide: Grok 4.6 API, pricing, and 500K context