# Grok 4.7: The Task Bill Depends on the Route

**Plutonous** | October 5, 2026 | 8 min read

> An October 2 employee report puts Grok's effort settings under scrutiny. API and Cursor context thresholds show why the same model can produce different bills.

Tags: Grok 4.7, xAI, AI Agents, API Pricing, Reasoning Effort, Cursor, Grok Build, Developer Tools

---

**TL;DR: On October 2, Grok Build employee @aksheyd reported CursorBench results of 41.6% at medium effort and $3.49 per task, versus 43.9% and $4.69 at high; these are attributed employee claims, not independently reproduced results.<sup><a href="#source-1">[1]</a></sup> Route choice also matters: xAI's API publishes a 200k long-context threshold, while Cursor documents higher billing only above 256k input tokens.<sup><a href="#source-2">[2]</a></sup><sup><a href="#source-4">[4]</a></sup>**

Grok 4.7 launched on September 21. This October 5 analysis concerns the operating decisions emerging after launch, rather than a new model release. xAI's announcement emphasizes persistence on difficult work.<sup><a href="#source-3">[3]</a></sup> Persistence has an accounting consequence: more thinking and repeated attempts can consume a budget even when the published token rate stays attractive.

The timely signal comes from an October 2 X thread by @aksheyd, whose profile identifies Grok Build at SpaceXAI. He describes complaints about usage limits, overthinking and early compaction, alongside changes to prompts and context settings.<sup><a href="#source-1">[1]</a></sup> That is a useful lead for investigation. It does not establish a universal defect, a subscription-policy change or a guaranteed saving.


### Why This Matters Now

A model name leaves crucial purchase decisions unspecified. Record the provider route, effort setting, speed tier and actual input length before comparing task invoices. Defaults can change the experiment before the first prompt is sent.


*Cover: newly generated editorial ink engraving of a routing switch, clockwork paths and a balance weighing time against completed work. It is illustrative artwork, not a product screenshot or measured evidence.*

## The Employee Signal: A Lower Bill Buys a Different Result

The employee's reported medium-to-high comparison has a $1.20 task-cost gap and a 2.3-percentage-point score gap. Dividing $1.20 by $4.69 gives **25.6% lower reported cost**, rounded to one decimal. That arithmetic describes his figures; it does not turn them into a controlled LLM Rumors evaluation.

The thread does not supply enough execution detail here to extrapolate across repositories, retry policies and task types. We have not reproduced its benchmark. A developer deciding whether medium is sufficient should ask which failures sit inside that score gap, not assume the cheaper setting preserves every important capability.

For a bounded documentation change, extra thought may add little value. For a difficult repair, a failed first attempt can make a cheaper setting expensive. Those are evaluation hypotheses, not measured Grok outcomes. The right local comparison keeps the same tasks and acceptance rules, then charges every retry to its original task.

xAI's reasoning reference makes that experiment concrete: Grok 4.7 supports `low`, `medium`, `high` and `xhigh`; the default is `high`, and reasoning cannot be disabled.<sup><a href="#source-10">[10]</a></sup> Explicitly set the effort in an API trial. Otherwise a comparison labeled “default versus medium” can become a comparison whose expensive side nobody recorded. Treat effort as a controlled input, not a remembered UI choice.

## The Route Boundary: 200k and 256k Are Different Contracts

The direct xAI API lists input/cached input/output rates of **$2/$0.50/$6** per million tokens, becoming **$4/$1/$12** at the long-context tier. That tier covers the whole request once its prompt reaches 200k. The US regional endpoint adds 10%.<sup><a href="#source-2">[2]</a></sup>

Cursor's own documentation instead sets its long-context billing boundary **above 256k input tokens** and lists a 500k maximum window. It also says Fast is the default speed tier on Pro and higher plans.<sup><a href="#source-4">[4]</a></sup> A 256k window therefore does not mean every request incurs a long-context premium. Measure the actual input, then apply the contract for the route used.

This discrepancy is practical, not semantic. Do not paste direct-API billing assumptions into a Cursor spreadsheet. Keep route-specific calculators, and recheck their source pages when defaults change. The same model label is insufficient evidence that two bills follow identical rules.

## The Speed Choice: Pay for Waiting Time Deliberately

Fast serves the same model at a premium through Cursor/Grok Build; it is absent from the public API.<sup><a href="#source-2">[2]</a></sup> Speed and effort are separate purchasing decisions.

We do not rank its speed against other models here. Hardware, concurrency, prompt length and latency conditions are not normalized. For an interactive repair, time saved may justify a premium; for an overnight queue, it may not. Measure elapsed time and accepted output together. The current model reference also marks Batch API unsupported, so an assumed batch discount is not a Grok 4.7 cost plan.<sup><a href="#source-5">[5]</a></sup>

## Astra Ultrafast: Speed Is a Service Tier, Not a Task Guarantee

OpenAI's Astra Ultrafast illustrates the same purchasing problem on another route. OpenAI claims **up to 8x faster token generation** than Astra Standard in Codex, explicitly not 8x faster task completion. App access requires Pro $500 or eligible Enterprise/Edu plans. Included usage runs at 8x Standard; credits and Enterprise pay-as-you-go run at 6x, subject to the workspace agreement.<sup><a href="#source-13">[13]</a></sup> Those billing multipliers and the speed claim describe different measurements.

The API has its own access contract: Astra Ultrafast is broadly available at initial rate limits using `gpt-6-astra` with `service_tier: "ultrafast"`. US residency and global processing are supported; non-US regional endpoints are excluded. OpenAI recommends WebSockets because connection overhead can reduce latency gains.<sup><a href="#source-14">[14]</a></sup> A Pro subscription requirement should not be imported into this API route.

OpenAI lists Astra Ultrafast short-context input/output at **$60/$300 per million tokens**, versus **$120/$450** for long context; cache reads and writes have separate rates.<sup><a href="#source-15">[15]</a></sup> These are API dollars, not the app's allowance multipliers. OpenAI Fast, formerly Priority, is a distinct tier.<sup><a href="#source-15">[15]</a></sup> None of those figures establishes whether Astra or Grok finishes a particular task more economically.

For a purchasing test, hold the task and acceptance criteria fixed, then measure how much elapsed time actually comes from generation. A workload dominated by tool execution, human review or external services may gain less from faster tokens. Compare each speed option against its own Standard baseline before making a cross-model decision. The business question is how much verified waiting time the premium removes.

## The Context Choice: Carry State Without Carrying Everything

In an [October 2 author follow-up](https://x.com/aksheyd/status/2105908838700650707), @aksheyd says the 256k setting automatically compacts before the full window and supports manual `/compact`. This qualifies the root post's context-cut story: window capacity, compaction timing and billing boundaries are separate controls. His explanation is a Grok Build claim, not evidence that Cursor uses the direct API's billing threshold.

xAI's compaction guide describes replacing conversation history with an opaque state item and recommends doing so before the context limit is exceeded.<sup><a href="#source-6">[6]</a></sup> That is an API mechanism, not proof that any particular Cursor or Build session preserved the right details.

Our recommendation is to define a checkpoint contract: unfinished requirements, failed approaches, exact files and validation results. After compaction, ask the agent to recover that state and inspect omissions. Keep authoritative requirements in durable project files. A smaller input is only economical if it avoids rediscovering or repeating work.

Cache reuse also needs evidence. xAI exposes cached-token counts separately for Chat Completions and Responses; its caching guide says cached tokens still count toward the total prompt length that triggers long-context pricing. It bills reasoning tokens at the completion rate.<sup><a href="#source-11">[11]</a></sup> A large cache hit therefore does not establish a short-context bill. Log the full input length and the reused portion separately, and include reasoning when measuring output consumption.

The public Grok Build repository exposes a terminal agent runtime with tools for reading, editing and shell execution.<sup><a href="#source-7">[7]</a></sup> This makes the harness part of the outcome. Tool descriptions, retrieval choices and permissions deserve a place in the evaluation record alongside effort.

## The Buying Decision: Evaluate the Whole Delivery Route

AWS's September 28 Bedrock post describes cross-Region inference profiles and Responses, Chat Completions and Converse access.<sup><a href="#source-8">[8]</a></sup> GitHub's September 21 Copilot announcement describes a gradual rollout and usage-based billing at provider list pricing.<sup><a href="#source-9">[9]</a></sup> Neither announcement establishes identical defaults, subscription allowances or regional economics across platforms.

For the direct API, xAI documents `usage.cost_in_usd_ticks` as the billed amount for each request, including server-side tools and applicable discounts; dividing by 10,000,000,000 converts it to dollars.<sup><a href="#source-12">[12]</a></sup> Sum those request amounts across the task rather than treating the last response as the session total. This gives an API evaluation a bill-backed numerator. It does not measure human repair or establish what a Cursor subscription will charge.

Choose a route for its integration and controls, then compare finished work inside that route. Keep separate rows for ordinary tasks and difficult tasks. Record tokens, speed tier, retries, elapsed time and human repair. This gives medium effort a fair trial without hiding expensive failures inside a cheaper average.


### The Invoice Needs a Route

The employee's benchmark figures are not a production forecast. Cursor's 256k billing boundary and xAI's 200k API threshold must stay attached to their respective routes. A model name alone cannot price a task.


The uncomfortable truth is that buying a capable model leaves much of the operating discipline unresolved. Grok 4.7 deserves an evaluation that prices how it is served, how long it thinks and what survives review. The valuable default is the one that finishes your work at a defensible total cost.


## Sources

<a id="source-1"></a>
1. [October 2 Grok Build task-cost thread](https://x.com/aksheyd/status/2105825185698066453)

<a id="source-2"></a>
2. [xAI API pricing](https://docs.x.ai/developers/pricing)

<a id="source-3"></a>
3. [Introducing Grok 4.7](https://x.ai/news/grok-4-7)

<a id="source-4"></a>
4. [Grok 4.7 in Cursor](https://cursor.com/docs/models/grok-4-7)

<a id="source-5"></a>
5. [Grok 4.7 model reference](https://docs.x.ai/developers/models/grok-4.7)

<a id="source-6"></a>
6. [Context compaction guide](https://docs.x.ai/developers/advanced-api-usage/context-compaction)

<a id="source-7"></a>
7. [Grok Build public repository](https://github.com/xai-org/grok-build)

<a id="source-8"></a>
8. [Grok 4.7 on Amazon Bedrock](https://aws.amazon.com/blogs/machine-learning/grok-4-7-is-now-available-on-amazon-bedrock/)

<a id="source-9"></a>
9. [Grok 4.7 in GitHub Copilot](https://github.blog/changelog/2026-09-21-grok-4-7-is-now-available-in-github-copilot/)

<a id="source-10"></a>
10. [Reasoning reference](https://docs.x.ai/developers/model-capabilities/text/reasoning)

<a id="source-11"></a>
11. [Prompt caching: usage and pricing](https://docs.x.ai/developers/advanced-api-usage/prompt-caching/usage-and-pricing)

<a id="source-12"></a>
12. [Cost tracking](https://docs.x.ai/developers/cost-tracking)

<a id="source-13"></a>
13. [Codex and ChatGPT Work speed](https://learn.chatgpt.com/docs/agent-configuration/speed)

<a id="source-14"></a>
14. [Ultrafast mode in the API](https://developers.openai.com/api/docs/guides/ultrafast-mode)

<a id="source-15"></a>
15. [OpenAI API pricing](https://developers.openai.com/api/docs/pricing)


*Last updated: October 5, 2026 (SGT)*

---

*Source: [LLM Rumors](https://www.llmrumors.com/news/grok-47-task-cost-context-routing)*
