FULL STORY

Grok 4.5 Launch: Coding Focus and High Cost-Performance

xAI released Grok 4.5, focusing on coding and agentic tasks. With high cost-performance and speed, it topped benchmarks and quickly integrated into popular developer tools.

2026-07-07 ~ 2026-07-16 · 20 episodes · 227 posts

Episode 1 · Musk Announces Grok 4.5 with 1.5T Parameters and Enhanced Coding (2026-07-07, 25 posts)

On July 8, Elon Musk announced that xAI would release the Grok 4.5 model to the public the following day, based on strongly positive feedback from beta testers. He claimed the model achieves Opus-level performance but is faster, more token-efficient, and cheaper. This official news corroborated a wave of recent online leaks.

Key Details and Rumor Roundup

Prior to official confirmation, several leakers and test accounts revealed extensive details. According to information compiled by @tetsuoai and @XFreeze, Grok 4.5 runs on a new V9 base model with 1.5 trillion parameters, three times the size of the previous V8-small (0.5T), making it xAI's largest model to date. The training focused heavily on coding and agentic tasks. Additionally, @nimaowji and @testingcatalog spotted traces of version 4.5 on the Grok web frontend, and @markk leaked that early access would be restricted to SuperGrok Heavy subscribers.

Deep Collaboration with Cursor

Multiple sources indicated a deep collaboration between xAI and the coding tool Cursor. According to internal memos and media reports relayed by @kimmonismus and @ns123abc, the two parties co-developed the model to directly compete with Opus 4.8 and GPT 5.5 in key areas. @haider1 added that the model was trained on Cursor data to boost agentic programming capabilities. However, @Angaisb noted that while they hold low expectations for xAI itself, they trust Cursor's ability to optimize it.

Future Model Roadmap

Alongside the Grok 4.5 announcement, Musk (@elonmusk) shared future development plans. He stated that the Grok Build harness and the 1.5T base model would see continuous daily improvements based on user needs, while the larger Grok 2T model will finish training this month and be made available to customers next month.

5 more related posts →

Episode 2 · Prediction Markets Strongly Price In Grok 4.4 Release (2026-07-07, 2 posts)

Polymarket traders are heavily betting that xAI will ship Grok 4.4 soon: one contract put the odds of a release by July 17 at 94%, while another gave a month-end release 84%. The figures signal strong market expectations, though no official launch has been confirmed.

Episode 3 · Grok 4.5 Released with Focus on Coding and Low Cost (2026-07-08, 61 posts)

xAI has officially released Grok 4.5, targeting coding and agentic use cases. The model enters the market with highly competitive pricing and has quickly appeared on major evaluation leaderboards and coding tools, sparking significant attention and positive feedback from the community.

Pricing and Availability

Grok 4.5 is priced at $2 per million input tokens and $6 per million output tokens. It is confirmed to be available on Grok Build, Cursor, and the SpaceXAI console. In Cursor, its Fast mode is priced at $4/M input and $18/M output tokens.

Benchmark Performance

In Artificial Analysis evaluations, Grok 4.5 scored 54 points, ranking 4th on the Intelligence Index, just behind certain Claude, GPT, and Fable series models. In evaluations of real knowledge work tasks, the average cost per task is around $0.49 to $1.12, taking about 12.4 minutes. Elon Musk retweeted claims that it ranks first in several benchmarks. Additionally, on the CritPt coding evaluation, its performance falls between Opus 4.6-7 and 4.8.

Feedback and Cost-Effectiveness

Multiple users and bloggers highlighted the model's high cost-effectiveness. Tests indicate its coding capability is comparable to GPT-5.5-xhigh but at half the cost, and it is about 17 times cheaper on real tasks than Opus 4.8. In Cursor testing, users felt it provided an excellent experience from ideation to implementation, acting like a faster, cheaper Opus 4.8. Analysts attribute this to the model being trained on high-quality coding trajectories and deeply integrated with Cursor data.

41 more related posts →

Episode 4 · xAI Launches Grok 4.5: Coding and Agent Focus to Rival Opus (2026-07-09, 55 posts)

On July 9, xAI officially launched the Grok 4.5 model. Described as the company's first model specifically trained for coding and agentic tasks, it was developed in collaboration with Cursor. The model is designed for real-world engineering tasks, excelling in large codebases, long multi-repository tasks, and multi-tool orchestration. Elon Musk stated that internal evaluations show Grok 4.5's capabilities are comparable to Claude Opus 4.7, but with faster speeds, lower costs, and maximum intelligence per unit of time and cost.

Core Specs and Pricing

Grok 4.5 features a 500K token context window, speeds of 80 tokens/s, and supports tool calling, structured outputs, and vision. The API is priced at $2 per million input tokens and $6 per million output tokens. Furthermore, the model is highly token-efficient; on the SWE Bench Pro, its average output token usage (15,954) was significantly lower than Claude Opus 4.8 max (67,020), saving approximately 4.2x.

Platform Availability

Grok 4.5 is now live and rolling out across multiple platforms. Users can access it via grok.com, Grok Build (requiring an update to version 0.2.92), Cursor, Hermes Agent, OpenClaw, and the xAI API, as well as gateways like OpenRouter. However, the rollout for the EU region is expected in mid-July.

35 more related posts →

Episode 5 · Grok 4.5 Receives Widespread Praise for Speed and Coding (2026-07-09, 13 posts)

The recent release of Grok 4.5 has sparked widespread discussion in the AI community, with many users giving highly positive reviews after hands-on testing. This marks a significant leap in the Grok model series' capabilities, establishing it as a serious contender among frontier models rather than just a niche product.

Core Experience and Performance Feedback

Multiple users (such as @tetsuoai, @markk, and @StarKnight12) unanimously agreed that Grok 4.5 performs "unexpectedly well." Regarding specific applications, @xiaohu noted that it performs close to top-tier models in simple tasks, writing, and frontend tasks, while also being extremely fast and offering cheap API pricing. After 24 hours of intensive use, @DanielFarinax claimed it provides a frictionless experience that crushes competitors, even prompting him to cancel his Claude Max subscription in favor of Grok Heavy. @chrisfirst also immediately felt a tangible improvement compared to Grok 4.

Coding Capabilities and Workflow Integration

Grok 4.5's programming abilities were particularly praised. @mariofilhoml gave explicit feedback that it is "surprisingly good" for coding, and community consensus suggests its coding skills are highly competitive. Additionally, @tylerbruno05 praised the build TUI experience, and @markk reported excellent performance when integrating the model within the Cursor AI code editor.

Benchmarks and Overall Positioning

In horizontal comparisons with other mainstream models, @HarveenChadha provided a rough ranking, suggesting the hype is real and placing its overall level close to GLM 5.2. Users like @maxpaperclips expressed satisfaction that Grok has finally shed its previous reputation as a "joke" and has become a truly serious and competitive model.

Episode 6 · Grok 4.5 Released, Ranks 6th on Vals Index (2026-07-09, 2 posts)

Grok 4.5 has been released and ranked 6th on the Vals Index with a score of 65.3%. This represents an impressive improvement of nearly 20 percentage points over its predecessor, highlighting significant advancements in the model's capabilities.

Episode 7 · Grok 4.5 Integrates into Notion (2026-07-09, 2 posts)

Notion has integrated Grok 4.5, allowing users to manage meetings, documents, and company knowledge directly within the app. Notion claims it is the most powerful Grok model they have tested, with significant improvements in search-intensive tasks.

Episode 8 · Musk Highlights Grok 4.5's High Performance and Low Cost (2026-07-09, 3 posts)

Elon Musk highlights Grok 4.5's optimal real-world ROI, combining frontier intelligence with lower accessibility barriers. Thanks to efficient architectures, the model maintains high performance even in low-effort mode while significantly saving computational usage.

Episode 9 · Grok 4.5 Tops Multiple Professional AI Benchmarks (2026-07-10, 5 posts)

According to the latest AI knowledge work benchmarks released by Artificial Analysis, Grok 4.5 has demonstrated strong professional capabilities, becoming the best-performing non-Anthropic model currently available. This achievement has drawn attention within the AI community, proving its competitiveness in complex tasks.

Key Details

Grok 4.5 has achieved leading scores in multiple professional work benchmarks. According to @kevinnbass, the model's scores significantly outperform other frontier models such as Opus 4.8 and GPT 5.5. Detailed data from @XFreeze points out that Grok 4.5 took first place in several细分 leaderboards, including AutomationBench-AA, Terminal-Bench v2, Harvey Legal Agent Benchmark, SWE Marathon, and SWE-Atlas.

Areas of Excellence

Multiple authors specifically highlighted Grok 4.5's outstanding performance in legal tasks. Whether in comprehensive evaluations or the specific Harvey Legal Agent Benchmark, the model has shown a prominent capacity for handling professional legal work.

Episode 10 · Grok 4.5 Tops Coding Benchmark with High Token Efficiency (2026-07-10, 3 posts)

xAI's Grok 4.5 paired with Grok Build scored 84 on the SWE-Atlas-QnA benchmark, tying for first with Codex GPT-5.6. The setup also demonstrated exceptional token efficiency, consuming significantly fewer tokens per task than mainstream competitors.

Episode 11 · Grok 4.5 Opens Free Tier and Integrates with Developer Tools (2026-07-10, 9 posts)

Between July 10 and 11, xAI's official Grok account announced that the Grok 4.5 model is officially available for a free tier trial. By simply using an X or SuperGrok account, users can access the model via the Grok Build platform and provide feedback. This move lowered the barrier to entry for developers and quickly attracted community attention.

Model Positioning and Developer Tool Integration

Elon Musk, resharing the official news, specifically emphasized that Grok 4.5 is a new Opus-class model featuring a "fast speed and low cost" profile, making it highly suitable for real-world programming and engineering tasks. To facilitate developer usage, SpaceXAI announced that Grok 4.5 can be accessed for free within the Grok Build CLI and Cursor. Community members even shared installation methods providing a direct curl command-line startup script.

Synchronous Grok Build Version Update

Alongside the opening of Grok 4.5, Grok Build received a version update to 0.2.95. According to developer @markk, this update includes several specific functional improvements, such as supporting the definition of default allowed commands for teams via the managedconfig.toml file, further enhancing the flexibility of team collaboration and tool management.

Episode 12 · Grok 4.5 Joins Perplexity as Orchestrator, Tops WANDR Benchmark (2026-07-11, 10 posts)

Recently, xAI's Grok 4.5 was officially integrated into Perplexity's system, becoming an orchestrator model for Perplexity Computer's Consumer Pro/Max subscribers. This collaboration has drawn significant attention due to Grok 4.5's dominant performance in internal benchmark tests.

Benchmark Data and Performance

In Perplexity's WANDR benchmark, Grok 4.5 was compared against five other orchestrator configurations. Results showed that Grok 4.5 achieved the highest score of 0.328, making it the best-performing frontier model, with a single trial cost of approximately $4.76. Perplexity's Arav Srinivas expressed being deeply impressed by its performance, while Elon Musk emphasized that the core value of Grok Build and Grok 4.5 lies in their genuine usefulness in real-world work scenarios.

Real-World Application Feedback

In terms of actual user experience, user @BWay124 tested Grok Build powered by Grok 4.5 (High). The user noted that Grok Build has reached a completely new level, adapting to any terminal, project, or codebase environment. By simply inputting a goal and removing barriers, it can autonomously complete subsequent tasks, such as smoothly executing database modifications and queries via autonomous CLI operations for the first time. However, the user also suggested adding support for seamlessly resuming conversations across terminals or codebases to improve cross-layer workflows.

Episode 13 · Grok 4.5 Shows Significant Improvements in Coding Capabilities (2026-07-11, 2 posts)

User tests indicate that Grok 4.5 has significantly improved its coding abilities, capable of generating complete code in one go. Although not necessarily top-tier intelligence, it executes well-defined coding tasks with exceptional speed and competence.

Episode 14 · Grok 4.5 Evaluations: Strong Cost-Performance in Mid-to-High Budgets (2026-07-11, 5 posts)

Recent tests and developer feedback indicate that Grok 4.5 offers an outstanding balance of cost and performance, demonstrating high usability particularly in specific budget ranges and practical production applications.

Key Details and Performance

The Xbow research team noted that while Grok 4.5 is not the strongest at all price points, it is the most powerful option they tested in the mid-to-high cost tier where cost is a concern but budgets are not strictly limited. In browser task evaluations, feedback suggests its performance surpasses GPT-5.6-Sol and is only slightly behind Opus. However, due to expensive cached inputs, its overall cost is only about 10% lower than Opus.

Developer Feedback and Production Use

Developer @zeeg shared his experience with safety evaluations, noting that Grok 4.5 performed well. Although he is currently still using Sonnet 4.6, he is highly likely to switch to Grok in the Warden production environment due to the trade-off between price and accuracy, praising its "excellent" accuracy at its current price point. Furthermore, a viewpoint shared by @amanmadaan corroborates this, stating that while single automated benchmarks cannot fully measure usability, Grok 4.5 consistently stands on the cost-performance Pareto frontier in practical tests.

Episode 15 · Grok 4.5 Gains Praise for Speed and Top-Tier Performance (2026-07-11, 2 posts)

Early users of Grok 4.5 report highly positive experiences, highlighting significantly faster response speeds. While not claiming it outright beats top-tier models like Claude Opus, testers note its intelligence is strong enough to compete directly with the industry's best.

Episode 16 · Grok 4.5 First Tests: Faster, Cheaper, and Entering Coding Workflows (2026-07-12, 8 posts)

After its release in mid-July, Grok 4.5 received a wave of first-hand feedback from developers and reviewers. The conclusions are highly consistent: compared to predecessors and peers, it is faster, cheaper, and sufficiently smart, entering coding workflows like Cursor and Amp with high cost-effectiveness. "Faster and more economical" has become the greatest common divisor of this feedback.

Benchmark Scores and Cost-Effectiveness

rasbt included Grok 4.5 and Meta's Muse Spark 1.1 in an updated comparison chart, placing the former on the Pareto frontier of cost-effectiveness. A video studio eval reposted by Elon Musk showed its score jumping from 6/33 to 23/33, with cost efficiency described as outstanding. Other reposts claim Grok 4.5 reached Claude Opus levels in browser usage scenarios, surpassing GPT-5.6-Sol and approaching Opus in one eval. However, due to high cached input costs, it is only about 10% cheaper than Opus overall, though slightly faster. zeeg stated that a weekend benchmark run didn't change his view, still considering Grok 4.5 perhaps the "best value," while noting GPT 5.6 Luna is also competitive on Warden's security bench.

Hands-on Tests in Coding Tools

Multiple developers expressed pleasant surprise testing Grok 4.5 in coding tools like Cursor and Amp: long tasks are very fast and usable, with Cursor pricing at about half of the original. Feedback indicates that, unlike using GPT 5.5 or Claude Opus, Grok 4.5 completes tasks so quickly that users need to return to the session more frequently, altering their workflows. bytebot noted reasonable usage under an $8/month X account plan, and tetsuoai summarized it as "faster, smarter, and cheaper."

Episode 17 · Musk Says Grok 4.5 Beats Fable on Some Coding Benchmarks (2026-07-13, 3 posts)

Elon Musk said Grok 4.5 scores slightly above Fable on some software benchmarks, while calling Fable a very good model. He also argued the publicly available Mythos/Fable version, which uses Claude as a fallback, is noticeably nerfed, fueling debate over its true performance.

Episode 18 · Grok 4.5 Tops Coding Q&A Benchmark with New Collaborative Release (2026-07-13, 3 posts)

Grok 4.5 has achieved the highest score on the SWE-Atlas-QnA benchmark, surpassing competitors. Additionally, SpaceXAI and Cursor jointly released a collaborative version of Grok 4.5, featuring near-frontier coding capabilities and significantly faster processing speeds.

Episode 19 · Grok 4.5 Draws Praise for Speed Lead (2026-07-14, 2 posts)

Posts argue that Grok 4.5’s standout advantage is speed rather than across-the-board capability gains. With throughput cited at about 80–112 TPS, it is described as faster than GPT-5.5 Sol and Claude Fable 5, enabling quicker iteration and a smoother workflow.

Episode 20 · Grok 4.5 shifts focus to coding and agent workflows (2026-07-15, 12 posts)

xAI’s Grok 4.5 is being discussed less as a general-purpose chatbot update and more as a model aimed at software engineering and agentic workflows. Based on the posts in this cluster, the main points of attention are its coding focus, long-running task support, and whether its rollout into tools like Grok Build and Cursor can turn it into part of everyday developer workflows.

Positioning and capabilities

Multiple reposted or relayed posts describe Grok 4.5 as xAI’s first model trained specifically for coding and agents. The messaging emphasizes practical engineering work rather than generic conversation: handling large codebases, working across multiple repositories over long durations, and using hundreds of skills, tools, and MCP. Several posts also repeat official-style claims that it is faster and more cost-efficient, with strong performance in coding and ongoing agentic work.

XFreeze relays Elon Musk’s explanation that early beta access went first to engineers at Tesla and SpaceX. According to that account, the model was improved using feedback from real engineering tasks, with the goal of strengthening real-world usability rather than only benchmark performance.

Rollout and access

Posts from tetsuoai and markk say Grok 4.5 has launched in Europe and can be selected in Grok Build through /model. markk later adds that it has also become available in Cursor, including in Europe. Another repost says the model is free to try.

Current information limits

At this stage, most of the concrete claims in the cluster come from reposts, relayed launch messaging, and early user impressions. There is still relatively little independent, systematic external evaluation in the materials provided here, so the most meaningful near-term signal will be whether developers actually keep using it in real coding workflows.