Grok 4.7 climbs the SWE-Together leaderboard after audit

As seen on the 24/7 Wall St. homepage on September 23, 2026.

Elon Musk @elonmusk Quote post

xAI's coding model climbed the SWE-Together leaderboard after maintainers re-ran every trial, and coding benchmarks are the scoreboard enterprise AI budgets increasingly follow.

Grok 4.7 moves up in ranking https://t.co/dDZqhkBJBV [Quoted @zhuokaiz]: After we fixed the weak spots exposed by Grok 4.7 (thank you, Grok), we audited every model we have run on the SWE-Together leaderboard for the same behavior, re-ran every trial that got through, and updated the rows. Here is what changed. We scanned the tool calls of all https://t.co/SKggd1AR6i https://t.co/32qJG7MB5x
  • Replies57
  • Reposts33
  • Likes178
Continue ReadingShow less

Elon Musk highlighted the rise on X after Zhuokai Zhao, who maintains the SWE-Together leaderboard, posted a detailed account of what triggered the reshuffle. Grok 4.7 had exposed weak spots in how trials were being evaluated, prompting the team to audit every model they had ever benchmarked on the leaderboard.

The maintainers did not simply patch Grok 4.7's results. They scanned the tool calls of every model on the board, re-ran every trial that had slipped through under the old methodology, and updated the rows across the board. That kind of comprehensive re-evaluation is rare and lends the revised standings more credibility than a typical incremental update.

Sponsored

Twelve Tabs, One Thesis

Your Research Resets Every Morning

The quote page in one tab. Filings in another. A chart you rebuilt from scratch, a transcript you never went back and found, a screener whose settings you will redo next week. Nothing you built yesterday is still there.

AlphaSpace replaces all of it with one screen you arrange yourself. Earnings calendar, estimate versus actual, the call transcript, live news, your own charts, every panel wired to whatever ticker you click. Close the browser and it is all still sitting there tomorrow.

See What a Built View Looks Like →

(Sponsor)

SWE-Together is a coding-focused benchmark, and coding performance is increasingly the metric enterprise buyers use to compare AI models when allocating software development budgets. A leaderboard climb that survives a full audit carries more weight than one posted under the original, looser conditions.

Musk's amplification of the result signals that xAI is leaning into benchmark transparency as a competitive differentiator. Whether Grok 4.7's ranking holds as other models face the same scrutiny is what to watch from here.