Grok 4.7 scores 1657 Elo on AA-Briefcase, closing in on Claude

As seen on the 24/7 Wall St. homepage on September 21, 2026.

xAI just closed most of the gap with Anthropic on the benchmark that tracks real professional work, a 111 Elo jump in one model generation. Enterprise agent contracts are the prize here.

Grok 4.7 joins the frontier on agentic knowledge work tasks. On AA-Briefcase, which evaluates models on realistic professional work tasks, Grok 4.7 scores 1657 Elo, up 111 from Grok 4.6 (high) and placing it just behind Claude Opus 5 and Claude Fable 5.1. Grok 4.7's improvement https://t.co/lBhd9h9mE2
  • Replies2
  • Reposts2
  • Likes26
Continue ReadingShow less

AA-Briefcase is a benchmark built around realistic professional work tasks, the kind of agentic knowledge work that enterprise customers actually pay for. It scores models on an Elo scale, so the gap between two scores reflects competitive distance in a meaningful, head-to-head sense.

Grok 4.7 lands at 1657 Elo on that benchmark, a gain of 111 points over Grok 4.6 (high). That is a substantial jump in a single model generation and puts xAI's latest just behind Claude Opus 5 and Claude Fable 5.1 in the rankings.

Sponsored

Twelve Tabs, One Thesis

Your Research Resets Every Morning

The quote page in one tab. Filings in another. A chart you rebuilt from scratch, a transcript you never went back and found, a screener whose settings you will redo next week. Nothing you built yesterday is still there.

AlphaSpace replaces all of it with one screen you arrange yourself. Earnings calendar, estimate versus actual, the call transcript, live news, your own charts, every panel wired to whatever ticker you click. Close the browser and it is all still sitting there tomorrow.

See What a Built View Looks Like →

(Sponsor)

xAI has moved from trailing Anthropic's frontier models by a wide margin to sitting immediately behind them on the benchmark most directly tied to enterprise agent deployments. One more comparable jump would put Grok at the top of that leaderboard.

Enterprise agent contracts are where the commercial stakes sit. Businesses procuring AI for professional workflows lean heavily on benchmarks like AA-Briefcase when evaluating vendors, which makes Grok 4.7's positioning a concrete competitive development for both xAI and Anthropic.