GPT-6 Sol gets better at coding and cuts cost in half, while Luna slips
As seen on the 24/7 Wall St. homepage on September 22, 2026.
Better coding scores at half the cost per task is the number that matters for anyone paying for agentic inference, and the Luna slip shows the gains are not uniform across the same model family.
GPT-6 Sol (max) gains 2 points in the Artificial Analysis Coding Agent Index at half the Cost per Task of its predecessor, while GPT-6 Luna (max) regresses by 2 points. https://t.co/HIdz1bx3UR
- Replies1
- Reposts3
- Likes25
Continue ReadingShow less
GPT-6 Sol in its max configuration picked up 2 points on the Artificial Analysis Coding Agent Index while delivering that performance at half the Cost per Task of its predecessor. For teams running agentic workloads at scale, that combination of a higher benchmark score and a lower cost per task is the kind of efficiency jump that changes procurement math.
The Luna side of the ledger tells a different story. GPT-6 Luna (max) gave back 2 points on the same index, a reminder that gains within a model family are not guaranteed to flow evenly across every variant. Sol and Luna are siblings, but their trajectories on this benchmark are moving in opposite directions.
Sponsored
Twelve Tabs, One Thesis
Your Research Resets Every Morning
The quote page in one tab. Filings in another. A chart you rebuilt from scratch, a transcript you never went back and found, a screener whose settings you will redo next week. Nothing you built yesterday is still there.
AlphaSpace replaces all of it with one screen you arrange yourself. Earnings calendar, estimate versus actual, the call transcript, live news, your own charts, every panel wired to whatever ticker you click. Close the browser and it is all still sitting there tomorrow.
See What a Built View Looks Like →
(Sponsor)
For anyone paying for agentic inference, cost per task is often the number that determines whether a model is deployable at production volume. A halving of that figure alongside a score improvement means Sol is cheaper while doing more. That is an unusual combination in benchmark releases, where price cuts typically come with some performance trade-off.
The Luna regression is worth watching in subsequent updates. A two-point drop is not catastrophic, but it signals that the optimization work that benefited Sol did not translate, and came at Luna's expense. Artificial Analysis is expected to update its index as new model versions are released, so the gap between the two variants will shift quickly.