Grok 4.7 · Community source · Independent measurement
On the AA-Briefcase Elo chart attached to this post, Grok 4.7 (xhigh) is 1657, behind Claude Fable 5.1 (max with fallback, 1678) and Claude Opus 5 (max, 1673).
On the AA-Briefcase Elo chart attached to this post, Grok 4.7 (xhigh) is 1657, behind Claude Fable 5.1 (max with fallback, 1678) and Claude Opus 5 (max, 1673). The post body also says that on this benchmark it trails only Anthropic models and ranks just behind Opus 5, at about half the latter’s cost per task; on AA-Briefcase-Lite, Analytical Quality Elo rises from Grok 4.6’s 1698 to 1994, and Presentation Elo falls from 1531 to 1499.
Tasks suitable for a judgment: Agentic knowledge work as defined by Artificial Analysis, and the public due-diligence scenario named in the post: building a market model and producing a target-assessment deck.
Tasks unsuitable for extrapolation: Coding, multimodal work, general conversation, and any benchmark that does not appear on this chart or in this body text. The composite Elo on the chart cannot be substituted for a score from another harness.
Applicable model versions: The body and the chart axis say Grok 4.7; the comparison bar is Grok 4.6. Finer checkpoints and API model IDs are not stated.
Test environment or client: The caption states AA-Briefcase v1.1, developed by Artificial Analysis. The specific API provider, client, and run date are not stated.
Reasoning tiers and parameters: On the chart, both Grok 4.7 and Grok 4.6 are xhigh. The API cost of the example slides is also marked xhigh. Other tiers are not stated.
The caption on the post’s first image states that AA-Briefcase v1.1 is an agentic knowledge-work benchmark developed by Artificial Analysis, and that AA-Briefcase Elo is a composite metric aggregating rubric pass rate, Analytical Quality Elo, and Presentation Elo, with higher being better. Each bar has an error bar above it, but the values at the ends of those bars were not read off one by one in this screenshot.
The body places Grok 4.7’s change relative to Grok 4.6 on AA-Briefcase-Lite: the author calls it a public due-diligence scenario, where the task is to build a market model and produce a target-assessment deck. Sample size, prompts, the judge model, and the number of runs are not stated.
Under the same permalink there are also two follow-up posts by the author. They only compare how the decks are written and do not give Elo again.
The visible body and links on this permalink do not include an artificialanalysis.ai article URL. The numbers below come only from the post’s English original and the first chart.
English original (the body after clicking “Show original”):
Grok 4.7 is behind only Anthropic models on AA-Briefcase, ranking just behind Opus 5 at ~50% of its Cost per Task Grok 4… This is a necessary excerpt; read the original source for full context.
Checked against the chart:
Composite AA-Briefcase Elo: Grok 4.7 (xhigh) 1657, Grok 4.6 (xhigh) 1555. The arrow on the chart points from 1657 to 1555.
The two bars above 1657 are both Anthropic models: Claude Fable 5.1 (max with fallback) 1678, Claude Opus 5 (max) 1673. The next bar down is Claude Opus 5 (xhigh) 1649. This is consistent with the body’s claim that it “trails only Anthropic models and ranks just behind Opus 5”; first place on the chart is Fable 5.1, not Opus 5.
The cost-per-task figure “about ~50% of Opus 5” appears only in the body. This Elo chart has no cost axis and does not give Opus 5’s cost per task in dollars.
1698 → 1994 and 1531 → 1499 are the AA-Briefcase-Lite Analytical Quality Elo and Presentation Elo in the body, not the composite Elo on the chart. 1657 cannot be subtracted directly from 1994.
API cost to produce the example decks: Grok 4.7 (xhigh) about $8, Grok 4.6 (xhigh) about $4.40. This is the cost of the example decks, not the “cost per task” above.
Bar values and axis labels that could be aligned in this pass, from left to right:
| Order | Visible axis label | AA-Briefcase Elo |
|---|---|---|
| 1 | Claude Fable 5.1 (max with fallback) | 1678 |
| 2 | Claude Opus 5 (max) | 1673 |
| 3 | Grok 4.7 (xhigh) | 1657 |
| 4 | Claude Opus 5 (xhigh) | 1649 |
| 5 | Qwen3.8 Max (0902) | 1640 |
| 6 | Muse Spark 1.3 (max) | 1597 |
| 7 | Qwen3.8-Flash-Next | 1597 |
| 8 | Claude Fable 5.1 (high with fallback) | 1592 |
| 9 | GPT-6 Astra (max) | 1569 |
| 10 | Grok 4.6 (xhigh) | 1555 |
| 11 | GLM-5.3 (max) | 1525 |
| 12 | Kimi K3 (max) | 1510 |
| 13 | GPT-5.6 Sol (max) | 1487 |
| 14 | GLM-5.3-Flash | 1459 |
| 15 | Step 5 Preview | 1432 |
| 16 | DeepSeek V4.1 Flash (max) | 1432 |
| 17 | Qwen3.8 27B (xhigh) | 1403 |
| 18 | GPT-5.6 Luna (max) | 1345 |
| 19 | GPT-5.6 Terra (max) | 1336 |
| 20 | K2 Horizon … A23B | 1303 |
| 21 | DeepSeek V4 Pro 0813 (max) | 1261 |
| 22 | Gemini 3.8 Flash (high) | 1202 |
The parenthetical tiers for rows 14 and 15 were not clearly readable on the axis. Row 20 can be read as K2 Horizon and A23B; the parameter-scale number in the middle is unstable across different crops, so no specific B count is recorded. Error bars are visible; the values at their ends were not recorded.
Follow-up posts by the author on the same page (not another set of scores):
https://x.com/ArtificialAnlys/status/2102170395130708064: Example 1. The task is comparable-benchmark analysis in a private-equity presentation template, with a step-by-step valuation chain. Grok 4.6 gives a short analysis and lands on the deal partner’s rough estimate; Grok 4.7 runs the valuation chain independently and marks the difference. Slide screenshots are attached; cells in the tables were not transcribed cell by cell.
https://x.com/ArtificialAnlys/status/2102170398339313801: Example 2. The task is an asset profile covering scale and financials. Grok 4.6 uses only the latest year, with no revenue-trend or currency discussion; Grok 4.7 reviews three years of history and concludes that growth is offset by currency depreciation, which affects the purchase decision. Slide screenshots are attached; cells in the tables were not transcribed cell by cell.
This post gives Artificial Analysis’s own snapshot of AA-Briefcase v1.1 composite Elo, plus a section of AA-Briefcase-Lite component Elo and the API cost of the example decks. In the visible bar order, Grok 4.7 (xhigh) ranks third, above Grok 4.6 (xhigh, 1555) on the same chart, and also above GPT-6 Astra (max, 1569) on the chart.
This is not a basis for saying that Grok 4.7 is close to Opus 5 on every task. The claim that cost per task is about half is not paired with a dollar unit price or a task definition. The rise in analytical quality and the small drop in Presentation Elo occur in the Lite scenario, and the deck API cost rises from about $4.40 to about $8. The example posts only illustrate differences in how two due-diligence slides are written.
On 2026-09-22 the permalink above was opened, “Show original” was clicked, and the English body was transcribed; the bar values and axis labels on the first chart were enlarged and read. The page could be read directly; no login wall or captcha was encountered. The post does not state the number of items, the prompts, the judging script, or the raw run logs, so 1657 cannot be rerun from this post alone. Slide screenshots in replies on the same page are qualitative examples only; table numbers that were not clearly read are not written into this note.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
X · Artificial Analysis (@ArtificialAnlys) · Original publication date 2026-09-22 · Site edit date 2026-09-22
Open original sourceGrok 4.7
Download the Tabbit client to check model access