Arena says Peter Gostev directly compared GPT-6 Sol with GPT-5.6 Sol under matched conditions: both used the same prompts and were set to the maximum reasoning level. The post says the evaluation covers generated results, total token counts, and wall-clock time.
The post also explicitly says that Arena scores for GPT-6 Sol and GPT-6 Luna are “coming soon.” It provides no Arena ranking or scores. This post should not be treated as an Arena leaderboard result.
This review opened the original post and its YouTube video, “GPT-6 Sol | First impressions,” through Tabbit. The video page was accessible and showed a duration of about 18 minutes 40 seconds. Its description confirms that the comparison used the same prompts and maximum reasoning level; the video shows examples of generated content. However, the complete prompts, paired raw outputs, per-item token counts, and timings were not available in the readable page text. Therefore, this review records no specific task results or numbers and does not infer which model is better from the on-screen examples.
The video description links to Peter's 3D prompt collection. Because this is not a video linked directly from the original post, it is outside the evidence scope of this entry.
This source is worth including with a limited scope: the same prompts, maximum reasoning level, and comparison dimensions are all stated explicitly in Arena's original post, so it serves as a record of a GPT-6 Sol versus GPT-5.6 Sol comparison setup. Since the item-level data was not verified, this entry supports only the comparison setup, not conclusions about capability, efficiency, or speed. Results can be added once item-level data from the video can be independently checked.
GPT-6 Sol