The author conducted a blind test by tasking GPT-5.6 Sol, Terra, Luna, and Claude Fable 5 with the same iPhone UI porting assignment under identical prompts and a one-hour time limit. The test demonstrates that real-world UI and code delivery hinges on spec comprehension and functional completeness, though the original post did not disclose Terra's itemized scores.
Task: Four models build the same iPhone UI / port the same app.
Compared models: GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, and Claude Fable 5.
Controlled conditions: The author stated publicly that identical prompts and an identical one-hour time limit were used in a blind evaluation. Video chapters also covered landscape mode, time parsing, external displays, auto-hiding controls, error lists, and remote Mac QA.
Output format: The X post linked to the author's video and chapters; the X post text did not provide the full prompt, code, itemized scorecards, or the final rankings for every model.
What is visible in the original post is the test design, not a fully reproducible prompt. Publicly disclosed test conditions:
Identical design documents, codebase, prompts, and a one-hour budget; four models independently port the same iPhone UI, followed by a blind evaluation.The video chapters revealed verifiable checkpoints: missing features, landscape mode, the "5pm" time feature, external displays, auto-hiding controls, error lists, and remote Mac QA.
In the chapters, the author recorded a score of 93 for Claude Fable 5; the text of the X post did not provide the scores for Terra, Sol, or Luna.
The author emphasized that only one model carefully read the design document and code thoroughly enough to deliver functionally usable features. The specific identity of that "one" model must be determined from the video reveal and code evidence, and cannot be assumed to be Terra or any other specific model solely from the X summary.
This provides a highly relevant testing lead regarding whether Terra is suitable for workflows featuring "existing design docs, existing codebases, and a need to nail UI details," though the currently visible data is insufficient to determine Terra's performance outcome. Future replications should treat requirement comprehension, functional completeness, and device QA as the primary evaluation metrics rather than judging solely by first-draft visual appearance.
The full prompt, repository, model configurations, run counts, itemized scores, and final video reveal were not disclosed in the text of the X post.
The one-hour time budget, the author's specific codebase, and manual scoring in the video influence the results; they cannot be extrapolated to all iOS or frontend tasks.
The Fable score of 93 is an isolated data point from the author's video chapters and should not be treated as a standardized benchmark score.
Prepare identical design documents, initial codebases, and access endpoints for the four models, fixing a one-hour budget and tool permissions.
For each model, preserve the full prompt, commit diffs, execution logs, and final build artifacts.
Blind the model identities, then score them across missing features, landscape mode, time parsing, external displays, auto-hiding controls, test pass rates, and human usability.
Report itemized results, failure root causes, and elapsed time; do not report only an aggregate score.
GPT-5.6 Terra