GPT-5.6 Luna · Community source · Editorial analysis
This evidence note covers “Agents on Rails: 8 Models, 21 Atomic Tasks” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
Google's index shows that the Rails team compared 8 models on 21 atomic tasks. The summary reports these results for Luna: 73% of runs succeeded with the default medium reasoning effort, the total cost of 63 runs was approximately 90 cents, and it was labeled the cheapest model.
Agents on Rails: We ran 8 models against 21 atomic tasks to see which were ... Cheapest: @OpenAI GPT-5.6 Luna. 73% of ru… This is a necessary excerpt; read the original source for full context.
The body of the X post did not expand in the Tabbit international edition, so this note preserves only the original excerpt visible in Google's index. The task list, success criteria, costs for the other models, and runtime environment are missing; therefore, the 73% figure must not be treated as a general coding pass rate.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
X · @rails (Ruby on Rails) · Original publication date Unknown · Site edit date 2026-09-20
Open original sourceGPT-5.6 Luna
Download the Tabbit client to check model access