The discussion reflects disagreement over the credibility of aggregate leaderboards, price cuts, and real-world usage time. It can serve as a checklist of model-selection questions, but not as an independent evaluation.
Environment: A Reddit community discussion whose title mentions Artificial Analysis Intelligence Index rankings; commenters used different subscriptions, model effort levels, and workflows.
Topics: Leaderboard weighting, real coding tasks, GPT-6 Luna and Sol pricing and performance, and Codex usage.
The visible post body does not include a benchmark data table, task inputs, or configuration. The comments do not use a consistent model effort level, prompt, task, or measurement method.
One commenter questioned the aggregate leaderboard, saying another model's rank did not match their personal experience and advising caution about relying on a single ranking.
Some commenters felt that Luna 6 high was cheaper and provided longer usage time. One said their own Codex usage time increased after using Sol 6 medium for orchestration and Luna 6 high as a worker. That commenter did not publish verifiable billing data or a same-task comparison.
Other users said they trusted hands-on model experience more than a single leaderboard. Views were mixed.
This discussion supports evaluating capability, price, usage limits, and workflow cost separately when choosing a model. The rankings and usage-time claims in the comments are personal reports; they cannot be used to estimate overall user performance or reproduce Artificial Analysis's evaluation.
No benchmark chart or raw data is visible in the post body, so the ranking in its title cannot be independently verified from this page alone.
The commenters are a self-selected sample and did not use standardized tasks, account tiers, or measurement methods.
Usage windows and subscription limits affect any claimed “sustainable work time.”
This record covers only the post and comments visible when the page was opened; it does not treat the title or comments as verified model scores.
Fix the Codex access point, model effort levels, and task set under the same account and billing period.
For Luna, Sol, and a baseline, record task completion rate, tokens used, cost, limit consumption, and wait time.
Publish the test tasks, configuration, and raw measurements; report aggregate leaderboard results separately from product usage limits.
GPT-6 Luna