Benchmark: Box Complex Work Eval, covering real document-driven tasks across twelve industries.
Task types: Reading source documents, checking numbers, due diligence, identifying errors in expert outputs, and quantitative analysis.
Comparison: GPT‑5.6 Sol, Terra, Luna, and GPT‑5.5; the page does not disclose the complete sample count or harness version.
Inputs consisted of document tasks such as multi-year financial forecasts, student-grade recalculations, retail product performance, energy operations reports, life-sciences test records, and clinical cases.
Key acceptance criteria were whether the chain of numbers remained consistent, the denominator was correct, reports were evidence-based, and high-risk judgment errors were avoided.
Box positions the model as an enterprise workflow option for future Box AI / Box AI Studio offerings.
| Industry task | GPT‑5.6 Sol | GPT‑5.5 |
|---|---|---|
| Financial Services | 76% | 71% |
| Public Sector | 74% | 63% |
| Retail | 72% | 66% |
| Energy | 61% | 54% |
| Life Sciences | 60% | 51% |
| Healthcare | 58% | 46% |
Overall, Sol scored about 1% higher than GPT‑5.5, but Box believes the gap is concentrated in more difficult, higher-consequence quantitative tasks.
Across data-analysis tasks, Sol scored 64% versus GPT‑5.5's 57%; the breakdown was 80% versus 63% in retail, 51% versus 27% in life sciences, 62% versus 50% in healthcare, and 78% versus 75% in finance.
For average latency, Terra was about 16% faster than Sol and Luna about 19% faster; in retail inventory/profit analysis, Luna returned a complete answer in roughly half Sol's time.
Sol's advantage looks more like getting the chain of numbers in source documents right and keeping it defensible, rather than leading uniformly across all knowledge work. Try Sol first for high-stakes quantitative analysis; route repetitive, structured, high-throughput tasks to Terra/Luna, using accuracy and latency together.
Box's internal benchmark, graders, sample size, random seeds, and original documents were not made public, so outsiders cannot fully rerun it.
Box is a potential product partner, creating publication-selection and scenario bias; the results cannot replace blind cross-vendor testing.
The page describes the overall improvement as “about 1%” but does not provide the total sample count or confidence interval.
Build a task set spanning twelve industries, retaining source documents, target answers, and numeric-checking rules.
Fix the context, tools, and output format for Sol, Terra, Luna, and GPT‑5.5.
Design separate graders for financial forecasts, denominator calculations, record recalculations, and clinical judgments.
Record task accuracy, latency, cost, and error types; do not record only the overall average score.
Conduct external validation with hidden documents and a second batch of industry tasks.
Box explicitly framed “reading source files, checking numbers, running due diligence, and reviewing expert outputs” as benchmark tasks, rather than a questionnaire about chat preferences.
The article gives Sol/GPT‑5.5 scores for six industries and scores for the data-analysis subset.
The conclusion is better suited to document-driven enterprise analysis; it does not mean Sol is likewise ahead in open-ended writing, frontend visuals, or long-horizon coding.
Latency is reported only as relative percentages, without service-region, concurrency, or token statistics.
This is Box's own evaluation and should not be combined directly with Artificial Analysis scores into a unified ranking.
Box summarized the overall result by saying that Sol “edges past” GPT‑5.5, but then emphasized that the advantage was concentrated in the most difficult quantitative tasks.
GPT-5.6 Sol