The official model card shows that Seed1.8 is strongest in search, GUI, video tools, and adjustable test-time compute, making it suitable for long-running Agents; however, its scores are vendor-reported and include internal tasks and comparisons with technical reports.
Suitable tasks: multi-step search and evidence synthesis, browser/desktop GUI, video time-window analysis, tool calling, and Agents that need reasoning adjusted to a budget.
Unsuitable tasks: directly extrapolating internal benchmark results to business success rates, or using only high-level thinking in pursuit of low latency and low cost.
Applicable model version: Seed1.8; the report uses four levels: no_think, think-low, think-medium, and think-high.
Applicable clients, Agents, or APIs: ByteDance Seed/Volcengine model services, and orchestration of Search, Code, GUI, and VideoCut tools implemented independently.
Recommended reasoning levels and parameters: try no_think/low first for simple extraction; use medium/high for complex retrieval, coding, and video reasoning, and record the tokens, steps, and success rate for each level.
Evaluation categories: basic language; multimodal vision/video; Agent search/coding/writing/tool/GUI; efficiency; and internal workflows inspired by real-world tasks.
Comparison models: GPT-5-high, Claude-Sonnet-4.5, Gemini-2.5-pro, Gemini-3-pro, and others; most non-Seed scores come from their respective technical reports.
Configuration: Sections 2.1–2.3 of the report primarily use think-high; Table 1's basic capabilities use no tools by default; Table 7 compares Seed1.8 with increased thinking.
Public tasks include AIME-25, LiveCodeBench v6, GPQA-Diamond, BrowseComp-en/zh, GAIA, SWE-bench Verified, Terminal Bench 2.0, BFCL-v4, OSWorld, Online-Mind2web, AndroidWorld, and a long-video collection.
The report supports the VideoCut tool: the model specifies a start and end time and 1–5 FPS; the tool resamples video frames before supplying them to the model for analysis.
The report does not disclose all questions, system prompts, sampling parameters, or per-question outputs; what can be reproduced is the task names, thinking levels, tool mechanism, and aggregate scores.
| Benchmark/task | Seed1.8 score | Notes |
|---|---|---|
| AIME-25 | 94.3 | Table 1, think-high, no tools |
| LiveCodeBench v6 | 79.5 | Table 1, Pass@1 |
| GPQA-Diamond | 83.8 | Table 1, Pass@1 |
| MMLU | 92.3 | Table 1, Pass@1 |
| Customer Support Q&A (internal) | 69.0 | Economic-value task |
| Complex Workflow (internal) | 54.6 | Multi-step execution according to an SOP |
| GAIA | 93.2 | Agentic search; the report says this is higher than GPT-5-high's 76.7 |
| BrowseComp-en / zh | 67.6 / 78.5 | Search |
| WideSearch | 63.8 | Search |
| SWE-bench Verified | No separate score is given in the Agent section of the main text | Listed only as an evaluation item; do not fill in the value from a search snippet |
| FinSearchComp | 56.2 | Financial retrieval and synthesis |
| XpertBench Finance / Law | 62.0 / 55.2 | Internal expert tasks |
| Benchmark | Seed1.8 | Seed1.8 + tool |
|---|---|---|
| OSWorld | 61.9 | — |
| Realbench | 49.1 | — |
| Online-Mind2web | 85.9 | — |
| AndroidWorld | 70.7 | — |
| CGBench | 62.4 | 65.9 |
| LVBench | 73.0 | 78.9 |
| ZeroVideo | 6.9 | 18.8 |
| Benchmark | Seed1.8 | Increased Thinking |
|---|---|---|
| AIME-25 Avg@10 | 94.3 | 97.3 |
| HMMT25(Feb) Avg@10 | 89.7 | 96.7 |
| LeetCodeBench(v6) Avg@8 | 79.4 | 84.7 |
| HLE full / text / vision | 21.4 / 22.1 / 17.3 | 25.6 / 26.4 / 20.2 |
Seed1.8's reusable advantage is putting search, vision, GUI, video tools, and multiple thinking levels behind the same Agent interface; VideoCut delivers a clear gain on long-video tasks. Pure coding or pure knowledge tasks should still be retested with the same harness and target models rather than selected solely on the basis of Agent search scores.
This is an official model card, combining vendor-reported results, internal benchmarks, and results cited from other technical reports.
The task definitions, graders, tool environments, and success criteria for complex workflows are not fully disclosed; table scores cannot be treated directly as business SLAs.
think-high/increased thinking increases compute; the model card emphasizes the quality-latency-cost tradeoff but does not provide a standardized token/USD table for each level.
The models in the report are named Seed1.8 and should not be conflated with different Volcengine/BytePlus snapshots of Seed1.8 or third-party agent results.
Fix the actual Seed1.8 endpoint, region, thinking level, and tool versions; first reproduce the tool-free AIME, coding, and vision baselines.
For search, GUI, and video tasks, disclose the inputs, tool schemas, web/application environments, and termination conditions.
Run no_think, low, medium, and high separately, recording accuracy, tool-call success rate, execution steps, total tokens, latency, and failure types.
For video tasks, fix VideoCut's time window and FPS, and report the difference with and without VideoCut enabled.
Use the same questions, tools, budgets, and repetition counts for external models, and place official self-reported results and measured results in separate columns.
The model card explicitly lists four thinking modes and extends Agent evaluations to search, coding, tool use, GUI, and video tool use; this article retains only aggregate data that can be checked directly on the page.
Doubao Seed 1.8