Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaDoubao Seed 1.8

Doubao-Seed-1.8: Agent, Multimodal, and Thinking-Level Benchmarks from the Official arXiv Model Card

Original source

arXiv

AuthorByteDance Seed

Source date2026-04-17

Tabbit curation2026-08-19

Read original

One-sentence takeaway

The official model card shows that Seed1.8 is strongest in search, GUI, video tools, and adjustable test-time compute, making it suitable for long-running Agents; however, its scores are vendor-reported and include internal tasks and comparisons with technical reports.

Use cases

  • Suitable tasks: multi-step search and evidence synthesis, browser/desktop GUI, video time-window analysis, tool calling, and Agents that need reasoning adjusted to a budget.

  • Unsuitable tasks: directly extrapolating internal benchmark results to business success rates, or using only high-level thinking in pursuit of low latency and low cost.

  • Applicable model version: Seed1.8; the report uses four levels: no_think, think-low, think-medium, and think-high.

  • Applicable clients, Agents, or APIs: ByteDance Seed/Volcengine model services, and orchestration of Search, Code, GUI, and VideoCut tools implemented independently.

  • Recommended reasoning levels and parameters: try no_think/low first for simple extraction; use medium/high for complex retrieval, coding, and video reasoning, and record the tokens, steps, and success rate for each level.

Test environment

  • Evaluation categories: basic language; multimodal vision/video; Agent search/coding/writing/tool/GUI; efficiency; and internal workflows inspired by real-world tasks.

  • Comparison models: GPT-5-high, Claude-Sonnet-4.5, Gemini-2.5-pro, Gemini-3-pro, and others; most non-Seed scores come from their respective technical reports.

  • Configuration: Sections 2.1–2.3 of the report primarily use think-high; Table 1's basic capabilities use no tools by default; Table 7 compares Seed1.8 with increased thinking.

Inputs/configuration

  • Public tasks include AIME-25, LiveCodeBench v6, GPQA-Diamond, BrowseComp-en/zh, GAIA, SWE-bench Verified, Terminal Bench 2.0, BFCL-v4, OSWorld, Online-Mind2web, AndroidWorld, and a long-video collection.

  • The report supports the VideoCut tool: the model specifies a start and end time and 1–5 FPS; the tool resamples video frames before supplying them to the model for analysis.

  • The report does not disclose all questions, system prompts, sampling parameters, or per-question outputs; what can be reproduced is the task names, thinking levels, tool mechanism, and aggregate scores.

Results

Basic and Agent summary

Benchmark/taskSeed1.8 scoreNotes
AIME-2594.3Table 1, think-high, no tools
LiveCodeBench v679.5Table 1, Pass@1
GPQA-Diamond83.8Table 1, Pass@1
MMLU92.3Table 1, Pass@1
Customer Support Q&A (internal)69.0Economic-value task
Complex Workflow (internal)54.6Multi-step execution according to an SOP
GAIA93.2Agentic search; the report says this is higher than GPT-5-high's 76.7
BrowseComp-en / zh67.6 / 78.5Search
WideSearch63.8Search
SWE-bench VerifiedNo separate score is given in the Agent section of the main textListed only as an evaluation item; do not fill in the value from a search snippet
FinSearchComp56.2Financial retrieval and synthesis
XpertBench Finance / Law62.0 / 55.2Internal expert tasks

GUI and video tools

BenchmarkSeed1.8Seed1.8 + tool
OSWorld61.9—
Realbench49.1—
Online-Mind2web85.9—
AndroidWorld70.7—
CGBench62.465.9
LVBench73.078.9
ZeroVideo6.918.8

Increased test-time compute

BenchmarkSeed1.8Increased Thinking
AIME-25 Avg@1094.397.3
HMMT25(Feb) Avg@1089.796.7
LeetCodeBench(v6) Avg@879.484.7
HLE full / text / vision21.4 / 22.1 / 17.325.6 / 26.4 / 20.2

Conclusions

Seed1.8's reusable advantage is putting search, vision, GUI, video tools, and multiple thinking levels behind the same Agent interface; VideoCut delivers a clear gain on long-video tasks. Pure coding or pure knowledge tasks should still be retested with the same harness and target models rather than selected solely on the basis of Agent search scores.

Limitations

  • This is an official model card, combining vendor-reported results, internal benchmarks, and results cited from other technical reports.

  • The task definitions, graders, tool environments, and success criteria for complex workflows are not fully disclosed; table scores cannot be treated directly as business SLAs.

  • think-high/increased thinking increases compute; the model card emphasizes the quality-latency-cost tradeoff but does not provide a standardized token/USD table for each level.

  • The models in the report are named Seed1.8 and should not be conflated with different Volcengine/BytePlus snapshots of Seed1.8 or third-party agent results.

Reproduction steps

  1. Fix the actual Seed1.8 endpoint, region, thinking level, and tool versions; first reproduce the tool-free AIME, coding, and vision baselines.

  2. For search, GUI, and video tasks, disclose the inputs, tool schemas, web/application environments, and termination conditions.

  3. Run no_think, low, medium, and high separately, recording accuracy, tool-call success rate, execution steps, total tokens, latency, and failure types.

  4. For video tasks, fix VideoCut's time window and FPS, and report the difference with and without VideoCut enabled.

  5. Use the same questions, tools, budgets, and repetition counts for external models, and place official self-reported results and measured results in separate columns.

Source excerpt or observation (for compliant short quotation only)

The model card explicitly lists four thinking modes and extends Agent evaluations to search, coding, tool use, GUI, and video tool use; this article retains only aggregate data that can be checked directly on the page.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Doubao Seed 1.8

Use and compare models in Tabbit

Doubao Seed 1.8

Related reviews

MediaPuter Developer

Doubao-Seed-1.8: Observations on the Puter Developer API and Search Agent Selection

CommunityX2025-12-19

Doubao-Seed-1.8: Post-release Multitasking Experience and Boundaries on X

Doubao Seed 1.8

Related prompts

CommunityGitHub / official ByteDance-Seed repository2025-12-19

Doubao-Seed-1.8: ByteDance's Official GitHub Code Agent Closed-Loop Workflow

MediaBytePlus official documentation

Doubao-Seed-1.8: BytePlus Multimodal Deep Thinking and Video Input Configuration