Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Doubao Seed 1.8 · Media / benchmark · Vendor report

Doubao-Seed-1.8: Agent, Multimodal, and Thinking-Level Benchmarks from the Official arXiv Model Card

The arXiv model card reports Seed 1.8 results for Agent, multimodal, and thinking-level tasks; the finding is limited to those tasks and is not a production success rate.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkVendor reportEdited 2026-09-20

Test conditions

Condition
The arXiv model card’s Agent, multimodal, and thinking-level tasks; interpret only within the paper version and reported setup.
Sample/date
Sample and repeat counts follow the card; reopened 2026-09-20, with undisclosed parameters left unknown.

Key data and applicable tasks

One-sentence takeaway

The official model card shows that Seed1.8 is strongest in search, GUI, video tools, and adjustable test-time compute, making it suitable for long-running Agents; however, its scores are vendor-reported and include internal tasks and comparisons with technical reports.

Use cases

  • Suitable tasks: multi-step search and evidence synthesis, browser/desktop GUI, video time-window analysis, tool calling, and Agents that need reasoning adjusted to a budget.

  • Unsuitable tasks: directly extrapolating internal benchmark results to business success rates, or using only high-level thinking in pursuit of low latency and low cost.

  • Applicable model version: Seed1.8; the report uses four levels: no_think, think-low, think-medium, and think-high.

  • Applicable clients, Agents, or APIs: ByteDance Seed/Volcengine model services, and orchestration of Search, Code, GUI, and VideoCut tools implemented independently.

  • Recommended reasoning levels and parameters: try no_think/low first for simple extraction; use medium/high for complex retrieval, coding, and video reasoning, and record the tokens, steps, and success rate for each level.

Test environment

  • Evaluation categories: basic language; multimodal vision/video; Agent search/coding/writing/tool/GUI; efficiency; and internal workflows inspired by real-world tasks.

  • Comparison models: GPT-5-high, Claude-Sonnet-4.5, Gemini-2.5-pro, Gemini-3-pro, and others; most non-Seed scores come from their respective technical reports.

  • Configuration: Sections 2.1–2.3 of the report primarily use think-high; Table 1's basic capabilities use no tools by default; Table 7 compares Seed1.8 with increased thinking.

Inputs/configuration

  • Public tasks include AIME-25, LiveCodeBench v6, GPQA-Diamond, BrowseComp-en/zh, GAIA, SWE-bench Verified, Terminal Bench 2.0, BFCL-v4, OSWorld, Online-Mind2web, AndroidWorld, and a long-video collection.

  • The report supports the VideoCut tool: the model specifies a start and end time and 1–5 FPS; the tool resamples video frames before supplying them to the model for analysis.

  • The report does not disclose all questions, system prompts, sampling parameters, or per-question outputs; what can be reproduced is the task names, thinking levels, tool mechanism, and aggregate scores.

Results

Basic and Agent summary

Benchmark/taskSeed1.8 scoreNotes
AIME-2594.3Table 1, think-high, no tools
LiveCodeBench v679.5Table 1, Pass@1
GPQA-Diamond83.8Table 1, Pass@1
MMLU92.3Table 1, Pass@1
Customer Support Q&A (internal)69.0Economic-value task
Complex Workflow (internal)54.6Multi-step execution according to an SOP
GAIA93.2Agentic search; the report says this is higher than GPT-5-high's 76.7
BrowseComp-en / zh67.6 / 78.5Search
WideSearch63.8Search
SWE-bench VerifiedNo separate score is given in the Agent section of the main textListed only as an evaluation item; do not fill in the value from a search snippet
FinSearchComp56.2Financial retrieval and synthesis
XpertBench Finance / Law62.0 / 55.2Internal expert tasks

GUI and video tools

BenchmarkSeed1.8Seed1.8 + tool
OSWorld61.9—
Realbench49.1—
Online-Mind2web85.9—
AndroidWorld70.7—
CGBench62.465.9
LVBench73.078.9
ZeroVideo6.918.8

Increased test-time compute

BenchmarkSeed1.8Increased Thinking
AIME-25 Avg@1094.397.3
HMMT25(Feb) Avg@1089.796.7
LeetCodeBench(v6) Avg@879.484.7
HLE full / text / vision21.4 / 22.1 / 17.325.6 / 26.4 / 20.2

Conclusions

Seed1.8's reusable advantage is putting search, vision, GUI, video tools, and multiple thinking levels behind the same Agent interface; VideoCut delivers a clear gain on long-video tasks. Pure coding or pure knowledge tasks should still be retested with the same harness and target models rather than selected solely on the basis of Agent search scores.

Limitations

  • This is an official model card, combining vendor-reported results, internal benchmarks, and results cited from other technical reports.

  • The task definitions, graders, tool environments, and success criteria for complex workflows are not fully disclosed; table scores cannot be treated directly as business SLAs.

  • think-high/increased thinking increases compute; the model card emphasizes the quality-latency-cost tradeoff but does not provide a standardized token/USD table for each level.

  • The models in the report are named Seed1.8 and should not be conflated with different Volcengine/BytePlus snapshots of Seed1.8 or third-party agent results.

Reproduction steps

  1. Fix the actual Seed1.8 endpoint, region, thinking level, and tool versions; first reproduce the tool-free AIME, coding, and vision baselines.

  2. For search, GUI, and video tasks, disclose the inputs, tool schemas, web/application environments, and termination conditions.

  3. Run no_think, low, medium, and high separately, recording accuracy, tool-call success rate, execution steps, total tokens, latency, and failure types.

  4. For video tasks, fix VideoCut's time window and FPS, and report the difference with and without VideoCut enabled.

  5. Use the same questions, tools, budgets, and repetition counts for external models, and place official self-reported results and measured results in separate columns.

Source excerpt or observation (for compliant short quotation only)

The model card explicitly lists four thinking modes and extends Agent evaluations to search, coding, tool use, GUI, and video tool use; this article retains only aggregate data that can be checked directly on the page.

What this supports

  • Supports interpreting the model-card results for Agent, multimodal, and thinking-level tasks within the paper’s task set.

What this does not support

  • The benchmark is bounded by the paper’s protocol and reported setup; it does not prove production success or cross-harness ranking.
  • Samples, versions, prices, and quotas are not controlled by this source and need separate verification.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

arXiv · ByteDance Seed · Original publication date 2026-04-17 · Site edit date 2026-09-20

Open original source

Doubao Seed 1.8

Compare Doubao Seed 1.8 in Tabbit

Download the Tabbit client to check model access

Related reviews

Doubao-Seed-1.8: Observations on the Puter Developer API and Search Agent SelectionThe Puter page records Seed 1.8 API and search-agent integration observations; provider routing, quotas, and latency are uncontrolled, so this is not native Ark reliability evidence.Doubao-Seed-1.8: Post-release Multitasking Experience and Boundaries on XZhihuFrontier’s post records post-release multitask boundaries for Seed 1.8; it has no fixed sample or reproducible protocol and cannot represent overall success.Doubao-Seed-1.8: BytePlus Multimodal Deep Thinking and Video Input ConfigurationFollow a task-specific guide for “Doubao-Seed-1.8: BytePlus Multimodal Deep Thinking and Video Input Configuration”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Doubao-Seed-1.8: ByteDance's Official GitHub Code Agent Closed-Loop WorkflowFollow a task-specific guide for “Doubao-Seed-1.8: ByteDance's Official GitHub Code Agent Closed-Loop Workflow”; prerequisites, steps, checks, fixes, and source boundaries are explicit.