Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
CommunityDeepSeek V4 Flash

I Ran DeepSeek V4 Flash on 8 Agent Harnesses (Reddit r/DeepSeek)

Original source

Reddit r/DeepSeek

Authoru/LimpComedian1317

Tabbit curation2026-08-19

Read original

Background

The author believes that model–harness fit (how well a model fits a framework) is real: the same model can perform very differently on different harnesses. After using OpenCode and Hermes to run DeepSeek V4 Flash in everyday work and seeing a clear cost difference, the author benchmarked DeepSeek V4 Flash (via OpenRouter) on 8 popular harnesses: 25 real-world automation tasks involving multiple applications (Slack, Sheets, Gmail, PostHog, and more).

Results table

HarnessPass rateMedian timeTool callsCost per success
Pi Agent66.7%132.2s443$0.028
Prime Agent62.5%*242.1s502$0.131
OMP56.7%272.4s390$0.103
Claude Code53.3%122.7s358$0.195
Codex53.3%245.0s448$0.081
DeepAgents53.3%187.1s353$0.045
Hermes Agent50.0%175.5s386$0.056+
OpenCode46.7%129.7s419$0.073

Key findings

Pass rate and tool calls:

  • Pi Agent had the highest pass rate at 66.7%; OpenCode had the lowest at 46.7%.

  • More tool calls do not necessarily produce better results: DeepAgents (353 calls) and Codex (448 calls) each passed 16 tasks; OMP (390 calls) passed 17; OpenCode (419 calls) passed only 14.

  • Prime Agent passed 15 of 24 valid runs and made the most tool calls (502).

Cost and tokens:

  • Claude Code had the highest cost per successful run at $0.195; Pi had the lowest at $0.028.

  • Claude Code and OMP each used about 742K tokens per task, but Claude Code was nearly twice as expensive: its cache-hit rate was only 1.5% (Codex 70%, OMP 57%).

  • Prime Agent used 1.4M tokens per task, the most of any harness; Hermes used the fewest, at about 192K.

Time:

  • Claude Code had the shortest median time at 122.7s; OpenCode took 129.7s; Pi took 132.2s; OMP was the longest at 272.4s, but passed one more task than Claude Code.

Conclusion

  • Pi Agent is the best harness for DeepSeek V4 Flash: it had the highest accuracy and was the cheapest.

  • Claude Code is the biggest money burner.

  • Model–framework fit (cache utilization and tool-call efficiency) has a significant impact on real-world cost and results.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

DeepSeek V4 Flash

Use and compare models in Tabbit

DeepSeek V4 Flash

Related reviews

OfficialDeepSeek API Docs2026-07-31

DeepSeek-V4-Flash: 0731 Benchmark Update and Harness Conditions

MediaMindStudio2026-08-01

DeepSeek-V4-Flash: Local Deployment, Quantization, and Agent Testing

MediaLightning AI blog2026-04-27

DeepSeek V4 Alters Everything We Knew About Price-Performance Math (Lightning AI)

MediaBenchLM.ai (model benchmarking and pricing tracking site)2026-07-31

DeepSeek V4 Flash 0731 Benchmarks, Pricing & Speed (BenchLM)

DeepSeek V4 Flash

Related prompts

OfficialDeepSeek API Docs2026-07-31

DeepSeek-V4-Flash: 0731 Thinking Levels and Tool-Calling Configuration

OfficialDeepSeek API Docs2026-07-31

DeepSeek-V4-Flash: Codex Responses API integration workflow

CommunityGitHub repository victorchen96/deepseekv4rolepalyinstruct

A Guide to Special Control Instructions for DeepSeek-V4 Role-Playing (Thinking-Mode Switching Guide)

CommunityReddit r/SillyTavernAI

DeepSeek V4 RP Guide — How to Switch Between Character Immersion & Pure Analysis Thinking Modes (Reddit r/SillyTavernAI)