Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

DeepSeek V4 Flash · Community source · Independent measurement

I Ran DeepSeek V4 Flash on 8 Agent Harnesses (Reddit r/DeepSeek)

A Reddit author compares eight harnesses on OpenRouter across 25 automation tasks: Pi Agent passes 66.7% versus OpenCode 46.7%, with about $0.028 versus $0.073 per successful task.

Community sourceIndependent measurementEdited 2026-09-20

Test conditions

Test/source conditions
OpenRouter; eight harnesses; 25 automation tasks; method follows the linked post
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh

Key data and applicable tasks

Background

The author believes that model–harness fit (how well a model fits a framework) is real: the same model can perform very differently on different harnesses. After using OpenCode and Hermes to run DeepSeek V4 Flash in everyday work and seeing a clear cost difference, the author benchmarked DeepSeek V4 Flash (via OpenRouter) on 8 popular harnesses: 25 real-world automation tasks involving multiple applications (Slack, Sheets, Gmail, PostHog, and more).

Results table

HarnessPass rateMedian timeTool callsCost per success
Pi Agent66.7%132.2s443$0.028
Prime Agent62.5%*242.1s502$0.131
OMP56.7%272.4s390$0.103
Claude Code53.3%122.7s358$0.195
Codex53.3%245.0s448$0.081
DeepAgents53.3%187.1s353$0.045
Hermes Agent50.0%175.5s386$0.056+
OpenCode46.7%129.7s419$0.073

Key findings

Pass rate and tool calls:

  • Pi Agent had the highest pass rate at 66.7%; OpenCode had the lowest at 46.7%.

  • More tool calls do not necessarily produce better results: DeepAgents (353 calls) and Codex (448 calls) each passed 16 tasks; OMP (390 calls) passed 17; OpenCode (419 calls) passed only 14.

  • Prime Agent passed 15 of 24 valid runs and made the most tool calls (502).

Cost and tokens:

  • Claude Code had the highest cost per successful run at $0.195; Pi had the lowest at $0.028.

  • Claude Code and OMP each used about 742K tokens per task, but Claude Code was nearly twice as expensive: its cache-hit rate was only 1.5% (Codex 70%, OMP 57%).

  • Prime Agent used 1.4M tokens per task, the most of any harness; Hermes used the fewest, at about 192K.

Time:

  • Claude Code had the shortest median time at 122.7s; OpenCode took 129.7s; Pi took 132.2s; OMP was the longest at 272.4s, but passed one more task than Claude Code.

Conclusion

  • Pi Agent is the best harness for DeepSeek V4 Flash: it had the highest accuracy and was the cheapest.

  • Claude Code is the biggest money burner.

  • Model–framework fit (cache utilization and tool-call efficiency) has a significant impact on real-world cost and results.

What this supports

  • Supports the observation that harness fit changes pass rate, tool calls, time, and cost per success for the same model.

What this does not support

  • Does not support a bare-model ranking, cross-version cost trend, or a universal success rate from one author’s task set.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Reddit r/DeepSeek · u/LimpComedian1317 · Original publication date Unknown · Site edit date 2026-09-20

Open original source

DeepSeek V4 Flash

Compare DeepSeek V4 Flash in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Pricing · English

DeepSeek V4 Flash Pricing: What You Pay in 2026

DeepSeek V4 Flash pricing changed with the V4.1 migration. See the current cache, peak-hour, output, and workload cost math before you budget.

Related reviews

DeepSeek-V4-Flash: 0731 Benchmark Update and Harness ConditionsThe official 0731 table reports Terminal Bench 2.1 82.7, DeepSWE 54.4, and Toolathlon Verified 70.3 under DeepSeek Harness minimal mode, max effort, top_p 0.95, and temperature 1.0.DeepSeek-V4-Flash: Local Deployment, Quantization, and Agent TestingMindStudio reports about 25–30 tokens/s on dual DGX Spark, estimates 168GB for 4-bit and 110GB for 3-bit, and still records live-data errors in two small projects.DeepSeek V4 Flash doesn't like us? (Reddit r/opencodeCLI)A user reports that changing 10 lines consumed 28% of quota after the increase, and a commit message raised it to 32%; no token ledger is supplied.DeepSeek V4 Flash is a monster! Cheap & Good, and so fast (Reddit r/opencodeCLI)An OpenCode user subjectively finds Flash cheap and fast enough to replace parts of a Claude/GPT workflow, while describing Pro as slower and costlier.Connect DeepSeek to Codex with the official configurationBack up local configuration, use the official script or minimal provider fields, and verify with a reversible task.Delegate in layers and synthesize a monograph with DSHThe source publishes a complete starting prompt with research, pushback, editing, and final synthesis targeting one cited Markdown artifact.Configure reasoning tiers and continue tool callsTurn low/high/max, tool results, and reasoning_content handoff into a checkable integration path.Choose a DeepSeek-to-Codex integration pathA case roundup comparing integration routes; it is not one complete copyable template.