Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaGemini 3.8 Flash

Gemini 3.8 Flash: Vals AI Harvey Legal Agent Benchmark Re-evaluation

Original source

Vals AI

AuthorVals AI

Source date2026-09-05

Tabbit curation2026-09-08

Read original

One-sentence takeaway

On Vals AI's Harvey's Legal Agent Benchmark v1 held-out test set, Gemini 3.8 Flash achieved a strict Task Pass Rate of 10.00% ± 2.42%, ranking 11th out of 59 systems; the page also shows a Criteria Pass Rate of 90.19% ± 0.36%. It passes most individual criteria, but only a small number of tasks pass every criterion, so 10% should not be interpreted as general legal capability or directly deliverable legal accuracy.

Task set and scope

  • Task subject: Completing legal work with files and tools, rather than answering isolated legal questions; inputs may include documents, spreadsheets, presentations, and file-system materials.

  • Data split: The page says there is currently a public set and a held-out test set; the results on this page come from the held-out set. The page's fallback auxiliary description uses 120 tasks as the denominator. This note records that as the task scale for the current run, but the page body does not disclose the task list item by item.

  • Coverage: The page charts provide 24 task types. Visible categories include Intellectual Property, Corporate M&A, Data Privacy/Cybersecurity, Banking Finance, Capital Markets, Corporate Governance, International Trade Sanctions, Real Estate, and Energy/Natural Resources.

  • Good for observing: End-to-end completion of long-chain file retrieval, cross-material synthesis, legal analysis, and review-ready work products.

  • Not suitable for extrapolation to: Chinese legal contexts, open-web legal research, arbitrary custom Agents, single-turn Q&A, real client deliverables, or conclusions for legal practice.

Method and tools

Vals AI says it uses Harvey's generation and scoring protocol and runs the benchmark in an environment with internet access disabled. Each task requires the Agent to produce legal work against task-specific criteria. The shared Agent harness recorded on the page includes:

  • Tools (6): Read File, Edit File, Write File, Glob, Bash, and Grep.

  • Skills (3): docx, pptx, and xlsx.

  • Scoring: Two LLM judges separately calculate task pass rate; a task passes only when 100% of its criteria pass, and the final score is the average of the two judges' task pass rates.

  • Evaluators: GPT 5.5 (medium reasoning) and Claude Sonnet 4.6 (unmodified).

  • Infrastructure: Vals AI uses its model library and Valkyrie Agent benchmark framework; the page says these are infrastructure changes that do not alter the model-performance metric.

Gemini 3.8 Flash page parameters

The run parameters in the page's embedded results data are: Google provider, temperature 1, reasoning effort high, and max output tokens 65,536; top_p and the model-specific harness were not reported. The page does not disclose Gemini's API snapshot, system prompt, random seed, retry strategy, or complete call logs.

Results data

Scoring metricGemini 3.8 FlashPage rankCost per testInput/output pricePage runtime
Task Pass Rate (all-pass)10.00% ± 2.42%11 / 59$3.66$1.50 / $7.50 per million tokens29 min 21 sec
Criteria Pass Rate90.19% ± 0.36%Not separately provided on the page$3.66$1.50 / $7.50 per million tokens29 min 21 sec

Adjacent results in the same Overall table on the page are: Kimi K3 at 10.83% (rank 9), Qwen 3.8 Max at 10.42% (rank 10), Claude Opus 4.8 at 9.58% (rank 12), and Gemini 3.7 Flash at 8.75% (rank 13). These are relative positions from the same page snapshot and do not mean the differences have reached statistical significance.

Conclusions

  1. High criterion-level pass rate, low task-level completion rate: There is a clear gap between the 90.19% criteria pass rate and the 10.00% all-pass task rate; missing one necessary criterion means the entire task does not pass, so the former cannot substitute for end-to-end task success rate.

  2. The ranking applies only to this experimental setup: 11/59 is the ranking under the Vals AI held-out set, six tools, three skills, two evaluators, and the page's parameter combination. It is not Gemini 3.8 Flash's overall ranking across all legal benchmarks.

  3. Tool use is part of the score: File-format handling, calls such as Bash/Read File, and the skill descriptions together make up the Agent environment; the 10.00% should not be directly applied to bare-model requests outside this harness.

Limitations and reproduction boundaries

  • Private tasks cannot be fully reproduced: The held-out set's per-task inputs, materials, rubric, and Gemini outputs are not publicly available on the page; the 120 tasks can only be seen in the page's auxiliary description and cannot be used to reconstruct the task set.

  • Strict all-pass is not partial correctness: One failed criterion makes the entire task fail; 10.00% does not mean that only 10% of the legal content was correct.

  • Parameters remain incomplete: The page gives the provider, temperature, reasoning effort, and max output tokens, but not the API snapshot, prompt, specific top_p value, seed, retry/failure handling, or model-specific harness.

  • The no-internet condition limits extrapolation: This run disabled internet access, making the result closer to closed client-matter file work; it does not cover online legal research or real-time regulatory updates.

  • Dynamic snapshot: The page is marked UPDATED 9/5/2026; the model list, prices, runtimes, rankings, and task data may continue to change. The figures in this note correspond only to the snapshot collected on 2026-09-08.

  • Evaluator error: The final score is the average of two LLM judges, GPT 5.5 and Claude Sonnet 4.6. The page does not provide Gemini's per-judge scores or differences from human review.

Reproduction recommendations

  1. Fix the same Harvey LAB v1 held-out/public split, 120-task denominator, task materials, rubric, and stopping conditions; if only the public set is accessible, label it separately as a public-set experiment.

  2. Fix google/gemini-3.8-flash, Google provider, temperature 1, reasoning effort high, and max output tokens 65,536, and record the actual API snapshot, top_p, prompt, retries, and failure handling.

  3. Use the same six tools, docx/pptx/xlsx skills, and no-internet environment, retaining every call trace and final file.

  4. Have GPT 5.5 medium and unmodified Claude Sonnet 4.6 score the binary criteria separately. Report both judges' task pass rate, criteria pass rate, and average; do not treat Criteria Pass Rate as All-Pass.

Original evidence and data

  • Original Vals AI Harvey's Legal Agent Benchmark page: https://www.vals.ai/benchmarks/hlab. The page shows v1, UPDATED 9/5/2026, the held-out test set, 59 systems, 6 tools, 3 skills, two LLM judges, and the no-internet run condition.

  • In the Overall results table on the same page, google/gemini-3.8-flash is listed at 10.00% ± 2.42%, rank 11/59, $3.656207/test, 1760.877 seconds; the page's embedded Criteria Pass Rate data is 90.187% ± 0.358%, rounded here to the displayed precision.

Source excerpt or observation (short quote for compliance only)

Vals AI summarizes the benchmark as testing Agents that “complete legal work using documents, spreadsheets, presentations, and file-system tools.” This precisely bounds the conclusions in this note: it measures end-to-end legal Agent tasks with files and tools, not general legal Q&A.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Gemini 3.8 Flash

Use and compare models in Tabbit

Gemini 3.8 Flash

Related reviews

OfficialGoogle Blog (The Keyword)2026-09-02

Gemini 3.8 Flash: Google’s Official Benchmarks and Reproduction Boundaries

MediaArtificial Analysis (official model pages, methodology, and release article; the official X account was used to discover and cross-check the release post)2026-09-02

Gemini 3.8 Flash: Artificial Analysis Intelligence, Speed, Pricing, and Latency

MediaAI IQ (AIIQ, Liberated Software LLC)2026-09-02

Gemini 3.8 Flash: AI IQ Capability Benchmarks and Task Boundaries

MediaVals AI2026-09-05

Vals AI Finance Agent v2: Professional Finance Agent Benchmark for Gemini 3.8 Flash

Gemini 3.8 Flash

Related prompts

OfficialGoogle AI for Developers / Google DeepMind2026-09-02

Gemini 3.8 Flash: Google’s Official Model Parameters and API Configuration

OfficialGoogle AI for Developers2026-06-10

Gemini 3.8 Flash: Google's Official Structured Prompting and Agent Workflow

OfficialGoogle AI for Developers

Gemini 3.8 Flash: Google's Official Function-Calling Configuration and Tool Workflow

OfficialGoogle AI for Developers2026-09-02

Gemini 3.8 Flash: Google's Official Structured Output Configuration