On Vals AI's Harvey's Legal Agent Benchmark v1 held-out test set, Gemini 3.8 Flash achieved a strict Task Pass Rate of 10.00% ± 2.42%, ranking 11th out of 59 systems; the page also shows a Criteria Pass Rate of 90.19% ± 0.36%. It passes most individual criteria, but only a small number of tasks pass every criterion, so 10% should not be interpreted as general legal capability or directly deliverable legal accuracy.
Task subject: Completing legal work with files and tools, rather than answering isolated legal questions; inputs may include documents, spreadsheets, presentations, and file-system materials.
Data split: The page says there is currently a public set and a held-out test set; the results on this page come from the held-out set. The page's fallback auxiliary description uses 120 tasks as the denominator. This note records that as the task scale for the current run, but the page body does not disclose the task list item by item.
Coverage: The page charts provide 24 task types. Visible categories include Intellectual Property, Corporate M&A, Data Privacy/Cybersecurity, Banking Finance, Capital Markets, Corporate Governance, International Trade Sanctions, Real Estate, and Energy/Natural Resources.
Good for observing: End-to-end completion of long-chain file retrieval, cross-material synthesis, legal analysis, and review-ready work products.
Not suitable for extrapolation to: Chinese legal contexts, open-web legal research, arbitrary custom Agents, single-turn Q&A, real client deliverables, or conclusions for legal practice.
Vals AI says it uses Harvey's generation and scoring protocol and runs the benchmark in an environment with internet access disabled. Each task requires the Agent to produce legal work against task-specific criteria. The shared Agent harness recorded on the page includes:
Tools (6): Read File, Edit File, Write File, Glob, Bash, and Grep.
Skills (3): docx, pptx, and xlsx.
Scoring: Two LLM judges separately calculate task pass rate; a task passes only when 100% of its criteria pass, and the final score is the average of the two judges' task pass rates.
Evaluators: GPT 5.5 (medium reasoning) and Claude Sonnet 4.6 (unmodified).
Infrastructure: Vals AI uses its model library and Valkyrie Agent benchmark framework; the page says these are infrastructure changes that do not alter the model-performance metric.
The run parameters in the page's embedded results data are: Google provider, temperature 1, reasoning effort high, and max output tokens 65,536; top_p and the model-specific harness were not reported. The page does not disclose Gemini's API snapshot, system prompt, random seed, retry strategy, or complete call logs.
| Scoring metric | Gemini 3.8 Flash | Page rank | Cost per test | Input/output price | Page runtime |
|---|---|---|---|---|---|
| Task Pass Rate (all-pass) | 10.00% ± 2.42% | 11 / 59 | $3.66 | $1.50 / $7.50 per million tokens | 29 min 21 sec |
| Criteria Pass Rate | 90.19% ± 0.36% | Not separately provided on the page | $3.66 | $1.50 / $7.50 per million tokens | 29 min 21 sec |
Adjacent results in the same Overall table on the page are: Kimi K3 at 10.83% (rank 9), Qwen 3.8 Max at 10.42% (rank 10), Claude Opus 4.8 at 9.58% (rank 12), and Gemini 3.7 Flash at 8.75% (rank 13). These are relative positions from the same page snapshot and do not mean the differences have reached statistical significance.
High criterion-level pass rate, low task-level completion rate: There is a clear gap between the 90.19% criteria pass rate and the 10.00% all-pass task rate; missing one necessary criterion means the entire task does not pass, so the former cannot substitute for end-to-end task success rate.
The ranking applies only to this experimental setup: 11/59 is the ranking under the Vals AI held-out set, six tools, three skills, two evaluators, and the page's parameter combination. It is not Gemini 3.8 Flash's overall ranking across all legal benchmarks.
Tool use is part of the score: File-format handling, calls such as Bash/Read File, and the skill descriptions together make up the Agent environment; the 10.00% should not be directly applied to bare-model requests outside this harness.
Private tasks cannot be fully reproduced: The held-out set's per-task inputs, materials, rubric, and Gemini outputs are not publicly available on the page; the 120 tasks can only be seen in the page's auxiliary description and cannot be used to reconstruct the task set.
Strict all-pass is not partial correctness: One failed criterion makes the entire task fail; 10.00% does not mean that only 10% of the legal content was correct.
Parameters remain incomplete: The page gives the provider, temperature, reasoning effort, and max output tokens, but not the API snapshot, prompt, specific top_p value, seed, retry/failure handling, or model-specific harness.
The no-internet condition limits extrapolation: This run disabled internet access, making the result closer to closed client-matter file work; it does not cover online legal research or real-time regulatory updates.
Dynamic snapshot: The page is marked UPDATED 9/5/2026; the model list, prices, runtimes, rankings, and task data may continue to change. The figures in this note correspond only to the snapshot collected on 2026-09-08.
Evaluator error: The final score is the average of two LLM judges, GPT 5.5 and Claude Sonnet 4.6. The page does not provide Gemini's per-judge scores or differences from human review.
Fix the same Harvey LAB v1 held-out/public split, 120-task denominator, task materials, rubric, and stopping conditions; if only the public set is accessible, label it separately as a public-set experiment.
Fix google/gemini-3.8-flash, Google provider, temperature 1, reasoning effort high, and max output tokens 65,536, and record the actual API snapshot, top_p, prompt, retries, and failure handling.
Use the same six tools, docx/pptx/xlsx skills, and no-internet environment, retaining every call trace and final file.
Have GPT 5.5 medium and unmodified Claude Sonnet 4.6 score the binary criteria separately. Report both judges' task pass rate, criteria pass rate, and average; do not treat Criteria Pass Rate as All-Pass.
Original Vals AI Harvey's Legal Agent Benchmark page: https://www.vals.ai/benchmarks/hlab. The page shows v1, UPDATED 9/5/2026, the held-out test set, 59 systems, 6 tools, 3 skills, two LLM judges, and the no-internet run condition.
In the Overall results table on the same page, google/gemini-3.8-flash is listed at 10.00% ± 2.42%, rank 11/59, $3.656207/test, 1760.877 seconds; the page's embedded Criteria Pass Rate data is 90.187% ± 0.358%, rounded here to the displayed precision.
Vals AI summarizes the benchmark as testing Agents that “complete legal work using documents, spreadsheets, presentations, and file-system tools.” This precisely bounds the conclusions in this note: it measures end-to-end legal Agent tasks with files and tools, not general legal Q&A.
Gemini 3.8 Flash