Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaGemini 3.8 Flash

Vals AI Finance Agent v2: Professional Finance Agent Benchmark for Gemini 3.8 Flash

Original source

Vals AI

AuthorVals AI

Source date2026-09-05

Tabbit curation2026-09-08

Read original

One-sentence takeaway

On the overall Vals AI Finance Agent v2 tasks, Gemini 3.8 Flash achieved 61.44% ± 0.13 Partial Credit, ranking No. 1 among the 58 systems listed on the page; its strict All-Pass score was 49.69% ± 0.42, ranking No. 2. It is suitable for evaluating an agent that retrieves, verifies, and completes multi-step financial analysis from public-company filings, but should not be interpreted as a measure of general intelligence or production investment-research accuracy.

Use cases

  • Tasks suitable for observation: Quantitative and qualitative analysis of public-company filings, market analysis, earnings analysis, disclosure analysis, comparables, precedent transactions, adjustments, and financial modeling.

  • Tasks not suitable for direct inference: General question answering, non-financial domains, real-time trading decisions, agents without equivalent tools and document context, or the stability and compliance of production systems.

  • Applicable model version: Gemini 3.8 Flash on the Vals AI model page; do not extend this to Gemini 3.7 Flash, Gemini 3.8 Flash Cyber, or other providers/reasoning tiers.

  • Applicable client, agent, or API: The shared default harness for Vals AI Finance Agent v2; this is not an end-to-end score for the Gemini App or an arbitrary self-built agent.

  • Recommended parameters: The page shows Google provider, temperature 1, top-p/top-k DEFAULT, max output 65,536, and reasoning effort HIGH as its defaults; it also warns that different benchmarks may use different providers/parameters, so these page defaults must not be treated as a complete Finance Agent v2 run log when reproducing the test.

Test environment and method

  • Target task: Simulate an entry-level financial analyst answering difficult questions based on public-company filings; questions typically require finding evidence across multiple filings or data sources, applying financial conventions, and keeping intermediate figures precise.

  • Shared default harness: Each agent can use six tools: edgar_search (SEC EDGAR API), web_search, parse_html_page, retrieve_information, calculator, and price_history.

  • Time limit: 2 hours per question; timeouts receive 0 points.

  • Runs and statistics: Each model was run 3 times; the page reports the mean and standard error (SEM) across the three runs.

  • Scoring: Each question consists of weighted checks and includes dealbreakers (critical facts/numbers). Partial Credit is the per-question average weighted by severity and gated by dealbreakers; if a critical item fails, the question receives 0. All-Pass is 100% only when every check passes; otherwise it is 0.

  • Evaluators: The Vals AI page lists three LLM evaluators: GPT-5.4, Gemini-3.1-Pro, and Claude Sonnet 4.6.

  • Question design: Questions require determinate answers, multi-source synthesis, domain specificity, and forensic precision (critical facts may be buried in footnotes, MD&A qualifiers, or accounting-policy disclosures).

Sample and coverage

Finance Agent v2 contains 927 expert-reviewed questions, divided as follows:

Data sectionSizeAvailability
Public27Open sample and agent harness accessible
Private Validation450Requires authorization/permission
Test450Kept private; published page results are based only on this section

The questions are divided into 9 analysis categories: General Qualitative Analysis, General Quantitative Analysis, Market Analysis, Comparables, Precedents, Adjustments, Earnings Analysis, Disclosure Analysis, and Financial Modeling. The page shows the highest accuracy by category ranging from 83.4% for General Qualitative to 34.5% for Financial Modeling, illustrating the benchmark's difficulty range from retrieval and summarization to multi-step modeling; these are the results of the leading model in each category, not Gemini 3.8 Flash's category-by-category scores.

Gemini 3.8 Flash results

Finance Agent v2 / Overall

Scoring metricAccuracyPage rankCost per testInput/output pricingPage runtime
Partial Credit61.44% ± 0.131 / 58$2.00$1.50 / $7.50 per million tokens3 min 21 sec
All-Pass49.69% ± 0.422 / 58$2.00$1.50 / $7.50 per million tokens3 min 21 sec

The Vals AI model page also lists Finance Agent (v2) at 61.44% ± 0.13, 1/58; the Vals Index in that page's overview is a separate aggregate metric and cannot replace the dedicated Finance Agent v2 score.

Conclusions

  1. Strong performance as a professional finance agent: In Vals AI's Finance Agent v2 Overall table, Gemini 3.8 Flash ranks first with 61.44% Partial Credit; the adjacent second-place model, Muse Spark 1.2, scored 60.60% ± 0.28, a very small gap that does not justify claiming broad cross-task leadership.

  2. Strict accuracy remains limited: All-Pass falls to 49.69% and ranks second; the gap between Partial Credit and All-Pass indicates that the model may capture key points while still missing peripheral checks or precise figures.

  3. Cost should be assessed across the full test: The page records $2.00 per test and separately lists input/output pricing of $1.50/$7.50 per million tokens; this is not a fixed bill for any API request and does not mean costs remain unchanged after switching providers.

Limitations and boundaries

  • Cannot fully reproduce every question: The public section contains only 27 questions; the 450 Test questions remain private, so the published 61.44%/49.69% cannot be recalculated from the public set alone.

  • The harness has a major effect: The six tools, document retrieval, retrieval method, calculator calls, and stopping conditions all contribute to the score; this score should not be applied to Gemini 3.8 Flash outside that harness.

  • The scoring is not a single accuracy rate: Partial Credit is affected by dealbreakers and severity weighting, while All-Pass is a binary measure requiring every check to pass; they should not be conflated into one “accuracy” figure.

  • Dynamic snapshot: The page is marked UPDATED 9/5/2026; the model list, prices, rankings, and model availability may change, and the figures in this article correspond only to the page snapshot collected on 2026-09-08.

  • Limited question scope: The questions focus on public-company filings and financial-analyst work; they do not cover visual understanding, general coding ability, Chinese-localized investment research, compliance judgments, or actual investment returns.

  • Page parameters are not a complete run log: The default provider/temperature/reasoning settings on the model page cannot prove that every benchmark run used exactly the same configuration; the Vals AI benchmark page and the target run logs should be treated as authoritative.

Reproduction conditions and steps

  1. Fix the Vals AI Finance Agent v2 page version, the exact model name Gemini 3.8 Flash, provider, temperature, top-p/top-k, max output tokens, reasoning effort, and pricing snapshot.

  2. Use the six tools in the shared default harness and the two-hour per-question limit; retain the trace of every search, page parse, information retrieval, calculator, and price-history call.

  3. First use the public 27 questions to verify the integration, then request access to the same Test split; do not claim to reproduce the page's 450-question results using public or self-authored questions.

  4. Run each question 3 times, and calculate Partial Credit and All-Pass using the same check weights, dealbreaker gating, severity rules, and three-evaluator setup, reporting the mean and SEM.

  5. Also record cost per test, input/output tokens, total runtime, timeouts, tool errors, and fallback behavior; if the provider or harness differs, name it as a new experiment and do not merge it with the ranking on this page.

Original evidence and data

  • Original Vals AI Finance Agent v2 page: https://www.vals.ai/benchmarks/fabv2; the page shows UPDATED 9/5/2026, 927 expert-reviewed questions, the shared harness, 6 tools, a 2-hour limit, 3 runs, and the definitions of Partial Credit/All-Pass.

  • Overall table on the same page: Gemini 3.8 Flash achieved Partial Credit of 61.44% ± 0.13, 1/58, $2.00/test, 3m21s; after switching to All-Pass, it was 49.69% ± 0.42, 2/58.

  • Original Vals AI model page: https://www.vals.ai/models/google_gemini-3.8-flash; the page lists the model release date, Google provider, token pricing, Finance Agent (v2) at 61.44% ± 0.13 and 1/58, and the page's default hyperparameters.

Source excerpt or observation (short quotation for compliance only)

Vals AI summarizes Finance Agent v2 as “Evaluating agents on core financial analyst tasks.” This accurately limits the test subject: it is a benchmark of professional financial-analysis workflows, not an overall assessment of Gemini 3.8 Flash's general capabilities.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Gemini 3.8 Flash

Use and compare models in Tabbit

Gemini 3.8 Flash

Related reviews

OfficialGoogle Blog (The Keyword)2026-09-02

Gemini 3.8 Flash: Google’s Official Benchmarks and Reproduction Boundaries

MediaArtificial Analysis (official model pages, methodology, and release article; the official X account was used to discover and cross-check the release post)2026-09-02

Gemini 3.8 Flash: Artificial Analysis Intelligence, Speed, Pricing, and Latency

MediaAI IQ (AIIQ, Liberated Software LLC)2026-09-02

Gemini 3.8 Flash: AI IQ Capability Benchmarks and Task Boundaries

MediaSimpleBench official leaderboard and project

SimpleBench: Gemini 3.8 Flash on Everyday Reasoning and Language-Trap Questions

Gemini 3.8 Flash

Related prompts

OfficialGoogle AI for Developers / Google DeepMind2026-09-02

Gemini 3.8 Flash: Google’s Official Model Parameters and API Configuration

OfficialGoogle AI for Developers2026-06-10

Gemini 3.8 Flash: Google's Official Structured Prompting and Agent Workflow

OfficialGoogle AI for Developers

Gemini 3.8 Flash: Google's Official Function-Calling Configuration and Tool Workflow

OfficialGoogle AI for Developers2026-09-02

Gemini 3.8 Flash: Google's Official Structured Output Configuration