Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Prompts and workflows

Gemini 3.5 Flash · configuration

Gemini 3.5 Flash: Thinking Levels and Gemini API Configuration

The official Thinking guide documents Gemini 3.5 Flash effort levels, preservation of parts and thought signatures across turns, and API parameter boundaries.

Source reviewed; not testedGemini 3.5 Flash client or API; confirm the live model ID, tools, permissions, and version before execution.

Prerequisites and inputs

  • task goal
  • source or reference material
  • runtime constraints
  • acceptance criteria

Complete templates

Editorial adaptation: task template

Tabbit editorial adaptation; not the original source prompt
For {{MULTI_TURN_TASK}}, pin {{THINKING_LEVEL}}, {{TEMPERATURE_POLICY}}, and {{MAX_OUTPUT_TOKENS}}. Preserve {{STATE_PARTS}} and {{TOOL_RESULTS}} verbatim and use {{INCOMPLETE_CHECK}} to detect truncation. If {{REPLAY_FAILURE}} appears, inspect the SDK and provider parser before drawing any quality conclusion.

Before running, fill every variable and return each value in the acceptance record.

Replace before running: {{MULTI_TURN_TASK}}, {{THINKING_LEVEL}}, {{TEMPERATURE_POLICY}}, {{MAX_OUTPUT_TOKENS}}, {{STATE_PARTS}}, {{TOOL_RESULTS}}, {{INCOMPLETE_CHECK}}, {{REPLAY_FAILURE}}

For {{MULTI_TURN_TASK}}, pin {{THINKING_LEVEL}}, {{TEMPERATURE_POLICY}}, and {{MAX_OUTPUT_TOKENS}}. Preserve {{STATE_PARTS}} and {{TOOL_RESULTS}} verbatim and use {{INCOMPLETE_CHECK}} to detect truncation. If {{REPLAY_FAILURE}} appears, inspect the SDK and provider parser before drawing any quality conclusion.

Read the source research notes

One-sentence takeaway

Gemini 3.5 Flash uses medium thinking by default. The API supports minimal, low, medium, and high to adjust reasoning depth; for multi-turn tool calls, preserve all parts and thought signatures returned by the model exactly as received.

Use cases

  • Suitable tasks: Coding, data analysis, document processing, function calling, and multi-step Agent workflows with the Gemini API.

  • Unsuitable tasks: Applying Gemini 2.5's thinkingBudget parameter directly to Gemini 3.5 Flash, or deleting signature parts from multi-turn function calls.

  • Applicable model version: gemini-3.5-flash.

  • Applicable client, Agent, or API: generate_content/generateContent in the Google Gen AI SDK, as well as REST generateContent.

  • Recommended reasoning levels and parameters: minimal or low for simple classification/fact tasks; medium for ordinary tasks; high for complex coding, mathematics, and planning. Enable thought summaries as needed, and record thought tokens in production at the same time.

Ready-to-use content

The following code is a runnable configuration template based on official API fields and official examples, with the model name explicitly replaced by gemini-3.5-flash; the model-name replacement is an adaptation prepared for this model and is not claimed to be verbatim code from the page.

Python: Set the thinking level

from google import genai
from google.genai import types

client = genai.Client()

response = client.models.generate_content(
    model="gemini-3.5-flash",
    contents="Analyze this sales data, first give the conclusion, then list the evidence and uncertainties.",
    config=types.GenerateContentConfig(
        thinking_config=types.ThinkingConfig(thinking_level="high")
    ),
)

print(response.text)

JavaScript: Medium thinking

import { GoogleGenAI, ThinkingLevel } from "@google/genai";

const ai = new GoogleGenAI({});
const response = await ai.models.generateContent({
  model: "gemini-3.5-flash",
  contents: "Turn these meeting notes into JSON with the topic, decisions, owner, and deadline.",
  config: {
    thinkingConfig: { thinkingLevel: ThinkingLevel.MEDIUM },
  },
});

console.log(response.text);

REST: Low-latency level

curl "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash:generateContent" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H 'Content-Type: application/json' \
  -X POST \
  -d '{
    "contents": [{"parts": [{"text": "Classify the following text as bug, feature, or question; return only the category."}]}],
    "generationConfig": {
      "thinkingConfig": {"thinkingLevel": "low"}
    }
  }'

Read thought summaries as needed

from google import genai
from google.genai import types

client = genai.Client()
response = client.models.generate_content(
    model="gemini-3.5-flash",
    contents="Compare the two approaches and give a recommendation and key risks.",
    config=types.GenerateContentConfig(
        thinking_config=types.ThinkingConfig(
            thinking_level="high",
            include_thoughts=True,
        )
    ),
)

for part in response.candidates[0].content.parts:
    if not part.text:
        continue
    print("Thought summary:" if part.thought else "Answer:")
    print(part.text)

Testing/workflow steps

  1. Establish a baseline with the default medium; rerun simple tasks with minimal/low and complex tasks with high.

  2. For each run, record usage_metadata.thoughts_token_count, output tokens, time to first token, total latency, and cost.

  3. To debug quality, temporarily enable include_thoughts=True to inspect summaries; in production, retain only necessary summaries/logs and do not treat internal reasoning as proof of facts.

  4. When using function calling or multi-turn conversations, pass back all response parts returned by the model exactly as received; do not concatenate, delete, or modify signed parts.

  5. Do not use Gemini 2.5's thinkingBudget to control 3.5 Flash; 3.5 Flash uses thinkingLevel, which supports minimal, low, medium, and high.

Original evidence and data

  • The official model page lists gemini-3.5-flash with an input limit of 1,048,576 tokens and an output limit of 65,536 tokens; it supports text, image, video, audio, and PDF input, with text output.

  • The official thinking table marks Gemini 3.5 Flash's default as medium and supports minimal, low, medium, and high; high can dynamically increase reasoning depth.

  • The official documentation states that after thinking is enabled, charges include output tokens and thought tokens; thought tokens can be read from thoughtsTokenCount/the corresponding SDK usage field.

  • Gemini 3 models may return thought signatures for various parts; the official recommendation is to pass all parts through unchanged, and signatures must not be dropped in multi-turn function calls in particular.

Scope and limitations

  • Some code examples on the documentation page use gemini-3.6-flash as the current example; the template above replaces it with this model ID while keeping the same field structure. Actual SDK/API availability must be verified in the target project.

  • minimal does not guarantee that thinking is completely disabled; the official Gemini 3.5 Flash page does not provide a thinkingBudget=0 configuration for strictly no-thinking semantics.

  • Thought summaries are summaries rather than complete internal reasoning, and cannot replace external evidence, testing, or human review.

  • The 1M input window and 65,536 output limit are API model-page specifications and do not mean that every client, plan, or agent framework exposes the same limits.

Source excerpt or observation (brief excerpt for compliance only)

The official table marks Gemini 3.5 Flash's default level as medium and describes high as deeper dynamic reasoning. The documentation also emphasizes that multi-turn requests must pass back signed parts unchanged—a configuration detail that custom Agent frameworks are especially likely to break.

Source and dates

Google AI for Developers · Source date: 2026-07-30 · Edited: 2026-09-20

Read the original source
Variable checklist

Still to replace: 8

{{MULTI_TURN_TASK}}{{THINKING_LEVEL}}{{TEMPERATURE_POLICY}}{{MAX_OUTPUT_TOKENS}}{{STATE_PARTS}}{{TOOL_RESULTS}}{{INCOMPLETE_CHECK}}{{REPLAY_FAILURE}}

Related prompts

Gemini 3.5 Flash: Structured Prompting, Grounding, and Agent System Instructions

Related reviews

Gemini 3.5 Flash: Appwrite Arena Comparison of Skill Context and Agent TasksGemini 3.5 Flash: Google's Official Follow-up Release Comparison of Efficiency and CapabilitiesGemini 3.5 Flash: A Community Field Report on Ten Saved Tasks and Five Repeated Runs

Read the full analysis

Overview · English

Gemini 3.5 Flash: What It Is, What It Costs, and Whether It Still Fits

A sourced Gemini 3.5 Flash overview covering its 1M context, multimodal tools, $1.50/$9 API pricing, legacy status, evidence limits and migration choices.

Gemini 3.5 Flash

Use Gemini 3.5 Flash in Tabbit

Run this guide in the environment listed above. Downloading does not transfer the template or establish model availability for your account.