Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Prompts and workflows

LongCat Flash Thinking · configuration

LongCat-Flash-Thinking-2601: Official SGLang/vLLM Deployment and MTP Configuration

Follow the official deployment notes to start MTP in SGLang or vLLM, hold concurrency and context constant, and measure first-token latency, generation speed, and tool-call parsing.

Source not verifiedGPU node, pinned SGLang/vLLM version, model weights, and MTP settings; a benchmark script is required.

Prerequisites and inputs

  • Backend and version
  • GPU count and precision
  • Concurrency and context settings
  • Fixed test prompts
  • Latency and parsing logs

Prerequisites

GPU node, pinned SGLang/vLLM version, model weights, and MTP settings; a benchmark script is required.

Task-specific steps

  1. Prepare Backend and version and pin the remaining inputs in the run log.

  2. Send one real request using the public format described by “LongCat-Flash-Thinking-2601: Official SGLang/vLLM Deployment and MTP Configuration”; save request, events, and return.

  3. Judge it against the acceptance rules; a model claim is not proof that a tool ran.

Task result

Produce a reproducible trace aligning input, arguments, tool return, and final answer.

Output and acceptance

Check response status, format, critical fields, and task result; retain raw errors.

Failure correction

Reduce to one tool, one turn, and the smallest schema before restoring fields; for service errors check endpoint, permission, model name, and timeout.

Source and boundary

Editorial adaptation of public guidance from GitHub, limited to this task and not a guarantee for another backend.

Read the source research notes

One-sentence takeaway

The official deployment guide provides single-node, multi-node, and Multi-Token Prediction (MTP) launch parameters for LongCat-Flash-Thinking-2601 on SGLang and vLLM, offering a starting point for reproducible deployment.

Use cases

  • Suitable tasks: Self-hosted 560B MoE weights, tool-Agent services, and inference deployments requiring tensor or expert parallelism.

  • Unsuitable tasks: Scenarios without multi-GPU nodes, without enough VRAM for FP8/BF16, or expecting to run directly on a single consumer GPU.

  • Applicable model versions: LongCat-Flash-Thinking-2601 and the official FP8 weights name.

  • Applicable client, Agent, or API: SGLang and vLLM; replace the master-node variables with values for the actual cluster.

  • Recommended inference tier and parameters: First validate the service with the single-node FP8 configuration, then enable multiple nodes using the official parameters; enable MTP only after confirming that the engine version supports it.

Ready-to-use content

The official guide says FP8 requires at least one 8×H20-141G node, while BF16 requires at least two 8×H800-80G nodes. The commands below come from the page and can be used to establish a minimal deployment configuration.

SGLang: Single-node FP8

python3 -m sglang.launch_server \
    --model meituan-longcat/LongCat-Flash-Thinking-2601-FP8 \
    --trust-remote-code \
    --attention-backend flashinfer \
    --enable-ep-moe \
    --tp 8

vLLM: Single-node FP8

vllm serve meituan-longcat/LongCat-Flash-Thinking-2601-FP8 \
    --trust-remote-code \
    --enable-expert-parallel \
    --tensor-parallel-size 8

Additional MTP parameters

SGLang:

--speculative-draft-model-path meituan-longcat/LongCat-Flash-Thinking-2601 \
--speculative-algorithm NEXTN \
--speculative-num-draft-tokens 2 \
--speculative-num-steps 1 \
--speculative-eagle-topk 1

vLLM:

--speculative_config '{"model": "meituan-longcat/LongCat-Flash-Thinking-2601", "num_speculative_tokens": 1, "method":"longcat_flash_mtp"}'

Test/workflow steps

  1. First confirm that the downloaded weights are the official FP8 or BF16 versions, and check that the SGLang/vLLM version includes the LongCat adaptation.

  2. Start the service on a single node and first run short-input, tool-free requests to check connectivity.

  3. Then add tool calls and long-context workloads, recording VRAM, throughput, time to first token, and expert-parallel stability.

  4. Add MTP only after the baseline is stable; record quality, latency, and error rate before and after enabling MTP separately.

  5. For multiple nodes, replace variables such as $MASTER_IP and $NODE_RANK, and retain the node topology, parallelism, and weight precision.

Original evidence and data

  • The official guide gives the hardware boundary as at least one 8×H20-141G node for FP8 and at least two 8×H800-80G nodes for BF16.

  • The SGLang single-node parameters are --enable-ep-moe --tp 8; the vLLM single-node parameters are --enable-expert-parallel --tensor-parallel-size 8.

  • The guide links to SGLang PR #9824 and vLLM PR #23991, and explicitly describes the current state as basic adaptations.

  • The MTP parameters use SGLang's NEXTN configuration and vLLM's longcat_flash_mtp method, respectively.

Applicability boundaries

  • The official commands are deployment templates, not independently measured results from this review; they provide no specific throughput, latency, available-VRAM, or service-concurrency data.

  • The 560B total parameter count and MoE expert parallelism make hardware, drivers, communication topology, and engine version critical variables; “at least one node” must not be interpreted as any eight-GPU machine running reliably.

  • $MASTER_IP and $NODE_RANK in multi-node commands must be filled in by the deployer and cannot be copied unchanged into production.

  • MTP may change generation behavior and speed. It must be regressed separately on the target task and should not be enabled by default.

Source excerpt or observation (short quote for compliance only)

The guide describes this adaptation as “basic adaptations,” so the commands should be treated as version-sensitive starting configurations rather than performance guarantees.

Source and dates

GitHub · Source date: Not disclosed · Edited: 2026-09-20

Read the original source
Variable checklist

No required variables

Related prompts

LongCat-Flash-Thinking-2601: Official Chat Template, Tool Calling, and Reasoning-History Configuration

Related reviews

LongCat-Flash-Thinking: API Alias Upgrade, Automatic Routing, and Service-Retirement BoundariesLongCat-Flash-Thinking-2601: Initial Reading and Deployment Observations from the LocalLLaMA CommunityLongCat-Flash-Thinking-2601: Heavy Thinking, Environmental Noise, and Agent Benchmarks

LongCat Flash Thinking

Use LongCat Flash Thinking in Tabbit

Run this guide in the environment listed above. Downloading does not transfer the template or establish model availability for your account.