LongCat Flash Chat · prompting-guide
Use the official Round format and LongCat tool-call tags for one weather or order lookup, with explicit parameters, call order, and final-answer checks.
Task: {{USER_REQUEST}}
Tool schema: {{TOOL_SCHEMA}}
Tool return: {{TOOL_RESULT}}
Acceptance: {{FINAL_ACCEPTANCE}}Replace before running: {{USER_REQUEST}}, {{TOOL_SCHEMA}}, {{TOOL_RESULT}}, {{FINAL_ACCEPTANCE}}
Transformers or a LongCat-compatible service; a reachable tool stub.
Prepare User request and pin the remaining inputs in the run log.
Send one real request using the public format described by “LongCat-Flash-Chat Official Chat Template and Tool-Calling Prompt”; save request, events, and return.
Judge it against the acceptance rules; a model claim is not proof that a tool ran.
Produce a reproducible trace aligning input, arguments, tool return, and final answer.
Check response status, format, critical fields, and task result; retain raw errors.
Reduce to one tool, one turn, and the smallest schema before restoring fields; for service errors check endpoint, permission, model name, and timeout.
Editorial adaptation of public guidance from GitHub (the official Meituan LongCat repository), limited to this task and not a guarantee for another backend.
Strictly assembling LongCat-Flash-Chat's round prefixes, conversation history, and <longcat_tool_call> XML wrapper according to the official template makes it possible to reproduce the multi-turn dialogue and function-calling format of its open-source weights.
Good for: Constructing message sequences for local/self-hosted inference, multi-turn dialogue, tool calling, and Agent harnesses.
Not for: Treating the XML template as an OpenAI/Anthropic API request schema, or rewriting special tokens without a tokenizer configuration.
Applicable model versions: The open-source Meituan LongCat-Flash-Chat weights; the API version has been upgraded, and the current legacy model service has been retired, so the integration surface must be confirmed first.
Applicable clients, Agents, or APIs: The official tokenizer/chat template, SGLang, and vLLM adapters; see the repository's Deployment Guide for specific deployment parameters.
Recommended inference tier and parameters: The model is a non-thinking foundation model; the repository does not provide a temperature/top-p combination that can be reproduced uniformly, so the deployer must fix these values.
[Round 0] USER:{query} ASSISTANT:SYSTEM:{system_prompt} [Round 0] USER:{query} ASSISTANT:SYSTEM:{system_prompt} [Round 0] USER:{query} ASSISTANT:{response}</longcat_s>... [Round N-1] USER:{query} ASSISTANT:{response}</longcat_s> [Round N] USER:{query} ASSISTANT:{tool_description}
## Messages
SYSTEM:{system_prompt} [Round 0] USER:{query} ASSISTANT:
## Tools
You have access to the following tools:
### Tool namespace: function
#### Tool name: {func.name}
Description: {func.description}
InputSchema:
{json.dumps(func.parameters, indent=2)}
For each function call, return a JSON object inside its own XML tag:
<longcat_tool_call>
{"name": <function-name>, "arguments": <args-dict>}
</longcat_tool_call>Read the actual chat template from the repository's tokenizer_config.json; do not rely only on the simplified example in the README.
First run a tool-free first turn and a multi-turn echo test to confirm that the special end marker and round concatenation match.
Serialize the tool schema into ## Tools, requiring the model to output only JSON wrapped in <longcat_tool_call>.
Have the harness parse each XML block, execute the function, return the result as the next-round USER/tool context, and then independently verify the final answer.
Record the model version, tokenizer, inference engine, context length, and tool-calling errors; do not mix the server API's JSON tool schema with the local template.
The official repository provides four templates: first turn, system prompt, multi-turn, and ToolCall.
The tool-calling convention uses <longcat_tool_call> XML tags, with the function name and parameter JSON inside; multiple function calls use multiple consecutive tags.
The official repository says that SGLang and vLLM adapters are available, with detailed deployment instructions in docs/deployment_guide.md.
The template defines the input/output convention for the open-source weights; compatibility with every third-party API gateway is not guaranteed.
The legacy API service has been retired according to the official Change Log; when calling through a new platform, use the current model list and API documentation as the authority.
The ToolCall template defines only the format. It does not guarantee correct tool selection, safe arguments, or successful completion of the final task; the execution layer must perform schema validation, permission isolation, and timeout handling.
The official template writes rounds as [Round N] USER... ASSISTANT and retains the <longcat_s> end marker in multi-turn exchanges.
The official tool format requires function calls to appear inside <longcat_tool_call> tags rather than as free-text descriptions.
GitHub (the official Meituan LongCat repository) · Source date: Not disclosed · Edited: 2026-09-20
Read the original sourceLongCat Flash Chat
Run this guide in the environment listed above. Downloading does not transfer the template or establish model availability for your account.