Qwen positions Qwen3.5-Plus as the hosted API version of Qwen3.5-397B-A17B, with strengths in native vision, multimodal Agents, tool calling, and a million-token context, but official benchmarks cannot replace a rerun with the same harness.
Suitable tasks: Image/video understanding, long documents and codebases, coding Agents, and tool-augmented production workflows.
Unsuitable tasks: Tasks that require fully reproducing the hosted API's internal tool strategy locally, or treating officially self-reported scores as fair cross-model comparisons.
Applicable model versions: Qwen3.5-Plus; the official blog also introduces the open-weight Qwen3.5-397B-A17B.
Applicable clients, Agents, or APIs: Qwen Chat, Alibaba Cloud Bailian API, and self-orchestrated tool-calling Agents.
Recommended reasoning level and parameters: The official blog does not disclose Plus's complete request parameters, system prompt, or harness; pin the API snapshot and tool set before testing.
Evaluation subjects: The Qwen3.5 series in Qwen's official release notes, including hosted Qwen3.5-Plus and open-weight Qwen3.5-397B-A17B.
Inputs/tasks: The official comparison covers natural language, reasoning, programming, Agent, and multimodal understanding categories; the models visible on the page include GPT-5.2, Claude 4.5 Opus, Gemini-3 Pro, Qwen3-Max-Thinking, and K2.5-1T-A32B.
Configuration: The official article does not fully disclose each prompt, sampling parameters, tool schema, number of repetitions, or evaluation harness.
Qwen3.5-397B-A17B has approximately 397B total parameters, with about 17B activated per forward pass; its architecture combines Gated Delta Networks linear attention with sparse MoE.
Qwen says its language and dialect coverage expanded from 119 to 201; the blog describes it as a native vision-language model and identifies Qwen3.5-Plus as the API version.
The official blog lists Qwen3.5-Plus with a 1M-token context window and mentions official tools and adaptive calling; the Alibaba Cloud model page further lists text/image/video inputs, function calling, and structured-output support.
The benchmark table visible on the page includes model columns and task categories, but the current body extraction retains only part of the natural-language rows in full; scores that could not be fully verified are not included here.
If a task needs visual input, long context, and Agent tools at the same time, Qwen3.5-Plus is worth prioritizing for a cost/latency baseline; for pure text or coding, rerun it against other models with the same prompt, tools, and output budget, rather than selecting it solely because of the official claim that it is “on par with frontier models.”
An official release is vendor-reported and lacks a complete, consistently public prompt, code, random seed, and per-question artifacts.
The open-weight baseline and hosted Plus API are different service forms, so their tools, context handling, and reasoning implementations may differ.
The blog contains dynamic example content, and not every table row can be extracted from the current page body; missing scores are explicitly marked as unverified.
Use the Alibaba Cloud Bailian qwen3.5-plus-2026-02-15 snapshot, recording region, tool allowlist, context length, temperature, and output limit.
Select a set of public tasks covering pure text, image understanding, function calling, and long-context retrieval; publish every input and expected output.
Repeat each item at least 3 times, recording success rate, tool-argument validity rate, latency, input/output tokens, and error type.
Rerun against the target comparison models under the same API protocol, equivalent output budget, and identical tool set; store the official release table separately from the measured table.
The official title defines Qwen3.5 as “Towards Native Multimodal Agents”; this article uses only the architecture, model versions, and capability positioning that can be checked directly on the page, without filling in benchmark numbers whose extraction is incomplete.
Qwen3.5 Plus