In a personal batch extraction of approximately 1,000 mixed JPG/PDF invoices, Seed2.1 Pro required manual correction on 3–4 rows per 100 after the same VLM prompt and post-processing, at an estimated cost of about one-third of the original frontier solution. Outputs involving amounts still require human review.
Suitable tasks: Evaluating candidate models for large-scale invoice/receipt OCR and structured extraction, especially in pipelines that can include human spot checks and failure reruns.
Unsuitable tasks: Unattended payment processing, financial posting, or direct automation on PDFs with complex scan backgrounds.
Applicable model version: Seed 2.1 Pro; the specific preview/snapshot was not disclosed.
Applicable client, agent, or API: Via the ZenMux provider router; this was not a controlled test of Ark's native API.
Recommended reasoning tier and parameters: Not disclosed; the author only said that the same VLM prompt and post-processing were reused.
Input: Approximately 1,000 invoices from the past several years, in a mixture of JPG and PDF formats.
Target fields: Structured columns such as name, amount, date, and tax ID.
Baseline: The frontier solution previously used by the author, with the same VLM extraction prompt and the same post-processing.
Human verification: Randomly checked 100 rows.
Routing: ZenMux; model-call parameters, the complete file set, and preprocessing scripts were not disclosed.
The original extraction prompt was not disclosed, so this article should not be presented as a reproducible prompt.
Preprocessing, chunking, retries, output schema, and provider parameters were not disclosed.
The author noted that an approximately 256K context window would limit long PDFs or large concatenated batches, requiring chunking; the specific splitting strategy was not disclosed.
The extraction output was generally usable: the author said totals matched and dates landed in the correct columns.
A random check of 100 rows required manual correction on approximately 3–4 rows; the original frontier solution typically required corrections on 1–2 rows per 100.
Some scanned PDFs with backgrounds were silently skipped and had to be rerun after adding a preprocessing hint.
Including retries, the author estimated the cost per document at about one-third that of the original solution.
The author still retains human review for all outputs involving amounts.
This report supports treating Seed2.1 Pro as a cost-optimization candidate for batch visual extraction; it does not support unattended financial extraction. Error rates, silent skips, and context limits should all be included in production acceptance criteria.
A single user, a single file set, and a provider router; no complete sample, random seed, raw output, or statistical confidence interval was disclosed.
A manual check of 100 rows cannot estimate the overall error rate across all layouts and scan qualities.
The cost is an approximation from the author's environment, and token usage, retry counts, and routing fees were not disclosed.
Build a de-identified, stratified invoice set, bucketed by JPG/PDF format, scan background, language, and layout.
Fix the same prompt, schema, preprocessing, chunking, retries, and provider, then run batch jobs for Pro and the reference model.
Calculate field-level accuracy separately for amounts, dates, tax IDs, and line items, while recording silent skips, retries, and per-document cost.
Set up human review and rejection/rerun rules for samples involving amounts, then decide whether to expand the scope of automation.
The author's key boundary is that “outputs involving amounts still require review.”
The report considers both the provider router's per-document cost and field errors, rather than looking only at the price of a single response.
Doubao Seed 2.1 Pro