MODEL PROFILE · InclusionAI · Vercel AI Gateway access
Ling 3.1 Flash
A hybrid reasoning model with a dated Gateway promotion. Choose the model identifier deliberately: the standard and free routes behave differently when the promotion ends.
Source reviewed: · Language model
Provider documentation summary. No hands-on results or independent ranking are claimed.
Documented facts
The following fields come from the official provider source, reviewed on the date above.
- Identity and access
- InclusionAI model; Vercel announced AI Gateway availability September 30, 2026. No inference request was made in this review.
- Architecture (provider description)
- 560B total parameters, 25B active per token; hybrid reasoning language model.
- Gateway context
- 262K tokens in the announcement. This is the hosted route’s limit, not a claim about every deployment.
- Intended use
- Coding, multi-step analysis and tool-using agents, according to Vercel; these are intended uses, not independently measured strengths.
- Model identifiers
- inclusionai/ling-3.1-flash and inclusionai/ling-3.1-flash-free.
- Promotion and expiry behavior
- Free through October 13, 2026, per the September 30 announcement. The standard ID begins billing after the promotion; the -free ID stops serving. Exact cutoff time and post-promotion rates are not established here.
A hosted limit is more useful than a family-wide headline
Vercel’s free-route catalog lists Novita AI as the provider, a rounded 262K context and 33K maximum output. Keep those displayed units: this page does not convert them into more precise token counts. Broader-context claims from another deployment cannot enlarge this route’s advertised limit. Confirm the selected provider and output allowance before planning a long-document task. Vercel free-route model catalog.
The practical decision
For a bounded evaluation, the free-suffixed identifier offers an explicit stopping behavior at promotion expiry. A service that must remain available needs a separate continuity decision: establish the later rate, approve a spending limit and test a fallback before relying on it. Neither choice guarantees that a particular account has access today. Keep an expiry reminder alongside the endpoint configuration so a temporary price does not become a permanent cost assumption.
What to record in your own comparison
- Compare saved answers on the same coding or document-analysis tasks, with the same tool permissions and acceptance criteria. A parameter count alone does not predict whether an answer is useful.
- Retain the exact identifier, hosting provider, reasoning configuration, input length and output limit. Record failures as well as accepted outputs.
- Check citations, code execution results and tool-call arguments independently. A fluent explanation is not evidence that the requested work was completed.
- Separate evaluation cost from operational cost. Even if model inference is temporarily free, tools, storage and other services can have their own charges; verify the applicable terms.
Limitations and unknowns
No hands-on quality, speed, reliability or comparative ranking is established. Exact modality coverage, downloadable weights, license and post-promotion token rates remain unverified by this profile. Do not inherit them from another Ling version. The catalog is access documentation, not proof that a request succeeds for your account.
Compare your saved results
Use the DualView editor to inspect files you already have. The accepted-output calculator uses your own rates and counts. Neither link runs a model or purchases generation.