MODEL PROFILE · Perplexity
Perplexity pplx-decider-v1-27b
Perplexity’s Qwen-derived decision model returns probabilities for defined choices. Review open weights, hosted API limits, pricing and benchmark caveats.
Source reviewed: · Multimodal language-model fine-tune · decision scoring
Provider documentation summary. No hands-on results or independent ranking are claimed.
Documented facts
The following fields come from the official provider source, reviewed on the date above.
- Identity
- pplx-decider-v1-27b is a decision-model fine-tune of Qwen3.8-27B.
- Weights and license
- The official Hugging Face repository publishes weights and declares Apache-2.0 licensing.
- Local use
- The model card requires Python 3.12+ and a CUDA GPU with room for about 49 GiB of weights plus working memory.
- Inputs and output
- The local example accepts text and images and illustrates choice and noul (yes/no) decisions with probabilities.
- Availability date
- Public weights and documentation verified on 2026-10-02 UTC. This review date is not a claimed launch date.
Hosted API contract
Perplexity documents POST /v1/decisions with model pplx-decider-v1-27b. It accepts text, JSON or images as state and 1–128 named questions. Types are noul (yes/no), choice (options) and score (ordered rubric). The input limit is under 262,144 tokens, including state, images and questions. Images require base64 PNG/JPEG/WebP data URLs; remote HTTP image URLs are rejected. It returns typed probabilities rather than generated replies, code or reasoning. For choice and score, confidence is distinct from the highest option probability. Official Decisions API guide.
Hosted pricing and local costs
The guide lists $0.04 per million input tokens and free output tokens. One million billed input tokens therefore costs $0.04; 25 million costs $1.00, assuming the listed rate and no other fees. This API rate is not a self-hosting cost estimate. Image tokens count as input. Official pricing and input accounting.
How to read the reported benchmark result
The model card reports 85.71% overall on an eleven-benchmark panel, measured through Perplexity’s API. Individual rows do not show a universal win: its JudgeBench result is 78.29% versus Jev’s 78.57%, and its JevBench public hard result is 70.30% versus 73.27%. These are publisher-reported results, not DualView measurements, and the API measurement route should remain attached to any comparison with local weights. Model card benchmark table and measurement route.
The practical decision
Choose the evaluation around the decision you need to make. A fixed set of labels can be assessed against agreed answers; an explanatory response needs a different rubric. Start with examples that people can label consistently, including cases where none of the ordinary options fits. Decide how uncertain results should be handled before choosing a cutoff. A probability is useful only when its meaning has been checked on the distribution of inputs you actually receive.
What to record in your own comparison
- Compare with another classifier or a language model assigned the same fixed-choice task. Keep the label descriptions, test examples and scoring rules identical; record any output-parsing failures when the alternative returns free-form text.
- Separate a development set used to tune questions and thresholds from a held-out evaluation set. Include rare labels, ambiguous examples and inputs outside the intended category so aggregate accuracy cannot hide those failures.
- Measure precision and recall for the decisions that matter, and check calibration by grouping similarly scored examples. Inspect whether an apparent improvement comes from changing the decision threshold rather than from better discrimination.
- For local-versus-hosted comparisons, retain the checkpoint revision, runtime, input representation and exact question text. Report hardware and serving cost separately from the hosted token bill. Do not assume the published API score transfers unchanged to your installation.
Limitations and unknowns
This profile is documentation research. No checkpoint was downloaded, no inference was run and no latency or calibration was measured. Local runtime requirements are the model card’s guidance, not a benchmark on our hardware. The cited local example and hosted endpoint need not expose identical features. An input ceiling alone does not establish reliable performance across the full length. Exact local context and output limits, quantized-memory requirements and numerical parity with the hosted API remain unverified.
Compare your saved results
Use the DualView editor to inspect files you already have. The accepted-output calculator uses your own rates and counts. Neither link runs a model or purchases generation.
Related model profiles
- GPT-6 Astra
- Claude Opus 5.5
- GLM-5.2
- Gemini Omni 1.1 Flash
- Qwen-Image-2.1
- GPT-6 Sol
- DeepSeek-V4.1-Flash
- Claude Sonnet 5.5
- Muse Video
- MiniMax-M3.1-Flash-Preview
- GPT-6.1 Sol
- Amazon Nova Reel v1:0
- Hy Image 3.5 Preview
- Amazon Nova Canvas v1:0
- Gemini 4 Argon
- LongCat Video Distilled I2V 720p
- Ling 3.1 Flash
- HeyGen Video
- Amazon Nova Reel v1:1
- HiDream-O1-Image
- Tavus Griffin-Lite
- FLUX 3 Image
- FLUX 3 FAST Edit Video