Vidu S2-Avatar vs S1: Resolution, References and Streaming Costs
Vidu S2-Avatar extends the S1 character-video model with documented 720p output and reference images that can change during a session. Those additions matter for interactive demonstrations and character experiences, but they do not establish a measured quality or latency advantage. This September 26 comparison separates the published upgrade from the access, cost and reliability questions an evaluator still needs to resolve.
An interactive character has to respond to what happens after a session begins. A recorded introduction can look convincing while revealing little about whether a character can adopt a new outfit, handle a referenced object or recover when a user changes direction. For that reason, the useful S1-to-S2 comparison concerns control during ongoing video generation as much as the dimensions of a frame.
This article compares Vidu S1 with Vidu S2-Avatar for that same interactive-character use case. Sources were checked September 26, 2026. The S1 paper was submitted July 3; the S2 paper followed September 10, and the provider update notice dates the S2 model introduction September 15. DualView has not run either model head to head. S2-Editing is a separate video-editing model and is not treated as an avatar substitute.
Documented model comparison
| Decision point | Vidu S1 | Vidu S2-Avatar |
|---|---|---|
| Core input and output | Custom character image and voice instructions; interactive generated video | Character image and persona; live interaction with text/voice and reference images |
| Documented resolution | 540p in the S1 paper | 720p in the S2 paper |
| Changing references | Dynamic reference updating is not established by the reviewed S1 abstract | Reference images can be updated during generation |
| Audio role | Voice controls digital-character video | Voice interaction is part of the avatar-video experience |
| Duration and limits | Paper describes continuous generation; current hosted session cap unverified | Streaming generation; current maximum session duration unverified here |
| Access and weights | Published paper; current independent S1 service availability unverified | Public demo/API integration links; downloadable weights not established |
| Comparable international cost | No current S1 rate verified on the reviewed pricing page | Avatar Real-Time: 1.5 credits per second; assumptions below |
What changes when a reference can arrive mid-session
The S2 paper identifies dynamic references as an addition to S1, alongside higher resolution and stronger instruction following. The product page describes object, outfit and background reference categories for the avatar experience. That creates a concrete distinction for a live product demonstration: the subject can be introduced after the character has already started responding. It is a documented control mechanism, not proof that a specific product label, shape or interaction will survive every generated frame.
An evaluator should separate three questions. Did the character use the newly supplied reference? Did it perform the requested action? Did its identity remain acceptable afterward? A successful background change alone does not answer the other two. If an existing S1 experience only needs a fixed character speaking against a fixed setting, the new control may have little value for that particular brief. A scenario with changing objects gives the addition a clearer purpose.
Resolution is a change; responsiveness still needs evidence
The S1 paper reports 540p generation and up to 42 frames per second. The current official research repository describes S2-Avatar at 720p and 25–42 frames per second. These are author-reported figures from different model descriptions. They do not form a controlled cross-model timing result, and this article does not translate them into a promise about how quickly a hosted character answers a user.
For an evaluation, record the output dimensions separately from the time between a completed instruction and the visible response. Playback can remain smooth while an action starts late; conversely, a prompt response can contain unstable detail. Examine faces, hands and small referenced objects at the intended display size. Keep the same framing and source character where practical. Increasing the available detail is useful only if the resulting stream satisfies the visual task, which cannot be established from resolution labels alone.
Public integration material does not establish self-hosted weights
The official integration repository now provides S2 Avatar and Editing examples, with separate pages for the two experiences. Its setup requires a Vidu API key and a regional API host. The described architecture uses a server for session creation and control requests, while the browser joins the RTC session to send and receive media. Running that example locally therefore still connects to the hosted service.
Neither the integration example nor the research repository establishes downloadable S1 or S2 model weights in the material reviewed here. The current demo links emphasize S2; a separately selectable S1 service was not verified. Treat S1 as the documented predecessor in this comparison, and confirm access before planning a fresh paired run. Existing S1 recordings can help explain a migration, but they should be labeled historical if settings, service revisions or recording conditions cannot be reproduced.
Normalize streaming cost without inventing an S1 discount
The international pricing page lists S2 Avatar Real-Time at 1.5 credits per second and a standard credit value of US$0.005. Multiplying those figures gives US$0.0075 per second, US$0.45 for 60 seconds or US$4.50 for ten minutes. These are arithmetic illustrations assuming every second is billable at that listed rate, excluding taxes, discounts and other charges. They are not observed invoices or a guarantee about session billing boundaries.
A current S1 rate was not verified on that same page, so no percentage saving or like-for-like cost advantage is claimed. Ask for the exact model identifier, region, billable-duration definition and concurrency terms before estimating a migration. Measure cost against usable session time as well as total elapsed time: failed starts, reconnects and discarded output can change the practical budget even when the advertised rate is unchanged. Those effects require actual billing records; none were collected for this article.
Choose the avatar model for the character task
The provider distinguishes interactive S2-Avatar from S2-Editing, which transforms incoming video streams using references. Both concern generated video, but their starting material and purpose differ. If the job is to keep an existing camera performance while changing its appearance, that belongs in an editing-model comparison. This article instead asks whether a generated character should move from its S1 baseline to the newer avatar model.
For a scripted introduction with no changing visual references, first identify the actual shortcoming in the existing output. It may be detail, character consistency or the response to a specific instruction. For a character that must react to new objects or clothing during a demonstration, S2 offers a directly relevant documented addition. In either case, set acceptance criteria before viewing provider examples. A pleasing sample from one scenario does not establish suitability for a different character, motion or reference sequence.
A useful paired review needs matched events and honest labels
A proposed evaluation can begin with one licensed character image and a short, fixed sequence of instructions: introduce the character, request a simple motion, present a reference and then return to the original activity. Repeat the sequence under the same practical conditions where both services are available. Record the actual model identifier, date, input assets and session settings. This is a suggested procedure, not an experiment performed for this article.
Review complete segments around each change, rather than selecting only attractive still frames. Align recordings by the instruction event and inspect identity, requested motion, reference use and recovery. Keep interruptions and failures visible. DualView can help inspect recorded video outputs side by side, but a comparison view cannot measure the network or service conditions that produced them. If current S1 access is unavailable, publish a clearly labeled S2 evaluation against the project requirements instead of manufacturing a matched benchmark from unrelated historical footage.
Open DualView video comparison to review your own video outputs.
Frequently asked questions
Is S2-Avatar a new release on September 26?
No. The sources were checked on September 26. The S2 paper was submitted September 10 and the official update notice records the model introduction September 15, 2026.
Does the S2 API example include model weights?
The reviewed example connects to the hosted Vidu service using an API key and RTC. Local execution of that application does not establish that the underlying model weights can be downloaded.
Can the two versions be ranked for latency from these sources?
No. Reported frame throughput and a resolution specification do not provide a matched measurement of instruction-to-response delay. A defensible ranking needs the same task, recorded conditions and actual observations.
Official sources
- Vidu S1 paper — July 3, 2026
- Vidu S2 paper — September 10, 2026
- Vidu S2 product description and FAQ
- Official Vidu S API integration repository
- Official Vidu S research repository
- Vidu international API pricing
- Vidu model update notice
No paid model generations were performed for this article. The proposed review procedure is editorial analysis, not an executed experiment.