In partnership with

Blu Dot surpasses 2,000% ROAS with self-serve CTV ads

Home furniture brand Blu Dot blew up on CTV with help from Roku Ads Manager. Here’s how:

After a test campaign reached 211,000 households and achieved 1,010% ROAS, the brand went all in to promote its annual sales event. It removed age and income constraints to expand reach and shifted budget to custom audiences and retargeting, where intent was strongest.

The results speak for themselves. As Blu Dot increased their investment by 10x, ROAS jumped to 2,308% and more page-view conversions surpassed 50,000.

“For CTV campaigns, Roku has been a top performer,” said Claire Folkestad, Paid Media Strategist, Blu Dot. “Comping to our other platforms, we have seen really strong ROAS… and highly efficient CPMs, lower than any other CTV partner we've worked with.”

Using Roku Ads Manager, the campaign moved from a pilot to a permanent performance engine for the brand.

AI video generation is moving from simple text-to-video systems toward multimodal production models. These systems can interpret combinations of text, images, video, audio, and reference material.

Alibaba’s Wan3.0 is part of this shift. The model is currently available in preview through Alibaba Cloud Model Studio. It combines several video-generation tasks within one model architecture.

What Wan3.0 can generate

Wan3.0 supports text-to-video, image-to-video, first-and-last-frame generation, and reference-based video generation.

A single generation can reach 30 seconds at 30 frames per second. This matters because many earlier AI video systems focused on much shorter clips. Longer generation gives the model more time to handle scene development, movement, dialogue, and transitions.

Its reference system is another important change. Wan3.0 can work with images, video clips, and audio files as source material. Alibaba also supports document and webpage inputs within the model workflow.

This suggests AI video generation is becoming less dependent on a single written prompt. Source materials can increasingly define characters, visual structure, movement, sound, and narrative context.

Seedance 2.5 takes a similar direction

ByteDance released Seedance 2.5 on July 31, 2026. It can generate up to 30 seconds of synchronized audio and video in one generation. It also supports repeated extensions.

Seedance accepts up to 30 images, 10 videos, and 10 audio clips as references. It also supports timestamp-based editing. Users can specify when actions, camera changes, or other events should occur.

Compared with Wan3.0, Seedance currently allows a larger number of individual media references. Wan3.0, however, extends the input concept toward documents and webpages.

Introducing The First Agentic CRM

Get revenue agents, workflows, and automations across every stage of your motion. Access customer data in real time through Attio's web app, MCP, API, and SDK.

Then Ask Attio anything about your business and get instant answers.

It's the CRM that runs the work behind every win.

Kling 3.0 focuses on integrated video and audio

Kuaishou released Kling 3.0 in February 2026. Its video models support generation up to 15 seconds.

Kling combines text, images, audio, and video within the same system. It also supports native audio generation across multiple languages, dialects, and accents. Text-to-video, image-to-video, reference-to-video, and video editing are included within the model family.

Veo 3.1 and Runway Gen-4.5

Google’s Veo 3.1 generates video with native audio. It accepts reference images for characters, objects, scenes, and visual styles. Google also provides scene extension and first-and-last-frame generation. Individual Veo generations are generally eight seconds in Google’s published evaluations.

Runway Gen-4.5 takes a narrower input approach. It currently supports text-to-video and image-to-video. Generation length ranges from two to ten seconds at 1280×720 resolution.

The research direction

These models show three connected changes in AI video research.

Generation length is increasing. Reference inputs are becoming more varied. Audio and video are increasingly generated within the same process.

Wan3.0 adds another direction by connecting video generation with documents and webpages. That could make source-based video creation more common in education, reporting, research communication, and business documentation.

The technical problem will then move beyond visual quality. Systems must also preserve source meaning, factual accuracy, character consistency, and temporal continuity across longer generated videos.

AI Insights. Real Growth. Higher GMV, Better Profits

The difference between growing stores and stagnant ones isn't more effort. It's better insights. StoreClaw analyzes your Shopify and Amazon data, surfaces your biggest growth opportunities, and helps you increase GMV while protecting profit. Start free with bonus tokens. No credit card required.