AWS has published a tutorial on deploying an image-to-video pipeline with the FLUX.2-klein-4B and Wan2.1-VACE-1.3B models on SageMaker AI
AWS has published a tutorial on deploying an image-to-video pipeline (FLUX.2-klein-4B + Wan2.1-VACE-1.3B) on SageMaker AI through a combination of a real-time and an asynchronous endpoint within the same AWS vLLM-Omni container.
Amazon Web Services has published the second installment of a tutorial series on deploying multimodal models using the AWS vLLM-Omni Deep Learning Container on Amazon SageMaker AI. This installment describes a pipeline that generates an image from a text prompt using the FLUX.2-klein-4B model and then animates it into a short video using the Wan2.1-VACE-1.3B model. Both models run within the same pinned container image, with model selection controlled by the SM_VLLM_MODEL environment variable, set separately for each endpoint.
The image endpoint is deployed as real-time (ml.g6.xlarge instance) and returns a base64-encoded PNG directly in the response, since the application needs the result immediately to assemble the next request. The video endpoint runs as asynchronous (ml.g6e.xlarge instance) – the input multipart request is stored in Amazon S3, and the resulting MP4 is written there as well once video generation completes, since this is a longer-running operation with an image-conditioned payload. According to the article, SageMaker Asynchronous Inference also allows smaller inputs to be sent directly (Body up to 128 000 bytes), but here InputLocation via S3 was used due to the size of the multipart request.
The default video settings are 17 frames and 30 diffusion steps; according to the authors, a lower number of steps (4) did not preserve the composition of the source image and is suitable only for a quick endpoint test. The solution includes both a command-line CLI workflow and an optional Streamlit interface. In a test run in the US East (N. Virginia) region, the ml.g6.xlarge image endpoint reached InService status in 9 minutes 30 seconds; the rest of the data from this measurement is not available in the source.
For details on setting instance quotas, pricing, and further deployment steps, see the source article.
Why it matters
The tutorial shows a concrete architecture for combining an image generative model and a video generative model within a single production infrastructure, including the decision between real-time and asynchronous processing based on latency and data size. For teams looking to deploy a similar feature on AWS, this shortens the time needed to design a solution, but they must still account for GPU instance costs and verify available quota in the region.
Two audiences, two different impacts
What this means
For individuals
Following the tutorial, a developer can build a custom text-image-video pipeline on SageMaker AI and choose between a real-time and an asynchronous endpoint depending on the latency and payload size of their model.
For a business
Companies building generative image/video pipelines on AWS gain a reference architecture combining real-time and asynchronous endpoints in a single container, but must factor in the costs of GPU instances (ml.g6.xlarge, ml.g6e.xlarge) and verify quota and pricing before deployment.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.