TL;DR: We ran MiniMax H3 image-to-video on a rented Vast.ai A100 SXM4 80GB. A 5-second, 576×1024 clip took about 86 seconds to generate and cost roughly $0.03 in GPU and disk time. The useful preparation happened before inference: put the source image on the exact output canvas, allow room for model weights, check inbound traffic charges, and plan a batch large enough to spread setup time across many clips. Download the results and destroy the instance when the session is done.
Why we rented an A100
We are a small content team that uses rented GPUs for AI video work. Renting makes sense for us when we have a queue of clips ready to run, especially with models whose weights and memory needs make a local setup awkward.
For this session, we used a Vast.ai A100 SXM4 80GB in Czechia. The GPU rate was $1.19 per hour, or about $1.25 per hour including disk. Our H3 image-to-video jobs used bf16 and a turbo 8-step LoRA. Each output was 576×1024, with 124 frames at 24 fps: approximately five seconds of video.
The same rental workflow can be used to evaluate models in the Wan family, but the measurements in this article are from our MiniMax H3 run. Different weights, inference scripts, and settings can change both memory use and speed. We would treat any timing estimate for another model as a fresh measurement.
If you need to open an account before trying the workflow, this is our Vast.ai referral link. Check the current offer and pricing yourself before renting; marketplace listings change.
Prepare the images before starting the meter
The input image was the first thing we learned to check. H3 stretched first_frame to the requested output dimensions. If the source image was on a different aspect ratio, the subject could look distorted before any motion began.
For a 576×1024 output, we placed the source image on a 576×1024 canvas ahead of time. That let us decide how to crop or pad it while we could inspect the result. We then checked the prepared image at the size the model would receive.
We also used the same image for the first and last frame when we wanted a loop. In our run, first_frame = last_frame produced clean loops. It is a useful constraint for short background motion or repeating shots, though the prompt and the image still need to describe motion that makes sense between those endpoints.
A prompt should focus on visual motion the model can plausibly supply. We tried asking for someone to “write on a board”; the resulting letters were gibberish. If readable text matters, plan to add it in a later editing step rather than relying on generated handwriting.
Before renting, we assembled the input images and prompts into a queue. This matters because setup time is paid time. Searching for images, fixing aspect ratios, and revising a shot list while the GPU is running make an inexpensive clip less inexpensive.
Select an offer with the full bill in mind
The hourly GPU rate is only one field in an offer. Our earlier test on an RTX PRO 6000 Blackwell 96GB showed why: the listing was $1.625 per hour, and a 56-minute session totaled $2.18. GPU time accounted for $1.47; inbound traffic for 245 GB of model weights added $0.64.
When browsing Vast.ai offers, we now inspect these fields together:
| Field | What we check |
|---|---|
gpu_name, num_gpus
|
The intended GPU and number of GPUs |
reliability |
At least 0.97 for our searches |
disk_space |
Enough room for the image, weights, inputs, and outputs |
inet_down_cost |
Zero when we expect to pull large weights |
cuda_vers, cuda_max_good, driver_version
|
Compatibility with the container image |
rentable, verified, geolocation
|
Availability and the offer details |
direct_port_count, inet_down
|
Connectivity details relevant to the setup |
For a large weight download, we provision at least 150 GB of disk and check the offer’s storage capacity before creating the instance. A 40 GB disk may be enough for a small model; it is a poor default for this job. Our H3 session downloaded 72 GB of weights in about eight minutes.
Driver compatibility deserves its own check. An A100 offer we tried in Sweden had driver version 535, which did not fit the image we intended to run. The GPU model alone did not tell us whether the container would work. We check driver_version and the CUDA fields against the image before committing to an offer.
Here is the search shape we use with the Vast.ai CLI:
vastai search offers \
'reliability>=0.97 num_gpus=1 disk_space>=150 inet_down_cost=0' \
-o 'dph+' \
--raw
Read the returned offer fields before choosing an ID. Sorting by hourly price helps narrow the list, but a cheap listing with paid inbound traffic or an incompatible driver can cost more time and money than expected. Keep API credentials in your environment or a protected file; do not paste them into a shared command, log, or article.
Bootstrap once, then run a batch
We built a public Docker image on GHCR with GitHub Actions. That build took about 11 minutes. Our onstart.sh downloaded only the weights selected through DL_GROUPS; for this session, the value was h3. The Vast on-start command was bash /workspace/bootstrap/onstart.sh.
That arrangement kept the container setup repeatable. It also let us separate the reusable software image from large model weights. The precise inference command depends on the checkout and weights you install, so verify it against the model’s own instructions rather than copying a command written for a different repository.
After selecting a suitable offer, the instance creation command has this form. Set OFFER_ID to the selected offer and IMG to the prepared container image before running it; onstart.sh is the local on-start file passed to Vast.ai.
: "${OFFER_ID:?set OFFER_ID to the selected offer ID}"
: "${IMG:?set IMG to the prepared container image}"
vastai create instance "$OFFER_ID" \
--image "$IMG" \
--disk 150 \
--label h3-batch \
--onstart onstart.sh \
--env '-e DL_GROUPS=h3' \
--ssh \
--direct
vastai show instances --raw
Once the instance is running, vastai ssh-url ID gives its SSH connection details. We wait for the on-start download to finish, confirm the expected weights are present, and run one prepared image first. That first clip checks the canvas, prompt, output path, and model setup before the rest of the queue spends GPU time.
Our measured H3 clip took about 86 seconds. At roughly $1.25 per hour with disk, that works out to about $0.03 of rental time for the generation itself. We used an nvfp4 text encoder on Ampere during this work, and it worked in our setup. That observation is specific to the configuration we ran, so we would still test a new image or weight set with one clip before launching a long batch.
Keep setup time from becoming clip time
We generated 82 hero clips in one session. The session cost $3.59 and took about three hours including setup. Dividing the full session bill by the outputs gives a higher per-clip figure than the roughly $0.03 generation cost, because downloads and other setup time share the same meter.
That distinction is useful when comparing rental with an API. The H3 API price in our notes was $0.08 per second of video. Five seconds at that price is $0.40. Our measured generation time on the A100 was much cheaper per clip, while the full session bill also included getting the machine ready. Check current API pricing pages before making a new purchasing decision; those prices may have changed.
Batching is the simplest way we found to make rental practical. Have the inputs ready, run a first-clip check, then keep the GPU occupied with the remaining queue. If you need only one output, setup and weight downloads will dominate its effective cost.
We used the same lesson on a separate LatentSync 1.6 lip-sync run. A queued self-test took 15.9 minutes in total, with about 11 minutes spent on setup. That is a different model and GPU, but it illustrates why we prepare work before starting a rental.
Download, verify, and destroy
A stopped Vast.ai instance can continue billing for storage. We left one stopped for days; its 90 GB disk cost about $0.0167 per hour, or roughly $0.40 per day. The balance went negative, the stopped instance could not be restarted with that balance, and we had never downloaded the trained result. That was a preventable loss.
Our end-of-session order is now fixed: copy outputs back, back them up, verify the copy, then destroy the instance. In one backup check, we verified that 5,342 files totaling 1.379 GiB matched locally. A successful copy command alone is less reassuring than checking both count and size.
We also use a $1 cost cap and a detached watchdog process set to destroy the instance after a deadline, retrying three times. It protects against a controller crash leaving the meter running. It does not replace checking that the output files arrived.
The CLI destroy command prompts for confirmation. For a scripted cleanup, we use the confirmed form below, then inspect the instance list:
: "${INSTANCE_ID:?set INSTANCE_ID to the instance being removed}"
echo y | vastai destroy instance "$INSTANCE_ID"
vastai show instances --raw
The old v0 instances listing once returned an error that looked like “0 instances” to us. We verify destruction against the current instances view as well: https://console.vast.ai/api/v1/instances/, authenticated with an Authorization: Bearer <key> header. Keep that key out of command output and shared logs.
What it cost us
On the A100 SXM4 80GB in Czechia, the GPU was $1.19 per hour and about $1.25 per hour with disk. A 5-second H3 image-to-video clip took about 86 seconds and cost roughly $0.03 in generation time. The 82-clip session cost $3.59 over about three hours including setup; 72 GB of weights downloaded in about eight minutes. In the earlier RTX PRO 6000 Blackwell test, 56 minutes cost $2.18, including $0.64 of inbound traffic for 245 GB of weights. Our reference H3 API price was $0.08 per second of video.
The workable routine for us is to prepare exact-aspect images, rent an offer whose driver and transfer costs fit the job, validate one clip, run the queue, and verify every output before destroying the instance. That routine matters more than the lowest hourly price on a listing.
Top comments (0)