A film from a prompt or a still
Text to video and image to video on Wan 2.2, in a full and a distilled "turbo" variant. Iterate on turbo at a third of the price, then run the locked shot through the full model at up to 1080p.
We run open-weight video models on our own infrastructure and meter them by the second — text to video, image to video, lip-sync, motion transfer, and your own LoRAs. The platform is Phosor.
Text to video and image to video on Wan 2.2, in a full and a distilled "turbo" variant. Iterate on turbo at a third of the price, then run the locked shot through the full model at up to 1080p.
Lip-sync driven by an audio track, motion transferred from a reference performance, and speech synthesis if you do not have a voice yet.
Scene variations, mannequin to model, pose changes, background replacement and in-image text localisation — at 2K or 4K.
Train a LoRA on your own dataset, or import .safetensors you
already own. Several can apply to a single request.
Custom ASIC and software co-design for zero-knowledge proving — more than 20x the throughput of a GPU prover on the same workload, delivered into production hardware with Intchains.
The gap between open weights and the closed frontier is mostly post-training and data curation, not architecture. Here is the evidence, and what it costs to close.
Every published reward model and preference-optimisation method for video generation, ordered by what it costs and what it has been measured to buy.
Open weights reached the top of the human-preference leaderboard in August 2026. What that changes, and the three constraints that stop it being a clean win.
Built with