7

Meta's Make-A-Video AI achieves a new, nightmarish state of the art

 1 year ago
source link: https://finance.yahoo.com/news/metas-video-ai-achieves-nightmarish-193019109.html
Go to the source link to view the article. You can view the picture content, updated content and better typesetting reading experience. If the link is broken, please click the button below to view the snapshot at that time.
neoserver,ios ssh client

Meta's Make-A-Video AI achieves a new, nightmarish state of the art

Devin Coldewey
Fri, September 30, 2022, 4:30 AM·3 min read

Meta's researchers have made a significant leap in the AI art generation field with Make-A-Video, the creatively named new technique for — you guessed it — making a video out of nothing but a text prompt. The results are impressive and varied, and all, with no exceptions, slightly creepy.

We've seen text-to-video models before — it's a natural extension of text-to-image models like DALL-E, which output stills from prompts. But while the conceptual jump from still image to moving one is small for a human mind, it's far from trivial to implement in a machine learning model.

Make-A-Video doesn't actually change the game that much on the back end — as the researchers note in the paper describing it, "a model that has only seen text describing images is surprisingly effective at generating short videos."

The AI uses the existing and effective diffusion technique for creating images, which essentially works backwards from pure visual static, "denoising" towards the target prompt. What's added here is that the model was also given unsupervised training (that is to say, it examined the data itself with no strong guidance from humans) on a bunch of unlabeled video content.

What it knows from the first is how to make a realistic image; what it knows from the second is what sequential frames of a video look like. Amazingly, it is able to put these together very effectively with no particular training on how they should be combined.

"In all aspects, spatial and temporal resolution, faithfulness to text, and quality, Make-A-Video sets the new state-of-the-art in text-to-video generation, as determined by both qualitative and quantitative measures," write the researchers.

It's hard not to agree. Previous text-to-video systems used a different approach and the results were unimpressive but promising. Now Make-A-Video blows them out of the water, achieving fidelity in line with images from perhaps 18 months ago in original DALL-E or other past generation systems.


About Joyk


Aggregate valuable and interesting links.
Joyk means Joy of geeK