In June, Runway debuted a new text-to-video synthesis mannequin referred to as Gen-3 Alpha. It converts written descriptions referred to as “prompts” into HD video clips with out sound. We’ve since had a probability to make use of it and wished to share our outcomes. Our exams present that cautious prompting is not as necessary as matching ideas possible discovered within the coaching information, and that reaching amusing outcomes possible requires many generations and selective cherry-picking.
An enduring theme of all generative AI fashions we have seen since 2022 is that they are often glorious at mixing ideas present in coaching information however are sometimes very poor at generalizing (making use of realized “information” to new conditions the mannequin has not explicitly been skilled on). That means they will excel at stylistic and thematic novelty however wrestle at basic structural novelty that goes past the coaching information.
What does all that imply? In the case of Runway Gen-3, lack of generalization means you would possibly ask for a crusing ship in a swirling cup of espresso, and supplied that Gen-3’s coaching information consists of video examples of crusing ships and swirling espresso, that is an “simple” novel mixture for the mannequin to make pretty convincingly. But should you ask for a cat ingesting a can of beer (in a beer industrial), it will usually fail as a result of there aren’t possible many movies of photorealistic cats ingesting human drinks within the coaching information. Instead, the mannequin will pull from what it has realized about movies of cats and movies of beer commercials and mix them. The result’s a cat with human hands pounding again a brewsky.
Just a few fundamental prompts
During the Gen-3 Alpha testing part, we signed up for Runway’s Standard plan, which gives 625 credit for $15 a month, plus some bonus free trial credit. Each era prices 10 credit per one second of video, and we created 10-second movies for 100 credit a piece. So the amount of generations we might make have been restricted.
We first tried a few requirements from our picture synthesis exams up to now, like cats ingesting beer, barbarians with CRT TV units, and queens of the universe. We additionally dipped into Ars Technica lore with the “moonshark,” our mascot. You’ll see all these outcomes and extra under.
We had so few credit that we could not afford to rerun them and cherry-pick, so what you see for every immediate is strictly the only era we acquired from Runway.
“A highly-intelligent individual studying “Ars Technica” on their laptop when the display explodes”
“industrial for a new flaming cheeseburger from McDonald’s”
“The moonshark leaping out of a laptop display and attacking a individual”
“A cat in a automobile ingesting a can of beer, beer industrial”
“Will Smith consuming spaghetti” triggered a filter, so we tried “a black man consuming spaghetti.” (Watch till the top.)
“Robotic humanoid animals with vaudeville costumes roam the streets accumulating safety cash in tokens”
“A basketball participant in a haunted passenger prepare automobile with a basketball court docket, and he’s enjoying in opposition to a workforce of ghosts”
“A herd of 1 million cats operating on a hillside, aerial view”
“video sport footage of a dynamic Nineties third-person 3D platform sport starring an anthropomorphic shark boy”



