Just someday after OpenAI revealed GPT-4o, which it payments as with the ability to perceive what’s happening in a video feed and converse about it, Google introduced Project Astra, a analysis prototype that options related video comprehension capabilities. It was introduced by Google DeepMind CEO Demis Hassabis on Tuesday at the Google I/O convention keynote in Mountain View, California.
Hassabis known as Astra “a common agent useful in on a regular basis life.” During an indication, the analysis mannequin showcased its capabilities by figuring out sound-producing objects, offering artistic alliterations, explaining code on a monitor, and finding misplaced gadgets. The AI assistant additionally exhibited its potential in wearable units, similar to good glasses, the place it might analyze diagrams, counsel enhancements, and generate witty responses to visible prompts.
Google says that Astra makes use of the digicam and microphone on a consumer’s system to supply help in on a regular basis life. By constantly processing and encoding video frames and speech enter, Astra creates a timeline of occasions and caches the knowledge for fast recall. The firm says that this permits the AI to establish objects, reply questions, and keep in mind issues it has seen which can be now not within the digicam’s body.
Project Astra: Google’s imaginative and prescient for the way forward for AI assistants.
While Project Astra stays an early-stage function with no particular launch plans, Google has hinted that a few of these capabilities could also be built-in into merchandise just like the Gemini app later this yr (in a function known as “Gemini Live”), marking a major step ahead within the growth of useful AI assistants. It’s a stab at creating an agent with “company” that may “suppose forward, cause and plan in your behalf,” within the phrases of Google CEO Sundar Pichai.
Elsewhere in Google AI: 2 million tokens
During Google I/O, the corporate unveiled a lot of AI-related bulletins, a few of which we might cowl in separate posts sooner or later. But for now, this is a fast overview.
At the highest of the keynote, Pichai talked about an “improved” model of February’s Gemini 1.5 Pro (similar model quantity, oddly) that’s coming quickly. It will function a 2 million-token context window, which suggests it could course of massive numbers of paperwork or lengthy stretches of encoded movies at as soon as. Tokens are fragments of knowledge that AI language fashions use to course of data, and the context window determines the utmost variety of tokens an AI mannequin can course of at as soon as. Currently, 1.5 Pro tops out at 1 million tokens (OpenAI’s GPT-4 Turbo has a 128,000 token window for comparability).
We requested AI researcher Simon Willison—who doesn’t work for Google however was featured in a promo video in the course of the keynote—what he considered the context window announcement. “Two million tokens is thrilling,” he replied by way of textual content whereas sitting within the keynote viewers. “But it is price retaining value in thoughts that $7 per million tokens means a single immediate might value you $14!” Google expenses $7 per million enter tokens for 1.5 on prompts longer than 150,000 tokens by means of its API.
Speaking of tokens, Google introduced that its beforehand introduced 1 million token context window for Gemini 1.5 Pro is lastly coming to Gemini Advanced subscribers. Previously, it was solely accessible within the API.
Google additionally introduced a brand new AI mannequin known as Gemini 1.5 Flash, which it billed as a light-weight, quicker, and cheaper model of Gemini 1.5. “1.5 Flash is the most recent addition to the Gemini mannequin household and the quickest Gemini mannequin served within the API. It’s optimized for high-volume, high-frequency duties at scale,” says Google.
Willison had a touch upon Flash as effectively: “The new Gemini Flash mannequin is promising there, it is meant to supply as much as 2m tokens at a cheaper price.” Flash prices $0.35 per million tokens on prompts as much as 128,000 tokens and $0.70 per million tokens for prompts longer than 128,000. It’s one-tenth the value of 1.5 Pro.
“35 cents per million tokens! That’s the most important information of the day, IMO,” Willison advised us.
Google additionally introduced Gems, which seems to be its tackle OpenAI’s GPTs. Gems are customized roles for the Google Gemini chatbot that can play a component that you just outline, permitting you to personalize Gemini in several methods. Google lists examples of potential Gems as “a health club buddy, sous chef, coding companion or artistic writing information.”
New generative AI fashions

Also at the Google I/O keynote on Tuesday, Google introduced a number of new generative AI fashions for creating photos, audio, and video. Imagen 3 is the newest in its line of picture synthesis fashions, which Google says is its “highest high quality text-to-image mannequin, able to producing photos with even higher element, richer lighting and fewer distracting artifacts than our earlier fashions.”
Google additionally confirmed off its Music AI Sandbox, which Google payments as “a set of AI instruments to rework how music may be created.” It combines its YouTube music undertaking with its Lyria AI music generator into instruments for musicians.
The firm additionally introduced Google Veo, which is a text-to-video generator that creates 1080P movies from prompts in a top quality that appears to match OpenAI’s Sora. Google says it’s working with actor Donald Glover to create an AI-generated demonstration movie that can debut quickly. It’s removed from Google’s first AI video generator, nevertheless it appears to be its most succesful to date.
The pattern video above, offered by Google, used the immediate, “A lone cowboy rides his horse throughout an open plain at lovely sundown, delicate mild, heat colours.”
Google says beginning at the moment, its new AI artistic instruments can be found to pick out creators in a personal preview solely however that wait lists are open.



