Google is hedging future risks in artificial intelligence and is desperately trying to prove it can keep up with rival Microsoft, who has unexpectedly emerged as the frontrunner.
Google CEO Sundar Pichai managed to generate excitement about his company’s AI technology at the recent IO conference. But many questions remained unanswered about the hardware that drives next-generation AI and Google’s future.
CEO Sundar Pichai’s keynote at Google IO has been carefully curated, and the presentation highlights how Google plans to incorporate AI to make its products smarter and easier for users to use. embossed. For example, Google Docs uses AI to help users create documents such as cover letters, saving time.
But Google fell short of Microsoft in hyping the hardware it’s using to power the future of AI.
The hardware and infrastructure that powers AI is becoming an important part of the conversation, but Google didn’t provide details about its upcoming TPU or other hardware that will train its upcoming generative AI technology.
Pichai spent a few minutes talking about upcoming base models, such as Gemini, which will drive search and other product offerings.
Google CEO Sundar Pichai on stage at Google IO 2023. Credit: EnterpriseAI.
“Gemini was built from the ground up to be multimodal and highly efficient…and built to enable future innovations. We’re already seeing great multimodal capabilities,” Pichai said in his keynote.
Pichai said Gemini is still in training. However, company representatives declined to comment on whether the company’s next-generation AI chip, dubbed the TPU v5, is in production and being used to train the Gemini.
The Gemini model comes in different sizes and, like the PaLM and PaLM-2, can be adapted to different applications. But training a Gemini can take time.
DeepMind develops technologies targeting vertical industries, such as models of nuclear energy and protein folding, and targets artificial general intelligence, a kind of universal digital brain that can answer any question. It’s unclear if Gemini will compete with GPT-4 or GPT-5.
Google’s silence on its own hardware contrasts with Microsoft’s fuss about the cost per transaction and response times of its Nvidia GPU-based Azure supercomputer.
Gemini is expected to be the successor to PaLM-2, also announced at Google IO. The PaLM-2 model is an upgrade of his previous PaLM model used in Bard, the first AI chatbot used in an experimental search product. The Bard chatbot, based on PaLM-2, is in beta and has been rolled out to users in 180 countries.
PaLM-2 is detailed in a research paper published this month that talks about fine-tuning the model to be less toxic and more accurate in reasoning, providing answers and coding. According to the research paper, the model was trained on TPU-v4 hardware.
The ability to release AI technology faster and respond to user queries faster is related to hardware capabilities. Microsoft is quickly moving to Nvidia’s H100 GPU after implementing his OpenAI’s GPT-4 technology in its search and productivity applications.
Google turned to its homegrown Tensor Processing Unit to train generative AI models. Google’s last AI chip, the TPU v4, was officially released in 2021, but it had been in testing for several years before that.
The company has a TPU v4 supercomputer with 4,096 of its AI chips. Google claims the supercomputer is the first to feature a circuit-switched optical interconnect, and the company has deployed hundreds of his TPU v4 supercomputers in its cloud services.
At the recent ISC trade show in Hamburg, keynote speaker Dan Reed shared an interesting fact. Training for GPT-4 could take him over a year of continuous exascale-class computing.
“Think of any computational model that we run that can justify spending a calendar year running it…so we can make money. That’s what’s driving it. And that , is also driving what’s happening in the world’s innovation space.”It’s the underlying technology,” said Reid, president and professor of computer science at the University of Utah.
Reed said the innovation is in the hands of a handful of companies with the best AI hardware, including Google, Amazon and Microsoft.
It’s not yet clear what hardware was used to train GPT-4, but timing suggests it was likely on an Azure supercomputer powered by Nvidia’s A100 and H100 GPUs.
Google’s AI applications are mostly tuned to run on TPUs. But Google has added AI hardware variety with the announcement of the A3 supercomputer, which will host up to 26,000 of his Nvidia H100 GPUs. This product is aimed at companies using his Nvidia GPUs in their AI infrastructure. His CUDA software stack at Nvidia dominates AI workloads.
Google hasn’t released Gemini’s specs, but it’s likely trained on TPU-v4, said Dylan Patel, principal analyst at specialist consulting firm Semianalysis.
Patel said training the model could take less than a year.
“They didn’t give us any parameters or tokens, but it’s based on estimates and speculation, and it’s probably going to take months,” Patel said.
Patel said TPU v5 is expected to be introduced this year.
“They increased this number to 1,024 for TPUv3 and 4,096 for TPUv4. Based on the trendline, we assume that the current generation TPUv5 can scale up to 16,384 chips without going through inefficient Ethernet. said the SemiAnalysis authors in a newsletter entry.
Hints about TPU v5 are out there, but they’re mired in controversy.
Google researchers informally announced (Via Twitter) revealed the existence of TPU-v5 in June 2021 and published a paper in Nature on how AI was used to design the chip.
Google claimed that an AI agent could be trained to floorplan the placement of various blocks and modules in the optimal locations on the chip. The agent will learn over time and will be able to place previously unseen modules on the chip.
“Our method produces placements that are better or comparable to human expert chip designers in less than 6 hours, but the most performant alternatives do not require human experts to participate in the loop. and takes weeks for every dozen blocks of modern chips,” Google said. said in a paper.
But academic researchers weren’t amused by Google’s claims. Google was criticized for keeping his research private and not open to public scrutiny. The company eventually put a limited amount of information he put on Github.
Andrew B. Kahn, a researcher at the University of California, San Diego, has attempted to reverse-engineer Google’s chip design techniques and found that in some cases, human chip designers and automated tools outperform Google’s AI-specific technology. I have found that it can be faster. He presented a paper on his findings at the International Symposium on Physical Design in March.
The announcement of AI products at the IO conference fits with Google’s overall message of responsible use of AI, which was a big part of Sundar Pichai’s keynote. But after executives on stage showed off cool Google Pixel smartphones and tablets, the secure AI message was quickly forgotten.
But executives said there are more risks to using AI properly, especially when it comes under user or government oversight. The key is to deploy AI well, not to rush it. Microsoft has done it.
For example, a program called Universal Translator can dub the voice of a person speaking in a video into another language and change the lip-sync to make it appear that the person is speaking that language natively. increase. However, there are also concerns that the technology could be used to create deepfakes, and Google is sharing the technology with a limited number of vetted partners to prevent abuse.
The company also showcased an emerging technology, multimodal AI capabilities. AI tools are mostly specialized in his one thing only. ChatGPT excels at text-to-text responses, while Dall-E specializes in text-to-image responses. Multimodal AI integrates chat, video, voice, and images into a single AI model. Gemini integrates multimodal capabilities into a universal model.
Google IO’s second keynote of the day, aimed at developers, talked about technologies such as WebGPU, which was thought to be a protocol to accelerate AI on the desktop. WebGPU built into the browser utilizes local hardware resources to accelerate AI on PCs and smartphones. It is considered the successor to his WebGL used for graphics.
Related
