This is a rush transcript. Copy may not be in its final form.
AMY GOODMAN: This is Democracy Now!, democracynow.org. I’m Amy Goodman in London, with Nermeen Shaikh in New York.
NERMEEN SHAIKH: We’re continuing to look at how artificial intelligence is rapidly changing the world. We’re joined now by a leader in the field of AI ethics, Timnit Gebru. She is the founder and executive director of the Distributed Artificial Intelligence Research Institute, or DAIR. Prior to that, she served as co-lead of the Ethical AI research team at Google. She was fired in 2020 for writing a paper warning about the dangers of large language models and raising issues of discrimination in the workplace. Timnit also co-founded Black in AI, a nonprofit that works to increase the presence, inclusion, visibility and health of Black people in the field of AI. Her forthcoming book is titled Deep Unlearning: The Rise of AI and the Radicalization of a Tech Idealist. She joins us now from Boston, Massachusetts.
Timnit, could you begin by explaining what AI means to you? You’ve said you think of AI as a discipline with a number of subspecialties. Explain.
TIMNIT GEBRU: Yes, and I’m glad that Heidy did some groundwork for us here in the earlier segment. But yes, for me, AI is a discipline with a collection of specialties that are all lumped into this term “AI,” and they can be very different. So, Heidy mentioned things like large language models, but also computer vision systems that have been around since the Vietnam War. And all of these things are bundled into the term “AI.”
And in fact, every — the things that are considered AI are not always considered AI. So, for example, the title of my forthcoming book is Deep Unlearning, and this is a play on the term “deep learning,” which is a technique which is currently, basically, synonymous with AI. But back in the day, deep learning was not considered AI, so much so that the researchers in that specialty were having their academic papers rejected by the prestigious AI conferences. And so, actually, deep learning was sort of a rebranding from another term, “artificial neural networks,” to avoid that stigma.
So, not only do we have this marketing term that is a hodgepodge of disparate techniques or tools, in fact, it’s not even — sometimes it’s techniques, sometimes it’s tools, sometimes it’s products. But also, not all of these things are always called AI, right? When one of these subspecialties sort of is really hyped up or has some sort of better performance on a specific benchmark, it gets elevated to AI. And then it gets sort of downgraded, not to AI. And there are other such examples, like expert systems of the ’80s. That was what was synonymous with AI in the ’80s, and now nobody calls expert systems AI.
So, what happens here is that it makes it so difficult to have conversations that are grounded and specific, and have specific conversations about what are the harms, what are the benefits, what approaches are good, what approaches are bad, because when I talk about some things that I’m really against, like, you know, ChatGPT kind of things, then people are like, “What about a particular system that is used for early diagnosis of cancer patients?” You know? And I have to say, “Well, that’s a completely different system. It doesn’t even have to be built in the same way that these big models are being built.”
AMY GOODMAN: Timnit Gebru, can you talk about co-founding Black in AI, but start off by talking about why you were fired from Google in 2020, the issues that you were raising there, and the position you had there?
TIMNIT GEBRU: So, yeah, I co-founded Black in AI way before I even joined Google. My current institute is called the Distributed AI Research Institute, which I founded after I was fired from Google. But I started Black in AI in, you know, around 2016, 2017, when I attended — you know, I started seeing simultaneously the lack of Black people in AI. And so, I would go to academic conferences, and you’d have about 5,000-6,000 people in those conferences, and only a couple of, one or two, you know, a handful of Black people. And at the same time, we had some of these systems, some of the kinds of systems that Heidy was talking about.
For instance, there was a ProPublica article from 2016 talking about a company purporting to determine someone’s likelihood of committing a crime again. And so, this company was saying that they built a software that can tell whether a person released from prison was likely to commit a crime again. And the ProPublica article was talking about how these systems were more likely even to label Black people as criminals, potential criminals, than white people.
And it was very scary, because judges were already using the outputs of these models in their decisions about bail, decisions about how long someone should be in prison for. So, the dichotomy of seeing the lack of Black people in the field and also the kinds of claims that were being made and the kinds of things these tools were being built for was very scary. And so, that’s kind of how I decided to found Black in AI.
So, by the time I started working at Google in 2018, I was a very known quantity. I had already, you know, founded and led Black in AI. I was working on uncovering all sorts of issues in the field. My collaborator Joy Buolamwini and I had written a very — a paper that was pretty — it made the rounds and was highlighted all over the place and even changed policy, showing for the first time that automated facial analysis tools, like face recognition, for instance, tools, were — had much higher error rates for darker-skinned women than lighter-skinned men. And the darker and darker the skin, and especially for women, the higher and higher the error rate. So, think about these tools being used on CCTV cameras to identify so-called criminals, etc., and the people who are of darker skin would be more likely to be wrongfully identified in this case. And so, I had already worked on these kinds — uncovering these kinds of issues.
And so, I was hired at Google in 2020 to co-lead a team called the Ethical AI research team, which was a small research team, with Meg Mitchell, who was also later fired. And there were many issues. I mean, it was one issue after another. This was in the middle of the Google walkout, where 20,000 women walked out because we found out that Google had paid Andy Rubin $91 million after allegations of sexual misconduct. And so, it was a combination of a whole bunch of issues, and it was at the height of Black Lives Matter, the Black Lives Matter movement in 2020.
And at this time, a company called OpenAI, that was getting very famous but not as famous as it is today, came out with a large language model called GPT-3, which later became the backbone of ChatGPT, but ChatGPT wasn’t released yet. And GPT-3 was really hyped up in the mainstream media, and OpenAI was making all sorts of claims about how powerful GPT-3 was. In fact, actually, they had said that its predecessor, GPT-2, was too dangerous to release because it’s so powerful. And so, the GPT-3 was a large language model. And we can just define what large language models are. They are models that are trained on vast amounts of textual data on the internet, and they’re trained to calculate the most likely sequences of text based on the training data. So, if you use large language models to generate text, you are kind of trying to generate the most likely sequences of text given your training data.
So, the fact that OpenAI got so much airtime for GPT-3 meant that all of these companies wanted to build similar models. And they wanted to build larger and larger models, which means that they were guzzling all of the data on the internet, and they were consuming, using huge amounts of computational power even then. And so, we were very worried about this race that was started. Every company wanted to have the largest model. And we were asking, “Why do we need the largest of anything? What problem are you trying to solve?” And we warned about — we wrote a paper, my collaborators and I, warning about the dangers of large language models. And we said, how big can large — “Can language models be too big?”
So, the first issue that we mentioned, the very first one, was the environmental and financial cost. And you can see how, six years later now, that is not only true, but even worse than we said in the paper. We talked about how the carbon footprint of these huge data — the necessary computational power necessitates huge carbon footprint. And even if you claim to have data centers with renewable energy, that energy is going away from heating people’s homes, for example, and towards training these models. And the financial cost that shuts out anybody who doesn’t have the wealth to participate in this kind of work.
We also warned about perpetuating hegemonic views. The claim was that because these models are large, which means that they are using huge data sets, then, oh, they have all of human knowledge in these data sets. But that’s not true. The internet represents hegemonic views. It does not represent views of everybody in the world. A lot of people are not even on the internet. We even know articles like Wikipedia are heavily biased. It’s overwhelmingly Western and male.
And we also warned about — and this is very related to what Heidy was talking about earlier. We talked about how, because large language models tend to output fluent and coherent text, this can be very deceiving. So, one example we gave was from 2017. And so, these large language models, you know, they could be used — they were used as systems in a whole bunch of other things like machine translation systems. And in 2017, a Facebook translate translated a Palestinian’s “good morning” into “attack them.” And, you know, “attack them,” there’s no grammatical error or anything like that, so it doesn’t even give you a cue that this translation might be incorrect. And this person was arrested and later released. And they didn’t even check to see what the untranslated version was when they arrested him. And we call this automation bias, the tendency to overtrust the outputs of automated systems.
And so, we warned about also people interacting with these kinds of texts output by these kinds of systems, attributing a mind behind whatever text they’re seeing. And so, you know, right now, these chatbots, like Claude or ChatGPT, are designed to make you believe that there’s some sort of superhuman brain behind whatever is outputting these texts. And that’s hugely dangerous. And even though we weren’t speaking about chatbots back then, we were already talking about the tendency of people to attribute a mind behind these texts, and so, which means that, you know, they can be deceived into believing that there’s something that the — something more than a model outputting text.
And the final thing we warned about was: What is the cost of pursuing this singular direction in terms of research and resources? What is the cost of shutting out all other possible futures, all other ways of building machine learning systems, which is a subspecialty of artificial intelligence, or any other kind of paradigm? Because it seemed like the whole field was only investing in this singular direction of building larger and larger models, that are guzzling all of the data — I mean, stealing so many people’s works — and based on exploited labor and killing the environment.
And so, of course, you know, I got fired after that. I mean, there’s details, but that was the icing on the cake, was writing that paper.
NERMEEN SHAIKH: So, Timnit, you know, we only have a couple of minutes left, but just to point out that, you know, the scale of which these AI bots are being used, very recently ChatGPT crossed a billion active users a month. Google Gemini has said now that it has done the same. But I want you very quickly to talk about the work that you do at DAIR, in particular with people working in these data centers and the conditions under which they work around the world, mostly in the Global South.
TIMNIT GEBRU: Yeah, we have a project called the Data Workers’ Inquiry project, and this one is working with data workers. And these are the people that label the data that is used in these models, and which is why we think that they’re superintelligent, right? These people have to painstakingly supply the data and label it. And they work in overwhelmingly exploitative conditions. They are paid maybe $1 an hour. They don’t have any breaks, I mean. And so, we have this project that you can look at where they perform research into their conditions. They organize. We helped them create an organization called the Data Labelers Association in Kenya.
And so, for me, it’s really important to make sure that the voices of the people who are negatively impacted are elevated and they’re the ones telling their stories. And it also shows us that we don’t have to do things. This doesn’t have to be the way. The path we’re on right now was never a preordained path that we had to be on. There were many other possible paths that we could have taken.
AMY GOODMAN: Timnit, I wanted to ask — you grew up in Addis Ababa, in Ethiopia. Your dad was an electrical engineer, your mom an economist. What led you, coming as a refugee to the United States, to get so deeply involved with, be a leading mind on AI right now? We just have a minute.
TIMNIT GEBRU: Yeah, I mean, actually, my book kind of talks about that journey, because I want people to understand that tech does not have to be the way it is right now. But I always loved math and physics, and I also thought it was a refuge from the messy world of war, and which is what made me leave Ethiopian politics. But I later learned that that’s not true, actually. The war and politics that I left is what drives the world of math and science and technology, as well. And if we don’t understand that, we’ll continue to harm our own people and, as engineers and scientists, build terrible and harmful products instead of ones that actually support our communities.
AMY GOODMAN: Timnit Gebru, I want to thank you so much for being with us, founder and executive director of the Distributed Artificial Intelligence Research — that’s DAIR — Institute. We look forward to interviewing you on your forthcoming book, titled Deep Unlearning: The Rise of AI and the Radicalization of a Tech Idealist. She also co-founded Black in AI. She previously served as co-lead of the Ethical AI research team at Google, fired in 2020 for writing a paper warning about the dangers of large language models and raising issues of discrimination in the workplace.
That does it for our show. Special thanks to our crew here in London — Julian Jones, Pablo Delbracio, Grace Garrett, Denis Moynihan, Hana Elias — and the whole team in New York. I’ll be next week in Middlebury, Vermont, and then we’re on to Madison, Wisconsin, and Chicago. I’m Amy Goodman in London, with Nermeen Shaikh in New York.
