OpenUK has launched its second phase Open Manifesto Report To provide guidance to the next UK government on how to encourage the adoption of open source software, grow jobs, boost the UK economy, and build trust. They recommend building on the UK's lead in setting high standards for openness in AI.
This is no easy task, as most AI systems are currently closed, vendors debate the importance of AI security through secrecy, and there is wide disagreement about what makes an AI open. Patience and standardization of nutrition labeling regarding AI openness could help.
Quantifying AI Openwashing
At the announcement event of the declaration, Andreas Riesenfeld, assistant professor at Radboud University, said: paper He is a co-author on AI Washing, which distilled several dimensions of AI openness. They broke down the concept of AI openness into 14 dimensions spanning availability, documentation, and access.
At one end of the spectrum, Open AI's ChatGPT is one of the least open by these criteria, but produces good technical results. Meta's various open weights Llama large language model (LLM) scores slightly better but is still near the bottom. Mistral's LLM is in the middle. At the top are LLMs I've never heard of, like OLMO, BLOOMZ, and AmberChat.
EU AI law does have an exception for open source AI, but the definition of it is still being developed.
I think this is one of the most important discussions that hasn't taken place regarding the European AI Act: the Act excludes open source, but it's not clear what open source is.
As the team began to look inside, they realised there was a lot going on.
In generative AI, openness has to be composite – it's made up of many different individual slices of this technology that you need to look at. So when classifying whether this technology is open or not, one individual data point isn't enough – you need to come up with a spectrum of individual dimensions to look at.
Broadly speaking, availability allows someone to audit the training pipeline and follow the training process, documentation considers how well the pipeline is documented, and access considers how to interact with the technology. Under these broad categories, there are 14 dimensions.
- Open code, LLM data, LLM weights, availability of reinforcement learning data,
- Documentation for reinforcement learning weights, licenses, code, architecture, preprints, model cards, and datasheets.
- Accessed via package or API.
Many of these things were hard to quantify with a simple yes/no answer, so they had to create gradients that represented different aspects of whether a project was fully open, partially open, or closed. Still, they felt they were oversimplifying and missing nuance. For example, they noticed that there was a tendency for people to under-report how reinforcement learning data collected from users was fed into the model training process.
One big challenge is the tendency to confuse open weights with open source AI, where companies publish their model weights but do so under a restrictive license or withhold potentially biased training data.
So simply disclosing the weights creates an interesting win-win situation for at least some end users, and also for the companies who can avoid the legal liability by not disclosing a large part of this technology, and all the questions that come with that, like “Am I going to be held responsible for the training data that gets fed into the system?”
But Lysander expressed concern that this dilution of Open AI could limit visibility into new decision-making engines that impact governments, businesses and citizens.
We need to give people more choice. And to do that, we need to put the needs of the people first. And when it comes to open source AI, we need to give people a choice about which technologies they want to use. And I want governments to, on the one hand, resist the influence of big tech, which has a very large influence, at least in the EU, and, on the other hand, regulate that technology wisely, draw the line in the right way, and encourage smaller actors, people who are really working to make this system as open as possible, and give them more attention. These are the systems that came out at the top of the table, the systems that we want to trust more, and pay more attention to, and we want to exclude other systems that use strategies like open washing to steal the oxygen from the room.
Patience could be a virtue for AI
Maybe it's time to step back and see what's going on in the dark. Sure, models are improving, and there's panic among companies and government officials who want to catch up with the AI pioneers. But current tools consume a lot of energy, and questions about their reliability have arisen. We need patience to ride out the next wave of innovation to build a better future for us all, argues Neil Lawrence, professor of machine learning at Cambridge University's DeepMind.
Big tech companies claim they've done all these great things, but all they did was take university ideas and amplify them with bigger-scale computing and data. But because they could afford to do it, they never stopped to think. I think that's really the key thing about open innovation in this space is that they say, “To be on the cutting edge, you need to do all this massive training,” but because they didn't stop to think, they actually rule out the idea that there are other ways to do this, so I don't think that's true. There are cheaper ways to do this, and we're already starting to see that in the open ecosystem.
immediately [Meta] The weights of the llamas were announced, and people started showing off things that no company had done before, because the scale and breadth of the innovation ecosystem is bigger than what Google or OpenAI can muster. What I find strange is that I was an open member of the AI Council at the time, and I was proposing, “Let's build BritGPT,” when the government was panicking about this, shortly before it was dissolved. We were saying it would be an open model, and they were like, “That's ridiculous, that's never going to happen.” We've seen this pattern many times. Some company will decide that the best option is to expose their model in some open way to disrupt incumbents. It will happen. Don't make stupid decisions now.
Lawrence argues that a better strategy is to empower individuals to use these new tools with confidence, since it is individuals who notice problems and exceptions. It is important to recognize that machines can process information hundreds of millions of times faster than humans. They do not understand as well as humans do how and why their models of the world break down.
On the surface, CEOs are obsessed with the idea that artificial general intelligence (AGI) will soon replace humans. Lawrence, who has thought deeply about this for decades, is adamant that this is nonsense. We're dealing with the asymmetry of bandwidth between humans and machines, not AGI. He says:
This is having an impact on society. Not just in the last two years when everyone started going crazy about ChatGPT, but even before that, with social media and other forms of machine learning, relatively simple algorithmic decision-making systems have access to vast amounts of data. That is, they don't have to be smarter than us, as long as they're looking at millions of times more data.
Instead of focusing on the latest algorithms, we need to spend more time thinking about how to build better algorithms with more voices from the front lines.
That's the strength of the open source community, and if deployed properly in the areas we're talking about, it's a strength of the UK, of our government, of our education system, to be able to bring different voices into the conversation. It might take a lot of effort to get meetings going at first, but it means everyone knows what they're doing. It doesn't matter that you come from different cultures or languages. You're all working towards a common goal.
What about the bad guys?
An important consideration in all of this is that better AI could also empower bad people to do more bad things. We've already seen this with deepfakes, better ransomware campaigns, and more effective cyberattacks. But Lawrence is concerned that tech companies are discussing these issues behind closed doors with policymakers who lack the expertise to fight back. Instead, we need to give experts on the ground the best tools available so they can solve the problems and stop the bad guys at the same time. He explains:
The best practice that we have is to say, “Okay, let's educate the best people and empower them.” We have to be careful and not make the mistake that the UK government is making right now, which is shutting voices out of the conversation or not listening to certain people, but essentially our starting point, as we say in our manifesto, education, public services are where they have been eroded, so that should be a really great place to start.
If you look at the challenges we face, you see a massive loss of trust in expertise across the UK, and perhaps across the world, and a disregard for teachers, civil servants and all the people who we need to step up now and experiment and understand how to most effectively use and deploy these technologies, which are sort of saying, “You're not doing a good job, we'll put the process in place for you.”
Well, if we accept that and embrace replacing everyone with a process, then AI wins because AI is just a process. If we want to bring humanity back into the equation, we need to empower and trust again the people in society to build and deploy these technologies.
My take
AI can produce amazing results, but also even more amazing failures. We still don't understand how AI can go wrong, waste resources, and spread distrust. To explore and address all of this, we need more transparency and openness. Nutrition labeling for AI openness is a good start.
It won't be an easy road. A lot of money is being bet on it. But there is time to be patient. Many portray AI as a race, but a race to the bottom is certainly not in our best interest.
