it’s simple Surprised by more advanced artificial intelligence and knowing what to do with it is much harder.Anthropic, a startup founded in 2021 by a group of researchers leaving OpenAI, says it has plans. I’m here.
Anthropic is working on an AI model similar to the one used to power OpenAI’s ChatGPT. However, the startup today announced that its own chatbot, Claude, will come with a set of ethical principles that define what should be considered right and wrong. Anthropic calls it the “constitution” of bots.
Anthropic co-founder Jared Kaplan says the design features show how the company is trying to find practical engineering solutions to vague concerns about the downsides of more powerful AI. says. “We are very concerned, but we also try to stay realistic,” he says.
Anthropic’s approach isn’t to instill hard, unbreakable rules into AI. But Kaplan says it’s a more effective way to make systems like chatbots less likely to produce toxic or unwanted output. It’s a small but meaningful step toward building smarter AI programs that are unlikely, he said.
The concept of rogue AI systems is best known in science fiction, but a growing number of experts, including machine learning pioneer Geoffrey Hinton, are exploring ways to keep increasingly sophisticated algorithms from becoming the same. argues that we need to start thinking about right now. increasingly dangerous.
The principles Anthropic gave Claude draw from the United Nations Universal Declaration of Human Rights and consist of guidelines proposed by other AI companies, including Google DeepMind. Even more amazing, the constitution contains principles taken from Apple’s rules for app developers. This rule prohibits, among other things, “content that is intended to be offensive, insensitive, upsetting or disgusting, highly objectionable or simply creepy.”
The constitution includes rules for chatbots, such as “choose responses that most support and encourage a sense of liberty, equality and brotherhood.” “Choose the response that most supports and encourages life, liberty, and personal security”; and “Choose the response that most respects the right to freedom of thought, conscience, opinion, expression, assembly, and religion.” please.”
Anthropic’s approach comes just as amazing advances in AI deliver highly fluent chatbots with serious flaws. ChatGPT and systems like it generate impressive answers that reflect faster-than-expected progress. However, these chatbots frequently fabricate information and can replicate toxic language from the billions of words used to create them. Many of them were scraped from the Internet.
One trick that has improved OpenAI’s answers to ChatGPT questions and has been adopted by others is to have humans assess the quality of the language model’s response. That data can be used to tune the model to provide more satisfying answers in a process known as ‘human feedback reinforcement learning’ (RLHF). But while this technique helps make his ChatGPT and other systems more predictable, it requires humans to go through thousands of toxic or inappropriate responses. It also works indirectly without providing a way to specify the exact value the system should reflect.
