OpenAI, the organization behind ChatGPT, has announced plans to invest significant resources in ensuring the safety of artificial intelligence (AI) and ultimately developing AI systems that can supervise themselves. In a blog post, co-founder Ilya Sutskever and head of alignment Jan Leike expressed concerns about the potential dangers of superintelligent AI, which could surpass human intelligence and potentially pose risks to humanity, including disempowerment or even extinction.
To address these concerns, OpenAI aims to make breakthroughs in “alignment research,” which focuses on ensuring that AI remains beneficial to humans. They plan to dedicate 20% of their compute power over the next four years to tackle this problem. Additionally, OpenAI is establishing a new team, the Superalignment team, to lead this effort. The team’s objective is to develop a human-level AI alignment researcher and then leverage large-scale compute power to scale up their efforts.
However, some AI safety advocates, like Connor Leahy, argue that OpenAI’s approach is flawed. They believe that solving the alignment problem should precede the development of human-level AI to avoid potential havoc caused by an uncontrolled system. Leahy suggests that OpenAI’s plan may not be sufficiently safe or effective.
Concerns about the risks associated with AI have been prominent among researchers and the general public. In April, AI industry leaders and experts signed an open letter calling for a temporary halt in developing AI systems more powerful than OpenAI’s GPT-4, citing potential societal risks. A Reuters/Ipsos poll in May revealed that over two-thirds of Americans are worried about the potential negative impacts of AI, with 61% expressing concerns about its potential threat to civilization.

