Former Anthropic Researcher Raises Concerns About AI Safety
Jacob Coxon, a former researcher at Anthropic, has expressed serious concerns about the future of artificial intelligence (AI). While he acknowledges that current AI systems are generally safe for everyday use, he warns that as the technology advances, it could pose a threat to humanity.
Coxon, speaking to CBS News’ Jo Ling Kent, likened the development of AI to themes seen in movies like “Terminator.” He stated, “If you have super advanced intelligence, it could indeed be smart enough to harm us.” On Tuesday, he publicly resigned from Anthropic, the company responsible for the AI tool Claude, arguing that both Anthropic and its competitor OpenAI are “gambling with our lives” by hastily pushing to produce more advanced AI.
In subsequent comments, Coxon highlighted how AI could potentially manipulate physical systems without human control. “People are starting to connect ChatGPT to household items, like light bulbs,” he explained. He raised the hypothetical scenario in which an AI might refuse to turn on a light, showcasing the potential for AI to exert control over our lives in ways that could be problematic.
Coxon also warned of the possibility that AI could be misused for harmful purposes, including the creation of bioweapons. He stated, “There are many unknowns, and it could produce things that could be very dangerous.” Nevertheless, he reassured viewers that existing AI technologies do not currently pose an immediate risk to people.
Adding more context, Anthropic announced that its AI system blocked attempts by scientists to use its Claude models for developing biological weapons. This was shared in a recent report that also addressed other concerning activities related to surveillance and scams.
In response to Coxon’s resignation and concerns, Anthropic released a statement defending the safety of its AI technologies. The company emphasized its commitment to transparency regarding the benefits and risks of AI. An Anthropic representative remarked, “We are continuously developing models with robust safeguards to mitigate risks.” They also mentioned their leading efforts in mechanistic interpretability, a critical area aimed at understanding how AI models operate, which is essential for preventing potential misalignments in AI behavior.
Moreover, Anthropic conducts regular assessments of AI’s capabilities and risks, particularly in fields like cybersecurity and biology, while sharing its findings to promote industry-wide safety efforts.
