AI Safety: Where’s the Puck Headed?

RSA Conference 2024 · South Stage Keynote

Overview

This panel discussion at RSAC 2024, moderated by Romka Siva, delved into the increasingly urgent and complex topic of AI safety, moving beyond the sensationalized "killer robot" narratives to explore its multifaceted implications. The session brought together a distinguished group of experts from AI research, cybersecurity, and hardware security to provide a comprehensive view of the challenges and considerations surrounding advanced artificial intelligence. The central theme explored was the rapid evolution of AI capabilities and the corresponding shift in how experts perceive the risks, moving from outright dismissal just a year prior to acknowledging a complicated and potentially existential threat.

Watch on YouTube

Visual summary for AI Safety: Where’s the Puck Headed?
Visual summary for AI Safety: Where’s the Puck Headed?

Key moments

  1. 0:00 Moderator's opening poll and changed perspective on AI risk
  2. 2:00 Academic researchers' growing concern about AI existential risk
  3. 2:20 Introduction of the expert panel for AI safety discussion
  4. 4:28 Dan Hendrickson defines AI safety and addresses 'killer robots'
  5. 5:35 AI safety as a broader sociotechnical risk management problem
  6. 6:30 Discussion on why the 'killer robot' narrative captures public attention

AI Safety: Where’s the Puck Headed?

Speakers: Romka Siva (Moderator), Dan Hendrickson, Bruce Schneier, Theodore Atkins, Daniel

Conference: RSAC 2024

YouTube: https://www.youtube.com/watch?v=m1yL78zAhT8

Overview

This panel discussion at RSAC 2024, moderated by Romka Siva, delved into the increasingly urgent and complex topic of AI safety, moving beyond the sensationalized "killer robot" narratives to explore its multifaceted implications. The session brought together a distinguished group of experts from AI research, cybersecurity, and hardware security to provide a comprehensive view of the challenges and considerations surrounding advanced artificial intelligence. The central theme explored was the rapid evolution of AI capabilities and the corresponding shift in how experts perceive the risks, moving from outright dismissal just a year prior to acknowledging a complicated and potentially existential threat.

The talk highlighted a significant change in sentiment, even among academic researchers, with a notable percentage now considering AI an existential risk. This shift underscores the critical need for a deeper understanding of AI safety, not merely as a technical problem but as a broad sociotechnical challenge. The discussion aimed to unpack what AI safety truly means, why it has garnered such intense attention, and how the security community, along with broader society, should prepare for its ongoing and future impacts.

The panel emphasized that the exponential trajectory of AI's development, exemplified by the leap from GPT-3 to GPT-4 in just two years, necessitates a proactive and comprehensive approach to risk management. This involves addressing both immediate harms and anticipated tail risks that could manifest across various timeframes. The insights shared are crucial for security professionals, policymakers, and AI developers alike, urging a move beyond simplistic solutions to embrace a holistic strategy for securing our future in an AI-driven world.

Background

▶ Watch: Moderator's opening poll and changed perspective on AI risk (0:00)

The conversation around AI safety has undergone a dramatic transformation in a remarkably short period. As moderator Romka Siva recounted, merely a year prior at RSA, the idea of "killer AI robots" was met with laughter and eye-rolls. Fast forward to RSAC 2024, and the sentiment has shifted profoundly, with Siva personally moving from a "hell no" to an "I think it's complicated" stance on the risks posed by AI systems. This personal evolution mirrors a broader trend within the scientific and security communities.

A striking data point presented was that 5% of academic researchers surveyed at a top neural information processing systems (NeurIPS) conference now believe that AI could pose an existential risk. This statistic underscores that concerns about AI safety are no longer confined to "the crazy and the influence peddlers" but are becoming a serious consideration for mainstream experts. The panel aimed to dissect this evolving perception, acknowledging that while the "killer robot" narrative has captured public and political imagination—even influencing executive orders after viewings of films like Mission: Impossible—the actual risks are far more nuanced and pervasive.

Dan Hendrickson, founder of the Center for AI Safety, articulated that the heightened attention stems from the exponential trajectory of AI development. He referenced the rapid progression from GPT-3 in 2020 to GPT-4 just two years later, noting that as the computational power invested in these systems grows by orders of magnitude, their capabilities become undeniable. This rapid advancement forces a reckoning with the technology's immense power and the potential for both beneficial and harmful applications. Hendrickson further clarified that AI safety is not solely an "AI thing" but rather a broader risk management problem, requiring a sociotechnical approach that extends beyond mere technical fixes to encompass human organizations, corporate incentives, and the secure allocation of foundational hardware. This perspective draws a parallel to the evolution of information security, which moved beyond thinking a "really good firewall" was sufficient to recognizing the need for complex, systemic solutions.

Key Findings

▶ Watch: Introduction of the expert panel for AI safety discussion (2:20)

The panel discussion crystallized several key findings regarding the current state and future trajectory of AI safety. Foremost among these is the consensus that AI safety is fundamentally a sociotechnical problem, not merely a technical one. Dan Hendrickson explicitly countered the notion that AI safety can be reduced to solving "jailbreaks for large language models," drawing an analogy to early information security where a "really good firewall" was once considered the panacea. This narrow view, he argued, fails to account for the intricate interplay of human behavior, organizational structures, economic incentives, and the broader societal context in which AI operates.

Another critical finding is the exponential growth in AI capabilities driven by continuous increases in computational power. The leap from GPT-3 in 2020 to GPT-4 in 2022 serves as a stark example of how rapidly AI systems are evolving. This rapid advancement is not just about incremental improvements but rather a qualitative shift in what AI can achieve, making the discussion of safety more urgent and less speculative. The panel underlined that as these capabilities continue to grow by "orders of magnitude," more people are compelled to recognize AI's profound power and inherent risks.

The discussion also delineated the spectrum of risks associated with AI, moving beyond the often-sensationalized "killer robot" scenarios, which the panelists generally considered "somewhat farther away." Instead, the focus was placed on more immediate and tangible threats, such as malicious use of AI. This includes ongoing harms that are already manifesting, as well as anticipated tail risks that could emerge in the near future (two to five years). These risks necessitate a proactive and comprehensive risk management framework, one that can adapt to evolving threats and consider both technical vulnerabilities and the human element. The panel stressed that securing AI systems requires attention to the entire ecosystem, from the algorithms themselves to the underlying hardware—like Nvidia's A100 GPUs—and ensuring these resources are not exploited by malicious actors.

Technical Deep Dive

▶ Watch: Dan Hendrickson defines AI safety and addresses 'killer robots' (4:28)

While the panel discussion was high-level, it provided a framework for understanding the technical dimensions of AI safety within a broader sociotechnical context. The core technical challenges, as interpreted from the discussion, revolve around ensuring the robustness, reliability, and security of AI systems throughout their lifecycle and deployment.

Dan Hendrickson's analogy between fixing "jailbreaks for large language models" and relying on a "really good firewall" in information security is particularly instructive. It suggests that while specific technical vulnerabilities like prompt injection or adversarial attacks on AI models are important to address, they represent only a fraction of the overall technical security surface. A true technical deep dive into AI safety must consider:

  1. Model Robustness and Alignment: This involves developing techniques to ensure AI models behave as intended, resist adversarial manipulation, and align with human values and safety objectives. The concept of "jailbreaks" falls into this category, representing instances where a model can be coaxed into bypassing its safety mechanisms. Technical solutions here might involve advanced reinforcement learning from human feedback (RLHF), red-teaming methodologies to discover vulnerabilities, and developing interpretable AI (XAI) techniques to understand model decision-making. However, the panel's point is that such technical fixes, in isolation, are insufficient without considering the broader system.
  1. Secure Compute and Hardware Infrastructure: Daniel, representing Nvidia, highlighted the critical role of hardware in AI safety. His involvement in "securing all the A100s" points to the necessity of a secure supply chain and robust hardware security for the foundational components of AI. This includes:
  • Secure boot processes for AI accelerators.
  • Hardware-level isolation and memory protection to prevent malicious code execution or data exfiltration.
  • Authenticity and integrity verification of AI chips and modules to prevent counterfeiting or tampering.
  • Trusted execution environments to protect sensitive AI models and data during inference and training.
  • Supply chain security measures to ensure that hardware components are secure from manufacturing through deployment. This prevents malicious actors from introducing backdoors or vulnerabilities at a fundamental level, which could then compromise the AI systems built upon them.
  1. Benchmarking and Evaluation: Dan Hendrickson's foundational contributions to "benchmarking AI systems" are directly relevant to AI safety. Robust benchmarking is a technical discipline that provides standardized methods for evaluating AI capabilities, performance, and crucially, safety. This involves:
  • Developing metrics and datasets to assess model robustness against adversarial attacks.
  • Creating evaluations for harmful outputs, biases, and unintended behaviors.
  • Establishing transparency and reproducibility standards for AI safety research.
  • Continuous monitoring and post-deployment validation of AI systems to detect emergent risks.
  1. Sociotechnical Integration: The panel implicitly argued for technical solutions that are designed with human and organizational factors in mind. This means technical safeguards should not operate in a vacuum but be integrated into a larger security framework that considers corporate incentives, ethical guidelines, and user interaction. For instance, designing AI systems with human-in-the-loop mechanisms for critical decisions or incorporating privacy-preserving AI techniques are technical approaches that directly address sociotechnical concerns. The challenge lies in developing technical architectures that are inherently resilient to both technical exploits and human misuse or organizational pressures.

In essence, the "technical deep dive" within the context of this panel implies a comprehensive approach to securing the entire AI stack—from the silicon up through the model architecture and its deployment environment—while acknowledging that purely technical "patches" are not sufficient without a holistic understanding of the system's human and organizational dimensions.

Demo / Proof of Concept

▶ Watch: AI safety as a broader sociotechnical risk management problem (5:35)

This panel discussion did not include a live demonstration or proof of concept. As a panel, the session focused on a high-level discussion of concepts, risks, and strategic approaches to AI safety rather than showcasing specific technical implementations or exploits.

Defensive Implications

▶ Watch: Discussion on why the 'killer robot' narrative captures public attention (6:30)

The panel's insights offer crucial guidance for security defenders navigating the rapidly evolving landscape of AI. The primary defensive implication is a mandatory shift from a purely technical, reactive security posture to a comprehensive sociotechnical risk management framework for AI. Defenders must recognize that securing AI extends far beyond traditional cybersecurity paradigms.

  1. Embrace a Sociotechnical Perspective: The most significant takeaway is that "information security equals a really good firewall" is an outdated concept for AI. Defenders must consider the entire ecosystem: not just the AI models themselves, but also the human organizations developing and deploying them, the corporate incentives driving their creation, and the societal impact. This means engaging with legal, ethical, and policy teams, understanding organizational culture, and influencing product development from a safety-first perspective.
  1. Proactive Risk Management for Tail Risks and Malicious Use: Security teams need to move beyond reacting to known vulnerabilities. The exponential growth of AI capabilities (e.g., GPT-3 to GPT-4) means new and unprecedented risks, including malicious use, will emerge rapidly. Defenders should establish processes for anticipating tail risks—low-probability, high-impact events—and developing contingency plans. This includes continuous threat intelligence gathering specific to AI, red-teaming AI systems, and scenario planning for novel attack vectors or system failures.
  1. Secure the AI Supply Chain, from Silicon to Software: Daniel's role at Nvidia underscores the criticality of hardware security. Defenders must ensure the integrity and security of the foundational compute infrastructure, such as Nvidia A100 GPUs, that powers AI. This involves:
  • Hardware Root of Trust: Verifying the authenticity and integrity of AI accelerators and components.
  • Supply Chain Transparency: Working with hardware vendors to understand and mitigate risks throughout the manufacturing and distribution process.
  • Secure Deployment: Implementing robust physical and logical security for AI data centers and edge devices.
  • Trusted Execution Environments: Leveraging hardware-based security features to protect AI models and data during training and inference.
  1. Invest in AI Benchmarking and Robustness: Leveraging expertise like Dan Hendrickson's in benchmarking, defenders should push for rigorous evaluation of AI systems. This includes:
  • Adversarial Robustness Testing: Actively seeking out and mitigating vulnerabilities to adversarial attacks and "jailbreaks."
  • Bias Detection and Mitigation: Implementing technical and process controls to identify and reduce harmful biases in AI outputs.
  • Performance and Safety Metrics: Developing and continuously monitoring metrics that go beyond accuracy to include safety, ethical compliance, and resilience.
  1. Foster Cross-Disciplinary Collaboration: Given the sociotechnical nature of AI safety, defenders cannot operate in silos. Collaboration with AI researchers, ethicists, legal experts, policymakers, and even social scientists is essential. Security professionals need to translate technical risks into broader organizational and societal impacts and work with diverse stakeholders to develop holistic solutions. This includes participating in industry standards bodies and contributing to best practices for AI development and deployment.

By adopting these defensive implications, security professionals can contribute to building a more secure and resilient future in an era increasingly shaped by powerful AI.

Key Takeaways

  • AI safety is a complex sociotechnical problem, not just a technical one. Addressing it requires considering human organizations, corporate incentives, and societal impacts, not just technical fixes like "jailbreak" solutions for LLMs.
  • AI capabilities are on an exponential trajectory. The rapid evolution from GPT-3 to GPT-4 demonstrates a significant and continuous increase in AI power, necessitating urgent and proactive safety measures.
  • Risks range from immediate malicious use to long-term "tail risks." While "killer robots" are distant, current concerns include misuse of powerful AI and anticipated, potentially existential, risks within the next 2-5 years.
  • Hardware and compute security are foundational to AI safety. Securing the underlying infrastructure, such as Nvidia's A100 GPUs, is critical to prevent malicious actors from compromising AI systems at their core.
  • Rigorous benchmarking and evaluation are essential for understanding and mitigating AI risks. Standardized methods for assessing AI capabilities, robustness, and safety are crucial for responsible development and deployment.
  • Defenders must adopt a holistic, proactive risk management approach. This involves moving beyond reactive security measures to anticipate threats, secure the entire AI supply chain, and collaborate across disciplines.

About the Speaker(s)

  • Romka Siva (Moderator): The moderator for this panel, Romka Siva, initiated the discussion by highlighting the significant shift in perception regarding AI risks, personally moving from dismissing concerns to acknowledging their complexity.
  • Dan Hendrickson: Founder for the Center for AI Safety, Dan Hendrickson is recognized on the Times 100 AI list and by ML researchers for his fundamental contributions to the field, including seminal papers on benchmarking AI systems and making robots. He is considered "the voice of AI and the Oracle, the embodiment of it."
  • Bruce Schneier: A prominent figure in the security community, Bruce Schneier is a New York Times best-selling author with 14 books, including Hacker's Mind, and is currently working on a new book about AI and democracy. He was brought to the panel to unpack the social ramifications of AI safety.
  • Theodore Atkins: As VP of Security at Google, Theodore Atkins famously described her job as "to keep the hackers away." She quite literally started security at Google, predating features like Gmail and unsubscribe buttons, making her a "wise internet elder" on the panel.
  • Daniel: A long-time veteran in the compute field, Daniel started his career as one of the first interns in India for Nvidia and is now the VP of Product Security at Nvidia, responsible for securing critical AI hardware like the A100 GPUs.

All talks from RSA Conference 2024