Reducing AI’s Blast Radius: How to Prevent Your First AI Breach

RSA Conference 2024 · West Stage Keynote

Overview

In an era where Generative AI (Gen AI) has rapidly transitioned from a niche technology to a mainstream enterprise tool, the question of data security has become paramount. Matt Radolec, Vice President of Incident Response and Cloud Operations at Varonis, delivers a compelling and urgent message at RSAC 2024: AI represents both the biggest opportunity and the biggest threat to organizational data. His talk, "Reducing AI’s Blast Radius: How to Prevent Your First AI Breach," meticulously dissects the inherent risks of AI adoption, particularly concerning sensitive, personal, and regulated data. Radolec emphasizes that while AI promises unprecedented productivity and innovation, its uncontrolled access to vast datasets can lead to catastrophic breaches, data corruption, and significant reputational and financial damage.

Watch on YouTube

Visual summary for Reducing AI’s Blast Radius: How to Prevent Your First AI Breach
Visual summary for Reducing AI’s Blast Radius: How to Prevent Your First AI Breach

Key moments

  1. 0:00 AI's Dual Nature: Opportunity and Threat to Data
  2. 2:40 Why Data is the True Impact of a Breach
  3. 3:20 Case Study: Protecting Alzheimer's Research Data from AI
  4. 4:50 Case Study: Safeguarding Critical City Infrastructure Data
  5. 6:15 Jensen Huang: AI is Fundamentally a Data Problem
  6. 8:00 Inspiring AI Use: Chat GPT Saves a Dog's Life
  7. 9:50 Security's Big Moment: Empowering Business with Safe AI

Reducing AI’s Blast Radius: How to Prevent Your First AI Breach

Speakers: Matt Radolec, Vice President, Incident Response and Cloud Operations, Varonis

Conference: RSAC 2024

YouTube: https://www.youtube.com/watch?v=1SxiO7R9kVU

Overview

In an era where Generative AI (Gen AI) has rapidly transitioned from a niche technology to a mainstream enterprise tool, the question of data security has become paramount. Matt Radolec, Vice President of Incident Response and Cloud Operations at Varonis, delivers a compelling and urgent message at RSAC 2024: AI represents both the biggest opportunity and the biggest threat to organizational data. His talk, "Reducing AI’s Blast Radius: How to Prevent Your First AI Breach," meticulously dissects the inherent risks of AI adoption, particularly concerning sensitive, personal, and regulated data. Radolec emphasizes that while AI promises unprecedented productivity and innovation, its uncontrolled access to vast datasets can lead to catastrophic breaches, data corruption, and significant reputational and financial damage.

Radolec’s presentation is a critical call to action for security professionals, urging them to shift their focus from traditional threat vectors to the core asset: data. Drawing on his extensive 15-year career in cybersecurity and his current role leading incident response at Varonis, he provides real-world examples of near-misses and actual breaches, illustrating how AI magnifies existing vulnerabilities in data access and control. The talk is designed to equip attendees with actionable insights, enabling them to proactively secure their data environments against the unique challenges posed by AI, ultimately aiming to reduce AI's "blast radius" and safeguard organizational secrets.

This article delves into Radolec's key arguments, exploring the technical intricacies of AI co-pilots, the critical role of Graph Connectors in Microsoft 365, and the pervasive problem of lax access control. It outlines the defensive strategies necessary to navigate the choppy seas of AI adoption, advocating for a data-centric security posture that leverages automation to combat AI-driven threats. By understanding the probabilities and focusing on controllable elements like security policies and data access, organizations can embrace the benefits of AI while mitigating the profound risks to their most valuable asset.

Background

▶ Watch: AI's Dual Nature: Opportunity and Threat to Data (0:00)

The modern enterprise, regardless of industry—be it healthcare, energy, finance, legal, manufacturing, higher education, or government—is fundamentally built upon and driven by data. Despite this universal dependency, Matt Radolec argues that the cybersecurity industry often misplaces its focus during breach discussions. Instead of prioritizing the "crown jewels" – the data itself – conversations typically revolve around actor names, MITRE ATT&CK techniques, malware variants, and CVEs. This misdirection, Radolec contends, overlooks the core truth: "Data is where the damage happens. Data is where we feel the pain of an attack." The advent of Generative AI only intensifies this pain, transforming existing data vulnerabilities into potential catastrophes.

The problem, Radolec explains, is not new, but AI acts as a potent accelerant. He recounts Varonis incident response investigations that underscore the critical need for data-centric security. One example involves a leading Alzheimer's research organization, which recognized the devastating potential of research data being stolen or, worse, corrupted and then fed into an AI for training. Such an event could have "grave consequences" for clinical research and patient care. Similarly, a major metropolitan utility provider, despite being under constant attack from sophisticated groups like Volt Typhoon, avoided significant disruption to critical infrastructure (water, traffic, electrical grid, voting systems) because of 24/7 data monitoring. These cases highlight that attackers are already targeting data, and AI introduces a new, powerful vector for both exfiltration and manipulation.

Echoing NVIDIA CEO Jensen Huang's assertion that "AI is a data problem," Radolec frames AI as the dominant theme of current cybersecurity discourse. The industry is grappling with fundamental questions: how to harness AI for good, how to thwart "evil AI," and crucially, how to prevent the inevitable "first AI data breach." The widespread adoption of AI, particularly AI co-pilots from tech giants like Google, Salesforce, and Microsoft, means "Pandora's Box is open." While AI offers immense potential for saving lives, boosting productivity, and driving innovation, realizing these benefits safely requires a profound shift in security strategy—one that acknowledges and controls the inherent risks associated with AI's insatiable appetite for data. The core challenge lies in understanding what can be controlled (security policies, access controls, incident response plans) versus what cannot (deepfakes, AI-generated phishing), and then focusing defensive efforts on the former.

Key Findings

▶ Watch: Case Study: Protecting Alzheimer's Research Data from AI (3:20)

Matt Radolec's talk distills several critical findings regarding the intersection of AI and data security, painting a stark picture of the challenges organizations face and the urgent need for a paradigm shift in defensive strategies.

Firstly, the proliferation of AI co-pilots (such as Google Gemini, Salesforce Einstein, and Microsoft Copilot) fundamentally changes the data security landscape. These co-pilots operate by combining an organization's proprietary business data with the vast general knowledge contained within a Large Language Model (LLM). This fusion means that if a co-pilot has access to data, it effectively grants the LLM, and by extension, any user interacting with that co-pilot, access to that data. Radolec emphatically states, "If you don't get access control right here, you will fail. You will have a data breach by AI."

Secondly, a major vulnerability point specifically identified is the implementation of Graph Connectors within Microsoft 365 Copilot. The Microsoft Graph acts as the access map for Copilot to business data. Graph Connectors allow organizations to extend Copilot's reach to virtually any external data source. The critical flaw lies in how access control is managed for these connectors: it is not "pass-through" like for native M365 data. Instead, developers must explicitly set the Access Control List (ACL) by hand. Radolec reveals a common, dangerous practice: "I'll bet you, if you have a graph connector, your permissions are set to everyone, just like I'm showing here." This default or easiest configuration becomes a wide-open door for data exfiltration, as demonstrated by a Varonis incident response case where a former employee leveraged lax access control on a database accessible via a co-pilot connector to steal customer account data for a competitor.

Thirdly, Radolec highlights that AI magnifies pre-existing, widespread issues with data access governance. Despite the industry's embrace of concepts like granular access control and Zero Trust, Varonis assessments consistently reveal a "huge and only going up" blast radius. The statistics are alarming: the average organization has at least 17 million files open to all employees and manages 40 million plus unique access control lists. Yet, users typically utilize only 1% of the permissions they are assigned. This extreme over-permissioning creates a vast attack surface that AI co-pilots can readily exploit, turning latent vulnerabilities into active threats.

Finally, the talk underscores the necessity of shifting security focus. Radolec likens endpoints to "ATMs" – they hold some money and warrant monitoring, but they are not the "crown jewels." The true crown jewels – the organization's "good, juicy, rich data" – reside in the "vaults" (databases, cloud object stores, file shares). Securing these vaults requires monitoring every transaction, detecting anomalies in data usage (akin to credit card fraud detection), and policing every prompt used with AI co-pilots. Furthermore, Radolec stresses that humans alone cannot win this battle; automation and AI are indispensable for combating AI threats effectively.

Technical Deep Dive

▶ Watch: Case Study: Safeguarding Critical City Infrastructure Data (4:50)

The technical core of Matt Radolec's presentation revolves around the mechanics of Generative AI (Gen AI), specifically how AI co-pilots interact with organizational data, and the critical security implications arising from these interactions. The speaker focuses heavily on the architecture and vulnerabilities associated with Microsoft's ecosystem, particularly Microsoft 365 Copilot and its reliance on the Microsoft Graph.

At its fundamental level, a Gen AI co-pilot, exemplified by offerings like Google Gemini, Salesforce Einstein, and Microsoft Copilot, functions by merging two distinct knowledge bases. The first is the Large Language Model (LLM) itself, which possesses a vast, general understanding of the world, derived from extensive training on internet data. The second is an organization's proprietary business data, encompassing everything from internal documents, emails, customer records, financial reports, and intellectual property. When a user inputs a prompt into a co-pilot, the system doesn't just consult the LLM; it dynamically combines the LLM's general knowledge with the specific context and information found within the organization's data environment.

This integration is where the security paradigm fundamentally shifts. The co-pilot, by design, has access to the same data that the user does, often without the user fully realizing the breadth of that access. This means that if an organization has lax data access controls, the co-pilot effectively inherits and magnifies those weaknesses. Any data that is over-permissioned or improperly secured becomes immediately discoverable and potentially exfiltratable through the co-pilot interface.

Radolec hones in on Microsoft 365 Copilot as a prime example of this challenge. The access map for Copilot to an organization's business data is the Microsoft Graph. The Graph is a unified API gateway that connects various Microsoft 365 services and data. While Copilot can access native M365 data (like SharePoint, OneDrive, Exchange) based on existing user permissions, its true power, and significant risk, lies in its ability to connect to external data sources via Graph Connectors.

Graph Connectors are mechanisms that allow organizations to ingest data from non-Microsoft 365 sources—such as on-premises file shares, third-party cloud storage, line-of-business applications, or custom databases—into the Microsoft Graph index. This makes the data discoverable and usable by Copilot. The critical technical detail highlighted by Radolec is how Access Control Lists (ACLs) are managed for these connectors. Unlike native M365 data, where permissions might "pass through" or be inherited, Graph Connectors require developers to explicitly set ACLs by hand. This manual process is prone to human error and, more alarmingly, often defaults to the path of least resistance. Radolec strongly suspects that many organizations, in an effort to quickly enable functionality, set these permissions to "everyone."

This "everyone" setting for a Graph Connector means that any user with access to the M365 Copilot could potentially query and retrieve sensitive data from the connected external source, regardless of their actual need or authorization for that specific data. Radolec illustrates this with a real-world incident response case: "An organization called us in to determine how a competitor got so much information from a particular database of theirs... The evidence was clear. Like a lot of companies, access control was pretty lax... They had widely adopted co-pilots and their connectors... one of their employees, former employees, I might add, was the first to realize this mishap. They grabbed the database of accounts, including the products they owned and how much they spend, and they left and took it to a competitor." This scenario perfectly exemplifies how a seemingly innocuous "everyone" permission on a Graph Connector, or even a database accessible by such a connector, can lead to a significant data breach when exploited by an AI co-pilot or an insider.

The underlying technical challenge, therefore, is not just about AI itself, but about the amplification of existing, systemic weaknesses in data governance. Organizations often struggle with granular access control and implementing true Zero Trust principles. The statistics cited by Radolec—an average of 17 million files open to all employees and 40 million+ unique ACLs to manage, with users only leveraging 1% of their assigned permissions—underscore this deep-seated problem. AI co-pilots, by providing an intuitive and powerful interface to query and synthesize this vast, often over-permissioned data, bypass traditional endpoint or perimeter security measures, making the data itself the primary attack surface. The technical defense, therefore, must shift to continuous monitoring of data transactions, anomaly detection at the data layer, and rigorous policing of AI prompts to prevent the weaponization of internal data by AI.

Demo / Proof of Concept

▶ Watch: Inspiring AI Use: Chat GPT Saves a Dog's Life (8:00)

While Matt Radolec's talk did not feature a live, interactive technical demonstration in the traditional sense, he effectively presented several compelling "proofs of concept" through real-world scenarios and visual aids that underscored his key arguments.

The most direct visual "demo" mentioned in the transcript relates to the critical vulnerability of Graph Connectors in Microsoft 365 Copilot. Radolec states, "I'll bet you, if you have a graph connector, your permissions are set to everyone, just like I'm showing here." This strongly implies that a screenshot or slide depicting a Graph Connector's access control settings, likely showing "everyone" as the default or common configuration, was displayed to the audience. This visual served as a stark illustration of how easily a critical security control can be misconfigured, creating a wide-open avenue for data exposure.

Beyond this implied visual, the core of Radolec's "proof of concept" lies in the concrete incident response investigations conducted by Varonis. He shared two primary examples:

  1. The Alzheimer's Research Organization: This case served as a hypothetical but highly plausible scenario of the "grave consequences" if critical clinical research and medical data were altered and then fed into an AI for training. While a breach was averted due to Varonis's 24/7 monitoring, it highlighted the specific risk of data corruption amplified by AI, which could lead to flawed research outcomes or patient harm. The implication is that without such monitoring, an AI could inadvertently or maliciously process compromised data, leading to disastrous, AI-driven outcomes.
  1. The Major Metropolitan Utility Provider: This example detailed repeated attempts by attackers (including sophisticated groups like Volt Typhoon) to infiltrate critical infrastructure systems (water, traffic, electrical grid, voting office). Radolec explicitly stated that "each time though, because they had eyes on their data, we squashed the attackers and the city was safe. No data was lost, none of the citizens felt the pain of an attack." This serves as a powerful testament to the effectiveness of data-centric security in preventing real-world catastrophic incidents, demonstrating how continuous monitoring of data access and anomalies can thwart advanced persistent threats before they impact critical systems, which AI could potentially interact with or leverage.

The most impactful "proof of concept" of an actual AI-related breach came from the Varonis incident response team's investigation into how a competitor gained sensitive customer information: "An organization called us in to determine how a competitor got so much information from a particular database of theirs... A few days after our initial assessment and investigation, the evidence was clear. Like a lot of companies, access control was pretty lax. They couldn't tell what cloud databases and object stores house critical data and which didn't. They had widely adopted co-pilots and their connectors... one of their employees, former employees, I might add, was the first to realize this mishap. They grabbed the database of accounts, including the products they owned and how much they spend, and they left and took it to a competitor." This real-world incident directly demonstrated how lax access control, exacerbated by the adoption of AI co-pilots and their Graph Connectors, could facilitate the exfiltration of highly sensitive business intelligence by an insider or a malicious actor exploiting the broad access granted to the AI. This story provides concrete evidence of an "AI breach" in action, where the AI ecosystem became the conduit for data theft.

Collectively, these examples, coupled with the implied visual of misconfigured Graph Connector permissions, served to validate Radolec's claims about the pervasive risks and the urgent need for a data-centric security approach in the age of AI.

Defensive Implications

▶ Watch: Security's Big Moment: Empowering Business with Safe AI (9:50)

The advent of AI, particularly Generative AI and AI co-pilots, necessitates a fundamental re-evaluation of defensive strategies. Matt Radolec outlines a clear path forward, emphasizing a shift from traditional security perimeters to a laser focus on data itself.

First and foremost, defenders must recalibrate their security priorities away from endpoints and towards data "vaults." Radolec uses the analogy of an ATM versus a bank vault: "Endpoints are to data like an ATM is to a bank. Sure, there's some money there. It's worth monitoring. But it's not where your crown jewels are. The crown jewels, the good, juicy, rich data, it's in your vaults." This means investing heavily in securing and monitoring cloud databases, object stores, file shares, and other repositories where sensitive and critical data resides.

To achieve this, organizations must implement continuous monitoring of every data transaction. Just as banks monitor every credit card transaction for fraud, security teams need to track who is accessing what data, when, from where, and how. This granular visibility is crucial for establishing baselines of normal behavior and subsequently identifying deviations.

Building on transaction monitoring, the next critical step is to detect anomalies in data usage. This involves leveraging advanced analytics and, ironically, AI itself, to identify patterns that suggest unauthorized access, data exfiltration, or misuse. If a user or an AI co-pilot starts accessing vast quantities of sensitive data outside their typical operational scope, or from an unusual location, this should trigger an immediate alert and investigation, much like a fraudulent credit card transaction. This "fraud detection" for data is paramount to catching AI-driven breaches in progress.

A particularly novel and crucial defensive measure in the age of AI co-pilots is to police every prompt. Since co-pilots combine business data with LLM knowledge based on user prompts, the prompts themselves become a vector for potential data exposure. Organizations need mechanisms to monitor, log, and potentially filter or audit the prompts users are inputting into AI systems, especially those connected to sensitive data. This helps identify attempts to extract or synthesize confidential information, even if the user has legitimate access to the underlying data.

Moreover, the talk underscores the enduring importance of granular access control and Zero Trust principles, but with renewed urgency in the context of AI. The pervasive issue of over-permissioning—evidenced by 17 million files open to all employees and users only utilizing 1% of their assigned permissions—must be addressed head-on. For Microsoft 365 Copilot, specifically, defenders must pay meticulous attention to Graph Connectors. It is imperative to move beyond the default "everyone" access control lists and manually configure explicit, least-privilege permissions for every connector. This requires collaboration between security teams and developers to ensure that connectors are secured by design, not as an afterthought. Regular assessments of data access, like those offered by Varonis and Microsoft, are vital to identifying and remediating these widespread permission issues.

Finally, Radolec makes a compelling argument that automation and AI are the only viable means to combat AI threats. Human security teams, even highly skilled ones, cannot keep pace with the scale, speed, and sophistication of AI-driven attacks or the sheer volume of data interactions. Organizations must invest in security tools and platforms that leverage AI to automate detection, response, and even proactive threat hunting, allowing their "own robots on their side" to fight against malicious AI. This means integrating AI into Managed Detection and Response (MDR), Security Orchestration, Automation, and Response (SOAR), and other security operations functions.

In summary, defending against AI's blast radius requires a strategic pivot: relentless focus on data, proactive monitoring and anomaly detection, vigilant prompt policing, rigorous access control implementation, and the strategic deployment of defensive AI and automation.

Key Takeaways

  • AI is a Data Problem: The core threat and opportunity of AI revolve around data. Security efforts must shift from focusing solely on endpoints or network perimeters to securing the "vaults" where an organization's most valuable data resides.
  • AI Co-pilots Amplify Access Control Risks: Generative AI co-pilots integrate an organization's business data with LLMs, making existing lax access controls a direct pathway for potential data breaches. If a co-pilot has access, the LLM and users interacting with it effectively gain access.
  • Graph Connectors are Critical Vulnerabilities: For Microsoft 365 Copilot, Graph Connectors, which link the co-pilot to external data sources, often default to "everyone" access due to manual ACL configuration. This creates a significant, easily exploitable exposure for sensitive data.
  • Widespread Over-Permissioning is Magnified by AI: Most organizations suffer from excessive data access permissions (e.g., 17 million files open to all employees). AI co-pilots provide an easy interface to exploit these pre-existing vulnerabilities, expanding the "blast radius" of potential data exposure.
  • Implement Data-Centric Security: Defenders must monitor every data transaction, detect anomalies in data usage (like credit card fraud detection), and "police every prompt" submitted to AI co-pilots to prevent misuse and exfiltration.
  • Automate AI Defense with AI: Human teams cannot effectively combat AI-driven threats at scale. Organizations must leverage automation and AI themselves to enhance detection, response, and overall data protection strategies.

About the Speaker(s)

Matt Radolec is the Vice President of Incident Response and Cloud Operations at Varonis, a leading provider of data security and analytics software. With an impressive career spanning over 15 years in cybersecurity, Radolec has dedicated his professional life to data protection. His experience includes safeguarding highly sensitive information, from classified State Department data to critical secrets at prestigious law firms like Wilmer Hale.

At Varonis, where he has been for nearly seven years, Matt leads several key organizations, including incident response, managed detection and response (MDR), cloud operations, and most recently, sales engineering. His expertise is not limited to technical leadership; he also hosts the popular podcast "State of Cybercrime," sharing insights and trends from the front lines of cybersecurity. Radolec's passion for data protection and his deep understanding of evolving cyber threats make him a compelling voice on the challenges and solutions for securing data in the AI era.

All talks from RSA Conference 2024