Data Honeytokens for the Cloud Era - Petrus Vasenius
Petrus Vasenius
Disobey 2026 · Main Stage
Overview
In an era where organizational data is increasingly decentralized across vast cloud ecosystems, traditional perimeter and identity-centric security controls often fall short when confronted with an authenticated, authorized user. Petrus Vasenius, Cloud Security Advisory Lead at Palo Alto Networks Unit 42, presented a compelling talk at Disobey, introducing the concept of data honeytokens as a robust defense mechanism against data exfiltration, particularly from insider threats or compromised legitimate credentials.

Key moments
- 0:00 Introduction to data honeytokens and client problem
- 2:30 The story of the sad PowerBI developer begins
- 6:15 PowerBI developer story conclusion and talk purpose
- 6:45 What you will learn: talk structure and agenda
- 8:00 The harsh reality: why traditional security fails
- 10:00 Why data honeytokens are crucial in the cloud era
Data Honeytokens for the Cloud Era
Speakers: Petrus Vasenius, Cloud Security Advisory Lead, Palo Alto Networks Unit 42
Conference: Disobey
YouTube: https://www.youtube.com/watch?v=QJN7IvN6cvk
Overview
In an era where organizational data is increasingly decentralized across vast cloud ecosystems, traditional perimeter and identity-centric security controls often fall short when confronted with an authenticated, authorized user. Petrus Vasenius, Cloud Security Advisory Lead at Palo Alto Networks Unit 42, presented a compelling talk at Disobey, introducing the concept of data honeytokens as a robust defense mechanism against data exfiltration, particularly from insider threats or compromised legitimate credentials.
Vasenius's presentation highlights a critical blind spot in modern cloud security: while organizations invest millions in EDRs, firewalls, and identity protection, an attacker with a valid token and active session can often exfiltrate sensitive data undetected until it's too late. Data honeytokens offer a proactive, asset-focused detection strategy, transforming the data itself into a trap. By strategically placing realistic, yet fake, sensitive files within cloud storage environments, organizations can detect and respond to unauthorized data access attempts with high fidelity and significantly reduce their Mean Time to Detect (MTTD).
This talk is crucial for cloud security professionals, incident responders, and security architects grappling with the complexities of data protection in dynamic cloud environments. It provides practical, actionable guidance on implementing data deception strategies without requiring significant budget outlays for new licenses or services, making it an accessible and impactful approach for enhancing cloud data security posture.
Background
▶ Watch: Introduction to data honeytokens and client problem (0:00)
The genesis of the data honeytoken concept arose from a common client dilemma: a recognized medium risk for data exfiltration from their cloud environment, coupled with a very limited budget and an unwillingness to purchase new security services. This scenario underscored a fundamental challenge: traditional cloud security controls often fail to detect malicious activity when performed by a valid, authenticated user. Such users, or attackers impersonating them, possess legitimate access tokens and sessions, allowing them to bypass many conventional perimeter and identity-based defenses.
To illustrate the gravity and mechanics of such an insider threat, Vasenius recounted a hypothetical, yet all too plausible, story of a "sad PowerBI developer." This developer, frustrated by years of missed promotions and financial struggles (exacerbated by online gambling), was enticed by an anonymous request on a Telegram channel offering Bitcoin for sensitive company data. Leveraging his decade-long tenure and intimate knowledge of the company's security controls and policies, he believed he could exfiltrate "bits and pieces of data" without detection. His browsing led him to a file named FI26_financial_estimations.XLSX in an archive folder. Unbeknownst to him, this file was a honeytoken, designed to trigger an immediate alert to the Security Operations Center (SOC) upon access. His subsequent download of the file led to his swift detection, disciplinary action, and termination, leaving him "even more miserable." This narrative powerfully demonstrates how insider motivations, combined with perceived low risk of detection, drive data exfiltration and how a targeted data deception strategy can effectively counter it.
The problem of undetected data exfiltration by valid users is further compounded by the evolving landscape of cloud data storage. Data is no longer confined to hardened SQL databases but is now "scattered all over the organization's infrastructure." In the Microsoft ecosystem, this can mean data residing in OneLake, hundreds of Fabric workspaces, Azure Data Lake Storage (ADLS), or even random SaaS exports. Traditional security solutions, if not meticulously configured, may miss short-lived access and service principal activity. While identity protection remains paramount, identities can be compromised or "leaky." This necessitates a shift in defensive mindset: rather than solely focusing on perimeter or identity, organizations must also protect the asset itself – the data.
Vasenius elaborated on the broader threat model for insider risk, encompassing various personas:
- Valid users with good history: Driven by curiosity or malicious intent, leveraging their legitimate, often highly privileged, access.
- External attackers wearing a valid employee session cookie: Compromised credentials granting an attacker the guise of a legitimate insider.
- Automated solutions, apps, or tools: Originally designed for one purpose, but repurposed or exploited to perform unauthorized actions.
A common thread across all these insider definitions is their possession of valid access, rendering many Data Loss Prevention (DLP) solutions ineffective. Vasenius also referenced Palo Alto Networks Unit 42's research into sophisticated insider threats, including state-sponsored actors like DPRK (North Korea). These groups demonstrate advanced tactics, infiltrating organizations, studying security policies and penetration test reports to understand their targets, and even participating in hiring processes to embed operatives. Their ultimate motivation is often financial, funding their regime. Protecting against such advanced threats demands more than traditional controls; it requires a multi-layered approach that includes behavioral monitoring and data-centric defenses.
Key Findings
▶ Watch: PowerBI developer story conclusion and talk purpose (6:15)
The core contribution of this talk is the introduction and practical application of data honeytokens for cloud environments. These are essentially honeypots in the data domain, designed to detect unauthorized access and exfiltration attempts by individuals or processes that possess valid, authenticated access. The key findings and contributions can be summarized as follows:
- Addressing the Valid Token Blind Spot: Data honeytokens directly tackle the critical security gap where traditional perimeter and identity-focused controls fail to detect malicious activities performed using legitimate user credentials or service principals. By shifting detection to the data layer, organizations gain visibility into actions that would otherwise go unnoticed.
- High-Fidelity Detection: When properly implemented and tuned, data honeytokens boast a 100% true positive rate. Since legitimate users or automated processes should have absolutely no reason to interact with these decoy assets, any access attempt is a strong indicator of compromise or malicious intent, minimizing false positives.
- Significant Reduction in Mean Time to Detect (MTTD): In complex cloud analytics platforms with vast amounts of legitimate data traffic, spotting anomalous data exfiltration can take months. Data honeytokens can drastically reduce this MTTD from months to minutes, enabling rapid response and containment of threats.
- Realism and Strategic Placement are Paramount: The effectiveness of data honeytokens hinges on their believability. They must appear authentic in terms of naming conventions, metadata, timestamps, and even content (at least superficially). They must also be strategically placed alongside real, sensitive data to maximize the chances of an adversary encountering them.
- Leveraging Cloud-Native Capabilities for Governance and Lineage: Integrating data honeytokens with cloud governance and data cataloging solutions, such as Microsoft Purview, provides enhanced visibility. Purview can not only classify these decoys but also track their movement within the cloud environment, offering crucial forensic insights beyond initial access.
- Cost-Effective Security Enhancement: The strategy emphasizes using existing cloud infrastructure and security monitoring tools (e.g., Azure Sentinel, Log Analytics) rather than requiring new, expensive licenses. This makes data honeytokens a budget-friendly yet powerful addition to an organization's security toolkit.
- Proactive Insider Threat Mitigation: By creating an environment where unauthorized data access is immediately detectable, data honeytokens serve as a powerful deterrent and detection mechanism against insider threats, whether driven by malice, curiosity, or external coercion.
These findings collectively present a compelling case for adopting data honeytokens as an essential component of a comprehensive cloud data security strategy, particularly against sophisticated insider threats and compromised credentials.
Technical Deep Dive
▶ Watch: What you will learn: talk structure and agenda (6:45)
The implementation of data honeytokens is a multi-faceted process that spans decoy creation, strategic deployment, integration with cloud governance tools, and robust detection engineering. Vasenius provided a detailed blueprint, primarily focusing on the Microsoft Azure ecosystem but with principles applicable to any cloud or on-prem environment.
What is a Data Honeytoken?
Fundamentally, a data honeytoken is a digital honeypot within the data domain. It can take various forms:
- Valid files: Such as Excel spreadsheets (
.xlsx), CSV files (.csv), or even.txtdocuments. - SQL objects: Decoy tables, views, or stored procedures within a database.
- Decoy credentials: Fake API keys, passwords, or access tokens embedded within files.
Crucially, these decoys must contain credible metadata, names, labels, and classifications to appear legitimate. The actual content, however, should be completely nonsense or rubbish, generated from sources like ChatGPT, ensuring no real sensitive information is ever exposed. The goal is to make the file pass a "sniff test," meaning an adversary opening it should find the first two to three pages of content convincing enough to proceed with exfiltration.
Creating Realistic Decoys
Realism is the cornerstone of this strategy:
- Naming Conventions: Decoy files must adhere to the organization's existing naming conventions. If an organization has a
FI26prefix for financial documents, the decoy should follow suit (e.g.,FI26_financial_estimations.XLSX). - Metadata and Timestamps: The file's metadata, especially the last modified date, must be authentic. A decoy in an "archive" folder should have a timestamp reflecting its archival nature, not a recent modification. For regularly updated folders, a script can periodically update the decoy's timestamp to prevent it from looking stale.
- Content: While the values are dummy, the file should have valid headers and a structure that mimics genuine data. Obfuscated or generated values maintain realism without containing actual sensitive information.
- Identifiers: Each decoy should embed a unique, hidden identifier, such as a GUID (Globally Unique Identifier). Other identifiers could include canary emails (email addresses designed to trigger alerts upon use), regex patterns, or chasing keys. These IDs are critical for detection and classification.
Deployment and Seeding
Decoys must be deployed strategically alongside real data to be effective:
- Placement: They should reside within relevant project folders, ADLS containers, Fabric workspaces, or Synapse databases, mimicking the location of actual sensitive data.
- Infrastructure as Code (IaC): Using tools like Terraform for deployment is recommended. This aligns with standard cloud resource provisioning processes and ensures consistency.
- Accessibility: Crucially, decoys must remain accessible to potential adversaries. This means configuring Access Control Lists (ACLs) to inherit parent folder permissions, ensuring the decoy isn't inadvertently locked down.
- Tagging: Implement internal tags (e.g.,
honeytoken:true) for governance and tracking by the security team. These tags should also appear authentic if an adversary inspects resource tags. - Sensitivity Classification: In the Microsoft ecosystem, asset metadata (which can be hidden within file properties) can be used to assign sensitivity classifications. This metadata can then be leveraged by Microsoft Purview for advanced governance and lineage tracking.
Advanced Scanning and Registration with Microsoft Purview
Microsoft Purview plays a pivotal role in enhancing the data honeytoken strategy by providing lineage and governance:
- Visibility and Movement Tracking: Purview can scan and register decoy files. If an attacker copies or moves a decoy file from one ADLS folder to another, Purview can detect and log this movement, providing crucial forensic evidence.
- Custom Classification Rules: A custom classification rule is created in Purview to specifically look for the unique GUID embedded within the decoy files. When Purview scans a file containing this GUID, it automatically tags the asset as a "honeytoken." This classification provides a higher indicator of compromise during forensic investigations.
- Addressing Common Purview Issues:
- Private VNets/Endpoints: If data sources are secured with private VNets, Purview scans require a managed v-net integration runtime to access them.
- Managed Identities Authentication: Ensure Purview's managed identity has the necessary permissions to scan the target file locations.
- Scan Latency: Large environments may experience scan latency, which needs to be accounted for in detection strategies.
Logging and Detection Engineering
Effective detection relies on ingesting relevant logs into a Security Information and Event Management (SIEM) solution, such as Azure Sentinel:
- Essential Logs:
- Storage Blob logs: For Azure storage accounts.
- Synapse audit data: For Azure Synapse Analytics.
- Activity logs: General Azure platform logs.
- Storage access logs: Detailed access logs for storage accounts.
- Office 365 audit logs: Crucial for OneLake and PowerBI environments, which may lack dedicated logging capabilities.
- Purview scan and insight logs: To track Purview's classification and lineage activities.
- Detection Rules (Azure Sentinel Examples):
- Detecting Read/Download Operations:
- This rule targets operations like
Get Blob,Copy Blob, andGet Blob Propertiesagainst specific decoy file URIs. - Relevant log fields to include:
TimeGenerated,AttackerIP,CallerAddress,Identity, andUserAgentHeader.
- Detecting SQL Queries Against Decoys:
- For SQL-based decoys, this rule monitors queries against a dedicated SQL pool, looking for access to specific decoy view or table names.
- Relevant log fields:
TimeGenerated,PrincipalName, andClientIP.
- Advanced Purview Sensitivity Log Detection:
- Once Purview is configured for lineage and classification, this rule monitors Purview's sensitivity logs for alerts related to "honeytoken" classified assets, especially if multiple tags or classification rules are triggered. It provides details on the location and operation.
This comprehensive technical approach ensures that data honeytokens are not just static files, but an integrated, dynamic part of the cloud security monitoring infrastructure.
Demo / Proof of Concept
▶ Watch: The harsh reality: why traditional security fails (8:00)
Petrus Vasenius walked through a pre-recorded demonstration illustrating the entire data honeytoken process within the Microsoft Azure ecosystem, from deployment to alert generation. The demo showcased a practical, real-world scenario of an insider attempting to exfiltrate data.
The demonstration began in the Azure portal, where a storage account was shown. Within this storage account, a container held various folders. The focus was on a finance_archive folder, which contained the strategically placed decoy file: FI26 merger target list.CSV. This file was designed to look legitimate but contained no actual sensitive data.
Next, the demo transitioned to the Microsoft Defender portal, specifically to Azure Sentinel for analytic rule configurations. Vasenius displayed an analytic rule query configured to detect specific activities against the defined decoy files. This rule specifically looked for get blob, get blob properties, and copy blob activities. The chosen attack tactic for this rule was "exfiltration discovery and collection," categorized under the MITRE ATT&CK framework. The rule frequency was set to run every 5 minutes, which is the recommended minimum threshold for production environments.
The demonstration then moved to Microsoft Purview, highlighting its data map functionality. Here, Vasenius showed that Purview had registered 21 assets, 3 data sources (including Azure Data Lake Storage, Synapse Analytics, and a Fabric instance), and conducted 3 scans. This illustrated Purview's role in providing broad visibility into data sources. Within the Purview environment, the custom classification rule was demonstrated, showing how a specific GUID embedded within the decoy files could be used to automatically classify them as "honeytoken" assets upon scanning. Vasenius also showed a poorly placed decoy file in a Fabric lakehouse environment, named LH_secret_projects_lakehouse, emphasizing the importance of realistic placement and context. He noted that the embedded GUID, while visible in the demo for illustration, should ideally be hidden deeper within the decoy file in a real-world scenario.
To simulate an adversary, the demo depicted an "insider" browsing the cloud storage. This insider identified the FI26 merger target list.CSV file as potentially valuable and proceeded to download it. The adversary then navigated to another folder, a "production" folder, where they found and downloaded a second decoy file.
Finally, the demo returned to the Microsoft Defender portal to view the resulting alerts. The "Alerts" page displayed multiple "Decoy access" alerts, with the most recent one at the top. Clicking on the alert properties revealed an incident file with detailed insights. The query results clearly showed that the keyboard_export.CSV (the second decoy) and the merger_target_list.CSV files had indeed been accessed and downloaded. Crucially, the alerts provided critical information such as the caller IP address and identity of the accessing entity, enabling immediate action and forensic investigation.
This demo effectively showcased the end-to-end process, proving that data honeytokens, when integrated with cloud-native monitoring and governance tools, can provide high-fidelity detection of unauthorized data access.
Defensive Implications
▶ Watch: Why data honeytokens are crucial in the cloud era (10:00)
Implementing data honeytokens has significant defensive implications, offering a potent, cost-effective layer of security against insider threats and compromised credentials. However, successful deployment requires adherence to specific rules of engagement and proactive tuning.
Rules of Engagement
- No Valuable Data: Decoy files must never contain any valuable or valid information that could be considered sensitive. This is paramount to avoid creating new data leakage risks.
- ETL Exclusion: Data honeytokens must be explicitly excluded from Extract, Transform, Load (ETL) pipelines. Integrating decoys into ETL processes could lead to them appearing in legitimate PowerBI dashboards or reports, causing confusion and compromising the strategy's integrity.
- Privacy Considerations: Since the data in honeytokens is entirely fake, there is no PII (Personally Identifiable Information) risk. This circumvents many privacy concerns associated with monitoring general user behavior, as the logging only monitors access to specific traps.
- Acceptable Logging: Monitoring access to specific decoy files is generally considered acceptable and falls within legitimate security monitoring practices, as it focuses on identifying malicious intent rather than broad user surveillance.
Tuning Alerts: Addressing Common Failures
Vasenius highlighted three common pitfalls and their solutions for effective alert tuning:
- Failure #1: The "Helpful Data Engineer"
- Problem: A data engineer, seeing a file named
secret.CSV, might innocently delete it or interact with it, generating a false positive. - Solution: Use scarier, more official naming conventions (e.g.,
temp_chunk_legal_hold_do_not_touch.CSV). If possible, deny access to the engineering team's group for these specific files (though Azure resource locks don't work at the file level, ACLs can). For automated jobs, use partition isolation so that automated processes naturally skip the decoy files.
- Failure #2: "50 Incidents in 5 Minutes"
- Problem: After turning on an alert, a flood of incidents occurs due to legitimate cloud services (e.g., Microsoft Defender for Storage, Purview scanner, Azure Backup) accessing the files.
- Solution: Identify legitimate noise by examining the
User-Agentheader andIdentityfields in the logs. Update Sentinel rules to exclude known legitimate scanners. For example,where Identity contains "PurviewScanner"orwhere UserAgentHeader contains "AzureStorageDefender". If Purview lineage is desired, the Purview scanner should not be excluded from logs, but rather understood as legitimate activity.
- Failure #3: "Backup Solution Reads File Every Night"
- Problem: A backup solution regularly reading the decoy file can trigger false "data exfiltration" alerts.
- Solution: Fingerprint the backup agent by looking for specific
User-Agentheaders (e.g.,Azure Backup,Veeam,Commvault) orIdentityfields (e.g.,Azure Backup Service). Update the detection rule to exclude these known backup activities.
Value Proposition and Strategic Recommendations
The value proposition of data honeytokens is compelling:
- Drastically Reduced MTTD: From months to minutes, enabling swift incident response.
- 100% True Positive Rate: When tuned correctly, any alert indicates genuine unauthorized access.
- Minimized False Positives: Legitimate users have no business touching these assets, making alerts highly reliable.
- Red Team Calibration: Engaging red teams or penetration testers to specifically seek out sensitive data can help calibrate the believability and effectiveness of the decoy strategy.
- Management Reporting: Clear, quantifiable metrics (MTTD reduction, true positive rate) make it easy to report the value of this strategy to management.
Ultimately, data honeytokens represent a strategic shift in defense from the perimeter or identity to the data itself. While existing defenses are still crucial, protecting the data directly provides a critical last line of defense. Vasenius's final recommendations include:
- Validate the Threat Model: Ensure data honeytokens are a worthwhile effort for the organization's specific context (e.g., large, complex organizations with many employees are ideal candidates).
- Start Small: Avoid trying to build a perfect system instantly. Begin by planting one believable decoy in a non-production ADLS container and configuring a single Sentinel alert. This allows for iterative learning and refinement before scaling up.
By adopting this approach, organizations can empower their data to "rat out insiders or intruders," significantly bolstering their cloud security posture.
Key Takeaways
- Traditional cloud security controls, focused on perimeter and identity, are often blind to malicious activity performed with valid user tokens or compromised legitimate credentials.
- Data honeytokens act as honeypots in the data domain, providing a high-fidelity mechanism to detect unauthorized access and exfiltration attempts by insider threats or external attackers.
- Realism is paramount for effective data honeytokens, requiring authentic naming conventions, timestamps, metadata, and superficially convincing content to pass an adversary's "sniff test."
- Leveraging cloud-native services like Microsoft Purview enhances the strategy by providing data lineage, governance, and the ability to track the movement and classification of decoy files.
- Proper alert tuning is crucial to minimize false positives from legitimate cloud services (e.g., backup solutions, security scanners) by filtering based on user agent headers and identities.
- This strategy can dramatically reduce Mean Time to Detect (MTTD) from months to minutes, achieve a 100% true positive rate, and significantly bolster an organization's defense posture against sophisticated data exfiltration attempts.
About the Speaker(s)
Petrus Vasenius is the Cloud Security Advisory Lead for the EMEA region at Palo Alto Networks Unit 42. With over 10 years of experience in the security industry, his primary role has been to bridge the gap between complex technical security problems and management communication. He specializes in helping organizations prepare against unknown threats and leading a team of highly skilled cloud security professionals. His expertise lies in translating intricate security challenges into actionable strategies for diverse organizational contexts.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Competent, well-structured talk on an underutilized defensive technique — data honeytokens in cloud environments — with solid Microsoft stack implementation detail and honest operational tuning advice. Nothing here is novel to anyone who's read the canary token literature or run deception tooling before, but Vasenius executes the 'here's how to actually build this on Azure' angle cleanly enough to be useful to a practitioner audience.
Heather Calloway (CISO) — SOLID
A technically competent and well-structured practitioner talk on data honeytokens for cloud environments. Vasenius delivers clear implementation guidance and a usable detection architecture, but the session operates almost entirely at the tool-and-tactic layer — it never engages the governance, accountability, or program-level questions that would make this relevant to the people who decide whether to fund and operationalize it.