Eavesdropping on Controller Acoustic Emanation for Keystroke Inference Attack in Virtual Reality

Shiqing Luo

Network and Distributed System Security (NDSS) Symposium 2024 · Day 3 · Physical Attacks

Overview

Virtual Reality (VR) is rapidly evolving from a niche gaming accessory into a pervasive computing platform, boasting over 171 million users and extending its utility across diverse sectors from healthcare to education. However, this expansion brings with it a burgeoning set of privacy challenges. VR systems, by their nature, collect and store vast amounts of personal data, including highly sensitive information such as passwords, credit card numbers, and biometric identifiers. The increasing reliance on VR for inputting such data necessitates robust security measures to protect users from sophisticated privacy breaches. This talk introduces Heimdall, a groundbreaking acoustic keystroke inference attack that exploits the distinct sounds emitted by VR controllers during user input.

Watch on YouTube · Slides

Visual summary for Eavesdropping on Controller Acoustic Emanation for Keystroke Inference Attack in Virtual Reality by Shiqing Luo
Visual summary for Eavesdropping on Controller Acoustic Emanation for Keystroke Inference Attack in Virtual Reality by Shiqing Luo

Key moments

  1. 0:00 Introducing Heimdall: a novel, flexible VR keystroke attack.
  2. 2:00 Understanding VR keystroke mechanics and user behavior.
  3. 2:20 Strong acoustic correlation between controller and key pressed.
  4. 2:50 Unique challenges for acoustic attacks in VR.
  5. 4:00 Threat model: public and insider attack scenarios.

Eavesdropping on Controller Acoustic Emanation for Keystroke Inference Attack in Virtual Reality

Speakers: Shiqing Luo

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=tNkaHmcIPAE

Overview

Virtual Reality (VR) is rapidly evolving from a niche gaming accessory into a pervasive computing platform, boasting over 171 million users and extending its utility across diverse sectors from healthcare to education. However, this expansion brings with it a burgeoning set of privacy challenges. VR systems, by their nature, collect and store vast amounts of personal data, including highly sensitive information such as passwords, credit card numbers, and biometric identifiers. The increasing reliance on VR for inputting such data necessitates robust security measures to protect users from sophisticated privacy breaches. This talk introduces Heimdall, a groundbreaking acoustic keystroke inference attack that exploits the distinct sounds emitted by VR controllers during user input.

Prior research has demonstrated the feasibility of inferring VR keystrokes using side-channel signals, primarily relying on video or wireless signal analysis. While these attacks proved that sensitive input could be compromised without direct access to the Head-Mounted Display (HMD), they suffered from significant practical limitations. They typically required strict positioning and line-of-sight to the controller, making them easily circumvented by users adopting secure physical layouts. Heimdall fundamentally overcomes these constraints by leveraging acoustic emanations, allowing a malicious recording device, such as a smartphone, to be placed anywhere around the victim, even in non-line-of-sight scenarios where the controller is obscured.

Developed by Shiqing Luo, Heimdall represents the first acoustic attack of its kind in VR, addressing unique challenges inherent to 3D interaction that are absent in traditional acoustic attacks on physical keyboards. These challenges include differentiating sounds in a 3D space, adaptively mapping keystroke sounds to keys despite varying recording placements, and handling the subtle, yet impactful, hand rotations that occur during VR typing. Through extensive evaluations involving 30 participants, Heimdall achieved a remarkable key inference accuracy of 96.51% and a top-5 accuracy ranging from 85.14% to 91.22% for passwords of 4 to 8 characters. Its demonstrated robustness across diverse practical conditions — including different smartphone-user placements, attack environments, hardware models, and victim conditions — establishes Heimdall as a highly practical and potent threat to VR user privacy, demanding immediate attention from the security community.

Background

▶ Watch: Introducing Heimdall: a novel, flexible VR keystroke attack. (0:00)

The foundation of the Heimdall attack lies in understanding user interaction within VR environments and the acoustic characteristics of VR controllers. VR devices predominantly utilize a controller-based keystroke method. When text input is required, a virtual QWERTY keyboard is rendered directly in front of the user within the HMD. Users navigate this virtual keyboard by physically moving a handheld controller, and a keystroke is registered by clicking a button on the controller. A critical observation for side-channel attacks is that during sensitive input sessions, the virtual keyboard remains fixed relative to the user, who generally maintains a stable physical location, primarily moving only the controller. The VR operating system meticulously tracks the controller's 3D position, and a click event confirms the key selection.

A core finding underpinning Heimdall is the strong correlation between the acoustic signal emitted from the controller and the specific key being entered. Studies confirmed that the controller's 3D positions for typing different keys are distinct, yet remarkably consistent for the same key across different users. This clustering, evidenced by a Silhouette Score of 0.64, indicates that each key effectively corresponds to a unique sound source location in 3D space. Furthermore, keystroke sounds for specific keys, such as 'M' and 'P', exhibit distinguishable acoustic shapes that remain consistent even when typed by different users. This cross-user similarity is crucial, as it enables an attacker to pre-collect a baseline mapping of keystroke sounds to keys, a prerequisite for the attack.

However, realizing an acoustic keystroke inference attack in VR presents several formidable challenges:

  1. 3D Sound Differentiation: Unlike 2D surfaces where time difference of arrival between two microphones might suffice for localization, VR keystrokes originate from various 3D positions. Standard smartphone microphones are omnidirectional, rendering them ineffective for differentiating directional signals in a 3D space based on simple phase differences. A novel approach for directional signal acquisition is required.
  2. Adaptive Mapping: The relative position and orientation between the malicious smartphone and the victim can vary significantly in real-world attack scenarios. This variability invalidates static baseline mappings, as the sound characteristics and its Direction of Arrival (DOA) for the same key will change with different placements. An adaptive mechanism is essential.
  3. Hand Rotation Errors: Users often introduce subtle wrist rotations during typing, which can generate keystrokes from similar controller positions despite targeting different keys. This can lead to overlapping DOAs for different keys, causing mapping ambiguities and inference errors.

Previous acoustic keystroke inference attacks on physical keyboards and touchscreens, while inspiring, cannot be directly applied to VR due to its unique 3D interaction modality. Other VR attacks relying on video or wireless signals are severely constrained by line-of-sight requirements, a limitation Heimdall explicitly aims to overcome. Research in sound localization often demands specialized microphone arrays or multiple recording devices. Heimdall, in contrast, seeks to achieve a balance between hardware simplicity (using a single smartphone) and localization precision.

Key Findings

▶ Watch: Understanding VR keystroke mechanics and user behavior. (2:00)

Heimdall introduces several key findings and contributions that significantly advance the state of VR side-channel attacks:

  • Novel Acoustic Attack Vector: Heimdall is the first documented acoustic keystroke inference attack targeting VR environments. Unlike prior video or wireless signal-based attacks, it exploits the distinct clicking sounds emitted by VR controllers, offering a fundamentally new and stealthier vector.
  • Placement Flexibility and Non-Line-of-Sight Capability: A critical innovation is Heimdall's ability to operate effectively regardless of the malicious smartphone's placement relative to the victim. It overcomes the severe line-of-sight and strict positioning constraints that limited previous attacks, functioning even when the controller is partially or completely obscured by the user's body. This makes it a far more practical and potent real-world threat.
  • High Inference Accuracy: The system achieves an impressive key inference accuracy of 96.51% and a top-5 password accuracy ranging from 85.14% to 91.22% for passwords between 4 and 8 characters. This level of accuracy demonstrates a significant privacy risk.
  • Technical Innovations for 3D Acoustic Sensing: Heimdall introduces three core technical modules to address the unique challenges of VR acoustic side channels:
  • Customized Directional Microphones: By attaching simple 5-centimeter porous plastic tubes to off-the-shelf smartphone microphones, Heimdall transforms them into directional sensors capable of differentiating sounds in 3D space.
  • Adaptive DOA-Key Mapping Scheme: This module intelligently adapts pre-collected baseline DOA-key mappings to varying attacker-victim placements by identifying consistent angular shifts in DOAs caused by rotation and translation.
  • Inter-Key Relation-based Calibration Model: To mitigate errors caused by subtle hand rotations, Heimdall incorporates a Hidden Markov Model (HMM) that leverages the time interval and controller moving direction between keystrokes, significantly enhancing inference robustness.
  • Robustness Across Diverse Conditions: Extensive evaluations confirm Heimdall's resilience against various practical impacts, including different smartphone-user placements (e.g., behind the victim, to the side), attack environments (laboratory, office, library, with background noise), hardware models (different smartphones and HMDs), and victim conditions (user fatigue, mid-term typing consistency). This robustness underscores its real-world applicability.
  • Identification of a Critical VR Privacy Vulnerability: This work highlights a previously unexplored and significant side-channel vulnerability in VR systems, underscoring the urgent need for developers and users to implement more robust privacy safeguards in this rapidly growing computing platform.

Technical Deep Dive

▶ Watch: Strong acoustic correlation between controller and key pressed. (2:20)

Heimdall's sophistication lies in its multi-faceted approach to overcoming the inherent challenges of acoustic keystroke inference in a 3D VR environment. The attack operates under a well-defined Threat Model and is structured into three interdependent technical modules.

Threat Model and System Overview

The threat model for Heimdall envisions a victim inputting sensitive information, such as passwords or security question answers, within a VR environment using a handheld controller. Concurrently, a malicious smartphone, strategically placed near the victim, records the acoustic emanations. Two primary attack scenarios are considered:

  1. Public Attack: The attacker places their smartphone in a public setting, like a library or airport, on a shared table.
  2. Insider Attack: The victim, familiar with the attacker, allows the smartphone to be placed nearby in a private setting such as an office or laboratory.

Crucially, Heimdall is stealthier than prior attacks because the victim's eyes are obscured by the HMD, reducing their vigilance towards nearby devices and granting the attacker greater flexibility in smartphone placement. The attacker's assumptions are minimal: access to the physical location, knowledge of the victim's VR HMD type (easily obtainable visually to determine virtual keyboard layout), but no assumption of line-of-sight to the HMD or controller, no need for precisely aligned wireless transceivers, and no requirement to install malware on the HMD. This means traditional countermeasures like secure physical layouts or avoiding camera views are ineffective.

Heimdall's system architecture comprises three core modules designed to address the challenges of 3D sound differentiation, adaptive mapping, and hand rotation errors:

  1. Directional Acoustic Signal Acquisition: Focuses on capturing and segmenting differentiable VR keystroke sounds.
  2. Adaptive DOA-Key Mapping Scheme: Identifies keystroke sounds based on their Direction of Arrival (DOA) and adapts pre-collected baseline DOAs to the specific attack scenario for key inference.
  3. Inter-Key Relation-based Calibration Model: Employs a Hidden Markov Model (HMM) to refine inference by considering temporal and directional relationships between consecutive keystrokes.

Directional Acoustic Signal Acquisition

The primary hurdle in acoustic keystroke inference in 3D is the omnidirectional nature of standard smartphone microphones. Heimdall addresses this through Microphone Customization. Inspired by human ear canals, off-the-shelf smartphone microphones are converted into directional sensors. This is achieved by attaching a 5-centimeter porous plastic tube to each of the smartphone's two microphones. These tubes feature uniformly distributed side slots, pierced with a 5-millimeter spacing. When a keystroke sound arrives from a specific DOA, it generates sub-signals with varying phases at these side slots. These sub-signals travel along multiple paths within the tube, interfering constructively or destructively, resulting in a uniquely distorted recording with a distinct phase and magnitude profile specific to that DOA. This physical modification allows for the differentiation of sounds in 3D space. The tubes are 3D-printed for ease of implementation, and recordings are made in stereo at a 48 kHz sampling rate. Theoretically, this setup enables Heimdall to differentiate sound sources with a minimum distance of 7 millimeters.

Following signal acquisition, Keystroke Sound Recognition is performed:

  1. Background Noise Removal: Raw acoustic signals are contaminated by ambient noises. Instead of simple bandpass filters, Heimdall employs wavelet denoising using a Maximal Overlap Discrete Wavelet Transform with a Daubechies 3 wavelet. This method is particularly effective at isolating transient, high-frequency controller clicks from persistent, low-frequency background noise, preserving the distinct click sounds.
  2. Controller Click Segmentation: Keystroke sounds are transient events. After denoising, specific click segments must be identified. A magnitude threshold is applied to the denoised signal, as controller clicks are highly salient. For signal periods exceeding this threshold, a peak is identified, and a 200 ms segment centered at that peak is extracted as a controller click. A pilot study optimized this threshold to 0.028, yielding an F1 score of 0.967 for click segmentation.
  3. Keystroke Session Recognition: Not all controller clicks are keystrokes. To differentiate sensitive keystroke sessions from other VR interactions (e.g., app launching, navigation), Heimdall observes distinct, temporally stable inter-click interval patterns during typing. A click frequency-based optimization (Equation 3 in the paper) is applied, with parameters configured as fmin = 0.2, fmax = 0.4, and tmin = 7.5 seconds for passwords of 4+ characters. This approach achieves a True Positive Rate (TPR) of 0.98 and a False Positive Rate (FPR) of 0.04.

Adaptive DOA-Key Mapping

The most significant challenge for practical deployment is the variability in smartphone placement relative to the user. Direct sound-key mapping is ineffective because raw acoustic signals change drastically with placement shifts. Heimdall’s key insight here is that while raw signals change, the varying Direction of Arrival (DOA) of a keystroke sound presents a traceable pattern under different placements. Even a 45-degree shift in smartphone position results in a consistent 45-degree shift in the azimuth of DOAs for keys like 'M' and 'P'. This consistent angular change allows for adaptive mapping.

The adaptive DOA-Key mapping procedure involves four steps:

  1. Collecting Baseline DOA-Key Mapping: In a controlled baseline placement, the attacker types each target key k (26 English letters, 10 digits). The controller's 3D coordinates (xk, yk, zk) are accessed from the VR OS, converted to the smartphone's coordinate system, and used to derive the baseline DOA for key k, denoted as dk,base = (φk, θk) (azimuth and altitude, Equation 4). This forms the baseline mapping: {k ∈ K, dk,base}.
  2. Deriving DOAs of Keystroke Sounds in the Attack: In an actual attack, direct access to the victim's controller coordinates is impossible. Instead, DOAs are derived from the recorded acoustic signals. The interaction of the keystroke sound with the porous tube microphones, coupled with environmental channel responses, influences the phase and magnitude of recorded signals (Yl, Yr for left/right mics). By relating Yl/Tl to Yr/Tr (Equation 6), where Tl(d) and Tr(d) are tube responses for different DOAs d, Heimdall can derive the DOA. Tl(d) and Tr(d) are pre-determined using the Maximum Length Sequence (MLS) method in an anechoic chamber (Equations 7 and 8). The DOA d* for a keystroke is found by identifying the d that maximally satisfies Equation 6 (Equation 9). This process yields ds,atk for each victim keystroke s (Equation 10).
  3. Updating DOA-Key Mapping: The relative displacement between the smartphone and user from baseline can be modeled as a combination of rotation (around the z-axis by γ) and translation (shifts δα, δy, δz along x, y, z axes). The goal is to find the optimal displacement parameters π = (δα, δy, δz, γ) that allow a subset of the displaced baseline DOAs dk,atk(π) to maximally match the measured victim DOAs ds,atk. This is solved by an optimization problem (Equation 14), searching δα, δy, δz from -1 meter to 1 meter (with a 0.05m step) and γ from 0 to 359 degrees (with a 1-degree step). This yields the updated DOA-Key mapping: {k ∈ K, dk,atk(π)}.
  4. Inferring Keystrokes in the Attack: Finally, to infer a victim keystroke s from its measured DOA ds,atk, Heimdall retrieves the key k from the updated DOA-Key mapping whose dk,atk(π) is most similar to ds,atk (Equation 15). This process is repeated for the entire input sequence.

Inter-Key Relation-based Calibration Model

Even with adaptive DOA-Key mapping, errors can occur, particularly when users rotate their wrist significantly, leading to minimal hand translation but substantial changes in controller orientation. This can cause different keys to have similar DOAs, leading to ambiguities (e.g., 'Z' keystrokes incorrectly mapped to 'Q'). To resolve these one-to-multiple mapping challenges, Heimdall introduces Inter-Key Relation-based Calibration.

The core insight is that despite hand rotation, the time interval and controller moving direction between two consecutive keystrokes are still strongly dependent on their inter-key distance on the virtual keyboard. For example, typing 'P' then 'O' will typically take less time and involve a different direction than 'P' then 'T'.

A. Modeling Keystroke Transition: This module models keystroke transitions as a Hidden Markov Model (HMM). The hidden states are keystroke pairs (e.g., A-B), and the observed outcomes are the inter-key time interval (o1) and controller moving direction (o2).

  • The HMM defines a start probability Po(q), transition probability P between hidden states, and observation probabilities P(o1|q) and P(o2|q).
  • Probability Distribution of Controller Moving Direction: Nine possible directions are considered (up, up-right, right, etc.). For a given keystroke transition t, four directions are possible: t, two directions 45 degrees away from t, and 'no-move', each assigned a 0.25 probability.
  • Probability Distribution of Time Interval: P(o1|q) is derived from 46 representative keystroke transitions with distinct inter-key distances. These distributions are typically Gaussian-like and unimodal.

B. Keystroke Mapping Correction: Using the trained HMM, Heimdall derives the probability of a keystroke sequence given the observed time intervals and moving directions.

  • Calibration Step 1: An exhaustive list of all possible keystroke sequences is generated and ranked in decreasing order of their HMM-derived probabilities.
  • Calibration Step 2: To maintain fidelity with the initial DOA-based mapping, sequences with a large Hamming distance from the initial mapping result are removed. A pilot study determined an optimal Hamming distance of 2. This ensures that calibrated sequences are both highly probable according to the HMM and sufficiently similar to the original DOA-based inference.

The output is a list of inference candidates, with the top-1 candidate being the initial DOA-based mapping result, and subsequent candidates derived from the highest-ranking calibrated sequences.

Demo / Proof of Concept

▶ Watch: Unique challenges for acoustic attacks in VR. (2:50)

The efficacy of Heimdall was rigorously evaluated through extensive experiments. The prototype system utilized a Samsung Galaxy S8 as the malicious smartphone and a Google Daydream as the primary VR HMD, both equipped with the customized directional microphones. For robustness validation, other devices like the Google Pixel 1, Samsung GearVR, and Oculus Rift were also tested.

Participants: 30 university students (18 males, 12 females, aged 19-30) were recruited. All participants provided informed consent and underwent a practice session to familiarize themselves with VR keystroke input.

Data Collection:

  • A custom Unity App featuring a standard QWERTY keyboard was developed for data collection.
  • Baseline DOA-Key Mapping: An attacker entered 26 English letters and 10 digits (10 trials each) in a controlled baseline placement. Controller coordinates were collected directly from the VR OS to derive dk,base.
  • HMM Modeling: The attacker entered 46 representative keystroke transitions (10 trials each) to train the HMM.
  • Attack Testing: 29 victims entered 45 different passwords (ranging from 4 to 8 characters) across 7 distinct smartphone-user placements. These placements covered various translations and rotations, crucially including non-line-of-sight scenarios (e.g., smartphone behind or to the right of the user). In total, the evaluation involved **29 victims 45 passwords 7 placements * 30 rounds, resulting in 274,050 attack attempts**.

Evaluation Metrics: The performance was assessed using top-w accuracy (successful if one of the 'w' candidates matched the password) and key inference accuracy (ratio of correctly inferred characters to total characters).

Evaluation Results:

  • Performance of Keystroke Inference:
  • All Characters: Heimdall demonstrated strong overall performance, with an average precision of 95.14% and recall of 96.29% across all alphanumeric characters.
  • Password Length: Top-1, 3, and 5 accuracy generally increased with password length, peaking at 6-character passwords (achieving 86.97% top-1, 88.43% top-3, and 91.58% top-5 accuracy). For 8-character passwords, top-5 accuracy remained a respectable 85.14%. The slight decrease for longer passwords was attributed to increased hand rotation and mapping errors, which the HMM helps mitigate.
  • Inter-key Distances: Heimdall maintained stable accuracy regardless of whether passwords had short, long, or mixed inter-key distances, confirming its ability to handle diverse and random inputs.
  • More Inference Candidates: Password inference accuracy significantly improved with more candidates. Heimdall achieved 95% accuracy within 20 attempts and converged to 98% with 50 candidates. This indicated that wrist rotation accounted for approximately 18% of initial mapping errors, with the remaining 2% due to DOA-Key mapping inaccuracies. This result highlights a serious threat, as even a small set of candidates can be used for effective brute-force attacks.
  • Smartphone-User Placement: A crucial finding was Heimdall's consistent performance across all 7 tested smartphone-user placements. Even in challenging non-line-of-sight scenarios (e.g., smartphone behind or to the right of the victim), an acceptable key inference accuracy of 92.06% was achieved. While accuracy was slightly lower when the smartphone was directly to the right due to similar DOAs, this robust performance across varied placements unequivocally confirmed Heimdall's placement flexibility.
  • Robustness of Keystroke Inference:
  • Attack Environments: Heimdall proved stable across different rooms (laboratory, office, library) and various noise conditions (static fans, air conditioners, moving people talking/walking). This resilience is attributed to the effective wavelet denoising and frequency-domain division approach employed, which robustly handles background noise and multipath distortion.
  • Recording Distances: As the distance between the victim and smartphone increased from 1m to 2.2m, key inference accuracy decreased from 96.51% to 84.22%. Despite the degradation, 85% accuracy at 2.2 meters still exposes significant information, allowing attackers to balance performance with stealthiness.
  • Smartphone Models: Only a negligible 1% decrease in accuracy was observed when using a Google Pixel 1 (44.1 kHz sampling rate) compared to the Samsung S8 (48 kHz), confirming broad applicability across various smartphone models.
  • VR HMD Models: Heimdall generalized well to different VR HMDs, showing negligible performance differences with Samsung GearVR and Oculus Rift.
  • Short-term Performance: User fatigue throughout the day (10 AM to 6 PM) resulted in less than 2% fluctuation in inference accuracy, indicating stable keystroke patterns.
  • Midterm Performance: Over five days, no significant variance was observed, suggesting users become familiar with VR typing, maintaining consistent patterns.

Defensive Implications

▶ Watch: Threat model: public and insider attack scenarios. (4:00)

Heimdall exposes a critical and previously underestimated side-channel vulnerability in VR systems, necessitating proactive defensive strategies. While the talk touched upon potential countermeasures, it also highlighted their inherent challenges:

  • Randomized Keyboard Layout: Intentionally disrupting the sound-key correlation by randomizing the virtual keyboard layout for each input session could prevent baseline mapping. However, this approach is highly cumbersome for users, significantly degrading the user experience, and is not natively supported by current VR operating systems. Implementing this would require substantial changes at the OS level or within individual applications.
  • Interrupting Keystroke Continuity: Victims could be advised to perform fake clicks or large, arbitrary hand movements to switch keyboard views or introduce noise into the acoustic data. While this might obfuscate sensitive input, it severely disrupts the smoothness of VR interaction and overall user experience, making it an impractical long-term solution for regular use.
  • Eliminating Keystroke Sounds: The most direct countermeasure would be to design VR controllers that produce minimal or no distinct acoustic emanations during button clicks. Using mechanical keyboards or designing quieter button mechanisms could reduce the acoustic side channel. However, most VR headsets lack native support for external keyboards, and even external keyboards can expose users to other attacks like shoulder-sniffing. A hardware redesign of the controller's input mechanism to make clicks acoustically indistinguishable or significantly quieter would be a more robust solution.

Beyond these direct suggestions, the findings of Heimdall strongly urge VR HMD and OS developers to prioritize acoustic side-channel resilience in future designs. This could involve:

  • Active Noise Cancellation (ANC) at the Controller: Integrating small, localized ANC systems directly into VR controllers could actively mask or cancel out the distinct click sounds before they propagate.
  • Sound Masking: VR HMDs could generate subtle, dynamic background audio noise during sensitive input sessions, specifically designed to mask the controller click frequencies without affecting the user's perception of the virtual environment.
  • Software-level Anomaly Detection: The VR OS could monitor controller movement and click patterns, flagging unusual sequences that might indicate an attack, though this would likely generate false positives.
  • Multimodal Input Obfuscation: Encouraging or forcing users to use alternative input methods (e.g., gaze-based selection with a click, voice input with obfuscation) for sensitive data could reduce reliance on the acoustically vulnerable controller clicks.

Ultimately, Heimdall demonstrates that relying solely on visual or wireless security for VR input is insufficient. A holistic approach that considers the acoustic dimension is crucial for safeguarding user privacy in the rapidly expanding VR ecosystem.

Key Takeaways

  • Heimdall is the first practical acoustic keystroke inference attack against VR systems, overcoming the line-of-sight and strict placement limitations of previous video/wireless-based attacks.
  • The attack leverages customized directional microphones (smartphones with porous plastic tubes) to accurately differentiate VR controller keystroke sounds in 3D space.
  • An adaptive DOA-Key mapping scheme intelligently accounts for varying smartphone-user placements, ensuring robust key inference regardless of recording device position.
  • A Hidden Markov Model (HMM)-based calibration model further refines accuracy by considering inter-key time intervals and controller movement directions, mitigating errors caused by subtle hand rotations.
  • Heimdall achieves high accuracy, with a 96.51% key inference accuracy and 85.14%-91.22% top-5 password accuracy for 4-8 character passwords, demonstrating a significant privacy vulnerability.
  • The attack is highly robust across diverse practical conditions, including different environments, recording distances (up to 2.2m), smartphone models, and VR HMDs, making it a potent real-world threat.

About the Speaker(s)

The talk was presented by Shiqing Luo. No additional biographical information, such as title or company, was provided in the transcript or metadata.

All talks from Network and Distributed System Security (NDSS) Symposium 2024