Eavesdropping on Black-box Mobile Devices via Audio Amplifier's EMR

Huiling Chen

Network and Distributed System Security (NDSS) Symposium 2024 · Day 2 · Audio & Voice Security · Audio & Voice Security

Overview

In an era where digital privacy is paramount, the talk "Eavesdropping on Black-box Mobile Devices via Audio Amplifier's EMR" by Huiling Chen unveils a critical and previously underestimated vulnerability. The research, embodied in a system dubbed Periscope, demonstrates a novel method for surreptitiously recovering private audio from mobile devices, even when users are diligently using headphones to maintain acoustic privacy. This work fundamentally challenges the assumption that headphones provide sufficient isolation against eavesdropping, revealing that the very act of plugging them in can inadvertently enhance a device's susceptibility to electromagnetic radiation (EMR) side-channel attacks.

Watch on YouTube · Slides

Visual summary for Eavesdropping on Black-box Mobile Devices via Audio Amplifier's EMR by Huiling Chen
Visual summary for Eavesdropping on Black-box Mobile Devices via Audio Amplifier's EMR by Huiling Chen

Key moments

  1. 0:00 Introduction: Eavesdropping via audio amplifier's EMR (Periscope)
  2. 4:00 Threat model and attacker's black-box capabilities
  3. 6:00 Feasibility study: EMR reflection of audio sounds

Eavesdropping on Black-box Mobile Devices via Audio Amplifier's EMR

Speakers: Huiling Chen

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=fOI5Ocp8jIM

Overview

In an era where digital privacy is paramount, the talk "Eavesdropping on Black-box Mobile Devices via Audio Amplifier's EMR" by Huiling Chen unveils a critical and previously underestimated vulnerability. The research, embodied in a system dubbed Periscope, demonstrates a novel method for surreptitiously recovering private audio from mobile devices, even when users are diligently using headphones to maintain acoustic privacy. This work fundamentally challenges the assumption that headphones provide sufficient isolation against eavesdropping, revealing that the very act of plugging them in can inadvertently enhance a device's susceptibility to electromagnetic radiation (EMR) side-channel attacks.

Periscope's core innovation lies in its ability to exploit unintentionally leaked EMRs emanating from a mobile device's internal audio amplifier. These radiations, which are typically weak and distorted, are shown to contain discernible information about the audio being processed. Crucially, the research identifies that plugged headphones act as effective antennas, significantly amplifying these EMR signals and extending the range from which an attacker can capture them. The system operates as a black-box attack, meaning it requires no prior knowledge of the target device's model, hardware, or any training data, making it highly adaptable and practical for real-world scenarios.

The significance of Periscope cannot be overstated. It exposes a widespread privacy threat affecting a broad spectrum of mobile devices—from smartphones to laptops—and a variety of applications, including private calls, voice messages, and confidential meetings. The recovered audio, even from low signal-to-noise ratio (SNR) EMRs, is demonstrated to be highly intelligible to human listeners and online speech-to-text tools, achieving a Word Error Rate (WER) as low as 7.44%. The researchers have proactively reported these findings to six leading mobile manufacturers, underscoring the urgency of addressing this pervasive vulnerability.

Background

▶ Watch: Introduction: Eavesdropping via audio amplifier's EMR (Periscope) (0:00)

The landscape of audio eavesdropping has historically been diverse, with various approaches attempting to capture sensitive vocal information. Early research focused on acoustic eavesdropping, analyzing direct sound waves or subtle vibrations caused by human speech. Examples include WiHear, which leveraged Wi-Fi signals perturbed by mouth and throat movements, and works that repurposed magnetic hard drives or speakers into microphones. However, these methods are largely rendered ineffective when a user employs headphones, as the headphones provide a substantial physical barrier, drastically reducing or eliminating measurable acoustic leakage.

A second category of attacks targets loudspeakers using non-acoustic side channels. This includes exploiting smartphone motion sensors (gyroscopes, accelerometers) to detect micro-vibrations, high-speed cameras to observe object vibrations, or laser beams and radio signals (LiDAR, Wi-Fi CSI, mmWave, RFID) to reconstruct sounds by analyzing airborne acoustic pressure perturbations. While innovative, these high-frequency radio signal attacks often suffer from susceptibility to environmental dynamics and frequently necessitate data-driven training models, which Periscope aims to circumvent. For instance, MagEar utilized near-field magnetic fluxes from speakers but required high volumes (80dB) and had a narrow effective angle (0-20 degrees).

Periscope distinguishes itself by building upon a third, more insidious category: EMR-based side-channel attacks. Previous studies have established that EMRs from electronic devices can inadvertently leak a wealth of sensitive data, ranging from cryptographic keys (e.g., AES-128) and screen contents to keystrokes and memory usage. More recent work, such as "Tempest Comeback," demonstrated the potential to extract audio processed by wireless devices from the EMR of mixed-SOC chips. Periscope extends this lineage by specifically and systematically targeting the EMR originating from the audio amplifier within headphone-plugged mobile devices. This particular attack vector, where headphones paradoxically amplify the vulnerability rather than mitigate it, was not adequately addressed by prior EMR side-channel research, highlighting Periscope's novel contribution to the field of privacy and security.

Key Findings

▶ Watch: Threat model and attacker's black-box capabilities (4:00)

The research presented in Periscope reveals several pivotal findings that redefine our understanding of mobile device audio privacy:

  1. Strong Correlation Between Audio and EMR: Feasibility studies conclusively demonstrated a robust correlation between the audio signals processed by a mobile device and the EMRs it emits. Cross-correlation analysis between original audio and captured EMRs consistently showed sharp peaks at lag=0 across diverse devices, confirming that audio content is indeed leaked via EMRs.
  2. Audio Amplifier as the EMR Source with Nonlinearity: The primary source of these exploitable EMRs was definitively identified as the mobile device's internal audio amplifier. Experiments involving two-tone and single-tone audio signals revealed the presence of not only fundamental frequencies but also distinct harmonic and cross-product frequencies in the EMR spectrum. These distortions are classic characteristics of amplifier nonlinearity, providing a unique spectral signature that Periscope leverages for audio recovery.
  3. Headphones Act as Antennas: A counter-intuitive but critical finding is that plugged headphones, including their conductive wires and microspeakers, do not merely provide acoustic isolation but actively function as antennas. This significantly enhances the strength of the EMRs emitted by the audio amplifier, leading to higher signal-to-noise ratios (SNRs) for eavesdropping and enabling audio recovery from longer distances than would otherwise be possible.
  4. Effective Black-Box Audio Recovery: Periscope successfully implements a training-free, black-box audio recovery scheme. This scheme, which relies purely on advanced signal processing techniques (DWT denoising and DBSCAN clustering), effectively overcomes the challenges of low SNR, environmental noise, and amplifier-induced distortions without requiring any prior knowledge or training data about the target device.
  5. High Intelligibility and Robustness: The recovered audio exhibits remarkable intelligibility, with Word Error Rates (WERs) as low as 7.44% (e.g., on a HW MateBook D14 at 100% volume). The audio is not only understandable to human listeners (STOI scores consistently above 0.7) but also accurately transcribable by online speech-to-text tools like Microsoft's API, and recognizable by music identification services like Shazam.
  6. Pervasive Vulnerability Across Diverse Scenarios: Periscope's effectiveness was validated across a wide array of attack scenarios. This includes 11 different mobile devices (6 laptops, 5 smartphones), 6 headphone models, varying EMR sensory distances (up to 1.05m direct sensing, 15-16m remote via Wi-Fi), different audio volumes (intelligible above 40%), diverse sensory angles (omnidirectional EMR), common physical obstacles (wood, plastic, concrete), varied environmental dynamics (office, subway, park), and different audio applications (YouTube, Zoom, Skype). This extensive evaluation confirms the broad applicability and robustness of the attack.

These key findings collectively establish Periscope as an alarming and practical threat, demonstrating that the very act of using headphones to secure audio privacy can, under certain circumstances, inadvertently create an enhanced pathway for sophisticated EMR-based eavesdropping.

Technical Deep Dive

Periscope's technical sophistication lies in its ability to identify, capture, and meticulously process weak, distorted electromagnetic radiations from mobile devices to reconstruct intelligible audio. This section details the underlying principles and the multi-stage recovery process.

Working Principle of Audio Circuits and EMR Generation

The foundation of Periscope's attack rests on understanding how audio is processed within a mobile device and how this process inherently generates EMR. Inside a mobile device's Printed Circuit Board (PCB), audio processing involves both digital and analog circuits. Digital circuits, including the CPU, memory, and Digital-to-Analog Converter (DAC), first decode audio files into digital signals. The DAC then converts these digital signals into an alternating analog current.

This analog current is subsequently routed through the analog audio circuits, primarily consisting of filters and amplifiers. Filters are responsible for removing unwanted noise, while the audio amplifier significantly strengthens the signal amplitude. This energized audio current then flows through the headphone jack and into the connected headphone, powering its voice coil to produce audible sound.

According to fundamental principles of electromagnetism, specifically Maxwell's equations and the Lorentz force law, these intense fluctuations of audio currents within the amplifier and associated circuits induce time-variant electromagnetic fields. These fields continuously radiate Electromagnetic Radiation (EMR) into the surrounding open space. The intensity and frequency characteristics of these EMRs are directly modulated by the audio signal, making them a potential side-channel for eavesdropping.

Feasibility Studies and EMR Characteristics

To validate the premise, the researchers conducted several feasibility studies:

  1. Correlation Analysis: Using a low-cost, ESP32-based sensory prototype (5.1cm long, < $10, 12-bit resolution, 10k samples/sec), EMRs were collected from various devices (Apple MacBook Pro 13, HW Mate 30, Lenovo Yoga 14s, OPPO R11st) playing speech from the Harvard corpus. A cross-correlation analysis between the original audio S(t) and the EMR readings E(t) revealed sharp peaks at lag=0 for all tested devices, confirming a strong linear correlation and the existence of audio leakage via EMRs. The prototype was found capable of restoring audio signals in the 0Hz to 5kHz range, sufficient for capturing intelligible human speech (85Hz-255Hz for vowels/consonants, 2kHz-4kHz for clarity).
  2. Headphone Impact: Experiments with a MacBook Pro 13 compared EMR spectrograms with functional headphones versus only the headphone plug inserted. While EMRs were present with just the plug (keeping the audio circuit active), the presence of the full headphone (wire and microspeaker) significantly enhanced EMR strengths. This led to the crucial finding that the headphone body acts as an antenna, boosting the device's EMRs and improving the SNR for eavesdropping, thereby extending the attack range.
  3. EMR Source and Amplifier Nonlinearity: To pinpoint the EMR source, a HW MateBook D14 was used to play a two-tone audio signal (f₁=130Hz, f₂=110Hz). The EMR frequency spectrum showed not only the original frequencies but also additional cross-product frequencies like 150Hz (2f₁-f₂), 410Hz (4f₁-f₂), and 610Hz (3f₁+2f₂). Similarly, a single 1kHz tone on a Lenovo Yoga 14s revealed harmonic components (e.g., 3kHz). These distortions (harmonics and cross-products) are hallmark properties of amplifier nonlinearity. The amplifier output S₀(t) can be modeled as a polynomial expansion of the input S(t) (S₀(t) = AS(t) + B₁(S(t))² + ...), where quadratic and higher-order terms generate these new frequencies. This confirmed the audio amplifier as the primary EMR source and highlighted the need to address these distortions for effective audio recovery.

Audio Recovery Scheme: Denoising, Extraction, and Generation

Periscope's black-box audio recovery scheme is a three-stage process:

A. Denoising

  1. Preprocessing:
  • Band-stop filters: To remove ambient EMRs, particularly 50Hz and its harmonics originating from electronic power cables, a sequence of band-stop filters is applied.
  • High-pass filter: A 20Hz high-pass filter eliminates low-frequency noises below the typical human hearing range.
  • Normalization: EMR measures are normalized to the range of -1 to 1 to improve signal quality.
  1. Environmental Noise Removal (DWT): To tackle random radiation noises from nearby electronics that often overlap with audio signals, Discrete Wavelet Transform (DWT) techniques are employed.
  • DWT decomposes EMRs into approximation coefficients (low-frequency components, large-scale characteristics) and detail coefficients (high-frequency components, including noise and fine details).
  • Approximation coefficients, crucial for speech intelligibility, are kept unchanged.
  • For detail coefficients, a multi-level wavelet decomposition is applied. A dynamic threshold method, based on the Birgé-Massart strategy, is recursively applied at each level to remove noise while preserving signal details. A three-level decomposition with a Daubechies Wavelet basis function is used.
  • The denoised signal is then reconstructed using inverse DWT. This significantly improves the SNR of the measured EMRs.

B. Audio Extraction (DBSCAN)

Even after denoising, the EMR signal still contains amplifier-induced distortions (harmonics, cross-products). Periscope leverages the observation that fundamental audio components typically form distinct, higher-amplitude groups in the spectrogram compared to distortions.

  • DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is applied to the EMR's Short-Time Fourier Transform (STFT) spectrogram. Each STFT bin (t, f, A) is treated as a point with time, frequency, and amplitude.
  • DBSCAN groups closely located points (bins) into clusters using Euclidean distance. These clusters are then categorized into audio signals, distortions, and white noise.
  • Audio signal clusters (C_audio) are identified by comparing their average amplitudes (Avgₙ). Clusters whose average amplitude is no smaller than γ times the maximum average amplitude (Amax) across all clusters are deemed audio signal clusters (C_audio = Cₙ Avgₙ ≥ γAmax). Empirically, γ = 0.1 is used.
  • Points outside the identified audio signal clusters (i.e., distortions and white noises) are neutralized by setting their amplitude to 0, effectively isolating the audio signal components.

C. Sound Generation

  • The processed EMR spectrogram, now clean and containing only the extracted audio components, is converted back into a temporal audio signal using the inverse STFT.
  • The resulting signals are then converted into standard .wav sound files using the Matlab signal processing toolbox.
  • Attackers can directly play these .wav files or input them into online speech-to-text recognition tools (e.g., Microsoft's Speech-to-Text API) for transcription, without needing specialized word/speech recognition models or device-specific training.

This intricate black-box approach, combining signal processing techniques like DWT and DBSCAN, allows Periscope to effectively recover intelligible audio from arbitrary mobile devices without prior knowledge, making it a highly potent and practical eavesdropping threat.

Demo / Proof of Concept

▶ Watch: Feasibility study: EMR reflection of audio sounds (6:00)

The practical implementation of Periscope showcases a miniaturized, stealthy system capable of both local EMR sensing and remote data transmission for audio recovery. The researchers provided a clear proof-of-concept through their prototype design and compelling real-world case studies.

Periscope System Design and Implementation

The Periscope prototype is designed for discreet deployment:

  1. EMR Sensor: The core sensing component is an ESP32 board, approximately 5.1cm long and costing less than $10. It is connected to a conductive wire to sense electric potential changes caused by EMRs. Configured for a 12-bit measurement resolution, it samples the received signal at 10k samples/sec, which is sufficient for capturing comprehensible audio.
  2. Raspberry Pi 4: This unit serves as the local processing and data forwarding hub. The ESP32 sensor connects to the Raspberry Pi 4, which stores the measured EMR data and handles preliminary processing if needed.
  3. Wi-Fi Access Point: A JDread 9600 portable access point provides Wi-Fi connectivity, enabling the EMR sensory device to transmit data wirelessly.
  4. Audio Recovery Server (Attacker's Laptop): A remote laptop receives the EMR measurements from the Raspberry Pi 4 via Wi-Fi. This is where Periscope's sophisticated audio recovery procedures (denoising, audio extraction, sound generation) are executed in real-time to output the recovered sounds.

The entire EMR sensory prototype (ESP32 + Raspberry Pi 4) costs approximately $85, making it an affordable and easily concealable tool. While the direct EMR sensing range is up to 1.05 meters, the Raspberry Pi's Wi-Fi transmission capability extends the effective attack distance significantly, allowing an attacker's laptop to be 15 to 16 meters away from the victim. This extended range and small footprint enable discreet deployment in various hidden locations, such as under a table, in a nearby bag, or even behind obstacles.

Real-world Case Studies

To demonstrate Periscope's practical efficacy, the researchers conducted two compelling real-world case studies, critically comparing its performance against a traditional hidden voice recorder:

Case 1: Study Room

  • Setup: A victim's iPhone SE2 was placed on a table, with the miniaturized EMR sensory device hidden underneath a 3cm thick wooden table, 25cm away, at a 30-45 degree angle. The audio volume was set to 40%. EMR measurements were continuously transmitted to the attacker's laptop, 5.5m away, for real-time recovery.
  • Results: Periscope achieved remarkable success. For phone calls and voice messages, STOI scores ranged from 0.71 to 0.73, and MOSNet scores from 2.45 to 2.63. Music segments had an average MOSNet of 2.57. These STOI scores, all above 0.7, indicate satisfactory intelligibility. Microsoft's speech-to-text API correctly transcribed speech, and Shazam successfully recognized music segments.

Case 2: Subway Cabin

  • Setup: A victim sat naturally, holding their smartphone. The EMR sensory device was concealed in a colleague's bag, 15cm away, at a 90-140 degree angle. Due to high ambient acoustic noise (94.8dB, compared to 60dB for normal conversation), the victim used 60% audio volume. The attacker's laptop was 6m away.
  • Results: Despite the challenging environment, Periscope performed exceptionally well. STOI scores ranged from 0.73 to 0.76, and MOSNet scores from 2.68 to 2.85. Again, these STOI values confirmed good intelligibility, with speech-to-text and music recognition tools successfully processing the recovered audio.

Comparison with Traditional Voice Recorder: In both scenarios, attempts to retrieve discernible audio using a concealed voice recorder proved entirely fruitless. This stark contrast highlights Periscope's strength: EMRs are inherently immune to acoustic noise and can penetrate common obstacles like wooden tables with minimal energy loss, whereas traditional recorders are severely hampered by acoustic isolation from headphones and high ambient noise levels. These case studies emphatically confirm Periscope as a highly effective and alarming threat, capable of recovering private audio remotely and covertly, even in challenging real-world conditions.

Defensive Implications

The revelations from Periscope underscore a significant privacy vulnerability that mobile device manufacturers and users must address. The researchers have taken proactive steps to disclose this threat and have identified effective countermeasures.

Disclosures and Industry Response

The Periscope team responsibly reported their eavesdropping threat to six leading mobile manufacturers: Apple, Lenovo, Huawei, Vivo, OPPO, and Dell. These disclosures included detailed explanations of the attack methodology, affected products, and comprehensive evaluation results. Encouragingly, Huawei has reproduced the attack results and is processing the report with high priority, indicating recognition of the severity. Lenovo has also committed to following up on the findings. The widespread nature of this vulnerability, potentially impacting a vast array of devices beyond those explicitly tested, led the researchers to file a vulnerability report to CVE (ID: CNVD-C-2023-85063), ensuring broader awareness within the security community.

Countermeasures

The most effective countermeasure identified to mitigate this EMR-based eavesdropping threat is the implementation of Electromagnetic (EM) shielding material within the device's audio processing circuits. Given that the primary frequency range of audio sounds is generally below 4kHz, readily available conductive materials like copper can be effectively employed.

To validate this, the researchers conducted an experiment by covering a Lenovo Yoga 14s laptop's audio circuits with a 1mm copper plate. The results were dramatic:

  • The Peak-Signal-to-Noise Ratio (PSNR) of the recovered speech plummeted significantly, from 24.5dB without shielding to a mere 4.25dB with shielding. This indicates a massive reduction in the quality and strength of the leaked EMR signal.
  • Correspondingly, the Word Error Rate (WER) for the recovered audio surged from an intelligible 15.23% to an almost incomprehensible 83.08%.

These findings unequivocally demonstrate that incorporating EM shielding, particularly around the audio amplifier and associated radiating circuits, can effectively impede EMR propagation and prevent the meaningful recovery of audio content. Manufacturers should consider integrating such shielding into their device designs as a standard security feature.

Potential Improvements for Attackers and Trade-offs

While the focus is on defense, the research also briefly touches on potential improvements for attackers. Increasing the EMR sensory range of the ESP32 board could be achieved by enlarging its antenna size or connecting it with a signal amplifier, following the Friis transmission equation. However, such enhancements would inevitably increase the overall size of the EMR sensory device, making it more challenging to conceal and more likely to be noticed by victims. This presents a critical trade-off between attack range and stealth, suggesting that while more powerful attacks might be technically feasible, they may sacrifice the crucial element of covertness.

Ultimately, Periscope serves as a stark reminder that physical isolation alone is insufficient for privacy, and engineers must consider the electromagnetic side-channels inherent in electronic design to truly secure user data.

Key Takeaways

  • Headphone-plugged mobile devices are vulnerable to EMR eavesdropping: Despite providing acoustic isolation, headphones do not prevent the leakage of private audio via electromagnetic radiation from internal audio amplifiers.
  • Audio amplifiers are the EMR source, and headphones act as antennas: The device's audio amplifier generates EMRs exhibiting characteristic nonlinear distortions. Crucially, plugged headphones significantly enhance these EMR strengths, making them easier to capture.
  • Black-box, training-free audio recovery is highly effective: Periscope's signal processing pipeline, utilizing DWT for denoising and DBSCAN for audio extraction from spectrograms, enables accurate and intelligible audio recovery without prior knowledge or training data for the target device.
  • High intelligibility and broad robustness across diverse scenarios: Recovered audio achieves WERs as low as 7.44%, is intelligible to humans (STOI > 0.7), and recognizable by speech-to-text tools. The attack is robust against varying devices, headphones, distances, volumes (effective above 40%), angles, common physical obstacles, environmental dynamics, and audio applications.
  • EM shielding is an effective countermeasure: Incorporating materials like a 1mm copper plate around audio circuits can drastically reduce EMR leakage, leading to a significant PSNR drop (from 24.5dB to 4.25dB) and a substantial increase in WER (from 15.23% to 83.08%), effectively preventing eavesdropping.
  • Challenges common privacy assumptions: This research highlights that relying solely on physical barriers (like headphones) for audio privacy is insufficient, necessitating a re-evaluation of security vulnerabilities in mobile device audio amplifier designs.

About the Speaker(s)

The talk was presented by Huiling Chen. Based on the provided transcript and metadata, further details regarding Huiling Chen's specific title or institutional affiliation were not included.

All talks from Network and Distributed System Security (NDSS) Symposium 2024