Research & Ideas / 2026-06-22 / 14 min read

How 3D Surround Sound Makes Sound Appear Outside the Head

Headphones still have only left and right drivers, yet sound can appear above, behind, or across a room. The explanation lies in the cues the brain uses for direction and how a system recomputes them as the head moves.

  • Audio
  • Acoustics
  • Technical principles

The moment that best reveals the difference between spatial-audio systems is not a sound circling the head once. It is what happens after the listener turns: does the source turn with the headphones, or remain where it was?

A voice can stay attached to the screen while room ambience remains at its original direction. The headphones move with the head, but the scene seems to belong to the room. The two ears are still receiving two signals; the brain has simply accepted a source that is not physically in the room.

The better question is therefore not how headphones imitate many loudspeakers, but which cues real sound gives the brain and how many of them a system reproduces.

Why two drivers can be enough

We mainly use two ears to locate sound in the real world. A sound from the left normally reaches the left ear first, while the head weakens some high frequencies before they reach the right ear. These differences are the interaural time difference (ITD) and interaural level difference (ILD).

They are useful for left-right judgments, but cannot explain the whole spatial image. A sound in front, behind, or at some vertical angles can produce similar binaural differences. If hearing only compared which ear received a sound earlier or louder, we would often confuse behind with in front.

The pinnae, head, and torso add the missing information. Before a sound enters the ear canal, these shapes reflect and block it. Different directions create different spectral changes. The brain learns from experience that one pattern tends to come from above and another from behind. The body itself is part of the positioning system.

HRTF: recording the geometry of hearing

The head-related transfer function (HRTF) describes how sound changes while travelling from a direction in space to the two ears. It includes changes in timing, phase, and spectrum, not just a louder left channel.

In a measurement, loudspeakers play test signals from many directions while miniature microphones near the ear canals record the response. The result is a map of directions. If a sound is filtered with the two responses belonging to a target direction and sent to headphones, the brain may interpret it as coming from there.

In digital audio the process is often expressed with impulse responses and convolution:

left output = source * left-ear directional impulse response
right output = source * right-ear directional impulse response

Here * means convolution. It gives each moment of the sound the character of a propagation path. HRTF describes direction; a room impulse response can add wall reflections, reverberation, and scale. Together they form a more complete auditory scene.

This is why ordinary left-right panning differs from binaural spatialization. Panning moves the image between the ears; HRTF tries to move it into a space made of front-back, up-down, and near-far relations.

Why front-back and up-down are easy to confuse

Left and right have relatively stable time and level differences. Front-back and vertical judgments depend more on spectral cues from the pinnae, and every pair of ears is different.

A generic HRTF temporarily lends the listener someone else’s ears. A filter that is accurate for the person measured may place peaks and notches in the wrong locations for someone else. The result can be an image inside the head, a front-back reversal, or a strange timbre instead of a source overhead.

That does not mean the listener cannot hear properly, nor necessarily that the headphones are poor. Spatial positioning is produced jointly by body shape and learned experience. Personalization tries to reduce the error of making everyone share one standard pair of ears. Camera-based scans of the head and ears are one practical approximation, even though they are not the same as measuring a complete HRTF in a laboratory.

Turning the head is the stricter test

If a virtual source is placed at the screen but rendering is fixed only relative to the head, turning makes the sound turn too. It stays in front of the listener but slides away from the screen in the room.

Head tracking estimates head pose and updates the relative direction between listener and source. When the head turns left, a source fixed in front of the room moves to the right in head coordinates, so the brain continues to hear it at the screen.

The difficulty is continuity. A correct but slow update makes the field drag behind; noisy pose data makes the source wobble; variable latency makes hearing and balance disagree. A static second can hide problems that several seconds of turning reveal: sensors, pose estimation, coordinate transforms, audio buffering, and real-time convolution all have to remain stable.

Useful tests include slowly turning while checking whether dialogue remains on the screen, closing the eyes and returning to the original pose to detect drift, and comparing fixed spatial audio with head-tracked audio. These separate externalization, HRTF matching, and dynamic coordinate consistency instead of reducing everything to “it sounds immersive.”

Distance is not just lower volume

Distance changes more than loudness. Direct sound becomes weaker relative to reflections, high frequencies may be absorbed by air and obstacles, and angular changes become smaller as a source moves away. Near sources produce stronger head-shadow and low-frequency differences.

If a system only turns down a distant sound, the result may be quiet but still stuck to the ear. Direct sound, early reflections, reverberation, spectral change, and room scale have to be handled together.

That is why spatial audio in games and films is more than a headphone effect. Content may be multichannel or made of audio objects with coordinates. The renderer needs the relation between source, listener, and environment. The final signal still has two waveforms, but they are the compressed output of a changing three-dimensional scene.

A small experiment without special equipment

Find a reliable binaural recording and listen with ordinary stereo headphones. With your eyes closed, judge left-right, front-back, and distance. The recording already contains directional cues from a dummy head or human ears, so it can produce an external image even without tracking.

Keep the audio unchanged and turn your head. The recorded field cannot recompute itself, so every source turns with you. Then repeat the movement with content that supports dynamic head tracking. Notice whether the source leaves the head, whether front and back are confused, and whether the position stays stable.

Many “8D music” tracks are only automated panning plus reverb. They can make sound circle the head, but they do not represent complete three-dimensional localization. Recording method, platform processing, and duplicate spatialization can all change the result.

The useful questions begin here

There is no single answer to how two headphone drivers simulate three-dimensional space. The problem includes binaural and pinna cues, HRTF measurement and personalization, reflections and distance, head tracking and world coordinates, and the latency needed to run the chain in real time.

Two questions remain especially useful: how much personalization is enough, and whether the brain can learn an initially mismatched HRTF after prolonged use. A good endpoint for this kind of interest research is not pretending that everything has been explained, but turning vague surprise into questions that can be checked, heard, and tested.

Sources