In reality, “HD” audio offers no benefit to the home listener.
24 bit, 192 kHz audio streams have been in use in recording studios for over 2 decades, but over the last several years or so, the format has been gaining popularity for use in home listening. However, it’s important for consumers to realise that, despite its ‘HD’ moniker, there is no ‘increased resolution’, and that the format offers no practical – nor perceptible – benefit for the listener.
The term ‘HD’ is deliberately misleading – when movies went ‘HD’, the increase in resolution resulted in a noticeable improvement in quality. So it seems to follow, then, that if you increase the ‘resolution’ of audio from the 16 bit, 44.1 kHz format stored on CDs, up to 24 bit, 192 kHz, naturally, the soundwaves will become more ‘detailed’ or ‘realistic’ than those held on CDs. So, of course, if you’re selling audio, you’ll want to give people the impression that we’re currently experiencing a resolution revolution, similar to the one previously seen with ‘HD’ movies. You may have even seen some convincing looking charts, showing ‘stepped’ digital waveforms, accompanying these claims.
Unfortunately, however, the digital representation of audio isn’t like it is with images or video – there are no pixels, and there is no ‘resolution’ to be increased.
To understand why these higher numbers don’t equate to higher resolution soundwaves, you should first understand what both of these numbers actually represent, and how analogue-to-digital audio conversion works, before I go into explaining why 16 bit, 44.1 kHz is all that humans will ever need.
So, let’s take a brief look at those numbers. In Pulse-code Modulation (PCM) audio (as used in CDs and most digital applications), the bit depth of an audio stream relates to its dynamic range – that is, the difference between the loudest part of an album or song, and the quietest. The difference between these two volumes is usually the greatest in symphonic music, while contemporary pop music is usually compressed to all buggery, and sits at more or less the same level for the whole album (the problems with doing this is best saved for another article).
The sample frequency (measured in kHz) of an audio stream is responsible for deciding the highest possible sound frequency apparent in a recording – it doesn’t, as most people assume, relate to the ‘resolution’ of the rebuilt soundwave. This warrants an immediate closer look, so let’s get back to bit depth later, and first focus on the way that analogue waveforms are represented in the digital realm, specifically, the Nyquist–Shannon sampling theorem responsible for this technology.
PCM vs DSD
It would be prudent to mention that the sampling method of encoding audio discussed here is what is used in Pulse Code Modulation (PCM) audio – a more or less ubiquitous standard, as it is the basis for CD audio, and most digital audio on PCs. The most notable non-PCM audio, however, is DSD audio (as used on SACDs), which uses Pulse Density Modulation (PDM). As such, the frequencies quoted for SACDs aren’t comparable to those on regular CDs or digital audio. You can dive into a more technical examination of PCM audio on Wikipedia.
Full-resolution audio at 44.1 kHz #
Sampling isn’t just about reinvigorating old disco records – it’s also the (admittedly, much dryer) process of converting an analogue signal into a numeric sequence. Crucially, for anything to be represented digitally, it has to be first converted into numbers.
The Nyquist–Shannon sampling theorem shows – in provable, mathematical terms – that any waveform containing no sound higher-pitched than B hertz can be completely (that is, fully and accurately) represented by a set of coordinates spaced 1/(2B) seconds apart. This translates to a sample frequency that’s double the maximum frequency of sound that needs to be represented.
For instance, if you want to represent a signal with a maximum frequency of 200 Hz, you require a sample rate of 400 Hz – that is, 400 numeric coordinates per second of sound.
What’s key to take away from this, is that any waveform – no matter how complicated – can be fully and accurately recreated by a set of coordinates that update at a frequency equal to double the frequency of the highest-pitch sound that you need to represent.
Given the well-established fact that the human ear can’t hear anything above 20 kHz (and most of us can’t even hear that high), and that the ‘standard’ audio held on a CD – stored at 44.1 kHz – can encode sounds up to 22.5 kHz, why would you need a 192 kHz audio file that can play back sounds ranging up to 96 kHz – nearly 5 times above the hearing ability of the human ear?
The answer, of course, is that you don’t.
In fact, given the way that amplifiers work, there’s a fair chance that inaudible sounds existing in the upper frequencies could have harmonic effects that will negatively affect the sound in the lower, audible frequencies, making ‘HD’ audio potentially sound worse than CD-quality.
Love 16 bits is all you need
#
So, now that we’ve established that 44.1 kHz PCM audio can fully and accurately represent the entire frequency spectrum of our hearing range, let’s take a look at bit depth. Now, if a waveform can be recreated by a set of coordinates that update at a specific frequency, how precise should those coordinates be? Ideally, the theorem requires infinite coordinates, but this isn’t practically possible. The precision of the coordinate is what is represented by bit depth, and it effectively defines the amplitude – or volume – of a waveform.
Strictly speaking, the bit depth actually defines the extent of the space that the waveform can exist in, which doesn’t so much define the maximum volume of a waveform (the volume knob on your amp defines that), but the maximum volume that it can represent, relative to the quietest. As mentioned earlier, this is known as its ‘dynamic range’.
A bit depth of 16 bits provides nominally 96 dB of dynamic range, according to the most common definition of 6x the number of bits. Why? Each bit doubles the number of possible values, therefore doubling the maximum possible volume. A doubling of volume translates to 6 dB, so 6 dB multiplied by the number of bits = total dynamic range.
But, how much dynamic range do we need? Surely, there are limits to both the quietest and loudest noise that we can hear? Indeed, there are. The human ear is surprisingly sensitive, being able to perceive noises as low in volume as around -8 dBSPL. To put this into perspective, and to paraphrase Monty from xiph.org, the hum from a 100W incandescent light bulb heard from a metre away is about 10 dBSPL, making it 18 dB louder. So, why have you never heard an incandescent light bulb humming to itself before? Because, most of the low-level noises that exist in the world are drowned out by incessant background noise.
Most broadcast studios sit at around 20 dBSPL, and these are exceptionally quiet environments, whereas most listening environments are closer to 50 dBSPL, or 30 dBSPL for headphones. We can safely say, then, that music will never get played in an environment where the background noise will be lower than 20 dBSPL. So, the practical noise floor (the level that defines the minimum perceptible signal) of the human ear, in terms of music consumption, actually sits at least 28 dB higher, and often 58 dB higher, than the -8 dBSPL that it can sense in isolated conditions.
There are upper limits to the volumes that the human ear can endure, too. Permanent hearing damage occurs at +130 dB in just seconds, and a jackhammer from a metre away is 100-110 dB, so there is no need for any sound on any recording (Despite what The Who believe) to go this far above the noise floor.
So, if we take the difference between quietest practical perceptible sound (aka the noise floor – around 20 db), and the loudest comfortable sound (around 110 dB), we have an idea of the absolute maximum range needed for a recording to capture every sound that we could hear. This translates to around 90 dB (of course, this range is far beyond any range needed for a recording).
Now, if we consider the possible dynamic range of 16bit audio, earlier stated as 96 dB, this is clearly above what is needed. So, if we have a recording that contains sounds at both the maximum level available in the range given to us by 16 bit audio, as well as the softest, and we set the volume of our amp to play the maximum sound at 110 dB (which is far louder than you’d ever want your home stereo), the recording could still play sounds as quiet as around 14 dB – below even the noise floor of a recording studio.
Strictly speaking, however, the bit depth of CD-quality audio actually provides an amplitude range even greater than 96 dB, thanks to dithering.
Because of the practical limits on the precision of the coordinate system, transferring an analogue waveform to digital does involve some amount of quantisation error (inaccuracies in capturing data). Dithering is a rather technical function that is designed to reduce the distortion caused by quantisation errors, but it also has the effect of lowering the noise floor (in this case, the quietest possible sound on a recording) and allow the bit depth to represent a dynamic range greater than its nominal 96 dB.
In fact, and without going into the technical details, dithering of a 16 bit recording allows a 120 dB difference between the loudest and the quietest possible sounds. This range is, in fact, far greater than what the human ear can perceive, and well over the even the most dynamic symphonic recordings, which sit at around 60 dB. To further expand on my earlier criticism of the compression of contemporary music, these recordings use as little as 12 dB of dynamic range.
To once again paraphrase Monty from xiph.org, 120 dB gives us a range so large that a jackhammer can be recorded alongside the mosquito that’s currently hiding behind your curtains.
Using any more than 16 bits only reduces the noise floor – the quietest perceptible sound – further, reducing it from completely imperceptible to completely imperceptible, but coming at the cost of storage space.
16 bits at 44.1kHz allows us to store sounds greater than our ears’ entire dynamic range – from imperceptible to deathly painful – as well as the entire range of frequencies that we could ever possibly hear – from the deepest bass to the shrillest keen – and, despite technology’s inexorable march forward, will be able to do so forever.
But my HD recording clearly sounds better!
Indeed, I thought the same thing, too. However, you should note that, often, the ‘HD audio’ versions of albums come from a newer master. So, when you compare your new 24/192 version of Dark Side of the Moon to your standard CD version, you’re actually listening to two entirely different versions of the album. Real world comparisons are notoriously difficult, but, before you can compare the difference between the formats, you really need to down-sample your ‘HD’ version down to 16/44.1 (using a high-quality re-sampler, like the SoX resampler in foobar2000), and use that version for the comparison. Better yet, scale this down-sampled version back up to 24/192, and see if you can tell the difference between the original and the new, ‘fake’ HD version, during a double-blind (A/B/X) test (there is a foobar2000 plugin that allows you to do this).
So, why 24/192 in studio environments? #
This article has tried to explain why 24 bit, 192 kHz audio offers no benefits to the home listener. Its popularity has perhaps come from its use in studio environments, where people have assumed that it is used there because it is of higher quality. There are definite practical advantages to using 24 bit 192 kHz audio in a studio environment, but ‘higher sound quality’ isn’t one.
Having a greater bit depth (translating to noise floor and dynamic range) allows the recording levels to be set with a little less precision than would otherwise be. If an instrument is accidentally recorded far too quiet, then bringing its volume up to the desired level won’t add perceptible noise, and if it’s recorded too loud, it’s less likely to cause clipping. For this reason, it’s also recommended to digitise vinyl at 24 bits, so you can set the levels perfectly afterwards.
A greater frequency response allows the recording to capture the full spectrum of sounds captured by the instruments, which can be of use, for example, when the recordings are to be sped up or slowed down.
Likewise, the broad frequency and lower noise floor also allow audio and mixing effects to be piled on top of each other, without introducing perceptible noise.
The extended range of 24 (or even 48) bits and 192 kHz acts, then, as a safety net for engineers, that maintains the purity of the sound throughout the many (often thousands of) applied engineering and mixing tweaks. Once the mix is complete, however, there is no need to keep all that empty space, and it should then be mixed down to the minimum size required for the human ear’s full spectrum: the tried and true 16 bits/44.1 kHz.
Keen to know more?
Most of what I covered here I learnt from Monty at Xiph.org’s excellent article entitled 24/192 Music Downloads …and why they make no sense. This article also goes into more detail on human hearing, dithering, amplifier distortion through harmonics, listening tests, some problems in contemporary music recording, and more, and also provides links to more sources and videos of practical experiments. Big thanks go out to Monty for opening my eyes (ears?) to the fallacy of the claims behind HD Audio.
A really concise overview of bit depth, dithering, and the noise floor (as well as a great discussion surrounding it) is available on the HeadFi forums at head-fi.org.
Finally, if you’re looking for more technical information on dithering (a topic I could only lightly cover here), part 1 of the paper Optimal Dither and Noise Shaping in Image Processing (courtesy of the University of Waterloo) offers a great technical overview.