Craft
A narrator and a music bed cannot share one loudness number
AES TD1008 cites listening panels that heard speech 2 to 3 dB louder than music at the same measured loudness, so its target moves with the speech fraction.
The Audio Engineering Society's TD1008 is dated 24 September 2021, and one sentence in it invalidates the way almost every marketing video gets mastered: formal tests with listening panels showed that speech normalized to the same BS.1770 Integrated Loudness as music is typically perceived 2 to 3 dB louder than the music. So it sets two targets. Speech at -18 LUFS, track-normalised music at -16 LUFS. A narrated video is both at once, so TD1008 hands over a formula rather than a number: Distribution Integrated Loudness = -16 - [2 x (SpeechPercentage / 100)] LUFS. Your target is a function of how much of the runtime is somebody talking.
Equal measured loudness is not equal heard loudness
TD1008 carries a document number, AESTD1008.1.21-9, a date of 24 September 2021, and a writing group of eleven that includes Scott Norcross, who wrote one of the two papers it cites for the 2 to 3 dB. It came out of the AES Study Group on Streaming Audio Loudness, co-chaired by Bob Katz and David Bialik, and supersedes TD1004. Not a forum opinion about ducking a bed.
The 2 to 3 dB is a citation rather than an assertion, and the two papers behind it disagree with each other and with TD1008. Scott Norcross of Dolby, paper 10448 at AES Convention 149 in October 2020, reports that subjects on average tend to match music content 3 to 4 dB higher than the speech-based reference signal when measured similar in loudness by ITU-R BS.1770. Fabian Begnert, Hakan Ekman and Jan Berg, paper 8489 at Convention 131 in 2011, normalised five types of broadcast program material to equal meter measured loudness, had subjects match them by ear, and put the maximum difference at 2.82 dB. Neither abstract page gives a sample size. Two tests, nine years apart, in a band from 2.82 to 4 dB, and the standard writes 2 to 3.
Look at what BS.1770 was calibrated against. ATSC A/85 records the ITU tests behind it: a three-member panel selected 48 test sequences of broadcast material, and a reference sequence consisting of English female speech was chosen to establish a target loudness level. Each was repeated at two amplitude levels, giving 96 monophonic sequences to match, and 97 subjects participated at five different test sites. A verification round used 20 subjects on the same 96. The anchor was a voice both times.
Every LUFS figure a platform quotes is a BS.1770 reading, and BS.1770 was calibrated against a voice.
TD1008 does not give you a target, it gives you a function
Table 1 puts speech at -18 LUFS with a +1 LU upper tolerance and track-normalised music at -16 LUFS with +0.2. The reasoning is stated: speech within streams is normalized to -18 LUFS to accommodate the gain of current player devices, and it is therefore additionally recommended that music be normalized to an average of -16 LUFS in operations where music and speech are separately normalized and played out automatically.
Then the document does what no platform spec does: it admits most real content is a mixture. A stream carrying speech at -18 LUFS and track-normalised music at -16 will measure between the two, and this depends on the relative proportion of speech and music in the stream. Hence the formula: Distribution Integrated Loudness = -16 - [2 x (SpeechPercentage / 100)] LUFS. The heading above it is No Separable Speech Anchor: a provider who cannot normalise speech and music separately. One rendered file with one mix is that case.
Run it on a real brief. A 60 second video carrying 45 seconds of narration is 75 per cent speech, so the target is -17.5 LUFS. Table 2 puts format names on three other points of that line: News/Talk and Drama at -18 LUFS, Mixed format and Sports at -17, Pop music at -16. TD1008 says the arithmetic need not be tight: because approximately 1 LU is considered a just noticeable loudness difference, this formula will still provide satisfactory results when the percentage of speech is roughly approximated. So counting the seconds of voiceover and dividing is the whole method. One row of Table 1 short-circuits even that for advertising: Interstitial content is set flat at -18 LUFS, 0.5 LU from what the formula gives this example, and it points the same way.
Plotted from the formula in AES TD1008, dated 24 September 2021: Distribution Integrated Loudness = -16 - [2 x (SpeechPercentage / 100)] LUFS. LUFS is a frequency-weighted, level-gated measurement of average power relative to digital full scale, not an acoustic measurement. Table 2 of the same document names formats on three of these points: Pop music at -16, Mixed format and Sports at -17, News/Talk and Drama at -18. Those percentages are the formula's own inverse, not a measurement of any format's speech content. The 75 per cent point is a declared example, 45 seconds of narration in a 60 second video, not a measured average of anything.
Apple ships the same premise as code, with the hinge at 17 and 25 per cent
Apple's HTTP Live Streaming documentation has a page called Adjusting anchor loudness, and its opening line states the same problem: adjust anchor loudness when measurements of speech-gated loudness for a full mix may be inaccurate, such as when speech activity is low. It takes four inputs: speech-gated loudness, speech activity from a speech detector, program loudness per ITU-R BS.1770-4, a revision behind the BS.1770-5 approved in November 2023, and the top value of loudness range per EBU Tech 3342.
The published Swift listing names its own thresholds in comments: ignore the speech loudness below a speech activity ratio of 0.17, keep it above 0.25. Below 17 per cent speech the model discards the speech measurement and computes anchor loudness from the other three; above 25 per cent it uses the speech measurement unmodified; in between, a gradual transition from the model to the measured, speech-gated loudness occurs as the speech activity increases. Speech activity is speech-gated duration over content duration.
The two mechanisms have nothing in common. AES draws a straight line from 0 to 100 per cent speech and publishes it as algebra. Apple fits a statistical model to a corpus it describes only as a large variety of long-form content, publishes the coefficients, and crossfades over an eight point window. What survives both is a premise no live platform spec carries.
The share of the runtime that is somebody talking is a parameter of the loudness target, not a detail of the edit.
Europe forbids exactly what the AES recommends, and both documents are current
EBU R 128, at V5 since November 2023, states in item (l) that the audio signal shall generally be measured in its entirety, without emphasis on specific foreground elements such as speech, music or sound effects. Supplement 1, the short form document, repeats that sentence, sets the target at -23.0 LUFS with a 0.2 LU tolerance and caps short-term loudness at -18.0 LUFS. Short form there means a programme typically shorter than 30 seconds, and the list names advertisements, promos, interstitials, stingers and bumpers. A marketing video is on it.
ITU-R Report BS.2434-0, from October 2018, tabulates the position across four regions. For short form content in North America under A/85, in Europe under R 128, in Japan under TR-B32 and in Australia under OP-59, anchor element measurement is Not permitted and full program mix measurement is Recommended. For long form the table flips, and A/85 recommends the anchor while permitting the mix only conditionally.
So one document tells you to derive your single number from the speech fraction, and four regional recommendations tell you not to measure the speech at all. The broadcast rule exists so a thirty second spot does not jump out of a programme stream. A/85 names the cost of its own choice: the louder elements of this type of material will increase the loudness measured with a long term integrated method, and consequently reduce the perceived Anchor Element loudness after normalization.
R 128 Supplement 2, from November 2023, lets the Distribution Loudness Level sit in the range of -20.0 to -16.0 LUFS where no metadata manages device gain, which puts the AES speech target of -18 in the middle of the EBU's own streaming range.
One document says derive your single number from the speech fraction. Four regional recommendations say do not measure the speech at all.
Every live spec you can actually read asks for one integrated number
Spotify's audio ads spec asks for an overall loudness of -16 LUFS, an integrated average with a 1.5 LUFS tolerance and a True Peak limit of -2.0 dBTP, so that the ad volume is consistent with content. Apple Podcasts: the overall loudness remains around -16 dB LKFS, with a 1 dB tolerance. Google's Assistant guidance: the average loudness should be -16 LUFS for stereo audio content, and -19 LUFS for mono, a 3 LU step that is its own trap.
Three specifications for speech over a bed, three single integrated numbers, and not one reference to how much of the runtime is speech. Apple asks podcasts for -16 LKFS while TD1008 puts speech at -18 and calls -16 the music figure. Neither is wrong. They are two bodies picking different points inside one argument.
The -14 LUFS that circulates in video advice belongs to none of this. It is Spotify's music normalisation figure, one of three settings a listener picks between with Loud at -11 dB LUFS and Quiet at -19 dB LUFS, and that page says nothing about spoken word at all.
Two LU sounds like nothing until it is priced against a hand on the volume control. A/85's Annex E puts the Comfort Zone, the loudness change a sample of listeners found acceptable, at +2.4 dB to -5.4 dB, and says a gain increase of two to three dB is enough to move a typical program out of it. TD1008 then draws its own boundary: it is not intended for sound-with-picture content, and it points to AES71-2018, which per the ITU report sends online video back to -24 LKFS internationally and -23 LUFS in Europe. So the -17.5 does not transfer to an OTT delivery spec. The premise does, and of everything quoted here only TD1008 and Apple's model move their answer when the narration does.
What to do about it
- Measure the narration stem on its own and put its LUFS figure in the brief, because a full mix reading is a different quantity with the same unit.
- Set the target from the speech fraction: -16 minus 2 times the share of the runtime that is voice, which puts a 60 second video with 45 seconds of narration at -17.5 LUFS.
- Decide up front whether the delivery is a broadcast spec or a phone, because -23.0 LUFS and -18 LUFS are 5 LU apart and one master cannot serve both.
Questions people actually ask
what lufs should a marketing video be?
AES TD1008 makes it depend on how much of the video is narration: -16 LUFS minus 2 times the speech fraction. A 60 second video with 45 seconds of voiceover is 75 per cent speech, which gives -17.5 LUFS integrated, with the maximum true peak held at -1 dBTP. If the piece is advertising or promotional material, Table 1 of the same document sets Interstitial content flat at -18 LUFS. If the delivery is a broadcast specification rather than a web player, EBU R 128 asks for -23.0 LUFS instead, and explicitly without emphasis on the speech.
is -14 lufs the right target for a video with a voiceover?
-14 LUFS is Spotify's music normalisation figure, and Spotify's own page presents it as one of three listener settings alongside Loud at -11 dB LUFS and Quiet at -19 dB LUFS. No loudness document in this area offers it as a target for speech over a music bed. AES TD1008 puts speech at -18 LUFS and track-normalised music at -16 LUFS, so a narrated mix belongs between those two, not above either.
why does music sound quieter than speech at the same lufs?
Listening panels put the difference between 2 and 4 dB. AES TD1008 reports that speech normalised to the same BS.1770 integrated loudness as music is typically perceived 2 to 3 dB louder than the music; Norcross measured subjects matching music 3 to 4 dB higher than a speech reference, and Begnert, Ekman and Berg found a maximum of 2.82 dB across five programme types. BS.1770 itself was calibrated against a reference sequence of English female speech, which is why a voice reads as the loud thing.
should i measure loudness on the voiceover or on the finished mix?
Both, because they answer different questions. TD1008 defines Speech Loudness as the integrated loudness of pure spoken voice that is not mixed with other elements, so that one has to be taken on the stem, while the number a distributor checks is the integrated loudness of the whole mix. For short form content, ITU-R Report BS.2434-0 records that all four broadcast regions call anchor element measurement Not permitted and full program mix measurement Recommended.
how do i work out the speech percentage of a video?
Divide the seconds of narration by the runtime, and TD1008 says roughly is good enough: because approximately 1 LU is considered a just noticeable loudness difference, the formula still provides satisfactory results when the percentage of speech is roughly approximated. Apple's HLS anchor loudness model reaches the same conclusion by another route, discarding the speech measurement below a speech activity ratio of 0.17 and trusting it above 0.25.
Sources
- Audio Engineering Society, AESTD1008.1.21-9, Recommendations for Loudness of Internet Audio Streaming and On-Demand Distribution
- ATSC, A/85:2013, Techniques for Establishing and Maintaining Audio Loudness for Digital Television
- ITU-R, Report BS.2434-0, Loudness in Internet delivery of broadcast-originated soundtracks
- EBU, R 128, Loudness normalisation and permitted maximum level of audio signals
- EBU, R 128 s1, Loudness Parameters for Short-form Content
- EBU, R 128 s2, Loudness in Streaming
- Apple Developer Documentation, HTTP Live Streaming, Adjusting anchor loudness
- AES E-Library, Norcross, Using ITU-R BS.1770 to Measure the Loudness of Music versus Dialog-based Content
- AES E-Library, Begnert, Ekman and Berg, Difference Between the EBU R-128 Meter Recommendation and Human Subjective Loudness Perception
- ITU-R, Recommendation BS.1770, Algorithms to measure audio programme loudness and true-peak audio level
- Spotify Advertising, Audio ads specs
- Apple Podcasts for Creators, Audio requirements
- Spotify for Artists, Loudness normalization
- Google, Assistant developer tools, Audio loudness
Every figure on this page comes from one of these. Where two of them measure the same thing differently, the article says so rather than picking the flattering one.