Keep background audio 20 decibels below speech
Music under a voiceover has to sit well under it, and the number is 20 decibels. G56 is the mixing technique for that, and 20 decibels means the speech arrives about four times louder than the bed rather than a shade louder. W3C files it as sufficient for its low-background-audio rule, which is Level AAA, so this is a quality bar you choose rather than a legal minimum you owe. Brief sounds fall outside it. A sting, a notification, a single effect is not background audio in the sense the rule means. What it is aimed at is a continuous bed running under speech for the length of a video, which is exactly where the difference between mixing and habit shows.
How we find it in an audit
This is measurement rather than opinion. We take the level of the background between phrases, take the level of the foreground speech, subtract one from the other, and check the gap is 20 or more. Explainer videos, product tours and testimonials are the usual candidates, because all three come off the same template with a bed underneath. Where the gap falls short, the finding carries the measured number, since a mix engineer can act on a number and cannot act on the word muddy.
How affected users experience it
Hearing loss does not turn the volume down evenly. It flattens the difference between the sound you want and the sound you do not, so a bed a few decibels down stops being background and becomes competition. Turning everything up brings the music along with it, which is why louder is not a fix. Anyone in a noisy room meets a milder version of the same problem, and anyone following speech in a second language meets another.
Passes vs. fails
Passes
The same explainer has the music mixed more than 20 decibels down, so the atmosphere survives and every word stays intact.
Fails
A product explainer runs its music bed 5 decibels under the voiceover, so half the words blur into it.
This guide is our interpretation of W3C technique G56: Mixing audio files so that non-speech sounds are at least 20 decibels lower than the speech audio content. W3C publishes its techniques as guidance rather than as the standard, and says so on every one of them. The success criterion is what conformance is measured against, and a technique is one documented way to meet it.