← All articles
Guides

The Complete Guide to Screen Recording Audio: System Sound, Microphone, and Mixing

Audio quality can make or break a screen recording. Learn how to capture system sound, mic input, and mix them perfectly for tutorials, demos, and presentations.

Audio Is Where Most Screen Recordings Fail

You’ve watched it before: a tutorial where the visuals are fine but the audio is hollow, clicky, echo-laden, or — worst of all — completely absent. The viewer gives up not because the content was bad but because understanding it requires too much effort.

Audio problems in screen recordings almost always trace back to setup, not equipment. A decent USB microphone in a treated room will always beat an expensive condenser mic in a reverberant one. Understanding how your system captures audio is the first step toward getting it right.

The Three Audio Sources in a Screen Recording

System audio is everything the computer plays through its speakers or headphones: app sounds, browser video, notification chimes, music. Capturing this alongside the screen gives your recording the same audio context the viewer would have if they were watching over your shoulder.

Microphone input is your voice, narrating what’s on screen. It arrives from an external microphone, a built-in laptop mic, a headset, or an audio interface.

Mixed output combines both. This is what most tutorial recordings aim for: your voice explaining the screen, with the system audio low in the mix (or muted for sections where it would be distracting).

Capturing System Audio on Each Platform

macOS doesn’t expose system audio to capture by default — a deliberate privacy decision. You need a virtual audio device like BlackHole or Loopback to create a channel that routes system audio back into your capture software. ScreenKit handles this automatically when system audio capture is enabled; it installs the necessary audio routing in the background.

Windows exposes system audio natively through WASAPI loopback. Any capture software can access it without additional drivers. In ScreenKit, enable “System audio” in the recording settings and select “What you hear” as the source.

Linux depends on your audio subsystem. On PulseAudio, a monitor source for the default sink gives you system audio; on PipeWire, a similar mechanism is available through the wireplumber routing layer.

Microphone Setup: The Variables That Matter Most

Distance. Stay 15–25 cm from the microphone. Closer than that and you pick up mouth sounds and breathing; farther away and the room acoustics start to dominate.

Room treatment. Hard surfaces reflect. Bookshelves, curtains, upholstered furniture, and acoustic foam absorb. Recording in a closet full of clothes produces surprisingly good results — the fabric absorbs reflections that would otherwise make your voice sound like it’s in a bathroom.

Gain staging. Set your microphone gain so your voice peaks around −12 dBFS and never exceeds −6 dBFS. Peaks above that clip the signal; levels far below it force you to amplify during editing, amplifying the noise floor along with the voice.

Mixing the Two Sources

When both system audio and microphone are captured simultaneously, you’re working with two separate audio tracks. Most screen recorders save them as separate tracks so you can adjust them independently in post.

A useful starting point: set the microphone track to 0 dB and the system audio track to −12 dB (about 25% of its original volume). This keeps your voice dominant while preserving the system audio as context rather than competition.

For sections where you’re playing video or demonstrating audio-heavy content, fade the microphone down or mute it entirely. Viewers find it easier to follow one thing at a time.

The One Setting Most People Miss

Monitor your microphone while you record. Put on headphones and listen to what you’re actually capturing — not what you assume you’re capturing. Background noise that’s invisible on a waveform can be obvious in headphones. A USB microphone that’s accidentally picking up fan noise from a laptop vent becomes apparent immediately when you hear yourself before you hit record.

Five seconds of monitoring before each session prevents the frustration of finishing a twenty-minute recording and discovering the audio was unusable throughout.