You're mid-sentence on a phone call when the little ring on your shelf lights up. Nobody said the wake word. Nobody even came close. And yet there it is, patiently waiting for a command you didn't mean to give.

So does this thing actually know what it's listening for?

It does. Imperfectly, fascinatingly, through a chain of acoustic sleight-of-hand that rewards about five minutes of your attention.

What the microphone is actually measuring

A wake word detector isn't doing speech recognition the way you'd imagine. It's not transcribing audio and checking whether the words spell out "Hey Siri" or "Alexa." Too slow. Too power-hungry to run continuously on a device the size of a hockey puck.

Instead, it's doing something closer to fingerprint matching.

The audio gets broken into tiny overlapping slices, each about 25 milliseconds long. Each slice becomes a spectrogram: a heat map of which frequencies are loud at that moment. Those spectrograms feed into a small neural network trained on thousands of hours of people saying the wake word across different accents, moods, and rooms. The network has learned what the shape of that phrase looks like in frequency space. Not words. A sound silhouette.

This is why the detector runs on a chip that draws less power than a nightlight. It's a narrow classifier, not a general listener, and that distinction matters more than most people realise.

The liveness problem, and the tricks used to solve it

Here's the part that actually earns its keep. A recording of you saying "Hey Google" produces spectrograms that look almost identical to you saying it live. Same words, similar frequencies. So why doesn't a phone playing your voice trigger your speaker every time?

Several things work together, and none of them is foolproof alone.

Room acoustics leave fingerprints. When you speak in a room, your voice bounces off walls, furniture, and the floor before hitting the microphone. Audio played through a phone speaker gets its own set of room reflections added on top of an already-recorded signal. The detector has learned, roughly, that doubled-reflection audio smells wrong. A voice that's passed through two acoustic environments carries a kind of double-exposure blur, like a photocopy of a photocopy, that a directly spoken word simply doesn't have.

Speaker frequency response is a giveaway. No consumer speaker reproduces sound flat across the full frequency range. A phone speaker rolls off sharply below about 200 Hz and adds its own coloration in the midrange. Live human speech has a bass presence and a low-frequency warmth that a small speaker physically cannot reproduce. The neural network, trained on enough replayed audio, has learned this coloration as a soft signal for "this came out of a speaker, not a mouth."

Multi-microphone arrays do spatial math. Many smart speakers use two, four, or even seven microphones in a ring. Sound from a real person hits each microphone at a slightly different time. The device calculates the direction. Audio blasted from a TV across the room arrives from a fixed, distant point with a different delay signature than a person standing nearby. Amazon's Echo devices have used this beamforming approach since the first generation, and it's one reason a far-field device is genuinely harder to spoof than a single microphone.

Consider two people who bought the same smart speaker on the same day. Priya keeps hers in a small home office and always addresses it directly from about a metre away. Marcus keeps his in an open-plan living room with a large TV two metres behind it. Marcus gets three or four false triggers a week. Priya gets almost none. Same device, same software, completely different acoustic geometry.

The part most people have backwards

The standard assumption is that these systems are detecting you specifically, like a voiceprint lock. They're not, at least not by default. A wake word detector doesn't know if it's your voice or your flatmate's. It only knows whether the sound matches the target phrase's acoustic pattern well enough. Voice Match features (Google's term for speaker-specific personalisation) are a separate layer that runs after the wake word fires, not part of the trigger itself.

And honestly, the security implications of that are underappreciated. A sufficiently high-quality recording, played through a good speaker in a quiet room, can absolutely fool these systems. Security researchers have demonstrated this repeatedly. The defenses above raise the bar. They do not build a wall.

Found your speaker waking up too often? Check what's behind the microphone. A TV playing voices at an angle that mimics a person's position in the room is the single most common culprit, and moving the speaker a metre to one side often fixes it completely.

The technology is genuinely clever. It's also, at its core, a probability threshold with a microphone attached. The dial between sensitivity and false alarms never fully satisfies either side, and every engineer who's shipped one of these things knows it.