The Crowd You're Not Looking At
You're sprinting through a packed football stadium, chasing a cutscene trigger. The stands heave with thousands of fans waving, chanting, doing the wave. It looks alive.
None of them exist.
Not in any meaningful computational sense. The moment you turned toward the pitch, most of that crowd stopped being simulated and became something closer to a very convincing painting. The engine made that call automatically, without asking you, in a fraction of a millisecond. And it was absolutely the right decision.
This is the invisible architecture underneath every crowd in every modern game: a tiered system of attention, deception, and ruthless prioritization. Once you see it, you can't unsee it in any busy street, packed arena, or festival scene you sprint through.
The Budget Problem No One Talks About
A single fully simulated NPC in a modern open-world game costs real processing power. Give it pathfinding, collision awareness, animation state tracking, audio reactivity, and the ability to respond to the player, and you're burning CPU cycles that could be doing something else. Something like rendering the building behind it, calculating your character's hair physics, or keeping the frame rate from collapsing.
Now multiply that by five thousand stadium fans.
The math doesn't work. It has never worked. So engines don't try.
Instead, they run a tiered attention budget sorted by proximity and visibility, operating roughly like a hospital triage system, except the urgency criteria is: how much will the player notice if this thing is wrong?
The NPC standing two meters in front of you, about to hand you a quest item, gets full simulation. Full pathfinding, full animation blending, full audio. The one forty meters away, partially occluded by a market stall, gets a simplified loop. The crowd filling the plaza behind that stall gets a particle system wearing a skin. They aren't NPCs at all. They're a visual effect with legs.
Three Tiers, One Illusion
Most engines converge on something like three distinct simulation levels, even if different studios name them differently.
The first tier is what developers sometimes call "full agents." These NPCs have real brains: behavior trees or utility AI systems that weigh options, react to stimuli, and update every frame. The player can interact with them, and the engine expects that. At any given moment in a dense urban scene, there might be a dozen of these running simultaneously. Not many more.
The second tier is scripted loopers. They follow pre-baked animation cycles, a market vendor gesturing at invisible produce, a pedestrian pacing a fixed short path. They respond to almost nothing. Walk into one, and the engine might play a collision response animation and stop there. They cost maybe a tenth of what a full agent costs.
The third tier is the crowd shader, the particle crowd, the impostor field. These are sprites, animated textures, or geometry instances that move in convincing but completely non-physical ways. They don't pathfind. They don't collide. They don't know you're there. In Assassin's Creed Unity, Ubisoft rendered crowds of what appeared to be thousands in revolutionary Paris using exactly this kind of layered impostor system, with full agents only materializing as players descended from rooftops toward street level.
The transition between tiers is the trick. Do it wrong and you get the infamous "NPC pop-in," a flat sprite snapping into a three-dimensional person as you approach. Do it right and the player never consciously notices that the crowd just gained sentience.
The Visibility Test Running Thirty Times Per Second
Proximity alone doesn't determine tier assignment. Visibility does.
Most modern engines run a continuous occlusion query: a rapid check asking whether a given object is actually in the player's view frustum (the cone of what the camera can see) and whether geometry is blocking it. An NPC standing directly behind a concrete pillar, ten meters away, might get demoted to tier two or even tier three instantly, because the engine has calculated that you cannot see it and are unlikely to see it in the next few frames.
This is why crowd simulation is fundamentally different from NPC simulation. Individual named characters need persistent state regardless of visibility: if a shopkeeper walks home at 6pm, they should be home whether or not you watched them leave. Crowd members are anonymous. No persistent state to maintain. The engine can safely despawn and respawn them, swap their tier, without breaking any narrative contract with the player.
Here's how that plays out in practice. You're in a dense city RPG and you duck into an alley. In the 0.3 seconds your camera faces a brick wall, the engine has already downgraded fourteen pedestrians behind you to scripted loopers and dissolved six more entirely. When you spin around, they're back, quietly upgraded, because the engine predicted your turn using your camera velocity and pre-promoted them a fraction of a second early. That pre-promotion window, sometimes called a "promotion buffer," is what separates smooth crowd systems from janky ones. It's a small thing. It's everything.
What People Assume Crowds Are Doing
Here's an assumption worth burning down: most players believe that NPCs in crowds are, on some level, living their own lives offscreen. That the vendor is still haggling with someone even when you're two districts away. That the stadium crowd is still chanting when you're in the locker room.
They're almost never doing that.
Offscreen crowd members in most games are either suspended entirely or running the absolute cheapest possible state update: a position tick and an animation frame increment, nothing more. Simulating the behavior of entities the player cannot observe and will never verify is wasted computation by definition. That's not a cynical shortcut. It's correct engineering.
Two players bought the same open-world game. One of them spent an hour trying to "catch" background NPCs doing something inconsistent by fast-traveling and spinning the camera quickly. He found several. The other played for sixty hours and never noticed anything wrong, because she was engaged with the game, not auditing it. The system isn't designed to survive forensic scrutiny. It's designed to survive normal play. For sixty-hour players, it works perfectly. The forensic ones are basically testing a magic trick by grabbing the magician's wrist, which, fair enough, but that's not what the trick is for.
The Sound Half of the Equation
Visual tiers don't run alone. Audio follows the same logic, and it matters more than most players realize.
A crowd sounds alive partly because of procedural audio layers: ambient crowd noise that scales with the size and excitement level of the simulated group, blended with occasional specific sounds (a laugh, a shout, a cough) triggered probabilistically rather than attached to any individual NPC. You hear a crowd. You don't hear four thousand individual voices.
When the visual tier drops a group of NPCs to impostors, the audio system usually doesn't change at all. The sound continues, sourced from the general crowd emitter, not from any individual. Audio is the crowd's alibi. Even when the visuals are a texture atlas doing a looping shimmy, the sound makes your brain file the whole thing under "real."
It's a little like how a restaurant feels busier when you can hear the kitchen.
The Frame Rate Is the Point
All of this exists for one reason: to give you a smooth, immersive sixty frames per second without making the machine catch fire.
The sophistication of crowd simulation in modern games isn't really about making NPCs smarter. It's about making them seem smarter than they are while spending as little as possible doing it. Every cycle saved on a background crowd member is a cycle available for something the player is actively watching. That trade-off, invisible and constant, is one of the genuinely underappreciated craft problems in game development.
So here's the question worth sitting with: next time you're in a crowded game area, are you actually looking at the crowd, or through it, toward your objective? The designers know the answer. The engine is counting on it.
The crowd doesn't need to be real. It needs to be real enough, for exactly as long as you're looking, and not a frame longer.