Qvantor logo QvantorInfrastructure, explained plainly
Games

How Online Game Servers Keep Millions of Players Connected

On a Friday night when a big patch drops, I can usually tell how well a studio planned its servers within about ten minutes.

Grid of server racks representing online game servers

On a Friday night when a big patch drops, I can usually tell how well a studio planned its servers within about ten minutes. Either the login screen spins and then lets everyone in, or the official status account starts posting apologies about queue times. The difference between those two nights is not luck. It is a stack of decisions made months earlier about how players get split up, where the machines live, and what happens when a few hundred thousand people all press Play at once.

I run small servers for my friends, so I live at the opposite end of the scale from a studio like Riot or Blizzard. But the problems rhyme. A Valheim server with six people and a battle royale with ten million daily players both have to answer the same three questions: who talks to which machine, how often does the game state update, and what breaks first when demand spikes.

Nobody plays on one giant server

The first thing worth understanding is that "the game server" is almost never one thing. When you queue for a match in Fortnite or Apex Legends, a matchmaking service, which is just a separate program, looks at your region, your skill rating and how long you have been waiting. It then picks a free match server process somewhere in a data center near you and hands your client its address. That match server holds maybe 60 to 100 players for twenty minutes, then shuts down and frees its slot for the next group.

So a game with millions of concurrent players is really running tens of thousands of small, short lived servers, plus a handful of long lived services behind them:

  • Login and account services that check who you are and what you own.
  • Matchmaking, which groups players and assigns match servers.
  • Inventory and progression databases that remember your skins, levels and currency between matches.
  • Chat, friends and party services, which keep running while you sit in the lobby.
  • The match or world servers themselves, where the actual shooting, building or questing happens.

Persistent online worlds split things differently. World of Warcraft has long divided players into realms, each a copy of the world with its own population, and later added cross realm zones so quiet realms do not feel empty. Final Fantasy XIV groups worlds into data centers, and when Endwalker launched in December 2021 the demand outran capacity badly enough that Square Enix temporarily paused sales of the game. That is what happens when the login side, not the combat side, becomes the bottleneck.

EVE Online is the famous exception. It runs its main universe as a single shard, so every player shares one galaxy. To survive huge fleet battles, CCP built a system called time dilation that slows the simulation in a busy star system rather than letting it collapse. Players hate the slow motion, but they hate a crashed node far more.

Why splitting players works better than buying bigger machines

A common assumption, and one I held when I started hosting, is that lag and crashes come from weak hardware, so the fix is a bigger box. For a single small server that is sometimes true. At scale it mostly is not.

The trouble is that game simulation does not grow in a straight line. If 10 players are in the same area, each one's actions may need to be sent to the other 9. At 100 players in the same area, each action may go to 99 others. The amount of state the server has to compare and send grows much faster than the player count. Throwing a faster CPU at that only moves the wall a little further away.

So big studios cap how many players share one simulation and spread the load sideways. Battle royale games cap the lobby. MMOs use instancing, where a dungeon gets its own private copy, and they use interest management, which means the server only tells you about players and objects close enough to matter. If someone is three kilometers away on the map, your client simply does not hear about them.

The trick to holding millions of players is making sure almost none of them ever need to know about each other.

Tick rate, regions and the speed of light

Every match server runs a loop, and each pass through that loop is called a tick. On each tick it reads player inputs, updates positions and physics, and sends the results back out. Valorant's servers run at 128 ticks per second, which Riot has written about publicly, while many other shooters run lower, often 20 to 64. Higher tick rates feel more precise, but each extra tick costs CPU time across every match on every machine, so studios choose carefully.

Region placement matters even more than tick rate. Light in fiber covers roughly 200 kilometers per millisecond, and real routes zigzag between cities, so a player in Lisbon connecting to a server in Frankfurt will rarely see less than 30 to 40 milliseconds round trip no matter how good the code is. That is why studios run servers in dozens of regions and why matchmaking often prefers a slightly worse skill match over a much worse connection. I go into what those milliseconds actually feel like in why low latency matters so much to players.

What happens on launch night

Launch spikes are the real test. Most of the year, player counts follow a gentle daily wave that peaks in the evening for each region. A new season or expansion turns that wave into a cliff.

Studios prepare in a few predictable ways:

  1. Load testing with bots that log in and play, sometimes millions of fake sessions, to find which service falls over first.
  2. Elastic match capacity, renting extra cloud machines for the first week and handing them back later.
  3. Login queues on purpose. A queue that says "you are number 4,812" is a pressure valve that protects the databases behind it.
  4. Staggered unlocks, releasing a patch region by region so not everyone hits the servers in the same hour.
  5. Kill switches that let operators turn off a broken feature, like a store page or a new mode, without taking the whole game down.

The queue is the part players complain about most, and it is the part I defend most. A login queue is the server saying no politely. The alternative is the server saying yes to everyone and then losing data for all of them.

The quiet services nobody thanks

Match servers get the attention, but the long lived services are where outages hurt. If the inventory database stalls, people finish a match and lose their rewards. If the friends service goes down, a squad cannot form even though every match server is healthy. Plenty of the worst launch nights I have watched were caused by an account or entitlement service, not by the game simulation.

These services also talk to the outside world. Account systems send verification and password reset emails, and when those messages land in spam during a launch, support tickets explode. I have seen friends locked out of a new game for an hour purely because a reset email never arrived. That side of the stack overlaps a lot with ordinary email infrastructure, and the email infrastructure section of this site covers how those messages get delivered.

What a small host can borrow from this

You do not need a data center to use the same ideas. When I moved our group's Minecraft server from one overloaded instance to two, a survival world and a separate creative world, the lag spikes nearly vanished. Same total hardware, but each simulation had less to track. I also set a soft cap of 12 players on survival, which felt strict until everyone noticed how much smoother it ran.

If you want to try hosting something yourself, start small and watch what actually breaks. Usually it is memory or disk, not raw CPU. My notes on what self hosting a game server teaches you walk through the first month of mistakes I made.

This week, pick a game you play online and look up where its servers are and what tick rate it uses. Then compare that to your ping in the scoreboard. Once you see the numbers side by side, the queues, the region locks and the occasional rubber banding all start to make a lot more sense.

MQ
Marisa Quan

Marisa hosts game servers for her friends and writes about latency, uptime and the people who keep servers alive.

More posts by Marisa

More in Games