Games are among the most demanding distributed systems in software engineering. Sub-50ms latency requirements, millions of concurrent users, complex state synchronisation, real-time fraud detection on micro-transactions, and the need to handle viral growth spikes that can 10x traffic within hours of a major feature drop. This guide covers the backend architecture Zyllo Tech designs for gaming and entertainment platforms.
How much latency can a multiplayer game tolerate?
Latency in games is a user experience issue, not just a technical one. A 200ms round-trip in a battle royale game makes it unplayable. The networking architecture must be designed around latency targets from day 1:
- WebSocket for persistent connections — eliminates HTTP handshake overhead for game state updates.
- UDP via QUIC protocol for real-time game data where some packet loss is acceptable (position updates) — use TCP/WebSocket for reliable events (purchases, achievements).
- Edge deployment on Cloudflare Workers or AWS Global Accelerator to minimise RTT — a player in Mumbai connecting to a Mumbai edge node instead of Singapore saves 40–60ms.
- Dedicated game servers (AWS GameLift, Agones on Kubernetes) for physics-authoritative games vs client-authoritative for mobile games.
- Regional matchmaking: never route players across continents — the latency makes the game unplayable.
How does skill-based matchmaking work?
Matchmaking is a real-time optimisation problem: find the best group of players given multiple competing objectives simultaneously:
- Skill-based matchmaking (SBMM): Elo, TrueSkill, or Glicko-2 rating systems — TrueSkill2 works better for team games.
- Latency constraints: only match players who can achieve <80ms P95 RTT to the same game server.
- Wait time vs quality trade-off: progressively relax skill constraints as wait time increases to prevent indefinite queuing.
- Queue priority for returning players after disconnect to restore them to the same match.
- Anti-smurf detection: flag accounts with suspicious skill progression patterns for manual review.
How do you keep game state in sync across clients?
- Server-authoritative model: the server is the source of truth for all game state — clients send inputs, receive state updates.
- Entity Component System (ECS) on the server for efficient state management at scale.
- Delta compression: only send changed state, not full world state on every tick — reduces bandwidth by 60–80%.
- Client-side prediction: clients predict movement locally, reconcile with server state when it arrives — eliminates perceived latency for movement.
- Lag compensation: server rewinds game state to the point in time the client fired a shot to calculate hit detection fairly.
- The client applies the input locally the instant it happens and tags it with a monotonically increasing input sequence number. The player sees the movement at 0ms — nothing waits for a round trip.
- The same input goes to the server, which simulates it on its own copy of the world at a fixed tick rate. The server is the only source of truth; the client's copy is a prediction it may have to give back.
- For hit detection the server rewinds: it reconstructs where every entity was at the moment the client actually fired, not where they are now. Skip this and high-ping players have to lead their shots to compensate for their own latency.
- The server broadcasts authoritative state along with the last input sequence number it processed for that client.
- The client reconciles — snap to the server state, then re-apply every input sent after that acknowledged sequence number. Discarding those unacknowledged inputs instead of replaying them is exactly what rubber-banding looks like.
const TICK_HZ = 30; // ~33ms per tick
function tick(world, clients) {
world.step(1 / TICK_HZ); // simulate all inputs queued this tick
for (const c of clients) {
const visible = world.entitiesVisibleTo(c); // interest management
const delta = diff(c.lastAckedSnapshot, visible);
if (delta.isEmpty) continue;
send(c, {
tick: world.tick,
ackInput: c.lastProcessedInputSeq, // what the client reconciles against
changes: delta, // only what moved
});
}
}
// Broadcasting the full world to every client is what caps concurrency
// long before CPU does: cost grows with players x entities, so the
// server dies at exactly the moment the game gets popular.In-Game Economy & Virtual Wallet
- Double-entry ledger for all virtual currency transactions — same principles as financial accounting.
- Atomic transactions: item purchase must debit currency AND credit inventory in a single ACID transaction — never update one without the other.
- Transaction idempotency: mobile payment flows retry on timeout; idempotency keys prevent double-crediting.
- Limited-edition item scarcity: Redis atomic DECR for item stock management during limited drops.
- Platform IAP (in-app purchase) integration: Apple StoreKit 2 and Google Play Billing with server-side receipt validation.
How do you stop cheating in a multiplayer game?
- Server-side validation of all game actions — never trust the client for anything consequential.
- Statistical anomaly detection: players with statistically impossible accuracy or movement patterns flagged for review.
- Velocity limits on economy actions: cap item purchases per account per minute.
- Device fingerprinting and account clustering to detect multi-accounting and account farms.
- Payment fraud: 3DS2 for high-value transactions, velocity checks, and ML fraud scoring on purchase patterns.
- Game Server P99 Latency: < 50ms
- Match Quality Score: 94/100
- Transaction Success Rate: 99.97%
- Cheat Detection Precision: 98.2%
