How the Elo Rating System Works

The Elo rating system predicts the outcome of a match from the gap between two players' ratings, then nudges each rating after the game. A larger rating means a higher expected win probability: a 100-point edge predicts about 64%, a 200-point edge about 76%. After each game your rating moves toward what the result says you deserved, scaled by a number called the K-factor.

The man behind the math

The system is named for Arpad Elo (1903 to 1992), a Hungarian-American physics professor at Marquette University and a strong chess master active in the U.S. Chess Federation from its founding in 1939. Before him, American chess used the Harkness system, devised around 1950 by USCF administrator Kenneth Harkness. It worked, but produced ratings many players found arbitrary. Elo, dissatisfied, built something with a genuine statistical basis. The USCF adopted his system at a 1960 meeting in St. Louis, and FIDE, the world chess body, followed in 1970. For roughly fifteen years afterward Elo did the calculations himself, which was feasible because fewer than 2,000 players were FIDE-rated at the time.

Elo was modest about precision. He compared rating a player to "the measurement of the position of a cork bobbing up and down on the surface of agitated water with a yardstick tied to a rope and swaying in the wind." That humility is worth remembering: a rating is an estimate, not a verdict.

Rating gaps as win probabilities

The heart of the system is the expected-score formula. For player A facing player B, with ratings RA and RB, the expected score (a number between 0 and 1, read as win probability when draws are ignored) is:

E_A = 1 / (1 + 10^((R_B - R_A) / 400))

This is a logistic, S-shaped curve. The number 400 is a scaling constant: every 400 points of advantage multiplies your expected-score ratio by ten. So a 400-point gap means roughly 10-to-1 odds, about 91% for the stronger player. The constant was chosen partly for historical continuity with the older system and partly because it makes base-10 logarithm arithmetic convenient. Plug in equal ratings and you get exactly 0.5, a coin flip, which is the sanity check the whole scale is built around.

The benchmarks people quote come straight out of that formula:

Rating gapExpected score (stronger player)Approx. odds
050%1 : 1
10064%~1.8 : 1
20076%~3 : 1
40091%10 : 1

Notice the curve flattens at the extremes. Even a 600-point favorite is not a sure thing, which is exactly why ratings keep being interesting rather than predetermined. You can experiment with these gaps and watch the points change on our Elo rating calculator.

Updating after a game: the K-factor dial

Prediction is only half the system. After a game, each rating updates by the same simple rule:

New = Old + K x (Actual - Expected)

Actual is your real result: 1 for a win, 0.5 for a draw, 0 for a loss. Expected is the probability the formula gave you. The difference is your surprise, and K is the dial that controls how strongly surprise moves your number.

Worked example: an 1800 plays a 1700. The formula gives the 1800 an expected score of 0.64 and the 1700 player 0.36. If the underdog wins with K = 32, they gain 32 x (1 - 0.36) ≈ 20 points and the favorite loses 32 x (0 - 0.64) ≈ 20 points. The match is zero-sum: every point one player gains, the other loses.

This is also why upsets pay big and expected wins pay little. Beat someone you were already supposed to beat and Actual minus Expected is small, so you gain almost nothing. Beat someone far above you and the gap is huge, so you leap. And because a draw scores 0.5, drawing a much stronger opponent still nets you points (your Actual beats your Expected), while drawing a much weaker one costs you. The math has no concept of "moral victory," only of beating expectations.

Different K-factors, different volatility

K is a design choice, and the major bodies disagree on it. FIDE uses K = 40 for newcomers (until they complete 30 rated games), 20 for established players under 2400, and 10 once a player reaches 2400, so elite ratings barely twitch. The older USCF system ran hotter: 32 below 2100, 24 from 2100 to 2400, 16 above (today's USCF formula also scales K with the number of games in an event). Chess.com sits in a similar 16 to 32 range. A high K means ratings chase your true strength quickly but bounce around; a low K means stability at the cost of slow response.

FIDE also applies a "400-point rule": when two players are more than 400 points apart, the gap is treated as exactly 400 for the calculation, so a single fluke loss to a much weaker player cannot crater a top rating. As of October 2025 this cap still applies to players rated below 2650, but for players rated 2650 and above the true difference is now used, a tweak aimed at curbing "rating farming" against far weaker opponents.

Why it spread far beyond chess

Elo originally assumed performance followed a normal (bell-curve) distribution. Later analysis suggested weaker players win slightly more often than that predicts, so the USCF and many sites switched to a logistic distribution; FIDE still uses Elo's original expectancy table. Either way, the core idea proved astonishingly portable, because all it needs is pairwise outcomes.

League of Legends launched with literal Elo, then in 2013 hid the number behind tiers and League Points while keeping a hidden Matchmaking Rating underneath that behaves the same way. Pure Elo struggles with team and multiplayer games, so successors emerged: Glicko adds a "rating deviation" that tracks how uncertain your rating is, Glicko-2 adds a volatility parameter for streaky players, and Microsoft's TrueSkill uses Bayesian inference to rate every player in a multiplayer lobby at once. Even AI got the treatment. The LMSYS Chatbot Arena ranked large language models with Elo from anonymous human preference votes, later moving to the related Bradley-Terry model for stability.

One crucial caveat ties it all together: Elo is relative, not absolute. A rating only means something inside its own pool, because the formula is calibrated against the opponents you actually face. A 1700 on Chess.com is not the same strength as a 1700 in FIDE, and USCF ratings tend to run 50 to 100 points above FIDE equivalents. The number measures where you stand among the people in your league, nothing more. If probability dialing intrigues you, the same expected-value logic powers the Kelly criterion for sizing bets.

Frequently Asked Questions

A 100-point rating advantage predicts about a 64% expected score for the stronger player. A 200-point gap raises that to roughly 76%, and a 400-point gap to about 91%, since every 400 points multiplies the odds ratio by ten.

The K-factor is the multiplier that controls how much a single result changes your rating. A high K (like 40) makes ratings move fast and stay volatile; a low K (like 10) keeps elite ratings stable. FIDE, the USCF, and Chess.com all use different K values.

Your rating change is K times (Actual minus Expected). Beating a much stronger opponent produces a large surprise, so you gain a lot. Beating someone you were already favored to beat produces almost no surprise, so you gain very little.

No. Elo is relative to the player pool it was calculated in, not an absolute skill score. Different systems calibrate against different opponents, so Chess.com, FIDE, and USCF ratings are not directly comparable; USCF ratings typically run 50 to 100 points above FIDE.