๐Ÿ”ข Hamming Distance Calculator

3Hamming distance
karolin

Compares two equal-length strings (or bit strings) character by character and counts the positions where they differ โ€” the Hamming distance. It's a fundamental distance metric used in error-detecting code theory, comparing DNA sequences, and similarity scoring in machine learning.

How to use

  1. Enter the two strings you want to compare (they must be the same length).
  2. The Hamming distance is calculated automatically, and differing characters are highlighted in red.
  3. If the strings have different lengths, an error message is shown instead.

How the calculation works

The Hamming distance between two strings (or bit strings) of equal length is the number of positions where they differ. Richard Hamming introduced it in 1950 in his work on error detection and correction. This tool compares the two strings character by character, counts the positions that differ and highlights them. Characters are compared as code points, so an emoji counts as one. If the lengths differ, the Hamming distance is undefined and the tool says so. With bit strings of 0s and 1s, it tells you how many bits are flipped.

Worked example

"karolin" and "kathrin" k a r o l i n k a t h r i n Differences at positions 3 (r/t), 4 (o/h) and 5 (l/r) Hamming distance: 3 Bit strings 1011101 and 1001001 Hamming distance: 2 (positions 3 and 5 differ)

Things to be aware of

  • Error-correcting codes such as Hamming codes keep valid code words at least 3 apart, so a single flipped bit can be detected and corrected.
  • To compare strings of different lengths, use the Levenshtein distance.
  • Upper and lower case are treated as different characters.

FAQ

What is Hamming distance?

When comparing two equal-length sequences (strings or bit strings), it's the total number of positions where the corresponding values differ.

What happens if the strings are different lengths?

Hamming distance is only defined between sequences of equal length, so mismatched lengths can't be calculated and show an error instead.

Where is this used?

It's widely used in error-detecting/correcting code theory, comparing mutation counts between DNA sequences, and similarity scoring for feature vectors in machine learning.