๐ Run-Length Encoding (RLE)
a3b3c4d1
Encodes and decodes run-length encoding (RLE), a simple compression scheme that replaces runs of identical characters with a "character + repeat count" format (e.g. aaa โ a3). Useful for learning how basic compression algorithms work, or understanding the principle behind fax and bitmap image compression.
How to use
- Enter text in the "Encode" tab to see the compressed result (e.g. aaa โ a3).
- Enter an RLE-format string in the "Decode" tab to reconstruct the original text.
- Click "Copy" to copy the result to your clipboard.
How the calculation works
Run-length encoding (RLE) is one of the simplest compression methods: each run of the same character is replaced with the character and how many times it repeats. Encoding reads the text from the start, counts how many times each character repeats and writes the character followed by the count. Decoding does the reverse, repeating each character by the number after it. Characters are handled as code points, so emoji encode and decode correctly. Because counts are written as digits, digits in the original text make it impossible to tell characters from counts. The tool warns you when that happens.
Worked example
aaabbbccccd โ a3b3c4d1 WWWWWWWWWWWWBWWWWWWWWWWWWBBB (28 characters) โ W12B1W12B3 (10 characters) Text without long runs gets longer: abc โ a1b1c1 (twice the length)
Things to be aware of
- RLE works best on data with long runs of the same value, such as black-and-white images, icons and fax data. Some old image formats (parts of BMP and PCX) used it.
- Ordinary text rarely repeats the same character, so RLE barely shrinks it.
- Practical compression such as ZIP's Deflate combines dictionary-based replacement of repeated patterns with variable-length codes based on frequency.
FAQ
Does this always shrink the input?
No. Strings with few repeated characters can actually get longer after encoding. RLE works especially well on data with long runs of identical values, like images with large solid-color areas.
Why does decoding fail with an error?
Input that isn't in the "character + digits" format (e.g. a character not followed by digits) can't be decoded.
Does this handle multi-byte characters like Japanese?
Yes, it processes text one character at a time, so Japanese characters are handled correctly.