Skip to content
JackSparrow414
Go back

Java Character Encoding and Decoding (Part 1)

Table of contents

Open Table of contents

Article body

The character values a computer can represent range only from 0 to 255. We use far more characters across different languages, so we need to convert them into something the computer understands: bits. With that introduction out of the way, let’s get to the point.

Let’s use an example to explore encoding and decoding in Java.

 For example, Java’s String.getBytes method converts a string to a byte array by calling StringCoding.encode, as shown below.

Java source of String.getBytes calling StringCoding.encode

Next, StringCoding calls into java.nio.charset, where lookupCharset() looks up the corresponding character encoding, such as UTF-8 or GBK.

StringCoding.encode looking up a charset and encoding through StringEncoder

After resolving the encoding name, it creates a StringEncoder and calls its encode method. At the lowest level, it calls Arrays.copyof(), which creates and returns a byte[].

That is the process of converting characters to bytes. Converting bytes back into characters follows a similar idea, so I won’t go through it again here.

Note: Java’s character-to-byte and byte-to-character conversions all follow this underlying process.


Share this post:

Continue this series

Encoding and Decoding in Java

  1. Java Character Encoding and Decoding (Part 1)You are here
  2. Java Character Encoding and Decoding (Part 2): URL Encoding

Comments

Questions, corrections, and experiences are welcome. Sign in with GitHub to comment; both language versions share this discussion.

Comments are available on the live site only.