Tagged articles

code points

3 articles · Page 1 of 1
IT Learning Made Simple
IT Learning Made Simple
Sep 20, 2026 · Fundamentals

How Chinese Characters Are Stored in Computers: Unicode & UTF-8 Deep Dive

This article explains how computers store Chinese characters using Unicode and UTF-8, covering the history of encoding chaos, Unicode code points, UTF-8 variable-length encoding rules, a step-by-step example of encoding '中', common mojibake causes, BOM, and programming best practices like using utf8mb4 in MySQL.

BOMCharacter EncodingMySQL utf8mb4
0 likes · 10 min read
How Chinese Characters Are Stored in Computers: Unicode & UTF-8 Deep Dive
AI Engineer Programming
AI Engineer Programming
Jun 25, 2026 · Fundamentals

A Programmer’s Intro to Unicode

This guide walks programmers through Unicode’s massive code space, its diverse scripts, encoding schemes like UTF‑8 and UTF‑16, combining marks, canonical equivalence, normalization forms, and grapheme clusters, explaining why the system is complex yet essential for global text handling.

Character EncodingUTF-16UTF-8
0 likes · 21 min read
A Programmer’s Intro to Unicode
Programmer DD
Programmer DD
Jul 22, 2020 · Fundamentals

Why Java’s char Can’t Represent All Unicode Characters – Understanding UTF‑16 and Code Points

This article explains how Java’s char type stores Unicode code units in UTF‑16, why its range of \u0000 to \uffff limits direct representation of newer Unicode characters, and how methods like String.length, getBytes, and code‑point APIs help handle multi‑byte characters such as emojis and rare Chinese glyphs.

StringUTF-16Unicode
0 likes · 10 min read
Why Java’s char Can’t Represent All Unicode Characters – Understanding UTF‑16 and Code Points