Character codes — ASCII and Unicode · 字符编码——ASCII 和 Unicode
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| Unicode/ˈjuːnɪkəʊd/ | 统一码 | tǒng yī mǎ |
| character set/ˈkærɪktə set/ | 字符集 | zì fú jí |
| ASCII/ˈæski/ | ASCII码 | ASCII mǎ |
| encoding/enˈkəʊdɪŋ/ | 编码 | biān mǎ |
| code point/kəʊd pɔɪnt/ | 码点 | mǎ diǎn |
| UTF-8/ˌjuː tiː ˈef eɪt/ | UTF-8编码 | UTF-8 biān mǎ |
The email that arrived as gibberish
- Through the 1990s, a message typed in Warsaw and read in Tokyo often arrived as a wall of nonsense. Both computers stored eight bits per character. They simply disagreed about what the top 128 patterns meant.
- Poland's code page put Polish letters there. Japan's put its own characters there. Nothing was corrupted in transit: the bytes arrived intact and were read against a different table.
- The fix was to stop having tables per country and give every character in every script one number, for ever. That is Unicode 统一码, and it is why the emoji you send arrives as the same picture.
- This lesson is how text becomes numbers: character sets 字符集, ASCII, Unicode, and the encodings that store them.
变成乱码的那封邮件
- 整个 1990 年代,在华沙敲下、在东京读到的消息常常变成一堵乱码。两台计算机都用八位存一个字符。它们只是对最高的 128 个模式的含义意见不合。
- 波兰的代码页在那里放波兰字母。日本的在那里放自己的字符。传输中什么都没损坏:字节完好地到达,却被对着另一张表来读。
- 解决办法是不再按国家各有一张表,而是给每种文字的每个字符一个永久的号码。这就是 Unicode(统一码),也是你发的表情符号能原样到达的原因。
- 这一课讲文字怎样变成数字:字符集(character set)、ASCII、Unicode,以及存储它们的编码。
A character set gives each character a number
- A computer stores no letters, only numbers. A character set is the set of characters a computer can represent, each with its own binary code.
- The number for one character is its code point 码点.
Ais 65,ais 97,0the digit is 48. - The code point is not the character's meaning; it is an agreed label. Text is readable only because both machines use the same character set.
The letter is the picture; the byte is the number
字符集给每个字符一个号码
- 计算机不存字母,只存数字。字符集是计算机能表示的字符的集合,每个字符有自己的二进制代码。
- 一个字符的号码是它的码点(code point)。
A是 65,a是 97,数字0是 48。 - 码点不是字符的含义;它是一个约定的标签。文字之所以可读,只因为两台机器用同一个字符集。

字母是那个图形;字节是那个数字
A character is stored as a number · 一个字符被存为一个数
Each character has a code number — 'A' is 65. Flip the bits to see that code in binary and hex, exactly how the computer holds it. · 每个字符都有一个代码数字——'A' 是 65。翻转这些位,看那个代码的二进制和十六进制,正是计算机保存它的方式。
A character set such as ASCII defines: · 一个像 ASCII 这样的字符集定义:
A character set maps each character to a number (its code point), which is what the computer actually stores. · 一个字符集把每个字符映射到一个数字(它的码点),那是计算机实际存储的东西。
ASCII
- ASCII ASCII码, the American Standard Code for Information Interchange, uses 7 bits, giving $2^7 = 128$ code points: the basic Latin letters, digits, punctuation, and 32 control codes such as carriage return.
- Extended ASCII uses 8 bits, giving 256 code points. The lower 128 are identical to ASCII; the upper 128 vary by region, which is exactly what broke that email.
- You are never asked to memorise a code, but you are asked for the counts: 7 bits, 128; 8 bits, 256.
ASCII
- ASCII(ASCII码,美国信息交换标准代码)用 7 位,给出 $2^7 = 128$ 个码点:基本拉丁字母、数字、标点,以及回车之类的 32 个控制码。
- 扩展 ASCII 用 8 位,给出 256 个码点。低 128 个与 ASCII 完全相同;高 128 个因地区而异,这正是弄坏那封邮件的东西。
- 从不要求你背下某个代码,但会要求你说出数量:7 位,128;8 位,256。
How many different code points does 7-bit ASCII have? · 7 位 ASCII 有多少个不同的码点?
7 bits give $2^7 = 128$ code points. · 7 位给出 $2^7 = 128$ 个码点。
How many different code points does extended ASCII have? · 扩展 ASCII 有多少个不同的码点?
Eight bits give 2^8 = 256. The lower 128 match plain 7-bit ASCII; the upper 128 vary by region. · 八位给出 2^8 = 256。低 128 个与普通 7 位 ASCII 相同;高 128 个因地区而异。
Unicode
- Unicode is a universal character set: one code point for almost every character in every script alive or dead, plus symbols and emoji, over 149,000 of them.
- The code point is separate from how it is stored. An encoding 编码 turns a code point into bytes.
- UTF-8 UTF-8编码 uses 1 to 4 bytes per character and is ASCII-compatible: the first 128 code points are one byte, identical to ASCII. UTF-16 uses 2 or 4 bytes; UTF-32 uses a fixed 4.
Unicode
- Unicode 是一个通用字符集:给几乎每种活着或死去的文字中的几乎每个字符一个码点,再加上符号和表情符号,超过 149,000 个。
- 码点与它怎样存储是分开的。编码(encoding)把码点变成字节。
- UTF-8(UTF-8编码)每个字符用 1 到 4 个字节,并且兼容 ASCII:前 128 个码点占一个字节,与 ASCII 完全相同。UTF-16 用 2 或 4 个字节;UTF-32 固定用 4 个。
Which is true of UTF-8? · 关于 UTF-8 哪个是真的?
UTF-8 is a variable-length Unicode encoding (1–4 bytes); its first 128 code points match ASCII, so plain ASCII text is valid UTF-8. · UTF-8 是一种可变长度的 Unicode 编码(1–4 字节);它的前 128 个码点与 ASCII 匹配,所以纯 ASCII 文本是有效的 UTF-8。
What is the relationship between Unicode and UTF-8? · Unicode 和 UTF-8 是什么关系?
The set assigns the numbers; the encoding decides how many bytes each number takes. UTF-16 and UTF-32 are other encodings of the same set. · 字符集分配号码;编码决定每个号码占多少字节。UTF-16 和 UTF-32 是同一集合的其他编码。
Worked example: compare ASCII and Unicode
- Give one advantage and one disadvantage of using Unicode instead of ASCII. [2]
- Advantage: Unicode represents far more characters, so text in any script, Chinese, Arabic, Greek, and emoji can be stored, and a single document can mix languages; files are portable because there is no per-country code page to disagree about.
- Disadvantage: for English-only text a Unicode file is usually larger, because a character may take more than one byte.
- Both halves are needed. "It has more characters" alone is one mark of two.
例题:比较 ASCII 和 Unicode
- 给出用 Unicode 代替 ASCII 的一个优点和一个缺点。[2]
- 优点:Unicode 能表示多得多的字符,所以任何文字——中文、阿拉伯文、希腊文——和表情符号都能存储,一份文档里可以混用多种语言;文件可移植,因为没有按国家各异、彼此不合的代码页。
- 缺点:对只有英文的文本,Unicode 文件通常更大,因为一个字符可能占不止一个字节。
- 两半都要有。单说"它字符更多"两分里只得一分。
A key advantage of Unicode over ASCII is that it: · Unicode 相对 ASCII 的一个关键优点是它:
Unicode covers nearly every writing system plus symbols and emoji — far beyond ASCII's basic English set. · Unicode 覆盖几乎每一种书写系统加上符号和表情符号——远超 ASCII 的基本英语集。
For English-only text, a Unicode file is usually larger than the same text in ASCII. · 对于只有英语的文本,一个 Unicode 文件通常比同样文本的 ASCII 更大。
Unicode encodings can use more bytes per character, so plain English text is usually larger than in 7-bit ASCII — the trade-off for universal coverage. · Unicode 编码每个字符可能用更多字节,所以纯英语文本通常比 7 位 ASCII 更大——通用覆盖的权衡。
Which are advantages of Unicode over ASCII? Select all · 所有 that apply. · Unicode 相对 ASCII 有哪些优点?选出所有适用的。
Size is the trade-off, not a benefit: for plain English a Unicode file is usually larger. · 大小是代价,不是好处:对纯英文,Unicode 文件通常更大。
Worked example: the size of a text file
- A message of 500 characters is stored in extended ASCII. How large is the file in bytes?
- Extended ASCII uses 8 bits, one byte, per character, so $500 \times 1 = 500$ bytes, or $500 \times 8 = 4000$ bits.
- The same message in UTF-32? Four bytes per character, so 2000 bytes.
- Show the bits-per-character, multiply, then convert once. Mixing bits and bytes halfway through is the usual lost mark.
例题:文本文件的大小
- 一条 500 个字符的消息用扩展 ASCII 存储。文件有多少字节?
- 扩展 ASCII 每个字符用 8 位即一个字节,所以 $500 \times 1 = 500$ 字节,或 $500 \times 8 = 4000$ 位。
- *同一条消息用 UTF-32 呢?*每个字符四个字节,所以 2000 字节。
- 写出每字符的位数,相乘,再一次性换算。中途混用位和字节是常见的失分点。
A 500-character message is stored in extended ASCII. How many bytes does it need (ignore any header)? · 一条 500 字符的消息用扩展 ASCII 存储。需要多少字节(忽略文件头)?
Eight bits, one byte, per character: 500 x 1 = 500 bytes, or 4000 bits. In UTF-32 the same text needs 2000 bytes. · 每个字符八位即一字节:500 x 1 = 500 字节,即 4000 位。同样的文本用 UTF-32 需要 2000 字节。
Reading a character in binary and hex
- Because a code point is just a number, it converts like any other.
Ais 65, which is0100 0001in binary and41in hexadecimal. - Two useful patterns: the digits
0to9run from 48, and lower case is 32 more than upper case, soa(97) isA(65) plus 32. That difference is a single bit. - A question that gives you
A= 65 and asks forEwants $65 + 4 = 69$, not a memorised table.
用二进制和十六进制读一个字符
- 因为码点只是一个数,它像任何数一样转换。
A是 65,二进制是0100 0001,十六进制是41。 - 两个有用的规律:数字
0到9从 48 开始;小写比大写大 32,所以a(97)是A(65)加 32。这个差正好是一位。 - 给出
A= 65 并问E的题,要的是 $65 + 4 = 69$,而不是背下来的表。
The ASCII code for A is 65. What is the ASCII code for E? · A 的 ASCII 码是 65。E 的 ASCII 码是多少?
The letters are consecutive, so E is four after A. Nothing has to be memorised beyond one anchor value. · 字母是连续的,所以 E 在 A 之后四位。除了一个锚点值,什么都不用背。
Marks that slip away
- ASCII is 7 bits and 128 code points; extended ASCII is 8 bits and 256. Do not write 8 bits for plain ASCII.
- Unicode is a character set; UTF-8 is an encoding of it. They are not two rival sets.
- UTF-8 is variable length, 1 to 4 bytes, not always one byte and not always four.
- The disadvantage of Unicode is file size for plain English, not "it is slower" or "it is harder to read".
容易丢掉的分
- ASCII 是 7 位、128 个码点;扩展 ASCII 是 8 位、256 个。不要给普通 ASCII 写 8 位。
- Unicode 是字符集;UTF-8 是它的一种编码。它们不是两个对立的集合。
- UTF-8 是变长的,1 到 4 个字节,不是永远一个,也不是永远四个。
- Unicode 的缺点是纯英文文本的文件大小,不是"它更慢"或"它更难读"。
You've got it
- a character set gives each character a code point; text is numbers plus an agreement about how to read them
- ASCII: 7 bits, 128 code points, English only · extended ASCII: 8 bits, 256, the top half varies by region
- Unicode: one code point for every script and emoji, stored by an encoding; UTF-8 is 1 to 4 bytes and ASCII-compatible
- Unicode gains characters, portability and mixed languages; it costs file size on plain English text
你掌握了
- 字符集给每个字符一个码点;文字就是数字加上关于怎样读它们的约定
- ASCII:7 位、128 个码点、只有英文 · 扩展 ASCII:8 位、256 个,高半部分因地区而异
- Unicode:每种文字和表情符号各有一个码点,由编码存储;UTF-8 是 1 到 4 字节且兼容 ASCII
- Unicode 换来字符、可移植性和多语言混排;代价是纯英文文本的文件更大