Data compression · データ圧縮
Why compress data
- Data takes up space to store and time to send.
- Compression makes a file smaller so it is cheaper to save and faster to share.
- There are two kinds: lossless and lossy.
データ圧縮の理由
- データを格納するにはスペースを、送信するには時間が必要です。
- 圧縮によりファイルが小さくなり、保存コストが削減され、共有速度が向上します。
- 2種類の圧縮方式があります:無損失圧縮と有損圧縮です。
Lossless compression
- Lossless makes a file smaller but keeps every bit of the data.
- When you open the file, you get back the exact original.
- It works by finding patterns and writing them in a shorter way.
無損失圧縮
- 無損失圧縮はファイルを小さくしますが、データのすべてのビットを保持します。
- ファイルを開くと、元のデータを正確に再構成できます。
- パターンを検出し、より短い表現で書き出すことで機能します。
Original: AAAAAAAA-BBB
Shorter: 8A-3B (8 A's, then 3 B's)
Open it: AAAAAAAA-BBB (exactly the same again)
Lossy compression
- Lossy makes a file much smaller by throwing away some data.
- It drops details that people can barely see or hear.
- You cannot get the exact original back — but it is "close enough".
有損圧縮
- 有損圧縮は一部のデータを削除することで、ファイルを大幅に小さくします。
- 人間がほとんど見えないまたは聞こえない詳細を省略します。
- 元のデータを完全には取り戻せませんが、「十分に近い」状態になります。
Photo (large) --lossy--> Photo (small)
A few colors and fine details are gone,
but your eye hardly notices.
The trade-off: size vs quality
- Lossless keeps full quality, but the file stays larger.
- Lossy gives a much smaller file, but quality goes down a little.
- You choose based on what matters more: perfect data or small size.
トレードオフ:サイズと品質
- 無損失圧縮は完全な品質を維持しますが、ファイルサイズは大きままです。
- 有損圧縮はファイルサイズを大幅に小さくしますが、品質が若干低下します。
- 完全なデータか小さなサイズか、どちらが重要かを基準に選択します。
When to use each
- Use lossless when every detail must be exact.
- Use lossy when a small drop in quality is fine and small size matters.
使用時機
- 全ての詳細が正確である必要がある場合は無損失圧縮を使用します。
- 品質のわずかな低下が許容され、サイズ的小型化が重要であれば有損圧縮を使用します。
Lossless: text, code, a .zip file, a spreadsheet
Lossy: photos (JPEG), music (MP3), video
Key idea
- Compression trades size against quality (or against work to undo it).
- Lossless = smaller and perfect; lossy = much smaller but not exact.
- Good engineers pick the right kind for the job.
重要な概念
- 圧縮はサイズと品質(または復元作業)とのトレードオフです。
- 無損失圧縮=小さく完全;有損圧縮=大幅に小さく不完全です。
- 優れたエンジニアはタスクに応じて適切な方式を選択します。
Run-length encoding
- Run-length encoding (RLE) is a simple lossless method.
- A run is a stretch of the same character repeated. RLE stores a count instead of the repeats.
- We will store each run as a pair
[character, count]inside a list.
ラン長符号化
- ラン長符号化(RLE)は単純な無損失圧縮手法です。
- 連続して同じ文字が並ぶ部分をランと呼びます。RLEは繰り返しを格納する代わりにカウントを記録します。
- 各ランをリスト内のペア
[character, count]として格納します。
"AAAB" -> [["A", 3], ["B", 1]] (3 A's, then 1 B)
[["A", 3], ["B", 1]] -> "AAAB" (decode it back — exact again)
Common mistakes
- Lossless compression can be reversed exactly; lossy throws away detail.
- More compression can mean lower quality.
よくあるミス
- 無損失圧縮は完全に逆転可能ですが、有損圧縮は詳細を失います。
- 圧縮率が高いほど、品質が低下する可能性があります。
Now you try
- Build RLE yourself: an encoder, a decoder, and a length helper.
- Each task checks your function on several inputs. Press Check answer.
あなたも試してみよう
- RLEを実装しましょう:エンコーダー、デコーダー、および長さ計算用の補助関数です。
- 各タスクでは複数の入力に対して関数を確認します。回答確認ボタンを押してください。
Lossless compression · 無損圧縮
Run-length encoding replaces a run of repeats with count + symbol. · ラン長エンコーディングは、連続する繰り返しの列をカウント+記号に置き換えます。
Write encode(text) for run-length encoding. Return a list of [character, count] pairs, one per run of repeats. Example: encode("AAAB") → [['A', 3], ['B', 1]]. For the empty string return []. · ラン長エンコーディング用encode(text)を書いてください。連続する繰り返しの列ごとに1つずつ、[character, count]のペアのリストを返します。例:encode("AAAB") → [['A', 3], ['B', 1]]。空の文字列に対しては[]を返します。
Click Run to see the output here. · 実行ボタンをクリックして出力を確認してください。
Write decode(pairs) that reverses the encoder: given a list of [character, count] pairs, rebuild the original string. Example: decode([['A', 3], ['B', 1]]) → 'AAAB'. For [] return ''. · エンコーダーの逆変換を行うdecode(pairs)を書いてください:[character, count]のペアのリストが与えられたら、元の文字列を再構築します。例:decode([['A', 3], ['B', 1]]) → 'AAAB'。[]に対しては''を返します。
Click Run to see the output here. · 実行ボタンをクリックして出力を確認してください。
Without decoding, write original_length(pairs) that returns how many characters the original text had — just add up the counts. Example: original_length([['A', 3], ['B', 1]]) → 4. · 復号化せずに、元のテキストが何文字あったかを示す original_length(pairs) を記述してください — カウントを足し合わせるだけでOKです。例:original_length([['A', 3], ['B', 1]]) → 4。
Click Run to see the output here. · 実行ボタンをクリックして出力を確認してください。