Data compression · 데이터 압축
Why compress data
- Data takes up space to store and time to send.
- Compression makes a file smaller so it is cheaper to save and faster to share.
- There are two kinds: lossless and lossy.
왜 데이터를 압축해야 하나요
- 데이터는 저장에 공간을 차지하고 전송에 시간이 걸립니다.
- 압축은 파일을 작게 만들어 저장 비용을 절감하고 공유 속도를 높입니다.
- 두 가지 유형이 있습니다: 무손실과 유손실입니다.
Lossless compression
- Lossless makes a file smaller but keeps every bit of the data.
- When you open the file, you get back the exact original.
- It works by finding patterns and writing them in a shorter way.
무손실 압축
- 무손실은 파일을 작게 만들지만 데이터의 모든 비트를 보존합니다.
- 파일을 열면 정확히 원본을 복원할 수 있습니다.
- 패턴을 찾아 더 짧은 방식으로 기록함으로써 작동합니다.
Original: AAAAAAAA-BBB
Shorter: 8A-3B (8 A's, then 3 B's)
Open it: AAAAAAAA-BBB (exactly the same again)
Lossy compression
- Lossy makes a file much smaller by throwing away some data.
- It drops details that people can barely see or hear.
- You cannot get the exact original back — but it is "close enough".
유손실 압축
- 유손실은 일부 데이터를 폐기하여 파일을 훨씬 작게 만듭니다.
- 사람이 거의 보거나 들지 못하는 디테일을 제거합니다.
- 정확한 원본을 복원할 수 없으나, "충분히 유사"한 수준입니다.
Photo (large) --lossy--> Photo (small)
A few colors and fine details are gone,
but your eye hardly notices.
The trade-off: size vs quality
- Lossless keeps full quality, but the file stays larger.
- Lossy gives a much smaller file, but quality goes down a little.
- You choose based on what matters more: perfect data or small size.
트레이드오프: 크기 대 품질
- 무손실은 완전한 품질을 유지하지만 파일 크기는 큽니다.
- 유손실은 파일을 훨씬 작게 하지만 품질이 다소 떨어집니다.
- 완전한 데이터인가 작은 크기인가, 무엇이 중요한지에 따라 선택합니다.
When to use each
- Use lossless when every detail must be exact.
- Use lossy when a small drop in quality is fine and small size matters.
사용할 때의 기준
- 모든 디테일이 정확해야 할 때는 무손실을 사용하십시오.
- 품질의 작은 하락은 허용되며 작은 크기가 중요할 때는 유손실을 사용하십시오.
Lossless: text, code, a .zip file, a spreadsheet
Lossy: photos (JPEG), music (MP3), video
Key idea
- Compression trades size against quality (or against work to undo it).
- Lossless = smaller and perfect; lossy = much smaller but not exact.
- Good engineers pick the right kind for the job.
핵심 개념
- 압축은 크기와 품질( 또는 복원 작업) 사이에서 타협을 찾습니다.
- 무손실 = 작고 완벽함; 유손실 = 훨씬 작지만 정확하지 않음.
- 좋은 엔지니어는 작업에 맞는 올바른 방식을 선택합니다.
Run-length encoding
- Run-length encoding (RLE) is a simple lossless method.
- A run is a stretch of the same character repeated. RLE stores a count instead of the repeats.
- We will store each run as a pair
[character, count]inside a list.
루스 인코딩(RLE)
- 루스 인코딩(RLE)은 간단한 무손실 방식입니다.
- 루스는 동일한 문자가 연속해서 나타나는 구간입니다. RLE는 반복 대신 개수를 저장합니다.
- 각 루스를 목록 내부의 쌍
[character, count]로 저장하겠습니다.
"AAAB" -> [["A", 3], ["B", 1]] (3 A's, then 1 B)
[["A", 3], ["B", 1]] -> "AAAB" (decode it back — exact again)
Common mistakes
- Lossless compression can be reversed exactly; lossy throws away detail.
- More compression can mean lower quality.
흔한 실수
- 무손실 압축은 정확하게 역변환 가능하지만, 유손실은 디테일을 잃습니다.
- 압축률이 높아질수록 품질이 낮아질 수 있습니다.
Now you try
- Build RLE yourself: an encoder, a decoder, and a length helper.
- Each task checks your function on several inputs. Press Check answer.
이제 직접 해보기
- RLE를 직접 구현해 봅시다: 엔코더, 디코더, 그리고 길이 도우미 함수.
- 각 과제는 여러 입력에 대해 함수를 검증합니다. 답안 확인 버튼을 누르십시오.
Lossless compression · 무손실 압축
Run-length encoding replaces a run of repeats with count + symbol. · Run-length encoding은 반복되는 연속을 개수 + 기호로 교체합니다.
Write encode(text) for run-length encoding. Return a list of [character, count] pairs, one per run of repeats. Example: encode("AAAB") → [['A', 3], ['B', 1]]. For the empty string return []. · run-length encoding을 위한 encode(text)을 작성하십시오. 반복되는 연속마다 하나씩 [character, count] 쌍의 list를 반환하십시오. 예시: encode("AAAB") → [['A', 3], ['B', 1]]. 빈 문자열의 경우 []을 반환하십시오.
Click Run to see the output here. · 출력을 보려면 '실행'을 클릭하세요.
Write decode(pairs) that reverses the encoder: given a list of [character, count] pairs, rebuild the original string. Example: decode([['A', 3], ['B', 1]]) → 'AAAB'. For [] return ''. · 엔코더를 반대로 하는 decode(pairs)을 작성하십시오: [character, count] 쌍의 list가 주어지면 원본 문자열을 복원하십시오. 예시: decode([['A', 3], ['B', 1]]) → 'AAAB'. []의 경우 ''을 반환하십시오.
Click Run to see the output here. · 출력을 보려면 '실행'을 클릭하세요.
Without decoding, write original_length(pairs) that returns how many characters the original text had — just add up the counts. Example: original_length([['A', 3], ['B', 1]]) → 4. · 복호화 없이, 원본 텍스트가 몇 개의 문자였는지 반환하는 original_length(pairs)을 작성하세요 — 단순히 카운트만 더하면 됩니다. 예: original_length([['A', 3], ['B', 1]]) → 4.
Click Run to see the output here. · 출력을 보려면 '실행'을 클릭하세요.