Data compression · ضغط البيانات
Why compress data
- Data takes up space to store and time to send.
- Compression makes a file smaller so it is cheaper to save and faster to share.
- There are two kinds: lossless and lossy.
لماذا ضغط البيانات
- البيانات تأخذ مساحة للتخزين ووقتاً للإرسال.
- الضغط يجعل الملف أصغر ليصبح أرخص للحفظ وأسرع للمشاركة.
- هناك نوعان: بدون فقدان ومع فقدان.
Lossless compression
- Lossless makes a file smaller but keeps every bit of the data.
- When you open the file, you get back the exact original.
- It works by finding patterns and writing them in a shorter way.
الضغط بدون فقدان
- الضغط بدون فقدان يصغر الملف لكنه يحافظ على كل بت من البيانات.
- عند فتح الملف، تحصل على نسخة مطابقة تماماً للأصل.
- يعمل عن طريق العثور على أنماط وكتابتها بطريقة أقصر.
Original: AAAAAAAA-BBB
Shorter: 8A-3B (8 A's, then 3 B's)
Open it: AAAAAAAA-BBB (exactly the same again)
Lossy compression
- Lossy makes a file much smaller by throwing away some data.
- It drops details that people can barely see or hear.
- You cannot get the exact original back — but it is "close enough".
الضغط مع فقدان
- الضغط مع فقدان يصغر الملف كثيراً عن طريق التخلي عن بعض البيانات.
- يتخلص من التفاصيل التي بالكاد يستطيع البشر رؤيتها أو سماعها.
- لا يمكنك استعادة النسخة الأصلية بدقة — لكنها "قريبة بما يكفي".
Photo (large) --lossy--> Photo (small)
A few colors and fine details are gone,
but your eye hardly notices.
The trade-off: size vs quality
- Lossless keeps full quality, but the file stays larger.
- Lossy gives a much smaller file, but quality goes down a little.
- You choose based on what matters more: perfect data or small size.
المقايضة: الحجم مقابل الجودة
- الضغط بدون فقدان يحافظ على جودة كاملة، لكن الملف يبقى أكبر.
- الضغط مع فقدان يعطي ملفاً أصغر بكثير، لكن الجودة تنخفض قليلاً.
- تختار بناءً على ما يهم أكثر: بيانات مثالية أم حجم صغير.
When to use each
- Use lossless when every detail must be exact.
- Use lossy when a small drop in quality is fine and small size matters.
متى تستخدم كل نوع
- استخدم الضغط بدون فقدان عندما يجب أن تكون كل التفاصيل دقيقة.
- استخدم الضغط مع فقدان عندما يكون الانخفاض الطفيف في الجودة مقبولاً والحجم الصغير مهمًا.
Lossless: text, code, a .zip file, a spreadsheet
Lossy: photos (JPEG), music (MP3), video
Key idea
- Compression trades size against quality (or against work to undo it).
- Lossless = smaller and perfect; lossy = much smaller but not exact.
- Good engineers pick the right kind for the job.
فكرة رئيسية
- الضغط يوازن بين الحجم والجودة (أو الجهد المبذول لعكسه).
- بدون فقدان = أصغر ومثالي؛ مع فقدان = أصغر بكثير لكن ليس دقيقاً.
- المهندسون الجيدون يختارون النوع المناسب لكل مهمة.
Run-length encoding
- Run-length encoding (RLE) is a simple lossless method.
- A run is a stretch of the same character repeated. RLE stores a count instead of the repeats.
- We will store each run as a pair
[character, count]inside a list.
ترميز طول التسلسل
- ترميز طول التسلسل (RLE) هو طريقة بسيطة بدون فقدان.
- التسلسل هو سلسلة من نفس الحرف المكرر. يخزن RLE عدداً بدلاً من المكررات.
- سنخزن كل تسلسل كزوج
[character, count]داخل قائمة.
"AAAB" -> [["A", 3], ["B", 1]] (3 A's, then 1 B)
[["A", 3], ["B", 1]] -> "AAAB" (decode it back — exact again)
Common mistakes
- Lossless compression can be reversed exactly; lossy throws away detail.
- More compression can mean lower quality.
أخطاء شائعة
- يمكن عكس الضغط بدون فقدان بدقة؛ أما الضغط مع فقدان فيتخلى عن التفاصيل.
- زيادة الضغط قد تعني انخفاض الجودة.
Now you try
- Build RLE yourself: an encoder, a decoder, and a length helper.
- Each task checks your function on several inputs. Press Check answer.
الآن جرب بنفسك
- قم ببناء RLE بنفسك: مشفر، مفكك، وأداة مساعدة للطول.
- تتحقق كل مهمة من دالتك على عدة مدخلات. اضغط تحقق من الإجابة.
Lossless compression · الضغط غير الخاسر
Run-length encoding replaces a run of repeats with count + symbol. · ترميز المسارات المتكررة (Run-length encoding) يستبدل سلسلة من التكرارات بـ العد + الرمز.
Write encode(text) for run-length encoding. Return a list of [character, count] pairs, one per run of repeats. Example: encode("AAAB") → [['A', 3], ['B', 1]]. For the empty string return []. · اكتب encode(text) لترميز المسارات المتكررة. أرجع قائمة من أزواج [character, count]، واحد لكل سلسلة تكرار. مثال: encode("AAAB") → [['A', 3], ['B', 1]]. للسلسلة الفارغة أرجع [].
Click Run to see the output here. · اضغط تشغيل لرؤية المخرجات هنا.
Write decode(pairs) that reverses the encoder: given a list of [character, count] pairs, rebuild the original string. Example: decode([['A', 3], ['B', 1]]) → 'AAAB'. For [] return ''. · اكتب decode(pairs) تعكس المشفر: معطى قائمة من أزواج [character, count]، أعد بناء السلسلة الأصلية. مثال: decode([['A', 3], ['B', 1]]) → 'AAAB'. بالنسبة لـ [] أرجع ''.
Click Run to see the output here. · اضغط تشغيل لرؤية المخرجات هنا.
Without decoding, write original_length(pairs) that returns how many characters the original text had — just add up the counts. Example: original_length([['A', 3], ['B', 1]]) → 4. · بدون فك التشفير، اكتب original_length(pairs) تُرجع كم عدد الأحرف كان فيها النص الأصلي — فقط اجمع العدادات. مثال: original_length([['A', 3], ['B', 1]]) → 4.
Click Run to see the output here. · اضغط تشغيل لرؤية المخرجات هنا.