Hash tables · Хэш-таблицы
Hash tables: fast lookup
- A hash table stores items so you can find them very fast — usually in one step.
- It is a list of slots. A hash function turns a key into a slot number.
- Instead of searching every item, you jump straight to the slot the key belongs in.
Хэш-таблицы: быстрый поиск
- Хэш-таблица хранит элементы так, чтобы находить их очень быстро — обычно за один шаг.
- Это список ячеек. Хэш-функция преобразует ключ в номер ячейки.
- Вместо поиска каждого элемента вы сразу переходите к ячейке, где должен находиться ключ.
A hash function
- A hash function takes a key and returns a slot index from
0tosize - 1. - The same key always gives the same slot, so you can find it again later.
- A simple one: add the character codes, then take
% sizeto stay in range.
Хэш-функция
- Хэш-функция принимает ключ и возвращает индекс ячейки от
0доsize - 1. - Один и тот же ключ всегда дает ту же ячейку, поэтому вы можете найти его снова позже.
- Простая функция: сложите коды символов, затем примените
% size, чтобы оставаться в диапазоне.
def hash_key(key, size):
total = 0
for ch in key:
total += ord(ch) # ord("A") is 65, ord("B") is 66, ...
return total % size
print(hash_key("cat", 10)) # a slot from 0 to 9
print(hash_key("cat", 10)) # same key -> same slot
print(hash_key("dog", 10))
Collisions
- Two different keys can hash to the same slot. That is a collision.
- A table has limited slots, so collisions are unavoidable as it fills up.
- We need a rule for what to do when the slot we want is already taken.
Коллизии
- Два разных ключа могут хэшироваться в одну и ту же ячейку. Это коллизия.
- В таблице ограниченное количество ячеек, поэтому коллизии неизбежны при заполнении.
- Нам нужно правило того, что делать, если нужная ячейка уже занята.
Linear probing
- Linear probing: if a slot is full, try the next slot, then the next, wrapping around.
- Keep stepping
(slot + 1) % sizeuntil you find a free slot. - Below,
AandFboth want slot 0, soFis pushed to slot 1.
Линейный пробинг
- Линейный пробинг: если ячейка занята, попробуйте следующую ячейку, затем следующую, с возвратом в начало.
- Продолжайте двигаться
(slot + 1) % size, пока не найдете свободную ячейку. - Ниже,
AиFоба хотят ячейку 0, поэтомуFпереносится в ячейку 1.
def hash_key(key, size):
total = 0
for ch in key:
total += ord(ch)
return total % size
def insert(table, key):
slot = hash_key(key, len(table))
while table[slot] is not None: # slot taken -> try the next one
slot = (slot + 1) % len(table)
table[slot] = key
return slot
table = [None] * 5
print(insert(table, "A")) # 0
print(insert(table, "F")) # 1 (A and F both hash to slot 0)
print(table) # ['A', 'F', None, None, None]
Looking up a key
- To find a key: hash it, then probe forward, comparing each slot to the key.
- Stop and return the index when you find it.
- If you reach an empty slot (or check every slot), the key is not there — return
-1.
Поиск ключа
- Чтобы найти ключ: хэшируйте его, затем выполняйте линейный пробинг вперед, сравнивая каждую ячейку с ключом.
- Остановитесь и верните индекс, когда найдете его.
- Если вы достигли пустой ячейки (или проверили все ячейки), ключа нет — верните
-1.
Why hash tables are fast
- With few collisions, insert and find take about one step — we call this
O(1). - As the table fills, probing gets longer, so it is wise to keep some slots free.
- In the worst case (everything collides) it slows to a linear scan,
O(n).
Почему хэш-таблицы быстрые
- При малом количестве коллизий вставка и поиск занимают около одного шага — мы называем это
O(1). - По мере заполнения таблицы пробинг становится длиннее, поэтому разумно оставлять некоторые ячейки пустыми.
- В худшем случае (все ключи collidят) скорость падает до линейного сканирования,
O(n).
Common mistakes
- A hash function maps a key to an index in the table.
- Two keys can collide at the same index — handle it, for example by chaining.
Распространенные ошибки
- Хэш-функция отображает ключ на индекс в таблице.
- Два ключа могут иметь коллизию в одном и том же индексе — обработайте это, например, используя цепочки.
Now you try
- Build the three parts: the hash function, insert with probing, and find.
- Each task checks your function on collisions and wrap-around cases.
- Press Check answer to test it.
Теперь попробуйте сами
- Постройте три части: хэш-функцию, вставку с пробингом и поиск.
- Каждая задача проверяет вашу функцию на случаях коллизий и возврата в начало (wrap-around).
- Нажмите Проверить ответ, чтобы протестировать её.
Hashing to a bucket · Хэширование в ячейку (корзину)
A hash function sends each key to a bucket; clashes chain. · Функция хэширования отправляет каждый ключ в ячейку; коллизии образуют цепочку.
Write hash_key(key, size). Add the character codes of key (use ord(ch)) and return the total % size, so the result is a slot from 0 to size - 1. Example: hash_key("AB", 10) is (65 + 66) % 10 = 1. The empty string gives 0. · Напишите hash_key(key, size). Сложите коды символов строки key (используя ord(ch)) и верните итоговую сумму % size, чтобы результат был индексом от 0 до size - 1. Пример: строка hash_key("AB", 10) дает (65 + 66) % 10 = 1. Пустая строка возвращает 0.
Click Run to see the output here. · Нажмите Запустить, чтобы увидеть результат здесь.
hash_key is provided. Write insert(table, key) using linear probing: go to the key's hash slot; while that slot is full (not None), step to (slot + 1) % len(table); put the key in the first free slot and return its index. · hash_key уже предоставлен. Напишите insert(table, key), используя линейную пробацию: перейдите в хеш-ячейку ключа; пока эта ячейка занята (не None), переходите к (slot + 1) % len(table); поместите ключ в первую свободную ячейку и верните его индекс.
Click Run to see the output here. · Нажмите Запустить, чтобы увидеть результат здесь.
hash_key is provided. Write find(table, key) that returns the index of key, or -1 if it is missing. Start at the hash slot and probe forward, comparing each slot. Stop at an empty slot (key not there). Make sure it ends even if the table is full — never loop forever. · Функция hash_key предоставлена. Напишите find(table, key), которая возвращает индекс элемента key или -1, если он отсутствует. Начните с ячейки хэша и зондируйте вперед, сравнивая каждую ячейку. Остановитесь при пустой ячейке (ключа там нет). Убедитесь, что функция завершается даже при полной таблице — никогда не зацикливайтесь бесконечно.
Click Run to see the output here. · Нажмите Запустить, чтобы увидеть результат здесь.