Lesson 1
How Redis picks its in-memory structure
HSET does not tell you what Redis just used to store that hash. For a small hash it is
one contiguous run of bytes; past some threshold it becomes a hash table with a pointer for every
field. Same command, same data type, several times the memory. OBJECT ENCODING is the
command that says which one you got.
One key, two layers
Every Redis key has a type that you see (TYPE returns hash,
set, zset, list, string) and an
encoding that is how it is really stored (OBJECT ENCODING). You choose the
type; Redis chooses the encoding, from the element count and the value lengths measured against
the configuration thresholds.
| Type | Compact encoding | Full encoding | Deciding thresholds |
|---|---|---|---|
hash | listpack | hashtable | hash-max-listpack-entries, hash-max-listpack-value |
set, all integers | intset | hashtable | set-max-intset-entries |
set with strings | listpack | hashtable | set-max-listpack-entries, set-max-listpack-value |
zset | listpack | skiplist | zset-max-listpack-entries, zset-max-listpack-value |
list | listpack | quicklist | list-max-listpack-size, plus a hard 8,192-byte limit |
string | int, embstr | raw | 44 bytes, not configurable |
112 combinations, measured
The rig builds each key with a single EVAL and then calls
OBJECT ENCODING, across four collection types × seven element counts (1, 127, 128,
129, 511, 512, 513) × four kinds of value (3-byte string, 64-byte string, 65-byte string,
integer). The thresholds are the Redis 7.4 defaults, except that
list-max-listpack-size is set to 128 so there is a positive number to compare against.
docker exec rdlab redis-cli EVAL "
for i=1,tonumber(ARGV[1]) do
redis.call('HSET', KEYS[1], 'f'..i, ARGV[2])
end
return redis.call('HLEN', KEYS[1])" 1 h513 513 v
docker exec rdlab redis-cli OBJECT ENCODING h513
# hashtable
These are the rows out of the 112 that say the most:
| Key | Elements | Value | ENCODING | MEMORY USAGE |
|---|---|---|---|---|
hash | 128 | 3-byte string | listpack | 1,584 |
hash | 511 | 3-byte string | listpack | 6,192 |
hash | 512 | 3-byte string | listpack | 6,192 |
hash | 513 | 3-byte string | hashtable | 28,816 |
hash | 1 | 64-byte string | listpack | 128 |
hash | 1 | 65-byte string | hashtable | 248 |
set | 512 | integers | intset | 1,328 |
set | 513 | integers | hashtable | 24,712 |
set | 128 | 3-byte string | listpack | 816 |
set | 129 | 3-byte string | hashtable | 6,280 |
zset | 128 | 3-byte string | listpack | 1,072 |
zset | 129 | 3-byte string | skiplist | 13,572 |
string | — | 44 characters | embstr | 96 |
string | — | 45 characters | raw | 112 |
string | — | 12345 | int | 48 |
set-max-listpack-entries is 128, yet a set of 200
integers is still intset rather than hashtable — because for an
all-integer set the threshold that applies is set-max-intset-entries (512). And when
it does pass 512 it goes straight to hashtable, skipping listpack
entirely. Add one non-integer member to that 200-member set and it becomes
hashtable immediately; add one to a 100-member set and it becomes
listpack.
The price of a 513th field
Two keys that differ by a single field:
| Key | Fields | ENCODING | MEMORY USAGE | Bytes per field |
|---|---|---|---|---|
h512 | 512 | listpack | 6,192 | 12.1 |
h513 | 513 | hashtable | 28,816 | 56.2 |
One extra field makes the key take 4.65 times the memory, and the cost per field
rises from 12.1 to 56.2 bytes. The reason is that a listpack stores the fields
back to back in one byte block, with no pointers and no buckets, while a hashtable
needs a bucket array plus a dictEntry for every field.
In practice: if your application holds millions of small hashes, keeping them under the threshold
is the largest saving available without changing a line of code. The reverse is also true —
raising hash-max-listpack-entries too far means every single-field read has to scan
the whole listpack.
One small detail worth recording: h511 and h512 both take 6,192 bytes.
That is not a sampling artefact of MEMORY USAGE — re-running with
SAMPLES 0, which counts everything instead of sampling, still gives exactly 6,192.
A listpack grows in steps, so sometimes one more field costs nothing.
Lists: the limit is in bytes, not elements
list-max-listpack-size defaults to -2, and a negative value means the
limit is a size rather than a count. But even when you set it to a large positive number, a hard
limit remains: one listpack node never exceeds 8,192 bytes.
Binary-searching for the boundary with list-max-listpack-size 100000:
| Element length | Last count still listpack | First count as quicklist | MEMORY USAGE at the boundary |
|---|---|---|---|
| 8 bytes | 818 | 819 | 8,240 |
| 64 bytes | 122 | 123 | 8,240 |
| 200 bytes | 40 | 41 | 8,240 |
All three boundaries stop at a MEMORY USAGE of 8,240 bytes: the 8,192-byte listpack
plus roughly 48 bytes of object header. The lab below recomputes the listpack size as 6 bytes of
header + 1 terminator + per element (contents + length prefix + backlen), and that
formula reproduces all three boundaries exactly.
Conversion only goes one way
Redis converts from the compact encoding to the full one when a threshold is crossed, and never converts back. Measured:
hash with 513 fields → hashtable
CONFIG SET hash-max-listpack-entries 1024 → still hashtable
HSET one more field → still hashtable
HDEL down to 512 fields → still hashtable
a NEW key with 513 fields (threshold 1024) → listpack
The same holds for sets: 129 strings become hashtable, and removing one to get back
to 128 leaves it hashtable.
CONFIG SET only affects keys created after it. Raising a threshold on a
running instance shrinks nothing that is already large; to reclaim the memory you have to rewrite
those keys from scratch, or reload the dataset from RDB/AOF after changing the configuration. That
is why threshold tuning belongs in the config file before startup, not in a command issued during
an incident.
The lab
Set the type, the element count, the value length and each threshold. The algorithm running in your browser is a reimplementation of the rules above, checked against 112 matrix rows, three list byte boundaries and four string cases on Redis 7.4.11 — all of them agree.
The key you want to try
The measured cases:
Configuration thresholds
Fields that differ from the default change colour.
What OBJECT ENCODING returns
Check yourself
A hash has 10 fields, one of which holds a 200-byte string. What is the encoding?
hashtable. The field count is nowhere near the 512 threshold, but hash-max-listpack-value is 64, and a single oversized value is enough to convert the whole key. This is the most common reason a small hash costs more memory than expected.
A set of 300 integers. What is the encoding, and what does it become if you add the string "x"?
Before: intset (300 ≤ 512). Adding a string means it is no longer all integers, so set-max-listpack-entries (128) now applies — 301 members exceed it, giving hashtable. If the set had only 100 integers, the same addition would have produced listpack.
You set list-max-listpack-size 10000. A list of 500 elements, 64 bytes each — what is the encoding?
quicklist. The element count is far below the threshold, but 500 × roughly 67 bytes is already over the hard 8,192-byte limit of one listpack node. The measured boundary at 64 bytes per element is exactly 122 elements.
An instance holds a hash with 20,000 fields. You run CONFIG SET hash-max-listpack-entries 50000. Does the memory drop?
No. Conversion is one-way: a key that is already hashtable stays hashtable, even after the threshold is raised and even after fields are added or removed. Only keys created after that command benefit.