Alex Rivera | Logout

How does django handle multiple memcached servers?

Asked 2011-07-29T16:36:20.640
21

In the django documentation it says this:

...

One excellent feature of Memcached is its ability to share cache over multiple servers. This means you can run Memcached daemons on multiple machines, and the program will treat the group of machines as a single cache, without the need to duplicate cache values on each machine. To take advantage of this feature, include all server addresses in LOCATION, either separated by semicolons or as a list.

...

Django's cache framework - Memcached

How exactly does this work? I've read some answers on this site that suggest this is accomplished by sharding across the servers based on hashes of the keys.

Multiple memcached servers question

How does the MemCacheStore really work with multiple servers?

That's fine, but I need a much more specific and detailed answer than that. Using django with pylibmc or python-memcached how is this sharding actually performed? Does the order of IP addresses in the configuration setting matter? What if two different web servers running the same django app have two different settings files with the IP addresses of the memcached servers in a different order? Will that result in each machine using a different sharding strategy that causes duplicate keys and other inefficiencies?

What if a particular machine shows up in the list twice? For example, what if I were to do something like this where 127.0.0.1 is actually the same machine as 172.19.26.240?

CACHES = {
    'default': {
        'BACKEND': 'django.core.cache.backends.memcached.MemcachedCache',
        'LOCATION
Edit
Report

1 Answer

5

I tested part of this and found some interesting stuff with django 1.1 and python-memcached 1.44.

On django using 2 memcache servers

cache.set('a', 1, 1000)

cache.get('a') # returned 1

I looked up which memcache server 'a' was sharded to using 2 other django setups each pointing at one of the memcache servers. I simulated a connectivity outage by putting up a firewall between the original django instance and the memcache server that 'a' was stored in.

cache.get('a') # paused for a few seconds and then returned None

cache.set('a', 2, 1000)

cache.get('a') # returned 2 right away

The memcache client library does update its sharding strategy if a server goes down.

Then I removed the firewall.

cache.get('a') # returned 2 for a bit until it detected the server back up then returned 1!

You can read stale data when a memcache server drops and comes back! Memcache doesn't do anything clever to try to prevent this.

This can really mess things up if you're using a caching strategy that puts things in memcache for a long time and depends on cache invalidation to handle updates. An old value can be written to the "normal" cache server for that key and if you loose connectivity and an invalidation is made during that window, when the server becomes accessible again, you'll read stale data that you shouldn't be able to.

One more note: I've been reading about some object/query caching libraries and I think johnny-cache should be immune to this problem. It doesn't explicitly invalidate entries; instead, it changes the key at which a query is cached when a table changes. So it would never accidentally read old values.

Edit: I think my note about johnny-cache working ok is crap. http://jmoiron.net/blog/i

answered 2011-10-28T18:12:56.067

Your Answer