Profile
Back to NewsBack
Dev.to 10 min
Reader Mode
Redis vs Dragonfly: A Hands-On Comparison

Redis vs Dragonfly: A Hands-On Comparison

5 hours ago

I have used Redis on almost every backend project I can remember. Caching, sessions, rate limiting, queues, the occasional bit of temporary state that felt too awkward to put in Postgres. At some point it stopped being a decision and became a default.

About 3 months ago, I came across Dragonfly, which keeps the Redis API but is built differently underneath. I was curious enough to run both on the same workload and see what happened. This post covers what I saw, what I learned about why each one works the way it does, and where I think each makes sense.

Disclaimer: This is not a sponsored post, and I have no affiliation with DragonflyDB. I wrote it because I was genuinely impressed by Dragonfly when I tried it. It performed very well for me. Whether it is the right choice still depends on your workload, so treat this as a guide for evaluating it, not a recommendation to switch.

A quick note on what Redis is

Calling Redis a cache undersells it. Over the years I have used it for sessions, counters, distributed locks, leaderboards with sorted sets, pub/sub, streams, and plain queues. The reason it works so well for all of that is that it gives you useful data structures with very cheap operations. A command is usually a single call:

GET user:123
INCR rate_limit:user:123
ZADD leaderboard 9000 adam

There is no SQL to parse and no query planner involved. That simplicity is a big part of why it is fast.

So, why is Redis single-threaded anyway?

Redis executes commands on one main thread, and this tends to get described as a limitation. I think that undersells the original design.

Redis was created in 2009. Servers back then typically had a handful of cores at most, and the main goal was a datastore that was simple, predictable, and very fast on that hardware. A single execution thread gave it several things:

  • No locks around data structures.
  • Commands run one at a time, so atomicity is easy to reason about.
  • Behavior is predictable, which makes it easier to debug.
  • The codebase stays small and understandable.

Given the hardware of the time, this was a sensible design, and it still holds up well. Node.js is a good comparison: being mostly single-threaded does not make it slow, and it avoids a whole category of concurrency bugs.

Redis has also kept evolving. It has supported I/O threading since version 6, and Redis 8 improved that further. Networking work can be spread across threads while command execution stays serialized, so the core model is preserved and the obvious bottleneck is reduced.

Redis is a good design for the constraints it was built around. Those constraints have shifted since then, which is where Dragonfly comes in.

What Dragonfly does differently

Dragonfly started from a different assumption: servers now have 16, 32, or more cores, and a single execution thread leaves most of them idle. It uses a multi-threaded, shared-nothing architecture. The dataset is split into shards, each shard is owned by one thread, and requests are routed to the thread that owns the data.

Redis

  Requests --> [ Main thread ] --> [ Data ]


Dragonfly

  Requests --> [ Thread 1 ] --> [ Shard A ]
           --> [ Thread 2 ] --> [ Shard B ]
           --> [ Thread 3 ] --> [ Shard C ]

The real implementation is more involved, especially for operations that touch several shards, but that is the idea. The gain comes from spreading work across every core on the machine.

Trying it out

The first thing I liked was how little effort it took to start. Dragonfly speaks the Redis protocol, so I pointed an existing connection string at it and my client libraries worked without changes.

I didn't run some elaborate benchmark suite for this article. Honestly, you should benchmark it against your own workload anyway. A synthetic GET/SET benchmark tells you very little about whether Dragonfly will improve your actual application.

The interesting part is what happens when your Redis workload becomes CPU or memory bound. That's where Dragonfly's multi-threaded architecture and memory efficiency start to matter.

If your Redis instance is barely using one core and isn't under memory pressure, there may be little reason to switch.

And that's fine. Not every piece of infrastructure needs to be replaced just because something newer exists.

Where Dragonfly is still rough

Once the initial excitement wore off, I went through the Dragonfly docs and GitHub issues to see what could go wrong. Here is what stood out. None of it ruled Dragonfly out for me, and all of it is worth knowing before you put real data on it.

Persistence is snapshot-only

Redis gives you RDB snapshots and AOF logging. Dragonfly's docs say plainly that AOF is not supported, and the stated reason is a lack of high-priority demand from the community. If you need every acknowledged write to survive a crash, that matters.

Snapshots have had their own problems for some users. One GitHub issue describes snapshots to S3 silently stopping for a week, after which a crash restored week-old data. It was closed as a duplicate of another issue, and I could not tell from the thread what the root cause was, so treat it as one user's report.

Lua scripts need a closer look

Dragonfly uses Lua 5.4, while Redis uses Lua 5.1. Most scripts will run fine, but a language version change is worth testing for.

The bigger difference is how keys are handled. By default, Dragonfly rejects scripts that touch keys they did not declare, and returns an error saying so. You can allow it with the allow-undeclared-keys flag, but the docs warn that Dragonfly then has to stop all other operations while the script runs. Scripts that build key names dynamically are the ones to watch if you migrate.

Multi-key operations work, with a coordination cost

I expected multi-key commands to be a weak spot, since the data lives on different threads. Dragonfly handles them with a transaction framework based on the VLL algorithm, and the project says this gives atomic multi-key operations without mutexes or spinlocks. So atomicity holds.

The cost is coordination. In the team's own write-up on transactions, multi-key operations take a shard lock, and other transactions on that shard wait while it is held. Single-threaded Redis does not pay this cost. I did not measure how much it matters, so if your workload leans on large multi-key commands or scripts, test that path specifically.

Clustering works differently

Dragonfly's cluster mode is built to look the same to Redis Cluster clients, but it does not self-manage the way Redis Cluster does. In Redis Cluster, nodes talk to each other to discover cluster state. Dragonfly Cluster uses centralized management, so you set it up and operate it differently. If your plan is to run dozens of nodes, Redis Cluster has the longer track record.

Modules

Dragonfly's own comparison guide notes that some advanced Redis modules may not have full compatibility yet. If your stack depends on modules such as search or time series, check your exact use case before planning a migration.

Bugs

Dragonfly is a younger project and it shows in the issue tracker. A few examples from the release notes, all listed as fixed:

  • A crash on the master when an older replica reconnected to a promoted master (#7491).
  • Replication divergence with hash field expiry commands like HEXPIRE and HSETEX on lazily expired fields (#7948).
  • A crash when Lua scripts called GET on large string values (#7934).
  • A cross-thread data race in command squashing (#7927).

There are also open or unresolved user reports, such as a SIGSEGV crash during load testing on a 31 GiB dataset in v1.28.0, which showed up on a replica running BGSAVE and on a newly promoted master during full resyncs. The reporter could not reproduce it on demand.

Every database has a bug list and Redis does too. The point is that Dragonfly's list is newer and shorter on history, and a lot of the replication and persistence paths are where the issues cluster. That is worth keeping in mind if this is your primary datastore.

Where Redis still comes out ahead

None of that makes Redis the worse choice in general. After using both, here is where I think Redis still wins.

Maturity. Redis has years of production use behind it, and the failure modes are well understood. When something breaks at 3 AM, there is a huge body of knowledge to lean on. Dragonfly is still building that history.

Durability options. AOF, RDB, and all the configuration around them are well tested. If you need stronger durability guarantees than periodic snapshots, Redis gives you more to work with today.

Ecosystem. Client libraries, framework integrations, monitoring tools, managed services on every major cloud, local dev setups, CI containers. All of it already supports Redis, and your team almost certainly knows it.

Behavior you can predict. Dragonfly is compatible with the protocol and most commands, but compatibility does not guarantee identical behavior. Scripting, clustering, persistence, and replication are the areas above where it differs. Test the specific commands you rely on before trusting it with anything important.

Enough headroom. If your Redis instance sits at 8% CPU and your app spends most of its time waiting on Postgres, Dragonfly will not make your app noticeably faster. A benchmark chart does not tell you what is limiting your system. Profiling does.

How I would choose

I would stay with Redis when:

  • the workload is comfortably within what Redis handles
  • the team already operates it and knows it well
  • I depend on a specific Redis feature, module, or managed service
  • I need AOF-level durability
  • familiarity is worth more than extra hardware efficiency

I would evaluate Dragonfly when:

  • the workload is CPU-heavy
  • I am running large multi-core machines
  • memory cost matters
  • Redis Cluster complexity is becoming a burden
  • vertical scaling looks more attractive than adding nodes
  • snapshot-based persistence is acceptable for the data

Replacing something that works has a cost, so it needs a measurable reason.

Ending

I don't think Dragonfly makes Redis obsolete. And I don't think Redis being older makes it a bad choice either.

They are solving a similar problem with different architectural decisions. Redis optimized heavily around simplicity and predictable command execution. Dragonfly is taking a different approach and making much better use of the hardware we have today.

That's what I found interesting about it.

The question isn't really which one is better. It's why was it built this way in the first place?

Once you understand that, a lot of the design decisions stop looking strange.

If Redis is working fine for you, keep using it. If you're starting to hit CPU, memory, or scaling limits, Dragonfly is absolutely worth putting next to it and seeing what happens with your workload.

Sometimes the best way to understand a piece of infrastructure is to stop reading benchmarks and just ask what problem its architecture was trying to solve.

Chat with me