Simulating DigitalOcean's $5 SSD Droplets
Back in 2014, DigitalOcean was experiencing hyper-growth. With hyper-growth obviously came fast-scaling problems and growing pains. With new paradigms, new user behaviors show up. In this case, people loved Droplets (what DigitalOcean called its virtual-machine product) and kept making new ones and deleting them and making them again and so on. Overall growth with a good amount of the Droplet population having short lifetimes.
A Footlong Subway Sandwich for $5, or an SSD Droplet
When DigitalOcean hit PMF, the founders Ben and Moisey had come up with a catchy differentiator. Back then, cloud was still a burgeoning new field and along with a bunch of undifferentiated VPS providers, AWS was offering free micro VMs. Linode was a traditional VPS offering click-click-click VMs with spinning disks.
DigitalOcean had the hubris of wanting to not be a VPS-provider anymore and instead become a pure-play cloud provider with its delusional sights set on AWS. To do so, the gimmick was that we'd sell 512MB Droplets (VMs) with 20GB of locally attached SSDs, for $5/month. Back then, no one else was doing this. AWS's free micro-VMs had network-attached storage which was dog slow, and other VPS providers had spinning disks. DigitalOcean was all flash, and for only the cost of a footlong Subway sandwich (Ben just loved to talk about the footlong sandwiches).
Hey man, you can have a VM with SSDs for the price of a footlong Subway! Ain't this cool! - probably something said by Ben Uretsky (cofounder and then-CEO of DigitalOcean)
Yes Ben, it's cool. And it was cool enough that DigitalOcean is now a multi-billion-dollar cloud provider.
Lots of Tiny Droplets
So here I came and we had all this demand for our crazy $5 SSD Droplets. These Droplets were to be scheduled on hypervisor nodes. These nodes were big fat servers with just a bunch of expensive SSDs in them: no spinning rust whatsoever. And we were packing a bunch of tiny annoying 512MB VMs on these huge hypervisors, and each of these VMs would want to run the same few Ubuntu or Debian or whatever Linux image, and then a tail of other images.
These disk images were reasonably large blobs (memory guesstimate of p50 at 1 GiB, p90 at 5 GiB) and had to come from somewhere, and the hypervisor's SSDs weren't it because those were exactly what we were selling. The disk images were stored on a fleet of storage servers and every time a customer would order a Droplet from us, we'd find a hypervisor for it, tell the hypervisor code "fetch this N GiB image and run it". And then because these $5 Droplets were selling like hotcakes, we'd have a bunch of people all trying to pull all these images on the fleet of hypervisors all at the same time, completely soaking up the network and leading to nobody getting any sandwich. This was all initially done over NFS which had no notion of what a disk image is and how to handle it like a disk image.
Embrace the Pain, Become the Pain
This was not great. We wanted to be great. A few of my buddies and I (Nick Van Wiggeren, Mac Browning, Justin Hines, Matt Layher) embarked on a mission to move away from NFS and create an image management service. This thing was going to handle the entire lifecycle of disk images and "know" what a disk image was and the various things we needed from a thing that knows about and serves disk images. Many interesting things came out of this overall project, being the first "modern HTTP/JSON Go service" that DigitalOcean built outside of its initial Perl monolith. And at the tail of this project was one particular problem that I want to talk about today. And I alluded to it earlier: how our customers fancied churning through Droplets at a very high rate.
Back at Shopify, Tobi had embraced BFCM and flash-sales as pathological user patterns that were to be celebrated and made normal and fast. Having just arrived at DigitalOcean from there, I wanted the same thing. I wanted people to be able to churn Droplets and for us to handle it well using our smarts and brains. I wanted us to be good at this. To delight our customers with the fastest Droplet creation experience we could provide, at scale and celebrate these flash sales of Droplets as an awesome pattern that we'd accelerate and make fast.
So how do we get these Droplets to start as fast as possible, when 99% of the startup time was spent downloading large disk images from a small set of storage nodes?
Caching When You Don't Want To
The naive and simplest solution to this is to prestage the disk images on the hypervisors. But then we didn't know exactly which images would be needed where, and the demand was mixed enough that aside from the latest Ubuntu, what would come next was not obvious without constant finagling and reactive "I think people will want this image next" from engineers who'd manually stage disk images on some hypervisors. This is labor intensive and kind of hacky. And also, all these prestaged disk images are using space on the hypervisors' SSDs that we want to sell, not use as disk image storage...
Then the other next obvious solution is to cache images as they're requested. So then we need to... reserve some amount of hypervisor disk space for the cache, which is also not what we want to do with these SSDs: we want to sell them for $5 a month.
In addition, these $5 Droplets were the golden goose at the time, and people were damn scared of messing this up and making the Droplets not work anymore. On top of it, the giant monolith code running on the hypervisor was a mess of Perl that absolutely no one wanted to touch. Basically doing stuff on these hypervisors was raising the cortisol levels of anyone in charge of it. And we needed the buy-in of the team handling that codebase.
So some of our degrees of freedom were constrained:
- We didn't want to waste SSD space on prestaging
- We didn't want to waste SSD space on caching
- We were scared of touching the hypervisors in general
Hidden Degree of Freedom
But here's a thing that wasn't obvious.
- We ultimately did not want to reserve SSD disk space for caching because we wanted to sell it.
- We had the most disk space available for caching when the hypervisor had no Droplets on it (0% sold).
- We had the least amount of disk space available for caching when the hypervisor was becoming full.
- We could afford to waste disk space when the hypervisor was unsold.
- We couldn't afford to waste disk space when the hypervisor was becoming sold-out.
- We had some institutional legitimacy in touching this code, since we were switching the hypervisor stack from NFS to HTTP and were introducing a new path.
So basically, caching was fine as long as we got out of the way as the hypervisor was filling up. The degree of freedom: use disk space for caching inversely-proportional to disk usage for Droplets.
Blobcache: A Self-Shrinking Cache
So I designed blobcache, the self-shrinking cache. This cache would use some fixed amount of the hypervisor's SSDs and monitor a buffer to always keep free space of some buffer size.
max_cache_space := 4 GiB
buffer_space := 2 GiB
min_free_space := 10 GiB
It would look at the disk free space and figure out how much cache it was allowed to keep. Whatever headroom was left above min_free_space + buffer_space was fair game, up to max_cache_space:
headroom := disk_free() - min_free_space - buffer_space
allowed := clamp(cache_space + headroom, 0, max_cache_space)
Then every few seconds, it would evict stuff until the cache fit in there:
for cache_space > allowed:
evict(least_recently_used())
As Droplets got created and wrote to the disk, disk_free() would go down, allowed would go down with it, and the cache would give its space back. On an empty hypervisor, the cache could grow to 4 GiB. On a full one, it would shrink to nothing. The buffer_space meant the cache started getting out of the way before Droplets could feel any pressure.
Then there's how blobs got into the cache. Blobs were content-addressed: you'd ask blobcache for sha256:ab12cd, not for "ubuntu-14.04". Same bytes, same key, so a blob could never go stale. On a hit, blobcache served the file straight from disk. On a miss, it would fetch the blob over HTTP from the storage fleet, and tee the response to a staging file while hashing it along the way:
resp := http.Get(storage_url(digest))
staging := os.CreateTemp(staging_dir)
hasher := sha256.New()
return io.TeeReader(resp.Body, io.MultiWriter(staging, hasher))
So the reader got its bytes as they arrived, and a miss cost nothing more than downloading from storage directly. The caching happened on the side, for free.
When the stream hit the end, blobcache would check the hash:
if hasher.Sum() != digest:
os.Remove(staging)
return ErrDigestMismatch
if staging.size() > allowed:
os.Remove(staging)
return io.EOF
evict_until_fits(staging.size())
os.Rename(staging, cache_path(digest))
return io.EOF
If the hash didn't match, the staging file got thrown away and the reader got an error in place of a clean EOF, a warning that the bytes it just read were bad. If the hypervisor had filled up in the meantime, the blob was served but not kept. Otherwise, the staging file got renamed into the cache. The rename was atomic, so the cache only ever held complete, verified blobs, and a crash mid-fill never left anything half-written behind.
And because it was really about storing content-addressable blobs, it wasn't just for disk images. It was for anything blob-shaped you needed to fetch on hypervisors.
Cool Design, Cold Feet
This was (in my humble opinion) a simple and elegant solution. But one thing it was going against was the fact that people were terrified of running anything new on hypervisors. It wasn't that we couldn't do anything, but we had to soothe the fears with some convincing arguments for why it would help and how it would not hurt. Switching from NFS to HTTP was already stressing people out, and telling them that I'd also use some of their precious SSDs to shove a cache in there was pushing it a bit.
At the same time, we had just gained a fancy new capability: centralized logging. And we also had just started logging everything in structured log format as beautiful JSON. And our hypervisor backend code (lovingly named DOBE) was now logging in JSON at just the right spots.
When you face tough times, sometimes you need a montage and sometimes you need a simulation. We had all the ingredients we needed:
- A system that generates structured events across many hypervisors.
- These historical events are saved somewhere queryable.
- These events record "when a given hypervisor has requested a given disk image".
It was time for a simulation to show that introducing this cache would achieve some hit-rate sufficient to justify the risk.
Simulating Two Days of $5 Droplets
On one Tuesday evening, I launched a query against our centralized logging system to fetch the list of all log lines where a hypervisor had requested a disk image to be downloaded for the last two days.
I then modelled a Cloud region by enumerating all the hypervisors that had been involved in the input data. This gave me a list of hypervisors:
hypervisor := [
// list of all hypervisors
]
type Hypervisor struct {
cache LRU // holds up to max_size bytes of disk images
}
And then I implemented the simulation model. Replaying the two days of events against the hypervisor fleet, what would have been the hit rate if:
- There was a cache using various
eviction_policy? - The cache had various
max_sizelimits?
And the simulation went like this:
- An event comes to schedule
dropletontohypervisorusingdisk_image_ubuntuof sizeN. - If
hypervisoralready hasdisk_image_ubuntu, record a cache hit to the output event stream. - If not, record a cache miss to the output event stream. Then pretend that
disk_image_ubuntuis now cached onhypervisor. Apply whatevereviction_policyto make space if needed. - Repeat across the whole set of historical events.
It's pretty simple, we basically pretend that the cache is there, and we replay the events that actually happened in the past, and we can then say:
If we had had this cache of
$sizewith$eviction_policy, the cache hit rate would have been N.
I reran the simulation at full speed for various $size to show that we would have what looked like a logarithm function in improvements in cache-hits: a fast increase in hit-rate until the knee of the curve, and the knee would be somewhere around 4 GiB of disk space kept for the cache. Beyond that, there was little change and that's what the proposal shipped with to alleviate worries about excessive disk utilization. But as you could guess, the run-away-before-the-buffer would have allowed basically using the entire disk as cache, but a cap was applied to keep people calm and pass the RFC review process.

By now it was Thursday, I had done all this work over a 48h burst of frenzy coding. I wrote my argument in an RFC, dropped it on Slack for others to read and review and went to sleep until Monday.
Warming Up to the Idea
And so my Monday when I opened my laptop again, blobcache was approved. I had also already implemented most of it.
So along with the design of blobcache shipped a simulation that showed what having blobcache would have gained us in the past. This wasn't just a "trust me it'll probably be good", it was a "if we had it, it would have helped us in this convincingly predictable manner, and behaved this way". This justified a trial on some few hypervisors, and after proving that the backend code was not flipping over in despair from the cache's disk usage and presence (there had been a worry that utilizing disk space for caching would have broken some obscure accounting for scheduling purpose), rolled out across the fleet.
It went on to live a happy life of high hit-rate and absorbing high-churn of Droplet for common disk images for a long many years. As far as I can tell, it was still running years after I left.
What Came Next
This was not the last modelling and simulation I did there. Rolling off the success of image management, Nick, Justin, and I moved on to introduce the first non-Droplet new DigitalOcean product: Volumes (network attached block storage). This involved more different types of simulations - flow networks.
As for blobcache, it was left alone until after Volumes. I briefly resumed work on a v2 that did gossip peer-to-peer and chunked the blobs into fixed sizes using merkle-trees - bittorrent style - and abandoned it mid-development since v1 was largely good enough, and new work was on the board. It was bittersweet, but I had the opportunity to start prototype work on a container runtime service that morphed into a Kubernetes product, the culmination of my services at DigitalOcean; the launch of DOKS on the road to the IPO.