CSIPE

Published

- 20 min read

A Deleted Container Can Leave Data Behind


Books by the author

Compare all 5

As an Amazon Associate I earn from qualifying purchases. Buying through these links costs you nothing extra and helps pay for the blog.

A fresh container should begin with an empty disk. In September 2026, a security researcher showed that some fresh Cloudflare containers could receive something else: fragments left in storage blocks previously used by another customer’s workload. The new customer could not choose whose data appeared, and the old customer’s running disk was not exposed. Still, the wall between tenants had a gap.

Cloudflare repaired the gap across its Containers fleet. The company says customers do not need to change any configuration, and its review of retained disk activity found no evidence of malicious exploitation. Those facts deserve equal weight with the flaw itself. This was a serious isolation failure, followed by a quick fix and an unusually useful public explanation.

The lasting lesson sits below the product name. A container boundary has a life before the workload starts and after it stops. Memory is cleared, storage blocks are reassigned, cached images are reused, credentials expire, and logs decide whether anyone can reconstruct what happened. If a security review examines only the running process, it misses the machinery that prepares and retires the process.

For platform teams, the practical question is simple: when one customer releases a resource, what proves that the next customer receives a clean one?

What Cloudflare disclosed

On 4 September 2026, Oren Yomtov of Accomplish reported a vulnerability affecting Cloudflare Containers and Cloudflare Sandboxes through Cloudflare’s bug bounty programme. Cloudflare published its account on 24 September after completing remediation. The affected services run customer workloads on shared infrastructure, and customers cannot select the physical host beneath a container (Cloudflare).

The researcher used a paid Workers account to create containers and inspect their writable root disks. Across six production placements, the test identified 5,614 directory blocks that did not belong to the researcher’s test filesystem. Checksum analysis distinguished 2,700 foreign directory inodes. Residual material appeared on 18 of 24 placements and 20 of 22 underlying nodes across four continents. The observed structures included directory entries, database pages, and complete SQLite database structures (Cloudflare).

Those numbers describe the controlled research, not a count of harmed customers. Cloudflare says the researcher supplied aggregate counts, offsets, sizes, and truncated hash prefixes rather than third-party content values. The researcher also confirmed that recovered data stayed confidential and was securely deleted after the report. As of 28 September, Cloudflare says it found no evidence that anyone else used this technique maliciously.

Independent reporting on 27 September matched the core account: a Workers Paid customer could recover residual blocks from containers that had previously occupied the same physical host, while the technique could not select a victim, target an active disk, or reliably produce useful data (BleepingComputer). The uncertainty matters. This was neither a public database that anyone could browse nor a demonstrated campaign stealing named customers’ records. It was a repeatable way for one tenant to cross into another tenant’s discarded storage.

Cloudflare’s timeline shows a rapid first response. The company confirmed the production condition at 18:45 UTC on 4 September, a little over three hours after the report. It merged the first runtime fix that evening, began the fleet rollout before midnight, and completed the initial rollout on 7 September. The researchers reported on 14 September that their proof of concept no longer worked. Cloudflare finished clearing pre-fix cached snapshots across the affected fleet on 19 September (Cloudflare).

That last date tells an important part of the story. Changing the setting stopped new unsafe allocations, but it did not clean every block already mapped into a running disk or cached image layer. The repair needed a second phase: retire existing container disks, drain hosts, restart virtual machines, and remove old image snapshots. A configuration fix protected the future. A cleanup campaign dealt with the past.

Customers have no vendor-directed patch to install. Cloudflare says remediation requires no customer-side configuration change. A team using the service should record the disclosure, decide whether its workloads carried data that changes its own risk assessment, and avoid inventing an emergency product change that the evidence does not support.

How a small write exposed an older block

The mechanism is easier to understand if we stop thinking in files for a moment. Storage systems often work with fixed-size blocks underneath the folders and filenames an application sees. A file can disappear from a directory while its old bytes remain in an unallocated block until the system overwrites or clears them.

Cloudflare Containers used Linux device-mapper thin provisioning, usually shortened to dm-thin, for writable root disks. Thin provisioning gives each workload a virtual disk but allocates physical storage only when the workload writes to a previously unused part of that disk. It saves capacity because an apparently large virtual disk does not need an equally large physical allocation on day one. The Linux kernel documentation describes the feature as mapping virtual blocks to a shared pool as writes occur (Linux kernel documentation).

The affected Cloudflare pools allocated physical storage in 64 kibibyte blocks. Their configuration also included skip_block_zeroing, an option whose name is unusually direct: it skips clearing a newly provisioned block before handing it to the new thin device. The kernel documentation lists this as an optional thin-pool feature. In a pool serving different customers, skipping that clearing step changed a performance setting into an isolation decision.

An untouched part of a new virtual disk still appeared to contain zeroes. That detail could make a normal check look healthy. Reading an unmapped region did not force dm-thin to assign physical storage, so the system simply returned zeroes. The older bytes became reachable only after a write caused an old physical block to be reassigned.

The proof of concept wrote 4 KiB into selected regions aligned to the 64 KiB storage blocks. That small write caused the pool to allocate a physical block. The new 4 KiB replaced its matching portion, but the remaining 60 KiB could still contain bytes from the previous owner because the allocation had not been cleared. A raw read of the disk could then see data that the new container had never written (Cloudflare).

Imagine a hotel replacing only the name card outside a room. A new guest opens the door, places one small bag in the corner, and then finds papers in the other cupboards. They cannot ask for a particular former guest’s room, and many rooms may already be empty. The allocation system still failed at its central promise: the new key should open clean space.

The filesystem made the recovered fragments easier to classify. Ext4 can attach checksums to directory blocks using values linked to a filesystem and inode. The researchers used those checksums to separate blocks created in their controlled test filesystem from blocks that came from somewhere else. They validated the method against 162 blocks they had deliberately created and deleted, then applied it to production placements. This was stronger evidence than spotting a stray readable string and guessing where it came from.

The same distinction matters in a platform test. “The disk looked empty when we mounted it” tests the file view. “Every newly allocated block is zeroed before a different tenant can read it” tests the storage boundary. The first can pass while the second fails.

The virtual machine wall did not cover the storage pool

Each Cloudflare container ran inside a dedicated virtual machine powered by Firecracker, according to the company’s incident report. That is a meaningful boundary. A separate virtual machine gives a workload its own guest kernel and makes many direct interactions with neighbouring workloads harder.

The stale blocks arrived through a layer shared below those virtual machines. The guest received a virtual root disk at /dev/vdc, but physical blocks came from a host pool that had served many customer accounts. The virtual machine could isolate one running guest from another while the allocator passed old bytes from one guest’s former disk into a later guest’s disk.

This is why security architecture cannot stop at product nouns. “Container,” “sandbox,” and “virtual machine” describe useful mechanisms, but none of them is a complete boundary diagram. The actual boundary includes the image cache, writable layers, host kernel, hypervisor, storage allocator, network path, control plane, identity service, snapshot lifecycle, and cleanup process.

Cloudflare’s Containers documentation presents the service as a way to run container images beside Workers for workloads that need custom runtimes or more resources (Cloudflare Containers documentation). A customer interacts with the service at that level. The customer does not set dm-thin pool flags on Cloudflare’s hosts. Shared responsibility therefore has a sharp edge here: customers choose what data and authority enter a workload, while Cloudflare owns the host-level mechanism that keeps one customer’s allocation separate from another’s.

Owning a layer does not mean keeping it invisible after failure. Cloudflare’s public account named the exact option, block sizes, partial-write behaviour, validation method, fleet cleanup, available telemetry, and limits of its conclusion. That gives customers enough information to assess the failure without pretending they can patch provider infrastructure themselves.

The incident also exposes a common testing blind spot. Teams tend to test whether a workload can escape while it is running. They try forbidden system calls, host paths, metadata endpoints, network routes, and management interfaces. Those tests matter. Yet a temporary workload also crosses boundaries during creation, snapshot restore, image preparation, scaling, suspension, destruction, and block reuse.

A good boundary test follows the resource through all of those states. Create tenant A’s workload and write marked data. Destroy it. Create tenant B’s workload on a reused pool. Trigger the same allocation paths a real filesystem would use. Read through both normal filesystem operations and the raw device interfaces available to the guest. Repeat enough placements to catch a condition that depends on scheduling. The Cloudflare research found residual material on 18 of 24 placements, which is exactly the kind of result a single happy-path test could miss.

The Secure Harness makes the same point for coding agents: the boundary has to cover every route by which authority or data can move, not merely the process labelled “agent.” Containers deserve the same discipline. A strong runtime wall does not repair a weak retirement path beneath it.

Deletion, deallocation, and sanitisation are different events

Developers often use “deleted” to describe several states that the storage system treats differently. A file can be removed from a directory. A filesystem block can be marked free. A virtual-disk mapping can be released. A physical block can return to a shared pool. The old bytes can remain readable until another operation overwrites or clears them.

None of this means deletion is pointless. It means deletion answers a naming and allocation question before it answers a confidentiality question. The operating system no longer presents the old file through its ordinary path, but a lower layer may still hold the bytes.

The security term for making target data infeasible to recover is sanitisation. In September 2025, the US National Institute of Standards and Technology published Revision 2 of its media sanitisation guidance. NIST defines sanitisation by the recovery effort left to an adversary, rather than by whether a file icon disappeared or a volume was detached (NIST SP 800-88 Rev. 2).

Cloud storage adds an abstraction problem. A customer may never see the physical device, and a provider may move or copy data through snapshots, caches, replicas, and replacement hardware. The customer cannot run a disk-erasure command against a server it does not control. It needs the provider to enforce safe reuse below the service interface and to describe how that enforcement is tested.

There are two separate moments to protect. First, storage should be safe before a new tenant receives it. Clearing a reused allocation is one way to achieve that. Second, old mappings and cached layers created under an unsafe policy may need retirement after the policy changes. Cloudflare’s response covered both moments: restore block zeroing for new allocations, then replace or clear the state that predated the repair.

Encryption can reduce the value of leftover bytes, but the details matter. If an application writes plaintext into its root filesystem, host-level disk encryption may protect against a stolen physical drive while still allowing the storage service to decrypt blocks for whichever guest receives them. The threat in this incident occurred after the platform had made a block available through its normal storage machinery. Encryption at the wrong layer does not restore tenant separation.

Application-level encryption with keys kept outside the container can narrow the consequence. A recovered database page may still be unreadable if the page was encrypted before storage and the old workload’s key is unavailable to the new workload. That approach carries operational costs, and it does not protect filenames, sizes, access patterns, or any plaintext the application writes to temporary space. It should support provider isolation, not excuse a broken allocator.

Secrets make the timing sharper. A long-lived application programming interface key written to a temporary file can remain useful months later. A short-lived workload credential may be useless by the time anyone recovers its bytes. Short lifetime does not prevent exposure, but it reduces the period in which exposed material carries authority. The practical model is layered: safe storage reuse, minimal sensitive data on ephemeral disks, encryption at a layer with separate keys, and credentials that expire quickly.

The lesson also applies to build runners, notebook services, preview environments, browser sandboxes, and coding-agent machines. These systems feel disposable because the interface has a “delete” button or a short time-to-live. Disposal is a workflow, not a label. Ask what happens to writable layers, caches, swap, snapshots, uploaded files, build artefacts, and logs after the user-facing object disappears.

What customers should do now

Cloudflare says the vulnerability is fixed and customers do not need to change a service setting. That should be the starting point. Rebuilding every application or rotating every company credential without evidence would consume attention while doing little to improve the boundary Cloudflare already repaired.

A sensible customer response has two tracks. One records and assesses this specific disclosure. The other uses the incident to improve design choices that remain valuable after the headline fades.

Start with scope. Confirm whether your organisation used Cloudflare Containers or Cloudflare Sandboxes before the fleet cleanup completed on 19 September 2026. Do not infer exposure merely because another Cloudflare product appears on an invoice. Record which applications ran there, what dates they ran, whether they wrote customer data or credentials to the writable root disk, and whether that material was encrypted by the application before storage.

Then classify consequence rather than guessing probability. A transcoding job that reads a public video and emits a public thumbnail has a different consequence from a sandbox processing private source files, health records, legal documents, or production database exports. A build container that receives a five-minute identity token differs from one that writes a year-long deployment key into its home directory.

Cloudflare says it found no evidence of malicious exploitation in the historical disk input/output telemetry it retained. Preserve that statement accurately. “No evidence in retained telemetry” is useful reassurance, but it is not the same claim as mathematical proof that no byte was ever observed. The report explains the signatures Cloudflare derived from the researcher’s write-and-read pattern and says only authorised validation matched them. Your incident note should carry both the reassuring result and the stated evidence boundary.

Customers with contractual, legal, or sector-specific reporting duties may need a direct provider conversation. Ask for the incident reference, affected-service scope, remediation dates, telemetry conclusion, and any assurance available under your agreement. The goal is a durable record, not a speculative breach declaration. Legal and privacy teams can then decide whether the nature of data processed in the service creates further obligations.

Credentials deserve a targeted review. If a workload placed long-lived secrets in its writable disk during the affected period, list those secrets and the authority each one carried. Consider rotation based on sensitivity, lifetime, evidence, and the cost of misuse. Do not rotate unrelated credentials merely to produce a larger incident spreadsheet. A narrow, completed rotation is better than an indiscriminate campaign that misses the one deployment key stored in a cache directory.

Finally, keep the provider’s “no customer action” statement separate from your own architectural actions. You do not need to alter Cloudflare’s fixed pool configuration. You may still choose to remove plaintext secrets from ephemeral disks, shorten token lifetimes, or add provider-disclosure intake to your platform process. Those improvements address your side of the boundary.

A practical storage-reuse review

The best response to this incident is a small review that produces evidence rather than a large policy document. One engineer should be able to trace a temporary workload from creation to retirement and show where sensitive bytes can remain.

  1. Draw the real data path. Begin with the input and follow it through the workload’s memory, writable root disk, attached volumes, temporary directories, image layers, caches, snapshots, logs, crash dumps, and exported artefacts. Mark which components belong to your team, your cloud provider, and another vendor. A box named “sandbox” is too coarse to answer where a deleted file goes.

  2. Name the isolation promise at each shared layer. For storage, ask whether new allocations are cleared, encrypted with tenant-separated keys, or both. For caches, ask whether entries are partitioned by tenant and how eviction works. For snapshots, ask when stale copies disappear. Turn “tenants are isolated” into a set of behaviours someone can test.

  3. Test reuse, not only first creation. A clean disk from a newly built host proves little about the hundredth workload on a busy host. In systems you control, write marked non-sensitive data as tenant A, destroy the resource, then allocate as tenant B and inspect every interface tenant B is allowed to use. Include partial writes, raw-device reads where permitted, snapshot restore, scale-down, and cache reuse. Run multiple placements.

  4. Reduce valuable residue. Keep secrets out of writable layers where a workload identity service or in-memory handoff can provide them when needed. Use scoped credentials with short expiry. Avoid copying full production datasets into test sandboxes when a representative subset will do. Encrypt especially sensitive application data with keys that are not stored beside the ciphertext.

  5. Collect a retirement receipt. For infrastructure you operate, record the sanitisation mechanism, configuration state, cleanup job result, failed hosts, cached-layer retirement, and validation test. For managed services, retain the provider’s assurance, contract language, incident notices, and dates. “Resource deleted” is an event. “Old data cannot be read by the next tenant” is the outcome you need to evidence.

  6. Prepare the disclosure path. Decide who receives a provider security notice, who maps it to internal services, who classifies the data involved, and who can request contractual detail. A notice sitting in a billing owner’s inbox for three days is a control failure even when the provider fixed the platform in three hours.

This review should end with a short list of owners and changes. If it produces forty abstract requirements and no test, it has become paperwork. The highest-value result may be one integration test for storage reuse, one change to stop writing a long-lived key to disk, and one inventory field that identifies which managed services process restricted data.

Platform builders should add one more test at the implementation layer. Configuration defaults and performance flags need security assertions. If a shared thin pool requires zeroing for tenant safety, a startup check should refuse the unsafe mode, and a reuse test should fail the release when old markers appear. A comment in an operations guide cannot compete with a configuration path that still accepts the dangerous state.

The same receipt should cover cleanup after a fix. Cloudflare correctly recognised that changing skip_block_zeroing could not sanitise blocks already mapped into old thin devices and cached layers. A repair test that creates only new disks might pass while old snapshots preserve the earlier condition. Remediation is complete when both new allocations and inherited state satisfy the boundary.

How to judge a provider’s response

Security disclosures can leave customers choosing between two bad instincts. One is to dismiss every fixed issue because the vendor says there was no observed abuse. The other is to treat every bug as proof that the whole service is untrustworthy. Neither helps an engineering team decide what to run next week.

Judge the response by evidence and control. Did the provider explain the mechanism clearly enough for independent scrutiny? Did it contain the issue quickly? Did it distinguish the immediate fix from cleanup of inherited state? Did it test the researcher’s method after remediation? Did it search available telemetry for similar activity and state the limits of that search? Did it tell customers whether any action was required?

Cloudflare’s report answers those questions better than many incident notices. It names the affected pool option, describes the 4 KiB write and 64 KiB allocation, publishes the research counts, distinguishes active disks from released blocks, gives a minute-level timeline, and explains why cached snapshots required more work. It also avoids claiming that telemetry can prove a universal negative.

There are still questions a high-sensitivity customer may reasonably ask through its account or assurance channel. How long did the affected configuration exist? Which product regions or service tiers used the relevant pools? What retention window supported the telemetry review? What independent validation now guards against recurrence? The public report may not answer every contractual question, and asking them does not require accusing the provider of hiding a confirmed breach.

A useful provider review separates architecture from response maturity. A mature response does not make the original flaw harmless. It does make the current state easier to assess and lowers the chance that customers repeat the same design mistake. Conversely, a product can have impressive isolation diagrams and still leave customers blind if its incident notice says only “an issue was resolved.”

Procurement teams often ask whether a service is “containerised” or “encrypted.” Those questions are too broad. Ask what protects one tenant when storage is reassigned, what keys exist at each layer, how cached images are retired, what a sandbox can read below its filesystem, and how the provider detects a cross-tenant access pattern. The answers reveal where the boundary actually lives.

For systems that handle highly sensitive data, design as though a provider layer can fail. Minimise plaintext, separate keys, keep credentials short-lived, and avoid giving a disposable workload more data than its task requires. This is defence in depth with a clear purpose. It reduces consequence without pretending customers can compensate for every host-level isolation flaw themselves.

The boundary includes the empty space

The most striking part of this incident was not that a container could read another running container. It could not. The failure lived in space the platform considered free.

That empty space still had history. A deleted thin volume returned its blocks to a shared pool, a small write assigned one of those blocks to a new tenant, and the unwritten remainder carried old bytes across the boundary. The running virtual machine did exactly what it had been allowed to do. The unsafe decision had already happened beneath it.

Cloudflare fixed the allocator, retired disks and cached snapshots created under the old condition, tested the researcher’s method, and found no malicious match in the telemetry it reviewed. Customers can take that result seriously without wasting the lesson. Resource retirement is part of resource isolation.

The durable control is plain: a shared platform must prove that reused storage is clean before a different tenant can read it. Pair that provider control with smaller data footprints, separate application keys, short-lived credentials, and a receipt for the cleanup path. The next boundary failure may sit in memory, a cache, a snapshot, or a log rather than a disk block. The review method still holds.

A sandbox includes the wall around a running process, the door that prepares the room, and the crew that clears it for whoever comes next.

For one calm, practical security email each month, join the newsletter on this site.

Sources