← Back to list

flexFS™ High-Performance File System Accelerates Cryo-EM Analysis

Benchmarks: how quickly cryogenic electron microscopy (cryo-EM) EMPIAR 10288 images can be analyzed using EFS, FSx for Lustre, & flexFS

David Freund · 2025-03-21 15:25 · 0 claps · 4.3 min read
#electron-microscopy #aws-fsx #aws-efs #flexfs #cryoem
Open on Medium ↗
Wiki topics: EVAL · Evaluation & Benchmarks PRO · Proteomics & Structure ☁️ · DevOps & Cloud

How to Accelerate Cryo-EM Analysis & Spend Less

Cryo-EM 3D model of EMPIAR 10288 Cannabinoid Receptor 1-G Protein Complex

Cryo-EM 3D model of EMPIAR 10288 Cannabinoid Receptor 1-G Protein Complex

Cryogenic electron microscopy (cryo-EM) enables scientists to visualize the 3D structure of biological molecules at near-atomic granularity, dramatically increasing our understanding of structures and mechanisms that govern life and revolutionizing structure-based drug discovery and design. Current cryo-EM microscopes can operate 24x7, generating 5 to 10 terabytes of data per day, 365 days a year –- a data rate that grows as cryo-EM hardware continues to improve. Increasingly powerful high-performance computing (HPC) clusters are needed to process and analyze that ever-increasing torrent of data.

Researchers are turning to cloud services such as Amazon Web Services (AWS), Google Cloud Platform, Oracle Cloud Infrastructure (OCI) and Microsoft Azure to perform cryo-EM analysis because they offer virtually limitless compute, network and storage resources. Better yet, those services are available on-demand, enabling teams to quickly procure — and release — resources as needed.

But compute power and storage capacity and networking alone are not enough. Cryo-EM HPC clusters also need an I/O infrastructure capable of feeding data fast enough to keep all those compute servers busy. Any time a server spends idle waiting for data is a waste of time, compute resources, and money.

Most cloud file systems have been designed with business computing in mind, keeping I/O latency (the time between a read or write request and the response, typically a small number of milliseconds) to a minimum. HPC workloads processing huge datasets, on the other hand, frequently require high throughput — gigabytes per second.

Some workloads require both low latency and high throughput. Cryo-EM analysis is one such case.

Amazon’s Elastic File System (EFS) is a POSIX-compliant file system that can grow and shrink storage capacity automatically based on the amount used. EFS latency characteristics are pretty good, but aggregate throughput tapers off dramatically once HPC clusters grow large enough. (There have been improvements over time; but at this writing, it’s still nowhere near enough to handle hundreds or thousands of servers.)

Another AWS offering, FSx for Lustre, is a parallel file system designed for HPC. The good news is you can scale throughput as much as needed. The bad news you increase throughput by manually increasing capacity in increments of 2.4TB, which can very quickly become wasteful and expensive. Decreasing Lustre capacity is even more onerous. And disruptive. You must create a new smaller volume, copy data, and delete the old volume. Manual planning, monitoring, and intervention that costs both time and money — and negates a major reason for using cloud in the first place.

So, what’s the alternative?

flexFS, a high-performance POSIX-compliant elastic file system that leverages the low cost, scale, and aggregate throughput of hyperscale object storage such as AWS S3, Azure Blob, Google Cloud Storage, OCI Object Storage, etc. — and transforms it into file storage that outperforms conventional cloud file systems — for significantly less cost. flexFS performance is tunable over a range from fully throughput-optimized to fully latency-optimized — or both. Customers such as Alnylam, Amgen, Bristol Myers Squibb and NASA have been using flexFS in production for HPC workloads, and for general-purpose file sharing by hundreds of concurrent users, for years.

Okay, but how well does flexFS accelerate cryo-EM analysis? Can it save me money? How much?

To answer those questions the creator of flexFS, Paradigm4, ran benchmarks on AWS measuring how quickly cryo-EM movies from the EMPIAR 10288 image set can be analyzed using EFS, FSx for Lustre, and flexFS. EMPIAR 10288 contains 476 GB of raw data comprising 2,756 multi-frame TIFF images, 40 frames per image. (Each frame is 3838 x 3710 unsigned 16-bit integer pixels, spaced: 0.86 Å x 0.86 Å apart.)

The benchmark team executed the automated steps typical of single-particle analysis workflows — Motion Correction, CTF Estimation, Particle Picking, 2D Classification, and 3D Reconstruction — using the RELION¹ analysis software suite running on a cluster of 8 AWS instances (g5.12xlarge), each processing the full set of EMPIAR 10288 electron-microscope movies. In other words, the cluster collectively processed the entire dataset 8 times, in parallel. The test was performed this way because most Cryo-EM projects tend to be larger and require an HPC cluster of multiple servers. flexFS was configured for both latency and throughput performance.

Here’s a summary of the results:

  • FSx for Lustre end-to-end analysis was 21% slower — but costs² 3.5 times as much — as flexFS.
  • EFS end-to-end analysis was 29% slower — but costs² 1.75–2 times³ as much — as flexFS.
  • flexFS ran each step of the cryo-EM workflow faster than either EFS or FSx for Lustre.

The following two graphs compare the performance and costs of flexFS vs FSx for Lustre and EFS

Cryo-EM Benchmark Results — Elapsed Time For Each Step, and End-to-End (minutes)

Cryo-EM Benchmark Results — Elapsed Time For Each Step, and End-to-End (minutes)

Monthly Cost for Different Capacities: flexFS, Amazon EFS and Amazon FSx for Lustre

Monthly Cost for Different Capacities: flexFS, Amazon EFS and Amazon FSx for Lustre

Cost relief provided by flexFS increases as data volume grows from 100 to 800 terabytes of data — an amount easily reached with a small number of cryo-EM datasets.

flexFS can be used as a drop-in replacement for Amazon EFS, FSx for Lustre, and other file systems on AWS, Google Cloud, Oracle Cloud Infrastructure, Microsoft Azure, and other cloud service providers — and easily serve thousands of concurrently active compute servers with minimal infrastructure or operational overhead.

Intrigued? You can find more details, including benchmark design, hardware configurations, and more at https://paradigm4.com/resource/cryo-em-analysis-benchmarking-flexfs-vs-lustre-vs-efs/.

For more about flexFS, check out flexfs.io, documentation site docs.flexFS.io, or contact Paradigm4 at lifesciences@paradigm4.com.

¹ RELION and CryoSPARC are the two major software suites commonly used to analyze cryo-EM datasets. By and large, each suite reads and writes the same file data. So observed file-system performance differences in one software suite can be expected to reflect differences in the other.

² Example pricing in AWS us-east-1 region (early 2025). Actual prices vary based on configuration.

³ AWS charges each time data is transferred to and from EFS storage


메타데이터
post_id
7c28da499efa
slug
flexfs-high-performance-file-system-accelerates-cryo-em-analysis-7c28da499efa
url
https://medium.com/@dfreund_21488/flexfs-high-performance-file-system-accelerates-cryo-em-analysis-7c28da499efa
canonical_url
https://medium.com/@dfreund_21488/flexfs-high-performance-file-system-accelerates-cryo-em-analysis-7c28da499efa
author_url
https://medium.com/@dfreund_21488
status
ok
fetched_at
2026-08-11 07:17:30