

Although it’s been almost 70 years since Watson, Crick, Wilkins, and Franklin elucidated the model of DNA (deoxyribose nucleic acid), research in genomics has picked up only in last two decades or so. The two important hurdles for the technology were the high cost of sequencing, and the lack of computational infrastructure to derive usable insights from the huge amount of information available in genomes.
As a measure of the amount of data that genomics can produce, consider that each somatic human cell contains a 2-meter-long DNA containing 3 billion nitrogenous bases wired along the backbone of sugar-phosphate molecules. Although it could theoretically be represented in 750 MB of data per human, whole genome sequences can practically occupy up to 100 GB per human, to account for real-life error-proofing redundancies.
Over the last 2 decades, research in genomics has more than compensated for the initial growth delays. Experts observing the exponential growth of data generated in genomics predict that growth to continue, surpassing even other large data stores such as online streaming and astronomy. They predict that the genomics industry may require 2 to 40 exabytes of storage capacity by 2025.
Figure 2: Growth of DNA Sequencing (© 2015 Stephens et al, Source)

Here are some possible reasons for the explosive increase in genomics research:
Genomics research is driving a large market. According to an estimate by Precedence Research, the global genomics market size was valued at US$20.06 billion in 2020 and is expected to hit more than US$72.13 billion by 2030.
As shown in Figure 3, genomics research can help many use cases across many industries:

We all know the core role that genomics research is playing in dealing with the COVID-19 pandemic. In last 14 months or so, there have been thousands of publications around COVID-19 genomics alone, not to mention countless webinars and other online meetups.
Here are some IT use cases for genomics workloads
FlexPod® solution from Cisco and NetApp can serve as the IT backbone of genomics research. NetApp and Cisco validate infrastructural hardware components, technologies, and software. This validation makes deployment and management of the IT infrastructure significantly less risky.
Just as nucleic acids, sugar, and phosphate are the building blocks of DNA and RNA, FlexPod is composed of building blocks from Cisco and NetApp. ONTAP provides the storage layer, and Cisco UCS blades or rack servers and MDS and Nexus switches form the compute and networking layers of a FlexPod unit.
FlexPod is a total solution for genomic data management that provides one seamless platform for simplicity and speed for genomics workloads. Figure 4 illustrates FlexPod’s value for genomics workloads.
Figure 4: FlexPod for genomics Infographic
In a nutshell, genomics expects IT infrastructure to solve these challenges:
