Home › DNA for Computational Biologists: A Complete Curriculum
DNA for Computational Biologists: A Complete Curriculum
Subject: DNA for Computational Biologists: A Complete Curriculum
21 chapters
Chapters
How to Use This Curriculum chill, uk spoken poetry, jazz, bass and piano, slow · 4:16 A practical orientation to the curriculum's structure, pacing, and recommended workflow, helping learners understand how to navigate the material and connect biological concepts with computational applications from the start.
Module 0 — Computational Environment Setup (Week 0) piano, classical, rap · 3:06 Before diving into genomic analysis, get your toolkit in order: install Python, Conda, and essential bioinformatics libraries, configure a command-line environment, and verify everything works through hands-on checks—laying the technical groundwork every subsequent module builds upon.
Untitled— Module 1 — The DNA Molecule (Weeks 1–2) melodic piano, duet male female · 3:38 Explore the chemical architecture of DNA—from nucleotide structure and base pairing to the double helix's geometry—building the foundational knowledge needed to understand how sequence data is generated and represented computationally.
Module 2 — Genome Biology & Organization (Weeks 3–4) chill, uk spoken poetry, jazz, bass and piano, slow · 4:22 Explore how genomes are structured and packaged, from chromatin organization and nucleosome positioning to repetitive elements, gene density, and the challenges these features pose for computational analysis. Listeners will gain a foundational understanding of genome architecture that underpins sequence alignment, annotation, and variant interpretation tasks in bioinformatics.
Module 3 — DNA Replication, Repair, Recombination & Mutation (Weeks 5–6) melodic piano, duet male female · 4:17 Explore how cells faithfully copy, fix, and occasionally reshuffle their genetic material, from replication fork mechanics and proofreading enzymes to repair pathways and the mutational signatures they leave behind. Listeners will gain the molecular foundation needed to interpret variant calling, mutation signatures, and genome stability analyses in computational biology workflows.
Module 4 — Laboratory Methods & How DNA Data Is Generated (Weeks 7–8) dark, ambient, mysterious, atmospheric · 5:01 Explore the wet-lab foundations behind sequencing data, from sample prep and library construction to the mechanics of Sanger, next-generation, and long-read platforms. Listeners will learn how sequencing errors, coverage, and read lengths originate at the bench, equipping them to make smarter decisions when analyzing genomic data downstream.
Module 5 — Sequence Data Formats & Tooling (Week 9) chill, uk spoken poetry, jazz, bass and piano, slow · 4:17 Explore the essential file formats—FASTA, FASTQ, SAM/BAM, and VCF—that underpin genomic data processing, along with the command-line tools and best practices bioinformaticians use to parse, convert, and manipulate sequence data efficiently. Listeners will gain practical fluency in navigating real-world genomic datasets and building reliable analysis pipelines.
Module 6 — Sequence Algorithms I: Alignment (Weeks 10–11) spoken word poetic, piano jazz · 3:46 Explore how algorithms like Needleman-Wunsch and Smith-Waterman compare DNA sequences to reveal evolutionary relationships and functional similarities, using dynamic programming to score matches, mismatches, and gaps. Listeners will learn to distinguish global from local alignment strategies and understand the computational foundations behind tools like BLAST.
Module 7 — Sequence Algorithms II: Assembly, k-mers & Sketching (Weeks 12–13) acoustic guitar, lyrical music, hand clapping, no drums · 4:03 Explore how genomes are reconstructed from short reads through de Bruijn graphs and k-mer decomposition, then discover how sketching techniques like MinHash enable rapid similarity comparisons across massive sequence datasets.
Module 8 — Variant Discovery & Genotyping (Weeks 14–16) chill, uk spoken poetry, jazz, bass and piano, slow · 4:19 Explore the computational pipeline for identifying genetic variants from sequencing data, covering key algorithms for calling SNPs and indels, filtering false positives, and assigning accurate genotypes across samples. Listeners will gain practical understanding of tools like GATK and variant quality metrics essential for distinguishing true biological variation from sequencing artifacts.
Module 9 — Variant Interpretation & Annotation (Week 17) chill, uk spoken poetry, jazz, bass and piano, slow · 4:28 Explore how raw genomic variants are translated into biological meaning, from functional annotation and effect prediction to navigating databases like ClinVar and gnomAD. Listeners will learn to classify variants by pathogenicity, interpret VEP and SnpEff outputs, and apply ACMG guidelines to distinguish disease-causing mutations from benign polymorphisms.
Module 10 — Population & Evolutionary Genomics (Weeks 18–20) acoustic guitar, lyrical music, hand clapping, no drums · 4:17 Explore how genetic variation arises and spreads through populations, covering allele frequencies, selection, drift, and the statistical methods used to detect signatures of evolution in genomic data. Listeners will learn to apply tools like Fst, dN/dS, and coalescent theory to real datasets, uncovering the evolutionary forces that shape genomes over time.
Module 11 — Statistical Genetics & Association (Weeks 21–22) acoustic guitar, lyrical music, hand clapping, no drums · 3:59 Explore how genetic variants connect to traits and disease through the statistical machinery of GWAS, from population structure and linkage disequilibrium to Manhattan plots and multiple testing correction. Listeners will learn to interpret association signals, understand heritability estimates, and grasp the mathematical foundations that separate true genetic effects from statistical noise.
Module 12 — Epigenomics & Functional DNA Assays (Weeks 23–24) dark, ambient, mysterious, atmospheric · 4:48 Explore how chromatin accessibility, histone modifications, and DNA methylation shape gene regulation beyond the sequence itself, using ATAC-seq, ChIP-seq, and bisulfite sequencing data. Listeners will learn to process and interpret these functional assays computationally, connecting epigenetic signals to gene expression patterns and regulatory element discovery.
Module 13 — Metagenomics, Pangenomics & Special Topics (Weeks 25–26) chill, uk spoken poetry, jazz, bass and piano, slow · 3:47 Explore how computational biologists analyze microbial communities directly from environmental samples and compare genomes across entire species using pangenome frameworks. Listeners will learn key techniques in metagenomic assembly, taxonomic classification, and pangenome construction, alongside emerging special topics shaping the future of genomic analysis.
Module 14 — Reproducibility, Scale & Capstone (Weeks 27–30) dark, ambient, mysterious, atmospheric · 4:19 Explores the engineering practices that separate one-off scripts from trustworthy computational biology pipelines—containerization, workflow managers, version control, and cloud-scale execution—before guiding learners through a capstone project that integrates the curriculum's genomics, statistics, and programming skills into a complete, reproducible analysis.
Appendix A — Mathematical & Statistical Foundations (reference track) piano, classical, rap · 3:15 A quick-reference deep dive into the probability, statistics, and linear algebra underpinning computational biology, covering concepts like likelihood, Markov chains, hidden Markov models, and eigenvectors as they apply to sequence analysis and genomic data. Listeners will walk away equipped to understand the mathematical logic behind common bioinformatics algorithms, from alignment scoring to phylogenetic tree building.
Appendix B — Molecular Biology Fast Track (for CS/math backgrounds) dark, ambient, mysterious, atmospheric · 4:07 A rapid-fire primer that maps core molecular biology concepts—genes, transcription, translation, and mutation—onto familiar computational analogies, giving programmers and mathematicians the biological vocabulary needed to work confidently with genomic data. Listeners will finish with a working mental model of the cell's information flow, enough to interpret sequencing outputs and biological literature without getting lost in jargon.
Appendix C — Core Tool Checklist acoustic guitar, lyrical music, hand clapping, no drums · 3:14 A quick-reference rundown of the essential software, libraries, and file formats every computational biologist should have ready to go, from sequence aligners to visualization tools. Listeners get a practical checklist for building or auditing their own bioinformatics toolkit before diving into real analysis work.
Appendix D — Core Reading List lo-fi, ambient, dreamy, relaxed · 4:53 A curated collection of essential papers, textbooks, and online resources that anchor each stage of the computational biology journey, giving learners a roadmap for deepening their expertise beyond the core curriculum.
Self-Assessment: You're "Fluent" When You Can… melodic piano, duet male female · 3:13 A practical checklist for gauging true fluency in computational biology, spelling out the concrete skills and problem-solving abilities you should possess before calling yourself proficient. Listeners will learn how to self-diagnose gaps in their knowledge and identify which concepts demand further practice before advancing.