Genetics in the Post-Genomic Era
by Ariel Yael
Audio version created with Paper2Audio.
Listen on Paper2Audio
Genetics in the Post-Genomic Era
Lecture #1
Genetics v. Genomics:
• Genetics:
○ Study of heredity
○ Involves study of functions and compositions of the single gene
Genomics:
Omics:
○ Study of the entirety of an organism's genes
- Organism's entire genetic makeup
- o Addresses all genes and their inter-relationships
○ Uses a combination of recombinant D.N.A, D.N.A sequencing methods, and bioinformatics to sequence, assemble, and analyze the structure and function of genomes
- Multi-omics combines data from multiple molecular layers to understand biological systems
- High-throughout technologies allow simultaneous measurements of D.N.A, R.N.A, proteins, and metabolites
• Genomics - profiles D.N.A
• Transcriptomics - measures R.N.A transcripts
• Proteomics - quantifies proteins
• Metabolomics - quantifies metabolites
Genomic Measurements:
Evolution of genomic measurement tools from bulk omics to spatial omics
Image summary: A diagram illustrating the progression of omics technologies across three stages. The first stage, Bulk omics, shows a mixed collection of blue circles, green hexagons, red squares, and orange stars. The second stage, Single-cell omics, shows these same shapes sorted into distinct groups by type. The third stage, Spatial omics, shows the sorted shapes reorganized into a cluster where they are grouped by type in specific spatial locations.
Functional Genomics:
- Study of how genes and intergenic regions of the genome contribute to different biological processes
- Studies genes or regions of genes on a “genome-wide scale” (i.e. all of multiple genes/regions at the same time
- To determine how the individual components of a biological system work together to produce a particular phenotype
• Dynamic expression of gene products in a specific context
• Mass data sets
- Example arrow at specific developmental stage or during disease
Functional Genomics Approaches:
Image summary: A diagram illustrating the four fields of "omics" that contribute to a person's phenotype. At the top, Genomics & epigenomics is defined as the study of DNA sequence and associated heritable biochemical modifications, and Transcriptomics is the study of RNA molecules present in a sample. At the bottom, Proteomics is the study of proteins present in a sample, and Metabolomics is the study of metabolites present in a sample. Arrows from all four fields point toward a central group of diverse human figures labeled Phenotype.
The Human Genome Project (H.G.P):
- 1990 to 2003
- To establish inventory of genes of human species (3 billion bases of human D.N.A)
• To find genes involved in disease
• To explain individual phenotypic variation
- Address the ethical, legal, and social issues (E.L.S.I) that may arise from the project
• Cost: approximately 2.7 billion U.S.D
Only 2 to 3 percent of the human genome is actually coding.
Lecture #2
Human Pangenome:
Integrate genomes from multiple individuals
• Reduces bias
• Improves disease variant discovery
Single Cell Analysis:
- Single cell - gene expression profiling with multiomics options
- Spatial - whole transcriptome in tissue context
• In situ - single cell spatial imaging
No correlation between the number of genes an organism has and its complexity
Image summary: An illustration of a fruit fly with a tan thorax and striped abdomen, accompanied by the text Fruitfly ~13,000.
Non-coding R.N.A:
• Genome is actively transcribed to produce thousands of non-coding transcripts
Central Dogma of Molecular Biology, Revised:
Image summary: A diagram illustrating two different genetic pathways. On the left, a protein coding gene (DNA) is transcribed into mRNA with the help of transcription factors, which is then translated into a protein. On the right, a miR gene (DNA) is transcribed into pre-miRNA and then processed into miRNA. The miRNA then acts to inhibit the translation of mRNA into protein, indicated by a red T-bar symbol blocking the path to the protein.
The role of mi R.N.A (micro R.N.A) is to block/regulate the step from m R.N.A to protein.
Epigenetics:
- Genetic changes that are heritable and do not depend on D.N.A sequence changes
- Inherited changes in gene expression through the modification of D.N.A and chromatin structure but not to the D.N.A sequence
• Principle mechanisms:
- D.N.A methylation
○ Histone modification
- Chromatin modifiers
• Examples:
- D.N.A methylation at specific gene
○ Histone modification at a promoter
- Silencing of one gene
Human ~20,000
Mouse ~20,000
Rice
~50,000
Plant
~25,500
Mustard ~25,500
Epigenomics:
• Global, genome-wide study of all epigenetic modifications across entire genome
• Focus on whole genome; systems-level patterns
- Large-scale data (N.G.S-based)
• Examples:
○ Mapping D.N.A methylation across all chromosomes
○ Profiling histone marks in all genes of a cell
- o Comparing epigenetic patterns between tissues or conditions
Epigenetic v. Epigenomics
Table summary: Epigenomics expands the scope of epigenetics from a single gene or locus to the entire genome. While epigenetics focuses on the mechanism of a specific gene, such as its methylation, and asks what regulates that gene using targeted assays, epigenomics examines the system-wide global regulatory landscape using high-throughput sequencing to create methylation maps of all genes.
Era of personalized/precision genomic medicine to to enable clinicians to quickly, efficiently, and accurately predict the most appropriate course of action for a patient.
Genome-Wide Association Study (G.W.A.S)
• Identifies genetic variants across entire genome that are associated with a trait or disease
Collects large group of people
Scans genome of single nucleotide polymorphisms (S.N.P's)
○ Compares frequencies - is it more common in people with trait?
○ Finds associations, not causation
Most variants have very small effects, still struggle with rare variants and non-European populations
• Combines big data, multi-omics, and functional biology
Gene therapy to any treatment in which a replacement gene is added to a person's body or a disease causing one is inactivated
- Can only be done in somatic cells, not germ cells due to ethical reasons
Lecture #3
Genome Editing:
• Methods for targeting sequences at higher efficiencies
- Using engineered nucleases that recognize specific regions in the genome, cleave them at high efficiencies, and allow permanent changes to occur at these regions
- crisper/Cas-9 based R.N.A-guided D.N.A endonucleases
Proteomics to large-scale study of proteins
• Quantifies thousands of proteins in one sample
• Detects post-translational modifications (P.T.M's)
• Mass spectrometry (M.S)-based proteomics
Data-independent acquisition (D.I.A)
• Single-cell proteomics (emerging)
Al and Genomics:
• Identify drug targets
• Massive datasets
Monogenic disease discovery:
- A mutant (variant) in one gene is responsible for the disease
• Pathogenic variant
- Inherited according to Mendelian Law
- Ways to identify genes:
- Candidate cloning
- Positional cloning
- Direct sequencing (used today)
Common diseases:
- Majority of morbidity and mortality in developed countries
- Disease, cancer, cardiovascular and coronary artery disease, mental health, and neurodegenerative disorders
• Increase in elderly
- Do not usually show a simple pattern of inheritance
- Genetic factors often multiple, interacting with each other and environmental factors in complex manner
Image summary: A diagram illustrating the spectrum of disease causation from genetic to environmental factors. A central horizontal arrow is divided into a red 'GENETIC' side and a white 'ENVIRONMENTAL' side. Examples of purely genetic diseases include Hemophilia, Osteogenesis imperfecta, Duchenne muscular dystrophy, Phenylketonuria, and Galactosemia. Examples of environmental diseases include Tuberculosis and Scurvy. Conditions falling in the middle, influenced by both, include Talipes, Pyloric stenosis, Dislocation of hip, Peptic ulcer, and Diabetes. The diagram further distinguishes between 'Rare Genetics simple (Unifactorial)' with high recurrence risk and 'Common Genetics complex (Multifactorial)' with low recurrence risk. A second arrow below maps conditions from 'Monogenic' (Cystic fibrosis) to 'Polygenic' (Infectious diseases), with Diabetes mellitus, hypertension, cerebrovascular and coronary artery disease, schizophrenia, common cancers, and congenital abnormalities listed as being due to both genetic and environmental factors.
Multigenic inheritance:
- Multiple mutations, in multiple different genes, that are responsible for developing the disease
○ Includes various auto-immune diseases, cancers, and heart diseases
• Does not follow rules of Mendelian inheritance
• Strong genetic component
- Locus heterogeneity
- Multi-hit model for disease risk
• Approaches: Q.T.L's, S.N.P's, trio-based exome sequencing Symbols for Pedigree:
Image summary: A diagram of standard symbols used in genetic pedigree charts. Symbols for individuals include a square for male, a circle for female, a diamond for unspecified sex, a red square or circle for affected individuals, a square or circle with a central dot for obligate carriers, and a square or circle with a vertical line for nonpenetrant carriers. Other markers include a diagonal line for deceased individuals, a circle with a slash for stillbirth, and a triangle for miscarriage. Relationship symbols include a horizontal line for marriage or union, a line with a slash for divorce, and double horizontal lines for consanguinity. Twin symbols are shown as inverted V-shapes: a horizontal bar connecting the V for monozygotic twins, a plain V for dizygotic twins, and a V with a question mark for unknown zygosity. An example pedigree shows two generations, with the proband indicated by a red circle and a red arrow.
Autosomal dominant inheritance:
- Vertical pedigree pattern, with multiple generations affected
- Each affected person normally has one affected parent
- Each child of an affected person has a 1 in 2 chance of being affected
- Males and females are equally affected and likely to pass the condition on
Autosomal recessive inheritance:
- Horizontal pedigree pattern, with 1 or more siblings affected. Often only single case
- Parents and children of affected people normally unaffected
- Each subsequent sibling of an affected child has a 1 in 4 chance of being affected
• Males and females equally affected
- Affected children can be product of consanguineous marriages
Image summary: Two diagrams illustrating genetic inheritance patterns between parents and offspring. The top diagram shows a red father and black mother; the father has one blue and one red chromosome, while the mother has two green chromosomes, resulting in offspring with varying combinations of these colors. The bottom diagram shows a black father and black mother; the father has one blue and one red chromosome, while the mother has two green chromosomes, resulting in offspring with different combinations of these colored chromosomes.
X-linked recessive inheritance:
- Never transmitted from father to son
- Affects mainly males (inherited from mother)
• Females can be carriers
• Affected males in pedigree linked through females
Identifying human disease genes:
1. Ascertainment of family with disease (Institutional review board or Helsinki)
2. Determining chromosomal location of gene (not relevant anymore because of N.G.S)
Image summary: A biological diagram illustrating the inheritance of sex chromosomes from two parents to their offspring. A male parent with X and Y chromosomes and a female parent with two X chromosomes are shown. The diagram uses arrows to trace the possible combinations of these chromosomes in offspring, resulting in four outcomes: two females with XX chromosomes and two males with YX chromosomes. One of the male offspring is highlighted in red, and a red mark is present on one of the female parent's X chromosomes and the corresponding X chromosome inherited by the red-colored male offspring.
3. Defining critical region and genes in that region (not relevant anymore because of N.G.S)
4. Evaluating candidate genes for mutation
5. Proving mutation causes disease
Step 1: Ascertainment of family with disease (Institutional review board or Helsinki)
The mandate of the N.I.H N.R.B is to provide ethical and regulatory oversight of research that involves human subjects
• Protects rights, welfare, and well-being of human research participants
• Ensuring compliance with relevant local, state, and federal laws and regulations
Employing the highest ethical standards for human research protections in all human subjects by adhering to the ethical principles outlined in the Belmont report
Step 2: Determining chromosomal location of gene to “Mapping” a genetic disease
Inherited in Mendelian manner
Sufficiently polymorphic so good chance that person that is heterozygous for disease also heterozygous for marker
• Easy to type on material from family
Available across whole genome at close intervals
• In the old days:
○ blood groups, protein electrophoretic variants, tissue types, R.F.L.P's
D.N.A polymorphisms
■ Microsatellites
Single nucleotide polymorphisms
***None of this needs to be done today!!
Microsatellites:
• Short tandem repeats
• Present in everyone at the same chromosomal location
• Number of repeat units varies between individuals
- Can determine how many allele in a population
- seph panel - panel of 100 to determine heterozygosity of allele
- Surrounded by unique sequence that can be used as P.C.R primers to amplify region
- Evaluate by gel electrophoresis
- Most frequent simple sequence in mammalian genomes is (C.A)n , occurs roughly 100,000 times in the mammalian genome
• Size varies between individuals
- 1 to 15 alleles of each marker in the population
Microsatellite marker
e.g. D.14.S.261
• Unique primer pair
• C.A repeat between primer pair
Image summary: A diagram of a DNA sequence fragment, 100-300 base pairs in length, featuring unique sequences at both ends and a repeating CACACACACACACACACACA sequence in the center that occurs approximately every 30 kb.
Single Nucleotide Polymorphism (S.N.P's):
- Germline substitutions of a single nucleotide at a specific position in the genome and are present in sufficiently large fraction of the population (1% or more). Closely spaced together
• Less informative than microsatellite markers (only 2 alternatives)
- Evaluate by electrophoresis or microarrays
L.O.D score leads to statistical test of linkage
- Estimate of whether two loci (the sites of genes) are likely to lie near each other on a chromosome and are therefore likely to be inherited together as a package
- The higher the L.O.D score, the stronger the evidence for linkage
Step 3: Defining critical region and genes in that region
- Narrow down chromosomal region with additional markers
Step 4: Evaluating candidate genes for mutation
• Prioritize genes to screen
Functional analysis
Pattern of expression (R.T-q.P.C.R, in situ hybridization, western blot protein analysis, immunohistochemistry)
o Family members or related genes - mutated in similar disease
o Animal models with similar phenotype
Sequence candidate genes
• Identify variant
• Mutation (pathogenic variant) or polymorphism?
Mouse and human genetic similarities
: A diagram of a human karyotype showing 22 pairs of autosomes and the X and Y sex chromosomes. Each chromosome is represented as a vertical bar divided into colored bands, with numbers labeling specific regions. A red oval highlights region 5 on chromosome 18.
Image summary: A karyotype diagram showing a set of human chromosomes numbered 1 through 22, plus X and Y sex chromosomes. A red circle highlights a specific region on chromosome 5.
Synteny implies physical co-localization of genetic loci on the same chromosome
Step 5: Proving variant causes disease (pathogenic)
1. Sequence D.N.A of unrelated, unaffected individuals
2. Check evolutionary conservation
3. Animal model - check if variant in same gene causes phenotype
4. Check gene expression (qRT-P.C.R)
5. Check protein expression (proteomics and western blot)
6. Evaluate 3D model of variant
7. Functional analysis (cell culture, mouse Tg/K.O)
{Lecture #4}
Sanger sequencing leads to 1977, complete genome of organism Sequencing - from low to high throughput
Deep Sequencing/N.G.S/M.P.S:
• Whole or exome sequencing
:Image summary: A line graph plotting the increase in DNA sequencing capacity from 1980 toward the future, measured in kilobases per day per machine. The capacity grows exponentially from 10 kilobases in 1980 to 1,000,000,000 kilobases in the future. Key technological milestones along the curve include manual slab gel and automated slab gel (gel-based systems), first-generation and second-generation capillary sequencers (capillary sequencing), microwell pyrosequencing, short-read sequencers (massively parallel sequencing), and the potential for single molecule sequencing.
• Gene-specific or region-specific capture
Sequence entire genome in 1 run
• Multiplex
• Less need for families
• Increase complexity to find mutations
Deep sequencing options:
Image summary: A diagram comparing three genomic sequencing methods. Whole Genome Sequencing shows continuous coverage across Gene 1, Gene 2, and Gene 3. Whole Exome Sequencing shows fragmented coverage specifically targeting the three genes. A Custom designed panel shows highly concentrated coverage on two specific Deafness Genes, with no coverage for the middle gene.
High-throughput sequencing:
• Next-generation sequencing
• Massively parallel sequencing
• Revolutionized human genetics
• From finding variants to interpreting variants
• Exome and whole genome sequencing
Long-read genomes
Image summary: A diagram comparing two methods of human genome sequencing. The left side describes generating a reference genome sequence, such as the Human Genome Project, which involves breaking the genome into large fragments, inserting them into clones, breaking those clones into small pieces, generating thousands of sequence reads to assemble the clone sequence, and finally assembling overlapping clones to establish a reference sequence. The right side describes generating a person's genome sequence, such as circa 2016, which involves breaking the genome into small pieces, generating millions of sequence reads, aligning those reads to an established reference sequence, and deducing the sequence to identify differences from that reference.
• Avoid cloning
• Able to stitch together due to human reference genome
• Re-sequenced to construct genome
Next generation sequencing steps:
D.N.A or R.N.A extraction
• Library preparation
• Sequencing
Data analysis & interpretation Explanation video of next generation sequencing steps
Rare Diseases:
• Rare diseases, although individually rare, are collectively common
Defined as one that affects fewer than 200,00 people in the U.S or less than 1 in 2,000 people in Europe
Most due to alterations in a single gene (R.G.D)
greater than 7000 R.G.D.S
{Sharing genetic data} to the goal is to help improve patient care
ClinGen to consortium of people sharing data, developing standards and curating knowledge; a platform for clinical genomic data sharing worldwide part of the N.I.H in the U.S. Consists of two databases:
• ClinVar to database for sharing variant level knowledge
- ClinGenome to database for sharing gene level knowledge
Image summary: A flow diagram illustrating the process of improving patient care through genomic medicine via ClinGen. Data from patients, clinicians, laboratories, and researchers is shared to answer three critical questions: whether a gene is associated with a disease (Clinical Validity), whether a variant is causative (Pathogenicity), and whether the information is actionable (Clinical Utility). This process builds a genomic knowledge base through ClinVar and ClinGen websites, ultimately leading to improved patient care.
Actionable arrow able to do something downstream
Source: Dr. Heidi Rehm
Variant classification:
: Table summary: The criteria for gene-disease association classifications based on a total point system of 0 to 18. Points are derived from Genetic Evidence, worth up to 12 points, and Experimental Evidence, worth up to 6 points. Calculated classifications are determined by the total score: LIMITED for 1 to 6 points, MODERATE for 7 to 11 points, and STRONG for 12 to 18 points. A DEFINITIVE classification requires a score of 12 to 18 points plus replication over time, defined as more than 2 publications with convincing evidence over more than 3 years. Other possible classifications include Disputed, Refuted, or No Evidence Reported, depending on contradictory evidence.
Nucleus Embryo to first genetic optimization software that helps parents pursuing I.V.F see and understand the complete genetic profile of each of their embryos
Today leads to 10 million base pairs left in the human genome
Telomere-to-Telomere (T.2.T) arrow first complete, gapless sequence of the human genome
Nullomers: short D.N.A sequences (11 to 18 base pairs) that are entirely absent from the human genome but can emerge through mutations
- CpG dinucleotides: the majority of nullomer-emerging mutations are located in regions with CpG sites, known as hotspots for mutation due to methylation
- Enrichment in functional genomic elements: these mutations frequently occur at transcription start sites, splice sites, and transcription factor binding sites, potentially disrupting gene regulation
- Association with early-replicating regions to nullomer-emerging mutations are enriched in early-replicating, G.C-rich, and coding regions, highlighting areas under strong selective constraints
- Nucleosomes: the mutations tend to cluster at nucleosome cores, suggesting chromatin structure influences nullomer emergence
- Alu repeats: recently evolved Alu repeats show high densities of these mutations, especially insertions, indicating evolutionary and regulatory significance
- Pathogenic potential: disease-causing mutations are significantly more likely to lead to the emergence of nullomers compared to benian mutations
- Rare variants and pathogenicity: nullomer-emerging mutations are frequently found among rare germline variants with high predicted pathogenicity (higher cadd scores)
: Image summary: A diagram titled Nullomer-Emerging Mutations in Human Genome illustrating how single base pair mutations can create a nullomer. It shows an original DNA sequence (CGACGTAACGT) undergoing 1bp insertions, deletions, or substitutions. One example shows a substitution of an 'A' for a 'C' resulting in the sequence CGACGTACCGT, labeled as a Nullomer. Another example shows a substitution of a 'T' for an 'A' resulting in the sequence CGACGTAACGA, labeled as Not Nullomer.
Image summary: A diagram titled Characterized in Functional Genomic Elements depicting a horizontal line representing a DNA sequence. A right-angled arrow indicates a transcription start site, marked above by a pink asterisk, which precedes a tan rectangular block representing a genomic element.
Image summary: A diagram titled Enriched in Pathogenic Mutations illustrating a genetic concept. It shows a DNA strand with a promoter arrow and a gene box. Above the strand are two star-shaped symbols, one green and one pink. Below the strand are two human silhouettes: a green figure with a checkmark icon and a pink figure with an X icon, representing the presence or absence of a specific trait or condition associated with the mutations.
Comparative Genomics: genome sequences of different species (human, mouse, and a wide variety of other organisms from yeast to chimpanzees) are compared
Zoonomia project to apply advances in D.N.A sequencing technologies to understand how genomes generate the tremendous wealth of animal diversity
• Sequencing of more than 240 mammal genomes
-
Number of Genes:
Image summary: A collection of six images comparing the approximate number of protein-coding genes across different organisms. The organisms and their associated numbers are: Fruitfly (~13,000), Mouse (~21,000), Plant (~25,500), Human (~20,000), Rice (~50,000), and Mustard (~25,500).
Number of genes is not associated with complexity of organism Human Genome Project: key findings approximately 20,000 genes in human beings
• About same as in mice (~21,000), twice as many as in roundworms (C.elegans)
• Human genes are more complex
Animal model to a living, non-human animal used during research and investigation of human disease, for the purpose of better understanding the disease without the added risk of causing harm to an actual human being during the process
Fruitfly ~13,000
Human ~20,000
Mouse ~21,000
Rice
~50,000
Plant ~25,500
Mustard ~25,500
Mus musculus to common lab mouse
- Ideal model organism for human disease
• Physiologically similar to humans
- Large reservoir of potential models of human disease
○ greater than 1,000 spontaneous mutations
- Radiation-induced
- Chemical-induced (e.g., E.N.U)
- ☐ Transgenics, K.O
• Development of high-resolution genetic and physical maps
Mouse genes helped isolate human genes:
Possible due to:
Detailed mouse genetic map
Map animal models for human disease on mouse genetic map Considerable conservation of synteny (homology)
Make predictions for subchromosomal location of gene in other species
○ Considerable conservation of coding D.N.A
Once human or mouse gene isolated, c.D.N.A probe can be used to screen D.N.A library to isolate orthologous gene
Replaced by high-throughput sequencing
In-bred strains to produced by at least 20 consecutive generations of brother-sister mating, resulting in animals that are essentially clones of each other at the genetic level
Lecture #5
Genetic Mapping in the Mouse:
- Monogenic traits: a character/trait determined by a single gene-use 2 inbred strains
• Multigenic traits: a hereditary characteristic that is specified by several genes
E.g., weight, high blood pressure, anxiety, and depression
Over 90% of the mouse and human genomes can be partitioned into corresponding regions of conserved synteny. At the nucleotide level, 40% of the human genome can be aligned to the mouse genome.
Graphical representation of conserved synteny relationships:
Image summary: A comparative map illustrating the synteny between the human genome and the mouse genome. The chart displays human chromosomes 1 through 22 and the X chromosome as vertical bars, which are color-coded in segments to match corresponding mouse chromosomes as defined by the provided legend. The segments indicate the homologous regions, showing how genetic sequences are rearranged across the different mouse chromosomes compared to their human counterparts.
In the entire genome of the mouse to 20,210 protein-coding genes, over a thousand more than predicted in the human genome (19,042 genes)
Mouse lineage-specific regions contain 3,767 genes drawn mainly from rapidly-changing gene families associated with reproductive function
Mouse mutants:
• Spontaneous mutations
• Radiation-induced mutants
• Chemical-induced mutants
• Transgenic mice
• Gene-targeted K.O or K.I
crisper/Casnine
Transgenic leads to mutation added into germ cells
E.N.U leads to chemical carcinogen
: Image summary: A scientific diagram illustrating four different genetic research methodologies using mouse models to study human disease. Natural Variation involves using inbred strains and collaborative crosses to identify disease traits. Reverse Genetics shows the process of manipulating DNA in a lab to delete a specific region, using embryonic stem cells and blastocysts to establish a new line with targeted deletions. Forward Genetics focuses on ENU mutagenesis, where males are injected with ENU to create mutations, followed by screening offspring for phenotypes that model human disease. Transgenics describes the process of manipulating DNA in a lab, injecting it into harvested fertilized eggs, and implanting those eggs into foster females to create founders with the transgene.
TECHNICAL APPROACHES FOR MOUSE MODELS OF HUMAN DISEASE
Monica J. Justice, Linda D. Siracusa and A. Francis Stewart
Mouse Models:
Phenotype-driven
- o Variation among inbred mouse strains
- Spontaneous mouse mutants
- ☐ Induced mouse mutants
Genotype-driven
- ☐ Transgenic mouse models
- Knock-out mouse models
- Knock-in mouse models
Transgenics:
- Mice with extra copies of a selected gene (normal or mutant, mouse, human or chimeric) inserted into the genome; or a new gene (like a human disease gene - for example, A.P.P driven under the Thyone promoter) inserted into the mouse genome
- Common for studying dominant diseases
- Suffer from position effect
• Are not driven by the endogenous promoter
- Are inserted to the genome as concatenomers
Knockouts:
Image summary: A diagram illustrating the process of creating a transgenic mouse. The process begins with the mating of a male and female mouse to produce a fertilized egg containing a male and female pronucleus. The egg is held by a holding pipet while a transgene is introduced via an injecting pipet. This modified egg is then implanted into a female mouse, which subsequently produces a litter of offspring, one of which is marked as transgenic.
- The c.D.N.A sequence is changed by homologous recombination, so that the protein is not synthesized (knock out mouse), a different gene is expressed instead of the original gene (knock in mouse), or a specific mutation (e.g., a mutation found in a human family) is introduced into the mouse genome
- Common for studying recessive diseases
Image summary: A diagram illustrating alternative splicing of a gene. The top line represents a gene with three segments: two blue outer segments and one pink middle segment. Dashed lines show two different splicing pathways: one that includes all three segments and another that skips the two blue outer segments, resulting in a single maroon segment on the bottom line.
Image summary: A diagram illustrating the process of gene targeting in a mouse embryonic stem cell. A circular DNA vector containing a targeted gene is inserted into a mouse embryonic stem cell. The vector then integrates into the mouse genome, replacing a portion of a mouse gene through homologous recombination, and the remaining vector DNA is excised.
Vector - Used to be the “gold standard” method for making a mouse model. The gene is driven by the endogenous promoter. Very reliable
Cre-Lox - Conditional Mutation
- This gene cassette encoding Cre recombinase is engineered in a separate mouse strain
- When Cre recombinase (red circle) is introduced, as a transgene by crossing into a mouse line carrying the targeted gene locus, the D.N.A between the loxP sites (red triangles) is removed, thereby inactivating the gene
- Avoid lethality by knocking out the gene, since it is site specific
Enu Chemical Mutagenesis:
Image summary: A scientific diagram illustrating the Cre-Lox recombination system in mice to achieve tissue-specific gene inactivation. It shows the cross between a mouse carrying a Cre recombinase gene controlled by tissue-specific promoter X and a mouse carrying conditional (floxed) alleles of gene Y. The molecular mechanism depicts the Cre protein recognizing and cutting the DNA at LoxP sites flanking gene Y. The resulting offspring mouse has gene Y inactivated specifically in tissue X.
☀️
N-ethyl-N-nitrosourea (E.N.U) leads to highly potent mutagen
• 1 new mutation per gene in 700 gametes
• Primarily point mutations
Genomically humanized mice:
• Targeted genomic humanization of mice
• Engrafted with human cells, tissues or genes, enabling them to replicate human immune responses, disease processes, or metabolism
• Creating “knock-in” mouse models
- Mouse sequenced is replaced by human orthologous D.N.A
Precision Medicine:
• New era of medical practice
Detailed genetic and other molecular information about a patient's disease routinely used
• For effective, patient-specific treatment
Integration of A.I
• 4 P's:
Personalized
○ Preventative
Predictive
○ Participatory
Pharmacogenomics: study of genetic variation and drug response; ability to metabolize a drug and the correlation with how well that drug will work, and the whole basis is the individual's genetic profile
Image summary: A diagram illustrating a patient group with the same diagnosis and prescription, divided into four sub-groups based on drug response. The groups are categorized as: drug toxic but beneficial, drug toxic but NOT beneficial, drug NOT toxic and NOT beneficial, and drug NOT toxic and beneficial.
Adverse drug reactions (A.D.R's):
• Estimated direct U.S costs leads to 136 billion per year
~6% of all new hospital admissions
4th leading cause of death
Medication dose is typically adjusted by weight, renal function, etcetera Medication is not adjusted by genetics. All of us are different and genetic variation influences our responses to drugs.
Drug-metabolizing enzymes:
- The molecular genetic basis for these inherited traits began to be elucidated in the late 1980s, with the initial cloning and characterization of a polymorphic human gene encoding the drug-metabolizing enzyme debrisoquin hydroxylase (C.Y.P.2.D.6)
Image summary: A pie chart showing the distribution of CYP enzymes, with CYP3A4/5 as the largest portion at 36%, followed by CYP2D6 at 19%, CYP2C8/9 at 16%, CYP1A2 at 11%, CYP2C19 at 8%, CYP2E1 at 4%, and CYP2B6 and CYP2A6 each at 3%.
- C.Y.P.2.D.6 is estimated to be the major metabolic pathway for approximately 20% of all prescribed drugs; 38% of A.D.R's.
- Genes are considered functionally “polymorphic” when allelic variants exist in the population.
Lecture #6
Gene therapy using crisper:
- crisper (clustered regularly interspaced short palindromic repeats) is a naturally-occurring, ancient defense mechanism found in a wide range of bacteria.
- 1980s - scientists observed a strange pattern in some bacterial genomes.
- One D.N.A sequence would be repeated over and over again, with unique sequences in between the repeats.
- crisper is one part of bacteria's immune system, which keeps bits of dangerous viruses around so it can recognize and defend against those viruses next time they attack.
- The second part of the defense mechanism is a set of enzymes called Cas (crisper-associated proteins), which can precisely snip D.N.A and slice out of invading viruses.
• Genes that encode for Cas always near crisper sequences
• There are a number of Cas enzymes - best known is Casnine
• Comes from Streptococcus pyogenes, bacteria that causes strep throat
crisper consists of two components:
• “Guide” R.N.A (g.R.N.A)
- Non-specific crisper-associated endonuclease (Casnine)
g.R.N.A is short synthetic R.N.A composed of a “scaffold” sequence necessary for Casnine-binding and user-defined approximately 20 nucleotide “spacer” or “targeting” sequence which defines a genomic target to be modified.
: A diagram of the CRISPR-Cas9 system showing a blue Cas9 protein complexed with an orange guide RNA (gRNA). The complex is bound to a double-stranded DNA sequence labeled Gene of interest, specifically targeting a region marked as PAM+Target, where the gRNA aligns with the target DNA sequence.
Can change genomic target of Casnine by changing targeting sequence present in the g.R.N.A Targeting efficiency and off-target mutations:
• Targeting efficiency, or the percentage of desired mutation achieved, is one of the most important parameters by which to assess a genome-editing tool
- Targeting efficiency of Casnine compares favorably with more established methods, such as tallen's or Z.F.N's
- In human cells, custom-designed Z.F.N's and tallen's only achieve efficiencies from 1% to 50%.
- Efficiency dependent on delivery method
• Off-target mutations
- Likely to appear in sites that have differences of only a few nucleotides compared to the original sequence, as long as they are adjacent to a P.A.M sequence.
- Mutations occur because Casnine can tolerate up to 5 base mismatches within the protospacer region or a single base difference in the P.A.M sequence
Donor D.N.A: a ss.D.N.A donor template molecule bearing point mutation and sg.R.N.A blocking mutations (6bp long silent mutation to prevent further ds.D.N.A break after ss.D.N.A integration)
- Donor D.N.A should have homology arm of 40 to 100 nucleotides on either arm flanking desired point mutation
For Casnine to function, required specific protospacer adjacent motif (P.A.M)
• Recognition of P.A.M by Casnine nuclease is thought to destabilize adjacent sequence, allowing interrogation of sequence by the cr.R.N.A, and resulting in R.N.A-D.N.A pairing when matching sequence is present
Non-coding genome and human disease:
Image summary: A diagram illustrating the CRISPR-Cas9 system. A Cas9 protein, represented by a light blue cloud, holds a guide RNA composed of a crRNA (blue) and a tracrRNA (yellow-green). The crRNA spacer sequence binds to a complementary protospacer sequence on a target DNA strand. Scissors icons indicate where the DNA is cleaved: once at the protospacer and once adjacent to the PAM (Protospacer Adjacent Motif) sequence, labeled NGG. An inset detail shows the secondary structure of the guide RNA, highlighting the spacer sequence and the hairpin loops of the tracrRNA.
- Epigenomics: study of the complete set of epigenetic modifications on the genetic material of a cell
- Epigenetic changes are genetic changes that are heritable that do not depend on the D.N.A sequence.
- Inherited changes in gene expression through the modification of D.N.A and chromatin structure but not the D.N.A sequence
Epigenetics: the study of heritable changes in genome function that occur without a change in D.N.A sequence
• Majority of the genome is not expressed
• Less than 3% of the expressed R.N.A in mammals is protein-coding R.N.A
Hierarchy of chromatin organization in an interphase nucleus:
Image summary: A biological diagram illustrating the levels of DNA packaging from a nucleus down to individual nucleosomes. It shows a progression from a nucleus containing higher-order chromatin, which unfolds into a series of nucleosomes consisting of DNA wrapped around histone cores. The diagram highlights specific components including histone tails, histone variants, and DNA modifications. It identifies five regulatory mechanisms: I. DNA modifications, II. Histone chaperones, III. Histone modifiers, IV. Histone readers, and V. Remodeling factors, the latter of which is shown utilizing ATP to alter the structure of the nucleosome.
Genomic Regulation:
: A scientific diagram illustrating the hierarchy of genetic regulation, from epigenetic modifications to the genetic regulatory network and non-coding RNA. The top section shows DNA wrapped around histones to form nucleosomes, highlighting DNA methylation, chromatin modifications, and DNase I hypersensitive sites. The middle section zooms into a DNA strand where transcription factors and transcription machinery bind to specific sites. The bottom section depicts a genomic region with long-range regulatory elements, promoter architecture, and a transcribed region producing protein-coding and non-coding transcripts. To the right, these elements are shown integrated into the larger structure of a chromosome, including long-range chromatin interactions.
- D.N.A methylation: includes methyl groups that attach to genes, altering their expression
- Histone modifications: transcription of gene is controlled by local chromatin structure, making chromatin more or less condensed and available for transcription
Principal mechanisms:
{Major types of epigenetics:}
• X-chromosome inactivation
• Imprinting
Transcription of genes controlled by local chromatin structure. D.N.A packaged by histones and other proteins into chromatin. There are two basic configurations of chromatin:
1. Heterochromatin: tightly packed and represses gene transcription
a. Highly condensed
b. At centromeres and telomeres
c. Contains repetitive sequences
d. Gene-poor
e. Replicated in late S phase
f. No meiotic recombination
2. Euchromatin: open and more variable
a. Less condensed
b. At chromosome arms
c. Contains unique sequences
d. Gene-rich
e. Replicated throughout S phase
f. Recombination during meiosis
The space between the genes:
- Methyl groups attach to genes, altering their expression
- Methyl groups attach to D.N.A, wrapped around histones that make up chromatin
- Non-coding R.N.A's
Image summary: A biological diagram illustrating various epigenetic and genetic regulatory elements on a DNA strand. The DNA is shown wrapped around histone proteins, with several key components labeled: histone modification, silencer/repressors, enhancers, insulators, promoters, and gene coding exons. The diagram shows transcription factors (TF) bound to enhancers and promoters, a co-activator protein interacting with these elements, and the presence of RNA Pol II at a promoter site. Other highlighted processes include DNA methylation and the production of non-coding RNA.
Topological associating domain (T.A.D) region: D.N.A sequences within the domain physically interact with each other much more frequently than with sequences outside of it
Lecture #7
Single Cell analysis:
- Tissues or tumors with:
- ☐ Different cellular elements
- Different transcriptomic programs
• Dissecting complexities from genomic, transcriptomic, and proteomic perspectives
- Determining molecular signatures of every cell and its destiny during the course of the disease
Image summary: This diagram illustrates the process of single-cell analysis for tumor characterization. A bulk tumor from a patient, which may lead to distant metastasis, is represented as a heterogeneous collection of cells including malignant cancer cell clones, stromal cells like CAFs and endothelial cells, and immune cells such as macrophages, NK cells, T-cells, and B-cells. These cells are fed into a funnel-like system where single-cell analysis is performed to isolate and categorize individual cell types. The analysis integrates data from genomics, transcriptomics, and proteomics to profile the diverse cell populations.
Bulk v. Single Cell Analysis:
• Bulk:
- Measures average gene expression levels across a population of cells
○ Useful for quantifying gene expression in homogenous populations
○ Cheaper, easier and more straightforward to analyze
• Single cell:
Image summary: A diagram comparing two methods of tumor analysis. The top row shows bulk tumor analysis, which is described as undetailed and incomplete, resulting in no identification of different cell populations. The bottom row shows dissociated tumor analysis, described as detailed and comprehensive, which allows for the identification of different cell populations, including a specific tumor propagating population.
- Measures the distribution of gene expression levels across a population of cells
○ Useful for studying cell-specific transcriptomic differences in heterogenous populations
- More difficult to collect and analyze (fresh tissue is required), but allows us to study new questions
Human cell atlas to collaboration to build a collection of maps that will describe and define the cellular basis of health and disease
Human pangenome to contains nearly full genomic data from 47 people whose ancestry traces to different populations around the world. This represents 94 human genomes because each person carries two copies, one from each parent
http://fightdipg.org/research-projects/
Gene Therapy:
Deliberate introduction of genetic material into human somatic cells for therapeutic, prophylactic, or diagnostic purposes
into humans
○ Genetically modified biological vectors (viruses)
○ Genetically modified stem cells
○ Oncolytic viruses
○ Nucleic acids associated with delivery vehicles
Naked nucleic acids
Antisense techniques
○ Genetic vaccines
○ R.N.A interference
Xenotransplantation of animal cells
Therapies containing an active ingredient synthesized following vector-mediated introduction of genetic sequence into target cells in-or ex-vivo
Used to replace defective or missing genes (as in cystic fibrosis) or to introduce broadly acting genetic sequences for treatment of multifactorial diseases (e.g. cancer)
As of 2025, globally for clinical use:
- 36 gene therapies have been approved (including genetically modified cell therapies)
• 36 R.N.A therapies have been approved
- 71 non-genetically modified cell therapies have been approved
Problems and Challenges of Gene Therapy:
Although each disorder affects fewer than 100,000 individuals in the U.S, in aggregate they represent an enormous cost to the health care system
• I treatment of the greater than 70,000 patients with sickle cell disease (S.C.D) in the U.S exceeds $1 billion per year
Organ transplants (heart, liver, lungs) cost on the order of $0.5 to $1.2 million in the first year
• F.D.A-approved immunotherapy combinations may approach $1 million per patient
Antisense oligonucleotides
: Image summary: A stylized 3D illustration of a DNA double helix, featuring two twisting strands in shades of yellow and purple connected by colorful horizontal rungs representing base pairs.
• Gene replacement
• Gene silencing – R.N.A.i, A.S.O
crisper-Casnine Genome Editing
: Image summary: A diagram illustrating the CRISPR-Cas9 gene-editing mechanism. The Cas9 protein, depicted as a large blue shape, holds a double-stranded target DNA molecule that has been unwound to expose a single strand. A guide RNA molecule, shown in yellow and orange, binds to the exposed target DNA strand. Two scissors icons indicate the locations where the Cas9 enzyme cuts both strands of the target DNA, adjacent to a labeled PAM sequence.
Splicing alterations
• Gene editing – crisper/Casnine, Base editor
Considerations for therapy:
• Mechanism
Phenotype
Genotype
• Molecular target
Off-target rate
• Sustained time
• Delivery form
• Delivery strategy
Antisense Oligonucleotides (A.S.O's):
- Single-stranded deoxyribonucleotide, which is complementary to the m R.N.A target
Delivery:
- Viral vectors
- Lipid and polymer nanoparticles
Goal of antisense approach is downregulation of a molecular target, usually achieved by
Adeno-associated virus
Over 200 clinical trials using A.A.V
Current A.A.V gene therapy: Spinal Muscular Atrophy Leber Congenital Amaurosis induction of R.N.ase H endonuclease activity that cleaves the R.N.A-D.N.A heteroduplex with a significant reduction of the target gene translation
Image summary: A diagram comparing Germline Gene Therapy and Somatic Gene Therapy. In Germline Gene Therapy, gene therapy is applied to an embryo, resulting in a human where all cells, including germline cells, are targeted. In Somatic Gene Therapy, gene therapy is applied to a human, targeting only somatic cells and not germline cells.
Rare Pediatric Disease (R.P.D) Designation and Voucher Programs to reduces F.D.A review and approval period for rare diseases affecting children
The toolbox of the genetic engineer:
Viral vectors
○ Gamma-retrovirus
○ Lentivirus
○ Adenovirus
Herpes simplex virus
○ Adeno associated virus
• Gene editing
○ crisper/Casnine
○ tallen
○ Zinc finger
- Nonviral vectors
○ Lipid nanoparticles
Cationic polymers
○ Naked D.N.A
• Oligonucleotides
Antisense oligonucleotides
○ sh.R.N.A
Viral vectors:
Figure A summary: A diagram of a recombinant adeno-associated virus (AAV) genome, consisting of a central DNA sequence flanked by two inverted terminal repeats (ITR) at each end. The central region contains two labeled gene segments: rep, shown in yellow, and cap, shown in red.
Figure B summary: A diagram of a recombinant single-stranded adeno-associated virus (ssAAV) consisting of a central blue transgene sequence flanked by two inverted terminal repeats represented as black bracket-like structures at each end.
Al in Genomics:
Image summary: A diagram illustrating recombinant adeno-associated virus (AAV) particles, represented by four distinct icons. Two icons depict hexagonal AAV structures, one light-colored and one dark blue, while the other two icons show spherical viral structures with surface projections, also in light and dark blue variations.
Table summary: A comparison of viral vectors for gene therapy, including ADENOVIRUS, AAV, gamma-RETIROVIRUS, and LENTIVIRUS. ADENOVIRUS stands out for its high transduction efficiency and the largest packaging capacity, ranging from 8 to 36 kb, though it has high immunogenicity and is non-integrating, leading to transient expression. AAV is the smallest at approximately 25 nm, has the lowest packaging capacity at 4.7 kb, and is the only vector listed with a BSL-1 biosafety level and low immunogenicity. Both gamma-RETIROVIRUS and LENTIVIRUS use ssRNA genomes and are integrating vectors that provide stable expression, but LENTIVIRUS can transduce both dividing and non-dividing cells, whereas gamma-RETIROVIRUS is limited to dividing cells. Finally, ADENOVIRUS and AAV are used for in vivo strategies, while gamma-RETIROVIRUS and LENTIVIRUS are used for ex vivo strategies.
• A.I technologies used to understand and analyze genomic data more effectively
• Massive data volumes
• High complexity
Pattern recognition
• A.I techniques used in genomics:
Machine learning: Decision trees, support vector machines, random forests
Deep learning: Convolutional neural networks, recurrent neural networks, transformers
Natural Language Processing: used for mining biomedical texts and D.N.A-as-language approaches
○ Unsupervised learning: clustering cell types, dimensionality reduction
Generative models: predicting genomic sequences, simulating mutations, or designing novel proteins
The future of genomics:
Therapy v. enhancement:
- Therapy: medical interventions that restore human functioning to species typical norms.
E.g., kidney dialysis, lasik eye surgery, angioplasty
• Enhancement: technologies to magnify human biological function beyond species typical norms
○ Ex. adding twenty I.Q points to someone with normal I.Q
You have reached the end of the document.