In a nutshell
Genome projects read the entire base sequence of an organism's DNA. Once you know the genome, you can work towards its proteome, but only cleanly in simple organisms.
This subtopic is about what sequencing gives us, why a bacterium's genome tells you its proteome while a human's does not, and one key application: finding antigens for vaccines.
Assumed knowledge: DNA, genes and chromosomes, DNA and protein synthesis.
Core content
What a genome project does
A genome is all the DNA in a cell or organism, not just the genes, and not "the DNA of a species".
Genome sequencing projects determine the entire base sequence of an organism's DNA.
- The genomes of a wide range of organisms, including humans, have now been sequenced.
- Sequencing methods are continuously updated and have become automated, so reading a genome is now far faster and cheaper than it used to be.
From genome to proteome in simple organisms
The proteome is the full range of proteins that a cell is able to produce.
In simple organisms such as bacteria, determining the genome lets you determine the proteome fairly directly. This is because a simple organism's DNA has:
- little or no non-coding DNA, and
- no introns interrupting its genes, so a gene is one continuous coding sequence.
So the base sequence maps straight onto the amino acid sequences of the proteins the organism can make, using the genetic code. Read the genome and you can read off the proteome.
Why a complex organism's genome does not give its proteome
In complex organisms such as humans, knowledge of the genome cannot easily be translated into the proteome. The spec gives two reasons, and there is a third worth knowing:
- Non-coding DNA. Much of a eukaryote's DNA does not code for a polypeptide. This includes introns within genes and non-coding sequences between genes. From the raw sequence alone you cannot simply tell which stretches are translated.
- Regulatory genes. Genes are switched on or off in different cell types and at different times, so not every gene is expressed. The genome does not tell you which proteins a particular cell actually makes.
- Alternative splicing (building on splicing of pre-mRNA, 3.4.2): the exons from one gene can be joined in different combinations, so a single gene can code for more than one polypeptide. One gene does not map to one protein.
Still don't get it? ยท why the genome does not equal the proteome in humans
Think of the genome as a giant recipe book that lists every dish a kitchen could ever make. The proteome is the set of dishes actually being cooked in one kitchen today. Knowing the whole book does not tell you today's menu.
Now build it up. In a bacterium the book is short: every entry is a plain, complete recipe with no padding, and the kitchen more or less cooks all of them. So the book basically is the menu, and reading the book (genome) tells you the meals (proteome).
In a human the book is huge and padded. Big chunks are notes, crossings-out and blank pages that are not recipes at all (non-coding DNA, including introns). The kitchen only cooks certain dishes depending on which kitchen it is and the time of day (regulatory genes switching genes on and off). And a single recipe can be cooked in several different ways (alternative splicing of one gene into different polypeptides). So you cannot read today's menu straight off the book.
Exam version: in complex organisms the presence of non-coding DNA and of regulatory genes means the genome cannot easily be translated into the proteome, whereas in simple organisms determining the genome does allow the proteome to be determined.
Simple versus complex organisms
| Feature | Simple organism (e.g. bacterium) | Complex organism (e.g. human) |
|---|---|---|
| Non-coding DNA | little or none | large amounts (introns and between genes) |
| Introns in genes | absent | present |
| Gene expression | genes largely all expressed | genes switched on and off by cell type and time |
| One gene codes for | one polypeptide | can be several (alternative splicing) |
| Genome to proteome | can be determined directly | cannot easily be determined |
Application: identifying antigens for vaccines
Determining a proteome has real medical value. The clearest example the spec names is vaccine production:
- Sequence the genome of a pathogen (a simple organism, so its proteome can be determined).
- From the proteome, identify proteins on the pathogen's surface that can act as antigens.
- Use those antigens in a vaccine to trigger an immune response, and so immunity, without needing the live pathogen.
An antigen is a molecule, usually a protein, that is recognised as foreign and stimulates an immune response.
Worked examples
Model 4-mark answer, "Explain why the genome of a bacterium allows its proteome to be determined, but this is not easily done for a human."
The lesson here is that a four-point "explain" needs four distinct, linked points, in two halves (the bacterium, then the human):
- In the bacterium there is little or no non-coding DNA and no introns, so the base sequence codes directly for the amino acid sequences.
- Therefore the genetic code can be used to work out the proteins (the proteome) from the genome.
- In the human there is non-coding DNA (introns and regulatory sequences), so not all the DNA codes for protein.
- Genes are regulated (switched on and off), and one gene can be spliced in different ways to give several polypeptides, so you cannot tell which proteins are made from the genome alone.
Common exam mistakes
- Defining a genome as "all the genes", "all the genes in a chromosome", or the DNA "in a species". The genome is all the DNA in a single cell or organism. The "species" version is explicitly rejected, and around 40% of students lose this mark.
- Confusing genome (all the DNA) with proteome (all the proteins a cell can make), and using the words interchangeably.
- Calling non-coding DNA "useless" or saying it "does nothing". It includes introns and regulatory sequences, which is exactly why the genome does not read straight off as the proteome.
- Writing that the human proteome "cannot be determined". The point is that it cannot easily be determined from the genome alone, because of non-coding DNA and gene regulation.
- Giving only "used to make vaccines" for the application, without the chain: determine the proteome, identify surface proteins that act as antigens, then use them as the antigens in the vaccine.
- Saying introns are "removed from the genome". Splicing happens to the pre-mRNA, not to the DNA; the introns are still in the genome.
Key definitions
- Genome - the complete set of DNA (all the DNA) in a cell or organism.
- Proteome - the full range of proteins that a cell is able to produce.
- Antigen - a molecule, usually a protein, that is recognised as foreign and stimulates an immune response.
- Non-coding DNA - DNA that does not code for a polypeptide (it includes introns and regulatory sequences).
Specification
- I can state that sequencing projects have read the genomes of a wide range of organisms, including humans.
- I can explain why determining the genome of a simpler organism allows its proteome to be determined.
- I can describe how determining a proteome is used to identify potential antigens for use in vaccine production.
- I can explain why, in complex organisms, non-coding DNA and regulatory genes mean the genome cannot easily be translated into the proteome.
- I can state that sequencing methods are continuously updated and have become automated.
Related notes
Ready to test yourself?
Put Using genome projects into practice with exam-style questions and full mark schemes.
Practise Using genome projects