Skip to main content
Biology LibreTexts

10: Research Project Week 5-7- Molecular Sequence Identification

  • Page ID
    124423
  • \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \( \newcommand{\dsum}{\displaystyle\sum\limits} \)

    \( \newcommand{\dint}{\displaystyle\int\limits} \)

    \( \newcommand{\dlim}{\displaystyle\lim\limits} \)

    \( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)

    ( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\id}{\mathrm{id}}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\kernel}{\mathrm{null}\,}\)

    \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\)

    \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\)

    \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)

    \( \newcommand{\vectorA}[1]{\vec{#1}}      % arrow\)

    \( \newcommand{\vectorAt}[1]{\vec{\text{#1}}}      % arrow\)

    \( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vectorC}[1]{\textbf{#1}} \)

    \( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)

    \( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)

    \( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)

    \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \(\newcommand{\longvect}{\overrightarrow}\)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)

    Microbial Characterization

    At this point in your research project, you have isolated an antibiotic producing microbe from the soil, grown it in isolation (away from the community of microbes with which it shares the soil. It is also possible that you haven't yet found an antibiotic producer and instead have selected another cool bacteria that you wish to learn more about. Either way, we are about to change directions from looking at the growth and capacity of many different bacteria to focusing specifically on one species. So . . . where do we go from here?

    We are still a long way off from marketing your potential antibiotic. Before we can use your antibiotic on patients we, the scientific community, must first produce large quantities of the molecule, ensure that it is effective in infected people, the antibiotic must not be harmful to eukaryotic/human cells, and it must continue to be active if ingested in pill form. Obviously, we won't be marketing your drug at the end of the semester! The first thing is that we need to know that this antibiotic is a NEW one— one that hasn't previously been discovered and is currently used.

    To identify this antibiotic, we are going to begin by determining which bacterial species is producing it. The reason for this is that we know which bacteria are known to produce particular antibiotics. For example, Acremonium chrysogenum is known to produce cephalosporin c (a β-lactam). If an active soil isolate can be classified as Acremonium or a close relative, we might predict that the antibiotic is cephalosporin and we would begin our testing there.

    Using Systematics to Identify Our Bacteria

    Systematics is a sub-field of biology that works to classify organisms in an organized and logical fashion. The living organisms are classified in a hierarchal system where each species is given a name and grouped with closely related species into a larger group called the genus and multiple genera (plural of genus) are grouped in one family. This classification system reflect evolutionary relatedness meaning that two species in the same genus are be more closely related to each other than two species in different genera.

    Table 1. Systematics Hierarchy. Important terminology/formatting note: The Genus and species names are always in italics, and Genus is capitalized while species is lowercase.

    Systematics Hierarchy
    Domain Bacteria
    Kingdom Eubacteria
    Phylum Firmicutes
    Class Bacilli
    Order Bacillales
    Family Staphylococcaceae
    Genus Staphylococcus
    Species aureus

    Historically, scientists used to classify bacteria on visible properties (cell shape, colony morphology, number or flagella, nutrient requirements, etc. However, sometimes these features are misleading and not particularly helpful in sorting such a diverse array of bacteria. Additional data, especially metabolic and genetic information is incredibly valuable in characterizing and classifying bacterial species. Using the genetic information, scientists have determined the universal tree of life (Figure 1). In our research on your antibiotic producer, we will attempt to identify the location of your species on this tree of life.

    Three-way branched tree of life showing the split from a common ancestor (LUCA) to three branches. Blue in the upper left represents bacteria. Green on the right represents archaea, and red on the lower left is Eukarya.

    Figure 1. The universal tree of life.

    With genetic (DNA sequence) information, scientists can compare the sequences of multiple organisms. Because two species evolved from a common ancestor, scientists can compare the similarities and differences of their DNA sequences to see how divergent they are. As the two new lineages evolved away from each other, their DNA sequences mutated and became less similar. The greater the difference in DNA sequence between two species, the greater the evolutionary distance between them. The more similar the two DNA sequences are, the more closely related these two species are.

    Although sequencing technology has improved dramatically, it is still not efficient or cost‐effective to sequence the entire genome of a new strain of bacterium to identify the species. Instead, scientists focus on a few genes possessed by all species of bacteria. The most important of these is the 16S rDNA gene. The 16S rRNA is an absolutely essential component of the bacterial ribosome, so all bacterial species have this gene. Although many other genes could be used for this analysis, the ribosomal 16S subunit is the gold standard because it is 1) present in all cells, 2) is not too large (~1500 base pairs in the gene), 3) has regions of incredible variability as well as regions that don't change much. An additional helpful feature is that so many bacterial species have sequenced 16S rRNA genes, there are many species to compare it to.

    Instead of sequencing the entire genome, a scientist typically uses a technique called PCR (polymerase chain reaction) to generate many identical copies of the 16S rDNA for the unknown species, sends this copied DNA to a sequencing company to determine the sequence of the gene, and compares that sequence with the publicly available established databases of all known bacterial 16S rDNA sequences. A “hit”, or match, in the database indicates that the species is likely known. Even if there are no hits, comparing the sequence of the 16S rDNA of the potentially new species gives valuable information about the identities of its closest relatives.

    Copying the DNA

    Polymerase Chain Reaction

    The process of DNA sequencing requires a large amount of a specific region of DNA to accurately determine the sequence. Just isolating the DNA from the bacteria does not provide a high enough concentration of DNA, so alternative methods must be used. The most common method for amplifying (aka copying) the DNA is a procedure called polymerase chain reaction (PCR).

    This technique uses a purified DNA Polymerase enzyme to copy small specific pieces of DNA exponentially and make millions of copies. PCR has made the field of molecular biology possible.

    A PCR reaction requires a few ingredients: 1) the template (DNA being copied), 2) a thermostable DNA polymerase, 3) a pair of specific single stranded DNA primers (to identify the particular region to copy), 4) free nucleotides, and 5) a buffer. The reaction mimics the DNA replication present in cells (Figure 2).

    First, as in DNA replication, the template is dsDNA, although it can be from virtually any source and does not necessarily have to be the genome of a cell. This template must be denatured, or separated into two single strands, so that it can be copied by synthesizing the complementary strand to each single strand. While cellular DNA replication uses a helicase enzyme to break hydrogen bonds between the bases and keep the two strands separate, PCR does the same thing by heating the dsDNA template.

    process of a PCR reaction. denaturation at 95C, annealing with ssDNA primers at 55-60C and extension with Taq Polymerase at 68-72C

    Figure 2. A single cycle of PCR. Each cycle of PCR replicates a region of the dNA template in vitro. First, the DNA is heated to denature the two strands. Primers anneal to the complementary regions of the template and provide the free 3’ –OH from which a thermostable DNA polymerase such as Taq polymerase can extend the complementary strand. The result is two double stranded DNA copies of the region between the forward and reverse primers.

    Second, in both types of replication a DNA polymerase enzyme synthesizes the new DNA strands in the 5' --> 3' direction. However, since the DNA polymerase used in PCR is exposed to the high heat necessary to denature the dsDNA template, it must be very heat stable and molecular biologists typically use the DNA Polymerase enzyme from a bacterium that lives in the very hot geyser waters of Yellowstone National Park. Since this bacterium is named Thermus aquaticus the enzyme is called Taq Polymerase.

    Third, in the cell, DNA Polymerase uses an RNA primer made by the enzyme primase as its starting point for replication. In contrast, when we do PCR we scientists provide a pair of DNA primers that are very specific for the section of the DNA that we want to copy. The sequence of the primers is designed to only bind to the DNA right around the gene being copied. As a result, Taq polymerase only copies the region between the pair of primers. It is important to note that to copy a region of DNA via PCR, a scientist must know the sequence of at least a small section of the DNA surrounding that target region.

    The final two ingredients of PCR are the nucleotides that get added by the enzyme and a buffer to maintain the appropriate pH and concentration of salts and cofactors. Note that the other pieces required for DNA replication in the cell are not required. No primase (we provide the primer instead), no DNA Pol I (there is no RNA to remove), no topoisomerase (the DNA being copied is short enough to not get tangled), no helicase (the heating separates the two strands instead).

    One additional important difference between in vivo DNA replication and PCR is that in PCR we only want to copy a small region of the DNA many many many times rather than just once in DNA replication. This process, called amplification, is done by cycling the temperature pattern of the reaction repeatedly. There are three different temperatures used for PCR and each plays a critical role in the copying process. First, the temperature is raised to near boiling (typically about 94°C) to

    separate the two strands of the DNA template. Second, the temperature is lowered to approximately 55–60°C so that the primers can stick to the complementary sequences in the template. Third,

    the temperature is raised slightly to 68–72°C and the polymerase synthesizes the complementary strands. This cycle of three temperatures is repeated multiple (about 25–40) times. With each cycle, the quantity of target DNA is theoretically doubled, so after n cycles with a starting point of x copies of template, there will be x2n copies of the target sequence (Figure 3).

    PCR products, doubling each round of amplifacation

    Figure 3. Amplification by PCR. Each cycle of PCR doubles the number of copies of the DNA between the forward and reverse primers. After n cycles, there are x2n molecules, where x is the initial number of molecules present.

    16s rDNA PCR

    Although the 16S rDNA gene is present in all bacteria, the sequences are slightly different in each organism as the bacteria have evolved away from each other. Some parts of the 16S rDNA gene have evolved slower than others and most species are nearly identical sequences at these regions. Other regions are divergent, meaning that mutations have accumulated at a higher rate and therefore these regions are less similar among species. Because we want to use the sequence of the 16S rDNA as a unique identifier of a species, these divergent regions are the ones whose sequence will be most helpful to us. Therefore, we will design primers to match the sequences that are unchanged and use those primers to copy the section that is variable. That will make the primers "universal," so that they work on almost all bacteria. The pair of primers we will use for this experiment is:

    oNS018: 5'-ACGCGTCGACAGAGTTTGATCCTGGCT-3'

    oNS019: 5'-CGCGGATCCGCTACCTTGTTACGACTT-3'

    These primers amplify the majority of the 16S rDNA gene in most species of bacteria. The exact size of the PCR may vary slightly depending on the species, but should be approximately 1500 bp.

    Separating the DNA by Size

    Gel Electrophoresis

    After amplifying a region of DNA by PCR, we need verify the size of the amplified fragment and estimate the quantity produced by the PCR. The most common way to do this is by gel electrophoresis, or “running a gel.”

    Gel electrophoresis is used to separate molecules based on their size and their charge as they move through a gel matrix. DNA has a net negative charge due to the phosphate groups in its backbone, so when an electric current is applied to a gel, DNA travels through the gel toward the positive electrode (anode). Agarose is a linear polymer extracted from seaweed that is composed of galactose sugars (Figure 4) and is typically used to separate DNA molecules. When melted agarose cools (or gels) it forms a three-dimensional mesh, like Jell-O. DNA moves through agarose at a rate inversely proportional the number of base pairs. Longer molecules of DNA migrate more slowly since they cannot worm their way through the pores of the gel as well as shorter DNA molecules. The rate at which the DNA travels depends also on the conditions of the gel itself (% agarose in the gel, amount of current applied). If a gel is run for a longer period of time, the DNA in that gel will migrate a greater distance.

    structure of agarose

    Figure 4. Structure of agarose. Agarose is a polymer of sugar molecules that connects to form a 3-dimensional mesh gel. Created by user:ideru, CC BY-SA 3.0, via Wikimedia Commons

    To run a gel, a sample of DNA is first mixed with a special loading dye and then loaded in a well or small hole at one end of the gel. This dye has two purposes. First, the dye contains glycerol which makes it heavier than the water buffer for the gel, so the sample sinks to the bottom rather than floating away. Second, it makes it colored so that it is easier to see and manipulate. The color also remains visible while the gel is running so that we can tell when the gel has run for long enough. If a gel is run too long, the DNA molecules will continue to migrate off the end of the gel and into the buffer.

    DNA can be visualized in a gel by staining it with a fluorescent DNA‐binding molecule. These DNA intercalate, or insert themselves, between the bases of double stranded DNA (Figure 5). When exposed to UV light, the dye fluoresces, which allow us to see where the DNA is located. Always use caution when working with these dyes because anything that binds to DNA has the potential of binding to YOUR DNA and may cause mutations and cancer. We will use the dye Gel Red, which stains the DNA well, but is relatively safe (much safer than earlier dyes). Despite its relative safety, scientists must also treat GelRed with caution because it is a potential mutagen and carcinogen. ALWAYS WEAR GLOVES when working with DNA stains and keep the work area clean. We will dispose of anything that came in contact with GelRed as hazardous (non-biological waste).

    structure of Gel Red

    Figure 5. Chemical structures of the DNA stain GelRed. The flat planar structure allows it to intercalate between the bases of the DNA double helix. From: Cacycle, Public domain, via Wikimedia Commons

    Often a DNA sample loaded on a gel contains a mixture of different sized DNA molecules. All the DNA molecules of a given size, say 1 kb, migrate through the gel at the same rate. Therefore, when the current is turned off, they will be at the same location and form a band (stripe) of DNA in the gel. DNA molecules of a different size, say 2 kb, migrate through the gel at the same rate as each other but at a different rate from the 1 kb molecules. Remember that larger pieces of DNA migrate more slowly through the gel matrix than smaller pieces of DNA. Thus, the 1 kb band will be farther from the well than the 2 kb band.

    To estimate the size of a band in a gel, we use a molecular weight standard called a DNA ladder (Figure 6). A ladder contains multiple bands of DNA fragments with known sizes. A scientist can compare the position of the DNA band to the ladder “eyeball” the size of the bands in a gel by looking where they are with respect to the different bands of the ladder. This ladder helps scientists measure the size of their DNA molecule.

    photograph of an agarose gel under UV illumination.

    Figure 6. Agarose gel. The wells where the DNA samples are loaded are visible near the top of the gel. The lane on the far left is the DNA ladder, which is used as a standard to compare the sizes of the bands in the two sample lanes. Larger fragments move more slowly from the wells and are closer to the top of the gel. Smaller fragments move more quickly and so are closer to the bottom of the gel. Samples are loaded into the first lane (ladder) second lane (positive control), third lane (negative control) and fourth lane (experimental sample).

    DNA Sequencing

    A variety of methods can be used to determine the sequence of a fragment of DNA; the method chosen depends on a variety of factors including type and number of samples, desired length of the sequencing read, etc. You will prepare your samples and we will mail them to a company for Sanger DNA sequencing. Sanger sequencing works by the selective incorporation of dideoxynucleotides (ddNTPs) by DNA polymerase. These ddNTPs are chain-terminating and marked with fluorescent tags to distinguish one base from the next. This allows all four bases to be terminators in the same reaction. The reaction then can be run through a capillary tube and the band type and order can be identified to determine the sequence of the DNA sequence. This information is recorded by the company in a trace chromatogram and the sequence data can be "read" off the chromatogram by reading the peaks in order (Figure 7).

    Sequencing results chromatogram diagram. Red = T, blue = C, green = A and black = G

    Figure 7. A portion of a trace chromatogram from an unknown sequence. Each peak corresponds to a base in that position in the sequence of the 16S rDNA.

    Protocols

    Protocol 1: Extract Genomic DNA

    Each group will extract the genomic DNA of one of your antibiotic producer. In consultation with your instructor you may also be extracting the DNA of one of the positive or negative control samples and/or an additional bacterial strain. Confirm your plan with your instructor before you begin.

    Materials
    • Plate cultures of your antibiotic producer
    • Sterile dH2O
    • Sterile 1.5 mL microfuge tubes
    • Pipettes and tips

    Protocol

    • Label your 1.5 mL microfuge tubes containing 200 µL of sterile water.
    • Using a micropipette tip, carefully touch the colony on the streak plate. A small, visible dab of cell that fills the end of the pipette tip will provide enough DNA template for the reaction.
    • Dip pipette tip into a microfuge tube the dH2O and gently swirl for 5–10 seconds to dislodge cells.
    • Vortex the tube for 30 seconds to suspend the colony completely in dH2O.
    • Lyse these cells by heating them at 95°C for 10 minutes.
    • Remove the lysed cells and store on ice until needed for the PCR.

    Protocol 2: Polymerase Chain Reaction

    Material
    • Genomic DNA from Protocol 1
    • PCR bead tubes (2-3 depending on number of samples)
    • Primer tubes (2-3 depending on number of samples)
    • Pipettes and tips

    Protocol

    • Obtain your PCR bead tubes, which contain Taq polymerase (heat-resistant enzyme) and other necessary reagents.
    • Label your tiny PCR bead tubes. This is critical—you need to tell yours apart from everyone else's and know what is in the next week.
    • Carefully and accurately add 15 μL of PCR primer mix into the PCR bead tube. The bead will start to dissolve and slightly effervesce.
    • Carefully and accurately add 10 μL of your lysed resuspended cells to this PCR bead tube. This will bring the volume of the reaction up to 25 μL.
    • Cap the PCR tube, vortex VERY gently.
    • Spin the tubes in a microfuge for a few seconds to collect all liquid at the bottom of the tube.
    • Transfer the tubes to the thermal cycler (PCR machine).
    • Select the appropriate program to start cycling (~2 hours).
    • Once cycling is complete, remove tubes and freeze until next week when you will run your agarose gel and prepare your sample for DNA sequencing. Be sure that your tube(s) are clearly labeled so that anyone can identify their contents.
    Table 1. The conditions for PCR amplification.
    Cycles Temperature Time Purpose
    1x 94°C 10 minutes Breaking open cells/ denaturing DNA
    36x

    94°C

    58°C

    72°C

    30 seconds

    30 seconds

    1.5 minutes

    Denaturation

    Annealing

    Elongation

    1x 72°C 10 minutes Final elongation (finish up any incomplete DNA pieces)
    1x 4°C forever Keep the samples cold until needed

     

    PCR Set Up Video

     

    Discarding Materials

    When you are finished running the PCR of your samples, please put these PCRs into the freezer rack provided by your instructor.

    While your PCR is running, discard the old plates from your experiment. You should put any old plates that you no longer need in the biohazard plates tub these plates would include your serial dilution plates, the initial master plate, and antibiotic test plates. Be sure that you have good pictures in your lab notebook of any plate you discard. In case your PCR does not work, be sure to keep your streak plate in the plate bag stored in the fridge.

    Any used sticks should go in the appropriate container to be sterilized for re-use. Any tubes, tips, and cotton swabs that you used can be thrown in the red biohazard bin. Please put any used ESKAPE pathogen tubes or plates in the labeled area and return all reagents and tools to the place that you got them.

    Using the 70% ethanol sprayers on your bench, spray down your bench space to disinfect it and wipe it clean with a paper towel.

    Protocol 3: Agarose Gel Electrophoresis

    You will need to run a small sample of each of the PCR reactions, including both controls, on a gel. Be sure to save most of the PCR from your favorite bacteria for sequencing. If you put it all on the gel, you won't have enough to sequence. Every gel must have at least 1 lane containing the DNA ladder. Plan which samples will go in each well of your gel. I suggest drawing a picture, so that you know what is where for certain! Your PCR product(s) from last week will have been stored in the freezer and will be returned to you.

    Materials
    • PCR Product(s)
    • 25 mL of 1% Agarose
    • GelRed DNA stain (10,000x GelRed)
    • 1x running buffer [Tris-borate-EDTA (TBE) or Tris-acetate_EDTA (TAE)]
    • 6x Loading dye
    • DNA ladder
    • Ice bucket with ice

    Protocol

    • Put a gel tray into the pouring tray with the shaded end where the comb will go.
    • Using gloves, measure 25 μL of GelRed stain and have it prepared.
    • Obtain a tube of 1% agarose from the hot water bath and pour it quickly into the gel tray. Each of these tubes contain 25 mL, so you don't need to measure them. NOTE: the agarose hardens very quickly, so work fast and be prepared before you remove the agarose from the hot water bath.
    • Add the 25 μL of Gel Red to your agarose in the tray, mixing it quickly with your pipette being sure to finish before it hardens, but don't make bubbles.
      • If there are any bubbles, use a clean pipet tip to pop or move them up onto the sides of the tray.
    • Allow the gel to cool at room temperature until it solidifies (~5 min).
      • Be careful not to bump the gel tray or table until it is completely solid or the gel will have waves that can affect how well the samples run.
      • The gel will appear opaque when it is solid.
    • Transfer 5 μL of each of your PCR products to a separate piece of parafilm. Label them so you know which is which.
      • Keep the original PCR product on ice until later. DO NOT USE IT ALL. If you make a mistake, ask for help!
    • Transfer 5 μL of the DNA Ladder to a different piece of parafilm. Label this one too.
    • Add 1 μL of 6x loading dye to each of your samples on the parafilm and mix well by carefully pipetting up and down.
    • After the gel is completely cool, remove the gel tray using gloves and insert the gel tray and gel in the gel box.
      • Be sure the end of the gel with the wells is connected to the negative electrode. This is VERY important as the DNA is negatively charged and will migrate to the positive electrode. If you run it backwards the DNA will run into the buffer and your DNA will be gone.
    • Pour in enough 1x running buffer so that it covers the surface of the gel and fills the chambers on either end of the gel box up to the faint line.
    • Carefully remove the comb by slowly lifting straight up.
    • Load 6 μL of the DNA ladder in the gel, as planned.
    • Load 6 μL of each of the controls.
    • Load 6 μL of your PCR sample with loading dye.
    • Close the lid of the gel box and plug in the gel box.
    • Turn on the power supply and set the voltage to 120 V.
    • Allow the gel to run until the loading dye is near the end of the gel, 20–40 min.
    • When the gel has run long enough, turn off the power supply and disassemble the gel box
      • The 1x running buffer should be collected to use again.
    • Carefully place gel into a carrying box.
    • With the assistance of your professor, photograph your gel under UV light.
    • Record the photograph of the gel in your lab notebook and be sure to label the contents of each lane as well as the size of all bands on the gel.

    Protocol 4: Purify Your PCR product

    While your gel is running, you will prepare your PCR sample for DNA sequencing. Although most of the DNA in your PCR reaction will be the product that you are interested in, there will also be some of the original template as well as leftover primers and dNTPs. All of these will get in the way of the DNA sequencing reaction, and need to be removed before sequencing. To do this, we will use an enzymatic method to specifically chew up and destroy single stranded DNA (primers) and individual nucleotides (dNTPs). These enzymes are active at 37°C. When it has completed its activity, we will denature it at 80°C so that it will no longer be active.

    Materials
    • PCR Project
    • 1.5 mL microfuge tube
    • Ice bucket with ice
    • ExoSap-IT reagent
    • PCR machine or heat blocks

    Protocol

    • Label your 1.5 mL tube.
    • Working with your tubes on ice, transfer 10 μL of your PCR product to a new 1.5 mL tube.
    • Add 4 μL of ExoSAP-IT reagent to your sample KEEP ExoSAP-IT ON ICE!.
    • Cap the PCR tube, vortex VERY gently, and then spin in a microfuge for a few seconds to collect all liquid at the bottom of the tube.
    • Return the ExoSAP-IT reagent to your instructor.
    • Transfer the tube to the heat block set at 37°C and leave them there for 15 minutes.
    • Transfer the tube to the other heat block set at 80°C for another 15 minutes.
    • Once cycling is complete, take the tube to your instructor to send for sequencing.

    If the heat blocks are not available, we will use the PCR machine instead. Follow these directions.

    • Label a 0.2 mL PCR tube.
    • Transfer your PCR + EXOSAP-IT to 0.2 mL tubes and gently mix.
    • Place these tubes in the PCR machine.
    • Select the appropriate program to start cycling (30 min). The cycle program is listed below in Table 2.
    • Once cycling is complete, take the tube to your instructor to send for sequencing.

     

     

    Table 2. The conditions for PCR clean-up.

    Temperature

    Time

    Purpose

    37°C

    15 minutes

    Enzymes hydrolyze primers and degrade dNTPs

    80°C

    15 minutes

    Enzymes are denatured

     

    Protocol 5: Submit Sample for Sequencing

    We will use a commercial company to sequence our 16S rDNA genes for us. To do this, we will need to package up each of the samples and the sequencing primer to prepare them all for sequencing.

    Materials

    Materials
    • ExoSAP-IT treated PCR Product
    • Bar codes
    • Sequencing spreadsheet

    Protocol

    • Once our PCR product has been purified, take it to your instructor.
    • Obtain one of the bar codes from your instructor and carefully place it around your tube.
    • Fill in the sequencing spreadsheet with the appropriate information associated with your bar code.
    • Give your sequencing tube to your instructor for submission.

    Protocol 6: Housekeeping and Organization

    It is critical as scientists that we carefully and accurately record and organize all of the products of our research. We will store your PCR samples until we get a good DNA sequence result back on your strain. We will also store your bacterial strain so that we can do additional research on it. Before leaving lab today, ensure that all of your materials are taken care of.

    Materials
    • ExoSAP-IT treated PCR Product
    • Original PCR Product
    • Master plates
    • Individual isolates
    • Sterile agar plates

    Protocol

    Clean up and organize all your material

    • Fully and carefully label your original (dirty) PCR product. Make sure that it is clear that this is the “dirty” one.
    • Give this PCR product tube to your instructor for safe keeping.
    • Look through all of the lab plates and identify all of the plates that belong to your group.
    • Using the most recent version of your master plate or individual isolate, restreak your antibiotic producer (or interesting strain) for individual colonies. Ensure that this plate is VERY well labeled.
      • Group Name
      • Date
      • Media Type
      • Incubation Temperature
      • Species name (genus and species)
      • List of ESKAPE relatives inhibited by this strain
    • Give your newly streaked plate to your instructor.
    • Dispose of all other plates in the biohazard bin.

    Protocol 7: DNA Sequence Analysis

    The portion of the genome that we submitted for DNA sequencing is the 16S rDNA gene which encodes one of the ribosomal RNAs. Since this particular gene is related in all living organisms, it makes this gene a particularly powerful way to identify, characterize, and classify an organism. Organisms that are more closely related share more DNA base pairs in this gene than organisms that are distantly related. We can therefore use this gene as a DNA fingerprint to identify your unknown bacterial isolate.

    However, each rDNA gene in bacteria is ~1500 base pairs long and it would be virtually impossible to read through the bases in your sequence and then compare them by hand to all other sequenced bacterial strains. Instead, we will use an online database (GenBank) and search algorithm (BLAST) to help to identify your microbe and its closely related species. BLAST or (Basic Local Alignment Search Tool) is a bioinformatics tool that allows us to navigate through huge databases and compare sequences with the collected library of published or submitted sequences (Altschul et al., 1990).

    Materials
    • 16S rDNA gene sequence file
    • Personal Computor (or lab computer)
    • BLAST online program at NCBI

    Protocol

    • Obtain your sequence files.
      • The sequence files for the whole class will be uploaded on Canvas
      • Locate the files that correspond to your unknown sample and the relevant controls. A description of the file names will be posted on the class website; if you have trouble determining which files are yours, contact your professor immediately.
      • Each sequence file is sent in different formats:
        • The .seq file is the sequence itself—the As, Ts, Gs, and Cs that were read by the sequencing machine. This is the file you will primarily work with.
        • The .ab1 file is the trace file. When the sequencing machine is gathering sequencing data, it appears as multiple peaks of different colors, which the machine reads and interprets as the As, Ts, Gs, and Cs that are in the .seq file. This is a great place to look quickly and see how your read "looks." Tall distinct peaks are a good thing, low or overlapping ones are not.
        • Some sequencing companies also send a .pdf file. The .pdf file has both pieces of information overlapped.
      • Locate your unknown sequence files and download them. Be sure to download all of the file types.
    • Copy and paste the sequence information from the .seq file into your lab notebook along with any additional relevant information supplied by the sequencing company.
    • Analyze your unknown sequence for the length and quality of the sequences.
      • Open the .seq file that you wish to examine.
      • First determine the length of your sequence
        • A good sequencing run should give you more than 500 base pairs of information.
        • An ok sequencing run would give you greater than 300 base pairs.
        • A bad/failed sequencing run will typically be less than 300 base pairs. This suggests that you did not have enough DNA in the sample or it was a mixture of two sequences. You will likely need to redo the experiment.
      • Second determine the quality of your sequences.
        • If the sequencer cannot determine the identity of a base, it will list an N for that position of the sequence.
        • The very beginning and end of a fragment of DNA cannot be decoded by sequencing, and so your sequence should start and end with a series of Ns. (Do not worry, this is normal and expected.)
          • Typically the first 60 or so bases will contain a large proportion of Ns.
          • After the first ~60 bases the sequence should consist of about 400–1000 bases of As, Gs, Ts, and Cs with ideally no Ns.
          • If the sequence you are examining contains more than 10 Ns between bases 60 and 500 the sequence is not considered to be of high quality.
          • A sequence with a few Ns may still be interpretable, but if there are too many (more than 30), you will likely have difficulty identifying the sequence.
    • Record the sequence quality in your notebook.
    • BLAST your unknown and control sequences.
      • Open the NCBI homepage (National Center for Biotechnology Information) homepage at www.ncbi.nlm.nih.gov.
        • NCBI is the primary site for molecular biology information. It is a part of the National Library of Medicine, as part of the National Institutes of Health. NCBI houses many important resources, including PubMed, the premier citation index for biomedical literature.
      • Click on BLAST under the list of popular resources on the right side of the NCBI landing page.
        • The Basic Local Alignment Search Tool (BLAST) is the most important way to search for DNA and protein sequences within a large collection of databases.
      • Select nucleotide BLAST under the Basic BLAST subheading.
      • Paste your sequence into the search box. At the top of the Standard Nucleotide BLAST page, you will see a section entitled Enter Query Sequence. In this section is a box labeled Enter accession number(s), gi(s), or FASTA sequence(s).
        • Copy and paste the sequence you wish to BLAST from the .seq file into this box
        • Select your entire sequence for BLAST analysis because BLAST will automatically remove the series of Ns at the beginning at and of your sequence (copy paste is your friend).
        • You can enter the sequence even if there are some Ns in this central sequence and the sequence is not of high quality.
      • Set search parameters. In the next section, entitled Choose Search Set, you will see a set of radio buttons and a pull down menu labeled Database.
    • On the pull down menu and/or selection circles, select 16S ribosomal DNA sequences (Bacteria and Archaea).
      • Leave all other options blank or set to their default value.
    • Run BLAST. Click on the blue button at the bottom of the page labeled BLAST.
      • You will be directed to a job processing page that will automatically update itself every few seconds until the search is complete.
      • Be patient; depending on network traffic this may take only a few seconds or several minutes.
    • Analyze your unknown BLAST results.
      • At the top of the BLAST results page, you will see a summary of the basic Query information followed by a Graphic Summary that contains a row of black, blue, green, purple, and red boxes with some red and/or other colored lines underneath.
      • We will not use the Graphic Summary for this analysis, so you can collapse this section by clicking on the (–) button to the left of the Graphic Summary section heading.
      • The Descriptions section gives a list of all the “hits” from your BLAST search. These are the species and strains for which there was significant sequence identity with your query sequence. The list is often quite lengthy!
    • Identify your Microbe's relatives! Scan down the Description column of this section, noticing the genus and species names for the top hits. Frequently the top several hits will be different species in the same genus.
      • This makes sense because two species that are closely related enough to be placed in the same genus likely also have very similar rDNA sequences to each other and therefore the same degree of homology with your query sequence.
      • Analyze the quality of your hits. The lower the E value, the stronger the match.
      • Notice also the values in the Max score column. The higher the Max score, the better the match.
      • After the Descriptions section is the Alignments section. This section shows an alignment of your query sequence with each of the hits listed in the descriptions section.
    • Analyze the first alignment. You will see two sequences aligned: your query sequence on top and the sequence of the best hit below it, labeled Sbjct. The alignment will likely be multiple rows long, so each row will have the Query on top and Sbjct below.
      • A vertical line between two bases means there is an identical match between the Query and Sbjct at that position. No line means there is a mismatch. A horizontal dash in the sequence means that there is an insertion/deletion of a nucleotide.
      • Notice how many mismatches there are in the alignment. A perfect hit will have none, but a very close hit may have a few.
      • Notice also what the mismatches are. If there is an N in your query sequence, then the mismatch may not “count” because an inability to determine that particular base in your sequence does not necessarily mean it doesn’t match with the Sbjct sequence.
      • BLAST usually automatically removes the many Ns at the beginning and end of the sequence, but there may be a few dispersed Ns at the end that weren’t removed automatically.
        • Optional: open the .ab1 file to examine the sequence trace file for your sequence. Find the N that shows a mismatch in the sequence alignment you are analyzing. Occasionally the automated sequencer is not able to determine the identity of a base but you can if you examine the trace by eye. You may need to use the y slider on the right side of the window to increase the amplitude of the peaks. Based on the color of the peak, can you determine with high confidence the identity of the peak? If there are two peaks overlapping, there may be contaminating DNA and no identification is possible. If you need assistance interpreting the sequence trace, consult your instructor.
          • If you can determine with high confidence the identity of the peak that the automated sequencer called an N, then you may edit your sequence to exchange that N for the correct nucleotide. Be sure you save the original and edited sequences as separate files.
    • Analyze the alignments the best several of the hits. Pay special attention to the 4–5 hits at the top of the list.
    • Be sure you record these results and observations in your notebook.
    • Make a figure and table for your notebook (and final project)
      • One figure of your notebook should be one or more alignments between your unknown bacterial sequence and the best BLAST hit(s).
      • Copy the alignment from the BLAST results page into Word.
      • Format the alignment appropriately to keep the spacing and information correct.
        • Use a monospace font (Courier) for this figure rather than a variable-width font such as Cambria or Times. If you do not, the spacing will be off and the vertical lines won't be in the right places.
        • Make the font type smaller (e.g. 8 or 10, so that a whole line will fit on one page)
      • At minimum you should have a figure with the alignment of the top BLAST hit. For your final project, you will evaluate the strength of the match based on how well the two sequences align.
      • If you wish to discuss multiple hits in your final project—for example if the top several BLAST hits are of similar strength be sure to include alignments for these too. They can be parts (a), (b), (c), etc. of the same figure.
      • When referring to specific mismatches, use the position number within the query sequence. For ease in determining this, the position numbers of the first and last base of each row are located next to the Query and Sbjct labels at the beginning and again at the end of the rows.

    Sequencing Analysis Instruction Video

    Discarding Material

    You are finished and have the sequence identity of your bacteria-congratulations!

    It is time to completely clean up all of your laboratory material. If your sequencing is successful, you can discard all of your DNA and PCR samples in the biohazard trash. You can also discard your plates in the biohazard plate bin. The only plate that you need to save is the final streak plate. Please be sure that this final streak plate is labeled well and give it to your instructor.

    Any used sticks should go in the appropriate container to be sterilized for re-use. Any tubes, tips, and cotton swabs that you used can be thrown in the red biohazard bin. Please put any used ESKAPE pathogen tubes or plates in the labeled area and return all reagents and tools to the place that you go them.

    Using the 70% ethanol sprayers on your bench, spray down your bench, spray down your bench space to disinfect it and wipe it clean with a paper towel.


    This page titled 10: Research Project Week 5-7- Molecular Sequence Identification is shared under a not declared license and was authored, remixed, and/or curated by Nora Sullivan.

    • Was this article helpful?