East River Notes
genomics

Sequencing 101

How genetic sequencing works and its impact on modern biotech.

What sequencing is and why it matters

  • DNA is a four-letter code (A, C, G, T), and the order determines the information.
  • Deciphering the sequence of DNA is critical to understanding biology.
  • Reading a human genome went from ~$100 million to a few hundred dollars over roughly two decades.
  • Bending the cost curve faster than Moore’s law has helped unleash new discoveries and technologies in biology and medicine.
  • Sequencing has become the backbone of biopharma research, cancer treatment, rare-disease diagnosis, prenatal screening, and disease outbreak tracking.

The simple version

Reading DNA became easier, faster, and cheaper – which ushered in major developments in researching, diagnosing, and tracking diseases.

What does it mean to read DNA?

DNA is the four-letter code every living cell runs on, and reading it means knowing the sequence of those letters.

DNA (deoxyribonucleic acid) stores the instructions for building and running an organism. Its shape is a double helix – two strands wound around each other, joined by rungs.

Each rung is a pair of chemical bases, and the pairing follows a fixed rule – adenine with thymine, cytosine with guanine (A–T, C–G).

Exhibit 1DNA structure and nucleotide bases
DNA has two complementary strandssupported by sugar-phosphate backbone Adenine&Thymine Cytosine&Guanine A always pairs with T, and C with G.

Source: East River Notes, from standard scientific literature. Schematic; the bases shown are illustrative.

The information is the order of those bases along a strand – the sequence (like binary code, but with four variables). A gene is a stretch of that sequence that codes for a product, usually a protein. The complete set is the genome – about three billion base pairs in humans. Sequencing is working out the order of the letters.

Why did reading a genome once cost millions?

The Human Genome Project ran from 1990 to 2003; producing that first sequence cost an estimated $500 million to $1 billion, and the program as a whole – which also funded mapping, model organisms, technology development, among other expenses – came to about $2.7 billion. Even after the first sequenced genome, reading a single genome still cost millions of dollars for years.

The reason sits in the method.

Sanger sequencing works by creating copies of the DNA strand at every possible length. These fragments are pushed through a gel or a narrow capillary as an electric current flows through it. The smallest fragments move fastest, so they arrive in order, shortest to longest. By adding dyes to the fragments and marking how far each traveled, you can sequence the DNA.

Although Sanger sequencing is scientifically sound and repeatable, reading three billion letters this way took a lot of time and money.

What made sequencing cheap?

The breakthrough was reading enormous numbers of fragments at the same time. The DNA is cut into short fragments, and those fragments are spread across a small glass chip whose surface has a lot of tiny wells – tens of billions on the highest-output chips.

A sequencing machine observes how the DNA is synthesized in each well. It reads billions of them at once. Reading enormous numbers of short fragments in parallel like this is called next-generation sequencing (NGS). This is what solved the bottleneck.

How does next-generation sequencing (NGS) actually work?

The predominant sequencing technology used today is called sequencing-by-synthesis (SBS). With this method, DNA is sequenced by observing as it replicates itself.

An enzyme rebuilds a DNA strand one letter at a time; each letter carries a colored tag; a camera records the color as each one locks into place.

Takeaway

Modern sequencing reads DNA by watching it get copied – an enzyme rebuilds one base pair at a time, and a camera records that sequence.

Two preps happen before the machine starts. Sample prep – getting clean DNA out of the sample (blood, saliva, tissue) and washing away everything else. Library prep – cutting DNA into short pieces and tagging the ends with adapters, which are handles that let each piece grab onto the chip surface.

Below are the highly simplified steps of the NGS workflow:

Step 1 – Get the DNA ready. Purify the DNA, cut it into short fragments a few hundred letters long, and tag both ends with adapters and a barcode.

Exhibit 2The DNA is cut into short pieces and tagged.
1 · Genomic DNA, extracted & purified cut into fragments 2 · Short fragments (~300 bases) ligate adapters + a barcode index to both ends 3 · Adapter-tagged double-stranded library denature: separate the two strands 4 · Denatured to single strands (ready to load onto the flow cell)

Source: East River Notes, from company presentations, filings, and other publicly available information.

Step 2 – Load the flow cell. The flow cell is a small glass chip whose surface is covered with tiny wells. The tagged fragments wash across the chip, and each settles into a well.

Exhibit 3Each piece lands in its own well on a patterned chip.
illustrative flow cell inside a lane: an ordered grid of wells
The flow cell is a glass chip patterned with multiple lanes. Zoom into a single lane and the surface is an ordered hexagonal grid of tiny wells.

Source: East River Notes, from company presentations, filings, and other publicly available information. Schematic; not to scale.

Step 3 – Copy each fragment into a cluster. In order to amplify the signals that the sequencer can detect, each fragment is copied over and over again, making a tight bundle of identical copies – a cluster.

Step 4 – Watch the copy being built, one letter at a time. Once the clusters have been generated, sequencing can begin. Each fragment is a single strand of DNA. Sequencing works by observing as each strand builds its complementary copy.

To start the process, two key ingredients are introduced: DNA polymerase (an enzyme that adds each new letter to the growing strand) and free-floating nucleotides (each carrying one of four fluorescent colors and a chemical cap).

To explain simply:

  1. nucleotides are added (each has a color tag and a chemical tag attached)
  2. DNA polymerase adds a nucleotide (A&T and C&G are the options)
  3. the chemical cap “pauses” the process after that one nucleotide is paired
  4. a laser shines across the flow cell; only the most recently paired nucleotide shines its fluorescent marker (each well shows a color, and that tells which of the four nucleotides was added for each well)
  5. another chemical is added, which snips off the color tag and the cap
  6. repeat
Exhibit 4Simplified view of SBS in action
new strand T A A T C G G C A G T A Polymerasethe copying enzyme T A C G T C G A T C CGT
Only the matching letter is added, and the cap stops the strand there. The whole flow cell (billions of wells) is photographed at once, then the chip is rinsed and the process repeats.
Note: the illustration shows one color per letter for simplicity; the latest generation of sequencers use two dyes and two images, with one base read as no color.

Source: East River Notes, from company presentations, filings, and other publicly available information.

The chemical cap that ensures only one nucleotide is added each time is critical. It allows all the wells to “pause” once each of them has added one nucleotide, so that a single photograph across the flow cell captures what each well synthesized at that given time. Without it, the enzyme would run the length of the strand and the picture would be a blur.

Once you collect the order of the colors for each well, it spells out the sequence.

Step 5 – Turn the reads into answers. One read is a string of about 150 letters, but that is only a small snippet of the gene from an unknown spot. On its own it means nothing. So software lines up every read against a reference genome – a standard, already-known human sequence. Once there are enough reads stacked on top of each other, using the reference genome like the cover picture on a jigsaw puzzle box, you can read off the sample’s sequence at every position.

Exhibit 5Stacking reads against reference genome
reference genome GATCCAGTTGCAAGGTCATCGGATTCAGGCATTGACCTGA TCCAGTTGCAAGGTCATCGG GGATTCAGGCATTGACC TTGCAAGGTCATCGGA AGGTCATCGGATTCAGGC CAGTTGCAAGGTCATCGGA ATCGGATTCAGGCATTGA CAAGGTCATCGGATT 7x Read depth = how many reads stacked on top of each other; the boxed base is covered 7x. A typical read depth ranges between 20x–100x.

Source: East River Notes, from standard scientific literature.

Sometimes, the piled-up reads all disagree with the reference genome in the same place. This is called a variant – and that is the key finding. It means that the sample’s DNA has a different sequence than a typical human genome. If all of the reads disagree, that means both chromosomes carry it (homozygous; came from both parents). If roughly half of the reads disagree, that means only one chromosome carries it (heterozygous; came from one of the parents).

Exhibit 6Identifying a variant
reference T G C A A G G T C A T C G G A T T C A G G C A A G G T C A T C A A A G G T C A T C A G A G G T C A T C A G A G G T C A T C A G A T G T C A T C A G A T T T C A T C A G A T T C C A T C A G A T T C A A T C A G A T T C A G T C A G A T T C A G G Reads agree with each other but not the reference – that position is a variant.

Source: East River Notes, from standard scientific literature.

How far and how fast did the cost fall?

The cost of sequencing fell faster than the cost of computing over the past ~20 years.

Exhibit 7Cost per human genome over time
$10 $100 $1K $10K $100K $1M $10M $100M 2001 2005 2010 2015 2020 2026 Moore’s law

Source: East River Notes, from NHGRI, DNA Sequencing Costs: Data (genome.gov), and other publicly available information. The 2022–2026 segment reflects current NovaSeq X cost estimates; NHGRI’s series ends May 2022.

What did cheaper sequencing unlock?

As sequencing became more accessible, the technology became the bedrock of research and clinical workflow.

Cheaper, faster sequencing changed what research and medicine could attempt. Work that had once been unaffordable became routine for scientists and clinicians.

Academic and biopharma researchers could run studies that were previously not possible and get precise data. Sequencing enabled the science community to concretely understand the role of certain genes and mutations. Sequencing technology underpins modern breakthroughs like gene therapy and CRISPR.

What started as a research tool for explaining biology became cheap and reliable enough to enter the clinic – to diagnose a cancer, guide its treatment, screen a pregnancy, track an outbreak.

To give a few examples, a liquid biopsy measures the amount of tumor DNA that is circulating in the bloodstream. This is notable because a) a blood draw is much easier and far less invasive, b) the test is highly repeatable and can track over time, and c) a cancer could be caught months earlier, instead of waiting until the tumor gets big enough to notice or to show up on scans.

Another example is prenatal screening. Non-Invasive Prenatal Testing (NIPT) screens for common chromosomal anomalies (e.g. Down Syndrome) and dozens of genetic disorders without requiring invasive procedures. Even before pregnancy, some may choose to do carrier screening, which determines if you or your partner carry genes for potential genetic disorders that could be passed down.

Why it matters

Sequencing DNA is now much more affordable and used broadly. But the lesson is not simply that sequencing is cheap. It is that sequencing unlocked and became the bedrock for much of modern biotechnology. You cannot study a gene, diagnose a mutation, target a therapy, or edit a sequence you have not first read.

Sequencing unlocked and became the bedrock for much of modern biotechnology.

Exhibit 8What sequencing enables
DNA sequencing Research explaining biology how genes work population studies how a tumor behaves Diagnosis finding the cause disease tracking cancer surveillance prenatal screening Treatment guiding what to do matching drug to tumor measuring effectiveness monitoring resistance Engineering building new therapies cell & gene therapy mRNA, siRNA, ASO clinical trials
A representative selection, not an exhaustive list of what sequencing enables.

Source: East River Notes, from company presentations, filings, and other publicly available information.

Takeaways

  • Reading DNA is step zero. DNA is the four-letter code of life, and almost everything in modern biotech – research, diagnosis, treatment – begins with being able to read it.
  • The barrier was throughput, and parallelism broke it. Identifying one letter was never the hard part. Reading billions of fragments at once is what took a genome from about $100 million to a few hundred dollars.
  • The cost fell faster than computing. At every order of magnitude, sequencing crossed from one domain into the next: landmark project, research tool, clinical test, routine blood-based test.
  • Sequencing became the foundation. Affordable, widely available sequencing enabled much of modern biotechnology, and is now woven through it.

One line to remember

Widely available DNA sequencing enabled new discoveries and underpins much of modern biotechnology.


General, educational, and informational research only, not tailored to your situation. Nothing here constitutes investment, legal, medical, or other professional advice; an offer to sell or a solicitation of an offer to buy any security; promotional or marketing material; or a recommendation. The author may hold positions in the securities or sectors discussed. Do your own research and consult a licensed professional. Full disclosures at www.eastrivernotes.com/disclosures.


Notes

The content and history described here are synthesized from company presentations, filings, and other publicly available information, together with standard scientific literature and the NHGRI's published cost data. Quantitative figures are rounded and estimated. Exhibits and texts are illustrative and are designed to simply explain complex topics. Some technical terms are deliberately simplified for a general reader without changing their underlying meaning. This primer favors durable concepts, over point-in-time statistics.

East River Notes