← Back to Book Detail

Chapter 6: Transcriptomics (6/6) -- Applied Bioinformatics

Browse
100%

Chapter 6: Transcriptomics

Chapter 6: Transcriptomics Broadly speaking, Transcriptomics is the study of transcriptomes, the sum total of all transcripts in a cell. Transcriptomics seeks to build transcriptome annotations, and to measure differential expression of transcripts from different tissue types or treatments. 6.1 High-throughout Sequencing (HTS) High-throughput sequencing (also known as deep sequencing) is a technology that has been developed in the late 20th century and continues to improve today. High-thoughtput sequencing has many applications, and most relevant for transcriptomics is deep sequencing of RNA, called RNA-seq. The word “deep” in deep sequencing refers to the depth of sequencing, characterized by: [latex]D = \frac{N \times L}{T}[/latex] where the depth [latex]D[/latex] is computed from the number of reads [latex]N[/latex], the length of the reads [latex]L[/latex], and the size of the transcriptome [latex]T[/latex]. The size of the transcriptome [latex]T[/latex] can be thought of the length of the union of all transcripts for a particular system. A depth of [latex]2\times[/latex] means that on average a location in the genome would have [latex]2[/latex] reads mapping to that location, assuming a uniform distribution of reads. This equation assumes that the reads are uniformly distributed, which is almost never true. Nevertheless it serves as a good approximation. High-throughput sequencing can produce hundreds of millions of reads per sequencing lane, and in many cases the lane is multiplexed to include multiple samples per lane. This technology has enabled scientists to study biological phenomena at a genome-wide scale, and has enabled the discovery of a number of properties of transcription. 6.2 RNA Deep Sequencing RNA deep sequencing is a method where a cDNA library is created for an RNA sample, and is sequenced using high-throughput sequencing, producing hundreds of millions of reads. Notably, there are different types of RNA-seq data sets. First, single-end reads involve the sequencing of one read per cDNA fragment, typically in the 5′ to 3′ direction. Paired-end reads have two reads per fragment, with the two paired-reads called “mates”. Often the first mate is sequenced in the direction of transcription, and the second mate is sequenced in the opposite 3′ to 5′ direction. This, however, can vary on the sequencing technology used. The manual for tophat 2 (https://ccb.jhu.edu/software/tophat/manual.shtml) provides the information on Figure 6.4. 6.2.1 Single-end Sequencing Single-end Sequencing produces one read per fragment, so it can be good for transcript quantification, but may not resolve differences in expression across splice variants or different isoforms of the same gene. Therefore, it can be good for quantifying small RNA expression, or expression at the gene-level when splice variants are not a concern. 6.2.2 Paired-end Sequencing Paired-end Sequencing produces two reads per fragment, and where typically a fragment size distribution is
← Previous Chapter Next Chapter →