sci_bio

Opening the Black Box: What a Gene Actually Does

Chapter summary, hard words and model exam answers.

Free online summary and notes. Read it here, no PDF download needed.

About the author

Science · CBSE Class 12 · NCERT Biology, Ch.5 (sections 5.5-5.10)

Summary

Every chapter in this thread has used the word gene freely without ever quite opening it up. A gene gets inherited, one copy from each parent. A gene can be dominant or recessive. A gene can mutate. A gene sits at a specific spot on a specific chromosome. All of that is true, and all of it treats a gene as a sealed box: something that clearly does something, without ever showing exactly what happens inside. This chapter finally opens that box. Cellular DNA is the information source a cell uses to build proteins, and it turns out to be a genuine multi-step manufacturing process, DNA copied into a working script, that script read in a strict code, and that code translated into an actual physical chain of amino acids, the protein itself. Follow that process all the way through, and something else falls out almost as a bonus: the very same molecular machinery that builds a protein also explains how forensic science can match a single strand of hair to one specific person out of billions.

The process of copying genetic information from one strand of DNA into RNA is called transcription, and it runs on the same complementary base-pairing logic as DNA replication, with one small chemical substitution: adenine now pairs with uracil instead of thymine, since RNA uses uracil in thymine's place. Transcription differs from replication in two important ways, though. Replication duplicates an organism's entire DNA; transcription copies only one specific segment. And replication copies both strands of the double helix; transcription copies only one. That second restriction is not arbitrary. If both strands acted as templates, they would produce two different RNA molecules with two different sequences, since complementary strands are not identical strands, and a single stretch of DNA would end up coding for two entirely different proteins at once, creating chaos rather than useful information. Worse, two simultaneously produced, mutually complementary RNA strands would simply pair up with each other into a double-stranded RNA, permanently unable to be translated into anything at all. A transcription unit, the actual stretch of DNA being copied, is defined by three parts: a promoter, where the enzyme RNA polymerase physically binds to begin the process, the structural gene itself, and a terminator, marking where copying stops.

A gene, especially in more complex organisms, rarely turns out to be one clean, uninterrupted stretch of coding sequence. The initial RNA copy typically contains alternating segments: exons, the parts that will actually appear in the finished, usable RNA, and introns, intervening sequences that get precisely cut out and discarded before that RNA ever leaves the nucleus, a process called splicing. Only after splicing removes the introns and stitches the remaining exons together in the correct order does the RNA become a mature, functional messenger ready to direct protein construction. This split-gene arrangement complicates the very idea of what a gene physically is, and it is not simply an inefficiency, most biologists now treat the presence of introns as a genuinely ancient feature of how genomes are built, one that current understanding is still working out the full significance of.

DNA and RNA are built from just four different bases, while proteins are built from twenty different amino acids. Somehow, a four-letter alphabet has to specify a twenty-item vocabulary, and the person who first worked out the arithmetic of how was, strikingly, not a biologist at all. George Gamow, a physicist, reasoned that a code using single bases could specify only four things, and a code using pairs of bases only sixteen, neither large enough to cover twenty amino acids. A code using three bases at a time, though, a bold proposal at the time, would generate sixty-four possible combinations, comfortably more than enough. Proposing this was one thing; actually proving it and working out which specific triplet corresponded to which specific amino acid was an entirely different, genuinely interdisciplinary challenge. Har Gobind Khorana developed chemical methods to synthesise RNA molecules with precisely defined, known sequences. Severo Ochoa contributed an enzyme capable of building RNA chains without needing a template at all, letting researchers manufacture RNA of exactly the composition they wanted to test. Marshall Nirenberg then built a cell-free system, actual protein-making machinery extracted from cells and kept functioning outside them, and fed it these synthetic RNAs one at a time, watching to see which amino acids got incorporated into the resulting protein. Piece by piece, this combined effort cracked what is now called the genetic code, a complete checkerboard mapping every one of the sixty-four possible three-base combinations to its corresponding amino acid, or, in three cases, to no amino acid at all.

The finished code turned out to follow several distinct rules, each with real consequences. It is degenerate, meaning most of the twenty amino acids are specified by more than one codon rather than just one, sixty-one codons in total covering only twenty amino acids between them. It is read continuously, in one direction, three bases at a time, with no gaps, punctuation or overlap between successive codons. It is, remarkably, nearly universal: the same codon specifies the same amino acid whether you are reading it in a bacterium or a human being, a shared vocabulary spanning nearly all of life, with only a handful of known exceptions, mostly inside mitochondria and in some single-celled protozoans. One codon, AUG, does double duty, both coding for the amino acid methionine and serving as the universal signal to start translation at all. And three codons, UAA, UAG and UGA, code for no amino acid whatsoever, functioning instead as stop signals marking exactly where a protein's construction should end.

Because the code is read three bases at a time with no punctuation to mark where one codon ends and the next begins, the position where reading starts matters enormously. Imagine a message written entirely in three-letter words with no spaces, something like RAMHASREDCAP, correctly read as RAM HAS RED CAP. Insert a single extra letter partway through and try reading it again in unbroken groups of three, and every word from that insertion point onward turns to nonsense, even though nothing about the letters themselves changed, only where each three-letter group now begins. Insert or delete three letters, or any multiple of three, instead, and the reading frame recovers immediately afterward, shifted by one word but otherwise intact. This is exactly what happens inside a real gene. Insertions or deletions of one or two bases are called frameshift mutations, and they corrupt every single codon downstream of the mutation site, typically destroying the resulting protein's function entirely. Insertions or deletions in multiples of three, by contrast, simply add or remove one or more whole amino acids from the protein without shifting anything that comes after, a far more limited kind of damage.

Even a perfectly cracked code raises an immediate practical problem: amino acids have no structural feature that lets them directly recognise a specific sequence of RNA bases, nothing about the physical shape of the amino acid methionine inherently connects it to the codon AUG. Francis Crick recognised this gap early and predicted there had to be some kind of adapter molecule, something able to read a codon on one end while carrying the correct, matching amino acid on the other. That adapter turned out to be transfer RNA, tRNA, a molecule that folds into a compact, characteristic shape, sometimes drawn as a simple clover-leaf though its real folded structure looks more like an inverted L. One end of a tRNA molecule carries a three-base sequence called an anticodon, complementary to one specific codon, while the opposite end is chemically bonded to that codon's matching amino acid. Every amino acid used in protein synthesis has at least one dedicated tRNA molecule built to carry it, and translation cannot begin at all without a specialised initiator tRNA specifically matched to the AUG start codon.

Translation itself is the process of stringing amino acids together into a growing protein chain, joined by peptide bonds, a reaction that requires energy, drawn from ATP, to first activate each amino acid and bind it to its matching tRNA. The physical site where all this happens is the ribosome, a two-part structure built from structural RNA and roughly eighty different proteins, existing as separate large and small subunits until an mRNA molecule brings them together to begin. The ribosome positions two charged tRNAs close enough for their amino acids to link, then shifts along the mRNA one codon at a time, releasing the tRNA that has delivered its cargo and welcoming the next. Translation begins precisely at the AUG start codon, recognised only by the initiator tRNA, and continues, codon by codon, amino acid by amino acid, until the ribosome reaches one of the three stop codons, at which point a release factor binds instead of another tRNA, the finished polypeptide chain detaches, and the ribosome's two subunits separate again, ready to begin the entire process over on a fresh mRNA.

Building a protein costs real energy, so it makes little sense for a cell to keep manufacturing an enzyme it currently has no use for. Consider Escherichia coli and its enzyme beta-galactosidase, which breaks the sugar lactose down into simpler sugars the bacterium can use for energy. If there is no lactose around to break down, that enzyme is simply wasted effort, and E. coli has evolved a genuinely elegant switch to avoid making it needlessly, discovered by Francois Jacob and Jacques Monod and known as the lac operon. Three structural genes, working together, handle lactose metabolism, but they sit under the control of a fourth, regulatory gene that constantly produces a repressor protein. That repressor normally clamps onto a specific site right next to the operon's promoter, physically blocking the enzyme RNA polymerase from transcribing the three structural genes at all. When lactose is actually present, it acts as an inducer, binding directly to the repressor and disabling it, releasing its grip on the promoter and finally allowing transcription, and enzyme production, to proceed. The moment lactose runs out, the repressor reasserts control and production shuts back down. It is, in effect, the enzyme's own substrate regulating whether the enzyme itself gets made at all, an efficient, self-correcting feedback loop that only produces exactly what the moment actually requires.

Everything covered so far in this chapter concerns a single gene at a time. The Human Genome Project, launched in 1990 and completed in 2003, asked a far larger question: what does the complete sequence of all human DNA actually look like, all roughly three billion base pairs of it? The scale involved is genuinely difficult to grasp. At the estimated starting cost of about three US dollars per base pair, the full project was projected to cost roughly nine billion dollars, and if the finished sequence were printed as plain text, one thousand letters to a page, one thousand pages to a book, it would fill some 3,300 separate books just to record the DNA from a single human cell. An international collaboration, coordinated by the United States Department of Energy and National Institutes of Health with major partners including the Wellcome Trust in the United Kingdom, pursued goals that went well beyond simply reading the sequence: identifying every human gene, storing and analysing the resulting data, and grappling directly with the ethical, legal and social questions the project's own findings were bound to raise. What the completed project actually found upended a great many prior assumptions. Rather than the eighty thousand to one hundred forty thousand genes many researchers had expected, the human genome turned out to contain only around thirty thousand. Fewer than two percent of all that DNA actually codes for protein at all. Over half of the genes identified still have no known function even today. And, perhaps most strikingly, any two human beings, chosen entirely at random from anywhere on Earth, share ninety-nine point nine percent identical DNA sequence.

If ninety-nine point nine percent of human DNA is identical from person to person, the genuinely useful information for telling two individuals apart has to live somewhere in that remaining fraction, and it turns out to concentrate specifically inside the vast stretches of repetitive, non-coding DNA the Human Genome Project found padding out most of the genome, the very regions long dismissed as functionally unimportant filler. Certain repetitive sequences, called satellite DNA, vary enormously in exactly how many times a short unit repeats, one person might carry twelve copies of a particular repeat at a given spot, another person forty, and that repeat count turns out to be almost as individually distinctive as an actual fingerprint. Alec Jeffreys developed the technique that exploits this, DNA fingerprinting, built around a specific category of highly variable repeat called Variable Number of Tandem Repeats, or VNTRs. Because the exact repeat count at several such locations is essentially unique to each individual, except identical twins, and because it is faithfully inherited from parents to children, comparing these patterns lets forensic scientists match a biological sample, blood, hair, skin, saliva, to one specific person with remarkable confidence, and lets the same basic comparison settle paternity disputes by checking whether a child's pattern is consistent with a claimed parent's own. The same repetitive DNA that helped make the Human Genome Project's completed map so much larger and messier than anyone originally expected turned out to be exactly what individual identification needed all along.

This chapter closes not just this one story but the entire Genetics and Heredity thread, and it is worth tracing the full line all the way back. Class 10 opened with a simple observation, a field of sugarcane clones next to a litter of visibly varied puppies, and asked why sexual reproduction produces so much more variety than asexual copying does. The answer led through independent assortment and the physical fact that genes sit on separate chromosomes. Class 12's earlier chapter pushed further still, into the many ways real inheritance complicates Mendel's clean rules, blending alleles, co-dominant alleles, linked genes, whole different chromosomal systems for deciding sex, and the disorders that result when any of this goes wrong. This chapter has gone one layer deeper than either, opening the gene itself to show the actual manufacturing process, DNA to RNA to protein, running underneath every single pattern the earlier chapters described, and then following that same molecular logic outward into one of its most practical modern applications, identifying one specific human being out of billions from nothing more than a fragment of their own DNA.

Hard words & meanings

transcriptionthe process of copying genetic information from one strand of DNA into RNA
exona coding segment of a gene that remains in the mature, processed RNA
intronan intervening sequence in a gene that gets cut out before the RNA is used
splicingthe process of removing introns and joining exons together in mature RNA
codona sequence of three RNA bases that specifies one amino acid or a start/stop signal
genetic codethe complete set of rules mapping each three-base codon to a specific amino acid
tRNAtransfer RNA, the adapter molecule that reads a codon and carries the matching amino acid
translationthe process of building a protein by reading mRNA codons at the ribosome
ribosomethe cellular structure, built from RNA and protein, where translation takes place
operona group of genes controlled together by a shared promoter and regulatory genes
repressora protein that blocks transcription by binding to an operator region
inducera molecule that switches on gene expression by disabling a repressor
VNTRVariable Number of Tandem Repeats, a highly variable repetitive DNA sequence used in DNA fingerprinting
🔒

Model exam answers, grammar & audio

You have read the summary. The board-ready model answers, grammar notes, one-touch audio and writing practice for this chapter are part of Lipi©.

Unlock free with any language course

See it, understand it, hear it read aloud, then write the exam answer with confidence, for a fraction of a tutor cost.