What a real sequence looks like
Five sequences worth knowing — and one that is not a gene
The box above opens on eighteen bases that somebody invented. These five are real: four reading frames taken from the sequence databases and one famous stretch of DNA that is not one. Click any of them to load it into the tool.
| Sequence | Bases | Length | GC | Translates to |
|---|---|---|---|---|
| Human β-globin, first eight codons | ATGGTGCATCTGACTCCTGAGGAG |
24 bp | 54.2% | Met-Val-His-Leu-Thr-Pro-Glu-Glu |
| The same, with the sickle-cell change | ATGGTGCATCTGACTCCTGTGGAG |
24 bp | 54.2% | Met-Val-His-Leu-Thr-Pro-Val-Glu |
| Green fluorescent protein, start | ATGAGTAAAGGAGAAGAACTTTTCACTGGA |
30 bp | 36.7% | Met-Ser-Lys-Gly-Glu-Glu-Leu-Phe-Thr-Gly |
| Preproinsulin signal peptide, start | ATGGCCCTGTGGATGCGCCTCCTGCCCCTGCTG |
33 bp | 69.7% | Met-Ala-Leu-Trp-Met-Arg-Leu-Leu-Pro-Leu-Leu |
| The Kozak context — not a reading frame | GCCACCATGG |
10 bp | 70.0% | Ala-Thr-Met |
Read the first two rows together and nothing else on this page matters as much. They differ at one base out of twenty-four — the twentieth — and the translation column turns that into Glu becoming Val at the seventh position. That single substitution is sickle-cell disease. Everything else here is scale and variety: GC runs from 36.7% in a jellyfish protein to 69.7% in the signal peptide of human insulin, so a sequence being GC-rich says more about which organism and which part of the gene than about what it does. And the last row is a warning about this tool rather than an example for it. The Kozak context is the signal that marks where translation should START; it is not itself a reading frame, and the tool cheerfully translates it to Ala-Thr-Met anyway. A calculator cannot tell you whether your question makes sense.
Problem solved in full
-
A three-letter codon for eighteen bases and 33.3% GC 5 steps
Eighteen bases, six codons, 33.3% GC. Work out why a three-letter codon is the shortest one that could possibly work, and what the leftover capacity buys.
-
The sequence divides into codons of three, so eighteen bases give six — and the last is TAA, a stop, which is why this reads as a complete short gene.
-
GC content counts the G and C bases, six of eighteen. It matters because G–C pairs have three hydrogen bonds to A–T's two, so a GC-rich sequence melts at a higher temperature.
-
Now the size question. Four bases taken two at a time give 16 combinations — fewer than the 20 amino acids that must be encoded, so pairs cannot work.
-
Three at a time gives 64, against 21 things to name counting the stop signal. Three is therefore the shortest workable word length, and it was deduced before it was observed.
-
That leaves 64 mapping onto 21, roughly three codons per meaning. The surplus is spent on redundancy, not on ambiguity: several codons name the same amino acid, but none names two.
Answer
The tool prints 18 bp, 6 codons and 33.3% GC. The redundancy is the part with consequences. Because synonyms usually differ in the third base, a mutation there often changes nothing at all — the code has a built-in error tolerance that is a property of the mapping, not of any repair machinery. The reading frame has the opposite property: delete one base and every codon after it is regrouped, turning the rest of the gene into nonsense. Same three-letter structure, and it makes one kind of error nearly free and the other catastrophic.
-
References (6)
- The code the tool tabulates: F. H. C. Crick, L. Barnett, S. Brenner and R. J. Watts-Tobin, "General Nature of the Genetic Code for Proteins." Nature 192, 1227–1232, 1961.
- The error-minimising structure, and how strong that claim actually is: S. J. Freeland and L. D. Hurst, "The Genetic Code Is One in a Million." Journal of Molecular Evolution 47(3), 238–248, 1998.
- The sequence library below: where its bases were fetched from, accession by accession: NCBI Reference Sequence Database — HBB (NM_000518.5) and INS (NM_000207) coding sequences, and GFP (GenBank M62653.1). Retrieved 2026-08-06.
- And the result the first two rows of that table exist to show: V. M. Ingram, "Gene mutations in human haemoglobin: the chemical difference between normal and sickle cell haemoglobin." Nature 180, 326–328, 1957 — one amino acid.
- The jellyfish protein that made cell biology visible: D. C. Prasher, V. K. Eckenrode, W. W. Ward, F. G. Prendergast & M. J. Cormier, "Primary structure of the Aequorea victoria green-fluorescent protein." Gene 111(2), 229–233, 1992.
- And the context that tells a ribosome where to start, which is the last row and not a gene: M. Kozak, "An analysis of 5′-noncoding sequences from 699 vertebrate messenger RNAs." Nucleic Acids Research 15(20), 8125–8148, 1987.